frontline learning research 5 special issue „learning through networks‟ (2014): 1 3 issn 2295-3159 corresponding author: anoush margaryan, anoush.margaryan@gcu.ac.uk doi: http://dx.doi.org/10.14786/flr.v2i2.123 1 | f l r introduction to the special issue ‘learning through networks’ anoush margaryan a a caledonian academy, glasgow caledonian university article received 26 june 2014 / revised 9 july 2014 / accepted 11 july 2014 / available online 15 july 2014 this special issue examines the role of networks in professionals‟ learning. by networks, we specifically mean personal professional networks, which may or may not be mediated by digital technology such as social media. the special issue is based on a symposium „learning through networks‟ held at the 2013 conference of the european association for research in learning and instruction (earli) in munich, germany. a special issue capturing current empirical and conceptual research in this area is timely. the importance of the social dimension of learning, in particular of learning from the experience of others, is firmly established in the literature. there are many different types of social formations that have been studied within the learning sciences– groups, teams, communities, collectives and increasingly also networks (dron and anderson, 2007; mccormick, fox, carmichael and procter, 2011). recently, the concept of „networked expertise‟ (hakkarainen, palonen, paavola, & lehtinen, 2008) has been put forward to characterise learning and development in professional contexts. yet, despite the growing recognition of the importance of networks in learning, it is not well understood what precisely is learned through networks, how it is learned, and what environmental factors organisational, social, structural or technological enable or constrain learning through networks (littlejohn and margaryan, 2013). identifying and analysing the mechanisms and factors of learning through networks are vital to our understanding of the contemporary professional learning. comprising contributions from researchers in four european countries (germany, netherlands, finland and the uk), this special issue brings together examples of emergent empirical and conceptual research in learning through networks and proposes recommendations to stimulate future research and development in this area. contextualized in different settings and using complementary approaches, the contributions collectively address the following overarching questions: 1. what is learned through networks and how it is learned? 2. what is the potential of applying a social network perspective to understanding the nature of learning through networks? a. margaryan 2 | f l r 3. what key factors – individual, structural, organisational, technological – impact professional's learning through networks and how they impact it? there are five contributions in this special issue. three of these contributions are reports of new empirical data. these papers draw on social network analysis (sna) and/or semi-structured interviews to examine the structure of networks of professionals and to identify what professionals learn through these networks. pataraia, falconer, margaryan, littlejohn and fincher examine personal networks of academics teaching in universities, analyzing the types of interactions that academics engage in and the implications of these interactions for their professional learning and improvement of their teaching practice. hytonen, palonen and hakkarainen on one hand, and rehm, gijselaers and segers on the other hand examine the impact of individuals‟ hierarchical positions within networks upon their opportunities for learning and knowledge sharing. these three empirical contributions are supplemented by a conceptual review paper by vaessen, van den beemt and de laat which draws on a synthesis of workplace learning, hrd, organisational and management science and learning science literatures to analyse key organisational factors impacting upon learning through networks and to propose how informal networked learning practices of professionals can be integrated within the formal organisational structures. the special issue concludes with a commentary, in which de laat and strijbos abstract and synthesize the key themes arising from the special issue contributions and outline a range of recommendations and directions for future research and development in the field. in his opening editorial in the first issue of frontline learning research lehtinen (2013) outlined the rationale: “…to develop a journal which would explicitly support innovative theoretical and methodological thinking and increase dynamics in the field.” (p.1). this special issue offers a number of innovative methodological and theoretical insights and ideas discussed in detail in de laat and strijbos (this issue). first, it contributes much needed empirical evidence about individual‟s networked learning practices at a range of levels from ego-networks and sub-networks to whole networks, elucidating the configurations and contents of these networks and their value for learning, development and improvement of professional practice. second, the special issue provides examples of application of social network analysis to study learning ties within a network rather than only the structure and dynamics of the network, generating new directions for future research. we hope you benefit from the contributions assembled here. acknowledgments i would like to thank frontline learning research and in particular erno lehtinen for the opportunity to publish this special issue and also to inneke berghmans and eva vanhee for their excellent editorial support. i am very grateful to the anonymous reviewers who closely engaged with the papers, providing feedback to the authors at a short notice. last but not least, thank you to the authors who contributed to this special issue. references dron, j., & anderson, t. (2007). collectives, networks and groups in social software for e-learning. in proceedings of world conference on e-learning in corporate, government, healthcare, and higher education, quebec. [online] www.editlib.org/index.cfm/files/paper_26726.pdf. http://www.editlib.org/index.cfm/files/paper_26726.pdf a. margaryan 3 | f l r hakkarainen, k., palonen, t., paavola, s., & lehtinen, e. (2008). communities of networked expertise. bingley, uk: emerald. lehtinen, e. (2013). frontline research in an accessible and flexible way. frontline learning research 1(1), 1-2. littlejohn, a., & margaryan, a. (2013). technology-enhanced professional learning: processes, practices and tools. london/new york: routledge. mccormick, r., fox, a., carmichael, p., & procter, r. (2011). researching and understanding educational networks. london/new york: routledge. frontline learning research 3 (2014) 78-82 issn 2295-3159 corresponding author: tomi silander, xerox research centre europe, www.xrce.xerox.com, tomi.silander@xrce.xerox.com ; petri nokelainen, university of tampere, www.uta.fi/edu, petri.nokelainen@uta.fi. http://dx.doi.org/10.14786/flr.v2i1.107 78 | f l r using new models to analyze complex regularities of the world: commentary on musso et al. (2013) petri nokelainen a , tomi silander b a university of tampere, finland b xerox research centre europe, france abstract this commentary to the recent article by musso et al. (2013) discusses issues related to model fitting, comparison of classification accuracy of generative and discriminative models, and two (or more) cultures of data modeling. we start by questioning the extremely high classification accuracy with an empirical data from a complex domain. there is a risk that we model perfect nonsense perfectly. our second concern is related to the relevance of comparing multilayer perceptron neural networks and linear discriminant analysis classification accuracy indices. we find this problematic, as it is like comparing apples and oranges. it would have been easier to interpret the model and the variable (group) importance’s if the authors would have compared mlp to some discriminative classifier, such as group lasso logistic regression. finally, we conclude our commentary with a discussion about the predictive properties of the adopted data modeling approach. keywords: artificial neural networks; commentary; model-fit; generative and discriminative models; algorithmic data modeling 1. introduction statistical methods are constantly developed, not only within statistics, but also in other disciplines, such as, physics, economics, bioinformatics, linguistics, and computer science. we therefore are very sympathetic to the attempts to promote new methods for analyzing educational data. however, also in this research field, for years there has been an emphasis on the predictive modeling, for example, to learn structures from the data (nokelainen, silander, ruohotie, & tirri, 2007; tirri, nokelainen, & komulainen, 2013) and to predict class membership (nokelainen & ruohotie, 2009; nokelainen, tirri, campbell, & walberg, 2007; villaverde, p. nokelainen & t. silander 79 | f l r godoy, & amandi, 2006). the recent boom of data analytics has further increased the efforts in this front. one methodological rationale behind this development is that predictiveness guards against over-fitting and serves as a natural criterion for the quality of the model. classical statistical literature was not emphasizing this aspect, since models were kept relatively simple to avoid over-fitting and to keep the calculations reasonable. in addition, much of the theory concerned asymptotic behavior in which case over-fitting is usually not an issue. increased computing power now allows more complicated models, such as bayesian, fuzzy and neural networks, to be used. while the increased flexibility brings benefits, there are also possibilities to make new kind of errors in the analysis. since we share the enthusiasm to promote new methods, we also feel that is of utmost importance to perform the analyses with these new methods using extremely high methodological standards. in this respect, we find some of the procedures followed in the recent article by musso, kyndt, cascallar and dochy (2013) problematic. before discussing about these issues in detail, we wish to indicate that we agree with edelsbrunner and schneider’s (2013) previous commentary on this article where they state that there are other data analysis techniques with similar properties than anns, but without the drawbacks. 2. fitting to the test data our first concern is the reported 100% classification accuracy in such a complex domain, and the lack of thorough discussion of this issue. multilayer perceptron (mlp) neural networks are universal function approximators (lek & guegan, 1999). with enough twisting of the parameters, one can use them to implement any classification rule (schittenkopf, deco, & brauer, 1997). consequently, the networks could in theory also be designed to explain the version of the dataset in which the gpa scores would be randomly assigned to the students. what knowledge does such a model (that can explain anything) extract from the real world? one persuasive answer does indeed lie in prediction. only the regularities help one to generalize beyond the training sample, that is, to predict. but here one needs to be very careful. to do this right, the data must first be split into two parts and then the model must be built using only the first part. the testing should be done with the second part of the data – the part that was not used in the model building process at all. the big question is: can we trust us to be able to refrain from “cheating” (using the test data)? in order to avoid this, it would be best to gather the test data after building the model, or to separate it from the training data in the very beginning, and give it to somebody else who will then, after the model has been built, test the accuracy of the model – once and for all! the paper by musso and her colleagues (2013) practically acknowledges that such a discipline was not rigorously followed. the network structure and learning parameters were adjusted to maximize the accuracy in test data. many models were tested to achieve this. even the division of the data into training and test samples was manipulated in order to “… maximize the training sample while preserving the appearance of all detected patterns in the testing sample …” (musso et al., 2013, 60). now, one cannot totally exclude the possibility that the authors actually promote this methodology as a sound one. take a maximally flexible model family, find the most parsimonious model that fits 100% to the data, and then analyze the model. but if that were the case, why torture oneself with the tedious manual work to find 100% fit (yes fit, not generalization) to the test data? it would be easier to just fit to the whole data set – but that would break the illusion of prediction. 3. comparison with the linear discriminant analysis we find that the authors’ decision to compare the model to the other models sets a very good example that should more often be followed in the educational research. such comparisons are widely used in machine learning (e.g., demšar, 2006). however, comparing the multilayer perceptron and the discriminant analysis raises some questions. behind the linear discriminant analysis is a linear discriminant model that defines a joint probability distribution for the whole 19-variate (18 independent variables + gpa p. nokelainen & t. silander 80 | f l r class) data vector. such joint probability distributions can be used for classification, since the conditional probability p(gpa-class | predictors) is proportional to the joint distribution p(gpa-class & predictors). these kinds of classifiers are usually called generative classifiers (e.g., xue & titterington, 2008), since they are based on the models that can be used to sample the whole (19-variate) data vectors. mlps are not generative classifiers, but so called discriminative classifiers. they are built to directly estimate the conditional distribution p(gpa-class | predictors) without modeling the relationships among the predictors. (reading the paper sometimes makes you feel that the authors claim otherwise.) while the linear discriminant model da1 used in the musso et al. (2013) paper has about 2*18 + 18*18 = 360 parameters, the neural network model has 18*15*2 = 540 parameters. the difference in number of parameters is not huge, but all the parameters of the mlp are used for modeling the conditional distribution, while the parameters in the linear discriminant model also take care of modeling the relationships between variables. since the linear discriminant is also a predictive classifier, one cannot but wonder why the confusion matrices for linear discriminants were not reported. those numbers surely would have fitted to the same space without any problem. on the other hand, it is plausible that any differences found are due to the other classifier being generative and the other one discriminative. it would have been much more meaningful to compare the mlp to some discriminative classifier such as a logistic regression, or better yet, some sparse version of it such as the group lasso with interaction terms (meier, van de geer, & bühlmann, 2008) that would make interpreting the model and the variable (group) importances much easier. furthermore, the musso et al. (2013) paper is very unclear about how the variable importances have been calculated. the attempt to follow the references only lead to the statements like “this has been implemented in software x” or to an unpublished technical report by one of the authors. 4. two cultures according to breiman (2001b), there are two statistical modeling cultures. the data modeling culture assumes that the data are generated by a given stochastic data model (such as linear or logistic regression). the algorithmic modeling culture treats the data mechanism as unknown, using, for example, decision trees and neural networks. although the first of these two cultures, focusing on data models, is still dominating, many fields outside statistics are rapidly adopting a wide variety of tools. neural networks are often considered as black-box models that do not offer a good explanation and understanding of the domain (correa, bielza, & pamies-teixeira, 2009). consequently, such models are sometimes hastily deemed as unsuitable for much of the science. we would like to take the opportunity to say a word for such black-box models along the lines expressed by a statistician, leo breiman (1928-2005). world may not be a simple place. while among the simple theories there are those who most closely approximate the complex reality, it is a priori possible that none of those simple theories, even the best of them, approximate the situation well. if the model does not predict well, one can argue that it has not captured the regularities of the world, so what insight would understanding and interpreting such a model offer us. (breiman, 2001b.) most of the statistical community would agree that only if the model is reasonably good (and we mean generalization, not just fit to the sample), interpretation makes sense. edelsbrunner and schneider (2013) indicate in their commentary on this article that whenever possible, more theory-driven data modeling techniques should be preferred. however, if we limit ourselves to the models that can be easily interpreted, we may end up discarding models that truly capture important regularities of the domain. there are two different strategies then to extract true knowledge from the world. the first one is a classical one in which we try to find a well predicting model among the easily interpretable ones. this is the path that should always be attempted. unfortunately, we suspect that it was not seriously pursued in the article by musso et al. (2013). it is also possible to try to build a well predicting model, even if it is not that easy to interpret, and then put more effort to squeeze out the knowledge from the model. one could argue that this is what happens, when you ask a doctor why she made the diagnosis she did. the answer will (only) be some approximation of the real reason. still doctors are considered useful. p. nokelainen & t. silander 81 | f l r we have an educated guess, that such a procedure is behind the independent variable importance measures featured in the article. naturally, such procedures should be carefully documented in order to understand what kind of information we have managed to extract from the model. the article leaves the impression of the claim that artificial neural networks were somehow especially good for inferring how different complex patterns of variables affect the outcome. however, the presented results list only univariate importance of variables. how could that possibly tell us anything relevant about complex patterns? neural networks are by no means the only black-box models that can be successful in the prediction. many ensemble learning based or motivated methods, for instance random decision forests (breiman, 2001a) and bayesian additive regression trees (chipman, george, & mcculloch, 2010), are among such models. ensemble methods have reached very high classification accuracies by using several (or growing a forest of) decision trees on the same data instead of a single-tree predictor. ever increasing data sizes (e.g., massive open online courses, moocs, may have 100 000 students with all their data gathered automatically to the digital form) and increasing computer power may well shift focus from small, simple and understandable models, to the big, complex black-box models. but hopefully some of that computing power can also be used to extract understandable (even if not always very close to truth) approximations of the true complex regularities of the world. key points artificial neural networks (ann) certainly provide interesting modeling possibilities for educational scientists, but they also set certain challenges for the design of the study and interpretation of the results. the article shows a very good example by comparing the results of ann to a conventional data modeling approach, but the comparison should have been made between two discriminative classifiers. ensemble methods provide a modern and powerful alternative to neural networks as they use predictions of several models built during learning process instead of using a single model. references breiman, l. (2001a). random forests. machine learning, 45, 5–32. doi:10.1023/a:1010933404324 breiman, l. (2001b). statistical modeling: the two cultures. statistical science, 16(3), 199–231. doi:10.1214/ss/1009213726 chipman, h. a., george, e. i., & mcculloch, r. e. (2010). bart: bayesian additive regression trees. the annals of applied statistics, 4(1), 266–298. doi:10.1214/09-aoas285 demšar, j. (2006). statistical comparison of classifiers over multiple data sets. journal of machine learning research, 7, 1–30. correa, m., bielza, c., & pamies-teixeira, j. (2009). comparison of bayesian networks and artificial neural networks for quality detection in a machining process. expert systems with applications, 36, 7270– 7279. doi:10.1016/j.eswa.2008.09.024 lek, s., & guegan, j. f. (1999). artificial neural networks as a tool in ecological modelling, an introduction. ecological modelling, 120, 65–73. doi:10.1016/s0304-3800(99)00092-7 meier, l., van de geer, s., & bühlmann, p. (2008). the group lasso for logistic regression. journal of the royal statistical society: series b, 70(part 1), 53-71. doi:10.1111/j.1467-9868.2007.00627.x musso, m. f., kyndt, e., cascallar, e. c., & dochy, f. (2013). predicting general academic performance and identifying differential contribution of participating variables using artificial neural networks. frontline learning research, 1, 42-71. doi:10.14786/flr.v1i1.13 p. nokelainen & t. silander 82 | f l r nokelainen, p., silander, t., ruohotie, p., & tirri, h. (2007). investigating the number of non-linear and multi-modal relationships between observed variables measuring a growth-oriented atmosphere. quality & quantity, 41(6), 869-890. doi:10.1007/s11135-006-9030-x nokelainen, p., & ruohotie, p. (2009). non-linear modeling of growth prerequisites in a finnish polytechnic institution of higher education. journal of workplace learning, 21(1), 36-57. doi:10.1108/13665620910924907 nokelainen, p., tirri, k., campbell, j. r., & walberg, h. (2007). factors that contribute or hinder academic productivity: comparing two groups of most and least successful olympians. educational research and evaluation, 13(6), 483-500. doi:10.1080/13803610701785931 schittenkopf, c., deco, g., & brauer, w. (1997). two strategies to avoid overfitting in feedforward networks. neural networks, 10(3), 505-516. doi:10.1016/s0893-6080(96)00086-x schneider, m., & edelsbrunner, p. (2013). modelling for prediction vs. modelling for understanding: commentary on musso et al. (2013). frontline learning research, 1(2), 99-101. doi:10.14786/flr.v1i2.74 tirri, k., nokelainen, p., & komulainen, e. (2013). multiple intelligences: can they be measured? psychological test and assessment modeling, 55(4), 438-461. doi:10.1007/978-94-6091-758-5_1 villaverde, j. e., godoy, d., & amandi, a. (2006). learning styles’ recognition in e-learning environments with feed-forward neural networks. journal of computer assisted learning, 22, 197–206. doi:10.1111/j.1365-2729.2006.00169.x xue, j-h., & titterington, d. m. (2008). comment on “on discriminative vs. generative classifiers: a comparison of logistic regression and naive bayes”. neural processing letters, 28(3), 169-187. doi:10.1007/s11063-008-9088-7 frontline learning research 7 (2014) 16 issn 2295-3159 corresponding author: frank fischer (frank.fischer@psy.lmu.de) & sanna järvelä (sanna.jarvela@oulu.fi) doi: http://dx.doi.org/10.14786/flr.v2i4.131 1 | f l r methodological advances in research on learning and instruction and in the learning sciences frank fischer a , sanna järvelä b a university of munich, germany b university of oulu, finland recent years have seen a dynamic growth of research communities addressing conditions, processes and outcomes of learning in formal and informal environments. two of them have markedly advanced the field: the community on research on learning and instruction that has been organized in the european association for research on learning and instruction (earli), and the learning sciences community, including the computer-supported collaborative learning community, organised in the international society of the learning sciences (isls). in this special issue we bring together excellent young researchers from these two communities who are currently contributing to advancing the methodology. we are convinced that the methodological developments in these two communities have a lot of commonalities as the core phenomena under investigation and the core questions are related to conditions, processes and outcomes of learning. common for both of these communities is that they have strong roots in cognitive science. however, we also assume that there are substantial differences in these methodological developments, as the foci of the two communities differ in important respects. most importantly, the learning sciences have strong theoretical roots in situative cognition and socio-cultural approaches focusing on learning activities in authentic contexts. the main assumption underlying this focus is that knowledge is represented in activity structures rather than solely in the head (greeno, 2006). therefore, removing the activities of their social and physical contexts into which they belong will change their nature and, hence, research would lead to invalid results, because only a part of the knowledge that is relevant for effectively participating in a practice can be investigated. given these assumptions, it comes as no surprise that learning sciences research focuses on learning in authentic activities in contexts rather than settings stripped off the context for reasons of control in the experimental studies. besides experiments and mixed-method approaches a core methodology that originated in the clear need for alternatives to deductive-experimental methods for early phases of such field research is design-based research with a cyclic process and the goal to improve a practice and to develop a modest and local theory. dbr has its origins in seminal papers by ann brown (1992) and by allan collins f. fischer & s. järvelä 2 | f l r (1992) as well as in influences coming from computer science (see hoadley & van haneghan, 2011). as knowledge is seen to be tied to activities in practices rather than to a single individual, units of analysis beyond the individual (e.g., network, or activity) are rather the rule than the exception in learning sciences research. explorations of different units of analysis are happening in both communities, of course, but they are more pronounced in the learning sciences community. due to the theoretical roots in socio-cultural thinking and situative cognition the relation of the social and material environment to individual cognition is at the core of theorizing in the learning sciences. this is perhaps most obvious in research on computersupported collaborative learning (see dillenbourg, järvelä & fischer, 2009). as the activities or practices are seen as the core medium of knowing, and the practices differ a lot between communities, domains and disciplines, research in the learning sciences has an important focus on disciplinary practices (e.g. herrenkohl & cornelius, 2013). as the use of tools is a key feature of any community, tool appropriation and use are important foci in learning sciences research. in the learning sciences, the concept of tool is often very broadly defined ranging from tools like scientific concepts to digital technologies. research in the learning and instruction community is characterized by a strong connection of basic research to applied field studies. the field has deeper roots into experimental psychology and general psychology of learning and motivation. traditionally, research on learning and instruction has focused on basic processes of cognition and learning and then applied these principles to teaching and learning practices. for example, understanding metacognitive processes in human learning (flavell, 1979) has led many research groups to making effective interventions to the classroom contexts (azevedo & hadwin, 2005). also research on self-regulated learning has tried to integrate empirical evidence on basic processes of cognition, motivation and emotion into broader applications and interventions in the classrooms, where teacher’s role, students’ activities and features of the learning environment have been synchronized to serve learning (e.g., dignath, buettner & langfeld, 2008). in recent years, basic research on learning and instruction has been helpful for designing powerful learning environments, where knowledge about student’s cognitive, motivational and emotional processes and their individual differences has been applied to instructional design. for example, knowledge on scientific reasoning and on worked-out examples has been applied in developing guidance for inquiry learning (mulder, lazonder & de jong, 2014) and collaborative learning (kollar, ufer, reichersdorfer, vogel, fischer & reiss, 2014). in the learning and instruction community one of the current strong emphases is on methodological orientations linking learning research to natural science brain research. the educational neuroscience movement seems to be more pronounced in research on learning and instruction than in the learning sciences. this is consistent with the deeper roots of learning and instruction research in general and experimental psychology, which has developed a strong neuroscience orientation over the last years. in addition, methodologies are being developed addressing the temporal characteristics of learning. in both communities, quantitative approaches to the analysis of temporal aspects of the learning process have been developed over the last years. it is argued that the explanatory power and the validity of the analyses can be improved dramatically by including the time information that has typically been neglected in many studies on individual and collaborative learning. in research on learning and instruction, this new focus has originated as a consequence of a conceptual shift, as molenaar (this volume, p. xx) puts it: “constructs formerly viewed as personal traits, such as self-regulated learning and motivation, are now conceptualized as a series of events that unfold over time”. there are several arguments in support for this point also in recent publications in the learning sciences (e.g., reimann, 2009). f. fischer & s. järvelä 3 | f l r there are four main potentials for innovation resulting from these developments for learning research, no matter if situated in research on learning and instruction or in learning sciences research. potential #1: increased gain in scientific understanding through more “messy studies” when investigating “real” learning in new fields. it seems inadequate to presume a purely deductive experimental approach in fields where the set of potentially influential variables is unknown. learning research is not an exception here, the same applies to other fields like, e.g. physics, where pioneering research at the edges of current scientific knowledge is more “messy” as well (wieman, 2014). dbr approaches, although still in their infancies, might well develop into a standard methodology for pioneering research on “real learning“ in authentic settings, also in research on learning and instruction. in this special issue, svihla (this volume) reports on recent developments in dbr that address the issues of scalability and generalizability: designbased implementation research (dbir). this might be a promising alternative approach to randomized trial approaches to implementation research in fields where the set of influential and to-be-controlled variables in real formal and informal learning environments is far from clear. because of its design focus, dbr and dbir might contribute to advancing learning research beyond generating new scientific knowledge: they might have the potential to build bridges into practice and increase the credibility and trustworthiness of learning research. an alternative approach is suggested by stegmann (this volume), who addresses the issue of control in studies of complex, collaborative learning environments. he argues for a more systematic use of nomological networks on the conceptual level in connection with as-controlled-as-possible empirical studies that include measures of learning processes as their methodological core. potential #2: more comprehensive understanding of learning phenomena through the use of methodologies that can handle multiple units of analysis and include process analyses. units like the activity, the group or the collective could become standard for questions that transcend the individual’s learning. it will be a challenge how to conceptually deal with this paradigm shift: talking about “learning“ also with respect to super-individual units. for example, should team learning be considered as a whole, or should the term “learning” be reserved for the individual and different concepts should be used to describe what is happening in activities or collectives? an even more far reaching question is to what extent phenomena on super-individual levels should be traced back (or be reduced as some would prefer to say) to the individual contribution, i.e. social phenomena are treated as a result of interacting individuals, and the phenomena can be fully explained by the individual contributions and reactions. increasingly there is research arguing that some social phenomena in contexts of learning cannot be reasonably reduced to the individuals involved (cress, held & kimmerle, 2013; eberle, stegmann & fischer, 2014; stahl, 2006). in this special issue, stegmann’s (this volume) work is additionally addressing this aspect. he describes measures of individual cognition and argumentative discourse in computer-supported small groups and exemplifies approaches to a synchronized analysis of individual cognition and group discourse to address the mutual impact. we argue that systematically employing other units of analysis in learning research than the individual would not only advance research on learning in context, but also help to build bridges into other social sciences that are sometimes hesitating because of the exclusivity of the individual-centric perspective of some learning researchers. potential #3: overcoming overreliance on self-reports: from personal constructs to series of interactions unfolding over time. many learning researchers are currently working on developing alternative conceptualisations of well-established psychological constructs such as self-regulation or motivation. there are shortcomings of relying solely on self-reports in questionnaires (e.g. zimmerman, 2008) to measure personal constructs, such as low predictive value for behaviour in real problem-solving situations. learning researchers have therefore begun to develop methodological approaches that use behaviour or interaction in problem-solving situations as indicators for these constructs. an example from research in the learning f. fischer & s. järvelä 4 | f l r sciences is dan hickeys work on disciplinary engagement in a discussion (filsecker & hickey, 2014) as a complementary measure of motivation. in this special issue, inge molenaar’s work is representing this broader issue. she focuses on the temporal characteristics of learning processes that are typically missed when only self-report measures are used or observational data is aggregated into frequencies over the whole learning process under consideration. also recent advances in the use of computer-generated trace data for understanding patterns and processes of students’ learning (malmberg, järvenoja & järvelä, 2013) have advanced the instructional design field for developing scaffolding and prompts for computer supported learning (järvelä & hadwin, 2013). potential # 4: building bridges between research on learning and cognitive neuroscience. there have been discussions if the gap between education and neuroscience might require a bridge too far. however, recent advances in cognitive neuroscience are encouraging. research on learning and instruction and in the learning sciences are increasingly interested in the biological basis of the learning phenomena under investigation and some of these ideas have already been applied e.g. to mathematics learning (hannula, lepola & lehtinen, 2010). in the learning sciences and the learning and instruction community there is increasing awareness of the possibilities to analyse processes that are not readily accessible for behavioural research. one can hope that in the future, researchers on learning and instruction and in the learning sciences will be able to successfully point out interesting learning phenomena to neuroscientists (varma, mccandliss & schwartz, 2008). these often complex and dynamic phenomena are typically highly challenging for contemporary neuroscientists. at the same time one can hope that researchers in learning and instruction as well as in the learning sciences would become more receptive for stimulations coming from unexplained phenomena in neuroimaging studies on cognition and learning. de smedt (this volume) addresses these questions and elaborates on some convincing examples from mathematics learning that give evidence for a productive interaction between research on learning and instruction and cognitive neuroscience. he argues that the successful interaction crucially depends on finding the right level of resolution or granularity when involving neuroscience methods. we argue that it is now a good point in time to start exploring this interaction from both research on learning and instruction and in the learning sciences more systematically. this would enhance the interface of learning research to the natural sciences. at this interface there is a considerable potential for innovation. conclusion research on learning and instruction and research in the learning sciences have seen considerable methodological advancements in recent years. although a certain specialisation can be seen due to differences in some of the basic assumptions we see good reasons for transferring these innovations between the research communities. we see four potentials for innovation for learning research resulting from these methodological developments: (1) increased gain in scientific understanding through more “messy studies” when investigating “real” learning in new fields, (2) more comprehensive understanding of learning phenomena through the use of methodologies that can handle multiple units of analysis and entail processes analyses, (3) overcoming overreliance on self-reports: from personal constructs of learning and motivation to series of interactions unfolding over time, and (4) building bridges between research on learning and cognitive neuroscience. the contributions to this special issue are each addressing one of these potentials. f. fischer & s. järvelä 5 | f l r references azevedo, r. & hadwin, a. f. (2005). scaffolding self-regulated learning and metacognition – implications for the design of computer-based scaffolds. instructional science, 33(5), 367–379. doi:10.1007/s11251-005-1272-9 brown, a. l. (1992). design experiments: theoretical and methodological challenges in creating complex interventions in classroom settings. journal of the learning sciences, 2(2), 141-178. doi:10.1207/s15327809jls0202_2 collins, a. (1992). toward a design science of education. in e. scanlon & t. o'shea (eds.), new directions in educational technology (pp. 15-22). new york: springer. doi:10.1007/978-3-642-77750-9_2 cress, u., held, c., & kimmerle, j. (2013). the collective knowledge of social tags: direct and indirect influences on navigation, learning, and information processing. computers & education, 60(1), 5973. doi:10.1016/j.compedu.2012.06.015 dignath, c., büttner, g. & langfeldt, h.-p. (2008). how can primary school students acquire self-regulated learning most efficiently? a meta-analysis on interventions that aim at fostering self-regulation. educational research review, 3, 101-129. doi:10.1016/j.edurev.2008.02.003 dillenbourg, p., järvelä, s., & fischer, f. (2009). the evolution of research on computer-supported collaborative learning. in n. balacheff et al. (eds). technology-enhanced learning (pp. 3-19). springer, the netherlands. doi:10.1007/978-1-4020-9827-7_1 eberle, j., stegmann, k., & fischer, f. (2014). legitimate peripheral participation in communities of practice: participation support structures for newcomers in faculty student councils. journal of the learning sciences, 23(2), 1-29. doi:10.1080/10508406.2014.883978 greeno, j. g. (2006). learning in activity. in r. k. sawyer (ed.), the cambridge handbook of the learning sciences (pp. 79-96). new york: cambridge university press. filsecker, m., & hickey, d. t. (2014). a multilevel analysis of the effects of external rewards on elementary students' motivation, engagement and learning in an educational game. computers & education, 75, 136-148. doi:10.1016/j.compedu.2014.02.008 flavell, j. h. (1979). metacognition and cognitive monitoring: a new area of cognitive-developmental inquiry. american psychologist, 34, 906-911. doi:10.1037/0003-066x.34.10.906 hannula, m., lepola, j. & lehtinen, e. (2010). spontaneous focusing on numerosity as a domain-specific predictor of arithmetical skills. journal for experimental child psychology, 107, 394-406. doi:10.1016/j.jecp.2010.06.004 herrenkohl, l. r. & cornelius, l. (2013). investigating elementary students' scientific and historical argumentation. journal of the learning sciences, 22(3), 413-461. doi:10.1080/10508406.2013.799475 hoadley, c. & van haneghan, j. (2011). the learning sciences: where they came from and what it means for instructional designers. in r. a. reiser & j. v. dempsey (eds.), trends and issues in instructional design and technology (3rd ed., pp. 53-63). new york: pearson. järvelä, s. & hadwin, a. (2013). new frontiers: regulating learning in cscl. educational psychologist, 48(1), 25-39. doi:10.1080/00461520.2012.748006 kollar, i., ufer, s., reichersdorfer, e., vogel, f., fischer, f., & reiss, k. (2014). effects of collaboration scripts and heuristic worked examples on the acquisition of mathematical argumentation skills of teacher students with different levels of prior achievement. learning and instruction, 32, 22-36. doi:10.1016/j.learninstruc.2014.01.003 malmberg, j., järvenoja, h. & järvelä, s. (2013). patterns in elementary school students’ strategic actions in varying learning situations. instructional science, 41, 933–954. doi:10.1007/s11251-012-9262-1 f. fischer & s. järvelä 6 | f l r mulder, y. g., lazonder, a. w., de jong, t. (2014). using heuristic worked examples to promote inquirybased learning. learning and instruction, 29, 56-64. doi:10.1016/j.learninstruc.2013.08.001 reimann, p. (2009). time is precious: variable-and event-centred approaches to process analysis in cscl research. international journal of computer-supported collaborative learning, 4(3), 239-257. doi:10.1007/s11412-009-9070-z stahl, g. (2006). group cognition. cambridge, ma: mit press. varma, s., mccandliss, b. d., & schwartz, d. l. (2008). scientific and pragmatic challenges for bridging education and neuroscience. educational researcher, 37, 140-152. doi:10.3102/0013189x08317687 wieman, c. e. (2014). the similarities between research in education and research in the hard sciences. educational researcher, 43(1), 12-14. doi:10.3102/0013189x13520294 zimmerman, b. j. (2008). investigating self-regulation and motivation: historical background, methodological developments, and future prospects. american educational research journal, 45(1), 166-183. doi:10.3102/0002831207312909 microsoft word 31-101-2-ce_salminen_final proof.docx frontline learning research 1 (2013) 72 80 issn corresponding author: jenni salminen, department of education, university of jyväskylä, p.o.box 35, jyväskylä, finland 40014, jenni.e.salminen@jyu.fi, t +358-40-805 4032, f +358-14-260 1761. doi case study on teachers’ contribution to children’s participation in finnish preschool classrooms during structured learning sessions jenni elina salminen a a university of jyväskylä, finland article received 15 may 2013 / revised 27 june 2013 / accepted 14 august 2013 / available online 27 august 2013 abstract the main aim of this study was to identify different teaching practices and explore the types of opportunities that they provide for children’s participation in four different finnish preschool classrooms for 6-year olds during structured learning sessions. observational data of four preschool teachers were analyzed according to the principles of qualitative content analysis. three themes of teachers’ practices were identified, which described the key practices through which teachers influence children’s participation, namely, through discussion and conversations; by referring to shared rules and managing the classroom; and through demonstrating pedagogical sensitivity and understanding towards children’s active participation. further, each teacher was observed implementing these practices in a unique combination in their classrooms, thus, creating different opportunities for participation. the four teachers showed a constructive, enabling, reserved or restrictive/unbalanced stance towards children’s participation. the results of this study highlight the importance of teachers’ pedagogically sensitive attitude as the key to children’s participation. given that the advantages of participation to learning and development are well established, the results also point to a need to evaluate the prevailing pedagogy and practices more closely from the perspective of participation. keywords: case-study; participation; preschool; teacher–child interactions; teaching practices. 1. introduction extensive research has suggested that one of the best ways to support learning is through encouraging active participation of children already in early childhood classroom contexts (e.g., pramling-samuelsson & j.e. salminen 73 | f l r sheridan, 2003; hännikäinen & rasku-puttonen, 2010). this study was set to explore teachers’ contribution to children’s participation, i.e., children’s right to experience respect and confidence in partnership with adults (cockburn, 2005; emilson & folkesson, 2006) in finnish preschool classrooms for 6-year old children. according to sociocultural approach interactional processes are the key elements for learning and development (vygotsky, 1978; mercer & littleton, 2007), so participation is also enabled in the interaction between teacher and children. participation demands that teacher values a child’s own ways of experiencing, understanding and exploring the world (pramling-samuelsson & sheridan, 2003), and that he or she is able to consider these practices as an important part of learning. further, genuine respect shown towards children by teachers has a significant impact on the relationships they build with children in care and educational backgrounds (laevers, 2005). thus, participation in educational settings can be seen to be contingent upon teachers’ decisions and ideas. through their professional role, teachers are the central figure in determining the learning opportunities available to children (hännikäinen, de jong, & rubinstein reich, 1997; pianta, 1999) and also how those children are encouraged to participate. according to recent studies the essential features that encourage children to participate are when teacher’s interest comes close to children’s own views (emilson & folkesson, 2006), when rules are negotiated and shared (e.g. bohn, roehrig, & pressley, 2004; hännikäinen, 2005) and when teachers provide children with a feeling of being part of the group and of being listened to (hännikäinen & rasku-puttonen, 2010; johansson & sandberg, 2010). in a previous study by salminen et al. (2013b), the contribution of teachers to the social life within preschool classrooms (i.e. for 6-year-olds) was explored through a ‘best-practices’ perspective. some of the practices that enhanced children’s participation included supporting children’s constructive and respectful friendships, working according to shared social rules in group contexts allowing individual children certain levels of leaderships and inviting children to contribute to simple decision making processes (salminen et al., 2013b). the inspiration for the current study was to extend these earlier findings, in particular those relating to participation. thus, i sought to investigate the naturally-occurring variation among a smaller sample of four finnish preschool teachers by identifying teachers’ key-practices and exploring the unique combinations of these practices that can be seen to provide ample support and opportunities for children to participate in different classroom contexts. in the field of participation studies, emilson and folkesson (2006) have studied how teachers’ control, in terms of classification and framing, affects children’s participation. the current study aimed to widen the perspective from teacher control to classroom interaction more broadly, since participation occurs in a socially shared network of interactions between adults and children. further, aim was to identify the ways in which teachers may affect children’s participation –– either by enhancing or preventing it –– during structured learning sessions. this was necessary, since a majority of the formal learning sessions (i.e., content driven purposeful sessions and about 45 minutes in length) in finnish preschool classrooms are constructed around teacher-led formats (e.g., hujala et al., 2012; salminen et al., 2013b). two related research questions were addressed. (1) what are the key practices by which finnish preschool teachers enable or disable children’s participation in a variety of classroom situations? (2) which combinations of teacher support do these key practices create for children’s participation in four different preschool classrooms? 2. methods 2.1 data the data for this study were collected as part of the large-scale ‘first steps’ follow-up study (lerkkanen et al., 2006). four finnish preschool teachers were selected as informants from the total of 49 of those participating in the ‘first steps’ follow-up study. in a previous study by salminen et al. (2012), the original 49 teachers were divided into four subgroups on the basis of observed classroom quality, as assessed j.e. salminen 74 | f l r with the classroom assessment scoring system (class; pianta, laparo, & hamre, 2008), utilizing the mixture modelling procedure of the mplus 5.0 statistical package. class is designed to measure the classroom level variables (i.e., observed indicators of classroom quality) in three domains: (1) emotional support, (2) classroom organization, and (3) instructional support, by rating each aspect numerically from 1 to 7. the profiles from which the cases of the current study were selected can be summarised as follows: profile 1 – highest quality (prevalence 53%); profile 2 – medium emotional, organizational, and instructional quality (prevalence 29%); profile 3 – medium to low emotional and instructional quality, medium organizational quality (prevalence 12%); and profile 4 – lowest quality (prevalence 6%). teachers for this study were selected to represent each of the four subgroups in order to investigate the maximum variation in practices among teachers as well as their relative representativeness throughout the whole dataset. further, a previous study by salminen et al. (2013a) partially utilized the same data of four teachers (with the exception of one teacher) in a case analysis that explored teachers’ instructional teaching practices. results from this work indicated that even the teachers at the higher end of the quality continuum employed only relatively low levels of the practices known to emphasise the role of active participation in children’s learning of deeper thinking skills. this was an important justification for further exploring the data for these four teachers: this time, more specifically from the perspective of participation. the qualitative observational data were collected through classroom observations in spring 2007, simultaneous with the live class observations. the observations were conducted on two different days during the morning assembly (i.e., times of more formal educational activities in the morning, before lunch, and nap time) and all of the teachers carried an mp3-player that recorded all teacher–child interactions. the length of each recording was, on average, 53 minutes. all of the recordings were transcribed, resulting 53 pages of transcribed text for the analysis of this study. 2.2 context and the participants of the study before beginning formal schooling at the age of 7 years, finnish children have a statutory right to receive a preschool education free of charge for 1 year. the core curriculum for preschool education (2000) serves as a binding guideline for preschool education throughout the country. nearly 100% of finnish 6year-old children attend preschool education (statistics finland, 2012; taguma, litjens, & makowiecki, 2012) despite its voluntary nature. all of the four teachers were finnish-speaking females, working in preschool classrooms with typical equipment and materials under the national guidelines provided by the core curriculum for preschool education (2000). of the four, diana and berta worked in larger groups of 22 and 24 children, respectively, with teacher’s aids in their classrooms; whereas cecilia and anna both worked in groups of seven children, with no teacher’s aids. however, in finnish preschool classrooms it is typical to divide large groups of children to smaller groups for the more formal learning sessions. hence, during the observed and recorded sessions, both diana and berta were working with smaller group of children (i.e., 8–10 children each). 2.3 data analysis data were analysed according to the principles of qualitative content analysis (patton, 2002; graneheim & lundman, 2004). the observational data for the four teachers were combined and analysed from the perspective of teachers’ practices through which teachers aimed to engage children to daily activities. these practices emerged during interactional episodes of varying lengths, and these episodes (each containing one or several meaningful interactional verbal and non-verbal expressions) were determined as the units of analysis for this study. the analytical process is illustrated in table 1. the first analytical interest of the study was in identifying certain commonalities in the practices of all four teachers. the episodes (i.e., units of analysis) were first combined into eight categories, which provided overarching concepts through which teachers’ practices could be further classified. each of the categories conceptualized teachers’ practices in relation to children’s participation without seeking individual patterns between teachers, but j.e. salminen 75 | f l r rather, by drawing together the practices in a more general level. second, the categories were revised and further combined to wider themes (i.e., pedagogical sensitivity and understanding; discussion and conversations; rules and management), which provided common and more generic denominators for the practice categories identified before. thus, these themes were generated on the basis of the practices that arose from the data of all four teachers, and can be seen to generally represent the key practices through which teachers either encourage or prevent participation of children during the structured learning sessions within this sample. table 1 identifying teachers’ key practices: describing the analytical process as the three themes represented general ways in which to deal with children’s participation, the second analytical interest was to further reflect the three themes (i.e., key practices) to each of the four individual teachers in order to determine which personal combinations of key practices characterized each of them. at this stage of the analysis i re-examined each teachers daily interaction with the children using the aspects provided by the three themes, and examples of individual ways to support children’s participation were gathered (e.g., how does this particular teacher use rules and management, discussions and establishes sensitivity in relation to children’s participation). as a result, each teacher was seen to represent a unique combination of the key practices, which created different opportunities for children’s participation. each teacher case was assigned with a descriptive name according to teachers’ prevailing stance towards children’s participation, namely: diana – constructive stance towards participation; cecilia – enabling stance towards participation; berta – reserved stance towards participation; anna – restrictive/unbalanced stance towards participation. these teacher cases and examples of the key practices will be introduced in the following paragraphs in detail. j.e. salminen 76 | f l r 3. results teachers’ key practices were displayed in unique combinations. these combinations created different learning environments and, thus, affected how children were encouraged to actively be part of a group, activities and the social network of their classrooms. the following results individually present the four teachers according to their unique combination of the key practices (i.e., combinations of pedagogical sensitivity and understanding; discussion and conversations; rules and management). 3.1 diana diana’s classroom was characterized by a constructive stance towards the children’s participation. this teacher was warm and respectful towards the children nearly all the time, establishing high pedagogical sensitivity. this was apparent as diana was well aware of the children’s needs and abilities, and she aimed to keep them engaged with the particular exercise or activities provided (e.g., by saying, “sam please tell the others”, or, “jonah, do you think you could tell what the number of the exercise at hands is?”, as well as, “please, alice, come here and help me to look for the missing syllable”). there were clearly established shared rules in the classroom, and as a result, teaching formed a logical and understandable entity that the children could easily follow, enjoy and participate in. the ways in which diana involved children in daily routines and activities consisted of subtle and delicate reminders of rules such as saying, “children, please listen, let’s listen to mandy for a moment more”, or by whispering softly, “raise your hand if you want to say something”. diana made an attempt to listen to children’s ideas: there were discussions on both academic and social issues. during these discussions diana made it easy for children to find a way to join in. for instance, she asked questions in a very whole-hearted manner, as if not only to hear the children but as if she was honestly pondering the same questions herself. for example, diana commented, “i really enjoyed the warmth of the sunshine today” and then asked, “but what do you think it has done to the snow outside?” when the teacher positioned herself at the children’s level like this it evoked very natural and easy participation from the children, and several such interactions occurred throughout the observed sessions. this type of behaviour, combined with provision of frequent opportunities for children to take turns to answer, for example in a show and tell, or to assist teacher in performing tasks, showed that diana was highly persistent and able in keeping children engaged in activities. her attitude towards the children’s ideas and comments showed she was aiming to understand what the children thought and were telling her. despite the fact that participation was occurring in a goal-oriented, teacher-led format all this time, children’s participation was nevertheless constructive (i.e., children were taken seriously and the classroom agenda was built on their active role). 3.2 cecilia cecilia’s classroom was characterized by an enabling stance towards children’s participation. cecilia repeatedly made children feel like she was listening to them and understood them (e.g., “i know you like these types of exercises, although they are a bit difficult”), indicating teacher’s pedagogical sensitivity. as she sensitively listened to the children, she was also able to monitor their needs and progress most of the time. however, every now and then she missed children’s hints. there were also clearly-established and shared rules in the classroom, which neither cecilia nor children had to be reminded of, and which made participation easier and also contributed to the coherence of the group. cecilia discussed subjects openly with the children throughout the observed sessions. she was, for instance, using children’s daily lives and own experiences efficiently as a tool to engage children in discussions. cecilia’s enabling stance towards participation was apparent when she used inviting questions during the learning sessions (e.g., “if you need to know what’s happening around the world, what types of sources of information can you think of?”, or, “today we are discussing of newspapers, do your mom or dad read the newspaper?”) as well as comments aimed at participation of individual children (e.g., “would you like to try to read this aloud andy?”). both diana and cecilia shared similar practices and personal warmth towards the children. however, throughout the observed sessions cecilia’s practices concerning children’s participation were slightly inconsistent; of the j.e. salminen 77 | f l r two teachers, cecilia’s attitude was less effective for truly understanding the children’s point of view. this was apparent as although cecilia provided children with opportunities to participate, she did not use children’s activity to construct the ideas to aid further learning as diana did and, thus, cecilia’s stance towards participation was enabling rather than constructive. 3.3 berta berta’s classroom was characterized by a reserved stance towards children’s participation. berta showed signs of ambivalent pedagogical sensitivity, since she seemed to be highly responsive towards children’s needs and aimed to achieve participation of the whole group, especially so during exercises and tasks (e.g., “roger’s answer was ‘a hat’. do you [saying to other children] think that roger’s answer was correct?”), but at other times she was less concerned about the children’s perspectives or about truly finding out their thoughts and ideas. the use of rules and management was structured, as berta was very efficient in teaching and managing the classroom. her teaching was logical and it was easy for children to comprehend. for instance, berta said, “you may come here and choose the word that corresponds with the picture, please use the pin and place the word beside the picture”. berta was talking to the children nearly all the time, however, she was restricting children’s participation to discussions by giving children rather short turns, and as a result the children usually only gave answers to the teacher’s questions or produced a few words or short sentences (e.g., “with which letter does the word peruna [potato] begin?”, or, “you are right, this is the face of the person, but could you be a bit more specific? which part of the face is the correct answer?”). as a consequence, berta’s reserved stance towards participation was most clearly apparent in the use of highly structured tasks that allowed only very few chances for children’s ideas or discussions to be used as a valuable way for children to learn and interact. 3.4 anna anna’s classroom was characterized by restrictive/unbalanced stance towards children’s participation. anna had occasional difficulties in monitoring the behaviour, needs and academic performance of the children. she was probably more aware of the children’s academic skills and needs (e.g., inviting children to goal-oriented tasks by using hints, or providing individual additional tasks) rather than their emotional needs (e.g., being unable to soothe restless children and assist them to participate in on-going activities), thus, establishing lower and unbalanced pedagogical sensitivity towards children’s emotional needs. the classroom in general was somewhat disorganized since anna’s practices were inefficient in managing her classroom. anna discussed topics with children, but due to their misbehaviour it was difficult to create an equal and content-driven discussion, when a majority of her time was used to discuss managerial issues. she was, in a sense, forced to cut down children’s turns at the expense of organization to be able to continue working. for instance, anna said, “it is not your turn to speak now”, or, “you are not allowed to speak until you sit quietly and still”, as well as, “you need to step outside unless you can’t be quiet”. as a consequence, autonomous opportunities were not provided to children and children’s participation was discontinuous or even restricted. 4. discussion in relation to the first research question, analysis of the four teacher cases indicated key practices in the four classrooms (i.e., pedagogical sensitivity and understanding; discussion and conversations; rules and management), that were related to children’s participation in preschool classrooms. in addition, the teacher cases showed four different combinations of teacher support, which created unique opportunities for children to participate in both the on-going activities and the social network of their classrooms. these can be discussed further as a response to the second research question. diana and cecilia had established a j.e. salminen 78 | f l r combination of (1) teacher’s pedagogical sensitivity and understanding towards children’s needs, (2) utilizing constructive and shared rules, and (3) involving children to conversations, whereas berta and anna provided fewer opportunities for children to participate. it was noted that both berta and anna had managerial practices that restricted active participation, but for very different reasons. the reason why children’s participation was infrequent in anna’s classroom was that management took too much time because the rules were not clear or shared, whereas in berta’s classroom, which was highly structured, participation did not occur on children’s terms and was thus reserved in nature. these observations indicate the importance of constructive classroom management and organization for children’s participation (see also emilson & folkesson, 2006): neither the lack of behavioural control nor too highly structured management are good ways of enhancing children’s participation. berta and diana shared similar well-managed rules in their classrooms, but the warmth and pedagogical sensitivity was different for these two teachers. diana was perceptive, and identified children’s needs, whereas berta was more concerned with working according to the plans she had made. in a previous study, sandberg and eriksson (2010) highlighted the importance of the intensive respectful discussions between teacher and children in encouraging children’s participation. in addition, constructive and coherent rules and management provide support for working as a group (bohn, roehrig, & pressley, 2004; hännikäinen, 2005). in light of the results of the present study, it seems that neither the intensive respectful discussions between teacher and children nor coherent rules and management alone can create practices that enhance children’s participation within these preschool classrooms. in order to be meaningful, participation requires a teacher’s pedagogical awareness and respectful attitude (pramling-samuelson & sheridan, 2003). this attitude enables teachers to see children’s participation as an important and usable way of learning in preschool. this is of great significance, since being a part of the group is one of the most meaningful things from children’s perspective too (e.g., einarsdottir, 2010). my study indicates that teachers enhance children’s participation through the simple daily routines and pedagogical choices that they make, an idea that is by hännikäinen & rasku-puttonen (2010). however, the findings of this study also showed aspects that may hinder participation in classrooms and unfortunately, such aspects included typical ways of working in preschool classrooms during formal content-driven and teacher-led learning sessions. such practices included working in a predominantly teacher-led format with relatively little control offered to children in deciding or determining what to do, or providing classroom management and rules that are too strict to allow frequent participation. within the classrooms studied, it was teachers’ determination and open-minded stance towards participation that seemed to make a positive difference. further studies are needed to widen the perspective from teachers’ practices to include child interviews or child observations, since in its current form this study cannot suggest how children experienced or perceived the different classroom environments and practices. moreover, it is noteworthy that participation takes different forms depending on the age of the children in the group as well as the cultural expectations (e.g., national curriculums, legislations) addressed within an educational setting. hence, it is necessary to raise a scientific discussion of the importance of children’s participation, and conduct studies in a variety of countries and contexts to gains deeper knowledge and understanding of how participation is experienced and what enhances it in different educational settings. the findings introduce exemplary practices for preschool education and for the discussion about the importance of teachers’ role in enhancing the active role of children in preschool classrooms. the results may provide both practical and educational implications for teachers in their daily work with children by promoting awareness of preschool teachers to the role of teaching practices and teacher–student interactions for children’s participation. since not all teachers were able to fully support children’s participation, this issue should be addressed more carefully in future research and teacher training, and also from the children’s perspective. j.e. salminen 79 | f l r keypoints observational data of four finnish preschool teachers were analysed according to the principles of qualitative content analysis. three themes indicated teachers’ key practices which were related to children’s participation: namely, discussion and conversation; rules and management; pedagogical sensitivity and understanding. combinations of key practices created ample opportunities for children to participate in each classroom. teachers showed a constructive, enabling, reserved or restricted/unbalanced stance towards children’s participation. acknowledgements this study has been carried out in the centre of excellence in learning and motivation research and is financed by the academy of finland (no. 213486 for 2006–2011) and other grants from the same funding agency (no. 213353 for 2005–2008, and no. 125811 for 2008–2009). the author gratefully acknowledges the personal support of prof. maritta hännikäinen and dr. pirjo-liisa poikonen in preparing the manuscript. references bohn, c. m., roehrig, a. d., & pressley, m. (2004). the first days of school in the classrooms of two more effective and four less effective primary-grades teachers. the elementary school journal, 104, 269– 287. cockburn, t. (2005). children’s participation in social policy: inclusion, chimera or authenticity? social policy and society, 4, 109–119. core curriculum for preschool education in finland. (2000). esiopetuksen opetussuunnitelman perusteet [the core curriculum for preschool education 2000]. helsinki, finland: national board of education. retrieved from http://www.oph.fi/download/123162_core_curriculum_for_pre_school_education_2000.pdf einarsdottir, j. (2010).children’s experiences of the first year of primary school. european early childhood education research journal, 18(2), 163–180. emilson, a., & folkesson, a.-m. (2006). children’s participation and teacher control. early child development and care, 176, 219–238. graneheim, u. h., & lundman, b. (2004). qualitative content analysis in nursing research: concepts, procedures and measures to achieve trustworthiness. nurse education today, 24, 105–112. hujala, e., backlund-smulter, t., koivisto, p., parkkinen, h., sarakorpi, h., suortti, o., … korkeakoski, e. (2012). esiopetuksen laatu 2012. [the quality of pre-primary education]. publications by the finnish education evaluation council 61. jyväskylä, finland. retrieved from http://www.edev.fi/img/portal/1354/julkaisu_61.pdf hännikäinen, m. (2005). rules and agreements and becoming a preschool community of learners. european early childhood education research journal, 13, 97–110. hännikäinen, m., de jong, m., & rubinstein reich, l. (1997). our heads are the same size! a study of quality of the child’s life in nordic day care centres. educational information and debate 107. malmö, sweden: malmö university school of education. hännikäinen, m., & rasku-puttonen, h. (2010). promoting children’s participation: the role of teachers in preschool and primary school learning sessions. early years: an international journal of research and development, 30, 147–160. j.e. salminen 80 | f l r johansson, i., & sandeberg, a. (2010). learning and participation: two interrelated key-concepts in the preschool. european early childhood education research journal, 18, 229–242. laevers, f. (2005). the curriculum as a means to raise the quality of early childhood education. implications for policy. european early childhood education research journal, 13(1), 17–29. lerkkanen, m.-k., niemi, p., poikkeus, a.-m., poskiparta, e., siekkinen, m., & nurmi, j.-e. (2006). the first steps study (alkuportaat). unpublished data. university of jyväskylä, finland. mercer, n., & littleton, k. (2007). dialogue and the development of children’s thinking: a sociocultural approach. new york, ny: routledge. patton (2002). qualitative research & evaluation methods (3rd ed.). thousand oaks, ca: sage. pianta, r. c. 1999. enhancing relationships between children and teachers. washington, dc: american psychological association. pianta, r. c., laparo, k. m., & hamre, b. k. (2008). the classroom assessment scoring system. manual, pre-k. baltimore, md: paul h. brookes. pramling-samuelsson, i., & sheridan, s. (2003). delaktighet som värdering och pedagogic [participation as valuation and pedagogy]. pedagogisk forskning i sverige 8, no. 1/2,70–84. salminen, j., lerkkanen, m.-k., poikkeus, a.-m., siekkinen, m., pakarinen, e., hännikäinen, m., poikonen, p.-l., & rasku-puttonen, h. (2012). observed classroom quality profiles of kindergarten classrooms in finland. early education and development, 23, 654–677. salminen, j., hännikäinen, m., poikonen, p.-l., & rasku-puttonen, h. (2013a). a descriptive case analysis of instructional teaching practices in finnish preschool classrooms. journal of research in childhood education, 27, 127–152. salminen, j., hännikäinen, m., poikonen, p.-l., & rasku-puttonen, h. (2013b). teachers’ contribution to the social life in finnish preschool classrooms during structured learning sessions. early child development and care, doi:10.1080/03004430.2013.793182. sandberg, a., & eriksson, a. (2010). children’s participation in preschool – on the conditions of the adults? preschool staff’s concepts of children’s participation in preschool everyday life. early child development and care, 180, 619–631. statistics finland. (2012). suomen virallinen tilasto (svt): esija peruskouluopetus. [official finnish statistics: preschool and primary education]. helsinki: tilastokeskus. retrieved from http://www.tilastokeskus.fi/til/pop/index.html taguma, m., litjens, i., & makowiecki, k. (2012). quality matters in early childhood education and care: finland. organisation for economic co-operation, & development (oecd). retrieved from http://www.oecd.org/edu/preschoolandschool/49985030.pdf vygotksy, l. s. (1978). mind in society. the development of higher psychological processes. cambridge, ma: harvard university press. frontline learning research vol.4 no. 4 special issue (2016) 30 38 issn 2295-3159 corresponding author: william r. penuel, school of education, university of colorado boulder, ucb 249, boulder, co 80305 usa. email: william.penuel@colorado.edu doi: http://dx.doi.org/10.14786/flr.v4i4.205 a social practice theory of learning and becoming across contexts and time william r. penuel, daniela k. digiacomo, katie van horne, ben kirshner university of colorado, united states article received 7 september / revised 9 august / accepted 12 september / available online 19 december abstract this paper presents a social practice theory of learning and becoming across contexts and time. our perspective is rooted in the danish tradition of critical psychology (dreier, 1997; mørck & huniche, 2006; nissen, 2005), and we use social practice theory to interpret the pathway of one adolescent whom we followed as part of a longitudinal study of interest-related learning. a social practice theory calls out the ways people pursue diverse concerns, become aware of new possibilities for action as they move across settings of practice, and learn as they adjust contributions to the flow of ongoing activity and to fit demands and structures of local institutions. it also highlights the ways that existing institutional structures of practice frame the choices people make about how and where to participate in activities. this perspective on learning is potentially transformative, in that it provides a way to promote equity by surfacing issues associated with linkages among settings of practice, networks of actors who support persons’ movement across settings, and diversities in structures of practices that shape opportunities to learn and become. keywords: social practice theory; learning; agency; equity mailto:william.penuel@colorado.edu http://dx.doi.org/10.14786/flr.v4i4.205 penuel et al | f l r 31 1. introduction education systems around the world are experimenting with different ways to prepare youth to participate in, and lead, the global “knowledge economy.” in the united states, where we live, this has meant a push for higher quality science, technology, engineering, and math (stem) learning opportunities, more stem graduates from universities, and greater access to stem learning for groups that have been historically excluded from quality schooling. although the equity-related goals of expanding stem access are laudable, we worry that too many stem policies and educational interventions rest on flawed, inaccurate understandings of what it means to develop and sustain a learning pathway. put briefly, where curriculum developers privilege narrowly defined maps of cognitive learning progressions required for deep understanding, we propose a more expansive view of pathways that embeds learning and development in the broader context of social practice. ours is a social, materialist conception of pathways that has roots within the danish school of critical psychology (dreier, 1997; mørck & huniche, 2006; nissen, 2005). learning pathways unfold over time across multiple settings. the temporal dimension of learning is always situated within the evolution of broader social practices and institutions. as such, learning is part of persons’ “changing participation in changing practices” (lave, 1996, p. 150). the spatial dimension involves a form of movement of actors within and across the different social contexts of their lives (gutiérrez, 2008). as practices evolve and spaces are reorganized, people pursue diverse and evolving concerns and imagine new possibilities for themselves, thereby becoming different people by gathering, adopting, and pursuing stances toward what they do and value (dreier, 2008). at the same time, political borders and historically rooted but persistent inequities shape who people become and where they can move; for many, these constraints are experienced as disruptions and threats to their well-being and social worlds that must be actively managed (gonzales, 2011; mørck, 2010). historically and still today, research on learning has proceeded from a different set of assumptions about the nature of learning pathways. for example, many studies of learning in neuroscience and psychology focus on brief experiments that take place in a single setting (e.g., mcdaniel, agarwal, huelser, mcdermott, & roediger iii, 2011). in education, where there is more research on complex classroom processes that unfold over time, few learning researchers have studied how these processes are embedded within changing social and political structures of schools and societies (lave & mcdermott, 2002; penuel, 2016). there is limited attention to how inequities are reproduced within educational systems through the production of failure, that is, by making educational success into a scarce resource and disproportionately identifying individuals from disenfranchised groups “disabled” or “failures” who should be held to account for their own plight (lave, 1996; varenne & mcdermott, 1998). in this conceptual paper, we argue for social practice theory to guide the study of learning pathways across setting and time, presenting a case study analysis that illustrates how it might be applied to the study of learning over time and across settings. social practice theory (dreier, 1999, 2008; holland & lave, 2009; lave, 2012) emphasizes the importance of tracing pathways of participation across varied contexts and over time. we elaborate this theory in the context of a longitudinal study of students’ interest-related learning. we argue that analyzing the visibility, continuities, and discontinuities of learning pathways from youth’s own point of view provides a valuable framework for diagnosing inequity in access to participation in valued social practices. the assumption is that structures of practice present both constraints and possibilities for action to persons; the task of the analyst is to identify specific conditions that are relevant to the lives of persons in practice, what those mean to persons and practices, and the reasons for persons’ actions (jefferson & huniche, 2009; mørck & huniche, 2006). penuel et al | f l r 32 2. studying learning pathways in social practices social practice theory begins with the premise that people participate in multiple and variable social contexts. participation is neither constant nor bound to a particular place: people “participate for longer or shorter stretches of time, on a regular or occasional basis and for various reasons in several contexts” (dreier, 2008, p. 38). the name pathway captures the ideas that movement and direction are both important aspects of the temporal and spatial dimensions of participation in practice (dreier), though social practice theory views these pathways in concrete, social, and material terms and not as idealized progressions (penuel, 2016). similarly, a learning pathway can be described in terms of how people move across borders between the familiar and the unfamiliar (gutiérrez, 2008) and in terms of a telos, that is, the direction of movement or change (lave, 1996). from the perspective of social practice theory, the movement and direction of learning is not chosen in isolation from available social and institutional structures, but in relation to them. as dreier (2008) writes, people must “take into account the structuring of social practice into particular contexts with particular links and use those links in directing their trajectories [pathways] across them” (p. 38). from the standpoint of social practice theory, the complexity and diversity of practices in which people participate is not necessarily a burden but is an enriching aspect of life. by moving across settings of social practice, people are able to pursue diverse concerns and become aware of new possibilities for action and arrangements for participation in practice (dreier, 2008). in addition, they are confronted with dilemmas and contradictions that motivate change and learning (engeström & sannino, 2010; mørck, 2010). people learn by adjusting their contributions to activities to one another (o'connor & allen, 2010) and to fit the demands and structures of local institutions (dreier, 2009). people also learn by inventing new ways to participate in practice, molding it into new cultural forms through our participation (calabrese barton & tan, 2009; gutiérrez, baquedano-lopez, & tejada, 2000). existing institutional structures of practice frame the choices people make about how and where to participate in activities. directing their learning pathways requires that people distribute their engagement across different settings, according to the suitability of each setting’s institutional arrangements for pursuing a particular concern and how the settings are linked to valued practices in other settings (dreier, 2008). these institutional arrangements themselves vary with respect to roles and possibilities for action, requirements for access to those roles, and persistent patterns of privilege, exclusion, and marginalization (lave & mcdermott, 2002). any one institutional setting for participation, then, is both a place for learning and a connection point in a more spatially and temporally extended pathway. for our analysis, then, we used social practice theory because it allowed us to make sense of youth’s suspended participation alongside attending carefully to youth’s oft-ambiguous relationship to contentious structures of practice and histories (holland & lave, 2009). social practice theory also demands that we pay “particular attention to differences among participants, and to the ongoing struggles that develop across activities around those differences” (holland & lave, 2009, p. 5). in this respect, a social practice account differs in emphasis from what might emerge from an analysis of activity systems from a cultural-historical activity theory (chat) perspective, because the unit of analysis is persons in practice. in chat, the unit of both analysis and intervention is the activity system (cole & engeström, 2006). at the same time, social practice theory has common roots with chat, in that both draw from accounts of human activity developed by marx and engels (marx & engels, 1848/1998). in addition, both share a commitment to collaborative engagement of researchers with practice to expand possibilities for learning and movement across contexts (gutiérrez, 2008; gutiérrez & vossoughi, 2010; jurow & shea, 2015; mørck, 2010).1 penuel et al | f l r 33 3. supporting and studying science learning pathways: the case of jerome a number of science education researchers have documented ways that youth are active participants in directing learning pathways in science in ways that are in close alignment with how dreier (2008) characterizes peoples’ efforts to distribute their engagements across settings (bricker & bell, 2014; crowley, barron, knutson, & martin, in press; polman & miller, 2010). as bell and colleagues (bell, tzou, bricker, & baines, 2012) write, “science learners need to figure out how to adapt their abilities, interests, and identities across a diverse set of locations on a routine basis as they attempt to accomplish their goals or respond to the interests of other social actors” (p. 270). they further argue that learners’ efforts are strongly shaped by competing value systems that valorize some practices over others, variations in supports available from guides, and others’ recognition (and misrecognition) of their racial and class identities and perceived abilities. as we illustrate through the case of jerome (a pseudonym), the particular pattern of engagements and stance toward science-related pursuits emerges from a kind of “dance of agency” (pickering, 1995) as he initiates movement toward science-related futures within a particular context that enables development toward some directions but not others, leading him to re-configure their engagements in relation to how that context responds to his initiative (see also carlone, 2004). 3.1. context for the case analysis for the duration of the study, jerome was a participant in a program called pathways into science (also a pseudonym) at a science museum in a large city in the u.s. west. the museum’s program is like many others in large science museums across the country that provide learning opportunities to youth from groups that are underrepresented in science. it is funded by a variety of private foundations and individual donors. youth participants must attend public schools in the city where the museum is located, be enrolled as a ninth or tenth grader, and commit to meeting the attendance requirements year-round and over multiple years for which the youth are eligible. the program is run as a paid internship in which youth serve as docents for the museum visitors and have opportunities to contribute to science investigations led by resident scientists. jerome fit the profile well of the type of youth the program seeks to serve. he had a strong interest in science and a strong work ethic that made him well suited to meet the heavy time commitments of the program. his mother was an immigrant to the united states, and during our interview with him, he identified himself as black, which he asserts made him stand out among his peers and at the museum as different from most other people. at the time of the interview he had just completed his junior year of high school. 3.2. approach to case development jerome was part of a larger study our team has been conducting for the past three years of youths’ experience of connected learning, an emerging, synthetic model of interest-related learning being investigated by a network of scholars in a range of settings (ito et al., 2013). as we defined it for this study, youth were engaged in interest-related learning if they could identify an activity they enjoyed doing, pursued it over a long period of time, and believed they were getting better at the activity over time or learning from it. the approach we took to analyzing youth learning was a case study approach, in which we focused on an interest-related pursuit as youth’s participation in it transformed over time and as they moved across settings. we drew on interview data we collected from a larger study of 54 youth who were aged 13-17, and through a process of collaborative data organization, coding, representation, and analysis, aimed to illustrate the utility of taking a social practice approach to the study of learning as movement across settings. our interview protocol included questions that elicited youth’s descriptions of their activities and purposes for participation, their current involvement in their activity, the networks (e.g. linkages and penuel et al | f l r 34 supports) they drew upon when participating in their activity, obstacles they experienced, and how they perceived the future as related to their participation. informed by our theoretical orientation to learning as movement, the interviews purposefully elicited youths’ perspectives on how their participation changed over time and across different settings. because a social practice perspective encouraged a focus on youth agency in distributing their engagements, we focused on how youth themselves viewed their actions, as well as the continuities and discontinuities they experienced as they moved across the settings they traverse (see also akkerman and bakker, 2011). to analyze the data, we began deductively with a set of high-level codes related to the broad outlines of social practice theory, attending specifically to how engagement in the activity transformed over time and across settings. parent codes such as ‘linkages/supports,’ ‘barriers,’ ‘possible futures,’ and ‘identities/roles’ allowed us to get a sense of how youth characterized the many disruptions and opportunities that became more or less relevant to their lives—a central aspect of a social practice approach to analysis. the coding was used as a first step in identifying themes related to how individuals distributed their engagements across multiple settings and over time. in the analysis for this conceptual article, we focused on themes derived from codes linked to youths’ experiences during their initial engagement with the activity and subsequent engagements with the activity, namely ‘initial relationship to engagement’ as it co-occurred with the child codes within ‘reason/stance/relationship to participation’ such as ‘career,’ ‘friends,’ ‘academics,’ ‘civic engagement,’ and/or ‘skill building and mastery.’ the focus on these codes was reflective of our desire to make sense of youths’ relationship to varied structures of practice, as well as their different stances toward interest-related pursuits. we then created data displays for each student that listed their descriptions of initial activity, history of involvement in activity, youth articulations of participation within and outside of the program, their future goals, and immediate outcomes, noting also the different settings in which activities took place. our analysis provided strong supporting evidence that rather than linear or straightforward, youths’ pathways were often characterized by a shifting and fluid distribution of engagement in a variety of settings over time. for this paper, we chose a case (jerome) that we concluded illustrated well the agency of young people in this process. 3.3. jerome's initial goals for participation in the pathways program jerome’s story is not one of a simple pipeline into a science career that is institutionally supported and easily navigable for him. he actively worked to set up and select his involvement in science-related programs, and he benefited from access to institutions that helped him to link engagements over time between contexts. before he first started at the science museum, he worked with another program that connected him with numerous different programs around the city that matched with his interest in science. jerome applied to many of these programs in attempt to find a place where he could explore his interest in science and be exposed to different kinds of sciences. he also decided against another program because it was too familiar, stating, “that’s a lot of people from my neighborhood, so i know them and i always like to get to know other people and see what different places and cities are like.” consistent with jerome’s goal to expand his social network, at the beginning he experienced discontinuities between his community and the museum. he said, “it was different because i wasn’t too accustomed to being around people who weren’t exactly from my neighborhood or my ethnicity.” jerome experienced this new context as discontinuous with his prior context but complementary to his goals, which he linked to a sense of racial difference within the program. over time, his sense of difference from others faded as he deepened relationships with others through retreats and a leveled role structure (e.g., youth at level one are mentored by youth at level two). as he has moved up this structure, jerome has taken on leadership roles in the program, joining the leadership council, giving him “more of a say in the program’s direction and working with others on how to improve the program.” penuel et al | f l r 35 3.4. emergent patterns of engagement an stances a big part of jerome’s learning within the paid intern program centered on learning how to contribute to the ongoing activities of the museum. interns’ teaching was organized around “stations” connected to natural history exhibits at the museum. he had to work three days a week on the floor of the museum, which he described as “nerve-wracking” since he disliked public speaking. but he said the push helped him build his public speaking skills, and because near-peers in the program—interns one level up in the program from him—pushed him to contribute: when we first started, they would try to encourage you to—you’re at a station and you’re talking about whatever your station is about and the other intern, they are more experienced, so they know what to talk about and to push you out there they will say, “oh, michael can tell you more about it.” he also got to take part in science investigations at the museum, where he gained a partial view of the ongoing work of scientists there. he came to realize through these experiences that he really liked working in labs with other people and coming up with a question that “you really never know what it is going to be in the end.” at the same time, he developed a stance toward science that pointed away from a career in basic science. “i couldn’t be a botanist,” he said, and then generalized to all scientists, saying, “i see what they do here and confined to the basement or downstairs. most scientists, they work alone and i don’t really like to work alone, but i would want to do it as volunteer work or helping out whatever.” though he had opportunities to work with scientists in the laboratory, it is not clear whether he has not accompanied them to places where he might gain a fuller view of range of activities that characterize scientific practice. jerome did anticipate doing something science-related for his future career and continuing to volunteer in a science museum as an adult, though these pathways are not visible or directly accessible to him, either through the program or through his social network. he expressed a mix of worry and confidence about his direction: i’ve never been exposed to the steps that it takes to be a doctor or physician or whatever. so i don’t know, is the workload going to be like crazy, but i’ve never experienced something that was like impossible. so i guess it’s possible. in addition, his map of colleges and universities was relatively incomplete, focused on highly competitive and well-known national universities like harvard and stanford on one end and a less competitive local public university. he believed that “it doesn’t matter which school you go to, it just matters how good you do at that school,” and that it’s not important to get caught up in the reputation of the school. 3.5. support from outside the program some of jerome’s confidence likely arose from the fact that jerome is strongly supported by his family members. jerome described his mother, grandmother, and older brother as all highly supportive of his involvement. his mother was especially supportive of jerome’s decision not to take the “sports route” his peers seemed to be on and that he saw was highly valued. she and other family members encouraged him to go his own way: they [older brother, grandmother] just tell me to do what i feel is best and are supportive. i was on the flyer for the program the first year i was in here and my mom made like fifty copies and sent them out. my grandma has one hanging in her room and she mailed it to other family members. they’re really happy and supportive in that kind of way, just happy for me in general. though jerome was happy that the recognition of his accomplishments followed him home, his personal stance toward sharing his accomplishments with others was more complex. he preferred not to talk penuel et al | f l r 36 about them, because he saw it as too self-centered. he said, “it’s weird to talk about yourself. i just do it and move on.” at the conclusion of our interview with him, we asked jerome to characterize his identity—is he a scientist, an intern, or something else? he easily characterized himself as a “student,” saying “i consider myself a student, i just feel like i’m constantly learning, not just in school, but in everyday life. in the city and just being alive…” jerome connected his identity as a student or an everyday learner to his interaction with people. he expressed confidence in his ability, and also a compulsion to know things, which he said makes him an especially good fit to the program he was in. 4. discussion and conclusion our analysis illustrates one way to use social practice theory to study transformations in youth participation in an interest-related pursuit over time and across settings. in jerome’s case, attending to the different stances toward participation in activity helped illuminate how he was thinking about his future. there was evidence of both continuity and discontinuity between his experience in the museum and his imagined future. as such, his interest in science may be more akin to a “line of practice” of the kind azevedo (2011, 2013) describes, in which a “line” or pathway can be discerned, though participation changes significantly over time. jerome’s opportunity to participate in and move across varied settings was central to his developing sense of future possibilities. social practice theory helped us to see jerome’s distribution of his engagements as central to his interest development. jerome highlighted opportunities to learn how to be a docent on the museum floor, as well as his participation in research as significant for his learning. he noted, for example, through his observation (“i see what they do here”) that most of the scientists worked alone, in basements, and contrasted that with his own enjoyment of collaboration, interaction, and helping others. although this may not be a completely accurate view of what is often highly collaborative work of scientists, this observation enabled jerome to channel his energies where he sees the greatest alignment with what he wants to be doing in his everyday work life. in its emphasis on how youth distribute engagement across different settings, social practice theory differs from traditional “pipeline” metaphors for pathways into stem, which tend to focus on content but not on movement across settings. this social practice analysis, which we have foregrounded in this paper, also shows that interaction with graduates of his program has already prepared him for some of the challenges that he may face in a few years. he learned from near-age peers about how challenging pre-med courses could be and the self-doubt that can creep in during that pre-med experience. jerome’s achievements to this point represent the best of what stem interventions seek to accomplish, in terms of broadening access to the field for someone with limited exposure to it, and cultivating deeper and more nuanced interest in the various pathways related to stem. he has a better sense of what he wants and what he doesn’t want, as well as the threats to achieving that goal. we contend that policy interventions that foreground these issues associated with linkages, networks, and practices represent a promising direction for the field. studying and supporting those interventions requires us, however, to move beyond traditional foci of learning research and look at learning through a social practice lens. we contend that using social practice theory in the analysis of youth’s learning is potentially transformative in this respect, for it highlights potential leverage points for transforming systems to enable broader participation in stem. penuel et al | f l r 37 keypoints a social practice theory offers a lens for interpreting interest-related learning pathways across settings and over time. a social practice account highlights learners’ agency as they pursue diverse concerns across a range of settings, as well as existing how structures of practice constrain agency and limit access to some settings. a social practice account provides a more expansive framework for understanding how and when persons and practices are mutually constituted in such a way as to broaden access to stem fields. acknowledgments funding for this research comes from the john d. and catherine macarthur foundation. any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the funder. references bell, p., tzou, c., bricker, l. a., & baines, a. d. (2012). learning in diversities of structures of social practice: accounting for how, why, and where people learn science. human development, 55, 269-284. doi:10.1159/000345315. bricker, l. a., & bell, p. (2014). “what comes to mind when you think of science? the perfumery!”: documenting science-related cultural learning pathways across contexts and timescales. journal of research in science teaching, 51(3), 260-285. doi: 10.1002/tea.21134. calabrese barton, a., & tan, e. (2009). funds of knowledge, discourses and hybrid space. journal of research in science teaching, 46(1), 50-73. doi: 10.1002/tea.20269. carlone, h. b. (2004). the cultural production of science in reform-based physics: girls' access, participation, and resistance. journal of research in science teaching, 41(4), 392-414. doi: 10.1002/tea.20006. cole, m., & engeström, y. (2006). cultural-historical approaches to designing for development. in j. valsiner & a. rosa (eds.), the cambridge handbook on sociocultural psychology (pp. 484-507). new york: cambridge university press. crowley, k., barron, b. j. s., knutson, k., & martin, c. k. (in press). interest and the development of pathways to science. in k. a. renninger, m. nieswandt, & s. hidi (eds.), interest in mathematics and science learning and related activity. washington, dc: american educational research association. dreier, o. (1997). subjectivity and social practice. aarhus, denmark: center for health, humanity, and culture. dreier, o. (1999). personal trajectories of participation across contexts of social practice. outlines: critical social studies, 1(1), 5-32. dreier, o. (2008). psychotherapy in everyday life. new york: cambridge university press. dreier, o. (2009). persons in structures of social practice. theory & psychology, 19(2), 193-212. doi: 10.1177/0959354309103539. engeström, y., & sannino, a. (2010). studies of expansive learning: foundations, findings and future challenges. educational research review, 5, 1-24. doi: 10.1016/j.edurev.2009.12.002. gonzales, r. g. (2011). learning to be illegal: undocumented youth and shifting legal contexts in the transition to adulthood. american sociological review, 76(4), 602-619. doi: 10.1177/0003122411411901. penuel et al | f l r 38 gutiérrez, k. d. (2008). developing sociocritical literacy in the third space. reading research quarterly, 43(2), 148-164. doi: 10.1598/rrq.43.2.3. gutiérrez, k. d., baquedano-lopez, p., & tejada, c. (2000). rethinking diversity: hybridity and hybrid language practices in the third space. mind, culture, and activity, 6(4), 286-303. doi: 10.1080/10749039909524733. gutiérrez, k. d., & vossoughi, s. (2010). lifting off the ground to return anew: mediated praxis, transformative learning, and social design experiments. journal of teacher education, 61(1-2), 100117. doi: 10.1177/0022487109347877. holland, d., & lave, j. (2009). social practice theory and the historical production of persons. actio: an international journal of human activity theory (2), 1-15. ito, m., gutiérrez, k. d., livingstone, s., penuel, w. r., rhodes, j. e., salen, k., schor, j., sefton-green, j., & watkins, s. c. (2013). connected learning: an agenda for research and design. irvine, ca: digital media and learning research hub. jefferson, a. m., & huniche, l. (2009). re(searching) for persons in practice: field-based methods for critical psychological practice research. qualitative research in psychology, 6(1-2), 12-27. doi: 10.1080/14780880902896507. jurow, a. s., & shea, m. (2015). learning in equity-oriented scale-making projects. journal of the learning sciences, 24(2), 287-307. doi: 10.1080/10508406.2015.1004677. lave, j. (1996). teaching, as learning, in practice. mind, culture, and activity, 3(3), 149-164. doi: 10.1207/s15327884mca0303_2. lave, j. (2012). changing practice. mind, culture, and activity, 19(2), 156-171. doi: 10.1080/10749039.2012.666317. lave, j., & mcdermott, r. p. (2002). estranged labor learning. outlines, 1, 19-48. marx, k., & engels, f. (1848/1998). the german ideology, including theses on feuerbach and introduction to the critique of political economy. amherst, ny: prometheus books. mcdaniel, m. a., agarwal, p. k., huelser, b. j., mcdermott, k. b., & roediger iii, h. l. (2011). testenhanced learning in a middle school science classroom: the effects of quiz frequency and placement. journal of educational psychology, 103(2), 399. doi: 10.1037/a0021782. mørck, l. l. (2010). expansive learning as production of community. in learning research as a human science. yearbook of the national society for the study of education (vol. 109, pp. 176-191). new york, ny: teachers college record. mørck, l. l., & huniche, l. (2006). critical psychology in a danish context. annual review of critical psychology, 5. nissen, m. (2005). the subjectivity of participation: sketch of a theory. international journal of critical psychology, 15, 151-179. o'connor, k., & allen, a.-r. (2010). learning as the organizing of social futures. in learning research as a human science. national society for studies in education, 109(1), 160-175. penuel, w. r. (2016). studying science and engineering learning in practice. cultural studies of science education, 11(1), 89-104. doi: 10.1007/s11422-014-9632-x. pickering, a. (1995). the mangle of practice: time, agency, and science. chicago, il: university of chicago press. polman, j. l., & miller, d. (2010). changing stories: trajectories of identification among african american youth in a science outreach apprenticeship. american educational research journal, 47(4), 879-918. doi: 10.3102/0002831210367513. varenne, h., & mcdermott, r. p. (1998). successful failure: the school america builds. new york: westview press. vygotsky, l. s. (1934/1978). mind in society: the development of higher psychological processes. cambridge, ma: harvard university press. vygotsky, l. s. (1987). thought and language (a. kozulin, trans.). cambridge: cambridge university press. frontline learning research 6 (2014) 46-55 issn 2295-3159 corresponding author: susan goldman (susan.goldman@gmail.com) doi: http://dx.doi.org/10.14786/flr.v2i4.117 46 | f l r perspectives on learning: methodologies for exploring learning processes and outcomes susan r. goldman learning sciences research institute university of illinois, chicago, usa article received 27 may 2014 / accepted 2 december 2014 / available online 23 december 2014 abstract the papers in this special issue were initially prepared for an earli 2013 symposium that was designed to examine methodologies in use by researchers from two sister communities, learning and instruction and learning sciences. the four papers reflect a common ground in advances in conceptions of learning since the early days of the “cognitive revolution” in the 1960s. this commentary shows the interdependence between advances in theory and advances in methodologies. four shifts in conceptions of learning are described. that these shifts are evident in the work of both communities suggests a blurring of the boundaries between the two. keywords: learning, collaboration, mixed methods s. goldman 47 | f l r the papers in this special issue were initially prepared for an earli 2013 symposium that was designed to examine methodologies in use by researchers from two sister communities, learning and instruction and learning sciences. a main goal was to explore what these methodologies might reveal about underlying conceptions of learning and potential common ground across the two communities. indeed, the four papers, taken together, reflect a common ground in advances in conceptions of learning since the early days of the “cognitive revolution” in the 1960s. the papers depict highly systematic, thoughtful, and rigorous approaches to studying learning as it is happening whether individually, in small groups or large; in classrooms, in workplaces, or in labs; face to face or virtually; in one moment in time or over extended periods. they also illustrate the interdependence between advances in theory and advances in methodologies. during the 30 year period from 1960 – 1990, the majority of studies of learning took place in one location: the laboratory; looked at individual cognition as a function of an operationally defined and restricted set of variables, in one time frame with occasional return visits. although some researchers were examining learning in the context of tasks students might be asked to do in school, many of the “learning” situations were set up as experiments that were highly constrained to maintain experimental control over “extraneous” variables; as well the tasks were frequently “toy” problems that had little relevance to classrooms or other contexts outside of cognitive theory and the academic settings in which the research was being conducted. consequently, it was difficult to see how findings from the lab could possibly have relevance to everyday learning. emphases were on manipulating characteristics of the materials, the task, or both and observing how people “solved” the tasks and their success at doing so with some emphasis on understanding how they had completed the tasks. some of the findings emerging from that research shaped instructional studies in which people were instructed to summarize sections of text, underline main ideas, break down tasks into subtasks before solving them, group words based on taxonomic categories to improve memory, and similar heuristics. published studies of this sort attest to the success of these approaches in building a cognitive theory of learning and problem solving that superceded extant behaviourist/empiricist views (cf. greeno, collins, & resnick, 1996). there was, however, little uptake of these theories by educational practitioners. in addition to changes in the conceptualization of learning and the kinds of questions being asked about learning, data analytic methodologies have come a long way since the 1960s. some of the (older) readers of this article will remember the days of doing anovas by hand, with “advances” marked by programmable wang calculators, and main frame programs that automated the process. of course, you had to batch process jobs, submitting stacks of punch cards containing the data (all the time living in fear that you would drop the deck and have to start again making sure they were in the right order) and then wait for the print out. turnaround varied from 10 minutes to 24 hours. over the past 30 years there have been huge advances in the technologies and data analytic applications available on devices that are small enough to carry around the way we once carried pads of paper and notebooks (not electronic ones). these technologies have expanded the ways we dare to think about analyzing our data, enabled us to collect and make sense of new forms of data, and automated or semi-automated analysis methods that we used to do solely by hand. currently, the learning sciences and learning and instruction communities operate with theoretical frameworks on learning that reflect more complex views of learning in four major ways. we now understand and attempt to study learning 1. in multiple and iteratively designed environments 2. over multiple time scales 3. occurring in social groups of multiple and collaborating individuals s. goldman 48 | f l r 4. with effects evident at multiple levels ranging from behavioural to neural. as a set, the papers in this special issue reflect these shifts in conceptions of learning and its investigation and provide us with methodological tools that enable us to rigorously investigate learning processes and outcomes despite the greater complexity of doing so. there is evidence of one or more of these shifts in conceptions of learning among researchers who identify with learning sciences, as well as among those who identify with learning and instruction, suggesting something of a convergence, or at least a blurring of the boundaries of the two communities. 1. learning in multiple and iteratively designed environments design-based research marked a pivotal shift in perspectives on learning and its study in classrooms. svihla (this issue) provided an excellent description of the goals of this research approach. up until the time that this approach to research on learning was introduced in the early 90s by ann brown (1992) and allan collins (1992), educational research in schools typically took the form of relatively short-term experiments that involved comparisons of the effects of different methods or materials on various cognitive skills. the studies were usually conducted by the researchers. teachers “cooperated” with the researchers in terms of providing access to their students for the duration of the study but were otherwise minimally involved in contributing to the instructional design or the materials. a major goal of this research was ascertaining which instructional methods were better than others for achieving largely cognitive objectives such as more accurate mathematics performance, better memory for new vocabulary, and better comprehension of text. accordingly, assessments were designed to measure changes in students’ performance as a function of having participated in the study either in the “experimental treatment” or the “control” group. along with these types of cognitive studies there were similarly designed studies that examined the impact on individuals of having worked in cooperative groups (e.g., johnson & johnson, 1999; see for review webb & palincsar, 1996). as svihla described, the goals of dbr reflected a fundamental shift to an emphasis on studying learning processes in situ as both social and interactional (collins, 1992). learning processes were studied in the context of designed learning environments developed through collaborations of researchers and practitioners and based on principles that constituted a learning theory. enactments of designs were objects of study for purposes of understanding how, with the understandings that emerged from close study of, and reflection on, the interactions and student work informing iterative refinement of the learning theory principles, and designs. svihla does an excellent job of depicting the ways in which dbr has developed since its initial introduction. suffice to say it made apparent the need for methodologies to capture processes occurring over multiple time scales and among individuals in social configurations. 2. learning processes over multiple time scales complex views of learning make it clear that processes occur over time, with different learning processes occurring at different time scales. molenaar (this issue) provided an excellent rationale for the need for temporal analysis methods. she described a variety of the issues involved in shifting from a focus on whether a particular construct has been learned or not to a focus on how that construct is learned, what that learning looks like at different time scales, and indeed what constructs are conceptualized as emerging over longer versus shorter time frames (cf. lemke, 2000). she referenced a variety of constructs that we now s. goldman 49 | f l r think of as emerging over events and across time but that used to be thought of as personality traits (e.g., motivation, persistence). she discussed various computational tools that can aid in segmenting, coding, and relating different time scales. these methodologies are critical to doing the analyses needed to understand learning over different time scales. a variety of issues face us as individual researchers and as a community as we apply methodologies for temporal analysis: what units of time are appropriate for particular constructs of interest, especially when multiple time scales operate in parallel? how do we determine the time scale most appropriate for tracing the emergence of a construct over time? or alternatively, how do we capture the interrelationships between events occurring over time but at different time scales? of potentially many patterns of events that might be extracted by pattern detection software, how do we determine which are psychologically meaningful and at what scale of time they are meaningful? equally necessary are new forms of representation that can assist us in conveying our findings to the broader community. molenaar (this issue) presented one form of representation. figure 1 illustrates a different form of representation in which we plotted the discourse moves of three students comprising a small group engaged in a science investigation (radinsky, goldman, doherty, & ping, 2010). this particular figure shows the moves for the first day of the investigation. we plotted similar representations for each day and then used the graphs to identify regions where there were clusters of moves across the three students that suggested there were interesting dialogic discourses occurring. we then “dove” into these segments of the discourse to determine the character of the “argumentation” in which the students were engaged and whether the claims and evidence being offered were similar or different later in the investigation versus earlier. as well, we considered how participation and roles in the discourse revealed dimensions of identity and positioning with respect to disciplinary competence. this form of representation was a useful analytic tool and with some refinement might be a useful way to represent the time course of argument development (radinsky, et al., 2010). molenaar highlighted the need to conceptualize different dimensions of time in order to define important temporal characteristics. she cited papers by bloome, et al. (2009) and lemke (2000) as informing this discussion. in addition, bahktin’s (1981) framings of time in relation to discourse, meaning, and learning will be a useful resource. we also need to (re)connect learning and development. indeed, the move from the more traditional educational research paradigms to dbr and related learning sciences methodologies creates a convergence between research on development and research on learning. that is, one distinction between development and learning had traditionally been the time frame over which phenomena of interest emerged. those that occurred over multiple years were called developmental, e.g., oral language; those over minutes or hours, learning, e.g., declarative knowledge such as “c – a – t” spells cat; or propositions such as earli is a professional research organization. arguably, a second distinction was whether the phenomenon emerged with or without formal instruction, the latter being deemed developmental phenomena and the former learning. for example, children develop oral language but learn to read with explicit instruction. because developmental psychologists have long been concerned with the study of change over time, they have developed techniques that examine change over relatively longer periods of times such as growth analysis (willett, 1989), as well as techniques that look at moment to moment change, such as sequential analysis (bakeman & quera, 1995) and microgenetic analysis (e.g., kuhn, 1995; siegler & stern, 1998). these methods are rich resources for examining learning over multiple time scales. s. goldman 50 | f l r 3. learning in social groups of multiple and collaborating individuals a core assumption of the learning sciences is that learning is social and interactional and takes place through situated activity (brown, collins, & duguid, 1989). collins, brown, and newman (1989) labeled the approach cognitive apprenticeship, reflecting the importance of observing the habits of mind as well as the actions of the more knowledgeable others in the community (cf. vygotsky, 1978). hence, discourse about activity and interaction with others engaged in the activity became a central focus for understanding learning. researchers with intellectual roots in a variety of disciplines have long relied on discourse among participants in a joint activity as a window into knowledge building processes of groups as well as individuals, and, along with gestures, into processes of learning through joint activity (gee, 1992; goodwin, 1994; hutchins, 1995; lave & wenger, 1991; sawyer, 2006; scardamalia & bereiter, 1991; schegloff, 1991, 2007). video and audio recordings have typically provided the raw data and various types of very time-consuming and intensive qualitative analyses have been used by researchers to provide evidence for claims about learning outcomes and processes. frequently and understandably given the labor-intensive nature of these analyses, the evidence provided in any one empirical report has tended to be based on relatively small data sets or corpora. many computer-supported collaborative learning (cscl) environments make available written traces of learning interactions that can also be mined to understand learning processes and outcomes for individuals and for groups. although initially these were also analyzed by humans using processes similar to those used for coding discourse that was transcribed from video and audio recordings, a number of computer-assisted methods have been developed that make the work of coding less time-consuming. stegmann (this issue) argued that to understand the mechanisms that produced enhanced learning outcomes in cscl, three hypotheses needed to be tested. these have been conceptualized as a “triangle of hypotheses:” “(a) instructional/technological support facilitates learning activities; (b) facilitated learning activities have positive effects on learning outcomes; and (c) mediated by learning activities, instructional/technological support has a positive effect on learning outcomes.” (stegmann, this issue, p. #, citing wecker, stegmann & fischer, 2013; fig. 1). it could be argued that these three hypotheses can be thought of as constituting an activity system (engeström, 1987, 2001) in which tools (technology support), activities, and performances of individuals and groups exist in interaction with one another and over time. stegman and colleagues argue that conceptualizations other than experimental designs are needed to establish relationships between learning tools, activities, and outcomes. they propose the use of nomological nets to ensure that direct and mediating relationships between the tools and outcomes can be tested. nomological nets specify what constructs are indexed to which observables over what time frames. as well interrelationships among constructs are specified. empirical evidence derived from collaborative activities constitute input to revisions and refinements of theoretically grounded nomological nets. these revisions may reflect mediational variables that become evident through indepth analyses of the discourse, of changes in the interactions and discourse over time, and at different ”units” of analysis (e.g., individual, dyad, small groups, entire activity system). stegmann (this issue) argued that the indepth analyses required to “test” initially specified nomological nets should take advantage of statistical techniques designed to detect patterns of interactions as they occur over time. these techniques require some form of quantified information; therefore, qualitative analyses need to be quantified. fortunately, there are a number of computational algorithms that can assist researchers in doing so. importantly, these systems assist researchers in parsing the input as well as counting instances of particular codes and discovering repeating sequences of codes. the construct specification required by nomological nets is one way of ensuring that codes, sequences of codes, and s. goldman 51 | f l r recurring patterns relate to theoretically meaningful constructs. thus nomological nets can assist researchers in determining whether “discovered” patterns have psychological validity and practical utility. construct specification in nets can also assist with an additional “sticky wicket” in efficient yet automated detection of meaningful patterns of interactions. essentially, over what time frame are patterns of interaction to be detected and at what levels? that is, if a pattern is detected in a series of successive turns does that pattern then become a “unit” that can act as input to a subsequent pattern analysis effort? how are such patterns related to constructs in the net? one might envision a series of intermediate level patterns being inferred from turn by turn coding. these intermediate levels (essentially patterned sequences of turns) are the units upon which further pattern detection analyses are conducted. determining and optimizing appropriate time scales over which patterns of code sequences are constituted depends on understanding the intentions and assumptions of specific designed learning environments, particularly what and when specific processes are expected; why, and how they support expected outcomes. although patterns and sequences can be detected automatically it will take human interpretive lenses and socio-cognitive theories to specify the constructs these index and their meaningfulness in the context of learning. the issue of levels is relevant not only to pattern detection but to learning in general, as reflected in the fourth aspect of a more complex view of learning. 4. the effects of learning are evident at multiple levels ranging from behavioural to neural. learning is “visible” at different levels. de smedt (this issue) is to be applauded for emphasizing the need for alignment between the level and topical focus of research questions and the methods selected to investigate the questions. he pointed out that if the research question is targeted at the macrolevel, behavioral methods would be most appropriate. cognitive neuroscience methods become appropriate for research questions focused on microlevel processes. the two most common cognitive neuroscience methods are electroencephalography (eeg) and functional magnetic resonance imaging (fmri). eeg methods provide temporal information about when particular processes are taking place and fmri methods provide spatial information about where in the brain processes are taking place. cognitive theories provide needed links between behavioral and neural levels. furthermore, echoing the point made above – that socio cognitive theory needs to guide the interpretation of patterns in interactions, de smedt (this issue) called for detailed cognitive theory of learning phenomena to provide needed links between behavioral and neural levels. he cited the cognitive theories and the behavioral data on which they are based as critical for interpretations of the information that results from the application of cognitive neuroscience methods. to demonstrate his claims, de smedt (this issue) illustrated three ways in which cognitive neuroscience methods elucidate mathematical instruction. these are convincing demonstrations of the value added of obtaining data on the same phenomenon at multiple levels and coordinating findings across levels. predictions can be pursued across levels by postulating what should be the case at one level based on manifestations at another level. the value added of using multiple methods to examine learning at multiple levels is not restricted to mathematics. for example, in the area of language acquisition, researchers have used eeg methods to establish predictive relationships between phonemic and word-level development. specifically, infants below six months of age are sensitive to phonetic contrasts in all languages; between six and 10 months, a perceptual narrowing process occurs that results in sensitivity to only those phonetic contrasts that matter in their native language. eeg methods have established that better neural discrimination of native language phonetic contrasts is associated with faster vocabulary development (kuhl & rivera-gaxiola, 2008). kuhl s. goldman 52 | f l r and colleagues have also used neural activation patterns to determine that the perceptual narrowing process occurs several months later for infants reared in two-language homes compared to those reared in monolingual homes. finally, neural indicators of phoneme learning demonstrate that social interaction plays a critical role in language acquisition (kuhl, 2007). in each of these cases, evidence from the neural level provides measures of far greater precision than could be obtained behaviorally. 5. summary and challenges more complex views of learning require more complex methodologies for addressing key questions about learning and the conditions that support it, including explicit instruction. the four papers in this special issue illustrate methodologies that assist with capturing the iterative design-based research process, learning processes and outcomes that occur at different time scales and levels, and make possible the formulation and testing of hypotheses that relate different levels to one another. the papers present examples of ways in which these methodologies are augmenting the knowledge base for understanding learning as it occurs across individuals as well as within individuals. as such they make valuable contributions to the field.moving forward, there are a number of areas that need attention in terms of further theoretical and methodological development. briefly, more emphasis needs to be devoted to formative assessment that provides opportunities to better facilitate instructional processes and outcomes. this includes the design and testing of tools for capturing learning interactions that are classroom, teacher, and student friendly. such tools would enable students and teachers to reflect on their learning processes as well as outcomes at much finer levels of detail than is currently feasible. ideally, researchers would develop and test various technology-based tools for accomplishing these goals and would then engage in “user testing” of tools that travel outside research labs and into the hands of teachers and learners. a type of tool that would be helpful in this process is one that enables visualizations of the ebb and flow of learning processes across people and across time. finally, the learning sciences community has tended to design within specific disciplines and fields; the learning and instruction community has tended to test principles and variables thought of as general across all learning situations. neither perspective has as yet come to grips with the tension between generalist and discipline-specific views of learning nor the limitations of each view. what is needed are studies that 1) embrace a disciplinary perspective but that also situate that discipline in the context of epistemological orientations and inquiry methods that have been adopted and developed within other disciplinary communities; and 2) examine the “fit” of principles, constructs, and explanatory mechanisms suggested by cognitive, developmental, and social psychological research to learning phenomena observed in designed learning environments. studies of the first type will advance our understanding of the general and idiosyncratic aspects of learning in different disciplines. studies of the second type will advance our understanding of explanatory mechanisms that have traction across a wide versus narrow band of learners and situations of learning. there are also aspects of learning processes and outcomes that need far more systematic and sustained research over shorter and longer time scales. specifically, we need to conduct systematic research on relationships among persistence, engagement, identity, learning processes and outcomes, within and across formal and informal contexts of learning. some of this research is currently being conducted; more of it needs to be conducted. the methodologies discussed in these papers can be synergistic with respect to tackling these challenges. s. goldman 53 | f l r keypoints contemporary views of learning depart from solely “in the head” views of learning. learning occurs in multiple and iteratively designed environments over multiple time scales. learning occurs in social groups of multiple and collaborating individuals with effects evident at multiple levels ranging from behavioral to neural. new methodologies are needed to capture the processes and outcomes of this complex perspective on learning. acknowledgments the writing of this paper was supported, in part, by the institute of education sciences, u.s. department of education, through grant r305f100007 to university of illinois at chicago. the opinions expressed are those of the authors and do not represent views of the institute or the u.s. department of education. references bakhtin, m. m. (1981). forms of time and of the chronotope in the novel: notes toward a historical poetics. in m. m. bakhtin the dialogic imagination: four essays, trans. by caryl emerson and michael holquist (pp. 8 – 258). austin, tx: university of texas press. bakeman, r., & quera, v. (1995). analyzing interaction: sequential analysis with sdis and gseq. ny: cambridge university press. bloome, d., beierle, m., grigorenko, m., & goldman, s. r. (2009). learning over time: uses of intercontextuality, collective memories, and classroom chronotopes in the construction of learning opportunities in a ninth-grade language arts classroom. language and education, 23(4), pp. 313-334. brown, a. l. (1992). design experiments: theoretical and methodological challenges in creating complex interventions in classroom settings. the journal of the learning sciences, 2, 141-178. brown, j. s., collins, a, & duguid, p. (1989). situated cognition and the culture of learning. educational researcher, 18, 32 – 42. collins, a. (1992). toward a design science of education. in e. scanlon & t. o’shea (eds.), new directions in educational technology (pp. 15-22). berlin: springer-verlag. collins, a., brown, j. s., & newman, s. e. (1989). cognitive apprenticeship: teaching the craft of reading, writing, and mathematics. in l. b. resnick (ed.), knowing, learning, and instruction: essays in honor or robert glaser (pp. 453-494). hillsdale, nj: lawrence erlbaum associates. de smedt, b. (this issue). advances in the use of neuroscience methods in research on learning and instruction. frontline learning research. engeström, y. (1987). learning by expanding: an activity-theoretical approach to developmental research. helsinki: orienta-konsultit. engestrӧm, y. (2001). expansive learning at work: toward an activity theoretical reconceptualization. journal of education and work, 14(1), 133-156. fischer, f., kollar, i., stegmann, k. & wecker, c. (2013). towards a script theory of guidance in computersupported collaborative learning. educational psychologist, 49(1), 56-66. garcia-sierra, a., rivera-gaxiola, m., percaccio, c. r., barbara t. conboy, b. t. romo, h., klarman, l., ortiz, s., kuhl, p. k. (2011). bilingual language learning: an erp study relating early brain responses to speech, language input, and later word production. journal of phonetics, 39, 546-557. gee, j. p. (1992). the social mind: language, ideology, and social practice. ny: bergin & garvey. http://www.sciencedirect.com.proxy.cc.uic.edu/science/article/pii/s0095447011000660 http://www.sciencedirect.com.proxy.cc.uic.edu/science/article/pii/s0095447011000660 s. goldman 54 | f l r goodwin, c. (1994), professional vision. american anthropologist, 96, 606–633. greeno, j. g., collins, a. m., and resnick, l. b. (1996). cognition and learning. in d. c. berliner and r. c. calfee (eds.), handbook of educational psychology (pp. 15-46). new york: macmillan. johnson, d. w., & johnson, r. t. (1999). making cooperative learning work. theory into practice 38, 6773. hutchins, e. (1995). cognition in the wild. cambridge, ma: mit press. kuhl, p. k. (2007). is speech learning “gated” by the social brain? developmental science 10, 110–120. kuhl, p. k. & rivera-gaxiola, m. (2008). neural substrates of language acquisition. annual review of neuroscience, 34, 511 – 534. kuhn, d. (1995). microgenetic study of change: what has it told us? psychological science, 6, 133-139. lave, j., & wenger, e. (1991). situated learning: legitimate peripheral participation. cambridge, england: cambridge university press. lemke, j.l. (2000). across the scales of time: artifacts, activities, and meanings in ecosocial systems. mind, culture and activity, 7(4), 273–290. molenaar, i. (this issue). advances in temporal analysis in learning and instruction. frontline learning research. radinsky, j. l., goldman, s. r. doherty, r. & ping, r. (2010, june). small group argumentation with visual data: negotiating what is seen and what it means. paper presented at the annual meeting of the american educational research association, denver, co. sawyer, r. k. (2006). analyzing collaborative discourse. in r. k. sawyer (ed.), cambridge handbook of the learning sciences (pp. 187 – 204). ny: cambridge university press. scardamalia, m., & bereiter, c. (1991). higher levels of agency for children in knowledge building: a challenge for the design of new knowledge media. journal of the learning sciences, 1, 37 -68. schegloff, e. a., (1991). conversation analysis and socially shared cognition. in l. b. resnick, j. m. levine, & s. d. teasley (eds.), perspectives on socially shared cognition (pp. 150 – 171). washington, dc: american psychological association. schegloff, e. a., 2007. sequence organization in interaction: a primer in conversation analysis. cambridge: cambridge university press. siegler, r. s., & stern, e. (1998). conscious and unconscious strategy discoveries: a microgenetic analysis. journal of experimental psychology: general, 127, 377-397. stegmann, k. (this issue). advances in the analysis of computer-supported collaborative learning processes. frontline learning research svihla, v. (this issue). advances in design-based research. frontline learning research. vygotsky, l. s. (1978). mind in society: the development of higher psychological processes. cambridge: harvard university press. webb, n. m., & palincsar, a. s. (1996). group processes in the classroom. in d.c. berliner & r.c. calfee (eds.), handbook of educational psychology, (pp. 841-873). ny: macmillan library reference. wecker, c., stegmann, k., & fischer, f. (2012). lernund kooperationsprozesse: warum sind sie interessant und wie können sie analysiert werden? [learning and cooperation processes in casebased learning. interesting issues and analysis approaches] report: zeitschrift für weiterbildungsforschung, 35(3), 30-41. doi:10.3278/rep1203w willett, j. b. (1989). some results on reliability for the longitudinal measurement of change: implications for the design of studies of individual growth. educational and psychological measurement, 49, 587-602. s. goldman 55 | f l r figure 1. representation of the discourse moves of three students engaged in a science inquiry task. (cf. radinsky, j. l., goldman, s. r. doherty, r. & ping, r. (2010, june). small group argumentation with visual data: negotiating what is seen and what it means. paper presented at the annual meeting of the american educational research association, denver, co.) frontline learning research 3 (2014) 22-30 issn 2295-3159 corresponding author: associate professor gavin t l brown, school of learning, development, & professional practice, faculty of education, the university of auckland, private bag 92019, auckland, 1142, new zealand, gt.brown@auckland.ac.nz http://dx.doi.org/10.14786/flr.v2i1.24 22 | f l r the future of self-assessment in classroom practice: reframing selfassessment as a core competency gavin t. l. brown a , lois r. harris b a university of auckland, new zealand b central queensland university, australia article received 23rd october 2013 / revised 7th november 2013 / accepted 26th february 2014 / available online 25th april 2014 abstract formative assessment policies and self-regulation theories argue that student selfassessment of their own work and processes are useful for raising academic performance and self-regulatory skills. however, research into student self-evaluation raises serious doubts about the quality of self-assessment as an assessment process and identifies conditions which must be met if students’ judgments are to be useful, valid, and reliable. this paper recommends that student self-assessment should no longer be treated as an assessment, but instead as an essential competence for self-regulation. as such, we describe a potential curriculum approach that could guide teachers to appropriate use of self-assessment tools. keywords: student self-assessment; compulsory schooling; curriculum; research synthesis 1. introduction student self-assessment is an evaluation of a student‟s own work products and processes in classroom settings. formative assessment (a.k.a., assessment for learning) policies argue that student self-assessment is useful for raising academic performance (black & wiliam, 2006). research evidence suggests that selfassessment does contribute positively to learning outcomes, but its effects are highly variable, with many threats to its validity (brown & harris, 2013). nonetheless, student self-assessment is strongly advocated as an important classroom practice (e.g., leahy, lyon, thompson, & wiliam, 2005). this paper responds to recent and seminal reviews and position papers on self-assessment (e.g., andrade, 2010; brown & harris, 2013; boud & falchikov, 1989; butler, 2011; falchikov & boud, 1989; dochy, segers, & sluijsmans, 1999; g. t. l. brown & l. r. harris 23 | f l r dunning, heath, & suls, 2004; ross, 2006), all of which have raised issues to consider in relation to what is needed for the future of self-assessment. we are increasingly persuaded that self-assessment is not a robust assessment practice and that its real place in schooling is as a teachable and learnable component of selfregulated learning. however, current manifestations of self-assessment advocacy do not provide wellinformed guidance to researchers or practitioners about self-assessment. hence, our goal is to first establish the need for a self-assessment curriculum and second to sketch out what that curriculum could look like. 2. self-assessment as assessment assessment practices, which contribute to decision-making, need to be demonstrably valid and reliable (messick, 1989). the usefulness of self-assessment for decision-making seems to depend, in part, upon whether the student can accurately or realistically judge the qualities of their own work. however, the realism or veridicality (i.e., truthfulness) of self-assessment are difficult to ascertain, since this can only be determined through comparison to other people‟s (e.g., teachers, peers, or parents) judgements or ratings or to performance on externally devised tests or examinations. as butler (2011) makes clear, there has been a stream of research around student self-assessment that has emphasised the need for realistic, veridical, or verifiably accurate self-assessment if it is to effectively contribute to achievement (e.g., brown & harris, 2013; boud & falchikov, 1989; falchikov & boud, 1989; dunning, heath, & suls, 2004; ross, 2006). in contrast, there is another stream of research that has claimed that realism or veridicality in self-assessment is moot, since the self-assessment process helps students develop greater awareness of the quality of their work and criteria by which their work can be evaluated (e.g., andrade, 2010). butler (2011) concludes that the research in student self-assessment indicates that inaccurate, but positively biased, self-assessment leads to improved outcomes; while, inaccurate, negatively biased, self-assessment has a negative impact on achievement. hence, while there is empirical and theoretical evidence for contrasting positions around the realism of student self-assessments, we take the position that if self-assessment is to contribute to highlyconsequential decision-making (e.g., teacher decisions about grouping, curriculum planning, or retention/promotion and student decisions about pursuing or dropping further study in a topic area), then it is necessary for self-assessments to be demonstrably realistic or truthful. the research evidence is robust that the agreement between student self-assessment and other measures (e.g., test scores, teacher judgements, or peer ratings) is moderate at best (brown & harris, 2013; falchikov & boud, 1989). correlations between (a) self-ratings and teacher ratings, (b) self-estimates of performance and actual test scores, and (c) student and teacher rubric-based judgments tend to range from r≈.20 to .80, with few studies reporting correlations greater than r > .60 (brown & harris, 2013). greater realism and sophistication of self-assessment is more evident among more experienced and more able students. furthermore, consideration of how teachers use student self-assessment in classroom contexts suggests that there are other important factors which threaten reliability and validity. there is robust evidence that when self-assessments are disclosed (e.g., traffic light self-assessments displayed to the teacher in front of the class), there are strong psychological pressures on students that lead to dissembling and dishonesty (harris & brown, 2013; cowie, 2009). students may intentionally disguise the truth in order to protect their reputations. other students will rely on construct-irrelevant and subjective criteria (e.g., “i made an effort” or “i‟m good at this”), rather than intended criteria in judging the quality of their performance and this is associated with lower accuracy in self-evaluations (brown & harris, 2013). there are many factors in the human condition that contribute to unrealistic self-assessments, including tendencies to (a) be unrealistically optimistic about one‟s own abilities, (b) believe that one is above average, (c) neglect crucial information, and (d) have deficits in required information (dunning, heath, & suls, 2004). while substantial advocacy of self-assessment has been promulgated, these studies suggest that assessment for learning policies as implemented may have overlooked important facets of how humans make judgements. further, there may be some psychological and social danger to students in current self-assessment practices. hence, awarding grades or basing educational interventions or changes based on unrealistic or construct-irrelevant self-assessments is untenable. if self-assessment processes lead students to conclude wrongly that they are good or weak in some domain and they base personal decisions on such false interpretations, harm could be done, even in classroom settings (e.g., task avoidance, not enrolling in future g. t. l. brown & l. r. harris 24 | f l r subjects) (ramdass & zimmerman, 2008). the quality of interpretations and decisions depends on realistic input (messick, 1989) and because students are generally unrealistic in self-assessments, the use of such data in any formal way for assessment is probably unwarranted. hence, in many ways student self-assessment fails the validity and reliability requirements of an assessment robust enough on which to (a) base changes to classroom practice, (b) calculate grades or scores, and (c) include in any reporting. 3. self-assessment as self-regulation the use of self-assessment within assessment for learning policies draws on self-regulation of learning theories which identify student capabilities to set targets and evaluate progress against criteria as a basis for meta-cognitively informed improvement of learning outcomes (zimmerman, 2008). self-regulation refers to self-directive and self-generated metacognitive, motivational, and behavioural processes through which individuals transform personal abilities into control of outcomes in a variety of contexts (zimmerman, 2001). in taking action, an individual requires the ability to understand and choose a rationale for taking action (i.e., motivation), the capacity to set future-oriented objectives, plans, or projects (i.e., goals), the capability to select means or methods for obtaining his or her goals, even in the face of adversity or boredom (i.e., strategies), and the proficiency to monitor progress (i.e., self-assess), and adjust strategy implementation as appropriate (i.e., regulation) (brown, et al., 2005). there is evidence that students can improve their self-regulation skills through self-assessment (i.e., set targets, evaluate progress relative to target criteria, and improve the quality of their learning outcomes) (andrade, du, & wang, 2008; andrade, du, & mycek, 2010; brookhart, andolina, zuza, & furman, 2004). furthermore, self-assessment is associated with improved motivation, engagement, and efficacy (griffiths & davies, 1993; klenowski, 1995; munns & woodward, 2006; schunk, 1996), reducing dependence on the teacher (sadler, 1989). it is also seen as a potential way for teachers to reduce their own assessment workload, making students more responsible for tracking their progress and feedback provision (sadler & good, 2006; towler & broadfoot, 1992). thus, consistent with self-regulation theory, self-assessment contributes to greater meta-cognitive skills associated with greater achievement. in reviewing literature around student self-assessment practices, brown and harris (2013) found diverse enacted practices which they grouped into three major categories (i.e., self-estimation of performance, self-rating, and rubric based judgements). these three categories contain a wide variety of procedures; for example, (a) using a model answer as a reference (hewitt, 2001), (b) integrating teacherevaluation with self-evaluation (olina & sullivan, 2002), (c) self-correction (harward, allred, & sudweeks, 1994), (d) using a computerized prompt system (daiute & kruidenier, 1985), (e) self-selected reinforcements or rewards, especially for achieving challenging goals (barling, 1980; miller, duffy, & zane, 1993; wall, 1982), (f) contributing to the design of a scoring rubric (sadler & good, 2006), or (g) judging the accuracy of answers to standardized test items (koivula, hassmén, & hunt, 2001). in contrast, inspection of recent pedagogical texts (e.g., absolum, 2006; clarke, 2005; harlen, 2007; taylor & nolen, 2005; weeden, winter, & broadfoot, 2002; wiggins & mctighe, 1998) indicates that a relatively narrow range of selfassessment techniques are suggested (e.g., rubrics, rating scales, including traffic lights, reflections on portfolios or series of tasks). further, we find no evidence of coherent packaging or sequencing of these activities or a theoretical model underpinning authors‟ recommendations of these practices. there seems to be an almost ad hoc, grab bag approach to the use of self-assessment. the templates given as examples in the texts we consulted included mainly checklists, rating scales (sometimes using smiley-face categories), and lists of possible prompting questions which focused on diverse outcomes including completion, compliance, effort, and attitude (e.g., mcmillan, 2001; stiggins, 2005), rather than necessarily upon the full-range of self-regulating behaviours. also, textbooks seemed to treat the topic generically, without specifying ages or stages of development which would be appropriate for the practices, templates, or examples that they provided; it appears to be left to the reader to judge if a particular activity or prompt is appropriate for his or her students. inspection of classroom practices of self-assessment reinforces this perception that teachers are applying self-assessment techniques with little thought as to potential threats to the validity of publicly displayed selfassessment (harris & brown, 2013; ross, rolheiser, & hogaboam-gray, 1998) and, certainly, have few g. t. l. brown & l. r. harris 25 | f l r concerns about the need to provide students structured support to enable them to use self-assessment realistically. this stands in stark contrast to the research literature that shows significant educational impact of self-assessment upon student learning, if students are systematically taught how to self-assess (daiute & kruidenier, 1985; glaser, kessler, palm, & brunstein, 2010; harward, allred, & sudweeks, 1994; mcdonald & boud, 2003; ramdass & zimmerman, 2008; ross, 2006; ross, hogaboam-gray, & rolheiser, 2002). hence, there is a clear need to give self-assessment techniques a semblance of order—identifying the ease or difficulty of implementation and use would provide a robust basis for developing a curriculum of self-assessment. self-assessment is an essential component of self-regulation and would appear to be a learnable competence. the advantage of situating self-assessment as a competence is that competencies usually have levels of development (e.g., ranging from novice to expert) (rychen & salganik, 2003) and, consequently, can be used as the basis for a teaching curriculum. 4. self-assessment as a curricular competence while the assessment for learning policy reforms have tried to move increasingly away from formal testing towards a more pedagogical understanding of „assessment‟, and despite advocacy for the use of selfassessment as a component of self-regulation, little attention has been put into formalising a self-assessment curriculum, in light of well-established research findings. insufficient attention has been given to curricular concerns, such as: (a) what self-assessment skills should be taught? (b) what is the developmental sequence for teaching self-assessment skills? (c) how should self-assessment skills be taught? (d) what are appropriate goals for teaching student self-assessment competence according to student age and ability? (e) what are useful criteria for evaluating student competence in self-assessment? (f) what are appropriate mechanisms by which student self-assessment reports could be evaluated, if required? this paper offers a first attempt into developing a curriculum for self-assessment as a component of self-regulation. we appreciate that school curricula are overloaded and do not advocate for the creation of a new curricular topic. instead, we suggest that within more general frameworks of teaching subject content and developing students as independent, life-long learners, the possession and implementation of a selfassessment curriculum is likely to be of great utility to teachers and students. as self-assessment is present in some forms already in most curricula, we are advocating for systematically organising and formalising what is, in most instances, already there. this would improve its impact and better match practices to student selfassessment abilities, allowing students opportunities to master more complex self-assessment and selfregulatory skills as they progress through school. lower performing and younger students need input (i.e., instruction and feedback) to master this key self-regulatory process. so what might that input look like? research has shown (brown & harris, 2013; ross, 2006) that realistic self-assessments are more likely when: (1) students are involved in the process of establishing criteria for evaluating work outcomes; (2) students are taught how to apply those criteria; (3) students receive feedback from others (i.e., teachers and peers) to help move students toward more accurate evaluations; (4) students are taught how to use other assessment data (e.g., test scores or graded work) to improve their work; (5) there is psychological safety when self-evaluation is used; (6) when rewards for accuracy are used; and (7) when students are required to explicitly justify to their peers their self-evaluations. these insights give us a basis for developing a curriculum that could guide the implementation of student self-assessment as a necessary competence for self-regulation. our first recommendation for a student self-assessment curriculum is to start with simple, concrete techniques before introducing complex, abstract techniques, including holistic, intuitive judgements about g. t. l. brown & l. r. harris 26 | f l r effort, satisfaction, or work quality. for very young students, even the act of estimating how many times they would be able to throw a bean bag into basket was difficult (powel & gray, 1995). nonetheless, powel and gray‟s (1995) technique is extremely simple and the realism of a student self-assessment can be objectively verified by the student using a tangible metric. hence, estimating how many items one might get right on a spelling list, math quiz, or vocabulary quiz are straightforward strategies which allow easy determination of the realism of the student self-assessment (jones, trap, & cooper, 1977; wan-a-rom, 2010). linking such estimates very close in time to the instructional moment also makes the task more concrete (barnett & hixon, 1997). even asking students to estimate how well they think they will do compared to their last known performance provides a concrete and personal reference point. at an intermediate stage, self-assessments supported by externally-sourced, yet explicit scaffolding of intended learning outcomes (e.g., models, computer-assisted prompts, teacher evaluations) as an adjunct or guide should be introduced. more realistic comparisons of the student‟s work quality to that of other students in the class may be feasible here, but such normative comparisons may not be a desirable curricular goal. it seems more useful to have students focus on comparing their work to that of established standards or against their previous performance rather than on how others are doing. nonetheless, techniques that allow greater autonomy in self-assessment (e.g., self-correction or self-rating of one‟s own work) should be introduced once students have demonstrated that they can assess their work realistically. at an advanced stage, rubrics or criteria, preferably developed in conjunction with the students to ensure that they have a deep understanding of the rubric progression, should be introduced. by this point, students should be able to be reasonably realistic making use of more holistic, and possibly intuitive, judgements of their work quality using rating scales or key point checklists. the research shows that the greatest learning gains come when students engage in a deeper analysis of their own work; however, getting students to that level of analysis is unlikely to be instantaneous. gradual introduction of more sophisticated self-assessment techniques seems highly desirable. throughout the development of self-assessment competence, the priority needs to be kept on realism in self-evaluation, regardless of the level of performance. we must help students avoid inappropriate negative bias in their self-assessments, which will mean helping highly able students accept that their work is actually exemplary or of a high standard. in contrast, while having an overly positive self-assessment does not have as many ill-effects (butler, 2011), realism has its own internal benefits. accurate self-monitoring contributes to the possibility of entering a growth-pathway in which students identify and respond to their weaknesses, instead of pursuing an ego-protection pathway in which students seek to maximise unmerited positive feelings about their work (boekaerts & corno, 2005). hence, teachers need to implement strategies that encourage and foster honest self-reflection. that will mean, at least for a time, permitting some selfassessments to remain private from the teacher, not forcing students to display realistic but negative selfassessments in front of classmates, and encouraging students to share their self-evaluations with trusted people (e.g., a best friend or a family member). this recommendation is not new; andrade (2010) has long advocated focusing on the self-regulatory effects of self-assessment rather than its veridicality. and certainly, it means not using self-assessments for grading, reporting, or accountability purposes. nonetheless, students generally want to understand if they have judged their own work appropriately and expect teachers to provide feedback and instruction (harris & brown, 2013; gao, 2009; peterson & irving, 2008). thus, insulating student self-assessment perpetually from the teacher would be counterproductive. hence, within a context of psychological safety, if teachers gain access to student selfassessments (e.g., those recorded alongside homework activities handed in to the teacher), it seems desirable for teachers to comment on the realism of student self-evaluations as an important learning objective in its own right. the goal is to foster realistic self-monitoring that is used to guide appropriate learning strategies (i.e., more sophisticated responses than „work harder‟ are needed). students need environments in which realism is prioritised and protected, even if it means, at first, teachers cannot easily ascertain what students think about their own learning. a self-assessment curriculum should also encourage students to explain the criteria they used to evaluate their own work. the intellectual sophistication required to justify an assessment is a significant factor in improving learning outcomes and metacognitive capability. in an environment of g. t. l. brown & l. r. harris 27 | f l r trust (e.g., a classroom with a warm supportive interpersonal climate), explaining one‟s reasoning for a selfassessment to a trusted peer is associated with improved learning outcomes (dunning, heath, & suls, 2004). it should come as no surprise that both teachers and students will need training before they can engage with self-assessment as a taught and learned competence. new professional development materials and courses are needed that go beyond the exhortation to use student self-assessment (e.g., leahy, lyon, thompson, & wiliam, 2005). these resources need to ensure teachers are aware of the theory and research base for self-assessment and provide techniques that are appropriately sequenced for the skill level students have in this competence. until teachers abandon a simple approach (e.g., using smiley-face self-rating scales for effort and satisfaction), it is unlikely self-assessment will fulfil its promise. once teachers have an appropriate understanding, they will need to train students in developing realistic self-evaluations for the explicit purpose of guiding their own learning. fortunately, the research evidence makes it abundantly clear that the quality of student self-assessment improves with training and that enhanced outcomes arise. while this paper provides a rough outline for the scope and sequence of a self-assessment curriculum, more research is needed to identify if there are ages or stages, below which, particular types of self-assessment are unrealistic for students to complete accurately. additionally, a proper curriculum would incorporate all parts of the self-regulation cycle (zimmerman, 2008) around the self-assessment practices proposed at particular levels. it is also important to invent new self-assessment practices which may better align with a complex model of self-regulation than current suggested self-reflection and self-assessment practices, most of which are focused at the end, rather than throughout, the learning cycle. we trust that this treatment of self-assessment as a self-regulating competence, rather than as an assessment practice, will contribute to improved classroom practice and professional development systems. we consider that a curriculum for self-assessment competence would be of great benefit to educational practice and trust that this first sketch will trigger significant developments. keypoints student self-assessment generally has a positive impact on academic performance, although it is not a robust assessment method in terms of validity and reliability. student self-assessment is an important aspect of and contributor to greater self-regulation of learning. student self-assessment needs a curricular framework to ensure it is an effective treated as a self-regulating competence. acknowledgments this paper was inspired by a presentation at the 2013 biennial meeting of the european association for research in learning & instruction, munich, germany. references absolum, m. (2006). clarity in the classroom: using formative assessment to build learning-focused relationships. auckland, nz: hachette livre new zealand. andrade, h. l. (2010). students as the definitive source of formative assessment: academic self-assessment and the self-regulation of learning. in h. l. andrade & g. j. cizek (eds.), handbook of formative assessment (pp. 90-105). new york: routledge. g. t. l. brown & l. r. harris 28 | f l r andrade, h. l., du, y., & mycek, k. (2010). rubric-referenced self-assessment and middle school students' writing. assessment in education: principles, policy & practice, 17(2), 199-214. doi: 10.1080/09695941003696172 andrade, h. l., du, y., & wang, x. (2008). putting rubrics to the test: the effect of a model, criteria generation, and rubric-referenced self-assessment on elementary school students' writing. educational measurement: issues and practice, 27(2), 3-13. doi: 10.1111/j.1745-3992.2008.00118.x brown, g. t. l., & harris, l. r. (2013). student self-assessment. in j. h. mcmillan (ed.). the sage handbook of research on classroom assessment (pp. 367-393). thousand oaks, ca: sage. brown, g. t. l., reddish, p., leeson, h.v., milfont, t.l., brychkova, l. & hattie, j. a. c. (2005, june). assessment of key competencies: a proposal for inclusion in e-asttle. asttle advisory. rep. 19, university of auckland, project asttle. barling, j. (1980). a multistage multidependent variable assessment of children's self-regulation of academic performance. child behavior therapy, 2(2), 43-54. doi: 10.1300/j473v02n02_03 barnett, j. e., & hixon, j. e. (1997). effects of grade level and subject on student test score predictions. journal of educational research, 90(3), 170-174. doi: 10.1080/00220671.1997.10543773 black, p., & wiliam, d. (2006). developing a theory of formative assessment. in j. gardner (ed.), assessment and learning (pp. 81-100). london: sage. boekaerts, m, & corno, l. (2005). self-regulation in the classroom: a perspective on assessment and intervention. applied psychology: an international review, 54(2), 199-231. doi: 10.1111/j.14640597.2005.00205.x boud, d., & falchikov, n. (1989). quantitative studies of student self-assessment in higher education: a critical analysis of findings. higher education, 18, 529-549. doi: 10.1007/bf00138746 brookhart, s. m., andolina, m., zuza, m., & furman, r. (2004). minute math: an action research study of student self-assessment. educational studies in mathematics, 57(2), 213-227. doi: 10.1023/b:educ.0000049293.55249.d4 butler, r. (2011). are positive illusions about academic competence always adaptive, under all circumstances: new results and future directions. international journal of educational research, 50(4), 251-256. doi: 10.1016/j.ijer.2011.08.006 clarke, s. (2005). formative assessment in the secondary classroom. abingdon, uk: hodder murray. cowie, b. (2009). my teacher and my friends helped me learn: student perceptions and experiences of classroom assessment. in d. m. mcinerney, g. t. l. brown & g. a. d. liem (eds.), student perspectives on assessment: what students can tell us about assessment for learning (pp. 85-105). charlotte, nc: information age publishing. daiute, c., & kruidenier, j. (1985). a self-questioning strategy to increase young writers' revising processes. applied psycholinguistics, 6(3), 307-318. doi: 10.1017/s0142716400006226 dochy, f., segers, m., & sluijsmans, dominique. (1999). the use of self-, peerand co-assessment in higher education: a review. studies in higher education, 24(3), 331-350. doi: 10.1080/03075079912331379935 dunning, d., heath, c., & suls, j. m. (2004). flawed self-assessment: implications for health, education, and the workplace. psychological science in the public interest, 5(3), 69-106. doi: 10.1111/j.15291006.2004.00018.x falchikov, n., & boud, d. (1989). student self-assessment in higher education: a meta-analysis. review of educational research, 59(4), 395-430. doi: 10.3102/00346543059004395 gao, m. (2009). students' voices in school-based assessment of hong kong: a case study. in d. m. mcinerney, g. t. l. brown & g. a. d. liem (eds.), student perspectives on assessment: what students can tell us about assessment for learning (pp. 107-130). charlotte, nc: information age publishing. glaser, c., kessler, c., palm, d., & brunstein, j. c. (2010). förderung der schreibkompetenz bei viertklässlern: spezifische und gemeinsame effekte prozessund egebnisbezogener prozeduren der selbstgregulation auf indikatoren der schreibleistung, strategiebeherrschung und selbstbewertung [improving fourth graders' self-regulated writing skills: specialized and shared effects of processoriented and outcomerelated self-regulation procedures on students' task performance, strategy use, g. t. l. brown & l. r. harris 29 | f l r and self-evaluation]. zeitschrift fur padagogische psychologie/ german journal of educational psychology, 24(3-4), 177-190. doi: 10.1024/1010-0652/a000015 griffiths, m., & davies, c. (1993). learning to learn: action research from an equal opportunities perspective in a junior school. british educational research journal, 19(1), 43-58. doi: 10.1080/0141192930190104 harlen, w. (2007). assessment of learning. los angeles: sage. harris, l. r., & brown, g. t. l. (2013). opportunities and obstacles to consider when using peerand selfassessment to improve student learning: case studies into teachers' implementation. teaching and teacher education, 36, 101-111. doi: 10.1016/j.tate.2013.07.008 harward, s. v., allred, r. a., & sudweeks, r. r. (1994). the effectiveness of our self-corrected spelling test methods. reading psychology, 15(4), 245-271. doi: 10.1080/0270271940150403 hewitt, m. p. (2001). the effects of modeling, self-evaluation, and self-listening on junior high instrumentalists' music performance and practice attitude. journal of research in music education, 49(4), 307-322. doi: 10.2307/3345614 jones, j, c., trap, j., & cooper, j. o. (1977). technical report: students' self-recording of manuscript letter strokes. journal of applied behavior analysis, 10(3), 509-514. doi: 10.1901/jaba.1977.10-509 klenowski, v. (1995). student self-evaluation processes in student-centred teaching and learning contexts of australia and england. assessment in education: principles, policy & practice, 2(2), 145-163. doi: 10.1080/0969594950020203 koivula, n., hassmén, p., & hunt, d. p. (2001). performance on the swedish scholastic aptitude test: effects of self-assessment and gender. sex roles: a journal of research, 44(11), 629-645. doi: 10.1023/a:1012203412708 leahy, s., lyon, c., thompson, m., & wiliam, d. (2005). classroom assessment minute by minute, day by day. educational leadership, 63(3), 18-24. mcdonald, b., & boud, d. (2003). the impact of self-assessment on achievement: the effects of selfassessment training on performance in external examinations. assessment in education: principles, policy & practice, 10(2), 209-220. doi: 10.1080/0969594032000121289 mcmillan, j. h. (2001). classroom assessment: principles and practice for effective instruction (2nd ed.). boston, ma: allyn & bacon. messick, s. (1989). validity. in r. l. linn (ed.), educational measurement (3rd ed., pp. 13-103). old tappan, nj: macmillan. miller, t. l., duffy, s. e., & zane, t. (1993). improving the accuracy of self-corrected mathematics homework. journal of educational research, 86(3), 184-189. doi: 10.1080/00220671.1993.9941157 munns, g., & woodward, h. (2006). student engagement and student self-assessment: the real framework. assessment in education: principles, policy and practice, 13(2), 193-213. doi: 10.1080/09695940600703969 olina, z., & sullivan, h. j. (2002). effects of classroom evaluation strategies on student achievement and attitudes. educational technology, research and development, 50(3), 61-75. doi: 10.1007/bf02505025 peterson, e. r., & irving, s. e. (2008). secondary school students‟ conceptions of assessment and feedback. learning and instruction, 18(3), 238-250. doi: 10.1016/j.learninstruc.2007.05.001 powel, w. d., & gray, r. (1995). improving performance predictions by collaboration with peers and rewarding accuracy. child study journal, 25(2), 141-154. ramdass, d., & zimmerman, b. j. (2008). effects of self-correction strategy training on middle school students' self-efficacy, self-evaluation, and mathematics division learning. journal of advanced academics, 20(1), 18-41. doi: 10.4219/jaa-2008-869 ross, j. a. (2006). the reliability, validity, and utility of self-assessment. practical assessment research & evaluation, 11(10), available online: http://pareonline.net/getvn.asp?v=11&n=10. ross, j. a., hogaboam-gray, a., & rolheiser, c. (2002). student self-evaluation in grade 5-6 mathematics: effects on problem-solving achievement. educational assessment, 8(1), 43-59. doi: 10.1207/s15326977ea0801_03 g. t. l. brown & l. r. harris 30 | f l r ross, j. a., rolheiser, c., & hogaboam-gray, a. (1998). skills training versus action research in-service: impact on student attitudes to self-evaluation. teaching and teacher education, 14(5), 463-477. doi: 10.1016/s0742-051x(97)00054-1 rychen, d. s., & salganik, l. h. (2003). a holistic model of competence. in d. s. rychen & l. h. salganik (eds.). key competencies for a successful life and a well-functioning society (pp. 41-62). cambridge, ma: hogrefe & huber. sadler, p. m., & good, e. (2006). the impact of selfand peer-grading on student learning. educational assessment, 11(1), 1-31. doi: 10.1207/s15326977ea1101_1 sadler, r. (1989). formative assessment and the design of instructional systems. instructional science, 18, 119-144. doi: 10.1007/bf00117714 schunk, d. h. (1996). goal and self-evaluative influences during children's cognitive skill learning. american educational research journal, 33(2), 359-382. doi: 10.2307/1163289 stiggins, r. j. (2005). student-involved assessment for learning (4th ed.). upper saddle river, nj: pearson education. taylor, c. s., & nolen, s. b. (2005). classroom assessment: supporting teaching and learning in real classrooms. upper saddle river, nj: pearson education. towler, l., & broadfoot, p. (1992). self-assessment in the primary school. educational review, 44(2), 137151. doi: 10.1080/0013191920440203 wall, s. m. (1982). effects of systematic self-monitoring and self-reinforcement in children's management of test performances. journal of psychology, 111(1), 129-136. doi: 10.1080/00223980.1982.9923524 wan-a-rom, u. (2010). self-assessment of word knowledge with graded readers: a preliminary study. reading in a foreign language, 22(2), 323-338. weeden, p., winter, j., & broadfoot, p. (2002). assessment: what's in it for schools? london: routledgefalmer. wiggins, g. p., & mctighe, j. (1998). understanding by design. alexandria, va: association for supervision and curriculum development. zimmerman, b. j. (2001). theories of self-regulated learning and academic achievement: an overview and analysis. in b. j. zimmerman & d. h. schunk (eds.). self-regulated learning and academic achievement: theoretical perspectives (2 nd edn.). mahwah, nj: lea. zimmerman, b. j. (2008). investigating self-regulation and motivation: historical background, methodological developments, and future prospects. american educational research journal, 45(1), 166-183. doi: 10.3102/0002831207312909 frontline learning research 2 (2013) 86-98 issn 2295-3159 corresponding author: katrien vangrieken, university leuven: centre for research on professional learning & development and lifelong learning, dekenstraat 2, bus 3772, 3000 leuven, belgium, katrien.vangrieken@student.kuleuven.be http://dx.doi.org/10.14786/flr.v1i2.23 86 | f l r team entitativity and teacher teams in schools: towards a typology katrien vangrieken a , filip dochy a , elisabeth raes a , eva kyndt a a university of leuven, belgium article received 17 april 2014 / revised 16 december 2014 / accepted 18 december 2014 / available online 20 december 2014 abstract in this article we summarise research that discusses „teacher teams‟. the central questions guiding this study are „how is the term „teacher team‟ used and defined in previous research?‟ and „what types of teacher teams has previous research identified or explored?‟. we attempted to answer these questions by searching literature on teacher teams and comparing what these articles present as being teacher teams. we attempted to further grasp the concept of teacher teams by creating a typology for defining different types of teacher teams. overall, the literature pertaining to teacher teams appeared to be characterised by a considerable amount of haziness and teacher „teams‟ mostly do not seem to be proper „teams‟ when keeping the criteria of a team as defined by cohen and bailey (1997) in mind. the proposed typology, characterising the groups of teachers by their task, whether they are organised disciplinary or interdisciplinary, whether they are situated within or cross grades, their temporal duration and degree of team entitativity or „teamness‟, appears to be a useful framework to further clarify different sorts of teacher „teams‟. keywords: teams; teacher teams; typology; entitativity 1. introduction: beyond ‘egg-crate’ schools – teams in schools overall, teaming in schools appears to be quite a challenge, not the least because of a long-standing culture of teacher isolation and individualism in schools (gajda & koliba, 2008). teachers may feel that their autonomy is threatened by collaboration and that conflicts that they previously tended to avoid come to the surface (somech, 2008). teachers appear to be predominantly confined in their classroom in which they work in isolation, as such creating what lortie (1975, in westheimer, 2008) calls „egg-crate‟ schools. k. vangrieken et al. 87 | f l r despite this prevalent resistance to collaborate, a lot of studies point out to positive effects of a team structure in schools. teaming in schools appears to be a broad and rather vague concept with varying interpretations in the literature. nonetheless, it is of vital theoretical and practical importance to clarify this concept. in order to be able to properly discuss teacher teams it is essential to have a clear view on what such teams actually are and whether it is warrantable to speak of „teacher teams‟ in general or whether there are different types of teacher teams. cohen and bailey (1997) already pointed at the importance of team types in discussing their results. thus, the first aim of this article is to look at how the term „teacher teams‟ is defined in previous research and what type of teams were explored in former scientific inquiry. the importance of this article is shown in the fact that there might be several types of teams in schools and that these might possess different levels of „team entitativity‟ (the degree to which a „team‟ actually is a team, the „teamness‟ of teams). a clear typology thus could be useful in order to be able to draw warranted conclusions from former research that are applicable to a specific subset of teams since different types of teams may have different characteristics and thus different conclusions (for practice) may be justified. aside from the description of a few rather vague categories, previous research on teacher teams seems to lack a clear typological framework in order to clearly conceptualise the complexity of the concept of teacher teams. thus, the second aim of the article is to present a typology for using the team concept in schools. in the following will be discussed what „teacher teams‟ are and a search for clarification in the ruling conceptual confusion concerning this sort of teams will be undertaken. 2. defining teams among the large number of existing definitions of „teams‟, the one formulated by cohen and bailey (1997) seems to be the most comprehensive and mostly used in research on teamwork and team learning (e.g. dochy, gijbels, raes, & kyndt, 2014; decuyper, dochy, & van den bossche, 2010). these authors described a team as follows: “a team is a collection of individuals who are interdependent in their tasks, who share responsibility for outcomes, who see themselves and who are seen by others as an intact social entity embedded in one or more larger social systems (for example, business unit or the corporation), and who manage their relationships across organizational boundaries” (p. 241). teams thus have to meet these six criteria. teams are seen as different from groups and are mostly defined more narrowly. as such, van den bossche, gijselaers, segers, and kirschner (2006) stated that „a team is more than a group of people in the same space, physical or virtual‟ (p.490). salas, burke, and cannon-bowers (2001) argued that teams differ from groups in task interdependence, structure and time span. in this sense, all teams are groups when groups are seen as sharing a common social categorisation and identity (raes, kyndt, decuyper, van den bossche, & dochy, submitted). but not all groups are teams since a team has characteristics that not necessarily have to be present to define a group. 3. on the use of the term ‘team’ the articles primarily using the term „team‟ show a considerable amount of diversity in the interpretation of this term which leads to a lack of conceptual clarity. few studies clearly define what they mean when they speak of a „team‟ or a „teacher team‟. several authors appeared to use the term „team‟ without specifying what they mean with this term or who or what these so-called „teams‟ include. a vast amount of studies were more exploratory in nature and did not start off by giving a definition of „teams‟ but described the teams under study (e.g. gunn & king, 2003; hackmann, petzko, valentine, clark, nori, & lucas, 2002; meirink, imants, meijer, & verloop, 2010; somech, 2005). even authors that focused on other denominations often seemed to use the term „team‟ k. vangrieken et al. 88 | f l r somewhere in their article: some studies started off by writing about „collaboration‟, „community‟, „department‟ or „critical friends group‟ and then later on referred to „teams‟ (mostly as a form of collaboration) without giving further explanation (e.g. achinstein, 2002; avila de lima, 2001; curry, 2008; datnow, 2011; dickinson, 2009; kelchtermans, 2006; leonard & leonard, 2003; lomos, hofman, & bosker, 2011; scribner, hager, & warne, 2002; visscher & witziers, 2004; williams, 2010). this mix up of different terms is confirmed by westheimer (2008) who mentioned that schools use different denominations to describe collaboration between teachers, one of them being „teams‟. there thus appears to be a misconception among teachers regarding the term „team‟ since it can be doubted whether what teachers define as being a „team‟ matches any of the criteria for teams that are mentioned in the contemporary team literature (smith, 2009). smith (2009) furthermore concludes that in the perception of teachers, teamwork is depicted as mere collaboration between friends. amidst this inaudibility that accompanies the use of the term „team‟, a few differentiations can be made in the different interpretations and uses of the term. this points at the importance of creating a teacher team typology: in order to be able to make clear and justified conclusions concerning teacher teams it is essential to clarify to which sort of teacher team these pertain. 4.1 existing teacher team categories and team typologies 4.1.1 existing teacher team categories some other authors already pointed at the diversity in the types of teams existing in schools. as such supovitz (2002) stated that teams can be organised in different ways: for example same grade level or vertical (cross grades), teams can loop (teachers stay with the same students for several years), the members can stay in fixed grade levels, or teams can have mixed configurations. pounder (1998) mentioned management teams or school advisory groups, special services teams, and interdisciplinary instructional teams. park, henkin, and robert (2005) distinguished three comparable types of school teams: governance teams, instructional teams and planning teams. governance teams do not have an instructional task but usually develop policies to meet specific needs of local communities. principals, teachers, parents, and community members are the primary decision makers (ellis & fouts, 1994, in park et al., 2005). according to buckley (2000, in park et al., 2005) instructional teams serve to realise flexible scheduling of instruction and higher integration of subject matters. this sort of teams can be organised according to grade level or subject. planning teams are organised to tackle specific problems, which can be temporary or more complex and long term. drach-zahavy and somech (2002) mentioned the fact that teams in schools serve different purposes and distinguished between management teams, instruction teams and pedagogic teams. the management teams are involved with administrative issues and participate in the management. the instruction teams gather around a subject area and their ultimate goal is to enhance teaching effectiveness. finally, the authors state that pedagogic teams consist of teachers that teach in the same class and these teams are focused on improving the pedagogic decisions on specific pupils. the teams in the study of tonso, jung, and colombo (2006) could be sorted out into administrative teams, grade level teaching teams (which were further divided into mixed content subteams) and social service teams. smith (2009) focused on science teachers conceptions of teams and teamwork and listed eight possible teams that can emerge in a school setting. management teams are charged with administrative issues, pedagogic teams are based on teachers teaching the same class, instructional teams are based on subject matter and serve to foster teacher effectiveness while interdisciplinary teams gather teachers from different subject areas who collaborate in teaching and learning. appraisal teams provide assistance in making sense of problem situations. in informational teams the members merely exchange information that is needed to perform the teaching profession, instrumental teams provide practical support and emotional teams form a supportive network with encouraging words and sympathetic understanding. as might be clear, the last three types of teams could be seen as less of an actual „team‟ compared to the first group. k. vangrieken et al. 89 | f l r thus, the existing categories seem to focus primarily on the task of the teacher team to distinguish different types (drach-zahavy & somech, 2002; park et al., 2005; pounder, 1998; tonso et al., 2006). only supovitz (2002) explicitly focused on the organisational differences between the teams. this study attempts to expand the focus to other constructs than the task domain including task as well as organisational features. 4.1.2 team typologies cohen and bailey (1997), devine, clayton, philips, dunford and melner (1999) and hollenbeck, beersma and schouteden (2012) presented typologies of teams in general (not focused on teacher teams in specific). cohen and bailey (1997) distinguished four types of teams: work, parallel, project and management teams. devine et al. (1999) presented a dimensional approach to a team typology using the dimensions product type and temporal duration. the crossing of these two typologies results in four team types: ad hoc project teams, ongoing project teams, ad hoc production teams and ongoing production teams. the article of hollenbeck et al. (2012) searched to transcend different existing team typologies relying on a dimensional scaling approach based on three underlying constructs: skill differentiation, authority differentiation and temporal stability. 4.2 transcending the different existing categorisations: a typology several authors thus pointed out to the existing diversity in teacher teams and distinguished different „categories‟ of teams. these different categorisations overlap to some extent and the teams mentioned in the literature appear to fit into these categories to a certain degree. for that reason, the abovementioned existing categories, together with typologies of teams in general (not teacher teams in specific) (devine et al., 1999; hollenbeck et al., 2012; cohen & bailey, 1997), will serve as a starting point for the typology that is made here. they will be supplemented with other important categories and dimensions that play an important role in the literature discussing teacher teams. overall, following defining features of a teacher team typology (presented in appendix 1 table 1) appear to be important. 4.2.1 task first, teacher teams may have tasks pertaining to governance or management. pounder (1998) stated that management teams may include representative teachers, school support staff and parents or community members. their main responsibility is advising the principal or other administrators in problem solving, planning and decision-making concerning school improvement. according to ellis and fouts (1994, in park et al., 2005) governance teams develop policies to meet specific needs of local communities. principals, teachers, parents and community members are the primary decision makers. thus although teachers may be part of such teams, other representatives are often included as well. secondly, instruction appears to be a very important task for teacher teams: overall, teacher teams show a primary focus on instruction and student learning. here instruction is seen as tasks teams perform that are directly related to student instruction. in order to create further clarification, two subtasks are distinguished here: instruction/teaching and planning of instruction. instruction/teaching includes all tasks of teachers directly pertaining to the instruction of a particular group of students. this includes collaborating on the instruction, evaluation, and follow-up of a particular group of students. the subdivision of planning of instruction is seen as collaboratively planning instruction in general and is not necessarily limited to a common group of students. tasks here entail in general the planning, coordinating and evaluating of curriculum (flowers, mertens, & mulhall, 2000; mertens & flowers, 2004; gunn & king, 2003; yisrael, 2008). it may also include planning considering student assignment (flexible grouping strategies) and scheduling (conley, fauske, & pounder, 2004), which are needed before the instructional process can start. a third task, problem-specific planning, is inspired by the typology of park et al. (2005) who stated that planning teams are responsible for tackling specific problems and can be of a temporary or a longer k. vangrieken et al. 90 | f l r lasting nature. smith (2009) spoke of appraisal teams who offer assistance in making sense of problem situations and as such can be related to this task category. although teams in the articles under study have an array of different decision-making responsibilities and do tackle specific problems, these specific tasks are mainly coupled with a more general task such as instruction for example. this type of task is thus not that clearly delineated from the other tasks. fourthly, the task of teacher teams may pertain to pedagogy, as it is one of the team-types distinguished by drach-zahavy and somech (2002) and smith (2009). this task can be related to supporting student learning and managing student behaviour (e.g. crow & pounder, 2000; supovitz, 2002; watson, 2005), to communication with parents (e.g. crow & pounder, 2000; flowers et al., 2000) or more general to a discussion of the teaching and pedagogy and the challenges experienced by teachers (e.g. havnes, 2009). a following and related task of teacher teams may include special or social services. pounder (1998) stated that special services teams are responsible for the evaluation, placement, and educational plans of exceptional students. they may include special education teachers, professional support staff, administrators, representative parents, and others. the responsibilities of this type of team are not limited to pure educational tasks but stretch further into the social and psychological functioning of students. both types of teams (special and social services) can be seen as similar to some extent (the social service team in the study of tonso et al. (2006) included a special education teachers for example) and might be integrated into one team in some schools and their tasks are thus seen as belonging to the same task category. sixthly, teacher teams can have tasks related to innovation and school reform (meirink et al., 2010). quite often teams are being associated with school reform or innovation. euwema and van der waals (2007) pointed to the fact that the environment of schools is increasingly dynamic and complex. and this will lead to a decrease in the predictability of developments causing an increased pressure on the ability of the school to adapt and innovate. the authors pointed to these developments as an important, although not the only, reason for organising schoolwork in teams. meirink et al. (2010) and meirink (2007) spoke of temporary „innovative teams‟ that are responsible for designing and experimenting with new teaching practices. overall, enhancing teacher collaboration appears to be a rather „recent‟ innovative attempt in a few countries, organising teachers in teams is one of the ways to accomplish this goal. some studies, such as watson (2005), spoke of learning teams, in which the learning of teachers is of central importance. this can entail learning of teachers considering the teaching practice, as such watson (2005) stated that the professional learning teams, sometimes referred to as (professional) learning communities (saunders, goldenberg, & gallimore, 2009; dickinson, 2009; cheng & ko, 2009; williams, 2010) or communities of practice (curry, 2008), in his study are involved in the implementation of a school improvement process. this is closely linked to the category „innovation‟ and shows that the boundaries between the different task categories can be blurred. overall the (professional) learning teams discussed in the literature are directed towards improving student performance. this thus forms a bridge between the task of learning and of instruction: teachers need to learn in order to improve their instruction and thus enhance student learning. finally, some studies appear to mention a mere material or practical „task‟ when discussing teacher teams. for example, main and bryer (2005) pointed at the sharing of physical space and/or resources as being part of the „task‟ of teacher teams and smith (2009) referred to instrumental teams who provide practical support. this clearly forms a rather infirm base for teacher teams and a grouping of teachers showing a mere material or practical base upon which to collaborate can be merely considered as working in proper „teams‟. smith (2009) furthermore pointed at informational teams and emotional teams. in the first, members merely exchange information that they need in order to perform the teaching profession. the latter provides a supportive framework with encouraging words and sympathetic understanding. it becomes clear that these tasks on themselves as well are not enough to justifiably speak of an actual „team‟ and as such they are mentioned in this task category. k. vangrieken et al. 91 | f l r 4.2.2 discipline level: disciplinary or interdisciplinary a second important distinction to be made is whether teacher teams are organised disciplinary or interdisciplinary (teachers teaching the same or different subjects). this can be linked to the dimension of skill differentiation mentioned by hollenbeck, et al. (2012): this means that members have more or less specialised knowledge or functional capacities that make them more or less difficult to replace. as such, in interdisciplinary teams teachers have expertise in different subject areas. in some school contexts this distinction may be less relevant. for example, in primary or elementary schools teachers are responsible for teaching all courses and are thus not specialists in one or more disciplines. in such contexts it appears irrelevant to speak of disciplinary or interdisciplinary since every teacher is responsible for all disciplines to be taught. an exception here could be when a different teacher who is responsible for teaching crafts or music, a special education teacher,... is included in the team. when teachers are not the only team members, „interdisciplinary‟ refers to the fact that the team is comprised of people from different professions (e.g. nurses, social workers, specialists,...). 4.2.3 grade level: cross or within grade level another important distinction that can be made in teams of teachers is the fact whether they are situated on a grade level (responsible for students in the same grade level) or not (responsible for students cross grades). pounder (1998) states that a common middle school structure appears to consist of interdisciplinary grade-level teams. as such it should not come as a surprise that quite a lot of the studies (of those who clarify these characteristics) referred to such teams in middle schools. 4.2.4 temporal duration considering temporal duration (whether the teams are designed temporarily or for a longer time period), there are only two studies explicitly referring to temporary teams (meirink, 2007; meirink et al., 2010). most other studies seem to refer to teams that are more long-term (a temporal duration of the collaboration is not given), except for drach-zahavy and somech (2002) and somech (2005) who mentioned that the teams under study already worked together for at least one year. 3.2.5 team entitativity a final and vital feature of teacher teams is a dimension that is captured in the term of „entitativity‟ (campbell, 1958). this terms covers the fact whether an aggregate of persons actually behaves as a system. according to campbell, entitativity includes „the perception that a social aggregate is a coherent, unified and meaningful entity‟ (haslam, rotschild, & ernst, 2004, p.65). it entails the degree of being a unity or a coherent whole and thus represents the interdependence that is present in groups or teams (campbell, 1985). ohlsson (2013) also states that teams possessing a strong level of interdependence see themselves as an actual team. in this article, team entitativity is conceived as the degree to which a collection of individuals is an actual team as described by cohen and bailey (1997). the criteria in the definition of cohen and bailey (1997) will serve as a basic measure of team entitativity. these six criteria entail: a collection of individuals; who are interdependent in their tasks; share responsibilities for outcomes; see themselves and are seen by others as an intact social entity; embedded in one or more social systems; and manage their relationships across organisational boundaries (what will be referred to as boundary crossing). the more criteria the teams meet and the stronger they fulfill them, the higher their degree of team entitativity or „teamness‟ will be‟. the concept of team entitativity is further elaborated upon in the review article of vangrieken, dochy and raes (submitted). k. vangrieken et al. 92 | f l r 4. conclusion this short article tried to answer the questions: „how is the term „teacher teams‟ used and defined in previous research?‟ and „what types of teacher teams has previous research identified or explored?‟. this study results in the following: first, starting from a comprehensive definition of „teams‟ that provides clear criteria from which can assessed whether groups can rightfully be called „teams‟ (cohen & bailey, 1997), we find that teacher teams in literature in most cases do not meet these criteria or at least often no effort is made to make definitions and characteristics of groups of teachers sufficiently explicit. a clear-cut unambiguous definition of teams in schools or teacher teams appears to lack. different authors discussing teacher teams tend to use different interpretations of the term „team‟ and seem to discuss different types of teams. most of the articles lack an insightful definition of what they mean when they use the term „team‟ which makes interpreting the results of their research quite challenging. moreover, no single description of teacher teams met all criteria of a team as described by cohen and bailey (1997). so, „teachers groups‟ appear to be mostly „groups‟ instead of highly entitative „teams‟. this finding is in line with a conclusion made by smith (2009) who stated that teams as they are usually defined outside education are perceived as dysfunctional in the experiences of science teachers, they do not exist or do not work in schools. smith (2009) furthermore stated that although the teachers in the study experience membership of multiple teams, it can be questioned whether these so called teams really exist in the meaning of „teams‟ as described in the conventional team literature. in the latter, a team is presumed to be much more than a collection of individual teachers who are gathered around their timetabled subjects, staffroom or science department (smith, 2009). as a consequence, it would be interesting to find out what criteria are really met by the so-called teacher teams in literature. it would be reasonable to argue that some teacher groups discussed in literature are more a „team‟ than others in the sense that they meet more of the aforementioned criteria (what is previously referred to as team entitativity). at this point, it is difficult to assess the degree of team entitativity of teams described in literature based on the current vague information in most studies. future studies on teacher teams should go deeper into the real origin and scope of the teams. a typology for teacher teams can be based on the following axes: task (governance/management, instruction, problem-specific planning, pedagogy, special/social services, innovation/school reform, learning, material/practical), discipline level (disciplinary or interdisciplinary), grade level (within or cross gradelevel), temporal duration (temporary or lasting) and team entitativity (low, moderate or high). as a consequence, an overarching typology is proposed in figure 1. k. vangrieken et al. 93 | f l r figure 1. typology. there appears to be a vast amount of variation in „teacher teams‟, with a variety in the task and organisation of the teams. the above discussed framework appears to be useful in trying to clarify what these „teams‟ consist of. by giving a specification of all of these distinctions, which lacks in a lot of studies, a rather clear description can be given of what sort of teacher team is under study. keypoints there appears to be a lack of clarity and a large variety within the concept of teacher teams. they can have a large diversity of tasks and can be organised in different ways. the lack of a clear description makes it difficult to draw warranted conclusions since it may not be justified to make generalisations across different types of teacher teams. when using the definition of cohen and bailey (1997) of teams to assess whether „teacher teams‟ can rightfully be called „teams‟, we conclude that groupings of teachers are mostly „groups‟ rather than „teams‟ since „teams‟ are often defined quite vaguely in team literature and these descriptions hardly ever meet all criteria mentioned in this definition. a typology for teacher teams based on the axes of task, discipline level, grade level, temporal duration and team entitativity is a useful framework to describe the sort of teacher team under study. this thus creates some clarity among the indistinctness surrounding the use of the term „teacher team‟. k. vangrieken et al. 94 | f l r references achinstein, b. (2002). conflict amid community: the micropolitics of teacher collaboration. teacher college record, 104(3), 421-455. retrieved from http://www.tcrecord.org/ avila de lima, j. (2001). forgetting about friendship: using conflict in teacher communities as a catalyst for school change. journal of educational change, 2, 97-122. retrieved from http://link.springer.com bertrand, l., roberts, r.a., & buchanan, r. (2006). striving for success: teacher perspectives of a vertical team initiative. national forum of teacher education journal-electronic, 16(3). retrieved from http://www.allthingsplc.info brouwer, p. (2011). collaboration in teacher teams (doctoral dissertation). retrieved from http://dspace.library.uu.nl/handle/1874/214140 brouwer, p., brekelmans, m., nieuwenhuis, l., & simons, r.j. (2012). fostering teacher community development: a review of design principles and a case study of an innovative interdisciplinary team. learning environments research, 15(3), 319-344. doi:10.1007/s10984-012-9119-1 campbell, d.t. (1958). common fate, similarity, and other indices of the status of aggregates of persons as social entities. behavioral science, 3(1), 14-25. retrieved from http://librilinks.libis.be carroll, t.g., & foster, e. (2008). learning teams: creating what‟s next. national commission on teaching and america‟s future. retrieved from http://nctaf.org cheng, l.p., & ko, h. (2009). teacher-team development in a school-based professional development program. the mathematics educator, 19(1), 8-17. retrieved from http://math.nie.edu.sg/ame/matheduc/ cohen, s.g., bailey, d.e. (1997). what makes teams work: group effectiveness research from the shop floor to the executive suite. journal of management, 23(3), 239-290. conley, s., fauske, j. & d.g. pounder (2004). teacher work group effectiveness. educational administration quarterly, 40(5), 663-703. doi:10.1177/0013161x04268841 crow, m.g., & d.g. pounder (2000). interdisciplinary teacher teams: context, design, and process. educational administration quarterly, 36(2), 216-254. doi:10.1177/0013161x00362004 curry, m. (2008). critical friends groups: the possibilities and limitations embedded in teacher professional communities aimed at instructional improvement and school reform. teachers college record, 110(4), 733-774. retrieved from http://www.tcrecord.org/ datnow, a. (2011). collaboration and contrived collegiality: revisiting hargreaves in the age of accountability. journal of educational change, 12(2), 147-158. doi:10.1007/s10833-011-9154-1 decuyper, s., dochy, f., & van den bossche, p. (2010). grasping the dynamic complexity of team learning: an integrative model for effective team learning in organisations. educational research review, 5, 111-133. doi:10.1016/j.edurev.2010.02.002 devine, d.j., clayton, l.d., philips, j.l., dunford, b.b., & melner, s.b. (1999). teams in organizations: prevalence, characteristics, and effectiveness. small group research, 30, 678-711. doi:10.1177/104649649903000602 dickinson, e.b. (2009). the impact of collaborative teacher teaming on teacher learning (specialist project). retrieved from http://digitalcommons.wku.edu dochy, f., gijbels, d., raes, e., & kyndt, e. (2014). team learning in education and professional organisations. in s. billet, c. harteis & h. gruber (eds.), international handbook of research in professional and practice-based learning. the netherlands: springer. drach-zahavy, a., & somech, a. (2002). team heterogeneity and its relationship with team support and team effectiveness. journal of educational administration, 40(1), 44 – 66. doi:10.1108/09578230210415643 euwema, m.c., & van der waals, j. (2007). teams in scholen. samen werkt het beter. leusden: bmc. flowers, n., mertens, s.b., & mulhall, p.f. (2000). what makes interdisciplinary teams effective? middle school journal, 31(4), 53-56. retrieved from http://limo.libis.be gajda, r., & koliba, c.j. (2008). evaluating and improving the quality of teacher collaboration: a fieldtested framework for secondary school leaders. nassp bulletin, 92(2), 133-153. doi:10.1177/0192636508320990 k. vangrieken et al. 95 | f l r gunn, j.h., & king, m.b. (2003). trouble in paradise: power, conflict and community in an interdisciplinary teacher team. urban education, 38(2), 173-195. doi:10.1177/0042085902250466 hackmann, d.g., petzko, v.n., valentine, j.w., clark, d.c., nori, j.r., & lucas, s.e. (2002). beyond interdisciplinary teaming: findings and implications of the nassp national middle level study. nassp bulletin, 86(632), 33-47. doi:10.1177/019263650208663204 haslam, n., rothschild, l., & ernst, d. (2004). essentialism and entitativity: structures of beliefs about the ontology of social categories. in v. yzerbyt, c.m. judd & o. corneille (eds.), the psychology of group perception: perceived variability, entitativity and essentialism (pp. 61-78). new york (ny): psychology press. havnes, a. (2009): talk, planning and decision‐making in interdisciplinary teacher teams: a case study. teachers and teaching: theory and practice, 15(1), 155-176. doi:10.1080/13540600802661360 hollenbeck, j.r., beersma, b., & schouten, m.e. (2012). beyond team types and taxonomies: a dimensional scaling conceptualization for team description. academy of management review, 37(1), 82-106. doi:10.5465/amr.2010.0181 kelchtermans, g. (2006). teacher collaboration and collegiality as workplace conditions. a review. zeitschrift für pädagogik, 52(2), 220-237. retrieved from http://www.pedocs.de leonard, l., & leonard, p. (2003). the continuing problem with collaboration: teachers talk. current issues in education, 6(15). retrieved from http://cie.asu.edu lomos, c., hofman, r.h., & bosker, r.j. (2011). the relationship between departments as professional learning communities and student achievement in secondary schools. teaching and teacher education, 27, 722-731. doi:10.1016/j.tate.2010.12.003 main, k. (2007). a year long study of the formation and development of middle school teaching teams (doctoral dissertation). retrieved from https://www120.secure.griffith.edu.au/rch/file/64a6473e-3a2bf149-bd30-6e2033bbef0f/1/02whole.pdf main, k., & bryer, f. (2005). what does a „good‟ teaching team look like in a middle school classroom? stimulating the “action” as participants in participatory research, 2, 196-204. retrieved from http://www.griffith.edu.au meirink, j.a. (2007). individual teacher learning in a context of collaboration in teams (doctoral dissertation). retrieved from https://openaccess.leidenuniv.nl meirink, j.a., imants, j., meijer, p.c., & verloop, n. (2010). teacher learning and collaboration in innovative teams. cambridge journal of education, 40(2), 161-181. doi:10.1080/0305764x.2010.481256 mertens, s.b., & flowers, n. (2004). research summary: interdisciplinary teaming. retrieved from http://www.nmsa.org ohlsson, j. (2013). team learning: collective reflection processes in teacher teams. the journal of workplace learning, 25(5), 296-309. doi:10.1108/jwl-feb-2012-0011 park, s., henkin, a.b., & egley, r. (2005).teacher team commitment, teamwork and trust: exploring associations. journal of educational administration, 43(5), 462 – 479. doi:10.1108/09578230510615233 pounder, d. (1998). chapter 5: teacher teams: redesigning teacher‟s work for collaboration. in d. pounder (ed.), restructuring schools for collaboration: promises and pitfalls (pp. 65-88). albany: state university of new york press. raes, e., kyndt, e., decuyper, s., van den bossche, p., & dochy, f. (submitted). group development and team learning: how development stages affect team-level learning behavior. human resource development quarterly. rone, b.c. (2009). the impact of the data team structure on collaborative teams and student achievement (doctoral dissertation). retrieved from http://proquest.umi.com salas, e., burke, c.s., & cannon-bowers, j.a. (2000). teamwork: emerging principles. international journal of management reviews, 2(4), 339-356. saunders, w.m., goldenberg, c.n., & gallimore, r. (2009). increasing achievement by focusing grade-level teams on improving classroom learning: a prospective, quasi-experimental study of title i schools. am educational research journal, 46(4). doi:10.3102/0002831209333185 k. vangrieken et al. 96 | f l r scribner, j.p., hager, d.r. & warne, t.r. (2002). the paradox of professional community: tales form two high schools. educational administration quarterly, 38(45). doi:10.1177/0013161x02381003 smith, g. (2009). if teams are so good... science teachers‟ conceptions of teams and teamwork (doctoral dissertation). retrieved from http://eprints.qut.edu.au somech, a. (2005). teachers‟ personal and team empowerment and their relations to organizational outcomes: contradictory or compatible constructs? educational administration quarterly, 41(2), 237266. doi:10.1177/0013161x04269592 somech, a. (2008). managing conflict in school teams: the impact of task and goal interdependence on conflict management and team effectiveness. educational administration quarterly, 44. doi:10.1177/0013161x08318957 supovitz, j.a. (2002). developing communities of instructional practice. teachers college record, 104(8), 1591-1626. retrieved from http://www.tcrecord.org tonso, k.l., jung, m.l. & m. colombo (2006). “it‟s hard answering your calling”: teacher teams in a restructuring urban middle school. research in middle level education, 30(1), 1-22. retrieved from http://www.eric.ed.gov van den bossche, p., gijselaers, w.h., segers, m., & kirschner, p.a. (2006). social and cognitive factors driving teamwork in collaborative learning environments: team learning beliefs and behaviors. small group research, 37(5), 490-521. doi:10.1177/1046496406292938 vangrieken, k., dochy, f., & raes, e. (submitted). teacher teams and collaboration: a review. visscher, a.j., & witziers, b. (2004). subject departments as professional communities? british educational research journal, 30(6), 785-800. doi:10.1080/0141192042000279503 watson, s. t. (2005). teacher collaboration and school reform: distributing leadership through the use of professional learning teams (doctoral dissertation). retrieved from http://edt.missouri.edu westheimer, j. (2008). chapter 41: learning among colleagues: teacher community and the shared enterprise of education. in m. cochran-smith, s. feiman-nemser, & j. mcintyre (eds.). handbook of research on teacher education. reston, va: assocation of teacher educators; lanham, md: rowman wigglesworth, m. (2011). the effects of teacher collaboration on students understanding: relating to high school earth science concepts. montana: lap lambert academic publishing. williams, m.l. (2010). teacher collaboration as professional development in a large, suburban high school (doctoral dissertation). retrieved from http://digitalcommons.unl.edu yisrael, s.b. (2008). a qualitative case study: the positive impact interdisciplinary teaming has on teacher morale (doctoral dissertation). retrieved from http://etd.ohiolink.edu appendix table 1. typology framework task governance/management 1 instruction (according to grade level or subject) instruction/teaching examining individual student work generated from common formative assessments (rone, 2009) developing instruction to address the academic needs of students (saunders et al., 2009) keep track of the progress and revise instruction (saunders et al., 2009) studying previous test data of students (bertrand, roberts, & buchanan, 2006) coordinating instruction, communication and assessment for a common group of students (flowers et al., 2000) 1 as mentioned earlier, management and special services teams were not the focal point of this study because of the fact that these often include other team members than just teachers and for that reason no literature considering this type of teams is discussed here. k. vangrieken et al. 97 | f l r developing and implementing interdisciplinary curriculum and teaching strategies based on the developmental needs of the children (crow & pounder, 2000) developing coordinated interventions and management strategies to tackle problems considering student learning (crow & pounder, 2000) team teaching (brouwer, 2011; brouwer, brekelmans, & nieuwenhuis, 2012; main, 2007; main & bryer, 2005) coherent curriculum development (organisation of education and discussing students) (brouwer, 2011; brouwer et al., 2012) planning instruction planning, coordinating and evaluating of curriculum and instruction across academic areas (yisrael, 2008; mertens & flowers, 2004) planning curriculum and developing assessments (gunn & king, 2003) realising common goals across different classes (main & bryer, 2005) set and share academic goals (saunders et al., 2009) collaboratively planning and administering assessment (main & bryer, 2005) development and implementation of the subject matter (somech, 2008) collaborating on instructional strategies (wigglesworth, 2011; supovitz, 2002) evaluating collaboratively constructed materials (wigglesworth, 2011) developing course syllabi and benchmark tests (bertrand et al., 2006) planning interdisciplinary teaching (havnes, 2009) coordinating individual subject-specific teaching (havnes, 2009) work together to plan, design, integrate and implement shared instructional methods, curricula and assessment targeted towards curricular and pedagogical alignment (watson, 2005) decision-making authorities considering curricular emphasis and coordination (conley et al., 2004) decision-making authorities considering student class assignment and flexible grouping strategies, student assessment (conley et al., 2004) decision-making authorities considering curricular and co-curricular scheduling (conley et al., 2004) problem-specific planning (tackle specific problems. temporary or long term) planning teams are responsible for tackling specific problems and can be of a temporary or a longer lasting nature (park et al., 2005) pedagogy developing coordinated interventions and management strategies to tackle problems considering student learning and/or behavior (crow & pounder, 2000) providing coordinated communication with parents (crow & pounder, 2000) building-wide support and intervention programs for students, monitor the effectiveness of these programs and make improvement recommendations (watson, 2005) communication (with families) (mertens & flowers, 2004) continually exploring their curricular and pedagogical strategies and the influences of these on student learning (supovitz, 2002) discussing teaching, practice, the challenges they experience as teachers, and pedagogy (havnes, 2009) decision-making authorities considering student management and behavioural interventions (conley et al., 2004) decision-making authorities considering coordinated parent communication (conley et al., 2004) special/social services 2 innovation and school reform 2 see 1 k. vangrieken et al. 98 | f l r designing and experimenting with new teaching practices (meirink et al., 2010) learning collaboratively learning (saunders et al., 2009; watson, 2005) sharing expertise and experience across generations (carroll & foster, 2008) material/practical these teams lack a shared task, but share for example resources. this collaboration is mostly realised for practicality reasons. budgetary allocation (main & bryer, 2005) sharing of resources and/or physical space (main & bryer, 2005) practical support (smith, 2009) discipline level interdisciplinary teachers from different subject areas are part of the team. disciplinary teachers from the same subject area or part of the team. grade level within-grade level teachers responsible for students from the same grade-level cross-grade level teachers responsible for students from different grade levels temporal duration temporary for a definite time/project lasting team entitativity low (1-2 criteria met) moderate (3-4 criteria met) high (5-6 criteria met) microsoft word baptista et al_publication.docx         frontline learning research vol.3 no. 3 special issue (2015) 55 67 issn 2295-3159 1 corresponding author: ana baptista, learning development, student services directorate, mile end library queen mary university of london, mile end road, london e1 4ns. phone: +44 (0)20 7882 2838, email a.baptista@qmul.ac.uk, doi: http://dx.doi.org/10.14786/flr.v3i3.147   the doctorate as an original contribution to knowledge: considering relationships between originality, creativity, and innovation ana baptistaa, liezel frickb, karri holleyc, marvi remmikd, jakob tesche, gerlese âkerlindf aqueen mary university of london, uk bstellenbosch university, south africa cuniversity of alabama, usa duniversity of tartu, estonia einstitute for research information and quality assurance, germany faustralian national university and university of canberra, australia article received 31 january 2015 / revised 17 july 2015 / accepted 3 august 2015 / available online 12 october 2015 abstract this article explores the meaning of originality in doctoral studies and its relationship with creativity and innovation. doctoral theses are expected to provide an original contribution to knowledge in their field all over the world. however, originality is not well defined. using the literature on concepts of originality as a foundation, this article shows that originality is not a concept commonly understood. creativity introduces a focus on the production of knowledge, which is not just novel but also meaningful. innovation is becoming of increasing importance in doctoral theses with the societal shift to knowledge-based economies and introduces the requirement of immediate relevance for economic purposes in doctoral education. while the three elements appear to be substantial building blocks of the potential contribution doctoral work can make in the 21st century, it is unclear the extent to which doctoral theses fulfil these expectations. the article discusses this problem with a focus on implications for doctoral education. keywords: doctoral education; originality; creativity; innovation baptista et al   | f l r       56   1. introduction the role of the doctoral thesis as an original contribution to knowledge has traditionally signalled a high level of intellectual output within the academic discipline. while considered an essential component of doctoral education, the nature of originality is typically ill-defined. commonly associated with the production of new knowledge, originality is increasingly seen as inherent to creativity and innovation (european universities association, 2010). however, how the three concepts of originality, creativity and innovation operate within the doctoral education process, independently and collectively, is unclear. in addition, questions remain over how and whether originality, creativity, and innovation may be facilitated in doctoral programs, even though these concepts are commonly found in policy documents and literature on doctoral education. nowotny, scott and gibbons (2001) suggest the production of knowledge within the knowledge society values creativity, application and flexibility, a process that is enhanced in the doctoral environment (walsh, anders & hancock, 2013). doctoral students form a key component of such knowledge production and are therefore directly influenced by how such notions – specifically originality, creativity and innovation – are defined and influence each other. this article therefore explores the meaning of originality in doctoral studies and the relationship with innovation and creativity. the aim is to provide insights into the nature of originality in doctoral education for 21st century knowledge societies. 2. originality the debate about the originality of doctorates dates back to the 19th century (mommsen, 1876). while originality has been a long-held requirement of doctorates, the publication of doctoral theses introduced in the 19th century helped to reduce fraud and enabled the assessment of originality by relevant disciplinary communities. for example, since the first uk doctorate was awarded in 1917, the degree has required “an output that constitutes original research as defined by the academic community into which the candidate wishes to be admitted” (qaa, 2011, p.12). this requirement places thesis examiners in a powerful brokerage position with responsibility to enact a judgement of originality on behalf of their respective academic community, although the assessment of appropriate degrees of originality differs substantially amongst examiners (clarke & lunt, 2014; denicolo, 2003; johnston, 1997). for over a century, the quality of originality has been considered essential to the doctoral thesis (european universities association, 2007, 2010; australasian qualifications framework advisory board, 2007; hornbostel, 2009; association of american universities 1998, in lovitts, 2005; new zealand qualifications authority, 2001; uk quality assurance agency for higher education, 2011) in order to achieve what is commonly now referred to as ‘doctorateness’ (wellington, 2013). the council for doctoral education of the european universities association (eua), as one example, recommends as the first principle of doctoral education that “the core component of doctoral training is the advancement of knowledge through original research” (2010, p.2). moving beyond a surface-level assessment of originality requires attention to the development of original thought and original work (clarke & lunt, 2014). for the former, new knowledge might be generated as a result of the doctoral thesis, or existing knowledge might be applied to result in a new understanding. for the latter, developing a musical score or a painting can indicate original work. not only are doctoral students required to assess and categorise existing bodies of knowledge through this process, but they also draw conclusions regarding knowledge and make decisions about implementation (simpkins, 1987, cf. lovitts, 2007). originality may be evident in the study’s design, the knowledge synthesis, the implications, or the way in which the research is presented (wellington, 2010). baptista et al   | f l r       57   this assessment emphasises the nuanced ways in which the outcome of originality might be achieved. applying existing methods to new data could result in incremental additions to the knowledge base, while the application of new methods, new questions, or new ideas could generate more substantial shifts in knowledge (lovitts, 2005). this variability underscores the emphasis on significance in doctoral research. whilst significance is not inherently a component of originality (johnston, 1997), it is important to note that original research within the context of doctoral education is expected to provide knowledge of significance to the field of study (tinkler & jackson, 2004). these varying perspectives on originality show that it does not have a universal definition, nor does it manifest in the same way in all doctoral work. originality is not only related to an outcome or product, but also to the overall process of producing an outcome. a doctoral student cannot achieve a product without undergoing a process that stimulates the creation of that product. what is deemed original may vary between disciplines, programmes and even individual projects. the originality of a dissertation can be expressed in a number of ways, and the kind of originality that is recognised and appreciated has traditionally been dependent on discipline (guetzkov, lamont & mallard, 2004; lamont, 2009; lovitts, 2007). disciplinary variation influences the assessment of originality. for example, clarke and lunt (2014) suggest that originality in science, technology, engineering and mathematics disciplines is defined by publishability, whilst in arts, humanities and social sciences it is related to intellectual originality. guetzkow and colleagues (2004) argue that natural sciences define originality “as the production of new findings and new theories”, while social sciences and humanities define it “much more broadly: as using a new approach, theory, method, or data; studying a new topic, doing research in an understudied area; or producing new findings” (p.190). disciplinary implications are evident for phd students’ perceptions and expectations about the phd as process and product, and also for the way students learn how to do research, and consequently what it means to be original. knowledge is rarely de-contextualised, and numerous factors influence the way an individual frames a question and chooses the path to answer that question. disciplines consist of old and emerging specialisms (kekäle, 2000), and how these different bodies of knowledge are defined and arranged determines the output (bailin, 1985). knowledge defined as old or emergent may intertwine to create a process or product that may be called original. delamont, atkinson and parry (2000, p.174) state that: “the originality of postgraduate research is always defined in terms of the essential tension between accepted prior knowledge and new discoveries or ideas”. disciplinary influences are evident in cultural norms including the research process (such as group projects or those led by a supervisor), the form of the thesis (such as monograph or articlebased), and the long-term impact on the field (such as future publication and citation impacts). thus, a definition of originality in doctoral degrees assumes different nuances in different contexts. numerous issues should be considered in addressing originality in doctoral education: • the interplay between old and new, i.e. that originality inevitably builds on existing knowledge and practices in some way; • disciplinary variation in originality; • the existence of degrees of originality, and the need for originality to be accompanied by significance; • the need to address originality in doctoral process as well as product, with associated implications for research training. both bennich-björkman (1997) and beghetto (2013) agree that originality can be defined as something that is new or novel, but originality does not necessarily have to be applicable or relevant. herein lies the difference between originality and creativity, as described below. baptista et al   | f l r       58   3. creativity along with the expectation of originality, doctoral research is strongly associated with creativity, commonly as a way in which students engage in the research process. for example, the australian qualifications framework (2013) specifies that doctoral graduates are required to demonstrate “the application of knowledge and skills with initiative and creativity”. thus creativity implies that a contribution (such as a doctoral thesis) needs to be both novel (original) and relevant (according to bennich-björkman, 1997) or applicable (according to beghetto, 2013). beghetto (2013) defines creativity as anything deemed as both original and task-appropriate within a particular socio-cultural-historical context – such as an academic discipline. the genealogy of creativity can be traced back to the greek word ‘krainein’, which means to fulfil. people who fulfil their potential, who express an inherent drive or capacity, can be seen as creative (evans & deehan, 1988). pope (2005, p.11) consequently defines creativity as “the capability to make, do or become something fresh and valuable with respect to others as well as ourselves”, which involves “a grappling deep within the self and within one’s relations with others: an attempt to wrest from the complexities and contradictions we have internalised”. this definition goes beyond creativity in the thesis production and process, to creativity of the person, i.e., the doctoral graduate themselves. this positions creativity as including the full realisation and expression of a person’s potential (lovitts, 2008; mackinnon, 1970) – thus ‘becoming doctorate’, a responsible and independent scholar (barnacle, 2005). assessing creativity requires attention to the intellectual context, including big c creativity, or that which brings about knowledge new to the human race, and pro c creativity, which occurs within a professional workspace (kaufman & beghetto, 2007). the disciplinary context adds another important variable, underscored by the key elements of motivation, independence, and intellectual challenge (jurisevic, 2011). bennich-björkman’s (1997) classification scheme (see table 1) offers further insights into the relationship between originality and creativity. table 1 classification of research contributions (adapted from bennich-björkman, 1997, p.25) is the contribution novel? yes no is the contribution relevant? yes creative cumulative no original replication the relationship between originality and creativity, according to bennich-björkman, is defined through novelty and relevance. in principle, relevance may be determined at individual, societal or economic levels (steinberg & lubart, 1999), but in the case of the doctorate, most commonly refers to the judgment of the disciplinary community in which the doctorate is produced. while creative work is expected to be relevant as well as novel, originality is expected only to be novel. by taking the focus off of immediate relevance, the pure concept of originality recalls blue skies research and an emphasis on the pursuit of knowledge for its own sake. this view of originality thus seems appropriate to the time when the expectation of original research was first introduced into the doctorate, with the rise of the modern university in the late 19th century. in addition to counterposing creativity and originality, bennich-björkman’s classification attempts to define knowledge production that is not original. cumulative research is characterised as being highly baptista et al   | f l r       59   relevant, in the sense of being valuable or useful to disciplinary communities, but not novel. this focus on relevance positions cumulative research as a valuable contribution to knowledge, but neither original nor creative. replication of research is positioned as neither novel nor relevant, but is nonetheless an important aspect of knowledge development that increases the reliability of research findings and thus trust in the outcomes – small studies may be replicated on a larger scale or with another sample, for instance. disciplinary differences matter, as cumulative research and replicative studies are not uncommon in many natural science doctorates. thus, the ‘in practice’ definition of originality in doctoral theses may be made as much on pragmatic grounds as on conceptual ones. the product of a creative endeavour demonstrates an original and appropriate contribution that has purpose and can be judged by some sort of external criteria (sternberg & lubart, 1999). a process-product distinction exists between creativity and originality, with the idea of a creative process underpinning an original product or outcome. this distinction has implications for the design of doctoral education, suggesting that originality in research outcomes may best be achieved by encouraging creative processes during the candidature, such as a creative learning environment or peer collaborations. the notion of fit for purpose that our discussion has highlighted as a key aspect of creativity raises questions such as fit for whom or what? such questions open the door to innovation being one of the drivers of research in the 21st century that also needs to be considered in the contributions doctoral work is expected to make. 4. innovation innovation has become an increasing expectation of doctoral studies as part of the global post world war ii economic shift from industrial and manufacturing based economies to technological and knowledge based economies (delanty, 2001; marginson & considine, 2000; rolfe, 2013). by definition, innovation involves the process of transforming an invention into practical application, and is most commonly associated with private industry (marsh, 2010). as the production of knowledge has come to be of increasing importance to national economies, university research is expected to better serve the needs of industry, through innovation in science and technology in particular. the term ‘innovation’ is most often found in economic discourses on production processes or products (marsh, 2010). governmental higher education policies place an emphasis on stronger links between industry and universities, and development of knowledge that can be exploited for economic benefit (delanty, 2001; henkel, 2000), bringing the concept of innovation firmly into the 21st century doctoral education. the lisbon declaration on the purpose of europe’s universities (2007) strongly links university research with innovation, emphasising the importance of universities’ “capacity for promoting cultural, social and technological innovation” (p.1) and that “to meet the challenges of the twenty-first century (...) [requires] technological and social innovation which will solve problems as they arise and ensure economic success” (p.2). thus, innovation as part of doctoral research privileges the production of knowledge that is economically useful, either in terms of technological advances or societal use. technological innovation is typically linked to marketable technologies, for example developing patents. social innovation would relate to applied research aimed at improving societal conditions or solving societal problems. examples are abundant in a variety of disciplines ranging from medicine (eg, curbing mother to child transfer of hiv/aids) to education (eg, improving literacy rates). in classical economic theory, innovators are considered creative entrepreneurs who successfully acquire monopoly positions with innovative products or production processes (schumpeter, 1912). innovation is defined as the practical application of a novel, and thus original idea, but it must be an idea with a potential application: “innovations of any kind start with some kind of creative enterprise, and the enterprise must produce work that is not just novel, but useful. innovation is the channelling of creativity so as to produce a creative idea and/or product that people can and wish to use” (sternberg, pretz & kaufman, 2003, p.158). baptista et al   | f l r       60   the doctorate is increasingly economically positioned as an important source of skilled and innovative knowledge workers, as required by a knowledge-based economy with a strong emphasis on research and development. this position has led to an exponential growth in the number of phds awarded internationally, especially in the natural sciences and engineering (cyranoski, gilbert, ledford, nayar & yahia, 2011), and a shift in expectations of employment post-phd away from academia and towards industry, government and private enterprise (auriol, 2010; enders, 2005). innovation has claimed a prominent place in defining a key purpose of the 21st century doctorate as preparing the candidate for a future career in either academe or industry, and developing skills for employability (wellington, 2013). the extent to which these developments have changed the conditions under which knowledge is produced in doctoral theses and science in general is unclear (geiger, 2004). the literature on thesis examiners shows hardly any expectation of innovation in doctoral theses in terms of developing applications for industry, though engineering is an exception here, where an application of existing methods to a problem from engineering practice is considered original, just as is the invention of new devices (lovitts, 2007, p.173). similarly, the conceptualisation of originality in economics, as the application of existing methods to a novel problem, is also considered original (lovitts, 2007, p.173). both disciplines consider practical problem solving as an original contribution. 5. implications for doctoral education risk is intrinsically linked to originality, creativity and innovation, and is thus an unavoidable element of doctoral education (frick, albertyn & bitzer, 2014). doctoral education is inherently risky given the requirement to produce original knowledge. the lisbon declaration (2007) argues that universities “should encourage a culture of risk-taking (...) in order to produce an institutional milieu favourable to creativity, knowledge creation and innovation” (p.3), reinforcing the idea that an original contribution requires a certain amount of risk-taking in choosing a topic and approach, due to the novelty aspect inherent to originality. students need to have “the courage and confidence to take risks, to make mistakes, to invent and reinvent knowledge, and to pursue critical and lifelong inquiries in the world, with the world, and with each other” (freire, 1970, cited in lin & cranton, 2005, p.458). mackinnon (1970) agrees that the courage to take risks is an important characteristic of creative endeavours – such as doctoral studies. however, balancing risk with originality, creativity and innovation may provide challenges for the supervisory relationship and the research process (brown, 2010; latham & braun, 2009). therefore, it is important not only to manage risk constructively, but also to understand how it manifests within doctoral education. byrnes, miller and schafer (1999) refer to four aspects that need consideration when defining risk that could be applied to doctoral education. firstly, risk is closely associated with goals, values and outcomes. hence, the importance of current debates about the purpose of a doctorate in a risk society full of uncertainties and changes (park, 2005, 2007), as well as the definition of supervisory and research responsibilities and roles that characterise doctoral students and supervisors. secondly, risk involves interplay between an individual’s subjective perception of risk and the perceptions of the larger community. different students and different supervisors may interpret risk differently, which may influence how they negotiate their relationship and study focus. thirdly, individual characteristics determine the extent of possible risk. for instance, a study may be less risky if the doctoral student has particular research and/or subject expertise. finally, context determines “who can take what risks and how” (hood, jones, pidgeon, turner, gibson & bevan-davies, 1992, p.136). for example, certain projects may become less risky if expert supervision and other resources are readily available. this conceptualisation of risk reflects significant forces that relate to elements in the context, relationships in the supervisory process, and individual characteristics of doctoral students. these forces are reflected in the broader literature on doctoral education, which highlights several factors that may affect the overall success of a doctorate, including: (i) characteristics of the doctoral candidate themselves; (ii) nature baptista et al   | f l r       61   of the doctoral supervision experienced; and (iii) institutional, departmental, disciplinary and external cultures. each of these factors is explored in more detail below. individual student characteristics can strongly impact on the originality of their work. for instance, doctoral education requires that students at times work independently in an uncertain environment. within this environment, healthy program cultures encourage risk-taking by students within the context of the field. the interpretation of risk is a process fraught with possible complications, particularly in terms of the expertapprentice relationship still prevalent between the supervisor and student. however, students who have been socialised in an undergraduate academic culture or a professional environment that promote novel ways of knowing will have a stronger foundation for originality. in addition to student characteristics, doctoral supervision is one of the most important influences on research student outcomes (latona & browne, 2001; seagram, gould & pyke, 1998). evans (2004) conceptualizes the role of the supervisor as that of risk manager and risk mitigator, acting as an intermediary between the demands of society, the discipline(s) involved, the institution and the doctoral candidate. frick, albertyn and bitzer (2014) report various strategies that supervisors use at different stages during the doctorate to support students and mitigate risk, including formulating clear expectations; determining and developing student capability, independence, analytical thinking skills, problem solving skills, integrative thinking skills, creativity, and expectations during the student selection phase; encouraging wide reading, critical debate, benchmarking, time for incubation of ideas, and challenging students during conceptualisation of the study; developing academic writing and methodological skills through incorporating expert input; supporting networking, colloquia, regular contact, communication, co-supervision and mentoring practices; and promoting peer review and writing for publication during the doctorate. they encourage further research that explores ways of balancing rather than controlling risk, while encouraging innovation in the doctoral education process. increased awareness of risk could lead supervisors to contain risk in a responsible manner. of course, it is not only the student who assumes the risk in terms of research, but also the supervisor. institutional, departmental, disciplinary and external cultures influence how faculty and students engage with a doctoral curriculum. backhouse (2009), frick (2012) and holligan (2005) point to cultural factors (including bureaucratic institutional systems, ethics and funding policies) as determinants of the extent to which risk-taking is possible in doctoral studies. for instance, a danger of the current emphasis on doctoral throughput in the minimum allocated time is that it may lead to avoiding the risk of choosing a complex and less defined problem. not all research that may be considered original requires lengthy periods of time, but nor can all research be contained within minimum, finite time periods. ultimately, the process of doctoral education is influenced by the various cultures in which such work takes place. in particular, how such cultures define novel knowledge outcomes is highly relevant. clearly, approaches to doctoral education that might encourage originality are patchy, making it difficult to design an educational agenda for the future when there are so many uncertainties and unpredictable changes embedded in doctoral (research) education and supervision, and when concepts that characterise this challenging high-level process overlap and seem somewhat blurred. but perhaps operating in a state of uncertainty, unpredictability and blurred boundaries is what the future of higher education is all about. 6. conclusions: insights into the nature of originality in doctoral research we can see from this examination of originality, creativity and innovation the extent to which all three concepts are often defined with reference to each other. clearly, these concepts share a focus on novelty in research. where the concepts differ is in the underlying purpose or intention for seeking novelty – with creativity it is disciplinary relevance or value, with innovation it is useful economic outcomes, whilst baptista et al   | f l r       62   with originality it is more blue skies knowledge seeking – but all three of these concepts may influence the way in which the potential contribution of doctoral work is seen. but whilst originality may be free of instrumental connotations, a doctorate is not. doctoral theses are expected to make not just an original but also significant contribution to the field, the implication being that there is little value in originality if it is not also significant. however, the determination of significance is context-dependent. what would be considered significant in the 19th century would likely be different to the 21st century, and in one discipline or subspecialisation different to another, for instance. it could be argued that creativity and innovation all incorporate originality, in the form of novelty in research. hence, it may be possible to have originality without creativity or innovation, but not vice versa. meanwhile, all three concepts can contribute to the development of the doctoral contribution in overlapping but different ways. conceptually, the links between these concepts can be displayed as follows: figure 1. the relationship between originality, creativity and innovation. in figure 1 we show that originality, creativity and innovation are related elements that can all contribute to the doctoral contribution, but that the emphasis shifts depending on the concept. as doctorateness seems to be a multi-faceted concept itself (wellington, 2013) this fluid emphasis may be useful to allow for (trans)disciplinary, programme and individual differences in what it means to be doctorate. baptista et al   | f l r       63   meanwhile, in the current economic and socio-political climate, the question of whether doctoral studies can or should be safe-guarded from instrumental requirements for applied relevance must be considered. doctoral theses call not just for originality, but originality that advances the field in a substantial way. just as the internal characteristics of the field change over a period of time, so does the external context which helps give shape to (and ultimately, contribute to a definition of) knowledge production. while this demand need not include the focus on economic benefits or relevance attached to innovation or creativity, it still places constraints on the type of originality considered appropriate for a doctoral thesis. appropriate approaches to developing originality as part of doctoral education need to be considered. although expectations of originality in doctoral work seem ubiquitous, there is little literature on design of curricula or pedagogical processes for supporting the development of originality. as described above, the concept remains vague to examiners and supervisors (clarke & lunt, 2014; lovitts, 2007). meanwhile, a common assumption seems to exist that the process of engaging in doctoral research will in and of itself lead to originality, as if through some magical process: “the goal of doctoral education is to cultivate the research mindset, to nurture flexibility of thought, creativity and intellectual autonomy through an original, concrete research project. it is the practice of research that creates this mindset” (european universities association, 2010, p.2). the unanswered question from this statement is how the practice of research cultivates these attributes, and in what ways doctoral education might intentionally foster these outcomes. such vague notions for ensuring the development of such a central expectation of doctoral education seem inappropriate in the context of the 21st century focus on higher education efficiency, accountability and quality assurance. considering the ways in which doctoral education can facilitate originality requires attention to the doctoral curriculum, i.e. process, as well as the thesis outcomes, i.e. product. 7. outlook in exploring the nature of originality, this article has linked different conceptualisations of novelty as applied to doctoral theses, showing that while originality appears to be the basic requirement, other expectations such as creativity and innovation, and associated criteria of usefulness and economic advancement have recently appeared on the agenda. this association suggests a new differentiation in the requirements for doctoral theses. however, the relation between these concepts is not yet fully clear. the question remains as to whether the differentiation of requirements for a doctoral thesis is just a mirror of changes affecting research and knowledge creation in general, or whether there are more nuanced issues to consider related to doctoral education specifically. as the doctorate is seen as the initial process in becoming a researcher, changing requirements for the doctorate will most likely affect the way knowledge creation operates in the future. higher education has experienced these changes before. as one example, the publication of doctoral theses is now commonplace, and many institutions offer open public access to theses produced by doctoral graduates. another example involves the development of the group dissertation for certain disciplines. these so-called ‘capstone projects’ not only encourage students to work collaboratively, but they often involve external stakeholders. the challenge of defining original research has implications for the nature of doctoral training, and specifically for the internal function of disciplines and for the relation between academic disciplines and society. future research should examine the extent to which these new requirements are part of institutional guidelines, supervisors’ expectations and doctoral students’ identity conceptualisations. an even more fundamental question is about the determination or assessment of originality. a troubling reality underscores the consideration of originality in doctoral education – to what extent have doctoral theses ever been shown to fulfil the requirement of an original and significant contribution to knowledge, apart from via the subjective judgments of examiners? with theses by publication becoming more widespread, new pathways for intra-individual replicapability of originality and in depth analysis baptista et al   | f l r       64   emerge, for example through the application of bibliometric tools and content analysis of citations. however, the question of which stakeholders should be involved in this assessment and what bibliometric indicators might be utilised are unresolved issues. another almost unquestioned theme in the extant literature is that originality arises out of the doctoral training process, be it an intensive supervisor-mentee relationship or more structured doctoral training conditions. this assumption is particularly noteworthy given that no valid database exists that can be used to demonstrate whether a doctoral thesis can be considered original, much less which experiences contribute to a doctoral student being able to perform such work. future research should take steps towards unpacking the relationship between doctoral training conditions and outcomes, in the sense of fulfilling the requirement of originality. the following questions offer ideas for future research: • what are doctoral program designers’ conceptualisations of originality? • how do these relate to conceptualisations of originality by supervisors, examiners and students? • which requirements can be achieved through better training, and which are dependent on individual characteristics of doctoral students, such as propensity for risk-taking? cross country and international comparisons could be valuable here; although the doctorate shares commonalities in the international context, the degree to which the doctorate is organised as a training process varies from country to country. this article has considered how originality builds on existing knowledge and practices by stimulating an interplay between old and new. how should doctoral curricula and the supervisory relationship explicitly develop students’ originality skills? it is incorrect to assume that all doctoral supervisors and those who design curricula at doctoral level at all higher education institutions possess originality skills themselves. additionally, formal structures at contextual and institutional levels, where doctoral education and supervision take place, as well as in national contexts stimulate both the definition of originality as well as the attitude towards research and knowledge. to tackle these questions, the research agenda for the future should open spaces for discussions about the place of originality in the supervisory relationship, curricula design, and the cultural environment that an institution and even a research group has to offer. disciplines should strengthen dialogues about the requirements for a doctoral thesis in their field, and research should supply these discussions with evidence based knowledge. simultaneously, a critical approach to the different discourses at different levels should be reviewed in the light of the most relevant and updated literature. these dialogic interactions between practices, perceptions and research may be a way of improving the overall experience students and supervisors will have in doctoral programs. keypoints in exploring the nature of originality, this article has linked different conceptualisations of novelty as applied to doctoral theses, showing that while originality appears to be the basic requirement, other expectations such as creativity and innovation, and associated criteria of usefulness and economic advancement have recently appeared on the agenda. the challenge of defining original research has implications for the nature of doctoral training, and specifically for the internal function of disciplines and for the relation between academic disciplines and society. further research must be carried out in order to shed light on possibly diverse ways of determining or assessing originality baptista et al   | f l r       65   references australian qualifications framework. (2013). aqf specification for the doctoral degree. retrieved in october 2014, from http://www.aqf.edu.au/wp-content/uploads/2013/05/14aqf_doctoral-degree.pdf auriol, l. (2010). careers of doctorate holders: employment and mobility patterns. sti working paper 201/4. paris: oecd. statistical analysis of science, technology and industry. backhouse, j. (2009). creativity within limits: does the south african phd facilitate creativity in research? journal of higher education in africa, 7(1/2), 265-288. bailin, s. (1985). on originality. interchange, 16(1), 6-13. barnacle, r. (2005). research education ontologies: exploring doctoral becoming. higher education research and development, 24(2), 179-188. bennich-björkman, l. (1997). organising innovative research: the inner life of university departments. oxford: iau press, pergamon. brown l. 2010. balancing risk and innovation to improve social work practice. british journal of social work, 40, 1211–1228. byrnes, j.p., miller, d.c., & schafer, w.d. (1999). gender differences in risk taking: a meta-analysis. psychological bulletin, 125(3), 367-383. clarke, g., & lunt, i. (2014). the concept of ‘originality’ in the ph.d.: how is it interpreted by examiners? assessment & evaluation in higher education, 39(7), 803-820. cyranoski, d., gilbert, n., ledford, h., nayar, a., & yahia, m. (2011). the phd factory. nature, 472, 277279. delanty, g. (2001). challenging knowledge – the university in the knowledge society. buckingham: srhe & open university press. delamont. s., atkinson, p., & parry, o. (2000). the doctoral experience. success and failure in graduate school. london: falmer press. denicolo, p. (2003). assessing the phd: a constructive view of criteria. quality assurance in education, 11(2), 84-91. enders, j. (2005). border crossings: research training, knowledge dissemination and the transformation of academic work. higher education, 49(1-2), 119-133. european universities association. (2007). lisbon declaration europe’s universities beyond 2010: diversity with a common purpose. brussels: eua. european universities association. (2010). salzburg ii recommendations european universities' achievements since 2005 in implementing the salzburg principles. brussels: eua. evans, p., & deehan, g. (1988). the keys to creativity. london: grafton. evans, t. (2004). risky doctorates: managing doctoral studies in australia as managing risk. paper presented at the australian association for research in education conference, melbourne, 28 november – 2 december 2004. frick, b.l. (2012). pedagogies for creativity in science doctorates. in a. lee & s. danby (eds.), reshaping doctoral education: programs, pedagogies, curriculum (pp. 113-127). london: routledge. frick, b.l., albertyn, r.m., & bitzer, e.m. (2014). conceptualising risk in doctoral education: navigating boundary tensions. in e.m. bitzer, r.m. albertyn, b.l. frick, b. grant & f. kelly (eds.), candidates, supervisors and institutions: pushing postgraduate boundaries. stellenbosch: sunmedia. geiger, r. (2004). knowledge and money: research universities and the paradox of the marketplace. stanford: stanford university press. guetzkov, j., lamont, m., & mallard, g. (2004). what is originality and the humanities and the social sciences? american sociological review, 69(2), 190-212. henkel, m. (2000). academic identities and policy change in higher education. london and philadelphia: jessica kingsley publishers. holligan, c. (2005). fact or fiction? a case history of doctoral supervision. educational research, 47(3), 267-278. baptista et al   | f l r       66   hood, c., jones, d.k.c., pidgeon, n.f., turner, b.a., gibson, r., & bevan-davies, c. (1992). risk management. in the royal society (eds.), risk: analysis, perception and management. london: the royal society. hornbostel, s. (2009). promotion im umbruch – bologna ante portas. in m. held, g. kubon-gilke & richard (eds.), bildungsökonomie in der wissensgesellschaft: band 8. jahrbuch normative und institutionelle grundfragen der ökonomik (pp. 213–240). marburg: metropolis verlag. johnston, s. (1997). examining the examiners: an analysis of examiners' reports on doctoral theses. studies in higher education, 22(3), 333-347. juriševič, m. (2011). postgraduate students’ perception of creativity in the research process. cepsjournal, 169. kaufman, j. c., & beghetto, r. a. (2009). beyond big and little: the four c model of creativity. review of general psychology, 13(1), 1. kekäle, j. (2000). quality assessment in diverse disciplinary settings. higher education, 40(4), 465-488. lamont, m. (2009). how professors think: inside the curious world of academic judgment. cambridge, ma: harvard university press. latham s. & braun m. 2009. closing the loop: innovation and decline. academy of management annual meeting, chicago, il. august 7-11. latona, k., & browne, m. (2001). factors associated with completion of research higher degrees. canberra: higher education division, department of education, training and youth affairs. lin l & cranton p. 2005. from scholarship student to responsible scholar: a transformative process. teaching in higher education, 10(4), 447–459. lovitts, b.e. (2005). being a good course-taker is not enough: a theoretical perspective on the transition to independent research. studies in higher education, 30(2), 137-154. lovitts, b.e. (2007). making the implicit explicit: creating performance expectations for the dissertation. sterling, va: stylus. lovitts, b.e. (2008). the transition to independent research: who makes it, who doesn’t, and why. the journal of higher education, 79(3), 296-325. mackinnon, d. (1970). creativity: a multi-faceted phenomenon. in j.d. roslansky (ed.), creativity (pp. 1732). amsterdam: north-holland. marginson, s., & considine, m. (2000). the enterprise university: power, governance and reinvention in australia. cambridge & melbourne: cambridge university press. marsh, i. (2010). innovation and public policy the challenge of an emerging paradigm. canberra: australian innovation research centre. mommsen, t. (1905 [1876]). die deutschen pseudoktoren. in o. hirschfeld (ed.), theodor mommsen: reden und aufsätze (pp. 402-409). berlin: weidmann. new zealand qualifications authority. (2001). national qualifications framework. wellington: new zealand qualifications authority. nowotny, h., scott, p., & gibbons, m. (2001). re-thinking science: knowledge and the public in an age of uncertainty (p. 12). cambridge: polity press. park, c. (2005). new variant phd: the changing nature of the doctorate in the uk. journal of higher education policy and management, 27(2), 189-207. park, c. (2007). redefining the doctorate. york: the higher education academy. pope, r. (2005). creativity: theory, history, practice. london: routledge. quality assurance agency for higher education. (2011). uk quality code for higher education: doctoral degree characteristics. gloucester: quality assurance agency for higher education. rolfe, g. (2013). the university in dissent. london and new york: routledge. schumpeter, j. a. (1997 [1912]). theorie der wirtschaftlichen entwicklung: eine untersuchung über unternehmergewinn, kapital, kredit, zins und den konjunkturzyklus (9th ed.). berlin: duncker und humblot. seagram, b., gould, j., & pyke, s. (1998). an investigation of gender and other variables on time to completion of doctoral degrees. research in higher education, 39(3), 319-335. baptista et al   | f l r       67   sternberg, r.j., & lubart, t.i. (1999). the concept of creativity: prospects and paradigms. in r.j. sternberg (ed.), handbook of creativity (pp. 3-15). cambridge: cambridge university press. sternberg, r., pretz, j., & kaufman, j. (2003). types of innovations. in l.v. shavinina (ed.), the international handbook on innovation (pp. 158-169). oxford: elsevier. tinkler, p., & jackson, c. (2004). the doctoral examination process. maidenhead: open university press. walsh, e., anders, k., & hancock, s. (2013). understanding, attitude and environment: the essentials for developing creativity in stem researchers. international journal for researcher development, 4(1), 19-38. wellington, j. (2010). making supervision work for you. london: sage. wellington, j. (2013). searching for 'doctorateness'. studies in higher education, 38(10), 1490-1503. dy research (3rd ed.). thousand oaks, ca: sage. frontline learning research 1 (2013) 12 issn 2295-3159 http://dx.doi.org/10.14786/flr.v1i1.57 1 | f l r editorial frontline research in an accessible and flexible way erno lehtinen an increasing number of new scientific journals have been founded in the last few years. a big part of these new publishing forums are open-access electronic-only journals. when starting a new journal it is important to carefully think about why this new journal is needed and which kind of journal it should be. the two previously founded journals of the european association for research on learning and instruction (earli) have been very successful. learning and instruction has established its role as one of the leading journals in education and educational psychology. it publishes theoretically and methodologically strong original articles. educational research review has opened new opportunities for publishing review articles, metaanalyses and theoretical papers, and the quickly increased impact factor indicates that it is also well trusted by the research community. then why the need to start the third earli journal frontline learning research (flr)? during the almost three decades of european scientific collaboration within earli the number of researchers and the quality of research in the field of learning and instruction have rapidly increased. also, the two previously founded earli journals have extended their influence far beyond the european countries and receive high-level submissions from all over the world. because of these developments there is a much larger number of excellent manuscripts out there in the earli community than the two existing journals are able to publish. we believe that many of these research papers are worth publishing. yet, it was not only the increase of the publication pressure that led earli to the decision to supplement the existing journals by founding frontline learning research. the main aim was to develop a journal which would explicitly support innovative theoretical and methodological thinking and increase dynamics in the field. accordingly, the emphasis of the journal will be on promoting educational and learning sciences as a multidisciplinary domain, drawing from cognitive, philosophical, sociological, psychological and pedagogical theoretical paradigms. while emphasising innovative and risk-taking approaches this new journal will follow the successful policy of the other earli journals by making sure that all manuscripts go through a serious and rigorous review process. it will be a big challenge for the editors and reviewers to combine these principles, however, we believe that it is feasible. authors submitting manuscripts are the key players in creating a novel publication culture for the journal. innovative ideas and risk-taking studies have a stronger impact if they are presented in a rigorous way that enables careful evaluation of theoretical arguments, methodological details, and conclusions. e. lehtinen 2 | f l r there are deep-going changes in scientific publication policies and it was the right time for earli to take these changes into account. although libraries of big and wealthy universities provide researchers with wide on-line access to scientific journals, the journal packages available in many universities are limited. in addition, many readers of scientific publications belong to universities or research institutions providing access to scientific journals. in order to increase the impact of scientific publishing in educational and learning sciences it is important to develop open-access forums, which are available for readers independently of the organisation in which they are working. frontline learning research is for this reason an open-access journal. it means that anyone using the internet can read it for free. researchers are guaranteed more flexible access to the journals articles and the open-access format also enables them to use the articles as teaching material in face-to-face and on-line courses. furthermore, it also means that the journal is easily accessible for practitioners. frontline learning research is an electronic-only journal. in daily research work on-line versions of scientific publications are used more and more frequently and researchers seldom go to libraries to read traditional hard copies of the journals. however, we still tend to think that the existence of a traditional paper version is a prerequisite for a high reputation scientific publication. in many traditional scientific fields (e.g. physics), however, the situation is rapidly changing and electronic-only journals can be found among the most highly ranked publishing forums. in planning flr we emphasised that the electronic-only format does not only mean that there are no hard copies available, but also that it opens up new opportunities that go beyond the possibilities of traditional printed journals. the electronic delivery form provides authors with a large variety of options for dynamic presentations, such as videos, simulations, hyperlinks and animations. in other words, this electronic journal makes it possible to demonstrate novel data collection processes and alternative analysis methods in a flexible way. the electronic-only format also allows more freedom for using varying types of articles. in flr we welcome short, regular and extended manuscripts. this allows very quick communication about new findings, while also enabling in-depth description of complex empirical data. slow review processes and a long publication lag are frustrating for researchers. in flr much attention is paid to the fast review process. as an electronic journal flr is also flexible in terms of articles published in individual issues and there is no need for publication lag. a fast review and publication schedule makes it possible to have intensive scientific discussions within the journal. the papers published in this first issue of the journal demonstrate some of the ideas we have about innovative and risk-taking research. the editorial team of the frontline learning research invites earli members and researchers elsewhere to participate in this collaborative enterprise to create a new innovative publishing culture for learning research. editor-in-chief erno lehtinen editors sanne akkerman, filip dochy, nikos papadouris and jan vermunt assistant editor inneke berghmans microsoft word niculescu et al_published.docx         frontline learning research vol. 3 no. 1 (2015) 1-17 issn 2295-3159   corresponding author: alexandra c. niculescu, educational research department, maastricht university school of business and economics, po box 616, 6200 md maastricht, netherlands. email address: a.niculescu@maastrichtuniversty.nl doi http://dx.doi.org/10.14786/flr.v3i1.136   exploring the antecedents of learning-related emotions and their relations with achievement outcomes alexandra c. niculescua, dirk tempelaara, amber dailey-hebertb, mien segersa, wim gijselaersa, a maastricht university, netherlands   b park university, united states   article received 2 december 2014/ revised 6 february 2015 / accepted 9 february 2015 / available online 24 march 2015 abstract recent work suggests that learning-related emotions (lres) play a crucial role in performance especially in the first year of university, a period of transition for most students; however, additional research is needed to show how these emotions emerge. we developed a framework which links a course-contextualized antecedent – academic control in pekrun’s (2006) control value theory of achievement emotions – with generic antecedents – adaptive and maladaptive cognitions and behaviors from martin’s (2007) motivation and engagement wheel framework – to explain a classical problem: the emergence of lres in a transition period. using a large sample (n = 3451) of first year university students, our study explores these two antecedents to better understand how four lres (enjoyment, anxiety, boredom and hopelessness) emerge in a mathematics and statistics course. through the use of path-modelling, we found that academic control has a strong effect on all four lres – with the strongest impact observed for learning hopelessness and secondary, for learning anxiety. academic control, on its turn, builds on contributions from adaptive and mal-adaptive cognitions. furthermore, adaptive cognitions have an impact on learning enjoyment (positive) and on boredom (negative). surprisingly though, the maladaptive behaviors impact positively learning enjoyment and negatively learning anxiety. following this, we predicted performance outcomes in the course and found again academic control as the main predictor, followed by learning hopelessness. overall, this study brings evidence that adaptive and maladaptive cognitions and behaviours act as important antecedents of academic control, the main predictor of lres and course performance outcomes. keywords: learning-related emotions; academic control; adaptive and non-adaptive cognitions and behaviors; academic achievement; first year of university. niculescu et al | f l r     2   1. introduction the first year experience of university is known as a transition period (baker & syrik, 1999; tinto, 1997), when students are confronted with novel situations over which they have low control, yet still hold high expectations for success (perry, hladkyj, pekrun, clifton, & chipperfield, 2005). these conditions typically create negative emotional reactions towards learning in academic situations (stupnisky, perry, hall, & guay, 2012), which can lead to voluntary withdrawal at the course level (ruthig et al., 2007) and overall poor performance across all courses taken at the university (hall, perry, ruthig, hladkyj, & chipperfield, 2006). such emotions, known as achievement related emotions, can have serious consequences on how students perform within a course (pekrun, goetz, frenzel, barchfeld, & perry, 2011). this is particularly true for mathematics and statistics courses, in which students experience high levels of negative emotions, especially in in learningor homework-related situations (dettmers et al., 2011; goetz et al., 2012). within these courses, negative emotions emerge from beliefs about a low capacity to influence outcomes (frenzel, pekrun, & goetz, 2007; pekrun, 2000), referred to as appraisals of control (pekrun, goetz, titz, & perry, 2002). at the same time, students come into these courses holding generic predispositions towards learning at university, such as adaptive and maladaptive cognitions and behaviours, which will also influence their experiences within a course (martin, 2007). although we know that emotions experienced in learningor homework-related situation are particularly important for performance (leone & richards, 1989; verma, sharma, & larson, 2002), additional research is needed in the first year of university to help us understand how these emotions emerge and how they can be influenced (putwain, sander, & larkin, 2013). such information can inform the design of educational interventions to create “emotionally sound” (astleitner, 2000) learning environments which can potentially improve academic achievement. the present study focuses on two different antecedents of achievement learning-related emotions: 1) the course contextualized antecedents (appraisal of control) and, 2) the generic antecedents towards learning at university (adaptive and maladaptive cognitions and behaviours). both antecedents need to be integrated, as they are complementary in providing information about the emergence of emotions in a course setting. direct antecedents are necessary for explaining the emergence of distinct emotions at a course level and distal antecedents can explain the individual differences that arise in the emergence of these emotions. finally, relations and implications for academic achievement are further discussed. 1.1 theoretical framework over the past twenty years we have seen a growing interest in, and increased research that explores the role of achievement emotions across various educational contexts and course settings. such research investigates different functions of academic emotions within a course, such as their effects on self-regulation (artino jr. & jones ii, 2012), learning engagement (ainley & ainley, 2011), learning choices (tempelaar, niculescu, rienties, gijselaers, & giesbers, 2012) and achievement (dettmers et al., 2011; goetz, frenzel, pekrun, & hall, 2006; goetz et al., 2012). the transition required in the first year of university involves several challenges which may include perceived competition and pressure to perform – both demanding heightened self-reliance and autonomy (perry, hladkyj, pekrun, & pelletier, 2001). since students are expected to engage in more individual self-study, the importance of achievement emotions in individual learningor homeworkrelated situations (as compared to the classroom setting, for example) is particularly important. these emotions are referred to in the literature as achievement learning-related emotions (pekrun, 2000). at the same time, a closer investigation of students’ experiences is necessary to clarify how learningrelated emotions (lres) emerge at the course level. 1.1.1 achievement emotions achievement emotions are defined as “emotions that are directly linked to achievement activities and outcomes” (pekrun et al., 2011, p. 37). in the control-value theory of achievement emotions (cvtae; pekrun, 2006), emotional experiences have a situational context, meaning that they can be niculescu et al | f l r     3   experienced in different academic situations within a course: 1) being in class, 2) taking tests and exams and, 3) studying outside of class (while learning or when preparing homework). of particular interest are the emotions experienced in learning-related situations as students seem to experience the most unpleasant emotions when compared with other academic situations, such as learning in the classroom (leone & richards, 1989). indeed, according to the cvtae, first year university students experience a variety of learning-related emotions, whether the emotions are positive or negative. 1.1.2 learning – related emotions and their course contextualized antecedents according to the control-value theory of achievement emotions (cvtae; pekrun, 2006), discrete learning-related emotions (lres) arise from the appraisal of achievement activities and outcomes. emotions that result from such appraisals can indirectly influence achievement outcomes. there are two dimensions of appraisals: control and value. the appraisal of control refers to a student’s belief about whether he/she has control over learning activities/outcomes; the appraisal of value describes the subjective value attributed to these activities/outcomes. these appraisals are considered direct antecedents of lres and are acquired at the course level (pekrun, 2006). control appraisals describe the perceived controllability of one’s own competency towards achievement activities and outcomes; as a general rule, low and high levels of control appraisals influence emotions differently (pekrun, 2000). for instance, low control leads to an increased level in negative emotions (e.g., learning anxiety) and a more elevated level of control favours a heightened experience of positive emotions (such as learning enjoyment). empirical evidence shows that the appraisal of control longitudinally relates to emotions (perry et al., 2001; perry et al., 2005), as well as to subsequent academic achievement in the first year of university (hall et al., 2006; ruthig et al., 2008; stupnisky et al., 2012). for instance, perry et al. (2001) found that students who reported higher levels of primary control also felt less bored (-.48) and less anxious (-.35) towards the course, and obtained higher final grades (.27). similar relations are shown by hall et al. (2006): correlations between primary control and several emotions (anger, regret, happiness and pride) are in the range of -.27 to .24; primary control relates positively to the final course grade (.21) as well as to cumulative gpa (.25). overall, this correlational evidence suggests relations between primary control, emotions and performance which are of moderate size (cohen, 1992). there are also documented gender differences in the beliefs students hold towards their abilities to perform in mathematics (female students tend to generally believe they are not very good at mathematics), with implications on how the two genders feel about this subject (robinson & clore, 2002; frenzel et al., 2007). finally, the implications of studying course specific antecedents of lres is relevant when explaining the development of emotions over time and, indirectly, for understanding their consequences on achievement. 1.1.3 generic antecedents of learning-related emotions there are also more general expectancies and predispositions towards learning at university students already hold when entering a course, which can be considered generic antecedents of lres and achievement. students enter a new course holding background characteristics (intelligence, personality, high school gpa etc.) but also possessing a set of adaptive and impeding cognitions, and adaptive and impeding behaviors, towards learning in the new setting of university (martin, 2007). therefore, we applied the ‘motivation and engagement wheel’ framework of martin (2007, 2009) as a model for distal antecedents of learning-related emotions (lres). the motivation and engagement wheel breaks down all motivation and engagement concepts into four categories: adaptive cognitions, adaptive behaviors, impeding cognitions, and maladaptive behaviors. these four categories each consist of two or three sub-dimensions. for adaptive cognitions, the dimensions consist of self-belief, valuing school, and learning focus. student’s confidence to do well in university, their belief that learning will be useful and relevant, and their interest in learning new topics/developing new skills, all contribute to various academic outcomes (martin, 2011). furthermore, the adaptive behavioral dimensions include persistence, planning, and task management. to date, a study of martin and marsh (2006) shows that self-efficacy, control, planning, low anxiety, and persistence predict enjoyment and class participation. conversely, the impeding or deactivating antipodes of the cognitions (that obstruct learning rather than enhance it) include anxiety, failure avoidance and uncertain control. the maladaptive behaviors are twofold: self-handicapping and disengagement. in turn, self-handicapping (as a niculescu et al | f l r     4   disruptive behaviour) can predict negative academic outcomes (martin, marsh, & debus, 2001). although the experience of the adaptive and mal-adaptive cognitions and behaviors can differ on average for female and male students (liem & martin, 2012), the concepts operating in this motivation and engagement wheel represent generic orientations that are relatively stable over contexts (martin, 2009). for this reason, in pekrun’s theory, such generic orientations can be integrated as distal antecedents of lres. although it may appear that some of the concepts (e.g. self-belief/efficacy, persistency and control) from the “motivation and engagement wheel” are closely related to the appraisal of control in the cvtae, it is important to ensure clarity (distinction) between them: while the distal antecedents are more trait-type of constructs, the direct antecedent (appraisal of control) is a subject specific type of appraisal. overall, the motivation and engagement concepts play an important role in students’ cognitive appraisals, in their emotions during learning, and in achievement outcomes (martin & marsh, 2006; martin, 2011). figure 1 summarizes the conceptual model used in our study. figure 1. the conceptual framework of the study to sum-up, the added value of integrating both direct and distal antecedents into one framework is to explain: 1) the emergence of distinct emotions through direct antecedents, and 2) through distal antecedents, the individual differences that arise in learning emotions when students enroll in a course. 1.1.4 learning – related emotions and academic performance while other settings have been extensively studied, such as the exam situation, few studies have investigated situations outside the class (putwain, larkin, & sander, 2013; schutz & pekrun, 2006; trautwein et al., 2009). recent research discusses students’ emotional experiences during individual learning activities such as mathematics homework (dettmers et al., 2011; goetz et al., 2012) in which the assignments are considered “emotionally charged activities” (dettmers et al., 2011, p. 25). in the homework situation students seem to experience the most unpleasant emotions when compared with other academic situations (leone & richards, 1989; verma, sharma, & larson, 2002). furthermore, learning – related emotions (lres) are of particular interest, as they demonstrate a strong relationship with achievement outcomes. while it is already known that positive emotions have a positive impact on academic performance (dettmers et al., 2011; pekrun et al., 2002), by focusing on the experience of unpleasant emotions during homework, dettmers et al. (2011) demonstrates how elevated anxiety and boredom levels shape effort and disengagement in study, to predict negative achievement in mathematics. considering the transition represented by the first year of university, more evidence is needed – particularly in this period – about niculescu et al | f l r     5   students’ emotional experiences in learning situations. to our best knowledge, only few studies (putwain, sander, et al., 2013) have addressed this issue in the first year of university context. to our best knowledge, we found only one study (tempelaar et al., 2012) which investigates how these emotions emerge and influence learning outcomes in the setting of an undergraduate introductory mathematics or statistics course. the present study builds further on the tempelaar et al. (2012) work to look how distinct lres emerge from course contextualized and generic antecedents and further, how they influence achievement outcomes in a first year university mathematics and statistics course. 1.2 research questions and hypotheses we have asked the following research questions: rq1. what role do distal and direct antecedents play in the development of lres? rq2. to what extent can the direct and distal antecedents together explain student performance at the course level? furthermore, we hypothesize: h1. the distal antecedents will have effects on both control appraisals and lres, with differential roles for adaptive and maladaptive distal antecedents. h2. the direct antecedents, control appraisals, will have an effect on lres. this effect will be different for positive versus negative (or neutral) lres. the control appraisals will influence positively enjoyment and negatively anxiety, boredom and hopelessness. h3. distal antecedents, direct antecedents and lres all explain student performance in the course. research hypotheses are graphically depicted in the figure 2, demonstrating the a priori structural model. to facilitate the reading of this conceptual model, all three negative emotions are taken together, as well as the two adaptive cognitions and behaviours, and the two maladaptive ones. figure 2. the hypothesized structural model niculescu et al | f l r     6   the hypothesized structural model expresses that adaptive cognitions and behaviours, academic control, positive emotion, and performance are all hypothesized to be positively related, whereas maladaptive cognitions and behaviours and negative emotions are hypothesized to be positively related amongst them, but negatively related with the first subset of variables. not explicit in this conceptual model is that distal antecedents are represented by second order factors of the motivation and engagement instrument, however allowing for path estimates being different from factor loadings. 2. method 2.1. sample and setting the participants were 3451 freshmen (19 years old on average, 62.5% male) enrolled over four consecutive academic years (10/11, 11/12, 12/13, 13/14) in a business and economics program at a european university. most students had an international background, a vast majority (77.4%) holding an international education diploma and one third of the sample had been previously educated in the field of mathematics (mathematical major specialization). the setting was a compulsory introduction course to mathematics and statistics, scheduled in the first term of the academic year. it had a duration of eight weeks out of which, seven weeks were scheduled for education and the last week was reserved for exams. 2.2. procedure in week two of the course students completed an online questionnaire concerning their adaptive and maladaptive cognitions and behaviors towards learning at university in general. in week four participants completed another online questionnaire, this time about their control appraisals and lres regarding the specific subject of the course. the timing was chosen to capture sufficient experience with the learning activities. in weeks three, five and seven of the course, voluntary mathematics and statistics quizzes were planned which, if performed successfully, added a bonus score to the final course grade. every week, students were expected to prepare homework assignments which, if solved, granted students bonus points. in week eight of the course, students participated in the written exam. all students included in this study provided informed consent for the data collected by means of online questionnaires and for use of their study results. 2.3. measures and variables we measured learning-related emotions through four scales: enjoyment, anxiety, boredom and hopelessness, of the achievement emotions questionnaire (aeq; pekrun et al., 2011). the enjoyment scale (10 items, e.g. “i enjoy accruing new knowledge”), anxiety scale (11 items, e.g. “i get tense and nervous while studying”), boredom scale (11 items, e.g. “the material bores me to death”) and hopelessness scale (11 items, e.g. “i feel hopeless when i think about studying”) were slightly re-phrased to match the specific situation of our course. for reasons of consistency in our research, all items were answered on a 7-point likert scale (1 = ‘completely disagree’ and 7 = ‘completely agree’). control appraisals were measured with the academic control scale (acs) of perry et al. (2001). academic control as described by perry et al. is a domain, course-specific measure of college students’ beliefs. the scale is composed of eight items, each answered on a 7-point scale (1 = ‘strongly disagree’ and 7 = ‘strongly agree’), e.g. “i have a great deal of control over my academic performance in this course”. niculescu et al | f l r     7   adaptive and maladaptive cognitions and behaviors were measured with the motivation and engagement scale (mes; martin, 2007). the mes consists of four scales and eleven subscales subsumed under the four scales. the adaptive cognition scale is composed of three sub-scales: self-belief (e.g. “if i try hard, i believe i can do my university work well”), valuing school (e.g. “learning at university is important for me”) and learning focus (e.g. “i feel very pleased with myself when i really understand what i’m taught at the university”). the second scale, adaptive behavior contains the following subscales: persistence (e.g. “if i can’t understand my university work at first, i keep going over until i do”), planning (e.g. “if i start an assignment i plan out how i am going to do it”) and study management (e.g. “when i study, i usually study in places where i can concentrate”). the third sub-scale, maladaptive (impeding) cognition includes the anxiety (e.g. “when exams and assignments are coming up, i worry a lot”), failure avoidance (e.g. “often the main reason i work at university is because i don’t want to disappoint others”) and uncertain control (e.g. “i am often unsure how i can avoid doing poorly at university”) sub-scales. finally, maladaptive behavior includes the self-handicapping (e.g. “sometimes i don’t study very hard before exams so i have an excuse if i don’t do as well as i hoped”) and disengagement (e.g. “i often feel like giving up at university”) sub-scales. academic achievement was measured with a performance portfolio consisting of three separate parts: mathperformance, statsperformance and bonusperformance. first, the two performance outcomes mathperformance and statsperformance were assessed in a final written exam which covered a mathematics component and a statistics component, graded separately. second, the bonusperformance represented the sum of bonus scored in quizzes and homework. quizzes, although optional, were available for both mathematics and statistics in an online format. some further bonus could be achieved by doing weekly homework, containing assignments for mathematics and statistics. finally, the three separate parts were summed in the qmperformance which represented the total score for the course. we accounted for any potential influences coming from gender (female and male) and level of introductory mathematics education (distinguishing between two tracks, mathmajor and mathminor) as control variables. 2.4. statistical analyses as a preliminary step in the analysis, the four cohorts were checked upon invariance of mean levels and correlation structures. next, beyond descriptive analyses, this study applies structural equation modeling. models were estimated with lisrel (version 8.8) using maximum likelihood (ml) estimation. to prevent capitalization on chance, rather conservative model building rules were adapted: p-values of 1% or less were required as a cutoff value for significance for the adoption of any structural path; correlated traits were only allowed for variables measured by the same instrument. as measurement model for the motivation and engagement constructs, a second order confirmatory factor model was postulated, with second the order factors adaptive cognitions, adaptive behaviors, impeding cognitions, and maladaptive behaviors (see martin, 2007). we identified both second order and first order latent factors for motivation and engagement variables, and in order to derive a parsimonious model, we based the relationships with lre’s and control appraisal on the second order factors. however, we allowed for differentiated effects of first order factors, by testing if first order factors would add predictive power to the already included second order factors. we report the chi-square and degrees of freedom values, the comparative fit index (cfi), the nonnormed fit index (nnfi, also known as tli) and the root mean square error of approximation (rmsea) as indicators of goodness of fit. hu and bentler (1999) suggested for cfi/tli values larger than .90 for a satisfactory fit and for rmsea values should not exceed .08 and preferably be .06 or lower. niculescu et al | f l r     8   3. results 3.1. preliminary analysis we checked the assumptions of normality through spss 22. values of skewness and kurtosis were in the expected range of chance fluctuations in that statistic for all scales. to make the performance measures equivalent over cohorts, we transformed exam scores into cohort specific z-scores. these transformed variables were used in all subsequent analyses. we provide descriptive statistics and reliabilities (table 1) – as well as measures for differences between gender and prior education track. all analyses were based on a subset of students for which background characteristics, lres variables and performance data were all available (3355 of the 3451 students, 97%). table 1. means (m), standard deviations (sd), cronbach’s alpha and test statistics for gender and prior mathematics education differences: t-value and cohen d-value m sd α gender difference math prior education t –value d–value t –value d–value adaptive cognitions: self-belief 5.82 0.73 0.73 1.08 0.04 2.86** 0.10 valuing school 5.84 0.67 0.67 -5.15 *** -0.18 1.65 0.06 learning focus 5.95 0.73 0.80 -9.65*** -0.34 -0.14 0.00 adaptive behaviors: planning 4.79 0.99 0.73 -9.73*** -0.34 0.15 0.01 study management 5.56 0.89 0.74 -9.04*** -0.32 -2.66* -0.09 persistence 5.34 0.85 0.78 -6.79*** -0.24 1.00 0.04 impeding cognitions: anxiety 4.50 1.27 0.83 -16.12*** -0.57 -6.07*** -0.21 failure avoidance 2.57 1.19 0.83 0.90 0.03 -1.45 -0.05 uncertain control 3.45 1.18 0.80 -5.418*** -0.19 -4.58*** -0.16 maladaptive behaviors: self-handicapping 2.43 1.08 0.81 5.68*** 0.32 -0.45 -0.02 disengagement 1.97 0.90 0.74 7.09*** 0.25 1.20 0.04 academic control 5.26 0.89 0.82 3.868*** 0.14 13.68*** 0.48 learning-related emotions anxiety 3.85 1.11 0.91 -11.41*** -0.40 -15.13*** -0.53 boredom 2.94 1.13 0.93 7.65*** 0.27 -4.44*** -0.16 hopelessness 3.01 1.22 0.94 -7.18*** -0.25 -17.08*** -0.60 niculescu et al | f l r     9   note: performance scores are normalized scores; concerning gender differences, a negative score represents female students; a positive score in the differences in previous math education represents math major. 3.2. bivariate correlations bivariate correlations are reported in table 2. due to the large number of manifest variables, the correlation table contains scale values rather than individual item values for the survey data based on the aeq, acs and mes instruments. the four performance measures are manifest variables too. table 2. correlations of scales of the aeq, asc, and mes instruments (1-16) and performance measures (17-20) enjoyment 4.11 0.92 0.85 -0.55 -0.02 10.40*** 0.37 performance outcomes math performance -1.03 -0.04 20.47*** 0.72 stats performance 1.68 0.06 11.87*** .042 bonus performance -6.70*** -0.24 11.73*** 0.41 qm performance -1.00 -0.04 18.41*** 0.65 niculescu et al | f l r     10   the signs of the bivariate correlations express the divide into adaptive and maladaptive constructs. adaptive cognitions and behaviours are positively correlated to 1) academic control, 2) the positive lre of enjoyment, and to 3) performance measures. correlations with performance measures are however weak, and not fully consistent for study management. correlations between academic control and enjoyment versus performance measures are stronger, and consistently positive. a reverse pattern exists for the maladaptive cognitions and behaviours: positively correlated to negative lres, negatively correlated to academic control, enjoyment and performance measures. however, within the motivation and engagement variables, anxiety is unique in that it acts as a maladaptive cognition dimension in relation to lres and performance. yet, it correlates weakly with other maladaptive mes variables, as well as with the adaptive constructs (learning focus, study management, and planning) but to a lesser degree. 3.3. structural models separate structural equation models were estimated for each of the four performance constructs, each of them having identical relationships between the motivation and engagement latent constructs, and the latent constructs based on lres and academic control. figure 3 contains the diagram of the structural part of the structural equation model (leaving out the measurement parts of the lre, academic control and motivation and engagement constructs for reasons of readability), having only the mathematics score in the exam as performance construct. it is relevant to mention that structural models for the other performance constructs deviate only in terms of the equation predicting the performance constructs, and these equations are provided at the end of this section, in table 3. all regression paths are expressed as standardized betas. structural models were estimated in two multi-group specifications: on the basis of gender, and on the basis of prior mathematics track in high school. both result in a rejection of invariant latent means, fully in line with the outcomes of the descriptive analyses: differences in mean scales between female and male students, and between students educated in the math major, versus math minor track, also show up as significant differences in latent means. however, at the stringent .01 significance level, no rejection of the hypothesis of invariant estimates in the variance-covariance structure was found: the structural relations appear to be the same for the subgroups. fit indices of both two-group models were nearly identical, with χ2 = 26,424 and 25,946 respectively, and identical measures for df = 9,030, cfi = .98, nnfi = .98, rmsea = .39, 95% ci rmsea = (.38, .39), for the structural models including the mathematics score as performance measure. niculescu et al | f l r     11   figure 3. path diagram of structural part with standardized estimates 3.3.1. testing hypotheses in h1 we expected that the distal antecedents will have effects on both control appraisals and lres. in agreement with the cvtae (pekrun, 2006), academic control plays a central role in the antecedentconsequence relationship of adaptive and mal-adaptive cognitions and behaviours, and lres. academic control is a pure cognitive construct: it builds on contributions from adaptive and maladaptive cognitions, excluding any behavioural influence. impeding cognitions as a whole have a strong negative impact on academic control. this is explained by the fact that impeding cognition is most strongly reflected by uncertain control (.76). at the same time, that effect is attenuated by the two paths of anxiety (.56) and failure avoidance (.68), which constitute the first order factor of impeding cognition. since behaviours, both of adaptive and maladaptive type, do not contribute to academic control, the relationships between behaviours and emotions are only direct ones. the paths originating from adaptive cognitions are fully in line with the hypotheses: positive impact on enjoyment (.13), negative impact on boredom (-.24). however, the maladaptive behaviours do play a rather remarkable role. although bivariate relations are all in the hypothesised direction (positive with negative emotions, negative with the positive emotion), within the full structural model, the additional impact of maladaptive behaviours on lres is positive for enjoyment (.40), whilst its impact on anxiety is negative (-.20). this is the resultant of a multiple relationship with colinearity amongst maladaptive cognitions and behaviours: for given levels of academic control and maladaptive niculescu et al | f l r     12   cognitions, the additional effect of maladaptive behaviours is adverse to the bivariate effect. gender differences may also contribute to these adverse effects: male students score much higher than female students on maladaptive behaviours, but at the same time demonstrate less emotion of anxiety and hopelessness. in h2 we assumed that control appraisals will influence positively enjoyment and negatively anxiety, boredom and hopelessness. as hypothesized and already shown in the bivariate relations analysis, academic control has indeed a strong effect on the four lres. these effects are positive for enjoyment and negative for all other three emotions. the strongest effect is observed for hopelessness (-.65). then, enjoyment and academic control and boredom and academic control respectively, relate rather weaker (.32, -.24). the relation between academic control and anxiety (-.54) is rather strong and has a negative direction: the students in our sample are on average high in academic control (m=5.26) which might result on a rather lower level of anxiety (m=3.85). in h3 we specified that the distal antecedents, direct antecedents and lres all explain student performance in the course. we notice a consistent and dominant role of academic control on performance. then, a secondary role of hopelessness, with a crucial exception: for the bonus score (which is composed of the digital homework and quizzes). this result is very plausible: for students high in hopelessness, it is rational to allocate relative high levels of time and effort to learning in the digital tool, given its intensive scaffolding. since the share of the bonus is much smaller in the overall score than the share of math and stats exam scores, in the overall score the negative impact of hopelessness is back. a remarkable role is played by enjoyment: it impacts performance, as expected, positively for math; nevertheless, it impacts performance negatively in stats. again, this finding can be regarded as very plausible, due to the different nature of mathematics and statistics education. students who like mathematics a lot tend to prefer it over statistics. evidence for this claim is indirect: t-test for independent groups indicates that students from the ‘math major’ track score different in enjoyment, hopelessness and anxiety, from students from the ‘math minor’ track. european ‘math major’ tracks focus on mathematics only, not on stats, and very often contain less statistics subjects than the ‘math minor’ track. since enjoyment has opposite impact on math and stats performance, it is no surprise that it drops out as explanatory variable in the total score, qm i performance. lastly, self-handicapping enters as explanatory variable in one performance category: bonus. again, this is very plausible: it requires discipline to do all the homework, so students high in self-handicapping will underperform. since bonus has only a small share in the total score, it is not visible for qm i performance. for a more detailed overview of each’s variable contribution in each of the four performance outcomes, the relations between these variables are provided in the equations below (coefficients for each independent variable are expressed in standardized betas): mathperfomance = 0.32*academiccontrol + 0.06*enjoyment – 0.10*hopelessness statsperfomance = 0.27*academiccontrol – 0.10*enjoyment – 0.13*hopelessness bonusperfomance = 0.24*academiccontrol + 0.09*enjoyment – 0.16*selfhandicapping qm1perfomance = 0.33*academiccontrol – 0.13*hopelessness 4. discussion recent work suggests that learning-related emotions (lres) play a crucial role in performance especially in the first year of university, a period of transition for most students; however, additional research is needed to show how these emotions emerge. to explain this classical problem, we developed a framework which links two types of antecedents of lres: 1) the course-contextualized academic control in the control value theory of achievement emotions (pekrun, 2006) as a direct antecedent and 2) the generic adaptive and maladaptive cognitions and behaviors from the motivation and engagement wheel framework (martin, niculescu et al | f l r     13   2007) as distal antecedents. we used this framework to predict learning achievements in a mathematics and statistics course. the main findings of this study bring forth the emergence of four distinct lres (enjoyment, anxiety, boredom and hopelessness) and the fact that they standalone from students’ individual performance. such findings are reassuring: although lres are important, they are not blocking students to perform academically. more importantly, the relations between lres and performance are rather weak when taking into account their antecedents. especially, in the mediational model comprising academic control, lres and performance, we see that academic control plays a central role in the development of the four lres investigated in our study as well as for what regards the performance outcomes in the course. the direct relationship between appraisals and performance strongly dominates the indirect relationship through lres. next, academic control has a strong effect on all of the four lres with the strongest impact observed for hopelessness and secondary, for anxiety. the model explaining the four lres is again of mediational type. beyond the indirect effect through the control appraisal, there are direct effects from the four second order motivation and engagement factors to the lres. in this part of the model, direct and indirect effects rather well balance in size. academic control, on one hand, builds on contributions from adaptive and mal-adaptive cognitions solely, where the main impact is explained by the uncertain control dimension of impending cognitions. on the other hand, adaptive cognitions have a positive impact on enjoyment and a negative one on boredom. where impeding cognitions confirm the hypotheses of positive relationship with the negative emotions, surprisingly though, the maladaptive behaviours impact the lres positively for enjoyment and negatively for anxiety. it seems that amongst students scoring high on maladaptive behaviour (amongst them an overrepresentation of male students), there exists a dislike of the learning activities (increased levels of boredom), but not of the learning content: high enjoyment, low anxiety. with respect to the implications on performance outcomes, the most consistent role is played by academic control; this is followed by hopelessness (with the exception played for bonus as detailed earlier). at last, an important role is also played by enjoyment: it has opposite impact for math (positive) and stats (negative) performance. our findings are consistent with earlier research on the central role of control appraisals in the emergence of achievement emotions (pekrun et al., 2002; perry et al., 2001) as well predicting performance at the course level (hall et al., 2006). this study also provides support for the positive relations between impeding cognitions and negative emotions (martin & marsh, 2006). conversely, it extends such evidence by showing maladaptive behaviours influencing positively enjoyment and negatively anxiety. we therefore extend on the control value theory of achievement emotions (pekrun, 2006) by integrating the distal antecedents of emotions from the motivation and engagement wheel framework (martin, 2007). most notably, to the knowledge of the authors, the study is the first of its kind in using an integrated framework to ultimately explain achievement outcomes in the first year at university. we have provided a new approach to understand students’ emotional experiences when they first enter a university study. in this respect, the two theories are complementary: on one side our results are an empirical validation of the cvtae; on the other side, the concepts operating in the mes could provide practical solutions on how to facilitate educational change in the classroom by using the influence these variables have in the experience of emotions. 4.1. additional findings although not the main focus of this study, we find interesting gender patterns and effects of prior education. they are described separately. first, in our descriptive analysis, we find gender patterns that match earlier research (martin, 2007). females score significantly higher on all adaptive dimensions, with one exception: self-belief, where no significant difference is found. statistical significance of gender differences is however inflated by the large sample size; effect sizes are in the .2 to .4 range, therefore, small in size. with regard to the maladaptive dimension, we find the same pattern as described by martin (2007): maladaptivity expresses itself stronger in the form of impeding cognitions in females, but in the form of maladaptive behaviours in males. the gender effect in anxiety is not only significant, but also medium in size, again in line with previous research (preckel, goetz, pekrun, & kleine, 2008). this divide between the niculescu et al | f l r     14   cognitive and behavioural aspects of maladaptivity repeats itself in the lres. it is in boredom, the behavioural aspect of neutral emotions (see pekrun et al., 2002), that males score higher than females, and in the cognitive aspects of the negative lres, anxiety and hopelessness, that females score higher. the last gender effect refers to academic control, where male students score higher than female students, in line with outcomes of selfconcept research (frenzel et al., 2007). the second effect we investigated refers to prior education: having been educated in high school in an advanced, rather than a basic mathematics track. the impact on the generic dimensions of motivation and engagement are quite small, as to be expected. students from the advanced track are higher in self-belief, but lower in study-management and anxiety; effect sizes are however very small. these findings contrast the impact of prior education on the lres and academic control: the largest effect size, .6, is observed for hopelessness; in rest we find medium size effects. these effects point in the direction that students from the advanced track are higher in enjoyment and academic control and lower in anxiety, hopelessness, and boredom. 4.2. limitations using a large sample, our study proposed a framework linking control appraisals (as direct predecessor) with motivation and engagement concepts (as distal predecessors) in an attempt to better explain the emerge and consequences of lres in a first year undergraduate mathematics and statistics course. however, we point out two limitations. first of all, our lres measures (assessed through self-reports) rely heavily on retrospective beliefs about emotions, which make them subject to the same biases as the self-appraisals (robinson & clore, 2002). at the same time, self-reports still remain the most reliable measure (zeidner, 1998) and, for that reason the most extensively used approach, which is able to capture in a non-invasive manner students’ emotional experiences in an educational setting. second, while in the present study we tried to answer how emotions emerge in an introductory course, an important question for future studies remains: how students’ emotions change over different courses in the first year at university. future work should employ the use of a longitudinal design, over a period of time and different course subject which could cover ideally an entire year of study. 4.3. recommendations for further research some general recommendations should be outlined. first, our results showed that amongst students scoring high on maladaptive behaviour, there is a dislike of the learning activities (increased levels of boredom), but not of the learning content: high enjoyment, low anxiety. we propose that they solve this tension by designing their own learning trajectories, participating at a lower level in homework and quizzes (as evident from the role of self-handicapping in explaining the bonus performance), and prepare independently for the exam. we mentioned earlier that the evidence gained in our study could potentially inform the design of educational interventions to improve academic achievement while, at the same time, support building emotionally sound learning environments. in this respect, a first aspect to consider would be that any educational interventions in the classroom should foster students’ sense of competency towards the specific learning activities required in a mathematics and statistics course. if such progress is acquired, then reinforcing – by means of feedback – the certainty of control over the activities and outcomes in which students engage is key. increasing students enjoyment and decreasing their hopelessness seems intuitive, still these measures should be regarded in context together with the factors from which they emerge, the maladaptive behaviours. if emotions are more difficult, and less desirable, to influence directly, addressing students maladaptive behaviours could be a reasonable solution. niculescu et al | f l r     15   4.4. conclusion it can be concluded from our study that next to personal factors that bring their contribution (especially in the development of academic control), it is the contextual experience in a course that shapes students’ emotional experiences and performance. besides all other known factors, emotions seem to play a central role in any learning process as an input and as a major educational outcome next to academic performance (pekrun, frenzel, goetz, & perry, 2007). therefore, learning about the factors that play a role in how these emotions develop – and how, in turn, they further influence academic outcomes – is crucial. good education should also care about how students feel and not only how well they can perform academically. keypoints academic control impacts strongest learning hopelessness adaptive cognitions impact both learning enjoyment and boredom maladaptive behaviours impact learning enjoyment and anxiety achievement outcomes are mainly predicted by academic control and learning hopelessness references ainley, m., & ainley, j. (2011). student engagement with science in early adolescence: the contribution of enjoyment to students’ continuing interest in learning about science. contemporary educational psychology, 36(1), 4–12. doi:10.1016/j.cedpsych.2010.08.001 artino jr., a. r., & jones ii, k. d. (2012). exploring the complex relations between achievement emotions and self-regulated learning behaviors in online learning. emotions in online learning environments, 15(3), 170–175. doi:10.1016/j.iheduc.2012.01.006 astleitner, h. (2000). designing emotionally sound instruction: the feasp-approach. instructional science, 28(3), 169–198. baker, r. w., & siryk, b. (1999). sacq student adaptation to college questionnaire (2nd ed.). los angeles: western psychological services. cohen, j (1992). a power primer. psychological bulletin, 112(1), 155-159. http://dx.doi.org/10.1037/00332909.112.1.155 dettmers, s., trautwein, u., lüdtke, o., goetz, t., frenzel, a. c., & pekrun, r. (2011). students’ emotions during homework in mathematics: testing a theoretical model of antecedents and achievement outcomes. students’ emotions and academic engagement, 36(1), 25–35. doi:10.1016/j.cedpsych.2010.10.001 frenzel, a. c., pekrun, r., & goetz, t. (2007). girls and mathematics--a “hopeless” issue? a control-value approach to gender differences in emotions towards mathematics. european journal of psychology of education, 22(4), 497–514. goetz, t., frenzel, a. c., pekrun, r., & hall, n. c. (2006). the domain specificity of academic emotional experiences. journal of experimental education, 75(1), 5–29. goetz, t., nett, u. e., martiny, s. e., hall, n. c., pekrun, r., dettmers, s., & trautwein, u. (2012). students’ emotions during homework: structures, self-concept antecedents, and achievement outcomes. learning and individual differences, 22(2), 225–234. doi:10.1016/j.lindif.2011.04.006 hall, n. c., perry, r. p., ruthig, j. c., hladkyj, s., & chipperfield, j. g. (2006). primary and secondary control in achievement settings: a longitudinal field study of academic motivation, emotions, and performance1. journal of applied social psychology, 36(6), 1430–1470. niculescu et al | f l r     16   hu, l., & bentler, p. m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives. structural equation modeling:a multidisciplinary journal, 6(1), 1–55. doi:10.1080/10705519909540118 ibm corp. released 2013. ibm spss statistics for windows, version 22.0. armonk, ny: ibm corp. jöreskog, k.g., sörbom, d. (1996). lisrel 8 user's reference guide. scientific software international: chicago. leone, c., & richards, h. (1989). classwork and homework in early adolescence: the ecology of achievement. journal of youth and adolescence, 18(6), 531–548. doi:10.1007/bf02139072 liem, g. a. d., & martin, a. j. (2012). the motivation and engagement scale: theoretical framework, psychometric properties, and applied yields. australian psychologist, 47(1), 3–13. doi:10.1111/j.1742-9544.2011.00049.x martin, a. j. (2007). examining a multidimensional model of student motivation and engagement using a construct validation approach. british journal of educational psychology, 77(2), 413–440. doi:10.1348/000709906x118036 martin, a. j. (2009). motivation and engagement across the academic life span: a developmental construct validity study of elementary school, high school, and university/college students. educational and psychological measurement, 69(5), 794–824. doi:10.1177/0013164409332214   martin, a. j. (2011). holding back and holding behind: grade retention and students’ non-academic and academic outcomes. british educational research journal, 37(5), 739–763. doi:10.1080/01411926.2010.490874 martin, a.j., marsh, h. (2006). academic resilience and its psychological and educational correlates: a construct validity approach. psychology in the schools, 43(3), 267 281. martin, a.j., marsh, h.w., debus, r.l. (2001). self-handicapping and defensive pessimism: exploring a model of predictors and outcomes from a self-protection perspective. journal of educational psychology, 93, 87–102. pekrun, r. (2000). a social-cognitive, control-value theory of achievement emotions. in j. heckhausen (ed.), motivational psychology of human development: developing motivation and motivating development. (pp. 143–163). new york, ny us: elsevier science. pekrun, r. (2006). the control-value theory of achievement emotions: assumptions, corollaries, and implications for educational research and practice. educational psychology review, 18(4), 315–341. doi:10.1007/s10648-006-9029-9 pekrun, r., elliot, a. j., & maier, m. a. (2006). achievement goals and discrete achievement emotions: a theoretical model and prospective test. journal of educational psychology, 98(3), 583–597. doi:10.1037/0022-0663.98.3.583 pekrun, r., goetz, t., frenzel, a. c., barchfeld, p., & perry, r. p. (2011). measuring emotions in students’ learning and performance: the achievement emotions questionnaire (aeq). contemporary educational psychology, 36(1), 36–48. doi:10.1016/j.cedpsych.2010.10.002 pekrun, r., goetz, t., titz, w., & perry, r. p. (2002). academic emotions in students’ self-regulated learning and achievement: a program of qualitative and quantitative research. educational psychologist, 37(2), 91–106. perry, r. p., hladkyj, s., pekrun, r. h., clifton, r. a., & chipperfield, j. g. (2005). perceived academic control and failure in college students: a three-year study of scholastic attainment. research in higher education, 46(5), 535–569. doi:10.1007/s11162-005-3364-4 perry, r. p., hladkyj, s., pekrun, r. h., & pelletier, s. t. (2001). academic control and action control in the achievement of college students: a longitudinal field study. journal of educational psychology, 93(4), 776.   preckel, f., goetz, t., pekrun, r., & kleine, m. (2008). gender differences in gifted and average-ability students: comparing girls’ and boys’ achievement, self-concept, interest, and motivation in mathematics. gifted child quarterly, 52(2), 146–159. niculescu et al | f l r     17   putwain, d. w., larkin, d., & sander, p. (2013). a reciprocal model of achievement goals and learning related emotions in the first year of undergraduate study. contemporary educational psychology, 38(4), 361–374. doi:10.1016/j.cedpsych.2013.07.003 putwain, d. w., sander, p., & larkin, d. (2013). using the 2 × 2 framework of achievement goals to predict achievement emotions and academic performance. learning and individual differences, 25(0), 80–84. doi:10.1016/j.lindif.2013.01.006 robinson, m. d., & clore, g. l. (2002). belief and feeling: evidence for an accessibility model of emotional self-report. psychological bulletin, 128(6), 934–960. doi:10.1037//0033-2909.128.6.934 ruthig, j. c., perry, r. p., hladkyj, s., hall, n. c., pekrun, r., & chipperfield, j. g. (2007). perceived control and emotions: interactive effects on performance in achievement settings. social psychology of education, 11(2), 161–180. doi:10.1007/s11218-007-9040-0 schutz, p. a., & pekrun, r. (2007). introduction to emotion in education. in p. a. schutz & r. pekrun (eds.), emotion in education. (pp. 3–10). san diego, ca us: elsevier academic press. stupnisky, r. h., perry, r. p., hall, n. c., & guay, f. (2012). examining perceived control level and instability as predictors of first-year college students’ academic achievement. contemporary educational psychology, 37(2), 81–90. doi:10.1016/j.cedpsych.2012.01.001 tempelaar, d. t., niculescu, a., rienties, b., gijselaers, w. h., & giesbers, b. (2012). how achievement emotions impact students’ decisions for online learning, and what precedes those emotions. the internet and higher education, 15(3), 161–169. doi:10.1016/j.iheduc.2011.10.003 tinto, v. (1997). colleges as communities: exploring the educational character of student persistence. journal of higher education, 68(6).   trautwein, u., schnyder, i., niggli, a., neumann, m., & lüdtke, o. (2009). chameleon effects in homework research: the homework–achievement association depends on the measures used and the level of analysis chosen. contemporary educational psychology, 34(1), 77–88. doi:10.1016/j.cedpsych.2008.09.001 verma, s., sharma, d., & larson, r. w. (2002). school stress in india: effects on time and daily emotions. international journal of behavioral development, 26(6), 500–508. doi:10.1080/01650250143000454 zeidner, m. (1998). test anxiety: the state of the art. springer. microsoft word korpershoek_publication.docx           frontline  learning  research  vol.4  no.  3  (2016)  28  -­‐  43   issn  2295-­‐3159       1  corresponding author: hanke korpershoek, university of groningen, grote rozenstraat 3, 9712 tg groningen, the netherlands. e-mail: h.korpershoek@rug.nl doi: http://dx.doi.org/10.14786/flr.v4i3.182   relationships among motivation, commitment, cognitive capacities, and achievement in secondary education hanke korpershoek university of groningen, the netherlands article received 5 june / revised 8 january / accepted 7 march / available online 10 may   abstract the aims of the present study were (1) to identify to what extent school motivation and school commitment contributed to the explanation of students’ academic achievement in addition to the effect of students’ cognitive capacities, (2) to find out whether school commitment mediated the relation between school motivation and academic achievement, and (3) to find out whether school motivation mediated the relation between school commitment and academic achievement. new in the field is that perspectives from two different research traditions were adopted, resulting in a selection of variables introduced by identity development theory and by motivational theories on achievement goals. the overall goal was to provide insight in the underlying structure of the relationships among these variables by providing new empirical evidence derived from a large student sample. a sample of more than 6,000 secondary school students from the netherlands was therefore used in the study. path models (structural equation models) were used to analyse the data. fit indices of the final model were satisfactory. this model included students’ cognitive capacities, three motivation factors (performance, social, and extrinsic motivation; mastery was excluded) and one commitment component (indepth exploration; the ‘commitment’ and ‘reconsideration of commitment’ components were excluded). the results showed small effects of performance (+), social (+), and extrinsic (-) motivation on academic achievement in addition to students’ cognitive capacities. a very small negative effect was found for in-depth exploration. in-depth exploration mediated the motivation – achievement relationships to a limited extent. suggestions for further research are discussed. keywords: school motivation; school commitment; cognitive capacities; academic achievement; identity development theory; achievement goal theory korpershoek       | f l r     29   1. introduction the purpose of the present study was to better understand the underlying structure of the relationships among school motivation, school commitment, and academic achievement of students in secondary education. school motivation is derived from the achievement goal framework. the school commitment construct follows from identity development theory, referring to students’ feelings of being committed to school. the relation between motivation and achievement has received ample attention in the literature (for recent meta-analyses using achievement goal theory see huang, 2012; hulleman, schrager, bodmann, & harackiewicz, 2010). however, within the widely used achievement goal framework, the focus is usually on a limited set of achievement goals (i.e. mastery and performance goals). maehr (1984) suggested that also social solidarity goals and extrinsic goals should be considered when studying achievement goals in educational settings, because students largely vary in their orientations toward learning. therefore, all four suggested achievement goals are investigated in this paper as indicators of students’ school motivation. the relation between school commitment and academic achievement has received far less attention in the literature. building and maintaining relationships with significant others in one’s environment is part of the identity development process (see e.g. klimstra, hale, raaijmakers, branje, & meeus, 2010; kroger, martinussen, & marcia, 2010; meeus, 2011). the school context is one of the most important life domains within which identity formation processes take place. students enter into various commitments by establishing meaningful relationships with peers and teachers. although it is plausible that the extent to which students feel committed influences students’ overall functioning at school, the literature on this topic is scarce. particularly the commitment construct as defined by identity development theory is not commonly used in educational studies. however, a wide variety of similar constructs (from various theoretical frameworks) have been used to explain student outcomes. that is, school commitment is conceptually related to school engagement (fredricks, blumenfeld, & paris, 2004), school membership (hagborg, 1998; wehlage, rutter, smith, lesko, & fernandez, 1989), school belonging (goodenow & grady, 1993), school relatedness (deci & ryan, 2002), and school connectedness (resnick et al., 1997; shochet, dadds, ham, & montague, 2006). the conceptually closest construct is ‘sense of school belonging’, which is explained further in the theoretical framework. prior studies have shown that students’ sense of school belonging is positively associated with school motivation (e. m. anderman, 2002; l. h. anderman & e. m. anderman, 1999; goodenow & grady, 1993; roeser, midgley, & urdan, 1996; ryan & powelson, 1991) and cognitive outcomes (anderman, 2003; goodenow, 1993; ma, 2003; osterman, 2000; roeser et al., 1996; pittman & richmond, 2007). based on these findings, it is expected that similar results can be found for the relationship between school commitment and academic achievement. all in all, the present study aims (1) to identify to what extent school motivation and school commitment contributed to the explanation of students’ academic achievement in addition to the effect of students’ cognitive capacities, (2) to find out whether school commitment mediated the relation between school motivation and academic achievement, and (3) to find out whether school motivation mediated the relation between school commitment and academic achievement. both school motivation and school commitment are, at least theoretically, malleable to some extent, thus insight into the (relative) contributions of these variables to students’ academic achievement is a relevant topic for educational practice. moreover, the multiple goal perspective that is adopted in this paper enables us to identify which achievement goals are related to more general academic achievement measures. this focus on general academic achievement is, in our view, important for educational practice, in addition to the more contextor domain-specific studies on student achievement. it is widely known that mastery goals are generally associated with favourable student achievement in class, but it is not clear whether this is also the case for students’ general academic achievement. in this paper, curriculum independent test scores on mathematics and reading comprehension are used as indicators of students’ general academic korpershoek       | f l r     30   achievement in the 9th grade of secondary education. these tests give an indication of students’ general academic functioning in secondary education. when relevant relationships are found between multiple achievement goals and students’ general academic achievement, these insights stress the importance of endorsing and stimulating various achievement goals in school. performance motivation, for example, may not be beneficial for students’ school grades in particular school subjects, but it may relate to students’ general academic achievement. the same line of reasoning applies to the impact of school commitment on student achievement. generally, positive effects are expected, but it is unclear whether these effects are contextor domain-specific or more general in nature. this paper addresses these issues by focusing on the effects of school motivation and school commitment on general academic achievement measures. some factors (e.g. performance motivation) might be weakly related to students’ grades in class, but show stronger relationships with general academic achievement in secondary education. as such, these factors can be seen as appropriate targets for intervention, because they are associated with students’ more general academic functioning. in paragraph 2, the school commitment and school motivation constructs are discussed in more detail before further explaining the present study. insights from various relevant theoretical frameworks are presented in order to clearly explain how the constructs were defined. 2. theoretical framework 2.1 the school commitment construct a fast-growing body of research now recognizes the significance of fulfilling basic psychological needs of students in education. self-determination theory (sdt) distinguishes the need for autonomy, competence, and relatedness which, when all three are supported, are associated with favourable outcomes. these needs specify ‘innate psychological nutriments that are essential for ongoing psychological growth, integrity, and well-being’ (deci & ryan, 2000, p. 229). the need for relatedness is suggested to facilitate the process of internalization, which means that people tend to internalize values and practices from contexts (and people within that context) in which they experience a sense of belonging (niemiec & ryan, 2009). the social context is therefore of major importance in facilitating growth processes such as growth in intrinsic motivation and integration of extrinsic motivation among students (deci & ryan, 2000). moreover, it is said that the need to belong precedes the desire for knowledge (e.g. deci & ryan, 2002). the need for relatedness is therefore seen as a basic and innate psychological need of people. closely linked to these statements about the need for relatedness is the so-called belongingness hypothesis, which states that human beings have a pervasive drive to form and maintain at least a minimum quantity of lasting, positive, and significant interpersonal relationships (baumeister & leary, 1995, p. 497). within the school context, this would imply that students generally have a pervasive drive (or in sdt an innate need) to form and maintain significant interpersonal relationships (e.g. with their teachers and peers). similarly, a sense of school belonging is conceptualized as ‘the extent to which students feel personally accepted, respected, included, and supported by others in the school social environment’ (goodenow & grady, 1993, p. 60-61). here we can already see that the need for relatedness, the pervasive drive to form and maintain interpersonal relationships, and the need to belong are closely related and, more importantly, are closely linked to identity development processes in the school context. faircloth (2012) stated that ‘identity can be seen as a type of ongoing negotiation of participation, shaped by – and shaping in response – the context(s) in which it occurs.’ (p. 186). the school context is therefore an important factor in shaping adolescents’ identity (eccles & roeser, 2011; lannegrand-willems, & bosma, 2006; rich & schachter, 2012). strongly grounded in the work of erikson (1950) and marcia korpershoek       | f l r     31   (1966, 1980, 1994), crocetti, rubini, and meeus (2008) developed a three-dimensional model of identity formation that can be used to assess adolescents’ identity formation processes in various life domains (e.g. the school). the model comprises three dimensions. the first dimension, commitment, is conceptualized as a choice made in an identity-relevant area and as the extent to which one identifies with that choice (crocetti et al., 2008, p. 218). it indicates whether a person feels committed to a certain relationship, for example, to friends or to school in general. meeus (1996) formerly defined commitment as the extent to which young people feel committed to, and derive self-confidence from, a positive self-image and confidence in the future from relationships (p. 585; see also bosma, 1985; meeus & dekovic, 1995; meeus, iedema, & maassen, 2002). recall that these definitions show remarkable overlap with the definition of school belonging. both refer to a malleable emotional state and both stress the importance of interpersonal relationships with significant others in obtaining a sense of school belonging or the feeling of school commitment. the second dimension, in-depth exploration, refers to the way in which adolescents deal with existing commitments and how much young people are actively engaged in investigating relationships. the third dimension, reconsideration of commitment, refers to the comparison between current commitments and other possible alternatives and also includes young peoples’ efforts to change present commitments because they are no longer satisfactory (crocetti et al., 2008, p. 209). together, the three dimensions can be used to characterize students’ (feelings of) commitment to the school in general. in the present study, crocetti et al.’s (2008) framework is used to measure students’ commitment to school. it has a strong theoretical basis and fits our idea that having a sense of commitment (or belonging) is an ongoing process of making and reconsidering commitments, thus interpersonal relationships with significant others such as teachers and peers (i.e. the school community). 2.2 the school motivation construct a broad range of motivational theories has attempted to unravel student motivation in educational settings, among others, achievement goal theory (agt; elliot & mcgregor, 2001) and personal investment theory (pi theory; maehr, 1984). motivational theories vary largely in how they define the concept of motivation and how motivation is operationalized. an oversimplified yet clear definition that can be drawn from agt and pi theory is that motivation refers to students’ general orientation towards learning. this general orientation involves cognitive aspects (e.g. adopting achievement goals) as well as noncognitive aspects (e.g. emotional reactions). for this paper, we focused on the adoption of achievement goals as indicators of students’ school motivation, because this approach takes a multiple goal perspective. it captures many different motivational dimensions (e.g. multiple achievement goals), which gives the opportunity to link students’ school commitment to various dimensions of students’ school motivation. agt emphasizes that students pursue different achievement goals in learning situations, such as mastery goals (focused on gaining knowledge and improving skills) and performance goals (focused on demonstrating their ability) (elliot & mcgregor, 2001). mastery-oriented students – those adopting (or striving towards) mastery goals – attempt to understand the topic at hand, gain knowledge, to improve their skills (e.g. tapola & niemivirta, 2008), which generally has a positive effect on students’ learning outcomes (huang, 2012). central to this orientation is the belief that effort leads to success (elliot & mcgregor, 2001). performance-oriented students – those adopting performance goals – are more focused on demonstrating their ability (e.g. tapola & niemivirta, 2008). one’s own ability is referenced against the performance of others (elliot & mcgregor, 2001). the effect of adopting performance goals is less straightforward; both positive and negative effects have been reported (e.g. huang, 2012). maehr (1984) suggested that also social solidarity goals and extrinsic goals should be considered when studying achievement goals in educational settings. maehr’s pi theory includes task goals (mastery), ego goals (performance), social solidarity goals, and extrinsic goals (see also king, ganotice, & watkins, 2014; king, mcinerney, & watkins, 2013; urdan & maehr, 1995). social goals can be referred to as socialgrounded reasons for studying, resulting from social concern and social affiliation (king & mcinerney, korpershoek       | f l r     32   2012). social-oriented students – those adopting social goals – are more focused on group learning, for example, studying for the sake of the group (covington, 2000). the relationship with academic achievement has not been studied frequently, though one can expect that the effect on academic achievement is at least positive. deci and ryan (2000) emphasize the importance of studying social goals that can affect achievement, in addition to examining more frequently addressed mastery and performance goals. extrinsic goals refer to the desire for external rewards such as praises and tokens. extrinsic-oriented students – those adopting extrinsic goals – attempt to gain external rewards in learning situations. external rewards then function as an incentive to continue one’s work or task (ryan & deci, 2000). some studies found negative effects of extrinsic motivation on cognitive outcomes (e.g. wolters, yu, & pintrich, 1996). however, as is the case with social goals, the relationship with academic achievement remains largely unclear. building on the theoretical frameworks of agt and pi theory, the inventory of school motivation was developed (ism; mcinerney, & sinclair, 1991; 1992; mcinerney & ali, 2006), in order to capture the four motivation dimensions, including mastery, performance, social, and extrinsic goals. these four motivation dimensions are used in the present paper. 2.3 relationships between the two constructs in a previous publication using the same dataset, latent cluster analysis was used to define student groups with different motivational profiles (korpershoek, kuyper, & van der werf, 2015). it was found that the student group with high scores on all motivation dimensions (i.e. adoption of mastery, performance, social, and extrinsic goals) also had high scores on school commitment. moreover, correlations between the four motivation dimensions and school commitment were all positive and small to medium in size (mastery .40; performance .17; social .32; extrinsic .23). there are also theoretical explanations why the associations are rather small. according to sdt, people tend to pursue goals, domains, and relationships that support their need satisfaction (deci & ryan, 2000). these authors state that relatedness plays a more distal role in the maintenance of intrinsic motivation than autonomy and competence, which more directly influence intrinsic motivation. it is not necessarily a prerequisite for intrinsic motivation, but a ‘needed backdrop’ that makes expression of the innate growth tendency of intrinsic motivation more likely (deci & ryan, 2000, p. 235). prior research also suggests that the two constructs are related to students’ academic achievement. school motivation is found to be a prominent predictor of school grades (e.g. brophy, 2004), but its relation with more objective academic achievement measures (e.g. curriculum independent achievement tests) is less straightforward. based on a meta-analysis of 84 studies, huang (2012) found correlations of .13 between mastery motivation and academic achievement and correlations of -.00 between performance motivation and academic achievement. correlations varying from -.02 to .09 were reported in korpershoek et al. (2015). korpershoek et al. (2015) also reported small and positive correlations between school commitment (as an overall construct) and academic achievement (.11 for reading comprehension and .13 for mathematics). having a sense of commitment (or belonging) is part of students’ basic psychological need satisfaction. it is therefore suggested to be an essential prerequisite for learning (and thus for academic achievement). 2.4 the present study an important question that follows from the theoretical framework is to what extent school commitment and school motivation are related, and to what extent they are related to students’ academic achievement. the goal of the path analyses conducted in this paper was to better understand the underlying structure of the relationships among these variables. three conceptual models were tested to identify to what extent school motivation and school commitment contributed to the explanation of students’ academic achievement in addition to the effect of students’ cognitive capacities. a measure of students’ cognitive capacities was included, because this is generally the strongest predictor of students’ academic achievement. korpershoek       | f l r     33   motivation and commitment were expected to show additive effects on academic achievement. the first model (model a) includes only direct effects on academic achievement, two other models also include indirect effects. the first mediation model (model b) includes mediation effects of school commitment on the relation between school motivation and academic achievement. theoretically, this model is the most plausible of the two because of the definition of school commitment used in this study. osterman (2000), for example, explains that in contexts in which students’ basis psychological needs (such as the need to belong) are met, students will function better (e.g. be more motivated) than in contexts in which their needs are not satisfied. the second mediation model (model c) includes mediation effects of school motivation on the relation between school commitment and academic achievement. there is no strong empirical support for the latter model, however, we sought to unravel the underlying structure of the relationships among motivation, commitment, and academic achievement. therefore, both mediation models were empirically tested. 3. method 3.1 participants the data used were collected as part of a large-scale study in secondary education in the netherlands, the so-called cool5-18 project (zijsling, keuning, kuyper, van batenburg, & hemker, 2009). the students included in the present paper were selected from a response group of 8,884 9th grade students (from 80 secondary schools throughout the netherlands) who had participated in the overall data collection. the students were on average 16 years old. in the netherlands, all students are expected to enter secondary education and obtain a secondary school diploma (track a or b, see below) or a secondary school diploma (track c) plus an addition diploma in senior secondary vocational education. students start 7th grade (year one of secondary education) in different educational tracks. the track placement is based on the primary school teachers’ recommendation. the lowest track is the preparatory secondary vocational education programme (track c, duration 4 years), which prepares students for senior secondary vocational education. this track is further divided into three sublevels. the senior general secondary education track (track b, duration 5 years) prepares students for higher professional education. the highest track, pre-university education (track a, duration 6 years) prepares students for university. thus, both tracks a and b prepare for higher education. the students in our sample pursued preparatory vocational secondary education (48%), senior general secondary education (27%), or pre-university education (25%). the sample included similar numbers of boys and girls (each 50%). 3.2 instruments 3.2.1 school commitment the school commitment scale was part of a paper-and-pencil questionnaire administered at the participating schools. we used an adapted version of the u-gids (utrecht-groningen identity development scale; crocetti et al., 2008). this instrument comprises three subscales: commitment (5 items), in-depth exploration (5 items), and reconsideration of commitment (3 items). sample items are: “my school gives me certainty in life” (commitment), “i think a lot about my school” (in-depth exploration), and “i often think it would be better to try to find a different school” (reconsideration of commitment; reversed scale), with answer categories ranging from 1 (strongly disagree) to 5 (strongly agree). the factor structure was confirmed in a factor analysis. the reliabilities of the subscales were: commitment (α = .86), in-depth exploration (α = .79), and reconsideration of commitment (α = .87). korpershoek       | f l r     34   3.2.2 school motivation a dutch version of the inventory of school motivation (ism) of mcinerney and ali (2006) was used. the questionnaire used here consisted of 32 items (see ali & mcinerney, 2004 for this subset of items) on a 5-point likert scale, ranging from 1 (strongly disagree) to 5 (strongly agree). the items were included in the same questionnaire as the items of the school commitment scale. factor analysis has confirmed the four factor structure suggested by the literature (mcinerney, dowson, & yeung, 2005; mcinerney, marsh, & yeung, 2003, see also korpershoek, xu, mok, mcinerney, & van der werf, 2015) and resulted in four reliable scales: mastery motivation (9 items, α = .77), performance motivation (7 items, α = .84), social motivation (7 items, α = .74), and extrinsic motivation (9 items, α = .86) in our sample. each of these four dimensions is based on two first order factors. mastery motivation is based on task (e.g. “i like to see that i am improving in my schoolwork”) and effort (e.g. “when i am improving in my schoolwork i try even harder”), performance motivation on competition (e.g. “i work harder if i’m trying to be better than others”) and social power (e.g. “i often try to be the leader of a group”), social motivation on social concern (e.g. “it is very important for students to help each other at school”) and affiliation (e.g. “i prefer to work with other people at school rather than alone”), and extrinsic motivation on praise (e.g. “at school i work best when i am praised”) and token (e.g. “i work hard in class for rewards from the teacher”). 3.2.3 cognitive capacities students’ score on an intelligence test was used as indicator of students’ cognitive capacities. students’ intelligence was estimated based on their performance on the so-called nscct intelligence test (“non-scholastic cognitive capacities test”; van batenburg & van der werf, 2004) which was adapted to the level of 9th grade students (see also zijsling et al., 2009). the test consists of 76 items including five topics: constructing figures, exclusion, series of numbers, categories, and analogies. the reliability of the test in the overall student sample was .91. 3.2.4 academic achievement two standardized achievement tests were used to assess the students’ achievements in mathematics and reading comprehension. the achievement tests were paper-and-pencil tests that were administered at the participating schools. the mathematics test was based on an item bank of 50 multiple choice questions, resulting in three different versions (with 11 anchored items) for students in different educational tracks. the reading comprehension text consisted of several short texts about which multiple choice questions were formulated. an item bank of 46 questions was used (with 11 anchored items). thus, different versions of the mathematics and reading comprehension tests with both anchored and unique items were used for students in the lower and higher educational tracks (for details see zijsling et al., 2009). for cool5-18 three versions of the test have been developed, two for track c students (one for the lowest two levels and one for the highest level within this track) and one for track a and b students. using a one-parameter logistic model (oplm; an item response model), the students’ scores were placed on one performance scale, indicating the percentage of items within the overall item test bank which a student was expected to answer correctly (between 0100%), regardless of the track they were in and regardless of the test version. the advantage of using this procedure is that the students’ scores can easily be compared across different test versions (e.g. when comparing the results of track a and track b students, which had taken the same test version). the reliability for the mathematics test was .94 and for the reading comprehension test it was .92. since we attempted to explain students’ academic achievement in general, a latent factor based on both test scores was included in the path models. 3.3 analyses structural equation modelling was applied to the data. models were estimated with mplus software (version 7) using maximum likelihood (ml) estimation. model fit indices reported are the chi-square and korpershoek       | f l r     35   degrees of freedom values, the root mean square error of approximation (rmsea), the standardized root mean square residual (srmr), the comparative fit index (cfi) and the tucker-lewis index (tli). adequate fit is found when the rmsea values are .06 or lower, srmr values are .08 or lower, and cfi/tli values are .95 or higher (hu & bentler, 1999). first, model a is presented, including only direct effects of the school motivation factors (i.e. four latent variables) and school commitment factors (i.e. three observed variables) on academic achievement. then, models b and c (the mediation models) are presented. insignificant paths (p > .01) will be deleted step-by-step to improve model fit. 4. results table 1 shows the correlations among all variables. table 1 correlations among all variables 1 2 3 4 5 6 7 8 9 1. cognitive capacities 2. mastery motivation .04 3. performance motivation .05 .36 4. social motivation .09 .47 .14 5. extrinsic motivation -.02 .51 .56 .36 6. commitment .11 .35 .14 .27 .17 7. in-depth exploration -.05 .35 .24 .26 .31 .30 8. reconsid. of commitment -.20 -.07 .08 -.10 .07 -.32 .09 9. reading comprehension .49 .07 .01 .08 -.03 .10 -.02 -.17 10. mathematics .70 .05 .10 .09 -.01 .13 -.04 -.21 .52 students’ cognitive capacities correlated highly with their scores on the mathematics test (r = .70) and moderately with their scores on the reading comprehension test (r = .49). all other correlations varied from -.01 to .56, with the highest correlations between performance and social motivation (r = .47), between mastery and extrinsic motivation (r = .51), between performance and extrinsic motivation (r = .56), and between the reading comprehension and mathematics scores (r = .52). finally, the correlations between the school motivation and school commitment components on the one hand and the achievement measures on the other hand were low (the highest correlation was -.21). all variables were initially included in the path models. the first path model (model a1) included direct effects of students’ cognitive capacities, school motivation (4 latent factors: mastery, performance, social, and extrinsic motivation), and school commitment (commitment, in-depth exploration, reconsideration of commitment) on students’ academic achievement. the model did not show adequate fit with regard to the rmsea (.084) and srmr (.091) values and the cfi (.881) and tli (.834) values. deleting the insignificant path from reconsideration of commitment to achievement (p = .167) in model a2 did not improve model fit: rmsea (.091), srmr (.098), cfi (.881), tli (.829). subsequently, the other insignificant path, that is, from mastery motivation to achievement (p = .026) was deleted in model a3. this model, now only including significant paths, also did not improve model fit (see table 2). korpershoek       | f l r     36   table 2 model fit results of models a3, b2, and c model a3 model b2 model c rmsea .090 [.087-.094] .057 [.053-.060] .126 [.122-.129] srmr .087 .034 .099 cfi .896 .967 .816 tli .846 .944 .702 aic 191957.338 215718.962 193420.437 bic 192181.762 215974.196 193665.437 χ2 (df) 1938.108 (35) 636.237 (26) 3395.381 (32) r2 .681 .677 .679 n 6639 7319 6639 note. full information maximum likelihood was used, therefore, the number of students included in the analysis varies per model. missing data are generally missing scores on the intelligence test, because not all schools administered this test. moreover, some individual students did not take the achievement tests or filled out the questionnaire (or had too many missing items to construct scale scores). subsequently, model b was constructed, including the direct effects from model a3 (3 out of 4 school motivation factors: performance, social, and extrinsic motivation; 2 out of 3 school commitment variables: commitment and in-depth exploration) and mediation effects of the school commitment variables on the motivation – achievement relationships. model b1 shows adequate fit: rmsea (.059), srmr (.037), cfi (.959), except for the tli value (.929). however, the one direct path was not significant, that is, from commitment to achievement (p = .215). model b2 therefore shows the results without this variable in the model (see table 2), which significantly improved model fit. the rmsea and srmr values are well below the cut-off values. the cfi value is above the cut-off value of .95 (hu & bentler, 1999), the tli value almost reaches the cut-off value (.944). model c, the model that included mediation effects of the school motivation factors on the school commitment – achievement relationships, did not fit the data (see table 2). model b2 appeared the best fitting model. figure 1 shows the corresponding path model. korpershoek       | f l r     37   figure 1 path model of model b2 (standardized estimates, standard errors between brackets) note. all paths are statistically significant at p < .001. the path from extrinsic motivation to in-depth exploration is significant at p < .01. the strongest predictor of academic achievement was students’ score on the intelligence test (an indicator of students’ cognitive capacities; .813). additionally, performance motivation (.155) and social motivation (.125) showed positive effects on students’ academic achievement. the desire to outperform others (performance motivation) and to learn together with others (social motivation) seems to progress students’ achievement. extrinsic motivation (e.g. learning for praise and tokens) was, however, associated with lower levels of academic achievement (-.161). the final model included one of the three subscales of school commitment, namely, in-depth exploration. referring to the extent to which students are actively engaged in investigating relationships and the way in which they deal with existing commitments, this variable was negatively related to academic achievement. the size of the effect was quite small (-.047), which indicates that this result needs to be interpreted with some caution. we will return to this issue in the discussion. stronger effects were found for the relationships between the motivational factors and in-depth exploration. higher levels of motivation (performance, social, and extrinsic) were associated with higher levels of in-depth exploration. that is, the higher one’s motivation, the more one thinks about and explores relationships at school. this was particularly the case for social motivation. the final model revealed small significant mediation effects of in-depth exploration on the motivation – achievement relationships, although we would like to stress that the relationship between indepth exploration and achievement was quite small to begin with. we tested the indirect effects of performance, social, and extrinsic motivation on achievement via in-depth exploration. these indirect effects korpershoek       | f l r     38   were negligible: performance motivation -.007 (se = .002; p < .01), social motivation -.014 (se = .004, p < .001), and extrinsic motivation -.004 (se = .002, p < .05). 5. discussion this study integrated insights from identity development theory and motivational theories on achievement goals in an educational context, using a large student sample. although the constructs that were used in this study have very different theoretical origins, the empirical findings underscore that school motivation (following motivational theories on achievement goals) and school commitment (following identity development theory) are related constructs among secondary school students. various school motivation factors (i.e. performance, social, and extrinsic motivation) and one school commitment component (i.e. in-depth exploration) each had unique effects on academic achievement in addition to the effect of students’ cognitive capacities. moreover, the school motivation factors were positively related to students’ in-depth exploration. educational studies attempting to explain students’ academic achievement should, therefore, integrate insights from these different theoretical perspectives in their explanatory models to further understand the direct and unique contributions of each of these variables. a positive direct effect was found for social motivation on students’ academic achievement (as suggested by covington, 2000 and deci & ryan, 2000) and a negative direct effect was found for extrinsic motivation (in line with findings presented by wolters et al., 1996), which suggests that it is relevant to study other achievement goals in addition to the more commonly addressed mastery and performance goals (see maehr, 1984). furthermore, a positive effect was found for performance motivation. performance-oriented students, thus those that, for example, responded that they worked harder when they tried to be better than others, had higher scores on the achievement tests than students with different orientations towards learning. for students’ general academic achievement, it seems beneficial to be (to some extent) oriented towards outperforming others. this finding is in contrast with the results of the meta-analysis of huang (2012), who did not find a significant relationship between performance motivation and achievement. a notable finding was that mastery motivation was the first factor that needed to be deleted from the model (see result section for details). the findings for mastery and performance motivation are in contrast with the results of the metaanalysis of huang (2012), who found positive relationships between mastery motivation and achievement but not between performance motivation and achievement. presumably, the study design is important here for the interpretation. when outperforming others is students’ general orientation toward learning, performance on a low stakes academic achievement test (which was used in this study) provides students with almost the same opportunities as performance on a high stakes test, namely outperforming others. when mastery is students’ general orientation toward learning, performance on a low stakes test does not imply that actual learning takes place. that is to say, the context does not ask for any learning activities such as trying to master the content. there were no consequences attached to the outcomes of the tests. a more methodological explanation is that the several motivation components were moderately correlated (which was allowed in the path model). particularly the correlations of social motivation with mastery and extrinsic motivation were moderately high, which may have resulted in smaller effects for each of these components. students are not mastery or performance-oriented, they often adopt various achievement goals in learning situations (see also korpershoek et al., 2015). only one of the three school commitment components was included in the final model. the higher students’ score on the in-depth exploration scale, the lower their general academic achievement. this would imply that thinking a lot about school and exploring one’s commitment to school is unfavourable for students’ general academic outcomes, which is not in line with theoretical notions discussed earlier in this paper. as already mentioned in the results section, the size of the effect was rather small (-.047), which is why this result should be interpreted with some caution. replication of the study is needed to validate these findings. the other two school commitment components (commitment and reconsideration of commitment) korpershoek       | f l r     39   were not included in the final model, indicating that those components were not related to students’ general academic achievement. as stated in the introduction, these factors may still be relevant for day-to-day functioning of students in class and presumably also for their school grades in more contextor domainspecific situations. the impact on general academic achievement could, however, not be confirmed. finally, although in-depth exploration mediated the motivation – achievement relationships, the indirect effects of performance, social, and extrinsic motivation on academic achievement via in-depth exploration were negligible. the final model that included these effects showed adequate model fit, but our data did not support the idea that one’s school commitment substantially mediated the motivation – achievement relationships. replication of the study is needed to validate these findings. notwithstanding these critical remarks, model b (including mediation effects of school commitment on the motivation – achievement relation) fitted the data much better than the theoretically less plausible model c (including mediation effects of school motivation on the school commitment – achievement relation). the study contributes to further theory development, particularly by highlighting that some motivational processes (such as adopting mastery goals) and some identity development processes (such as making commitments to people in one’s environment) are presumably more important for situation-specific school contexts then for general school contexts. that is, in our models, mastery motivation did not show a meaningful relationship with our general academic achievement measures (r < .10), but the correlations between mastery motivation and two school commitment components (commitment and in-depth exploration) were meaningful (both r = .35). these latter findings are more in line with theory (e.g. deci & ryan, 2000; osterman, 2000), because these relationships suggest that motivational processes and students’ identity development processes go, to some extent, hand in hand. model b (including mediation effects of school commitment on the motivation – achievement relation) fitted these theoretical notions, although the relationship between in-depth exploration and achievement we found was quite unexpected. however, in our study, we used curriculum-independent test scores to measure students’ academic achievement rather than situation-specific achievement measures (e.g. student achievement on a domain-specific test in a specific course in secondary education), which might explain this finding. based on our results, one could argue that the theories that we studied to explain differences in student achievement appear less applicable to this more general school context. an important suggestion for further theory development with regard to agt (elliot & mcgregor, 2001) and pi theory (maehr, 1984) is, therefore, to see how and to what extent these motivational theories on achievement goals can capture more general motivational patterns among adolescents in addition to more situation-specific contexts such as classroom learning. additionally, it might be worthwhile to examine different ways to operationalize school motivation (i.e. more situation-specific versus more in general) when studying students’ general academic achievement. with regard to educational practice, the finding that social motivation is positively associated with students’ general academic achievement, suggests that social motivation is a suitable target for intervention. although the contribution of this variable to the explanation of students’ general academic achievement is relatively small compared to the effect of students’ cognitive capacities, it showed a meaningful relationship. stimulating students’ social concern, for example, by emphasizing that it is important to help each other at school, may create an atmosphere in which students stimulate each other’s’ learning processes. in a similar vein, the findings show that students’ often prefer to work in groups rather than alone (social affiliation). the positive association between social motivation and academic achievement suggests that group work may stimulate student learning. in addition to validating the findings and confirming the final model in future studies, we suggest investigating differential effects on students’ academic achievement. that is, for particular student groups (e.g. for underperforming students) some relationships may be stronger than for other student groups, but more research is needed to investigate this (e.g. by using multigroup analysis). additionally, the addition of other variables in the model, for example, school engagement (see osterman, 2000) and self-efficacy (see hejazi, shahraray, farsinejad, & asgary, 2009) is a relevant topic for future research. various studies propose that the effect of sense of school belonging (conceptually related to school commitment) does not directly influence student achievement, but influences student engagement and self-efficacy beliefs, which in korpershoek       | f l r     40   turn affects achievement. an important limitation of this paper is that cross-sectional data were used, therefore eliminating the opportunity to examine cause-effect relationships. that is, the findings confirmed various significant associations, but it is likely that the relationships work both ways. for example, high academic achievement may have a positive impact on students’ motivation as well. further research in therefore needed to understand how these relationships develop over time (e.g. using cross-lagged models). notwithstanding this limitation, the main contribution of this paper lies in the empirically-funded argument that the integration of insights from identity development theory and motivational theories enhances our general understanding of student learning and student achievement in secondary school. keypoints this paper adopted insights from two different theories, namely identity development theory and achievement goal theory various motivation and school commitment components were significantly related to students’ academic achievement cognitive capacity was the strongest predictor of academic achievement among 9th grade secondary school students the final model included small effects of performance (+), social (+), and extrinsic (-) motivation on students’ academic achievement in-depth exploration mediated the motivation – achievement relationships to a limited extent references ali, j., & mcinerney, d. m. (2004). multidimensional assessment of school motivation. paper presented at the 3rd self research conference, berlin, germany. anderman, e. m. (2002). school effects on psychological outcomes during adolescence. journal of educational psychology, 94, 795-809. doi:  10.1037//0022-0663.94.4.795 anderman, l. h. (2003). academic and social perceptions as predictors of change in middle school students’ sense of school belonging. journal of experimental education, 72, 5-22. doi:   10.1080/00220970309600877 anderman, l. h., & anderman, e. m. (1999). social predictors of changes in students’ achievement goal orientations. contemporary educational psychology, 24, 21-37. doi: 10.1006/ceps.1998.0978 batenburg, th. a. van, & werf, m.p.c. van der. (2004). nscct: niet schoolse cognitieve capaciteiten test. voor groep 4, 6 en 8 in het basisonderwijs. verantwoording, normering en handleiding. groningen, the netherlands: gion. baumeister, r. f., & leary, m. r. (1995). the need to belong: desire for interpersonal attachments as a fundamental human motivation. psychological bulletin, 117, 497-529. doi: 10.1037//00332909.117.3.497 bosma, h. a. (1985). identity development in adolescence. coping with commitments. unpublished doctoral dissertation. groningen, the netherlands: university of groningen. brophy, j. (2004). motivating students to learn. mahwah, nj: erlbaum. covington, m. v. (2000). goal theory, motivation, and school achievement: an integrative review. annual review of psychology, 51, 171-200. doi:  10.1146/annurev.psych.51.1.171 crocetti, e., rubini, m., & meeus, w. (2008). capturing the dynamics of identity formation in various ethnic groups. development and validation of a three-dimensional model. journal of adolescence, 31, 207-222. doi: 10.1016/j.adolescence.2007.09.002 korpershoek       | f l r     41   deci, e. i., & ryan, r. m. (2000). the ‘what’ and ‘why’ of goal pursuits: human needs and the selfdetermination of behavior. psychological inquiry, 11, 227-268. doi: 10.1207/s15327965pli1104_01 deci, e. l., & ryan, r. m. (eds.). (2002). handbook of self-determination theory research. rochester, ny: university of rochester press. eccles, j. s., & roeser, r. w. (2011). schools as developmental contexts during adolescence. journal of research on adolescence, 21, 225-241. doi: 10.1111/j.1532-7795.2010.00725.x elliot, a. j., & mcgregor, h. a. (2001). a 2x2 achievement goal framework. journal of personality and social psychology, 80, 501-519. doi: 10.1037//0022-3514.80.3.501 erikson, e. h. (1950). childhood and society. new york: norton. faircloth, b. s. (2012). “wearing a mask” vs. connecting identity with learning. contemporary educational psychology, 37, 186-194. doi: 10.1016/j.cedpsych.2011.12.003 fredricks, j. a., blumenfeld, p. c., & paris, a. h. (2004). school engagement: potential of the concept, state of evidence. review of educational research, 74, 59-109. doi: 10.3102/00346543074001059 goodenow, c., & grady, k. e. (1993). the relationship of school belonging and friends’ values to academic motivation among urban adolescent students. the journal of experimental education, 62, 60-71. doi:   10.1080/00220973.1993.9943831 hagborg, w. j. (1998). an investigation of a brief measure of school membership. adolescence, 33, 461468. hu, l., & bentler, p. m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives. structural equation modeling, 6, 1-55. doi: 10.1080/10705519909540118 huang, c. (2012). discriminant and criterion-related validity of achievement goals in predicting academic achievement: a meta-analysis. the journal of educational psychology, 104, 48-74. doi: 10.1037/a0026223 hulleman, c. s., schrager, s. m., bodmann, s. m., & harackiewicz, j. m. (2010). a meta-analytic review of achievement goal measures: different labels for the same constructs or different constructs with similar labels? psychological bulletin, 136, 422-449. doi: 10.1037/a0018947 klimstra, t.a., hale, w.w., raaijmakers, q.a.w., branje, s.j.t. & meeus, w.h.j. (2010). identity formation in adolescence: change or stability? journal of youth and adolescence, 39, 150-162. doi: 10.1007/s10964-009-9401-4 king, r. b., & mcinerney, d. m. (2012). including social goals in achievement motivation research: examples from the philippines. online readings in psychology and culture, unit 5. retrieved from http://scholarworks.gvsu.edu/orpc/vol5/iss3/4. doi:  10.9707/2307-0919.1104 king, r. b., ganotice, f. a., & watkins, d. a. (2014). a cross-cultural analysis of achievement and social goals among chinese and filipino students. social psychology of education, 17, 439-455. published online first may 2014. doi: 10.1007/s11218-014-9251-0 king, r. b., mcinerney, d. m., & watkins, d. a. (2012). competitiveness is not that bad…at least in the east: testing the hierarchical model of achievement motivation in the asian setting. international journal of intercultural relations, 36, 446-457. doi: 10.1016/j.ijintrel.2011.10.003 korpershoek, h., kuyper, h., & van der werf, m. p. c. (2015). differences in students’ school motivation: a latent class modelling approach. social psychology of education, 18, 137-163. doi:10.1007/s11218014-9274-6 korpershoek, h., xu, j. k., mok, m. m. c., mcinerney, m. d., & van der werf, m. p. c. (2015). testing the multidimensionality of the inventory of school motivation in a dutch student sample. journal of applied measurement, 16, 41-59. kroger, j., martinussen, m., & marcia, j. e. (2010). identity status change during adolescence and young adulthood: a meta-analysis. journal of adolescence, 33, 683-698. doi: 10.1016/j.adolescence.2009.11.002 lannegrand-willems, l., & bosma, h. (2006). identity development-in-context: the school as an important context for identity development. identity, 6, 85-113. doi: 10.1207/s1532706xid0601_6 korpershoek       | f l r     42   ma, x. (2003). sense of belonging to school: can schools make a difference? the journal of educational research, 96, 340-349. doi: 10.1080/00220670309596617 maehr, m. l. (1984). meaning and motivation: toward a theory of personal investment. in c. ames & r. ames (eds.), research on motivation in education, vol. 1 (pp. 115-144). new york: academic press. marcia, j. e. (1966). development and validation of ego-identity status. journal of personality and social psychology, 3, 551-558. doi: 10.1037/h0023281 marcia, j. e. (1980). identity in adolescence. in j. adelson (ed.), handbook of adolescent psychology (pp. 159-187). new york: wiley. marcia, j. e. (1994). the empirical study of ego identity. in h. a. bosma, t. l. g. graafsma, h. d. grotevant, & d. j. de levita (eds.), identity and development. an interdisciplinary approach (pp. 6780). thousand oaks, ca: sage publications, inc. mcinerney, d. m., & ali, j. (2006). multidimensional and hierarchical assessment of school motivation: cross-cultural validation. educational psychology, 26, 595-612. doi: 10.1080/01443410500342559 mcinerney, d. m., & sinclair, k. e. (1991). cross-cultural testing: inventory of school motivation. educational and psychological measurement, 51, 123-133. doi: 10.1177/0013164491511011 mcinerney, d. m., & sinclair, k. e. (1992). dimensions of school motivation: a cross-cultural validation study. journal of cross-cultural psychology, 23, 389-406. doi: 10.1177/0022022192233009 mcinerney, d. m., dowson, m., & yeung, a. s. (2005). facilitating conditions for school motivation: construct validity and applicability. educational and psychological measurement, 65, 1046-1066. doi: 10.1177/0013164405278561 mcinerney, d. m., marsh, h. w., & yeung, a. s. (2003). toward a hierarchical goal theory model of school motivation. journal of applied measurement, 4, 335-357. meeus, w. (1996). studies on identity development in adolescence: an overview of research and some new data. journal of youth and adolescence, 25, 569-598. doi: 10.1007/bf01537355 meeus, w. (2011). the study of adolescent identity formation 2000-2010: a review of longitudinal research. journal of research on adolescence, 21, 75-94. doi: 10.1111/j.1532-7795.2010.00716.x meeus, w., & dekovic, m. (1995). identity development, parental and peer support in adolescence: results of a national dutch survey. adolescence, 30, 931-944. meeus, w., iedema, j., & maassen, g. h. (2002). commitment and exploration as mechanisms of identity formation. psychological reports, 90, 771-785. doi: 10.2466/pr0.90.3.771-785 niemiec, c. p., & ryan, r. m. (2009). autonomy, competence, and relatedness in the classroom. applying self-determination theory to educational practice. theory and research in education, 7, 133-144. doi: 10.1177/1477878509104318 osterman, k. f. (2000). students’ need for belonging in the school community. review of educational research, 70, 323-367. doi: 10.3102/00346543070003323 pittman, l. d., & richmond, a. (2007). academic and psychological functioning in late adolescence: the importance of school belonging. the journal of experimental education, 75, 270-290. doi: 10.3200/jexe.75.4.270-292 resnick, m. d., bearman, p. s., blum, r. w., bauman, k. e., harris, k. m., jones, j., tabor, j., beuhring, t., sieving, r. e., shew, m., ireland, m., bearinger, l. h., & udry, j. r. (1997). protecting adolescents from harm: findings from the national longitudinal study on adolescent health. the journal of the american medical association, 278, 823-832. doi: 10.1001/jama.278.10.823 rich, y., & schachter, e. p. (2012). high school identity climate and student identity development. contemporary educational psychology, 37, 218-228. doi: 10.1016/j.cedpsych.2011.06.002 roeser, r. w., midgley, c., & urdan, t. c. (1996). perceptions of the school psychological environment and early adolescents’ psychological and behavioral functioning in school: the mediating role of goals and belonging. journal of educational psychology, 88, 408-422. doi: 10.1037/0022-0663.88.3.408 ryan, r. m., & deci, e. l. (2000). self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. american psychologist, 55, 68-78. doi: 10.1037//0003066x.55.1.68 korpershoek       | f l r     43   ryan, r., & powelson, c. (1991). autonomy and relatedness as fundamental to motivation and education. journal of experimental education, 60, 49-66. doi: 10.1080/00220973.1991.10806579 shochet, i. m., dadds, m. r., ham, d., & montague, r. (2006). school connectedness is an underemphasized parameter in adolescent mental health: results of a community prediction study. journal of clinical child and adolescent psychology, 35, 170-179. doi: 10.1207/s15374424jccp3502_1 tapola, a., & niemivirta, m. (2008). the role of achievement goal orientations in students’ perceptions of and preferences for classroom environment. british journal of educational psychology, 78, 291-312. doi: 10.1348/000709907x205272 urdan, t. c., & maehr, m. l. (1995). beyond a two-goal theory of motivation and achievement: a case for social goals. review of educational research, 65, 213-243. doi: 10.3102/00346543065003213 wehlage, g. g., rutter, r. a., smith, g. a., lesko, n., & fernandez, r. r. (1989). reducing the risk: schools as communities of support. new york: falmer press. wolters, c. a., yu, s. l., & pintrich, p. r. (1996). the relation between goal orientation and students’ motivational beliefs and self-regulated learning. learning and individual differences, 8, 211-238. doi: 10.1016/s1041-6080(96)90015-1 zijsling, d., keuning, j., kuyper, h., batenburg, th. van, & hemker, b. (2009). cohortonderzoek cool518. technisch rapport eerste meting in het derde leerjaar van het voortgezet onderwijs [cohort study cool5-18. technical report of the first wave in the 9th grade of secondary education]. groningen/arnhem, the netherlands: gion/cito.   frontline learning research 1 (2013) 323 issn 2295-3159 corresponding author: suparna sinha, graduate school of education rutgers university, 10 seminary place, new brunswick, nj 08901, suparna.sinha@gse.rutgers.edu http://dx.doi.org/10.14786/flr.v1i1.1 3 | f l r conceptual representations for transfer: a case study tracing back and looking forward suparna sinha a , steven gray b , cindy e. hmelo-silver a , rebecca jordan a , catherine eberbach a , ashok goel c , spencer rugaber c a rutgers university, united states of america b university of hawaii, united states of america c georgia institute of technology, united states of america article received 10 march 2013 / revised 15 may 2013 / accepted 23 may 2013 / available online 27 august 2013 abstract a primary goal of instruction is to prepare learners to transfer their knowledge and skills to new contexts, but how far this transfer goes is an open question. in the research reported here, we seek to explain a case of transfer through examining the processes by which a conceptual representation used to reason about complex systems was transferred from one natural system (an aquarium ecosystem) to another natural system (human cells and body systems). in this case study, a teacher was motivated to generalize her understanding of the structure, behaviour, and function (sbf) conceptual representation to modify her classroom instruction and teaching materials for another system. this case of transfer was unexpected and required that we trace back through the video and artefacts collected over several years of this teacher enacting a technology-rich classroom unit organized around this conceptual representation. we provide evidence of transfer using three data sources: (1) artefacts that the teacher created (2) in-depth semi-structured interview data with the teacher about how her understanding of the representation changed over time and (3) video data over multiple years, covering units on the aquatic ecosystem and the new system that the teacher applied the sbf representation to, the cell and body. borrowing from interactive ethnography, we traced backward from where the teacher showed transfer to understand how she got there. the use of the actor-oriented transfer and preparation for future learning perspectives provided lenses for understanding transfer. results of this study suggest that identifying similarities under the lens of sbf and using it as a conceptual tool are some primary factors that may have supported transfer. keywords: transfer; technology; teacher learning; systems thinking s. sinha et al. 4 | f l r 1. introduction the aim of transfer research is to identify instructional conditions that prepare learners to apply what they have learned to new contexts. as designers of learning environments, we seek to create tools to facilitate transfer. we argue that one such tool is the use of conceptual representations to organize instruction by allowing students to develop a means to think about conceptual elements in a more generalised way (liu & hmelo-silver 2009). in addition, our prior research suggests that use of certain conceptual representations can promote understanding of complex systems. helping students and their teachers develop an understanding of complex systems is a difficult yet important component of scientific literacy (sabelli 2006). given the ubiquity of complex systems in the natural world, transferring ideas about complex system learning in one context to another is critical for the development of scientific thought. in many cases the behaviour of system components can affect its overall function, through emergent processes and localized interactions (jacobson & wilensky s2006). these interactions are often dynamic and invisible which make them difficult for learners to understand and present instructional challenges for teachers (feltovich et al. 2001; hmelo-silver et al. 2007). here we define systems thinking as being able to understand how bounded phenomena arise through considering the interactions and relationships among these interdependent structures, behaviours, and functions (hmelo-silver et al., 2007; nrc, 2012). there is evidence to suggest that students find it especially challenging to think about: (1) the interactions between visible and invisible structures, (2) the effect of their dynamic behaviours on overall functions, and (3) being able to extend their thinking beyond direct causality of complex systems (grotzer & bell-basca 2003; hogan 2000; hogan & fisherkeller 1996; jacobson & wilensky, 2006; leach et al. 1996; reiner & eilam 2001). in the research presented here, we investigate an unexpected case of transfer in a teacher as the learner who had been involved in a long-term classroom research project and appropriated the conceptual representation from the researcher-developed units to develop new instruction. this is particularly notable because learning about complex systems is often difficult (hmelo-silver et al., 2007). although our research focuses on the use of conceptual representations as a tool for learners, it also appears that it can be a tool for teachers to deepen their own understanding of complex systems (liu & hmelo-silver 2009; goel et al., 1996). specifically we discuss how structure-behaviour-function (sbf) served as a conceptual representation that promoted transfer across different complex systems (goel et al., 1996). structures are defined as the components of a system, behaviours as the mechanisms or processes that occur within a system and functions as system outcomes (goel et al., 1996; machamer et al., 2000). we developed technological tools using the sbf representation that make these features of complex systems salient (hmelo-silver et al., 2007; liu & hmelo-silver 2009; vattam et al. 2011). our study draws attention to a teacher‘s journey of understanding sbf as a conceptual tool, using it in the context of a technologyintensive science curriculum and her initiative to appropriate sbf as a conceptual representation beyond what we designed it for and use it meet local curricular needs. 2. research goals this study focuses on two main research questions: 1. how does a middle school science teacher develop her understanding of sbf as a representational tool? 2. how does generalization of sbf prepare her to make sense of a new complex system? specifically the focus of this study is to understand the means by which the teacher takes up opportunities to generalise her understanding of sbf as a representational tool to view similarities between two systems; one provided by researchers and one designated by the teacher. to understand the conditions that facilitated transfer, we need to view it through a lens that magnifies this teacher‘s learning trajectory. to s. sinha et al. 5 | f l r focus on the dynamic nature of transfer, we did not see a traditional model of transfer as a productive lens. traditional transfer researchers consider decontextualised expert knowledge, independent of how learners construe meaning in situations (cobb & bowers, 1999; greeno, 1997). because our objective was to highlight the processes the teacher used to understand and transfer a conceptual representation, we needed to consider alternative transfer models. such models should illuminate the interactions that were meaningful and engaging for the teacher and subsequently, led her to generalize her learning experience. 2.1 transfer through alternative lenses we consider transfer from both an actor oriented approach (aot; lobato 2004, 2006) and a preparation for future learning perspective (bransford & schwartz 1999) to investigate a teacher as a learner applying knowledge in a new curricular unit. lobato (2003, 2006) proposes that shifting from the observer‘s (expert‘s) perspective to considering how the actor (learner) perceives similarities between the new problem scenarios to prior experiences is a useful tool to understand transfer. evidence for transfer from this perspective is found by scrutinizing a given activity for any indication of influence from previous activities. moreover, we investigate how a greater understanding of sbf representations might have contributed to transfer from a preparation for future learning (pfl; bransford & schwartz, 1999) perspective. the pfl perspective focuses on the strategies used by learners in knowledge rich environments and their ability ―to learn a second program as a function of their previous experiences‖ (bransford & schwartz, 1999, p. 69). this provides a framework for evaluating the quality of particular kinds of learning experiences and the feedback they provide. feedback is a powerful factor in preparing students to make sense of instructional materials, to help them in knowledge construction and as a result facilitate transfer of skills needed to unpack novel problems (moreno, 2004; tan & biswas, 2006). like other alternative perspectives on transfer (e.g., konkola, tuomi-grohn, lambert, & ludvigsen, 2007), the classroom context and activity is an important factor in promoting transfer. we add to the transfer literature by exploring use of the sbf conceptual tool for abstracting systems thinking. that is, the conceptual tool can be used to make sense of complex systems by thinking about macro and micro level connections either independently or at multiple levels of intersections. we make the conjecture that sbf as a conceptual tool can serve as a focusing phenomenon, which makes it suitable for integrating the aot and pfl lenses of transfer as we describe in the next section. in this study, we investigate how the experiences that led to successful generalization of sbf as a conceptual tool prepared the teacher to keep refining her systems thinking. 2.2 supporting transfer through focusing phenomena lobato et al (2003) propose that focusing phenomena supports transfer by prompting students to generalize their learning. as a concept they define focusing phenomena as "observable features of the classroom environment that regularly direct attention to certain mathematical properties or patterns" (p.2). they attribute a combination of factors such as curriculum materials, artefacts, teacher‘s instructions as important for directing and focusing students' attention towards the intended content. in the context of this study, we extend the notion of focusing phenomena to science. we propose that sbf serves as focusing phenomena (figure 1) to advance systems thinking. it helps the teacher focus her attention on understanding connections between multiple structures, their functional roles within the complex system and the behaviours they exhibit. here we consider the importance of generalizing sbf as a tool for transfer. from an aot perspective, sbf as a focusing phenomena highlights what is similar between two complex systems i.e. the aquatic ecosystem (introduced by the researcher) and human digestive system (introduced by the teacher). it helps concretize the idea that biological systems are similar to ecosystems in terms of interacting at multiple levels. using this framework affords the teacher opportunities to focus on the s. sinha et al. 6 | f l r connections that exist between various organs of the digestive system. specifically, it directs the teacher‘s attention to the ways that ―structure and function in biological systems are causally related through behavioural mechanisms‖ (hmelo-silver et al., 2007, p. 308). the teacher‘s understanding of sbf in the classroom mirrors her understanding of systems thinking. this is important for us, as researchers, as it lets us trace the teacher‘s learning trajectory. from a pfl perspective, thinking in terms of sbf prepares learners to understand that behaviours are mechanisms and processes that enable structures to achieve their functions in biological systems (bechtel & abrahamson, 2005; machamer, darden, & craver, 2000). in the remainder of the paper, we present a case study that considers how several aspects of the learning environment influenced the teacher‘s generalization of sbf as a conceptual tool. figure 1. sbf as focusing phenomena. 3 a case of transfer: the instructional context this study is part of a larger research program, which is a technology-intensive curriculum unit centred on an aquarium based aquatic ecosystem. the curriculum provides multiple opportunities for learners to develop and deepen their understanding of sbf as a conceptual tool. first, technological tools such as the reptools toolkit (hmelo-silver et al., 2011) and the aquarium construction toolkit (act; vattam et al., 2011) were designed: (1) to help learners think about aquatic ecosystems in terms of structures, the functions they perform within the system and the behaviours they exhibit to perform the functions, (2) teach about the aquarium ecosystems using sbf as a conceptual tool for a period of 4 years, and (3) engage in active discussions about the concept and ways to teach it with the research team present daily in the classroom and at the annual professional development workshops. 3.1 sbf tools the reptools toolkit includes a function-oriented hypermedia (hmelo-silver et al., 2007; 2009; liu & hmelo-silver, 2009) organized in terms of sbf representation and net logo computer simulations (wilensky & reisman, 2006). the hypermedia (figure 2) introduces the aquarium system with a focus on functions and provides linkages between structural, behavioural and functional levels of aquariums. it is organized around what, how, and why questions which correspond to structures, behaviors, and functions. http://www.youtube.com/watch?v=y0n9jectfuu&feature=youtube_gdata http://reptools.rutgers.edu/%20startpage.html s. sinha et al. 7 | f l r figure 2. aquarium hypermedia. two netlogo simulations allow learners to explore macroscopic processes of fish reproduction (i.e., the fishspawn simulation, figure 3a) as well as microscopic processes (the nitrification simulation, figure 3b) that represent the chemical and biological processes in the aquarium. the simulations provide a context for learners‘ investigation of the aquatic ecosystem. they afford opportunities for designing experiments, manipulating variables, making predictions, and discussing conflicts between predictions and results. each simulation allows learners to explore key features that are relevant to the process of fish spawn or nitrification cycle. figure 3a. macro levelfish spawn simulation. http://reptools.rutgers.edu/revisedfishspawnmodel.html http://reptools.rutgers.edu/revisednitrificationmodel.html s. sinha et al. 8 | f l r figure 3b. micro level – nitrification simulation. the second component to the learning environment, act is designed to promote construction of sbf models (vattam et al., 2011). models can be constructed either in a table (figure 4a) or graph (figure 4b) format. the model table focuses learners‘ attention on thinking about various structures in an ecosystem. the three column table affords the opportunity for learners to think about the structural components, their multiple behaviours and functions. this is valuable because learners get an opportunity to understand both individual mechanisms in the system and the meta-level concepts related to complex systems. figure 4a. sample act model table. s. sinha et al. 9 | f l r figure 4b. sample act model graph. the act model graph is a platform for learners to create models of their evolving understanding of ecosystem processes in terms of sbf. as students read through the hypermedia, generate and test their hypotheses with the simulations, they integrate the critical structures with their behaviours and functions in act models. 3.2 methods we used a case study approach to characterize how a science teacher, ms. y, appropriated her understanding of sbf as a representational tool and applied it to make sense of a new complex system. case study methodology allowed us to use multiple data sources to study this complex phenomenon in context (e.g., stake, 1998; yin, 2009). borrowing from interactional ethnography (castanheira, green, & yeager, 2009) we began at the end—the sbf hypermedia that ms. y constructed. the unit of analysis for this case is the individual teacher in her classroom context over several years. through this approach, we used multiple sources of data to trace the social and cognitive events that occurred over time and led ms. y to see sbf as a tool she could appropriate for her teaching practice. although this was not an ethnography we borrowed the logic of this inquiry approach to understand how an individual within a social context constructed particular knowledge over time (bridges, botelho, green, & chau, 2012). 3.3 context ms. y taught seventh grade science at a public middle school in north east united states. she had been teaching science for 26 years and had a bachelor‘s degree in elementary education. this study was part of a larger 4-year study focused on teaching middle school science students about aquatic ecosystems. ms. y participated in annual professional development (pd) workshops. the pd focused on concepts related s. sinha et al. 10 | f l r to aquatic ecosystem, analysis in terms of sbf and the technological tools that she would need to use in her classroom. during the pd, ms. y. had the chance to share her pedagogical challenges and experiences, such as difficulties in using the software or teaching about sbf as a conceptual tool. ms. y had been using the reptools and act in an aquarium curriculum for four years when she informed us that she wanted to develop her own instructional tools using the sbf representation to teach about cell and human body systems. this prompted her to collaborate with her colleague, another science teacher, ms. t. together they used microsoft power point to create a human body system presentation, modelled after the function-centred aquatic hypermedia. we refer to it as the teacher-created hypermedia. given their limitations in terms of technical knowledge in designing a hypermedia similar to the one we had created, the teachers hyperlinked key words in their power point presentation and follow up questions to point to relevant slides. ms. t also taught seventh grade science in the same school. she was a new teacher with one year of teaching experience. ms. t had a science education background. while she collaborated with ms. y, she also attended the annual pd and implemented the same technology intensive curriculum on aquatic ecosystem in her classroom. each teacher taught four diverse seventh grade classes with approximately twenty-five students in each section. during the curriculum implementation the students were grouped together in small heterogeneous groups. 3.4 data sources we had three primary sources of data. first was the artefact that the teacher created (this indeed was the impetus for our research). second, we conducted an hour-long semi-structured interview with the two teachers, ms. y & ms. t. finally we collected video data of classroom interactions. these videos were drawn from classroom data from a long-term (i.e., four year) research project. these helped us to understand: (1) why the teacher transferred her generalizations of sbf representations to new instructional domains and (2) how she transferred these understandings. we interviewed ms. y & ms t approximately two months after ms. y completed teaching about both systems. the primary focus of the interview was to understand how she conceived the idea of extending the computer-based representational tools beyond what was expected from her, the influence of her prior knowledge during this process, and her attempts to prepare herself to solve new challenges. following powell, francisco and maher‘s (2003) recommendations for video analysis, we reviewed video data to identify critical events. in an attempt to trace and track the nature of ms. y‘s generalizations of sbf we selected representative clips of critical events from her classroom that demonstrated evidence of her developing understanding and generalization of sbf representations as a tool to teach about another complex system. these video clips included whole class discussions that ms. y had with her students while: (1) introducing the sbf representation for the aquatic ecosystem in year 3 (i.e., the year before she created the digestive system unit), (2) introducing the sbf representation for the aquatic ecosystem in year 4 i.e. the year she employed the digestive system unit, and (3) explanation of sbf representations and modelling of the digestive system unit. we viewed a total of nine clips that consisted of three classroom interactions for each of the three kinds of whole class discussions. 3.5 analysis we examined classroom interactions that highlight ms. y‘s learning trajectory with sbf as a representational tool. the video data were analysed using interaction analysis (ia; jordan & henderson 1995), which involved collaborative viewing of video clips by six members of the interdisciplinary research team. we successively conducted nine ia sessions to collaboratively review the selected video clips, describe observations, and generate hypotheses. any differences in opinions were resolved by discussions. s. sinha et al. 11 | f l r this helped ensure the trustworthiness of our interpretations through the initial independent interpretations of the ia session participants and the subsequent discussions. during the ia sessions we focused our attention on two specific aspects of ms y‘s practice. first, we paid attention to patterns and variations in the ways that she introduced the sbf as a conceptual tool in relation to the aquatic ecosystem across the four years. specifically, we examined her explanation of the concept, the analogies she presented and whether or not she sought help from any external resources, such as researchers in the classroom or ms t. second, we focused on how she introduced sbf as a conceptual tool in the context of the human body unit. at this time we made comparisons between the ways the topic was introduced in the aquatic ecosystem with the human body system. we also looked for similarities in terms of analogies. in particular, we wanted to understand if and how her prior knowledge of sbf prepared her to discuss this particular complex system with ease and confidence. to gain a holistic perspective of the teacher‘s journey we also examined the interview transcript. we looked for themes related to the mechanisms by which transfer occurred in the ways in which the teacher constructed similarities between aquarium and digestive systems. this allowed us to triangulate the teacher‘s perspective with the ia and artefact analysis. 4. findings based on our analysis of the interview and video data we identified themes related to aot or pfl perspectives. these findings helped strengthen our understanding of the processes ms. y used to generalize sbf as a representational tool and observe how it prepared her for the transfer. the aot perspective provided a framework to trace ms. y‘s evolving understanding of using the sbf lens as a tool to make sense of aquatic ecosystem. the pfl perspective demonstrated how ms. y transferred and used her knowledge of sbf to make sense of a complex system that was outside the scope of our research. 4.1 tracing and tracking ms y’s understanding of sbf from an aot lens 4.1.1 orientation to the sbf representation led by the teacher ms y‘s journey began with using the act tool. the act technology enabled construction of sbf representations using the model table (figure 4a). the tool introduced the students and ms. y to the language of sbf representations. initial data analysis of the whole class video revealed that the teacher‘s introduction of the sbf representation played a critical role in students‘ conceptual understanding of the complex system. she presented the idea that the sbf representations captured interconnected entities within a complex system while completing the act table: 1. ms. y: alright, so the first thing yours say is fish right? so, lets go back and tell me what is the behaviour of the fish? 2. student: releases waste. 3. ms y. ok. so the fish releases waste. right? alright, so it, it releases what kind of waste? 4. students (in unison): ammonia. 5. ms. y: right, so you have that in there right? now. what is the function? 6. students (in unison): remove toxins from the body. 7. ms. y: okay. so we want to get these things out of the fishes‘ body. now next, the next one is what? 8. students: ammonia. s. sinha et al. 12 | f l r 9. ms. y: ammonia. so, put, put ammonia here. alright, so now, what is, what is the behaviour of the ammonia? what‘s it do if you look at it in the tank? 10. student: water? 11. ms. y: yeah, it‘s just floating around right? what‘s its function do? it‘s food for bacteria. so it has its purpose right? so the next one on our list which is blank on yours will be what? in this excerpt, ms. y drew the students' attention to the functions and behaviours of various structures present in the aquarium. the students identified structures such as fish (turn 1) and ammonia (turn 8). next she prompted them to think about their behaviours and functions. in turn 2, the students responded that the behaviour of the fish is to release waste. she pushed them to think in detail about the kind of waste (turn 3) and the function or overall purpose of this behaviour (turn 5). in turn 11, she clearly articulated that structures have a function within complex systems. although this is a somewhat mechanical application, it also allowed her to begin to see how the sbf lens might serve as a tool for understanding systems. we speculate that this discussion prepared both the students and ms y. to use the sbf conceptual representation to understand the interconnectivity between various structures within complex systems. this initial understanding of sbf as a representation may have prepared ms. y to appropriate sbf as tool when she collaborated with her colleague to create a new learning tool i.e. (the teacher-created hypermedia). 4.1.2 teacher-created hypermedia just as the orientation to sbf was the starting point, the artefact that ms. y created at the other end bound the case study. ms y., in collaboration with her colleague ms. t, created new hypermedia in the form of an interactive powerpoint of the cell and body systems mirroring the aquarium hypermedia developed by the research team (figures 5a and 5b). the teachers‘ hypermedia outlined the different structures in the system along with orienting why and how questions. the how questions were directed towards behaviours of system components and the why questions focused on functions. the teachers created this hypermedia as a learning resource to help students connect cell systems to larger body systems. the research team did not plan either the body system hypermedia or the use of modelling these systems using the act software; the teachers did this of their own volition. figure 5a. researcher-developed hypermedia. figure 5b. teacher-developed hypermedia. the development of the cell hypermedia demonstrated multiple ways by which ms. y generalized and transferred her understanding of sbf as a conceptual tool. first, understanding the sbf of the aquatic ecosystem prepared her to teach it better in successive years and second, she was able to modify the learning environment (i.e., by changing them physically–from an aquarium hypermedia to a cell hypermedia and by seeking resources) into something that was more compatible with her current goals. s. sinha et al. 13 | f l r 4.1.3 identifying similarities through sbf representations ms. y‘s initiative to extend and appropriate our research and develop additional classroom instruction suggested that the sbf representation was becoming a tool for her to see similarities across complex systems. adopting an aot perspective helped us understand how she constructed similarities between what she had been teaching for several years (the aquatic ecosystem) to the current unit she developed (cell and body systems). this perspective helped us recognize which connections she made, on what basis, and how and why those connections were productive (lobato, 2004). for example, consider ms. y‘s response when asked about the utility of their hypermedia during the interview session: right, and it's a hard concept to get. so, what we were thinking about is like the kids actually think when they eat food it breaks down and then leaves the body. they don't get that the food has to go to the cells and the cell actually works and creates energy from this food and then there's a waste and it sends that back to the body for it to be excreted. so we're trying to give them not only the names of the parts and what each part does individually but how it needs to work-...and we're doing the behaviour not only of the cell itself but behaviour of all the systems and then the behaviour of the whole body. and the cells are all part of that whole body. this highlights that ms. y understood that the cells were an integral part of the body systems and could not be taught in isolation. earlier, she noted that systems in the body are not disconnected and have complex mechanisms that allows for higher order operation. this provided evidence that she now understood how structures within a system perform multiple behaviours in order for it to function effectively. the ia results showed how ms. y introduced the sbf representation and refined her thinking over multiple years. 4.1.4 refining the sbf representation as a conceptual tool from an aot perspective we needed to track ms. y‘s transition from her initial naïve ideas about sbf representations to a more expert conception. the results from the ia indicated that ms. y‘s understanding of the sbf representation as a conceptual tool changed. she used several distinct strategies to introduce the topic of complex systems ranging from discrete (i.e., in years 1, 2 and 3), to acknowledging complexity (in year 4), and finally providing a systems perspective with her new cell/body unit. in the first three years, she introduced to the sbf representation to her students by mentioning the new terminology being used to understand the aquatic system. however, she introduced structures, behaviours, and functions as discrete constructs. in year 4, she espoused a coherent view of sbf representation as a conceptual unit. later that year, while introducing sbf in the context of the unit on cells and body she explained sbf as a system, complete with nested and interconnected subsystems. 4.1.4.1 year 3: sbf representations ms. y‘s early introduction to sbf representation suggested a focus on linear connections. this was shown by the way in which she filled out the act sbf table (figure 4a) in front of the classroom. as a way to connect ideas about sbf she drew clear conceptual lines between one structure at a time and all the behaviours exhibited by that structure as the following example shows: we just named them all yesterday. the heater, the fish, the plants. those things are called the ‗structure‘. the next word we're gonna use is ‗behaviour‘. the behaviour is what the fish do. what do the things do in the tank? and the next word we're gonna use is ‗function‘ okay? so what i want to do today is to start with structure and behaviour. so, i made a chart and the first column is the structure, or the parts. so everyone write down one of the things in the fish tank is fish and the second column i wrote was behaviour, and the third column i wrote was function. we're going to start with this second column that is behaviour. when i ask you the behaviour of something, i want to s. sinha et al. 14 | f l r know is what does it do? "what do fish do?" swims, eats, breathes, and poops. okay, all fish swim. that is their behaviour okay. they swim. what else do fish do? here ms. y. described the meaning of the term ―behaviour‖ somewhat superficially as ―what fish do‖ rather than the more expert mechanistic view. she established linear connections between the structure (fish) and the multiple behaviours (swims, eats, breathes, poops) that this structure exhibits. after promoting an understanding of the behaviour exhibited by the structure (fish), she then drew another relationship between each individual behaviour in the last column to indicate the behaviour‘s function. 4.1.4.2 year 4: sbf representations are interconnected over time, ms y‘s introduction to the sbf representation became richer and more complex. in the excerpt below taken from a whole class discussion in year 4 she described structures, behaviours and functions as interconnected entities within a system, rather than discrete elements on a worksheet: 1. ms. y: okay, now, let's do the filter. i'm gonna do the filter with you and then you're gonna do one on your own. all right, so what does the filter do? what does the filter do? jim what does the filter do? 2. jim: um, cleans out the tank 3. ms. y: cleans the tank. or cleans the ―what part of the tank?‖ 4. jim: the water in the tank? 5. ms. y: all right, so the filter will clean the water. okay? now, why does it clean the water? 6. jim: so it can put more oxygen into the water? 7. ms. y: no. that's another thing that it does. it actually, because it's spinning around, because it's spinning like this, it's actually, one of the things it does…is it adds oxygen to the water. now, this part here, why does it do it? first of all, i want to stop right here. the filter is this big grey thing here. right? now, first of all, how does it work? what's this big tube doing? [points to picture of filter on the screen.] 8. pat: sucking up the water 9. ms. y: sucking up the water. then the water comes up here, right? and it gets sucked up and it goes back here and it pours back down. when it flushes back over that's when the oxygen from the air can get pulled back into the water. okay, so howyou said it cleans the waterhow does it do this? 10. pat: well, it has the filter. the filter has like chemicals and stuff. 11. ms. y: what do you think is in this bag? 12. pat: bad stuff 13. ms. y: well, eventually the bad stuff is going to get in here, but actually there's charcoal in here, gravel in here. and then when the water flows through it, can it catch all the big chunks? maybe the fish faeces and stuff like that? so, and then see how it spins back down here? water splashes and it's pulling in the oxygen. so now, all right so now, why does it clean the water? what is the point of cleaning the water? after turn 13, the class went on to discuss the fish and the plants, how the filter aerates the tank and how it affects the whole system. in turns 3 and 5 when ms. y discussed the behaviours (the mechanism that cleans water in tank) and function of the filter (by collecting faeces from fish) she was guiding students‘ answers to structure, behaviour, and function simultaneously and filling in the chart appropriately, stressing relationships rather than focusing on any one aspect in isolation. turns 6-12 show that ms. y used student response to generate more questions that linked what and why questions throughout her classroom discussion, highlighting the system complexity. 4.1.4.3 year 4: sbf representations at multiple levels of complex systems s. sinha et al. 15 | f l r later in the same year, when introducing her unit on the cells to the class, ms y emphasized that sbf works as a whole across multiple levels of complex systems. as the next excerpt shows, she did so not directly, but more discretely through leading questions: 1. ms. y: eventually what we want [the researchers] to do for us is allow us to model systems within systems. what happens if i can click on the cell and zoom in on that and put the cell parts in there? because they don't have the ability to zoom right in on that one part, are there any ideas on how to connect the cell through modelling to the other body systems? because you also want to go and look at the function. what do you think? 2. lucia: umm, what about if you like umm put a picture of the cell. 3. ms y: yeah but i want to drive everything to the cell because that's, you know, the whole body operates to get things to the cell you know that right? but then i also want to show what the cell does inside once you send the food there. so how can i show that part…on this graph? okay. you know how this is a system. the body parts and the cell is its own little mini system, how can i show the stuff inside the cell? should i circle all the mitochondria right around the cell? or should i pull the cell out and make that part separate? … these demonstrate how ms. y refined her thinking about sbf as a conceptual tool. whereas earlier, her focus was primarily in working with the aquatic ecosystem, she later introduced a new level of complexity by introducing the idea that there exists multiple ‗mini systems‘ within the human body system. she still focused largely on structures but she also made connections to behaviours and functions. in addition, she helped students understand that one structure may have multiple behaviours and functions (in turns 1 and 3). comparing her sbf representation of the cell system here to that of the aquatic ecosystem in the earlier unit, she presented it to the class as a coherent system rather than discrete sbfs. in addition, when applying the sbf representation to the cell, ms. y introduced a meta-perspective by explicitly explaining that the task was to represent their ideas through modelling (in turn 1). moving away from the isolated task provided in earlier (i.e., filling out the table by first listing structure followed by behaviour, and then function), ms. y explained that the students were organizing their knowledge in model graph. by placing emphasis on the modelling tool and providing students with the starting point of the structure, the cell, ms. y explained that the task was to develop a representation of their ideas about the human body system, using the table to organize their ideas and providing the students with leading questions that she had provided earlier when talking about the sbf representation in the aquarium unit. this transition suggested that ms. y was an active learner herself. she frequently asked questions to the research team and ms. t, to refine her understanding. this practice of asking questions had two effects. first, it helped ms. y identify and address the gaps in her understanding, which prepared her for future learning. second, it shed light on the processes that she as an actor (learner) used to construct similarities between the aquatic ecosystem and cell system. 4.2 experiences to promote transfer from a pfl perspective 4.2.1 recognition of teacher as a learner in the interview, ms. y indicated that since the beginning of her involvement in the project, her knowledge continually developed. she explained that she was the primary source by which information was passed from the research team to the students and that over time she felt that she became more competent in this role. in the interview, she acknowledged her lack of mastery over the content and was aware that she refined her ideas of the sbf representation and the aquarium unit which lead to development of the new unit: okay, my knowledge of this still develops every year because it‘s knowledge that [research team leader] had and ityou knowwas her angle on something and then i had to try to understand what s. sinha et al. 16 | f l r was going on in her head. so it's taken me many years of practice and talking to [research team leader], talking to [researchers in the room], to kind of get this. and i still do not feel like i'm really solid on it, but i get it more and more each year. these statements demonstrated that ms. y saw herself as a learner in her classroom as she was looking critically at her current knowledge and beliefs. this experience prepared her to deepen her understanding of the content, and revise her ideas as she gathered new information. 4.2.2 collaboration facilitates generalisation the collaboration aspect was beneficial during the inception, design, and construction of the teacher created hypermedia. together they went beyond our research agenda by using sbf as a conceptual tool to create a power point presentation of human body systems. it afforded opportunities for sense making and focus on critical aspects of complex systems while working with the tools (figure 6). as ms. y talked about the creation of the cell hypermedia, she revealed that she was highly motivated to do so because of the potential for feedback and interaction with ms. t. for example, when asked how the idea came about and the variables that affected the development of the new tool, ms. y responded: so then i kind of realized that what i needed to do was give her [ms. t.] my idea and then hear from her what she would add to that and in turn that wouldi would take what she added into my lesson, so one of us throws out like a main idea and then the other one builds upon that main idea and then we get a better idea. and that's how i think that the hypermedia came along. because this whole concept has been in my head for a long time, about how kids don't understand the whole body and the cells connection to the body. so i talked about it with ms. t and then she started talking about making a hypermedia and then we went back and forth on how we we're going to do it. figure 6. affordances of the learning environment that promote sbf thinking. from a pfl perspective, people seeking multiple viewpoints about issues may be one of the most important ways to prepare them for future learning (bransford & schwartz, 1999). it is clear from this excerpt that ms. y. felt it useful that she could exchange her ideas and collaborate in the creation of the new hypermedia with ms. t. this finding suggests that ms. y. was able to see the possibilities for transferring her understanding of the sbf representation. however, this transfer was dependent on the idea of using hypermedia itself as a s. sinha et al. 17 | f l r way to organize complex content in addition to the sbf representations. our next set of results focus on elaborating how she used the aquatic hypermedia to guide her thinking about designing for another complex system. 4.2.3 appropriating salient features of the aquarium hypermedia when asked about what parts of the hypermedia she found useful in her own development, ms. y felt that working with the same aquarium hypermedia for four years allowed her to incorporate some of the key features in the hypermedia she created. although her hypermedia did not possess the technological and conceptual sophistication of the aquarium hypermedia, it prepared her for refining her understanding along a trajectory of increasing expertise. this process was important from an aot perspective as it enabled her to see the connections between two situations by identifying the salient features from the earlier hypermedia environment (lobato, 2004). it is notable that she transferred other features of the hypermedia structure beyond sbf, including the use of guiding questions as well as the use of short pieces of text accompanied by simple and relevant graphics: i would say that i definitely liked how each question lead to another question because that's how we modelled ours was every question gave an answer but then lead to another question and another question and another question…. we also used just short pieces of information because i think the kids get bored if you put too much it's overwhelming. we used pictures and then we also had it not only lead to different the next one and the next one but it bounced back sometimes a design in the hypermedia too. from the interview it is clear that ms. y drew upon relevant features of the aquarium hypermedia. although her rationale for keeping a short text was different from what we had in mind while designing the aquarium hypermedia, this process of experimentation also helped her clarify her own thinking about the concepts that she is placing within the new hypermedia contexts (bransford et al., 1990). 4.2.4 approaching act to model a new system in addition to appropriating aspects of the aquarium hypermedia, ms. y also appropriated the act tool so that students could model body systems in the same fashion as they had for the aquarium system (figures 7a and 7b). figure 7a. digestive system act table view. s. sinha et al. 18 | f l r figure 7b. digestive system act graph view. the following excerpt highlights ms.y‘s journey of trying to understand how to use sbf as a conceptual tool, the act technology itself and feel comfortable using it to teach by herself: at first she (research team leader) came and she was just testing the kids‘ knowledge and that i was not really involved. and then … we originally started talking about the cell and the body as that was an area she worked in, and then she got the idea of the respiratory system because that slowly developed into … the net logo and the hypermedia. back then structure, function and behaviour i think for me was all just disjointed. all the pieces were here and i was just trying to keep up with her. and then … the act program helped a lot because it sort of put everything together for me in the end, like okay, here's all the knowledge that the kids have been getting along the way, here is proof that they got it. and for me it was just a slow process of absorbing everything and you know kind of understanding it until i could you know turnkey it and then we could turn around and together make another hypermedia with it. this exemplified the importance of the act software as a capstone to allow for students‘ and ms. y‘s understanding of the new system be made explicit. in the interview ms. y recalled that in the beginning of the research program (i.e., years 1 and 2), her understanding of the framework was ―disjointed‖. she attributed the act modelling toolkit to prepare her to create the human body system hypermedia. it appeared to help her think about interconnections between structures, their functions and visible behaviours. this example from the interview, and the classroom task of modelling body systems in act, indicated that ms. y possessed the confidence to organize the new ideas generated by her hypermedia into sbf terms using the tool and the importance. additionally it also highlighted her ability to appropriate the act tool as the final classroom task to evaluate knowledge generated by the hypermedia as a way to organize student ideas about complex systems. 4.2.5 preparing to ask sbf oriented questions a critical aspect of transfer of the sbf framework involved being able to make sense of the new complex system in terms of "what", "how" and "why" questions. the act modelling table (see figure 4a) prepared learners to think about the aquatic ecosystem in terms of sbf by answering questions related to "what", "how" and "why." it was evident that questions related to "what" pertain to visible and invisible structures that determine key variables of the aquatic ecosystem. because the learner had to only identify s. sinha et al. 19 | f l r relevant components in the first column, it involved an important but superficial level of system understanding, unless it led the learner to consider why and how it performs specific actions in context to the aquatic ecosystem. video analysis in year 1 revealed that although ms. y discussed the role of functions and behaviours, she was more comfortable labelling the aquatic ecosystem in terms of its relevant components. this was apparent, as she would begin the class with "what" questions. if the students gave her the expected answer she would make an attempt to elaborate on it. but when the students gave incorrect answers, she just ignored the response. as a result the students were not encouraged to share their confusion with the class in terms of why they thought so and how they came to the conclusion. during the year we observed that ms. y consistently asked more "what" questions. this prompted the students to give single word responses. the students also noticed that the teacher expected them to give short answers that did not call for detailed explanations. this indicated that ms. y was hesitant to open the discussion for an in-depth systems thinking conversation that focused on sbf relations. it was likely that at that stage her idea about complex systems was focused on identifying relevant structures. we observed a slightly different trend in year 2. although the "what" questions dominated the whole class discussions, students were also asked to think about possible interactions or connections between structures. as the students identified such relationships, ms. y led the discussion on ―how‖ questions by writing down behaviours that connected structures. video analysis indicated that in years 3 and 4, ms. y appeared to be confident in discussing the aquatic ecosystem in terms of a complex system, interconnected by visible and invisible components as this next example shows: 1. ms. y: yes, anybody have something else, let‘s put another living thing in there. what do you have? 2. jaden: microorganisms. 3. ms. y: okay. so what are microorganisms? 4. jaden: they clean up the waste. 5. ms. y: what do you mean they clean up waste? 6. jaden: they eat. 7. ms. y: ok, the next problem is function. these particular structures do a particular behaviour and that behaviour fits in a little bit more into the whole picture. think why does it need to do this behaviour for it, why do the fish need to swim? this excerpt shows that ms. y opened the discussion by asking the class to identify structures connected to the aquatic ecosystem. next she drew their attention to thinking about their behaviours. as soon as the class discussed some behaviours, she asked them to think about behaviours in context to their functional role in the aquarium based aquatic ecosystem. ms. y was able to build upon her prior understanding of sbf as a conceptual tool. 5. discussion as we seek to understand transfer, we must address questions related to the ―what‖ and ―how‖ of transfer. that is, we need to articulate the exact nature of the content or ―what‖ is being transferred. equally important is identifying the mechanisms or the ―how‖ that is responsible for this transfer to occur. we suggest that we can accomplish these goals through the integration of aot and pfl perspectives on transfer. we used aot to reach backwards and see how the similarities were constructed, whereas pfl allowed a look forward at how applying sbf prepared ms. y for her future learning and practice. the case study findings showcase how different perspectives on transfer allowed us to understand how participation in a research project driven by principles of learning empowered a teacher to appropriate these tools in her own practice, going beyond the research project context. s. sinha et al. 20 | f l r this case study suggests that sbf as a conceptual tool has potential for making sense of complex systems. we propose that using sbf as a focusing phenomena (lobato et al., 2003) is a mechanism that facilitates transfer. sbf was a lens through which ms. y could see the relationship between systems and prepare her to learn about new systems. our findings demonstrated the processes adopted by the teacher to generalize her understanding of sbf. this included an initial superficial engagement with sbf that she deepened and refined over several years and her own reflectiveness in seeing herself as a learner. in addition we discussed the influence of the social environment and technological affordances that appear to prepare her for transfer. the additional viewpoints of ms. t. and the conversations with the research team suggest that collaboration is important in preparing for transfer. having a general-purpose tool that she could repurpose to use for a new unit was instrumental in this process. finally, she was able to use the hypermedia that the research team had created as a worked example that allowed her to explore the content and how sbf could be applied to a new domain. from a pfl perspective, these results shed light on specific processes and challenges that ms. y had to overcome. specifically we were keen to understand what it took for a teacher to acquire mastery over using a conceptual tool in one context and be prepared to use it to solve a problem in a different context. the findings indicate that the sbf representation focused the teachers‘ attention on the behavioural connections and functional roles of components within complex systems. it prepared them to think about the actions or ―how‖ components behave within a complex system in relation to their overall functions. both teachers reported that this was useful when they started working on creating the hypermedia on digestive system. although the teacher-constructed hypermedia lacked the technical sophistication of the researcher created hypermedia, the teachers made productive use of a technology they were familiar with, a power point presentation. the teachers also successfully incorporated key features of the aquarium hypermedia such as leading questions, short descriptions and use of images. their interview responses indicated that their prior experience with the aquarium hypermedia drew their attention to these features. this prepared them to be efficient and effective with their own hypermedia design. both these processes (i.e., creation of the new hypermedia and thinking in terms of behaviours, in addition to structures and functions) were vital as ms. y was able to revise her knowledge and beliefs, which set the stage for her to analyse and appreciate critical features of the new information presented to her (bransford et al. 1990; moore & schwartz, 1998). this process of analysing her beliefs and strategies also highlights the active nature of transfer, which is an important part of pfl. the initiative she took in applying her sbf representation understanding to teaching a new unit demonstrates her ability to revise and rethink the current situation to suit her current goals. from a pfl perspective this is valuable as it reveals the importance of activities and practices that are beneficial for ―extended learning‖ rather than on one-shot task performances (bransford & schwartz, 1999). our study also extends the transfer literature by proposing new ways for understanding teacher learning trajectories. as we observed ms. y‘s transition over multiple years, our focus was on the processes she followed during this transition rather than assessing mastery over content knowledge. in terms of learning trajectories, our results highlight the fact that ms. y was looking critically at her knowledge and gradually developed a deeper understanding in that content area. data analysis from earlier years revealed a limited understanding of the sbf representation as a conceptual tool. however, she actively sought resources (fellow colleague, ms. t and researchers present in the classroom) to help her understand the interconnections between multiple structures, their functions in the system and visible and invisible behaviours. her increasing confidence in the content area, coupled with collaboration, resulted in her being highly motivated to extend the research tools to other areas of her classroom practice. this case study provides an existence proof that aot and pfl can be used to explain a single case of transfer. it is important however to consider the limitations from a single case (yin, 2009). although we cannot rule out all possible rival explanations, we triangulated data from multiple data sources and included researchers with a range of disciplinary backgrounds and experience in the interaction analysis. other members of the research team who were not involved in the ia sessions reviewed the examples and interpretations that were presented here. we acknowledge that further research in complex classroom environments is needed in order to generalize these findings. because of the importance of the social interactions and feedback that ms. y received from teaching her students (e.g., okita & schwartz, in press), s. sinha et al. 21 | f l r it is unlikely that a purely cognitive explanation could account for these results. the analysis presented in this study suggests the possibilities of extending research on alternative approaches to transfer (lobato, 2006; bransford & schwartz, 1999; van oers, 1998). these new approaches to transfer suggest a much more complex and dynamic process than traditional cognitive accounts. our results also suggest that different theoretical frameworks can be productively integrated in providing accounts of transfer. in our case, teacher adoption and appropriation of a learning framework was an exciting by-product of scholarly research because it provides evidence that classroom innovations can be appropriated and sustained. keypoints sbf as a focusing phenomena is a mechanism that facilitates transfer. it acts as a lens through which the learner can see the relationship between systems and prepare them to learn about new systems. there are possibilities of extending research on alternative approaches to transfer (lobato, 2006; bransford & schwartz, 1999; van oers, 1998). these new approaches to transfer suggest a much more complex and dynamic process than traditional cognitive accounts. different theoretical frameworks can be productively integrated in providing accounts of transfer. acknowledgements this research was funded by institute of education sciences (ies) grant # r305a090210. conclusions or recommendations expressed in this paper are those of the authors and do not necessarily reflect the views of ies. we thank the participating teachers without whom this work would not be possible. an earlier version of this research was presented at the 2010 international conference of the learning sciences. references bechtel, w., & abrahamson, a. (2005). explanation: a mechanist alternative. studies in the history and philosophy of biological and biomedical sciences, 36, 421-441. bransford, j. d., vye. n„ kinzer, c, & risko, v. (1990). teaching thinking and content knowledge: toward an integrated approach. in b. f. jones & l. idol (eds.), dimensions of thinking and cognitive instruction: implications for educational reform. hillsdale, nj: erlbaum. bransford, j. d. & schwartz, d. l. (1999). rethinking transfer: a simple proposal with multiple implications. review of research in education, 24, 61-100. bridges, s., botelho, m., green, j. l., & chau, a. c. m. (2012). multimodality in problem-based learning (pbl): an interactional ethnography. in s. bridges, c. mcgrath & t. l. whitehill (eds.), problem-based learning in clinical education (pp. 99-120). dordrecht netherlands: springer. castanheira, m. l., green, j. l., & yeager, e. (2009). investigating inclusive practices: an interactional ethnographic approach. in k. kumpalainen, c. e. hmelo-silver & m. césar (eds.), investigating classroom interaction: methodologies in action (pp. 145-178). rotterdam: sense publishers. cobb, p., & bowers, j. s. (1999). cognitive and situated learning perspectives in theory and practice. educational researcher, 28, 4-15. darling-hammond, l., & mclaughlin, m. w. (1995). policies that support professional development in an era of reform. phi delta kappan, 76, 597-604. feltovich, p. j., coulson, r. l., & spiro, r. j. (2001). learners‘ (mis)understanding of important and s. sinha et al. 22 | f l r difficult concepts. in k.d. forbus & p. j. feltovich (eds.), smart machines in education: the coming revolution in educational technology (pp. 349–375). menlo park, ca: aaai/mit press. greeno, j.g. (1997). response: on claims that answer the wrong questions. educational researcher, 26, 517. goel, a. k., gomez de silva garza, a., grué, n., murdock, j. w., recker, m. m., & govinderaj, t. (1996). towards designing learning environments -i: exploring how devices work. in c. fraisson, g. gauthier & a. lesgold (eds.), intelligent tutoring systems: lecture notes in computer science. ny: springer. grotzer, t.a., & bell-basca, b. (2003). how does grasping the underlying causal structures of ecosystems impact students‘ understanding? journal of biological education, 38, 16-28. hmelo-silver, c. e., jordan, r., honwad, s., eberbach, c., sinha, s., goel, a., rugaber, s., & joyner, d. (2011). foregrounding behaviors and functions to promote ecosystem understanding. proceedings of hawaii international conference on education (pp. 2005-2013). honolulu hi: hice. hmelo-silver, c. e., marathe, s., & liu, l. (2007). fish swim, rocks sit, and lungs breathe: expert-novice understanding of complex systems. journal of the learning sciences, 16, 307-331. hmelo-silver, c. e., liu, l., & jordan, r. (2009). visual representation of a multidimensional coding scheme for understanding technology-mediated learning about complex natural systems. research and practice in technology enhanced learning environments, 4, 253-280 hogan, k., & fisherkeller, j. (1996). representing students‘ thinking about nutrient cycling in ecosystems: bidimensional coding of a complex topic. journal of research in science teaching, 33, 941– 970. hogan, k. (2000). exploring a process view of students' knowledge about the nature of science. science education, 84, 51-70. jacobson, m. j., & wilensky, u. (2006). complex systems in education: scientific and educational importance and implications for the learning sciences. journal of the learning sciences, 15, 11-34. jordan, b., & henderson, a. (1995). interaction analysis: foundations and practice. journal of the learning sciences, 4, 39-103. konkola, r., tuomi-grohn, t., lambert, p., & ludvigsen, s. (2007). promoting learning and transfer between school and workplace. journal of education and work, 20, 211-238. leach, j., driver, r., scott, p. & wood-robinson, c.: 1996, ‗children‘s ideas about ecology 3: ideas about the cycling of matter found in children aged 5–16. international journal of science education, 18, 129-142. liu, l., & hmelo-silver, c. e. (2009). promoting complex systems learning through the use of conceptual representations in hypermedia. journal of research in science teaching, 9, 1023-1040. lobato, j., ellis, a. b., & munoz, r. (2003). how focusing phenomena in the instructional environment support individual students generalizations. mathematical thinking and learning, 5, 1-36. lobato, j. (2004). abstraction, situativity, and the ―actor-oriented transfer‖ perspective. in j. lobato (chair), rethinking abstraction and de contextualization in relationship to the ―transfer dilemma.‖ symposium conducted at the annual meeting of the aera, san diego, ca. lobato, j. (2006). alternative perspectives on the transfer of learning: history, issues, and challenges for future research. the journal of the learning sciences, 15, 431-449. machamer, p., darden, d., & craver, c. f. (2000). thinking about mechanisms. philosophy of science, 67, 1-25. moore, j. l., & schwartz, d. l. (1998). on learning the relationship between quantitative properties and symbolic representations. in proceedings of the international conference of the learning sciences (pp. 209-214). mahwah, nj: erlbaum. moreno, r., (2004). decreasing cognitive load for novice students: effects of explanatory versus corrective feedback in discovery-based multimedia. instructional science, 32, 99-113. national research council. (2012). a framework for k-12 science education practices, crosscutting concepts, and core ideas. washington, dc. okita, s., & schwartz, d. l. (in press ). learning by teaching human pupils and teachable agents: the importance of recursive feedback journal of the learning sciences. powell, a. b., francisco, j., & maher, c. a. (2003). an analytical model for studying the development of learners' mathematical ideas and reasoning using videotape data. journal of mathematical behaviour, s. sinha et al. 23 | f l r 22, 405-435. reiner, m., & eilam, b. (2001). conceptual classroom environment: a system view of learning. international journal of science education, 23, 551-568. sabelli, n. (2006). complexity, technology, science, and education. journal of the learning sciences, 15, 510. stake, r. e. (1998). case studies. in n. k. denzin & y. s. lincoln (eds.), strategies of qualitative inquiry (pp. 86-109). thousand oaks ca: sage. tan, j., & biswas, g., (2006). the role of feedback in preparation for future learning: a case study in learning by teaching environments. intelligent tutoring systems, 2, 370-381. van oers, b. (1998). from context to contextualizing. learning and instruction, 8, 473-488. vattam, s., goel, a., rugaber, s., hmelo-silver, c., jordan, r., gray, s., & sinha, s. (2011) understanding complex natural systems by articulating structure-behavior-function models. educational technology & society, 14, 66-81. wilensky, u. & reisman, k. (2006). thinking like a wolf, a sheep or firefly: learning biology through constructing and testing computational theories – an embodied modelling approach. cognition and instruction, 24, 171-209. yin, r. k. (2009). case study research: design and methods (fourth ed.). thousand oaks ca: sage. frontline learning research 2 (2013) 70-85 issn 2295-3159 corresponding author: stellan ohlsson, university of illinois, chicago, stellan@uic.edu http://dx.doi.org/10.14786/flr.v1i2.58 70 | f l r beyond evidence-based belief formation: how normative ideas have constrained conceptual change research stellan ohlsson a a university of illinois at chicago, chicago, united states article received 4 september 2013 / revised 13 december 2013 / accepted 13 december 2013/ available online 20 december 2013 abstract the cognitive sciences, including psychology and education, have their roots in antiquity. in the historically early disciplines like logic and philosophy, the purpose of inquiry was normative. logic sought to formalize valid inferences, and the various branches of philosophy sought to identify true and certain knowledge. normative principles are irrelevant for descriptive, empirical sciences like psychology. normative concepts have nevertheless strongly influenced cognitive research in general and conceptual change research in particular. studies of conceptual change often ask why students do not abandon their misconceptions when presented with falsifying evidence. but there is little reason to believe that people evolved to conform to normative principles of belief management and conceptual change. when we put the normative traditions aside, we can consider a broader range of hypotheses about conceptual change. as an illustration, the pragmatist focus on action and habits is articulated into a psychological theory that claims that cognitive utility, not the probability of truth, is the key variable that determines belief revision and conceptual change. keywords: belief formation; belief revision; cognitive utility; conceptual change; descriptive vs. normative inquiry; pragmatism cognitive scientists pride themselves on their interdisciplinary approach, drawing upon anthropology, artificial intelligence, evolutionary biology, linguistics, logic, neuroscience, philosophy, psychology, and yet s. ohlsson 71 | f l r other disciplines in their efforts to understand human cognition. interdisciplinary research strategies have paid off in the natural sciences. for example, in the middle of the 20 th century, research on the border between biology and chemistry resulted in spectacular advances, including the determination of the structure of dna (watson & crick, 1953). it is plausible that an interdisciplinary approach will pay off in the study of cognition as well. but the cognitive sciences exhibit principled differences that might get in the way of interdisciplinary efforts. inquiry into cognition was originally rooted in the desire for human betterment. the first cognitive disciplines, including logic, epistemology, and linguistics, were normative disciplines. logicians wanted to systematize valid inferences, as opposed to whatever inferences, including fallacious ones, that people make; philosophers sought to identify criteria for certain knowledge, as opposed to describe all knowledge 1 ; and early linguists were more concerned with codifying correct grammar than with cataloguing grammatical errors. historically, these and related disciplines mixed normative and descriptive elements in a way that is quite foreign to the contemporary conception of a natural or social science. in this respect, they resembled aesthetics, ethics, and legal scholarship more than biology, chemistry, and physics as practiced since the scientific revolution (butterfield, 1957; osler, 2000). because the normative disciplines were historically prior, the concepts, practices, and tools of normative inquiry became part of the intellectual infrastructure of the self-consciously descriptive sciences like neuroscience and experimental psychology that became established in the latter half of the 19 th century. concepts like abstraction, association, and imagery are obvious examples of such imports. this intellectual inheritance helped the new sciences get started, in part by suggesting questions and problems (how and when are associations formed?). but normative and descriptive disciplines are different enough in their goals and methods so that it is reasonable to ask whether that inheritance has had a negative impact as well. in this article, i argue that certain normative ideas have led the study of cognition in general and cognitive change in particular down an unproductive path. the types of cognitive changes i have in mind are those that psychologists call conceptual change, belief revision, and theory change (carey, 2009; duit & treagust, 2003; nersessian, 2001; thagard, 1992; vosniadou, baltas & vamvakoussi, 2007) for purposes of this article, i use these terms as near-synonyms. when a collective label is needed, i call them nonmonotonic change processes (ohlsson, 2011). the argument proceeds through seven steps. in section 1, i elaborate on the distinction between normative and descriptive inquiry. section 2 highlights the role of normative concepts in what i call the ideal-deviation paradigm, a particular style of research that is common in cognitive psychology. in section 3 i show that this research paradigm is also present, albeit implicitly, in conceptual change research. from a normative perspective, people ought to base their concepts and beliefs on evidence and revise them when they are contradicted by new evidence. the assumption -sometimes implicit, sometimes explicit -that people‟s cognitive systems are designed to operate in this way has focused researchers‟ attention on the deviations of human behaviour from normatively correct belief management (using the latter term as a convenient shorthand for “belief formation and belief revision”). but the normative perspective is irrelevant to the scientific study of cognitive change; hence, so are the deviations between the norms and actual cognitive processing. but if the deviations are irrelevant, so are our explanations for them. to go beyond the current state of conceptual change research, we need to make explicit the influence of the normative perspective, identify the constraints it has imposed on theory development, and relax those constraints. when we cultivate a resolutely descriptive stance, the space of possible theories of conceptual change expands. as a first step towards a new theory, section 4 argues that the notion that people form concepts and beliefs on the basis of evidence might be fundamentally incorrect. the problem of why people do not revise their misconceptions when confronted with contradictory evidence then dissolves, and other questions move to the foreground. in section 5, i outline an approach to conceptual change that is inspired by the pragmatist notion that concepts and beliefs are tools for successful action. according to this perspective, the key variable that drives conceptual change is not the strength of the relevant evidence or the probability of truth, but cognitive 1 indeed, many philosophers insist that unless knowledge is true and certain, it does not qualify as knowledge. s. ohlsson 72 | f l r utility. section 6 answers two plausible objections to this view, and section 7 outlines some of its implications. section 8 recapitulates the argument. although the concept of utility is not itself new, the critique of conceptual change research as mired in normative ideas is stated here for the first time, and the conjecture that utility can replace probability of truth as the key theoretical variable in conceptual change has not been previously proposed. 1. normative versus descriptive inquiry descriptive sciences aim to provide accurate theories about the way the world is. as i use it, the word “descriptive” does not stand in contrast to “theoretical” or “explanatory” but encompasses all the empirical and theoretical practices of the natural and social sciences as we now conceive them; the term “empirical” could have been used instead, but is sometimes understood as standing in contrast to “theoretical.” the goal of descriptive science is to provide an account of reality that is intersubjectively valid and reflects the world as it is, independent of human judgments or wishes. descriptive sciences are essentially concerned with adapting theories and concepts to data. the descriptive sciences constitute what we today call “science.” normative disciplines, in contrast, investigate how things ought to be. they are essentially concerned with conformity to standards of goodness. a large portion of what we have in our heads consists of more or less explicit normative knowledge. aesthetics, epistemology, ethics, etiquette, law, literary criticism, logic, rhetoric, and several other disciplines ask what the appropriate standards are, or should be, in some area of human endeavour; how one decides whether some instance does or does not conform to the relevant standards; and why particular instances conform, or fail to conform. the state of theory in these disciplines varies widely, from formalized theories of valid inferences in logic to the obviously culture-dependent rules of good manners, and the highly controversial theories of literary criticism. as these nutshell definitions are meant to illustrate, descriptive and normative disciplines are so different from each other that the distinction seems impossible to overlook. but the separation of normative and descriptive inquiry was in fact long in coming; for example, psychology was included among “the moral sciences” well into the 19 th century. the distinction was not fully articulated and accepted in western thought until the 20 th century, supported by, among other influences, the logical positivists‟ emphasis on the distinction between fact and value. (for criticisms of the distinction, see, e.g., köhler, 1938/1966, putnam, 2002, and others.). it is nevertheless anachronistic to think of earlier generations of scholars as having confused descriptive and normative inquiry. the situation is better described by saying that they had not yet distinguished them. astronomy provides an example of research in an era when the distinction was not yet fully articulated (kuhn, 1957; margolis, 1987, 1993). some ancient astronomers adopted the normative idea that planets ought to move in perfectly circular orbits, because the heavenly bodies were perfect beings and perfect beings ought to move in perfect orbits and the circle is the most perfect geometric figure. astronomers then spent two millennia explaining the deviations of the observed planetary orbits from the normatively specified orbits using the ptolemaic construct of epicycles, instead of exploring other hypotheses about the geometric shape of the orbits (frank, 1952). research in the natural sciences is no longer constrained by normative ideas in this way. section 2 shows that psychology, in contrast, has not yet outgrown its normative inheritance. s. ohlsson 73 | f l r 2. the ideal-deviation paradigm normative principles have generated a psychological research paradigm that i refer to as the idealdeviation paradigm. although only a subset of psychological research conforms to this paradigm, the paradigm has had a strong and largely negative impact on research in cognitive psychology in general and research on conceptual change in particular. a line of research that follows this paradigm proceeds through the following general steps (examples to follow): (a) choose a normative theory. how ought the mind carry out such-and-such a process, or perform such-and-such a task? (b) construct or identify a situation or task environment in which that theory applies, and derive its implications for normatively correct behaviour. (c) recruit human subjects and observe their behaviour in the relevant situation. (d) describe the deviations of the observed behaviour from the normatively correct behaviour. (e) hypothesize an explanation for the observed deviations. (f) test the explanation in further empirical experiments. readers who are familiar with cognitive psychology will have no difficulty in thinking of instances of the ideal-deviation paradigm. the prototypical example is research on logical inference (evans, 2007). in this area, researchers originally used logic as developed by logicians – primarily the logics of syllogistic and propositional inferences – as the relevant normative theory. the reasoning problems presented to human subjects include wason‟s famous 4-card task (a.k.a. the selection task; wason & johnson-laird, 1972). propositional logic prescribes a particular pattern of responses to this task, and the deviations of human responses from the prescribed pattern have been replicated in dozens, perhaps hundreds, of experimental studies. researchers have proposed and debated a wide range of explanations for the observed deviations (johnson-laird, 2006; klauer, stahl, & erdfelder, 2007). research on decision making is a second instance of the ideal-deviation paradigm. the subjective expected utility (seu) theory and the mathematics of probability provide a normative theory for how to choose among competing options. when people are confronted with choices that involve probabilistic outcomes in laboratory settings, they deviate from the normatively correct behaviour in a variety of ways. errors like availability and representativeness are examples. in the former, human judgments about the probability of an event (e.g., an airline crash) are influenced by the ease with which the person can retrieve an example of such an event from memory. in the latter case, human judgments are influenced by the similarity of a sample to the population from which the sample was drawn. in this field, too, researchers have proposed, debated, and experimentally investigated multiple explanations for these and other observed deviations (kahneman, 2011). the key point for present purposes is that the ideal-deviation paradigm mixes normative and descriptive elements in a way that is foreign to the way we now think of scientific research. to highlight this point, imagine biochemists in the 1950s deciding that a particular protein molecule ought to fold itself into such-and-such a three-dimension structure, for, say, aesthetic reasons. imagine also that they observe that the actual shape of the molecule deviates from this normatively specified structure, and then spend their time and theoretical energy explaining why the protein deviates in such-and-such a way from the supposedly correct structure, instead of explaining why it folds together the way it actually does. no such investigation could survive peer review for a contemporary chemistry journal. in short, given our current conception of scientific research, there is no justification for the normative element in the ideal-deviation paradigm. a descriptive theory of how people think or learn must be based on accounts of the actual processes occurring in people‟s heads when they draw inferences, make decisions, and s. ohlsson 74 | f l r revise their knowledge, regardless of whether those processes are similar to, or different from, normatively correct processes. comparing empirical observations to a normative theory contributes nothing to that enterprise. section 3 argues that normative conceptions are nevertheless at the centre of contemporary research in conceptual change. 3. ideal-deviation in conceptual change the ideal-deviation paradigm has strongly impacted psychological research on belief revision, conceptual change, theory change, and related processes. the impact is not immediately obvious, because the relevant normative theory is less precise and less explicit than the normative theories that underpin studies of logical reasoning and decision making. the normative theory of belief management can be summarized in four principles: principle 1: grounding. beliefs and concepts ought to be based on evidence. in this context, “based on” means “derived from.” the derivation is typically understood to be some form of induction across qualitative observations and/or aggregation of quantitative data. to adopt a belief for which one has no evidence is deplorable, even irresponsible, and a belief that is not grounded in evidence is dismissed as a guess, prejudice, or mere speculation. principle 2: graded conviction. beliefs ought to be held with a conviction that is proportional to the strength of the relevant evidence. for example, hearsay provides weaker evidence than direct observation; a anecdote provides weaker evidence than a study based on a representative sample; and a correlational study provides weaker evidence for a causal relation than an experimental study. the strength of one‟s convictions ought to reflect such differences in the nature and extent of the relevant evidence. for purposes of quantitative comparisons, the conviction with which a belief is held can be conceptualized as an estimate of its probability of being true. principle 3: belief-belief conflicts. when two beliefs or informal theories contradict each other, the person ought to choose to believe the one that is backed by the stronger evidence. the theory with the strongest support ought to have priority in the control of behaviour, including both discourse and action. to hold contradictory beliefs (p & not-p) is to be inconsistent and hence irrational. principle 4: belief-evidence conflicts. when beliefs are contradicted by new evidence, they ought to be revised so as to be consistent with both the old and the new evidence. failure to do so makes a person “closed minded”, “irrational”, “rigid minded”, or a victim of “robust misconceptions.” these four principles are mere common sense; this is how a rational agent ought to manage his or her beliefs. there seems to be little gain in giving such vacuous verities the status of principles. but my purpose is to make explicit what is normally too embedded in our conceptual infrastructure to be visible. elements of the normative theory of belief management, masquerading as descriptive statements, can be found throughout the cognitive sciences. for example, allport (1958/1979) proposed the contact theory of racial prejudice. the key idea was that negative racial stereotypes would be diminished if a person with such a stereotype were subjected to frequent contacts with members of the relevant ethnic group. the hypothesis was that the contacts would provide evidence against the negative stereotypes and pave the way for other, more positive opinions. in the philosophy of science, kuhn (1970) described theory change as a consequence of the accumulation of anomalies. in educational psychology, posner et al. (1982) hypothesized that students have to be dissatisfied with their current beliefs about scientific phenomena before they are prepared to revise them, and that being confronted with evidence to the contrary is the key source of dissatisfaction. in developmental psychology, gopnik and meltzoff (1997) embraced principle 4, designating belief-evidence conflicts as the main drivers of cognitive change: “theories may turn out to be inconsistent with the evidence, and because of this theories change.” (p. 39) although other processes are involved as well, the processing of counterevidence is the most important: “theories change as a result of a number of different epistemological processes. one particularly critical factor is the accumulation of counterevidence to the theory.” (p. 39) s. ohlsson 75 | f l r paradoxically, cognitive scientists confidently assert these variations of principle 4, while they simultaneously and in parallel assert that people deviate from principle 4. in discipline after discipline, researchers have observed that people do not always and necessarily revise their beliefs when confronted with contradictory evidence. the predictions of the contact theory of racial prejudice were not verified and the theory had to be reformulated (pettigrew, 1998). likewise, strike and posner (1992) found that students retain their misconceptions even after instruction that is directly aimed at confronting those misconceptions with contradicting evidence. “one of the most important findings of the misconception literature…is that misconceptions are highly resistant to change” (strike & posner, 1992, p. 153). in a review paper, limón (2001) wrote that “…the most outstanding result of the studies using the cognitive conflict strategy is the lack of efficacy for students to achieve a strong restructuring and, consequently, a deep understanding of the new information.” (p. 364) indeed, the deviation of student behaviour from principle 4 is the very phenomenon that created conceptual change as a distinct field of research, at least within educational research. consistent with the ideal-deviation paradigm, researchers have responded to the finding that people do not (necessarily) adapt their beliefs to contradictory evidence by proposing various explanations for this deviation. for example, rokeach (1960, 1970) proposed that belief systems have a hierarchical structure, and that change becomes more and more difficult as one moves from the periphery to the centre. as a result, most changes are peripheral and central principles are hardly ever affected by evidence. political and religious principles are cases in point. the philosopher imre lakatos has proposed a similar theory to explain theory change in science (lakatos, 1980). festinger (1957/1962) launched a long-lasting line of research in social psychology that centred on a set of mechanisms for reducing what he called cognitive dissonance. cognitive mechanisms for dissonance reduction process contradictory evidence without any fundamental revision of the relevant beliefs. more recently, cognitive psychologists have added yet other explanations. the category shift theory of chi (2005, 2008) and co-workers explains the robustness of misconceptions as a consequence of the inheritance of characteristics from the (frequently inappropriate) ontological category to which a phenomenon has been assimilated. vosniadou and brewer (1992) and vosniadou and skopeliti (2013) explain the deviations as a consequence of the synthesis of prior (and frequently inaccurate) mental models into more comprehensive (but sometimes equally inaccurate) mental models. sinatra and co-workers have added motivational and emotional variables as additional sources of explanation (broughton, sinatra, & nussbaum, 2013; sinatra & pintrich, 2003). yet other perspectives on conceptual change have been proposed (see, e.g., rakison & poulin-dubois, 2001; shipstone, 1984). ohlsson, (2011, chap. 9) provides a more extensive comparative analysis of these and related types of explanations. in short, although the normative theory of belief formation is less explicit than the normative theories that underpin studies of logical reasoning and decision making, research on conceptual change follows closely the ideal-deviation paradigm. the basic structure of conceptual change research is that (a) students ought to revise their misconceptions when confronted with contradictory evidence, (b) the empirical evidence indicate that they do not in fact do so, and therefore (c) we need to explain why they do not do so. but this research enterprise is only meaningful if one accepts principles 1 4 as relevant for the study of conceptual change. section 4 prepares for a new approach to conceptual change by arguing that people do not base their concepts and beliefs on evidence. if so, principles 1-4 are irrelevant for understanding conceptual change. 4. the irrelevance of evidence at first glance, the normative theory of belief management seems highly relevant for understanding human cognition. if people do not adapt their beliefs to reality, how do they get through their day? surely the deviations from rational belief management uncovered in various areas of cognitive research are relatively minor slips of a fundamentally rational cognitive system for building and maintaining a veridical belief base? such slips might be due, for example, to cognitive capacity limitations or emotional biases. s. ohlsson 76 | f l r this view is plausible but difficult to evaluate. we know very little about how people form and revise beliefs in natural settings, because there are few relevant empirical studies. what follows are some informal observations and examples. in conjunction, they suggest a radical conclusion: the principle that people base their beliefs on evidence might be fundamentally incorrect rather than an optimistic idealization or a partial truth. an adult person has a large belief base in memory, at least if the term “belief” is applied broadly enough to include not only the deep principles that tend to be the object of analysis, but also local, concrete facts. for example, i have multiple beliefs about the public transportation system in the city where i live: that there are buses and subway trains; that there are multiple subway lines; where they go; how long a trip is likely to take; how much it costs; the location of stations; and so on. this small domain of experience is likely to encompass several hundreds, perhaps even thousands, of beliefs, most of which are likely to be accurate. the view that a belief ought to be derived from, or based on, observational evidence works well with respect to such concrete, particular matters. for example, the belief that there is a subway station at the corner of x and y streets might very well be acquired by no more complicated a process than walking down x street and encountering that very station at the crossing with y street. such routine belief formation events can plausibly be attributed to direct observation and in that sense conforms to principles 1 4. the direct observation account of belief formation quickly runs into difficulties when the belief is general. for example, most adults have a variety of beliefs about economical, political, and social affairs. informal observations indicate that a significant proportion of such beliefs are not based on any evidence whatsoever. will austerity economics stimulate the economy or depress the markets by robbing consumers of their ability to consume? quite a few adults are prepared to offer a point of view about this issue, and, just as obviously, very few of them have access to relevant quantitative data or other observational evidence. this is not an isolated instance. consider the range of controversial socio-political and economic issues in the public discourse: gun control, surveillance by intelligence organizations, same-sex marriage, drone strikes on foreign soil, the benefits of universal health care – a large proportion of adults have beliefs regarding many of such issues, but almost none of those beliefs are based on evidence. that is, a person who holds a belief on such an issue did not, as a rule, induce it from multiple historical examples or derive it from statistical data or other types of observational evidence. most people cannot give any coherent or detailed account of why, how, or even when they adopted any particular belief. if people operated with principle 1, they ought to answer almost every question about socio-political and economic issues by saying, “i don‟t know; i don‟t have an opinion on that; i don‟t have enough information.” if general beliefs are not formed by induction from observations, how, by what processes are they formed instead? informal reflection on everyday life suggests that we form general beliefs by accepting what someone else tells us, either in face-to-face conversation or via media. the notion of evidence does not enter into this belief formation process in any prominent way, because we do not normally and as a rule question or doubt what we are being told. it is enough to hear someone say it for us to encode it as veridical. gabbay and woods (2001) call this the ad ignorantiam rule: “human agents tend to accept without challenge the utterances and arguments of others except where they know or think they know or suspect that something is amiss.” (p. 150) the reason for this rule is probably that we tend to communicate with people we trust, and access sources that we have already judged as reliable. but this does not support the normative principle, because it is not obvious that we base our judgments about the trustworthiness of a source on anything that would qualify as evidence. another hypothesis is that belief formation is internal to the cognitive system. many of our beliefs appear to arrive in the belief base as consequences of already adopted beliefs. for example, i believe that public education is an essential social institution. i also believe that nations that invest in education will fare better than those that do not. it would be an exaggeration to say that i have evidence for the second belief. after all, what counts as evidence as to what will happen in the future? it seems more accurate to say that i have adopted the second belief because it follows from the first. if public education is essential, nations underfund it at their peril. intra-mental derivations of this sort can hardly be characterized as evidence-based. first, the question of evidence is merely pushed one step backwards, because the derived belief cannot be s. ohlsson 77 | f l r said to be evidence-based unless the beliefs it is derived from are themselves evidence-based. second, the internal derivations are influenced by factors that are themselves unrelated to truth, such as a desire for consistency, instrumental gain, and various types of biases. consider next principle 2, that beliefs ought to be held with graded convictions that reflect the strength of the evidence. are people sensitive to the relative strength of the evidence? that is, do people in general and as a rule hold their beliefs more strongly when they are supported by more evidence and less strongly when the support is weaker? a thorough answer to this question would require extensive data collection and some way of measuring the strength of the relevant evidence. however, it is noteworthy that the beliefs that people hold with the greatest conviction tend to be their religious beliefs, and church leaders and followers alike insist that religious beliefs are, and should be, based on faith, not evidence. the fact that faith-based beliefs are held more strongly than other classes of beliefs is inconsistent with the idea that our brains are programmed, in some deep and fundamental way, to base beliefs on evidence. other examples of conviction levels that do not seem to reflect the strength of the available evidence include those that pertain to beliefs regarding climate change and the value of vaccinations. at one time, it was rational to be sceptical regarding the reality of climate change; now, the evidence is overwhelming (oreskes, 2004). nevertheless, some people continue to believe that the climate is not changing. the controversy over vaccinations exhibits a similar pattern. although caution was once rational, there are now multiple, large-scale studies that show conclusively that there is nothing wrong with the common vaccines that are given to children, or with the way they are administered. people do not get sick from vaccines; they get sick from germs. however, anti-vaccine activists continue to claim that vaccines are harmful and they have many followers (offit, 2011). finally, consider principle 3, namely that theory-theory conflicts are to be resolved with reference to the relative strengths of the evidence for the competing theories. do people consistently side with the view that has the strongest evidence? consider the issue whether human beings are fundamentally evil and require discipline in order to behave themselves, or fundamentally good, so all they need is an opportunity to blossom in a natural way. every news story about yet another serial killer is evidence for the former view; every heartwarming news story about someone who goes out of their way to make a difference for people around them is evidence for the latter view. every war produces novel atrocities, but every natural catastrophe – forest fire, hurricane, tsunami – generates a fresh batch of stories about individual heroism and self-sacrifice. anybody who attends to the news has as much evidence for one view as for the other. given that there is much evidence for either view, those of us who hold a strong opinion on the issue of human nature must have resolved the conflict between these two theories at least partially on the basis of something other than the evidence. to summarize, informal observations suggest that people do not, in general, induce their beliefs from observational evidence. although we often base concrete beliefs about particular objects and events on direct observation, we appear to form general beliefs through ubiquitous encoding of communications by trusted sources and by deriving them from other, already adopted beliefs. these hypothetical but plausible belief formation processes are not inductive in nature, and the contribution of what we normally call evidence to each is weak. in addition, people show few signs of holding their beliefs with a conviction that is proportional to the strength of the supporting evidence, resolve conflicts among competing beliefs by comparing the relative strength of the supporting evidence, or to revise their beliefs when they encounter contradictory evidence. these observations suggest that the normative theory of belief management, taken as a descriptive theory, is fundamentally wrong rather than merely an optimistic idealization. but if so, why do the observed deviations of human behaviour from principles 1-4 deserve our attention? the consequence of abandoning principles 1-4 is that conceptual change researchers no longer need to explain why misconceptions are robust in the face of contradictory evidence. if there is no reason to expect students to revise their beliefs when confronted with new evidence, then the absence of such revisions is not puzzling. explanations for why misconceptions are robust become obsolete, not in the sense of being falsified, but in the sense of being answers to a question we do not need to ask. the problem of why people do not revise their beliefs is not so much solved as dissolved. however, the task of formulating a scientific s. ohlsson 78 | f l r theory of belief formation and belief revision that can support effective pedagogical practices remains. in section 5, i propose that some of the ideas put forward by the american pragmatist philosophers can serve as a starting point for a new approach to conceptual change. 5. a pragmatist approach at the end of the 19 th century and the beginning of the 20 th , american scholars, lead by william james, charles sanders peirce, and john dewey, tried to reformulate the classical philosophical problems about knowledge, meaning, and truth in terms of action instead of observation. they claimed that the meaning of a concept or belief resides in the set of actions or “habits” to which it gives rise. “the essence of belief is the establishment of a habit, and different beliefs are distinguished by the different modes of action to which they give rise.” (peirce, 1878, p. 129-130) the truth of a belief is tied to the outcomes of executing those habits. they stopped short of claiming that, “what works is what is true”, but some of their contemporaries stated their ideas even more boldly than they did themselves (schiller, 1905). pragmatism did not flourish as a philosophical, i.e., normative, theory. its impact faded after the demise of its most charismatic leaders. although it is once again receiving serious attention from philosophers (stich, 1983), my purpose is not to revive philosophical pragmatism. instead, i intend to mine this strand of thought for an approach to cognitive change that does not begin with the assumption that people decide what to believe by estimating the probability of truth. the question is what they estimate instead. 5.1 cognitive utility as the basis for cognition the pragmatist emphasis on action fits well with psychological theories of cognition. there is broad consensus on certain general features of what cognitive psychologists have come to call the cognitive architecture, i.e., the information processing machinery that underpins the higher cognitive processes (polk & seifert, 2002). at the centre of the cognitive architecture there is a limited-capacity working memory, connected to separate long-term memory stores for declarative and practical (skill) knowledge. the working memory receives input from sensory systems, and holds information that is being processed in reasoning and decision making. the purpose of the cognitive system is to generate behaviour that satisfies the person‟s current goal. in the process, the system makes endless, lightening quick choices: which goal to pursue next (planning); which part of the environment to attend to next (attention allocation); which interpretation of perceptual input to prefer (perception); which memory structure to activate next (retrieval); which inference to carry out next (reasoning); and which change, if any, to make in the system‟s knowledge base at any given time (learning). the pragmatist stance invites the hypothesis that the variable that guides the never-ending choices is the cognitive utility of the relevant knowledge structures. to articulate this idea, imagine that each knowledge structure (concept or belief) in memory is associated with a numerical value that measures its past usefulness. when there is a choice to be made among knowledge structures, the one with the higher utility is preferred and gets to control discourse and action. if a knowledge structure is instrumental in generating a particular action, and if that action is successful, then the utility of that knowledge structure is adjusted upwards; if the action is unsuccessful, it is adjusted downwards. each application of a knowledge structure is an opportunity for that structure to accrue utility (or to loose some of it, in the case of unsuccessful action). over time, the value of the cognitive utility associated with a knowledge structure will become stabilized at some asymptotic value that estimates its usefulness in general. the distribution of utility values over the belief base represents the person‟s experience of the world, as filtered through action rather than perception. the cognitive utility hypothesis is not novel. a construct of this sort has been incorporated into the act-r model of the cognitive architecture proposed by john r. anderson and co-workers (anderson, s. ohlsson 79 | f l r 2007). in act-r, cognitive skills are encoded in sets of goal-situation-action rules (skill elements) that specify an action to be considered when certain conditions are satisfied by the current situation. in each cycle of operation, the architecture retrieves all the rules that have their conditions satisfied. it then selects one of those rules to be executed; that is, its action is taken. the action usually changes the current situation, and the cycle starts over with a renewed evaluation of which rules have their conditions satisfied in the changed situation. in act-r, the utility u of rule i determines its probability of being selected for execution. in simplified form, that probability is given by eq. (1) prob(i) = u(i) / σ u(1, 2,…i,…j), where σ u(1,2,…i,…j) is the sum of the utilities of the rules for which the conditions are satisfied by the current situation. the probability that a particular rule i will be selected is thus proportional to how much of the total utility represented by all the currently satisfied rules it accounts for. the probability of being chosen for execution is thus a dynamic quantity that depends on context and that changes from moment to moment as cognitive processing unfolds. if rule i is selected and executed on operational cycle n, its utility is adjusted upwards or downwards, depending on the outcome. the adjustment is given by the equation eq. (2) u(i, n) = u(i, n-1) + α[r(i, n) – u(i, n-1)], in which i is the relevant rule, n is the operational cycle, and r is the reward or feedback from the environment about the success of the executed action (reinforcement in the behaviourist sense). the magnitude r(i, n) – u(i, n-1) is the reward the rule realized in cycle n, r(i, n), over and above the utility it already possessed in the previous cycle of operation, u(i, n-1). the rate parameter α controls the proportion of that reward increment that is to be added to the current utility of the rule, u(i, n-1), to compute the utility of the rule in the following cycle, u(i, n). the reader is recommended to consult the original source for further technical details (anderson, 2007, pp. 159-164). in the act-r theory, utility values are associated with skill elements (rules), and there is a separate system of theoretical quantities that pertain to the learning and application of declarative knowledge elements. to make the utility construct relevant for belief formation we have to hypothesize that utility values are associated with declarative knowledge structures (beliefs, concepts, informal theories) instead of (or in addition to) skill elements. furthermore, the relation between cognitive utility and belief (subjective truth) has to be specified. one possible hypothesis is that there is a threshold such that, when the utility of a particular belief rises above that threshold, the person feels that the belief is true. a cognitive system that operates in this way would be significantly different from act-r and other cognitive systems described in the cognitive literate (polk & seifert, 2002). a key question is what degree of utility new information will be assigned when it is first encoded into memory. at the outset, the new knowledge structure has no track record of supporting successful action, so one might decide that its initial utility is zero. this causes a paradox: if it is zero, it will always have lower utility than any competitor with even a modest track record, so it will never be activated or chosen, and therefore never have an opportunity to accrue utility. to an outside observer, it will appear as if the learner did not encode the new information, because his or discourse and action continue to be guided by other knowledge structures. there are multiple solutions to this theoretical problem. in anderson‟s act-r theory, the initial value is indeed set to zero, but a knowledge structure (rule) can be created multiple times, and each time the utility value is increased. other solutions are possible. the initial value can be hypothesized to be random, or equal to the mean of the utility values of all knowledge structures in memory. there might be situations in which s. ohlsson 80 | f l r competing older knowledge structures do not apply, but the new one does, and those situations afford the newer knowledge with opportunities to accrue utility. many scenarios that seem like straightforward instances of truth-based processing are equally or better understood in terms of utility. for example, suppose that my eyes itch. i might have dry eyes, or i might suffer from an allergy attack. i decide to take an antihistamine pill. the itch disappears. in a logic-inspired analysis, the belief that i am suffering from an allergy outbreak is a hypothesis the truth of which is unknown. the connection between the belief that i have an allergy and the prediction that the itch will disappear is a step-by-step chain of inferences. the disappearance of the itch is an observation that verifies the hypothesis, and my estimate of the probability that i have an allergy increases as specified by, for example, bayesian principles. this account has weaknesses. one weakness is that i am not aware of any lengthy reasoning process to arrive at a testable prediction. the process that connects the belief “i have an allergy attack” with the fact that my itch stopped is a process of problem solving and planning (what should i do about my itchy eyes?), not a process of propositional inference. another weakness is that the envisioned process is an instance of a logical fallacy: if p, then q in conjunction with q does not imply p. this is popper‟s classical critique of verificationism. but if the truth-based account has a logical fallacy at its core, how can people function? the utility-based account avoids this problem by postulating a direct link between the action outcome and the relevant belief: the action of taking the antihistamine worked, so my disposition to act on the allergy belief in the future is increased. as the example illustrates, the difference between an account in terms of evidence, inference, and truth, on the one hand, and an account in terms of utility and action, on the other, can be subtle. how does that difference affect how we view conceptual change? the pragmatist stance focuses attention on action, the output side of the cognitive system, instead of perception, the input side. what the learner does matters more than what he or she hears or sees. passive reception of information will not in and of itself have any cognitive consequences. unless the learner retrieves a knowledge structure and uses it to decide what to do next, that knowledge structure cannot accrue utility and hence might remain dormant, even though the new information has been encoded accurately. in the pragmatist perspective, new information does not replace the old. in a logic-based theory, two different beliefs can be mutually incompatible, which implies that a person cannot embrace both. the earth is either round or flat; it is impossible to believe both assertions at once. however, the fact that knowledge structure i has utility u(i) is not incompatible with the fact that knowledge structure j has utility u(j). the belief that the earth is flat might be useful for mapmaking purposes, while the belief that the earth is round might be more useful for the purpose of circumnavigation. many tasks in real life admit of multiple solutions, varying with respect to goal satisfaction, efficiency, and range of applicability. evaluating beliefs with respect to their cognitive utility is thus very different from evaluating them with respect to their truth. falsification by contradictory evidence is, in principle, a one-shot affair. a single application of modus tollens is logically sufficient to bring down a belief and even an entire theory. but utility-based belief revision is necessarily a gradual matter. once the utility rises to the point where a new knowledge structure is chosen to be the basis for action on at least some occasions, belief change is contingent on the outcomes of the resulting actions. the utility of structure i might be steadily raising with each application, while the utility of some competing knowledge structure j is gradually dropping. eventually, the utility of the newer knowledge structure will surpass that of the older, competing structures, and rise above the threshold of belief. if changes in utility values are incremental, then this process is necessarily gradual. the most radical difference between a truth-based and a utility-based account of cognition pertains to the trigger of conceptual change. in the truth-based account, it is the failure of the older knowledge that drives belief revision. change happens because already acquired concepts and beliefs have been found to be false, triggering dissatisfaction and a search for more veridical concepts and beliefs to replace them. if there is no failure, there is no push for change. in the utility-based account, on the other hand, new information need not wait for falsification or dissatisfaction with prior beliefs. it is the success of the newer concepts and beliefs that drives the change. no dissatisfaction with the old belief is required, only a recognition that the s. ohlsson 81 | f l r newer belief is an even more useful basis for action. change is driven by success, not failure (ohlsson, 2009, 2011). however, before the utility-based perspective can be adopted, some plausible objections must be dealt with; this is the task of section 6. 6. two objections the purpose of this section is to address two objections that must have occurred to the reader. the first is that evolution through natural selection ought to have pushed human cognition in the direction of the normative theory of belief management (principles 1-4), and the second is that the behaviour of scientists appears to conform to the normative theory. 6.1 natural selection for truthfulness? one might argue that the shift from estimates of the probability of truth to estimates of cognitive utility is unimportant. after all, how can a belief be useful unless it is, in fact, true? if only true beliefs are useful, then the selective pressures that drove the evolution of human cognition must have pushed the belief management processes in the learner‟s head to conform at least approximately to principles 14. how could our hunter-gatherer ancestors have survived unless their beliefs corresponded to reality? the instrumental value of veridicality in the struggle for survival implies that the human cognitive architecture is designed to derive beliefs from evidence. but natural selection cannot have operated directly on the truthfulness of beliefs. the probability of surviving long enough to mate and to raise the resulting offspring to reproductive age is a function of how the individual behaves, not on how he or she thinks. what mattered during human evolution cannot have been the truth of beliefs per se, but the effectiveness of human behaviour. consistent selection in the direction of effective action would create a utility-based rather than truth-based system. the distinction between truth and utility would be of minor importance, if the two were perfectly correlated. however, false beliefs can lead to successful action. for example, it does not matter what belief one has about the causes of severe weather, as long as that belief implies that when storm clouds gather, it is time to seek shelter. the belief that lightening is a sign of the anger of the gods and the belief that it is an electrical discharge are equally good reasons to get out of the way. the belief that a certain medical condition is caused by an evil spirit and that the spirit can be exercised by ingesting a certain herb can be as successful as an account of the disease in terms of bacteria, white blood cells, etc., if the relevant herb contains traces of, for example, an antibiotic substance. an even stronger example is provided by the 14 th century physicist buridan‟s impetus theory of mechanical motion (claggett, 1959; robin & ohlsson, 1989). a central principle in this theory says that to keep an object in motion requires the continuous application of force, the opposite of the principle of inertia that is at the centre of newtonian mechanics. however, the impetus principle holds on the surface of the earth due to the universal presence of friction. if the goal is to keep an object moving, or to make it move further or faster, the impetus concept is as useful a guide to action as the theory that physicists teach (apply more force). in short, truth and utility are only partially correlated, and evolution has no way of selecting for the truth of beliefs directly, but only for the success of an individual‟s struggle for survival. evolutionary considerations thus support rather than contradict the hypothesis that utility is the key variable in belief formation. s. ohlsson 82 | f l r 6.2 the behaviour of scientists the reader would be excused for thinking that the present author is engaged in a self-defeating enterprise: to use evidence and arguments to make the reader believe that people do not use evidence and arguments when deciding what to believe. this article is itself an attempt to base belief in this matter on evidence. more generally, scientists do base their theories on evidence and scientists are people, so it seems unreasonable to claim that this is not a common cognitive capability. gopnik and meltzoff (1997) has emphasized this connection between the procedures of scientific knowledge creation and individual belief formation: “the central idea of [our] theory is that the processes of cognitive development in children are similar to, indeed perhaps even identical with [sic], the processes of cognitive development in scientists.” (p. 3) indeed, they have stated their hypothesis quite clearly: “…the most central parts of the scientific enterprise, the basic apparatus of explanation, prediction, causal attribution, theory formation and testing, and so forth, is not a relatively late cultural invention but is instead a basic part of our evolutionary endowment.” (pp. 20-21) the consequence is that cognition can be explained with normatively correct processes such as bayesian inference (gopnik et al., 2004). the utility-based view explored in this article does not deny that people can acquire the higher-order cognitive skills needed to engage in the methods and procedures of science. it does claim that those methods and procedures are acquired. scientists are professional theorizers; they engage in belief formation (a.k.a. hypothesis testing) deliberately and on purpose, with a high degree of awareness. to be able to do this, they undergo a multi-year training process called graduate school. they are supported by a wide variety of tools such as special-purpose statistical software that embody the principles of the normative view of belief management. furthermore, scientific research takes place within a social context, the scientific discipline, that enforces adherence to the normative theory. for example, a scientist who revises his or her theory to improve its fit to empirical data (principle 4) is more admired than someone who continues to advocate a favourite theory in the face of counterevidence. the behaviour of scientists shows that people can acquire the high-level skills needed to function at least approximately as prescribed by the normative theory. but this does not imply that the basic processes of the cognitive architecture conform to the normative theory. to cast the procedures of science as a description of conceptual change in the individual is to confuse two levels of description: the level of the basic processes of the cognitive architecture (the “basic part of our evolutionary endowment”), on the one hand, and the level of acquired higher-order strategies and skills, on the other. the arguments put forth in this paper concern the basic processes. i know of no reason to believe that “the most central parts of the scientific enterprise, the basic apparatus of explanation, prediction, causal attribution, theory formation and testing” is part of our “evolutionary endowment.” the late arrival of science in human history, its invention by one culture at one time, and the extensive training individuals need to conduct scientific research make it highly implausible that anything like the “basic apparatus” of science is among our “evolutionary endowment.” instead, the cognitive apparatus of science is precisely “a relatively late cultural invention.” the relation between the basic processes of cognitive change and the procedures of science is the opposite of the one claimed by gopnik and meltzoff (1997). rather than scientific practices explaining how cognitive change happens in children and lay adults, the relationship should be construed the other way around: a theory of the basic cognitive processes should explain how it is possible to acquire the higherorder strategies for belief management that approximate the normative theory in principles 1-4. the utilitybased perspective have other implications as well, three of which are discussed in section 7. 7. implications if we adopt the utility-based perspective, what follows? from the point of view of basic research on conceptual change, it implies a re-evaluation of existing theoretical constructs, methodologies, and applications. traditionally, research on conceptual change and belief formation has been perception-centric: the focus has been on what the learner sees and hears, and how he or she processes the perceived s. ohlsson 83 | f l r information. the utility-based perspective, in contrast, implies a need to focus on what the learner does, when and where he or she succeeds or fails, and on what information is activated, retrieved, and used to guide action. a learning trajectory is primarily to be defined in terms of tasks undertaken, and only secondarily in terms of information encountered. as a side effect of such a re-focusing, the traditional concepts, tools, and puzzles regarding truth inherited from philosophy and logic will become comparatively less important. the perception-centric bias of cognitive research in general and cognitive studies for education in particular is driven, in part, by the practicalities of psychological experimentation. the experimenter controls the subject‟s task environment, so he or she can create complex but well-specified conditions and contrasting situations by varying the stimulus. such variations can easily be described in research reports. the subjects‟ behaviours, on the other hand, are only easy to report and interpret in an intersubjectively valid way if they consist of simple, easy-to-record events, like pushing a button or placing a mark on a rating scale. the pragmatist perspective implies that this style of empirical inquiry runs the risk of eliminating from the researchers‟ consideration the central subject matter of cognition, namely complex, temporally extended, hierarchically structured, and dynamically coordinated sequences of actions in the service of human goals and objectives. the pragmatist perspective implies a need for a period of methodological innovation in which researchers develop new techniques to record and interpret complex behaviours. from the point of view of instructional application, the utility-based account poses multiple challenges: how to stimulate students to encode knowledge that they have no reason to believe, and that is only tangentially relevant for their own action? how to design situations in which the new knowledge presented in the course of instruction, but not their prior knowledge, applies, so that the new knowledge can accrue utility? how to provide learners with multiple opportunities to apply new knowledge without resorting to mind numbing drill and practice? these questions are quite different from the questions of why misconceptions are robust, what evidence will convince a student that his or her misconceptions are in fact inaccurate, or how to train students to pay attention to evidence, so pursuing them will likely lead educational researchers in novel directions. 8. conclusion throughout the history of science, interdisciplinary work has often been innovative and path breaking. at the beginning of conceptual change research, there was every reason to believe that drawing upon a variety of disciplines was a productive way to proceed. however, researchers (including the present author) overlooked the distinction between normative and descriptive disciplines, fell into the ideal-deviation paradigm, and spent their theoretical energies explaining the main observed deviations from the normative theory: that students do not revise their prior conceptions when confronted with counterevidence. but the normative idea that people ought to base their beliefs on evidence is irrelevant for the empirical study of cognition. there is little or no evidence that people base any of their beliefs on evidence, and considerable evidence that they do not. if they do not, then it is no surprise that science courses fail to impact students‟ beliefs about scientific phenomena, and efforts to explain this supposed phenomenon are unnecessary. to make progress in understanding conceptual change, researchers need to adopt a resolutely naturalistic approach that makes no normatively inspired assumptions about belief formation and belief revision. the pragmatist view that cognition evolved to support successful action and that beliefs are evaluated on the basis of their cognitive utility instead of their probability of being true is an alternative starting point for conceptual change research. the utility-based perspective implies that action is necessary for conceptual change, that old and new beliefs are not mutually incompatible, that conceptual change is necessarily gradual, and that change is not driven by the failures of misconceptions but by the successes of better ideas. a research program to articulate this perspective would replace the traditional perception-centric bias of psychological research with an action-centric approach that forefronts the cognitive consequences of complex actions. s. ohlsson 84 | f l r references allport, g. w. (1958/1979). the nature of prejudice (2 nd ed.). reading, ma: addison-wesley. anderson, j. r. (2007). how can the human mind occur in the physical universe (pp. 159-165)? oxford up. broughton, s. h., sinatra, g. m., & nussbaum, e. m. (2013). “pluto has been a planet my whole life!” emotions, attitudes, and conceptual change in elementary students‟ learning about pluto‟s reclassification. research in science education, 43, 529-550. butterfield, h. (1957). the origins of modern science 1300-1800 (revised ed.). indianapolis, in: hackett. carey, s. (2009). the origin of concepts. new york: oxford university press. chi, m. t. h. (2005). commonsense conceptions of emergent processes: why some misconceptions are robust. the journal of the learning sciences, 14, 161-199. chi, m.t.h. (2008). three types of conceptual change: belief revision, mental model transformation, and categorical shift. in s.vosniadou (ed.), handbook of research on conceptual change (pp. 61-82). hillsdale, nj: erlbaum. claggett, m. (1959). the science of mechanics in the middle ages. madison, wisconsin: university of wisconsin press. duit, r., & treagust, d. f. (2003). conceptual change: a powerful framework for improving science teaching and learning. international journal of science education, 25(6), 671-688. evans, j. st. b. t. (2007). hypothetical thinking: dual processes in reasoning and judgment. new york: psychology press. festinger, l. (1957/1962). a theory of cognitive dissonance. stanford, ca: stanford university press. frank, p. (1952). the origin of the separation between science and philosophy. proceedings of the american academy of arts and sciences, 80(2), 115-139. gabbay, d., & woods, j. (2001). the new logic. logic journal of the interest group in pure and applied logics, vol. 9, pp. 141-174. gopnik, a., & meltzoff, a. n. (1997). words, thoughts, and theories. cambridge ma: mit press. gopnik, a., glymour, c., sobel, d. m., schulz, l. e., & kushnir, t. (2004). a theory of causal learning in children: causal maps and bayes nets. psychological review, 111, 3-32. johnson-laird, p. n. (2006). how we reason. new york: oxford university press. kahneman, d. (2011). thinking, fast and slow. new york: farrar, straus, and giroux. klauer, k. c., stahl, c., & erdfelder, e. (2007). the abstract selection task: new data and an almost comprehensive model. journal of experimental psychology: learning, memory, and cognition, 33, 680-703. kuhn, t. s. (1957). the copernican revolution: planetary astronomy in the development of western thought. new york: random house. kuhn, t. s. (1970). the structure of scientific revolutions (2 nd ed.). chicago, il: university of chicago press. köhler, w. (1938/1966). the place of value in a world of facts. new york: liveright. lakatos, i. (1980). philosophical papers (vol. 1): the methodology of scientific research programmes). cambridge, uk: cambridge university press. limón, m. (2001). on the cognitive conflict as an instructional strategy for conceptual change: a critical appraisal. learning and instruction, 11, 357-380. margolis, h. (1987). patterns, thinking, and cognition: a theory of judgment. chicago, il: university of chicago press. margolis, h. (1993). paradigms and barriers: how habits of mind govern scientific beliefs. chicago, il: university of chicago press. nersessian, n. j. (2008). creating scientific concepts. cambridge, ma: mit press. offit, p. a. (2011). deadly choices: how the anti-vaccine movement threatens us all. new york: basic books. ohlsson, s. (2009). resubsumption: a possible mechanism for conceptual change and belief revision. educational psychologist, 44, 20-40. ohlsson, s. (2011). deep learning: how the mind overrides experience. new york: cambridge university press. s. ohlsson 85 | f l r osler, m.j., (ed.), (2000). rethinking the scientific revolution. cambridge, ny: cambridge university press. oreskes, n. (2004). beyond the ivory tower: the scientific consensus on climate change. science, 306, 1686. peirce, c. s. (1878). how to make our ideas clear. popular science monthly, vol. 12, pp. 286-302. [reprinted in n. houser and c. kloesel (eds.), the essential peirce: selected philosophical writings (vol. 1, pp. 124-141). bloomington, in: indiana university press.] pettigrew, t. f. (1998). intergroup contact theory. annual review of psychology, vol. 49, pp. 65-85. polk, t. a., & seifert, c. m., (eds.), (2002). cognitive modeling. cambridge, ma: mit press. posner, g. j., strike, k. a., hewson, p. w., & gertzog, w. a. (1982). accommodation of a scientific conception: toward a theory of conceptual change. science education, 66, 211-27. putnam, h. (2002). the collapse of the fact/value dichotomy; and other essays. cambridge, ma: harvard university press. robin, n., & ohlsson, s. (1989) impetus then and now: a detailed comparison between jean buridan and a single contemporary subject. in d. e. herget (ed.), the history and philosophy of science in science teaching. proceedings of the first international conference (pp. 292-305). tallahassee: florida state university, science education & dept. of philosophy. rakison, d. h., & poulin-dubois, d. (2001). developmental origin of the animate-inanimate distinction. psychological bulletin, 127(2), 209-228. rokeach, m. (1960). the open and closed mind. new york: basic books. rokeach, m. (1970). beliefs, attitudes, and values: a theory of organization and change. san francisco, ca: jossey-bass. schiller, f. c. s. (1905). the definition of „pragmatism‟ and „humanism‟. mind, 14, 235-240. shipstone, d. m. (1984). a study of children‟s understanding of electricity in simple dc circuits. european journal of science education, 6, 185-198. sinatra, g. m., & pintrich, p. r., (eds.), (2003). intentional conceptual change. mahwah, nj: lawrence erlbaum. stitch, s. p. (1983). from folk psychology to cognitive science: the case against belief. cambridge, ma: mit press. strike, k. a., & posner, g. j. (1992). a revisionist theory of conceptual change. in r. a. duschl and r. j. hamilton (eds.), philosophy of science, cognitive psychology, and educational theory and practice (pp. 147-176). new york: state university of new york press. thagard, p, (1992). conceptual revolutions. princeton, nj: princeton university press. vosniadou, s., baltas, a., & vamvakoussi, x., (eds.), (2007). reframing the conceptual change approach to learning and instruction. amsterdam, the netherlands: elsevier science. vosniadou, s., & brewer, w.f. (1992). mental models of the earth: a study of conceptual change in childhood. cognitive psychology, 24, 535-585. vosniadou, s., & skopeliti, i. (2013). conceptual change from the framework theory side of the fence. science & education, doi 10.1007/s11191-013-9640-3. wason, p. c., & johnson-laird, p. n. (1972). psychology of reasoning: structure and content. london, uk: b. t. batsford. watson, j. d., & crick, f. h c. (1953). a structure for deoxyribose nucleic acid. nature, vol. 171, pp. 737738. frontline learning research 3 (2014) 31-49 issn 2295-3159 corresponding author: liisa postareff, p.o.box 9, 00014 (centre for research and development of higher education) university of helsinki, finland, liisa.postareff@helsinki.fi http://dx.doi.org/10.14786/flr.v2i1.63 31 | f l r explaining university students’ strong commitment to understand through individual and contextual elements liisa postareff a , sari lindblom-ylänne a & anna parpala a a university of helsinki, finland article received 20th september 2013 / revised 14th february 2014 / accepted 20th february 2014 / available online 25th april 2014 abstract since the late 1970s numerous studies have explored students‟ approaches to learning (referred to as the „sal‟ tradition). these studies have provided valuable evidence of students‟ study strategies and intentions at the university. since extensive research already exists on students‟ approaches to learning, there is a need to move forward and analyse student learning from new perspectives. in the present in-depth qualitative study, we analyse interviews of 34 students who scored extremely highly on the deep approach scale in a pre-test in our previous quantitative study (lindblom-ylänne, parpala & postareff, 2013) and thus are likely to have a strong commitment to understand, and a „disposition to understand for oneself‟ which is a recently introduced, yet unexplored phenomenon (see entwistle & mccune, 2009; mccune & entwistle, 2011). we identified several individual and contextual elements which provided explanations for the students‟ high scores on the deep approach, as well as for the increase, decrease or stability in their deep approach during one course. the results showed that most students showed a strong commitment to understand, but those whose deep approach sharply decreased during the course showed less commitment and their descriptions revealed problems with, for example, study skills, time management and regulation of learning. however, contextual elements such as the students' experiences of the course teaching and their interest in the course content did not clearly provide explanations for the changes in the deep approach. elements of a 'disposition to understand for oneself‟ clearly emerged among students whose deep approach did not decrease, or decreased only slightly. keywords: approaches to learning, disposition to understand for oneself, commitment to understand, higher education l. postareff et al. 32 | f l r 1. introduction numerous studies which have been conducted within the sal (students approaches to learning; see lonka, olkinuora, & mäkinen, 2004) tradition have identified three qualitatively different approaches to learning: the deep and surface approaches (e.g., biggs, 1987; entwistle & ramsden, 1983; marton & säljö, 1976, 1997) and strategic approach or organised studying (e.g. biggs, 1987; entwistle & ramsden, 1983). a deep approach to learning has been shown to be related to high-quality learning outcomes (diseth 2003; watters & watter 2007), and therefore university students are encouraged to aim at constructing meaning and develop deep understandings of the study content. however, research has shown that students vary to a great extent with regard to their approaches to learning in that some are more likely to adopt deep approaches while others will rely more on surface learning. moreover, most students‟ approaches seem to vary depending on the context (e.g. vermunt, 1998; nieminen et al., 2004, lindblom-ylänne et al., 2012). recent research has further suggested that some students continually aim at a deep understanding and thus have a disposition to understand for oneself (entwistle and mccune 2009; mccune and entwistle, 2011). the disposition to understand is a more consistent and stronger form of the „intention to understand‟ found in a deep approach to learning. students with such a disposition aim at reaching a full and satisfying understanding of what they study. entwistle and mccune suggest that a disposition to understand is an important characteristic for students to develop at university in order to cope with the uncertainty and complexity of society in the future. the disposition to understand contains three central elements: 1) a welldeveloped use of learning strategies which concentrate on relating ideas, the critical use of evidence and attention to detail; 2) a willingness to devote the necessary time, effort, and concentration to apply the learning strategies effectively; and 3) an alertness to the learning context (entwistle & mccune 2009; mccune & entwistle, 2011). these three elements are similar to the ones perkins and tishman (2001) identified when exploring „thinking dispositions‟. thinking dispositions are stable ways of reacting to situations and thinking critically, and are comprised of three components needed in carrying out intellectual tasks: willingness to apply effort, ability to perform the task effectively, and alertness to situations in which thinking is required. entwistle and mccune (2009) took this avenue, when analysing students‟ approaches to learning and disposition to understand. the disposition to understand is a broader concept than the deep approach to learning. while a deep approach describes the student‟s intention to understand the content of study and the use of effective learning strategies, such as relating ideas and using evidence (e.g. entwistle & ramsden 1983; marton & säljö, 1997), the disposition to understand focuses more broadly on a specific discipline as a whole. a student with a strong disposition to understand for oneself shows an emotional commitment to continuously strive towards understanding and to monitor the development of understanding the contents of study (entwistle & mccune 2009; mccune & entwistle, 2011). empirical studies on the disposition to understand for oneself are rare. still, in a recent study entwistle and mccune (in press) analysed nearly 2000 students from undergraduate courses and aimed to identify students who showed high and consistent scores for a deep approach to learning, as well as for organised effort and monitoring studying. they found one cluster of students whose scores were high on all these scales and remained stable over time, thus showing characteristics of a disposition to understand. the other clusters also scored high on the deep approach, but in addition showed surface elements, or lower scores on organised effort or monitoring studying. to our knowledge, this quantitative study is the only empirical one conducted on students‟ disposition to understand. what especially differentiates a disposition to understand from a deep approach is that the former is characterised as more stable (entwistle and mccune 2009; mccune & entwistle, 2011) while the latter has been characterised as more contextual and changeable (e.g., vermunt, 1998; nieminen et al., 2004, lindblom-ylänne et al., 2013). however, opposite views have also been presented, as the deep approach has been claimed to be a relatively stable construct (e.g., lietz & matthews, 2010; zeegers, 2001) but the mainstream conception seems to rely on the original view of the contextual and dynamic nature of the approaches presented by marton and säljö already in 1976. a deep approach to learning seems to be difficult to induce since a number of quantitative studies have shown a decrease in the deep approach after varying l. postareff et al. 33 | f l r study periods. some studies have shown a decrease in the deep approach after a three-year study period (biggs, 1987; watkins & hattie, 1985) and others have shown a decrease after shorter course units (e.g., lindblom-ylänne et al., 2013). evidence shows that inducing a deep approach to learning might be difficult even in student-centred learning environments (e.g. gijbels, segers, & struyf, 2008; struyven, dochy, janssens, & gielen, 2006). however, in a number of quantitative studies it has been shown that satisfaction with the quality of a course as well as a student-centred approach taken by teachers has been shown to enhance the application a deep approach (trigwell, prosser, & waterhouse., 1999; see also baeten, kyndt, struyven, & dochy, 2010). in addition, individual reasons such as intrinsic motivation, self-confidence, strong self-efficacy beliefs and openness to experience have been shown to enhance the adoption of a deep approach (baeten et al., 2010; kyndt, dochy, cascallar, & struyven, 2011). the changes in a deep approach and the individual and contextual factors affecting its adoption have thus been explored in a number of quantitative studies, but to our knowledge no qualitative research combining these two elements has been carried out. this study presents an in-depth qualitative analysis of the factors explaining the stability or changes in a deep approach to learning among students who scored extremely highly on a deep approach scale in the pre-test. the existence of the disposition to understand for oneself is also explored. thus the study provides both methodologically and theoretically a fresh perspective on research on student learning in higher education. 1.2 aims of the study the study aims to more deeply understand the results from previous research of ours which concentrated on analysing with quantitative methods group-level changes in one‟s deep approach during different courses (lindblom-ylänne et al., 2013). the pre-test/ post-test results showed large individual variation in all disciplinary contexts in terms of the amount and direction of change in students‟ deep approach to learning. the present study concentrates on analysing interviews of students who scored the highest on a deep approach scale in the pre-test. the aim is to identify the individual and contextual elements which are related to a strong commitment to understand as shown in the pre-test, and also to the stability or decrease in one‟s deep approach during the course as shown in the post-test. in addition, the study provides a new perspective on analysing approaches to learning as it focuses on exploring the existence of the disposition to understand for oneself. thus the second aim is to analyse how a disposition to understand for oneself emerged from the interviews. previous work on the disposition to understand for oneself has not been based on coherent evidence, and has not been empirically conducted with diverse samples of students (see entwistle & mccune, 2009). thus the present study aims to provide empirical evidence of the disposition to understand through an in-depth qualitative analysis of student interviews. 2. materials and methods the study adopts a mixed-methods approach as it has its roots in the quantitative results, which are then explained through analysing student interviews. mixed methods research is a type of research in which elements of qualitative and quantitative research approaches are combined in order to gain a deep and broad understanding of the phenomenon under study (see e.g. johnson, onwuegbuzie & turner, 2007). cresswell (2009) emphasises that the use of mixed methods approach has its roots in pragmatism, in which the most important thing is to focus attention on the research problem itself and use pluralistic approaches to derive knowledge about the problem. an advantage of using mixed methods is that the results from one method can help develop or inform other methods. in the present study we adopted a sequential procedure through which we attempted to elaborate on and expand the quantitative findings with the qualitative ones (see creswell, 2013). in addition, the study adopts, along with a traditional variable-oriented technique, a person-oriented approach, which assumes that human behavior is affected by several factors and the interplay between these l. postareff et al. 34 | f l r several factors forms unique profiles of individuals (see vanthournout et al., 2013). the person-oriented approach is sometimes considered as an opposite to variable-oriented techniques, but they should be more viewed as complementary techniques (vanthournout et al. 2013). in the present study both techniques are used to provide complementary information (see section 2.3). 2.1 participants and contexts of the study the participants of the study were selected among 277 bachelor students from the university of helsinki, who took part in our previous quantitative study (lindblom-ylänne et al., 2013). the 34 students were selected on the basis of having scored extremely highly on a scale measuring their deep approach to learning. the students were categorised into three groups on the basis of how their deep approach to learning changed between two measurements (at the beginning and at the end of a course). the present study focuses on analysing interviews of the students in these three change groups. in the following chapter, we provide some insight into our previous quantitative study and into the procedure of creating the change groups in order to create an understanding of the inventory results and of the differences between the three groups of students. a more detailed description of the quantitative analyses is provided in our previous study (lindblom-ylänne et al., 2013). 2.1.1 background for selecting the participants for the previous studies, the data were collected from 10 bachelor-level courses in five different disciplines. the students of the courses completed the approaches to studying and learning inventory (alsi; see section 2.2.) at the beginning and at the end of a lecture course. at the beginning of the course the students were asked to consider how they had studied in their major up until then, and at the end of the course they were asked to focus on how they had studied in that particular course. their scores at the beginning and at the end of the course were used to explore changes in their approaches to learning between the two measurements at the beginning of the courses the mean of the deep approach scale in each discipline varied from 3.42 to 3.66 on a scale of 1 to 5, and at the end of the courses the range was from 3.07 to 3.24. in each discipline, the decrease in the deep approach between the two measurements was statistically significant. a paired-samples t-test showed that the t-values ranged from 2.25 to 4.86 and the p-value varied between 0.001 and 0.028. among bioscience students the decline was the lowest being -0.15, while among mathematics students the decline was the highest being -0.35. after exploring the differences between the two measurements at the group level, we wanted to explore the changes at a more individual level. we computed the deep approach change variables by subtracting the students‟ scale scores at the end of the course from their scale scores at the beginning of the course. the magnitude and the direction of the change served to create five groups of change (see table 1): strong increase, increase, no change, decrease and strong decrease. the distributions of the change variables were explored in detail in order to decide upon the best cutting points for the change groups. our decision was to create the change groups on the basis of likert-scale point changes, which is a procedure followed previously by lindblom-ylänne, trigwell, nevgi & ashwin (2006). the benefit of using change variables is that they mirror the absolute changes instead of relational changes. our purpose was to focus on the changes of individual students regardless of other students‟ values. we used a quarter of a likert scale as the cutting point between the change groups in order to separate different types of changes in as detail as possible. a half of a likert scale would have been too robust, since a decrease of 0.15 in the deep approach was already statistically significant. we also considered categorisation on the basis of standard deviation or median split, but these are based on relational values and they would have rendered a comparison of the changes between the three approaches impossible (which was the focus in our previous study, see lindblom-ylänne et al., 2013). for the present study, the 34 students were divided into three subgroups on the basis of the change variable categories: 1) deep approach remains high (includes the „no change‟ and the „increase‟ groups), 2) slight decrease in one‟s deep approach and 3) sharp decrease in one‟s deep approach. l. postareff et al. 35 | f l r table 1 change variable categories the direction of change differences in scores sharp increase in the approach to learning 0.50 or higher slight increase in the approach to learning from 0.25 to 0.49 no change from -0.24 to 0.24 slight decrease in the approach to learning from -0.25 to -0.49 sharp decrease in the approach to learning -0.50 or lower after creating the change groups, the students were divided into four ranked percentile groups based on their deep approach scores at the beginning of the course. the 34 students who were selected for the present study were categorised into the highest-ranked percentile group at the beginning of the course. their deep approach score at the beginning of the course was between 4.00 and 5.00 which is remarkably higher than the mean score of the deep approach in any of the disciplines. thus the selected students were among the highest scoring students in the deep approach scale. while the students‟ deep approach scores at the beginning of the course varied between 4.00 and 5.00, the variation at the end of the course was between 1.5 and 5.00. nine students were categorised into the „deep approach remains high‟ group. of them, four students showed increase in the deep approach, while five students‟ scores remained exactly the same. seven students were categorised into the „slight decrease in deep approach‟ group. these students‟ deep approach scores at the beginning of the course varied between 4.00 and 5.00, and all of them scored 0.25 lower on the deep approach scale at the end of the course. the „sharp decrease in the deep approach‟ group consisted of 18 students, whose deep approach scores at the beginning of the course varied between 4.00 and 5.00, and at the end of the course between 1.5 and 4.50. the decrease in their deep approach was between -0.5 and -3.25. however, in only five cases was the decrease more than -1.00. 2.1.2 participants and contexts of the 34 students, 21 were female and 13 male. they ranged in age from 19 to 43 years, the mean age being 27. in finland, students‟ mean age is higher than in most european countries because students graduate from upper secondary school later and thus enter university at an older age than in most european countries. however, the mean age of 27 years is higher than the average among bachelor students. two of the students were minoring in the courses while the rest were major students. the 34 students were attending a compulsory bachelor-level course in their own discipline (except for the two students who were minor students), and each was interviewed after completing the course. the students participated in one of 10 courses representing five disciplines: three courses in bioscience (n=7), two in educational sciences (n=12), one in mathematics (n=2), three in theology (n=10) and one in veterinary medicine (n=3). the courses are presented in table 2. the courses were designed for second year students, except for courses on educational sciences which were designed for first year students. all ten courses were lecture courses that included both lecturing and activating assignments for the students, with the nature of the assignments varying. the courses lasted from 6 to 13 weeks and were worth between 3 and 10 credits. eight courses included a written exam at their conclusion, while one course included a learning diary and an oral exam, and one included a drama-type exam. an effort was made to select as similar courses as possible from the point of view of the students‟ role in order to minimise the effect of the course itself on the results. however, one course was based on group activities, and the exam was a group exam. in another course, the exam was a group exam in drama form. thus two courses differed from the others because they were based more on group activities, while the rest were based on individual tasks and exams. in all courses the students attended in lectures and completed some activating tasks. l. postareff et al. 36 | f l r table 2 the discipline and course of the participants discipline course students (n) course type and assessment bioscience course 1 3 lectures, written exam at the end course 2 3 lectures, written exam at the end course 3 1 lectures, written exam at the end educational sciences course 1 7 peer group working, short lectures, oral exam and learning diary course 2 5 lectures, written exam at the end, essay mathematics course 1 2 lectures, calculations, written exam at the end theology course 1 3 lectures (including lots of discussions) written exam at the end course 2 6 lectures, written exam at the end course 3 1 lectures (including much discussions), drama-type exam at the end veterinary medicine course 1 3 lectures, written exam at the end 2.2 materials the students were interviewed on a voluntary basis once the courses had concluded. they were told beforehand about the content and purpose of the interview, and at the beginning of the interview they were allowed to ask questions about the research or the interview. the interviewees were told about the confidentiality of the interviews and that they cannot be identified at any point. the interviews were held in autumn 2009, autumn 2010, or spring 2011. conducted by the first author and two research assistants, the interviews lasted from 35 to 75 minutes and were transcribed verbatim. the interviews focused on the students‟ descriptions of their intentions and goals related to studying and learning at the university, their learning processes and practices in general as well as during the course they had just completed, and their experiences of studying and learning in the specific course they had recently attended. the interviews were deep and open in nature, with each of them covering the above-mentioned theme. for our previous study (lindblom-ylänne et al., 2013), which formed the basis for selecting the students for the present one, the students filled in a revised version of the alsi (entwistle & mccune, 2004; parpala & lindblom-ylänne, 2012) which contains scales measuring students‟ approaches to learning. the deep approach scale consists of four items, which are presented in table 3. items 1 and 2 measure students‟ learning strategies while items 3 and 4 focus on their intentions. in the interviews, the focus was similarly on students‟ study strategies and intentions. l. postareff et al. 37 | f l r table 3 items on the deep approach scale item 1 ideas i‟ve come across in my academic reading set me off on long chains of thought. item 2 i look carefully at evidence to reach my own conclusion about what i‟m studying. item 3 i try to relate new material, as i am reading it, to what i already know on the topic. item 4 i try to relate what i learn in one course to what i have learned in other courses. 2.3 analyses qualitative content analysis was selected as the analysis method for the interview data. in the first phase, we used inductive content analysis, in which themes are allowed to emerge from the data without any theoretical assumptions (see elo & kyngäs, 2007; schilling, 2006). inductive content analysis was used to analyse how the students described their studying and learning (both generally in their university studies and in the specific course), as well as their study experiences during the specific course. three steps typical of inductive content analysis were carried out: data reduction, grouping and conceptualisation (see patton, 1990; flick, 2002). these three steps represented the variable-oriented technique (see vathournout et al., 2013) as the aim was to identify all factors related to students‟ learning and study experiences regardless of the individuals. the first step was data reduction, in which all descriptions related to these issues were identified from the interview transcripts. this was done by the first author independently. the second step was to group similar descriptions under same categories (e.g., all descriptions related to students‟ motivation were placed under the same category). this was done by the first and second author independently, and the identified categories were compared and discussed. after an in-depth discussion, the identified descriptions were placed under seven categories. the third step, conceptualisation, included finding a concept for each of the seven categories which describe the content and nature of each category. for example, the category including description related to emotions and attachment was conceptualised as „emotional commitment‟. this was done in collaboration with all three authors. in the second phase of the analysis a person-oriented approach was adopted (see vanthournout et al., 2013). each of the seven categories was investigated in more depth within the following three groups of students: „deep approach remains high‟, „slight decrease in deep approach‟, and „sharp decrease in deep approach‟ in order to identify similarities and differences within students in the same group and between students in different groups. for example, we explored how students showing sharp decrease in their deep approach described their motivation, emotional commitment, and the remaining five categories, and compared their descriptions to other students‟ descriptions. this phase was conducted by the first author, but the final results were obtained through a thorough discussion with all three authors. the third phase of the analysis was deductive content analysis, in which existing theories are utilised in analysing the data (see elo & kyngäs, 2007; schilling, 2006). in this phase we analysed, through adopting a person-oriented approach, how the three central elements of the disposition to understand emerge in each students‟ interview: 1) a well-developed use of learning strategies which concentrate on relating ideas, the critical use of evidence and attention to detail; 2) a willingness to devote the necessary time, effort, and concentration to apply the learning strategies effectively; and 3) an alertness to the learning context (see entwistle & mccune 2009; mccune & entwistle, 2011). this phase was conducted independently by the first two authors. the findings of both authors were compared and discussed together. the inter-rater agreement was high, although several discussions were needed to obtain the final results. to give an example, the „emotional commitment‟ category was discussed in depth to determine which elements would be included in it. l. postareff et al. 38 | f l r 3. results and discussion in each of the three groups, elements related to the high level of a deep approach at the beginning of the course were analysed. in addition, students‟ descriptions of studying and learning in the specific courses were analysed separately in each group. in each of the groups the elements related to stability or changes in one‟s deep approach could be categorised under seven different themes: 1) students‟ motives, intentions and study strategies, 2) organised studying and regulation of learning, 3) emotional commitment, 4) experiences of challenge, 5) interest in the course content, 6) devoting time and effort to studying during the course and 7) experiences of the course teaching. these were identified in each of the three groups. finally, elements of a „disposition to understand for oneself‟ were identified. in what follows, the seven themes identified in the student interviews are described and discussed within each of the following categories of change: „remains high‟, „slight decrease‟ and „sharp decrease‟. finally, the results concerning a disposition to understand are presented and discussed. 3.1. individual and contextual elements related to one’s deep approach to learning both individual and contextual elements related to the stability or changes in one‟s deep approach to learning were identified. however, clearly distinguishing between individual and contextual elements was challenging because, for example, the category „devoting time and effort to studying during the specific course‟ combined the individual‟s effort and the context. the range of elements is presented from individual to more contextual ones below. 3.1.1. students‟ motives, intentions and study strategies when describing their studying and learning in general, the nine students whose deep approach scores remained high described having a strong intrinsic motivation to study at the university. they said that it is not enough for them to just pass courses, but that they aim at a deep understanding of the subject matter and developing themselves as persons. marton, dall‟alba and beaty (1993) found a similar category, „changing as a person‟, and van rossum, deijkers and hamer (1985) labelled another similar category as „self realisation‟. the nine students‟ descriptions revealed that their conceptions of learning were sophisticated, including, for example, conceptions of learning being about relating ideas and combining new information with their previous knowledge. some of the students also stated that when they are able to explain the subject to someone else in their own words, they feel they have learned well. these students‟ descriptions revealed that they all had developed good and functional study skills. they all concentrated on the big picture instead of details, and explained that they had formed a larger picture of the learned material for themselves. the students‟ descriptions revealed that they go through deep thinking processes while studying. these students‟ descriptions of their studying were therefore in line with the inventory results in that they clearly reflected an adoption of a deep approach to learning: their intention was to learn deeply and to form a coherent whole from the subject matter, and they used strategies which enabled deep-level learning (see e.g. entwistle, 2009). the seven students showing a slight decrease in their deep approach to learning described their studying and learning at the university very similarly to the students showing no deep approach decrease. in addition, all seven students in this group seemed to have a strong intrinsic motivation towards their university studies. they emphasised that learning is about broadening one‟s understanding and observing things from new perspectives. integrating new information with previous knowledge was also emphasised. all six students‟ descriptions revealed that they had good study skills as did the students showing no deep approach decrease. the students stated, for example, that they analyse the subject matter from diverse perspectives and explain the central concepts or content in their own words. most mentioned that they search for extra material by themselves and concentrate on what they find challenging. in addition, most also mentioned using a variety of learning strategies (e.g. mind maps, notes, explaining things in their own words). none of the students described difficulties in their learning. two of the seven mentioned preparing for lectures beforehand through familiarising themselves with the content. thus the students showing a slight l. postareff et al. 39 | f l r deep approach decrease were aiming at a deep understanding, and used effective strategies to accomplish this, being in this sense very similar to the students showing no decrease. the 18 students showing a sharp deep approach decrease differed more from the two other student groups. firstly, six of them described more extrinsic motivators (such as earning a degree), although they also mentioned that they like studying at the university and that in most occasions they are also interested in course content. secondly, not all students in this group described applying a deep approach to learning as strongly as those in the two other groups. for example, one characterised her study process as mainly memorising things. although most students‟ descriptions reflected elements of deep learning, only three described themselves learning as deeply as students in the two other groups. moreover, three students mentioned uncertainty regarding their own way of learning as well as feelings of incompetence. two of these students described their learning in a very theoretical manner, rather than in their own words; thus their awareness of the elements of deep learning as students in educational sciences might have contributed to their high score on the deep approach scale at the beginning of the course. the third student stated that she would like to „learn how to learn‟. these three students seemed to have problems with their study skills. however, the other students‟ interviews did not reflect such problems. these results imply that the more varying motives, intentions and study strategies among some of the students in this group might be related to a decrease in their deep approach to learning. while students showing a slight or no decrease had a commitment to learn, the interviews of students showing a sharp decrease did not reflect such a clear commitment. there were no clear differences between students participating in different courses with regard to their intentions, motives and study strategies, except that two students in a an educational sciences course described their learning in a way which implied that they were aware of theories of learning when describing their own studying. however, this difference was related more to the students‟ discipline than to the course itself. 3.1.2. regulation of learning the descriptions of the nine students whose deep approach scores remained high revealed that they all had good self-regulation skills. self-regulation refers to processes in which students plan, monitor, control and regulate their own learning (vermunt, 1998). self-regulation resembles organised studying (entwistle & mccune, 2004), making the two concepts partly overlap. for example, time-management skills can be related to organised studying or self-regulated learning. all nine students described setting goals for their own learning, and studying regularly instead of only before deadlines or exams. they wanted to learn the course content deeply, and they read additional material or consulted their teachers or peers when they had difficulties in understanding the content. they all attended lectures regularly, although attendance was voluntary in all courses. thus they clearly assumed responsibility for their own learning. scheduling studies beforehand was also emphasised by some of the students. they all described having good time-management skills although one student mentioned sometimes having difficulties in getting started. some of these students emphasised that they concentrate on the most relevant content and study effectively in order to avoid an overload of work. the interview results support the results of previous studies showing that selfregulation is related to a deep approach to learning (e.g. lonka & lindblom-ylänne, 1996; heikkilä & lonka, 2006; heikkilä et. al., 2011; vermunt & van rijswijk, 1988). self-regulation skills are also related to students‟ study pace, study success and well-being (e.g. heikkilä et. al., 2011; rytkönen et. al., 2012). the seven students showing slight decrease in the deep approach also seemed to have good selfregulation skills. their descriptions implied that they were aware of what they were supposed to learn and were able to focus their attention on the relevant content. they all studied on a regular basis, but a very organised way of scheduling own studies was mentioned by only one student, which clearly differentiated her from the three students showing no decrease in the deep approach. three of these students mentioned that they do not attend lectures regularly, but instead devote their time to reading the course material. however, none of these students‟ interviews reflected clear problems with self-regulation skills. the 18 students whose deep approach decreased sharply clearly differed from the two other groups with regard to self-regulation skills. only five students‟ descriptions reflected good self-regulations skills. the remaining 13 students‟ descriptions did not reflect severe problems in self-regulation, but most of these l. postareff et al. 40 | f l r students had not, for example, set goals for their own learning, and four of them expected concrete guidance and support from the teacher. thus a lack of regulation or external regulation characterised these students more than self-regulated learning. externally regulated students rely on teachers, other students or study material for guidance, while lack of regulation relates to difficulties in self-regulation. students who lack self-regulation skills are unsure about how they should study or may find it difficult to assess whether they have sufficiently learned the subject matter (vermunt, 1998; vermunt & verloop, 1999). furthermore, the 13 students in this group did not describe organising their studies systematically, and their time management skills were not as good as those whose deep approach did not decrease. two students had more severe problems concerning time-management and organising their studies effectively. for example, one described trying to find a rhythm and routine in her studies and avoid doing things right before the deadlines. these problems in self-regulation skills and time management are likely to explain the decrease in these students‟ deep approach, and may also reflect a combination of unorganised studying and applying a deep approach in an overly sense (see parpala et al., 2010). there were no differences between students participating in the different courses with regard to their self-regulation skills. 3.1.3. emotional commitment all nine students whose deep approach remained high described being committed to their studies. however, six of those students described stronger emotions related to studying their major subject, i.e. having a strong attachment to and respect for the subject they were studying. they also mentioned having a desire to learn more about the discipline and enjoying the university experience as well as pride in studying at the university. a recent quantitative study similarly showed a relationship between university students‟ positive emotions, such as pride, and a deep approach to learning (trigwell, ellis & han, 2011). of the seven students showing a slight deep approach decrease, only one mentioned strong emotions and an attachment to studying her discipline, or a desire to do so. enjoyment of learning was mentioned by four students in this group, but none of the seven described their studying in a negative light. the results show that these students did not have such a strong emotional commitment towards studying as did their peers whose deep approach did not decrease. all 18 students showing a sharp deep approach decrease described enjoying their studies in general, but a majority of these students mentioned it being very context-dependent. in some courses they are very committed, but in others they simply want to do the minimum in order to pass the course. only one of the 18 students described having a strong attachment or desire with respect to her studies. she stated that it is important for her to be part of the scientific community. to conclude, enjoyment of learning was mentioned by all 34 students, but clear differences were noted between the students‟ emotional commitment with regard to the changes in their deep approach to learning. again, there were no differences between students participating in the different courses with regard to their emotional commitment. 3.1.4. devoting time and effort to studying during the specific course eight of the nine students whose deep approach remained high during the courses devoted considerable time to studying the course content, and they described studying actively in all the courses they take. most of them said that they did not have to invest much time in preparing for the exam because they had studied thoroughly throughout the entire course. some mentioned that the course content was challenging, or that the lecturer proceeded too quickly, which compelled them to devote even more time and effort than normally. only one student mentioned that she did not invest time to study the course content until a few days before the exam. to conclude, active and regular participation in courses was common to all students, except for one, whose deep approach remained on a high level. of the seven students whose deep approach decreased slightly, five described investing time and effort in studying during the course. they actively participated in the lectures and spent a significant amount of time reading the study material. however, two of the students described being less active during the course. one of them stated that he did not have enough time to study on a regular basis and the other student l. postareff et al. 41 | f l r stated that she did not have to invest that much time or effort in her studying because the course did not provide much new information. only two of the 18 students showing a sharp decrease mentioned studying regularly and actively during the course. of the remaining 16 students, nine clearly stated that they invested less time and effort in the course than they normally would. seven of them felt that the course they participated in was not very demanding and they did not have to invest much time or effort in studying, and two students described that their own activity decreased because the course was based on group activities. the other seven students reported that they more generally tend to study actively only when exams approach. one stated that her weak study skills prevented her from studying more effectively during the course. one student worked during the evenings which left her little time for studying. the results are clear in that investing time and effort in studying during the course was related to the stability of one‟s deep approach to learning. the less students described being active during the course, the more their deep approach to learning declined. however, the reasons for investing less time varied, being related to particular study habits, weak study skills, working along with studying, or to the challenges the course presented. most of the students who felt that the course was not very demanding and therefore invested little time and effort, were students from the same theology course. otherwise no differences were noted between students participating in the different courses as to how much time or effort they devoted to studying. 3.1.5. experiences of challenge the nine students whose deep approach remained high during the course reported that the courses challenged them positively in one way or another. for four students the course content was challenging, which made them invest more time and effort in studying. five students did not describe their course as very challenging, but they wanted to thoroughly learn the content, and challenged themselves by reading extra material or analysing the content from different points of view. thus a challenging learning environment or a student‟s own desire to thoroughly learn the content, seemed to maintain these three students‟ deep approach at a high level. the students showing a slight deep approach decrease experienced the demands of the courses in different ways. three of the seven students said that the course did not offer them enough of a challenge. for instance, one described the course being more about pondering the course material from different perspectives, because the content was already so familiar. she would have hoped to learn more new information during the course. on the other hand, four students stated that the courses were in some ways too challenging. for example, one mentioned that the course books were too difficult and that she had to read them many times until gaining some kind of understanding. another student stated that the course covered too much information, and that he did not always know what he was supposed to study although he felt that he was able to follow the course well. thus the students showing a slight deep approach decrease experienced both too many and too few challenges, but only to a slight degree. despite facing challenges or too few of them, most of them actively participated in the lectures and put effort into studying during the course. of the 18 students showing a sharp deep approach decrease, the descriptions of eight indicated that the courses were not challenging enough. some considered that the content was easy, and some thought that the course provided little new information. conversely, four students mentioned that the courses were too challenging. two of them had difficulties understanding the course content, while two mentioned that they had difficulties in forming their own view of the content because they acted in groups, which was quite challenging for them. the remaining six students mentioned neither too many nor too few challenges. students from one theology course stated more often than students from the other courses that the course was not challenging enough, but in the other courses students varied more with respect to how they experienced the challenges the course presented. l. postareff et al. 42 | f l r these results suggest that the course „fit‟ is important for maintaining one‟s deep approach at a high level. either too many or too few challenges seem to lower the level of a deep approach among most students. kyndt, dochy, struyven and cascallar (2011) similarly showed that task complexity might hinder the application of a deep approach to learning. similarly, mccune and entwistle (2011) have emphasised that students need to experience challenging teaching–learning environments that systematically encourage students to focus on personal understanding. however, the results of the present study imply that some students seemed to be able to maintain their deep approach at a relatively high level although the course did not offer many challenges. this accord with the results of lindblom-ylänne and lonka (1999), who identified a cluster of students who were meaning-oriented and who independently found their own way, being immune to the effects of the teaching-learning environment. 3.1.6. interest in course content four of the nine students whose deep approach remained high mentioned that the course content was not very interesting. however, two stated that they became more interested in the content during the course. the other mentioned that the way the course was taught promoted her interest, and another expressed that discovering links between the course content and his previous experiences from working life increased his interest. more generally, this student stated that doing things properly is important to him, whether or not he is interested, and that he wanted to succeed and be proud of himself. the remaining five students found the courses more interesting than the two other students, although three of them did not express a strong interest in the course. despite the level of interest not being high among all of them, all these students showed a commitment to learn and wanted to succeed in their studies. kyndt et al. (2011) showed that the more student is motivated to study for autonomous reasons such as find a course pleasant, the more they will be inclined to use a deep approach to learning. so although not all students whose deep approach remained at a high level throughout the course described being interested in the course content, they seemed to have a more general motivation to learn for autonomous reasons. previous research similarly suggest that even though university students may find some content initially uninteresting and their studying may be based on extrinsic motivation, some are through self-regulation processes able to generate their own thoughts, feelings and actions to meet uninteresting study demands (ryan & deci, 2000; hidi & ainley, 2008). also students‟ whose deep approach to learning slightly declined varied with regard to how interesting they found the course content. six of the seven students described being interested in the content, with only one of them expressing a strong interest. one student said that the course did not interest her much and that her goal was simply to complete it. she also expressed being more interested in the course content at the beginning of the course, but that her level of interest decreased as the course progressed. ten of the 18 students showing a sharp deep approach decrease stated that they found the content of the course interesting. however, most of them were not interested beforehand in the content, but became so during the course. on the other hand, eight students found the course content less interesting, but only two of these students mentioned that their goal was only to pass the course because the content was not interesting or useful to them. students‟ interest in the course content did not explain the changes in their deep approach, since both the students whose deep approach decreased, and those whose deep approach did not, described different levels of interest in the course content. however, the students whose deep approach remained high described more often than the others that they create links between different courses, which implies that they try to find the meaning of the courses even if the content does not particularly interest them. furthermore, these students would invest time and effort in studying even though a course was not of great interest to them, because they would want to understand the content deeply. interestingly, there were no differences between students participating in the different courses with regard to their interest in the course content. 3.1.7. experiences of the course teaching the nine students whose deep approach remained high described the teaching of the course in a positive manner. however, only two of them expressed that they were extremely satisfied with the teaching l. postareff et al. 43 | f l r and thought that the teaching enhanced their learning. most students stated that the teaching was fine and that the teacher was pleasant, but they did not mention that the teaching would have considerably enhanced their learning. one student stated that no matter what the teaching is like, he always studies the same way. the students whose deep approach slightly decreased described their experiences of the teaching in different ways. two of the seven students described the course teaching in a slightly negative manner. one of them mentioned that the lectures were frustrating, but that she valued the discussions with other students outside the lectures. the other student expressed that the lecturer was agreeable but that the lectures concentrated too much on discussions, with little new information being presented. three students reported more positive experiences. for example, they mentioned that the teacher had structured the lectures well, was genuinely interested in his students‟ learning and that there was supportive interaction during the lectures. however, none of the three said that the teaching significantly supported their learning. two student‟s experiences of the teaching were rather neutral. they said that the lectures were traditional and not so useful, although they mentioned that the teacher of the course was good. ten of the 18 students showing a sharp deep approach decrease were rather satisfied with the teaching of the course. however, only three of these said that the teaching was exceptionally good and the rest expressed milder positive responses. three students described that the teacher of the course was very pleasant, but still they considered that the teaching did not significantly enhance their learning. five students had more negative experiences of the teaching: two stated that they did not understand what the teacher had said and that the teaching was boring, and three were not completely satisfied with the teaching method because it was new to them and they did not find themselves comfortable with it. these three students were from a course in educational sciences, where the students studied in groups throughout the whole course. most of their peers enjoyed the group method, but these three students had some difficulties with it. otherwise, no clear differences could be detected between students participating in different courses in their experiences of the course teaching. interestingly, these results imply, that the experiences of the quality of teaching were not related to deep approach stability or decrease. most students whose deep approach decreased sharply were satisfied with the teaching and some whose deep approach remained high throughout the course described some negative experiences related to the teaching. therefore the results support the existence of students who are „immune‟ to the teaching-learning environment, as suggested in previous studies by lindblom-ylänne & lonka (1999), who showed that the study practices of some meaning-oriented students remain unaffected by the learning environment. the results of baeten et al. (2010), however, showed that students who are satisfied with the quality of a course are more likely to employ a deep approach than students who are less satisfied. a larger sample of interviewees would be needed to explore in more depth the relationship between students‟ experiences of teaching and their approaches to learning. students with lower deep approach scores might be more sensitive to the quality of teaching and the effects of the teaching-learning environment. this could not be confirmed in the present study since all students scored highly at the beginning of the course and thus the sample was highly selected. another limitation is that the students were categorised into the three change groups according to their scores within only one course. however, during the first measurement the students were asked to consider how they have studied so far, and the interviews were used to improve the reliability of the questionnaire data. thus the use of a mixed-methods approach enabled a deep and more reliable investigation of the elements affecting students‟ studying and learning and their disposition to understand. a further limitation concerns the use of different cohorts of students: the data was collected during three different semesters and from both first and second year students, which might affect the results since students representing different cohorts might have diverse experiences of their learning environment. we are able to address some of these limitations in our other studies, since the data collected for our large research project includes a rich variety of qualitative data, e.g. stimulated recall data on assessment of student learning and students‟ exam papers, observation and video data from the courses as well as large interview data from both students and teachers. l. postareff et al. 44 | f l r 3.2. elements of a ‘disposition to understand for oneself’ the disposition to understand for oneself, as defined by entwistle and mccune (entwistle & mccune 2009; mccune & entwistle, 2011), could not be analysed in detail from the interviews because these focused broadly on studying and learning at the university and on the specific course rather than explicitly on the disposition to understand. however, the strength of our interviews was that they were openended and deep in nature and the students were able to thoroughly describe the elements of their studying and learning they considered to be important. some elements of having a disposition to understand clearly emerged from the data although the students were not specifically asked about them. a central element of the disposition to understand is a well-developed use of learning strategies which concentrate on relating ideas, critically using of evidence and attention to detail. these types of learning strategies were identified among most of the interviewed students, although some students who showed sharp decrease in the deep approach described a more narrow use of learning strategies. one component of well-developed learning strategies is a broader focus on the discipline as a whole instead of individual courses or blocks of content. this type of broader focus could be clearly identified among five of the nine students whose deep approach remained high and among three of the seven students whose deep approach slightly decreased. most of the students whose deep approach sharply decreased described a welldeveloped use of learning strategies, but aiming at gaining a broader view of the discipline as a whole was not emphasised. one student, whose deep approach remained high, described this type of broader focus on the discipline as follows: ” … i enjoy studying and learning hugely. i am very attached to my own major subject, but i would also like to explore what else i can learn here outside my major subject because i want to learn things deeply and broadly. i want to gain thinking and writing skills as well and absorb information from all possible sources.” (male student, educational sciences) secondly, willingness to devote the necessary time, effort and concentration to apply the learning strategies effectively is an essential element of a disposition to understand. again, the students whose deep approach remained high and most of those showing a slight decrease put a considerable amount of time and effort into studying during the course. however, only a few of the students whose deep approach sharply decreased devoted a good deal of time and effort to studying the content. in the following quotation a student whose deep approach did not decrease describes how he devoted time and effort during the course: “i attended lectures regularly and took notes. then, at home i looked at the materials we studied in many different books and i compared it… i tried to combine the new information with the old. i did this during the whole course, which was good, because i was able to keep on track all the time. i didn‟t read only for the examination, but evenly throughout the course.” (male student, biosciences) furthermore, some of the students showed an „alertness to the learning context‟ which is the third important element in a disposition to understand. it is defined as “alertness that monitors the learning processes and strategies in relation to the demands of the task, along with alertness to opportunities provided by the teaching, and indeed the whole learning environment, to further one's understanding”. elements of this type of alertness could be found in five students‟ interviews whose deep approach did not decrease, in two students‟ interviews whose deep approach slightly decreased and in one student‟s interview whose deep approach sharply decreased. however, the data was somewhat thin with regard to analysing alertness to the learning context, which prevented us from further investigating this element. the following quotation offers an example of how alertness to the learning context emerged in our data: “this bioscience course was very basic, but one had to learn the content deeply. therefore i did a lot of work and studied the content very well… once you do that, it helps you later on to understand things. in some courses i have to invest less effort and i don‟t necessarily understand everything so deeply, but the material in this course needed to be studied well… i l. postareff et al. 45 | f l r also study mathematics, physics, chemistry and biochemistry, and it‟s awesome to notice that some physics and chemistry matters can be combined with biology, and that mathematics is needed in all of them. i want to understand all of these fields.” (male student, biosciences) an important finding of our study was that the level of interest towards the content of the courses was not always related to deep approach changes. a low level of interest also characterised students whose deep approach remained high or only decreased slightly. these results suggest that such students have a will to learn even though they might find the course content less interesting. entwistle and mccune (entwistle & mccune 2009; mccune & entwistle, 2011) suggest that students with a disposition to understand for oneself have a continuing desire to adopt effortful, deep approaches across a wide range of contexts, and to reach the most satisfying understanding possible. the continuing desire to learn is illustrated in the following quotation: ”i have this constant hunger for information, i always find new things which i want to explore in more depth and i feel i need to find out more about this and that, and sometimes i end up on a number of different paths.” (female student, educational sciences) these students also expressed wanting to understand things deeply for their own purposes. entwistle and mccune (entwistle & mccune 2009; mccune & entwistle, 2011) note that students showing a disposition to understand feel strongly that they need to understand for themselves, and that they want to demonstrate the depth of their understanding, for example in examination answers. our results revealed that some of the students showing a slight or no change in their deep approach described that they study as long as it takes to understand the content deeply, and that they explained the content to themselves in their own words, as one student put it: “i want to understand as deeply as possible, my head can‟t take pure memorisation. i do mind maps and then i explain the things in my own words to the walls.” an attempt to understanding for oneself becomes evident in the following quotation from another student: “i read the course books in a very self-oriented way and i concentrate on the things that are important to me.” mccune and entwistle suggest that a students‟ strong commitment to understand may indicate that it has become part of that student's sense of identity as a learner, and so represents a much more stable characteristic than a deep approach. the strong feelings students express suggest that it has become a part of the students‟ sense of identity as learners (entwistle & mccune, 2009). the interview quotations presented in this article contain a substantial number of words implying strong feelings (such as „hugely‟, „hunger for information‟) which indicate that these students have a strong commitment and disposition to understand. in our data, four of the students showing no deep approach decrease, and one student showing a slight decrease, described this type of strong commitment. as well, the following quotation demonstrates strong feelings („really‟, „overly‟) when a student describes his studying, which also indicates a solid commitment to understand: “in every course i take i really want to learn, and not just pass the course. i feel that i really want to understand and use the information. sometimes there are courses which at first don‟t seem to be very important, but then i find myself being overly enthusiastic once i get involved with the content….“ (male student, theology) a challenge related to developing a disposition to understand is that it is, according to mccune and entwistle (entwistle & mccune 2009; mccune & entwistle, 2011), a more stable characteristic and less changeable through specific experiences. this was supported by the results of the present study showing that particularly the students whose deep approach did not decrease during the course were unaffected by the teaching-learning environment. mccune and entwistle emphasise that students need to experience challenging teaching–learning environments that systematically encourage a focus on personal l. postareff et al. 46 | f l r understanding. our results support this view by showing that having too few or too many challenges was mostly related to a decrease in one‟s deep approach, while positive challenges were related to the stability, or even an increase of one‟s deep approach. a student whose deep approach decreased sharply describes her perceived lack of challenges in the following way: “i could have invested more in studying. but somehow i felt that i would remember these things well enough … i felt that i wouldn‟t have to write them down in my learning diary immediately after the lectures. i only started the learning diary two days before the deadline.” (female student, theology) thus the central elements of a disposition to understand for oneself clearly emerged from the interviews, but our results imply that only about half of the students whose deep approach remained high, and a few whose deep approach slightly decreased, showed a disposition to understand for oneself. however, a more thorough analysis would require different types of interview questions, which more thoroughly would focus on the disposition to understand. nevertheless, the deep and open nature of the interviews made it possible to analyse elements of the disposition to understand in the current study as well. in the interviews the students broadly described their studying and learning and these descriptions revealed elements related to the disposition to understand. the mean age of the three students whose deep approach did not decrease was 30 years. this implies that a stronger commitment to understand might be related to student age as well. previous studies have also shown that older students are more likely to adopt a deep approach to learning than their younger peers (e.g. gow & kember, 1990). 4. conclusions in the present study we were able to identify the individual and contextual elements which were related to the stability of or changes in one‟s deep approach to learning. we identified that the students whose deep approach to learning decreased sharply during the course described problems in their studying and it seemed that at least some of them had exaggerated their deep approach level at the beginning of the course when answering the questionnaire. thus they did not show as strong commitment to understand as the students whose deep approach did not decrease or decreased only slightly. the students whose deep approach remained high or decreased only slightly described their studying and learning very similarly, and both individual and contextual elements were identified that logically explained the questionnaire results. the students whose deep approach decreased sharply clearly differed from those showing only slight or no changes in their deep approach with respect to the individual elements. for example, they described more problems in their self-regulation skills, time-management skills and study strategies. however, these students did not differ from the others in terms of their experiences of the teaching or their interest in the course. it therefore seems that the individual elements explained sharp decreases more than the contextual elements did. however, some interviews clearly showed that a lack of challenges or too many challenges decreased the deep approach level. in general, students showing a lack of interest in course content might be inclined to exhibit a disposition to understand for oneself when courses are challenging them in a positive way. adjusting the level of the course appropriately, then, seems to be a key element in course design. these elements should be further examined among students scoring lower on the deep approach scale. the results of the study confirm the strength of using in-depth qualitative analysis when examining students‟ approaches to learning and why they may vary or change. in addition, the study provided new information of the relationship between the deep approach to learning and a disposition to understand for oneself. our future research will focus on analysing the interviews of students scoring lower on the deep approach to learning scale. l. postareff et al. 47 | f l r keypoints elements explaining stability or change in the deep approach to learning were explored. individual elements explained the stability or change more than the contextual elements did. not all students showed a strong commitment to understand despite their high score on the deep approach scale. elements of a „disposition to understand for oneself‟ were identified among some students. references baeten, m., kyndt, e., struyven, k., & dochy, f. (2010). using student-centred learning environments to stimulate deep approaches to learning: factors encouraging or discouraging their effectiveness. educational research review, 5, 243-260. doi:10.1016/j.edurev.2010.06.001 biggs, j. (1987). student approaches to learning and studying. camberwell, vic: australian council for educational research. creswell, j. (2009). research design: qualitative, quantitative and mixed methods approaches (3rd edition). london: sage publications. diseth, a. (2003). personality and approaches to learning as predictors of academic achievement. european journal of personality, 17, 143–155. doi: 10.1002/per.469 elo, s. & kyngäs, h. (2007). the qualitative content analysis process. journal of advanced nursing, 62 (1), 107-115. doi: 10.1111/j.1365-2648.2007.04569.x entwistle, n. (2009). teaching for understanding at university: deep approaches to learning and distinctive ways of thinking. basingstoke, hampshire: palgrave macmillan. entwistle, n. j., & mccune, v. (in press). the disposition to understand for oneself at university: integrating learning processes with motivation and cognition. british journal of educational psychology. entwistle, n. j., & mccune, v. (2009). the disposition to understand for oneself at university and beyond: learning processes, the will to learn and sensitivity to context. in l-f. zang & r. j. sternberg (eds.), perspectives on the nature of intellectual styles (pp. 29-62). new york: springer. entwistle, n. & mccune, v. (2004). the conceptual bases of study strategies inventories in higher education. educational psychology review, 16 (4), 325-345.doi: 10.1007/s10648-004-0003-0 entwistle, n., & ramsden, p. (1983). understanding student learning. london: croom helm. flick, u. (2002). an introduction to qualitative research. 2nd ed. london: sage publications. gijbels, d., segers, m., & struyf, e. (2008). constructivist learning environments and the (im)possibility to change students‟ perceptions of assessment demands and approaches to learning. instructional science, 36, 431–443. doi: 10.1007/s11251-008-9064-7 gow, l. & kember, d. (1990). does higher education promote independent learning? higher education 19, 307-322. doi: 10.1007/bf00133895 hailikari, t., postareff, l,. tuononen, t., räisänen, m. & lindblom-ylänne, s. (in press). students‟ and teachers‟ perceptions of fairness in assessment. in c. kreber, c. anderson, n. entwistle, & j. mcarthur (eds), advances and innovations in university assessment and feedback. the edinburgh university press. heikkilä, a., & lonka, k. (2006). studying in higher education: students‟ approaches to learning, selfregulation, and cognitive strategies. studies in higher education, 31, 99-117. doi: 10.1080/03075070500392433 heikkilä, a., niemivirta, m., nieminen, j., & lonka, k. (2011). interrelations among university students‟ approaches to learning, regulations of learning, and cognitive and attributional strategies: a person oriented approach. higher education, 61, 513-529. doi: 10.1007/s10734-010-9346-2 hidi, s. & ainley, m. (2008). interest and self-regulation: relationships between variables that influence learning. in d. schunk & b. j. zimmerman (eds.), motivation and self-regulated learning (pp. 77110). theory, research, and applications. new york: taylor & francis. l. postareff et al. 48 | f l r johnson, r.b, onwuegbuzie, a.j., & turner, l.a. (2007). toward a definition of mixed methods research, journal of mixed methods research, 1(2), 112-133. doi: 10.1177/1558689806298224 kyndt, e., dochy, f., cascallar, e., & struyven, k. (2011). the direct and indirect effect of motivation for learning on students‟ approaches to learning, through perceptions of workload and task complexity. higher education research & development, 30, 135-150. doi: 10.1080/07294360.2010.501329 kyndt, e., dochy, f., struyven, k., & cascallar, e. (2011). the perception of workload and task complexity and its influence on students‟ approaches to learning. european journal of psychology of education, 26, 393-415. doi: 10.1007/s10212-010-0053-2 lietz, p., & matthews, b. (2010). the effects of college students‟ personal values on changes in learning approaches. research in higher education, 51, 65–87. doi: 10.1007/s11162-009-9147-6 lindblom-ylänne, s., & lonka, k (1999). individual ways of interacting with the learning environment are they related to study success? learning and instruction, 9, 1-18. doi: 10.1016/s09594752(98)00025-5 lindblom-ylänne, s., parpala, a., & postareff, l. (2013). challenges in analysing change in students‟ approaches to learning. in v. donche, j. richardson, j. vermunt, & d. gijbels (eds), learning patterns in higher education. routledge. lonka, k., & lindblom-ylänne, s. (1996). epistemologies, conceptions of learning, and study practices in medicine and psychology. higher education, 31, 5-24. doi: 10.1007/bf00129105 lonka, k., olkinuora, e., & mäkinen, j. (2004). aspects and prospects of measuring studying and learning in higher education. educational psychology review, 16 (4), 301-323. doi: 10.1007/s10648-0040002-1 marton, f. & säljö, r. (1976). on qualitative differences in learning: i. outcome and process. british journal of educational psychology, 46, 4-11. doi: 10.1111/j.2044-8279.1976.tb02980.x marton, f. & säljö, r. (1997). approaches to learning. in f. marton, d. hounsell & n. entwistle, (eds.) the experience of learning (2 nd ed., pp. 39-58). edinburgh, uk: scottish academic press. marton, f., dall‟alba, g., & beaty, e. (1993). conceptions of learning. international journal of educational research, 19, 277-300. mccune, v., & entwistle, n. (2011). cultivating the disposition to understand in 21 st century university education. learning and individual differences, 21, 303-310. doi: 10.1016/j.lindif.2010.11.017 nieminen, j., lindblom-ylänne, s., & lonka, k. (2004). the development of study orientations and study success in students of pharmacy. instructional science, 32, 387–417. doi: 10.1023/b:truc.0000044642.35553.e5 parpala, a., & lindblom-ylänne, s. (2012). using a research instrument for developing quality at the university. quality in higher education, 18, 313-328. doi: 10.1080/13538322.2012.733493 parpala, a., lindblom-ylänne, s., komulainen, e., litmanen, t., & hirsto, l. (2010). students‟ approaches to learning and their experiences of the teaching-learning environment in different disciplines. british journal of educational psychology, 80, 269-282. doi: 10.1348/000709909x476946 patton, m.q. (1990). qualitative evaluation and research methods (2nd ed.). london: sage publications. perkins, d. n., & tishman, s. (2001). dispositional aspects of intelligence. in j. m. collis & s. messick (eds.), intelligence and personality (pp. 233−258). mahwah, nj: lawrence erlbaum. ryan, r. m. & deci, e. l. (2000). intrinsic and extrinsic motivations: classic definitions and new directions. contemporary educational psychology, 25, 54-67. doi: 10.1006/ceps.1999.1020 rytkönen, h., parpala, a., lindblom-ylänne, s., virtanen, v., & postareff, l. (2012). factors affecting bioscience students‟ academic achievement. instructional science, 40, 241-256. doi: 10.1007/s11251-011-9176-3 schilling, j. (2006). on the pragmatics of qualitative assessment: designing the process for content analysis. european journal of psychological assessment, 22 (1), 28-37. doi: 10.1027/1015-5759.22.1.28 struyven, k., dochy, f., janssens, s., & gielen, s. (2006). on the dynamics of students‟ approaches to learning: the effects of the teaching/learning environment. learning and instruction, 16, 279–294. doi: 10.1016/j.learninstruc.2006.07.001 trigwell, k., ellis, r. a., & han, f. (2012). relations between students‟ approaches to learning, experienced emotions and outcomes of learning. studies in higher education, 37, 811-824. doi: 10.1080/03075079.2010.549220 l. postareff et al. 49 | f l r trigwell, k., prosser, m., & waterhouse, f. (1999). relations between teachers‟ approaches to teaching and students‟ approaches to learning. higher education, 37, 57-70. doi: 10.1023/a:1003548313194 van rossum, e. j., deijkers, r., & hamer, r. (1985). students‟ learning conceptions and their interpretation of significant educational concepts. higher education, 14, 617–641. doi: 10.1007/bf00136501 vanthournout, g., donche, v., gijbels, d. & van petegem, p. (2013). (dis)similarities in research on learning approaches and learning patterns. in d. gijbels, v. donche, j.t.e. richardson & j.d. vermunt (eds.), learning patterns in higher education. dimensions and research perspectives (pp. 11-32). routledge. watters, d., & watters, j. (2007). approaches to learning by students in the biological sciences: implications for teaching. international journal of science education, 29, 19–43. doi: 10.1080/09500690600621282 vermunt, j. d. (1998). the regulation of constructive learning processes. british journal of educational psychology, 68, 149-171.doi: 10.1111/j.2044-8279.1998.tb01281.x vermunt, j. d.., & van rijswijk, f.a.w.m. (1988). analysis and development of students‟ skill in selfregulated learning. higher education, 170, 647-682. vermunt, j. d., & verloop, n. (1999). congruence and friction between learning and teaching. learning and instruction, 9, 257-280. doi: 10.1016/s0959-4752(98)00028-0 watkins, d. a., & hattie, j. (1985). a longitudinal study of the approach to learning of australian tertiary students. human learning, 4, 127—142. zeegers, p. (2001). approaches to learning in science: a longitudinal study. british journal of educational psychology, 66, 59-71. doi: 10.1348/000709901158424 microsoft word reed et al_published.docx ! ! ! ! ! frontline learning research vol. 3 no. 1 (2015) 36 54 issn 2295-3159 ! ! ! self-beliefs mediate math performance between primary and lower secondary school: a large-scale longitudinal cohort study helen c. reeda, paul a. kirschnerb & jelle jollesa a vu university amsterdam, the netherlands b welten institute, open university of the netherlands article received 10 december / revised 19 february / accepted 5 april / available online 21 april abstract it is often argued that enhancement of self-beliefs should be one of the key goals of education. however, very little is known about the relation between self-beliefs and performance when students move from primary to secondary school in highly differentiated educational systems with early tracking. this large-scale longitudinal cohort study examines the extent to which academic self-efficacy (i.e., how confident students are that they will be able to master their schoolwork) and math self-concept (i.e., students’ perceived math competence) mediate the relation between math performance at the end of primary school (grade 6) and the end of lower secondary school (grade 9) in such a system. the study involved 843 typically-developing students in the netherlands. self-efficacy and math self-concept were measured with self-report questionnaires. math performance was measured with nationally validated tests. the relation between math performance in grade 6 and in grade 9 was uniquely mediated by both self-efficacy in grade 6 and math self-concept in grade 9, but in opposing directions. math self-concept was the most influential mediator, explaining nearly a quarter of the total effect of grade 6 math performance on grade 9 math performance. unexpectedly, high selfefficacy in grade 6 was negatively related to grade 9 math performance, particularly for girls and high-track students. these findings suggest that self-efficacy may not necessarily be a protective factor in highly differentiated early tracking educational systems and may need to be actively managed when students move to secondary school. keywords: self-beliefs; self-efficacy; math self-concept; math performance; school transition; educational tracking corresponding author: helen c. reed, department of educational neuroscience and learn! research institute, faculty of psychology and education, vu university amsterdam, van der boechorststraat 1, 1081 bt amsterdam, the netherlands; e-mail: hc.reed@vu.nl doi http://dx.doi.org/10.14786/flr.v3i1.139 ! reed!et !al ! ! | f l r ! ! 37! 1. introduction most students hold beliefs about their own capabilities and competence in accomplishing academic tasks. do these so-called self-beliefs affect the relation between students’ performance at the end of primary school and the end of lower secondary school? this question is especially relevant in systems that make use of early educational tracking to stratify students according to scholastic ability. tracking is based on the premise that homogeneous classes allow curriculum and instruction to be directed towards the common needs of groups of similar ability and that this leads to maximum learning for all (chmielewski, dumont, & trautwein, 2013; hanushek & wößmann, 2006). tracking is considered highly differentiated when students are stratified into different schools or educational programs with little or no contact between them. in educational systems with early tracking (e.g., the netherlands (see box 1), germany, belgium (flanders), singapore), track placement in lower secondary school depends to a large extent on performance at the end of primary school or even earlier. an important assumption is that there is a substantial degree of stability in performance between the end of primary school and lower secondary school. if this were not the case, then discrepancies between track placement and students’ actual performance would soon render these systems ineffectual. the stability of this relation could, however, be affected by student variables that depress or elevate performance in secondary school relative to expectations at the moment of track assignment. of particular concern are students whose performance in secondary school falls below expectation. students who fail in their designated track are often retained or drop down to a lower track, which is reported to be detrimental to student outcomes (brophy, 2006; jacob & lefgren, 2009; oecd, 2012). it may be possible to prevent this happening when more is known about student variables that affect the stability of the relation between performance in primary and secondary school. within this context, the present study investigates the extent to which the relation between math performance at the end of primary school (i.e., grade 6) and the end of lower secondary school (i.e., grade 9) in a highly differentiated early tracking educational system is mediated by student self-beliefs relating to their academic functioning at school. box 1: educational tracking in the netherlands in the netherlands, educational tracking is implemented early in secondary school. track placement depends largely on performance at the end of primary school and is based on school grades and/or the results of a school placement test, as well as study skills, concentration, motivation, application, etcetera. once track placement has been determined sometimes after an initial orienting period there is little academic contact (e.g., shared classes) between tracks. there are three main tracks: pre-university (preparing the most able students for university; 6 years duration; around 20% of students), higher general secondary (preparation for professional higher education; 5 years duration; 20%), and prevocational (theory-oriented or practice-oriented preparation for vocational education; 4 years duration; 55%). in addition, around 5% of students are in special needs education or receive training in low-level practical skills for entry to the workforce. 1.1 self-beliefs in school settings a large body of research indicates that positive self-beliefs are strongly related to higher academic performance, as we review presently. it is therefore worrying that many students experience a decline in self-beliefs between primary and secondary school (jacobs, lanza, osgood, eccles, & wigfield, 2002; liu, wang, & parkins, 2005), especially in the domain of mathematics. for example, the most recent cycle of the trends in international mathematics and science study reported that only just over a tenth of 8th graders are confident in their mathematics ability compared to a third of 4th graders (mullis, martin, foy, & arora, 2012). ! reed!et !al ! ! | f l r ! ! 38! the causes of this decline appear to be manifold. when students move from primary to secondary school, they are confronted with many factors (e.g., different learning and assessment goals, demands and conditions; relationships with peers and teachers; biological and neurological changes of adolescence) that can affect their beliefs about their ability to do well in the new school environment (cauley & jovanovich, 2006; fenzel, 2000; sakiz, pape, & hoy, 2012; schunk & meece, 2006; urdan & schoenfelder, 2006). the early years of secondary school occur at a crucial developmental period in early to mid-adolescence. on the one hand, adolescents have greater need for autonomy, feelings of competence, social connectedness, and positive relations with peers and adults; while at the same time they have heightened sensitivity to social comparisons, peer influence and emotional support or the lack thereof (cauley & jovanovich, 2006; osterman, 2000; sakiz et al., 2012; schunk & meece, 2006). on the other hand, secondary schools are more anonymous and more regimented than primary schools, there is stronger emphasis on testing and grades, teachers are perceived as more controlling and distant and less supportive and fair, and schoolwork is more plentiful and more demanding (cauley & jovanovich, 2006; osterman, 2000; sakiz et al., 2012). from a developmental perspective, demands made on adolescent learners could diverge from their neurocognitive capacities to meet them. for instance, many academic areas (e.g., science and mathematics, language learning and problem-solving) require higher-order thinking skills that depend on neural networks which show considerable individual variability in maturation during adolescence (e.g., crone et al., 2009; dumontheil, houlton, christoff, & blakemore, 2010). if students are required to think in ways that exceed their developmental capabilities, frustration, disillusionment, and decreased feelings of competence can result (cauley & jovanovich, 2006). furthermore, secondary school students are often expected to regulate their own learning at a time when their behavioural control is compromised by a heightened sensitivity to motivational cues (somerville & casey, 2010). in short, a mismatch between students’ developmental needs and capacities and the secondary school environment can lead to reduced motivation, engagement, interest in school, and beliefs about their ability to succeed (cauley & jovanovich, 2006; sakiz et al., 2012; schunk & meece, 2006; urdan & schoenfelder, 2006). the present study focuses on two of the most influential and widely studied types of self-beliefs, namely self-efficacy and self-concept, both of which arise from the perception and appraisal of oneself in relation to prior experience (huang, 2011; marsh & martin, 2011; valentine, dubois, & cooper, 2004). within the school context, self-efficacy refers to what individuals expect and believe they will be able to accomplish in academic tasks with whatever abilities and skills they may have (bandura, 1997; bong & skaalvik, 2003; schunk & meece, 2006). it is typically measured by asking individuals to judge how confident they are that they will be able to master their schoolwork or perform representative tasks. self-concept represents an individual’s evaluation of their actual functioning or competence in general or in a specific domain (bong & skaalvik, 2003; marsh & martin, 2011). it is typically measured by asking individuals to indicate the extent to which they endorse statements as “i am good at (a particular subject area)”. thus, the conviction that one will be able to pass a test if one studies for it is a self-efficacy judgment, while the belief that one is not good at math is a self-concept judgment. self-efficacy and self-concept are not always clearly distinguished in the literature. nonetheless, a comprehensive review by bong and skaalvik (2003) identified important differences between the constructs. these include the extent to which they are influenced by goals and designated standards, social norms, and/or internal comparisons (e.g., comparing one’s own performance in different domains or across time); whether they are oriented to the future (i.e., what one believes one could achieve) or to the past (i.e., what one has actually achieved); and whether they are changeable or stable across time. in these terms, self-efficacy is argued to be heavily goal-referenced, somewhat normatively referenced, future-oriented and temporally changeable. by comparison, self-concept is both normatively and ipsatively referenced, pastoriented and more stable across time. despite these differences, self-efficacy and self-concept share important similarities. for one, they are both shaped by individuals’ prior experiences and performance (bong & skaalvik, 2003; möller, pohlmann, köller, & marsh, 2009; möller, retelsdorf, köller, & marsh, 2011; schunk & meece, 2006). for ! reed!et !al ! ! | f l r ! ! 39! example, self-efficacy is strengthened by successful experiences and undermined by repeated failures, while self-concept in particular academic areas (e.g., mathematics, languages, science) is influenced by students’ achievement in these areas over time (möller et al., 2009; möller et al., 2011; skaalvik, & skaalvik, 2002). another important antecedent of both constructs is the appraisal of significant others such as parents and teachers, which can influence and/or reinforce individuals’ views of themselves (bong & skaalvik, 2003). thus, for instance, when teachers express that a student will succeed and is good at certain things, this can contribute to the student’s own expectation of success and positive appraisal of his/her abilities. self-efficacy and self-concept are also both influenced by comparisons in relation to personally relevant external frames of reference, notably similar peers (möller et al., 2009; möller et al., 2011; skaalvik, & skaalvik, 2002). thus, students’ self-efficacy beliefs may be influenced by the performance of similar classmates on particular tasks: when classmates are successful, students may become more confident that they too will succeed on the tasks in question, and when classmates are unsuccessful, students may become less confident of success. similarly, students’ self-concept in a particular subject area is shaped through comparing their own achievements to those of their classmates: if their math achievement is higher than that of their classmates, math self-concept is generally also higher. in the initial years of secondary school, peer comparisons are particularly influential because students are unfamiliar with many of the tasks and learning environments and have few sources of information other than their friends with which to gauge their own experiences (schunk & meece, 2006). these issues are complicated in highly differentiated early tracking educational systems where students move from heterogeneous primary school classrooms into ability-homogeneous tracks in secondary school. under these circumstances, peer comparisons are affected by the so-called ‘big-fish-little-pond’ effect (marsh, 1991; marsh & hau, 2003). this refers to the phenomenon that performance of higher ability students in mixed ability groups is higher than most of their classmates, which elevates self-judgments in comparison to others. however, their performance may be only average or below-average in groups whose performance standards are set by high ability students; self-judgments are then likely to be lower. the reverse is true for lower ability students. thus, the change in reference peer group after the move from primary to secondary school in early tracking systems is likely over time to depress self-beliefs in higher tracks and increase them in lower tracks. this is particularly so where there is little academic contact between students in different tracks (as in the netherlands), so that within-track as opposed to across-track comparisons become dominant (chmielewski et al., 2013; liu et al., 2005). investigating self-beliefs within a system of educational tracking therefore requires careful consideration of the effects of changes in reference group. 1.2 self-beliefs and math performance previous research has demonstrated strong relationships between self-beliefs and academic performance generally (caprara, vecchione, alessandri, gerbino, & barbaranelli, 2011; huang, 2011; marsh & martin, 2011; oecd, 2013; schunk & meece, 2006; valentine et al., 2004) as well as between math-related self-beliefs and math performance specifically (chiu & klassen, 2010; ferla, valcke, & cai, 2009; ireson & hallam, 2009; möller et al., 2009; möller et al., 2011; skaalvik & skaalvik, 2006; steinmayr & spinath, 2009; valentine et al., 2004). self-beliefs and performance are more strongly related when measured at the same level of specificity (bong & skaalvik, 2003; valentine et al., 2004). thus, general selfbeliefs such as the belief that one will be able to master one’s schoolwork are less strongly related to math performance than the specific belief that one is good (or not good) at math. importantly, and as a point of departure for the present research, reciprocal effects between math performance and self-beliefs have been demonstrated in longitudinal studies (marsh & martin, 2011; möller et al., 2011; pajares & schunk, 2001). these studies indicate that: (a) math performance at an earlier time point affects math performance at a later time point; (b) math performance influences students’ self-beliefs; and (c) students’ self-beliefs affect math performance. while these studies have established that self-beliefs ! reed!et !al ! ! | f l r ! ! 40! mediate the relation between math performance at successive time points presumably by means of mutual reinforcement there is currently little research that examines these effects in highly differentiated early educational tracking systems spanning the period bridging primary and secondary school. as noted, the change in reference group when students move from heterogeneous primary school classrooms to homogeneous secondary school classrooms can profoundly affect students’ self-beliefs through the mechanism of peer comparison. thus, research still needs to resolve the role of self-beliefs in this situation. finally, it is possible that the relation between self-beliefs and math performance could be moderated by sex (valentine et al., 2004). boys and girls differ in self-beliefs in several academic areas, including mathematics (herbert & stipek, 2005; ireson & hallam, 2009; jacobs et al., 2002; preckel, goetz, pekrun, & kleine, 2008; schunk & meece, 2006). moreover, girls report lower self-belief in their math competence than boys, even when performance levels are equal (else-quest, hyde, & linn, 2010; oecd, 2013). thus, self-beliefs could have different effects on math outcomes for boys and girls. 1.3 the present study the present study addresses these issues by examining the extent to which self-efficacy (i.e., how confident students are that they will be able to master their schoolwork) and math self-concept (i.e., students’ perceived math competence) mediate the relation between math performance at the end of primary school (i.e., grade 6) and the end of lower secondary school (i.e., grade 9) in a highly differentiated early tracking system. this is investigated in a multiple mediator model reflecting the reciprocal effects identified above and including self-belief measures in grade 6 and grade 9. furthermore, the study examines whether these relations are moderated by educational track and/or sex. the study draws on a large sample of typically-developing students who participated in a nationally representative, longitudinal cohort study in the netherlands. next to the longitudinal design, a strength of the study is that math performance was measured with validated, standardised national tests rather than school grades, which are known to suffer from variability in assessment and grading practices (bowers, 2011). the measures used here can therefore be considered a more reliable proxy for math performance. furthermore, performance was standardised within relevant reference peer groups. this is a crucial point, given the importance of these frames of reference in shaping self-beliefs. the large-scale longitudinal design combined with the use of validated measures allows strong inferences to be drawn about the relations of interest within the context of highly differentiated early educational tracking. the results could therefore be of considerable value in identifying factors that could affect students’ ability to maintain the levels of secondary school performance that are expected in their designated track. 2. methods this study comprises secondary analysis of data from the first and second cohort measurements of the cool5-18 study (cohort research on educational careers), a large-scale, nationally representative, longitudinal cohort study into the determinants of the cognitive and social-emotional development of children and adolescents in the netherlands1. the cool5-18 datasets are available for third-party use, as in the present study. the first cohort measurement included n = 11,609 grade 6 students from 550 primary schools. the second measurement included n = 21,384 grade 9 students from 151 secondary schools. a total of n = 2,646 students from 355 primary schools and 143 secondary schools participated in the first measurement when in grade 6 and in the second measurement when in grade 9. participants took several cognitive tests at each measurement, including a math test. they also completed self-report questionnaires ! reed!et !al ! ! | f l r ! ! 41! that included scales from externally validated questionnaires on topics including self-efficacy and school functioning. parents/caregivers completed a demographic questionnaire and schools provided administrative data (e.g., age, sex, educational track). the following paragraphs describe the participants, instruments and data relevant to the present study. 2.1 participants individuals were selected when they had participated in the cool5-18 study in both grade 6 and grade 9, when they had dutch nationality and when complete data were available for sex, educational track, both math tests (i.e., in grade 6 and grade 9), and the hypothesised mediators (i.e., self-efficacy and math self-concept). in addition, students had to be aged between 14.5 and 15.5 years at grade 9 measurement. an age-restricted window was chosen in order to have a relatively homogeneous sample of typically-developing students. accelerated and delayed students were excluded, as these students differ from their classmates in several respects relating to self-beliefs that could confound the results. for example, delayed secondary school students have significantly lower self-beliefs about their ability to do well in school (martin, 2011), while accelerated students in dutch lower secondary school have more positive self-beliefs about their school abilities and their math ability in particular (hoogeveen, van hell, & verhoeven, 2009). of the n = 969 students for whom the required data were available, 78 (8%) delayed students and 35 (3.6%) accelerated students were excluded. another 13 (1.3%) students were excluded as age was unknown. the final sample comprised n = 843 students (47% male (n = 394); mage = 14.9 years, sdage = 0.3). of these, n = 329 (39%) were in a ‘low’ track (i.e., pre-vocational education), n = 235 (28%) were in a ‘medium’ track (i.e., higher general secondary education) and n = 279 (33%) were in a ‘high’ track (i.e., preuniversity education). the students came from 188 primary schools and 101 secondary schools. 2.2 grade 6 instruments and data 2.2.1 math performance participants were administered a validated, standardised, norm-referenced math test for grade 6 (m8, 2002 version) developed by the dutch central institute for educational measurement. the test contained 107 items covering: (1) numbers and number relations; (2) arithmetic fact fluency; (3) mental arithmetic; (4) multiple operations; (5) fractions; (6) proportions; (7) percentages; (8) measurement; (9) geometry; (10) time. raw test scores were converted to proficiency scores (range: 54-160) according to standard procedure. one case with an input error was excluded from analysis. as indicated, students’ self-beliefs are influenced by comparison of their own achievements relative to relevant reference peer groups. thus, proficiency scores were standardised to denote individual performance relative to performance levels of these reference groups. in grade 6 (before stratification), class or school can be considered a relevant reference group. in the cool5-18 dataset, distribution of participants across classes was uneven, so scores were standardised per school within the whole grade 6 sample. this approach effectively nests students within schools. the whole sample (i.e., before exclusion of participants on grounds of missing data, nationality or age) was used for standardisation to keep reference groups intact. the standardised scores were used as the grade 6 math performance measure. the correlation between the standardised and unstandardised scores was high (r = .82, p < .001). 2.2.2 self-efficacy self-efficacy was measured by the academic efficacy scale of the patterns of adaptive learning scales (pals; midgley et al., 2000; urdan & midgley, 2003) from the student questionnaire. this instrument has strong psychometric properties and strong predictive and concurrent validity for both primary and secondary school students (anderman, urdan, & roeser, 2003) and is therefore highly suitable for the ! reed!et !al ! ! | f l r ! ! 42! present purpose. the self-efficacy scale contains six items (e.g., “i'm certain i will be able to master the skills taught in school this year” and “i'm certain i could figure out how to do even the most difficult classwork”), rated on a 5-point likert-type scale with choice options ranging from ‘not at all true’ to ‘very true’. items (in dutch) were coded from 1 to 5, with higher scores indicating higher self-efficacy. scale internal reliability was acceptable (cronbach’s α = .78). self-efficacy was calculated as the average of the six items. 2.3 grade 9 instruments and data 2.3.1 math performance participants were administered a validated, norm-referenced math test developed by the dutch central institute for educational measurement. test items were drawn from an item-bank of 60 items and administered in three test versions comprising 30 multiple-choice items on arithmetic, proportions, geometry and mathematical relationships. an example (translated from the original dutch) is: a group of 5 men buys one lottery ticket between them every month. a group of 8 women does the same. there is one lottery draw per month. if a prize is won by one of the tickets, then the prize is shared out among the group members: among the 5 members of the men’s group and among the 8 members of the women’s group. in april, the ticket bought by the men’s group won a prize of €100,000 and the ticket bought by the women’s group won a prize of €200,000. each man then received an amount of money and each woman received another amount of money. the amount that each man received was: a 4/5 times b 5/8 times c 5/4 times d 8/5 times the amount that each woman received. as not all participants were administered the same test version, their test scores would not be comparable under standard scoring procedures. thus, items were analysed using the one-parameter logistic model (oplm; verhelst & glas, 1995) from item response theory. when the oplm holds for a collection of test items, a student’s skill level can be estimated from every subset of items in this case, each test version. the oplm was used to translate raw test scores to skill-scores that in turn were translated to bank-scores on a scale of 0 to 100% (hambleton, swaminathan, & rogers, 1991). the bank-score indicates individual mastery level (e.g., a bank-score of 70 means that the student is expected to answer 70% of the total item-bank correctly) and is directly comparable across participants and test versions. in grade 9 in the netherlands, relevant reference groups are class or the school/track combination within which classes are embedded. again, distribution of participants across classes in the cool5-18 dataset was uneven, so bank-scores were standardised per school and track within the whole grade 9 sample to denote performance relative to this reference group. this approach nests students within school and educational track. the whole sample (i.e., before exclusion of participants on grounds of missing data, nationality or age) was used for standardisation to keep reference groups intact. the standardised scores were used as the grade 9 math performance measure. the correlation between the standardised and unstandardised scores (r = .56, p < .001) was lower than that between the standardised and unstandardised grade 6 measures. this is consistent with the fact that standardised grade 9 scores were relative to scores of students of similar ability level (i.e., track) rather than being relative to scores of students of all ability levels, as in grade 6. ! reed!et !al ! ! | f l r ! ! 43! 2.3.2 self-efficacy self-efficacy was measured by the academic efficacy scale of the patterns of adaptive learning scales from the student questionnaire, coded as described above. scale internal reliability was acceptable (cronbach’s α = .83). self-efficacy was calculated as the average of the scale items. 2.3.3 math self-concept math self-concept was measured by the item: “i am good at arithmetic and math” (in dutch) from the student questionnaire, with choice options ‘disagree’, ‘partly agree’ and ‘agree’. the item was coded from 1 to 3, with higher scores indicating a higher competence judgment. single-item measures are frequently used in research on self-beliefs, for example by having participants indicate an anticipated exam grade (e.g., vancouver & kendall, 2006). a single omnibus measure can be as psychometrically sound and effective as multiple-item measurement scales in self-report questionnaires (gardner, cummings, dunham, & pierce, 1998; robins, hendin, & trzesniewski, 2001) and can eliminate item redundancy and variance due to spurious correlations between highly related items. for example, möller et al. (2011) measured math self-concept with three items (‘‘math is one of my best subjects’’; ‘‘in math, i do quite well’’; ‘‘in math, i usually get good grades’’). the cronbach’s alphas of this scale at different time points were extremely high (.90 to .91), which may indicate item redundancy (streiner, 2003). descriptive statistics for these measures in the final sample are shown in table 1. note that mean standardised scores for math performance need not be zero as standardisation was performed within the full cool5-18 samples, which included students who did not meet the inclusion criteria for the final sample of the present study. table 1 descriptive statistics main variables sex educational track total male female low medium high n=843 n=394 n=449 n=329 n=235 n=279 m sd m sd m sd m sd m sd m sd grade 6: math performancea 0.25 0.91 0.46 0.89 0.06 0.89 -0.35 0.77 0.33 0.71 0.88 0.75 self-efficacy 3.71 0.58 3.80 0.56 3.63 0.59 3.50 0.57 3.78 0.57 3.90 0.51 grade 9: math performanceb 0.10 0.95 0.30 0.92 -0.08 0.95 0.27 0.94 0.04 0.88 -0.06 1.00 self-efficacy 3.48 0.64 3.61 0.62 3.37 0.63 3.41 0.64 3.46 0.60 3.59 0.65 math self-concept 2.13 0.77 2.26 0.72 2.01 0.79 2.07 0.76 2.07 0.76 2.24 0.77 notes. a standardised within full cool5-18 grade 6 sample (n=11,609); b standardised within full cool5-18 grade 9 sample (n=21,384). 2.4 analysis analyses were performed in ibm spss statistics 20® (α = .05). preliminary glm analyses with posthoc comparisons (bonferroni correction) were performed to establish the extent to which between-subjects differences (i.e., track and sex) and within-subjects temporal differences (i.e., between grade 6 and grade 9) were present for the main variables (i.e., standardised math performance, self-efficacy, ! reed!et !al ! ! | f l r ! ! 44! math self-concept). age was included as a covariate. note that the temporal analysis could not be performed for math self-concept as it was only measured in grade 9. for the main analysis, a multiple mediator model2 determined the extent to which the effect of grade 6 math performance on grade 9 math performance is mediated by self-efficacy and math self-concept. this model3 is depicted in figure 1, assuming the direction of effects between math performance, self-efficacy and math self-concept presented in the introduction. as multicollinearity could affect the outcomes of the analysis (hayes, 2013) and be particularly misleading when comparing effects of self-efficacy and self-concept (marsh, dowson, pietsch, & walker, 2004), variance inflation factors (vif) and tolerances were first calculated. all vifs were below 2.5 and all tolerances were above 0.40, indicating absence of multicollinearity. then, hayes’ (2013) bootstrapping method4 was used to estimate the indirect effects of the hypothesised mediators with age as a covariate as well as confidence intervals for these effects. an indirect effect is significant if the 95% confidence interval does not contain zero. effect size was calculated as the ratio of the indirect effect to the total effect of grade 6 math performance on grade 9 math performance. simple contrasts between each pair of proposed mediators identified the most influential mediator overall. finally, moderated mediation analyses tested whether the strength of the indirect effects was conditional on sex and/or track. the conditional indirect effect of a specific mediator estimates the indirect effect of that mediator at specified values of the moderator. for dichotomous moderators (e.g., sex), these values represent the two groups. for moderation by track, conditional indirect effects were estimated for the (a) low versus medium tracks; (b) low versus high tracks; and (c) medium versus high tracks. the so-called index of moderated mediation (imm) tests the equality of the conditional indirect effects in the groups being compared. when the index is not significant, these effects are equivalent. 3. results 3.1 preliminary analyses: between-subjects and temporal differences there was a large effect of track, a medium-large effect of sex and a small effect of age (table 2). males scored higher than females on all variables. tracks also differed on all variables. in grade 6, self-efficacy and math performance were lowest in the lowest track and math performance was highest in the highest track (all pbonf < .001). in grade 9, self-efficacy and math self-concept were higher in the highest track than the lowest track (pbonf = .002 and .02 respectively). although mean standardised scores (standardised relative to school and track) could be expected to be zero in all tracks, math performance was higher in the lowest track (pbonf < .03). this apparent anomaly is due to the exclusion of delayed students, most of whom were in the lowest track and had lower math performance than the other low track students. consequently, the mean of the final low track sample was higher than zero. this has no further significance for the study findings. age did not affect self-efficacy but did affect math self-concept and both math performance measures (in grade 6 and grade 9): these variables were lower for older students (r = -.10, -.12 and -.11, respectively; p < .01). temporal differences took the form of two time x track interactions: for self-efficacy (f(2,836) = 11.89, p < .001, ηp 2 = .03), with the lowest track showing the smallest decline, and for math performance (f(2,836) = 250.87, p < .001, ηp 2 = .38). in the lowest track, math performance in grade 9 was higher than in grade 6, while the reverse was true for the two higher tracks. this is consistent with the shift in reference group: many lower ability students have higher scores in grade 9 relative to students of similar ability than in grade 6 relative to students of all ability levels, while the converse is true for higher ability students. ! reed!et !al ! ! | f l r ! ! 45! table 2 between-subjects comparisons main variables wilks’ λ f (df1, df2) p ηp 2 sex .88 22.21 (5,832) < .001 .12 track .52 64.15 (10,1664) < .001 .28 sex*track .99 0.53 (10,1664) .87 .00 age (covariate) .98 3.48 (5,832) .004 .02 sex: self-efficacy g6 24.22 (1,836) < .001 .03 self-efficacy g9 32.63 (1,836) < .001 .04 math self-concept g9 24.72 (1,836) < .001 .03 math performance g6 78.79 (1,836) < .001 .09 math performance g9 34.91 (1,836) < .001 .04 track: self-efficacy g6 44.85 (2,836) < .001 .10 self-efficacy g9 6.20 (2,836) .002 .01 math self-concept g9 4.43 (2,836) .012 .01 math performance g6 225.05 (2,836) < .001 .35 math performance g9 9.64 (2,836) < .001 .02 3.2 mediation analysis the bootstrapping estimates for the multiple mediator model are presented in table 3 and figure 1. moderation estimates are presented in table 4. the total model explained 11% of variance in grade 9 math performance (f(2,840) = 60.60, p < .001). there were significant total and direct effects of grade 6 math performance on grade 9 math performance and the total indirect effect through the hypothesised mediators was also significant. grade 6 self-efficacy and grade 9 math self-concept each uniquely mediated the relationship between grade 6 math performance and grade 9 math performance, but in different directions. grade 9 math self-concept was the most influential mediator, explaining 23% of the total effect, while grade 6 self-efficacy had a smaller, negative relation with grade 9 math performance. grade 9 self-efficacy had a positive relationship to both math performance measures, but its indirect effect was not significant. moderation by sex. there were no sex differences in any of the indirect effects as none of the imms were significant, though the imm for grade 6 self-efficacy was nearly so. specifically, the negative indirect effect of grade 6 self-efficacy was significant only for females. thus, the relation between grade 6 math performance and grade 9 math performance via the hypothesised mediators was similar for both sexes, but the negative relation with high self-efficacy at the end of primary school tended to affect females in particular. moderation by track. the indirect effect of grade 9 math self-concept was significant in all tracks. the indirect effect of grade 6 self-efficacy was not significant in the low or medium tracks and was borderline significant in the high track. the indirect effect of grade 9 self-efficacy was significant only in the low track. the imms indicated one difference between tracks: the indirect effect of grade 9 self-efficacy was greater in the low track than in the medium track. ! reed!et !al ! ! | f l r ! ! 46! table 3 bootstrapping results mediation analysis estimate boot se es 95% ci b se t p total effect (c path) 0.33 0.03 10.48 <.001 direct effect (c’ path) 0.28 0.03 8.30 <.001 age (covariate) -0.23 0.11 -2.21 .03 indirect effects: total indirect 0.06 0.02 0.17 0.02 0.10 self-efficacy g6 -0.03 0.01 0.09 -0.06 -0.00 a path 0.22 0.02 10.37 <.001 b path -0.14 0.06 -2.37 .02 self-efficacy g9a 0.01 0.01 0.03 -0.00 0.02 a path 0.11 0.02 5.01 <.001 b path 0.08 0.05 1.50 .13 math self-concept g9 0.08 0.01 0.23 0.05 0.11 a path 0.24 0.03 8.13 <.001 b path 0.33 0.05 7.30 <.001 contrasts: m1-m2 -0.04 0.02 -0.07 -0.01 m1-m3 -0.11 0.02 -0.15 -0.07 m2-m3 -0.07 0.02 -0.10 -0.04 notes. a when math self-concept g9 is omitted, estimates for the b path of self-efficacy g9 (b=0.22, se=0.05, t=3.96, p<.001) and the indirect effect of self-efficacy g9 (est=0.02, se=0.01, es=0.07, ci=0.01-0.04) are significant; 5000 bootstrap samples; α = .05; es (effect size) = magnitude(indirect effect/total effect); m1 = self-efficacy g6; m2 = self-efficacy g9; m3 = math self-concept g9; estimated values are rounded to 2 decimal places (e.g., -0.0016 is reported as -0.00). figure 1. multiple mediator model with bootstrapping estimates for indirect, direct and total effects. ! reed!et !al ! ! | f l r ! ! 47! table 4 bootstrapping results moderated mediation analysis conditional indirect effects moderation by sex males females imm estimate boot se 95% ci estimate boot se 95% ci estimate boot se 95% ci self-efficacy g6 -0.01 0.02 -0.04 0.03 -0.05 0.02 -0.09 -0.02 -0.05 0.03 -0.10 0.00 self-efficacy g9 0.01 0.01 -0.00 0.02 0.00 0.01 -0.01 0.02 -0.00 0.01 -0.02 0.02 math self-concept g9 0.05 0.02 0.02 0.10 0.08 0.02 0.05 0.13 0.03 0.03 -0.02 0.08 moderation by track low track medium track high track estimate boot se 95% ci estimate boot se 95% ci estimate boot se 95% ci self-efficacy g6 -0.01 0.02 -0.04 0.03 0.01 0.02 -0.02 0.04 -0.02 0.01 -0.05 ±0.00 self-efficacy g9 0.02 0.01 0.00 0.06 -0.01 0.01 -0.04 0.01 0.01 0.01 -0.01 0.04 math self-concept g9 0.06 0.02 0.03 0.12 0.09 0.03 0.04 0.16 0.10 0.03 0.05 0.17 imm: low versus medium imm: low versus high imm: medium versus high estimate boot se 95% ci estimate boot se 95% ci estimate boot se 95% ci self-efficacy g6 -0.01 0.02 -0.06 0.03 0.02 0.02 -0.03 0.06 -0.03 0.02 -0.07 0.01 self-efficacy g9 0.03 0.02 0.00 0.07 0.01 0.02 -0.02 0.05 0.01 0.01 -0.01 0.05 math self-concept g9 -0.02 0.04 -0.10 0.05 -0.04 0.04 -0.12 0.03 0.02 0.04 -0.07 0.10 notes. 5000 bootstrap samples; α = .05; estimated values are rounded to 2 decimal places (e.g., -0.0049 is reported as -0.00). ! reed!et !al ! ! | f l r ! ! 48! 4. discussion this study investigated the extent to which self-beliefs mediate the relation between math performance at the end of primary school (i.e., grade 6) and the end of lower secondary school (i.e., grade 9) in a highly differentiated early tracking educational system. the study involved 843 typicallydeveloping students who participated in a large-scale, nationally representative, longitudinal cohort study in the netherlands. in interpreting the results, it is important to note that self-beliefs are shaped by comparisons with relevant reference groups (möller et al., 2009; möller et al., 2011; schunk & meece, 2006) and that math performance was standardised on the same basis. while grade 6 students compare themselves to classmates of all ability levels (i.e., a heterogeneous reference group), the highly differentiated tracking structure of dutch secondary education means that grade 9 students, who are established in ability-homogeneous tracks, compare themselves to classmates in the same track as themselves. the corresponding change in reference group is likely over time to depress self-beliefs as well as relative math performance in higher tracks and increase them in lower tracks (chmielewski et al., 2013; liu et al., 2005; marsh, 1991; marsh & hau, 2003). indeed, exactly this pattern was found for math performance and despite a general decline in self-efficacy from grade 6 to grade 9 the lowest track showed a much smaller decline than the other two tracks. self-efficacy in grade 6 and math self-concept in grade 9 both uniquely mediated the relation between math performance in grade 6 and in grade 9, but self-efficacy in grade 9 only added to the mediation effects in the lowest track. it should be noted that the mediation analysis method used here focuses on the unique contribution of each proposed mediator. although there was no excessively high relation between the measures of self-efficacy and math self-concept in grade 9, the existing degree of overlap clearly diminished the unique contribution of the former when the latter was taken into account (see note a of table 3). math self-concept was the most influential mediator, explaining nearly a quarter of the total effect of math performance in grade 6 on math performance in grade 9. the finding that math-specific self-beliefs (here, math self-concept) are more influential than general self-beliefs (here, self-efficacy) is consistent with previous research (bong & skaalvik, 2003; valentine et al., 2004). although causality cannot be determined from these data even with the longitudinal design, the findings suggest that higher math performance at the end of primary school may positively influence math self-concept which, in turn, may be conducive to math performance in lower secondary school. this is in line with previous research demonstrating reciprocal effects between math self-concept and performance, which shows that self-concept influences outcomes (thus, performance is improved by enhancing self-concept) and outcomes influence self-concept (thus, self-concept is enhanced by developing stronger skills) (marsh & martin, 2011; möller et al., 2011). unexpectedly, higher self-efficacy in grade 6 was negatively related to grade 9 math performance in the highest track and for girls. with the same caveat regarding causality, this could mean that, when these students are confident about their academic abilities at the end of primary school, this may lead to lower math performance at the end of lower secondary school. these findings run counter to the large body of research indicating that self-efficacy has a positive influence on performance (ferla et al., 2009; schunk & meece, 2006; skaalvik & skaalvik, 2006; valentine et al., 2004). several explanations are plausible. as discussed in the introduction, self-efficacy is shaped by several factors, including repeated successes or failures as well as appraisals by significant others. thus, students who have completed primary school with ease evidenced by repeated successes and reinforced by parents and teachers may enter secondary school expecting to succeed at academic tasks. this could particularly be the case for high ability students, who are often successful in primary school with comparatively little effort. however, these students may have difficulty changing this approach in secondary school, for example spending less time on schoolwork than is necessary (cf. vancouver & kendall, 2006). given the more exacting demands and conditions of secondary school particularly in higher tracks this approach is likely to produce lower performance. ! reed!et !al ! ! | f l r ! ! 49! additionally, disparities between learning environments in primary and secondary school could mean that learning strategies that have served well and brought success in primary school may be less effective or even counterproductive in secondary school. thus, students who persist in using such strategies could be at a disadvantage when dealing with schoolwork in secondary school. for example, students who habitually make use of rote-learning strategies (e.g., for learning multiplication tables) or standard algorithms for problem solving are likely to encounter difficulties when required to master concepts and solve more complex, novel problems in secondary school (mayer, 2002). notably, students with unrealistically high self-efficacy are often overconfident of their study methods and are unwilling to change them (schunk & pajares, 2004). furthermore, students who enter secondary school believing they will be successful face a harder ‘reality check’ when confronted with more demanding environments. this may produce distress that diverts attention away from learning and towards re-establishing well-being (boekaerts, 2006). initial problems encountered after school transition could set students on a downward path that they may not easily recover from. in any case, higher self-efficacy at the end of primary school may not necessarily be a protective factor if not appropriately managed when students move to secondary school. previous research reported sex differences in math-related self-beliefs (else-quest et al., 2010; herbert & stipek, 2005; ireson & hallam, 2009; jacobs et al., 2002; oecd, 2013; preckel et al., 2008; schunk & meece, 2006). in the present study, boys also had higher self-beliefs than girls but the patterns of relationships between self-beliefs and math performance were largely similar for both sexes. nonetheless, the negative effect of grade 6 self-efficacy on later math performance was significant only for girls, suggesting that the mechanisms proposed above could be less influential for boys, at least in typicallydeveloping students. boys have been reported to have a more positive adaptation to secondary school than girls, who are more susceptible to stress and distress during this period (akos & galassi, 2004; cauley & jovanovich, 2006). furthermore, gender differences in mathematical problem solving strategies have been found, with girls having a greater propensity for following rules and standard algorithms (leedy, lalonde, & runk, 2003; zhu, 2007). as noted, though these strategies may bring success in primary school, they may not be conducive to more complex mathematical thinking and learning later on. 5. future research this study has a number of strengths that contribute to understanding the relation between self-beliefs and math performance: specifically, the large-scale longitudinal design, the use of validated self-report and performance measures, and the inclusion of students’ external frames of reference (i.e., peer group comparisons). nonetheless, certain issues not addressed here should be investigated in future research. the negative relation between high self-efficacy at the end of primary school and later math performance was not significant for typically-developing boys. however, this relation could be stronger in underachieving or failing (i.e., delayed) boys. boys are known to overestimate their capabilities (pajares, 2002) and are also overrepresented among underachieving students and school dropouts (driessen & van langen, 2010; lamb, markussen, teese, sandberg, & polesel, 2011). it seems likely that unrealistic self-beliefs could contribute to these outcomes. thus, the mediating effects of self-beliefs on performance in delayed students should be examined in future research. furthermore, math self-concept was not measured in grade 6. assuming a degree of overlap between self-efficacy and math self-concept in grade 6, as in grade 9, it would be of interest to isolate the effects of self-efficacy in grade 6 when a concurrent measure of math self-concept is included. additional longitudinal studies with repeated measurements are needed to confirm whether the effects found here reflect causal influences. as it is often argued that enhancement of self-beliefs should be one of the key goals of education (marsh & martin, 2011; möller et al., 2009; oecd, 2013; schunk & meece, 2006), it is important to determine their impact in educational systems with highly differentiated ! reed!et !al ! ! | f l r ! ! 50! early tracking. the present findings suggest that, in these systems, students’ self-efficacy beliefs may need to be managed during the transition between primary school and the early years of secondary school. if initiatives to improve self-beliefs do not regard the realities that students face and their ability to adapt learning strategies to different environments, this could be detrimental to performance. in fact, unrealistically high self-beliefs have been linked to lower performance (chiu & klassen, 2010; vancouver & kendall, 2006). finally, while the study took account of students’ external frames of reference, an internal comparison process is also recognised in the literature, whereby students compare their own achievements across several domains. these comparisons may attenuate or inflate self-concept in a particular domain, independent of actual performance (möller et al., 2009; möller et al., 2011; skaalvik & skaalvik, 2002). future research including both frames of reference would complement other work investigating these issues in early tracking systems (e.g., möller et al., 2009; möller et al., 2011). keypoints self-beliefs mediate math performance between primary and lower secondary school in a highly differentiated early tracking educational system math self-concept explains a quarter of the total effect of earlier math performance on later math performance self-efficacy at the end of primary school has a negative relation with later math performance, particularly for girls and high-track students high self-efficacy may not necessarily be a protective factor in highly differentiated early tracking educational systems acknowledgments the authors thank the creators of the cool5-18 datasets, particularly greetje van der werf and hans kuypers. the datasets were obtained from the data archiving network services (dans) website (http://www.dans.knaw.nl). cool5-18 (2007/8): • sco-kohnstamm instituut amsterdam; its radboud universiteit nijmegen; cito arnhem; gion ru groningen • cohortonderzoek onderwijsloopbanen van 5-18 jaar cool 5-18 basisonderwijs 2007; eerste meting basisonderwijs 2007 (2007-09-01, 2008-04-30) • persistent identifier: urn:nbn:nl:ui:13-icz-r75 cool5-18 (2010/11): • gion ru groningen, cito arnhem, sco-kohnstamm instituut amsterdam, its radboud universiteit nijmegen • cohortonderzoek onderwijsloopbanen van 5-18 jaar cool 5-18 voortgezet onderwijs klas 3 2010/11 (2012-11-09) • persistent identifier: urn:nbn:nl:ui:13-y9jp-e0 ! reed!et !al ! ! | f l r ! ! 51! references akos, p. & galassi, j. p. (2004). gender and race as variables in psychosocial adjustment to middle and high school. the journal of educational research, 98, 102-108. doi:10.3200/joer.98.2.102-108 anderman, e. m., urdan, t., & roeser, r. (2003). the patterns of adaptive learning survey: history, development and psychometric properties. indicators of positive development conference. march 1213, 2003, washington, d.c. bandura, a. (1997). self-efficacy: the exercise of control. new york: freeman. boekaerts, m. (2006). self-regulation and effort investment. in k. a. renninger & i. e. sigel (eds.), handbook of child psychology : vol. 4. child psychology in practice (6th ed., pp. 345-377). hoboken, nj: john wiley & sons. bong, m., & skaalvik, e. m. (2003). academic self-concept and self-efficacy: how different are they really? educational psychology review, 15, 1-40. doi: 10.1023/a:1021302408382 bowers, a. j. (2011). what's in a grade? the multidimensional nature of what teacher-assigned grades assess in high school. educational research and evaluation: an international journal on theory and practice, 17, 141-159. doi:10.1080/13803611.2011.597112 brophy, j. (2006). grade repetition. education policy series no. 6. belgium/paris: international academy of education and international institute for educational planning, unesco. caprara, g. v., vecchione, m., alessandri, g., gerbino, m., & barbaranelli, c. (2011). the contribution of personality traits and self-efficacy beliefs to academic achievement: a longitudinal study. british journal of educational psychology, 81, 78-96. doi:10.1348/2044-8279.002004 cauley, k. m., & jovanovich, d. (2006). developing an effective transition program for students entering middle school or high school. the clearing house: a journal of educational strategies, issues and ideas, 80, 15-25. doi:10.3200/tchs.80.1.15-25 chiu, m. m., & klassen, r. m. (2010). relations of mathematics self-concept and its calibration with mathematics achievement: cultural differences among fifteen-year-olds in 34 countries. learning and instruction, 20, 2-17. doi:10.1016/j.learninstruc.2008.11.002 chmielewski, a. k., dumont, h., & trautwein, u. (2013). tracking effects depend on tracking type: an international comparison of students' mathematics self-concept. american educational research journal, 50, 925-957. doi:10.3102/0002831213489843 crone, e. a., wendelken, c., van leijenhorst, l., honomichl, r. d., christoff, k., & bunge, s. a. (2009). neurocognitive development of relational reasoning. developmental science, 12, 55-66. doi:10.1111/j.1467-7687.2008.00743.x driessen, g., mulder, l., ledoux, g., roeleveld, j., & van der veen, i. (2009). cohortonderzoek cool5-18: technisch rapport basisonderwijs, eerste meting 2007/08 [cohort study cool5-18: technical report primary education, first cohort measurement 2007/08]. nijmegen: its/amsterdam: sco-kohnstamm instituut, the netherlands. driessen, g., & van langen, a. (2010). de onderwijsachterstand van jongens. omvang, oorzaken en interventies [the educational disadvantage of boys. extent, causes and interventions]. nijmegen, the netherlands: its radboud university nijmegen. dumontheil, i., houlton, r., christoff, k. & blakemore, s-j. (2010). development of relational reasoning during adolescence. developmental science, 13, f15-f24. doi:10.1111/j.1467-7687.2010.01014.x else-quest, n-m., hyde, j. s, & linn, m. c. (2010). cross-national patterns of gender differences in mathematics: a meta-analysis. psychological bulletin, 136, 103-127. doi:10.1037/a0018053 fenzel, l. m. (2000). prospective study of changes in global self-worth and strain during transition to middle school. journal of early adolescence, 20, 93-116. doi:10.1177/0272431600020001005 ferla, j., valcke, m., & cai, y. (2009). academic self-efficacy and academic self-concept: reconsidering structural relationships. learning and individual differences, 19, 499-505. doi:10.1016/j.lindif.2009.05.004 gardner, d. g., cummings, l. l., dunham, r. b., & pierce, j. l. (1998). single-item versus multiple-item measurement scales: an empirical comparison. educational and psychological measurement, 58, ! reed!et !al ! ! | f l r ! ! 52! 898-915. doi:10.1177/0013164498058006003 hambleton, r. k., swaminathan, h., & rogers, h. j. (1991). fundamentals of item response theory. newbury park, ca: sage. hanushek, e. a., & wößmann, l. (2006). does educational tracking affect performance and inequality? differences-in-differences evidence across countries. the economic journal, 116(510), c63-c76. doi:10.1111/j.1468-0297.2006.01076.x hayes, a. f. (2013). introduction to mediation, moderation, and conditional process analysis. new york, ny: guilford press. herbert, j., & stipek, d. (2005). the emergence of gender differences in children’s perceptions of their academic competence. journal of applied developmental psychology, 26, 276-295. doi:10.1016/j.appdev.2005.02.007 hoogeveen, l., van hell, j. g., & verhoeven, l. (2009). self-concept and social status of accelerated and nonaccelerated students in the first 2 years of secondary school in the netherlands. gifted child quarterly, 53, 50-67. doi:10.1177/0016986208326556 huang, c. (2011). self-concept and academic achievement: a meta-analysis of longitudinal relations. journal of school psychology, 49, 505-528. doi:10.1016/j.jsp.2011.07.001 ireson, j., & hallam, s. (2009). academic self-concepts in adolescence: relations with achievement and ability grouping in schools. learning and instruction, 19, 201-213. doi:10.1016/j.learninstruc.2008.04.001 jacob, b. a., & lefgren, l. (2009). the effect of grade retention on high school completion. american economic journal: applied economics, 1(3), 33-58. doi:10.1257/app.1.3.33 jacobs, j. e., lanza, s., osgood, d. w., eccles, j. s., & wigfield, a. (2002). changes in children’s selfcompetence and values: gender and domain differences across grades one through twelve. child development, 73, 509-527. doi:10.1111/1467-8624.00421 lamb, s., markussen, e., teese, r., sandberg, n., & polesel, j. (2011). school dropout and completion: international comparative studies in theory and policy. dordrecht, the netherlands: springer science+business media b.v. leedy, m. g., lalonde, d., & runk, k. (2003). gender equity in mathematics: beliefs of students, parents, and teachers. school science and mathematics, 103, 285-292. doi:10.1111/j.1949-8594.2003.tb18151.x liu, w. c., wang, c. k., & parkins, e. j. (2005). a longitudinal study of students’ academic self-concept in a streamed setting: the singapore context. british journal of educational psychology, 75, 567-586. doi:10.1348/000709905x42239 marsh, h. w. (1991). failure of high-ability high schools to deliver academic benefits commensurate with their students' ability levels. american educational research journal summer, 28, 445-480. doi:10.3102/00028312028002445 marsh, h. w., dowson, m., pietsch, j., & walker, r. (2004). why multicollinearity matters: a reexamination of relations between self-efficacy, self-concept, and achievement. journal of educational psychology, 96, 518-522. doi:10.1037/0022-0663.96.3.518 marsh, h. w., & hau, k-t. (2003). big-fish-little-pond effect on academic self-concept: a cross-cultural (26-country) test of the negative effects of academically selective schools. american psychologist, 58, 364-376. doi:10.1037/0003-066x.58.5.364 marsh, h. w., & martin, a. j. (2011). academic self-concept and academic achievement: relations and causal ordering. british journal of educational psychology, 81, 59-77. doi:10.1348/000709910x503501 martin, a. j. (2011). holding back and holding behind: grade retention and students’ non-academic and academic outcomes. british educational research journal, 37, 739-763. doi:10.1080/01411926.2010.490874 mayer, r. e. (2002). rote versus meaningful learning. theory into practice, 41, 226-232. doi:10.1207/s15430421tip4104_4 midgley, c., maehr, m. l., hruda, l. z., anderman, e., anderman, l., freeman, k. e., et al., (2000). manual for the patterns of adaptive learning scales (pals). ann arbor, mi: university of michigan. ! reed!et !al ! ! | f l r ! ! 53! möller, j., pohlmann, b., köller, o., & marsh, h. w. (2009). a meta-analytic path analysis of the internal/external frame of reference model of academic achievement and academic self-concept. review of educational research, 79, 1129-1167. doi:10.3102/0034654309337522 möller, j., retelsdorf, j., köller, o., & marsh, h. w. (2011). the reciprocal internal/external frame of reference model: an integration of models of relations between academic achievement and selfconcept. american educational research journal, 48, 1315-1346. doi:10.3102/0002831211419649 mullis, i. v. s., martin, m. o., foy, p., & arora, a. (2012). timss 2011 international results in mathematics. chestnut hill, ma: timss & pirls international study center. oecd. (2012). equity and quality in education: supporting disadvantaged students and schools. oecd publishing. doi:10.1787/9789264130852-en oecd. (2013). pisa 2012 results: ready to learn students’ engagement, drive and self-beliefs (volume iii). pisa, oecd publishing. doi:10.1787/9789264201170-en osterman, k. f. (2000). students' need for belonging in the school community. review of educational research, 70, 323-367. doi:10.3102/00346543070003323 pajares, f. (2002). gender and perceived self-efficacy in self-regulated learning. theory into practice, 41, 116-125. doi:10.1207/s15430421tip4102_8 pajares, f., & schunk, d. h. (2001). self-beliefs and school success: self-efficacy, self-concept, and school achievement. in r. j. riding & s. g. rayner (eds.), self perception (pp. 239-265). westport, ct: ablex publishing. preckel, f., goetz, t., pekrun, r., & kleine, m. (2008). gender differences in gifted and average-ability students: comparing girls’ and boys’ achievement, self-concept, interest and motivation in mathematics. gifted child quarterly, 52, 146-159. doi:10.1177/0016986208315834 robins, r. w., hendin, h. m., & trzesniewski, k. h. (2001). measuring global self-esteem: construct validation of a single-item measure and the rosenberg self-esteem scale. personality and social psychology bulletin, 27, 151-161. doi:10.1177/0146167201272002 sakiz, g., pape, s. j., & hoy, a. w. (2012). does perceived teacher affective support matter for middle school students in mathematics classrooms? journal of school psychology, 50, 235-255. doi:10.1016/j.jsp.2011.10.005 schunk, d. h., & meece, j. l. (2006). self-efficacy development in adolescence. in f. h. pajares & t. c. urdan (eds.), self-efficacy beliefs of adolescents (pp. 71-96). greenwich, ct: information age publishing. schunk, d. h., & pajares, f. (2004). self-efficacy in education revisited: empirical and applied evidence. in d. m. mcinerney & s. van etten (eds.), big theories revisited (pp. 115-138). greenwich, ct: information age publishing. skaalvik, e. m., & skaalvik, s. (2002). internal and external frames of reference for academic self-concept. educational psychologist, 37, 233-244. doi:10.1207/s15326985ep3704_3 skaalvik, e. m., & skaalvik, s. (2006). self-concept and self-efficacy in mathematics: relation with mathematics motivation and achievement. in a. p. prescott (ed.), the concept of self in education, family and sports (pp. 51-74). new york: nova science publishers. somerville, l. h., & casey, b. j. (2010). developmental neurobiology of cognitive control and motivational systems. current opinion in neurobiology, 20, 236-241. doi:10.1016/j.conb.2010.01.006 steinmayr, r., & spinath, b. (2009). the importance of motivation as a predictor of school achievement. learning and individual differences, 19, 80-90. doi:10.1016/j.lindif.2008.05.004 streiner, d. l. (2003). starting at the beginning: an introduction to coefficient alpha and internal consistency. journal of personality assessment, 80, 99-103. doi:10.1207/s15327752jpa8001_18 urdan, t., & midgley, c. (2003). changes in the perceived classroom goal structure and pattern of adaptive learning during early adolescence. contemporary educational psychology, 28, 524-551. doi:10.1016/s0361-476x(02)00060-7 urdan, t., & schoenfelder, e. (2006). classroom effects on student motivation: goal structures, social relationships, and competence beliefs. journal of school psychology, 44, 331-349. doi:10.1016/j.jsp.2006.04.003 ! reed!et !al ! ! | f l r ! ! 54! valentine, j. c., dubois, d. l., & cooper, h. (2004). the relation between self-beliefs and academic achievement: a meta-analytic review. educational psychologist, 39, 111-133. doi:10.1207/s15326985ep3902_3 vancouver, j. b., & kendall, l. n. (2006). when self-efficacy negatively relates to motivation and performance in a learning context. journal of applied psychology, 91, 1146-1153. doi:10.1037/0021-9010.91.5.1146 verhelst, n. d. , & glas, c. a. w. (1995). the one parameter logistic model. in g. h. fischer & i. w. molenaar (eds.), rasch models: foundations, recent developments, and applications (pp. 215-238). new york: springer-verlag. zhu, z. (2007). gender differences in mathematical problem solving patterns: a review of literature. international education journal, 8, 187-203. zijsling, d., keuning, j., naayer, h., & kuyper, h. (2012). cohortonderzoek cool5-18: technisch rapport meting vo-3 in 2011 [cohort study cool5-18: technical report grade 9 cohort measurement in 2011]. groningen, the netherlands: gion. appendix: methodological footnotes 1 the cool5-18 study was commissioned by the netherlands organisation for scientific research and the ministry of education, culture and science, and was carried out by a broad consortium of research and assessment organisations in the netherlands. full descriptions of participants, methods and procedures are provided in the technical reports (driessen, mulder, ledoux, roeleveld, & van der veen, 2009; zijsling, keuning, naayer, & kuyper, 2012). 2 a mediation model is a type of structural equation model, referring to a sequence of relations in which an independent variable affects a dependent variable by influencing intervening (i.e., mediator) variables. the order of the variables must be established on theoretical, logical or procedural grounds (hayes, 2013). 3 the ai paths represent the effect of grade 6 math performance on the proposed mediators. the bi paths represent the effect of the proposed mediators on grade 9 math performance, partialling out the effect of grade 6 math performance. path c represents the total effect of grade 6 math performance on grade 9 math performance and path c’ represents the direct effect of grade 6 math performance on grade 9 math performance after controlling for the proposed mediators. the specific indirect effect of grade 6 math performance on grade 9 math performance through a particular mediator (i.e., the unique ability of the mediator to mediate the effect of grade 6 math performance on grade 9 math performance conditional on the other mediators) is the product of the two paths linking grade 6 math performance to grade 9 math performance via that mediator (i.e., ai*bi). the total indirect effect of grade 6 math performance on grade 9 math performance is the sum of the specific indirect effects. the total effect of grade 6 math performance on grade 9 math performance (path c) is the sum of the direct effect and all of the specific indirect effects. 4 the bootstrapping method is implemented in hayes’ process macro (obtained from http://www.afhayes.com/spss-sas-and-mplus-macros-and-code.html). a strength of this procedure is that it does not make assumptions about the sampling distribution of the indirect effects or force choices about estimation or constraint of residual covariances. it resamples thousands of times from the dataset and estimates the indirect effects in each resample, thereby providing an empirical approximation of and confidence intervals for these effects. bias-corrected confidence intervals were used, as indirect effects usually have a skewed distribution. a heteroscedasticity-consistent standard error estimator was used, which reduces the likelihood that inference validity is compromised by any potential violation of homoscedasticity. model 4 in the process macro was used to estimate the indirect effects of the hypothesised mediators. model 59 was used for the moderated mediation analyses. frontline learning research 6 (2014) 1-25 issn 2295-3159 corresponding author: heidi hyytinen, institute of behavioural sciences, the university of helsinki, 00014 university of helsinki, finland, email: heidi.m.hyytinen@helsinki.fi doi: http://dx.doi.org/10.14786/flr.v2i4.124 1 | f l r the complex relationship between students’ critical thinking and epistemological beliefs in the context of problem solving heidi hyytinen a , katariina holma b , auli toom a , richard j. shavelson c , sari lindblom-ylänne a a university of helsinki, finland b university of eastern finland, finland c sk partners, llc & graduate school of education, stanford university, usa article received 2 may 2014 / revised 14 june 2014 / accepted 27 july 2014 / available online 24 september 2014 abstract the study utilized a multi-method approach to explore the connection between critical thinking and epistemological beliefs in a specific problem-solving situation. data drawn from a sample of ten third-year bioscience students were collected using a combination of a cognitive lab and a performance task from the collegiate learning assessment (cla). the cognitive-lab data were analysed using thematic analysis. the findings showed that students’ epistemological beliefs were interwoven into their critical thinking: students used critical thinking as a tool (1) for enhancing understanding and (2) for determining truth or falsehood. based on this classification, students could be placed in one of two qualitative profiles, either (1) thorough processing or (2) superficial processing. the results indicated that students who showed superficial processing palmed off justification for knowing on authoritative figures. in contrast to previous studies these students did not consider knowledge to be absolutely certain or unquestionable. the findings also show that students with thorough processing believed knowledge to be tentative and fallible, but did not share the relativist view of knowledge where any claim counts because all knowledge is relative. all ten students shared a fallibilist view of knowledge. keywords: critical thinking; epistemological beliefs; cognitive lab; relativism; fallibilism h.hyytnen et al. 2 | f l r 1. introduction critical thinking has been singled out as one of the most important skills for citizens of the twentyfirst century (halpern, 2014). mastering critical thinking is thus a goal that can be found in almost every higher education curriculum today. however, recent studies have raised concerns that even though most students make significant progress in learning concepts and procedures during their university studies, some students show little if any growth in critical thinking (arum & roksa, 2011a, 2011b; bok, 2006; pascarella, blaich, martin & hanson, 2011). in the field of higher education, research on critical thinking has generally focused on the development of critical thinking skills (e.g. arum & roksa, 2011a; heijltjes, van gog, leppink & paas, 2014). researchers have also highlighted the importance of understanding critical thinking as a social activity (e.g. arum & roksa, 2011b; kuhn, 2005; moore, 2004; 2013). in this exploratory study we provide a multidimensional framework for analysing critical thinking by combining theoretical aspects from philosophical, educational and psychological approaches. in our view the concept of critical thinking is closely connected to the concepts of ‘knowledge’ and ‘knowing’. furthermore, we assume that critical thinking cannot be formulated by referring to skills alone, but also always involves a disposition to use these skills adequately (see bailin & siegel, 2003; holma, 2014; siegel, 1988). previous research on critical thinking and personal epistemology has frequently applied quantitative multiple-choice tests, questionnaires or qualitative interviews (see e.g. australian council of education research, 2001; heijltjes, van gog, leppink & paas, 2014; greene & yu 2014; lahtinen & pehkonen, 2013; tremblay, lalancette & roseveare, 2012). recently, many researchers have questioned the reliability and adequacy of self-report questionnaires (greene & yu, 2014; elby & hammer, 2001). as a result, researchers have stated that there is a need for studies that assess the performance of students directly (e.g. elby & hammer, 2001; hofer, 2004; stes, min-leliveld, gijbels and van petegem 2009). at the same time researchers have also assumed that one assessment method is not enough to evaluate complex cognitive processes such as reasoning (e.g. baartman, bastiaens, kirschner & vleuten, 2007; dierick & dochy, 2001; maclellan, 2004). this study responds to current concerns by exploring students’ critical thinking as well as their epistemological beliefs, as elaborated upon below, in a problem-solving situation to which we applied a multi-method qualitative approach. a think-aloud method was used as the students worked through an openended performance task. our aim is to identify and understand qualitative differences in the critical thinking of students and in their beliefs about knowledge, as well as in their personal relationships. 2. critical thinking in university-level studies critical thinking is often ‘regarded as fundamental aim of education’ (bailin & siegel, 2003, p.188; cf. dewey, 1910). in a university context critical thinking has an essential role and is an important component of the learning outcomes (bok, 2006). critical thinking is defined as a process that enables an individual to make an informed decision about conflicting claims (ennis, 1991; fisher, 2011; bailin & siegel, 2003). it is purposeful, reasoned and reflective thinking (ennis, 1991; american philosophical association, 1990). a critical thinker knows how to assess the strength of evidence and the reasons that are relevant to the particular context or type of task, and also shows the disposition to draw on these skills (bailin & siegel, 2003; scheffler, 1965, halpern, 2014). critical thinking is seen as a skilful activity in which a person may be more or less proficient (fisher, 2011; scheffler, 1965). definitions of critical thinking typically include a list of the thinking skills that characterise an ideal critical thinker. for example, fisher (2011) lists the following: the ability to identify the elements in a reasoned case, especially reasons and conclusions; the abilities to identify and evaluate assumptions; the abilities to clarify and interpret expressions and ideas; to be able to judge the acceptability, especially the credibility, of claims; to evaluate arguments, analyse, evaluate and produce explanations; to be able to analyse, evaluate, and make decisions; to draw inferences and produce arguments (see also halpern, 2014). university studies require all of these abilities. h.hyytnen et al. 3 | f l r however, many philosophers have argued that critical thinking cannot be conceptualised merely by referring to a prescribed set of skills (bailin & siegel, 2003; holma, 2014; fisher, 2011; siegel, 1988, scheffler, 1965; see also halpern, 2014). it may be that a person has acquired the skills, but does not use them (fisher, 2011). as holma (2014) has pointed out, it is not enough for students to have critical thinking skills; they also need to use these skills effectively. thus, critical thinking always involves both the essential skills or abilities and the disposition to use them (bailin & siegel, 2003, holma, 2014; siegel, 1988). previous studies have called attention to the fact that students’ critical thinking skills do not always develop during university studies (arum & roksa, 2011a; bok, 2006; pascarella, blaich, martin & hanson, 2011). arum and roksa (2011b) demonstrated in their longitudinal study that a large number of university students showed no significant improvement in a range of critical thinking skills, such as reasoning and problem solving. however, a recent study by heijltjes and colleagues (2014) has shown that the combination of explicit instruction and practice has proven successful in improving students’ performance in reasoning skills. 3. knowledge and knowing in critical thinking critical thinking demands a comprehensive use of different types of knowledge (bok, 2006; ennis, 1991). there is a reciprocal relationship between ‘critical thinking’, ‘knowledge’ and ‘knowing’; on the one hand, students need knowledge about a phenomenon before they can think about it critically (halpern, 2014); on the other hand, students must have the necessary skills to evaluate that knowledge. the concepts of ‘knowledge’ and ‘knowing’ are thus substantial aspects of conceptualising critical thinking. there are several different definitions and classifications of the concept of knowledge. for example, philosophical epistemologists usually differentiate amongst three types of knowledge: propositional knowledge, procedural knowledge and knowledge by acquaintance (everitt & fisher, 1995; ichikawa & steup, 2012), although there is no consensus on the interpretation of knowledge or on the number of types of knowledge (fenstermacher, 1994). for our purposes the distinction between propositional and procedural knowledge has theoretical importance. propositional knowledge is defined as knowing that ‘such-and-such is the case’. this is sometimes referred to as factual or declarative knowledge. propositional knowledge (i.e. ‘knowing that’) is usually distinguished from procedural knowledge (i.e. ‘knowing how’) (ryle, 1949). in philosophical discussions propositional knowledge is related to such epistemological concepts as truth, justification, reason and evidence (ryle, 1949; scheffler, 1965, see also niiniluoto, 1999; shope, 2004). scheffler (1965) argued that the ‘knowing that’ attributes of a person may reveal his epistemological orientations, such as the criteria for justifying knowing. empirical research on personal epistemology focuses particularly on these personal orientations. procedural knowledge, meaning ‘knowing how’ to do something (knowing how to analyse, knowing how to swim, etc.; see everitt & fisher, 1995; shope, 2004), is related to possessing a skill (scheffler, 1965). in this sense critical thinking represents procedural knowledge, which is consistent with the other aspect of critical thinking mentioned above. however, several researchers have assumed that procedural knowledge always involves some propositional knowledge (i.e. everitt & fisher, 1995; smith 2002; markowitsch & messerer, 2007). for example, if a person knows how to play chess, he will probably know certain facts (e.g. rules) about playing chess. smith (2002) has emphasized that an individual has a certain skill only when his performance reflects both procedural and propositional knowledge. in sum, critical thinking involves a disposition to think critically, having the necessary propositional knowledge about a phenomenon and having the thinking skills (i.e. procedural knowledge) to evaluate that knowledge (cf. halpern, 2014). h.hyytnen et al. 4 | f l r 4. students’ epistemological beliefs as premises of critical thinking the term ‘personal epistemology’ or, alternatively, ‘epistemological belief’ is defined as an individual’s views of the nature of knowledge and knowing. the term also includes a view of one’s personal beliefs as a knower (pintrich, 2002; hofer, 2004). the concept of ‘personal epistemology can be described along a continuum from less sophisticated to more sophisticated’ ways of knowing (kaartinen-koutaniemi & lindblom-ylänne, 2012, p. 2) or a progress ‘from a state of simple, absolute certainty into a multifaceted, evaluative system’ (west, 2004, p. 61). during this process the individual changes from a passive recipient of knowledge to an active participant in constructing and evaluating knowledge (hofer & pintrich, 2002; kuhn, 2005; king & kitchener, 2004). over time epistemological beliefs develop more and more toward relativistic beliefs (hofer & pintrich, 1997, 2002). previous research on personal epistemology has found that the ability to think critically is embedded in a progression of epistemological beliefs (i.e. king & kitchener, 2004; kuhn & weinstock, 2002; kuhn, 1999; 2005). several researchers have hypothesised that students with weak critical thinking skills have an absolute view of knowledge. when students move on to the most developed epistemological level, their critical thinking tends to improve as well (bok, 2006; kuhn, 1999; kuhn & weinstock, 2002). it has also been demonstrated that students’ epistemological beliefs play an important role in their ability to evaluate the credibility of competing claims (barzilai & zohar, 2012). whether instruction has any influence on the development of epistemological beliefs is currently under discussion (e.g. valanides & angeli, 2005; lahtinen & pehkonen, 2013). however, there is evidence that not all university students reach the most highly developed level of personal epistemology (kuhn & weinstock, 2002; kaartinen-koutaniemi & lindblom-ylänne, 2012; king & kitchener, 2004; perry, 1970). king and kitchener (2004) have found that only advanced doctoral students consistently show the highest level of epistemological beliefs. furthermore, kaartinen-koutaniemi and lindblom-ylänne (2008, 2012) have shown that there is a considerable variation in personal epistemology among final-year master’s students. their results also showed variations between students in different age groups, study phases and disciplines (see also hofer, 2006; muis, bendixen & haerle, 2006). in addition, researchers have assumed that students’ epistemological beliefs may vary within the same discipline or domain (hammer & elby, 2003; greene & yu, 2014). 5. critical thinking and different conceptions of knowledge as the brief review above indicates, the literature of personal epistemology makes a distinction between a lower level of epistemological beliefs, in which knowledge is perceived as consisting of unchanging facts and is acquired directly from external authorities, and higher level epistemological beliefs, in which knowledge is seen as uncertain and constructed by the individual himself (kuhn & weinstock, 2002; hofer, 2005; valanides & angeli, 2005). several researchers have stated that students with higherlevel epistemological beliefs have better critical thinking skills than students with lower level epistemological beliefs (king & kitchener, 2004; kuhn & weinstock, 2002; kuhn, 1999; 2005). recently, holma and hyytinen (2014) have argued that there are several conceptual problems in this kind of hierarchical theory of knowledge (see also elby & hammer, 2001). in this section we focus on three conceptions of knowledge identified in the review of the literature on epistemology. these conceptions, specifically relativism, metaphysical realism and fallibilism, have theoretical importance for conceptualising critical thinking. a relativist position implies that all knowledge is relative to the person who believes or that all interpretations, theories and beliefs are equally right. because all beliefs are equally right, there is no reason to compare and evaluate different beliefs—all beliefs are equally justified (holma, 2012; holma & hyytinen, 2014). the problem of relativism becomes clear when it is related to the concept of critical thinking (holma & hyytinen, 2014). given that relativism allows people to construct their own ‘personal truths’, critical thinking turns out to be unnecessary (bleazby, 2011). for example, there is no need to evaluate ideas or h.hyytnen et al. 5 | f l r search for alternatives, because all ideas are equally trustworthy and justifiable (bleazby, 2011; holma & hyytinen, 2014). therefore, the idea that critical thinking presupposes the relativist view of knowledge is untenable. metaphysical realism is an epistemological position that assumes that ‘our knowledge and symbol systems [i.e. theories] directly reflect the structure of reality’ (holma, 2004, p. 421; putnam, 1981). the literature of personal epistemology seems to understand realism as metaphysical realism (see e.g. kuhn 2005; kuhn & weinstock, 2002; see also holma & hyytinen 2014), and furthermore, it appears to connect with metaphysical realism the assumption of the possibility of the certainty of human knowledge. as king and kitchener (2004) put it, knowledge is ‘obtained with certainty by direct observation’ (p. 7). 1 in the context of metaphysical realism, critical thinking turns out to be pointless. fallibilism is an epistemological position that implies that all our beliefs are liable to error (reed, 2002; niiniluoto, 1999; holma, 2012). contrary to relativism, fallibilism does not assume that all beliefs or theories are equally right. it presumes the possibility of improving our current conceptions, theories or beliefs. as holma (2012, p. 399) aptly states of fallibilism, ‘this position, like the belief that all human knowledge is uncertain, coheres with the evolutionary understanding of knowledge: the bodies of knowledge we now have may be mistaken and thus [are] possible subjects for revision, but they have, nevertheless, survived the process of evolution to this point; as such, they provide the best available starting point for choices and action of the present moment concerning further inquiry’ (see also peirce, 1934). from this point of view, epistemological fallibilism fits the presumption of critical thinking. previous research on personal epistemology lacks the notion of epistemological fallibilism. 1 king and kitchener (2004) do not call the lowest level of reflective thinking realism. however, in their model they maintain that, at the most limited level of thinking, knowledge is certain and is obtained from direct observation (p.7). this position fits metaphysical realism. h.hyytnen et al. 6 | f l r table 1. summary of the key concepts of this study concept description critical thinking process that enables an individual to make an informed decision between conflicting claims. it involves skills and dispositions (e.g. attitude and motivation) to evaluate the reliability and relevance of evidence, to identify arguments, to analyse, interpret and synthesise data from a variety of sources, to draw valid conclusions and address opposing viewpoints). 1 critical thinking also involves ‘knowing how to do something’ (procedural knowledge) and ‘knowing that’ (propositional knowledge). 2 epistemological beliefs students’ thoughts/beliefs about the nature of knowledge and the nature of knowing, including personal beliefs about themselves as knowers. 3 metaphysical realism the idea that human beliefs are direct copies of reality. the belief that all human knowledge is certain is connected to this epistemological position. 4 relativism the view that all knowledge is relative to the person who believes or that all interpretations/beliefs are equally correct. because all beliefs are equally correct, there are no means for comparing different beliefs. 5 epistemological fallibilism the view that human knowledge is uncertain. in contrast to relativism, it presumes the possibility of improving our current conceptions, theories or beliefs, seeking criteria for evaluating, comparing and justifying these beliefs or theories. 5 1 based on bailin & siegel (2003); ennis (1991); fisher (2011); fisher & scriven (1997), siegel (1988). 2 based on scheffler (1965); cf. also ryle (1949). 3 based on pintrich (2002). 4 based on holma (2004); putnam (1981). 5 based on holma (2012); holma & hyytinen (2014); peirce (1934). table 1 provides a summary of the definitions of the key concepts in this study. with this broader framework we are able to pin down different areas in critical thinking and epistemological beliefs, which have been shown to be vital for conceptualising these phenomena in prior studies or theorizations. although the conventions of critical thinking and epistemological beliefs are commonly embodied in social practices (e.g. arum & roksa, 2011b; elby & hammer, 2001; kuhn, 2005), the underlying dimensions (i.e. evaluating the reliability and relevance of evidence, identifying arguments, analysing information, addressing opposing viewpoints, reasoning) are relevant in each scientific discipline. moreover, in line with previous studies we expected that students’ epistemological beliefs and critical thinking might vary within the same discipline (see greene & yu, 2014; see also bailin & siegel, 2003). in our study we focused on the qualitative differences in critical thinking and personal epistemological beliefs by examining ten third-year university students’ thinking and performance in a cognitively-demanding authentic problem-solving situation. the aims of this study are twofold: to identify and describe qualitative differences in third-year university students’ critical thinking skills and epistemological beliefs in a problem-solving situation, and to analyse the interconnections between students’ h.hyytnen et al. 7 | f l r personal epistemologies and critical thinking skills. to achieve these aims, we formulated the following research questions: (1) how are critical thinking and epistemological beliefs presented in a problem-solving situation in a specific group of third-year university students? (2) how do critical thinking and epistemological beliefs vary from one individual to the next? 6. research methods and materials 6.1 participants this study was conducted with ten third-year bioscience students drawn from the fields of biological and environmental sciences in a research-intensive university in finland. the target population consisted of all third-year bioscience students in this particular university. first, we selected 40 students at random (approximately one-half of the target population). then we invited all students selected to participate in our study. ten out of 40 students volunteered. seven of the participants were female and three male. the students’ ages varied from 22 to 29, the mean age being 24. all came from a homogeneous cultural background, and all shared the same first language (finnish). in addition, the students had the same national high school certificate and had enrolled in the same bachelor’s study programme. the participants were at the same phase of their studies, that is, near the end of their bachelor’s studies, with the exception of one student whose study pace had been slower. during their university careers, the students had participated in lectures, practical laboratories, seminars, field courses and web-based teaching. we are aware that the sample size is too small for generalization. however, the purpose of this study is to deepen understanding of critical thinking and epistemological beliefs, for example, so as to describe how these phenomena vary across individuals in this specific group of students. 6.2 procedures for this study we collected a large body of data for each participant using a multi-method approach (johnson, onwuegbuzie & turner, 2007), including think-aloud protocol, interviews and a collegiate learning assessment (cla) performance task. the data collection was carried out in the spring of 2010 and consisted of ten cognitive labs. the students came to a classroom and were given the details of the study. the students spent two to three hours reading and responding to the performance task. in responding to the task, the students were asked to verbalise their thoughts (to ‘think aloud’). in the course of carrying out the task while thinking aloud, the students were also asked to write a memorandum addressing critical issues in the task and recommending —and justifying— a course of action. following the task, the students were interviewed about their processes in carrying out the task. students were also asked questions about critical thinking, knowledge and knowing. details of the procedures are provided below in appropriate sections. 6.2.1 collegiate learning assessment (cla) the collegiate learning assessment (cla) instrument for assessing college-level critical thinking skills used in this study was developed by the council for aid to education (cae). the cla is a standardised, open-ended test and it measures analytical reasoning, problem solving and written communication. unlike most standardised tests used in measuring critical thinking, the version of the cla used here did not include any multiple-choice questions (klein, benjamin, shavelson & bolus, 2007). the cla consists of two elements: a set of performance tasks and a set of analytical writing tasks (shavelson, 2010). only the performance task was used in this study. recent studies have found that open-ended problems with no obvious solution provide an opportunity for students to reflect on their beliefs about knowledge (barzilai & zohar, 2012; ferguson & bråten, 2012). for example, in a problem-solving situation students would need to determine the trustworthiness, and relevance, of different types of information h.hyytnen et al. 8 | f l r presented to them, co-ordinate various pieces of information related to the problem and consider the underlying assumptions and claims (shavelson, 2010). the cla performance task presents a realistic situation or problem and includes directions, openended questions and a document library containing reading material. in order to respond to the task, the students need to read, organise, synthesise and analyse information (which might be reliable/unreliable; relevant/irrelevant to the completion of the task; see shavelson, 2010) from multiple documents (for example letters, memos, summaries of research reports, articles, diagrams, graphs, maps, interview notes). in doing these activities the students need to assess their confidence in information taken from various sources, including the relevance of the source, and thereby deal with conflicting information. they then need to decide on a course of action and provide a reasoned explanation and justification for their course, drawing on supporting information from the document library (klein et al., 2007; shavelson, 2010). they also have to argue for and against alternative explanations. the specific performance task used in this study is proprietary and consequently cannot be described here. an example of a representative cla performance is presented in figure 1. adapted from r. shavelson, 2010, measuring college learning responsibly: accountability in a new era. stanford, ca: stanford university press, p. 38. figure 1. an example of a cla performance task. 6.2.2 cognitive labs the purpose of cognitive labs is to study the cognitive processes that students use when they complete different tasks. students are asked to report their thoughts verbally as they carry out a task (see johnstone, bottsford-miller & thompson, 2006). in this study cognitive labs were divided into three parts: (1) instruction and training, where the researcher explained what the cognitive lab was about and trained the students to think aloud with a short warm-up task; (2) ‘think-aloud’, where the students talked aloud while completing the cla performance task; and (3) a follow-up interview. the cognitive lab for each student was video-recorded and lasted two to three hours. to ensure the consistency of cognitive labs, a script of directions and the same training task and the interview questions for each student were used. the videos h.hyytnen et al. 9 | f l r were recorded with two cameras and a table microphone. the cognitive workshop produced the following materials: video data, content logs (see below), written test answers and transcribed interview data. the neutral type of think-aloud protocol conducted by ericsson and simon (1993) in which students were not interrupted while they were performing a task was used in this study. the think-aloud method makes it possible to collect data about a student’s ongoing thinking processes whilst he or she is working on a task (ericsson & simon, 1993; cotton & gresty, 2006; van someren, barnard & sandberg, 1994). we assume that students’ ‘knowing-that’ attributions (e.g. ‘scientific knowledge is true’) may reflect their epistemological orientations and reveal their criteria for justifying beliefs (see scheffler, 1965). moreover, in some cases the think-aloud method makes it possible to explore critical thinking in action, especially in situations that simulate real-world circumstances. immediately after the task was performed, a follow-up interview was conducted. the aim of the interview was to gain more detailed information about the processes and knowledge that the students used to complete the task and to probe students’ beliefs about knowledge and knowing. for example, the students were asked questions about how they dealt with conflicting information, how they decided which information to use, what sources of information in documents from the documents library they trusted and why, and how they usually evaluate knowledge. 7. data analysis the data were analysed using a qualitative thematic analysis with an abductive approach (timmermans & tavory, 2012; haig 2005). an abductive strategy means that the themes identified from the data were linked to the theoretical understanding based on previous studies. abduction is a process that combines things which one had not previously associated by creating a new interpretation, that is, the relationship of a new combination of study features (timmermans & tavory, 2012). hence, the analysis process was nonlinear, moving back and forward amongst all the data, data items, analysed qualities and understanding of the phenomenon based on prior studies. the first and fifth authors were responsible for the analysis, but the final results were obtained through a thorough discussion with all authors. the data were processed in such a way that the participants could not be identified. the analysis included four phases (figure 2) that represented the unique combination of datagrounded and theory-driven phases, as well as phenomenon and individual-level analyses. during the first phase, video recordings were initially indexed with the elan program, which allows the addition of as many tiers and annotations on the video stream as needed (see lausberg & sloetjes, 2009; max planck institute for psycholinguistics, 2012). the purpose of indexing was to make the large video data set easier to handle. in this study the indexing tiers corresponded to the parts of cognitive labs including training, thinkaloud methods and interviews. in addition, students’ interviews from the videos were transcribed. h.hyytnen et al. 10 | f l r *based on braun & clarke (2006, p. 87). figure 2. a visualisation of the analysis process. after the indexing, content logs were created for each video in which accurate descriptions and summaries of events were systematically recorded. transcriptions of relevant sections of verbalisations of students’ critical thinking and epistemological beliefs (e.g. whenever a student evaluated the quality and reliability of the information in a document or where a student reached a conclusion based on her or his analysis) and nonverbal acts (e.g. a student did not read in detail or skipped over the document) were also included in the log (cf. table 1). the second phase of the analysis was the data coding (see table 2 for definitions). this phase was theory-driven, meaning that the features guiding the coding were based on prior studies (see table 1). the coding focused on the following qualities: the process by which the student approached the task and solved h.hyytnen et al. 11 | f l r the problem, the knowledge that the student used to carry out the task, the critical thinking exhibited, and epistemological beliefs. these different qualities were coded systematically across the entire data set and within the data items such as the transcribed interviews and the think-aloud videos of each person. by this means, all the data items from one student, including the video data, content log, written test answers and transcribed interviews, were coded and analysed separately, after which data from all students were combined and compared (see table 3 for an example of the codes). all extracts were labelled with a student code (s1-s10) and a method code (i= interview, t=think aloud, w= written test answer). the data examples were translated into english. table 2. data sources and focal points of coding data sources coding features video data, content logs, transcribed interviews 1. the process: how does the student approach the task and solve the problem? video data, students’ written answers, content logs, transcribed interviews 2. what knowledge/information does the student use to solve the task? 2.1 what kind of knowledge/information did the student use? 2.2 why? 2.3 how does the student use that knowledge/information? video data, students’ written answers, content logs, transcribed interviews 3. critical thinking 3.1 how does the student identify, analyse and evaluate information, ideas and arguments? 3.2 how does the student judge the acceptability (especially the credibility) of documents? 3.3 how does the student interpret data/ graphs/ maps? 3.4 how does the student recognise the relationship between assumptions? 3.5 how does the student evaluate background information? 3.6 how does the student make a decision? 3.7 how does the student identify reasons and come to a conclusion? 3.8 how does the student produce explanations and arguments? h.hyytnen et al. 12 | f l r video data, students’ written answers, content logs, transcribed interviews 4. epistemological beliefs 4.1 what does the student think about knowledge, knowing and the credibility of knowledge? 4.2 how does the student determine the trustworthiness, acceptability and justification of different types of information? 4.3 how does the student describe herself or himself as a knower? table 3. an example of codes data extract coded for you could consider this a good argument; the expert has gone [to the place where events took place] to see for himself (s9t) 4.1 what does the student think about knowledge, knowing and the credibility of knowledge? 4.2 how does the student determine the trustworthiness, acceptability and justification of the different types of information? yeah, i don’t believe the chair of the stakeholder group] is completely off the mark either. [reliability] is just always case-specific. (s8i) 4.1 what does the student think about knowledge, knowing and the credibility of knowledge? this just seems scientific somehow. (s6t) 4.1 what does the student think about knowledge, knowing and the credibility of knowledge? in the third phase the codes and coded extracts were grouped under potential themes, and all the relevant data were gathered under each theme (braun & clarke, 2006). we identified a variety of preliminary themes on the basis of the codes. during the analysis, the preliminary themes were defined and combined several times. in the end two main themes and two subthemes remained (see figure 2). the final themes were refined, labelled and cross-checked to see if they worked in relation to the coded extracts and the entire data set. the focus of the thematic analysis was the variation of study features on the phenomenon level. after completing the thematic analysis, we found that the students could be placed in different profiles based on our themes as well as on patterns of behaviour and cognition observed. this phase focused on the variation of study features at the individual level. thereafter, we conducted final descriptions, interpretations and revisions of the results. the results of thematic analysis show how critical thinking and epistemological beliefs manifested themselves in this particular group of students, whereas the student profiles describe how these phenomena vary across individuals. 8. results in the thematic analysis two main themes were identified: (1) flexibility in critical thinking and (2) variation in critical thinking and epistemological beliefs. the two themes emerged from exploring the students’ critical thinking from different perspectives. the ways in which the themes were related differed amongst the participants, which further allowed us to identify student profiles. we identified two main profiles, and on the basis of their characteristic features we labelled them as (1) thorough processing and (2) h.hyytnen et al. 13 | f l r superficial processing. the results are described using a combination of identified themes and student profiles. 8.1 flexibility in critical thinking students showed various skills in their ability to adapt their thinking and their performance flexibility to the demands of the task. there was clear variation in the students’ ability to change their actions or ways of critical thinking, in which we identified both rigidity and flexibility. flexibility meant that the students could modify their actions and processes and change their behaviours as needed, whereas rigidity refers to situations in which students could not change their processes or look at things from a new perspective or adjust to new evidence in a problem-solving situation. students who were able to make changes in their actions showed open-mindedness and an inquiring attitude. in the following extract, one student describes how he adjusted his performance and ended up analysing and interpreting the documents correctly: i approached this assignment maybe a little too much as if i had simply copied what they say here in these papers and put them down in my answer. but then when i started thinking, like about my own views on the topics, then right off in [question] number one, it took me a really long time to answer this question. (s8i) on the other hand, there were students who could not adjust their thinking or performance. some of these students said that they always act in the same way: well, i’m always like this time-management catastrophe. like in exams and everything, especially exams, it always feels like i run out of time. and in general i notice that in all comprehension and analysis assignments and things like that, they always take me a really long time. (s5i) 8.2 variation in critical thinking and epistemological beliefs students showed various aims in the problem-solving situation. some students tried to understand the complex situation, whereas others tried to find the right answer to the problem. students also varied in their critical thinking, including (a) their disposition and ability to identify, analyse, evaluate and interpret information; (b) their ideas and arguments in judging the acceptability of documents; (c) their abilities to recognise relationships between assumptions; (d) their abilities to make a reasoned decision; and (e) their abilities to produce explanations and arguments. in addition, students’ epistemological beliefs varied. some students claimed that only through scientific knowledge we can arrive at truth. however, other students expressed the idea that both objective and subjective knowledge can hold the highest epistemic status. we found that critical thinking emerged as a tool for understanding knowledge and determining the goodness and reliability of knowledge; thus, students’ epistemological beliefs were interwoven into their critical thinking. within this theme we found that students used critical thinking either a) as a tool for enhancing understanding or b) as a tool for determining truth or falsehood. based on this difference, students could be classified in one of two qualitative student profiles, either (1) thorough processing or (2) superficial processing. the profiles captured the diversity of the students’ abilities and dispositions to think critically. in addition, these two profiles characterised the variation in how students viewed the nature and limitations of knowledge and knowing, and especially in how they determined what is needed to evaluate knowledge as true or justified and how they acquired and used the knowledge in the problem-solving situation (see table 4). the phrase ‘acquiring knowledge’ here emphasises the dominant way that students used to obtain knowledge in a problem-solving situation. the students classified in the profile called ‘thorough processing’ demonstrated an ability to carry out a deep processing of the content of the documents. these students saw knowledge as fallible and contextual. similarly, the students in the profile called ‘superficial processing’ expressed the idea that knowledge is fallible, yet they did not consider the contextual nature of knowledge at all. in the problem h.hyytnen et al. 14 | f l r solving situation they did make a serious effort to analyse, interpret or synthesise the information in the materials. the thorough processing profile is further divided into two sub-profiles: (1a) reasoning in order to reach conclusions and (1b) intuition. likewise, the second profile, ‘superficial processing’, also consisted of two sub-profiles: (2a) referring to an argument made by authoritative specialists or experts and (2b) trust in scientific method and proof. we describe the characteristics of the profiles and sub-profiles below and provide details pertaining to variation in academic thinking. table 4. the nature of knowledge and acquiring knowledge in two qualitatively different student profiles of critical thinking sub-theme student profile epistemological beliefs acquiring knowledge (sub-profile) critical thinking as a tool for enhancing understanding thorough processing both objective and subjective knowledge can hold the highest epistemic status. knowledge is fallible, relative and contextual. reasoning in order to reach a conclusion intuition critical thinking as a tool for determining truth or falsehood superficial processing knowledge may reach truth only if it is produced by a reliable process, that is, using empirical methods. objective knowledge holds the highest possible epistemic status, but is fallible. some theories may be false. referring to an argument made by authoritative specialists/experts trust in scientific method and proof 8.3 profile 1: thorough processing students the students (n=5) who deeply analysed the content of the documents created their own understanding of the problem-solving situation. for them, critical-thinking skills were tools to deepen and enhance understanding. these students believed that theories and beliefs could be understood in relation to some context, as the following extract shows: yeah, i don’t believe the chair of the stakeholder group] is completely off the mark either. [the reliability of] knowledge is just always context-specific. (s8i) these students considered it possible to improve current theories and beliefs. these students were thus open to new evidence that could disprove a previously-held position or belief. for them, scientific knowledge is probably reliable. they believed that both objective and subjective knowledge could attain the highest epistemic status, meaning that subjective perceptions (e.g. their own perceptions or information obtained from someone else) could also be reliable. these students thought that the credibility of knowledge could be affected by vested interests or bias, for example. although these students emphasised their own role in constructing knowledge, they did not believe that all knowledge is constructed or generated by human minds. from the epistemological perspective, these students took the fallibilist position. h.hyytnen et al. 15 | f l r the students who belonged to the ‘thorough processing’ profile were further divided into two subprofiles on the basis of how they acquired knowledge and reached conclusions in the problem-solving situation and how flexible they were in changing their actions or ways of thinking. the first sub-profile was called (1a) reasoning in order to reach conclusions and the second was called (1b) intuition. 8.3.1 reasoning in order to reach conclusions two students endeavoured to reach conclusions by reasoning. these students analysed connections across the information presented in the different documents. they also clarified and interpreted different claims and ideas that were presented in the documents. on the basis of their own analyses, they synthesised information, reached a clear decision or conclusion, provided arguments for their decision and explained why this decision was the best in light of all the issues brought up in the documents. in the following example one student describes her analytical process: ‘somehow i knew how to read beyond the documents’ (s4i). these students were also able to adjust their thinking in line with new evidence and make changes in their actions. these students justified conclusions with good reasons (e.g. reliable and valid evidence) and considered themselves as active and responsible knowers, as the following extracts show: but maybe i wouldn’t, like, start criticising right away; somehow, i’d have to start looking into, you know, on what basis they arrived at these figures. (s10i) for instance, using this graph is fine, but i think it’s been, you know, clearly misinterpreted here in the text. (s4i) these students created their own understandings of the situation on the basis of their analyses. they used the materials for the analysis and evaluation process in a way that went beyond the obvious. for example, they identified, analysed, evaluated and interpreted all the major facts and ideas presented in the documents. they consciously excluded some information in the documents because of contradictory evidence. in addition, they were able to distinguish relevant claims from irrelevant ones. these students also judged the reliability of the documents, evaluated presuppositions and analysed connections between claims. furthermore, they produced different explanations, identified reasons, produced arguments and drew inferences. these students further identified and used several criteria in evaluating reliability: corroborating claims from different sources, evaluating the context in which the claim was made, exploring who interpreted the data and evaluating the presuppositions. moreover, these students considered the ethical aspects of knowledge: knowledge and information shape human beings’ worldview. i’ve just gotten the impression about newspapers, about the media too, that it somehow has the effect that the opinions [presented in them] are so strong that maybe you don’t analyse it so clearly. so like even helsingin sanomat [finland’s largest daily] really has, somehow it seems that they have a pretty strong, you know, bias... you know that even if it’s neutral in a way, then the fact the issues they raise in it, in a way that already affects what information is raised, and what... that it like really powerfully shapes people’s worldview. (s4i) 8.3.2 intuition three students justified conclusions by intuition. these students created their own understanding of a situation. however, they did not select materials or question any information: they used all the information in the documents, such as empirical knowledge, expert opinions, reports, maps, experiences of an inhabitant, recommendations, letters and second-hand knowledge. these students acquired knowledge in a rather uncritical way. they rarely evaluated the reliability of documents. indeed, these students did not have clear criteria for evaluating the reliability or relevance of information. they just trusted their intuition: this just seems scientific somehow. (s6t) h.hyytnen et al. 16 | f l r i don’t know how i should formulate this, but i’ll start by saying that when i read, for instance... or when i’m taking classes, i don’t spend a whole lot of time wondering if some piece of information is reliable or not. (s6i) these students started to analyse and interpret thoroughly all information presented in the documents. they identified all major facts and ideas. they also considered different decisions or explanations, but could not explain what decision was the best or why. there were too many options available. because the students did not reach clear conclusions, they did not present any arguments for accepting the conclusion either. these students showed an inability to adjust their thinking to new evidence or make changes in their actions. 8.4 profile 2: superficial processing students common to all students in the second main profile was that they processed the materials in the problem-solving task superficially: they did not make a serious effort to analyse, interpret or synthesise the information in the materials. this profile consisted of five students who used critical thinking as an instrument for determining truth or falsehood. their goal was to find the right answer to the problem. in contrast to the ‘thorough processing’ profile, students in this profile believed that knowledge is trustworthy only if it was produced through a reliable process, for example, by using empirical methods or consulting suitable experts. for these students, scientific and verified knowledge is the most reliable, because that kind of knowledge is based on evidence, and it is unbiased and objective. the students believed that subjective knowledge is predominantly untrustworthy. however, these students considered empirical knowledge (which holds the highest epistemic status) to be fallible too, not absolutely certain. they believed that some theories might be false and that it is possible to improve current conceptions and theories. from the epistemological perspective, these students also took the fallibilist position. the analysis indicated varying problems in critical thinking, such as problems in evaluating information, reasoning and reaching conclusions. some of these students also had little motivation to think critically. characteristic of the students in this profile was that they focused on isolated details. they took knowledge for granted. in other words, they accepted knowledge (particularly scientific knowledge) as true without question. these students were further divided into two sub-profiles according to how they acquired knowledge in a problem-solving situation, trusting either (2a) an argument by authoritative specialists or experts and (2b) verified empirical evidence or testimony. 8.4.1 referring to an argument by authoritative specialists or experts two students were categorised in this sub-profile. these students trusted authorities in acquiring knowledge. they saw themselves as uncertain knowers. these students believed that if a person who is said to be an authority on something makes an argument about that something, then the argument should be trustworthy and therefore usable. the right answers can be reached by consulting the right expert. these students repeated arguments and conclusions as these were presented in the documents. they drew on empirical knowledge and expert opinions, that is, arguments from authoritative sources. these students had difficulties in evaluating information. they focused on details and took in all the information they were presented without question. they picked up isolated and obvious details from the materials for each question. the students did not properly analyse, evaluate or interpret the information presented in the documents; they just jumped to conclusions. they disregarded and seriously misinterpreted important information. they also had problems in reasoning and reaching a conclusion. in order to make decisions or arguments, these students reproduced lists of isolated details from documents. they did not provide any reasons or explanations for their decisions. moreover, they did not identify alternative solutions. these students presented some unreliable claims as being credible. in the interview one student representing this sub-profile said that she has had similar problems in learning: h.hyytnen et al. 17 | f l r creative comprehension and, like, reaching a synthesis of overall concepts is really challenging for me. like, for instance, it’s really hard to study for exams, because i’d be more than happy to read the book, but then i don’t really grasp the key message and structure that it’s trying to communicate. acquiring data independently and, like, learning information that way is challenging. so, for instance, i haven’t done all my exams. i haven’t done them because, i’ve tried to start [studying for] them lots of times, but then some, how would you put it, if listening is auditory, then learning from text is pretty hard for me. this third year, which is currently underway, has been, like, really hard. i’ve really haven’t gotten many credits. i don’t feel i’ve accumulated the amount of information i should have or could have in three years. that the pieces of information are discrete and still pretty scattered in my head at the moment. (s5i) these students expressed the view that knowledge is always uncertain, but they did not consider themselves capable of evaluating knowledge. these students named a few external criteria for evaluating knowledge (such as an authority, expert opinion, publication, openness, journal citations). however, in practice they did not know how to use these criteria independently. both of these students gave authoritative experts the responsibility for evaluating knowledge, as the following example demonstrates: i don’t know what the right approach is in order to grasp those overall concepts from that huge mass of teensy-weensy details. because a candidate has to read a huge number of articles to find the ones that are, like, related to one’s own topic and all. so it’s really hard when you’re, like, reading an article to judge why this one might be better than that one. so. but i got a tip from my supervisor that i should pay attention to the reliability of the journal. to be honest, it’s the research articles, the ones we have at the university, that are actually the only ones we’re told we can cite. and then it’s like... they’re easy to evaluate based on which publications are more credible. and on the web on [sic] science, they have this one like... what is it, like an indicator that they have, just based on the number of citations and other factors, of the accuracy of the research data… it’s hard! in a way, to make that distinction between what’s true and what isn’t. at least i don’t have the know-how to say what’s true and what isn’t. (s5i) both students expressed the view that in a real-life situation they would seek help from other people, such as authoritative specialists (e.g. a university teacher) or other students. for example, one student representing this sub-profile said several times that she needed co-operation with other students to solve the task: i haven’t really had to do anything like this before. that it’s pretty hard in a way. there are so many points and, you know, perspectives here. i haven’t even had to think about stuff like this at the university, then it’s really like new for me, or you know. the assignment was pretty difficult. this might have been more interesting as a group assignment. like there would have been, you know, interaction, and then maybe it would have generated more thoughts somehow. (s7i) 8.4.2 trust in scientific method and proof the three students comprising this sub-profile were very critical. they all selected documents roughly based on empirical evidence, excluding more than half of the documents provided. these students were aware of their own behaviour: i eliminate some of the documents right away, for instance, email exchanges and letters, because the people haven’t investigated the matter; the text was written based on a gut feeling. (s3t) these students expressed the view that scientifically and empirically verified knowledge is the most reliable. they knew that corroboration from other reliable and related sources improves credibility. in the problem-solving situation they only trusted and used arguments by scientific authorities. these students described themselves as ‘error seekers’ in the interviews. the following examples illustrate the view of the students in this sub-group: h.hyytnen et al. 18 | f l r i trust exam books and articles a lot, yes. the difference between the two is that books can often, you know, be unreliable. plus the fact that, at least when they’re academic, it has a lot to do with when they were written, because things move so fast. that i have this one book for my thesis that i was just looking at, it’s got tons of mistakes. so, like, you just have to find them yourself. but with articles, probably those, and then of course depending on the journal. that maybe some article in science: i consider them pretty reliable. nowadays, i’m a little too sceptical about all kinds of things. i question a lot more these days than i used to. (s1i) you can get the first impression of reliability, of course, from the kind of source it was published in. in other words, i wouldn’t swallow some iltalehti [a finnish tabloid] headline on some scientific subject without thinking it over properly first. but having a reference to those academic publications, and as far as how i’ve drawn those conclusions myself after having read the article, not based on some newspaper headline, then that would be at least important in terms of first impressions. and… well, even if you read a scientific article, if it doesn’t agree at all with what you’ve learned about the topic earlier, then of course you’d have good cause to suspect those research results quite a bit. but the source is what i’d probably consider as the main thing. (s3i) although these students describe themselves as critical, they did not evaluate information from reliable sources in the problem-solving situation. for example, they did not recognise that two sources, which included empirical or verified knowledge, were biased. they analysed and interpreted information superficially and focused on isolated details. they did not interpret the documents they selected nor did they consider presuppositions. in order to draw conclusions these students mainly reproduced details from the documents. they did not identify alternative solutions or conclusions or approaches to the problem. nor did they provide any reasons or explanations for their own conclusions. they identified only a few claims that were presented in the documents and disregarded many relevant aspects of those claims. as a result they had problems reaching a conclusion. in the following example one student described the situation as follows: ‘at least i wouldn’t draw any conclusions based on those [documents]’ (s1t). all these students thought that there was one definite answer to the problem. one student in this sub-profile emphasised that she does not have any disposition or motivation to express reasons for or against some idea in a test situation or in everyday life: in everyday life it’s rare that, if you’re discussing something, it’s rare that anything like this happens. or i never, really rarely discuss anything argumentatively in any way. in real life i simply don’t like it, discussing issues. (s9i) 8.5 summary of the results figure 3 combines the two main themes in order to form a comprehensive picture of participants’ critical thinking. students who had several problems in critical thinking, yet had flexibility coped with the demands of the task. for example, two students had problems evaluating documents and did not form a general picture of the situation presented in the documents. because the students were struggling with the demands of the task, they selected documents and reproduced arguments and conclusions just as these were presented in the materials. eventually, the students reached a limited conclusion. on the other hand, there were students who were skilled in specific critical-thinking skills, such as analysing and interpreting information, but lacked other abilities, such as evaluating conflicting claims or producing explanations. these students could neither reach a conclusion nor were they able to determine the weaknesses of alternative solutions. in addition, these students were unable to change their actions or thinking; for example, they were not flexible in time management. these students somehow ‘over-analysed’ the problem, and, in the end, they failed in the problem-solving process. h.hyytnen et al. 19 | f l r in sum, the aspect that distinguished the participants were the differences in 1) aims, 2) the skills and disposition to think critically, 3) epistemological beliefs, 4) acquiring knowledge and 5) the skill of flexibility in adapting thinking and performance to the demands of the task. figure 3. summary of results. 9. conclusions f le x ib il it y i n c ri ti ca l th in k in g critical thinking as a tool for enhancing understanding generating personal understanding through ‘thorough processing’ epistemological beliefs: both objective and subjective knowledge can hold the highest epistemic status. knowledge is fallible and contextual. r ig id ity in critica l th in k in g reaching a well-reasoned solution + figuring out how to complete multidimensional tasks and planning action + defining the problem, evaluating, analysing, interpreting information, identifying alternative reasons, considering relationships between assumptions and ultimately reaching a reasoned conclusion. endless weighing of the different options + defining the problem, identifying ideas, analysing, and interpreting information problems in time management, decision-making, reaching reasoned conclusions, evaluating knowledge and judging the acceptability of information. problems in producing explanations reaching a limited solution identifying only a few ideas problems in evaluating, analysing and interpreting information problems in decision-making and reaching conclusions + searching for alternative ways to complete a task, changing one’s own routines or seeking help from authoritative specialists problems in reaching a conclusion identifying only a few ideas problems in evaluating, analysing and interpreting information problems in decision-making, reaching conclusions and producing explanations expectation that a problem has a definite, right answer disposition to think critically may be low critical thinking as a tool for determining truth or falsehood seeking the right answer through ‘superficial processing’ epistemological beliefs: objective knowledge can reach the highest possible epistemic status. knowledge is fallible. h.hyytnen et al. 20 | f l r even though the number of participants in this study was small, the variety in the students’ critical thinking was evident. our results showed that after three years of university study, students’ critical-thinking skills and epistemological beliefs differed greatly, and eight out of ten volunteer students had some problems in critical thinking (cf. arum & roksa, 2011; pascarella, blaich, martin & hanson, 2011). the multi-method approach effectively revealed the variety of problems that university students may encounter. while many problems were related to the lack of disposition or skill, such as an inability to evaluate the credibility of documents, examine presuppositions, make interpretations, develop a personal perspective or generate arguments or conclusions, some of the problems were related to an inability to modify the whole criticalthinking process in a flexible manner. these findings corroborate the ideas of fisher (2011) and scheffler (1965), who suggested that individuals may be more or less skilled at different critical thinking abilities. in other words, a student may have the ability to identify and evaluate information, for example, yet at the same time struggle with other abilities, such as arriving at a conclusion, adjudicating conflicting claims or producing arguments. therefore, it is clear that unilateral instructions concerning critical thinking are difficult to provide. the findings of this study support the idea that students’ epistemological beliefs were interwoven into their critical thinking (cf. kuhn, 2005; kuhn & weinstock, 2002). critical thinking emerged as a tool for understanding and determining the relevance and reliability of knowledge. students who showed superficial processing believed that objective knowledge (i.e. scientific and verified knowledge) has the highest possible epistemic status. although it is sensible to trust in scientific and empirical knowledge more than in personal opinions, the problem was that these students accepted scientific knowledge without question: they did not analyse, evaluate or interpret the information contained in the documents they were given. they acquired knowledge by appealing in equal measure to authoritative opinion, trusting in verified empirical evidence and listening to testimonies. these students palmed off a justification for knowing on authoritative experts. contrary to the results of many previous studies (e.g. kaartinen-koutaniemi & lindblom-ylänne, 2012; king & kitchener, 2004; kuhn, 1999, 2005; kuhn & weinstock, 2002), our main finding was that the students who appealed to authorities, testimonies or empirical evidence did not believe that knowledge is absolutely certain or unquestionable. nor did these students share the view that beliefs accurately represent or correspond to reality. in effect, the students did not share a sense of metaphysical realism. instead, these students claimed that scientific theories are uncertain, but probably true. the findings also show that the students who believed that knowledge is contextual and relative did not share a relativist view of knowledge. this finding is also contrary to the findings of earlier studies (e.g. lahtinen & pehkonen, 2012). conversely, all of the students saw knowledge as fallible. the students believed that it is possible to seek criteria for evaluating, comparing and justifying beliefs or theories. although some students struggled with evaluating knowledge, all of them saw current conceptions and theories as a starting point for further inquiry. they were thus fallibilist in the epistemological sense. this study further shows that students’ belief in themselves as critical thinkers and knowers is not necessarily equivalent to how they perform. thus, we assume, along with previous studies (elby & hammer, 2001; greene & yu, 2014), that the self-reported assessment method is not enough to gauge these kinds of complex processes. the present small-scale qualitative study has provided a unique picture of the critical thinking and personal epistemological beliefs of ten third-year bioscience students. furthermore, this study has educational significance by revealing problems in these students’ critical-thinking skills and by describing the role of students’ conception of knowledge in the process of thinking critically. through a multifaceted approach, it was also possible to deepen understanding of the emphases and gaps in the prevailing empirical research on critical thinking and personal epistemology. however, the findings of this study should not be interpreted as an accurate prediction of the target population. the findings of this study rather illustrate the nature of the phenomenon being studied, and how the different aspects of critical thinking and epistemological beliefs are intertwined and contribute to it together. this study involved a small, homogeneous sample of students in one discipline only. owing to these limitations, more communication between the theoretical, empirical and methodological perspectives is required to increase understanding of this complex phenomenon in the different spheres. h.hyytnen et al. 21 | f l r keypoints in this exploratory study we provide a multidimensional framework for analysing critical thinking by combining theoretical aspects from philosophical, educational and psychological approaches. in this exploratory study we provide a multidimensional framework for analysing critical thinking by combining theoretical aspects from philosophical, educational and psychological approaches. a large body of data for each participant (n=10) was collected using multiple methods, including think-aloud protocol, interviews and a collegiate learning assessment (cla) performance task. the result shows that students’ epistemological beliefs were interwoven into their critical thinking: students used critical thinking as a tool for enhancing understanding and for seeking a right answer none of students shared an absolutist view of knowledge. none of students shared a relativist view of knowledge. all students shared a fallibilist view of knowledge. acknowledgements the first author was financially supported by a scholarship from the alfred kordelin foundation. the authors are grateful to their colleagues for helpful comments on an earlier version of this article in manuscript, as well as to the students for their participation. we are grateful to roger benjamin at the council for aid to education for permitting us to use a performance task from the collegiate learning assessment and to viivi virtanen (phd) and mikael kivelä (ma) for helping us in data collection. references the american philosophical association. (1990). critical thinking: a statement of expert consensus for purposes of educational assessment and instruction "the delphi report". committee on pre-college philosophy. millbrae, ca: the california academic press. angeli, c., & valanides, n. (2009). instructional effects on critical thinking: performance on ill-defined issues. learning and instruction, 19, 322334. arum, r., & roksa, j. (2011a). limited learning on college campuses. society, 48, 203207. doi 10.007/s12115-011-9417-8 arum, r., & roksa, j. (2011b). academically adrift. limited learning on college campuses. chicago: the university of chicago press. the australian council for educational research. (2001). graduate skills assessment. summary report. 01/e occasional paper series. higher education division. retrieved from http://www.acer.edu.au/documents/gsa_summaryreport.pdf baartman, l. k. j., bastiaens, t. j., kirschner, p. a., & van der vleuten, c. p. m. (2007). evaluating assessment quality in a competence-based education: a qualitative comparison of two frameworks. educational research review 2, 114129. bailin, s., & siegel, h. (2003). critical thinking. in n. blake, p. smeyers, r. smith, & p. standish (eds.), the blackwell guide to the philosophy of education (pp. 181193). blackwell publishing. h.hyytnen et al. 22 | f l r barzilai, s., & zohar, a. (2012). epistemic thinking in action: evaluating and integrating online sources. cognition and instruction 30, 3985. bleazby, j. (2011). overcoming relativism and absolutism: dewey’s ideals of truth and meaning in philosophy for children. educational philosophy and theory 43, 453466. doi: 10.1111/j.14695812.2009.00567.x bok, d. (2006). our underachieving colleges. a candid look at how much students learn and why they should be learning more. princeton, nj: princeton university press. braun, v., & clarke, v. (2006). using thematic analysis in psychology. qualitative research in psychology 3, 77101. cotton, d., & gresty, k. (2006). reflecting on the think-aloud method for evaluating e-learning. british journal of educational technology 37, 4554. dewey, j. (1910). how we think. boston: d.c. heath & co. dierick, s., & dochy, f. (2001). new lines in edumetrics: new form of assessment lead to new assessment criteria. studies in educational evaluation 27, 307329. elby, a., & hammer, d. (2001). on the substance of a sophisticated epistemology. science education, 554567. ennis, r. (1991). critical thinking: a streamlined conception. teaching philosophy 14, 524. ericsson, k. a., & simon, h. a. (1993). protocol analysis: verbal reports as data. cambridge, ma: mit press. everitt, n., & fisher, a. (1995). modern epistemology: a new introduction. new york, mcgraw-hill. fenstermacher, g. d. (1994). the knower and the known: the nature of knowledge in research on teaching. review of research in education 20, 356. ferguson, l. e., & bråten, i. (2012). students’ profiles of knowledge and epistemic beliefs: changes and relations to multiple-text comprehension. learning and instruction 25, 4961. fisher, a., & scriven, m. (1997). critical thinking: its definition and assessment. edgepress and centre for research in critical thinking. university of east anglia. fisher, a. (2011). critical thinking: an introduction. second edition. cambridge: cambridge university press. green, j. a., & yu, s. b. (2014). modelling and measuring epistemic cognition: a qualitative reinvestigation. contemporary educational psychology 39, 1228. haig, b. d. (2005). abductive theory of scientific method. psychological methods 10, 371388. halpern, d. f. (2014) thought and knowledge. fifth edition. ny: psychology press. hammer, d., & elby, a. (2003). tapping epistemological resources for learning physics. the journal of the learning sciences 12 (1), 5390. heijltjes, a., van gog, t., leppink, j., & paas, f. (2014). improving critical thinking: effects of dispositions and instructions on economics students’ reasoning skills. learning and instruction 29, 3142. hofer, b. k. & pintrich, p. r. (2002). personal epistemology: the psychology of beliefs about knowledge and knowing. new jersey: lawrence erlbaum associates. hofer, b. k. (2004). epistemological understanding as a metacognitive process: thinking aloud during online searching. educational psychologist 39, 4355. hofer, b. k. (2005). the legacy and the challenges: paul pintrich’s contributions to personal epistemology research. educational psychologist 40, 95105. hofer, b. k. (2006). domain specificity of personal epistemology: resolved questions, persistent issues, new models. international journal of educational research 45, 8595. holma, k. (2004). plurealism and education: israel scheffler’s synthesis and its presumable educational implications. educational theory 54, 419430. holma, k. (2012). fallibilist pluralism and education for shared citizenship. educational theory 62, 397409. holma, k. (2014). the critical spirit: emotional and moral dimensions critical thinking. under review. holma, k., & hyytinen, h. (equal contribution) (2014). the philosophy of personal epistemology. under review. h.hyytnen et al. 23 | f l r ichikawa, j. j., & steup, m. (2012). the analysis of knowledge. in edward n. zalta (ed.), the stanford encyclopedia of philosophy (winter 2012 edition). retrieved from http://plato.stanford.edu/archives/win2012/entries/knowledge-analysis/>. johnson, r. b., onwuegbuzie, a. j., & turner, l. a. (2007). toward a definition of mixed methods research. journal of mixed methods research 1, 112133. johnstone, c. j., bottsford-miller, n. a., & thompson, s. j. (2006) using the think aloud method (cognitive labs) to evaluate test design for students with disabilities and english language learners. technical report 44. minneapolis, mn: university of minnesota, national centre on educational outcomes. retrieved from http://www.cehd.umn.edu/nceo/onlinepubs/tech44/ kaartinen-koutaniemi, m., & lindblom-ylänne, s. (2008) personal epistemology of psychology, theology and pharmacy students: a comparative study. studies in higher education 33, 179191. kaartinen-koutaniemi, m., & lindblom-ylänne, s. (2012). personal epistemology of university students: individual profiles. education research international 2012, 18. doi:10.1155/2012/807645 king, p. m., & kitchener, k. s. (1994). developing reflective judgment: understanding and promoting intellectual growth and critical thinking in adolescents and adults. san francisco: jossey-bass. king, p. m., & kitchener, k.s. (2004). reflective judgment: theory and research on the development of epistemic assumptions trough adulthood. educational psychologist 39, 518. klein, s., benjamin, r., shavelson, r., & bolus, r. (2007). the collegiate learning assessment, facts and fantasies. evaluation review 31, 415439. doi: 10.1177/0193841x07303318. klein, s., freedman, d., shavelson, r., & bolus, r. (2008). assessing school effectiveness. evaluation review 32, 511525. kuhn, d. (1999). a developmental model of critical thinking. educational researcher, 28, 1625. kuhn, d. (2005). education for thinking. harvard university press. kuhn, d., & weinstock m. (2002). what is epistemological thinking and why does it matter? in barbara k. hofer & paul, r. pintrich (eds.) personal epistemology: the psychology of beliefs about knowledge and knowing, 121144. new jersey: lawrence erlbaum associates. lahtinen, a-m, & pehkonen, l. (2013). ‘seeing things in a new light’: conditions for changes in the epistemological beliefs of university students. journal of further and higher education, 37, 397415. lausberg, h., & sloetjes, h. (2009). coding gestural behavior with the neuroges-elan system. behavior research methods, instruments, & computers, 41(3), 841849. retrieved from http://www.springerlink.com/content/d53722q3k3314374/ maclellan, e. (2004). how convincing is alternative assessment for use in higher education. assessment & evaluation in higher education 29, 311321. markowitsch, j., & messerer, k. (2007). practice-oriented methods in teaching and learning in higher education. in p. tynjälä, j. välimaa, & g. boulton-lewis (eds.), higher education and working life: collaborations, confrontations and challenges (pp. 177194). oxford: elsevier. max planck institute for psycholinguistics. (2012). elan. [software]. available from http://tla.mpi.nl/tools/tla-tools/elan/ moore, t. (2004). the critical thinking debate: how general are general thinking skills. higher education research & development 23, 318. moore, t. (2013). critical thinking: seven definitions in search of a concept. studies in higher education, 38, 506522. muis, k. r., bendixen, l. d., & haerle, f. c. (2006). domain-generality and domain-specificity in personal. epistemology research: philosophical and empirical reflections in the development of a theoretical framework. educational psychology review 18, 3–54. niiniluoto, i. (1999). critical scientific realism. oxford: oxford university press. pascarella, e. t., blaich, c., martin, g. l., & hanson, j. m. (2011) how robust are the findings of academically adrift? change: the magazine of higher learning 43, 20–24. peirce, c. s. (1934). pragmatism and pragmaticism, vol. 5 of collected papers of charles sanders peirce. cambridge, massachusetts: harvard university press. h.hyytnen et al. 24 | f l r perry, w. g. (1970). forms of intellectual and ethical development in the college years. harvard university. pintrich, p. r. (2002). future challenges and directions for theory and research on personal epistemology. in b. k. hofer & p. pintrich (eds.), personal epistemology: the psychology of beliefs about knowledge and knowing (pp. 121144). new jersey: lawrence erlbaum associates. putnam, h. (1981). reason, truth and history. cambridge: cambridge press. reed, b. (2002). how to think about fallibilism. philosophical studies 107, 143157. ryle, g. (1949). the concept of mind. london: hutchinson’s university library. scheffler, i. (1965). conditions of knowledge. an introduction to epistemology and education. glenview: scott, foresman and company. shavelson, r.j. (2010). measuring college learning responsibly: accountability in a new era. stanford, ca: stanford university press. shope, r. k. (2004). the analysis of knowing. in i. niiniluoto, m. sintonen, & j. woleński (eds.) handbook of epistemology (pp. 356). dordrecht: kluwer academic. siegel, h. (1988). educating reason: rationality, critical thinking, and education. ny: routledge. smith, g. (2002). are there domain-specific thinking skills? journal of philosophy of education 36, 207227. van someren, m.w., barnard, y. f., & sandberg, j. a. c. (1994). the think aloud method. a practical guide to modelling cognitive processes. department of social science informatics. university of amsterdam. london: academic press. states, a., min-leliveld, m., gijbels, d., & petegem, p. van (2010). the impact of instructional development in higher education: the state-of-the-art of the research. educational research review 5, 25–49. timmermans, s., & tavory, i. (2012). theory construction in qualitative research: from grounded theory to abductive analysis. sociological theory 30, 167186. tremblay, k., lalancette, d., roseveare, d. (2012) ahelo feasibility study report. volume 1  design and implementation. organization for economic co-operation and development (oecd). retrieved from http://www.oecd.org/edu/skills-beyond-school/ahelofsreportvolume1.pdf valanides, n. and angeli, c. 2005. effects of instruction on changes in epistemological beliefs. contemporary educational psychology 30, 314330. west, e. j. (2004). perry’s legacy: models of epistemological development. journal of adult development 11, 6170. footnotes 1 king and kitchener (2004) do not call the lowest level of reflective thinking realism. however, in their model they maintain that, at the most limited level of thinking, knowledge is certain and is obtained from direct observation (p.7). this position fits metaphysical realism. frontline learning research 6 (2014) 15-24 issn 2295-3159 corresponding author: i.molenaar@pwo.ru.nl doi: http://dx.doi.org/10.14786/flr.v2i4.118 15 | f l r advances in temporal analysis in learning and instruction inge molenaar a a radboud university nijmegen article received 8 june 2014 / revised 23 september 2014 / accepted 23 september 2014 / available online 23 december 2014 abstract this paper focuses on a trend to analyse temporal characteristics of constructs important to learning and instruction. different researchers have indicated that we should pay more attention to time in our research to enhance explanatory power and increase validity. constructs formerly viewed as personal traits, such as self-regulated learning and motivation, are now conceptualized as a series of events that unfold over time. this raises new questions with regard to the temporal characteristics of these constructs and their dynamic interplay with learner and context characteristics. even though the value of analyzing temporal characteristics is becoming evident, a number of challenges need to be tackled in order to make progress in the field of learning and instruction. first, we need to be aware of the paradigm shift that temporal analysis entails. second, a common understanding of different dimensions of time and the position of temporal characteristics therein can facilitate our time-related research dialogue. third, a better understanding how to answer time-related questions with appropriate methodological approaches needs to emerge. fourth, researching temporal characteristics requires procedures and guidelines for segmenting time units. fifth, temporal data are mostly collected at the micro level, whereas most theory is defined at a macro level; consequently we need to bridge these differences in the granularity used between collecting, coding and theorizing to enhance meaning making. finally, so far, most examples of time-related research are exploratory or comparative studies; the next step is to move toward confirmative studies, which constitute the “holy grail” of temporal analysis. keywords: temporal analysis; learning and instruction; time; methodologies i. molenaar 16 | f l r 1. introduction learning is defined as the acquisition of skills and knowledge and can be recognized through changes in the learners’ behaviour (mayer 2008; zimmerman 2002). the concept of time is innate to learning, as it takes time to acquire skills and knowledge and to signal changes in behaviour. in learning and instruction research, we mostly capture time in preand post-test designs. as such we often focus on a narrow concept of time, reducing the temporal characteristics of learning to the changes between preand post-tests which reduces validity and explanatory power of our research. currently, technological advancements increase our ability to gain traces of learners while they are learning, which is an important facilitator to overcome this limited focus on time in our field (greene & azevedo, 2010; reimann, 2009; winne, 2010). a steadily growing group of researchers is raising questions that address how different constructs act and develop over time (bannert et al., 2014, greene & azevedo, 2010; molenaar & chiu, 2014; riemann, 2009; schmitz, 2006; wise & chiu, 2011). with this growing interest in temporal characteristics of constructs at the heart of learning and instruction research, the need for temporal analysis is becoming more prevalent. this paper focuses on the developing trend in learning and instruction research to analyze temporal characteristics of different constructs. the rationale for temporal analysis in our field is discussed as well as the fact that temporal analysis entails a deviation from our main research approach, changing our analysis from characteristics of students to attributes of. learning activities (riemann, 2009; schmidt, 2006). researchers have conceptualized temporal characteristics of learning and instruction constructs in many ways in their research, leading to a diverse set of dimensions of time driving research questions. it is argued that a conceptual framework of temporal characteristics can support transparency and enhance comparability in the field. lastly, a number of challenges are discussed that we, as a field, need to overcome to successfully engage in temporal analysis. 2. the rationale for temporal analysis a number of researchers in the field of computer-supported learning (kapur, 2011; reiman, 2009;) and self-regulated learning (greene & azevedo, 2010; schmitz, 2006; schoor & bannert, 2012) indicate that we should pay more attention to time in the learning process. existing research methods do not “fully” utilize the temporal information embedded in the data collected (kapur, voiklis & kinzer, 2008; wise, perera, hsiao, speer & marbouti, 2012). this reduces the explanatory power of the analysis performed and limits the validity of the conclusions drawn (akhras & self, 2000; reimann, 2009). for example, kuvalja and colleagues (2014) show that self-directed speech and self-regulatory behaviour of children with a specific language impairment does not differ in frequency; neither the number of self-directed speech and selfregulatory events during learning, nor the order between the these events as detected by sequential lag analysis differed, but there was a difference between the two groups in the co-occurrence of self-directed speech and self-regulatory behaviour as detected by temporal patterns analysis (magnusson, 2000). this indicates that without proper temporal analysis, existing differences between groups of learners cannot be detected. moreover, a number of constructs, such as self-regulated learning and motivation, that were traditionally viewed as a trait of the learner are now conceptualized as a series of events (bannert et al. 2014; greene & azevedo, 2009; schmitz, 2006). driving this conceptual change are indications that self-report data have little relation with the actual student behaviour during learning (veenman, 2011). these findings point towards the need for new conceptualisations of these constructs. a temporal conceptualisation viewing i. molenaar 17 | f l r self-regulated learning as a series of events that act differently over time and changing contexts, might overcome these issues (azevedo et al., 2010; hadwin & järvelä, 2011). for example, malmberg and colleagues (2014) show that students use different strategies and learning patterns when working on an illstructured task compared to a well-structured task. a series of events can be perceived as a process that unfolds over time in a certain order (reimann, 2009). for example, self-regulated learning processes of successful students show a cyclical order among different strategies that repeat over time (bannert et al. 2014). moreover, molenaar and chiu (2014) found strong positive predictive relations between different learning activities during collaborative learning over time. the changed conceptualization of constructs raises new questions with regard to the characteristics of these constructs and their dynamic interplay with the learning context. finally, an emerging interest is in connecting different levels of analysis (hollenstein, 2013; suthers, teplovs, de laat, oshima & zeini 2011). this investigates of how macro-level phenomena can emerge from and/or be constrained by different micro-level dynamics. for example, chiu (2008) found that microcreativity in a group’ mathematical solutions can be sparked by a discourse pattern, namely a wrong idea followed by disagreement among the group members. temporal analysis can help develop an understanding of how patterns unfold, providing insights into how learning is taking place (chiu, 2008; wise & chiu, 2013). take together, the argument for temporal analysis is driven by the realisation that without careful attention for temporal characteristics of constructs in learning and instruction research, we are reducing the significance of our research and are unable to explain important aspects of learning and instruction. 3. a paradigm shift as touched upon in the introduction, it is important to understand that advanced temporal analysis entails a deviation from the traditional research paradigm used in learning and instruction (reimann, 2009; schmitz, 2006). often the variable-based approach is applied, which focuses on the analysis of variance between independent and dependent variable(s). in contrast, the event-based approach looks at events analysing the (dynamic) relations between the events (reimann, 2009). this approach focuses on researching the nature of these relations and their development over time. this reveals the temporal characteristics of a construct and/or how different constructs interplay over time. for example, it can indicate how a discussion among learners unfolds over time. consistency and change in the behaviour of constructs can be investigated by specifying these temporal characteristics (schmitz, 2006). yet often reviewers in learning and instruction immediately ask the next question: can we explain learning performance from temporal characteristics of the constructs? this question embodies the ”holy grail” of temporal analysis and often constitutes a connection between our traditional variable-based approach and the event-based approach. however, few (perhaps none) researchers have so far reached the “holy grail”. moreover, many of those initially aiming for this connection started to grow a realization that there are valuable questions to be answered within analysing temporal characteristics themselves. an example of such a research question is: which sequences of learner actions (discuss, elaborate, summarize) occur during collaborative learning? an example of a research question combining temporal characteristics with learning performance is: which sequences of learners’ actions during collaborative learning influence learning performance positively? overall temporal analysis in learning and instruction is innate to our intuitive understanding of learning, but the operationalization of this understanding entails a deviation of our i. molenaar 18 | f l r traditional research paradigm. consequently, the nature of the questions addressed with temporal analysis varies from our characteristic research questions in learning and instruction. after all, time is a highly complex construct that has been debated on from physics to philosophy. also within educational research, conceptualizations of different time scales (lemke, 2000) and the use of time in classrooms (bloome et al., 2009; mercer, 2008) have been discussed. still, there is no framework that conceptualizes dimensions of time and positions different temporal characteristics within these dimensions. research questions, therefore, focus on different dimensions of time and address conceptually different temporal characteristics of constructs. in the next section, different dimensions of time important for learning and instruction research are highlighted. 4. different dimensions of time so far, in our field when addressing temporal characteristics, we have encountered mainly frequency analysis indicating the number of occurrences of a variable during a particular time window. this provides insights into the prevalence of a construct during learning. for example, students receiving scaffolds during learning apply more metacognitive activities compared to students that do not receive scaffolds (molenaar et al., 2011). although informative, frequency analyses provide limited insights into the individual time-related characteristics of the constructs researched. even though this analysis showed that students perform more metacognitive activities, we do not know the importance of their position in the learning process, their duration or the rate at which these metacognitive activities occur during learning. thus frequency analyses treat the learning process as one holistic unit, ignoring the individual time-related characteristics of constructs. using the individual time-related characteristics allows for the analysis to illustrate how events occur within the flow of continuous events in a particular time window. examples are analyzing the significance of the position of events, the duration of particular events and the rate of particular events within the learning process (molenaar & wise, in prep). for example, planning at the start of a learning task was found to be more productive for learning compared to planning latter on (moos et al. 2008). also students monitor at a higher intensity and longer in more difficult tasks compared to easier tasks (iiskala et al. 2010). the dimension of time described above conceptualizes how constructs behave in a continuous flow of events by examining the individual time-related characteristics of these events within the flow. another dimension of time in contrast to analysing events in a continuous flow, is analysing relative arrangements of multiple events in time. here the focus does not lie on the individual time-related characteristics of events in a time window, but on how events are organized among each other. examples are both reoccurring processes and non-reoccurring processes (molenaar & wise, in prep). an example of a reoccurring process is the cyclical notion of self-regulated learning, which suggests that orientation, planning, monitoring and evaluation follow each other (hadwin & jarvela, 2013; zimmerman, 2002). nonreoccurring transitions occur only once, for example, students who learn how to read progress from spelling letters into the automatic detection of words (verhoeven, 2004). apart from reoccurring and non-reoccurring patterns which both indicate a form of regular change, irregular change is another form of an arrangement of multiple events that can be investigated. the notion of productive failure where collaborating students seem to engage in chaotic interaction in the beginning of their collaboration is an example of irregular change (kapur, 2009). this seemingly unstructured process is of essential importance for their later learning. the dimension of time described above conceptualizes how constructs behave in in relative arrangements of multiple events by examining the organisation among these events. i. molenaar 19 | f l r without claiming that the above is a complete overview of temporal characteristics useful for the field of learning and instruction, a clear distinction can be made between two dimensions of time, i.e. focusing on individual events within the continuous flow of events or on relative arrangements of multiple events (molenaar & wise, in prep). in order to push the conceptual understanding of time in our field, a conceptual framework of looking at time and positioning temporal characteristics therein is important for learning and instruction research to articulate and classify time-related research questions. such a framework can support conceptual clarity among researchers engaging in temporal analysis and organize and deepen debates. furthermore, it can be used as a roadmap to articulate temporal research questions, unravelling temporal characteristics of different constructs. 5. an illustrative example of temporal analysis in order to illustrate the above, i provide an example of a temporal analyses used to research socially regulated learning. during collaborative learning, students support one another’s learning as they discuss, elaborate, argue, confirm and regulate one another’s activities. we know that regulative activities such as metacognitive (i.e., planning, monitoring) and relational activities (i.e., confirming, engaging) contribute significantly to students’ learning (molenaar et al., 2011). yet, we know very little about how the group’s socially regulative activities influence students’ cognition at a micro level during collaborative learning. therefore, we explored how sequences of students’ cognitive, metacognitive and relational activities affect the likelihood of subsequent cognitive activities during collaborative learning and whether these relationships differ across time (molenaar & chiu, 2014). the data are from 18 triads (54 students) engaged in 51.338 conversation turns over 6 hours of learning activities. the triads collaborated face-to-face while working in a computer based learning environment. the primary school students were in grades 4, 5, and 6, and aged between 10 and 12. statistical discourse analysis, content and discourse analysis were used to analyse the learning activities. during content analysis, each turn in the conversation was coded as cognitive (higher or lower cognition), metacognitive (orientation, planning, monitoring and evaluation) or relational (confirm, deny, engage), procedural or off task activities. then, statistical discourse analysis (sda) was used to examine the sequential relations predicting lower and higher cognition (chiu & koo, 2005). i. molenaar 20 | f l r figure 1. path diagram of standardized final multivariate outcome, multilevel cross-classification of low cognition component. solid lines indicate positive effects. dashed lines indicate negative effects. thicker lines indicate larger effect sizes. *p < .05, **p < .01, ***p < .001. (molenaar & chiu, 2014; reproduced with permission) we found that high cognitive, low cognitive, metacognitive and relational activities in recent conversation turns were linked to the likelihood of low cognition in a conversation turn (see figure 1). metacognitive activities in the form of planning (in the previous conversation turn or -1), monitoring (-1), evaluating (2 conversation turns ago or -2), monitoring (-2), summarizing (-3) and monitoring (-3) all increased the likelihood of low cognition, while orientation (-2) reduced it. higher cognitive activities in either of the last two conversation turns or low cognition in any of the last six conversation turns also increased the likelihood of low cognition. lastly, relational activities in the form of confirming and engaging in any of the last two conversation turns increased the likelihood of low cognition. this example analyzes temporal characteristics of arrangement of multiple events to understand how these events act within the learning process. this type of analysis illustrates how different learning activities alternate and fluctuate among collaborating students and emerge into socially regulated learning. the findings show recurrent sequential relationships between cognitive, metacognitive and social relational activities. moreover, this analysis indicates that these patterns are rather stable over time. even though these analyses reveal important information about micro-level temporal interaction among learning activities, an often received question is: “what do these relations among learning activities mean for learning, i.e. which sequences should we encourage with instructional designs?” this question embodies the “ holy grail” and has not been addressed yet. although it is an important question, this inquiry clearly indicates the need for a paradigm shift within our field. we need to learn to value results of temporal analysis in their own right, providing important information about constructs in learning and instruction and taking steps to defining micro level temporal theories of how constructs behave over time. i. molenaar 21 | f l r 6. challenges apart from creating the awareness of the need for temporal research questions, there are a number of other challenges that need to be addressed to forward temporal analysis in the field of learning and instruction. as discussed in section 4, time can be conceptualised differently in our research (bloome et al., 2009; lemke, 2000; mercer, 2008). a conceptual framework to articulate different dimensions of time to frame temporal characteristics and related research questions could enhance conceptual clarity and provide ground for in-depth debate about time-related characteristics of individual events in the continues flow of events or about the arrangements of multiple events over time. second, although there are many emerging methods such as visualizations (reimann, 2009), sequential lag analysis (bakeman and gottman, 1997), statistical discourse analysis (chiu & khoo, 2005), temporal pattern analysis (magnusson, 2000), markov modeling (biswas, kinnebrew & segedy, 2012), data mining (robero et al. 2010), and dynamic systems (hollenstein, 2013) used to explore time and order in learning processes, we are only starting to explore the commonalities and differences among these methods. understanding about these techniques, as well as which learning and instruction questions can be answered by their application, is required. comparing different methods can enhance our understanding of temporal characteristics of constructs in learning and instruction (e.g., via triangulation) and methodological issues (e.g., which method is most appropriate for specific research questions?). collaboration among researchers is needed to create guidelines and to work towards a methodological framework for temporal analysis. third, when performing time-related research, we always “cut in time” i.e., we make an artificial division in time units. this segmentation of time can be approached differently, that is at the level of instructional units, time units or units of time in which a construct is acting homogeneously. for example, determining the time window based on the frequency of occurrence of low cognition in the group discourse (molenaar & chiu, 2014). choices made about segmentation have important implications for the results, and therefore, clear guidelines towards determining time windows should be formulated. fourth, granularity of our time related-research is an issue. the level at which we collect and code is often at a micro level capturing very small units, such as events from electronic learning environments or utterances in a dialogue. our theories are usually defined at a macro level, explaining how different constructs act. these different levels of granularity between coding and theory are a challenge for meaning making. aggregation of micro level variables to more macro level constructs can be a solution to this issue. yet, as with segmentation, decisions about granularity used in analysis also impacts results profoundly and should therefore follow clear procedures to ensure quality standards. moreover, combinations of different research traditions, such as ethnographical approaches and data-mining methods, can help make connections between macro level theory and micro level coding. a number of researchers have already indicated the need for micro level temporal theories of constructs to support temporal analysis (azevedo, 2014; bannert et al. 2014; molenaar & chiu, 2014; molenaar & järvelä, 2014; molenaar et al., 2011; kuvulja et al. 2014; winne, 2014). finally, until now, mainly exploratory studies have been done and there is a request from our community to move toward to the holy grail, that is to establish that particular temporal characteristics contribute to learning performance in particular ways. on the one hand, the holy grail will help confirm the value of temporal analysis for the field of learning and instruction. yet, as indicated above, linking these analysis to learning performance is challenging. collaboration among researchers is needed to overcome these issues and create guidelines to work towards a uniform approach for event-based methods to enhance our understanding of the temporal characteristics of learning and instruction. i. molenaar 22 | f l r 7. conclusion in the field of learning and instruction, there is an intuitive belief that temporality is important to comprehend learning. in order for us, as a field, to make progress in understanding the temporal aspects of learning, a number of challenges need to be overcome. the field needs to be aware that temporal analysis often departs from the traditional research approach. in order to enhance this advancement, the field must embrace a different kind of research question specifically related to temporal aspects of learning and instruction. keypoints time deserves more attention in learning and instruction research temporal analysis entails a paradigm shift addressing a different type of research question a conceptual framework of time can support framing temporal characteristics and research questions we need to advance our understanding of methodologies, time segmentation and meaning making of temporal analysis acknowledgments the thinking in this paper reflects idea’s developed and discussed during the various workshops “it’s about time”. all participants in these workshops have contributed to the construction of this understanding and especially my conversations with alyssa wise and ming ming chiu. references akhras, f. n., & self, j. a. (2000). modeling the process, not the product, of learning. in s. p. lajoie, computers as cognitive tools, volume two: no more walls (pp. 3-28). mahwah, nj: lawrence erlbaum associates. azevedo, r. (2014). issues in dealing with sequential and temporal characteristics of self-and sociallyregulated learning. metacognition and learning, 9(2), 217-228. http://dx.doi.org/10.1007/s11409014-9123-1 bakeman, r., & gottman, j. m. (1997). observing interaction: an introduction to sequential analysis. cambridge: cambridge university press. http://dx.doi.org/10.1017/cbo9780511527685 bannert, m. (2006). effects of reflection prompts when learning with hypermedia. journal of educational computing research, 4, 359-375. http://dx.doi.org/10.2190/94v6-r58h-3367-g388 bannert, m., reimann, p. & sonnenberg, c. (2014). process mining techniques for analysing patterns and strategies in students' self-regulated learning. metacognition and learning, vol 9, 161-185. http://dx.doi.org/10.1007/s11409-013-9107-6 segedy, j. r., kinnebrew, j. s., & biswas, g. (2012). supporting student learning using conversational agents in a teachable agent environment. in the future of learning: proceedings of the 10th international conference of the learning sciences (icls 2012) (vol. 2, pp. 251-255). bloome, d., beierle, m. grigorenko, m. & goldman, s. (2009). learning over time: uses of intercontextuality, collective memories, and classroom chronotopes in the construction of learning opportunities in a ninth-grade language arts classroom. language and education, 23(4), pp. 313334. http://dx.doi.org/10.1080/09500780902954257 i. molenaar 23 | f l r chiu, m. m., & khoo, l. (2005). a new method for analyzing sequential processes: dynamic multi-level analysis. small group research, 36, 600-631. http://dx.doi.org/10.1177/1046496405279309 chiu, m. m. (2008). flowing toward correct contributions during groups' mathematics problem solving: a statistical discourse analysis. journal of the learning sciences, 17 (3), 415 463. http://dx.doi.org/10.1080/10508400802224830 goldstein, h. (1995). multilevel statistical models. sydney: edward arnold. günther, c., & van der aalst, w. (2007). fuzzy mining: adaptive process simplification based on multiperspective metrics. in g. alonso, p. dadam & m. rosemann (eds.), international conference on business process management (bpm 2007) (pp. 328-343). berlin: springer. greene, j. a. & azevedo, r. (2010). the measurement of learners’ self-regulated cognitive and metacognitive processes while using computer-based learning environments. educational psychologist, 45, 203 – 209. http://dx.doi.org/10.1080/00461520.2010.515935 hadwin, a.f., & järvelä, s. (2011). introduction to a special issue on social aspects of self-regulated learning: where social and self meet in the strategic regulation of learning. teachers college record, 113(2), 235-239 hollenstein, t. (2013).state space grids: depicting dynamics across development.new york: springer. http://dx.doi.org/10.1007/978-1-4614-5007-8 iiskala. t., vauras, m., lehtinen, e., & salonen, p. (2011). socially shared metacognition within primary school pupil dyads’ collaborative processes. learning and instruction, 21, 379-393. http://dx.doi.org/10.1016/j.learninstruc.2010.05.002 järvelä, s. & hadwin, a. (2013). new frontiers: regulating learning in cscl. educational psychologist, 48(1), 25-39. http://dx.doi.org/10.1080/00461520.2012.748006 lemke, j.l. (2000). across the scales of time: artifacts, activities, and meanings in ecosocial systems. mind, culture and activity, 7(4), 273–290. http://dx.doi.org/10.1207/s15327884mca0704_03 kapur, m., voiklis, j., & kinzer, c. (2008). sensitivities to early exchange in synchronous computersupported collaborative learning (cscl) groups. computers and education, 51, 54-66. http://dx.doi.org/10.1016/j.compedu.2007.04.007 kapur, m. (2011). temporality matters: advancing a method for analyzing problem-solving processes in a computer-supported collaborative environment. international journal of computer-supported collaborative learning (ijcscl), 6,(1), 39-56. http://dx.doi.org/10.1007/s11412-011-9109-9 kennedy, p. (2008). a guide to econometrics. cambridge: blackwell. kinnebrew, j. s., segedy j.r. & biswas, g. (2014). analyzing the temporal evolution of students' behaviors in open-ended learning environments. metacognition and learning, vol 9, 217-228. http://dx.doi.org/10.1007/s11409-014-9112-4 kuvalja, m., verma, m. & whitebread, d. (2014). patterns of co-occurring non-verbal behavior and selfdirected speech; a comparison of three methodological approaches. metacognition and learning, vol 9, 87-111. http://dx.doi.org/10.1007/s11409-013-9106-7 magnusson, m. s. (2000). discovering hidden time patterns in behavior: t-patterns and their detection behavior research methods, instruments, & computers: a journal of the psychonomic society, inc, 32(1), 93–110. malmberg, j., järvelä, s. & kirchner, p. (2014). elementary school students’ strategic learning: does tasktype matter? metacognition and learning, vol 9, p. 113-136. http://dx.doi.org/10.1007/s11409-0139108-5 mayer, r.e. (2008). learning and instruction. pearson; new jersey. mercer, n. (2008) the seeds of time: why classroom dialogue needs a temporal analysis. journal of the learning sciences, 17, 1, 33-59. http://dx.doi.org/10.1080/10508400701793182 molenaar, i., van boxtel, c.a.m. & sleegers, p.j.c. & roda, c. (2011). attention management for selfregulated learning: atgentschool. in c. roda (ed), human attention in digital environments, cambridge university press: cambridge, 259 280.ttp://dx.doi.org/10.1017/cbo9780511974519.011 molenaar, i., chiu, m. m., van boxtel, c. & sleegers, p. j.c. (2011). scaffolding of small groups’ metacognitive activities with an avatar. international journal of computer-supported collaborative learning, 6, 601-624. http://dx.doi.org/10.1007/s11412-011-9130-z i. molenaar 24 | f l r molenaar, i., van boxtel, c.a.m & sleegers, p.j.c. (2011). metacognitive scaffolding in an innovative learning arrangement. instructional science, vol 39(6), 785-803. http://dx.doi.org/10.1007/s11251010-9154-1 molenaar, i & chiu m.m. (2014). dissecting sequences of regulation and cognition: statistical discourse analysis of primary school children’s collaborative learning. metacognition and learning, vol 9, 137160. http://dx.doi.org/10.1007/s11409-013-9105-8 molenaar, i. & järvelä, s. (2014). sequential and temporal characteristics of self and social regulated learning. metacognition and learning, vol 9, p. 75-85. http://dx.doi.org/10.1007/s11409-014-9114-2 molenaar, i & wise, a.f. (in prep). concepts of time: a framework for thinking about temporal aspects of learning. moos, d. c., & azevedo, r. (2008). self-regulated learning with hypermedia: the role of prior domain knowledge. contemporary educational psychology, 33(2), 270-298. http://dx.doi.org/10.1016/j.cedpsych.2007.03.001 reimann, p. (2009). time is precious: variableand event-centred approaches to process analysis in cscl research, international journal of computer-supported collaborative learning, 3, 239-257. http://dx.doi.org/10.1007/s11412-009-9070-z schegloff, e. a., 2007. sequence organization in interaction: a primer in conversation analysis. cambridge: cambridge university press. http://dx.doi.org/10.1017/cbo9780511791208 schmitz, b. (2006). advantages of studying processes in educational research. learning and instruction. 16, 433-449. http://dx.doi.org/10.1016/j.learninstruc.2006.09.004 schoor, c. & bannert, m. (2012). exploring regulatory processes during a computer-supported collaborative learning task using process mining. computers in human behavior. 28(4), 13211331. http://dx.doi.org/10.1016/j.chb.2012.02.016 suthers. d., teplovs, c., de laat, m., oshima, j., & zeini, s. (2011). connecting levels of learning in networked communities. workshop conducted at the 9th international conference on computer supported collaborative learning, july 9, 2011, hong kong. robero, c., ventura, s., pechenizkiy, m., & baker, r. (eds.). (2010). handbook of educational data mining. boca raton: chapman&hall/crc. veenman, m.v.j. (2011). learning to self-monitor and self-regulate. in r. mayer,& p. alexander (eds.), handbook of research on learning and instruction. new york: routledge. weinberger, a., & fischer, f. (2006). a framework to analyze argumentative knowledge construction in computer-supported collaborative learning. computers & education, 46, 71-95. http://dx.doi.org/10.1016/j.compedu.2005.04.003 winne, p. h. (2014). issues in researching self-regulated learning as patterns of events. metacognition and learning, 229-237. http://dx.doi.org/10.1007/s11409-014-9113-3 wise, a. f., & chiu, m. m. (2011). analyzing temporal patterns of knowledge construction in a role-based online discussion. international journal of computer-supported collaborative learning, 6(3), 445470. http://dx.doi.org/10.1007/s11412-011-9120-1 wise, a. f., perera, n., hsiao, y. , speer, j., & marbouti, f. (2012). microanalytic case studies of individual participation patterns in an asynchronous online discussion in an undergraduate blended course. the internet and higher education, 15(2), 108-117. http://dx.doi.org/10.1016/j.iheduc.2011.11.007 zimmerman, b. j. (2002). becoming a self-regulated learner: an overview. theory into practice, 42(2), 64-70. http://dx.doi.org/10.1207/s15430421tip4102_2 microsoft word südkamp et al_publication.docx         frontline learning research vol.3 no. 2 (2015) 1-26 issn 2295-3159 1  corresponding author: anna südkamp, emil-figge-str. 50, 44227 dortmund, germany, phone: +49 231 755 6570, fax: +49 231 755 6572, e-mail: anna.suedkamp@tu-dortmund.de doi: http://dx.doi.org/10.14786/flr.v3i2.130   competence assessment of students with special educational needs—identification of appropriate testing accommodations anna südkampa, steffi pohlb, & sabine weinertc atu dortmund university, germany bfreie universität berlin, germany cuniversity of bamberg, germany article received 23 october 2014 / revised 16 march 2015 / accepted 18 may 2015 / available online 1 june 2015 abstract including students with special educational needs in learning (sen-l) is a challenge for largescale assessments. in order to draw inferences with respect to students with sen-l and to compare their scores to students in general education, one needs to assure that the measurement model is reliable and that the same construct is measured for different samples and test forms. in this article, we focus on testing the appropriateness of competence assessments for students with sen-l. we specifically asked how the reading competence of students with sen-l may be assessed reliably and comparably. we thoroughly evaluated different testing accommodations for students with sen-l. the reading competence of n = 433 students with sen-l was assessed using a standard reading test, a reduced test version, and an easy test version. also, n = 5,208 general education students and a group of n = 490 lowperforming students were tested. results show that all three reading test versions are suitable for a reliable and comparable measurement of reading competence in students without sen-l. for students with sen-l, the accommodated test versions considerably reduced the amount of missing values and resulted in better psychometric properties than the standard test. they did not, however, show satisfactory item fit and measurement invariance. implications for future research are discussed. keywords: students with special educational needs; testing accommodations; reading competence; large-scale assessment südkamp et al   2 | f l r       1. introduction large-scale assessments generally aim at drawing inferences about individuals’ knowledge, competencies, and skills (popham, 2000). today, educational assessments play an important role as they inform students, parents, educators, policymakers, and the public about the effectiveness of educational services (pellegrino, chudowsky, & glaser, 2f001). using results from large-scale assessments, researchers can study factors influencing the acquisition and development of competencies and derive strategies on the improvement of educational systems. often, assessments are meant to serve even more ambitious purposes such as supporting student learning (chudowsky & pellegrino, 2003). assessing students’ domain-specific competencies (e.g., reading competence, mathematical competence) is a key aspect of most large-scale assessments today (weinert, 2001). in this study, we focus on the assessment of competencies of students with special educational needs (sen) in large-scale assessments. while national large-scale assessments like the national assessment of educational progress (naep) in the united states and international assessments like the programme for international student assessment (pisa) have established sophisticated methods for the assessment of students without sen, testing students with sen has proven to be challenging. in order to inform strategies for the assessment of students with sen, we evaluate whether and if so, how students with sen may be tested reliably and comparably to general education students. for this purpose, students with and without sen were tested with accommodated and non-accommodated test versions. on the level of the single items, we carefully checked the reliability and comparability of the test scores obtained with the different test versions as reliability and comparability are necessary prerequisites for drawing meaningful inferences from large-scale assessments. 1.1 assessing reading competence of students with sen large-scale assessments usually aim at describing the abilities of students within a country across the whole spectrum of the educational system or even across countries. this also includes students with sen. in our notion, students with sen include all students who are provided with special educational services due to a physical or mental impairment. in germany, special schools are established for students with sen. the special school system—in turn—is highly differentiated itself. there are special schools for students with special educational needs in learning, visual impairments, hearing disability/impairment, specific language/speech impairments, physical handicaps/disabilities, severe intellectual impairment/disability, emotional and behavioral difficulties, comprehensive sen, and students with health impairment. so far, comparatively little is known about the educational careers of students with sen and their development of competencies across the life span (heydrich, weinert, nusser, artelt, & carstensen, 2013; ysseldyke et al., 1998). however, there is evidence that for students with sen, reading problems pose one of the greatest barriers to success in school (kavale & reece, 1992; swanson, 1999). learning to read is a tedious process requiring psycholinguistic, perceptual, cognitive, and social skills (gee, 2004). beyond the basic acquisition of the alphabet system (i.e., letter-sound correspondence and spelling patterns), reading expertise implies phonological processing and decoding skills, linguistic knowledge (vocabulary, grammar), and text comprehension skills (durkin, 1993; verhoeven & van leeuwe, 2008). according to kintsch (2007), text comprehension can be seen as a combination of text-based processes that integrate previous knowledge to a mental representation of the text. it is thus a form of cognitive construction in which the individual takes an active role. text comprehension entails deep-level problem-solving processes that enable readers to construct meaning from text and derives from the intentional interaction between reader and text (duke & pearson, 2002; durkin, 1993). on average, students with sen show lower reading performance in large-scale assessments than students without sen (thurlow, 2010; thurlow, bremer, & albus, 2008; ysseldyke et al., 1998). for example, for the naep 1998 reading assessment in grades 4 and 8, lutkus, mazzeo, zhang, and jerry (2004) südkamp et al   3 | f l r       report lower average scale scores for students with sen compared to students without sen. within the german kess study (bos et al., 2009) reading competence of seventh graders in special schools was compared to the reading competence of fourth graders in general education settings. results demonstrated that fourth grade primary school students outperformed students with sen in seventh grade in reading competence, the difference being about one third of a standard deviation. drawing on data from a three-year longitudinal study, wu et al. (2012) found that, compared to their general education peers, students receiving special educational services were more likely to score below the 10th percentile for several years in a row. in light of these findings, different reasons for the low performance of students with sen have been discussed (abedi et al., 2011). first, some students with sen have difficulties related to the comprehension of text (e.g., lack of knowledge of common text structures, restricted language competencies, inappropriate use of background knowledge while reading; gersten, fuchs, williams, & baker, 2001). reading problems of students with sen in upper elementary and middle school are likely to be complex and heterogeneous resulting, for example, from a lack of phonological processing and decoding skills, a lack of linguistic knowledge (vocabulary, grammar), and a lack of text comprehension skills, or from a combination of problems in these areas. second, lower performance could be attributed to a lack of opportunities to learn and to low teacher expectations (woodcock & vialle, 2011). third, there could be barriers for students with disabilities in large-scale assessments that lead to unfair testing conditions (pitoniak & royer, 2001). according to thurlow (2010), a combination of all these factors is likely. taking the norm of test fairness seriously, largescale studies try to ensure that students with disabilities will not be confronted with unfair testing conditions. that is why testing accommodations are often employed for students with sen. 1.2 providing students with sen with testing accommodations the provision of testing accommodations for individuals with disabilities is a highly controversial issue in the assessment literature (pitoniak & royer, 2001; sireci, scarpati, & li, 2005). generally, testing accommodations are defined as changes in test administration that are meant to reduce construct-irrelevant difficulty associated with students’ disability-related impediments to performance. according to the standards for educational and psychological testing, accommodations comprise “any action taken in response to a determination that an individual’s disability requires a departure from established testing protocol. depending on circumstances, such accommodation may include modification of test administration processes or modification of test content” (american educational research association, 1999, p. 110). note that some authors differentiate between accommodations and modifications: while accommodations are not meant to change the nature of the construct being measured, modifications result in a change in the test and equally affect all students taking it (hollenbeck, tindal, & almond, 1998; tindal, heath, hollenbeck, almond, & harniss, 1998). in this article, however, we use the definition of the standards for educational and psychological testing. due to the many types of disabilities, various accommodations have been provided when testing students with sen. accommodations include, for example, modification of presentation format—including the use of braille or large-print booklets for visually-impaired examinees and the use of written or signed test directions for hearing-impaired examinees—and modification of timing, including extended testing time or frequent breaks (koretz & barton, 2003). in the 1998 naep reading assessment, a sample of students with varying disabilities and students with limited english proficiency were assigned to the following accommodations based on their individual needs: one-on-one testing, small-group testing, extended time, oral reading of directions, signing of directions, use of magnifying equipment, and use of an aide for transcribing responses. changes in the test bear the possibility that they alter the construct measured. if accommodated tests for students with sen measure a different construct than the standard test for general education students, the competence scores between the two student groups are not comparable. thus, it is utterly important to test whether test accommodations result in reliable and comparable competence measures (borsboom, 2006; südkamp et al   4 | f l r       millsap, 2011). lutkus et al. (2004) address the issue of whether the naep reading construct remains comparable for accommodated versus non-accommodated students by analyzing differential item functioning (dif). dif exists when subjects with the same trait level have a different probability of endorsing an item. only very few items were found to have statistically significant dif for the focal group (accommodated students) versus the reference group (non-accommodated students), which indicated measurement invariance across subgroups being assessed with different tests. in contrast, koretz (1997) did find indications of dif as 13 of 22 common items showed strong dif when comparing item difficulty for students with sen tested with accommodations and students without sen tested under standard conditions using data from the kentucky instructional results information system assessment. in pisa 2012, samples of students with sen were also tested with accommodated test versions (a shortened test version of the standard test and a test version including easier items). here, the results on the psychometric properties of the accommodated test versions are still to be published (müller, sälzer, mang, & prenzel, 2014). in sum, the results concerning the use of testing accommodations are inconsistent. a major concern remains that in some cases accommodations may alter the test to the extent that accommodated and non-accommodated tests are no longer comparable (abedi et al., 2011; bielinski, thurlow, ysseldyke, freidebach, & freidebach, 2001; cormier, altman, shyyan, & thurlow, 2010). however, one drawback of the studies by lutkus et al. (2004) and koretz (1997) is that in the analyses, different accommodations are not distinguished although different accommodations may have different effects. another disadvantage is that students with sen are often compared to students without sen at the same grade level. here, students with sen and students without sen differ not only in terms of their sen status but also in terms of their expected achievement level. the study by yovanoff and tindal (2007) is one of the rare studies using an alternative comparison group where students with sen in grade 3 are compared to students without sen in grade 2. another issue is that comparisons usually involve students without sen receiving the standard test and students with sen receiving the accommodated test versions. by doing this, possible dif may be due to both, testing accommodations and problems of testing students with sen. in order to disentangle the appropriateness of test accommodations from the testability problems of students with sen, the effects of test accommodations should separately be tested in a group of students without sen. in the same vein, pitoniak and royer (2001) identify three major challenges for research on testing accommodations: variability in examinees, variability in accommodations, and small sample sizes (also see geisinger, 1994). in the present study, we approach these challenges by focusing on students with special educational needs in learning (sen-l), by focusing on specific accommodations appropriate for students with sen-l, and by using a study design that incorporates a group of low-achieving students without sen for evaluating the appropriateness of testing accommodations. 1.3 testing students with special educational needs in learning (sen-l) while providing students with physical, hearing, and visual impairments with testing accommodations is rather accepted, pitoniak and royer (2001) stress the importance of studying the effects of testing accommodations on test validity (or comparability), especially when testing students with learning disabilities. in this study, we focus on students with sen-l in germany, who comprise all students, who are provided with special educational services due to a general learning disability1. in germany, students are assigned to the sen-l group when their learning, academic achievement, and/or learning behavior are impaired (kmk, 2012) and when students cognitive abilities are below normal range (grünke, 2004). in contrast to students with sen-l, students with (specific) learning disabilities (e.g., a reading disorder) are not necessarily impaired in their general cognitive abilities. in germany, the decision of whether a student                                                                                                                           1 as for the term “learning disabilities”, the term sen-l is not clearly defined. note that we refer to a heterogeneous group of students with multifaceted etiology. südkamp et al   5 | f l r       has special educational needs in learning is based on a diagnostic procedure and made collaboratively by parents, teachers, consultants, and school administrations. about 78% of the sen-l students in germany (kmk, 2012) do not attend regular schools but attend special schools with specific programs and trainings tailored to those who are unable to follow school lessons and subject matter in regular classes. in fact, students with sen-l compose the largest group of students with special educational needs in germany (kmk, 2012). comparably, students with learning disabilities compose the largest group of students with disabilities in the unites states (cortiella & horowitz, 2014; us department of education, 2013). our assumption is that the acceptance of testing accommodations for students with sen-l is low, because the disabilities of students with sen-l (e.g., information processing restrictions) are very likely to interfere with the construct that is to be measured (e.g., reading literacy). in turn, respective testing accommodations are likely to be construct-relevant. there are two test accommodations typically implemented for students with sen-l: extended test time and “out-of-level” testing. extended test time is usually implemented in order to compensate for information-processing restrictions in students with sen-l. in his review on the appropriateness of extended time accommodations for students with sen—including students with sen-l among others—lovett (2010) identified two studies with a serious amount of differentially functioning items, while dif was negligible in one other study. a prominent hypothesis regarding extended test time is the differential boost hypothesis (fuchs, fuchs, eaton, hamlett, & karns, 2000), which states that students with sen benefit more from extended time than students without sen. in their review on test accommodations for students with sen including 14 studies on extended time, sireci et al. (2005) conclude that students with sen as well as students without sen benefit from extended test time. in only one of the reviewed studies students with sen benefited more from extended test time than students without sen. another common method is to provide students with sen-l with an out-of-level test, which was originally meant for testing younger children (thurlow, elliott, & ysseldyke, 1999). similarly, alternate assessments that test lower-level reading and mathematical skills or skills that are precursory to reading and numerical literacy can be applied (zebehazy, zigmond, & zimmerman, 2012). both methods aim at avoiding undue frustrations for students with sen-l and at improving the accuracy of measurement. critics of out-oflevel and alternate assessments argue that students with sen-l are faced with low expectations due to the assessment and are prevented from taking the standard tests, and thus consider the assessments to be inappropriate for accountability assessment. nevertheless, thurlow et al. (1999) consider out-of-level testing a good opportunity to test students with sen, if one can make sure that a common scale across different disparate grade levels is available. such a common scale may be achieved by using methods of item response theory (irt), given that the items measure the same construct. when scaling oregon’s early reading alternate assessment onto the first general statewide benchmark reading assessment in grade 3, yovanoff and tindal (2007) identified good psychometric properties of the alternate assessment and no severe dif between students with sen (grade 3) and students without sen (grade 2). however, data that support either the use or nonuse of out-of-level testing or alternate assessments is still rare (minnema, thurlow, bielinski, & scott, 2000; see the study by zebehazy et al., 2012, which focuses on visually impaired students, for an exception). 2. research questions as prior research has shown, testing competencies of students with sen-l represents a challenge for large-scale assessments (thurlow, 2010). assessing competencies of students with sen-l with tests that have been developed for students without sen-l may fail to result in satisfying item fit measures and may be associated with differential item functioning, which impedes the opportunity to compare the competence scores of students with and without sen. the present study aims to evaluate different strategies of testing südkamp et al   6 | f l r       students with sen-l. generally, we address the question of whether and how satisfying item fit measures and measurement invariant test scores can be obtained for students with sen-l in large-scale-assessments. we evaluate whether standard tests developed for students without sen-l and testing accommodations for students with sen-l result in reliable and comparable measures of reading competence. if a reliable and comparable measurement of reading competence can be achieved, substantial research on the competence level, predictors of reading competence and competence development, as well as group differences may be investigated. in this study, two major research questions are addressed: first, we investigate whether a reduction in test difficulty and a reduction of the number of items lead to test results comparable to students tested without accommodations. we approach these questions by testing students in general education, for whom reliable and valid competence scores can be obtained using a standard reading test. as the accommodated test versions are targeted towards a lower competence level, we did not use the whole group of students in general education, but focused on the subgroup of low-achieving students. secondly, we explore whether these accommodations are suitable for testing students with sen-l. 3. method 3.1 sample and design we collected data within the german national educational panel study (neps). the neps is a national, large-scale longitudinal multicohort study that investigates the development of competencies across the lifespan (blossfeld & von maurice, 2011; blossfeld, von maurice, & schneider, 2011). the study aims at providing high-quality, user-friendly data on competence development and educationally relevant processes for an international scientific community (barkow et al., 2011). between 2009 and 2012, six representative start cohorts (aßmann et al., 2011) were sampled, including about 60,000 individuals from early childhood to adulthood. specific target groups include migrants (kristen et al., 2011) and students with sen-l (heydrich et al., 2013). all participants are accompanied on their individual educational pathways through a collection of data on competencies (weinert et al., 2011), learning environments (bäumer, preis, roßbach, stecher, & klieme, 2011), educational decisions (stocké, blossfeld, hoenig, & sixt, 2011), and educational returns (gross, jobst, jungbauer-gans, & schwarze, 2011). following the principles of universal design (dolan & hall, 2001; thompson, johnstone, anderson, & miller, 2005), the neps aims at providing a basis for fair and equitable measures of competencies for all individuals. in the present study, we used data from three different studies of students in fifth grade. these studies comprise a) a representative sample of general education students (main sample), b) a sample of students with sen-l, and c) a group of students in the lowest academic track (lat). the response rate in these studies was 55%, 45%, and 63%, respectively. in the main sample there were n = 5,208 general education students, including n = 700 students in the lowest academic track (see aßmann, steinhauer, & zinn, 2012, for more information on the neps main sample). on average, these students were mage = 10.95 (sdage = .53) years old and 48.3% were female (0.7% had a missing response on age, 0.2% had a missing response on gender). about 24.1% of the students reported that they spoke a language other than german at home. the sample of students with sen-l draws on a feasibility study with n = 433 students who were recruited at special schools for children with sen-l in germany. students in this sample were mage = 11.41 (sdage = .63) years old and 43.3% were female (0.7% had a missing response on gender). in this sample, about 30.1% of the students reported that they spoke a language other than german at home. südkamp et al   7 | f l r       in this feasibility study, we applied two accommodated test versions that aimed at a) reducing the difficulty of the test and b) reducing the test length (and thereby increasing the testing time per item). in order to discern whether test items do not function properly because the accommodations change the test construct or whether students with sen-l still have problems with the test, we implemented a group of low achieving students without sen. this group consisted of a separate sample of n = 490 students enrolled in the lowest academic track, or hauptschule. students in this sample were mage = 11.28 (sdage = .63) years old and 48.4% were female. about 29.8% of the students in the lat spoke a language other than german at home. focusing on this sample, we evaluated whether the accommodated test versions yield reliable test scores and whether they assess the same construct as the standard reading test. students without sen were tested as for this group it has already been shown that reliable and valid competence assessment can be obtained using the standard reading test. thus, we could investigate the impact of the testing accommodations and disentangled testing problems resulting from badly-constructed accommodated test versions and testing problems resulting from the assessment of students with sen-l. we restricted our sample to students in the lowest academic track, because the accommodated test versions were targeted towards a lower competence level. for students in general education in higher academic tracks, the test accommodations would be too easy and, as a consequence of such low test targeting, could result in aberrant response patterns (due to motivation problems) as well as in low item discriminations (due to the low variability in item responses). implementing this group of low-achieving students allowed us to investigate whether the accommodated test versions generally result in reliable and comparable measures of competence. all students were tested in the middle of fifth grade in november and december 2010. data were collected by the international association for the evaluation of educational achievement (iea) data processing and research center (dpc). students participated in the study voluntarily, so student and parental consent was necessary. each student who participated in the study received 5 euros. 3.2 measures and procedures within all three samples, reading literacy as well as mathematical competence was assessed. the orientation towards the functionality and everyday relevance of the competencies studied is one central aspect of the neps framework for the assessment of competencies. it draws on the concept of literacy in international comparative studies with a focus on enabling participation in society (see oecd, 1999). in this study, we focus on the assessment of reading literacy. within the neps, the reading competence assessment focuses on text comprehension. all reading tests are developed based on a framework for the assessment of reading competence (gehrer, zimmermann, artelt, & weinert, 2013). this framework has been developed based on theoretical and pragmatic considerations that take earlier concepts and studies of reading competence within large-scale assessments into account. the most important dimensions within the framework are text types, cognitive requirements, and task formats. concerning text types, texts with commenting, information, literacy-aesthetic, instruction, and advertising functions are included. in turn, cognitive requirements range from finding information in the text, drawing text-related conclusions, and reflecting and assessing. across all age groups, the items in the test are either simple multiple choice (mc) items, complex mc items, or matching items. complex multiple-choice (cmc) items present a common stimulus followed by a number of mc questions with two response options each. matching (ma) items consist of a common stimulus followed by a number of statements, which require assigning a list of response options to these statements (see gehrer, zimmermann, artelt, & weinert (2012) for a full description of the framework including information on text types, cognitive requirements, item formats, and example items). südkamp et al   8 | f l r       3.2.1 standard reading test the standard reading test was designed for students enrolled in the regular school system. it was developed based on the conceptual framework sketched above. students were asked to read five different texts and answer questions focusing on the content of these texts (gehrer, zimmermann, artelt, & weinert, 2013). the test for students in fifth grade included a text about a continent (information function), a recipe (instruction function), an invitation (advertising function), a critical statement on a societal topic (commenting function), and a fictive story about a famous character (literacy-aesthetic function). in the analysis of the standard reading test, 56 items were included; however, subtasks of complex mc and matching items were treated as single items. so when combined, there were 33 questions in the standard reading test, which students had to complete within 30 minutes. for testing general education students, the test has shown good psychometric properties (pohl, haberkorn, hardt, & wiegand, 2012). 3.2.2 reading test with accommodations based on the standard reading test, two accommodated test versions were administered in this study. as mentioned above typical testing accommodations for students with sen-l include extended testing time and “out of level” testing. within the neps, time for testing a domain-specific or domain-general competence is limited to 30 minutes. under this restriction, we decided to develop one accommodated test version by reducing test length (reduced test), resulting in an increased test time per item. one text and its respective nine items, plus an additional 10 hard items were removed. the text on the societal topic and the items were removed, because the items showed to be comparatively difficult in prior item analyses in samples of general education students. in order to facilitate scaling of the different test versions on a same scale, an anchor item design (e.g., kolen & brennan, 2004) was used for linking the different test versions. for this design a sufficient number of items need to be the same in all test forms. therefore, in the reduced test four texts and 37 items remained the same as in the standard reading test and functioned as anchor items in this design. we refer to the term “anchor item” when an item is the same in the standard test and in the accommodated test versions. as a result of reducing the length of the test, one text function was left out in the reduced test version (the commenting function). still, the anchor items represented all three cognitive requirements. note that while this accommodation mainly served to reduce test length, it also reduced item difficulty. we decided to develop a second accommodated test version (easy test) that mainly aimed at reducing the difficulty of the standard test. therefore, three texts and their respective 37 items from the standard reading test were removed (the text about the continent, the critical statement on a societal topic, and the fictive story about a famous character). these texts and its respective items were replaced with three texts and 23 items that had been developed for younger children in grade 3—including a text on the human body (information function), a short story about a family (literacy-aesthetic function), and an invitation (advertising function). this procedure can be considered as some sort of “out-of-level” testing. however, two texts remained the same as in the standard reading test as we used an anchor item design. based on prior item analysis in samples of general education students in grade 5, five especially difficult items were eliminated from these texts. this procedure resulted in 12 overlapping items in the standard reading test and the easy test version. these items were used as anchor items in this design. in sum, the easy test version included 35 items. overall, 5,208 general education students including 700 students from the lowest academic track were tested with the standard reading test. students with sen-l took the standard reading test (n = 176), the reduced test (n = 173), or the easy test (n = 84) by random assignment. the additional sample of n = 490 students from the lowest academic track was randomly assigned to the reduced test (n = 332) and the easy test (n = 158). note that the standard reading test was not administered to this sample of students in the lowest academic track. for investigating the appropriateness of the standard reading test for students in the südkamp et al   9 | f l r       lat, the subsample of the main sample of general education students attending schools of the lowest academic track were used (n = 700). in order to control for fatigue and acquaintance effects, the order of the different tests was rotated within the booklet in almost all test versions. the standard test and the easy test were administered either before or after a mathematics test. for the analyses, due to sample size issues, the test order was ignored and the different conditions were analyzed together. due to sample size limitations, there was no rotation of the position of the reduced test; the reduced test was only administered before the mathematics test. for the comparison of estimated item difficulties with general education students, data from these students refer to the same position within the booklet as data of students with sen-l or the students in the lat. so no bias is to be expected from test position. 4. analyses 4.1 the model we scaled the data within the framework of item response theory (irt). in accordance with the scaling procedure for competence data in the neps (pohl & carstensen, 2012; 2013), we used a rasch model (rasch, 1960) estimated in conquest (wu, adams, wilson, & haldane, 2007). in this model a unidimensional measurement model with equal loadings across items is proposed. various fit indices are available that describe the psychometric properties of the tests. as described above, the reading test included complex mc and matching items. these items consisted of a set of subtasks that were aggregated to a polytomous variable in the final scaling model in the neps. when aggregating the responses on the subtasks to a single polytomous super-item, we lose information on the single subtasks. since in this study we were interested in the fit of the items, we treated the subtasks of complex mc and matching items as single dichotomous items in the analyses. as such, we could not account for possible local item dependence within each set of subtasks. we applied the rasch model to every test version (standard test, reduced test, easy test) and sample (students with sen-l, students in the lat). 4.2 item fit in order to investigate whether the standard test and the accommodated reading tests reliably measured reading competence, we evaluated different fit measures. these included the weighted mean square (wmnsq; wright & masters, 1982), item discrimination, point-biserial correlation of the distractors with the total score and the empirically approximated item characteristic curve (icc). all of these measures provide information on how well the items fit a unidimensional rasch model. as wu (1997) showed, fit statistics depend on the sample size. the larger the sample size, the smaller the wmnsq and the greater the t-value. thus, since the group of students with sen-l differs in sample size from the group of students in the lat, we considered different evaluation criteria for the interpretation of the wmnsq. in this study, we report item discrimination, which describes the pointbiserial correlation of the item with the total score (i.e., relative number of correct responses on the total number of valid responses). a well-fitting item should have a high positive correlation—that is, subjects with a high ability should score higher on the item than subjects with a low ability. for an easier interpretation, we report the discrimination not only in absolute values, but classify the item fit regarding the discrimination into acceptable item fit (discrimination > .2), slight misfit (discrimination between .1 and .2) südkamp et al   10 | f l r       and strong misfit (discrimination < .1). furthermore, point-biserial correlations of incorrect response options and the total score are evaluated. the correlations of the incorrect responses with the total score allow for a thorough investigation of the performance of the distractors. a good item fit would imply a negative or zero correlation of the distractor with the total score. distractors with a high positive correlation may indicate an ambiguity in relation to the correct response. finally, empirically approximated item characteristic curves (icc) were considered. these describe whether the number of correct responses corresponds to the theoretical implied response probability at each competence level. 4.3 measurement invariance reading scores of students with sen-l versus students in general education can only be compared when the tests are measurement invariant—that is, when there is no differential item functioning (dif). measurement invariance is furthermore a necessary assumption for linking the different test forms. when measurement invariance holds—and thus there is no dif—the probability of endorsing an item is the same for students with sen-l and those without sen-l who have the same ability. the presence of dif is an indication that the respective reading test measures a different reading construct for both target groups, and thus that the reading scores between the target groups may not be compared. we tested dif for each test version (standard, reduced, easy) and each target group (students with sen-l, students in the lat) by comparing the estimated item difficulties in the respective test version and target group to the estimated item difficulty of the same items for students in the main sample of the neps. students with sen-l as well as students in the lat were, thus, compared to general education students in the main sample. there is one exemption: the group of students in the lat was not tested with the standard reading test. in order to estimate dif for that group on the standard test, we used the data of the students in the lowest academic track of the main sample of general education students. for this, we separated the main sample into students in the lowest academic track and students in other tracks and compared the estimated item difficulty between both groups. we estimated dif in a multi-facet irt model, estimating separate item difficulties for general education students and for the respective target group. in line with the benchmarks chosen in the neps (pohl & carstensen, 2012), we considered absolute differences in item difficulties greater than 0.6 to be noticeable and absolute differences greater than 1 to be strong dif. note that these benchmarks serve here as an orientation for interpretation. to get a thorough picture, we also report the absolute dif value. also note that while in the standard test dif may be investigated for all items, dif in the reduced test and the easy test may only be investigated for the anchor items. in the reduced test and the easy test there are anchor items that allow linking of the different test versions. as described above, there are 37 anchor items in the reduced test and 12 anchor items in the easy test. 5. results in the following we will first represent the occurrence of missing values in each test form. then we will present item fit for the different test forms and samples, followed by a further investigation of reasons for item misfit. in a next step, results on the comparability of test scores are presented. the results on item fit and measurement invariance are then considered together for evaluating the appropriateness of the different test forms for assessing competencies of students with sen-l. südkamp et al   11 | f l r       5.1 missing responses table 1 depicts the mean of the relative amount of different kinds of missing responses for each of the target groups and test versions. similar to the main study (pohl et al., 2012), there is a large number of missing responses—on average, up to 19% of the items are missing. the amount of missing responses is larger in the students with sen-l group than in the group of students in the lat for all test versions and all types of missing responses. comparing the different test versions, the lowest number of omitted items is found in the easy test version. this is probably due to the fact that the easy test version contains many easy items and that omission of items is related to the difficulty of the item (see, e.g., pohl, gräfe, & rose, 2014). the lowest number of not reached items is found in the reduced test version. thus, the reduction of texts and items to work on within the given assessment time does increase the number of items reached. the lowest number of invalid missing responses occurs in the reduced test version. this is likely because the reduced test version contains fewer matching items; this is the item format with the largest number of invalid responses (pohl et al., 2012). table 1 averages of the relative frequency of missing responses type of missing response test sen-l m lat m omitted standard 6.72 4.78 reduced 5.20 2.70 easy 2.01 0.92 not reached standard 10.46 9.45 reduced 3.90 1.04 easy 5.63 3.44 invalid standard 1.13 0.44 reduced 0.48 0.18 easy 1.59 0.18 total number of missing responses standard 18.31 14.67 reduced 9.58 3.92 easy 9.22 4.53 note. sen-l = special educational needs in learning; lat = lowest academic track. 5.2 item fit 5.2.1 standard test first we analyzed item fit for the standard reading test for students with sen-l and students in the lowest academic track. overall, item discrimination is relatively small for students with sen-l. the mean item discrimination is .25 (it is .34 in the lowest academic track). four items show a slight misfit (discrimination between .1 and .2) and 10 items a strong misfit (discrimination less than .1). in the lowest academic track, there is only one item with a strong misfit and nine items with a slight misfit. evaluation of further fit measures for students with sen-l confirms these results. table 2 depicts the number of misfitting items for the wmnsq, icc, and point-biserial correlations. summarizing these results, there is a large amount of items in the standard test that do not fit. eap-reliability of competence südkamp et al   12 | f l r       scores for students in the lowest academic track is sufficiently high (rel = 0.823), while it is considerably lower for students with sen-l (rel = 0.652).we can conclude that students with sen-l may not be tested appropriately with the standard reading test. in contrast, fit indices in the lowest academic track indicate a relatively good item fit that is comparable to the fit found in the main sample of general education students (see pohl et al., 2012 for the results in the main study). the results indicate that the test is appropriate not only for the main sample including students attending higher academic tracks but also for low-performing students. table 2 number of items with misfit indicated by weighted mean square (wmnsq), item characteristic curve (icc), and point-biserial correlations fit measure test sen-l lat wmnsq standard 7 7 reduced 2 1 easy 1 3 icc standard 15 9 reduced 17 2 easy 12 1 point-biserial correlations standard 21 3 reduced 14 1 easy 5 0 note. sen-l = special educational needs in learning; lat = lowest academic track. 5.2.2 reduced test the item discriminations of the items in the reduced test version indicate a better item fit for students in the lat than for students with sen-l. for both target groups, the reduced test shows better item fit indices than the standard test version. for students with sen-l, there are six items with a slight misfit (discrimination between .1 and .2) and five items with a strong misfit (discrimination below .1). note that— not necessarily—the items showing misfit in the standard reading test, also show low discriminations in the reduced test. this may indicate that problems with testing of students with sen-l do not necessarily lay in the specificity of the items, but may reflect other aspects of testing. the mean item discrimination is .28. in contrast, for students in the lat the mean item discrimination is .47 and there is only one item with a slight misfit and one item with a strong misfit. note that the item with the strong misfit was also problematic in the main sample. evaluation of the wmnsq, the iccs, as well as of the point-biserial correlations of the responses (see table 2) corroborates these findings. the results show that the items in the reduced test version have a good item fit for students in the lat. they have, however, an insufficient fit in the students with sen-l group. nevertheless, the item fit in the students with sen-l group is better for the reduced test than for the standard test. as in the standard test, eap-reliability was sufficiently high for students in the lowest academic track (rel = 0.850) but it was not sufficient for students with sen-l (rel = 0.525). 5.2.3 easy test the items in the easy test fit the data for both target groups better than the standard test. for students with sen-l there are only four items with a slight misfit and three items with a strong misfit. the mean item discrimination for students with sen-l is .30, while it is .46 for the students in the lat. in the students of the lat group, there is no item with an unsatisfactory discrimination. also the other fit measures evaluated (see table 2) show that the items in the easy test version fit the model in the group of students in the lat südkamp et al   13 | f l r       but show some misfit in the students with sen-l group. the eap-reliability for students in the lowest academic track was high (rel = 0.877), while it was not satisfactory for students with sen-l (rel = 0.600) compared to the other two test versions, the easy test version shows the best model fit for students with sen-l. 5.3 investigation of item misfit we further investigated the occurrence of item misfit based on test characteristics. we did not find any systematic relationship between item misfit and the different dimensions of the conceptual framework of the reading test (text function, cognitive requirements, and item format). however, we did find a relationship between item misfit and item difficulty. 5.3.1 standard test the correlation of the item difficulty estimated in the main sample—thus being independent of the measurement model in the sen-l group—and item discrimination within the students with sen-l group is -.492. the more difficult an item, the lower is the discrimination. this may be an indication of disadvantageous test targeting—that is, inappropriate item difficulties for this target group. the items in the standard test are too difficult for students with sen-l (mean item difficulty with the mean of the reading ability set to zero = 0.58 logits), while item difficulties match the abilities of the students of the lowest academic track well and are in fact rather easy (mean item difficulty = -0.41 logits). here, the correlation between item difficulty estimated in the main sample and item discrimination for students in the lowest academic track is -.324. note that since the measurement model of the standard test in the lowest academic track was estimated based on a subsample of the main sample, estimated item difficulty is not independent of the estimated item discrimination in the sample of students of the lowest academic track in the main sample. 5.3.2 reduced test in the group of students in the lat item fit of the reduced test is not substantively correlated with item difficulty (cor = -0.06) and is considerably negatively correlated in the students with sen-l group (cor = -.43). within students in the lat, there is no relationship between item difficulty and item misfit, while in the students with sen-l group, items with high difficulty show larger item misfit. this may also be a result of the small variance in item discrimination in the group of students in the lat for this test version. test targeting shows that the reduced test is still too difficult for students with sen-l (mean item difficulty = 0.43 logits) but too easy for students in the lower academic track of general education (mean item difficulty = -1.03 logits). 5.3.3 easy test since most of the items in the easy test are not part of the standard test, we did not compute correlations between item difficulty and item fit. however, we did investigate test targeting. in test targeting, the easy test version is also too easy for students in the lat (mean item difficulty = -0.99 logits) and too hard for students with sen-l (mean item difficulty = 0.61 logits). note that the easy test version is even more difficult than the reduced test version. 5.4 measurement invariance 5.4.1 standard test table 3 shows the absolute differences in estimated item difficulties first, between general education students and students with sen-l and second, between students in the lowest academic track and students in südkamp et al   14 | f l r       other tracks of the main sample taking the standard test version. for students with sen-l, negative values in the table indicate a higher item difficulty compared to general education students while positive values indicate lower item difficulty. for students in the lowest academic track, negative values indicate a higher item difficulty for these students compared to students in other tracks in the main sample and positive values indicate a lower item difficulty. table 3 differential item functioning (dif) in the different test versions and student groups differential item functioning item difficulty sen-l lat standard reduced easy standard reduced easy reg50110 -1.909 -1.010 -0.942 -0.304 -0.256 reg50121 -2.814 -1.678 -1.200 -0.320 0.246 reg50122 -2.063 -0.926 -0.800 -0.276 -0.360 reg50123 -2.078 -0.444 -0.848 -0.222 -0.072 reg50124 -2.236 -0.510 -0.930 -0.140 -0.246 reg50125 -2.202 -1.018 -0.752 -0.260 -0.442 reg50126 -1.793 -0.652 -0.512 0.018 0.858 reg50127 -2.173 -0.714 -1.234 -0.414 -0.362 reg50130 -0.805 -0.850 -0.090 -0.068 -0.106 reg50140 -0.148 -0.382 -0.288 -0.132 -0.024 reg50150 0.874 -0.400 0.200 reg50161 0.542 -1.688 -1.388 -0.536 -0.302 reg50162 0.149 -1.020 -0.130 -0.310 -0.686 reg50163 0.035 -0.348 0.422 -0.342 -0.330 reg50164 -0.076 -1.320 -0.774 -0.466 -0.116 reg50165 0.048 -0.302 -0.042 -0.428 -0.368 reg50170 2.351 0.570 -0.294 reg50210 -1.411 -1.054 -0.566 -0.564 -0.352 -0.574 -0,304 reg50220 1.436 1.200 1.602 1.360 0.576 0.414 0,490 reg50230 -1.187 -0.926 -0.814 -0.850 -0.148 0.006 -0.148 reg50240 0.050 -0.082 0.146 0.232 -0.094 -0.044 -0.226 reg50250 0.667 0.164 0.344 -0.096 0.134 0.170 -0.018 reg50261 -1.352 -0.318 -0.204 reg50262 1.924 0.580 -0.038 reg50263 2.159 0.172 0.088 reg50264 2.167 -0.290 0.188 reg50265 2.195 0.724 0.180 reg50266 2.221 1.016 0.116 reg50310 -0.867 -0.824 -1.254 -0.756 -0.318 -0.142 -0.444 reg50320 -1.425 -0.982 -0.870 -0.798 -0.464 -0.196 -0.066 reg50330 -1.185 -1.654 -1.632 -1.020 -0.440 -0.106 -0.154 reg50340 -0.158 -0.570 0.378 0.026 -0.186 0.078 0.030 reg50350 0.838 0.082 0.310 0.420 0.102 0.028 0.222 reg50360 -0.887 -0.844 -0.324 -0.324 -0.274 -0.062 -0.130 (continued) südkamp et al   15 | f l r       item difficulty sen-l lat standard reduced easy standard reduced easy reg50370 0.140 -0.256 0.318 -0.058 0.020 0.288 -0.100 reg50410 0.885 0.370 0.206 reg50421 -0.481 0.042 0.342 reg50422 -0.225 0.380 0.666 reg50423 0.243 1.268 0.536 reg50430 2.371 0.772 0.080 reg50452 0.531 1.586 0.590 reg50440 1.922 1.264 0.374 reg50451 0.183 1.526 0.716 reg50460 1.356 0.436 0.100 reg50510 -0.898 -0.532 -0.618 -0.594 0.044 reg50521 -0.313 -0.052 0.878 -0.366 -0.058 reg50522 -0.635 -0.156 0.966 -0.080 0.416 reg50523 -0.004 0.366 1.066 0.214 0.250 reg50524 -0.634 0.256 0.872 -0.318 0.704 reg50530 1.487 0.770 0.206 reg50540 0.064 -0.262 0.748 -0.428 -0.096 reg50551 -0.035 0.064 -0.756 0.030 -0.090 reg50552 1.135 0.188 1.214 -0.184 -0.108 reg50553 0.385 0.210 0.540 -0.452 0.214 reg50560 1.125 0.716 0.938 0.624 0.666 reg50570 0.515 -0.334 0.392 -0.330 0.206 note. sen-l = special educational needs in learning; lat = lowest academic track. the results clearly show measurement invariance for students in the lowest academic track and large differences in estimated item difficulties for students with sen-l. for students in the lowest academic track, of the 56 items there is no item with strong dif (absolute difference in item difficulties greater than 1) and only three items with slight dif (absolute difference in item difficulties between 0.6 and 1). for students with sen-l there are 12 items with slight dif and 14 items with strong dif. the results indicate that measurement invariance holds for students in the lowest academic track but that the test measures a different construct for the group of students with sen-l compared to general education students. thus, reading test scores for students with sen-l are not comparable to test scores for general education students. 5.4.2 reduced test table 3 also shows dif for the accommodated test versions. in the reduced test, for students with sen-l, 15 out of 38 items have slight dif and eight items have strong dif. only 15 items show no considerable dif. thus, the measurement of reading competence with the reduced test is different from that of general education students with the standard test. this does, however, not seem to be a result of the test accommodation. within the group of students in the lat measurement invariance holds as only three items show slight dif. the results indicate that for students with sen-l the measurement model, and thus, the measured construct, is different from that of students in general education. 5.4.3 easy test in the easy test, for students with sen-l three out of twelve anchor items show noticeable dif and two items show strong dif. there are only seven items with no noticeable dif. in contrast, in the lat group there is no noteworthy dif in the easy test and only four items show slight dif in the reduced test. südkamp et al   16 | f l r       while measurement invariance may be assumed for the students in the lat, it does not hold for students with sen-l. again, differences in the measurement model do not seem to be induced by the test accommodation, but rather reflect a specific testing problem of students with sen-l. 5.5 item fit and measurement invariance considering both criteria—item fit and measurement invariance—how many items with good psychometric properties are left within the different groups and test versions? is it possible to construct a test out of well-fitting items? figure 1 shows the discrimination and dif of the items in the standard test version for students with sen-l (a) and for students of the lowest academic track (b). the grey lines give the rules of thumb for the evaluation of the items. items within discrimination > .2 and absolute dif < 0.6 have no noticeable misfit or dif. items within .2 > discrimination > .1 and 0.6 < absolute dif < 1 have noticeable but not considerable misfit and/or dif. items with discrimination < .1 and absolute dif > 1 have considerable misfit and/or dif. these items should not be used for testing. figure 1a) shows that a considerable amount of items do not meet the fit and dif criteria in the sen-l group. only 22 out of 56 items show good fit and dif indices. thirteen items show a slight misfit in at least one of the two criteria and 21 items exceed at least one of the criteria for a strong misfit or large dif. there are obviously not many items left that meet the criteria of a good test. for students of the lowest academic track (figure 1b), there is only one item with a slight misfit in either of the two criteria and seven items with a strong deviation from at least one of the two criteria. thus, there are 48 items that meet the criteria of a good test in the lowest academic track group of the main sample. a) students with sen-l südkamp et al   17 | f l r       b) students in the lowest academic track of the main sample figure 1. discrimination and differential item functioning of the items in the regular test. sen-l = special educational needs in learning. in the reduced test (see figure 2a), for students with sen-l, 13 out of 38 items show a strong misfit and/or dif, 16 items show a slight deviation from at least one of the two criteria and only nine items are suitable for testing considering both criteria. there are a high number of items that may not be used on a test. again, in the lat group the items fit both criteria very well (figure 2b). only one out of 38 items needs to be excluded due to strong misfit or dif, and only three items show a slight misfit and/or dif. thirty-four items meet the criteria of fit and measurement invariance. the low dif values in the lat group provide evidence in support of the argument that reducing the test length (i.e., increasing the testing time per text and item) does not threaten the comparability of the results. thus, reducing test length may be an appropriate accommodation. however, this accommodation is not sufficient to reliably and comparably measure reading competence for students with sen-l. a) students with sen-l südkamp et al   18 | f l r       b) students in the lowest academic track figure 2. discrimination and differential item functioning of the items in the reduced test. sen-l = special educational needs in learning. since there are only 12 items in the easy test that may be tested for dif, we refrained from plotting the different evaluation criteria for this test version. it may, however, be concluded that from the 12 items, there are four with a slight misfit or dif and two with a strong one. only six of the 12 anchor items meet the criteria of fit and dif. since linking may only be done using 12 items, losing six items due to fit and dif problems raises questions as to the appropriateness of this accommodated test version for the group of students with sen-l. as a comparison, in the lat group there is only one of these 12 items with a slight misfit and one with a strong misfit. the results in the lat group are an indication that reducing the difficulty of the test does result in reliable and comparable reading competence measures. however, this test accommodation is not appropriate enough for assessing students with sen-l. 6. discussion the present research dealt with the question of how competencies of students with sen-l may be assessed reliably and comparably to general education students. we assessed the reading competence of students with sen-l using a standard reading test, a reduced reading test, and an easy reading test. we used a group of low-achieving students without sen to test whether the test accommodations alter the measured construct. the results showed that all three reading test versions are suitable for a reliable and comparable measurement of reading competence in students without sen. reducing both test length and item difficulty resulted in reliable measures that are comparable to those of a standard test for general education students. for students with sen-l, the accommodated test versions considerably reduced the amount of missing values. they did not, however, show a satisfactory item fit and measurement invariance. although the testing accommodations increase item fit and measurement invariance for students with sen-l as compared to using a standard reading test, there are still many items unsuitable for a reliable and comparable assessment of reading competence in students with sen-l. thus, the competence scores assessed by the tests in this study are neither suitable for a substantive interpretation of the competence level of students with südkamp et al   19 | f l r       sen-l, nor may they be used for a valid comparison of competence levels between students with sen-l and students in general education. concerning the testing accommodations implemented in this study, the reduced test primarily aimed at compensating for information-processing restrictions in students with sen-l (e.g., for slow processing speed) while the easy test primarily aimed at adapting the test to a reduced competence level in reading (by reducing test difficulty in general) thereby improving the accuracy of measurement and avoiding undue frustrations for students with sen-l. since we showed—within the group of students in the lat—that the items in the accommodated test versions have a good fit, we may conclude that the misfit in the group of sen-l students is not due to badly constructed items or to the fact that the test versions changed the measured construct. misfit of items in the sen sample must be due to problems in testing this specific target group. our analyses on test targeting showed that even the accommodated test versions are too difficult for students with sen-l. since item fit became better for accommodated versions, which were composed of easier items than the standard test, we hypothesize that a further reduction in item difficulty may help to improve testing of students with sen-l. this hypothesis is corroborated by the negative correlation of item difficulty and discrimination. still, both testing accommodations focus on general problems faced by students with sen-l when reading (slow processing speed, reduced competence level in reading). in future research, it would be desirable to identify more specific reading problems of students with sen-l that can be addressed in testing accommodations. another explanation for item misfit in the sample of students with sen-l may lay in the test-taking behavior (such as guessing or item omission, see pohl, südkamp, hardt, carstensen, & weinert, 2015). it is also possible that differences in item fit between the students in the lat and the students with sen-l are due to differences in school curricula. comparing the three test versions—the standard test, the reduced test, and the easy test—in the lat group, the accommodated test versions resulted in better competence measures than the standard test. for students with sen-l, the easy test showed the best results regarding item fit, test targeting, and dif. since in the reading test, items are grouped to sets belonging to different texts, constructing a reading test from wellfitting and measurement invariant items is a difficult encounter. this is different in other competence domains of the neps that do not have such a strong testlet structure (see weinert et al., 2011, for a description of the tests). 6.1 strengths and limitations studying the effects of testing accommodations not only in groups of students with sen-l but also in groups of students in general education (here: low-performing students), is a promising approach to the identification of appropriate testing accommodations. in many previous studies, accommodated test versions were only applied to students with sen. thus, one could not disentangle whether low psychometric properties of accommodated tests and change of the measured construct were due to testing accommodations or testability problems of students with sen. using the lat group allowed us to investigate whether the applied testing accommodations generally provide reliable and measurement invariant measures of reading competence. with the results in the group of lat students, we ruled out the premise that misfit and measurement invariance for students with sen-l is due to changes in the measured construct resulting from a reduction in test length or reduction in item difficulty. considering the wide range of competence levels of students in general education, students in the lat are the group of students without sen being closest in competence level to students with sen. thus, the accommodated test versions—that are targeted towards students with sen-l—will still be better targeted to students in the lat than to all students in general education. the study’s strength also lies in the use of a sophisticated methodological approach and the evaluation of various measures of item fit in addition to differential item functioning. when using methods of irt, other studies on the assessment of students with sen mainly report dif but leave out information on item fit in the sample of students with sen (abedi, leon, & kao, 2008; bolt & ysseldyke, 2008). südkamp et al   20 | f l r       considering the group of students with sen-l, using data from a relatively large representative sample allows us to draw credible conclusions. however, our samples of students with sen-l and students in the lat group considerably differed in their size. there were about twice as many students in the lat group compared to the students with sen-l group. for some testing conditions the sample was comparatively small. for example, only 84 students with sen-l were assessed with the reduced test version. due to the large number of missing responses, there were items with just 52 valid responses. fit and dif measures may, as a consequence, be unreliable. we tried to account for this in the evaluation of the fit and dif criteria. one might also argue that the group of students with sen-l is still a highly heterogeneous one, including, for example, students with different performance and ability profiles in the cognitive domain. compared to prior research, however, the target population is rather homogeneous as students with sen in areas other than learning (e.g., those with physical impairments) are precluded. other studies investigated appropriateness of competence assessments on even more heterogeneous groups of students (e.g., lutkus et al., 2004, including students with disabilities in general). possible testing problems may, however, only occur for students with specific disabilities (e.g., for students with sen-l, but not for students with visual impairments) or for specific testing accommodations. analyzing the whole group of students with disabilities and running analyses across all types of testing accommodations may mask possible testing effects. in our study we focused on a specific group of students with sen and analyzed different testing accommodations separately. item misfit and dif do not need to be caused by all students with sen-l, but only by a certain group of students. however, we did not account for interindividual differences within our samples in this study. in ongoing research, we (pohl et al., 2015) use a person-based approach and try to empirically identify groups of students with sen-l whose assessment is especially challenging. here, we assume that individual student characteristics (e.g., individual test taking strategies, cognitive performance profiles) are related to testability2. 6.2 implications and future research incorporating easy instead of hard items in the test version (e.g., as done in the easy test version), is methodologically seen a form of adaptive testing. adaptive testing is currently discussed in large-scale studies such as the naep (xu, sikali, oranje, & kulick, 2011), the programme for international student assessment (pisa; pearson, 2011), and the neps (pohl, 2014). if better test targeting is one of the key issues for testing students with sen-l, adaptive testing procedures for general education students may well be extended to include students with sen-l. one way to systematically reduce difficulty in reading tests for students with sen might be a reduction in grammatical and lexical complexity of texts and items (abedi et al., 2011). in upcoming feasibility studies within the neps, seventh graders with sen-l will be tested with a standard reading test that is reduced in grammatical and lexical complexity. in another feasibility study in grade 3 we will examine the effects of newly developed test instructions on students’ test performance, missing values, and invalid answers, as well as on their motivation, and test anxiety. there are numerous and manifold arguments for the inclusion of students with sen in large-scale assessments. however, the issue of whether students with sen-l may be assessed reliably and comparably in large-scale assessments —and if so how—remains to be an important and complex question. in our study, we aim to present a sophisticated design and a comprehensive methodological approach to these questions                                                                                                                           2 in the present study, differences in test taking in students with and without sen-l might also be caused by differences in school curricula. this alternative hypothesis could be tested by comparing students with sen-l attending general education and special schools. however, in germany only few students with sen-l attended general education schools at the time of data collection and these students often differ in individual as well as in social background characteristics from students attending special schools. südkamp et al   21 | f l r       and to shed light on them. we think that the systematic identification of specific testing accommodations for groups of students with sen is a promising approach. keypoints so far, data on the acquisition and development of competencies of students with special educational needs in learning (sen-l) are rare. assessing competencies of students with special educational needs within large scale assessments is challenging. this study addresses the question of whether and how satisfying item fit measures and measurement invariant test scores can be obtained for students with sen-l in large-scaleassessments. testing accommodations may result in reliable and to the standard test comparable competence measures. the investigated testing accommodations helped to some extent to increase the testability of students with sen-l. the systematic identification of further appropriate testing accommodations is a promising approach to the assessment of students with sen-l. acknowledgments this paper uses data from the national educational panel study (neps). from 2008 to 2013, neps data were collected as part of the framework program for the promotion of empirical educational research funded by the german federal ministry of education and research (bmbf). as of 2014, the neps survey is carried out by the leibniz institute for educational trajectories (lifbi) at the university of bamberg in cooperation with a nationwide network. we especially thank cordula artelt, claus h. carstensen, jana heydrich, lena nusser, and markus messingschlager for their contribution to this study. our thanks also go to the staff of the neps administration of surveys and to the methods group. we would also like to thank the anonymous reviewers for their comments on earlier versions of the manuscript and erika fisher for copy editing services. references abedi, j., leon, s., & kao, j. (2008). examining differential item functioning in reading assessments for students with disabilities. (cresst report 744). los angeles, ca: university of california, los angeles, national center for research on evaluation, standards, and student testing. abedi, j., leon, s., kao, j., bayley, r., ewers, n., herman, j., & mundhenk, k. (2011). accessible reading assessments for students with disabilities: the role of cognitive, grammatical, lexical, and textual/visual features (cresst report 785). los angeles, ca: university of california, los angeles, national center for research on evaluation, standards, and student testing. südkamp et al   22 | f l r       american educational research association, american psychological association, & national council on measurement in education (1999). standards for educational and psychological testing. washington, dc: american educational research association. aßmann, c., steinhauer, h. w., kiesl, h., koch, s., schönberger, b., müller-kuller, a., … blossfeld, h.-p. (2011). sampling designs of the national educational panel study: challenges and solutions. zeitschrift für erziehungswissenschaft, 14, 51-65. doi:10.1007/s11618-011-0181-8 aßmann, c., steinhauer, h. w., & zinn, s. (2012). weighting the fifth and ninth grader cohort samples of the national educational panel study, panel cohorts (technical report). bamberg, germany: university of bamberg national educational panel study, retrieved from https://www.nepsdata.de/portals/0/neps/datenzentrum/forschungsdaten/sc3/1-0-0/sc3_sc4_1-00_weighting_en.pdf. bäumer, t., preis, n., roßbach, h.-g., stecher, l., & klieme, e. (2011). education processes in life-coursespecific learning environments. zeitschrift für erziehungswissenschaft, 14, 87-101. doi:10.1007/s11618-011-0183-6 barkow, i., leopold, t., raab, m., schiller, d., wenzig, k., blossfeld, h.-p., & rittberger, m. (2011). remoteneps: data dissemination in a collaborative workspace. zeitschrift für erziehungswissenschaft, 14, 315-325. doi: 10.1007/s11618-011-0192-5 bielinski, j., thurlow, m. l., ysseldyke, j. e., freidebach, j., & freidebach, m. (2001). read-aloud accommodations: effects on multiple-choice reading and math items (nceo technical report 31). minneapolis, mn: university of minnesota, national center on educational outcomes. blossfeld, h.-p., & von maurice, j. (2011). education as a lifelong process. zeitschrift für erziehungswissenschaft, 14, 19-34. doi:10.1007/s11618-011-0179-2 blossfeld, h.-p., von maurice, j., & schneider, t. (2011). the national educational panel study: need, main features, and research potential. zeitschrift für erziehungswissenschaft, 14, 5-17. doi:10.1007/s11618-011-0178-3 bolt, s. e., & ysseldyke, j. (2008). accommodating students with disabilities in large-scale testing: a comparison of differential item functioning (dif) identified across disability types. journal of psychoeducational assessment, 26, 121-138. doi:10.1177/0734282907307703 borsboom, d. (2006). the attack of the psychometricians. psychometrika, 71, 425-440. doi: 10.1007/s11336-006-1447-6 bos, w., bonsen, m., gröhlich, c., guill, k., may, p., rau, a., et al. (2009). kess 7: kompetenzen und einstellungen von schülerinnen und schülern—jahrgangsstufe 7 [kess 7: competencies and attitudes of students in grade 7]. hamburg, germany: behörde für bildung und sport. chudowsky, n., & pellegrino, j. (2003). large-scale assessment that support student learning: what will it take? theory into practice, 42, 75-83. doi:10.1207/s15430421tip4201_10 cormier, d. c., altman, j., shyyan, v., & thurlow, m. l. (2010). a summary of the research on the effects of test accommodations: 2007-2008 (technical report 56). minneapolis, mn: university of minnesota, national center on educational outcomes. cortiella, c., & horowitz, s. h. (2014). the state of learning disabilities: facts, trends and emerging issues. new york: national center for learning disabilities. dolan, r. p., & hall, t. e. (2001). universal design for learning: implications for large-scale assessment. ida perspectives, 27, 22-25. duke, n. k., & pearson, p. d. (2002). effective practices for developing reading comprehension. in a. e. farstrup & s. j. samuels (eds.), what research has to say about reading instruction (pp. 205–242). newark, de: international reading association. durkin, d. (1993). teaching them to read. boston, ma: allyn and bacon. fuchs, l. s., fuchs, d., eaton, s. b., hamlett, c. l., & karns, k. m. (2000). supplementing teacher judgments of mathematics test accommodations with objective data. school psychology review, 29, 65–85. südkamp et al   23 | f l r       gee, j. p. (2004). reading as situated language: a sociocognitive persepective. in r. b. ruddell & n. j. unrau (eds.), theoretical models and processes of reading (pp. 116-132). newark: international reading association. gehrer, k., zimmermann, s., artelt, c., & weinert, s. (2013). neps framework for assessing reading competence and results from an adult pilot study. journal of educational research online, 5, 50-79. gehrer, k., zimmermann, s., artelt, c., & weinert, s. (2012). the assessment of reading competence (including sample items for grade 5 and 9) [scientific use file 2012, version 1.0.0.] bamberg: university of bamberg, national educational panel study. geisinger, k. f. (1994). psychometric issues in testing students with disabilities. applied measurement in education, 7, 121-140. doi:10.1207/s15324818ame0702_2 gersten, r., fuchs, l. s., williams, j. p., & baker, s. (2001). teaching reading comprehension strategies to students with learning disabilities: a review of research. review of educational research, 71, 279-320. doi:10.3102/00346543071002279 gross, c., jobst, a., jungbauer-gans, m., & schwarze, j. (2011). educational returns over the life course. zeitschrift für erziehungswissenschaft, 14, 139-153. doi:10.1007/s11618-011-0195-2 grünke, m. (2004). lernbehinderung [learning disabilities]. in lauth, g., grünke, m., & brunstein, j. (eds.). interventionen bei lernstörungen [interventions to learning deficits](pp. 65-77). göttingen: hogrefe. heydrich, j., weinert, s., nusser, l., artelt, c., & carstensen, c. h. (2013). including students with special educational needs into large-scale assessments of competencies: challenges and approaches with the german national educational panel study (neps). journal of educational research online, 5, 217240. hollenbeck, k. tindal, g. almond, p. (1998). teachers’ knowledge of accommodations as a validity issue in high-stakes testing. the journal of special education, 32, 175-183. kavale, k. a., & reece, j. h. (1992). the character of learning disabilities. learning disability quarterly, 15, 74-94. doi: http://dx.doi.org/10.2307/1511010 kintsch, w. (2007). comprehension: a paradigm for cognition. cambridge, uk: cambridge university press. kmk – sekretariat der ständigen konferenz der kultusminister der länder in der bundesrepublik deutschland [standing conference of the ministers of education and cultural affairs of germany] (2012). sonderpädagogische förderung in schulen 2001–2010 [special education in schools 2001– 2010]. retrieved from http://www.kmk.org/fileadmin/pdf/statistik/komstat/dokumentation_sopaefoe_2010.pdf kolen m. j., & brennan r. l. (2004). test equating, scaling, and linking. new york, ny: springer-verlag. koretz, d. m. (1997). the assessment of students with disabilities in kentucky (cse technical report 431). los angeles, ca: cresst/rand institute on education and training. koretz, d. m., & barton, k. e. (2003). assessing students with disabilities: issues and evidence (cse technical report 587). los angeles, ca: university of california, center for the study of evaluation. kristen, c., edele, a., kalter, f., kogan, i., schulz, b., stanat, p., & will, g. (2011). the education of migrants and their children across the life course. zeitschrift für erziehungswissenschaft, 14, 121-137. doi:10.1007/s11618-011-0194-3 lovett, b. j. (2010). extended time testing accommodations for students with disabilities: answers to five fundamental questions. review of educational research, 80, 611-638. doi:10.3102/0034654310364063 lutkus, a. d., mazzeo, j., zhang, j., & jerry, l. (2004). including special-needs students in the naep 1998 reading assessment part ii: results for students with disabilities and limited-english proficient students (research report ets-naep 04-r01). princeton, nj: ets. millsap, r. e. (2011). statistical approaches to measurement invariance. new york, ny: routledge. minnema, j., thurlow, m., bielinski, j., & scott, j. (2000). past and present understandings of out-of-level testing: a research synthesis. (out-of-level testing project report 1). minneapolis, mn: university of minnesota, national center on educational outcomes. retrieved from http://education.umn.edu/nceo/onlinepubs/oolt1.html südkamp et al   24 | f l r       müller, k., sälzer, c., mang, j., & prenzel, m. (2014, march). kompetenzen von schülerinnen und schüler mit besonderem förderbedarf. ergebnisse aus dem pisa 2012 förderschul-oversample [competencies of students with special educational needs. results from the pisa 2012 oversample of special schools]. paper presented at the conference of the german association for empirical educational research, frankfurt, germany. oecd – organisation for economic co-operation and development. (1999). measuring student knowledge and skills: a new framework for assessment. paris, france: oecd. pearson (2011, october 7th). pearson to develop framework for oecd’s pisa students assessment for 2015 [pearson announcement]. retrieved from http://www.pearson.com/news/2011/october/pearson-todevelop-frameworks-for-oecds-pisa-student-assessment-f.html?article=true pellegrino, j., chudowsky, n., & glaser, r. (2001). knowing what students know: the science and design of educational assessment. washington, d. c.: national academy press. pitoniak, m. j., & royer, j. m. (2001). testing accommodations for examinees with disabilities: a review of psychometric, legal, and social policy issues. review of educational research, 71, 53-104. doi:10.3102/00346543071001053 pohl, s. (2014). longitudinal multi-stage testing. journal of educational measurement, 50, 447-468. doi: 10.1111/jedm.12028 pohl, s., & carstensen, c. h. (2012). neps technical report: scaling the data of the competence test (neps working paper no. 14). bamberg, germany: university of bamberg, national educational panel study. pohl, s., & carstensen, c. h. (2013). scaling the competence tests in the national educational panel study—many questions, some answers, and further challenges. journal for educational research online, 5, 189-216. pohl, s., gräfe, l., & rose, n. (2014). dealing with omitted and not reached items in competence tests evaluating approaches accounting for missing responses in irt models. educational and psychological measurement, 74, 423-452. doi: 10.1177/0013164413504926 pohl, s., haberkorn, k., hardt, k., & wiegand, e. (2012). neps technical report for reading—scaling results of starting cohort 3 in fifth grade (neps working paper no. 15). bamberg, germany: university of bamberg, national educational panel study. pohl, s., südkamp, a., hardt, k., carstensen, c. h., & weinert, s. (2015). testability and test-taking behavior of students with special educational needs in large-scale assessments. manuscript submitted for publication. popham, w. j. (2000). educational measurement. boston, ma: allyn and bacon. rasch, g. (1960). probabilistic models for some intelligence and attainment tests. copenhagen: nielsen & lydiche (expanded edition, chicago, university of chicago press, 1980). ritchey, k. d., silverman, r. d., schatschneider, c., & speece, d. l. (2015). prediction and stability of reading problems in middle childhood. journal of learning disabilities, 48, 298-309. doi:10.1177/0022219413498116 sireci, s. g., scarpati, s. e., & li, s. (2005). test accommodations for students with disabilities: an analysis of the interaction hypothesis. review of educational research, 75, 457-490. doi:10.3102/00346543075004457 stocké, v., blossfeld, h.-p., hoenig, k., & sixt, m. (2011). social inequality and educational decisions in the life course. zeitschrift für erziehungswissenschaft, 14, 103-199. doi:10.1007/s11618-011-0193-4 swanson. (1999). reading research for students with ld: a meta-analysis of intervention outcomes. journal of learning disabilities, 32, 504-532. doi:10.1177/002221949903200605 thompson, s. j., johnstone, c. j., anderson, m. e., & miller, n. a. (2005). considerations for the development and review of universally designed assessments (nceo technical report 42). minneapolis, mn: university of minnesota, national center on educational outcomes. thurlow, m. l. (2010). steps toward creating fully accessible reading assessments. applied measurement in education, 23, 121-131. doi:10.1080/08957341003673765 südkamp et al   25 | f l r       thurlow, m. l., bremer, c., & albus, d. (2008). good news and bad news in disaggregated subgroup reporting to the public on 2005-2006 assessment results (technical report 52). minneapolis, mn: university of minnesota, national center on educational outcomes. thurlow, m., elliott, j., & ysseldyke, j. (1999). out-of-level testing: pros and cons (policy directions no. 9). minneapolis, mn: university of minnesota, national center on educational outcomes. retrieved from http://education.umn.edu/nceo/onlinepubs/policy9.htm tindal, g., heath, b., hollenbeck, k., almond, p., & harniss, m. (1998). accommodating students with disabilities on large-scale tests: an experimental study. exceptional children, 64, 439–450 u.s. department of education, national center for education statistics. (2013). digest of education statistics, 2012 (nces 2 014-015). verhoeven, l., & van leeuwe, j. (2008). prediction of the development of reading comprehension: a longitudinal study. applied cognitive psychology, 22, 407-423. doi:10.1002/acp.1414 weinert, f. e. (2001). concept of competence: a conceptual clarification. in d. s. rychen, l. & h. salganik (eds.), defining and selecting key competencies (pp. 45-66). seattle: hogrefe & huber. weinert, s., artelt, c., prenzel, m., senkbeil, m., ehmke, t., & carstensen, c. h. (2011). development of competencies across the life span. zeitschrift für erziehungswissenschaft, 14, 67-86. doi:10.1007/s11618-011-0182-7 woodcock, s., & vialle, w. (2011). are we exacerbating students’ learning disabilities? an investigation of pre-service teachers’ attributions of the educational outcomes of students with learning disabilities. annals of dyslexia, 61, 223-241. doi:10.1007/s11881-011-0058-9 wright, b. d., & masters, g. n. (1982). rating scale analysis: rasch measurement. chicago, il: mesa press. wu, m. (1997). the development and application of a fit test for use with marginal maximum likelihood estimation and generalized item response models (unpublished doctoral dissertation). melbourne, australia: university of melbourne. wu, m., adams, r. j., wilson, m., & haldane, s. (2007). conquest 2.0. [computer software] camberwell, australia: acer press. wu, y.-c., liu, k. k., thurlow, m. l., lazarus, s. s., altman, j., & christian, e. (2012). characteristics of low performing special education and non-special education students on large-scale assessments (technical report 60). minneapolis, mn: university of minnesota, national centre on educational outcomes. xu, x., sikali, e., oranje, a., & kulick, e. (2011, april). multi-stage testing in educational survey assessments. paper presented at the annual meeting of the national council on measurement in education (ncme), new orleans, la. yovanoff, p., & tindal, g. (2007). scaling early reading alternate assessments with statewide measures. exceptional children, 73, 184-201. ysseldyke, j. e., thurlow, m. l., langenfeld, k. l., nelson, r. j., teelucksingh, e., & seyfarth, a. (1998). educational results for students with disabilities: what do the data tell us? (technical report 23). minneapolis, mn: university of minnesota, national center on educational outcomes. zebehazy, k. t., zigmond, n., & zimmerman, g. j. (2012). ability or access-ability: differential item functioning of items on alternate performance-based assessment tests for students with visual impairments. journal of visual impairment & blindness, 106, 325-338. table of footnotes 2 as for the term sen-l, the term “learning disabilities” is not clearly defined. note that we refer to a heterogeneous group of students with multifaceted etiology. 3 in the present study, differences in test taking in students with and without sen-l might also be caused by differences in school curricula. this alternative hypothesis could be tested by comparing südkamp et al   26 | f l r       students with sen-l attending general education and special schools. however, in germany only few students with sen-l attended general education schools at the time of data collection and these students often differ in individual as well as in social background characteristics from students attending special schools. microsoft word andres et al_publication.docx         frontline learning research vol.3 no. 3 special issue (2015) 5 22 issn 2295-3159 1 corresponding author: lesley andres, 2125 main mall, university of british columbia | vancouver, bc v6t 1z4, phone +1 604 822 8943, fax +1 604 822 4244, email lesley.andres@ubc.ca, doi: http://dx.doi.org/10.14786/flr.v3i3.177   drivers and interpretations of doctoral education today: national comparisons lesley andresa1, søren s. e. bengtsenb, liliana del pilar gallego castañoc, barbara crossouardd, jeffrey m. keefere, kirsi pyhältöf a university of british columbia, canada b aarhus university, denmark c university of caldas, colombia d university of sussex, uk e new york university, usa f university of oulu and university of helsinki, finland article received 17 may 2015 / revised 17 may 2015 / accepted 16 june 2015 / available online 14 august 2015 abstract in the last decade, doctoral education has undergone a sea change with several global trends increasingly apparent. drivers of change include massification and professionalization of doctoral education and the introduction of quality assurance systems. the impact of these drivers, and the forms that they take, however, are dependent on doctoral education within a given national context. this paper is frontline in that it contributes to the literature on doctoral education by examining the ways in which these global trends and drivers are being taken up in policies and practices by various countries. we do so by comparing recent changes in each of the following countries: canada, colombia, denmark, finland, the uk, and the usa. each country case is based on national education policies, policy reports on doctoral education (e.g., oecd and eu policy texts), and related materials. we use the same global drivers to examine educational policies of each country. however, depending each national context, these drivers are framed in considerably different ways. this raises questions about (1) their comparability at a global level and (2) the universality of the phd. also we find that this global-local nexus reveals unresolved tensions within the national doctoral educational frameworks. keywords: doctoral education, higher education policy, massification, professionalization, quality assurance                                                                                                                             andres et al   | f l r       6   1. introduction globally, research and researchers are viewed increasingly as critical to social and economic competitiveness and societal health (e.g., uk council for science and technology, 2007; european commission, 2014). it follows that over the past quarter of a century, the education of future researchers, principally through doctoral education, has become increasingly valued. as doctoral education shifts from the periphery (e.g., available to a small elite) to a more mainstream trajectory of the total educational experience, it is undergoing a sea change. several global trends and related drivers of such changes can be identified. the forms that the drivers take, however, and their impacts, are dependent on the specific contexts of doctoral education in a given national context. our paper contributes to the literature on doctoral education by examining the ways in which these global trends and drivers are being taken up in policies and practices by various countries. depending on priorities, path dependencies, and openness to change, global trends play out in different ways in given countries. however, countries are also influenced by wider historical, economic, and cultural geopositioning. in this paper, we highlight how drivers and trends have manifested themselves in individual countries. in a six country comparative case study – approach canada, colombia, denmark, finland, the uk, and the usa – we address the following question: what recent changes related to doctoral education in relation to the three drivers and trends identified above can be identified in each country? to address this question, document based cases of doctoral education in each country are presented below. particular attention has been paid to the identification of the most recent policy changes in doctoral education and the ways in which the changes are taken up. drawing on our analysis of each context, we conclude by proposing future research agendas for examining doctoral education. the case countries were selected because they present different cultural geopositionings and traditions of doctoral education, ranging from the more structured and course work based model of the usa to the less structured model in nordic countries. also, the cases present variation in terms of the extent to which the higher education system in a given country is teaching-oriented, – for example, colombia as a highly teaching-oriented system and finland more research oriented – their emphasis on performance based management (e.g., the uk and usa presenting highly performance-based systems, denmark, finland and canada being at the middle and colombia being at the other end), and whether a country’s higher education system is in the process of developing (colombia), recently developed (finland and denmark) or well developed (canada, usa & uk) (shin, 2010; shin & jung, 2014). first, we begin with an overview of the key drivers, followed by country specific descriptions. 2. global drivers of doctoral education core global drivers affecting doctoral education have been identified in the research literature (e.g., kehm, 2006) and in various policy reports (oecd, 2010; 2014; department for education and skills, 2003). these trends include massification of doctoral education, professionalization of doctoral education and careers, and the development of various quality assurance systems. 1.1. massification of doctoral education worldwide, the number of doctoral students and the number of doctoral degree holders has increased significantly. since 2000, the proportion of those who have earned doctoral degrees has risen by 38% from 154,000 new graduates in 2000 to 213,000 new doctoral graduates in 2009 in oecd countries (auriol, misu & freeman, 2013; oecd, 2014). on average in 2009, 1.6% of young people, compared to 1% in 2000, in oecd countries have earned doctoral degrees (oecd, 2014). although graduation rates for women in 2012 (1.5%) at the doctoral level are still somewhat lower than those of men (1.7%), in several countries the expected proportion of women who are expected to graduate is larger based on increased number of women andres et al   | f l r       7   currently undertaking the doctoral studies (oecd, 2014). massification of doctoral education has also increased researcher mobility. in 2010, worldwide about 3.6 million students were enrolled as international students in tertiary education (auriol, misu, & freeman, 2013) and it is assumed that this number will continue to grow (moguerou & di pietrogiacomo, 2008; & rizen & marconi, 2011). in addition to a highly educated work force, rapid increases in the number of doctoral degree holders have resulted in an unequal balance across disciplines. for instance, the number of doctoral degree holders in the majority of oecd countries is significantly higher is natural sciences than in humanities (oecd, 2014). also, there is considerable variation in gender representation of doctoral degree holders across countries; as such, these figures mask substantial differences in the gender balance across different disciplines. at the phd level, education, health, and welfare and the humanities continue to be female dominated; male phds are predominant in science, mathematics and computing, and particularly in engineering, manufacturing, and construction (oecd, 2012). there is some evidence that outside of academia, labour markets have not been able to fully absorb these highly qualified individuals kehm (2006). in general, however, high employment rates between 93 to 99% have been reported among individuals possessing doctoral degrees. in most countries, employment rates of male doctoral degree holders slightly exceed those of females and male doctoral degree holders have higher earnings than their female counterparts (auriol, misu, & freeman, 2013). 1.2. professionalization of doctoral education considering the rise in the number of doctoral degree holders, it is evident that not all will be able to pursue careers in academia, nor should they be assumed to desire this. based on a comparison of oecd countries, doctoral degree holders in the natural sciences and engineering are more likely to be engaged in research, while social scientists are likely to find more opportunities in non-research occupations (auriol, misu, & freeman, 2013). given that research skills are also now seen as being valuable to a broad range of employment sectors, a current driver is therefore the perceived need to better prepare doctoral students to work outside of academia through emphasizing more strongly the acquisition of “generic skills” in doctoral education (eua, 2009; 2010; fiske, 2011; gilbert, balatti, turner & whitehouse, 2004; oecd, 2012). doctoral degree holders are considered to have the potential to contribute to economic growth, advancement, and diffusion of knowledge and technologies, and to solve societal and environmental problems (auriol, schaaper & felix, 2012). research, particularly in engineering, sciences, and medicine, is expected to result in innovations that will increase national competitiveness. also, researchers are expected to participate in turning scientific discoveries into patents and innovations. hence, fostering an entrepreneurial culture by instilling the skills and attitudes needed for creative enterprises is suggested to be a central part of 21st century researcher competence (oecd, 2010). this is driven by (1) an increased number of doctoral students, (2) an agenda to create “free flow of knowledge,” (3) accountability demands, such as reducing the time spent earning the degree, and (4) the goal of lowering levels of attrition among doctoral students. for instance, in europe the berlin communiqué, 2003 and bucharest communiqué, 2012 have espoused professionalization of doctoral education including emphasising learning generic skills yet, a comparison of 19 oecd countries shows that government policies typically emphasise general researcher development, employability of researchers in academia, and improving research work rather than explicitly transferable skills in doctoral education (oecd, 2012). somewhat paradoxically, this agenda sits alongside a perceived need to develop more comparable and structured doctoral programs, which suggests increasing standardisation and routination of programs of study. 1.3. quality assurance in knowledge-based economies, knowledge production has become a commoditized and strategic resource (fernandez-zubieta & guy, 2010; kehm, 2006). the impact of global competition has resulted in a greater emphasis on evaluating the quality of research (adras 2011). frequent evaluation is seen as a means to meet the demands of greater transparency to the public and accountability of research organizations andres et al   | f l r       8   (edler, georhiou, blind & uyrra 2012). many western countries have adopted higher education policies such as systematic benchmarking and research evaluation of universities, including doctoral education, as a means of quality assurance (e.g. buela-casal, gutierrez-matinez, bermudez-sanchez & vadillo-munzo 2007). principal methods used in quality assurance are peer review, high volume bibliometric data (geuna & martin 2003), or a combination of these methods (e.g., informed peer review). quality assurance has resulted in the burgeoning of global ranking schemes that have contributed to the intensification of institutional hierarchies. also, the role of strategic alliances and competitive advantages – among market areas, countries, universities, and even individuals – has become an increasingly important asset in research. as knowledge producers, doctoral students are recognized as increasingly important societal and economic assets. a downside of this is that practices such as poaching highly qualified people who travel abroad from developing countries to earn doctoral credentials is on the increase (auriol, schaaper & felix, 2012; oecd, 2014). 3. research design the paper focuses on the exploration of global drivers of doctoral education and their local manifestation by using a comparative case study strategy (yin, 2012). each country case is based on national education policies, policy reports on doctoral education (e.g., oecd and eu policy texts), and related materials. based on similarities and differences in terms of recent changes in the area of doctoral education in each country (hsieh & shannon, 2005), changes related to massification, professionalization and quality assurance were most frequently reported. accordingly, our comparison focuses on addressing these three trends. 4. country cases each country invoked different ways in which trends have unfolded. hence, each country case was analysed according to its most predominant trends. to provide readers a systematic overview, we conclude this analysis by summarising the findings in table 1. 4.1. canada in recent years, it has been recognized that canada needs more individuals educated at the phd level. according to the conference board of canada (2014), “highly skilled people [i.e., phd graduates] are key to the creation, commercialization, and diffusion of innovation” (p. 1). yet, since 1998, canada has earned a “d” in the multi-country rankings of phd graduates. in 2010, canada was ranked 15th out of 16 in terms of numbers of graduated phds. this suggests the need toward, rather than away from massification of doctoral programs and graduates. however, coordinated efforts to change the course of phd education are difficult because of canada’s decentralized education system. in terms of phd studies, responsibility for education – including higher education – rests with the provinces,. the primary influence of the federal government on increasing the number of phd graduates is through the awarding of doctoral scholarships. regarding phd funding, rather than providing moderate scholarships to many students, currently the trend is to award a select few with “winner take all” super-scholarships (frank, 1999; tamburri, 2013) the federal government has moved away from a more equitable playing field to one of promoting academic “stars” housed in institutions of “excellence,” which seems to be at odds with the goal of increasing the number of phd graduates. other types of funding, for example, by the universities themselves (e.g., through teaching assistantships) and faculty research grants, are not guaranteed and are disproportionally available across andres et al   | f l r       9   disciplines. hence, some students may spend their entire doctoral careers with little or no financial support. several drivers for the need to re-imagine the phd can be found in both policy documents and in the academic literature, including lengthy time to completion, limited or uneven funding opportunities, disappointing completion rates, allegedly antiquated forms of assessment (i.e., the traditional doctoral dissertation), oversupply in some disciplines, demand for skilled workers – highly qualified personnel (hqp) – in a knowledge society, and a poor employment outlook within academia (elgar, 2003; institute for the public life of arts and ideas, 2013, tamburri, 2013). however, the demand for highly qualified personnel could be argued to be the strongest driver of change. the policy headlights appear to be aimed most strongly on changes that will produce labour market-ready workers – in other words, professionalisation of the phd – who will be employed outside of the tenure track framework. however, the discourse around preparation for the labour force and related “skill” acquisition is rather is messy and often contradictory. labour marketready skills can include critical thinking, creativity, and effective communication skills. others believe that internships, professional development programs, partnerships with businesses and industry external to the university are needed to expand the skill repertoires of phd students. one recent report that emerged out of a re-imagining exercise, provided the following criticism: “rather than simply supplementing the student experience with additional opportunities, doctoral programs need to re-think their pedagogical aims and methods at the most fundamental level” (ubc graduate and postdoctoral studies, 2014). however, the absence of national quality assurance mechanisms beyond implicit checks and balances within and among universities (e.g., comprehensive examinations, examination of the dissertation by external assessors) create challenges for re-imagining exercises. in terms of massification of the phd within the canadian context, it is paradoxical that (1) more phd graduates are required; (2) scholarships are awarded to a small proportion of phd students; and (3) there appears to be a glut of phds in terms of employability within academia. hence, the re-imagining process will be a long and contentious process in canada. time will tell whether re-imagining the phd as just-in-time training for the workforce can in any way successfully supplant previous educational ideals such as newman’s notion of education as an end in itself or humboldt’s conceptualization of bildung – that is, cultivation of the entire individual. 4.2. colombia in the 1990s, with the introduction of law 30 (general law of education, 1994), colombia experienced the second highest increase of latin american countries, at 150%, in university (undergraduate and graduate) attendance. however, this increase lagged behind the mean achieved by oecd countries in the same period of time; only 6% of the population in colombia continued their studies and entered to phd programs. in the 1960s and 1970s, the need to promote doctoral studies was identified and one of the first attempts of the government to ameliorate the problem took place in 1968 with the creation of the national institute to promote science and technology (colciencias). additionally, to encourage high quality assurance of future phd candidates from that time onward, the national ministry of education (república de colombia, ministerio de educación nacional, 2010) invested large amounts of money to train colombian doctoral students abroad. however, as in other countries, such a mobility policy generated considerable “brain drain” and most students remained in their host countries because of better professional opportunities. massification of doctoral programs has occurred in many developed countries. however, this was not the case in colombia as national doctorate programs only began to appear some decades ago. thus, between 1986 and 1990, only nine programs were in existence; between 1997 and 2001 this increased to 14 doctoral programs (national council of acreditation cna, 2010). during that time only 2% of university professors held doctoral degrees. in the year 2001 for instance, only 26 individuals had completed doctoral degrees in colombian universities; that is, a very low rate of only four graduates per 1,000,000 people. the world bank (2003) predicted that globalization and economic growth policies would positively affect growth, professionalization, and the development of tertiary education in colombia during 2001 and would andres et al   | f l r       10   lead to a greater number of people graduating with phds. in 2008, around 100 people had graduated from doctoral programs. today, there are 92 doctoral programs in colombia officially reported by the national council of accreditation, cna (national system of innovation in higher education; 2008; unesco-ibe, 2011), with more in natural sciences and mathematics, social sciences, education, and humanities than in engineering, health sciences, and economics (jaramillo, 2009). of these, 52% of doctoral programs are offered by private institutions. hence, doctoral studies are still available only for a small elite and the low availability of doctoral programs in some areas has led some professionals to choose a doctorate not with the goal of mastering an area related to their own field, but only in order to gain access to good jobs. additionally, there is a lack of employment opportunities after graduation because funds provided by government for financing state universities and the opening of places for full-time faculty are not enough to meet national demand. to assure the quality of programs, the government has adopted strategies such as creating and designing regulations (curricular, administrative and academic) and regulatory institutions (cesu, snies, cndm, cna, icfes, among others). however, with so many institutions assigned to assure quality, overlap of functions has the potential to interfere negatively with the flow and development of doctoral programs which differ a great deal from one another (brunner, 2001). all of the work undertaken regarding colombian doctorate education has led to gradual and positive academic development. however, tensions regarding the existing dichotomy between promoting the creation of more doctorates while not addressing the parallel necessity of creating opportunities for employment of alumni exist. the other tension has to do with giving more importance to the regulation of programs rather than for the preparation of academic communities to develop new ways to teach and conduct research. 4.3. denmark in response to the rapid increase in doctoral students at danish universities during the 1990s, denmark created its first graduate schools in 1996. the university act of 2007 required the establishment of graduate schools at all danish universities. the purpose of mandatory graduate schools was to enhance the quality of doctoral education, including optimizing completion rates and standardizing doctoral education across universities (danish ministry of higher education and science, 2014). with the finance act of 2005 and the globalization agreement of 2006, the danish government decided to double the annual enrollment rate of doctoral students from 1,200 in 2003 to 2,400 in 2010. since then, universities have maintained high enrollment rates and today around 2,400 doctoral students are enrolled annually (danish ministry of higher education & science 2015a). the development of doctoral education in denmark is part of a wider european trend of more closely aligning research and doctoral education at the local universities with national and international “policy making and regulation through qualifications framework, benchmarking and evaluation” (fortes, kehm, & mayekiso, 2014, p. 100). together with most of the nordic countries, doctoral education in denmark has been reformed recently “involving a clear trend towards programmed teaching (a more heavy reliance on generic phd courses for example,) and all of the countries are participating in the bologna process for the creation of ehea, the european higher education area” (gudmundsson, 2008, p. 86), which is a body “meant to ensure more comparable, compatible, and coherent systems of higher education in europe” (european higher education area, 2015). as fortes, kehm and mayekiso (2014) point out, the tendency towards increase in “quality assurance at the european level should not be underestimated” (p.100) in terms of the fact that policy making at the european level highly influences and informs national policies on doctoral education in denmark. fortes, kehm and mayekiso highlight that despite the fact that the locus of doctoral education and its curricular content is a national issue, the european commission “acts as a true policy entrepreneur” (p. 100) by specifying agendas and encouraging regulation at the european level. in denmark, the ministry urges universities to ensure that their doctoral programs promote interdisciplinary training and the development of transferrable skills, thus meeting the needs of the wider employment market (gudmundsson, 2008, p. 77). however, at the same time the ministry states that “[o]verregulation of andres et al   | f l r       11   doctoral programs should be avoided” as doctoral education is seen as “a source for human capital for research but is also an extremely important part of the research itself” (p. 77). the danish ministry of higher education and science foregrounds the importance of the european qualifications framework (eqf) and the discourse of lifelong learning with the aim to align the quality and level of doctoral education internationally (danish ministry of higher education & science 2015b). with the eqf, it is possible to compare educational systems, increase mobility across borders, and more fully to internationalize danish universities. this can be said to increase competition among universities, which is seen in the benchmarking systems and the global ranking systems in relation to which the danish universities navigate. the eqf’s effect on doctoral education in denmark has been to promote formalised generic skills and competences within research, development, and teaching at universities. the goals advanced by the ministry focus on “better quality and better cohesion in higher education; even more quality and relevance in research; increased use and dissemination of knowledge and technology; improve[ment] of internationalisation of higher education, research and innovation; increased innovation in businesses, public institutions and higher education, and effective administration of education support and grants” (european commission, 2014). this development points to some potential tensions including a dual focus on wider employment for the market and development of deep research skills necessary for academic environments specifically, together with an increased focus on internationalization and mobility and while attempting to build strong research environments at home universities in denmark. also, the dual goal of increasing training programs and support systems to anchor doctoral education more closely to the home institutional structure and the wish to enhance mobility and independence of individual doctoral students creates another tension. 4.4. finland massification of doctoral education has been driven by the needs of a knowledge economy and national innovation policy and has been promoted systematically by the ministry of education and culture (mec) that provides the primary source of funding for the universities in finland. accordingly, between the 1990s and 2010 the number of doctoral degrees completed annually tripled. currently, about 1600 doctoral degrees are awarded annually. the number of degrees completed yearly is highest in medicine, natural, and technical sciences. although half of doctoral degrees are awarded to women, there are still some gendered disciplinary differences (auriol, misu & freeman, 2013; kota-national data base, 2009; puhakka & rautapuro, 2013). doctoral education has become more mainstream and at the same time researcher mobility has become increasingly important in national doctoral education policy. one result is an increased number of international doctoral students. to promote this inflow, the mec provides financial support to universities to attract international doctoral students earning their degrees in finland. however, the proportion of foreigners in doctoral training is still relatively low (14.8%). also, the outflow of finnish doctoral students is slightly higher than the inflow of international doctoral students studying in finland (garam, 2013). the need to provide a highly skilled workforce for labour markets and the need to improve the quality of doctoral education has led to increasing professionalization of doctoral education (niemi et al, 2011; the graduate school working group, 2012). this resulted in the introduction of more structured forms of doctoral education, that is, the launching of a doctoral school system funded by the academy of finland (finnish ministry of education, 1997). however, by 2010 only about 50% of the doctoral student population studied in these selected doctoral schools. in 2011, a national graduate school system reform was implemented that reversed this and as a result, most universities adopted a single graduate school model to support systematic doctoral education. now all doctoral students belong to a doctoral school in their university and to one of the university’s doctoral programs. there are no tuition fees, but funding for doctoral studies is not automatically provided by, for example, the universities, projects, or foundations for the doctoral students. as a result, some students receive little or no financial support. despite taking a stance towards a more structured system, doctoral studies are still highly research intensive rather than course centred (niemi et al, 2011). to promote the attractiveness and predictability of researcher careers, a four stage researcher career model (first stage being completion of doctoral degree, followed by 2-5 year post andres et al   | f l r       12   doctoral fellow that paves the way for becoming an independent researcher, and finally professorships and research directorships in the final stage) has been introduced (academy of finland, 2010). also, a tenure track system that aims to promote the shift between stages three and four has been introduced. the employment rate of the doctoral degree holders is extremely high 97.6% (treuthardt, & nuutinen, 2012) and the majority (about 80%) work at the universities or research institutions in finland (sainio, 2010; the graduate school working group, 2012). this may explain why, despite the emphasis on learning transferable skills in doctoral education policy documents (academy of finland, 2010; oecd, 2012), efforts to ensure and support work/life relevance have still remained somewhat minor at universities (niemi et al, 2011). the bologna process and adaptation to the european qualifications framework (eqf) to increase the potential to promote international mobility and to facilitate equal participation in european doctoral programs (berlin communiqué, 2003; bucharest communiqué, 2012; european commission, 2014) has resulted in the enhancement of quality assurance in finnish doctoral education (the graduate school working group, 2012) and engagement in international benchmarking and global ranking systems. quality assurance developments have included setting the target doctoral completion time at four years of full-time study; however, time to graduation has remained almost unchanged at six to seven years (sainio, 2010), also, launching the finnish higher education evaluation council that carries out audits of quality systems of the universities and assists universities in thematic and research evaluations, including doctoral education, is another development. 4.5. united kingdom even before its inclusion in the bologna qualifications framework, uk doctoral education had emerged as an area of some interest to policy makers. this phenomenon can be related to the growing significance attached to the knowledge economy and to doctoral education as a training ground for professional researchers, both within and outside of the academy. although the data presented by the uk’s higher education statistics agency (hesa) on students and qualifiers (higher education statistics agency (n.d.) suggests that the number of doctoral graduates in the uk has tripled from 7,000 in 1994-5 to 22,000 in 2012-13, early concerns emerged during this period (e.g. harris, 1996; national committee of inquiry into higher education, 1997) about whether doctoral education was producing the highly skilled knowledge workers required by the knowledge economy, particularly in science, technology, engineering, mathematics, and medicine, that is, the so-called stemm subjects. doctoral education was considered to be overspecialised and not providing training in generic skills relevant to industry and commerce. in addition to questions about whether its assessment mode (a doctoral thesis judged in a viva voce examination) was fair (morley 2004; morley et al, 2002) but also appropriate (park, 2007), given the wider range of skills acquisition expected within the doctorate, other concerns included low and lengthy completion rates, low numbers entering stemm subjects, and gender biases in these disciplines (harris, 1996; institute of employment research, 2003). a key concern during this period has therefore been to intensify quality assurance of doctoral education. quality assurance agency for higher education, qaa (2004) introduced national guidelines regarding the frequency of doctoral supervision meetings, who can be a doctoral supervisor, the monitoring of student progress (overlapping uncomfortably with immigration-related monitoring of international students), and use of completion rates as a quality assurance measure. new institutional roles and practices (e.g., specialist consultants, specialized software for institutional monitoring of doctoral education, and new academic specialisations such as doctoral pedagogy) have evolved in response to these regulatory demands. uk doctoral education has also seen a strong emphasis on researcher training, framed in a discourse of individual skills and competences. a review by roberts (2002) was largely prompted by concerns about the supply of scientists and engineers and found that the phd provided “inadequate training – particularly in the more transferable skills” (p.10). having been constituted in 2005 to evaluate the impact of the “roberts” funding stream that was then created to support such training, the sector working group on the evaluation andres et al   | f l r       13   of skills development of early career researchers, known as the “rugby team” also promulgated the concept of “early career researcher” (ecr), defined as encompassing the first 10 years of a researcher’s postgraduate career (rugby team, 2006). their work also informed the constitution of vitae, a nationallyfunded body that promotes but also shapes uk researcher training through instruments such as its “researcher development framework” (rdf) (vitae, 2010), a text that continues to reflect the language of skills and competences. vitae is now promoting its rdf to european audiences and more widely, projecting the uk as a leader in doctoral education provision. maintaining a high level of international postgraduate admissions (currently around one third of the annual intake) is a further important priority for heis (universities uk, 2014). uk research council support for doctoral research has also become more focused. whereas in the past, applicants from a wide range of universities could apply for doctoral studentships, these are now awarded through a national network of “doctoral training centres (dtcs),” accredited by the research councils to award “mres” degrees (a structured masters’ degree devoted to research methods). in the social sciences, there are only 21 dtcs, so many universities (particularly “newer” universities) are excluded from accessing these studentships. this raises potential equity questions which require further research, as does the intensification of a research “training” agenda that aspires to incorporate a wider range of skills, but within a timeframe whose boundaries are more firmly regulated. 4.6. united states with greater numbers pursuing doctorates than ever before, the notion of a traditional research phd is expanding. the federal survey of earned doctorates (sed) reported that there were 52,760 earned research doctorates (phds) awarded from 421 doctoral granting institutions in 2013. this represents a 3.5% increase from 2012; in 2012 the rate had increased 4.2% from the previous year. fifty-eight percent of earned doctorates were in science and engineering, with the remainder being in the social sciences, humanities, and education (national science foundation, 2014). with these increases, fields such as the humanities continue to produce more doctorates than can be absorbed by available research careers (june, 2014; lederman, 2014). also, these figures mask the growth of professional or practice doctorates, including the edd (education, including educational administration), psyd (psychology), or dm (management). this double growth in doctorates exemplifies a massification of the credential, typified in disciplinary areas that require individuals to have doctoral credentials. this suggests an inflation of educational requirements with questionable value or unjustified educational costs, commonly without their mapping on to a societal or personal return on investment. although most formal educational institutions expect their researchers to have earned phds, it is not universally mandated. disciplinary bodies are beginning to acknowledge that the status quo of research doctorates solely for the purpose of preparing learners to continue on to academic rather non-academic careers is problematic (neem, 2014). for example, the american historical association is seeking to broaden career options for those who will not be able to obtain academic positions; academic positions will eventually be one of only several potential career opportunities or directions (grafton & grossman, 2011; jaschik, 2014). the 2014 report of the modern language association has as its first recommendation the need to redesign doctoral programs away from only academic careers. the goal of the mla is to “align [careers] with the learning needs and career goals of current and future students and to bring degree requirements in line with the ever evolving character of our fields” (mla task force on doctoral study in modern language and literature, 2014, p. 13). this is increasingly addressed by university career placement offices that help research students find positions outside academia (patel, 2015). lacking a central oversight body, doctoral regulations regarding program content, degree specifics, and university requirements are guided by 37,000 combinations of institutional, disciplinary, state, or national accreditation criteria (u.s. department of education office of postsecondary education (ope), n.d.). related to the number of disciplinary certification bodies and proprietary information among programs, it is difficult at best to try to compare data across programs and degrees to determine successful andres et al   | f l r       14   outcomes, speak to activities of early career researchers, or even track career paths (national science foundation, n.d.; sinche, 2014). with ambiguous quality assurance, it should not be surprising that there is nearly a 50% rate of doctoral attrition, including those in a limbo of decade-long abd (all but dissertation / defended) student status (yesko, 2014). given that less than 30% of u.s. faculty now work with tenure or are full-time on a tenure track (mla task force on doctoral study in modern language and literature, 2014), the growing population of casual and adjunct instructors, specifically those with doctoral degrees, will further invite investigation over educational quality. endemic challenges of fairness in pay and labour related to the increase of faculty in temporary or contract positions result in time spent ensuring future teaching contracts rather than engaging in research or university / disciplinary service. given the pragmatic nature of american doctoral training, current efforts focused on saving money through defunding education while eliminating full-time permanent faculty by relying increasingly on contingent labour point to a challenging future. 5. discussion from the individual country cases we have revealed three main issues: (1) what is happening on the ground? (2) the consequences of an increased formalization of doctoral education, and (3) the global-local nexus. 5.1. what is happening on the ground? in keeping with our recognition of the necessary recontextualisation of any policy narrative, our comparative study points to the need to examine more fully “what is happening on the ground” in order to understand more adequately how the different global trends play out in the institutional environments in specific countries. our comparative analysis demonstrates that the links between global (international) and local (national, institutional) levels of doctoral education are not similar across countries. even though the countries considered in this paper do subscribe to the same global trends on the policy level, there are many differences on the national and institutional levels. this, we argue, makes comparisons among systems of doctoral education at the global level difficult and fraught with uncertainties and potential inequalities. more research should be undertaken into unlocking the potential for understanding more fully the diverse and complex nature of doctoral educational practice worldwide. to fully understand the character and consequences of global trends within doctoral education, one needs to take into account the level of integration that always takes place at the local level. not only do countries differ when it comes to interpreting and understanding the meaning and relevance of global drivers such as massification, professionalization, and quality assurance within doctoral education, but individual institutions (universities) also face the task of integrating the global drivers into their own specific educational contexts and frameworks. 5.2. formalization of doctoral education as seen across the different country cases, even though there is a tendency to increase the numbers of doctoral programs, at the same time the aim is to consolidate them within doctoral “schools” and to enlarge the size of graduate schools within their institutions – hereby also increasing the level of formal training expected within doctoral education. the first issue relates back to the global trends of massification and quality assurance, while the latter issue is linked to the global trend of professionalization of doctoral education. as visible in the country cases, as the number of doctoral students have increased over the years, this has been met with the response of structuring doctoral education more “tightly” organisationally and demanding more formal procedures for how to develop and evaluate the performance done both by doctoral andres et al   | f l r       15   students and their supervisors. this has been described as the development of a generic doctoral curriculum (green 2009) and a “transdisciplinary doctorate” (willetts, mitcell, abeysuriya, & fam, 2012) which is promoted in order to ensure educational relevance for the job market and to safeguard the quality of doctoral education globally. the aim of foregrounding and developing the generic dimension of the phd across disciplines creates tension in relation to the desire to at the same time strengthen research environments at the disciplinary level, to maintain the strong disciplinary focus of the phd, and to resist its over-regulation (gudmundsson 2008). 5.3. the global-local nexus despite presenting country cases at the level of global drivers of doctoral education, we are aware that even if similarities exist at the level of policy, how these policies play out at local levels will always involve a process of recontextualization (bernstein, 2001). during our discussion of the meaning of the global drivers seen from individual national perspectives, it becomes apparent that although some of the same discourses and semantics are being used across different countries, the national, or local, meanings vary greatly. in addition to the shifts which recontextualisation necessarily involves, other factors which come into play include the size of the universities, the variation of gender, ethnicity, and age in student population, and the underlying political-economic conditions in each country. in a similar vein, teichler (2004) has pointed to the fact that “nations and strategic policies of national governments continue to play a major role in setting the frames for international communication, cooperation and mobility as well as for international competition. therefore, the frequent use of the term ‘globalization’ might be based on misunderstandings” (21). for example, this is specifically seen in the variation across countries regarding the meaning and management of “massification.” in some countries, massification seems to imply that the specific country “opens up the gate” with the simple aim of increasing the total number of people with phd credentials, as in the cases of colombia, denmark and finland. however, we see in the cases of canada, usa and the uk that massification is also about generating hierarchies within the doctoral system itself – creating a difference between the so called eliteand super-scholarship holders and the rest, thus pointing to equity issues within the phd system, which needs further scrutinizing. this calls for further research into what we call the “global-local nexus” of doctoral education. this nexus can be seen in several of the country cases where goals of increased internationalization of the doctorate and enhancing mobility among universities on a global scale stand alongside goals of strengthening research environments at the home institutions and the desire to allocate resources to enhance doctoral learning environments. also, this affects the very nature of the phd degree. originating as a universal degree with universal credentials, the increasing focus on internationalization and mobility paradoxically makes visible how diverse, complex, and in some cases incomparable, the phd degree has become. promotion of doctoral student mobility and concommitent alignment of different research programs and structures of different doctoral schools have become exceedingly difficult and has the potential to create many problems and unwanted strain for individual doctoral students and universities alike. this calls for further discussion about whether the phd degree is, still, really a universal degree or if it has transformed into a culturally and regionally contextual educational phenomenon. notwithstanding these distinctions, national and local priorities are not always aligned and the breadth of the doctoral experiences covered here, primarily those involving research doctorates, do not always transfer to professional doctorates that in some national contexts may often focus on more local priorities. the point emerging from this paper is that understandings of global and local levels of doctoral education are deeply linked, as global drivers saturate local doctoral education and supervision practice. more in depth understandings is needed regarding how this is played out at local institutional levels and also if and how these local practices relate back to global and political levels of doctoral education. more specifically, further research about the following is required: a) how global trends, drivers, and strategies for doctoral education play out in local national settings and how such global drivers are integrated locally in specific teaching and learning environments at specific universities; andres et al   | f l r       16   b) more awareness and discussion about the “universality” of the phd degree. in an era where mobility regarding doctoral education policy is on the agenda, more attention should be given to what is actually possible to transfer across national arenas; c) what possibilities and challenges do the infrastructures of graduate schools bring with them in relation to doctoral education. we need to examine the everyday workings of graduate schools to learn more about what forms of organisation are at work within the broader higher education system. this paper focused on six national sets of policies regarding research doctorates. it was beyond the scope of this paper to address the varied complexities of doctoral study in other nations. without attending to all national and international trends, including those in australia, new zealand, and the asian and african regions, the scope of this study is necessarily limited. we hope that our attempt to initiate these discussions will serve the purpose of highlighting what can only be thought of as a expanding area of study. andres et al   | f l r       17   table 1 summary of recent changes of doctoral education in canada, colombia, denmark, finland, uk and usa driver country massification professionalization quality assurance canada • aim to increase number of phds • poor employment outlook within academia • funding for elite students and institutions • structuring doctoral fellowships to be in line with economic and social trends • emphasis on labor market ready skill acquisition, internships, and partnerships with industry/business colombia • investment in educating phds abroad to increase number of degree holders • rapid increase in the number of doctoral programs providing degrees • 150% increase in number of phds • launching doctoral programs • increase in regulation and regulatory institutions • defunding education • cutting permanent places for faculty denmark • doubling the annual enrollment of doctoral students • emphasizing learning of generic skills and interdisciplinarity • investing in developing teaching at the university • harmonizing doctoral degree according to european standards (i.e. third cycle of bologna process) • launching graduate schools • adopting benchmarking and raking systems finland • increase in number of doctoral degree holders • awarding universities for attracting international students completing the phds • launching doctoral schools and programs • harmonizing doctoral degrees according to european standards (adopting bologna qualifications) • introducing four stage researcher career model and tenure track system • launching international doctoral programs • adopting benchmarking, and international evaluation systems uk • increase in number of doctoral students • development of professional doctorates • development of legislation and charters to address inequalities related to gender and race inequalities • emphasis on training: in skills and competences • providing national funding for generic skills training • more structured preparation for phd entry degree • contract researcher career system • awarding national funding for phd scholarships through networks of accredited doctoral training centres • stronger regulation of a host of doctoral education issues usa • increase in both professional and research doctorates • aim to increase number of phds amongst african american and hispanic leaners • introducing professional doctorate degrees • emphasizing labor market ready skills also in the training of research doctorates • defunding education • cutting permanent faculty • increasing contingent labor force andres et al   | f l r       18   keypoints the national, or local, meanings of doctoral education vary greatly. we question whether the phd degree is a universal degree or if it has transformed into a culturally and regionally contextual educational phenomenon. comparison among systems of doctoral education on the global level difficult and fraught with uncertainties and potential inequalities. references academy of finland. (2010). get ahead in our career. get a doctorate. retrieved from http://aka.smartpage.fi/en/doctors_career/. andras, p. (2011). research: metrics, quality, and management implications. research evaluation, 20(2), 90–106. doi: 10.3152/095820211x12941371876265. auriol, l., misu, m., & freeman, r. a. (2013). careers of doctorate holders: analysis of labour market and mobility indicators. oecd science, technology and industry working papers, 2013/04. oecd publishing. doi: 10.1787/5k43nxgs289w-en. auriol, l., schaaper, m., & felix, b. (2012). mapping careers and mobility of doctorate holders: draft guidelines, model questionnaire and indicators – third edition. oecd science, technology and industry working papers, 2012/07. oecd publishing. doi: 10.1787/5k4dnq2h4n5c-en. berlin communiqué. (2003). realising the european higher education area. conference of ministers responsible for higher education in 33 european countries (september). bernstein, s. (2001). the compromise of liberal environmentalism. new york: columbia university press. brunner, j. j. (2001). globalización y el futuro de la educación: tendencias, desafíos, estrategias. análisis de prospectivas de la educación en américa latina y el caribe. santiago de chile: unesco. retrieved from http://www.rmm.mineduc.cl/usuarios/jvill1/file/futuroedunesco.pdf. bucharest communiqué. (2012). making the most of our potential: consolidating the european higher. education area. retrieved from http://www.ehea.info/uploads/%281%29/bucharest%20communique%202012%281%29.pdf. buela-casal, g., gutiérrez-martínez, o., bermúdez-sánchez, m. p., & vadillo-muñoz, o. (2007). comparative study of international academic rankings of universities. scientometrics, 71(3), 349– 365. doi: 10.1007/s11192-007-1653-8. conference board of canada. (2014). how canada performs. retrieved from http://www.conferenceboard.ca/hcp/provincial/education/phd.aspx. danish ministry of higher education and science (2014). official website. retrieved from http://ufm.dk/en?set_language=en&cl=en danish ministry of higher education and science (2015a). official website. retrieved from http://ufm.dk/uddannelse-og-institutioner/videregaende-uddannelse/universiteter/ph-duddannelse/ph-d-skoler. danish ministry of higher education and science (2015b). official website. retrieved from http://ufm.dk/uddannelse-og-institutioner/anerkendelse-ogdokumentation/dokumentation/kvalifikationsrammer/europaeisk-kvalifikationsramme-eqf. department for education and skills (dfes). (2003). the future of higher education. london: hmso. edler, j., georhiou,l., blind, k., & uyrra, e. (2012). evaluating the demand side: new challenges for evaluation. research evaluation, 21, 33–47. doi:10.1093/reseval/rvr002. elgar, f. j. (2003). phd degree completion in canadian universities: final report. halifax: graduate students association of canada. retrieved from http://www.researchgate.net/profile/frank_elgar/publication/236595361_phd_degree_completion_in_ canadian_universities_final_report/links/02e7e5182b86db33e0000000.pdf. andres et al   | f l r       19   eua (european university association). (2009). collaborative doctoral education: university-industry partnerships for enhancing knowledge exchange. eua, brussels. retrieved from http://www.eua.be/eua-work-and-policy-area/research-and-innovation/doctoral-education/doc-careers/. eua (european university association). (2010). salzburg ii recommendations: european universities’ achievements since 2005 in implementing the salzburg principles. eua, brussels. retrieved from http://www.eua.be/news/10-1028/eua_publishes_recommendations_for_continued_reform_of_doctoral_education.aspx. european commission. (2014). european research area progress report. brussels, european commission. retrieved from http://ec.europa.eu/research/era/eraprogress_en.htm european higher education area. (2015). official website. retrieved from http://www.ehea.info. fernandez-zubieta, a. & guy, k. (2010). developing the european eesearch area: improving knowledge flows via researcher mobility. jcr scientific and technical reports. european commission. retrieved from http://erawatch.jrc.ec.europa.eu/erawatch/opencms/information/reports/countries/eu/report_0020. finnish ministry of education (1997). tutkijakoulut suomessa 1995–1998. tutkijakouluissa annettavan opetuksen ja ohjauksen laadun arviointi [doctoral schools in finland 1995–1998.]. opetusministeriö, koulutusja tiedepolitiikan osasto. fiske, p. (2011). what is a phd really worth? nature, 472, 381. doi:10.1038/nj7343-381a. fortes, m., kehm, b. m., & mayekiso, t. (2014). evaluation and quality management in europe, mexico, and south africa. in m. nerad & b. evans (eds.), globalization and its impacts on the quality of phd education (pp. 81–110). springer: rotterdam. frank, r. h. (1999). higher education: the ultimate winner-take-all market? ithica, ny. retrieved from http://digitalcommons.ilr.cornell.edu/cheri/2/. garam, i. (2013) kansainvälinen liikkuvuus yliopistoissa ja ammattikorkeakouluissa 2013. tietoja ja tilastoja -raportti 2/2014. centre for international mobility cimo. retrieved from http://www.cimo.fi/instancedata/prime_product_julkaisu/cimo/embeds/cimowwwstructure/32368_tieto a_ja_tilastoja-raportti_2_2014.pdf. general law of education. (1994). [law 115, 1994]. do: 41.214. colombian congress. geuna, a., & martin, b. (2003). university research evaluation and funding: an international comparison, minerva, 41(4), 277–304. doi: 10.1023/b:mine.0000005155.70870.bd. gilbert, r., balatti, j., turner, p., & whitehouse, h. (2004). the generic skills debate in research higher degrees. higher education research & development, 23(3), 375–388. doi: 10.1080/0729436042000235454 grafton, a. t., & grossman, j. (2011). no more plan b: a very modest proposal for graduate programs in history. historians.org. retrieved from http://historians.org/publications-and-directories/perspectiveson-history/october-2011/no-more-plan-b. green, b. (2009). challenging perspectives, challenging practices. doctoral education in transition. in boud, d. & lee, a. (eds.). changing practices of doctoral education. london & new york: routledge. gudmundsson, h. k. (2008). nordic countries. in nerad, m., & heggelund, m. (eds.). toward a global phd? forces & forms in doctoral education worldwide. seattle & london: university of washington press. harris, m. (1996). review of postgraduate education. report for higher educational funding council for england, committee of vice chancellors and principals, and standing conference of principals (bristol). higher education statistics agency (n.d.) retrieved from https://www.hesa.ac.uk/component/datatables/, accessed 29 april 2015 hsieh, h.-f., & shannon, s. e. (2005). three approaches to qualitative content analysis. qualitative health research, 15(9), 1277–1288. doi:10.1177/1049732305276687. institute for employment research. (2003). bulletin: women in science, engineering and technology. university of warwick: warwick. institute for the public life of arts and ideas. (2013). white paper on the future of the phd in the humanities. montreal: mcgill university. andres et al   | f l r       20   jaramillo s., h. (2009). la formación de posgrado en colombia. revista iberoamericana de ciencia tecnología y sociedad, 5(13), 131–155. retrieved from http://www.scielo.org.ar/scielo.php?script=sci_arttext&pid=s185000132009000200008&lng=es&nrm=iso. issn 1850-0013. jaschik, s. (2014). a broader history ph.d. inside higher ed. retrieved from https://www.insidehighered.com/news/2014/03/20/historians-association-and-four-doctoral-programsstart-new-effort-broaden-phd. june, a. w. (2014). doctoral degrees increased last year, but career opportunities remained bleak. the chronicle of higher education. retrieved from http://chronicle.com/article/doctoral-degreesincreased/150421/ kehm, b. m. (2006). doctoral education in europe and north america. a comparative analysis. in u. teichler (ed.), the formative years of scholars. wenner-gren international series, vol. 83 (pp. 67– 78). london: portland press. kota-national data base. (2009). retrieved from https://kotaplus.csc.fi/online/transfer.do. lederman, d. (2014). doctorates up, career prospects not. inside higher ed. retrieved from https://www.insidehighered.com/news/2014/12/08/number-phds-awarded-climbs-recipients-jobprospects-dropping. mla task force on doctoral study in modern language and literature. (2014). report of the mla task force on doctoral study in modern language and literature, 1–41. moguérou, p., & di pietrogiacomo, m. p. (2008). stock, career and mobility of researchers in the eu. jcr scientific and technical reports. european comission. retrieved from http://erawatch.jrc.ec.europa.eu/erawatch/opencms/information/reports/countries/eu/report_mig_0011 morley, l. (2004). interrogating doctoral assessment. international journal of educational research, 41(2), 91–97. morley, l., leonard, d., and david, m. (2002). variations in vivas: quality and equality in british phd assessments. studies in higher education, 27(3), 263–273. doi:10.1080/03075070220000653. national committee of inquiry into higher education (the dearing report). (1997). higher education in the learning society. hmso: london. national science foundation (nsf). (n.d.). early career doctorates project (forthcoming). nsf.gov. retrieved from http://www.nsf.gov/statistics/srvyecd/. accessed 1 november 2014. national science foundation. (2014). doctorate recipients from u.s. universities 2012: survey of earned doctorates. national center for science and engineering statistics. retrieved from http://www.nsf.gov/statistics/srvydoctorates/. neem, j. (2014). ministers, not m.b.a.s. inside higher ed. retrieved from https://www.insidehighered.com/views/2014/10/03/humanities-phd-calling-not-vocational-trainingessay. niemi, h., aittola, h., harmaakorpi, v., lassila, o., svärd, s., ylikarjula, j., hiltunen, k., & talvinen, k. (2011). tohtorikoulutuksen rakenteet muutoksessa: tohtorikoulutuksen kansallinen seurantaarviointi [changing structures of doctoran education: national follow-up evaluation of doctoral education]. korkeakoulujen arviointineuvoston julkaisuja, 15. oecd (2010). the oecd innovation strategy: getting a head start on tomorrow. paris: oecd. oecd (2012). transferable skills training for researchers: supporting career development and research. oecd publishing. doi: 10.1787/9789264179721-en. oecd (2014). education at a glance 2014: oecd indicators. oecd publishing. doi: 10.1787/eag-2014en. park, c. (2007). redefining the doctorate. higher education academy, york. retrieved from http://eprints.lancs.ac.uk/435/1/redefiningthedoctorate.pdf. patel, v. (2015). new job on campus: expanding ph.d. career options. the chronicle of higher education. retrieved from http://chronicle.com/article/new-job-on-campusexpanding/151105/?key=sg16ivrgnicwm3bnzglanjhroh1snbwjanfjbcgkbllxfg%3d%3d andres et al   | f l r       21   puhakka, a., & rautopuro, j. (2013). sumusta nousee riski – tieteentekijöiden liiton jäsenkysely [the finnish union of university researchers and teachers membership survey]. joensuu: grano. retrieved from http://tieteentekijoidenliitto.fi/materiaali/jasenkyselyraportit. quality assurance agency for higher education (qaa). (2004). code of practice for the assurance of academic quality and standards in higher education. section 1: postgraduate research programmes. gloucester: qaa. república de colombia, ministerio de educación nacional (men) (2010). consejo nacional de acreditación (cna). lineamientos para la acreditación de alta calidad de programas de maestría y doctorado. 2010, bogotá. retrieved from http://cmsstatic.colombiaaprende.edu.co/cache/binaries/articles-186363_lineam_myd.pdf?binary_rand=7259. rizen, j., & marconi, g. (2011). internationalization in european higher education. international journal of innovation science, 3(2), 83–100. doi: 10.1260/1757-2223.3.2.83. roberts, g. (2002). set for success. final report of sir gareth roberts review. london: hmso. rugby team. (2006). evaluation of skills development of early career researchers a strategy paper from the rugby team. uk grad programme roberts policy forum. birmingham: uk grad. sainio, j. (2010). asiantuntijana työmarkkinoille. vuosina 2006 ja 2007 tohtorin tutkinnon suorittaneiden työllistyminen ja heidän mielipiteitään tohtorikoulutuksesta [as an expert to the labour market. employment of those receiving a doctorate in 2006 and 2007, and their opinions about doctoral education]. aarresaaren julkaisusarja. retrieved from https://www.aarresaari.net/uraseuranta/julkaisut. shin, j. c. (2010). impacts of performance-based accountability on institutional performance in us. higher education, 60(1), 47-68 doi: 10.1007/s10734-009-9285-y. shin, j. c. & j. jung (2014). academic job satisfaction and job stress across countries in changing academic environments. higher education, 67, 603–620. doi: 10.1007/s10734-013-9668-y. sinche, m. (2014). tracking ph.d. career paths. inside higher ed. retrieved from https://www.insidehighered.com/advice/2014/10/27/essay-importance-tracking-phd-career-paths. tamburri, r. (2013). the phd is in need of revision. university affairs. retrieved from http://www.universityaffairs.ca/features/feature-article/the-phd-is-in-need-of-revision/. teichler, u. (2004). the changing debate on internationalisation of higher education, higher education, 48 (1), 5–26. doi: 10.1023/b:high.0000033771.69078.41. the graduate school working group, (2012). towards quality, transparency and predictability in doctoral training. the graduate school working group’s suggestions for doctoral training development. academy of finland. retrieved from http://www.aka.fi/en-gb/a/academy-offinland/academy-publications/other-publications/ the world bank. (2003). colombian tertiary education in the context of reform in latin america. colombia: tertiary education paving the way for reform vol. 1. policy briefing. the world bank, report no. 23935-co. treuthardt, l., & nuutinen, a. (2012). the state of scientific research in finland 2012. publications of academy of finland 7/2012. retrieved from http://www.aka.fi/en-gb/a/decisions-and-impacts/thestate-of-scientific-research-in-finland/previous-reviews/the-state-of-scientific-research-in-finland20121/ ubc graduate and postdoctoral studies. (2014). re-imagining the phd: new forms and futures for graduate education. symposium summary report. vancouver: university of british columbia. retrieved from http://www.ligi.ubc.ca/sites/liu/files/news/ubcgraduatesymposiumreport2014.pdf. uk council for science and technology. (2007). pathways to the future: the early career of researchers in the uk. cst, london. unesco-ibe. (2011). world data on education vii ed. 2010/11. retrieved from http://www.ibe.unesco.org //. universities uk. (2014). international students in higher education: the uk and its competition. london: uuk. u.s. department of education office of postsecondary education (ope). (n.d.). the database of accredited postsecondary institutions and programs. retrieved from http://ope.ed.gov/accreditation. accessed 19 october 2014. andres et al   | f l r       22   vitae. (2010). researcher development statement. cambridge: vitae and crac. willetts, j., mitcell, c., abeysuriya, k., & fam, d. (2012). creative tensions. negotiating the multiple dimensions of a transdisciplinary doctorate. in lee, a. & danby, s. (eds.). reshaping doctoral education. international approaches and pedagogies. london & new york: routledge. yesko, j. (2014). an alternative to abd. inside higher ed. retrieved from https://www.insidehighered.com/advice/2014/07/25/higher-ed-should-create-alternative-abd-statusessay yin, r. k. (2012). applications of case study research (3rd ed.). thousand oaks, ca: sage. microsoft word seiz et al_publication.docx           frontline learning research vol.3 no. 1 (2015) 55 77 issn 2295-3159   corresponding author. johanna seiz, institute of psychology, department of educational psychology, goethe university, theodor-w.-adorno-platz 6, 60629 frankfurt/main, germany. e-mail: seiz@psych.uni-frankfurt.de doi: http://dx.doi.org/10.14786/flr.v3i1.141 when knowing is not enough – the relevance of teachers’ cognitive and emotional resources for classroom management johanna seiza, thamar vossb, mareike kuntera agoethe university of frankfurt, germany buniversity of tübingen, germany article received 22 december 2014 / revised 13 march 2015 / accepted 26 march 2015 / available online 13 may 2015 abstract this study expands the discussion on teacher competence by investigating the relevance of teachers’ combined cognitive resources and emotional resources for effective classroom management. while research on teacher qualification stresses the importance of knowledge for effective teaching, research on teacher stress focuses on their emotional functioning, often without connection to their in-class behaviour. drawing on findings from health psychology showing that high levels of emotional exhaustion can impair cognitive performance, we hypothesised that teachers’ pedagogical/psychological knowledge would predict their classroom management behaviour only when their level of emotional exhaustion was low. we administered a test to assess the pedagogical/psychological knowledge of 205 secondary school teachers, measured their emotional exhaustion, and assessed their classroom management using ratings of their 4,672 students obtained one year later. data were analysed using latent moderation analyses, a novel statistical approach that rarely has been employed in research on learning and instruction. our findings confirmed our hypotheses and indicated an interaction between teachers’ cognitive resources and emotional resources, which together predict their classroom management behaviour. thus, the new theoretical and empirical integration of two distinct areas of teacher quality broadens our understanding of teacher resources necessary for effective instruction. we argue that teacher education should acknowledge the interplay of the different resources teachers have and help them develop their emotional resources to ensure effective instruction. keywords: classroom management; teacher competence; emotional exhaustion; professional knowledge     seiz et al | f l r     56 1. introduction there has been considerable debate in educational research about which qualities make teachers effective (e.g., roehrig et al., 2012). from a subject-specific perspective, professional knowledge as a cognitive resource is essential (shulman, 1986, 1987); however, some authors (e.g., jennings & greenberg, 2009) stress the importance of teachers’ emotional resources. in this study we expand the discussion on teacher competence by investigating the relevance of teachers’ combined cognitive resources and emotional resources for effective classroom management. teaching is a complex activity (doyle, 2006; helsing, 2007) in two respects. first, classrooms have unique characteristics (doyle, 2006). for instance, the multitude of tasks which all require an adequate response from the teacher, reflects considerable multidimensionality. as many tasks occur simultaneously, teachers need appropriate monitoring and management skills. unexpected disruptions can occur in the classroom and put constant pressure on the teacher and the teaching task (doyle, 2006). second, teachers need to employ several skills for effective instruction (e.g. baumert et al., 2010). they must choose instructional tasks and appropriate methods, establish rules and structures to manage the class, and provide students with emotional as well as individual learning support (baumert et al., 2010; pianta & hamre, 2009). all these demands and practices occur simultaneously and are interconnected. some authors argue that efficient classroom management supports learning-related activities as it structures the learning environment (doyle, 2006; ophardt & thiel, 2008). the widely-used observation tool classroom assessment scoring system (class) considers classroom organisation to be one dimension in its framework, as it is believed to be relevant to students’ academic and social development (pianta & hamre, 2009). taking this into account, we focus in this study on classroom management as an important part of instructional quality. evertson and weinstein (2006) defined classroom management as “the actions teachers take to create an environment that supports and facilitates both academic and social-emotional learning” (p. 4). this definition subsumes various dimensions of teacher behaviour. empirically effective and therefore central dimensions of classroom management include monitoring students’ behaviour, preventing disturbances, establishing rules, and quickly intervening during disruptions (marzano, marzano, & pickering, 2003). monitoring involves continuously observing students, which enables the teacher to prevent or detect, and possibly react quickly to, disruptions (doyle, 2006; kounin, 1970). establishing rules in the classroom is an important part of classroom management (emmer & evertson, 2013), as these help students to regulate their behaviour. in addition, reacting adequately and quickly to disruptions is crucial (marzano et al., 2003). further dimensions of effective classroom management focus on the quality of student-teacher relationships and the maintenance of instructional flow (doyle, 2006; pianta, 2006). empirical evidence shows that classroom management is crucial for students in various groups and in different domains (wang, haertel, & walberg, 1993). effective classroom management is a strong predictor for students’ academic outcomes (e.g. wang et al., 1993). yet classroom management also is related to non-cognitive outcomes such as students’ motivation and interest in (fauth, decristan, rieser, klieme, & büttner, 2014; kunter, baumert, & köller, 2007), and satisfaction with, school (nie & lau, 2009). further, effective classroom management can result in better studentteacher relationships (de jong et al., 2014). however, effective classroom management is challenging, especially for young teachers, who often do not feel well prepared for the task (liston, whitcomb, & borko, 2006). to summarise, effective classroom management is crucial for students yet challenging for teachers, as it requires pedagogical, social and emotional competence as well as the ability to react quickly and appropriately in critical situations. given the complexity of classroom management, the     seiz et al | f l r     57 question arises as to which resources—in the sense of personal prerequisites—teachers need to manage their classrooms effectively. 1.1 necessary resources for effective classroom management there is considerable discussion about the prerequisites for providing high quality instruction and managing the classroom effectively. in the following section we introduce two views, one stressing the importance of teachers’ professional knowledge and the other stressing the relevance of teachers’ emotional resources. 1.1.1 professional knowledge – the importance of teachers’ cognitive resources one particular cognitive resource that often is considered a prerequisite for high quality instruction is professional knowledge (depaepe, verschaffel, & kelchtermans, 2013; shulman, 1986, 1987). within this discussion professional knowledge is understood as specialised knowledge shared within a community of professionals. research has shown that subject matter related knowledge such as subject-specific content knowledge and subject-specific pedagogical content knowledge are important for processing and communicating content related tasks (depaepe et al., 2013; krauss et al., 2008). subject matter related knowledge is an important predictor for cognitive activation and student achievement (baumert et al., 2010; hill, rowan, & ball, 2005). regarding classroom management, subject-unspecific knowledge such as pedagogical/ psychological knowledge, meaning the teachers’ knowledge of creating and improving classroom situations and interactions, is of great importance. such knowledge includes that of classroom management strategies, teaching methods, classroom assessment and dealing with students’ heterogeneity (park & oliver, 2008; voss, kunter, & baumert, 2011; voss, kunina-habenicht, & kunter, 2015). pedagogical/psychological knowledge subsumes declarative and procedural knowledge (voss et al., 2011). the importance of this knowledge was indicated in a recent study in which teachers’ pedagogical/psychological knowledge was shown to be associated with the quality of their instruction, including classroom management (voss, kunter, seiz, hoehne, & baumert, 2014). helping teacher candidates develop classroom management skills is therefore an essential part of teacher education (emmer & stough, 2001). the assumption that teachers need cognitive resources such as professional knowledge for effective instruction thus seems well established. however, considering the great challenge that teaching may present, other researchers have argued that emotional resources are another important asset for teachers. 1.1.2 the importance of teachers’ emotional resources in their prosocial classroom model, jennings and greenberg (2009) claim that teachers’ social and emotional resources are prerequisites for effective teaching and especially for classroom management. following their model, teachers with sufficient emotional resources are better capable of dealing with diverse challenges in their classrooms such as effectively managing their classrooms. in the model it is assumed that effective classroom management leads to an optimal classroom climate with positive social, emotional and academic outcomes for students (jennings & greenberg, 2009). teachers’ emotional resources clearly are an important topic to investigate (sutton, 2005; sutton & wheatley, 2003). teaching is an emotionally challenging profession and teachers need to be able to regulate their emotions (roeser et al., 2013). teachers’ emotional resources have often been analysed within health psychology, focusing on how negative emotions evolve; yet few studies have analysed the relationship between teacher emotions and instructional behavior. keller, chang, becker,     seiz et al | f l r     58 goetz, and frenzel (2014) showed that emotional exhaustion was related to teachers’ emotional experience in their classrooms. highly exhausted teachers reported increased feelings of anger and less enjoyment during instruction. further, teacher emotions were associated with student-rated instructional quality (frenzel, goetz, stephens, & jacob, 2009). in a study testing the assumption that teachers’ emotional resources are relevant to their instructional behaviour, klusmann, kunter, trautwein, lüdtke, and baumert (2008) found that teachers who were able to balance their emotional engagement attained better instructional quality and their students reported greater motivation. additionally, students’ and teachers’ emotions seem to be related (becker, goetz, morger, & ranellucci, 2014): students who witnessed their teachers enjoying instruction also felt more enjoyment in class. summing up, teacher emotions are a relevant resource for effective instruction. 1.1.3 combining the perspectives: the interplay of cognitive and emotional resources to date, researchers have investigated the cognitive and emotional resources of teachers mostly in separate studies stemming from different theoretical traditions (e.g. brouwers & tomic, 1999; depaepe et al., 2013; skaalvik & skaalvik, 2011), neglecting a possible joint relevance for effective teaching, especially for classroom management. in our study, we combine both perspectives. although we agree that cognitive resources such as professional knowledge are crucial for effective instruction, we argue that due to the complexity of teaching, teachers will be able to profit from their cognitive resources only if they also possess a sufficient amount of emotional resources. thus, one the one hand, in this study we consider teachers’ pedagogical/psychological knowledge as an example of their cognitive resources. on the other hand, we consider emotional exhaustion as a central aspect of teachers’ emotional resources (klusmann et al., 2008). emotional exhaustion is the feeling of being drained or experiencing chronic fatigue and a low level of energy (maslach & leiter, 1999; schwarzer, schmitz, & tang, 2000), and it is the core component of burnout syndrome (maslach, schaufeli, & leiter, 2001). many studies have shown that teachers generally report higher levels of emotional exhaustion than other professionals although significant differences among teachers exist (e.g., hakanen, bakker &, schaufeli, 2006; unterbrink et al., 2007). the interdependence of cognitive and emotional resources already has been empirically demonstrated in research on health psychology. studies comparing the cognitive functioning of highly exhausted adults and non-exhausted adults have shown that those with high levels of exhaustion had impaired cognitive functioning (kleinsorge, diestel, scheil, & niven, 2014; sandström, rhodin, lundberg, olsson, & nyberg, 2005). in a study by van der linden, keijsers, eling, and schaijk (2005) a non-clinical sample of exhausted teachers performed significantly lower on cognitive performance tasks than a sample of non-exhausted teachers. feuerhahn et al. (2013) investigated the relation between emotional exhaustion and multiple indicators of performance using a sample of teachers with varying degrees of exhaustion. they found that emotional exhaustion was related to cognitive impairment. in a follow-up investigation six months later, emotional exhaustion at the first testing period predicted impairment ratings at the second testing period. however, emotional exhaustion at the second testing period was not predicted by cognitive impairment at the first testing period, meaning that emotional exhaustion leads to cognitive impairment, rather than the other way around. most of these studies were framed within information processing theory (e.g., feldon, 2007; mayer, 2012; sweller, van merrienboer, & paas, 1998) which assumes that emotional exhaustion limits information processing capacity and thus leads to poorer performance on cognitive performance tasks. applying this to teachers, who need sufficient information processing capacities to be able to use their cognitive resources in challenging classroom situations (feldon, 2007), one might assume that high levels of emotional exhaustion could drain processing capacities limiting the access to professional knowledge.     seiz et al | f l r     59 further theories and approaches can be used to support our argument. ego depletion theory assumes that self-regulation is based on a limited amount of resources (baumeister, gailliot, dewall, & oaten, 2006) and that each act of self-control exhausts these resources and leads to a state of ego depletion. subsequent attempts at self-control or volition will fail due to a lack of available resources. studies supporting ego depletion theory showed that after acts of self-regulation (e.g. emotional or attentional regulation) performance was impaired in tasks demanding high-order cognitive functioning (johns, inzlicht, & schmader, 2008; schmeichel, vohs, & baumeister, 2003). it could be argued that teachers with a high level of emotional exhaustion are in a state of ego depletion because their selfregulatory efforts have used up resources for further acts of volition (e.g., knowledge-based decisions concerning classroom management). 1.2 the present study in this study we analyse the interaction between teachers’ cognitive resources and emotional resources for classroom management behaviour. we argue that knowledge (as a cognitive resource) and emotional exhaustion (as an emotional resource) are interconnected when it comes to predicting teachers’ behaviour as emotional exhaustion might limit capacities to process knowledge. classroom management behaviour such as monitoring or preventing disturbances relies on cognitive resources as it requires quick reactions to the unforeseen (e.g., feldon, 2007). we hypothesize that the successful application of knowledge in challenging classroom situations requires sufficient information processing capacity, but that high emotional exhaustion will reduce these processing capacities, thus limiting teachers’ access to knowledge. only when teachers possess sufficient emotional resources will they have enough capacities to apply knowledge-based strategies to manage the classroom. to our knowledge, this is the first study that combines cognitive and emotional resources of teachers to predict their in-class teaching behaviour. 1.3 hypotheses methodologically, we thus investigate whether teachers’ emotional exhaustion moderates the relation between their professional knowledge and their classroom management behaviours, as indicated by their prevention of disturbances and their monitoring behaviour. we hypothesise as follows: [1] pedagogical/psychological knowledge will not predict: a) classroom disturbances when the level of emotional exhaustion is high. b) monitoring behaviour when the level of emotional exhaustion is high. [2] when the level of emotional exhaustion is low, pedagogical/psychological knowledge will relate: a) negatively to classroom disturbances. b) positively to monitoring behaviour.     seiz et al | f l r     60 2. method 2.1 design and sample the data used in this study were derived from a larger longitudinal study investigating the development of secondary school mathematics teacher candidates’ professional competence during and after the practical induction phase. the practical induction phase is mandatory in germany and follows university studies. during this phase teacher candidates are placed in schools where they observe instruction and gradually start their own teaching. in addition, they attend courses on general principles of teaching. two assessments of this study were used for this analysis. the first assessment involved 568 participants and was conducted at the end of the participants’ induction phase. in this assessment, pedagogical/psychological knowledge and emotional exhaustion were assessed. the aim of the second assessment was to gather data on instructional quality (rated via students) of the participants after they had taken over full teaching responsibilities. therefore, the second assessment was conducted 14 months after the end of the induction phase to ensure that participants were already established as teachers. in this assessment 205 teachers and their students still participated. in our study we aimed at predicting studentrated quality of classroom management using prior teacher resources. therefore we used this subsample of 205 teachers of the second assessment as our sample of analysis. our sample was 61% women and the average age of the participants was 28.4 years (sd = 3.74) at the first assessment. participants had on average 14 months of teaching experience when the data were collected at the second assessment. germany has a tracked school system with a high, an intermediate and a low track. these different school types were represented in the sample; however, the sample was slightly skewed as 61.3% of the participants taught the highest school track. in 2013/2014, 47% of all students in germany attended the highest school track (statistisches bundesamt, 2014). we analysed the demographic variables (age, sex and school type) and self-reports on motivation and exhaustion of the dropouts from the two different samples of the longitudinal study. participants of the second assessment taught more in the higher school track, were more enthusiastic and satisfied with their jobs, and showed less emotional exhaustion. thus, the generalisability of our results may be somewhat compromised. in addition, 4,672 students from grades 7 to 10 participated in the second assessment and were included in our analyses. on average 12 students rated the classroom management of each teacher. teachers were allowed to have up to five classes participate in the ratings. however, ratings from all the classes of each teacher were combined, as they revealed high correlations across classes and our focus was on the teacher. a different analysis of this data focussing on the importance of pedagogical/psychological knowledge for general instructional quality based on a different teacher sample already has been published (voss et al., 2014). the focus of this investigation is the relevance of the interplay between different teacher resources and how this interplay affects classroom management, which has not yet been the subject of investigation. including emotional exhaustion as a moderator expands existing research and allows testing more differentiated hypotheses on the relevance of teachers’ professional knowledge.     seiz et al | f l r     61 2.2 measures we applied confirmatory factor analysis and structural equation modeling. the scales and items described represent the multiple indicators for the latent factors. table 1 provides an overview of the descriptive data and the reliabilities of the measures based on the raw dataset. the remainder of the analysis refers to the latent dimensions of the variables. table 2 provides an overview of the fit indices of the measurement models; appendix a displays information on factor loadings of the indicators on the respective factors. table 1 psychometric properties of study variables variables items m sd icc1 icc2 adm α   missing in %   teacher ratings pedagogical/psychological knowledge 39 73.37 11.45 ― ― ― .79 19.5 emotional exhaustion 4 72.02 1.64 ― ― ― .81 8.3 student ratings classroom disturbances 6 72.17 1.74 .33 .92 .70 ― .6 monitoring 3 72.85 1.65 .23 .87 .64 ― .7 note. student ratings based on teacher mean scores. icc = intraclass correlation, adm = average deviation index, averaged across all classes of each teacher. 2.2.1 pedagogical/psychological knowledge to assess teachers’ pedagogical/psychological knowledge we employed a test that had been used and validated in previous studies (voss et al., 2011; voss et al., 2014). the test consists of four scales measuring knowledge of classroom management, teaching methods, classroom assessment and students’ heterogeneity. test construction and validation analysis indicated that the scales are well represented by a second order factor expressing general pedagogical/psychological knowledge (voss et al., 2011); thus we used the scales as indicators of one latent factor. altogether, the measure consists of 39 items across the four subscales including multiple-choice, short-answer and video-based items (voss et al., 2011). the multiple-choice items assessed declarative knowledge whereas procedural knowledge also was assessed using video-based items on classroom management. pedagogical/psychological knowledge as measured by this test has proven to be differentiable from discriminant constructs such as general reasoning ability, pedagogical content knowledge and teacher beliefs about mathematics learning and teaching (voss et al., 2011). 2.2.2 emotional exhaustion we used an established german version (enzmann & kleiber, 1989) of the maslach burnout inventory (maslach, jackson, & leiter, 1996) to assess teachers’ state of emotional exhaustion. the     seiz et al | f l r     62 instrument consists of four items and participants rated their agreement with statements (e.g., “i often feel exhausted at school”) on a 4-point response scale (1 = strongly disagree, 4 = strongly agree). table 2 fit indices of individual and combined measurement models without interaction term model χ2 df p cfi rmsea srmr (between) srmr (within) teacher ratings emotional exhaustion 2 8.57 2 .01 .97 .03 .03 ― pedagogical/ psychological knowledge 6.85 2 .03 .94 .03 .04 ― student ratings classroom disturbance 173.42 18 .00 .98 .04 .03 .03 monitoring .00 0 1.00 1.00 .00 .00 .00 measurement models without interaction term model 1 277.93 94 .00 .98 .02 .05 .03 model 2 75.24 49 .01 .98 .01 .06 .00 note. cfi = comparative fit index; rmsea = root-mean-square error of approximation; srmr = standardized root-mean-square residual. dashes indicate nonavailable data. the monitoring model is saturated. 2.2.3 classroom management there are several methods to assess instructional quality: teacher ratings, student ratings or ratings of external observers (lüdtke, robitzsch, trautwein, & kunter, 2009). we measured the quality of classroom management with students’ ratings to avoid shared method variance (podsakoff, mackenzie, lee, & podsakoff, 2003) and because several studies have indicated that students are a reliable and valid source for judging instructional quality (e.g., fauth et al., 2014; lüdtke, trautwein, kunter, & baumert, 2006). research suggests that teacher and student ratings of classroom management are highly congruent (e.g. kunter & baumert, 2006). students responded to all classroom management items using a 4-point likert scale (1 = strongly disagree, 4 = strongly agree). classroom disturbances were assessed using six items (e.g., “in mathematics class it takes a long time at the beginning of the lesson until the students have settled down and started working”), giving examples of wasting time in class and student disruptions. a high score on this scale represented a high rate of classroom disturbances. monitoring was assessed using three items (e.g., “in mathematics our teacher always knows what is going on in the classroom”). both scales were developed in a previous project (baumert, gruehn, heyn, köller, & schnabel, 1997) and have been validated in several other studies (e.g., kunter et al., 2007). two-level confirmatory factor     seiz et al | f l r     63 analysis and model difference tests revealed a significantly better fit for a two-factor model of the two aspects of classroom management than for a global factor. to estimate whether the individual student ratings can be conceptualised as indicators for behaviour on the teacher level, we followed recommendations by marsh et al. (2009). the reliability and agreement of the ratings on the teacher level was calculated using intra-class correlations (icc) and the average deviation index (adm) of the manifest scales (lüdtke et al., 2006). the icc1 indicated the amount of variance among groups; in our case it reflected differences in classroom management ratings among teachers. the icc2 described the reliability of the group-mean rating of the whole scale, taking into account the number of raters. it can be interpreted in a similar manner as cronbach’s alpha. the adm is a means for assessing agreement within the group. it represents the average individual deviation from the group mean and is expressed in the metric of the original scale. there were substantial differences in ratings of classroom disturbances (icc1 = .33) and monitoring (icc1 = .23) among the teachers in our sample. both scales showed good reliability on the class level (classroom disturbances icc2 = .92; monitoring icc2 = .87). the adms were at .70 for classroom disturbances and at .64 for monitoring, indicating good agreement, with average individual ratings differing less than one point of the scale from the group mean. 2.2.4 control variable school type was included as a control variable on the teacher level (dummy coded: high track versus lower tracks). 2.3 statistical analysis our data has a hierarchical structure with students being nested in teachers. we therefore analysed the data using multilevel modeling, which overcomes the violation of the independence of observations and produces correct standard errors (hox, 2010). teacher resources were assessed on the teacher level. ratings of classroom management were assessed on the student level. we chose the teacher level and not the class level as our unit of analysis, since our focus is on the relevance of teacher resources. we combined multilevel modeling with structural equation modeling, thus correcting measurement errors. all constructs were estimated as latent factors with multiple indicators using mplus (muthén & muthén, 1998-2010). in our analysis, classroom management was modeled as a latent factor simultaneously on the individual level and the teacher level. with this doubly latent approach we followed the recommendations by marsh et al. (2009), correcting measurement error on both levels as well as sampling error on the teacher level. to test our hypotheses that the relation between pedagogical/psychological knowledge and classroom management is moderated by teachers’ exhaustion, we used the latent moderation structural equation approach (lms; klein & moosbrugger, 2000) implemented in mplus. by using latent predictors and calculating the interaction term of latent predictors we overcame the problem of manifest moderation analysis, in which the multiplicative term is affected particularly by measurement error (klein, 2000). the lms approach corrects measurement error in the predictor terms as well as in the multiplicative interaction term, leading to unbiased estimates for interaction effects. following the suggestion of klein and moosbrugger (2000), the latent factors were entered as predictors and then a multiplicative term of these two latent factors was formed, resulting in the following equation for the between-level (schermelleh-engel, kerwer, & klein, 2014): η b = α + γ 1b ξ 1b + γ 2b ξ 2b + γ 3b ξ 1b ξ 2b + ζ b (1)     seiz et al | f l r     64 this analytical approach is relatively new, computationally intensive and has rarely been applied in research on learning and instruction. we calculated two separate models for each aspect of classroom management due to the computational complexity. the rate of missing values was acceptable for most variables (0.7 % for student ratings; 8.3 % for emotional exhaustion) except for pedagogical/psychological knowledge (19.5 %; see table 1). this test was conducted only in the first assessment. the high percentage of missing data for the knowledge scores emerged as not all teachers participating in the second assessment (our sample of analysis) had completed the knowledge test in the first assessment. we analysed the selectivity of teacher respondents vs. non-respondents regarding demographic variables and teachers’ emotional exhaustion. because there were no significant differences between these groups and thus no indication for systematic missing values (schafer & graham, 2002), we used the effective full information maximum likelihood (fiml) algorithm (enders & bandalos, 2001) to estimate missing values in the following analysis. all significance testing was undertaken at the .05 level. for the calculation of practical effect sizes of multilevel analysis the following formula was employed (reyes, brackett, rivers, white, & salovey, 2012): 𝛿 = ! !!!!!! . while γ is the association between predictor and outcome variable, 𝜏!! and 𝜎! are the betweenand within-group variances of the outcome variable (from the unconditional model). reyes et al. (2012) states that δ can be interpreted similarly to cohen’s (1988) d. 3. results 3.1 preliminary analysis we conducted zero-order correlations on the teacher level between all latent factors involved in the analysis (see table 3).     seiz et al | f l r     65 table 3 latent standardized correlations of the study variables variables 1 2 3 4 5 teacher ratings 1 pedagogical/psychological knowledge ― -.18 -.10 .11 .22* 2 emotional exhaustion ― .02 .04 -.05 student ratings 3 classroom disturbances ― -.78* -.11 4 monitoring ― -.29* 5 school type ― note. school type is dummy-coded: high track versus low track. * p < .05. 3.2 results of the latent moderation models after our preliminary analysis, we calculated two separate models of latent interaction. pedagogical/psychological knowledge and the moderator emotional exhaustion were entered as predictors. then the multiplicative term of the two latent factors pedagogical/psychological knowledge and emotional exhaustion was added as the third predictor. the dependent variables were either students’ ratings of classroom disturbance or monitoring-ratings. we controlled for school type by including it as an additional predictor in the models. although fit indices for the lms approach have not yet been developed (see table 2 for fit indices for the measurement models without interaction term), it is possible to test the models with interaction effect against models without interaction effect using log likelihood differences, which are χ-distributed (klein & moosbrugger, 2000). the results of the difference tests revealed that the models with interaction term fit the data significantly better than models without interaction term, indicating that a significant interaction effect existed in both models (see table 4).     seiz et al | f l r     66 table 4 latent regression on teachers’ classroom management behavior with pedagogical/psychological knowledge as predictor and emotional exhaustion as moderator model 1 model 2 classroom disturbances monitoring variable b (se) δ b (se) δ intercept 2.24*(.04) 2.63*(.03) school type -.08*(.08) -.11 -.22*(.06) -.63 pedagogical/psychological knowledge -.03*(.06) -.04 .06*(.04) .17 emotional exhaustion .06*(.05) .09 -.04*(.04) -.12 ppk x ee .11*(.01) .16 -.10*(.03) -.29 r² .07 .14 note. model fit indices for lms not yet provided by mplus. ppk = pedagogical/psychological knowledge; ee = emotional exhaustion; b = unstandardised regression coefficient; se = standard error; δ = effect size. * p < .05. the effect sizes for the interaction effects can be considered small. following recommendations by aiken and west (1991) we plotted the interactions using different levels of the moderator. the three slopes represent different levels of the moderator emotional exhaustion (one standard deviation below the mean, the mean, and one standard deviation above the mean; see figures 1 and 2). additionally, we tested whether the simple slopes differed significantly from zero, meaning that the slope for the chosen value of the moderator was significant. since the interactions were disordinal, there can be no valid interpretation of the main effects (aiken & west, 1991). for the prediction of monitoring there also was a significant interaction between knowledge and emotional exhaustion. testing the simple slopes revealed that only the slope for a large amount of knowledge and a low level of emotional exhaustion was significant (see figure 2), indicating that only teachers with a high level of knowledge experiencing a low level of exhaustion showed better monitoring.     seiz et al | f l r     67 * p < .05. figure 1. interaction effect of pedagogical/psychological knowledge (ppk) and emotional exhaustion (ee) on classroom disturbances * p < .05. figure 2. interaction effect of pedagogical/psychological knowledge (ppk) and emotional exhaustion (ee) on monitoring 4. discussion the aim of this study was to analyse the joint relevance of teachers’ cognitive and emotional resources for classroom management. by analysing the distinct interplay of these resources we extended research in the area of teacher competence. our results indicate that neither cognitive nor 2,00 2,10 2,20 2,30 2,40 2,40 2,50 2,60 2,70 2,80 c la ss ro om d is tu rb an ce s ppk (-1sd) ppk (+1sd) ee (-1 sd) ee (mean) ee (+1 sd) * ppk (-1sd) ppk (+1sd) ee (-1 sd) ee (mean) ee (+1 sd) m on ito rin g *     seiz et al | f l r     68 emotional resources alone are linked to students’ ratings of classroom management, as there were no significant bivariate correlations or main effects. still, significant interaction effects illustrate that teachers’ cognitive and emotional resources interact. the results of both interaction models reflect the hypothesised mechanism of interplay between the resources: only the combination of knowledge and a low level of emotional exhaustion is associated with ratings of effective classroom management (a low level of classroom disturbances or a high level of monitoring). these results confirm hypotheses 2a and 2b. however, knowledge does not predict classroom management when the level of emotional exhaustion is high or average (hypotheses 1a and 1b). our results indicate that pedagogical/psychological knowledge alone may not be sufficient for effective classroom management but rather that cognitive and emotional resources are synergistic: only the combination of resources results in better classroom management. these results could be interpreted as potential support for our theoretical argumentation following information processing theory. a high level of emotional exhaustion may influence teachers’ information processing capacity and consequently teachers will not be able to process their knowledge extensively. in a similar vein the results can be interpreted through the lens of ego depletion theory. teachers experiencing high emotional exhaustion need to intensively regulate their emotions during instruction (näring, briët, & brouwers, 2006). this emotional labour may deplete volitional resources for consecutive higher-order cognitive activities, like applying professional knowledge in challenging classroom situations. no matter which theoretical approach is followed, processing of knowledge fails if teachers are highly exhausted, and classroom management is less effective. our study integrated several innovative aspects. first, we combined two theoretical approaches to teacher competence, which have not yet been brought together empirically. through analysing the interaction of cognitive and emotional resources we aimed to detect relevant psychological processes influencing teacher behaviour. second, with our test of teachers’ pedagogical/psychological knowledge we introduced an objective and direct measure of teachers’ cognitive resources, and thus went beyond subjective or distal measures (e.g., course work) to assess teacher knowledge. third, we applied an advanced methodological approach by using latent moderation analysis with multilevel data which rarely has been applied in educational research but overcomes problems of measurement error of the multiplication term (klein, 2000). 4.1 limitations and areas for future research some limitations of this study need to be considered. first, the causal direction of our argumentation and interpretation of our results needs further proof. due to our longitudinal design and the temporal ordering of our variables we concluded that the interplay between knowledge and emotional exhaustion has an effect on later classroom management and that prior teacher resources cannot be affected by later classroom management problems with the classes that provided the ratings; however, we were not able to control prior levels of classroom management. there are several studies indicating that teacher stress and emotional exhaustion may be a consequence of problems with classroom management, and thus reciprocal effects seem likely (e.g., chaplain, 2008; dicke, parker, marsh, et al., 2014). further, problems with classroom management may also impact student’s functioning and behavior (helmke & renkl, 1993; luckner & pianta, 2011), which then may influence teachers’ in-class experiences and thus affect teachers’ resources in return. the relation of teacher resources, classroom management and student functioning is much more complex and our study was only able to focus on some of these relations. more research and different designs are needed to disentangle the different relations, especially between classroom management and teachers’ emotional resources. studies using cross-lagged designs could help researchers approach this question.     seiz et al | f l r     69 as our results remained stable when using ratings of emotional exhaustion of the second assessment, they can also support our argumentation. however, the fact that the time interval between the first and the second assessment was 14 months needs to be considered an additional limitation as pedagogical/psychological knowledge is likely to still change after the induction phase. further, another study based on a different teacher sample showed a direct association between pedagogical/psychological knowledge and classroom management (voss et al., 2014), which contrasts with our findings. this study assessed teacher knowledge data at the beginning of the induction phase, using a slightly different subsample. we conducted several analyses in order to interpret these differences. as participants did not differ substantially and the knowledge test was invariant across measurements we conclude differences in results to be on the conceptual level. apparently, during the evolvement of the induction phase and the beginning of regular teaching emotional functioning becomes more important, explaining our findings of interaction effects and our lack of main effects. further, in our study, we combined two approaches to assessing teacher resources focusing on their professional knowledge and emotional exhaustion. however, there are other relevant aspects of teacher competence such as motivational orientations, and other domains of their cognitive resources such as beliefs (baumert & kunter, 2013). some researchers already have approached the question as to how different competence aspects influence each other (dicke et al., 2014; klusmann, kunter, voss, & baumert, 2012). however, we argue that instead of analysing these associations with regard to teacher variables as outcomes it would be interesting to study the additional impact of these interplays on instructional or student outcomes. in general, alternative explanations for the results might be applied. for instance, it would be possible that teachers high in pedagogical/psychological knowledge are very self-efficacious regarding their classroom management. these favourable motivational orientations could also help to apply knowledge during instruction, resulting in effective classroom management (morris-rothschild & brassard, 2006). regarding the generalisability of our results we need to point out some specific characteristics of our sample. first, our sample was not representative. second, our sample consisted of secondary school mathematics teachers and their students. as our research question was not subject-specific we would expect similar results in samples of teachers of other subjects. third, our sample included teachers with relatively little teaching experience. since we based our arguments on information processing theory, our results need to be interpreted with caution: research has shown differences in information processing between expert and novice teachers (swanson, o’connor, & cooney, 1990; wolff et al. 2014). more experienced teachers possess more automatised routines and schemas which claim less information processing capacity (e.g., feldon, 2007). further, differences between experts and novices in terms of classroom management exist in their perceptions of classroom events, in that novices have problems noticing simultaneous class events (van den bogert, van bruggen, kostons, & jochems, 2014). this could be interpreted in a way that that classroom management claims more information processing capacity from novices than from experts (sabers, cushing, & berliner, 1991). according to such findings it is possible that the joint relevance of teacher resources analysed in this study could apply especially to teachers with little experience and thus may be overestimated in our sample. also, it could be that experienced teachers’ emotions differ from those of less experienced teachers when reacting to classroom management situations (sutton & wheatley, 2003). it would be highly recommendable for future studies to explore this relation with samples of more experienced teachers.     seiz et al | f l r     70 4.2 theoretical and educational implications teaching is challenging and our results show that resources of successful teachers interact in complex ways. we argue that this expanded view on teacher resources is highly relevant for teacher education and pedagogical practice. teacher education aims to prepare students for their professional career, yet the understanding of teacher resources focuses foremost on teachers’ professional and practical knowledge (korthagen & kessels, 1999). also, teacher selection programs often focus on knowledge, yet our results indicate that knowledge alone might not be sufficient. we argue for a combined approach in teacher education that focuses on the development of professional knowledge as well as on teachers’ emotional resources. several authors highlight the importance of acknowledging teaching as an emotional practice (chang, 2009; sutton & wheatley, 2003).teachers’ emotional resources are also relevant for students. students profit from having warm and highly supportive student-teacher interactions in regard to their academic development, self-regulation and their executive control (e.g. roorda, koomen, spilt, & oort, 2011; williford, whittaker, vitiello, & downer, 2013). as emotions often emerge in various classroom situations regulation of emotions is especially relevant for effective classroom management (sutton & wheatley, 2003). nevertheless, little is known about teachers’ emotional processes in such situations (chang, 2009). chang (2013) showed that the appraisal of classroom incidents involving problematic student behaviour is related to unpleasant emotions, which are associated with burnout. this association between negative emotions and burnout was in turn mediated by different coping strategies. chang (2009) argues that teachers should learn to regulate their emotions by using reappraisal techniques and coping mechanisms. another way to help teachers deal with their emotions has emerged. mindfulness training programs equip teachers with techniques to integrate mindfulness skills in the classroom (flook, goldberg, pinger, bonus, & davidson, 2013) and thereby cope with stress more effectively (roeser et al., 2013). after completing a mindfulness training program, participants showed fewer burnout symptoms, performed better on attentional tasks and even organised their classrooms better than those in a control group (roeser et al., 2013). it would seem beneficial to incorporate such training on emotion regulation in teacher education and professional development programs. helping teachers understand their emotions and enhance their competence in regulating them certainly would not replace teachers’ professional knowledge; however, knowledge of classroom management strategies may help teachers prevent later exhaustion (dicke et al., 2015; klusmann et al., 2012). we argue that teacher education and continuing professional development programmes would profit from broadening their scopes and acknowledging the relevance of cognitive and emotional aspects of teacher competence and their potential interplay. helping teachers address their emotions during teacher education and continuously supporting them in doing this through professional development would have two benefits: synergies between teachers’ cognitive and emotional resources may be promoted, enabling teachers to make the most use of their knowledge in the classroom; and, in the long run, work-related stress and burnout may be lessened or even avoided.     seiz et al | f l r     71 keypoints teachers’ cognitive and emotional resources interact. teachers’ knowledge is not related per se to ratings of classroom management. teachers’ knowledge predicts classroom management only when emotional exhaustion is low. acknowledgments this study used data from the coactiv-r research project which was funded by the max planck society’s strategic innovation fund (2008–2010). we would like to thank patricia alexander for her helpful comments on a previous version of this paper. references aiken, l. s., & west, s. g. (1991). multiple regression. testing and interpreting interactions. newbury park, ca: sage publications baumeister, r. f., gailliot, m., dewall, c. n., & oaten, m. (2006). self-regulation and personality: how interventions increase regulatory success, and how depletion moderates the effects of traits on behavior. journal of personality, 74(6), 1773–1802. doi: 10.1111/j.1467-6494.2006.00428.x baumert, j., gruehn, s., heyn, s., köller, o., & schnabel, k. u. (1997). bildungsverläufe und psychosoziale entwicklung im jugendalter (biju). dokumentation, band 1. skalen längsschnitt i, welle 1-4 [learning processes, educational careers and psychological development in adolescence and young adulthood (biju). documentation, vol. 1, scales, waves 1-4]. berlin, germany: max planck institute for human development. baumert, j., & kunter, m. (2013). the coactiv model of teachers’ professional competence. in m. kunter, j. baumert, w. blum, u. klusmann, s. krauss, & m. neubrand (eds.), cognitive activation in the mathematics classroom and professional competence of teachers (vol. 8, pp. 25–48). new york, ny: springer. baumert, j., kunter, m., blum, w., brunner, m., voss, t., jordan, a., . . . tsai, y.-m. (2010). teachers’ mathematical knowledge, cognitive activation in the classroom, and student progress. american educational research journal, 47(1), 133–180. doi: 10.3102/0002831209345157 becker, e., goetz, t., morger, v., & ranellucci, j. (2014). the importance of teachers' emotions and instructional behavior for their students' emotions: an experience sampling analysis. teaching and teacher education, 43, 15–26. doi: 10.1016/j.tate.2014.05.002 brouwers, a., & tomic, w. (1999). teacher burnout, perceived self-efficacy in classroom management, and student disruptive behaviour in secondary education. curriculum and teaching, 14(2), 7–26. doi: 10.7459/ct/14.2.02 chang, m.-l. (2009). an appraisal perspective of teacher burnout: examining the emotional work of teachers. educational psychology review, 21(3), 193–218. doi: 10.1007/s10648-009-9106-y chang, m.-l. (2013). toward a theoretical model to understand teacher emotions and teacher burnout in the context of student misbehavior: appraisal, regulation and coping. motivation and emotion, 37(4), 799– 817. doi: 10.1007/s11031-012-9335-0 chaplain, r. p. (2008). stress and psychological distress among trainee secondary teachers in england. educational psychology, 28(2), 195–209. doi: 10.1080/01443410701491858     seiz et al | f l r     72 cohen, j. (1988). statistical power analysis for the behavioral science (2nd ed.). hillsdale, nj: lawrence erlbaum associates. de jong, r., mainhard, t., van tartwijk, j., veldman, i., verloop, n., & wubbels, t. (2014). how preservice teachers' personality traits, self-efficacy, and discipline strategies contribute to the teacher– student relationship. british journal of educational psychology, 84, 294–310. doi: 10.1111/bjep.12025 depaepe, f., verschaffel, l., & kelchtermans, g. (2013). pedagogical content knowledge: a systematic review of the way in which the concept has pervaded mathematics educational research. teaching and teacher education, 34, 12–25. doi: 10.1016/j.tate.2013.03.001 dicke, t., parker, p. d., holzberger, d., kunina-habenicht, o., kunter, m., & leutner, d. (2015). beginning teachers' efficacy and emotional exhaustion: latent changes, reciprocity, and the influence of professional knowledge. contemporary educational psychology, 41, 62–72. dicke, t., parker, p. d., marsh, h. w., kunter, m., schmeck, a., & leutner, d. (2014). self-efficacy in classroom management, classroom disturbances, and emotional exhaustion: a moderated mediation analysis of teacher candidates. journal of educational psychology, 106(2), 569–583. doi: 10.1037/a0035504 doyle, w. (2006). ecological approaches to classroom management. in c. m. evertson & c. s. weinstein (eds.), handbook of classroom management (pp. 97–125). mahwah, nj: lawrence erlbaum. emmer, e. t., & evertson, c. (2013). classroom management for middle and high school teachers. boston, ma: pearson. emmer, e. t., & stough, l. m. (2001). classroom management: a critical part of educational psychology, with implications for teacher education. educational psychologist, 36(2), 103–112. doi: 10.1207/s15326985ep3602_5 enders, c. k., & bandalos, d. l. (2001). the relative performance of full information maximumlikelihood estimation for missing data in structural equation models. structural equation modeling: a multidisciplinary journal, 8(3), 430–457. doi: 10.1207/s15328007sem0803_5 enzmann, d., & kleiber, d. (1989). helfer-leiden: stress und burnout in psychosozialen berufen [helpers' suffering: stress and burnout in psycho-social professions]. heidelberg, germany: roland asanger verlag. evertson, c. m., & weinstein, c. s. (2006). classroom management as a field of inquiry. in c. m. evertson & c. s. weinstein (eds.), handbook of classroom management (pp. 3–15). mahwah, nj: lawrence erlbaum. fauth, b., decristan, j., rieser, s., klieme, e., & büttner, g. (2014). student ratings of teaching quality in primary school: dimensions and prediction of student outcomes. learning and instruction, 29(0), 1–9. doi: 1016/j.learninstruc.2013.07.001 feldon, d. f. (2007). cognitive load and classroom teaching: the double-edged sword of automaticity. educational psychologist, 42(3), 123–137. doi: 10.1080/00461520701416173 feuerhahn, n., stamov-roßnagel, c., wolfram, m., bellingrath, s., & kudielka, b. m. (2013). emotional exhaustion and cognitive performance in apparently healthy teachers: a longitudinal multi-source study. stress and health, 29(4), 297–306. doi: 10.1002/smi.2467 flook, l., goldberg, s. b., pinger, l., bonus, k., & davidson, r. j. (2013). mindfulness for teachers: a pilot study to assess effects on burnout, and teaching efficacy. mind, brain, and education, 7(3), 182– 195. doi: 10.1111/mbe.12026 frenzel, a., goetz, t., stephens, e., & jacob, b. (2009). antecedents and effects of teachers' emotional experiences: an integrated perspective and empirical test. in p. a. schutz & m. zembylas (eds.), advances in teacher emotions research: the impact on teachers' lives (pp. 129–146). new york, ny: springer. hakanen, j. j., bakker, a. b., & schaufeli, w. b. (2006). burnout and work engagement among teachers. journal of school psychology, 43, 495–513. doi: 10.1016/j.jsp.2005.11.001     seiz et al | f l r     73 helsing, d. (2007). regarding uncertainty in teachers and teaching. teaching and teacher education, 23(8), 1317–1333. doi: 10.1016/j.tate.2006.06.007 helmke, a., & renkl, a. (1993). unaufmerksamkeit in grundschulklassen: problem der klasse oder des lehrers? [inattention in primary classrooms: problem of the class or of the teacher?]. zeitschrift für entwicklungspsychologie und pädagogische psychologie, 25(3), 185–205. hill, h. c., rowan, b., & ball, d. l. (2005). effects of teachers’ mathematical knowledge for teaching on student achievement. american educational research journal, 42(2), 371–406. doi: 10.3102/00028312042002371 hox, j. j. (2010). multilevel analysis. techniques and application (2nd ed.). east sussex: routledge. johns, m., inzlicht, m., & schmader, t. (2008). stereotype threat and executive resource depletion: examining the influence of emotion regulation. journal of experimental psychology: general, 137(4), 691–705. doi: 10.1037/a0013834 jennings, p. a., & greenberg, m. t. (2009). the prosocial classroom: teacher social and emotional competence in relation to student and classroom outcomes. review of educational research, 79(1), 491–525. doi: 10.3102/0034654308325693 keller, m., chang, m.-l., becker, e., goetz, t., & frenzel, g. (2014). teachers’ emotional experiences and exhaustion as predictors of emotional labor in the classroom: an experience sampling study. frontiers in psychology, 5(1442). doi: 10.3389/fpsyg.2014.01442 klein, a. (2000). moderatormodelle: verfahren zur analyse von moderatoreffekten in strukturgleichungsmodellen [moderator models: methods for the analysis of moderator effects in structural equation models]. hamburg, germany: kovač. klein, a., & moosbrugger, h. (2000). maximum likelihood estimation of latent interaction effects with the lms method. psychometrika, 65(4), 457–474. doi: 10.1007/bf02296338 kleinsorge, t., diestel, s., scheil, j., & niven, k. (2014). burnout and the fine-tuning of cognitive resources. applied cognitive psychology, 28(2), 274–278. doi: 10.1002/acp.2999 klusmann, u., kunter, m., trautwein, u., lüdtke, o., & baumert, j. (2008). teachers' occupational wellbeing and quality of instruction: the important role of self-regulatory patterns. journal of educational psychology, 100(3), 702–715. doi: 10.1037/0022-0663.100.3.702 klusmann, u., kunter, m., voss, t., & baumert, j. (2012). berufliche beanspruchung angehender lehrkräfte: die effekte von persönlichkeit, pädagogischer vorerfahrung und professioneller kompetenz [occupational stress of beginning teachers: the effects of personality, pedagogical experience and professional competence]. zeitschrift für pädagogische psychologie, 26, 275–290. doi: 10.1024/1010-0652/a000078 korthagen, f., & kessels, j. p. (1999). linking theory and practice: changing the pedagogy of teacher education. educational researcher, 28(4), 4–17. doi: 10.3102/0013189x028004004 kounin, j. s. (1970). discipline and group management in classrooms. new york, ny: holt, rinehart & winston. krauss, s., brunner, m., kunter, m., baumert, j., blum, w., neubrand, m., & jordan, a. (2008). pedagogical content knowledge and content knowledge of secondary mathematics teachers. journal of educational psychology, 100(3), 716–725. doi: 10.1037/0022-0663.100.3.716 kunter, m., & baumert, j. (2006). who is the expert? construct and criteria validity of student and teacher ratings of instruction. learning environments research, 9(3), 231–251. doi: 10.1007/s10984-006-90157 kunter, m., baumert, j., & köller, o. (2007). effective classroom management and the development of subject-related interest. learning and instruction, 17, 494–509. doi: 10.1016/ j.learninstruc.2007.09.002 liston, d., whitcomb, j., & borko, h. (2006). too little or too much: teacher preparation and the first years of teaching. journal of teacher education, 57(4), 351–358. doi: 10.1177/0022487106291976     seiz et al | f l r     74 luckner, a. e., & pianta, r. c. (2011). teacher–student interactions in fifth grade classrooms: relations with children's peer behavior. journal of applied developmental psychology, 32, 257–266. doi: 10.1016/j.appdev.2011.02.010 lüdtke, o., robitzsch, a., trautwein, u., & kunter, m. (2009). assessing the impact of learning environments: how to use student ratings of classroom or school characteristics in multilevel modeling. contemporary educational psychology, 34(2), 120–131. doi: 10.1016/j.cedpsych.2008.12.001 lüdtke, o., trautwein, u., kunter, m., & baumert, j. (2006). reliability and agreement of student ratings of the classroom environment: a reanalysis of timss data. learning environments research, 9(3), 215– 230. doi: 10.1007/s10984-006-9014-8 marsh, h. w., lüdtke, o., robitzsch, a., trautwein, u., asparouhov, t., muthén, b., & nagengast, b. (2009). doubly-latent models of school contextual effects: integrating multilevel and structural equation approaches to control measurement and sampling error. multivariate behavioral research, 44(6), 764–802. doi: 10.1080/00273170903333665 marzano, r. j., marzano, j. s., & pickering, d. j. (2003). classroom management that works researchbased strategies for every teacher. alexandria, va: association for supervision and curriculum development. maslach, c., jackson, s. e., & leiter, m. p. (1996). maslach burnout inventory manual (3rd ed.). palo alto, ca: counsulting psychologists press. maslach, c., & leiter, m. p. (1999). teacher burnout: a research agenda. in r. vandenberghe & a. m. huberman (eds.), understanding and preventing teacher burnout: a sourcebook of international research and practice (pp. 295–303). cambridge, uk: cambridge university press. maslach, c., schaufeli, w. b., & leiter, m. p. (2001). job burnout. annual review psychology, 52, 397– 422. doi: 10.1146/annurev.psych.52.1.397 mayer, r. e. (2012). information processing. in k. r. harris, s. graham, t. urdan, c. b. mccormick, g. m. sinatra, & j. sweller (eds.), apa educational psychology handbook, volume 1: theories, constructs, and critical issues (pp. 85–99). washington, dc: american psychological association. morris-rothschild, b. k., & brassard, m. r. (2006). teachers´ conflict management styles: the role of attachment styles and classroom mangement efficacy. journal of school psychology, 44, 105–121. doi: 10.1016/j.jsp.2006.01.004 muthén, l. k., & muthén, b. o. (1998-2010). mplus user's guide. sixth edition. los angeles, ca: muthén & muthén. näring, g., briët, m., & brouwers, a. (2006). beyond demand–control: emotional labour and symptoms of burnout in teachers. work & stress, 20(4), 303–315. doi: 10.1080/02678370601065182 nie, y., & lau, s. (2009). complementary roles of care and behavioral control in classroom management: the self-determination theory perspective. contemporary educational psychology, 34, 185–194. doi: 10.1016/j.cedpsych.2009.03.001 ophardt, d., & thiel, f. (2008). klassenmanagement als basisdimension der unterrichtsqualität [classroom management as a basic dimension for instructional quality]. in m. k. w. schweer (ed.), lehrer-schüler-interaktion [teacher-student-interaction] (2nd ed., pp. 259–282). wiesbaden, germany: vs verlag für sozialwissenschaften. park, s., & oliver, j. s. (2008). revisiting the conceptualisation of pedagogical content knowledge (pck): pck as a conceptual tool to understand teachers as professionals. research in science education, 38, 261–284. doi: 10.1007/s11165-007-9049-6 pianta, r. c. (2006). classroom management and relationships between children and teachers: implications for research and practice. in c. m. evertson & c. s. weinstein (eds.), handbook of classroom management (pp. 685–709). mahwah, nj: lawrence erlbaum. pianta, r. c., & hamre, b. k. (2009). conceptualization, measurement, and improvement of classroom processes: standardized observation can leverage capacity. educational researcher, 38(2), 109–119. doi: 10.3102/0013189x09332374     seiz et al | f l r     75 podsakoff, p. m., mackenzie, s. b., lee, j. y., & podsakoff, n. p. (2003). common method biases in behavioral research: a critical review of the literature and recommended remedies. journal of applied psychology, 88(5), 879–903. doi: 10.1037/0021-9101.88.5.879 reyes, m. r., brackett, m. a., rivers, s. e., white, m., & salovey, p. (2012). classroom emotional climate, student engagement, and academic achievement. journal of educational psychology, 104(3), 700–712. doi: 10.1037/a0027268 roehrig, a. d., turner, j. e., arrastia, m. c., christesen, e., mcelhaney, s., & jakiel, l. m. (2012). effective teachers and teaching: characteristics and practices related to positive student outcomes. in k. r. harris, s. graham, t. urdan, s. graham, j. m. royer, & m. zeidner (eds.), apa educational psychology handbook, volume 2: individual differences and cultural and contextual factors (pp. 501– 527). washington, dc: american psychological association. doi: 10.1037/13274-020 roeser, r. w., schonert-reichl, k. a., jha, a., cullen, m., wallace, l., wilensky, r., . . . harrison, j. (2013). mindfulness training and reductions in teacher stress and burnout: results from two randomized, waitlist-control field trials. journal of educational psychology, 105(3), 787–804. doi: 10.1037/a0032093 roorda, d. l., koomen, h. m., spilt, j. l., & oort, f. j. (2011). the influence of affective teacher–student relationships on students’ school engagement and achievement: a meta-analytic approach. review of educational research, 81(4), 493–529. doi: 10.3102/0034654311421793 sabers, d. s., cushing, k. s., & berliner, d. c. (1991). differences among teachers in a task characterized by simultaneity, multidimensionality, and immediacy. american educational research journal, 28(1), 63–88. doi: 10.2307/1162879 sandström, a., rhodin, i. n., lundberg, m., olsson, t., & nyberg, l. (2005). impaired cognitive performance in patients with chronic burnout syndrome. biological psychology, 69(3), 271–279. doi: 10.1016/j.biopsycho.2004.08.003 schafer, j. l., & graham, j. w. (2002). missing data: our view of the state of the art. psychological methods, 7(2), 147–177. doi: 10.1037//1082-989x.7.2.147 schermelleh-engel, k., kerwer, m., & klein, a. (2014). evaluation of model fit in nonlinear multilevel structural equation modeling. frontiers in psychology, 5, 1–11.doi: 10.3389/fpsyg.2014.00181 schmeichel, b. j., vohs, k., & baumeister, r. f. (2003). intellectual performance and ego depletion: role of the self in logical reasoning and other information processing. journal of personality and social psychology, 85(1), 33–46. doi: 10.1037/0022-3514.85.1.33 schwarzer, r., schmitz, g., & tang, c. (2000). teacher burnout in hong kong and germany: a crosscultural validation of the maslach burnout inventory. anxiety, stress and coping, 13, 309–326. doi: 10.1080/10615800008549268 shulman, l. s. (1986). those who understand: knowledge growth in teaching. educational researcher, 15(2), 4–14. doi: 10.2307/1175860 shulman, l. s. (1987). knowledge and teaching: foundations of the new reform. harvard educational review, 57(1), 1–22. skaalvik, e. m., & skaalvik, s. (2011). teacher job satisfaction and motivation to leave the teaching profession: relations with school context, feeling of belonging, and emotional exhaustion. teaching and teacher education, 27, 1029–1038. doi: 10.1016/j.tate.2011.04.001 statistisches bundesamt (2014). bildung und kultur. allgemeinbildende schulen. [education and culture. secondary schools]. wiesbaden: statistisches bundesamt. sutton, r. e. (2005). teachers' emotions and classroom effectiveness: implications from recent research. the clearing house, 78(5), 229–234. doi: 10.2307/30189914 sutton, r. e., & wheatley, k. f. (2003). teachers' emotions and teaching: a review of the literature and directions for future research. educational psychology review, 15(4), 327–358. doi: 10.1023/a:1026131715856     seiz et al | f l r     76 swanson, h. l., o’connor, j. e., & cooney, j. b. (1990). an information processing analysis of expert and novice teachers’ problem solving. american educational research journal, 27(3), 533–556. doi: 10.3102/00028312027003533 sweller, j., van merrienboer, j. j. g., & paas, f. g. w. c. (1998). cognitive architecture and instructional design. educational psychology review, 10(3), 251–296. doi: 10.1023/a:1022193728205 unterbrink, t., hack, a., pfeifer, r., buhl-grießhaber, v., müller, u., wesche, u., . . . bauer, j. (2007). burnout and effort–reward-imbalance in a sample of 949 german teachers. international archives of occupational and environmental health, 80, 433–441. doi: 10.1007/s00420-007-0169-0 van den bogert, n., van bruggen, j., kostons, d., & jochems, w. (2014). first steps into understanding teachers' visual perception of classroom events. teaching and teacher education, 37, 208–216. doi: 10.1016/j.tate.2013.09.001 van der linden, d., keijsers, g. p. j., eling, p., & schaijk, r. v. (2005). work stress and attentional difficulties: an initial study on burnout and cognitive failures. work & stress, 19(1), 23–36. doi: 10.1080/02678370500065275 voss, t., kunina-habenicht, o., & kunter, m. (2015). stichwort pädagogisches wissen von lehrkräften: empirische zugänge und befunde [teachers’ pedagogical knowledge: empirical approaches and findings]. zeitschrift für erziehungswissenschaft. doi: 10.1007/s11618-015-0626-6 voss, t., kunter, m., & baumert, j. (2011). assessing teacher candidates’ general pedagogical / psychological knowledge: test construction and validation. journal of educational psychology, 103(4), 952–969. doi: 10.1037/a0025125 voss, t., kunter, m., seiz, j., hoehne, v., & baumert, j. (2014). die bedeutung des pädagogischpsychologischen wissens von angehenden lehrkräften für die unterrichtsqualität [the impact of teachers’ general pedagogical and psychological knowledge on instructional quality]. zeitschrift für pädagogik, 60(2), 184–201. wang, m. c., haertel, g. d., & walberg, h. j. (1993). toward a knowledge base for school learning. review of educational research, 63(3), 249–294. doi: 10.3102/00346543063003249 williford, a. p., whittaker, j. e. v., vitiello, v. e., & downer, j. t. (2013). children’s engagement within the preschool classroom and their development of self-regulation. early education and development, 24, 162–187. doi: 10.1080/10409289.2011.628270 wolff, c. e., van den bogert, n., jarodzka, h., & boshuizen, h. p. a. (2015). keeping an eye on learning: differences between expert and novice teachers’ representations of classroom management events. journal of teacher education, 66(1), 68–85. doi: 10.1177/0022487114549810     seiz et al | f l r     77 appendix a standardized factor loadings of latent factors factors and indicators factor loadings (within level) factor loadings (between level) emotional exhaustion i often feel exhausted at school. ― .79 as a whole, i feel overworked. ― .65 i often notice how listless i am at school. ― .70 i sometimes feel really depressed at the end of a school day. ― .75 pedagogical/ psychological knowledge teaching methods ― .72 classroom management ― .35 classroom assessment ― .56 students’ heterogeneity ― .56 classroom disturbance in mathematics teaching is very often interrupted. .75 .99 in mathematics students talk among themselves the whole time. .76 .99 in mathematics students mess around the whole time. .70 .98 in mathematics it takes a very long time at the start of the lesson until the students have settled down and started working. .60 .95 in mathematics a lot of lesson time is wasted. .63 .95 in mathematics the lesson often starts late. .42 .82 monitoring in mathematics our teacher always knows what is going on in the classroom. .44 .87 in mathematics our teacher always checks our homework thoroughly. .60 .68 in mathematics our teacher makes sure that we pay attention. .45 .95 note. all loadings were significant at p < .05. microsoft word kärner et al_publication.docx ! ! ! ! ! frontline!learning!research!vol.5!no.!1!(2017)!16! 1). the items form a two-factor solution that accounts for 85.29 % of the total variance (compared to the one-factor solution, accounting for only 43.17 %). both items relating to students’ stress experience (r = .73) and both items of situational coping (r = .68) are correlated moderately to each other. for further analysis we used the factor scores as estimated values of the factors “situational coping” and “students’ stress experience”. 4.2.5 classroom demands classroom demands—as “objective” characteristics of the classroom context—were operationalised by the amount of student-centred learning and by the quality of cognitive challenge during education. student-centring is defined by phases of individual or group work where learners work independently from the teacher on complex problems. it was assessed via video-based timesampling analysis using a defined category-system we adopted from seidel et al. (2001). time intervals of 15 seconds each were coded. afterwards, the single coded 15-second intervals were aggregated to 10-minute intervals via sum scores, synchronising the context conditions—in terms of “micro-segments” of classroom context—and the person-related data. to assess the reliability of the codings, one third of the videos were coded by two independent coders, finding a satisfactory cohen’s kappa of .73. overall we found a mean of m = 3.16 minutes (sd = 3.63, min. = 0, max. = 9.5) studentcentring per 10 minutes of education. cognitive challenge during class was coded by ratings of two independent coders: they coded one third of the videos (between-coder-correlation r = .82) on the basis of the observed classroom discussion and the learning material the students worked on. in order to assess content-related difficulty, we referred to the curriculum of the corresponding subject “economic business processes” and to bloom’s taxonomy of learning objectives. time intervals of 1 minute each were coded, and we used a four point likert-type scale based on bloom’s taxonomy to assess the complexity of learning contents (0 = “applying”, 1 = “analysing”, 2 = “synthesizing”, 3 = “evaluating”; cf. bloom et al., kärner&et &al & & | f l r ! ! 24! 1956). afterwards, the single coded 1-minute intervals were arithmetically aggregated to 10-minute intervals. an overall mean of m = 1.50 (sd = .82, min. = .44, max. = 3.00) of cognitive challenge in the observed lessons was found. assessing the factorial structure of the amount of student-centred learning and the cognitive challenge during education, we applied an exploratory factor analysis with varimax rotation and referred to the kaiser criterion (eigenvalue > 1). the items form a one-factor solution that accounts for 87.55 % of the total variance, with both variables correlated moderately to each other (r = .75). afterwards, we used the factor scores as estimated values of the factor “classroom demands”. 4.3 statistical analysis 4.3.1 previous analysis pearson product-moment correlations were calculated in order to identify multicollinearity and to check characteristics of the independent variables. 4.3.2 multilevel analysis against our theoretical background, it seems to be crucial to measure time-varying states and objective context conditions in a synchronic way in addition to relatively enduring characteristics. appropriate interrelationships can be investigated using multilevel analytic methods, as they provide the opportunity to simultaneously analyse different hierarchical data levels. in this context, longitudinal data can be seen as hierarchical data, with repeated measurements nested within persons (bryk & raudenbush, 1992; goldstein et al., 1994; heck & thomas, 2009; hox, 2002; nezlek, 2007). scollon et al. (2003) point out that multilevel modelling is useful in analysing continuously sampled data because multiple data points are nested in a single individual. in our case, students’ states are not only nested within persons but also within situations that are defined as “micro-segments” of the classroom context prevailing at the time of measurement. therefore we applied a cross-classified multilevel model (cf. heck et al., 2010). cross-classification considers that the multiple state-measures are not only nested within persons but that they also belong to observation units of the classroom context at level 2 corresponding to the time of measurement (cf. goldstein, 1994; hill & goldstein, 1998). most of the existing approaches in researching person-situation interactions analyse how variations in situations affect individuals depending on their dispositions via analysis-of-variance or multivariate regression analysis (carver & scheier, 2008; cronbach & snow, 1977). but these approaches do not usually consider the hierarchical data structure implied by multiple measures, that is important because investigating trait-treatment interactions should focus on analysing how persons show changes in an outcome variable over time or across situations (yeh, 2012). multilevel methods have the advantage of combining the different approaches within a simultaneous analysis, therefore providing a potential method for investigating person3.0.co;2-q little, r. j. a. & rubin, d. b. (2002). statistical analysis with missing data. new jersey: wiley. looser, r. r., metzenthin, p., helfricht, s., kudielka, b. m., loerbroks, a., thayer, j. f., & fischer, j. e. (2010). cortisol is significantly correlated with cardiovascular responses during high levels of stress in critical care personnel. psychosomatic medicine, 72 (3), 281–289. doi: 10.1097/psy.0b013e3181d35065 luria, r. e. (1975). the validity and reliability of the visual analogue mood scale. journal of psychiatric research, 12 (1), 51–57. doi: 10.1016/0022-3956(75)90020-5 malik, m., bigger, j. t., camm, a. j., kleiger, r. e., malliani, a., moss, a. j., & schwartz, p. j. [task force of the european society of cardiology and the north american society of pacing and electrophysiology] (1996). heart rate variability: standards of measurement, physiological interpretation, and clinical use. european heart journal, 17 (3), 354–381. michaud, k., matheson, k., kelly, o., & anisman, h. (2008). impact of stressors in a natural context on release of cortisol in healthy adult humans: a meta-analysis. stress, 11 (3), 177–197. doi: 10.1080/10253890701727874 nezlek, j. b. (2007). a multilevel framework for understanding relationships among traits, states, situations and behaviours. european journal of personality, 21 (6), 789–810. doi: 10.1002/per.640 peugh, j. l. & enders, c. k. (2005). using the spss mixed procedure to fit cross-sectional and longitudinal multilevel models. educational and psychological measurement, 65 (5), 717–741. doi: 10.1177/0013164405278558 podsakoff, p. m., mackenzie, s. b., lee, j.-y., & podsakoff, n. p. (2003). common method biases in behavioral research: a critical review of the literature and recommended remedies. journal of applied psychology, 88 (5), 879–890. doi: 10.1037/0021-9010.88.5.879 pruessner, j. c., kirschbaum, c., meinlschmid, g., & hellhammer, d. h. (2003). two formulas for computation of the area under the curve represent measures of total hormone concentration versus time-dependent change. psychoneuroendocrinology, 28 (7), 916–931. doi: 10.1016/s0306-4530(02)00108-7 raessler, s., rubin, d. b., & schenker, n. (2008). incomplete data: diagnosis, imputation, and estimation. in e. d. de leeuw, j. j. hox, & d. a. dillman (eds.), international handbook of survey methodology, (pp. 370–386). new york: lawrence erlbaum. raudenbush, s. w. & bryk, a. s. (2002). hierarchical linear models. applications and data analysis methods, second edition. london: sage. rieser, s., fauth, b., decristan, j., klieme, e., & büttner, g. (2013). the connection between primary school students’ self-regulation in learning and perceived teaching quality. journal of cognitive education and psychology, 12 (2), 138–156. doi:10.1891/1945-8959.12.2.138 kärner&et &al & & | f l r ! ! 40! rösler, u., gebele, n., hoffmann, k., morling, k., müller, a., rau, r., & stephan, u. (2010). cortisol – ein geeigneter physiologischer indikator für belastungen am arbeitsplatz? [cortisol – a useful physiological parameter for work-related stress?] zeitschrift für arbeitsund organisationspsychologie, 54 (2), 68–82. doi: 10.1026/0932-4089/a000011 rost, d. h., sparfeldt, j. r., & schilling, s. r. (2007). disk-gitter mit skslf-8. differentielles schulisches selbstkonzept-gitter mit skala zur erfassung des selbstkonzepts schulischer leistungen und fähigkeiten [scale for assessing self-concept of school related achievement and skills]. göttingen: hogrefe. roubinov, d., hagan m. j., & luecken l. j. (2012). if at first you don't succeed: the neuroendocrine impact of using a range of strategies during social conflict. anxiety, stress, & coping: an international journal, 25 (4), 397–410. doi: 10.1080/10615806.2011.613459 rubin, d. b. (1987). multiple imputation for nonresponse in surveys. hoboken, nj: wiley-ieee. saxbe, d. e. (2008). a field (researcher’s) guide to cortisol: tracking hpa axis functioning in everyday life. health psychology review, 2 (2), 163–190. doi: 10.1080/17437190802530812 schunk, d. h. & zimmerman, b. j. (eds.) (1994). self-regulation of learning and performance. issues and edicational applications. hillsdale: lawrence erlbaum. scollon, ch. n., kim-prieto, ch., & diener, e. (2003). experience sampling: promises and pitfalls, strengths and weaknesses. journal of happiness studies, 4 (1), 5–34. doi: 10.1023/a:1023605205115 seidel, t., dalehefte, i. m., & meyer, l. (2001). videoanalysen – beobachtungsschemata zur erfassung von „sicht-strukturen“ im physikunterricht [video-based analysis – observation patterns for assessing conditions in classroom in physics education]. in m. prenzel, r. duit, m. euler, m. lehrke, & t. seidel (eds.), erhebungsund auswertungsverfahren des dfg-projekts „lehr-lern-prozesse im physikunterricht – eine videostudie“, (pp. 41–58). kiel: ipnmaterialien. seifried, j. (2004). fachdidaktische variationen in einer selbstorganisationsoffenen lernumgebung. eine empirische untersuchung im rechnungswesenunterricht [subject-didactical variations in a self-organized learning environment. an empirical study in accounting education]. wiesbaden: deutscher universitäts-verlag. sembill, d. (1984). modellgeleitete interaktionsanalysen im rahmen einer forschungsorientierten lehrerausbildung – am beispiel von untersuchungen zum „kaufvertrag“ [modell-driven interaction analysis within the scope of research oriented teacher education by the example of the topic “purchase agreement”]. dissertation, georg-august-universität göttingen, germany. sembill, d. (1999). selbstorganisation als modellierungs-, gestaltungsund erforschungsidee beruflichen lernens [self-organisation as idea for modelling, design, and research of vocational learning]. in t. tramm, d. sembill, f. klauser, & e. g. john (eds.), professionalisierung kaufmännischer berufsbildung. festschrift zum 60. geburtstag von frank achtenhagen [professionalisation of commercial vocational education. festschrift to the 60th birthday of frank achtenhagen], (pp. 146–174). frankfurt a. m.: peter lang. sembill, d. (2004). abschlussbericht an die deutsche forschungsgemeinschaft im rahmen des schwerpunktprogramms „lehr-lern-prozesse in der kaufmännischen erstausbildung“ [final report to the german research foundation within the scope of the priority programme “teaching and learning processes in business education”]. http://www.uni-bamberg.de/fileadmin/uni/fakultaeten/sowi_lehrstuehle/wirtschaftspaedagogik/dateien/forschung/forschungsprojekte/prozessanalysen/dfg-abschlussbericht_sole.pdf. accessed 10 july 2013. sembill, d., rausch, a., & kögler, k. (2013). non-cognitive facets of competence – theoretical foundations and implications for measurement. in k. beck & o. zlatkin-troitschanskaia (eds.), from diagnostics to learning success: proceedings in vocational education and training, (pp. 199–212). rotterdam: sense. sembill, d., seifried, j., & dreyer, k. (2008). pdas als erhebungsinstrument in der beruflichen lehrlern-forschung – ein neues wundermittel oder bewährter standard? eine replik auf henning kärner&et &al & & | f l r ! ! 41! pätzold [pdas as a tool for data collection in adult education research – new panacea or approved standard? a reply to henning pätzold]. empirische pädagogik, 22 (1), 64–77. sembill, d., wolf, k. d., wuttke, e., & schumacher, l. (2002). self-organized learning in vocational education – foundation, implementation, and evaluation. in k. beck (ed.), teachinglearning processes in vocational education. foundations of modern training programms, (pp. 267–295). frankfurt a.m.: peter lang. shavelson, r. j., hubner, j. j., & stanton, g. c. (1976). self-concept: validation of construct interpretations. review of educational research, 46 (3), 407–441. doi: 10.1037/00223514.45.1.173 shedletsky, r. & endler, n. s. (1974). anxiety: the state-trait model and the interaction model. journal of personality, 42 (4), 511–527. doi: 10.1111/j.1467-6494.1974.tb00690.x sleight, p. & bernardi, l. (1998). sympathovagal balance. circulation, 98 (23), 2640. doi: 10.1161/01.cir.98.23.2640 snow, r. (1989). aptitude-treatment interaction as a framework for research on individual differences in learning. in p. ackerman, r. j. sternberg, & r. glaser (eds.), learning and individual differences, (pp. 13–59). new york: freeman. spangler g., pekrun, r., kramer, k., & hofmann, h. (2002). students’ emotions, physiological reactions, and coping in academic exams. anxiety, stress, & coping: an international journal, 15 (4), 413–432. doi: 10.1080/1061580021000056555 spielberger, c. d. (1989). state-trait anxiety inventory: bibliography (2nd ed.). palo alto, ca: consulting psychologists press. spss (2005). linear mixed-effects modeling in spss: an introduction to the mixed procedure. armonk, new york: ibm. retrieved from http://www.spss.ch/upload/1126184451_linear%20mixed%20effects-%20modeling%20in%20spss.pdf. stadler, t., evans, ph., hucklebridge, f., & clow, a. (2011). associations between the cortisol awaking response and heart rate variability. psychoneuroendocrinology, 36 (4), 454–462. doi: 10.1016/j.psyneuen.2010.07.020 sweller, j. (1988). cognitive load during problem solving: effects on learning. cognitive science, 12 (2), 257–285. doi: 10.1207/s15516709cog1202_4 sweller, j. (1994). cognitive load theory, learning difficulty, and instructional design. learning and instruction, 4 (4), 295–312. doi: 10.1016/0959-4752(94)90003-5 sweller, j., ayres, p., & kalyuga, s. (2011). cognitive load theory. new york: springer. takahashi, t., ikeda, k., ishikawa, m., kitamura, n., tsukasaki, t., nakama, d., & kameda, t. (2005). anxiety, reactivity, and social stress-induced cortisol elevation in humans. neuroendocrinology letters, 26 (4), 351–354. twisk, j. w. r. (2006). applied multilevel analysis. a practical guide. cambridge: cambridge university press. vygotskii, l. s. (1978). mind in society: the development of higher psychological processes. cambridge: harvard university press. wagner, w., göllner, r., helmke, a., trautwein, u., & lüdtke, o. (2013). construct validity of student perceptions of instructional quality is high, but not perfect: dimensionality and generalizability of domain-independent assessments. learning and instruction, 28, 1–11. doi: 10.1016/j.learninstruc.2013.03.003 wild, k. p. (2001). die optimierung von videoanalysen durch zeitsynchrone befragungsdaten aus dem experience sampling. in s. v. aufschnaiter & m. wenzel (eds.), nutzung von videodaten zur untersuchung von lehr-lern-prozessen. aktuelle methoden empirischer pädagogischer forschung [using video data for investigating teaching and learning processes. current methods of empirical pedagogical research], (pp. 61–74). münster: waxmann. wild, k. p. & krapp, a. (1996). die qualität subjektiven erlebens in schulischen und betrieblichen lernumwelten: untersuchungen mit der erlebens-stichproben-methode [the quality of emotional experiences in vocational and company learning environments: studies with the experience-sampling-method]. unterrichtswissenschaft, 24 (3), 195–216. kärner&et &al & & | f l r ! ! 42! wüst, s., federenko, i., hellhammer, d. h., & kirschbaum, c. (2000). genetic factors, perceived chronic stress, and the free cortisol response to awakening. psychoneuroendocrinology, 25 (7), 707–720. doi: 10.1016/s0306-4530(00)00021-4 wuttke, e. (2000). lernstrategien im lernprozeß. analysemethode, strategieeinsatz und auswirkungen auf den lernerfolg [learning strategies in learning processes: method of analysis, use of strategies, and effects on learning]. zeitschrift für erziehungswissenschaft, 3 (1), 97–110. doi: 10.1007/s11618-000-0007-6. yeh, yu-chu (2012). aptitude-treatment interaction. in n. m. seel (ed.), encyclopedia of the sciences of learning, (pp. 295–298). new york: springer. yerkes, r. m. & dodson, j. d. (1908). the relation of strength of stimulus to rapidity of habitformation. journal of comparative neurology and psychology, 18 (5), 459–482. doi: 10.1002/cne.920180503. zierer, k. & seel, n. m. (2012). general didactics and instructional design: eyes like twins. a transatlantic dialogue about similarities and differences, about the past and the future of two sciences of learning and teaching. springerplus, 17, 1–15. doi: 10.1186/2193-1801-1-155. microsoft word catrysse et al_publication.docx frontline learning research vol.4 no. 1 (2016) 1-16 issn 2295-3159 mapping processing strategies in learning from expository text: an exploratory eye tracking study followed by a cued recall catrysse leena1, gijbels davida, donche vincenta, de maeyer svena, van den bossche pieta, gommers lucib a university of antwerp, belgium b university of st. gallen, switzerland article received 22 july / revised 12 december / accepted 12 december / available online 27 january abstract this study starts from the observation that current empirical research on students’ processing strategies in higher education has mainly focused on the use of self-report instruments to measure students’ general preferences towards processing strategies. in contrast, there is a rather limited use of more direct and online observation techniques to uncover differences in processing strategies at a task specific level. we based our study on one of the most influential studies in the domain of students’ approaches to learning (sal) (marton, dahlgren, säljö, & svensson, 1975). in our exploratory experiment we used eye tracking followed by a cued recall to investigate how students use processing strategies in learning from expository text. nineteen university students participated in the experiment. results suggested that students in the deep condition did not look longer at the essentials in the text compared with students in the surface condition, but that they processed them in a more deep way. in our sample, students in the surface condition looked longer at facts and details and also reported repeating these facts and details more often. we suggest that the combination of eye tracking followed by a cued recall is a promising tool to investigate students’ processing strategies since not all differences in processing strategies are reflected in overt eye movement behaviour. the current methodology allows researchers in the domain of sal to complement and extend the present knowledge base that has accumulated through years of research with self-report questionnaires and interviews on students’ general preferences towards processing strategies. keywords: processing strategies; expository text; eye tracking; cued recall; higher education 1 corresponding author: catrysse leen, faculty of social sciences, department of training and education sciences, research group edubron. gratiekapelstraat 10, 2000 antwerpen, belgium. e-mail: leen.catrysse@uantwerpen.be doi: http://dx.doi.org/10.14786/flr.v4i1.192 catrysse  et  al       2 | f l r 1. introduction learning from text is one of the most essential skills in our modern society and the ability to understand challenging texts is an important key to success in education and beyond (mason, tornatora, & pluchino, 2013; mcnamara, 2004; moss, schunn, schneider, mcnamara, & vanlehn, 2011). one of the research traditions that is interested in how students learn from text is the domain of student approaches to learning (sal) (gijbels, donche, richardson, & vermunt, 2014; lonka, olkinuora, & mäkinen, 2004; richardson, 2000). research in the sal domain is founded on the seminal studies by marton and his colleagues in the 1970s in sweden (marton et al., 1975). they investigated how students went about reading academic texts in experimental situations by conducting retrospective interviews (marton et al., 1975; richardson, 2000). a distinction was made between deep processing strategies and surface processing strategies, which has been influential in the later development of self-report questionnaires to quantify individual differences in students’ processing strategies (biggs, 1987; entwistle & mccune, 2004). up till now, empirical studies in the sal field have mainly been focused on the use of self-report instruments such as interviews and questionnaires to uncover differences in students’ general preferences towards processing strategies. although these offline measures are claimed to be reliable and valid at this general level, many authors argue that the results are poor indicators of the actual processing at a task specific level (perry & winne, 2006; samuelstuen & braten, 2007; veenman, 2005; veenman, bavelaar, de wolf, & van haaren, 2014). recently, there has been a plea for the use of more direct and online measurement tools when it comes to describe students’ processing strategies (richardson, 2013). in the present study we will therefore use eye tracking to map individual differences in cognitive processing followed by a cued recall. eye tracking provides a unique opportunity to study processing strategies in a level of detail that no other measures can provide (lai et al., 2013; van gog & jarodzka, 2013). in what follows we will describe how different processing strategies can be manipulated in experimental designs by the assessment demands, and how eye tracking followed by a cued recall can be useful to investigate differences in processing strategies. 2. uncovering differences in processing strategies processing strategies refer to cognitive activities a student applies whilst studying (vermunt & vermetten, 2004). in general, two main types of processing strategies are described in the literature namely deep and surface processing strategies (gijbels et al., 2014). research in the sal domain showed that deep processors try to comprehend what the author wants to say about a certain topic, try to understand the overall meaning of the text, try to relate the message to a wider context and to prior knowledge, identify the main ideas and adopt a critical angle to the conclusion. in contrast, surface processors direct their attention towards learning the text itself, focus more on specific comparisons, focus on the parts of the text in sequence, memorize details and definitions, remember introductory sentences and list points (biggs & tang, 2007; entwistle & ramsden, 1982; marton et al., 1975; richardson, 2000). 2.1. processing strategies and task demands in the 1960s, rothkopf (1966) introduced the concept of mathemagenic activities, which refers to activities that stimulate students to actively engage in learning. the use of adjunct questions in written texts is one example of these mathemagenic activities. one possible type of an adjunct question is the inserted post question, which is placed within the text and follows the text passage containing the information needed. these questions result in a change in the processing strategy on subsequent text passages. they steer students attention to a specific type of information in the text (hamaker, 1986; rothkopf, 1966). catrysse  et  al       3 | f l r similarly, researchers in the sal domain agree that one of the most salient contextual variables to influence processing strategies is the assessment method (baeten, kyndt, struyven, & dochy, 2010; gielen, dochy, & dierick, 2003; marton et al., 1975; scouller, 1998; scouller & prosser, 1994; segers, nijhuis, & gijselaers, 2006). research showed that how students learn is influenced by their initial preference for a processing strategy (baeten et al., 2010), but they can shift between deep and surface processing strategies according to the assessment demands, also known as the backwash-effect of assessment (baeten et al., 2010; gielen et al., 2003; segers et al., 2006). in contrast to adjunct question research (hamaker, 1986; rothkopf, 1966), research in the sal domain evaluated the effect of the assessment method at the end of a text or study process, without inserting questions in the text or interrupting the study process. in the experiments of marton et al. (1975), students were asked to read three texts and to prepare for answering some questions on the content after reading them. the questions they received after the first two texts were the only indication on how to behave during reading the third text. students in the deep condition received questions at a deep level (e.g., making a summary statement), while students in the surface condition received reproduction-oriented questions. after studying the third text, a semi-structured interview was conducted to gather data on the effect of the experimental manipulation on the levels of processing. the results of the interviews suggested that students tended to adapt the intended level of processing (marton et al., 1975; richardson, 2000). this study was the first study in the sal domain to confirm the possibility to manipulate students’ levels of processing by appropriate questions or prompts. it shows that the level of processing depends on the expected form of assessment (richardson, 2000). another study of scouller and prosser (1994) suggested that the assessment method influences processing strategies. their research showed that multiple-choice questions led to more surface processing strategies. also research of scouller (1998) investigated how students perceived two assessment methods namely multiple-choice examination and an assignment essay and which processing strategies they used. the findings were in line with scouller and prosser (1994), multiple-choice examination was perceived as assessing lower levels of intellectual abilities and students indicated to engage in more surface processing strategies. an assignment essay was perceived as testing higher-level intellectual abilities and students engaged in more deep processing strategies. a last study of segers et al. (2006) showed that students who perceive the demands on a deep level, to demonstrate a thorough understanding and integration of knowledge, are more likely to employ deep processing strategies. in contrast, students who perceive the demands of assessment on a surface level, to acquire passive acquisition and reproduction of details, are expected to employ more surface processing strategies such as rote learning and concentrating on facts and details. 2.2. processing strategies and eye tracking online measures to map cognitive processing strategies include the think aloud method, observation of behaviour and eye movement measurement (schellings, 2011; veenman, 2011). the think aloud method provides a rich source of data, but it is intrusive and can alter the processing itself (ericsson & simon, 1993; veenman, 2005). the main limitation of the observation of behaviour is that it cannot detect covert cognitive processes (veenman, 2005). according to hyönä and lorch (2004) eye tracking is an attractive method for studying cognitive processing strategies in comparison with other online measures because eye tracking collects several indices of processing simultaneously and does not disrupt normal processing. there are two theoretical assumptions that make the relation between eye movement and cognitive processing clear: the immediacy assumption and the eye-mind assumption (just & carpenter, 1980). the immediacy assumption states that information processing is not postponed and takes place when the information is encountered. the eye-mind hypothesis explains that eye movements are closely linked to the focus of attention as students process the information in the text. therefore, eye movements can be used to trace cognitive processing when learning from text (hyönä, lorch, & rinck, 2003; just & carpenter, 1980). in eye tracking research, the movement of the eyeball is recorded and these movements are related to a stimulus. this allows us to investigate to what parts of the text a student allocates visual attention and for how long (holmqvist et al., 2011; van gog & jarodzka, 2013). a distinction is made between two main catrysse  et  al       4 | f l r measures namely fixations and saccades. during fixations the eye is almost completely still and information can be extracted from the text. in contrast, during saccades the focus of visual attention is moved to another location and the eye is rapidly moving between fixations, as a result students are not able to extract information from text during saccades (holmqvist et al., 2011; lai et al., 2013; van gog & jarodzka, 2013). although eye tracking methodology seems a promising tool to investigate students’ processing strategies, we could not find studies that examine eye movement behaviour that results from using different cognitive processing strategies such as deep and surface processing. in another related research field, namely research in reading comprehension, they already adopted the eye tracking methodology (hyönä, lorch, & kaakinen, 2002; ponce & mayer, 2014; rayner, 1998). more specifically, the perspective driven text comprehension framework states that the allocation of visual attention is influenced by the reading perspective and this reading perspective shapes the cognitive processing in learning from text (kaakinen & hyönä, 2005, 2007, 2010; kaakinen, hyönä, & keenan, 2002). a reading perspective refers to the mental frame from which the reader approaches a text and this perspective makes parts of the text more important to the reader than others (hyönä et al., 2003; kaakinen & hyönä, 2007). kaakinen and hyönä (2007) gave the example that when you read a travel guide in order to find information about a specific country (e.g., finland), you will approach the text with a specific reading perspective. this reading perspective is thus content related. alternatively, processing strategies correspond to the different aspects of the learning material on which the learner focuses (richardson, 2000). so students with different processing strategies focus on the same content but search for other types of information (e.g., facts and details vs. essences) (schellings, van hout-wolters, & vermunt, 1996). research that investigates the influence of reading perspective on eye movements showed that there is more time spent on relevant words or facts in the text than on irrelevant words or facts (kaakinen & hyönä, 2007; kaakinen et al., 2002). next to that, relevant words attracted more refixations than irrelevant words (kaakinen & hyönä, 2007). research of kaakinen and hyönä (2005) indicated that the extra time spent on relevant information is used to rehearse this information in order to encode it to memory. particularly relevant for research on learning from text is that these refixations reflect purposeful and effortful strategic eye behaviour (ariasi & mason, 2010). eye tracking is an interesting method to investigate cognitive processes, but to reduce the amount of inferences required by the researcher, eye movement data should be combined with other data such as verbal reports (hyönä, 2010; van gog & jarodzka, 2013). recent studies have already applied the think aloud method to obtain verbal reports on students’ processing strategies during reading and learning from text (dinsmore & alexander, 2012, 2015). concurrent reporting while learning from text can affect the eye movement patterns, and therefore cued retrospective reporting offers a valuable alternative in combination with eye tracking (van gog & jarodzka, 2013). besides recording the eye movement, the eye tracking software allows replaying the records of eye movements. using this eye movement pattern as a memory cue, it may help learners to recover how they encoded and interpreted elements in the text (hyönä, 2010; penttinen, anto, & mikkilä-erdmann, 2012; van gog, paas, & van merrienboer, 2005). because of the small delay after processing the text and the presentation of the memory cue, students are still able to report on their cognitive processes (veenman, 2005, 2011). for this reason we chose to use cued retrospective reporting to triangulate with eye movement measures. 3. present study our study aims to extend current research on processing strategies by using eye tracking methodology followed by a cued recall to map differences in processing strategies. this more direct and online way of measuring processing strategies allows to learn more about the actual processing behaviour of students while learning from expository text. catrysse  et  al       5 | f l r as stated above processing strategies shape what information is looked for in a text and what information is perceived as relevant (kaakinen & hyönä, 2005). next to that, research using self-report measures suggested that deep processors focus more on essences and surface processors focus more on details and definitions (biggs & tang, 2007; entwistle & ramsden, 1982; marton et al., 1975; richardson, 2000). based on findings from research on perspective driven text comprehension (kaakinen & hyönä, 2008) and the sal domain (lonka et al., 2004), we suggest the following hypotheses for students in the deep condition (after receiving guiding questions at a deep level) and students in the surface condition (after receiving reproduction-oriented questions): a) hypothesis 1: students in the deep condition focus their attention longer on the essentials (e.g., key phrases and words) in the text compared to students in the surface condition. b) hypothesis 2: students in the deep condition, more often return back to essences compared to students in the surface condition. c) hypothesis 3: students in the surface condition focus their attention longer on facts and details (e.g., names) compared to students in the deep condition. d) hypothesis 4: students in the surface condition, more often return back to facts and details compared to students in the deep condition. 4. method 4.1. participants twenty-eight students (age range: 18-25) enrolled at the university of antwerp (belgium), participated on a voluntary basis. participants were randomly divided in either the deep condition (dc, n = 14) or the surface condition (sc, n = 14). unfortunately, data of nine respondents could not be used due equipment failure and problems with eye tracking calibration. therefore, data of 19 students were considered in the statistical analyses (table 2). all participants had normal or corrected-to-normal vision and dutch was their native language. table 1 participant characteristics dc sc n 12 7 gender male 5 5 female 7 2 4.2. materials in order to test our hypotheses, we based our experimental design on the seminal studies by marton et al. (1975). in their experiments they induced either a deep or surface processing strategy by giving students questions after they studied an academic text. in our experiment, students were asked to study a series of three expository texts (± 800 words) on a topic they were not familiar with, namely research on happiness. the texts were taken from the dutch version of ‘the world book of happiness’ (bormans, 2010). catrysse  et  al       6 | f l r after processing each text they received a number of evaluation questions on the preceding text (figure 1). students in the deep condition received questions at a deep level (e.g., give a summary of the text). in contrast, students in the surface condition received reproduction-oriented questions (e.g., in which country was the research discussed in the text conducted?). so in both conditions students processed the same learning content, but received different questions. in the original study, marton et al. (1975) interviewed and tested the students after the third text and concluded that in the surface condition, students adopted more surface processing strategies while students in the deep condition adopted more deep processing strategies. similarly, in our study we analysed the eye tracking data and cued recalls from the third text. figure 1. experimental design. 4.3. eye tracking eye movements were collected using the tobii tx300 eye tracker (dark pupil tracking), manufactured by tobii technology (stockholm, sweden). it is integrated into a 23-inch tft monitor with a maximum resolution of 1920 x 1080 pixels. the camera samples data at the rate of 300 hz and registration was binocular. tobii tx300 does not require a head stabilization system and allows for more freedom of head movement (37 x 17 cm). gaze accuracy is 0.4° and gaze precision is 0.15°, as reported by the hardware producer. the eye tracker latency is between 1.0 and 3.3 milliseconds. data were recorded with tobii-studio (3.2) software. before starting the experiment, students were seated about 60 cm from the screen for the eye tracking calibration. a five point calibration procedure was used in which students needed to track five red calibration dots on a plain, grey background. areas of interest (aoi’s) define regions in the text that the researcher is interested in gathering data about (holmqvist et al., 2011). with regard to our hypotheses we are interested in key phrases and keywords for the deep condition and in details and facts for the surface condition. six volunteers (master students in educational sciences) read the text in a pilot study to determine the key phrases and keywords. in total 15 deep aoi’s (e.g., a topic sentence with summary statements) and three surface aoi’s (e.g., name of a country) were marked. there were only parts of the text defined as aoi’s, so not the whole text was covered with aoi’s. the total size of the text was 1490 x 1087 pixels, the smallest aoi was 47 x 31 pixels and the biggest aoi was 684 x 71 pixels. the complete text could be seen on the screen, so scrolling was not needed. in line with hyönä et al. (2002) first pass fixation time, look back fixation time and total fixation time were analysed at the level of aoi’s. an overview of the definitions is given in table 2 (holmqvist et al., 2011; hyönä et al., 2003). students were able to process the text in a self-paced manner and therefore we calculated relative duration measures. next to that, aoi’s differed in size because they sometimes contained phrases, while others consisted of only words. therefore, aoi measures were normalized by calculating the reading depth measure (holmqvist et al., 2011; holmqvist & wartenberg, 2005; holsanova, holmqvist, & rahm, 2006). this reading depth measure is defined by the total time spent in an aoi per cm2 and is an indication of how densely an aoi is processed. so for the three measures described in table 2 we calculated relative measures and reading depth measures. catrysse  et  al       7 | f l r table 2 overview of eye tracking measures and their definitions measure definition first pass fixation time the time spent in an aoi when it was visited for the first time. a visit can consist of more fixations. it reflects early processing and object recognition. look back fixation time duration of all the regressions back to an aoi. it reflects delayed processing, for example to integrate information. total fixation time the time spent in an aoi during the whole trial, it is the sum of the first pass fixation time and the look back fixation time in that aoi. the fixation indices were calculated for either the group of deep aoi’s or the group of surface aoi’s. we used the tobii fixation filter for fixation identification, which is an implementation of a classification algorithm proposed by olsson (2007). it uses a velocity threshold (35 pixels/window) and a distance threshold (35 pixels). for all the measures, the means and standard deviations were calculated. to compare students in both conditions, we used non-parametric tests due to the small sample sizes (van gog et al., 2005). therefore the medians together with the first and third quartile were calculated as well. relative measures and reading depth measures for the eye movement measures were compared for students in both conditions using mann-whitney u tests. we reported the exact two-tailed significance. also in line with van gog et al. (2005), we used a less stringent significance level of 0.10 to avoid type ii error and to increase power. 4.4. cued recall after the eye tracking experiment, a cued recall was conducted. after processing the third text, the experimenter informed students that they would watch the replay of eye movements of the third text together. the cued recall was conducted by using gaze videos produced by tobii-studio software (3.2). in the cued recall, a video showed the text and a moving red dot representing the point of fixation. the bigger the dot, the longer the fixation lasted. students saw their gaze videos at the same speed they processed the text. the interviewer instructed students to watch the video and to tell the interviewer what they were thinking during processing the text. the interviewer also stated that she would occasionally stop the video and ask questions about the reading process, such as ‘here you fixated a lot, what where you doing?’ or ‘here you are going back in the text, what were you doing?’. catrysse  et  al       8 | f l r table 3 coding scheme for the cued recall analysis strategy example dc sc surface processing 66 (65,3%) 45 (100%) rereading i tried to understand that part so i was rereading it. skimming now i am reading it again and just scanning for important words in the text. guessing meaning word in context that was gnp, i was wondering what the meaning of that word was. rehearsing those countries, i was trying to remember them. connecting to prior text i realise that i go back a lot in the text and that is because i am trying to link parts of the text. connecting to the research task i guess the first paragraph was going to give an overview about the rest of the task, so i thought that was important. detecting mistakes in the text i was looking at the ‘n’ that was missing in that word. deep processing 35 (34,7%) 0 (0%) questioning i was wondering what they meant with that phrase. paraphrasing first, they name something and then you know a summation is coming. second, they talk about cross national comparisons, … connecting to personal experiences you try to process the text critically and to take you own findings and personal experience into account. interpreting and elaborating what i do most of the time is reading the text and then trying to analyse what i just read. in this way i get a better picture of what the text is about. the cued recalls were transcribed from the audiotapes. next to that, we linked comments of the cued recalls to the part of the text that was discussed. the cued recalls were coded based on an initial set of ten codes developed in a study of dinsmore and alexander (2015). specifically, comments were coded as either a surface or deep processing strategy (table 3). after coding the interviews deductively, we added one extra code in the surface processing category namely “detecting mistakes in the text”. transcripts were coded with the qualitative analysis software package nvivo 10. two judges (authors lc and lg) coded the cued recalls and an inter-rater agreement of 73% was reached, which is considered as substantial. we compared the number of coded utterances in each condition between the two categories (table 3). we first analysed the data on a general level and looked for differences between students in both conditions. we also analysed the data at a more fine-grained level to see whether the reported strategies are linked to aoi’s and to examine differences at the aoi level between groups. 5. results table 4 shows the means and standard deviations. standard deviations for the measures in the deep condition are higher than in the surface condition. this may be an indication that students in the deep condition differ more from each other. when we look at the cued recall results of students in the deep catrysse  et  al       9 | f l r condition, some students pointed out that they sometimes took a pause to integrate processed information instead of looking back. this may also be an indication that there are two types of students in the deep condition, on the one hand students who process information immediately and take a pause to integrate information and on the other hand students who need to look back to parts in the text to integrate this information and to encode it to memory. “sometimes i keep staring at the text, because i try to visualize it for myself” (r7, dc) “i sometimes have the feeling that when i am staring at a word that i am not processing that word but that i am just taking a moment to think about what i have read” (r3, dc) “sometimes i have the feeling that i am staring at something in the text, to process the things i just read before” (r5, dc) the most reported processing strategy in the cued recalls, is the surface processing strategy and more specifically rereading. students in both groups indicated that they reread parts of the text the most. only students in the deep condition reported deep and surface processing strategies. students in the surface condition only reported surface processing strategies. deep processing strategies are reported on a more general level and are not linked to certain phrases, paragraphs or aoi’s in the text. “when you know you will need to answer questions after reading the text, you try to read the text critically and i always try to take into account my personal experiences and findings.” (r5, dc) “i first think about what i read in the text, before i proceed with the next part. i try to make a summary for myself of what i read in the previous parts.” (r3, dc) table 4 means and standard deviations essentials facts and details dc sc dc sc m sd m sd m sd m sd fpft r 2.69 2.22 3.23 1.48 0.40 0.21 0.38 0.25 fpft rd 45.53 23.77 57.50 25.19 80.60 24.49 79.02 51.30 lbft r 12.65 3.63 11.22 2.69 1.09 0.49 1.94 0.77 lbft rd 285.66 192.02 200.75 46.39 268.20 236.77 401.22 168.88 tft r 15.35 3.01 14.46 2.85 1.50 0.57 2.32 0.84 tft rd 331.19 189.72 285.25 43.45 348.80 234.62 480.24 185.68 fpft = first pass fixation time; lbft = look back fixation time; tft = total fixation time; r = relative measure; rd = reading depth measure. we compared the total reading time of students in both conditions with a mann-whitney u test, but no significant differences were found (u = 41, p = 0.97). so students in both conditions spent on average the same amount of time on processing the text. catrysse  et  al       10 | f l r 5.1. essentials in the text table 5 shows the medians and quartiles for the essentials in the text for students in both conditions. we conducted mann-whiney u tests on all these measures but no significant differences were found between students in both groups. table 5 first quartile, median and third quartile for relative measures and reading depth measures. dc sc mann-whiney u q1 mdn q3 q1 mdn q3 u p fpft r 1.53 1.92 2.86 2.13 2.57 4.43 27 0.227 fpft rd 27.65 41.14 45.17 41.18 48.55 75.67 30 0.340 lbft r 10.93 12.92 15.88 9.39 10.60 12.69 56 0.261 lbft rd 173.39 215.14 354.49 170.57 178.85 215.59 53 0.385 tft r 13.97 15.63 18.26 12.42 13.02 16.41 53 0.385 tft rd 209.29 261.89 396.90 235.15 257.08 267.00 45 0.837 results from the cued recalls indicate that both students in the deep and surface condition reread essentials in the text. a reason for rereading is that they did not really understand essential parts of the text. the motivation to better understand these essential parts in the text is only reported by students in the deep condition. these results suggest that students in the deep condition reread these parts at a deeper level to get a better understanding. “i am rereading a lot, i read something fast and then i think whether i understood it and no i did not, so then i go back again” (r3, dc) “i was trying to understand that part better, so that is why i was rereading it over and over again” (r2, dc) both groups indicated skimming the text after reading it for the first time to look back at the essential parts of the text. “what i often do when i finished reading, is rereading only the essential parts of the text” (r4, dc) “i am just scanning quickly to see if i missed important words in the text” (r16, sc) a final finding from the cued recall results is that both groups guessed the meaning of keywords in context, when they did not understand the word. overall, cued recall results are in line with results from eye tracking, in that no big differences are found between both groups when processing essential parts in the text. “here, that was a difficult word, elitist, i tried to understand the meaning in the text” (r8, dc) “some keywords i do not know, i need to think about them or see the context to understand them” (r19, sc) catrysse  et  al       11 | f l r 5.2. facts and details in the text table 6 shows the medians and quartiles for facts and details in the text for students in both conditions. students in the surface condition spent relatively more time on facts and details when they looked back at them and also during the whole experiment. next to that, these students read the facts and details with more depth than students in deep condition when they look back at them and during the whole experiment. table 6 first quartile, median and third quartile for relative measures and reading depth measures. dc sc mann-whiney u q1 mdn q3 q1 mdn q3 u p fpft r 0.23 0.38 0.54 0.19 0.39 0.54 46 0.773 fpft rd 66.34 78.06 102.04 38.77 82.39 112.36 44 0.902 lbft r 0.74 1.04 1.41 1.74 2.28 2.41 17 0.036 lbft rd 160.56 178.28 270.98 329.09 423.92 529.68 20 0.068 tft r 1.01 1.47 1.95 2.23 2.62 2.86 15 0.022 tft rd 241.65 276.45 351.52 440.90 479.24 611.01 18 0.045 cued recall results showed that students in the surface condition repeated facts and details in the text, while students in the deep condition did not. other coding categories did not show a link with processing facts and details in the text. again we can see a clear link between the eye movement measures and the results from the cued recalls. “the names of those countries, i really tried to remember those” (r14, sc) “those four countries, i memorized them” (r19, sc) “i tried to remember the name of the author, i thought that would be important” (r17, sc) 6. conclusion and discussion this exploratory study aimed at extending current research on processing strategies during learning from expository text. research in the sal domain is mostly based on students’ self-reports of processing strategies at a general level in which the context of learning is not taken into account (dinsmore & alexander, 2012; gijbels et al., 2014). by looking at the actual processing behaviour of students while learning from expository text, this study makes a first preliminary contribution to the field by using a more direct and online measurement tool at a task specific level that takes the context explicitly into account. it is the first experimental study to explore students’ cognitive processing strategies at a task specific level using objective online measures. most of the research using online measures is based on the think aloud method, which can alter the processing itself (veenman, 2005). by using eye tracking methodology followed by a cued recall this problem is circumvented, in that this method does not demand students to manage cognitive load of the task completion and self-reports of strategies at the same time (samuelstuen & braten, 2007). catrysse  et  al       12 | f l r in our study we manipulated the task demands to steer processing strategies. results from the cued recalls indicated that this manipulation was successful as students in the deep condition reported a combination of surface and deep processing strategies, while students in the surface condition only reported surface processing strategies. this is in line with previous research that showed that demands on a deep level, to demonstrate a thorough understanding, lead to more deep processing strategies whereas demands on a surface level, to acquire passive acquisition of facts and details, lead to more surface processing strategies (marton et al., 1975; richardson, 2000; scouller, 1998; scouller & prosser, 1994; segers et al., 2006). results of the cued recalls indicated that students in both conditions processed facts and details and essential parts in the text but they did it in a different way. these results are similar to results from think aloud studies in which processing strategies were examined without manipulating task demands (dinsmore & alexander, 2012, 2015; penttinen et al., 2012). based on the eye movement data, we cannot confirm the first and second hypothesis that stated that students in the deep condition focus their attention longer on essentials in the text compared to students in the surface condition and that they return more back to them. both groups of students spent time on processing the essentials in the text. although we could not find differences between groups based on their eye movement data, results from the cued recalls indicated that students in the deep condition reread the essentials in the text to understand them better. this motivation to better understand these parts is related to a deep way of processing (biggs & tang, 2007; entwistle & ramsden, 1982). students in the surface condition did not report this motivation. these descriptive findings indicate that students in our sample processed the text in a different way but more substantive research is needed to further explore found differences in overt eye movement behaviour. in contrast with research from the angle of perspective driven text comprehension, these essential parts do not seem to be perceived as more relevant by students in the deep condition (kaakinen & hyönä, 2005, 2007, 2008), they are just processed in a more deep way. another interesting finding from the cued recall results is that some students in the deep condition indicated that they took a pause at some places in the text to integrate the processed information instead of actively looking back. other students in the deep condition reported actively looking back at these essential parts in the text. also the higher standard deviations for students in the deep condition may be an indication of these differences. it is in line with other research that shows that building the necessary links to incorporate text information to the developing memory representation can be achieved mentally or can result in overt behaviour in which students actively reread essential parts (hyönä et al., 2003; kaakinen & hyönä, 2008). so, based on these preliminary findings, we suggest that some deep processors actively return back to essentials to encode it to memory, while others take a pause to integrate the new information without looking back to this information. further research is needed to confirm these findings. regarding the third and fourth hypothesis, the results indicated that students in the surface condition indeed looked longer at facts and details and returned more back to them. it seems that students in the surface condition switch to strategic processing by paying more attention to relevant parts, namely facts and details (kaakinen & hyönä, 2007). research of kaakinen and hyönä (2005) showed that the extra time spent on relevant information is used to rehearse this information in order to encode it to memory. results from both eye tracking and cued recalls indicate that facts and details are more repeated in order to encode into memory in the surface condition (kaakinen & hyönä, 2007). only students in the surface condition reported repeating facts and details, while students in the deep condition did not report learning activities like that. although our findings suggest that eye tracking followed by a cued recall is a fruitful way to investigate processing strategies, we want to stress the preliminary nature of this study because of some limitations. an important limitation of this study is the small sample size. due to equipment failure or problems with eye tracking calibration, the sample size decreased at the onset of this study. because of this smaller sample size we decided to use non-parametric tests and deepened the results obtained by a cued recall. we also raised the significance level to increase power due to the smaller sample size (van gog et al., 2005). the findings from this study can serve as a baseline for further research in which larger samples can be used to increase power without adjusting the significance level. another limitation of this study is that we catrysse  et  al       13 | f l r used a between groups design. reading times and online processing strategies can vary among adult readers (hyönä et al., 2002; kaakinen & hyönä, 2008). therefore we suggest for further research to use a within groups design in which students use both processing strategies to take this individual variability into account. another way to understand the significance of individual variability is to include control variables such as reading ability, interest in the topic and prior knowledge about the topic (fox, 2009; mason et al., 2013). by increasing the sample size and using a within subjects design, more complex statistical analysis can be conducted to confirm our preliminary findings. in this way it will be possible to give more generalized statements regarding processing strategies as measured by eye tracking. a last limitation is that students needed to process the text on a computer screen to be able to use the eye tracking. by doing this it does not reflect the natural setting in which students habitually process learning contents. despite the limitations, this study was able to show that eye tracking followed by a cued recall is a promising tool to examine students’ processing strategies. an important finding from our study is that it is valuable to combine eye tracking with a cued recall, because differences in processing strategies not always lead to overt eye movement behaviour (hyönä et al., 2003; kaakinen & hyönä, 2008). by using a cued recall we were able to uncover differences in processing strategies that were not reflected in eye movement behaviour. based on our preliminary findings, the combination of eye tracking and a cued recall seems to be a promising tool to further investigate cognitive processing strategies when learning from text. students in the deep condition do not seem to look longer at essentials and do not seem to return more back to them, but processed them in a more deep way then students in the surface condition. results suggest that students in the surface condition looked longer at facts and details and did return more back to them. this first exploratory eye tracking study in the sal domain is an important illustration on how processing strategies can be further examined beyond the use of self-report questionnaires. in our opinion it would be worthwhile to use this innovative eye tracking methodology in multi-method designs to triangulate it with often used self-report measures to look for convergent or divergent validity. in our study we steered students’ processing strategies by task demands. although research indicated that it is possible to influence processing strategies by manipulating this contextual variable (baeten et al., 2010; gielen et al., 2003; marton et al., 1975; scouller, 1998; scouller & prosser, 1994; segers et al., 2006), it would be interesting to combine it with these self-report measures in order to examine a more natural way of processing behaviour. next to that, using multiple sources of data is important to develop a comprehensive understanding of how we can adequately measure students’ processing strategies. eye tracking methodology followed by a cued recall in the sal domain can also deepen the conceptual underpinnings on what constitutes deep and surface processing of learning contents. keypoints eye tracking followed by a cued recall is a promising tool to uncover differences in students’ processing strategies while learning from expository text. students in the deep condition do not look longer at the essentials, but they process them in a more deep way by trying to understand these parts better. students in the surface condition look longer at facts and details and try to rehearse these parts. references ariasi, n., & mason, l. (2010). uncovering the effect of text structure in learning from a science text: an eye-tracking study. instructional science, 39(5), 581-601. doi: 10.1007/s11251-010-9142-5 baeten, m., kyndt, e., struyven, k., & dochy, f. (2010). using student-centred learning environments to stimulate deep approaches to learning: factors encouraging or discouraging their effectiveness. educational research review, 5(3), 243-260. doi: 10.1016/j.edurev.2010.06.001 catrysse  et  al       14 | f l r biggs, j. (1987). student approaches to learning and studyin. research monograph. melbourne: australian council for educational research. biggs, j., & tang, c. (2007). teaching for quality learning at university: open university press / mcgraw-hill education. bormans, l. (2010). geluk. the world book of happiness. tielt: lannoo. dinsmore, d. l., & alexander, p. a. (2012). a critical discussion of deep and surface processing: what it means, how it is measured, the role of context, and model specification. educational psychology review, 24(4), 499-567. doi: 10.1007/s10648-012-9198-7 dinsmore, d. l., & alexander, p. a. (2015). a multidimensional investigation of deep-level and surfacelevel processing. the journal of experimental education, 1-32. doi: 10.1080/00220973.2014.979126 entwistle, n., & mccune, v. (2004). the conceptual bases of study strategy inventories. educational psychology review, 16(4), 325-345. doi: 10.1007/s1064800400030 entwistle, n., & ramsden, p. (1982). understanding student learning. new york: nichols publishing company. ericsson, k. a., & simon, h. a. (1993). protocol analysis. verbal reports as data. massachusetts: massachusetts institute of technology. fox, e. (2009). the role of reader characteristics in processing and learning from informational text. review of educational research, 79(1), 197-261. doi: 10.3102/0034654308324654 gielen, s., dochy, f., & dierick, s. (2003). evaluating the consequential validity of new modes of assessment: the influence of assessment on learning, including pre-, postand true assessment effects. . in m. segers, f. dochy & e. cascallar (eds.), optimising new modes of assessment: in search of qualities and standards (pp. 37-54). the netherlands: kluwer academic publishers. gijbels, d., donche, v., richardson, j. t. e., & vermunt, j. d. (eds.). (2014). learning patterns in higher education. dimensions and research perspectives. . london: routledge. hamaker, c. (1986). the effects of adjunct questions on prose learning. review of educational research, 56(2), 212-242. doi: 10.3102/00346543056002212 holmqvist, k., nyström, m., andersson, r., dewhurst, r., jarodzka, h., & van de weijer, j. (2011). eye tracking : a comprehensive guide to methods and measures. oxford ; new york: oxford university press. holmqvist, k., & wartenberg, c. (2005). the role of local design factors for newspaper reading behaviour an eye-tracking perspective. lund university cognitive studies (vol. 127). lund university. holsanova, j., holmqvist, k., & rahm, h. (2006). entry points and reading paths on newspaper spreads: comparing a semiotic analysis with eye-tracking measurements. visual communication, 5(1), 65-93. doi: 10.1177/1470357206061005 hyönä, j. (2010). the use of eye movements in the study of multimedia learning. learning and instruction, 20(2), 172-176. doi: 10.1016/j.learninstruc.2009.02.013 hyönä, j., & lorch, r. f. (2004). effects of topic headings on text processing: evidence from adult readers’ eye fixation patterns. learning and instruction, 14(2), 131-152. doi: 10.1016/j.learninstruc.2004.01.001 hyönä, j., lorch, r. f., & kaakinen, j. k. (2002). individual differences in reading to summarize expository text: evidence from eye fixation patterns. journal of educational psychology, 94(1), 44-55. doi: 10.1037//0022-0663.94.1.44 hyönä, j., lorch, r. f., & rinck, m. (2003). eye movement measures to study global text processing. in j. hyönä, r. radach & h. deubel (eds.), the mind's eye: cognitive and applied aspects of eye movement research. amsterdam: elsevier science. just, m. a., & carpenter, p. a. (1980). a theory of reading: from eye fixations to comprehension. pyschological review, 87(4), 329-354. doi: 10.1037/0033-295x.87.4.329 kaakinen, j. k., & hyönä, j. (2005). perspective effects on expository text comprehension: evidence from think-aloud protocols, eyetracking, and recall. discourse processes, 40(3), 239-257. doi: 10.1207/s15326950dp4003_4 catrysse  et  al       15 | f l r kaakinen, j. k., & hyönä, j. (2007). perspective effects in repeated reading: an eye movement study. memory & cognition, 35(6), 1323-1336. doi: 10.3758/bf03193604 kaakinen, j. k., & hyönä, j. (2008). perspective-driven text comprehension. applied cognitive psychology, 22, 319-334. doi: 10.1002/acp.1412 kaakinen, j. k., & hyönä, j. (2010). task effects on eye movements during reading. journal of experimental psychology: learning, memory and cognition, 36(6), 1561-1566. doi: 10.1037/a0020693 kaakinen, j. k., hyönä, j., & keenan, j. m. (2002). perspective effects on online text processing. discourse processes, 33(2), 159-173. doi: 10.1207/s15326950dp3302_03 lai, m.-l., tsai, m.-j., yang, f.-y., hsu, c.-y., liu, t.-c., lee, s. w.-y., . . . tsai, c.-c. (2013). a review of using eye-tracking technology in exploring learning from 2000 to 2012. educational research review, 10, 90-115. doi: 10.1016/j.edurev.2013.10.001 lonka, k., olkinuora, e., & mäkinen, j. (2004). aspects and prospects of measuring studying and learning in higher education. educational psychology review, 16(4), 301-323. doi: 10.1007/s10648–004– 0002–1 marton, f., dahlgren, l. o., säljö, r., & svensson, l. (1975). the göteborg project on non-verbatim learning. göteborg: university of göteborg. mason, l., tornatora, m. c., & pluchino, p. (2013). do fourth graders integrate text and picture in processing and learning from an illustrated science text? evidence from eye-movement patterns. computers & education, 60(1), 95-109. doi: 10.1016/j.compedu.2012.07.011 mcnamara, d. s. (2004). sert: self-explanation reading training. discourse processes, 38(1), 1-30. moss, j., schunn, c. d., schneider, w., mcnamara, d. s., & vanlehn, k. (2011). the neural correlates of strategic reading comprehension: cognitive control and discourse comprehension. neuroimage, 58(2), 675-686. doi: 10.1016/j.neuroimage.2011.06.034 olsson, p. (2007). real-time and offline filters for eye tracking. kth royal institute of technology. penttinen, m., anto, e., & mikkilä-erdmann, m. (2012). conceptual change, text comprehension and eye movements during reading. research in science education, 43(4), 1407-1434. doi: 10.1007/s11165-012-9313-2 perry, n. e., & winne, p. h. (2006). learning from learning kits: gstudy traces of students' self-regulated engagements with computerized content. educational psychological review, 18, 211-228. doi: 10.1007/s10648-006-9014-3 ponce, h. r., & mayer, r. e. (2014). an eye movement analysis of highlighting and graphic organizer study aids for learning from expository text. computers in human behavior, 41, 21-32. doi: 10.1016/j.chb.2014.09.010 rayner, k. (1998). eye movements in reading and information processing: 20 years of research. psychological bulletin, 124(3), 372-422. doi: 10.1037//0033-2909.124.3.372 richardson, j. t. e. (2000). researching student learning. buckingham: open university press and srhe. richardson, j. t. e. (2013). research issues in evaluating learning pattern development in higher education. studies in educational evaluation, 39(1), 66-70. doi: 10.1016/j.stueduc.2012.11.003 rothkopf, e. z. (1966). learning from written instructive materials: an exploration of the control of inspection behavior by test-like events. american educational research journal, 3, 241-249. doi: 10.3102/00028312003004241 samuelstuen, m. s., & braten, i. (2007). examining the validity of self-reports on scales measuring students' strategic processing. britisch journal of educactional psychology, 77(pt 2), 351-378. doi: 10.1348/000709906x106147 schellings, g. l. m. (2011). applying learning strategy questionnaires: problems and possibilities. metacognition and learning, 6(2), 91-109. doi: 10.1007/s11409-011-9069-5 schellings, g. l. m., van hout-wolters, b., & vermunt, j. d. (1996). individual differences in adapting to three different tasks of selecting information form texts. contemporary educational psychology, 21, 423-446. doi: 10.1006/ceps.1996.0029 catrysse  et  al       16 | f l r scouller, k. m. (1998). the influence of assessment method on students' learning approaches: multiple choice question examination versus assignment essay. higher education, 35, 453-472. doi: 10.1023/a:1003196224280 scouller, k. m., & prosser, m. (1994). students' experiences in studying for multiple choice question examinations. studies in higher education, 19(3), 267-279. doi: 10.1080/03075079412331381870 segers, m., nijhuis, j., & gijselaers, w. (2006). redesigning a learning and assessment environment: the influence on students' perceptions of assessment demands and their learning strategies. studies in educational evaluation, 32, 223-242. doi: 10.1016/j.stueduc.2006.08.004 van gog, t., & jarodzka, h. (2013). eye tracking as a tool to study and enhance cognitive and metacognitve processes in computer-based learning environments. in r. azevedo & v. a. w. m. m. aleven (eds.), international handbook of metacognition and learning technologies. new york: springer. van gog, t., paas, f., & van merrienboer, j. j. g. (2005). uncovering expertise-related differences in troubleshooting performance: combining eye movement and concurrent verbal protocol data. applied cognitive psychology, 19(2), 205-221. doi: 10.1002/acp.1112 veenman, m. v. j. (2005). the assessment of metacognitive skills: what can be learned from multi-method designs? in c. artett & b. moschner (eds.), lernstrategien und metakognition. implikationen für forschung und praxis (pp. 77-99). münster: waxmann. veenman, m. v. j. (2011). alternative assessment of strategy use with self-report instruments: a discussion. metacognition and learning, 6(2), 205-211. doi: 10.1007/s11409-011-9080-x veenman, m. v. j., bavelaar, l., de wolf, l., & van haaren, m. g. p. (2014). the on-line assessment of metacognitive skills in a computerized learning environment. learning and individual differences, 29, 123-130. doi: 10.1016/j.lindif.2013.01.003 vermunt, j. d., & vermetten, y. j. (2004). patterns in student learning: relationships between learning strategies, conceptions of learning, and learning orientations. educational psychology review, 16(4), 359-384. doi: 10.1007/s10648-004-0005-y frontline learning research 5 (2014) 1 28 issn 2295-3159 corresponding author: dorit alt doritalt@014.net.il http://dx.doi.org/10.14786/flr.v2i3.68 1 | f l r the construction and validation of a new scale for measuring features of constructivist learning environments in higher education dorit alt a a kinneret college on the sea of galilee, israel article received 4 december 2013 / revised 11 february 2014 / accepted 5 june 2014 / available online 13 june 2014 abstract this study was aimed at mapping features of constructivist activities in higher education settings, constructing and validating a new scale for measuring their presence in lecture face-to-face based environments (lbe), seminars (sm), and distance learning environments (dle). a mix-method approach was implemented in three phases. the first phase was aimed at qualitatively analysing classroom observational activities as experienced by students, in order to learn about actual instantiations of the theoretical constructivist features. the results foregrounded eight categories: 'knowledge construction', 'authenticity', 'multiple perspectives', 'prior knowledge', 'in-depth learning', 'teacherstudent interaction', 'social interaction' and 'cooperative dialogue'. the second phase was aimed at developing a questionnaire, based on the descriptions gathered in phase 1. the third quantitative phase was used to validate the developed questionnaire (constructivist learning in higher education settings scale [clhes]) by using structural equation modelling. in addition, students' academic self-efficacy had been chosen as a criterion variable in order to further assess construct validity of the clhes. lastly, a multivariate analysis of covariance was applied to allow the characterisation of differences between the learning settings in regard to the clhes eight factors and academic self-efficacy. the scales were submitted to 597 undergraduate third-year college students. according to the main results: construct validity of the new scale has been confirmed; teacher-student and student-student interactions were positively connected to self-efficacy for learning; and sm were perceived as generally more constructivist when compared with the other learning environments. implications of these findings and directions for future research are discussed. keywords: constructivism; academic self-efficacy, higher education d. alt 2 | f l r 1. introduction educational practice is continually subjected to renewal needs, due mainly to the growing proportion of information communication technology, social changes, globalisation of education, and the pursuit of quality. the accelerating rate of social change puts a premium on adaptability to the emerging requirements of present society such as communication and cooperation skills, and ability to critically select, acquire, and use knowledge (quisumbing, 2005; wegerif & de laat, 2011). these types of renewal needs require developing updated instructional practices that could integrate knowledge with the personal transferable skills (pellegrino & hilton, 2012). in order to meet the demands of 21th century learning needs, the creation of learning environments based on the constructivist pedagogy is suggested to engage learners in knowledge construction carried out by social negotiated tasks in real-world contexts while enhancing students' ability to regulate their learning (de kock, sleegers, & voeten, 2004). the constructivist approach has taken a leading theoretical position and has become a powerful driving force in the dynamic relationship between teaching methods and learning processes. however, despite the growing attention paid to constructivist pedagogic challenges in the context of learning environments, the instructional principles of this theory, which are aimed at directing the nature of educational processes, still need to be clarified (gijbels, van de watering, dochy, & van den bossche, 2006). nonetheless, during the past two decades, attempts to map instructional constructivist principles of educational materials and learning environments have yielded a few results in the field of university teaching (fraser, treagust, williamson, & tobin, 1987; tenenbaum, naidu, jegede, & austin, 2001). for example, tenenbaum et al. (2001) defined and empirically examined seven key features of constructivist learning environments: (1) arguments, discussions, debates, (2) conceptual conflicts and dilemmas, (3) sharing ideas with others, (4) materials and resources targeted toward solutions, (5) motivation toward reflections and concept investigation, (6) meeting students‟ needs, and (7) making meaning, real-life examples. however, alt (in press) maintains that this scale could be further elaborated to include additional perceptions on a wider range of theoretical dimensions that are important to the current situation in higher education setting. for example, understanding the students' prior knowledge (meyer, 2004); constructing environments for teaching and learning that are decompartmentalised (minick, stone, & forman, 1993); and engaging students in a self-regulated learning, in which they can set their own goals, mediate new meanings from existing knowledge, and form an awareness of current knowledge structures (hakkarainen, lipponen & järvelä, 2002). therefore, constructing a new scale for measuring a wider range of constructivist features in university learning environments is central for this study. other scales, such as the approaches to study inventory (asi) or the approaches to learning and studying inventory (alsi) (entwistle & ramsden, 1983), and the student process questionnaire (r-spq-2f) (biggs, kember, & leung, 2001), were used to measure constructivist learning by means of students' approaches to learning. these studies were based on the assumption that constructivist learning environments are aimed at fostering a deep (rather than surface) approach to learning (lea, stephenson, & troy, 2003; tiwari et al., 2006). approaches to learning refer to how students perceive themselves going about learning in a specific learning situation, and focus on how intention and process are combined in students' deep or surface learning (biggs et al., 2001). it has been recognised that these approaches to learning are not characteristics of learners but are determined by a relation between a learner and a context, and that students adjust their approaches to learning depending on the requirements of the task (evans, 2014). however, the nature of learning tasks and contexts has changed dramatically in the last decade in terms of depth and range of curricula and the diversity of settings (e.g., distance learning), thus the depth of learning in constructivist environments could currently refer to diversified requirements of those d. alt 3 | f l r environments, pertaining to the process of 'learning to learn', learning to gain an internal control for learning, and learning how to cooperate within communities of enquiry (de kock et al., 2004). therefore, assessing constructivist features implementation in current higher education learning contexts is of importance and lies at the core of the present study. moreover, both teacher and student are assumed to be jointly responsible for the outcome, the teacher for structuring the enabling conditions, the learner for engaging them, thus an approach to learning is described as the nature of the relationship between student, context, and task (biggs et al., 2001). however, the learning approaches scales seem to put emphasis on the learners, disregarding some significant theoretical components of learning patterns such as students' perceptions of the learning context that could affect their learning engagements (cano & garcía-berbén, 2014). in order to bridge the gap between theory and empirical study, this study will assess the relations between three learning dimensions: students' constructive learning activity perceptions, teacher-student engagements and students' social activity. finally, current studies have suggested that constructivist learning environments do not always promote students' deep learning, and point to several factors that limit the effectiveness of those learning settings (baeten, kyndt, struyven, & dochy, 2010; gijbels, segers, & struyf, 2008; kyndt, dochy, & cascallar, 2014). for example, kyndt et al. (2014) maintain that these learning environments demand too much from the students in terms of workload and task complexity, in these cases inducing a deep approach to learning could be difficult. based upon those studies, it seems important to detect possible relations between the learners and their social learning environment that could encourage them to become selfregulatory and support their confidence and ability to excel in complex tasks required for constructivist learning. hence, this mix-method study represents an effort to map features of constructivist learning environments, construct and validate a new scale for measuring facets of constructivist learning and asses their perceived implementation in several higher education learning contexts. moreover, since previous studies have consistently link students' academic self-efficacy (bandura, 1997) to learning settings based on the constructivist theory (dorman & adams, 2004; dorman, fisher, & waldrip, 2006), this psychological outcome has been chosen as a criterion variable in order to further assess construct validity of the new scale. this study could detect effective constructivist practices in university learning settings and measure their connection to self-efficacy for learning. revealing interrelations among several constructivist practices could provide practical implementations, informed by the constructivist theory, for higher education teaching practices. finally, the potential differences between various forms of contemporary learning settings: lecture based environments, seminars and distance learning environments, and the assessment of the use of constructivist activities in these settings, will be addressed in this study. such comparative examination could demonstrate how different constructivist activities could be applied in various settings as well as challenge the positive effect attributed to constructivist based environments on academic self-efficacy. 2. theoretical framework 2.1 the constructivist pedagogy constructivism is a view of learning that perceives the individual as an active and responsible agent in his/her knowledge acquisition process (brooks & brooks, 1999). this view is shared by cognitive constructivism and social constructivism. however, while cognitive constructivism is concerned with the individual's construction of knowledge, social constructivism stresses the collaborative processes in d. alt 4 | f l r knowledge building (windschitl, 2002). these epistemological emphases are exemplified by bakhtin (1984, 1986). for bakhtin (1984), meaning is a product of dialogues: "truth is not born nor is it to be found inside the head of an individual person; it is born between people collectively searching for truth, in the process of their dialogic interaction" (p. 110). several essential factors of the social constructivist pedagogy are indicated by theorists and practitioners (packer & goicoechea, 2001; popkewitz, 1998; steffe & gale, 1995). these features may be grouped around three key tenets of the constructivist learning environment in line with de kock et al.'s (2004) classification: constructive activity, teacher-student interaction and social activity, as further described below. the first tenet (constructive activity) pertains to the process of 'learning to learn'. this principle is based on several educational practices. first is the idea that learning occurs during sustainable participation in inquiry practices focused on the advancement of knowledge. this process, consists of a so-called predict observeexplain procedure (white & gunstone, 1992) where learners hypothesise, test their hypothesis, explain observations as a way of verifying hypothesis, and later discuss discrepancies between the hypothesis and the outcome. in this format, learners‟ participation throughout the lesson will be through predicting, observing and explaining the learning process. in this process, learners are required to actively make meaning from information; they cannot be passive consumers of conceptualisations, analyses and conclusions of others. however, although university teaching is claimed to have a special task to support students in adopting ways of thinking and producing new knowledge anchored in scientific inquiry practices (gellin, 2003; resnick, 1987), stahl (2011) argues that students' habits of learning are still overwhelmingly skewed toward passive acquisition of knowledge from authority sources rather than from collaborative inquiry activities. authenticity is another dimension of the constructive activity tenet. authentic experiences allow the individual to construct mental structures that are viable in meaningful situations. since learning is contextual, knowledge construction should occur in situations that are real rather than contrived (dolittle & camp, 1999). situating learning in a real world task ensures that learning is personally interesting, and provides the students with opportunities to think at the level of sophistication they are likely to encounter in the real world (erstad, 2011). lahn (2011) maintains that more attention should be paid to contextual variables that provide learners with a wide range of authentic experiences, and scaffolds that support an effective reorganisation of knowledge, while conceiving learners as active designers of their learning environment. providing multiple perspectives and representations of a content, is another dimension of the constructive activity tenet. the constructivist learning encourages the student to examine a phenomenon from several points of view (perspectives). when students are able to examine an experience from multiple perspectives, their understanding and adaptability are increased. in this process they are forced to go beyond everyday ethical contemplation by developing dialogue and multiple perspectives as well as drawing on available resources (lund & hauge, 2011). this practice provides students with multiple opportunities to develop a more viable model of their learning and social experiences (dolittle & camp, 1999). another dimension of the constructive activity first tenet refers to the idea that content and skills should be understood within the framework of the learner's prior knowledge (dochy & alexander, 1995). teachers should be able to ascertain their students' prior knowledge and teach accordingly. by understanding the student's mental structures, teachers can clarify incomplete or erroneous prior knowledge, determine the method of instruction necessary in a particular topic area, create effective experiences and plan independent activities, and assess materials adapted to the student (meyer, 2004). teachers should also create environments for teaching and learning that are decompartmentalised, by integrating individual, social and d. alt 5 | f l r institutional processes, as stressed by minick et al. (1993): "...one cannot develop a viable socio-cultural conception of human development without looking carefully at the way these institutions develop, the way they are linked with one another, and the way human social life is organised within them" (p. 6). hence, contrary to the traditional ideology of teaching and learning, which relies mainly upon learning opportunities that are the mere “spelled out” transmission of dominant knowledge, according to the new interdisciplinary approach, experiences retrieved from the past could offer mediations to decipher present experience, and lessons learned from prior inquiry could be turned towards a creative future (perret-clermont & perret, 2011). this approach is considered an efficient way to help teachers and learners deal with acquiring knowledge that grows at exponential proportions within change processes (jacobs, 1989). the second tenet (teacher-student interaction) is one of the main conceptual pillars of the constructivist pedagogy. this principle stresses on the self-regulated learner, and on shifting the external control over the learning process, as used in conventional and wellstructured learning settings, to the student's internal control for learning. in these processes, students should be encouraged to become selfregulatory, self-mediated, and self-aware (de kock et al., 2004). students are given opportunities to actively engage in self-regulated learning processes, including setting their own goals, mediating new meanings from existing knowledge, and forming an awareness of current knowledge structures (hakkarainen et al., 2002). the teacher role is to engage students in a self-regulated learning, often referred to as meta-cognition (brown, 1987), and encourage students to set their own goals while emphasising collaboration and negotiation. the teacher should also provide scaffolding during the learning process, while encouraging and guiding students to reflect on their own learning processes, rather than acting as a knowledge conduit (järvelä, hurme, & järvenoja, 2011). king (2002) describes this learning as a deliberate process during which learners focus on their performance and think carefully about the thinking that led to particular actions, what happened and what they are currently learning from the experience, in order to better perform in the future. according to the final tenet (social activity), learning is a social activity in which individual learning processes are affected by personal characteristics as well as by external social factors, and meaning is constructed from the interaction between existing knowledge and social situations (vygotsky, 1978). this principle highlights the cooperative nature of the learning process aimed at fostering a dialogic thinking (schwarz, 2009; schwarz & de groot, 2011; wegerif, 2007). the dialogic interpretative framework implies that pedagogic practices should be able to sustain more than one perspective simultaneously. this pedagogy has been described by wegerif and de laat (2011) in terms of moving learners into the space of dialogue. this process includes the promotion of communities of enquiry and dialogue skills through the use of forums of alternative voices, and the induction of students into real dialogues across cultural differences. järvelä et al. (2011) maintain that successful engagement in such collaborative and dialogic learning involves core processes of self-regulated learning, effective use of learning strategies to participate in collaborative interactions, meta-cognitive control, and regulation of motivation and emotions. 2.2 features of constructivist learning in higher education environments although the conventional lecture form has been consistently associated with the traditional one-way traffic instruction, based on objectivist philosophical assumptions, nave (1991) implies that several constructivist activities could be implemented in university lecture based settings. she distinguishes a conventional lecture from an 'open-text' lecture. in a conventional lecture, learners simply absorb new materials, without being allowed to raise questions. in contrast, an 'open-text' lecture allows the teacher to manoeuvre his/her ways from time to time, present the material from multiple points of view, and use varied d. alt 6 | f l r examples which are relevant to the students' world. during these activities, teachers can promote dialogic processes in the classroom. nave (1991) maintains that this complex and challenging approach necessitates qualified teachers who have the special skills required for this 'open-text' instructional design. another higher education environment is the distance learning, defined as a planned activity that occurs in a different place from the teacher, far from the designated learning place, using special techniques for designing online courses (barak & dori, 2009). the philosophy of constructivism seems to have crucial implications for learning and instructional design in distance learning settings. in the neo-vygotskian sociocultural theory, technology is seen as a facilitator of dialogic spaces where students can use networks to creative learning (wegerif & de laat, 2011). with the rapid growth of distance learning courses, it seems worthwhile to examine how distance learning settings support the use of constructivist activities. additional learning environment, based on the constructivist pedagogical approach, is the researchbased seminar. seminars include intense study relating to the student's major, typically have significantly fewer students per professor than normal courses, and are generally more specific in topic of study. these settings are conceived as excellent ways by which a community of learners could be built, interdisciplinary research-based (i.e. inquiry-based) settings could be promoted, and student-centred activities, where students themselves could take a key role in creating the research/learning link, could be fostered (lueddeke, 2003). despite the many theoretical appeals of comparing between traditional learning environments and constructivist based environments, few are the empirically based studies. for example, tynjälä (1999) showed how students in a constructivist learning environment acquire more diversified knowledge when compared with students in a traditional teaching setting. however, the potential differences between various forms of contemporary learning settings and the assessment of the use of constructivist activities in these settings are yet to be explored. such comparative examination could demonstrate how different constructivist activities could be applied in various settings. 2.3 academic self-efficacy an important psychological outcome addressed in previous research concerning constructivist teaching and learning, is academic self-efficacy (bandura, 1977, 1986). studies have stressed that academic self-efficacy is a positive predictor of academic achievement (carroll et al., 2009), and of self-motivation for academic attainment (bandura, 1997), therefore measuring the potential contribution of different learning environments to this psychological outcome is of importance. academic self-efficacy refers to personal judgements of one‟s ability to succeed at an academic task on a designated level or to attain a specific academic goal (bandura, 1997; linnenbrink & pintrich, 2002). accordingly, self-efficacy competence includes behavioural actions as well as the cognitive skills necessary for performance in a specific domain, and has been defined as “an individual‟s confidence in their ability to organise and execute a given course of action to solve a problem or accomplish a task” (eccles & wigfield, 2002, p. 110). according to bandura (1997), learners with the same level of cognitive skill development could differ in their intellectual performances due to the strength of their perceived self-efficacy. previous studies (dorman & adams, 2004; dorman et al., 2006; loyens, rikers, & schmidt, 2008; van dinther, dochy, & segers, 2011), link self-efficacy competence to the psychosocial learning environment that students experience in their schools and classrooms, and report a consistent contribution of the constructivist learning environment to students' academic self-efficacy. donche, coertjens, van daal, de maeyer and van petegem (2014) showed how academic self-efficacy has a positive direct effect on first year d. alt 7 | f l r university students' deep learning engagement. dorman and adams (2004) suggest that the potential of the constructivist learning environment in explaining academic self-efficacy should be recognised. 2.4 the present study this study attempts at first, mapping features of actual constructivist learning instantiations in higher education settings, second, constructing and validating a new scale for measuring those features, third, assessing the constructivist features implementation in different higher education settings, and fourth, measuring their effect on self-efficacy for learning. this study's main research questions were formulated as: q1. to what extent do students' perceptions of the presence of constructivist learning practices in their classes contribute to their academic self-efficacy? which perceived constructivist practices are connected to students' academic self-efficacy? q2. which learning environment sufficiently reflects an assemblage of constructivist tenets, and promotes academic self-efficacy? figure 1. demonstrates the theoretical structure of the proposed theoretical framework. figure 1. model 1. the theoretical structure of the proposed framework. 3. method a mix qualitative and quantitative research method, applied in three phases, was used to address the research aims and questions. creswell (2007) emphasised the superiority of a mixed-method research design in exploratory research. this method builds upon the synergy that exists between the qualitative-quantitative research continuum thus allowing to reinforce research construct validity and to expand the understanding of an explored phenomenon. 3.1 phase 1 the first phase was aimed at gathering and analysing classroom observational activities as experienced by students, in order to learn about actual instantiations of the theoretical constructivist features. this phase used a qualitative methodology to analyse the gathered materials according to the categorical scheme suggested by theory, while allowing for additional meaningful categories identification. d. alt 8 | f l r 3.1.1 participants and material gathering procedure phase 1 included 62 undergraduate third-year students from one major college in israel, (12.5% male students 84.6% female students). their distribution with respect to faculties was as follows: education15 students, criminology – ten students, sociology – 12 students, management – four students, economy – five students, behavioural sciences – eight students, political sciences four students, and communication four students. participants were asked to keep observation diaries of their learning activities in one of the following courses: a seminar (sm), a lecture based environment course (lbe) or a distance learning environment course (dle). since the following analysis procedure involved both deductive and inductive category applications, a prescribed general format of the diary was given, and three theoretical foci were suggested to assist observations: learning activity, teacherstudent interaction and social activity. there was also a selfreflection section in the diary. 3.1.2 analysis of the study materials in line with the deductive approach, a categorical scheme suggested by the theoretical perspective was defined (see the independent variable shown in fig. 1). the inductive approach allowed identifying additional meaningful categories. according to strauss (1987), both these aspects of inquiry are absolutely essential throughout the analysis. thus, both logically derived categories and those that have "serendipitously" arisen from the data may find their way into the research (merton, 1968). students' observations were analysed by four raters; all are experts in the research area of constructive learning. inter-rater cohen's kappa (k) reliability (cohen, 1960), which is commonly assessed in psychological research, was used. the raters were asked to categorise the students' observation reports according to the theoretical scheme. the k values were interpreted as follows: k < 0.20 poor agreement; 0.21 < k < 0.40 fair agreement; 0.41 < k < 0.60 moderate agreement; 0.61 < k < 0.80 good agreement; 0.81 < k < 1.00 very good agreement. results of 0.61 < k < 1 were considered acceptable for the purposes of the current study. the raters were also asked to report on new identified categories. 3.2 phase 2: questionnaire development this phase was aimed at developing a questionnaire that could assess constructivist activities in various educational settings. the students' descriptions gathered in the qualitative research (phase 1), where formulated as short items by three instructional design experts in the research area of constructive learning. for example, the following description of dle: "assignments were given during this course on moodle (modular object-oriented dynamic learning environment). this allowed me preparing the required work when i chose to; i could progress at my own pace" was phrased as: 'in this course, the teacher considered my learning pace' (c12). each item was given a likert-type score ranging from 1 = not at all true to 5 = completely true. consequently, a 41-item scale was submitted to 78 undergraduate third-year students in order to assess the clarity of the items. accordingly, five items were excluded due to unclear phrasing. the new scale (hereinafter: constructivist learning in higher education settings scale [clhes]) included 36 items. 3.3 phase 3 this quantitative phase was used to validate the developed questionnaire by using structural equation modelling (sem) (bentler, 2006; mcdonald & ho, 2002). in addition, since previous studies have d. alt 9 | f l r consistently link students' academic self-efficacy to constructivist learning settings, this psychological outcome had been chosen as a criterion variable to further assess construct validity of the new scale. additional aim of this phase was to test the research questions. 3.3.1 the criterion variable: academic self-efficacy an eight-item (g1 – g8) scale derived from the motivated strategies for learning questionnaire (mslq) (pintrich, smith, garcia, & mckeachie, 1993) was used to assess perceived academic competence in the students' learning environments. the mslq was originally designed to measure college undergraduates‟ motivation and self-regulated learning perception and learning strategies. the mslq is modular, thus allows using the sub-scales separately, as has been the case in the present study, which used only the academic self-efficacy sub-scale. all items were scored on a 5-point likert scale with anchors of 1 = strongly disagree to 5 = strongly agree. for example, 'i'm certain i can master the skills being taught in this course.' (cronbach's alpha = 0.89). 3.3.2 participants the clhes and mslq were submitted to 597 undergraduate third-year students (15.4% males and 84.6% females) from one major college in israel, of whom 37.5% were jewish students and 62.5% muslim students, with a mean age of 24.5 (sd=4.7) years. based on the report of the central bureau of statistics (2011) and the council for higher education (2009) in israel, the gender and ethnicity breakdown of northern galilee college students, majoring mainly in social sciences studies, is 20% males and 80% females of whom 40% jewish, 55% muslim, and 5% belonging to other religions, thus the current study's sample represents, to some extent, the gender and ethnicity breakdown of regional colleges located in the northern galilee. the distribution of the participants with respect to course settings (course groups) was as follows: 29.1% lbe students (enrolled in three randomly selected courses), 40.2% seminar course students (sm), (enrolled in eight randomly selected courses), and 30.7% dle students (enrolled in three randomly selected courses). the sample reflected the faculty enrollment breakdown of the campus, composed as follows: education – 63%, criminology – 12.8%, sociology – 7.9%, management 7.5%, economy – 4.3%, behavioural sciences 2%, political sciences 19. %, and communication – 0.6%. 3.3.3 procedure the clhes was administered to the participants near the end of their courses at the second semester of the third year of studies. the students were told that the purpose of the study was to examine their perceptions of the course. prior to obtaining participants' consent it was specified that the questionnaires were anonymous and that no pressure would be applied should they choose to return the questionnaire unfilled or incomplete (the overall response rate was 87%; 34 questionnaires were excluded due to incomplete response). finally, participants were assured that no specific identifying information about the courses would be processed. the scale items were originally generated in hebrew, and were translated into english and back translated by professional editors for the purpose of this paper. 4. findings 4.1 phase 1. qualitative study results table 1 presents the categories and several examples from the students' reports. in line with the theoretical framework, five categories have been recognised from the reports: knowledge construction, d. alt 10 | f l r authenticity, multiple perspectives, prior knowledge and teacherstudent interaction. an additional category of in-depth learning has emerged from the analysis. moreover, the theoretical category of social activity has been divided into two distinctive sub-categories: social interaction and cooperative dialogue, as further described below: 1) knowledge construction is described as multiple opportunities given to students to investigate real problems, raise questions and search for possible explanations while using various methodological approaches. 2) in-depth learning. this category pertains to the extent to which students are given opportunities to deeply explore a certain subject matter, rather than engaging them in a surface learning. 3) authenticity, deals with giving relevant meaning to the learned concepts and addressing real life and interesting events which are related to the studied topic. 4) the multiple perspectives category refers to presenting complex ideas from several points of view. 5) prior knowledge primarily deals with connecting the subject materials to other courses' topics. 6) teacherstudent interaction refers to the teacher role which includes guidance toward reflection on learning processes. 7) social interaction includes a variety of learning activities with other students, not necessarily during a lesson. 8) cooperative dialogue refers to dialogical activities during the lesson in which students can express opinions and original ideas. it can be learned from table 1 that the pedagogical principles introduced in the theoretical framework and in the analysis were associated with various course formats. for example, the following example shows how authentic real life examples are integrated in a lecture based course: "this course, entitled 'social roles', deals with the family life span, especially with men's and women's roles in different societies, for example, conflict situations within the family. the examples given in class reflect real situations from our daily life." a reversed description (rv) is a report in which students describe a lack of a constructive related activity in the learning environment, for example, the following report exemplifies how the teacher does not implement dialogical activities during a lecture based lesson: "when students want to comment on a specific issue that has been taught in class, the teacher explains that they have no right to do so, since "much better scholars than them have investigated the issue". eventually, everyone silently obeys the teacher." table 1 categories and examples from students' reports. note: seminars (sm), lecture based environments (lbe), distance learning environments (dle), reversed description (rv) category examples knowledge  in this course we have investigated an interesting issue related to parents' d. alt 11 | f l r construction empowerment in educational processes, with relation to different cultural needs. this inquiry required interviewing parents; some of them were parents of children with special needs. we also interviewed educational teams in order to find ways to enrich parental involvement in schools and communities.(lbe)  i want to explore how teenagers from different cultures experience their adolescence period. in order to find an answer to my question, i have to interview parents from different ethnic groups.(sm)  this course involved a field work. we went to kindergartens in our city and explored how different theoretical approaches can be applied in real situations. the conclusions of our experiences were later discussed in the class.(lbe)  students have presented their research work in class. they have described the whole process from the start: stated their research question, described the preferred methodology, presented the data analysis, research findings and conclusions.(sm) in-depth learning  this course required preparing a project regarding the skills of the school counsellor. this was really an intensive work that included a deep study of this topic. (sm)  the teacher shows us power point presentations loaded with complex figures i cannot understand. he moves from one topic to another, sometimes i really get confused.(rv) (lbe)  the main goal [of this course] is the final exam. we study in order to pass the exam. there was no enriching beyond the concepts required for the exam. we could not ask questions during classes in order to deepen our understanding, since "there is no time for questions". (rv)(lbe)  sometimes i get very interested in a subject raised by the teacher, at this point, disappointedly, she moves on to another subject. i feel that the quantity is much more important for her than the quality. (rv)(lbe) authenticity  the teacher uploads assignments to the course website. these assignments concern current educational issues. we are also required to search for news and to find items regarding the studied material. (dle)  this course, entitled 'social roles', deals with the family life span, especially with men's and women's roles in different societies, for example, conflict situations within the family. the examples given in class reflect real situations from our daily life. (lbe)  one of the requirements of this course was conducting a research assignment related to problems which arab women are confronted with when leaving their close environment sphere towards academic studies, and the obstacles they encounter when they get back to their villages to work. this is an interesting issue; i was highly motivated to take part in this investigation. (sm)  one of the topics was the history of the maccabiah [an international jewish athletic event]. we have studied the subject through protocols of interviews with past athletes, newspapers articles and stories related to the history of this event.(lbe) multiple perspectives  the subject of this lesson was 'sexual assault'. each student could present his or her attitude. different perspectives were brought up by the students. one of them argued that women "bring it upon themselves" and should dress in a more modest manner. others disagreed and argued that religious girls in their d. alt 12 | f l r villages, although dressed by the religious code, were sexually abused. (lbe)  in this course we talk about different codes of norms of several religions: jewish, muslim, christian and druze. at first, every student introduced his/her tradition regarding the dressing code, then, we asked each other questions regarding for example, the origin of these codes, and the obstacles arise within a multicultural society with relation to these codes. (lbe)  in this lesson we have discussed the subject of 'egalitarian division of labour within the family'. some female students were against the idea of equal sharing, one of them argued that her husband is working hard and this is enough labour for him, and that from her point of view women should take care for domestic issues only. other students strongly opposed this position. maybe their different cultures effect their point of view.(lbe) prior knowledge  the main topic dealt with the transition to parenthood. this subject was related first, to my previous experience as a mother, and second, to many subjects such as psychology, childhood era, conflicts in the family, which i have learned during the past year.(lbe)  in this lesson we learned about ethics in research. the teacher showed us videos of the milgram's experiment on obedience to authority figures. i have learned about this experiment in a psychology related course earlier this year, however, this moral perspective has broadened my knowledge. (lbe)  one of the discussed topics was on unmarried couples who choose to have a parenting agreement. this issue raised many important aspects that were related to several course materials i had previously studied, such as: parents and parenting, the child's security and needs. (lbe) teacher student interaction  one of my assignments was to present a theme with relation to the studied material. the teacher encouraged me to search for papers, she has given me a general guidance on how and where to find scientific materials related to my subject.(lbe)  assignments were given during this course on moodle (modular objectoriented dynamic learning environment). this allowed me preparing the required work when i chose to; i could progress at my own pace. (dle)  the teacher knows every single student by his/her name. she always encourages me. after my class presentation, she sent me an email in which she had appreciated my progress and added some comments on how to improve my learning process. (sm)  in this course the assignments are given in a way which allows me to organise my schedule in a flexible manner.(dle) social interaction  during this course arab and jewish students have cooperated on multiple occasions. for example, the hebrew language is very difficult for non-native speakers, so in many occasions during a cooperative in -class or out-class work, jewish students helped arab students correcting spelling mistakes and improving oral presentations.(lbe)  i have kept downloading materials from the website, nothing else was needed. i was not required to work with others, frankly, i did not know the students participating in the course .(rv)(dle)  the teacher encourages us to use the forum. she raises questions and asks us to comment and hold a debate. however, in practice, it seems that many students d. alt 13 | f l r invest their time in answering her questions, and do not pay any attention to students' comments.(rv)(dle) cooperative dialogue  the discussed subject was conflict in the family with relation to the "coming out of the closet" issue. a female student shared her private experience in this context with us. people got excited, students in this class come from different cultures, some of them religious, and therefore very different voices were heard. (lbe)  although defined as a lecture based course, discussions were held in every lesson. for example, the jewish ancient law of halitza was discussed. according to this law, a jewish widow would need to marry her brother-in-law unless he freed her in a ceremony known as halitza. many students wished to say something about it. some argued that this ceremony is no longer valid even in orthodox communities. others suggested that this is another example of an anti-feminist realty imposed by religion. through these dialogues i have become more interested in the studied material.(lbe)  when students want to comment on a specific issue that has been taught in class, the teacher explains that they have no right to do so, since "much better scholars than them have investigated the issue". eventually, everyone silently obeys the teacher. (rv)(lbe) 4.2 phase 2. descriptive statistics, internal consistency and construct validity of the clhes table 2 presents the clhes factors, sub-factors, item descriptions (as derived from phase 2) and internal consistencies (cronbach‟s alpha). items 10, 25, 30 were excluded from the analysis due to low item loading results (< .30) found in the structural equation modelling (fig.2). each of the eight factors showed a very high internal consistency. table 3 provides descriptive statistics for the clhes factors (n = 597). table 4 displays the bivariate correlation analysis results among the clhes factors and between these factors and the academic self-efficacy criterion variable. convergent validity has been shown by positive statistically significant correlations between all factor pairings. meaning, the measures of the constructivist factors that theoretically are related to each other are in fact observed to be related to each other. the generally moderate correlations among the dimensions suggest that the factors are, to some extent, independent each from the other. finally, as can be learned from table 4, the correlation coefficients shown between the clhes factors and the academic self-efficacy variable are lower than the amongconstructivistfactor coefficients. therefore, discriminant validity of the clhes scale may be confirmed. these conditions were posited by campbell & fiske (1959) as evidence supporting construct validity. d. alt 14 | f l r table 2 the clhes questionnaire: factors, sub-factors, item descriptions and internal consistencies (cronbach’s alpha) factors and subfactors item cronbach‟s alpha constructive activity (f1) knowledge construction (a1) c1. in this course, i was given opportunities to investigate real problems (five items) .93 c2. during this course, i was given opportunities to raise questions about complex problems c3. during this course, i was given opportunities to search for possible explanations for real problems c4. i was asked to analyse data regarding a significant problem i have raised during this course c5. during this course, i was asked to draw conclusions from a research work, in which i have participated constructive activity (f1) in-depth learning (a2) c6. in this course, i have learned skills with which i can deeply explore a subject of interest to me (four items, item c10 was omitted due to a low loading result) .87 c7. i could examine in depth a major issue in this course c8. in this course, i have focused on a central subject which i was required to deeply understand c9. in this course, i have learned how to deeply investigate a certain subject c10. in this course, we "jump" from one subject to another without examining any subject in depth* constructive activity (f1) authenticity (a3) c16. this course addressed interesting situations in reality (five items) .87 c17. the course focused on giving relevant meaning to the learned concepts c18. the course addressed real life and interesting events c19. the course was rich with real-life examples that interested me c20. the course did not addressed real life examples* constructive activity (f1) c21. in this course, ideas were presented from several points of view (four items, item c25 was omitted due c22. i have learned about complex real issues in this course d. alt 15 | f l r multiple perspectives (a4) c23. i have realised that the reality is complex and multi – dimensional, in this course to a low loading result) .81 c24. in this course, i had to question and criticise accepted ideas c25. in this course, ideas were presented from only one perspective, and were not allowed to be criticised* constructive activity (f1) prior knowledge (a5) c26. this course dealt with subjects i have learned in other courses (four items, item c30 was omitted due to a low loading result) .85 c27. the subjects learned in this course were related to prior knowledge i have gained c28. things i have learned in this course have helped me understand issues i have learned in other courses c29. the subjects in this course were related to diverse contents of knowledge c30. the subjects in this course were not related to other things i have learned in other courses* teacherstudent interaction (f2) c11. in this course, the teacher allowed me to think about my learning and how to improve it (five items) .91 c12. in this course, the teacher considered my learning pace c13. in this course, i could set myself some learning goals c14. in this course, the teacher encouraged me to think about my learning and ways to improve it c15. in this course, the teacher made me think about the advantages and disadvantages of my learning social activity (f3) social interaction (h1) c31. this course included a variety of learning activities with other students (three items) .88 c32. i was given opportunities to learn with other students in this course c33. i could collaborate with other students in this course social activity (f3) cooperative dialogue (h2) c34. arguments and discussions were held during this course (three items) .89 c35. it was possible to express original ideas in this course c36. in this course, i could express my opinion, even when it was different from other students * reversed items d. alt 16 | f l r table 3 descriptive statistics for the clhes measured factors kurtosis skewness sd mean factor -0.815 -0.31 1.11 3.11 knowledge construction (a1) -0.26 -0.54 0.99 3.41 in-depth learning (a2) 0.43 -0.76 0.86 3.59 authenticity (a3) .40 -0.54 0.79 3.41 multiple perspectives (a4) 0.36 -0.62 0.87 3.42 prior knowledge (a5) -0.24 -0.55 0.95 3.33 teacherstudent interaction (f2) -0.62 -0.35 1.09 3.13 social interaction (h1) 0.02 -0.63 0.99 3.48 cooperative dialogue (h2) table 4 bivariate correlation matrix for the eight factors of the clhes scale and academic self-efficacy factors 1 2 3 4 5 6 7 8 academic selfefficacy 1 knowledge construction (a1) .775 ** .557 ** .589 ** .409 ** .589 ** .458 ** .495 ** .336 ** 2 in-depth learning (a2) .604 ** .623 ** .497 ** .663 ** .501 ** .465 ** .364 ** 3 authenticity (a3) .686 ** .535 ** .623 ** .435 ** .455 ** .309 ** 4 multiple perspectives (a4) .533 ** .628 ** .488 ** .520 ** .302 ** 5 prior knowledge (a5) .546 ** .423 ** .380 ** .328 ** 6 teacherstudent interaction (f2) .457 ** .436 ** .385 ** 7 social interaction (h1) .595 ** .286 ** 8 cooperative dialogue (h2) .291 ** p < .01** d. alt 17 | f l r 4.3 phase 3 4.3.1 testing the first research question structural equation modelling (sem) was employed to test the first research question (q1), and to further assess the construct validity of the clhes, using a confirmatory factor analysis. data used for the sem were analysed with the maximum likelihood method. three fit indices were computed in order to evaluate model fit: χ2(df), (p > .05), cfi (> 0.9), and rmsea (< 0.08). the structural model (fig. 2) refers to the combined measurement and path models. the measurement model includes the following factors: first, the constructive activity (f1) latent variable accompanied by five latent variables: knowledge construction (a1) with five observed items (c1 – c5); indepth learning (a2) with four observed items (c6 – c9); authenticity (a3) with five observed items (c16 – c20); multiple perspectives (a4) with four observed items (c21 – c24); and prior knowledge (a5) with four observed items (c26 – c29); second, the teacherstudent interaction (f2) latent variable accompanied by five observed variables (c11 – c15); third, the social activity (f3) latent variable accompanied by two latent variables: social interaction (h1) with three observed items (c31 – c33) and cooperative dialogue (h2) with three observed items (c34 – c36). the path model was constructed as follows: three paths were specified between the latent factors f1 – f3 and the criterion latent variable of academic self-efficacy (se) which was accompanied by eight observed items (g1 – g8). the goodness of fit of the data to the model yielded to sufficient fit results (χ2 = 2079.36, df = 766, p = .000; cfi = .926; rmsea = .054). the results showed positive low significant coefficients between the teacherstudent interaction (f2) factor and the criterion variable of academic self-efficacy (β = .23, p < .01) and between the social activity (f3) factor and the criterion variable (β = .22, p < .05). an insignificant coefficient result was indicated between the constructive activity (f1) factor and the dependent variable. as shown in fig. 2, the clhes factors together explained 36% of the academic self-efficacy criterion variable variance. 4.3.2 testing the second research question in order to test the second research question (q2), a multivariate analysis of covariance (mancova) with bonferroni pair-wise comparisons and wilks' lambda criterion was applied to allow the characterisation of differences between the course groups (lbe, sm and dle) in regard to a linear combination of the multiple eight dependent factors of the clhes. in addition, an analysis of covariance (ancova) with bonferroni pair-wise comparisons was used to assess betweencourse group differences on the academic self-efficacy variable. the variables of gender (1 = male, 2 = female) and cultural group (1 = jewish, 2= muslim) were entered as covariates to neutralise any significant confounding effect in the analyses of variance. table 5 shows the mean scores, standard deviations, f values, wilks' lambda and partial eta-squared statistics of the analyses. results indicated significant differences between the course groups regarding the combination of the multiple clhes factors and separately on each of them. all the betweengroup differences were accompanied by moderate to large effect sizes, when small, moderate, and large effects are reflected in values of ηp2 equal to .0099, .0588, and .1379, respectively (cohen, 1969, pp. 278–280; richardson , 2011, p. 142). d. alt 18 | f l r d. alt 19 | f l r figure 2. the structural model, with standardised parameter estimates (n= 597). note: *p < .05 **p < .01 ***p < .001. table 5 mean scores, sd, f values, wilks' lambda, partial eta-squared statistics (ηp 2 ) and bonferroni pair-wise comparisons of the three course groups (lbe, sm and dle) on the eight clhes factors and the academic self-efficacy variable. the numbers of the pair-wise comparisons indicate: 1=the lowest mean result, 2= in between, 3= the highest mean result, identical numbers indicate insignificant between-group differences. course groups sm dle lbe dependent variables m sd m sd m sd f ηp 2 factors of the clhes scale wilks' lambda statistic (main effect) 27.90*** .277 anova knowledge construction (a1) 3.87 0.70 2.99 0.99 2.20 0.90 183.75*** .384 pair-wise comparisons 3 2 1 in-depth learning (a2) 3.98 0.65 3.35 0.90 2.68 0.97 115.10*** .281 pair-wise comparisons 3 2 1 authenticity (a3) 3.95 0.66 3.40 0.74 3.27 1.00 40.56*** .121 pair-wise comparisons 3 1 1 multiple perspectives (a4) 3.71 0.67 3.35 0.71 3.06 0.87 34.06*** .104 pair-wise comparisons 3 2 1 prior knowledge (a5) 3.68 0.73 3.43 0.83 3.04 0.95 24.92*** .078 pair-wise comparisons 3 2 1 teacherstudent interaction (f2) 3.72 0.79 3.32 0.83 2.83 1.01 45.06*** .133 pair-wise comparisons 3 2 1 social interaction (h1) 3.38 1.04 3.41 0.94 2.49 1.04 39.40*** .118 d. alt 20 | f l r pair-wise comparisons 3 3 1 cooperative dialogue (h2) 3.82 0.81 3.32 0.99 3.18 1.08 24.91*** .078 pair-wise comparisons 3 1 1 covariate effect gender .020 cultural group .067 academic self-efficacy pair-wise comparisons 4.07 0.04 3.90 .05 3.76 0.05 10.69*** .035 3 3 1 covariate effect gender .000 cultural group .020 note: p < .05 * p < .01** p < .001*** as presented in table 5, salient betweengroup differences were indicated for the factors: knowledge construction (a1) (ηp2 = .384) and in-depth learning (a2) (ηp2 = .281). on each factor, the lowest mean result was indicated for the lbe group and the highest for the sm group. somewhat lower effect sizes were found for three factors: teacherstudent interaction (f2) (ηp2 = .133) the lowest mean result was indicated for the lbe group and the highest for the sm group; authenticity (a3) (ηp2 = .121), with a significant higher score shown for the sm group compared with the other groups; and social interaction (h1) (ηp2 = .118) the lowest mean result was indicated for the lbe group and the highest results were shown for the sm and dle groups. the relatively lowest effect sizes were found for three factors: multiple perspectives (a4) (ηp2 = .104), prior knowledge (a5) (ηp2 = .078), on each factor, the lowest mean result was indicated for the lbe group and the highest for the sm group; and cooperative dialogue (h2) (ηp2 = .078) with a significant higher score indicated for the sm group compared with the other groups. regarding the academic self-efficacy variable, differences were found between the three groups, accompanied by a low effect size (ηp2 = .035) the highest results were indicated for the sm and dle groups and the lowest for the lbe group. d. alt 21 | f l r 5. discussion the overarching goals of this study were to map features of constructivist learning environments, construct and validate a new scale for measuring the presence of those features in different higher education settings, by using a mix-method approach. 5.1 the qualitative analysis consistent with previous theoretical research (de kock et al., 2004) this research revealed three key tenets of the constructivist learning environment: constructive activity, teacher-student interaction and social activity. regarding the constructive activity tenet, the results foregrounded five categories: knowledge construction, authenticity, multiple perspectives, prior knowledge, and in-depth learning. this research elaborates the body of literature by adding the sub-category of in-depth learning which emerged from the content analysis. this facet pertains to the extent to which students are given opportunities to deeply explore a certain subject matter, in order to seek a clearer understanding of the learning materials, in contrast to surface learning which is confined to rote learning and memorising facts. although in-depth learning is not a new concept, this research has empirically demonstrated its relation to constructive activities in higher education settings. moreover, the theoretical category of social activity has been divided into two distinctive facets: cooperative dialogue and social interaction. social interaction includes a variety of learning activities with other students, not necessarily during a lesson, whereas cooperative dialogue refers to dialogical activities during the lesson in which students can express opinions and original ideas. another finding regarding the qualitative research was that some constructivist pedagogical principles are associated with lecture based courses. for example, according to the students' reports, teachers of lecture based courses have used real-life examples during their lectures. some students reported on dialogical activities during lectures in which students could express opinions and original ideas. these findings were partially corroborated by the quantitative analysis results according to which, lbe and dle were perceived by the students to be equally consistent with the authenticity and cooperative dialogue constructivist features. although, the quantitative analyses have revealed that lbe are generally less consistent with other examined constructivist features compared with sm and dle formats, these findings may imply that some constructivist features can be applied in lecture based environments, in accordance with nave (1991). 5.2 the quantitative analysis phase – perceptions of the learning environments the main result of this phase showed that students perceive sm learning environments as more constructivist when compared with perceptions held by other course groups (lbe and dle). since sm settings are conceived as excellent ways by which constructivist activities could be fostered (lueddeke, 2003), this finding could have been expected, and thus could further validate the new scale. additional findings showed that dle are generally perceived as more constructivist than lbe, and less constructivist when compared with sm environments. however, no differences were shown between dle and lbe in authenticity and cooperative dialogue activities. although technology is seen as a facilitator of dialogic spaces (wegerif & de laat, 2011), according to this research findings, it may be inferred that this practice is inadequately applied by teachers. researchers (e.g., östlund, 2008) argue that guaranteeing collaboration for learning can be difficult to achieve in dle. in order to achieve this goal, d. alt 22 | f l r learners should be encouraged to use the forum, and teachers should stimulate interaction by creating assignments in which the learners can be actively engaged in discussion. nonetheless, the factor social interaction, which includes a variety of learning activities with other students, was similarly applied in dle and sm, compared with lbe. this could suggest that students of dle courses tend to be more engaged in off-line cooperative activities than during 'on-line' dialogues. 5.3 the quantitative analysis phase academic self-efficacy and perceptions of the learning environments additional important findings regard the criterion variable of academic self-efficacy. this study's empirical model indicates that stimulating meta-cognitive and reflective aspects of learning, through teacherstudent interaction, could bolster the students‟ confidence in their ability to accomplish a task. studies indicate that students who develop strong academic self-efficacy beliefs are better able to manage their learning, and consequently are more likely to successfully complete their education and be better equipped for a variety of occupational options in today's competitive society (bandura, barbaranelli, caprara, & pastorelli, 2001). accordingly, this study suggests that educators should be aware of the importance of pursuing this affective outcome by motivating the students to think reflectively, regarding the individuals' learning process. through this process of evaluating their own performance as learners, students could become active participates in their development (king, 2002), and consequentially, as suggested by this study, more confident in their ability to execute assignments. the social activity factor was found to be the second positive predictor of academic self-efficacy. this factor deals with the need to encourage interaction and collaboration among students. interaction is perceived to be one of the most important components of the learning experience, in which students are given sufficient opportunities to express themselves and to share their own experiences with others (dewey, 1938; tenenbaum et al., 2001; vygotsky, 1978). a recent study shows that effective cooperative learning communities support knowledge acquisition (wyatt et al., 2010). the present research indicates that social interaction could also benefit academic self-efficacy. a plausible explanation could be that interactions with others allow the learners to reflect on their own work and to make independent use of their results thus being able to perform more effectively as suggested by vygotsky (1978) and bandura (1986). moreover, encouraging interaction and collaboration among students could have provided sufficient opportunities for students to observe other group members. such vicarious experience could be gained in collaborative assignments provided by the learning environment, and could affect students' perceptions of their own ability to perform (bandura, 1997). moreover, students who worked together could have been encouraged to share their views and evaluations of other students in their group. having them identify the strengths of others, rather than their weaknesses, might have benefited their self-efficacy beliefs (schunk & miller, 2002). the present study stresses the importance of facilitating cooperative tutorial study groups not only in order to create a well-functioning environment, but also to nurture self-efficacious learners in higher education studies. it should be noted that according to this study's result, both sm and dle courses were more positively associated with increased self-efficacy for learning compared with lbe courses. this result could be theoretically explained by the firm contribution attributed to the philosophy of constructivism to learning and instructional design in distance learning settings and research-based seminar (lueddeke, 2003; wegerif & de laat, 2011). empirically, this result could be explained by the sm and dle emphasis on interpersonal interactions compared with the lbe courses, according to the participants' report. d. alt 23 | f l r lastly, the factor constructive activity was not found to be significantly connected to the selfefficacy dependent variable. it could be inferred that the social interaction dimensions of the learning environments are more prominent in explaining self-efficacy for learning. nonetheless, the positive high connections found between the three tenets of constructive activity, teacher-student interaction and social activity could suggest an indirect connection between constructive activities and academic self-efficacy through increased interpersonal interactions. 5.4 limitations and directions for future research first limitation is that the clhes scale constructed and validated in this study could be further elaborated. for example, this scale did not include characteristics of assessment as components of the constructivist learning environment. assessment is considered part of the fabric of classrooms to which students attach importance. assessment tasks that do not match student learning could lower the confidence of students for successfully performing academic tasks (dorman et al., 2006). thus further research is needed to examine this mediator measure with relation to higher education. second, future research should also consider expanding the model tested here with additional variables that could be related to learning activities such as, academic motivation psychological variables. these variables could be related to learning setting perceptions and academic self-efficacy, therefore assessing them in conjunction with the present study examined constructs is of importance and could allow measuring additional constructivist environments' effects on psychological constructs. third limitation concerns the cross-sectional nature of the data which can prevent definitive statements about causality. definitive proof of mediation will also require longitudinal data (cole & maxwell, 2003). it should be further acknowledged that alternate models might explain the relationships in these data as well as the one tested in this study. in fact, many relationships in the model are likely reciprocal. for example, although the analysis implies that the self-efficacy construct is mainly informed by the teacher-student interaction factor, it is equally plausible that teachers may become more involved with self-efficacious students. despite such possibilities, the path model could represent a reasonable, theoretically grounded structure of the relations between the examined factors. however, researchers should extend this work with longitudinal paradigms. lastly, this study was conducted in a single country, meaning that the results cannot necessarily be generalised. therefore, larger population studies are needed to validate these findings, and more research on this topic needs to be undertaken before the associations between the perceived learning environment and self-efficacy belief are more clearly understood. despite its limitations, this study underscores the importance of interpersonal relationships to students' psychological outcomes, specifically, the significant roles of teacher-studentand student-student relationships in enhancing academic self –efficacy are recognised in this study. keypoints a qualitative analysis of classroom observational activities has foregrounded eight factors: 'knowledge construction', 'authenticity', 'multiple perspectives', 'prior knowledge', 'in-depth learning', 'teacherstudent interaction', 'social interaction' and 'cooperative dialogue'. based on the qualitative analysis results, the constructivist learning in higher education settings scale [clhes] was developed. d. alt 24 | f l r construct validity of the clhes was confirmed by using structural equation modelling. teacher-student interactions and student-student social activities were positively connected to self-efficacy for learning. seminars (sm) were perceived as generally more constructivist when compared with lecture based environments (lbe) and distance learning environments (dle). acknowledgments this research was supported by a grant from the maslovaty foundation for the advancement of education on morals and society, founded by dr. nava maslovaty of blessed memory. references alt, d. (in press) assessing the contribution of constructivist based academic learning environment to academic self-efficacy. learning environments research. baeten, m., kyndt, e., struyven, k., & dochy, f. (2010). using student-centred learning environments to stimulate deep approaches to learning: factors encouraging or discouraging their effectiveness. educational research review, 5, 243-260. doi: 10.1016/j.edurev.2010.06.001 bakhtin, m. (1984). problems of dostoevsky’s poetics. (c. emerson, ed. & trans.). minnieapolis: university of michigan press. bakhtin, m. (1986). speech genres and other late essays. austin, tx: university of texas press. bandura, a. (1977). social learning theory. new york: prentice hall. bandura, a. (1986). the explanatory and predictive scope of self-efficacy theory. journal of social and clinical psychology, 4, 359-373. doi: 10.1521/jscp.1986.4.3.359 bandura, a. (1997). self-efficacy: the exercise of control. new york: freeman. bandura, a., barbaranelli, c., caprara, g. v., & pastorelli, c. (2001). self-efficacy beliefs as shapers of children‟s aspirations and career trajectories. child development, 72(1), 187-206. doi: 10.1111/1467-8624.00273 barak, m., & dori, y. j. (2009). enhancing higher order thinking skills among in-service science teachers via embedded assessment. journal of science teacher education, 20(5), 459 474. bentler, p. m. (2006). eqs 6 structural equations program manual. encino, ca: multivariate software, inc. biggs, j. b., kember, d., & leung, d.y.p. (2001). the revised two factor study process questionnaire: rspq-2f. british journal of educational psychology, 71 ,133-149. doi: 10.1348/000709901158433 brooks, j. g., & brooks, m. g. (1999). in search of understanding: the case for constructivist classrooms. alexandria, va: association for supervision and curriculum development. brown, a. l. (1987). metacognition, executive control, selfregulation and other mysterious mechanisms. in f. weinert & r. kluwe (eds.), metacognition, motivation and understanding (pp. 65–115). hillsdale, nj: lawrence erlbaum. campbell, d. t., & fiske, d. w. (1959). convergent and discriminant validation by the multitraitmultimethod matrix. psychological bulletin, 56, 81-105. doi: 10.1037/h0046016 cano, f., & garcía-berbén, a-b. (2014). university students' achievement goals and approaches to learning in mathematics: a re-analysis investigating 'learning patterns'. in d. gijbels, v. donche, j. t. e. richardson, & j. d. vermunt (eds.), learning patterns in higher education: dimensions and research perspectives (pp. 163 – 186). london and new york: routledge and earli. carroll, a., houghton, s., wood, r., unsworth, k. l., hattie, j., gordon, l., & bower, j. (2009). selfefficacy and academic achievement in australian high school students: the mediating effects of academic aspirations and delinquency. journal of adolescence, 32(4), 797-817. doi: 10.1016/j.adolescence.2008.10.009 http://www.routledge.com/books/search/author/david_gijbels/ http://www.routledge.com/books/search/author/vincent_donche/ http://www.routledge.com/books/search/author/john_t._e._richardson/ http://www.routledge.com/books/search/author/john_t._e._richardson/ http://www.routledge.com/books/search/author/jan_d._vermunt/ d. alt 25 | f l r central bureau of statistics. (2011). women in higher education. retrieved december 21, 2012 from http://www1.cbs.gov.il/www/publications/desc_exp/women.pdf (hebrew) cohen, j. (1960). a coefficient of agreement for nominal scales. educational and psychological measurement, 20, 37-46. doi: 10.1177/001316446002000104 cohen, j. (1969). statistical power analysis for the behavioural sciences. new york: academic press. cole, d. a., & maxwell, s. e. (2003). testing mediational models with longitudinal data: questions and tips in the use of structural equation modeling. journal of abnormal psychology, 112, 558–577. council for higher education. (2009). planning and budgeting committee 34/35 report. retrieved july 20, 2010 from http://www.che.org.il/download/files /contents_1.pdf (hebrew) creswell, j. w. (2007). educational research (3rd ed.). thousand oaks, ca: sage. de kock, a., sleegers, p., & voeten, m. j. m. (2004). new learning and the classification of learning environments in secondary education. review of educational research, 74(2), 141-170. doi: 10.3102/00346543074002141 dewey, j. (1938). experience and education. new york: touchstone. dochy, f. j. r. c., & alexander, p. a. (1995). mapping prior knowledge: a framework for discussion among researchers. european journal of psychology of education, 10, 225-242. doi: 10.1007/bf03172918 donche, v., coertjens, l., van daal, t., de maeyer, s., & van petegem, p. (2014). understanding differences in student learning and academic achievement in first year higher education: an integral research perspective. in d. gijbels, v. donche, j. t. e. richardson, & j. d. vermunt (eds.), learning patterns in higher education: dimensions and research perspectives (pp. 214 – 231). london and new york: routledge and earli. doolittle, p. e., & camp, w. g. (1999). constructivism: the career and technical education perspective. journal of vocational and technical education, 16(1), 23-46. retrieved from http://scholar.lib.vt.edu/ejournals/jvte/v16n1/doolittle.html dorman, j. p., & adams, j. (2004). associations between students' perceptions of classroom environment and academic efficacy in australian and british secondary schools. westminster studies in education, 27, 69 – 85. doi: 10.1080/0140672040270106 dorman, j. p., fisher, d., & waldrip, b. (2006). classroom environment student` perceptions of assessment, academic efficacy and attitude to science. a lisrel analysis. in d. l. fisher & m. s. khine (eds.), contemporary approaches to research on learning environments: world views (pp. 1 -28). singapore: world scientific. eccles, j. s., & wigfield, a. (2002). motivational beliefs, values, and goals. annual review of psychology, 53, 109-132. doi: 10.1146/annurev.psych.53.100901.135153 entwistle, n. j., & ramsden, p. (1983). understanding student learning. london: croom helm. erstad, o. (2011). weaving the context of digital literacy. in s. ludvigsen, a. lund, i. rasmussen & r. säljö (eds.), learning across sites: new tools, infrastructures and practices (pp. 295-310). london: routledge. evans, c. (2014). exploring the use of a deep approach to learning with students in the process of learning to teach. in d. gijbels, v. donche, j. t. e. richardson, & j. d. vermunt (eds.), learning patterns in higher education: dimensions and research perspectives (pp. 187 – 213). london and new york: routledge and earli. fraser, b. j., treagust, d. f., williamson, j. c., & tobin, k. g. (1987). validation and application of the college and university classroom environment inventory (cucei). in b. j. fraser (ed.), the study of learning environments (pp. 17-30). perth, western australia: curtin university of technology. gellin, a. (2003). the effect of undergraduate student involvement on critical thinking: a metaanalysis of the literature 1991–2000. journal of college student development, 44, 745–762. doi:10.1353/csd.2003.0066 gijbels, d., segers, m., & struyf, e. (2008). constructivist learning environments and the (im)possibility to change students‟ perceptions of assessment demands and approaches to learning. instructional science, 36, 431–443. doi:10.1007/s11251-008-9064-7 http://www.routledge.com/books/search/author/david_gijbels/ http://www.routledge.com/books/search/author/vincent_donche/ http://www.routledge.com/books/search/author/john_t._e._richardson/ http://www.routledge.com/books/search/author/jan_d._vermunt/ javascript:__dolinkpostback('detail','mdb%257e%257eaph%257c%257cjdb%257e%257eaphjnh%257c%257css%257e%257ejn%2520%252522westminster%2520studies%2520in%2520education%252522%257c%257csl%257e%257ejh',''); javascript:__dolinkpostback('detail','mdb%257e%257eaph%257c%257cjdb%257e%257eaphjnh%257c%257css%257e%257ejn%2520%252522westminster%2520studies%2520in%2520education%252522%257c%257csl%257e%257ejh',''); http://www.routledge.com/books/search/author/david_gijbels/ http://www.routledge.com/books/search/author/vincent_donche/ http://www.routledge.com/books/search/author/john_t._e._richardson/ http://www.routledge.com/books/search/author/jan_d._vermunt/ d. alt 26 | f l r gijbels, d., van de watering, g., dochy, f., & van den bossche, p. (2006). new learning environments and constructivism: the students' perspective. instructional science, 34(3), 213-226. doi:10.1007/s11251-005-3347-z hakkarainen, k., lipponen, l., & järvelä, s. (2002). epistemology of inquiry and computer supported collaborative learning. in t. koschmann, n. miyake & r. hall (eds.), cscl2: carrying forward the conversation (pp. 129–156). mahwah, nj: erlbaum. jacobs, h. h. (1989). interdisciplinary curriculum: design and implementation. alexandria, va: asdc. järvelä, s., hurme, t.-r., & järvenoja, h. (2011). selfregulation and motivation in computersupported collaborative learning environments. in s. ludvigsen, a. lund, i. rasmussen & r. säljö (eds.), learning across sites: new tools, infrastructures and practices (pp. 330-345). london: routledge. king, t. (2002, july). development of student skills in reflective writing. paper presented at the 4 th world conference of the international consortium for educational development in higher education, perth, australia. doi: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.136.2518 kyndt, e., dochy, f., & cascallar, e. (2014). students' approaches to learning in higher education: the interplay between context and student. in d. gijbels, v. donche, j. t. e. richardson, & j. d. vermunt (eds.), learning patterns in higher education: dimensions and research perspectives (pp. 249 – 272). london and new york: routledge and earli. lahn, l. c. (2011). professional learning as epistemic trajectories. in s. ludvigsen, a. lund, i. rasmussen & r. säljö (eds.), learning across sites: new tools, infrastructures and practices (pp. 53-68). london: routledge. lea, s., stephenson, d., & troy, j. (2003). higher education students' attitudes to student-centred learning: beyond 'educational bulimia'. studies in higher education, 28, 321-333. doi:10.1080/03075070309293 linnenbrink, e. a., & pintrich, p. r. (2002). motivation as an enabler for academic success. school psychology review, 31, 313-327. retrieved from http://prof.usb.ve/jjramirez/pregrado/ccp114/ccp114%20motivacion%20linnenbrink%20y% 20pintrich.pdf loyens, s. m. m., rikers, r. m. j. p., & schmidt, h. g. (2008). relationships between students‟ conceptions of constructivist learning and their regulation and processing strategies. instructional science, 36, 445–462. doi:10.1007/s11251-008-9065-6 lueddeke, g. r. (2003). professionalising teaching practice in higher education: a study of disciplinary variation and „teaching-scholarship‟. studies in higher education, 28, 213-228. doi: 10.1080/0307507032000058082 lund, a., & hauge, t. e. (2011). changing objects in knowledge-creation practices. in s. ludvigsen, a. lund, i. rasmussen & r. säljö (eds.), learning across sites: new tools, infrastructures and practices (pp. 206-221). london: routledge. mcdonald, r. p., & ho, m.-h. (2002). principles and practice in reporting structural equation analyses. psychological methods, 7, 64 – 82. doi: 10.1037/1082-989x.7.1.64. 64 merton, r. k. (1968). social theory and social structure. new york: free press. meyer, h. (2004). novice and expert teachers' conceptions of learners' prior knowledge. science education, 88(6), 970 983. doi: 10.1002/sce.20006 minick, n., stone, c. a., & forman, e. a. (1993). introduction: integration of individual, social, and institutional processes in accounts of children's learning and development. in e. a. forman, n. minick, & c. a. stone (eds.), contexts for learning: sociocultural dynamics in children's development (pp. 3 15). new york: oxford university press. nave, h. (1991). in favour of the frontal teaching. the israeli ministry of education: the division for curriculum planning and development. retrieved from http://www.education.gov.il/tochniyot_limudim/sifrut/asi12020.htm (hebrew) östlund, b. (2008). prerequisites for interactive learning in distance education: perspectives from swedish students. australasian journal of educational technology, 34, 42-56. retrieved from http://www.westga.edu/~distance/ojdla/fall163/ekstrand164.html packer, m. j., & goicoechea, j. (2001). sociocultural and constructivist theories of learning: ontology, not just epistemology. educational psychologist, 35, 227–241. doi: 10.1207/s15326985ep3504_02 http://www.routledge.com/books/search/author/david_gijbels/ http://www.routledge.com/books/search/author/vincent_donche/ http://www.routledge.com/books/search/author/john_t._e._richardson/ http://www.routledge.com/books/search/author/jan_d._vermunt/ http://www.routledge.com/books/search/author/jan_d._vermunt/ d. alt 27 | f l r pellegrino, j. w., & hilton, m. l. (eds.). (2012). education for life and work: developing transferable knowledge and skills in the 21st century. washington, d.c: the national academies press. perret-clermont , a.-n., & perret, j.-f. (2011). a new artifact in the trade: notes on the arrival of a computer supported manufacturing system in a technical school. in s. ludvigsen, a. lund, i. rasmussen & r. säljö (eds.), learning across sites: new tools, infrastructures and practices (pp. 87-102). london: routledge. pintrich, p. r., smith, d., garcia, t., & mckeachie, w. (1993). reliability and predictive validity of the motivated strategies for learning questionnaire (mslq). educational and psychological measurement, 53, 801-813. doi: 10.1177/0013164493053003024 popkewitz, t. s. (1998). dewey, vygotsky and the social administration of the individual: constructivist pedagogy as systems of ideas in historical spaces. american educational research journal, 35, 535– 570. doi:10.3102/00028312035004535 quisumbing, l. r. (2005). education for the world of work and citizenship: towards sustainable future societies. prospects: quarterly review of comparative education, 35, 289–301. doi: 10.1007/s11125-005-4266-0 resnick, l. (1987). education and learning to think. washington, dc: national academy press. richardson, j. t. e. (2011). eta squared and partial eta squared as measures of effect size in educational research. educational research review, 6, 135–147. doi:10.1016/j.edurev.2010.12.001 schunk, d. h., & miller, s. d. (2002). self-efficacy and adolescents‟ motivation. in f. pajares & t. urdan (eds.), academic motivation of adolescents (pp. 29-52). greenwich, ct: information age. schwarz, b. (2009). argumentation and learning. in n. mullermirza & a.n. perret-clermont (eds.), argumentation and education – theoretical foundations and practices (pp. 91–126). new york and london: springer. schwarz, b., & de groot, r. (2011). breakdowns between teachers, educators and designers in elaborating new technologies as precursors of change in education to dialogic thinking. in s. ludvigsen, a. lund, i. rasmussen & r. säljö (eds.), learning across sites: new tools, infrastructures and practices (pp. 261-277). london: routledge. stahl, g. (2011). social practices of group cognition in virtual match teams. in s. ludvigsen, a. lund, i. rasmussen & r. säljö (eds.), learning across sites: new tools, infrastructures and practices (pp. 190-205). london: routledge. steffe, l. p., & gale, j. (1995). constructivism in education. mahwah, nj: lawrence erlbaum associates. strauss, a. l. (1987). qualitative analysis for social scientists. cambridge: cambridge university press. tenenbaum, g., naidu, s., jegede, o., & austin, j. (2001). constructivist pedagogy in conventional oncampus and distance learning practice: an exploratory investigation. learning and instruction 11, 87–111. doi:10.1016/s0959-4752(00)00017-7 tiwari, a., chan, s., wong, e., wong, d., chui, c., wong, a., & patil, n. (2006). the effect of problembased learning on students‟ approaches to learning in the context of clinical nursing education. nurse education today, 26, 430-438. doi: http://dx.doi.org/10.1016/j.nedt.2005.12.001 tynjälä, p. (1999). towards expert knowledge? a comparison between a constructivist and a traditional learning environment in the university. international journal of educational research, 33, 355–442. doi: http://dx.doi.org/10.1016/s0883-0355(99)00012-9 van dinther, m., dochy, f., & segers, m. (2011). factors affecting students‟ self-efficacy in higher education. educational research review, 6(2), 95–108. doi: 10.1016/j.edurev.2010.10.003 vygotsky, l. s. (1978). mind and society: the development of higher mental processes. cambridge, ma: harvard university press. wegerif, r. (2007). dialogic, education and technology: expanding the space of learning. new york: springer. wegerif, r., & de laat, m. (2011). using bakhtin to rethink the teaching of higher-order thinking for the network society. in s. ludvigsen, a. lund, i. rasmussen & r. säljö (eds.), learning across sites: new tools, infrastructures and practices (pp. 313-329). london: routledge. white, r. t., & gunstone, r. (1992). probing understanding. london: the falmer press. d. alt 28 | f l r windschitl, m. (2002). framing constructivism in practice as the negotiation of dilemmas: an analysis of the conceptual, pedagogical, cultural, and political challenges facing teachers. review of educational research, 72(2), 131–175. doi: 10.3102/00346543072002131 wyatt, t. h., krauskopf, p. b., gaylord, n. m., ward, a., huffstutler-hawkins, s., & goodwin, l. (2010). cooperative m-learning with nurse practitioner students. nursing education perspectives, 31(2), 109-113. doi: http://dx.doi.org/10.1043/1536-5026-31.2.109 frontline learning research 1 (2013) 24 41 issn 2295-3159 corresponding author: ming fai pang, the university of hong kong, pangmf@hku.hk, t (852) 28592428, f (852) 28585649 http://dx.doi.org/10.14786/flr.v1i1.16 24 | f l r meanings are acquired from experiencing differences against a background of sameness, rather than from experiencing sameness against a background of difference: putting a conjecture to the test by embedding it in a pedagogical tool ference marton a , ming fai pang b a university of gothenburg, sweden b the university of hong kong, hong kong sar, china article received 26 march 2013 / revised 31 may 2013 / accepted 30 june 2013 / available online 27 august 2013 abstract in helping learners to make a novel meaning their own, such as when helping children to understand what a word means or teaching students a new concept in school, we frequently point to examples that share the aimed-at meaning but differ otherwise. this type of approach rests on the assumption that novel meanings can be acquired through the experience of sameness against a background of difference. this paper argues that this assumption is unfounded and that the opposite is the case: we make novel meanings our own through the experience of differences against a background of sameness. we put this conjecture to the test in an experimental study by embedding it in a computer game and the results support the conjecture. . keywords: variation theory; phenomenography; discernment; critical experiment 1. the conjecture this paper is about a conjecture and how it is put to the test. the conjecture is actually the title of the paper and we first briefly describe the theory that elaborates its implications, together with some previous results. after that we report on a study which is meant to be a critical test of it. we call this conjecture—and the system of corollaries that it implies—somewhat immodestly the variation theory (of learning) (marton, forthcoming; marton & tsui, 2004). f. marton & m.f. pang 25 | f l r 1.1 the origin of meaning it is commonly believed that a child or an adult for that matter can learn the meaning of a word by observing a number of examples of what the word refers to, that share this meaning but differ in other ways. for example, we point to a dog and say dog, point to another dog and say dog, point to a third dog and say dog, and then expect the child to understand what the word ―dog‖ means (refers to), i.e., a certain kind of animal. in an experimental context such a learning event could look like the following: ―…children might be shown a red fuzzy triangle labeled ―wug‖, a blue bumpy triangle labeled ―wug‖, a green scratchy triangle labeled ―wug‖ and then at the test be asked to pick out a ―wug‖ (a yellow squishy triangle) among two or three objects‖ (vlach et al, 2008). now, if a child has noticed previously that there are different geometric forms of which triangle is one, it is most likely that she will see that the three things are different, but they are all triangles regardless if she has learned that they are called ―triangle‖. hence she will identify the yellow squishy triangle in the test as a ―wug‖. but if she has never noticed triangles previously, or geometric forms in general, she will not see any triangles at all. in consequence, she will not be able to see what the different cases have in common. there is no way of learning the idea of triangle in such an experimental context, if you have not come across that idea earlier. but if you have, you might be able to learn that triangles are called ―wug‖ in the actual context. in the same way, no child can learn that dogs are a kind of animals, without coming across other animals than dogs. the idea of triangle derives from how it differs from other geometric forms, and the idea of dog derives from how it differs from other animals. providing different examples of the same thing is not only the most common method of helping young children to build a vocabulary, but probably also the most common method of teaching concepts, principles, and problem-solving methods in school. stigler and hiebert (1999) describe such an approach as the typical way of teaching mathematics in u.s. schools, and the highly authoritative volume how people learn urges teachers to provide ―…many examples in which the same concept is at work.‖ (bransford, brown, & cocking, 2000, p 20) looking at cases that are the same in one respect but differ in others to determine what they have in common is called induction. according to fodor (1980), this is the only idea that exists to explain how novel meanings (concepts) are learned, and it simply does not work, for the reasons already cited. it follows then that there is no explanation of how we learn, find, create, or appropriate new meanings. hence, by default, fodor concludes that meanings (concepts) are innate. in our view, however, even if the concept (meaning) of ―dog‖ were innate, you would never be able to separate that meaning from the meaning of ―animal‖ if you had never encountered any animals other than dogs. regardless of whether dogs were then called dogs or animals, the meaning of ―dog‖ would be exactly the same as the meaning of ―animal.‖ hence, you still would not have acquired the meaning of ―dog.‖ nor of ―animal‖ for that matter. the meaning of dog has to be learned, and this happens by coming across dogs, as well as other animals. similarly, if we lived in an entirely green world, then we would be unable to notice the greenness of everything. hence, whether or not concepts (meanings) are innate, we must encounter alternatives to them if we are to be able to notice and grasp these concepts. awareness of a particular number presupposes awareness of other numbers (or at least one other number), and awareness of a particular color presupposes awareness of other colors (or at least one other color). you cannot possibly understand what chinese is simply by listening to different people speaking chinese if you have never heard another language, and you cannot possibly understand what virtue is by inspecting different examples of the same degree of virtue. nor can you understand what a ―linear equation‖ is by looking only at linear equations. you cannot arrive at a novel meaning through induction, but you can through contrast. in induction, the focused meaning, i.e., the one that you are trying to help another to make his or her own (e.g. ―chinese‖) is kept invariant, while the other features of the same entity (e.g. words) vary. in contrast, it is just the other way around. the focused meaning (e.g. language) varies, while other features (e.g. words) are invariant. instead of saying different words in the same language (chinese), you say the same word in different languages (one of which is chinese). f. marton & m.f. pang 26 | f l r inductive learning is a frequent research topic, not the least in the field of machine learning (e.g. michalski, 1983). as the conjecture being put to the test in our study is about how novel meanings are acquired, and as it states that they are not acquired through induction, we will leave those studies aside here. 1.2 earlier attempts to put the conjecture to the test most work on variation theory has been carried out in the form of learning studies, the inspiration for which is the japanese lesson study, which came to wider attention through the publication of stigler and hiebert’s (1999) best-selling book the teaching gap. in this type of study, a group of teachers teaching a particular subject at a particular level together choose an object of learning (something to be learned) that is vitally important for students’ continued learning and that has earlier been found to present difficulties for them. the teachers plan a lesson together, and one of them carries it out usually in his or her own class while the others observe. afterwards, the group analyzes and discusses what happened in the classroom. the learning study is a hybrid form of lesson study and design experiment. it is a theory-based research undertaking whose important components include exploration of students’ ways of making sense of the object of learning before and after the lesson(s). a learning study usually comprises three cycles, each building on the conclusions of the previous. finally, a learning study is documented, frequently in publishable form. while lesson study is primarily an arrangement for in-service training of the participating teachers, learning study is primarily teachers’ research, the results of which are supposed to be widely shared with other teachers. the variation theory of learning has so far been the theoretical point of departure for the studies carried out. the model was originally developed right after the turn of the millennium in hong kong, and subsequently spread to other countries, notably to sweden. our estimate is that nearly 1000 such studies have been carried out by now (cf. lo, 2009). the main (quantitative) results of the studies published to date can be summarized as follows.  in nearly all of the studies, students’ results were better after the lesson(s) than before (lo, pong, & chik, 2005). (although this may appear self-evident, it is not. unfortunately, there are many school lessons in which students learn nothing, or at least not what the teacher had hoped they would.)  students with weaker learning prerequisites usually learn the most. hence, not only does the average rise, but the spread diminishes (lo et al., 2005).  in cases in which what the students had learned was observed not only immediately after the lesson but also on a later occasion, the results were often found to be better at the later time (thus indicating a content-specific ―learning to learn‖ effect) (holmqvist, gustavsson & wernberg, 2008).  results on national achievement tests increased for classes that had participated in several learning studies, an effect that in all likelihood was mediated by changes in teachers’ regular ways of teaching (maanula, 2011).  when the same object of learning is dealt with in a learning study and in a lesson study by groups of equally well-qualified teachers, the quality of learning turns out to be strikingly higher in the former (marton & pang, 2006, 2008; pang, 2010; pang & marton, 2003, 2005, 2007).  when the three cycles of a learning study are compared, the results from the third are usually better than those from the second, and those from the second are usually better than those from the first (lo, 2009). john elliot, one of the founders of the ―action research‖ movement in education, has evaluated two large-scale learning study projects carried out in hong kong. he concluded: ―the evaluation gathered convincing evidence of the positive impact of the process on teachers’ and students’ learning …. learning study is focused on realizing new kinds of pedagogical roles. from the evidence gathered in this evaluation it has enormous potential in this respect‖. (elliott, 2004) f. marton & m.f. pang 27 | f l r it seems, in other words, that the learning study approach has been something of a success story. what about our conjecture? has it been supported in learning study research? in our learning studies, every lesson was initially planned to be consistent with variation theory, and hence consistent with our conjecture. differences between cycles were related to differences between different interpretations of the same ideas. although this approach may be a good way to improve lessons, it is not really suitable for testing a theoretical conjecture. accordingly, we carried out a few studies using comparison groups, controlling for the assumed generally positive effects of the co-operative lesson study model. two groups of teachers, randomly selected for the two conditions (i.e., a learning study and lesson study condition), agreed on a particular object of learning. together, they explored their students’ understanding of that object, and planned a lesson on the basis of what they found and on their previous experience of teaching the same object of learning. one of the teachers then carried out the lesson, while the others observed. after the lesson, the group again explored students’ understanding of the object of learning, and the lesson was analyzed in light of the results. a researcher was present as a resource person during both the discussions and lessons. the only difference between the two conditions was that in the learning study group, the researcher introduced variation theory, which he did not do in the lesson study group. although he participated in the discussions in both groups, he tried to act in a reactive rather than active (initiating) manner. the focus of the studies was a comparison of students’ results under the two conditions in relation to a comparison of the patterns of variation and invariance brought about in those conditions (marton & pang, 2006, 2008; pang & marton, 2003, 2005, 2007). although the results showed dramatic differences to the advantage of the learning study (and hence the theory on which it is based, as these patterns were controlled by the teachers and by the students, of course these comparisons had to be post hoc. to sharpen the comparison of patterns of variation and invariance, the researcher must be able to ascertain exactly what patterns are being compared. in quasi-experimental comparisons, such as that described here, there are usually no consecutive cycles. even if a researcher tries to be as blind to the two conditions as possible, we can hardly claim that he or she has succeeded completely. in our case, the ―theory group‖ may have had an advantage beyond that originating in the theory itself. furthermore, the comparisons were made between the conditions in terms of the patterns of variation and invariance observed by the researcher, which means that they were post hoc, as noted, and hence the matter of empirical support for variation theory is not entirely straightforward. 1.3 there are no teaching experiments a fair number of studies have been published in recent years in which the outcomes of learning have been found to be systematically related to the patterns of variation and invariance inherent in the conditions of learning. the lived object of learning (learning outcome) in these studies has generally been found to be related to the enacted object of learning (teaching and classroom interaction) in ways entirely consistent with our conjecture. the outcomes of learning, and differences therein, can be made sense of in terms of the patterns of variation and invariance or the differences in these patterns that are inherent in the conditions of learning (see, for instance, fraser, allison, coombes, case, & linder, 2006; fraser & linder, 2009; linder, fraser & pang, 2006; marton & pang, 2006, 2008; pang, linder & fraser, 2006; pang & lo, 2012; pang & marton, 2003, 2005, 2013). if lessons are to provide stronger evidence, then they must be defined in advance, and their effects on learning must also be predicted in advance. kullberg (2010) carried out an interesting study in which she instructed teachers to teach particular objects of learning in terms of the critical features identified and patterns of variation and invariance employed in previous successful studies. the teachers were familiar with variation theory, according to which critical features and patterns of variation and invariance are powerful tools for communicating ways of handling a certain object of learning. even when kullberg’s (2010) results supported her expectations, however, there were several cases in which the enacted pattern of variation and invariance differed from that expected. although in some cases, the teacher had failed to open up dimensions f. marton & m.f. pang 28 | f l r of variation to make it possible for the students to discern certain critical features, in others, the students opened up dimensions of variation that they were not supposed to under their specific condition, but that were critical for learning. in such cases, the class was meant to serve as a control, and the unpredicted changes may have strengthened or weakened the results. 1.4 a critical experiment the only possible way to ensure that what is being compared is what we want it to be seems to be to build a pattern of variation and invariance into pedagogical tools: texts, tasks, examples, illustrations, problems, and the like. variation and invariance as far as the conditions of learning are concerned can then be defined in terms of the relationships between the constituent parts of the pedagogical tools that are used. a study of this kind was carried out by ki and marton (2003). they investigated how non-native speakers of cantonese could be helped to learn to attend to both the tonal and segmental (the sound but not the tone) aspects of cantonese words simultaneously to identify their meanings. cantonese is a tonal language in which the distinctions between six tones are of vital importance. the difficulty that speakers of non-tonal languages have when they try to learn it is not so much their inability to distinguish between two juxtaposed tones (stagray & downs, 1993) as their inability to link variation in pitch at the word level to variation in word meanings. variation in pitch exists in all languages, but its significance in non-tonal languages is at the sentencerather than word-level. learning to pay attention to differences in pitch at the word level as a cue to differences in word meanings requires reorganization of the attentional field. ki and marton (2003) employed a set of nine words grouped in two ways. in the first, they were grouped to constitute three triplets, each characterized by one tone (the same within each triplet, but differing from the other two). in the second, three segments were grouped to constitute three triplets, each characterized by one segment and three different tones (see figure 1). segmental1 segmental2 segmental3 tone1 word11 word12 word13 tone2 word21 word22 word23 tone3 word31 word32 word33 segmental1 segmental2 segmental3 tone1 word11 word12 word13 tone2 word21 word22 word23 tone3 word31 word32 word33 figure 1. three triplets characterized by one segment and three different tones (if read by column) and by one tone and three different segments (if read by rows). the participants’ task was to learn to identify the meaning of the word they heard by selecting its english equivalent. if we consider each triplet as a sub-task, then to be able to come up with the meaning of f. marton & m.f. pang 29 | f l r each word, a participant must be able to differentiate between the three words. if the three words in the sub-task have a tone in common, then the participant must learn to distinguish between the three different segments and link them to the three different meanings. if, instead, the three words in the sub-task have a segment in common, then the participant must learn to distinguish between the three different tones and link them to the three different meanings. hence, when the segments vary, you learn segments, and when the tones vary, you learn tones. the two ways of grouping the words can be seen as a comparison between two patterns of variation and invariance, that is, as induction and contrast from the point of view of tones. if we believe that language learners learn tones (i.e., differentiate between them) best if we offer them different examples of the same tone, then we group words into triplets, within which each has the same tone but a different segment, and ask learners to compare them. if we believe instead as our conjecture suggests that meaning (in this case, ―the meaning of tones‖) derives from variation, then we group the words into triplets, within which the tones differ but the segment is the same, and ask learners to compare them. in ki and marton’s (2003) study, the participants clearly learned to distinguish words more effectively by means of tones in the condition in which the tones were varied during the lesson and the segment remained the same, than in the condition in which the tone was invariant and the segments varied. the study thus demonstrated that learning is more effective under the contrast condition than under the induction condition, as predicted by our main conjecture (see also guo & pang, 2011). this was the first critical experiment in which it was put to the test. 1.5 another way of putting the conjecture to the test above, we have argued in agreement with fodor (1978, 1980) that induction is the most common means of trying to help others to acquire novel meanings, but it is certainly not the only one. in our own studies of the teaching and learning of economics (pang & marton, 2003, 2005; marton & pang, 2006, 2008), we found that teachers frequently used neither induction, nor contrast. they differed from the teachers using variation theory by not only varying the focused aspect but also varying the unfocused aspect. the teachers not using variation theory actually used more variation than the teachers using variation theory. the comparison between induction and contrast mentioned above and being the first critical test of the conjecture, can be illustrated in the following form: induction contrast focused aspect unfocused aspect focused aspect unfocused aspect i v v i in relation to the tone learning experiment described in the previous section, induction means that the participants learn one tone at a time in three different runs. in each run the tone is the same in every task, while the segments (the unfocused aspect) vary. in the case of contrast, there are three runs too, but in each run the segment is the same in every task, while the tone (the focused aspect) varies. this is one way of putting the conjecture to the test. contrast is aligned to the theory, induction is not. but what if the object of learning is cantonese words (and not only tones)? then we have two focused aspects (tonal and segmental). according to the theory, they should vary one at a time. but there is a third aspect, not new for the learners, hence unfocused. this is the meaning aspect of the words represented by pictures and english words in the experiment. it is not independent from the other two aspects: when one or both vary, the meaning varies too. the second way of putting the conjecture to the test is to compare the case when the two focused aspects vary one at a time, followed by both varying simultaneously (to bring the different aspects of the words together), with the case of having the focused aspects varying simultaneously f. marton & m.f. pang 30 | f l r from the beginning. the former pattern of variation and invariance is consistent with the conjecture, the latter is not. this is exactly the comparison that ki, ahlberg and marton (2006) carried out (see figure 2), demonstrating that the participants in the condition that was consistent with the main conjecture learned better than those in the condition that was not. moreover, the conjecture was built into the pedagogical tools they used in the study, a computer-administered program that afforded variation, invariance, and feedback to the participants. in the first experiment, one of the aspects, tone, was considered focused (what is to be learned) and the other, segment, was considered unfocused. in this second critical experiment in which the conjecture was put to the test, both aspects were considered focused (they had to be learned). tone segment meaning tone segment meaning v i v v v v i v v v v v v v v v v v figure 2. comparing patterns of variation and invariance consistent with (left) and not consistent with (right) the conjecture. discerning an aspect amounts to separating it from other aspects. two aspects can be distinguished from each other if one varies and the other is invariant. furthermore, if there are two focused aspects that learners are expected to learn to discern, then they should be varied one at a time, rather than simultaneously. if we want these learners to relate the two aspects, then we should vary them simultaneously, but only after they have been discerned. in the second experiment carried out by ki, ahlberg and marton (2006), there was a third aspect, meaning, that was assumed to be recognized by the learners (they were expected to make sense of the pictures representing the meanings). this aspect is a function of the other two aspects and cannot be kept invariant when any of the other aspects vary; nor does it interfere with the experience of variation in the other aspects, of which it is a function. the first critical experiment showed that letting the focused aspect (that which is to be learned) vary, while keeping the unfocused aspect (that which has already been learned) invariant yields better learning than keeping the focused aspect invariant and letting the unfocused aspect vary. the conjecture was thus supported. in the second critical experiment it was shown that in the case of two focused aspects varying one at a time, and then varying both simultaneously, yields better learning than letting both focused aspects vary from the beginning, even when an unfocused aspect, which is a function of the two focused aspects varies at the same time. the conjecture was supported again. it was put to the test in a third critical experiment in a study reported in the next section. in this case, in addition to the two focused aspects (demand and supply) and the unfocused aspect being a function of the two (price), there was an additional unfocused aspect involved (good) independent of the two focused aspects which according to the theory was supposed to remain invariant. again, two patterns of variation and invariance one consistent with, and one not consistent with our main conjecture were compared in terms of their effect on learning. and the conjecture was supported once again. this is the empirical contribution of the present paper. f. marton & m.f. pang 31 | f l r 2. the study 2.1 understanding pricing the point of departure for this study was an earlier study in which 10-year-old children were taught to discern price as a function of demand and supply (lo, lo-fu, chik & pang, 2005). that study, in turn, built on an earlier study of qualitatively different ways of understanding price and pricing (dahlgren, 1978). in both of these studies, with minor differences, it was found that most children and many adults see price as a function of the attributes of goods. for instance, if something is expensive, then it is because it is big, beautiful, tastes good, etc. price is thus seen as an attribute of the good in question, and linked to its other attributes, not as a function of market conditions (notably demand and supply), as economics tells us that it is. some see price as a function of demand only, and others as a function only of supply. for others still, price is a function of both demand and supply, or rather of the relationship between the two, which is roughly in accordance with the canonical conceptualization of price in classical, liberal economics. we use the expression ―learning to see something in a certain way‖ as synonymous with ―making a novel meaning your own‖ or ―appropriating a meaning.‖ all three refer to the capability to discern certain aspects of a phenomenon and focus on them simultaneously. what then are ―ways of seeing something‖? they are categories of description used to depict the various appearances of something or the different ways in which it is experienced (or its different meanings). the research specialization of phenomenography (marton, 1981; marton & booth, 1997; marton & pang, 2008) is the study of categories of description depicting appearances, experiences, and meanings. it posits that if a learner exhibits a certain way of seeing something, then this does not imply that he or she has that way of seeing (as a mental representation, for instance). what then does it imply? it implies that he or she is seeing or has seen a particular phenomenon in a particular way under particular circumstances. further, the fact that he or she is so seeing implies that he or she is able or has been able to see that particular phenomenon in that particular way under the given particular circumstances. accordingly, what we might wish to explore is the extent to which the same person can see the same phenomenon in the same way under different circumstances. if he or she can, then this could be interpreted as demonstrating that he or she has separated the particular way of seeing this particular phenomenon from the particular circumstances. becoming an ―expert‖ frequently amounts to being able to see particular phenomena in particular ways under widely varying circumstances (cf. chi, feltovich, & glaser, 1981; goodwin, 1994; marton & booth, 1997, p. x; sandberg, 1994). hence, phenomenography does not tell you what individuals’ ways of seeing something are. it tells you how their ways of seeing something vary (between people under the same circumstances and/or within people under different circumstances). the different categories of description together constitute the outcome space (of how the particular phenomenon might be experienced). as previously mentioned, studies have established four categories of description that together constitute the outcome space of the experience of price. 2.2 making it possible to learn to see price in a more powerful way are the different ways of seeing price equally powerful? we do not believe that we can always or even most of the time find a universal ordering of how valid and powerful different ways of seeing the same thing are. in a planned economy, and according to marxist economics, for example, price is not a function of demand and supply. however, we can delimit a set of contexts and settle for ordering the different options within that set. we could thus argue that it is better to enable learners to see something in an additional way that we believe to be powerful in certain contexts, that is than not doing so. accordingly, we may try to help learners to see something in a new way, that is, in a way that they have previously been unable to. although we can certainly try, we can never be certain of success. at best, we can ascertain that this new way of seeing might have been instilled, that is, that under the conditions f. marton & m.f. pang 32 | f l r given, it is possible that the learners learned to discern certain critical features, which is exactly what lo et al. (2005) did in five primary school classes (grade 4) in hong kong in the context of a learning study. the aim of each lesson in this study was the same: to enable the students to see price as a function of demand and supply in novel situations. after the lesson, a novel question was used to probe their way of seeing price. 2.3 the enacted object of learning a double-lesson was used to help the students to learn to discern demand and supply, and the relationship between the two, as determinants of price. during the lesson, the students formed groups and participated in an auction of four items (a mechanical dinosaur, a doll, a dinosaur card, and a stationery set). the auction was repeated several times, with variations. to encourage the students to focus on and discern the critical aspects of supply and demand separately, changes were made in supply (by varying the number of items available) while demand was kept invariant, and then changes were made in demand (by varying the purchasing power through changes in the auction money afforded the groups) while supply was kept invariant. after each auction, they were asked what would be a reasonable price for a new, limited-edition mechanical dinosaur if people had more money to spend. after the groups had written their answers on a worksheet, the teacher engaged the class in a discussion of the case of supply going down and demand going up. did the teachers who took part in this study achieve their goal? if so, to what extent did they do so? as can be seen from table 1, their attempts were not especially successful, with the possible exception of class 4b (see the frequencies for category d, considered the canonical conception here). table 1 distribution of conceptions in preand post-tests in learning study carried out by lo et al. (2005) class 4a class 4b class 4c class 4d class 4e conceptions of price pretest posttest pretest posttest pretest posttest pretest posttest pretest posttest a. attributes of the good 6.1% 15.2% 7.7% 0.0% 0.0% 9.7% 17.9% 0.0% 6.5% 9.7% b. demand 39.4% 45.5% 64.0% 28.2% 77.5% 71.0% 50.0% 78.6% 51.5% 48.4% c. supply 0.0% 6.0% 2.6% 7.7% 3.2% 3.2% 10.7% 3.6% 19.4% 6.5% d. demand and supply 3.0% 9.1% 10.3% 61.5% 3.2% 12.9% 7.1% 14.2% 9.7% 22.6% e. other noneconomic reasons 3.0% 0.0% 7.7% 2.6% 0.0% 0.0% 3.6% 0.0% 3.2% 0.0% unclassified 48.5% 24.2% 7.7% 0.0% 16.1% 3.2% 10.7% 3.6% 9.7% 12.8% rather than ask whether (and why or why not) seeing price in terms of demand and supply is too difficult for 10-year-old children, we are more eager to understand the striking difference in results between f. marton & m.f. pang 33 | f l r class 4b and the other classes. (as is shown in table 1, while the frequency of the target conception (d) increased after the lesson from 10.3 to 61.5% in class b, it increased from about 6 to about 15% in the other classes). did something happen in this class that did not in the others? or did something happen in all of the classes except 4b? prompted by the same curiosity, lo et al. (2005) did indeed come up with an interpretation for the discrepancy in their results: the necessary conditions for discerning a simultaneous variation in demand and supply were present only in class 4b, which was the only class in which the unfocused aspect (the good, i.e. the item for auction) was invariant throughout the entire sequence of variation and invariance in the focused aspects (demand and supply). differences of this kind (the focused aspect varying and the unfocused aspect remaining invariant versus both aspects varying) have also been found in two other studies, and in both cases were linked to rather dramatic differences in what the participants had learned (i.e., according to the outcome measures) (marton & pang, 2006; pang & marton, 2003). the conjecture that we want to put to the test here has two component parts: what is expected to vary in sequence (the focused aspects) and what is expected to remain invariant (the unfocused aspect) throughout. in the study reported here, we wanted to compare two conditions: one consistent with the second component part (the unfocused aspect remaining invariant throughout) and one not consistent with it. could we replicate the findings of the aforementioned study (lo et al., 2005), which served as our point of departure, with the same difference built into pedagogical tools? figure 3 shows the comparison carried out. demand supply meaning good demand supply meaning good v i v i v i v v i v v i i v v v v v v i v v v v figure 3. comparing patterns of variation and invariance, consistent (left) and not consistent (right) with the conjecture. 2.4 design of the study to reduce the number of factors that could affect the outcome, we tried to build the pattern of variation and invariance (which we assumed to be necessary) into the task structure of the learning resources in such a way that the entire experiment would be an interaction between students and the auction game tool: the computer. students were invited to attempt to achieve the object of learning by using two different computerized learning resources during an independent learning session that lasted approximately one and a half hours and was held in the multi-media learning center of the participating school. in line with lo et al. (2005) study, in both learning resources, the economic principle to be dealt with was the determination of the market price through the interaction of supply and demand. an auction game was used to embody the variation in the dimensions of supply and demand. to test whether it is crucial to keep the auction item in question invariant, so as to enable students to focus on and discern the critical aspect of the interaction between supply and demand more readily and effectively, the two learning resources were identical in all respects but: one resource made use of the same product (i.e., boxes of candy) throughout the auction game, whereas the other featured different products within and across each round. seventy-eight grade 4 students from four classes of one school in hong kong participated in the study. within each class, students were randomly divided into two groups, with each given one of the two learning resources. to minimize the teacher effect, learning took place in an autonomous manner, with the students involved playing the computerized auction game on their own, although the researcher gave a fiveminute summary at the end of the session to remind the students of the key learning points. (note that it was impossible for the researcher to know under which condition each student was working. the only difference between the two conditions was that students in the same multi-media learning center received one or the other of the two versions of the learning resource, with the distribution of the two completely randomized.) f. marton & m.f. pang 34 | f l r to obtain students’ existing understanding of the object of learning before they engaged with the learning resources and to form a baseline for comparison of the learning outcomes of the two groups, a pretest was administered to all students. then, immediately after the independent learning session, they were required to complete a post-test to allow evaluation of their mastery of the object of learning. in both tests, the students were asked to consider a problem relating to a real-life scenario embodying the principle in question, i.e., the interaction of supply and demand in determining the market price of a good. they were also asked to elaborate upon the factors they had considered in setting that price. the questions in the preand post-tests were essentially identical, except that the product in question varied. mirroring lo et al. (2005) study, a hot dog and a box of biscuits were used. students who were asked a question about the hot dog in the pre-test were asked about a box of biscuits in the post-test, and vice versa. the preand post-test questions were as follows: have you ever tried the hot dogs (biscuits) sold in the school shop? do you know how much they cost? maybe you know or you don’t know. anyway, just for your information, hot dogs are (a box of biscuits is) now sold at hk$5. suppose that you are the new owner of the shop. what price would you set for a hot dog (box of biscuits)? would you set the current price, or a different price? what would you consider when you set the price? the students’ answers were analyzed and described in terms of the aforementioned set of four categories of understanding. 2.5 the learning resources to build a relevance structure (marton & booth, 1997, p. 143) that would enable students to appropriate the object of learning, they were given the task of bidding on goods for an upcoming new year’s celebration through the computerized auction game. in the first round, students were introduced to the basic rules and operation of the game. each student was given hk$400 in auction money and asked to bid for and thus try to obtain as many items as possible from the nine being auctioned, which were displayed on screen with their base prices shown. each round of the auction came to an end after three minutes or once the student had used up all of his or her money, whichever came first. the average prices of the goods auctioned were then calculated and shown to the student so that he or she could associate possible changes in those prices with changes in the conditions of each round of the auction, such as the amount of auction money provided, the number of goods to be auctioned, or both. as previously noted, the only difference between the two learning resources was that for the ―different goods‖ group the nine items, which included different kinds of snacks such as potato chips, chocolate bars, biscuits, and so on, differed both within each round and between rounds, whereas for the ―same goods‖ group the nine items were all the same, i.e., every item was a box of candy. in the second round (see figures 4 and 5), to bring students’ focal awareness to bear upon the dimension of demand, demand was deliberately varied (by varying students’ purchasing power by changing the amount of auction money they were given) while the supply of goods was kept invariant. each student’s auction money was cut by hk$200, thus diminishing their purchasing power and demand for goods. the supply of goods for auction, however, remained invariant, with the number of items kept at nine. everything was identical for both learning resources except that the nine items for auction remained invariant in the ―same goods‖ design (the same nine boxes of candy as in the first round, whereas the type of goods varied in the ―different goods‖ design, changing from the nine kinds of snacks in the first round to nine kinds of soft drink in the second.) f. marton & m.f. pang 35 | f l r day 2 group a record of the average price of items for auction day 1 day 2 day 3 day 4 amount of auction money hk$400 hk$200 number of items auctioned average price of items for auction now the shop has closed. are you happy with the items that you have obtained? reflect on the auctions on days 1 and 2 and complete the following task. task: compare today’s average item price with yesterday’s. what have you found? figure 4. same goods design (round 2). day 2 group b record of the average price of items for auction day 1 day 2 day 3 day 4 amount of auction money hk$400 hk$200 number of items auctioned average price of items for auction now the shop has closed. are you happy with the items that you have obtained? reflect on the auctions on days 1 and 2 and complete the following task. task: compare today’s average item price with yesterday’s. what have you found? figure 5. different goods design (round 2). f. marton & m.f. pang 36 | f l r in the third round, to help students to shift their focal awareness to the dimension of supply, supply was deliberately varied while demand was kept invariant. the number of items for auction was reduced from nine to seven, whereas the amount of auction money remained the same (hk$200). however, the only but critical difference between the two learning resources was that all seven items in the ―same goods‖ design remained boxes of candy, whereas the seven items used in the ―different goods‖ design now differed from those in the two previous rounds, with participants being asked to consider different kinds of balls in this round. unlike the classroom study carried out by lo et al. (2005), we introduced a fourth auction round in which variation was introduced in both the demand and supply of goods in a simultaneous manner. our purpose was to help students to focus on the dimensions of both in determining the market price of a good. to this end, the auction money given to students was increased from hk$200 to hk$400, and the number of items to be auctioned was reduced from seven to six. this round thus involved a simultaneous variation in the supply of goods and variation in purchasing power (demand for the goods), the aim of which was to enable students to discern the critical aspects of experiencing price and pricing. as before, the only difference between the two learning resources was that the six items in the ―same goods‖ design remained the same, whereas a new set of items (six different kinds of decorations) was introduced in the ―different goods‖ design. lastly, similar to the procedure in the earlier study lo et al. (2005), students were asked a question (for instructional purposes) about what would happen to the price if the supply were increased and purchasing power decreased. in the current study, they were invited to predict the direction of change in the market price, that is, whether the price would go up or down, if the amount of auction money was decreased from hk$400 to hk$100 while the number of items to be auctioned increased from six to 11. as noted, the learning session concluded with a five-minute summary delivered by the researcher to remind students of the key learning points in the computerized learning resources. he simply read the following powerpoint slides to the two groups of students at the same time. 1. (slide 1) ―compare the auction game on days one and two. as the auction money given to you on day two was less than that on day one, your income decreased. when your income decreased, your purchasing power also decreased. this made your demand for goods decrease. as the supply of goods on day two was the same as that on day one, the average price of goods was lower.‖ 2. (slide 2) ―compare the auction game on days two and three. as the auction money given to you on day three was the same as that on day two, your income and purchasing power remained unchanged. your demand for goods also remained unchanged. as the cost of production, such as the prices of raw materials, electricity, and labor increased, the supply of goods decreased. as a result, the average price of goods on day three was higher than that on day two.‖ 3. (slide 3) ―compare the auction game on days three and four. as the auction money given to you was more than that that on day three, your income and purchasing power increased, and your demand for goods also increased. at the same time, the increase in the cost of production made the supply of goods decrease. as demand increased and supply decreased, the average price of goods on day four was higher than that on day three.‖ 4. (slide 4) ―the price of a good is determined by its supply and demand. the supply of a good is affected by its cost of production, such as the prices of raw materials, electricity, and labor, whereas the demand for a good is affected by people’s income and purchasing power. when businesses set the price of a good, they need to consider the factors affecting supply and demand at the same time.‖ f. marton & m.f. pang 37 | f l r 3. results and findings the results presented in table 2 show that students who belonged to the group using the learning resource with the ―same goods design‖ outperformed their counterparts using the learning resource with the ―different goods design‖ in the post-test, in which statistically significant difference was observed between the two groups ( 2 = 10.36, p = 0.03 (< 0.05); effect size = 0.32). (note, in particular, the relative frequencies for the target understanding category d in the post-tests for the two conditions.) table 2 distribution of conceptions, preand post-test conception of price ―same goods design‖ group (40 students) ―different goods design‖ group (38 students) occurrence percentage occurrence percentage pre-test post-test pre-test post-test pre-test post-test pre-test post-test a. attributes of the good 10 1 25.0% 2.5% 4 2 10.5% 5.3% b. demand 14 11 35.0% 27.5% 14 15 36.8% 39.5% c. supply 5 9 12.5% 22.5% 2 5 5.3% 13.2% d. demand and supply 5 16 12.5% 40.0% 5 6 13.2% 15.8% e. other noneconomic reasons 6 3 15.0% 7.5% 13 10 34.2% 26.3% total 40 40 100% 100% 38 38 100% 100% 2 = 7.92 (df = 4) (p = 0.09, i.e., p > 0.05) – pre-test 2 = 10.36 (df = 4) (p = 0.03, i.e., p < 0.05) – post-test we can see in table 2 that while the frequency of the target conception (d) increased after the lesson from 5 to 16 (of 40) under conditions consistent with the conjecture, it increased from 5 to 6 (of 40) under conditions not consistent with the conjecture. 4. conclusions only in a restricted sense was this study a replication of lo et al. (2005) investigation. we wanted to find out if invariance or variation in an unfocused aspect can really have such a strong impact on the learning of the focused aspects as was interpreted to be the case in the previous study. the question could be answered in the affirmative and the conjecture was thus supported. it should be noted that in both the original and follow-up studies, variation was restricted. when the supply was invariant, demand went down (instead of going up in one case and down in another), and when demand was invariant, the supply went down (instead of going up in one case and down in another). in the f. marton & m.f. pang 38 | f l r last round, only one of the four combinations of demand (up/down) and supply (up/down) was realized. we decided not to include all four combinations as it would have made the task too difficult for such young participants. in all of the circumstances considered, it is of course possible that the results would have differed had the students been exposed to more of the possible differences among the patterns of variation and invariance. quite a few of the students in the comparison group managed to learn to discern the critical features of pricing, even though the good was not invariant. in general, even if there is variation in several dimensions, learners may be able to block out all dimensions but one, that on which they happen to focus. there is an interesting twist concerning how this question appeared in the experiment, however. as noted, in the target group, a number of items of the same type of good were offered in each round at the same base price. in the comparison group, the same number of items as in the target group were offered at the same base price. further, whereas the target group considered the same type of good both within and between rounds, the comparison group considered different goods in each round. the difference between the conditions was illusory, however: all the relevant factors (the number of items available, the amount of money participants had, and the base price of goods) were exactly the same. the only element that differed was the irrelevant labels placed on the goods. the only thing those in the comparison group had to do was to separate what was relevant for their decisions from what was not, and bracket the latter. if all of them had done so, then the conditions for the two groups would have been the same. as we can see from the differences in outcome, however, this was not the case. our finding that the comparison group was affected by the irrelevant differences in the item labels implies that quite a few learners in that group failed to separate relevant information from irrelevant information, and therefore failed to see the former. the main contribution of this study is the support it provides for our conjecture: if both the focused and unfocused aspects of the object of learning vary, then it is more difficult to discern the focused aspects and relate them to one another than if the unfocused aspect remains invariant while the focused aspects vary. we found this to be true in the current study, even though the unfocused aspect was completely redundant. however, the conjecture also addresses the question of how we can acquire new meanings (or how we can learn to see certain things in certain ways). as mentioned earlier, fodor (1978, 1980), and others, claim that there is no answer to this question and, in fact, there cannot be any. meanings are innate. however, we argue that regardless of whether meanings (concepts) are innate, or of the sense in which they are (or are not) innate, we have to learn to discern them as aspects of the world around us, and for this to happen, there are necessary conditions. these necessary conditions are specific to particular meanings and to learners’ particular experiential history. they can be formulated in terms of patterns of variation and invariance among instances that do and do not have that particular meaning. our conjecture is thus very straightforward, as is the way in which it can be put to the test. we simply have to create the necessary conditions in one case and ensure that they are absent in another, as pang and marton (2003) did in their aforementioned study. then, we can compare the two cases and determine whether, as expected, all participants in the first case learn the target meaning, whereas none of those in the second do. if these are indeed the results, then the conjecture is strengthened. obviously, this is not what happened. even if we can demonstrate that contrast is more powerful than induction as far as the learning of new meanings is concerned, we cannot demonstrate that new meanings cannot be learned through induction. after all, some learners seem to learn in that condition too, and certainly not all learners will learn even if all possible steps are taken to make it possible for them to do so. the relationship between what is learned, on the one hand, and the conditions of learning, on the other, is stochastic rather than deterministic. but why is this so? returning to the experiment reported in this paper, beyond the fact that the target principle (price as a function of the relationship between demand and supply) was made explicit to both groups, there is a more general answer to the foregoing question. our conjecture concerns the pattern of variation and invariance as experienced by the learner, whereas a pattern of variation and invariance that can be controlled by the researcher refers to the patterns as seen by the researcher. what might the relationship between the two look like? one condition of experiencing variation is that there is variation to be experienced. making sure that f. marton & m.f. pang 39 | f l r this condition is met is the first step toward making learning possible (which in our view is what teaching is all about). however, variation can also be experienced because of previous experiences. experienced variation is thus not necessarily the experience of what is present in the learning situation as seen by the observer. on the other hand, even if there is variation, it is not necessarily experienced by all learners. in conclusion, when comparing two randomly selected groups, we would expect more learners to experience variation if it is present than if it is not. however, experiencing variation not only concerns the variation to be experienced in a relevant dimension; it also presupposes invariance in other dimensions. in other words, variation can be experienced only against a background of invariance. in this sense, experienced variation is a function of invariance, and, as previously stated, experienced variation is also a function of variation. our intention with the present study was to illustrate that learning (in the sense of the discernment of the necessary features of a phenomenon) is a function of experienced variation (by the learner), which is a function of both variation and invariance (as seen by the observer). this we did, with a focus on the latter (invariance). we have thus shown that introducing redundant information (different goods) that is correlated with a variation in critical aspects (a change in demand and supply) significantly reduces the likelihood of learners being able to discern the critical features of the object of learning. a seemingly subtle difference between two conditions, both representing 90 minutes of pedagogical effort, is proved to play a key role in what the students managed to learn. keypoints this paper addresses one of the oldest unsolved mysteries of learning: how do we make novel meanings our own? the answer suggested is: by discerning, separating and bringing together the critical aspects of what we learn about. a critical aspect can be discerned and separated through the experience of variation in that aspect against the background of invariance in other respects. we have put our conjecture to the test by embedding it in a pedagogical tool. the conjecture was supported. acknowledgments the research reported here was financially supported by the swedish research council. we also want to thank the two reviewers of our paper for their excellent input. references chi, m. t. h., feltovich, p. j., & glaser, r. (1981). categorization and representation of physics problems by experts and novices. cognitive science, 5(2), 121-152. doi: 10.1207/s15516709cog0502_2 dahlgren, l. o. (1978). effects of university education on the conception of reality. reports from the institute of education, university of goteborg. gothenburg: institute of education, university of gothenburg. elliott, j. (2004). the independent evaluation of the pips project. hong kong: hong kong institute of education. fodor, j. a. (1978). the language of thought. hassocks, sussex: harvester. f. marton & m.f. pang 40 | f l r fodor, j. a. (1980). fixation of belief and concept acquisition. in m. piattelli-palmarini (ed.), language and learning: the debate between jean piaget and noam chomsky (pp. 142-162). london: routledge & kegan paul. fraser, d., allison, s., coombes, h., case, j., & linder, c. (2006). using variation to enhance learning in engineering. the international journal of engineering education, 22(1), 102-108. fraser, d., & linder, c. (2009). teaching in higher education through the use of variation: examples from distillation, physics and process dynamics. european journal of engineering education, 34(4), 369381. goodwin, c. (1994). professional vision. american anthropologist, 96(3), 606-633. doi: 10.1525/aa.1994.96.3.02a00100 guo, j.p., & pang, m. f. (2011). learning a mathematical concept from comparing examples: the importance of variation and prior knowledge. european journal of psychology of education, 26(4), 495-525. doi: 10.1007/s10212-011-0060-y holmqvist, m., gustavsson, l., & wernberg, a. (2008). variation theory: an organizing principle to guide design research in education. in a. e. kelly, r. a. lesh & j. y. baek (eds.), handbook of design research methods in education: innovations in science, technology, engineering, and mathematics learning and teaching (pp. 111-130). new york: routledge. ki, w. w., ahlberg, k., & marton, f. (2006). computer-assisted perceptual learning of cantonese tones. paper presented at the 14th international conference on computers in education, beijing, china: asia-pacific society for computers in education (apsce). ki, w. w., & marton, f. (2003). learning cantonese tones. paper presented at the earli biennial conference 2003, padova, italy. kullberg, a. (2010). what is taught and what is learned: professional insights gained and shared by teachers of mathematics. doctoral dissertation, university of gothenburg, acta universitatis gothoburgensis, göteborg. linder, c., fraser, d., & pang, m. f. (2006). using a variation approach to enhance physics learning in a college classroom. the physics teacher, 44(9). 589-592. lo, m. l. (2009). building a teacher learning network for developing the ability to teach for learning. paper presented at the 13th biennal conference of earli, amsterdam, the netherlands. lo, m. l., lo-fu, y. w., chik, p. m. p., & pang, m. f. (2005). two learning studies. in m. l. lo, w. y. pong & p. m. p. chik (eds.), for each and everyone: catering for individual differences through learning studies (pp. 75-116). hong kong: hong kong university press. lo, m. l., pong, w. y., & chik, p. m. p. (eds.). (2005). for each and everyone: catering for individual differences through learning studies. hong kong: hong kong university press. maanula, t. (2011). resultat från nationella prov i matematik m m. available from tuula.maunula@telia.com (unpublished manuscript). marton, f. (1981). phenomenography—describing conceptions of the world around us. instructional science, 10(2), 177-200. doi: 10.1007/bf00132516 marton, f. (forthcoming). necessary conditions of learning. new york: routledge. marton, f., & booth, s. (1997). learning and awareness. mahwah, n.j.: l. erlbaum associates. marton, f., & pang, m. f. (2006). on some necessary conditions of learning. journal of the learning sciences, 15(2), 193-220. doi: 10.1207/s15327809jls1502_2 marton, f., & pang, m. f. (2008). the idea of phenomenography and the pedagogy for conceptual change. in s. vosniadou (ed.), international handbook of research on conceptual change (pp. 533-559). london: routledge. marton, f., & tsui, a. b. m. (2004). classroom discourse and the space of learning. mahwah, nj: lawrence erlbaum associates. michalski, r. (1983). a theory and methodology of inductive learning. artificial intelligence, 20, 111-161. pang, m. f. (2010). boosting financial literacy: benefits from learning study. instructional science, 38(6), 659-677. doi: 10.1007/s11251-009-9094-9 f. marton & m.f. pang 41 | f l r pang, m. f., linder, c., & fraser, d. (2006). beyond lesson studies and design experiments: using theoretical tools in practice and finding out how they work. international review of economics education, 5(1), 28-45. pang, m. f., & lo, m. l. (2012). learning study: helping teachers to use theory, develop professionally, and produce new knowledge to be shared. instructional science, 40(3), 589–606, doi: 10.1007/s11251011-9191-4 pang, m. f., & marton, f. (2003). beyond "lesson study'': comparing two ways of facilitating the grasp of some economic concepts. instructional science, 31(3), 175-194. doi: 10.1023/a:1023280619632 pang, m. f., & marton, f. (2005). learning theory as teaching resource: enhancing students’ understanding of economic concepts. instructional science, 33(2), 159-191. doi: 10.1007/s11251-005-2811-0 pang, m.f., & marton, f. (2007).the paradox of pedagogy. the relative contribution of teachers and learners to learning. iskolakultura, 1(1), 1-29. pang, m. f., & marton, f. (2013). interaction between the learners’ initial grasp of the object of learning and the learning resource afforded. instructional science, doi: 10.1007/s11251-013-9272-7 sandberg, j. (1994). human competence at work: an interpretative approach. göteborg, sweden: bas. stagray, j. r., & downs, d. (1993). differential sensitivity for frequency among speakers of a tone and a non-tone language. journal of chinese linguistics, 21(1), 143-163. stigler, j. w., & hiebert, j. (1999). the teaching gap: best ideas from the world's teachers for improving education in the classroom. new york: free press. vlach, h.a., sandhofer, c.m. & kornell, n. (2008). the spacing effect in children's memory and category induction. cognition, 109, 163-167. frontline learning research 2 (2013) 111 issn 2295-3159 corresponding author: sandra van aalderen-smeets, po box 217, 7500 ae, enschede, the netherlands, sandra.vanaalderen@utwente.nl http://dx.doi.org/10.14786/flr.v1i2.27 3 | f l r investigating and stimulating primary teachers’ attitudes towards science: summary of a large-scale research project juliette walma van der molen a , sandra van aalderen-smeets a * a university of twente, the netherlands * both authors contributed equally article received 13 may 2013 / revised 19 september 2013 / accepted 19 september 2013 / available online 20 december 2013 abstract attention to the attitudes of primary teachers towards science is of fundamental importance to research on primary science education. the current article describes a large-scale research project that aimed to overcome three main shortcomings in attitude research, i.e. lack of a strong theoretical concept of attitude, methodological flaws in attitude research, and ineffective interventions. the research project included (a) the development of a new theoretical framework for teachers’ attitudes towards (teaching) science, (b) a new validated survey instrument (the das) to measure the different underlying components of primary teachers’ attitudes toward teaching science, and (c) an in-service professional development training course based on the previously developed theoretical framework. the framework of attitude consists of three dimensions: cognitive beliefs, affect, and perceived control, each consisting of several subcomponents. by means of the survey instrument we investigated the effects of the attitude focussed training course. the course aimed to improve attitude by creating awareness about teachers’ own attitudes, stimulating their scientific attitudes and curiosity, and training inquiry and thinking skills. the course refrained from providing recipe-like example lessons, materials, or methods. using a pre-post, experimental-control design we showed that the course significantly improved the affective and perceived control dimension of attitude. teachers enjoyed teaching science more, showed increased self-efficacy, and felt less dependent on external factors. this project shows that genuine attitude improvements of primary teachers can be accomplished by attitude focussed professional development. keywords: science education; attitude towards science; professional development; inquiry based learning. j. walma van der molen & s. van aalderen-smeets 4 | f l r 1. introduction the study on attitude towards science has received considerable attention over the last decades (osborne, simon, & collins, 2003; osborne & dillon, 2008). in our society, we are increasingly dependent on science and technology in all kinds of ways. despite this, a large section of the population has little scientific or technical knowledge and attitudes towards science and technology are not very positive. although this lack of interest often only really manifests itself when young people make their choice of subjects at secondary school, most pupils have already formed stereotypical images of and negative attitudes about science and technology before the age of 14 (osborne & dillon, 2008; tai, qi liu, maltese, & fan, 2006; turner & ireson, 2010). international research (e.g., jarvis & pell, 2004) shows that this negative image of science subjects is also found among primary school teachers. many primary teachers feel insufficiently capable of providing education in the field of science. they find it difficult to deal with pupils’ questions in this area and prefer to fall back on standard textbooks or highly structured materials or exercises. when this type of practice is the norm, it is no wonder that pupils’ attitudes with regard to science and technology are difficult to change for the better. primary schools and their teachers therefore play a crucial role in determining the attitudes and images of students towards science and improving primary teachers’ attitudes towards science is one of the major challenges in today’s science education (haney, czerniak, & lumpe, 1996; osborne, simon & collins, 2003; osborne & dillon, 2008). research has shown that (pre-service) primary teachers’ scientific literacy is low, that their attitudes towards science are mostly negative, and that primary teachers share a number of characteristics that impede the stimulation of science learning and of positive attitudes towards science among their pupils (harlen & holroyd, 1997; jarvis & pell, 2004; tosun, 2000; yates & goodrum, 1990). professional development should therefore pay explicit attention to improving the attitude of (preservice) primary teachers towards science (haney, czerniak, & lumpe, 1996). however, although primary teachers’ attitudes toward science have been investigated widely, scientific progress in this field has been slow. in our view, there are three important and in part related reasons for this poor development. first, until recently the theoretical conceptualization of the construct of primary teachers’ attitudes towards science was poorly articulated, both in research and in educational change projects (barmby, kind, & jones, 2008; bennett et al., 2001; coulson, 1992; osborne et al., 2003; pajares, 1992). many studies provided incomplete definitions (or no definition at all) for the construct of attitude, failed to explicate the components of attitude that they measured, or did not distinguish between attitudes towards science and other related concepts (e.g., opinions or motivation). a second reason for the slow scientific progress in research on primary teachers’ attitudes towards science is the lack of reliable and valid attitude measuring instruments that accommodate to necessary theoretical and statistical standards (gardner, 1995; reid, 2006). a recent review of the literature points at important flaws in the methodology of a majority of studies, such as weak psychometric properties and failure to pilot-test, validate, and evaluate the instrument according to current psychometric standards (van aalderen-smeets & walma van der molen, 2012). a third reason for slow progress in primary teachers’ attitude development may be found in the interventions that are directed at teachers’ professional development. most professional development projects that aim to improve science education in primary school focus on improving classroom didactics and provide a collection of standardized, recipe-like science lessons. although this might improve the knowledge of teachers regarding science content or pedagogical content knowledge, it does not automatically lead to improvements in their attitudes toward science. j. walma van der molen & s. van aalderen-smeets 5 | f l r 2. recent research to remedy these shortcomings in research and in the professional development of primary teachers’ attitudes towards science, we established a large-scale project over the past three years that included (a) the development of a new theoretical framework for teachers’ attitudes towards (teaching) science, (b) a new validated survey instrument to measure the different underlying components of primary teachers’ attitudes towards teaching science, and (c) an in-service professional development training course that was based on the previously developed theoretical framework. the project was based on the contention that only when teachers possess a positive attitude towards teaching science and towards inquiry-based learning, they will be motivated and able to seek and use science content in their lessons, to use inquiry-based learning in class, and to affect pupils’ scientific attitudes and their attitudes towards science in a positive manner. the present article presents an overview of the results of this integrated project. 2.1 theoretical framework the development of a new attitude framework implied disentangling the construct of primary teachers’ attitudes towards science. as described elaborately in the theoretical article that resulted from this project (van aalderen-smeets & walma van der molen, 2012), we aimed to explicate and structure the range of underlying components or dimensions of primary teachers’ attitudes towards science. the framework was based on an extensive review of previously used concept definitions of the construct of primary teachers’ attitudes towards science and we related these components to general psychological attitude theories, such as the tripartite model of attitudes (e.g., eagly & chaiken, 1993) and the theory of planned behaviour (e.g., ajzen & fishbein, 1980). this resulted in a framework of attitude consisting of the following three main dimensions: cognitive beliefs, affect, and perceived control. cognitive beliefs refer to teachers’ beliefs and opinions about (a) the relevance of science and science education, (b) beliefs about the relative difficulty of teaching science, and (c) gender stereotypical beliefs regarding science and science teaching. the second dimension of affect contains the independent subcomponents of (a) enjoying (teaching) science and (b) anxiety related to (teaching) science. the third dimension, perceived control, refers to the amount of control teachers perceive to have over (teaching) science and it consists of (a) self-efficacy (an internal sense of control, such as the perceived capacity to teach science) and (b) perceived dependency on context factors (beliefs about the extent to which a teacher is dependent on external factors to teach science, such as the availability of teaching-methods or materials, enough time, or other resources). the outcomes of our review of concepts suggested that, in addition to internal beliefs and feelings associated with self-efficacy, the beliefs and feelings that teachers have about external (i.e., contextual) factors are closely related to teachers’ sense of being in control. in our view, the perception of teachers regarding their dependency on context factors (e.g., their belief that they can teach science only if their school ensures the availability of the proper materials and sufficient preparation time) is an indispensable component of a complete theoretical framework of primary teachers’ attitudes toward science. the development of the theoretical framework provided a new theoretical basis for measuring primary teachers’ attitudes towards science and for interventions aiming to improve their attitude, two issues that were pursued in the studies described below. 2.2 validated survey instrument based on the theoretical exercise described above, we developed a new measurement instrument: the dimensions of attitudes towards science questionnaire (das). after construction of the first version of the das, we investigated its validity and reliability by means of a qualitative in-depth focus group study and a quantitative survey study (asma, walma van der molen, & van aalderen-smeets, 2011; van aalderensmeets & walma van der molen, 2013a). using the theoretical framework as a basis for the development of a new attitude instrument ensured that the complete range of relevant attitude dimensions and subcomponents was incorporated in the instrument. in addition to the subscales that correspond to the components of the theoretical framework of attitude, scales measuring teachers’ views on science and their intended science teaching behaviour were included in the questionnaire. j. walma van der molen & s. van aalderen-smeets 6 | f l r the pilot-tested das questionnaire was distributed digitally to in-service and pre-service teachers. a total of 556 respondents returned a complete questionnaire (80% female, mean age 31 years). the das instrument was evaluated at multiple levels; i.e. the validity of the overall structure of the instrument was investigated by confirmatory factor analysis, the internal consistency of the subscales was determined by cronbach’s alpha coefficient, and to assess the discriminating ability of each item, we looked at the standard deviations of each item and their response-range. our results supported the validity, reliability, and discriminating ability of the das instrument. the resulting factor structure corresponded to the underlying theoretical model. the obtained seven-factor solution confirmed the hypothesis that the das questionnaire is measuring the seven underlying sub-components of the attitude framework. furthermore, the results of the internal consistency analyses showed high internal consistency in all seven subscales. also, regression analyses showed that scores on the subscales of affect and perceived control were predictive of the scores on intended behaviour, indicating predictive validity. finally, all items in the revised das instrument showed large response variation, indicating a strong ability to discriminate between respondents displaying different beliefs, feelings, and thoughts toward teaching science. these results show that the das instrument is a valid, reliable, and comprehensive survey tool, which is able to measure a complex and difficult motivational concept (for a complete description of the instrument, see: van aalderen-smeets & walma van der molen, 2013a). the das instrument thus proves to be a promising instrument within the field of science education and teacher training at primary school level. it can be utilized as a research instrument for effect studies of training courses and other interventions aiming to professionalize primary teachers. furthermore, it can serve as a diagnostic tool for adapting training courses and interventions to the individual needs of preand inservice teachers. and finally, it can be used as a coaching tool for making primary teachers aware of their own view of science and their (changed) attitudes toward teaching science. by means of these different uses, the das instrument could become a highly valuable instrument for making progress within the field of science education in primary schools. 2.3 professional development our in-service training course was also based on the different underlying components in our theoretical framework and consisted of six 3-hour meetings spread over six months (walma van der molen, van aalderen-smeets & groot koerkamp, 2011). the course focused on creating awareness about teachers’ attitudes towards teaching science, awareness about their views on science, and awareness about their attitude towards inquiry-based methods of learning. these attitudes were reflected upon and challenged by means of assignments, exercises, questioning, information transfer, and research activities like experiments. in addition, each of these course elements was accompanied by activities that stimulated teachers’ own curiosity, inquisitiveness, critical thinking, reflection and metacognition, and higher-order thinking. most importantly, the training course refrained from providing recipe-like example lessons, pre-structured materials, or predefined methods. during the meetings, teachers engaged in coursework that prepared them for take-home assignments. during the final meeting, participants presented a science and inquiry-based project that they had developed and carried out with their pupils (see appendix for a general overview of the core elements of the course). to test the effectiveness of our course, we used a pre-post test, experimental-control group design to asses changes in teachers’ attitudes toward teaching science over time (experimental group n = 49, control group n = 45). the experimental group participated in the training course, while the control group did not receive any formal training. however, the control group did consist of teachers that reported to be interested in science education. we used our ‘dimensions of attitude toward science’ (das) instrument to measure quantitative changes in teachers’ attitudes towards teaching science. in addition, we used qualitative, open ended, self-report measures to investigate changes in scientific attitude, perceptions of science, and attitude towards inquiry-based methods of learning. j. walma van der molen & s. van aalderen-smeets 7 | f l r two out of three attitude components showed significant improvements (see table 1), indicating that the training course had a positive effect on attitude towards teaching science (van aalderen-smeets & walma van der molen, 2013b). participants in the course gained a more positive attitude in the affective and perceived control dimension of attitude, compared to the control group. this means they enjoyed science teaching more, showed increased self-efficacy, and felt less dependent on recipe-like, standardized methods, top down instruction, and the availability of pre-organized projects and materials. participating teachers in the training course did show a significant improvement on the remaining attitude components (less anxiety when teaching science, believing science teaching is more relevant, and less stereotypical beliefs), but this improvement was not significantly different from changes in the control group, even though the changes in the control itself were not significant, see table 1. this could be due to the relatively high interest in science and engagement with science of the control group, i.e., because they engaged in science related teaching and activities in between the preand post-test, they improved their attitudes slightly. on the open-ended questions, teachers indicated that, after participation in the course, they found science to be less complex and to use inquiry-based methods of teaching more often in both science-related lessons and in other school subjects. in addition, teachers’ responses showed enhancement in their scientific attitude, i.e., they became more curious, more critical, and more explorative. furthermore, teachers’ perceptions and expectations of their pupils changed; teachers reported to be surprised by the excellent achievement of some pupils during science lessons (for a more detailed report about this effect study, see van aalderen-smeets & walma van der molen, 2013b). table 1. attitude toward teaching science; comparison of attitude scores between trained and control group on pre and post-test. mean difference scores of the trained and control group are presented in the left columns. the right columns present the results of an anova analysis for each component of professional attitude (significant effects are printed in bold). trained group control group mdiff sd mdiff sd f value (1, 104) p eta cognition relevance .24a .59 .08 .64 1.9 .17 .02 gender -.27a .82 -.06 .87 1.6 .21 .01 affect enjoyment .53a .80 .14 .90 5.6 .02b .05 anxiety -.39a .77 -.15 .95 2.1 .16 .02 perceived control self-efficacy .40a .55 .11 .51 7.6 .01b .07 context dependency -.98a 1.10 .09 .98 27.8 (1,102) .01b .21 mdiff = mean difference score (t2-t1), sd = standard deviation, a significant improvement within group (paired t-test between pre-test and post-test within group) b significant difference between improvements of trained and control group (anova interaction effect) this study indicates that focusing explicitly on primary teachers’ attitudes in a training course does improve the beliefs, feelings, and perceived control primary teachers have regarding science education and inquiry based methods of learning. j. walma van der molen & s. van aalderen-smeets 8 | f l r 3. future directions this large-scale research project provides valid tools and new approaches for improving and assessing primary teachers’ attitude towards (teaching) science and opens the door for several future directions in attitude research. the attitude effects reported here are short-term effects based on self-reports. further research is needed to investigate the long-term effects of teachers’ attitude change and changes in their actual teaching. furthermore, more research is needed on the effects of improved teacher attitudes on their pupils’ or students’ attitude toward science and their career choices in their future school career. the results presented here are not only relevant for primary education. the gained knowledge about improving attitudes towards science can be applied in interventions, professional development, and research aiming to improve secondary school students’ attitudes towards science as well. in addition, the approach taken in this research project, i.e., building a theoretical framework, then constructing a valid instrument, and testing the effects of a new intervention, may be followed for other lines of research, such as research on attitudes towards inquiry based learning or on attitudes towards the use of digital media in education. the framework presented in this article is an essential new step toward a convergence of the research in this field. only when researchers are aware of the complexity of the construct of teachers’ attitudes toward science, when explicit and substantiated decisions have been made regarding which components and objects should be measured, and when methodologically sound instruments and interventions are used, can scientific progress be achieved in research on teachers’ attitudes. future research is needed to investigate the various aspects of the proposed framework, including the relationships between the components and the weights of the various components and sub-attributes in predicting behavioural intention. the investigation of these aspects is a prerequisite for gaining further insight into the dynamics of primary teachers’ attitudes toward science and for the development of interventions that are better suited to improve specific aspects of teachers’ attitudes. keypoints professional development of primary teachers in science education should pay explicit attention to attitude improvements. attitude research should be based on a theoretical model, such as the framework of attitude towards science. primary teachers feel more in control over science teaching when they have gone through the inquiry process themselves. primary teachers’ attitudes towards teaching science can be improved by stimulating their own curiosity, scientific attitudes and thinking skills. sound theoretical and methodological attitude research should be on the frontline of research in science learning. acknowledgements contract grant sponsor: platform beta technology in the netherlands. references ajzen, i., & fishbein, m. (1980). understanding attitudes and predicting social behavior. englewood-cliffs, nj: prentice-hall. j. walma van der molen & s. van aalderen-smeets 9 | f l r asma, l.j.f., walma van der molen, j.h., & van aalderen-smeets, s.i. (2011). primary teachers’ attitudes towards science: results of a focus group study. in m.j. de vries, h. van keulen, s. peters, & j.h. walma van der molen (eds.). professional development for primary teachers in science. the dutch vtb-pro project in an international perspective (pp. 89–105). rotterdam: sense. barmby, p., kind, p.m., & jones, k. (2008). examining changing attitudes in secondary school science. international journal of science education, 30, 1075–1093. doi: 10.1080/09500690701344966. bennett, j., rollnick, m., green, g., & white, m. (2001). the development and use of an instrument to assess students’ attitude to the study of chemistry. international journal of science education, 23, 833–845. doi: 10.1080/09500690010006554. coulson, r. (1992). development of an instrument for measuring attitudes of early childhood educators towards science. research in science education, 22, 101–105. doi: 10.1007/bf02356884. eagly, a., & chaiken, s. (1993). the psychology of attitudes. belmont, ca:wadsworth group/thomson learning. gardner, p.l. (1995). measuring attitudes to science: unidimensionality and internal consistency revisited. research in science education, 25, 283–289. doi: 10.1007/bf02357402. haney, j.j., czerniak, c.m., & lumpe, a.t. (1996). teacher beliefs and intentions regarding the implementation of science education reform strands. journal of research in science teaching, 33, 971–993. doi: 10.1002/(sici)1098-2736(199611)33:9,971::aid-tea2.3.0.co;2-s. harlen, w., & holroyd, c. (1997). primary teachers’ understanding of concepts of science: impact on confidence and teaching. international journal of science education, 19, 93–105. doi: 10.1080/0950069970190107. jarvis, t., & pell, a. (2004). primary teachers’ changing attitudes and cognition during a two-year science inservice program and their effect on pupils. international journal of science education,26, 1787– 1811. doi: 10.1080/0950069042000243763. osborne, j., simon, s., & collins, s. (2003). attitudes towards science: a review of the literature and its implications. international journal of science education, 25, 1049–1079. doi: 10.1080/0950069032000032199. osborne, j., & dillon, j. (2008). science education in europe: critical reflections (a report to the nuffield foundation). london: the nuffield foundation. retrieved from http://www.pollen-europa.net/pollen dev/images editor/nuffield report.pdf. pajares, m.f. (1992). teachers’ beliefs and educational research: cleaning up a messy construct. review of educational research, 62, 307–332. doi: 10.3102/00346543062003307. reid, n. (2006). thoughts on attitude measurement. research in science & technological education, 24, 3– 27. doi: 10.1080/02635140500485332. tai, r. h, liu, c. q, maltese, a. v, & fan, x. (2006). planning early for careers in science. science, 312, 1143-1145. doi: 10.1126/science.1128690. turner, s., & ireson, g. (2010). fifteen pupils’ positive approach to primary school science: when does it decline? educational studies, 36, 119-141. doi: 10.1080/03055690903148662. tosun, t. (2000). the beliefs of preservice elementary teachers towards science and science teaching. school science and mathematics, 100, 374–379. doi: 10.1111/j.1949-8594.2000.tb18179.x. van aalderen-smeets, s.i., walma van der molen, j.h., & asma, l.j.f. (2012). primary teachers’ attitude toward science: a new theoretical framework. science education, 96, 158–182. doi: 10.1002/sce.20467. van aalderen-smeets, s. i. & walma van der molen, j. h. (2013a). measuring primary teachers’ attitudes toward teaching science: development of the dimensions of attitude towards science (das) instrument. international journal of science education, 35, 4, 577-600. doi:10.1080/09500693.2012.755576. van aalderen-smeets, s. i. & walma van der molen, j. h. (2013b). improving primary teachers’ attitudes toward science by attitude focussed professional development. (submitted). j. walma van der molen & s. van aalderen-smeets 10 | f l r walma van der molen, j.h., aalderen-smeets, van, s.i., & groot koerkamp, e. (2011). cursusboek talentontwikkeling, wetenschap en techniek: professionalisering voor basisschoolleerkrachten [coursebook talent development, science, and technology: professional development for primary teachers]. knowledge center for science and technology (kwto). yates, s., & goodrum, d. (1990). how confident are primary school teachers in teaching science? research in science education, 20, 300–305. doi: 10.1007/bf02620506. appendix schematic overview of the science education-training course for primary teachers (walma van der molen, van aalderen-smeets & groot koerkamp, 2011). content elements take-home assignments attitude towards (teaching) science scientific attitudes scientific skills personal development in class 1 creating awareness about view of science introduction on attitude toward science stimulating teachers’ curiosity and amazement about everyday items keeping a diary of amazement identifying and challenging pupils views about and perceptions of science linking meeting 1 to 2: from amazement and curiosity to formulating research questions 2 challenging cognitive beliefs about the relevance of science and stereotypical gender beliefs regarding science stimulating curiosity, inquisitiveness and question asking, dealing with scientific uncertainty and ambiguity formulating research questions and hypotheses evaluating a science education method or related medium (website, tv) with attitudinal criteria how many difficult questions can you think of? linking meeting 2 to 3: from research questions to research design and enjoying science 3 challenging the affective component of attitude, i.e. enjoyment and anxiety stimulating a critical attitude choosing research method and design, and conducting research in the classroom self-observation: searching for opportunities to integrate science in your existing lessons conducting research in the classroom; from research question to experiment linking meeting 3 to 4: not only hands-on but minds-on; stimulating academic thinking skills 4 being persistent, creative and original creative and higher order thinking skills; developing a thinking lesson improving an existing science method stimulating creative thinking in children j. walma van der molen & s. van aalderen-smeets 11 | f l r linking meeting 4 to 5: independence of context factors and being in control of science teaching 5 stimulating selfefficacy and perceived control using different perspectives to solve problems stimulating reflective and metacognitive thinking skills developing and teaching a science lesson linking meeting 5 to 6: from feeling in control to actually teaching science 6 summary of training course participant’s presentations of visual reports on science lessons note that this figure provides a schematic overview. several additional attitudinal elements are interwoven in the course. also spontaneous questions and comments from the participants that came up during the course were explained in terms of attitude or related to attitude toward science and scientific attitude. the first four columns are aimed at the personal development of the teacher. the in class, take home assignments are explicitly aimed at the interaction with pupils. microsoft word gonzalez-ocampo et al_publication .docx ! ! ! frontline learning research vol.3 no. 3 special issue (2015) 23 38 issn 2295-3159 corresponding author: dr. janice malcolm, reader in higher education, centre for the study of higher education, uelt building, university of kent, canterbury ct2 7nq, uk, (+44) 01227 824579, cshe@kent.ac.uk doi: http://dx.doi.org/10.14786/flr.v3i3.191! ! the curriculum question in doctoral education gabriela gonzález%ocampoa, margaret kileyb, amélia lopesc, janice malcolmd, isabel menezese, ricardo moraisf, viivi virtaneng a ramon llull university, spain b the australian national university, and university of newcastle, australia c university of porto, portugal d university of kent, united kingdom e school of economics and management, universidade católica portuguesa (porto), portugal f university of helsinki, finland article received 18 july 2015 / revised 18 july 2015 / accepted 23 july 2015 / available online 25 september 2015 abstract the landscape of doctoral education has changed immensely during the last decades. different transnational policies, different publics, different purposes and different academic careers all contribute to the need for a new understanding of this underresearched field. our focus is on explicit curriculum analysis to undertake intentional and meaningful change, especially in terms of the processes and outcomes of doctoral education. we draw on research on doctoral education, as well as the emerging literature on early career researchers (ecrs) and on professional learning, and consider how the concept of curriculum can help us think differently about doctoral education, particularly in relation to processes and outcomes. finally, we suggest a research agenda for developing the curricula of doctoral education. keywords: doctoral education; curriculum; processes, outcomes, professional learning !gonzalez)ocampo!et !al ! | f l r ! ! 24! 1. introduction in recent years there has been a burgeoning of research interest into the experiences of phd/doctoral students and supervisors, although much of this work is limited to specific models and contexts of doctoral education (e.g., gardner, 2007; golde, 2005; ives & rowley, 2005; mcalpine, paulson, gonsalves, & jazvac-martek, 2012; pyhältö, vekkaila, & keskinen, 2012; scaffidi & berman, 2011; vekkaila, pyhältö, & lonka, 2013). there has been a clear evolution from the individual focus of the “master-apprenticeship” model to more structured programmes, an increasing number of candidates, growing internationalisation of the academy, and the emergence of new types of phds which reconfigure the relationship between research and practice (brew & peseta, 2004; pearson, evans, & macauley, 2008; walker, golde, jones, bueschel, & hutchings, 2008). as the quality of research and supervision are increasingly recognised as decisive in the process of doing a phd (and the products that emerge from it), the question of “what a phd really is” is also under discussion, leading wellington (2013) and others to explore the possible meanings of “doctorateness”. this research field has struggled to keep pace with proliferating phd formats and diverging practices. doctoral contexts such as “practice as research”, professional doctorates, phds based entirely on publications, etc. are less well understood than more traditional formats, and are stimulating increasing research interest (e.g. the carnegie project on the education doctorate http://cpedinitiative.org/; kot & hendel, 2012; nelson, 2006). if we look at researcher education more broadly conceived, there has been very little work on the extended ‘adolescence’ of academic researchers, or on the experiences and trajectories of researchers in other professional fields beyond the academy; mcalpine, amundsen, and turner (2013) offer one of the few contributions in this area. this raises questions about how far doctoral education succeeds (or perhaps has ever succeeded) in providing appropriate professional preparation and enhancement to those for whom an academic career is a clear motivation. even more so than in undergraduate education, the doctoral student has commonly been seen as an apprentice member of a disciplinary community. the phd degree, once considered the pinnacle of academic achievement, is increasingly regarded as a kind of entry-level global academic passport offering junior scholars access to an insecure career (pearson et al., 2008). yet as we have seen, the proliferation of types of doctorate has been partly driven by the demands of careers outside academia. the development of professional doctorates, and the realisation that many, or even most, phd graduates will experience careers outside the academic labour market have given impetus and legitimacy to the inclusion of employability skills in doctoral education (baker & lattuca, 2010). these developments have been justified in terms of harmonisation and flexibility, and have borrowed heavily from skills models used in vocational education (see e.g., vitae); as yet we have little evidence of how effective they are at meeting their multiple (and often unclear) purposes. in practice, this skills-oriented approach to doctoral education tends to meet resistance from subscribers to a more purist view of the university, according to which academic freedom is not compatible with external standardisation initiatives (kiley, 2014). however these changes also raise important questions about academic judgments and assessment processes at the doctoral level, which are far from standardised and remain, for the most part, poorly understood. these profound changes are occurring in a context where we have scarcely begun to explore the nature of the formal and informal curriculum of doctoral education (eua, 2007). the introduction of the idea of curriculum in phd programmes necessitates urgent discussion among educators from different backgrounds and with different perspectives on doctoral education. we need a clearer understanding of how the curriculum of doctoral education works, how it can be developed to meet changing needs, and how its outcomes can be appropriately assessed. in this paper our goal is to explore whether an explicit curriculum approach can help us make sense of existing research and practices regarding the processes and outcomes of doctoral education. our starting point is that, whether we acknowledge it or not, the curriculum is inevitably there; and adopting an explicit curriculum approach will help us to disclose the tensions between the formal/informal, open/hidden, and standardised/pluralised dimensions in doctoral education and brought to our attention by enders (2002). these dimensions are summarised in table 1 and contribute to a research agenda that allows us to develop more nuanced and useful understandings of the doctoral education curriculum. we blend the contributions of an ecological and socio-constructivist perspective (e.g. bronfenbrenner, 1979, 1986; lave, 1988; vygotsky, !gonzalez)ocampo!et !al ! | f l r ! ! 25! 1978) with curricular viewpoints grounded in the policy cycle of stephen ball (1994) that frame a vision of the curriculum as a contextual and social-cultural phenomenon entailing a continuous meaning-making process in which diverse interpretations struggle to emerge (lopes & lópez, 2010). however, before addressing the above in the context of doctoral education we suggest that it might be helpful as background to outline a more standard approach to curriculum, the sort of approach we might see in texts related to coursework degrees at the tertiary level (e.g. kiley, 2014; print, 1987). while the starting point in curriculum is generally contested, one might begin with examining the aims for the course or program. this is where a question such as: what is the teacher/ faculty/ university aiming to achieve with this course? engaging staff in answering this question often uncovers many of the implicit, as well as explicit, views held by participants. at the doctoral level, asking what might appear to be such a simple question is likely to highlight a wide and complex set of responses. again, while curriculum development is rarely linear, for the sake of argument, the next question that can be asked is: what knowledge, skills and attitudes is it expected that the learner will be able to demonstrate following engagement in this program? this stage is often termed “learning outcomes” and at the doctoral level it is again contested with comments ranging from employment skills through to higher level cognitive skills and an original contribution to knowledge. a logical next step in light of having determined the potential learning outcomes is the identification of the possible learning content and activities that are to be provided to allow the learner to engage in appropriate learning. in some cases this is referred to as the syllabus. again at the doctoral level, the notion of content for learning is wide and varied often depending on country, discipline, and type of doctorate being undertaken. linked to the content is the consideration of pedagogy. it is during this stage in curriculum development that questions are posed regarding the teaching approaches to be used. until recently pedagogy was a term that was rarely used in relation to doctoral education (boud & lee, 2009). rather, there was an assumption that the supervisor/mentor/adviser would work with the candidate in private and mysterious ways until the candidate had achieved a level of doctorateness (trafford & leshem, 2009). the concept of achieving doctorateness brings us to the next stage of curriculum development, that is, assessment. in much of the curriculum literature there is discussion of the concept of the aligned curriculum, that is, where the assessment strategies closely align with the espoused learning outcomes (biggs, 2003). at the doctoral level there are various practices ranging from the inclusion of the results of coursework in the assessment through to examination of the written thesis only, or the inclusion of assessment of the candidate’s performance in an oral examination. the final stage in this formalized model of curriculum development is evaluation where the various stages, activities and outcomes are evaluated in an ongoing fashion. following this formal and somewhat stylised discussion of curriculum we now discuss the concept of the curriculum in doctoral education in more sophisticated and complex ways, and then address the processes and outcomes of doctoral education. the analysis of research and practice suggests that there is a strong need for a research agenda that will help reconfigure the notion of curriculum in doctoral education. 2. the curriculum in doctoral education as noted above, it is relatively unusual to speak of “curriculum” in relation to doctoral education. jones, in his review of 40 years of research on doctoral education (jones, 2013), does not use the term at all, though it is implicit in several of the themes he identifies, such as programme design, doctoral writing and research, and socialisation; this obliquity is echoed too in calma and davies’ (2015) review of the history of one key higher education journal. the fact that the curriculum in doctoral education is not explicitly discussed does not make it less significant. however it may hinder our recognition of how the curriculum !gonzalez)ocampo!et !al ! | f l r ! ! 26! can generate and reproduce inequalities, and of the need for change and adaptation to the new challenges of doctoral education. moreover, a focus on the curriculum must acknowledge the particularities of doctoral education, and in particular the possible incompatibility between current tendencies towards regulation and structure, and the flexibility and plurality inherent in doctoral education (enders, 2004; pearson et al., 2010). this is only possible if we recognise the curriculum as the unacknowledged “elephant in the room”. in theories of formal education, the curriculum is often understood as a structured selection of propositional knowledge and/or skills which learners need to acquire in order to meet the aims and objectives of the learning programme (e.g. eraut, 2000; print, 1987). with aims and outcomes clearly defined and made explicit, it is then possible to “align” these with appropriate learning activities and assessment strategies (biggs, 2003). however, as colley, hodkinson, and malcolm (2003) argue, “formal” learning activities are only ever one strand of any learning situation and cannot be extricated from the social context in which they take place. although there may be broad learning objectives for higher education programmes, it is expected that learners will develop a degree of self-direction, and will inevitably emerge from the process with differing understandings of the academic content and with varied mastery of research skills. thus educational researchers have turned increasingly to a range of alternative approaches to the curriculum to analyse what is learned at university, and how it is learned (e.g. brennan et al., 2009). doctoral education specifically entails a further shift of emphasis away from the standardised formal curriculum (see table 1), and towards a highly complex set of structures, practices and expectations from which doctoral students and their supervisors create new and unpredictable learning. this complex and pluralised perspective, seeks to address the diversity of training needs and career preferences by adjusting not only to labour concerns but also to students, supervisors and administrators. however, social and labour claims for specific needs may lead into the development of standardised programs. thus, curricula in doctoral education may struggle to find a balance between these two perspectives that contribute to defining their orientation. therefore, a curricular perspective cannot ignore core changes and challenges doctoral education entails. pearson et al. (2010) point out that in the current context of doctoral education, “opportunities for researchers, or employees with enhanced research skills, now arise inside universities and in non-university settings where knowledge and professional industries develop their capacity to carry out work that draws on specialist knowledge and research skills (e.g. contract research, university administration, school teaching, nursing and business)” (p. 348). discussing the dilemma between standardisation and pluralism, described in table 1, the same authors advise that “any attempt to resolve [it] must draw on a fully accurate and up to date picture of the contemporary doctoral experience and address the goals, motivation and expectations of the increasingly diverse doctoral population. particularly important is recognition that the connection and integration of work and learning is an issue for research education, as for other forms of higher education” (pearson et al., 2010, p. 349). a curricular perspective on doctoral education may then take into account the new ecology of doctoral education, considering that students’ experience is framed by (and frames) what happens at the different levels of the ecological system. in bronfenbrenner’s (1979, 1986) ecology of human development, for example, curriculum can be seen as an interaction system constituted by different “nested” subsystems: microsystem (in this context, what happens within typical classes in doctoral programs), mesosystem (what happens within universities, research centres or professional industries), exosystem (educational policies regarding doctoral education), macrosystem (cultural models in a certain period, such as representations of doctoral education and the mandates of universities or the significance of professional phds) and chronosystem (changes that result from specific non-normative events, such as the bologna process in europe). this is surely a good departure point that must be reinforced with two additional features: on the one hand, the representations, contents and meanings that mark the interactions between each ecological system, and on the other, the experience of students in their journey through the doctorate. ball’s policy cycle (1994), particularly as reinterpreted by lopes and macedo (2011) is helpful in addressing the first feature. ball’s studies (1989, 1994; ball, bowe, & gold., 1992) focus on micro-political processes and the need to articulate macro and micro levels in curricular studies. lopes and macedo (2011) insist on the non-hierarchical character of the policy cycle in the field of the curriculum, emphasising the circularity of its three contexts: !gonzalez)ocampo!et !al ! | f l r ! ! 27! the context of influence, i.e., of policy-producing; the context of policy text production, and the context of policy practices. this perspective assumes that the curriculum is in itself the struggle for meaning (lopes, 2012) and reveals how the context of policy practices can drive, and is driven by, the other contexts. broadly social-constructivist and situated perspectives on learning can be helpful in identifying significant features of the curriculum of doctoral education that are relevant to understanding the journey of doctoral students (e.g. brown, collins, & duguid, 1989; greeno, collins, & resnick, 1996; lave, 1988; rogoff, 1998). the influence of vygotsky (1978) is apparent in a number of alternative theorisations of learning processes, and indeed vygotsky’s social-cultural approach to learning offers the basis for rich understandings of the contextual and relational dimensions of the curricula. where learning is seen as situated, knowledge is immersed in and generated by the activities, relationships, tools, contexts and culture that occur in daily activities. this implies recognising the collective, participative and social nature of cognition (rogoff, 1998), and this emphasis on social engagement and communication has significant implications in a context where the phd has increasingly become an interactional rather than a solitary endeavour. the notion of communities of practice is of particular interest here; the development of an identity as a researcher can be clearly understood as legitimate peripheral participation through engagement in research activities within a research group or disciplinary community. this attention to how “informal” practices and messages are produced and conveyed has been taken up in the literature of learning in the workplace, and theories of social learning developed to explain workplace practices have increasingly been applied to educational settings as well (e.g. lave & wenger, 1991; billett, 2009). doctoral students, from this perspective, are situated as both learners and emerging practitioners in the discipline, increasingly inhabiting the identity and responsibilities of professional disciplinary researchers in an academic workplace (and the extent of these responsibilities varies in different national systems of doctoral education) or highly qualified and innovative professionals in a hybrid academic and professional context. indeed, an emphasis on inclusion in research communities or networks and the creation of collaborative knowledge-sharing environments appears to be a significant trend in doctoral education (johnson, lee, & green, 2000; malfroy, 2005; pyhältö, stubb, & lonka, 2009). this view of doctoral education emphasises the pluralised approach to curriculum and specifically the impact of the social context in which training take place (table 1). the “landscape” metaphor proposed by clandinin and connelly (1995) can also be useful in analysing doctoral students’ experiences as they construct their identities as researchers. within this metaphor, learning involves a double transaction (biographical and relational) that results from the relationships between people, places, and things, and this view of the “landscape” of professional development as being inherently relational (in itself made up of relations), provides a gateway for relating the study of identity to the study of curriculum (lopes & pereira, 2012). recent work on socio-material understandings of learning (e.g. fenwick, edwards, & sawchuk, 2012) suggests that the “landscape” metaphor can be extended to include all of the actors and practices present in a learning setting – social, material, technological, pedagogic, symbolic – and a close attention to their multiple, complex connections and interactions. the fact that the profile of doctoral students and doctoral programs has changed also implies that issues of identity development will also change (baker & lattuca, 2010). the consensual current distinction within curriculum theory, between “formal”, “informal” and “hidden” curricula (pacheco, 1996) seems to assume a special relevance here (see table 1). the formal curriculum refers to qualifications frameworks, course syllabi; the informal curriculum relates to what is really done through teaching and learning processes, such as readings and discussion, interactions with researchers in the context of classes; and the hidden curriculum represents the unintended learning, often in regard to class and gender roles, social expectations, etc., that emerges from structures, relationships and practices in the educational setting, revealing the pedagogy of the learning context, rather than its intended content (apple, 1971). doctoral education clearly involves codified objectives of degree programmes, as well as a complex web of structures, practices and expectations far beyond the more explicit/formal dimension. solem, hopwood, and schlemper (2011) explore what kind of events made students feel an “academic and belonging to a departmental community” (p. 10) and conclude that mostly these are “informal events [that] include conversations [and] social events” (p. 12): some doctoral students mentioned joint !gonzalez)ocampo!et !al ! | f l r ! ! 28! coffee meetings or lunches as significant experiences. however, these events might be experienced very differently by different students. margolis and romero (1998) find “patterns of interaction with intended and unintended consequences that make it particularly difficult for students of color, women, and students from working-class background to survive and thrive in graduate school” (p. 2). gender relations also appeared relevant in the study by solem et al. (2011), with women expressing more extreme evaluations of support that interfered in their perception of progress in their own work; international students and non-white minorities also seem to report more troubles and feelings of isolation. margolis and romero consider apple and king’s (1977) notion of the weak (related to professionalism) and strong (related to socialisation) hidden curriculum, concluding that the formal curriculum (e.g. affirmative action policies) often contradicts these hidden dimensions at the expense of successful experiences for minority students. implicit in all of these alternative approaches to understanding learning as social and situated, is the fundamental problematisation of any notion of a stable set of knowledge and skills to be learned and assessed. guerin (2013) argues that “rhizomatic” models of knowledge structures as proposed by deleuze and guattari (1980) may be a more appropriate way to understand knowledge-content and research cultures at doctoral level: “in effect, this alternative model acts as a licence to try out new combinations of ideas. thus, a rhizomatic research culture is characterised by heterogeneity, multiplicity, proliferation, flexibility, non-linearity, connection and non-hierarchical networks” (guerin, 2013, p. 139, emphasis added). alternative conceptions of knowledge as emergent in social practices (e.g. hager, lee, & reich, 2012), socio-material assemblages (fenwick & nerland, 2014) and hybrid or interdisciplinary research fields (clausen, pohjola, sapprasert, & verspagen, 2012), all offer further possible starting points for a more nuanced analysis of the complexities of the processes and outcomes in the doctoral curriculum. however, some departmental cultures seem to emphasise the phd as a solitary endeavor which students should be able to cope individually (solem et al., 2011). in the next two sections, we turn our attention first to a more detailed discussion of the processes and experience of the doctoral curriculum, and then to the assessment of its outcomes. 3. doctoral education processes – how the curriculum is experienced the analysis of the lived curriculum of doctoral education should firstly consider doctoral students’ experiences during their candidature. recent studies suggest there is quite a high variation in how ecrs experience the doctoral study process, but there are also strong indications that good progress and satisfaction with doctoral education are more likely where candidates experience factors such as good supervisory relationships, belonging to an academic community, and/or being able to contribute new knowledge in science (ives & rowley, 2005; zhao, golde, & mccormick, 2007; overall, deane, & peterson, 2011). it is clearly difficult to identify what emerges from the formal or informal curriculum, or to distinguish formal from informal learning within student experiences. however, some results suggest that when doctoral students talk about their most meaningful experiences, they tend not to emphasise formal studies or other activities that might be seen as constituting the formal curriculum (virtanen & pyhältö 2012, vekkaila et al. 2013); anderson & anderson (2012) also indicate that the curriculum does not always work as intended. from the perspective of doctoral students, it seems, the curriculum appears undefined and lacking in focus, but further research is needed to explore specific conceptions about the curriculum and its manifestations in different contexts. a wide range of activities influences students’ experiences during their doctoral journey, these activities shed light about the different manifestations of curriculum (see table 1). the way in which curriculum is experienced goes beyond institutional policies; beliefs and expectations have a main role, which can create tensions between students’ expectations and supervisors and administrators’ perspectives about doctoral training. !gonzalez)ocampo!et !al ! | f l r ! ! 29! a recent study on postdoctoral researchers (postdocs) who had already successfully completed their doctoral studies suggests that career planning should ideally have been included in their doctoral education from the beginning of the doctoral study process. these postdocs also stressed that formal study and other academic activities should have been designed with a view to supporting their future careers. these findings are in line with those of scaffidi and berman (2011) who argue that for postdocs to have the best chances of prospering in academia, industry, or elsewhere, they need to plan their future careers strategically. analysing the experiences and conceptions of post-doctoral researchers (pitcher & åkerlind, 2009) is essential to promoting their future career development after the phd; thus rethinking the curriculum of doctoral studies is vital not only from the perspective of doctoral students themselves, but also from that of higher education researchers. åkerlind argues for “varied” and “flexible” provision to enable postdocs to make “informed career decisions” (åkerlind, 2009). others have proposed a reconceptualisation of postdoctoral research pathways to produce a better “fit” between training and professional interests and skills (berman, juniper, pitman, & thomson, 2008). thus a review of the curriculum of both doctoral and postdoctoral preparation is acknowledged as an essential task. a “hybrid curriculum” model to address the connections among university, profession and workplace, is proposed by lee, brennan, and green (2009) as a way of adapting the curriculum for diverse doctoral needs. this idea has also engendered further studies reviewing the purposes of doctoral education, and taking into account the changing needs of the “knowledge economy” in academic, professional, social and labour domains. this questioning of assumed and hitherto tacit purposes has also encouraged the development of alternatives to traditional doctoral programmes, such as practitioner or professional doctorates for those who are engaged in leading practice and introducing change in tandem with their academic research (lester, 2004). utilising research on networking learning, and on students’ socialisation in disciplinary communities and in other professional fields (e.g. vaessen, van den beemt, & de laat, 2014; boden, borrego, & newswander, 2011) could also strengthen the development of interdisciplinary curriculum structures, enabling ecrs to construct and assume their professional roles taking broader labour market needs into account. studies of the academic transitions experienced by junior researchers could also deepen our understanding of the academic and professional practices needed to offer more appropriate training and support to ecrs, enabling them to make the transition from doctoral education to other careers (mcalpine & emmioğlu, 2014). where the focus is clearly on preparation for an academic career, the quality of supervision emerges as key to supporting doctoral students’ developmental processes (roulston, preissle, & freeman, 2013). in this context the supervisory relationship is of fundamental importance to how students experience the “doctoral journey” (pyhältö et al., 2012; zhao et al., 2007; mcalpine et al., 2013); students’ learning experiences and satisfaction are closely related to the nature of the relationship developed between students and supervisors, so the role of the supervisor is critical to constructive doctoral preparation (lee, 2008). solem et al. (2011) emphasise how “timely, proactive, and supportive advising and mentoring from faculty, peers, and program committees” (p. 13) are essential elements for preventing difficulties. yet the practice of supervision (and often of pedagogy more generally) only becomes a developmental focus after students have completed their thesis, thus presenting a clear obstacle to their development as future academics. as mcalpine et al. (2013) point out, this means that doctoral supervision is a long-term and collective process, and this needs to be acknowledged in the structuring of the curriculum. existing research on doctoral students reveals a high degree of variation in the experience of doctoral study processes (mcalpine & mckinnon, 2013) and further work is needed in order to understand how the curriculum shapes and influences these experiences, particularly with regard to the study of the experiences that are promoted in formal, informal and hidden curriculum and how these experiences affect students’ training as well as the role of supervisors and administrators. this could include longitudinal studies to examine how doctoral programmes are currently developing and how far this development aligns with changes in industry and the employment market. this could then inform discussions of how far the doctoral curriculum and the training of doctoral students can or should be adapted to meet the changing and multiple purposes of the phd. the academic and professional socialisation and disciplinary networking of doctoral !gonzalez)ocampo!et !al ! | f l r ! ! 30! students also merit more extensive study; this remains a relatively under-researched area (anderson & anderson, 2012), despite its key importance to students and to their future careers. 4. the outcomes of doctoral education – assessment and employability the question of how the outcomes of doctoral education are assessed cannot be avoided in any discussion of the doctoral curriculum, particularly in the light of the ongoing diversification of programmes and career paths. in this section we consider two of the outcomes of doctoral education: assessment and employability. in spite of commonalties in terms of formality and structure, assessment varies significantly by discipline, country, institution, and supervisor. in addition, the “core competences” of a phd may serve both academic and non-academic careers; these multiple purposes have complex implications which are not yet fully understood, and which may not be susceptible to standardised or comprehensive solutions. the final examination is only one aspect of the complex assessment processes occurring at the doctoral level. for example, we have forms of assessment at entry to a doctorate, and ongoing assessment during the candidature. depending on the country or the disciplinary context, this may take formal shape through the marking and grading of coursework, or structural milestones such as confirmation of candidature seminars, annual reports of progress, mid-term and final seminars. informal assessment occurs throughout candidature as judgments are made by the supervisory team on the quality of writing and thinking candidates display, and peers reach verdicts on the quality of research papers submitted to journals and conferences. these various strategies vary by institution and country. for example in some systems an advisory committee additional to the supervisor/s will have an overview of the quality of the candidate’s work and progress, and meet to assess key milestones. some institutions have developed rubrics to use for assessing these various milestones. others require candidates to provide reflective essays on learning, or to develop a portfolio, or to produce a number of peer-reviewed publications prior to completion. all of these assessment strategies support the expectation of experienced examiners that the thesis they are about to examine is passable (golding, sharmini, & lazarovitch, 2014; mullins & kiley, 2002). despite the variety of formal and informal assessment strategies employed during candidature, the most common formal assessment at doctoral level is the final examination, known by a number of different names and exhibiting a wide range of types (hartley, 2000; morley, leonard, & david (2002). variations in vivas: quality and equality in british phd assessments, 2002). for example, in parts of europe and scandinavia, following examination and approval of the written thesis, the candidate publicly defends her/his thesis before an audience of academics and others. this process is in stark contrast to the uk model where the written thesis is generally examined by one internal and one or two external examiners, and then a private viva voce is held, in some cases in the presence of a neutral chair who oversees the process. while an oral examination in held in canada this is generally a semi-public affair, often with four to five supporters joining the candidate. the us model is different again: the candidate has a “committee” with whom they interact on occasions throughout their candidature, and when the supervisor thinks the candidate is ready, the committee conducts a private oral examination where the candidate “defends” the thesis. a very different model exists in australia and south africa, where the written thesis is the sole examinable item (although universities offer the option of an oral if the examiner requires one). a high level of confidentiality is maintained; the candidate does not know who the examiners are, and the examiners are generally unaware of each other’s identity, and do not discuss the work among themselves. each university has a process for bringing together the various reports into a single recommendation, as a journal editor might do with reviewers’ reports (kiley, 2009). given the diversity of approaches to assessment, in the complex settings of various approaches to curriculum in doctoral education (table 1) one particular question arises: what is being examined? when the !gonzalez)ocampo!et !al ! | f l r ! ! 31! written work is examined, one could argue that it is the candidate’s demonstrated ability to be a researcher that is being assessed, judged by the quality of the research and its presentation. with the oral component, it is arguable that other qualities are being assessed, such as the candidate’s broader knowledge of the discipline, and their ability to deal with challenges to their work. however, in view of curriculum considerations and the substantial international developments in doctoral education outlined above, we suggest that there may be other assessable outcomes of the doctoral learning experience which are not yet fully developed, and are not currently the focus of formal assessment. international research in this area is in its early stages; we suggest that it is time to reconsider formal and informal types of assessment for future academic researchers, as well as for those in, or aiming for, other kinds of professional employment. future research will need to take into account the specifics of such forms of assessment in terms of the demands of different disciplines, sub-disciplines, academic and professional fields, and will also need to recognise the significance of local settings and histories. at the simplest level we are asking: are we assessing the candidate or their research? and is this assessment formal, informal or a mixture of both? finally, the question of how the curriculum of doctoral education enhances employability is of key interest, and not only from the doctoral students’ point of view. academic communities – both universities and disciplinary organisations – have increasingly been concerned to support career development for early career researchers and diversify their employment opportunities, recognising that their training is often predicated on the assumption that will pursue an academic or research-only career (åkerlind, 2005). yet there is still little reliable evidence regarding the employability and the career pathways of ecrs, particularly in relation to careers in industry and other non-academic settings, and in the increasingly international labour market. this situation calls for a clearer understanding of multiple doctoral pathways and a review of curriculum structures within doctoral education that might facilitate diverse transitions. the tacit assumption of many supervisors, also implicit in many doctoral programmes and in the popular press (economist, 2010; cyranoski et al., 2011), is that the phd is a training ground for the next generation of academics. this encourages graduates to aspire to, and apply for, academic research positions (manathunga, pitt, & critchley, 2009), though only a minority will get a position in academia. this situation constricts the scope of academic training and skill development by focusing on a narrow range of labour market possibilities, and promotes a perception that many doctoral graduates have effectively “failed”. this problem accentuates the relevance of exploring the changing relationships between university and social and professional spheres (lee et al., 2009), and ensuring that ecrs are aware of and willing to pursue options other than the academic role. this in turn requires the development of new academic cultural practices (boud & tennant, 2006) based on a much clearer understanding of the ‘fit’ between the doctoral curriculum and the doctoral labour market. 5. developing a research agenda this paper has explored some emerging themes in doctoral education from a curricular perspective. this focus on the curriculum is significant not only because it might help to uncover existing tensions, but also because it allows us to face and reinterpret current challenges to doctoral education by undertaking intentional and meaningful change, especially in terms of the processes and outcomes of doctoral education. whilst recognising that knowledge and practices in this field are situated in historical and cultural contexts, we suggest here a number of possible themes for a future research agenda: 1. the diversity of training programmes developed for researchers around the world calls for a review. we need to improve our understanding of the historical context of current curriculum models and their impact on the training and experience of doctoral students and ecrs. 2. despite the extensive research already conducted on the changes in doctoral education, in terms of public policy, internationalisation, formats, etc. there is a need for more research on how these changes are being dealt with at the level of the formal, the informal and the hidden curriculum. !gonzalez)ocampo!et !al ! | f l r ! ! 32! 3. in order to avoid the unintended and perverse reproduction of inequalities, we need to explore the central role of departmental cultures and practices (involving both weak and strong elements of the hidden curriculum) in the integration and progression of doctoral students, and the diverse ways in which these are perceived by students from different backgrounds. 4. networking and professional socialisation have become increasingly important strategies in the development of doctoral students as researchers. these elements need to be explored as part of the doctoral curriculum, and supported by research on the roles of communities of practice and networks in supporting the construction of early career researchers’ identity. 5. in the light of the issues addressed in this paper, there is clearly a need for more research on the process of “becoming a supervisor”, and a review of the training and support available to doctoral supervisors and examiners. 6. assessment is a core curricular process in doctoral education, and yet there is very little research evidence on assessment practices (compared to, for example, the extensive literature on assessment in undergraduate education). our understanding of assessment needs to incorporate critical analysis of formal and informal practices and the variety of purposes which they fulfil. the fluidity of the “knowledge economy” presents new challenges to traditional forms of assessment, raising the possibility of replacing or extending traditional examinations with more flexible assessment models more appropriate to the diversity of ecrs’ academic and professional futures. 7. the current evidence on the destinations of ecrs illustrates the need for further research on the new relationships developing between universities and the labour market. from an international perspective there is a lack of evidence on the employability, career aspirations and mobility of ecrs, particularly those who do not follow academic careers. 8. the new demands of the labour market suggest a need to address the competencies of ecrs and a critical appraisal of the career pathways enabled through doctoral and postdoctoral education. this paper has been shaped very much by the interests and experiences of its diverse group of authors, and we recognise that consequently, any proposed research agenda is likely to be partial and incomplete. we welcome further discussion of the themes raised here and wider contributions to this important debate. !gonzalez)ocampo!et !al ! | f l r ! ! 33! keypoints the phd has become a “global academic passport”, although doctoral education practices are increasingly diverse; we argue for the need of an explicit discussion of what constitutes the “doctoral curriculum”, including its formal, informal and hidden dimensions. review of the doctoral curriculum should consider how phd students experience the curriculum, including identity as researchers, supervision, insertion in research networks, and the role of departmental cultures. review of the doctoral curriculum requires further research on assessment practices and the preparation of supervisors and examiners, and a consideration how these can be improved. review of the doctoral curriculum needs to take account of the multiple purposes of the phd and the divergent professional pathways of doctoral graduates, both inside and outside the academy. acknowledgments this work was funded (in part) by national funds through the fct – fundação para a ciência e a tecnologia (portuguese foundation for science and technology) within the strategic project of ciie, with the ref. “pest-oe/ced/ui0167/2014”. references åkerlind, g. s. (2005). postdoctoral researchers: roles, functions and career prospects. higher education research and development, 24(1) 21-40. åkerlind, g. s. (2009). postdoctoral research positions as preparation for an academic career. international journal for researcher development, 1(1), 84-96. anderson, s., & anderson, b. (2012). preparation and socialization of the education professoriate: narratives of doctoral student-instructors. international journal of teaching and learning in higher education, 24(2), 239–251. apple, m. w. (1971). the hidden curriculum and the nature of conflict. interchange, 2(4), 27–40. apple, m. w., & king, n. r. (1977). what do schools teach?. in r. h. weller (ed.), humanistic education (pp. 29-63). berkeley, ca: mccutchan. baker, v. l., & lattuca, l. r. (2010). developmental networks and learning: toward an interdisciplinary perspective on identity development during doctoral study. studies in higher education, 35(7), 807827. ball, s. (1989). la micropolítica de la escuela: hacia una teoria de la organización escolar. barcelona: paidós. ball, s. (1994). education reform: a critical and post-structural approach. buckingham: open university press. ball, s., bowe, r., & gold, a. (1992). reforming education and changing school: case studies in policy sociology. london and new york: routledge. berman, j., juniper, s., pitman, t., & thomson, c. (2008). reconceptualising post-phd research pathways: a model to create new postdoctoral positions and improve the quality of postdoctoral training in australia. australian universities' review, 50(2), 71-77. biggs, j. (2003). teaching for quality learning at university: what the student does (2nd ed.). buckingham: srhe and open university press. billett, s. (2009). conceptualizing learning experiences: contributions and mediations of the social, personal and brute. mind, culture and activity, 16(1), 32-47. !gonzalez)ocampo!et !al ! | f l r ! ! 34! boden, d., borrego, m., & newswander, l. k. (2011). student socialization in interdisciplinary doctoral education. higher education, 62(6), 741–755. boud, d., & lee, a. (eds.). (2009). changing practices of doctoral education. abbingdon: routledge. boud, d., & tennant, m. (2006). putting doctoral education to work: challenges to academic practice. higher education research & development, 25(3), 293–306. brennan, j., edmunds, r., houston, m., jary, d., lebeau, y., osborne, m., & richardson, j. t. e. (2009). improving what is learned at university: an exploration of the social and organisational diversity of university education. london: routledge. brew, a., & peseta, t. (2004). changing postgraduate supervision practice: a programme to encourage learning through reflection and feedback. innovations in education and teaching international, 41(1), 5-22. bronfenbrenner, u. (1979). the ecology of human development: experiments by nature and design. cambridge, ma: harvard university press. bronfenbrenner, u. (1986). ecology of the family as a context for human development: research perspectives. developmental psychology, 22(6), 723–742. brown, j. s., collins, a., & duguid, p. (1989). situated cognition and the culture of learning. educational researcher, 18(1), 32-41. calma, a., & davies, m. (2015). studies in higher education 1976-2013: a retrospective using citation network analysis. studies in higher education, 40(1), 4-21. clandinin, d. j., & connelly, f. m. (1995). teacher’s professional knowledge landscapes. new york: teachers college press. clausen, t., pohjola, m., sapprasert, k., & verspagen, b. (2012). innovation strategies as a source of persistent innovation. industrial and corporate change, 21(3), 553–585. colley, h., hodkinson, p., & malcolm, j. (2003). formality and informality in learning: a report for the learning and skills research centre. london: lsrc. cyranoski, d., gilbert, n., ledford, h., nayar, a., & yahia, m. (2011). the phd factory. nature, 472, 276279. deem, r., & o’brehony, k. (2000). doctoral students’ access to research cultures: are some more equal than others?. studies in higher education 25 (2) 149-165. deleuze, g., & guattari, f. (1980). mille plateaux: capitalisme et schizophrénie. paris: minuit. economist (2010). the disposable academic: why doing a phd is often a waste of time. the economist, dec 16. enders, j. (2002). serving many masters: the phd on the labour market, the everlasting need of inequality, and the premature death of humboldt. higher education, 44, 493-517. enders, j. (2004). research training and careers in transition: a european perspective on the many faces of the ph.d. studies in continuing education, 26(3), 419–429. eraut, m. (2000). non-formal learning, implicit learning and tacit knowledge. in f. coffield (ed.), the necessity of informal learning (pp. 12-31). bristol: policy press. eua (2007). doctoral programmes in europe’s universities: achievements and challenges. brussels: european universities association. fenwick, t., edwards, r., & sawchuk, p. (2012). emerging approaches to educational research: tracing the socio-material. london: routledge. fenwick, t., & nerland, m. (eds.). (2014). reconceptualising professional learning: sociomaterial knowledges, practices and responsibilities. london: routledge. gardner, s. k. (2007). ‘i heard it through the grapevine’: doctoral student socialization in chemistry and history. higher education, 54, 723–740. golde, c. m. (2005). the role of department and discipline in doctoral student attrition: lessons from four departments. journal of higher education, 76(6), 669–700. golding, c., sharmini, s., & lazarovitch, a. (2014). what examiners do: what thesis students should know. assessment & evaluation in higher education, 39(5), 563-576. greeno, j. g., collins, a. m., & resnick, l. b. (1996). cognition and learning. in d. c. berliner & r. c. calfee (eds.), handbook of educational psychology (pp. 15-45). new york: macmillan. !gonzalez)ocampo!et !al ! | f l r ! ! 35! guerin, c. (2013). rhizomatic research cultures, writing groups and academic researcher identities. international journal of doctoral studies, 8, 137–150. hager, p., lee, a., & reich, a. (eds.). (2012). practice, learning and change: practice-theory perspectives on professional learning. springer. hartley, j. (2000). nineteen ways to have a viva: appendix 2. psypag quarterly newsletter, 35, 22-28. ives, g., & rowley, g. (2005). supervisor selection or allocation and continuity of supervision: ph.d. students' progress and outcomes. studies in higher education, 30(5), 535–555. johnson, l., lee, a., & green, b. (2000). the phd and the autonomous self: gender, rationality and postgraduate pedagogy. studies in higher education, 25(2), 135-147. jones, m. (2013). issues in doctoral studies: forty years of journal discussion. where have we been and where are we going?. international journal of doctoral studies, 8(6), 83–104. kiley, m. (2009). rethinking the australia doctoral examination process. australian universities' review, 51(2), 32-41. kiley, m. (2014). coursework in australian doctoral education: what’s happening, why and future directions? final report. sydney: office for learning and teaching. kot, f. c., & hendel, d. d. (2012). emergence and growth of professional doctorates in the united states, united kingdom, canada and australia: a comparative analysis. studies in higher education, 37(3), 345-364. lave, j. (1988). cognition in practice: mind, mathematics, and culture in everyday life. cambridge, uk: cambridge university press. lave, j., & wenger, e. (1991). situated learning: legitimate peripheral participation. cambridge: cambridge university press. lee, a. (2008). how are doctoral students supervised? concepts of doctoral research supervision. studies in higher education, 33(3), 267-281. lee, a., brennan, m., & green, b. (2009). re-imagining doctoral education: professional doctorates and beyond. higher education research & development, 28(3), 275–287. lester, s. (2004). conceptualizing the practitioner doctorate. studies in higher education, 29(6), 757–770. lopes, a. c. (2012). a qualidade da escola pública: uma questão de currículo?. in m. taborda, l. faria filho, f. viana, n. fonseca, & r. lages (orgs.), a qualidade da escola pública no brasil (pp. 13-29). belo horizonte: mazza edições. lopes, a. c., & lópez, s. b. (2010). a performatividade nas políticas de currículo: o caso do enem. educação em revista, 26(1), 89-110. lopes, a. c., & macedo, e. (2011). contribuições de stephen ball para o estudo de políticas de currículo. in s. ball & j. mainardes (orgs.), políticas educacionais: questões e dilemas (pp. 249-283). são paulo: cortez. lopes, a., & pereira, f. (2012). everyday life and everyday learning: the ways in which pre-service teacher education curriculum can encourage personal dimensions of teacher identity. european journal of teacher education, 35(1), 17-38. malfroy, j. (2005). doctoral supervision, workplace research and changing pedagogic practices. higher education research & development, 24(2), 165-178. manathunga, c., pitt, r., & critchley, c. (2009). graduate attribute development and employment outcomes: tracking phd graduates. assessment & evaluation in higher education, 34(1), 91–103. margolis, e., & romero, m. (1998). the department is very male, very white, very old, and very conservative: the functioning of the hidden curriculum in graduate sociology departments. harvard educational review, 68(1), 1-33. mcalpine, l., & emmioğlu, e. (2014). navigating careers: perceptions of sciences doctoral students, postphd researchers and pre-tenure academics. studies in higher education, 1–17. mcalpine, l., & mckinnon, m. (2013). supervision the most variable of variables: students perspectives. studies in continuing education, 35(3), 265-280. mcalpine, l., amundsen, c., & turner, g. (2013). identity trajectory: reframing early career academic experience. british educational research journal, 40(6), 952-969. !gonzalez)ocampo!et !al ! | f l r ! ! 36! mcalpine, l., paulson, j., gonsalves, a., & jazvac-martek, m. (2012). untold doctoral stories in the social sciences: can we move beyond cultural narratives of neglect?. higher education research and development, 31(4), 511–523. mills, d., & paulson, j. (2014). making social scientists, or not? glimpses of the unmentionable in doctoral education. learning and teaching, 7(3), 73-97. morley, l., leonard, d., & david, m. (2002). variations in vivas: quality and equality in british phd assessments. studies in higher education, 27(3), 263-273. mullins, g., & kiley, m. (2002). it's a phd, not a nobel prize: how experienced examiners assess research theses. studies in higher education, 27(4), 369–386. nelson, r. (2006). practice-as-research and the problem of knowledge. performance research, 11(4), 105116. overall, n. c., deane, k. l., & peterson, e. r. (2011). promoting doctoral students research self-efficacy: combining academic guidance with academic support. higher education research & development, 30(6), 791–805. pacheco, j. a. (1996). currículo: teoria e praxis. porto: porto editora. pearson, m., evans, t., & macauley, p. (2008). growth and diversity in doctoral education: assessing the australian experience. higher education, 55(3), 357-372. pearson, m., kiley, m., evans, t., macauley, p., palmer, n., & pike, m. (2010). pathways to the phd in australia: a symposium. in m. kiley (ed.), quality in postgraduate research: educating researchers for the 21st century (p. 285). adelaide sa: cedam, the anu. pitcher, r., & åkerlind, g. s. (2009). post-doctoral researchers’ conceptions of research: a metaphor analysis. international journal for researcher development, 1(2), 160-172. print, m. (1987). curriculum development and design. sydney: allen & unwin. pyhältö, k., vekkaila, j., & keskinen, j. (2012). exploring the fit between doctoral students and ‘supervisors’ perceptions of resources and challenges vis-a-vis the doctoral journey. international journal of doctoral studies, 7, 395–414. pyhältö, k., stubb, j., & lonka, k. (2009). developing scholarly communities as learning environments for doctoral students. international journal for academic development, 14(3), 221-232. qaa. (2011). doctoral degree characteristics. the quality assurance agency for higher education. http://www.qaa.ac.uk/en/publications/documents/doctoral_characteristics.pdf rogoff, b. (1998). cognition as a collaborative process. in w. damon, d. khun, & r. s. siegler (eds.), handbook of child psychology (5th ed., vol. 2) (pp. 679–743). new york: wiley. roulston, k., preissle, j., & freeman, m. (2013). becoming researchers: doctoral students’ developmental processes. international journal of research & method in education, 36(3), 252-267. scaffidi, a. k., & berman, j. e. (2011). a positive postdoctoral experience is related to quality supervision and career mentoring, collaboration, networking and a nurturing research environment. higher education, 62(6), 685–698. solem, m. n., hopwood, n., & schlemper, b. (2011). experiencing graduate school: a comparative analysis of students in geography programs. professional geographer, 63(1), 1-17. thesis whisperer blog http://thesiswhisperer.com/ accessed 19 june 2015. trafford, v., & leshem, s. (2009). doctorateness as a threshold concept. innovations in education and teaching international, 46(3), 305-316. vaessen, m., van den beemt, a., & de laat, m. (2014). networked professional learning: relating the formal and informal. frontline learning research, 5, 56-71. vekkaila, j., pyhältö, k., & lonka, k. (2013). focusing on doctoral students’ experiences of engagement in thesis work. frontline learning research, 1(2), 10–32. virtanen, v., & pyhältö, k. (2012). what engages doctoral students in biosciences in doctoral studies?. the psychologist, 3(12a), 1231–1237. vitae (uk) researcher development framework. retrieved from: https://www.vitae.ac.uk/researchersprofessional-development/about-the-vitae-researcher-development-framework-planner accessed 19 june 2015. !gonzalez)ocampo!et !al ! | f l r ! ! 37! vygotsky, l. s. (1978). mind in society: the development of higher psychological processes. london: harvard university press. walker, g., golde, c., jones, l., bueschel, a., & hutchings, p. (2008). the formation of scholars: rethinking doctoral education for the twenty-first century. san fransisco: jossey-bass. wellington, j. (2013). searching for 'doctorateness'. studies in higher education, 38(10), 1490-1503. zhao, c-m., golde, c. m., & mccormick, a. c. (2007). more than a signature: how advisor choice and advisor behaviour affect doctoral student satisfaction. journal of further and higher education, 31(3), 263–281. gonzalez(ocampo-et -al | f l r ! 38! table 1 dimensions of various approaches to curriculum, and the specific themes/questions arising from them dimensions in 1 dimensions in 2 arising themes/questions 1 formal – 2 informal refers to qualifications frameworks, course syllabi aims and learning outcomes defined activities: workshops, supervision, seminars, conferences -includes regulations for candidature -e.g. affirmative action policies relates to what is really done through teaching and learning processes, such as readings and discussion, interactions with researchers in the context of classes. -activities: peer interaction, dialogues in academic community -impact of the social context in which training take place the role of academic practices in learning outcomes: -peer learning -social media (e.g. thesis whisperer blog) -departmental practices (e.g. golde, 2005) disciplinary networking (e.g. deem & brehony 2000) allocation of teaching duties/other work professional conventions/ expectations in particular subject areas 1 open – 2 hidden refers to such contents in doctoral training that are defined but variable in individual level, e.g., prescribed reading, research methods provision, seminars etc. which doctoral candidates are expected to attend. -learners’ degree of self-direction and the social context in which training take place embedded refers to unintended learning, often in regard to class and gender roles, social expectations, etc., that emerges from structures, relationships and practices in the educational setting, revealing the pedagogy of the learning context, rather than its intended content (apple 1971) what is students’ role (active/passive) in developing their doctoral training? -departmental practices (e.g. mills & paulson 2014) -dyadic dynamics in the supervisory relationship (including gender etc., plus reputational/prestige issues which are very intangible) (e.g. mcalpine & mckinnon, 2013, johnson et al., 2000) 1 standardised – 2 pluralized refers to systems such as phd programmes (i.e. with prescribed taught elements preceding thesis), and also skills programmes, e.g. vitae researcher development framework. -an inflexible system -intended learning outcomes laid down in policy documents (e.g. qaa) refers to a highly complex set of structures, practices and expectations from which doctoral students and their supervisors create new and unpredictable learning. -flexible -impact of the social context in which training take place -learners’ degree of embedded self-direction the purpose of doctoral degrees in relation to working career and employment (enders, 2004) microsoft word tulis et al_publication.docx         frontline learning research vol.4 no. 2 special issue (2016) 12– 26 issn 2295-3159   corresponding author: maria tulis, university of augsburg, department of psychology, universitätsstraße 10, 86159 augsburg, germany. email address: maria.tulis@phil.uni-augsburg.de doi: http://dx.doi.org/10.14786/flr.v4i2.168 learning from errors: a model of individual processes maria tulisa, gabriele steuera, markus dresela auniversity of augsburg, germany article received 4 may / revised 16 october / accepted 2 march / available online 6 april abstract errors bear the potential to improve knowledge acquisition, provided that learners are able to deal with them in an adaptive and reflexive manner. however, learners experience a host of different—often impeding or maladaptive—emotional and motivational states in the face of academic errors. research has made few attempts to develop a theory that focuses on learning from errors (with the exceptions of the theory of impasse-driven learning and the theory of negative knowledge) and, in particular, a theoretical framework that focuses on antecedent motivational processes. by integrating theories of self-regulated learning, volition, attributions, and appraisals, we propose a model that highlights individual processes that are characteristic of this specific learning phenomenon. more precisely, our theoretical framework aims to explain how emotional, motivational and self-regulatory processes—influenced by personal and contextual conditions—interact in order to facilitate or impede adaptive dealing with errors and appropriate metacognitions and cognitive activities. our objective is to provide a framework that allows for the systematic integration of various aspects that have been targeted in previous research and to guide and stimulate future research on learning from errors. as a first evidence for validation, we summarise research findings that address specific parts of the proposed model. keywords: learning from errors; self-regulation; motivation; emotion tulis  et  al       | f l r     13   1. learning from errors: a specific learning phenomenon in order to facilitate learning—the development of knowledge, metacognitive skills and autonomy— learners should be challenged with tasks that refer to skills and knowledge just beyond their current level of mastery (vygotsky, 1978). errors are a natural by-product of attempting challenging learning tasks and they may, in particular, provide learning opportunities (van lehn, 1988). recent research findings in educational psychology and contemporary cognitive psychology (e.g. cyr & anderson, 2014; van lehn, siler, murray, yamauchi, & baggett, 2003) give reason to revisit ancient wisdoms like “mistakes are the stepping stones for learning” or “you can always learn from your mistakes”. based on empirical findings, the consistent key argument is that errors initiate explanation and reflection processes in which deficient concepts are contrasted with correct concepts in order to establish accurate mental models (see also chi, 1996; kapur, 2008; oser & spychiger, 2005; siegler, 2002). however, as van lehn et al. (2003) put it, ”a learning opportunity is only an opportunity to learn”. accordingly, empirical findings consistently point to the importance of metacognitive support (e.g. keith & frese, 2005; künsting, kempf, & wirth, 2013). for example, westermann and rummel (2012) found that metacognitive support during student collaboration on difficult learning content and discussions of their wrong solutions lead to better learning outcomes. in addition to metacognitive processes, motivational processes obviously play a particularly important role for successful learning from errors. experiences of errors and impasses are accompanied by a host of different emotional and motivational states which facilitate or impede persistent learning engagement, the use of appropriate metacognitions, and cognitive activities. it can be assumed that poor learners are characterised by the experience of deactivating emotions following errors (for more details see section 2) and an inability to regulate their motivation and the respective emotions adaptively. in other words—as with learning in general (cf. kanfer & ackerman, 1989) but particularly after making errors—learning from one’s own errors through (self-) explanation basically requires motivational forces in order to persist after setbacks, to correct the error at hand, and to reflect on the underlying misconceptions. surprisingly,   educational research has paid little attention to learning from errors. a theoretical framework that addresses error-related learning processes in terms of emotional experiences, motivational changes, self-regulation, metacognitive activities, and cognitions is lacking. in order to explain why some learners show adaptive reactions and learning gains after errors while others fail to do so, such a model needs to simultaneously explain individual differences with motivational self-regulatory processes (inextricably bound to emotions) as well as the learners’ prerequisites and conditions (i.e. dispositions, motivational beliefs and orientations, knowledge, abilities or skills) in interaction with characteristics of the learning environment and the context. we propose a model with perceived errors as the events that initiate selfregulation. it systematically integrates personal determinants, contextual conditions and situational processes that are specific for learners dealing with errors. within this framework, we integrated components of previous models and built on the central assumptions of established theories (for another attempt to integrate different motivational theories, but not specifically adjusted to processes following errors, see de brabander & martens, 2014). in particular, we included models that contribute to explain individual processes following errors—all of them further addressed in the next sections: the transactional stress/coping model based on primary and secondary appraisals (lazarus & folkman, 1984), aspects of volition theory (kuhl, 1985, 2000), feedback loops (carver & scheier, 1998), self-regulation models (boekaerts, 2006; winne & hadwin, 1998) and theories of impasse or error-driven learning (de leeuw & chi, 2003; kolodner, 1983, 1997; minsky, 1997; oser & spychiger, 2005; van lehn, 1988). findings from studies on error management (heimbeck, frese, sonnentag, & keith, 2003; keith & frese, 2005) and error-related beliefs or attitudes (rybowiak, garst, frese & batinic, 1999; tulis & ainley, 2011; tulis, steuer & dresel, subm.) complete our proposed model. tulis  et  al       | f l r     14   1. 1 current state of research within the behaviouristic paradigm, and for a long time in the field of cognitive psychology, it was assumed that errors should be avoided because they would interfere with correct information and thus hinder the recall of correct answers (e.g. ayers & reder, 1998). in contrast, contemporary research provides empirical evidence for the fundamental role of errors in learning: overcoming impasses through reflection on errors and (self-) explanation of the underlying misconceptions has been shown to be important for learning progress since these processes help to establish accurate mental models (kapur 2008; mathan & koedinger, 2005; oser & spychiger, 2005; siegler, 2002; van lehn et al., 2003). based on a comprehensive literature review, we found different approaches that have been adopted in educational research to investigate the role of errors in learning: alongside research on classroom error management and error climate (e.g. tulis, 2013; steuer, rosentritt-brunn & dresel, 2013), individual responses to errors have been examined under different perspectives: for instance, there is a large body of research on (error) feedback and its impact on learning and achievement (for a meta-analysis see bangert-drowns, kulik, kulik & morgan, 1991; for an overview see mory, 1996). however, most of these studies did not address learning from errors per se. going deeper into this issue, a line of research has investigated students’ error patterns from a diagnostic perspective (for mathematics: clements, 1980; fiori & zuccheri, 2005; resnick, 1984) and has elaborated on error-types and taxonomies (e.g. frese & zapf, 1994). more recent studies have focused on learning from erroneous examples (eichelmann, narciss & schnaubert, 2013; große & renkl, 2007). for example, große and renkl (2007) found that incorrect solutions lead to enhanced learning outcomes if learners have favourable prior knowledge. including errors in worked examples motivated these learners to explain what was wrong and why, and it fostered elaborations on the correct solutions. their results underpin the positive relationship between transfer performance and the generation of self-explanations when learning with incorrect solutions. other researchers have concentrated on learning from errors with (intelligent) tutors (mathan & koedinger, 2005; van lehn et al., 2003). mathan and koedinger (2005) focused on learners’ error-detection and errorcorrection skills and how these can be supported. the authors provide evidence that feedback which allows students to detect, correct and reflect on their own errors fosters learning at a faster rate, conceptual understanding, and (transfer) performance. similarly, but in another setting (i.e. collaborative learning environments), research on productive failure has emphasised the benefits of delaying instruction in order to enable reflection on incorrect solution attempts by students (kapur, 2008; westermann & rummel, 2012). van lehn and colleagues (2003) investigated the conditions of successful learning episodes within their framework of impasse-driven learning. in particular, they studied tutorial dialogues between students and expert tutors. the results suggest that impasses and errors are strongly associated with learning. reaching impasses and clarifying errors turned out to have stronger effects on effective learning than when a tutor modelled the correct action. finally, some researchers have addressed learners’ attitudes towards making errors (rybowiak et al., 1999) and implemented the positive function of errors for learning in a training condition (gully, payne, koles & whiteman, 2002; kanfer & ackerman, 1996; keith & frese, 2005). in these studies, the positive function of errors was prompted to participants while practising a task and the participants were encouraged to make errors. however, error-trainings had better effects on performance if they were combined with instructions providing metacognitive techniques supporting cognitive and emotional self-regulation (keith & frese, 2005) or if individuals were higher in ability, higher in openness to experience, or lower in conscientiousness (gully et al., 2002). in summary, there is a growing research interest in the specific phenomenon “learning from errors, but a theoretical framework that allows an integration of these different perspectives is lacking. in addition to the above-outlined findings regarding the individual preconditions and their interaction with training efforts to enhance successful learning from errors, learners’ adaptive reactions to errors—their antecedents and consequences—have been considered to a minor degree. particularly little attention has been paid to differences in learners’ emotional and motivational responses to errors and their significance for subsequent learning processes. in this regard, we present four different theoretical perspectives on dealing with errors in tulis  et  al       | f l r     15   learning contexts in the following sections. their theoretical assumptions build the basis for our proposed model described afterwards. 1. 2 perspectives on individual dealing with errors a first perspective to explain individual differences in learners’ reactions to errors can be derived from research on stress and coping (cf. boekaerts, 2010). lazarus and folkman (1984) proposed two cognitive appraisal processes which determine if a situation is perceived as stressful. first, in a primary appraisal process an (error) situation is interpreted along a continuum ranging from irrelevant, benignpositive, not harmful to challenging, threatening or harmful. the secondary appraisal process further evaluates the situation and determines which coping resources are available and whether the individual can apply them effectively. finally, the situation and coping strategies are monitored and evaluated, and the primary and secondary appraisals are modified if necessary. numerous studies have shown that appraisal processes—operating automatically or conscious and volitional—determine emotional experiences (lazarus, 1991). altogether, appraisal theory appears to constitute a proper basis for describing emotional states, motivational changes and self-regulatory processes following errors. a second perspective that is strongly related to learners’ reactions to errors stems from research on reactions to (success and) failure. literature review reveals an impressive body of research that has been proven to explain differences in individuals’ (affective) reactions to failure based on different theoretical foundations (for an overview see elliot & dweck, 2005), such as achievement goal theory (for an overview see maehr & zusho, 2009), or attribution theory (weiner, 1986). for example, mastery-oriented students with a focus on skill development and individual improvement do not necessarily feel threatened by failure when faced with a difficult task, but rather perceive setbacks as an opportunity for learning and mastery (e.g. dweck & leggett, 1988). causal beliefs of the importance of effort for success were found to mediate the relationship between mastery orientation and retained positive affect after errors were made (tulis & ainley, 2011). in contrast, performance avoidance goals have been shown to be associated with increased negative affect following failure experiences and lower preference for difficult tasks (e.g. elliott & dweck, 1988). clifford’s (1984) theory of constructive failure also emphasised learners’ differences in affective experiences following errors: students who are focused on the task rather than on themselves were less likely to fear failure and to feel negative emotions. rather, they were more likely to invoke positive thoughts and further appraisals of “challenge” (cf. boekaerts, 1993). finally, volition theory (kuhl, 1985, 2000) has broached the issue of the interplay between emotion, motivation, metacognition, and cognition in the face of failure: besides cognitive control—in terms of metacognitive activities directed towards keeping attention and effort on the task—emotion and motivation control (i.e. self-regulatory processes to keep negative emotions and other intrusive thoughts at bay during task engagement) can be assumed to mediate the effectiveness of learning from errors. in the context of learning situations, empirical studies by kanfer and ackerman (1996) provide evidence that emotion control is most critical when the task is likely to appear most daunting to the learner—a likely situation after making errors. important to note is that, although failure and errors are interrelated constructs, they are not the same: errors are usually defined as an unintended discrepancy between a current and a desired state, or as a deviation from a given standard (e.g. frese & zapf, 1994). “failure” implies more than just this perceived discrepancy. in contrast to errors, failure experiences constitute a more global miss of a goal with a greater focus on the subsequent consequences (cf. zhao & olivera, 2006). above all, not every error is necessarily interpreted as failure. whether an error is evaluated as failure or not depends on situational aspects (e.g. social norms) and personal characteristics of the learner, such as self-concept of ability. bandura (1997), as well as eccles and her colleagues (e.g. eccles & wigfield, 2002), concluded that efficacy expectations or perceptions of self-competence are a major determinant of a person’s willingness to invest more effort if the task becomes challenging—hence, also following errors. thirdly, regarding theories on learning from errors in a narrower sense, a perspective on dealing with errors stems from organisational psychology and technology based learning. within the field of organisational psychology, rather economised working models developed for empirical studies in the field of tulis  et  al       | f l r     16   workplace learning have been proposed (zhao, 2011; bauer, gartmeier, & harteis, 2012; van dyck, van hooft, de gilder, & liesveld, 2010). researchers have either primarily focused on personal characteristics that may facilitate or impede effective learning from errors at work (e.g. components of an error specific attitude, rybowiak et al., 1999) or they have focused exclusively on contextual features, such as the organisational error climate. as an exception, oser and colleagues (e.g. oser & spychiger, 2005) introduced the concept of “negative knowledge” (cf. minsky, 1997) in the context of academic learning. it represents knowledge about false facts and inappropriate action strategies that labels incorrect concepts as wrong and helps to prevent the repetition of errors in similar situations. similarly, kolodner (1983, 1997) emphasised that individuals use knowledge about formerly experienced errors in new situations. comparably, van lehn (1988) suggested that impasses pave the way for learning from the subsequent explanation and therefore are even necessary for learning processes. however, these theoretical explanations primarily consider cognitive processes, and they  do not cover emotional, motivational, and self-regulatory processes following errors as antecedents of successful learning from errors. in order to bridge this gap, contemporary models of self-regulated learning which propose recursive processes including emotional/motivational functioning as well as metacognitions and cognitive activities (e.g. boekaerts, 1999; pintrich, 2000; schmitz, 2001; zimmerman, 2008) appear to constitute a fourth perspective on differences in learners’ reactions to errors. more specifically, we explicate three selfregulation models in the following which provide a proper basis for describing motivational and selfregulatory processes following errors with different focal points (boekaerts & niemivirta, 2000; carver & scheier, 1998; winne & hadwin, 1998). first, the “dual processing self-regulation model” (boekaerts & niemivirta, 2000; boekaerts, 2006) provides a framework addressing the importance of affective experiences and the learners’ competences to regulate their emotions and motivation following errors. two main goal priorities which are pursued by selfregulative activities are distinguished: (1) the “mastery/growth pathway” and (2) the “well-being pathway”. learners who want to reach a specific subgoal in order to improve skills or gain knowledge (e.g. analyse the causes of the error at hand) initiate activities in the mastery/growth pathway because they value that goal and feel competent enough to commit energy to its pursuit. on the other hand, learners who are primarily concerned with the anticipated threat to their self-worth and the negative consequences of errors initiate activities in the well-being pathway. importantly, it is assumed that learners can switch to the mastery/growth pathway by using adaptive emotional and motivational regulation strategies (boekaerts, 2006). another theoretical model that can be used to explain both learners’ emotional as well as behavioural changes following errors was introduced by carver and scheier (1998). the authors have focused on the role of feedback control processes during self-regulation. the core construct in their model is a discrepancyreducing feedback loop (or a discrepancy-enlarging loop in the case of an avoidance situation): if a discrepancy between a current state/situation (input function) and a goal/standard (reference value) is detected, adjustments are made in an output function in terms of behavioural changes. for example, a learner may invest more effort to identify the error causes or she may seek further information in the learning material after the perception of an error. parallel to this behaviour-guiding loop, carver and scheier (1998) described the affect-creating feedback loop which operates automatically and simultaneously. it is assumed to monitor the rate of progress of behaviour discrepancy reduction over time. hence, the theoretical model of carver and scheier (1998, 2013) provides an appropriate framework for behavioural reactions as well as the origins and functions of emotions that are experienced after errors. finally, the model suggested by winne and hadwin (1998) highlights the ongoing evaluation of potential discrepancies between products and standards of the learning process. in their model, the authors describe   four basic phases—task definition, goal setting and planning, studying tactics, and adaptations to metacognition (for an overview see also perry & winne, 2006)—in terms of the interaction of personal and contextual conditions, products (i.e. learning behaviour and outcomes) compared with standards (i.e. the optimal end state of each phase) and the learner’s goals through metacognitive evaluation processes all these tulis  et  al       | f l r     17   aspects are types of information that are used or generated during learning. a mismatch between products and standards is assumed to initiate further learning operations, the use of metacognitive strategies and/or the revision of the conditions and standards. the output, or performance, is the result of recursive processes that cascade back and forth, altering conditions, standards, operations, and products as needed. thus, the model represents a “recursive, weakly sequenced system” (winne & hadwin, 1998, p. 281) and it primarily addresses cognitive and metacognitive activities. therefore, it perfectly augments the proposed theoretical framework for learning from errors presented in the next section, which considers not only emotional and motivational but also cognitive and metacognitive processes and learning activities. in summary, most of the theories outlined above focus on self-regulatory processes in general, but a sufficiently elaborated model with respect to perceived errors as initiating points for self-regulation is lacking.   previous research was not able to adequately explain individual differences in error-specific emotional and motivational self-regulation. regarding prior research that aimed to investigate learning from errors, it is striking that self-regulatory processes—in particular motivational processes—have only been addressed sketchily, although there is a common agreement on their importance in the face of setbacks. a theoretical perspective for relating personal and contextual conditions and motivational (self-regulatory) processes following errors is needed to explain and systematically investigate error-related learning phenomena. previous empirical findings suggest that such a model must address three issues: first, affective and motivational reactions to errors, as well as cognitive and behavioural reactions specifically adjusted to the error in question have to be included (dresel, schober, ziegler, grassinger, & steuer, 2013; tulis, grassinger, & dresel, 2011). secondly, such a theory must take into account critical characteristics of errorsituations and indicate how these characteristics affect the potential contribution of individual dispositions, orientations, abilities or skills and current motivational states to adaptive learning behaviour. finally an integrated framework must consider the effects of interactions between personal determinants, contextual conditions, and situational processes. prior research and working models have tended to focus either on personal preconditions or the context, depending on the researcher’s primary concern. we attempt to overcome these existing shortcomings and to expand previous approaches by providing a framework which integrates proven theories of self-regulation, volition, motivation, emotion, and cognition. 2. individual reactions to and learning from errors: a process model the purpose of this section is to introduce a framework that includes the above presented theoretical perspectives and to provide a model which can provide an explanation of individual differences, situational influences and sequenced processes following error-experiences as antecedents for successful learning from errors (see figure 1). learning from errors is an effortful activity. our understanding of learning from errors includes a detailed analysis of the error causes in order to identify and explain potential misconceptions, a self-evaluation of the underlying knowledge and its modification, as well as the correction of the error in question (e.g. dresel, et al., 2013). prior to these metacognitions and cognitive activities, learners have to deal with changes in affect and motivation after the perception of an unintended discrepancy between a current state and a desired outcome or a given standard. more specifically, the perception of an error represents the (“bottom-up”) starting point in our model (see “error feedback/detection of an error” marked with an asterisk in figure 1) which induces a sequence of processes (indicated with bold arrows in figure 1). yet irrespective of the type of error or its causes, this is assumed to trigger direct reactions in terms of affect based on primary appraisals of the situation (see “direct reactions towards errors” in figure 1). in line with lazarus (1991), primary appraisals are directed towards the assessment of the relevance of this unintended discrepancy/goal incongruence to the learner and subsequent affect acts as a signal for this personally relevant deviation from an implicit or explicit standard. based on these primary appraisals, different emotions such as surprise, frustration, anger or boredom may be tulis  et  al       | f l r     18   experienced. for example, a self-confident, high achieving learner may first experience surprise after error feedback, whereas a low achiever may experience frustration at first sight of an error—in the event that both learners value the task and aim to master the task. this first emotional reaction might be vague, maybe not as easy to categorize as a specific emotion, but in any case we would expect an observable change in arousal. we assume that primary reactions are followed by more indirect reactions towards the error at hand (see “indirect/secondary reactions towards errors” in figure 1) including secondary appraisals directed at the assessment of controllability and personal resources to deal with the error (cf. lazarus, 1991). it is further assumed that these secondary appraisal processes change or intensify the primary emotional reaction and the learners’ subsequent motivation. analogous to top-down processing (i.e. knowledge or expectations are used to guide processing), further selfand task related appraisals, such as causal attributions (weiner, 1986) are made. these, in turn, might evoke attribution-dependent emotions other than the learner’s primary emotional states—or the learner’s primary emotional states might be intensified. it can be expected that not all types of errors lead to the same processes and subsequent learning. for example, in contrast to careless mistakes (e.g. slips, caused by attentional problems, or lapses, caused by memory failures) only knowledgeand rule-based errors might bear a potential for learning (for this taxonomy see reason, 1990). the nature of emotional and motivational changes is likely affected by the type of the error at hand. thus, at this stage of the model, we presume that the error-type has an impact on the learners’ secondary appraisals, the subsequent selfand task-related motivation, and the respective learning actions. in the next step (see “emotional and motivational regulation” in figure 1), these changes in self and task-related motivation—and emotional states—are assumed to trigger emotional and motivational regulation processes (cf. boekaerts, 2003). depending on personal characteristics of the learner, these errorrelated regulation processes may become necessary to maintain learning motivation. for example, overthinking the value of the task, the use of social resources, efficacy self-talk, or cognitive reappraisal may help to reassure the learner to proceed with the task despite setbacks (e.g. wolters, 2003). some learners may be more concerned with emotion-focused coping (lazarus, 1993) to avoid a threat to self-worth and restore their well-being (cf. “well-being pathway”, boekaerts, 2006), others may focus on strategies to re-direct attention and learning activities in order to master the task (boekaerts, 2006; kuhl, 2000). hence, we assume that learners actively (i.e. consciously or automated) use emotional and motivational regulation strategies following errors to activate and sustain their cognitive, metacognitive and affective functioning (butler & winne, 1995; wolters, 2003). we presume that adaptive (and effective) emotional and motivational selfregulation (gross, 1998; schwinger, steinmayr & spinath, 2009; wolters, 1998) provides the basis for the use of appropriate metacognitive activities, cognitive strategies and learning behaviour to adequately reflect on the underlying misconception (subsumed under “learning process” in figure 1). however, the regulation strategies that learners may use can also be dysfunctional: the use of maladaptive strategies following errors, such as distraction, suppression or rumination (e.g. gross, 1998; knollmann, 2006) may impede a detailed self-explanation of errors and their respective correct counterparts. furthermore, it can be assumed that—as for the use of learning strategies—some regulation strategies may be appropriate for certain learning contexts whereas the same strategies might be dysfunctional in other contexts (engelschalk, steuer & dresel, 2015). in any case, inappropriate or failed regulation strategy use may result in the experience of negative deactivating emotions, such as hopelessness or boredom, which are held to be detrimental for motivation (e.g. pekrun, goetz, daniels, stupnisky, & perry, 2010). thus, emotions are not only assumed to act as a signal after the perception of a discrepancy, but they are also assumed to be an indicator of the learners’ current motivation. consequently, they guide subsequent learning behaviour (e.g. in terms of persistence, attention focus, or information seeking) and they serve as a monitoring instrument for goal pursuit (carver & scheier, 1990). hence, emotions are assumed to moderate learning processes and we regard the presence of activating (or epistemic) emotions as a necessary condition for persistent task engagement in the face of obstacles and for learning from errors in general. tulis  et  al       | f l r     19   figure 1. process model of individual reactions to and learning from errors. it is important to note that individuals’ learning behaviour following errors, their emotional and motivational experiences and regulation strategies, and their subsequent metacognitions and cognitive activities are all assumed to be influenced by personal characteristics as well as contextual features which interact continuously with one another throughout the entire learning process. learners continuously appraise the learning conditions against the background of their individual dispositions, skills and abilities (e.g. prior knowledge or topic-interest), and their motivational beliefs such as self-concept of ability or goal orientation (for an overview see schunk, meece & pintrich, 2013). previous findings indicate that the effectiveness of error encouragement training might depend on such individual differences (e.g. gully et al., 2002). contextual conditions include characteristics of the task (e.g. an enquiry-based learning task versus a routine task), the learning context (e.g. practice versus testing situation), and the interpersonal aspect of dealing with errors in social learning environments which may facilitate or impede learning from errors (i.e. error climate). although located at the starting point in our model, personal and contextual conditions impact later processes as well (indicated with dashed arrows in figure 1). their interaction is affected by previous learning experiences and outcomes which are integrated in a broader social and cultural context. “learning from errors” (marked with an asterisk in figure 1) takes place in terms of reflection and self-explanation processes based on respective metacognitive activities, the use of appropriate cognitive strategies, and learning behaviour adapted to the new situation (see “learning process” in figure 1). finally, this should result in the modification of the underlying knowledge, improved skills and performance gains (see “learning outcome” in figure 1) which are expected to have reciprocal effects on the learners’ personal tulis  et  al       | f l r     20   conditions, and hence on the interpretation of subsequent error-situations (indicated with backwards directed arrows in figure 1). in order to validate the proposed processes, different stages/components of the model and their relations or sequenced effects need to be analysed: in particular, (1) the impact and interplay between different personal and situational conditions on individual reactions to errors and the use of error-specific adaptive regulation strategies, (2) the proposed changes in motivation and emotion and their function for further self-regulation and learning behaviour, (3) the relevance of metacognitive and cognitive activities for error-related learning processes, and (4) the necessity of affective-motivational functioning to provide a basis for such activities. 3. empirical evidence and open research questions so far we have provided some evidence for the assumed functions of emotions at the stage pertaining to direct/primary and indirect/secondary reactions towards errors, the use of error-related regulation strategies, and the influence of selected personal conditions and contextual factors on individual responses to errors (dresel, et al., 2013, steuer et al., 2013; tulis, 2013; tulis & ainley, 2011; tulis & dresel, 2013; tulis & fulmer, 2013; tulis et al., 2011; tulis et al., subm.). in the present section we summarise the findings of three studies with different foci, namely individual determinants of adaptive dealing with errors (study 1), motivational (self-regulation) processes following errors and their impact on subsequent learning behaviour (study 2), and the dimensions of error climate and their impact on students’ responses to errors (study 3). study 1—located in the “person í situation” part of figure 1—focused on individual components that may facilitate a learner’s adaptive reaction to errors (tulis et al., subm.). previous studies (e.g. dresel et al., 2013; tulis & ainley, 2011; tulis et al., 2011) have already demonstrated a positive relationship between students’ (more stable) motivational orientations (i.e. positive self-concept of ability, mastery goal orientation, adaptive error-related beliefs) and emotional/motivational reactions following errors. based on these results tulis et al. (subm.) tested a tripartite classification of adaptive individual dealing with errors in terms of a cognitive, an affective, and a behavioural component. more specifically, the authors analysed the distinctiveness of 614 students’ self-reported beliefs about errors as learning opportunities from students’ affective-motivational reaction tendencies that facilitate persistence and engagement despite setbacks, and students’ behavioural reaction tendencies including metacognitive activities and error-related learning behaviour (dresel et al., 2013). the results—obtained with confirmatory factor analyses demonstrating a good fit to the data—provided evidence for three distinct factors. in addition, the authors analysed their relationship to other motivational beliefs, and, whether these components of adaptive individual dealing with errors may differ between the scholastic domains of mathematics, english and german. correlational findings suggested domain-specificity for the three components. thus, error-related beliefs, habitualised affective-motivational and behavioural responses to errors might be acquired domain-specifically. further research is needed before any conclusions can be drawn, but the results point to the likelihood that students’ dealing with errors may not be differentiated along the verbal and mathematics continuum as, for instance, the academic self-concept is (e.g. marsh, walker & debus, 1991). furthermore, the findings emphasise the differentiation of personal conditions in terms of rather proximal beliefs in addition to less error-specific motivational beliefs, such as mastery goal orientation. study 2 (tulis & dresel, 2013, august) addressed motivational changes (see “emotional and motivational regulation” in figure 1), undergraduate students’ motivational and emotional self-regulation following errors and its effects on learning behaviour (see “learning process” in figure 1) in a computerbased learning setting. data were collected during two time intervals—the same study design was implemented in both studies with the only difference that in study 2a, we additionally conducted stimulated tulis  et  al       | f l r     21   recall interviews immediately after the learning session to examine participants’ use of various emotional and motivational regulation strategies following errors whereas -self-reported regulation strategies were assessed on-task after error feedback in study 2b. regarding the hypothesised motivational changes after making errors measured with on-task state items, we found a substantial decrease in students’ motivation. repeated-measures manovas with the three components of adaptive individual dealing with errors (high versus low levels) as between-subject factors indicated a stronger decline in task-related motivation, situational interest and enjoyment, and perceived competence for students’ low in action adaptivity, affective motivational adaptivity, and adaptive beliefs about errors, respectively. interview data pointed out a rich variety of different emotional and motivational regulation strategies that are used following errors, ranging from “proximal goal setting”, and the “use of social resources” (i.e. asking someone for help; most reported) to “self consequating” (i.e. motivating oneself by self-reinforcement for having reached a particular goal; least reported). in addition, also maladaptive strategies, such as “rumination” and “suppression” were reported in study 2a. however, when measured on-task (study 2b)—immediately after error feedback— appraisal based strategies, such as cognitive reappraisal (i.e., having a positive view on making errors as a natural part of learning) and mastery self-talk (i.e., thinking of the potential of errors for personal improvement) were more prominent, as was the use of maladaptive strategies. thus, according to our model, a decrease in student motivation triggered the use of emotional and motivational regulation strategies, and personal characteristics served as a buffer. regression analyses further emphasized differential associations between these strategies and adaptive learning activities after error feedback: mastery self-talk and reappraisal were found to facilitate an in-depth analysis of the error at hand, whereas distraction negatively predicted the reflection of the underlying misconceptions. logistic regression results indicated positive associations between proximal goal setting and students’ actual persistence. in summary, our findings emphasised the importance of motivational self-regulation for subsequent engagement following errors, and hence the proposed function of emotional and motivational regulation strategies for subsequent learning processes and learning behaviour. furthermore, they provided first evidence for a differentiation between adaptive and maladaptive error-related strategies. the findings of study 3 (steuer et al., 2013) are based on a questionnaire-study with 1,116 students from 56 sixth and seventh grade classrooms. this study focused on contextual conditions and it is located in the “person í situation” part as well as related to the dashed arrows of figure 1 that indicate the influence of characteristics of the learning environment on individual learning behaviour. study 3 provided evidence for eight theoretically and empirically distinguishable subdimensions of error climate and their impact on students’ individual dealing with errors. steuer et al. (2013) further demonstrated that classroom error climate has an impact on students’ affective-motivational and action adaptivity of error reactions, which, in turn, were positively associated with students’ self-reported effort. hence, according to our proposed model, the results supported the assumed association between personal conditions and characteristics of the social learning environment as well as their influences on individual learning behaviour following errors. 4. contribution to theory development and implications for future research taken together, our findings corroborate the assumed interplay between personal and contextual conditions as well as the importance of functional emotional and motivational self-regulation for adaptive dealing with errors. supported by some preliminary empirical evidence, the proposed model provides a more complete understanding of the motivational processes following errors in interaction with personal and contextual conditions. it gives several indications of how learners’ adaptive reactions to errors—a necessary precondition for learning from errors—can be supported. however, besides motivational processes, further research is needed to address the cognitive processes specifically related to effective learning from errors (see “learning process” in figure 1). research findings on conceptual change, cognitive conflicts, impasse driven learning and productive failure may provide the basis for further investigations using on-task tulis  et  al       | f l r     22   measurements (e.g. eye-tracking). another important issue raised by previous findings (e.g. keith & frese, 2005) concerns metacognitive activities—also part of the proposed framework that needs to be specified in future research. finally, our findings raised some methodological issues for future research: retrospective measurements (even if the time interval is short) might not offer adequate insights into actual and transient task-specific regulation processes, especially strategies involving cognitive change. therefore, future studies should differentiate between strategies learners tend to use to regulate their motivation during learning (e.g. assessed with questionnaires) and the actual strategy learners use (measured on-task). in summary, our model contributes to current research on motivation in several ways: (1) it expands current theories of self-regulated learning because it highlights perceived errors as initiating points for selfregulatory processes, (2) it provides a solid foundation for the analysis of motivational processes compatible with almost all contemporary theories of motivation, and (3) it enables the examination of personal, contextual and situational conditions and their interactions as well as their potential impact on error-related learning processes. finally, our proposed model provides a unified framework specifically adjusted to the phenomenon of learning from errors—a growing but still barely investigated field of educational research. previous findings and future research can be easily integrated into the present framework in order to specify the antecedents and processes of effective learning from errors. finally, the major implication for the future research practice is the process-related view on learning. keypoints a theoretical framework specifically adjusted to the phenomenon of learning from errors is introduced. changes in motivation trigger emotional and motivational self-regulation processes. individual differences are explained by personal and situational conditions, emotional, motivational, metacognitive and cognitive processes. contemporary theories of motivation are integrated in the model. antecedents and processes of successful learning from errors are specified. references ayers, m. s., & reder, l. m. (1998). a theoretical review of the misinformation effect. predictions from an activation-based memory model. psychonomic bulletin and review, 5, 1–21. doi:10.3758/bf03209454. bandura, a. (1997). self-efficacy: the exercise of control. new york: w. h. freeman. bangert-drowns, r. l.; kulik, c.; kulik, j. a. & morgan, m. t. (1991). the instructional effect of feedback in test-like events. review of educational research, 61, 213–238. doi:10.3102/00346543061002213. bauer, j., gartmeier, m., & harteis, c. (2012). human fallibility and learning from errors at work. in j. bauer & c. harteis (eds.), human fallibility: the ambiguity of errors for work and learning (pp. 155– 169). dordrecht: springer. doi:10.1007/978-90-481-3941-5. boekaerts, m. (1993). being concerned with well-being and with learning. educational psychologist, 28, 149–167. doi:10.1207/s15326985ep2802_4. boekaerts, m. (1999). self-regulated learning: where we are today. international journal of educational research, 31, 445–457. doi:10.1016/s0883-0355(99)00014-2. boekaerts, m. (2003). towards a model that integrates motivation, affect, and learning. british journal of educational psychology monograph, series ii (2), development and motivation: joint perspectives, 173–189. boekaerts, m. (2006). self-regulation and effort investment. in e. sigel & k. a. renninger (eds.), handbook of child psychology, child psychology in practice, 4 (pp. 345–377). new jersey: wiley. tulis  et  al       | f l r     23   boekaerts m. (2010), coping with stressful situations: an important aspect of self-regulation. in p. peterson, e. baker, & b. mcgaw (eds.), international encyclopedia of education (pp. 570–575). oxford: elsevier. boekaerts, m., & niemivirta, m. (2000). self-regulation in learning: finding a balance between learning goals and ego-protective goals. in m. boekaerts, p.-r. pintrich, & m. zeidner (eds.), handbook of selfregulation (pp. 417–450). san diego: academic press. doi:10.1016/b978-012109890-2/50042-1. butler, d. l., & winne, p. h. (1995). feedback and self-regulated learning: a theoretical synthesis. review of educational research, 65, 245–281. doi:10.3102/00346543065003245. carver, c. s., & scheier, m. f. (1990). origins and functions of positive and negative affect: a control process view. psychological review, 97, 19–35. doi:10.1037/0033-295x.97.1.19. carver, c. s. & scheier, m. f. (1998). on the self-regulation of behavior. new york: cambridge: university press. doi:10.1017/cbo9781139174794. carver, c. s., & scheier, m. f. (2013). goals and emotion. in m. d. robinson, e. r. watkins, & e. harmonjones (eds.), guilford handbook of cognition and emotion (pp. 176–194). new york: guilford press. chi, m. t. h. (1996). constructing self-explanations and scaffolded explanations in tutoring. applied cognitive psychology, 10, 33–49. doi:10.1002/(sici)1099-0720(199611)10:7%3c33::aidacp436%3e3.3.co;2-5. clements, m. a. (1980). analyzing children's errors on written mathematical tasks. educational studies in mathematics, 11, 1–21. doi:10.1007/bf00369157. clifford, m. m. (1984). thoughts on a theory of constructive failure. educational psychologist, 19, 108–120. doi:10.1080/00461528409529286. cyr, a.-a., & anderson, n. d. (2015). mistakes as stepping stones: effects of errors on episodic memory among younger and older adults. journal of experimental psychology: learning, memory, and cognition. doi:10.1037/xlm0000073. de brabander, k., & martens, r. (2014). towards a unified theory of task-specific motivation. educational research review, 11, 27–44. doi:10.1016/j.edurev.2013.11.001. de leeuw, n., & chi, m. t. h. (2003). self-explanation, enriching a situation model or repairing a domain model? in g. sinatra & p. r. pintrich (eds.), intentional conceptual change (pp. 55–78). mahwah: new jersey: erlbaum. dresel, m., schober, b., ziegler, a., grassinger, r., & steuer, g. (2013). affektiv-motivational adaptive und handlungsadaptive reaktionen auf fehler im lernprozess [affective-motivational adaptive and action adaptive reactions on errors in learning processes]. zeitschrift für pädagogische psychologie, 27, 255– 271. doi:10.1024/1010-0652/a000111. dweck, c. s., & leggett, e. l. (1988). a social-cognitive approach to motivation and personality. psychological review, 95, 256–273. doi:10.1037/0033-295x.95.2.256. eccles, j. s., & wigfield, a. (2002). motivational beliefs, values, and goals. annual review of psychology, 53, 109–132. doi:10.1146/annurev.psych.53.100901.135153. eichelmann, a., narciss, s., & schnaubert, l. (2013, august). learning from errors through tasks-withtypical-errors. paper presented at the 15th biennial conference of the european association for research on learning and instruction (earli), munich, germany. elliott, e., & c. dweck. (1988). goals: an approach to motivation and achievement. journal of personality and social psychology, 54, 5–12. doi:10.1037/0022-3514.54.1.5. elliot, a. j., & dweck, c. s. (eds.). (2005). handbook of competence and motivation. new york, ny: guilford. engelschalk, t., steuer, g. & dresel, m. (2015). wie spezifisch regulieren studierende ihre motivation bei unterschiedlichen anlässen? ergebnisse einer interviewstudie [situation-specific motivation regulation: how specifically do students regulate their motivation for different situations?]. zeitschrift für entwicklungspsychologie und pädagogische psychologie, 47, 14–23. doi:10.1026/00498637/a000120. fiori, c., & zuccheri, l. (2005). an experimental research on error patterns in written subtraction. educational studies in mathematics, 60, 323–331. doi:10.1007/s10649-005-7530-6. tulis  et  al       | f l r     24   frese, m., & zapf, d. (1994). action as the core of work psychology: a german approach. in h. c. triandis, m. d. dunette, & l. m. hough (eds.), handbook of industrial and organizational psychology (vol. 4, pp. 271–340). palo alto: consulting psychologists. gross, j. j. (1998). the emerging field of emotion regulation: an integrative review. review of general psychology, 2, 271–299. doi:10.1037/1089-2680.2.3.271. große, c. s., & renkl, a. (2007). finding and fixing errors in worked examples: can this foster learning outcomes? learning and instruction, 17, 612–634. doi:10.1016/j.learninstruc.2007.09.008. gully, s. m., payne, s. c., koles, k. l., & whiteman, j.-a. k. (2002). the impact of error training and individual differences on training outcomes: an attribute–treatment interaction perspective. journal of applied psychology, 87, 143–155. doi:10.1037/0021-9010.87.1.143. heimbeck, d., frese, m., sonnentag, s., & keith, n. (2003). integrating errors into the training process: the function of error management instructions and the role of goal orientation. personnel psychology, 56, 333–362. doi:10.1111/j.1744-6570.2003.tb00153.x. kanfer, r., & ackerman, p. l. (1989). motivation and cognitive abilities: an integrative/aptitude-treatment interaction approach to skill acquisition. journal of applied psychology, 74, 657–690. doi:10.1037/0021-9010.74.4.657. kanfer, r., & ackerman, p. l. (1996). a self-regulatory skills perspective to reducing cognitive interference. in i. g. sarason, b. r. sarason & g. r. pierce (eds.), cognitive interference: theories, methods, and findings (pp. 153–171). mahwah, nj: erlbaum. kapur, m. (2008). productive failure. cognition and instruction, 26, 379–424. doi:10.1080/07370000802212669. keith, n., & frese, m. (2005). self-regulation in error management training: emotion control and metacognition as mediators of performance effects. journal of applied psychology, 90, 677–691. doi:10.1037/0021-9010.90.4.677. knollmann, m. (2006). kontextspezifische emotionsregulationsstile. entwicklung eines fragebogens zur emotionsregulation im lernkontext mathematik [emotion regulation in learning contexts: development of a questionnaire measuring emotion regulation during math learning]. zeitschrift für pädagogische psychologie, 20, 113–123. doi:10.1024/1010-0652.20.12.113. kolodner, j. (1983). towards an understanding of the role of experience in the evolution from novice to expert. international journal of man-machine studies, 19, 497–518. doi: 10.1016/s00207373(83)80068-6. kolodner, j. (1997). educational implications of analogy: a view from case-based reasoning. american psychologist, 52, 57–66. doi:10.1037/0003-066x.52.1.57. kuhl, j. (1985). volitional mediators of cognitive-behavior-consistency; self-regulatory processes and action versus state orientation. in j. kuhl & j. beckmann (eds.), action control. from cognition to behavior (pp. 101–128). berlin: springer. kuhl, j. (2000). the volitional basis of personality systems interaction theory: applications in learning and treatment contexts. international journal of educational research, 33, 665–703. doi:10.1016/s08830355(00)00045-8. künsting, j., kempf, j., & wirth, j. (2013). enhancing scientific discovery learning through metacognitive support. contemporary educational psychology, 38, 349–360. doi:10.1016/j.cedpsych.2013.07.001. lazarus, r. s. (1991). emotion and adaptation. oxford: oxford university press. lazarus, r. s. (1993). why we should think of stress as a subset of emotion? in l. goldberger & s. breznitz (eds.), handbook of stress: theoretical and empirical aspects (2nd ed., pp. 21–39). new york: the free press. lazarus, r. s., & folkman, s. (1984). stress, appraisal, and coping. new york: springer. maehr, m. l., & zusho, a. (2009). achievement goal theory: the past, present, and future. in k. r. wentzel & a. wigfield (eds.), handbook of motivation at school (pp. 77–104). new york: routledge. marsh, h. w., walker, r., & debus, r. (1991). subject–specific components of academic self–concept and self–efficacy. contemporary educational psychology, 16, 311–345. doi:10.1016/0361-476x(91)90013b. tulis  et  al       | f l r     25   mathan, s. a., & koedinger, k. r. (2005). fostering the intelligent novice: learning from errors with metacognitive tutoring. educational psychologist, 40, 257–265. doi:10.1207/s15326985ep4004_7. minsky, m. (1997). negative expertise. in p. j. feltovich, k. m. ford, & r. r. hoffman (eds.), expertise in context (pp. 515–521). menlo park: aaai press/mit press. mory, e. h. (1996). feedback research. in: d. h. jonassen (ed.), handbook of research for educational communications and technology. a project of the association for educational communications and technology, 919–956. new york: macmillan. oser, f., & spychiger, m. (2005). lernen ist schmerzhaft. zur theorie des negativen wissens und zur praxis der fehlerkultur [learning is painful. on the theory of negative knowledge and the practice of error culture]. weinheim, germany: beltz. pekrun, r., goetz, t., daniels, l. m., stupnisky, r. h., & perry, r. p. (2010). boredom in achievement settings: control-value antecedents and performance outcomes of a neglected emotion. journal of educational psychology, 102, 531–549. doi:10.1037/a0019243. perry, n. e., & winne, p. h. (2006). learning from learning kits: gstudy traces of students’ self-regulated engagements with computerized content. educational psychology review, 18, 211–228. doi:10.1007/s10648-006-9014-3. pintrich, p. r. (2000). the role of goal orientation in self-regulated learning. in m. boekaerts, p. r. pintrich, & m. zeidner (eds.), handbook of self-regulation (pp. 452–502). san diego, ca: academic press. schunk, d. h., meece, j. r., & pintrich, p. r. (2013). motivation in education: theory, research, and applications (4th ed.). upper saddle river, nj: prentice hall. reason, j. t. (1990). human error. cambridge: cambridge university press. doi:10.1017/cbo9781139062367. reason, j. t. (1995). understanding adverse events: human factors. quality in health care, 4, 80–89. doi:10.1136/qshc.4.2.80. resnick, l. b. (1984). beyond error analysis: the role of understanding in elementary school arithmetic. in h. n. cheek (ed.), diagnostic and prescriptive mathematics: issues, ideas, and insights (pp. 214). kent, oh: research council for diagnostic and prescriptive mathematics. rybowiak, v., garst, h., frese, m., & batinic, b. (1999). error orientation questionnaire (eoq): reliability, validity, and different language equivalence. journal of organizational behavior, 20, 527–547. doi:10.1002/(sici)1099-1379(199907)20:4%3c527::aid-job886%3e3.0.co;2-g. schmitz, b. (2001). self-monitoring zur unterstützung des transfers einer schulung in selbstregulation für studierende [self-monitoring to support the transfer of a self-regulation instruction for students]. zeitschrift für pädagogische psychologie, 15, 181–197. doi:10.1024//1010-0652.15.34.181. schwinger, m., steinmayr, r., & spinath, b. (2009). how do motivational regulation strategies affect achievement: mediated by effort management and moderated by intelligence. learning and individual differences, 19, 621–627. doi: 10.1016/j.lindif.2009.08.006. siegler, r. s. (2002). microgenetic studies of self-explanation. in n. granott & j. parziale (eds.), microdevelopment. transition processes in development and learning (pp. 31–58). cambridge: cambridge university press. doi:10.1017/cbo9780511489709.002. steuer, g., rosentritt-brunn, g., & dresel, m. (2013). dealing with errors in mathematics classrooms: structure and relevance of perceived error climate. contemporary educational psychology, 38, 196– 210. doi:10.1016/j.cedpsych.2013.03.002. tulis, m. (2013). error management behavior in classrooms: teachers’ responses to students’ mistakes. teaching and teacher education, 33, 56–68. doi:10.1016/j.tate.2013.02.003. tulis, m., & ainley, m. (2011). interest, enjoyment and pride after failure experiences? predictors of students' state-emotions after success and failure during learning mathematics. educational psychology, 31, 779–807. doi:10.1080/01443410.2011.608524. tulis, m. & dresel, m. (2013, august). motivational and emotional self–regulation and adaptive learning activities after errors. paper presented at the 15th biennial conference of the european association for research on learning and instruction (earli), munich, germany. tulis  et  al       | f l r     26   tulis, m., & fulmer, s. m. (2013). students’ motivational and emotional experiences and their relationship to persistence during academic challenge in mathematics and reading. learning and individual differences, 27, 35–47. doi:10.1016/j.lindif.2013.06.003. tulis, m., grassinger, r., & dresel, m. (2011). adaptiver umgang mit fehlern als aspekt der lernmotivation und des selbstregulierten lernens von overachievern [adaptive handling of errors as an aspect of learning motivation and self-regulated learning of overachievers]. in m. dresel & l. lämmle (eds.), motivation, selbstregulation und leistungsexzellenz [motivation, self-regulation and achievement excellence] (pp. 29–51). münster, germany: lit. tulis, m., steuer, g., & dresel, m. (subm.). components of adaptive individual dealing with errors during academic learning. van dyck, c., van hooft, e. a. j., de gilder, t. c., & liesveld, l. c. (2010). proximal antecedents and correlates of adopted error approach: a self-regulatory perspective. journal of social psychology, 150, 428–451. doi:10.1080/00224540903366743. van lehn, k. (1988). toward a theory of impasse-driven learning. in h. mandl & a. lesgold (eds.), learning issues for intelligent tutoring systems, 19-41. new york: springer. van lehn, k., siler, s., murray, c., yamauchi, t., & baggett, w. (2003). why do only some events cause learning during human tutoring? cognition and instruction, 21, 209–249. doi:10.1207/s1532690xci2103_01. vygotsky, l. s. (1978). mind in society: the development of higher mental processes. cambridge, ma: harvard university press. weiner, b. (1986). an attributional theory of motivation and emotion. psychological review, 92, 548–573. doi:10.1007/978-1-4612-4948-1. westermann, k., & rummel, n. (2012). delaying instruction: evidence from a study in a university relearning setting. instructional science, 40, 673–689. doi:10.1007/s11251-012-9207-8. winne, p. h. & hadwin, a. f. (1998). studying as self-regulated learning. in d. j. hacker, j. dunlosky & a. c. graesser (hrsg.), metacognition in educational theory and practice (pp. 279–306). hillsdale, nj: erlbaum. wolters, c. a. (1998). self-regulated learning and college students' regulation of motivation. journal of educational psychology, 90, 224–235. doi:10.1037/0022-0663.90.2.224. wolters, c. a. (2003). regulation of motivation. evaluating an underemphasized as-pect of self-regulated learning. educational psychologist, 38, 189–205. doi:10.1207/s15326985ep3804_1. zimmerman, b. j. (2008). investigating self-regulation and motivation: historical background, methodological developments, and future prospects. american educational research journal, 45, 166– 183. doi:10.3102/0002831207312909. zhao, b. (2011). learning from errors: the role of context, emotion, and personality. journal of organizational behavior, 32, 435–463. doi:10.1002/job.696. zhao, b., & olivera, f. (2006). error reporting in organizations. academy of management review, 31, 1012– 1030. doi:10.5465/amr.2006.22528167. microsoft word bronkhorst & de kleijn_publication.docx             frontline  learning  research  vol.4  no.  3  (2016)  75  -­‐  91   issn  2295-­‐3159       challenges and learning outcomes of educational design research for phd students dr. larike h. bronkhorst1 & dr. renske a.m. de kleijn utrecht university, the netherlands article received 10 october / revised 2 march / accepted 11 april / available online 17 may   abstract educational design research (edr) is described as a complex research approach. the challenges resulting from this complexity are typically described as procedural, whereas edr might also be challenging for different reasons, specifically for early career researchers. yet, challenging experiences may be noteworthy in the process of learning to do research and becoming a researcher. to explore this issue further, we engaged in a collaborative self-study, and conducted a narrative cross-case analysis of two phd candidates’ experiences of engaging in edr, focusing on challenges and learning outcomes. we find indications that the challenges of edr might be related to edr’s relatively new and minority position in educational sciences and the role a (early career) researcher needs to assume in edr. retrospectively, the challenges appear closely related to learning outcomes, which are described in terms of a more profound understanding of research (quality) and of oneself as a researcher. as such, insights gained by self-study of research practices provide a complementary perspective to existing literature on edr and becoming a researcher. keywords: educational design research; phd learning; doctoral education; self-study                                                                                                                           1 corresponding author: heidelberglaan 1, 3584 cs utrecht, utrecht, the netherlands, email: l.h.bronkhorst@uu.nl doi: http://dx.doi.org/10.14786/flr.v4i3.198 bronkhorst  &  de  kleijn       | f l r     76   1. introduction educational design research is described as a challenging research approach (e.g. collins, joseph, & bielaczyc, 2004). different aspects of this complexity are discussed in literature, typically tracing the origin of this complexity back to the multiple aims of educational design research, namely contributing to the general understanding of teaching and learning and creating a viable contextualized design to solve a local problem (anderson & shattuck, 2012). some stress how educational design research might be especially challenging for early career researchers (herrington, mckenney, reeves, & oliver, 2007), defined as researchers with up to ten years of experience (andres, bengtsen, castaño, crossouard, keefer, & pyhältö, 2015), including phd students. the challenges of educational design research for early career researchers are described as procedural, whereas case studies (e.g. akkerman, bronkhorst, & zitter, 2013) and experience suggest that educational design research might also be challenging for different reasons. while some studies suggest that such challenges can lead to phd students experiencing dissonance (e.g. wisker, robinson, trafford, creighton, & warns, 2003), others suggest that while challenges can be burdensome for phd students, they can also be experienced as empowering (stubb, pyhältö, & lonka, 2011), benefitting the learning process involved in becoming a researcher (hall & burns, 2009). being early career researchers with experience in educational design research, we conducted a selfstudy exploring phd candidates’ experiences with educational design research in terms of the challenges as well as the learning outcomes. self-study is an unconventional and relatively unknown method, gaining popularity in research on teacher education (zwart, smit, & admiraal, 2015) as a powerful way of providing insights complementary to those gained from other research methods (bullough & pinnegar, 2001; loughran, 2007). as such, this article can be appreciated as a potentially thought-provoking example of using self-study methodology for studying researcher practices, illustrating the methodology’s potential and pitfalls to critically analyse early career researchers’ developing research practices. 1.1 educational design research (edr) the origin of educational design research (edr) is often traced back to the works of brown (1992). different authors use different terms, including design research, design-based research (or the abbreviation dbr) and design experiments; also, the specific methods used in edr studies differ (engeström, 2011a; reeves, herrington, & oliver, 2005). barab and squire (2004, p.2) typify edr as “a series of approaches, with the intent of producing new theories, artifacts, and practices that account for and potentially impact learning and teaching in naturalistic settings”. in their review of the last decade of edr research, anderson and shattuck (2012) characterize edr as research situated in a real educational context, concentrating on testing a significant intervention, in collaboration with practitioner(s), informed by theories and an assessment of the local context as well as practices in other contexts. accordingly, edr uses mixed methods, involves multiple iterations to perfect the design, and requires a collaborative partnership between researcher(s) and practitioner(s), as it focuses on theory development and overcoming a problem in local practice, typically culminating in design principles. edr is generally acclaimed as it is assumed to have the potential to enhance theoretical knowledge development while also having public educational value (van den akker, 1999) and therefore resonates with wider calls for increasing the relevance and impact of educational research (anderson & shattuck, 2012). baptista, frick, holley, remmik, tesch, and åkerlind (2015) describe how these calls for relevance are also being voiced in relation to research conducted as part of phd dissertations, where the usefulness of the knowledge gained by the research is increasingly considered. increasing attention to and application of edr is demonstrated by the special issues devoted to this topic by leading educational research journals, including journal of the learning sciences (barab & squire, 2004), educational researcher (kelly, 2003) and educational psychologist (sandoval & bell, 2004). bronkhorst  &  de  kleijn       | f l r     77   more and more, edr is not only applauded, but also critically assessed (svihla, 2014). several authors have pointed at potential weaknesses in edr methodology (dede, 2004; shavelson, phillips, towne, & feuer, 2003), questioning edr’s potential to draw causal claims in natural settings and edr’s tendency to generalize small scale studies (kelly, 2004). akkerman, bronkhorst, and zitter (2013) maintain that pursuing concurrent goals in edr requires immediate and sometimes intuitive actions and decisions, as accepted systems of quality assurance are lacking. barab and squire (2004) highlight how edr requires researchers to carefully balance insider and outsider perspectives: “if a researcher is intimately involved in the conceptualization, design, development, implementation, and re-searching of a pedagogical approach, then ensuring that researchers can make credible and trustworthy assertions is a challenge” (p.10). additionally, several scholars have addressed the challenges of analysis in edr, given the large amounts of data that are usually collected (collins, joseph, & bielaczyc, 2004; kelly, 2004). herrington, mckenney, reeves, and oliver (2007) have argued that specifically for beginning researchers, edr might be too challenging. for one, the longitudinal nature of most design studies might extend the four-year period in which most phd students are expected to complete their studies (evans, 2010). but even if the data collection itself can be completed, the richness of the data collected might extend the time and especially the skill needed for analysis. these challenges are presented and interpreted as procedural and thereby manageable, as can be deduced from anderson and shattuck’s (2012) solution in terms of multi-year multi-actor research agendas. such solutions help advance edr, but fail to appreciate the experience of engaging in edr and its challenges for early career researchers. for instance, a case study of edr conducted by a phd candidate stresses the consequences in terms of feelings of insecurity and misfit of a seemingly procedural challenge—namely assuring quality in edr studies (akkerman et al., 2013). recently, castelló, kobayaski, mcginn, pechar, vekkaila, and wisker (2015) have called attention to feelings of misfit as a phd student, as they impact present and future professional aspirations and can lead students to abandoning the field of learning and instruction. others have also cautioned against the consequences of feelings of dissonance as a phd student (wisker et al., 2003). 1.2 learning to do research in contrast, lee and roth (2003) argue that the learning potential of research actually lies in working through the challenges engaging in research is likely to generate. others have also hinted at the learning potential of challenging circumstances for learning to do research, especially for early career researchers (e.g. haigh, 2012; hopwood, 2010). in such cases, learning is conceptualized as participation (gonzálezocampo et al., 2015) and learning outcomes are studied in terms of becoming a researcher, emphasizing not only the technical, but also the personal nature of learning to do research (barnacle, 2005). “transitioning […] to the role of researcher is not as simple as acquiring a new set of skills and expanding one’s knowledge of scholarship” (hall & burns, 2009, p.53). departing from this conceptualization, the biographical or narrative process of learning is usually studied, as well as the open-ended and potentially transformative outcomes for researcher identity of engaging in research (e.g. lee & roth, 2003). such studies are in line with widespread calls for more in-depth, longitudinal research to explore what it means to become a researcher (hall & burns, 2009; stubb et al., 2011). in terms of specific learning outcomes, herrington and colleagues (2007) also report potential learning outcomes of conducting edr specifically for phd students. most importantly, phd students can learn to see practitioners as partners in research instead of beneficiaries of the outcomes of their research, thereby potentially increasing the impact of educational research on educational practice. moreover, phd students might benefit from learning about the ways in which edr differs from other research methods. also stressing the importance of awareness of the implications of methodological decisions made during research, newbury (2002, p.156, emphasis added) states that “[p]erhaps most important is research students’ exposure to alternative approaches”. similarly, pallas (2001) emphasizes how preparing doctoral students for an essentially unpredictable future entails acquainting them with epistemological diversity as consumers bronkhorst  &  de  kleijn       | f l r     78   (i.e. reading research grounded in diverse epistemologies) as well as producers (i.e. engaging in research with diverse epistemologies). summarizing, despite the fact that the challenges and learning potential of engaging in edr for early career researchers are acknowledged, there is still a great deal to explore about the actual experience of engaging in edr. in this study, we explore what challenges and learning outcomes phd students experience when engaging in edr, as such insight could not only support early career researchers and their supervisors in making their edr studies successful, but also provide an in-depth insider perspective, informing a wider audience about what it means to become an (educational design) researcher. 2. methods 2.1 context of the study cognizant of the differences in early career researcher education across countries (andres, bengtsen, castaño, crossouard, keefer, & pyhältö, 2015), we detail the specific characteristics of phd trajectories in educational and learning sciences in the netherlands, where this study was conducted. first of all, phd trajectories are jobs; phd candidates do not pay tuition, but are paid for doing research. as such, the phrase ‘phd student’ is not used in dutch, as candidates are not seen as students, although they are supervised by at least one full professor and one so-called ‘daily supervisor’ – usually an assistant or associate professor. additionally, the phd dissertation consists of four semi-independent articles, written in english and preferably published internationally. hence, although coursework is included in this trajectory, it usually does not extend throughout the trajectory, as the focus of the four-year trajectory is the research project. consequently, a completed (research) master’s degree is a prerequisite to enter a four-year phd trajectory. 2.2 self-study given the exploratory and interpretative research aim, and conceptualizing phd learning as participation, we chose to conduct a collaborative self-study, also referred to as auto-ethnography (e.g. holt, 2003). self-study is typified as a methodology for studying professional practices that stems from a desire “to be more fully informed about the nature of a knowledge of practice” (loughran, 2007, p.14). in the field of learning and instruction, self-study is an approach mainly used in research on teacher education and teaching (zwart, smit, & admiraal, 2015). self-study has become recognized as a powerful methodology to promote critical reflective attitudes, understand the relationship between theory and practice more profoundly, and develop knowledge from an insider perspective (loughran, 2007; petrarca & bullock, 2014; williams & ritter, 2010). the increasing use of self-study, reflected in the creation of the self-study journal studying teacher education, can be understood in the light of current debates on methodology—more specifically, debates on how to take into account the contextualization of human thinking and acting, and the importance of both outsider and insider understanding in research on teacher education (e.g., hamilton, smith, & worthington, 2008) and in educational research in general (e.g., maxwell, 2004a). 2.3 participants the authors of this paper are the two participants in this self-study. for clarity, we use pseudonyms and refer to them in the third person, and only use ‘we’ to refer to ourselves as authors of this article. at the time of the study, both participants were in the penultimate year of their phd trajectory. they had started their phd trajectories at the same university department in 2008, having completed the same research bronkhorst  &  de  kleijn       | f l r     79   master’s degree in two consecutive cohorts. the full professors who supervised their phd research differed, but the phd candidates had the same daily supervisor. mary was 25 years old at the time of data collection. in her edr study, she originally aimed to design a tool to support the goal-relatedness of the master’s thesis supervision by conducting group discussion meetings and individual interviews with five supervisors with a locally good reputation (see also de kleijn, meijer, brekelmans, & pilot, 2015). erica was 28 years old. she studied how student teachers’ meaning-oriented learning and deliberate practice could be fostered by collaboratively re-designing two year-long courses in the teacher education program with two pairs of teacher educators, based on design principles developed in prior research2 (see also bronkhorst, meijer, koster, & vermunt, 2011; bronkhorst, meijer, koster, akkerman, & vermunt, 2013). mary and erica informally discussed their progress in their edr studies and it seemed that they had quite different experiences, which they found striking given their similar background. this triggered the desire for a more systematic exploration of their experiences in edr by means of a collaborative self-study. 2.4 interviews we assumed that a probing interview might help in the explication of challenges and learning outcomes, as self-narratives have been shown to be a powerful methodology for self-studies (haigh, 2012). such interviews require well established interview skills and can benefit from the interviewer and interviewee being acquainted (lichtman, 2006). therefore, the interviews were conducted by their daily supervisor, christine, as she knew both mary and erica and had extensive experience in conducting qualitative research in general and open interviews specifically. christine was asked to conduct individual in-depth interviews with the phd candidates revolving around broadly defined topics: (1) the edr research that they had conducted; (2) the challenges they had experienced in their research and how they had dealt with these; and (3) the learning outcomes in the process of becoming a researcher that they attributed to engaging in edr. these themes were to be addressed longitudinally (i.e. in terms of their development over time). as a potential learning outcome of engaging in edr concerns a different relationship with participants (herrington et al., 2007), mary and erica had interviewed the participants in their edr studies upon completing their studies. selected fragments concerning the edr participants’ perception of the phd candidates from these interview were provided to christine, to inform her about the participants’ perspectives. christine designed an open interview structure (see table 1) adhering to this input and these guiding principles. she explicitly used her knowledge of the phd candidates to have them explicate more. both interviews lasted about an hour and a half and were fully transcribed.                                                                                                                           2 although departing from design principles, the approach to the collaborative design in erica’s study can be characterized as a formative intervention (see also bronkhorst et al., 2013; penuel, 2014). bronkhorst  &  de  kleijn       | f l r     80   table 1. interview themes and example questions interview themes example questions engaging in edr can you explain your reasoning for designing the study the way you did? would you (still) call it (design) research and why? challenges experienced i can recall that this was challenging, at times. can you tell me some more about that? how did you deal with this challenge? was this a conscious choice? becoming a researcher how would you describe yourself as a researcher? what did you learn by engaging in edr? additional probes used for all themes would you do/have done it differently in the future/past? can you give an example? 2.5 analysis a cross-case analysis of experiences of engaging in edr was performed. this was preceded by a within-case, narrative, connective analysis (maxwell, 2004b) of each phd candidate individually. first, drawing on critical incident technique (meijer, de graaf, & meirink, 2010), in both interviews fragments were selected where challenges and/or (learning) outcomes were substantially discussed. we verified our selections by scrutinizing the transcripts for words that indicated emotions (e.g. ‘doubt’), struggles (e.g. ‘difficult’) and/or words that indicated changes (e.g. ‘different’) or time differences (e.g. ‘now’). secondly, we traced each of these key experiences backwards and forwards: in the transcripts, we identified the processes by which they came about and how they subsequently developed, as well as factors or individuals that had influenced their origin, development or outcome. in order to triangulate the findings from these interviews with the phd candidates, the interviews with the participants of both phd candidates’ edr studies were also scrutinized for confirming and disconfirming evidence. thirdly, we compared and discussed our individual findings from the previous steps until a consensus on the relevant themes and their relationships was reached. the quality of this last step was enhanced by a ‘peer-debriefing’ (guba, 1981), in which a colleague, unfamiliar with the study, read the data and analysis and critically questioned the initial findings. based on these steps, we created two descriptions which are presented chronologically in the results section in order to increase legibility and understanding for readers. these descriptions are based on and contain illustrative quotes from the interview transcripts of the interviews with the phd candidates and of the data from their participants. we used these descriptions for the cross-case analysis. the cross-case analysis was sensitized by our theoretical framework, focusing on developing views on edr methods and quality, the role of a researcher in edr, and the process of becoming a researcher. bronkhorst  &  de  kleijn       | f l r     81   3. results 3.1 mary although an edr study had been included in mary’s research plan from the start, she kept postponing it. her hesitation mainly resulted from the lack of guidelines, protocols or general conventions mary found in the literature on edr. she herself considered reliability, in terms of reproducibility, a key quality criterion for scientific research, for which clear and shared conventions were necessary. she generally preferred to be in control of what she was doing: “this is what makes me insecure and creates chaos in my head. because there are no guidelines to hold on to and ‘everything goes’. and that i find very difficult. i obviously need boundaries and limitations.” additionally, at that time she thought research was about answering questions and proving or demonstrating theories. yet, when an edr study was no longer necessary for completing her thesis, she made the conscious decision to engage in edr. she knew that edr was well outside her comfort zone and thus would be fairly challenging for her, but she wanted to be(come) an all-round researcher and, being a phd student, she counted on her supervisors’ support. for her edr study, mary invited expert thesis supervisors to three collaborative design meetings. she prepared these meetings extensively, but she reasoned that she could not completely control (“board up”) the research, as she sought her participants’ expertise on the topic of thesis supervision, for which she needed their ownership, expertise and creativity. this conscious lack of control over how the process unfolded was a recurring challenge for her, before and between these meetings. she also experienced an ethical dilemma in relation to the time investment she asked of her participants: “how will i ever tell them that they invested three times two hours, almost an entire working day? how will i ever tell them that i am not able to write [an article] about it? that it eventually does not lead to something that will be part of my dissertation. that was the biggest stress [factor], so to say.” during the collaborative design meetings, she was immediately confronted with things that did not go as she had imagined, but she surprised herself by being able to adapt to unexpected circumstances successfully: “i thought: ‘i want to understand what is happening here.’ […] at that time i did not worry at all, as in: ‘gosh, where is this going?’ it was more in looking back and when i started preparing the next meeting, that i panicked. because that brought me back in the research mind set. but during the meetings i was mainly curious.” the fact that she saw how she could relate her participants’ contributions to the literature greatly supported this process, as did their general enthusiasm: “they were also really captivated by the issue and also appreciated discussing it so much that my fear of ‘they are wasting their time here’ lessened a little. and i also really observed during those meetings that [engaging in the study] brought about things for them as well.” the planning, time-management and enthusiasm was also recognized by one of her participants, who indicated in the final interview: “well, i really liked it, the way that you handled it. i mean, you are enthusiastic but also focused and flexible in the way you handled things. i found that very pleasant.” (edr participant) these meetings, the data collection, became a collaborative exploration and mary’s role as a researcher changed in this process; she was seen more as an expert on the literature of supervision, which contrasted with her earlier experiences with questionnaires and interviews, where she felt like a “nitwit”. bronkhorst  &  de  kleijn       | f l r     82   yet, it also meant that sometimes others took on the leadership role and determined the course of action in the meetings. moreover, mary did not only ask questions, but was also asked questions in return. she saw this as a sign of ownership on the part of her participants, necessary for the research, and welcomed it as such. looking back, engaging in edr brought about a number of changes for mary. first, her view on research and research quality has shifted. she would no longer claim reproducibility is a key criterion for scientific research. instead, transparency has become crucial and the dialogue with theory is now her first and foremost connotation with scientific research. consequently, she would design future studies differently, leaving room to deviate from her original plan. similarly, she would now say that research is about understanding and asking questions, rather than answering them. by engaging in edr, she now has a clearer view on it, but she still perceives confusion between different perspectives on edr: “i kind of have the impression that we have yet to agree on what [educational design research] actually is. or at least that there are a lot of different perspectives on it. so it has become clearer for me what i would consider good design research.” as her edr was actually a quest for understanding and did not result in a design, she doubts whether she has actually engaged in edr, an idea with which she now feels comfortable: “i don’t mind that at all. that was not my first priority, or goal. i consider such a design a means. what this resulted in—what i had not imagined in advance—is the [increased] understanding.” now, she would question the time investment asked of participants in completely controlled research: “can you ask people to…what does it entail to ask people to participate in such a highly structured interview, in which there is hardly any room for their own input? in such a way that you as a researcher are only ‘taking’?” moreover, she now sees the dialogue with theory, present already during data collection, as a key characteristic for research quality in general. these shifting perspectives also concern how she sees herself as a researcher. she no longer believes in the need to choose between the paradigms or ‘teams’ she perceived before, but feels comfortable as a multi-faceted and most of all curious researcher. she prefers certain types of research, including a preference for control and statistics, without considering these to be better types of research. “i immediately get the jibbers when i am assigned to a team, whether that is qualitative or quantitative. i like to think that the [research] question is leading, so to speak. the research design is then a means to answer the question, to put it like that.” she now knows she is capable of doing different types of research to satisfy her curiosity. 3.2 erica erica’s research goal was to study an intervention in educational practice, by having the educators (i.e. the participants in her study) experiment and explore the effects of that experimentation. the two main guiding principles for her research were that, first, educators should have agency in shaping educational design, as she believed that researcher control is not desirable nor possible in general, and, second, that the agency and expertise of the educators would actually make the design better. “so in that sense, i hoped that they would try out things for me, but i also hoped that the things they would then try out would be enriched with what they knew. so not just [based on] my theoretical knowledge.” bronkhorst  &  de  kleijn       | f l r     83   based on the literature she had read about edr, she thought her research would be a somewhat structured research endeavour. she started her project combining ideas resulting from her master’s with her intuitive ideas about how “the world works”. initially, these ideas were almost contradictory. “if you look at my research design, it really falls between two stools. the one is how i was educated, with large scale and more objective instruments on which i would have no further influence. and [on the other hand] apparently my intuitive ideas about how such a thing works and which data you need to collect for that.” in her research design, she mainly attributed the valuable expertise required in this research to the educators, as she felt she had very little knowledge of the educational context, also calling herself a “nitwit”. upon engaging in the research, erica noticed how the process took its own turn, which she— contrary to her expectations—could not really characterize as structured nor cyclical, because every decision was grounded in a prior decision. moreover, she noticed how she herself felt reluctant to exert agency. at some point, her participants asked her to share more of her knowledge and expertise, as they indicated in their interview: “at some point we said: […] ‘we want to hear your opinion. […] especially [as it is] a different perspective. that is the added value.’” (edr participant) this made her realize that her research would benefit from combining different expertise, including her own. she became more proactive and felt more at ease with steering the collaborative design according to her own agenda. she noticed how she recognized possibilities in which she could influence the educators’ engagement, relying on communicative techniques she normally did not associate with research (e.g. purposeful small talk) and which she had acquired elsewhere. “i had not considered [….] that i would apply those things. and that i would consider that they are part of research, [which] i had not really imagined. i thought it would be […] much less interaction, or actually, much less personal.” this led to ethical dilemmas as she questioned if it “was allowed” to act in this way deliberately in research. “but because i feel as though the values of education and research are different, and people always think that research is objective, and spotless and ethical and responsible, that makes it feel worse when you apply these [communicative techniques] as a researcher.” all in all, this meant that her research design relied more and more on her intuition and less on what she assumed other educational researchers would prefer. intuitively she thought her research “could not be done differently”, but rationally she feared the response of the educational and learning sciences community at large. especially at conferences, erica often realized how her research differed from research presented: “actually at each research conference i attended, where a different type of research was presented, each time i thought: ‘i do that completely differently…why do i do that so differently?’ also because i had trouble indicating what it was exactly that i was doing, and why i thought it was important.” this in turn led to doubts about herself as a researcher, which she hid from her supervisors, along with specific details about what she was really doing in her research. “during the year i had my doubts [about] if i was a researcher after all, or someone who participated. or someone who was very meaningful for the educators, but who doesn’t amount to anything in the research context.” her participants recognized these doubts, as they indicated during their interview: “in that sense, i did not envy her [erica], the past year. i was aware that she put herself in a difficult position, by choosing this [research] approach. […] because it means that she has a lot to legitimize in the research domain.” (edr participant) bronkhorst  &  de  kleijn       | f l r     84   in dealing with these challenges, she relied on two strategies. first, she made an effort to explicate her intuition(s), which was rather implicit at the start of her study. putting her rationale for the research in writing enabled her to have faith in the study’s value for research alongside its value for educational practice. secondly, she purposefully shared this vision, first at conferences and later also with her supervisors. the feedback was positive, which made her trust her intuition more. “in that sense the aera was also important. [...it was where] the keynote of engeström [2011b] took place and then i felt supported in: okay, it might not be common what i assert, but there are people who’d like to hear about it.” as such, she feels that perhaps her vision on research has not changed much, but she now understands how more collaborative research designs, and the active role assigned to participants in these, benefit research next to educational practice, which was a crucial outcome for her. in terms of research and research quality, she assigns increased importance to (ecological) validity: measuring what one intends to measure in context. she describes how she might not feel comfortable doing educational research in decontextualized settings any more. in general, she would now argue that research control over a situation, often employed to yield reliability, might produce unnatural behaviour. she prefers to rely more on transparency and the available knowledge on human behaviour in designing her studies: “in that sense i started thinking less about how a research intervention ought to be and more about, ‘what do we know about how people learn?’” looking back, erica now thinks that her perspective on edr or intervention research in general may not be shared by all, but is supported by some whom she holds high. consequently, she would now call herself a researcher, albeit one with a specific perspective, namely: ‘we have as much to learn from practice as they do from us’. this in turn has increased her confidence towards the research community and her participants. 3.3 cross-case comparison 3.3. 1 challenges of edr mary and erica both describe how engaging in edr was challenging, especially with respect to two archetypal characteristics of edr: the cyclical process, including multiple iterations, and the relationship with participants. both describe how their experience with edr differs from how edr is presented in the literature, specifically as being simplified and structured. mary had expected this in advance, which was one of the main reasons for postponing her edr study, whereas erica was somewhat surprised, despite the fact that the complexity of edr is described in the literature. apparently, reading about edr’s complexity— including the necessity of ongoing decisions and the challenges of balancing the multiple stakes and stakeholders involved—does not fully account nor prepare for the experience of edr. only when actually engaging in the research did it become clear to both phd students what a cyclical research process entailed, indicating that learning by doing seems necessary for learning to conduct edr. despite having taken rather different approaches, erica and mary both experienced ethical dilemmas with respect to their role in relation to the role of the participants. the nature of their dilemmas differed. erica struggled with the fact that she purposely invested in social interaction with her participants in order for them to be committed to her study, which she found ethically questionable in conducting research. the participants in her edr study do not mention having noted such behaviour when describing the working relationship, which they felt was appropriate. mary, on the other hand, struggled with the fact that her participants invested time in her study while she herself was not even sure about whether the study would result in a chapter in her dissertation. in contrast, the participants in her edr study only mentioned how inspiring their participation in the edr study had been for them. bronkhorst  &  de  kleijn       | f l r     85   3.3.2 becoming a researcher edr necessitated dealing with challenges and the resulting feeling of misfit. mary had anticipated that she personally would not be a good fit with edr; she questioned whether she would be able to cope with scant guidelines about how to carry out the study, mainly in light of her prior experience with more controlled research and personal preference for structured activities. erica experienced a lack of fit between her edr approach and what she considered to be the “general research community” as she questioned whether her collaborative approach to research, initially mainly informed by previous experiences outside academia, would be accepted and understood by the general educational research community, where other standards appeared to exist. conducting their edr studies involved working through these challenges and dealing with their insecurities, thereby fuelling reflections on becoming and being a researcher. ironically in light of the insecurities, the most salient outcome is that in the end both phd students describe a development from seeing themselves as ‘nitwits’ or novice researchers to being seen as an expert in their specific fields of study. in both cases, the participants in their edr studies played an important part in the transition, as they explicitly asked the phd candidates to take an expert role and to share their knowledge, even before the candidates themselves felt comfortable doing so. 3.3.3 quality of educational research next to reflections on their own expertise, the challenges experienced in engaging in edr triggered contemplations on the quality of educational research. both phd candidates described that after having engaged in edr, they regarded transparency as one of the most important criteria for educational research, more than replicability. this appeared to result from the fact that replicability was impossible to achieve in edr, which would imply that their edr was not scientific. yet, in hindsight, the phd candidates do not question their research, as the ongoing dialogue with theory and the general quest for understanding made it scientific, in their opinion (and the acceptance of their papers in international journals supports that assertion). both phd candidates do, however, question the feasibility of replicability when it is operationalized in terms of complete researcher control. 4. conclusion and discussion this study departed from contrasting findings of previous studies concerning the challenges of phd students when engaging in educational design research (edr). for one, the challenges of edr are described as procedural, but also as fundamental in the literature. moreover, the consequences of dealing with such challenges are evaluated as burdensome, but also as empowering, benefitting the process of becoming a researcher. to explore these contrasting findings, this explorative study examined the experiences of phd candidates engaging in edr, focusing on challenges and learning outcomes. our findings show that the challenges experienced concerned two typical aspects of edr, namely the cyclical nature of an edr process and the role of participants (e.g., collins et al., 2004). more specifically, our analysis shows how the cyclical nature of edr required ongoing decisions within limited time. yet, it was not the procedural aspect of making these decisions, but the candidates’ awareness of the significant implications of these decisions for the research (cf. newbury, 2002) that was experienced as challenging. similarly, our findings reaffirm the potential edr to support phd students in learning to see practitioners as partners in research, rather than beneficiaries of the outcomes of their study (herrington et al., 2007). yet, for the phd candidates not interacting with participants, but balancing the multiple stakes and stakeholders involved – including the research community was troubling and caused a lot of doubt and insecurity. it follows that while the iterative nature and the role of the participants are manageable for phd bronkhorst  &  de  kleijn       | f l r     86   candidates engaging in edr, the (perceived) conflicts with accepted quality standards of research can make them disruptive (as suggested by akkerman, et al., 2013). the phd candidates’ experiences in edr contained both ambiguity and novelty, a combination that often proves to be quite challenging (cf. wisker et al., 2003). however, the analysis illustrates how working through the challenges actually made the engagement in edr educative, as this necessitated learning about alternative perspectives on and approaches to research and research quality. this resulted in a more elaborated perspective on edr and research quality in general, and a more pronounced understanding of what kind of researcher the phd candidates hoped to become. finding that the learning outcomes are intertwined with the challenges extends previous findings on phd learning, wherein phd experiences were categorized as burdensome or empowering (stubb et al., 2011), by illustrating that burdensome experiences can become empowering over time. more specifically, our findings show that in the end both phd candidates describe having developed a more refined perspective on research quality, alongside a new understanding of what it means to be a researcher. the latter finding is in line with a conceptualization of phd learning as a process of participation (gonzález-ocampo et al., 2015) and learning outcomes in terms of becoming (hall & burns, 2009). this resonates with pallas (2001) who, among others, stresses the importance of learning about and from epistemological diversity in research preparation. hall and burns (2009, p.61) draw attention to how learning about multiple and even contrasting epistemological approaches can inform phd students about the researcher they want to become, as “becoming a professional researcher requires students to negotiate new identities and reconceptualize themselves both as people and professionals in addition to learning specific skills” (p.49). generally speaking, the results of this study indicate that describing the challenges of edr for early career researchers in terms of the necessary time and technical skills needed for the data collection and analysis, is insufficient. alternatively, we would ascribe at least some of edr’s complexity and learning potential for phd students to two other aspects: edr’s relatively new and minority position in educational sciences, and the role a researcher needs to assume to make edr a success. first, edr was not a mainstream research design in the netherlands at the time of study and was only marginally discussed in doctorate curricula (cf. wilhelm, craig, glover, allen, & huffman, 2000, for a similar discussion of qualitative research). the phd candidates therefore did not have experience in designing, conducting or even reading about edr. moreover, some edr characteristics differ from what the phd candidates learned about research, as edr aims to generate hypotheses rather than to prove them; is interactive rather than objective or distant; and relies on emergent instead of completely controlled designs (collins et al., 2004). secondly, the relationship with the participants in edr also differs in many respects from other types of research, notably from those in which these phd candidates were educated. the relationship with participants is a cornerstone of edr and required mary and erica to take on a specific role as a researcher. our findings indicate that two aspects of this role can be quite challenging. first, in edr the participants and researchers are assumed to collaborate. it follows that part of the control over the research process is handed over to the participants (edwards, sebba, & rickinson, 2007). sharing the control over the research process is precisely what differentiates interventions that seek to build on participant agency and those that solely recognize researcher expertise, according to engeström (2011a). as the implications of shared research control have only recently become topic of debate, the phd candidates had few examples or guidelines to build on. secondly, the participants acknowledged and addressed the phd researchers as knowledgeable or even experts on their research topic. this expert role not only disagreed with the phd candidates’ perception of themselves as novices and their role as junior in the research domain, it it also implied that they as researchers might have had influence over the research content – something they have learned to avoid at all cost. these differences in perspectives highlight the merit of checking participant perspectives on the research, also as a way of quality assurance (see also bronkhorst et al., 2013). additionally, our analysis suggests that not only the causes, but also the experience of edr’s complexity, especially for early career researchers, is not fully represented entirely in the literature. the edr literature describes how edr is and should be a cyclical process and a collaborative exploration for both bronkhorst  &  de  kleijn       | f l r     87   researcher and participants, but the specifics of such processes and collaborations are typically not discussed extensively. akkerman and colleagues (2013) even suggest that there might be a difference between how edr studies are reified as structured while “a lot of design researchers in praxis act differently” (p.422). for those learning to do research, being aware of potential differences between edr practice and its reification might prove valuable. 4.1 limitations in light of calls for more in-depth research (stubb et al., 2011), attending to experiences of engaging in edr from a contextualized insider perspective was our main argumentation for choosing to conduct a self-study. while our findings highlight self-study’s possibilities for increased understanding of research practices, there are some pitfalls that deserve attention. for one, seeing as the content of the interviews was partly determined by the phd candidates, ways in which edr was not experienced as challenging or did not carry learning potential have received limited attention. in terms of interpretation of findings, distinguishing personally relevant findings from findings relevant for the field requires an outsider perspective in self-study. therefore, we asked a colleague to be involved in verifying our analysis, supporting us in being less biased. seeing as the anonymous reviewers also played a vital role in interpreting our findings from a wider perspective, we underscore the importance of debating self-study methods and findings publicly—in line with calls from the self-study community of teacher education (see, for instance, loughran, 2007). 4.2 implications the findings illustrate how the experience of engaging in research in general, and edr in particular, deserves more attention and support, which holds significant implications for designing and supervising early career researchers. in both cases studied, the edr experience evolves from challenging to carrying learning potential, but not without effort. therefore, for those debating whether or not edr is (too) challenging for phd students, we would like to stress how our study shows that engaging in edr can challenge phd students to develop and extend their methodological competence, to adapt and develop methodologies and perceive methodology as a field of study in itself. few would not consider these to be valuable outcomes. evans (2010), for instance, holds that they align with an extended understanding of professionality of early career researchers and newbury (2002) claims that they lie at the core of methodological reflexivity, which he considers key to researcher preparation. however, for these challenges to be educative, our findings suggest there needs to be space, not only in terms of time and resources, but also conceptually and methodologically, and support for phd students to work through such challenges. unfortunately, such space is not always available in our age of efficiency and accountability (biesta, 2010). in terms of support, our findings also draw attention to how engaging in edr plays out differently in light of phd students’ participation in other communities, echoing the results of castelló and colleagues (2015). for supervisors, it is important to realize that conducting edr as an early career researcher brings about a range of insecurities and challenges that can differ depending on researcher personalities, previous experiences with research, and participation in other communities. given the important role supervisors play in facilitating the successful completion of a phd trajectory (gonzález-ocampo et al., 2015), we suggest that supervisors need to adapt their supervision to the specific challenges that phd students experience, referred to as “adaptivity” in the context of in the context of master’s thesis supervision (de kleijn, bronkhorst, meijer, pilot, & brekelmans, 2014). yet, phd students and supervisors should not avoid insecurities altogether, as they can carry learning potential. a final implication concerns the methodology used in our study. typically, the professional practices under scrutiny in self-study concern teaching, with the exception of studies focusing on self-study as a method. there are very few examples in which research practices are the object of study (for an example, see bronkhorst, van rijswijk, meijer, koster, & vermunt, 2013). in research on teacher education, self-study bronkhorst  &  de  kleijn       | f l r     88   has become recognized as a powerful methodology for teachers and teacher educators to promote local professional (knowledge) development, as well as develop relevant knowledge for the field from an insider perspective (petrarca & bullock, 2014; williams & ritter, 2010). our experiences in this study suggest that invoking self-study to understand research practices can enrich our understanding in more than one way. not only by means of our findings, foregrounding aspects of engaging in edr undisclosed in the literature, but also by critically analysing and openly discussing the challenges involved in conducting research and becoming an educational researcher. keypoints studies phd students’ experiences of engaging in educational design research. finds specific edr challenges and learning outcomes pertaining to early career researchers. relates these challenges to edr’s position in educational sciences. advances self-study as a way for phd students to learn about research practices. acknowledgements the authors would like to thank prof. dr. paulien meijer for her role in data collection and dr. harmen schaap for his valuable help with the data analysis. references akkerman, s. f., bronkhorst, l. h., & zitter, i. (2013). the complexity of educational design research. quality & quantity, 47(1), 421-439. doi: 10.1007/s11135-011-9527-9 anderson, t., & shattuck, j. (2012). design-based research: a decade of progress in education research? educational researcher, 41(1), 16-25. doi: 10.3102/0013189x11428813 andres, l., bengtsen, s. s., castaño, l. g., crossouard, b., keefer, j. m., & pyhältö, k. (2015). drivers and interpretations of doctoral education today: national comparisons. frontline learning research, 3(3), 1-18. doi: 10.14786/flr.v3i3.177 barab, s., & squire, k. (2004). design-based research: putting a stake in the ground. the journal of the learning sciences, 13(1), 1-14. doi: 10.1207/s15327809jls1301_1 barnacle, r. (2005). research education ontologies: exploring doctoral becoming. higher education research & development, 24(2), 179-188. doi: 10.1080/07294360500062995 baptista, a., frick, l., holley, k., remmik, m., tesch, j., & åkerlind, g. (2015). the doctorate as an original contribution to knowledge: considering relationships between originality, creativity, and innovation. frontline learning research, 3(3), 51-63. doi: 10.14786/flr.v3i3.147 biesta, g. j. (2010). why ‘what works’ still won’t work: from evidence-based education to value-based education. studies in philosophy and education, 29(5), 491-503. doi: 10.1007/s11217-010-9191-x bronkhorst, l. h., meijer, p. c., koster, b., & vermunt, j. d. (2011). fostering meaning oriented learning and deliberate practice in teacher education. teaching and teacher education, 27, 1120-1130. doi: 10.1016/j.tate.2011.05.008 bronkhorst, l. h., meijer, p. c., koster, b., akkerman, s. f., & vermunt, j. d. (2013). consequential research designs in research on teacher education. teaching and teacher education, 33, 90-99. doi: 10.1016/j.tate.2013.02.007 bronkhorst  &  de  kleijn       | f l r     89   bronkhorst, l. h., van rijswijk, m. m., meijer, p. c., koster, b., & vermunt, j. d. (2013). university teachers’ collateral transitions: continuity and discontinuity between research and teaching. infancía y aprendizaje, 36, 293-308. doi: 10.1174/021037013807532972 brown, a.l. (1992). design experiments: theoretical and methodological challenges in creating complex interventions in classroom settings. journal of the learning sciences, 2(2), 141-178. doi: 10.1207/s15327809jls0202_2 bullough, r. v., & pinnegar, s. (2001). guidelines for quality in autobiographical forms of self-study research. educational researcher, 30(3), 13-21. doi: 10.3102/0013189x030003013 castelló, m., kobayaski, s., mcginn, m., pechar, h., vekkaila, j., & wisker, g. (2015). researcher identity in transition: signals to identify and managesspheres of activity in a risk-career. frontline learning research, 3(3), 35-50. doi: 10.14786/flr.v3i3.149 collins, a., joseph, d., & bielaczyc, k. (2004). design research: theoretical and methodological issues. journal of the learning sciences, 13(1), 15-42. doi: 10.1207/s15327809jls1301_2 dede, c. (2004). if design-based research is the answer, what is the question? a commentary on collins, joseph, and bielaczyc; disessa and cobb; and fishman, marx, blumenthal, krajcik, and soloway in the jls special issue on design-based research. journal of the learning sciences, 13(1), 105-114. doi: 10.1207/s15327809jls1301_5 edwards, a., sebba, j., & rickinson, m. (2007). working with users: some implications for educational research. british educational research journal, 33(5), 647-661. doi: 10.1080/01411920701582199 engeström, y. (2011a). from design experiments to formative interventions. theory & psychology, 21, 598628. doi:10.1177/0959354311419252 engeström, y. (2011b). intervening to shape the future. keynote given at the annual aera conference, new orleans, lo. evans, l. (2010). developing the european researcher: ‘extended’ professionality within the bologna process. professional development in education, 36, 663-677. doi:10.1080/19415251003633573 gonzález-ocampo, g., kiley, m., lopes, a., malcolm, j., menezes, i., morais, r., & virtanen, v. (2015). the curriculum question in doctoral education. frontline learning research, 3(3), 19-34. doi: 10.14786/flr.v3i3.191 guba, e. g. (1981). criteria for assessing the trustworthiness of naturalistic inquiries. educational technology research and development, 29(2), 75-91. doi:10.1007/bf02766777 haigh, n. (2012). historical research and research in higher education: reflections and recommendations from a self-study. higher education research & development, 31(5), 689-702. doi: 10.1080/07294360.2012.689955 hall, l., & burns, l. (2009). identity development and mentoring in doctoral education. harvard educational review, 79(1), 49-70. doi: 10.17763/haer.79.1.wr25486891279345 hamilton, m. l., smith, l., & worthington, k. (2008). fitting the methodology with the research: an exploration of narrative, self-study and auto-ethnography. studying teacher education, 4(1), 17-28. doi: 10.1080/17425960801976321 herrington, j., mckenney, s., reeves, t., & oliver, r. (2007). design-based research and doctoral students: guidelines for preparing a dissertation proposal. in c. montgomerie, & j. seale (eds.), proceedings of edmedia 2007: wolrd conference on education multimedia, hypermedia & telecommunications (pp. 4089-4097). chesapeake, va: aace. holt, n. l. (2003). representation, legitimation, and autoethnography: an autoethnographic writing story. international journal of qualitative methods, 2(1), 18-28. doi: 10.1177/160940690300200102 hopwood, n. (2010). doctoral students as journal editors: non-formal learning through academic work. higher education research & development, 29(3), 319-331. doi: 10.1080/07294360903532032 kelly, a. e. (2003). theme issue: the role of design in educational research. educational researcher, 32(1), 3-4. doi: 10.3102/0013189x032001003 kelly, a. e. (2004). design research in education: yes, but is it methodological? journal of the learning sciences, 13(1), 115-128. doi: 10.1207/s15327809jls1301_6 bronkhorst  &  de  kleijn       | f l r     90   kleijn, r. a. m. de, bronkhorst, l. h., meijer, p. c., pilot, a., & brekelmans, m. (2014). understanding the up, back, and forward-component in master's thesis supervision with adaptivity. studies in higher education, 1-17. doi: 10.1080/03075079.2014.980399 kleijn, r. a. m. de, meijer, p. c., brekelmans, m., & pilot, a. (2015). adaptive research supervision: exploring expert thesis supervisors' practical knowledge. higher education research & development, 34(1), 117-130. doi: 10.1080/07294360.2014.934331 lee, s., & roth, w. (2003). becoming and belonging: learning qualitative research through legitimate peripheral participation. forum: qualitative social research, 4(2), available online at: http://www.qualitative-research.net/index.php/fqs/article/view/708 (accessed october 26, 2012). lichtman, m. (2006). qualitative research in education. a user's guide. thousand oaks, london, new delhi: sage publications, inc. loughran, j. (2007). researching teacher education practices responding to the challenges, demands, and expectations of self-study. journal of teacher education, 58(1), 12-20. doi: 10.1177/0022487106296217 maxwell, j. a. (2004a). causal explanation, qualitative research, and scientific inquiry in education. educational researcher, 33(2), 3-11. doi: 10.3102/0013189x033002003 maxwell, j. a. (2004b). using qualitative methods for causal explanation. field methods, 16(3), 243-264. doi:10.1177/1525822x04266831 meijer, p. c., de graaf, g., & meirink, j. (2011). key experiences in student teachers’ development. teachers and teaching: theory and practice, 17(1), 115-129. doi: 10.1080/13540602.2011.538502 newbury, d. (2002). doctoral education in design, the process of research degree study, and the ‘trained researcher’. art, design and communication in higher education, 1(3): 149-159. doi: 10.1386/adch.1.3.149 pallas, a. m. (2001). preparing education doctoral students for epistemological diversity. educational researcher, 6-11. doi: jstor.org/stable/3594455 penuel, w. r. (2014). emerging forms of formative intervention research in education. mind, culture, and activity, 21(2), 97-117. doi: 10.1080/10749039.2014.884137 petrarca, d., & bullock, s. m. (2014). tensions between theory and practice: interrogating our pedagogy through collaborative self-study. professional development in education, 5, 265-281. doi:10.1080/19415257.2013.801876 reeves, t. c., herrington, j., & oliver, r. (2005). design research: a socially responsible approach to instructional technology research in higher education. journal of computing in higher education, 16(2), 96-115. doi:10.1007/bf02961476 shavelson, r. j., phillips, d. c., towne, l., & feuer, m. j. (2003). on the science of education design studies. educational researcher, 32(1), 25-28. doi: 10.3102/0013189x032001025 sandoval, w. a., & bell, p. (2004). design-based research methods for studying learning in context: introduction. educational psychologist, 39(4), 199-201. doi:10.1207/s15326985ep3904_1 svihla, v. (2014). advances in design-based research. frontline learning research, 2(4), 35-45. doi: 10.14786/flr.v2i4.114 stubb, j., pyhältö, k., & lonka, k. (2011). balancing between inspiration and exhaustion: phd students' experienced socio-psychological well-being. studies in continuing education, 33(1), 33-50. doi: 10.1080/0158037x.2010.515572 van den akker, j. (1999). principles and methods of development research. in j. van den akker, n. nieveen, r. m. branch, k. l. gustafson & t. plomp (eds.), design methodology and developmental research in education and training (pp. 1-14). the netherlands: kluwer academic publishers. wilhelm, r. w., craig, m. t., glover, r. j., allen, d. d., & huffman, j. b. (2000). becoming qualitative researchers: a collaborative approach to faculty development. innovative higher education, 24(4), 265-278. doi: 10.1023/b:ihie.0000047414.56668.5b bronkhorst  &  de  kleijn       | f l r     91   williams, j., & ritter, j. k. (2010). constructing new professional identities through self-study: from teacher to teacher educator. professional development in education, 36, 77-92. doi:10.1080/19415250903454833 wisker, g., robinson, g., trafford, v., creighton, e., & warnes, m. (2003). recognising and overcoming dissonance in postgraduate student research. studies in higher education, 28(1), 91-105. doi: 10.1080/03075070309304 zwart, r. c., smit, b., & admiraal, w. f. (2015). a closer look at teacher research: a review study into the nature and value of research conducted by teachers. pedagogische studien, 92(2), 131-149. frontline learning research 5 special issue „learning through networks‟ (2014) 72-80 issn 2295-3159 corresponding author: maarten de laat, open university of the netherlands, welten institute, valkenburgerweg 177, 6419 at, heerlen, the netherlands, maarten.delaat@ou.nl doi: http://dx.doi.org/10.14786/flr.v2i2.122 72 | f l r unfolding perspectives on networked professional learning: exploring ties and time maarten de laat a , jan-willem strijbos b a open university of the netherlands, welten institute, the netherlands b ludwig-maximilians-university of munich, department of psychology, germany article received 26 june 2014 / revised and accepted 27 june 2014 / available online 15 july 2014 abstract networked learning and learning networks are commonplace concepts in most contemporary discourse on learning in the 21st century. this special issue provides a collection of studies that address the need for a growing body of empirical work to extent the limited understanding of the use and benefits of networks in relation to learning and professional development. in this article we attempt to offer a synthesis of the studies presented in this special issue and reflect on their findings. the studies in this issue present a rich combination of networked professional learning research addressing issues related to the composition and structure of learning networks, their content and activities, showing how multi-faceted research in the field of networked learning really is. based on the findings and methods used in the articles in this issue, we articulate some recommendations for further research. the recommendations are focused on the need for advanced multi-level analysis to understand the complexity of learning ties, the need for employing a multi-method research approach to triangulate and contextualize findings, the need to conduct process and time-based analysis and finally the need to further develop a theory and toolkit for applying social network analysis in the context of networked learning. keywords: networked professional learning, networked learning, professional development, informal-formal learning. m. de laat & j. w. strijbos 73 | f l r 1. introduction the domain of networked learning research has been around for some time. the term has been used predominantly in the uk, where the research by steeples and jones (2002) and goodyear, banks, hodgson and mcconnell (2004) played a central role in the early days. originally there was a strong focus on higher education, but nowadays a networked learning approach to understanding learning practices has extended to include learning in formal, non-formal and informal settings (hodgson, de laat, mcconnell, & ryberg, 2014). according to hodgson et al. (2014), networked learning refers to learning through connections between learners, learners and their tutors, and a learning community and its learning resources. within networked learning, learners have always been seen as proactive and engaging agents. many contemporary perspectives on networked learning derive from critical and humanistic traditions (dewey, 1916; freire, 1970; illich, 1971; mead, 1934) positing that learning is social, takes place in communities and networks, is a shared practice, involves negotiation, and requires dialogue (hodgson, mcconnell, & dirckinck-holmfeld, 2012). often, digital technology is used to support networked learning processes (goodyear et al., 2004). the field of networked learning aims to understand the pedagogical values and beliefs underpinning networked learning in order to advance teaching and learning practices and the design of technologies supporting such practices. the focus is on understanding how relations between learners influence teaching and learning in physical, online and/or blended settings. this special issue is an example of how a networked approach to learning has spread beyond education, since all the articles address questions around professional development, in this paper termed “networked professional learning”. this special issue forms an important and timely collection of articles especially because there is a strong interest in the promise and value of networked professional learning. there is considerable consensus that professionals organise and carry out their own professional development effectively through their own social networks and communities (cross & parker, 2004; duguid, 2005; hargreaves & fullan, 2012; weinberger, 2011; wenger, 1998). however, we lack empirical evidence about people‟s specific networked learning experiences. in particular, it is not well-understood how professionals build and maintain networked connections for learning, what the composition of these networks is, whether and how these learning relationships create value, and how to assess the outcomes of learning through networks in the context of professional development. research brought together in this special issue advances our understanding of networked professional learning, allowing us to reflect on and contribute to networked learning theory and helping us to develop and facilitate networked learning in practice. each article investigates learning from a relational point of view, in formal, informal or mixed settings. this final article offers a reflection on the unfolding perspectives and research presented by the articles collected in this special issue. 2. exploring networked professional learning the articles in this special issue are focused on understanding how social networks influence and impact professional development in networks and communities. vaessen, van den beemt and de laat (this issue) present a conceptual literature review to uncover some underlying mechanisms and factors that influence usage of networked learning in the context of teacher professional development. they explicitly explore the tension between formal and informal learning. they argue that the increased complexity of work requires professionals to use their networks to access and/or develop knowledge and expertise to stay up to date and function successfully. understanding the role and impact of these informal social networks on professional development can foster a better relationship – if necessary – with the traditional, yet dominant, formal professional development activities informed by acquisition and transfer of knowledge via expertdriven, pre-planned courses. vaessen et al.‟s literature review provides a broader framework for understanding professional development through participation in social networks, setting the context for the other articles in the special issue that examine particular aspects of “networked professional learning” in greater detail. m. de laat & j. w. strijbos 74 | f l r pataria, falconer, margaryan, littlejohn and fincher (this issue) investigate academics‟ learning through their personal professional networks. pataraia et al. build on roxå and mårtensson‟s (2009) research on teacher networks in academic contexts, focusing on conversations about teaching. more specifically, pataria et al. examined whether the composition of personal networks (i.e., the proximity of people with whom one is connected) and characteristics of interactions in these networks (i.e., what is exchanged and how it is valued) may support change of teaching practice in universities. this important descriptive research showed that the networks of academics were small, discipline-specific and strongly localised. based on data from interviews from two studies, they conclude that the academics interacted most frequently with closely proximate colleagues, typically from the same discipline. these findings support the notion that homophily (degree of similarity) influences establishment of ties and the development of networks. hytonen, palonen and hakkarainen (this issue) investigated network patterns and structures that contribute to professionals‟ cognitive centrality within a network. the context of their study was a professional training course in the field of energy efficiency. cognitive centrality was based on ties that represented who people contacted for professional advice over the course of twelve months; as such the networked ties constitute the product of the networked learning. more specifically, hytonen et al. examined the central actors within the network and their learning connections in order to identify possible factors that could explain cognitive centrality. their study showed that cognitive centrality is influenced by several factors, such as personal characteristics, expertise, and organisation that the actor represents, but a single decisive factor could not be found. these findings emphasise the complexity of social learning, suggesting that learning is highly contextualized and situated. rehm, gijselaers and segers (this issue) examined the transferability of knowledge in relation to the hierarchical network positions of members of an online community of learners during a professional development training program. rehm et al. addressed the notion that participants‟ hierarchical positions within the organization can have an effect on the collaborative processes within communities of learning. they showed that higher inand out-degree and centrality scores were associated with higher hierarchical positions within the organization. their longitudinal analysis indicated that these trends were established relatively early on during the professional development programme. the studies present a rich combination of networked professional learning research addressing issues related to the composition and structure of learning networks, their content and activities, showing how multi-faceted research in the field of networked learning really is. based on the findings reported and methods used, the following sections articulate some recommendations for further research. 3. unfolding networked professional learning the articles in this special issue provided us with snapshots of networked professional learning and details about the constitution of the learning networks in a variety of contexts. combined these articles challenge the naïve view that large(r) networks with many ties, or very elaborate networks with many ties of a specific type (e.g., weak vs. strong), are better and/or preferable by default. needless to say professionals take part in and maintain many networked relationships, but in essence networks are always about something, focused on a particular problem or shared interest. the “whole might be greater than the sum”, but the merit of the research presented in this special issue is to understand how particular (sub)networks or networked activity that professionals take part in contribute to their learning. for example, pataraia et al., and to some extent also hytonen et al., clearly show that professionals maintain many networked relationships with a variety of people for a number of reasons. done from an ego perspective, this work shows that professionals use their relations for exchanging information and ideas, to talk about work-related problems and to seek advice. rather than focussing on the impact and effects of networking in general it is very important to understand in great detail “what goes on in particular networks” and see how participation in networks affects learning. although the articles present findings at different levels of analysis and network scale (ego-personal network, sub-network, and/or whole network), there are interesting connections between the findings of m. de laat & j. w. strijbos 75 | f l r these different studies to be reflected upon. in the following synthesis we will try to address these differences in theory, method and network scale and explore if and how these different levels can be connected. boud and hager (2012) highlighted the importance of uncovering the ways in which people participate in social settings (networks) through which they seek to co-create knowledge and become a better professional. all articles in this issue take a social perspective on learning. in reflecting on professionals‟ network positions and the role of these networks in the learning processes, these studies draw on the “participation” metaphor of learning (as opposed to the acquisition metaphor, see sfard, 1998, for an elaborate discussion on these metaphors). while rehm et al. concentrate on the transfer of knowledge amongst members of a community of learners, they too position professional learning as a process of collaborative knowledge creation in social networks. social participation and network building is predominantly seen as an informal activity promoted, for example, through professional autonomy (cross & parker, 2004; kessels, 2012). however, vaessen et al. specifically argue that the underlying mechanisms for networked learning are found in both formal and informal settings. they further argue that networked learning is the most effective when located within work practices. in the workplace, learning is collaborative and situated within social relationships. networked learning is most effective in work settings in which professionals have high levels of autonomy, trust, openness and accountability and where these is an organisational culture of management promoting collaboration, discursive and open communication, and bottom-up learning and change. the findings presented by vaessen et al. are to some extent reflected in the studies by pataraia et al., hytonen et al., and rehm et al., but criticised as well. for example, the finding by pataraia et al. that academics‟ teaching networks were small, discipline-specific and strongly localised reveals that the establishment of connections with others is influenced by proximity, homophily, as well as perceived relevance and anticipated value of these connections. pataraia and colleagues‟ empirical data seems to suggest that academics‟ teaching networks are predominantly formed around strong ties. in a similar vein, findings by hytonen et al. show that cognitive centrality of core participants is affected by a multitude of factors, including personal characteristics (e.g., expertise, engagement), openness, and their organisational background. finally, rehm et al. show that characteristics of the formal work setting – i.e. people‟s hierarchical position – influence interaction patterns in an informal setting. the findings are in line with vaessen et al. in the sense that a similar structure (hierarchy) in the formal setting affects networked learning ties in the informal setting, but not necessarily as intended (although rehm et al. do not comment on this aspect). rehm et al. concluded that more senior professionals could draw more actively upon the input of colleagues allowing less senior participants to gradually move towards the centre of a network. likewise, vaessen et al. concluded that the network(s) transcend organisational boundaries, while rehm et al. indicate that this process may also benefit from some facilitation and/or intervention. both agree that management may need to promote networked learning by opening up organisational structures where management and community members can learn together. finally, vaessen et al. indicate that informal networks thrive in open practices, in which strong and weak ties co-exist (granovetter, 1973). such open network practices and “culture of learning” that is facilitated and promoted by the management appear especially relevant for professional learning (price, 2013). open practices consist of networks that are collections of individuals across organisational, spatial and disciplinary boundaries, who come together to create and share a body of knowledge (de laat, schreurs, & nijland, 2014). open networks focus typically on developing, distributing and applying knowledge (pugh & prusak, 2013). open network members connect around a common goal and share social and operational norms. they typically participate out of common interest and of a shared purpose rather than because of contract, quid pro quo or hierarchy. they are not bound or confined by shared identities and knowledge and meaning is not retained in the way in which it is done in communities of practice. the relationship between the members is much more loose and dynamic, yet effective in the creation of new ideas. open network practices offer professionals a more dynamic platform to connect with relevant peers who can help them to stay up to date than communities of practice do. a further feature of such open network practices is that they m. de laat & j. w. strijbos 76 | f l r are self-directed and non-hierarchical. wellman‟s (2002) notion of networked individualism emphasizes the point that professionals have a great ability to act on their own, to solve their problems and organise their lives, but they do this in a networked way with the help of friends and other relationships. the diversity of sources in a professionals‟ network is also echoed in the findings by pataraia et al. and hytonen et al. although rather implicitly, the articles in this special issue suggest several avenues for further research. in the next subsections we will discuss some directions for further research. 3.1 need to clarify the “what” and “why” of “learning tie” there is a clear need to develop theoretically-based and differentiated qualifications of the meaning of a tie, that is to investigate the “what” and “why” of a tie. in this special issue, pataraia et al., hytonen et al. and rehm et al. explicitly unfold the meaning of a tie. pataria et al. and hytonen et al. focused on both the “what” (content of a tie) and “why” (explanation for a tie or structure of personal/ego-network or the entire network), whereas rehm et al. focused only on the “why”. furthermore, networked learning ties can be treated both as relations that connect people as well as outcomes of relations (haythorntwaite & de laat, 2012). in the first instance, the tie refers to relational ties used for learning, such as a student learning from a teacher, students or professionals learning from peers, or novice professionals learning from experts. an example of networked learning ties as outcomes is when a group collectively acquires a competence in a certain domain that helps them to deal with new situations. as relational ties can represent both the process and the product of learning, there is a clear need to separate them or at least treat each tie as a compound construct consisting of several layers of process and product components. for example, at the individual level, a tie may consist of 30% on-going learning activities, 40% current project work, 20% personal bonds, and 10% status. a multi-layered perspective on ties allows for (multilevel) multiple regression approaches to understand the multifaceted nature of ties. the conceptualisation of ties as multi-layered constructs also opens up new directions regarding how these ties can be afforded, fostered, and facilitated through social interaction, design for learning, and technology. 3.2 need to examine networks at multiple levels and the interplay between levels combined the articles in this special issue cover all levels of scale possible. pataria et al. investigated the individual level in terms of academics‟ personal teaching networks and the characteristics of their interactions with colleagues. in a slightly different way, hytonen et al. examined the individual level to understand both the structure and heterogeneity of central participants‟ personal networks. they analysed the entire network to identify which other participants (alters) connected to the cognitively-central actors and to examine the associated network clusters and the degree of collaboration within these. finally, rehm et al. investigated networks as communities of learners, adopting a whole network analysis approach to explore the positions within these online communities based on participants‟ rank and hierarchical position within the organization. although these articles cover the range of possible levels – personal (ego), sub-network (community or larger cluster) and entire network – the explicit comparison or investigation of the interplay between various levels was not attempted. it is conceivable, for example, that an individual‟s personal network may be low in density, yet the individual may hold a key brokering position in the entire network. examining the interplay between levels might be a promising direction in future research in professional networked learning. referring back to the issue of the “whole being greater than the sum”, research on the interplay of levels will help to uncover how to potentially assess the nature of learning ties for the individual, a particular network and the organization. within human resource development (hrd) – especially from a formal management point of view – there is interest in monitoring and assessing networked learning in order to validate and award it. the immediate response seems to be on trying to assess networks, similar to registration of attendance of professional development programmes, rather than focussing on the value that is created through networks and communities (wenger, trayner, & de laat, 2011). multi-level research on the value of learning ties can help assess the outcome of networked professional learning in relation to different stakeholders. m. de laat & j. w. strijbos 77 | f l r 3.3 need for extending the methodological toolkit as there are different ways to conceptualise learning ties, there are different analysis techniques to study them. for example, who learns from whom, what do learners learn from each other, the kinds of interactions between learners, the direction of ties, flow of resources, and the frequency of interactions. several of these aspects are related to communication and information patterns, whereas others directly deal with learning itself. networked learning studies often address these aspects, however some reflection on how we may be more critical and cautious about the way in which network analysis is used to understand learning ties is required. a popular method for studying networked (professional) learning is the use of social network analysis (sna). two studies in this issue applied sna (hytonen et al.; rehm et al.) to understand the network structures or dynamics. sna has become rather popular for trying to understand learning ties, but we have to remain cautious about its application. sna was developed to understand for example the flow of information or communication across networks – i.e., more factual data. if person a passes something on to person b, traditional sna assumes that person b has received it. however, when learning is concerned this assumption may not hold. first, the extent to which whatever was passed on was received may be uncertain. second, the contribution of information to the actual learning process of the receiver is uncertain. hence, the network theory or operationalization of indicators behind the tests that researchers conduct may have different implications. does density in a communication network imply the same as density in a learning network? what does the shortest path mean in terms of “learning”? are all sna indicators by default useful indicators of learning? a more advanced theory of sna is needed to guide studies on “social network learning analysis” (snla). sna is a very flexible method, but it requires a solid theoretical framework to enable interpretation of findings. in the absence of a solid theoretical framework of learning through networks, researchers rely on conceptualisation from related research domains. when applying analysis techniques that reflect a different theoretical orientation, researchers risk type i and ii errors. furthermore, despite the ease with which network visualisations can be produced from sna, such visualisations should be approached with more restraint when interpreting the network structures. another approach would be the application of multilevel analyses, discussed by rehm et al. (this issue). an example is the recent study by eberle, stegmann and fischer (2014), who investigated legitimate peripheral participation (a well-known construct introduced by lave and wenger (1991) to describe learning processes in communities of practice), in terms of support structures used to foster newcomers‟ participation. they applied a 2-level model, which included 14 student councils (communities) and 68 newcomers. they found that exposure time (duration of community membership) and the support structure of “accessibility of community knowledge” positively predicted the newcomers‟ participation, whereas community size and “recruitment strategies” negatively predicted participation. finally, the instruments and methods applied in the articles in this special issue reflect that a multimethod approach is required to investigate networked professional learning and obtain a more complete understanding of the nature of “learning” reflected by the ties and the indicators that sna offers. a potential direction would be the combination of sna, content analysis of communication, and a contextual analysis through interviews (de laat & lally, 2003). the contributions by pataraia et al. and hytonen et al. are examples of such a contextualised approach to understanding network structures. rehm et al. acknowledge that their study would have benefitted from content analysis to help uncover how the “what” of the tie might have impacted network position and exchange of knowledge. 3.4 need to examine networked learning over time over the past decade, the issue of time has slowly developed into a more focal point of research on interactive learning processes. an early contribution in this respect is the work by de laat and lally (2003), who identified changes in both interactive and tutoring patterns within a community of learners, by distinguishing between the early, middle and end phase of the community experience. similarly, the notion of time is receiving more attention in the domain of (small) group collaborative learning, where learning is studied longitudinally, in terms of sequences of actions (suthers, dwyer, medina, & vatrapu, 2010), specific m. de laat & j. w. strijbos 78 | f l r timeframes such as days, weeks or months (arrow, henry, poole, wheelan, & moreland, 2005; reimann, 2009), or in terms of activities in a learning environment over time (schümmer, strijbos, & berkel, 2005). vaessen et al., pataraia et al., and hytonen et al. implicitly refer to issues of time. vaessen et al. describe professional development as an “ongoing process”, arguing that “networking skills need to be developed over time”. pataraia and colleagues‟ data were collected over a 12-months time span. they argue that “the temporal component of interactions determines the strength of ties”, and that the networks are not only influenced by proximity and/or discipline, but also have a “historical or temporal component”. hytonen et al. analysed data collected after a 12-months period following a training programme. their study focused on small group and community level aspects, but the development of the networks of cognitively central participants was not part of their aim. in contrast, rehm et al. explicitly adopted a longitudinal analytical lens when analyzing reply structures in online communities collected over a 14-week time span in terms of two blocks of about six weeks. their analysis showed that the more central positioning of senior management was established relatively early on and persisted – in fact slightly increased – over time. 4. closing remarks the articles comprising this special issue have advanced our understanding of networked professional learning. the empirical studies provided detailed accounts of the structure and focus of networked learning at various levels (ego, sub-network and whole network). they improve our understanding of the characteristics of networked learning and contribute to a much-needed empirical knowledge base in this area of research. the literature review provided by vaessen et al. helps to broaden our horizon as well as situating the findings of the other articles, by offering the mechanisms that influence networked professional learning. the three empirical studies, although addressing different levels of scale, reinforce and supplement each other. for example, where pataraia et al. find that professionals maintain multiple networks for their development, hytonen et al. identify several sub-networks centralized around key actors. the emergence of these personal networks hinges on expertise, interest, enthusiasm, competency, familiarity, organizational background, as well as hierarchy and formal organizational relationships and structures. simultaneously the studies (implicitly) generated some directions for future research that we elaborated upon: (a) the need to clarify the “what” and “why” of “learning tie”, (b) the need to examine networks at multiple levels and the interplay between levels, (c) the need for extending the methodological toolkit, and (d) the need to examine networked learning over time. the importance of professional autonomy and cross-boundary collaboration that seems to foster networked professional learning brings the emergence of open practices into perspective. both professionals and organizations are increasingly becoming aware that knowledge and innovation processes are not bounded by the organizational context and that boundary crossing becomes an important aspect of professional development. this raises further questions about how to monitor, promote and assess networked professional learning. the naïve view of “the more contacts the merrier” is too simplistic. studies in this special issue have shown that important features and mechanisms of networks are personal (probably small and localized), centralized around shared topics interests and key members, driven by professional autonomy and negotiate both informal and formal settings influenced by hierarchical organizational structures. these findings shed some light on how networks operate and create value, based on which knowledge about how to facilitate and manage networked professional learning can be inferred. m. de laat & j. w. strijbos 79 | f l r keypoints the studies in this issue present a rich combination of networked professional learning research addressing issues related to the composition and structure of learning networks, their content and activities, showing how multi-faceted research in the field of networked learning really is. need for advanced multi-level analysis to understand the complexity of learning ties need for employing a multi-method research approach to triangulate and contextualize findings need to conduct process and time-based analysis need to further develop a theory and toolkit for applying social network analysis in the context of networked learning references arrow, h., henry, k. b., poole, m. s., wheelan, s., & moreland, r. (2005). traces, trajectories, and timing. in m. s. poole & a. b. hollingshead (eds.), theories of small groups: interdisciplinary perspectives (pp. 313-367). thousand oaks, ca: sage. boud, d., & hager, p. (2012). re-thinking continuing professional development through changing metaphors and location in professional practice. studies in continuing education, 34(1), 17-30. doi: 10.1080/0158037x.2011.608656 cross, r. l., & parker, a. (2004). the hidden power of social networks: understanding how work really gets done in organizations. boston, ma: harvard business school press. de laat, m., & lally, v. (2003). complexity, theory, and praxis: researching collaborative learning and tutoring processes in a networked learning community. instructional science, 31(1-2), 7-39. doi: 10.1023/a:1022596100142 de laat, m. f., schreurs, b., & nijland, f. (2014). communities of practice and value creation in networks. in. r. f. poell, t. rocco & g. roth (eds.), the routledge companion to human resource development (pp. 249-257). new york: routledge. dewey, j. (1916). democracy and education: an introduction to the philosophy of education. new york: the macmillan company. duguid, p. (2005). the art of knowing: social and tacit dimensions of knowledge and the limits of the community of practice. the information society, 21(2), 109-118. doi: 10.1080/01972240590925311 eberle, j., stegmann, k., & fischer, f. (2014). legitimate peripheral participation in communities of practice: participation support structures for newcomers in faculty student councils. the journal of the learning sciences, 23(2), 216-244. doi: 10.1080/10508406.2014.883978 freire, p. (1970). pedagogy of the oppressed. new york: continuum. goodyear, p., banks, s., hodgson, v., & mcconnell, d. (eds.) (2004) advances in research on networked learning. dordrecht, the netherlands: kluwer academic publishers. granovetter, m. s. (1973). the strength of weak ties. the american journal of sociology, 78(6), 1360-1380. http://www.jstor.org/stable/2776392 hargreaves, a., & fullan, m. (2012). professional capital: transforming teaching in every school. new york: teachers college press. haythornthwaite, c., & de laat, m. f. (2012). social network informed design for learning with educational technology. in a. olofson & o. lindberg (eds.). informed design of educational technologies in higher education: enhanced learning and teaching (pp. 352-374). hershey, pa: igi-global. hodgson, v., de laat, m. f., mcconnell, d., & ryberg, t. (2014). researching design, experience and practice of networked learning: an overview. in v. hodgson, m. f. de laat, d. mcconnell & t. ryberg (eds.). the design, experience and practice of networked learning (pp. 1-26). dordrecht, the netherlands: springer. m. de laat & j. w. strijbos 80 | f l r hodgson, v., mcconnell, d., & dirckinck-holmfeld, l. (2012). the theory, practice and pedagogy of networked learning. in l. dirckinck-holmfeld, v. hodgson & d. mcconnell (eds.), exploring the theory, pedagogy and practice of networked learning (pp. 291-305). new york: springer. illich, i. (1971). deschooling society. manchester, uk: pelican books. kessels, j. w. m. (2012). leiderschapspraktijken in een professionele ruimte [leadership practice in a professional space]. heerlen, the netherlands: ruud de moor centrum, open university of the netherlands. lave, j., & wenger, e. (1991). situated learning: legitimate peripheral participation. cambridge, ma: university press. mead, g. (1934). mind, self & society from the standpoint of a social behaviorist. chicago, il: university of chicago press. price, d. (2013). open: how we’ll work, live and learn in the future [ebook, kindle edition]. crux publishing. pugh, k., & prusak, l. (2013). designing effective knowledge networks. retrieved october 24, 2013, from http://sloanreview.mit.edu/article/designing-effective-knowledge-networks reimann, p. (2009). time is precious: variableand event-centered approaches to process analysis in cscl research. international journal of computer-supported collaborative learning, 4(3), 239-257. doi: 10.1007/s11412-009-9070-z roxå, t., & mårtensson, k. (2009). significant conversations and significant networks: exploring the backstage of the teaching arena. studies in higher education, 34(5), 547-559. doi: 10.1080/03075070802597200 schümmer, t., strijbos, j. w., & berkel, t. (2005). measuring group interaction during cscl. in t. koschmann, d. suthers & t. w. chan (eds.), computer supported collaborative learning 2005: the next 10 years! (pp. 567-576). mahwah, nj: lawrence erlbaum associates. sfard, a. (1998). on two metaphors for learning and the dangers of choosing just one. educational researcher, 27(2), 4-13. doi: 10.3102/0013189x027002004 steeples, c., & jones, c. (2002). networked learning: perspectives and issues. london: springer. suthers, d., dwyer, n., medina, r., & vatrapu, r. (2010). a framework for conceptualizing, representing, and analyzing distributed interaction. international journal of computer-supported collaborative learning, 5(1), 5-42. doi: 10.1007/s11412-009-9081-9 weinberger, d. (2011). too big to know: rethinking knowledge now that the facts aren’t the facts, experts are everywhere, and the smartest person in the room in the room. new york: basic books. wenger, e. (1998). communities of practice: learning, meaning, and identity. cambridge, ma: cambridge university press. wenger, e., trayner, b., & de laat, m. (2011). telling stories about the value of communities and networks: a toolkit. heerlen, the netherlands: ruud de moor centrum, open university of the netherlands. wellman, b. (2002). little boxes, glocalization, and networked individualism. in m. tanabe, p. van den besselaar & t. ishida (eds.), digital cities ii: computational and sociological approaches (pp. 1025). berlin: springer. microsoft word castello et al_publication.docx           frontline learning research vol.3 no. 3 special issue (2015) 1-4 issn 2295-3159   corresponding author: montserrat castelló, facultat de psicologia, ciències de l’educació i l’esport. blanquerna. universitat ramon llull, císter 34. 08022. barcelona. phone: +34932533000, fax: +34932533031, email: montserratcb@blanquerna.url.edu doi: http://dx.doi.org/10.14786/flr.v3i3.197   trends influencing researcher education and careers: what do we know, need to know and do in looking forward montserrat castellóa, lynn mcalpineb and kirsi pyhältöc auniversity of ramon llull, spain buniversity of oxford, uk cuniversity of oulu and university of helsinki, finland article received 3 august 2015 / revised 19 august 2015 / accepted 20 august 2015 / available online 23 october 2015 abstract earli sig 24, researcher education and careers (sig-reac), was founded because increasing interest has emerged within the earli community into understanding different aspects of doctoral and post-phd researcher educational and career development. this special issue brings together the outcome of our first scholarly discussion at the sig-reac inaugural meeting in september 2014 in barcelona. the goal of each of the five co-authored papers is to make visible what has been overlooked, and to attend to methodological considerations in order to draw out future lines of research. as a collection, the papers address multiple levels and issues of researcher education: establishing the multifaceted phenomenon that is researcher education and careers and providing key concepts that others might take up, e.g., informal/invisible curriculum; the personal as a sphere of activity that may collide with the sphere of work; drivers of education that can provide cross-national points of comparison. further, by identifying gaps in the literature, these papers together lay out an ambitious research agenda in a number of areas related to researcher education. in the process, they provide an extensive list of references well worth exploring since they represent the knowledge networks of over thirty researchers. in this editorial paper the sig-reac is presented, and the characteristics of the papers, their limitations and some future challenges of researcher education are discussed. keywords: researcher education, career development; post phd education; phd education; cross-cultural research castelló  et  al       | f l r     2   earli sig 24, researcher education and careers (from now on sig-reac), was founded because increasing interest has emerged within the earli community into understanding different aspects of doctoral and post-phd researcher educational and career development. this special issue brings together the outcome of our first scholarly discussion at the sig-reac inaugural meeting in september 2014 in barcelona. our goal was to construct a richer, more comprehensive view of researcher education and careers: to begin to address the theoretical and methodological challenges underlying research and theory development in this area in order to create a shared agenda for the future. the meeting (and the preparation for it) launched collaborative writing that challenged us collectively to make transparent different theoretical perspectives, methods and methodologies. our goal was to negotiate these differences in order to articulate a commonly understood research agenda. while we shared an interest in examining the experiences of early career researchers we come from a variety of locations: geographic, disciplinary, career stage and intellectual tradition. when researchers from different theoretical, methodological, and national spaces want to do ‘real work’, it takes time to really understand each other and negotiate new understandings. therefore, the preparation for the sig-reac meeting included participants writing individual positions papers in which they addressed the following questions: what are the emerging trends in the research environment essential to better/more fully understand early career researcher (ecr) experience? what do we learn about ecr experience of the emerging trends by looking across the fields of academic communication, sociology of work, pedagogy? what are the gaps? what has been overlooked? what different methods and methodologies have been used across the three fields? which of these has been productive? what has been overlooked? in this way, we had the opportunity before the meeting to read each other’s thoughts and begin to get a sense of the richness and diversity in the group as well as common concerns, conceptions or methodologies. the preparation for the sig-reac meeting also included launching pre-discussions via moodle, based on reading each other’s papers. altogether 31 scholars from fourteen different countries participated in the sig meeting, where we launched co-writing, and worked together in small groups intensively for two days. post-meeting, this face-to-face work shifted to virtual exchanges and the special issue represents the results of our continued discussion over ten months. the special issue consists of five co-authored papers. the goal of each is to make visible what has been overlooked, and to attend to methodological considerations in order to draw out future lines of research. each of the papers addresses a specific aspect of researcher education and careers in order to develop a future research agenda: • drivers and interpretations of doctoral education today contributes to the literature on researcher education by examining the ways in which core global trends and drivers of higher education emerge in different guises at national levels. the paper compares recent doctoral education changes in the following countries – canada, colombia, denmark, finland, uk, and the usa – to provide insights on how global trends translate into local policies. by using the same global drivers as criteria across national boundaries, it is possible to see how educational policies are formed in considerably different ways. this raises questions about the universality of the phd. in the discussion, a research agenda for comparative studies is discussed. • the curriculum question in doctoral education begins by stating that although a global trend in researcher education has been developing more systematic doctoral education to enhance the quality of research and researchers, the value of a curricular perspective has remained largely unexplored both theoretically and empirically. it is argued that adopting an explicit curriculum approach is significant not only because it might help to disclose the tensions, but also because it allows us to face and reinterpret current challenges to doctoral education. first the concept of the curriculum in doctoral education is discussed and tensions between the formal/informal, open/hidden, and standardised/pluralised dimensions of curriculum are discussed. then, processes –how the curriculum is experiencedand outcomes – assessment and employabilityof doctoral education are addressed. finally, a research agenda drawing on notions of curriculum to help reconfigure doctoral education is proposed. castelló  et  al       | f l r     3   • the doctorate as an original contribution to knowledge: considering relationships between originality, creativity, and innovation explores the meaning of originality in doctoral studies and its relationship with creativity and innovation. the paper opens up discussion about the taken-for-granted traditional expectation of ‘originality’ as an outcome of doctoral research. it does so by juxtaposing ‘originality’ with the notions of ‘innovation’, and ‘creativity.’ by exploring the similarities and differences among the concepts, the paper provides insight into both the possible meanings of ‘originality’ in research as well as the utility of the term in the context of 21st century knowledge societies. some future research steps are suggested to move towards unpacking the relationship between doctoral training conditions and outcomes, in the sense of fulfilling the requirement of originality. • mentoring: a review of early career researcher studies describes the result of a focused literature review of studies on early career researcher as a base for further inquiry into mentoring, given the frequent reference to mentoring as a source of support for early career researchers, e.g., eu concordat on researchers. the most striking finding of this analysis was the unand underconceptualized nature of empirical studies. there is much research to do, first, to better inform our conceptualization of early career researcher mentoring and, second, to better understand the value of specific aspects of mentoring support. • researcher identities in transition: signals to identify and manage spheres of activity in a risk-career argues that changes in ‘knowledge societies’ mean researchers are now embarked upon what could be defined as a ‘risk-career.’ this paper uses a framework of researcher identity produced by analysing spheres of activity and individuals’ ability to identify and interpret external signals (expectations, constraints and opportunities) to account for theoretical assumptions about researcher identity. it is argued that applying the framework to empirical examples of tensions in identity construction provides the basis for future research to unravel the complex interplay between signals and spheres of activity when dealing with the tensions and struggles of becoming a researcher. as a collection, the papers address multiple levels and issues of researcher education: establishing the multifaceted phenomenon that is researcher education and careers and providing key concepts that others might take up, e.g., informal/invisible curriculum; the personal as a sphere of activity that may collide with the sphere of work; drivers of education that can provide cross-national points of comparison. further, by identifying gaps in the literature, these papers together lay out an ambitious research agenda in a number of areas related to researcher education. in the process, they provide an extensive list of references well worth exploring since they represent the knowledge networks of over thirty researchers. still, there are limitations represented in this special issue. while it has explored in depth a number of issues, we are mindful there remains much to explore. for instance, they mostly focus on doctoral experience, as does much of the research in this area. so we encourage ourselves and other researchers to pay greater attention to postdoctoral experience, both in and out of academia. for instance, the vertical transition from doctoral student to post-doctoral researcher still remains largely uncharted, as do horizontal transitions e.g. from academia to other types of careers. we know that internationally, more than half of phd graduates leave academia whether by choice or lack of opportunity (barnacle & dall’alba, 2011). what appears to be emerging internationally is a range of alternate academic positions: contract teaching, contract post-phd research, and increasingly teaching-only lecturer positions, as well as administrative positions related to research and teaching. in the non-academic context, emerging types of employment include business, government, ngos, banking, industry, and previously unknown positions, e.g., start-ups. unfortunately we know little of the experience of individuals in any of the three fields, e.g., the extent to which they have the skills needed, their satisfaction with their employment, what range of genre they use. this is especially the case as regards a theoretical perspective since most of the available evidence is non-theorized survey data. such studies are needed to gain better understanding of the complexity of researcher careers. as well, postdoctoral supervision is also an underexplored issue that deserves more research interest. post-phd researchers consistently report they do not receive supervisory support to develop as researchers, further that they are even discouraged from seeking out professional development opportunities themselves. castelló  et  al       | f l r     4   as long as such individuals are not conceived as becoming researchers, the supervisory attitudes they report are unlikely to change concluding remarks this special issue maps some of the uncharted terrain of inquiry into researcher education and careers. national developments in researcher education are affected by the global forces which, however, take different forms in national and local contexts. this became particularly apparent to us at our barcelona meeting where we represented fourteen different national contexts. there is still an insufficient understanding of how and in which forms global trends (which we collectively believe we understand) are translated into the local practices of researcher education, and their effect on doctoral education and academic work (which we collectively may not understand, though believe we do). accordingly, our overall conclusion is the need for well-designed international comparative studies so that as researchers we can gain a concrete understanding of the effects of global developments for researcher education and careers. we hope that the papers in this special issue evoke curiosity, provoke discussions and stimulate both theoretical and especially empirical research on researcher education and careers. the various approaches, empirical evidence and challenges identified in the papers highlight the importance of and the need for further research into this fascinating area. we look forward to lively discussion, commentaries and research papers addressing the new terrains in this area of research and encourage you to join us in earli sig 24, researcher education and careers (sig-reac). references barnacle, r., & dall’alba, g. (2011). research degrees as professional education? studies in higher education, 36(4), 459-470. http://dx.doi.org/10.1080/03075071003698607 cantwell, b. (2011). academic in-sourcing: international postdoctoral employment and new modes of academic production. journal of higher education policy and management, 33(2), 101-114. http://dx.doi.org/10.1080/1360080x.2011.550032 evans, l. (2011). the scholarship of researcher development: mapping the terrain and pushing back boundaries. international journal for researcher development, 2(2), 75-98. http://dx.doi.org/10.1108/17597511111212691 laudel, g., & glaser, j. (2008). from apprentice to colleague: the metamorphosis of early career researchers. higher education, 55, 387-406. http://dx.doi.org/10.1007/s10734-007-9063-7 frontline learning research 3 (2014) 50-63 issn 2295-3159 corresponding author: markus gebhardt, school of education, tu münchen, arcisstraße 21, 80333 münchen, germany, markus.gebhardt@tum.de http://dx.doi.org/10.14786/flr.v2i1.73 50 | f l r basic arithmetical skills of students with learning disabilities in the secondary special schools: an exploratory study covering fifth to ninth grade markus gebhardt a , fabian zehner a , marco g. p. hessels b a tu münchen, germany b university of geneva, switzerland article received 27th november 2013 / revised 17th march 2014 / accepted 17th march 2014 / available online 25th april 2014 abstract the mission of german special schools is to enhance the education of students with special educational needs in the area of learning (sen-l). however, recent studies indicate that students with sen-l from special schools show difficulties in basic arithmetical operations, and the development of basic mathematical skills during secondary special school is not warranted. this study presents a newly developed test of basic arithmetical skills, based on already established tests. the test examines the arithmetical skills of students with sen-l from fifth to ninth grade. the sample consisted of 110 students from three special schools in munich. testing took place in january and june 2013. the test shows to be an effective tool that reliably and precisely assesses students’ performance across different grades. the test items can be used without creating floor and ceiling effects among fifth to ninth grade students with sen-l. the items’ conformity to the dichotomous rasch model is demonstrated. the students’ skills turn out to be very heterogeneous, both overall and within grades. many of the students do not even master basic arithmetical skills that are taught in primary school, although achievement improves in higher grades. keywords: arithmetical skills; curriculum based measurement; special needs gebhardt et al. 51 | f l r 1. schooling of students with sen the schooling of children with special educational needs (sen) is a controversial issue in school policies (european agency for development in special needs education, 2007). it has been shown that students in integrative educational settings show superior school performance (particularly in mathematics) and, in the long run, show greater social skills than students in special schools (baker, wang, & walberg, 1995; carlberg & kavele, 1980; eckhart, haeberlin, sahli lozano, & blanc, p., 2011; haeberlin, blanc, eckhart, & sahli-lozano, 2012; haeberlin, bless, moser, & klaghofer, 1991; merk, 1982; wang & baker, 1986). longitudinal research among students with sen in german-speaking regions showed a delay in school achievement of at least two years compared to children of a corresponding grade in a regular school (haeberlin et al., 1991). the hamburg school trials showed that the performance gap appeared in second grade and increased up to fourth grade, even in classes with particularly good inclusive care (hinz, katzenbach, rauer, schuck, wocken, & wudtke, 1998). cross-sectional studies confirm these findings (tent, witt, bürger, & zschoche-lieberum, 1991; wocken, 2000, 2005; wocken & gröhlich, 2007). seventh grade students with sen-l in special schools did not accomplish the requirements of fifth grade students in a general-education secondary school (hauptschule; wocken, 2000). in germany in 2010, however, only 22% of the students with sen and 23% of the students with sen in the area of learning (sen-l) were in integrative settings (sekretariat der ständigen konferenz der kultusminister der länder in der bundesrepublik deutschland, 2010). nevertheless, the integration rate is rising slowly. in the usa, the statistics about the school performance of students with sen draw a similar picture. in the special education elementary longitudinal study (seels; schiller, sandford, & blackorby, 2008), children with sen between the ages of 10 and 17 (n=5400) were observed over a period of six years. results showed that 60% of students with learning disabilities (ld) in segregated settings and 32% of students with ld in integrative classes achieved the lowest performance level in mathematics (lower than the 20 th percentile; schiller et al., 2008). in secondary school, the performance gap between the students with and without sen continues to widen. in ninth grade, the delay ranges from 3 to 4.9 years on average for students with ld, 1 to 3 years for students with emotional disturbance and more than five years for students with intellectual disabilities (blackorby, chorost, garza, & guzman, 2003). the individual growth over three school years varies widely, but in general, there are no significant differences in the magnitude of growth between the students with different types of sen (blackorby et al., 2003). this kind of longitudinal study is missing in the german speaking countries. 2. identification of students with sen-l in germany in almost all school systems, children with sen are identified to give them a legal right to additional resources and support in school, but the concepts of ld vary widely from country to country. as a consequence, the size of the population of children with diagnosed ld is different in any given country (sideridis, 2007). in the usa, for example, 5% of the student population is classified as having ld (hallahan, lloyd, kauffman, weiss, & martinez, 2005). in germany, 3% of all students are identified as students with sen-l (kmk statistics, 2010). these students have basic difficulties in various learning areas. traditionally, in german-speaking countries, next to pervasive difficulties in school learning, an iq below 85 (but above 70, thus excluding intellectual disability) was considered as the most effective diagnostic criterion of sen-l, since this allowed a general “objective” assessment of a child’s cognitive performance without using school indicators (grünke, 2004). the categorization of students with sen-l in germany is similar to the international definition of ld by lloyd, keller, and hung (2007). this definition refers to significant academic difficulties in school, for which neither other disabilities (e.g., sensory impairment, intellectual disability or emotional and behavioral disorders) nor lack of schooling can be found as cause (lloyd et al., 2007). students with a diagnosed dyslexia or dyscalculia are not identified as students with sen in germany (büttner & hasselhorn, 2011). identification of students with sen-l and, therefore, the gebhardt et al. 52 | f l r allocation of special educational resources to the school only applies to children with severe learning difficulties (klauer & lauth, 1997; schröder, 2008). since the diagnosis of sen-l appears not caused by somatic-medical reasons, but rather by the specific criteria of a given school system, the diagnosis of sen-l is under constant legitimacy pressure. iq testing has been criticized since the 1970s (bundschuh, 2010), both by psychologists and, especially, by teachers and educational practitioners, and consequently, iq is no longer used as the sole indicator of sen-l in present governmental recommendations in germany. nevertheless, many researchers still regard low intellectual abilities as the most important aspect of diagnosing sen-l (kretschmann, 2006) and recommend the administration of a language-free iq test in addition to standardized academic achievement tests as part of the diagnostic process (kany & schöler, 2009; kottmann, 2006). we hope that the instrument under construction that is presented in this article will provide an additional means for improved objective diagnosis of sen-l in the future. 3. basic mathematical skills one third of the students with sen-l, who have graduated from special schools, cannot handle numbers adequately and also have great trouble solving simple division tasks (lehmann & hoffmann, 2009). students show problems with the understanding of word problems, division, the decimal system, and the doubling or halving of numbers (moser opitz, 2007). the lack of elementary arithmetic skills is mainly responsible for mathematical difficulties in secondary school. basic mathematical skills require knowledge of quantity and numbers as well as operation rules (ehlert, fritz, arndt, & leutner, 2013; ennemoser, krajewski, & schmidt, 2011). a cross-sectional study by krajewski and ennemoser (2010) showed that basic skills are not only acquired in elementary school, but also trained in secondary school classes. however, the level of mastery of these basic skills of students in different school tracks is very diverse. high school fifth graders in gymnasium (grammar school) show better mastered basic skills than students in the eighth grade of hauptschule (lower track of secondary school; ennemoser et al., 2011). only one study exists in integrative classes which includes students with sen. an austrian study carried out in urban integrative classes showed that the level of basic skills was also very heterogeneous (gebhardt, schwab, schaupp, rossmann, & gasteiger-klicpera, 2012). even pupils without sen-l had great difficulties in basic arithmetic. as a matter of fact, more than 30% of the regular students (without sen) in fifth grade scored more than one standard deviation below the mean on a standardized school test (lower than the 16 th percentile). students with sen-l were able to solve tasks regarding additions and subtractions, but had significant problems with tasks concerning multiplications and divisions in the number range up to 10,000 (gebhardt et al., 2012). in german-speaking regions, research on the academic performance of students with sen-l is mostly performed in intervention studies (hecht, sinner, kuhl, & ennemoser, 2011; moog, 1993, 1995; moog & schulz, 1997, 2005; sinner & kuhl, 2010). these studies generally observed significant effects immediately following the interventions, but follow-up results again showed large differences between students with sen-l and regular students with learning difficulties. when the training in basic mathematical skills ended, the students with sen-l regressed to the same low level they showed before the intervention (hecht et al., 2011; sinner & kuhl, 2010). all intervention studies used grade based standardized school-tests, which were constructed with classical test theory. however, when overlooking these various studies, which show the specific difficulties of students with sen-l, it would be very useful to have one diagnostic tool that addresses the various arithmetic sub-skills and that is specifically tailored to this special population. 4. research question special needs students show a oneto three-year delay in their development of basic arithmetic skills. the problem with standardized school tests is that they were developed and standardized for average students in the regular curriculum and, as a consequence, have difficulty displaying the academic growth of gebhardt et al. 53 | f l r students with sen-l. adapting such tests raises challenges with respect to the measurement’s discriminatory power (e.g., ceiling and floor effects). another possibility is to use curriculum-based measurements (cbm) to examine academic growth of students with sen (deno, 2003). tests that are actually available were constructed with classical test theory. however, to measure academic progress, item response theory would be the better option (klauer, 2011; wilbert & linnemann, 2011) since these models avoid certain methodological flaws that are associated with tests constructed with classical test theory (such as unreliability of the change scores and incomparability of the scale units of the subsequent measures). our goal is to longitudinally assess the students’ arithmetic skills and to evaluate the achievements of students of different ages, both criterion-based and norm-based. this can be achieved by using instruments that show conformity to specific models from item response theory. assessing basic arithmetical skills, the instrument developed in the longitudinal study on student development in integrative classes silke (schulische integration im längsschnitt – kompetenzentwicklung bei schülerinnen mit und ohne spf in der sekundarstufe i; academic integration in a longitudinal study – development of competences of students with and without sen in secondary schools; gebhardt, 2013; gebhardt, schwab, krammer, & gasteiger-klicpera, 2012; gebhardt, schwab, schaupp et al, 2012; schwab, 2013), is used in this study to assess sen-l students in separated special schools. in contrast to students without sen, students in these special secondary schools are still explicitly taught in elementary arithmetical skills and these need to be addressed in the test. the aims of this pilot study, hence, are the following: − apply the instrument assessing basic arithmetical skills to assess the arithmetical skills of a sample of sen-l students and evaluate the scale’s conformity to the dichotomous rasch model. − explore the instrument’s characteristics regarding discriminatory power, as well as classical psychometric criteria. − explore the basic arithmetical achievement of students with sen-l in special schools, especially in respect to its development across the secondary school grades (cross-sectional), across one school year (longitudinal), as well as the interaction between these two factors. 5. method 5.1 design and sample the study was carried out in three special schools in munich in january and june 2012, which constitute the middle (t1) and the end (t2) of the school term, respectively. at both times of measurement, 62 male and 48 female students (n = 110) with sen-l from fifth to ninth grade were tested with the same instruments. at t1, students were 13.9 years old on average (sd = 1.6). students took tests in groups in sessions of about 15 to 20 minutes, but they could take as much time as needed. if a student did not answer an item, the test administrator reminded the student to do his very best to do so. as all items comprise free response formats guessing behavior can be neglected. table 1 shows the distribution of the sample across grades. gebhardt et al. 54 | f l r table 1 distribution of participants across school grades grade n female male age 5 20 (18%) 35% 65% 11.9 (0.6) 6 23 (21%) 48% 52% 13.1 (0.7) 7 14 (13%) 43% 57% 13.8 (0.6) 8 33 (30%) 48% 52% 15.0 (0.6) 9 16 (15%) 38% 62% 16.0 (0.6) total 110 (100%) 44% 56% 13.9 (1.6) 5.2 instruments on the basis of the arithmetic tests eggenberger rechentest 3+ (ert 3+; holzer, schaupp, & lenart, 2010) and ert 4+ (schaupp, lenart, & holzer, 2010), an instrument was devised that consists of the ert-scales, as well as additional, newly constructed items to handle the large heterogeneity in the target population and to avoid floor and ceiling effects. the ert was originally designed to assess arithmetical skills at the end of the third (3+) and the fourth grade (4+) of elementary school. ennemoser et al. (2011) differentiate arithmetic skills into knowledge of quantity as well as numbers and operation rules. in the currently devised instrument this differentiation is reflected in its subtests: knowledge of quantity is represented by the subtests writing numbers from dictation and number series; numbers and operation rules is represented by the subtests basic arithmetical skills and word problems. for the adapted instrument, the 12 items of the ert 4+ subtest number series were used, which measures knowledge about the place-value system. furthermore, the subtest basic numeracy (comprising 13 items) was used, dealing with addition, subtraction, multiplication and division. the placeholder task is another subtest taken from ert 4+, consisting of 6 items in which 2 numbers are given and the student has to find the third (e.g., ___ + 8 = 21). the subtest word problems comprise 9 items and was taken from ert 3+ to match the students’ levels and to avoid floor effects. table 2 presents the final instrument with its four subtests. table 2 subtests of the final instrument before item-selection procedure subtest origin n items basic arithmetical skills ert 4+: basic numeracy 13 ert 4+: placeholder 6 constructed by authors 15 word problems ert 3+: word problems 9 p re cu rs o rs number series ert 4+ number series 12 constructed by authors 2 writing numbers from dictation constructed by authors 14 5.3 analyses to test the subtests’ unidimensionality, the data were checked for conformity to the dichotomous rasch model. this means, all items pertaining to the same subtest were scaled in one model. then, to check the models’ conformity with regard to specific objectivity, the independence of item parameters across subsamples was evaluated. these subsamples were chosen using two split criteria: raw score median (thus creating two achievement groups) and gender (kubinger, 2005). andersen’s likelihood ratio test (lrt; gebhardt et al. 55 | f l r andersen, 1973), which is based on conditional maximum likelihood estimates, was used to indicate items’ conformity or non-conformity. for testing the items’ fit to the model, the so-called waldtest was used, which indicates the item parameter’s deviance from the model while taking the estimates’ standard error into account (fischer & scheiblechner, 1970). all analyses reported in this article were conducted with the software r (r core team, 2013) and more specifically the package erm (mair, hatzinger, & maier, 2012) which was used for estimating item parameters and calculation of goodness of fit tests, as well as the package pp (reif, 2012) for estimating person parameters. to analyze students’ ability and development in arithmetical skills, the person (ability) parameters were estimated using the item parameters from t1. these allowed to estimate person parameters for t1 as well as for t2 and, consequently, to map these abilities on one scale. in this case, warm maximum likelihood estimates were used, as these allow for the estimation of extreme abilities, especially regarding possible 0 scores in the sen-l group. 6. results 6.1 scaling and item-selection procedure the scaling process was based on the data of t1 and afterwards crosschecked with the data of t2, taking into account its interdependency. after removing two items from the subtest word problems and one item from each of the other subtest, all items showed conformity to the dichotomous rasch model. the subsequent quasi-cross-validation using t2 data was also successful. only for word problems and writing numbers from dictation the gender effect reached significance, but all other tests were not significant. table 3 presents the statistical values of the final rasch models for the four subtests. the andersen lrts showed to be not significant for the final selection of items (1% level of significance was chosen to avoid accumulation of type-i-errors; cf. kubinger, 2005), which indicates conformity to the dichotomous rasch model, both with respect to t1 and t2 data. as the andersen lrt uses cml-estimates, item parameters could not be estimated for items that were solved by all or never solved in the subsamples (the number of items is labeled with na in table 3). table 3 statistical values of the final rasch models for the four subtests split criterion lrt ² df  2 α=.01 p items removed na basic arithmetical skills t1 raw score median 42.6 28 48.3 .04 1 3 gender 45.8 31 52.2 .04 0 t2 raw score median 31.1 28 48.3 .32 3 gender 23.5 32 53.5 .86 0 number series t1 raw score median 9.4 10 23.2 .50 1 2 gender 16.2 11 24.7 .14 1 t2 raw score median 10.8 8 20.1 .22 4 gender 22.6 11 24.7 .02 1 word problems t1 raw score median 6.6 3 11.3 .08 2 3 gender 12.1 5 15.1 .03 1 t2 raw score median 5.0 4 13.3 .28 2 gender 14.6 5 15.1 .01 1 gebhardt et al. 56 | f l r writing numbers from dictation t1 raw score median 8.8 6 16.8 .18 1 6 gender 12.5 11 24.7 .33 1 t2 raw score median 16.8 7 18.5 .02 5 gender 25.7 12 26.2 .01 0 note. all tests show to be not significant at 1% level, indicating conformity to the rasch model. the nacolumn indicates the number of items that could not be evaluated due to 0% or 100% correct in the subsample. to illustrate the results of the item-selection procedure, a graphical representation of the model check of the subtest basic arithmetical skills at t2 is shown in figure 1. nearly all items are situated in the region of acceptable deviance, which is indicated by the gray control line. acceptable deviance is defined in regard to the standard error of estimations in the respective area on the logit scale (cf. wright & stone, 1999). furthermore, the standard errors of the estimations appear to be in an acceptable range (min = 0.2, mean = 0.3, max = 0.8, across all subtests and both times of measurement). gebhardt et al. 57 | f l r figure 1. graphical model checks of the subtest basic arithmetical skills (top) and number series (bottom) by raw score-median (left) and gender (right). the gray line indicates the limit of acceptable deviance for single items (cf. text). finally, the total instrument with the four subtests comprising 33, 8, 12 and 13 items, respectively, also showed conformity to the dichotomous rasch model. table 4 shows that the items present a wide range of difficulty levels, both overall and across grades in nearly every subtest, leading to a reliable assessment across a broad range of ability. only the subtest writing numbers from dictation shows a more narrow range of item difficulty for 9 th graders, which might lead to a small ceiling effect for these students. table 4 proportion correct within subtests across grades, including all selected items. basic arithmetic number series word problems writing numbers grade lo hi m lo hi m lo hi m lo hi m 5 6 7 8 9 .00 .00 .00 .00 .00 .85 .91 .93 .97 1.00 .27 .39 .46 .54 .68 .00 .04 .14 .21 .44 1.00 1.00 1.00 .97 1.00 .40 .57 .70 .67 .85 .00 .00 .00 .03 .19 .65 .91 1.00 1.00 1.00 .24 .30 .41 .44 .62 .11 .30 .36 .42 .81 .95 1.00 1.00 1.00 1.00 .50 .69 .79 .81 .97 note. lo = lowest value, hi = highest value, m = mean the subtest reliabilities (cronbach α) are presented on the diagonal of table 5. the reliabilities vary from .72 to .92, which is above the conventional cut-off-value of .80, except for the subtest word problems, of which the reliability is still very acceptable. it should be mentioned that items that function conform the rasch model are, as such, internally consistent because unidimensionality is included in the theoretical formulation of the model. table 5 further reports high inter-correlations between the subtests, ranging from .64 between number series and writing numbers from dictation to .74 between number series and word problems. gebhardt et al. 58 | f l r table 5 reliabilities and inter-correlations between the subtests at t1 (1) (2) (3) (4) basic arithmetic skills (1) .92 .75 .74 .72 number series (2) .86 .67 .64 word problems (3) .72 .66 writing numbers (4) .85 note. the subtests’ cronbach α is presented on the diagonal. 6.2 basic arithmetical achievement of students with sen-l students’ achievement, in the form of their person (ability) parameter, was very heterogeneous in every subtest and in every grade. person parameters referring to the subtest basic arithmetical skills showed standard deviations from 1.7 (on the logit scale) in grade six to 2.4 in grade nine. the dispersion in achievement did not show a trend across grades in terms of reduced or increased standard deviations. linear regression shows that achievement in every subtest at t1 is predicted by grade (α = .05), with effects ranging from β = 0.47 in word problems to β = 0.58 in basic arithmetical skills. these relations were also significant at t2, but decreased in effect size, which were now ranging from β = 0.37 in word problems to β = 0.41 in writing numbers from dictation. the moderate relationships between grade and ability confirm the instrument’s developmental validity. however, it must be noted that students from grades 7 and 8 showed very similar levels of achievement in every subtest and at both measurement points, except for basic arithmetical skills, in which 8 th grader scored 0.7 logits higher than 7 th graders at t1, but this difference vanished at t2. when shifting from cross-sectional analysis to a longitudinal analysis of the development of achievement from t1 to t2, further differences between the subtests become evident. two subtests appeared to group together with regard to development of mean achievement: in the basic arithmetical skills and the writing numbers for dictation subtests, students from lower grades somewhat improved over time, while those from higher grades regressed (see figure 2). in the other two subtests, number series and word problems, students from every grade improved over time. however, these are descriptive tendencies and in terms of significance only number series showed a longitudinal main effect (d = 0.22). an anova for repeated measurements shows that the factor time plays a significant role, f(1, 101) = 8.6, p = .00, η² = .08. although the interaction term did not reach significance, especially students from grade five increased in their achievements (+1.3 logits). a significant interaction effect between development and grade was found in basic arithmetical skills: f(1, 101) = 3.9, p = .01, η² = .14. this indicates that students in lower grades improve their basic arithmetical skills over time while those in higher grades do not, or even drop in performance (d = 0.34 for 5 th graders, d = 0.20 for 6 th graders, d = -0.12 for 7 th graders, d = -0.53 for 8 th graders, d = -0.41 for 9 th graders). gebhardt et al. 59 | f l r figure 2. development of person (ability) parameter distributions from t1 to t2 of the two subtests basic arithmetical skills (left) and number series (right), for each grade separately. research in the field of special education is particularly interested in the students’ performance variations. table 6 shows the students’ mean ability parameters on all 4 subtests and for all grades separately, in the context of temporal development. the fifth and sixth graders show improvement on all subtests. however, the mean scores of students in 7 th , 8 th and 9 th grade decreased in basic arithmetical skills and writing numbers for dictation, but remained stable or improved on number series and word problems. overall, a regression-to-the-mean-effect was found. i.e., students with low scores tended to improve their scores whereas students with high scores tended to show a decrease at t2. this was confirmed by weak to moderate negative correlation between learning gains (t2 t1) and achievement at t1 in basic arithmetic skills (r = -.55), number series (r = -.38), word problems (r = -.38) and writing numbers (r = -.35). table 6 mean (m) values and standard deviations (sd) of person parameters per grade at t1 and t2 basic arithmetic skills number series grade m t1 m t2 sd t1 sd t2 m t1 m t2 sd t1 sd t2 5 -2.0 -1.4 2.2 1.7 -1.0 0.4 2.4 1.9 6 -0.6 -0.4 1.0 1.4 0.7 1.2 1.8 2.2 7 -0.2 -0.4 1.6 2.2 1.6 1.7 1.1 1.1 8 0.5 -0.2 1.5 1.3 1.4 1.9 1.8 1.8 9 1.7 1.2 1.1 1.2 3.2 3.3 1.5 1.5 word problems writing numbers f. dictation grade m t1 m t2 sd t1 sd t2 m t1 m t2 sd t1 sd t2 5 -2.4 -1.9 2.2 2.5 0.3 0.6 2.2 2.6 6 -1.6 -1.3 1.7 1.7 2.0 2.3 2.0 2.0 7 -0.7 -0.3 1.9 2.4 3.0 2.6 1.5 1.8 8 -0.4 -0.4 1.7 2.3 3.0 2.4 1.8 2.0 9 1.0 1.3 2.4 2.1 4.8 4.3 1.0 1.3 gebhardt et al. 60 | f l r 7. discussion the instrument described in this article showed conformity to the dichotomous rasch model. it also did not show remarkable ceiling or floor effects and, thus, allowed to measure basic arithmetical performance of students with sen-l in special schools. only the newly constructed subtest writing numbers from dictation showed a somewhat narrow range of item difficulties for 9 th graders. this is not unexpected, since these students should already have acquired the basic competence of knowledge of quantity (krajewski & ennemoser, 2010). it would further be questionable if additional, more difficult items would measure the same construct. two items of the subtest word problems, which was taken from the ert 3+, had to be rejected and this scale should be further improved. nevertheless, the instrument showed similar results as those found in the silke study in integrative classes (gebhardt, 2013; gebhardt, schwab, schaupp et al., 2012; schwab, in press) and allowed a first exploration of the basic performance of students with sen-l in special schools. generally, students with sen-l lag several years behind their peers without sen. they are still learning what the other students learn in primary school and especially the basics of multiplication and division are taught to them in secondary school (see also moser opitz, 2007). the inter-correlations of the subtest showed that the performance levels were similar across the subtests and, empirically, it would be sufficient to describe a student with only one scale score, indicating arithmetical ability. however, since the subtest scores are indicative of the development of different arithmetical skills, these should provide support for fitting an appropriate arithmetic curriculum of students with sen-l. thus, the results should help improve the construction of real curriculum based measurement of arithmetic for students with learning disabilities. the instrument discriminated between the grades. although the grade level showed medium effects on all subtests at t1 and t2, the heterogeneity of student performance within the grades was very large. this means that it is necessary to have different mathematical problems with varying levels of complexity available to be able foster the mathematical abilities of all students (moser opitz, 2007). similar findings were described previously in several intervention studies (hecht et al., 2011; moog & schulz, 1997, 2005; sinner & kuhl, 2010), but until now, the arithmetical performance of students with sen had not been measured with a rasch scaled standardized test. one important finding of the longitudinal results was that students from every grade improved on the subtests number series and word problems, while only the 5 th and 6 th graders improved on the subtests writing numbers from dictation and basic arithmetical skills. this might be explained by the fact that the curriculum in 5 th and 6 th grade includes teaching basic arithmetical skills, whereas the curriculum of grades 7 to 9 prepares the students for vocational training. in these grades, basic skills are no longer explicitly trained, but instead, new operations such as fractions are introduced. as the old skills are not explicitly consolidated, basic arithmetic skills (including writing numbers form dictation) and from 3 rd grade in primary school may again become a challenge for students in the 9 th grade of special schools (see, e.g., steiner, 2009). another factor influencing the results, might be that the special school students who are performing well in 5 th and/or 6 th grade can attain integrative classes in 7 th grade. since such students “disappear” to other classes or schools, the cross-sectional data presented here cannot be interpreted in the same way as real longitudinal data. the present data must be viewed as giving explorative information, also when considering the relatively small sample that was included in this study. a much larger sample must be tested to draw stronger conclusions. finally, the development of basic arithmetical skills in this study was relatively limited. this underlines the challenge of teaching basic arithmetical skills in special schools and the, currently, rather limited success. instruments such as the one presented in this article, that allow the continuous measurement of a series of arithmetical skills in secondary special education, may help to further develop evidence based interventions that are tailored to the needs of the students. when measurement and intervention are adapted to the needs of the students, they can jointly help in improving the students’ arithmetic abilities. gebhardt et al. 61 | f l r keypoints students with special educational needs from special schools show difficulties in basic arithmetical operations. a newly developed rasch scaled instrument allows the reliable measurement of basic arithmetical skills of students with sen-l in secondary education. students’ skills turn out to be very heterogeneous, both overall and within grades. many students do not even master arithmetical skills that are taught in primary school, although achievement improves in higher grades. references andersen, e. b. (1973). a goodness of fit test for the rasch model. psychometrika, 38(1), 123–140. doi: 10.1007/bf02291180 baker, e. t., wang, m. c. & walberg, h. j. (1995). the effect of inclusion on learning. educational leadership, 52(4), 33–35. blackorby, j., chorost, m., garza, n., & guzman, a. m. (2003). the academic performance of secondary school students with disabilities. in u.s. department of education (eds.), the achievement of youth with disabilities during secondary school. a report from the national longitudinal transition study 2. menlo park, ca: sri international. retrieved from http://www.seels.net/designdocs/seels_w1w3_final.pdf bundschuh, k. (2010). einführung in die sonderpädagogische diagnostik (7th ed.). münchen: e. reinhardt. büttner, g. & hasselhorn, m. (2011). learning disabilities: debates on definitions, causes, subtypes and responses. international journal of disability, development and education, 58(1), 75–87. doi: 10.1080/1034912x.2011.548476 carlberg, c. & kavale, k. (1980). the efficacy of special versus regular class placement for exceptional children: a meta-analysis. the journal of special education, 14(3), 295–309. doi: 10.1177/002246698001400304 deno (2003). curriculum-based measurment. journal of special education, 37, 184–192. doi:10.1177/00224669030370030801 eckhart, m., haeberlin, u., sahli lozano, c., & blanc, p. (2011). langzeitwirkungen der schulischen integration. [long-term effects of school integration]. bern, switzerland: haupt verlag. ehlert, a., fritz, a., arndt, d. & leutner, d. (2013). arithmetische basiskompetenzen von schülerinnen und schülern in den klassen 5 bis 7 der sekundarstufe. journal für mathematik-didaktik, 34(2), 237– 263. doi:10.1007/s13138-013-0055-0 ennemoser, m., krajewski, k., & schmidt, s. (2011). entwicklung und bedeutung von menge-zahlenkompetenzen und eines basalen konventionsund regelwissens in der klasse 5 bis 9. zeitschrift für entwicklungspsychologie und pädagogische psychologie, 34(4), 228–242. doi: 10.1026/00498637/a000055 european agency for development in special needs education (2007). assessment in inclusive settings. key issues for policy and practice. odense, denmark: european agency for development in special needs education. fischer, g. h., & scheiblechner, h. h. (1970). algorithmen und programme für das probabilistische testmodel von rasch. [algorithms and programs for rasch’s probabilistic test model.]. psychologische beiträge, 12, 23–51. gebhardt, m. (2013). integration und schulische leistungen in grazer sekundarstufenklassen. eine empirische explorative pilotstudie. wien: lit verlag. gebhardt, m., schwab, s., krammer, m., & gasteiger-klicpera, b. (2012). achievement and integration of students with special needs (sen) in the fifth grade. journal of special education and rehabilitation, 13, 7–19. doi: 10.2478/v10215-011-0022-6 http://dx.doi.org/10.2478/v10215-011-0022-6 gebhardt et al. 62 | f l r gebhardt, m., schwab, s., schaupp, h., rossmann, p., & gasteiger-klicpera, b. (2012). heterogene gruppen in mathematischen grundfertigkeiten: eine explorative erkundung der fähigkeiten im grundrechnen in integrationsklassen der 5. schulstufe. zeitschrift für inklusion (online), (1-2). available from www.inklusion-online.net/index.php/inklusion/article/view/155/147 grünke, m. (2004). lernbehinderung. in g. w. lauth, & m. grünke (eds.), interventionen bei lernstörungen. förderung, training und therapie in der praxis (pp. 65–77). göttingen: hogrefe, verl. für psychologie. haeberlin, u., blanc, p., eckhart, m., & sahli-lozano, c. (2012, may). intégration scolaire d'enfants en difficultés d'apprentissage: effets à long terme. information sur la recherche éducationnelle, csre, n° 12:021. [school integration of children with learning difficulties: long-term effects. information about educational research, csre, n° 12:021]. retrieved may 31, 2013 from www.skbfcsre.ch/de/bildungsforschung/datenbank/. haeberlin, u., bless, g., moser, u., & klaghofer, r. (1991). die integration von lernbehinderten: versuche, theorien, forschungen, enttäuschungen, hoffnungen. bern: haupt. hallahan, d. p., lloyd, j. w., kauffman, j. m., weiss, m. p., & martinez. (2005). learning disabilities: foundations, characteristics, and effective teaching. needham heights: allyn & bacon. hecht, t., sinner, d., kuhl, j., & ennemoser, m. (2011). differenzielle effekte eines trainings der mathematischen basiskompetenzen bei kognitiv schwachen grundschülern und schülern der förderschule mit dem schwerpunkt lernen – reanalysen zweier studien. empirische sonderpädagogik, (4), 308–323. hinz, a., katzenbach, d., rauer, w., schuck, k. d., wocken, h. & wudtke, h. (1998). die integrative grundschule im sozialen brennpunkt: ergebnisse eines hamburger schulversuchs. hamburg: hamburger buchwerkstatt. holzer, n., schaupp, h., & lenart, f. (2010). eggenberger rechentest (ert 3+): diagnostikum für dyskalkulie für das ende der 3. schulstufe bis mitte der 4. schulstufe. bern: huber. kany, w., & schöler, h. (2009). diagnostik schulischer lernund leistungsschwierigkeiten: ein leitfaden mit einer anleitung zur gutachtenerstellung. stuttgart: kohlhammer. klauer (2011). lernverlaufsdiagnostik – konzept, schwierigkeiten und möglichkeiten. empirische sonderpädagogik, (3), 207–224 klauer, k. j., & lauth, g. w. (1997). lernbehinderungen und leistungsschwierigkeiten bei schülern. in f. e. weinert (eds.), enzyklopädie der psychologie, psychologie des unterrichts und der schule (pp. 701–738). göttingen: hogrefe. kubinger, k. d. (2005). psychological test calibration using the rasch model – some critical suggestions on traditional approaches. international journal of testing, 5(4), 377–394. lauth, g. w., & grünke, m. (eds.). (2004). interventionen bei lernstörungen: förderung, training und therapie in der praxis. göttingen: hogrefe, verl. für psychologie. kottmann, b. (2006). selektion in die sonderschule: das verfahren zur feststellung von sonderpädagogischem förderbedarf als gegenstand empirischer forschung. bad heilbrunn: klinkhardt. krajewski, k., & ennemoser, m. (2010). entwicklung mathematischer basiskompetenzen in der sekundarstufe. empirische pädagogik, 24(4), 353–370. kretschmann, r. (2006). diagnostik bei lernbehinderungen. in u. petermann, & f. petermann (eds.), diagnostik sonderpädagogischen förderbedarfs (pp. 139–162). göttingen: hogrefe. lehmann, r., & hoffmann, e. (2009). berliner erhebung arbeitsrelevanter basiskompetenzen von schülerinnen und schüler und schüler mit dem förderbedarf „lernen“. münster: waxmann. lloyd, j. w., keller, c., & hung, l. (2007). international understanding of learning disabilities. learning disabilities research & practice, 22(3), 159–160. doi: 10.1111/j.1540-5826.2007.00240. mair, p., hatzinger, r., & maier, m. j. (2012). erm: extended rasch modeling. r package version 0.15–1. merk (1982). lernschwierigkeiten – zur effizienz von fördermaßnahmen an grundund lernbehindertenschulen. heilpädagogische forschung, 1, s. 53–69. moog, w. (1993). schwachstellen beim addieren – eine erhebung bei lernbehinderten sonderschülern. zeitschrift für heilpädagogik, 44, 534–554. gebhardt et al. 63 | f l r moog, w. (1995). flexibilisierung von zahlbegriffen und zählhandlungen – ein übungsprogramm. heilpädagogische forschung, 21(3), 113–121. moog, w., & schulz, a. (1997). das dortmunder zahlbegriffstraining – lernwirksamkeit bei rechenschwachen grundschülern. sonderpädagogik, 27(2), 60–68. moog, w., & schulz, a. (2005). zahlen begreifen: diagnose und förderung bei kindern mit rechenschwäche (2. überarb. aufl.). weinheim: beltz. moser opitz, e. (2007). rechenschwäche – dyskalkulie: theoretische klärungen und empirische studien an betroffenen schülerinnen und schüler. bern: haupt. r core team. (2013). r: a language and environment for statistical computing. r foundation for statistical computing: vienna, austria. reif, m. (2012). pp: person parameter estimation. r package version 0.2. schaupp, h., lenart, f., & holzer, n. (2010). eggenberger rechentest (ert 4+): diagnostikum für dyskalkulie für das ende der 4. schulstufe bis der mitte der 5. schulstufe. bern: huber. schiller, e., sanford, c. & blackorby, j. (2008). the achievments of youth with disabilities during secondary school: a report from the national longitudinal transition study 2 (u.s. department of education, hrsg.). retrieved from http://www.seels.net/info_reports/seels_learndisability_%20spec_topic_report.12.19.08ww _final.pdf. schröder, u. (2008. lernbehindertenpädagogik: grundlagen und perspektiven sonderpädagogischer lernhilfe (2nd ed.). stuttgart: kohlhammer. schwab, s. (2013). schulische integration, soziale partizipation und emotionales wohlbefinden in der schule ergebnisse einer empirischen längsschnittstudie. berlin: lit-verlag. sekretariat der ständigen konferenz der kultusminister der länder in der bundesrepublik deutschland (2010). sonderpädagogische förderung in schulen 1999 bis 2008: dokumentation nr. 189 – märz 2010. retrieved from http://www.kmk.org/fileadmin/pdf/statistik/dok_189_sopaefoe_2008.pdf. sideridis, g. d. (2007). international approaches to learning disabilities: more alike or more different? learning disabilities research & practice, 22(3), 210–215. doi: 10.1111/j.1540-5826.2007.00249.x sinner, d., & kuhl, j. (2010). förderung mathematischer basiskompetenzen in der grundstufe der schule für lernhilfe. zeitschrift für entwicklungspsychologie und pädagogische psychologie, 42(4), 241– 251. steiner, g. (2009). forgetting while learning: a plea for specific consolidation. journal of cognitive education and psychology, 8, 117–127. tent, l., witt, m., bürger, w., & zschoche-lieberum, c. (1991). ist die schule für lernbehinderte überholt? heilpädagogische forschung, (1), 289-320. wang, m. c. & baker, e. t. (1985-86). mainstreaming programs: design features and effects. the journal of special education, 19(4), 503–521. wilbert, j., & linnemann, m. (2011). kriterien zur analyse eines tests zur lernverlaufsdiagnostik. empirische sonderpädagogik, 3, 225–242. wocken, h. (2005). andere länder, andere schüler?: vergleichende untersuchungen von förderschülern in den bundesländern brandenburg, hamburg und niedersachen. potsdam. retrieved from http://bidok.uibk.ac.at/library/wocken-forschungsbericht.html wocken, h. (2007). fördert förderschule? eine empirische rundreise durch schulen für „optimale förderung“. in i. demmer-dieckmann, & a. textor (eds.), integrationsforschung und bildungspolitik im dialog (pp. 35–60). bad heilbrunn: klinkhardt. wocken, h. & gröhlich, c. (2009). kompetenzen von schülerinnen und schülern an hamburger förderschulen. in: w. bos, m. bonsen & c. gröhlich (hrsg.), kess 7 – kompetenzen und einstellungen von schülerinnen und schülern an hamburger schulen zu beginn der jahrgangsstufe 7 (pp. 133–142). münster: waxmann. wright, b.d., & stone, m.h. (1999). measurement essentials. wide range inc.: wilmington. retrieved from: http://www.rasch.org/measess/me-all.pdf frontline learning research 5 (2014) 28-45 issn 2295-3159 corresponding author: http://dx.doi.org/10.14786/flr.v2i3.96 28 | f l r scientific reasoning and argumentation: advancing an interdisciplinary research agenda in education frank fischer a , ingo kollar a , stefan ufer b , beate sodian a , heinrich hussmann c , reinhard pekrun a , birgit neuhaus d , birgit dorner e , sabine pankofer e , martin fischer f , jan-willem strijbos a , moritz heene a & julia eberle a,d a ludwig maximilians university of munich, department of psychology, germany b ludwig maximilians university of munich, department of mathematics, germany c ludwig maximilians university of munich, department of informatics, germany d ludwig maximilians university of munich, department of biology, germany e katholische stiftungsfachhochschule münchen university of applied sciences, germany f ludwig maximilians university of munich, university hospital, institute for medical education, germany a-f munich center of the learning sciences, germany article received 24 february 2014 / revised 1 april 2014 / accepted 19 may 2014 / available online 16 june 2014 abstract scientific reasoning and scientific argumentation are highly valued outcomes of k-12 and higher education. in this article, we first review main topics and key findings of three different strands of research, namely research on the development of scientific reasoning, research on scientific argumentation, and research on approaches to support scientific reasoning and argumentation. building on these findings, we outline current research deficits and address five aspects that exemplify where and how research on scientific reasoning and argumentation needs to be expanded. in particular, we suggest to ground future research in a conceptual framework with three epistemic modes (advancing theory building about natural and social phenomena, artefact-centred scientific reasoning, and science-based reasoning in practice) and eight epistemic activities (problem identification, questioning, hypothesis generation, construction and redesign of artefacts, evidence f. fischer et al. 29 | f l r generation, evidence evaluation, drawing conclusions as well as communicating and scrutinizing scientific reasoning and its results). we further propose addressing the domain specificities and domain generalities of scientific reasoning and argumentation as well as approaches for facilitation. finally, we argue for investigating the role of epistemic emotions, the role of the social context, and the influence of digital technologies on scientific reasoning and argumentation. keywords: scientific reasoning; argumentation; epistemic emotions; collaboration; technology 1. problem to participate in the knowledge society and to benefit from the unprecedented open access to a vast volume of scientific knowledge requires a broad set of skills and abilities that have lately been labelled as 21 st century skills (e.g., trilling, & fadel, 2009). these include skills and abilities to use scientific concepts and methods to understand how scientific knowledge is generated in different scientific disciplines, to evaluate the validity of science-related claims, to assess the relevance of new scientific concepts, methods, and findings, and to generate new knowledge using these concepts and methods. the acquisition of these complex competencies is considered a main goal and outcome of k-12 and higher education. however, contemporary knowledge about what constitutes these competencies and how they can be facilitated is scattered over different research disciplines. in order to develop a better understanding of these competencies, we propose to build on three existing strands of research. first, research on the development of scientific reasoning (e.g., koslowski, 2012); second, research looking at the processes and products of scientific argumentation (e.g., chinn & clark, 2013) from the fields of educational psychology, education, as well as science education and other subject education disciplines. third, there is a broad range of approaches to support and facilitate scientific reasoning and argumentation (sra) in educational contexts (e.g., furtak, seidel, iverson, & briggs, 2012). in this article, we will first provide an overview of the main topics and key findings of these three strands of research. building on these findings, we outline the deficits of existing research and address five aspects that exemplify where and how research on sra needs to be expanded. 2. key findings of previous research 2.1 development of scientific reasoning research on scientific reasoning amongst laypeople has its roots in developmental psychology. inhelder and piaget (1958) assumed that scientific rationality was a model of the ideal human reasoning, that is, a person who reflects on theories, builds hypothetical models of reality, critically and exhaustively tests for all possible main and interaction effects between variables, and objectively and systematically evaluates evidence with respect to a claim. in a series of studies they showed that the scientific reasoning of preadolescent children was severely deficient, whereas significant improvement took place in adolescence. these findings led them to claim the stage of “formal operational thought” as the highest stage of cognitive development. this view has since been heavily criticised, as it neither adequately captures adult reasoning nor its development (kuhn & franklin, 2006). neither the lay adult nor professional scientists conform to a model of domain-general, ideal scientific rationality. rather, adult reasoning abilities are heavily dependent on domain-specific knowledge and context (e.g., kruglanski & gigerenzer, 2011). this is found for laypersons, but professional scientists f. fischer et al. 30 | f l r are equally influenced by their prior knowledge and theoretical biases (dunbar, 1995). similarly, children’s scientific reasoning is context and task dependent and does not differ fundamentally from adult scientific reasoning (koslowski, 1996, 2012; see zimmerman, 2000, 2007). the “layperson as scientist” metaphor, which focuses on processes of intentional knowledge seeking to test theories and hypotheses and to evaluate evidence with respect to a hypothesis or theory (kuhn & franklin, 2006), has proved to be a productive framework for research into scientific reasoning. however, broad models of scientific reasoning that incorporate early competencies are only now emerging (kuhn, & franklin, 2006; sodian & bullock, 2008). for example, kuhn (1991) showed that differentiation of theory and evidence poses a major problem for many lay adults in complex, real-world argumentation. however, even young elementary school children can differentiate hypothetical beliefs from evidence and identify a conclusive research design to test a hypothesis (sodian, zaitchik & carey, 1991). third graders distinguish controlled from confounded experiments (bullock, & ziegler, 1999). even pre-schoolers possess basic data evaluation competencies (koerber, sodian, thoermer, & nett 2005; koerber b& sodian, 2009). thus, neither children nor adults appear to lack a basic understanding of the relationship between hypothetical beliefs and empirical evidence. rather, in complex theory evaluation tasks, both children and adults appear to lack an understanding of mechanisms, as well as methodological knowledge to provide and judge evidence-based arguments (e.g., koslowski, 2012). a meta-conceptual understanding of the nature of scientific knowledge has been identified as a major source of developmental progress. understanding progresses from an undifferentiated level 1 (science as activities and effects) through an intermediate level 2 (science as providing explanations via testable claims) to a level 3 understanding (science as a cyclical and cumulative process of theory, testing, and revision), with children rarely displaying level 2 and even adults rarely articulating a coherent level 3 understanding (e.g., carey & smith, 1993). however, even the nature of elementary school students’ science understanding can be improved through instructional support (e.g., sodian, jonen, thoermer, & kircher, 2006). moreover, an advanced meta-conceptual understanding of science in childhood has been found to predict strategy acquisition in adolescence (bullock, sodian, & koerber, 2009). recent attempts in developmental research with elementary school students support a model of scientific reasoning as a complex set of interrelated abilities, consisting of four major components: “understanding the nature of science”, “understanding theories”, “designing experiments”, and “interpreting data” (e.g., koerber, sodian, kropf, mayer, & schwippert, 2011). apart from general cognitive abilities, student’s problem-solving skills and spatial abilities have been shown to have a major impact on these scientific reasoning competencies. moreover, scientific reasoning has been shown to be a separate construct from measures of intelligence and reading skills in elementary school students (mayer, sodian, koerber, & schwippert, 2014). 2.2 scientific argumentation while developmental research is mainly interested in the developmental trajectories of an individual’s scientific reasoning, educational and science education research on scientific argumentation has focused on the externalised processes and products of scientific reasoning within social contexts (e.g., the science classroom; osborne, 2010). the interest in scientific argumentation is sparked by the view that argumentation relates to the learning of core content and acquisition of general argumentation skills (chinn & clark, 2013). previous research strived for two main goals: (a) identification of students’ deficits during their engagement in scientific argumentation in social contexts, and (b) design and development of effective scaffolding approaches to improve students’ argumentation. with respect to students’ deficits in scientific argumentation, some studies focused on the structural quality of student-generated arguments, for example on the use of evidence (e.g., mcneill, 2011), qualifiers (stegmann, wecker, weinberger, & fischer, 2012) or warrants (kollar, fischer, & slotta, 2007). a recurring finding has been that students tend to make claims without justifications. in socio-scientific debates, they typically do not spontaneously refer to scientific concepts and information (sadler, 2004). other studies have shown that students often have problems producing arguments of high content quality (e.g., kelly, & takao, f. fischer et al. 31 | f l r 2002). a third set of studies revealed that students often exhibit a poor dialogic or social quality of argumentation as reflected in the social exchange and co-construction of arguments. for example, students have been found to refrain from challenging others’ arguments (weinberger, stegmann, & fischer, 2010). this might be related to the recurring finding that students have difficulties recognising contrasting argumentative positions (sadler, 2004) and are often not successful in integrating different perspectives of different learners within a group or community (noroozi, weinberger, biermans, mulder, & chizari, 2012). 2.3 intervention studies how students can effectively be supported in their acquisition of sra-related skills has been subject to a large body of intervention-based research, including long-term and short-term interventions, technologybased and teacher-based scaffolding, laboratory as well as field studies, and studies at the school and university levels (e.g., kollar et al., 2007; mcneill, lizotte, krajcik & marx, 2006). overall, this research shows that sra can be substantially advanced by making it an explicit topic of instruction (see osborne, 2010). this applies to both increasing students’ abilities to engage in activities of scientific knowledge generation (or epistemic activities) and helping them develop a more sophisticated understanding of the nature of science. current research on instructional approaches focuses on immersing learners into scientific practices (see cavagnetto, 2010) which typically involves student engagement in research-related activities and debates. three prototypical instructional approaches are inquiry learning, problem-based learning, and design-based learning. inquiry learning engages students in more or less authentic activities of hypothesis formulation, generation of evidence, and drawing conclusions (chinn & malhotra, 2002). inquiry learning proved to be an effective instructional approach to advance science learning, especially when combined with teacher-led activities (e.g., furtak et al., 2012). similarly, in problem-based learning, students are confronted with complex problems and expected to find explanations and solutions that are based on scientific concepts and methods (e.g., dochy, segers, van den bossche & gijbels, 2003). design-based learning (e.g., kolodner, 2007) engages students in inter-linked cycles of research and design with the goal of arriving at an optimal design of a concrete product, such as a miniature car that that can go from one end of the classroom to the other. in all of these approaches that aim to immerse students into authentic sra processes, it has been found crucial to provide students with structural support. this scaffolding may be directed at individual learners, small groups and whole classrooms. for individual learning, hints, prompts, sentence starters, and guiding questions that help students focus their attention on the critical aspects of sra have been found to be effective (see quintana, reiser, davis, krajcik, fretz, duncan, & soloway, 2005). a hypothesis scratchpad, for example, helped students formulate better hypotheses than students whose hypothesis formation was unscaffolded (van joolingen, & de jong, 1993). for small-group collaboration, several studies showed that the quality of sra can be raised substantially through collaboration scripts (see fischer, kollar, stegmann, & wecker, 2013), which assign roles to learners and sequence their epistemic activities. for instance, a social-discursive peer review script has been shown to enhance student argumentation. detailed process analyses revealed that social-discursive argumentation during the peer review processes mediated the effects of scaffolding by the script on the improvement of (individual) argumentation skills (stegmann et al., 2012). a related form of structuring collaboration is peer assessment (e.g., cho, schunn, & wilson, 2006; strijbos, & sluijsmans, 2010), which is also a crucial aspect of the contemporary scientific process. peer assessment can be used to help collaborators uncover incongruence in their respective sra processes when scrutinising scientific claims and evidence. the incongruence can subsequently foster refinement of target processes through critical reflection (nicol, thomson, & breslin, 2014). finally, studies demonstrated that teachers can be successfully empowered to help students gain scientific argumentation skills (e.g., erduran, simon, & osborne, 2004). for instance, research on classroom scripts has shown that epistemic activities can be facilitated if teachers combine scaffolding at different social levels in the classroom (plenary, group, individual; e.g., mäkitalo-siegl, kohnle, & fischer, 2011). moving even beyond the boundaries of the classroom, knowledge building communities have been successfully implemented in schools around the globe to engage students in argumentative processes to jointly construct knowledge in the classroom (scardamalia & bereiter, 2006). f. fischer et al. 32 | f l r 3. deficits of prior research and directions for advancing studies on sra research on the development of scientific reasoning, as well as research on scientific argumentation, has substantially progressed over the last two decades (see nussbaum, 2011; zimmerman, 2007). however, there are still important research gaps which leads us to argue for more systematic and interdisciplinary research on sra. we propose that future research should (a) expand the range of epistemic modes and epistemic activities, (b) investigate domain-specific aspects of sra more systematically, (c) examine the role of emotions in sra, (d) consider the social context of sra in a more systematic way, and (e) explore the influence of digital technologies on sra. each of these suggestions is more closely elaborated upon in the following. 3.1 expanding the range of epistemic modes and epistemic activities 3.1.1 epistemic modes people engage in sra with different motivations. for example, a researcher may strive to contribute to theory building in a domain while practitioners try to find solutions for problems in their professional practice by applying scientific concepts or methods. we argue that these different motivations have not yet been systematically reflected in research on sra in educational contexts. stokes (1997) suggested a widely accepted classification according to which approaches to scientific reasoning vary in their primary goals along two orthogonal dimensions: understanding and use. pure basic research is characterised by its primary goal of advancing scientific understanding of natural and social phenomena, regardless of its usefulness in practice. stokes used nils bohr’s scientific approach – with no emphasis on the use and societal uptake of his theoretical advances – to characterise this type of research. in contrast, pure applied research emphasises the use of scientific knowledge without the aim of advancing theory building and understanding. stokes exemplified this kind of research with the work of thomas a. edison, who brought electricity to a whole country by using scientific knowledge and methods, but without being concerned about generalisation and theory building beyond this practical challenge. a third class that stokes (1997) identified is the scientific approach that combines the goals of understanding and use, which he termed “use-inspired basic research” and exemplified with louis pasteur’s work. pasteur started from problems in practice (e.g., how to make food last longer), conducted systematic research to solve them, but simultaneously strived for a generalised theoretical explanation. we suggest that stokes’ classification of research approaches can be used to inform the differentiation of three distinct modes of sra. in a first mode (1) sra can be used to advance theory building about natural and social phenomena. when learners apply this mode, they aim to generate and test hypotheses to develop and improve scientific theories and explanations about social and natural phenomena. that way, this epistemic mode will help support student learning of the scientific knowledge of a domain, how it is created, and how students themselves can contribute to knowledge creation by engaging in scientific research. a second sra mode may be labelled (2) science-based reasoning and argumentation in practice. in this mode, learners aim at developing solutions for contextualised problems using scientific concepts, theories, and methods. based on information about the problem and the state-of-research as they know it, learners generate one or more solution approaches and evaluate them in light of scientific knowledge and methods, but also based on standards of the practice under consideration. that way, learners take over the role of scientifically knowledgeable practitioners rather than that of basic researchers. for example, teacher education students may develop a concept to help 4th graders improve with respect to their reading abilities, based on both practice-based observations of the possibly poor reading abilities of their students and on prior scientific theories and empirical studies on how to effectively support students with reading difficulties (e.g., reciprocal teaching; palincsar & brown, 1984). another example is the application of mathematics to solve practical problems (e.g., predicting the development of sprint world records by describing historical data with an appropriate mathematical function), typically referred to as “mathematical modelling” (galbraith, f. fischer et al. 33 | f l r henn, & niss, 2007). the difference between science-based reasoning in practice and problem solving is that the result is not only the solution of a problem, but also an argument based on scientific theory. the third sra mode we would like to introduce is called (3) artefact-centred sra. this mode is realized when students engage in circular processes which involve the concurrent development of an artefact and a scientific theory or explanation for why the artefact works or does not work (i.e., why a given problem can or cannot be solved by the use of the artefact), through repeated cycles of prototype design, testing, and analysis of test results. for example, kolodner (2007) reports on a science curriculum unit during which students are supposed to build miniature cars from a given set of materials. based on concepts from physics (e.g., friction and force), the students’ task is to design a car that would travel from one end of the classroom to the other. that way, the students’ reasoning and argumentation resembles that of researchers in engineering and technology. this mode of scientific reasoning differs from “science-based reasoning in practice” with respect to the thrust towards generalisation and theory building. nevertheless, in educational contexts, both modes have the potential to address student competence of understanding and engagement in scientific knowledge creation activities, as well as their competence to address practical problems through application of scientific concepts and methods. 3.1.2 epistemic activities the three epistemic modes imply an extended notion of sra that also calls for considering a comprehensive set of scientific activities. students in educational contexts need to learn how these activities work and how to engage in them. we suggest distinguishing eight epistemic activities that all may be fulfilled in sra in all of the three epistemic modes. yet, both the weight that is attributed to each activity in each of the three modes and the way these activities are performed within each of the three modes may differ. in the following, we describe each of these activities along with one example of how the activity may be performed in one of the three epistemic modes. (1) problem identification. many scientific reasoning processes are driven by concrete problems. according to the three epistemic modes, such problems might be practical real-world problems (see kolodner, 2007), but also scientific problems that cannot be solved with the available theoretical concepts and methods. becoming aware that available explanations do not appropriately explain phenomena is a starting point for both the advancement of science as an abstract set of knowledge, and for the individual learner advancing his or her understanding of the world. thus, to engage in sra, one first needs to perceive a mismatch or shortcoming concerning the available explanation of a particular problem. during this epistemic activity, a problem representation is built from an analysis of the situation. a medical student may for example be confronted with a patient who reports a diverse set of illness symptoms (exemplifying the epistemic mode science-based sra in practice). based on medical knowledge, which in medical experts typically is encapsulated in so-called “illness scripts” (charlin, boshuizen, custers, & feltovich, 2007), the student will try to identify which parts of the patient’s descriptions are relevant for the diagnostic process and which are not. that way, the actual biomedical problem is gradually concretized and then determines further action. (2) questioning. based on the representation developed during problem identification, one or more initial questions are identified for the subsequent reasoning process (see white & frederiksen, 1998). later on, this question might be refined to allow for a systematic search of evidence. to exemplify how a math student may be confronted with questioning in the epistemic mode of advancing theory building about natural and social phenomena we refer to the following famous problem formulated by euler in 1741 (seven bridges of königsberg; solution proved by hierholzer & wiener, 1873): in a given arrangement of points and lines between these points (e.g., a set of crossings and streets in a city), how can we determine if an “euler-walk” along adjacent lines is possible, which passes each line exactly once (e.g., a sightseeing walk through the city)? the problem here is a classification problem (how to describe objects with a given property). (3) hypothesis generation. during hypothesis generation, students derive possible answers to the question from plausible models, available theoretical frameworks or empirical evidence they are aware of f. fischer et al. 34 | f l r (klahr & dunbar, 1988). if the student’s prior knowledge does not allow for predictions, the question might be refined or – alternatively – an exploratory approach of evidence generation may be adopted to derive a hypothesis based on patterns in this evidence. this process involves formulating the hypothesis according to scientific standards. in biology, a learner may for example aim at developing an answer to the question how the memory of honey bees develops. based on prior research, the learner may hypothesize that glutamate plays a role in this process, since glutamate has been shown to be important for human memory development. to substantiate this hypothesis, further search for corresponding literature may be necessary, e.g. concerning the question whether glutamate has also been found in other insects. (4) construction and redesign of artefacts. scientific reasoning often includes the construction of some kind of artefact, be it the development of a prototype object by an engineer or an axiomatic system describing a new mathematical structure. typically, this construction will be based on current theoretical knowledge. following its construction, the artefact is submitted to a test in an authentic environment (see kolodner, 2007). for example, teacher students may have the task to develop a computer-based collaborative learning environment that would effectively scaffold the interaction of small groups of learners in order to raise the individuals’ learning outcomes (exemplifying the epistemic mode of artefact-centred sra). for that purpose, a prototype of the learning environment (e.g., based on the collaboration script approach; fischer et al., 2013) may be built that – based on theoretical reasoning and prior empirical evidence – seems promising to achieve this goal. (5) evidence generation. evidence generation includes various approaches. one approach is to conduct hypothetico-deductive experimental studies that refer to the systematic, theory-driven variation of one or more variables by the learner in consecutive trials, while repeatedly observing the same outcome variables. evidence generation may also follow an inductive approach of observing, comparing and describing phenomena to draw conclusions about structures and functions, for example in evolutionary biology or sociology. another approach is observing the synchronous or sequential co-occurrence of phenomena, which is frequently applied in the natural sciences (e.g., when studying climate models), but also in the social sciences (e.g., in longitudinal studies). finally, most natural and social sciences use deductive reasoning – within more or less elaborate theories – to generate evidence in favour or against a claim. in the mathematic example by euler described above, a first approach to gather (exploratory, in the mathematical sense preliminary) evidence would be to study single examples of point-line configurations and test if they admit an euler-walk. comparing configurations which admit such a walk and some which do not, might lead to a first hypothesis about the characteristic difference between the two (hypothesis generation). studying more, and perhaps extreme, examples will add further (still preliminary) evidence to support the hypothesis, maybe leading to its revision or refinement. finally, starting from a set of basic assumptions on such line configurations (described by the axioms of mathematical graph theory), a deductive chain of arguments can be constructed that shows that configurations admitting an euler-walk have the hypothesized property, and vice versa. constructing such a line of deductive arguments, which derive that a conjecture follows from the axioms of a mathematical theory, is actually the main mode of evidence generation in mathematics. nevertheless, also other kinds of evidence play a major role in mathematical reasoning, such as counter-examples that disprove a general conjecture (e.g., zazkis, & chernoff, 2008). (6) evidence evaluation. the aim of evidence evaluation is to assess the degree to which a certain piece of evidence supports a claim or theory. what counts as evidence will differ both with respect to the epistemic mode in which sra is realised and with respect to the domain under study. observational studies (shafto, kemp, bonawitz, coley, & tenenbaum, 2008), for example, might be considered the best available evidence in one discipline (e.g., astronomy) but less valuable than experimental studies in another (e.g., psychology, engineering; kolodner, 2007). deductions from a theoretical framework constitute the crucial acceptance criterion in mathematics, whereas in psychology or in natural sciences they serve an auxiliary role as predictions about the outcomes of an experiment from theoretical assumptions. even though an “experimentum crucis” is not viable in most disciplines, cumulated evidence from several experimental or observational studies is necessary to sustain a claim. an example from medical education in the epistemic mode of science-based sra in practice would be a medical student aiming to find the right diagnosis for a f. fischer et al. 35 | f l r patient’s health problem in a case-based simulation environment. evidence evaluation in this example may refer to the accumulating evidence from the patient´s history, physical examinations and additional lab and technical tests. optimally, this evidence is interpreted in light of candidate diagnoses that have already been set up during hypothesis generation. here the development of encapsulated, experiential knowledge in the form of illness scripts (charlin et al., 2007) has been identified as crucial in order to arrive at a sound evaluation of the collected evidence. (7) drawing conclusions. since different kinds of evidence can be generated within the scientific reasoning process, drawing conclusions is not restricted to reconsidering an initial claim in light of experimental results. different pieces of evidence must often be integrated by weighing each single piece according to the method by which it was generated and by the rules and criteria of the discipline. in the case of a teacher student developing a scaffolded computer-supported learning environment, drawing conclusions means to critically analyse data and observations from an experiment or a field trial in which the environment was used and to derive consequences for whether the environment (or specific features of it) needs to be re-designed or may be used as originally planned in further trials. to arrive at such a conclusion, typically a multitude of data sources needs to be considered (e.g., individual knowledge tests, verbal protocols, data on students’ motivation). (8) communicating and scrutinising. individual scientific reasoning processes and their results are typically shared with and scrutinised by others (shavelson & towne, 2002). persons involved in scientific reasoning are more or less constantly involved in conversations and discussions in work groups or peer groups. these interactions might influence scientific reasoning from problem identification to knowledgebased interventions in practice situations. thus, social-discursive and dialogic argumentation is an integral component of many scientific reasoning processes and should be included when analysing and facilitating sra in educational contexts (e.g., clark, sampson, weinberger & erkens, 2007; sampson & clark, 2009). in the biology example on the memory of honey bees, communicating and scrutinising may play a double role. on the one hand, if groups of learners work on the honey bee problem, communication within the team is necessary to secure that the research process is carried out in a rigorous way, including arriving at a sound explanation for the phenomenon under investigation. on the other hand, the research process and outcomes are typically shared with the broader community, e.g. in the form of plenary presentations. 3.2 domain-specific aspects of sra need to be investigated more systematically while research on sra focused on commonalities across domains, investigations on the differences of sra between disciplines have been rare (e.g., herrenkohl & cornelius, 2013). in addition, the set of domains under consideration has so far been small and seemingly arbitrary. one crucial question is what role domain-specific conceptual knowledge plays for successful sra (e.g., chinnappan, ekanayake & brown, 2011; schunn & anderson, 1999). domain-specific conceptual knowledge is, for example, necessary to build a mental representation of the problem situation and to identify aspects of the situation that offer scientifically accessible questions. moreover, the process of scientific reasoning is different across domains, with respect to both nature and weight of the epistemic activities to be displayed. for example, engineers enact the epistemic activity of “problem identification” by starting their design process with a clear problem for which the initial stage, the solution stage, and the constraints are all well-defined. natural scientists and social scientists do not necessarily have such well-defined initial and solution stages – for them, thus, the epistemic activities “questioning” and “hypothesis generation” play a major role. regarding the epistemic activity of “evidence evaluation”, scientific disciplines vary considerably in what is regarded as acceptable evidence to support a scientific claim. while many natural sciences rely upon hypothetico-deductive methods, many social sciences accept inductive comparisons as methods of evidence evaluation. in (pure) mathematics the only acceptable evidence is a chain of deductive arguments within a theory. all other kinds of evidence are regarded as informal. thus, transferring criteria for evidence evaluation from one discipline to another appears problematic. moreover, it is unclear whether exposure to one domain-specific approach of scientific reasoning influences the nature of evidence evaluation skills in other domains (given that k-12 education, as well as teacher education, immerses students in various domains). f. fischer et al. 36 | f l r although the nature of epistemic activities varies across disciplines, approaches to foster student’s scientific reasoning have typically focused on single domains and developed in different directions. while research from developmental psychology and science education has predominantly focused on hypothesis and evidence generation and evaluation processes, research from mathematics education focused on metacognitive aspects to improve students’ self-regulated problem-solving (for example when searching for mathematical proofs, chinnappan & lawson, 1996). despite the fact that the existence of domain-dependent differences concerning sra can hardly be doubted, we contend that the three epistemic modes and the eight epistemic activities are of relevance to a broad range of disciplines. in other words, there may also be skill aspects of sra that are similar across domains (such as skills for structuring a problem situation, experimentation or deductive reasoning). however, since disciplines might differ substantially in the relative weights of the modes and activities and thus in the specific knowledge, skills and attitudes that students are supposed to develop when learning sra, a more representative selection of disciplines seems key for investigating their particularities in future research. finally, existing approaches to facilitation have typically proven effective for only one specific domain, in the context of one epistemic mode, in referring to only some specific epistemic activities, and in focusing on only some specific learning prerequisites. the extent to which the approaches to facilitation are domain-specific is an important question, but the extent to which they can be generalised across epistemic modes, domains, epistemic activities, and different learners is an important question as well (see klahr, zimmerman & jirout, 2011). future research should thus invest effort in identifying domain-specific and domain-general aspects of sra and their facilitation. 3.3 the role of emotions in sra requires investigation cognition is intricately interwoven with emotions. emotions are defined as systems of interrelated component processes, including subjective, physiological, and behavioral components (e.g., uneasy and nervous feelings, physiological activation, and anxious facial expression in anxiety; shuman & scherer, 2014). cognitive appraisals of situational demands and one’s competencies are known to shape human emotion. emotions, in turn, are prime drivers of motivation to solve problems and can profoundly impact the quality and outcomes of cognitive processes (e.g., moors, ellsworth, scherer & fijda, 2013; pekrun, 2006). it seems likely that this is also true for sra. without emotions such as surprise, curiosity triggered by contradictory findings, joy about solving scientific problems, or pride in one’s accomplishments, scientists would likely not be motivated to engage in scientific discovery, and students would lack motivation to learn science (pekrun, hall, goetz & perry, in press). furthermore, these emotions are known to regulate attention, memory processes, and different modes of cognitive problem solving, such as analytical versus holistic ways to approach problems, which are critically important for sra (fiedler & beier, 2014). systematic research examining the links between emotions and scientific reasoning, however, is largely lacking as yet (see sinatra, broughton & lombardi, 2014). we propose that five groups of emotions that seem to be relevant for scientific reasoning should be investigated. (1) epistemic emotions. as noted, epistemic activities such as generating hypotheses, are at the core of scientific reasoning in a broad range of domains. typically, these activities are accompanied by emotions triggered by the epistemic quality of problem-related information and mental activity. a prototypical case is cognitive incongruity triggering surprise, awe, curiosity, confusion, or joy when the incongruity is resolved. as proposed by philosophers (brun, doğuoğlu, & kuenzle, 2008; morton, 2010), these emotions can be called epistemic emotions (pekrun, & stevens, 2011). (2) achievement emotions. achievement emotions are emotions that relate to activities or outcomes that are judged according to competence-related standards of quality (pekrun, 2006). in many learning situations, scientific reasoning activities and the outcomes of these activities are judged for their achievement quality. depending on the perceived importance of success and failure, scientific reasoning can induce strong achievement emotions, such as hope and pride or anxiety, shame, and hopelessness. f. fischer et al. 37 | f l r (3) topic emotions. during scientific reasoning, emotions can be triggered by the contents of the problem to be solved. an example is the anxiety experienced when dealing with issues of climate change or genetically modified food. in contrast to epistemic emotions, topic emotions do not directly pertain to the process of scientific reasoning, however, they can strongly influence engagement in reasoning (ainley, 2006). (4) social emotions. scientific reasoning is often situated in social contexts. by implication, scientific reasoning can induce a multitude of social emotions related to other people. these emotions include both social achievement emotions, such as admiration, envy, contempt, or empathy related to the success and failure of others, as well as non-achievement emotions, such as love or hate in relationships with collaborators in the reasoning process (weiner, 2007). (5) incidental emotions and moods. when engaging in scientific reasoning, a person can continue experiencing emotions that relate to external events, such as current stress, or problems in their family. these emotions do not relate to the reasoning process itself, but have the potential, nonetheless, to strongly influence the quality of reasoning and learning to reason, such as a student’s worries about their parents’ divorce being brought into the science classroom. all five classes of emotion can play a role in all epistemic activities. however, it seems likely that different emotions are more typical for some of these activities than for others. for example, epistemic emotions are likely to be triggered by mental activities that can involve impasses and cognitive incongruity, such as “problem identification” or “evidence evaluation”, whereas social emotions are of primary importance in collaborative reasoning processes and for the communication of the results of scientific reasoning. furthermore, emotions of all five classes can profoundly influence the scientific reasoning process and its outcomes. the impact of these emotions on reasoning can be mediated by various cognitive and motivational processes, e.g. intrinsic and extrinsic motivation to engage, or deep versus shallow information processing strategies (e.g., clore & huntsinger, 2009). as a consequence, positive activating emotions in reasoning may typically support high-quality reasoning, whereas some negative emotions may be detrimental. however, for many emotions and task conditions, the effects on reasoning performance are likely to be more complex. thus we argue that studying the role of emotions in and during sra is an important task for future research. 3.4 the social context of sra should be considered more systematically scientific reasoning and argumentation are typically situated in a social context (dunbar, 1995). some epistemic activities are collaborative in nature, such as discussing the results of scientific reasoning with peers or communicating them to the broader public. other epistemic activities are not collaborative in nature but may benefit from collaboration (chi, 2009; duschl, 2008). we propose two strands for future research that bear the potential to improve our understanding of the social aspects of sra. (1) collaborative knowledge construction. extensive research has been carried out on the cognitive and social mechanisms of knowledge construction in groups and collectives. research on knowledge construction in pairs and small groups has often been conducted in a joint problem-solving paradigm. this line of research focuses on how pairs and groups, in contrast to individuals, work on complex science-related problems (e.g., okada & simon, 1997) and on how groups develop joint strategies and norms for sra beyond just learning the domain content associated with the task (e.g., roschelle & teasley, 1997). research on dialogic education (wegerif, 2007) and argumentative classroom discourse (osborne, 2010) focuses on the structure and content of discussions in groups and collectives, and on the conditions for evolving (scientific) quality of the argumentation in these discussions. in contrast to the perspectives on joint problem solving and dialogic argumentation that analyse the micro-mechanisms of knowledge construction, research on communities of practice emphasises processes of knowledge creation, participation and identity in collectives of people sharing goals or interests (lave & wenger, 1991). in knowledge community approaches, domain knowledge acquisition by individuals is rather seen as a by-product of qualitative f. fischer et al. 38 | f l r changes in the participation pattern, from legitimate peripheral to more “core” participation. research is needed on which forms of participation in epistemic activities of certain scientific communities effectively advance students’ sra skills. (2) distributed, shared and collective cognition. approaches to distributed and shared cognition share the assumption that reasoning in real world tasks cannot be understood by just focussing on isolated individuals. in real world tasks, individuals collaborate with others on solving problems and making decisions, but they also use tools that allow them to act much more intelligently than they would be able to without. the distributed cognition perspective suggests a systemic perspective for the analysis of complex social and socio-technical tasks (e.g. salomon & perkins, 1998). research on transactive memory systems (wegner, 1987) addresses the cognitive interdependence that develops when group members collaborate for some time and specialise in specific areas of which the other members are aware. a transactive memory system is thus characterised by the collaborative division of labour for learning, remembering, and communication of knowledge (e.g., hollingshead, gupta, yoon, & brandon, 2011), which seems crucial for most epistemic activities. the shared mental models perspective (e.g. mohammed & dumville, 2001; wu & keysar, 2007) addresses the question which kind of knowledge (e.g., knowledge on task vs. knowledge on team) is needed and the extent to which group members need overlapping (shared) information as opposed to unique (or unshared) information to perform well as a team. research on cognitive convergence (teasley et al., 2008) or knowledge convergence (fischer & mandl, 2005) focuses on the similarity and dissimilarity of cognitive representations in collaborative situations, as well as their changes through collaboration. in the context of sra it is an interesting open question to which extent divergent vs. convergent cognitive representations of different individuals in a group are supportive for different epistemic activities. it seems plausible to hypothesise that divergent knowledge in a group is specifically supportive in epistemic activities such as “evaluating evidence” and “scrutinising arguments”. furthermore, a recurring result from prior research is that the knowledge learners acquire through collaboration is surprisingly dissimilar (miyake, 1986). this might be especially relevant for educational settings where students engage in collaborative learning to develop sra skills. research on expert-layperson communication (bromme, jucks & runde, 2005) has shown that large differences in domain expertise may have detrimental effects on communication and understanding. measures to support expert-layperson communication have shown positive effects (e.g., nückles & stürz, 2006). in the context of sra these knowledge differences exist, e.g., between scholars acting as teachers and students in their early years but also in the context of communicating scientific outcomes to wider audiences. an open question is how different disciplines try to overcome detrimental effects and make optimal use of large knowledge differences between scholars and students. 3.5 the influence of digital technologies on sra needs further research studies show that digital technologies affect reasoning and learning contingent to the way that they are used. for instance, a study by sparrow, liu and wegner (2011) revealed that digital technologies increasingly become “external memories” integral to people’s reasoning. their findings show that the availability of externally stored information changes cognitive processing dramatically, depending on the person’s assumptions on later accessibility. it is plausible to assume that the availability of digital technology affects sra in a similar way. moreover, this should be true for all three epistemic modes, but in different ways. when advancing theory about natural and social phenomena, technology is typically used for data collection and visualisation (e.g., computer simulations; gijlers & de jong, 2009) as well as for analysis, including not only statistical analysis but also analysis based on language and logic (e.g., rosé, wang, arguello, stegmann, weinberger & fischer, 2008). when applying science-based reasoning in practice, technology is often used to provide access to the knowledge base and theories in the respective domain (e.g., sparrow, et al., 2011). in the epistemic mode of artefact-centred scientific reasoning, technology often acts as the core enabler for prototypes or simulating features of a design artefact (e.g., wiethoff, schneider, rohs, butz & greenberg, 2012). across the three modes, research has generated evidence that communication and collaboration can be substantially enhanced by digital technologies (see stegmann et al., 2012). furthermore, awareness tools, i.e. tools capable of capturing and mirroring the quality of group processes via external f. fischer et al. 39 | f l r representations, have shown strong potential to support scientific argumentation (janssen & bodemer, 2013; streng, stegmann, boring, böhm, fischer & hussmann, 2010). we propose that future research should investigate how technologies shape sra. firstly, research should more systematically address how available and easily accessible technologies influence scientific reasoning in the different epistemic modes and activities. a co-evolutionary perspective on the mutual influence of technology development and scientific reasoning seems promising, for example, how access to scientific information through the internet affects the sra of practitioners. secondly, research should investigate the effects of technological tools specifically designed to facilitate certain epistemic activities in sra. prior research on computer simulations and computer-supported collaboration will be informative for the formulation of design principles for the development of technology-based scaffolds. 4. conclusions sra is considered one of the core competences in knowledge societies (trilling & fadel, 2009). knowledge of the structure and generality of these competences, their emotional, social and technological conditions, and how they can be facilitated appears key for a promising re-design of curricula and interventions in schools, higher education and vocational practice to foster the development of sra. as a starting point for the necessary interdisciplinary research we suggest the following broad definition of sra: scientific reasoning and argumentation include the knowledge and skills involved in different epistemic activities (problem identification, questioning, hypothesis generation, construction of artefacts, evidence generation, evidence evaluation, drawing conclusions as well as communicating and scrutinising scientific reasoning and its results) in the context of three different epistemic modes (advancing theory building about natural and social phenomena, science-based reasoning in practice, and artefact-centred scientific reasoning,). scientific reasoning and argumentation are assumed to consist of domain-specific as well as domain-general components, and depend on emotional, social, instructional/facilitative, and technological conditions. we proposed a research agenda on the analysis and facilitation of sra in educational contexts, which significantly broadens our perspective beyond basic experimental research. based on stokes’ (1997) model of scientific knowledge production, we suggested three epistemic modes of sra: (1) advancing theory building about natural and social phenomena, (2) science-based reasoning in practice, and (2) artefact-centred scientific reasoning. in a broad range of domains, all three epistemic modes play a role. students thus need to learn to understand how scientific knowledge is developed in their domains of study, and how it can be applied to address practical problems. to an extent differing vastly between domains and study programmes, students are also expected to learn to participate in processes of scientific research (trilling & fadel, 2009). we further identified eight epistemic activities, of which some have only received marginal or narrowly focused consideration in research on sra, mainly in the experimental paradigm: (1) problem identification, (2) questioning, (3) hypothesis generation, (4) construction and redesign of artefacts, (5) evidence generation, (6) evidence evaluation, (7) drawing conclusions and (8) communicating and scrutinising. we do not claim that this process typology is exhaustive, and do not intend to conceal that others have developed alternative typologies (e.g., van joolingen, de jong, lazonder, savelsbergh & manlove, 2005; white & frederiksen, 1998). instead it is proposed as a starting point for an interdisciplinary research agenda, to be modified in further theoretical discussion and based on findings of empirical studies. based on this framework, we suggest five further areas in research on sra that require more systematic investigation. first, research should investigate the differences between disciplines regarding how epistemic modes and activities are employed and to what extent knowledge generated within them is considered as evidence for or against theories. we suggest that it is crucial to advance our understanding of sra by determining which aspects are domain-general and which aspects are specific for a single domain or group of domains (see schunn & anderson, 1999). f. fischer et al. 40 | f l r second, commonalities and differences between disciplines are also likely to exist with respect to measures of intervention and facilitation. on the one hand, some of the interventions developed for a specific domain and context might prove generalizable to some extent to other contexts and domains. on the other hand, domain-independent instructional approaches might well be differentially effective in different domains (see klahr et al., 2011). in addition, we suggest building a coherent conceptual framework for integrating the diverse research findings from intervention research across domains. chi’s (2009) icap model might be a promising starting point in this respect to integrate the available evidence and guide future research on sra interventions. icap classifies learning activities based on their underlying cognitive processes into interactive, constructive, active and passive. the model predicts the best learning outcomes for interactive learning activities, followed by constructive, active and passive activities. third, research on sra displays a strong cognitive bias. however, it seems likely that most scientific reasoning processes are triggered, modulated, or followed by emotions (see shuman, & scherer, 2014). thus far there is no systematic research on emotions in the context of sra, which is striking because, for example, curiosity is widely regarded as a major driving force for any scientific endeavour (pekrun & stevens, 2011). fourth, scientific reasoning is increasingly recognised as a social epistemic practice rather than a purely individual activity (dunbar, 1995). however, prior research on sra has examined the social context in which sra appears in a rather unsystematic way. therefore, we suggest considering constructs of research fields that are advanced in this perspective, such as peer assessment (cho et al., 2006; strijbos & sluijsmans, 2010) or research on collaboration scripts (fischer et al., 2013), as starting points for addressing the social aspects of scientific reasoning. fifth, recent years have seen an expansion of digital technology in nearly every sector of society, including research and related fields of practice. we argue that the effects of digital technologies on sra practices need to be examined more systematically. important questions include how digital technologies are used to support scientific reasoning and how technologies can be designed to support students in sra (see for example gijlers & de jong, 2009). given the amount of research in the fields of scientific reasoning and scientific argumentation described at the outset of this article, the field might benefit from an integrative view that combines the so far largely separated strands of research. with concerted and interdisciplinary research efforts, we strongly believe that we may achieve a better understanding of what sra skills are, how they develop and how their development can be supported effectively. the outcomes of this research may subsequently inform educational practice to help educate citizens who are able to participate in science-related societal debates and make more systematic use of scientific knowledge and skills. keypoints we review research in three areas – developmental research on scientific reasoning, research on argumentation, and research on interventions on scientific reasoning and argumentation. the article proposes a framework for scientific reasoning and argumentation meant as a starting point for interdisciplinary research. the framework includes three epistemic modes: advancing theory building about phenomena, artefact-centred scientific reasoning, and science-based reasoning in practice. we distinguish eight epistemic activities (e.g., generating evidence) relevant in all three epistemic modes. we argue that differences between disciplines as well as the roles of emotions, the social context, and digital technologies have been neglected and are promising foci for interdisciplinary research on scientific reasoning and argumentation. f. fischer et al. 41 | f l r acknowledgements this work was funded by the elite network of bavaria. references ainley, m. (2006). connecting with learning: motivation, affect and cognition in interest processes. educational psychology review, 18(4), 391-405. doi: 10.1007/s10648-006-9033-0 bromme, r., jucks, r., & runde, a. (2005). barriers and biases in computer-mediated expert-laypersoncommunication. in r. bromme, f.w. hesse, & h. spada (eds.), barriers and biases in computermediated knowledge communication (pp. 89-118). new york: springer. brun, g., doğuoğlu, u., & kuenzle, d. (eds.). (2008). epistemology and emotions. aldershot, uk: ashgate. bullock, m., sodian, b., & koerber, s. (2009). doing experiments and understanding science. development of scientific reasoning from childhood to adulthood. in w. schneider, & m. bullock (eds.). human development from early childhood to early adulthood: findings from a 20 year longitudinal study (pp. 173-198). new york, nj: psychology press. bullock, m., & ziegler, a. (1999). scientific reasoning: developmental and individual differences. in f. e. weinert, & w. schneider (eds.). individual development from 3 to 12: findings from the munich longitudinal study (pp. 38-54). cambridge: cambridge university press. carey, s., & smith, c. (1993). on understanding the nature of scientific knowledge. educational psychologist, 28(3), 235-251. doi: 10.1207/s15326985ep2803_4 cavagnetto, a. r. (2010). argument to foster scientific literacy. a review of argument interventions in k-12 science contexts. review of educational research, 80(3), 336-371. doi: 10.3102/0034654310376953 charlin, b., boshuizen, h. p., custers, e. j., & feltovich, p. j. (2007). scripts and clinical reasoning. medical education, 41(12), 1178-1184. doi: 10.1111/j.1365-2923.2007.02924.x chi, m. t. (2009). active, constructive, interactive: a conceptual framework for differentiating learning activities. topics in cognitive science, 1(1), 73-105. doi: 10.1111/j.1756-8765.2008.01005.x chinn, c., & clark, d. b. (2013). learning through collaborative argumentation. in c. e. hmelo-silver, c. a. chinn, c. k. k. chan, & a. m. o'donnell (eds.), international handbook of collaborative learning (pp. 314-332). new york: routledge. chinn, c. a., & malhotra, b. a. (2002). epistemologically authentic inquiry in schools: a theoretical framework for evaluating inquiry tasks. science education, 86(2), 175-218. doi: 10.1002/sce.10001 chinnappan, m., ekanayake, m. b., & brown, c. (2011). specific and general knowledge in geometric proof development. saarc journal of educational research, 8, 1-28. chinnappan, m., & lawson, m. j. (1996). the effects of training in the use of executive strategies in geometry problem solving. learning and instruction, 6(1), 1-17. doi: 10.1016/s09594752(96)80001-6 cho, k., schunn, c. d., & wilson, r. w. (2006). validity and reliability of scaffolded peer assessment of writing from instructor and student perspectives. journal of educational psychology, 98(4), 891901. doi: 10.1037/0022-0663.98.4.891 clark, d. b., sampson, v., weinberger, a., & erkens, g. (2007). analytic frameworks for assessing dialogic argumentation in online learning environments. educational psychology review, 19(3), 343-374. doi: 10.1007/s10648-007-9050-7 clore, g. l., & huntsinger, j. r. (2009). how the object of affect guides its impact. emotion review, 1, 3954. doi: 10.1177/1754073908097185 dochy, f., segers, m., van den bossche, o., & gijbels, d., (2003). effects of problem-based learning: a meta-analysis. learning and instruction, 13(5), 533-568. doi: 10.3102/00346543075001027 http://dx.doi.org/10.1016/s0959-4752%2896%2980001-6 http://dx.doi.org/10.1016/s0959-4752%2896%2980001-6 http://psycnet.apa.org/doi/10.1037/0022-0663.98.4.891 f. fischer et al. 42 | f l r dunbar, k. (1995). how scientists really reason: scientific reasoning in real-world laboratories. in r. j. sternberg, & j. davidson (eds.), mechanisms of insight (pp. 365-395). cambridge ma: mit press. duschl, r. (2008). science education in three-part harmony: balancing conceptual, epistemic, and social learning goals. review of research in education, 32(1), 268-291. doi: 10.3102/0091732x07309371 erduran, s., simon, s., & osborne, j. (2004). tapping into argumentation: developments in the application of toulmin's argument pattern for studying science discourse. science education, 88(6), 915-933. doi: 10.1002/sce.20012 euler, l. (1741). solutio problematis ad geometriam situs pertinentis. commentarii academiae scientiarum petropolitanae, 8, 128-140. fiedler, k., & beier, s. (2014). affect and cognitive processes. in r. pekrun, & l. linnenbrink-garcia (eds.), international handbook of emotions in education (pp. 36-55). new york: taylor & francis. fischer, f., kollar, i., stegmann, k., & wecker, c. (2013). toward a script theory of guidance in computersupported collaborative learning. educational psychologist, 48(1), 56-66. doi: 10.1080/00461520.2012.748005 fischer, f., & mandl, h. (2005). knowledge convergence in computer-supported collaborative learning: the role of external representation tools. journal of the learning sciences, 14(3), 405-441. doi: 10.1207/s15327809jls1403_3 furtak, e. m., seidel, t., iverson, h., & briggs, d. c. (2012). experimental and quasi-experimental studies of inquiry-based science teaching. a meta-analysis. review of educational research, 82(3), 300329. doi: 10.3102/0034654312457206 galbraith, p. l., henn, h.-w., & niss, n. (2007). modelling and applications in mathematics education. new york, nj: springer. gijlers, h., & de jong, t. (2009). sharing and confronting propositions in collaborative inquiry learning. cognition and instruction, 27(3), 239-268. doi: 10.1080/07370000903014352 herrenkohl, l. r. & cornelius, l. (2013). investigating elementary students' scientific and historical argumentation. journal of the learning sciences, 22(3), 413-461. doi: 10.1080/10508406.2013.799475 hierholzer, c., & wiener, c. (1873). ueber die möglichkeit, einen linienzug ohne wiederholung und ohne unterbrechung zu umfahren. mathematische annalen, 6(1), 30-32. doi: 10.1007/bf01442866 hollingshead, a. b., gupta, n., yoon, k., & brandon, d. p. (2011). transactive memory theory and teams: past, present, and future. in e. salas, s. m. fiore, & m. p. letzky (eds.), theories of team cognition: cross-disciplinary perspectives (pp. 421-455). new york: routledge. inhelder, b. & piaget, j. (1958). the growth of logical thinking from childhood to adolescence: an essay on the construction of formal operational structure. london: routledge & kegan pau. janssen, j., & bodemer, d. (2013). coordinated computer-supported collaborative learning: awareness and awareness tools. educational psychologist, 48(1), 40-55. doi: 10.1080/00461520.2012.749153 kelly, g. j., & takao, a. (2002). epistemic levels in argument: an analysis of university oceanography students' use of evidence in writing. science education, 86(3), 314-342. doi: 10.1002/sce.10024 klahr, d., & dunbar, k. (1988). dual space search during scientific reasoning. cognitive science, 12(1), 148. doi: 10.1207/s15516709cog1201_1 klahr, d., zimmerman, c., & jirout, j. (2011). educational interventions to advance children’s scientific thinking. science, 333(6045), 971-975. doi: 10.1126/science.1204528 koerber, s., & sodian, b. (2009). reasoning from graphs in young children. preschoolers’ ability to interpret and evaluate covariation data from graphs. journal of psychology of science & technology, 2(2), 73-86. doi: 10.1891/1939-7054.2.2.73 koerber, s., sodian, b., kropf, n., mayer, d., & schwippert, k. (2011). die entwicklung des wissenschaftlichen denkens im grundschulalter. theorieverständnis, experimentierstrategien, dateninterpretation. zeitschrift für entwicklungspsychologie und pädagogische psychologie, 43(1), 16-21. doi: 10.1026/0049-8637/a000027 koerber, s., sodian, b., thoermer, c., & nett, u. (2005). scientific reasoning in young children: preschoolers' ability to evaluate covariation evidence. swiss journal of psychology 64(3), 141-152. doi: 10.1024/1421-0185.64.3.141 http://psycnet.apa.org/doi/10.1024/1421-0185.64.3.141 f. fischer et al. 43 | f l r kollar, i., fischer, f., & slotta, j. d (2007). internal and external scripts in computer-supported collaborative inquiry learning. learning & instruction, 17(6), 708-721. doi: 10.1016/j.learninstruc.2007.09.021 kolodner, j. l. (2007). the roles of scripts in promoting collaborative discourse in learning by design. in f. fischer, i. kollar, h. mandl, & j. m. haake (eds.), scripting computer-supported collaborative learning cognitive, computational and educational approaches (pp. 237-262). new york: springer. koslowski, b. (1996). theory and evidence: the development of scientific reasoning. cambridge, ma: mit press/bradford books. koslowski, b. (2012). scientific reasoning: explanation, confirmation bias, and scientific practice. in g. j. feist, & m. e. gorman (eds.), handbook of the psychology of science (pp. 151-192). new york, nj: springer. kruglanski, a. w., & gigerenzer, g. (2011). intuitive and deliberate judgments are based on common principles. psychological review, 118(1), 97-109. doi: 10.1037/a0020762 kuhn, d. (1991). the skills of argument. new york: cambridge university press. kuhn, d., & franklin, s. (2006). the second decade: what develops (and how)? in d. kuhn, & r. siegler (eds.), handbook of child psychology: vol. 2. cognition, perception, and language (pp. 517-550). hoboken, nj: wiley. lave, j., & wenger, e. (1991). situated learning: legitimate peripheral participation. cambridge, uk: cambridge university press. mäkitalo-siegl, k., kohnle, c., & fischer, f. (2011). computer-supported collaborative inquiry learning and classroom scripts: effects on help-seeking processes and learning outcomes. learning and instruction, 21(2), 257-266. doi: 10.1016/j.learninstruc.2010.07.001 mayer, d., sodian, b., koerber, s., & schwippert, k. (2014). scientific reasoning in elementary school children: assessment and relations with cognitive abilities. learning and instruction, 29, 43-55. doi: 10.1016/j.learninstruc.2013.07.005 mcneill, k. l. (2011). elementary students' views of explanation, argumentation, and evidence, and their abilities to construct arguments over the school year. journal of research in science teaching, 48(7), 793-823. doi: 10.1002/tea.20430 mcneill, k. l., lizotte, d. j., krajcik, j., & marx, r. w. (2006). supporting students' construction of scientific explanations by fading scaffolds in instructional materials. journal of the learning sciences, 15(2), 153-191. doi: 10.1207/s15327809jls1502_1 miyake, n. (1986). constructive interaction and the iterative process of understanding. cognitive science, 10, 151-177. doi: 10.1207/s15516709cog1002_2 mohammed, s., & dumville, b. c. (2001). team mental models in a team knowledge framework: expanding theory and measurement across disciplinary boundaries. journal of organizational behavior, 22(2), 89-106. doi: 10.1002/job.86 moors, a., ellsworth, p. c., scherer, k. r., & frijda, n. h. (2013). appraisal theories of emotion: state of the art and future development. emotion review, 5(2), 119-124. doi: 10.1177/1754073912468165 morton, a. (2010). epistemic emotions. in p. goldie (ed.), the oxford handbook of philosophy of emotion (pp. 385–399). oxford, united kingdom: oxford university press. nicol, d., thomson, a., & breslin, c. (2014). rethinking feedback practices in higher education: a peer review perspective. assessment & evaluation in higher education, 39(1), 102-122. doi: 10.1080/02602938.2013.795518 noroozi, o., weinberger, a., biemans, h. j., mulder, m., & chizari, m. (2012). argumentation-based computer supported collaborative learning (abcscl): a synthesis of 15 years of research. educational research review, 7(2), 79-106. doi: 10.1016/j.edurev.2011.11.006 nückles, m., & stürz, a. (2006), the assessment tool: a method to support asynchronous communication between computer experts and laypersons. computers in human behavior, 22(5), 917-940. doi: 10.1016/j.chb.2004.03.021 nussbaum, m. (2011). argumentation, dialogue theory, and probability modeling: alternative frameworks for argumentation research in education. educational psychologist, 46(2), 84-106. doi: 10.1080/00461520.2011.558816 http://dx.doi.org/10.1016/j.learninstruc.2007.09.021 http://psycnet.apa.org/doi/10.1037/a0020762 http://dx.doi.org/10.1016/j.learninstruc.2010.07.001 http://dx.doi.org/10.1016/j.learninstruc.2013.07.005 http://dx.doi.org/10.1016/j.edurev.2011.11.006 http://dx.doi.org/10.1016/j.chb.2004.03.021 f. fischer et al. 44 | f l r okada, t., & simon, h. a. (1997). collaborative discovery in a scientific domain. cognitive science, 21(2), 109-146. doi: 10.1207/s15516709cog2102_1 osborne, j. (2010). arguing to learn in science: the role of collaborative, critical discourse. science, 328(5977), 463-466. doi: 10.1126/science.1183944 palincsar, a. s., & brown, a. l. (1984). reciprocal teaching of comprehension-fostering and comprehension-monitoring activities. cognition and instruction, 1(2), 117-175. doi: 10.1207/s1532690xci0102_1 pekrun, r. (2006). the control-value theory of achievement emotions: assumptions, corollaries, and implications for educational research and practice. educational psychology review, 18, 315-341. doi: 10.1007/s10648-006-9029-9 pekrun, r., hall, n. c., goetz, t., & perry, r. p. (in press). boredom and academic achievement: testing a model of reciprocal causation. journal of educational psychology. pekrun, r., & stephens, e. j. (2011). academic emotions. in k. r. harris, s. graham, t. urdan, s. graham, j. m. royer, & m. zeidner (eds.), apa educational psychology handbook (vol. 2, pp. 3-31). washington, dc: american psychological association. quintana, c., reiser, b. j., davis, e. a., krajcik, j., fretz, e., duncan, r. g., & soloway, e. (2004). a scaffolding design framework for software to support science inquiry. journal of the learning sciences, 13(3), 337-386. doi: 10.1207/s15327809jls1303_4 roschelle, j., & teasley, s. d. (1997). the construction of shared knowledge in collaborative problem solving. in c. o'malley (ed.), computer supported collaborative learning (vol. 128, pp. 69-97). berlin: springer. rosé, c. p., wang, y. c., arguello, j., stegmann, k., weinberger, a., & fischer, f. (2008). analyzing collaborative learning processes automatically: exploiting the advances of computational linguistics in computer-supported collaborative learning. international journal of computersupported collaborative learning, 3, 237-271. doi: 10.1007/s11412-007-9034-0 sadler, t. d. (2004). informal reasoning regarding socio-scientific issues: a critical review of research. journal of research in science teaching, 41(5), 513-536. doi: 10.1002/tea.20009 salomon, g., & perkins, d. n. (1998). individual and social aspects of learning. review of research in education, 23, 1-24. doi:10.3102/0091732x023001001 sampson, v., & clark, d. (2009). the impact of collaboration on the outcomes of scientific argumentation. science education, 93(3), 448-484. doi: 10.1002/sce.20306 scardamalia, m., & bereiter, c. (2006). knowledge building: theory, pedagogy, and technology. in k. sawyer (ed.), cambridge handbook of the learning sciences (pp. 97-119). new york: university press. schunn, c. d., & anderson, j. r. (1999). the generality/specificity of expertise in scientific reasoning. cognitive science, 23(3), 337-370. doi: 10.1016/s0364-0213(99)00006-3 shafto, p., kemp, c., bonawitz, e. b., coley, j. d., & tenenbaum, j. b. (2008). inductive reasoning about causally transmitted properties. cognition, 109(2), 175-192. doi: 10.1016/j.cognition.2008.07.006 shavelson, r. j., & towne, l. (eds.). (2002). scientific research in education. washington, dc: national academic press. shuman, v., & scherer, k. r. (2014). concepts and structures of emotions. in r. pekrun, & l. linnenbrinkgarcia (eds.), international handbook of emotions in education (pp. 13-35). new york: taylor & francis. sinatra, g. m., broughton, s. h., & lombardi, d. (2014). emotions in science education. in r. pekrun, & l. linnenbrink-garcia (eds.), international handbook of emotions in education (pp. 415-436). new york: taylor & francis. sodian, b., & bullock, m. (2008). scientific reasoning – where are we now? cognitive development, 23(4), 431-434. doi: 10.1016/j.cogdev.2008.09.003 sodian, b., jonen, a., thoermer, c. & kircher, e. (2006). die natur der naturwissenschaften verstehen: implementierung wissenschaftstheoretischen unterrichts in der grundschule. in m. prenzel, & l. allolio-näcke (eds.), untersuchungen zur bildungsqualität von schule. abschlussbericht des dfg-schwerpunktprogramms (s. 147-160). münster: waxmann. http://dx.doi.org/10.1016/s0364-0213%2899%2900006-3 http://dx.doi.org/10.1016/j.cognition.2008.07.006 http://dx.doi.org/10.1016/j.cogdev.2008.09.003 f. fischer et al. 45 | f l r sodian, b., zaitchik, d., & carey, s. (1991). young children's differentiation of hypothetical beliefs from evidence. child development, 62, 753-766. doi: 10.1111/j.1467-8624.1991.tb01567.x sparrow, b., liu, j., & wegner, d. (2011). google effects on memory: cognitive consequences of having information at our fingertips. science, 333(6043), 776-778. doi: 10.1126/science.1207745 stegmann, k., wecker, c., weinberger, a., & fischer, f. (2012). collaborative argumentation and cognitive elaboration in a computer-supported collaborative learning environment. instructional science, 40(2), 297-323. doi: 10.1007/s11251-011-9174-5 stokes, d. e. (1997). pasteur’s quadrant: basic science and technological innovation. washington, dc: brookings institution press. streng, s., stegmann, k., boring, s., böhm, s., fischer, f., & hussmann, h. (2010). measuring effects of private and shared displays in small-group knowledge sharing processes. in e. hvannberg, m. k. lárusdóttir, a. blandford, & j. gulliksen (eds.), proceedings of the 6th nordic conference on human-computer interaction (nordichi 2010) (pp. 789-792). new york, ny: acm. strijbos, j. w., & sluijsmans, d. (2010). unravelling peer assessment: methodological, functional, and conceptual developments. learning and instruction, 20(4), 265-269. doi: 10.1016/j.learninstruc.2009.08.002 teasley, s., fischer, f., dillenbourg, p., kapur, m., chi, m., weinberger, a., & stegmann, k. (2008). cognitive convergence in collaborative learning. in proceedings of icls 2008 (vol. 3, pp. 360– 367). international society of the learning sciences. trilling, b., & fadel, c. (2009). twenty-first century skills. learning for life in out times. san francisco: jossey-bass. van joolingen, w. r., & de jong, t. (1993). exploring a domain through a computer simulation: traversing variable and relation space with the help of a hypothesis scratchpad. in d. towne, t. de jong, & h. spada (eds.), simulation-based experiential learning (pp. 191-206). berlin: springer. van joolingen, w. r., de jong, t., lazonder, a. w., savelsbergh, e. r., & manlove, s. (2005). co-lab: research and development of an online learning environment for collaborative scientific discovery learning. computers in human behavior, 21, 671-688. doi: 10.1016/j.chb.2004.10.039 wegerif, r. (2007). dialogic education and technology: expanding the space of learning. new york: springer. wegner, d. m. (1987). transactive memory: a contemporary analysis of the group mind. in b. mullen & g. r. goethals (eds.), theories of group behavior (pp. 185-208). new york: springer. weinberger, a., stegmann, k., & fischer, f. (2010). learning to argue online: scripted groups surpass individuals (unscripted groups do not). computers in human behavior, 26(4), 506-515. doi: 10.1016/j.chb.2009.08.007 weiner, b. (2007). examining emotional diversity in the classroom: an attribution theorist considers the moral emotions. in p. a. schutz, & r. pekrun (eds.), emotion in education (pp. 75-88). san diego, ca: academic press. white, b. y., & frederiksen, j. r. (1998). inquiry, modelling, and metacognition: making science accessible to all students. cognition and instruction, 16(1), 3-118. doi: 10.1207/s1532690xci1601_2 wiethoff, a., schneider, h., rohs, m., butz, a., & greenberg, s. (2012). sketch-a-tui: low cost prototyping of tangible interactions using cardboard and conductive ink. in s. n. spencer (ed.), proceedings of the sixth international conference on tangible, embedded and embodied interaction (pp. 309-312). new york: acm. wu, s., & keysar, b. (2007). the effect of information overlap on communication effectiveness. cognitive science, 31(1), 169-181. doi: 10.1080/03640210709336989 zazkis, r., & chernoff, e. j. (2008). what makes a counterexample exemplary? educational studies in mathematics, 68, 195-208. doi: 10.1007/s10649-007-9110-4 zimmerman, c. (2000). the development of scientific reasoning skills. developmental review, 20, 99-149. doi: 10.1006/drev.1999.0497 zimmerman, c. (2007). the development of scientific thinking skills in elementary and middle school. developmental review, 27, 172-223. doi: 10.1016/j.dr.2006.12.001 http://dx.doi.org/10.1016/j.learninstruc.2009.08.002 http://dx.doi.org/10.1016/j.chb.2004.10.039 http://dx.doi.org/10.1016/j.chb.2009.08.007 http://dx.doi.org/10.1006/drev.1999.0497 frontline learning research 3 (2014) 83-101 issn 2295-3159 corresponding author: jessie de naeghel, henri dunantlaan 2, 9000 ghent, belgium, jessie.denaeghel@ugent.be http://dx.doi.org/10.14786/flr.v2i1.84 83 | f l r strategies for promoting autonomous reading motivation: a multiple case study research in primary education jessie de naeghel a , hilde van keer a , ruben vanderlinde a a department of educational studies, ghent university, belgium article received 5 february 2014 / revised 16 february 2014 / accepted 26 april 2014 / available online 11 june 2014 abstract it is important to reveal strategies which foster students’ reading motivation in order to break through the declining trend in reading motivation throughout children’s educational careers. consequently, the present study advances an underexposed field in reading motivation research by studying and identifying the strategies of teachers excellent in promoting fifth-grade students’ volitional or autonomous reading motivation through multiple case study analysis. data on these excellent teachers were gathered from multiple sources (interviews with teachers, sen coordinators, and school leaders; classroom observations; teacher and student questionnaires) and analysed. the results point to the teaching dimensions of autonomy support, structure, and involvement – as indicated by self-determination theory – as well as to reading aloud as critical strategies to promote students’ autonomous reading motivation in the classroom. a school culture supporting students’ and teachers’ interest in reading is also an essential part of reading promotion. the theoretical and practical significance of the study is discussed. keywords: reading motivation; reading promotion; primary education; case studies 1. introduction competence in reading is essential for functioning adequately in today’s society. in this respect, it is crucial to encourage students’ high-quality forms of reading motivation and, therefore, to stimulate them to read more frequently (de naeghel, van keer, vansteenkiste, & rosseel, 2012; wigfield & guthrie, 1997) and master important reading skills (de naeghel et al., 2012; becker, mcelvany, & kortenbruck, 2010; wang & guthrie, 2004). unfortunately, research indicates that intrinsic reading motivation declines as children go through school (guthrie & wigfield, 2000). hence, it is important to uncover strategies which de naeghel et al. 84 | f l r foster students’ “love of reading” in order to break through the declining trend in reading motivation throughout children’s educational careers. reading motivation research indicates that teachers can play a crucial role in sustainably stimulating their students to read for pleasure and information (gambrell, 1996; guthrie & cox, 2001; guthrie, mcrae, & klauda, 2007; guthrie et al., 2006; santa et al., 2000). moreover, encouraging students’ willingness to read can be considered as a critical part of a high-quality education (de naeghel et al., 2012; guthrie & cox, 2001; guthrie et al., 2007), which can equip children from different socioeconomic backgrounds with the necessary reading competencies to be successful in today’s society (oecd, 2004). furthermore, teachers’ activities to promote their students’ volitional or autonomous reading motivation are of importance for achieving equal opportunities for all children, as teachers reach the majority of children independent of their socioeconomic background. in this respect, studying teachers excellent in promoting autonomous reading motivation can reveal critical strategies to promote reading motivation in education. mohan, lundeberg, and reffitt (2008) even explicitly encourage further research on excellent reading teachers. as teachers’ self-reports on their reading instruction do not always correspond with their actual behaviour (pressley, rankin, & yokoi, 1996) and, hence, observations of classroom teaching are explicitly encouraged (mohan et al., 2008), it is essential to study what exactly occurs in classrooms from different methodological perspectives to enhance data triangulation. therefore, a multiple case study research approach has been applied in the current study with an embedded mixed-method design (i.e., mix of quantitative and qualitative research approaches in which the emphasis is placed on the qualitative data; creswell & plano, 2007) to portray the strategies applied by teachers excellent in the promotion of highquality forms of reading motivation. in this respect, the study advances an underexposed field in reading motivation research through the study of what exactly occurs in the classroom practice of teachers excellent in promoting autonomous reading motivation, aiming to identify critical strategies to stimulate students’ willingness to read. moreover, it contributes to classroom practice by formulating practical guidelines for teachers and schools. 1.1 autonomous and controlled reading motivation several studies underline the multidimensional nature of reading motivation (e.g., baker & wigfield, 1999; de naeghel et al., 2012; watkins & coffey, 2004), indicating that children can be motivated for a variety of reasons. in line with the self-determination theory (sdt; ryan & deci, 2000), which is a contemporary and promising motivation theory with a rich and continuously emerging empirical basis, de naeghel et al. (2012) differentiate between qualitatively different types of reading motivation. particularly, autonomous and controlled types of reading motivation are distinguished. autonomous reading motivation, on the one hand, refers to engaging in reading activities for their own enjoyment (e.g., pleasure, interest) or because of their perceived personal significance and meaning (e.g., personal value, importance). on the other hand, controlled reading motivation is defined as reading to meet internal feelings of pressure (e.g., guilt, fear, pride) or to comply with external demands (e.g., expectations, reward, punishment). the present study will especially focus on autonomous reasons for reading, as autonomous reading motivation is associated with more positive outcomes, including higher leisure-time reading frequency, more reading engagement, and better reading comprehension. conversely, controlled reading motivation is related to less frequent reading in leisure time and lower reading comprehension scores (becker et al., 2010; de naeghel et al., 2012). 1.2 promoting reading motivation in the classroom the sdt formulates general guidelines to facilitate autonomous motivation (ryan & deci, 2000). particularly, conditions or teaching dimensions supporting students’ basic psychological needs for autonomy (i.e., the experience of a sense of volition or psychological freedom), competence (i.e., the experience of being confident and effective in action), and relatedness (i.e., the experience of feeling connected to and accepted by others) are argued to encourage students’ autonomous motivation to engage in activities de naeghel et al. 85 | f l r (skinner & belmont, 1993; ryan & deci, 2000; see figure 1). in this respect, it should be noted that the need for autonomy refers to the experience of being the initiator of one’s own behaviour or being selfdetermined and hence differs from acting independently without making an appeal to others (deci & ryan, 1987). the teaching dimensions distinguished in sdt are frequently studied in education in general (e.g., skinner & belmont, 1993; sierens, vansteenkiste, goossens, soenens, & dochy, 2009) as well as in physical education in particular (e.g., chatzisarantis & hagger, 2009; tessier, sarrazin, & ntoumanis, 2008), but less explicitly in primary education and in research on reading motivation. moreover, previous sdtbased research especially adopted a quantitative approach (e.g., chatzisarantis & hagger, 2009; sierens et al., 2003; skinner & belmont, 1993). hence, the focus on qualitative methods in the present study adds value to the sdt literature. the first teaching dimension, autonomy support, refers to giving students age-appropriate choices, recognising and connecting with children’s interests, offering rationales, taking the students’ perspective, and providing students with opportunities to take the initiative during learning activities (reeve, 2002; sierens, 2010; skinner & belmont, 1993). several studies confirm that autonomy-supportive teacher behaviour facilitates autonomous motivation (e.g., soenens & vansteenkiste, 2005) and positive learning outcomes, such as deep-level learning (e.g., vansteenkiste et al., 2005) and performance (e.g., black & deci, 2000). figure 1. teaching dimensions supporting students’ basic psychological needs and hence encouraging autonomous motivation (sdt; based on reeve, 2009). the second teaching dimension, structure, primarily fosters children’s need for competence. structure concerns clearly communicating expectations, responding consistently, providing optimal challenges, offering help and support, and providing positive feedback (reeve, 2002; sierens, 2010; skinner & belmont, 1993). research indicates that structuring by providing optimal challenges and providing positive feedback is positively associated with volitional or autonomous motivation (mouratadis, vansteenkiste, lens, & sideridis, 2008; vallerand & reid, 1984). third, the teaching dimension associated with children’s need for relatedness is involvement or “the quality of the interpersonal relationship with teachers and peers” (skinner & belmont, 1993, p. 573). teachers are involved with their students when they invest personal resources, express affection, and enjoy time with their students (reeve, jang, carrell, jeon, & barch, 2004). involvement is positively related to students’ behavioural and emotional engagement in the classroom (skinner & belmont, 1993). literature explicitly focusing on the encouragement of reading motivation (e.g., edmunds & bauserman, 2006; gambrell, 2011; gaskins, 2008) formulates strategies relating to the significance of providing choices and recognising interests (i.e., autonomy support), scaffolding and positive feedback (i.e., structure), and helping one another and interaction about books (i.e., involvement) as well. consequently, the value of the general teaching dimensions of autonomy support, structure, and involvement is acknowledged involvement relatedness = the experience of feeling connected to and accepted by others psychological need satisfaction autonomous motivation autonomy support structure autonomy = the experience of being self-determined competence = the experience of being confident and effective in action de naeghel et al. 86 | f l r in reading motivation studies and is therefore useful as a frame of reference to explore how teachers specifically encourage autonomous reading motivation in their classrooms. although research on instructional programs focusing on promoting reading motivation in late primary classrooms is relatively rare (guthrie et al., 2007), one instructional program did receive a lot of attention in the research literature, namely concept-oriented reading instruction (cori; e.g., guthrie & cox, 2001; guthrie et al., 2007; guthrie, wigfield, & vonsecker, 2000; wigfield et al., 2008). cori combines reading strategy instruction, conceptual knowledge in science, and support for students’ reading motivation. the theoretical justification for practices which influence children’s motivation in cori (e.g., providing students with age-appropriate choices linked to personal interests, providing collaborative support to stimulate interpersonal interaction) comes in part from the abovementioned sdt teaching dimensions (guthrie, 2004; guthrie et al., 2000). however, it should be noted that the adoption of sdt in reading motivation research to study the enhancement of students’ autonomous reading motivation remains rather limited and fragmented. above and beyond the significance of the sdt teaching dimensions of autonomy support, structure, and involvement the literature stresses the importance of teachers acting as reading models, valuing reading and sharing the “love of reading” to enhance their students’ reading motivation (gambrell, 1996; pecjak & kosir, 2008). teachers’ reading aloud is in this respect considered an effective strategy to stimulate students’ reading for enjoyment (fisher, flood, lapp, & frey, 2011; gambrell, palmer, codling, & mazzoni, 1996; pecjak & kosir, 2008). middle school students, for example, explicitly corroborate the value of their teachers’ reading out loud (ivey & broaddus, 2001). the literature, however, reveals contrasting results with respect to the effectiveness of reading aloud in early childhood education (e.g., morrow & gambrell, 2002; meyer, wardrop, linn, & hastings, 1993). in this respect, lane and wright (2011) emphasise that especially a systematic approach to reading aloud (e.g., dialogic reading; whitehurst et al., 1999) yields important academic benefits for children (e.g., increasing vocabulary, listening comprehension, word-recognition skills). since teachers are part of a broader school environment or community, it can be argued that the school culture can support and foster teachers’ and students’ willingness to invest in reading. in this respect, taylor, pearson, clark, and walpole (2010) indicate that effective schools indeed prioritise reading at both the class and school level. nevertheless, the role of the school and the specific school culture is still underexposed in reading motivation research. daniels and steres (2011) argue that schools’ prioritising of reading as a school-wide goal and hence fostering a climate in which teachers and students are expected and stimulated to read will positively influence students’ engagement. particularly, they encourage the allocation of a specific time for students to read self-selected books during the school day, support for teachers and administrators to read and discuss their reading with students, teachers’ professional development on literature, and investment in classroom libraries. moreover, literature underlines the role which literacy coaches can play in professionally supporting teachers to reflectively consider and improve the quality of classroom reading instruction and student learning. often, literacy coaches coordinate and support the literacy program of a school as well (steckel, 2009; vanderburg & stephens, 2010; walpole & blamey, 2008). 1.3 aim of the present study the present study is innovative in a number of ways. this study extends previous sdt research by applying sdt in research on primary school students and reading motivation. moreover, whereas numerous sdt-based studies relied solely on quantitative research, the present study adopts an embedded mixedmethod approach. this study also builds on the literature on reading motivation by studying reading aloud (fisher et al., 2011; gambrell et al., 1996; pecjak & kosir, 2008) and by exploring the critical role of the school’s reading culture for teachers’ classroom practices (daniels & steres, 2011; taylor et al., 2010). de naeghel et al. 87 | f l r the present study aims at contributing to theory on strategies to promote autonomous reading motivation and at offering guidelines for teachers’ classroom practice. in this respect, this study explores whether sdt’s teaching dimensions (i.e., autonomy support, structure, and involvement; reeve, 2002; skinner & belmont, 1993), reading aloud, and the reading culture at school can be identified as valuable strategies and stimulating contexts for the promotion of autonomous reading motivation in late primary classrooms. to pursue this goal, teachers excellent in promoting autonomous reading motivation were selected for a multiple case study research, as reading research explicitly expresses a need for further research on excellent reading teachers (mohan et al., 2008). 2. methodology 2.1 design a multiple case study research design (yin, 1989) was chosen, since on the one hand it affords an excellent way to identify and describe how teachers promote autonomous reading motivation and on the other hand it contributes to the establishment of theory on the promotion of autonomous reading motivation. also, the present study is regarded as an embedded mixed-method design (cresswell & plano, 2007). 2.2 teacher selection the present study is part of a broader research project on reading motivation and the promotion of reading motivation in flemish (belgium) late primary education. this study questioned 1270 fifth-grade students and their 67 teachers. on the basis of this large-scale enquiry, three teachers were selected for the present case study research, mrs. k, mrs. s, and mr. t (see table 1), according to two criteria. first, in an open-ended teacher questionnaire the three selected teachers self-reported applying several reading promotion strategies in their classroom (e.g., book promotion, reading aloud, small-group reading activities) and engaging in reading projects at the school level (e.g., school library, book club). second, their students reported high levels of recreational autonomous reading motivation on the self regulation questionnaire (srq)-reading motivation (see data collection section for a description of the instrument and table 1 for more detailed background information on the selected teachers; mrs. k’s class: m = 4.10, sd = 0.71, mrs. s’s class: m = 3.98, sd = 1.01, and mr. t’s class: m = 4.14, sd = 0.55; sample mean of all classes [n = 67] = 3.63, sd = 0.99; de naeghel et al., 2012). these two criteria reflect the selected teachers' excellence in terms of encouraging autonomous reading motivation. the three selected teachers agreed to participate in the present study. 2.3 data collection for the three selected teachers, qualitative and quantitative data regarding the class and school context were collected from multiple sources to enhance data triangulation. first, semi-structured teacher interviews were conducted which questioned their own reading motivation, their perception of their students’ reading motivation, and the practice of activities at class and school level to promote reading motivation. additional semi-structured interviews were conducted with special educational needs (sen) coordinators and school leaders to explore the role of the school in promoting students’ willingness to read. sen coordinators are members of the school team with both a supportive function towards students and teachers and a coordinating function aimed at optimising the school’s sen policy. second, field notes were taken by the researcher during at least two classroom observations of different reading activities in each class. third, two questionnaires were administered to teachers and their students to assess their reading motivation (srqreading motivation, de naeghel et al., 2012) and execution/perception of teaching dimensions (i.e., autonomy support, structure, and involvement; teacher as a social context (tasc) questionnaire, belmont, skinner, wellborn, & connell, 1988). fourth, school documents (e.g., the school website and inspectorate reports) were analysed. de naeghel et al. 88 | f l r 2.3.1 measurement scales students’ autonomous reading motivation was measured with the srq-reading motivation (de naeghel et al., 2012). each of the eight items of the autonomous reading motivation subscale was administered twice, with regard to motivation for recreational reading on the one hand (e.g., “i read in my free time, because it is important for me to read”) and motivation for academic reading on the other hand (e.g., “i read for school, because it is important for me to read”). in this respect, recreational reading referred to reading in students' leisure time and academic reading was defined as reading at school and for homework. items were scored on a five-point likert scale, ranging from one (disagree a lot) to five (agree a lot). the eight-item subscales had a good internal consistency with cronbach’s α = .90 and cronbach’s α = .92 respectively. the three teachers completed a slightly adapted version of the srq-reading motivation which measured autonomous reading motivation in general (i.e., without distinguishing between the recreational and academic context) and leaving out some less age-related items (e.g., “i have to prove myself that i can get good reading grades”). students’ perception of the teaching dimensions of autonomy support (e.g., “my teacher gives me a lot of choices about how i do my schoolwork”), structure (e.g., “my teacher doesn’t make clear what he/she expects of me in class”), and involvement (e.g., “my teacher likes me”) were assessed with the short version of the tasc questionnaire (belmont et al., 1988; sierens, vansteenkiste, goossens, soenens, & dochy, 2009). the eight-item subscales structure and involvement had an acceptable internal consistency, with cronbach’s α = .67 and cronbach’s α = .75 respectively. regarding autonomy support, four items were deleted, since they raised questions during administration and were found to be too difficult for fifth-graders. this resulted in a four-item subscale with an acceptable internal consistency, cronbach’s α = .62. items were scored on a five-point likert scale, ranging from one (disagree a lot) to five (agree a lot). teachers completed an adapted version of the tasc teacher questionnaire (belmont et al., 1988), which measured their execution of autonomy support, structure, and involvement in interaction with their students. 2.4 data analysis data analysis consisted of two phases, a vertical and a horizontal analysis. in the vertical analysis qualitative and quantitative data on each teacher were collected and a within-case analysis was performed (miles & huberman, 1994). the interview transcripts, school documents, and field notes were labelled with descriptive codes (summarising the content of text fragments) and subsequent interpretative codes (reflecting concepts from the theoretical framework). we designed the coding scheme starting with the three teaching dimensions as described in sdt (i.e., autonomy support, structure, and involvement; ryan & deci, 2000) and further developed it in the light of the interpretative data. text fragments with the same codes were clustered and interpreted with the use of the conceptual framework of this study. moreover, teacher and student questionnaires (srq-reading motivation, de naeghel et al., 2012; tasc, belmont et al., 1988) were analysed with spss 18. the analysis of the qualitative and quantitative data resulted in a case-specific report for each teacher which presented the data in the same format. in the second phase, the case-specific reports were subject to cross-site or horizontal analysis (miles & huberman, 1994) in which the cases were systematically compared for similarities and differences. to safeguard the quality of the data analysis, the intermediary results, interpretations, and conclusions were critically discussed by the researchers. 3. results 3.1 vertical analysis data presented in the three case-specific reports are structured around the same themes: (1) context and teacher profile, (2) classroom design (i.e., the availability of reading material, reading promotion material, etc.) aimed at reading promotion in the class, (3) classroom strategies (i.e., teaching dimensions: autonomy support, structure, and involvement; and reading aloud), and (4) school-level strategies on reading de naeghel et al. 89 | f l r motivation. the selection of these themes was based both on theory and empirical evidence (de naeghel & van keer, 2013; daniel & steres, 2011; fisher et al., 2011; gambrell et al., 1996; marinak & gambrell, 2007; mullis, martin, kennedy, & foy, 2007; reeve, 2002; sierens, 2010; skinner & belmont, 1993; steckel, 2009). in the case-specific reports the source of results is mentioned in parentheses. table 1 presents background information on the three selected teachers, their classes, and schools. table 1 background information on the three selected teachers, their classes, and schools mrs. k mrs. s mr. t teacher gender female female male age 49 35 34 teaching experience 28 years 10 years 14 years class number of students 16 19 17 mean student age 10.75 (0.31) 10.91 (0.40) 10.82 (0.45) school educational network subsidised private (roman catholic) subsidised private (roman catholic) community district type city city rural note: standard deviation in parentheses. 3.1.1 promotion of autonomous reading motivation in mrs. k’ s classroom context and teacher profile. mrs. k is a 49-year-old teacher with 28 years of teaching experience. she teaches fifth grade in a small school located just outside the city. there are 16 students in her class, who are on average 11 years old. mrs. k spends about 100 minutes a week on reading instruction. she uses “taalsignaal” as a teaching manual for the dutch language lessons. her preferred teaching methods are whole-class instruction, small-group instruction, and independent work. mrs. k hesitates to call herself a motivated reader, since she does not spend a lot of time reading novels. on the other hand, she is interested in journals, newspapers, informative books, etc. for gathering information [teacher interview] and reports that she is an autonomously motivated reader [table 2, element a]. classroom design. approximately 40 journals and 60 informative books are on the shelves. the children’s book week (i.e., a national reading project) theme “secrets” is illustrated on the bulletin board and books by anthony horowitz are displayed on a small table [observation 1]. classroom strategies. autonomy support effected by affording choices, offering rationale, and taking the students’ perspective is not so prominent in mrs. k’s teaching style [appendix 1, elements a, c, and d]. she discusses various text genres and text fragments provided in the manual in a systematic way, posing rather standard questions: who?, what?, what about?, etc. in her opinion, the manual offers fascinating texts and nice illustrations with the potential to promote reading pleasure [appendix 1, element b]. although both mrs. k and the school leader consider writing book reviews a questionable motivational strategy, students are required to write 10 reviews of self-selected reading material (i.e., six novels, one informative book, one comic book, and two poems) following an imposed format [appendix 1, element a]. she is enthusiastic about “panel reading” as instructional practice which implies discussing and presenting informative texts in small groups. it gives students opportunities to be more self-determined [appendix 1, element e]. after finishing their appointed tasks, students have the opportunity to read self-selected books or journals individually [appendix 1, elements a and e]. she provides structure by communicating her expectations [appendix 1, element f], offering students support when needed [appendix 1, element h], and providing positive feedback [appendix 1, element i]. mrs. k is greatly involved in interpersonal relationships with her students. she takes time for and expresses enjoyment in the interactions with her de naeghel et al. 90 | f l r students [appendix 1, element j]. the greater attention to structure and involvement compared with autonomy support is reflected in higher scores on the related subscales in the teacher survey [table 2, element b]. her students say they perceive more structure and involvement than autonomy support, but these remain moderate [table 2, element d]. moreover, her students report moderate levels of autonomous reading motivation [table 2, element c]. next to these sdt teaching dimensions, mrs. k acknowledges the value of reading aloud to promote children’s reading motivation. she does not invest a lot of time in it, however. further, mrs. k engages in national reading projects [teacher interview and observation 1]. in the teacher interview she said: “in the children’s book week, i read aloud every day. but otherwise … i don’t have time for it, to my regret.” school-level strategies. mrs. k’s school has a large library, founded and run by the school leader. the library is open during lunch break and puts narrative as well as informative books at students’ disposal. the collection is frequently updated to stimulate students’ curiosity. in this respect, the school leader tries to pass on his “love of reading” to children and their parents by creating a reading culture at school. moreover, he promotes national reading projects [school leader interview]. table 2 teachers’ and students' autonomous reading motivation and execution/perception of teaching dimensions mrs. k mrs. s mr. t teacher a. autonomous reading motivation a 4.00 5.00 3.88 b. execution of teaching dimensions autonomy support a 3.88 3.63 3.88 structure a 4.14 4.00 4.00 involvement a 4.63 3.75 4.63 students c. autonomous reading motivation a mean recreational reading motivation a 3.15 (0.72) 3.63 (.73) 3.98 (.46) mean academic reading motivation a 3.11 (0.88) 3.43 (.80) 3.91 (.57) d. perception of teaching dimensions mean autonomy support a 3.56 (0.67) 2.56 (0.68) 3.76 (0.27) mean structure a 3.52 (0.52) 3.69 (0.43) 3.77 (0.22) mean involvement a 3.45 (0.57) 3.68 (0.65) 3.92 (0.43) note: a subscale scores range from one to five, with five indicating a higher score. standard deviation in parentheses. 3.1.2 promotion of autonomous reading motivation in mrs. s’s classroom context and teacher profile. mrs. s is a 35-year-old teacher with 10 years of teaching experience. she teaches languages, social studies, and sciences half-time in a small school located in the city. there are 19 students in her class, who are on average 11 years old. mrs. s spends about 130 minutes a week on reading instruction. she uses “taalsignaal” as a teaching manual for the dutch language lessons. her preferred teaching methods are whole-class instruction, small-group instruction, and independent work. mrs. s is an autonomously motivated reader, devouring novels, informative books, comics, etc. in her free time as well as for her professional development [teacher interview and table 2, element a]. classroom design. narrative and informative books are on the shelf at the back of the classroom. the collection is often renewed with books from the public library, depending on the themes discussed in the social studies and sciences lessons. approximately 500 narrative and informative books are located in a small library room nearby the classroom [teacher interview and observation 1]. classroom strategies. mrs. s provides autonomy support especially by fitting in with students’ interests [appendix 1, element b], offering rationales [appendix 1, element c], and providing students with de naeghel et al. 91 | f l r opportunities to be initiators of their own behaviour [appendix 1, element e]. for instance, her students are tutors for their third-grade peers in a reading project combining direct instruction in reading comprehension strategies and cross-age peer tutoring to practice their reading skills with self-selected books. in this respect, she explicitly discusses with her students why being a good tutor and using reading strategies is important. mrs. s, the sen coordinator, and the students experience these opportunities to read together as motivating [observation 1, teacher interview, and sen coordinator interview]. after finishing their appointed tasks, students have the opportunity to work independently on additional material or to read self-selected books [appendix 1, elements a and c]. moreover, fifth-grade students write one book review on a self-chosen book (i.e., design a new cover, write a summary, and make a drawing [appendix 1, element a]). mrs. s and the sen coordinator underline the importance of providing students with fascinating texts to promote reading pleasure. in mrs. s’s opinion, the manual does not offer enough interesting texts to practice reading comprehension. therefore, mrs. s often brings new reading material from the public library into the classroom to stimulate students’ willingness to read [appendix 1, element b]. it should be noted, however, that giving students choices occurs primarily during peer tutoring sessions [appendix 1, element a]. mrs. s provides structure by having a clear plan of the day, communicating her expectations [appendix 1, element f], and providing student support [appendix 1, element h]. as part of the reading peer tutoring project she gives direct instruction in reading comprehension strategies, supporting students’ reading competence [appendix 1, element h]. mrs. s further invests a lot in interpersonal relationships with her students. she cares about how students do in class and takes their needs into account as much as possible. in other words, she is involved [appendix 1, element j]. data from the teacher survey illustrate that she especially provides structure and to a somewhat lesser extent is involved with her students and supports their autonomy [table 2, element b]. students’ reports indicate that her students perceive more structure and involvement than autonomy support [table 2, element d]. furthermore, her students report moderate levels of autonomous reading motivation [table 2, element c]. besides implementing the sdt teaching dimensions mrs. s reads aloud frequently. in the teacher interview she said: “i bring books to read aloud to stimulate them. … reading aloud is just for fun. no questions afterwards.” mrs. s further engages in national reading projects [teacher interview]. school-level strategies. as mentioned above, the reading project combining direct instruction in reading comprehension strategies and cross-age peer tutoring is organised across different grades. not only fifth and third grade, but also sixth and second, and fourth and first grade read together in this school reading project. currently, the teachers themselves are responsible for running the project, coordinated by mrs. s. in the early stages of the project, the school leader and sen coordinator were more involved [teacher, school leader, and sen coordinator interview]. furthermore, the sen coordinator reads picture books in all grade levels to introduce new school projects [sen coordinator interview]. finally, there is a study group, in which mrs. s takes part, which works out new ideas regarding reading and reading promotion in staff meetings [school leader interview]. 3.1.3 promotion of autonomous reading motivation in mr. t’s classroom context and teacher profile. mr. t is 34 years old and has 14 years of teaching experience. he teaches fifth grade in a small private school in the countryside. there are 17 students in his class, who are on average 11 years old. mr. t spends about 120 minutes a week on reading instruction. he uses “taalmakker” as a teaching manual for the dutch language lessons. his preferred teaching methods are whole-class instruction and independent work. mr. t especially reads to gain knowledge. he prefers short passages in newspapers and journals, comics, and children’s books [teacher interview] and reports that he is an autonomously motivated reader [table 2, element a]. classroom design. two bookshelves filled with approximately 50 narrative and informative books and two boxes with comics are put at students’ disposal. a bean-bag seat and the step in front of the classroom provide a reading spot [observation 1]. to expand the number of books in the class library, mr. t asks parents to donate comic books that are no longer read at home and to give a book to the class as a birthday gift instead of sweets. each week one student gets the role of librarian by lottery [teacher interview]. de naeghel et al. 92 | f l r classroom strategies. mr. t provides autonomy support by affording choices [appendix 1, element a], fitting in with students’ interests [element b], offering rationales [element c], taking students’ perspective [element d], and providing students with opportunities to be initiators of their own behaviour [element e]. more specifically, he tries to teach reading in a meaningful context (e.g., making a class garden, solving puzzles, keeping abreast of current events [appendix 1, element c]). in his opinion, the manual offers fascinating texts for teaching reading comprehension. in addition, he brings newspaper and journal articles to the class to study, a tradition which is copied by his students [appendix 1, element b]. mr. t likes to challenge his students with group assignments (e.g., making a picture book, searching for key words in various text passages [appendix 1, element e]. he asks his students to make a drawing of a self-selected book during holidays, which is then presented in the classroom [appendix 1, element a]. after finishing their appointed tasks, students have the opportunity to read or draw [appendix 1, element e]. he provides structure by passing on his expectations [element f], providing optimal challenges [element g], offering support to his students [element h], and giving them constructive feedback [appendix 1, element i]. moreover, he is greatly involved in interpersonal relationships with his students. mr. t attaches great importance to creating a respectful classroom atmosphere and listening to students’ personal stories [appendix 1, element j]. in the teacher survey mr. t reports that he is highly involved with his students [table 2, element b]. students’ reports confirm they experience autonomy support, structure, and involvement [table 2, element d]. moreover, the students report that they are autonomously motivated to read [table 2, element c]. next to implementing the sdt teaching dimensions, mr. t reads aloud each friday afternoon to create a stimulating reading atmosphere. in the teacher interview he stated: “… children really enjoy it. i create a nice reading atmosphere, reading expressively and immersing myself in the book … and by doing so the interest of students in books certainly grows.” his students are involved in the selection of the book and each finished book results in a creative project (e.g. a play, a scale model of the village described in the book). furthermore, mr. t engages in national reading projects [teacher interview]. school-level strategies. mr. t’s school pays a lot of attention to reading. his school organises an overall reading project from kindergarten to sixth grade. the project was set up by mrs. l, the school’s sen coordinator and literacy coach (steckel, 2009; walpole & blamey, 2008), and the school leader. the project’s theme is a story about a boy, “jonah sprout,” who meets all kinds of letters during a boat trip. his boat (an old yard wagon) comes ashore in the school’s playground. in the school “jonah sprout” is represented by a puppet [sen coordinator and school leader interview]. mrs. l, the literacy coach, acts as a pioneer for all reading activities at school. she promotes the children’s book week, the reading aloud week, and poetry day (national reading projects). during staff meetings she illustrates possible activities and provides teachers with the necessary reading material. the introduction and closure of all reading activities is a collective school event. each activity is introduced by a play with “jonah sprout” in the leading role and closed with a presentation of reading activities of each grade. moreover, mrs. l organises a book club for students of fifth and sixth grade in “jonah sprout”'s boat. during book club time, books are discussed and approached in a creative way (e.g., reading and cooking a recipe, improvising the end of a story [sen coordinator interview]). 3.2 horizontal analysis 3.2.1 classroom strategies for promoting autonomous reading motivation sdt’s teaching dimensions. in line with more general sdt research, the teaching dimensions of autonomy support, structure, and involvement (reeve, 2002; skinner & belmont, 1993) could be identified as critical strategies promoting autonomous reading motivation in particular and this in each of the three cases. the selected teachers, however, especially differ in the extent and manner of the autonomy support they provide. de naeghel et al. 93 | f l r as mentioned above, autonomy support primarily nurtures students’ need for autonomy or selfdetermined behaviour (ryan & deci, 2000). students’ autonomy is first supported by giving students ageappropriate choices (appendix 1, element a; reeve, 2002; skinner & belmont, 1993). in the three cases this is mainly reflected in opportunities to select books for independent reading and book reviews. whereas mrs. k provides an imposed format for the book reviews, mrs. s and mr. t allow more creativity and personal input. in addition, mr. t occasionally provides choices between different assignments. in all three cases, however, there are still opportunities to enlarge the number of choices regarding what students read and how they engage in and complete reading tasks (gambrell, 2011). second, the three teachers recognise the importance of fitting in with students’ interests to promote autonomous reading motivation (appendix 1, element b; reeve, 2002; skinner & belmont, 1993). in this respect especially, mrs. s and mr. t bring supplementary reading material into the classroom related to topics studied in social studies and sciences, students’ social environment, or the news. furthermore, each of the three teachers has a classroom library, containing narrative and informative books, and sometimes comics or journals, which are at students’ disposal during independent reading. third, students’ autonomy is encouraged by the offer of rationales (appendix 1, element c; skinner & belmont, 1993). mrs. s clearly explains to her students why she teaches certain topics. mr. t, on the other hand, tries to offer a rationale by teaching reading in a meaningful context. in contrast, mrs. k does not seem to invest a lot of effort in this strategy. fourth, taking the students’ perspective was only explicitly observed in mr. t’s classroom (appendix 1, element d; reeve, 2002; skinner & belmont, 1993). finally, the three selected teachers apply various instructional strategies, such as panel reading, group work, cross-age peer tutoring, and independent reading, that allow students to be more self-determined or volitional and therefore fulfil the need for autonomy and encourage autonomous motivation for reading (appendix 1, element e; reeve, 2002; sierens, 2010; skinner & belmont, 1993). in general, mr. t provides the highest level of autonomy support by providing choices, fitting in with students’ interests, teaching reading in a meaningful context, taking the students’ perspective, and providing opportunities to his students to be initiators of their own behaviour [appendix 1, elements a to e]. from the student questionnaires it can be noted that his students corroborate to perceive the highest level of autonomy support [table 2, element d] and, moreover, report the highest level of recreational and academic autonomous reading motivation [table 2, element c]. the fact that mr. t’s students indicate not only the highest perceived autonomy support but also the highest level of autonomous reading motivation is certainly an argument in favour of his autonomy-supportive teacher behaviour (reeve, 2002; skinner & belmont, 1993). according to the interpretative data, mrs. k, in contrast, appears to be the least autonomy-supportive of the three participating teachers. a closer look at the results of the teacher and student questionnaire suggests that mrs. k and mr. t report equal practice of autonomy-supportive behaviour [table 2, element b]. furthermore, mrs. k's students perceive a higher level of autonomy support than mrs. s’s students [table 2, element], although her students do report lower levels of recreational and academic autonomous reading motivation [table 2, element c]. this finding illustrates how differences in research methods (i.e., interpretative or quantitative) can lead to different perspectives and conclusions, as detailed observation and questioning of stakeholders (i.e., interpretative methods) and information gathering by surveys (i.e. quantitative methods) probably do not address the research questions in exactly the same manner. nevertheless, these methods can jointly help to create a fuller and more nuanced picture of what exactly happens in the classroom. the teaching dimension structure, which promotes students’ need for competence (reeve, 2002; sierens, 2010; skinner & belmont, 1993), is more or less equally addressed by the three case study teachers. all three communicate their expectations to the students, offer help and support, and provide positive feedback [appendix 1, elements f to i]. in addition, mrs. s invests the most time in explicitly teaching reading comprehension strategies to foster students’ competence in reading [appendix 1, element h]. mr. t invests the most in providing optimal challenges by giving stimulating group tasks [appendix 1, element g]. de naeghel et al. 94 | f l r results of the teacher questionnaire corroborate the roughly equal levels of structure in their classrooms [table 2, element b], although the students of mrs. s and mr. t experience structure related to teaching practices slightly more in their classrooms [table 2, element d]. the teaching dimension of involvement, associated with the need for relatedness (reeve et al., 2004; skinner & belmont, 1993), is most prominent in the teaching style of the three selected teachers. all three invest a lot in interpersonal relationships with their students by explicitly making time to listen to students’ personal stories and interests, expressing enjoyment in the interaction with their students, and taking students’ needs into account [appendix 1, element j]. furthermore, mrs. s and her school’s sen coordinator explicitly point to the importance of reading together as a motivating strategy, confirming the relevance of involvement between students (reeve et al., 2004; skinner & belmont, 1993) and opportunities to collaborate as in cori (guthrie & cox, 2001). according to the teachers’ responses in the teacher questionnaire, mrs. k and mr. t seem to be most highly involved with their students. moreover, mr. t receives the highest score on involvement from his students [table 2, elements b and d], corroborating the qualitative interview and observational data [appendix 1, element j]. reading aloud. next to the teaching dimensions of autonomy support, structure, and involvement, reading aloud is recognised as an important strategy to promote autonomous reading motivation in the three cases (pecjak & kosir, 2008). in particular, mrs. s and mr. t often read aloud to stimulate their students’ reading behaviour, whereas mrs. k reports that she generally does not have enough time for it. further, mr. t explicitly indicates that he creates a stimulating reading atmosphere and involves his students in the selection of the reading material. 3.2.2 school-level strategies for reading promotion the three participating teachers belong to schools which recognise the importance of reading. in mrs. k’s school the presence of a large library and the dedication of the school leader to managing the library communicate to teachers, students, and parents how strongly reading is appreciated by the school, and, hence, that encouraging reading is significant. in mrs. s’s school, mrs. s plays a prominent role herself in coordinating a reading project which combines direct instruction in reading comprehension strategies with cross-age peer tutoring across different grades and in participating in a study group on reading and reading promotion. moreover, the sen coordinator of her school reads picture books in all grade levels. the school leader and literacy coach of mr. t’s school organise an overall reading project from kindergarten to sixth grade. additionally, the literacy coach supports the teachers in promoting reading in their classroom and organises a book club for fifth and sixth graders. in sum, each of the three teachers belongs to a school that palpably acknowledges the importance of reading and therefore confirms that a school culture focusing on school-wide reading has potential to encourage teachers’ and students’ engagement (daniels & steres, 2011) and motivation for reading. 4. discussion and conclusion in order to break through the declining trend in reading motivation throughout children’s educational careers, it is important to identify strategies which enable teachers to encourage students’ autonomous reading motivation. in this respect, the present study furthers an underexposed field in reading motivation research by studying and identifying the strategies of teachers excellent in the promotion of volitional or autonomous reading motivation. sdt formulates general guidelines or teaching dimensions to facilitate autonomous motivation (ryan & deci, 2000). these general teaching dimensions of autonomy support, structure, and involvement could be identified as critical strategies to promote autonomous motivation for reading in the classroom practice of the selected teachers. in this respect, the present study points to the theoretical significance of adopting these teaching dimensions in reading motivation research, as the sdt teaching dimensions have rarely been explicitly studied in the specific context of reading motivation before and on the basis of our results appear to be transferable and relevant to this research area. it should be noted that the participating de naeghel et al. 95 | f l r teachers more or less equally addressed the teaching dimensions of structure and involvement, whereas they differed particularly in the extent and manner of the autonomy support they provided. this indicates that even some of the selected teachers apparently invest less in or have more difficulties with supporting their students’ autonomy and suggests autonomy support is a powerful strategy with opportunities for growth. next to the significance of the sdt teaching dimensions, the results confirm the relevance of reading aloud as an effective classroom strategy to stimulate students’ willingness to read (fisher et al., 2011; gambrell et al., 1996; pecjak & kosir, 2008). further research is, however, needed to collect more detailed information on teachers’ specific approach to reading aloud (lane & wright, 2011). what is of interest as well is that the teachers considered as excellent in promoting autonomous reading motivation belong to schools that invest in reading at school level, underlining the importance of a school-wide interest in and attention to reading (daniels & steres, 2011; taylor et al., 2010). as the role of the school and school culture is still underexposed in reading motivation research, follow-up studies could enlarge their focus to how schools (e.g., school members [teachers, school leaders, literacy coaches, etc.], policy, projects, and curriculum) contribute to a supportive reading environment in order to formulate additional guidelines for school practice. the identified strategies for promoting autonomous reading motivation are of particular importance for teaching practice and accordingly for teachers’ professional development in both pre-service and inservice training. considering the significant influence of the home environment on students’ reading motivation (swalander & taube, 2007), teachers can play a crucial role in positively motivating all of their students to read (gambrell, 1996; santa et al., 2000). in this way, they invest in equipping their students with the necessary reading competencies to be successful in today’s society, striving for equal opportunities for all. further, the identified strategies are valuable as tools for reflection on and improvement of teachers’ and schools’ reading promotion approach and practice. first, it appears that the sdt teaching dimensions (reeve, 2002; skinner & belmont, 1993) can be implemented and integrated relatively easily in classroom practice, as they merely involve a change of attitude and awareness of sdt’s frame of reference. in this respect, teachers can make their own reading activities more supportive of autonomous reading by applying the sdt teaching dimensions (e.g., providing choices between different reading materials, offering rationales for learning activities, providing positive feedback to their students) without having to make time-consuming changes to their reading curriculum. as mentioned above, teachers can invest particularly in making their reading activities more autonomy-supportive (e.g., providing choices between different activities, matching students' interests, taking the students’ perspective), as even teachers indicated as excellent in promoting autonomous reading motivation still have opportunities for growth. additionally, as reading aloud remains an important and valuable activity in late primary education, teachers can invest more time in reading aloud in class to stimulate children’s interest in reading. they can underline its significance for instance by scheduling reading aloud in the plan for the week. moreover, teachers and schools can be inspired by the described school-level reading strategies to further their own school-wide reading policy. this study focused on the strategies of teachers considered to be excellent in promoting autonomous reading motivation. it can be expected that the identified strategies will be less explicitly present in the daily classroom practice and schools of teachers who are less excellent or even rather poor at promoting autonomous reading motivation. hence, these strategies can function as guidelines to improve their reading activities. nevertheless, further research should offer insight into the classroom practices of teachers who are less than excellent in promoting autonomous reading motivation and explore possibilities to improve their skills through teacher training. in sum, the present study points to the theoretical and practical significance of adopting sdt’s teaching dimensions (i.e., autonomy support, structure, and involvement) as well as to reading aloud as critical strategies to encourage students’ autonomous reading motivation in the classroom. moreover, a school culture supporting students' and teachers' interest in reading is essential. de naeghel et al. 96 | f l r keypoints this study extends sdt research by applying sdt in research on primary school students and reading motivation and by adopting an embedded mixed-method design this study contributes to reading motivation research by identifying the strategies of teachers excellent in the promotion of reading motivation this study indicates autonomy support, structure, and involvement as critical strategies to promote autonomous reading motivation in the classroom this study confirms the relevance of reading aloud as an effective classroom strategy to stimulate students’ willingness to read this study builds on the literature on reading motivation by highlighting the critical role of the school’s reading culture for teachers’ practices acknowledgements this research was supported by a grant from the special research fund of ghent university (bijzonder onderzoeksfonds universiteit gent). references baker, l., & wigfield, a. (1999). dimensions of children's motivation for reading and their relations to reading activity and reading achievement. reading research quarterly, 34, 452-477. doi:10.1598/rrq.34.4.4 becker, m., mcelvany, n., & kortenbruck, m. (2010). intrinsic and extrinsic reading motivation as predictors of reading literacy: a longitudinal study. journal of educational psychology, 102, 773-785. doi: 10.1037/a0020084 belmont, m., skinner, e., wellborn, j., & connell, j. (1988). teacher as social context: a measure of student perceptions of teacher provision of involvement, structure, and autonomy support [tech. rep. no. 102]. rochester, ny: university of rochester. black, a. e., & deci, e. l. (2000). the effects of instructors' autonomy support and students' autonomous motivation on learning organic chemistry: a self-determination theory perspective. science education, 84, 740-756. doi: 10.1002/1098-237x(200011)84:6<740::aid-sce4>3.0.co;2-3 chatzisarantis, n. l. d., & hagger, m. s. (2009). effects of an intervention based on self-determination theory on self-reported leisure-time physical activity participation. psychology & health, 24, 29-48. doi: 10.1080/08870440701809533 creswell, j. w., & plano, c. v. l. (2007). designing and conducting mixed methods research. thousand oaks, calif.: sage publications. daniels, e., & steres, m. (2011). examining the effects of a school-wide reading culture on the engagement of middle school students. research in middle level education online, 35, 1-13. deci, e. l., & ryan, r. m. (1987). the support of autonomy and the control of behavior. journal of personality and social psychology, 53, 1024-1037. doi: 10.1037//0022-3514.53.6.1024 de naeghel, j., van keer, h., vansteenkiste, m., & rosseel, y. (2012). the relation between elementary students’ recreational and academic reading motivation, reading frequency, engagement, and comprehension: a self-determination theory perspective. journal of educational psychology, 104, 10061021. doi: 10.1037/a0027800 de naeghel et al. 97 | f l r de naeghel, j., & van keer, h. (2013). the relation of student and class-level characteristics to primary school students’ autonomous reading motivation: a multilevel approach. journal of research in reading, 36, 351-370. doi: 0.1111/j.1467-9817.2013.12000.x edmunds, k. m., & bauserman, k. l. (2006). what teachers can learn about reading motivation through conversations with children. the reading teacher, 59, 414-424. doi: 10.1598/rt.59.5.1 fisher, d., flood, j;, lapp, d., & frey, n. (2004). interactive reading-alouds: is there a common set of implementation practices? the reading teacher, 58, 8-17. doi:10.1598/rt.58.1.1 gambrell, l. b. (1996). creating classroom cultures that foster reading motivation. the reading teacher, 50, 14-25. gambrell, l. b. (2011). seven rules of engagement: what's most important to know about motivation to read. the reading teacher, 65, 172-178. doi: 10.1002/trtr.01024 gambrell, l. b., palmer, b. m., codling, r. m., & mazzoni, s. a. (1996). assessing motivation to read. reading teacher, 49, 518-533. doi: 10.1598/rt.49.7.2 gaskins, i. w. (2008). ten tenets of motivation for teaching struggling readers – and the rest of the class. in r. fink & s. j. samuels (eds.), inspiring reading success. interest and motivation in an age of highstakes testing (pp. 98-116). newark, de: international reading association. guthrie, j. t. (2004). classroom contexts for engaged reading: an overview. in j. t. guthrie, a. wigfield & k. c. perencevich (eds). motivating reading comprehension. concept-oriented reading instruction (pp. 1-24). mahwah, nj: lawrence erlbaum associates. guthrie, j. t., & cox, k. e. (2001). classroom conditions for motivation and engagement in reading. educational psychology review, 13, 283-302. doi: 10.1006/ceps.1999.1044 guthrie, j. t., mcrae, a., & klauda, s. l. (2007). contributions of concept-oriented reading instruction to knowledge about interventions for motivations in reading. educational psychologist, 42, 237-250. doi: 10.1080/00461520701621087 guthrie, j. t., & wigfield, a. (2000). engagement and motivation in reading. in m. l. kamil, p. b. mosenthal, p. d. pearson, & r. barr (eds.), handbook of reading research: volume iii (pp. 403-422). mahwah, nj: lawrence erlbaum associates. guthrie, j. t., wigfield, a., humenick, n. m., perencevich, k. c., taboada, a., & barbosa, p. (2006). influences of stimulating tasks on reading motivation and comprehension. journal of educational research, 99, 232-245. doi:10.3200/joer.99.4.232-246 guthrie, j. t., wigfield, a., & vonsecker, c. (2000). effects of integrated instruction on motivation and strategy use in reading. journal of educational psychology, 92, 331-341. doi:10.1037/00220663.92.2.331 ivey, g., & broaddus, k. (2001). “just plain reading”: a survey of what makes students want to read in middle school classrooms. reading research quarterly, 36, 350–377. doi: 10.1598/rrq.36.4.2 lane, h. b., & wright, t. l. (2011). maximizing the effectiveness of reading aloud. the reading teacher, 60, .668-675. doi: 10.1598/rt.60.7.7 marinak, b.a. & gambrell, l.b. (2007). choosing and using informational text for instruction in the primary grades. in b.j. guzzetti (eds.) literacy for a new millennium: early literacy. (pp. 141-154). westport, ct: praeger. meyer, l.a., wardrop, j.l., linn, r.l., & hastings, c.n. (1993). effects of ability and settings on kindergartners’ reading performance. journal of educational research, 86, 142–160. doi: 10.1080/00220671.1993.9941153 miles, m., & huberman, m. (1994). qualitative data analysis: an expanded sourcebook. thousand oaks, ca: sage. mohan, l., lundeberg, m. a., & reffitt, k. (2008). studying teachers and schools: michael pressley's legacy and directions for future research. educational psychologist, 43, 107-118. doi: 10.1080/00461520801942292 morrow, l.m., & gambrell, l.b. (2002). literature-based instruction in the early years. in s.b. neuman & d.k. dickinson (eds.), handbook of early literacy research (pp. 348–360). new york: guilford. mouratadis, m., vansteenkiste, m., lens, w., & sideridis, g. (2008). the motivating role of positive feedback in sport and physical education: evidence for a motivational model. journal of sport & exercise psychology, 30, 240-268. de naeghel et al. 98 | f l r mullis, i. v. s., martin, m. o., kennedy, a. m., & foy, p. (2007). iea’s progress in international reading literacy study in primary school in 40 countries. chestnut hill, ma: timms & pirls international study center, boston college. organisation for economic co-operation and development [oecd]. (2004). learning for tomorrow’s world. first results from pisa 2003. paris: oecd. pecjak, s., & kosir, k. (2008). reading motivation and reading efficiency in third and seventh grade pupils in relation to teachers' activities in the classroom. studia psychologica, 50, 147-168. pressley, m., rankin, j., & yokoi, l. (1996). a survey of instructional practices of primary teachers nominated as effective in promoting literacy. elementary school journal, 96, 363-384. doi: 10.1086/461834 reeve, j. (2002). self-determination theory applied to educational settings. in e. l. deci, & r. m. ryan (eds.), handbook of self-determination research (pp. 183-203). rochester, ny: university of rochester press. reeve, j. (2009). understanding motivation and emotion. hoboken, nj: john wiley & sons, inc. reeve, j., jang, h., carrell, d., jeon, s., & barch, j. (2004). enhancing students' engagement by increasing teachers' autonomy support. motivation and emotion, 28, 147-169. doi:10.1023/b:moem .0000032312.95499.6f ryan, r. m., & deci, e. l. (2000). self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. american psychologist, 55, 68-78. doi:10.1037/0003-066x.55.1.68 santa, c. m., williams, c. k., ogle, d., farstrup, a. e., au, k. h., baker, b. m., et al. (2000). excellent reading teachers: a position statement of the international reading association. journal of adolescent & adult literacy, 44, 193-199. soenens, b., & vansteenkiste, m. (2005). antecedents and outcomes of self-determination in three life domains: the role of parents' and teachers' autonomy support. journal of youth and adolescence, 34, 589-604. doi: 10.1007/s10964-005-8948-y sierens, e. (2010). autonomy-supportive, structuring, and psychologically controlling teaching: antecedents, mediators, and outcomes in late adolescents [unpublished dissertation]. leuven: kuleuven. sierens, e., vansteenkiste, m., goossens, l., soenens, b., & dochy, f. (2009). the interactive effect of perceived teacher autonomy support and structure in the prediction of self-regulated learning. british journal of educational psychology, 79, 57-68. doi: 0.1348/000709908x304398 skinner, e. a., & belmont, m. j. (1993). motivation in the classroom reciprocal effects of teacherbehavior and student engagement across the school year. journal of educational psychology, 85, 571581. doi: 10.1037/0022-0663.85.4.571 steckel, b. (2009). fulfilling the promise of literacy coaches in urban schools: what does it take to make an impact? the reading teacher, 63, 14-23. doi: 10.1598/rt.63.1.2 swalander, l., & taube, k. (2007). influences of family based prerequisites, reading attitude, and selfregulation on reading ability. contemporary educational psychology, 32, 206-230. doi:10.1016/j.cedpsych.2006.01.002 taylor, b. m., pearson, p. d., clark, k., & walpole, s. (2010). effective schools and accomplished teachers: lessons about primary-grade reading instruction in low-income schools. the elementary school journal, 101, 121-165. doi: 10.1086/499662 tessier, d., sarrazin, p., & ntoumanis, n. (2008). the effects of an experimental programme to support students' autonomy on the overt behaviours of physical education teachers. european journal of psychology of education, 23(3), 239-253. doi: 10.1007/bf03172998 vanderburg, m., & stephens, d. (2010). the impact of literacy coaches: what teachers value and how teachers change. the elementary school journal, 111, 141-163. doi: 10.1086/653473 vansteenkiste, m., simons, j., lens, w., soenens, b., & matos, l. (2005). examining the motivational impact of intrinsic versus extrinsic goal framing and autonomy-supportive versus internally controlling communication style on early adolescents' academic achievement. child development, 76, 483-501. doi:10.1111/j.1467-8624.2005.00858.x walpole, s., & blamey, k. l. (2008). elementary literacy coaches: the reality of dual roles the reading teacher, 62, 222-231. doi: 10.1598/rt.62.3.4 http://www.vopspsy.ugent.be/pdfs/download.php?own=mvsteenk&file=bjep2009.pdf http://www.vopspsy.ugent.be/pdfs/download.php?own=mvsteenk&file=bjep2009.pdf de naeghel et al. 99 | f l r watkins, m. w., & coffey, d. y. (2004). reading motivation: multidimensional and indeterminate. journal of educational psychology, 96, 110-118. doi: 10.1037/0022-0663.96.1.110 whitehurst, g.j., zevenbergen, a.a., crone, d.a., schultz, m.d., velting, o.n., & fischel, j.e. (1999). outcomes of an emergent literacy intervention from head start through second grade. journal of educational psychology, 91, 261–272. doi: 10.1037/0022-0663.91.2.261 wigfield, a., guthrie, j. t., perencevich, k. c., taboada, a., klauda, s. l., mcrae, a., & barbosa, p. (2008). role of reading engagement in mediating effects of reading comprehension instruction on reading outcomes. psychology in the schools, 45, 432-445. doi:10.1002/pits.20307 yin, r. k. (1989). case study research: design and methods. thousand oaks, ca: sage. appendix 1 examples of the execution of sdt’s teaching dimensions in mrs. k’s, mrs. s’s, and mr. t’s classroom reading activities examples and illustrations mrs. k mrs. s mr. t strategies to promote autonomous reading motivation autonomy support a. providing choices students write ten reviews of self-selected reading material, i.e., six novels, one informational book, one comic book, and two poems, following an imposed format. “i oblige them a little, since there are children who wouldn’t read anything otherwise. … they have still read something and maybe it motivates them to choose a book on their own … but possibly it lets them take a dislike to reading.” [teacher interview] during a group assignment students have the opportunity to choose their group members. [observation 1] students choose books or journals for independent reading. [observation 1] students write one book review on a self-chosen book (e.g., design a new cover, write a summary, and make a drawing [teacher interview]). students choose reading books during cross-age peer tutoring sessions and independent reading. [observation 1] every holiday students make a drawing of a self-selected book, which is presented in class afterwards. [teacher interview] students are involved in the selection of the book that is read aloud. [teacher interview] during a group assignment students have the opportunity to choose between different text passages, e.g. newspaper, children’s newspaper, difficult sentences, … [observation 1] during a group assignment students have the opportunity to choose their partner. [observation 2] students choose books or comics for independent reading [observation 1] b. fitting interests according to mrs. k the manual offers fascinating texts and nice illustrations. [teacher interview] students are very enthusiastic about the mystery theme of the group assignment. [observation 1] mrs. s often brings new reading material from the public library into the classroom, since the manual does not offer many interesting texts in her opinion. [teacher interview and observations] “a story should be exciting, certainly for that age! … and some children, not all of in mr. t’s opinion the manual offers fascinating texts. [teacher interview] mr. t brings newspaper and journal articles to the class to study with his students. [teacher interview and observation 1] “i search for ageappropriate things (to read) such as first love. the de naeghel et al. 100 | f l r them of course, are interested in all kinds of details about famous historical figures.” [teacher interview] “we try to relate lessons to real-world experiences, if possible … to involve the children and stimulate their interests. … however, i still have to impose the material that i have to teach.” [teacher interview] “i think the most important thing in motivating children is matching their interests.” [sen coordinator interview] giggling, the familiarity, that’s great … i especially start from reality. … in short, … situations from their environment.” [teacher interview] “don’t force them. show them that there is something about their interests, perhaps a journal, an informative book, a novel, ... there is something for everyone.” [teacher interview] c. offering rationale “clearly indicating lesson goals … i don’t do that.” [teacher interview] mrs. s discusses with her students why being a good tutor and using reading strategies is important. [observation 1] “i try to communicate why we do certain things. for instance, when i started class this morning i clearly indicated what we would do and why. it motivates them to engage in the activity.” [teacher interview] mr. t teaches reading in a meaningful context, e.g., making a class garden, solving puzzles, reading about current events, … [teacher interview] “… offering a whole range of possibilities, preferably as integrated as possible … makes them realize the relevance …” [school leader interview] d. taking students’ perspective reciting some difficult sentences, a boy stumbles over his words. students laugh. the boy feels mocked, sits down on the ground, and starts crying. mr. t lets him know that it is okay, accepts his emotional outburst, gives him some time, and talks to him during playtime. [observation 1] e. initiator of own behaviour during “panel reading” students discuss and present informative texts in small groups. “the students really like panel reading.” [teacher interview] after finishing the appointed tasks, students have the opportunity to read. [teacher interview] the timing during the first group assignment is very restrictive. students have seven minutes to find the answer to some questions on the blackboard concerning the author of a book. “still three minutes. … still 30 seconds! … stop!” [observation 1] fifth-graders are tutors for their third-grade peers in a reading project combining direct instruction in reading comprehension strategies and cross-age peer tutoring. “… in reading comprehension children often work together … most children enjoy working together.” [teacher interview and observation 1] “certainly reading together … stimulating by reading together.” [sen coordinator interview] after finishing the appointed tasks, students can work on additional material or read a book. “… additional material to do on their own mr. t provides challenging group tasks such as making a picture book, creating a play, making a class garden, … “in the spring we make a class garden. students look for a step-by-step plan to make the garden.” [teacher interview and observations] after finishing the appointed tasks, students have the opportunity to read or draw. [teacher interview] de naeghel et al. 101 | f l r has to be fun and motivating and can be completed at their own speed.” [teacher interview] structure f. communicating expectations mrs. k communicates step by step what the children are expected to do in the group assignment. first, make three groups of three and two groups of four. go with your group to a computer and open the dutch webpage of wikipedia. search for an answer to the following questions. … [observation 1] “formulate this in a sentence, please.” [observation 1] “planning is very important for children, … knowing first we will do this, and afterwards that, …” [teacher interview] mr t clearly communicates how to fulfil the group assignment. first, go to your group. second, choose a group leader. third, turn over the sheet with the assignment. … [observation 1] g. providing optimal challenges mr t provides challenging group tasks such as making a picture book, creating a play, making a class garden, … [teacher interview and observations] h. offering help and support the teacher drops hints to help the children find the right answer to the riddles and questions in the group assignment. [observation 1] mrs. s provides direct instruction on reading comprehension strategies. [teacher interview] after reading the text, she helps the children to answer the more difficult questions. for instance, she rereads a certain passage aloud. [observation 2] the teacher drops hints on how to decode the mysterious title of one of the assignments on the blackboard. [observation 1] mr. t suggests strategies for the social studies and sciences’ test. [observation 1] i. providing positive feedback “well read.” [observation 1] “very good, brief and to the point.” [observation 1] “and giving positive feedback. great! stimulate. okay, doesn’t matter, chin up.” [teacher interview] involvement (j.) the children may whisper the answer of the riddles or questions in mrs. k’s ear. [observation 1] when mrs. k reads a book, she asks the children to come and sit around her with cushions. [observation 1] “i believe the school library encourages a strong exchange between students. … some children say to each other: i read a nice book. you should certainly read it too! ” [teacher interview] “who is already thinking about his future?” children enthusiastically tell the teacher their dreams about future professions. mrs. s takes time to listen to her students’ stories. [observation 2] “i try to take the children into account as much as possible.” [teacher interview] “i listen to their story, their interests, their favourite books … i encourage their participation.” [teacher interview] “we become equal, respecting each other. i respect them, they respect me.” [teacher interview] “i listen if there is something they want to tell me … i go to a soccer game, a dance show which my students are taking part in. i show my interest in more than the regular lessons, e.g. how was soccer or rope skipping? i have a little chat with them in the playground. i just make sure that they like me and vice versa.” de naeghel et al. 102 | f l r [teacher interview] one of the students talks about difficulties at home. mr. t listens carefully and gives moral support. [observation 1] frontline learning research 6 (2014) 35-45 issn 2295-3159 corresponding author: vanessa svihla, organization, information & learning sciences, university of new mexico, albuquerque nm 87131, usa. email: vsvihla@unm.edu doi: http://dx.doi.org/10.14786/flr.v2i4.114 35 | f l r advances in design-based research vanessa svihla a a university of new mexico, usa article received 26 may 2014 / revised 1 november 2014 / accepted 11 november 2014 / available online 23 december 2014 abstract design-based research (dbr) is a core methodology of the learning sciences. historically rooted as a movement away from the methods of experimental psychology, it is a means to develop ―humble‖ theory that takes into account numerous contextual effects for understanding how and why a design supported learning. dbr involves iterative refinement of both designs for learning and theory; this process is illustrated with retrospective analysis of six dbr cycles. calls for educational research to parallel medical research has led learning scientists to strive for more specific standards about what constitutes dbr and what makes it desirable, especially regarding robustness and rigor. a recent trend in dbr involves efforts to extend the reach through scalability. these developments potentially endanger the designerly nature of dbr by orienting focus toward generalizability, meaning researchers must be vigilant in their pursuit of understanding how and why learning occurs in complex contexts. keywords: design-based research; learning sciences; research methods v. svihla 36 | f l r 1. overview of design-based research design-based research (dbr) is a core methodology of the learning sciences. begun as a movement away from experimental psychology, dbr was proposed as means to study learning amidst the ―blooming, buzzing confusion‖ of classrooms (brown, 1992, p. 141). it is a way to develop theory that takes into account numerous contextual effects for understanding how and why a design supports learning; these theories are ―humble in that they target domain-specific learning processes‖ (cobb, confrey, disessa, lehrer, & schauble, 2003, p. 9). dbr involves iterative refinement of both designs for learning and theory (brown, 1992; collins, 1992; the design-based research collective, 2003). this paper outlines the methodological standards for conducting dbr, illustrated with an example, and describes recent advances. 2. methodological standards for conducting design-based research dbr is sometimes conflated with mixed methods or action research; this, paired with calls for educational research to parallel medical research has led learning scientists to strive for more specific standards about what constitutes dbr and what makes it desirable, especially regarding robustness and rigor. this section details current methodological standards for conducting dbr. 2.1. a collaborative effort conducted in context dbr is typically conducted as a team of researchers, designers, and practitioners with intensive planning and debriefing sessions throughout the process. rather than a wholly researcher-driven process, practitioners generally have greater ownership of the process. working collaboratively, they identify a practical problem (reeves, 2006) that is then investigated through literature review, learning theory, and question posing. the intervention instantiates this learning theory into the design. because learning is understood to be a process, and because dbr seeks understanding of how learning occurs, process data are prioritized in dbr, such as video records and artifacts of student work. this approach allows researchers to be opportunistic when something surprising or emergent occurs. the notion that emergence plays a central role in dbr is a shift away from the more positivistic origins in which variables are well-known a priori (collins, 1992). dbr allows for intervention while yet valuing the importance of social interaction rather than social isolation (collins, joseph, & bielaczyc, 2004). this resonates the basic belief by learning scientists that learning is a fundamentally social, interactional process. by occurring in classrooms rather than in laboratories, dbr also allows for testing of designs and theory that address ―the complexity that is a hallmark of educational settings‖ (cobb et al., 2003, p. 9). the challenge is to apply lessons learned in context to a broader range of settings (barab & squire, 2004). because dbr may not be replicated in the classical sense, given strong ties to context, it is critical to share the design along with thick description (barab & squire, 2004). 2.2. iterative cycles refine the design and the theory because of the contextual nature of dbr, some view dbr as a means to generate, but not validate conjectures about learning (sandoval, 2004); however, because such conjectures are made visible in designs for learning, they become testable through iterative refinement. simply conducting one study in the field does not qualify, although it may be reported as one cycle in a longer dbr effort. iterative refinement across contexts allows conjectures to become robust (disessa & cobb, 2004) by placing theory ―in harm’s way‖ (cobb et al., 2003, p. 10). the development of interactive learning assessments (ilas) illustrates the iterative refinement process (mckay, cantarero, svihla, yakes jimenez, & castillo, 2014; phillips et al., 2009; svihla et al., v. svihla 37 | f l r 2010; svihla, phillips, et al., 2009; svihla, vye, et al., 2009; svihla et al., 2013; yakes et al., 2013). ilas place the learner in an authentic, professional role giving advice to virtual clients. ilas were first developed in response to a call for high school biology assessments that did not pause learning, but instead assessed students as they learned; more specifically, we aimed to assess how students used resources to solve problems that were new to them. because this call came from an organization interested in using our designs for all schools in one state, we faced early challenges; our design decisions were driven by the need for scalability. this led us to seek school partners to test our designs, but meant that we neglected some of the contextual influences that are typical of dbr. initially, we did not involve instructors in the design process extensively, but we did debrief with them to inform redesign. we partnered with subject matter experts (e.g., a genetic counselor or dietitian) who helped ensure the problems reflected authentic professional practices, as this was central to our humble theory. our designs for and theory of learning evolved through six iterations (figure 1 & table 1). we initially provided authentic, real-world problems posed by virtual clients and access to resources as a way to support students to solve complex problems. students took on the role of interns and gave counsel to virtual (avatar) clients. our first design succeeded in supporting learning, but was too open-ended to be a useful assessment at scale. beginning with iteration 2, we designed more specified sequences of questions and provided feedback from a virtual supervisor. we found the ilas supported learning and provided useful data for assessment, but the student experience was too linear and scaffolded. we moved to a new setting—a university nutrition program seeking to provide students with ways to learn about professional practices prior to internships as a means to recruit and retain diverse students (svihla et al., 2013). with this different motivation driving our work, we sought to bring instructors more centrally into the role of designers of cases. to offset the linear feel of the cases, we sought to support greater agency, providing opportunities to make choices among story-like branches. instructors found it cumbersome to design such cases. instead of distancing the instructors from the design process, we changed how we instantiated agency into the design, creating short story-like loops; students could explore as many or as few of the loops as they liked. in these versions, students learned content and professional practices, and they enjoyed the opportunity to explore further according to their level of interest. retrospective analysis of dbr cycles provides an opportunity to ―see order, pattern, and regularity‖ in messy, complex settings (disessa & cobb, 2004, p. 84) and supports the development of ―useful, generalizable theories‖ (edelson, 2002, p. 112). this analysis includes considering the conditions for success (dede, 2004) and highlights the need to report failures (o'neill, 2012). retrospective analysis of the six iterations – across varied contexts (rural, urban; high school, university; biology, nutrition) – highlights areas where our theory is robust: students consistently learned by taking on real roles and solving challenges posed by virtual clients. this hinged on our ability to place students in roles they could understand; when the role was further from their experience, the addition of vignettes of the virtual supervisor explaining the role bridged this gap. the distance between student figure 1. refinement of humble theory of learning instantiated in interactive learning assessments v. svihla 38 | f l r experience and professional role also affected feedback given to students. for high school students, it was hard to design feedback that did not seem schoolish, lowering the authenticity. in contrast, the university students found the opportunity to see an expert answer and compare it to their own answers to be an authentic learning activity. v. svihla 39 | f l r table 1. design-based iterations in the development of interactive learning assessments iteration participants and setting role taken by students and problem implementation main findings sample design decisions 1 biology students (n=34) at a rural southern us high school as a conservation geneticist, student advises developers on conservation of two bird species one case completed as think aloud task with researcher students learned about genetics and saw what they were learning as relevant increase scaffolding, create diagnostic yet authentic multiple choice questions 2 biology students (n=24) at a rural southern us high school as an intern genetic counselor, student counsels couple worried about potential for having a baby with sickle cell disease one case completed as think aloud task with researcher students didn’t understand what an internship was, but did learn genetics content from the case provide explicit guidance about internship 3 biology students (n=48) at an affluent west coast us suburban high school as an intern genetic counselor, student counsels couple worried about potential for having a baby with sickle cell disease one case completed in class session students who moved quickly from reading the problem to searching for information struggled; teacher unsure how/when to use case add generate ideas step and reflective prompts, make less linear; add teacher-asdesigner 4 undergraduate nutrition students (n=15) at a southwestern us research university as an intern dietitian, student counsels family about nutrition needs of child with down syndrome one case completed as an online assignment students learned and retained content; designing branches was burdensome for instructor replace branches with loops 5 graduate nutrition students (n=14) at a southwestern us research university as an intern dietitian, student counsels pregnant woman about gestational diabetes one case completed as an online assignment, one in-class discussion session students learned and retained content; instructor could design and teach with the case create more cases 6 undergraduate nutrition students (n=25) at a southwestern us research university as an intern dietitian, student counsels a range of clients on various nutrition topics seven cases completed in place of class meetings, plus seven in-class discussion sessions students learned and retained content; instructor developed more student-centered practice investigate ways to make branching design feasible v. svihla 40 | f l r 3. extensibility of dbr: design-based implementation research (dbir) in the earlier example of ilas, the initial goal was to help bring about statewide systemic change by providing a new way to embed assessment within learning. this driver necessitated changes to traditional dbr. when we changed settings, we also changed the role of the instructors from informants and consumers to designers of cases; this shift reflected our goal to help bring about smaller scale yet systematic change within a university program. in the first set of high school iterations, instructors were uncertain about how to use the cases. in the first iterations in the university setting in which the cases were designed by instructors, the cases were treated as homework, supplemental to in-class lectures. in the most recent iteration, the same instructors replaced lectures with the cases and further supplemented them with discussion (mckay et al., 2014). what we first viewed as a better assessment tool evolved into a tool for instructors to test their ideas about learning, resulting in more learner-centered teaching. 3.1. design-based implementation research recently, others have similarly sought ways to expand the reach of dbr, such as through ―implementation paths‖ that could lead the way to scaling a design (bielaczyc, 2013), seeking to develop learning theory that can be adapted to contexts (barab & squire, 2004), and design-based implementation research (dbir, fishman, penuel, allen, cheng, & sabelli, 2013; penuel & fishman, 2012). dbir includes ―(a) a focus on persistent problems of practice from multiple stakeholders’ perspectives; (b) a commitment to iterative, collaborative design; (c) a concern with developing theory related to both classroom learning and implementation through systematic inquiry; and (d) a concern with developing capacity for sustaining change in systems‖ (penuel, fishman, haugan cheng, & sabelli, 2011, p. 331). in one example of dbir, researchers partnered with four districts to develop a theory of action around improving mathematics instruction (cobb, jackson, smith, sorum, & henrick, 2013); the partnership lasted four years through cycles of data collection and analysis focused on the strategies as implemented. in each cycle, they documented the intended strategies, recorded how they were actually enacted, and made recommendations based on analysis. in order to support and maintain the relationship between researchers and practitioners, the team used two means of data collection and analysis: first, they prioritized providing usable evidence for the districts to evaluate the impact of their policies; second, they iteratively tested their theory of action to refine it. in addition to being guided by and refining a theory of action, they created an interpretive framework; this tool was used to evaluate and guide design decisions prior to, during and after implementation. following the four cycles of implementation, they began retrospective analysis to further test and refine their theory of action. this example highlights many parallels with dbr, including collaborative and contextual work with a focus on refining design and theory through iterative refinement and retrospective analysis. it also highlights the different scale at which dbir is conducted, involving many districts, schools, and classrooms, and a focus on creating sustainable change. by working at this scale, the research is more easily generalizable; by testing conjectures across four districts, they were able to learn about strategies that were effective across districts given specific conditions. because the target of their design was tied to how districts could support improved mathematics instruction, they were able to identify and refine strategies that were ineffective. for instance, school leaders had been receiving contentindependent professional development to guide their feedback to mathematics teachers; however, this process uncovered that they were not able to distinguish between high and low quality enactments of the mathematics. by recommending school leaders instead receive content-based professional development, they were able to design a sustainable, lasting change. dbir researchers emphasize the practical nature of their work, from problem to design to theory (dolle, gomez, russell, & bryk, 2013). this approach takes a broader view of the context and attends to usability by jointly considering how to change larger entities or systems (e.g., school districts) and how to support their ability to sustainably adapt designs (penuel & fishman, 2012). dbir has only begun to be taken up, bringing focus on scalability and sustainability, while respecting teachers and avoiding trying to ―teacher-proof‖ the materials of reform, for instance, through productive adaptation. v. svihla 41 | f l r 3.2. productive adaptation one approach to dbir is in teachers’ productive adaptations of curricula; this means staying faithful to the original intent of the design, reproducing ―invariant principles‖ across sites while being responsive to local contexts (kirshner & polman, 2013). in particular, focusing on maintaining or increasing – rather than reducing – the complexity and students’ engagement can support productive adaptations (debarger, choppin, beauvineau, & moorthy, 2013). dialogic interactions between teachers and researchers can support productive adaptions (kirshner & polman, 2013), but deliberate support – and spaces – are needed to ensure these are frequent enough and sustained (donovan, snow, & daro, 2013). related to this, it is also important to attend to power relationships and ownership of problems of practice; researchers bring different cultural norms and may have status not afforded to practitioners. deliberately viewing this as a cultural exchange, in which researchers and practitioners can trade ideas, can mitigate these challenges (penuel, coburn, & gallagher, 2013). in some cases, district support for any particular program, professional development, or curriculum may be taken as another in a sequence of top-down mandates, and therefore meet with resistance at school sites (borko & klingner, 2013). this highlights the importance of attending to influences across levels in the system in which research is occurring. because of this systems approach, not all dbir research occurs within schools or formal settings; though less common, dbir research has been conducted in communities, as a means to identify issues that might prevent youth from being successful and address them in creative, cross-institutional ways (mclaughlin & london, 2013). such approaches are important because dbr has been critiqued for not sufficiently attending to equity and social justice (confrey, 2005), though some work has sought this out (e.g., barab, dodge, thomas, jackson, & tuzun, 2007). 4. are dbr and dbir designerly? although design-based, not all dbr and dbir appear to be designerly (cross, 2001), explicitly applying design process by seeking needs, optimizing the design, and evaluating a solution in light of identified needs (edelson, 2002). because the targets of dbr are designs for learning and theories of learning, potential needs may be found both in review of research and in the world. needs are sometimes implicit and the design process left to the reader’s imagination (e.g., ―the tool was designed to scaffold learning of argumentation‖). aiming at scalability can strip the contextualist, designerly aspects from dbr, but committing to novel usability—and therefore a focus on context-can mitigate this. dbir focuses on design at scale, which would suggest a less designerly approach; yet, the emphasis on working in partnership with practitioners to support sustained change has helped focus dbir research on worldly needs. as these methods continue to evolve and incorporate bigger systems and big data, there are many opportunities for looking across streams of related data, such as logfiles and videos. these offer ways to evaluate the influence and refinement of designs for learning and of learning theories that are contextual and adaptive to the systems in which they reside. 4.1 credibility of design-based research concerns have been raised previously about the credibility of educational research in general (levin & o'donnell, 1999; national research council, 2002), urging researchers to employ methodologies influenced by medical research. in such approaches, tests of efficacy (whether the treatment works under optimal conditions) and effectiveness (whether the treatment works under real world conditions) ―are often conflated‖ (sloane, 2008, p. 625). influenced by this, discussions about dbr have focused on robustness, rigor and validity, grounded in experimental perspectives, an odd choice given the contextual, qualitative work that is commonly done with dbr. however, trustworthiness and credibility – as applied in qualitative methods – have also been considered (barab & squire, 2004), resulting in other ways to evaluate dbr: methodological alignment means the ―research methods we use actually test what we think they are testing‖ (hoadley, 2004, p. 203). edelson holds that dbr should not be evaluated by the same standards as v. svihla 42 | f l r traditional approaches because the goals differ; instead, ―novelty and usefulness‖ of the theory developed should be applied (2002, p. 118). 4.2 new types of data with the increasing popularity of big data and the relatively common use of technology-enhanced learning, some have included these new types of data in dbr studies. for instance, complex statistical modeling has recently taken the traditional place of qualitative approaches (markauskaite, 2010; markauskaite & reimann, 2008), arguing this approach avoids selection and confirmation bias. others remain skeptical about finding usable evidence of learning from big data, citing examples of contextual, interactional ―in-room‖ events that are not logged automatically (stevens, 2013); such events may explain successes and failures of designs in important ways. as an example, a long period of activity on a logfile might indicate a range of activities: a student spending a long time diligently reading the screen; a student absent from the activity, wandering the class out of boredom; or a teacher interaction in response to a reflective question by the student. these tell us very different things about how the design is or is not supporting learning, and do not, on average, provide useful design information. to deal with this issue, others rely on a combination of video and logfiles. for instance, researchers first analyzed classroom and video data to redesign a feedback feature that students rarely used (segedy, kinnebrew, & biswas, 2012). they then analyzed data from students’ interactions with the technology using hidden markov modeling, to evaluate the impact of their design decisions, leading to further refinement of both the design and theory guiding their work. similarly, in our research, we have leveraged data from logfiles, field notes, student performance, and videos of implementations to test and inform design decisions (svihla & linn, 2012a, 2012b). for instance, based on review of video and logfiles and student performance, we chose to add a step to an instructional unit to support students to interpret interactive visualizations, but we feared students might use a guess-and-check approach as a result. by examining logfiles, we found that most students revisited an earlier step seeking information, rather than guessing. this led us to more closely examine logfiles for particular patterns of activities, such as revisiting steps from earlier activities. though the primary theory guiding that work was well developed, the instantiation of it in the particular context and for the particular curricular goals was not, resulting in a much more humble, localized version that incorporated new ideas about how students revisit prior curricula to support their learning. keypoints design-based research (dbr) is the core methodology of the learning sciences the purpose of dbr is to develop designs for learning and learning theory through iterative refinement and retrospective analysis typically, dbr involves qualitative data; recently, some researchers have begun using ―big data‖ to make design refinements and build theory design-based implementation research (dbir) is a recent trend that involves scaling dbr to support change in larger systems, such as school districts acknowledgments the author would like to acknowledge support from the usda/nifa hispanic-serving institutions (hsi) education grants program (#2012-38422-19836). v. svihla 43 | f l r references barab, s. a., dodge, t., thomas, m. k., jackson, c., & tuzun, h. (2007). our designs and the social agendas they carry. the journal of the learning sciences, 16(2), 263-305. doi: 10.1080/10508400701193713 barab, s. a., & squire, k. (2004). design-based research: putting a stake in the ground. journal of the learning sciences, 13(1), 1-14. doi: 10.1207/s15327809jls1301_1 bielaczyc, k. (2013). informing design research: learning from teachers' designs of social infrastructure. journal of the learning sciences, 22(2), 258-311. doi: 10.1080/10508406.2012.691925 borko, h., & klingner, j. (2013). supporting teachers in schools to improve their instructional practice. national society for the study of education yearbook, 112(2), 274-297. brown, a. l. (1992). design experiments: theoretical and methodological challenges in creating complex interventions in classroom settings. the journal of the learning sciences, 2(2), 141-178. doi: 10.1207/s15327809jls0202_2 cobb, p., confrey, j., disessa, a. a., lehrer, r., & schauble, l. (2003). design experiments in educational research. educational researcher, 32(1), 9-13. doi: 10.3102/0013189x032001009 cobb, p., jackson, k., smith, t., sorum, m., & henrick, e. (2013). design research with educational systems: investigating and supporting improvements in the quality of mathematics teaching and learning at scale. national society for the study of education yearbook, 112(2). collins, a. (1992). toward a design science of education. in e. scanlon & t. o’shea (eds.), new directions in educational technology (pp. 15-22). berlin: springer-verlag. collins, a., joseph, d., & bielaczyc, k. (2004). design research: theoretical and methodological issues. journal of the learning sciences, 13(1), 15-42. doi: 10.1207/s15327809jls1301_2 confrey, j. (2005). the evolution of design studies as methodology. the cambridge handbook of the learning sciences, 135-151. cross, n. (2001). designerly ways of knowing: design discipline versus design science. design issues, 17(3), 49-55. doi: 10.1162/074793601750357196 debarger, a. h., choppin, j., beauvineau, y., & moorthy, s. (2013). designing for productive adaptations of curriculum interventions. national society for the study of education yearbook, 112(2). dede, c. (2004). if design-based research is the answer, what is the question? a commentary on collins, joseph, and bielaczyc; disessa and cobb; and fishman, marx, blumenthal, krajcik, and soloway in the jls special issue on design-based research. journal of the learning sciences, 13(1), 105-114. doi: 10.1207/s15327809jls1301_5 disessa, a. a., & cobb, p. (2004). ontological innovation and the role of theory in design experiments. journal of the learning sciences, 13(1), 77-10327. doi: 10.1207/s15327809jls1301_4 dolle, j., gomez, l. m., russell, j., & bryk, a. s. (2013). more than a network: building professional communities for educational improvement. national society for the study of education yearbook, 112(2), 443-463. donovan, m. s., snow, c., & daro, p. (2013). the serp approach to problem-solving research, development, and implementation. national society for the study of education yearbook, 112(2), 400-425. edelson, d. (2002). design research: what we learn when we engage in design. journal of the learning sciences, 11(1), 105-121. doi: 10.1207/s15327809jls1101_4 fishman, b., penuel, w. r., allen, a., cheng, b. h., & sabelli, n. (2013). design-based implementation research: an emerging model for transforming the relationship of research and practice. national society for the study of education yearbook, 112(2), 136-156. hoadley, c. m. (2004). methodological alignment in design-based research. educational psychologist, 39(4), 203-212. doi: 10.1207/s15326985ep3904_2 kirshner, b., & polman, j. l. (2013). adaptation by design: a context-sensitive, dialogic approach to interventions. national society for the study of education yearbook, 112(2), 215-236. levin, j. r., & o'donnell, a. m. (1999). what to do about educational research's credibility gaps? issues in education, 5(2), 177-229. doi: 10.1016/s1080-9724(00)00025-2 v. svihla 44 | f l r markauskaite, l. (2010). digital media, technologies and scholarship: some shapes of eresearch in educational inquiry. the australian educational researcher, 37(4), 79-101. doi: 10.1007/bf03216938 markauskaite, l., & reimann, p. (2008). enhancing and scaling-up design-based research: the potential of e-research. paper presented at the proceedings of the 8th international conference on international conference for the learning sciences-volume 2. mckay, t., cantarero, a., svihla, v., yakes jimenez, e., & castillo, t. (2014, june 23-27). becoming a professional through virtual practice. paper presented at the 11th international conference of the learning sciences (icls2014), boulder, co. mclaughlin, m., & london, r. a. (2013). taking a societal sector perspective on youth learning and development. national society for the study of education yearbook, 112(2), 192-214. national research council. (2002). scientific research in education. washington, dc: the national academies press. o'neill, d. k. (2012). designs that fly: what the history of aeronautics tells us about the future of designbased research in education. international journal of research & method in education, 35(2), 119140. doi: 10.1080/1743727x.2012.683573 penuel, w. r., coburn, c. e., & gallagher, d. j. (2013). negotiating problems of practice in research– practice design partnerships. national society for the study of education yearbook, 112(2), 237255. penuel, w. r., & fishman, b. j. (2012). large‐ scale science education intervention research we can use. journal of research in science teaching. doi: 10.3102/0013189x11421826 penuel, w. r., fishman, b. j., haugan cheng, b., & sabelli, n. (2011). organizing research and development at the intersection of learning, implementation, and design. educational researcher, 40(7), 331-337. doi: 10.3102/0013189x11421826 phillips, r., gawel, d. j., svihla, v., brown, m., vye, n. j., & bransford, j. d. (2009). new technology supports for authentic science inquiry, practice, and assessment in the classroom. paper presented at the aera, san diego. reeves, t. c. (2006). design research from a technology perspective. educational design research, 1(3), 5266. sandoval, w. a. (2004). developing learning theory by refining conjectures embodied in educational designs. educational psychologist, 39(4), 213-223. doi: 10.1207/s15326985ep3904_3 segedy, j., kinnebrew, j., & biswas, g. (2012). supporting student learning using conversational agents in a teachable agent environment. in j. van aalst, k. thompson, m. j. jacobson & p. reimann (eds.), the future of learning: proceedings of the 10th international conference of the learning sciences (icls 2012) – volume 2, short papers, symposia, and abstracts (pp. 251-255). sydney, australia: isls. sloane, f. c. (2008). randomized trials in mathematics education: recalibrating the proposed high watermark. educational researcher, 37(9), 624-630. doi: 10.3102/0013189x08328879 stevens, r. (2013, 6/12-6/14). big data, interaction analysis, and everything in between. paper presented at the games, learning, society 9.0, madison, wi. svihla, v., gawel, d. j., brown, m., moore, a., vye, n. j., & bransford, j. d. (2010). 21st century assessment: redesigning to optimize learning. in k. gomez, l. lyons & j. radinsky (eds.), learning in the disciplines: proceedings of the 9th international conference of the learning sciences (icls) (vol. 2, pp. 474-475). chicago, il: international society of the learning sciences. svihla, v., & linn, m. c. (2012a). a design-based approach to fostering understanding of global climate change. international journal of science education, 34(5), 651-676. doi: 10.1080/09500693.2011.597453 svihla, v., & linn, m. c. (2012b). distributing practice: challenges and opportunities for inquiry learning. in j. van aalst, k. thompson, m. j. jacobson & p. reimann (eds.), the future of learning: proceedings of the 10th international conference of the learning sciences (icls 2012) – volume 1, full papers (pp. 371-378). sydney, australia: isls. svihla, v., phillips, r., gawel, d. j., vye, n. j., brown, m., & bransford, j. d. (2009). a tool for 21st century learning and assessment. in a. dimitracopoulou, c. o'malley, d. suthers & p. reimann v. svihla 45 | f l r (eds.), cscl practices: proceedings of the 8th international conference on computer supported collaborative learning (cscl 09) (vol. 2, pp. 46-48). rhodes, greece: international society of the learning sciences. svihla, v., vye, n. j., brown, m., phillips, r., gawel, d. j., & bransford, j. d. (2009). interactive learning assessments for the 21st century education canada, 49(3), 44-47. svihla, v., yakes, e., castillo, t., cantarero, a., valdez, i., & dominguez, n. (2013). interactive learning assessment: providing context and simulating professional practices proceedings of games, learning, society 9. the design-based research collective. (2003). design-based research: an emerging paradigm for educational inquiry. educational researcher, 32(1), 5–8. doi: 10.3102/0013189x032001005 yakes, e., cantarero, a., mckay, t., svihla, v., castillo, t., valdez, i., & hertel, j. (2013). interactive learning assessment: simulating professional practices. nacta journal, 57(supplement 1). frontline learning research 5 special issue „learning through networks‟ (2014) 4 14 issn 2295-3159 corresponding author: nino pataraia, 58 port dundas road, g4 0hg, glasgow, uk, nino.pataraia@gcu.ac.uk doi: http://dx.doi.org/10.14786/flr.v2i2.89 4 | f l r ‘who do you talk to about your teaching?’: networking activities among university teachers nino pataraia a , isobel falconer a , anoush margaryan a , allison littlejohn a , sally fincher b a caledonian academy, glasgow caledonian university, glasgow, scotland, uk b university of kent, kent, england, uk nino pataraia, 58 port dundas road, g4 0hg, glasgow, uk, nino.pataraia@gcu.ac.uk article received 15 february 2014 / revised 26 april 2014 / accepted 28 june 2014 / available online 15 july 2014 abstract as the higher education environment changes, there are calls for university teachers to change and enhance their teaching practices to match. networking practices are known to be deeply implicated in studies of change and diffusion of innovation, yet academics’ networking activities in relation to teaching have been little studied. this paper extends the current limited understanding, building on roxå and mårtensson’s work (2009) and extending it from sweden to the uk and usa. it is based on two separate studies, one from the share project led by the university of kent, and one from glasgow caledonian university, exploring the composition of personal networks, and the characteristics of interactions in order to understand the networking practices which may support change of teaching practice. we conclude that academics’ personal teaching networks are mainly discipline-specific and strongly localised. this contrasts with the research networks found by becher and trowler (2001) and may reduce innovation, although about half the respondents also had external contacts that might support creativity. keywords: networks; interactions; conversational partners; higher education; academics n. pataraia et al. 5 | f l r 1 background as the higher education environment changes, there are calls for university teachers to change and enhance their teaching practices to match (e.g. european commission, 2009). if in the past learning, adult education and professional development were largely associated with formal education and training (kyndt, dochy, & nijs, 2009; tynjälä, 2008), nowadays it is becoming recognised that learning is lifewide and can take place at work or elsewhere (skule, 2004). furthermore, education scholars argue that teaching knowledge is frequently experientially acquired, and change in teaching occurs through adoption and adaptation of new practices learnt about informally (eraut; 1994; 2004; knight, 2006). thomson (2013) argues that “academics are able to learn about teaching through informal conversation, and for some issues, and even individuals, it may be a more appropriate means for learning about teaching than formal academic development” (p. 205). despite the fact that the significance of informal aspects of academics‟ learning about teaching is becoming recognised, there is still little insight into how and when academics engage in informal learning for enhancing their practice (thomson, 2013). given that a network represents a locus for informal interactions, offering a medium for the exchange of resources and experience, capacity building and collaborative development of knowledge (koper, rusman, & sloep, 2005; powell, koput, & smith-doerr, 1996; tynjälä & nikkanen 2009), academics‟ interactions about teaching are grounded and discussed in the context of networks. a network comprises a set of actors (“nodes”) and a set of relations (“ties” or “edges”), between the nodes (wasserman & faust, 1994). common objectives for interaction bring network participants together (paavola, lipponen, & hakkarainen, 2002). network members may be connected either directly or indirectly, and their connections can be either informal (trust-based), or formalized through contracts. the ties may comprise flows of various types, such as flows of information, materials, financial resources, services, and social support (monge & contractor, 2003). granovetter (1973) differentiated between strong and weak ties, describing strong ties in terms of the time and emotions invested in the relationship. examples of strong ties include friendship and familial relationships, which facilitate the transfer of tacit, sensitive and complex knowledge (burt, 1992; reagans & mcevily, 2003). weak ties, by contrast, encompass a more restrained investment of time and intimacy. granovetter suggested that weak ties serve as bridges between otherwise disconnected social groups and are more important in disseminating new, non-redundant information and resources than strong ties. in order to understand different properties of networks, it is useful to draw on a range of network theories. homophily and proximity theories are particularly important for scrutinising and interpreting the likelihood of establishing and/or dissolving network ties. according to proximity theory (monge & contractor, 2003, p.303), “people communicate most frequently with those to whom they are physically closest and proximity increases the opportunities for individuals to observe and learn more about one another, thereby creating conditions favourable for the development of communication ties”. rogers (2003) asserted that communication is usually most effective between individuals who are similar, or homophilous, in some respect. proximity theory implies that those who are physically close and communicate frequently are more likely to become homophilous, thus leading to the development of rogers‟s (2003) conditions for effective communication. nevertheless, recent technological developments have greatly affected the spatial and social structure of groups, communities and other entities, offering easy access to new information/knowledge/resources and sustainment of communication ties (wellman, 2001). the advent of ubiquitous virtual networking raises the question of whether the concept of proximity is still relevant. while homophily and proximity theories are useful for understanding the formation of network ties, social capital theory helps to evaluate the value of social networks. social capital theory explicates that individuals invest in forming social relationships in order to acquire access to rich resources, namely emotional and professional support, expertise, valuable new connections, and different type of capital (knowledge, human, social and learning) (wenger, trayner, & de laat, 2011). previous research has emphasised the importance of networking, along with other forms of social exchange, for both individual and organisational learning (katz, earl, & jaffar, 2009; trinkle, 2009; tynjälä, 2008). scholars have concluded that networks facilitate dissemination of good teaching practices (coburn & russell, 2008). engagement in networks offers new ways of thinking about educational quality and enhances teachers‟ knowledge, potentially altering their thinking and classroom practice (hargreaves, 2003). furthermore, networks have been recognised as a key instrument for sustained teacher learning and n. pataraia et al. 6 | f l r professional development (katz et al., 2009). through networking, individuals form and maintain useful relationships with others who can, potentially, provide work-related support (forret & dougherty, 2004). in addition, networks equip teachers with a sense of empowerment, provide emotional support, and encourage engagement in teaching (baker-doyle, 2011). nevertheless, it is worth highlighting that these arguments have been largely derived from research in school teaching contexts. pioneering investigations of educational networks have primarily focused either on teachers‟ learning in the context of secondary education (mccormick et al., 2011) or on academics‟ research and departmental networks (becher & trowler, 2001; pifer, 2010). for example, mccormick et al. (2011) examined the role of networks in school teachers‟ learning, suggesting that application of network theories would lead to a better understanding of educational networks. several studies have documented that informal interactions contribute to enhancement of teaching practice (schuck, aubusson, & buchanan, 2008; thomson, 2013). however, many of these studies have examined centrally-organised, formal networks stressing network coordinators‟ viewpoints on the overall value of networks for teachers‟ professional development (kerr et al., 2003). therefore, there is still little insight into what role personal networks play in supporting the professional development of teachers (baker-doyle, 2011). it is worth emphasising that there is even less understanding of personal networks at he teacher level. hence, this paper aims to extend the limited understanding of academics‟ teaching networks by focusing on personal, egocentric networks where, “the network is perceived by the individual at its centre” (wellman, 1998, p.19). we explore the composition of academics‟ personal networks, and also the nature, frequency, venue and characteristics of interactions in order to understand how academics‟ networks may support learning and change of teaching practice. furthermore, coburn & russell (2008) have emphasised that previous studies have ignored the content of teachers‟ interactions. this research responds to this call by investigating themes of participating academics‟ interactions. a number of previous studies examined academics‟ self-initiated networks. most notably, pifer (2010) explored the networking behaviour of academics in the us universities. she found that academics relied on their departmental colleagues for instruction, mentoring, professional opportunities, support with writing grant application and publications, and general support and friendship. pifer‟s study showed that “departmental characteristics, such as proximity, disciplinary influence, and the culture of the department, appeared to influence the interactions of academics” and academics tended to “cultivate relationships and exchange resources with colleagues they perceived to be like them, and less likely to interact with colleagues they perceived to be different from them” (ibid, p. 227–230). however, pifer‟s work focused solely on networks within single departments. she identified the need for further research into other types of academic networks. this paper addresses this gap by examining relationships both within and beyond the department. similarly, roxå and mårtensson (2009) investigated academics‟ networks in a swedish university. drawing on a socio-cultural perspective, they explored the conversations that teachers have with their colleagues. they presumed that some of these conversations could have an influence on teachers to develop new understanding of teaching or even significantly alter their conceptions of teaching. to test the reliability of their assumption, they asked 106 faculty members in sweden from a range of disciplines to reflect on their conversations about teaching. they discovered that “academics relied on a network of a few significant others as they constructed, maintained, or changed their understanding of the teaching and learning reality” (2009, p. 214). on average, participants reported ten conversational partners, which accords with becher and trowler‟s observations of the smaller research network (2001). furthermore, their research revealed that although the participants found their conversational partners anywhere in the same or other departments, disciplines, or institutions, or outside academia the proportion of conversational partners was higher within the department than in other locations. this study extends the work of roxå and mårtensson (2009) by examining a wider range of aspects of conversations about teaching within networks. the research is guided by the following research questions: who do academics talk to about their teaching? what are the main themes (content) of academics‟ conversations? with what frequency and where do academics‟ conversations take place? what factors motivate academics to network and what value do they perceive in their interactions? n. pataraia et al. 7 | f l r the data presented in this paper were derived from two interrelated studies from the pilot phase: the share project longitudinal study 1 and the “academics‟ networking practices” (anp) project 2 at glasgow caledonian university. the share project, at the university of kent, comprised a number of separate studies, which broadly aimed to investigate with whom academics discuss their teaching practice. more precisely, the share project longitudinal study was concerned with the exploration of the setting, nature and value of academics‟ interactions related to teaching. examination of these topics informed the anp project in terms of its methodological approach and research objectives. the overarching aim of the anp project was to uncover further how social interactions and the structure of personal networks influence academics' learning, affecting their behaviour and supporting change in teaching practice. 2 methodology we applied the analytical method of social network analysis (sna). this method is specifically designed to examine the patterns, causes and consequences of established relationships between different individuals (scott & carrington, 2011). however, sna falls short of revealing the motivation behind individuals‟ actions within their networks. since several authors have recommended application of different forms of data collection for breadth and depth of understanding and also for corroboration of network processes (kilduff & tsai, 2007; mehra, kilduff, & brass, 1998), this study integrated both quantitative and qualitative approaches. 2.1 study 1: share project longitudinal study as part of a more extensive questionnaire, longitudinal study of 18 academics in computing, mathematics and technology subjects, 14 provided a free-text written response regarding their teachingrelated interactions. study 1 drew on convenience sampling. the response rate was 83%. two explicit inclusion criteria were used: 1. potential study participants had to be teaching in math/computing/technology area, and 2. participants would be eager to participate in two interventions a year over a period of three years. within the mathematics/computing/technology constraint, they were chosen to represent a variety of institutional contexts, experience and reputation for innovation. the given study examined the composition of academics‟ teaching networks along with the frequency, nature and content of interactions. 2.2 study 2: semi-structured interviews within the anp project to probe findings from the share project further, eleven academics representing three institutions and five disciplines, namely engineering-2/11; life sciences-4/11; education-2/11; social sciences-1/11; humanities-2/11, were interviewed. interviewees for study 2 were drawn using convenience sampling. the response rate was 100%. the main criterion for the selection was that the potential interviewee had to be an innovative/excellent teacher. the interviews lasted one to one and a half hours and were audio recorded and transcribed. interview protocol and interview questions can be accessed at: https://drive.google.com/file/d/0b2to0roh4ibxxzrsawz4ckj5nuk/edit?usp=sharing. 2.3 data analysis 2.3.1 studies 1 and 2 the same techniques of analysis were applied to the 14 written responses from study 1 and 11 interview transcripts from study 2. descriptive statistics, using spss software, focused on describing the characteristics of the sample along with the number of contact types/categories, the frequency and themes of interaction about teaching. given that variables of interest were categorical (qualitative) by nature, frequencies were utilised to obtain descriptive statistics (pallant, 2010). the research questions were used to define initial coding classes for written data and further classes were created as themes emerged. emergent classes were developed by two independent researchers, then compared and contrasted. checks for consistency and reliability were carried out and the final list of codes was refined. overall, eight classes were created: 1. contact category (table 1); 2. level of experience; 3. disciplinary affiliation; 4. frequency of 1 http://www.sharingpractice.ac.uk/homepage.html 2 http://www.gcu.ac.uk/networkedinnovation/ https://drive.google.com/file/d/0b2to0roh4ibxxzrsawz4ckj5nuk/edit?usp=sharing http://www.sharingpractice.ac.uk/homepage.html http://www.gcu.ac.uk/networkedinnovation/ n. pataraia et al. 8 | f l r interaction; 5. venue of interaction; 6. nature of interaction; 7. preferred method of interaction; 8. content of interaction. the purpose of these thematic categories was to organise data into meaningful units of analysis. table 1 shows the different categories of contacts enumerated by academics. table 2 outlines the categories within the classes frequency, nature and themes of conversations. table 1 contact categories (top row) and types within each category ‘family’ ‘in department’ ‘in institution’ ‘friends’ ‘elsewhere’ family member, profession not specified family member teaching family member non-teaching departmental colleague, role not specified colleagues teaching same or companion modules support staff students: current academics teaching in other departments, same institution (discipline not specified) academics teaching in other, but related disciplines/departments central support staff friends, profession not specified friends teaching friends nonteaching professional relationships outside the institution, role not specified formal relationships; collaborations (i.e., co-authors) non-academic relations former colleagues students: former and prospective table 2 categories within the frequency, nature of conversation, and theme classes frequency of interactions the nature of conversations themes ns-not specified=0 once a term-yearly=1 fortnightly-several times per term=2 2 weekly-fortnightly=3 daily=4 ns-not specified=0 formal=1 informal=2 unspecified=0 learning, curriculum design; projects for students=1 students experience/progress=2 research and developing teaching=3 approach to teaching=4 feedback to students/students’ assessment=5 tips and ideas for teaching=6 problems with students=7 n. pataraia et al. 9 | f l r administration/management=8 concerns with institutional environment=9 other=10 interview data were classified, summarized and visualized using nvivo 9. initially, interview transcripts were read to uncover the key themes; subsequently, data were broken down into discrete parts, closely examined, and compared for similarities and differences. from content analysis, themes, such as contact categories; nature, content, intensity and venue of interactions; motivating factors, and also the value of networking, emerged (babbie, 2007). 3 results and discussion the presentation of results is structured around our four research questions. 3.1 research question 1: who do academics talk to about their teaching? in order to understand the configuration and composition of academics‟ teaching networks, information about teaching-related interactions was gathered. each participant was free to name as many contacts as they wished, located across different settings and representing diverse categories of relationships, namely department/institutional/external colleagues, friends and/or family members. it is worth mentioning that each academic could identify more than one contact type under each category (e.g. “staff directly involved in the course i teach” and “postgraduates who teach”; these two different types of contact would still appear “in department” category). since no boundaries were predefined and also no temporal or numerical constraints were introduced for capturing academics‟ significant teaching-related interactions, we presume that enumerated contacts represent members of participants‟ personal networks rather than of their tightly-knit communities. results revealed that academics discuss their teaching with diverse types of contact. however, when asked “who do you talk to about your teaching”, participants tended to name departmental colleagues first before mentioning other types of connection. interviewee 5, specialising in life science, emphasised that “everything i do, i discuss with others, here, in the departmental level”. similarly interviewee 8, representing life sciences, highlighted having close interactions with the departmental programme team while designing new, or amending old, courses. overall, the majority of teaching-related contact types fell under the category of department. “elsewhere” and “institution” represented the second and the third most frequently quoted categories, followed by “family” and “friends”. only two out of eleven academics from study 2 prioritised interactions outside their institution. these two a-typical cases were experienced teachers from the discipline of education. it has to be noted that some respondents identified individual contacts (eg. “the director of teaching”), while others named only types of contact (eg. “other instructors in my department”), giving no precise idea how many individuals within each contact type they talk to. therefore, analysis is at the level of contact type, rather than individuals. findings from this research suggest that common interests, namely joint projects, goals, problems, mutual commitment (“we actually sit on the same committee, we teach on the same course, we are on the same project”), trust and good personal relations played an essential role in cultivating and maintaining connections with others, encouraging open discussions and idea exchange in regards to teaching. since network studies normally rely on a simple name generator question, such as “who do you talk to about specific topic”, data derived from these two studies were sufficient to capture participants‟ contacts distributed across diverse settings, determining the configuration and the basic size of academics‟ teaching network. in sum, findings suggest that academics‟ interactions are concentrated in, but not confined to, departments, spreading more weakly across and outside academia. this observation is in line with roxå and mårtensson‟s finding in sweden that, “academics‟ conversational partners could be found anywhere: within their discipline, in other universities or outside academia” and with their diagram showing a higher proportion of contacts within the department (2009, p. 551, diagram on page 552). as mentioned above, there were only two interviewees who had teaching networks that focused strongly outside their department and institution. it seems likely that their teaching and research networks were inseparable and shared the n. pataraia et al. 10 | f l r characteristics of research networks comprised of dispersed contacts (becher & trowler, 2001). given that respondents tended to list informal interactions first and in greater numbers than formal, it can be presumed that they attach greater significance to the informal. this concurs with roxå and mårtensson‟s (2009) finding for teaching networks in sweden, and bears out knight (2006), thomson (2013) and eraut‟s (1994; 2004) claims that teachers‟ learning is informal. the small significant research networks observed by becher and trowler (2001) were also informal. although our data did not measure the absolute size of respondents‟ teaching networks, the indications are that they were small, sparse and simultaneously informal. 3.2 research question 2: what are the main themes of academics’ conversations about teaching? this research sheds light on the content of academics‟ interactions, examining the flow of different types of resources, advice, information and support within personal networks. data revealed that conversations about teaching varied in terms of their content across diverse types of contact. table 3 illustrates the themes discussed across the five contact categories: table 3 distribution of themes discussed across five categories of contacts themes discussed with different categories of contacts ‘family’ 'in department' ‘in institution’ ‘friends’ ‘elsewhere’ learning, curriculum design; projects for students 1 11 3 0 6 students experience/progress 2 6 3 0 3 research and developing teaching 0 1 1 0 1 approach to teaching 1 4 3 1 1 feedback to students/students‟ assessment 0 7 0 1 4 tips and ideas for teaching 1 2 2 0 1 problems with students 2 5 2 1 1 administration/management 1 9 2 0 3 concerns with institutional environment 2 2 0 0 0 other 3 5 4 1 3 student-related issues and concerns with the institutional environment formed a high proportion of conversations with family “... content of modules, how things are going, irritating admin regulations, marking woes, and odd events”. in addition, family members offered emotional support: “[my wife] provides a valuable balance that helps me to deal with tough situations. it‟s not really directly to do with teaching, but it is certainly a huge help with part of my job”. inside their departments, respondents discussed a wider variety of themes, as detailed in table 3. problems, concerns about students and administrative issues were discussed with administrative staff (five respondents) and people who provided teaching support (two respondents). whereas, curriculum design, projects for students and approaches to teaching were discussed with people whose teaching participants supervised (seven respondents). students‟ experience, progress, feedback, assessment and problems, were discussed with current students, mainly during tutorials and classes (seven respondents). one participant indicated that students‟ opinion was “invariably good source of feedback, insight into teaching practices”. beyond the department, but within the institution, conversations were not discipline-specific. general pedagogical approaches, assessment tools, curriculum design and problems associated with students were discussed with academics from other departments. the conversations with institutional colleagues occurred in a formal setting, normally during seminars and training events. interactions with people from support departments addressed educational research and development of teaching, students‟ experience, administration/management and use of technology (five respondents). three respondents discussed approaches to teaching, students‟ issues, assessment and feedback with their friends. while some stated sharing and testing new teaching ideas or seeking advice for teaching-related challenges from friends, others n. pataraia et al. 11 | f l r specified that their conversations with friends were general and entailed sharing funny stories. with colleagues from other institutions, academics compared and contrasted their professional and teaching environments and discussed prospects for collaborations. the external colleagues tended to be either from the same discipline or at least share similar research interests. course content, teaching approaches, learning process and students, in particular their changing expectations, progress, and issues, were the main themes of conversations. overall, the depth of conversations varied across different contacts, yet appearing more comprehensive with departmental colleagues in comparison with peers from other departments or institutions. 3.3 research question 3: with what frequency and where do academics’ conversations take place? results indicated variations between participants in terms of the regularity of their interactions about teaching. some engaged in task specific interactions, such as struggling with a particular aspect of teaching or designing a new course, while others took part in regular, informal talks around various aspects of their practice. for interviewee 4, specialising in social sciences, networking is a natural way of working and an integral part of her everyday professional life: “my whole practice is based on this idea of collaboration and networking, because it is how i work; you know, it‟s a personal preference, i am not a lone scholar”. in written responses, respondents specified the frequency of their conversations either in quantitative or qualitative terms for 72 out of 105 interactions. table 4 illustrates the distribution of frequencies across different contact categories where this was specified quantitatively, and table 5 shows the distribution where frequency was specified qualitatively. table 4 distribution of quantitatively specified frequencies (n=14) frequencies reported in quantitative terms contact types ‘family’ ‘in department’ ‘in institution’ ‘friends’ ‘elsewhere’ total once a term-yearly 0 6 1 1 3 11 fortnightly-several times per term 2 7 1 1 2 13 1/2 weekly-fortnightly 0 10 0 0 1 11 daily 0 1 0 0 0 1 table 5 distribution of qualitatively specified frequencies (n=14) frequencies specified in a qualitative way contact types ‘family’ ‘in department’ ‘in institution’ ‘friends’ ‘elsewhere’ total very occasionally 0 2 1 0 1 4 sometimes 3 9 3 2 4 21 frequently 3 2 0 0 0 5 when change is required 0 1 0 0 0 1 table 4 and 5 suggest that participants talk about their teaching with colleagues in the department regularly, half-weekly or several times per term. the content analysis of written responses and interview transcripts revealed that interactions about teaching were ad hoc, taking place during lunch and coffee breaks, and more frequently during the teaching term (verified by five respondents). a detailed analysis of written data revealed that within the department, participants talked most frequently with colleagues teaching the same or a companion module, or whose teaching they supervised, namely teaching assistants. interactions with colleagues teaching the same or a companion module were mainly face-to-face, spontaneous, casual in nature, and took place in common rooms or corridors. some selectivity was evident in interviewee 2‟s (humanities) statement that, “i‟ve got a couple of colleagues here i often talk to about teaching... so yes, a fair amount of, probably two or three people out of 40, ... they tend to be people you can n. pataraia et al. 12 | f l r talk to or you feel are on a same sort of wave length as you are,”. similarly, interviewee 10, specialising in engineering, mentioned talking with some colleagues far more frequently than with others. overall, participants emphasized talking more with those with whom they were on friendly terms. despite the fact that interactions with colleagues from other departments, from subject networks, other he institutions, industry or employers, were mentioned, the majority of interviewees indicated a lower frequency of such interactions, occurring occasionally, once a term-yearly basis: “maybe a couple of times a year, depending on if there‟s an event” (8/11). interactions with institutional colleagues occurred at universitywide events, mainly face-to-face, but email, phone, chat and online platforms, were used with physically distant colleagues. since the most frequent interactions were with departmental colleagues these are likely to be discipline-specific. proximity theory appears useful for interpreting the greater frequency of interactions about teaching, while the emphasis on discipline points to the evidence of homophily. this observation highlights that physical proximity still plays an influential role in activating and sustaining network ties, and also for developing trust and rapport with peers despite the widespread popularisation of technologies. if, following granovetter (1973), frequency of interaction is taken as a measure of strength of tie, then the study suggests that academics tend to have strong teaching ties with people within the department, and far weaker ties with people outside their institution. it appears that respondents rely mainly on close, localised connections when dealing with teaching matters. however, since some academics maintained weak ties, such contacts could represent a source of radically novel teaching ideas, bringing complementary knowledge to personal teaching networks (granovetter, 1973). finally, results point to the fact that not only the temporal component of interactions determines the strength of ties, but also the significance of a conversational partner (i.e., friendship). 3.4 research question 4: what factors motivate academics to network and what value do they perceive in their personal networks? in addition to exploring the composition and the basic size of networks along with the content, frequency, venue and nature of academics‟ interactions, this study expands understanding of the incentives for networking and the benefits obtained through personal teaching networks. table 6 summarises results across all of the interviewees: table 6 motivation for networking and the benefits obtained through personal networks motivation for networking benefits obtained through networks access to new teaching ideas good personal relationships access to disciplinary knowledge professional guidance access to new learning opportunities prompt feedback access to diverse resources solidarity and the sense of community access to professional and emotional support confidence findings suggest that personal networks provide not only access to new teaching ideas, learning opportunities and diverse resources, but also the exposure to diverse viewpoints and a wide pool of expertise within networks enriches academics‟ knowledge base and challenges their conceptions of teaching. through interactions, participants keep track of others‟ work, sometimes triggering their motivation to adopt or experiment with new things: “i find out what other people are doing; looking at what someone else is doing and then changing my teaching is one of the things that i would do” (interviewee 9). sometimes, interviewees adopted ideas without much alteration; at others they adapted new concepts to their own context, “you can take something that someone is using to teach in a particular context and you maybe like n. pataraia et al. 13 | f l r the idea, but it doesn‟t fit with your students or with what you teach. so what you could do is take that idea and you can change it until it does fit with your students.” furthermore, interviewees indicated that the network offered a sense of security, comfort and reliability. participants particularly valued availability of prompt feedback, especially when facing a specific teaching-related issue. by discussing problems with peers, academics could easily develop useful solutions. in addition, personal networks represented a locus for testing new ideas: “if you are planning some changes to your course, you‟ll often try it out on them first, before you go to the larger group, just to make sure you don‟t make a complete fool of yourself” (interviewee 2). in sum, findings showed that through personal networks academics acquire various kinds of resources (new ideas and teaching materials), share knowledge and experience with one another as speculated by social capital theory (wenger, trayner, & de laat, 2011). the interviewees appreciated these as benefits that provided incentives for networking. overall, respondents used their personal networks for exchanging ideas, discussing teaching-related problems and obtaining professional advice. academics‟ teaching networks thus conform with the network functions proposed by tynjälä and nikkanen (2009), koper et al. (2005) and paavola et al. (2002), as discussed in the introduction. 4 conclusion understanding academics‟ learning is important as in today‟s society lifelong learning is becoming the benchmark of all professional fields. given that academics are the key agents in transforming educational practices, scientific knowledge about from whom or how they learn and also in what ways their professional development can be supported is of key importance. this research specifically unpacks the interactions that influence and enhance academics‟ teaching practices, examining their networking in terms of its nature, processes and outcomes. this study can be of interest not only to the academics themselves, but also to the wider university staff, especially those who are responsible for professional development, and national bodies interested in teaching and learning (for instance, higher education academy and seda). the small size and variation of the sample limit the generalisability of the findings. nevertheless, some tentative conclusions are drawn below. these have implications for understanding the ways in which academics develop understanding of teaching, acquire new knowledge, skills and dispositions in regards to teaching, and also how change in instruction might be supported. nevertheless, further testing and verification of the results through additional empirical research are highly recommended. despite the fact that personal networks relating to teaching are valued by academics, in most cases these are strongly localised. there is little evidence of personal networks extending beyond immediate (faceto-face) contacts. even if other means were utilized to contact external colleagues, the ties were weaker, the intensity of interactions less frequent, the content of conversations less comprehensive, and generally considered less significant. two interpretations are possible for these observations. first, that teaching practice is a highly contextualised activity (in contrast to research), so meaningful interactions are likely to be with those who understand the local context, namely institutional regulations/politics, departmental culture, students – such people often share the same building, have mutual commitments and/or similar interests, and face to face contact is easy. second, that face to face contact could be the most effective way for sharing teaching practice and also for acquiring prompt feedback, hence significant interactions are likely to be with those who are geographically close. the data, though, may show some research bias: the prompt “who do you talk to about your teaching?” could have predisposed respondents to think in terms of face-toface interactions. investigation of the ways in which academics network about teaching through other media, could establish the circumstances under which face to face contact is significant in supporting changes in practice. the local focus implies densely connected networks where the majority of members know each other considerably well. tushman and anderson (1986) suggest that members of such networks are less exposed to radically new ideas and also less likely to absorb knowledge created elsewhere. nonaka and takeuchi (1995) agree, advocating being open to external resources and diverse sources of information to avert pressures for social conformity, and „not invented here‟ syndrome. however, ruef (2002) suggests that a diverse network may support creativity, through flow of information via weak ties, and adoption of the resultant innovation through strong ties. about half of respondents in this research had the diverse networks that might support effective innovation in teaching according to ruef‟s model. n. pataraia et al. 14 | f l r the majority of significant ties, for most respondents, appear to be with others from the same discipline, whether within the department or external to the institution. this implies that disciplinary networks may be more effective in supporting change than generalised intra-institutional networks. however, the share project respondents all came from computing, mathematics or technology, and this has heavily weighted the sample. further research should test whether this conclusion applies to disciplines with less technical content where teaching approaches may transfer more easily across disciplinary boundaries. that the teaching networking practices of those whose research discipline is education may be a-typical requires further investigation as it has implications for the applicability of any conclusions drawn from the study of such networks. academics‟ connections did not appear time or context specific, since respondents maintained contact both with current colleagues and with those from previous institutions. this implies a historical or temporal component of networks which are thus not entirely explained by proximity or discipline. moreover, there was a wide diversity in intensity of networking relations, but only within the department interactions appeared to be regular in nature. the dynamics of teaching network formation and maintenance, and the impact this has on the types of flow warrant further investigation. given that personal networks offered new teaching ideas, learning opportunities, diverse resources, and also shaped academics‟ perceptions about teaching, it can be presumed that personal networks play an influential role in academics‟ professional development. furthermore, since previous network literature has been dominated largely by quantitative research (filliettaz, 2011; rijt, bossche, & segers, 2012), this research project addresses the methodological gap by adding a much-needed qualitative perspective on academics‟ teaching-related interactions and network processes within their personal networks. by examining the depth of academics‟ interactions about teaching, this study addresses yet another gap concerning the content of interactions (coburn & russell, 2008). the current studies examined the static snapshot of participants‟ teaching networks. therefore, future studies should consider scrutinising how academics‟ network composition changes over time and what factors cause changes in their network structure. furthermore, the current research did not explore in what ways personal characteristics, namely age, gender, experience level, disciplinary domain or institutional culture influence academics‟ networking behaviours. hence, future research should consider investigating the influence of these characteristics on the patterns of networking. finally, this paper made partial use of social network analysis by elaborating on the basic structure of academics‟ networks, along with the frequency, content and the value of teaching-specific interactions. nevertheless, another research paper reports the detailed sna analysis on the project data, outlining the impact of ego, ego-alter and alter-alter characteristics on the patterns and nature of relationships formed by academics (pataraia et al., 2014). keypoints academics‟ teaching networks are localised, marked with strong ties. personal networks offer a wide range of benefits, namely new information, ideas and support. academics‟ personal connections do not appear time-bound. acknowledgements we are grateful to all participants in these two studies, and to the national teaching fellowship scheme, which funded the share project. n. pataraia et al. 15 | f l r references babbie, e. (2007). the practice of social research. belmont, ca: thomson higher education. baker-doyle, k. j. (2011). the networked teacher: how new teachers build social networks for professional support. new york, ny: teachers college press. becher, t., & trowler, p. (2001). academic tribes and territories: intellectual enquiry and the cultures of disciplines. buckingham, england: open university press. burt, r. s. (1992). structural holes: the social structure of competition. cambridge, ma: harvard university press. coburn, c. e., & russell, j. l. (2008). district policy and teachers‟ social networks. educational evaluation and policy analysis, 30(3), 203–235. eraut, m. (2004). informal learning in the workplace. studies in continuing education, 26(2), 247–273. eraut, m. (1994). developing professional knowledge and competence. london, england: routledge. european commission (2009). council conclusions of 12 may 2009 on a strategic framework for european cooperation in education and training (et 2020) [official journal c 119 of 28.5.2009]. retrieved from http://europa.eu/legislation_summaries/education_training_youth/general_framework/ef0016_en.htm. filliettaz, l. (2011). asking questions...getting answers: a sociopragmatic approach to vocational training interactions. pragmatics and society, 2(2), 234–259. fincher, s., & tenenberg, j. (2011). a commons leader„s vade mecum. university of kent press available at: http://www.cs.kent.ac.uk/people/staff/saf/share/papers/bt_111049_vadexmecum_final.pdf forret, m. l., & dougherty, t.w. (2004). networking behaviors and career outcomes: differences for men and women? journal of organizational behavior 25(3), 419–437. granovetter, m. s. (1973). the strength of weak ties. american journal of sociology, 78(6), 1360–1380. hargreaves, a. (2003). teaching in the knowledge society: education in the age of insecurity. new york, usa: teachers' college press. katz, s., earl, l. m., & jaffar, s. b. (2009). building and connecting learning communities: the power of networks for school improvement. thousand oaks, ca: corwin press. kerr, d., aiston, s., white, k., holland, m., & grayson, h. (2003). review of networked learning communities. maidenhead, uk: national foundation for educational research. kilduff, m., & tsai, w. (2007). social networks and organisations. los angeles, ca: sage. knight, p. (2006). quality enhancement and educational professional development. quality in higher education, 12(1), 29–40. koper, r., rusman, e., & sloep, p. (2005). „effective learning networks‟. article. retrieved from http://dspace.ou.nl/handle/1820/304. kyndt, e., dochy, f., & nijs, h. (2009). learning conditions for non-formal and informal workplace learning. journal of workplace learning, 21(5), 369–383. mccormick, r., fox, a., carmichael, p., & procter, r. (2011). researching and understanding educational networks. london, england: routledge. mehra, a., kilduff, m., & brass, d. j. (1998). at the margins: a distinctiveness approach to the social identity and social networks of underrepresented groups. academy of management journal, 41(4), 441–452. monge, p. r., & contractor, n. s. (2003). theories of communication networks. new york, ny: oxford university press. nonaka, i., & takeuchi, h. (1995). the knowledge-creating company: how japanese companies create the dynamics of innovation. new york, ny: oxford university press. paavola, s., lipponen, l., & hakkarainen, k. (2002). epistemological foundations for cscl: a comparison of three models of innovative knowledge communities. in g. stahl (ed.), proceedings of the conference on computer-supported collaborative learning: foundations for a cscl community (pp. 24–32). cscl ‟02. hillsdale, nj: erlbaum. international society of the learning sciences. retrieved from http://dl.acm.org/citation.cfm?id=1658616.1658621. http://www.cs.kent.ac.uk/people/staff/saf/share/papers/bt_111049_vadexmecum_final.pdf http://dl.acm.org/citation.cfm?id=1658616.1658621 n. pataraia et al. 16 | f l r pallant, j. (2010). spss survival manual: a step by step guide to data analysis using spss. crows next, australia: allen & unwin. pataraia, n., margaryan, a., falconer, i., littlejohn, a., & falconer, j. (2014). discovering academics‟ key learning connections: an ego-centric network approach to analysing learning about teaching. journal of workplace learning, 26(1), 56–72. pifer, m. (2010). such a dirty word: networks and networking in academic departments (unpublished doctoral dissertation). pennsylvania state university, usa. powell, w.w., koput, k.w., & smith-doerr, l. (1996). interorganizational collaboration and the locus of innovation: networks of learning in biotechnology. administrative science quarterly, 41(1), 116–145. reagans, r., & mcevily, b. (2003). network structure and knowledge transfer: the effects of cohesion and range. administrative science quarterly, 48(2), 240–267. rijt, j. van der, bossche, p. v. den, & segers, m. s. r. (2013). understanding informal feedback seeking in the workplace: the impact of the position in the organizational hierarchy. european journal of training and development, 37(1), 72–85. rogers, e. m. (2003). diffusion of innovations (5th ed.). new york, ny: the free press roxå, t., & mårtensson, k. (2009). significant conversations and significant networks: exploring the backstage of the teaching arena. studies in higher education, 34(5), 547–559. ruef, m. (2002). strong ties, weak ties and islands: structural and cultural predictors of organizational innovation. industrial and corporate change, 11(3), 427–449. schuck, s., aubusson, p., & buchanan, j. (2008). enhancing teacher education practice through professional learning conversations. european journal of teacher education, 31(2), 215–227. scott, j., & carrington, p. j. (2011). the sage handbook of social network analysis. london, uk: sage publications ltd. skule, s. (2004). learning conditions at work: a framework to understand and assess informal learning in the workplace. international journal of training and development, 8(1), 8–20. thomson, k. e. (2013). the nature of academics‟ informal conversation about teaching (unpublished doctoral dissertation). the university of sydney, australia. retrieved from http://ses.library.usyd.edu.au/bitstream/2123/9166/1/ke-thomson-2013-thesis.pdf. trinkle, c. (2009). twitter as a professional learning community. school library monthly, 26(4), 22–23. tynjälä, p. (2008). perspectives into learning at the workplace. educational research review, 3(2), 130– 154. tynjälä, p., & nikkanen, p. (2009). transformation of individual learning into organizational and networked learning in vocational education. in m. stenström., & p. tynjälä (eds.). towards integration of work and learning: strategies for connectivity and transformation (pp. 117–135). dordrecht, the netherlands: springer. tushman, m.l., & anderson, p. (1986). technological discontinuities and organizational environments. administrative science quarterly, 31(3), 439–465. wasserman, s., & faust, k. (1994). social network analysis: methods and applications. cambridge, england: cambridge university press. wellman, b. (2001). physical place and cyberplace: the rise of personalised networking. international journal of urban and regional research, 25(2): 227−52. wellman, b. (1998). networks in the global village: life in contemporary communities. boulder, co: westview press. wenger, e., trayner, b., & de laat, m. (2011). promoting and assessing value creation in communities and networks: a conceptual framework. heerlen, the netherlands: ruud de moor centrum, open university. retrieved from http://www. knowledge-architecture. com/downloads/wenger_trayner_delaat_value_creation. pdf frontline learning research 3 (2014) 64-77 issn 2295-3159 corresponding author: siân e. jones, department of psychology, social work, and public health, oxford brookes university gipsy lane ox3 0bp +44 (0)1865 48371, sianjones@brookes.ac.uk http://dx.doi.org/10.14786/flr.v2i1.80 64 | f l r bullying and belonging: teachers’ reports of school aggression siân emily jones a , antony s.r. manstead b , andrew g. livingstone c a oxford brookes university, united kingdom b cardiff university, united kingdom c university of exeter, united kingdom article received 17th january 2014 / revised 24th february 2014 / accepted 3rd march 2014 / available online 25th april 2014 abstract research on bullying has confirmed that social identity processes and group-based emotions are pertinent to children’s responses to bullying. however, such research has been done largely with child participants, has been quantitative in nature, and has often relied on scenarios to portray bullying. the present paper departs from this methodology by examining group processes in qualitative reports of bullying provided by teachers. fifty-one teachers completed an internet-based survey about a bullying incident at a school where they worked. thematic analysis of survey responses concerned two core themes in the reports: (a) children ganging up on another child and (b) children sticking together to protect each other. there was evidence that children act in specific ways, in line with social identity processes, in order to support or resist bullying. there was also evidence that teachers understand bullying to be a group phenomenon. the implications of these findings for anti-bullying interventions are discussed. keywords: bullying; teachers; group processes; social identity 1. introduction bullying can happen in any setting where power relations exist (smith &brain, 2000). of particular concern in this paper is bullying in schools, because research indicates that bullying is a common experience for such children. for example, representative research shows that 28% of students aged 12-18 years reported being bullied during the school year (roberts, zhang, truman, & snyder, 2012). the effects of bullying are serious: targets may suffer higher rates of anxiety, depression, physical health problems, and social maladjustment (espelage, low, & de la rue, 2012). such negative consequences may last into s. e. jones et al. 65 | f l r adulthood (e.g., hunter, mora-merchan, & ortega, 2004; olweus, 1994). as these effects touch both perpetrators and targets (gini & pozzoli, 2009) and those who witness it (nishina & juvonen, 2005), it is important to reduce incidences of bullying. the finding that those who witness bullying are susceptible to negative consequences points to the ways in which bullying may be understood as a group process. indeed, recent research supports a framing of bullying in these terms. since the publication of atlas and pepler‟s (1998) observational study, which revealed that peers were present in 85% of all bullying episodes on a school playground, a burgeoning research literature has confirmed that it is helpful to regard bullying as a group process. for example, espelage, holt, and henkel (2003) used peer nomination techniques (for a review see hymel, vaillancourt, mcdougall, & renshaw, 2002) to identify peer groups of middle school children, and followed them longitudinally for a year. they found that members of peer groups that engaged in bullying increased their own bullying behaviours over time. additionally, using peer nomination techniques as part of the participant-role approach, it has been shown that peers may form groups that work collectively to resist bullying: sainio, veenstra, huitsing, and salmivalli (2011) found that targets who had one or more classmates defending them when they were bullied were less anxious, less depressed, and had higher self-esteem than undefended targets, even when the frequency of the bullying incidents was taken into account. in line with the above research findings, in recent years the zeitgeist in terms of responses to bullying in schools has changed from a focus at the level of the individual to interventions focused at the school level (for a review of school/class-wide interventions, see horne, stoddard, & belle, 2007). horne et al. (2007) note that a common feature of these group-level interventions is that they work at the whole school or class level, as well as targeting those directly affected by a bullying incident. as such, these interventions focus on social skills training of individuals, but do not address the peer/friendship group dynamics identified by researchers, and discussed in greater detail below. indeed, although much research has been directed at a group-level understanding of children‟s responses to bullying, comparatively little research has looked at the group-level nature of teachers‟ responses. in light of this, this paper aims to look at how groups are represented in teachers‟ responses to bullying. 1.1 a social identity account of bullying empirical work looking at bullying as a group process has used social identity theory (sit; tajfel & turner, 1979) as a means of understanding why children might work in groups to (a) bully, and (b) overcome bullying. this theory proposes that a person‟s group memberships are an important part of their identity – their social identity – and, as a consequence, group members will try to enhance their own self-esteem by seeking to maintain a positive image of their group. the more strongly one identifies with a given group membership, the more likely one is to act on behalf of the (positive image of) the group; in other words, the more likely one is to enhance one‟s social identity. the group image is epitomised, according to sit, by a set of group norms to which its members are expected to adhere (turner, 1999). as such, group members are likely to be rewarded for adherence to group norms, or rejected by the group when they fail to adhere to them (morrison, 2006). building on this, it was hypothesized (e.g., jones, haslam, york, & ryan, 2008; jones, livingstone, & manstead, 2011, 2012; nesdale, 2007) that bullying might be a set of behaviours that is motivated by social identity processes, including levels of ingroup identification, and adherence to group norms. in line with this hypothesis, a number of studies have indicated the role of social identity processes in maintaining bullying. these studies have been mainly conducted using the minimal group paradigm (tajfel, billig, bundy, & flament, 1971), in which children are assigned to a group at random (but ostensibly on the basis of some activity, such as a dot-estimation task) and their responses to hypothetical intergroup events are recorded (see dunham, baron & carey, 2011, for a review of minimal group research with children). ojala and nesdale (2004) demonstrated that children understand the need for group members to behave normatively, even if this involves bullying. they gave children scenarios to read, and found that children understood that story characters who engaged in bullying would be rejected by a group with an anti-bullying norm, but accepted by a group with a pro-bullying norm. evidence from jones et al. (2008), using the minimal group paradigm, showed that children encouraged to identify with a perpetrating group in a scenario s. e. jones et al. 66 | f l r concluded that one bullying child from that group was deserving of punishment for a bullying incident, whereas third party group members concluded that the whole of the perpetrating group was punishable. furthermore, nesdale, durkin, maass, kiesner, and griffiths (2008) showed, in a minimal group study, that children‟s intentions to engage in bullying were greater when they were assigned to a group that had a norm of outgroup-disliking, rather than a norm for outgroup-liking. in later research, jones et al. (2011) showed that children who identify highly with a target feel more anger on behalf of that target – they “stick together” with a target of bullying, while children who identify with a bullying group express more pride – and want to be friends with the bullying children. thus, social identity processes might account for children‟s responses to bullying, in terms of a need to maintain a positive ingroup image, and to adhere to ingroup norms. 1.2 teachers’ responses to bullying despite research showing that group processes might be involved in bullying, little research effort has been spent examining teachers‟ awareness of processes underlying bullying (nesdale & pickering, 2006). this lack of research attention is problematic in light of the finding from a study by whitney and smith (1993), which found that less than half of teachers intervened when a pupil was being bullied. this is despite the fact that it is a recommended government policy for children to be actively encouraged to talk to adults about bullying, to see that it is stopped (department for children, schools and families, 2007). more worryingly, teacher intervention in bullying decreases in likelihood as pupils get older (o‟moore, kirkham, & smith, 1998), and incidences of bullying increase with age (horne et al., 2007). one possible reason for lack of intervention is lack of awareness or understanding of a situation as bullying. fekkes et al. (2005) showed that a substantial number of both teachers and parents were unaware that the child was being bullied; for classmates this figure was lower. teachers did not speak to bullies, only to the bullied children. children indicate that verbal and psychological bullying is more prevalent than physical bullying, yet few teachers recognize these incidents or identify them as bullying (hazler, miller, carney, & green, 2001). boulton (1997), investigating teachers‟ definitions of and attitudes towards bullying, found that one in four teachers did not regard name-calling, spreading rumours or social exclusion as bullying. khoury-kassabri (2009) argued that in many cases school staff do not have the ability to determine who the victims and bullies are, and do not make an effort to distinguish each student‟s role in the bullying situation. thus, a student‟s involvement in bullying, in whatever role, is associated with being verbally or physically punished by teachers. also, in some instances, students who are involved in violent acts (even as victims) are perceived as disrupting the learning process, which might increase the probability of being punished. yoon and kerber (2003) investigated teacher attitudes via their responses to various bullying scenarios. they found that when teachers are unaware of the extent of bullying or when they did not consider the behaviour to be serious, they exhibited passive attitudes towards bullying and did not intervene or did not do so effectively. because non-physical acts of bullying are easier to hide, teachers must be aware of the symptomatic bullying behaviours (yoon & kerber, 2003). in a study by nicolaides, toda, and smith (2002), trainee teachers were reasonably accurate in their estimates of the frequency of bullying in school and the extent of teacher intervention. they were unaware that self-reports of victimization decline with age. in addition, they believed that girls and boys were equally likely to be bullies and that bullies have low selfesteem and lack social skills. these trainee teachers saw their role as instrumental in reducing bullying in the classroom. whether or not they will be effective in that role is contingent on a number of factors. further to this, bauman and del rio (2005) used a questionnaire assessing knowledge, attitudes and beliefs about bullying on a sample of 82 trainee teachers in the united states. participants had some accurate knowledge as well as some beliefs and attitudes that would not be consistent with effective teacher behaviours towards students involved in bullying. only 6 per cent mentioned repetitive behaviour and 28 per cent included power imbalance in their definitions. these are the two elements that are unique to bullying vis-à-vis s. e. jones et al. 67 | f l r aggression. boulton et al. (2014) found that willingness to intervene by teachers corresponded to the type of bullying portrayed. in a similar vein, hazler et al. (2001) reported that teachers frequently label any physical conflict as bullying, even when it is not, and show less concern for and intent to intervene in situations with the potential for social or emotional harm. the teachers were interested in further training. teachers‟ views and beliefs about bullying inform their anti-bullying action or inaction. yoon and kerber (2003)‟s research shows that teachers are less likely to intervene in bullying if they are unsympathetic to victims or believe that getting involved is unnecessary. kochenderfer-ladd and pelletier (2008) showed that avoidant beliefs (“children would not be bullied or picked on if they avoided mean children”) were predictive of separating students which was then associated both directly and indirectly (via reduced revenge seeking) with lower levels of peer victimization. teachers who held normative beliefs (“bullying is normative behaviour that helps children learn social norms”) about bullying were not likely to intervene. holt and keyes (2004) found that 27 % of teachers agreed with the statement, „a little teasing doesn‟t hurt‟. research suggests that teachers are aware of the group-level nature of bullying. yubero and navarro (2006) found that teachers believed that girls employ bullying tactics planned in advance, with the objective of creating unease in their relationships “in order to obtain a more advantageous position within their group.” (p. 499, emphasis ours). a vignette study by nesdale and pickering (2006) examined the impact on teachers‟ reactions to children‟s aggression of three variables, two of which were related to the aggressors and one was related to the teachers. teachers each read a scenario that described an aggressive episode committed by a group of boys against a boy from another class. the aggressors were either good or bad children, who were either popular or unpopular with their classroom peers. in addition, the scenario manipulated the teachers‟ social identity, in terms of the strength of their identification with the class to be either high or low. analysis of the teachers‟ ratings revealed a consistent negative response from the teachers towards the aggressors versus the victim. however, the teachers‟ responses were also influenced by the aggressors‟ goodness and popularity, and the teachers‟ class identification. 1.3. the present study given this, and that empirical research shows that social identity processes are relevant to bullying, it seems timely to explore whether teachers‟ narratives about bullying include mention of the role of groups. we sought to examine teachers‟ accounts of school bullying, with a particular focus on the way in which bullying involving more than two children was described. owing to the paucity of previous research on teachers‟ perceptions of bullying, this study was exploratory in nature. we used qualitative research methods as a means to explore the way in which teachers represented bullying episodes among pupils, and as a way of investigating the content of the bullying episodes and the approaches that were used to deal with them. qualitative research methods thus enabled us to consider a range of bullying episodes in order to determine whether there was any evidence that the group processes that have been investigated empirically are echoed in teachers‟ reports of school bullying. accordingly, teachers were invited to complete an internet-based survey of their experiences of children‟s bullying at a school where they had worked. through a series of open-ended questions, they were asked to recall the details of a bullying incident. 2. methods 2.1 data collection and participants following ethical approval, teachers were invited to take part in an online survey (hosted by survey monkey). to encourage participation, links to the survey were hosted on anti-bullying sites, social networking sites, and on discussion forums aimed at teachers. one hundred and fifty-six teachers responded to the questionnaire. responses from 51 teachers (25% of the total number of respondents) were sufficiently s. e. jones et al. 68 | f l r complete (i.e., these participants had answered, in a meaningful way, at least one open-ended question concerning the bullying incident) to be included in analyses. of these, 32 were female and 15 were male (four unknown). thirteen teachers taught at primary schools, 35 at secondary schools (three unknown). all teachers taught at state schools. in the interests of anonymity, no further demographic information about participants was gathered. 2.2 children and schools participants provided data concerning the children involved in the bullying incident and the schools in which these incidents took place. 2.2.1 age of children bullying incidents were reported among children between 6-7 years, up to 17-18 year-olds. bullying was most frequently reported among 11-13 year-olds, (14 cases) and was not reported among 4-6 year-olds. this information is reported in figure 1. figure 1. the number of bullying incidents reported by participants as a function of age group. 2.2.2. size the modal school size was over 1000 pupils (n = 13), while the modal class size was 20-29 pupils (n = 22). bullying incidents were most frequently reported in this sample in schools with over 1000 students where the class size was between 20-29 pupils. 2.3 questionnaire items three questionnaire items concerned the details of a bullying incident that had occurred at a school in which they had worked. open-response questions asked for details about (1) the reporting of the bullying incident, (2) the nature of the bullying, and (3) the extent to which children involved in the bullying were 0 1 2 3 4 5 6 7 4-5 years 5-6 years 6-7 years 7-8 years 8-9 years 9-10 years 10-11 years 11-12 years 12-13 years 13-14 years 14-15 years 15-16 years 16-17 years 17-18 years frequency of cases a ge g ro u p o f c h ild re n s. e. jones et al. 69 | f l r familiar to each other. following this were closed questions about the age of the children involved, sex of the teacher, school type, school and class size, and about whether the school had an anti-bullying policy. 2.4 data analysis strategy all usable data from open-response items were transferred to nvivo, and then submitted to a thematic analysis. two themes used to inform the analysis were guided by the extant research (see jones et al., 2011) on social identity processes: 1) children ganging up on another child, (condoning and joining in the bullying) and (2) children sticking together with the target (supporting the target and/or reporting the bullying). the analysis first involved organizing the data into categories according to the number of perpetrators involved. of the 51 incidents reported, seven involved only two children (one perpetrator and one target) and 44 cases involved more than one perpetrator. because the focus is on group processes in bullying, subsequent analyses concentrated on the latter 44 cases. data from these cases were coded under descriptive categories, such as “school journey” or “cyberbullying” in order to reduce the data to analyzable form (coffey & atkinson, 1996). extracts from the data were coded for each category to ensure that later abstractions would „fit‟ the data (straus & corbin, 1998). these descriptive categories were then arranged around the two primary themes, reflecting the nature of the bullying and the processes involved in reporting it, as indicated in the teachers‟ reports. illustrative extracts of each primary theme are reported below. 3. results 3.1 primary themes the following primary themes were examined in analysis of the teachers‟ reports: (1) children ganging up on another child, and (2) children sticking together. these are outlined in figure 2, and in more detail below, along with illustrative extracts. in parentheses immediately following each extract is the participant number, participant sex, and the age of the children involved in the bullying. figure 2. themes and sub-themes in the data (number of cases categorized in this theme in parentheses). support from peers (33) ganging up sticking together multiplicity of place (7) multiplicity of perpetrators (26) multiplicity of methods (28) support from school (41) s. e. jones et al. 70 | f l r 3.1.1 ganging up particularly common in teachers‟ accounts of bullying involving more than one perpetrator was the way in which children were seen as „ganging up‟ on their target. this theme could be divided into three subthemes. the first concerned the multiplicity of the perpetrators doing the bullying: “i discovered that a group of girls in my class were bullying one particular child ... there were about 7 or 8 involved altogether” (p30, female, 10-11 years old). “a year 8 boy [was] repeatedly called homophobic names by a number of class peers” (p22, female, 12-13 years old). “the [bullying] group involved two girls and four boys” (p4, male, 12-14 years old). in a few instances the ganging up by multiple perpetrators was directed at a group-level characteristic in the target, like race or sexuality: “one boy at a lunch table directed the word "nigger" at one of our black students…the white students at the table had been directing racial comments at the black student for quite some time” (p44, female, 12-13 years old). “there was a case when teenagers were harassing a student who was perceived to be gay” (p45, female, 12-13 years old) “a boy repeatedly called homophobic names by a number of class peers” (p22, female, 12-13 years old). the majority of the bullying occurred between perpetrators and a target who were members of the same class group, and who were sometimes described as close friends before the bullying started, but who would then gang up on a target: “they appeared to be good friends at the start of the year and sat next to each other in class. they certainly had several classes together” (p 25, female, 1112 years old). “bullying between girls that had been friends ... the main three girls had been close friends” (p2, female, 15-16 years old). “same class, close friends” (p3, female, 11-12 years old) “…the target student had previously been good friends with the bullies... children involved were in some of the same classes” (p8, female, 17-8 years old). “same class... child being bullied was friends with those showing bullying behaviour” (p14, male, 9-11 years old). ganging up was also apparent in the multiplicity of methods (the second sub-theme) that were used to bully the target according to many reports: “name-calling, nasty comments, bringing student to tears, getting others to ignore student, hiding student’s possessions” (p28, male, 13-14 years old). “the bullying was mostly gossiping, rumour-spreading and withdrawing friendships (also encouraging others to withdraw friendships)” (p2, female, 15-16 years old). s. e. jones et al. 71 | f l r “bullying included name-calling, throwing small objects [and] trying to split up friendship groups” (p19, unknown, 11-13 years old). among the reports, it was uncommon for one „type‟ of bullying to be administered to a target. also prevalent was that bullying occurred not just at school, but in multiple places (the third sub-theme): “bullying began in school and then moved to outside school and through e-mail and im [instant messaging]” (p29, female, 12-14 years). “bullying spilled over into extra-curricular activities” (p14, male, 9-11 years old). “the bullying took place mostly at home but intimidation followed in school” (p31, unknown, 17-18 years old). “this happened in school and continued out of school” (p32, female, 12-13 years old). “happened in school halls at first but carried over to homes” (p37, female, 14-16 years old). the effects of „ganging up‟ were seen in the emotional experiences of the targets, as reported by the teachers: “the target had been devastated by the bullying.” (p4, male, 12-14 years old) “they [parents] said he was very distressed and did not want to return to class as he was too afraid.” (p5, female, 15-16 years old) “name calling (about appearance)... is what upset the girl. (p10, male, 11-12 years old) thus, bullying is construed as a set of activities whereby a group of children „gang up‟ on another child, as illustrated by the multiplicity of the perpetrators involved, the acts that take place, the spaces they take place in, and the way in which children can turn upon former friends, with negative emotional reactions sometimes directly induced by the perpetrators, and often evident in the targets‟ responses. 3.1.2 sticking together in parallel with „ganging up‟ on the part of the perpetrators, in the majority of cases children who found themselves to be the target of bullying were supported by their peers. peers often showed solidarity with the target, independently of support of adults, in reporting the bullying to a teacher: “children (friends of the bullied) approached me and told me about what had happened, giving me names of the bullies, also of other children who could corroborate their story.[t]hey had not approached any other teachers or informed their parents” (p19, unknown, 11-13 years old). “a child reported the bullying – a friend of the child reported it” (p3, female, 11-12 years old). “his friend (not the target) reported to me an incident of verbal and physical bullying of the pupil” (p17, female, 13-14 years old). “five of the boy’s friends were all supportive of the bullying claims and spoke to the teacher about it” (p26, female, 14-15 years old). s. e. jones et al. 72 | f l r peers also encouraged targets to report bullying for themselves, because they saw the bullying behaviour as illegitimate: “she was supported by a small number of peers who had encouraged her to complain and felt her treatment was unfair” (p10, male, 11-12 years old). in one case alternative friendship groups were effective in dissipating negative effects of bullying: “[he] found a different friendship group that seemed to be more effective than the school intervention” (p20, male, 11-12 years old). there is evidence, then, that some children who are aware of bullying going on in their class appraise the situation as unfair, and work together as a group to „stick by‟ the target in order to overcome the bullying. beyond this, there was evidence in the teachers‟ responses that the school stuck together to deal with the bullying, often in line with a whole school policy: “in this instance i spoke to the whole class as well as the girls involved. i also did my next class assembly on bullying so that it was kept in the forefront of their minds” (p30, female, 10-11 years old). “whole year group received a number of anti-homophobia forum theatre and in-class support resources “(p22, female, 12 -13 years old). “there was a whole year 7 assembly on cyberbullying and how it was easy for comments to have an effect. there was also a pse [personal and social education] session on cyberbullying that linked in with this” (p25, female, 1112 years old). “in all the tutor groups we reminded students about the college’s zero tolerance policy towards bullying” (p9, female, 16 -17 years old). thus, not only children, but staff members were seen here to “stick together” to promote an antibullying message to pupils. 4. discussion the vast majority of cases that were reported by teachers for this research involved more than a twoperson perpetrator-target dyad. the data presented above provided a more nuanced picture of the ways in which social identity processes might be relevant to the problem of school bullying than that provided by previous experimental work (e.g., jones et al., 2011; nesdale et al., 2008), which has focused on strength of identification and group norms . specifically, it emerged that bullying in groups has a substantial intragroup dynamic, with bullying sometimes occurring among former friends. this bullying took multiple forms, and happened in multiple spaces. despite this, there was evidence that children work together in groups to overcome bullying. 4.1 social identity and bullying this research lends support to a social identity-based account of bullying. there was evidence in the teachers‟ accounts that children form groups in order to bully, and that bullying is based on characteristics of group membership (e.g., sexuality, race). there was also evidence that children form supportive groups around targets of bullying, and that children are encouraged to identify with school-level group norms s. e. jones et al. 73 | f l r surrounding peer victimization. these findings are thus in line with scenario-based research (e.g., jones et al., 2011, 2012; nesdale et al., 2008) showing children‟s tendency to follow group norms surrounding bullying, and to identify with, and behave in line with, their friendship groups. a novel insight for research looking at social identity processes in bullying is that bullying occurs between children who were former friends. situations were described by teachers whereby two or more children would target someone who was previously perceived to be part of their friendship group. notwithstanding possible misconceptions by teachers regarding friendship groups, or that this sample was self-selected, and likely to be unrepresentative of all bullying incidents in a school, or specific time period, this finding is consistent with recent research by mishna, wiener, and pepler (2008), whose interview data showed that children were sometimes targeted by their friends. this finding prompted the authors to pose further research questions concerning how friendships might become bullying relationships, as well as how children deal with such bullying. from a social identity perspective, one might also ask about the group dynamics entailed in such bullying. jetten, branscombe, spears, and mckimmie (2003) coined the term peripheral group members to describe new group members, or those who represent the group‟s prototype less well. it may be the case that the children who are bullied from within friendship groups are peripheral group members who want to become closer to the friendship group, but are bullied because they are unsure of the norms of that group. or, relatedly, is bullying within groups a way of policing friendship group norms, such that those who are bullied are those members who fail to conform to such norms? alternatively, is it the case that each friendship group contains multiple alliances between children such that the group is made up of one superordinate, and several subordinate groups, between which bullying occurs? these are all questions that could be addressed in future research. 4.2 teachers’ views this study shows that teachers are aware of a group-level nature to bullying. here, teachers reported which were the targets and perpetrators of the bullying, as well as the “group of girls and boys” who surround and support the perpetrators and targets. additionally, the teachers recognized that targets were often supported by friends in reporting what had happened. this is in line with yubero and navarro (2006), who found that teachers also showed awareness of the “relational” nature of bullying. the findings are also consistent with those of nesdale and pickering (2006) in showing that social identity concerns, regarding the schools norms about bullying (seen in their adherence to school policy) often came to the fore. indeed, teachers‟ responses to the bullying seemed overwhelmingly to stem from a need to ensure that key messages concerning bullying were understood at a group level: extensive group-level interventions were executed, in order to reinforce anti-bullying messages. nonetheless, the question regarding the extent to which these work in harmony with or at cross purposes to other aspects of the school‟s ethos remains open. it is not clear whether the anti-bullying strategies noted above are part of a coherent norm-based strategy, or an ad-hoc reaction to the bullying. thus, from a social identity perspective, it would be interesting to consider more carefully, and in a larger-scale study, with a representative sample of teachers, the processes of formation, dissemination, and acceptance of school-wide anti-bullying norms among school pupils and staff. 4.3 practical implications the research reported here has implications both for research into bullying and for practice. for researchers, it is apparent that one bullying episode is not always of a single type (e.g., verbal bullying, physical bullying, emotional bullying, or cyberbullying) as classified in the literature (e.g., rigby, 2007). although rigby recognized that these forms of bullying may co-occur, scenario-based research, such as jones, manstead and livingstone‟s (2009) work on cyberbullying, or hitti, mulvey, rutland, abrams, and killen‟s (in press) work on social exclusion, has typically focused on just one form of bullying. it may be advisable in future research to represent various forms of bullying as happening concurrently, in order to represent more accurately the ways in which children „gang up‟ on a peer. similarly, given the evidence reported above that children often show a supportive response to targets of bullying, this type of reaction could be investigated in scenario-based research: specifically, when there are children in support of a target, and children in support of a perpetrator, what determines bystanders‟ reactions? it should also be noted that s. e. jones et al. 74 | f l r previous research, to our knowledge, has only focused on one understanding of these scenarios (i.e., what teachers or children think). here, we assessed teachers‟ views. it would have been interesting to triangulate these with children‟s or parents‟ views about these same instances of bullying. this would have compromised anonymity, but would certainly be feasible in the context of scenario-based research. resolving mismatches and omissions in reporting of bullying could provide another route to intervention. at a practical level, this study points to a potential avenue for intervention in terms of teachers‟ responses to bullying. while the bullying described frequently happened among groups of children, current interventions do not focus on the group dynamics among perpetrating children that might have led to and sustained the bullying. thus, future interventions could seek to raise teachers‟ awareness of group dynamics, as outlined by social identity research, and of the (group-based) emotional responses of children other than the target. in this way, teachers might be better attuned to the group dynamics of the classroom and thereby be better positioned to „nip bullying in the bud‟ before it escalates. 4.4 conclusions the main aim in this research was to explore how teachers described bullying episodes in which they have been involved, with a particular focus on the role of the group in perpetrating, dealing with and stopping these bullying episodes. the qualitative analysis employed here was well suited to this aim. although it does not allow us to make conclusive statements regarding the broader picture of group bullying, for example concerning how commonly bullying episodes involve the group, or the specific characteristics of those children who are involved in group bullying, it does permit exploration of the content of bullying episodes. previous scenario-based research had shown that social identity concerns may be relevant to bullying. what is evident from the present study is that children bully in groups and work together to resist bullying. the teachers‟ reports also provide insight into the specific activities that children engage in in order to bully or support other children. the research could therefore be used as a basis for (a) helping teachers to understand better the nature of bullying, and (b) researchers to represent the group processes that children engage in a more realistic and more nuanced way in their empirical work. keypoints bullying may be understood as a group phenomenon. social identity theory gives a framework for how peer group processes might maintain or resist bullying. much work on bullying in groups has been scenario-based experimental research, while interventions work at the school or class, rather than at the friendship group level. this study asks teachers for accounts of bullying. the teachers provided rich accounts of bullying that evidence the group processes that might undergird its support or resistance, and which point to ways in which bullying might be addressed at the peer (friendship) group level. acknowledgements the first author gratefully acknowledges support from the economic and social research council (award number: pta-031-2006-00548). the third author would like to thank the leverhulme trust (ecf/2007/0050) for their support. we are also grateful to guida de abreu for her comments on an earlier draft of this manuscript, and to the children who took part in this research, and to the school, teachers, and parents who allowed them to do so. s. e. jones et al. 75 | f l r references atlas, r. s., & pepler, d. j. (1998). observations of bullying in the classroom. journal of educational research, 92, 86–89. doi: 10.1080/00220679809597580 bauman, s., & del rio, a. (2005). knowledge and beliefs about bullying in schools: comparing pre-service teachers in the united states and the united kingdom. school psychology international, 26(4), 428-442 doi: 10.1177/0143034305059019 boulton, m. j. (1997). teachers' views on bullying: definitions, attitudes and ability to cope. british journal of educational psychology, 67,(2) 223-233. doi: 10.1111/j.2044-8279.1997.tb01239.x boulton, m.j., hardcastle, k., down, j. simmonds, j., & fowles, j. a. (2014). a comparison of pre-service teachers‟ responses to cyber versus traditional bullying scenarios: similarities and differences and implications for practice. journal of teacher education. 65 ,(2)145-155. doi:10.1177/0022487113511496 coffey a & atkinson p (1996). making sense of qualitative data: complementary strategies. thousand oaks ca: sage. department for children, school, and families (2007). safe to learn: embedding anti-bullying work in schools. retrieved on 03/31/2011 from: http://www.teachernet.gov.uk/publications dunham, y., baron., a.s., & carey, s. (2011). consequences of "minimal" group affiliations in children child development, 82(3), 793-811. doi: 10.1111/j.1467-8624.2011.01577.x espelage, d.l., low, s. & de la rue, l. (2012). relations between peer victimization subtypes, family violence, and psychological outcomes during early adolescence. psychology of violence, 2, 313-24 fekkes, m., pijpers, f. i. m. and verloove-vanhorick, s. p. (2005). bullying: who does what, when and where? involvement of children, teachers and parents in bullying behavior. health education research: theory and practice, 20: 81-91. hazler, r. j., miller, d. l., carney, j. v. & green, s. (2001). adult recognition of school bullying situations. educational research 43, 133–46. hitti, a., mulvey, k. l., rutland, a., abrams, d. & killen, m. (2013). when is it ok to exclude a member of the ingroup?: children‟s and adolescents‟ social reasoning. social development, doi:org/10.1111/sode.12047 horne, a.m., stoddard, j.l., & bell, c.d. (2007). group approaches to reducing aggression and bullying in school. group dynamics theory, research and practice, 11 (4), 262-271.doi:10.1037/10892699.11.4.262 holt, m. k. & keyes, m. a. (2004). teachers‟ attitudes towards bullying. in d. l. espelage and s. m. swearer (eds). bullying in american schools: a social-ecological perspective on prevention and intervention, (pp. 121–40). mahwah, nj: erlbaum. hunter, s.c., mora-merchán, j.a., & ortega, r. (2004). the long-term effects of coping strategy use in the victims of bullying. the spanish journal of psychology, 7 (1), 3-12. hymel, s., vaillancourt, t., mcdougall, p., & renshaw, p.d. (2002). peer acceptance and rejection in childhood. in p.k. smith & c.h. hart (eds.), blackwell handbook of childhood social development (pp. 265–284). malden, ma: blackwell. jetten, j., branscombe, n. r., spears, r., & mckimmie, b. m. (2003). predicting the paths of peripherals: the interaction of identification and future possibilities. personality and social psychology bulletin, 29, 130-140. doi: 10.1177/0146167202238378 jones, s. e., haslam, s. a., york, l., & ryan, m. k. (2008). rotten apple or rotten barrel? social identity and children‟s responses to bullying. british journal of developmental psychology, 26(1), 117–132. doi:10.1348/026151007x200385 jones, s.e., manstead, a.s.r., & livingstone, a.g. (2012). fair-weather or foul-weather friends? group identification and children‟s responses to bullying. social psychology and personality science, 3(4), 414-420. doi: 10.1177/1948550611425105 http://www.teachernet.gov.uk/publications http://dx.doi.org/10.1111/sode.12047 s. e. jones et al. 76 | f l r jones, s.e., manstead, a.s.r., & livingstone, a.g. (2011). ganging up or sticking together: group processes and children‟s responses to bullying. british journal of psychology, 102 (1), 71-96. doi: 10.1348/000712610x502826 jones, s.e., manstead, a.s.r.,& livingstone, a.g. (2009). birds of a feather bully together: group processes and children‟s responses to bullying. british journal of developmental psychology, 27, 853873. doi: 10.1348/026151008_390267 khoury-kassabri, m. (2009). the relationship between staff maltreatment of students and bully-victim group membership. child abuse & neglect, 33, 914-923. kochenderfer-ladd, b., & pelletier, m. (2008). teachers' views and beliefs about bullying: influences on classroom management strategies and students‟ coping with peer victimization. journal of school psychology 46, 431-453. doi: 10.1016/j.jsp.2007.07.005 mishna, f., wiener, j., & pepler, d. (2008). some of my best friends: experiences of bullying within friendships. school psychology international, 29(5), 549-573. doi:10.1177/0143034308099201 morrison, b. (2006). school bullying and restorative justice: toward a theoretical understanding of the role of respect, pride and shame. journal of social issues, 62,(2) 371-392 doi: 10.1111/j.15404560.2006.00455.x nesdale, d. (2007). peer groups and children's school bullying: scapegoating and other group processes. european journal of developmental psychology, 4, 388-392.doi: 10.1080/17405620701530339 nesdale, d., durkin, k., maass, a., kiesner, j., & griffiths, j. (2008). effects of group norms on children's intentions to bully. social development, 17, 889–907. doi: 10.1111/j.1467-9507.2008.00475.x nesdale,d.,& pickering,k. (2006) teacher‟s reactions to children‟s aggression. social development, 15,(1)109-127. doi: 10.1111/j.1467-9507.2006.00332.x nicolaides, s., toda, y. & smith, p. k. (2002). knowledge and attitudes about school bullying in trainee teachers. british journal of educational psychology 72, 105–18. doi: 10.1348/000709902158793 ojala, k., & nesdale, d. (2004). bullying and social identity: the effects of group norms and distinctiveness threat on attitudes towards bullying. british journal of developmental psychology, 22, 19–35. doi: 10.1348/026151004772901096 olweus, d. (1994) annotation: bullying at school: basic facts and effects of a school based intervention program. journal of child psychology and psychiatry and allied disciplines, 35, 1171–1190.no doi o'moore, m., kirkham, c., & smith, m. (1998) bullying in schools in ireland : a nationwide study. irish educational studies, 17, 255 – 271. rigby, k. (2007). bullying in schools and what to do about it (updated, revised). melbourne: australian council for education research. roberts, s., zhang, j., truman, j., & snyder, t. d. (2012). indicators of school crime and safety: 2011 (pub no. nces 2012-002/ncj 236021). washington, dc: u.s. department of education and u.s. department of justice. retrieved from http://nces.ed.gov/pubs2012/2012002.pdf sainio, m., veenstra, r., huitsing, g., & salmivalli, c. (2011). victims and their defenders: a dyadic approach. international journal for behavioral development 35, 144-151. doi: 10.1177/0165025410378068 smith, p.k., & brain, s. (2000). bullying in schools: lessons from two decades of research. aggressive behavior, 26, 1-9. doi: 10.1002/(sici)1098-2337 strauss, a., & corbin, j. (1998). basics of qualitative research: techniques and procedures for developing grounded theory. thousand oaks, ca: sage. tajfel, h., billig, m. g., bundy, r. p., & flament, c. (1971). social categorization and intergroup behavior. european journal of social psychology, 1, 149-177. tajfel, h., & turner, j. (1979). an integrative theory of intergroup conflict. in w.g. austin & s. worchel (eds.) the social psychology of intergroup relations. (pp. 7-24). monterey, ca: brooks cole. turner, j. c. (1999). some current issues in research on social identity and self-categorization theories. in n. ellemers, r. spears, & b. doosje (eds.) social identity: context, commitment, content. (pp. 6-34). oxford: blackwell. http://nces.ed.gov/pubs2012/2012002.pdf s. e. jones et al. 77 | f l r whitney, i. & smith, p.k. (1993) a survey of the nature and extent of bullying in junior/middle and secondary schools. educational research, 35, 3–25. yoon, j. s., & kerber, k. (2003). bullying: elementary teachers‟ attitudes and intervention strategies. research in education 69, 27–35. yubero, s., & navarro, r. (2006): student‟s and teachers‟ views of gender-related aspects of aggression. school psychology international, 27, 488-512. doi: 10.1177/0143034306070436 frontline learning research 6 (2014) 26-45 issn 2295-3159 corresponding author: christoph könig, institute of educational science, university of regensburg, universitätsstrasse 31, d-93051 regensburg, germany. e-mail: christoph.koenig@ur.de doi: http://dx.doi.org/10.14786/flr.v2i4.109 26 | f l r a change in perspective – teacher education as an open system christoph könig a , regina h. mulder a a university of regensburg, germany article received 30 april 2014 / revised 5 june 2014 / accepted 11 september 2014 / available online 24 september 2014 abstract teacher education is the environment for the learning and instruction of prospective teachers. its structure, components, and contents shape the development of relevant competences which enable prospective teachers to be effective in the classroom. but its relevance is questioned because respective research, characterised by inconclusive results, does not offer explanations about the reasons why certain teacher education programmes are more effective than others in the development of relevant competences. one reason for the lack of explanations can be found in the way research assesses the effectiveness of teacher education. this might be due to problems regarding the conceptualisations of teacher education, as well as to the inherent selection and nonrandom allocation problems in research on the relation between teacher education and student achievement. in this paper we respond to claims for an organisational perspective on teacher education and develop such a new perspective. accordingly, we provide these claims with an adequate theoretical foundation and develop an organisational model of teacher education based on open systems theory. besides being one of the first integrative organisational models of teacher education, it is among the first models which illustrate the relations and interdependencies of systems, its different parts, and its different levels, and enables researchers to investigate these interdependencies. the development of this model is further based on an alteration of the input variables of the concept of teacher quality. moreover, the model has consequences for the notion of teacher education effectiveness. we illustrate these changes, and discuss them and the model with respect to possible areas of further research. keywords: teacher selection; teacher allocation; teacher education effectiveness; open system; positive matching c. könig & r.mulder 27 | f l r 1. introduction teacher education is the environment for the learning and instruction of prospective teachers. its structure, components, and contents shape the development of relevant competences which enable prospective teachers to be effective in the classroom. these competences comprise cognitive, motivational, volitional, and social abilities and skills necessary for effective teaching (weinert, 2001). but its relevance is questioned because respective research, characterised by inconclusive results, does not offer explanations about the reasons why certain teacher education programmes are more effective than others in the development of relevant competences (boyd, grossman, lankford, loeb, & wyckoff, 2009; harris & sass, 2011; yeh, 2009). one reason for the lack of explanations can be found in the way research assesses the effectiveness of teacher education. most studies compare graduates from different teacher education programmes with regard to differences in the achievement of students in schools; this approach has relatively high demands concerning methodology and conceptualisations of teacher education (boyd, grossman, hammerness, lankford, loeb, ronfeldt, & wyckoff, 2012; morge, toczek, & cakroun, 2010). however, this dominant approach and the conceptualisations of teacher education in these studies do not fully grasp the complexity of teacher education, especially the interplay between different components and the learning and instruction of prospective teachers. four specific aspects illustrate the problems associated with the way research currently investigates teacher education effectiveness. the first two aspects are directly related to teacher education conceptualisations. first, many studies conceptualise teacher education as an individual teacher attribute. they use narrow sets of variables, for example the degree and certification status, as proxies for competences which teachers bring into the classroom (harris & sass, 2011). even structural features or policies of teacher education, for example the selection procedures or the structure of learning opportunities, are considered such individual teacher attributes (little & bartlett, 2010). these kinds of conceptualisations may not adequately reflect the relation between organisational aspects of teacher education and the behaviour of individuals, e.g. the use of learning opportunities by prospective teachers during initial teacher training. what happens at the level of the individual prospective teacher, that is, his learning processes, is embedded in the structure of teacher education. harris and sass (2011) labelled this aspect the “inherent selection problem”. second, most studies directly relate the aforementioned narrow sets of indicators for teacher education to the achievement of students in schools. however, as konold, jablonski, nottingham, kessler, byrd, imig, berry, and mcnergney (2008, p. 310) argue, “[…] there is little to be learned by examining the long jump between teacher characteristics and pupil learning. […]”. few studies take into account the full complexity of the relation between teacher education, teacher characteristics (such as their competences), teacher behaviour, and student achievement. especially the relation between teacher behaviour and student achievement is neglected (connor, son, hindman, & morrison, 2005). an effect size of 0.91 for teacher behaviour measured by classroom observations on student achievement, found by schacter and tum (2004), illustrates the importance of teacher behaviour. the „long jump‟ disregards this relation, and does not take into account the distinction between teacher quality (characteristics teachers possess) and teaching quality (their teaching practice). thus, it hinders the identification of teacher characteristics which are important for effective teaching. the other two aspects are related to potential sources of bias in current estimates of the effectiveness of teacher education (harris & sass, 2011). third, one source of bias is the variation in the development of relevant competences across teacher education programmes (boyd, et al., 2009). this variation may not be attributed only to a better provision of opportunities to learn, but also to a better selection of prospective teachers (denzler & wolter, 2009). structural features of the selection procedures may shape unobserved characteristics of prospective teachers which influence their learning (kennedy, 1998). individual conceptualisations of teacher education lack explanatory power with regard to such organisational aspects. fourth, another source of bias is the nonrandom allocation of teachers to schools. a prominent manifestation of this problem is positive matching. students in schools with high socioeconomic status have better access to highly qualified teachers (in terms of paper qualifications), compared to students in schools with a lower socioeconomic status (luschei & carnoy, 2010; loeb, kalogrides, & beteille, 2012). only few studies investigate relevant structural features of the teacher labour market with regard to their influence on teacher distributions (goldhaber, 2007; c. könig & r.mulder 28 | f l r winters, dixon, & greene, 2012). current individual conceptualisations of teacher education do not allow for explanations of the development of positive matching, because they address this problem when the allocation of teachers to schools has already happened. hence, it remains unknown why teachers bring their competences into schools and classrooms in such a systematic way. in this paper we address these issues and argue that, with a change in perspective on teacher education, some of them may be attenuated. this change in perspective is based on three premises: (1) a rearrangement of teacher education and teacher characteristics within the concept of teacher quality, accompanied by a clear distinction between teacher quality and teaching quality (goe & strickler, 2008). (2) an organisational approach to teacher education modelling teacher education as a system, which focuses on structural features relevant for the selection of teacher education candidates and prospective teachers, the development of relevant competences, and for the allocation of teachers to schools. (3) a change in the notion of teacher education effectiveness, which is due to the rearrangement of the teacher quality concept and the organisational approach to teacher education. the aim is to develop an organisational model of teacher education which allows researchers to take into account (1) the relation between teacher education and its context, as well as (2) the interplay between teacher education and prospective teachers. the development is oriented along the ecological framework of teacher education proposed by zeichner and conklin (2008) and specifically focuses on the admission process and the institutional and labour market context of teacher education. grossman and mcdonald (2008) identify these contexts as being important influences on the policy and practice of teacher education, and argue that in order to gain new insights research should incorporate these contextual conditions. moreover, the model provides a theoretical basis for explanations of learning and instruction of prospective teachers which is embedded in a teacher education system (zeichner, 2005). given the lack of research on organisational level we make use of system and organisational theories in order to characterise teacher education as a system. however, the reliance on these theories might be an advantage because, as grossman and mcdonald (2008) state, broadening the theoretical basis of research on teacher education might facilitate new insights and explanations of teacher education policy and practice. eventually, the model will provide researchers with a new theoretical basis for research in order to reach a better understanding of learning and instruction of prospective teachers, because it illustrates the connections between different (organisational and individual) levels and systems, as well as the interdependencies of individual and organisational learning. these new insights might further be used for policies aimed at the facilitation of learning and instruction of prospective teachers. 2. the prerequisite rearranging components of the teacher quality concept goe and strickler (2008) conceptualise teacher quality as a multidimensional concept consisting of three interrelated dimensions. they conceive of teacher qualifications (understood as degrees, majors, and other paper qualifications) and characteristics (such as their competences) as input variables, teacher behaviour as process variable, and teacher effectiveness as output variable which is commonly measured by standardised student test scores. in accordance with other authors they emphasise that teacher quality and teaching quality are two different aspects, and that they should be modelled accordingly (goe & strickler, 2008; konold et al., 2008). however, as we already mentioned in the introduction, many studies on the relation between teacher education and student achievement disregard this distinction. the interrelations between the different concepts are as follows. teacher qualifications and characteristics (such as their competences) have an influence on the behaviour of teachers, that is, what they do and can do in the classroom (teaching quality). following weinert‟s (2001) definition of competence, teacher characteristics constituting teacher quality comprise cognitive abilities and skills, for example knowledge about and mastery of subject-didactics and a repertoire and understanding of multiple models of teaching, as well as motivational, volitional, and social aspects such as commitment to a continued professional development after initial teacher training, love of children, collaboration with colleagues, and reflection over practice (hopkins, 2008). teacher quality translates into teaching quality. at the same time, with teaching being an experience good and social practice (jovanovic, 1979), teaching quality influences teacher quality. for c. könig & r.mulder 29 | f l r example, reflection over practice, collaboration with colleagues, and a high commitment to continued professional development enables teachers to refine their practice and to further develop their competences after their initial teacher training. eventually, the interplay between teacher and teaching quality is an important influencing factor for student achievement and, consequently, directly related to student achievement. what becomes obvious is that teacher characteristics (such as their competence) have no direct relation to student achievement. their effect on student achievement is mediated by the respective teacher behaviour, that is, they only have an indirect effect on student achievement. this indirect relation is also disregarded by many studies (for example marshall & sorto, 2012). differences in teacher characteristics may lead to differences in what teachers are able to do in the school and in the classroom, and in turn to differences in student achievement. as of yet the specifics of these pedagogical mechanisms are unclear (baumert, kunter, blum, brunner, voss, jordan, klusmann, krauss, neubrand, & tsai, 2010). the unclear picture is due to a negligence of the indirect effect of teacher quality on student achievement. hence, a first prerequisite for the change in perspective on teacher education involves acknowledging this indirect relation. this is accompanied by shifting the focus to the relation between teacher and teaching quality. this may be a way to identify specific teacher characteristics which are relevant for effective teaching. teacher qualifications and characteristics are frequently used interchangeably. however, they are two distinct concepts. teacher qualifications are frequently used in studies as proxies for what the teacher did during initial teacher training (harris & sass, 2011). but teacher characteristics, such as their competence, are a consequence of teacher qualifications, that is, of what they did during initial teacher training. in other words, what teachers did during their initial teacher training, and why, has consequences for what they bring into the school and the classroom, and where. jackson (2010) showed that the quality of teacher-student matches accounts for up to 40 percent of what is usually attributed to a teacher effect on student achievement. hence, the second prerequisite involves a clear distinction between teacher qualifications and teacher characteristics. however, with individual level conceptualisations of teacher education, which mix up teacher qualifications and teacher characteristics, we cannot explain what a prospective teacher actually does during initial teacher training and why, where he ends up teaching, what he is able to do in the classroom, and eventually how his behaviour affects student achievement. having teacher education disentangled from teacher characteristics, and having it identified as starting point for the complex chain between the resulting teacher characteristics (such as their competence), teacher behaviour, and student achievement, we are now in a position to model teacher education as a system of structured learning opportunities, including structural elements governing the selection of prospective teachers and the allocation of teachers, which is embedded in multiple institutional contexts (zeichner, 2006). 3. a different perspective – teacher education as an open system the organisational model of teacher education described in this section is based on open systems theory (katz & kahn, 1978). despite being a rather old model, up to this date it still remains “the most systematic introduction of open system concepts into organisation theory” (scott & davis, 2007, p. 90), and is furthermore the theoretical basis for much of current organisational research (schneider & somers, 2006; martz, 2013). katz and kahn (1978) were among the first recognising the dependency of organisations and their environment, as well as the linkage between psychological and structural/economic aspects of organisations. compared to other currently used open system models, for example contingency theory (lawrence & lorsch, 1967), it is the aforementioned linkage between individual and organisation which makes open systems theory an appropriate framework for teacher education systems. compared to current further developments of open system models, for example complex adaptive systems (stacey, 1995), open systems theory provides a more accessible framework due to the comprehensiveness of its core components. however, the main reason for choosing open systems theory was the fit of its theoretical propositions with the characteristics of teacher education systems (bess & dee, 2008) also show the c. könig & r.mulder 30 | f l r usefulness of this theory for educational organisations in their application of open system theory to higher education). first, it explicitly takes into account the relations and exchanges between different systems. this is important because the teacher education system is not an isolated entity, but is embedded in multiple contexts, for example higher education and the teacher labour market (grossman & mcdonald, 2008). in this part of the framework we are able to model which individuals choose teacher training, and where teachers bring their characteristics (such as their competence) to the school and the classroom. second, it explicitly takes into account the dependencies and interplay of system and prospective teachers. this is important for modelling the use of available learning opportunities by prospective teachers. in this part of the framework we integrate what the prospective teacher does during initial teacher training. 3.1 teacher education from the point of view of open systems theory an open teacher education system consists of a sequence of structured learning opportunities provided to prospective teachers within the system. the sequence and structure of the learning opportunities constitute an environment where the learning of prospective teachers is situated in a gradually growing participation in teaching practice (korthagen, 2010). the active use of these opportunities leads to the development of competences required for effective teaching. the use of learning opportunities by prospective teachers is labelled as, in open system terms, patterned activities of individuals and describe the core of the interplay between system and prospective teachers (katz & kahn, 1978). thus, what happens within the teacher education system is seen as an active developmental process, rather than just a transmission of declarative knowledge (zeichner, 1983). what prospective teacher do, and how successful their professional development is during initial teacher training depends on the characteristics they bring into the teacher education system. at the same time, the learning opportunities provided by the teacher education system require certain individual characteristics. if teacher education candidates or prospective teachers do not meet these requirements, the utilisation of learning opportunities, as a part of their professional development, becomes suboptimal and may even get cancelled prior to graduation (blömeke, 2009). thus, for an open teacher education system control over entry is essential (which is also called boundary maintenance; scott & davis, 2007). the selection function plays a key role in this regard, and is defined as the selection and sorting of teacher education candidates and prospective teachers (musset, 2010; van de werfhorst & mijs, 2010). it is based on the characteristics of the candidates and prospective teachers. an optimal selection function avoids adverse selection in terms of characteristics which hinder a successful utilisation of learning opportunities as a part of the professional development of prospective teachers. given the connection of an open teacher education system to its context (scott & davis, 2007), we have to consider what happens immediately after initial teacher training. the degree to which prospective teachers successfully use the learning opportunities during initial teacher training influences the competence they bring into schools and classrooms. this is a second component of the connection between teacher education and the education system. this allocation function is defined as the assignment of teachers to schools (parsons, 1951), which has long been based on the assumption that schools and teaching position are equivalent across districts and regions (johnson & kardos, 2008). however, jackson (2010) could show that there are teacher-school combinations which lead to better student achievement. thus, it matters where teachers bring their competence into the classroom. an optimal allocation function provides teacher-school matches that minimise teacher turnover and attrition. in sum, the general characteristics of an open teacher education system closely resemble the three common functions of education systems, which constitute an input-transformation-output-model (kast & rosenzweig, 1972): the selection and sorting of candidates and prospective teachers (input/selection), the provision of learning opportunities for students situated in a gradually growing participation in teaching practice to develop relevant competences (transformation/instruction), and the allocation of qualified teachers to schools (output/allocation). c. könig & r.mulder 31 | f l r 3.2 the selection and allocation functions in order to establish and maintain the selection and allocation processes, the open teacher education system develops respective structural elements (katz & kahn, 1978; wang, coleman, coley, & phelps, 2003). these structural elements are arranged in subsystems governing the selection and sorting of prospective teachers, and the allocation of teachers to schools. these structural elements comprise institutional structures and administrative regulations for control over and socialisation of prospective teachers and teachers (maaz, hausen, mcelvany, & baumert, 2006). they allow screening out individuals when they do not meet the requirements of teacher education or a given teaching position in a school. both functions are closely connected to the context of the teacher education system, because they govern the transitions of individuals into and out of initial teacher training. thus, the arrangements of structural elements can be understood as transition systems (van der velden & wolbers, 2007). as such, they are means for the teacher education system to react to policy changes in the immediate context, namely the education system and the teacher labour market. an example for such reactions is a change in the selection mechanisms of a teacher education system given a shortage of teachers in the teacher labour market (blömeke, 2006). 3.2.1 general characteristics of the selection function the selection function governs the admission of teacher education candidates at entry into, and the sorting of prospective teachers within the teacher education system. by means of the aforementioned control and socialisation elements, the selection function provides information about (1) the aptitude of teacher education candidates for teaching, and (2) about the success of prospective teachers in their use of learning opportunities. moreover, socialisation mechanisms initiate the transfer of professional role expectations and norms from teacher education to the prospective teacher and support the professional development of the prospective teacher (saks, uggerslev, & fassina, 2007). this information can be used by prospective teachers in order to judge his attitude to and aptitude for teaching. furthermore, it enables prospective teachers to reflect on their practice in order to determine how to improve his teaching. moreover, the information provided by the selection function serves also as relevant feedback for the system for admission and progression decisions, in order to reduce the variability in the use of learning opportunities, which is due to variability in individual characteristics (scott & davis, 2007). with its control and socialization mechanisms, the selection function serves both the prospective teachers and the teacher education system in determining if a given prospective teacher can progress to the next developmental stage. while the information provided by the function is at first only a rough estimate of how well a given candidate might do, the information becomes more detailed when the actual development of the prospective teacher is assessed. it is important to note that it is only possible to select individuals who (are able to) make themselves available (grodsky & jackson, 2009). thus, variability in individual characteristics can be found either in the candidate pool or the prospective teachers. the structural elements constituting, and in turn influencing the success of the selection function, can be assigned to and described with three dimensions. first, the capacity of the teacher labour market influences the number and characteristics of the candidates. this comprises the accessibility of teacher education and the attractiveness of teaching. second and third, the comprehensiveness of available information about candidates and students and the level of integration of students into teaching influence the number and the characteristics of the prospective teachers. 3.2.2 structural elements of the selection function we begin with the structural elements constituting the capacity of the teacher labour market. the theoretical rationale of the respective structural elements is based on rational choice and supply and demand models (sicherman & galor, 1990; ehrenberg & smith, 2011). given that initial teacher training is an educational choice among others they postulate that individuals analyse educational alternatives by weighing costs against benefits. when the costs of a given educational alternative are higher than individual resources, individuals will opt for another alternative. rational choice models emphasise two core aspects relevant for c. könig & r.mulder 32 | f l r characteristics of the candidate pool: structure and status. based on these core aspects, the length and level of initial teacher training and the occupational status of teaching are structural elements of the capacity of the teacher labour market. while the influence of the length and level is ambiguous, a high occupational status of teaching attracts a greater number of teacher training candidates and increases the candidate pool. countries with a highly attractive teaching profession do not have teacher supply problems (schwille & dembele, 2007). however, with an increased candidate pool it is more likely that the variability in individual characteristics is increased as well. furthermore, characteristics of the student population affect the number of available teaching positions, that is, the demand of teachers. for example, an increased number of students in the education system affects the student-teacher ratio, which in turn influences teacher demand. while this aspect has no direct influence on the candidate pool, it affects the control mechanisms at entry into initial teacher training. educational decisions and the selection process are characterised by an asymmetric distribution of information (van der velden & wolbers, 2007). imperfect information about candidates and prospective teachers is problematic for systems, because they rely on signals (stiglitz, 1975). lack of information increases the risk of admitting and progressing teacher education candidates and prospective teachers who are not successfully using the learning opportunities, or else show an insufficient development. hence, structural elements influencing the comprehensiveness of information available to the teacher education system are admission and assessment procedures, which are based on respective criteria. these criteria determine which individual characteristics are required for entry into initial teacher training and for teaching. students with required characteristics utilise learning opportunities successfully and are more likely to graduate. while the admission procedures are implemented in order to collect information about teacher education candidates, the assessment procedures are implemented in order to monitor prospective teachers with respect to their use of learning opportunities as part of their professional development. moreover, the assessment procedures serve as feedback and possibility for the prospective teachers to reflect on their development and teaching practice. the comprehensiveness of information increases if the admission and assessment procedures exhibit certain characteristics. according to baartman, bastiaens, kirschner and van der vleuten (2006) the characteristics of such assessment procedures within a competence-based approach to teacher education comprise fitness for purpose, comparability and reproducibility of results, acceptability and transparency. moreover, the fairness, cognitive complexity, meaningfulness, and authenticity of the procedures are relevant, besides their costs and efficiency and their consequences (admission and progression decisions). especially admission procedures are closely linked to the demand of teachers. the literature frequently discusses solutions to teacher shortages in form of reduced entry requirements for initial teacher training (blömeke, 2006). the sequence, rigor, and the aforementioned quality-characteristics of procedures and their criteria increase the comprehensiveness of information about candidates and prospective teachers. this is especially important when the candidate pool is large. socialisation mechanisms serve as means to help prospective teachers to take on new roles and simultaneously stress the social aspects of the learning processes. these structural elements reduce the uncertainty of students about expectations and requirements about teaching when entering teacher education. furthermore, the respective structural elements situate the learning of prospective teachers in a social environment, where they are guided and supported in their professional development (korthagen, 2010). one structural element is internal support. it gives access to structured forms of support, either with guidance by experienced teachers or sequenced in clearly defined courses. the other is field experience. it describes opportunities for field experiences prior to entering the teaching profession, and directly influences the transfer of professional role expectations and norms. the level of integration of the selection function is high when a prospective teacher receives frequent internal support, as well as several possibilities to make relevant field experiences. the structural elements of the selection function and their assignment to their respective dimensions are summarised in table 1. 3.2.3 general characteristics of the allocation function the selection function governs the transition of trained teachers from initial teacher training into the teaching profession. thus, it is related to the allocation of teachers to schools. by means of the c. könig & r.mulder 33 | f l r aforementioned control and socialisation elements, the allocation functions provides information (1) for schools about the characteristics of trained teachers, and (2) for trained teachers about characteristics of teaching positions in schools. the socialisation mechanisms initiate the transfer of school specific role expectations and norms. they serve as information for schools about how well a trained teacher is able to integrate into the specific school context. this is a relevant feedback for schools in order to make recruitment decisions. these decisions result in teacher-school matches (lankford & wyckoff, 2010). similarly, the information is at first only a rough estimate of the characteristics of teachers, but becomes more detailed by an increasing amount of time between the first assignment and the definite recruitment decision (liu & johnson, 2006). due to varying success regarding the use of learning opportunities variability in teacher competences is likely. for example, despite having obtained the same degree, trained teachers still can vary in their acquired cognitive, motivational, volitional, and social skills (van der velden & wolbers, 2007). thus, it is difficult for schools to distinguish between teachers who are suited for a given teaching position, and those who are not. hence, the structural elements constituting, and in turn influencing the success of the allocation function, can be assigned to and described with three dimensions. the first dimension is control over the recruitment process. this dimension includes the level of control, as well as the actual utilisation of the level of control with adequate recruitment procedures. a more direct control over recruitment, combined with various recruitment measures may facilitate staffing (liu & johnson, 2006). the control over the recruitment process is directly connected with the second dimension, namely the comprehensiveness of information which is available to schools and teachers about each other. with an increased comprehensiveness of information it is possible to make more informed recruitment decisions. third, the level of integration of teachers into schools influences the smoothness of the transition into the specific teaching position. 3.2.4 structural elements of the allocation function the starting point are structural elements constituting the comprehensiveness of available information about teachers and their characteristics. similarly to the selection function, the allocation process is characterised by an asymmetric distribution of information (van der velden & wolbers, 2007). signals for teachers‟ characteristics and structural factors of the recruitment process attenuate the lack of information (stiglitz, 1975). lack of information about teacher characteristics increases the risk of recruiting the “wrong” teacher and increases the risk of teacher turnover. signals are provided by certification requirements which trained teachers have to fulfil. however, certification requirements and respective teacher test scores are only weak signals of teachers‟ knowledge and skills (goldhaber, 2007). thus, another structural element for information about beginning teachers is probationary periods. with probationary periods, where teachers are monitored regarding their performance, the definite recruitment decision can be delayed, and more information about a teacher can be collected (staiger & rockoff, 2010). however, it is important to note that not only the length of the probationary period is relevant, but also its implementation. probationary periods may be successful only if they provide trained teachers with a well-established and supportive environment (oecd, 2011). examples of respective aspects are, for example, faculty collaborative periods, meeting with supervisors, classroom assistance, or a reduced workload (ingersoll & strong, 2011), within which teachers are enabled to reflect on their practice. probationary periods may be combined with induction measures. in sum, the comprehensiveness of information is high if the allocation function includes certification requirements combined with elaborate probationary periods for teachers. however, the influence of the level of information on the allocation process depends on the control over the recruitment process. as mentioned before, control over the recruitment process comprises the level of and utilisation of this control. the level of control is indicated by the degree of school autonomy regarding recruitment decisions. a direct control over recruitment decisions might facilitate the staffing of schools (liu & johnson, 2006). it may be hindered when there are central authorities or union regulations governing the recruitment process. such regulations may not adequately consider school specific needs regarding personnel and can be understood as constraints interfering with school based recruitment. thus, the level of control over recruitment decisions can be distinguished between school based or local recruitment, a recruitment controlled by regional or central authorities, or a recruitment which is coordinated c. könig & r.mulder 34 | f l r between local and central authorities. however, the level of control alone is not sufficient to characterise control over the recruitment process. several studies have found that although schools have a high degree of autonomy in staffing decisions, they only utilise a small set of recruitment procedures during recruiting teachers (balter & duncombe, 2008; staiger & rockoff, 2010). respective recruitment procedures might include for example interviews and supervised sample lessons. in sum, control over the recruitment process is adequate only if a school-based recruitment is complemented by a variety of recruitment procedures. at the same time, such control has a positive influence on the comprehensiveness of information about trained teachers (liu & johnsons, 2006). table 1 the functions, their dimensions, and their respective structural elements function dimension structural elements context capacity of the teacher labour market length of teacher education level of teacher education occupational status of teaching student population selection comprehensiveness of information about candidates & prospective teachers admission procedures assessment procedures admission criteria assessment criteria level of integration of prospective teachers internal support field experiences allocation control over the recruitment process school autonomy union regulations recruitment procedures comprehensiveness of information about trained teachers certification probationary periods level of integration of teachers into schools teacher mentoring teacher induction socialisation mechanisms serve as means to help teachers to take on school-specific roles and norms. first, the beginning teacher learns the requirements of a role or teaching position (functional aspect); second, he integrates into the social structure of the school (inclusion aspect). over time they get accustomed to the specific organisational characteristics and can adapt to them. similarly to the selection function, these structural elements reduce the uncertainty of teachers about expectations and requirements when they start teaching in a given school. moreover, they offer possibilities for teachers to reflect on their practice in order to improve their teaching. as such the socialisation mechanisms are means to foster teacher professional development after initial teacher training (ingersoll & strong, 2011). structural elements related to the level of integration are teacher induction and teacher mentoring. they are means to make the teachers acquainted to the specific characteristics of a given school. it includes a formalised system to support teachers. teacher mentoring is personal guidance provided by a senior teacher at a school. it varies from single meetings to formalised programmes involving frequent communications between teacher and mentor. teacher induction and mentoring also influences teacher retention, thus decreasing teacher shortages and turnover (wang, odell, & schwille, 2010). schools are more frequently required to provide teachers with school-specific learning opportunities (ingersoll & strong, 2011). the level of integration varies according the comprehensiveness of induction and mentoring measures. the structural elements of the allocation function and their assignment to their respective dimensions are summarised in table 1. c. könig & r.mulder 35 | f l r 4. a change in notion – a different view of teacher education effectiveness we already mentioned that the change in perspective on teacher quality and teacher education requires a different notion of teacher education effectiveness. morge et al. (2010) distinguish three levels of validation of teacher education, depending on the specific outcome variable which is evaluated. the first level comprises teacher thinking and teacher knowledge as primary outcome. the effectiveness of teacher education is assessed by the level of cognitive and non-cognitive characteristics of teachers, that is, their knowledge and motivational, volitional, and social skills which they acquired during initial teacher training. however, at this first level the link between these characteristics and the instructional practice of teachers is not included (morge et al., 2010). the second level includes this link, i.e. the effectiveness of teacher education is assessed with respect to the behaviour of the teachers. while the first level only allowed to ask what teachers know, the second level extends this question to what they are able to do in the school and in the classroom. the third level further extends the concept of teacher education effectiveness. here, teacher education effectiveness is a question of what teacher is able to do in schools and in the classroom, and how this affects student achievement. current notions of teacher education effectiveness involve primarily the third level of validation. however, with the narrow teacher education conceptualisations which directly relate distal variables to student achievement, we cannot expect to gain reliable estimates of the effect of teacher education on student achievement (konold et al., 2008). furthermore, we cannot investigate if teachers who participated in initial teacher training behave in ways which positively affect student learning (morge et al., 2010; konold et al., 2008). the organisational model of teacher education as an open system, however, may be a way to investigate this question. in this regard, a change in notion of teacher education effectiveness, that is, a focus on the second level of validation, might be a necessary step. in the following we illustrate this change in notion and focus. the starting point is teacher competence as outcome of teacher education. thus, we focus on the first level of teacher education validation. teacher competence depends on the utilisation of learning opportunities by prospective teachers. as already mentioned, the learning process situated in a gradually growing participation in teaching practice requires specific individual characteristics (tillema, 1994). teacher education is effective if it provides learning opportunities, based on specific curricula, which provide prospective teachers with the possibility to develop competences necessary for effective teaching. given that the characteristics of prospective teachers depend on the effectiveness of the selection function in sorting them, the notion of teacher education effectiveness is extended: a teacher education system is only effective if (1) it provides prospective teachers with information about their development, with which they can reflect on their practice, and additionally if (2) the system screens out prospective teachers who are likely to fail. besides this individual outcome of teacher education, we also have an organisational outcome. a successful utilisation of learning opportunities by students implies higher success rates (gansemer-topf & schuh, 2006). hence, a comprehensive notion of teacher education effectiveness includes selection effects on the use of learning opportunities and, thus, the professional development of prospective teachers, and an organisational aspect in terms of success rates. moreover, the competences of prospective teachers are related to their teaching practice. in other words, teacher quality may only become visible through the associated teaching quality (mulder, messmann, & gruber, 2009). this means that in order to assess teaching quality it is necessary to consider the competences of the (prospective) teachers, and vice versa. classroom observations during initial teacher training, along with guided support by experienced teachers and room for reflection on their teaching practice, may facilitate an assessment of prospective teachers‟ readiness to teach and teaching quality, given the consensus on effective teaching practices (akiba, letendre, & scribner, 2007). however, classroom observations require the teachers‟ reflections on their teaching, that is, explications of the reasons why they did what they did. this may be a way to unravel the connection between teacher and teaching quality, and thus a possible clarification of the mechanisms with which teachers translate their competence into effective teaching. including what a teacher is able to do in a real classroom in a school, and how this affects student achievement in the concept of teacher education effectiveness is difficult. each school, even each classroom, is a unique social system (johnson & kardos, 2008). hence, specific contextual characteristics of schools, c. könig & r.mulder 36 | f l r for example their facilities and equipment, or the leadership style of the principal, may influence how well teachers are able to translate their knowledge into effective teaching. moreover, where teachers bring their characteristics into schools and in the classroom depends on the specific characteristics of the allocation function. each teacher effect on student achievement involves a complex interplay between recruitment decisions, school and classroom characteristics, and the behaviour of the teacher in the schools and in the classroom. given that it is still unclear how teachers translate their knowledge into effective teaching (baumert et al., 2010; croninger, rice, rathbun, & nishio, 2007), it is questionable if an effect of teacher education on student achievement can be identified. as a consequence, the assessment of teacher education effectiveness remains a question of the development of competences necessary for effective teaching, and thus remains on the second level of validation. figure 1. the organisational model of teacher education as an open system. rectangles depict the dimensions of the selection and allocation function, as well as contextual conditions in the education system/teacher labour market. black arrows illustrate the transition of an individual through teacher education into schools, from teacher education candidate over prospective teacher to a trained teacher in a school. gray arrows and boxes show the consequence of the use of learning opportunities by prospective teachers on their competence and success rates, and the consequences of specific teacher distributions (teacher turnover and positive matching). from an organisational point of view it is nevertheless possible to relate the allocation function to specific manifestations of teacher distributions, such as the positive matching between teachers and schools. it is a peculiarity of the allocation in the context of education systems that a successful allocation is not only a question of balancing supply and demand, but to a greater degree a question of students‟ equal access to highly qualified teachers. hence we have an organisational indicator for the effectiveness of the allocation function: the degree to which its structural arrangement of elements attenuates positive matching of teachers to schools. c. könig & r.mulder 37 | f l r in sum, based on the changes in the teacher quality concept and the organisational perspective on teacher education as an open system, the notion of teacher education effectiveness receives a narrower, but more meaningful and distinct focus. the inclusion of organisational indicators for the effectiveness of the selection and allocation functions allow for an investigation of teacher education effectiveness on a different level. an interesting aspect in this regard is the relation between higher success rates of the teacher education system and the impact of the allocation function on positive matching, because higher success rates imply a higher number of teachers available for allocation. hence, the organisational model allows investigating the relation between the functions as well. the complete organisational model is visualised in figure 1. 5. discussion – the model’s value in research on teacher education in this paper we addressed four shortcomings of current research on the relation between teacher education and student achievement, namely the conceptual, the complexity, the inherent selection, and the non-random allocation problem (konold et al., 2008; harris & sass, 2011). the aim was to develop an organisational model of teacher education which provides researchers with a new, alternative perspective on teacher education practice. this perspective enables researchers to investigate the relation between teacher education and its context (for example the teacher labour market and the education system), the interaction of different systemic levels, as well as the interdependencies of individual and organisational development. the development was based on three specific premises. first, an alteration of the input variables of the teacher quality concept. this involved a clear distinction between teacher education as an antecedent of teacher characteristics, that is, teacher education directly influences teacher competences relevant for teaching. second, a change in perspective away from teacher education as an individual teacher characteristic to a model of teacher education as an open system. within this model, we outlined the role of the selection function for prospective teachers‟ professional development, and the role of the allocation function for different manifestations of the non-random allocation of teachers to schools, for example positive matching. third, as a consequence of the change in perspective, we illustrated an associated change in the notion of teacher education effectiveness. this concept was refocused on the development of competences of prospective teachers, and extended with two organizational indicators of effectiveness. this narrower focus is necessary because of the complex interplay between school and classroom characteristics and what teachers are able to do in the school and in the classroom, which may hinder the identification of a definite teacher education effect on student achievement. the relative underspecification of the learning opportunities in the model is intentional. in contrast to the elements of the selection and allocation functions, it is difficult to identify generic elements of learning opportunities which are comparable across institutional or national settings. although there is some convergence in the design of learning opportunities, there is still a great variety in elements of learning opportunities (paine & zeichner, 2012). moreover, research shows that some of the more generic characteristics such as the length and structure of teacher education are unrelated to teacher education effectiveness (zeichner, 2006). however, in order to make the model useful for, for example, cross-country comparisons it is necessary to keep the model as generic as possible. the underspecification of the learning opportunities provided by a teacher education system might be interpreted as an opportunity for researchers to take into account country-specific characteristics of the learning opportunities in their own studies. hence, researchers are able to fill this gap in the model with characteristics of learning opportunities in their respective samples. the model as a whole imposes high requirements on the collection, amount, and quality of data. this limitation applies to all aspects mentioned in this section. although recent international comparative studies such as teds-m and talis provide new databases, available data might not be sufficient to test the model as a whole. thus, it might be more reasonable to concentrate on specific aspects of the model, such as the relation between selection and student characteristics, the relation between allocation and positive matching, or the relation between student teachers and their use of learning opportunities. nevertheless, the model c. könig & r.mulder 38 | f l r outlined in this paper might serve as a foundation for more elaborate and comprehensive data collection in future studies on teacher education systems. we hope that the organisational model of teacher education will provide a theoretical basis which initiates new research leading to new insights and a better understanding of teacher education policy and practice, especially with regard to the identification of teacher characteristics relevant for teaching, the selection of teacher education candidates and prospective teachers, and the positive matching between teachers and schools. in the following sections we will discuss the usefulness of our model in the context of three possible areas of research. 5.1 identification of teacher characteristics relevant for effective teaching we already mentioned in the introduction that research on the relation between teacher education and student achievement is unsuccessful at identifying teacher characteristics relevant for effective teaching. besides the inherent selection problem, that is, the unobserved characteristics which influence what a teacher did during his initial teacher training, this is further due to the distal conceptualisations of teacher education used in current studies. these conceptualisations, for example the certification status of teachers, are selected because of their relevance for policies concerning the teacher labour market (goldhaber, 2007; harris & sass, 2011). however, these conceptualisations might gain meaning if the aforementioned unobserved characteristics are made observed, and their relations to effective teaching are established (this is in line with the focus on the second teacher education validation level described in section four). we argue that our model can provide a means in order to accomplish these tasks. our model explicitly states relations between characteristics of prospective teachers and their use of learning opportunities provided by the system. with teaching being an experience good, the identification of relevant characteristics requires accurate information about what prospective teachers are able to do in the classroom, that is, classroom observations of prospective teachers which are supported by guided reflection on teaching practice (morge et al., 2010). these classroom observations and possibilities for reflection may be integrated in a more refined concept of the assessment procedures. the authenticity of the assessment procedures may be the core aspect with regard to the identification of relevant characteristics, because it is a more direct way of assessing how well prospective teachers are able to translate the contents of their initial teacher training into effective teaching behaviour (darling-hammond & snyder, 2000). the performance scores derived from these observations, as well as information about the reflections of the teachers, may then be related to a set of characteristics prospective teachers possess. the identification of relevant teacher characteristics further has positive consequences for the selection and sorting of prospective teachers during initial teacher training. the selection and sorting of teacher education candidates and prospective teachers are still based on rather gross measures, such as the grade point average or subject-specific grades in secondary education (blömeke, 2009). with an increased authenticity of assessment in the context of the selection function, and with the associated more accurate information about prospective teachers, the identified characteristics can in turn be used as more refined and accurate admission and assessment criteria. hence, our model not only allows addressing the inherent selection problem on individual, but also on organisational level. it has to be noted that the identification and use of the identified characteristics is an iterative process and requires a significant amount of time, that is, longitudinal models. however, our model is flexible enough to allow for such extensions. the identified characteristics of prospective teachers may be of limited use for the identification of what a teacher is able to do in a school and in a real classroom, given school-specific contexts influencing their practice. 5.2 research on teacher distributions and the teacher body given the explicit modelling of the allocation function, which is integrated into our model of teacher education as an open system, researchers are enabled to investigate consequences of different approaches to allocating teachers to schools. for example, it may be investigated how certification requirements affect the pool of teachers who choose to teach. there are already studies concerning this problem (for example c. könig & r.mulder 39 | f l r angrist & guryan, 2008). however, they investigate this feature of the allocation function isolated from other relevant features, and isolated from the teacher labour market context. an isolated investigation of these features may not suffice for explanations of different teacher distributions. for example, boyd et al. (2012) conclude that, while some teacher education programmes produce teachers with higher student achievement gains than others, these effects are eliminated when their attrition rate is taken into account. another example are the results a simulation study conducted by rothstein (2012). it showed that changing the quality of the teaching force through selection is only successful if at the same time teacher evaluation systems and increased teacher salaries are introduced. this illustrates the need for possibilities for an integrated rather than isolated investigation of selection and allocation effects, which our model provides. moreover, schools depend on the amount of available information about teachers in order to make informed recruitment decisions. these decisions seem to rely on only weak and noisy signals (goldhaber, 2007). thus, it is frequently argued that for an acquisition of reliable specific information, an assessment of teachers based on actual classroom performance is necessary (goldhaber & liddle, 2011). staiger and rockoff (2010) suggest that tenure should be delayed until a sufficient amount of information is collected. as long as indicators of teacher education do not adequately capture what teachers do during their initial teacher training (cf. the respective description in section 5.1), mismatches between teachers and schools are to be expected which lead to teacher turnover. in light of the change in the notion of teacher education effectiveness, a stronger reliance on actual classroom performance of teachers in the context of recruitment seems reasonable. our model allows for an investigation of the influence of different approaches to recruiting teachers and their relation to teacher turnover, taking into account contextual conditions of the teacher labour market. it has to be noted that the model in his current state captures only the structural prerequisites of recruitment decisions. however, our model can be easily extended to include the individual recruitment (or transfer) decisions of teachers and principals within the context of a given configuration of an allocation function. the relations and research questions outlined in the previous sections may also be investigated by cross-country comparisons of teacher education systems, for example a comparison of credential-based and information-based allocation functions (van de werfhorst, 2011). comparisons of different approaches to allocating teachers to schools need to consider not only quantitative, but also qualitative aspects of, for example, recruitment procedures or probationary periods. these qualitative aspects not only include the variety of the different procedures, but also the actual utilisation of these procedures by principals, school boards, or other entities responsible for staffing decisions. thus, when collecting data, researchers may not only rely on institutional data provided by administrative datasets or official documents, because this might only cover the „espoused allocation‟. in order to gain a complete picture of the qualitative aspects, it might be necessary to actually ask principals or school boards about the actual utilisation of the procedures in order to capture the „allocation in use‟ (a similar distinction can be found in cannata, 2010). covering only one of these two procedures may lead to biased estimates of the relation between allocation approaches and teacher distributions. 5.3 cross-country and cross-institutional comparisons of teacher education systems it is important to consider that teacher education practice, as well as learning of prospective teachers during initial teacher training, depend on country-specific characteristics of teacher education systems and contextual conditions present in education systems and teacher labour markets (paine & zeichner, 2012). depending on the point of view, our model enables researchers to investigate not only cross-country, but also cross-institutional differences in teacher education practice. cross-country, as well as cross-institutional analyses involve three overarching steps: (1) the choice and inclusion of contextual information in the model; (2) modeling the interrelation between functions, dimensions, or structural elements; (3) and modeling the interrelation between prospective teachers and the system. in its current form, the focus is on the general education system, or teacher labour market, as the immediate context of teacher education. it has to be kept in mind that this is not the only context teacher education is embedded in. depending on the researcher‟s point of view, the institutional, political, or societal context might be considered the immediate context of teacher education (grossman & mcdonald, 2008). c. könig & r.mulder 40 | f l r the choice of contextual information relates to the decision of the researcher to compare different teacher education programmes (for example, university-based versus school-based teacher education; concurrent versus consecutive), or to compare teacher education systems in different countries. when comparing teacher education programmes, the primary context is the institutional context. thus, respective information relates to higher education, for example the degree of integration of the teacher education programme into universities. when comparing teacher education systems, the primary context is the education system or teacher labour market. respective information relates, according to our model, to the supply and demand of teachers in the education system. it is possible to include contextual characteristics as background information, or else, information about group membership in a multigroup model. for example, comparing teacher education programmes in this multigroup framework allows investigating the differential effect of teacher education variables across different educational levels (huang & moon, 2009). for example, the importance of obtaining a degree for student achievement seems to differ across elementary, middle, and high school levels (phillips, 2010). the differential relevance is explained by the generalist/specialist distinction between elementary, middle, and high school teacher education; the importance of subject-specific degrees increases with education level, where teachers are more often trained to be specialists. hence, there seem to be differential effects of different teacher education programmes on teacher characteristics. other possibilities to include contextual information are cross-classification approaches or multilevel models, depending on the quality and detail of available data. the interrelation of the functions, the dimensions of the functions, and even the structural elements constituting the functions might complicate cross-country or institutional comparisons of teacher education systems. with these interrelations it becomes difficult to pinpoint the influencing factors of competence (development) of the prospective teachers, as well as of positive matching, or more general teacher-school matches. however, it can be argued that it is especially this interrelation which renders the possibility of a single influencing factor of teacher education effectiveness improbable. consequently, our model allows the investigation of the influence of configurations of functions, dimensions, and structural elements on the different aspects of teacher education effectiveness. this might be a more appropriate approach to research on teacher education, especially in light of the complex nature of teacher education systems. these interrelations can be accounted for depending on the availability of data and on the focus on either outcomes or processes. with our characterisation of the selection and allocation functions, it is possible to construct empirical typologies of their structural arrangements. in this case, the structural elements are then treated as indicators of their respective dimensions. for example, the assessment procedures and their criteria are indicators of the comprehensiveness of information available about prospective teachers. in a similar manner, school autonomy, recruitment procedures, and union regulations are indicators of control over the recruitment process. based on the structural elements composite measures can be constructed for each dimension. in a further step these composite measures can be used in latent class or cluster analyses in order to identify different approaches to selecting teacher education candidates and prospective teachers, as well as different approaches to allocating teachers. these different profiles can be investigated with regard to their associated organisational outcomes, that is, to success rates of the teacher education system or to different distributions of teachers in the education system. similar approaches have been taken in the context of institutional dimensions of education systems and the relation between education and labour market outcomes (hofman, hofman, & gray, 2008; bol & van de werfhorst, 2011). another possibility for cross-country comparisons in a multigroup framework is focusing on processes rather than outcomes, that is, focusing on the interplay between use of learning opportunities and development of competence rather than on comparisons of mean competence levels. such questions are suited best for a multiple group structural equation modelling approach. the different configurations of both functions can be used as background variables to select countries with similar or different levels of information, integration, or labour market capacities. next, these countries can be compared in differences in the relation between characteristics of prospective teachers, their use of learning opportunities, and the development competences. depending on the comprehensiveness of this learning model, differences in the relations are attributed to differences in the configuration of the functions. c. könig & r.mulder 41 | f l r interrelations may further be specified as interaction effects or cross-classifications of the structural elements in multilevel models. this might be suited if the researcher wants not only to compare different teacher education systems or programmes, but also to identify the influencing factors on for example competence development of prospective teachers. the aforementioned multigroup model can be extended to a multigroup multilevel model. on the organisational level we have the specific structure and characteristics of the learning environment, cross-classified with characteristics of the selection function and contextual conditions in the education system. the individual level comprises, for example, characteristics of prospective teachers and information about their use of the learning opportunities. the different programmes or systems can easily be integrated into the multigroup approach by specifying the multilevel model for each educational level (for example elementary, middle, or high school level). the aforementioned relationships can then be compared across programmes or systems. any difference in coefficients across the groups informs us about the differential effect of teacher education on competence development across teacher education programmes or systems. with this approach, is it not necessary to keep contextual information constant, because it is directly included in the model. moreover, modern structural equation modeling programmes allow the specification of cross-level interactions. with these interaction it is not only possible to investigate top-down (from the system to the prospective teacher), but also bottom-up processes (from the prospective teacher to the system), or else, to investigate the relation between individual and organisational development more closely. 6. conclusion to sum up, it can be stated that the organisational perspective on teacher education as an open system can contribute to existing research by raising awareness with regard to the interrelations of the different parts of a teacher education system, and the interplay between system and individual prospective teachers. with its focus on the selection and sorting of teacher education candidates and prospective teachers, and on the allocation of teachers to schools in the education system, it offers a framework which facilitates a better understanding of these processes and their relation with teacher education effectiveness. additionally it is flexible enough to allow for further developments and extensions, for example the continuing professional development of teachers once they are in the teaching profession, and offers a framework in which researchers are able to integrate own studies and projects. in the end, the model may lead to substantive new insights which facilitate informed and effective policies in order to make teacher education practice more effective, both for prospective teachers and for the system itself. keypoints an organisational model of teacher education is developed. the model illustrates the dependencies of teacher education and its context. the model illustrates the interplay of individual and organisational development. the model includes characterisations of the selection and allocation functions. the model offers various opportunities for further research on teacher education. references akiba, m., letendre, g. k., & scribner, j. p. (2007). teacher quality, opportunity gap, and national achievement in 46 countries. educational researcher, 36(7), 369-387. doi:10.3102/0013189x07308739 c. könig & r.mulder 42 | f l r angrist, j., & guryan, j. (2008). does teacher testing raise teacher quality? evidence from state certification requirements. economics of education review, 27, 483-503. doi:10.1016/j.econedurev.2007.03.002 baumert, j., kunter, m., blum, w., brunner, m., voss, t., jordan, a., ... tsai, y.-m. (2010). teachers' mathematical knowledge, cognitive activation in the classroom, and student progress. american educational research journal, 47, 133–180. doi:10.3102/0002831209345157 bess, j. l. & dee, j. r. (2008). understanding college and university organization: theories for effective policy and practice; volume i: the state of the system. sterling, va: stylus publishing. blömeke, s. (2006). struktur der lehrerausbildung im internationalen vergleich. ergebnisse einer untersuchung zu acht ländern. zeitschrift für pädagogik, 52, 393-416. http://nbnresolving.de/urn:nbn:de:0111-opus-44668 blömeke, s. (2009). predicting educational and occupational success in teacher training and subject-specific degrees – on the predictive validity of cognitive and psycho-motivational selection criteria. zeitschrift für erziehungswissenschaft, 12, 82-110. doi:10.1007/s11618-008-0044-0 bol, t., & van de werfhorst, h. (2011). signals and closure by degrees: the education effect across 15 european countries. research in social stratification and mobility, 29, 119-132. doi:10.1016/j.rssm.2010.12.002 boyd, d., grossman, p., lankford, h., loeb, s., & wyckoff, j. (2009). teacher preparation and student achievement. educational evaluation and policy analysis, 31, 416-440. doi:10.3102/0162373709353129 boyd, d., grossman, p., hammerness, k., lankford, h., loeb, s., ronfeldt, m., & wyckoff, j. (2012). recruiting effective math teachers: evidence from new york city. american educational research journal, doi:10.3102/0002831211434579 cannata, m. (2010). understanding the teacher job search process: espoused preferences and preferences in use. teachers college record, 112, 2889-2934. http://www.tcrecord.org id number: 16011, date accessed: 1/23/2014 4:13:37 pm connor, c. m., son, s. h., hindman, a. h., & morrison, f. j. (2005). teacher qualifications, classroom practices, family characteristics, and preschool experience: complex effects on first graders‟ vocabulary and early reading outcomes. journal of school psychology, 43, 343-375. doi:10.1016/j.jsp.2005.06.001 croninger, r. g., rice, j. k., rahbun, a., & nishio, m. (2007). teacher qualifications and early learning: effects of certification, degree, and experience on first-grade student achievement. economics of education review, 26, 312-324. doi:10.1016/j.econedurev.2005.05.008 darling-hammond, l., & snyder, j. (2000). authentic assessment of teaching in context. teaching and teacher education, 16, 523-545. doi:10.1016/s0742-051x(00)00015-9 denzler, s., & wolter, s. (2009). sorting into teacher education: how the institutional setting matters. cambridge journal of education, 39, 423-441. doi:10.1080/03057640903352440 ehrenberg, r. g., & smith, r. s. (2011). modern labor economics (11th edition). amsterdam: prentice hall. gansemer-topf, a., & schuh, j. (2006). institutional selectivity and institutional expenditures. research in higher education, 47, 613-142. doi:10.1007/s11162-006-9009-4 goe, l., & strickler, l. (2008). teacher quality and student achievement: making the most of recent research. washington, dc: national comprehensive center for teacher quality. (eric document reproduction service no. ed520769) goldhaber, d. (2007). everyone‟s doing it, but what does teacher testing tell us about teacher effectiveness? journal of human resources, 42, 765-794. doi:10.3368/jhr.xlii.4.765 goldhaber, d., & liddle, s. (2011). the gateway to the profession: assessing teacher preparation programs based on student achievement. seattle: cedr. http://nbn-resolving.de/urn:nbn:de:0111-opus-44668 http://nbn-resolving.de/urn:nbn:de:0111-opus-44668 c. könig & r.mulder 43 | f l r grodsky, e., & jackson, e. (2009). social stratification in higher education. teachers college record, 111, 2347-2384. http://www.tcrecord.org id number: 15713, date accessed: 9/16/2013 3:53:50 pm grossman, p., & mcdonald, m (2008). back to the future: directions for research in teaching and teacher education. american educational research journal, 45, 184-205. doi:10.3102/0002831207312906 harris, d. n. & sass, t. r. (2011). teacher training, teacher quality and student achievement. journal of public economics, 95, 798-812. doi:10.1016/j.jpubeco.2010.11.009 hofman, r. h., hofman, w. h. a., & gray, j. m. (2008). comparing key dimensions of schooling: towards a typology of european school systems. comparative education, 44, 93-110. doi:10.1080/03050060701809508 hopkins, d. (2008). a teacher’s guide to classroom research. maidenhead: mcgraw-hill. huang, f. l., & moon, t. r. (2009). is experience the best teacher? a multilevel analysis of teacher characteristics and student achievement in low performing schools. educational assessment evaluation and accountability, 21, 209-234. doi:10.1007/s11092-009-9074-2 ingersoll, r.m., & strong, m. (2011). the impact of induction and mentoring programs for beginning teachers: a critical review of the research. review of educational research, 81, 201-233. doi:10.3102/0034654311403323 jackson, c. k. (2010). match quality, worker productivity, and worker mobility: direct evidence from teachers. nber working paper 15990. johnson, s. m., & kardos, s. m. (2008). the next generation of teachers: who enters, who stays, and why. in m. cochran-smith, s. feiman-nemser & d.j. mcintyre (eds.), handbook of research on teacher education (pp. 445-467). new york: routledge. jovanovic, b. (1979). job matching and the theory of turnover. journal of political economy, 87, 972-990. http://www.jstor.org/stable/1833078 kast, f. e., & rosenzweig, j. e. (1972). general systems theory: applications for organization and management. the academy of management journal, 15, 447-465. http://www.jstor.org/stable/255141 katz, d., & kahn, r. l. (1978). the social psychology of organizations. new york: wiley. kennedy, m. (1998). learning to teach writing: does teacher education make a difference? new york: teachers college press. konold, t., jablonski, b., nottingham, a., kessler, l., byrd, s., imig, s., … mcnergney, r. (2008). adding value to public schools – investigating teacher education, teaching, and pupil learning. journal of teacher education, 59, 300-312. doi:10.1177/0022487108321378 korthagen, f.j.a. (2010). situated learning theory and the pedagogy of teacher education: towards an integrative view of teacher behavior and teacher learning. teaching and teacher education, 26, 98106. doi:10.1016/j.tate.2009.05.001 lankford, h., & wyckoff, j. (2010). teacher labor markets: an overview. in d.j. brewer & p.j. mcewan (eds.), economics of education (pp. 235-242). london: elsevier. liu, e., & johnson, s. (2006). new teachers‟ experiences of hiring: late, rushed, and information-poor. educational administration quarterly, 42, 324-360. doi:10.1177/0013161x05282610 loeb, s., kalogrides, t., & beteille, t. (2012). effective schools: teacher hiring, assignment, development, and retention. education finance and policy, 7, 269-304. doi:10.1162/edfp_a_00068 luschei, t., & carnoy, m. (2010). educational production and the distribution of teachers in uruguay. international journal of educational development, 30, 169-181. doi:10.1016/j.ijedudev.2009.08.004 little, j., & bartlett, l. (2010). the teacher workforce and problems of educational equity. review of research in education, 34, 285-328. doi:10.3102/0091732x09356099 maaz, k., hausen, c., mcelvany, n., & baumert, j. (2006). keyword: transitions in the educational system. zeitschrift für erziehungswissenschaft, 9, 299-327. doi:10.1007/s11618-006-0053-9 c. könig & r.mulder 44 | f l r marshall, j. h., & sorto, a. m. (2012). the effects of teacher mathematics knowledge and pedagogy on student achievement in rural guatemala. international review of education, 58, 173-197. doi:10.1007/s11159-012-9276-6 martz, w. (2013). evaluating organizational performance: rational, natural, and open system models. american journal of evaluation, 34, 385-401. doi:10.1177/1098214013479151 morge, l., toczek, m-c., & chakroun, n. (2010). a training programme on managing science class interactions: its impact on teachers„ practises and on their pupils achievement. teaching and teacher education, 26, 415-426. doi:10.1016/j.tate.2009.05.008 musset, p. (2010). initial teacher education and continuing training policies in a comparative perspective: current practices in oecd countries and a literature review on potential effects. oecd working papers no. 48. paris: oecd publishing. oecd (2011). teachers matter: attracting, developing and retaining effective teachers. pointers for policy development. paris: directorate for education, education and training policy division. http://www.oecd.org/edu/school/48627229.pdf paine, l., & zeichner, k. (2012). the local and the global in reforming teaching and teacher education. comparative education review, 56, 569-583. doi:10.1086/667769 parsons, t. (1951). the social system. london: routledge. phillips, k. j. r. (2010). what does „highly qualified‟ mean for student achievement? evaluating the relationships between teacher quality indicators and at-risk students‟ mathematics and reading achievement gains in first grade. elementary school journal, 110, 464-493. http://www.jstor.org/stable/10.1086/651192 rothstein, j. (2012). teacher quality policy when supply matters. nber working paper 18419. saks, a., uggerslev, k., & fassina, n. (2007). socialization tactics and newcomer adjustment: a metaanalytic review and test of a model. journal of vocational behavior, 70, 413-446. doi:10.1016/j.jvb.2006.12.004 schacter, j., & thum, y. m. (2004). paying for high and low-quality teaching. economics of education review, 23, 411–430. doi:10.1016/j.econedurev.2003.08.002 schneider, m., & somers, m. (2006). organizations as complex adaptive systems: implications of complexity theory for leadership research. the leadership quarterly, 17, 351-365. doi:10.1016/j.leaqua.2006.04.006 schwille, j., & dembele, m. (2007). global perspective on teacher learning: improving policy and practice. paris: unesco international institute for educational planning. scott, w. r., & davis, g. f. (2007). organizations and organizing. rational, natural, and open system perspectives. upper sadle river: pearson. sicherman n., & galor o., (1990). a theory of career mobility. journal of political economy, 98, 169-192. http://www.jstor.org/stable/2937647 staiger, d. o., & rockoff, j. e. (2010). searching for effective teachers with imperfect information. journal of economic perspectives, 24, 97-118. doi:10.1257/jep.24.3.97 stiglitz, j. e. (1975). the theory of screening, education and the distribution of income. american economic review, 65, 283-300. http://www.jstor.org/stable/1804834 tillema, h. h. (1994). training and professional expertise: bridging the gap between new information and pre-existing beliefs of teachers. teaching and teacher education, 10, 601-615. doi:10.1016/0742051x(94)90029-9 van de werfhorst, h. g. (2011). skills, positional good or social closure? the role of education across structural-institutional labour market settings. journal of education and work, 24, 521-548. doi:10.1080/13639080.2011.586994 c. könig & r.mulder 45 | f l r van de werfhorst, h. g., & mijs, j. j. b. (2010). achievement inequality and the institutional structure of educational systems: a comparative perspective. annual review of sociology, 36, 407-428. doi:10.1146/annurev.soc.012809.102538 van der velden, r., & wolbers, m. h. j. (2007). how much does education matter and why? european sociological review, 23, 65-80. doi:10.1093/esr/jcl020 wang, a., coleman, a., coley, r., & phelps, r. (2003). preparing teachers around the world. princeton: educational testing service. wang, j., odell, s. j., & schwille, s. a. (2008). effects of teacher induction on beginning teachers„ teaching. a critical review of the literature. journal of teacher education, 59, 132-152. doi:10.1177/0022487107314002 weinert, f. e. (2001). concept of competence: a conceptual clarification. in d. s. rychen, & l. h. salganik (eds.), defining and selecting key competencies (pp. 45-65). seattle, wa: hogrefe & huber. winters, m. a., dixon, b. l., & greene, j. p. (2012). observed characteristics and teacher quality: impacts of sample selection on a value added model. economics of education review, 31, 19-32. doi:10.1016/j.econedurev.2011.07.014 yeh, s.s. (2009). the cost-effectiveness of raising teacher quality. educational research review, 4, 220232. doi:10.1016/j.edurev.2008.06.002 zeichner, k. (1983). alternative paradigms of teacher education. journal of teacher education, 34, 3-9. doi:10.1177/002248718303400302 zeichner, k. (2005). a research agenda for teacher education. in m. cochran smith, & k. zeichner (eds.), studying teacher education (pp. 737-761). mahwah: lawrence erlbaum. zeichner, k. (2006). studying teacher education programs: enriching and enlarging the inquiry. in c.f. conrad, & r.c. serlin (eds.), the sage handbook for research in education (pp. 79-95). thousand oaks: sage. zeichner, k., & conklin, h.g. (2008). teacher education programs as sites for teacher preparation. in m. cochran-smith, s. feiman-nemser, d. mcintyre, & k. demers (eds.), handbook of research on teacher education (pp. 269-289). new york: routledge. microsoft word goldberg & schwarz_publication.docx           frontline  learning  research  vol.4  no.  4  special  issue  (2016)  7  -­‐  19   issn  2295-­‐3159       harnessing emotions to deliberative argumentation in classroom discussions on historical issues in multi-cultural contexts tsafrir goldberga, baruch b. schwarzb a university of haifa, israel bthe hebrew university of jerusalem, israel article received 18 september / revised 25 january / accepted 16 march / available online 11 may   abstract this theoretical paper is about the role of emotions in historical reasoning in the context of classroom discussions. peer deliberations around texts have become important practices in history education according to progressive pedagogies. however, in the context of issues involving emotions, such approaches may result in an obstacle for historical clairvoyance. the expression of strong emotions may bias the use of sources, compromise historical reasoning, and impede argumentative dialogue. coping with emotions in the history classrooms is a new challenge in history education. in this paper, we suggest that rather than attempting to foster positive emotion only or to avoid emotions all together, we should look at ways of engaging with emotion in history teaching. we present examples of peer deliberations on charged historical topics according to three pedagogical approaches that address emotions in different ways. the protocols we present open numerous questions: (a) whether facilitating engagement with own and the other's emotions may lead to better processing of information and better deliberation of a historical question; (b) whether promoting national pride boosts reliance on collective narratives; and (c) whether adopting a critical teaching approach eliminates emotions and biases. based on these examples and findings in social psychology, we bring forward working hypotheses according to which we suggest that instead of dodging emotional issues, teachers should harness emotions – not only positive but also negative ones, to critical and productive engagement in classroom activities. keywords: teacher identity; knowledge building; inquiry learning; problem-centered pedagogy; technology tsafrir goldberg, dept. of learning, instruction & supervision, university of haifa, 199 abba khoushi rd., haifa, 31905, israel. tgoldberg@edu.haifa.ac.il. doi: http://dx.doi.org/10.14786/flr.v4i4.211 goldberg  &  schwarz       | f l r     8   1. introduction: positions towards the role of emotions in history learning emotions are responses to internal or external events with particular significance for the organism. they include verbal, physiological, behavioral, and neural mechanisms (fox, 2008). emotional experience involves the coordination and synchronization of bodily symptoms, action tendencies, and feelings, driven by appraisal processes (scherer, 2005). although the inclusion of cognitive appraisal (the evaluation of events and objects) in emotions is controversial, cognition is considered either as part of the emotion experience or as interacting with emotion. in his celebrated book, descartes' error: emotion, reason, and the human brain, neurologist antónio damásio (1994) explains that emotions guide (or bias) behavior and posits that rationality requires emotional input. he argues that rené descartes' "error" was the dualist separation of mind and body, rationality and emotion. researchers from other domains express comparable claims: for biologists maturana and varela (1987), cognition, language and mood or emotion are inextricable. for semiotician radford (2015), all emotions and motivations are inherently social and culturally constructed, and do not necessarily obstruct thinking. the question for the educationalist is then how to handle emotions when aiming to foster rational reasoning. this issue is particularly challenging in historical reasoning: since psychologists showed that strong emotions and loyalties hinder rational thinking in general (bless & fiedler, 2006), some researchers in history education suggest avoiding strong and negative emotions holding a sway over cognition (foster, 2013). moreover, politicians and decision makers are often adamant to avoid negative emotions in history classes for ideological reasons (evans, avery, & pederson, 1999). in other words, although research has shown the intricate relations between rationality and emotion, many opt lessening bursts of strong emotions in history classes. descartes comes in again through the backdoor. in the present paper, we claim that handling and capitalizing on emotions as resources for learning is a major goal in history education in the 21st century. we focus on face-to-face deliberative argumentation about highly loaded historical issues, a context that exacerbates emotions in the case of history (schwarz & goldberg, 2013) and fosters reasoning and learning in general (schwarz & asterhan, 2010). we provide examples of teaching approaches designed to engage emotions in different ways and exemplify how these emotions affect deliberative discussions of historical topics in productive or counter-productive directions. we rely on these examples to articulate hypotheses on the beneficial and detrimental roles of emotions in deliberative historical argumentation. 1.1 "don't get emotional now"   historical research has long strove for impartiality, and even when presumptuous aspirations for objectivity were abandoned, emotionality and partisanship are still considered hindrances (haskell, 1990). school history teaching, while attuned to various other goals besides promoting norms of academic historical practice, is also shifting to a growing extent to emphasizing rational disciplinary thinking and discourse (barton, 2009; national center for history in the schools, 2010). in a survey of leading history education experts, only one of the ten intellectually challenging core practices of history teaching they recommend refers (tangentially) to learner emotions ("connect to personal/cultural experience"). even then, the emphasis is on helping learners properly distance themselves from personal reaction and views (fogo, 2014). some educationalists are even more explicit: foster's (2013) review of teaching controversial issues advises eschewing the highly emotive topics in favor of more distant events, in order to allow learners to better focus on disciplinary practices. accordingly, the majority of teachers tend to avoid charged and emotive issues in history teaching (levstik, 2000). issues arousing strong (negative) emotions are often considered "taboo" and are formally or covertly sanctioned (evans et al., 1999). evasion is more frequent with topics that shed unpleasant light on learners' in-group, and may elicit collective shame or guilt which are aversive emotions (helmsing, 2014; wohl, branscombe, & klar, 2006). indeed, historical issues bearing on identity are especially emotion intensive and the emotions they may raise are nor solely positive (mccully, 2006). goldberg  &  schwarz       | f l r     9   zembylas and kambani (2012) claim the emotional risk and complexity of teaching controversial historical issues are especially threatening in societies divided by intergroup conflicts. such a risk arises due to history's role in learners' identity formation, constructing a meta-narrative in which learners position themselves (goldberg, porat, & schwarz, 2006). this may be the reason curriculum policy makers tend to restrict learners' encounter with out-group historical narratives (bar-tal, 2007; goldberg & gerwin, 2013; hilton & liu, 2008). even educational initiatives designed to engage students with their nation's multicultural history such as euroclio's initiatives for the new post soviet democracies, have been criticized for a tendency towards harmonization and towards the evasion from conflictual episodes (maier, 2011). history education experts claim learners should check personal inclinations and emotions lest they be prone to bias and "presentism" (davis, yeager, & foster, 2001; wineburg, mosborg, & porat, 2001). in the context of emotionally charged topics that are more salient in collective memory learners' evidence evaluation tended to be more biased (goldberg, schwarz, & porat, 2008). when confronted with accounts of their nation's history that posed threat to their national pride, patriots demonstrated biased processing of historical information (miron, branscombe, & biernat, 2010). thus, emotions and issues arousing strong emotions seem to threaten “good” rational and disciplinary oriented learning. however, can identity and emotions be side tracked without losing essential aspects of history teaching? some educationalists harness emotions to history learning. for example, zembylas and kambani (2012) call for a focus on the emotional side of teaching (contested) history, to purposefully engage learners with discomforting emotions in reference to sensitive historical topics. for them, empathizing is a strategy to be practiced (zembylas, 2004; zembylas, 2013). britzman (2000) refers to the importance of engaging with learners' emotions when teaching about historical collective trauma. none of the above researchers is interested in history teaching as an end in itself, though. their educational aim concerns reconciliation between adversaries. in the learning sciences, there is now a general recognition of the importance of relating curriculum to learners' identity and community history and of engaging them in advocacy and social critique to boost motivation for learning (thompson, 2014; varelas, 2012). barton and mccully (2010) stress that encountering conflicting historical perspectives through critical disciplinary inquiry may help students engage in an internally persuasive dialog about the past and achieve a tolerant and receptive identity. instead of dodging the role of emotions and of identity, gottlieb, wineburg, and zakai (2005) claim one should acknowledge identity influences on historical understanding, and accept the fact that individuals apply different critical standards to historical sources central and peripheral to group identity. muller mirza and colleagues (muller mirza et al., 2014) point to the importance of "secondarisation" of emotion, the process in which individuals reflect on their emotions and generalize them into more abstract concepts. such a process is essential for handling contentious intercultural topics and for academic achievement. bar-on and adwan (2006) advocate structuring history teaching to help learners acknowledge their own and the other's emotions and collective identity. they assume this would help deliberating contentious issues of the past in a more productive and reasoned way and motivate engagement with it. however, the effects of affirmation of sentiments relating to national identification and collective narrative have not been explored so far. we present here a study enabling a comparison between three approaches to history learning – disciplinary critical inquiry, mutual narrative acknowledgement and patriotic apologetic teaching, and by such we inquire about relation between emotion and learning in history. to do so, we focus on inter-group deliberative discussions (or deliberative argumentation) on a "hot" historical topic: deliberative argumentation is a propitious context for learning (schwarz & asterhan, 2010); hot historical topics are lieux de memoire where identity and national identification (or national pride) are susceptible to emerge. our working hypothesis is that emotion and identity do not constitute obstacles to deliberative discussions or disciplinary practice by themselves. to check our hypothesis, we compare examples of peer deliberations with different structuring of engagement with emotions. in the last part of the paper, we rely on these examples to articulate hypotheses on the role of emotions in deliberative argumentation. goldberg  &  schwarz       | f l r     10   1.2 feeling and discussing the past: learners' deliberative discussions learners' emotions were addressed through three approaches to history teaching (for full details see goldberg and ron's (2014) description of procedure). the first is an authoritative single-narrative approach aligns with declared national history teaching goals such as acquiring factual knowledge of main events and enhancing students' commitment to the state and their collective identity (israeli ministry of education, 2015). instruction according to this approach was based on a textbook chapter written under direct governmental supervision to produce a clear account stressing the righteousness of israel (domke, urbach, & goldberg, 2009; yaron, 2009). teaching was an "initiation-recitation-evaluation" session with a powerpoint presentation, directed at getting the "right answer" in a short quiz and instilling pride in one's nation. this conventional lower order thinking type of teaching appears to be very common in social studies classrooms (saye & social studies inquiry research collaborative (ssirc), 2013) and aligns with helmsing's examples of reasoning that does not challenge national pride (2014). the second approach is empathetic dual-narrative. it aims at arousing feelings of empathy for the other and mutual affirmation of collective sentiments. instruction was based on excerpts from a dual narrative history textbook created by jewish and palestinian teachers (adwan & bar-on, 2004; bar-on & adwan, 2006). learners were driven to empathetic attention to the emotions and values of adversary narrators and reflection on the reactions they arouse. this practice aligns to some degree with the process of secondarisation, in which learners reflect on emotions and produce more generalized understanding of their social aspect (mirza, grossen, de diesbach-dolder, & nicollin, 2014). finally, the critical inquiry approach aims at modelling disciplinary practice and developing critical thinking skills (reisman, 2012). instruction was based on the use of conflicting sources accompanied with information allowing inferences as to context, goal and bias of authors. teachers coached students in sourcing and corroboration (wineburg, 2001) and explicitly attempted to instil impartiality in the encounter with evidence (instructing them to read like a "swedish" [i.e. neutral] historian). table 1 succinctly summarizes the emotional emphases and practices of the three approaches. jewish and arab israeli students from diverse schools were randomly allocated to study the topic of the 1948 war ("war of independence") in one of the three approaches. two weeks later, participants were paired by teaching approach into jewish-arab dyads; they engaged in deliberative discussion of the war and the causes of the palestinian refugee problem, a topic bound to raise emotional reactions. all discussions were audiotaped and transcribed. we marked all discussion episodes that included direct references to history or identity, or use of historical disciplinary practices (sourcing, contextualization, perspective taking, causal explanation) (lee & ashby, 2000; wineburg, 2001). two researchers coded ten of the sixty discussion, and arrived at 75% agreement. differences were discussed and resolved. the rest of discussions were analyzed separately. we bring forth four episodes we deem representative of potential effects of teaching approaches on the relations of emotion and learning in history. goldberg  &  schwarz       | f l r     11   table 1 teaching approaches, emotional emphases and practices approach emotional emphasis practices conventional authoritative instilling pride in one's nation getting the "right answer" presentation textbook reading and summary. exam empathetic dual-narrative mutual affirmation conflict resolution empathy nonjudgmental listening to collective narratives. identifying with emotions and values critical disciplinary inquiry induction into disciplinary practice instill impartiality critical analysis of conflicting sources synthesis of sources figure 1 shows a first protocol in the critical-inquiry condition. at the beginning of the discussion (not presented here), the jewish participant devoted almost four times as many words to analysis and evaluation of the sources as his arab peer (416 vs. 111 words). could this stress on the analytical disciplinary approach come at the price of feeling attached to the collective and its history? the protocol shows that the jewish discussant, who attempts an implicit distancing from jewish identity ("israeli is enough for me"), also expresses a disinterest in jewish history ("it doesn't arouse interest in me…i did it for the exam…i don't care what i study"). by contrast, his peer declares herself a palestinian arab, and stresses her glorifying view of her people ("the palestinians especially are very very very smart"). she also endorses emphatically learning the history of "my arab country", by which she learns "very very wise things", and which she would "love others to study" it too. the contrast between the two learners suggests that some degree of identification and perhaps even glorification seems to motivate learning one's own national history. furthermore, it seems that the arab student who is more enthusiastic to learn about her own groups history also views learning the others' history more positively ("even when i study jewish history…it's good, nice that i learn more"). could the positive effect that feelings of pride and identification have on motivation to learn in-group history, be generalized to the out-group? is there an essential relation between disidentification and emphasis on disciplinary practice? these questions are complex and do not seem to have general answers. goldberg  &  schwarz       | f l r     12   figure 1. an excerpt of a discussion in a critical-inquiry condition that suggests that national identification boosts the motivation to study own and others’ national history. a: we in school study history, jewish history, the jewish undergrounds and…but you don't study arab history…why? j: because we live in one state…state defined with one nation, and which has one curriculum… a: i live in this country and define myself as palestinian arab, ok? you define yourself as jewish, ok? j: israeli a: jewish israeli j: whatever….israeli is good enough for me a: for you it is enough to study jewish history, i define myself as arab i will learn only the arab… j: call me narrow… but the fact i study something connected to me, supposedly connected to me, the truth is it doesn't, it doesn't arouse interest in me. i don't care which history. i did it for the exam. a: not me. i, even when i study jewish history…it's good, nice that i learn more and more. but i study jewish history and i would love the others to study mine… j: i agree with you a: you know the arabs, the palestinians especially … are very very very smart…so, i learn about my arab country, i very very very study very wise things. j: first, l agree with you…i wouldn't mind studying another history…but personally, if you ask me, i don't care what history i would have studied. goldberg  &  schwarz       | f l r     13   figure 2 shows an excerpt of a discussion between two discussants in the narrative empathetic condition. they took the perspective of the other, and frequently expressed feelings. figure 2. two discussants in the narrative empathetic condition exemplify they took the perspective of the other, and frequently expressed feelings it is noteworthy that both attune themselves to the suffering of the palestinians, an orientation that apparently also facilitated a more collaborative atmosphere. perspective taking is used for laying foundations for mutual trust, for discussing the possible (although to some degree counterfactual) decisions of leaders and for agreeing they would not fight as their predecessors had. the second part of figure 2 shows another phase in which the jewish discussant invites his palestinian peer to relate to her and her peoples' feelings, with which he appears to empathize. both cases of perspective taking show awareness of the difference and distance between the historical agents and the learners (as evinced by the use of the third person "them”), and of the action of perspective taking ("that's how i imagine myself if i was in their place at the time"). emotion here is used as a venue into the disciplinary practice of perspective taking. it is a cognitive tool, directed at reconstructing the historical agents' consciousness. this approach also aligns to some degree with the movement between unicity and genericity, which drives forward the process of "secondarisation" in dealing with tense social phenomena (muller mirza et al., 2014). it is worth comparing this discussion to a conversation in the conventional-authoritative condition. figure 3 shows an arab discussant responding in highly contentious and emotional manner to his jewish peer's reconciliatory counterfactual speculation. both discussants do not demarcate themselves from historical figures, but identify and "merge" with them through using first and second person plural, in what a: i have a question; if you were one of the great men in the country or in israel, would do you think you would have done? j: got it. i'd try to compromise on one decision a: which is? come on… j: seems to me like the un saidto split the land into two parts so there would be peace and they wouldn't fight…that's it…what do you thinkwhat would you have done? a: me too, i think may be we could have reached a peace agreement… … j: just a question right out of my headhow do you think the palestinians felt after they were deported? a: fear. no mother, no land, no one to turn to, that's how i imagine myself if i was in their place at the time…nothing to do, no power, no army, no leaders to supervise them or lead them, nothing but themselves going where the arab states or un tells them to go. j: no one to lead them. a: and that's scary, right. goldberg  &  schwarz       | f l r     14   seems like a clear expression of collective memory. discussants do not refer to the text they studied, nor do they analyse critically the evidence it contained. they rely on religious or mythical backings rather than on historical ones, impeding further deliberation of the question. thus, it is in the context of teaching aimed at conveying a clear undisputed narrative that learning is emotionally disrupted and the past is disputed with no reliance on the discipline of history. the jewish discussant, who initially adopted a more rational and more collaborative perspective, feels forced to gradually adopt a confrontational model. figure 3. two discussants in the conventional-authoritative condition entertain a contentious interaction. let us exemplify now another discussion in the critical-disciplinary condition. figure 4 shows an arab discussant who demonstrates his feeling of loyalty to his community as mediated through the family. in spite of his declared impartiality ("you know, i don't distinguish between the texts") he points quite clearly to his being an arab as the reason for his preference for arab historian's excerpt. this preference is accompanied by what seems like a confirmation bias – the tendency to view evidence consistent with prior opinions as more reliable. the jewish participant on the other hand shows far lower preference for in-group member's sources and founds his criticism of sources on the practices he studied in the preparatory session – sourcing and contextualization. we can see here how national identity can bias historical practices such as evidence evaluation. however, the critical inquiry approach encourages participants to reflect on their evaluation of evidence and expose their bias. it also seems that, at least for a member of the dominant ethnic group, the critical inquiry approach promotes a more balanced and impartial disciplinary practice. j: that's what i think should have happened; we could get along, you know what, even two separate states, one state for two people and joint leadership and everything. no need to deport, you know…at most we could have asked for some more territory so it would be enough for two states. a: but you didn't! you deported and murdered j: and i'm saying, in my view there shouldn't have been deportation a: and you killed us and murdered us and stole and all this j: i didn't do it and i think a: you didn't do it but they did j: i think it was a mistake and it shouldn't have been done in no way…there were mistakes on both sides in the same way we fought over our state earlier, and you fought for your state. we both wanted something, basically the same. each in his direction. a: ok. but you said you needed a state for the jews, but why in palestine? j: because i said it, this place is sacred for me too, it is, like my granddad’s granddad’s granddad he had here, they found graves here, it's a place my ancestors lived in same as your ancestors…this place is important to us, also sacred for us too…i'm not a religious person and wouldn't go by the edicts of judaism in most cases but in the same way a palestinian thinks this place is sacred for the palestinian people, it's a sacred place for the jewish people. goldberg  &  schwarz       | f l r     15   figure 4. two discussants in the critical-inquiry condition demonstrating identity motivated and disciplinary practice. 2. discussion this paper initiates a reflection on the effects of emotions on oral argumentation for highly charged historical topics. the three pedagogical approaches exemplified in the protocols modelled different ways students handled emotions. the conventional authoritative single-narrative approach instils pride in one's nation and appears to delegitimize doubt and perspective taking. the empathetic dual-narrative approach facilitates mutual affirmation and increases the use of historical perspective taking, though not the use of critical thinking. the critical inquiry approach draws learners to reflect and expose the relation of their identity and emotions to disciplinary practices, and to some degree helps overcome it. our claims cannot count as conclusions but as working hypotheses for further research. moreover, this paper outlined the implications of various theoretical and empirical perspectives on the role of emotions in historical reasoning, fleshing them with actual discussion excerpts. we set forth the working hypotheses that these implications suggest. firstly, history teaching that legitimizes the complex emotions arising from encounter with outgroup perspectives by promoting strategic empathy and reflection on emotion (mccully, 2006; zembylas, 2013), appears to promote productive deliberative discussions. this is perhaps because it affords more chances for mutual gestures helping maintain dialogue or because discomforting emotions help participants take the perspective of historical agents in troubled times (zembylas & mcglynn, 2012). engagement with emotion according to this approach seems to foster a nuanced and conscious use of perspective taking, though not necessarily better handling of evidence. secondly, a direct attempt to expose the relation of identity-related emotions to historical practices, helps develop learners' internally persuasive dialogue and reflection about evidence (barton & mccully, 2010). this may promote critical thinking practices such as evidence evaluation. however, it does not insure curbing emotionally driven biases since evidence is used (perhaps with more restraint and awareness) in relation to identity needs and emotions (goldberg, 2013; gottlieb et al., 2005). both these approaches promote a productive merging of emotion and cognition or reasoning (mingers, 1991; radford, 2015). by contrast, the conventional teaching approach, neither challenges nor acknowledges the role of collective emotion in learning (barton, 2009). however, it appears a: our opinion, i, because i'm an arab, you know, i don't distinguish between the texts, that of the jew or the arab, but because my parents lived through it, they were in the time of 1948, i give more credibility to the text of written by the arab. i read there many things i already knew. j: your parents told you similar things? a: yeah…but what i read in the jewish author's text was new, so that's why… j: ok. personally, i too really believe the arab author more. a: what? j: i too really believe the arab author more because the israeli author writes… from a role of propaganda kind of, and from within the ministry of foreign affairs and attempts to explain israel to the world, so it's in his interest to be in favor of the israeli side. goldberg  &  schwarz       | f l r     16   to enhance or unleash its (negative) effect, both on learning and on deliberative discussion (bless & fiedler, 2006; hilton & liu, 2008). this may hint that relating to emotions holds a promise for better processing of information or deliberation of a historical question. in summary, the examples presented suggest that emotion and identity do not necessarily constitute obstacles to deliberative discussions or disciplinary practice by themselves, in line with baker et al.'s (2013) ideas. however, students in each approach engaged their emotions and learning differently. it appears that instructional practices moderate and influence the relations of emotion and reasoning. facilitating empathetic listening, nurturing national glorification, or attempting to hold emotion at bay, may each lead to a different way of arguing about the past. we currently undertake an experimental study that involves the systematic comparison of discussions as well as of learning outcomes in the three approaches. the analyses we undertake will hopefully confirm and sharpen our working hypotheses. meanwhile, we believe that it is possible to rely on the above examples to draw some tentative conclusions on the teaching of history. history was introduced as a core discipline in schools in the 19th century in order to bring students to believe that they belong to a nation and to foster national pride (ferro, 2004). we alluded to the experts' emotions-free list of core teaching practices and skills that reflect a substantial shift to the adoption of the norms of history as a critical rational discipline (fogo, 2014). this shift does not pay attention, though, to the emotions history nurtures, arouses and is motivated by. the cognitive practices of history are nested within, colored by and interact with emotion (maturana & varela, 1987) as collaborative learning and disciplinary oriented practices take the lead in history teaching, the role of emotions becomes ever more important. the protocols we presented suggest that while simply fostering the glorification of the nation impedes historical deliberation, concurring teaching approaches that bring to the fore alternative narratives and engage with strong emotions may actually help handle such influences. therefore, educators should not dodge emotions in their teaching but, on the contrary, should capitalize on them to boost historical reasoning. the first place for bringing forward these strong emotions is of course the discussion. this setting risks bringing to the surface strong emotions (like anger), leading to breakdowns. however, the preparatory practices of engagement with both in-group and out-group perspectives, acknowledging and evaluating emotional overtones, seem to tone down contentious reactions when engaging in argumentation across groups. with appropriate framing and instruction, students develop their capacity to handle emotions in discussions (muller mirza et al., 2014). they speak with each other, and deliberating together the historical roots of their conflict, even if their discourse was sometimes biased. we suggest, then, that the core practices of history teaching should also include addressing common opinions held by the different stakeholders on the issue, relating to the emotional states of the different historical actors and to the emotional reactions, of the learners (bar-on & adwan, 2006) and facilitating small group discussions across groups by helping the handling of emotions. these additional practices may promote the role of history in helping learners become citizens engaged in productive deliberation of their contentious past and their shared present. furthermore, we believe that research on historical understanding and reasoning should change both its prescriptive and its analytic stance to the role of emotions. first, if we wish to address the complexity of goals and needs history education addresses in reality (and not in an idealized rational expert model), it would serve researchers well to give emotion a more central role and treat it without disdain (barton, 2009). second, instead of relating to emotions and loyalties anecdotally, as indications of diverse identities, biased cognition or novice practice, there should be much to gain in exploring it proactively. methodologically this means tracking and documenting emotion, whether through self-report or other implicit and observational methods now accessible through emotions research (see baker et al., 2013). analytically, we should start looking at emotions as promoters, factors and even as desirable outcomes of learning in history (goldberg, 2013). this does not mean, of course, that we should ignore the inhibiting role of emotions in some cases, and that we should eliminate from class activity a detached critical-disciplinary approach. our message is goldberg  &  schwarz       | f l r     17   that emotions are precious resources for history education, but that teachers should learn when and how to capitalize on them. references adwan, s., & bar-on, d. (2004). shared history project: a prime example of peace building under fire. international journal of politics, culture, and society, 17(3), 513-521. baker, m. j., andriessen, j. e., & järvelä, s. (eds.). (2013). affective learning together. london: routledge. bar-on, d., & adwan, s. (2006). the psychology of better dialogue between two separate but interdependent narratives. israeli and palestinian narratives of conflict: history’s double helix, 205224. bar-tal, d. (2007). socio-psychological foundations of intractable conflicts. american behavioral scientist, 50(11), 1430-1453. barton, k. c., & mccully, a. w. (2010). "you can form your own point of view": internally persuasive discourse in northern ireland students' encounters with history. teachers college record, 112(1), 142181. barton, k. c. (2009). the denial of desire: how to make history education meaningless. in l. symcox, & a. wilshcut (eds.), national history standards: the problem of the canon and the future of teaching history (pp. 265-282). charlotte, nc: information age. bless, h., & fiedler, k. (2006). mood and the regulation of information processing and behavior. in j. p. forgas (ed.), affect in social thinking and behavior (pp. 65-84). new york: psychology press. britzman, d. p. (2000). if the story cannot end: deferred action, ambivalence, and difficult knowledge. in r. simon, s. rosenberg & c. eppert (eds) between hope and despair: pedagogy and the remembrance of historical trauma, (pp. 27-55). lanham, md: rowan and littlefield. davis, o. l., yeager, e. a., & foster, s. j. (2001). historical empathy and perspective taking in the social studies. lanham, md: rowman & littlefield. domke, e., urbach, c., & goldberg, t. (2009). building a state in the middle east [bonim medina bamizrach hatichon]. jerusalem, israel: zalman shazar center. eid, n. (2010). the inner conflict: how palestinian students in israel react to the dual narrative approach concerning the events of 1948. journal of educational media, memory, and society, 2(1), 55-77. evans, r. w., avery, p. g., & pederson, p. v. (1999). taboo topics: cultural restraint on teaching social issues. the social studies, 90(5), 218-224. ferro, m. (2004). the use and abuse of history: or how the past is taught to children. london: routledge. fogo, b. (2014). core practices for teaching history: the results of a delphi panel survey. theory & research in social education, 42(2), 151-196. fox, e. (2008). emotion science: cognitive and neuroscientific approaches to understanding human emotions. hampshire, uk: palgrave macmillan. goldberg, t., & gerwin, d. (2013). israeli history curriculum and the conservative liberal pendulum. international journal of historical teaching, learning and research, 11(2), 111-124. goldberg, t. (2013). “it's in my veins”: identity and disciplinary practice in students' discussions of a historical issue. theory & research in social education, 41(1), 33-64. goldberg, t., & ron, y. (2014). ‘look, each side says something different’: the impact of competing history teaching approaches on jewish and arab adolescents’ discussions of the jewish–arab conflict. journal of peace education, 11(1), 1-29. goldberg, t., porat, d., & schwarz, b. b. (2006). “here started the rift we see today”: student and textbook narratives between official and counter memory. narrative inquiry 16(2), 319–347. goldberg, t., schwarz, b. b., & porat, d. (2008). living and dormant collective memories as contexts of history learning. learning and instruction, 18(3), 223-237. doi:http://dx.doi.org/10.1016/j.learninstruc.2007.04.005 goldberg  &  schwarz       | f l r     18   gottlieb, e., wineburg, s., & zakai, s. (2005). when history matters: epistemic switching in the interpretation of culturally charged texts. eleventh biennial meeting of the european association of learning and instruction, haskell, t. l. (1990). objectivity is not neutrality: rhetoric vs. practice in peter novick's that noble dream. history and theory, 29(2), 129-157. helmsing, m. (2014). virtuous subjects: a critical analysis of the affective substance of social studies education. theory and research in social education, 42(1), 127-140. hilton, d. j., & liu, j. h. (2008). culture and intergroup relations: the role of social representations of history. in r. m. sorrentino, & y. susumu (eds.), handbook of motivation and cognition across cultures (pp. 343-368). london, uk: academic press. israeli ministry of education (2015). history curriculum for the jewish secular public schools. retrieved march 2016 from http://cms.education.gov.il/educationcms/units/mazkirut_pedagogit/history/tochnitlimudimvt/talt ashah.htm lee, p., & ashby, r. (2000). progression in historical understanding among students ages 7-14. in p. n. stearns, p. seixas & s. wineburg (eds.), knowing, teaching, and learning history: national and international perspectives (pp. 199-222). new york, ny: new york university press. levstik, l. s. (2000). articulating the silences: teachers' and adolescents' conceptions of historical significance. in p. n. stearns, p. seixas & s. wineburg (eds.), knowing, teaching, and learning history: national and international perspectives (pp. 284-305). new york, ny: new york university press. maier, r. (2011). how we lived together in georgia in the 20th centuryexternal evaluation euroclio/matra project (2008—2011). retrieved from http://www.euroclio.eu/new/index.php/component/docman/doc_download/1080-how-we-livedtogether-in-georgia-in-the-20th-century-external-review maturana, h. r., & varela, f. j. (1987). the tree of knowledge: the biological roots of human understanding. boston: new science library/shambhala publications. mccully, a. (2006). practitioner perceptions of their role in facilitating the handling of controversial issues in contested societies: a northern irish experience. educational review, 58(1), 51-65. mingers, j. (1991). the cognitive theories of maturana and varela. systems practice, 4(4), 319-338. miron, a. m., branscombe, n. r., & biernat, m. (2010). motivated shifting of justice standards. personality & social psychology bulletin, 36(6), 768-779. doi:10.1177/0146167210370031 muller mirza, n., grossen, m., de diesbach-dolder, s., & nicollin, l. (2014). transforming personal experience and emotions through secondarisation in education for cultural diversity: an interplay between unicity and genericity. learning, culture and social interaction, 3(4), 263-273. national center for history in the schools. (2010). historical issues. retrieved from http://www.nchs.ucla.edu/history-standards/historical-thinking-standards/5.-historical-issues radford, l. (2015). of love, frustration, and mathematics: a cultural-historical approach to emotions in mathematics teaching and learning. in b. pepin, & b. roesken-winter (eds.), from beliefs to dynamic affect systems in mathematics education (pp. 25-49). new york: springer. reisman, a. (2012). reading like a historian: a document-based history curriculum intervention in urban high schools. cognition and instruction, 30(1), 86-112. saye, j., & social studies inquiry research collaborative (ssirc). (2013). authentic pedagogy: its presence in social studies classrooms and relationship to student performance on state-mandated tests. theory & research in social education, 41(1), 89-132. schwarz, b. b., & asterhan, c. s. (2010). argumentation and reasoning. in k. littleton, c. wood & j. kleine staarman (eds.), international handbook of psychology in education (pp. 137-176). london: emerald group publishing. schwarz, b. b., & goldberg, t. (2013). “look who’s talking”: identity and emotions as resources to historical peer reasoning. in m. j. baker, j. e. andriessen & s. järvelä (eds.), affective learning together (pp. 272-292). london: routledge. goldberg  &  schwarz       | f l r     19   sorek, t. (2011). the quest for victory: collective memory and national identification among the arabpalestinian citizens of israel. sociology, 45(3), 464-479. thompson, j. (2014). engaging girls’ sociohistorical identities in science. journal of the learning sciences, 23(3), 392-446. varelas, m. (2012). identity construction and science education research: learning, teaching, and being in multiple contexts. springer science & business media. wineburg, s. s. (2001). historical thinking and other unnatural acts: charting the future of teaching the past temple university press. wineburg, s. s., mosborg, s., & porat, d. (2001). what can forrest gump tell us about students' historical understanding? social education, 65(1), 55-58. wohl, m. j., branscombe, n. r., & klar, y. (2006). collective guilt: emotional reactions when one's group has done wrong or been wronged. european review of social psychology, 17(1), 1-37. doi:10.1080/10463280600574815 yaron, m. (2009). history superintendent's requested corrections to the textbook: building a state in the middle east.[personal correspondence to author] zembylas, m. (2004). emotion metaphors and emotional labor in science teaching. science education, 88(3), 301-324. zembylas, m. (2013). critical pedagogy and emotion: working through ‘troubled knowledge’ in posttraumatic contexts. critical kambani studies in education, 54(2), 176-189. zembylas, m., &, f. (2012). the teaching of controversial issues during elementary-level history instruction: greek-cypriot teachers' perceptions and emotions. theory & research in social education, 40(2), 107133. zembylas, m., & mcglynn, c. (2012). discomforting pedagogies: emotional tensions, ethical dilemmas and transformative possibilities. british educational research journal, 38(1), 41-59. frontline learning research 6 (2014) 25-33 issn 2295-3159 corresponding author: karsten stegmann, department psychology, lmu munich, leopoldstrasse 13, 80802 munich, germany, email: stegmann@lmu.de doi: http://dx.doi.org/10.14786/flr.v2i4.112 25 | f l r constructing nomological nets on the basis of process analyses to strengthen cscl research karsten stegmann a a department psychology, lmu munich, germany article received 26 may 2014 / revised 6 august 2014 / accepted 29 september 2014 / available online 23 december 2014 abstract due to the nature of collaborative learning, realising perfectly controlled experiments often requires an unreasonable amount of resources and sometimes it is not possible at all. against this background, i propose to augment as good as feasible experimental design with a nomological net of relations between instructional support (intervention), learning processes and learning outcomes. nomological networks are known from construct validity. in construct validity, the relations between variables (e.g. group differences, correlation matrices) are used to provide evidence for the validity of a measure. by adding multiple process and outcome variables together with the corresponding relations between intervention, process and outcome, the validity of causal relations found can be strengthened. i suggest adopting quality criteria from good research designs to evaluate the nomological nets. the resulting net needs to be (1) theory grounded, (2) situational, (3) feasible, (4) redundant, and (5) efficient. by making these nomological nets explicit and by designing them according to the presented criteria, cscl research becomes more potent: the risk of inconclusive results is reduced while results that form a consistent nomological net can be interpreted with a stronger confidence, even if the experimental design has some flaws. if this becomes standard in cscl research, it can be expected to contribute significantly better to knowledge accumulation in this area of research. keywords: construct validity; nomological net; research design; computer-supported collaborative learning (cscl) k. stegmann 26 | f l r 1. the role of process analyses in cscl research approaches to computer-supported collaborative learning (cscl) are mainly based on three assumptions. first, collaborative learning outperforms (under particular circumstances, e.g. with specific support) other methods when it comes to learning outcomes. usually, specific collaborative activities like argumentation (e.g. clark, d'angelo, & menekse, 2009), transactive co-construction (e.g. molinari et al., 2013; weinberger & fischer, 2006), reciprocal teaching (palincsar & brown, 1984) and collaborative concept mapping (van boxtel, van der linden, roelofs, & erkens, 2002) are considered to be positively related to individual cognitive processes of learning. second, computer support enables both certain learning activities (e.g. simulation-based inquiry learning; de jong & van joolingen, 1998) and more direct support for certain activities (e.g. scaffolds as an inherent, but adaptive, component of the learning environment; cf. koschmann, 1994). technology enables natural systems and phenomena that would otherwise be invisible and therefore impossible to be experienced (e.g. the heart of an engine or magnetism on an atomic level; cf. fischer, lowe, & schwan, 2008). technology also enables us to facilitate learning processes by different means, e.g. by making various resources accessible (e.g. osborne & hennessey, 2003), scaffolding specific individual processes like the construction of single arguments (stegmann, wecker, weinberger, & fischer, 2012), or offering ways to communicate and collaborate (wegerif, 2002). third, the combination of collaborative learning and technology can have positive interaction effects that go beyond the simple combination of main effects. on the one hand, the quality of collaborative learning processes is lifted through adaptive scaffolds that positively moderate the positive effects of collaborative learning. on the other hand, the effects of technology functions (like access to various resources) on learning outcomes are boosted through collaborative learning (cf. weinberger, stegmann, & fischer, 2010). set against this background, cscl research aims to provide knowledge about how technology can support collaborative learning processes (and thereby learning outcomes on an individual as well as a group level; cf. stahl, 2006) most effectively. on the one hand, the problems that may arise through collaboration or the use of technology have to be minimised, while, on the other, the use of technology resources and collaborative learning processes has to be optimised. the effect of cscl on learning outcomes is therefore mediated by processes that occur during the collaborative learning phase. this general model can be described in a triangle of hypotheses (cf. wecker, stegmann, & fischer, 2012; fig. 1): (a) instructional/technological support facilitates learning activities; (b) facilitated learning activities have positive effects on learning outcomes; and (c) mediated by learning activities, instructional/technological support has a positive effect on learning outcomes. figure 1: general triangle of hypotheses in cscl research. to test this triangle of hypotheses and to allow researchers to infer causal-effect relations, three conditions must be fulfilled (cf. cook & campbell, 1979): (a) when the causing variable varies, the affected variable must vary too (covariation); (b) the cause must occur before the effect occurs (temporal precedence); and (c) no plausible alternative explanations exist. while the first two conditions can be reached in cscl research rather easily, the third condition is very difficult to reach. cscl research often takes place in field-like settings and even studies with a rather higher level of control (e.g., weinberger, marttunen, laurinen, & stegmann, 2013) are much less controlled than classical psychological experiments. k. stegmann 27 | f l r the adherence to instructional advice, for example, is usually not enforced. testing the effect of an intervention on learning activities is, therefore, experimentally a variation check, but semantically the test of whether the way the instruction is realised is able to induce the intended behaviour. due to the nature of collaborative learning, realising perfectly controlled experiments requires an unreasonable amount of resources and sometimes it is not possible at all. just imagine a jigsaw experiment (for a detailed description of the jigsaw method see aronson, 1978). in a jigsaw script, the content to be learnt is split into, for example, four subtopics. groups of four prepare one of the four topics and finally four new groups are formed with one learner from each of the previous groups and learners teach their subtopic to the other group members. in this experiment, individual learning, unscripted collaborative learning and collaborative learning using the jigsaw method are compared. in the individual learning condition, 32 subjects are enough if a large effect with 80% power is expected. in the condition with unscripted collaborative learning in groups of four, the number of subjects might be optimally, due to nestedness of data, 128 learners in 32 groups. in the jigsaw condition, 512 subjects are required due to the fact that 16 subjects learn collaboratively together in 32 groups. according to maas and hox (2005), a number of more than 50 groups is needed for acceptable statistical multilevel analyses. with 64 groups across two conditions, this criterion is fulfilled. finally, this simple one-factorial design with three conditions requires 672 subjects to find effects with large effect size. and still it would not be free of confounded factors, e.g. the type of support is confounded with group size. while a condition with unscripted learners in groups of 16 would be possible, a condition with jigsaw script with groups of four is not possible. this example illustrates the inherent problems of cscl research in excluding plausible alternative explanations through experimental design. against this background, i propose to augment as good as feasible experimental design with a nomological net of relations (hypotheses) between instructional support (intervention), learning processes and learning outcomes. 2. criteria for nomological nets as the basis of high quality cscl research nomological networks are known from construct validity (cf. cronbach & meehl, 1955). in construct validity, the relations between variables (e.g. group differences, correlation matrices) are used to provide evidence for the validity of a measure. i like to utilise this idea to validate the (causal) relation between interventions, mediators and outcome variables. the smallest nomological net possible comprises just two variables and one relation, but does not yet additionally validate a causal relation. the net, however, becomes stronger the more ties and knots are part of the net. the stronger the net, the stronger the confidence in the validity of causal relations between variables. the smallest net that can increase the confidence in a causal relation in cscl research is the triangle of hypotheses described previously with three knots and three relations. by adding multiple process and outcome variables together with the corresponding relations between intervention, process and outcome, the validity of causal relations found can be strengthened. the development of such a nomological net requires some general quality criteria that allow evaluation of the net. i suggest adopting quality criteria from good research designs as described, for example, by trochim and land (1982). according to these authors, the nature of good designs is (1) theory grounded, (2) situational, (3) feasible, (4) redundant, and (5) efficient. in the following sections, i provide a short explanation of the criteria and some illustrating examples from cscl research for each criterion. 2.1 theory grounded the nomological net needs to be theory grounded. for each of the relations, a directional effect needs to be explicable by a theory and may be backed up by previous empirical findings. the question, for example, concerning the extent to which learners mutually influence one another has attracted considerable attention in cscl research. particular focus has been placed on the degree to which groups of learners share a mutual understanding via social interaction. accordingly, attempts have been made to quantify this process, which is referred to as knowledge convergence, based on analyses of text-based knowledge-building processes. mäkitalo-siegl and colleagues (2012) traced so-called knowledge pieces through collaboration. k. stegmann 28 | f l r the transfer of knowledge pieces from one learner to another was measured by comparing the knowledge pieces mentioned by single learners before, during and after collaboration. the relations in the corresponding nomological net would be that collaborative learners share more knowledge pieces during collaboration and that these shared knowledge pieces are known better by group members after collaboration. weinberger, stegmann and fischer (2007) presented an approach with additional quantitative measures of the convergence of prior knowledge, collaborative processes and acquired knowledge based on fine-grained (i.e. at the level of inferences) analyses of text-based data sources (pretest, text-based online discussion, post-test). the authors suggest using the variation coefficient of the number of different inferences within a group of learners (analogous to knowledge pieces) as an indicator of knowledge divergence. applying these measures provided insight into the relationship between the processes and outcomes of collaborative knowledge construction (e.g., weinberger, stegmann, & fischer, 2005; zottmann, et al., 2013). the nomological net may comprise, for example, the relation that learners with high divergence during online discussions (i.e. contributing different as opposed to identical inferences) are more likely to share knowledge after collaboration than learners with high convergence during online discussions. in real life, however, not all assumed relations show up as expected in empirical research. the strength of the net is, of course, stronger if all hypotheses made a priori test successfully. in practice, a net needs to be adopted post hoc to explain the results at hand. in these cases, additional explanatory variables and relations may be added to achieve a consistent net. the new relations need, of course, to be theory grounded as well. by adding further mediators, moderators and/or suppressors, the results can form a consistent net again. as a by-product, the adaption of a net is a further development of the initial theory. 2.2 situational the manipulated/measured variables as well as the relations between them need to be situational defined. the interventions, the process variables and, to a large extent, the outcome variables in cscl research are highly situational, i.e. depend on the situation, content and context at hand. the intervention is usually one realisation out of an endless number of possible realisations of a specific theoretical principle (e.g. an implementation of the jigsaw method; cf. wecker, 2013). the process and outcome variables may be termed rather general (such as “content quality”, “quality of argumentation”; cf. stegmann, weinberger, & fischer, 2007), but the concrete operationalisation requires the inclusion of bottom-up criteria that derive from the (learning) material and raw process data. in most cases, especially in the case of process variables, general (i.e. less situational) measures of skills or competences (like those applied in pisa studies; cf. kobarg, prenzel, & seidel, 2011) are not suitable, because they measure features (abilities) of persons, not activities. it is, therefore, necessary to apply measures with a high content validity. such measures usually need to be developed individually for each situation. it is necessary to quantify the qualities of collaborative activities with respect to multiple quality dimensions. a detailed description of such a multidimensional approach for the qualitative coding of online discussions (maqcod) can be found in stegmann and fischer (2011). collaborative (learning) activities are, for example, often analysed in terms of the content quality or quality of the argumentation (e.g. weinberger & fischer, 2006). depending on the dimension, the grain size of the analysis needs to be defined (cf. strijbos, martens, prins, & jochems, 2006). the quality of the argumentation, for example, could be defined for the entire discourse, single messages or even single arguments. which grain size is most suitable depends on the theoretically defined relationship between the quality dimension to be analysed and the intended learning outcome. if, for example, the theoretical model assumes that formulating arguments with grounds and warrants is a core collaborative learning activity, single arguments rather than complete conversations need to be the focus. along with grain size, the categories need to be defined. single arguments can, for instance, be coded in terms of whether they are grounded or not. the definitions of the dimensions, grain size and categories per dimension need to be carefully documented. this may comprise: segmentation rules and examples for the application of these rules; the names of dimensions and categories; rules about when to assign a specific category; and examples when a category applies and when it does not. this documentation forms the basis for an objective, reliable k. stegmann 29 | f l r and content-valid coding of learning activities and, thereby, for the inclusion of process variables in a nomological net. 2.3 feasible the inclusion of process variables in a nomological net requires that it is feasible to extract the variable from the recorded activities during cscl. it is, for example, problematic to measure the depth of cognitive processing just by analysing written cscl discourse data. researchers may argue that a sophisticated, well-elaborated argument can be regarded as an indicator of deep cognitive processing. the argument, however, might be just “copied” from a different source (e.g. a learning partner or prior knowledge) without deep cognitive processing. such a measure, therefore, does not have sufficient content validity to be included in a nomological net with the intended function. this is not an argument against the variable “depth of cognitive processing” in general. the requirement is to measure the variable in a contentvalid way, i.e. as directly as possible. if data sources such as think-aloud protocols are available (e.g. stegmann, wecker, weinberger, & fischer, 2012), it might be adequate to include such a variable in a nomological net. 2.4 redundant like in a cockpit of an aeroplane, central components of the nomological net might be redundant, i.e. implemented several times. in cscl research, this redundancy is often regarded as a methodological challenge rather than a strength of the research design. an inherent feature of collaborative learning is the nestedness of learners in groups and in time. learners are part of a group and thereby features of the group affect learning. in addition, the knowledge and skills of the single learner as well as of the group change over time (wise & chiu, 2011). the activities of learners in a group are affected not only by the initial features of the single learners and the collaborative learning phase, but also by the activities that the single learner and the group performed previously. furthermore, instructional support such as collaboration scripts affects the relationship between previous and current activities. the nestedness of learners in groups and in time is not only an issue if researchers aim to understand why learners learn in a certain way; learners and groups also change over time. the script theory of guidance (stog; fischer, kollar, stegmann, & wecker, 2013), for example, assumes that instructional support for collaborative learning needs to be adapted consecutively to ensure the optimal fit between the skills of the single learners/the group and the instructional support. as a result, a nomological net may include relations that reflect such ideas but add time as a moderator of the effect of an intervention on the process of collaborative learning. as already raised in the jigsaw study example, (quantitative) research on cscl often requires many more participants due to the issue that learners who learned in groups cannot be regarded as independent observations. from the viewpoint of the nomological net, the relation, for example, of an intervention at group level on a specific process is redundant within a group. for an intervention that aims to facilitate, for example, argumentation sequences (e.g., stegmann, weinberger, & fischer, 2011; jeong, clark, sampson & menekse, 2010), a positive relation between the intervention and the number of argument-counterargumentsyntheses sequences may be added to the nomological net at group level. on an individual level (i.e. the level of group members), positive relations between the intervention and the number of contributed counterarguments and syntheses may be added to the net. furthermore, scaffolds examined in cscl research often focus on specific processes that occur multiple times during collaborative learning. if an intervention such as a collaboration script aims to support the quality of each single argument (cf. stegmann, wecker, weinberger, & fischer, 2012) with respect to grounds and warrants, the effect is expected to show up on the level of each argument (as the probability that a claim is supported by ground and/or warrant), on the level of each individual learner (as the share of grounded/warranted claims contributed by an individual), and on the group level (as a higher argumentative quality of the discussion). k. stegmann 30 | f l r 2.5 efficient an important aspect, finally, is the efficiency of the testing of the relations specified in the nomological net. efficiency is in general determined by two factors: the usefulness of a result and the extent of resources required to reach the result. this criterion seems to be contradictory to the examples for the previously described criteria. these criteria require rather qualitative analyses of processes on multiple levels including time series. many more resources (i.e. technology to record data, space to archive the data, manpower to develop coding manuals and to analyse data) need to be spent. in cscl research, however, the digital learning environment in which learning activities under examination usually take place can reduce the amount of resources. especially assessing and analysing data are supported by technology. technologies like ibeacon or active rfid chips allow learners to be traced as well as the interaction between them and artefacts to learn from in the context, for example, of museums. eberle and colleagues (2013), for example, traced the activities of conference participants using active rfid chips to examine the relation between interaction between conference participants, planned future collaboration right after the conference and collaborative publications two years after the conference. in scenarios with computer-mediated communication, the communication can easily be recorded. the opportunity to log data in technology-enhanced learning environments can easily produce a large amount of data that exceeds the limits for meaningful human analyses. the development in the area of machine learning technology, however, enables researchers to train algorithms to – supervised or unsupervised – analyse digital data according to multiple dimensions such as quality of argumentation, content quality or emotions. to apply these algorithms to data at hand, in a first step, features of the learning processes need to be extracted. in the case of written discourse data, for example, the number of specific words or word pairs, the punctuation or the line length are extracted. this step can be easily performed using tools such as taghelper (rosé et al., 2008) or lightside (mayfield & rosé, 2012). in a second step, these features are used in conjunction with a human coding that serves as training material to build models that are able to measure the respective quality. recently, mu and colleagues (2012) presented the acodea framework, which may serve as a blueprint on how to apply this technology in cscl research. the empirical results presented by mu and colleagues (2012) show that this procedure enables objective analyses of texts that were not previously used for training to be conducted. the results obtained were at the same level as those produced by human coding, and were in some cases even better than those produced by interhuman objectivity. while the described application of technology in the research process contributes to a reduction of required resources, i further argue that the usefulness of the results is increased by testing a comprehensive nomological net in comparison to results not embedded in a net. testing the three types of hypotheses of the general triangle of hypotheses (cf. fig. 1) increases the probability that a study will produce useful results (and not just because three times more hypotheses are tested). if all of the three hypotheses are significant, this can be regarded as a validation of the underlying theoretical model. however, if one or two of the three hypotheses fail (e.g. the relationship between learning activities and learning outcomes), but the others are significant, the findings provide a starting point for explanations that may improve the initial theoretical model. it is only if all of the hypotheses fail to be significant that the empirical results will be completely inconclusive regarding generalisability and causal relations. nevertheless, a more in-depth analysis of learning activities and post hoc adaption of the nomological net still provides insights into the mechanisms of learning, regardless of significant effects on learning outcomes. 3. conclusion the general structure of cscl research can be described using the introduced triangle of hypotheses. therefore, nomological nets are an inherent feature of cscl research. by making these nomological nets explicit and by designing them according to the presented criteria, the research becomes more potent: the risk of inconclusive results is reduced while results that form a coherent nomological net can be interpreted with a stronger confidence even if the experimental design has some flaws. this, however, k. stegmann 31 | f l r is by no means an argument to conduct studies with an easily improvable experimental design or to skip experimental variation completely. an as good as possible experimental design is the basic prerequisite for the nomological net to contribute to strengthening the confidence in causal interpretations of effects. the suggestion to use a nomological net as described is, nevertheless, not limited to quantitative research approaches. some if not all relations might be examined with qualitative methods. the effect on the confidence in the interpretation is the same as in quantitative methods: it increases. actually, i would expect quantitative and qualitative methods to be used in a complementary way to form nomological nets in cscl research. the explication of the nomological net, therefore, should become obligatory in reports and presentations on research in cscl. studies that aim to provide evidence for causal relations need to report effects on processes and outcomes, not either or. the processes have to be analysed in a way that ensures content validity. the (statistical) analyses have to make use of the multilevel structure of the process data. new technologies have to be applied to cope with the vast amount of data. if this becomes standard in cscl research, it can be expected to contribute significantly better to knowledge accumulation in this area of research. keypoints realising perfectly controlled experiments often requires unreasonable amount of resources and sometimes it is not possible at all. by adding multiple process and outcome variables together with the according relations between intervention, process and outcome into a nomological net, the validity of causal relations found can be strengthened. the resulting nomological net needs to be (1) theory grounded, (2) situational, (3) feasible, (4) redundant, and (5) efficient incorporating nomological nets reduce the risk of inconclusive results while results that form a consistent nomological net can be interpreted with a stronger confidence, even if the experimental design has some flaws. references aronson, e. (1978). the jigsaw classroom. london: sage. clark, d. b., d'angelo, c. m. & menekse, m. (2009). initial structuring of online discussions to improve learning and argumentation: incorporating students' own explanations as seed comments versus an augmented-preset approach to seeding discussions. journal of science education and technology, 18, 321-333. doi:10.1007/s10956-009-9159-1 cook, t. d., & campbell, d. t. (1979). quasi-experimentation: design and analysis for field setting. ma: houghton mifflin. cronbach, l. j., & meehl, p. e. (1955). construct validity in psychological tests. psychological bulletin, 52(4), 281. doi:10.1037/h0040957 de jong & van joolingen (1998). scientific discovery learning with computer simulations of conceptual domains. review of educational research, 68, 179-201. doi:10.3102/00346543068002179 eberle, j., stegmann, k., lund, k., barrat, a., sailer, m., & fischer, f. (2013). fostering learning and collaboration in a scientific community – evidence from an experiment using rfid devices to measure collaborative processes. in n. rummel, m. kapur, m. nathan, & s. puntambekar, s. (eds.), to see the world and a grain of sand: learning across levels of space, time, and scale: cscl 2013 conference proceedings volume 1 — full papers & symposia (pp. 169-175). international society of the learning sciences. k. stegmann 32 | f l r jeong, a., clark, d. b., sampson, v. d., & menekse, m. (2010). sequential analysis of scientific argumentation in asynchronous online discussion environments. in s. puntambekar, g. erkens & c. hmelo-silver (eds.), analyzing interactions in cscl: methodologies, approaches and issues. berlin: springer. doi:10.1007/978-1-4419-7710-6_10 fischer, f., kollar, i., stegmann, k., & wecker, c. (2013). toward a script theory of guidance in computersupported collaborative learning. educational psychologist, 48(1), 56-66. doi:10.1080/00461520.2012.748005 fischer, s., lowe, r. k., & schwan, s. (2008). effects of presentation speed of a dynamic visualization on the understanding of a mechanical system. applied cognitive psychology, 22(8), 1126-1141. doi:10.1002/acp.1426 kobarg, m., prenzel, m., & seidel, t. (2011). an international comparison of science teaching and learning. further results from pisa 2006. münster: waxmann verlag. koschmann, t. d. (1994). toward a theory of computer support for collaborative learning. the journal of the learning sciences, 3(3), 219-225. doi:10.1207/s15327809jls0303_1 maas, c. j. m., & hox, j. (2005). sufficient samples sizes for multilevel modeling. methodology, 1, 86–92. doi:10.1027/1614-1881.1.3.86 mayfield, e., & rosé, c. p. (2012). lightside: open source machine learning for text accessible to nonexperts. in m. d. shermis & j. burstein (eds.), handbook of automated essay grading (pp. 124135). new york: routledge. mäkitalo-siegl, k., stegmann, k., frete, a., & streng, s. (2012). orchestrating computer-supported collaborative learning: effects of knowledge sharing and shared knowledge. in s. abramovich (ed.), computers in education (pp. 75-91). commack, ny: nova science publishers. molinari, g., chanel, g., betrancourt, m., pun, t, & bozelle, c. (2013). emotion feedback during computer-mediated collaboration: effects on self-reported emotions and perceived interaction. in n. rummel, m. kapur, m. nathan, & s. puntambekar, s. (eds.), to see the world and a grain of sand: learning across levels of space, time, and scale: cscl 2013 conference proceedings volume 1 — full papers & symposia (pp. 336-343). international society of the learning sciences. mu, j., stegmann, k., mayfield, e., rosé, c. & fischer, f. (2012). the acodea framework: developing segmentation and classification schemes for fully automatic analysis of online discussions. international journal of computer-supported collaborative learning, 7(2), 285–305. doi:10.1007/s11412-012-9147-y osborne, j., & henessy, s. (2003). literature review in science education and the role of ict: promise, problems and future directions. bristol: nesta futurelab. retrieved from http://www.futurelab.org.uk/resources/publications_reports_articles/literature_reviews /literature_review380 palincsar, a. s., & brown, a. l. (1984). reciprocal teaching of comprehension-fostering and comprehension-monitoring activities. cognition & instruction, 1(2), 117-175. doi:10.1207/s1532690xci0102_1 rosé, c. p., wang, y. c., arguello, j., stegmann, k., weinberger, a., & fischer, f. (2008). analyzing collaborative learning processes automatically: exploiting the advances of computational linguistics in computer-supported collaborative learning. international journal of computer-supported collaborative learning, 3(3), 237-271. doi:10.1007/s11412-007-9034-0 stahl, g. (2006). group cognition. cambridge, ma: mit press. stegmann, k. & fischer, f. (2011). quantifying qualities in collaborative knowledge construction: the analysis of online discussions. in s. puntambekar, g. erkens & c. hmelo-silver (eds.), analyzing interactions in cscl: methods, approaches and issues (pp. 247-268). new york: springer. doi:10.1007/978-1-4419-7710-6_12 stegmann, k., weinberger, a., & fischer, f. (2011). aktives lernen durch argumentieren: evidenz für das modell der argumentativen wissenskonstruktion in online-diskussionen [active learning by argumentation: evidence for the model of argumentative knowledge construction in online discussions.]. unterrichtswissenschaft, 39(3), 231–244. k. stegmann 33 | f l r stegmann, k., wecker, c., weinberger, a. & fischer, f. (2012). collaborative argumentation and cognitive elaboration in a computer-supported collaborative learning environment. instructional science, 40(2), 297-323. doi:10.1007/s11251-011-9174-5 stegmann, k., weinberger, a., & fischer, f. (2007). facilitating argumentative knowledge construction with computer-supported collaboration scripts. international journal of computer-supported collaborative learning, 2(4), 421-447. doi:10.1007/978-0-387-36949-5_12 strijbos, j.-w., martens, r. l., prins, f. j., & jochems, w. m. g. (2006). content analysis: what are they talking about? computers & education, 46(1), 29-48. doi:10.1016/j.compedu.2005.04.002 trochim, w., & land, d. (1982). designing designs for research. the researcher, 1(1), 1-6. van boxtel, c., van der linden, j., roelofs, e., & erkens, g. (2002). collaborative concept mapping: provoking and supporting meaningful discourse. theory into practice, 41(1), 40-46. doi:10.1207/s15430421tip4101_7 wecker, c. (2013). how to support prescriptive statements by empirical research: some missing parts. educational psychology review, 25(1), 1-18. doi:10.1007/s10648-012-9208-9 wecker, c., stegmann, k., & fischer, f. (2012). lernund kooperationsprozesse: warum sind sie interessant und wie können sie analysiert werden? [learning and cooperation processes in casebased learning. interesting issues and analysis approaches] report: zeitschrift für weiterbildungsforschung, 35(3), 30-41. doi:10.3278/rep1203w wegerif, r. (2002). thinking skills, technology and learning: a review of the literature for nesta futurelab. bristol: nesta futurelab. retrieved from: http://www.futurelab.org.uk/resources/publications_reports_articles/literature_reviews /literature_review394 weinberger, a., & fischer, f. (2006). a framework to analyze argumentative knowledge construction in computer-supported collaborative learning. computers & education, 46(1), 71-95. doi:10.1016/j.compedu.2005.04.003 weinberger, a., marttunen, m., laurinen, l., & stegmann, k. (2013). inducing socio-cognitive conflict in finnish and german groups of online learners by cscl script. international journal of computersupported collaborative learning, 8(3), 333-349. doi:10.1007/s11412-013-9173-4 weinberger, a., stegmann, k., & fischer, f. (2005). computer-supported collaborative learning in higher education: scripts for argumentative knowledge construction in distributed groups. in t. koschmann, d. d. suthers & t.-w. chan (eds.), computer supported collaborative learning 2005: the next 10 years! proceedings of the international conference on computer supported collaborative learning 2005 (pp. 717-726). mahwah, nj: lawrence erlbaum. doi:10.3115/1149293.1149387 weinberger, a., stegmann, k., & fischer, f. (2007). knowledge convergence in collaborative learning: concepts and assessment. learning and instruction, 17(4), 416-426. doi:10.1016/j.learninstruc.2007.03.007 weinberger, a., stegmann, k., & fischer, f. (2010). learning to argue online: scripted groups surpass individuals (unscripted groups do not). computers in human behavior, 26(4), 506-515. doi:10.1016/j.chb.2009.08.007 wise, a. f., & chiu, m. m. (2011). analyzing temporal patterns of knowledge construction in a role-based online discussion. international journal of computer-supported collaborative learning, 6(3), 445-470. doi:10.1007/s11412-011-9120-1 zottmann, j., stegmann, k., strijbos, j. w., vogel, f., wecker, c., & fischer, f. (2013). computersupported collaborative learning with digital video cases in teacher education: the impact of teaching experience on knowledge convergence. computers in human behavior, 29(5), 2100-2108. doi: 10.1016/j.chb.2013.04.014 frontline learning research 5 (2014) 140-166 issn 2295-3159 corresponding author: ismo t. koponen, department of physics, p.o. box 64, fi-00014 university of helsinki, finland. ismo.koponen@helsinki.fi doi: http://dx.doi.org/10.14786/flr.v2i3.120 140 | f l r a systemic view of the learning and differentiation of scientific concepts: the case of electric current and voltage revisited ismo t. koponen, tommi kokkonen department of physics, university of helsinki, finland article received 12 february 2014 / revised 14 april 2014 / accepted 29 june 2014 / available online 3 july 2014 abstract in learning conceptual knowledge in physics, a common problem is the incompleteness of a learning process, where students’ personal, often undifferentiated concepts take on more scientific and differentiated form. with regard to such concept learning and differentiation, this study proposes a systemic view in which concepts are considered as complex, dynamically evolving structures. the dynamics of the concept learning and differentiation is driven by the competition of model utility in explaining the evidence. based on the systemic view, we introduce computational model, which represents the essential features of the conceptual system in the form of directed graph (dgm), where concepts are nodes connected to other conceptual elements (nodes) in the graph. the results of a dgm are then compared to the empirical findings to identify differentiation between concepts of electric current and voltage based on a re-analysis of previously published empirical findings on upped secondary school students’ learning paths in the context of dc circuits. the comparison shows that the model predicts and explains many relevant, empirically observed features of the learning paths of concept learning and differentiation, such as: 1) contextdependent dynamics, 2) the persistence of ontological shift and concept differentiation, and 3) the effects of communication on individual learning paths. the systemic view and the dgm model based on it make these generic features of interest in concept learning and differentiation understandable and show that these features are associated with the guidance of theoretical knowledge. finally, we discuss briefly the implications of the results on teaching and instruction. keywords: concept learning; concept differentiation; ontological shift; complex system, directed graph model i. t. koponen & t. kokkonen 141 | f l r 1. introduction learning scientific concepts is a demanding and lengthy process, in which the learner’s initial and personal concepts and conceptions gradually change towards more scientific concepts in that they are part of an extensive and coherent knowledge system (theory), which regulates and constrains their use. previous research (lee & law 2001; reiner, slotta, chi & resnick, 2000; smith, carey & wiser, 1985) has raised the notion that learners seldom use concepts in the same sense as they are used in scientific knowledge. one particular but central question is proper concept differentiation. when two closely related concepts are linked to the same phenomenon, novice learners do not always properly understand them as different concepts. rather, the concepts are confused and used in undifferentiated ways (lee & law, 2001; reiner et al., 2000; smith et al., 1985). the aim of the learning process then is to produce a clearer and more scientific understanding not only of how such concepts differ, but also of how they are related, a process referred to here as concept differentiation. concept differentiation has often been discussed from the viewpoint of ―ontological shift‖, which views that the ontological attributions are at the centre of concept learning, and changes in those attributions are the main mechanisms behind differentiation (chi & slotta, 1993; chi, 2005, 2008). this position finds support in the notion that ontological commitments in concept development are deeply rooted in the psychological aspects of concepts (murphy, 2004; keil 1989). however, the ontological shift view has been criticized for overemphasising the role of static ontologies (gupta, hammer & redish, 2010) and failing to pay proper attention to the role of theory in learning (ohlsson, 2011). in addition, when studying concept differentiation, one should understand that concepts must be shared and be communicable to other learners. communicating and sharing of concepts is closely related to problem of knowledge convergence, discussed mostly in cases of the explanations convergence and seeking consensus and a common way to understand concepts and terms (jeong & chi, 2007; weinberger, 2007). however, how communication affects the learning of scientific concepts and the differentiation process or, in general, which stages or steps of the learning paths communication could possibly affect, remains unclear. consequently, our understanding of the learning path in concept differentiation remains partially incomplete. one promising way to remedy this lack of understanding views learners’ concepts as complex structures and the learning process itself as a systemic process consisting of different conceptual elements and where those elements interact (brown & hammer, 2008; koponen & huttunen, 2013). the present study proposes a new way of synthesising different views by focusing explicit attention on concept learning and concept differentiation, so that the synthesis takes into account aspects of interest for personal concepts, such as the role of ontological attributions, and aspects relevant to scientific concepts, such as the communicability of concepts and their constrained, law-like use. such synthesis, referred to here as the systemic view, sees concepts as complex structures. different stages of concept learning, with partially differentiated concepts, are then seen as partial projections of the structure in different real situations; the projections are partial and incomplete mappings of more complete systems. on the level of personal concepts, the systemic view stems from recent views of the heterogeneity of concepts, which emphasise the diversity of roles of concepts in different cognitive processes (machery, 2009). on the level of scientific concepts, the systemic view borrows much from the ―dynamic frames‖ view of scientific concepts, where both ontological attributions and theoretical, law-like (nomic) knowledge are considered central to concept development (anderssen & nersessian, 2000; andersen, barker & chen, 2006). in the systemic view, the learning process also requires a driving force or mechanism; this study suggests that the utility of models, through which the concepts are used, and the competition of models based on utility provides that mechanism (ohlsson, 2009, 2011). the systemic model is applied here to discuss and simulate concept differentiation and its generic features in one empirically well-studied case; the concepts of electric current and voltage. the generic features of interest are the robustness of certain simple forms of the concepts (often called as misconception or intuitive conceptions), the strong context dependence of these conceptions, the occurrence of ontological shift and its persistence once achieved, and the role of theoretical knowledge in concept differentiation and ontological shift. this study focuses on developing the theoretical background of the systemic model. to that i. t. koponen & t. kokkonen 142 | f l r end, we further develop the directed graph model that we recently introduced (koponen, 2013). we embody the theoretical model by using re-analysis of empirical data of nine students’ learning processes, in groups of three (koponen & huttunen, 2013). we introduce a simulation model, based on directed graphs, to model the learning path and to reproduce the most important generic features of the empirical findings of concept differentiation. finally, we discuss some interesting implications for teaching that the model raises. the model presented here supports the view that ontological shift is not the primary agent in learning scientific concepts; rather, it stems from theoretical learning, driven by model utility. this means that instead of focusing on ontological training and on developing instructional methods based on it, attention should focus on how theoretical knowledge is introduced and applied in the learning process. another important notion is the role of context in learning and how students are gradually introduced to more demanding tasks. the model results show that overly complex tasks cannot promote learning if students lack sufficiently advanced concepts; yet overly simple tasks lead to stagnation, where a learner gets stuck on simple models and unsophisticated concepts. what is needed is a learning path that is progressive and which demands use of complex models and concepts. according to the view presented here, the learner needs to receive theoretical knowledge through instruction and to see its utility in complex enough situations, thus avoiding ―overlearning‖ of simple cases. this emphasises not only the teacher’s role, but also the importance of variation in contexts in which the knowledge is applied. these notions, based on the systemic view and on its computational embedding, therefore have direct practical consequences for how one should design learning paths and the role of teacher in them. 2. concept learning and differentiation: theoretical underpinnings pre-scientific concepts are often idiosyncratic, context dependent and difficult to communicate. of course, scientific concepts, as used by advanced learners and experts, not only share some aspects with ―personal‖ pre-scientific concepts, but also differ from them in important ways. one of the most important differences is that scientific concepts often refer to categories (entities or objects) that − like models − are themselves purely conceptual rather than categories within the reach of experience or observation, and their use is law-like (nomological) and constrained (andersen & nersessian, 2000; andersen, barker & chen, 2006; hoyningen-huene, 1993). nevertheless, personal and scientific concepts share features, in particular on the level of how theory or theory-like knowledge structures the sets of attributes that characterise the concepts. concept differentiation is a process where the sets of attributes that characterise and typify concepts become structured so that no other concept shares the same set and values of its attributes. in the case of scientific concepts, this requires that a given concept have law-like (i.e. nomological) relationships to other concepts. it is through these features that the concepts acquire the sharp descriptive power they hold in scientific theories (andersen & nersessian, 2000; andersen et al., 2006; hoyningen-huene, 1993.) at the core of learning scientific concepts is a transformation process where personal, individual concepts which are meaningful to a learner himself or herself, but not easily communicable or meaningful to other learners, are transformed into concepts that are more communicable and, where consensus exists, how they can be used in relation to other concepts. this latter way of using concepts is already scientific in that normative and law-like rules govern the use of the concepts (the nomological use of concepts), and the attributes that characterise them are sharply identified. these brief notions suggest that a suitable theoretical underpinning must link the individual learning and use of concepts to the shared, public use of concepts; intraand interpersonal levels of concept use and learning must be coupled. 2.1. concepts: personal and shared discussions of the learning process, where a student learns and acquires scientific concepts and becomes a fluent user of such concepts, must focus on differences in the ways in which concepts are i. t. koponen & t. kokkonen 143 | f l r understood when seen from the viewpoint of an individual’s personal cognition and learning, and when concepts are discussed as they are shared and used in scientific communities. in the former case, concepts are personal, often un-explicated, and seen as strongly context dependent (carey, 2010; gopnik & meltzoff, 1997), whereas in the latter case, concepts are explicated elements of scientific theories, and their proper use is constrained by the knowledge system as a whole (andersen, 2006; andersen & nersessian, 2000; hoyningen-huene, 1993). to highlight this difference between personal and shared scientific concepts, we use the terms ―intrapersonal concepts‖ and ―interpersonal concept‖. of these, the interpersonal concepts can be shared on the level of small and local groups, such as study groups in learning, or on the level of extended and global groups, such as scientific communities, in which case interpersonal concepts are simply called scientific concepts. the learning process, where an individual’s intrapersonal concepts acquire scientific character and are transformed into interpersonal ones, involves epistemic dimensions (the use of concepts in context of explanation) and communicative dimensions, where concepts are used in communication and consensus finding for what is explained and how. a description of the learning process from intrapersonal to interpersonal concepts requires one to have a model of concepts, which, at one end of the continuum, takes a form of an intrapersonal concept and in another end, as an interpersonal, scientific concept. 2.1.1 intrapersonal concepts in psychology and cognitive science, two important viewpoints of interest here regarding intrapersonal concepts are concepts as prototypes and concepts as theories (machery, 2009; murphy, 2004; smith & medin, 1981). in the concepts-as-prototypes view, the prototype represents a certain class of entities or objects to which the concept refers, and the prototype is understood as a body of knowledge about the properties of the members in that class. however, such properties are assumed to be only statistical or probabilistic, and are not strictly necessary or sufficient by themselves to determine membership (machery, 2009; murphy, 2004). the statistical or probabilistic knowledge contained in the prototypes can be about either 1) the typicality of the category or 2) its cue-like properties. in both cases, a set of properties or attributes and their values indicate how likely or significant a given property is in regard to the identification of the concept (smith & medin, 1981). in concept learning, where concepts develop and are transformed, as in, for example, a concept combination process, new concepts emerging from the combination process inherit some − but not necessarily all − of the properties of the ancestor prototypes (murphy, 2004). a reverse process of concept differentiation can be understood as a process where new concepts inherit partial or split sets of the properties of the original concepts. from the viewpoint of concepts as prototypes, an important part of concept learning is to learn the concept’s ontological attributions, which determine to which ontological categories the concept refers (keil, 1989; murphy, 2004). the ontological shift theory of conceptual change (chi & slotta, 1993; reiner et al., 2000; slotta & chi, 2006) addresses the way in which learners associate substanceand process-like attributes with the concepts they use. according to the ontological shift theory, many students’ learning difficulties originate from a misconceived ontological class (chi & slotta, 1993; slotta & chi, 2006). the view of concepts-as-theories is based on psychological research, which views conceptual knowledge as theory-like (carey, 2010; gopnik & meltzoff, 1997). the concepts-as-theory view focuses on the role of causal knowledge in the categorisation process and in concept learning. concepts, in this view, are first and foremost carriers of causal knowledge about the properties of the members of classes to which the concepts refer. therefore, causal knowledge is considered crucial for concept recognition and differentiation. quite often, the role of causal knowledge is discriminative with regard to the attributes or properties attached to a concept (machery, 2009; murphy, 2004; rehder, 2003). these two different views of intrapersonal concepts can be thought of as two different ways to use concepts (machery, 2009), thus reflecting the multifaceted aspects of concepts. such multifacetedness can considered as a sign of a real cognitive difference between the various ways of using concepts (machery, 2009) or as different projections (or mappings) of a more integrated, complex and generic system that projects differently in different real situations (danks, 2010). here, we adopt and further develop the latter viewpoint of concepts as systems projecting differently in different situations. the aspects of greatest interest i. t. koponen & t. kokkonen 144 | f l r in developing such a systemic view are: 1) attributes and sets of attribute values (as in the prototype view) and 2) causal and theoretical knowledge and its role in distinguishing the attributes in concept combination (as in the concepts-as-theories view). 2.1.2 interpersonal concepts in learning, concepts must be shared with other learners, instructors and teachers; concepts must be interpersonal. when concepts are shared, there must be common agreement of referents of the concepts, the ways in which they refer to and the ways to use concepts; there must be certain norms of usage. in particular, when concepts are scientific concepts, they are ―public‖ in that members (scientists) of institutional groups (scientific communities) share these intrapersonal concepts. crucial in to this is not only to agree on the norms, but also to link the norms to accepted verification methods, such as observations, experiments and models (andersen et al., 2006; andersen & nersessian 2000; hoyningen-huene 1993). there are relatively few attempts to discuss scientific concepts so that connection is made to a psychological understanding of concepts. one notable exception, however, is a view which sees scientific concepts as dynamic frames embracing conceptual knowledge (andersen et al., 2006). the dynamic frame view assumes that advanced scientific concepts are acquired by the same process of categorisation as everyday concepts. the categorisation of interest here is how different exemplar-type problems fall into the same classes based on how different types of models serve in solving those problems. then, the characteristic (but not defining) features of concepts emerge from the reference to classes of models, or clusters of models. scientific concepts, where the models form the classes relevant for learning concepts, are also regulated and constrained by certain rules for applying concepts in construction of models; the norms guide how to use the concepts. the dynamic frames incorporate the attributes of concepts, in much the same way as in the prototype theory, and the theoretical knowledge as in the theory-theory approach, but now theoretical knowledge has a role of organising the attributes and imposing constraints on their co-variation (andersen & nersessian, 2000; andersen et al., 2006). 2.3. concept learning as convergence process the focal point of this study is the individual learner’s process of learning scientific concepts, where personal concepts are transformed into scientific concepts. however, because learning takes place in a community of students and teachers, we must also understand how communication affect the learning process. collaborative learning and sharing ideas in small groups has been shown to enhance student learning. such learning is beneficial when the members’ knowledge supplements others’, but differs only slightly from it; members show some knowledge equivalence and knowledge sharing within the group (jeong & chi, 2007; weinberger, 2007). here, however, re-analysed and revisited the empirical data contain little information about the communication, and although the effect is evident, very simple models serve here to estimate the effects of communication on concept differentiation. 2.4. systemic view the systemic view sees concepts as part of a knowledge system, where the ―concept‖ as a part of the operation of the system may have a plurality of appearances and project differently in different contexts, yet the parts of the system remain unchanged. recently, some have suggested a different but related type of systemic view. it employs the ideas of dynamic complex systems, where robust and persistent conceptual patterns can arise in emergent fashion from interactions of elemental pieces of the dynamic system (brown & hammer, 2008). these interactions can thus give rise to a full spectrum of different projections of concepts, some of which are simple and some, complex. the systemic view is also adopted here, so concepts are considered functional parts of the system, affected by the system and its evolution. the systemic view requires specific representations of concepts which can capture their complex, multifaceted and dynamic nature. a suitable model of such concepts should address at least the following features: i. t. koponen & t. kokkonen 145 | f l r 1) attributes and sets of attribute values as in prototype and dynamic frame views. 2) theoretical knowledge in role of constraining and guiding the use of concepts. 3) models as they connect to the development of scientific concepts. 4) model competition and utility as mechanisms affecting the evolution of concepts. requirements 1-2 are essential to retaining a connection to a psychological view of concepts and concept learning, which understand intrapersonal concepts and the transition from intrato interpersonal concepts. requirements 2-4 are essential to describing how intrapersonal concepts develop or change into scientific concepts. in what follows, we introduce just such a systemic model in section 3 and then, in section 4, discuss how empirical results concerning the differentiation of the concepts electric current and voltage can be embedded within it. finally, in section 5, we present a computational embedding of the systemic view and use the computational model to simulate the process of concept differentiation. 3. systemic view on concept differentiation the systemic view of concept learning and differentiation sees concepts as constructs which, in the one hand, take the form of intrapersonal concepts, and on the other hand, the form of interpersonal concepts. such constructs are embedded in a conceptual system which evolves and affects the constructs as part of the system’s own evolution. the evolution of the concept system is changes in the connectedness of the concepts and in the strength of those connections. in that change, models play a central role, because through models, the concepts become projected onto actual, real situations. 3.1. structure: constructs concepts as complex structures are called here c-constructs. c-constructs are first and foremost connected to sets of attributes, where connecting links carry information about attributes and the strength of associations with those attributes. other elements of the system carry knowledge of regularities and relate concepts to each other in different ways: causally, through constrained determination (constrained covariation in a law-like manner) by constraining the use of concepts (e.g. conservation laws). these schemes are called determination constructs or, in shorthand, as d-constructs. cand d-constructs are the most elemental conceptual constructs of the systemic model, and as such, they offer no explanations or predictions by themselves. the task of explaining or predicting falls on models, which utilise the cand d-constructs as their constituents. models project concepts (c-constructs) onto phenomena to be explained, and through the success or failure of this projection, c-constructs are altered. the models, called here as m-constructs, are also conceptual constructs, but unlike cand dconstructs, are context dependent. the relationship of c-constructs’ to characteristic attributes is familiar from the prototype theory of concepts and is not only essential in describing personal concepts, but also important in describing scientific concepts. the basic level of attributes consists of simple and unstructured sets of attributes {a1, a2, ..., ak}, but more structured combinations can fall under more general schemes. these more general sets are subordinated under a more general (e.g. constraining) scheme, which here is typically a d-construct. in the course of concept development, these attributes are inherited, although some of the inherited attributes can be discarded as the concept evolves. these features resemble the dynamic framework approach to concepts (andersen & nersessian, 2000; andersen et al., 2006). the attributes are not strictly mutually exclusive, especially when c-constructs are not used together. however, the more important it is to use two c-construct together, the more difficult it is to maintain dissonant attributes; c-constructs are differentiated with regard to their attributes, as they should if they represent scientific concepts. d-constructs are general schemes which relate c-constructs to each other, typically in the form of causal connections or in the form of constrained determination (i.e. constrained co-variance with no causal i. t. koponen & t. kokkonen 146 | f l r dependence). in some cases, d-construct can simply constrain how a single c-construct can be applied. therefore, these constructs are essentially the carriers of theoretical knowledge (c.f. machery, 2009; rehder, 2003). the d-construct is the general template of the form of determination and is largely independent of the context, yet it prescribes how on can legitimately apply c-constructs to a given context through models. dconstructs play a crucial role in discriminating between attributes, because d-constructs connect cconstructs. through d-constructs the dissonances between attribute associations are revealed. m-constructs are designed so that they serve as models which explain phenomena or their selected properties. they use c-constructs, because concepts are needed to build models (nersessian, 2008). in some cases, m-constructs are also related to d-constructs, which then specify the relationship between cconstructs when more than one c-construct is involved. m-constructs are the basic vehicles for explaning or matching predictions with observable features of phenomena or, if one so wants, to select certain features of phenomena which fall under the explanatory power of a given m-construct. on the most advanced level, mconstructs are full-fledged scientific models. on the most basic level, m-constructs are simple and even selfexplanatory. however, in either cases, m-constructs are evaluated only against observational evidence {e1, e2, ..., ek}, which may either lend it support or lead to its rejection/inhibition. m-constructs compete against other available m-constructs in providing the most likely explanation of the evidence. figure 1. a schematic diagram (right) of cand d-constructs and their connections to attributes {a1, a2, ..., ak}, and (left) c-, dand m-constructs connected to each other and to sets of evidence {e1, e2, ..., ek}. links can be congruent (solid line) or dissonant (dashed line). the systemic view sees the knowledge system as a connected network of c-, dand m-constructs, where connections between these constructs can continuously change when the evidence changes. of course, the system can also reach stable states so that there are no changes in connections when additional evidence is available. changes in connections are based on locally effective rules, but the total effect depends on the global state of the system as whole (due to connectedness of the system). different concepts can then be expressed as different relational structures of the pieces or as different constellations of elements. within the systemic view, connections can be a type of positive constraint so that a connection strengthens the role of a given element. the connections can also be a type of negative constraint so that the connections weaken the role of the element. identifying connections and determining whether they are negative or positive must be based on empirical evidence. 3.2 dynamics: model competition and utility the evolution of the concept system is driven by the utility of m-constructs in explaining evidence (c.f. henderson, goodman, tenenbaum & woodward, 2010; ohlsson, 2009). the explanatory power of the m-construct changes with the changing amount of evidence to be explained. because different m-constructs can explain the same evidence, m-constructs compete against each other. if the context is simple and only i. t. koponen & t. kokkonen 147 | f l r little evidence need be explained, one can achieve this by using simple and only partially correct models that correspond to simple m-constructs. then, utility of the simple m-construct is better than utility of more complex ones (e.g. scientific ones), and the simpler ones are therefore more likely to be adopted. however, with the increasing complexity of the context and greater amounts of evidence to be explained, the complex models (m-constructs), which explain more, gain utility and become adopted. thus, it is important to note that in learning, the adoption of a model is a question not only of its correctness, but also of its utility (ohlsson, 2009, 2011). it is assumed here that different models and evidence to be explained are known in advance. many of the models may be inactive and much of the evidence unknown for the learner in the initial stages of learning. from point of view of modelling the learning, one can assume finite collection of possible models, some of them active and some inactive (cf. henderson et al., 2010). 3.3 concept convergence in a learning situation where concept learning and differentiation take place, learners share concepts in the group discussions within small groups. in learning, knowledge convergence is often considered crucial to forming shared, public concepts (weinberger et al., 2007; jeong & chi, 2007). the empirical data of concept differentiation indicate that knowledge also converges during this process; the ways in which one uses and understands the concepts become similar, at least to certain degree (koponen & huttunen, 2013). in the systemic view, this kind of knowledge convergence means that the ways in which the c-constructs link to other elements of knowledge and their attributes among the learning group become more similar during the learning process. here, the knowledge convergence discussed only to the extent that it concerns concept differentiation and via communication in small groups of three. therefore, in what follows, we concentrate on simple triadic communication patterns (discussed in more detail in section 4) and assume that the effect of convergence takes place mainly through utility of models. the communication is described simply as a consensus-based knowledge sharing where, through communication, all group members always adopt the model with the strongest utility. there is no threshold effect on adoption of the model. such a convergence model exaggerates the effect of communication on learning, but it is an adequate model for estimating the maximal expected effect of communication on concept differentiation. 4. empirical findings revisited: electric current and voltage research on learning scientific concepts and concept differentiation has been conducted in several ways and on different topics, but perhaps most extensively on the concepts electric current and voltage (cohen, eylon & ganiel; 1983; shipstone, 1984; engelhardt & beichner, 2004; koumaras, kariotoglou & psillos, 1997; lee & law 2001; mcdermott & shaffer, 1992; reiner et al. 2000;). some of the studies have focused on students’ explanatory models (cohen et al., 1983; engelhardt & beichner, 2004; mcdermott & shaffer, 1992; koumaras et al., 1997; shipstone, 1984), while some other studies have focused on ontological attributions (lee & law, 2001; reiner et al., 2000). the general outcome of these studies is that the concepts electric current and voltage are often mixed with personal, intuitive concepts or conceptions (or intrapersonal concepts) and, furthermore, are poorly differentiated. the impact of numerous empirical studies on deeper theoretical understanding of concept learning and differentiation, however, has been relatively modest for at least two reasons. first, although these studies have identified brought a variety of different types of models, intuitive conceptions and ontological attributions, they have failed to abstract from the empirical details general and generic features which could provide a broad enough theoretical perspective to understand the relationship between different views. therefore, for lack of a sufficiently broad theoretical perspective, discussions have often focused on the differences of a preferred theoretical perspective over that of some other perspective, rather than trying to provide a more integrated, progressive and broader theoretical framework that makes the partial accounts understandable (see e.g. chi & brem, 2009; gupta et al., 2010; ohlsson, 2009). we believe that many empirical findings can be captured within the systemic model when suitably idealised to reveal the essential generic features behind the multitude of details. furthermore, the systemic view can help us to understand i. t. koponen & t. kokkonen 148 | f l r how different aspects of concept learning are related. in what follows, we focus only on features of interest to advanced learners, typically those on an upper-secondary school level or first-year university level. 4.1. empirical results revisited and re-interpreted the purpose of the present work is to provide a new theoretical framework to discuss concept differentiation when learners’ concepts take on a scientific character. rather than report new empirical results, this study is re-uses and re-analyses already published empirical data (koponen & huttunen, 2013). the empirical data consists of nine students’ (upper secondary school) interviews about their conceptions of electric current and voltage in dc circuits. the students built dc circuits, observed their behaviour, and then proposed explanations for the observed brightness of light bulbs. the interviews were transcribed and analysed to identify the models students use to explain the behaviour. the nine students discussed their explanations in groups of three. the study consisted of three different contexts i-iii: i: light bulbs in series. the participants compared two variants (a single light bulb and two light bulbs) in terms of the brightness of the bulbs. this comparison produces evidence e1 and e2. ii: light bulbs in parallel. the first variant is again involves a single light bulb. the second variant involves two light bulbs in parallel. comparing the two variants yields evidence e1’ and e2’. iii: comparison of the brightness of light bulbs in series (i) and in parallel (ii). in the first variant, participants compare the brightness of light bulbs in series, and parallel circuits to the one-bulb case only. in the second variant, participants compare series and parallel cases to each other. this produces evidence e1’’ and e2’’. all six different types of evidence are referred to as an evidence set e = {e0, e1, e2, e0’,e1’ e2’, e0’’, e1’’, e2’’}, with e0, e0’ and e0’’ representing observations of the brightness of a single light bulb in each context (the brightest light bulb). further details about the empirical setup, design and excerpts from the student interviews are reported by koponen and huttunen (2013). these empirical studies reveal some common features, which answer the following questions: 1. what are the models students use to make predictions and explanations? 2. what are the determination (constraining or causal) schemes students employ as part of their models? 3. what attributes do students associate with the concepts, models or determination schemes they use? 4. how does communication in a small group (3 students) affect the relationships between concepts, models and determination schemes? the results of the re-analysis serve here to construct idealised sets of the mand d-constructs in contexts i-iii, with a summary of the results in table 1. a summary of the attributes revealed by the analysis is in table 2. m-constructs m1 and m2 are well-known electric current-based intuitive models found in many empirical studies (see koponen & huttunen, 2013, and references therein), while constructs m1’ and m2’ represent corresponding models, but are based on voltage (undifferentiated from current). these appear in relatively fewer cases, but are taken into account here. constructs m3 and m3’ are partially correct explanations, which take into account the role of components in determining the current. construct m3’, however, appears only once in the empirical data. construct m4 is the correct scientific model based on ohm’s law (d3) and kirchhoff’s laws i (d1) and ii (d2) which correctly differentiates between electric current and voltage. i. t. koponen & t. kokkonen 149 | f l r table 1 the mand d-constructs inferred from the empirical study (koponen & huttunen 2013) construct construct m1 the battery as a source of current. m1’ the battery as a source of voltage. m2 m1+ components consume current. m2’ m1’+ components consume voltage. m3 m1 + voltage over components creates current. m3’ m1’ + current over components creates voltage. m4 model based on ohm’s law + kirchhoff’s laws ki and kii. d0 constraining laws: conservation (of ―electricity‖ or current). d1 constraint: current is conserved in junctions/branches (kirchhoff i). d2 constraint: voltages in a closed loop equal zero (kirchhoff ii). d3 ohm’s law: u = ri or u/i = r. table 2 attributes a1-a9 inferred from the empirical study, with key word(s) used to characterise and identify each attribute. attribute key word attribute key word a1 stored a2 contained a3 consumed a4 conserved a5 degraded or diminished a6 divided and diminished a7 maintained a8 partitioned and conserved a9 generated, supported 4.2. representations as directed graphs most of the relevant elements found in the interviews can now be represented according to the systemic view and by using a directed graph model (dgm) to relate different c-, dand m-constructs to sets of attributes and evidence. the dgm is a representation, where connections between different elements are related through directed links which are either congruent or dissonant. the links and their direction provide information on how the elements interact. this has the advantage that dgm can serve as a computational template (koponen, 2013). an example of how dgm relates to different constructs and the most important links connecting them appears in figure 2. the links shown as solid lines are mutually supporting, congruent links; dissonant links, shown as dotted lines, point out contradictions. congruent links were recognised on the basis of how students combined these elements in different situations. the recognition of dissonant links was more problematic. most of the dissonant links are the interviewers’ interpretation of unavoidable logical contradictions rather than notions expressed by the students themselves (koponen & huttunen, 2013). different students’ conceptions can now be visualised as graphs with different node strengths. some typical students’ conceptions a-d found in the interviews and represented in this way appear in figure 3. cases a and b are the most common in contexts i and ii, while c and d usually occur only in context iii. d occurred in only two of the nine cases studied, while c (or constellations close to it) occurred in four cases (koponen & huttunen, 2013). an important aspect of the dgm representations is that they represent the students’ understanding as a constellation of c-, dand m-constructs and associated attributes with various strengths. of course, these strengths are idealisations of the systemic model, which only phenomenologically represents the apparent importance of a given construct as it can be identified in interviews, and only partial i. t. koponen & t. kokkonen 150 | f l r information about such strengths are available from the empirical data. nevertheless, such fine grained representations of students’ conceptions contain more information than do traditional ways based on written descriptions only. figure 2. directed graph model of all essential c-, dand m-constructs based on the empirical results, as reported in table 2. cand dconstructs are linked to attributes {a1, a2, ..., ak}, m-constructs are linked to sets of evidence {e1, e2, ..., ek}. links can be congruent (solid line) or dissonant (dashed line). construct c1 is current, and c2 is voltage. i. t. koponen & t. kokkonen 151 | f l r figure 3. some examples a-d of typical graphs representing students’ conceptions as projected on the dgm shown in figure 2 (these graphs are sub-graphs of the dgm). nodes are classified into three classes: strong, s > 0.7 (large circle); average, 0.3 < s < 0.7 (medium circle); and weak s < 0.3 (bullet). construct c1 is current, and c2 is voltage. 4.3. effects of communication the information about communication acts between the students (as it is available from the interviews) can serve to construct idealised communication patterns between students and to temporally locate the effects of communication on the students’ choices of models and attributions. analyses of data on the individual students’ conceptions have been published previously (koponen & huttunen, 2013), but data on communication is unpublished. a summary of the changes in models and how communication takes place is given in table 3. here, the re-analysed data, ordered in temporal sequences to reveal the communication acts, allows the identification of changes in cand m-constructs. unfortunately, the original data offer no detailed information on changes in sets of attributes. the results in table 3 show that the communication patterns in groups g1 and g2 are reciprocal in that all students exchange information in all directions. nevertheless, some one-person dominated patterns are evident, where a single student (s3 in g1 and s6 in g3) is more active than others. formally, such communication patterns between students p, p’ and p’’ can be modelled as a triad (see figure 4). with group g3, communication takes place reciprocally between all students and the communication pattern is dense. unfortunately, the empirical data here do not permit a more detailed analysis of the communication patterns. in what follows, the effect of communication is modelled as a triad; in one case as relatively sparse and one-person dominated, and in other case as dense and reciprocal. i. t. koponen & t. kokkonen 152 | f l r table 3 evolution of nine students’ conceptions 1-9 in groups g1-g3 of three students (s1-s3, s4-s6 and s7-s9) in contexts i, ii and iii, as idealised in terms of the dgm. communication events are shown as directed dyads i→ j from student i to student j, or as reciprocal dyads i ↔ j. group g1, with students 1-3 group g2, with students 4-6 group g3, with students 7-9 context s1 s2 s3 s4 s5 s6 s7 s8 s9 i a a (d) a a (d) a c a,c 1←3 2←3 3→1,2 4←5,6 5←6 6↔4 7←8,9 8↔9 9↔7,8 a c d c,b b d a,c c,a a,c ii 1↔3,2 2↔3,1 3↔1,2 4←6 5↔6 6→4,5 7↔8,9 8↔7,9 9↔7,8 c (d) d c b d 1←3 3→1 4←6 5←4 6↔4 c d d c c d c c,a a,c iii 1↔3,2 2↔2,3 3↔1,2 4→5,6 5←4 6←4 7↔8,9 8↔7,9 9↔7,8 d d d d d d c c c figure 4. patterns of students’ communication. the thickness of the arrows denotes the amount of communication. the communication pattern on the left is dominated by student p, while for students p’ and p’’ communication is sparse. the communication pattern on the right is reciprocal and dense between all students p, p’ and p’’. 5. computational embedding of systemic view in terms of dgm the directed graph model (dgm) can serve as a computational template; as a computational embedding of the systemic view to produce generically similar features found in empirical situations. in what follows, we briefly describe the computational features of such embedding; with similar type updating rules (see appendix) that have previously been introduced and motivated elsewhere (koponen, 2013). computational embedding transforms the qualitative notions contained in the systemic view into computational rules, quantifies the roles of model utility and theoretical guidance. concept learning and the degree of concept differentiation are monitored through two quantities: theoricity t and separability s. theoricity t describes the theoretical complexity of the concept, while separability s is connected to differentiation and, thus, to ontological shift. the pair of values (s,t) then specifies the learning path. i. t. koponen & t. kokkonen 153 | f l r in the dgm, c-, dand m-constructs are nodes connected by directed links. each node has a dynamically evolving strength, which determines its effect on the other nodes to which it is connected and, thus, the dynamics of the system. m-constructs are also connected to sets of evidence (see figure. 2). node strengths are updated after comparing m-constructs with the evidence, which means obtaining new evidence or reconsidering existing evidence. the dgm also has a memory effects in that the new strengths of the links and nodes and depend recursively on the previous values. furthermore, the simulations take into account also effects of communication between learners. in computational embedding information contained in one graphs affects the strengths of the nodes and thus the dynamics of another graph. to describe the state of the system and to characterise the evolution of the concepts, we must define several quantities in terms of node strengths and links. below is a short overview of these quantities. a complete description of the update rules and their definitions in terms of link and node strengths are given separately in the appendix. the details given in the appendix are not essential for understanding in general level how the model works, but the mathematical details give are needed to fully appreciate how the memory effects arise through connectivity from the global state of the network. the most important of the quantities are the theoricity t and separability s of c-constructs, which serve to specify the learning paths. theoricity t is a measure of the theoretical complexity of a c-construct that roughly describes the number of paths from c-constructs to m-constructs while taking into account the strengths of the links and nodes (see the appendix for details). separability s describes the degree of dissimilarity between c-constructs with regard to the different attributes associated with them (see table 2). if two c-constructs are connected to completely different sets of attributes, s is at its maximum value. both quantities are defined in a range from 0 to 1 so that t = 1 means full theoretic complexity (corresponding to the scientific use of given concept) and s = 1 means complete differentiation. the dynamics of the dgm depend crucially on the utility u of m-constructs, and the utility is the basis for model selection (the strengths of m-constructs depend on their utility, see appendix). first and foremost, the utility is proportional to the ratio of explained evidence to the theoretical complexity t of the c-construct while taking into account the relevant strengths of nodes and links. if the model explains most of the available evidence and its theoricity is low, the utility will be high and the model will be favoured in explanations. with more evidence to explain, the less complex models will generally explain less, thereby their utility is reduced. the extent to which the model takes into account conflicting evidence can be controlled with the parameter k, which also controls the effect of d-constructs on utility. if k is set to a high value, the state of the system is heavily guided by evidence and theoretical knowledge (i.e. d-constructs). with low values for k, conflicting evidence and theoretical information will be more or less ignored, thus (since d-constructs describing causal conservation laws will be less important), favouring simpler models. the importance of evidence (whether conflicting or not) can be also adjusted by weakening the links between m-constructs and evidence. one can also alter the order in which one encounters the evidence. in practice, parameter k is related to the potential of an individual student to make use of theoretical knowledge to construct explanatory models. learning also depends on the state of an individual student’s initial knowledge. the state of initial knowledge is taken into through the initial strengths of the different models m1-m4, usually so that simple models (such as m1 and m2) have high a priori strengths (they are then the preferred models), while complex models have low initial strengths. also, the initial strength of d-constructs affects how strongly theoretical knowledge will guide the learning process, an effect taken into account through parameter d’. in addition, students differ in their attentiveness to evidence, so that part of the evidence receives more weight than some other parts of the evidence. this is taken into account by giving weights to the evidence also. finally, the sequence of evidence and the order one encounters the evidence (i.e. the training sequence) affect the dynamics of the learning paths. the values of these parameters and their initial values serve to model the individual learners’ initial knowledge and their potential to make use of theoretical knowledge. in addition to the individual characteristics described above, communication between individuals affects the dynamics of learning paths. in the dgm, communication can also affect the strengths of the mconstructs. assuming pairwise communication between students, the stronger m-constructs affect the weaker ones so that the lower value is increased by the communication impact factor c. a value c = 1 means i. t. koponen & t. kokkonen 154 | f l r complete adoption of the highest utility models in communication, and c = 0 means ignoring completely the information provided through communication. the appendix explains the details of the communication model as part of the dgm. the computational model is idealisation and takes into account only the roughest features of concepts, models and communication. however, the model is constructed to include the most essential generic features and, as such, is capable of providing important insight into how different parts of a conceptual system and its internal connectivity affect concept differentiation. 6. results: dgm simulations of concept differentiation the directed graph model (dgm) and simulations based on it must make understandable the following generic features of concept learning and differentiation: 1) context-dependent dynamics of concept learning and differentiation (learning paths). the students’ conceptual states (as shown in figure 4) are context dependent in that they appear mostly in given contexts i-iii, with a given set of evidence (or observations) to be explained, and the state changes with the changes in set of evidence. 2) the dynamics and persistence of ontological change in attributions. changes in ontological attributions are indicative of concept differentiation. when it takes place, it leads robust and persistent learning outcome. 3) the effect of communication on concept differentiation. in two of three cases, a given group has a student with a more sophisticated conception and more differentiated concepts than the two other students have, but who eventually partially adopt that sophisticated conception. of course, the learning process entails many other details, but these generic features 1-3 are the most important and interesting ones that any model of concept learning and differentiation should explain. in what follows, we concentrate on simulating just such a process of concept learning and differentiation by using the dgm and monitoring the learning process through the theoricity t and separability s of concepts. 6.1. model parameters and initial conditions the dgm allows parameterisation of many different initial stages. the initial stages are described through the initial strength of m-constructs and d-constructs, and through the strength of the evidence. for initial model strengths, we studied here cases where m1 and m2 are strong models (strength 1.0 0.75), and m3 is of moderate strength (strength 0.5 0.25), and other m constructs are weak (strength 0.25 0.05). these cases are interesting, because they provide information on how initial, rather unsophisticated models such as m1 and m2 evolve during the learning process towards sophisticated models such as m4, and how concept differentiation relates to this change. this is also the learning path of most practical interest, a path from intuitive to scientific concepts. for initial d-construct strengths, we studied cases where d1, d2 and d3 are of equal strengths d’, varying from 1.0 to 0.4. in addition to initial values for d’, theoretical knowledge operates through congruent and dissonant connections, which can be tuned by parameter k (see the appendix and table 4) so that value k = 1 denotes the strongest guidance and, k = 0, no guidance at all. in addition to these parameterisations, the dynamics of the dgm and the learning paths depend on what we call here the training sequence, meaning evidence and the order in which one encounters it. the training sequences are constructed to correspond to empirical contexts i-iii (see section 4.1) so that each sequence iiiiii consists of evidence {e0, e1, e2, e0’, e1’ e2’, e0’’, e1’’, e2’’} where each element e0, e1, … is associated with strength e, specifying how much attention one pays to the evidence. if e = 1, then evidence is taken fully into consideration, but for 0.0 < e < 1.0, only partially. in the simulations, each event is repeated n times, (n = 3 or n = 4) and the sequence is then reversed in order to verify the i. t. koponen & t. kokkonen 155 | f l r permanence of learning (i.e. no reduction of values t and s for the reversed sequence, and no ―hysteresis‖ effect). thus, the computation consists of training sequences of form {nx(e’,e’,e’); nx(e’’,e’’,e’’); nx(e’’’,e’’’,e’’’); nx(e’’’,e’’’,e’’’); nx(e’’,e’’,e’’); nx(e’,e’,e’)}, with n = 3 consisting of 54 events and for n = 4 of 72 events. the training sequences of that form are completely specified by n and the set of values o = (e’,e’’,e’’’). in some cases, to test the hysteresis, an additional 38 events are added in random sequence, denoted by r. in summary, the parameters that specify the initial conditions are n and o, and the parameters affecting the dynamics are k and d’. 6.2. simulations of personal learning paths the personal (individual, without communication) learning paths are studied first in a case, where initial conditions favour models m1 and m2 with initial strengths of 0.75, but where model m2’ also has a substantial strength of 0.5. the learning paths begin from unsophisticated models, which closely correspond to patterns such as a and b (see figure 3), and then progress towards a more sophisticated patterns of type d. the learning paths of concept differentiation are monitored through the evolution of theoricity t and separability s. for comparison, estimates of the values of t and s corresponding to the empirical cases idealised as graphs a-d (see figure 3) are: t = 0.35 0.45, s = 0.10 0.20 for a; t = 0.35 0.45, s = 0.55 0.65 for b; t = 0.55 – 0.70, s = 0.60 – 0.70 for c; and t = 0.90 – 1.00, s = 0.95 – 1.00 for d. the learning paths are shown in figure 5 for training sequence parameterisations o1=(1.0,1.0,1.0), o2=(1.0,0.8,0.8) and o3=(1.0,0.5,0.1), with n = 3. the positions, where sequences corresponding to contexts i, ii and iii end, are denoted. theoretical guidance is studied for strong guidance k = 1.0 and 0.8, and for weaker guidance k = 0.50, while parameter d’ (for d-constructs) ranges from 1.0 to 0.4. the evolution of m-construct strengths, which reflects the competition between models that must explain more evidence, is shown in figure 6. the situation shown in figures 5 and 6 is asymmetric with respect to c1 (current) and c2 (voltage), with c1 always having a higher theoricity t than c2. this asymmetry stems from asymmetry in the initial strengths of m1 and m2, as shown in figure 6. this corresponds to the most frequent empirical situation in which students initially favour current-based models over voltage-based ones. one can also interpret the results in reverse way, with c2 having a higher theoricity and the roles of m1 and m2 reversed. however, this situation where voltage-based model is initially preferred over current-based model seldom occurs in empirical cases. for strong theoretical guidance (k = 1.0 or 0.8), together with close attention to observations (training sequences o1 and o2), learning and concept differentiation are successful. in such cases (figure 5, in the upper row, two cases on the left) learning is complete for concept c1 (current), which is fully scientific (t = 1) and completely differentiated (s = 1) from concept c2 (voltage). concept c2 is nearly scientific (t = 0.6), and with some extra training (shown in grey in figure 5), it rapidly becomes a fully scientific (t = 1) concept. the learning paths are step-wise, with clearly distinguishable stable stages in theoricity t with increasing separability s. a sequence corresponding context i is already enough for relatively advanced differentiation (i.e. ontological shift), although theoricity t may remain low. there appears to be a threshold of s = 0.7-0.8, which one can reach even with moderate development in theoricity. this threshold shows that one can achieve nearly complete separability and good differentiation (i.e. nearly complete ontological shift) even though the learning is otherwise still incomplete. when theoretical guidance decreases, k = 0.5 and d’ = 0.60 (figure 5, lower row, middle) or k = 0.8 and d’ = 0.4 (figure 5, upper row, right), the theoricity of concepts c1 and c2 remains low for the training sequence (black dots), but again, with extra training (grey dots), improvement is possible. this trend shows that, eventually, even moderate theoretical guidance is effective, but then more training is needed. however, the order one encounters the evidence in training is not crucial, if the context iii is involved. when theoretical guidance is low (k = 0.5, d’ = 0.4) and little attention focuses on evidence in case of context iii, very little learning takes place, irrespective of the amount of training. this situation is shown in figure 5 in the lower right corner. the evolution of m-constructs in figure 6 corresponds to learning paths in figure 5. the initially dominant m-constructs m1 and m2 remain dominant until the end of the sequence corresponding to context i. t. koponen & t. kokkonen 156 | f l r i. during the sequence corresponding to context ii, models m1’ and m2’ also become active and grow stronger. however, when sequence iii begins, with strong theoretical guidance, the initial models cannot compete with m4 (fully scientific model), which eventually dominates when the sequence corresponding to context iii ends. this does not occur with low theoretical guidance (figure 6, right column). however, with extra training, m4 is eventually enforced in cases of moderate theoretical guidance also (figure 6, lowest row), but not for the least guidance (figure 6, lower right corner). it is noteworthy that m3’ is never activated, which is in line the empirical finding that such a voltage-based model is only seldom encountered. figure 5. the theoricity t and separability s of concepts c1 (bullets) and c2 (boxes) in the case of six different learning paths with given parameters k and d’ that control the strength of theoretical guidance. the upper row shows cases where k ≥ 0.8 is always relatively high but d’ varies from 1.0 to 0.4. in the lower row, k also varies from a high value of 0.8 to a lower value of 0.5. the initial values of the model strengths and strengths of the observations are different in cases shown in the left, middle and right columns (corresponding model strengths are shown in figure 6). left column: initial values of model strengths favour models m1 and m2’ with strengths of 0.5, while other models have a weaker but equal strength of 0.25. the observations of events i-iii are strong (link strengths have a value of 1). middle column: model strengths as in the left column, but m3, m3’ and m4 are reduced to 0.15, observations i-ii are strong (1), but iii is only moderately strong (0.75). right column: otherwise similar to the middle column, but the observations in case iii are weak (0.10). the training sequence from i to iii (end points of each sequence are marked in the figure), with three repetitions for each event appear in black dots. the training sequence testing the permanence of learning from i to iii, then back from iii to i, and one random sequence appears in grey dots. the values corresponding to the empirical results of configurations a-d (see figure 3) are indicated (two symbols for each are located in the pairs of the lowest estimated and highest estimated values for t and s). i. t. koponen & t. kokkonen 157 | f l r figure 6. model evolution of the learning paths shown in figure 5 with parametrisations for k and d and different training sequences o1, o2 and o3 as indicated. the color represents the strength of a given model. the number of steps in the simulation appears on the vertical axis thus indicating the ordering of the sequence. learning paths with slightly different initial conditions from the cases in figure 5 are shown in figure 7. in the cases shown in the upper row, the m-constructs m3 and m4 are slightly weaker than in the case shown in figure 5. in the cases shown in the lower row, m-constructs m1 and m2 are of nearly equal strengths, which makes the initial stage more symmetric with respect to c1 and c2. in both cases, the attention to events corresponding to contexts ii and iii becomes weaker from left to right, represented as parameterisations o3 (as in figure 5) and o4 = (1.0,0.8,0.5). in addition, the training sequences is now such that each event is reproduced three (n = 3, black dots) or four times (n = 4, grey dots). the corresponding evolution of m-constructs is shown in figure 8. compared to the cases shown in figure 5, one can observe some interesting differences. in the case shown in figure 7, in the upper left corner, learning consists mostly of ontological shift through end of the sequence corresponding to context ii. after that, when the sequence corresponding context iii begins, theoricity t rapidly increases because m4 rapidly gains strength (see figure 8) due to strong theoretical guidance. eventually, when the training sequence ends, learning is again complete. the training sequence with n = 4 leads to higher theoricity t of concepts in i and ii, but interestingly, to a slower increase in theoricity in iii than in cases with n = 3 because with n = 4, m3’ grows strong during i and ii, which slows the adoption of m4. this ―overlearning‖ effect is most pronounced in the case shown in figure 7, in the upper right corner, where theoretical guidance is low and attention paid to events is also low. in this case, more frequent repetition of events with n = 4 leads to deterioration of the learning results, and learning stagnates on the low theoricity t of c1 and c2. ontological shift, however, advances and eventually, separability s = 0.8 is reached. such a situation corresponds to what occurs in real learning; too much focus on overly simple tasks, which reinforces unsophisticated models, may lead to the persistent and robust use of under-developed models and conceptions. i. t. koponen & t. kokkonen 158 | f l r figure 7. three deterministic cases of learning paths with a given k. the figures show theoricity t and separability s of c-constructs c1 (bullets) and c2 (boxes) from figure 2 and indicates the values corresponding to empirical results a-d (see figure 3). construct c1 is current, and c2 is voltage. the initial conditions favour models m1, and m2’ (voltage based). figure 8. model evolution of the learning paths in figure 7 with parameterisations for k and d and different training sequences o3 and o4 as indicated. the darkness represents the strength of a given model. the numbers of steps in the simulation appear on the vertical axis, thus indicating the ordering of the sequence. the lower row in figure 7 shows some interesting situations, where repetition temporarily leads to the deterioration of learning results, when simple situations recur after the sequence iiiiii. we briefly refer to this as ―hysteresis‖ in learning. eventually, however, (figure 7, lower row, two cases on left) i. t. koponen & t. kokkonen 159 | f l r complete learning with t = 1 and s = 1 occurs and learning becomes permanent, no longer affected by further repetitions. in the case of weak theoretical guidance and weak learning from events (figure 7, lower right corner), incomplete learning occurs, with stable learning resulting at t = 0.4 and s = 0.6. this is again due to ―overlearning‖ of incomplete models, which prevents the adoption of the more sophisticated model m4. similar results can also be observed for other cases with moderate k and moderate attention to observations; repeating simple situations i and ii many times before encountering more complex situation iii, may reinforce the incomplete models m1-m3 or m1’-m3’ so much that further development can no longer take place. this shows that repetitions of training sequences can have detrimental consequences on learning if initial theoretical guidance is too low. the examples discussed above are asymmetrical situations, where concept c1 (current) is the favoured concept, while c2 (voltage) is initially less developed, and remains largely as is during further evolution. this is the most common situation in learning, although the roles can sometimes be reversed. however, because the dgm is symmetrical with respect to c1 and c2 (see figure 2), a reversed situation where c2 has stronger theoricity than c1 is quite to similar if c1 and c2 simply switch roles. also, the symmetrical situations can occur can closely follow the results in figures 5 and 7, with the learning paths for c1 and c2 then simply overlapping. 6.3. simulations of communication effects the effects of communication on learning paths are simulated by using the sparse and dense communication pattern between members p, p’ and p’’ in a group of three (see figure 4), and two impact factors c = 0.2 and c = 0.75 for communication. the effect of communication is tested on cases where theoretical guidance is strong or moderate (figures 5 and 7, upper row, in the right column). the results of learning paths are shown in figure 9, and the evolution of m-constructs in figure 10. these figures show that even the effect of dense communication with a high impact c = 0.75 only moderately affects the learning paths. the most obvious effect is that if one member p in the group has a learning path which is strongly theoretically guided, thus reaching high values for t and s, the other cases tend to learn from that specific case and improve their learning and, consequently, reach higher values for t and s than without communication. eventually, members p’ and p’’ who are less successful (figures 5 and 7, in the lower right corner) than member p also achieve complete learning owing to the communication. this happens equally well for sparse and low-impact communication as for dense and high-impact communication. of course, this occurs only in cases with one successful learner in the group. these features appear to be in concordance with the empirical findings, although the empirical findings presently allow no more detailed comparisons. in the learning model, which is simply biased toward adopting the strongest model, the good learning result may temporarily worsen (see figure 9, upper right corner). however, this is a transient effect, and the learning path eventually evolves toward complete learning. i. t. koponen & t. kokkonen 160 | f l r figure 9. learning paths with the effect of communication taken into account. different figures represent paths with different parameterisations for k and d and different training sequences o1, o2, o3 and o4. figure 10. model evolution of the learning paths in figure 9 with different parameterisations for k and d and different training sequences o1, o2 and o3. the darkness represents the strength of a given model. the number of steps in the simulation appears on the vertical axis, thus indicating the ordering of the sequence. i. t. koponen & t. kokkonen 161 | f l r 6.4. summary of simulation results the results based on the dgm agree with the following central empirical findings of concept learning and differentiation: 1. context-dependent dynamics. this is apparent in the strong dependence of paths on the learning sequences. complete learning takes place only in sufficiently rich contexts (e.g. case iii), whereas in narrow contexts (e.g. cases i and ii), learning is moderate or incomplete. this is a consequence of model competition and the greater utility of complex models in complex contexts. 2. the persistence of ontological shift and concept differentiation (s ≈ 1). in the dgm, this is a direct consequence of the guidance of d-constructs and their ―memory effect‖, retaining the memory of successful applications of d-constructs. the persistence of the ontological shift agrees with the empirical findings. however, the ontological shift in attributions is not a driving force of concept learning, but an outcome of a learning process driven by theoretical knowledge. 3. communication affects individual learning paths and enables less advanced members of the group to adopt more advanced m-constructs from the most advanced member of the group. thus communication improves learning, although the effect in the cases studied here is not particularly strong. in summary, the dgm model reproduces the generic features of interest in concept learning and differentiation, and demonstrates that these features are associated with the guidance of theoretical knowledge, model utility and the memory effects of success in using models. 7. discussion and conclusions the model presented here is based on the systemic view, where concepts are viewed as complex, dynamically evolving structures. the model is constructed to capture generic aspects of concept learning and differentiation as exemplified in the case of learning two closely related scientific concepts – here, electric current and voltage. the generic features of most interest in need of explanation are: 1) the robustness of certain simple ways to use concepts to provide explanations in simple situations, a phenomenon usually assigned to robustness of misconceived ontological classes, 2) the context-dependent dynamics of change and requirement to encounter complex enough situations to effect in the change, and 3) the robustness of ontological shift once it has occurred. we suggest here that in order to understand these features and the dynamics of the change, we must develop a rich and complex enough model of concepts. on the one hand, the model of concepts takes a form of simple and nearly self-explanatory concepts, but on the other hand, a form of complex structures, dependent on other concepts. the systemic view is embodied by the use of the well-known case of the differentiation of electric current and voltage as concepts describing the behaviour of simple dc circuits. in that, we use here reanalysed empirical data. the re-analysed (and partly re-interpreted) empirical results are then represented by using different conceptual elements, or constructs: c-constructs, which stand for concepts, d-constructs for causal schemes and law-like theoretical schemes, and m-constructs, which are model-like structures that use cand d-constructs as integrated parts. as a formal representational model for these constructs and their mutual relationships, we introduce a directed graph model (dgm). in the dgm, concepts are nodes in the graph connected by directed links to other conceptual elements. the dgm serves as a computational template to simulate concept learning and differentiation and their dynamics. the stability of certain properties of concepts, traditionally considered robust ―misconceptions‖, and their dependence on contexts is now seen as related to the complex interplay of different conceptual elements. change is driven by competition between m-constructs (models) and by how available evidence governs it. however, how this is reflected in the theoryand attribute-relatedness of concepts depends on how those m-constructs employ concepts (i.e. how the concept projects onto the actual evidence). thus, d-constructs (theoretical knowledge) are central. all these aspects are recognised in current i. t. koponen & t. kokkonen 162 | f l r cognitively oriented views of concept learning, but are usually discussed separately or as unrelated views. the present study strongly suggests unifying these views and treating concepts as complex, multifaceted and dynamic structures. finally, the present work suggests that the theoretical background developed for research on concept development has important implications for the ways in which researchers and instructors view the learning process and how, on this basis, they design teaching solutions. the results point to the crucial role of theoretical knowledge in guiding concept learning and, furthermore, show that ontological shift, while an important part of learning, is not the primary driving force of learning, but is rather a consequence of more fundamental changes in the conceptual system. this suggests that theoretical structures and model utility should receive more attention in designing instructional solutions. on the other hand, it is clear that initial conceptions need not be actively ―unlearned‖; they can instead serve as a natural and useful starting point for the transformation. for teaching and instruction, one important message lies in the role of the training sequence in determining learning paths. the details of the training sequence, and their repetitions, do matter in the initial stages of learning if the guidance of theoretical knowledge is low. then, too much repetition of overly simple situations to explain may lead to ―overlearning‖ of unsophisticated concepts and models, and effectively prevent the acquisition of more advanced model. the results of the simulations, interpreted within the theoretical framework of the systemic model put forward here, suggests that designing specific training sequences which help one to ―unlearn‖ unsophisticated models is unnecessary; rather, what is needed is a training sequence which gradually and at suitable stages of learning introduces more challenging learning situations, where the utility of more advanced and scientific concepts and models becomes apparent. the theoretical positions discussed and suggested here directly impact learning and instruction by clarifying the degree to which degree ontological shift drives the learning process and to which degree it should be considered a consequence of more fundamental, theory-driven learning process. also, the question of to what extent differentiation and concept learning take place through the evolution of existing structures, and to what extent the learner must receive these structures instead of constructing them receives clarification. briefly, if the systemic view is correct, it suggests that ontological shift takes place, but is a consequence of theory-driven learning. the learner must receive complex theoretical structures through instruction and see their utility in complex enough situations to warrant adopting them. such results are practical in that they guide teaching and the development of teaching solutions; they provide support to some of the well-known teacher-centred solutions (the role of the teacher in providing models to organise new knowledge and in familiarising the students with complex theoretical models), while showing the indispensability of a rich context and context variation in the construction of explanatory models and the role of predictions and observations in learning. these notions, even without detailed suggestions for training sequences and instructional solutions, demonstrate that the choices to employ a certain theoretical framework to understand concept learning and differentiation are not neutral. rather, they have fundamental consequences for how learning and instruction are conceived, how their purposes and goals are viewed, and how our attention is guided towards crucial generic features of learning and its dynamics. keypoints concepts are considered complex structures, which are projected differently in different contexts. concept differentiation can be modelled when embedded within a systemic view on concepts. theoretical guidance and theoretical schemes are crucial for concept differentiation. ontological shift is a consequence of theory-guided learning process. i. t. koponen & t. kokkonen 163 | f l r robust misconception are stable dynamic states of the concept system attention must be paid on training sequences in learning. too frequent use of overly simple situations in training will stagnate the concept learning in robust states corresponding misconceptions. references andersen, h. barker, b. and chen, x. (2006). the cognitive structure of scientific revolutions. cambridge, ma: cambridge university press. andersen, h. and nersessian, n. j. (2000). nomic concepts, frames, and conceptual change. philosophy of science, 67, s224-s241. brown, d. e., & hammer, d. (2008). conceptual change in physics. in s. vosniadou (ed.), international handbook of research on conceptual change (pp. 127–154). new york: routledge. carey, s. (2010). the origin of concepts. new york, ny: oxford university press. chi, m. t. h., & slotta, j. d. (1993). the ontological coherence of intuitive physics. cognition and instruction, 10, 249-260. chi, m. t. h. (2005). commonsense conceptions of emergent processes: why some misconceptions are robust. the journal of the learning sciences, 14, 161-199. doi: 10.1207/s15327809jls1402_1. chi, m. t. h. (2008). three types of conceptual change: belief revision, mental model transformation, and categorical shift. in s. vosniadou (ed.), international handbook of research on conceptual change (pp. 35–60). new york, ny: routledge. chi, m. t. h., & brem, s. k. (2009). contrasting ohlsson's resubsumption theory with chi's categorical shift theory'. educational psychologist, 44, 58 — 63. doi: 10.1080/00461520802616283. cohen, r., eylon, b., & ganiel, u. (1983). potential difference and current in simple electric circuits: a study of students’ concepts. american journal of physics, 51, 407-412. danks, d. (2010). not different kinds, just special cases. behavioral and brain sciences 33, 208-209.doi: 10.1017/s0140525x1000052x engelhardt, p. v., & beichner, r. j. (2004). students’ understanding of direct current resistive electrical circuits. american journal of physics, 72, 98-115. doi: 10.1119/1.1614813. gopnik, a., & meltzoff, a. n. (1997). words, thoughts, and theories. cambridge, ma: mit press. gupta, a., hammer, d., & redish, e. f. (2010). the case for dynamic models of learners’ ontologies in physics. the journal of the learning sciences, 19, 285-321. doi: 10.1080/10508406.2011.537977. henderson, l., goodman, n. d., tenenbaum, j. b., & woodward, j. f. (2010). the structure and dynamics of scientific theories: a hierarchical bayesian perspective. philosophy of science, 77, 172–200. hoyningen-huene, p. (1993). reconstructing scientific revolutions: thomas s. kuhn’s philosophy of science. chicago, il: the university of chicago press. jeong, h & chi, m. t. h. (2007). knowledge convergence and collaborative learning. instructional science, 35, 287–315. doi: 10.1007/s11251-006-9008-z. keil, f. c. (1989). concepts, kinds and conceptual development. cambridge, ma: mit press. koponen, i. t. (2013). systemic view of learning scientific concepts: a description in terms of directed graph model. complexity, 19, 27-37. doi: 10.1002/cplx.21474. koponen i. t. and huttunen l. (2013). concept development in learning physics: the case of electric current and voltage. science & education, 22, 2227-2254. doi: 10.1007/s11191-012-9508-y. koumaras, p., kariotoglou, p. & psillos, d. (1997). causal structures and counter-intuitive experiments in electricity. international journal of science education, 19, 617–630. lee, y., & law, n. (2001). explorations in promoting conceptual change in electrical concepts via ontological category shift. international journal of science education, 23, 111149. machery, e. (2009). doing without concepts. oxford: oxford university press. murphy, g. l. (2004). the big book of concepts. cambridge, ma: mit press. nersessian, n. (2008) creating scientific concepts. mit press: cambridge, ma. i. t. koponen & t. kokkonen 164 | f l r ohlsson, s. (2009). resubsumption: a possible mechanism for conceptual change and belief revision. educational psychologist, 44, 20-40. doi: 10.1080/00461520802616267. ohlsson, s. (2011). deep learning: how the mind overrides experience. cambridge, ma: cambridge university press. rehder, b. (2003). categorization as causal reasoning. cognitive science, 27, 709–748. reiner, m., slotta, j. d., chi, m. t. h., & resnick, l. b. (2000). naive physics reasoning: a commitment to substance based reasoning. cognition and instruction, 18, 1-34. shipstone, d. m. (1984). a study of children’ s understanding of electricity in simple dc circuits. european journal of science education, 6, 185198. slotta, j. d., & chi, m. t. h. (2006). helping students understand challenging topics in science through ontology training. cognition and instruction, 24, 261-289. smith, c., carey, s., & wiser, m. (1985). on differentiation: a case study of the development of the concept of size, weight and density. cognition, 21, 177-237. smith, e. e., & medin, d. l. (1981). categories and concepts. cambridge ma: harvard university press. weinberger, a., stegmann, k., fischer, f. (2007) knowledge convergence in collaborative learning: concepts and assessments. learning and instruction, 17, 416-426. appendix: dgm update rules the dynamics of the dgm is determined by the update rules of the node strengths and weights of connecting links between the nodes. in the dgm, c-, dand m-constructs are nodes, which are connected by directed links. each node i has dynamically evolving strength si, which determines its effect on the other nodes to which it is connected and, thus, the dynamics of the system. node strengths s (the subscript is omitted if not essential) are updated after each ―event‖ e in the set of all events e, which means obtaining new evidence or reconsidering the evidence (i.e. any kind of comparison with the evidence). variable e is treated as a running index that keeps track of the encounter with evidence. the strength of the previous step s(e-1) is then updated to a new value s(e-1)  s(e) that corresponds to evidence e. the congruent link between nodes i and j is described by the value aij =1, while dissonant link has aij = -1. the following quantities are then defined entirely in terms of node and link weights. 1. theoricity t is the theoretical complexity of the c-construct. the more there are d-constructs and m-constructs, which are connected to c-constructs, the greater is the theoretical complexity of the cconstruct (i.e. the greater is its theoricity). quantitatively, within the dgm, theoricity t can be quantified as the number of directed paths from the c-construct to the m-construct and to their respective strengths. in some of the paths, the d-constructs are also involved, which increases their theoricities. theoricity tc of the c-construct at node c ϵ c (index c refers to c-constructs) is: the first term represents one-step paths from the models (m ϵ m) to the c-construct. the second term represents two-step paths through the d-construct (d ϵ d). note here that sc = 1. the theoricity of a model, needed in what follows as part of the dynamic update rules for the dgm, is defined similarly, but now sc  sm and inverted directed paths from the model to the cand d-constructs are counted (see table iv). 2. separability s measures the degree of dissimilarity between the set of attributes associated with two c-constructs represented by nodes c and c’. it is defined in regard to attributions only, as in the i. t. koponen & t. kokkonen 165 | f l r prototype theories. separability is operationalised as a suitably normalised number of unshared elements so that for fully differentiated concepts, s = 1, while for similar concepts, s = 0, defined as: here, the element aac represents attributions (i.e. the values of attributes a, to be defined later) linked to c-construct c, and n is a normalisation factor . in calculating separability, n takes into account the total strength of the attributions so that s ≈ 1 represents strong attributions with totally dissimilar attributions, while s << 1 can represent either totally similar or weak attributions. 3. utility u of the models in providing explanations depends on the ratio of explained facts to the model’s theoricity t. the utility of model m ϵ m is defined as where e is an event in set e, and e’ is a node which groups events together (see figure 2). dissonant (negative) links are denoted by a’ij. the first term in the sum represents direct congruent paths to the models (i.e. explanations), the second term direct dissonant paths, the third term two-step paths through e’, and the last term two-step paths through d-constructs. parameter k ϵ [0,1] controls the effect of dissonant links and the d-constructs on utility. if the model explains most of the available evidence and its theoricity t is low, its utility is high. however, with a growing set of evidence to explain, models with low theoricity generally fail, thereby reducing their utility. utility is the basis for model comparison and selection. the updating rules for the node strengths determine the dynamics of the graph and, consequently, the evolution of quantities u, t, and s, which depend dynamically on the node strengths. the updating rules also contain a memory effect, and the new values s(e) with evidence e depend recursively on the past values s(e-1), or s in shorthand. the updating rules are defined as follows: 4. the update rule for strength sm of m-constructs m is based on bayesian-type selection criteria (koponen, 2013; henderson et al., 2010) and depends on the utilities of the models. each m-construct has a certain expected plausibility or probability sm, which is updated to sm(e) when more evidence e in the form of observations becomes available. the new plausibility with evidence e is evaluated according to the bayesian rule so that if um(e) is the utility of model m when e is known, the model strength is then updated to where the sum in the denominator ensures its normalisation. generally, the more complex cconstructs provide more alternatives for m-constructs, so the bayesian rule favours ―simple‖ m-constructs. however, this may change when observations accumulate and more complex m-constructs explain more. the initial conditions of the system are its prior strengths and utilities, which must be deduced from available empirical data (based e.g. on interviews). 5. strength sd with evidence e depends on the connections of the d-construct to other d-constructs and m-constructs. the update rule for it is i. t. koponen & t. kokkonen 166 | f l r where the first and the last sums take into account the fact that dissonant connections reduce strength, while the second sum takes into account the fact that connections to successful m-constructs increase strength. the factor (1-sd) takes into account the ―memory effect‖; the use of given d-construct in successful m-constructs increase the value of sd, which does not decrease again. this models the known effect that the successful application of theoretical knowledge increases confidence in that knowledge. 6. attribute strengths sa are updated by taking into account one-step congruent (second sum) and dissonant (first sum) paths from the models, where parameter k controls the weight of dissonant paths (compare to utility). attributions aac are then defined on the basis of attribute strengths, where the first term represents two-step paths, and the last term, three-step paths from the attributes to the c-constructs through the d-constructs (see figure 2). attributions serve as a basis for calculating the separability s of a given pair of concepts (see definition 3 above). 7. the effect of communication on the dynamic of the dgm operates through the strengths of the mconstructs. at each step, where pairwise communication between students p and p’ is assumed, the stronger m-constructs affect the weaker ones. this is done by setting new values to the sm and s’m of p and p’, respectively, so that the larger value of them remains unchanged, but the lower value increases by factor c max{0,sm s’m}, where c is the communication impact factor. the update takes place before applying the bayesian rule in step 4. this means that the update rule can be interpreted as learning by adopting better explanatory models before comparing utilities. the update rules for the strengths of the mand d-constructs drive the dynamic evolution of the graph. table 4 summarises these strengths and other quantities defined in 1-6. table 4 definitions of quantities t, s and u, as well as update rules for node strengths s, are given for dand mconstructs and their attributes. the attributes of the c-construct are given by aac, and the separability of c and c’ by scc’. subscripts c, m, d and a denote c-, mand d-constructs and attributes, respectively. congruent links between i and j are denoted by aij = 1, while for dissonant links, a’ij = -1. name definition theoricity tk if k=m then i=c ; if k=c then i=m strength of d sd utility um strength of m sm strength of a sa attribution of c aac separability scc’ where microsoft word castello et al_publication.docx         frontline learning research vol.3 no. 3 special issue (2015) 39 54 issn 2295-3159   corresponding author: montserrat castelló, facultat de psicologia, ciències de l’educació i l’esport. blanquerna. universitat ramon llull, císter 34. 08022. barcelona. phone: +34932533000, fax: +34932533031, email: montserratcb@blanquerna.url.edu doi: http://dx.doi.org/10.14786/flr.v3i3.149     researcher identity in transition: signals to identify and manage spheres of activity in a risk-career montserrat castellóa, sofie kobayashib, michelle k. mcginnc, hans pechard, jenna vekkailae, & gina wiskerf auniversitat ramon llull, spain buniversity of copenhagen, denmark cbrock university, canada dalpen-adria-universität klagenfurt, austria euniversity of helsinki, finland funiversity of brighton, uk article received 31 january 2015 / revised 14 june 2015 / accepted 9 july 2015 / available online 28 september 2015   abstract within the current higher education context, early career researchers (ecrs) face a ‘risk-career’ in which predictable, stable academic careers have become increasingly rare. traditional milestones to signal progress toward a sustainable research career are disappearing or subject to reinterpretation, and ecrs need to attend to new or reimagined signals in their efforts to develop a researcher identity in this current context. in this article, we present a comprehensive framework for researcher identity in relation to the ways ecrs recognise and respond to divergent signals across spheres of activity. we illustrate this framework through eight identity stories drawn from our earlier research projects. each identity story highlights the congruence (or lack of congruence) between signals across spheres of activity and emphasises the different ways ecrs respond to these signals. the proposed comprehensive framework allows for the analysis of researcher identity development through the complex and intertwined activities in which ecrs are involved. we advance this approach as a foundation for a sustained research agenda to understand how ecrs identify and respond to relevant signals, and, consequently, to unravel the complex interplay between signals and spheres of activity evident in struggles to become researchers in a risk-career environment. keywords: researcher identity; identity development; signals; spheres of activity; risk-career castelló et al   | f l r         40   1. introduction the position of early career researchers (ecrs) has always been challenging and involves many difficulties that must be conquered in order to secure personally and intellectually satisfying positions and a strong sense of self as a researcher. however the situation has become particularly acute over the past few decades as higher education systems have been confronted with changing worldwide circumstances due to the requirements of the knowledge society and various economic and political constraints (cantwell, 2011; winter, 2009). changes are especially dramatic with respect to the nature of researcher education and identity development for ecrs who struggle with the demands of global mobility, the lack of stable or permanent positions, and the need to consider alternative careers (introduction, this issue). ecrs are now embarked upon what we define as a ‘risk-career’ (weber, 1947), rather than, as previously, a relatively more predictable academic career. in this changing context, traditional milestones that enabled ecrs to build their identities are disappearing or subject to reinterpretation. ecrs need to identify or reinterpret signals (yorke, 2009) from institutions and academic communities. signals related to expectations, constraints, and opportunities may cue performance and progress toward professional skill development and potential career directions. although studies focusing on identity development or identity trajectories have grown exponentially in recent years, research in the field has not yet resulted in a comprehensive framework that integrates identity and signals or offers a comprehensive way to analyse researcher identity as it unfolds across the different systems or spheres of activity in which ecrs participate. the specific aim of this article is to explore researcher identity in relation to the signals ecrs perceive across different spheres of activity as they attempt to manage a risk-career. our overarching purpose is to offer a comprehensive framework useful for analysing how signals can be identified and used to build a researcher identity in a risk-career, one where career trajectories are less certain than they were. consistent with the position presented in the first article of this special issue (introduction), we assume the definition of early career researchers (ecrs) presented by the earli special interest group researcher education and careers to include individuals with up to 10 years of research experience, which means doctoral students, postdoctoral researchers, newly-hired lecturers, as well as professionals in universities and other employment. globally, stable academic careers have dwindled and a range of alternative academic positions has emerged: contract teaching, contract postdoctoral research, teaching-only lecturer positions, and administrative positions related to research, teaching, or student services. in the non-academic context, emerging types of and contexts for employment include business, government, non-governmental organisations, banking, industry, and previously unknown entities (e.g., start-up companies). the existing research literature base provides little information about the experiences of individuals facing uncertain employment or alternative academic positions. prior studies shed light on quite narrow aspects, such as international postdoctoral employment in enterprise modes of academic production (cantwell, 2011; porfilio, gorlewski, & pineo-jensen, 2013), or critical interactions that shape careers for early academics and new teaching staff (hemmings, hill, & sharp, 2013). little is known about required competencies, employment satisfaction, the range of skills required, and, most specifically, ways to formulate a researcher identity in this changing environment. it is a paradox that on the one hand research and advanced education is of ever-growing importance for knowledge-based economies, while on the other hand the attractiveness of academic working conditions is decreasing and it is becoming more difficult for ecrs to embark on stable careers. ecrs are exposed to contradictory signals about expectations, constraints, and opportunities in relation to their careers. the knowledge-based economy boosts an expansion in training positions for researchers (cyranoski, gilbert, ledford, nayar, & yahia, 2011), which signals to potential students that it is worthwhile to start doctoral training. however, once they become students, these individuals may learn that the increase in training positions is not matched by an increase in stable jobs for researchers. they realise, sometimes too late, that castelló et al   | f l r         41   they have chosen a risk-career in which they might face the danger of precarious positions where they may or may not feel they can contribute as researchers. in a risk-career, traditional mechanisms are fading for individuals to identify as ‘members’ of a collective, and for others to attribute or acknowledge such membership (castelló & iñesta, 2012). ecrs are positioned differently to those already established within their fields and may hold various competing interests and identity constructions (archer, 2008). ecrs are not only ‘becoming’ but also ‘unbecoming’ (archer, 2008), meaning that they are not always recognised by others in terms of the dominant structures and practices. ecrs may also unbecome by their own choice as a possible form of resisting the dominant practices (archer, 2008; danaher, 2015; pyhältö & keskinen, 2012). although there is a vast body of literature from the field of higher education about professional identity, academic identity, authorial or writing identity, emotional identity, and other related concepts, many studies lack a clear definition of what identity means, how the notion is operationalised or analysed, and the underlying theoretical and methodological assumptions (trede, macklin, & bridges, 2012). moreover, it is common that studies do not focus on identity as a whole, but rather tend to conceptualise it as a multidimensional construct that can be applied to different activities and systems in which particular experiences are developed. there is a need for an integrative and comprehensive framework to identify and analyse signals and changes in identity. in order to understand such mechanisms, we first present a comprehensive framework of the notion of researcher identity, produced by analysing spheres of activity related to researcher and career development to account for theoretical assumptions about researcher identity in a risk-career; and second we illustrate this framework through eight identity stories drawn from our earlier research projects. 2. a comprehensive framework for the study of researcher identity in a changing environment we conceptualise researcher identity to be a dynamic and social process that develops through participation in different disciplinary and academic communities. this conceptualisation also implies that researcher identity is relational and discursively constructed through a recursive and iterative process of subject positioning, which involves a process of self or subject constructions that influence the ways people interpret the present and learn for the future (harré, moghaddam, cairnie, rothbart, & sabat, 2009; holland, lachicotte, skinner, & cain, 2001; lave & wenger, 1991; sutherland & taylor, 2011). therefore, researcher identity should not be considered a static product but a continuous process of identification, which can be described in terms of development (baker & lattuca, 2010) or an ‘identity-trajectory’ (mcalpine, amundsen, & turner, 2014) that accounts for both the continuity of stable personhood over time and a sense of ongoing change. this conceptualisation presents identity development as a route by which a newcomer becomes part of a community (golde, 1998; lave & wenger, 1991; sweitzer 2009). however, socialisation could also be considered a two-way process (mcdaniels, 2010) in which an individual actively explores possibilities for differentiation and negotiation with a community to find balance between institutional and structural positioning (archer, 2008), and to create space for personal autonomous actions in a changing environment (clegg, 2008). according to this broad sociocultural conceptualisation of researcher identity, it is important to account for the particular activities and interactions that characterise the different communities in which ecrs participate. ecrs interact with and engage in multiple communities, and these different communities shape the activities and positions that ecrs adopt. in the current complex higher education work conditions, crossing boundaries is one of the requirements of researchers and this includes personal, disciplinary, national, and professional positions related to research, teaching, administration, and leadership (boden, borrego, & newswander, 2011; holley, 2010; mcalpine & amundsen, 2009; sweitzer, 2009). castelló et al   | f l r         42   we propose the notion of spheres of activity as a helpful construct to characterise and explain the prototypical activities of different communities in which ecrs tend to be engaged. a particular sphere, as it works as a system, is shaped by rules, artefacts, and specific divisions of labour (engeström & sannino, 2010) and by the actions that individuals and communities develop to achieve outputs. actions, although performed by individuals, are also socially organised within communities, which accounts for recurrent actions shared by a group of individuals. at the same time, each of the spheres in which an individual participates can be shaped by different communities. therefore, notions of spheres of activity and communities are not synonymous. spheres can be considered domains or fields of participation in life or in human activity. communities are defined by the types of social actions that are developed by different groups of individuals within each sphere of activity. for instance, the learning sphere includes several communities (e.g., a community of peers participating in regular doctoral courses or seminars, or a community of phd students in a research team working with—and learning from—more senior researchers). in the case of ecrs, we distinguish at least three related spheres of activity that affect identity development (camps & castelló, 2013), as illustrated in figure 1. some representative activities of a particular sphere are emphasised or have more relevance at the beginning of the process of be(com)ing a researcher (e.g., completing set requirements for a doctoral program), whereas other activities (e.g., publishing or securing research funding) may be more common throughout or at advanced stages of researcher development. figure 1. spheres of activity for ecrs. the learning activity sphere is characterised by those more or less formal situations in which ecrs are situated as students in learning environments of different communities. these situations include seminars and doctoral courses, some aspects of supervisor relationships, and the increasing variety and number of development activities assessed for doctoral students and probationary or apprentice research staff (mcalpine, jazvac-martek, & hopwood, 2009). displaying these activities has to do with redefining the student identity developed in previous stages, since roles, outputs, and artefacts differ as students advance in their doctoral journeys toward the status of and possible employment as researchers. moreover, ecrs should also learn, usually implicitly, ways of acting, values, and practices that are prototypical of relevant disciplinary communities. institutional expectations of an increase in interdisciplinary work make this learning of disciplinary activity even more complex for ecrs faced with contradictory and fuzzy signals regarding appropriate actions and expected outputs. castelló et al   | f l r         43   the professional activity sphere is shaped by prototypical activities defining the professional communities to which ecrs belong or aim to belong when they finish their journeys. at times, these activities overlap with others from the learning sphere, which is common during doctoral study, particularly for those ecrs who aim to develop academic careers. in these cases, participating in scientific events, applying for grants and funding, presenting research, or teaching courses could simultaneously serve as learning and professional activities toward acquiring a university position. for ecrs advancing research careers in professional settings outside academia, the scenario is still more complex, since they must understand and participate in two—or more—distinct professional communities. doctoral students who are already employed in professional roles within or outside the academy may experience particular challenges deciding when and how to prioritise the learning activity or professional activity sphere. a third sphere of activity accounts for personal, family, and social activities that are variably related to learning and professional activities, especially in terms of values and aims, and the need to develop a researcher identity aligned with one’s personal intentions. there is a complex, dynamic interplay between ecrs and the spheres in which they are involved (pyhältö, nummenmaa, soini, stubb, & lonka, 2012; vekkaila, pyhältö, & lonka, 2013a) and hence, between the signals perceived across these spheres. one way to understand the dynamics contributing to researcher identity for ecrs is to explore them in terms of congruence (fit) (see edwards, 2007) or lack of congruence (misfit) between individuals and their environment (castelló, iñesta, & corcelles, 2013; pyhältö et al., 2012; vekkaila et al., 2013a). ideally, a constructive congruence is formed among the signals from the overlapping spheres. for instance, the personal or the professional sphere outside academia may function as a source of support for ecrs’ activities and goals in the learning sphere (such as earning a doctorate and aspiring toward a researcher career) (vekkaila, pyhältö, & lonka, 2013a, 2013b). on the other hand, activities in the personal and professional spheres may compete with the academic learning activities and often-distant goals (e.g., publishing articles and books, securing a permanent position at the university) by providing rival interests and prioritising short-term goals (vekkaila et al., 2013a, 2013b). moreover, depending on the context and domain, there are likely to be tensions or contradictions among the signals across these spheres. for instance, in much doctoral education, professional and learning spheres are highly intertwined. pursuing a doctorate entails conducting—and learning to conduct—research work and increasingly writing and publishing—and learning to write and publish—articles (castelló et al., 2013; pyhältö et al., 2012; vekkaila, pyhältö, hakkarainen, keskinen, & lonka, 2012). through numerous interactions across the spheres, individuals’ positioning, the environments, and the relations among them are constantly evolving and being re-negotiated (camps & castelló, 2013; lave & wenger, 1991). intersections across spheres are multiple and unavoidable, which may illuminate synergies or contradictions for ecrs who are striving to make sense of the signals and transitions they encounter as they formulate a researcher identity. discovering and sharing the changes in rules and recommended actions in each sphere could enhance ecrs’ awareness of new signals crucial for researcher identity development in the 21st century. it might also be useful to explain transitions between communities and what these transitions imply for the processes involved in researcher identity construction (castelló et al., 2013; giddens, 1991; goffman, 1967; strandler, johansson, wisker, & claesson, 2014; wisker & robinson, 2012). 3. identity stories earlier work conducted individually by team members on various projects concerned issues related to ecrs’ identity, engagement, sense of belonging, writing, metacognition, wellbeing, and resilience, among other topics. the risk-career, signals during researcher career development, researcher identity, and the need to manage contradictions and tensions emerged as main themes during our discussion of these previous castelló et al   | f l r         44   research projects. in order to illustrate the proposed comprehensive framework, we consulted our existing datasets to build indicative identity stories exemplifying ecrs’ experiences and trajectories in the light of their recognition of and response to signals affecting the development of researcher identity in the context of a risk-career. data were drawn from projects about doctoral education in finland (pyhältö et al., 2012; pyhältö, stubb, & lonka, 2009; vekkaila et al., 2012, 2013a) and in the united kingdom (wisker et al., 2010), and studies about being academics and researchers in canada (mcginn, 2012a, 2012b). this analysis is based upon interview transcripts from a total of 83 ecrs across three countries: finland (35), the united kingdom (30), and canada (18). all interviews were gathered according to the research ethics clearance procedures in the respective jurisdictions with care to protect the rights of the (potentially vulnerable) ecrs. to reduce risks of harm and ensure compliance with accepted procedures, individual team members worked directly with the original interview transcripts from their respective projects and did not share raw data with others. all the interviews included information about how the ecrs perceived themselves. the finnish participants were asked to discuss significant positive and negative turning points during their doctoral journeys and how they perceived themselves at these points (vekkaila et al., 2012, 2013a). the interviews conducted in the united kingdom aimed to draw out the participants’ experiences and identify transitions, turning points, and key learning moments within their doctoral journeys (wisker et al., 2010), whereas the participants in canada were asked explicitly to describe their perceptions of themselves as academic researchers (mcginn, 2012a) and more generally in academe (mcginn, 2012b). we selected interviews that presented (a) the strongest and clearest expressions of researcher identity, that is, participants’ perceptions of themselves as researchers; and (b) evidence of ecrs’ recognition of and response to signals from different spheres of activity while managing a risk-career. team members re-analysed their respective interviews using our jointly constructed framework of identity, spheres, and signals. based on these new analyses, we prepared drafts of identity stories from original interview transcripts, which were then reviewed by the whole research team. initially we started with a large number of identity stories; however, after several careful readings, we collectively selected a final set of eight identity stories on the basis of their potential capacity for illustrating the ways signals are emerging and interpreted by ecrs in the current higher education context of risk-careers. in our selections, we specifically sought diversity in terms of countries of origin and location, fields of research, and researcher career systems. although this analysis emphasises this limited set of just eight identity stories, the issues addressed were prevalent across our international datasets and not limited to these eight ecrs. 4. discussion of the identity stories identity stories from mari, jaakko, dan, wang, siiri, elaine, aatu, and kenneth (these are pseudonyms) provide a wide range of examples of development of researcher identity, spheres, and signals and represent empirical examples of researcher identity and tensions during early-stage identity construction. in each story, we highlight relevant spheres, the congruence (or lack of congruence) between signals coming from one sphere and another sphere (or among signals coming from multiple spheres), and the ways in which ecrs respond to (or, equally important, miss) these signals. we also specify the nature of different types of signals, ranging from those perceived as implicit to more explicit ones, as well as characteristics of responses in terms of agency and identity construction. for ease of reference, we provide an overview of the eight identity stories in table 1. castelló et al   | f l r         45   table 1 overview of identity stories pseudonym gender country english as first (l1) or second (l2) language interviewed in first (l1) or second (l2) language position at time of interview mari f finland l2 l1 doctoral student jaakko m finland l2 l1 doctoral student dan m uk (originally israel) l2 l2 lecturer wang m canada (originally china) l2 l2 assistant professor siiri f finland l2 l1 doctoral student elaine f canada l1 l1 assistant professor aatu m finland l2 l1 doctoral student kenneth m canada l1 l1 educational developer (doctoral studies on hold) 5. signals of congruence congruence between individuals’ perceptions of and interpretations of signals in the learning activity sphere and those in the professional activity sphere can reinforce ecrs’ identity as researchers. such congruence was evident for two finnish doctoral students from the natural sciences. participating in international conferences and networking across universities is an increasingly common requirement for establishing a successful researcher career, and therefore these professional academic activities are also activities and goals in the learning activity sphere. mari reported, “the most significant turning point in the final phase of doctoral studies was the meeting of the international researcher whose research had inspired me from the beginning of my studies…. we also talked that i could visit her group and conduct my post-doc project there.” her success in networking internationally and establishing connections with an international researcher were signals that strengthened her identity as a scientist and prompted her to make active decisions in terms of her further research and future in academia. a similar fit was evident in jaakko’s story. he also strengthened his identity as a researcher by reading signals from the international arena: “i presented my results, got encouragement from others, and i learnt what they did.” moreover, jaakko moved beyond the learning activity sphere, where dependence on a supervisor is common, into the professional activity sphere through co-authoring an article with other international researchers: “in this article the main responsibility of writing was shared between me and another, more senior scientist…. if my supervisor would have been involved in this i would not have such an independent role in writing the article and collaborating with others.” such signals enabled jaakko to develop an identity consistent with moving castelló et al   | f l r         46   toward the post-doctoral phase of his academic career: “and they are also indicators in my cv showing that i have the competence to work with others outside my own group.” figure 2. interpretation of signals from spheres of activity for mari and jaakko. dan’s identity story illustrates the importance of congruence between the perceived signals from the professional and the personal activity spheres. brought up in an orphanage in israel, dan had neither family ties nor money. he developed an internal sense of values and determination to succeed, becoming a physical education instructor for young men with difficult histories involving crime, poverty, and lack of education. his research was on physical education training. his achievement during a phd in the united kingdom was intellectual and personal; it affected his sense of himself as a professional success and a role model: “my wife fell in love with me more, she appreciates me more, especially my father in law. my students too—they appreciate me and they were my catalyst for my research. i dedicated this research to my students while other people usually dedicate their phd to their family…. the close society—family and colleagues— appreciate this, including my students who are proud of their lecturer who is perceived as a role model. a role model in the practical area and in the cognitive area—a doctor and a professional. usually, when i publish or when i participate in conferences, professional development courses or workshops then the title is meaningful.” the fit between the personal and the professional activity spheres for dan applied to his personal values and family as well as his research and teaching. figure 3. dan’s interpretation of signals from spheres of activity. castelló et al   | f l r         47   6. solving tensions and incongruences such congruence between the perceived signals coming from the personal and the professional activity spheres was also important for wang, but his identity story is more complex with tensions within the professional sphere between two competing communities with different values and rules. wang completed his doctoral degree, was hired at a canadian university, and had recently transferred to a second canadian university. both positions were as a tenure-track assistant professor in education. prior to his doctoral studies, he was employed as a professor in his home country, china, where he received awards as a researcher and was selected for an international scholarship. his first appointment in canada was at a research-intensive university where there were extensive pressures to secure research funding: “this is a bombard, this unspoken language.” he was happy to have transitioned to a less competitive environment where his national research grant and strong publication record allowed him to stand out rather than trail behind colleagues. perceived improvements in his general health and wellbeing indicated the importance for him of experiencing greater congruence between his personal and professional activity spheres. figure 4. wang’s interpretation of signals from spheres of activity. another identity story revealed perceived signals in the personal activity sphere that represented a misfit with one community of the professional activity sphere and a fit with another community of this same professional activity sphere. siiri, a finnish doctoral student in behavioural sciences, was involved in doctoral studies part-time while working full-time outside academia. initially, her doctoral studies and professional work outside the university were strongly interconnected and her employer encouraged her to conduct a thesis; however, this congruence diminished over time: “there was this organisational change and my time to conduct the thesis disappeared…. then, i was not able to write the articles, i have dozens of conference posters but those cannot be included in the doctoral thesis.… if the situation would have stayed the same i probably would have a stronger identity as a researcher.” within the professional activity sphere, a tension developed between her academic and her non-academic activities, and in time the signals from her non-academic professional life increasingly diminished the importance of earning the doctorate: “in this field the academic degrees are of course one way of gaining expertise but the other, as appreciated, way is through conducting the work.… in the beginning the expectations motivated me to pursue the doctorate but now when i am an acknowledged expert in my field without the phd it has decreased my motivation to pursue it.” siiri was dedicated to her professional career outside academia and had developed her identity as an expert by relying on signals coming from her non-academic professional life. therefore, she had gradually become distanced and alienated from her identity as an academic researcher. instead, she valued her identity as an expert, and in that sense, she experienced congruence between her personal values and her professional life outside academia. castelló et al   | f l r         48   figure 5. siiri’s interpretation of signals from spheres of activity. elaine’s story also illustrates a misfit between perceived signals coming from the personal activity sphere and those from the different communities of the professional activity sphere, in which she participated, particularly regarding personal values. for elaine, completing a phd and securing a position as a tenure-track professor in education in canada actually diminished her sense of self-esteem. prior to entering academe as a mature student, she had research experience through her work outside academia, including publishing and evaluating research funding bids. during those early experiences, her personal sense of identity as a researcher was reinforced in the ways that others treated and referenced her: “it wasn’t only my own identity but being recognised by others as being a researcher in the field … certainly the recognition by others and which i think had started off … by having my first study being published in [an academic journal].” these early, positive expressions led her to doctoral studies and an academic position, but she felt the signals of success in academia had shifted to particular kinds of dissemination (peer-reviewed journal articles) rather than the public policy work she had done. this shift undermined the pleasures she once associated with research and her confidence as a researcher. her identity story is a clear example of tensions within the professional activity sphere and particularly of the ways contradictions between communities within and outside academia can interfere with identity development. elaine’s expectation that her professional experience would contribute and enhance her academic activity had not been upheld, which undermined her researcher identity. figure 6. elaine’s interpretations of signals from spheres of activity. aatu’s story illustrates a misfit in the perceived signals coming from all three spheres. at the beginning of his doctoral studies, aatu, a finnish doctoral student in behavioural sciences, was eager and inspired to pursue his doctoral research: “i thought that now i will pursue my thesis, i planned my castelló et al   | f l r         49   publications.… no problems with that because i had such excellent data…. and the beginning was excellent, i wrote three conference papers and then i presented my results in conferences… and i thought that i had reached a new level as a researcher.” however, the signals aatu interpreted from the journal peer-review process were different: “this same paper that included the same information and structure as the paper that got accepted in the previous conference and got positive review comments got now crushing review comments from the journal.… i thought that there was no logic in this system.” aatu’s story focused on new expectations in the academic professional activity sphere and the learning activity sphere: increasingly doctoral students are required to publish during their doctoral studies, but they are still in the process of learning what is involved in writing and publishing papers. the learning experiences involved in the publication process were not in congruence with aatu’s initial expectations, which resulted in misfits and tensions between his personal intentions and the signals he received from the professional activity sphere, causing him to struggle with his identity as a researcher: “i wondered if i could get any permanent position from the university with my publication list.… maybe i was not capable to play this ‘science game.’ i think that this is not worth it at all…. do i even want to play this game anymore?” the sense he made of the signals from the learning activity sphere and the professional activity sphere made him increasingly alienated from research and cynical towards science. aatu identified some signals and requirements defining a research career but they were not things he considered meaningful and worth striving toward; they were inconsistent with priorities in his personal activity sphere. figure 7. aatu’s interpretation of signals from spheres of activity. the final identity story illustrates both fits and misfits among the perceived signals from the professional, learning, and personal activity spheres. kenneth had placed his doctoral studies in the humanities on hold temporarily while employed in a full-time, non-tenure-track position at a canadian university. within his doctoral program, he had perceived himself on the margins both with regards to his theoretical interests and his evolving interests in pedagogy: “i was loving my teaching, loving it. so that in one of my comprehensive exams i added … some things on pedagogy… and at that time maybe it should have clicked that i was, you know, interested in maybe writing and exploring that further.” he had been feeling like a “loser” because his phd was unfinished. this feeling of powerlessness lifted, however, when he accepted a full-time teaching development position. he felt he could effect change in the teaching profession, and this in turn left him feeling positive about himself and more confident he would eventually finish his degree: “i feel like i am beginning to be able to effect some of that change and it’s a cool feeling and so i know now that i’ll finish my phd.” he felt a strong sense of belonging within the higher education castelló et al   | f l r         50   pedagogical community and even began to see teaching and teaching development as a prospective career choice. he felt respected and included by other academics in his teaching role, and saw this as an affirming space for himself. he did not feel similarly encouraged by others to do research, which undermined his researcher identity. figure 8. kenneth’s interpretation of signals from spheres of activity. 7. conclusion higher education now offers increasingly precarious career prospects for ecrs. in this contribution, we first offered a framework to account for the notion of researcher identity, which provides a new comprehensive way to analyse researcher identity development within the complex and intertwined spheres of activity in which ecrs are involved. second, in combining across and re-scrutinising data from a range of previous research projects, we explored the usefulness and potential of this framework by means of illustrating ways in which ecrs were aware of, and responded to, signals about their career trajectories, which, in turn, were connected to researcher identity (castelló et al., 2013; pyhältö et al., 2012; vekkaila et al., 2013a). recognition of and response to these differing signals is an important aspect of an ecr’s identity-trajectory (mcalpine et al., 2014) within the context of a risk-career. in the current higher education landscape, ecrs are faced with ever-increasing and possibly conflicting demands to advance toward research careers they fear may not materialise. rather than anticipating stable research careers in academic institutions, ecrs are now pressured to consider how they might contribute and find satisfaction through alternative academic or perhaps non-academic careers. consciously attending to the signals present across spheres of activity may provide ecrs with a sense of agency within the uncertainty of a risk-career environment. we were heartened by the extent to which this new framework applied across the diverse researcher education and researcher career systems in our various international contexts and the different disciplinary fields and professional settings for ecrs involved in our interviews, but we also acknowledge that this conceptualisation requires further testing and analysis with new data to assess its wider transferability along the various career trajectories ecrs face globally. moreover, since data we used came from our previous castelló et al   | f l r         51   work, the situations and signals we have been able to identify might not be fully representative of the emerging tensions and pressures that ecrs are facing within the current context of a risk-career. assessing and refining the provided comprehensive framework, identifying signals emerging in and across different spheres of activity, and helping ecrs to identify and respond to these signals are important issues that deserve recognition and focused attention in the efforts of the researcher education and careers sig to advance a shared research agenda for exploring ecrs’ identity development in the 21st century. we propose the following future research emphases as ones that have the greatest potential for consolidating the comprehensive conceptual framework introduced here and facilitating ecrs’ identity harmonic development: • original data specific to ecrs’ spheres of activity, perceived signals, and tensions associated with a risk-career are needed in order to discuss and further illuminate the complex considerations discussed in this paper and advance knowledge about the nature and development of researcher identity. • cross-cultural analyses of ecrs’ experiences across contexts could lead to better understandings of the ways ecrs identify and respond to relevant signals in their various contexts, and, consequently, could unravel the complex interplay between signals and spheres of activity when dealing with tensions and struggling to become researchers in a risk-career environment. in the current globalised context, such cross-cultural analyses could extend to include situations were ecrs pursue opportunities in other cultures and countries. • in the changing scenarios facing higher education systems worldwide, the study of the ways ecrs deal with the perceived continuities, and especially discontinuities, among spheres of activity could help to identify and theorise the conflicting signals that systems are producing and to provide ecrs with tools to better interpret and respond to these signals. • longitudinal analyses of changes ecrs face as they progress from admission to graduation and into initial appointments within and beyond academe will be particularly useful to understand transitions, trajectories, and the varying signals between and among spheres of activity. • more generally, we encourage researchers who focus on the ways specific activities (e.g., writing, supervisory interactions, teaching or publishing, among others) contribute to ecrs’ identity development should attempt to situate their conceptual and methodological assumptions in relation to a comprehensive framework of identity development, such as the one provided in this article, in order to make diverse research data integration possible. keypoints in the current higher education context, early career researchers (ecrs) face a ‘risk-career,’ in which they must identify and interpret new or emergent ‘signals’ in their efforts to develop a researcher identity. the proposed comprehensive framework for researcher identity emphasises ecrs’ recognition and response to signals across spheres of activity. identity stories drawn from prior studies illustrate the congruence (or lack of congruence) between and among signals across different spheres of activity, and the varied ways ecrs respond to (or miss) these signals. the framework and identity stories are intended to offer exemplars to assist ecrs, supervisors, and university managers to identify issues and manage risk-careers. castelló et al   | f l r         52   acknowledgements data for this analysis and some of theoretical underpinnings were drawn from earlier projects funded by the finnish cultural foundation, the academy of finland, the higher education academy of uk, the higher education funding council for england, the social sciences and humanities research council of canada, the human resources and skills development canada job creation program, the spanish ministerio de economía y competitividad (dgicyt (cso2013-41108-r) and our respective institutions. this paper arose from a working group established for the inaugural meeting of the earli special interest group researcher education and careers in barcelona, spain, october 2014. montserrat castelló served as coordinator for the group and primary author for this paper. all other authors’ names are listed in alphabetical order. references archer, l. (2008). younger academics’ constructions of ‘authenticity’, ‘success’ and professional identity. studies in higher education, 33, 385–403. doi:10.1080/03075070802211729 baker, v. l., & lattuca, l. r. (2010). developmental networks and learning: toward an interdisciplinary perspective on identity development during doctoral study. studies in higher education, 35, 807–827. doi:10.1080/03075070903501887 boden, d., borrego, m., & newswander, l. k. (2011). student socialization in interdisciplinary doctoral education. higher education, 62, 741–755. doi:10.1007/s10734-011-9415-1 camps, a., & castelló, m. (2013). la escritura académica en la universidad [academic writing at university]. redu. revista de docencia universitaria, 11(1), 17–36. retrieved from http://www.redu.net cantwell, b. (2011). academic in-sourcing: international postdoctoral employment and new modes of academic production. journal of higher education policy and management, 33, 101–114. doi:10.1080/1360080x.2011.550032 castelló, m., & iñesta, a. (2012). texts as artifacts-in-activity: developing authorial identity and academic voice in writing academic research papers. in m. castelló & c. donahue (eds.), university writing: selves and texts in academic societies (pp. 179–200). bingley, uk: emerald. castelló, m., iñesta, a., & corcelles, m. (2013). learning to write a research article: ph.d. students’ transitions toward disciplinary writing regulation. research in the teaching of english, 47, 442–477. clegg, s., (2008). academic identities under threat? british educational research journal, 34, 329–345. doi:10.1080/01411920701532269 cyranoski, d., gilbert, n., ledford, h., nayar, a., & yahia, m. (2011). the phd factory. nature, 472, 276– 279. doi:10.1038/472276a danaher, p. (2015). forms of capital and transition pedagogies: researching to learn among postgraduate students and early career academics at an australian university. to appear in c. guerin, c. nygaard, & p. bartholomew (eds.), learning to research—researching to learn. oxfordshire, uk: libri. edwards, j. r. (2007). the relationship between person–environment fit and outcomes: an integrative theoretical framework. in c. ostroff & t. a. judge (eds.), perspectives on organizational fit (pp. 209– 258). san francisco, ca: jossey-bass. engeström, y., & sannino, a. (2010). studies of expansive learning: foundations, findings and future challenges. educational research review, 5, 1–24. doi:10.1016/j.edurev.2009.12.002 giddens, a. (1991). modernity and self-identity: self and society in the late modern age. cambridge, uk: polity. castelló et al   | f l r         53   goffman, e. (1967). interaction ritual: essays on the face-to-face behaviour. london, uk: allen lane. golde, c. m. (1998). beginning graduate school: explaining first‐year doctoral attrition. new directions for higher education, 1998(101), 55–64. doi:10.1002/he.10105 harré, r., moghaddam, f. m., cairnie, t. p., rothbart, d., & sabat, s. r. (2009). recent advances in positioning theory. theory & psychology, 19, 5–31. doi:10.1177/0959354308101417 hemmings, b., hill, d., & sharp, j. g. (2013). critical interactions shaping early academic career development in two higher education institutions. issues in educational research, 23, 35–51. retrieved from http://www.iier.org.au/ holland, d. c., lachicotte, w., jr., skinner, d., & cain, c. (2001). identity and agency in cultural worlds. cambridge, ma: harvard university press. holley, k. (2010). doctoral student socialization in interdisciplinary fields. in s. k. gardner & p. mendoza (eds.), on becoming a scholar: socialization and development in doctoral education (pp. 97–112). sterling, va: stylus. lave, j., & wenger, e. (1991). situated learning: legitimate peripheral participation. new york, ny: cambridge university press. mcalpine, l., & amundsen, c. (2009). identity and agency: pleasures and collegiality among the challenges of the doctoral journey. studies in continuing education, 31, 109–125. doi:10.1080/01580370902927378 mcalpine, l., amundsen, c., & turner, g. (2014). identity-trajectory: reframing early career academic experience. british educational research journal, 40, 952–969. doi:10.1002/berj.3123 mcalpine, l., jazvac-martek, m., & hopwood, n. (2009). doctoral student experience in education: activities and difficulties influencing identity development. international journal for researcher development, 1, 97–109. doi:10.1108/1759751x201100007 mcdaniels, m. (2010). doctoral student socialization for teaching roles. in s. k. gardner & p. mendoza (eds.), on becoming a scholar: socialization and development in doctoral education (pp. 29–44). sterling, va: stylus. mcginn, m. k. (2012a). being academic researchers: navigating pleasures and pains in the current canadian context. workplace: a journal for academic labor, 21, 14–24. retrieved from http://ojs.library.ubc.ca/index.php/workplace/ mcginn, m. k. (guest ed.). (2012b). belonging and non-belonging: costs and consequences in academic lives [special issue]. workplace: a journal for academic labor, 19. retrieved from http://ojs.library.ubc.ca/index.php/workplace/ porfilio, b. j., gorlewski, j. a., & pineo-jensen, s. (guest eds.). (2013). the new academic labor market and graduate students [special issue]. workplace: a journal for academic labor, 22. retrieved from http://ojs.library.ubc.ca/index.php/workplace/ pyhältö, k., & keskinen, j. (2012). doctoral students’ sense of relational agency in their scholarly communities. international journal of higher education, 1(2), 136–149. doi:10.5430/ijhe.v1n2p136 pyhältö, k., nummenmaa a. r., soini, t., stubb, j., & lonka, k. (2012). research on scholarly communities and development of scholarly identity in finnish doctoral education. in t. s. ahola & d. m. hoffman (eds.), higher education research in finland: emerging structures and contemporary issues (pp. 337–357). jyväskylä, finland: jyväskylä university press. pyhältö, k., stubb, j., & lonka, k. (2009). developing scholarly communities as learning environments for doctoral students. international journal for academic development, 14, 221–232. doi:10.1080/13601440903106551 strandler, o., johansson, t., wisker, g., & claesson, s. (2014). supervisor or counsellor?—emotional boundary work in supervision. international journal for researcher development, 5, 70–82. doi:10.1108/ijrd-03-2014-0002 sutherland, k., & taylor, l. (2011). the development of identity, agency and community in the early stages of the academic career. international journal for academic development, 16, 183–186. doi:10.1080/1360144x.2011.596698 castelló et al   | f l r         54   sweitzer, v. (2009). towards a theory of doctoral student professional identity development: a developmental networks approach. the journal of higher education, 80, 1–33. doi:10.1353/jhe.0.0034 trede, f., macklin, r. & bridges, d. (2012). professional identity development: a review of the higher education literature. studies in higher education, 37, 365–384. doi:10.1080/03075079.2010.521237 vekkaila, j., pyhältö, k., hakkarainen, k., keskinen, j., & lonka, k. (2012). doctoral students’ key learning experiences in the natural sciences. international journal for researcher development, 3, 154–183. doi:10.1108/17597511311316991 vekkaila, j., pyhältö, k., & lonka, k. (2013a). experiences of disengagement—a study of doctoral students in the behavioral sciences. international journal of doctoral studies, 8, 61–81. retrieved from http://www.informingscience.org/journals/ijds/ vekkaila, j., pyhältö, k., & lonka, k. (2013b). focusing on doctoral students’ experiences of engagement in thesis work. frontline learning research, 1(2), 10–32. doi:10.14786/flr.v1i2.43 weber, m. (1947). science as a profession. in h. h. gerth & c. w. mills (eds.), from max weber: essays in sociology (pp. 129–156). london, uk: kegan. winter, r. (2009). academic manager or managed academic? academic identity schisms in higher education. journal of higher education policy and management, 31, 121–131. doi:10.1080/13600800902825835 wisker, g., morris, c., cheng, m., masika, r., warnes, m., trafford, v., robinson, g., & lilly, j. (2010). doctoral learning journeys: final report. retrieved from http://about.brighton.ac.uk/clt/research/cltresearch/doctoral-learning-journeys/ wisker, g., & robinson, g. (2012). picking up the pieces: supervisors and doctoral “orphans.” international journal for researcher development, 3, 139–153. doi:10.1108/17597511311316982 yorke, m. (2009). faulty signals? inadequacies of grading systems and a possible response. in g. joughin (ed.), assessment, learning and judgement in higher education (pp. 65–84). new york, ny: springer science+business media. microsoft word boeren et al_publication.docx           frontline learning research vol.3 no. 3 special issue (2015) 68 80 issn 2295-3159   corresponding author: dr ellen boeren, university of edinburgh, moray house school of education, simon laurie house, holyrood road , edinburgh eh8 8aq. phone: 0044-(0)131-651 6233, email: ellen.boeren@ed.ac.uk doi: http://dx.doi.org/10.14786/flr.v3i3.186   mentoring: a review of early career researcher studies ellen boerena, irina lokhtina-antonioub, yusuke sakuraic, chaya hermand. lynn mcalpinee a university of edinburgh, uk b university of leicester, uk c university of tokyo, japan c university of pretoria, south africa e university of oxford, uk article received 14 june 2015 / revised 14 july 2015 / accepted 16 july 2015 / available online 12 october 2015 abstract this paper reviews 23 journal articles on ‘mentoring’ in the context of early career researchers, defined as those in academia with less than 10 years of experience from the start of their phd. achieving a better understanding of mentoring is important since within the higher education context new dynamics have created expectations towards more supportive mechamisms for ecrs. in order to better understand the benefits of mentoring for ecrs careers and psychosocial well-being, it is important to understand (1) the core definitions of mentoring used in research, (2) the research methodologies that are applied to research mentoring, (3) the empirical evidence showing the value of mentoring and (4) the remaining gaps for which future research will be needed. results of the review lead to the following conclusions: there is much research to do, first, to better inform our conceptualization of ecr mentoring and, second, to better understand the value of ecr mentoring support. a research agenda is outlined. keywords: mentoring; early career researchers; review paper       boeren et al 69 | f l r     1. introduction this paper presents the results of a review of mentoring papers that appeared in leading higher education journals in the past ten years. over the past years, new dynamics have emerged in the context of higher education globally that have created both expectations and aspirations towards supportive mechanisms of early career researchers’ (ecrs) professional development. in this paper, ecrs are defined as researchers in academia with less than 10 years of experience from the start of their phd studies, congruent with the definition used by the european commission. why is mentoring an important topic in relation to ecrs? internationally, ecrs in academia are challenged as regards access to resources, supportive interactions and lack of transparent career prospectives (the european commision 2011). related to this is the underlying pressure experienced by ecrs in terms of their opportunities for research and development (sauermann & roach, 2012; åkerlind, 2005; vitae, 2011) and international mobility (jepsen et al. 2014; mellors-bourne et al., 2013; kehm, 2007) required to enhance their career prospects and secure stable positions. moreover, academic workplaces have been transformed; that in turn, has lengthened the learning trajectories of ecrs (bonetta, 2011) and made them in some respects more complex (shuster, 2009). the above reports have pointed out the learning challenges ecrs perceive in developing their intellectual independence and scholarly profiles (gardner, 2008 for doctoral students; laudel & glaser, 2008 for postdocs). these reports also make relatively frequent mention of the value of mentors and mentoring (mullen & forbes, 2000; hemmings, 2012) as do reports of institutional practices to support ecrs (debowski, 2012). in this context, mentoring broadly can be situated in an array of complex supportive mechanisms including co-working and networking that lead to ecrs’ personal development, adaptation and integration as members of their scholarly community (e.g., baker et al., 2014). on the one hand, this includes informal mentoring through interactions between academics at different career stages. on the other hand, this includes formal mentoring programmes organised and structured at the institutional level. the role of mentoring in relation to ecrs corresponds to the more general literature on mentoring which focuses on ‘career’ and ‘psychosocial’ functions as the two major functions of support between mentors and mentees, contributing to (1) increasing the chances for promotion and higher salaries, building a network of professional collaborators (career function) as well as (2) achieving higher levels of confidence and social skills (psychosocial function) (ragins & kram, 2007). not only the specific function of mentoring, but also the organisation in which mentoring takes place might also have a significant effect on how mentoring is carried out, and which outcomes of mentoring are experienced. it is the specific higher education context we are interested in, and how mentoring gets discussed in the higher education literature as regards ecrs. four specific aims were formulated for the review. first of all, we wondered the extent to which ecr mentoring was conceptually constituted since a review on mentoring spanning 30 years of formal mentoring programs in the fields of education, business and medicine (ehrich et al., 2004) noted the absence of conceptual frameworks. so, we undertook to explore the conceptual tools and definitions used in the post-ehrich literature on ecr mentoring, starting from 2005. secondly, not only were we interested in the definitions and conceptual frameworks used by scholars, but we also wished to document the methodological tools they used to measure the impact of their definitions of mentoring. thirdly, we also analysed the extent to which empirical evidence would provide insight into how best to support ecrs’ development, i.e., what to avoid, since ehrich et al. (2004) had reported some negative consequences related to, for instance, the lack of training of mentors or mismatch of expertise or personality. there was also some evidence, again from non-higher education contexts (eby, 2008), that the effects of mentoring could be quite small. this led us to explore the nature of the evidence that would suggest mentoring could be a solution for the existing problems with career development and retention of ecrs, in       boeren et al 70 | f l r     particular whether mentoring could be used as a tool that contributes to ecrs’ professional development, including their competence (linden et al., 2013) and professional confidence in a range of key academic practices. finally, this review analysis aimed to identify gaps in the current literature and to explore the recommendations scholars have made for future research. in other words, our goal was to provide a research agenda for further inquiry into mentoring in relation to ecrs. 2. research question our overall question was ‘what does the ecr literature-research say about mentoring?’ specific research questions, summarizing the aims in the previous four paragraphs were: • what is the range of ways in which ecr mentoring is defined or conceptually presented? • what methodological tools are used, i.e., the range of ways in which mentoring is measured? • what evidence (or counter evidence) is there of the value of ecr mentoring? • to what extent does the literature point to future research? the answers to these questions provided a means to assess the extent to which mentoring was a robust workable construct in examining ecr experience. 3. method scope of review: as we undertook the study, we noted two related fields of study on mentoring, mentioned in the introduction: • the informal field of learning and acquiring research skills from interactions with more experienced researchers usually working in the same context • the structured, institutionalized programs of mentoring which are designed to support the needs of special groups like women, newly hired staff, or minority groups if they were directed at ecr. search process: in terms of the scope of the review, we decided to include journal articles published in the isi top ranked higher education journals in the past 10 years only (2005-2014) as these are supposed to be the most influential ones in the field. papers had to be published post-ehrich review, that is 2004. we reviewed papers that appeared in higher education, journal of higher education, research in higher education, review of higher education, and studies in higher education. these journals were likely the ones that he researchers, developers and policy makers would go to in seeking information about ecr mentoring. we also included the international journal for academic development and international journal for researcher development since these two journals are highly referenced in the field of academic development and thus need to be taken into account in an academic development-related review exercise. while the review has been limited to these journals, we feel confident in having made a sound selection of the major journals in the field. the keywords ‘mentor(ing)’ combined with ‘early career researchers’, ‘post-docs’, ‘doctoral students’, and variants had to appear in the title and/or the abstract of the article. we distributed the search task amongst the group of authors, with the search producing 23 papers. the distribution of papers according to journals can be found in table 1, full references of the 23 papers are included in the reference list at the end of the paper.       boeren et al 71 | f l r     analysis: the analysis framework drew upon boote and beile’s (2005) literature review scoring rubric, which they developed based on hart’s (1999) previous work. the five review categories boote and beile constructed are (1) coverage (reasons for inclusion or exclusion), (2) synthesis (state of the field, ambiguities, new perspectives), (3) methodology (methodologies and research techniques), (4) significance (practical and scholarly significance) and (5) rhetoric (level of coherence and structure). boote and beile’s work was specifically undertaken to increase scholar’s awareness of the literature review stage of a research project and is well-cited. in order to synthesize the selected literature, we created an excel file with the following expanded sub-categories of all but the last category in boote and beile’s (2005) rubric: (1) nature of the article: empirical or theoretical, (2) gap identified by authors, (3) question or purpose of the article, (4) conceptual framework for study, (5) pedagogical intervention (if there was one), (6) data collection method, (7) sample and nature of participants (8) country, disciplines (9) key empirical findings, if any (10) conceptual representation of results (11) practical and pedagogical implications (given our interest in ecr development), (12) suggestions for future research, (13) core references used by authors, and (14) reviewers notes/critique. in general, it can be argued that the sequence from exploring the literature, identifying a gap, spelling out research questions, explaining methodology, explaining and discussing results, and drawing conclusions with recommendations for future policy and practice, is perceived as a standard structure following which social sciences journal articles are written (see shon, 2012). this structure is also reflected in the sequence of our four research questions, focussing on (1) conceptual frameworks and definitions, (2) methodological approaches, (3) empirical evidence and (4) recommendations for future research. the articles were distributed across the group of authors. we each read and then summarized the papers; each author separately wrote a description of the emerging findings and his/her interpretation of them. these were used by two of the authors to create a first draft of the findings and conclusions. the draft was then reviewed and edited by the other authors. table 1: papers included in review (journals listed in alphabetic order) journal year authors title country article keywords higher education 2014 lechuga a motivation perspective on faculty mentoring: the notion of ‘‘non-intrusive’’ mentoring practices in science and engineering us faculty mentoring motivation discipline 2014 van der weijden, belder, van arensbergen & van den besselaar how do young tenured professors benefit from a mentor? effects on management, motivation and performance. the netherlands mentorship academic careers research management human resources motivation performance 2011 bell & treleaven looking for professor right: mentee selection of mentors in a formal mentoring program. australia academic development flexible mentoring mentor-mentee choice pairing process 2011 lechuga faculty-graduate student mentoring relationships: mentors’ perceived roles and responsibilities. us faculty graduate students mentoring higher education 2011 scaffidi & a positive postdoctoral australia postdocs       boeren et al 72 | f l r     berman experience is related to quality supervision and career mentoring, collaborations, networking and a nurturing research environment. mentoring collaborations networking research environment international journal for academic development 2012 saito when a practitioner becomes a university faculty member: a review of literature on the challenges faced by novice ex-practitioner teacher educators. professional development faculty member ex-practitioner teacher educator 2012 weaver, robbie, kokonis & miceli collaborative scholarship as a means of improving both university teaching practice and research capability. australia academic development mentoring scholarship of teaching and learning 2011 cox the impact of communities of practice in support of early-career academics. us early-career academics academic development program transformative learning community of practice faculty learning community 2011 remmik, karm haamer & lepp early-career academics’ learning in academic communities. estonia early career academics professional learning professional identity community of practice 2010 hubball, clarke & poole ten‐year reflections on mentoring sotl research in a research‐intensive university. canada mentoring scholarship of teaching and learning (sotl) sotl research outcomes 2009 foote & solem toward better mentoring for early career faculty: results of a study of us geographers. us early career faculty mentoring doctoral education 2008 kamvounias, mcgrath‐ champ & yip ‘gifts’ in mentoring: mentees' reflections on an academic development program. australia mentee mentor mentoring gift 2005 mathias mentoring on a programme for new university teachers: a partnership in revitalizing and empowering collegiality. uk international journal for researcher development 2014 baker, pifer& griffin mentor-protégé fit. us mentoring mentor-protégé fit doctoral education student–faculty mentoring relationships academic identity 2014 browning, thompson & developing future research leaders. australia early career researchers researcher development       boeren et al 73 | f l r     dawson evaluation research leaders track record studies in higher education 2013 gilmore, maher, feldon & timmerman exploration of factors related to the development of science, technology, engineering, and mathematics graduate teaching assistants' teaching orientations. us graduate teaching assistant teaching orientation teacher beliefs graduate student education graduate student development graduate student mentoring 2011 lindén, ohlin & brodin mentorship, supervision and learning experience in phd education. sweden mentorship phd students phd supervision learning outcomes professional development 2010 hopwood doctoral experience and learning from a sociocultural perspective. uk doctoral education doctoral practices academic practice sociocultural perspectives doctoral study 2008 kamler rethinking doctoral publication practices: writing from and beyond the thesis. australia the journal of higher education 2012 noy & ray graduate students' perceptions of their advisors: is there systematic disadvantage in mentorship? us 2009 patton my sister’s keeper: a qualitative examination of mentoring experiencesamong african american women in graduate and professional schools us the review of higher education 2014 main gender homophily, ph.d. completion, and time to degree in the humanities and humanistic social sciences. us 2013 o’meara, knudsen & jones the role of emotional competencies in facultydoctoral student relationships. us       boeren et al 74 | f l r     4. results as stated above, the main aim of this paper is to generate insight into the current academic literature on ecr mentoring; the findings are structured around the four research questions. 4.1 what is the range of ways in which ecr mentoring is defined or conceptually presented? in order to answer this question, we first explored the nature of the articles, the gaps identified by the authors and the specific research questions in these papers, as these elements could be expected to be related to the conceptual frameworks and definitions authors had drawn upon in developing their research study. nature of articles: our initial search of the journals confirmed ehrich’s (2004) outcome that while mentoring was frequently referred to, it was rarely studied. in fact, we found more articles that referred to mentoring than those which studied mentoring as the core business of their research project. of the 23 articles that studied mentoring and formed the basis of the review, 19 were empirical studies, four (mathias, 2005; kamvounias et al, 2008; hubball et al., 2010; bell & treveanor, 2011) of which evaluated programs that had a mentoring element. four articles were non-empirical in nature (baker et al., 2014; cox, 2013; saito, 2013; weave et al., 2013). gap identified: in examining the ‘gap’ that the authors were attempting to address, we noted that a definitional representation of mentoring was rare. for instance, most of the studies addressing doctoral experience used as a starting point that the supervisor (referred to as advisor in the us) was equivalent to a mentor, though baker et al. (2014) noted the ambiguity in the roles of supervisor, advisor, and mentor. so, while half of the studies explicitly named the ‘gap’ as the need to understand mentoring better, most appeared to assume a shared understanding of mentoring between authors and readers with the focus of each study mainly directed to a given situation in a specific context, e.g., support for teaching assistants. question/purpose: the research questions underlying the studies were formulated to answer question about (1) experiences of mentoring and (2) mentoring relationships. the bulk of the studies (kamler, 2008; patton, 2009; hopwood, 2009; lechuga, 2011; baker et al., 2012; noy & ray, 2012; gilmore et al., 2013; linden et al., 2013; main, 2014) addressed mentoring in the context of doctoral education, answering a wide range of research questions in relation to career advice, teaching and supervisory relationships. this group was followed by a substantial minority on mentoring related to teaching development with reference to ecrs, though not necessarily defining who they were in terms of their length of research or academic experience. lastly, only one addressed postdoctoral experience (scaffidi & berman, 2011). in general, across all papers reviewed, two main purposes were thus found. (1) many articles focused on gaining better insight into the way ecrs experience mentoring and whether they get something out of it in terms of their own learning process and professional development. examples include patton’s article (2009) on experiences of african american women in academia, kamler’s research (2008) on mentoring experiences in relation to academic writing, or mathias’ paper (2005) on specific mentoring experiences in relation to participation in the postgraduate certificate of academic practice (the uk’s officially recognised higher education teaching qualification). (2) another cluster of papers focused on the specific relationships that are being built between mentors and mentees. for instance, lechuga (2011) focused on mentors’ responsibilities and the relationships they built with faculty-graduate students. kamvanious et al’s research (2008) explored the idea of ‘gifts’ in mentoring, and how mentees want to give something back to their mentors. research by o’meara et al. (2013) explored the ‘emotional landscape’ of relationships between mentors/advisors and doctoral students. conceptual frame and core references: having identified the nature, gaps and purposes of these articles, the next step was to explore the conceptual frameworks used by these authors. in general, conceptual framing of the studies in relation specifically to mentoring was minimal. rather, papers tended to draw on general theories of learning and faculty development largely rooted in socio-constructivist       boeren et al 75 | f l r     perspectives, e.g., communities of practice (cox, 2013), learning (linden et al., 2013), emotional competence (o’meara et al., 2013), scholarship of teaching (gilmore et al., 2013; weave et al., 2013) or were firmly empirical (e.g. mathias et al. 2005; browning et al., 2014). two empirical studies stood out for their efforts to frame mentoring: van der weijden et al. (2004) and linden et al. (2013). linden et al. (2013) used a typology of learning outcomes related to mentoring in the business context (lankau and scandura 2007). van der weijden et al. (2014) drew on the meta-analysis of mentoring programs in a range of fields referred to earlier (ehrich et al, 2004). as well, baker et al. (2014) in their conceptual paper proposed a model based on the notion of professional, relational and personal fit, rather than similarity, between student and supervisor. given the diversity of stances taken in these studies, it was hard to discover a consistent pattern of common core references to conceptions of mentoring. to conclude, as a general answer to this research question, the most striking finding of this analysis was a confirmation of the findings in the earlier non-higher education review (ehrich et al., 2004): the generally under-conceptualized nature of mentoring in empirical studies on ecr experience. we would encourage researchers undertaking future studies of ecr mentoring to explicitly explore the value of different conceptual frameworks of mentoring, perhaps beginning with erich et al.’s meta-analysis. 4.2 what methodological tools are used, i.e., the range of ways in which mentoring is measured? in order to answer this research question, we explored the specific settings in which data were collected, by this we mean, the nature of the participants, their disciplines, the national location of the study, as well as the ways in which data were collected and analyzed, distinguishing principally between quantitative and qualitative methodologies. country/disciplines/participants: a majority of the papers represented research in english-speaking countries, with more in north america and australia than in the uk. there were ten north american (nine us gilmore et al., 2013; foote & solem, 2009; main, 2014; noy & ray, 2012; baker et al., 2014; o'meara et al., 2013; cox, 2011; lechuga, 2014; patton, 2009, one canada – hubball et al., 2010); eight australia and five eu (three in continental europe: sweden – linden et al., 2013; the netherlands – van der weijden et al. 2014; estonia – remmik, 2011); two in the uk (mathias, 2005; hopwood, 2010). as to disciplinary context, four focused on stem fields (scaffaldi & berman, 2011; gilmore et al., 2013; van der weijden et al., 2014; lechuga, 2014), one focused specifically on humanist disciplines (main, 2014) and the remainder represented participants from a range of disciplines, although most within social sciences. data collection/participants: as to the methods, qualitative and mixed methods were used more than solely quantitative studies. all, but one, of the quantitative studies were based on surveys, while main (2014) conducted an analysis of pre-existing large data sets. as for the qualitative studies, the four papers evaluating programs used semi-structured interviews, student work, program documents and sometimes focus groups (hubball et al., 2010; bell & treleaven 2011; kamvounias et al., 2008; mathias, 2005). the other qualitative studies were based on semi-structured interviews, with one also using focus groups, and another interviews over time (kamler, 2008). of the two mixed methods studies (foote & solem, 2009; gilmore et al., 2013), one used interviews which were analyzed both quantitatively and qualitatively; the other used interviews, followed by a survey. participant numbers in the qualitative studies tended to be quite small though foote & solem (2009) used focus groups with 46 ecrs and hopwood (2010) 33 in focus groups and interviews. the quantitative studies also varied considerably in size from 86 (van der weijden et al., 2014) to several thousands (main, 2014). as a general answer to this research question, we concluded that while a mix of qualitative and quantitative approaches were used, most studies (regardless of the methods) were based on one-time data collection with small numbers of participants. one study (kamler, 2008) stood out in studying participant experience longitudinally which we view an innovative approach which might be emulated in future studies. further, the tools used in the studies rarely were designed to capture experience related to the specific mentoring activities under study. we suggest future studies could develop tools to better capture the       boeren et al 76 | f l r     experience of specific elements of mentoring. lastly, the majority of studies were based on self-report; future research might move beyond this way of collecting data. 4.3 what evidence (or counter evidence) is there of the value of ecr mentoring? evidence of the value of mentoring was searched for in the results and conclusion sections of the papers under review. apart from the nature of the results, we also explored the way in which they were formulated, in order to search for a conceptual representation of the results, which could form a strong conceptual basis for future research. findings: key findings (if any), conceptual representation of results while most papers reported positive experiences and relationships in mentoring, it is important to recognise the influence of a range of factors on mentoring: e.g. mathias (2005) concluded that mentoring provided within a postgraduate course resulted in several positive experiences, though much of the effect depended on the successful match between mentor and mentee. still, given the research was undertaken in different contexts and within different disciplines, it is difficult to draw an overall conclusion which indicates either a positive or negative effect of mentoring. for example, lechuga (2014) proposes that some mentoring relations that are acceptable in the social sciences may be considered “intrusive” in science and engineering. still, at first sight, the cumulative results of these studies would appear to confirm the earlier non-higher education literature reviews: eby et al. (2008) argued that the effects of mentoring seem quite small, and that mentoring does not always lead to positive experiences; ehrich et al. (2004) that mentoring can, in fact, have negative consequences. some authors, like patton (2009) have already reflected on the lack of robust conceptual frameworks emerging from studies of mentoring, such as the emphasis on the paternal, male representation of mentoring and lack of critical studies on the topic. we agree that a more critical attitude is required among researchers in the field. as a general answer to this research question, and as noted earlier, there was a stronger focus on the positive outcomes of mentoring rather than any negative ones. given that the earlier reviews also noted this, a key aspect of any future research needs to be a careful seeking after possible negative effects as well as whether the effects are worth the time and money invested. to continue to propagate the notion that mentoring is important for ecr success without sufficient evidence seems counter-productive. further analysis of both short-term and long-term effects of mentoring on ecr development as well as the existing challenges is recommended. furthermore, apart from trying to position mentoring as something ‘positive versus negative’, it might be worthwhile to control for a wide range of other factors such as age, gender, subject, type of university, etc., in order to account for other direct or indirect effects of mentoring experiences and relationships. 4.4 to what extent does the literature point to future research? the papers pointed towards the future in two ways. on the one hand, several papers made recommendations for future policy and practice, which were mainly concentrated around actions to increase the importance of mentoring and awareness of what good mentoring among mentors consists of. both foote and solem (2009) and gilmore et al. (2013) reflected on the notion of inclusiveness and involvement of mentors in their mentoring practices with students. baker et al. (2014) recommended having more advanced reflections on the fit between doctoral students and the supervisor in order to increase the effectiveness of mentoring. as well, apart from reflecting on what needed to happen in future mentoring, some papers formulated recommendations for future research, such as the need for a better understanding of what is causing good mentoring (kamvounias et al., 2008; noy & ray, 2012; o’meara et al., 2013; baker et al., 2014; van der weijden et al., 2014), as well as the need to enlarge research in terms of countries and disciplines (patton, 2009; bell & treveaven, 2011). as a general answer to this research question, we       boeren et al 77 | f l r     suggest there is room for conducting further research on mentoring, in a broader and more diversified way than it has been conducted until now. 5. significance we undertook this review to assess the state of the literature on ecr mentoring in the past 10 years, post-ehrich (2004) review, to provide a base for future inquiry. we also wanted to consider possible policy implications since the eu has created an imperative for institutions to address the career development needs of post-docs/ early career researchers within the bologna strategy, the result of which has been a proliferation of institutional mentoring. we conclude there is much research still to be done that can better inform our conceptualization and implementation of ecr mentoring support and the development of mentoring programs. below we make specific recommendations for future research drawn from our review. a future research agenda we suggest a key goal is a more robust conceptualisation of ecr mentoring including a welldefined representation of the learning process. it should include at a minimum, starting with the recommendations at the top of this list: • examine the theoretical awareness of mentoring (including organisational and individual obstacles) that exist among ecrs and their mentors. these findings would help to capture the complexity of ecr mentoring support. • study ecrs who are actively engaged in structuring informal learning situations that meet their specific needs at different times, and that as mentees they are free to choose one or more mentors. this fits with the idea of mentoring as support towards self-organisation and self-development, in which early career researchers gain the skills to grow towards independence (gardner, 2008; laudel & glaser, 2008). • further examine the effects of mentorship relationships by looking at different elements of mentoring (especially the functions of mentoring), since there is evidence of some negative consequences related to the mismatch of expertise and/or personality (e.g. ehrich et al. 2004). • seek evidence of how the literature on mentoring is connected with growing ecr confidence and competence as independent researchers and scholars; this would mean linking mentoring to the range of abilities essential for ecrs to develop in relation to research, teaching, management, leadership, intercultural skills, publishing, media use, and expectations regarding social engagement (e.g., debowski, 2012). • explore the potential of a trans-organizational conceptualization of mentoring that addresses transfer across institutions and countries, as mobility and intercultural learning form important aspects of early career researcher experience today (horta, 2009). • consider ways to link mentoring for ecr development to other fields including education (not higher education), business and medicine (ehrich et al., 2004) in order to build up the required resource base and decide on suitable strategies and benchmarks. • examine organizational measures for supportive mentoring systems by integrating formal learning with informal mentoring to support cooperation between less and more experienced researchers and thus encourage collegial relationship across scholarly communities in activities such as publishing, research organization, data collection. this review analysis was conducted post-ehrich review, exploring the period from 2005 to 2014. overall, we consider ehrich et al.’s earlier assessment of the education and business reports on mentoring programs to still hold true – the need to: a) attend to mentors’ experiences our analyses showed a relatively weaker focus on mentors, and implied the need to examine both mentees and mentors’ experiences ; b) pay       boeren et al 78 | f l r     more attention to negative outcomes few authors reported negative outcomes of mentoring, and it is necessary to understand how the system malfunctions in those cases ; c) move beyond the data that have generally been collected (self-report process data) for example, observation approaches have not been used in our articles reviewed and seek evidence of impact on actual behaviour performance. and, in order to assess the value of mentoring, attention should be given to opportunities to collect longitudinal data. keypoints more research is recommended to better inform the conceptualization of mentoring a future research agenda needs to explore formal as well as informal aspects of mentoring it would be recommended to explore mentoring using a range of research methods references åkerlind, g. (2005). postdoctoral researchers: roles, functions and career prospects. higher education research & development, 24(1), 21-40. doi:10.1080/0729436052000318550 bonetta, l. (august 26, 2011). postdocs: striving for success in a tough economy. science careers from the journal science. debowski, s. (2012). strategic research capacity building: investigating higher education researcher development strategies in the united kingdom, united states and new zealand: the winston churchill memorial trust of australia. eby., l., allen, t., evans, s., ng, t., & dubois, d. (2008). does mentoring matter? a multidisciplinary meta-analysis comparing mentored and non-mentored individuals. journal of vocational behaviour, 72(2), 254-267. doi: 10.1016/j.jvb.2007.04.005 ehrich, l., hansford, b., & tennent, l. (2004). formal mentoring programs in education and other professions: a review of the literature. educational administration quarterly, 40(4), 518-540. doi: 10.1177/0013161x04267118 gardner, s. (2008). what's too much and what's too little?" the process of becoming an independent researcher in doctoral education. journal of higher education, 79(3), 326-350. doi 10.1353/jhe.0.0007 hansford, e.c., ehrich, l.c. & tennent, l. (2004). formal mentoring programs in education and other professions: a review of the literature. educational administration quarterly, 40(4), 518-540. doi: 10.1177/0013161x04267118 jepsen, d.m., sun, j. j-m., budhwar, p.s., klehe, u-c., krausert, a., raghuram, s. & valcour, m. (2014) ‘international academic careers: personal reflections’. the international journal of human resource management, 25(10), 1309–1326. doi:10.1080/09585192.2013.870307 hemmings, b. (2012). sources of research confidence for early career academics: a qualitative study. higher education research and development, 31(2), 171-184. doi:10.1080/07294360.2011.559198 kehm, b. (2007). the changing role of graduate and doctoral education as a challenge to the academic profession: europe and north america compared, in kogan, m. &teichler, u. (eds.) key challenges to the academic profession, unesco forum on higher education research and knowledge, 111-124. laudel, g., & glaser, j. (2008). from apprentice to colleague: the metamorphosis of early career researchers. higher education, 55, 387-406. doi: 10.1007/s10734-007-9063-7 lindén, j., ohlin, m. &brodin, e.m. (2013). mentorship, supervision and learning experience in phd education. studies in higher education, 38 (5), 639-662. doi:10.1080/03075079.2011.596526       boeren et al 79 | f l r     mellors-bourne, r., metcalfe, j., & pollard, p. (2013). ‘what do researchers do? early career progression of doctoral graduates. retrieved from http://www.vitae.ac.uk/cms/files/upload/what-do-researchers-doearly-career-progression-2013.pdf. mullen, c., & forbes, s. (2000). untenured faculty: issues of transition, adjustment and mentorship. mentoring & tutoring: partnership in learning, 8(1), 31-46. doi:10.1080/713685508 ragins, b.r. & kram, k.e. (2007). the roots and meaning of mentoring. in b.r. ragins & k.e. kram (eds). the handbook of mentoring at work: theory, research, and practice (pp. 3-15). thousand oaks: sage. sauermann, h., & roach, m. (2012). science phd career preferences: levels, changes and advisor encouragement. plos one, 7(5), e36307. doi: 10.1371/journal.pone.0036307 schuster, s. (2009). bambed commentary: post-phd education. biochemistry and molecular biology education, 37(6), 381-382. doi: 10.1002/bmb.20337 shon, p.c.h. (2012). how to read journal articles in the social sciences. a very practical guide for students. london: sage. the european commission (2011) towards a european framework for research careers, [on-line], available at www.ec.europa.eu/pdf/research_policies/towards_a_european_framework_for_research_careers_fin al.pdf (accessed 18 december 2014) vitae. (2011). principal investigators and research leaders survey. london, uk. papers reviewed baker, v., l., pifer, m., j., & griffin, k., a. (2014). mentor-protégé fit. international journal for researcher development, 5(2), 83-98 http://dx.doi.org/10.1108/ijrd-04-2014-0003 bell, a., &treleaven, l. (2011). looking for professor right: mentee selection of mentors in a formal mentoring program. higher education, 61(5), 545-561. doi10.1007/s10734-010-9348-0 browning, l., thompson, k., & dawson, d. (2014). developing future research leaders. international journal for researcher development, 5(2), 123-134. http://dx.doi.org/10.1108/ijrd-08-2014-0019 cox, m. d. (2011).the impact of communities of practice in support of early-career academics. international journal for academic development, 18(1), 18-30. doi:10.1080/1360144x.2011.599600 foote, k. e., & solem, m. n. (2009). toward better mentoring for early career faculty: results of a study of us geographers. international journal for academic development, 14(1), 47-58. doi:10.1080/13601440802659403 gilmore, j., maher, m. a., feldon, d. f., & timmerman, b. (2013). exploration of factors related to the development of science, technology, engineering, and mathematics graduate teaching assistants' teaching orientations. studies in higher education, 39(10), 1910-1928. doi:10.1080/03075079.2013.806459 hopwood, n. (2010). doctoral experience and learning from a sociocultural perspective. studies in higher education, 35(7), 829-843. doi:10.1080/03075070903348412 hubball, h., clarke, a., & poole, g. (2010). ten-­‐year reflections on mentoring sotl research in a research-­‐ intensive university. international journal for academic development, 15(2), 117-129. doi:10.1080/13601441003737758 kamler, b. (2008). rethinking doctoral publication practices: writing from and beyond the thesis. studies in higher education, 33(3), 283-294. doi:10.1080/03075070802049236 kamvounias, p., mcgrath-­‐champ, s., & yip, j. (2008). ‘gifts’ in mentoring: mentees' reflections on an academic development program. international journal for academic development, 13(1), 17-25. doi:10.1080/13601440801962949 lechuga, v. (2011). faculty-graduate student mentoring relationships: mentors’ perceived roles and responsibilities. higher education, 62(6), 757-771. doi: 10.1007/s10734-011-9416-0 lechuga, v. (2014). a motivation perspective on faculty mentoring: the notion of “non-intrusive” mentoring practices in sciences and engineering. higher education, 68, 909-926 doi10.1007/s10734-014-9751-z       boeren et al 80 | f l r     lindén, j., ohlin, m., &brodin, e. m. (2011). mentorship, supervision and learning experience in phd education. studies in higher education, 38(5), 639-662. doi:10.1080/03075079.2011.596526 main, j. b. (2014). gender homophily, ph.d. completion, and time to degree in the humanities and humanistic social sciences. the review of higher education, 37(3), 349-375. doi:10.1353/rhe.2014.0019 mathias, h. (2005). mentoring on a programme for new university teachers: a partnership in revitalizing and empowering collegiality. international journal for academic development, 10(2), 95-106. doi:10.1080/13601440500281724 noy, s., & ray, r. (2012). graduate students' perceptions of their advisors: is there systematic disadvantage in mentorship? the journal of higher education, 83(6), 876-914. o’meara, k., knudsen, k., & jones, j. (2013). the role of emotional competencies in faculty-doctoral student relationships.the review of higher education, 36(3), 315-347. doi: 10.1353/rhe.2013.0021 patton, l. d. (2009). my sister’s keeper: a qualitative examination of mentoring experiences among african american women in garaduate and professional schools. the journal of higher education, 80(5), 510537 doi:10.1353/jhe.0.0062 remmik, m., karm, m., haamer, a., &lepp, l. (2011).early-career academics’ learning in academic communities.international journal for academic development, 16(3), 187-199. doi:10.1080/1360144x.2011.596702 saito, e. (2012). when a practitioner becomes a university faculty member: a review of literature on the challenges faced by novice ex-practitioner teacher educators. international journal for academic development, 18(2), 190-200. doi:10.1080/1360144x.2012.692322 scaffidi, a., k., & berman, j., e. (2011). a positive postdoctoral experience is related to quality supervision and career mentoring, collaborations, networking and a nurturing research environment. higher education, 62(6), 685-698. doi10.1007/s10734-011-9407-1 van der weijden, i., belder, r., van arensbergen, p., & van den besselaar, p. (2014). how do young tenured professors benefit from a mentor? effects on management, motivation and performance. higher education, 1-13. doi10.1007/s10734-014-9774-5 weaver, d., robbie, d., kokonis, s., &miceli, l. (2012).collaborative scholarship as a means of improving both university teaching practice and research capability.international journal for academic development, 18(3), 237-250. doi:10.1080/1360144x.2012.718993   frontline learning research 2 (2013) 53-69 issn 2295-3159 corresponding author: martin salaschek, university of münster, fliednerstrasse 21, 48149 münster, germany, martin.salaschek@uni-muenster.de http://dx.doi.org/10.14786/flr.v1i2.51 53 | f l r web-based progress monitoring in first grade mathematics martin salaschek a , elmar souvignier a a university of münster, germany article received 6 august 2013 / revised 27 november 2013 / accepted 11 december 2013 / available online 20 december 2013 abstract the purpose of our research was to examine a web-based tool for mathematics progress monitoring in first grade. the newly developed assessment tool uses several robust indicators and curriculum-based measures forming three competences (basic precursors, advanced precursors, and computation) to determine comprehensive early numeracy skills in general education. 373 students completed a total of eight online tests every two or three weeks. results indicate that delayed alternate-form reliability was adequate (rm = .78). repeated measures analyses with post hoc comparisons were used to ascertain the sensitivity to assess learning growth. all three competences showed linear growth rates that were significant over time, but only computation and overall scores produced dependable increases from test to test. predictive validity was determined using two standardised school achievement tests (end of first grade, end of second grade). results indicate high predictive validity of the first four online tests (rm = .67, rm = .66 for 6 months and 18 months prediction). correlations with teacher ratings of their students' skills confirmed this pattern. results from student and teacher questionnaires indicate that the students were able to conduct the tests independently and that a three-week interval was adequate for regular-education use. teachers declared to use the progress monitoring results diversely for classroom purposes. we conclude that the use of a web-based assessment setting with diverse measures is beneficial with respect to psychometric properties and feasibility for frequent use in general education. keywords: early numeracy; mathematics; progress monitoring; web-based assessment 1. introduction learning progress assessment aims at providing teachers with information about learning growth, and using diagnostic information for individualised instruction has been shown to result in higher learning gains (connor, morrison, & petrella, 2004; stecker, fuchs, & fuchs, 2005). especially in first grade, results from kim, petscher, schatschneider, and foorman (2010) show that the slope of learning is highly predictive for future achievement. however, stecker et al. note that teachers need assistance in interpreting and m. salaschek & e. souvignier 54 | f l r successfully using progress monitoring results. progress monitoring tools should therefore provide educators with reliable and comprehensive feedback about students' skills. for successful implementation in regulareducation classrooms, high utility and feasibility is additionally required. this can be achieved with highly automated assessment and feedback systems. traditional progress monitoring tools reliably and validly assess students' performance, but are time-consuming because they usually require face-to-face assessment. in addition, most tools for first grade consist of only a few different curricular tasks, making it difficult for educators to use results for adjustments in classroom work. in the present study, we examined psychometric properties and utility of a web-based progress monitoring tool for first-graders. the tool assesses early mathematics competences comprehensively and allows students to work on the tests independently without teacher aid. 1.1 early numeracy and later mathematical achievement early numeracy plays a vital role for the development of later mathematics performance and general school achievement (aunola, leskinen, lerkkanen, & nurmi, 2004; duncan et al., 2007). thus, much research in the past decade has focused on the identification of relevant skills that children should be proficient in when entering school (berch, 2005; gersten, jordan, & flojo, 2005; jordan, kaplan, oláh, & locuniak, 2006; koponen, aunola, ahonen, & nurmi, 2007; methe, begeny, & leary, 2011; missall, mercer, martínez, & casebeer, 2012). certain number sense abilities seem to form precursors or even gateways for further mathematical achievement, but the definition of number sense remains vague (cf. berch, 2005, for an overview). unlike reading, in which well-defined precursors (such as phonological awareness) have been identified, numeracy seems to develop from a diverse set of mental processes which evolve during childhood. the triple-code model of number processing (dehaene & cohen, 1995; dehaene, 1992, 2011) describes three systems involved in different aspects of number processing (i.e., for nonverbal semantic representations; for verbal representations; and for written numerals) derived from a biological viewpoint. these systems develop independently, and pathways are used for communication when solving mathematical problems. developmental models like the model of early mathematical development, which describes three levels of successional skills (krajewski & schneider, 2009; krajewski, 2008), take up a more growth-oriented stance. in krajewski's model, skills at the second level represent the linking of number words with quantities. these skills proved to be particularly predictive for mathematical achievement at the end of primary school (krajewski & schneider, 2009). 1.2 progress monitoring in early mathematics students at risk of not reaching educational goals can be identified by assessing progress of essential skills, such as curricular abilities and number sense skills, which have been described as "gateway" skills for further mathematical development (clarke, baker, smolkowski, & chard, 2008, p. 48). subsequently, suitable interventions can be implemented. educators can use tools to monitor learning progress over time and thereby identify students who do not improve (at an acceptable rate). assessment tools for this purpose should reliably assess students’ performance level and its development, so that students at risk of not reaching curricular goals can be identified. furthermore, diagnostic information about curricular competences should be provided, which teachers can use for instructional changes. implementation should be efficient and as effortless as possible such that general classroom work is not hindered (förster & souvignier, 2011). one progress monitoring approach for this purpose is curriculum-based measurement (cbm; see deno, 2003, for an overview). in cbm, short tests of important curricular competences are conducted regularly. for early mathematics, the psychometric properties of several cbm tests have been discussed in the literature recently (e.g., chard et al., 2005; clarke et al., 2011; seethaler & fuchs, 2011). much of the recent early mathematics cbm research focuses on a set of measures known as tests of early numeracy m. salaschek & e. souvignier 55 | f l r (ten). ten measures have demonstrated high levels of reliability and predictive value for later mathematics performance in a number of studies during kindergarten and first grade general education (e.g., baglici, codding, & tryon, 2010; chard et al., 2005; clarke & shinn, 2004; missall et al., 2012). ten consist of four measures: (1) oral counting, assessing the ability to count orally; (2) number identification, assessing the ability to verbally identify a written number between 0 and 20; (3) quantity discrimination, assessing the ability to identify the larger of two visually presented numbers; and (4) missing number, assessing the ability to name the missing number from a string of three numbers, with one of the three numbers missing. however, there are several issues still to be worked on if these measures shall serve as a basis for instructional changes in the classroom: first, as methe (2012, p. 68) notes, ten measures "struggle to capture more exact knowledge deficits" because they lack close relation to curricula. results are therefore hard to interpret by educators. measures that relate more closely to specific curricular goals might make it easier for educators to use the diagnostic information for classroom work or further interventions. second, reliability and predictive validity results of the four single measures vary from study to study (see missall et al., 2012, for an overview); missall et al. (l.c., p. 96) ascertain that a combination of several measures seems to result in elevated technical adequacy. as a consequence, the authors call for progress monitoring tools which assess early mathematics more comprehensively. third, with the recent exception of a study by hampton et al. (2012), most studies report results from only two or three data points and interpolate learning growth between them. this procedure does not allow a timely evaluation of individual learning growth and also leaves the possibility of non-linear growth patterns. this aspect is especially relevant in the light of low (interpolated) weekly growth rates that often do not exceed 0.30 points per week (foegen, jiban, & deno, 2007). low average growth rates make it more difficult to interpret stagnating scores as at-risk. finally, ten measures are time-consuming to implement because two of the measures (oral counting and number identification) require students to verbalize their answers and therefore can only be assessed in one-on-one settings. in general education, the time and effort needed are reasons why educators usually do not utilise early mathematics progress monitoring at all or regularly enough to make quick instructional adjustments possible. 1.3 aims of the study in our study we aim to approach the aforementioned issues with a web-based progress monitoring tool for first grade mathematics which is feasible for frequent use in general education. the tool intends to assess mathematics skills comprehensively and includes both precursor and curricular competences. that way, educators are enabled to make inferences about students' strengths and weaknesses for classroom work or intervention. assessment time needs to be low and the retrieval and use of results as effortless as possible. psychometric properties of the test concept should be sufficient for dependable estimations of students' shortterm and long-term curricular achievements and for the detection of learning growth. students should work on the tests in a motivated manner to obtain valid results. these aims lead to the following research questions: (1) does the progress monitoring tool assess students' performance reliably? (2) as measures of concurrent and predictive criterion validity, do the progress monitoring test scores correlate significantly with results from standardised achievement tests and teacher ratings of students' mathematics performance? (3) are learning gains represented in the test scores? i.e., can increases in test scores be observed when testing frequently? (4) do teachers and students rate the tool and its implementation feasible for frequent use in general education? m. salaschek & e. souvignier 56 | f l r 2. method 2.1 participants and setting two consecutive studies were conducted with a total of 373 first-grade students in 18 regulareducation classrooms (see table 1 for demographics). the studies took place in rural and urban areas of germany. eight progress monitoring tests were conducted in both studies in intervals of either two weeks (study 1, november 2010 to march 2011) or three weeks (study 2, november 2011 to may 2012). figure 1 provides an overview of the time structure and main dependent variables of the two studies. in study 1, a number of additional measures was obtained: three different standardised paper-pencil tests (pp1-pp3) were conducted, assessing relevant curricular competences of each time point. pp1 was conducted immediately before the first progress monitoring test, pp2 immediately after the last progress monitoring test. eight of the 10 classrooms in study 1 (148 students) participated in a follow-up paper-pencil test approximately 14 months later at the end of second grade (pp3). teacher ratings of students' overall mathematical competence were obtained before each of the three school achievement tests. at the end of first grade, teachers were also surveyed about the feasibility of the web-based progress monitoring tool and their use of the results. students completed a short questionnaire about the progress monitoring test before pp2. purpose of study 1 was to obtain detailed information about the tests' validity. study 2 was then conducted to inspect reliability and sensitivity to learning in an extended time-frame. in preparation of study 2, single items were revised pertaining to difficulty and parallelism after study 1. because of student mobility or sick absentees, some data were missing (progress monitoring tests: 0%11%, mmissing = 1.8%; paper pencil tests: 0%-3.6%, mmissing = 1.7%; teacher ratings: 4.5%-23.2%, mmissing = 12.6%). we used multiple imputation with five imputed data sets to handle missing test data (newton et al., 2004). unbiased results can be expected from multiple imputation when data are missing at random (mar; see schafer & graham, 2002, for a discussion of the term) or when auxiliary variables are included in the imputation model which closely relate to the missing data (collins, schafer, & kam, 2001). given the number of strongly correlated variables in our study designs, we assumed that our inclusive multiple imputation model produced results that are not meaningfully biased. where applicable, coefficients reported in the results section were obtained by combining the imputed data sets using the formulas reported by rubin (1987, 1996). table 1 demographics of study participants study 1 study 2 n 220 153 sex girls boys 51% 49% 46% 54% migration background 22% 9% age at first progress monitoring test 6.68 years 6.72 years note. migration background was defined via language(s) spoken at home. students who spoke another language than german at home were categorized as having a migration background. m. salaschek & e. souvignier 57 | f l r figure 1. schematic overview of the time structure of study 1 and study 2. study 1 was conducted from november 2010 to june 2012, study 2 was conducted from november 2011 to may 2012. pp = paper pencil test. 2.2 progress monitoring measures progress monitoring tests consisted of nine measures in three competences with a total of 52 problems (table 2 provides an overview of the measures used in the progress monitoring test in both studies). the tests were completely computerised, and students received detailed audio instructions before each new set of tasks via headphones to eliminate the influence of reading skills. all tasks were in multiple choice format, in which students clicked on the solution they thought to be correct. tests were untimed, and the children worked on them independently without teacher instruction. results were computed as percentage correct, and educators could access results (graphs and tables) at student and classroom level immediately after a test was completed by a student. results could be compared with class means or overall mean scores of all participating classrooms in the study, and results differing more than one standard deviation from the mean could be highlighted. during the two-week/three-week interval of each test, classrooms could choose to test all students during one class period (if computer rooms were available) or consecutively on computers in the classroom, e.g., during self-study periods. a time frame of two weeks per test was initially chosen for particularly close monitoring of learning growth. intervals were extended to three weeks in study 2 as a response to teacher feedback. the test emphasized the gateway role of number sense by assessing two sets of precursor skills, basic precursors and advanced precursors. both competences were closely related to the triple-code model (dehaene & cohen, 1995) and krajewski and schneider's model of early mathematical development (krajewski & schneider, 2009). precursor measures were complemented by relevant curriculum-based computation skills. all measures included questions of varying difficulty to differentiate between weaker and stronger students. four parallel versions (a-d) of the test were created by using item-cloning algorithms for task creation and the selection of distractors (cf. clause, mullins, nee, pulakos, & schmitt, 1998): for every task, attributes that define its difficulty were identified and held constant in the parallel tests (e.g., for an addition task, the size of the second summand and whether crossing the tens boundary was necessary). throughout the school year, each of the four tests was conducted twice to obtain eight data points (sequence a-d, a-d). basic precursors aimed at assessing fundamental skills that students should be proficient at when entering school. basic precursors contained the measures number discrimination (similar to the ten measure quantity discrimination), symbol quantity discrimination, and number identification (also similar to the corresponding ten measure). advanced precursors aimed at more sophisticated precursor skills, which usually partly develop before school entrance and should soon be mastered during school. advanced precursors contained the measures number sequence 1/number sequence 2 (similar to the ten measure missing number and the next number task used by hampton et al., 2012) and number line, which assesses the extent to which a linear mental number line is developed (see siegler & booth, 2004, for a discussion). study 1 study 2 progress monitoring tests 1-8 (3-week intervals) may june grade 2 pp1 progress monitoring tests 1-8 (2-week intervals) pp2 pp3 nov. dec. jan. feb. mar. apr. m. salaschek & e. souvignier 58 | f l r computation aimed at the main curricular arithmetic goals of german first grade, i.e., handling numbers in the range of 1-20. computation contained addition and subtraction tasks as well as equation problems with dice. table 2 description of progress monitoring measures competence/measure no. of items range example problem distractors task description basic precursors 20 number discrimination 8 1-500 64 | 38 select the larger number symbol quantity discrimination 6 1-10 select the picture with more shapes number identification 6 1-100 audio: "28" 82 | 27 | 72 | 28 | 38 select the number that was given via audio advanced precursors 17 number sequence 1 4 1-20 19, 18, ? 15 | 20 | 16 | 17 select the missing number (steps of 1) number sequence 2 4 1-20 4, 6, ? 10 | 8 | 9 | 7 select the missing number (steps of 2) number line 9 1-20 audio: "12" select the number line that has a mark at the position of the number that was given via audio computation 15 addition 5 1-20 6 + 5 = ? 9 | 10 | 11 | 13 select the correct solution subtraction 4 1-20 15 8 = ? 7 | 9 | 23 | 5 select the correct solution equation 6 1-10 4 + 4 | 7 + 3 | 4 + 3 select the problem with the same solution as the dice problem note. all measures contained problems of varying difficulty, e.g., lower or higher numbers. detailed task descriptions were provided via headphones in language suitable for children. 2.3 criterion measures the three paper-pencil achievement tests in study 1 were selected with reference to their curricular adequacy of the given time points. e.g., at the beginning of grade 1, an achievement test suitable for whole classrooms cannot yet test curricular competences which are only expected to develop during the school year. for this reason, the osnabrück test of number concept development (otz; van luit, van de rijt, &hasemann, 2001) was chosen as pp1. the otz is suitable for children age 4.5 to 7.5 and assesses m. salaschek & e. souvignier 59 | f l r precursor skills such as counting, sorting, and comparing quantities. at the end of first grade, the german mathematics test for first grade (demat 1+; krajewski, küspert, & schneider, 2002) was chosen as end-ofyear criterion (pp2). the demat 1+ was developed following models of early mathematical development, but mainly assesses curricular goals from first grade, e.g., addition/subtraction in the range of 1-20 and (de)composition of numbers. at the end of second grade, the german mathematics test for second grade (demat 2+; krajewski, liehm, & schneider, 2004) was chosen for inspecting long-term predictive validity (pp3). the demat 2+ assesses the main curricular goals from second grade, e.g., basic arithmetic operations in the range of 1-100, number properties, and geometry problems. paper pencil tests were groupadministered within one 45-minute period in all classrooms 1 . all paper-pencil data were collected and put in by trained university students. results were calculated automatically from raw test answers to prevent scoring errors. before each paper pencil test, teachers were asked to rate each of their students' overall mathematic competence on a 7-point likert scale. 2.4 usability and practicality for study 1, several measures of feasibility of the progress monitoring tests were assessed. students were surveyed about the computer tests after completion of all eight probes, asking (1) how they liked the tests, and (2) how they would like to do more tests in the next school year. a 5-point likert scale using smiley faces was used as answer format. additionally, as a measure of direct usability, the time needed to complete each test was logged by the test system. finally, all 10 teachers from study 1 completed a survey about implementation time and their usage of test results. 3. results study 1 3.1 internal reliability we computed the internal reliability for total scores and the three competences. mean reliability of total scores was .86 and varied within a narrow range, demonstrating good overall internal consistency. reliabilities of the single competences were lower: while advanced precursors showed satisfactory reliability, coefficients of basic precursors and computation ranged from low to acceptable (see table 3). 1 otz tasks were slightly adjusted to allow group administration (no german standardised paper pencil test that originally allows group administration was available). for demat 1+ and demat 2+, one task was omitted that had not been introduced in any of the participating classes at the time of testing. thus, overall results are not directly comparable to the reference sample reported by the test authors. m. salaschek & e. souvignier 60 | f l r table 3 internal consistencies of progress monitoring overall scores and competence scores  progress monitoring overall score basic precursors advanced precursors computation time 1 .84 .65 .72 .71 time 2 .86 .60 .78 .71 time 3 .85 .62 .79 .69 time 4 .85 .55 .81 .74 time 5 .87 .65 .83 .74 time 6 .86 .64 .80 .76 time 7 .88 .66 .82 .79 time 8 .88 .65 .84 .79 m .86 .63 .80 .74 3.2 concurrent and predictive validity 3.2.1 school achievement tests 2 as a measure of concurrent validity, correlations between the progress monitoring tests and grade 1 fall pp1 scores were moderate, with .40 ≤ r ≤ .50. to assess the progress monitoring tests' capacity to predict later mathematics performance early in the school year, correlations between the first four tests and grade 1 spring pp2 scores were calculated. coefficients were higher, with .64 ≤ r ≤ .71, indicating strong predictive validity for the end-of-year performance. correlations between the first four progress monitoring tests and pp3 scores at the end of grade 2 were only slightly lower, with .61 ≤ r ≤ .68. later progress monitoring tests related to the pp2 and pp3 scores to a somewhat lesser degree (see table 4). 2 our study design resulted in data with a hierarchical structure (students nested in classrooms), and some intra-class correlations (icc) suggested that error variances may be underestimated if this was not accounted for (the mean icc for all progress monitoring and paper pencil tests was .08). we therefore performed multi-level modelling (using mplus 7.11) in addition to single-level modelling for all correlational analyses in both studies. concerning correlations, the maximum absolute difference between the methods in study 1 and 2 was .04 and .03, respectively. the mean difference of all correlation coefficients was <.01 and .01, respectively, with multi-level mean correlations being marginally higher in study 2. furthermore, there was no meaningful difference in the mean standard error (mdiff < .01; the single maximum absolute difference was .03), and all p levels were identical. because of the relatively small number of classrooms and because single-level results are slightly more conservative, we report results from single-level analyses. m. salaschek & e. souvignier 61 | f l r table 4 concurrent and predictive validity of progress monitoring scores measure 1 2 3 4 5 6 7 8 9 10 1. time 1 2. time 2 .74 3. time 3 .70 .80 4. time 4 .67 .74 .76 5. time 5 .62 .69 .69 .73 6. time 6 .64 .67 .74 .77 .73 7. time 7 .59 .59 .70 .76 .74 .80 8. time 8 .54 .59 .66 .68 .68 .75 .76 9. pp1 .41 .50 .47 .44 .45 .47 .43 .40 10. pp2 .64 .66 .65 .71 .62 .58 .59 .61 .46 11. pp3 a .61 .68 .65 .68 .51 .56 .57 .50 .42 .76 note. all correlation coefficients were statistically significant at an alpha level of p < .001. pp = paper pencil test a n = 148 3.2.2 teacher ratings teachers' ratings of their students' mathematical ability were correlated with the progress monitoring test scores (see table 5). results initially revealed low to moderate correlations between the progress monitoring scores and ratings provided at the beginning of grade 1 (teacher rating 1; .29 ≤ r ≤ .42). correlations with ratings provided at the end of grade 1 were substantially higher (teacher rating 2; .54 ≤ r ≤ .64) and remained stable for ratings provided at the end of grade 2 (teacher rating 3; .54 ≤ r ≤ .66), indicating high predictive validity. m. salaschek & e. souvignier 62 | f l r table 5 correlations between progress monitoring scores and teacher ratings of students' mathematical ability, provided at grade 1 fall, grade 1 summer, and grade 2 summer progress monitoring teacher ratings 1 2 3 time 1 .39 .60 .60 time 2 .40 .62 .66 time 3 .42 .64 .66 time 4 .37 .59 .62 time 5 .38 .54 .58 time 6 .34 .60 .59 time 7 .29 .54 .58 time 8 .37 .56 .54 note. all correlation coefficients were statistically significant at an alpha level of p < .01. 3.3 usability and practicality median test time for the first progress monitoring test was 15.48 minutes (sd = 4.81). later test times were considerably lower and declined continuously, from 13.85 minutes for test 2 (sd = 4.37) to 8.20 minutes for test 8 (sd = 3.81). the difference between the first test and all other tests was partly due to initial starting introductions to the test (approx. 1 minute) and to the students' unfamiliarity with the system. in the survey about the progress monitoring tests, students rated the tests highly, with mean scores of 4.28 (sd = 1.05) on the question, "how did you like the tests?" and 4.34 (sd = 1.13) on the item, "would you like to do the tests again next school year?" (on a smiley faces scale from 1, very unhappy to 5, very happy). 4% and 7% of the students rated the items negatively (scale points 1 or 2), opposed to 71% and 78% positive ratings (scale points 4 or 5). the 10 teachers who participated in study 1 gave similar estimations in the questionnaire provided after completion of the progress monitoring tests. on the 4-point likert scale (disagree to agree), all teachers agreed that, "most of the students had fun completing the tests" (m = 3.70). the same distribution of answers was found for the item, "the students were able to conduct the tests independently". nine teachers stated that the added benefit of the tool was worth the additional timely effort (m = 3.10). moreover, these teaches stated that they would continue to use the system in the next school year (m = 3.60) and recommend the program to fellow colleagues (m = 3.50). teachers declared that they used the progress monitoring results diversely for classroom purposes. apart from obtaining general performance information at student and class level (100%, 70% agreement, respectively), teachers found the information especially useful when they were previously unsure of a student's performance (70% also used the system for this purpose). most teachers adjusted their estimate of students' performance for some students (80% agreement) and claimed to have at least sometimes given weaker or stronger students adjusted exercises based on progress monitoring test results (70%, 90% agreement for weaker or stronger students, respectively). eight teachers stated that supplementary education for weak students was offered at their schools, and information from the progress m. salaschek & e. souvignier 63 | f l r monitoring tests was used for designing the supplementary education at six of these schools. a majority of respondents also found the information important for communicating about performances with students, parents and fellow teachers (90% agreement). the main concern of several teachers participating in the study was the two-week time frame per test in that study. they wished for three-week testing intervals to allow more time for analysing and working with the results. 4. results study 2 while study 1 evaluated the test's validity as well as its usability and practicality, study 2 focused on the reliability and sensitivity to learning. with respect to the different aims of the two studies, analyses also differed between the studies. additionally, given the extended test intervals and because some of the test items were adjusted concerning their difficulty for study 2, results differ slightly from study 1. 4.1 alternate-form reliability we calculated the delayed alternate-form reliability for each adjacent test (t1 × t2, t2 × t3, … t7 × t8). coefficients ranged from r = .71 to .83 (m = .78), which is a sign for parallelism across tests. parallelism is also indicated by the pattern of correlations between non-adjacent tests (see table 6), which decreased only slightly with increasing amount of time between the probes (e.g., test 1 × test 4). table 6 delayed alternate-form reliability of progress monitoring scores, study 2 progress monitoring 1 2 3 4 5 6 7 1. time 1 2. time 2 .71 3. time 3 .65 .74 4. time 4 .68 .76 .81 5. time 5 .67 .71 .78 .82 6. time 6 .60 .60 .64 .74 .77 7. time 7 .57 .63 .67 .74 .77 .79 8. time 8 .59 .67 .69 .69 .75 .76 .83 note. correlations of same test forms are printed in bold. all correlation coefficients were statistically significant at an alpha level of p < .001. 4.2 sensitivity to learning the test's overall capacity to assess learning gains was determined by calculating growth rates in test scores using linear regression for the eight tests. weekly growth rates were obtained by dividing the resulting slopes by 3 because of the three-week time frame of each test. weekly increases in overall scores of 1.0 percent points could be observed (see table 7; descriptive statistics for study 1 are listed in the appendix), m. salaschek & e. souvignier 64 | f l r with larger weekly gains for advanced precursors and computation skills than for basic precursors. smaller basic precursors gains are mainly due to the symbolic quantity discrimination task which revealed ceiling effects from the first probe (see figure 2). table 7 descriptive statistics and growth rates for competences, study 2 overall score basic precursors advanced precursors computation progress monitoring m sd m sd m sd m sd time 1 62.1 11.7 79.4 13.0 55.9 19.0 46.1 15.4 time 2 66.5 14.0 79.1 12.6 65.6 21.9 50.6 19.1 time 3 70.8 14.5 83.1 12.4 66.6 23.1 59.1 19.5 time 4 74.3 14.8 84.6 10.9 73.0 23.0 61.9 22.9 time 5 75.6 13.6 86.1 10.5 72.5 20.5 65.0 20.7 time 6 78.1 14.5 87.0 11.2 74.3 20.6 70.6 21.5 time 7 82.8 13.4 90.1 9.6 80.2 21.3 76.0 21.3 time 8 81.2 14.4 88.5 11.1 77.8 22.6 75.2 22.6 growth rate 1.0 0.5 1.0 1.5 note. all scores as percentage correct. growth rates are weekly growth rates, calculated as slopes of linear regressions of the 8 tests divided by 3 (because of the three-week delay between each test in study 2). figure 2. growth rates for single measures in study 2 (n = 153). statistical significance of growth rates for overall scores was examined by conducting repeatedmeasures analyses of variance. mauchly's test revealed a violation of sphericity (p < .001). thus, greenhouse-geisser corrections were used (greenhouse & geisser, 1959). results indicate an effect of time, f(5.50, 836.18) = 137.73, p < .001, η² = .48. there was also a significant effect of time for the three single m. salaschek & e. souvignier 65 | f l r competences basic precursors, f(6.22, 945.13) = 35.14, p < .001, η² = .19; advanced precursors, f(6.04, 917.67) = 51.47, p < .001, η² = .25; and computation, f(5.63, 855.82) = 96.95, p < .001, η² = .39. post hoc tests were performed to analyse for significant increases from test to test. all six increases in total scores from test 1 to test 7 were significant (see table 8). however, scores decreased from test 7 to test 8. for basic precursors and advanced precursors, 4 and 3 of the six increases from test 1 to 7, respectively, were significant (p < .05) as well as all six increases for computation scores. decreases from test 7 to 8 were significant only for advanced precursors, t(152) = 1.69, p = .049. table 8 comparisons of mean differences in progress monitoring scores for study 2 comparisons mean score difference (sd) t df p time 1 – time 2 -2.26 (5.20) -5.37 152 < .001*** time 2 – time 3 -2.25 (5.32) -5.23 152 < .001*** time 3 – time 4 -1.80 (4.73) -4.72 152 < .001*** time 4 – time 5 -0.67 (4.46) -1.87 152 .031* time 5 – time 6 -1.33 (5.01) -3.28 152 .001** time 6 – time 7 -2.45 (4.77) -6.18 152 < .001*** time 7 – time 8 0.86 (4.23) 2.24 152 .014* 5. discussion the current study extends the research on progress monitoring for young students by using an automated assessment tool that allows frequent tests in regular-education settings and provides educators with detailed information about students' skills. the primary goal of the study was to determine the adequacy of the newly-developed progress monitoring tool. first-grade students work independently on the short online tests, so that diagnostic information about students' performance and progress is obtained with minimal instructional time. the tool uses a combination of robust indicator and curriculum sampling approaches to comprehensively assess nine short measures of mathematic performance forming three competences. static scores and longitudinal psychometric properties were investigated alongside feasibility and usefulness for instructional changes. first, with regard to reliability, the overall scores of the progress monitoring tests showed good internal consistencies within a narrow range. consistencies of individual competence scores – particularly basic precursors and computation – were considerably lower. low coefficients for basic precursors may be due to ceiling effects; computation consistencies were larger for later tests, which may indicate that the three measures within the competence set are distinct skills at first. the distribution of difficulties (see figure 2) contributes to this interpretation. correlations between adjacent tests as a measure of delayed alternate-form reliability were strong, which indicates reliable assessment of students' performance despite the young age of the students. increasing adjacent-test correlations after test 3 (see table 6) argue that frequent tests are advantageous. second, progress monitoring tests 1 to 4 were closely related to the paper pencil results and teacher ratings at the end of first and second grade (pp2 and pp3). noteworthy is the stability of the predictions over time, which indicates that the progress monitoring tests in the first half of the school year assess skills m. salaschek & e. souvignier 66 | f l r particularly important for long-term mathematics success. somewhat lower correlations between tests 5 to 8 and the standardised tests pp2 and pp3 may be because – as indicated in figure 2 – some children showed ceiling effects at the end of the school year. some ceiling effects are a desired result because test items are designed to represent end-of-year competence goals, which several students typically already reach earlier in the school year. yet, reduced variance of progress monitoring tests is likely to result in a slight reduction of correlations with standardised measures of mathematical competence. progress monitoring results were less closely related to the paper pencil test at the beginning of grade 1, which merely assessed precursor abilities and was only moderately predictive of the results of the later paper pencil achievement tests (see table 4). moderate predictive value was also observed for the first performance ratings by the teachers, who had known their students for about two months at that time (correlations between teacher rating 1 and pp2/pp3 were r = .44 and .43, respectively). thus, in addition to the detailed results on precursor abilities from standardised tests (e.g., otz), the progress monitoring tests can provide teachers with information about students' abilities vital for long-term learning growth. third, the tests proved to be sensitive to learning growth with increasing scores from progress monitoring test 1 to 8 in all competences. however, some scores decreased in the last test, an occurrence which has also been observed in other progress monitoring research when frequent tests were conducted (förster & souvignier, 2011; hampton et al., 2012). for progress monitoring tests 1 to 7, all test-to-test increases were significant for overall scores and computation. for basic precursors and advanced precursors – skills that were expected to be mastered before or soon after school entrance – higher overall scores than for computation were observed, and only some of the increases were significant. thus, growth patterns of these two single competences should be interpreted with caution and over longer time periods. finally, several measures of feasibility and usefulness of the tool showed adequate results. the time that students needed to complete a test was low, and the students were able to work on the tests independently. the remaining implementation effort was justified in the eyes of the teachers, a precondition for frequent and beneficial use. teachers also stated that they used the results in diverse ways for classroom purposes and individualised instructions, although the exact scope of instructional changes remains unknown. to conclude, the study at hand addresses a number of issues that were discussed in previous research. by including measures from two approaches, robust indicators and curriculum sampling, the progress monitoring tool provided teachers with performance information about tasks which are directly related to classroom work. at the same time, the combination of different measures proved to be reliable and highly predictive of students' shortand long-term performance. overall scores increased from test to test for all but the last data point, enabling teachers to judge their students' progress and implement necessary interventions rapidly. low testing times and concise results views provide an adequate basis for use in general education. 5.1 limitations at least five limitations should be taken into account when generalising the findings of this study. first, although the participating classrooms were selected from rural and urban areas in different school districts, all schools were in the same federal state, and results could differ in other regions of germany. second, the differing test intervals and slightly adjusted test items between study 1 and 2 limit the comparability of results between the studies. third, no direct measure of parallel-forms reliability was obtained because different test forms were not administered at the same time. all test items were designed using detailed algorithms to ensure similar difficulties, and narrow-ranging reliability coefficients (a) for adjacent tests in study 2 and (b) for predictive validity in study 1 suggest some degree of parallelism. nonetheless, parallelism of the test concept should be assumed with caution until direct parallel-forms reliability has been determined. fourth, slightly larger test score increases in the first few progress monitoring tests (when students are still somewhat unfamiliar with the computer tests) may indicate some degree of retest effects. however, m. salaschek & e. souvignier 67 | f l r large differences in the slopes of different measures (cf. figure 2) and teachers' ratings of the usability of the tests for children suggest that this effect is small. finally, the added value of the basic precursors competence for the majority of students remains questionable. basic precursors scores showed ceiling effects early, with low internal consistencies and limited increases over time. the competence was included in the test as a measure for skills which students should already have acquired before school entrance. teachers should therefore pay special attention to students who do not reach high basic precursors scores. 5.2 implications for research and practice several different competences were included in the test concept at hand to provide teachers with detailed information about students' strengths and weaknesses, as recommended by methe (2012). overall scores were highly predictive of the students' long-term learning outcome, and teachers stated to utilise the information for individualised instruction and supplementary education. single competence scores in part showed lower levels of internal consistency and sensitivity to learning growth than desired. teachers should thus prefer overall test scores when making high-stakes educational decisions. results of the nine single measures can be used at individual level to detect specific deficiencies that prevent a student from advancing in other competence areas. all in all, general education teachers can use the progress monitoring tool to reliably and quickly assess different aspects of their students' mathematics performance and the development over time. a review by stecker et al. (2005) showed that the use of progress monitoring tools resulted in higher learning gains specifically if educators were provided with diverse information about student competences, which they then utilised for individualised instruction. most participating teachers in our study stated that they used the results to adjust their classroom work. however, the extent and success of these adjustments have not been assessed. we recommend two fields of interest for further research in this domain. first, the specific contribution of single competences for the performance of different groups of students remains to be determined. for low-performing students, certain precursor cut-off scores may provide a more accurate risk estimation of long-term mathematics success than total scores. second, it remains largely unexplored how teachers systematically use progress monitoring information to enhance student learning. although the tool at hand includes several measures that are directly related to the curriculum, the review by stecker et al. (2005) suggests that teachers need additional support with "translating" diagnostic information into improved classroom work. keypoints web-based progress monitoring is used for highly automated documentations of learning progress scores of progress monitoring tests are highly predictive of mathematics performance at the end of first and second grade first-grade students worked on the tests independently and with high satisfaction the short tests with nine different measures in three competences were sensitive to learning growth, showing test-to-test increases teachers stated to use progress monitoring results diversely for individualised instruction m. salaschek & e. souvignier 68 | f l r references aunola, k., leskinen, e., lerkkanen, m.-k., & nurmi, j.-e. (2004). developmental dynamics of math performance from preschool to grade 2. journal of educational psychology, 96(4), 699–713. doi:10.1037/0022-0663.96.4.699 baglici, s. p., codding, r. s., & tryon, g. (2010). extending the research on the tests of early numeracy: longitudinal analyses over two school years. assessment for effective intervention, 35(2), 89–102. doi:10.1177/1534508409346053 berch, d. b. (2005). making sense of number sense: implications for children with mathematical disabilities. journal of learning disabilities, 38(4), 333–339. doi:10.1177/00222194050380040901 chard, d. j., clarke, b., baker, s. k., otterstedt, j., braun, d., & katz, r. (2005). using measures of number sense to screen for difficulties in mathematics: preliminary findings. assessment for effective intervention, 30(2), 3–14. doi:10.1177/073724770503000202 clarke, b., baker, s., smolkowski, k., & chard, d. j. (2008). an analysis of early numeracy curriculumbased measurement: examining the role of growth in student outcomes. remedial and special education, 29(1), 46–57. doi:10.1177/0741932507309694 clarke, b., nese, j. f. t., alonzo, j., smith, j. l. m., tindal, g., kame’enui, e. j., & baker, s. k. (2011). classification accuracy of easycbm first-grade mathematics measures: findings and implications for the field. assessment for effective intervention, 36(4), 243–255. doi:10.1177/1534508411414153 clarke, b., & shinn, m. r. (2004). a preliminary investigation into the identification and development of early mathematics curriculum-based measurement. school psychology review, 33(2), 234–248. clause, c. s., mullins, m. e., nee, m. t., pulakos, e., & schmitt, n. (1998). parallel test form development: a procedure for alternate predictors and an example. personnel psychology, 51(1), 193–208. retrieved from http://search.ebscohost.com/login.aspx?direct=true&db=buh&an=487650&lang=de&site=ehost-live collins, l. m., schafer, j. l., & kam, c.-m. (2001). a comparison of inclusive and restrictive strategies in modern missing data procedures. psychological methods, 6(4), 330–351. doi:10.1037/1082989x.6.4.330 connor, c. m., morrison, f. j., & petrella, j. n. (2004). effective reading comprehension instruction: examining child x instruction interactions. journal of educational psychology, 96(4), 682–698. doi:10.1037/0022-0663.96.4.682 dehaene, s. (1992). varieties of numerical abilities. cognition, 44(1-2), 1–42. doi:10.1016/00100277(92)90049-n dehaene, s. (2011). the number sense. how the mind creates mathematics (2nd ed.). new york, ny: oxford university press. dehaene, s., & cohen, l. (1995). towards an anatomical and functional model of number processing. mathematical cognition, 1(1), 83–120. deno, s. l. (2003). curriculum-based measures: development and perspectives. assessment for effective intervention, 28(3-4), 3–11. doi:10.1177/073724770302800302 duncan, g. j., dowsett, c. j., claessens, a., magnuson, k., huston, a. c., klebanov, p., … japel, c. (2007). school readiness and later achievement. developmental psychology, 43(6), 1428–1446. doi:10.1037/0012-1649.43.6.1428 foegen, a., jiban, c. l., & deno, s. l. (2007). progress monitoring measures in mathematics. the journal of special education, 41, 121–139. förster, n., & souvignier, e. (2011). curriculum-based measurement: developing a computer-based assessment instrument for monitoring student reading progress on multiple indicators. learning disabilities: a contemporary journal, 9(2), 65–88. gersten, r., jordan, n. c., & flojo, j. r. (2005). early identification and interventions for students with mathematics difficulties. journal of learning disabilities, 38(4), 293–304. greenhouse, s. w., & geisser, s. (1959). on methods in the analysis of profile data. psychometrika, 24(2), 95–112. m. salaschek & e. souvignier 69 | f l r hampton, d. d., lembke, e. s., lee, y.-s., pappas, s., chiong, c., & ginsburg, h. p. (2012). technical adequacy of early numeracy curriculum-based progress monitoring measures for kindergarten and first-grade students. assessment for effective intervention, 37(2), 118–126. doi:10.1177/1534508411414151 jordan, n. c., kaplan, d., oláh, l. n., & locuniak, m. n. (2006). number sense growth in kindergarten: a longitudinal investigation of children at risk for mathematics difficulties. child development, 77(1), 153–175. doi:10.2307/3696696 kim, y.-s., petscher, y., schatschneider, c., & foorman, b. (2010). does growth rate in oral reading fluency matter in predicting reading comprehension achievement? journal of educational psychology, 102(3), 652–667. doi:10.1037/a0019643 koponen, t., aunola, k., ahonen, t., & nurmi, j.-e. (2007). cognitive predictors of single-digit and procedural calculation skills and their covariation with reading skill. journal of experimental child psychology, 97(3), 220–41. doi:10.1016/j.jecp.2007.03.001 krajewski, k. (2008). prävention der rechenschwäche. [the early prevention of math problems]. in w. schneider & m. hasselhorn (eds.), handbuch der pädagogischen psychologie (pp. 360–370). göttingen: hogrefe. krajewski, k., küspert, p., & schneider, w. (2002). demat 1+. deutscher mathematiktest für erste klassen. [german mathematics test for first grades]. göttingen: beltz test. krajewski, k., liehm, s., & schneider, w. (2004). demat 2+. deutscher mathematiktest für zweite klassen. [german mathematics test for second grades]. göttingen: hogrefe. krajewski, k., & schneider, w. (2009). early development of quantity to number-word linkage as a precursor of mathematical school achievement and mathematical difficulties: findings from a fouryear longitudinal study. learning and instruction, 19(6), 513–526. doi:10.1016/j.learninstruc.2008.10.002 methe, s. a. (2012). innovations and future directions for early numeracy curriculum-based measurement: commentary on the special series, part 2. assessment for effective intervention, 37(2), 67–69. doi:10.1177/1534508411431256 methe, s. a., begeny, j. c., & leary, l. l. (2011). development of conceptually focused early numeracy skill indicators. assessment for effective intervention, 36(4), 230–242. doi:10.1177/1534508411414150 missall, k. n., mercer, s. h., martínez, r. s., & casebeer, d. (2012). concurrent and longitudinal patterns and trends in performance on early numeracy curriculum-based measures in kindergarten through third grade. assessment for effective intervention, 37(2), 95–106. doi:10.1177/1534508411430322 newton, h. j., baum, c., clayton, d., franklin, c., garrett, j. m., gregory, a., … royston, p. (2004). multiple imputation of missing values. the stata journal, 4(3), 227–241. retrieved from http://www.stata-journal.com/article.html?article=st0067 rubin, d. b. (1987). statistical analysis with missing data (4th ed.). new york, ny: wiley. rubin, d. b. (1996). multiple imputation after 18 + years. journal of the american statistical association, 91(434), 473–489. doi:10.1080/01621459.1996.10476908 schafer, j. l., & graham, j. w. (2002). missing data: our view of the state of the art. psychological methods, 7(2), 147–177. doi:10.1037/1082-989x.7.2.147 seethaler, p. m., & fuchs, l. s. (2011). using curriculum-based measurement to monitor kindergarteners’ mathematics development. assessment for effective intervention, 36(4), 219–229. doi:10.1177/1534508411413566 siegler, r. s., & booth, j. l. (2004). development of numerical estimation in young children. child development, 75(2), 428–444. doi:10.1111/j.1467-8624.2004.00684.x stecker, p. m., fuchs, l. s., & fuchs, d. (2005). using curriculum-based measurement to improve student achievement: review of research. psychology in the schools, 42(8), 795–819. doi:10.1002/pits.20113 van luit, h., van de rijt, b., & hasemann, k. (2001). osnabrücker test zur zahlbegriffsentwicklung [osnabrück test of number concept development]. göttingen: hogrefe. frontline learning research 5 (2014) 92 114 issn 2295-3159 corresponding author: andria andiliou, academic staff developer, university of bristol, andria.andiliou@bristol.ac.uk http://dx.doi.org/10.14786/flr.v2i3.87 92 | f l r creative solutions and their evaluation: comparing the effects of explanation and argumentation tasks on student reflections andria andiliou a , p. karen murphy b a university of bristol, united kingdom b the pennsylvania state university, usa article received 12 february 2014 / revised 24 march 2014 / accepted 15 june 2014 / available online 27 june 2014 abstract creative problem solving which results in novel and effective ideas or products is most advanced when learners can analyze, evaluate, and refine their ideas to improve creative solutions. the purpose of this investigation was to examine creative problem solving performance in undergraduate students and determine the tasks that support critical selfevaluations of creative solutions by comparing alternative types of reflective tasks. participants (n = 103) first provided demographic information and responded to individual difference measures (i.e., divergent thinking, need for cognition, and beliefs about creative outcomes) and then read a problem scenario in which they assumed the role of a high school teacher who was asked to design a creative college preparatory course. following, participants completed either an explanation reflective task or an argument based reflective task. finally, participants evaluated their proposed course by rating it on characteristics that describe the originality and effectiveness of creative solutions. findings confirmed the role of divergent thinking as a positive predictor of the originality of a creative solution, whereas, need for cognition, and academic major were positive predictors of the effectiveness of a creative solution. participants rated their creative solutions differentially depending on their beliefs and the type of reflective task. those whose beliefs aligned better with conceptualizations of creative outcomes assessed more positively the originality and effectiveness of their solution. the findings indicate that the argumentation task could potentially promote reflective and critical thinking about a creative solution as participants who completed the argumentation task evaluated their solution more conservatively. keywords: creative problem solving; creativity beliefs; self-evaluation; reflection; argument diagrams a.andiliou 93 | f l r 1. introduction creative problem solving is manifested in everyday life situations as well as in academic contexts (diakidoy & constantinou, 2001). it represents a goal directed cognitive process that results in the production of original and effective solutions when no obvious solution method is available (antiliou, 2012). consider the example of a mud engineer trying to find an innovative solution to prevent oil leaks from an underwater well, the example of a community center manager attempting to design activities that address the needs of a diverse community, or a group of preschool children trying to improvise the rules of a game to accommodate more players. in all cases, the situations call for novel but still effective solutions. many countries around the world have identified the development of creative thinking across subject areas as a core student learning outcome (diakidoy & constantinou, 2001) and have pushed for problembased approaches that provide students with opportunities to construct creative solutions to authentic problems. however, students are often overwhelmed when they are asked to submit creative assignments or generate creative solutions. arguably, one possible explanation for these negative reactions is that students have to rely on their creativity beliefs, but they are uncertain about the characteristics of creative outcomes. consequently, they are unsure about how creative their proposed solution is or how to determine the creativeness of their solution. unfortunately, the role of beliefs has not been adequately explored with respect to creative problem solving performance. as such, the first goal of the present study was to examine the contribution of cognitive and affective individual difference variables including an individual’s beliefs on creative performance. besides creative performance per se, researchers have explored the self-evaluations of the proposed solutions and research evidence suggests that students’ evaluations of self-generated solutions are superficial rather than reflective (runco & chand, 1994). in addition, students tend to engage in case-building about a solution. instead of critically judging a solution, students argue and justify their solutions by discounting potential obstacles and consequences or curtailing the importance or extensiveness of the problem (byrne, shipman, & mumford, 2010; daily & mumford, 2006; nussbaum, 2008). in order to promote more critical reasoning and reflective evaluations in problem solving, researchers examined the effectiveness of structure supports such as prompts (chen & bradshaw, 2007; ge & land, 2003), directions (feretti, macarthur, & dowdy, 2000; nussbaum & sinatra, 2003; nussbaum & kardash, 2005), cases (choi & lee, 2009; hernandez-serano & jonassen, 2003), visual representations (nussbaum, 2008; nussbaum & schraw, 2007), collaborative reasoning, and argumentation tools and tasks (cho & jonassen, 2002). these types of structure supports engage students in thinking about other perspectives, opinions, and approaches to the problem. this is particularly the case when the structure support involves argumentation as a mechanism to elaborate, make explicit the reasoning underlying the problem solving and to foster reflection about a solution (andriessen, 2006; oh & jonassen, 2007). a type of structure support, argumentation tools, were found to significantly improve students’ argumentation skills and group problem solving with some marginal effects on individual problem solving (cho & jonassen, 2001; oh & jonassen, 2007; uribe, klein, & sullivan, 2003). however, past research has not specified how argumentation tools and specifically, argument diagrams influence the self-evaluation of a solution when the problem calls for creative solutions which are original and effective. thus, the second goal of this study was to address this gap in the literature by investigating the effects of an argument diagram on the self-evaluation of a creative solution that participants forwarded to a course design problem. 1.1 creative problem solving several models have been proposed to describe creative problem solving including: the simplex model of creative process (basadur et al., 1994), the creative problem solving framework (isaksen et al., 1994), and the model of creative thought (mumford et al., 1991). in the simplex model the problem solver moves in cycles of ideation and evaluation that occur in different phases of problem solving during which the learner generates and formulates the problem, solves the problem and implements the relevant, appropriate, and original ideas (runco & chand, 1994). according to the creative problem solving a.andiliou 94 | f l r framework, the learner needs to understand the problem, generate solution ideas, and plan for action by developing solutions that could be effectively implemented (treffinger, 1995). in mumford et al. (1991) model of creative thought the learner combines and reorganizes categories or concepts to develop a new understanding of the problem (ideas), which are then evaluated and implemented. these aforementioned models illustrate that creative problem solving evolves across several cognitive subprocesses initiated by the construction of a problem space, the generation of ideas, and the evaluation of a selected solution. creative problem solving is a form of ill-structured problem solving, which results in the production of original and effective solutions (antiliou, 2012). drawing on a review of the empirical literatures of creative and ill-structured problem solving, certain individual difference variables were expected to have an effect on the creativity of a solution with respect to its originality and effectiveness. specifically, divergent thinking that is the ability to generate multiple ideas was found to be predictive of creative problem solving performance (e.g., hunter et al., 2008; reiter-palmon et al., 2009; 1997). in addition, research evidence indicated that need for cognition that represents an individual’s tendency to engage in and enjoy effortful cognitive endeavours (cacioppo, petty, & kao, 1984) also, predicts creative problem solving performance (butler et al., 2003; hunter et al., 2008; osburn & mumford, 2006). students’ domain knowledge and their beliefs about the characteristics of creative outcomes can potentially exhibit an influence on creative problem solving performance. research evidence suggests that problem solvers’ perceptions of a task and their domain knowledge impacts the search for relevant information, the representation of the problem, and the evaluation of potential solutions (jonassen, 1997; voss et al., 1991). also, evidence indicates that knowledge of the important concepts and principles of a domain contributes in better performance on ill-structured tasks (shin, jonassen, & mcgee, 2003; voss & post, 1988) and serves as the foundation of creative solutions (weisberg, 2006). evaluation is a component of problem solving and it represents the metacognitive process during which problem solvers reflect and assess a proposed solution. evaluation is essential for ill-structured problems that require a creative solution because it is the process by which problem solvers can determine whether a proposed solution meets the creative criteria of originality and effectiveness. according to voss and her colleagues (1981) argumentation is a means for problem solvers to evaluate more analytically a solution by elaborating and clarifying the solution and identifying its limitations. researchers have experimented with graphic organizers such as argument diagrams in order to promote critical and reflective thinking in writing tasks. for example, nussbaum and schraw (2007) found that argument diagrams supported better integration of arguments and counterarguments in writing tasks. for the present study, our second aim was to determine whether an argument task (i.e., argument diagram) promotes more reflective critical self-evaluations of a potentially creative solution in comparison to an explanation task. 1.2 explanation and argumentation for critical thinking explanation and argument tasks have been used to promote and assess understanding, critical thinking, conceptual change, and problem solving (reznitskaya, anderson, & kuo, 2007; jonassen & kim, 2010; nussbaum & sinatra, 2003; willey & voss, 1999). explanation is a constructive learning activity during which learners elaborate and clarify an idea by explaining it to oneself and it was found to lead to enhanced learning, more accurate self-assessments, and more effective problem-solving (fonseca & chi, 2011). when learners elaborate they generate inferences and integrate information with prior knowledge. the self-explanation effect was found to be positive both for learners with low and high prior knowledge. a possible interpretation of this result is that for individuals with low knowledge, self-explaining allows them to generate inferences to fill their knowledge gaps and for learners with high prior knowledge, selfexplaining allows them to repair their existing mental models (chi, 2000; fonseca & chi, 2011). research findings indicate that self-explanation can be a powerful learning strategy due to the underlying cognitive mechanisms that allow learners to identify and remedy knowledge gaps by generating inferences and to develop and repair their knowledge representation models. thus, when learners move beyond simply knowledge-telling with summaries or paraphrased statements and are engaged in self-explanation through inference generation and knowledge integration they seem to gain a deeper understanding. however, even if a.andiliou 95 | f l r self-explanation to one-self represents a constructive learning activity it was found to be somewhat less effective in learning and problem solving tasks when compared to more interactive learning activities such as responding to question prompts, explaining to someone else, and discussing with a partner to generate collaborative explanations. in the present study, within the context of a problem solving task an explanation prompt was compared with an argument task to examine the degree to which the tasks promote reflective thinking when evaluating a proposed creative solution. argumentation was conceptualized by kuhn (1991) as the cognitive process of formulating and weighting the arguments for and against a course of action, a point of view, or a solution to a problem. argumentation skills are comprised of the skill to generate reasons, offer evidence, and provide counterarguments and rebuttals. three theoretical frameworks have been applied to analyze argumentation based on rhetorical and dialectical arguments in educational settings. rhetorical arguments are put forward to persuade or convince others about a claim or proposition without consideration to alternative positions (toulmin, 1958). dialectical arguments are based on the dialogue between supporters of alternative positions during a dialogue game or a discussion (jonassen & kim, 2010). through dialectical argumentation within an individual or within a group, individuals resolve differences, compromise between multiple opinions, and convince on the advantages of a position. researchers in the learning sciences draw primarily on three theoretical approaches to analyze and evaluate the quality of argumentation: toulmin’s rhetorical argumentation framework, pragma-dialectics (van eeemeren & grootendorst, 1992) and walton’s (2000) dialogue theory. toulmin has proposed an argument scheme to describe the structure of effective argumentation that includes a sequential set of components: a claim that expresses the position, facts or opinions that serve as data in support of the claim, a warrant as justification, and elaborative elements such as a backing, qualifier, and rebuttal to potential counterclaims. although toulmin’s framework is useful for analyzing rhetorical argumentation to determine the soundness and effectiveness of a line of reasoning of an individual (andriessen, 2006), it has two primary limitations: its complexity (e.g., warrants are often implicit) and its focus on the perspective of one proponent (leitão, 2003; van eemeren & grootendorst, 1999) instead of argumentation as “a discourse phenomenon” (andriessen, 2006) especially with reference to educational contexts. two theoretical models that are more applicable to the dialectical nature of argumentation as it is manifested in educational contexts are the pragma-dialectics (van eeemeren & grootendorst, 1992) and dialogue theory (walton, 2000). based on the pragma-dialectics, argumentation is a means of resolving differences of opinions through critical discussions that evolve in four stages. first, people present their positions at the confrontation stage, they assume their roles and agree on procedures at the opening stage, they defend and challenge during the argumentation stage and at the concluding stage they decide who has prevailed the critical discussion. another collaborative view of argumentative discourse is conceptualized in walton’s dialogue theory (walton, 2000) in which walton argues that argumentation is a goal-directed and interactive dialogical activity during which individuals reason together about arguments to generate one proposed solution. walton (2000) identified specific forms of dialogue (e.g., information seeking, negotiation, persuasion, inquiry) along with argumentation schemes comprised of critical questions and moves to model and support argumentation. educators can draw on the schemes suggested in dialogue theory to plan, organize, and evaluate classroom and online discussions, and use argumentation as a vehicle for critical thinking and problem solving. in the present study an overarching critical question was used to stimulate student argumentation with an imaginary group of stakeholders with the purpose of exploring the potential of a creative solution to an authentic problem. 1.2.1 supporting argumentation researchers have documented that acquiring the skills to argue effectively is challenging both for adolescents and young adults (felton & kuhn, 2001; reznitskaya et al. 2001). in order to engage students in argumentation and promote the development of argumentation skills educational researchers have designed and experimented with argumentation supports in the contexts of reading, writing, and problem solving tasks. among these argumentation supports are directions, computerized and face-to-face collaborative argumentation, and visual argumentation aids. a.andiliou 96 | f l r goal directions have been used as a means to promote reflective argumentation in problem solving and writing tasks. for example, nussbaum and sinatra (2003) asked undergraduate students who provided a wrong answer to a physics problem in which they had to predict the path of a falling object to counter-argue by providing reasons why a person would hold an opposing position. the researchers found that students who proposed counterarguments had a more integrated understanding of the problem situation and the important underlying concepts. the effectiveness of directions in supporting argumentation varied based on the goal they conveyed. when students were instructed to persuade an audience instead of explaining their position or solution, evidence indicated that they engaged in case-building, the overall quality of writing was poorer with fewer counterarguments but more reasons in justification of their position. alternatively more specific directions that guided students to generate complete arguments were more effective in facilitating student argumentation. when nussbaum and kardash (2005) gave goal directions that varied in generality (i.e., opinion, reason, counterargue/rebut), they found that the group that received more specific directions to persuade by generating reasons, evidence, counterclaims and rebuttals, produced writing of better overall quality, more balanced, and the participants offered more counterarguments and rebuttals. in a follow-up experiment, nussbaum and kardash (2005) compared the effects of two types of goal directions (e.g., express an opinion or persuade an audience) and the effects of a two-sided non-refutational text. undergraduate students who were directed to express an opinion and read the text produced essays of better overall quality, wrote more elaborative arguments and offered more counterarguments in comparison to those who were directed to persuade as the text stimulated students’ thinking. directions to persuade had a significant negative effect on the overall quality of argumentation only for students who did not read the text. researchers raised caution about persuasion directions as it is possible that students’ rely on a misconception that they are more effective in convincing an audience by elaborating on their position than raising counterarguments (ferretti, macarthur, & dowdy, 2000; nussbaum & kardash, 2005). in order to promote more balanced and reflective reasoning researchers have utilized other structure supports such as collaborative reasoning discussions and visual argumentation aids including computer-scaffolding tools, outlines and diagrams. researchers have investigated the effect of collaborative argumentation both computerized and faceto-face to facilitate students’ critical reasoning and argumentation within the context of problem solving tasks. in two exemplar studies researchers examined the effects of argumentation scaffolds and question prompts on the quality of argumentation, the group problem solving performance and transfer to individual problem solving (cho & jonassen, 2003; oh & jonassen, 2007). typical argumentation scaffolds included sentence openers that helped to explicate a solution, agree or disagree with a solution, put forward evidence, and elaborate on the solution. in addition, guidance questions functioned as scaffolds (e.g., how can you verify the accuracy or value of your solution?) in the collaborative discussion environments. researchers found that the argumentation scaffolds improved the quality of the discussion in terms of the number of argument components including claims on how to solve the problem and evidence to support the solution (cho & jonassen, 2003; oh & jonassen, 2007). there was also improvement in the overall quality of problem solving subprocesses including problem definition, selection of relevant information, hypothesis generation and testing, solution development and evaluation. however, in both studies the researchers did not detect significant transfer effects of argumentation scaffolding on individual problem solving. thus, suggesting that learners may need long-term and more comprehensive opportunities for extended engagement in collaborative problem solving to effectively transfer and apply argumentation skills in individual problem solving. collaborative discourse was also effective in facilitating argumentation when combined with instruction on basic argumentation concepts and reading of multiple texts. in a study of middle school students, martunen and laurinen (2006) found that after reading three texts and participating in pair conversations on the topic of genetically modified organisms, the student-constructed argumentation diagrams were more elaborative and reflective as they included more themes and arguments. moreover, kim (2001) found that incorporating a metacognitive group monitoring activity in collaborative reasoning discussions had contributed in more dialogic and reflective student writing. in addition, the counterarguments and rebuttals increased in the post-discussion essays and the essays provided evidence that students were attentive to their reasoning by reflecting and evaluating their position. in conclusion, guided a.andiliou 97 | f l r opportunities in which learners participate in argument-based discourse facilitate development and internalization of argumentation knowledge and skills, and improve both the quality of the arguments and the peer dialogues. visual argumentation aids such as argument diagrams have been utilized by researchers to promote coherent and organized argumentation with well-integrated arguments and counterarguments in support of a final position. in a series of studies nussbaum and schraw (2007) examined the effectiveness of a graphic organizer to guide more balanced and reflective argumentation. they found that both instruction about the criteria of a good argument and the graphic organizer improved the quality of writing, increased the number of counterarguments, and the overall integration score. however, the students who used the graphic organizer preferred to apply refutation as an integration strategy in comparison with students who received criteria instruction and primarily used weighing and synthesizing opposing perspectives into a creative position. in a follow-up study, that aimed to facilitate students to become more metacognitively reflective and explore perspectives on an issue before integrating them into a final position, nussbaum (2008) modified the graphic organizer to an argumentation vee diagram (avd). in this study nussbaum examined whether an elaborative intervention that utilizes the diagram with instruction on how to integrate arguments and counterarguments, and discussion of the criteria of evaluating the strength of arguments and counterarguments results in better argumentation and has a transfer effect. the experimental group improved their writing in terms of integration over three sessions using most frequently the synthesis strategy but there was no significant transfer effect to a task in which the diagram was removed. in another study, when fifth graders collaboratively used an argument diagram, they generated more coherent arguments than when they collaborated to list pro-con positions (scwarz, neuman, & biezuner, 2000). thus, the studies provide evidence that the argument diagrams have the potential to stimulate consideration of counterarguments and facilitate more elaborated and coherent argumentation but more research is needed to determine whether the use of diagrams enhances reflective and critical thinking about complex issues. virtual graphic tools have also been used to support student argumentation and engagement in critical discussions. typically computerized argumentation diagrams have the capacity to represent both the components of an argument and relations of support and disagreement (jonassen & kim, 2010). in their study of the vcri argumentation tool, munneke, van amelsvoort, and andriessen (2003) examined the role of argumentative diagrams that were constructed in advance individually or collaboratively during an electronic discussion in supporting student interaction with the purpose of writing a collaborative text on genetically modified organisms. the researchers found that diagrams that were constructed in advance helped students to focus their subsequent discussions on argumentation and were used as information sources and collaborative diagrams were also used for note-taking to summarize the discussion. thus, the diagram helped to maintain focus and functioned as an aid for organizing and maintaining coherence during the discussion. munneken and colleagues (2003) noted though, that even if diagrams stimulated collaborative discussions still argumentation was one-sided as most diagrams were very unbalanced. another study conducted by easterday, aleven, and scheines (2007) provided further evidence of the effectiveness of argument diagrams as a graphic organizer for argumentation. learners in this experimental study who analyzed public policy problems using a causal diagram organized better their perceptions of the arguments in comparison with students who only read about the problem in a text. however, students who used the diagramming tool learned more about constructing causal arguments as they were engaged in a more constructive activity while using the tool to formulate their arguments. as newell and colleagues (2011) argued in their review of the studies on teaching and learning argumentation, the diagrams printed or virtual help learners manage the complexities of argumentation and especially the task of considering alternative perspectives and integrating arguments with counterarguments but more evidence is needed to support whether they facilitate more critical and reflective thinking. 1.3 purpose of the study the purpose of the study was to examine creative problem solving performance in undergraduate students and compare how alternative tasks (e.g., explanation or argumentation) support reflective selfevaluations of creative solutions. two research questions guided our investigation: a.andiliou 98 | f l r how do individual differences in divergent thinking, need for cognition, beliefs about creative outcomes, and academic major impact the creativity of a solution with respect to its (a) originality and (b) effectiveness? to what extent does a reflective task (i.e., an explanation task or an argumentation task) differentially support the students’ self-evaluation of their creative solution? based on the review of literature the following five hypotheses were forwarded: hypothesis 1.1. students who are strong divergent thinkers and high in need for cognition will propose creative solutions that are both original and effective. hypothesis 1.2. students who conceptualize creative solutions as both original and effective will develop a solution that is highly effective and may or may not be as original. these students will also evaluate their creative solutions more positively. hypothesis 1.3. students who possess more extensive prior knowledge, based on their academic major, will develop highly effective solutions. hypothesis 2.1. for students who complete the argumentation task, the effectiveness of the proposed creative solution will strongly and positively predict the self-evaluation of the solution with respect to its effectiveness. hypothesis 2.2. for students who respond to the explanation task, the effectiveness of their proposed creative solution will be less predictive of their self-evaluation of the effectiveness of the solution. 2. method the purpose of this study was to explore creative problem solving performance in undergraduate students and compare alternative tasks that support reflective self-evaluations of their proposed creative solutions. the study was designed based on a single-factor, between groups design with two comparison groups (i.e., explanation or argumentation task). 2.1 participants for this study participants were recruited from an undergraduate educational psychology course at a public research university in the united states. the completion rate was 82% with 103 volunteers completing the study. the sample was comprised of primarily sophomores (52%) and freshmen (30%), the majority were females (n=88), and more than half of the participants were education majors (57%) in comparison to 43% of non-education students (e.g., communication sciences and disorders, kinesiology). the demographics were comparable to most introductory courses required for teacher certification. 2.2 measures 2.2.1 demographics respondents completed a demographic cover page in which they provided background information including their academic major, academic classification, courses they completed in preparation for the transition to college, and courses they enrolled or completed pertaining to curriculum and instruction. participants also listed and described their teaching experiences. 2.2.2 divergent thinking the two divergent thinking tasks were derived from the tasks in guilford’s consequences’ test a’ (christensen, merrifield, & guilford, 1953). for each task students had two minutes to generate as many possible results to each of these hypothetical scenarios: (a) what would happen if a new invention makes it unnecessary for people to eat? (b) what would happen if a new invention makes it unnecessary for people to a.andiliou 99 | f l r sleep? the responses were scored for ideational fluency operationalized as the number of distinct valid ideas recorded by a respondent excluding any duplicates or irrelevant ideas due to a misinterpretation of the scenario. on average, for the two divergent thinking tasks participants generated m1=6.11(2.46) ideas and m2=5.52(2.27). due to their marginal internal consistencies (α=.63), the two scores were entered as separate divergent thinking indicators for data analysis. 2.2.3 beliefs questionnaire a 28-item likert scale designed for the purposes of this study was administered to gauge participants’ beliefs about creative outcomes with reference to a creative course. twelve items targeted characteristics of a creative course related to its (a) originality (i.e., innovative, unusual, original, novel, unique, and imaginative) and (b) effectiveness (i.e., successful, affordable, effective, implementable, goaldirected, and feasible). these characteristics are recurring terms that describe creative outcomes in the extant literature of creativity and creative problem solving. the remaining 16 items were distracters. an example of an item on the belief scale is: “creative high school courses are implementable.” participants rated the belief scale items with a score ranging from not very (0) to very (5) to indicate how typical the characteristic is of a creative course. a factor analysis with a principal axis factoring (paf) and a promax rotation was conducted with the 12 characteristics of creative courses to determine the underlying structure of the belief scale. the promax rotation was selected because it is a type of rotation that aids the interpretation of the factor analysis results when the factors are believed to be correlated as in this case (r12=.52). three factors were extracted with eigenvalues of 37.066, 10.688, and 6.996 respectively and they exceeded the criterion of 1.0 based on the kaiser-guttman rule (guttman, 1954; kaiser, 1960). however, only two underlying factors were detected in the scree plot. the characteristic affordable was the only one that loaded on the 3rd factor, which explained 6.996 % of the data. this characteristic was the only item that targeted financial aspects of a creative course and this is possibly why this characteristic failed to load on the two first factors that represented the effectiveness and originality of a course. thus, this item was removed and a second factor analysis was conducted with the 11 items. two factors emerged from the final factor analysis with eigenvalues of 4.828 and 1.497, which explained 39.984% and 9.917% of the variation in the data, respectively. as evidenced in table 1, nine items had loadings greater than the harman criterion value of .40. the two detected factors represent underlying characteristics of creative courses with the first factor representing the effectiveness dimension and the second factor representing the originality dimension of a creative course (r12 =.57). characteristics that underlie the effectiveness of a creative course included successful, effective, innovative, implementable, and feasible. characteristics that underlie the originality dimension of a creative course included the characteristics imaginative, unique, novel, and original. a belief scale with the nine items was formulated with acceptable internal consistency (α=.87). the internal consistency of the two component subscales were α1= .84 for effectiveness and α2=.81 for originality. the composite score for the entire belief scale ranged from 0 to 45 with higher scores indicating beliefs in agreement with current conceptualizations of creative outcomes in the literature. the average belief score was m=26.89(7.29) suggesting that participants’ beliefs were in moderate alignment with these conceptualizations. table 1 coefficients for the factor analysis with promax rotation for the beliefs scale characteristic effectiveness originality successful .998 -.153 effective .959 -.195 innovative .605 .070 implementable .509 .181 feasible .440 .169 a.andiliou 100 | f l r imaginative .004 .812 unique .081 .757 novel .095 .619 original .233 .619 unusual -.231 .367 goal-directed .325 .307 eigenvalues 4.828 1.497 percentage of variance 39.984 9.917 note. factor loadings >.40 are in boldface. 2.2.4. need for cognition scale the 18-item abbreviated need for cognition scale (cacioppo, petty, & kao, 1984, p.306) was administered (α=.79) to assess participants’ tendency to engage in and enjoy effortful cognitive endeavours. an example item from the scale is the following: “i would prefer complex to simple problems”. participants rated the statements with a score ranging from not very much (0) to very much (5). for the scoring of the scale, a composite need for cognition score was calculated and it ranged from 0 to 90. on average, participants manifested moderate to low need for cognition m=48.5(10.46). 2.2.5 solution self-evaluation questionnaire finally, participants evaluated their creative course on a 16-item likert scale questionnaire developed for the study, which consisted of two distracter items and 14 items that represented criteria of a creative solution with respect to its originality (i.e., innovative, unusual, original, imaginative, novel, unique, and risky) and effectiveness (i.e., effective, successful, affordable, implementable, goal-directed, feasible, and organized). these items reflected descriptive characteristics of creative outcomes identified in the extant theoretical and empirical literature of creativity and creative problem solving. participants rated how creative their proposed solution was based on the aforementioned characteristics on a scale ranging from not very (0) to very (5). a factor analysis with a principal axis factoring (paf) and a promax rotation was conducted with the 14 items after the distracters were removed. two factors emerged with eigenvalues equal to 6.086 and 2.452, which explained 40.58% and 14.01% of the data, respectively. the factor intercorrelation was moderate (r12=.49). table 2 summarizes the loadings on the two factors based on the pattern matrix. based on the results of a factor analysis two subscales were formulated: the originality selfevaluation and the effectiveness self-evaluation subscale with seven items each. both subscales ranged from 0 to 35 and had acceptable internal consistency indices of α1=.87 and α2=.88 respectively. οn average, participants evaluated their proposed course solution low in originality m=18.96(6.73) and moderate in effectiveness m=26.69(5.38). table 2 coefficients for the exploratory factor analysis with promax rotation for the self-evaluation scale characteristic course effectiveness course originality effective .81 .04 successful .78 .10 affordable .74 -.30 organized .73 .10 goal-directed .69 .00 implementable .69 -.09 feasible .65 .04 unique .08 .86 imaginative -.02 .80 a.andiliou 101 | f l r unusual -.13 .75 original .13 .68 novel .06 .68 risky -.40 .60 innovative .34 .56 eigenvalues 6.09 2.45 percentage of variance 40.58 14.01 note. factor loadings >.40 are in boldface. 2.3 procedure participants completed the study through an online survey system (qualtrics) that randomly assigned them to a condition either the explanation (n1=53) or the argumentation (n2=50) task. the study was selfpaced as students completed it in one sitting at their own pace. students first provided demographic information and responded to two counterbalanced divergent thinking tasks. following, they completed a beliefs questionnaire and the need for cognition scale. then all participants read the same problem scenario and in response to it they developed a creative course as a solution to the problem described in the scenario. following participants responded to a reflective task (i.e., an explanation or an argumentation task) about their proposed creative course. finally, all participants evaluated the creativity of their course by rating it on a scale with a set of characteristics that describe the originality and effectiveness of a creative solution. 2.4 problem solving task the problem scenario was originally developed by hunter and his colleagues (2008) for a study of undergraduate students’ idea generation and problem solving. the specific scenario was selected for two reasons: (a) it had been previously used with undergraduate students and has yielded acceptable interrater agreement scores (0.70-0.80) with respect to the originality and quality scores assigned to the solution and (b) the embedded problem solving task is ill-structured as it requires students to: extract the important and relevant information from the scenario, identify the parameters and constraints for solving the problem, apply personal beliefs about creative courses and creative teaching, draw on their knowledge and experiences to define the problem, make judgments, and establish criteria for evaluation. the scenario required participants to assume the role of a high school teacher asked to design a creative college preparatory course for the high school’s seniors to better prepare them for college and reduce the college dropout rate among this high school’s graduates. the final paragraph in the problem scenario explained the task: “in her description of the requirements for the course, the principal makes one point very clear, the senior prep course needs to be a creative high school course designed to prepare the high school students for college. she emphasized that you need to take a creative approach in designing and teaching the course. the principal has asked you to 1) identify the overall goal of the course and 2) list and describe the specific learning activities that you will include in the course.” also, for the purposes of the study we modified the problem scenario in two ways. first, any descriptions or conceptualizations of a creative course were removed so that participants rely on their own beliefs and understandings of a creative course. still we emphasized in the scenario that the problem solver needs to take a creative approach in designing and teaching the course. second, in the final paragraph we identified the two specific tasks that participants had to complete after reading the scenario. we conducted two pilot studies followed by two focus group discussions in order to gather evidence for the comprehensibility and the face validity of the problem scenario and the problem solving task. the pilot study participants were representative of the sample (i.e., non-education and education majors) and in general the students found the directions clear and the scenario understandable. they also acknowledged the authenticity of the problem scenario as they pointed that it challenged them to provide a solution to a real life a.andiliou 102 | f l r problem: the fact that high school students are not prepared for the academic, social, and emotional challenges of the transition to college. pilot study participants said that once they read the problem scenario they had to pause and reflect on what their needs were when they moved to college and what is important for a student to succeed in college. overall, the authenticity of the problem scenario and its relevance to students’ recent college transition experiences seems to have motivated participants to engage with the task as they agreed on the importance of designing a high school course to prepare students for college. 2.4.1. coding a coding scheme was developed to summarize the responses that participants provided to the problem solving task. specifically, the scheme was used to code the learning activities participants suggested for their creative high school course. an iterative procedure was followed to develop the coding scheme and establish its validity. the researcher and an independent coder (coder a) applied a keyword content analysis approach to identify the task-relevant units within a response. a task-relevant unit was defined as any distinct task-relevant statement that captured learning activities that participants generated for their high school course. a learning activity was defined as any learning experience, enactive (i.e., actual doing) or vicarious (i.e., students observe, listen or are engaged in other ways), designed for the learners to attain an instructional goal such as the acquisition of information, knowledge, skills, abilities, attitudes and strategies (antiliou, 2012, p. 66). the first author began by reading all of the responses to generate an initial set of coding categories to summarize the responses to the question prompt that participants recorded. this initial review of responses revealed that participants recorded learning activities as well as assessment activities and other instruction and course design elements such as materials, educational technology, and learning goals. following, the researcher provided directions to another colleague (coder a) to develop independently her version of the coding scheme. the directions included the problem scenario, the problem solving task, a set of coding guidelines and an example of a coded response. then, coder a proceeded to read the entire set of responses to independently generate a second version of the coding scheme. in the two discussions that followed between the first author and coder a, the two independent coders analyzed, compared, and synthesized the two alternative coding schemes to generate a merged coding guide that included a coding scheme and a set of guidelines with definitions, assumptions, and decision rules. there was consensus that the coding scheme should capture task-relevant units that identified not only learning and assessment activities but also other instruction/course design elements (e.g., instructional materials or educational technology, etc.). a total of 47 coding categories were included in the coding scheme and they are summarized under ten overarching categories including discussion-based activities, problem solving activities, experiential learning activities, reading and writing assignments (see table 3 for the complete list of categories). the 47 coding categories represented learning or assessment activities, as well as other instruction/course design elements. six of the 47 coding categories were further divided into four additional specific subcodings. for example, the category modelling had two subcodes: instructor model and other models. also, the category expository writing had two subcodes: extended and brief. in the case of coding categories with more specific subcodes, the coding was done by applying the more specific subcode. the coding guide was used for a trial coding (10%) followed by a discussion to resolve potential differences, refine the scheme, and clarify the decisions rules. following the development and validation of the coding scheme, the intercoder agreement for the reliability of the coded responses was examined. the researcher and coder a independently coded another 20% of the responses for this purpose. in each response, the coders (a) identified the total number of taskrelevant units and (b) coded the task-relevant units including the learning activities or other instruction/course design elements. a total of n=349 valid task-relevant units were recorded by the participants. the intercoder agreement for the total number of task-relevant units was α=.79 and for the type of code assigned to a taskrelevant unit was α=.72. both indices were above the moderate criterion .70 selected for the conservative a.andiliou 103 | f l r kalpha coefficient of intercoder agreement. in a discussion that followed, the coders first resolved disagreements on the number of task-relevant units and then disagreements on the assigned codes in order to reach consensus. 2.4.2. scoring the creativity of a course that participants proposed was operationalized with respect to the average originality and effectiveness of the valid task-relevant units. each valid task-relevant unit that a participant recorded was scored for its originality and effectiveness. among the valid task-relevant units, 315 units were classified as learning or assessment activities. another 34 task-relevant units represented instruction/course design elements, which referred to aspects of instruction or course design including materials, educational technology, the structure of the course, and the learning environment. originality: originality was defined as the rareness of occurrence of a task-relevant unit within the pool of valid units (n=349) generated by all participants. the originality score (x) assigned to a valid taskrelevant unit (i) was the rareness proportion of the specific code within the pool of (a) learning/assessment activity units or (b) instruction/course elements, depending on the nature of the coded unit. for example, a task-relevant unit with the code instructor modelling appeared 55 times so its proportion of occurrence within the pool of learning/assessment activities was 55/315=0.18 and its rareness of occurrence was 10.18=.82. similarly, a task-relevant unit with the code educational technologies appeared 6 times within the pool of the 34 instruction/course design elements, thus its proportion of occurrence was 6/34=0.18 and its rareness of occurrence was 1-0.18=.82. the average originality score for the solution proposed by a participant (j) was: average originality = . the average originality score (x) is the sum total of the rareness proportion (x) for every ith valid task-relevant unit divided by the number of valid task-relevant units (m) recorded by each participant. effectiveness: the potential effectiveness of a learning activity was defined as the degree to which a learning or assessment activity or other instruction/course design element could contribute to the smooth transition and academic success during the first years of college. an effectiveness rubric was developed to operationalize and score each task-relevant unit (see appendix). the rubric was developed by drawing on the literatures of instructional design, and college transition and persistence and instructional design (eggen & kauchak, 2010; goldbrick-lab et al., 2007; louie, 2007; pritchard et al, 2007; roe clark, 2005). the effectiveness scores ranged from inadequate (0) to strong (4) effectiveness. the effectiveness of a learning activity or other task-relevant unit was considered strong if (a) it targeted important information, knowledge, abilities, skills, or strategies for smooth transition and academic success in the first years of college and (b) it strongly aligned (i.e., directly relevant) with the identified overall goal of the course. examples of potentially effective activities include those that targeted writing, note taking, and test taking skills but also coping strategies and interpersonal skills. in order to establish the reliability of the effectiveness scores another colleague was trained to serve as a second rater using a subset of the pilot data. following, the researcher and the second rater independently, scored two subsets of the data to reach an acceptable interrater agreement level (α=.82). when all valid task-relevant units were scored for their potential effectiveness, an average effectiveness score for a solution was estimated by applying the following formula: average effectiveness = . for each participant (j), the average effectiveness score (y) is the sum total of the effectiveness score (y) for every ith valid task-relevant unit that each participant generated divided by the number of valid units (m) proposed by each participant (j). 2.5.1 reflective task in both experimental conditions participants completed one of two alternative post problem solving tasks. students in the explanation condition (n1=53) were directed to “provide an explanation of their high a.andiliou 104 | f l r school course to the school board members”. the explanation task required a short written response to this prompt. students in the argumentation condition (n2=50) completed an argument diagram. the argumentation diagram is a modified argumentation vee diagram1 (nussbaum, 2008) adapted (a) for the online administration of the study and (b) for constraining participants to use weighing as the integration strategy between the arguments and counterarguments. figure 1 presents the argumentation diagram administered in this study. the overarching question inquired whether the proposed course is a potentially creative course. the participants generated reasons in favour of their creative course and corresponding potential objections of the school board. then they were directed to reread the reasons and objections and decide for each pair whether the reason or objection was stronger. participants could offer up to 5 pairs of reasons and objections. 1 two pilot studies (n1=19; n2=9) and focus group discussions were conducted to ensure that the modified argumentation diagram is comprehensible and that participants are able to complete the diagram. please, contact the first author for information on the pilot studies. a.andiliou 105 | f l r figure 1. the argumentation diagram utilized in the study. 3. results the purpose of the study was to examine problem solving performance and identify reflective tasks that better support students’ self-evaluations of their proposed creative solutions. participants completed online a set of individual difference measures before responding to the problem solving task in which they assumed the role of a high school teacher who was asked to design a creative college preparatory course for the high school senior students. the participants identified the overall goal of their high school course and generated specific learning activities for the course. following they completed an explanation or argument reflective task and rated the creative course in terms of its originality and effectiveness. a summary of the creative solutions that participants forwarded is followed by the presentation of the descriptive statistics and corresponding statistical models performed to answer the two research questions. 3.1 the creative solutions participants designed a creative course to reduce the high college dropout rate among the high school graduates and better prepare them for the transition to college. participants listed and described specific learning activities that they would implement in their course. among the most widely referenced learning activities within the pool of valid task-relevant units generated by the respondents were instructor led activities in which the instructor or another more experienced individual (e.g., guest speaker) was responsible for providing instructional support such as presenting content and sharing experiences. equally popular were activities that were based on experiential learning, such as simulations, fieldtrips, and student presentations. table 3 frequency of occurrence of overarching categories within valid task-relevant units (n=349) overarching category frequency of occurrence f percentage of occurrence % discussion 18 5.16 warm up 4 1.15 instructor led 94 26.93 problem solving 6 1.72 experiential learning 93 26.65 research 13 3.72 writing assignments 58 16.62 reading assignments 3 0.86 classroom assessment 26 7.45 instruction/course design 34 9.74 other learning activities included writing (i.e., expository, persuasive, reflective, organizational aids) and reading assignments (e.g., textbooks, articles, or reports), discussions including student-centered discussions, debates, and discussions with experts. moreover, participants identified research activities, for example searching information about an academic topic, and searching about potential careers and colleges, and learning activities based on problem solving (e.g., decision making). classroom assessments such as formative, summative, and diagnostic assessments were included in participants’ proposed learning activities (n1=26; 7.45%). several students (n2=34; 9.74%) suggested other instructional or course design elements such as materials and educational technologies. it is possible that these students had interpreted the prompt more broadly than intended such that they provided ideas about how they would organize the course and plan instruction to attain the goals of the creative course. a.andiliou 106 | f l r 3.2 predictors of creative solutions in the first research question we examined the extent to which individual difference variables including divergent thinking, need for cognition, beliefs about creative outcomes, and academic major impact the creativity of a proposed solution in terms of its average originality and potential effectiveness. across the sample, the mean average originality score was high m=0.9(0.09) and the mean average effectiveness score was moderate m=3.23(0.59). descriptive statistics for the three continuous predictors are presented in table 4. with respect to their academic major, 57% of the sample were education majors and 43% non-education majors (e.g., communications, kinesiology, or human development). participants exhibited moderate divergent thinking ability and on average they generated six valid ideas. considerable more variability was evident in participants’ need for cognition which was moderate to low. table 4 means and standard deviations of predictors of creative solutions variables m sd range divergent thinking (i) 6.11 2.42 divergent thinking (ii) 5.54 2.23 need for cognition 48.50 10.50 0-90 beliefs 26.92 7.31 0-45 moreover, the students’ beliefs about creative outcomes mean score was m=26.89(7.29), which indicates that participants’ beliefs were somewhat in alignment with conceptualizations in the literature. participants rated high characteristics of creative outcomes pertaining to their effectiveness [i.e., feasible m=3.98(1.24), effective m=3.42(1.04), and successful m=3.23(1.12)]. they also acknowledged as important characteristics those describing the originality of a creative course, for example, innovative m=3.25(1.21) and imaginative m=3.02(1.05). this result suggests that participants took into consideration the context of schooling and appreciated not only the originality but also the effectiveness of a course as an important quality of a creative course. two regression models were conducted to determine the predictors of the average originality and effectiveness of a solution since the correlation between the two outcome variables of average originality and potential effectiveness was non-significant (r=.07, p=.5). due to the violation of the assumption of the residuals, instead of a multiple regression, an ordinal regression model was conducted to determine the predictors of average solution originality. thus, the dependent variable was transformed into an ordinal variable with three levels of average originality (i.e., low, moderate, high) to determine the cumulative odds ratio of proposing a solution of high originality. high average originality (≥.94) was manifested by 52 participants, moderate average originality (.86≤ y ≤.93) was exhibited by 30 participants, and 21 participants scored low (≤.85) in average originality. table 5 predictors of solution originality variable estimate wald p confidence intervals threshold low -.828 0.47 .49 [-3.19, 1.53] moderate .572 0.23 .63 [-1.78, 2.93] parameter divergent thinking (ii) 0.20* 4.82 .03 [.02, 0.38] beliefs 0.01 .03 .47 [-0.05, 0.02] a.andiliou 107 | f l r academic major 0.50 0.02 .90 [-0.73, 0.83] need for cognition -0.01 0.51 .86 [-0.05, 0.02] the initial full ordinal regression model was non-significant. the non-significant predictors were removed stepwise and the ordinal regression model reached significance (-2ll=67.66, χ2(4) =5.08, p=0.02) with divergent thinking (task ii) being the only significant predictor (table 5). divergent thinking positively predicted average solution originality and for each unit increase in divergent thinking participants had lower cumulative odds of developing a solution of poorer originality (low or moderate) by a factor of 0.82. a multiple linear regression with the same individual difference variables as predictors was performed (see table 6) to determine their effect on the average effectiveness of creative solutions. the full model was significant but explained a modest amount of variation [f(4,97)=3.51, p=0.01, r2=0.13]. academic major and need for cognition positively predicted the average effectiveness of a creative solution. for education majors, a solution was on average 0.24 (p=.02) more effective in comparison to a solution proposed by a non-education major. in addition, for each unit of increase in need for cognition there was a 0.21 (p=.03) increase in solution effectiveness. table 6 predictors of solution effectiveness variable b β p confidence intervals constant 2.72** <.001 2.18-3.27 divergent thinking (ii) 0.03 0.12 .22 -0.02-0.07 beliefs -0.01 -0.1 .32 -0.02-0.01 academic major 0.23 0.24 .02* 0.04-0.43 need for cognition 0.01 0.21 .03* 0.001-0.02 3.3 creative solution self-evaluation in the second research question we explored the effect of two alternative reflective tasks: the explanation and the argumentation task on the self-evaluations of the creative solution, with respect to its originality and effectiveness. a multivariate multiple regression (mmr) was performed with four predictors as covariates and the type of reflective task (1=explanation, 2=argumentation) as the fixed factor in the model. the covariates included beliefs about creative outcomes, academic major, and the average assigned originality and effectiveness score. the mmr analysis was conducted since the two outcome variables namely the effectiveness and originality self-evaluations were significantly and positively correlated (r12=.38, p<.001). table 7 predictors of creative solutions self-evaluations hotelling’s trace f p partial η 2 observed power intercept .20 11.67 .001 .20 .99 a.andiliou 108 | f l r beliefs .27 17.31 .001* .27 1.00 average originality .03 1.34 .27 .03 .28 average effectiveness .001 .001 .99 .001 .05 academic major .02 .77 .47 .02 .18 condition .11 5.77 .004* .11 .86 the mmr model was significant [f(2,94) =11.67, p=<.001, η2 =.20]. the type of reflective task [f(2,93) =5.77, p=.004, η2 =.11] and beliefs about creative outcomes [f(2,93) =17.31, p=<.001, η 2 =.27] had a significant effect on the self-evaluations (see table 7). specifically, the type of reflective task significantly and positively predicted the effectiveness self-evaluations [f(1,94) =11.23, p<.001, η 2 =.11]. participants in the argumentation condition evaluated their creative course by 2.81 [p=.001, 95% (1.14, 4.47)] points lower than participants in the explanation condition when all other predictors were equal. thus, the argument diagram was a structure support that seems to have promoted more conservative self-evaluations about the proposed creative solution. participants’ beliefs about the characteristics of creative outcomes were a significant positive predictor of the self-evaluations of originality [f(1,94) =14.63, p<.001, η 2 =.14] and effectiveness [f(1,94) =28.74, p<.001, η 2 =.23] of a forwarded creative solution. in fact, participants whose beliefs better aligned with the current conceptualizations of creative outcomes evaluated higher the creativity of their solution both in terms of its originality and effectiveness. 4. discussion students across education levels are challenged to acquire complex cognitive skills including creative thinking. in this study we examined the individual difference variables that contribute to creative performance in problem solving with respect to the originality and effectiveness of a proposed creative solution. in addition, we attempted to address a gap in the literature related to the effect of argumentation tasks on the self-evaluation of creative solutions. the major contribution of the present study is the development of the creative solution selfevaluation questionnaire which is a reliable rating scale that can be administered by teachers and used by students to evaluate creative solutions, ideas, and products with respect to a set originality and effectiveness criteria. however, the self-evaluation scale needs to be further validated to determine whether it yields the same underlying structure for creative outcomes across fields as it is also possible that additional criteria have to be met for an outcome to be judged as creative in a different field since influential individuals in each field evaluate ideas based on some consensus about the contribution of an idea in the field (antiliou, 2010; csikszentmihalyi, 1999). our investigation also contributes to the research efforts to identify the cognitive and affective variables that predict the creative performance of novices. the findings of the study aligned with research findings regarding the predictors of creative performance. divergent thinking was reported as a predictor of creative problem solving (diakidoy & constantinou, 2001; hunter et al., 2008; reiter-palmon et al., 1997), and this ability to generate various, distinct responses to a divergent thinking task was found in this study to be the single significant predictor of the originality of a proposed creative solution. the effectiveness of a creative solution was predicted by affective and cognitive variables, specifically, need for cognition and academic major. the findings add to the existing evidence, which show that individuals with high need for cognition perform more effectively when solving complex problems (butler et al., 2003; nair & ramnarayan, 2000; osburn & mumford, 2006). however, it is worrisome that participants reported moderate to low need for cognition since this cognitive disposition to enjoy effortful and challenging endeavours represents a prerequisite for lifelong learning and continued professional a.andiliou 109 | f l r development especially for future educators. academic major served as a prior knowledge proxy and it positively predicted the effectiveness of the creative solution. in the future, researchers who aim to examine creative problem solving in specific disciplines could administer measures of domain knowledge such as the pedagogical/psychological (ppk) knowledge measure of general pedagogical knowledge (voss, kunter, baumer, 2011) instead of relying on proxies of prior knowledge, which was a limitation in this study. in the present study we also aimed to investigate the type of tasks that support more reflective selfevaluations of creative solutions. the findings of the study provide some indication that argumentation tasks facilitate more critical self-evaluations of the effectiveness of creative solutions. participants who completed the argument diagram rated the effectiveness of their course more conservatively in comparison to those who responded to the explanation prompt, possibly because the argument diagram provided a structure support for students to elaborate, reflect more deeply and to critically analyze their proposed solution by considering alternative perspectives held by other stakeholders (jonasssen & kim, 2010; nussbaum & sinatra, 2003; suthers, 2001). further, andriessen cited baker (2004) to argue that argumentation is a mechanism through which students not only provide explanations but also prepare a justification to explicitly describe their rationale, which fosters better reflection. munneke (2004) and colleagues also argued that as a knowledge representation tool a diagram explicitly presents the structure of argumentation, thus, providing an overview and making components and perspectives more visible. the fact that participants were more conservative about their solutions after completing the argumentation diagram provides some evidence for voss’s (1981) idea that argumentation is a mechanism that allows students to not only elaborate and clarify the solution but also identify potential limitations, thus, becoming more critical of their solution. however, more research evidence from a study based on a think aloud procedure is needed to provide stronger evidence on how students reflect on their solutions and whether the argument diagram itself promotes more reflective selfevaluations. given that in the present study the design of the argument diagram guided students to apply the weighing argument-counterargument integration strategy, a follow-up study with a think aloud methodology would allow for a more authentic assessment of the integration strategies (e.g., synthesis, refutation, and minimization) that students choose to apply in tasks that require a creative solution that realizes benefits and minimizes disadvantages. the findings of the study confirm the important role of beliefs as an affective variable that impacts problem solving with regard to the self-evaluations of a proposed creative solution rather than on creative performance per se. in fact, participants whose beliefs about the characteristics of creative outcomes aligned better with current conceptualizations in the literature rated both the originality and effectiveness of their solutions more positively. this finding signals the need for educators to pay more attention to affective dimensions of learning including students’ beliefs since they inform critical thinking such as the selfevaluations of solutions. teachers need to provide students with opportunities to explicate, contradict, and enrich their beliefs through classroom discussions, encounters with creative individuals, and exposure to examples of creative work across domains. when practitioners realize that students’ beliefs are narrow or naïve they can also administer rating scales in advance to provide students with criteria for their selfevaluations. the finding also confirms that ontological beliefs about the nature of creative outcomes play an important role in the self-evaluation process. educational researchers have shown interest in examining how epistemological beliefs impact problem solving performance (lodewyk, 2007; muis, 2008; oh & jonassen, 2007) but further research can be conducted in other knowledge domains by using approaches such as think aloud protocols, interviews, and classroom discourse to provide additional evidence on the role of ontological beliefs in creative problem solving in which learners have to draw on their creativity beliefs to define the problem and establish criteria to evaluate a potentially creative solution. thinking as argument is implicated in the beliefs that people hold, the judgments they make, and the conclusions they come to; it arises every time a significant decision must be made (jonassen & kim, 2010, p.439). drawing on the findings of this study, we encourage educators who aim to facilitate students’ critical thinking to use argument-based tasks in the form of diagrams to support students in generating, organizing, and evaluating their arguments and counterarguments in order to make more reflective evaluations during problem solving. a.andiliou 110 | f l r keypoints creative solutions were operationalized as original and effective and innovative procedures were used to measure these dimensions creative outcomes. the theoretical frame emerged from a literature review, which integrate two lines of research namely ill-structured and creative problem solving. the findings confirmed that argumentation diagrams can support reflective critical evaluations beyond writing tasks but in problem solving as well. a reliable self-evaluation scale was developed to assess characteristics of creative solutions which students and teachers can use to evaluate creativity references anderson, l.w., & krathwohl, d.r. (eds.). (2001). a taxonomy of learning, teaching, and assessment: a revision of bloom's taxonomy of educational objectives. new york: longman. andriessen, j. (2006). arguing to learn. in: k. sawyer (ed.) handbook of the learning sciences (pp.443459). cambridge: cambridge university press. andiliou, a. & murphy, p. k. (2010). examining variations among researchers’ and teachers’ conceptualizations of creativity: a review and synthesis of contemporary research, educational research review, 4(3), 201-219. antiliou, a. (2012). the effect of an argumentation diagram on the self-evaluation of a creative solution. (unpublished doctoral dissertation). the pennsylvania state university, university park, pa. basadur, m., runco, m. a., & vega, l. a. (2000). understanding how creative thinking skills, attitudes and behaviors work together: a causal process model. journal of creative behavior, 34(2), 77-100. butler, a. b., scherer, l. l., & reiter-palmon, r. (2003). effects of solution elicitation aids and need for cognition on the generation of solutions to ill-structured problems. creativity research journal, 15(23), 235-244. doi:10.1207/s15326934crj152&3_13. byrne, c. l., shipman, a. s., & mumford, m. d. (2010). the effects of forecasting on creative problemsolving: an experimental study. creativity research journal, 22(2), 119-138. cacioppo, j. t., petty, r. e., & kao, c. f. (1984). the efficient assessment of need for cognition. journal of personality assessment, 48(3), 306-307. chen, c., & bradshaw, a. c. (2007). the effect of web-based question prompts on scaffolding knowledge integration and ill-structured problem solving. journal of research on technology in education, 39(4), 359-375. chi, m.t.h. (2000). self-explaining expository texts: the dual processes of generating inferences and repairing mental models. in r. glaser (ed.), advances in instructional psychology, hillsdale, nj: lawrence erlbaum associates. cho, k., & jonassen, d. h. (2003). the effects of argumentation scaffolds on argumentation and problem solving. educational technology research and development, 50(3), 5-22. christensen, p. r., merrifield, p. r., & guilford, j. p. (1953). consequences form a-1. beverly hills, ca: sheridan supply. csikszentmihalyi, m. (1999) implications of a systems perspective for the study of creativity. in r. j. sternberg, (ed.), handbook of creativity (pp. 313-335). new york: ny cambridge university press. dailey, l.r. & mumford, m.d. (2006). evaluative aspects of creative thought: errors in appraising the implications of new ideas. creativity research journal, 18(3), 367-384. diakidoy, i. n., & constantinou, c. p. (2001). creativity in physics: response fluency and task specificity. creativity research journal special issue: commemorating guilford's 1950 presidential address, 13(3-4), 401-410. eggen, p. & kauchak, d. (2010). educational psychology: windows on classrooms (8 th ed.). new jersey: pearson education. a.andiliou 111 | f l r easterday, m.w., aleven, v., & scheines, r. (2007). tis better to construct or to receive? effect of diagrams on analysis of social policy. in r. luckin, k. r. koedinger, & j. greer (eds.), proceedings of the 13 th international conference on artificial intelligence in education (pp. 93-100). amsterdam: ios. felton, m., & kuhn, d. (2001). the development of argumentative discourse skill. discourse processes, 32(2&3), 135-153. ferretti, r. p., macarthur, c. a., & dowdy, n. s. (2000). the effects of an elaborated goal on the persuasive writing of students with learning disabilities and their normally achieving peers. journal of educational psychology, 92(4), 694-702. fonseca, b. a. & chi, t. h. (2011). instruction based on self-explanation. in r. e. mayer & p. alexander (eds.) handbook of research for learning and instruction. new york, ny: routledge ge, x., chen, c., & davis, k. a. (2005). scaffolding novice instructional designers' problem-solving processes using question prompts in a web-based learning environment. journal of educational computing research, 33(2), 219-248. ge, x., & land, s. m. (2003). scaffolding students' problem-solving processes in an ill-structured task using question prompts and peer interactions. educational technology research and development, 51(1), 2138. doi:10.1007/bf02504515. goldbrick-lab, s., carter f. d., & wagner, r. w. (2007). what higher education has to say about the transition to college? teachers college record, 109(10), 2444-2481. hunter, s. t., bedell-avers, k. e., hunsicker, c. m., mumford, m. d., & ligon, g. s. (2008). applying multiple knowledge structures in creative thought: effects on idea generation and problem-solving. creativity research journal, 20(2), 137-154. isaksen, s. g. & treffinger, d. j. (1985). creative problem solving: the basic course, buffalo, ny: bearly limited. jonassen, d. h. (1997). instructional design models for well-structured and ill-structured problem-solving learning outcomes. educational technology: research & development, 45(1), 65-94. jonassen, d.h., & kim, b. (2010). arguing to learn and learning to argue: design justifications and guidelines. educational technology: research & development, 58, 439-457. kim, s. (2001). the effects of group monitoring on transfer of learning in small group discussions. unpublished doctoral dissertation, university of illinois at urbana-champaign. kuhn, d. (1991). the skills of argument. cambridge, uk: cambridge university press. leitão, s. (2003). evaluating and selecting counterarguments. written communication, 20, 269-306. lodewyk, k. r. (2007). relations among epistemological beliefs, academic achievement, and task performance in secondary school students. educational psychology, 27(3), 307-327. louie, v. (2007). who makes the transition to college? why we should care, what we know and what we need to do. teachers college record, 109(10), 2222-2251. marttunen, m., & laurinen, l. (2006). collaborative learning through argument visualisation in secondary school. in s. n. hogan (ed.), trends in learning research. (pp. 119-138). hauppauge, ny, us: nova science publishers. muis, k. r. (2008). epistemic profiles and self-regulated learning: examining relations in the context of mathematics problem solving. contemporary educational psychology, 33, 177-208. mumford, m.d., & mobely, m. i., uhlman, c. e., reiter-palmon, r., & doares, l. m. (1991). process analytic models of creative thought. creativity research journal, 4, 91-122. munneke, l. van amelsvoort, m., & andriessen, j., (2003). the role of diagrams in collaborative argumentation-based learning. international journal of educational research, 39, 113-131. nair, k. u., & ramnarayan, s. (2000). individual differences in need for cognition and complex problem solving. journal of research in personality, 34(3), 305-328. newell, g. e., beach, r., smith, j., & vanderheide, j. (2011). teaching and learning argumentative reading and writing: a review of research. reading research quarterly, 46(3), 273-304. nussbaum, e. m. (2008). using argumentation vee diagrams (avds) for promoting argumentcounterargument integration in reflective writing. journal of educational psychology, 100(3), 549-565. nussbaum, e. m., & schraw, g. (2007). promoting argument-counterargument integration in students’ writing. journal of experimental education, 76, 59-92. a.andiliou 112 | f l r nussbaum, e. m., & kardash, c. m. (2005). the effects of goal instructions and text on the generation of counterarguments during writing. journal of educational psychology, 97, 157-169. nussbaum, e. m., & sinatra, g. m. (2003). argument and conceptual engagement. contemporary educational psychology, 28, 384-395. doi:10.1016/s0361-476x(02)00038-3 oh, s., & jonassen, d. h. (2007). scaffolding online argumentation during problem solving. journal of computer assisted learning, 23(2), 95-110. osburn, h. k., & mumford, m. d. (2006). creativity and planning: training interventions to develop creative problem-solving skills. creativity research journal, 18(2), 173-190. pritchard, m. e., wilson, g., & yamnitz, b. (2007). what predicts adjustment among college students? a longitudinal panel study. journal of american college health, 56(1), 15-21. reiter-palmon, r., illies, m. y., cross, l. k., buboltz, c., & nimps, t. (2009). creativity and domain specificity: the effect of task type on multiple indexes of creative problem-solving. psychology of aesthetics, creativity, and the arts, 3(2), 73-80. reiter-palmon, r., mumford, m. d., o'connor boes, j., & runco, m. a. (1997). problem construction and creativity: the role of ability, cue consistency and active processing. creativity research journal, 10(1), 9-23. reznitskaya a., anderson, r. c., mcnurlen, b., nguyen-jahiel, k., archondidou, a., & kim, s. y. (2001). influence of oral discussion on written argument. discourse processes, 32(2&3), 155-175. roe clark, m. (2005). negotiation the freshman year: challenges and strategies among first-year college students. journal of college student development, 46(3), 296. runco, m. a., & chand, i. (1994). problem finding, evaluative thinking, and creativity. in m. a. runco (ed.), problem finding, problem solving, and creativity. (pp. 40-76). westport, ct, us: ablex publishing. scwarz, b. b., neuman, y., & biezuner, s. (2000). two wrongs may make a right ...if they argue together! cognition and instruction, 18(4), 461-494. shin, n., jonassen, d. h., & mcgee, s. (2003). predictors of well-structured and ill-structured problem solving in an astronomy simulation. journal of research in science teaching, 40(1), 6-33. suthers, d. d. (2001). towards a systematic study of representational guidance for collaborative learning discourse. journal of universal computer science, 7(3), 254–277. toulmin, s. e. (1958). the uses of argument. cambridge, england: cambridge university press. uribe, d., klein, j. d., & sullivan, h. (2003). the effect of computer-mediated collaborative learning on solving ill-defined problems. educational technology research & development, 51(1), 5-19. van eemeren, f., & grootendorst, r. (1999). developments in argumentation theory. in j. andriessen & p. coirier (eds.). foundations of argumentative text processing (pp. 43-57). amsterdam: amsterdam university press. van eemeren, f. h., & grootendorst, r. (1992). argumentation, communication, and fallacies: a pragmadialectical perspective. hilisdale, nj: lawrence erlbaum associates. voss, j. f. & post, t. a. (1988). on the solving of ill-structured problems. in m. t. h. chi, r. glaser & m. j. farr (eds.), the nature of expertise (pp. 261-285). hillsdale, nj: lawerence erlbaum associates. voss, j. f., wolfe, c. r., lawrence, j. a., & engle, r. a. (1991). from representation to decision: an analysis of problem solving in international relations. in r. j. sternberg, & p. a. frensch (eds.), complex problem solving: principles and mechanisms. (pp. 119-158). hillsdale, nj, england: lawrence erlbaum associates. voss, t., kunter, m., & baumert, j. (2011). assessing teacher candidates’ general pedagogical and psychological knowledge: test construction and validation. journal of educational psychology, 103(4), 952-969. walton, d. (2000). the place of dialogue theory in logic, computer science, and communication studies. synthese, 123, 327-346. walton, d. n. (1996). argumentation schemes for presumptive reasoning. mahwah, nj: laurence erlbaum associates. weisberg, r. w. (2006). expertise and reason in creative thinking: evidence from case studies and the laboratory. in j. c. kaufman & j. baser (eds.) creativity and reason in cognitive development (pp. 742). new york: cambridge university press. a.andiliou 113 | f l r wiley, j., & voss, j. f. (1999). constructing arguments from multiple sources: tasks that promote understanding and not just memory for text. journal of educational psychology, 91(2), 301-311. a.andiliou 114 | f l r appendix effectiveness scoring rubric table 8 effectiveness scoring rubric for the coded task-relevant units score descriptor 4-strong the learning activity or other instruction/course design element  targets important information, knowledge, abilities, skills and strategies for academic success in college or smooth transition to college.  strongly aligns with the overall goal of the course. 3-moderate (weak/strong) or (strong/weak) the learning activity or other instruction/course design element  targets somewhat important information, knowledge, abilities, skills and strategies for academic success in college or smooth transition to college.  strongly aligns with the overall goal of the course. the learning activity or other instruction/course design element  targets important information, knowledge, abilities, skills and strategies for academic success in college or smooth transition to college.  weakly aligns with the overall goal of the course. 2-weak (weak/weak) the learning activity or other instruction/course design element  targets somewhat important information, knowledge, abilities, skills and strategies for academic success in college or smooth transition to college.  weakly aligns with the overall goal of the course. 1-insufficient (weak/inadequate) or (inadequate/weak) the learning activity or other instruction/course design element  targets somewhat important information, knowledge, abilities, skills and strategies for academic success in college or smooth transition to college.  does not align with the overall goal of the course. the learning activity or other instruction/course design element  does not target information, knowledge, abilities, skills and strategies for academic success in college or smooth transition to college.  weakly aligns with the overall goal of the course. 0-inadequate (inadequate/inadequate) the learning activity or other instruction/course design element  does not target information, knowledge, abilities, skills and strategies for academic success in college or smooth transition to college.  does not align with the overall goal of the course. frontline learning research 1 (2013) 81 96 issn 2295-3159 corresponding author: kerry lee, national institute of education, 1 nanyang walk, singapore 637616, kerry.lee@nie.edu.sg, t +65 6219 3888, f +65 6896 9845 http://dx.doi.org/10.14786/flr.v1i1.49 81 | f l r longer bars for bigger numbers? children’s usage and understanding of graphical representations of algebraic problems kerry lee, kiat hui khng, swee fong ng, jeremy ng lan kong national institute of education, nanyang technological university, singapore article received 7 june 2013 / revised 19 july 2013 / accepted 16 august 2013 / available online 27 august 2013 abstract in singapore, primary school students are taught to use bar diagrams to represent known and unknown values in algebraic word problems. however, little is known about students’ understanding of these graphical representations. we investigated whether students use and think of the bar diagrams in a concrete or a more abstract fashion. we also examined whether usage and understanding varied with grade. secondary 2 (n = 68, mage = 13.9 years) and primary 5 students (n = 110, mage = 11.1 years) were administered a production task in which they drew bar diagrams of algebraic word problems with operands of varying magnitude. in the validation task, they were presented with different bar diagrams for the same word problems and were asked to ascertain, and give explanations regarding the accuracy of the diagrams. the küchemann algebra test was administered to the secondary 2 students. students from both grades drew longer bars to represent larger numbers. in contrast, findings from the validation task showed a more abstract appreciation for how the bar diagrams can be used. primary 5 students who showed more abstract appreciations in the validation task were less likely to use the bar diagrams in a concrete fashion in the production task. performance on the küchemann algebra test was unrelated to performance on the production task or the validation task. the findings are discussed in terms of a production deficit, with students exhibiting a more sophisticated understanding of bar diagrams than is demonstrated by their usage. keywords: algebra; pre-algebra; graphical representation; mathematical understanding k. lee et al. 82 | f l r 1. introduction singapore has performed well in recent international tests of mathematics (mullis, martin, gonzalez, & chrostowski, 2004; oecd, 2010). perhaps because of this, there has been much interest in the singapore mathematics curriculum, with some schools in other countries having been reported to have adopted her curriculum (e.g., hu, 2010). the wisdom of such cross-country adoption aside, one peculiar feature of the singapore curriculum is that algebraic thinking is introduced early. unlike many countries where algebra is introduced in the secondary or high school years, mathematical problems with an algebraic structure are taught in the senior primary years (grades 4 – 6). algebra is recognised widely as an important pillar for both academic and economic success (national council for teachers of mathematics, 2000; national mathematics advisory panel, 2008). primary school children have been shown capable of exhibiting algebraic thinking (e.g. carpenter & levi, 2000; carraher, schliemann, brizuela, & earnest, 2006; ng & lee, 2009; swafford & langrall, 2000; warren & cooper, 2005; warren & cooper, 2009). however, whether this understanding is similar to that of older students has not been studied widely. certainly, algebra can be difficult; even for college students. some of the documented difficulties include intrusion of arithmetic reasoning (e.g. khng & lee, 2009; ng, 2003; stacey & macgregor, 1999), difficulties translating word problems to equations (e.g. capraro & joffrion, 2006; duru, 2011; hefferman & koedinger, 1997), problems with the concept of equivalence (e.g. hunter, 2007; kieran, 1981; knuth, stephens, mcneil, & alibali, 2006; steinberg, sleeman, & ktorza, 1991) and poor understanding of the concept of variables (e.g. küchemann, 1978). in this study, we focused on an important pedagogical device for providing children with earlier access to algebraic problems. in the latter part of primary 4 (grade 4, ~10 years old), children in singapore are introduced to algebraic or start-unknown word problems. instead of symbolic algebra, children are taught a graphical heuristic in which they draw bar diagrams to represent known and unknown quantities (ng & lee, 2005). we examined children’s usage and understanding of this heuristic. 1.1 an early start to learning algebra this graphical heuristic, also called the model method, provides students with access to problems that would otherwise require symbolic algebra. three different types of graphical models are commonly taught: (a) part-whole, (b) comparison, and (c) multiplication and division models (ng & lee, 2009). to illustrate its operation, take for example a simple question. mary and john have 6 marbles altogether. john has 2 more marbles than mary. how many marbles does mary have? figure 1 shows the graphical and letter symbolic approach to the problem. in the graphical approach, students draw rectangular bars to represent the number of marbles carried by mary and john. the difference between the two quantities is shown by drawing one bar longer than the other. because the quantitative difference is specified in the question, john’s bar is drawn longer and the quantity represented by the difference in length -the difference unit -is labelled as 2. with this graphical representation, children typically proceed with a variety of arithmetic strategies, such as unwinding or guess-and-check (nathan & koedinger, 2000), to arrive at the solution. k. lee et al. 83 | f l r (a) solution by the model method (b) solution by letter-symbolic algebra 6 – 2 = 4 2 units = 4 1 unit = 2 mary has 2 marbles. let the number of marbles mary has be x. john has x + 2 number of marbles. x + x + 2 = 6 2x = 4 x = 2 mary has 2 marbles. figure 1. model method and the symbolic algebra approach to the question “mary and john have 6 marbles altogether. john has 2 more marbles than mary. how many marbles does mary have?” the bar labelled “2” in (a) represents the difference unit. formal algebraic notations in the form of letter symbols and related expressions were not introduced till some time into grade 6. although both the model method and symbolic algebra require students to translate information in a word problem to an alternative representation, there are some fundamental differences. symbolic algebra requires students to work directly with unknown quantities. an equation comprising both known and unknown values is formed using forward operations and the unknown is solved by constructing a series of equivalent expressions. students operate on the equation in a way that maintains symmetric equivalence across the two sides of the equals sign. in contrast, the graphical approach can trace its roots to pedagogical tools used in the early primary school years. beginning in lower primary, familiar objects and pictures (e.g., pictures of bears and dolls) are used to depict known quantities as an aid to understanding arithmetic word problems. with the graphical approach, a standardized representational tool is used to represent both known and unknown quantities. of course, unknown quantities cannot be depicted exactly. instead, children are taught to draw a diagram with the unknown quantities depicted by bars of arbitrary length, which are constrained by the given quantities and their quantitative relations. computation of solution is effected using arithmetic procedures in which students work only with known quantities. in other words, the unknown is solved by the direct application of arithmetic operations on known values (e.g., backward operations such as unwinding). unlike symbolic algebra, with the graphical approach, the equals sign is generally used as a directive to calculate instead of representing equality (see khng & lee, 2009, for more details). some recent works on the use of the model method were motivated by parents’ and teachers’ concerns that the model method may confuse students when it comes time for them to learn letter symbolic algebra, with some concerned that the two approaches may draw on different cognitive processes (e.g., kwokwc, 2011; lim, 2007). the model method is taught system-wide and has been part of the national curriculum in singapore for over a decade. for this reason, it is difficult to evaluate these claims using standard programme evaluation methodology. in two recent studies, lee and his colleagues used functional magnetic resonance imaging techniques to examine the cognitive underpinnings of these two approaches (lee et al., 2007; lee et al., 2010). the findings showed substantial overlap between the two approaches, but symbolic algebra activated more strongly areas associated with attention and working memory engagement. these findings suggest that symbolic algebra is more demanding of cognitive resources than the model method. from a pedagogical viewpoint, introducing the model method prior to symbolic algebra can thus be interpreted as being consistent with their respective cognitive demands. k. lee et al. 84 | f l r 1.2 the model method, letter symbols, and variables regarding the cognitive factors that influence children’s success in algebra, there has been a number of recent studies on both domain-general and domain-specific correlates of algebraic performance (e.g., fuchs et al., 2012; lee, ng, bull, pe, & ho, 2011; tolar, lederberg, & fletcher, 2009; wei, yuan, chen, & zhou, 2012). on the specific influence of using diagrams, there is a large body of research that examined whether students learn better when text is accompanied by diagrams (mayer, 1989, 2002) or when students generate the diagrams themselves (meter & garner, 2005). of particular relevance are a number of studies conducted by koedinger and his colleagues. they investigated the use of “picture algebra”, a strategy similar to the model method. they found that by using this strategy, even students in grade 6 were successful in solving algebraic problems that are known to be challenging for older students (koedinger & terao, 2002). the strategy was also found to be effective for lower achieving pre-algebra students (booth & koedinger, 2010). although we now have some information on the efficacy of the model method and some of the contexts in which they are more likely to be efficacious, an important issue on which we know little is how students understand or perceive these graphical representations. compared to pedagogical practices in earlier grades in which familiar objects are used to depict operands in arithmetic in a one-to-one manner, the model method involves a greater degree of abstraction. instead of familiar and discrete objects (e.g., bears or dolls), bars of different lengths are used. nonetheless, given students’ earlier experiences, it is possible that they retain a concrete way of thinking about the bars, such that a bar of a certain length is deemed capable of holding, say, ten bears and ten bears alone. in teaching the model method, teachers generally ask children to use relatively longer bars for bigger numbers (ng & lee, 2009). more important than absolute length is that, within a problem, children are taught that the lengths of the bars should preserve the quantitative relations between the protagonists. this is especially when more than two protagonists are involved in a problem. how students understand these graphical representations is important because when they learn symbolic algebra, a concept that they should understand is that letter symbols (e.g., x and y) denote variables. although the model method gives children earlier access to algebraic questions without the use of letter symbols, the bars that are used to depict quantities play a similar role as do letter symbols in algebraic equations. both serve to depict relations between known and unknown quantities. if children regard the bar diagrams in a concrete manner, one potential drawback is that they may over-generalise and regard x and y as depicting unknown constants. a related concern was raised by dede (2004), who argued that the preponderance of questions that use letter-symbols to represent unknowns could result in students developing a restricted view of the roles and functions of variables. the concept of a variable is important in algebra, but it is also difficult, perhaps because it has different meanings or usages. usiskin (1988) argued that variables are used in different ways: (a) as an unknown or a constant, (b) a pattern generalizer, (c) an argument or parameter, or (d) an arbitrary mark on paper. when students are first introduced to letter symbols, they typically encounter them as unknowns in solve-for-x questions, where students are asked to find a solution for the letter symbols (e.g., x + y = 47, x + 13 = y. what is the value of x?). as dede (2004) argued, one concern is that given the preponderance of experiences with this usage of letter symbols, children come to take this as the norm and neglect to entertain other ways in which variables are used. indeed, when asked to simplify algebraic expressions (e.g., 3x + 5x 24), children tended to find a solution for x instead (philipp, 1992). a similar difficulty was reported by akgün and özdemir (2006). in their study, students presented with x + 2 = 2 + x attempted to solve for x when they were told to report all the values that x can assume. kuchëmann (1978) investigated students’ understanding of the use of letter symbols and found most secondary school students treated them as concrete placeholders or “shorthand names” (macgregor & stacey, 1997) (e.g., p in 3p as pears instead of number of pears). only a small number of students displayed an understanding of letters as representing specific unknowns. an even smaller number considered the letters to be generalized numbers. k. lee et al. 85 | f l r 1.3 the present study to understand better the utility of teaching the method model, we focused on students’ understanding of these graphical representations. specifically, we investigated whether students use and think of the bar diagrams in a concrete or a more abstract fashion. we asked children to perform two tasks. in the production task, we examined how they drew model representations for algebraic word problems with operands of varying quantities. we defined concrete usage as varying the length of bars across questions in accordance to the magnitude of operands. abstract usage was indicated by the lack of a consistent relation between the length of the bars and the magnitude of operands. we also manipulated the sequence in which increases in the magnitude of operands were presented. because a sequential increase may overly focus the children’s attention on changes in magnitude and attenuate any tendency to adopt a more abstract strategy, we also presented the increases in a random sequence. in the validation task, we assessed the same children’s understanding by asking them to judge and explain whether several presented graphical representations were drawn correctly. the representations contained drawings that fall into either the concrete or abstract pattern as defined above. children with more sophisticated understanding were expected to know that though bars of particular lengths depict unknown constants within the context of each question, across questions, the same bars can be used to depict different quantities. in fact, it is when letter symbols are considered in this sense that they are considered variables. we tested primary 5 (grade 5) children who have not been taught symbolic algebra and secondary 2 (grade 8) children who have been taught both the model method and symbolic algebra. the primary school children would have used the model method for a year, whereas the secondary school children would have been introduced to symbolic algebra some time during since primary 6. in an interview study conducted with ten primary 5 students, most of the children showed some abstract understanding of how the bar diagrams should be used (ng & lee, 2008). typical of their responses was statements suggesting that the absolute size of the bars do not matter. with their added experience with algebraic problems, we expected the secondary school children to display more abstract usage and awareness. to examine how our measures of children’s understanding were related to other tests of algebraic understanding, we administered the küchemann algebra test (brown, hart, & kuchemann, 1985) to the secondary school children. this is a standardized measure that gauges children’s understanding of variables expressed in the letter symbolic format. it was given only to secondary school children as the primary school children have not been exposed to letter symbols. 2. method 2.1 participants and design the experiment was based on a 4 (magnitude band: one, tens, hundreds, versus thousands) × 2 (question sequence: increasing value versus randomised) × 2 (grade: primary 5 versus secondary 2) full factorial split-plot design. magnitude band served as the only within-subject variable. a total of 68 secondary 2 students (mage =13.9, sd = 0.45) and 110 primary 5 students (mage = 11.1, sd = 0.60) from 5 schools (2 secondary and 3 primary) of mixed abilities and social economic status participated in the study. all schools were government funded, located in the western region of singapore, and followed the national mathematics curriculum. all children participated with parental consent. 2.2 task and materials we designed a web-based program comprising a production task followed by a validation task. participants’ inputs on the computer interface were logged on a server. in addition, the secondary 2 students were administered the küchemann algebra test (brown et al., 1985). k. lee et al. 86 | f l r 2.2.1 production task the children were presented with algebra word problems on the computer screen and were asked to draw model diagrams for these problems using the graphical tools provided onscreen. a computerized interface provided a standardized interface both for problem administration and data collection. the children started with an online palette that contained a variety of specially designed drawing and labelling tools. the children viewed and worked on one problem at a time. they were not required to solve the problem, just to draw the bar diagrams in the same format that they would with pen and paper. there was no limit to the size of the bars that could be drawn. each child completed 12 problems. all the problems were of the same structure, but varied in the name of the protagonists and referents (e.g., mary versus jane, marbles versus cupcakes). the specific magnitude band in each problem varied across four scales: ones, tens, hundreds, and thousands. three problems were given within each scale. the size of the operand relating to the difference unit -our dependent measure of interest -in these three problems was drawn from the smaller, medium, and larger range of each magnitude band; with one question from each range. in figure 1a, for example, we depicted a question with a small difference unit (“2”) from the magnitude band of “ones”. a question with a large difference-unit operand (e.g., 7), still from the magnitude band of ones, would read something like “mary and john have 9 marbles altogether. john has 7 more marbles than mary…” a question with a large difference unit (e.g., 70) from the magnitude band of tens would have been presented as “john has 70 more marbles than mary…” to reduce stimuli specific effects, we developed two parallel sets of problems that were identical in structure and differed only in specific quantities. in both sets, questions were administered either in randomized order or in accordance to the magnitude of the difference-unit operands. the children’s drawings were recorded by the computer program, which also logged the pixel length of the bar diagrams. the pixel length of the difference unit or the section of the bar that represents the difference between the two protagonists (see figure 1a) was used as the dependent variable. of particular interest was whether students varied the length of the difference unit across magnitude bands. that is, did they draw bars that were much longer to represent “john has 70 more marbles” as compared to “john has 7 more marbles”? because we did not impose an upper limit on the length of the bar, there was also no upper cap on the range for the dependent variable. in table 1, we provided both the confidence intervals and the range from the observed data. 2.2.2 validation task participants were presented with two sets of 3 questions. for the first set of validation questions, each question comprised a word problem followed by two model diagrams differing in whether proportionality in the length of the bars was maintained, across questions. in one diagram, the bars were drawn with longer bars for larger numbers. in the other, proportionality was not maintained. in other words, the diagrams differed in whether the bars were drawn in a more concrete or abstract manner. participants were asked to choose which diagram (one or the other, or both) was correct. they were also asked to explain their selection by selecting a response from four multiple choice options (see appendix a). for the second set of questions, each question contained two word problems, each accompanied by two diagrams, said to be drawn by student a and student b respectively. the two word problems in each question were structurally equivalent, differing only in the magnitude of the operands. student a’s diagrams demonstrated a more abstract usage of the bars: the size of the bars were identical for operands of different sizes. student b’s diagrams demonstrated a more concrete usage: the size of the bars differed in accordance to changes in the size of the operands. participants were asked to indicate if the two students were correct. a maximum of 9 marks could be obtained in the validation task (see appendix a for scoring criteria) with higher marks indicating a more abstract understanding of the bar diagrams. k. lee et al. 87 | f l r 2.2.3 küchemann algebra test the küchemann algebra test from the chelsea diagnostic mathematics tests (brown et al., 1985) is a standardized measure of how children interpret letter symbols used in algebra. up to six different common interpretations have been identified in the literature: (i) letter numerically evaluated, (ii) letter not used or ignored, (iii) letter used as an object or abbreviation, (iv) letter used as a specific unknown, (v) letter used as a generalised number, (vi) letter used as a variable. the test places children on one of four levels of understanding based on the type and complexity of their interpretations. for example, a student using any of the first three interpretations and can only answer very simple questions is classified as level 1. a student using the same interpretations, but who is able to answer more structurally complex questions is classified as level 2. understanding letter symbols as referring to specific unknowns qualifies classification at level 3. this is deemed a basic level of understanding required for symbolic algebra. level 4 requires students to demonstrate an understanding of letter symbols as generalised numbers. of interest was whether secondary 2 students’ attainment on the küchemann test correlated with their performance on the production task. that is, did students with higher scores on the küchemann test show less tendency to adjust the length of the bars in accordance to the magnitude of the operands? 2.3 procedure for the computerised tasks, participants from the same school were tested together in a single group session in their school computer laboratories. the secondary 2 students completed the küchemann algebra test in an additional session in a classroom. 3. results to examine whether children’s performances on the production task differed across the various experimental conditions, we subjected the data to a 4 (magnitude band: ones, tens, hundreds, versus thousands) × 2 (question sequence: increasing value versus randomised) × 2 (grade: primary 5 versus secondary 2) repeated measures multivariate analysis of variance. pixel length of the difference unit drawn for questions with smaller, medium, versus larger operands within each magnitude band served as the dependent measures. in addition to the main independent variables, we entered which of the two parallel forms children were administered to take account of potential differences in performance across the two forms. descriptive statistics can be found in table 1. k. lee et al. 88 | f l r table 1 mean pixel length, standard deviation, and confidence intervals for the production task size of difference unit operand small medium large order of presentation/ magnitude band m sd 95% ci m sd 95% ci m sd 95% ci primary 5 increasing (n = 55) ones 33 (19) [28, 38] 51 (29) [43, 59] 111 (56) [95, 127] tens 44 (25) [38, 51] 57 (29) [49, 65] 116 (59) [100, 132] hundreds 50 (28) [42, 57] 60 (28) [52, 68] 122 (68) [103, 140] thousands 54 (35) [44, 63] 59 (27) [52, 66] 106 (62) [89, 123] randomized (n = 55) ones 35 (22) [29, 41] 53 (34) [44, 62] 74 (54) [59, 88] tens 39 (22) [33, 46] 54 (26) [47, 61] 104 (53) [89, 118] hundreds 48 (26) [41, 55] 54 (23) [48, 61] 101 (77) [80, 122] thousands 59 (37) [49, 69] 65 (35) [55, 74] 123 (63) [106, 140] secondary 2 increasing (n = 36) ones 36 (16) [31, 42] 59 (26) [50, 67] 133 (57) [114, 153] tens 47 (28) [38, 56] 61 (25) [53, 70] 124 (51) [106, 142] hundreds 53 (29) [43, 63] 69 (30) [58, 79] 121 (53) [103, 138] thousands 59 (32) [48, 70] 74 (29) [64, 84] 125 (48) [109, 141] randomized (n = 32) ones 36 (18) [30, 43] 59 (29) [49, 70] 111 (70) [85, 136] tens 52 (25) [43, 61] 65 (25) [56, 74] 138 (73) [111, 164] hundreds 53 (24) [45, 62] 72 (27) [62, 82] 134 (84) [103, 164] thousands 70 (38) [56, 84] 81 (39) [67, 95] 150 (71) [125, 176] notes. 1. m and sd refers to the means and standard deviations of the bars drawn for the difference unit (du). 2. 95% ci refers to the 95% confidence interval. it is calculated from the sample mean and is used as an indication of the precision of the estimate. typically, the narrower is the range, the more precise the estimate. 3. values for the m, sd and the 95% ci are rounded to the nearest pixel unit. 4. values for the du range from 5 to 356. length of the difference unit drawn by the students was affected by magnitude band, f(9, 150) = 12.85, p < .001, ηp 2 = .44. as can be noted in both table 1 and figure 2, children generally drew longer bars for larger operands. however, this was qualified by an interaction with question sequence, f(9, 150) = 4.13, p < .01, ηp 2 = .20. univariate tests showed that the interaction effect was not significant for either the smaller or medium sized operands. for these operands, there were significant and strong linear trends across the four magnitude bands, .39 > ηp 2 > .14, regardless of sequence of presentation. in other words, children uniformly used longer bars for small and medium sized operands, regardless of magnitude band. for large operands, a strong linear and increasing trend across the four magnitude bands was found when question sequence was randomised, f(1, 80) = 61.93, p < .01, ηp 2 = .44. there were no differences in the length of the bars, across the four magnitude bands, when questions with large operand sizes were presented as part of a sequence in which the size of operands was ordered (see figure 2). k. lee et al. 89 | f l r figure 2. mean performance on the production task by magnitude band and order of question presentation for smaller, medium and larger numbers within each band. although there was a significant main effect associated with grade, f(3, 156) = 4.05, p < .01, ηp 2 = .07, with the older children drawing longer bars, it did not enter into interaction with other variables. for secondary 2 students, we also tested the relation between their performance on the production task and the küchemann algebra test. as there were only 2 children in the küchemann level 1 category, levels 1 and 2 were combined. approximately 22% of the students attained level 1 – 2, 40% level 3, and 38% level 4. a 4 (magnitude: ones, tens, hundreds, vs. thousands) × 3 (küchemann: level 1 – 2, level 3 vs. level 4) repeated measures analysis of variance showed no significant interaction between magnitude band and level of algebraic understanding on the length of the difference unit drawn by the children. children with more advanced understanding of the use of letters in algebra adjusted the length of bars according to magnitude in a similar manner as did students with a more basic level of algebraic understanding. the students performed well on the validation task with 62% scoring the maximum nine marks. oneway analysis of variance indicated that there were no significant age differences in performance on the validation task. amongst the secondary 2 students, performances also did not differ across levels of understanding on the küchemann algebra test. we also examined whether performance on the production task was related to performance on the validation task. performance on the production task, relative to magnitude band, was indexed by a performance coefficient. this was derived for each individual by fitting a line of best fit using the length of their bar diagrams and the four magnitude conditions as the two axes. a larger performance coefficient indicates a greater propensity to adjust the length of the bars according to magnitude band. there was a small but significant correlation between the performance coefficient and the validation task score but only for the younger age group (r = 0.26, p < .007). primary 5 students with a higher score on the validation task were less likely to adjust the length of the bars according to the magnitude of the operands. 4. discussion findings from the production task showed that children drew longer bars when the magnitude of the operands increased. the only condition in which children did not do this was when they were presented with larger numbers, and only when questions were presented in order of operand magnitude. these findings k. lee et al. 90 | f l r suggest that the children used the bars in a concrete fashion, but their usage is tempered by affordances in the question set. recall that we defined concrete usage as drawing longer bars to denote larger operands, with abstract usage being indicated by the lack of a consistent relation between the length of the bars and magnitude of the operands. children are less likely to engage in a concrete fashion when changes in the magnitude of operands across questions are more salient. findings from the validation task point to a different conclusion. more than half of the children scored full marks and demonstrated awareness that the absolute size of the bars were not essential to the accuracy of the model representations. in contrast to findings from the production task, this finding shows that the majority of children have quite sophisticated understanding of how the graphical representations should be interpreted. indeed, for the younger children, there was a significant correlation between performances on the production and validation tasks. those who showed better understanding on the validation task were more likely to produce models that conformed to our definition of abstract depiction. our data provide no definitive information on why the children’s performance on the production task was less sophisticated than their performance on the validation task. the findings point to another case of children knowing and understanding more than what they can do. this is a common phenomenon in the development of complex skills. in the development of memory strategies, for example, children tend not to deploy skills or strategies spontaneously, but are able to do so successfully when either instructed or are given explicit prompts (flavell, 1970; harnishfeger & bjorklund, 1990). here, both the younger and older children seem to be producing graphical depictions in a concrete manner despite having fairly sophisticated understanding. however, once the manipulation of problem size becomes apparent, they are able to deploy their knowledge accordingly. an alternative version of this explanation, which focuses on the affordances of the production task, is that the students approached the task with what is most familiar. in the production task, the children were not given specific instructions or directives on how the bar diagrams should be drawn. although all the children have had extensive practice with the use of such graphical representations, it is possible that they drew the bars in a more concrete manner because this was what they did, and had to do, with arithmetic problems. findings from the validation task suggest that many children have some appreciation of the fact that the absolute length of the bars across questions is unimportant. although this is just one aspect in understanding the role of variables in algebraic equations, from a pedagogical viewpoint, it is comforting to know that the use of the model method is not overtly associated with erroneous thinking regarding the nature of what is being represented. what is somewhat worrisome is that the secondary school children’s performance on the production task was no different from the primary school children’s. if the younger children’s performance resulted from a production deficiency, the older children, with more experience with such questions, should have been able to use the heuristic in a more abstract fashion. although speculative, one explanation for why this was not observed is related to the way in which algebra is taught. in secondary schools, students are taught symbolic algebra and use letter symbols to represent unknown values. in some schools at least, there is little discussion of differences and similarities in using bar diagrams versus letter symbols to represent known and unknown quantities (ng, lee, ang, & khng, 2006). it is perhaps this lack of explicit linkage that resulted in some lingering confusion, which is reflected in the children’s performance on the production task. nonetheless, given their performance on both the validation task and the küchemann test, their performance on the production task should not be viewed as a major deficit. one pedagogical approach that may further benefit student is to emphasize the different situations under which the two approaches are best suited. although not specifically focused on the model method, previous research have shown that students are more successful when they use more concrete or grounded representations to solve simple algebra questions, but are more successful with more abstract, symbolic representations with more complex problems (koedinger, alibali, & nathan, 2008; koedinger & nathan, 2004). findings from the küchemann algebra test showed that the majority of our secondary 2 students demonstrated a level of understanding that is deemed sufficient to enable them to engage in further studies in algebra. one challenging aspect of the findings is that performance on the küchemann test was not related to performances in either the production or validation tasks. the küchemann test focuses on how children interpret letter symbols in algebra. in contrast, our measures are focused on children’s usage and understanding of the bar diagrams used to represent algebraic questions. although an understanding of the k. lee et al. 91 | f l r use of letter symbols to represent unknowns should help children’s performances in our tasks, one interpretation of the findings is that there is a lack of transfer between understanding the notion of variables when represented as letter symbols versus when represented in the form of bars. an alternative interpretation is that the two types of tasks map onto aspects of algebraic understanding that are more disparate than we anticipated. further research on linkages between these aspects of algebra may help bridge the gap between what is taught in the primary and secondary curricula. 5. conclusion the main aim of this study was to understand how primary and secondary school students use and understand the bar diagrams used for solving algebraic questions. findings from the production task showed that children generally drew longer bars for bigger operands. although this finding shows that the children are using the graphical representations in a more concrete fashion, findings from the validation task suggest that their understanding is more sophisticated. in the validation task, both the younger and the older children demonstrated understanding that the bars can be used in an abstract manner and the length of the bars need not be tied to the size of the operands. the mismatch between findings from the production and validation tasks is interpreted as evidence of a production deficit. although it was surprising that our primary and secondary school students performed in a similar fashion in the production task, when the totality of findings are considered, we do not think the evidence points to a major deficit. despite the production finding, for the secondary students, performances on the production and validation tasks were not correlated. furthermore, the great majority of secondary school students showed quite sophisticated understanding on both the validation task and the küchemann test. on the other hand, for the primary students, the negative correlation between the production and validation tasks suggests that examining the way these students depict problems with different operands does provide some indication of their understanding. primary school teachers may wish to use similar sets of questions, presented in a randomised order, to gain additional insight on students’ facility with algebraic concepts. it should be noted that what is important is not the length of the bars produced for each individual question, but the pattern of responses to changes in the magnitude of operands that provide insight to children’s understanding. discussing how bars of the same length can be used to represent operands of different magnitudes may also help teachers make explicit the conceptual connections between the bars and letter symbols used in algebra. although we were motivated by concerns regarding the role of the model method in the curriculum, this study is not an evaluation of the curriculum, nor did we evaluate whether learning the model method aids in the acquisition of letter symbolic algebra. instead, the findings provide some answers to how children use and understand the model method, which may assist policy makers and curriculum designers when they evaluate its role in the curriculum. to that end, we were encouraged by the sophistication of the children’s responses in the validation task and view these findings as being supportive of the way in which algebraic problem solving is taught. keypoints unlike many other countries, algebraic word problems are introduced in the primary school years in the singapore mathematics curriculum. this study examined how children understand and use bar diagrams that are used to give them earlier access to such problems. both grade 5 and 8 students showed an abstract understanding of the bar diagrams. however, they tended to use the diagrams in a more concrete fashion. discrepancy between what the students produced versus what they understood is indicative of a production deficit. k. lee et al. 92 | f l r discussing how bars of the same length can be used to represent operands of different magnitudes in algebraic questions may help teachers make explicit the conceptual connections between the bar diagrams and letter symbols. acknowledgements the work was supported by grants from the centre for research in pedagogy and practice (#crp 9/05 kl). views expressed in this article do not necessarily reflect those of the national institute of education, singapore. we thank the students who participated in this study and the school administrators who provided access and assistance. appendix a here are some word problems. student a and student b drew the models for these problems. you have to pick if student a is correct, or student b is correct, or if both are correct. you can do this by checking the box next to the options. mary has some marbles. john has 30 marbles more than mary. they have 150 marbles altogether. how many marbles has mary? student a drew the following model. student b drew the following model. who is correct?  student a  student b  both students a and b check the box next to the reason that best explains your choice.  student a is correct because the numbers are small. therefore, the rectangles are short.  student b is correct because the numbers are big. therefore, the rectangles are long.  both students a and b are correct because the size of the rectangle does not matter.  both students a and b are wrong because the models are wrongly drawn. note. one mark was awarded if the student chose “both students a and b” and no marks were given for choosing either “student a” or “student b”. one mark was awarded for choosing the answer “both students a and b are correct because the size of the rectangle does not matter” and no marks were given for any other responses. k. lee et al. 93 | f l r student a and student b drew models for the 2 questions below. q.1 mary has some marbles. john has 30 marbles more than mary. they have 150 marbles altogether. how many marbles has mary? q.2 mary has some marbles. john has 300 marbles more than mary. they have 1500 marbles altogether. how many marbles has mary? student a drew: student b drew: note. one mark was awarded for choosing “yes” for both students and no marks were given for any other responses. references akgün, l., & özdemir, m. e. (2006). students' understanding of the variable as general number and unknown: a case study. the teaching of mathematics, 9(1), 45-51. is student a correct? yes no is student b correct? yes no k. lee et al. 94 | f l r booth, j. l., & koedinger, k. r. (2010). facilitating low-achieving students’ diagram use in algebraic story problems. in s. ohlsson & r. catrambone (eds.), proceedings of the 32nd annual meeting of the cognitive science society (pp. 1649-1654). austin, tx: cognitive science society. brown, m., hart, k., & kuchemann, d. (1985). chelsea diagnostic mathematics tests and teacher's guide. windsor: nfer-nelson publishing company ltd. capraro, m. m., & joffrion, h. (2006). algebraic equations: can middle-school students meaningfully translate from words to mathematical symbols? reading psychology, 27(2-3), 147-164. doi: 10.1080/02702710600642467 carpenter, t. p., & levi, l. (2000). developing conceptions of algebraic reasoning in the primary grades. (res. rep. 00-2). madison, wi: national center for improving student learning and achievement in mathematics and science. retrieved from http://ncisla.wceruw.org/publications/reports/rr-002.pdf carraher, d. w., schliemann, a., brizuela, b. m., & earnest, d. (2006). arithmetic and algebra in early mathematics education. journal for research in mathematics education, 37(2), 87-115. dede, y. (2004). the concept of variable and identification its learning difficulties. educational sciences: theory & practice, 4(1), 50. duru, a. (2011). middle school students’ reading comprehension of mathematical texts and algebraic equations. international journal of mathematical education in science and technology, 42(4), 447468. doi: 10.1080/0020739x.2010.550938 flavell, j. h. (1970). developmental studies of mediated memory. in h. w. reese & l. p. lipsitt (eds.), advances in child development and child behavior (vol. 5). new york: academic press. fuchs, l. s., compton, d. l., fuchs, d., powell, s. r., schumacher, r. f., hamlett, c. l., et al. (2012). contributions of domain-general cognitive resources and different forms of arithmetic development to pre-algebraic knowledge. developmental psychology, 48(5), 1315-1326. doi: 10.1037/a0027475 harnishfeger, k. k., & bjorklund, d. f. (1990). children’s strategies: a brief history. in d. f. bjorklund (eds.), children’s strategies: contemporary views of cognitive development. hillsdale, nj: lawrence erlbaum associates. hefferman, n., & koedinger, k. r. (1997). the composition effect in symbolizing: the role of symbol production versus text comprehension. in m. g. shafto & p. langley (eds.), proceedings of the nineteenth annual conference of the cognitive science society (pp. 307-312). mahwah, nj: lawrence erlbaum associates. hu, w. (2010). making math lessons as easy as 1, pause, 2, pause ... the new york times. retrieved from http://www.nytimes.com/2010/10/01/education/01math.html?_r=0 hunter, j. (2007). relational or calculational thinking: students solving open number equivalence problems. in j. watson & k. beswick (eds.), proceedings of the 30th annual conference of the mathematics education research group of australasia (vol. 2, pp. 421-429). adelaide: merga. khng, k. h., & lee, k. (2009). inhibiting interference from prior knowledge: arithmetic intrusions in algebra word problem solving. learning and individual differences, 19(2), 262-268. doi: 10.1016/j.lindif.2009.01.004 kieran, c. (1981). concepts associated with the equality symbol. educational studies in mathematics, 12(3), 317-326. doi: 10.1007/bf00311062 knuth, e. j., stephens, a. c., mcneil, n. m., & alibali, m. w. (2006). does understanding the equal sign matter? evidence from solving equations. journal for research in mathematics education, 37(4), 297-312. koedinger, k. r., alibali, m. w., & nathan, m. j. (2008). trade-offs between grounded and abstract representations: evidence from algebra problem solving. cognitive science: a multidisciplinary journal, 32(2), 366-397. doi: 10.1080/03640210701863933 koedinger, k. r., & nathan, m. j. (2004). the real story behind story problems: effects of representations on quantitative reasoning. journal of the learning sciences, 13(2), 129-164. koedinger, k. r., & terao, a. (2002). a cognitive task analysis of using pictures to support pre-algebraic reasoning. in c.d. schunn & w. gray (eds.), proceedings of the twenty-fourth annual conference of the cognitive science society (pp. 542-547). mahwah, nj: lawrence erlbaum associates. k. lee et al. 95 | f l r küchemann, d. (1978). children's understanding of numerical variables. mathematics in school, 7(4), 2326. kwokwc. (2011). not able to use algebra in primary level. retrieved from http://www.kiasuparents.com/kiasu/forum/viewtopic.php?f=27&t=18728 lee, k., lim, z. y., yeong, s. h., ng, s. f., venkatraman, v., & chee, m. w. (2007). strategic differences in algebraic problem solving: neuroanatomical correlates. brain research, 1155 (june), 163-171. doi: 10.1016/j.brainres.2007.04.040 lee, k., ng, s. f., bull, r., pe, m. l., & ho, r. h. m. (2011). are patterns important? an investigation of the relationships between proficiencies in patterns, computation, executive functioning, and algebraic word problems. journal of educational psychology, 103(2), 269-281. doi: 10.1037/a0023068 lee, k., yeong, s. h. m., ng, s. f., venkatraman, v., graham, s., & chee, m. w. l. (2010). computing solutions to algebraic problems using a symbolic versus a schematic strategy. zdm, 42(6), 591-605. doi: 10.1007/s11858-010-0265-6 lim, b. t. (2007). can algebra be used to solve psle maths problems, the strait times. retrieved from http://www.moe.gov.sg/media/forum/2007/forum_letters/20070217.pdf macgregor, m., & stacey, k. (1997). students' understanding of algebraic notation: 11-15. educational studies in mathematics, 33(1), 1-19. doi: 10.1023/a:1002970913563 mayer, r. e. (1989). systematic thinking fostered by illustrations in scientific text. journal of educational psychology, 81(2), 240. mayer, r. e. (2002). multimedia learning. psychology of learning and motivation, 41, 85-139. meter, p., & garner, j. (2005). the promise and practice of learner-generated drawing: literature review and synthesis. educational psychology review, 17(4), 285-325. doi: 10.1007/s10648-005-8136-3 mullis, i. v. s., martin, m. o., gonzalez, e. j., & chrostowski, s. j. (2004). timss 2003 international mathematics report: findings from iea's trends in international mathematics and science study at the fourth and eighth grades. chestnut hill, ma: boston college. nathan, m. j., & koedinger, k. r. (2000). teachers' and researchers' beliefs about the development of algebraic reasoning. journal for research in mathematics education, 31(2), 168-190. doi: 10.2307/749750 national council for teachers of mathematics. (2000). principles and standards for school mathematics. reston, va: national council for teachers of mathematics. national mathematics advisory panel. (2008). foundations for success: the final report of the national mathematics advisory panel. washington, dc: u.s. department of education. ng, s. f. (2003). how secondary two express stream students used algebra and the model method to solve problems. the mathematics educator, 7(1), 1-17. ng, s. f., & lee, k. (2005). how primary five pupils use the model method to solve word problems. the mathematics educator, 9(1), 60-84. ng, s. f., & lee, k. (2008). as long as the drawing is logical, size does not matter. the korean journal of thinking & problem solving, 18(1), 67-82. ng, s. f., & lee, k. (2009). model method: singapore children's tool for representing and solving algebra word problems. journal for research in mathematics education, 40(3), 282-313. ng, s. f., lee, k., ang, s. y., & khng, f. (2006). model method: obstacle or bridge to learning symbolic algebra. in w. bokhorst-heng, m. osborne & k. lee (eds.), redesigning pedagogies (pp. 227-242). ny: sense. oecd (2010). pisa 2009 results: executive summary. retrieved from http://www.oecd.org/pisa/pisaproducts/46619703.pdf philipp, r. (1992). the many uses of algebraic variables. mathematics teacher, 85, 557-561. stacey, k., & macgregor, m. (1999). learning the algebraic method of solving problems. the journal of mathematical behavior, 18(2), 149-167. steinberg, r. m., sleeman, d. h., & ktorza, d. (1991). algebra students' knowledge of equivalence of equations. journal for research in mathematics education, 22(2), 112-121. swafford, j. o., & langrall, c. w. (2000). grade 6 students' preinstructional use of equations to describe and represent problem situations. journal for research in mathematics education, 31(1), 89-112. k. lee et al. 96 | f l r tolar, t. d., lederberg, a. r., & fletcher, j. m. (2009). a structural model of algebra achievement: computational fluency and spatial visualisation as mediators of the effect of working memory on algebra achievement. educational psychology: an international journal of experimental educational psychology, 29(2), 239-266. usiskin, z. (1988). conceptions of school algebra and uses of variables. in a. f. coxford & a. p. schulte (eds.), the ideas of algebra (pp. 8-19). reston, va: national council of teachers of mathematics. warren, e., & cooper, t. (2005). introducing functional thinking in year 2: a case study of early algebra teaching. contemporary issues in early childhood, 6(2), 150-162. warren, e., & cooper, t. j. (2009). developing mathematics understanding and abstraction: the case of equivalence in the elementary years. mathematics education research journal, 21(2), 76-95. wei, w., yuan, h. b., chen, c. s., & zhou, x. l. (2012). cognitive correlates of performance in advanced mathematics. british journal of educational psychology, 82(1), 157-181. doi: 10.1111/j.20448279.2011.02049.x microsoft word zander et al_publication.docx frontline learning research vol.3 no. 4 (2015) 1-13 issn 2295-3159 corresponding author: dr. steffi zander, faculty of art and design, chair of instructional design. geschwisterscholl-straße 7, 99423 weimar. germany. phone: +49(0)3643 / 58 32 29, fax: +49(0)3643 / 58 32 48, email: steffi.zander@uni-weimar.de doi: http://dx.doi.org/10.14786/flr.v3i4.161 does personalisation promote learners’ attention? an eye-tracking study steffi zander, maria reichelt, stefanie wetzel, frauke kämmerer, sven bertel bauhaus-universität weimar, germany article received 29 marc 2015 / revised 6 may 2015 / accepted 24 june 2015 / available online 5 november 2015 abstract the personalisation principle is a design recommendation and states that multimedia presentations using personalised language promote learning better than those using formal language (e.g., using ‘your’ instead of ‘the’). it is often assumed that this design recommendation affects motivation and therefore allocation of attention. to gain further insight into the processes underlying personalisation effects we conducted an eye tracking experiment with 37 german university students who were presented with either personalised or formal learning materials. we examined group differences in attention allocation parameters (fixation rate, mean fixation duration, transition count, reading depth). the eye-tracking data was combined with self-reports concerning motivation, cognitive load, and learning outcomes. eye-tracking data revealed a higher reading depth for the main picture areas of interest in the personalised condition. additionally, participants found the personalised version more appealing and inviting. for learning outcomes, there was a positive effect of personalisation. however, after bonferroni correction effects and therefore the pattern expected did not reach significance. the results are discussed in regard to their importance for methodological and practical implications for instructional design. keywords: multimedia learning; personalisation effect; motivation and learning; eyetracking; mixed-methods zander'et 'al ' ' ' | f l r 2 1. personalisation effects in multimedia learning learning of multimedia content can be affected by modest changes to the wording of learning materials. for example, learning is promoted by personalising formal texts, that is, by changing ‘the’ to ‘your’, ‘one’ to ‘you’, and then including direct comments to learners. this approach is known as the personalisation principle, which assumes that multimedia presentations using personalised language promote learning better than those that use formal language. further, based on empirical findings (cf. moreno & mayer, 2000, 2004), even modest changes create personalisation effects. several studies have revealed personalised language effects, mostly for transfer and retention, but also for motivation (interest and intrinsic motivation) and for perceived cognitive load (difficulty and invested mental effort). however, results of existing studies are not consistent with regard to motivation and cognitive load. one reason for this is that the variables underlying personalisation effects do not allow for reasoning based on simple causal chains (ginns, martin, & marsh, 2013). the following chapter reflect the existing theoretical framework of the personalisation principle. 1.1. theoretical framework although personalisation effects are well documented, it remains unclear which processes are responsible for the beneficial effects on learning outcomes. these effects have mainly been investigated using subjective measures at the end of learning phases, whereas objective process-oriented measures have largely been overlooked. in the field of multimedia learning, several theoretical approaches have been proposed to explain personalisation effects (reichelt, kämmerer, niegemann, & zander, 2014). these include social agency theory (mayer, 2005; mayer, 2009), the effect of stronger familiarity (moreno & mayer, 2000a) and the self-reference effect (rogers, kuiper, & kirker, 1977). overall, the assumptions within these approaches can be subsumed into two basic underlying processes: (1) facilitation of cognitive processing and (2) focusing of cognitive processing driven by a personalised language style. ‘facilitation’ and `focusing´ here refer to the notion that personalised messages act as a social cue. in a cognitive view, the cue activates other internal cues that enable learners to more easily connect new information to internal structures of the self, via self-referencing processes (rogers et al., 1977). in a motivational-emotional view, this social cue causes a feeling of social presence or familiarity (mayer, 2005, 2009; mayer & moreno, 2000). in turn, these processes result in a higher intrinsic motivation, situational interest and in decreased perceived cognitive load during learning (moreno & mayer, 2000); they also facilitate the encoding, organisation, and elaboration of relevant information (moreno & mayer, 2000; rogers et al. 1977), ultimately leading to improved learning outcomes. in our current study personalisation effects were investigated using an objective, process-oriented measure (eye-tracking analysis) to test whether the positive effects of a personalised language style can be traced back to differences in the allocation of attention resources and therefore to deeper and more focused processing (due to personalised language style). differences in allocation of attentional resources were measured by eye-tracking parameters including fixation duration, fixation count, and transitions between different sources of information (e.g. text and images). 1.2. current studies previous research confirming the personalisation principle shows that people learn better from multimedia presentations when words are presented in a personalised language style rather than a formal style (mayer, 2009). however, not all studies have demonstrated an effect of personalisation on motivation, cognitive load, and learning outcomes (for an overview see ginns et al., 2013). table 1 gives an overview of research into the personalisation principle, listing authors and results for several dependent variables. the studies reported in table 1 investigated learning outcomes mainly as a result of language style, revealing positive effect of personalised language on transfer performance (except kurt, 2011) and retention (except kurt, 2011; mayer et al., 2004; reichelt et al., 2014; schworm & stiller, 2012). together, these findings zander'et 'al ' ' ' | f l r 3 support the concept of focused processing being driven by personalised messages. the other two variables of interest, motivation (interest) and cognitive load (perceived difficulty of the material and subjective mental effort), have been applied only in select studies. consequently, no empirically proven pattern has been revealed concerning the assumption that these variables facilitate processing. furthermore, in their meta-analysis of these inconsistent findings, ginns et al. (2013) showed a diversity of effects and reported several variables with the potential to moderate the effect of personalised language on motivation, retention, and transfer performance. as a consequence, these authors suggested using more fine-grained methods to analyse the effect. for example, only the study of reichelt et al. (2014) was a mixed-methods-study combining experiments with think-aloud method to examine the personalization effect (see table 1, column 3). however, ginns et al. (2013) as well as reichelt et al. (2014) emphasize the potential of a multi-methods-design to gain more information on underlying processes why personalization effects occur. 1.3. research gaps the overview of measurements applied in personalisation studies (table 1) shows that study results are based mainly on subjective self-report of learners. to date, self-report instruments are the only method that has been used to test the personalisation principle. the application of alternate methods is desirable (ginns et al., 2013) and a potentially useful approach would be to measure the allocation of attentional resources (using eye-tracking methods) during multimedia presentations as this is seen to reveal information regarding the focusing approach. indeed, this method has already shown that eye movements are an indicator of depth and/or direction of information processing in multimedia learning, according to manipulations of visual or audio-visual characteristics of the learning material. for example, de koning, björn b., tabbers, rikers, and paas (2010) investigated cognitive processing during learning of animations containing visual cues, while johnson and mayer (2012) examined the processing of spatially contiguous and non-contiguous textual and pictorial information. these findings, together with those of moreno & mayer (2000), suggest that effective processing across text and pictures might also be promoted by using a personalised style (as a social cue), thus promoting focused information processing. therefore, eye-tracking methodology was applied in the present study to analyse whether personalisation affects the allocation of attentional resources and thus provides support for the personalisation principle. 1.4. aims and research questions of the study the literature review showed that the effect of personalized language on attentional processes were not considered so far. therefore, to fill the research gaps identified above, our eye-tracking study aimed to investigate (1) the impact of personalised language on motivation, cognitive load and learning outcomes, and their possible relation to (2) the processes of allocation of visual attention resources as an indicator of deeper and more focused information processing (driven by personalised messages). we therefore examined the following research questions: (1) does personalised learning material promote learning processes better than formal text versions (in terms of motivation, cognitive load, and learning outcomes)? (2) which attention processes underlie the personalisation effect? based on the literature review (e.g., ginns et al., 2013; kartal, 2010; mayer et al., 2004), we hypothesised that a personalised language style increases learners’ intrinsic motivation, reduces cognitive load, and improves their learning outcomes (hypothesis 1). based on current theoretical models (reichelt et al., 2014, keller, 2009), we further assumed that learners who received a personalised version of a multimedia presentation would allocate their attention resources with more focus on the relevant areas of the learning material than would learners who received a formal version. this difference should be reflected in (a) increased fixation counts (as a parameter for task difficulty), increased duration of fixations (as a parameter for the amount of effort to process complicated texts, rayner & pollatsek, 1989), and a higher number of transitions between text and pictorial information (as a parameter for the process of connecting and integrating information, holsanova, holmberg, & holmqvist, 2009) in the relevant areas for learners who receive personalised presentations (hypothesis 2). zander'et 'al ' ' ' | f l r 4 table 1 overview of results for key personalisation studies author and year of publication transfer retention interest intrinsic motivation mental effort cognitive load task difficulty friendliness helpfulness moreno and mayer (2000a) +1 + moreno and mayer (2004b) + + + + + moreno and mayer (2004b) + -2 03 0 kartal (2010) + + + + + ginns and fraser (2010) 0 + kurt (2011) 0 0 schworm and stiller (2012) + 0 rey and steib (2013) + + 0 reichelt et al. (2014) 0 + + + 1 (+) means that the effect of personalization on this variable (e.g., transfer) was positive 2 (-) means that the effect of personalization on this variable (e.g., transfer) was negative (in favour of formal texts) 3(0) means that were no differences between formal and personalized condition zander'et 'al ' ' ' | f l r 5 2. the eye-tracking study: methods and materials 2.1 participants and design participants were 37 college students (mean age = 25.03, sd = 3.436; male = 21) at bauhausuniversität weimar and the university of erfurt in germany. the participants received either a personalised (n = 19) or a formal (n = 18) version of a computer-based program about typical weather phenomena. to test the influence of domain specific prior knowledge, we used a kruskal-wallis test. the test showed a non-significant result (χ2 = 0.001, p = 0.975); therefore, we assumed an equal distribution of prior knowledge in the experimental groups can be assumed. 2.2 learning material the multimedia learning material consisted of a combination of static pictures and on-screen text, presented on seven slides and with a total duration of approximately 10 minutes. in accordance with mayer (2009), we used various techniques for creating a personalised style. personalisation of the formal text was achieved by replacing impersonal articles with possessive pronouns and third person constructions with second person constructions. only the text was personalised. table 2 shows examples of this manipulation. table 2 examples of personalized and formal text versions formal style personalized style the task is… the picture shows a tropical storm… your task is… you can see a picture of a tropical storm… 2.3 procedure 2.3.1 measurements the pre-test phase consisted of a task description in either formal or personalized language style, a questionnaire on learners’ initial motivational state (qcm, dimension situational interest, rheinberg, vollmeyer, and burns, 2001), and a prior knowledge test. after this, the eye-tracker was adjusted and the learning phase began. after completing the learning phase, participants rated (1) how inviting and personally appealing they perceived the language style (for manipulation check), (2) their intrinsic motivation based on the questionnaire by isen and reeve (2005), and (3) their perceived cognitive load (as a measure of perceived difficulty, koch, seufert, and brünken, 2008). ratings were provided on a 7-point-likert scale (“i disagree” to “i agree”). (4) following this, participants gave responses on the retention and transfer test (learning outcome). the investigations were conducted in a computer laboratory at bauhaus-universität weimar. to capture the eye movements of the participants, we used an sr research eyelink ii head-mounted eye tracker. the participants were placed 55 – 60 cm in front of a 24-inch monitor. fixations, saccades and blinks were recorded at 250 hz for the dominant eye of each participant. a linear drift correction (to the zander'et 'al ' ' ' | f l r 6 screen centre) was implemented after presentation of each stimulus screen. a second calibration was completed before presentation of the learning material (using a 9-point calibration). 2.3.2 data analyses for gaze data analysis, we measured fixation rate, overall fixation duration, and average fixation duration for pre-defined areas of interest on the stimulus screens. to divide the screen into areas of interest (aois), we used analysis based on the expected findings regarding personalisation effects (hypotheses based method, see figure 1). figure 1. predefined aois based on hypotheses. (picture modified for eye-tracking analysis based on the following source: http://www2.klett.de/sixcms/media.php/76/karte_desertifikation.jpg). moreover, to verify the resulting aois cluster analysis (data driven method, see figure 2) was used. figure 2 shows an example screenshot for the formal style (right) and the personalised style (left). both approaches revealed similar partitions of the stimulus screens. however, we decided to analyse our gaze data based on the predefined aois (figure 1) because the granularity level of the cluster analysis was too high. figure 2. screenshots of the extracted aois developed by cluster analysis. zander'et 'al ' ' ' | f l r 7 additionally, we considered fixation transitions between pairs of aois (i.e. a fixation on one aoi followed by a fixation on another aoi, irrespective of transition direction). aoi locations and extent were set based on pre-defined hypotheses (e.g. encompassing all text or an entire diagram). locations and extent were verified for each stimulus screen according to a post-hoc cluster analysis based on the dbscan algorithm (with 4 gaze samples minimum per cluster and 35 pixels maximum distance between gaze samples) and by visual inspection of heat maps; less than 5% of fixations were found to lie outside of the aois. in the first analysis step, we examined the data concerning fixations and transitions for the aois. to accomplish this we analysed fixations and transitions for the picture and text for both the personalized and formal stimulus groups (figure 1, aoi_text and aoi_picture). for subsequent steps of the analysis, the main text and picture aois were further subdivided into smaller aois; this process was hypothesis-driven. 3. results we used mann-whitney u tests for the statistical analysis because prior testing (shapiro-wilk test) revealed that data were not normally distributed. effect size is reported using r. overall we analysed 19 dependent variables (see table 3 and table 4) which increases the type i error. therefore we used bonferroni correction to calculate a new alpha level. we will describe the results considering both alpha levels, namely the corrected (α = 0.002) and the uncorrected (α = 0.05). 3.1. manipulation check comparisons of the participant perceptions of the two language styles (formal and personalised) revealed that participants found the personalised presentation more appealing (u = 101, z = -2.207, p = 0.034, r = -0.363) and inviting (u = 82.5, z = -2.780, p = 0.006, r = -0.457) than the formal presentation. however, the results are not statistically significant after bonferroni correction. 3.2. hypothesis 1: personalized language, motivation, and learning outcomes table 3 provides an overview of medians for each variable test in hypothesis 1. although, the descriptive data show a trend towards the expected effect in favour of the personalised version, the results for the motivational variables showed no significant difference between learners who viewed a personalised text compared with those who viewed a formal version. this non-significant effect was found for both situational interest (u = 138, z = -1.008, p = 0.313, r = -0.166) and intrinsic motivation (u = 142, z = -0.883, p = 0.377, r = -0.145). further, the descriptive analysis confirms the assumption that learners who viewed personalised learning material estimated their cognitive load to be lower than did those who learned with a formal version. however, there were also no significant differences between ratings for learners’ cognitive load (u = 150.5, z = -0.625, p = 0.532, r = -0.103) for formal and personalised presentations. the trend in descriptive data was also reflected in the learning outcome variables. there were differences (α = 0.1, uncorrected) in learners’ retention of personalised vs. formal presentation materials (u = 119, z = -1.654, p = 0.098, r = -0.271), such that participants who viewed a personalised computer-based program showed superior retention compared to those who viewed a formal version. again, after applying bonferroni correction the results are not significant. for the transfer test, the differences were not significant (u = 168.5, z = -0.082, p = 0.935, r = -0.013). zander'et 'al ' ' ' | f l r 8 table 3 medians for all variables for formal and personalised conditions (hypothesis 1) variables and measures formal personalised manipulation check perceived personal appeal of language style perceived inviting character of language style 15.11 14.08 22.68 23.66 motivation initial motivation (before) intrinsic motivation (after) 17.17 17.39 20.74 20.53 cognitive load (based on perceived difficulty) 20.14 17.92 learning outcome retention transfer 16.11 18.86 21.74 19.13 3.3. hypothesis 2: eye-tracking analysis table 4 shows all median data for gaze analysis for both personalised and formal presentation styles. in the first instance the results over all screens are reported. afterwards, data for single screen 3 are presented to accentuate the findings on a more fine-grained level. the eye-tracking analyses revealed several group differences (formal vs. personalisation) in fixation rate and reading depth. the fixation rate was higher in the personalised condition than in the formal condition whereas average fixation duration on the main text aois was greater in the formal condition than in the personalised condition. lower fixation rates indicate greater task difficulty (minoru nakayama, koji takahashi, & yasutaka shimizu, 2002) while greater average fixation duration indicates more effortful cognitive processing (rayner & pollatsek, 1989) that is necessary for more complicated texts. thus, these values indicate that the personalized text was easier to understand and more easily processed by the learners. however, none of the observed differences reached significance, neither using the uncorrected alpha level of 0.05 (there were marginal differences at the 10% level), nor the corrected alpha level of 0.002. for the picture aspect of the stimuli, results were that participants demonstrated greater reading depth for the main picture aois in the personalised condition than in the formal condition. reading depth is defined as the accumulated time spent looking at the aoi divided by the aoi area in cm2. this measure indicates how much of the text has been read or how much of a picture has been examined (holmqvist et al., 2011). the higher value in the personalised condition suggests more intensive observation of the picture than in the formal condition. this is supported by the descriptive finding that the number of transitions between text and picture aois tends to be greater for the personalised learning material. a greater number of transitions between aois with semantic relations indicates better connection and integration of the presented information (holsanova et al., 2009). zander'et 'al ' ' ' | f l r 9 the assumptions are moreover fostered by combining these results with data on retention. therefore, each screen analysis was inspected to verify the findings. comparisons of personalised vs. formal presentations for the individual screens in corresponding pairs revealed similar patterns of differences for the fixation rate on legends, average fixation durations on text aois, reading depth on pictures, and transitions between several map components. many of these differences were statistically significant at the 0.1 alpha level but not at the corrected 0.002 alpha level. for example, reading depth for the picture on screen 3 (figure 2) was higher for personalised than for formal learning material. further, the average fixation duration on text aois differed between personalised and formal presentations. both differences reached significance at the 10% level. the situation is again comparable when comparing aois created by subdividing the main picture aoi into legend and picture proper. in this case, there were greater numbers of transitions between the legend and picture proper and a higher fixation rate on the legend in the personalised condition than in the formal condition. results for the individual screen 3 support the findings for the combined screen analysis, indicating that learners who viewed the personalised learning materials paid more attention to the pictorial material than did those who viewed the formal learning materials. this finding is supported by the retention results and by the finding that learning time for screen 3 differed between formal and personalised presentations, with greater learning time for the personalised presentation on a descriptive level. as an additional measure this finding indicates a deeper processing of information under the personalised condition. those results have to be interpreted and discussed with caution. zander'et 'al ' ' ' | f l r 10 table 4 medians and statistical results for gaze data for two conditions, formal and personalised (hypothesis 2) gaze data parameters formal (median) personalised (median) u z p r gaze data (all screens combined) fixation rate 3.73 4.03 123 -1.66 0.099 -0.27 average fixation duration (ms) on text aoi 176.01 158.44 123 -1.66 0.099 -0.27 average fixation duration (ms) on picture aoi 44.53 42.38 163 -0.497 0.633 -0.08 reading depth (s/cm2) on text aoi 150.72 156.67 160 -0.585 0.573 -0.09 reading depth (s/cm2) on picture aoi 24.11 31.37 109 -2.076 0.038 -0.34 transitions (per s) between text and picture aoi 0.102 0.118 148 -0.936 0.361 -0.15 gaze data (screen 3) reading depth (s/cm2) on text aoi 135.376 141.820 138 -1.228 0.228 -0.20 reading depth (s/cm2) on picture aoi 25.98 38.68 122 -1.696 0.093 -0.28 average fixation duration (ms) on text aoi 178.90 168.00 121 -1.725 0.087 -0.28 average fixation duration (ms) on picture aoi 46.08 46.86 151 -0.848 0.409 -0.14 fixation rate on legend (aoi_legend) 0.23 0.27 119 -1.783 0.077 -0.29 transitions between legend and picture proper (aoi_legend & aoi_map) 0.04 0.09 100.5 -2.324 0.019 -0.38 zander'et 'al ' ' ' | f l r 11 4. discussion the presented study aimed to investigate whether the effect of personalised language style on learning outcomes can be associated with motivational and cognitive load issues and differences in the pattern of attention allocation based on gaze pattern analyses. on a descriptive level, our findings confirm the assumption that personalisation affects learners’ motivation and their perceived cognitive load; however, the results were non-significant. for learning outcomes, there was a positive effect of personalisation for retention but not for transfer. these somewhat inconsistent findings are in accordance with previous studies (e.g. ginns et al., 2013). the inconsistency in findings makes it necessary to further investigate underlying processes. therefore, we conducted more detailed examinations of the allocation of attention resources and other explanatory variables. for the former, the eye-tracking data show the expected pattern of results with a greater number of transitions between main aois in the textual and pictorial information, along with higher fixation duration on the main picture aois. these results provide evidence that mayer’s (2009) proposition that people engage in more focused processing of personalised learning material than of formal material. unexpectedly, in our study, this result was found only for pictorial information but did not extend to textual information. learners of personalised material may pay attention not only to the text but also to the picture, which in turn is reflected in a higher transition count between text and picture and higher fixation count and duration. this assertion is supported by the data for individual screen 3, which revealed a difference with regard to learning time spent on the screen, with longer learning time for the personalised presentation than for the formal presentation. why do these findings occur and what limitations should be considered when interpreting the results? overall, the results of our study confirm that the combination of methods was a fruitful approach for clarifying a very complex set of interacting variables. to add to this approach, we suggest that, in future studies, gaze data should be combined with retrospective interviews (van gog et al. 2005, cued retrospective reporting) while learners view the gaze distributions. this would make it possible to obtain more finegrained information concerning motivation and cognitive load from reflections about the learning process. in the same vein, measuring learning outcomes after each screen presentation should provide a better match between gaze behaviour and learning results. future research could also include the analysis of the sample for any relevant differences, particularly with regard to their educational disciplines. such differences may act as moderator variables (ginns et al., 2013; reichelt et al., 2014). as main limitation, the sample size should be discussed. the sample was small, suggesting a lack of explanatory power, especially with regard to data on learning outcomes, motivation and cognitive load. however, a small sample size is typically for eye-tracking studies (goldberg & wichansky, 2003) and it is justified by the individual surveys. to give generalized statements regarding motivation and learning, a larger sample size is needed. hence, our investigation should be replicated with more participants to increase the power. although our findings suggest support for personalisation theory, the data should be interpreted with caution. in particular, the data on learning outcomes, especially for transfer, frequently did not reach statistically significant levels. several limitations may be responsible for these inconsistencies in our results. to begin with, because our hypotheses included learning material as a whole, we did not focus on pictorial information. to measure learning outcomes with regard to the pictorial information would have required the implementation of explicit pictorial tasks. future studies should contain more detailed pictorial analyses. moreover, with regard to procedures and physical aspects of the study, our eye tracker was head mounted and had to be re-calibrated after every screen presentation. both of these circumstances may have affected the availability of attention resources for learning and understanding. for example, participants had to concentrate on sitting still and were subject to interruptions of the learning process during the recalibrations. future studies should apply less intrusive methods of recording gaze data. zander'et 'al ' ' ' | f l r 12 another possible limitation is linked to the time required for learning. one important and unanswered question is whether the personalisation effect (for fixation, transition, learning outcome) can vanish if the learning time (duration of presentation) increases. this question should be tested in future studies to determine the practical implications for instructional design and especially for the improvement of design principles in multimedia learning environments. keypoints eye-tracking measures can be applied to study the effects of personalisation of learning material on learning outcomes. the combination of eye movement data and self-report reveals that personalised learning material may be processed more deeply than formal material. eye-tracking data suggest that people engage in more focused processing of personalised learning material than formal learning material. references de koning, b. b., tabbers, h. k., rikers, r. m., & paas, f. (2010). attention guidance in learning from a complex animation: seeing is understanding? learning and instruction, 20(2), 111–122. doi: 10.1016/j.learninstruc.2009.02.010 ginns, p., & fraser, j. (2010). personalisation enhances learning anatomy terms. medical teacher, 32(9), 776–778. doi: 10.3109/01421591003692714 ginns, p., martin, a. j., & marsh, h. w. (2013). designing instructional text in a conversational style: a meta-analysis. educational psychology review, 25(4), 445–472. doi: 10.1007/s10648-013-9228-0 goldberg, j.h. & wichansky, a.m. (2003). eye tracking in usability evaluation: a practitioner’s guide. in: radach, ralph, jukka hyona, and heiner deubel, eds. the mind's eye: cognitive and applied aspects of eye movement research (pp. 493-516). elsevier. holmqvist, k., nyström, m., andersson, r., dewhurst, r., jarodzka, h., van de weijer, j. (2011). eye tracking: a comprehensive guide to methods and measures. oxford, new york: oxford university press. holsanova, j., holmberg, n., & holmqvist, k. (2009). reading information graphics: the role of spatial contiguity and dual attentional guidance. applied cognitive psychology, 23(9), 1215–1226. doi: 10.1002/acp.1525 isen, a. m., & reeve, j. (2005). the influence of positive affect on intrinsic and extrinsic motivation: facilitating enjoyment of play, responsible work behavior, and self-control. motivation and emotion, 29(4), 297–325. doi: 10.1007/s11031-006-9019-8 johnson, c. i., & mayer, r. e. (2012). an eye movement analysis of the spatial contiguity effect in multimedia learning. journal of experimental psychology. applied, 18(2), 178–191. doi: 10.1037/a0026923 kartal, g. (2010). does language matter in multimedia learning? personalisation principle revisited. journal of educational psychology, 102(3), 615. koch, b., seufert, t., & brünken, r. (2008). one more expertise reversal effect in an instructional design to foster coherence formation. in j. zumbach, n. schwartz, t. seufert, & l. kester (eds.), beyond knowledge: the legacy of competence (pp. 207–215). dordrecht: springer netherlands. doi: 10.1007/978-1-4020-8827-8_29 kurt, a. a. (2011). personalisation principle in multimedia learning: conversational versus formal style in written word. turkish online journal of educational technology-tojet, 10(3), 185–192. zander'et 'al ' ' ' | f l r 13 mayer, r. e. fennell, s., farmer, l., & campbell, j. (2004). a personalization effect in multimedia learning: students learn better when words are in conversational style rather than formal style. journal of educational psychology, 96, 389-395. mayer, r. e. (2009). multimedia learning (2nd ed.). cambridge, new york: cambridge university press. mayer, r. e. (2005). principles of multimedia learning based on social cues: personalisation, voice, and image principles. in r. e. mayer (ed.), the cambridge handbook of multimedia learning (pp. 201– 212). cambridge: cambridge university press. moreno, r., & mayer, r. e. (2000). engaging students in active learning: the case for personalized multimedia messages. journal of educational psychology, 92(4), 724. moreno, r., & mayer, r. e. (2004). personalized messages that promote science learning in virtual environments. journal of educational psychology, 96(1), 165–173. retrieved from http://www.apa.org/journals nakayama, m., takahashi, k., & shimizu, y. (2002). the act of task difficulty and eye-movement frequency for the 'oculo-motor indices'. in proceedings of the 2002 symposium on eye tracking research & applications (pp. 37–42). new orleans, louisiana: acm. doi: 10.1145/507072.507080 rayner, k., & pollatsek, a. (1989). the psychology of reading: new jersey: lawrence erlbaum associates. reichelt, m., kämmerer, f., niegemann, h. m., & zander, s. (2014). talk to me personally: personalisation of language style in computer-based learning. computers in human behavior, 35, 199–210. doi:10.1016/j.chb.2014.03.005 rey, g. d., & steib, n. (2013). the personalisation effect in multimedia learning: the influence of dialect. computers in human behavior, 29(5), 2022–2028. doi: 10.1016/j.chb.2013.04.003 rheinberg, f., vollmeyer, r., & burns, b. d. (2001). fam: ein fragebogen zur erfassung aktueller motivation in lernund leistungssituationen [qcm: a questionnaire to assess current motivation in learning situations]. diagnostica, 47(2), 57–66. rogers, t. b., kuiper, n. a., & kirker, w. s. (1977). self-reference and the encoding of personal information. journal of personality and social psychology, 35(9), 677–688. doi: 10.1007/bf01172724 schworm, s., & stiller, k. d. (2012). does personalisation matter? the role of social cues in instructional explanations. intelligent decision technologies, 6(2), 105–111. doi: 10.3233/idt-2012-0127 van gog, t., paas, f., & van merriënboer, j. j. g. (2005). uncovering expertise-related differences in troubleshooting performance: combining eye movement and concurrent verbal protocol data. applied cognitive psychology, 19, 205-221. doi: 10.1002/acp.1112 frontline learning research 5 special issue „learning through networks‟ (2014) 38-55 issn 2295-3159 corresponding author: martin rehm, department of educational media & knowledge management, university duisburg-essen, forsthausweg 2, 47057 duisburg, germany, phone: +49-203-3794323, email: martin.rehm@uni-due.de doi: http://dx.doi.org/10.14786/flr.v2i2.85 38 | f l r effects of hierarchical levels on social network structures within communities of learning martin rehm a , wim gijselaers b , mien segers b a university duisburg-essen, germany b maastricht university, the netherlands article received 12 february 2014 / revised 14 may 2014 / accepted 15 may 2014 / available online 15 july 2014 abstract facilitating an interpersonal knowledge transfer among employees constitutes a key building block in setting up organizational training initiatives. with practitioners and researchers looking for innovative training methods, online communities of learning (col) have been promoted as a promising methodology to foster this kind of transfer. however, past research has only provided limited data from actual organizations and largely neglected characteristics that constitute a major obstacle to such collaborative processes, namely participants’ hierarchical levels. the current study addresses these shortcomings by providing empirical evidence from 25 col of an online training program, provided for 249 staff members of a global organization. using social network analysis, we are able to show significant differences in participants’ network behaviour and position based on their hierarchical rank. this translates into higher inand out-degree network ties, as well as centrality scores among participants from higher up the hierarchical ladder. finally, based on a longitudinal analysis of all indicated network measures, our results indicate that the main trend develops predominately during the first half of the training program. by incorporating these insights into the implementation of future col, it is not only possible to anticipate participants’ behaviour. our findings also allow to draw conclusions about how collaborative activities within col should be designed and facilitated, in order to provide participants with a valuable learning experience. keywords: social learning networks; longitudinal analysis; centrality; hierarchical levels m. rehm et al. 39 | f l r 1. introduction researchers have stipulated that organizations are transactive knowledge systems, where the vast majority of knowledge is stored in the heads of individual employees (cross, borgatti, & parker, 2001). consequently, it has been suggested that facilitating an interpersonal knowledge transfer among employees constitutes a key building block in setting up organizational training initiatives (argote & ingram, 2000). this notion is further supported by researchers who suggested that knowledge is being created while collaborating in social networks composed of diverse groups of people (e.g. hakkarainen, palonen, paavola, & lehtinen, 2004; paavola, lipponen, & hakkarainen, 2004). in practice, this process of connecting people greatly builds upon the extensive use of electronic communication tools, such as asynchronous discussion forums. these types of communication channels have been proposed by scholars to effectively enable the establishment and development of new ways in which training can build upon networked communities (e.g. venkatraman, 1994). yet, organizations cannot assume that once a technology is introduced and the appropriate structure has been designed the rest will follow. instead, previous research has established that for social (learning) networks to achieve their intended goals, a clear understanding is needed of how existing organizational structures influence not only the adoption of electronic communication tools, but also their implementation (zack & mckenney, 1995). with practitioners and researchers starting to increasingly look for new approaches to design and implement organizational training programs (yamnill & mclean, 2001), online collaborative learning has received a growing amount of attention in recent years (brower, 2003). in the context of this study, we consider (online) collaborative learning as a setting where “[participants] are working in groups on a shared task or problem, in which they are expected to have equal contributions and participation” (de laat, lally, simons, & wenger, 2006, p. 103). one promising methodology that has been developed within this framework is the concept of online communities of learning (col). being defined as groups of people “engaging in collaborative learning and reflective practice involved in transformative learning” (paloff & pratt, 2003, p. 17), col have been proposed to foster the effective exchange of knowledge and experience between members of an organization‟s workforce (e.g. stacey, smith, & barty, 2004). moreover, online communities, like col, have been considered as an almost ready-made laboratory for analysing collaboration in social (learning) networks over time (haythornthwaite, 2001). in order to conduct these types of analysis, numerous researchers have suggested social network analysis (sna) as a valuable tool for describing and understanding whether and how members of a (learning) network interact with each other (e.g. daradoumis, martínez-monés, & xhafa, 2004; de laat, lally, lipponen, & simons, 2007). according to aviv, erlich, ravid and geva (2003) a social network can be defined as “a group of collaborating (and/or) competing entities that are related to each other” (p. 4). sna has been used to analyse various networks from several academic domains, ranging from social sciences, communication studies, economics, to computer networks and different other fields (aviv et al., 2003). moreover, garton and colleagues (2006) specifically suggest using sna methods in the context of online learning networks. when considering their structure and development, and following the seminal work of erdös and rényi (1960), social networks should evolve according to the concept of random graph theory. in essence, the underlying supposition of this theory is that while some participants of a network might get in touch with more people than others, on average everyone should have made the same amount of contacts, similar to a random distribution of connections. in other words, all participants of a network should have an equal chance of making connections (rienties, tempelaar, giesbers, segers, & gijselaers, 2012). however, if everyone did indeed have equal chances of getting connected with others, why can we then observe so many biased networks in the real world (barabási, 2003)? more specifically, based on numerous studies of newly emerging online communities, researchers have found that a small minority of participants (15%) is gravitating around the centre of their community‟s activity, while a considerable larger group (40%) is barely engaging into communication with their colleagues (e.g. cross, laseter, parker, & velasquez, 2006). in order to explain these observed patterns, some researchers have referred to the fact that communication is an inherently social act (pearce, 1976). new m. rehm et al. 40 | f l r tools and methodologies can only reach their full potential, if organizers fully understand how existing social relationships influence communication patterns and participants‟ behaviour therein (wellman, 2001). moreover, de laat and lally (2003) stipulated that the social and contextual frameworks in which the learning takes place have a considerable influence on how participants behave and perform within online learning networks. furthermore, the nature of social networks, as well as their development over time, is significantly affected by the background characteristics of their individual members (e.g. barabasi & albert, 1999). yet, past research has largely been concerned with the static features of online communities (panzarasa, opsahl, & carley, 2009). while this offers preliminary insights on the overall processes that take place within these communities, it lacks a more refined picture of how social relationships might develop over time (e.g. aviv et al., 2003; haythornthwaite, 2001). additionally, the vast amount of research has neglected a particular background characteristic that can have a severe effect on the underlying learning processes, namely participants‟ hierarchical levels (carley, 1992; griffith & neale, 2001; romme, 1996). the present study addresses these shortcomings by providing empirical evidence from 25 col of an online training program that was provided for 249 staff members of a global organization. each col consisted of 7 – 13 participants and was centred on asynchronous discussion forums, where participants from different parts of the organization‟s hierarchical ladder collaboratively enhanced their knowledge and skills. in order to analyse whether participants‟ network behaviour was influenced by their hierarchical level, social network analysis (sna) was employed. based on the resulting findings of our study, organizers of col will able to anticipate (groups of) individuals holding crucial positions and design actions targeted at participants who tend to be situated more towards the fringe of the network (hatala, 2006). moreover, incorporating our findings into the design and implementation strategies of future col will allow a more refined setup that contributes to employees‟ learning experience and can foster the knowledge creation within an entire organization. 2. effects of hierarchical levels on social network structures within col one of the key elements of online (learning) communities is that they allow for an open dialogue between participants (amin & roberts, 2006). yet, when considering the findings and experiences from real-life communities within organizations, there is increasing evidence that information flows are constrained by underlying organizational structures, such as departments, units and hierarchical levels (e.g. cross, laseter, parker, & velasquez, 2004). one possible explanation for this finding has been put forth by authors like drazin (1990), who stipulated that professionals might not join communities with the intention of learning. instead, individuals would primarily engage into discussion with colleagues, in order to secure their role, and gain access to and control over information. holmqvist (2009) indicated that all organizational learning processes are subject to the influence of a dominant individual or group of individuals. similarly, van der krogt (1998) postulated that “[…] powerful work actors will attempt to influence both the work and the learning network” (p. 170). furthermore, yates and orlikowski (1992) argued that top management will spent more time proactively setting the tone, as they are concerned with losing control of online groups, which could potentially feed through to the real world. considering the role of middle management, bird (1994) advocated that they would act as a “nexus between the real and the ideal” (p. 333). in practice this would result in members of this hierarchical level to “translate” information from one level to the next, providing clarifications and elaborating on shared information. focusing on the lower end of the hierarchical ladder, edmondson (2002) has shown that lower level management is particularly concerned about how colleagues perceive them and their work. consequently, they tend to limit their interaction with colleagues from higher hierarchical levels. additionally, members of this group have been suggested to be more passive in discussions within training programs (nembhard & edmondson, 2006). fox (2000) has described this situation as being “caught in a dilemma” (p.856). on the one hand, individuals would like to establish a reputation of being knowledgeable. on the other hand, they also need to consider the existing rules of conduct. sutton and colleagues (2000) follow this notion and propose that members from lower hierarchical levels will mainly try to blend in while not upsetting the status quo. in practice, this then translates into activities such as flattering, where lower level management frequently contacts their colleagues from higher hierarchical levels (bird, 1994). m. rehm et al. 41 | f l r regarding the overall structure of a (learning) network, it has been established that the position of individuals within such a network is related to their access to valued resources (e.g. ibarra & andrews, 1993; sparrowe, liden, wayne, & kraimer, 2001). casciaro (1998) noted that occupying high-level positions within an organization provides individuals with an intrinsic attraction to lower level management. studying three research centres of an italian university, the author implied that, given their position within the organization, higher level management has privileged access to (vital) information and knowledge sources that are relevant for all employees. moreover, this power can create a type of vortex, where lower level management is trying to get connected and, over time, stay in contact with higher level management (krackhardt, 1990). additionally, borgatti and cross (2003) have argued that lower level management, with only constrained access to valued resources, will be less likely to be contacted for information. as a result, they should hold more peripheral network positions. johnson-cramer, parise and cross (2007) have found empirical evidence for this argument. in their study of a consumer electronic company, they were able to show that higher level management held more central positions in the organization‟s information sharing network. on the contrary, lower level management primarily occupied positions at the outer fringe of the same network. based on these considerations, and taking into the suggestions of previous studies that called for more longitudinal research (e.g. haythornthwaite, 2001), we formulate three research hypotheses: hypothesis 1 (h1): over time, participants' propensity to actively contact other colleagues will be positively influenced by their hierarchical level. hypothesis 2 (h2): over time, participants’ ability to attract connections from other colleagues will be positively related to their hierarchical level. hypothesis 3 (h3): over time, the higher a participant’s hierarchical level, the higher her degree of centrality within col. 3. organisational setting the data was collected from an online training program that aimed at enhancing the capacity and skills of a global organization‟s staff, operating in the sector of economic development. overall, the organization has more than 7.000 employees, operates in 126 countries worldwide, and has its headquarters located in northern america. the training program was delivered twice over a time-span of 14 weeks and covered five pre-defined content modules on the general topic of economics. operating in a fast changing environment, where new analyses and solutions are needed to address old problems, the organization wanted to embrace these developments by training their management staff accordingly. participants engaged into two types of learning activities, namely self-study and collaborative learning. the self-study element included (multimedia) learning materials, such as web lectures and online quizzes. during the collaborative learning activities, which constituted the backbone of the training program, participants discussed real-life tasks via asynchronous discussion forums. the forums were nested in dedicated col that consisted of 10 – 15 randomly assigned participants. each of the five content modules had a separate task, which were discussed within dedicated forums in chronological order. participation in these forums was obligatory and assessed by academic staff members, who facilitated the col. more specifically, a team of two academic staff members was assigned to one col each. these facilitators graded participants‟ contributions, facilitated the discussions, and provided technical assistance. in practice, this could take the form of encouraging discussions and notifying participants when the communication departed too much from the intended focus of the discussion. before engaging with their assigned col, all facilitators were trained on how to work with col and received elaborate guidelines for all collaborative learning activities. additionally, regular meetings were scheduled where facilitators could discuss their experiences and streamline their behaviour and actions towards participants. next to the obligatory, content-driven discussion forums, participants also had the opportunity to exchange private information and socialize via a so-called “café-talk” forum. upon successful completion, participants could attain a certificate of m. rehm et al. 42 | f l r participation, together with academic credits that were based on the european credit transfer and accumulation system (ects). 4. method 4.1 participants and sampling overall, 337 participants were randomly assigned to 30 col. however, the present study analyses a subset of 25 col and 249 participants (73.88%). the underlying reason for this smaller subset is twofold. on the one hand, we had incomplete datasets for some participants. on the other hand, we discovered that some col were biased, in the sense that not all applicable hierarchical levels were represented. consequently, we dropped the applicable col from the analyses. the remaining 25 col had an average of 9.96 members (sd = 1.72, range = 7 – 13), the average age was 43.92 (sd = 7.33, range = 27 – 58), 54.61 percent of the participants were female, and more than 80 nationalities were represented. the educational backgrounds of participants were categorized into master‟s (71.37 %), phd‟s (14.51 %), bachelor‟s (7.26 %), to other degrees (6.85 %). particular examples of the latter category included, health sciences and international law. following the official job categories of the organization in question, participants‟ could be subdivided into “low” (n = 82, 32.93 %), “middle” (n = 93, 37.35 %) and “high” hierarchical levels (n = 74, 29.71 %). 4.2 data collection procedure following the work of daradoumis and colleagues (2004), and based on the collected log-files and user statistics from the underlying discussion forums, we subdivided the data according two different types of network links, namely indirect and direct links. indirect links refer to passive connections that took the form of reading a colleague‟s contributions, but not replying to them. this type of activity was separately recorded in the log-files captured via read-networks. in case a participant actively reacted to another col member‟s contribution and replied, this established a direct link, created another applicable entry in the logfile, and was included in reply-networks. based on this distinction it was then possible to make inferences about the type of learning actions underlying a certain network connection. 4.3 data collection instruments participants reported their own hierarchical level via the training‟s official registration form. the indicated options were subject to the organization‟s official job categories. based on the target group of the training program, three main categories were identified, namely ”low”-, “middle”and “high”-level hierarchical levels. generally, representatives of the “low” group were associated with project level work, contributing to sub-parts of the overall product. members of the “middle” group were leaders of such projects. finally, participants from the “high” group were responsible for departments and often entire regions in which the organization was operating. 4.4 data analysis procedure the analyses of this study focus on data from individual participants. however, these participants were distributed over different col. depending on the specific composition of a particular col, with respect to participants‟ hierarchical levels, this could have led to different dynamics and results. as a result, the validity of comparing across different learning networks might have been reduced. hence, in order to account for possible differences in group compositions across col, we employed the shannon equitability index (magurran, 1988). the index ranges from 0 to fa1 and indicates the percentage share of diversity in relation to the maximal possible diversity within a given col. focusing on participants‟ hierarchical levels as a source of diversity, the average score for the investigated 25 col was .44 (sd = .05, range = .35 – .55). based on this value and the low standard deviation, we concluded that the col represented comparable sample for our analysis. m. rehm et al. 43 | f l r all network statistics were computed with the help of ucinet 6.357 (borgatti, everett, & freeman, 2002). the visualization of an exemplary col network, in terms of sociograms, was conducted with the help of the incorporated visualization software netdraw (borgatti, 2002). the underlying data was based on the log-files and user statistics from the discussion forums within the different col. in order to determine the basic nature of the networks‟ structure, we measured the col network density scores. the density measure is based on the amount of actual ties, divided by the amount possible ties within a col. consequently, it provides an indication of how well-connected participants within a particular col are (hanneman & riddle, 2005). the amount and nature of an individual‟s network connections was determined via the concept of freeman degree centrality, including inand out-degree measures. in-degree network connections indicate how often and by how many colleagues a particular individual was contacted from within a col. more specifically, in the context of the reply-networks, the measure captures how often an individual has been replied to by their colleagues. when considering the read-networks, it reveals how frequent an individual‟s contributions were read by her colleagues. generally, a high amount of in-degree connections has been attributed to prominent participants within (learning) networks, with whom others would like to be connected (hanneman & riddle, 2005). therefore, this constituted our main variable to check our second research hypothesis. the out-degree measure accounts for all those links that originate from a focal individual and summarizes how often that individual contacted her colleagues within the col. when distinguishing between replyand read-networks, the out-degree captures how often a participant has replied to their colleagues and read their contributions, respectively. scholars have often equated a high level of out-degree connections with influential participants, who are able and willing to shape discussions (hanneman & riddle, 2005). consequently, this measure formed the basis for testing the validity of our first research hypothesis. for the analysis of our third research hypothesis, we combined the results of the previous analyses. more specifically, taking into account that we were dealing with multiple col, we determined participants overall centrality on the basis of the normalized number of inand out-degree ties, which allowed to control for the different sizes of the individual col (hanneman & riddle, 2005). in contrast to the more general, nominal network measures, these particular values provided more profound insights on how an individual‟s network ties affected their overall network position within their col. in order to test for the parametric assumption of normality of the data‟s distribution, kolmogorovsmirnov tests (k-s) were conducted. the results revealed a violation of the normality assumption for all measured variables, which translated into statistically significant k-s results at the .01 level. consequently non-parametric tests were used to examine the research hypotheses. more specifically, correlations were determined with the spearman‟s rho measure (rs). in order to assess whether mean differences in the chosen network measures between the different hierarchical levels could be observed, we employed kruskal-wallis tests (h). jonckheere-terpstra tests (j-t) were used to identify whether the potential main effect, as assessed by h, exhibited any possible linear trends. the results of this provided valuable information on how the different hierarchical levels differed in their network measures. the occurrence of possible patterns within the underlying h-test results was determined by post-hoc mann-whitney (u) tests. being designed to only measure differences between two independent conditions, the u-test results were corrected by the bonferroni method. as a result, our adjusted critical value of significance was .016 for this part of the analysis. in order to cater for the longitudinal nature of the data and to test for any possible changes in participants‟ network measures over time, a range of wilcoxon signed rank test were used. the chosen points in time for the longitudinal study were based on the work of previous studies, who conducted similar research on networked learning within teacher education (de laat et al., 2007). the authors of these studies chose for the beginning, the middle and the end phases of online (learning) community. in the context of this study, we decided to subdivide the overall duration of the underlying col of 14 weeks into six time intervals of about two weeks each. this allowed to capture a short “transition period”, during which the focus of the discussions changed from one content module to the next. during this timeframe, participants rounded-up the discussion of the previous module and started preparing for the next one. following the work of de laat and colleagues (2007), out of the six time intervals, we then considered intervals 1 (beginning), 3 (middle) and 6 (end) for our analysis. finally, we also estimated the effect size of our findings. however, the vast majority of effect size measures are only suitable for parametric data (snyder & lawson, 1993). consequently, we followed m. rehm et al. 44 | f l r the suggestion of rosenthal (1991) and approximated the effect size (r) on the basis of the u-results. this measure takes on values from 0 to 1, where small, medium and large effects are associated with .10, .30 and .50, respectively (cohen, 1992). 4.5 control measures although the focus of this research is on the impact of hierarchical levels, we acknowledge that this aspect might only explain parts of possible observed differences between participants. consequently, we controlled for age, gender, educational background, prior knowledge, culture and motivation for attending the training, which have been suggested to influence online collaborative learning. with respect to age, some researchers have suggested that older employees tend to participate less in online training activities (e.g. garavan, carbery, o'malley, & o'donnell, 2010). additionally, other empirical studies have been able to show that age similarity had the potential to trigger emotional conflicts within groups, resulting in lower participation rates (pelled, eisenhardt, & xin, 1999). regarding gender, im and lee (2004) stipulated that if males dominate women in a regular face-to-face environment, this is also likely to carry over to an online environment. in contrast, joinson (2001) was able to show that online training environments had an equalizing effect on participants. when considering participants‟ educational background and prior knowledge, previous studies have highlighted the potential impact of participants‟ prior knowledge on their behaviour within learning initiatives (dochy & mcdowell, 1997). even more so, there has been a growing consensus that individuals‟ prior knowledge constitutes an important variable in participants‟ activity patterns (dochy, segers, & buehl, 1999). if a participant already possesses a considerable amount of prior knowledge about a certain topic, it can be expected that she will be more comfortable in contributing to discussions, thereby positively influencing her general activity and performance levels. participants‟ cultural background has also been suggested to have an impact on participants‟ behave (jehn & bezrukova, 2004). more specifically, researchers like pelled and colleagues (1999) suggested that some cultures tend to exhibit more competitive behaviours than others. hence, representatives of a more competitive culture are also more likely to proactively engage into conversations, trying to shape discussions and thereby achieve higher potential benefits. finally, numerous studies have highlighted the importance of motivation on participants‟ behaviour within the context of online learning (e.g. rienties, tempelaar, van den bossche, gijselaers, & segers, 2009). for example, yang and colleagues (2006) conducted research in online learning environments and discovered that motivation was positively related with how learners perceive each other. consequently, when participants share a similar level of motivation when starting a training program, they tend to “get along” better, which in turn affects their network behaviour (e.g. they connect more often). in this study, participants‟ age, gender, educational background and culture, as assessed by participant‟s country of birth, were self-reported as part of the training programs official registration form. for educational background, participants were asked to indicate their highest attained educational degree, including bachelor, master, phd and other (e.g. vocational training). prior knowledge was measured via a diagnostic test, consisting of 25 multiple choice questions. all five pre-defined content modules were assessed based on five dedicated questions each. these questions were created by academic experts and related to the working environment of the participants. the response rate for the test was 88.76 % and the internal consistency of participants‟ answers was acceptable (cronbach α = .81) (cortina, 1993). participants‟ motivation for attending the training, were approximated based on a previously developed instrument (rienties et al., 2009; rienties, tempelaar, waterval, rehm, & gijselaers, 2006). the questionnaire consisted of 24 questions, subdivided into four categories, and was administered with a 7-point likert scale ranging from 1 (not true for me at all) to 7 (completely true for me). the applicable categories for this study were (the number of questions are reported in brackets): “reasons to join the training” (6), and “expectations and goals” (10). the response rate was 88.51 % and the internal consistency was again acceptable (cronbach α = .95) (cortina, 1993). m. rehm et al. 45 | f l r 5. results overall, while the vast majority of posts were placed in the forums of the five content modules (86%), only few contributions were shared in the “café-talk” forums (14%). in order to visualise the underlying data, figure 1 represents a graphical depiction of the final readand reply-network of an exemplary col. a first glance already indicated a great amount of divergence between these two types of networks. participants were highly connected and exhibited very similar communication patterns with respect to their reading behaviour (fig. 1a). however, considerable differences prevailed regarding whether and how participants replied to each other (fig. 1b). furthermore, a closer look at the figure also revealed a first preliminary sign that participants behaviour and network position were related to their hierarchical level within the organization. an overall picture of the longitudinal nature of our data is depicted in figure 2, which captures the average density values of the col across time. as can be seen from the applicable figure, the average density per time interval of the read-networks is about 10-times higher than those of the reply-networks. yet, while the average density of the read-networks declined over time, the reply-networks increased in terms of density. nonetheless, at the end of the col, the average density for the read-networks remained considerably higher at a value of 62.27 (range = 26.36 – 86.36), as compared to a final value of 11.54 (range = 0 – 28.21) for the reply-networks. a) b) figure 1. read (a) and reply (b) network of an exemplary community of learning. the layout of the figure has been determined using iterative metric multidimensional scaling. the different hierarchical levels are denoted as: “low” – light circle; “middle” – grey square; “high” – dark diamond figure 2. longitudinal data on average density scores for the communities of learning. m. rehm et al. 46 | f l r 5.1 hypotheses 1 & 2 table 1 summarizes the results of participants‟ overall inand out-degree network ties for both types of networks. as can be seen from the table, all measures for the read-networks were statistically insignificant, which led us to reject research hypotheses 1 and 2 for these types of network. in contrast, our kruskal-wallis tests clearly indicated significant differences between hierarchical levels and the degree with which participants‟ either replied to their colleagues, or attracted replies from others. moreover, the jonckheere-terpstra tests showed a clear trend that the amount of both inand out-degree ties were both positively related to participants‟ hierarchical level. additionally, an investigation of the underlying patterns revealed that the observed differences were especially pronounced between the “low” and “high” groups (in-degree: u = 2,261.50, p < .01; out-degree: u = 2,338.00, p < .05), which is also reflected in the observed effect sizes (rin-degree = -.23; rout-degree = -.20). the results of our longitudinal analysis are represented in table 2. as participants‟ behaviour within the read-networks did not show any signs of statistically significant differences, these networks were neglected from the analysis. our results indicated a significant increase of inand out-degree ties for the “middle” and “high” groups over the entire duration of the col. the “low” group did not exhibit a common, noticeable trend. moreover, the evidence indicated that the increases for the “middle” and “high” groups were mainly situated in the first half of the col. during the second half, only members of the “high” group showed significant signs of continued contact-seeking with their colleagues. taken together, these findings indicate that, over time, higher level management was contacted more frequently than lower level management (h1). moreover, our evidence also supported the supposition that over the duration of the col, participants from higher hierarchical levels were more likely to actively contact other col members, than lower level management (h2). table 1 results of kruskal-wallis and jonckheere-terpstra tests for (nominal) inand out-degree network measures m. rehm et al. 47 | f l r table 2 results of wilcoxon signed ranked test for (nominal) inand out-degree measures (reply-networks). 5.2 hypotheses 3 similarly to the previous findings, we again found no significant differences between hierarchical levels within the read-networks. however, as can be seen from table 3, our results for the reply-networks did again sketch another picture. more specifically, the kruskal-wallis tests revealed significant inand outdegree centrality measure differences between hierarchical levels. another set of jonckheere-terpstra tests was then conducted to determine a possible underlying trend. the results showed that whether participants hold a central position within their network was significantly and positively influenced by their hierarchical level. in order to determine the pattern of the main effect, we conducted another range of mann-whitney tests. similarly to hypotheses one and two, the most pronounced difference was again found between the “low” and “high” groups (in-degree: u = 2,202.50, p < .01; rcentrality-in = -.23; out-degree: u = 2,234.50, p < .05; rcentrality-out = -.24). for the longitudinal analysis, based on the described results, we again decided to focus on the replynetworks. table 4 summarizes the main results of the applicable analyses. as in the case of the more general network statistics, we did not find any significant results for the “low” group. in contrast, participants from the “middle” and “high” groups attained higher inand out-degree centrality measures throughout the duration of the col. however, the main acceleration for this development again appeared to be situated in the first half of the col. taking into account that the read-networks did again not yield any significant results, we did not find any support for the notion that, over time, higher level management will hold more central positions in their col network, compared to their colleagues from lower positions (h3). however, based on the statistically significant findings for the reply-networks, we accepted our third research hypothesis for these types of col networks. table 3 results of kruskal-wallis and jonckheere-terpstra tests for (normalized) inand out-degree network measures m. rehm et al. 48 | f l r table 4 results of wilcoxon signed ranked test for (normalized) network measures (reply networks). 5.3 control measures the investigation of whether participants differed in terms of age, gender, educational background, prior knowledge, culture, or motivation for attending the training, subject to their hierarchical levels, revealed no significant results. however, we also conducted a separate correlation analysis, where we investigated any possible, underlying relations between all variables included in this study. as can be seen from table 5, in terms of our dependent and control variables, participants‟ hierarchical level was positively correlated with age. a closer look at the control variables revealed that age (reply-networks: in-degree), gender (read-networks: in-degree) and prior knowledge (read-networks: out-degree) were positively correlated with some of the network measures. hence, in order to incorporate this finding in our analysis, we conducted a separate partial correlation analysis between hierarchical levels and the chosen network measures, while holding age, gender and prior knowledge constant. the results are presented in table 6. while hierarchical levels continued to be significantly correlated with network measures, a more refined picture emerged. more specifically, the potential influence of hierarchical levels now seemed to be mainly applicable for the out-degree measures. moreover, the partial correlation analysis showed this to be true for both the replyand read-networks. consequently, when interpreting the main results of this research, these findings need to be taken into account. moreover, a closer look at the results also revealed that all measured network statistics were highly and significantly correlated with each other. in other words, if an individual participant attained a high amount of in-degree ties, for example, she would also be very likely to initiate a high amount out-degree ties and achieve a comparatively high degree of centrality within her col. as we have been able to show that hierarchical levels have a strong effect on each one of these measures, this provided additional support for our supposition that hierarchical levels have a significant impact on network structures within col. 6. discussion the purpose of this study was to determine whether and to what extend participants‟ hierarchical levels influence the network structures of col. we thereby were able to address a number of shortcomings in current research and contributed to the discussion about how existing organizational structures can affect training initiatives. in order to investigate the relationship between hierarchical levels and network structures, we employed social network analysis and conducted a longitudinal study to test for our research. m. rehm et al. 49 | f l r in the context of the investigated read-networks, we did not find any evidence for individuals‟ hierarchical levels influencing their network behaviour. however, when considering the reply-networks, our results clearly indicated that higher level management attracted more attention, contacted more colleagues, and attained more central positions within their col, as compared to their colleagues from lower level positions. additionally, based on our longitudinal analyses of all network measures, we were able to show that the overall impact generally increased over time, and in particular during the first half of the training program. table 5 overview of correlation coefficients between hierarchical level, control variables and network measures. table 6 correlation coefficients for hierarchical levels and network measures (controlling for age, gender and prior knowledge). m. rehm et al. 50 | f l r in terms of the read-networks, which capture passive connections between participants (daradoumis et al., 2004), this can be considered as a preliminary indication that col have the potential to stimulate an interpersonal knowledge transfer among participants (argote & ingram, 2000). however, the observed range of density scores across the different col varied considerably. moreover, while the average overall density score of 62.27 can be regarded as reasonable, there still remains a considerable gap to be filled in order to achieve a situation where “everyone reads everything”. regarding the reply-networks, we were able to validate our second research hypothesis, which stated that over time, participants‟ ability to attract connections from other colleagues will be positively related to their hierarchical level (h2). this supports the work of krackhardt (1990), who suggested the existence of a vortex that allows higher level management to attract more attention and connections from their colleagues. additionally, our evidence suggested that higher level management will proactively set the tone in online discussions (h1), which confirms the work of yates and orlikowski (1992). we were also able to show that higher level management held central positions, while lower level management was located more towards the fringe of their col (h3) (borgatti & cross, 2003). finally, when conducting longitudinal analyses of the underlying data, our results indicated that the observed general patterns increased over the duration of the col (e.g. bird, 1994; sutton et al., 2000). additionally, this positive trend was particularly pronounced during the first half of the training program, which appears as a kind of “initiation phase”. however, we also discovered that this trend was not statistically significant for the “low” group. this finding can be considered as support for the work of nembhard and edmondson (2006), who suggested that members of this group generally tend to be more passive in discussions within training programs. additionally, it could also be attributed to the importance of the “initiation phase”. once members from the “middle” and “high” group have established their comparatively more central role within their col, it seems as if the “low” group is content with the situation. alternatively, it could also be that members of the “high” group convey such an “imposing message”, trying to lead the group and becoming (more) central to the discussions, that representatives of the “low” group rather not change their behaviour and become more active. furthermore, when reinvestigating the potential influence of hierarchical levels on the chosen network measures, while incorporating our control variables, an even more refined picture emerged. our results indicate that age, gender and prior knowledge seem to have a mediating role in determining participants‟ network measures. more specifically, participants‟ hierarchical background mainly affected their out-degree behaviour, e.g. the degree with which they reply to colleagues in discussions. additionally, this effect was applicable for both the replyand the read-networks, which suggests two main conclusions for higher level management. first, members of this group really try to set the tone and actively try to shape the discussions. second, higher level management more carefully followed the discussions by reading the contributions of their colleagues from lower hierarchical levels. considering these findings, we can draw conclusions about how collaborative learning activities within col should be designed and facilitated, in order to provide participants with a valuable learning experience. for example, acknowledging the considerable influence of hierarchical levels, organizers can device targeted interventions that increase the potential benefits of col (cross et al., 2006). more specifically, higher level management could be stimulated to actively draw upon the input of their colleagues, thereby allowing participants from lower level management to gradually move towards the centre of the col network. in practice, this could be achieved via two possible approaches. on the one hand, facilitators could try to foster a (more) active exchange of information between members of these two opposite parts of the organization. the potential benefit of this approach would be that connections between participants would be initiated and supported by an external party. this in turn could relax underlying norms and regulation that govern how members from different hierarchical levels communicate with each other. alternatively, participants could be asked to complete assignments that build upon a type of mentoring system. with higher level management occupying more central positions, these participants could take their colleagues from lower hierarchical levels “by the hand” and actively include them in the discussions. this could create a pull-effect, whereby participants, who generally tend to occupy positions towards the fringe of a learning network, are drawn closer towards the centre. this not only has the potential to make them a more integral part of the col. it also would provide them with better opportunities to share their knowledge and m. rehm et al. 51 | f l r insights. using the analogy of kozlowski and colleagues (2009), they could thereby more easily contribute their piece to the puzzle, which can enhance the success of the entire organization. finally, considering the longitudinal findings of our research, we have highlighted the importance of the “initiation phase” within col. during the beginning stages of the learning process, participants get to know each other‟s background characteristics, including professional experience and prior knowledge. additionally, participants will also exchange either directly (as part of their introduction to the col), or indirectly (by making appropriate references) information about their hierarchical levels. this in turn will significantly influence their behaviour towards each other throughout the col. consequently, facilitators of such communities should pay specific attention to this initiation process, in order to be able to possibly intervene in the discussions and assist the central participants to engage the entire group into the discussions. 7. conclusions 7.1 limitations the current study exhibits two main limitations that should be taken into account when considering our results. first, we have based our social network statistics purely on observed links between participants. in contrast, previous studies have also commonly incorporated familiarity measures in the context of social network analysis (e.g. krackhardt, 1990). these measures allow to control for the degree with which participants might already be acquainted with each other. this in turn could influence the comfort level of participants‟ and thereby affect their behaviour within col. second, connections between participants did not take into account the content of the shared information. consequently, network ties between individual participants might have reflected personal commonalities that have no direct link with the actual content of the training and are therefore difficult to control for by organizers of similar initiatives. 7.2 future research building upon the findings of this study, future research should conduct (hierarchical) multilevel regression modelling (goldstein, 1995). our results indicate that age, gender, and prior knowledge also had an effect on participants‟ network behaviour. consequently, in order to incorporate these findings and to further contribute to our understanding of whether and how hierarchical levels are transferred into the network structures of col, future studies should consider modelling a larger set of explanatory variables simultaneously. moreover, future research should conduct a content analysis (ca) of the underlying discussions forums within col. this approach is widely accepted to assess the quality of learning processes and outcomes (de laat & lally, 2003) and allows to draw a more refined picture of the actual level of content and knowledge that has been exchanged between participants. moreover, by mapping the ca results against the findings of a sna analysis, it would be possible to provide detailed insights about who has been in contact with whom, what they talked about, and whether this has had an impact on their network position (de laat et al., 2007). additionally, future research should incorporate the role of facilitators into the analysis of col. previous research has suggested that online learning communities must be cherished and protected in order to become an effective educational resource (paloff & pratt, 2003). in other words, facilitators‟ involvement can have a considerable influence on how learning networks develop and evolve over time (anderson, rourke, garrison, & archer, 2001). yet, although a considerable amount of research has already investigated how online facilitation can affect learning processes, the vast majority of these studies has focused on the context of higher education (berge, 1995; de laat et al., 2006; garrison, anderson, & archer, 2010) and largely neglected the field of training within organizations. by investigating the role of facilitators in col, it would be possible to provide profound insights that can serve as a springboard for facilitators to design and implement an effective teaching strategy for col. consequently, the quality of learning process could be further augmented. m. rehm et al. 52 | f l r keypoints we assess the impact of hierarchical levels on online learning networks. the higher the hierarchical level, the higher the connectedness of participants. the higher the hierarchical level, the higher the centrality of participants. our findings are particularly strong for the first half of the networks‟ duration. references amin, a., & roberts, j. (2006). communities of practice: varieties of situated learning‟. eu network of excellence dynamics of institutions and markets in europe (dime). http://www.dimeeu.org/files/active/0/amin_roberts.pdf anderson, t., rourke, l., garrison, d. r., & archer, w. (2001). assessing teaching presence in a computer conferencing context. journal of asynchronous learning networks, 5(2), 1-17. argote, l., & ingram, p. (2000). knowledge transfer: a basis for competitive advantage in firms. organizational behavior and human decision processes, 82(1), 150-169. doi: 10.1006/obhd.2000.2893 aviv, r., erlich, z., ravid, g., & geva, a. (2003). network analysis of knowledge construction in asynchronous learning networks. journal of asynchronous learning networks, 7(3), pp. 1-23. barabási, a.-l. (2003). linked : how everything is connected to everything else and what it means for business, science, and everyday life. from http://quijote.biblio.iteso.mx/dc/ver.aspx?ns=000149871 barabasi, a. l., & albert, r. (1999). emergence of scaling in random networks. science, 286(5439), 509512. berge, z. l. (1995). facilitating computer conferencing: recommendations from the field. educational technology, 15(1), 22-30. bird, a. (1994). careers as repositories of knowledge a new perspective on boundaryless careers. journal of organizational behavior, 15(4), 325-344. doi: 10.1002/job.4030150404 borgatti, s. p. (2002). netdraw: graph visualization software. harvard, ma: analytic technologies. borgatti, s. p., & cross, r. (2003). a relational view of information seeking and learning in social networks. management science, 49(4), 432–445. borgatti, s. p., everett, m. g., & freeman, l. c. (2002). ucinet for windows: software for social network analysis. harvard, ma: analytic technologies. brower, h. h. (2003). on emulating classroom discussion in a distance-delivered obhr course: creating an on-line learning community. academy of management learning & education, 2(1), 22-36. carley, k. (1992). orgabnizational learning and personnel turnover. organization science, 3(1), 20-46. casciaro, t. (1998). seeing things clearly: social structure, personality, and accuracy in social network perception. social networks, 20(4), 331-351. doi: 10.1016/s0378-8733(98)00008-2 cohen, j. (1992). statistics a power primer. psychology bulletin, 112, 155–159. cortina, j. m. (1993). what is coefficient alpha? an examination of theory and applications. journal of applied psychology, 78(1), 98-104. cross, r., borgatti, s. p., & parker, a. (2001). beyond answers: dimensions of the advice network. social networks, 23(3), 215-235. doi: 10.1016/s0378-8733(01)00041-7 cross, r., laseter, t., parker, a., & velasquez, g. (2004). assessing and improving communities of practice with organizational network analysis. paper presented at the the network roundtable at the university of virginia, virginia. cross, r., laseter, t., parker, a., & velasquez, g. (2006). using social network analysis to improve communities of practice. california management review, 49(1), 32-60. daradoumis, t., martínez-monés, a., & xhafa, f. (2004). an integrated approach for analysing and assessing the performance of virtual learning groups groupware: design, implementation and use (vol. 3198, pp. 289-304): springer berlin / heidelberg. http://www.dime-eu.org/files/active/0/amin_roberts.pdf http://www.dime-eu.org/files/active/0/amin_roberts.pdf http://quijote.biblio.iteso.mx/dc/ver.aspx?ns=000149871 m. rehm et al. 53 | f l r de laat, m., & lally, v. (2003). complexity, theory and praxis: researching collaborative learning and tutoring processes in a networked learning community. instructional science, 31(1-2), 7-39. de laat, m., lally, v., lipponen, l., & simons, r.-j. (2007). investigating patterns of interaction in networked learning and computer-supported collaborative learing: a role for social network analysis. computer-supported collaborative learning, 2, 87-103. de laat, m., lally, v., simons, r.-j., & wenger, e. (2006). a selective analysis of empirical findings in networked learning research in higher education: questing for coherence. educational research review, 1(2), 99-111. dochy, f., & mcdowell, l. (1997). assessment as a tool for learning. studies in educational evaluation, 23(4), 279-298. dochy, f., segers, m., & buehl, m. m. (1999). the relation between assessment practices and outcomes of studies: the case of research on prior knowledge. review of educational research, 69(2), 145-186. doi: 10.3102/00346543069002145 drazin, r. (1990). professionals and innovation structural functional versus radical structural perspectives. journal of management studies, 27(3), 245-263. doi: 10.1111/j.14676486.1990.tb00246.x edmondson, a. c. (2002). the local and variegated nature of learning in organizations: a group-level perspective. organization science, 13(2), 128-146. doi: 10.1287/orsc.13.2.128.530 erdös, p., & rényi, a. (1960). on the evolution of random graphs. publications of the mathematical institute of the hungarian academy of sciences, 5, 17-61. fox, s. (2000). communities of practice, focault and actor-network theory. journal of management studies, 37(6), 853-867. garavan, t. n., carbery, r., o'malley, g., & o'donnell, d. (2010). understanding participation in elearning in organizations: a large-scale empirical study of employees. international journal of training and development, 14(3), 155-168. doi: 10.1111/j.1468-2419.2010.00349.x garrison, d. r., anderson, t., & archer, w. (2010). the first decade of the community of inquiry framework: a retrospective. internet and higher education, 13(1-2), 5-9. doi: 10.1016/j.iheduc.2009.10.003 garton, l., haythornthwaite, c., & wellman, b. (2006). studying online social networks. journal of computer-mediated communication, 3(1), 0-0. doi: 10.1111/j.1083-6101.1997.tb00062.x goldstein, h. (1995). multilevel statistical models. sydney: edward arnold. griffith, t. l., & neale, m. a. (2001). information processing in traditional, hybrid, and virtual teams: from nascent knowledge to transactive memory. research in organizational behavior, vol 23, 23, 379421. doi: 10.1016/s0191-3085(01)23009-3 hakkarainen, k., palonen, t., paavola, s., & lehtinen, e. (2004). communities of networked expertise: professional and educational perspectives. amsterdam: elsevier. hanneman, r. a., & riddle, m. (2005). introduction to social network methods. riverside, ca: university of california. hatala, j. p. (2006). social network analysis in human resource development: a new methodology. human resource development review, 5(1), 45-71. doi: 10.1177/1534484305284318 haythornthwaite, c. (2001). exploring multiplexity: social network structures in a computer-supported distance learning class. information society, 17(3), 211-226. holmqvist, m. (2009). complicating the organization: a new prescription for the learning organization? management learning, 40(3), 275-287. doi: 10.1177/1350507609104340 ibarra, h., & andrews, s. b. (1993). power, social-influence, and sense making effects of network centrality and proximity on employee perceptions. administrative science quarterly, 38(2), 277303. doi: 10.2307/2393414 im, y., & lee, o. (2004). pedagogical implications of online discussion for preservice teacher training. journal of research on technology in education, 36(2), 155-170. jehn, k. a., & bezrukova, k. (2004). a field study of group diversity, workgroup context, and performance. journal of organizational behavior, 25(6), 703-729. johnson-cramer, m. e., parise, s., & cross, r. l. (2007). managing change through networks and values. california management review, 49(3), 85-109. m. rehm et al. 54 | f l r joinson, a. n. (2001). self-disclosure in computer-mediated communication: the role of self-awareness and visual anonymity. european journal of social psychology, 31(2), 177-192. doi: 10.1002/ejsp.36 kozlowski, s. w. j., chao, g. t., & jensen, j. m. (2009). building an infrastructure for organizational learning: a multilevel approach. in e. salas & s. w. j. kozlowski (eds.), learning, training, and development in organizations. new york, ny, united states of america: routledge. krackhardt, d. (1990). assessing the political landscape structure, cognition, and power in organizations. administrative science quarterly, 35(2), 342-369. doi: 10.2307/2393394 magurran, a. e. (1988). ecological diversity and its measurement. princeton, nj, usa: princeton university press. nembhard, i. m., & edmondson, a. c. (2006). making it safe: the effects of leader inclusiveness and professional status on psychological safety and improvement efforts in health care teams. journal of organizational behavior, 27(7), 941-966. doi: 10.1002/job.413 paavola, s., lipponen, l., & hakkarainen, k. (2004). models of innovative knowledge communities and three metaphors of learning. review of educational research, 74(4), 557-576. doi: 10.3102/00346543074004557 paloff, r., & pratt, k. (2003). the virtual student: a profile and guide to working with online learners. san francisco: jossey-bass. panzarasa, p., opsahl, t., & carley, k. m. (2009). patterns and dynamics of users' behavior and interaction: network analysis of an online community. journal of the american society for information science and technology, 60(5), 911-932. doi: 10.1002/asi.21015 pearce, w. b. (1976). the coordinated management of meaning: a rules-based theory of interpersonal communication. in g. r. miller (ed.), explorations in interpersonal communication (pp. 17-35). beverly hills, ca: sage publications. pelled, l. h., eisenhardt, k. m., & xin, k. r. (1999). exploring the black box: an analysis of work group diversity, conflict, and performance. administrative science quarterly, 44(1), 1-28. doi: 10.2307/2667029 rienties, b., tempelaar, d., giesbers, b., segers, m., & gijselaers, w. (2012). a dynamic analysis of social interaction in computer mediated communication: a preference for autonomous learning. interactive learning environments. rienties, b., tempelaar, d., van den bossche, p., gijselaers, w., & segers, m. (2009). the role of academic motivation in computer-supported collaborative learning. computers in human behavior, 25(6), 1195-1206. doi: 10.1016/j.chb.2009.05.012 rienties, b., tempelaar, d., waterval, d., rehm, m., & gijselaers, w. (2006). remedial online teaching on a summer course. industry and higher education, 20(5), 327-336. romme, a. g. l. (1996). a note on the hierarchy-team debate. strategic management journal, 17(5), 411417. rosenthal, r. (1991). meta-analytic procedures for social research. newbury park, ca: sage. snyder, p., & lawson, s. (1993). evaluating results using corrected and uncorrected effect size estimates. journal of experimental education, 61(4), 334-349. sparrowe, r. t., liden, r. c., wayne, s. j., & kraimer, m. l. (2001). social networks and the performance of individuals and groups. academy of management journal, 44(2), 316-325. doi: 10.2307/3069458 sutton, r., neale, m. a., & owens, d. (2000). technologies of status negotiation: status dynamics in email discussion groups: stanford university, graduate school of business. van der krogt, f. j. (1998). learning network theory: the tension between learning systems and work systems in organizations. human resource development quarterly, 9(2), 157-177. doi: 10.1002/hrdq.3920090207 venkatraman, n. (1994). it-enabled business transformation: from automation to business scope redefinition. sloan management review, 35(2), 73-87. wellman, b. (2001). computer networks as social networks. computers and science, 293(5537), 20312034. doi: 10.1126/science.1065547 yamnill, s., & mclean, g. n. (2001). theories supporting transfer of training. human resource development quarterly, 12(2), 195. doi: 10.1002/hrdq.7 m. rehm et al. 55 | f l r yang, c.-c., tsai, i., kim, b., cho, m.-h., & laffey, j. m. (2006). exploring the relationships between students' academic motivation and social ability in online learning environments. the internet and higher education, 9(4), 277-286. yates, j., & orlikowski, w. j. (1992). genres of organizational communication a structurational approach to studying communication and media. academy of management review, 17(2), 299326. doi: 10.2307/258774 zack, m. h., & mckenney, j. l. (1995). social-context and interaction in ongoing computer-supported management groups. organization science, 6(4), 394-422. doi: 10.1287/orsc.6.4.394 frontline learning research 1 (2013) 42 71 issn 2295-3159 corresponding author: mariel f. musso, katholieke universiteit leuven / universidad argentina de la empresa, mariel.musso@hotmail.com http://dx.doi.org/10.14786/flr.v1i1.13 42 | f l r predicting general academic performance and identifying the differential contribution of participating variables using artificial neural networks mariel f. musso ab , eva kyndt ac , eduardo c. cascallar ad , filip dochy a a katholieke universiteit leuven, belgium b universidad argentina de la empresa, argentina c university of antwerp, belgium d assessment group international, usa / belgium article received 8 march 2013 / revised 2 july 2013 / accepted 16 july 2013 / available online 27 august 2013 abstract many studies have explored the contribution of different factors from diverse theoretical perspectives to the explanation of academic performance. these factors have been identified as having important implications not only for the study of learning processes, but also as tools for improving curriculum designs, tutorial systems, and students’ outcomes. some authors have suggested that traditional statistical methods do not always yield accurate predictions and/or classifications (everson, 1995; garson, 1998). this paper explores a relatively new methodological approach for the field of learning and education, but which is widely used in other areas, such as computational sciences, engineering and economics. this study uses cognitive and non-cognitive measures of students, together with background information, in order to design predictive models of student performance using artificial neural networks (ann). these predictions of performance constitute a true predictive classification of academic performance over time, a year in advance of the actual observed measure of academic performance. a total sample of 864 university students of both genders, ages ranging between 18 and 25 was used. three neural network models were developed. two of the models (identifying the top 33% and the lowest 33% groups, respectively) were able to reach 100% correct identification of all students in each of the two groups. the third model (identifying low, mid and high performance levels) reached precisions from 87% to 100% for the three groups. analyses also explored the predicted outcomes at an individual level, and their correlations with the observed results, as a continuous variable for the whole group of students. results demonstrate the greater accuracy of the ann compared to traditional methods such as discriminant analyses. in addition, the ann provided information on those predictors that best explained the different levels of expected performance. thus, results have allowed the identification of the specific influence of each pattern of variables on different levels of academic performance, providing a better understanding of the variables with the greatest impact on individual learning processes, and of those factors that best explain these processes for different academic levels. keywords: predictive systems; academic performance; artificial neural networks m. f. musso et al. 43 | f l r 1. introduction many studies have explored the contribution to the explanation of academic performance with the use of various different variables and from diverse theoretical perspectives (e. g. bekele & mcpherson, 2011; fenollar, roman, & cuestas, 2007; kuncel, hezlett, & ones, 2004; miñano, gilar, & castejón, 2008). many factors have been identified as having important implications not only for the study of learning processes, but also as tools for improving of curriculum designs, tutorial systems, and students‟ academic results (miñano et. al., 2008; musso & cascallar, 2009a; zeegers, 2004). from this previous body of research, it has become apparent that the accurate prediction of student performance could have many useful applications for positive outcomes of the learning process and lead to advances in learning theory. for example, it could be helpful to identify students at risk of low academic achievement (musso & cascallar, 2009a; ramaswami & bhaskaran, 2010). this prediction could serve as an early warning of future low academic performance and guide interventions that could prove beneficial for such students. similarly, being able to understand the role of different intervening variables that influence performance for all and for each category of performance level, would be a significant contribution to improve the approach to teaching and better understand learning processes. many previous studies have focused on the prediction of academic performance (e.g., hailikari, nevgi, & komulainen, 2008; krumm, ziegler, & buehner, 2008; turner, chandler, & heffer, 2009). many of the studies about academic performance have considered grade point average (gpa) as the best summary of student learning, not only because of its strong prediction of performance for other levels of education (e. g. kuncel et al., 2004, 2005), but also for other life outcomes as salary (roth & clarke, 1998), and job performance (roth, be vier, switzer, & schippman, 1996). the prediction of academic performance has been carried out with different methodological approaches. the first and most common approach found in the educational literature, has to do with the use of traditional statistical methods, such as discriminant analysis and multiple linear regressions (braten & stromso, 2006; vandamme, meskens & superby, 2007). a second approach can be found in various studies which have used structural equation modelling (sem) to compare theoretical models to data sets and/or to test different models of academic performance (fenollar et al., 2007; miñano et al., 2008; ruban & mccoach, 2005). these traditional approaches – that are tools widely used to predict gpa, to orient selection, placement, and/or classification of the academic process –failed to consistently show the capacity to reach accurate predictions or classifications in comparison with artificial intelligence computing methods (everson, chance, & lykins, 1994; kyndt, musso, cascallar, & dochy, 2012, submitted; lykins & chance, 1992; maucieri, 2003; weiss & kulikowski, 1991). therefore, a third approach to the “prediction of academic performance” that we can find in recent literature involves machine learning techniques, such as methods using artificial neural networks (ann). this method has been used and proven useful in several other fields, such as business, engineering, meteorology, and economics. it is considered an important method to classify potential outcomes and is well regarded as an excellent pattern-recognizer (detienne, detienne, & joshi, 2003; neal & wurst, 2001; white & racine, 2001). recent work in the field of computer sciences has started to apply this methodology to large data banks of nation-wide educational outcomes (abu naser, 2012; croy, barnes, & stamper, 2008; fong, si, & biuk-aghai, 2009; kanakana, & olanrewaju, 2011; maucieri, 2003; mukta & usha, 2009; pinninghoff junemann, salcedo lagos, & contreras arriagada, 2007; ramaswami & bhaskaran, 2010; zambrano matamala, rojas díaz, carvajal cuello, & acuña leiva, 2011; walczak, 1994). this methodology has also recently been used with various applications in educational measurement, in conjunction with other theoretical models of different constructs such as self-regulation of learning (cascallar, boekaerts & costigan, 2006; everson et al., 1994; gorr, 1994; hardgrave, wilson, & walstrom, 1994), reading readiness (musso & cascallar, 2009a); and performance in mathematics (musso & cascallar, 2009b; musso, kyndt, cascallar, & dochy, 2012). the application of predictive systems, with the emergence of new methodologies and technologies, have made it possible to assess a wide range of data and student performances in order to evaluate their current and future performance without the need for traditional testing (boekaerts & cascallar, 2006; cascallar et al., 2006). this methodological approach using ann can lead to the possible implementation of continuous assessment in the context of intelligent classrooms (birenbaum et al., 2006). m. f. musso et al. 44 | f l r existing databases together with the constant monitoring of student performance could provide a continuous evaluation in real time of the students‟ progress. the interrelationship between many of the variables participating in the complex and multi-faceted problem of academic performance are not clearly understood, and they are often related in nonlinear ways. ann have demonstrated to be a very effective approach to address situations with these characteristics and to be able to classify and predict outcomes under those conditions with a high level of accuracy, especially when large data sets are available. this approach also allows the researcher to consider a large number of variables simultaneously and make use of their interrelationships without the usual parametric constraints. these advantages would allow researchers in the learning sciences to better understand the complex patterns of interactions between the variables at different levels of academic performance, not just for the prediction of performance but also to understand the participating factors that could be related to these outcomes. several previous studies using ann have addressed the classification of outcomes into different levels of performance, for different academic purposes: a) diagnostic purposes in order to identify those students most in need of support at the beginning of their primary school, regarding their readiness for learning to read (musso & cascallar, 2009a), and b) identifying students with low expected writing performance at the vocational secondary school level in order to provide support prior to their first year, and thus avoiding possible failure (boekaerts & cascallar, 2011). in these and other possible applications, the early detection of future low performance, and more targeted interventions, would decrease the negative experience of failure, and it would provide an important diagnostic tool for effective interventions. this approach would improve the chances of achieving successful outcomes, particularly for students identified as being “at-risk”. detecting and understanding the most significant variables that are the best indicators of the future low performers would be an important tool for management of school resources and planning remediation programs at all levels of an educational system. similarly, knowing the best indicators of the future high performers, would allow first of all the understanding of many of the factors leading to these positive outcomes. it would also allow an accurate selection of those students who could be assigned to advance programs, fellowships and/or be the object of talent searches. the accurate placement of students in different courses or programs according to how they are expected to perform would prevent possible failure, as well as providing the opportunity to offer challenging tasks for students expected to be among the high performers. in addition, a better understanding of the interrelationships between the variables leading to different levels of performance, would allow the fine-tuning of instructional approaches to the individual and/or group needs using the information provided by an ann approach. some authors have shown that traditional statistical methods do not always yield accurate predictions and/or classifications (bansal, kauffman & weitz, 1993; everson, 1995; duliba, 1991). preliminary research using ann for prediction, selection, and classification purposes suggests that this method may improve the validity and accuracy of the classifications, as well as increase the predictive validity of educational outcomes (everson et al., 1994; hardgrave et al., 1994; perkins, gupta, tammana, 1995; weiss & kulikowski, 1991). this paper explores this new methodological approach using a large amount of data collected from the students (including both cognitive and non-cognitive measures) in order to design predictive models using artificial neural networks (ann). the ann models in this research study can identify those predictors that could best explain different levels of academic performance in three different performance groups which cover all the range of performances, as well as making accurate classifications of the expected level of performance for each subject. data about individual differences in basic cognitive variables were collected, since they are strongly related to the student‟s achievement (colom, escorail, chin shih, & privado, 2007; grimley & banner, 2008). although it has been argued that considering students‟ cognitive ability can lead to a relatively strong prediction of academic performance (colom et al., 2007), this prediction could be strengthened by including background and non-cognitive predictors. as chamorro-premuzic & arteche (2008) discuss, combining both cognitive ability and non-cognitive measures can provide a broader understanding of an individual‟s likelihood to succeed in academic settings, with models that predict such m. f. musso et al. 45 | f l r performance at least one academic year in advance of the actual measure being obtained (grade-point average, gpa). in addition, discriminant analyses (da) was used to analyse the same data in order to compare the predictive classificatory power of both methodologies. to better understand the rationale for this research, it is useful to review some of the main constructs included as predictors in this study, and to explain the quite novel methodology introduced from the family of predictive systems, that is, the machine learning modelling technique of artificial neural networks (ann). 2. theoretical considerations 2.1 working memory and academic performance intelligence and the g-factor are the most frequently studied factors in relation to academic achievement and the prediction of performance (miñano et al., 2012). there is a large body of research that shows a strong positive correlation between g and educational success (e.g., kuncel, hezlett, & ones, 2001; linn & hastings, 1984). the g-factor is defined, in part, as an ability to acquire new knowledge (e.g., cattell, 1971; schmidt, 2002; snyderman & rothman, 1987). although the g-factor is not the same construct as working memory (wm), several studies have demonstrated a high correlation between these measures (heitz et al., 2006; unsworth, heitz, schrock, & engle, 2005). following the early study of daneman and carpenter (1980) on individual differences in working memory capacity (wmc) and reading comprehension, further research has shown the importance of wmc as a domain-general construct (conway, cowan, bunting, therriault, & minkoff, 2002; conway & engle, 1996; engle & kane, 2004; feldman barrett, tugade, & engle, 2004; kane et al., 2004), including the prediction of average scores over several academic areas (colom et al., 2007). similarly, a large body of literature shows wmc as a very important construct in several areas and several studies have shown its importance in a wide range of complex cognitive behaviours such as comprehension (e.g., daneman & carpenter, 1980), reasoning (e.g., kyllonen & christal, 1990), problem solving (welsh, satterleecartmell, & stine, 1999) and complex learning (kyllonen & stephens, 1990; kyndt, cascallar, & dochy, 2012; st clair-thompson & gathercole, 2006). wmc is an important predictive variable of intellectual ability and academic performance, consistent over time (e.g. engle, 2002; musso & cascallar, 2009a; passolunghi & pazzaglia, 2004; pickering, 2006). working memory is a paradigmatic form of cognitive control that explains how this cognitive control occurs, and which involves the active maintenance and executive processing of information available to the cognitive system, combining the ability to both maintain and effectively process information with minimal loss (jarrold & towse, 2006). it is crucial for the processing of information within the cognitive system, it has a limited capacity and it differs between individuals (conway et al., 2005). the literature seems to indicate two fundamental approaches according to the interpretation of working memory and executive control. traditional perspectives represent working memory and executive control as separate modules (e.g., baddeley, 1986). the perspective taken in this research coincides with another view that understands working memory and executive control as constituting two sides of the same phenomenon, an emergent property from the neuro-cognitive architecture (anderson, 1983, 1993, 2002, 2007; anderson et al., 2004; hazy; frank & o‟reilly, 2006). 2.2 attention and academic performance attention as a cognitive construct has been studied from different theoretical and methodological approaches (e.g., posner & rothbart, 1998; redick & engle, 2006; rueda, posner, & rothbart, 2004). it is evident that our cognitive system is constantly receiving a variety of inputs form the environment. all these inputs are competing for the limited resources of the cognitive system, and requiring our “attention”. m. f. musso et al. 46 | f l r however, because human cognitive capacities are limited in their ability to process information simultaneously (gazzaniga, ivry, & mangun, 2002), it is the shifting of the processing capacity and selection of stimuli to attend to, which constitute the basic aspects of our attentional system (redick & engle, 2006). this shifting and selection of incoming information is the function of the attentional system, which allows us to redirect our attention to the relevant aspects of the environmental information for the task or goals at hand. this study adopts the framework of posner and petersen (1990) who described three different and semiindependent attentional networks: orientation, alertness and executive attention. the orienting network allows the selection of information from sensory input, the alerting network refers to a system that achieves and maintains an alert state, and executive attention or executive control is responsible for resolving conflict among responses (fan, mccandliss, summer, raz, & posner, 2002). the efficiency of these three attentional networks can be quantified by reaction time measures (fan et al., 2002). redick and engle (2006) and unsworth et al. (2005) have found that individual differences in working memory capacity are related to those in attentional control, thus establishing that the executive control mechanism is closely related to working memory capacity. several studies have shown the importance of attention as a predictor of general academic performance (gsanger, homack, siekierski, & riccio, 2002; kyndt et al., 2012, submitted; riccio, lee, romine, cash and davis, 2002), reading (landerl, 2010; lovett, 1979), mathematical performance (fernandez-castillo & gutiérrez-rojas, 2009; fletcher, 2005; musso et al., 2012), and written expression (reid, 2006). the research on learning disorders has found that attentional problems are negatively associated to academic achievement (jimmerson, dubrow, adam, gunnar, & bozoky 2006). 2.3 learning strategies and academic performance the estimated level of contribution of basic cognitive processes to the determination of academic achievement has shown considerable variation, which ranges from a moderate to a medium-high effect (castejón & navas, 1992; navas, sampascual, & santed, 2003). consequently, the studies focusing on the prediction of academic performance have increasingly included the so-called non-cognitive variables such as motivation, attributions, self-concept, effort, goal orientation, etc. (e.g., fenollar et al., 2007; pintrich, 2000). learning strategies (ls) have been defined as student‟s actual behaviours, in a specific context, to engage in a task (biggs, 1987). other researchers describe ls as any thoughts or behaviours that help the students to acquire new information and integrate this new information with their existing knowledge (weinstein & mayer, 1986; weinstein, palmer, & schulte, 1987; weinstein, schulte & cascallar, 1982). ls also help students retrieve stored information. examples of ls include summarizing, paraphrasing, imaging, creating analogies, note-taking, and outlining (weinstein et al., 1987). previous research has provided support for the mediating role of learning strategies (dupeyrat & marine, 2005; fenollar et al., 2007; simons, dewitte, & lens, 2004). fenollar et al. (2007) have compared a theoretical model, where achievement goals and self-efficacy were hypothesised to have direct effects on academic performance, to a mediating model where such effects were mediated through study strategies. results from the study showed that achievement goals and self-efficacy have no direct effects on performance, and they suggest that the mediating model provides a better fit to the data (fenollar et al., 2007). 2.4 artificial neural networks and performance conceptually, a neural network is a computational structure consisting of several highly interconnected computational elements, known as neurons, perceptrons, or nodes. each “neuron” or unit carries out a very simple operation on its inputs and transfers the output to a subsequent node or nodes in the network topology (specht, 1991). neural networks exhibit polymorphism in structure and parallelism in computation (mavrovouniotis & chang, 1992), and it can be represented as a highly interconnected structure of processing elements with parallel computation capabilities (grossberg, 1980, 1982; rumelhart, hinton, & m. f. musso et al. 47 | f l r williams, 1986; rumelhart, mcclelland, & the pdp research group, 1986). in general, an ann consists of an input layer (which can be considered the independent variables), one or more hidden layers, and an output layer that is comparable to a categorical dependent variable (cascallar et al., 2006; garson, 1998). all ann process data through multiple processing entities which learn and adapt according to patterns of inputs presented to them, by constructing a unique mathematical relationship for a given pattern of input data sets on the basis of the match of the explanatory variables to the outcomes for each case (marshall & english, 2000). thus, neural networks construct a mathematical relationship by “learning” the patterns of all inputs from each of the individual cases used in training the network, while more traditional approaches assume a particular form of relationship between explanatory and outcome variables and then use a variety of fitting procedures to adjust the values of the parameters in the model. during the training phase, anns generate a predicted outcome for each case, and when this prediction is incorrect the network makes adjustments to the weights of the mathematical relationships among the predictors and with the expected outcome, weights that are represented in the hidden layers of the network. the predicted output is a continuous variable with a specific value for each case (or subject) which includes information on the probability of belonging to each of the categorical classifications requested by the developer of the ann. according to this architecture, the ann finally recognizes patterns and classifies the cases presented into the requested outcome categories, depending on the target question, and given the individual probability values for each case. this information is generated by the network through many iterations, gradually changing and adjusting the weights for all the interrelationships between the units after each incorrect prediction. during this training process, the network becomes increasingly accurate in replicating the known outcomes from the test cases. the neural network continues to improve its predictions until one or more of the pre-determined stopping criteria have been met. these stopping criteria can be, for example, a minimum level of accuracy, learning rate, persistency, number of iterations, amount of time, etc. once trained, the network is tested with the remaining cases in the dataset, which is considered a form of validation of the network (testing phase), by observing how the weights in the model, now fixed to those obtained in the training phase, predict classes of outcomes in a new set of data of which outcomes are known to the experimenter but not to the ann system. afterwards it can also be applied to predict future cases where the outcome is still unknown (cascallar et al., 2006). in addition, with complementary techniques in predictive stream analysis, the neural network approach allows us to determine the predictive power of each of the variables involved in the study, providing information about the importance of each input variable (cascallar et al., 2006; garson, 1998). predictive stream analyses (cascallar & musso, 2008), based in this case on neural network (ann) models, have several strengths: (a) because these are machine learning algorithms, the assumptions required for traditional statistical predictive models (e.g., ordinary least squares regression) are not necessary. as such, this technique is able to model nonlinear and complex relationships among variables. ann aim to maximize classification accuracy and work through the data in an interactive process until maximum accuracy is achieved, automatically modelling all interactions among variables; (b) anns are robust, general function estimators. they usually perform prediction tasks at least as well as other techniques and most often perform significantly better (marquez, hill, worthley, & remus, 1991); (c) ann can handle data of all levels of measurement, continuous or categorical, as inputs and outputs. because of the speed of microprocessors in even basic computers, anns are more accessible today than when they were originally developed. current research has shown that neural network analysis substantially improves the validity of the classifications and increases the accuracy and predictive validity of the models, in education and other fields (kyndt et al., 2012, submitted; musso & cascallar, 2009b; perkins et al., 1995). the ann learns by examining individual training cases (subjects/students), then generating a prediction for each student, and making adjustments to the weights whenever it makes an incorrect prediction. information is passed back through the network in iterations, gradually changing the weights. as training progresses, the network becomes increasingly accurate in replicating the known outcomes. this process is repeated many times, and the network continues to improve its predictions until one or more of the stopping criteria have been met. a minimum level of accuracy can be set as the stopping criterion, although m. f. musso et al. 48 | f l r additional stopping criteria may be used as well (e.g., number of iterations, amount of processing time). once trained, the network can be applied, with its structure and parameters, to future cases (validation or holdout sample) for further validation studies and programme implementation (lippman, 1987). as long as the basic assumptions of the population of persons or events that the ann used for training is constant or varies slightly and/or gradually, it can adapt and improve its pattern recognition algorithms the more data it is exposed to in the implementations. the class of ann models used in this research can be compared with the more traditional discriminant analysis approach. both of these methods derive classification rules from samples of classified objects based on known predictors. this general approach is called „supervised learning‟ since the outcomes are known and relationships are modelled or „supervised‟ according to these outcomes (kohavi & provost, 1998). but, there are significant differences in the algorithms and procedures for both analyses, such as the fact that while discriminant analysis assumes linear relationships, neural network analysis does not. in terms of comparisons with another common statistical method used in educational research, linear regression, it is important to note that although neural networks can address some of the same research issues as regression it is inherently a different mathematical approach (detienne et al., 2003). there is another family of predictive systems which are “unsupervised” (e.g., kohonen networks), in which the patterns presented to the network are not associated with specific outcomes; it is the neural network itself that derives the commonalities between the predictors, grouping cases into classes on the basis of these similarities. thus, these analyses can be used to explore the data from a different perspective and learn the grouping of cases based on these predictor commonalities instead of being focused on predictions or individual outcomes (cascallar et al., 2006; kyndt et al., 2012, submitted). neural networks excel in the classification and prediction of outcomes; especially when large data sets are available that are related in nonlinear ways, and where the intercorrelation between variables is not clearly understood. these properties of anns clearly make them particularly suitable for social science data where they can simultaneously consider all variables in a study (garson, 1998). moreover, the assumptions of normality, linearity and completeness that are made by methods such as multiple linear regression (kent, 2009), and that are often very difficult to establish for social science data, are not made in neural network analysis. neural networks can work with noisy, incomplete, overlapping, highly nonlinear and noncontinuous data because the processing is spread over a large number of processing entities (garson, 1998, kent, 2009). in this regard it can be said that neural networks are robust and have wide non-parametric application. there is also evidence that neural models are robust in the statistical sense, and also robust when faced with a small number of data points (garson, 1998). very few studies within the educational literature have used neural network analysis or any other type of predictive system (e.g., cascallar et al., 2006; cascallar & musso, 2008; musso & cascallar, 2009a; pinninghoff junemann et al., 2007; wilson & hardgrave, 1995). 2.5 ann processing and measures to evaluate the neural network system performance in order to evaluate the performance of the neural network system, there are a number of measures used which provide a means of determining the quality of the solutions offered by the various network models tried. the traditional measures include the determination of actual numbers and rates for true positive (tp), true negative (tn), false positive (fp), and false negative (fn) outcomes, as products of the ann analysis. in addition, certain summative evaluative algorithms have been developed in this field of work, to assess overall quality of the predictive system. these overall measures are: recall, which represents the proportion of correctly identified targets, out of all targets presented in the set, and is represented as: recall = tp/(tp + fn); and precision which represents the proportion of correctly identified targets, out of all identified targets by the system, and is represented as: precision = tp/(tp + fp). two other measures, derived from signal-detection theory (roc analysis), have also been used to report the characteristics of the detection sensitivity of the system. one of them is sensitivity (similar to recall: the proportion of correctly identified targets, out of all targets m. f. musso et al. 49 | f l r presented in the set), and which is expressed as sensitivity = tp/(tp + fn). the other is specificity, defined as the proportion of correctly rejected targets from all the targets that should have been rejected by the system, and which is expressed as specificity = tn/(tn + fp). all the traditional measures are typically represented in what is called a “confusion matrix” representing all four outcomes. in addition, the evaluation of ann performance is also carried out with another summative measure, which is used to account for the somewhat complementary relationship between precision and recall. this measure is defined as f1, and is defined as f1 = (2 * precision * recall)/(precision + recall). such a definitional expression of f1 assumes equal weights for precision and recall. this assumption can be modified to favour either precision or recall, according to the utility and cost/benefit ratio of outcomes favouring either precision or recall for any given predictive circumstance. 2.6 objectives and research questions the objective of this study is to identify patterns of variables that will allow a correct predictive classification of three levels of general academic performance (gap) into: low, middle and high gap, measured by the grade-point-average (gpa). this was achieved by taking into consideration basic cognitive processes (working memory capacity; alerting, orienting and executive attention), learning strategies, and family-social background factors. the idea behind this paper is to explore new approaches to obtain predictive classifications of learning outcomes, without the use of one specific test, using a large number of variables (cognitive and non-cognitive) that could better capture the true complex composite of influences participating in the actual observed outcomes from individual students. in addition, it is another objective of the research to explore the differences in the patterns predicting each level of performance (low, middle and high performance) to inform future research into the causal factors generating and participating in those sets of identified variables and that could explain different levels of performance using artificial neural networks. of course, previous academic performance could have been taken into account to facilitate the predictive classification, but this was purposely avoided for two reasons: as a proof-of-concept that other variables are sufficient to predict academic performance, and to highlight more clearly the weight that each of these other variables has in the determination of a student‟s academic performance. in order to explore the differences in the patterns predicting each level of performance, three artificial neural network (ann) models were developed. two of them to predict the students who would be in each of the extreme performance levels (low 33% and high 33% of gpa) in order to analyse the differences between the patterns of variables having the most predictive weight for each group, and thus providing information on the potentially different processes involved in those low and high performance outcomes. a third ann was developed, capable of accurately producing a predictive classification for the three levels of performance simultaneously (low 33%, middle 33%, and high 33%). this final ann model was capable of finding the common patterns that could predict simultaneously all performance groups. the relative importance of the predictors for each network was also analysed. the predictive capability of each ann was systematically improved by modifying the parameters that determine the rate of learning, the persistence, momentum, and stopping criteria, and the type of functions used for weight adjustments. precision, sensitivity, specificity and accuracy of the three networks were obtained. in addition, the correlation between the individual prediction for each student and the actual observed gpa was established, and proved to be very high. the main research questions of this study are: how accurately can different levels of academic performance in higher education be predicted by working memory capacity, attentional networks, learning strategies and background variables when used as inputs in a neural network model? what is the relative importance of the predictor variables and the observed differences for each performance level category? m. f. musso et al. 50 | f l r 3. method 3.1 participants the total sample included 864 university students, of both genders (male 45.4%; female 54.6%), ages between 18 and 25 (mage = 20.38, sd = 3.78), recently enrolled in the first year in several different disciplines (psychology, engineering, medicine, law, social communication, business and marketing), in three private universities in argentina, during the 2009-2011 academic years. in all, 67.8% of the sample was 17 to 20 years old, 24.7% was 21-25 years old, and 7.5% was older than 25 years. the students in the sample came from private religious secondary schools (48.5%), private non-religious schools (19%) , private bilingual schools (15.4%), public secondary schools (15%), and 2.1% from international community schools. all student data (predictors) was collected at the beginning of the corresponding academic year, and the dependent variable (gpa) was collected at the end of the same academic year. an 80% math accuracy criterion was imposed for all participants in the automated operation span (unsworth et al., 2005). therefore, they were encouraged to keep their math accuracy at or above 80% at all times (to insure that the interfering task was actually being performed). as a consequence of this criterion, 78 participants were excluded from the analyses. the final sample consisted of 786 students. 3.2 instruments 3.2.1 attention network test (ant) (fan et al., 2002) this computerized task provides a measure for each of the three anatomically defined attentional networks: alerting, orienting, and executive. the ant is a combination of the cued reaction time (posner, 1980) and the flanker test (eriksen & eriksen, 1974). the participant saw an arrow on the screen that, on some trials, was flanked by two arrows to the left and two arrows to the right. participants were asked to determine when the central arrow points left or right, by two mouse buttons (leftright). they were instructed to focus on a centrally located fixation cross throughout the task, and to respond as quickly and accurately as possible. during the practice trials, but not during the experimental trials, subjects received feedback from the computer on their speed and accuracy. the practice trials took approximately 2 minutes and each of the three experimental blocks took approximately 5 minutes. the whole experiment took about twenty minutes. the measure for (general) attention is the average response time regardless of the cues or flankers. to analyse the effect of the three attentional networks, a set of cognitive subtractions described by fan et al. (2002) were used. the efficiency of the three attentional networks is assessed by measuring how response times are influenced by alerting cues, spatial cues, and flankers (fan et al., 2002). the alerting effect was calculated by subtracting the mean response time of the double-cue conditions from the mean response time of the no-cue conditions. for the orienting effect, the mean response time of the spatial cue conditions (up and down) were subtracted from the mean response time of the center cue condition. finally, the effect of the executive control (conflict effect) was calculated by subtracting the mean response time of all congruent flanking conditions, summed across cue types, from the mean response time of incongruent flanking conditions (fan et al. 2002). the test-retest reliability of the general response times (in this study used as a measurement of general attention), calculated by fan et al. (2002) equaled .87. the test-retest reliability of the subtractions is less good. the executive control is the most reliable (r=.77), followed by the orienting network (r=.61). the alerting network showed to be the least reliable (r=.52) (fan et al. 2002). m. f. musso et al. 51 | f l r 3.2.2 automated operation span (unsworth et al., 2005) this is a computer-administered version of the ospan instrument (unsworth et al., 2005) that measures working memory capacity. the responses were collected via click of a mouse button. first, participants receive practice and secondly, the participants perform the actual experiment. the practice sessions are further broken down into three sections. the first practice is a simple letter span task. they see letters appear on the screen one at a time. in all experimental conditions, letters remain on-screen for 800 milliseconds (ms). then, participants must recall these letters in the same order they saw them from a 4 x 3 matrix of letters (f, h, j, k, l, n, p, q, r, s, t, and y) presented to them. recall consists of clicking the box next to the appropriate letters; the recall phase is untimed. after each recall, the computer provides feedback about the number of letters correctly recalled. next, participants practice the math portion of the experiment. participants first see a math operation (e.g. (1*2) + 1 = ?). once the participant knows the answer they click the mouse to advance to the next screen. participants then see a number (e.g. “3”) and are required to click if the number is the correct solution by clicking on “true” or “false.” after each operation participants are given feedback. the math practice serves to familiarize participants with the math portion of the experiment, as well as to calculate how long it takes a given person to solve the math problems, establishing an individual baseline. thus, it attempts to account for individual differences in the time it takes to solve math problems. this is then used as an individualized time limit for the math portion of the experimental session. the final practice session has participants perform both the letter recall and math portions together, just as they will do in the experimental block. the participants first are presented with a math operation, and after they click the mouse button indicating that they have solved it, they see the letter to be recalled. if the participants take more time to solve the math operations than their average time plus 2.5 sd, the program automatically moves on and counts that trial as an error. this serves to prevent participants from rehearsing the letters when they should be solving the operations. participants complete three practice trials, each of set size 2. after the participant completes all of the practice sessions, the program moves them on to the real trials. the real trials consist of 3 sets of each set-size, with the set-sizes ranging from 3 to 7 letters. this makes for a total of 75 letters and 75 math problems. subjects are instructed to keep their math accuracy at or above 85% at all times. during recall, a percentage in red is presented in the upper right-hand corner. subjects are instructed to keep a careful watch on the percentage in order to keep it above 85%. this study reports the absolute ospan score (the sum of all perfectly recalled sets) that is interpreted as the measure of overall working memory capacity, and one reaction time score (operations). the task takes approximately 20–25 minutes to complete (unsworth et al., 2005). this measure of working memory capacity has a high correlation with other measures of working memory and general intelligence, as ospan and raven progressive matrices. in addition, aospan has a good test-retest reliability (r = .83) and an adequate internal consistency (α=.78) (unsworth et al., 2005). 3.2.3 learning strategies questionnaire (lassi; weinstein et al.,1987; weinstein & palmer, 2002; weinstein et al., 1982). the original version is a 77-item questionnaire with 10 scales that assesses the students' awareness about, and use of, learning and study strategies related to skill, will, and self-regulation components of strategic learning. these scales and their corresponding internal consistency coefficients reported in the users‟ manual (weinstein & palmer, 2002), are as follows: attitude scale (α = .77), motivation scale (α = .84), time management scale (α = .85), anxiety scale (α = .87), concentration scale (α = .86), information processing scale (α = .84), selecting main ideas scale (α = .89), study aids scale (α= .73), self-testing scale (α = .84), and test strategies scale (α = .80). the present study used a spanish-version (strucchi, 1991), which was slightly modified in some semantic and grammatical aspects for the local sample. the exploratory factor analysis determined a matrix with five factors that explained 37.52% of the variance. factor 1 related to “cognitive resources/cognitive processing” (α = .871; 13 items; r 2 = 18.03%); factor 2, related to “time management” (α = .807; 10 items; r 2 = 8.404%); factor 3, dealing with “processing of information and generalization” (α = .783; 8 items; r 2 = 4.567%); factor 4 which is related to “anxiety management” (α = .60; 5 items; r 2 = 3.431%); and factor 5, which involves the construct of “study m. f. musso et al. 52 | f l r techniques and use of help” (α = .728; 7 items; r 2 = 2.685%). students gave responses on a likerttype scale, from 1 (never) to 5 (always). 3.2.4 background information basic background information of each student used in the analyses was: gender, highest level of education of mother and father (not completed primary schoolprimary schoolsecondary schoolgraduated universitypost-graduate), occupation of parents, and secondary school from which the student graduated (public private religious school private non-religious school bilingual school foreign community) 3.2.5 academic performance academic performance was measured by the grade point average (gpa) of all courses (different subjects depending on the discipline) at the end of each of the academic years. all course grades which are used by the universities to calculate the overall gpa are obtained using university-wide criteria for the interpretation and assignment of final scores in each course, from which the gpa was calculated. the gpa information was collected from official records at the end of the first academic year for each student, at each of the participating universities, and they all are in a scale from 0 to 10 (with 10 indicating best performance). 3.3 analyses procedure the ann model used was a backpropagation multilayer perceptron neural network, that is, a multilayer network composed of nonlinear units, which computes its activation level by summing all the weighted activations it receives and which then transforms its activation into a response via a nonlinear transfer function, which establishes a relationship between the inputs and the weights they are assigned. during the training phase, these systems evaluate the effect of the weight patterns on the precision of their classification of outputs, and then, through backpropagation, they adjust those weights in a recursive fashion until they maximize the precision of the resulting classifications. ann parameters and variable groupings, as well as all other network architecture parameters, were adjusted to maximize predictive precision and total accuracy. confusion matrices have been determined for each ann, as well as roc analyses for the evaluation of sensitivity and specificity parameters. parameters such as learning rate (the rate at which the ann “learns” by controlling the size of weight and bias changes during learning), momentum (adds a fraction of the previous weight update to the current one, and is used to prevent the system from converging to a local minimum), number of hidden layers, stopping rules (when the network should stop “learning” to avoid over-fitting the current sample), activation functions (which define the output of a node given an input or set of inputs to that node or unit), and number of nodes were specified and varied in the model construction phase in order to maximize the overall performance of the network model. 3.4 architecture of the neural networks according to the objectives of this research, three different neural networks (ann) were developed as predictive systems for the gpa of the students in this study. ann1 was developed to maximize the predictive classification of the lowest 33% of students, which would be scoring the lowest average gpa at the end of the academic year. ann2 was developed to maximize the predictive classification of the highest 33% of students, which would be scoring the highest gpa. ann3 was developed to predict the classification of students into the three levels of expected gpa at the same time. the data set was partitioned into a training set and a testing set for each ann, and for each network, training and testing samples were chosen at random by the software, from the available set of cases. one suggested criterion is that the number of m. f. musso et al. 53 | f l r training inputs (cases) should be at least 10 times the number of input and middle layer neurons in the network (garson, 1998). similarly, it is suggested that about 2/3 (or 3/4) of the cases in the available data set be used for the training phase in order to include a set of cases representing most of the patterns expected to be present in the data (patterns represented by the vector for each case). the remaining 1/3 or 1/4 of the data is used for the testing phase of the network. the specific architecture of each of the three neural networks developed is as follows: ann1 (maximizing the prediction for the low 33% performance group): all cognitive variables, learning strategies, and background variables were introduced in the analysis. they were used for the development of the vector-matrix containing all predictor variables for each student. the resulting network contained all the input predictors, with a total of 18 input units (reaction time operation, reaction time math, reaction time problem, orienting attention, alerting attention, executive control, absolute aospan, processing of information/ generalization, study techniques and use of help, anxiety management, time management, cognitive resources/cognitive processing, gender, mother's occupation, father's occupation, secondary school from which the student graduated, highest level of education completed by father, and highest level of education completed by mother). the model built contained one hidden layer, with 15 units. the output layer contained a dependent variable with two units (categories corresponding to “belongs to lowest 33%” or “belongs to highest 67 %”). in terms of the architecture of the network, a standardized method for the rescaling of the scale dependent variables was used. the hidden layer had a hyperbolic tangent activation function which is the most common activation function used for neural networks because of its greater numeric range (from -1 to 1) and the shape of its graph. the output layer utilized a softmax activation function that is useful predominantly in the output layer of a clustering system, converting a raw value into a posterior probability. the output layer used the cross-entropy error function in which the error signal associated with the output layer is directly proportional to the difference between the desired and actual output values. this function accelerates the backpropagation algorithm and it provides good overall network performance with relatively short stagnation periods (nasr, badr, & joun, 2002). the training was carried out with the „online‟ methodology (one case per cycle), with an initial learning rate of 0.4, and momentum equal to 0.9. the optimization algorithm was gradient descent (which takes steps proportional to the negative of the approximate gradient of the function at the current point), and the minimum relative change in training error was 0.0001. ann2 (maximizing the prediction for the high 33% performance group): all cognitive, learning strategies, and background variables were introduced in the analysis. they were used for the development of the vector-matrix containing all predictor variables for each student. the resulting network contained all the input predictors, with a total of 18 units (reaction time operation, reaction time math, reaction time problem, orienting attention, alerting attention, executive control, absolute aospan, processing of information/generalization, study techniques and use of help, anxiety management, time management, cognitive resources/cognitive processing, gender, mother's occupation, father's occupation, secondary school from which the student graduated, highest level of education completed by father, and highest level of education completed by mother). the model built contained one hidden layer, with nine units, and an output layer with two units (categories corresponding to “belongs to highest 33%” or “belongs to lowest 67%”). in terms of the architecture of the network, a standardized method for the rescaling of scale dependent variables was used. the hidden layer had a hyperbolic tangent activation function. the output layer utilized a softmax activation function. cross-entropy was chosen as the error function. the dataset was partitioned into training set and testing set. the training was carried out with the „online‟ methodology, with an initial learning rate of 0.5, and momentum equal to 0.7. the optimization algorithm was gradient descent, and the minimum relative change in training error was 0.0001. ann3 (maximizing the simultaneous prediction for all the performance groups: low 33% middle 33% high 33%, simultaneously): all cognitive, learning strategies and background variables were introduced in the analysis. they were used for the development of the vector-matrix containing all predictor m. f. musso et al. 54 | f l r variables for each student. the resulting network contained all the input predictors, with a total of 19 input units (reaction time operation, reaction time math, reaction time problem, orienting attention, alerting attention, executive control, absolute aospan, processing of information/ generalization, study techniques and use of help, anxiety management, time management, cognitive resources/cognitive processing, gender, mother's occupation, father's occupation, secondary school, highest level of education completed by father, and highest level of education completed by mother, ln of attention total rt). the model built contained one hidden layer, with 20 units, and one output layer with three units (categories corresponding to “belongs to low 33%”, “belongs to middle 33%” or “belongs to high 33%” of the performance groups). in terms of the architecture of the network, a standardized method for the rescaling of scale dependent variables was used. the hidden layer and the output layer both had a hyperbolic tangent activation functions. a standardized method for the rescaling of covariates was used. sum of squares was chosen as error function. the dataset was partitioned into training set and testing set. the training was carried out with the „online‟ methodology, with an initial learning rate of 0.4, and momentum equal to 0.8. the optimization algorithm was gradient descent, and the minimum relative change in training error was 0.0001. the software used was spss v.19 – neural network module, for the development and analysis of all predictive models in this study. two development phases of the predictive system were carried out: training of the network and testing of the network developed. during the training phase several models were attempted, and several modifications of the neural network parameters were explored, such as: learning persistence, learning rate, momentum, and other criteria. these tests continued until achieving desired levels of classification, maximizing the benefits of the model chosen. in these analyses both precision and recall, as outcome measures of the network, were given equal weight. there was no need to trim the number of predictor inputs in the three models. the validation procedure used was the leave-one-out methodology. 3.5 discriminant analyses discriminant analyses (da) were carried out using the same data and the same categories of gpa used in the neural networks analyses. da1 was performed to discriminate between the students belonging to the lowest 33% of gpa and contrasting them against those not in that category. da2 was focused on identifying students in the highest 33% of academic performance versus those not in that group, and da3 was calculated to discriminate the students belonging to each one of the three levels of gpa performance. in order to give every variable the opportunity to contribute significantly to the prediction, a stepwise discriminant analysis was calculated for each category including all independent variables. in addition, we calculated three discriminant analyses, one for each category including the independent variables of the maximised neural networks of each category. 4. results 4.1 descriptive data the final sample included 786 university students from several disciplines (psychology, engineering, medicine, law, social communication, business and marketing), in three private universities, during the 2009-2011 academic years. descriptive statistics of the cognitive variables and learning strategies are presented in table 1 (cognitive variables) and table 2 (learning strategies). m. f. musso et al. 55 | f l r table 1 descriptive statistics for attentional networks, general reaction time, working memory capacity (absolute aospan) and reaction time operation alerting attention orienting attention executive control ln of attention total rt absolute aospan (sum of perfectly recalled sets) ln rt operation n 786 786 786 786 786 786 mean 34.40 44.01 102.54 6.20 27.88 7.01 sd 22.14 22.90 41.68 .11 14.83 .20 skewness .25 .24 3.31 .67 .25 .46 kurtosis 1.96 5.01 26.14 .98 -.510 .45 minimum -78.00 -77.67 19.00 5.92 0 6.50 maximum 123.83 213.83 558.00 6.74 68 7.75 note: ln of attention total rt: logarithm of attention total reaction time (measure of attention network test) ln rt operation: logarithm of reaction time operation (measure of aospan) table 2 descriptive statistics for each factor of learning strategies (lassi) cognitive resources/cog nitive processing time management processing of information/ generalization anxiety management study techniques and use of help n 756 756 756 756 756 mean -.02 .00 .01 .00 -.01 sd 1.09 1.12 1.11 1.15 1.14 skewness .24 .18 -.37 .35 -.67 kurtosis -.16 -.21 -.07 -.41 -.03 minimum -2.87 -2.86 -4.61 -2.53 -4.24 maximum 3.85 3.30 2.56 3.57 2.22 4.2 neural network analyses ann1 was designed to predict the performance group corresponding to the lowest 33% of predicted gpa. it included 82.4 % of the participants (n = 632) in the training phase and 17.6% (n = 111) in the testing phase. after training, ann1predicting the group with the low 33% of academic performance – was able to reach 100% correct identification of the students that belong to the target group (lowest 33%) (see figure 1).the precision of ann1 equalled 1 on a maximum of 1. the sensitivity of the network equalled 1, and the specificity (defined as the proportion of correctly rejected targets from all the targets that should have been rejected by the system) was equal to 1. the area under the curve equalled .877. m. f. musso et al. 56 | f l r prediction of academic performance 33% lowest (target group) others observed academic performance 33% lowest (target group) 100% 0% others 0% 100% figure 1. testing phase of the neural network predicting the lowest 33% of academic performance scores. in general, several tables (3-5) show the actual predictive weights of the variables that the anns used in the prediction of future academic performance for each of the groups (low 33%, high 33% and the whole sample). the “importance” column can be interpreted as the actual predictive weight of each variable, and the “normalized importance” column represents the percent of predictive weight for each variable (in each group‟s analysis) with respect to the variable with the greatest predictive weight for the group in question, which is assigned a 100%. table 6 summarizes the actual predictive weights of the variables, grouped by construct: background variables (i.e., parents‟ education, parents‟ occupation, type of secondary school), basic cognitive variables (i.e., working memory capacity, attentional networks), reaction time variables (i.e., operations, attentional), and learning strategies/motivation variables (i.e., study techniques, time management, anxiety management). it allows an easier comparison of the sources of predictive weights by area between the various student groups and also for the total sample. table 3 shows the actual predictive weight of each input, and the normalised importance of the different variables for the ann1 predictive classification. these results indicate that the learning strategies regarding cognitive processes, reaction time (rt), and time management were the most important predictors. all reaction times are converted to natural logarithms (ln) of the actual rt. table 3 relative importance of the most predictive variables included in the model for the predictive classification of the lowest 33% of scores in academic performance low 33% group independent variable importance variables importance normalized importance cognitive resources/cognitive processing 0.092 100.00% ln reaction time math 0.083 90.80% time management 0.080 87.30% secondary school from which the student graduated 0.066 71.50% father's occupation 0.065 70.90% executive control 0.062 67.60% mother's occupation 0.058 63.70% ln reaction time problem 0.058 62.80% m. f. musso et al. 57 | f l r absolute aospan (sum of perfectly recalled sets) 0.055 60.50% anxiety management 0.051 55.40% alerting attention 0.050 54.40% ln reaction time operation 0.048 52.40% orienting attention 0.048 52.10% study techniques and use of help 0.046 51.70% processing of information/ generalization 0.043 46.50% gender 0.040 43.70% highest level of education completed by mother 0.030 32.60% highest level of education completed by father 0.025 27.10% ann2 was designed to predict the performance group corresponding to the highest 33% predicted gpa. it included 77.9% of the students in the training phase (n= 614) and 22.1% in the testing phase (n= 136). after training, ann2 reached an accuracy of 100 % (see figure 2). the precision of ann2 equalled 1 on a maximum of 1. the sensitivity of the network equalled 1, and the specificity amounted to 1. the area under the curve equalled .788. prediction of academic performance 33% highest (target group) others observed academic performance 33% highest (target group) 100% 0% others 0% 100% figure 2. testing phase of the neural network predicting the highest 33% of academic performance scores. the most important variables for the prediction of ann2 (high 33%) were reaction time, mother‟s occupation, type of secondary school, father‟s occupation and executive control (executive attention measure) (see table 4). m. f. musso et al. 58 | f l r table 4 relative importance of the most predictive variables included in the model for the predictive classification of the highest 33% of scores in academic performance high 33% group independent variable importance variables importance normalized importance ln of reaction time operation 0.084 100.00% mother's occupation 0.081 97.10% secondary school from which the student graduated 0.081 96.10% father's occupation 0.076 90.10% executive control 0.072 86.40% alerting attention 0.062 73.90% processing of information/ generalization 0.055 65.10% orienting attention 0.054 64.10% study techniques and use of help 0.053 62.30% highest level of education completed by father 0.051 60.70% ln of reaction time math 0.049 58.50% anxiety management 0.047 55.60% highest level of education completed by mother 0.044 52.80% absolute aospan (sum of perfectly recalled sets) 0.044 52.70% time management 0.044 52.20% cognitive resources/cognitive processing 0.037 44.70% ln of reaction time problem 0.033 39.90% gender 0.033 39.60% both networks showed interesting differences in the pattern of relative normalized importance of those variables with the highest participation in the predictive model. for the low performing group in terms of general gpa (those predicted to be in the lowest 33% of scores), several learning strategies related to cognitive processes, reaction time (wmc and attentional networks functioning), and time management were most important in providing predictive weights for a correct classification. on the other hand, results from the predictive model for those students expected to be in the highest 33% of the general gpa scores, the top three predictors with the most significant participation were background variables involving mother‟s and father‟s occupation, type of secondary school, and overall reaction time of the cognitive and attentional processes. ann3, which was designed to predict the three gpa performance groups simultaneously, used 82.8% of the students (n=710) for the training phase, and 17.2% (n=122) for the testing phase. after maximizing the training procedures, the accuracy in the testing phase reached 87.5% for the lowest 33%, 100% for the middle 33%, and 100% for the highest 33% (see figure 3). the precision of ann3 equalled .875 on a m. f. musso et al. 59 | f l r maximum of 1. the sensitivity of the network equalled 1, and the specificity amounted to .50. the areas under the curve were .658 for the low 33%, .583 for the middle 33%, and .637 for the high 33%. prediction of academic performance 33% lowest middle 33% 33% highest observed academic performance low 33% 87.5 % 10% 2.5% middle 33% 0% 100% 0% high 33% 0% 0% 100% figure 3.testing phase of the neural network predicting the three levels of academic performance scores (low 33%middle 33%high 33%). the most important variables for the prediction of ann3 were orienting attention, learning strategies related to the cognitive resources and information processing, time management, and executive control (executive attentional network) (see table 5). table 5 relative importance of the most predictive variables included in the model for the predictive classification of the three levels of academic performance all 3 groups gpa (low 33% mid 33% high 33%) independent variable importance variables importance normalized importance orienting attention 0.087 100.00% cognitive resources/cognitive processing 0.076 86.86% time management 0.074 84.92% executive control 0.073 83.30% father's occupation 0.071 81.80% mother's occupation 0.070 79.91% ln of attention total reaction time 0.067 77.25% alerting attention 0.067 76.63% ln of reaction time math 0.061 70.14% processing of information/ generalization 0.050 57.20% ln of reaction time operation 0.043 49.64% study techniques and use of help 0.041 46.55% ln of reaction time problem 0.040 46.13% m. f. musso et al. 60 | f l r anxiety management 0.038 43.89% gender 0.032 36.67% highest level of education completed by father 0.031 35.73% absolute aospan (sum of perfectly recalled sets) 0.031 35.09% highest level of education completed by mother 0.026 29.88% secondary school 0.024 27.29% 4.3 maximizing the ann models all ann models were developed so as to maximize the accuracy of the classification. the number of units in the hidden layers was determined by optimizing the ability of the hidden nodes to store the necessary weight information, while avoiding the over-determination that would result from an excessive number of units. while greater number of units would have given the model greater flexibility, it would have increased complexity at the cost of decreasing generalizability to the testing sample. similarly, not enough units would not have produced a proper fit with the data and would have reduced the power of the model. therefore, various models were developed in order to find the proper balance and maximize the predictive power for each model. in all models, the training and testing samples were selected at random from the existing data and the proportions were adjusted in order to maximize the training sample while preserving the appearance of all detected patterns in the testing sample, so as to be able to appropriately test the model. other parameters that were varied in order to maximize the performance of the networks were learning rate and momentum. the variations in the learning rate parameter allowed the control of the amount of weight and bias change during the training of the network. different problem conditions find better solutions with different size of changes in the architecture of the network. regarding the momentum, it was used to prevent the network from converging too early to a local minimum, and conversely to avoid overshooting the global minimum of the function; thus, it is important to avoid having a value which is too large for the momentum (it can overshoot), or too low (it can get stuck in a local minimum). balancing these parameters maximizes the solution, and if correctly identified provide a stable and reliable solution as the ones that were found in this study. 4.4 predictive contribution by categories of variables besides studying the contribution of each variable individually for each neural network developed to classify the various expected performance levels (low performers, high performers, and three performance groups simultaneously), the contribution of each category or set of variables (background, basic cognitive processes, total reaction times for wmc operations and attentional networks, and learning strategies/motivation) was analysed for each ann developed, and the total predictive weight for each category of variables, as well as their average, was determined. table 6 and figure 4 show that in terms of predictive weight, the most important variables when estimating the levels of predicted gpa performance for all three groups simultaneously, are the background factors (e.g., socio-economic status proxy data, type of secondary school, occupation and education of parents, etc.), but when comparing the two extreme predicted performance groups, it is interesting to note that specific patterns involving different variables are evident for low and high expected academic performance: learning strategies/motivation had a stronger predictive weight for students expected to be in the lowest 33% of gpa performance; on the other hand, for students predicted to belong to the highest 33% of gpa performance, background variables and some of the cognitive processing variables were those carrying the most predictive weight. m. f. musso et al. 61 | f l r table 6 comparative predictive weight contribution for the three levels of academic performance by each of the categories of predictor variables low 33% mid 33 high 33% mean predictive weight of each area background 28.40% 25.40% 36.60% 30.13% basic cognitive 21.50% 25.70% 23.20% 23.47% reaction time total 18.90% 21.10% 16.60% 18.87% learning strategies/motivation 31.20% 27.80% 23.60% 27.53% 100% 100% 100% figure 4. comparison of predictive weight levels for the three levels of academic performance by categories of predictor variables. 4.5 initial analysis of individual continuous estimates of future academic performance while most of this study has been centered around the successful development of models to categorize expected levels of performance (which can be varied according to the problem situation), it is also important and useful to demonstrate that this machine learning approach can be used to predict individual specific outcomes (not just relatively broad performance categories). although these performance categories can be very useful, as has been indicated for the identification and possible intervention in specific groups of high achievers or low achievers (i.e., learning disabilities, non-readiness for some specific task such as reading), and they can be used very effectively for targeted interventions in learning situations, it is also important to 0% 5% 10% 15% 20% 25% 30% 35% 40% low 33% middle 33% high 33% background basic cognitive reaction time total learning strategies/motivation m. f. musso et al. 62 | f l r be able to understand the underlying phenomenon at the individual level, considering performance a continuous variable. for this reason, the predicted gpa-category (low-middle-high) probability values assigned by the network to each individual student were used to analyze their correlation with the observed gpa, as compared to the predicted value, in the context of the ann3 model, in which the whole sample of students was simultaneously classified in the three levels of expected performance. that is, the probability value for each student of belonging to a given category (all students received a certain probability of belonging to each of the outcome groups, as determined by the ann), was correlated with the gpa actually obtained by each student. results were indicative of a high degree of correlation between those measures. the three predicted groups of low, mid, and high performance had an actual observed gpa mean of 3.88 (sd = 1.21, n = 327), 5.67 (sd = .33, n = 243), and 7.28 (sd = .78, n = 294), respectively. all these average gpa means were significantly different from each other (p < .000). within each one of the performance levels, the correlation of the ann individual predicted value with the actual gpa was: low 33%, r = .78; high 33%, r = .73, and for the whole sample of students, at all three levels, the correlation of the ann predicted values with the observed gpa was r = .86. further studies will continue to explore these individual relationships, but as they are, they confirm a high level of correlation between the actual gpa and the expected values assigned by the ann. 4.6 discriminant analyses (da) da1 focused on the attempted predictive classification of students expected to be in the lowest 33% of gpa average, compared to the rest of the students. one of the restrictions of this analysis has to do with the assumption of equality of covariance matrices that, in this case, is not violated (box‟s m = 5.253, f = .871, p= .515). gender, wmc and cognitive resources/learning strategies, were able to discriminate between the two groups of students, but not the rest of the variables, that were included in the ann1. the squared canonical correlation (cr²) gives the amount of variation between the groups that is explained by the discriminating variables, which in this case was quite low (wilk‟s λ = .896, χ² = 84.786, df = 3, p = .001, cr² = .323). da2 was carried out to attempt to discriminate between students expected to be in the highest 33% of gpa average, compared to the 67% of the rest of the students. the same independent variables that were used in the ann2 were entered in this analysis. results show that the independent variables were not able to discriminate between both groups of students. the box‟s m statistic is not significant (box‟s m = 11.813, f = .781, p = .700), meaning that the assumption of equality of covariance matrices is not violated. in this analysis the squared canonical correlation indicated that the strength of the function is very low (wilk‟s λ = .926, χ² = 58.694, df = 5, p = .001, cr² = .271). only gender, highest level of education of the father, wmc, and cognitive resources, and time management among the learning strategies set, were variables that entered significantly in this model. da3 was carried out with the same variables as those used to develop ann3, in order to predict the expected gpa performance level of the three groups of academic performance simultaneously. the assumption of equality of covariance matrices was not violated (box‟s m = 7.522, f = .623, p = .824). in this case, only gender, cognitive resources within the learning strategies set and wmc were significant for the model, and participated in the discrimination between the students in the three groups. but the model explained a very low and non-significant proportion of the variance (wilk‟s λ = .998, χ² = 1.791, df = 2, p = .408, cr² = .048). m. f. musso et al. 63 | f l r 5. discussion and conclusions the purpose of this study was to show the applicability and the effectiveness of the ann approach to the predictive classification of students in the full range of academic performance (gpa), as well as to identify and understand the importance of the variables for each level (low, middle and high) of expected gpa. this methodology, using a predictive system, was chosen as it is very effective under conditions of very complex and great amount of data, in which a large number of variables interact in various complex and not very well understood patterns. the results attained in this study have allowed the identification of the specific influence of each input set of variables on different levels of academic performance (high and low performance), on one hand, and common processes across all students, on the other hand. one important contribution of this predictive approach is the finding that the same variables have different effects in each group of students, defining specific patterns for each performance level. although the contribution of each variable in a particular pattern carries a relatively small predictive weight, it is the combined effect of the pattern of variables which explains a lower or higher academic performance model. among the student group with the lowest 33% of academic performance, two main predictors are learning strategies components (cognitive resources/cognitive processing and time management). the importance of learning strategies as a mediating factor in a model predicting academic performance has been shown in different studies (dupeyrat & marine, 2005; fenollar, et al., 2007; simons et al., 2004; weinstein & mayer, 1986; weinstein et al., 1987; weinstein et al., 1982). however, this study added the contribution of a complex pattern of variables for a particular group of students, identifying specific learning strategies that help the classification of students in a low performance group (i.e., thoughts or behaviours that help to use imagery, verbal elaboration, organization strategies, and reasoning skills). included in this set are learning strategies that help build bridges between what they already know, and what they are trying to learn and remember (i.e., knowledge acquisition, retention, and future application). in addition, variables related to speed of processing involved in wmc functioning have an important predictive weight for the determination and modelling of the low performance group. other studies that have used ann have also found that basic cognitive processing variables such as wmc and executive attention carried the most predictive weight in the low performance group of students (kyndt et al., 2012, submitted; musso & cascallar, 2009a; musso et al., 2012). moreover, the literature has indicated the positive association between wmc and academic achievement (gathercole, pickering, knight, & stegmann, 2004; riding, grimley, dahraei, & banner, 2003). regarding the relative importance of each variable, if we compare the relative role of wmc and other cognitive resources between the low and high performance groups, wmc and cognitive resources were far more important for lower gpa students. the fact that their importance for the prediction is much greater for the lower performing group is greatly due to the fact that all members of the high group had higher levels of wmc and cognitive resources, therefore not providing the necessary information to the network. on the other hand, it was an identifying characteristic of the low performing group which had consistently lower values of wmc and cognitive resources. remediation programmes, tutorial systems and instruction methods should consider these specific learning strategies, cognitive processing characteristics and wmc resources, in order to provide basic support to students at risk. such informed interventions would improve the possibilities of successful academic achievement for the at-risk groups, including those with particular learning difficulties. background variables together with reaction time measures and attentional executive control are the most important predictors for the highest academic performance group, as indicators of both efficiency in the processing and of adequate selection of information. social background variables, such as educational level of the parents, have been found to be significant in a previous ann study (pinninghoff, junemann et al., 2007), and these results have been replicated in this study. the executive control mechanism is responsible for resolving conflicts among responses (fan et al., 2002). this attentional system has been closely related to working memory capacity (redick & engle, 2006), and was found to mediate and compensate wmc deficits for certain tasks (musso et al., 2012). other attentional networks seem to be much less discriminating among students who reach certain threshold levels needed for high academic performance. these findings have m. f. musso et al. 64 | f l r significant implications in the way that the learning process can be addressed for students identified as potential high achievers. for this group, promoting learning through the use of metacognitive strategies, complex processing, and targeted teacher feedback would be an important way of maximizing their potential performance. regarding methodological implications, these results demonstrate the greater accuracy of the ann approach compared to other traditional methods such as da. other studies have also made use of multilayer perceptron artificial neural networks, with positive results for the analysis of educational data (abu naser, 2012; croy et al., 2008; fong, et al., 2009; kanakana, & olanrewaju, 2011; mukta & usha, 2009; ramaswami & bhaskaran, 2010; zambrano matamala, et al., 2011). however, the present study has been able to maximize the precision obtained in the predictive classification of overall academic performance through the careful adjustment of network parameters and algorithms, producing highly accurate results with minimal misclassifications. similarly, the initial study of the correlation between the ann probabilities of performance level assigned to each individual student, with the actual gpa observed, shows a significant degree of correlation between the two measures (r = .86 for the whole sample), with performance as a continuous variable. further studies will refine the technique to maximize these individual results. the results of the da confirm the lack of significant linear relationships between the independent variables analysed in this study and academic performance. neural network models have an important advantage in this respect, as they are able to model nonlinear and complex relationships among variables with greater precision and accuracy. even though the assumptions required for traditional statistical predictive models (e.g. equality of covariance matrices) were not violated for the three stepwise discriminant analyses that were performed, the amount of variance explained was low in all three da analyses. none of these analyses were able to discriminate with sufficient accuracy between the different levels of expected academic performance. when we compare these results with the anns modelled in this study, it can be concluded that anns are much more robust, and perform significantly better than other classical techniques, as prior studies have also indicated (everson et al., 1994; marquez et al., 1991). this study has shown the power of this predictive approach using anns to model future overall academic performance in higher education, specifically in academic admissions and/or placement. to put the current results in perspective, if we consider one of the best known and most reliable tests currently in use, the sat from the college board, it has been found (kobrin, patterson, shaw, mattern, & barbuti, 2008) that all sections of the sat taken together, even with the more recent addition of a writing score, can predict at best 28% of the variance of the first-year college gpa for the average population of students. if we add to the sat results the information of the gpa obtained in secondary education, the overall prediction is of only 38% of the variance of first-year college gpa (kobrin et al, 2008). with the current ann models, it has been possible to correctly classify 100% of student performance in the categories examined, that is, 100% of the students were correctly classified, and our research currently continues into the development of new predictive models, with much larger data sets, to classify students in much narrower bands of expected performance having already attained 98-99% accuracy in models for quintals of student performance distributions. in addition, work will also continue for the prediction of specific expected gpa results for each individual student. in conclusion, the current predictive systems approach facilitates and maximizes the identification of those factors (or predictors) of the learning processes which participate in varying degrees in the modelling of different levels of performance in academic outcomes in higher education. if we can identify specific profiles of students, focusing on the most important variables, this opens major possibilities for the improvement of assessment procedures and the planning of pre-emptive interventions. given that this methodology allows for the accurate prediction of actual academic performance at least one academic year in advance to it actually being measured (gpa), it has implications for the application of these methods in educational research and in the implementation of diagnostic “early-warning” programmes in educational settings. these results also inform cognitive theory and help in the development of improved automated tutoring and learning systems. although some of the variables involved, such as educational level of the m. f. musso et al. 65 | f l r parents, are impossible to alter in their effects on academic performance at the time of the assessment, they do inform policy and indicate the weight that many social and environmental factors influence future academic performance. this methodological and conceptual approach allows us to consider a large number of variables simultaneously and select those which are most relevant and allow a greater degree of intervention to improve student performance, including early intervention programmes for students in need of special support. the capacity to very accurately classify expected student performance, which is also what tests attempt to do, without the performance sampling issues of traditional testing, and using a much broader spectrum of all factors influencing a student‟s overall performance, is a major advantage of the anns methodology. in fact, it also represents a more valid approach to educational assessment due to its overall accuracy and the breadth of the constructs considered to classify the expected performance. traditional assessments are not sufficient for more complex assessments or for assessment systems that intend to serve multiple direct and indirect purposes, in complex educational situations (mislevy, 2013; mislevy, steinberg, & almond, 2003) in this respect, this new approach allows for the conceptualization and development of new modes of assessment which could facilitate breaking away from traditional forms of testing while at the same time improving the quality of the assessment process (segers, dochy & cascallar, 2003). finally, the use of ann together with other methods as cluster analyses and kohonen networks could contribute to the study of the specific patterns of those variables which influence the learning process for each level of performance. in fact, a major observation resulting from the data in this study is that variables contribute to the prediction in relatively small proportions, and it is the joint effect of many contributing variables that could cause significant changes in performance. in other words, there is no “magic bullet”, rather the accumulation of effects from all these various sources that produces significant changes in outcomes. these results provide an insight into learning questions from a different perspective and one that has important implications for educational policy and education at large. keypoints this approach provides a more contextualized and encompassing new mode of assessing expected performance without some of the pitfalls found in traditional forms of testing. anns are a powerful tool to model future academic performance, specifically in academic diagnostic evaluations for placement and early-warning assessments. this methodology demonstrates that variables impacting the outcome of the learning process are embedded in specific large-scale patterns which determine their degree of influence and direction of their effects. a predictive systems approach is a valuable method to study the specific patterns of variables influencing the learning process at each level of expected performance, to better understand the determinants of learning outcomes and ways to improve them with early interventions. references abu naser, s. s. (2012). predicting learners performance using artificial neural networks in linear programming intelligent tutoring system. international journal of artificial intelligence & applications (ijaia), 3(2), 65-73 anderson, j. r. (1983). the architecture of cognition. cambridge, ma: harvard university press. anderson, j. r. (1993). rules of the mind. hillsdale, nj: lawrence erlbaum associates. anderson, j. r. (2002). spanning seven orders of magnitude: a challenge for cognitive modeling. cognitive science, 26, 85–112. m. f. musso et al. 66 | f l r anderson, j. r. (2007) how can the human mind occur in the physical universe? new york: oxford university press. anderson, j. r., bothell, d., byrne, m. d., douglass s., lebiere, c., &yulin, q. (2004). an integrated theory of the mind. psychological review, 111(4), 1036–1060. baddeley, a. d. (1986). working memory. oxford: clarendon press. bansal, a., kauffman, r. j., & weitz, r. r. (1993). comparing the modeling performance of regression and neural networks as data quality varies: a business value approach. journal of managemnet informations systems, 10(1), 1132. bekele, r., & mcpherson, m. (2011).a bayesian performance prediction model for mathematics education: a prototypical approach for effective group composition. british journal of educational technology, 42(3), 395–416. biggs, j. (1987). study process questionnaire manual. melbourne, australia: australian council for educational research. birenbaum, m., breuer, k., cascallar, e., dochy, f., dori, y, ridgway, j, wiesemes, r. (2006), & nickmans, g. (editor). a learning integrated assessment system. educational research review, 1, 61-67. boekaerts, m., & cascallar, e. (2006). how far have we moved toward the integration of theory and practice in self-regulation? educational psychology review, 18(3), 199-210. boekaerts, m. & cascallar, e. c. (2011). predicting and explaining writing outcomes: neural network methodology at work. symposium: predicting academic performance with the use of predictive systems analysis. proceedings of the biennial conference of the european association for research on learning and instruction (earli). exeter, uk, 30 august – 3 september 2011. braten, i. & stromso, h. (2006). epistemological beliefs, interest, and gender as predictors of internet-based learning activities. computers in human behavior, 22, 1027-1042. cascallar, e. c., boekaerts, m., & costigan, t. e. (2006) assessment in the evaluation of selfregulation as a process, educational psychology review, 18(3), 297-306. cascallar, e. c., & musso, m. f. (2008). classificatory stream analysis in the prediction of expected reading readiness: understanding student performance. international journal of psychology, proceedings of the xxix international congress of psychology icp 2008, 43(43/44), 231-.231. castejón, j. l., & navas, l. (1992). determinantes del rendimiento académico en la educación secundaria. un modelo causal. [determinants of academic achievement in secondary education. a causal model]. análisis y modificación de conducta, 18(61), 697-728. cattell, r. b. (1971). abilities: structure, growth and action. boston: houghton mifflin. chamorro-premuzic, t., & arteche, a. (2008). intellectual competence and academic performance: preliminary validation of a model. intelligence, 36, 564-573. colom, r., escorial, s., chun shih, p., & privado, j. (2007).fluid intelligence, memory span, and temperament difficulties predict academic performance of young adolescents. personality and individual differences, 42, 1503-1514. conway, a. r. a., cowan, n., bunting, m. f., therriault, d., & minkoff, s. (2002). a latent variable analysis of working memory capacity, short term memory capacity, processing speed, and general fluid intelligence. intelligence, 30, 163183. conway, a. r. a., & engle, r.w. (1996). individual differences in working memory capacity: more evidence for a general capacity theory. memory, 4, 577-590. conway, a. r. a., kane, m. j., bunting, m. f., hambrick, d. z., wilhelm, o., & engle, r. w. (2005).working memory span tasks: a methodological review and user‟s guide. psychonomic bulletin & review, 12(5), 769-786 croy, m., barnes, t., & stamper, j. (2008). towards an intelligent tutoring system for propositional proof construction. in a. briggle, k. waelbers, and p. brey (eds.), computing and philosophy (pp. 145215). amsterdam, the netherlands: ios press. daneman, m., & carpenter, p. a. (1980).individual-differences in working memory and reading. journal of verbal learning and verbal behaviour, 19, 450 466. m. f. musso et al. 67 | f l r detienne, k. b., detienne, d. h., & joshi, s. a. (2003). neural networks as statistical tools for business researchers. organizational research methods, 6, 236-265. duliba, k. a. (1991) contrasting neural nets with regression in predicting performance in the transportation industry. proceedings of the twenty-fourth annual hawaii international conference on system sciences, 4. dupeyrat, c., & marine, c. (2005). implicit theories of intelligence, goal orientation, cognitive engagement, and achievement: a test of dweck's model with returning to school adults. contemporary educational psychology, 30(1), 43-59. engle, r.w. (2002). working memory capacity as executive attention. current directions in psychological science, 11, 19-23. engle, r.w., & kane, m. j. (2004).executive attention, working memory capacity, and a two-factor theory of cognitive control. in b. ross (ed.), the psychology of learning and motivation (pp. 145-199). newyork, ny: elsevier. eriksen, b. a., & eriksen, c.w. (1974). effects of noise letters upon the identification of a target letter in a non search task. perception and psychophysics, 16, 143-149. everson, h. t. (1995). modelling the student in intelligent tutoring systems: the promise of a new psychometrics. instructional science, 23(5-6), 433-452. everson, h. t., chance, d., & lykins, s. (1994). exploring the use of artificial neural networks in educational research. paper presented at the annual meeting of the american educational research association, new york. fan, j., mccandliss, b. d., summer, t., raz, a., & posner, m.i. (2002).testing the efficiency and independence of attentional networks. journal of cognitive neuroscience, 14(3), 340-347. feldman barrett, l., tugade, m. m., & engle, r. w. (2004). individual differences in working memory capacity and dual-process theories of mind. psychological bulletin, 130, 553-573. fenollar, p., roman, s., & cuestas, p. j. (2007). university students‟ academic performance: an integrative conceptual framework and empirical analysis. british journal of educational psychology, 77, 873891. fernandez-castillo, a., & gutiérrez-rojas, m. e. (2009). selective attention, anxiety, depressive symptomatology and academic performance in adolescents. electronic journal of research in educational psychology, 7(1), 49-76. fletcher, j. m. (2005). predicting math outcomes: reading predictors and comorbidity. journal of learning disabilities, 38(4), 308-312. fong, s., si, y.-w., & biuk-aghai, r. p. (2009). applying a hybrid model of neural network and decision tree classifier for predicting university admission. proceedings of the 7th international conference on information, communication, and signal processing (icics2009), pp. 1-5, macau, china, ieee press. garson, g. d. (1998). neural networks. an introductory guide for social scientists. london: sage publications ltd. gathercole, s. e., pickering, s. j., knight, c., & stegmann, z. (2004).working memory skills and educational attainment: evidence from national curriculum assessments at 7 and 14 years of age. applied cognitive psychology, 18, 1-16. gazzaniga, m., ivry, r., & mangun, g. (2002).cognitive neuroscience: the biology of the mind (2nd ed.). new york, ny: w.w. norton grimley, m., & banner, g. (2008).working memory, cognitive style, and behavioural predictors of gcse exam success. educational psychology, 28(3), 341-351. grossberg, s. (1980). how does the brain build a cognitive code? psychological review, 87, 151. grossberg, s. (1982). studies of mind and brain: neural principles of learning, perception, development, cognition and motor control. boston: reidel press. gsanger, k., w., homack, s., siekierski, b., & riccio, c. (2002).the relation of memory and attention to academic achievement in children. archives of clinical neuropsychology, 17(8), 790. hailikari, t., nevgi, & a., komulainen, e. (2008). academic self-beliefs and prior knowledge as predictors of student achievement in mathematics: a structural model. educational psychology, 28(1), 59-71. http://ieeexplore.ieee.org/xpl/mostrecentissue.jsp?punumber=882 http://ieeexplore.ieee.org/xpl/mostrecentissue.jsp?punumber=882 http://www.sciencedirect.com/science?_ob=gatewayurl&_method=citationsearch&_urlversion=4&_origin=sdtoptwofive&_version=1&_piikey=s0361476x04000256&md5=b17b12e34e6786bfaf70ce74613ab07b http://www.sciencedirect.com/science?_ob=gatewayurl&_method=citationsearch&_urlversion=4&_origin=sdtoptwofive&_version=1&_piikey=s0361476x04000256&md5=b17b12e34e6786bfaf70ce74613ab07b m. f. musso et al. 68 | f l r hardgrave, b. c., wilson, r. l., & walstrom, k. a. (1994).predicting graduate student success: a comparison of neural networks and traditional techniques. computer and operations research, 21(3), 249-263. hazy, t. e., frank, m. j., & o‟ reilly, r. c. (2006). banishing the homunculus: making working memory work, neuroscience 139, 105–118. heitz, r. p., redick, t. s., hambrick, d. z., kane, m. j., conway, a. r. a., & engle, r. w. (2006). working memory, executive function, and general fluid intelligence are not the same. behavioral and brain sciences, 29, 135-136. jarrold, c., & towse, j. n. (2006). individual differences in working memory. neuroscience, 139, 39-50. jimmerson, s. r., dubrow, e. h., adam, e., gunnar, m., & bozoky, i. k. (2006).associations among academic achievement, attention, and andrenocortical reactivity in caribbean village children. canadian journal of school psychology, 21, 120-138. kanakana, g., & olanrewaju, a. (2011).predicting student performance in engineering education using an artificial neural network at tshwane university of technology, proceedings of the isem, stellenbosch, south africa. kane, m. j., hambrick, d. z., tuholski, s.w., wilhelm, o., payne, t.w., & engle, r.w. (2004). the generality of working memory capacity: a latent variable approach to verbal and visuospatial memory span and reasoning. journal of experimental psychology: general, 133, 189-217. kent, r. (2009). rethinking data analysis – part two. some alternatives to frequentist approaches. international journal of market research, 51, 181-202. kobrin, j. l., patterson, b. f., shaw, e. j., mattern, k. d., & barbuti, s. m. (2008). validity of the sat for predicting first-year college grade point average. college board research report 2008-5.new york: the college board. retrieved from http://research.collegeboard.org/rr2008-5.pdf. kohavi, r. & provost, f. (1998).glossary of terms. machine learning, 30(2–3): 271–274. kuncel, n. r., hezlett, s. a., & ones, d. s. (2001). a comprehensive meta-analysis of the predictive validity of the graduate record examinations: implications for graduate student selection and performance. psychological bulletin, 127(1), 162-181. kuncel, n. r., crede, m., thomas, l. l., klieger, d.m., seiler, s.n., & woo, s.e. (2004). a meta-analysis of the pharmacy college admission test (pcat) and grade predictors of pharmacy student success. annual conference of the american psychological society, chicago, il. kuncel, n. r., hezlett, s. a., & ones, d. s. (2004). academic performance, career potential, creativity, and job performance: can one construct predict them all? journal of personality and social psychology, 86(1), 148-161. kuncel, n. r., crede, m., thomas, l. l., klieger, d. m., seiler, s. n., & woo, s. e. (2005). a meta-analysis of the pharmacy college admission test (pcat) and grade predictors of pharmacy student success. american journal of pharmaceutical education, 69(3), 339-347. krumm, s., ziegler, m., buehner, m. (2008). reasoning and working memory as predictors of school grades. learning and individual differences, 18 (2), 248-257. kyllonen, p. c., & christal, r. e. (1990). reasoning ability is (little more than) working-memory capacity?! intelligence, 14, 389-433. kyllonen, p. c., & stephens, d. l. (1990).cognitive abilities as determinants of success in acquiring logic skill. learning and individual differences, 2, 129-160. kyndt, e., cascallar, e., & dochy, f. (2012). individual differences in working memory capacity and attention, and their relationship with students‟ approaches to learning. higher education, 64(3), 285297. kyndt, e., musso, m., cascallar, e., & dochy, f. (2012, submitted). predicting academic performance: the role of cognition, motivation and learning approaches. a neural network analysis.journal of further and higher education. landerl, k. (2010). temporal processing, attention, and learning disorders. learning & individual differences, 20(5), 393-401. linn, r. l., & hastings, c. n. (1984). a meta-analysis of the validity of predictors of performance in law school. journal of educational measurement, 21, 245-259. http://www.psychology.gatech.edu/renglelab/publications/2006/heitzetal_bbs_2006.pdf http://www.psychology.gatech.edu/renglelab/publications/2006/heitzetal_bbs_2006.pdf http://www.psychology.gatech.edu/renglelab/publications/2006/heitzetal_bbs_2006.pdf http://internal.psychology.illinois.edu/~nkuncel/gre%20meta.pdf http://internal.psychology.illinois.edu/~nkuncel/gre%20meta.pdf http://internal.psychology.illinois.edu/~nkuncel/gre%20meta.pdf http://internal.psychology.illinois.edu/~nkuncel/academic_performance%20-%20in%20jpsp,%20by%20kuncel.pdf http://internal.psychology.illinois.edu/~nkuncel/academic_performance%20-%20in%20jpsp,%20by%20kuncel.pdf http://internal.psychology.illinois.edu/~nkuncel/academic_performance%20-%20in%20jpsp,%20by%20kuncel.pdf http://internal.psychology.illinois.edu/~nkuncel/kuncel%20et%20al%202005%20-%20pcat%20-%20ajpe.pdf http://internal.psychology.illinois.edu/~nkuncel/kuncel%20et%20al%202005%20-%20pcat%20-%20ajpe.pdf http://internal.psychology.illinois.edu/~nkuncel/kuncel%20et%20al%202005%20-%20pcat%20-%20ajpe.pdf m. f. musso et al. 69 | f l r lippman, r. (1987). an introduction to computing with neuralets. ieee assp magazine, 3(4), 4-22. lovett, m. w. (1979). the selective encoding of sentential information in normal reading development. child development, 50(3), 897. lykins, s., & chance, d. (1992). comparing artificial neural networks and multiple regression for predictive application, proceedings of the eight annual conference on applied mathematics, edmond ok, 155169 marquez, l., hill, t., worthley, r., & remus, w. (1991). neural network models as an alternative to regression. proceedings of the ieee 24th annual hawaii international conference on systems sciences, 4, 129-135. marshall, d. b., & english, d. j. (2000).neural network modelling of risk assessment in child protective services. psychological methods, 5(1), 102-124. maucieri, l. p. (2003). predicting behavior with an artificial neural network: a comparison with linear models of prediction (january 1, 2003). etd collection for fordham university, ny, usa. retrieved from http://fordham.bepress.com/dissertations/aai3098134. mavrovouniotis, m. l. & chang, s. (1992).hierarchical neural networks. computers & chemical engineering, 16(4), 347-369. miñano, p., gilar, r., & castejón, j. l. (2012) a structural model of cognitive-motivational variables as explanatory factors of academic achievement in spanish language and mathematics. anales de psicología, 28(1), 45-54. mislevy, r. j. (2013). measurement is a necessary but not sufficient frame for assessment. measurement, 11, 47–49, 2013 mislevy, r. j., steinberg, l. s., & almond, r. a. (2003). on the structure of educational assessments. measurement: interdisciplinary research and perspectives, 1, 3–67. mukta, p., & usha, a., (2009). a study of academic performance of business school graduates using neural network and statistical techniques. expert systems with applications, 36(4), 7865-7872. musso, m. f., & cascallar, e. c. (2009a). new approaches for improved quality in educational assessments: using automated predictive systems in reading and mathematics. journal of problems of education in the 21 st century, 17, 134-151. musso, m. f., & cascallar, e. c. (2009b).predictive systems using artificial neural networks: an introduction to concepts and applications in education and social sciences. in m. c. richaud & j. e. moreno (eds.).research in behavioural sciences (volume i), (pp. 433-459). argentina: ciipme/conicet. musso, m. f., kyndt, e., cascallar, e. c., & dochy, f. (2012). predicting mathematical performance: the effect of cognitive processes and self-regulation factors. education research international.vol. 12. nasr, g. e., badr, e. a., & joun, c. (2002). cross entropy error function in neural networks: forecasting gasoline demand. flairs-02 proceedings of the aaai. retrieved from http://www.aaai.org/papers/flairs/2002/flairs02-075.pdf navas, l., sampascual, g., & santed, m. a. (2003). predicción de las calificaciones de los estudiantes: la capacidad explicativa de la inteligencia general y de la motivación. [prediction of students‟ performance scores: the role of the general intelligence and motivation. journal of general and applied psychology], 56(2), 225-237. neal, w., & wurst, j. (2001). advances in market segmentation. marketing research, 13(1), 14-18. passolunghi, m. c., & pazzaglia, f. (2004). individual differences in memory updating in relation to arithmetic problem solving. learning and individual differences 14(4), 219-230. perkins, k., gupta, l. & tammana (1995). predict item difficulty in a reading comprehension test with an artificial neural network. language testing, 12(1), 34-53. pickering, s. j. (2006). working memory and education. usa: academic press. pinninghoff junemann, m. a., salcedo lagos, p. a., & contreras arriagada, r. (2007).neural networks to predict schooling failure/success. in j. mira & j.r. ´alvarez (eds.), iwinac 2007, part ii, lncs 4528(pp. 571–579). berlin / heidelberg: springer-verlag. pintrich, p. r. (2000). the role of goal orientation in self-regulated learning. in m. boekaerts, p.r. pintrich, & m. zeidner (eds.), handbook of self-regulation (pp. 452–502). san diego, ca: academic press. m. f. musso et al. 70 | f l r posner, m. i. (1980). orienting of attention. quarterly journal of experimental psychology, 41a, 19-45. posner, m. i., & petersen, s. e. (1990). the attention system of the human brain. annual review neuroscience. 13, 25-42. posner, m. i., & rothbart, m. k. (1998). attention, self-regulation and consciousness. philosophical transactions of the royal society of london. series b, biological sciences, 353, 1915–1927. ramaswami, m. m., & bhaskaran, r. r. (2010). a chaid based performance prediction model in educational data mining. international journal of computer science issues, 7(1), 10-18. redick, t. s., & engle, r.w. (2006).working memory capacity and attention network test performance. applied cognitive psychology, 20, 713-721. reid, r. (2006). self-regulated strategy development for written expression with students with attention deficit/ hyperactivity disorder. exceptional children, 73(1), 53-67. riccio, c. a., lee, d., romine, c. cash, d., & davis, b. (2002).relation of memory and attention to academic achievement in adults. archives of clinical neuropsychology, 18(7), 755-756. riding, r. j., grimley, m., dahraei, h., & banner, g. (2003).cognitive style, working memory and learning behaviour and attainment in school subjects. british journal of educational psychology, 73, 749-769. roth, p. l., be vier, c. a., switzer, f. s., & schippmann, j. s. (1996). meta-analyzing the relationship between grades and job performance. journal of applied psychology, 81, 548-556. roth, p. l., & clarke, r. l. (1998). meta-analyzing the relation between grades and salary. journal of vocational behavior, 53, 386-400. ruban, l. m., & mccoach, d. b. (2005). gender differences in explaining grades using structural equation modeling. the review of higher education, 28, 475-502. rueda, m. r., posner, m. i., & rothbart, m. k. (2004). attentional control and self regulation. in r.f. baumeister & k.d. vohs (eds), handbook of self regulation: research, theory, and applications, new york: guilford press, 14: 283-300. rumelhart, d., hinton, g. & williams, r. (1986). learning representations by back-propagating errors. nature, 323, 533536. rumelhart, d. e., mcclelland, j. l., & the pdp research group. (1986). parallel distributed processing: explorations in the microstructure of cognition. volume i. cambridge, ma: mit press. schmidt, f. l. (2002). the role of general cognitive ability and job performance: why there cannot be a debate. human performance, 15, 187–210. segers, m., dochy, f., & cascallar, e. (2003).optimizing new modes of assessment: in search of qualities and standards.the netherlands: kluwer academic publishers. simons, j., dewitte, s., & lens, w. (2004). the role of different types of instrumentality in motivation, study strategies, and performance: know why you learn, so you'll know what you learn! british journal of educational psychology, 74, 343-360. snyderman, m., & rothman, s. (1987). survey of expert opinion on intelligence and aptitude testing. american psychologist, 42(2), 137-144 specht, d. (1991). a general regression neural network. ieee transactions on neural networks, 2(6), 568576. st clair-thompson, h. l., & gathercole, s. e. (2006). executive functions and achievements in school: shifting, updating, inhibition, and working memory. the quarterly journal of experimental psychology, 59(4), 745-759. strucchi, e. (1991). inventario de estrategias de aprendizaje y de estudio. [learning strategies inventory and study]. buenos aires: psicoteca. turner, e. a., chandler, m., & heffer, r. w. (2009). influence of parenting styles, achievement motivation, and self-efficacy on academic performance in college students. journal of college student development, 50, 3, 337-346. unsworth, n., heitz, r. p., schrock, j. c., & engle, r. w. (2005). an automated version of the operation span task. behavior research methods, 37(3), 498-505. vandamme, j. p., meskens, n., & superby, j. f. (2007). predicting academic performance by data mining methods.education economic, 15(4), 405-41. m. f. musso et al. 71 | f l r walczak, s. (1994). categorizing university student applicants with neural networks. ieee international conference on neural networks, 6, 3680-3685. weinstein, c. e., & mayer, r.e. (1986). the teaching of learning strategies. in m.c. wittrock (ed.), handbook of research on teaching (3rd ed.). macmillan, new york. weinstein, c. e. & palmer, d. r. (2002). lassi: user’s manual (2 nd edition). clearwater, fl: h&h publishing company, inc. weinstein, c. e., palmer, d. r., & schulte, a. c. (1987).learning and study strategies inventory. clearwater, fl: h & h publishing company, inc. weinstein, c. e., schulte, a. c, & cascallar, e. c. (1982). the learning and studies strategies inventory (lassi): initial design and development. technical report, us army research institute for the social and behavioural sciences, alexandria, va. weiss, s. m. & kulikowski, c. a. (1991). computer systems that learn. san mateo, ca: morgan kaufmann publishers. welsh, m.c., satterlee-cartmell, t., & stine, m. (1999). towers of hanoi and london: contribution of working memory and inhibition to performance. brain cognition, 41(2), 231-242. white, h. & racine, j. (2001): statistical inference, the bootstrap, and neural network modelling with application to foreign exchange rates. ieee transactions on neural networks: special issue on neural networks in financial engineering, 12, 657-673. wilson, r. l. & hardgrave, b. c. (1995). predicting graduate student success in a mba program: regression vs. classification. educational and psychological measurement, 55, 186-195. zambrano matamala, c., rojas díaz, d., carvajal cuello, k., & acuña leiva, g. (2011). análisis de rendimiento académico estudiantil usando data warehouse y redes neuronales. [analysis of students‟ academic performance using data warehouse and neural networks] ingeniare. revista chilena de ingeniería, 19(3), 369-381. zeegers, p. (2004). student learning in higher education: a path analysis of academic achievement in science. higher education research & development, 23(1), 35-56. frontline learning research 5 (2014) 46-63 issn 2295-3159 corresponding author: elisabeth wegner, institute for educational science, university of freiburg, rempartstr. 11, d-79089 freiburg, germany. e-mail: elisabeth.wegner@ezw.uni-freiburg.de http://dx.doi.org/10.14786/flr.v2i3.83 46 | f l r student teachers’ perception of dilemmatic demands and the relation to epistemological beliefs elisabeth wegner a , nora anders a , matthias nückles a a university of freiburg, germany article received 24 january 2014 / revised 3 march 2014 / accepted 4 june 2014 / available online 14 june 2014 abstract teaching is characterized by contradictory demands, resulting in teaching dilemmas. for example, to promote the continuous learning of students, teachers need to set up rules and control them, which in turn can undermine students’ intrinsic motivation. teachers have to become aware of these contradictions and need to understand that not all aspects of good teaching can be maximized at the same time. an adequate representation of the dilemmatic nature of problems of teaching is therefore crucial for judging different teaching situations. also, an adequate epistemological understanding is needed. we assessed student teachers’ (n = 122) perceptions of demands in teaching in general and in regards to specific situations, as well as their epistemological beliefs. perception of demands in general influenced the judgment of specific situations, but there was also a situation-specific component. epistemological beliefs were related to the perceptions of demands in general, especially in situations in which the dilemmatic content was highly visible. together, findings suggest that epistemological beliefs shape the perception of demands in teaching in general, and that the perception of demand in general again influences perception in specific situations. keywords: dilemmas in teaching; epistemological beliefs; teacher decision making; reflection wegner et al. 47 | f l r 1. introduction can teachers ―force‖ students to be motivated? can they adapt instruction to students‘ individual needs and treat them equally at the same time? a number of researchers have pointed out that there are several aspects of teaching that are in conflict with each other (e.g. berlak & berlak, 1981; helsper, 2004; lampert, 1985). therefore, teachers need to continuously decide between equally desirable goals, even though deciding for one goal reduces the possibility of reaching another goal because both options cannot be maximized at the same time. dilemmatic demands, as well as uncertainties and role-conflicts, have also been linked to the high rate of teachers that retire early from their jobs (e.g. schwab & iwanicki, 1982). teacher candidates have been shown to have difficulties in dealing with these kinds of dilemmatic demands (e.g. harrington, 1995; levin, 2002; schoen, 2005). teachers and teacher candidates expect that more knowledge about pedagogy could solve dilemmatic problems (lampert, 1985; fenstermacher, 1994). therefore, the perception of demands in teaching should be related to beliefs about the nature of pedagogical knowledge or knowledge in general, that is, epistemological beliefs. also, there is evidence that some kinds of dilemmas are more apparent then others (levin, 2002; wegner & nückles, 2011). therefore, the awareness of dilemmatic demands might be situation-specific. the question of the role of epistemological beliefs in the perception of demands in teaching, and the question of situation-specificity of the perception of demands in teaching have important consequences for the development of measures for fostering awareness of dilemmatic demands. therefore, we examined in our study (1) how teacher students perceive the demands in teaching in general and how they judge specific dilemmatic teaching situations, (2) how the general perception of demands relates to judgment of specific situations, and (3) which role epistemological beliefs in general and in regards to pedagogy play in the perception of demands in teaching and the judgment of specific teaching situations. we will at first outline what we mean by dilemmatic demands and characterize teaching as dealing with ill-structured problems, then we will summarize the (sparse) research on teacher candidates‘ dealings with dilemmatic demands, and afterwards we will outline the relation of perception of demands in teaching to epistemological beliefs. finally we will present evidence from our study suggesting that the general perception of demands in teaching is related to epistemological beliefs, and also influences the judgment of specific teaching situations, but that there is also a situation-specific component in the judgment of teaching situations. 1.1 dilemmas in teaching and their sources dilemmas in teaching can be tracked down to multiple sources, such as insufficient resources, too many tasks, administrative hierarchies, and badly organized departments that can force teachers to choose between equally necessary actions (e.g. berlak & berlak, 1981; cuban, 1992; windschitl, 2002). other dilemmas stem from the multiple roles to which teachers are assigned within the educational system (e.g. schwab & iwanicki, 1982). for example, educational institutions typically fulfill the function of both educating and assessing students at the same time. because students‘ grades in school or university greatly impact the future lives of the students, students will usually try to get as good grades as possible. therefore, the double demand of educating and assessing can present teachers with the dilemma that they want students to indicate if they have problems, but because the teacher has the power to fail them, students might decide to conceal their problems from the teacher (helsper, 2004). resource dilemmas and role conflicts are a frequent, but not necessarily inherent aspect of teaching, because they might be overcome by a different organizational structure or a better allocation of resources. however, there are other dilemmatic demands that cannot be resolved since they are part of the very nature of teaching. these genuine teaching dilemmas are located ―in the idea of teaching, constituting contradictions or contradicting demands of ideals that are equally relevant and can equally claim validity‖ (helsper, 2004, p. 61, translated by author). the following teaching dilemmas are relevant in almost all educational settings: wegner et al. 48 | f l r the dilemma of self-regulation: how much should a teacher guide students to foster learning (bräu, 2008; labaree, 2000)? teachers need to guide students in learning, provide structure and feedback in order to facilitate learning. at the same time, these supporting actions reduce opportunities for students to learn in a self-regulated way, to develop their own approaches to learning, and to learn to give feedback for themselves (e.g. windschitl, 2002). also, too much structure leads to pressure which easily reduces intrinsic motivation (deci & ryan, 1985; labaree, 2000). the dilemmas of self-regulation have been discussed with regard to many different kinds of learning environments, such as computer aided learning (koedinger & aleven, 2007) or collaborative settings (dann, 2002). the dilemma of didactic structure: how should teachers arrange the learning contents? should they arrange the contents according to the substantive structure of the subject, that is the key principles, theories and explanatory frameworks of the discipline (schwab, 1964), or should they arrange the material according to problems and situations (e.g. geddis & wood, 1997)? while the systematic approach facilitates the understanding of the subject, there is the risk of ―inert knowledge‖, which is not available for students if they have to solve complex problems as encountered in real life settings (renkl, mandl, & gruber, 1996). on the other hand, the problem-based approach facilitates the transfer of knowledge to real-life situations, because the knowledge is acquired in a way that corresponds to situations in which it could potentially be used. nevertheless, arranging learning contents in a problem-based fashion makes it potentially more difficult for students to grasp the substantive structure of the subject (albanese & mitchell, 1993). assessment dilemmas: which reference standard should assessment follow? linking assessment to individual growth fosters intrinsic motivation and values the individuals‘ progress, but on the other hand, it would be unfair if students did not receive the same grade for the same output, thus creating a dilemma between criterion-based norm and individual-based norm (hager, gonczi, & athanasou, 1994; pearson, destefano, & garcia, 1998). another dilemma in assessment is the interdependence of validity and reliability (brookhart, 1994): reliable measurement of achievement needs clear criteria. this often leads to tests that ask students to reproduce knowledge rather than to demonstrate their ability to apply it (e.g., multiple choice questions). assessments of learning outcomes that allow for higher validity, such as essays or scientific writing, have usually a lower reliability because they are less standardized and assessment is more prone to multiple biases. heterogenity dilemma: how should teachers deal with the heterogeneity regarding students‘ prior knowledge, interests and needs? optimal teaching calls for respecting the individual and his or her needs, but at the same time teachers need to treat all students equally (e.g. ball, 1993; brodie, 2010; lampert, 1985; osborne, 1997). the dilemma of professional relationship: how closely or distanced should teachers relate to their students? teachers share with other professions the challenge that they have to maintain a professional relationship, that is, they have to build a relationship without emotional involvement. they need to be neutral and need authority, but at the same time they need to create a positive climate and relationship. this creates a tension between proximity and distance (labaree, 2000). often several teaching dilemmas and structural aspects interact in creating a dilemmatic situation for a teacher. also, sometimes several teachers are involved in a dilemma and have to face the consequences of the decision, for example in assessment dilemmas. other decisions are just dilemmatic for specific situations (for example, didactic structure of one lesson), while other decisions reach out further (for example, arrangement of contents for a whole term). teaching dilemmas can be amplified by diverging expectations of students and teachers (barcelos, 2001), especially if neither learners nor teachers are aware of the dilemmatic nature of the demands in teaching. 1.2 teaching as an ill-structured problem but what do teacher students need to learn in order to deal with dilemmatic demands? teaching can be viewed as an ―ill-structured problem‖ (nespor, 1987, p. 324), that is, ―a problem for which there are conflicting assumptions, evidence, and opinion which may lead to different solutions‖ (kitchener, 1983, p. wegner et al. 49 | f l r 223). the first crucial step in dealing with this kind of problem is to come to an adequate representation of the problem space. this means, before one can start solving the problem, one needs to determine whether the problem is solvable at all, which goals might be pursued, which strategies there are to deal with it, and by which criteria these strategies might be judged. the representation of the problem space is the frame for any further cognition, such as the actual determination of the goals and the actual selection of strategies (kitchener, 1983). which kind of representation a person develops about a problem is also influenced by their epistemological beliefs, (i.e.beliefs about the nature of knowledge and knowing), because one's beliefs about the available knowledge for dealing with a problem also influence the perception of whether a problem can be solved at all. for example, a person who expects pedagogical knowledge to be stable and simple will be more likely to expect all problems in teaching to be solvable than a person who believes that pedagogical knowledge is imprecise and permanently changing. therefore, teachers need to develop an adequate representation of the problems of teaching, that is, develop awareness for dilemmatic demands of teaching, in order to be able to act in dilemmatic teaching situations. also, they need adequate beliefs about the knowledge that is available to solve dilemmatic problems. 1.3 awareness for dilemmatic demands of teaching even though there has been a substantial amount of publications on the problem of teaching dilemmas (e.g. ball, 1993; berry, 2007; cuban, 1992; geddis & wood, 1997), there are few publications that look at teachers‘ awareness of the dilemmatic demands of teaching. lampert (1985) distinguishes four perspectives on dilemmatic demands. in the perspective of ―opposing camps‖, there is little or no awareness for the dilemmatic aspects of teaching. there is one right answer, and teachers with deviant opinions have to be convinced that they are wrong. the perspective of teachers besieged by expectations accepts dilemmatic demands, but teachers are described as helpless and troubled by these demands. the origin of the conflicts is mainly seen in the organizational structure of the educational system. therefore, dilemmas can be solved by changes in the system. however, this does not help with genuine teaching dilemmas. the perspective of teachers as technical production managers and cognitive information processors holds the idea that dilemmas are created by too little knowledge. therefore, researchers have to discover the rules of how to teach, and if teachers implement these rules correctly, all problems will be eliminated. a more refined version of this view accepts the complexity of teaching. in this approach, one has to specify conditions under which circumstances which teaching behavior is appropriate. if the resulting rules are implemented correctly, problems will disappear. this view seems to be especially attractive to pre-service teachers, political decision makers and the public (fenstermacher, 1994). such ―technical rationality‖ has been criticized repeatedly by educational researchers, teacher educators and practitioners (e.g. calderhead, 1989; hatton & smith, 1995; schön, 1983). therefore, lampert puts forward the view of the teacher as dilemma manager. in this perspective, teachers realize that there are dilemmas that cannot be resolved, but only managed by reflecting on different options and weighing arguments against each other. up to now, only a few empirical studies have been conducted on how teachers or teacher candidates conceive of such dilemmas in general, and how they judge different teaching situations. lack of research might also be due to the fact that most studies used qualitative methods such as interviews or writing tasks. such methodologies are, on the one hand, appropriate given the complexity of the research question and the multitude of different kinds of dilemmas. on the other hand, qualitative methods typically limit the research to small samples. for example, schoen (2005) reports that in a sample of 10 pre-service teachers in field placements, all participants experienced dilemmas regarding students‘ discipline (e.g. ―how can a teacher keep control in the classroom without being oppressive?‖). more than half of them struggled with dilemmas between ―teacher-directed‖ and ―student-centered‖ instruction, as well as with dilemmas in dealing with heterogeneity among students. also quite frequent were dilemmas resulting from the need to prepare students for high stakes testing (such as college entrance tests), while at the same time wanting to promote complex understanding. pre-service teachers faced dilemmas in the development of a personal identity (such as developing a professional relationship to their students without too much emotional involvement) and feeling torn between the demands of field supervisors and their teaching education institution, as well as wegner et al. 50 | f l r their own goals. similarly, levin (2002) asked 12 pre-service elementary school teachers to reflect on dilemmas they encountered in their field placements. most dilemmas revolved around the relationship with their cooperating teachers and students, or classroom management concerns. none of the pre-service teachers connected their dilemmas with structural, moral, social or political issues. levin concludes that preservice teachers were only ―beginning to see the complexity and ambiguity of teachers‘ work‖ (p. 215). also, this indicates that some kinds of dilemmas are more visible than others. harrington (1995) examined how student teachers‘ ability in making reasoned decisions on exemplary dilemmas developed within one semester. participants were given dilemmatic, ill-structured cases and had to identify important issues of the case, the priority of the issues at stake, and discuss different perspectives in interpreting the case. also, they had to propose solutions, analyze different consequences of the solution and add critique to their own solution and analysis. special emphasis was put on including different perspectives on the case. 65% of the participants had difficulties in identifying the ill-structured nature of the dilemmatic cases. they failed to make connections between the different issues they had identified and addressed the issues only in isolation. figures improved substantially during the course, thus indicating the need to support pre-service teachers‘ decision making skills. 1.4 perceptions of demands and epistemological beliefs as pointed out above, perceptions of demands in teaching should be related to epistemological beliefs, because beliefs about the domain of pedagogy in general should influence perception of pedagogical problems. epistemological beliefs have been described in different ways. some researchers describe epistemological beliefs in the form of different dimensions, such as structure, certainty and sources of knowledge, as well as control and speed of knowledge acquisition (schommer, 1994; hofer & pintrich, 1997). trautwein and lüdtke (2007) found two dimensions of epistemological beliefs, relativism (―scientific knowledge can change‖) and dualism (―there is just one truth‖). the two dimensions were not independent of each other, but correlated negatively (-.36). similarly, stahl and bromme (2007) described two negatively correlated dimensions, stability and texture. other researchers (king & kitchener, 1994; kuhn, 1991; kuhn, cheney, & weinstock, 2000), who have investigated the development of epistemological beliefs, have described a stage-like development of epistemological beliefs. individuals start out from absolutistic stages (―there is only one truth‖), develop into relativistic stages (―there is no truth but only opinions‖), and eventually reach the highest, evaluatistic stage (―knowledge is subjective but can be justified to various degrees‖). krettenauer (2005) argues that the distinction between dimensional models and stage models is a result of different methodological approaches. interviews bring out the stage-like qualities of development of epistemological beliefs, whereas questionnaires focus on inter-individual differences in regards to certain dimensions of epistemological beliefs at a given point in time (see also hofer & sinatra, 2010). therefore, when assessing epistemological beliefs, one has to choose the methodological approach in consideration of the goal of the assessment. both the dimensional models as well as the stage-models of epistemological beliefs assume that epistemological beliefs are the same in all domains. however, reviews have come to the conclusion that epistemological beliefs also have a strong domain-specific component (e.g buehl, alexander & murphy, 2002; muis, bendixen & haerle, 2006). muis, bendixen and haerle (2006) state that epistemological beliefs are influenced by the socio-cultural context. academic knowledge is situated in another socio-cultural context than everyday knowledge, and also the academic contexts differ between each other. therefore, individuals‘ epistemological beliefs can differ depending on whether they relate to everyday knowledge or to knowledge in academia, and they can also differ in relation to different domains. according to muis, bendixen and haerle, beliefs from different socio-cultural contexts influence each other reciprocally. within these contexts, beliefs develop stage-like from absolutistic via relativistic into evaluatistic stages. against this background, how does the perception of demands relate to epistemological beliefs? at present, there exists a paucity of empirical evidence. wegner and nückles (2011) studied academics‘ awareness for dilemmatic demands in teaching in higher education. in an interview study they assessed the argumentative reasoning of 36 academics with regard to four dilemmatic scenarios. the authors identified wegner et al. 51 | f l r five different perspectives on the scenarios, which mirrored both kuhn‘s stages of epistemological development (kuhn, 1991) and the types of dealing with dilemmas, as lampert (1985) had described them. interviewees adopting an absolutistic perspective did not see any dilemma in the scenarios. interviewees adopting a technological perspective acknowledged the complexity of the problem, but made decisions based on heuristics and clear rules, thus ignoring the dilemma. similarly, academics with a relativistic perspective denied the principally dilemmatic nature in the scenario, because they argued that each individual teacher has his/her own approach to teaching. the academics with the most advanced perspectives recognized that there was a dilemma. under the general evaluatistic perspective, academics acknowledged complexity and were aware that an easy, general answer is not possible. academics adopting a dilemma management perspective additionally stated that dilemmatic demands have to be weighed against each other and that the problem can only be solved by making reflected decisions between equally desirable goals. wegner and nückles also found that interviewees had different perspectives in different scenarios, and that some scenarios were perceived as dilemmatic by more participants than others. this indicates that the perception of demands was specific to the situation and that dilemmas vary in their visibility. schoen (2005) further analyzed in their above-mentioned study of ten pre-service teachers how they dealt with teaching dilemmas they experienced in their field placement. she also determined different levels in dealing with dilemmas based on king and kitchener‘s (1994) stages of development in reflective judgment, ranging from ―knowledge as limited to concrete observations‖ to ―knowledge as the outcome of reasonable inquiry‖. reflective judgment level was linked to the perception of dilemmas and teachers‘ classroom activities regarding these dilemmas. generally, pre-service teachers showed medium levels of reflective judgment, thus indicating the need for improving the awareness for genuine teaching dilemmas. contrarily to the wegner and nückles study, each teacher was assigned to one level of reflective judgment, i.e. no situation-specific component was determined. both studies, wegner and nückles (2011) as well as schoen (2007), suggest that there are structural similarities in the development of the perceptions regarding the dilemmatic nature of demands in teaching and in the development of epistemological beliefs as described in the stage models by kuhn (1991) as well as by king and kitchener (1994), but that the perception of demands in teaching is a different construct. however, both studies leave important questions open for further research. none of the studies assessed epistemological beliefs separately. therefore, no conclusions about the kind of relation between epistemological beliefs and the perception of demands in teaching can be drawn from these studies. also, the studies differ in regards to whether they describe situation-specific or general aspects of the perception of demands. 1.5 relations between epistemological beliefs, general perception of demands in teaching, and the judgment of different teaching situations from the review of literature it can be concluded that teachers need to be aware of the dilemmatic nature of teaching in order to make reflected decisions. the perception of demands in teaching shape the way teachers deal with a concrete teaching dilemma and thereby their ability to make reflected decisions. also, the perception of demands in teaching is related to epistemological beliefs, especially to beliefs in the domain of pedagogy, but is nevertheless a different construct. additionally, the perception of demands might vary based on the situation. for example, there might be a difference between dilemmas that are restricted to one teaching situation (e.g. choice of contents for one lesson), and dilemmas that are more visible because they have further consequences (e.g. choice of contents for a whole term). figure 1 summarizes the assumed relations between epistemological beliefs, the general perception of demands in teaching, and the judgment of different teaching situations based on the model on muis, bendixen and haertle (2006). general and domain-specific beliefs are taken together in the graphic in order to aid in clarity. wegner et al. 52 | f l r figure 1. relations between epistemological beliefs, general perception of demands in teaching, and the judgment of different teaching situations, adopted from muis, bendixen and haerle (2006), p. 31. the arrows a, b and c denote different hypotheses (see section 2, scope of the study). 2. scope of the study in our study, we aimed to examine (1) the perception of demands in teaching in general, (2) the relation of the general perception of demands in teaching to the judgment of specific teaching situations, and (3) the relation of epistemological beliefs to the perception of demands in general as well as to the judgment of specific teaching situations. based on muis, bendixen and haerle (2006), we assume that epistemological beliefs influence perceptions of demands in general. therefore we expect medium correlations between general perception of demands in teaching and general epistemological beliefs, and slightly higher correlations to epistemological beliefs in the domain of pedagogy (hypothesis a; see arrow a in fig. 1). also, perception of demands in general influences the judgment of specific situations, but there is also an influence of the situational context. specifically, we expected differences in judgment of situations with high visibility and with weak visibility of the dilemma. due to the dependence on the situational context, we expected medium correlations of the judgment of different teaching situations with the perception of demands in teaching in general (hypothesis b). finally, we expected the correlation between general epistemological beliefs and situation-specific measures of perception of demands to be only low, because of the strong dependence on the context. again, we assumed the correlation to epistemological beliefs in the domain of pedagogy to be slightly higher than general beliefs (hypothesis c). wegner et al. 53 | f l r 3. method 3.1 sample one hundred twenty-two teacher students preparing for teaching in college-track high schools (―gymnasium‖) took part in the study. all of them filled in the questionnaires in a paper-and-pencil version at the end of a lecture on pedagogy. participants were 22.2 years old (sd = 3.3) on average; 59% were female. half of the participants (50.2%) already had had teaching experience in a field placement, lasting at least 3 months. 3.2 material 3.2.1 general perceptions about demands in teaching for the development of the questionnaire on demands in teaching, in a first step, a broad range of statements capturing different beliefs about the general nature of demands in teaching were collected based on the literature, mirroring the different perspectives on demands as outlined by lampert (1985), wegner and nückles (2011) and schoen (2005, for examples see table 2). items were piloted with a small number of teacher students, until finally 30 items were included in the questionnaire. participants were asked to rate each statement on a 6-point scale (―i don‘t agree at all – i mostly don‘t agree – i rather don‘t agree – i rather agree – i mostly agree – i completely agree‖). 3.2.2 judgment of different teaching situations based on krettenauer (2005), we developed a format of assessment in which for each item two positions were described that were related to a dilemmatic decision in teaching (e.g. teacher a says: ―i rigidly check homework because students otherwise don‘t do their assignments.‖ teacher b says: ―i usually don‘t check homework. students need to learn that they are responsible for their own learning.‖). to make sure that participants actively thought about the statements, they were asked to indicate which statement reflected their opinion most. afterwards they were asked to rate four different judgments on a 6-point scale. these judgments were developed according to kuhn‘s (1994) stages of epistemological development, lampert‘s (1985) differentiation between different perspectives on dilemmas and wegner and nückles (2011) findings (table 1). a complete sample item is given in figure 2. the final version of the questionnaire contained eight different scenarios relating to different dilemmatic decisions: • decisions related to the dilemma of self-regulation (opposing statements about regulation within cooperative learning tasks in the classroom, opposing statements about monitoring self-regulated learning tasks such as homework in general) • one decision related to the heterogeneity dilemma (opposing statements in regards to the choice of tasks for a heterogeneous group) • two decisions related to the dilemma of didactic structure (problem-centered vs. contentcentered approaches in a chemistry class, opposing approaches to the choice of contents in history classes) • two decisions related to assessment dilemma (comparison of two students according to individual vs. criterion based norm; opposing statements about the adaptation of grading to students‘ individual situations) • one decision related to the dilemma of professional relationship (opposing statements about contact with students outside school) we varied the visibility of the dilemmas by varying whether the decision had only consequences for one specific situation (that is, choice of tasks for a group, the regulation within cooperative learning tasks in the classroom, problem-centered vs. content-centered approaches, contact with students outside school), or whether the decision had further consequences for future situationsor for other people as well (e.g. both of the assessment dilemmas, choices of contents for history classes, control over homework). wegner et al. 54 | f l r table 1. selection of judgments for the scenarios kuhn (1994): epistemological beliefs lampert (1985): dealing with wegner & nückles (2011) statement absolutistic stage opposing camps absolutistic perspective ―it is absolutely clear what is right” teachers as technical production managers technological perspective “there should be clear rules for what to do in this situation” relativistic stage relativistic perspective “everyone thinks something else. you have to develop your own style” evaluatistic stage dilemma manager evaluatistic perspective “both teachers have good reasons. one needs to weigh the options carefully” figure 2. sample item from the questionnaire on judgment of teaching situations 3.2.3 epistemological beliefs for the assessment of epistemological beliefs both in general as well as in regards to the domain of pedagogy, we chose questionnaires instead of interviews because we wanted to describe inter-individual differences in beliefs, and not individual belief structures. general epistemological beliefs were assessed by a german questionnaire on epistemological beliefs, containing the two dimensions, ―dualism‖ (sample item: wegner et al. 55 | f l r ―if two scientists have a different opinion on a matter, one of them has to be wrong.‖) and ―relativism‖ (sample item: ―scientific insights that seem true today can turn out to be wrong‖, trautwein & lüdtke, 2007). domain-specific epistemological beliefs in the area of pedagogical knowledge were assessed by using the questionnaire on connotative aspects of epistemological beliefs (caeb, stahl & bromme, 2007). the caeb aims at measuring connotative aspects of beliefs that are difficult to express. participants have to rate pairs of adjectives that represent a semantic differential (such as ―strong – weak‖) on a 7-point rating-scale. the caeb comprises two dimensions that are similar to the scales of trautwein and lüttke (2007). the dimension of ―texture‖ is related to the factor ―dualism‖ and contains items that describe the accuracy and structure of knowledge in a given domain (e.g. ―knowledge in pedagogy is… preciseimprecise‖, ―structured unstructured.‖). the dimension of ―stability‖ is related to the factor ―relativism‖ and describes the stability and dynamics of knowledge (e.g. ―knowledge in pedagogy is… stable unstable‖, ―dynamic static‖). 4. results 4.1. general perception of demands in teaching at first, we analyzed the general perception of demands in teaching. for this purpose, we first determined the factorial structure of the construct of general perception of demands in teaching. we performed an exploratory factor analysis (principal component analysis, pca), because we did not expect a certain number of factors due to the complexity of the construct. findings on epistemological beliefs suggest that different factors of the perception of demands as a form of epistemic thinking are correlated with each other (e.g. stahl & bromme, 1997; krettenauer, 2005, see above). therefore we used oblique rotation (promax). neither the scree plot nor the eigenvalue criterion yielded a clear picture of the number factors. therefore, we ran factor analyses with 3, 4 and 5 factors. the three-factor solution yielded the best result, explaining altogether 37.3% of the variance. we labeled the factors ―simple demands‖, ―subjective demands‖, and ―complex demands‖ (see table 2). items which had loadings < .3 were excluded. the scale of simple demands had the lowest mean values, with the complex demands scale having highest mean values. this shows a generally high awareness for the complexity of demands in teaching. internal consistency as measured by cronbach‘s ɑ was good. also, the three factors were inter-correlated. the complex demands factor was correlated negatively with the factors subjective demands and simple demands. simple and subjective demands were correlated positively (see table 3). table 2. characteristics of the scales on perception of demands highest loading item (factor loading) cronbach‘s ɑ m (sd) number of items simple demands “it is clear to teachers how they have to fulfill their task” (.692) .716 2.51 (0.56) 9 subjective demands “teachers with a good personality don’t have to think about their teaching” (.755) .710 2.45 (1.04) 8 complex demands “when planning a lesson, there are a lot of aspects that have to be considered” (.711) .700 4.65 (0.53) 6 wegner et al. 56 | f l r table 3. inter-correlation of the three factors representing perceptions of teaching demands in general 1 2 3 1 simple demands 1 .40** -.48** 2 subjective demands 1 -.20* 3 complex demands 1 note: ** p < .01, * p < .05 4.2 medium correlation between perception of demands in teaching and epistemological beliefs (hypothesis a) to determine the relation between the general perception of demands and epistemological beliefs, we calculated the two factors of the caeb according to stahl and bromme (2007), as well as the two scales on general epistemological beliefs (―relativism― and ―dualism―) according to trautwein and lüdtke (2007). epistemological beliefs in the domain of pedagogy were correlated with perceptions of demands in teaching (table 4). the factor of ―texture‖ correlated negatively with perception of demands as simple, and positively with the perceptions of demands as complex. this means that persons who perceived knowledge in pedagogy as rather well structured were also likely to perceive demands as simple. the factor of ―stability‖ was positively correlated with perceptions of the demands as simple, and negatively with demands as complex. also, general epistemological beliefs that knowledge is stable and simple were also correlated with perception of teaching as simple. all correlations were significant, but only at a small to medium degree. taken together, these results support our hypothesis that epistemological beliefs are related to the perception of demands in teaching, but that the perception of demands in teaching is a separate construct. nevertheless, there was no difference between domain-specific and general epistemological beliefs. 4.3 general perception of demands and situation specificity of judgments of teaching demands (hypothesis b) 4.3.1 situation specific aspects next we analyzed how the perception of demands in general was related to judgment of specific teaching situations. for this purpose, we calculated in a first step a general measure across all kinds of situations. because for each scenario, participants had to rate the same four strategies on a 6-point scale, we calculated means for each strategy across the eight scenarios. the strategy of reflective decision making was rated highest, whereas simple decision making received the lowest values (see table 5), indicating that students were in general aware of the dilemmatic content of the decisions. the scores were correlated systematically: simple decisions correlated positively with clear rules and negatively with own style and reflective decision making. own style also correlated negatively with clear rules and positively with reflective decision making (see table 6). we could not find any differences in regards to demographic measures or field experience. internal consistency over the scenarios was low to medium, ranging from .449 (reflective decision making) to .691 (clear rules). this indicates that there is some consistency across the situations, but also a situation-specific component in the judgment of teaching situations. wegner et al. 57 | f l r table 4. correlation between epistemological beliefs and general perception of demands domain specific epistemological beliefs: pedagogy general epistemological beliefs texture stability relativism dualism m (sd) 4.29 (.75) 3.80 (.43) 1.89 (.47) 1.92 (.48) simple demands -.24** .29** .27** .30** subjective demands -.04 -.00 .13 .27* complex demands .29** -.32** -.05 -.18 table 5. characteristics of the scales on judgment of different teaching situations scale prototypic statement m (sd) min max cronbach‘s ɑ simple decisions ―it is absolutely clear what is right‖ 2.67 (.68) 1.00 4.71 .618 clear rules ―there should be clear rules what to do in this situation.‖ 3.54 (.81) 1.29 5.71 .691 own style ―that is just a matter of opinion. you‘ve got to develop your own style.‖ 3.83 (.65) 2.00 5.86 .635 reflective decision making ―one needs to weigh the options carefully.‖ 4.56 (.60) 2.71 6.00 .449 table 6. intercorrelation between the four scales 1 2 3 4 1 simple decisions 1 .33 ** -.25 ** -.42 ** 2 clear rules 1 -.26 ** -.05 3 own style 1 .45 ** 4 reflective decision making 1 note: ** p < .01, * p < .05; next we compared situations with high and low visibility of the dilemma. judgments differed between scenarios in which the decision had consequences for one instance only (weak visibility of the dilemma), and scenarios in which the decision had further reaching consequences as well as consequences for other teachers (high visibility of the dilemma, see fig. 3). we analyzed differences between both kinds of scenarios by four one-factorial anovas with repeated measurement, with weak vs. high visibility of the dilemma as within-subject factor and the strategy under consideration as the dependent measure. wegner et al. 58 | f l r dilemmas with weak visibility were rated to a higher degree as simple decisions than those with high visibility, f(122, 1) = 7.959, p=. 01, partial η2= .062, but there were no differences between the two kinds of dilemmas in regards to the rating of reflective decision making, f(122, 1) = 1.997, p=. 160, ns, partial η2= .016. however, in situations in which the dilemmatic content was highly visible, because rather large consequences or consequences for other teachers were to be expected, clear rules were rated higher than in dilemmas with weak visibility. for the development of an own style, the result was reversed: for the dilemmas with low visibility, the development of an own style was rated higher than for the dilemmas with high visibility (difference between the items for clear rules: f(122,1) = 121.062, p = .000, partial η2= .50; for own style: f(122,1) = 74.547, p = .000, partial η2= .38). this seems adequate to the situations, because the highly visible dilemmas contained scenarios with consequences for others, which clear rules might help to minimize. figure 3. mean ratings of teaching situations in regards to decisions with individual and with school-wide consequences. 4.3.2. relation of the general perception of demands to the judgment of specific situations we analyzed how the general perception of demands related to the judgment of specific situations. we found systematic relations (see table 7). perceiving general demands in teaching as simple was associated with positive judgment of simple decisions in specific situations. general perception of demands as subjective was related mildly to simple decisions as well as to developing one‘s own style in teaching. interestingly, perception of demands as complex was associated most strongly with a positive appreciation for the establishment of rules, and only mildly with the strategy of reflective decision making. this indicates that students might wish for a reduction in the complexity of situations. we checked whether the correlation patterns were different for situations with weak and with high visibility of the dilemma. in both types of situations, simple demands were correlated significantly with simple decisions, and complex demands with the establishment of clear rules. only for situations in which the dilemma was highly visible, positive correlations between subjective demands and simple decisions as well as with the development of an own style, and negative correlations with the establishment of clear rules were significant. from this pattern of results, we can conclude that general perceptions of demands do influence the judgment of teaching situations, but there also is a situation-specific component. the influence of general perceptions of demands seems to be somewhat stronger for situations in which consequences are to be expected for other teachers, that is, for situations in which the content is experienced as particularly dilemmatic. taken together, (a) the medium internal consistency of the four scales on judgment of specific teaching situations, (b) the differences between dilemmas with high and with weak visibility, and (c) the wegner et al. 59 | f l r medium correlation of general perception of demands with the judgment of the teaching situation can be seen as an indicator that there is both a personal as well as a situation-specific component of perception of demands, thus confirming hypothesis b. table 7 correlations between perception of demands and judgment of teaching situations (n=122). weak = scenarios with weak visibility of the dilemma, high = scenarios with high visibility of the dilemma, all = all scenarios. simple demands subjective demands complex demands weak all high weak all high weak all high simple decisions .23 * .29 ** .26 ** .10 .23 ** .28 ** -.02 -.04 -.01 clear rules .10 -.03 -.14 -.00 -.15 -.26 ** .24 ** .36 ** .37 ** own style -.03 -.03 -.03 .17 .23 ** .22 * .14 .07 -.00 reflective decision making -.06 -.15 -.18 * -.07 -.13 -.14 .19 * .15 .07 note: ** p < .01, * p < .05; 4.4 relation of judgment of teaching situations to epistemological beliefs (hypothesis c) last, we checked the relation between epistemological beliefs and the judgment of teaching situations. we only found small or no correlations between judgment of specific situations and general epistemological beliefs or epistemological beliefs in the domain of pedagogy (table 8). as with the other scales, neither gender, nor subject of study, nor field experience as teacher had an impact on the epistemological beliefs. again, for the highly visible dilemmas the relationship between epistemological beliefs was more pronounced than for dilemmas with weak visibilty. this indicates that epistemological beliefs have only a minor influence on the judgment of specific teaching situations. table 8. means and sd for the epistemological beliefs. correlations between epistemological beliefs and perception of demands as well as strategies domain specific epistemological beliefs: pedagogy general epistemological beliefs texture stability relativism dualism m (sd) 4.29 (.75) 3.80 (.43) 1.89 (.47) 1.92 (.48) simple decisions -.19 * .07 .15 .19 * clear rules -.11 .02 .10 .04 own style .18 -.16 -.02 .02 reflective decision making .15 -.01 .03 -.02 note: ** p < .01, * p < .05 wegner et al. 60 | f l r 5. discussion in our study, we examined how teacher students perceive demands in teaching in general and in specific situations, and how these perceptions relate to epistemological beliefs in general and in the domain of pedagogy. epistemological beliefs correlated with perception of demands in teaching in general, but only mildly with judgment of specific teaching situations. general perception of demands in teaching influenced the judgment of specific situations, especially in situations in which the dilemma was highly visible. also, there was a situation-specific component in the judgment of situations, as indicated by the medium to low internal consistency of the judgments across all teaching situations. taken together, the results can be interpreted in such a way that epistemological beliefs shape the general perception of demands, and that the general perception of demands shapes the way different teaching situations are judged. the influence is especially strong in situations in which the dilemma is especially visible. therefore, it is important to help teacher students to understand the dilemmatic content of specific situations as well as to develop a differentiated perspective on teaching in general. however, these results are based on correlations and cannot be interpreted as causal relations. longitudinal designs are needed to further support our hypothesis. generally, teacher students showed a high awareness for dilemmatic demands. in regards to specific situations, reflective decision making was rated as the best way to deal with the situation, whereas simple decisions received the lowest rating. links between perceptions of demands in general with the judgment of teaching situations yielded an interesting pattern. for both kinds of scenarios (weakly vs. highly visible dilemmas), general perceptions of the demands as simple were related to judgment of dilemmatic situations as simple, but also a complex representation of the demands in teaching led to a positive judgment of the establishment of clear rules. this was interesting, because rules can help to reduce the complexity of teaching (e.g. koedinger, booth & klahr, 2013). we conclude that especially teacher students who experience teaching as a very complex task wish to be supported in difficult teaching situations by clear directions for dealing with the situation. however, this can be problematic, because rules can prevent teachers from acting deliberately and reflectively in such situations (lampert, 1985). therefore, teacher students should be prevented from thinking about rules in the form of a rigid, technological perspective, but rather be supported as thinking of them as a guideline or heuristic. generally, the ability to deal with contradicting demands is one of the core competences of teachers (e.g. berlak & berlak, 1981; labaree, 2000) that has received little attention by empirical researchers. the present study gives first insights into student teachers‘ perceptions of demands in teaching. because the results are only based on self-reports in questionnaires, we cannot make any inferences about actual decision making in dilemmatic situations. however, the study is a first step in the exploration of teachers‘ ability in dealing with this kind of demands. research on teachers‘ dealing with contradictory demands should therefore be put on the research agenda. future research should be especially directed to the question of how this ability can be fostered and which kind of interventions are most helpful in making teacher students aware that teaching is not merely a question of heuristics and simple answers, but that the challenge in teaching is to manage dilemmas by reflected decision making (lampert, 1985; nückles & wegner, 2013). keypoints teachers have to realize that demands in teaching are contradictory in order to be able to make reflected decisions perception of demands in teaching has a situation-specific as well as a general component, and is related to epistemological belief, especially in situations with strong dilemmatic content. perception of demands as complex promotes reflective judgment of teaching situations, but also the wish for implementing clear rules for everyone. wegner et al. 61 | f l r references albanese, m. a., & mitchell, s. (1993). problem-based learning: a review of literature on its outcomes and implementation issues. academic medicine, 68(1), 52-81. doi:10.1097/00001888-199301000-00012 ball, d. l. (1993). with an eye on the mathematical horizon: dilemmas of teaching elementary school mathematics. the elementary school journal, 93(4), 373-397. barcelos, a. m. f. (2001). the interaction between students' beliefs and teachers' beliefs and dilemmas. in b. johnston & s. irujo (eds.), research and practice in language teacher education. selected papers from the first international on language teacher education (pp. 69-86). minneapolis: university of minnesota. berlak, a., & berlak, h. (1981). dilemmas of schooling: teaching and social change: london: routledge. berry, a. (2007). reconceptualizing teacher educator knowledge as tensions: exploring the tension between valuing and reconstructing experience. studying teacher education, 3(2), 117-134. doi: 10.1080/17425960701656510 bräu, k. (2008). die betreuung selbstständigen lernens—vom umgang mit antinomien und dilemmata. [supporting self-regulated learning – of dealing with antinomies and dilemmas].in: breidenstein, g. & schütze f.: paradoxien in der reform der schule. ergebnisse qualitativer sozialforschung. [paradoxes of the school reform. results of qualitative research.] (pp. 179-199) wiesbaden: vs verlag für sozialwissenschaften. brodie, k. (2010). pressing dilemmas: meaning-making and justification in mathematics teaching. journal of curriculum studies, 42(1), 27-50. doi:10.1080/00220270903149873 brookhart, s. m. (1994). teachers' grading: practice and theory. applied measurement in education, 7(4), 279-301. doi:10.1207/s15324818ame0704_2 buehl, m. m., alexander, p. a., & murphy, p. k. (2002). beliefs about schooled knowledge: domain specific or domain general?. contemporary educational psychology, 27(3), 415-449. doi: 10.1006/ceps.2001.1103 calderhead, j. (1989). reflective teaching and teacher education. teaching and teacher education, 5(1), 4351. doi: 10.1016/0742-051x(89)90018-8 cuban, l. (1992). managing dilemmas while building professional communities. educational researcher, 21(1), 4-11. doi:10.3102/0013189x021001004 dann, h.-d. (2008). lehrerkognitionen und handlungsentscheidungen. [teachers‗ cognitions & decision making]. in: schweer, m. k. w. (ed.): lehrer-schüler-interaktion. inhaltsfelder, forschungsperspektiven und methodische zugänge [teacher-student-interaktion. content areas, perspectives for research and methodology] (pp. 177-207). opladen: leske und budrich deci, e.l. & ryan, r.m. (1985). intrinsic motivation and self-determination in human behavior. new york: plenum press. fenstermacher, g. (1994). the knower and the known: the nature of knowledge in research on teaching. review of research in education, 20, 3-56. geddis, a. n., & wood, e. (1997). transforming subject matter and managing dilemmas: a case study in teacher education. teaching and teacher education, 13(6), 611-626.doi: 10.1016/s0742051x(97)80004-2 hager, p., gonczi, a., & athanasou, j. (1994). general issues about assessment of competence. assessment & evaluation in higher education, 19(1), 3-16.doi: 10.1080/0260293940190101 harrington, h. l. (1995). fostering reasoned decisions: case-based pedagogy and the professional development of teachers. teaching and teacher education, 11(3), 203-214. doi: 10.1016/0742051x(94)00027-4 hatton, n., & smith, d. (1995). reflection in teacher education: towards definition and implementation. teaching and teacher education, 11(1), 33-49. doi: 10.1016/0742-051x(94)00012-u helsper, w. (2004). antinomien, widersprüche, paradoxien: lehrerarbeit – ein unmögliches geschäft? [antinomies, contradiction, paradoxes: working as a teacher – an impossible job?] in b. kochpriewe, kolbe fritz-ulrich, & j. wildt (eds.), grundlagenforschung und mikrodidaktische reformansätze zur lehrerbildung [fundamental research and microdidactic reform ideas in teacher education] (pp. 49-98). bad heilbrunn/obb.: klinkhardt. wegner et al. 62 | f l r hofer, b. k., & pintrich, p. r. (1997). the development of epistemological theories: beliefs about knowledge and knowing and their relation to learning. review of educational research, 67(1), 88140. doi: 10.3102/00346543067001088 hofer, b. k., & sinatra, g. m. (2010). epistemology, metacognition, and self-regulation: musings on an emerging field. metacognition and learning, 5(1), 113-120. doi: 10.1007/s11409-009-9051-7 king, p. m., & kitchener, k. s. (1994). developing reflective judgment: understanding and promoting intellectual growth and critical thinking in adolescents and adults. jossey-bass: san francisco. kitchener, k. s. (1983). cognition, metacognition and epistemic cognition: a three-level model of cognitive processing. human development, 26(4), 222–232. doi: 10.1159/000272885 koedinger, k. r., & aleven, v. (2007). exploring the assistance dilemma in experiments with cognitive tutors. educational psychology review, 19(3), 239-264. doi: 10.1007/s10648-007-9049-0 koedinger, k. r., booth, j. l., & klahr, d. (2013). instructional complexity and the science to constrain it. science, 342(6161), 935-937. doi: 10.1126/science.1238056 krettenauer, t. (2005). die erfassung des entwicklungsniveaus epistemologischer überzeugungen und das problem der übertragbarkeit von interviewverfahren in standardisierte fragebogenmethoden. [measuring the developmental level of epistemological beliefs and the problem of transfering interview procedures to standardized questionnaire methods] zeitschrift für entwicklungspsychologie und pädagogische psychologie, 37(2), 69–79. doi: 10.1026/00498637.37.2.69 kuhn, d. (1991). the skill of argument. cambridge: cambridge university press. kuhn, d., cheney, r., & weinstock, m. (2000). the development of epistemological understanding. cognitive development, 15(3), 309–328.doi: 10.1016/s0885-2014(00)00030-7 labaree, d. f. (2000). on the nature of teaching and teacher education: difficult practices that look easy. journal of teacher education, 51(3), 228-233. doi:10.1177/0022487100051003011 lampert, m. (1985). how do teachers manage to teach? perspectives on problems in practice. harvard educational review, 55(2), 178-195. levin, b. b. (2002). dilemma-based cases written by preservice elementary teacher candidates: an analysis of process and content. teaching education, 13(2), 203-218. doi: 10.1080/1047621022000007585 muis, k. r., bendixen, l. d., & haerle, f. c. (2006). domain-generality and domain-specificity in personal epistemology research: philosophical and empirical reflections in the development of a theoretical framework. educational psychology review, 18(1), 3-54.doi: 10.1007/s10648-006-9003-6 nespor, j. (1987). the role of beliefs in the practice of teaching. journal of curriculum studies, 19(4), 317–328. doi: 10.1080/0022027870190403 osborne, m. d. (1997). balancing individual and the group: a dilemma for the constructivist teacher. journal of curriculum studies, 29(2), 183-196. doi:10.1080/002202797184125 pearson, p., destefano, l., & garcia, g. (1998). ten dilemmas of performance assessment. in: harrison, c. & salinger, t. (eds), assessing reading: theory and practice (pp. 21-49). london: routledge. renkl, a., mandl, h., & gruber, h. (1996). inert knowledge: analyses and remedies: educational psychologist. educational psychologist, 31(2), 115-121. doi: 10.1207/s15326985ep3102_3 schoen, l. (2005). learning to make sense of the dilemmas of teaching practice: an exploration of preservice teachers. online submission journal citation: ph.d. dissertation. boston: boston college. schommer, m. (1994). synthesizing epistemological belief research: tentative understandings and provocative confusions. educational psychology review, 6(4), 293–319. doi: 10.1007/bf02213418 schön, d. (1983). the reflective practitioner: how professionals think in action: basic books. schwab, j.j. (1964). the structure of disciplines: meanings and significance. in: ford, g.w. & pugno, l. (eds.). the structure of knowledge and the curriculum. chicago: rand mcnally. schwab, r. l., & iwanicki, e. f. (1982). perceived role conflict, role ambiguity, and teacher burnout. educational administration quarterly, 18(1), 60-74. doi:10.1177/0013161x82018001005 stahl, e., & bromme, r. (2007). the caeb: an instrument for measuring connotative aspects of epistemological beliefs. learning and instruction, 17(6), 773–785. doi: 10.1016/j.learninstruc.2007.09.016 wegner et al. 63 | f l r trautwein, u., & lüdtke, o. (2007). epistemological beliefs, school achievement, and college major: a large-scale longitudinal study on the impact of certainty beliefs. contemporary educational psychology, 32(3), 348-366. doi:10.1016/j.cedpsych.2005.11.003 wegner, e.& nückles, m. (2011). die wirkung hochschuldidaktischer weiterbildung auf den umgang mit widersprüchlichen handlungsanforderungen. [impact of professional development on dealing with contradictory demands in teaching]. zeitschrift für hochschulentwicklung, 6 (3), 172-188. windschitl, m. (2002). framing constructivism in practice as the negotiation of dilemmas: an analysis of the conceptual, pedagogical, cultural, and political challenges facing teachers. review of educational research, 72(2), 131-175. doi: 10.3102/00346543072002131 microsoft word ritella et al_publication.docx spe             frontline  learning  research  vol.4  no.  4  special  issue  (2016)  48  -­‐  55   issn  2295-­‐3159       corresponding author: giuseppe ritella, cradle, institute of behavioural sciences, university of helsinki, siltavuorenpenger 5a, p.o. box 9, 00014 university of helsinki. email: gritella@gmail.com doi: http://dx.doi.org/10.14786/flr.v4i4.210 theorizing space-time relations in education: the concept of chronotope giuseppe ritella1a, maria beatrice ligoriob & kai hakkarainena a university of helsinki, finland b university of bari, italy article received 15 september / revised 21 december / accepted 21 december / available online 19 january abstract due to ongoing cultural-historical transformations, the space-time of learning is radically changing, and theoretical conceptualizations are needed to investigate how such evolving space-time frames can function as a ground for learning. in this article, we argue that the concept of chronotope – from greek chronos and topos, meaning time and place/space – lends itself well to reach this aim. in particular, we outline three features of chronotope: 1) its analytical focus includes the examination of the potential interdependency between space and time; 2) it allows us to examine space and time as social constructions, negotiated in dialogical interaction; 3) it involves the analysis of both the material organization and the discursive negotiation of space and time. we use examples from our own studies and from relevant literature to illustrate how these features of the concept allow us to examine the role that space-time relations play in educational practice. finally, we draw our conclusions and briefly introduce the theoretical and methodological challenges to be addressed for a full development of the concept. keywords: space-time, ecological perspectives on learning, co-construction of knowledge, agency, materiality, trans-contextuality, chronotope                                                                                                                             ritella  et  al         | f l r       49   1. introduction this article is framed as a theoretical contribution to a special number of the journal whose purpose is to discuss how existing conceptualizations of educational practices should be redefined to address emerging learning issues. the aim of the article is twofold: 1) to argue for the relevance of analysing spacetime relations as an emerging issue in research on learning, and 2) to propose chronotope as a useful theoretical concept, allowing us to uncover the time-space relations in local investigation sites. the rationale of our argumentation is that we live in an historical moment in which both technological innovations and educational reforms are triggering deep changes in space-time relations in learning. the introduction of continuously evolving virtual spaces and the implementation of pedagogical approaches such as the flipped classroom (flipped learning network, 2014), connected learning (ito et al., 2010) and place-based learning (van eijck & roth, 2010) entail the transformation of the spatial and temporal organization of learning. for example, in a school using the flipped learning approach, lessons can be shared through internet-based software and studied at home, while in the classroom the students can engage in collaborative learning. the time-space organization of learning and teaching change, based on the pedagogical use of technology. historically, such transformations of space-time have been crucial in changing educational practices: “the organization of a serial space was one of the great technical mutations of elementary education. it made it possible to supersede the traditional system (a pupil working for a few minutes with the master, while the rest of the heterogeneous group remained idle and unattended). by assigning individual places it made possible the supervision of each individual and the simultaneous work of all. it organized a new economy of the time of apprenticeship. it made the educational space function like a learning machine, but also as a machine for supervising, hierarchizing, rewarding.” (foucault, 1977, p. 155) accordingly, in order to understand what current transformations imply for learning, we need conceptual and analytical tools that allow examining the space-time relations that emerge in the empirical sites of investigation. we claim that the concept of chronotope can be productively used to reach this aim. crafted from the ancient greek words chronos and topos, meaning time and place/space, chronotope was devised by mikhail bakhtin (1981) both to examine the space-time patterns characterising literary genres and to develop a framework for the cultural analysis of space-time (holquist, 1982). recently, educational research has shown increasing interest in this concept. following bakhtin, this literature is based on the assumption that space and time are interdependent social constructions rather than independent given realities (van eijck & roth, 2010). a literature review of the uses of the concept is beyond the scope of this short article. instead, by drawing on our own studies and on some of the relevant literature, we discuss how chronotope can be adopted for enriching our understanding of learning in the twenty-first century. first, we will briefly introduce the main features of chronotope as a conceptual tool for the analysis of space-time frames. second, we will refer to three significant socio-cultural studies that have addressed issues related to space-time relations. we have selected these three studies because we think that they are valuable contributions to sociocultural research; they are also relevant for developing a theory of space-time in education. in sum, we seek to demonstrate that we can gain further insight about learning processes by carrying out a specific analysis of the (discursive) social negotiation and bodily-material organization of space-time. finally, theoretical and methodological challenges will be discussed. ritella  et  al         | f l r       50   2. the concept of chronotope three features of chronotope make its use fruitful for examining space-time relations: a) its analytical focus includes the examination of the potential interdependency between space and time; b) it allows us to examine space and time as social constructions, negotiated in dialogical interaction; c) it involves the analysis of both the material organization and the discursive negotiation of space and time. in the following paragraphs, we will briefly discuss each feature of chronotope, illustrating through examples how it contributes to make salient situated time/space arrangements. first, space and time are interdependent when considered for the analysis of human action, making it crucial to examine space and time in a coordinated way. a clear example of this interdependence can be found in the different space-time organization of a literature review carried out by means of a search engine such as google books and the same review carried out in a physical library. the absence of a digital search engine implies that one has to take the time to visit the library, consult the physical book records, often organised alphabetically by author names, titles or fields/topics of the books. an alternative is to ask for help from the library staff. all of that must be done before going to find the actual physical books on the shelves and then finding the relevant text within each book. these actions should be carried out in this order, unless one decides to engage in a (most likely ineffective) exploration of the shelves, opening and skimming systematically those books having interesting titles. in contrast, some digital search engines enable a search for books in any location connected with the internet through a full-text search of one or more keywords and obtaining immediate access to the pages where the searched keyword appears. from that point, a researcher has the chance to form an idea of the contents of interest, and then decide whether to read the full text and if so, how carefully. analysing these two cases by using chronotope implies to examine spatial relations (e.g. location of events, spatial arrangements of workspaces, etc.) and temporal relations (e.g. duration, temporally ordered sequences of actions, rate of recurrence of events, etc.) in a coordinated way. thus, it allows recognizing that the introduction of a virtual space containing a digital search engine in the workspace involves a transformation of the entire temporal structure and duration of the activity. the other way around, temporal limitations can affect the selection of tools and the organization of the spaces of the activity. for example, in one of our investigations (ligorio & ritella, 2010), we used the concept of chronotope to analyse how a group of teachers changed their bodily positions around the computer throughout a session of collaborative problem solving. we identified different spatial arrangements of bodies and technology and different tempos of the activity. the analysis showed that one particular spatial arrangement, where all the teachers were standing around the same computer, was realized in order to speed up the accomplishment of the task when the end of the session was approaching. in observing that space and time are interdependent, we do not mean that all transformations of space involve a transformation of time and vice versa. it is possible that some changes in space are neutral with regard to the organization of time and vice versa. for example, changing the spatial organization of the objects on a desk does not necessarily implies changes in the temporal organization of the activity. however, theoretically it is important to acknowledge the existence of such interdependence, as described in the two examples above. it is a goal for researchers to determine under which conditions such interdependence is relevant for educational practice. second, the concept of chronotope was introduced by bakhtin as a concept for the cultural analysis of space and time (holquist, 1981). this implies considering all the different voices involved in social processes, in contrast to the “philosophical monopolization” of scientific discourse, which considers space and time as given realities external to human experience (van eijck & roth, 2010). therefore, the task set by using the concept of chronotope is not just to measure physical spaces and time intervals according to the ritella  et  al         | f l r       51   consolidated scientific paradigm. the scientific understanding of space-time is just one of the possible voices to be considered for understanding space-time relations, and the other voices – for instance, the voices of the participants – should not be silenced in the analysis. thus, there is a need to consider dialogue in the analysis of space-time. in other words, the concept of chronotope is devised to examine space and time as they are socially negotiated in dialogue. indeed, they are not simply ‘given’ to the participants involved in an activity: meanings associated with both space and time are socially negotiated, and the organization of activities in space-time can be highly flexible. van eijk and roth (2013), for example, examined an environmental education project in which students from aboriginal canadian communities engaged in nature conservation practices in a marine park. by using the lens of chronotope, the authors managed to understand how the aboriginal students and the collaborators to the project constructed the material space of the marine park in dialogical interaction. in particular, they illustrated how science education activities carried out during the project involved conflicting notions of space and time derived from the perspective of natural science and from the local culture. analysing these tensions allowed the researchers 1) to identify some contradictions in the project and 2) to trace how the students experienced the science education activities while developing their cultural identities as aboriginals. third, as stated by many authors (hirst & vadeboncouer, 2006; van eijck & roth, 2010; matusov, 2009), chronotope concerns both the immaterial, semiotic, worlds of discourse and narratives, and “patterns of organization of space and time” (lemke, 2004) that are enacted through the movement of bodies and objects. therefore, chronotopic analysis takes into account also the bodily-material aspect of space-time. given such ground, many scholars have used chronotope to investigate space-time at the boundary between material and discursive processes. for example, brown and reshaw (2006) discuss how students express their agency by actively shaping the space-time contexts of the classroom, drawing on past, present and future temporal relations though discursive interaction. in particular, they analysed: 1) how a student initially built her private space-time within the classroom by using a library shelf and a desk, and in a second moment removed those barriers to actively enter into the shared space-time of collaborative activity; 2) how the students discursively shifted between different space-times while explaining and justifying their ideas and developing their identities. another good example is the study by hist and vadeboncoeur (2006). here the authors analyse how the re-engagement in schooling by a dropout student was mediated by the construction of a dynamic spatial network, which involved movements between material spaces (a learning centre and the student’s home) and the concomitant reframing of the student-teacher relationship. in both these studies, the use of the concept of chronotope allowed to uncover how space-time arrangements affected the learning processes. in the present article, chronotopes are defined as “socially emergent” (sawyer, 2005) units of spacetime, where both discursive and material aspects of space-time relations are considered. in line with the literature, human cognition and learning are not conceived as located within the boundaries of the mind, but are distributed in the space-time context of the activity. contexts – including space-time – emerge from a continuous process of social negotiation engaged in by learners (bateson, 1972; cole, 1996; duranti & goodwin, 1992). during learning activities, participants individually or jointly attend to various physical and symbolic spaces: they organise their workspace, co-ordinate their efforts, and perceive space-time constraints and opportunities related to the technological tools used and to the institutional regulation of space and time. following bakhtin, we consider these spatial and temporal processes to be fused, requiring a co-ordinated analysis. examining only space or only time could bias our understanding, given the reciprocal impact they can have on each other. ritella  et  al         | f l r       52   3. sketching a conceptualization of chronotope for 21st-century learning   as argued above, chronotope is a concept specifically intended for the analysis of space-time relations. given the features of chronotope that we have presented above, the goal we set for chronotopic analysis is to investigate: 1) how patterns of space-time organization are involved in different learning activities, different schools systems, different pedagogical approaches; 2) how participants make sense of space-time patterns in dialogical interaction; 3) how participants’ discursive negotiation of space-time is related to their bodily-material organization of space and time. this topic is not new to education. space and time are ubiquitous categories of human experience, and this is reflected in the research literature. indeed, most studies in the learning sciences involve references to space and/or time co-ordinates, although it is rare to find studies that specifically address space and time, or better put, space-time, as the primary focus of investigation. below, we will discuss how specific attention to space-time attained through the concept of chronotope can enrich and extend the knowledge derived from three investigations in which the categories of space and time were relevant. first, silseth and arnseth (2011; 2015) provide an interesting example by examining learning across sites from a dialogical perspective. the authors analyse how ideas and perspectives that emerge in out-ofschool situations, and external representations produced in the past, are mobilized in situated interaction. these resources contribute to creating connections between different situations of learning. space and time are highly relevant for this topic of investigation, which concerns learning processes taking place in multiple locations across extended periods of time. however, silseth and arnseth seem to consider space and time as a background against which the learning takes place, failing to consider them as analytical foci. we argue that the role played by the organization of space-time in this process can be examined by using the concept of chronotope. for example, in a previous case study (ritella & ligorio, 2016) we studied the collaborative sensemaking of a group of professionals working on the design of a web-platform. we discussed how the space-time organization of the activities that took place before face-to-face meetings of the group is connected with the emergence of ideas and viewpoints during subsequent meetings. we found that providing an online space for writing individual notes foregrounded the emergence of the personal perspectives of the participants during the upcoming meeting. in contrast, arranging a physical meeting resulted in the emergence of a collective perspective by a subgroup of participants. thus, we argued that chronotope helped to uncover how the space-time organization of activities affects sensemaking across multiple locations and extended periods of time. second, engle (2006) shows that the way teachers discursively frame the context of learning, including the definition of temporal boundaries, affects students’ transfer of knowledge across different contexts. in particular, the author analyses how a teacher framed learning episodes as building on previous ones or as relevant for the future. the continuous references to the past and the future helped the students to consider these episodes as interconnected, thus supporting the transfer of knowledge across contexts. the author states that space is also relevant for analysing the framing of context, and further research is needed to understand its role. the concept of chronotope, we argue, could be employed to direct analytical attention to both space and time relations, and analyse them in a co-ordinated way, adding further insights. ryan (2011), for example, used the concept of chronotope to examine how students discursively conceive the space-time of university. this study shows that space and time jointly contribute to define the students’ orientation towards academic life. in particular, some students depicted the university as a site of mass education with large lecture theatres, no permanent space for student groups and limited time for individual meetings with teachers because of the busy life of the academic staff. put together, time limits and spatial arrangements of university buildings generated a conception of the university as a potentially distant service provider, which encouraged students to spend most of their time off-campus. thus, using the concept of chronotope allowed detecting how the discursive construction of space-time can have an effect on the students’ academic practices. ritella  et  al         | f l r       53   a last example is the research by jornet and roth (2015), who discussed how students make sense of multiple material representations of scientific phenomena across time. in particular, the authors trace the students’ bodily and pragmatic actions in interaction with representations of scientific phenomena, and discuss how these relate to the students’ experiences and interpretations. we recognise that jornet and roth significantly discuss such a relationship by means of their analysis. however, we claim that this topic of investigation requires also a specific analysis of space-time relations that was missing in their study. the material representations used by the students are distributed in space and they are picked up at different times. conceptualizing space-time relations as chronotopes might allow researchers to uncover how the students spatially arrange their bodies around, and attend to, multiple representations at different times. in a previous study (ritella, ligorio & hakkarainen, 2015), we have used the concept of chronotope to examine how a group of teachers managed the resources available in the context during a session of collaborative problem solving (ps). we interpreted, diachronically, the alternation of 1) events in which the participants explored the space and actively searched for resources in the environment, and 2) events characterized by a focus on a stable set of resources. one of the findings of this study was that in some phases of ps the first type of events was dominant. these moments were associated with: (a) the introduction of a new task, (b) the use of a software suite not yet mastered by the teachers or (c) a change in the configuration of participation (i.e. the participants physically changed their positions in the room, or changed the set of tools used) realised in conjunction with a difficulty in trying to solve the problem. we expect that patterns of this kind could be found also while analysing the students’ use of multiple representations. for example, it could be found that material representations play a crucial role in some phases of educational activities and/or that the way in which they are spatially organized affects their usage by the students. thus, a better understanding of when the students explore the space around them and when they attend to each representation could give us further insights on how they interact with the material environment during learning practices. in sum, in this section we have shown that a specific focus on space-time can yield additional insights for the analysis of learning. for instance, we could see how learning may be affected by the psychical organization of the space within which students interact and how the perception of time constrains influences how the task is perceived. as we will further discuss in the next section, the concept of chronotope has great potentials for examining learning practices as they unfold in space and time, even though some challenges have still to be tackled for the full deployment of the concept. 4. implications and challenges for chronotopic analysis in this paper, we argue that the concept of chronotope can enrich the (dialogical) understanding of learning practice. we have briefly presented this concept and discussed how it could be used to enhance and extend the findings of investigations that implicitly or explicitly address space and time relations in learning. we define chronotope as the emergent configuration of temporal and spatial relations in educational practices. to provide some examples of how chronotopes can be fruitfully used to analyse learning practices, we have discussed how the discursive/bodily/material organization of space-time is connected to: a) the learning processes taking place at multiple locations across extended periods of time, as theorised by silseth and arnseth (2011); b) the participants’ discursive framing of situations of learning, as outlined by engle (2006); c) the use of multiple representations during a learning activity, as examined by jornet and roth (2010). we believe that the range of applications of chronotope extends far beyond the processes we have discussed here. surely, considering how space-time is organised by participants may be useful in designing learning tasks, especially when technology is involved. indeed, contemporary digital environments offer multiple types of resources; (re)organising them in time and space during an activity is not a trivial task. a ritella  et  al         | f l r       54   complex orchestration is needed to use all resources effectively. the same applies in designing training situations for teachers. the appropriation of technology from the teachers’ side is not only a matter of understanding the technical features. developing awareness of how the time-space is transformed by technology may help teachers in improving their educational practices. the changes that technology introduces go beyond the local classroom situation. as in a cascade effect, these changes ultimately cause modification in the larger society. renshaw (2014) pointed out that the school systems can be characterised by different chronotopes in different historical periods. we are aware that the conceptualization sketched in this paper is not yet fully developed and has some limitations. one of them is that current methodological tools are not yet adequate to grasp the organization of space-time in all its complexity. indeed, the organization of space and time applies to different units of analysis. both micro-processes, such as the situational co-ordination of a group of students, and macro-level processes, such as the historical development of school systems, involve patterns of organization of space and time. this suggests that the operationalization of chronotope and the methods used should be adapted to the unit of analysis in each investigation. moreover, another challenge is that the negotiation of space and time can often be implicit and difficult to detect. in one of our studies (ritella, ligorio & hakkarainen, accepted), while negotiating the meaning of a task set by teachers, the students broadly discussed the negotiation of time, but the discussion about space was marginal during the observed interaction. one possible interpretation was that space had been taken for granted and did not emerge clearly in the students’ discourse. however, this could be attributed also to a methodological limitation. the students may have engaged in discussions about space during breaks or informal meetings, when the researcher did not observe the interactions. therefore, a more comprehensive research design should be planned, able to make explicit the conceptions of space and time while preserving the privacy of the participants. based on the discussion here presented, we argue that continuing to pursue this frontline concept is important for advancing our understanding of contemporary learning practices because space-time relations have undergone profound transformations. acknowledging the ongoing transformations of space and time in education involves theoretical and methodological challenges to research. we believe that the concept of chronotope, thanks to its focus on space and time as interconnected social constructions, which are discursively negotiated and bodily-materially enacted by participants, lends itself well to addressing these challenges. aknowledgments the authors are grateful to feldia loperfido and antti rajala for their comments on a previous version and to the editors and reviewers for constructive criticism that have allowed the development of a stronger article. references bakhtin, m. (1981). the dialogic imagination. four essays by m. m. bakhtin. austin: university of texas press. bateson, g. (1972). steps to an ecology of mind: collected essays in anthropology, psychiatry, evolution, and epistemology. chicago: the chicago university press. brown, r., & renshaw, p. (2006). positioning students as actors and authors: a chronotopic analysis of collaborative learning activities. mind, culture and activity, 13(3), 244–256. doi: 10.1207/s15327884mca1303_6 cole, m. (1996). cultural psychology: a once and future discipline. cambridge: harvard university press. ritella  et  al         | f l r       55   duranti, a., & goodwin, c. (eds.). (1992). rethinking context: language as an interactive phenomenon. cambridge: cambridge university press. van eijck, m., & roth, w. m. (2010). towards a chronotopic theory of “place” in place-based education. cultural studies of science education, 5(4), 869-898. doi: 10.1007/s11422-010-9278-2 van eijck, m., & roth, w. m. (2013). place and chronotope. in imagination of science in education (pp. 133-162). springer netherlands. engle, r. a. (2006). framing interactions to foster generative learning: a situative explanation of transfer in a community of learners classroom. the journal of the learning sciences, 15(4), 451-498. doi: 10.1207/s15327809jls1504_2 flipped learning network (fln). (2014) the four pillars of f-l-i-p™ foucault, m. (1977) discipline and punish: the birth of the prison. vintage. hirst, e., & vadeboncoeur, j. a. (2006). patrolling the borders of otherness: dis/placed identity positions for teachers and students in schooled spaces. mind, culture, and activity, 13(3), 205-227. doi: 10.1207/s15327884mca1303_4 holquist, m. (1982). bakhtin and rabelais: theory as praxis. boundary 2, 5-19. ito, m., gutiérrez, k., livingstone, s., penuel, b., rhodes, j., salen, k., & watkins, s. c. (2013). connected learning: an agenda for research and design. irvine, ca: digital media and learning research hub. jornet, a., & roth, w. m. (2015). the joint work of connecting multiple (re) presentations in science classrooms. science education, 99(2), 378-403. doi: 10.1002/sce.21150 lemke, j. (2004). learning across multiple places and their chronotopes. in aera 2004 symposium, april (pp. 12-16). ligorio, m. b., & ritella, g. (2010). the collaborative construction of chronotopes during computersupported collaborative professional tasks. international journal of computer-supported collaborative learning, 5(4), 433-452. doi: 10.1007/s11412-010-9094-4 matusov, e. (2009). journey into dialogic pedagogy. nova science publishers. renshaw, peter d. (2013). classroom chronotopes privileged by contemporary educational policy: teaching and learning in testing times. in phillipson, s., kelly y. l. ku and shane n. phillipson (ed.), constructing educational achievement: a sociocultural perspective (pp. 57-69). oxon, uk: routledge. ritella, g., ligorio, m. b., & hakkarainen, k. (2015). the role of context in a collaborative problem-solving task during professional development. technology, pedagogy and education, 25(3), 395-412. doi: 10.1080/1475939x.2015.1062412 ritella, g., & ligorio, m. b. (2016). investigating chronotopes to advance a dialogical theory of collaborative sensemaking. culture & psychology, 22(2), 216-231. doi: 10.1177/1354067x15621475 ritella, g., ligorio, m. b., & hakkarainen, k. (accepted) interconnections between the discursive negotiation of space-time and the interpretation of a collaborative task. learning, culture and social interaction. ryan, m. (2011). productions of space: civic participation of young people at university. british educational research journal, 37(6), 1015-1031. doi: 10.1080/01411926.2010.517827 sawyer, r. k. (2005). emergence: societies as complex systems. cambridge: cambridge university press. silseth, k., & arnseth, h. c. (2011). learning and identity construction across sites: a dialogical approach to analysing the construction of learning selves.culture & psychology, 17(1), 65-80. doi: 10.1177/1354067x10388842 silseth, k., & arnseth, h. c. (2015). frames for learning science: analyzing learner positioning in a technology-enhanced science project. learning, media and technology, 1-20. doi: 10.1080/17439884.2015.1100636     microsoft word iiskala et al_publication.docx ! ! ! ! ! ! frontline learning research vol. 3 no. 1 (2015) 78 111 issn 2295-3159 ! corresponding author: tuike iiskala, assistentinkatu 5, 20014 university of turku, finland, tuike.iiskala@utu.fi doi: http://dx.doi.org/10.14786/flr.v3i1.159 ! ! socially shared metacognitive regulation in asynchronous cscl in science: functions, evolution and participation a tuike iiskala, b simone volet, c erno lehtinen, & d marja vauras a, c, d department of teacher education & centre for learning research, university of turku, fin-20014, finland b school of education, murdoch university, 6150 murdoch, australia article received 12 january 2015 / revised 23 march 2015 / accepted 5 may 2015 / available online 20 may 2015 abstract the significance of socially shared metacognitive regulation (ssmr) in collaborative learning is gaining momentum. to date, however, there is still a paucity of research of how ssmr is manifested in asynchronous computer-supported collaborative learning (cscl), and hardly any systematic investigation of ssmr’s functions and evolution across different phases of complex collaborative learning activities. furthermore, how individual students influence group regulatory effort is not well known and even less how they participate in ssmr over the entire collaborative learning process. the multimethod, in-depth case study presented in this article addresses these gaps by scrutinizing the participation of a small group of students in ssmr in asynchronous computer supported collaborative inquiry learning. the networked discussion, consisting of 640 notes, was used as baseline data. the sets of notes, which formed nine ssmr threads, were identified and their functions analyzed. several analytical methods, including social network analysis, were used to investigate various aspects of individual participation. the findings show that some ssmr threads lasted over an extended period, and they sometimes intertwined or overlapped. furthermore, ssmr threads were found to play different functions, mainly inhibiting the perceived inappropriate direction of the ongoing cognitive process. finally, ssmr was found in all phases of the process – but with some variation. the use of different analytical methods was critical as this provided a variety of complementary insights into students’ participation in ssmr. the value of using multiple, rigorous analytical methods to understand ssmr’s significance over the entire course of an asynchronous cscl activity is discussed. keywords: socially shared metacognitive regulation; collaborative processes; asynchronous cscl; inquiry learning; multi-method approach iiskala'et'al' ' ' | f l r ! ! 79! 1. introduction over the last decade, there has been growing shift in the study of learning as a socially and culturally embedded (individual or collaborative) activity (salomon & perkins, 1998). this has led many researchers to argue that a focus on individual regulation of learning is insufficient for understanding learning that takes place in social contexts and, in particular, collaborative learning environments (e.g. efklides, 2008; grau & whitebread, 2012; hogan, 2001; järvelä & hadwin, 2013; molenaar & chiu, 2014). on the grounds that collaborative learning environments bring together “multiple self-regulated agents [who] socially regulate each other’s learning” (volet, summers, & thurman, 2009a, p. 129) and at times engage in genuinely shared regulatory processes (salonen, vauras, & efklides, 2005), research on social regulation in collaborative learning needs to integrate interpersonal processes and individual cognition (greeno, 2006; vauras & volet, 2013). this approach is critical since, according to winne, hadwin and perry (2013), the quality of learning in collaborative settings is not only dependent on individuals’ distributed regulation but also on the coand shared regulation that emerges in situated interactions. this paper contributes to and extends the growing literature on the significance of social regulation in collaborative learning by presenting the findings of an in-depth, micro-level and multi-method exploratory study of the socially shared metacognitive regulation of one small group of students engaged in an asynchronous, computer-supported collaborative inquiry learning process. although research on computersupported collaborative learning (cscl) is extensive, there is still a paucity of fine-grained empirical studies on the manifestations, functions, and evolution of socially shared metacognitive regulation processes as they unfold over the duration of collaborative inquiry learning. it is argued that this insight is essential to understanding fully the significance of socially shared metacognitive regulation in collaborative learning and the role of individual contributions in the process. this dual focus on group-level, socially shared regulatory processes and individual-level contributions to the group regulatory effort, required a combination of qualitative and social network analytical methods. 1.1 socially shared metacognitive regulation (ssmr) socially shared metacognitive regulation (ssmr thereafter) refers to participants’ goal-directed, consensual, egalitarian and complementary regulation of joint cognitive processes in the collaborative learning context (see iiskala, vauras, & lehtinen, 2004; iiskala, vauras, lehtinen, & salonen, 2011). this definition focuses on the regulation aspects of metacognition (i.e. see brown, 1987, on regulation of cognition; see brown & deloache, 1983, on metacognitive skills) geared towards achieving goals (see wertsch, 1977). this can refer to activities such as identifying task requirements and expectations (e.g. what has to be done); planning (e.g. time allocation); keeping track of the process and mindfully changing it if needed; monitoring comprehension (e.g. questioning the direction of the cognitive process) or evaluating the quality of the task outcome (see veenman & beishuizen, 2004). hence, in ssmr, students jointly regulate their ongoing cognitive learning process towards the common goal. in contrast, knowledge co-construction involving processing the content knowledge (e.g. gathering additional information or establishing concepts’ scientific meaning) and task co-production involving generating the tangible outcome (e.g. justifying the features of the outcome such as ‘content knowledge x’ causes ‘content knowledge y’ and, hence, they are connected to each other) are not conceptualized as metacognitive regulation but rather cognitive activity (see khosa & volet, 2014). research on ssmr originates from the metacognition research tradition which focuses on regulation of cognition, whereas research on self-regulation and self-regulated learning is further concerned with behavior and motivation regulation (see dinsmore, alexander, & loughlin, 2008). overall, the literature on regulatory processes displays a confusing instability in the use of terms that originate from different research traditions (dinsmore et al., 2008). hence, the term ssmr makes the focus on cognition regulation quite explicit. ssmr provides a solid conceptual basis on which to gain insights into how collaborating dyads or groups regulate their cognitive activity and, at a fine-grained level, their understanding of the specific iiskala'et'al' ' ' | f l r ! ! 80! requirements of tasks and content knowledge. while the concept of ssmr is solidly grounded in the metacognition tradition, the research focus is on the regulation component of metacognition, in other words, the executive function of regulating the cognitive activity (see brown, bransford, ferrara, & campione, 1983; flavell & miller, 1998) and, therefore, on the procedural component of metacognition (veenman, 2011). in turn, the term ‘socially shared’ seeks to capture collective, mutually shared metacognitive regulation that unfolds on the social plane when a group engages in co-constructing knowledge and understanding of the task and its content. in their comprehensive review of different forms of regulation, winne and colleagues (2013) and hadwin, järvelä and miller (2011) defined coand socially shared regulation by describing the regulation of collaborative study processes in general, but they did not elaborate on micro-level, unfolding metacognitive regulation processes involved in problem solving and knowledge construction in real time. to date, empirical ssmr studies are still scarce as compared to the comprehensive research on metacognitive regulation in individual learning. studying social regulation in collaborative learning – where the core unit of analysis is the group’s learning process rather than solely individual’s learning process – presents important conceptual and methodological challenges. a number of researchers have proposed that, in the context of collaborative learning, metacognitive regulation can be best understood as both an individual and a social process, in which participants’ metacognitive regulatory activities are interdependent and often shared (e.g. efklides, 2008; hadwin et al., 2011; iiskala et al., 2004; volet, vauras, & salonen, 2009b). as a consequence, conceptualizations of metacognitive regulation during collaborative learning – and derived analytical approaches in empirical work – must consider selfand socially-regulated systems as operating concurrently (volet et al., 2009b; winne et al., 2013). 1.2 ssmr in face-to-face collaborative learning ssmr manifestations in face-to-face collaborative learning have identified tasks and contexts across age groups. this includes, for example, studies of face-to-face mathematical problem-solving processes in high-achieving primary school student dyads (vauras, iiskala, kajamies, kinnunen, & lehtinen, 2003; iiskala et al., 2004, 2011; volet, vauras, khosa, & iiskala, 2013), young children working in naturalistic educational settings (whitebread, bingham, grau, pino pasternak, & sangster, 2007; whitebread et al., 2009), veterinary science students’ working on authentic clinical cases (volet et al., 2009a; summers & volet, 2010; volet et al., 2013), elementary school student triads’ collaborative learning in different instructional design environments (molenaar, 2011) and primary school students’ collaborative activities in science classes (grau & whitebread, 2012). ssmr has been found to play different functions in the flow of collaborative learning processes. for example, a study of collaborative mathematical problem-solving processes of high-achieving dyads revealed that ssmr either facilitated (through confirming or activating) the perceived appropriate direction in the dyads’ cognitive activities or, alternatively, inhibited (by slowing down, changing or stopping) the perceived inappropriate direction of activities (iiskala et al., 2011). therefore, ssmr’s functions either promote the continuation of appropriate cognitive learning processes or, conversely, prevent the ongoing cognitive process from moving in an adverse direction. ssmr’s overall function, therefore, is to ensure that the group cognitive activity develops in the appropriate direction. as a result, the quality of the ongoing cognitive process is constantly monitored and controlled (e.g. whitebread et al., 2009). understanding ssmr functions at the micro level, therefore, is essential to gaining greater insight into ssmr’s significance in collaborative learning. one under-researched aspect of ssmr in face-to-face collaborative learning is its evolution over an extended period and across phases of problem solving or inquiry learning processes. systematic investigations of the emergence are rare, as are studies of the subsequent evolution over, and the metacognitive regulation of, time and through different phases of collaborative learning processes (see azevedo, 2014; molenaar & järvelä, 2014). for example, grau and whitebread’s study (2012) revealed that iiskala'et'al' ' ' | f l r ! ! 81! individual self-regulated behavior increased within the groups during the process but the number of shared regulation episodes varied across sessions. rogat and linnenbrink-garcia (2011) also found no evidence of a systematic change in regulation quality within small groups of sixth-grade students. lajoie and lu’s (2012) findings for medical students indicated that the timing of metacognitive activities varied so that engagement in these activities early in the problem solving process led to a greater percentage of executive processes in latter sessions, especially in teams supported by technology, as compared to non-technology supported teams. rogat and linnenbrink-garcia (2011) also suggested that sub-processes, such as planning and monitoring, mutually influence each other and overlap; for example, students monitor the group’s understanding of the task directions and plan enactment. these studies’ findings provide preliminary support for the importance of better understanding ssmr manifestations and functions in the context of evolving cognitive learning processes. 1.3 ssmr in asynchronous cscl empirical studies of ssmr in asynchronous cscl environments are still scarce. research conducted in cscl environments have tended to adopt a more general view of socially shared regulation of motivation and learning strategies, paying limited attention to the detailed metacognitive regulatory processes taking place in real time during the course of learning activities but, at the same time, incorporating other forms of regulation. for example, järvelä and hadwin (2013) located their approach within the self-regulated learning tradition, which takes into consideration motivation and emotions regulation. network mediated asynchronous communication is a particularly challenging environment for groups to engage in ssmr because the multiple verbal and non-verbal channels of reciprocal interaction available in face-to-face situations are limited. however, cscl environments provide learners with some new tools, such as structuring supports, mirroring and metacognitive tools, as well as guiding steps that can support reciprocal regulation of learning processes (soller, martinez-monés, jermann, & muehlenbrock, 2005; see also järvelä & hadwin, 2013). a major advantage of studying ssmr in asynchronous cscl is that part of the individuals and the group’s metacognitive thinking become ‘visible’ through computer-based traces of interactions. this means students can return to previous thinking and follow earlier trains of thoughts. hence, some researchers (e.g. hurme, palonen, & järvelä, 2006) have explored how engagement in joint discussion metacognition can vary between participants in secondary school dyadic activities. they found dyads that monitored and evaluated their ongoing discussion achieved a more optimal position in the communication network than other dyads. prinsen, volman and terwel’s (2007) study of primary-school classes, also conducted in an asynchronous cscl environment, revealed evidence of mutual engagement in discussing, monitoring, and evaluating group processes and in instructing fellow students, but limited time was spent regulating the cognitive activities’ direction (6%). the authors interpreted their findings in terms of the task’s wellstructured nature. since ssmr of difficult learning or problem solving can be highly demanding, some students can reasonably be expected to engage more than others do in the social regulation of group cognitive process. this was illustrated in palonen and hakkarainen’s (2000) study, in which they re-analyzed previous data from 11 to 12year-old students’ computer-supported learning, using social network analysis (sna). interestingly, the findings obtained with sna revealed that averageand high-achieving female students dominated social interactions in cscl environments and took on the main responsibility for collaborative knowledge building. this contradicted the researchers’ original qualitative content analysis, which found that the culture of inquiry was rather homogenous across students. sna, therefore, appears particularly well suited to exploring the presence of relationships patterns among interacting units (see wasserman & faust, 1994). this type of investigation has been successfully applied beyond single participants to collective-level analysis in elementary and secondary schools (toikkanen & lipponen, 2011). iiskala'et'al' ' ' | f l r ! ! 82! these researches show that in-depth explorations of within-group differences in ssmr contributions benefit from a combination of analytical methods. for toikkanen and lipponen (2011), combining sna and qualitative analysis can potentially provide more accurate results and a stronger foundation for interpretation. 1.4 ssmr in asynchronous cscl inquiry over time and across phases: a case study group inquiry learning provides a useful research setting to study ssmr in asynchronous cscl because this is a challenging, research-like approach with the goal of processing and deepening explanatory knowledge, rather than merely acquiring factual knowledge (e.g. hakkarainen, 2003; scardamalia & bereiter, 1994). a number of studies have shown that inquiry learning stimulates students’ active engagement in learning processes that require metacognition to monitor understanding (e.g. khosa & volet, 2014; white & frederiksen, 1998). in pedagogically well planned and technology-supported inquiry learning environments, even young students are able to engage in constructive peer interaction and knowledge building processes that go beyond the typical cognitive demands of regular school learning (e.g. hakkarainen, 2003; zhang, scardamalia, reeve, & messina, 2009). hence, some prior research implies that joint engagement in demanding, asynchronous cscl inquiry can trigger ssmr. 1.5 research aims and expectations the first aim of the present study was to explore the manifestation and functions of socially shared metacognitive regulation (ssmr) in an asynchronous cscl environment of a small group’s inquiry in science learning. the second aim was to examine the ssmr’s evolution through different phases of the group’s asynchronous cscl inquiry process. the third aim was to determine how individual group members participate in ssmr during asynchronous cscl inquiry and how specific individuals’ contributions influence the group’s regulatory effort. since the study was exploratory in nature, no predictions could be made about the ssmr’s manifestation and functions in an asynchronous cscl environment of inquiry learning. in light of the high cognitive demand imposed by inquiry learning on school-aged students, evidence of ssmr was expected throughout all phases. however, the ssmr’s functions in different phases could not be predicted. regarding how individual group members would participate in ssmr and how specific individuals’ contributions would influence the group’s regulatory effort, since individuals are self-regulating agents and asynchronous exchange allows time for reflection, some members were expected to play a pivotal regulatory role in initiating and sustaining the process. alternatively, they could inhibit the current direction of the group’s cognitive process. the investigation’s exploratory nature at the individual level (third aim) led to the use of a combination of several methods of analysis. 2. method 2.1 participants this case study focused on a small group of four 12-year-old girls (iina, piia, sonja and julia). the girls were attending an urban finnish primary school with high standards and a good reputation. the class from which the small group was selected, had a language enriched curriculum, and an entrance examination was used to select students for program admission. because of the school’s high criteria for admission, the students could be described as above average in their overall academic skills. however, the students had only some experience of cscl and inquiry learning, and, from this point of view, their usual learning context iiskala'et'al' ' ' | f l r ! ! 83! could be described as traditional. the students’ parents gave written consent for them to take part in the study, and the students themselves were willing to participate. the small group studied was formed at the beginning of the process under the teacher’s guidance, based on the students’ similar interests in the topic of inquiry chosen for the science-based activity. these four students were also disposed to working together. 2.2 collaboration environment the theoretical ideas underpinning the learning process as a collaborative inquiry, as proposed by hakkarainen (2003) (see also brown & campione, 1994, 1996; brown et al., 1993; hakkarainen & sintonen, 2002; scardamalia & bereiter, 1994, 2006), formed the conceptual basis of the present study’s collaborative inquiry environment. in this study, the collaborative inquiry process was applied to the scientific exploration of the universe. the teacher introduced the subject at the process’s beginning, after which each small group was responsible for collecting further scientific information from other sources, such as textbooks, experts and the internet, and working together to develop meaning out of this material. each group worked during the process on complex, ill-defined questions that they had set for themselves at the process’s beginning. the whole process involved five phases: i – setting up the research question; ii – constructing a hypothesis; iii – developing a work plan; iv – searching for and processing knowledge; v – summarizing findings and concluding (see the appendix for a full description). at the beginning of each phase, the students were reminded to collaborate with each other and to regulate their group’s learning process. at the activity’s beginning, students were introduced, in a face-to-face situation, to the expectations for the collaborative inquiry process. they were informed that they needed to work like scientists, and then the process’s different phases were presented. in addition, the students were encouraged to make their thinking visible to each other by sharing their ideas and jointly regulating the group inquiry process. at the start of each phase, the researcher presented the current phase’s aim to the entire class, once again in a faceto-face environment. after that, in each phase, the students collaborated within their small groups in the asynchronous cscl environment workmates (for detailed information about workmates, see nurmela, lehtinen, & palonen, 1999). although the students had access to workmates at all times and, hence, the chance to write notes outside their science class, they collaborated mainly during the time specified by their teacher during class. some discussions were close to synchronous (chat-like) discussions, whereas, in other cases, longer time lags appeared between notes. all the students were acquainted with the workmates environment, as they had used it a few times beforehand. in workmates, students can send written notes to each other, and they can reply to one another’s notes. all the notes sent during a particular phase are visible on the computer screen, and students can also return to previous phases’ notes. data for this study were the four group members’ discussions in this asynchronous cscl environment. 2.3 data analysis to address the aim of exploring students’ ssmr engagement, the analysis unit was a thread, that is, a sequence of interconnected notes. when a student posted a message to the public space where other students in the small group could see it and react, this was counted as a note. a sequence of interconnected notes was classified as an ssmr thread if students, through their notes, demonstrated evidence that they were jointly regulating the progress of their inquiry learning process towards their common goal (e.g. student 1: “we want to address the task’s goal but we don’t yet have enough knowledge of this issue because…”, student 2 (reacting to student 1’s note): “hmm, let’s think about this issue more in detail. so far, we have evidence that … but we still don’t have knowledge about…”, student 3 (reacting to student 1’s and/or student 2’s note(s)): “so, let’s try to fill in these gaps before concluding. first, we have to find out … next, we have to…” the interaction continues in further notes). hence, small group’s learning process was identified through the students’ regulatory acts, whenever a minimum of two students were involved in the iiskala'et'al' ' ' | f l r ! ! 84! process, and the students’ reciprocal notes were interdependent, together affecting the course of the ongoing cognitive process. furthermore, although each thread had to involve a minimum of two notes, no upper limit was set to the number of notes included. consistent with iiskala and colleagues (2004), each note’s meaning, whether it was a part of an ssmr thread or not, was thus dependent on the flow of other preceding or subsequent notes, which means these notes were interconnected. each note included in an ssmr thread had to be a reaction to some previous note(s) or had to be followed by a note(s) reacting to it. the notes in the ssmr thread did not necessarily follow each other immediately, but other notes could appear in between or in parallel that were unconnected to the ssmr threads. each ssmr thread was analyzed in terms of its overall facilitative regulatory function in the development of the overall inquiry learning process, as adapted from iiskala and colleagues (2011). hence, depending on the group’s perceived appropriateness of their cognitive process’s direction, the ssmr’s specific function was either to ‘continue’ or ‘inhibit’ their evolving cognitive process. both functions sought to achieve a better performance and understanding of the task. a continue ssmr thread captured situations when the group confirmed that the direction of their current thinking was appropriate or they activated further cognitive processes in the same direction. for example, the students identified gaps in their group’s thinking and, on this basis, planned to maintain their current direction, as this would bring them closer to their common goal. in contrast, an inhibit ssmr thread captured situations when the group slowed down, changed or stopped its cognitive process’s direction because the students perceived it as inappropriate and decided to re-think their approach. for example, some interconnected notes revealed how the students jointly interrupted their current train of thought and re-oriented this thinking. if a student’s cognitive note (e.g. a proposal) triggered an ssmr thread, the first regulatory note reacting to it was analyzed as the thread’s starting point. the end was signalled by the note that ended the ssmr process, and the thread displayed its executive function of continuing or inhibiting the direction of the group’s thinking process. 2.4 inter-coder agreement determining a thread is complex, and, as argued by strijbos and stahl (2007), represents an interpretation of the discussion’s structure. inter-coding, therefore, was a necessary part of the analysis. it was conducted in two steps. first, after the principal coder had reviewed all the notes and identified the ssmr threads, the other coder, a scholar with extensive experience in learning research, independently analyzed the threads in the context of all notes. the second coder’s task was to assess whether or not, in his opinion, the threads identified by the principal coder represented ssmr. the second coder could also suggest adding or removing note(s) to a thread, and propose new threads. this researcher was also asked to define the threads’ function according to the coding scheme. only a few disagreements arose between the two coders, and these were resolved in discussions. if the disagreement was whether or not the proposed thread represented ssmr, the thread was automatically abandoned. this happened in only two instances. in addition, one thread, whose function was originally defined as ‘activation’ by the principal coder, was, after a discussion, re-classified as ‘stop’. all other ssmr threads were agreed upon by both coders. overall, this means that eight of 11 threads were accepted as such, two threads were removed and one thread’s function was re-classified. hence, all nine threads illustrated in the present article were recognized as ssmr threads and their function were fully accepted by both coders. the inter-coder agreement was 82% in the identification of ssmr threads, and 89% in classifying these threads according to their function. second, related to sna and research aim 3, the inter-coder agreement was calculated on the basis of how notes outside the ssmr threads were connected with each other. after the principal coder had finished the coding process, two phases of the inquiry learning process (i.e. ‘setting up the research question’ and ‘developing a work plan’) were randomly selected, and 33% (n = 189) of all 566 notes outside the threads were coded by the second, trained coder. the coders agreed on 233 of 263 connections, so an agreement of iiskala'et'al' ' ' | f l r ! ! 85! 89% (cohen’s κ = .88) was reached, which can be considered substantial (see landis & koch, 1977, p. 165). all disagreements were resolved by negotiation, and, conservatively, only connections that both coders had agreed on before the negotiation were used in analysis. 3. results 3.1 ssmr manifestation and functions in an asynchronous cscl environment of a small group’s inquiry process a general overview of the findings is provided first. table 1 summarizes the main features of the nine ssmr threads, including their function and behavioral indicators, the phase of the inquiry process in which they were located, and the number of notes and participants. table 1 summary of ssmr threads thread function phase number of notes number of students names of students a inhibit/ stop setting up the research question 5 3 iina, piia, sonja b inhibit/ stop setting up the research question 2 2 iina, sonja c continue/ confirm setting up the research question 8 3 iina, piia, sonja d inhibit/ slow constructing a hypothesis 15 4 iina, piia, sonja, julia e inhibit/ slow constructing a hypothesis 3 2 piia, julia f continue/ activate constructing a hypothesis/ developing a work plan 14 4 iina, sonja, piia, julia g inhibit/ slow searching for and processing knowledge/ summarizing findings and concluding 10 3 iina, sonja, julia h inhibit/ change summarizing findings and concluding 8 3 iina, sonja, julia i continue/ activate summarizing findings and concluding 9 4 iina, piia, sonja, julia iiskala'et'al' ' ' | f l r ! ! 86! as shown in table 1, both specific ssmr functions (continue or inhibit the current direction of the cognitive process), and all five behavioral indicators of these functions (confirm, activate, slow down, change and stop) were found in the data. ssmr threads that continued or inhibited the cognitive process’s direction were found across all phases of the inquiry learning process, but the majority (six out of nine) of the threads functioned to inhibit the process’s perceived inappropriate or overly rushed direction (stop, slow or change). most of the threads (seven out of nine) were contained within one phase, and two of the threads spread over two phases. only three of the nine threads displayed the participation of all four students, but only two threads had only two students contributing. table 1 presents the threads in groups of three (a, b, c / d, e, f / g, h, i) because each set of three included threads that were connected to each other. this will be discussed below (in section 3.2), when addressing the second research aim. figure 1 presents a comprehensive and dynamic visual display of all the data for this small group, over the evolution of its asynchronous cscl inquiry process. this figure captures the emergence of the nine ssmr threads (a–i) and their respective functions in the inquiry process’s different phases and among all written notes. the figure also shows how each individual contributed to ssmr threads and other non-ssmr activities. therefore, the figure identifies the four girls, each phase of the inquiry process, the number of lessons, all the written notes with symbols and the timeline (as hour:min). for example, in iina’s row, a symbol ‘●’ features at 9:31, located in lessons 3–4 of the ‘setting up the research question’ phase, which means that iina wrote one note at that time. however, this note is not a part of any ssmr threads because no arrow (→) links this note to other notes. as illustrated in the figure, all notes that are part of an ssmr thread are linked by arrows, which show the progression of the girls’ thoughts throughout the thread. moreover, the letter (a–i) illustrates to which ssmr thread that note belongs. for instance, five notes appear in ssmr thread a (inhibit/stop), two from iina, one from sonja and two from piia. furthermore, the notes generated during the same phase of the inquiry process are assigned the same symbol. for example, the symbol ‘●’ is used to represent the ‘setting up the research question’ phase. moreover, the girls could also return to some topic (e.g. setting up the research question) later during the inquiry process (e.g. sonja wrote a note regarding setting up a research question during the ‘constructing a hypothesis’ phase, which the symbol ‘●’ indicates in lessons 7–8 at 11:14). altogether, 74 out of 640 notes (12%) were classified as belonging to ssmr threads. all nine ssmr threads (a–i) shown in figure 1 are also qualitatively illustrated at the micro level (see section 3.2). in summary, as illustrated in figure 1, ssmr in the small group’s asynchronous cscl inquiry learning process was manifested as follows: a) ssmr threads appeared mainly at the start and the end of the overall inquiry learning process, with intensive and quiet periods where no ssmr took place. this stresses the importance of looking at what students are doing at the process’s different phases and what may trigger ssmr. b) some ssmr threads overlapped or intertwined (a, b, c / d, e, f / g, h, i), which may have played an important role in the evolution of the overall inquiry learning process. c) because of the asynchronous nature of the cscl process, some ssmr threads took place over an extended period, and other notes that were not part of the ssmr thread could be found in between ssmr notes. this highlights the importance of studying ssmr manifestations over a long period when dealing with asynchronous cscl process. iiskala'et'al' ' ' | f l r ! ! 87! iiskala'et'al' ' ' | f l r ! ! 88! iiskala'et'al' ' ' | f l r ! ! 89! 3.2 evolution of ssmr across different phases of the group’s asynchronous cscl inquiry learning process as shown in table 1 and illustrated by figure 1, the ssmr’s distribution varied across phases of the inquiry learning process. ssmr threads were found especially in the first two phases, ‘setting up the research question’ and ‘constructing a hypothesis’, and then when ‘summarizing the findings and concluding’, with limited ssmr during ‘developing a work plan’ and ‘searching for and processing knowledge’. the ssmr threads found in each phase are presented in turn, with a narrative describing their behavioral indicators and their functions in the inquiry process. 3.2.1 phase i: setting up the research question three inter-related threads (a, b and c) were found in this stage, illustrating the function of inhibiting the perceived undesirable direction of the group’s current cognition process and, alternatively, the function of continuing in a promising direction. figure 2 illustrates how the three ssmr threads interact. both threads a and b illustrate the specific ssmr functions that inhibit current undesirable thinking directions. in these threads, the girls realize that the research questions they had in mind would not address the overall aim of their inquiry process. in thread a, three of the girls (iina, sonja and piia) think about a potential research question to address the project’s aim, taking into account their prior knowledge. the ssmr thread stops, with the group concluding that the theme is too complex. this means that, through ssmr, the girls stopped developing their initial question and, accordingly, rejected their related thinking for this question. in thread b, two of the girls (iina and sonja) consider ambivalently why it is possibly not judicious to pursue their second proposed idea as a research question. the thread leads them to stop (a behavioral indication of inhibiting ssmr functions). in contrast to threads a and b, thread c reveals how the girls, through ssmr, confirmed their new question’s feasibility (a behavioral indication of continuing ssmr functions) and, thus, continued to pursue their current thinking direction. interestingly, thread c’s starting point was the same as for threads a and b, namely, that a potential research question was scrutinized. the third research question was eventually adopted through ssmr in thread c. three of the girls (sonja, iina and piia) were involved in this last thread, which unveils how they regulated the suitability of their new theme against the research aim and how they scrutinized possibilities for upcoming efforts. iiskala'et'al' ' ' | f l r ! ! 90! figure 2. inhibition function of stop (a and b) and continue function of confirm (c) of ssmr threads in ‘setting up the research question’. iiskala'et'al' ' ' | f l r ! ! 91! as illustrated in figure 2, threads a, b and c followed each other consecutively, but partly overlapped as well, each thread playing a specific function in the evolving inquiry learning process. threads a and b demonstrate how the girls discarded some of their ideas by regulating their respective usefulness, while thread c revealed how the group moved forward with another idea. thus, the group stopped or confirmed the direction of their ongoing thinking process through ssmr. in summary, the analysis of the three ssmr threads showed that, in the first phase of the asynchronous cscl inquiry process that sought to define the research question, some ssmr threads’ function was to inhibit (stop as the behavioral indicator) the direction of their cognitive process, since they perceived this as going in an unwanted direction. another ssmr thread enabled the group to continue (confirm) their line of thought. 3.2.2 phases ii and iii: constructing a hypothesis and developing a work plan three inter-related threads (d, e and f) were found in these two phases, once again illustrating both continuing and inhibiting functions in the girls’ current cognitive process. consistent with the evolving inquiry process, after setting up its research question, the small group proceeded to construct a hypothesis to address that question. the question itself, which had been confirmed in the previous phase, was subsequently elaborated as follows: ‘how can it be concluded that an asteroid will hit the earth, and what will then happen to the earth?’ some sub-questions were also generated: ‘where does an asteroid come from?, ‘will the whole earth be destroyed?’, ‘how can the hit be prevented?’ and ‘when will the hit be projected to happen?’ figure 3 illustrates the three ssmr threads (d, e and f) found during the phases of ‘constructing a hypothesis’ and ‘developing a work plan’. both threads d and e are illustrations of ssmr inhibiting undesirable directions of the group’s cognitive process (slowing down as the behavioral indicator), in this case, due to the group realizing their ideas’ limitations. in these two ssmr threads, both taking place in the phase of constructing the hypothesis, the girls question their inquiry process. these particular ssmr threads captures a slow down in the group’s thinking process, rather than a change in their thoughts’ overall direction, since their hypothesis remains the same. by stating their uncertainty and slowing down their thinking process the students became aware of their ideas’ inadequacies and the unresolved issues, thus they realized the importance of not rushing and making decisions in a hurry. in thread d, the girls’ notes show that, although they were aware that a reliable hypothesis could not be constructed, they were nevertheless unable to revise their line of thought. all four girls participated in this ssmr thread, which highlights its importance. subsequently, thread d focused on some aspects of the hypothesis, whereas thread e briefly brought to light difficulties, while considering the hypothesis, with how to construct a situation at a general level. in thread e, the regulation process was shared by only two of the girls (julia and piia). thus, these threads had to be analyzed separately. in turn, thread f illustrates another behavioral indicator of ssmr’s continuing function, namely, the activation of other ideas while pursuing a single line of thought. this thread spreads over the phases of ‘constructing a hypothesis’ and ‘developing a work plan’. in this thread, the girls identified gaps in their thinking by an ongoing regulation of the kind of knowledge they thought was needed. the notes show how the knowledge they thought was required was co-constructed through ssmr. in this thread, the girls demonstrated their ability to regulate their cognitive process jointly towards a common goal and to coconstruct a shared awareness of the knowledge they needed and the sources they could access. finally, thread f led the girls to organize their plan’s implementation, with a division of labor based on identified gaps in their knowledge. all four girls participated in this thread, which was pivotal for their subsequent work plan and for searching for and processing knowledge, as illustrated in figure 1, but which involved minimal ssmr. iiskala'et'al' ' ' | f l r ! ! 92! iiskala'et'al' ' ' | f l r ! ! 93! iiskala'et'al' ' ' | f l r ! ! 94! figure 3. inhibition function of slow (d and e) and continue function of activate (f) of ssmr threads in ‘constructing a hypothesis’ and ‘developing a work plan’. a key feature of ssmr threads d, e and f is that they were closely intertwined. in thread d, which features the group slowing down their thinking process (slow down as a behavioral indicator of the process being somewhat inhibited), the girls questioned their hypothesis and hesitantly pursued the development of this hypothesis. partly coinciding with thread d, the girls disclosed in thread e that they were uncertain about how to proceed and thus slowed down and reconsidered their options. in contrast, thread f illustrates how the group was eventually able to move forward by activating the ideas they had reconsidered and found suitable in threads d and e. taken in combination, these ssmr threads illustrate how the group regulated the progress of their inquiry process and eventually agreed on what knowledge they needed. this provided evidence that previous threads could influence subsequent threads, an aspect of ssmr that may be iiskala'et'al' ' ' | f l r ! ! 95! characteristic of asynchronous cscl processes in which students can go back to previous discussions. furthermore, this influence became visible only through analysis of cscl processes over an extended period. 3.2.3 phases iv and v: searching for and processing knowledge and summarizing findings and concluding after developing a work plan, the group moved to the final phases of its cscl inquiry process, namely ‘searching for and processing knowledge’ and, finally, ‘summarizing the findings and concluding’. three ssmr threads (g, h and i) identified in these phases are shown in figure 4. in the course of searching for and processing knowledge and summarizing findings and concluding, one ssmr thread (g) appeared to slow down the group’s inquiry learning process substantially. in thread g, the group discussed if the exact time of the asteroid hit could be determined and whether this information should be provided when summarizing the findings and concluding. the thread shows that the group’s ssmr of this issue prevented overly hasty decisions (slowing down as the behavioral indicator of inhibiting function) over an extended period and across several phases of the inquiry process. notably, thread g, which slowed down the process of searching for and processing knowledge, as well as summarizing the findings and concluding, took place over a much longer period than threads d and e (discussed in section 3.2.2). their respective function was also to inhibit undesirable or hastily selected directions in the group’s cognitive process, but they were contained in a single phase. all of the girls, except piia, participated in this extended ssmr thread. in addition, ssmr thread h appeared during the ‘summarizing findings and concluding’ phase, which led to a change in the group’s course of action (another behavioral indicator of inhibiting functions). in this thread, the group regulated their searching for, and processing of, knowledge in relation to the research question. the students became aware of a mismatch: the knowledge they had activated did not answer the research question. as a consequence, the group changed its research question so that this would correspond better to the knowledge they had gathered during the phase of ‘searching for and processing knowledge’. thread h shows quite explicitly how the girls changed their previous research question to a new one by consensual regulation. sonja, julia and iina were all actively involved in this discussion. the final ssmr thread (i) functioned to move the inquiry process forward (a continue function), with all the girls jointly engaged in shared regulation of their idea, activating all three concepts (i.e. an asteroid, a meteorite and a meteoroid) that were needed to explain their summary findings and conclusion. piia’s first note (10:39) appeared to have been discounted when the other three girls changed the research question. however, her second note (10:44), where she repeated the thought in her first note by clarifying what she meant, showed that she was closely monitoring the group’s cognitive process. as she was not satisfied with the other girls’ changes in the research question (thread h), she posted a new note starting a new thread during which the group eventually agreed to include the difference between the three different concepts in their findings and conclusion. iiskala'et'al' ' ' | f l r ! ! 96! iiskala'et'al' ' ' | f l r ! ! 97! iiskala'et'al' ' ' | f l r ! ! 98! figure 4. inhibition functions of slow (g) and change (h) and continue function of activate (i) of ssmr threads in ‘searching for and processing knowledge’ and ‘summarizing findings and concluding’. a most interesting characteristic of threads g, h and i, was their deep intertwined nature. during the phase of ‘searching for and processing knowledge’, only one ssmr thread (g) appeared. in that thread, the girls regulated the issue of whether having the exact time was important, as well as if this would be reliable. however, at that time, they had not yet realized that their inquiry would not address their research questions. thus, because of insufficient regulation while ‘searching for and processing knowledge’, need for regulation emerged later (threads h and i), at the time the findings were summarized and a conclusion was drawn. hence, together, threads g and h demonstrate how insufficient regulation in one phase can necessitate making up for lost time in a latter phase. threads h and i also displayed strong interconnections, with piia’s neglected note in thread h triggering thread i. iiskala'et'al' ' ' | f l r ! ! 99! 3.3 individual group members’ participation in ssmr during an asynchronous cscl inquiry process and the influence of specific individual contributions on the group’s regulatory effort 3.3.1 distribution of students’ notes contributing to ssmrs the first analysis was to determine if the four students’ contributions to the ssmr effort was similar or different. although ssmr participation appeared to vary across students, iina participated the most (25 out of 74 notes or 34% of the group’s notes in the ssmr threads), followed by julia (22 notes or 30%), sonja (18 notes or 24%) and, finally, piia (9 notes or 12%). table 2 displays the number (and percentage) of each student’s notes in the ssmr threads and, nonssmr (other) interactions and their totals, within the whole asynchronous cscl inquiry process. a cross tabulation and a pearson’s chi-square test were performed. cramer’s v correlation was used to test associations (see siegel & castellan, 1988). the relationship between the number of notes within and outside ssmr threads across students was not different, χ2(3, n = 640) = 3.79, v = .08, ns (see table 2). z-tests with bonferroni correction also showed that the differences in the proportions of notes between the students were not significant. table 2 frequencies (and percentages) of each student’s notes student notes iina piia sonja julia total other 201 (89)a 110 (92)a 132 (88)a 123 (85)a 566 (88) ssmr 25 (11)a 9 (8)a 18 (12)a 22 (15)a 74 (12) total 226 (100) 119 (100) 150 (100) 145 (100) 640 (100) note. other = notes outside ssmr threads. ssmr = notes that are included in ssmr threads. each subscript denotes a subset of the process phases whose column proportions do not differ significantly from each other at the .05 level. 3.3.2 distribution of students’ notes in ssmr threads that functioned to continue or inhibit the direction of the group’s evolving cognitive process based on perceived appropriateness table 3 displays the number (and percentage) of each students’ notes contributing to ssmrs that played different functions. table 3 frequencies (and percentages) of each student’s notes in terms of ssmr functions (continue, inhibit) student function iina piia sonja julia total continue 13 (52)a 4 (44)a 9 (50)a 5 (23)a 31 (42) inhibit 12 (48)a 5 (56)a 9 (50)a 17 (77)a 43 (58) total 25 (100) 9 (100) 18 (100) 22 (100) 74 (100) note. each subscript denotes a subset of the process phases whose column proportions do not differ significantly from each other at the .05 level. iiskala'et'al' ' ' | f l r ! ! 100! as can be seen, individuals’ participation in ssmr threads that played a different overall function seemed to vary across students. iina, piia and sonja participated with an equal number of notes, contributing to ssmrs with the functions of continuing and inhibiting the direction of the group’s cognitive process, but julia contributing three times as many notes, contributing more to ssmr inhibiting the group’s process (slow down, stop or change) than to ssmr continuing the process. the relationship between students’ notes in ssmr threads and the function category they contributed to (continue, inhibit) was however, not statistically significant based on pearson’s chi-square test and cramer’s v correlation, χ2(3, n = 74) = 4.88, v = .26, ns. the non-significance was confirmed, using z-tests with bonferroni correction. however, julia’s strong contribution to making the group stop, slow down or change the direction of their thinking is noteworthy. it contrasts with iina’s contributions to the ssmr, which were as frequent as julia’s (respectively 22 and 25 notes out of 74) but which differed qualitatively in terms of the nature of each girl’s contributions to the group’s regulatory effort. 3.3.3 distribution of students’ notes contributing to ssmrs the inquiry process’s different phases table 4 displays the number (and percentage) of each student’s notes contributed to ssmrs in different phases. table 4 frequencies (and percentages) of each student’s notes in different phases student phase iina piia sonja julia total setting up the research question 7 (47) 3 (20) 5 (33) 0 (0) 15 (100) constructing a hypothesis 5 (19) 5 (19) 5 (19) 12 (43) 27 (100) developing a work plan 3 (60) 0 (0) 1 (20) 1 (20) 5 (100) searching for and processing knowledge 2 (33) 0 (0) 1 (17) 3 (50) 6 (100) summarizing findings and concluding 8 (38) 1 (4) 6 (29) 6 (29) 21 (100) as can be seen, some students appeared to make a substantial contribution to ssmr in some phases and less (sometimes none) in others. for example, julia made no contribution to ssmr in the ‘setting up the research question’ phase but played a dominant role in the phases ‘searching for and processing knowledge’ (50% of the notes in this phase) and ‘constructing a hypothesis’ (43%). another example is piia, who contributed to ssmr almost exclusively in the first two phases (with no contributions in phases iii and iv and one single contribution in phase v), while iina contributed substantially throughout all phases (respectively, 47% of all notes in phase i, 19% in phase ii, 60% in phase iii, 33% in phase iv and 38% in phase v). a significant relationship was found in the distribution of students’ notes by phases (fisher’s exact test = 17.17, p < .10). a number of observations can be made based on these findings. first, they reveal that piia, who was identified as the group member who contributed least to the ssmr effort overall (see section 3.3.2), was in fact active in the early phases of the inquiry learning process, and she contributed as much as iina and sonja iiskala'et'al' ' ' | f l r ! ! 101! towards the group’s ssmr of the construction of their hypothesis (5 notes each). one may wonder why piia hardly contributed to the group regulatory effort after that and whether the other girls’ perhaps more assertive statements may have played a role, in particular, julia’s questioning approach illustrated in ssmr threads d, f, g and h (figures 3 and 4). second, this analysis also shows that julia, who played such a dominant role overall in inhibiting the group cognitive process moving in an unwanted direction, had in fact not contributed at all in the first phase of ‘setting up the research question’. given that the original research question was eventually modified, understanding why julia did not participate in the first phase may be necessary in order to understand her subsequent dominance in regulating the process of constructing the hypothesis. 3.3.4 connections between students across the whole cscl inquiry process sna was used to explore further individuals’ contributions to ssmr, offering a different analytical approach. consistent with sna, each note was considered in relationship to other notes; that is to say, each note was studied for whether it reacted (respondent or, in sna terms, ‘in-degree’) to other notes (initiator or, in sna terms, ‘out-degree’). the directional information generated by sna, therefore, provided useful information on how each individual was positioned in reference to her peers during the joint regulation process. density, centrality and centralization measures were established based on these data. density is a group level measure that indicates how lively interaction is among the students. it measures the average strength of connections in the network (wassermann & faust, 1994). in the data, six dyad ties were present (sonja-julia, sonja-iina, iina-julia, iina-piia, piia-sonja, piia-julia), which can be divided into 12 relationships in actor level analysis, distinguishing between the initiator and respondent roles. in turn, the actor level analysis produced 60 connections between students in the ssmr threads, generating a density value of 5 (60/12) for the connections within the ssmr threads. while density indicates how lively the interaction is, centrality is an individual level measure for this, showing, who has the most or least contacts to other students (see table 5). using scott’s (1991) approach, centrality was measured with freeman’s degree (i.e. number of initiator or out-degree and respondent or in-degree notes). the results indicate that each student initiated and reacted on average to 15 notes (initiator sd 4.95, respondent sd 4.89) within the ssmr threads, with important individual differences. for example, as shown on the left hand side of table 5, pia was less at the center of the connections than other students (see table 5). table 5 frequency of connections between students and number of ties between pairs of students within the nine ssmr threads frequency of connections ties between pairs of students respondent pair initiator iina piia sonja julia total sonja-julia 16 iina 2 5 8 15 sonja-iina 15 piia 3 2 2 7 iina-julia 15 sonja 10 3 7 20 piia-iina 5 julia 7 2 9 18 piia-sonja 5 total 20 7 16 17 piia-julia 4 total 60; m=10 this means that piia was less verbally active in the joint regulatory process. this finding is consistent with the differences tested with pearson’s chi-square tests (see sections 3.3.1 and 3.3.3), which were significant at the p<.10 level. furthermore, sonja’s notes appeared to be the most frequently reacted to iiskala'et'al' ' ' | f l r ! ! 102! by others in the group. however, although sonja did not generate as many notes (18) as iina and julia (respectively, 25 and 22) (see section 3.3.1), sna indicates that her notes produced the most reactions, which implies an important role in the ssmr process. finally, centralization is a group level measure of how equally interaction is divided among the participants. the measure of centralization indicates how tightly connections are organized within this small group (see scott, 1991; wassermann & faust, 1994) – the higher the percentage, the greater the difference between students. in a valued graph, the percentage can go over 100%. the results indicate that this small group was only very little centralized within the ssmr process (initiator 22.2%, respondent 22.2%), that is, the students’ interactivity levels did not differ very much. the six dyad ties between the four students (ignoring person initiator and respondent roles) also were computed and compared. as shown on the right hand side of table 5, the dyad ties that did not include piia were always stronger (15, 15, 16) and those involving piia were always weaker, regardless of who was the other partner (5, 5, 4). 4. discussion the aim of the present study was to explore socially shared metacognitive regulation (ssmr) in a small group of students’ collaborative inquiry process in an asynchronous cscl environment. although a substantial body of empirical research exists on asynchronous cscl, in-depth micro-level studies of ssmr are still scarce. this study’s main contribution is a novel micro-level analysis of ssmr processes and the demonstration of this methodology’s use in a case study. previous studies have dealt with these processes mainly on a theoretical or a much more generalized empirical level. this study provides evidence that ssmr manifestations are not only found in face-to-face collaborative learning environment but also in asynchronous cscl environment (see also hurme, merenluoto, & järvelä, 2009; volet et al., 2013), which is important since previous research pointed to the challenge of developing analytical methods that can be used reliably and validly across different contexts, such as cscl environments (beers, boshuizen, kirschner, & gijselaers, 2007). this study’s findings are consistent with other studies (e.g. grau & whitebread, 2012; hurme et al., 2009; molenaar, 2011; whitebread at al., 2007; volet et al., 2009a), which have documented the emergence of ssmr in a range of collaborative learning contexts. however, as revealed in this study, ssmr manifestations can look different in computer-supported asynchronous environments, when compared to face-to-face contexts in which ssmr takes place in real-time. in the asynchronous environment studied, we found evidence of ssmr threads that lasted over an extended period and notes within threads that did not follow each other immediately but were interspersed with non-ssmr notes and sometimes intertwined with other threads.!this kind of non-linearity can be found in face-to-face discussions as well (iiskala et al., 2004, 2011), but, in an asynchronous environment, this kind of longitudinal regulation is much more dominant. the present study also examined each ssmr thread’s function in the flow of cognitive process. similar to what has been reported in face-to-face collaborative learning contexts (e.g. iiskala et al., 2011), evidence was found that ssmr can play different functions. some appeared to continue (i.e. activate, confirm) the current direction of the inquiry learning process, whereas others appeared to inhibit (i.e. slow, change, stop) its current direction, more specifically inhibiting its perceived ineffective or overly hasty progression. while these findings somewhat mirror the ssmr manifestations found in face-to-face mathematical word-problem solving by high-achieving student dyads (iiskala et al., 2011), the dominant functions differed. in iiskala and colleagues’ (2011) study, the most common ssmr function was to confirm the direction of the cognitive flow (i.e. continuation of the problem-solving process). in contrast, substantial evidence appeared in the present study of inhibitive ssmr functions, mainly slowing down, stopping or changing the inquiry process’s direction when this was perceived as inappropriate or rushed at that point in iiskala'et'al' ' ' | f l r ! ! 103! time. hence, the findings of this in-depth case analysis indicate that, in an asynchronous cscl environment, small groups can regulate of their cognitive activity’s direction so that inappropriate processes are reassessed. in this study, limited evidence surfaced that the group was using ssmr to develop an appropriate way to proceed towards their common goal. one reason for this could be that, in computer-supported environments, small groups have the possibility of deepening their knowledge through long discussions (e.g. zhang, scardamalia, lamon, messina, & reeve, 2007). in the present study, however, both short and relatively long ssmr threads were identified. when threads contained only a few notes, as is often the case in cscl, the discussions soon ended (see lipponen, rahikainen, lallimo, & hakkarainen, 2003). this implies, that in short threads, ssmr is not sustained over long periods. notably, although ssmr threads can be short, consisting of only a few notes, the discussions (e.g. on cognitive topics) in which ssmr intervenes can be longer. this means, that ssmr, for example, can change the direction of the group’s thinking process even during a rather short ssmr thread. in addition, although ssmr might be visible for only short periods, cognitive implementation based on ssmr can last longer. furthermore, the distinct findings in face-to-face and asynchronous cscl environments could be due to the nature of the students’ tasks. in iiskala and colleagues’ (2011) study, the problems had a single correct answer, although not always achieved by straightforward arithmetic calculations, which could lead the students to decide to check whether or not their solutions to the problems were correct. this was not the case in the present study because the problems were more multidimensional and ill-defined. ill-defined questions could be one of the reasons the small group in the present study spent more time regulating their inquiry learning process (12% of all notes) than has been reported in other studies also conducted in asynchronous cscl environments. for example, prinsen and colleagues (2007) reported that only 6% of students’ time was spent on regulation, and they speculated that the tasks’ well-structured nature could be the cause. although ill-defined questions and multidimensional tasks can create opportunities for ssmr in collaborative inquiry learning, the possibility of alternative explanations of variations in ssmr functions will have to be investigated more systematically, for example, changing the group size (e.g. dyad vs. small group). one under-examined issue in previous research is how ssmr evolves over the duration of an extended collaborative learning activity. in the present study conducted in an asynchronous cscl environment, ssmr was found in all phases of the process, but with some variation in the number of threads across phases, as well as the number of notes within threads. for example, fewer notes appeared within ssmr threads in the phase of developing a work plan than during constructing a hypothesis and summarizing findings and concluding. hence, different phases triggered students to regulate the process to a different extent. this is consistent with lajoie and lu (2012), who found that the frequency of regulation varied with the learning process’s different phases. furthermore, in asynchronous collaboration, the present study found that ssmr threads were either contained in one phase or developed and evolved over several phases. these findings suggest that, although ssmr can emerge across a range of collaborative learning contexts, its manifestation over time can look different, depending on the mode of collaborative learning (cscl vs. face-to-face). this indicates that simply extrapolating the findings from face-to-face environments to virtual learning environments is not possible. differences in ssmr engagement across phases could also be related to variations in each phase’s level of difficulty. for example, developing a work plan could be easier than constructing a hypothesis or summarizing findings and concluding, since planning is a generic activity that may not require as much content knowledge. this interpretation is consistent with other studies (e.g. efklides, papadaki, papantoniou, & kiosseoglou, 1998; iiskala et al., 2004, 2011; prins, veenman & elshout, 2006), which revealed that metacognition is more likely to appear when tasks are reasonably difficult. this may also help explain research findings (e.g. khosa & volet, 2014) pointing to a relationship between ssmr and construction of high-level content knowledge. hence, in order to foster productive ssmr in school learning activities, an optimal difficulty level of tasks may be important, that is, sufficiently challenging but within the students’ zone of proximal development (see vygotsky, 1978). one original contribution of the present research was to investigate students’ participation in ssmr using different analytical methods. this approach proved valuable as it provided complementary insights iiskala'et'al' ' ' | f l r ! ! 104! into group members’ ssmr participation. for example, no individual differences appeared in the number of ssmr notes themselves. however, sna revealed that group members reacted the most frequently to one particular student’s notes. our findings concur with palonen and hakkarainen’s (2000) study, in which sna revealed information on students’ participation in cscl not revealed through other analyses. hence, examining who has the largest number of notes in ssmr threads appears to be important, but so is pinpointing whose notes caused the other students to react the most. this finding suggests that ssmr is somewhat similar to the notion of distributed expertise (see brown et al., 1993), in which members of a community of practice are recognized as critically interdependent, their expertise is shared between members and meaning negotiated within the group. another interesting finding in the present study was that one student had fewer connections within the ssmr threads than her three fellow students did, even though all four were above average students in terms of their cognitive skills. the range of individual and contextual aspects that have an impact on ssmr participation needs to be better understood. the present study also highlighted the different roles students play in different ssmr threads. the finding that one student played a dominant role in the ssmr threads that slowed down, changed or stopped the inquiry process’s direction, and a minor role in the threads that continued the process, suggests that some students make different contributions from a metacognitive point of view. this supports the reasonable assumption that trying to inhibit the cognitive process’s flow because those involved perceive that it is overly hasty or that it is going in an inappropriate direction is more demanding metacognitively than confirming the group activity continues. previous cscl research (e.g. prinsen et al., 2007) also suggested that learners’ characteristics and position among classmates might affect their participation in asynchronous cscl. therefore, how individual differences, and possibly individual position within groups, play out in regard to more or less demanding contributions to ssmr will need to be examined in future research. other individual differences emerged, in this study, focused on in which phases students participated most and in which they participated less or not at all. these findings highlight that analyzing individuals’ total ssmr engagement during inquiry learning processes – as if these are single entities – is insufficient. scrutinizing how each individual’s ssmr engagement evolves in different phases is also important – similar to analyses undertaken at the group level. variations in individuals’ participation across phases, or in specific tasks carried out in different phases, should generate new insights into the emergence of ssmr, including the role of group dynamics. studying students’ contributions at different phases of problem-solving or inquiry processes may also reveal whether and how high-quality collaboration is sustained. this understanding is vital, since rogat and linnenbrink-garcia (2011) found that high-quality regulation is characterized by frequent high-quality interactions. similarly, based on their study of networked learning, toikkanen and lipponen (2011) suggested that students’ participation needs to be distributed among many students. one additional aspect to keep in mind when studying ssmr is that, sometimes, all group members and, at other times, only a sub-group is involved. how this aspect relates to the quality of the group process and outcome will need to be examined further. summers and volet’s (2010) study provides support for this suggestion as they found evidence, in their most successful group, that half of all the contributions of every single group member reflected high-level engagement. the qualitative examples presented in the present study showed how a student’s suggestion was ignored by the other students in the first instance but was taken into consideration later and eventually became the starting point of a new ssmr thread. this is in line with molenaar’s (2011) study, which reported ‘ignored metacognitive activities’ that represented situations in which other group members ignored a peer’s input. these findings provide support for the assumption that ssmr is more than simply the sum of individuals’ metacognitive regulation and, thus, cannot be reduced to each individual’s level. instead, as in all social systems (see salomon & globerson, 1989; vauras, salonen, & kinnunen, 2008), including small groups, the reactions and subsequent interactions between participants play a crucial role. the present study also provided support for the importance of using more than one method to analyze ssmr data, not only because different methods provide complementary insights but also because iiskala'et'al' ' ' | f l r ! ! 105! they prevent researchers from drawing conclusions from insufficient information. the use of sna in this study created new insights into ssmr by revealing that, although all students participated in ssmr to some degree, the distribution of their participation was unequal and appeared to vary according to the method of analysis. however, sna centralization percentages showed that the small group studied was minimally centralized, and only around some students, which was consistent with the results of the cross tabulation, which did not show any individual differences. calls for the use of multi-method designs when studying metacognitive skills (veenman, 2005) and computer-based learning environments (azevedo, 2005) have been issued, and these methodological approaches need to pursue vigorously. the use of multi-method designs and alternative analytical methods is particularly critical when studying ssmr in cscl since the additional information provided by nonverbal indicators in face-to-face ssmr is unavailable. this was particularly important in the present study, which relied on an in-depth analysis of a single case. as an example, sna is typically used in the analysis of members’ connections, providing critical insights into the nature of individual contributions, more specifically, individual positioning within interactions. as toikkanen and lipponen (2011) have suggested, the combination of qualitative analysis and sna provides more accurate information about students’ participation and, thus, a stronger foundation for interpretation. when technology is introduced in classrooms, it does not just add a new element in the existing pedagogical practice and environment, but has larger consequences (salomon, 1994). information and communication technology, such as an environment supporting computer-mediated collaboration, can be seen as an affordance for new forms of interaction and, simultaneously, as a challenge or constraint for conventional forms of pedagogical communication, and this can result in intended and un-intended consequences (kirschner, strijbos, kreins, & beers, 2004; suthers, 2006). in this article, we were able to deal with only some of these, but a more in depth analysis is needed in future studies of the multifarious consequences of information and communication technology in the regulation of social interaction. furthermore, in the present study, it was not possible to test how well the analysis method could be used in varying situations because only one group was under scrutiny. moreover, during the working process, direct face-to-face discussions could not be avoided among the participants. therefore, it was not possible to control what kind of regulation happened in these discussions. in addition, when the entire analysis is based on written notes, little information is available about students’ actual metacognitive experiences, which have proved to be important parts of ssmr processes in face-to-face situations (iiskala et al., 2011). finally, while the study of a single group limits the generalization that can be made from the findings, a number of directions for future research emerged. for example, follow-up studies with multiple groups will need to incorporate more versatile indices (e.g. toikkanen & lipponen, 2011) in order to reveal group characteristics, as well as similarities and differences between small groups. another research direction is to explore further the nature of, and relationship between, consecutive and concurrent ssmr threads in asynchronous cscl. calls have been made for sequential analyses (molenaar & järvelä, 2014) and analyses of individual participation across and within episodes or threads (panadero & järvelä, in press), which will provide further insights into ssmr. these analyses need also to focus on the entire flow of problem-solving or inquiry learning process, rather than analyses of single, isolated examples of ssmr, extracted from the process. by scrutinizing a small group’s ssmr during their entire inquiry learning process, the present study has made a unique contribution by revealing the distinct impact of ssmr on particular phases and showing the ways in which some ssmr functions are more frequent than others are in specific phases. iiskala'et'al' ' ' | f l r ! ! 106! keypoints in asynchronous cscl, ssmr manifests in threads of varied length, which can intertwine or overlap and affect the evolution of the learning process. intensive periods of ssmr and periods of limited ssmr take place in asynchronous cscl. in asynchronous cscl, ssmr has different functions, with a dominance of ssmr functions that inhibit the process’s perceived inappropriate direction. combining analytical methods, including sna, provides vital complementary insights into how students influence group regulatory efforts. scrutinizing the entire cscl process is essential to revealing the relationship between ssmr threads. acknowledgments we wish to thank the participating teacher and students who unfortunately must remain anonymous. we also wish to thank senior researcher tuire palonen for her help with social network analysis, statistician eero laakkonen for his advice regarding the quantitative analyses and graphic designer sirpa lehti for preparing figures 1, 2, 3 and 4. the research was supported by grant no. 274117 from the council of cultural and social science research, academy of finland, awarded to the fourth author. references azevedo, r. (2005). computer environments as metacognitive tools for enhancing learning. educational psychologist, 40, 193–197. doi: 10.1207/s15326985ep4004_1 azevedo, r. (2014). issues in dealing with sequential and temporal characteristics of selfand sociallyregulated learning. metacognition and learning, 9, 217–228. doi: 10.1007/s11409-014-9123-1 beers, p. j., boshuizen h. p. a., kirschner, p. a., & gijselaers, w. h. (2007). the analysis of negotiation of common ground in cscl. learning and instruction, 17, 427–435. doi:10.1016/j.learninstruc.2007.04.002 brown, a. l. (1987). metacognition, executive control, self-regulation, and other more mysterious mechanisms. in f. e. weinert & r. h. kluwe (eds.), metacognition, motivation, and understanding (pp. 65–116). hillsdale, nj: erlbaum. brown, a. l., ash, d., rutherford, m., nakagawa, k., gordon, a., & campione, j. c. (1993). distributed expertise in the classroom. in g. salomon (ed.), distributed cognitions. psychological and educational considerations (pp. 188–228). cambridge: university press. brown, a. l., bransford, j. d., ferrara, r. a., & campione, j. c. (1983). learning, remembering, and understanding. in p. h. mussen (ed.) & j. h. flavell & e. m. markman (vol. eds.), handbook of child psychology: vol. 3. cognitive development (4th ed., pp. 77–166). new york: john wiley & sons. brown, a. l., & campione, j. c. (1994). guided discovery in a community of learners. in k. mcgilly (ed.), classroom lessons: integrating cognitive theory and classroom practice (pp. 229–270). cambridge, ma: the mit press. brown, a. l., & campione, j. c. (1996). psychological theory and the design of innovative learning environments: on procedures, principles and systems. in l. schauble & r. glaser (eds.), innovations in learning new environments for education (pp. 289–325). mahwah, nj: lawrence erlbaum. brown, a. l., & deloache, j. s. (1983). metacognitive skills. in m. donaldson, r. grieve, & c. pratt (eds.), early childhood development and education (pp. 280–289). new york: blackwell. iiskala'et'al' ' ' | f l r ! ! 107! dinsmore, d. l., alexander, p. a., & loughlin, s. m. (2008). focusing the conceptual lens on metacognition, self-regulation, and self-regulated learning. educational psychology review, 20, 391– 409. doi: 10.1007/s10648-008-9083-6 efklides, a. (2008). metacognition: defining its facets and levels of functioning in relation to self-regulation and coregulation. european psychologist, 13, 277–287. efklides, a., papadaki, m., papantoniou, g., & kiosseoglou, g. (1998). individual differences in feelings of difficulty: the case of school mathematics. european journal of psychology of education, 8, 207–226. flavell, j. h., & miller, p. h. (1998). social cognition. in w. damon (ed.) & d. kuhn & r. s. siegler (vol. eds.), handbook of child psychology: vol. 2. cognition, perception, and language (5th ed., pp. 851– 898). new york: wiley & sons. grau, v., & whitebread, d. (2012). self and social regulation of learning during collaborative activities in the classroom: the interplay of individual and group cognition. learning and instruction, 22, 401–412. doi:10.1016/j.learninstruc.2012.03.003 greeno, j. g. (2006). learning in activity. in r. k. sawyer (ed.), the cambridge handbook of the learning sciences (pp. 79–96). new york: cambridge university press. hadwin, a. l., järvelä, s., & miller, m. (2011). self-regulated, co-regulated, and socially shared regulation of learning. in b. j. zimmerman & d. h. schunk (eds.), handbook of self-regulation of learning and performance (pp. 65–84). new york, ny: routledge. hakkarainen, k. (2003). emergence of progressive-inquiry culture in computer-supported collaborative learning. learning environments research, 6, 199–220. hakkarainen, k., & sintonen, m. (2002). the interrogative model of inquiry and computer-supported collaborative learning. science & education, 11, 25–43. hogan, k. (2001). collective metacognition: the interplay of individual, social, and cultural meanings in small groups’ reflective thinking. in f. columbus (ed.), advances in psychology research, 7 (pp. 199– 239). huntington, ny: nova science publishers. hurme, t-r., merenluoto, k., & järvelä, s. (2009). socially shared metacognition of pre-service primary teachers in a computer-supported mathematics course and their feelings of task difficulty: a case study. educational research and evaluation, 15, 503–524. doi: 10.1080/13803610903444659 hurme, t.-r., palonen, t., & järvelä, s. (2006). metacognition in joint discussions: an analysis of the patterns of interaction and the metacognitive content of the networked discussions in mathematics. metacognition and learning, 1, 181–200. doi: 10.1007/s11409-006-9792-5 iiskala, t., vauras, m., & lehtinen, e. (2004). socially-shared metacognition in peer learning. hellenic journal of psychology, 1, 147–178. iiskala, t., vauras, m., lehtinen, e., & salonen, p. (2011). socially shared metacognition of dyads of pupils in collaborative mathematical problem-solving processes. learning and instruction, 21, 379–393. doi:10.1016/j.learninstruc.2010.05.002 järvelä, s., & hadwin, a. (2013). new frontiers: regulating learning in cscl. educational psychologist, 48, 25–39. doi: 10.1080/00461520.2012.748006 khosa, d. k., & volet, s. (2014). productive group engagement in cognitive activity and metacognitive regulation during collaborative learning: can it explain differences in students’ conceptual understanding? metacognition and learning, 9, 287–307. doi: 10.1007/s11409-014-9117-z kirschner, p., strijbos, j-w., kreins, k., & beers, p.j. (2004). designing electronic collaborative learning environments. educational technology research and development, 52, 47–66. lajoie, s. p.& lu, j. (2012). supporting collaboration with technology: does shared cognition lead to coregulation in medicine? metacognition and learning, 7, 45–62. doi: 10.1007/s11409-011-9077-5 landis, j. r., & koch, g. g. (1977). the measurement of observer agreement for categorical data. biometrics, 33, 159–174. lipponen, l., rahikainen, m., lallimo, j., & hakkarainen, k. (2003). patterns of participation and discourse in elementary students’ computer-supported collaborative learning. learning and instruction, 13, 487– 509. doi:10.1016/s0959-4752(02)00042-7 iiskala'et'al' ' ' | f l r ! ! 108! molenaar, i. (2011). it’s all about metacognitive activities: computerized scaffolding of self-regulated learning. doctoral dissertation, university of amsterdam. molenaar, i., & chiu, m. m. (2014). dissecting sequences of regulation and cognition: statistical discourse analysis of primary school children’s collaborative learning. metacognition and learning, 9, 137–160. doi: 10.1007/s11409-013-9105-8 molenaar, i., & järvelä, s. (2014). sequential and temporal characteristics of self and socially regulated learning. metacognition and learning, 9, 75–85. doi: 10.1007/s11409-014-9114-2 nurmela, k. lehtinen, e., & palonen, t. (1999). evaluating cscl log files by social network analysis. in c. hoadley & j. roschelle (eds.), computer support for collaborative learning (pp. 434–444). stanford university: palo alto, ca. palonen, t., & hakkarainen, k. (2000). patterns of interaction in computer-supported learning: a social network analysis. in b. fishman, & s. o’connor-divelbiss (eds.), fourth international conference of the learning sciences (pp. 334–339). mahwah, nj: erlbaum. panadero, e., & järvelä, s. (in press). reviewing findings on socially shared regulation of learning. european psychologist. prins, f. j., veenman, m. v. j., & elshout, j. j. (2006). the impact of intellectual ability and metacognition on learning: new support for the threshold of problematicity theory. learning and instruction, 16, 374– 387. doi:10.1016/j.learninstruc.2006.07.008 prinsen, f., volman, m. l. l., & terwel, j. (2007). the influence of learner characteristics on degree and type of participation in a cscl environment. british journal of educational technology, 38, 1037– 1055. doi:10.1111/j.1467-8535.2006.00692.x rogat, t. k., & linnenbrink-garcia, l. (2011). socially shared regulation in collaborative groups: an analysis of the interplay between quality of social regulation and group processes. cognition and instruction, 29, 375–415. doi: 10.1080/07370008.2011.607930 salomon, g. (1994). differences in patterns: studying computer enhanced learning environments in s. vosniadou, e. de corte, & h. mandl (eds.), technology-based learning environments: psychological and educational foundations (pp. 79–88). berlin: springer-verlag. salomon, g., & globerson, t. (1989). when teams do not function the way they ought to. journal of educational research, 13, 89–99. doi:10.1016/0883-0355(89)90018-9 salomon, g., & perkins, d. n. (1998). individual and social aspects of learning. review of research in education, 23, 1–24. salonen, p., vauras, m., & efklides, a. (2005). social interaction – what can it tell us about metacognition and coregulation in learning? european psychologist, 10, 199–208. scardamalia, m., & bereiter, c. (1994). computer support for knowledge-building communities. the journal of the learning sciences, 3, 265–283. scardamalia, m., & bereiter, c. (2006). knowledge building: theory, pedagogy, and technology. in r. k. sawyer (ed.), the cambridge handbook of the learning sciences (pp. 97–115). new york: cambridge university press. scott, j. (1991). social network analysis: a handbook. london: sage. siegel, s., & castellan, n. j., jr. (1988). nonparametric statistics for the behavioral sciences (2nd ed.). new york: mcgraw–hill book. soller, a., martinez-monés, a., jermann, p., & muehlenbrock, m. (2005). from mirroring to guiding: a review of state of the art technology for supporting collaborative learning. international journal of artificial intelligence in education, 15, 261–290. strijbos, j-w., & stahl, g. (2007). methodological issues in developing a multi-dimensional coding procedure for small-group chat communication. learning and instruction, 17, 394–404. doi:10.1016/j.learninstruc.2007.03.005 summers, m., & volet, s.e. (2010). group work does not necessarily equal collaborative learning: evidence from observations and self-reports. european journal of psychology of education, 25, 473–492. doi: 10.1007/s10212-010-0026-5 iiskala'et'al' ' ' | f l r ! ! 109! suthers, d. d. (2006). technology affordances for intersubjective meaning making: a research agenda for cscl. computer supported collaborative learning, 1, 315–337. doi: 10.1007/s11412-006-9660-y toikkanen, t., & lipponen, l. (2011). the applicability of social network analysis to the study of networked learning. interactive learning environment, 19, 365–379. doi: 10.1080/10494820903281999 vauras, m., iiskala, t., kajamies, a., kinnunen, r., & lehtinen, e. (2003). shared-regulation and motivation of collaborating peers: a case analysis. psychologia: an international journal of psychology in the orient, 46(1), 19–37. vauras, m., salonen, p., & kinnunen, r. (2008). influences of group processes and interpersonal regulation on motivation, affect and achievement. in m. maehr, s. karabenick, & t. urdan (eds.), social psychological perspective: vol 15. advances in motivation and achievement (pp. 275–314). new york, ny: emerald group. vauras, m., & volet, s. (2013). the study of interpersonal regulation in learning and its challenge to the research methodology. in s. volet & m. vauras (eds.), interpersonal regulation of learning and motivation: methodological advances (pp. 1–13). london: routledge. veenman, m. v. j. (2005). the assessment of metacognitive skills: what can be learned from multi-method designs? in c. artelt & b. moschner (eds.), lernstrategien und metakognition: implikationen für forschung und praxis (pp. 77–99). berlin: waxmann. veenman, m. v. j. (2011). learning to self-monitor and self-regulate. in r. e. mayer, & p. a. alexander, (eds.), handbook of research on learning and instruction (pp. 197–218). new york, ny: routledge. veenman, m. v. j., & beishuizen, j. j. (2004). intellectual and metacognitive skills of novices while studying texts under conditions of text difficulty and time constraint. learning and instruction, 14, 621–640. doi:10.1016/j.learninstruc.2004.09.004 volet, s., vauras, m., khosa, d., & iiskala, t. (2013). metacognitive regulation in collaborative learning: conceptual developments and methodological contextualizations. in s. volet & m. vauras (eds.), interpersonal regulation of learning and motivation: methodological advances (pp. 67–101). new york, ny: routledge. volet, s., summers, m., & thurman, j. (2009a). high-level co-regulation in collaborative learning: how does it emerge and how is it sustained? learning and instruction, 19, 128–143. doi:10.1016/j.learninstruc.2008.03.001 volet, s., vauras, m., & salonen, p. (2009b). selfand social regulation in learning contexts: an integrative perspective. educational psychologist, 44, 215–226. doi: 10.1080/00461520903213584 vygotsky, l. s. (1978). mind in society. the development of higher psychological processes. cambridge, ma: harvard university press. wasserman, s., & faust, k. (1994). social network analysis: methods and applications. new york: cambridge university press. wertsch, j. w. (1977, may). metacognition and adult-child interaction. paper presented at northwestern university annual conference on learning disabilities, evanston, il. (eric document reproduction service no. ed180610). white, b. y., & frederiksen, j. r. (1998). inquiry, modeling, and metacognition: making science accessible to all students. cognition and instruction, 16(1), 3–118. whitebread, d., bingham, s., grau, v., pino pasternak, d. p., & sangster, c. (2007). development of metacognition and self-regulated learning in young children: role of collaborative and peer-assisted learning. journal of cognitive education and psychology, 6, 433–455. whitebread, d., coltman, p., pino pasternak, d., sangster, c., grau, v., bingham, s. et al. (2009). the development of two observational tools for assessing metacognition and self-regulated learning in young children. metacognition and learning, 4, 63–85. doi: 10.1007/s11409-008-9033-1 winne, p. h., hadwin, a., & perry, n. e. (2013). metacognition and computer-supported collaborative learning. in c. e. hmelo-silver, c. a. chinn, c. k. k. chan, & a. o’donnell (eds.), the international handbook of collaborative learning (pp. 462–479). new york, ny: routledge. iiskala'et'al' ' ' | f l r ! ! 110! zhang, j., scardamalia, m., lamon, m., messina, r., & reeve, r. (2007). socio-cognitive dynamics of knowledge building in the work of 9and 10-year old. educational technology research and development, 55, 117–145. doi: 10.1007/s11423-006-9019-0 zhang, j., scardamalia, m., reeve, r., & messina, r. (2009). design for collective cognitive responsibility in knowledge-building communities. the journal of the learning sciences, 18, 7–44. doi: 10.1080/10508400802581676 iiskala'et'al' ' ' | f l r ! ! 111! appendix description of the process (adapted from hakkarainen, 2003) order of phase in cscla phase lessons introduction face-to-face/ csclb content creating the context 1–2 introduction − giving instructions for the process − watching the documentary video of the universe − discussing interests based on the video and previous knowledge/experiences − forming small groups i setting up the research question 3–4 introduction cscl − giving instructions for the explanation-seeking research questions − setting up of small group’s research questions ii constructing a hypothesis 5–8 introduction cscl − giving instructions for constructing hypothesis/es − constructing small group’s hypothesis for research questions iii developing a work plan 9–10 introduction cscl − giving instructions for making a work plan − making small group’s work plan iv searching for and processing knowledge 11–15 introduction cscl − giving instructions for inquiring − searching in small group for knowledge and processing it v summarizing findings and concluding 16–20 introduction cscl − giving instructions for summarizing the findings and making conclusions − summarizing small group’s findings and drawing conclusions common discussion 21–22 csclc − listening to students outside the small group comment on the summarized findings and conclusions of the small group and the small group’s students reactions to the other students’ comments a phases i–v were analyzed. b introductions were face-to-face and not analyzed; ssmr was analyzed only during cscl. c the two last lessons (21– 22) were not analyzed because, in this phase, all students in the class connected to each other independently and, thus, the focus was not only on the selected small group’s process but also on individual students. frontline learning research 2 (2013) 33-52 issn 2295-3159 corresponding author: robert klassen, department of education university of york, york uk yo10 5dd, robert.klassen@york.ac.uk doi 33 | f l r measuring teacher engagement: development of the engaged teachers scale (ets) robert m. klassen a,b , sündüs yerdelen c , tracy l. durksen b a university of york, uk b university of alberta, edmonton, canada c middle east technical university, ankara, turkey and kafkas university, kars, turkey article received 27 june 2013 / revised 10 december 2013 / accepted 10 december 2013 / available online 20 december 2013 abstract the goal of this study was to create and validate a brief multidimensional scale of teacher engagement—the engaged teachers scale (ets)—that reflects the particular characteristics of teachers’ work in classrooms and schools. we collected data from three separate samples of teachers (total n = 810), and followed five steps in developing and validating the ets. the result of our scale development was a 16-item, 4-factor scale of teacher engagement that shows evidence of reliability, validity, and practical usability for further research. the four factors of the ets consist of: cognitive engagement, emotional engagement, social engagement: students, and social engagement: colleagues. the ets was found to correlate positively with a frequently used work engagement measure (the uwes) and to be positively related to, but empirically distinct from, a measure of teachers’ self-efficacy (the tses). our key contribution to the measurement of teacher engagement is the novel inclusion of social engagement with students as a key component of overall engagement at work for teachers. we propose that social engagement should be considered in future iterations of work engagement measures in a range of settings. keywords: teachers; engagement; scale validation; motivation 1. introduction a recurring theme of recent educational debate in public and research circles is the critical importance of providing all students with access to teachers who are highly engaged in their work (economist intelligence unit, 2012; pianta, hamre, & allen, 2012; rimm-kaufman & hamre, 2010; staiger & rockoff, 2010). klassen et al. 34 | f l r although work engagement research in business settings is thriving (bakker, albrecht, & leiter, 2011; sonnentag, 2003), the same attention has not been paid to the construct in education, at least partly due to the absence of context-relevant tools. building an understanding of teachers‘ engagement at work is vital: research shows that teachers‘ attitudes and motivation levels are transmitted to students (roth, assor, kanatmaymon, & kaplan, 2007). however, the most frequently used measure of work engagement (bakker et al., 2011)—the utrecht work engagement scale (uwes)—is designed for research involving workers in the business sector, and sharply contrasting work environments may demand dimensions of work engagement not currently covered in existing measures. shuck and colleagues noted ―an essential first step (to advance development of work engagement research) is a context-specific, conceptual exploration of the construct of employee engagement in relation to other well-researched job attitude(s)‖ (shuck, ghosh, zigarmi, & nimon, 2013, p. 11). thus, the purpose of this article is to report the design and validation of a teacher engagement scale that reflects the particular context and demands experienced by teachers working in classroom settings, and to explore the scale in relation to teachers‘ self-efficacy and to the frequently used work engagement scale, the uwes. work engagement is a motivation concept that refers to the voluntary allocation of personal resources directed at the range of tasks demanded by a particular vocational role (christian, garza, & slaughter, 2011). two core conceptual dimensions—energy and involvement—underpin work engagement (bakker et al., 2011), with three domains of engagement often posited: physical, emotional, and cognitive (e.g., saks, 2006). in some cases, these three domains are subsumed under a higher-order engagement construct, whereby the individual domains are experienced simultaneously or holistically (e.g., rich, lepine, & crawford, 2010; sonnentag, 2003). the relationship of engagement to burnout has been debated. in the view of some, engagement is the opposite of burnout, representing the other end of the continuum that stretches from fully engaged (low burnout) to not engaged (high burnout). recent research using the oldenburg burnout inventory (olbi; demerouti, mostert, & bakker, 2010), which simultaneously measures the energy and identification dimensions of engagement/burnout using positively and negatively worded items, provides equivocal results about the relationship of burnout and engagement. the creators of the olbi found that the identification dimension of burnout seemed to be opposite of the dedication dimension of engagement, whereas the energy dimensions of burnout (exhaustion) and engagement (vigour) operated as separate, but related, dimensions. existing engagement measures—such as the olbi and uwes—have the advantage of measuring engagement in a broad variety of settings, but have not been created to examine engagement in specific contexts, like teaching. creating a tailor made teacher engagement measure offers the advantage of including content that reflects the unique characteristics of teachers and the teaching context. engagement is considered to be relatively stable, with some fluctuations over time, reflecting both trait-like and state-like components (dalal, brummel, wee, & thomas, 2008; schaufeli, salanova, gonzalez-roma, & bakker, 2002). macey and schneider‘s (2008) review of the engagement literature and subsequent conceptualization of the construct suggests work engagement reflects the dispositions (feelings of energy) that lead to engaged behaviours (acting in an energetic fashion). engagement reflects motivational forces (e.g., intrinsic reasons for behaviour), but is conceptually distinct from these forces and from the ensuing behaviours (schaufeli & salanova, 2011); for example, the related construct of work commitment refers to an attitude of attachment to a job or career (e.g., meyer, allen, & smith, 1993; saks, 2006), but is conceptually separate from the feelings of energy during work time that defines engagement. work commitment refers to an attitude about work; work engagement refers to the degree of attention and absorption in work activities (shuck et al., 2013). work engagement has also shown discriminant validity from job attitudes (christian et al., 2011), and job involvement and satisfaction (rich et al., 2010). engagement has been shown to be related to self-efficacy; that is, beliefs in the capabilities to accomplish tasks in particular domains. xanthopoulou, bakker, demerouti, and schaufeli (2007) found that self-efficacy (along with optimism and organizational-based self-esteem) served as workplace resources that predicted engagement. in education settings, teachers‘ self-efficacy has been shown to be a potent motivational force associated with commitment to teaching and (inversely) to quitting intention (klassen & chiu, 2011), and to be robustly related to teacher resilience (gu & day, 2007). although there are close relationships between engagement and other work-related motivation constructs, there is support for empirical and conceptual klassen et al. 35 | f l r distinctiveness, and exploring the nomological web of relationships among key related variables results in a more nuanced picture of how people behave in the workplace. schaufeli and colleagues operationalised work engagement in their creation of the uwes (e.g., schaufeli, bakker, & salanova, 2006), and defined work engagement as an affective-cognitive state, not targeted at any particular work event or task. however, questions remain about the robustness of its factor structure (e.g., klassen et al., 2012; shimazu et al., 2008; sonnentag, 2003), and its item content may not be relevant for all contexts. for example, although the uwes has been used with teachers (e.g., bakker & bal, 2010; hakanen, bakker, & schaufeli, 2006), the scale content ignores the particular conditions associated with teachers‘ work. in particular, the uwes and other work engagement scales do not reflect the dimension of social engagement with students, a dimension which perhaps uniquely defines the act of teaching (jennings & greenberg, 2009). the work of teaching involves a level of demand for social engagement—energy devoted to establishing relationships—that is rarely found in other professions (e.g., pianta et al., 2012; roorda, koomen, spilt, & oort, 2011) and that is not included in other conceptual definitions of engagement (i.e., the uwes). although workers in many settings must engage socially with colleagues, teaching uniquely emphasises energy spent on the establishment of long-term, meaningful connections with the clients of the work environment (i.e., students) in a way that characterises the job of teaching. in fact, researchers propose that teacher-student relationships may play the primary role in fostering student engagement and positive student outcomes (davis, 2003; klassen, perry, & frenzel, 2012; pianta et al., 2012; wang, 2009). teachers who devote energy to forming warm and nurturing relationships with their students tend to experience higher levels of well-being, and less emotional stress and burnout (jennings & greenberg, 2009). to be sure, workers in other professions such as health (e.g., physicians, nurses, psychologists) or business (e.g., sales representatives), may form deep and meaningful relationships with their patients or clients, but rarely do workers in these fields spend the number of hours that most teachers spend with their students. like workers in other professions, teachers form social relationships with colleagues during work, but the emphasis on social relationships with students characterises the heart of the work of teaching; in fact, the opportunity to work closely with students is a strong motive for many teachers entering the profession (e.g., watt & richardson, 2007). measuring teachers‘ work engagement without capturing social engagement with students ignores one of the most important aspects of teacher engagement. shuck‘s recent review of work engagement (2011) concludes that the construct remains in a state of evolution, with disciplinary bridges needed between disparate communities of research. as educational psychologists, we question the fit of business-oriented work engagement models and measures to educational contexts, and see a clear need for a context-specific engagement measure tailored to the work performed by teachers. in this article, we address this need by creating and testing the engaged teacher scale (ets), in which workplace (i.e., classroom) engagement, comprising context-responsive physical, cognitive, and emotional dimensions (e.g., rich et al., 2010), is combined with social engagement with students and colleagues to represent teachers‘ overall engagement. 1.1 current study the goal of the study was to create and validate a usable (i.e., brief) scale of teacher engagement. we followed five steps involving three samples of teachers (total n = 810) in developing and validating the ets. in step 1 we developed item content, and received critical feedback from a focus group of experts. in steps 2 through 5 we collected data from three independent samples and conducted a series of statistical analyses designed to reduce the item pool, explore the factor structure, and examine the construct validity of the emerging scale. the result of our five steps is a 16-item, 4-factor scale of teacher engagement that shows evidence of reliability, validity, and usability for future research. klassen et al. 36 | f l r 2. step 1 step 1 consisted of creation of an item pool, and generation of feedback about the content of the item pool. to begin, our team of researchers (i.e., the three authors who represent disparate backgrounds— psychology, education, and educational psychology—and three countries) reviewed the existing literature and created and adapted item content through a process of generation, discussion, and revision. a comprehensive literature search revealed a number of theory-driven work engagement measures (e.g., rich, 2006; saks, 2006; schaufeli, et al., 2006; shuck, 2010; thomas, 2006; wang & qin, 2011). theoretical guidance from research by rich et al. (2010), kahn (1990, 1992), and schaufeli et al. (2006) provided the foundation for the dimensions of engagement (physical, cognitive, and emotional; or vigour, absorption, and dedication for the uwes). we also drew from teacher-student relatedness research (davis, 2003; klassen et al., 2012; pianta et al., 2012; wang, 2009) for generation of social engagement items. item development included adaptation of items from existing measures (e.g., at my work, i feel bursting with energy was adapted to when teaching, i feel bursting with energy), and creation of new items guided by theory (e.g., in class, i care about the problems of my students was an item reflecting social engagement: students). the proposed structure of the ets is presented in figure 1, with an over-arching engagement factor, and five second level dimensions: physical, cognitive, emotional, social: students, and social: colleagues. after reviewing the literature, an initial survey of 56 items was created and presented to 13 educational psychology graduate students, nine of whom were practicing teachers, during a graduate-level seminar. following an introduction to the engagement literature (e.g., discussion of the uwes; schaufeli et al., 2006), the students were given instructions to provide feedback on the content, wording, and plausibility of the initial item list. small groups (2-4 students) were formed to provide feedback on one dimension after which the students participated in a large group discussion of the item content. the items and item content were revised based on the feedback and discussion, with the resulting survey consisting of 48 items representing five factors. figure 1 presents the hypothesised dimensions of the ets, with initial number of items for each dimension, and item examples for each of the five dimensions. figure 1. hypothesised dimensions for the engaged teachers scale (ets). the number of initial items identified with each dimension is listed in parentheses, with example items listed in the following row. e n g a g e d t e a c h e r s c a le physical (7) i devote a lot of energy to teaching. cognitive (7) while teaching, i get absorbed in my work. emotional (12) i really put my heart into teaching. social: students (11) i connect well with my students. social: colleagues (11) i am accessible to my colleagues. klassen et al. 37 | f l r 3. step 2 in step 2, we administered the emergent 48-item measure to a sample of 224 practicing teachers, and analyzed the data using principle components analysis (pca) for item reduction purposes. although the use of pca has been criticised as a means of extracting factors (e.g., velicer & jackson, 1990), it is a preferred method for item reduction (conway & huffcutt, 2003; henson & roberts, 2006; matsunaga, 2010). 3.1 participants and procedures data for step 2 were collected at a compulsory teacher conference 1 in an urban/suburban setting with a population of about 1,000,000 in western canada. participants were volunteers who were recruited in the exhibition hall of the conference during breaks between professional development sessions. consenting teachers completed the paper-and-pencil survey on-site while research assistants kept notes on any verbal feedback offered during data collection. the sample for step 2 consisted of 224 teachers (74.6% female) between the ages of 23 and 65 years (m = 40.73 years). participants‘ highest level of education was reported as: undergraduate degree (73.4%), master‘s degree (22.5%), doctorate degree (0.9%), and 3.2% unspecified. most participants were employed full-time (84.8%) in urban 2 (77.5%), suburban (20.3%), and rural (2.3%) canadian schools. participants‘ school settings were elementary (43.3%), middle (17%), secondary (28%), and multiple (9%), with a mean class size of 26.6 students. participants typically rated the socioeconomic status of most students in their class as low to average (67.9%), with 26.7% reported as average-high to high (5.4% varied or unknown). teaching experience ranged from 0 to 38 years, with a mean of 13.42 (sd = 9.79) years of total teaching experience, and a mean of 5.05 years at their current school. most participants (48.7%) were early career (≤10 years experience), with 23.7% at mid-career stage (11-20 years), and 25.6% with more than 20 years of experience. before conducting analyses, we examined item correlations, and subsequently excluded three items from further analysis due to non-significant correlations with the other variables, leaving 45 items. we used pca with promax rotation (kappa set at 4) in order to derive a smaller number of items for subsequent steps. 3.2 step 2 results results of pca revealed several items that did not load on theoretically consistent components, as well as items that clearly loaded on more than one component. for example, the item ―i burst with energy while teaching‖ loaded on a component with items characterizing emotional engagement; however, the item was intended to characterise physical engagement. furthermore, items that did not load on components with an adequate number of items (at least three) were excluded. since the purpose of the pca in this step was not to explore the factor structure but to reduce items, the main focus of the analysis was item reduction. hence, rather than examining the number of components, we examined the emergence of principal components and the magnitude of component loadings, with a minimum component loading set at > .50. after inspecting conceptual fit of the items and the item loadings for each component, six items from three components and five items from one component were retained for further analyses. the loading of these items ranged between .61 and .98. in total, four components were extracted and retained, with a total of 23 items. items on two components—tentatively labelled as cognitive and physical engagement—did not extract separately as initially hypothesised. since we hypothesised physical engagement as an important facet of work engagement, we created an additional two items representing each of physical and cognitive engagement items for further analysis, resulting in 27 items available for analysis in step 3. 1 attendance at one of the regional annual two-day teacher conventions is mandatory for all of the approximately 30,000 public school teachers in the province. 2 the term ―urban‖ in a canadian context typically connotes geographical location (i.e., a large city or town), not sociological context (i.e., socioeconomic status level or ethnicity) as is sometimes the case in u.s.-based research. klassen et al. 38 | f l r 4. step 3 in step 3 we administered the emergent 27-item version of the scale to a new sample of 265 teachers and conducted exploratory factor analysis (efa) to test the scale‘s factor structure. 4.1 participants and procedures participants were recruited in a similar fashion to step 2, in a multi-district compulsory teacher conference at a different urban setting (population ~1,100,000) in the same western canadian province. the step 3 sample consisted of 265 teachers (68.7% female) between the ages of 21 and 68 years (m = 40.37 years). demographics—ses, teaching level, and teaching experience—were similar to those in step 2, with additional demographic information available from the authors. 4.2 step 3 results the 27 items from step 2 were analyzed using efa with principle axis factoring and promax rotation (kappa set at 4). results of the efa were first examined in terms of the appropriateness of the existing data for factor analysis. the kaiser-meyer-olkin measure of sampling adequacy was .92, suggesting that the data were appropriate for factor analysis. additionally, bartlett‘s test of sphericity,  2 (351) = 4402.20, p < .05, indicated that the population correlation matrix was not an identity matrix and suitable for factor analysis (field, 2009). we next followed three approaches to determine the number of factors to be retained. first, we examined kaiser‘s eigenvalues > 1.0 and scrutiny of the screen test. retaining factors with eigenvalues > 1.0 resulted in five factors and yielded 66.27% of the variance in respondents‘ scores. examination of the scree plot suggested four or five factors. although the eigenvalues > 1.0 rule and screen test are commonly used methods for determining number of factors, both are criticised for lack of reliability (e.g., ledesma & valero-mora, 2007; velicer & jackson, 1990). second, parallel analysis—based on statistical rather than mechanical rules—was used as an alternative and more accurate test to determine number of factors (ledesma & valero-mora, 2007; o‘connor, 2000; zwick & velicer, 1986). results from the parallel analysis suggested retention of four factors. third, efa was performed to compare 4and 5-factor solutions. only the 4-factor solution yielded interpretable factors. with the 5-factor solution, one item, ―in class, i am accessible to my students‖ created a factor by itself. in the 4-factor solution, this item loaded inappropriately (i.e., theoretically unjustifiable) on the factor that was extracted by cognitive engagement items. therefore, this item was excluded from the scale and the 4-factor solution was retained. as in step 2, cognitive and physical engagement items did not produce separate factors; since cognitive items dominated the content, we labelled the factor cognitive engagement. examining the factor pattern coefficients with the cut-off point set at .70 resulted in eight more items eliminated from the scale. however, two borderline-case items with coefficients between .50 and .70 were retained since the item content made the factors more representative in terms of the construct being measured. two items with redundant content were considered: ―at school, i value the relationships i build with my colleagues,‖ and ―at school, i value spending time with my colleagues.‖ we excluded the latter item due to lower factor loading (.82 versus .92 for the former item). as a result of these procedures, the scale was reduced to 16 items with four items in each of four factors. table 1 lists the pattern and structure coefficients of items for the related factors. the final version of the ets with item content of each engagement dimension is presented in the appendix. the efa resulted in four factors accounting for 71.31% of the variance in the respondents‘ scores. the first factor was named emotional engagement (ee), accounting for 40.25% of the variance in the correlation matrix. the other three factors were social engagement: colleagues (sec), cognitive engagement (ce), and social engagement: students (ses) accounting for 13.84%, 9.56%, and 7.66% of the variance, respectively. correlations between factors ranged from .33 to .62. cronbach‘s alpha coefficients for the ee, sec, ce and ses factors were .89, .85, .85, and .84, respectively. klassen et al. 39 | f l r table 1 factor pattern and structure coefficients in descending order (efa, promax rotation) for the four-factor model of ets item content factor ee sec ce se 10 i love teaching .95 (.89) 2 i am excited about teaching .80 (.81) 5 i feel happy while teaching .72 (.83) 13 i find teaching fun .70 (.76) 9 at school, i value the relationships i build with my colleagues .88 (.83) 7 at school, i am committed to helping my colleagues .83 (.83) 12 at school, i care about the problems of my colleagues .79 (.82) 1 at school, i connect well with my colleagues .57(.58) 11 while teaching i pay a lot of attention to my work .82 (.82) 8 while teaching, i really ―throw‖ myself into my work .77 (.80) 15 while teaching, i work with intensity .76 (.76) 4 i try my hardest to perform well while teaching .65 (.71) 14 in class, i care about the problems of my students .87 (.82) 16 in class, i am empathetic towards my students .79 (.83) 6 in class, i am aware of my students‘ feelings .75 (.73) 3 in class, i show warmth to my students .53 (.65) note. factor structure coefficients were included in the parenthesis. ee = emotional engagement, sec = social engagement: colleagues, ce = cognitive engagement, ses = social engagement: students. 5. step 4 in steps 4 and 5 we administered the final version of the scale to 321 teachers and analyzed the data using firstand second-order confirmatory factor analyses (cfa) for the purpose of testing construct validity. in particular, step 4 was performed to validate the factor structure of the ets. 5.1 participants and procedures data were collected at compulsory teachers‘ convention in an adjacent province. demographic information was similar to the samples in steps 2-3 and is available from the authors. klassen et al. 40 | f l r 5.2 step 4 results a series of cfas was performed in step 4 to test the factor structure of the ets. first, we performed cfa on the 16 items and 4 factors (model 1). second, we tested models with and without social engagement by testing models that excluded factors representing social engagement with students (ses, model 2) and social engagement with colleagues (sec, model 3). finally, a second-order cfa was performed to examine whether the four first-order factors could be explained by a second-order teacher engagement (te) factor (model 4). we used lisrel 8.80 (jöreskog & sörbom, 2006) with simplis command language to conduct cfa. we used a series of fit indices to evaluate the model fit in addition to the conventional use of chisquare (see kline, 2005): comparative fit index (cfi), normed fit index (nfi), goodness-of-fit index (gfi), and root mean square error of approximation (rmsea). since the level of missing data was low (1.8%), we replaced missing values with means (tabachnick & fidel, 2007). data were checked for multivariate normality through inspection of univariate and multivariate outliers (kline, 2005), with eight cases excluded as a result. skewness and kurtosis values were checked and absolute values were found within the ranges .40 1.0 and .03 .45, respectively. the maximum likelihood approach was selected to estimate the parameters of the model (chou & bentler, 1995). 5.2.1 model 1: four first-order factors the 16-item scale was subjected to first-order cfa to test the four-factor structure of ets. results demonstrated a good fit to the data ( 2 (98) = 292.67, p < .05; cfi = .97; gfi = .90; nfi = .96; rmsea = .08; 90% ci = .07, .09). standardised parameter estimates for each item of the four-factor ets model are listed in table 2. as presented in the table, all of the standardised estimates (ranging from .66 to .85) were significant and above a cut-off value of .50 (hair, black, babin, anderson, & tatham, 2010). table 3 presents the correlations (phi estimates) among the four factors. as seen in the table, correlations ranged between .49 and .73, and were significant at the p < .01 level. internal consistencies of each subscale of ets were examined, with cronbach‘s alpha coefficients at .84, .87, .83, and .79 for ce, ee, ses, and sec, respectively. table 4 presents the means, standard deviations, and reliability coefficients for the four factors. these findings supported our initial prediction of a first-order factor structure for teacher engagement. since we proposed the novel hypothesis that social engagement was a dimension of teacher engagement, we tested models 2 and 3 that examined the validity of including social engagement dimensions with students and colleagues in our model of teacher engagement. klassen et al. 41 | f l r table 2 standardised parameter estimates for the first-order factor solution for the ets (model 1) item content factor λ 4 i try my hardest to perform well while teaching ce .72 8 while teaching, i really ―throw‖ myself into my work ce .80 11 while teaching i pay a lot of attention to my work ce .75 15 while teaching, i work with intensity ce .74 2 i am excited about teaching ee .78 5 i feel happy while teaching ee .75 10 i love teaching ee .85 13 i find teaching fun ee .80 3 in class, i show warmth to my students ses .71 6 in class, i am aware of my students‘ feelings ses .69 14 in class, i care about the problems of my students ses .74 16 in class, i am empathetic towards my students ses .81 1 at school, i connect well with my colleagues sec .66 7 at school, i am committed to helping my colleagues sec .68 9 at school, i value the relationships i build with my colleagues sec .85 12 at school, i care about the problems of my colleagues sec .66 note. ce = cognitive engagement, ee = emotional engagement, ses= social engagement: students, sec = social engagement: students. all coefficients were significant, p < .05. table 3 factor correlations (phi estimates) of model 1 2 3 4 1. ce .73** .73** .49** 2. ee .64** .53** 3. ses .52** 4. sec note. ce = cognitive engagement, ee = emotional engagement, ses= social engagement: students, sec = social engagement: students. **p < .001. klassen et al. 42 | f l r table 4 means, standard deviations, and reliability coefficients for factors of ets factors mean sd α te (composite) 5.07 .56 .91 ce 5.16 .65 .84 ee 5.05 .73 .87 ses 5.26 .60 .83 sec 4.80 .80 .79 note. te = teacher engagement, ce = cognitive engagement, ee = emotional engagement, ses= social engagement: students, sec = social engagement: colleagues. 5.2.2 model 2: three first-order factors, ses excluded model 2 was constructed to test if a 3-factor structure without ses provided a better fit to the data than the full 4-factor structure. the purpose of this procedure was to examine the contribution of teacher‘ social engagement with students to explain their general work engagement. this model showed good fit to the data ( 2 (51) = 155.65, p < .05; cfi = .97; gfi = .93; nfi = .96; rmsea = .08; 90% ci = .07, .09). model 2 was compared to model 1 using the chi-square difference test. the δ 2 value of 137.02 (δdf = 47) was significant, indicating that model 2 was a significantly poorer fit for the data than model 1. 5.2.3 model 3: three first-order factors, sec excluded in model 3, we excluded the social engagement: colleagues (sec) factor from the 4-factor ets. the model was compared with model 1 to test the role of teachers‘ relationship with colleagues in teacher engagement. although model 3 showed an adequate fit to the data ( 2 (51) = 179.33, p < .05; cfi = .97; gfi = .91; nfi = .96; rmsea = .09; 90% ci = .08, .11), the chi-square difference test between model 1 and model 3 revealed a significantly poorer fit for the model 3 data (δ 2 = 113.34, δdf = 47). thus we concluded that social engagement with students and peers were viable dimensions with which to measure teacher engagement. 5.2.4 model 4: second-order factor the high reliabilities and intercorrelations found in the first-order factor structure of ets suggested the possibility of a second-order factor. therefore, a second-order cfa was conducted to examine whether the four-factor ets could be represented by a superordinate factor labelled teacher engagement. figure 2 presents the first order and second order models in graphic format. the fit indices for the second-order factor ( 2 (100)= 296.94, p < .05; cfi = .97; gfi = .89; nfi = .95; rmsea = .08; 90% ci = .07, .09) suggested that the hypothesised model fit the data well. as shown in table 5, all first-order factors significantly loaded on the second-order factor and their standardised coefficients were above the .50 cut-off suggested by hair et al. (2010). a chi-square difference test conducted between models 1 and 4 revealed no significant difference, suggesting the viability of an underlying single factor in addition to valid use of the four subscale scores. a summary of the goodness of fit indices for the four models is presented in table 6. thus, results suggested that using the four-factor or single factor models was viable for measuring teacher engagement. klassen et al. 43 | f l r table 5 standardised parameter estimates for the second-order factor solution for the ets (model 4) second-order factor first-order factors γ te ce .88 te ee .82 te ses .82 te sec .61 note. te = teacher engagement, ce = cognitive engagement, ee = emotional engagement, ses= social engagement: students, sec = social engagement: colleagues. all coefficients were significant, p < .05. table 6 goodness of fit indices for the four models model χ 2 df χ 2 /df rmsea cfi gfi nfi model comparison δ χ 2 δ df four first-order factor model (model 1) 292.67 98 2.99 .08 .97 .90 .96 three first-order factor model, ses excluded (model 2) 155.65 51 3.85 .08 .97 .93 .96 2 vs. 1 137.02** 47 three first-order factor model, sec excluded (model 3) 179.33 51 3.52 .09 .97 .91 .96 3 vs. 1 113.34** 47 one second-order factor model (model 4) 296.94 100 2.97 .08 .97 .89 .95 4 vs. 1 4.27 2 note. df = degrees of freedom; rmsea = root mean square error approximation; cfi = comparative fit index, gfi = goodness-of-fit index, nfi = normed fit index. **p < .001. klassen et al. 44 | f l r figure 2. first-order and second-order factor structures for ets (ee = emotional engagement, sec = social engagement: colleagues, ce = cognitive engagement, ses= social engagement: students, te = teacher engagement). 6. step 5 in step 5 we conducted canonical and zero-order correlation analyses to further test the construct validity of the scale. canonical correlation analysis examines commonalities in sets of factors from different variables by providing linear combinations of each set of the factors (hair, et. al., 2010). we examined correlations between the ets and two related measures: the utrecht work engagement scale (uwes; schaufeli et al., 2006) and the teacher sense of efficacy scale (tses; tschannen-moran & woolfolk hoy, 2001), a teacher motivation variable that taps teachers‘ expectancies of success in the classroom. the tses consists of three factors: self-efficacy for student engagement (se), instructional strategies (is), and classroom management (cm). the scale has been shown to be valid in a range of settings and to be related to positive teacher outcomes such as teacher commitment and inversely with quitting intention (klassen & chiu, 2011). schaufeli et al. (2006) found professional efficacy to be strongly related to work engagement across international contexts. 6.1 procedure the sample consisted of the same 321 participants described in step 4. to start, we used cfa to ensure the factor structure of the uwes and tses with our sample. results for the 3-factor uwes showed adequate fit to the data ( 2 (24)= 78.51, p < .05; cfi = .98; gfi = .95; nfi = .97; rmsea = .08; 90% ci = ce item4 item8 item11 item15 ee item2 item5 item10 item13 ses item3 item6 item14 item16 sec item1 item7 item9 item12 ce item4 item8 item11 item15 ee item2 item5 item10 item13 ses item3 item6 item14 item16 sec item1 item7 item9 item12 te model 1 (first-order, four factors) model 4 (second-order, one factor) klassen et al. 45 | f l r .06, .10). all factor loadings were significant and internal consistencies of each subscale raged from .74 to .78. results for the 3-factor tses also indicated good model fit ( 2 (41)= 112.90, p < .05; cfi = .98; gfi = .94; nfi = .97; rmsea = .07; 90% ci = .06, .09). all factor loadings were significant, with reliability coefficients above .80. 6.2 step 5 results the relationship of the ets with the tses and uwes scales was assessed through canonical correlation analyses (see table 7). the first canonical analysis (ets and tses) yielded three canonical variate pairs. a canonical correlation of .58 (33% overlapping variance),  2 (12) = 149.02, p < .001, was found for the first canonical variate, and .25 (6% overlapping variance),  2 (6) = 22.03, p < .05, for the second canonical variate. while the first two pairs of canonical variates accounted for the significant relationship, the  2 test was not statistically significant for the third pair. since the overlapping variance for the second variate was very low (i.e., < 10%, see tabachnick & fidell, 2007), only the result of the first pair is reported. as shown in table 7, with a cut off value set at .30 (tabachnick & fidell, 2007), all variables had significant relationship with the first canonical variate. thus, the first canonical analysis suggests positive relationships between all teacher engagement variables and teacher self-efficacy variables. table 7 correlations, standardised canonical coefficients, canonical correlations, percentages of variance, and redundancies between self-efficacy and engagement variables first canonical variate variables correlation coefficient set 1 (tses) student engagement -.84 -.57 instructional strategies -.89 -.66 classroom management -.50 -.13 percent of variance .59 redundancy .19 set 2 (ets) ce -.87 -.45 ee -.75 -.23 ses -.89 -.56 sec -.39 -.17 percent of variance .57 redundancy .19 canonical correlations .58 note. ce = cognitive engagement, ee = emotional engagement, ses= social engagement: students, sec = social engagement: colleagues. klassen et al. 46 | f l r the second canonical correlation analysis was performed between the set of ets variables and the set of uwes variables (vigour, dedication, and absorption). this second analysis yielded only two of the three variates as significant (see table 8). the first canonical correlation was .73 (i.e., 53% overlapping variance),  2 (12) = 286.92, p < .001. the second canonical correlation was .37 (14% overlapping variance),  2 (6) = 46.37, p < .05. in light of the high overlapping variance and the modest overlapping variance of the second correlation, only the first variate was taken into account. all factors had significant relationships (i.e., above .30 as suggested by tabachnick & fidell, 2007) with the first canonical variate suggesting a positive relationship between the ets and uwes variables. therefore, based on the results of two canonical correlation analyses, it can be concluded that teachers with high engagement scores on ets tend to gain high score on the tses and uwes. the zero-order correlation matrix (table 9) confirms this finding with all pairs of factors showing significant relationships. cognitive engagement showed the strongest correlations with absorption (r = .63, uwes) and student engagement (r = .48, tses); emotional engagement was most strongly related to dedication (r = .67, uwes) and student engagement (r = .39, tses); social engagement: students showed the strongest relationship with absorption (r = .42, uwes) and instructional strategies (r = .45, tses); and social engagement: colleagues showed the strongest relationship with dedication (r = .37, uwes) and student engagement (r = .26, tses). table 8 correlations, standardised canonical coefficients, canonical correlations, percentages of variance, and redundancies between the uwes and ets subscales first canonical variate variables correlation coefficient set 1 (uwes) vigour -.80 -.09 dedication -.94 -.58 absorption -.89 -.43 percent of variance .77 redundancy .41 set 2 (ets) ce -.86 -.46 ee -.93 -.65 ses -.61 -.06 sec -.53 -.07 percent of variance .57 redundancy .30 canonical correlations .73 note. ce = cognitive engagement, ee = emotional engagement, ses= social engagement: students, sec = social engagement: colleagues. klassen et al. 47 | f l r table 9 zero-order correlation coefficients between ets variables and uwes and tses variables uwes tses vigour dedication absorption instructional strategies student engagement classroom management ce .43** .54** .63** .38** .48** .28** ee .59** .67** .55** .38** .39** .22** ses .32** .41** .42** .45** .44** .27** sec .31** .37** .33** .15** .26** .24** note. ce = cognitive engagement, ee = emotional engagement, ses= social engagement: students, sec = social engagement: colleagues. **p < .001 7. discussion recent discussions about ways to improve social and educational outcomes have focused on the critical role played by teachers. rarely before has so much emphasis been placed on understanding the psychological make-up of effective teachers (rimm-kaufman & hamre, 2010; staiger & rockoff, 2010). from a psychological viewpoint, effective teaching is dependent on teachers who are motivated: fully engaged in their work, and engaged not just cognitively and emotionally, but also socially. our study‘s aim was to respond to the call for better understanding of teacher engagement by creating a reliable, valid, and usable multi-dimensional measure of work engagement that was specifically targeted at the work carried out by teachers in classrooms and schools. from a measurement perspective, the findings from this research provide support for the reliability and validity of the ets. in particular, the item statistics and reliabilities of the ets are very good, and the four factors represent appropriate measures of the internal structure of teacher engagement. furthermore, the analyses show that the ets factors are discrete, reliable, and valid. in general, the results suggest that measures of teacher engagement should incorporate the component factors of engagement, and that the factors are related to an overarching engagement factor. from a theoretical perspective, the findings show that social engagement with students and with colleagues should be considered as important dimensions of teacher engagement, alongside cognitive and emotional dimensions of engagement. our primary contribution to future research is in the creation of a four-factor teacher engagement measure that is practical (i.e., brief), valid, reliable, and that reflects the context of educational settings. our multiple steps of analyses resulted in a robust measure that correlates positively with a frequently used work engagement measure (the uwes), and is positively related to, but empirically distinct from, teachers‘ self-efficacy. the inclusion of social engagement is novel for conceptualizing and measuring work engagement, but the conceptual framework for work engagement is still developing (bakker et al., 2011; shuck et al., 2013), and conceptualizations that challenge how engagement is defined across contexts may contribute to a more general understanding of how the construct operates in diverse vocational settings. we know that social engagement with students is a fundamental aspect of teachers‘ work (e.g., pianta et al., 2012), and perhaps reflects the key mechanism through which student development is influenced. although conceptualizations of engagement that consist of dimensions of physical, cognitive, and emotional energy and involvement at work have been conventionally proposed, the results from our study suggest that social engagement—with students and with colleagues—forms an important dimension of overall engagement for teachers. we klassen et al. 48 | f l r suggest that a dimension representing social engagement is worth considering for future iterations of work engagement measures applied to a wide range of vocational settings. we failed to find separate domains of physical and cognitive engagement in our samples of teachers, and the question remains whether physical engagement is separable from cognitive, emotional, and social dimensions of teacher engagement. hakanen et al. (2006) proposed vigour (physical engagement) and dedication (emotional engagement) as the core dimensions of engagement in their study of the uwes with a group of teachers, but they did not test the hypotheses by including a cognitive dimension in their analyses. we did not find clear support for the separation of physical and cognitive engagement dimensions, and propose that for teachers, the line between the two is blurred. for example, we labelled ―i try my hardest to perform well while teaching‖ and ―while teaching, i really ‗throw‘ myself into my work‖ as examples of cognitive engagement, but the demands of individual teachers‘ classroom work may determine the relevance of particular dimensions for teachers. a teacher of young children may need to physically interact with students (crouching down, tying shoes, performing actions during music sessions) more often than a high school history teacher, thus increasing the salience of the physical engagement dimension for some teachers. hakanen et al. describe the physical job demands and resources that can be associated with engagement, but the level of physical demands for teachers varies as a function of the setting. further work should focus on teasing apart teachers‘ physical and cognitive engagement by exploring the two dimensions in a wider range of contexts. more work is needed to understand how engagement is fostered in teachers, and especially how the specific dimensions—emotional, cognitive, social, and perhaps physical engagement—develop through teacher training and into professional practice. research from related constructs such as teacher resilience (gu & day, 2007), self-efficacy (klassen & chiu, 2010, 2011), and commitment (collie, shapka, & perry, 2011) have shown that teacher motivation constructs change in predictable ways over the course of a career. the ets provides a way of measuring individual facets of engagement and how the facets change over time: for example, a teacher may exhibit high levels of social engagement at the beginning of a career but lower levels of cognitive engagement. we know that teacher engagement changes over even brief periods of time: recent research has shown that global teacher engagement shows weekly within-person variability in starting teachers (bakker & bal, 2010; durksen & klassen, 2012), with commitment to the profession mirroring the pattern of change in engagement. the job-demands resources model (jdr; e.g., bakker, hakanen, demerouti, & xanthopoulou, 2007; hakanen et al., 2006) provides a general way of conceptualizing the drivers of engagement, but details about how job resources in classrooms and schools—supportive climate, transformational leadership, access to information, job control—can be targeted at fostering specific engagement dimensions have not yet been studied. multilevel analyses of teacher engagement may provide insight into how engagement might be shared in a school, and how teachers working together transmit their engagement amongst themselves, and to their students. psychosocial research in a range of vocational contexts has shown that workers regularly share beliefs, emotions, and motivational patterns, and that social interaction influences individual psychology (e.g., bandura, 1997). 7.1 limitations although we found strong psychometric properties of the ets, and collected data from three independent samples of teachers, there are some clear limitations. the participants were all working in two western provinces in canada, and were largely female, and thus the samples offer only limited representativeness to other populations. the data collected were cross-sectional, and although engagement is said to possess state and trait characteristics (e.g., schaufeli & salanova, 2011), engagement fluctuates over time (e.g., bakker & bal, 2010). external validity of the measure is limited by the correlational nature of the study design, and no objective measure of teaching effectiveness or of student achievement was used as an outcome measure, a clear direction for future research. teacher engagement may lead to positive teacherstudent interactions, increased student engagement, and eventually to increased student achievement, but the evidence base needs developing. it must also be considered that the relationships among teacher engagement, klassen et al. 49 | f l r teacher-student interactions, student engagement, and student achievement are reciprocal: it is likely that teacher engagement both influences, and is influenced by, positive experiences of teacher-student interaction. although the ets focuses on in-classroom and in-school engagement, for some (but not all) teachers, out-of-school activities involving parents and the community form an important component of their social engagement. other measures, such as the olbi and uwes measure engagement more broadly, and their use may be preferable for cross-professional comparisons and to capture teachers‘ non-classroom related engagement. 7.2 conclusions and future research understanding teacher engagement is critical to understanding the psychological processes underlying effective teaching. our aim was to create a measure of teacher engagement that reflects the particular features of working in classrooms and in schools, and especially the social interactions shared by teachers and students. an important step in understanding effective teaching is to conceptualise and measure teacher engagement, and we hope that the ets can be useful in this regard, but our knowledge of how teachers‘ selfreports of engagement are reflected by behaviours in real classrooms is limited. although data from observation systems (e.g., the class from pianta et al., 2012) provide some insight into how engaged and effective teachers behave, such methods still leave interpretation of teachers‘ behaviours to the presence of external observers sitting in classes for relatively brief periods of time. further study is needed to identify the behavioural indicators of teacher engagement, and how these behaviours develop individually and collectively, and change over time. bakker and bal‘s 2010 study on weekly fluctuations of teacher engagement provides a useful starting point, but examining work engagement using finer-grained time spans may provide a valuable way forward in understanding teachers and teaching. creation of the ets may be a useful point of departure for better understanding teacher engagement, and by extension, student engagement and learning. keypoints we created and validated a 4-factor 16-item measure of teacher engagement: the engaged teachers scale (ets). the five steps of development resulted in a multidimensional measure that is practical (i.e., brief), valid, and reliable for use in education settings. the four factors were cognitive engagement, emotional engagement, social engagement: students, and social engagement: colleagues. acknowledgements the authors would like to gratefully acknowledge research funding from the social sciences and humanities research council of canada provided to the first author. references bandura, a. (1997). self-efficacy: the exercise of control. new york: w.h. freeman. bakker, a. b., albrecht, s. l., & leiter, m. p. (2011). key questions regarding work engagement. european journal of work and organizational psychology, 20, 4-28. doi:10.1080/1359432x.2010.485352 klassen et al. 50 | f l r bakker, a., & bal, m. (2010). weekly work engagement and performance: a study among starting teachers. journal of occupational and organizational psychology, 83, 189-206. doi:10.1348/096317909x402596 bakker, a. b., hakanen, j. j., demerouti, e., & xanthopoulou, d. (2007). job resources boost work engagement, particularly when job demands are high. journal of educational psychology, 99, 274– 284. doi: 10.1037/0022-0663.99.2.274 christian, m. s., garza, a. s., & slaughter, j. e. (2011). work engagement: a quantitative review and test of its relations with task and contextual performance. personnel psychology, 64, 89-136. doi:10.1111/j.1744-6570.2010.01203.x chou, c. p., & bentler, p. m. (1995). estimates and tests in structural equation modeling. in r. h. hoyle (ed.), structural equation modeling: concepts, issues, and applications (pp. 37-55), thousand oaks, ca: sage. collie, r. j., shapka, j. d., & perry, n. e. (2011). predicting teacher commitment: the impact of school climate and social-emotional learning. psychology in the schools, 48, 1034-1048. doi:10.1002/pits.20611 conway, j. m., & huffcutt, a. i. (2003). a review and evaluation of exploratory factor analysis practices in organizational research. organizational research methods, 6, 147-168. doi:10.1177/1094428103251541dalal, r. s., brummel, b. j., wee, s., & thomas, l. l. (2008). defining employee engagement for productive research and practice. industrial and organizational psychology, 1, 52-55. doi:10.1111/j.1754-9434.2007.00008.x davis, h. a. (2003). conceptualizing the role and influence of student– teacher relationships on children‘s social and cognitive development. educational psychologist, 38, 207–234. doi:10.1207/ s15326985ep3804_2 demerouti, e., mostert, k., & bakker, a. b. (2010). burnout and work engagement: a thorough investigation of the independency of both constructs. journal of occupational health psychology, 15, 207-234. doi:10.1037/a0019408 durksen, t. l., & klassen, r. m. (2012). pre-service teachers‘ weekly commitment and engagement during a final training placement: a longitudinal mixed methods study. educational and child psychology, 29, 32-46. economist intelligence unit (2012). the learning curve: lessons in country performance in education. retrieved from: http://thelearningcurve.pearson.com/ field, a. (2009) discovering statistics using spss (3rd ed.). london: sage. gu, q., & day, c. (2007). teachers‘ resilience: a necessary condition for effectiveness. teaching and teacher education, 23, 1302-1316. doi:10.1016/j.tate.2006.06.006 hair, j. f., black, b., babin, b., anderson, r. e., & tatham, r. l. (2010). multivariate data analysis. (7th ed.). upper saddle river, nj: prentice-hall. hakanen, j. j., bakker, a. b., & schaufeli, w. b. (2006). burnout and work engagement among teachers. journal of school psychology, 43, 495-513. doi:10.1016/j.jsp.2005.11.001 henson, r. k., & roberts, j. k. (2006). use of exploratory factor analysis in published research: common errors and some comment on improved practice. educational and psychological measurement, 66, 393-416. doi: 10.1177/0013164405282485 jennings, p. a., & greenberg, m. t. (2009). the prosocial classroom: teacher social and emotional competence in relation to student and classroom outcomes. review of educational research, 79, 491-525. doi:10.2307/40071173 jöreskog, k. g. & sörbom, d. (2006). lisrel 8.80 for windows [computer software]. lincolnwood, il: scientific software international, inc. kahn, w. a. (1990). psychological conditions of personal engagement and disengagement at work. academy of management journal, 33, 692-724. doi: 10.2307/256287kahn, w. a. (1992). to be fully there: psychological presence at work. human relations, 45, 321-349. doi: 10.1177/00187267920450040 klassen, r. m., al-dhafri, s., mansfield, c. f., purwanto, e., siu, a., wong, m. w., & woods-mcconney, a. (2012). teachers‘ engagement at work: an international validation study. journal of experimental education, 80, 1-20. doi: 10.1080/00220973.2012.678409 http://psycnet.apa.org/doi/10.1037/0022-0663.99.2.274 http://psycnet.apa.org/doi/10.1037/a0019408 klassen et al. 51 | f l r klassen, r. m., bong, m., usher, e. l., chong, w. h., huan, v. s., wong, i. y., & georgiou, t. (2009). exploring the validity of the teachers‘ self-efficacy scale in five countries. contemporary educational psychology, 34, 67-76. doi:10.1016/j.cedpsych.2008.08.001 klassen, r. m., & chiu, m. m. (2010). effects on teachers‘ self-efficacy and job satisfaction: teacher gender, years of experience, and job stress. journal of educational psychology. 102, 741-756. doi:10.1037/a0019237 klassen, r. m., & chiu, m. m. (2011). the occupational commitment and intention to quit of practicing and pre-service teachers: influence of self-efficacy, job stress, and teaching context. contemporary educational psychology, 36, 114-129. doi:10.1016/j.cedpsych.2011.01.002 klassen, r. m., perry, n. e., & frenzel, a. c. (2012). teachers‘ relatedness with students: an underemphasized component of teachers‘ basic psychological needs. journal of educational psychology, 104, 150-165. doi: 10.1037/a0026253 kline, r. b. (2005). principles and practice of structural equation modeling (2nd ed.). new york: guilford. ledesma, r. d., & valero-mora, p. (2007). determining the number of factors to retain in efa: an easy-touse computer program for carrying out parallel analysis. practical assessment, research & evaluation, 12(2), 2-11. retrieved from http://pareonline.net/pdf/v12n2.pdf macey, w. h., & schneider, b. (2008). the meaning of employee engagement. industrial and organizational psychology, 1, 3-30. doi: 10.1111/j.1754-9434.2007.0002.x matsunaga, m. (2010). how to factor-analyze your data right: do‘s, don‘ts, and how-to‘s. international journal of psychological research, 3, 97-110. retrieved from http://www.redalyc.org/articulo.oa?id=299023509007 meyer, j. p., allen, n. j., & smith, c. a. (1993). commitment to organizations and occupations: extension and test of a three-component conceptualization. journal of applied psychology, 78, 538–551. doi: 10.1037/0021-9010.78.4.538 o‘connor, b. p. (2000). spss and sas programs for determining the number of components using parallel analysis and velicer‘s map test. behavior research methods, instruments, & computers, 32, 396402. doi: 10.3758/bf03200807 pianta, r. c., hamre, b. k., & allen, j. p. (2012). teacher-student relationships and engagement: conceptualizing, measuring, and improving the capacity of classroom interactions (pp. 365-386). in s. l. christenson, a. l. reschly, & c. wylie (eds.), handbook of research on student engagement. dordrecht, netherlands: springer. doi:10.1007/978-1-4614-2018-7 rich, b. l. (2006). job engagement: construct validation and relationships with job satisfaction, job involvement, and intrinsic motivation (doctoral dissertation, university of florida). retrieved from econis, proquest umi 3228825 dissertation publishing. rich, b. l., lepine, j. a., & crawford, e. r. (2010). job engagement: antecedents and effects on job performance. the academy of management journal, 53, 617-635. doi:10.5465/amj.2010.51468988 rimm-kaufman, s. e., & hamre, b. k. (2010). the role of psychological and developmental science in efforts to improve teacher quality. teachers college record, 112, 2988-3023. retrieved from http://www.tcrecord.org/library roorda, d. l., koomen, h. m. y., spilt, j. l., & oort, f. j. (2011). the influence of affective teacherstudent relationships on students‘ school engagement and achievement: a meta-analytic approach. review of educational research, 81, 493-529. doi:10.3102/0034654311421793 roth, g., assor, a., kanat-maymon, y., & kaplan, h. (2007). autonomous motivation for teaching: how self-determined teaching may lead to self-determined learning. journal of educational psychology, 99, 761-774. doi: 10.1037/0022-0663.99.4.761 saks, a. m. (2008). the meaning and bleeding of employee engagement: how muddy is the water? industrial and organizational psychology, 1, 40-43. doi: 10.1111/j.1754-9434.2007.00005.x schaufeli, w. b., bakker, a. b., & salanova, m. (2006). the measurement of work engagement with a short questionnaire: a cross-national study. educational and psychological measurement, 66, 701-716. doi: 10.1177/0013164405282471 schaufeli, w., & salanova, m. (2011). work engagement: on how to better catch a slippery concept. european journal of work and organizational psychology, 20, 39-46. doi: 10.1080/1359432x.2010.515981 http://pareonline.net/pdf/v12n2.pdf http://psycnet.apa.org/doi/10.1037/0021-9010.78.4.538 http://www.tcrecord.org/library klassen et al. 52 | f l r schaufeli, w.b., salanova, m., gonzalez-roma, v., & bakker, a. b. (2002). the measurement of engagement and burnout: a two-sample confirmatory factor analytic approach. journal of happiness studies, 3, 71–92. doi: 10.1023/a:1015630930326 shimazu, a., schaufeli, w.b., kosugi, s., suzuki, a., nashiwa, h., kato, a., et al. (2008). work engagement in japan: development and validation of the japanese version of the utrecht work engagement scale. applied psychology: an international review, 57, 510-523. doi: 10.1111/j.14640597.2008.00333.x shuck, m. b. (2010). employee engagement: an examination of antecedent and outcome variables (doctoral dissertation, florida international university). retrieved from http://digitalcommons.fiu.edu/etd/235 shuck, b. (2011). four emerging perspectives of employee engagement: an integrative literature review. human resource development review, 10, 304-328. doi:10.1177/1534484311410840 shuck, b., ghosh, r., zigarmi, d., & nimon, k. (2013). the jingle jangle of employee engagement: further exploration of the emerging construct and implications for workplace learning and performance. human resource development review, 12, 11-35. doi:10.1177/1534484312463921 sonnentag, s. (2003). recovery, work engagement, and proactive behavior: a new look at the interface between non-work and work. journal of applied psychology, 88, 518-528. doi: 10.1037/00219010.88.3.518 staiger, d. o., & rockoff, j. e. (2010). searching for effective teachers with imperfect information. journal of economic perspectives, 24, 97-118. doi:10.1257/jep.24.3.97 tabachnick, b. & fidell, l. s. (2007). using multivariate statistics (5th ed.). boston: pearson. thomas, c. h. (2006). clarifying the concept of work engagement: construct validation and an empirical test. (doctoral dissertation, the university of georgia). retrieved from http://hdl.handle.net/10724/9113 tschannen-moran, m. & woolfolk hoy, a. (2001). teacher efficacy: capturing an elusive construct. teaching and teacher education, 17, 783-805. retrieved from http://www.sciencedirect.com/science/article/pii/s0742051x01000361# velicer, w. f., & jackson, d. n. (1990). component analysis versus common factor analysis: some further observations. multivariate behavioral research, 25(1), 97-114. doi: 10.1207/s15327906mbr2501_12 wang, m.-t. (2009). school climate support for behavioral and psychological adjustment: testing the mediating effect of social competence. school psychology quarterly, 24, 240–251. doi:10.1037/a0017999 wang, y., & qin, j. (2011). the structure of preschool teachers‘ work engagement survey in china. international conference on social science and humanity, 5, 464-468. retrieved from http://www.ipedr.com/vol5/no2/103-h10248.pdf watt, h. m. g., & richardson, r. w. (2007). motivational factors influencing teaching as a career choice: development and validation of the fit-choice scale. the journal of experimental education, 75, 167-202. doi:10.3200/jexe.75.3.167-202 xanthopoulou, d., bakker, a.b., demerouti, e., & schaufeli, w.b. (2007). the role of personal resources in the job demands-resources model. international journal of stress management, 14, 121-141. doi: 10.1037/1072-5245.14.2.121 zwick, w. r., & velicer, w. f. (1986). comparison of five rules for determining the number of components to retain. psychological bulletin, 99, 432-442. doi: 10.1037/0033-2909.99.3.432 http://psycnet.apa.org/doi/10.1037/0021-9010.88.3.518 http://psycnet.apa.org/doi/10.1037/0021-9010.88.3.518 http://www.sciencedirect.com/science/article/pii/s0742051x01000361 http://www.ipedr.com/vol5/no2/103-h10248.pdf http://psycnet.apa.org/doi/10.1037/0033-2909.99.3.432 microsoft word prast et al_publication.docx ! ! ! ! ! frontline learning research vol.3 no. 2 (2015) 90-116 issn 2295-3159 readiness-based differentiation in primary school mathematics: expert recommendations and teacher self-assessment emilie j. prasta, eva van de weijer-bergsmaa, evelyn h. kroesbergena, & johannes e.h. van luita a utrecht university, the netherlands article received 8 april 2015 / revised 18 june 2015 / accepted 3 august 2015 / available online 21 august 2015 abstract the diversity of students’ achievement levels within classrooms has made it essential for teachers to adapt their lessons to the varying educational needs of their students (‘differentiation’). however, the term differentiation has been interpreted in diverse ways and there is a need to specify what effective differentiation entails. previous reports of low to moderate application of differentiation underscore the importance of practical guidelines for implementing differentiation. in two studies, we investigated how teachers should differentiate according to experts, as well as the degree to which teachers already apply the recommended strategies. study 1 employed the delphi technique and focus group discussions to achieve consensus among eleven mathematics experts regarding a feasible model for differentiation in primary mathematics. the experts agreed on a fivestep cycle of differentiation: (1) identification of educational needs, (2) differentiated goals, (3) differentiated instruction, (4) differentiated practice, and (5) evaluation of progress and process. for each step, strategies were specified. in study 2, the differentiation self-assessment questionnaire (dsaq) was developed to investigate how teachers self-assess their use of the strategies recommended by the experts. while teachers (n = 268) were moderately positive about their application of the strategies overall, we also identified areas of relatively low usage (including differentiation for high-achieving students) which require attention in teacher professional development. together, these two studies provide a model and strategies for differentiation in primary mathematics based on expert consensus, the dsaq which can be employed in future studies, and insights into teachers’ self-assessed application of specific aspects of differentiation. keywords: differentiation; mathematics; primary school; teacher self-assessment; delphi method corresponding author: emilie prast, faculty of social and behavioural sciences, department of pedagogical and educational sciences, heidelberglaan 1, 3584 cs utrecht, the netherlands, e.prast@uu.nl doi: http://dx.doi.org/10.14786/flr.v3i2.163 prast&et &al & & & | f l r ! ! 91! 1. introduction every day, primary school teachers are faced with the task of teaching students of diverse academic ability and achievement levels. therefore, teachers should adapt their lessons to the diverse educational needs of their students (corno, 2008). such adaptations are often promoted using the term differentiation or differentiated instruction, defined by tomlinson et al. (2003, p.120) as “an approach to teaching in which teachers proactively modify curricula, teaching methods, resources, learning activities, and student products to address the diverse needs of individual students and small groups of students to maximise the learning opportunity for each student in a classroom”. the international trend towards inclusive education makes the need for differentiation especially urgent. within response to intervention models, general education teachers are required to provide both universal support i.e., a good general education for all students (tier 1) and targeted support (tier 2) such as small-group instruction for struggling students (fuchs & fuchs, 2007; mcleskey & waldron, 2011). small-group or individual interventions carried out by an educational specialist (tier 3) are only available for a limited number of students whose problems persist despite the supports provided by the general education teacher. thus, general education teachers have the primary responsibility for providing a good education to all students, regardless of their achievement level. attending to the educational needs of students with a broad range of ability and achievement levels is a challenge for teachers. successful differentiation requires advanced subject matter knowledge, pedagogical skills and classroom management skills (vantassel-baska & stambaugh, 2005). consequently, a need for professional development in the area of differentiation has been identified repeatedly (johnsen, haensly, ryser, & ford, 2002; van den broek-d’obrenan et al., 2012; vantassel-baska et al., 2008). to design effective professional development programmes, it is important to know what teachers should do in their day-to-day teaching to differentiate their lessons for students of diverse achievement levels. what constitutes best practice? in two studies, we investigated how teachers should differentiate according to experts as well as the degree to which teachers already apply the recommended strategies. the focus was exclusively on mathematics since strategies for differentiation may vary across subject areas. moreover, domain-specific guidelines or strategies tend to be more concrete and may therefore provide stronger guidance to teachers. differentiation is an umbrella term that may be used to refer to one or several of a variety of instructional modifications. it may involve modifications of the content (what students learn), the process (how they learn it), or the product of learning (how students demonstrate their learning) (tomlinson, 2005). various student characteristics may serve as a ground for differentiation. for example, tomlinson et al. (2003) distinguish between differentiation by student readiness (representing the current level of knowledge and skills in the subject area), learning profile (a student’s preferred ways of learning, such as a preference for visual input) and interest (topics about which the student wants to learn more). in the current study, the focus is on differentiation by student readiness. readiness is influenced by a child’s natural ability as well as learning experiences and is reflected in the child’s current knowledge and skill level. the importance of differentiation by student readiness is supported by the theoretical constructs of the zone of proximal development (vygotsky, 1978), challenge-skill balance (csikszentmihalyi, 1990), aptitude-treatment interaction (cronbach & snow, 1977), and adaptive teaching (corno, 2008). vygotsky (1978) stated that learning occurs when a child engages in activities that fall within its zone of proximal development (zpd), i.e. that are slightly more difficult than what the child already masters independently. when children within one classroom have widely varying readiness levels, their zones of proximal development also differ. a task that is just within reach for average-achieving students (i.e. in their zpd) may be too difficult for children with lower readiness levels when the gap between existing knowledge and skills and the task is too big. conversely, children with higher readiness levels may already master the task and in this case they are not challenged to reach beyond what they can already do. this implies that children within the same classroom may need different instructional treatments to work in their zpd. to work in the prast&et &al & & & | f l r ! ! 92! zpd, the skill level of the students should be in balance with the difficulty level of the tasks. such a challenge-skill balance may result in effective and engaged learning, while tasks that are much too difficult or too easy may lead to frustration, boredom, and withdrawal from learning (csikszentmihalyi, 1990).! additionally, certain characteristics of the learning environment may be useful for some learners but not for others, depending on the aptitude of the student (cronbach & snow, 1977). because of the variation in student aptitudes and the resulting diversity of educational needs, teachers should adapt education to the needs of their students (corno, 2008). what these theories have in common is the idea that students with different readiness levels have different educational needs and that instruction should be matched to these needs, which is exactly what differentiation aims to do. roy, guay, and valois (2013) took a step towards clarification of the term differentiation by identifying two main components of readiness-based differentiation: academic progress monitoring and instructional adaptations. ideally, the developments in students’ achievement or understanding are closely followed, for example using frequent formal or informal tests, and adaptations are then made to ensure a good fit between the readiness of the student and the instruction. most approaches to differentiation include these two components in some way. nevertheless, the way in which progress is monitored and the nature of instructional adaptations strongly vary across intervention studies (e.g. brown & morris, 2005; mcdonald connor et al., 2009; reis, mccoach, little, muller, & kaniskan, 2011; tieso, 2005; ysseldyke & tardrew, 2007). students’ achievement may be measured with standardised, curriculum-based, or informal assessments. in some cases, the results of these assessments are used to determine the instructional treatment for an extended period of time (weeks or months) whereas other interventions continuously monitor progress and adapt the instructional treatment accordingly. adaptations may be at the level of individual students or subgroups of students. when grouping is used, such groups may be between-class or within-class, fixed or flexible (tieso, 2003). adaptations may entail modification of the amount of instruction, the content or type of instruction, the content or type of independent practice tasks, or combinations of these elements. given the diverse interpretations of the term differentiation, there is a need to specify what effective differentiation entails. one line of research has examined the effects of various types of ability grouping. the best results are obtained when students can switch between groups based on changes in their educational needs (the progress monitoring component of differentiation) and when instruction is tailored to the needs of the students in the groups (the instructional adaptations component) (kulik & kulik, 1992; lou et al., 1996; slavin, 1987; tieso, 2003). when these conditions are met, homogeneous within-class ability grouping has demonstrated positive effects on student achievement across multiple studies (kulik & kulik, 1992; lou et al., 1996; slavin, 1987; tieso, 2005). in contrast, slight negative effects of within-class ability grouping in primary school were found across three studies in which variations in instructional treatment were not explicitly described (deunk, doolaard, smale-jacobse, & bosker, 2015). so, it seems to be important to use the grouping arrangement as a means to provide the different subgroups with the instruction that they specifically need, i.e. to differentiate instruction. another issue in the literature on ability grouping is the potential existence of differential effects, i.e. different effects for students of different ability levels. while slavin (1987) reported a higher median effect size for low-ability students than for average-ability and highability students, other reviews have found different patterns with smaller or even negative effects for lowachieving students (deunk et al., 2015; kulik & kulik, 1992; lou et al., 1996). more research is necessary to determine in which situations such differential effects may arise. a recent review (deunk et al., 2015) examined the effects of various readiness-based differentiation practices on student achievement. although the authors aimed to include all high-quality studies published about this topic since 1995, only sixteen studies about differentiation in primary school could be included. most of these sixteen studies were either too narrow (ability grouping without explicit instructional differentiation) or too broad (interventions in which differentiation was one of several components of a comprehensive school reform initiative) to be informative about the effects of differentiation on student achievement. however, promising results were obtained with two computerised interventions for prast&et &al & & & | f l r ! ! 93! differentiation: individualizing student instruction (mcdonald connor, morrison, fishman, schatschneider, & underwood, 2007; mcdonald connor et al., 2011a; mcdonald connor et al., 2011b) and accelerated math (ysseldyke et al., 2003; ysseldyke & bolt, 2007). the individualizing student instruction programme provides the teacher with recommendations about the amount and type of literacy instruction needed by individual students based on their scores on a computerised test. accelerated math is a technological application which continuously monitors students’ progress, adapts practice tasks to students’ individual skill level, and informs the teacher when students struggle with certain types of problems. both of these interventions, which clearly contain both components of differentiation (progress monitoring and instructional adaptations), have demonstrated significant positive effects across multiple studies. prior research has shown that there is room for improvement in teachers’ implementation of differentiation. the dutch inspectorate of education recently found that adequate adaptations to diverse educational needs are only made at about half of the schools (van den broek-d’obrenan et al., 2012). in us middle schools, both teachers and students reported low usage of differentiation strategies (moon, callahan, tomlinson, & miller, 2002). in a recent study on canadian elementary schools, teachers self-reported moderate use of differentiation strategies, but strategies requiring more time to implement were used relatively infrequently (roy et al., 2013). similarly, studies about adaptations for students with learning disabilities found that teachers tend to implement typical adaptations which can be easily implemented for all students rather than specialised adaptations, i.e. adaptations targeted at the unique educational needs of individual students (mcleskey & waldron, 2002, 2011; scott, vitale, & masten, 1998). however, a recent study carried out in finland found that teachers do provide more individual support to struggling students (nurmi et al., 2013). for high-achieving or gifted students, low levels of differentiation have generally been found (reis et al., 2004; westberg, archambault, dobyns, & salvin, 1993; westberg & daoust, 2003). in sum, prior studies have generally found low to moderate use of differentiation strategies, although the degree of implementation of differentiation seems to vary depending on the specific strategies for differentiation examined, the targeted population of students, and perhaps also the country in which data are collected. specialised adaptations as well as adaptations targeted specifically at high-achieving students seem to be used relatively infrequently. in conclusion, there is a clear need to apply differentiation based on differences in students’ readiness and teachers could use some help in doing this. the literature shows that differentiation should include progress monitoring and instructional adaptations. however, the ways in which this can be done effectively are less clear. promising results have been obtained with two computerised interventions. however, high-quality research about the achievement effects of interventions in which differentiation is mainly implemented by the teacher himself is scarce. there is a need for general guidelines for differentiation that can be applied in a wide array of schools, independently from curricular methods or technological applications. therefore, study 1 sought to achieve consensus among a consortium of mathematics experts about a feasible model and associated strategies for differentiation. study 2 linked the results of study 1 to teachers’ daily practice by examining how teachers self-assess their use of the strategies for differentiation recommended by the experts. 2. study 1 2.1 aims study 1 the aim of study 1 was to operationalise the concept of differentiation by achieving consensus among a consortium of mathematics experts about a coherent set of strategies for differentiating primary school mathematics education. the result of the consensus procedure needed to be feasible for use by general education teachers in daily mathematics teaching. additionally, it needed to be applicable in diverse schools, independent from curricular method. prast&et &al & & & | f l r ! ! 94! expert consensus procedures can be valuable when scientific literature provides insufficient information to make complex decisions (landeta, 2006) and have been applied before to achieve consensus about effective teaching (teddlie, creemers, kyriakides, muijs, & yu, 2006). while several individual experts have made recommendations for differentiation in primary mathematics in books and journals for practitioners, consensus among various experts could provide a more solid foundation. for differentiation in mathematics, teacher educators with expertise in the didactics of primary mathematics are the relevant group of experts. teacher educators may have gained practical knowledge regarding the effectivity and feasibility of diverse strategies for differentiation. making use of this experiential knowledge has the potential to complement the scientific literature and strengthen the link between theory and practice. 2.2 method study 1 2.2.1 participants the consortium of experts was designed to include distinguished dutch pre-service and in-service teacher educators with a professional focus on mathematics education. to be eligible for participation, potential members had to be experts in their field, as demonstrated by their (1) experience in providing preservice or in-service teacher training about teaching mathematics (2) regular presence as invited speaker at educational conferences and (3) role as a consultant to the ministry of education, culture and science to discuss new educational policy. the senior authors approached potential candidates with these criteria in mind. all experts who were invited to participate agreed to join the consortium. this resulted in a consortium of eleven experts (seven men, four women) representing eight large national and regional institutes for preand in-service teacher training spread across the netherlands. the members had experience in at least two of the following areas: in-service teacher training for mathematics (m = 8.6 years, sd = 8.5 years), pre-service teacher training for mathematics (m = 5.4 years, sd = 6.3 years), carrying out educational evaluation studies (m = 25.0 years, sd = 21.2 years) and teaching (m = 5.7 years, sd = 5.4 years). the current daily work of the consortium members included educating pre-service teachers, providing professional development for in-service teachers, and guiding schools in the implementation of new educational approaches including differentiation. 2.2.2 consensus procedure focus group discussions (liamputtong, 2011) and the delphi method (hasson, keeney, & mckenna, 2000) were used to investigate the experiential knowledge of the experts on differentiated mathematics education systematically. focus group discussions are structured discussions with a group of persons involved in the topic in which certain roles (e.g. a discussion leader, a timekeeper and a secretary) and rules (e.g. only on-topic contributions) are specified and adhered to. the delphi technique entails the repeated administration of a questionnaire in order to achieve consensus among experts. after the first round of administration, the initial responses are presented anonymously to the participants, who are then asked to fill out the questionnaire again. this procedure is repeated until consensus (specified with a consensus criterion) is reached. the order of focus group discussions and delphi rounds in the current study is presented in figure 1. the whole procedure took place between november 2011 and january 2012. in the first three-hour focus group discussion, the experts were invited to share their knowledge, prompted by eight core questions about differentiation (see figure 1). these questions were deliberately left open to elicit broad input. no particular theoretical perspective was chosen a priori apart from the assumption that student readiness would be an important ground for differentiation (see questions 6 and 7). rather, the questions were asked from a practical point of view (what does and does not work in practice and how can this be improved). in principle, the questions were discussed one by one in the listed order, but in practice, the discussion sometimes moved back and forth between the various questions because of their high prast&et &al & & & | f l r ! ! 95! interrelatedness. after the first focus group discussion, the first author restructured the meeting minutes in terms of (initial) answers to the eight core questions. second, based on this input, the researchers constructed an online delphi questionnaire (round 1). during the first focus group discussion, one of the experts had listed five general themes that are central to differentiation: organisation, goals, instruction, practice, and learning styles. these themes were used to structure the delphi questionnaire. the theme ‘differentiation in kindergarten’ was added to account for aspects of differentiation specific to kindergarten (kindergarten is integrated in the dutch primary school system). for each theme, the first author summarised the main ideas of the focus group discussion and proposed this to the other authors. apart from some minor changes, the other authors agreed that these summaries accurately reflected what had been said in the discussion. these summaries were included in the delphi questionnaire as one-paragraph concepts for differentiation (see appendix 1). for each theme, statements about specific elements of the concept were also included (for example: ‘the low-achieving subgroup profits from extended instruction’). the experts rated their agreement with the concepts and with the specific statements on a likert scale ranging from 1 (do not agree at all) to 5 (fully agree). additionally, open questions prompted participants to provide any comments they had. third, a round 2 delphi questionnaire was developed which included only those questions on which no overwhelming consensus had been reached in round 1. the consensus criterion for round 1 was that all responses should be at one end of the scale (i.e. either 4 and 5 or 1 and 2), with a maximum of one neutral response (3). the questions on which consensus had not yet been reached were presented to the participants again accompanied by a bar chart of the responses in round 1. additionally, comments provided by the participants in round 1 were included as new questions in round 2. using a more lenient consensus criterion of maximally three neutral responses and the rest at one end of the scale, the researchers determined for which items consensus was achieved in round 2. fourth, the researchers presented the results of the delphi questionnaire to the experts during the second focus group discussion which lasted two hours. the items on which consensus had not been achieved were discussed to clarify misunderstandings (especially about the open comments provided by the participants in round 1) and resolve conflicting opinions. fifth, the first author reviewed the meeting minutes of the focus group discussions and the responses to the delphi questionnaire to synthesise the input, resulting in a proposed model for implementing differentiation. the other authors agreed with this model. sixth, the proposed model for differentiation was sent to all consortium members and discussed during a third one-hour meeting of the focus group. 2.2.3 attendance rates of consortium members of the eleven consultants, six (54.5%) attended the first two focus group discussions and completed the two delphi questionnaires and four (36.4%) completed three out of four components (i.e. either both discussions and one questionnaire or both questionnaires and one discussion). one participant only completed the delphi questionnaire, after being informed about the content of the first focus group discussion in a separate meeting with the researchers. all members received the proposed model for differentiation by email and were given the opportunity to send any comments or questions, and six participants (54.5%) attended the third meeting in which the cycle was discussed. prast&et &al & & & | f l r ! ! 96! figure 1. consensus procedure. focus group discussion 1 delphi round 2 delphi round 1 focus group discussion 2 focus group presentation of results!! determine consensus synthesise results questions asked in focus group discussion 1: 1. what should differentiation in primary maths education look like? 2. what could be improved compared to the current situation on primary schools regarding differentiation in maths? 3. what does a teacher who successfully differentiates education do? 4. what are challenges in implementing differentiation in maths education? 5. which types of instruction could be distinguished and for which pupils would these types of instruction be appropriate? 6. are there other aspects than readiness level based on which a teacher could differentiate (for example: visual or verbal instruction)? 7. what should a teacher do differently with students of different readiness levels? 8. how often should additional or specific types of instruction be provided? in focus group discussion 2, issues on which consensus had not yet been reached were discussed and resolved. content of the delphi questionnaires themes and consensus criteria are reported in table 1 round 1 questionnaire: • summarising concepts (rated on five-point scale) • specific statements (rated on five-point scale) • comments (open answer) round 2 questionnaire: • round 1 concepts and statements on which consensus had not been reached in round 1 • comments provided by the participants in round 1 included as new specific statements determine consensus create delphi questionnaire prast&et &al & & & | f l r ! ! 97! 2.3 results study 1 in round 1 of the delphi questionnaire, the experts agreed with the concepts for differentiation in instruction, differentiated goals, differentiated practice, differentiation based on learning styles, and differentiation in kindergarten which had been formulated based on the first focus group discussion. positive consensus on the remaining concept (organisation of differentiation) was reached in round 2. table 1 provides an overview of the degree of consensus achieved on the specific statements in the two rounds of the delphi questionnaire. consensus was reached on 35 items in the first round and on an additional 25 original items in the second round, amounting to consensus on 74.1% of original items after two rounds. regarding the new statements that were derived from the open comments in round 1, consensus was reached on 46.0% of these statements in round 2. the items on which consensus had not been reached after the second delphi questionnaire were discussed in the second focus group discussion. differences in interpretation of certain items were resolved and consensus was reached about the main issues. items on which no consensus had been reached in the delphi questionnaires often concerned issues about which the experts were unsure or had no pronounced opinion, including the importance of specific elements (e.g. videotaped instruction, mind maps, games, student choice) and preference for certain grouping formats (e.g. pairs or small groups). in the second focus group discussion, the overall conclusion about these elements and formats was that they all have their merits and that the choice is dependent on the situation, but that they are not crucial for differentiation. table 1 overview of consensus on statements in the delphi questionnaire original statements new statements in round 2 theme subtheme no. of statements consensus after round 1 consensus after round 2a no. of statements consensus after round 2 organisation of differentiation 11 4 7 14 2 differentiation in instruction general 5 3 4 14 4 whole-class instruction 9 2 6 10 6 subgroup instruction for low-achieving students 10 3 8 7 3 subgroup instruction for high-achieving students 12 6 8 4 3 differentiated goals 10 4 8 23 10 differentiated practice 14 5 9 6 3 differentiation based on learning styles 3 1 3 25 15 differentiation in kindergarten 7 7 7 10 6 total 81 35 60 113 52 a total amount of items on which consensus was reached, including items on which consensus had already been reached in round 1 the experts approved the model for differentiation which was created based on their input. the model, dubbed the cycle of differentiation, consists of the following five steps: identification of educational needs, differentiated goals, differentiated instruction, differentiated practice, and evaluation of progress and process (see figure 2). a distinction is made between instruction and practice. instruction refers to moments during which the teacher provides instruction to the whole class, subgroups of students, or individual students, whereas practice refers to moments during which students work on tasks, individually or in groups. prast&et &al & & & | f l r ! ! 98! these two can happen simultaneously, for example when the teacher provides instruction to a subgroup while other students are working on practice tasks. in the following paragraphs, we describe the key recommendations for each step in the cycle of differentiation provided by the experts in the consensus procedure. figure 2. cycle of differentiation. organisation is placed centrally in the cycle, because successful implementation of differentiation depends on a facilitative organisational structure and good classroom management. a key organisational characteristic of the model for differentiation agreed upon by the experts is the assignment of students to subgroups based on achievement level, allowing teachers to make instructional adaptations for subgroups of students with similar educational needs. only the remaining individual educational needs that are not met within the subgroup call for individual accommodations. the first step in the cycle of differentiation is the identification of educational needs. initially, the teacher should assign students to subgroups (typically a low-achieving, an average-achieving and a highachieving subgroup) based on their results on standardised tests and curriculum-based tests. in the course of the schoolyear, teachers continuously gather new and more detailed information about students’ educational needs, for example with the analysis of daily work, informal observations and diagnostic conversations. the subgroups should be flexible, i.e. students should be able to switch groups based on changes in their educational needs. based on the educational needs of the students, differentiated goals should be set. overarching objectives (the material that students should master at the end of primary school) and lesson goals (goals for a specific lesson) are distinguished. overarching objectives should not only be formulated for averageachieving students but also specifically for low-achieving and high-achieving students (see also appendix 1, differentiation in goals). the overarching objectives should be translated into concrete lesson goals, which are provided mostly by the curriculum. however, only some of the mathematics curricula available in the netherlands differentiate lesson goals for three achievement levels. when the curriculum does not differentiate goals sufficiently, the teacher should formulate challenging but realistic lesson goals for all subgroups. identification of educational needs differentiated goals differentiated instruction differentiated practice evaluation of progress and process organisation prast&et &al & & & | f l r ! ! 99! based on the educational needs and the goals that have been set, the teacher differentiates instruction through broad whole-class instruction, subgroup instruction tailored to the needs of that subgroup, and individual adaptations. during whole-class instruction, the teacher should serve a broad range of educational needs by varying the difficulty level of questions, stimulating all children to think about the answer to a question by giving thinking time, teaching at various levels of abstraction, and using several input modalities (e.g. visual, verbal, tactile). in subgroup instruction, the teacher should adapt instruction to the educational needs of low-achieving and high-achieving students. it is assumed that, in general, low-achieving students need more guidance (e.g. explicit instruction) and instruction at lower levels of abstraction (e.g. using blocks to represent and calculate a sum) while high-achieving students need more exploratory instruction about advanced content with a focus on conceptual understanding (e.g. the relation between multiplication and division). to the extent possible, the teacher should also take into account individual differences during subgroup instruction and while giving individual feedback. in the practice phase of the lesson, the subgroups need quantitatively and qualitatively different tasks. for the low-achieving subgroup, completing all regular tasks is often not realistic, so the tasks that are crucial for mastery of the objectives for low-achieving students should be selected (the remaining tasks can still be completed when students have time left). for the high-achieving subgroup, the regular material should be compacted. for the most popular dutch mathematics curricula, guidelines exist to inform the teacher which tasks can be skipped by high-achieving students. in other cases, the teacher should remove most of the repetitive tasks and select the tasks that are crucial steps towards mastery of the objectives. the time freed up by compacting should be spent on enrichment. some curricula provide enrichment tasks and these tasks can be used in some cases, but they are often not sufficiently challenging for very high-achieving students. therefore, supplemental enrichment curricula should also be used. technological applications such as mathematics websites and instructional computer programmes can also be valuable tools for individual differentiation, provided that they are used deliberately for additional practice in areas the student does not master yet or for enrichment at an appropriate challenge level. the final step in the cycle of differentiation is the evaluation of progress and learning process. based on daily work and achievement tests, the teacher should evaluate whether the students have met the lesson goals. regarding process, the teacher should evaluate whether the applied adaptations of instruction and practice had the desired effect. for example, when a teacher has intentionally taught the low-achieving subgroup at a lower level of abstraction, the teacher should evaluate whether this was helpful for these students. to gauge the effectiveness of the accommodations made, the teacher may supplement achievement results with informal measures such as observations or diagnostic conversations. the evaluation phase informs the teacher about students’ current achievement level and about instructional approaches that work for these students, completing the cycle and serving as new input for the identification of educational needs. 2.4 concluding summary study 1 the aim of study 1 was to operationalise the concept of differentiation by achieving consensus among a consortium of mathematics experts about a coherent set of strategies for differentiating primary school mathematics education. a combination of focus group discussions and the delphi technique was used to investigate the experiential knowledge of eleven experts in mathematics education systematically. consensus was reached on all summarising concepts and on the majority of specific statements from the delphi questionnaire. the input from the experts was synthesised into a cycle of differentiation consisting of the following five steps: identification of educational needs, differentiated goals, differentiated instruction, differentiated practice, and evaluation of learning progress and process. for each step, strategies were specified, providing teachers with concrete guidelines for implementing differentiation in primary school mathematics. prast&et &al & & & | f l r ! ! 100! 3. study 2 3.1 aims study 2 study 2 linked the results of study 1 to teachers’ daily practice by investigating teachers’ selfassessed use of the strategies for differentiation recommended by the experts. therefore, we developed the differentiation self-assessment questionnaire (dsaq) which covers the recommended strategies in five subscales corresponding with the five steps of the cycle of differentiation (see also section 3.2.2). since the dsaq was newly developed, we also aimed to investigate its statistical properties, including its factor structure and relation with other scales. the development of a new instrument was necessary to ensure coverage of the broad set of strategies recommended by the experts in study 1. another recently developed instrument to measure teachers’ selfreported use of differentiation is the differentiated instruction scale (dis; roy et al., 2013). this instrument was based on a similar theoretical framework and there is overlap between the content of the items of the dis and the dsaq. however, the dis has only twelve items and is not sufficiently specific to measure all strategies recommended by the experts. regarding progress monitoring, for example, the dis includes the rather general item ‘analyse data about students’ academic progress’ while the dsaq distinguishes between the different types of progress monitoring recommended by the experts, ranging from standardised tests to diagnostic interviews. thus, the added value of the dsaq is that it is a detailed measure of the specific strategies recommended by the experts in study 1. the first aim of the current study was to examine the factor structure of the dsaq. the literature reviewed in the introduction indicates that effective differentiation entails two components: progress monitoring and instructional adaptations. the steps in the cycle of differentiation reflect these components: identification of educational needs and evaluation of progress and process involve progress monitoring, while differentiated goals, instruction, and practice involve instructional adaptations. for the dis, roy et al. (2013) found that a model with these two factors provided a better fit to the data than a model in which all items loaded on one general differentiation factor. therefore, we investigated whether the dsaq has a similar factor structure by comparing the fit of a two-factor model with one factor for progress monitoring and one factor for instructional adaptations to the fit of a one-factor model. the second aim was to examine the convergent and divergent validity of the dsaq by investigating its relationship with other teacher self-report scales. teacher self-efficacy is a multidimensional construct which comprises teachers’ perceived ability to perform various aspects of teaching (skaalvik & skaalvik, 2007; tschannen-moran & woolfolk hoy, 2001). theoretically, self-assessed usage of differentiation as measured by the dsaq should be more closely related to aspects of self-efficacy related to differentiation than to other aspects of teacher self-efficacy. specifically, we expected stronger correlations with scales that measure teachers’ self-efficacy for instruction to students of diverse achievement levels, for adapting education to individual students’ needs, and for self-assessed prerequisite knowledge for differentiation (which would support convergent validity) than with scales that measure teachers’ self-efficacy for motivating students, for coping with changes and challenges, and for classroom management (which would support divergent validity). the third and main aim was to investigate teachers’ self-assessed use of the strategies for differentiation recommended by the experts. besides examining teachers’ overall usage, we also aimed to identify strategies which were relatively infrequently used. such information may provide starting points for teacher professional development by indicating in which areas teachers perceive most room for improvement. based on the literature reviewed in the theoretical background as well as on the input from the experts in study 1, we hypothesised that average scores would be low to moderate and that specialised strategies aimed at individual students’ unique educational needs as well as strategies targeted specifically at high-achieving students would be used relatively infrequently. prast&et &al & & & | f l r ! ! 101! 3.2 method study 2 3.2.1 participants and procedure the sample consisted of 268 primary school teachers working at 31 schools participating in a largescale project about differentiation. schools were informed about the project through flyers and advertisements and could register themselves for participation on a project website. the schools were located in rural and urban areas spread across the netherlands and were diverse in terms of school size, student population, religious background, and mathematics curriculum used. all 325 teachers of grade 1 through 6 of the participating schools were invited by email to fill out an online questionnaire containing the dsaq and related scales. the questionnaire was administered at the beginning of the 2012 2013 school year. a total of 268 teachers (83%) completed the questionnaire and gave informed consent. the remaining teachers did not give informed consent (n = 3), completed the questionnaire only partly (n = 7) or did not respond at all (n = 47). on average, participants had 15.6 years of teaching experience (range 0 – 40 years). seventyone teachers (26%) taught a multigrade class. fifty-four teachers (20%) worked full-time, whereas most teachers worked two, three or four days a week (61, 81 and 55 teachers respectively). 3.2.2 instruments the dsaq was developed to examine how teachers assess their use of the strategies recommended by the experts in study 1. each subscale represents one step of the cycle of differentiation and covers core strategies for differentiation belonging to that step. subscales and sample items are provided in table 2. organisational aspects of the model for differentiation were not captured in a separate scale but were partly covered in the subscales corresponding to each step of the cycle. in an earlier pilot study, a pilot version of the dsaq had been completed by 27 teachers recruited at four schools. based on the analysis of the internal consistency of the pilot version, which was acceptable to good, some adaptations were made in the final version of the dsaq. the internal consistency of the final version, obtained in the current sample (n = 268), is reported in section 3.3.1. to assess the convergent and divergent validity of the dsaq, subscales from two well-established multidimensional teacher self-efficacy scales were selected. the norwegian teacher self-efficacy scale (skaalvik & skaalvik, 2007) was developed with special attention for adapting education to individual educational needs. it consists of six subscales with acceptable to good internal consistency which load on six primary factors, which in turn load on a second-order factor for general teacher self-efficacy (skaalvik & skaalvik, 2007). for the current study, the subscales for instruction which emphasises instruction to students of diverse achievement levels and adapting education to individual students’ needs1 were selected to assess the convergent validity of the dsaq while the subscales for motivating students and for coping with changes and challenges were selected to assess the divergent validity. the subscales were translated into dutch and the wording was adapted to make the items domain-specific for mathematics instruction. to further examine the divergent validity, the subscale for classroom management from the wellestablished ohio state teaching efficacy scale (ostes; tschannen-moran & woolfolk hoy, 2001) was administered in dutch translation (goei, bekebrede, & bosma, 2011). the ostes consists of three subscales which load on a first-order factor and on a general teaching efficacy factor. the subscales have demonstrated high internal consistency and can be used independently with in-service teachers (tschannenmoran & woolfolk hoy, 2001). as a third potential support for convergent validity, a self-assessment scale about prerequisite knowledge for differentiation was adapted from an informal scale that had already been used to assess the !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 1 compared to the four-item ntses subscale for adapting instruction to individual students’ needs, the added value of the dsaq is that it provides a more detailed measure of self-assessed use of a range of differentiation strategies. prast&et &al & & & | f l r ! ! 102! level of prior knowledge in professional development programmes (nationaal expertisecentrum leerplanontwikkeling, 2010). teachers self-assess the extent to which they already possess the knowledge necessary for implementing differentiated instruction. 3.2.3 analyses because we wanted to compare the fit of two specific models based on theory and previous findings, we used confirmatory factor analysis (cfa) to investigate the factor structure of the dsaq. we first tested a one-factor model in which all dsaq subscales loaded on one general differentiation factor. second, we tested a two-factor model with one factor for progress monitoring (subscales identification of educational needs and evaluation of progress and process) and one factor for instructional adaptations (subscales differentiated goals, differentiated instruction and differentiated practice). version 7.3 of the mplus statistical package (muthén & muthén, 1998-2012) was used. model fit was evaluated with the chi-square statistic, the comparative fit index (cfi), the tucker-lewis index (tli), the root mean squared error of approximation (rmsea), and the standardised root mean square residual (srmr). values above .95 for the cfi and tli and values below .06 and .08 for the rmsea and srmr, respectively, indicate good model fit (hu & bentler, 1999). the maximum likelihood estimator was used. in the standardised solution, the variance of the factors was fixed to 1 so all factor loadings could be estimated freely. correlational analyses were performed to assess the convergent and divergent validity of the dsaq. we expected moderate to strong positive correlations with prerequisite knowledge for differentiation, adapting education to individual students’ needs, and instruction (for students of all achievement levels). for the latter two scales, we expected that correlations would be lower for the factor progress monitoring than for the factor instructional adaptations, because progress monitoring does not focus on the instructional phase. regarding divergent validity, we hypothesised that dsaq scores would be less strongly related to self-efficacy for motivating students, coping with changes and challenges, and classroom management, although we still expected positive correlations because these dimensions of teacher self-efficacy can be helpful when implementing differentiation. to identify areas of relatively low use, we compared the means of all single items to the mean of their factor. if a mean was more than one standard deviation below the mean of the factor, it was classified as relatively low. prast&et &al & & & | f l r ! ! 103! table 2 sample items and descriptive statistics (n = 268) of the administered scales scale sample item response options no. of items α m sd dsaq identification of educational needsa i analyse the answers on curriculumbased tests to assess a student’s educational needs 1 = does not apply to me at all, 5 = fully applies to me 5 .69 3.64 .55 differentiated goalsa i set extra challenging goals for high-achieving students 1 – 5 as above 6 .79 3.78 .55 differentiated instructiona i adapt the level of abstraction of my instruction to the educational needs of the students 1 – 5 as above 7 .72 3.81 .42 differentiated practicea i select the most important elaboration activities for very lowachieving students 1 – 5 as above 8 .72 3.46 .55 evaluation of progress and processa i use diagnostic conversations to evaluate whether specific students have met the lesson goals 1 – 5 as above 7 .86 3.56 .57 additional scales instructionb how certain are you that you can explain central themes in mathematics so that even the lowachieving students understand? 1 = not certain at all, 4 = absolutely certain 4 .74 3.13 .37 adapting education to individual students’ needsb how certain are you that you can adapt instruction to the needs of low-achieving students while you also attend to the needs of other students in class? 1 4 as above 4 .78 2.91 .44 coping with changes and challengesb how certain are you that you can manage instruction regardless of how it is organised (working with subgroups, multigrade classes with 3 grades, etc.)? 1 4 as above 4 .76 2.99 .43 motivating studentsb how certain are you that you can get students to do their best even when working with difficult problems? 1 4 as above 4 .76 2.95 .42 classroom managementc how much can you do to control disruptive behavior in the classroom? 1 = nothing, 9 = very much 7 .92 7.17 .77 prerequisite knowledge for differentiationd i know the different solution strategies that are used by children 1 = does not apply to me at all, 5 = fully applies to me 10 .84 3.79 .41 a newly deloped dsaq-scales b adapted from skaalvik and skaalvik (2007) c taken from goei, bekebrede, and bosma (2011) d adapted from nationaal expertisecentrum leerplanontwikkeling (2010) prast&et &al & & & | f l r ! ! 104! 3.3 results study 2 3.3.1 properties of the dsaq: internal consistency and factor structure as reported in table 2, the internal consistencies of the dsaq subscales were acceptable to good (streiner, 2003). the results of the confirmatory factor analysis indicate that the one-factor model in which all five subscales loaded on a general differentiation factor did not fit the data well: χ2(5) = 55.126, p < .001; rmsea = .193 (90% ci .149 .241); cfi = .912; tli = .824; srmr = .050. the two-factor model had a good fit: χ2(4) = 5.637, p = .228; rmsea = .039 (90% ci .000 .107); cfi = .997; tli = .993; srmr = .017. for the factor progress monitoring, standardised factor loadings were .84 (se = 0.03, r2 = .71) for identification of educational needs and .85 (se = 0.03, r2 = .73) for evaluation of progress and process. for the factor instructional adaptations, standardised factor loadings were .77 (se = .04, r2 = .59) for differentiated goals, .75 (se = .04, r2 = .56) for differentiated instruction, and .74 (se = .04, r2 = .54) for differentiated practice (p < .001 for all factor loadings). the correlation between the factors was .78 (p < .001). since the two-factor model provided a better fit, the two factor scores (average of the subscale scores comprising that factor) were used in subsequent analyses. 3.3.2 convergent and divergent validity: correlations with other scales the correlations between the two dsaq factors and related scales are reported in table 3. in support of convergent validity, the correlation with prerequisite knowledge for differentiation was strong for both factors. as hypothesised, correlations with self-efficacy for instruction and self-efficacy for adapting education to individual students’ needs were moderate to strong for the factor instructional adaptations and somewhat lower for progress monitoring. regarding divergent validity, the correlations with motivating students and classroom management were less strong, although still in the moderate range. contrary to expectations, coping with changes and challenges correlated strongly with the factor instructional adaptations. table 3 correlations (p < .001) between dsaq factor scores and related scales scale progress monitoring instructional adaptations r 95% ci r 95% ci selected for convergent validity prerequisite knowledge for differentiation .62 .54 .68 .70 .64 .76 instruction (to students of all achievement levels) .40 .30 .49 .47 .37 .56 adapting education to individual students’ needs .38 .28 .48 .56 .48 .64 selected for divergent validity motivating students .30 .19 .40 .37 .27 .47 classroom management .34 .23 .44 .40 .30 .50 coping with changes and challenges .42 .31 .52 .58 .49 .65 3.3.3 distribution of dsaq scores: mean scores and infrequently reported strategies the mean factor scores were 3.60 (sd = 0.52) for progress monitoring and 3.68 (sd = 0.43) for instructional adaptations. with a range from 1.83 to 4.86 for progress monitoring and from 2.45 to 4.94 for instructional adaptations, the factor scores were normally distributed at the high end of the scale. table 2 provides the means and standard deviations of all subscales. taken together, the mean factor and subscale scores reflect moderate to high self-assessed use of differentiation strategies. prast&et &al & & & | f l r ! ! 105! table 4 provides the means and standard deviations for each item of the dsaq. five items numbers 3.7, 4.2, 4.4, 4.8, and 5.5 had a mean score at least one standard deviation below the mean of their factor. two of these the use of diagnostic conversations to evaluate whether the learning goals have been met and the adaptation of type of practice to students’ needs reflect specialised strategies because they involve the refined diagnosis of and adaptation to individual students’ needs. other specialised strategies (items 1.5 and 5.7) also had somewhat lower means, although these means were within one standard deviation of the factor mean. the three remaining infrequently reported items concerned adaptations for high-achieving students, namely additional on-level instruction or guidance, curriculum compacting, and the use of computer programmes for additional challenge. nevertheless, two other strategies targeted at highachieving students (items 2.5 and 4.5) were frequently reported. table 4 means and standard deviations of dsaq items (scale range 1 5) dsaq item m sd subscale 1: identification of educational needs 1.1 i analyse the answers on curriculum-based tests to assess a student’s educational needs 4.02 0.77 1.2 i analyse the answers on standardised tests to assess a student’s educational needs 3.49 0.91 1.3 i assess specific students’ educational needs based on daily maths work 3.75 0.72 1.4 i assess specific students’ educational needs based on (informal) observations during the maths lesson 3.76 0.77 1.5 if necessary, i conduct diagnostic conversations to analyse the educational needs of specific students 3.20 0.90 subscale 2: differentiated goals 2.1 i set different goals for the children, dependent on their achievement level 3.62 0.79 2.2 i set extra challenging goals for high-achieving students 3.57 0.83 2.3 i set well-considered minimum goals for very low-achieving students 3.75 0.76 2.4 i know the opportunities for differentiation offered by the curriculum 4.03 0.68 2.5 i use the opportunities the curriculum offers for differentiation for high-achieving students 3.88 0.84 2.6 i use the opportunities the curriculum offers for differentiation for low-achieving students 3.83 0.82 subscale 3: differentiated instruction 3.1 i adapt the level of abstraction of instruction to the needs of the students 3.95 0.55 3.2 i adapt the modality of instruction (visual, verbal, manipulative) to the needs of the students 3.82 0.62 3.3 i adapt the pace of instruction to the needs of the students 3.95 0.56 3.4 i deliberately ask open-ended questions during whole-class instruction 3.82 0.67 3.5 i deliberately ask questions at various difficulty levels during whole-class instruction 3.69 0.73 3.6 i regularly provide low-achieving children with additional instruction (extended instruction, pre-teaching) 4.25 0.64 3.7 i regularly provide high-achieving students with additional instruction or guidance at their level, in a group or individually 3.20 0.92 prast&et &al & & & | f l r ! ! 106! table 4 (continued) dsaq item m sd subscale 4: differentiated practice 4.1 i vary different types of practice during the maths lesson (e.g. individual or group work, solution spoken, written or drawn) 3.53 0.78 4.2 i adjust different types of practice to the needs of the students in the classroom (e.g. having a specific child complete exercises on the computer because this child learns more in this way) 3.04 0.83 4.3 i select the most important tasks for very low-achieving students 3.73 0.73 4.4 i use curriculum compacting for high-achieving students 3.20 1.25 4.5 i provide high-achieving students with enrichment tasks 4.00 0.87 4.6 i also use computer programmes or maths websites in my maths lessons 3.68 0.97 4.7 i use computer programmes and/or maths websites to offer children focused practice in a skill that they do not sufficiently master 3.32 0.96 4.8 i use computer programmes and/or maths websites to offer specific children additional challenge in the maths lesson 3.15 1.05 subscale 5: evaluation of progress and process 5.1 i use scores on standardised and curriculum-based tests to evaluate whether the learning goals have been met 4.04 0.73 5.2 i analyse the answers on curriculum-based tests to evaluate whether the learning goals of that unit have been met 4.06 0.72 5.3 i regularly evaluate whether all students have met the learning goals based on their daily maths work 3.75 0.85 5.4 i evaluate whether all students have met the lesson goals based on (informal) observations during the maths lesson 3.45 0.86 5.5 i conduct diagnostic conversations to evaluate whether specific students have met the lesson goals 2.85 0.87 5.6 i evaluate whether the type of instruction and practice chosen by me were effective for the majority of the students in the class 3.44 0.77 5.7 i evaluate whether a specific type of instruction was effective for specific students 3.32 0.80 3.4 concluding summary study 2 study 2 investigated teachers’ self-assessed implementation of differentiation using the dsaq. the first goal was to examine the psychometric properties of the dsaq. the subscales of the dsaq were internally consistent and loaded on two correlated but distinct factors: progress monitoring and instructional adaptations. confirmatory factor analysis demonstrated that this two-factor structure provided a better fit than a one-factor model, which converges with the findings reported by roy et al. (2013). the second goal was to examine the convergent and divergent validity of the dsaq. the pattern of correlations between the dsaq and other scales supported its convergent and divergent validity. as expected, strong to moderate correlations with prerequisite knowledge for differentiation, adapting education to individual students’ needs, and instruction were found. as hypothesised, the correlations with the scales selected for testing the divergent validity were lower, except for the correlation with coping with changes and challenges which was unexpectedly strong. the third and main goal was to examine teachers’ perceived usage of the strategies recommended by the experts. with factor means in the moderate to high range, teachers assessed their use of differentiation strategies more highly than we had expected. five items with relatively low means were identified. in prast&et &al & & & | f l r ! ! 107! support of our hypothesis, these items concerned specialised strategies and strategies targeted at highachieving students. 4. general discussion teachers are required to implement differentiation for students of diverse achievement levels. however, the term differentiation had been used in diverse ways and the literature did not provide sufficient information regarding the most effective strategies to provide teachers with general guidelines for implementing differentiation. to fill this gap, study 1 operationalised the concept of differentiation by achieving consensus among a consortium of experts about a model and strategies for differentiation in primary school mathematics. study 2 investigated the degree to which dutch teachers already implement the strategies suggested by the experts. study 1 resulted in a model for differentiation consisting of five steps: identification of educational needs, differentiated goals, differentiated instruction, differentiated practice, and evaluation of progress and process. these steps reflect the two core components of differentiated instruction identified by roy et al. (2013). progress monitoring is captured by the steps of identification of educational needs and evaluation of progress and process. the component of instructional adaptations is represented by the steps of differentiated goals, instruction, and practice. study 2 demonstrated that a two-factor model in which the subscales of the dsaq load on these two factors provides a better fit than a one-factor model. our findings converge with the findings reported by roy et al. (2013), supporting the idea that progress monitoring and instructional adaptations are two distinct but related components of differentiation. new in this study is expert consensus on how progress should be monitored and how goals, instruction and practice should be adapted to the learning needs of students with diverse achievement levels. regarding progress monitoring, the experts recommended to use standardised and curriculum-based tests first to divide students over achievement groups. more refined and informal measures such as the analysis of daily work should be used frequently to monitor short-term progress, to diagnose unique educational needs, and to determine whether a (temporary) adjustment of the groups is necessary. compared to technological applications which tend to make use of one or two types of assessment to monitor progress (e.g. mcdonald connor et al., 2009; ysseldyke & tardrew, 2007), the experts recommended a broader range of strategies and indicated how they can be used together. the strategies have complementary purposes: while relatively formal and standardised tests are useful to get an overview of what a student can do, more informal and qualitative measures such as diagostic conversations and the analysis of daily work provide valuable information about why a student struggles with a certain problem and what the student needs. the use of within-class homogeneous achievement groups provides the opportunity to tailor subgroup instruction to similar educational needs and has demonstrated positive effects (kulik & kulik, 1992; lou et al., 1996; slavin, 1987; tieso, 2005). in line with slavin (1987), the experts stressed the importance of flexibility, i.e. allowing students to switch between groups based on changes in their educational needs. the literature indicates that the effects of within-class ability grouping may depend upon student achievement level, with smaller or even negative effects for low-achieving students (deunk et al., 2015; lou et al., 1996). nevertheless, the experts clearly perceived small-group instruction as a good way to provide low-achieving students with the instruction they specifically need. also, students are only grouped for part of the lesson and participate in the whole-class instruction for students of all ability levels as well. future research should establish whether these conditions ensure that low-achieving students also profit from this type of within-class flexible ability grouping. regarding instructional adaptations, the experts recommended a coherent set of strategies to differentiate goals, instruction and practice. this comprehensive approach is somewhat broader than technology-based interventions which have tended to focus on differentiation of either instruction (individualizing student instruction) or practice (accelerated math). many of the strategies recommended prast&et &al & & & | f l r ! ! 108! by the experts are supported by previous research, including the adaptation of practice tasks to the skill level of the student (ysseldyke & tardrew, 2007), the use of explicit instruction and visual representations for low-achieving students (gersten et al., 2009) and the use of compacting, enrichment, and instruction at challenge level for advanced students (rogers, 2007). to use teachers’ time efficiently, the experts recommended to teach the whole class when possible, to use subgroups when the diverse educational needs of subgroups require this, and to serve remaining unique educational needs individually. thus, the experts recommended both universal supports (supports for all students such as varying the difficulty level of questions in broad whole-class instruction) and targeted supports (supports specifically for low-achieving and high-achieving students including small-group instruction and differentiation in practice tasks). the experts also recommended some adaptations to individual students’ educational needs (e.g. the adaptation of type of practice to the preference of specific students), but they realised that such specialised adaptations were advanced and primarily suitable for teachers who already master basic strategies for differentiation. to link the advice provided by the experts to teachers’ daily practice, study 2 investigated teachers’ self-reported usage of the recommended strategies. overall, dsaq scores were moderate to high, exceeding the expectations we had based on previous studies. perhaps, the different context (primary schools in the netherlands versus middle schools in the united states) can explain the discrepancy with the low use of differentiation strategies reported by moon et al. (2002). our findings are more similar to those of a recent study with canadian primary school teachers in which moderate usage was reported (roy et al., 2013). nevertheless, the moderately high self-assessments in the current study seem discrepant with the finding of the dutch inspectorate of education that adequate adaptations to students’ diverse educational needs are only made at about half of the schools (van den broek-d’obrenan et al., 2012). also, the experts in study 1 clearly perceived a need for professional development about differentiation. perhaps, the inspectors of education and the experts from our consortium have high standards for the quality of implementation which are not captured by the dsaq. teachers might also overestimate their implementation. refined observational studies are necessary to examine whether teachers’ high self-assessed usage of differentiation strategies can be confirmed by external observers. in line with previous studies (mcleskey & waldron, 2002, 2011; reis et al., 2004; scott et al., 1998; westberg et al., 1993; westberg & daoust, 2003), specialised studies and strategies targeted at highachieving students were used relatively infrequently. two specialised strategies the use of diagnostic conversations to evaluate whether the learning goals have been met and adaptation of the type of practice to specific students’ needs were relatively infrequently reported. this corresponds with the view expressed by the experts that individual-level differentiation is advanced and primarily suitable for teachers who already implement group-based strategies for differentiation successfully (van groenestijn, borghouts, & janssen, 2011). three strategies targeted at high-achieving students curriculum compacting, the use of computer programmes for additional challenge, and targeted instruction for these students were used infrequently. the difference between the use of instruction targeted at high-achieving students versus low-achieving students is especially striking. perhaps, teachers are not aware that high-achieving students also need guidance when working on sufficiently challenging enrichment tasks (vantassel-baska & stambaugh, 2005). many teachers do implement some differentiation in practice tasks. however, there is still a lot of room for improvement, since it seems that only few teachers use a complete approach including challenging goals, curriculum compacting, enrichment tasks and on-level guidance. low usage of differentiation for high-achieving students has repeatedly been attributed to a lack of the specific attitudes, knowledge, and skills this requires (latz, speirs neumeister, adams, & pierce, 2009; megay-nespoli, 2001; vantasselbaska & stambaugh, 2005). many teacher educators feel that initial teacher training does not adequately prepare teachers to differentiate instruction for high-achieving students (schram, van der meer, & van os, 2013). thus, it seems important that this topic receives sufficient attention in teacher training and professional development programmes. prast&et &al & & & | f l r ! ! 109! the following limitations should be taken into account. first, the results of consensus procedures are inherently restricted by the participating experts. the risk that other experts might have provided different input cannot be eliminated but was diminished in this study by recruiting experts working for several different institutions for both pre-service and in-service teacher training. second, some experts missed some components of the procedure (a focus group discussion or a round of the delphi questionnaire). this limitation was compensated for by the repetitive nature of the procedure: participants who missed one component could still provide comments and additional input in the subsequent component. third, it is possible that teachers provided socially desirable answers since a self-report questionnaire was used in study 2. the variability between the items provides an indication that teachers did not simply rate themselves highly on all items to create a favourable impression. nevertheless, we state again that observational studies are necessary to investigate how self-reported use is related to observed use of differentiation strategies. fourth, it is unknown whether non-responders differed from teachers who did respond to the questionnaire, although the response rate of 83% is quite good. a strong combination of methodologies was used. study 1 employed an innovative methodology which combined the advantages of two methods: focus group discussions are suitable for creating shared understanding and generating ideas, while the delphi procedure gives all participants an equal and anonymous say in the systematic evaluation of those ideas. this combination was fruitful and efficient and we recommend it for future research. moreover, the collaboration with experts who were familiar with the daily practice of teaching enhanced the feasibility of the findings. based on their experience in various settings, the experts perceived these strategies as effective and feasible. in addition to this expert perspective, study 2 examined the results of study 1 from a teacher perspective. the fact that the teachers in study 2 reported to use most of the strategies recommended by the experts in study 1 shows that teachers acknowledge the need to differentiate and that the recommended strategies are largely compatible with teachers’ daily practice. at the same time, the discrepancy between the teachers’ and experts’ perception of the degree of application of differentiation opens up an interesting avenue for future research. thus, the inclusion of two complementary sources of information provides a richer perspective on differentiation. despite these methodological advantages, future research is necessary to test empirically whether the implementation of the strategies recommended by the experts leads to higher student achievement. the results of the current studies contribute to scientific research as well as to educational practice. although the experts in study 1 departed from a practical rather than a theoretical perspective, the elements of the cycle of differentiation overlap with elements of more general didactical models, such as van gelder’s didactical analysis model (van gelder, oudkerk pool, peters, & sixma, 1973) and de corte’s didaxology (de corte, geerligs, lagerweij, peters, & vandenberghe, 1976). apparently, effective readiness-based differentiation is consistent with the principles of general good teaching, with the addition that each stage of teaching needs to be differentiated. a first theoretical implication is that, rather than studying specific elements (i.e. differentiated practice) in isolation, it seems promising to move towards an integral view of differentiation which involves all stages of teaching. second, the experts clearly endorsed the view that students of different achievement levels have different educational needs and need different treatments at least part of the time, echoing the aptitude-treatment interaction literature (cronbach & snow, 1977). this emphasises the need to consider the potential variation between students in the design and analysis of educational intervention studies: what works for high-achieving students, may not work for low-achieving students and vice versa. at the practical level, the model and strategies for differentiation recommended by the experts can be used in teacher training and professional development, for which purpose they have also been published in a dutch journal for practitioners (van de weijer-bergsma & prast, 2013). current educational policies require teachers to implement differentiation and our results provide teachers with concrete advice on how to do this. the cycle of differentiation can be used as a framework to structure professional development about differentiation. it shows teachers that differentiation requires attention at all stages of teaching in one coherent approach. the recommended strategies provide teachers with practical suggestions for each step (these can be found in section 2.3, appendix 1, and also in the dsaq-items listed in table 4). the focus on prast&et &al & & & | f l r ! ! 110! mathematics promoted the concreteness of the results, since domain-specific guidelines can be applied directly without the need to transfer general principles. for example, the general guideline that advanced learners should be adequately challenged was made ready for use by providing achievement criteria for selecting high-performing students, suggestions for increasing task difficulty, guidelines for compacting, and a list of supplemental enrichment curricula. nonetheless, the principles behind these concrete recommendations, including the cycle of differentiation, seem to be applicable in other domains as well. future research could examine whether and how our results extend to other domains. study 2 provides researchers and practitioners with a new tool, the dsaq. researchers can use it, for example, as a preand post-assessment in intervention studies or to investigate teachers’ self-assessed implementation in other countries. in professional development, the dsaq can inform trainers about areas in which teachers perceive most room for improvement. theoretically, study 2 builds on the existing literature by providing support for the two-dimensional structure of differentiation. moreover, this study is the first to investigate the self-reported use of a broad range of strategies for differentiation in mathematics in the netherlands. the identified areas of low usage have practical implications, including the need to pay sufficient attention to differentiation for high-achieving students in teacher training and professional development.! keypoints effective combination of the delphi technique and focus group discussions to achieve consensus among experts use of experts’ practical knowledge to enhance feasibility of the findings a model and strategies for implementing differentiation in primary school mathematics, usable in teacher professional development a new questionnaire (the dsaq) to measure teachers’ self-assessed implementation of differentiation strategies, usable in future research teacher self-assessment indicating moderate to high usage and identification of relatively infrequently used strategies acknowledgements this work is part of the research programme ‘every child deserves differentiated mathematics education’, which is financed by the netherlands organisation for scientific research (nwo), grant number 411-10-753. the nwo was not involved in the collection, analysis, interpretation, or reporting of the data. we thank the consortium members for their valuable input during the consensus procedure. we are also grateful to the teachers who completed the questionnaire. finally, we thank the anonymous reviewers for their useful comments on a previous version of this article. prast&et &al & & & | f l r ! ! 111! references brown, j., & morris, d. (2005). meeting the needs of low spellers in a second-grade classroom. reading & writing quarterly, 21, 165-184. doi:10.1080/10573560590915969 corno, l. (2008). on teaching adaptively. educational psychologist, 43, 161-173. doi:10.1080/00461520802178466 cronbach, l. j., & snow, r. e. (1977). aptitudes and instructional methods: a handbook for research on interactions. new york: irvington. csikszentmihalyi, m. (1990). flow: the psychology of optimal experience. new york: harperperennial. de corte, e., geerligs, c., lagerweij, n., peters, j., & vandenberghe, r. (1976). beknopte didaxologie [concise didaxology]. groningen, the netherlands: wolters-noordhoff. deunk, m., doolaard, s., smale-jacobse, a., & bosker, r. j. (2015). differentiation within and across classrooms: a systematic review of studies into the cognitive effects of differentiation practices. groningen, the netherlands: gion. expertgroep doorlopende leerlijnen taal en rekenen (2008). over de drempels met taal en rekenen: consolideren, onderhouden, uitbreiden en verdiepen [overcoming barriers in mathematics: consolidating, maintaining, using and deepening]. enschede, the netherlands: expertgroep doorlopende leerlijnen taal en rekenen. retrieved from http://www.taalenrekenen.nl/referentiekader/rel_doc/downloads/rekenrapport.pdf fuchs, l., & fuchs, d. (2007). a model for implementing responsiveness to intervention. teaching exceptional children, 39(5), 14-20. doi:10.1177/004005990703900503 gelderblom, g. (2007). effectief omgaan met verschillen in het rekenonderwijs [dealing with differences in mathematics education effectively]. amersfoort, the netherlands: cps. gersten, r., chard, d. j., jayanthi, m., baker, s. k., morphy, p., & flojo, j. (2009). mathematics instruction for students with learning disabilities: a meta-analysis of instructional components. review of educational research, 79, 1202-1242. doi:10.3102/0034654309334431 goei, s. l., bekebrede, j., & bosma, t. (2011). teachers' sense of self efficacy scale: meningen van leraren, experimentele versie [teachers' sense of self efficacy scale: teacher’s opinions, experimental version]. amsterdam, the netherlands: onderwijscentrum vrije universiteit. hasson, f., keeney, s., & mckenna, h. (2000). research guidelines for the delphi survey technique. journal of advanced nursing, 32, 1008-1015. doi:10.1046/j.1365-2648.2000.t01-1-01567.x hu, l., & bentler, p. m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives. structural equation modeling: a multidisciplinary journal, 6, 1-55. doi:10.1080/10705519909540118 janson, d., & noteboom, a. (2004). compacten en verrijken van de rekenles voor (hoog)begaafde leerlingen in het basisonderwijs [compacting and enrichment of the mathematics lesson for highly able students in primary education]. enschede, the netherlands: slo. johnsen, s. k., haensly, p. a., ryser, g. a. r., & ford, r. f. (2002). changing general education classroom practices to adapt for gifted students. gifted child quarterly, 46, 45-63. doi:10.1177/001698620204600105 kulik, j. a., & kulik, c. c. (1992). meta-analytic findings on grouping programs. gifted child quarterly, 36, 73-77. doi:10.1177/001698629203600204 landeta, j. (2006). current validity of the delphi method in social sciences. technological forecasting and social change, 73, 467-482. doi:10.1016/j.techfore.2005.09.002 latz, a. o., speirs neumeister, k. l., adams, c. m., & pierce, r. l. (2009). peer coaching to improve classroom differentiation: perspectives from project clue. roeper review, 31, 27-39. doi:10.1080/02783190802527356 liamputtong, p. (2011). focus group methodology: principles and practice. london, uk: sage. prast&et &al & & & | f l r ! ! 112! lou, y., abrami, p. c., spence, j. c., poulsen, c., chambers, b., & d'appolonia, s. (1996). within-class grouping: a meta-analysis. review of educational research, 66, 423-458. doi:10.3102/00346543066004423 mcdonald connor, c., morrison, f. j., fishman, b., giuliani, s., luck, m., underwood, p. s., . . . schatschneider, c. (2011a). testing the impact of child characteristics x instruction interactions on third graders' reading comprehension by differentiating literacy instruction. reading research quarterly, 46, 189-221. doi:10.1598/rrq.46.3.1 mcdonald connor, c., morrison, f. j., fishman, b. j., schatschneider, c., & underwood, p. (2007). algorithm-guided individualized reading instruction. science, 315, 464-465. doi:10.1126/science.1134513 mcdonald connor, c., morrison, f. j., schatschneider, c., toste, j. r., lundblom, e., crowe, e. c., & fishman, b. (2011b). effective classroom instruction: implications of child characteristics by reading instruction interactions on first graders’ word reading achievement. journal of research on educational effectiveness, 4, 173-207. doi:10.1080/19345747.2010.510179 mcdonald connor, c., piasta, s. b., fishman, b., glasney, s., schatschneider, c., crowe, e., . . . morrison, f. j. (2009). individualizing student instruction precisely: effects of child x instruction interactions on first graders' literacy development. child development, 80, 77-100. doi:10.1111/j.14678624.2008.01247.x mcleskey, j., & waldron, n. l. (2002). inclusion and school change: teacher perceptions regarding curricular and instructional adaptations. teacher education and special education, 25, 41-54. doi:10.1177/088840640202500106 mcleskey, j., & waldron, n. l. (2011). educational programs for elementary students with learning disabilities: can they be both effective and inclusive? learning disabilities research & practice, 26, 48-57. doi:10.1111/j.1540-5826.2010.00324.x megay-nespoli, k. (2001). beliefs and attitudes of novice teachers regarding instruction of academically talented learners. roeper review, 23, 178-182. doi:10.1080/02783190109554092 moon, t. r., callahan, c. m., tomlinson, c. a., & miller, e. m. (2002). middle school classrooms: teachers' reported practices and student perceptions (no. rm02164). storrs, ct: the national research center on the gifted and talented. muthén, l. k., & muthén, b. o. (1998-2012). mplus user’s guide (7th ed.). los angeles, ca: muthén & muthén. nationaal expertisecentrum leerplanontwikkeling (2010). professionele ontwikkeling: tussenmeting project 'als je merkt dat het werkt' [professional development: midtime measurement project 'when you notice it works']. enschede, the netherlands: slo. nurmi, j., kiuru, n., lerkkanen, m., niemi, p., poikkeus, a., ahonen, t., . . . lyyra, a. (2013). teachers adapt their instruction in reading according to individual children's literacy skills. learning and individual differences, 23, 72-79. doi:http://dx.doi.org/10.1016/j.lindif.2012.07.012 reis, s. m., gubbins, e. j., richards, s., briggs, c. j., jacobs, j. k., eckert, r. d., . . . schreiber, f. j. (2004). reading instruction for talented readers: case studies documenting few opportunities for continuous progress. gifted child quarterly, 48, 315-338. doi:10.1177/001698620404800406 reis, s. m., mccoach, d. b., little, c. a., muller, l. m., & kaniskan, r. b. (2011). the effects of differentiated instruction and enrichment pedagogy on reading achievement in five elementary schools. american educational research journal, 48, 462-501. doi:10.3102/0002831210382891 rogers, k. b. (2007). lessons learned about educating the gifted and talented: a synthesis of the research on educational practice. gifted child quarterly, 51, 382-396. doi:10.1177/0016986207306324 roy, a., guay, f., & valois, p. (2013). teaching to address diverse learning needs: development and validation of a differentiated instruction scale. international journal of inclusive education, 17, 11861204. doi:10.1080/13603116.2012.743604 schram, e., van der meer, f., & van os, s. (2013). omgaan met verschillen: (g)een kwestie van maatwerk [responding to differences: (not) a matter of customisation] (no. an2.6547.542). enschede, the netherlands: slo. prast&et &al & & & | f l r ! ! 113! scott, b. j., vitale, m. r., & masten, w. g. (1998). implementing instructional adaptations for students with disabilities in inclusive classrooms: a literature review. remedial and special education, 19, 106-119. doi:10.1177/074193259801900205 skaalvik, e. m., & skaalvik, s. (2007). dimensions of teacher self-efficacy and relations with strain factors, perceived collective teacher efficacy, and teacher burnout. journal of educational psychology, 99, 611625. doi:10.1037/0022-0663.99.3.611 slavin, r. e. (1987). ability grouping and student achievement in elementary schools: a best evidence synthesis. review of educational research, 57, 293-336. doi:10.3102/00346543057003293 streiner, d. l. (2003). starting at the beginning: an introduction to coefficient alpha and internal consistency. journal of personality assessment, 80, 99-103. doi:10.1207/s15327752jpa8001_18 teddlie, c., creemers, b., kyriakides, l., muijs, d., & yu, f. (2006). the international system for teacher observation and feedback: evolution of an international study of teacher effectiveness constructs. educational research & evaluation, 12, 561-582. doi:10.1080/13803610600874067 tieso, c. l. (2003). ability grouping is not just tracking anymore. roeper review, 26, 29-36. doi:10.1080/02783190309554236 tieso, c. l. (2005). the effects of grouping practices and curricular adjustments on achievement. journal for the education of the gifted, 29, 60-89. tomlinson, c. a. (2005). how to differentiate instruction in mixed-ability classrooms (2nd ed.). upper saddle river, nj: pearson education. tomlinson, c. a., brighton, c., hertberg, h., callahan, c. m., moon, t. r., brimijoin, k., . . . reynolds, t. (2003). differentiating instruction in response to student readiness, interest and learning profile in academically diverse classrooms: a review of literature. journal for the education of the gifted, 27, 119-145. doi:10.1177/016235320302700203 tschannen-moran, m., & woolfolk hoy, a. (2001). teacher efficacy: capturing an elusive construct. teaching and teacher education, 17, 783-805. doi:10.1016/s0742-051x(01)00036-1 van de weijer-bergsma, e., & prast, e. j. (2013). gedifferentieerd primair rekenonderwijs volgens experts: de resultaten uit een delphi-onderzoek [differentiated primary math education according to experts: results from a delphi study]. orthopedagogiek: onderzoek en praktijk, 52, 336-349.! van den broek-d’obrenan, v., van cauwenberghe, c., van dongen, d., drewes, i., knuver, a., lincklaen arriëns, k., . . . de vries, b. (2012). de staat van het onderwijs: onderwijsverslag 2010-2011 [the state of education: educational report 2010-2011]. utrecht, the netherlands: inspectie van het onderwijs. retrieved from http://www.onderwijsinspectie.nl/binaries/content/assets/onderwijsverslagen/2012/onderwijsverslag_2 010_2011_printversie.pdf van gelder, l., oudkerk pool, t., peters, j., & sixma, j. (1973). didactische analyse [didactical analysis]. groningen, the netherlands: wolters-noordhoff. van groenestijn, m., borghouts, c., & janssen, c. (2011). protocol ernstige rekenwiskunde-problemen en dyscalculie [protocol severe mathematics difficulties and dyscalculia]. assen, the netherlands: van gorcum. vantassel-baska, j., feng, a. x., brown, e., bracken, b., stambaugh, t., french, h., . . . bai, w. (2008). a study of differentiated instructional change over 3 years. gifted child quarterly, 52, 297-312. doi:10.1177/0016986208321809 vantassel-baska, j., & stambaugh, t. (2005). challenges and possibilities for serving gifted learners in the regular classroom. theory into practice, 44, 211-217. doi:10.1207/s15430421tip4403_5 vygotsky, l. s. (1978). mind in society: the development of higher psychological processes. cambridge, ma: harvard university press. westberg, k. l., archambault, f. x., dobyns, s. m., & salvin, t. j. (1993). the classroom practices observation study. journal for the education of the gifted, 16, 120-146. doi:10.1177/016235329301600204 westberg, k. l., & daoust, m. e. (2003). the results of the replication of the classroom practices survey replication in two states. retrieved from http://www.gifted.uconn.edu/nrcgt/newsletter/fall03/fall03.pdf prast&et &al & & & | f l r ! ! 114! ysseldyke, j., & bolt, d. m. (2007). effect of technology-enhanced continuous progress monitoring on math achievement. school psychology review, 36, 453-467. retrieved from http://www.nasponline.org/publications/spr/abstract.aspx?id=1847 ysseldyke, j., spicuzza, r., kosciolek, s., & boys, c. (2003). effects of a learning information system on mathematics achievement and classroom structure. the journal of educational research, 96, 163-173. doi:10.1080/00220670309598804 ysseldyke, j., spicuzza, r., kosciolek, s., teelucksingh, e., boys, c., & lemkuil, a. (2003). using a curriculum-based instructional management system to enhance math achievement in urban schools. journal of education for students placed at risk, 8, 247-265. doi: 10.1207/s15327671espr0802_4 ysseldyke, j., & tardrew, s. (2007). use of a progress monitoring system to enable teachers to differentiate mathematics instruction. journal of applied school psychology, 24, 1-28. doi:10.1300/j370v24n01_01 appendix 1: summarising concepts included in the delphi questionnaire what follows are the translations of the concepts as they were included in the delphi questionnaire. background information that might be relevant for non-dutch readers is given in the footnotes. organisation the starting point is convergent differentiation2. students are assigned to one of three subgroups based on standardised tests and / or curriculum-based tests. if curriculum-based tests are used to assess what students already master, the test score of the previous unit can be used, but an alternative is to use the end-ofunit test of the upcoming unit as a pretest. the teacher can change the grouping arrangement for a certain unit or lesson based on test scores. during mathematics classes, whole-class instruction, instruction to one of the subgroups (of low-achieving or high-achieving students) and independent practice are alternated. average achievers take part in the whole-class instruction and receive individual feedback or guidance during the time for independent practice. differentiation in instruction during whole-class instruction, the teacher serves different levels and educational needs to the extent that this is possible. the teacher can do this by teaching at different levels of abstraction and showing the connection between these different levels. the teacher should ask questions at varying difficulty levels, implying that some questions may be too easy or too difficult for some of the students in the class. during instruction to a subgroup, the teacher takes into account the educational needs of the students in that specific subgroup. for example, the teacher spends more time on lower levels of abstraction when teaching the lowachieving subgroup, while instruction to the high-achieving subgroup is mainly at a high level of abstraction. additionally, it is assumed that the low-achieving subgroup needs more guidance (more direct instruction) than the high-achieving subgroup (more exploratory instruction). to the extent possible, the teacher also takes into account individual differences within a subgroup. for example, the teacher can accommodate to a student’s need to verbalise a solution strategy himself, or a student’s need for visualisation. an additional strategy for differentiation in instruction is the use of instructional videos. !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 2 in the netherlands, a distinction is often made between convergent and divergent differentiation (gelderblom, 2007). in convergent differentiation, all student in a classroom work on roughly the same topics at the same time (even if they might engage in the topic at varying levels of complexity). in divergent differentiation, different students work on different learning goals and topics at the same time. prast&et &al & & & | f l r ! ! 115! differentiation in goals a strong awareness of the learning trajectories and accompanying educational goals is essential for a good lesson. for differentiated education, this means that different goals are set for different students. goals are differentiated primarily at the subgroup level. highly competent teachers can also differentiate goals on an individual basis based on their professional insight. for the low-achieving subgroup, the objective is to master the fundamental level (1f)3 at the end of primary school. for the average-achieving subgroup, the objective is to master the target level (1s) at the end of primary school. for the high-achieving subgroup, mastery of the target level is a minimum requirement, but additionally, more advanced goals (for example regarding logical reasoning) are set for these students. the goals for the end of primary school are converted into specific lesson goals for the three subgroups based on the curriculum and the professional insight of the teacher. these lesson goals should be both ambitious and realistic. the teacher keeps in mind the lesson goals while preparing and teaching his lesson. after the lesson, the teacher evaluates whether lesson goals have been met. differentiation in the practice phase the different subgroups need quantitatively and / or qualitatively different practice tasks. from the tasks that the curriculum offers, the tasks at the minimum and fundamental level are most important for low-achieving students. the high-achieving subgroup can skip a large proportion of the tasks at minimum and fundamental level. existing guidelines for compacting4 can be used to select the tasks that high-achieving students do need to do. high achieving students spend the time that is freed up by compacting the regular material on enrichment. the enrichment tasks provided in the regular curriculum are often not sufficiently challenging, especially for gifted students. additional enrichment should be provided for these students. such enrichment may include assignments for which students have to carry out research or use information from different sources. besides the adaptation (selection and supplementation) of tasks, practice can also be differentiated during instruction to subgroups. for example, the teacher could use the extended instruction for lowachieving students to solve the exercises together step-by-step, while a discussion of the big ideas behind a certain task may be more useful in the high-achieving subgroup. differentiation based on learning styles the educational needs of different students may also vary within subgroups. for example, students may have a preference for certain formats (whole-class instruction, working together, working individually) and certain input modes (visual or verbal, written or spoken). it has been mentioned repeatedly that it is important for some students to express the content themselves or to have the content explained to them by another student. teachers need to be aware of these differences between students and learn how they can vary their instruction, tasks and formats to accommodate various educational needs. especially during the instruction to subgroups the teacher can accommodate to individual educational needs, provided that he is able to identify what kind of instruction or type of task a student needs to understand the content. !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 3 in the dutch educational system, overarching objectives (comparable to the common core state standards employed in the us) have been defined at two levels: the fundamental level (1f) that should be reached by all students and the target level (1s) that should be reached by about 65% of the students (expertgroep doorlopende leerlijnen taal en rekenen, 2008). 4 an educational advisory company has published guidelines for compacting the most popular dutch mathematics curricula (janson & noteboom, 2004). children for whom the material should be compacted receive an additional booklet with an overview per lesson of the exercises they should do and the exercises they can skip. prast&et &al & & & | f l r ! ! 116! differentiation in kindergarten when the files of students with problematically low mathematics achievement in primary school are examined, it often turns out that problems with preparatory mathematics were already detected in kindergarten but that no or insufficient action has been taken to tackle those problems in the meantime. factors that may play a role in this lack of action are beliefs of the teacher (‘the child is not ready for preparatory mathematics’ or ‘children of this age should be allowed to play’), inadequate communication to the teacher of the next grade, and lack of knowledge of ways to tackle low achievement in (preparatory) mathematics. in order to respond more quickly to early signals of problems with acquiring preparatory mathematics skills, teachers should (a) set more specific and ambitious goals (what should the child be able to do at the beginning of grade 1?), (b) be more knowledgeable about levels of abstraction and be able to demonstrate the connections between various levels of abstraction, (c) realise that learning can take place in the process of playing if the activity is well adapted to the child’s educational needs, and (d) that certain children need some additional instruction, also in kindergarten. additionally, more attention should be given to providing extra challenge to students with highly developed preparatory mathematics skills. table 5 table of footnotes no. footnote 1 compared to the four-item ntses subscale for adapting instruction to individual students’ needs, the added value of the dsaq is that it provides a more detailed measure of self-assessed use of a range of differentiation strategies. 2 in the netherlands, a distinction is often made between convergent and divergent differentiation (gelderblom, 2007). in convergent differentiation, all student in a classroom work on roughly the same topics at the same time (even if they might engage in the topic at varying levels of complexity). in divergent differentiation, different students work on different learning goals and topics at the same time. 3 in the dutch educational system, overarching objectives (comparable to the common core state standards employed in the us) have been defined at two levels: the fundamental level (1f) that should be reached by all students and the target level (1s) that should be reached by about 65% of the students (expertgroep doorlopende leerlijnen taal en rekenen, 2008). 4 an educational advisory company has published guidelines for compacting the most popular dutch mathematics curricula (janson & noteboom, 2004). children for whom the material should be compacted receive an additional booklet with an overview per lesson of the exercises they should do and the exercises they can skip. ! frontline learning research 6 (2014) 7-14 issn 2295-3159 corresponding author: bert de smedt, leopold vanderkelenstraat 32 box 3765, b-3000 leuven, belgium. e-mail: bert.desmedt@ppw.kuleuven.be doi: http://dx.doi.org/10.14786/flr.v2i4.115 7 | f l r advances in the use of neuroscience methods in research on learning and instruction bert de smedt a a faculty of psychology and educational sciences, university of leuven, belgium article received 27 may 2014 / revised 19 september 2014 / accepted 11 november 2014 / available 23 december 2014 abstract cognitive neuroscience offers a series of tools and methodologies that allow researchers in the field of learning and instruction to complement and extend the knowledge they have accumulated through decades of behavioral research. the appropriateness of these methods depends on the research question at hand. cognitive neuroscience methods allow researchers to investigate specific cognitive processes in a very detailed way, a goal in some but not all fields of the learning sciences. this value added will be illustrated in three ways, with examples in field of mathematics learning. firstly, cognitive neuroscience methods allow one to understand learning at the biological level. secondly, these methods can help to measure processes that are difficult to access by means of behavioral techniques. finally, and more indirectly, neuroimaging data can be used as an input for research on learning and instruction. this paper concludes with highlighting the challenges of applying neuroscience methods to research on learning and instruction. frontline: cognitive neuroscience offers a series of tools and methodologies that allow researchers in the field of learning and instruction to complement and extend the knowledge they have accumulated through decades of behavioral research. the appropriateness of these methods depends on the research question at hand. keywords: cognitive neuroscience; methods; mathematics learning; educational neuroscience b. de smedt 8 | f l r 1. introduction non-invasive brain imaging methods, such as event-related potentials (erp) or functional magnetic resonance imaging (frmi), represent a series of tools and methodologies that allow researchers in the field of learning and instruction to complement and extend the knowledge they have already accumulated through decades of behavioral research (e.g., cacioppo, berntson, & nusbaum, 2008; de smedt, ansari, et al., 2011). the potential application of these methods depends on the research question at hand. after a brief discussion of relevant brain imaging methods, i will use the field of mathematics learning to illustrate three ways in which these methods can be used in research on learning and instruction. this paper concludes with highlighting some challenges of applying such methods to research on learning and instruction. 2. neuroscience methods when considering different methods that are used by neuroscientists to study the structure and function of the brain, it is important to point out that neuroscience is a very broad field that includes a variety of disciplines ranging from cellular and molecular neuroscience to cognitive neuroscience (e.g., squire et al., 2013). i restrict the focus here to cognitive neuroscience and its methods (ward, 2006), because this sub-field of neuroscience is the closest to research on learning and instruction, given its focus on the neural mechanisms that underlie human cognition and behavior. a detailed description of these cognitive neuroscience methods is beyond the scope of this contribution and excellent introductions are provided by ward (2006) and dick, lloyd-fox, blasi, elwell, & mills (2014). sometimes, psychophysiological measures, such as skin conductance, heart rate or eyemovement data, are also denoted as neuroscience methods. although these methods tap into the nervous system, they are not direct measures of brain structure or function and therefore they are not considered here. the transmission of information in the brain from one cell to the other occurs through electrical signals, and this electrical activity of the brain can be captured by methods such as electroencephalography (eeg), which requires a cap of electrodes to be mounted on the head of a participant, and magnetoencephalography (meg) (ward, 2006). on one hand, the advantage of these methods is that they can measure the activity of the brain in response to a particular stimulus (i.e. event-related activity) at a very accurate temporal scale and they are particularly suited to investigate when a process is taking place. on the other hand, a large number of stimuli of a particular type (typically a few dozens) are needed in order to reliably estimate the brain signal in response to that stimulus. another series of methods are magnetic resonance imaging (mri) techniques, which use large magnetic fields and the magnetic properties of hydrogen atoms in brain tissue or in blood to visualize brain structure and brain function, respectively (ward, 2006). these data are acquired in a specific and very noisy environment, the mri scanner, in which participants have to lie still and are not allowed to move more than a few millimeters. this category of methods can investigate the structure of the brain, i.e. the gray or white matter, and how this structure is related to performance or changes as a result of learning. interesting examples are provided by supekar et al. (2013), who showed that the size of the hippocampus predicted the performance gains in response to one-on-one math tutoring and by keller b. de smedt 9 | f l r and just (2009), who showed that intensive remedial reading instruction resulted in changes in white matter in poor readers. mri also allows us to investigate brain function, a technique that is called functional mri or fmri, which is one of the most common techniques used in cognitive neuroscience (ward, 2006). functional mri is an indirect way of assessing the brain’s activity and measures the level of oxygen in the blood. the assumption is that an increase in oxygen level is the result of the vascular system’s response to an increase in brain activity. mri methods are very accurate on a spatial scale and are particularly suited to investigate where in the brain a particular process is taking place. due to the practical constraints of the mri-environment (e.g., noise, no movement) the type of tasks that participants can complete is limited, yet progress is being made over the last years to use more complex tasks, such as playing video games (anderson et al., 2011) and even face-to-face interaction (e.g., redcay et al., 2010). it is crucial to point out that the measures reviewed above, i.e. signals indicating brain structure or function, can only be meaningfully interpreted by linking them to cognitive theories (e.g., cacioppo et al., 2008; de smedt, ansari, et al., 2011). furthermore, the collection of behavioral data represents a necessary step in most studies in cognitive neuroscience (e.g., ward, 2006). in all, a detailed cognitive theory of the phenomenon under investigation is crucial to design and interpret cognitive neuroscience data and to apply cognitive neuroscience methods to the field of learning and instruction. 3. application to research on learning and instruction how can the cognitive neuroscience methods reviewed above advance the field of learning and instruction? this depends on the research question at hand, and only some but certainly not all types of research questions in the field of learning and instruction might benefit from the use of cognitive neuroscience methods. stern and schneider (2010) provided a nice analogy for determining when these cognitive neuroscience tools and theories could be appropriate. they compared this issue with the use of a digital road map. when using a digital road map for looking at the field learning and instruction, the appropriate resolution of the map depends on what the map viewer is looking for, alleys (micro-level) vs. highways (macro-level), and users can zoom in and out between different levels of resolution. some questions only focus at the broader context of learning (macro-level), as is the case in large-scale research on educational systems, and are at a low level of resolution. others aim to unravel the very specific cognitive processes that underlie learning, and this requires a map at very high resolution (micro-level). it is at this micro-level of understanding of such specific cognitive processes that cognitive neuroscience methods can be applied in the field of learning and instruction. i will use the field of mathematics learning to illustrate three ways in which cognitive neuroscience methods can be useful for research in learning and instruction (see also de smedt et al., 2010; de smedt, ansari, et al. 2011; de smedt & grabner, 2015). 3.1 understanding learning at the biological level neuroimaging data allow us to examine at the biological level how people learn. such data can provide converging evidence for findings that have been obtained through psychological and educational research. this convergence of findings from different research methodologies has the b. de smedt 10 | f l r potential to provide a better and more complete understanding how typical and atypical learning takes place (e.g., de smedt, ansari, et al., 2011; lieberman, schreiber, & ochsner, 2003). for example, how do people acquire and apply different strategies to solve elementary arithmetic problems, such as 5 + 9 or 4 × 3? decades of behavioral research have revealed that these problems are either solved by using fact retrieval from declarative memory or by using procedural strategies, such as counting, and developmental data indicate that children develop an increasing reliance on arithmetic fact retrieval, while the use of procedures to solve such elementary problems decreases over time (e.g., siegler, 1996). research in cognitive neuroscience is now beginning to understand on how this learning of arithmetic is reflected at the neural level (e.g., arsiladou & taylor, 2011; zamarian, ischebeck, & delazer, 2009). in a series of studies, we have tried to investigate this issue with eeg (de smedt, grabner, & studer, 2009; grabner & de smedt, 2011; grabner & de smedt, 2012). in these studies, adults had to solve a series of addition, subtraction and multiplication problems, while their brain activity was recorded with eeg, and they had to verbally report on a trial-by-trial basis on the strategies they used to solve the presented problems. these studies had two aims. first, we wanted to verify whether these two types of strategies were reflected in different brain activity patterns and whether fact retrieval training resulted in changes in brain activity that reflected a shift in strategy use. second, we aimed to test if cognitive neuroscience methods, such as eeg, can be used as a way of methodological triangulation to further validate the use of verbal report data. these data are typically used in behavioral research to investigate strategy use but their validity has been debated (e.g., kirk & ashcraft, 2001). the eeg data revealed different patterns of activity for the two types of strategies: oscillations in the theta band (3–6 hz) were associated with fact retrieval whereas oscillations in the lower alpha band (8–10 hz) were related to procedural strategies (grabner & de smedt, 2011). when we trained participants in using fact retrieval strategies, we were also able to show that the well-known behavioral shift from procedural strategies to fact retrieval as a function of training was also reflected in specific changes in brain activity, i.e. training-related activity increases in the theta band and decreases in the lower alpha band (grabner & de smedt, 2012). combining verbal strategy reports with reaction times on specific problem types and neuroimaging data allowed us to further examine the validity of these verbal reports. this type of methodological triangulation confirmed that verbal strategy reports are a valid way to capture strategies in mental arithmetic. in all, this convergence of findings obtained by different research methods at behavioral and biological levels provides a more solid empirical ground for our theories on strategy development. 3.2 measuring difficult-to-access processes neuroimaging data can provide a level of analysis and measurement that cannot be accessed by behavioral data alone. examples of this application can be observed in the study of individual differences between learners and in understanding the origins of atypical development. de smedt, holloway, & ansari (2011) used fmri to investigate brain activity in 10-12-year-old children during addition and subtraction and compared children with low and average levels of arithmetical competence, who significantly differed in their performance on a standardized arithmetic fluency test. although both groups of children did not differ in a simple calculation task at the behavioral level (i.e. accuracy, speed) during the acquisition of the fmri data, the authors observed significant group differences in brain activity in the right intraparietal sulcus, a brain region that is known to play a key role in the processing of numerical magnitudes: children with low levels of arithmetical competence showed higher activation in this region during the solutions of problems with a relatively small b. de smedt 11 | f l r problem size. the interpretation of these data in the context of neurocognitive theories of numerical magnitude processing and arithmetic development (e.g., ansari, 2008; butterworth, varma, & laurillard 2011) suggests the use compensatory strategies and generates predictions that should be further exploited in subsequent research. for example, it might be that the children with low arithmetical competence in the study of de smedt, holloway, et al. (2011) continued to rely to a greater extent on quantity-based strategies (such as counting or procedural calculation) on those problems that children with relatively higher arithmetical competence already retrieved from their memory, a possibility that should be evaluated in subsequent research. in all, this indicates that brain imaging data can uncover subtle processing differences between groups of learners that may not be detected through the measurement of behavioral data alone, illustrating the high resolution level which cognitive neuroscience methods are able to capture. 3.3 input for research on learning and instruction studies in cognitive neuroscience can also have an indirect impact on research in learning and instruction, by drawing our attention to specific fine-grained cognitive processes that are implicated in different types of learning (see aue, lavelle, & cacioppo, 2009 for a similar rationale in the field of psychology). such data have the potential to generate new hypotheses that can be tested in research on learning and instruction. for example, neuroimaging studies on how the brain processes numbers have revealed that the intraparietal sulci (ips) are consistently active whenever we have to perform numerical and arithmetical tasks and that this structure supports the processing of numerical magnitudes (e.g., ansari, 2008; dehaene, piazza, pinel, & cohen, 2003). brain imaging studies in children with developmental dyscalculia, a learning disorder that is characterized by severe and persistent difficulties in acquiring mathematical competencies, point to structural and functional abnormalities in the ips in these children (e.g., butterworth et al., 2011; price & ansari, 2013 for a review). this all suggests that the processing of numerical magnitudes is potentially a key to successful mathematical development and this processing might be compromised in developmental dyscalculia (dd). this suggestion has fueled a large number of psychological and educational studies that have empirically confirmed this hypothesis at the behavioral level (see de smedt, noël, gilmore, & ansari, for a review), by consistently showing that individuals with dd have significant impairments in their ability to compare (symbolic) numbers. more broadly, these studies have also furthered our understanding of individual differences in typical mathematical development, as the ability to compare (symbolic) numbers is predictive of subsequent mathematical development (see de smedt et al., 2013, for a review). this research has impacted on studies in the field of learning and instruction, through the development and evaluation of specific interventions (e.g., de smedt et al., 2013) and diagnostic instruments that can be used for the screening and early identification of at-risk children (nosworthy, bugden, archibald, evans, & ansari, 2013). it is important to point out that even if these studies do not collect measures of brain activity or structure, they rely to some extent on insights gleaned from cognitive neuroscience studies. used in this way, cognitive neuroscience data might set the stage for new educational research and it can, albeit indirectly, enhance our understanding of learning. 4. challenges the application of cognitive neuroscience methods to research on learning and instruction also imposes some challenges and caveats that one needs to be aware of (see also ansari, de smedt, & b. de smedt 12 | f l r grabner, 2012; de smedt & grabner, 2015), which are not specific to the domain of mathematics learning. these challenges deal with the issue of external validity or generalizability as well as the scope of biological data and explanations (e.g., beck, 2010). it is important to point out that most of the existing studies in cognitive neuroscience involved adult participants and that these methods are not so easy to apply in children. this is because the acquisition of neuroimaging data is very sensitive to movement and motion artefacts in children often negatively impact on data acquisition. progress is being made in the reduction of such artefacts, for example by training children to keep still when such data are being collected (de bie et al., 2010). at a more theoretical level, cognitive neuroscience findings obtained in adult participants cannot be readily generalized to the developing brain and the learning of children and adolescents, as the human brain undergoes massive structural and functional changes throughout childhood and adolescence (ansari, 2010). the tasks used in most cognitive neuroscience studies are very elementary and differ from the rich and complex tasks that are typically solved in everyday learning environments and that are used in research learning and instruction. such complex tasks cannot be easily administered in cognitive neuroscience studies for various reasons. as indicated above, there are practical constraints related to the laboratory environment in which neuroimaging data are being collected. in order to obtain reliable data on brain activity during a particular task, a large number of trials of the same task need to be presented. these tasks need to be very elementary, because the larger the number of cognitive processes in a particular task, the more difficult it will be to disentangle these cognitive processes physiologically. one way to resolve this is to correlate data acquired in very constrained laboratory settings to ecologically valid measures of learning (see price, mazzocco, & ansari (2013), for an example). one important caveat deals with the scope of a neuroscientific data and explanations. there might be an inappropriate belief that neuroscientific data are more convincing, informative and valid than behavioral data (beck, 2010). on the contrary, knowledge gained through cognitive neuroscience methods should be considered at the same level of data obtained by standard behavioral methods in learning and instruction. there should be no knowledge hierarchy, but an appreciation of multiple sources of data to better understand how learning takes place and how it can be fostered (de smedt, ansari, et al., 2011). 5. conclusion the application of neuroscience methods to research on learning and instruction depends on the level of the research question. when interested in very specific low-level processes, neuroimaging data have the potential to help understanding learning at the biological level, to measure processes that are difficult to access via behavioral data and to generate and test hypotheses for educational phenomena that can be subsequently investigated via behavioral research on learning and instruction. keypoints the application of neuroscience methods to research on learning and instruction depends on the level of the research question. b. de smedt 13 | f l r cognitive neuroscience methods allow researchers to investigate specific cognitive processes in a very detailed way, a goal in some but not all fields of the learning sciences neuroimaging data have the potential to understand learning at the biological level, to measure process that are difficult to access via behavioral data and to generate hypotheses for subsequent research on learning and instruction. acknowledgments this work is partially supported by grant goa 2012/010. references anderson, j. r., bothell, d., fincham, j. m., anderson, a. r., poole, b., & qin, y. l. (2011). brain regions engaged by partand whole-task performance in a video game: a model-based test of the decomposition hypothesis. journal of cognitive neuroscience, 23, 3983-3997. doi: 10.1162/jocn_a_00033 ansari, d. (2008). effects of development and enculturation on number representation in the brain. nature reviews neuroscience, 9, 278-291. doi: 10.1038/nrn2334 ansari, d. (2010). neurocognitive approaches to developmental disorders of numerical and mathematical cognition: the perils of neglecting the role of development. learning and individual differences, 20, 123-129. doi:10.1016/j.lindif.2009.06.001 ansari, d., de smedt, b., & grabner, r. (2012). neuroeducation a critical overview of an emerging field. neuroethics, 5, 105-117. doi: 10.1007/s12152-011-9119-3 arsalidou, m., & taylor, m. j. (2011). is 2 + 2 = 4? meta-analyses of brain areas needed for numbers and calculations. neuroimage, 54, 2382-2393. doi: 10.1016/j.neuroimage.2010.10.009 aue, t., lavelle, l. a., & cacioppo, j. t. (2009). great expectations: what can fmri research tell us about psychological phenomena? international journal of psychophysiology, 73, 10-16. doi: 10.1016/j.ijpsycho.2008.12.017 beck, d. m. (2010). the appeal of the brain in the popular press. perspectives on psychological science, 5, 762-766. doi: 10.1177/1745691610388779 butterworth, b., varma, s., & laurillard, d. (2011). dyscalculia: from brain to education. science, 332, 1049-1053. doi: 10.1126/science.1201536 cacioppo, j. t., berntson, g. g., & nusbaum, h. c. (2008). neuroimaging as a new tool in the toolbox of psychological science. current directions in psychological science,17, 62-67. doi: 10.1111/j.1467-8721.2008.00550.x de bie, h. m. a., boersma, m., wattjes, m. p., adriaanse, s., vermeulen, r. j., oostrom, k. j., huisman, j., veltman, d. j., & delemarre-van de waal, h. a. (2010). preparing children with a mock scanner training protocol results in high quality structural and functional mri scans. european journal of pediatrics, 169, 1079-1085. doi: 10.1007/s00431-010-1181-z de smedt, b., ansari, d., grabner, r. h., hannula, m. m., schneider, m., & verschaffel, l. (2010). cognitive neuroscience meets mathematics education. educational research review, 5, 97105. doi:10.1016/j.edurev.2009.11.001 de smedt, b., ansari, d., grabner, r.h., hannula-sormunen, m., schneider, m., & verschaffel, l. (2011). cognitive neuroscience meets mathematics education: it takes two to tango. educational research review, 6, 232-237. doi:10.1016/j.edurev.2011.10.003 de smedt, b., & grabner, r. h. (2015). applications of neuroscience to mathematics education. in a. dowker & r. cohen-kadosh (eds.) the oxford handbook of mathematical cognition. oxford: oxford university press. doi: 10.1093/oxfordhb/9780199642342.013.48 de smedt, b., grabner, r. h., & studer, b. (2009). oscillatory eeg correlates of arithmetic strategy use in addition and subtraction. experimental brain research, 195, 635-642. doi b. de smedt 14 | f l r 10.1007/s00221-009-1839-9 de smedt, b., holloway, i. d., & ansari, d. (2011). effects of problem size and arithmetic operation on brain activation during calculation in children with varying levels of arithmetical fluency. neuroimage ,57, 771-781. doi:10.1016/j.neuroimage.2010.12.037 de smedt, b., noël, m. p., gilmore, c., & ansari, d. (2013). the relationship between symbolic and non-symbolic numerical magnitude processing skills and the typical and atypical development of mathematics: a review of evidence from brain and behavior. trends in neuroscience and education, 2, 48-55. doi: 10.1016/j.tine.2013.06.001 dehaene, s., piazza, m., pinel, p., & cohen, l. (2003). three parietal circuits for number processing. cognitive neuropsychology, 20, 487-506. doi: 10.1080/02643290244000239 dick, f., lloyd-fox, s., blasi, a., elwell, c., & mills, d. (2014). neuroimaging methods. in d. mareschal, b. butterworth, & a. tolmie (eds.) educational neuroscience. (pp. 13-45). malden, ma: wiley-blackwell. grabner, r. & de smedt, b. (2011). neurophysiological evidence for the validity of verbal strategy reports in mental arithmetic. biological psychology, 87, 128-136. doi:10.1016/j.biopsycho.2011.02.019 grabner, r. h., & de smedt, b. (2012). oscillatory eeg correlates of arithmetic strategies: a training study. frontiers in psychology, 3(428), 1-11. doi: 10.3389/fpsyg.2012.00428 keller, t. a., & just, m. a. (2009). altering cortical connectivity: remediation-induced changes in the white matter of poor readers. neuron, 64, 624-631. doi: 10.1016/j.neuron.2009.10.018 kirk, e. p., & ashcraft, m. h. (2001). telling stories: the perils and promise of using verbal reports to study math strategies. journal of experimental psychology-learning memory and cognition, 27, 157-175. lieberman, m. d., schreiber, d., & ochsner, k. n. (2003). is political cognition like riding a bicycle? how cognitive neuroscience can inform research on political thinking. political psychology, 24, 681-704. doi: 10.1046/j.1467-9221.2003.00347.x nosworthy, n., bugden, s., archibald, l., evans, b., & ansari, a. (2013). a two-minute paper-andpencil test of symbolic and nonsymbolic numerical magnitude processing explains variability in primary school children's arithmetic competence. plos one, 8, e67918. doi: 10.1371/journal.pone.0067918 price, g. r., & ansari, d. (2013). dyscalculia. in o. dulac & m. lassonde (eds.) handbook of clinical neurology (pp. 241-244). london: elsevier. price, g. r., mazzocco, m. m. m., & ansari, d. (2013). why mental arithmetic counts: brain activation during single digit arithmetic predicts high school math scores. journal of neuroscience, 33, 156-163. doi: 10.1523/jneurosci.2936-12.2013 redcay, e., dodell-feder, d., pearrow, m. j., mavros, p. l., kleiner, m., gabrieli, j. d. e., & saxe, r. (2010). live face-to-face interaction during fmri: a new tool for social cognitive neuroscience. neuroimage, 50, 1639-1647. doi:10.1016/j.neuroimage.2010.01.052 siegler, r. s. (1996). emerging minds: the process of change in children's thinking. new york, ny: oxford university press. squire, l. r., berg, d., bloom, f. e., du lac, s., ghosh, a., & spitzer, n. c. (2013). fundamental neuroscience (4th ed.). oxford, uk: academic press. stern, e., & schneider, m. (2010). a digital road map analogy of the relationship between neuroscience and educational research. zdm the international journal on mathematics education, 42, 511-514. doi 10.1007/s11858-010-0278-1 supekar, k., swigart, a. j., tenison, c., jolles, d. d., rosenberg-lee, m., fuchs, l., & menon, v. (2013). neural predictors of individual differences in response to math tutoring in primarygrade school children. proceedings of the national academy of sciences, 110, 8230-8235. doi: 10.1073/pnas.1222154110 ward, j. (2006). the student's guide to cognitive neuroscience. new york, ny: psychology press. zamarian, l., ischebeck, a., & delazer, m. (2009). neuroscience of learning arithmetic: evidence from brain imaging studies. neuroscience and biobehavioral reviews, 33, 909-925. doi: 10.1016/j.neubiorev.2009.03.005 frontline learning research 6 (2014) 67-81 issn 2295-3159 corresponding authors: eduardo cascallar, ku leuven, leuven, belgium, cascallar@msn.com and mariel musso, national research council (conicet), argentina and ku leuven, leuven, belgium, mariel.musso@hotmail.com doi: http://dx.doi.org/10.14786/flr.v2i5.135 67 | f l r modelling for understanding and for prediction/classification the power of neural networks in research eduardo cascallar ab , mariel musso acd , eva kyndt a and filip dochy a a university of leuven, belgium b assessment group international, usa / belgium c national research council (conicet)/ciipme, argentina d universidad argentina de la empresa, argentina article received 28 november 2014 / revised 18 january 2015 / accepted 18 january 2015 / available online 30 january 2015 abstract two articles, edelsbrunner and, schneider (2013), and nokelainen and silander (2014) comment on musso, kyndt, cascallar, and dochy (2013). several relevant issues are raised and some important clarifications are made in response to both commentaries. predictive systems based on artificial neural networks continue to be the focus of current research and several advances have improved the model building and the interpretation of the resulting neural network models. what is needed is the courage and open-mindedness to actually explore new paths and rigorously apply new methodologies which can perhaps, sometimes unexpectedly, provide new conceptualisations and tools for theoretical advancement and practical applied research. this is particularly true in the fields of educational science and social sciences, where the complexity of the problems to be solved requires the exploration of proven methods and new methods, the latter usually not among the common arsenal of tools of neither practitioners nor researchers in these fields. this response will enrich the understanding of the predictive systems methodology proposed by the authors and clarify the application of the procedure, as well as give a perspective on its place among other predictive approaches. keywords: artificial neural networks; response to commentaries; methodology; data modelling cascallar et al 68 | f l r research is the process of going up alleys to see if they are blind. marston bates two articles, edelsbrunner and, schneider (2013), and nokelainen and silander (2014) comment on musso, kyndt, cascallar, and dochy (2013). several relevant issues are raised and some important clarifications need to be made in response to both commentaries. this response will enrich the understanding of the predictive system methodology proposed by the authors and clarify the application of the procedure, as well as give a perspective on its place among other predictive approaches. edelsbrunner and schneider (2013) in their commentary on musso, kyndt, cascallar and dochy (2013) argue that artificial neural networks (anns) should only be used as exploratory modelling techniques, in spite of being powerful statistical modelling tools with demonstrated ability to improve outcomes of classifications and predictions over traditional statistical methods (marquez, hill, worthley, & remus, 1991). garson (1998, pp. 11-14) cites more than thirty-five articles which have shown the ability of anns to outperform traditional techniques in specific circumstances. in addition, haykin (1994, pp. 4-5) summarizes some of the main favourable properties of anns which explain their advantages over traditional methods. the reasons edelsbrunner and schneider (2013) argue for their rather strong position are centred on two main arguments: (a) that the output from anns cannot be fully translated into a meaningful set of rules because of a lack of accessibility to the input-output relationships, and (b) that there is a lack of equivalent statistical parameters in anns when compared to more traditional statistical techniques. these are the two fundamental misconceptions that will be addressed. one of the essential requirements for development and advancement in science is the willingness and vision to explore new conceptualizations and methods. in particular, as is the case in the study by musso et al. (2013), the ability to bring together data from interdisciplinary domains (e.g., decuyper, dochy, & van den bossche, 2010), and to use new methodologies for analyses that are commonly applied in other disciplines such as business, finance, and the social sciences (aldeek, 2001; detienne, detienne, & joshi, 2003; laguna & marti, 2002; neal & wurst, 2001; nguyen & cripps, 2001; white & racine, 2001, and others as stated in musso et al., 2003). the literature still shows relatively few studies applying neural networks in education and in educational assessment in particular (everson, chance, & lykins, 1994; wilson & hardgrave, 1995), although anns have been shown to improve the validity and the accuracy of the predictions and/or classifications, and also improve the predictive validity of test scores (everson et al., 1994; perkins, gupta, & tamanna, 1995; weiss & kulikowski, 1991). more recently, several studies have shown the applicability and use of this methodology in education (e.g., cascallar, boekaerts, & costigan, 2006; kyndt, musso, cascallar, & dochy, 2011; kyndt, musso, cascallar, & dochy, 2015; musso & cascallar, 2009a; musso & cascallar, 2009b; musso, kyndt, cascallar & dochy, 2012; musso et al., 2013; pinninghoff junemann, salcedo lagos, & contreras arriagada, 2007; ramaswami & bhaskaran, 2010; zambrano matamala, rojas díaz, carvajal cuello, & acuña leiva, 2011). these recent studies have used anns both for prediction/classification as well as for the understanding of the underlying variables involved in the educational outcomes studied. now it cascallar et al 69 | f l r is important to show that recent advances in ann analysis have addressed the main concerns expressed in edelsbrunner & schneider (2013). first, the concerns regarding the presumed “opacity” of ann in terms of their input-output relationships will be addressed. the authors undermine their own estimate of the value of anns as a “promising technique” by essentially arguing that it is contrary to good scientific practice for theory-building given the presumed “opaque” nature of their internal structure which makes interpretation difficult if not impossible. the often and now quite outdated argument of anns as “black boxes” (cf. benitez, castro & requena, 1997) is therefore raised once again. however, these arguments are raised ignoring the vast amount of research that has been going on in this field to overcome this initial drawback of predictive systems analyses (e.g., frey & rusch, 2013; intrator & intrator, 2001; lee, rey, mentele, & garver, 2005; tzeng & ma, 2005; yeh & cheng 2010). considering the nature and centrality of modelling in science, as was clearly presented by frigg and hartmann (2006), models can perform two different representational functions, which are not mutually exclusive as scientific models. first, they can be a representation of an aspect or selected part of the world, what they call the “target system”. in this case, what can be modelled are either phenomena or data. the second notion of modelling is the representation of a theory in that it represents its rules, laws and axioms. clearly, anns contribute to the construction of better representational models consisting of “models of data” (suppes, 1962). in particular, this contribution is based on ample research that has been crucial in making the link between anns representations and their relationship to the obtained outputs. as an anecdote, it is interesting and revealing that edelsbrunner and schneider (2013) cite the paper of benitez, et al. (1997) which presents an addition to the usual ann techniques which according to benitez et al. (1997) provide “such an interpretation of neural networks so that they will no longer be seen as black boxes” (p. 1156), which clearly contradicts the use of the article of benitez et al. (1997) as supporting the “black box” unique perception of anns. the proposed approach, in this case is based on the determination of the equality between multilayered perceptron anns, precisely the one used by musso et al. (2013), and fuzzy rule-based systems. the operator derived from this equivalency concept results in the transformation of fuzzy rules into a format which can be easily understood. thus, the knowledge generated by the ann after the learning process is finished can be more easily and clearly explained, “so that they can no longer be considered as black boxes” (benitez et al., 1997, p. 1156), while retaining all the advantages and power of the anns as very efficient computing representations as automated knowledge acquisition procedure models, and as universal approximators (ripley, 1996). in fact, west, brockett, and golden (1997) state that neural networks “are a well-defined adaptive gradient search procedure for parameter fitting in a complex nonlinear model, and not a „black box‟ at all” (p. 389). in addition, the efforts to develop better and more comprehensive visualisation techniques for the complex interactions in an ann, such as those suggested by tzeng and ma (2005) have contributed to open the “black box” and help the researcher in determining underlying dependencies between inputs and outputs of a neural network. as a consequence, they do not only facilitate the design of efficient anns, but also enable the use of anns for problem solving. it is true that cascallar et al 70 | f l r visualisation is not explanation, but they are powerful tools to guide the refinement of neural network structures for problem solving (e.g., classification tasks) using anns or other machine learning models. another significant addition to the literature which “opens the box” in ann analyses is the concept of structured neural network (snn) techniques used for modelling (lee, rey, mentele, & garver, 2005). in this approach, the actual construction of the network is based on existing contextual and theoretical knowledge to assist in the design of the ann structure of inputs. in fact, a similar approach was followed by musso et al. (2013), by populating the inputs based solely on solid theoretical constructs derived from previous cognitive, motivational, and sociodemographic research and models, avoiding blind data mining techniques (hand, mannila & smyth, 2001), and based on the factor analysis and structural equation modelling (sem) of several variables to determine their potential weight in the problem. cause-and-effect relationships have been traditionally modelled, among others, by sem and partial least squares (pls) approaches. but these procedures have their own shortcomings. in pls, there is no theoretical rationale for all indicators to have the same weighting (haenlein & kaplan, 2004), and the pls procedure does not take into account the fact that some indicators may be more reliable than others and should, therefore, receive higher weights (chin, marcolin, & newsted (2003). in addition, there is the difficulty of interpreting the loadings of the independent latent variables in pls (which are based on cross-product relations with the response variables). regarding sem several authors also point out some issues that require attention from the researcher or that are still awaiting further research (lei & qiong wu, 2007; schermelleh-engel, kerwer, & klein, 2014; weston & gore, 2006). among the issues noted with sem are possible data problems, such as missing data, non-normality of observed variables, or multicollinearity; estimation problems that could be due to data problems or identification problems in model specification; or interpretation problems due to unreasonable estimates. these potential problems have led to suggestions involving the development of “mixture pls” models (hahn, johnson, herrmann, & huber, 2002), hierarchical bayesian methods in sem models (ansari, jedidi, & jagpal, 2000) and new ways of evaluating fit in non-linear multilevel structural equation models (schermelleh-engel et al., 2014). even if nonlinear sem and pls models could handle asymmetric relationships, they still do not solve the problems associated with large data and complex interactions. the snn approach takes into account these complexities and non-linearity in data sets, while maintaining the advantages of the ann general model. another significant addition to the battery of approaches that researchers have explored to eliminate the “black box” risk of anns is the inclusion of sensitivity analysis for each of the variables in the model (kim & ahn, 2009) in order to extract the necessary information for model validation and process optimisation, from the relationships between inputs and outputs in the ann. this method, based on the relative importance (ri) parameter estimate improves on garson‟s (1991) use of relative importance weights, and uses sensitivity analysis to determine the causal importance of the input variables on the outputs. the sensitivity is a measure of the increase in the error of the predicted value as each variable is excluded from the model, and demonstrates systematically the degree of influence on the network weights of each participating variable. the ri methods used in both classification and prediction models are another evidence of the fallacy of the https://www.researchgate.net/researcher/2045138797_karin_schermelleh-engel/ https://www.researchgate.net/researcher/2045138797_karin_schermelleh-engel/ cascallar et al 71 | f l r view of neural networks as black-boxes beyond human understanding. incidentally, kim and ahn (2009) also compared the results from the ann analysis with logistic regression and classification and regression trees (cart) analyses, with ann models obtaining better results in both training and testing sets of data. other authors (e.g., blackard & dean, 1999) have compared anns absolute accuracy and relative accuracy compared to predictions based on discriminant analysis (da) models, with a consistent finding that ann models outperformed the da models. a very interesting comparison of methods to accurately assess the contribution of variables in ann architectures has been reported by olden, joy, and death (2004). the authors compare nine different methods for quantifying variable importance in anns using simulated data with known properties. the use of simulated data, when the true importance of the variables is known, provides a solid base for future developments in this field, which are not possible with natural data as is the case with gevrey, dimopoulos, and lek (2003). the nine methodologies studied by olden et al. (2004) included: connection weights, garson‟s algorithm, partial derivatives, input perturbation, sensitivity analysis, forward stepwise addition, backward stepwise elimination, improved stepwise selection 1, and improved stepwise selection 2 (see olden et al., 2004 for details on these methods). the results indicated that the connection weights approach showed the best overall performance both in terms of accuracy (degree of similarity between true and estimated variable ranks) and precision (degree of variation in accuracy), when estimating the true importance of all the variables in the ann. partial derivatives, input perturbation, sensitivity analysis and both versions of the improved stepwise selection methods showed moderate performance in the simulations. when estimating the actual ranks, the connection weights approach once again was the method which exhibited the best performance. in addition, olden and jackson (2002) reviewed a randomisation approach to better evaluate and understand the contribution of predictors in ann analysis. they conclude by stating: “thus, by coupling this new explanatory power of neural networks with its strong predictive abilities, anns promise to be a valuable quantitative tool to evaluate, understand, and predict ecological phenomena” (olden & jackson, 2002, p. 135). all of these examples demonstrate that using the appropriate techniques, the complexity of an ann does not need to translate into “opacity”, and researchers are not limited in their ability to gain insight into the explanatory factors of the prediction and classification processes performed efficiently by anns. studies such as olden et al. (2004), gevrey et al. (2003), and lek, belaud, baran, dimopoulos, and delacoste (1996), are but the beginnings of a vast number of applications that have “opened the box” in ann analysis. in addition, regularisation approaches have been used to enhance the interpretation of ann results (intrator & intrator, 2001), and the estimation of interaction effects in anns was used and demonstrated by donaldson and kamstra (1999). therefore, contrary to what has been pointed out by edelsbrunner and schneider (2013) and quoted by golino and gomes (2014), the ann approach offers the potential to examine the complex relationships amongst its components. an additional important advantage of ann analysis refers to the need to capture the complexity of the interaction of various factors in the understanding of also complex phenomena (agrawal, 2001). it is difficult to find large-n studies with a large set of variables, particularly in the social and educational sciences. so, most studies attempt to develop causal models based on a cascallar et al 72 | f l r very limited set of variables, without the capacity to encompass a large number of predictors, and therefore not providing the possibility to observe their complex interactions (boekaerts & cascallar, 2006; cascallar et al., 2006). a resulting problem is that meta-analyses trying to find general statistical correlations face very serious problems as interactions between the factors analysed are not known, which in turn leads to wrong estimations of relevance. related to this problem is the fact that in all studies that knowingly or unknowingly exclude a relevant factor, the importance of all other variables shifts dramatically. this effect has been noted in very diverse fields ranging from natural resource estimation to self-regulated learning (agrawal & chhatre, 2006; boekaerts & cascallar, 2006). studies which only take into account a few variables, in rather simple designs, and do not consider very important but complex interactions with a larger number of participating factors can and do often show contradictory results. this should not be considered a trivial problem for the conceptualisation of various effects and phenomena in every scientific field (boekaerts & cascallar, 2006). frey and rusch (2013) present an interesting study in the area of social-ecological systems which uses anns with an analytic approach that produces an open architecture in which it is possible to establish the input-output relationships which edelsbrunner and schneider (2013) seem to perceive are unachievable for anns. these analyses suggested by various authors (thrush, coco & hewitt, 2008; yeh & cheng 2010) make the relationships among the various input-output variables explicit. the second main argument regarding problems associated with the ann methodology, as claimed by edelsbrunner and schneider (2013), has to do with the lack of some statistical parameters in anns. this ignores the evidence that there has also been an abundance of research to provide the ann model with equivalent information. there have been increasing efforts for some time, to embed anns in general statistical frameworks (cheng & titterington, 1994), with bridle (1992) comparing and blending anns with markov-chain models, and applying bayesian approaches and methods in the modelling of neural networks (mackay, 1992). more recently, he and li (2011) provide an interesting example of such work. they used the standard backpropagation algorithm derived in vector form, and they were successful in determining the confidence interval and prediction intervals for the ann, while also exploring which neural network structural characteristics had more of an impact on such parameters. in particular, when the levenberg-marquardt backpropagation algorithm is used to train a neural network, since the jacobian matrix has been calculated to update the weights and biases of the neural network, the confidence interval with the corresponding confidence level can be computed to evaluate the predictive capability of the ann. in addition, on similar topics, zapranis and livanis (2005) state that given that anns are a good example of consistent non-parametric estimators with powerful universal approximation properties, they require that the development and implementation of neural network applications has to be based on established procedures for estimating confidence and especially prediction intervals. they go on to review the main state-of-the-art approaches for the construction of confidence and prediction intervals, and evaluate their strengths and weaknesses. after comparing them in a controlled simulation, the authors suggest that a combination of bootstrap and maximum likelihood approaches are superior to analytic approaches when constructing the prediction intervals (zapranis & livanis, 2005). on the other hand, other authors propose the construction of confidence intervals for neural networks based on least squares cascallar et al 73 | f l r estimations and using the linear taylor expansion of the nonlinear model output, which also detects ill-conditioning of ann candidates and can estimate their performance (rivals and personnaz, 2000). in terms of the comparison between anns and logistic regression, in neural network analysis the purpose of the hidden layer is to map a set of patterns, which are linearly non-separable in the input space, into the so-called image-space in the hidden layer, where these patterns may become linearly separable. as in logistic regression, decision surfaces in the neural networks are hyperplanes in the input space. the key difference, though, between neural networks and logistic regression is that each hidden neuron (other than the bias neuron) produces an output that corresponds to a distinct, discriminating hyperplane in the input space. when these are weighted, summed, and transformed at an output neuron, the resulting output corresponds very closely to a multidimensional step function. it is found that the boundaries of regions of similar probability are defined by the discriminating hyperplanes, which crisscross the input space (dreiseitl & ohnomachado, 2002). given the vast number of practical applications already mentioned in the original article by musso et al. (2013), it is unfortunate that edelsbrunner and schneider (2013) choose to exemplify an unrealistic example of application of anns in a contrived situation in which a student is eliminated from a programme based on a neural network classification. anns, like any other methodology provides the researcher or applied scientist with information. as we have already shown from the literature cited, in the case of anns there are a number of methods to establish the necessary input-output relationships and to determine the confidence and prediction intervals provided by an ann. therefore, the contrived diagnostic example provided by edelsbrunner and schneider (2013, pp. 100) shows an underestimation/misinterpretation of the potential of anns. furthermore, poor advice is always a problem, as would be the case in this example, with the unfortunately frequent decision-making of students‟ career paths determined by a single-point examination. on the other hand, a trusted result from a properly constructed and tested ann could provide valuable diagnostic, educational, and public policy information. in fact, the research carried out by some of these authors (cascallar et al., 2006; kyndt et al., 2011, 2015; luft, gomes, priori & takase, 2013; musso & cascallar, 2009a; musso et al., 2012, 2013) provides examples of useful diagnostic models in the educational field. it is a false dichotomy to present modelling for understanding versus modelling for prediction. in reality, both are achievable and in fact they should be integrated for the advancement of the field and the success of each application. much insight has been gained by integrating understanding with predictive and classification models. as is good practice in various fields, especially in applied statistics and mathematical modelling, the various approaches constitute a toolbox that the professional has available in order to apply the best method for the problem at hand. the fact that our article (musso et al., 2013) demonstrated the use of anns in a given academic application is not meant to be exclusionary. on the contrary, the field requires the integration of mathematical modelling and statistical techniques. regarding the comments in nokelainen and silander (2014) on the article by musso et al. (2013), they can be summarized in two main points. the first point questions whether the methodology used was rigorous in its procedures, and the second suggests comparing the neural cascallar et al 74 | f l r network results with those obtained from another discriminative classifier in addition to the comparison to a generative classifier such as discriminant analysis. it is very important to clarify that the data reported in musso et al. (2013) rigorously followed the standards established by the message understanding conferences (muc) (grishman & sundheim (1996). as is clearly stated in the musso et al. (2013) article, “the training and testing samples were selected at random from the existing data and the proportions were adjusted in order to maximize the training sample while preserving the appearance of all detected patterns in the testing sample, so as to be able to appropriately test the model” (p. 60). the two samples were chosen at random, precisely to avoid what nokelainen and silander (2014) put forward. these authors seem to have misinterpreted the sections on analyses procedures and architecture of the neural network (musso et al., 2013, pp. 52-54) in which the process is described in detail, and they completely misjudge when they state that “the paper by musso and her colleagues (2013) practically acknowledges that such a discipline was not rigorously followed.” (nokelainen & silander, 2014, p. 79). it is clearly stated in the above mentioned sections the way in which the sample was divided, the complete independence of the randomly selected training and testing subsets, and the criteria followed to determine the proportions of cases in each of the two subsets. ironically, the procedures followed coincide with those suggested by (nokelainen & silander, 2014, p. 79). let us state unequivocally that both subsets of cases in the training and testing samples were analyzed separately. in addition, all training of the neural network model was carried out on the training sample, as well as all parameter adjustments, until the desired level of precision was attained. then, the model was independently tested on the testing sample, capturing the generalization of the network structure and the learning parameters. none of the model building took place on the testing sample as nokelainen and silander (2014) incorrectly assume. thus, the performance of the model with the testing subset actually provides an indication of the generalization of the model, not just “fit” as nokelainen and silander (2014, pp. 79) also incorrectly state. a related comment regarding the “ethical standards” of the musso et al. (2013) paper is truly surprising. do nokelainen and silander (2014) truly believe or imply that the authors could not “refrain from cheating (using the test data)” (nokelainen & silander (2014, p. 79) in developing the model? if so, it is alarming, because they are making a serious assumption regarding the authors or at best an implication of ignorance of basic rules of science and of this methodology in particular. their fear of “cheating” and their implication that the testing sample analysis should be carried out by different researchers because of this assumed temptation to cheat could be extended to all research in all areas and all statistical methods. it is precisely part of the scientific method to follow any scientific finding with careful replications, not simply to avoid cheating, but to truly evaluate the generalizability of scientific results. it does not mean that we cannot trust researchers, at least a priori, with carrying out an ethically sound analysis. if not, all findings, including theirs, would be in question. certainly, the musso et al. (2013) article followed careful and rigorous methodological procedures. if their question has to do with the perfect classification obtained, it is the product both of the appropriate modelling process carried out, and of the granularity of the expected results given the available data; it should be noted that the correlation between the individual gpa scores of the cascallar et al 75 | f l r students in the whole testing sample and their predicted score (with data from one year in advance), was .86 (musso et al., 2013, p. 64). regarding the suggestion to use other discriminative classifiers, such as logistic regression, to compare with the results obtained with the neural network model, it is a good suggestion which has already been carried out in the literature (kim & ahn, 2009), and it has been found that neural networks obtained better classification results. in fact, some of the authors in musso et al. (2013) already have carried out such analyses in research currently underway, with the same results favourable to neural networks (musso, boekaerts, segers, & cascallar, in preparation). the field of machine learning research and the related predictive systems is in constant development and new advances are introduced at a rapid pace (monteith, carroll, seppi, & martinez, 2011). several methods have been suggested to improve the performance of machine learning algorithms and of neural network methods in particular, some of them using bayesian approaches which have shown excellent potential (aires, prigent, & rossow, 2004; orre, lansner, bate, & lindquist, 2000). we share the view expressed by nokelainen and silander (2014) that continued research in this field should be pursued, and ensemble methods (rokach, 2010), such as those involving bootstrap aggregating (sahu, runger, & apley, 2011), and bayesian model combination (monteith et al., 2011), together with multiple classifier systems (roli, giacinto, & vernazza, 2001) are among those that should continue to be considered in certain applications. in conclusion, we can state that as was very accurately stated by anders and korn (1996) in their work on model selection in neural networks, the process of model selection in ann can be informed by statistical procedures and methods. statistical methods can improve the model building and the interpretation of anns. what is needed is the courage and open-mindedness to actually explore new paths and new methodologies which can perhaps sometimes unexpectedly provide new conceptualisations and tools for theoretical advancement and practical applied research. this is particularly true in the fields of educational science and social sciences, where the complexity of the problems to be solved requires the exploration of proven methods and new methods, the latter usually not among the common arsenal of tools of neither practitioners nor researchers in these fields. keypoints artificial neural networks are powerful mathematical modelling tools for classification and prediction. advances in artificial neural network methodologies have made them more transparent and useful, avoiding the original “black box” characteristics in their early development. there is a long history with significant recent advances which has achieved strong ties between traditional statistical constructs with their equivalent in artificial neural networks. artificial neural networks are a useful methodology that can advance our understanding of phenomena when modelling for understanding and modelling for classification/predictions are combined. cascallar et al 76 | f l r artificial neural networks are an additional important tool in the researcher‟s toolbox which can be particularly useful to tackle highly complex and large data sets with interactions among the variables which are not fully understood. references agrawal, a. (2001). common property institutions and sustainable governance of resources. world development, 29, 1649-1672. doi: 10.1016/s0305-750x(01)00063-8 agrawal, a., & chhatre, a. (2006). explaining success on the commons: community forest governance in the indian himalaya. world development, 34, 149-166. doi: 10.1016/j.worlddev.2005.07.013 aires, f., prigent, c., & rossow, w. b. (2004). neural network uncertainty assessment using bayesian statistics: a remote sensing application. neural computing, 16, 2415-2458. doi: 10.1162/0899766041941925 al-deek, h. m. (2001). which method is better for developing freight planning models at seaports – neural networks or multiple regression? transportation research record, 1763, 9097. doi: 10.3141/176314 anders, u., & korn, o. (1996). model selection in neural networks. zew discussion papers, 96-21. retrieved from http://hdl.handle.net/10419/29449 ansari, a., jedidi, k., & jagpal, h. s. (2000). a hierarchical bayesian methodology for treating heterogeneity in structural equation models. marketing science, 19, 328-347. doi: 10.1287/mksc.19.4.328.11789 benitez, j. m., castro, j. l., & requena, i. (1997). are artificial neural networks black boxes? ieee transactions on neural networks, 8, 1156-1164. doi: 10.1109/72.623216 blackard, j. a. & dean, d. j. (1999). comparative accuracies of artificial neural networks and discriminant analysis in predicting forest cover types from cartographic variables. computers and electronics in agriculture, 24, 131–151. doi: 10.1016/s0168-1699(99)00046-0 boekaerts, m., & cascallar, e. c. (2006). how far have we moved toward the integration of theory and practice in self-regulation? educational psychology review, 18, 199-210. doi: 10.1007/s10648-0069013-4 bridle, j. s. (1992). neural networks or hidden markov models for automatic speech recognition: is there a choice? in p. laface (ed.), speech recognition and understanding: recent advances, trends and application (pp. 225-236). new york: springer. cascallar, e. c., boekaerts, m., & costigan, t. e. (2006) assessment in the evaluation of selfregulation as a process. educational psychology review, 18, 297-306. doi: 10.1007/s10648-006-9023-2 http://dx.doi.org/10.1016/s0305-750x(01)00063-8 http://dx.doi.org/10.1016/j.worlddev.2005.07.013 http://dx.doi.org/10.1162/0899766041941925 http://dx.doi.org/10.3141/1763-14 http://dx.doi.org/10.3141/1763-14 http://hdl.handle.net/10419/29449 http://dx.doi.org/10.1287/mksc.19.4.328.11789 http://dx.doi.org/10.1109/72.623216 http://dx.doi.org/10.1016/s0168-1699(99)00046-0 http://dx.doi.org/10.1007/s10648-006-9013-4 http://dx.doi.org/10.1007/s10648-006-9013-4 http://dx.doi.org/10.1007/s10648-006-9023-2 cascallar et al 77 | f l r cheng, b., & titterington, d. m. (1994). neural networks: a review from a statistical perspective. statistical science, 9, 1, 2-54. doi: 10.1214/ss/1177010638 chin, w. w., marcolin, b. l., & newsted, p. r. (2003). a partial least squares latent variable modelling approach for measuring interaction effects: results from a monte carlo simulation study and an electronic-mail emotion/adoption study. information systems research, 14, 189–217. doi: 10.1287/isre.14.2.189.16018 decuyper, s., dochy, f., & van den bossche, p. (2010). grasping the dynamic complexity of team learning: an integrative model for effective team learning in organisations. educational research review, 5, 111-133. doi: 10.1016/j.edurev.2010.02.002 detienne, k. b., detienne d. h., & joshi, s. a. (2003). neural networks as statistical tools for business researchers. organizational research methods, 6, 236-265. doi: 10.1177/1094428103251907 donaldson, r. g., & kamstra, m. (1999). neural network forecast combining with interaction effects. journal of the franklin institute, 336b, 227-236. doi: 10.1016/s0016-0032(98)00018-0 dreiseitl, s., & ohno-machado, l. (2002). logistic regression and artificial neural network classification models: a methodology review. journal of biomedical informatics, 35, 352–359. doi: 10.1016/s1532-0464(03)00034-0 edelsbrunner, p., & schneider, m. (2013). modelling for prediction vs. modelling for understanding: commentary on musso et al. (2013). frontline learning research, 2, 99-101. everson, h. t., chance, d., & lykins, s. (1994, april). exploring the use of artificial neural networks in educational research. paper presented at the annual meeting of the american educational research association, new orleans, louisiana. frey, u. j., & rusch, h. (2013). using artificial neural networks for the analysis of social-ecological systems. ecology and society, 18, 40.doi:10.5751/es-05202-180240. frigg, r. & hartmann, s. (2006). models in science. in e. n. zalta (ed.), the stanford encyclopaedia of philosophy. summer 2006 edition. stanford, ca: stanford university press. garson, g. d. (1991). interpreting neural-network connection weights. ai expert, 6, 47-51. garson, g. d. (1998). neural networks. an introductory guide for social scientists. london: sage publications ltd. gevrey, m., dimopoulos, i., & lek, s. (2003). review and comparison of methods to study the contribution of variables in artificial neural network models. ecological modelling, 160, 249-264. doi: 10.1016/s0304-3800(02)00257-0 golino, h. f., & gomes, c. m. (2014). four machine learning methods to predict academic achievement of college students: a comparison study. manuscript submitted for publication. http://dx.doi.org/10.1214/ss/1177010638 http://dx.doi.org/10.1287/isre.14.2.189.16018 http://dx.doi.org/10.1016/j.edurev.2010.02.002 http://dx.doi.org/10.1177/1094428103251907 http://dx.doi.org/10.1016/s0016-0032(98)00018-0 http://dx.doi.org/10.1016/s1532-0464(03)00034-0 http://dx.doi.org/10.1016/s1532-0464(03)00034-0 http://dx.doi.org/10.1016/s0304-3800(02)00257-0 http://dx.doi.org/10.1016/s0304-3800(02)00257-0 cascallar et al 78 | f l r grishman, r., & sundheim, b. (1996). message understanding conference 6: a brief history. in: proceedings of the 16th international conference on computational linguistics (coling), i, copenhagen, 466–471. haenlein, m., & kaplan, a. (2004). a beginner's guide to partial least squares analysis. understanding statistics, 3, 283–297. doi: 10.1207/s15328031us0304_4 hahn, c., johnson, m. d., herrmann, a., & huber, f. (2002). capturing customer heterogeneity using a finite mixture pls approach. schmalenbach business review, 54, 243269. hand, d., mannila, h., & smyth, p. (2001). principles of data mining. cambridge, ma: mit press. haykin, s. (1994). neural networks: a comprehensive foundation. new york: macmillan. he, s., & li, j. (2011). confidence intervals for neural networks and applications to modeling engineering materials. in c. l. p. hui (ed.), artificial neural networks – application. shanghai, china: intech. doi: 10.5772/16097 intrator, o., & intrator, n. (2001). interpreting neural-network results: a simulation study. computational statistics and data analysis, 37, 373–393. doi: 10.1016/s0167-9473(01)00016-0 kim, j., & ahn, h. (2009). a new perspective for neural networks: application to a marketing management problem. journal of information science and engineering, 25, 1605-1616. kyndt, e., musso, m., cascallar, e., & dochy, f. (2011, august). predicting academic performance in higher education: role of cognitive, learning and motivation. symposium conducted at the 14th earli conference, exeter, uk. kyndt, e., musso, m., cascallar, e., & dochy, f. (2015, in press). predicting academic performance: the role of cognition, motivation and learning approaches. a neural network analysis. in v. donche & s. de maeyer (eds.), methodological challenges in research on student learning. antwerp, belgium: garant. laguna, m., & marti, r. (2002). neural network prediction in a system for optimizing simulations. iie transactions, 34, 273-282. doi: 10.1080/07408170208928869 lee, c., rey, t., mentele, j., & garver, m. (2005). structured neural network techniques for modeling loyalty and profitability. proceedings of the thirtieth annual sas® users group international conference. cary, nc: sas institute inc. lei, p. w., & qiong wu, q. (2007). introduction to structural equation modelling: issues and practical considerations. items – instructional topics in educational measurement fall 2007, ncme instructional module, 33-43. lek, s., belaud, a., baran, p., dimopoulos, i., & delacoste, m. (1996). role of some environmental variables in trout abundance models using neural networks. aquat. living resour, 9, 23-29. doi: 10.1051/alr:1996004 http://dx.doi.org/10.1207/s15328031us0304_4 http://dx.doi.org/10.5772/16097 http://dx.doi.org/10.1016/s0167-9473(01)00016-0 http://dx.doi.org/10.1080/07408170208928869 http://dx.doi.org/10.1051/alr:1996004 http://dx.doi.org/10.1051/alr:1996004 cascallar et al 79 | f l r luft, c. d. b., gomes, j. s., priori, d., & takase, e. (2013). using online cognitive tasks to predict mathematics low school achievement. computers & education, 67, 219-228. doi: 10.1016/j.compedu.2013.04.001 mackay, d. j. c. (1992). a practical bayesian framework for backpropagation networks. neural computation, 4, 448472. doi: 10.1162/neco.1992.4.3.448 marquez, l., hill, t., worthley, r., & remus, w. (1991). neural network models as an alternative to regression. proceedings of the ieee 24th annual hawaii international conference on systems sciences, 4, 129-135. doi: 10.1109/hicss.1991.184052 monteith, k., carroll, j., seppi, k., & martinez, t. (2011). turning bayesian model averaging into bayesian model combination. in: proceedings of the international joint conference on neural networks (ijcnn) 2011, 2657–2663. musso, m. f., & cascallar, e. c. (2009a). new approaches for improved quality in educational assessments: using automated predictive systems in reading and mathematics. journal of problems of education in the 21st century, 17, 134-151. musso, m. f. & cascallar, e. c. (2009b).predictive systems using artificial neural networks: an introduction to concepts and applications in education and social sciences. in m. c. richaud & j. e. moreno (eds.). research in behavioural sciences (volume i), (pp. 433-459). buenos aires, argentina: ciipme/conicet. musso, m. f., kyndt, e., cascallar, e. c., & dochy, f. (2012). predicting mathematical performance: the effect of cognitive processes and self-regulation factors. education research international. vol 2012, article id 250719, 13 pages. doi: 10.1155/2012/250719 musso, m. f., kyndt, e., cascallar, e. c., & dochy, f. (2013). predicting general academic performance and identifying differential contribution of participating variables using artificial neural networks. frontline learning research, 1, 42-71. doi: 10.14786/flr.v1i1.13 musso, m. f., boekaerts, m., segers, m., & cascallar, e. c. (in preparation). a comparative analysis of the prediction of student academic performance. neal, w., & wurst, j. (2001). advances in market segmentation. marketing research, 13, 14-18. nguyen, n., & cripps, a. (2001). predicting housing value: a comparison of multiple regression and artificial neural networks. journal of real estate research, 22, 313-336. nokelainen, p. & silander, t. (2014). using new models to analyse true complex regularities of the world: commentary on musso et al. (2013). frontiers in psychology, 3, 78-82. doi: .org/10.14786/flr.v2i1.107. olden, j. d., & jackson, d. a. (2002). illuminating the ''black box'': a randomization approach for understanding variable contributions in artificial neural networks. ecological modelling, 154, 135150. doi: 10.1016/s0304-3800(02)00064-9 http://dx.doi.org/10.1016/j.compedu.2013.04.001 http://dx.doi.org/10.1016/j.compedu.2013.04.001 http://dx.doi.org/10.1162/neco.1992.4.3.448 http://dx.doi.org/10.1109/hicss.1991.184052 http://dx.doi.org/10.14786/flr.v1i1.13 http://dx.doi.org/10.1016/s0304-3800(02)00064-9 cascallar et al 80 | f l r olden, j. d., joy, m. k. & death, r. g. (2004). an accurate comparison of methods for quantifying variable importance in artificial neural networks using simulated data. ecological modelling, 178, 389-397. doi: 10.1016/j.ecolmodel.2004.03.013 orre, r., lansner, a., bate, a., & lindquist, m. (2000). bayesian neural networks with confidence estimations applied to data mining. computational statistics & data analysis, 34, 473-493. doi: 10.1016/s0167-9473(99)00114-0 perkins, k., gupta, l., & tamanna (1995). predict item difficulty in a reading comprehension test with an artificial neural network. language testing, 12, 34-53. doi: 10.1177/026553229501200103 pinninghoff junemann, m. a., salcedo lagos, p. a., & contreras arriagada, r. (2007). neural networks to predict schooling failure/success. in j. mira & j. r. alvarez (eds.), nature inspired problemsolving methods in knowledge engineering, (part ii), (pp. 571–579). berlin/heidelberg: springerverlag. doi: 10.1007/978-3-540-73055-2_59 ramaswami, m. m., & bhaskaran, r. r. (2010). a chaid based performance prediction model in educational data mining. international journal of computer science issues, 7, 10-18. roli, f., giacinto, g., & vernazza, g. (2001). methods for designing multiple classifier systems. in j. kittler & f. roli (eds.), multiple classifier systems, (pp. 78-87). berlin/heidelberg: springer-verlag. doi: 10.1007/3-540-48219-9_8 ripley, b. d. (1996). pattern recognition and neural networks. cambridge: cambridge university press. doi: 10.1017/cbo9780511812651 rivals, i., & personnaz, l. (2000). construction of confidence intervals for neural networks based on least squares estimations. neural networks, 13, 463-484. doi: 10.1016/s0893-6080(99)00080-5 rokach, l. (2010). ensemble-based classifiers. artificial intelligence review, 33, 1-39. doi: 10.1007/s10462-009-9124-7 sahu, a., runger, g., apley, d. (2011). image denoising with a multi-phase kernel principal component approach and an ensemble version. ieee applied imagery pattern recognition workshop, 1-7. schermelleh-engel, k., kerwer, m., & klein, a. g. (2014). evaluation of model fit in nonlinear multilevel structural equation modelling. frontiers in psychology, 5, article 181, 1-11. doi: 10.3389/fpsyg.2014.00181. suppes, p. (1962). models of data. in e. nagel, p. suppes & a. tarski (eds.), logic, methodology and philosophy of science: proceedings of the 1960 international congress. stanford: stanford university press, 252-261. thrush, s. f., coco, g., & hewitt, j. e. (2008). complex positive connections between functional groups are revealed by neural network analysis of ecological time series. american naturalist 171, 669-677. doi: 10.1086/587069 http://dx.doi.org/10.1016/j.ecolmodel.2004.03.013 http://dx.doi.org/10.1016/s0167-9473(99)00114-0 http://dx.doi.org/10.1016/s0167-9473(99)00114-0 http://dx.doi.org/10.1177/026553229501200103 http://dx.doi.org/10.1007/978-3-540-73055-2_59 http://dx.doi.org/10.1007/3-540-48219-9_8 http://dx.doi.org/10.1007/3-540-48219-9_8 http://dx.doi.org/10.1017/cbo9780511812651 http://dx.doi.org/10.1016/s0893-6080(99)00080-5 http://dx.doi.org/10.1007/s10462-009-9124-7 http://dx.doi.org/10.1007/s10462-009-9124-7 http://dx.doi.org/10.1086/587069 cascallar et al 81 | f l r tzeng, f. y., & ma, k. l. (2005). intelligent feature extraction and tracking for visualizing large-scale 4d flow simulations. in dvd proceedings of the international conference for high performance computing, networking, storage and analysis (sc '05). november, 2005. weiss, s. m., & kulikowski, c. a. (1991). computer systems that learn. san mateo, ca: morgan kaufmann publishers. west, p. m., brockett, p. l., & golden, l. l. (1997). a comparative analysis of neural networks and statistical methods for predicting consumer choice. marketing science, 16, 370-391. doi: 10.1287/mksc.16.4.370 weston, r., & gore, p. a. (2006). a brief guide to structural equation modeling. the counseling psychologist, 34, 719-751. doi: 10.1177/0011000006286345 white, h., & racine, j. (2001). statistical inference, the bootstrap, and neural network modelling with application to foreign exchange rates. ieee transactions on neural networks, 12, 657-673. doi: 10.1109/72.935080 wilson, r. l., & hardgrave, b. c. (1995). predicting graduate student success in an mba program: regression versus classification. educational and psychological measurement, 55, 186-195. doi: 10.1177/0013164495055002003 yeh, i. c., & cheng, w. l. (2010). first and second order sensitivity analysis of mlp. neurocomputing, 73, 2225-2233. doi: 10.1016/j.neucom.2010.01.011 zambrano matamala, c., rojas díaz, d., carvajal cuello, k., & acu-a leiva, g. (2011). análisis de rendimiento académico estudiantil usando data warehouse y redes neuronales. [analysis of students' academic performance using data warehouse and neural networks] ingeniare. revista chilena de ingeniería, 19, 369-381. doi: 10.4067/s0718-33052011000300007 zapranis, a., & livanis, e. (2005). prediction intervals for neural network models. proceedings of the 9th wseas international conference on computers (iccomp'05). world scientific and engineering academy and society (wseas). stevens point, wisconsin, usa. http://dx.doi.org/10.1287/mksc.16.4.370 http://dx.doi.org/10.1287/mksc.16.4.370 http://dx.doi.org/10.1177/0011000006286345 http://dx.doi.org/10.1109/72.935080 http://dx.doi.org/10.1109/72.935080 http://dx.doi.org/10.1177/0013164495055002003 http://dx.doi.org/10.1177/0013164495055002003 http://dx.doi.org/10.1016/j.neucom.2010.01.011 http://dx.doi.org/10.4067/s0718-33052011000300007 microsoft word von der linden et al_publication.docx           frontline  learning  research  vol.3  no.  4  (2015)  37-­‐  55   issn  2295-­‐3159     effects of a short strategy training on metacognitive monitoring across the life-span nicole von der linden1, elisabeth löffler, wolfgang schneider university of würzburg, germany article received 29 july / revised 20 november / accepted 21 november / available online 18 january abstract the present study was conducted to explore the potential positive influence of a short strategy training on metacognitive monitoring competencies covering a life-span approach. participants of four age groups (3rd-grade children, adolescents, younger and older adults) concluded a paired-associate learning task. additionally, they gave delayed judgments-of-learning (jols), that is, they rated their certainty that they would later be able to recall specific details correctly, and confidence judgements (cjs), that is, they rated their certainty that the provided answers in the recall test were correct. half of the participants underwent a short strategy training in order to enhance their recollection of contextual details thus providing a diagnostic basis for forming metacognitive judgements. results revealed significant gains in memory performance after completing the strategy training. moreover, a positive effect of the strategy training on jols and cjs differentiation and accuracy could be detected. effects were most pronounced for children and older adults. participants who had completed the strategy training also reported a decrease of familiarity-based metacognitive judgments and were able to identify memories for which no reliable cues existed more easily than participants in the control condition. accordingly, improvements in monitoring performance seemed to be due to a shift in underlying cues. in sum, this study integrates traditional aims from the relatively separately existing lines of metacognitive research in the developmental and cognitive literature and adds to understanding and improving monitoring judgments in a lifetime sample. keywords: metamemory, judgments-of-learning, confidence judgments, monitoring, life-span, strategy instruction                                                                                                                           1  corresponding author: nicole von der linden, department of psychology, university of würzburg, röntgenring 10, 97070 würzburg, germany. phone: +49(0)931 / 31 89067, fax: +49(0)931 / 31 2763, email: linden@psychologie.uni-wuerzburg.de doi: http://dx.doi.org/10.14786/flr.v3i4.196   von  der  linden  et  al         | f l r     38   1. introduction accurate metacognitive monitoring plays an important role in many everyday situations as well as in learning contexts (schneider, 2010; son & metcalfe, 2000). in daily routines, metacognitive monitoring is for instance relevant when one has to decide about whether one has memorized the departure time of one’s train, or whether one has taken appropriate notes of a lecture. moreover, subjective monitoring judgments influence learning behaviors, especially the selection of to-be-studied items and the allocation of study time (see son & metcalfe, 2000, for a review). structured learning situations are not only important to children and young adults but also for life-long learning which has gained importance in recent years. yet, life-span perspective is still rare in the metacognitive monitoring literature. metacognitive research has traditionally been conducted within two main but separate lines of research: a developmental perspective (flavell, 1999) which focuses on the changes of metacognitive abilities during life and a cognitive tradition (for an overview cf. dunlosky & metcalfe, 2009; koriat & levy-sardot, 1999) which tries to explore mechanisms underlying metacognitive monitoring processes and its consequences for regulation of learning typically in an adult sample only. in the following, we present a study which is one of few existing attempts to combine both perspectives concerning metacognitive monitoring processes. firstly, the present study aimed at exploring developmental trajectories of monitoring abilities in a life-span sample from early school age to older adulthood. a second purpose was to improve participants’ monitoring competencies by reinforcing highly diagnostic cues in paired-associate learning situations (mccabe & soderstrom, 2011; robinson, hertzog, & dunlosky, 2006) and to investigate the role of familiarity and recollection-based cues for monitoring processes. in learning situations two aspects of monitoring are of special interest: judgments of learning (jol) and confidence judgments (cj). according to nelson and narens’ (1990) seminal model of procedural metamemory, jols provide subjective information about the degree to which encoded information has been mastered and can be potentially recalled during a future memory test (nelson & narens, 1990). findings from studies using different age groups suggest that even young children can effectively monitor their learning progress under certain circumstances. on the one hand, the results indicate that immediate jols are typically inaccurate and also represent overestimations of one’s actual performance. remarkably, this is true not only for children of different ages but also for adults. immediately after studying new information, judgments about its future recall seem severely biased by the false belief that information currently in shortterm memory can be easily recalled some minutes later. obviously, this bias operates similarly in participants of different ages. on the other hand, however, even young children can make rather accurate assessments of the subsequent recallability of items when this judgment is somewhat delayed, that is, when it takes place a minute or two after studying the item. in other words, even young children seem to have a good feeling for which items will be recallable and which will not when long-term memory information has to be accessed for the jol (schneider, 2015). confidence judgments (cjs) concern retrieval monitoring and are typically made after a response is given to indicate how sure participants are about the correctness of an answer. cjs are thought to reflect a substantive sense of certainty that arises from the strength of the memory that is being retrieved, and this sense of certainty has been interpreted as an indicator of memory accuracy (ghetti, lyons, lazzarin, & cornoldi, 2008; roebers, 2002). metacognitive monitoring judgments are commonly believed to be based on multiple cues. this accessibility view has been proposed for immediate and delayed jols (koriat, 1997; metcalfe & finn, 2008; toth, daniels, & solinger, 2011) as well as for cjs (kelley & jacoby, 1996; kelley & sahakyan, 2003). the accuracy of those judgments depends on whether accessible cues are diagnostic of memory performance or not (dunlosky & metcalfe, 2009). cues are highly diagnostic if they influence metacognitive judgments and recall performance in a similar way. thus far multiple sources involved in the construction of monitoring judgments have been postulated. for example, research investigating immediate jols emphasizes encoding fluency as a major base, which in turn depends on different factors such as the concreteness and the frequency of the items (begg, duft, lalonde, melnick, & sanvito, 1989) as well as familiarity of items von  der  linden  et  al         | f l r     39   (nelson & narens, 1990). delayed jols have been linked to retrieval fluency (benjamin & bjork, 1996) as well as success and ease of item retrieval (nelson, narens, & dunlosky, 2004). similar to jols, for cjs a number of different cues have been discussed to influence accuracy, among them perceived ease (zakay & tuvia, 1998) and vividness of retrieval (robinson, johnson, & robertson, 2000). different attempts have been made to categorize these various types of cues (koriat, 1997; kelley & jacoby, 1996). in recent years the literature has begun to discuss the distinction between familiarity-based and recollection-based cues for different metacognitive judgments (daniels, toth, & hertzog, 2009; mccabe & soderstrom, 2011; metcalfe & finn, 2008; toth et al., 2011). recollection is typically defined as the consciously controlled intentional use of memory that allows for the retrieval of qualitative details of a past event. this process is frequently associated with the subjective experience of vivid remembering. familiarity, by contrast, usually refers to experiences of prior events that may arise from activated semantic representations. the relative contribution of recollection and familiarity may differ from task to task. memory tasks that require participants to recall or recognize details about the target item rely heavily on recollection. so far, the literature lacks a systematic examination of the role of recollection and familiarity cues across different monitoring indicators and across the life-span. evidence suggests that both immediate (daniels et al., 2009; toth et al., 2011) and delayed jols (metcalfe & finn, 2008) are influenced by recollection and familiarity processes. yet, delayed jols are mainly based on cues related to recollection such as target retrievability (for an overview see metcalfe & finn, 2008), whereas familiarity processes, such as processing fluency, have been identified as primary cues for immediate jols in younger adults (matvey, dunlosky, & guttentag, 2001; rhodes & castel, 2008). to our knowledge, the role of familiarity and recollection processes underlying jols in children has not been examined to date. concerning older adults, first evidence suggests that they have more problems with monitoring recollection processes than younger adults (daniels et al., 2009). several findings support the idea that recollection processes increase the accuracy of monitoring processes. first, as noted above, delayed jols have been shown to be more accurate than immediate jols for children, younger, and older adults (connor, dunlosky, & hertzog, 1997; koriat & shitzer-reichert, 2002; nelson & dunlosky, 1991; schneider,visé, lockl, & nelson, 2000). this can be explained by the fact that for delayed jols participants actively assess long-term memory (recollection processes), which is more predictive of recall than short-term memory used for immediate jols. additionally, with delayed jols participants seem to rely more on idiosyncratic cues of encoding and remembering than with immediate jols (koriat, 1997). idiosyncratic cues refer to personal, item-specific details, for instance, images or associations. providing a rich basis of idiosyncratic cues (e.g. through a strategy training) should facilitate the identification of recollection-based memories. further evidence for the importance of recollection processes for accurate monitoring in younger adults is provided by the following fact: focusing participants’ attention to cues connected with target retrievability enhanced jol accuracy compared to immediate jols (mccabe & soderstrom, 2011). furthermore, for recollection-based memories higher (daniels et al., 2009; toth et al., 2011) and more accurate immediate jols (toth et al., 2011) could be found than for familiarity-based memories. in sum, more research on the role of familiarity and recollection processes for jols is needed, not only for children but also in terms of comparisons of broader age ranges. yet, existing evidence points to the fact that the retrieval of contextual information seems to provide a reliable cue for later memory performance in different age groups. similarly, cjs seem to be based on recollection (analytic) or familiarity (non-analytic) components (kelley & jacoby, 1996; kelley & sahakyan, 2003). recollection processes have been identified as playing a major role for cj accuracy (kelley & sahakyan, 2003) and accuracy losses with increasing old age have been linked to recollection impairments (kelley & sahakyan, 2003; wong, cramer, & gallo, 2012). for children the role of recollection and familiarity processes underlying cjs has not been addressed yet. remember-know-judgments which are positively correlated with cjs (holmes & weaver, 2010) also emphasize the role of recollection and familiarity processes for metacognitive judgments. remember-knowjudgments distinguish whether a memory content is associated with specific contextual information von  der  linden  et  al         | f l r     40   (recollection) or is familiarity-based, with participants unable to retrieve the personal encounter with a memory detail. although there is consensus that participants base their monitoring judgments on various cues, they do not always take into consideration the factors which are most predictive of memory performance (koriat, 1997; touron, hertzog, & speagle, 2010). as discussed above, recollection-based cues are considered to be important and also highly diagnostic for both jols and cjs. therefore, a training program that aims at strengthening the accessibility of those cues should enhance monitoring accuracy (mccabe & soderstrom, 2011). yet, to our knowledge no systematic training in this area has been carried out. a training should be beneficial for all age groups but especially for older adults, among whom deficits in monitoring processes have been linked to problems with recollection processes for jols (toth et al., 2011), cjs (shing, werklebergner, li, & lindenberger, 2009; wong et al., 2012) and feeling-of-knowing judgments (souchay, bacon, & danion, 2006). children should also particularly profit from such a training procedure as developmental progression in monitoring skills seems to be influenced by improved retrieval processes for both jols (koriat & shitzer-reichert, 2002) and cjs (roderer & roebers, 2011). the study presented here was designed to fill this gap in the literature and to explore the effect of recollection and familiarity based cues on metacognitive monitoring. although the role of recollection and familiarity-based processes for metacognitive processes has received more attention in recent years, available studies have focused on one type of metacognitive judgment only and involve one or at most two age groups (daniels et al., 2009; souchay et al., 2006). especially empirical studies with children are scarce. therefore the design included a life-span perspective to account for the life-long significance of learning. moreover, two different types of metacognitive judgments (jols and cjs) were included in the present study in order to allow for direct comparison within one sample. multiple but partly different cues seem to underlie delayed jols and cjs (koriat, 1997, 2012) for they occur in different stages of the learning process. compared to jols, strengthening recollection processes may have a somewhat greater effect on cjs as the longer interval to the learning stage might otherwise foster the reliance on familiarity, especially in older adults (shing et al., 2009). consequently, it is of interest to compare different monitoring indicators yet such studies are very rare in the literature (leonesio & nelson, 1990). specifically, participants of four age groups (early school age to later adulthood) were asked to complete a paired-associate learning task and to give delayed jols (which are more accurate than immediate jols across all of the different age groups; see above) and cjs. a paired-associate task was chosen in order to ensure comparability with related studies and because for this stimulus material ample evidence exists for the effectivity of strategy trainings across all included age groups (see below). half of the participants underwent a strategy training in order to enhance their recollection of contextual details during jol, cj and test collection, thus providing a diagnostic basis for forming metacognitive judgments. to ensure that the training was transferable to rehearsing processes in everyday life and for different age groups a short instruction in mental imagery was chosen. in paired-associate learning, mental imagery has proven to be the most efficient way of processing (richardson, 1998), and it effectively improves recall performance in different age groups from first grade on to older age (richardson, 1998; verhaeghen, marcoen, & goossens, 1992; willoughby, porter, belsito, & yearsley, 1999). even more relevant to our study, some recent studies provide first evidence for the fact that strategy use successfully improves monitoring processes for both jols and cjs although this research does not specifically explore the cues underlying monitoring judgments and hardly ever an explicit strategy training was done. hertzog, sinclair, and dunlosky (2010) have shown that spontaneous strategy use (e.g. mental imagery), which was not induced but only accessed after jol collection, substantially influenced jols and jol resolution, that is, the accuracy with which a person can monitor the relative recallability of different items, in adults aged 18 to 81. robinson et al. (2006) instructed but not trained younger and older adults to use a mental imagery strategy when memorizing pairs of items. they found that the size of the jols and recall performance were positively correlated with strategy use in both age groups and mental imagery was identified as a diagnostic cue for jols. to our knowledge no comparable studies exist for children. thus in von  der  linden  et  al         | f l r     41   jols, so far no attempts have been made to directly train subjects of a broad age range to apply an imagery strategy. concerning cjs, nietfeld and schraw (2002) trained college students to use various strategies for probability tests. as a result, subjects benefitted from the instructions both in terms of performance and monitoring accuracy (cjs). besides shing et al. (2009) showed that participants from 10 to 75 years of age benefitted from strategy training in terms of their cjs by enlarging the difference in cjs provided after hits compared to false alarms. in accordance with the literature we expected a positive effect of strategy training on both recall processes and metacognitive processes (jols and cjs) in all age groups. as our study is the first to include a life-span approach to investigate the influence of recollection processes on metacognitive monitoring, developmental effects were of special interest. we proposed that children and older adults would benefit most from the strategy training: production deficits concerning strategy use are most pronounced in children and older adults (naveh-benjamin, brav, & levy, 2007; pressley & levin, 1977), and recall performance increases during childhood declines in older adulthood (weinert & schneider, 1996). although generally little developmental progression is found for jols, recent evidence suggests that under certain circumstances deficits in recollection processes may play a role in lower jol accuracy in older adults (daniels et al., 2009; toth et al., 2011). as for cjs, their accuracy has been shown to improve over the primary school years (roebers, von der linden, howie, & schneider, 2007), and they seem to be influenced by retrieval processes (roderer & roebers, 2010). deficits in older adults’ cjs have also been linked to deteriorated recollection processes (kelley & sahakyan, 2003). additionally, we aimed to compare the effects of our training on different monitoring indicators as the importance of recollection processes might vary in different stages of the learning process. since we proposed that a strategy training should be effective mainly due to enhanced accessibility of recollection-based cues, we additionally asked participants to classify the basis of their recall as recollection, familiarity or no memory (rfn-judgments). these classifications were successfully introduced for jols by daniels et al. (2009) and toth et al. (2011). we extended the use of rfn-judgments to cjs. 2. method 2.1 sample a total of 160 (85 male, 75 female) participants of four age groups (40 children in 3rd grade, 40 adolescents in 7th and 8th grade, 40 younger adults between 19 and 26 years of age and 40 older adults between 60 and 75 years) took part in our study. this sample size surpasses the required number of n = 132 participants as determined by an a-priori power-analysis which was conducted with the premise to detect medium-sized effects according to cohen (d = .25). they were recruited via contacting their schools directly and via newspaper and internet advertisements. children and adolescents received small gifts, whereas the other participants got 10-euro vouchers or were paid in cash. subjects’ mean ages were 8.38 (sd = 0.49) for the children, 12.73 (sd = 0.72) for the adolescents, 22.75 (sd = 2.02) for the younger adults and 68.40 (sd = 4.08) for the older adults. 2.2 materials the learning items consisted of pairs of concrete german nouns from different semantic categories (e.g. zoo animals, furniture, clothing etc.). to vary the difficulty, one half of the pairs represented two words von  der  linden  et  al         | f l r     42   from the same category and the other half of the pairs comprised words from two different categories. the item list for children and older adults consisted of 45 word pairs in the study phase and 60 word pairs in the recognition phase (the latter including 30 pairs identical to the study phase, 15 newly matched pairs and 15 completely new pairs). item pairs for adolescents included 54 word pairs in the study phase and 72 items in the recognition phase. the numbers for younger adults were 60 and 80 item pairs respectively. four practice pairs not included in the analysis preceded each single phase of the experiment. the appearance of word pairs as identical or recombined was counterbalanced among the subjects. the order of presentation was randomized as well. 2.3 procedure the consent of the parents and of the school was obtained for the children and adolescents before the beginning of the study. participants were tested individually in quiet rooms in the school or in the laboratory. half of the subjects of each age group were randomly assigned to the strategy-instruction condition (experimental group). they received instructions on visual imagery. the test administrator first explained the advantages of memorizing word pairs as one interconnected image emphasizing the importance of integration of the two images. the explanation was facilitated by two drawings, one of a frog carrying a banana and one of a candle burning a letter. then the subjects in the experimental condition had to practice this visual imagery strategy by means of ten word pairs which were different from those used in the experiment. they were given feedback on the quality of their imagery and were asked to imagine another combined image if necessary. the instruction was standardized for all participants with the restriction that the wording in children’s was slightly simplified. participants of all age groups reported that they easily understood the strategy. the participants in the experimental group were instructed to use the visual imagery strategy while memorizing the items. in the control group no strategy instruction was given. the word pairs were presented on a computer screen with presentation rates of 8 seconds per item pair for children and older adults, 6 seconds for adolescents, and 2.5 seconds for younger adults. the presentation rates were adapted in order to control for baseline difficulty between the age groups. subjects were instructed to concentrate on the pairs because they later would have to recognize them and to indicate whether the word pair had appeared in the study phase or not. in the jol phase, each left noun of the item pair (stimulus) was presented on the screen in the same order as in the learning phase. to avoid relearning only the stimulus was shown. subjects were asked to indicate the likelihood of recognizing the word pair in about 30 minutes. jols were rated on a thermometer scale from 0 (very unsure) to 100 (very sure) successfully used in previous studies (koriat, ackerman, lockl, & schneider, 2009; koriat & shitzer-reichert, 2002). in the recognition phase, word pairs of each type (i.e., either identical, recombined, or new) were presented. participants had to indicate by opting for “yes” or “no” whether they thought that the item had appeared in exactly this combination in the studying phase. after that, the item disappeared from the screen in order to avoid distraction from the presented item pair, and participants were asked to indicate how they generated the yes-or-no-decision. they had to decide between three options: a) “i can remember the word pair very well” (recollection), b) “the word pair seems familiar to me” (familiarity), c) “i cannot remember the word pair at all” (no memory). finally, subjects had to indicate on a hot-cold-scale equivalent to that used for jols how sure they were that the given answer was correct (cj). at the end of the session, the test administrator asked whether the subjects in the experimental condition had applied the strategy while memorizing the items. participants in the control group were asked whether they had employed any strategy, and if so, to specify the strategies and to indicate how often they were used. von  der  linden  et  al         | f l r     43   3. results a preliminary analysis assessing the effect of gender did not reveal any systematic differences between male and female participants. thus data were collapsed across this variable. scheffé tests were used as a post-hoc follow-up on main effects. the level of significance was set to p < 0.05. in a first step of analysis, we assessed memory performance in terms of the percentage of correctly recognized items, that is, either identical items correctly recognized as “old” or recombined or completely new items correctly classified as “new” items. next, we analyzed jols and cjs as indicators of metacognitive monitoring. finally, we will report changes in the rfn judgments as cues for monitoring processes. results are reported as a function of age group and experimental condition in order to examine the influence of cognitive development and strategy instruction on recognition rates and metacognitive monitoring. 3.1 recognition rates the first column of table 1 shows the mean proportion of correctly recognized items as a function of age group and experimental condition, that is overall recognition rates. an anova with age group and experimental condition as between-subject factors revealed a main effect of age group (f(3,152) = 5.34; p < .01; η2 = .10). a post-hoc analysis indicated that younger adults performed significantly better than children (.80 vs. .71 correct, respectively). furthermore, a main effect of the experimental condition was found (f(1,152) = 21.82; p < .001; η2 = .13), indicating that those participants who had received the strategy instruction recognized significantly more word pairs correctly than participants who had not received such an instruction (.80 vs. .72, respectively). the second to fourth column of table 1 splits the recognition rates into percentages depending on the type of word pair: that is, whether the word pair in the recognition phase was identical to that in the study phase or whether it was recombined or a completely new word pair. inferential statistics were conducted separately for each word pair in order to facilitate the interpretation. for the identical word pairs, an anova with age group and experimental condition as between-subject factor revealed a significant main effect of age group (f(3,152) = 5.14; p < .01; η2 = .09). a post-hoc analysis showed that children (.65) recognized fewer of the identical items correctly than younger (.77) and older adults (.76). furthermore, the main effect of the experimental condition reached significance (f(1,152) = 5.24; p < .05; η2 = .03) with subjects in the experimental group (.74) recognizing more items correctly than subjects in the control group (.69). for the recombined word pairs, only the main effect of strategy instruction reached the significance level (f(1,152) = 5.24; p < .001; η2 = .11), with subjects in the strategy instruction group recognizing more of the recombined word pairs correctly (.75) than subjects who had not received the strategy instruction (.61). as for the new word pairs, a significant main effect of strategy instruction was found as well (f(1,152) = 7.04; p < .01; η2 = .04). those participants who had received the strategy instruction recognized more of the new word pairs correctly (.94) than those who had not been thus instructed (.88). von  der  linden  et  al         | f l r     44   table 1 recognition rates as a function of age group, experimental condition, and type of word pair age group type of word pair overall identic word pairs recombined word pairs new word pairs children control group experimental group .65 (.06) .77 (.11) .60 (.16) .71 (.18) .54 (.19) .72 (20) .88 (.13) .91 (.10) adolescents control group experimental group .70 (.11) .81 (.09) .64 (.17) .76 (.12) .63 (.20) .76 (.15) .89 (.11) .94 (.09) younger adults control group experimental group .79 (.09) .81 (.12) .75 (.10) .79 (.15) .73 (.19) .75 (.23) .92 (.10) .93 (.10) older adults control group experimental group .74 (.12) .80 (.11) .79 (.14) .73 (.18) .53 (.28) .77 (.22) .85 (.21) .95 (.08) standard deviations are in parentheses. 3.2 metacognitive monitoring 3.2.1 mean jols before correct vs. incorrect responses figure 1 shows participants’ mean jol ratings as a function of the correctness of the subsequent response, age group, and experimental condition. an anova with correctness of response as within-subject factor and age group and experimental condition as between-subject factors revealed a significant main effect of correctness of response (f(1,151) = 135.38; p < .001; η2 = .47): subjects gave higher jols before correct (57.98) than before incorrect responses (46.78). in addition, a significant interaction between the factors correctness of response and experimental condition was found (f(1,151) = 6.22; p < .05; η2 = .04). furthermore, the triple interaction between correctness of response, experimental condition, and age group attained a significant level (f(3,151) = 3.43; p < .05; η2 = .06). in order to examine the direction of the interactions post hoc, we analyzed the experimental and the control group data separately. for subjects in the experimental condition, an anova with correctness of response as within-subject factor and age group as between-subject factor revealed a main effect of correctness of response (f(1,75) = 89.07; p < .001; η2 = .54) with mean jols being higher before correct (59.58) than before incorrect responses (45.98). for the participants in the control condition the main effect of correctness of response was also significant (f(1,76) = 47.36; p < .001; η2 = .38). furthermore, for the subjects in the control condition a significant interaction between correctness of response and age group was found (f(3,76) = 5.93; p < .01; η2 = .19). subsequent analyses revealed that only the adolescents and the younger adults distinguished between correct and incorrect responses given that it was only for these two age groups that the factor correctness of response turned out to be significant (children: f(1,19) = 0.66; p = .426; η2 = .03; adolescents: f(1,19) = 19.82; p < .001; η2 = .51; younger adults: f(1,19) = 66.06; p < .001; η2 = .78; older adults: f(1,19) = 4.43; p = .05; η2 = .19). von  der  linden  et  al         | f l r     45   figure 1. mean jols preceding correct vs. incorrect answers as a function of age group and experimental condition. 3.2.2 jol accuracy in order to assess jol accuracy as a function of age group, question format and experimental condition, goodman-kruskal gamma correlations between jols and recall performance were computed for each participant, and then averaged for each single cell in the experimental design. gamma correlations are considered to be the most appropriate measure of metacognitive accuracy (nelson, 1984) and are commonly used in the contemporary literature (nelson & dunlosky, 1991; schneider et al., 2000). a positive correlation indicates that higher jols were given for items that were recalled correctly than for those recalled incorrectly. table 2 shows mean gamma correlations for jols as a function of age group and experimental condition. one-tailed t-tests revealed that all gamma correlations were different from zero for almost all 30 40 50 60 70 children adolescents younger adults older adults control condition correct answers incorrect answers von  der  linden  et  al         | f l r     46   groups. the only exception concerned the children in the control group whose mean gamma correlations for jols were not significantly different from zero. table 2 mean gamma correlations as a function of age group and experimental condition age group jols cjs children control group experimental group .03 (.24) .32 (.28) .18 (.20) .38 (.29) adolescents control group experimental group .20 (.19) .30 (.19) .33 (.21) .42 (.21) younger adults control group experimental group .32 (.15) .27 (.23) .42 (.20) .35 (.23) older adults control group experimental group .14 (.26) .29 (.25) .38 (.26) .49 (.26) standard deviations are in parentheses. an anova with age group and experimental condition as between-subject factors revealed a significant main effect of strategy instruction (f(1,151) = 11.36; p < .001; η2 = .07). in addition, the anova showed a significant interaction between age group and experimental condition (f(3,151) = 3.58; p < .05; η2 = .07). in order to examine the direction of this effect, the mean gamma correlations of each age group were tested individually using univariate anovas. in children, the main effect of experimental condition was significant (f(1,38) = 12.25; p < .01; η2 = .24) with children in strategy-instruction group having higher gamma correlations (.32) than those in the control group (.03). in older adults, the results pointed into the same direction (experimental group: 29; control group: .14): the main effect of experimental condition was just short of being significant (f(1,38) = 3.29; p = .077; η2 = .08). 3.2.3 mean cjs after correct vs. incorrect responses differentiation in cjs was analyzed in the same way as for jols (cf. figure 2): mean cj ratings after correct and incorrect responses respectively were calculated for each subject. as for jols, an anova with correctness of response as within-subject factor, and age group and experimental condition as betweensubject factors was conducted. the anova revealed a main effect of age group (f(3,151) = 6.26; p < .001; η2 = .11). post-hoc tests according to scheffé’s procedure showed that younger adults gave significantly lower mean cjs (65.56) than the other age groups (children: 76.10, adolescents: 74.59, older adults: 75.00), regardless whether the answer was correct or not. in addition, a main effect of correctness of response was found (f(3,151) = 188.75; p < .001; η2 = .56). subjects of all age groups gave higher ratings after correct (79.02) than after incorrect responses (66.60). finally, the interaction between correctness of response and age group reached the significance level (f(3,151) = 5.61; p < .001; η2 = .10). to further explore the direction of the effect, separate anovas for cjs after correct and incorrect answers were conducted. for correct answers, the main effect of age group reached significance (f(3,156) = 3.42; p < .05; η2 = .06). subsequent post-hoc analyses showed that older adults (82.57) were more confident after correct responses than younger adults (74.25). for the cjs after incorrect answers, a significant main effect of age group was found as well (f(3,155) = 7.89; p < .001; η2 = .13). here, post-hoc analyses showed that younger adults gave lower cjs (57.33) after incorrect responses than the other three age groups (children: 72.63; adolescents: 69.00; older adults: 67.42). von  der  linden  et  al         | f l r     47   figure 2. mean cjs after correct vs. incorrect answers as a function of age group and experimental condition 3.2.4 cj accuracy cj accuracy was assessed in the same way as jol accuracy. mean gamma correlations are displayed in table 2. all gamma correlations were different from zero (using one-tailed t-tests). an anova with age group and experimental condition as between-subject factors revealed a main effect of age group (f(3,151) = 2.93; p < .05; η2 = .06). post-hoc tests according to scheffé showed a significant difference between children (.28) and older adults (.44). furthermore, a main effect of strategy instruction was found (f(1,151) = 4.75; p < .05; η2 = .03). subjects in the strategy instruction condition had higher mean gamma correlations (.41) than participants without such instruction (.33). 3.3 rfn-judgments in addition, the rfn-judgments subjects made were contrasted. we compared the percentage of how often each option was picked out because age groups differed in quantity of items. table 3 shows the percentage of each rfn-judgment as a function of age group and experimental condition. an anova was 40 50 60 70 80 90 children adolescents younger adults older adults experimental condition correct answers incorrect answers 40 50 60 70 80 90 children adolescents younger adults older adults control condition correct answers incorrect answers von  der  linden  et  al         | f l r     48   conducted with rfn-judgment as within-subject factor and age group and experimental condition as between-subject factor. we found a significant main effect of rfn-judgment (f(1,152) = 29.99; p < .001; η2 = .17). paired contrasts revealed that subjects chose “familiarity” (.22) less often than “recollection” (.40) and “no memory” (.38). furthermore a significant interaction between rfn-judgment and experimental condition was found (f(1,152) = 4.12; p < .05; η2 = .03). separate anovas for each judgment showed that the strategy instruction had a significant effect only for the answer “familiarity” (f(1,158) = 12.63; p < .01; η2 = .07), and “no memory” (f(1,158) = 5.54; p < .05; η2 = .03): subjects in the experimental group chose the “familiarity” option less often (.19) than subjects in the control condition (.26), and the “no memory” option more often (.41) than the control group (.34). table 3 rfn-judgment in percentage of chosen option as a function of age group and experimental condition age group recollection familiarity no memory children control group experimental group .41 (.24) .35 (.23) .27 (.15) .19 (.11) .32 (.25) .45 (.21) adolescents control group experimental group .41 (.19) .42 (.18) .30 (.16) .20 (.11) .29 (.18) .38 (.20) younger adults control group experimental group .35 (.16) .41 (.22) .24 (.10) .22 (.11) .40 (.17) .38 (.17) older adults control group experimental group .42 (.21) .40 (.20) .22 (.11) .15 (.09) .36 (.16) .45 (.22) standard deviations are in parentheses. 3.4 spontaneous and instructed strategy use in a last step of analysis, we assessed the outcomes for the strategy use questionnaire. table 4 shows how many participants of each age group and experimental condition reported having used the visual imagery strategy. participants’ open responses were categorized as visual imagery by two independent raters (kappa = .95). we compared the use of the visual imagery strategy in the experimental and the control group. an anova with age group and experimental condition as between-subject factors revealed a significant main effect of age group (f(3,154) = 10.97; p < .001; η2 = .18) with post-hoc analysis showing that younger adults (.83) applied mental imagery more often than the other age groups (children: .48; adolescents: .53; older adults: .55). furthermore, we found a significant main effect of experimental condition (f(1,154) = 221.46; p < .001; η2 = .63): participants who received the strategy instruction used mental imagery much more often than participants in the control condition (.28 vs. .95). the interaction between age group and experimental condition was also significant (f(3,154) = 9.38; p < .001; η2 = .16). separate analyses carried out for participants in the experimental group on the one hand and participants in the control group on the other hand showed there was no main effect of age group for participants who received the strategy instruction, indicating that the participants were able to transfer the training to the learning process. for subjects in the control group, the anova revealed a main effect of age group (f(3,74) = 12.30; p < .001; η2 = .35). a post-hoc analysis showed that the percentage of younger adults (.65) spontaneously applying a mental imagery strategy was significantly higher than that of the other age groups. more specifically, none of the von  der  linden  et  al         | f l r     49   children, only 5% of the adolescents, and about 25% of the older adults applied such a strategy spontaneously. table 4 percentage of subjects reporting the use of visual imagery as a function of age group and experimental condition age group children adolescents younger adults older adults control group experimental group .00 .95 .05 1.00 .65 1.00 .25 .85 4. discussion the present study is among the first to explore metacognitive monitoring skills across the life-span, and also to investigate the effects of a memory strategy training principally suited to improve skills in this domain. thus our study combined traditional interests of the developmental and cognitive literature on metacognition. in particular, we focused on possible positive effects of strengthening highly diagnostic cues (koriat, 1997). this was achieved by training half of our participants in visual imagery before memorizing item pairs, and by assessing the training effects on monitoring quality, that is, jol and cj differentiation and accuracy. we postulated that instructing subjects to connect idiosyncratic content to items should lead them to rely less on familiarity but to focus on recollection processes when monitoring their performance. an important innovative aspect of our study was its life-span perspective in that four age groups (children in third grade, adolescents in seventh and eighth grade, younger and older adults) were included in the sample. especially for children only very few studies exist that have explored the cues underlying monitoring judgments but for all included age-groups more research on the effects of familiarity and recollection-based cues is needed. first, the results show that our manipulation of task difficulty across the age groups was successful: recognition rates in both the experimental and the control condition were of comparable height across the four age groups. thus, the task was suitable for participants from primary school to older adulthood. secondly, we found significant gains in recognition performance in subjects who underwent the strategy training. this effect was most pronounced for recombined and new item pairs but still substantial for the overall data. these results thus point to the fact that our experimental manipulation was successful and are in line with many previous findings: it has been shown for various age groups from primary school to older adulthood that visual imagery is an efficient strategy in paired-associate learning (richardson, 1998; verhaegen et al., 1992), and that its instruction leads to superior recall performance compared to spontaneous use (shing et al., 2009). also in accordance with the literature, self-reported use of visual imagery was higher in the experimental than in the control group, and substantial spontaneous use of visual imagery was only reported by young adults. we acknowledge that we cannot rule out the possibility that pre-training differences had an effect on memory performance. yet, participants in our study were randomly assigned to experimental and control group which substantially reduces the risk of imbalance in potentially confounding factors. although our sample was at the lower end of recommended sample sizes for randomization (bortz & döring, 2009), the expected positive results of the strategy training on recognition performance and the reported low level of von  der  linden  et  al         | f l r     50   strategy use in the control group speak for a true effect of the strategy training. however, in future research the inclusion of a pre-training measurement would further substantiate the results. in accord with our hypothesis, we found evidence that both jol differentiation between later correct and incorrect answers and jol accuracy as measured by goodman-kruskal gamma correlations were enhanced by our strategy training in certain age groups. concerning jol differentiation, in the experimental group subjects of all age groups differentiated between correct and incorrect answers compared to the control group where only adolescents and younger adults gave higher jols before correct than before incorrect responses. additionally, adolescents in the experimental condition descriptively showed a more pronounced discrimination between later correct and incorrect answers, as compared to those in the control condition. concerning jol accuracy, the analysis revealed that only children’s accuracy improved significantly by the strategy instruction. although there was a similar tendency in the group of older adults, only a marginal effect of strategy training was found. similarly, adolescents’ and younger adults’ gamma correlations were not enhanced by the training. we also detected the expected positive effect of visual imagery on cj quality. for cj accuracy, we found a significant main effect of strategy instruction, implying that all age groups benefitted from the strategy instruction. concerning differentiation, training effects were found for children and older adults; here, the difference between correct and incorrect answers was about twice as high in the experimental condition than in the control condition. in contrast, adolescents and younger adults showed about the same amount of differentiation in both conditions. the strategy training had positive effects on both jols and cjs which were most pronounced for children and older adults for both monitoring indicators. yet, impact on cjs could be detected in a broader age range than for jols. this points to the fact that both jols and cjs rely on cues involved in our strategy instruction but at the same time draw onto different sources. for cjs information from the retrieval process might be most significant and it is possible that cues based on recollection processes are even more important for accurate judgments than for jols. further research is needed to clarify this point. in sum, the results confirm our predictions that a strategy training can improve the quality of metamemory monitoring judgments. this finding is in line with outcomes of other studies that have shown a positive influence of strategy use on prospective and retrospective monitoring judgments for children, younger and older adults (hertzog et al., 2010; nietfield & schraw, 2002; robinson et al., 2006; shing et al., 2009) and expands the existing literature by training strategy use explicitly and by the inclusion of two monitoring judgments. the developmental trends found in our study were also as expected: that is, children and older adults benefitted most from strategy instruction in terms of enhancing both their jol and cj quality, followed by the adolescents. these results are in accordance with developmental trends in regard to production deficits concerning strategy use (naveh-benjamin et al., 2007), and of recollection processes in general (ghetti & angelini, 2008). this outcome also emphasizes the practicability of our training. it proved to be effective yet was simple enough to be understood by elementary school children, and could also be successfully acquired by older adults in very short time. still, the strategy training did not account for much monitoring improvements in younger adults, with the exception of cj accuracy. one possible explanation for this finding is the high level of spontaneous use of visual imagery reported by young adults in the control group. it appears likely that the short strategy training did not greatly improve the already high level of strategy use in young adults. obviously, young adults showed high competence to memorize and to monitor their recall performance in paired-associate learning without further instruction. it is possible and should be investigated in further research that more pronounced effects of a strategy training would be found on more complex tasks. support for this assumption comes from studies where gains from a strategy training in cj accuracy could be shown for a comprehensive problem-solving task (nietfield & schraw, 2002). von  der  linden  et  al         | f l r     51   presumably further reasons were responsible for the fact that we were not able to confirm the positive effects of the strategy instruction for all age groups and for all indicators of metacognitive monitoring. one possible cause is the influence of the memory paradigm. a recognition task was chosen in this study in order to explore the basis of memory and monitoring processes by collecting rfn-judgments. specifically, participants were asked to rate the quality of their recognition memory as recollection, familiarity or no memory as an indication of the mode of action of our strategy training. yet, it seems possible that recognition processes make differentiation of recollection and familiarity-based cues more difficult than recall, given that no active memory retrieval was necessary. recall processes seem to offer more cues to increase the accuracy of monitoring judgments, as compared to recognition (buratti & allwood, 2012). other research showed a positive effect of strategy use on monitoring accuracy required active memory recall (hertzog et al., 2010; robinson et al., 2006). although shing and colleagues (2009) found positive effects of a strategy training on cjs in a recognition task, they used more complicated stimuli (malay word pairs) than those used in this study. furthermore they collected metacognitive measures of calibration and not resolution as done here. another explanation for this unexpected outcome could be that we collected delayed jols which have been shown to be more accurate than immediate jols (nelson & dunlosky, 1991). this seems to be due to the fact that delayed jols in all age groups are based on active assessment of long-term memory (recollection processes) instead of short-term and long-term memory as in immediate jols. yet, we wanted to explore a possible add-on effect to maximize the quality of monitoring judgments. this in turn makes it more difficult to show an effect than with immediate jols which are commonly used in many studies (daniels et al., 2009; hertzog, fulton, mandiwala, & dunlosky, 2013; robinson et al., 2006). a third possibility is that the effects of a short strategy intervention as used here are generally limited. possibly results of a longer intervention could exceed the promising findings of our training in all age groups and for both monitoring indicators. this issue would be worth to be explored in a follow-up study. yet in general, a strategy instruction proved to be a promising starting point to influence monitoring processes in different age groups and across various monitoring indicators. as a mode of operation we proposed that the strategy instruction should be effective due to a shift in accessible cues. specifically, we assumed that improvement should be due to the fact that now sources of monitoring judgments should be less familiarity-based cues and increasingly recollection-based cues. rfn-judgments confirm that as predicted the number of familiarity based judgments was significantly reduced in the experimental group. this trend was accompanied by more “no memory” responses. we assume that subjects in the experimental group profited from the instruction in that they were able to decide for which memories no reliable cues existed. in such cases, the answer “no memory” was correctly given. participants could have used recall of interactive imagery to discriminate recollection states: if they can recall something about the image created to memorize a word pair, they are more confident that they base their delayed-jol on a diagnostic cue. strategy recall seems to have a similar effect on feeling-of-knowing judgments (hertzog, fulton, sinclair, & dunlosky, 2014). thus, although the number of recollection-based memories as perceived by the participants could not be enhanced, the strategy training seems to have increased participants’ awareness of possible cues and enabled them to distinguish more securely between real memories and no memories. this contrast of correct recall and no retrieval at the time of the jol has been shown to be the most important source for high accuracy of delayed-jols (nelson et al., 2004). it appears likely that the number of guesses which probably fell into the familiarity category could be successfully reduced by our training. given that no age effects were found, the instruction seems to be effective in a similar way across all age groups in that it reduces the impact of familiarity. in sum, the present study yields evidence that a strategy training is a suitable means to improve prospective and retrospective monitoring processes throughout the life-span, especially in children and older adults. the instruction used in this study proved to be an economic procedure that could be successfully applied in different age groups. the training was simple enough to be easily mastered by both elementary school children and older adults. therefore a transfer to everyday-life situations seems possible. von  der  linden  et  al         | f l r     52   the findings of the present study emphasize the significance of recollection-based cues as well as its distinction from other cues for metacognitive monitoring processes and encourage further research in this direction. especially an expansion of our findings to more complex stimuli like texts or films and the investigation of the effect of a more elaborate strategy training are of interest. this would allow to  further test relevance for every-day life. additionally, it would be interesting to explore the effects of a strategy training on a larger samples as our study included relative small sample sizes per group and to follow up long-term effects of the training. we do not exclude the possibility that multiple cues underlie and can significantly influence monitoring processing. yet, along with recent research (ghetti et al., 2008; mccabe & soderstrom, 2011; metcalfe & finn, 2008; toth et al., 2011) we assume that exploring the role of recollection and familiarity processes in mediating the accuracy of monitoring judgments is a promising issue for future research. improving monitoring processes is of great importance as it is very closely linked to memory performance. thus successful monitoring represents a very valuable competence in different learning contexts. our results demonstrate that monitoring occurs from childhood on, but that there is still room for improvement at every age level. at the same time, the findings also illustrate that there are very economical ways to improve metacognitive monitoring in different age-groups. they thus indicate a direction which is worth to be pursued in future research. keypoints approaches to improve metacognitive monitoring in a broad age range covering the life-span are still very rare. integration of traditional developmental and cognitive research questions, as in this study, are scarce in the metacognitive literature. our results show that a short training in visual imagery enhances both memory and metamemory performance, especially in children and older adults. improvements in monitoring seem to be associated to a use of more reliable cues after the strategy training. acknowledgments this research was conducted as part of a research project on the development of procedural metacognitive knowledge across the life-span and financed by the german research foundation (dfg-gz. schn 315/45-1). we wish to thank all participating children, adolescents and adults as well as teachers, principals and parents for their cooperation. references begg, i., duft, s., lalonde, p., melnick, r., & sanvito, j. (1989). memory predictions are based on ease of processing. journal of memory and language, 28, 610-632. doi: 10.1016/0749-596x(89)90016-8 benjamin, a. s., & bjork, r. a. (1996). retrieval fluency as a metacognitive index. in l. reder (ed.), implicit memory and metacognition (pp. 309-338). hillsdale, nj: erlbaum. doi: 10.1037/00963445.127.1.55 von  der  linden  et  al         | f l r     53   bortz, j., & döring, n. (2009). forschungsmethoden und evaluation für humanund sozialwissenschaftler. heidelberg: springer. buratti, s., & allwood, c. m. (2012). the accuracy of meta-metacognitive judgments: regulating the realism of confidence. cognitive processing, 13, 243-253. doi: 10.1007/s10339-012-0440-5 connor, l. t., dunlosky, j., & hertzog, c. (1997). age-related differences in absolute but not relative metamemory accuracy. psychology and aging, 12 (1), 50-71. doi: 10.1037/0882-7974.12.1.50 daniels, k. a., toth, j. p., & hertzog, c. (2009). aging and recollection in the accuracy of judgments of learning. psychology and aging, 24, 494-500. doi: 10.1037/a0015269 dunlosky, j., & metcalfe, j. (2009). metacognition. thousand oaks, ca: sage publications, inc. flavell, j. h. (1999). cognitive development: children’s knowledge about the mind. annual review of psychology, 50, 21-45. doi: 10.1146/annurev.psych.50.1.21 ghetti, s., & angelini, l. (2008). the development of recollection and familiarity in childhood and adolescence: evidence from the dual-process signal detection model. child development, 79, 339-358. doi: 10.1111/j.1467-8624.2007.01129.x ghetti, s., lyons, k. e., lazzarin, f., & cornoldi, c. (2008). the development of metamemory monitoring during retrieval: the case of memory strength and memory absence. journal of experimental child psychology, 99, 157-181. doi: 10.1016/j.jecp.2007.11.001 hertzog, c., fulton, e. k., mandviwala, l., & dunlosky, j. (2013). older adults show deficits in retrieving and decoding associative mediators generated at study. developmental psychology, 49, 1127-1131. doi: 10.1037/a0029414 hertzog, c., fulton, e. k., sinclair, s. m., & dunlosky, j. (2014). recalled aspects of original encoding strategies influence episodic feelings of knowing. memory and cognition, 42, 126-140. doi: 10.3758/s13421-013-0348-z hertzog, c., sinclair, s. m., & dunlosky, j. (2010). age differences in the monitoring of learning: crosssectional evidence of spared resolution across the adult life span. developmental psychology, 46, 939948. doi: 10.1037/a0019812 holmes, a. e., & weaver iii, c. a. (2010). eyewitness memory and misinformation: are remember/know judgments more reliable than subjective confidence? applied psychology in criminal justice, 6, 47-61. kelley, c. m. & jacoby, l. l. (1996). adult egocentrism: subjective experience versus analytic bases for judgment. journal of memory and language, 35, 157-175. doi: 10.1006/jmla.1996.0009 kelley, c. m., & sahakyan, l. (2003). memory, monitoring, and control in the attainment of memory accuracy. journal of memory and language, 48, 704-721. doi: 10.1016/s0749-596x(02)00504-1 koriat, a. (1997). monitoring one's own knowledge during study: a cue-utilization approach to judgments of learning. journal of experimental psychology: general, 126, 349-370. doi: 10.1037/00963445.126.4.349 koriat, a. (2012). the self-consistency model of subjective confidence. psychological review, 119, 80113. doi: 10.1037/a0025648 koriat, a., ackerman, r., lockl, k., & schneider, w. (2009). the memorizing effort heuristic in judgments of learning: a developmental perspective. journal of experimental child psychology, 102, 265-279. doi: 10.1016/j.jecp.2008.10.005 koriat, a., & levy-sadot, r. (1999). processes underlying metacognitive judgments: information-based and experience-based monitoring of one's own knowledge. in s. chaiken, & y. trope (eds.), dual-process theories in social psychology (pp. 483-502). new york, ny: guilford press. koriat, a., & shitzer-reichert, r. (2002). metacognitive judgments and their accuracy: insights from the processes underlying judgments of learning in children. in m. izaute, p. chambres, p.-j. marescaux (eds.), metacognition: process, function, and use (pp. 1-17). new york: kluwer. leonesio, r. j., & nelson, t. o. (1990). do different metamemory judgments tap the same underlying aspects of memory? journal of experimental psychology, 16, 464-470. doi: 10.1037/02787393.16.3.464 von  der  linden  et  al         | f l r     54   matvey, g., dunlosky, j., & guttentag, r. (2001). fluency of retrieval at study affects judgments of learning (jols): an analytic or nonanalytic basis for jols? memory & cognition, 29, 222-233. doi: 10.3758/bf03194916 mccabe, d. p., & soderstrom, n. c. (2011). recollection-based prospective metamemory judgments are more accurate than those based on confidence: judgments of remembering and knowing (jorks). journal of experimental psychology: general, 140, 605-621. doi: 10.1037/a0024014 metcalfe, j., & finn, b. (2008). familiarity and retrieval processes in delayed judgments of learning. journal of experimental psychology, learning, memory & cognition, 34, 1084-1097. doi: 10.1037/a0012580 naveh-benjamin, m., brav, t. k., & levy, o. (2007). the associative memory deficit of older adults: the role of strategy utilization. psychology and aging, 22, 202-208. doi: 10.1037/0882-7974.22.1.202 nelson, t. o. (1984). a comparison of current measures of the accuracy of feeling-of-knowing predictions. psychological bulletin, 95, 109-133. doi: 10.1037/0033-2909.95.1.109 nelson, t. o., & dunlosky, j. (1991). when people's judgments of learning (jols) are extremely accurate at predicting subsequent recall: the "delayed-jol effect". psychological science, 2, 267-270. doi: 10.1111/j.1467-9280.1991.tb00147.x nelson, t. o., & narens, l. (1990). metamemory: a theoretical framework and new findings. in g. bower (ed.), the psychology of learning and motivation: advances in research and theory (vol. 26, pp. 125141). new york: academic press. nelson, t. o., narens, l., & dunlosky, j. (2004). a revised methodology for research on metamemory: prejudgment recall and monitoring (pram). psychological methods, 9, 53-69. doi: 10.1037/1082989x.9.1.53 nietfeld, j. l., & schraw, g. (2002). the effect of knowledge and strategy training on monitoring accuracy. the journal of educational research, 95(3), 131-142. doi: 10.1080/00220670209596583 pressley, m., & levin, j. r. (1977). developmental differences in subjects' associative-learning strategies and performance: assessing a hypothesis. journal of experimental child psychology, 24, 431-439. doi: 10.1016/0022-0965(77)90089-3 rhodes, m. g., & castel, a. d. (2008). metacognition and part-set cuing: can inference be predicted at retrieval. memory and cognition, 36, 1429-1438. doi: 10.3758/mc.36.8.1429 richardson, j. t. e. (1998). the availability and effectiveness of reported mediators in associative learning: a historical review and an experimental investigation. psychonomic bulletin & review, 5, 597-614. doi: 10.3758/bf03208837 robinson, a. e., hertzog, c., & dunlosky, j. (2006). aging, encoding fluency, and metacognitive monitoring. aging, neuropsychology, and cognition, 13(3-4), 458-478. doi: 10.1080/13825580600572983 robinson, m. d., johnson, j. t., & robertson, d. a. (2000). process versus content in eyewitness metamemory monitoring. journal of experimental psychology: applied, 6(3), 207-221. doi: 10.1037/1076-898x.6.3.207 roderer, t., & roebers, c. m. (2010). explicit and implicit confidence judgments and developmental differences in metamemory: an eye-tracking approach. metacognition and learning, 5, 229-250. doi: 10.1007/s11409-010-9059-z roebers, c. m. (2002). confidence judgments in children’s and adults’ event recall and suggestibility. developmental psychology, 38, 1052-1067. doi: 10.1037/0012-1649.38.6.1052 roebers, c. m., von der linden, n., schneider, w., & howie, p. (2007). children’s metamemorial judgments in an event recall task. journal of experimental child psychology, 97, 117 137. doi: 10.1016/j.jecp.2006.12.006 schneider, w. (2010). metacognition and memory development in childhood and adolescence. in h. s. waters & w. schneider (eds.), metacognition, strategy use, and instruction (pp. 54-81). new york: guilford press. schneider, w. (2015). memory development from early childhood through emerging adulthood. new york: springer. von  der  linden  et  al         | f l r     55   schneider, w., visé, m., lockl, k., & nelson, t. o. (2000). developmental trends in children's memory monitoring: evidence from a judgment-of-learning (jol) task. cognitive development, 15, 115-134. doi: 10.1016/s0885-2014(00)00024-1 shing, y. l., werkle-bergner, m., li, s.-c., & lindenberger, u. (2009). committing memory errors with high confidence: older adults do but children don't. memory, 17, 169-179. doi: 10.1080/09658210802190596 son, l. k., & metcalfe, j. (2000). metacognitive and control strategies in study-time allocation. journal of experimental psychology: learning, memory, and cognition, 26, 204-221. doi: 10.1037/02787393.26.1.204 souchay, c., bacon, e., & danion, j. (2006). metamemory in schizophrenia: an exploration of the feelingof-knowing state. journal of clinical and experimental neuropsychology, 28, 828-840. doi: 10.1080/13803390591000846 toth, j. p., daniels, k. a., & solinger, l. a. (2011). what you know can hurt you: effects of age and prior knowledge on jol accuracy. psychology and aging, 26, 919-931. doi: 10.1037/a0023379 touron, d. r., hertzog, c., & speagle, j. z. (2010). subjective learning discounts test type: evidence from an associative learning and transfer task. experimental psychology, 327-337. doi: 10.1027/16183169/a000039 verhaeghen, p., marcoen, a., & goossens, l. (1992). improving memory performance in the aged through mnemonic training: a meta-analytic study. psychology and aging, 7, 242-251. doi: 10.1037/08827974.7.2.242 weinert, f. e., & schneider, w. (1996). entwicklung des gedächtnisses . in d. albert & k.-h. stapf (hrsg.), enzyklopädie der psychologie, themenbereich c, serie ii, band 4 (s. 433-487). göttingen: hogrefe. willoughby, t., porter, l., belsito, l., & yearsley, t. (1999). use of elaboration strategies by students in grades two, four and six. the elementary school journal, 99(3), 221-231. wong, j. t., cramer, s. j, & gallo, d. a. (2012). age-related reduction of the confidence-accuracy relationship in episodic memory: effects of recollection quality and retrieval monitoring. psychology and aging, 27(4), 1053-1065. doi: 10.1037/a0027686 zakay, d., & tuvia, r. (1998). choice latency times as determinants of post-decisional confidence. acta psychologica, 98, 103-115. doi: 10.1016/s0001-6918(97)00037-1 frontline learning research 2 (2013) 12-34 issn 2295-3159 corresponding author: jenna vekkaila, department of teacher education, faculty of behavioural sciences, university of helsinki, p.o. box 9 (siltavuorenpenger 7), 00014 helsinki, finland. e-mail: jenna.vekkaila@helsinki.fi http://dx.doi.org/10.14786/flr.v1i2.43 12 | f l r focusing on doctoral students’ experiences of engagement in thesis work jenna vekkaila a , kirsi pyhältö a , kirsti lonka a a university of helsinki, finland article received 13 june 2013 / revised 18 september 2013 / accepted 19 september 2013 / available online 20 december 2013 abstract little is known about what inspires students to be involved in their doctoral process and stay persistent when facing challenges. this study explored the nature of students’ engagement in the doctoral work. altogether, 21 behavioural sciences doctoral students from one top-level research community were interviewed. the interview data were qualitatively content analysed. the doctoral students described their engagement in terms of experiences of dedication and efficiency. they rarely reported experiences of absorption. the primary sources of their engagement in their thesis work were increased sense of competence and relatedness. in addition, three qualitatively different forms of engagement in doctoral work including adaptive engagement, agentic engagement and work-life inspired engagement were identified from the doctoral students’ descriptions. further, there was a variation among the students in terms of what forms of engagement they emphasised in different phases of their doctoral studies. this study contributed to the literature on doctoral student engagement by opening the nature of engagement at the interfaces of studying and working by shedding light on the dual role of doctoral students as both students and professional researchers. moreover, this study broke down the complexity of engagement by identifying qualitatively different experiences and sources of engagement. the results encourage designing such engaging learning environments for doctoral students that promote their experiences of being competent researchers and integrated into their scholarly community. keywords: engagement; doctoral education; doctoral experience; scholarly community j. vekkaila et al. 13 | f l r 1. introduction doctoral studies are about learning in terms of research work and becoming an acknowledged researcher in a scholarly community. this takes place at the interfaces of studying and working. conducting doctoral research can be seen as both academic work and studying. doctoral students take their first steps as professional researchers by carrying out doctoral research and teaching undergraduates, which both can be considered to be academic work (brew, boud, & namgung, 2011; golde, 1998; turner & mcalpine, 2011). however, doctoral students also take courses in the role of a student (brew et al., 2011; golde, 1998; turner & mcalpine, 2011). such dual role at the interfaces of studying and working are nowadays required also more generally in life-long training to various professions in business, industry and government by the wider knowledge economy (boud & tennant, 2006; bourner, bowden, & laing, 2001; park, 2005) where solving complex, ill-defined problems (alexander, 1992; lonka, 1997) is constantly increasing. although doctoral students are highly competent and successful based on their academic backgrounds, earning the doctorate is always a highly challenging process. for instance, in doctoral education literature students‟ experienced distress (e.g., hyun, quinn, madon, & lustig, 2006; kurtz-costes, helmke, & ülküsteiner, 2006; toews et al., 1997) and remarkably high attrition rates varying from 30% to 50% (gardner, 2007; golde, 2005; lovitts, 2001; mcalpine & norton, 2006) depending on the contexts have been identified as huge challenges. especially in social sciences high attrition rates among doctoral students are a major concern (lovitts, 2001; lovitts & nelson, 2000; mcalpine & norton, 2006; nettles & millet, 2006). so-called “soft” or “ill-defined” domains such as the social and behavioural sciences are characterised by relatively loose theoretical structure and target of interest as well as unspecific strategies of inquiry (alexander, 1992; biglan, 1973a, 1973b). in such domains researchers often define and are involved in their own individual projects (e.g., lovitts, 2001). therefore, individualistic research structure may promote the idea of independent thinkers (chiang, 2003). however, it can also entail separation, which, in turn, is likely to promote negative experiences (e.g., chiang, 2003; lovitts, 2001) and consideration of interrupting doctoral studies (e.g., stubb, pyhältö, & lonka, 2011). in order to find ways to support doctoral student persistence research on doctoral education has for a long time focused on attrition and negative experiences (e.g., golde, 1998, 2005; lovitts, 2001; vassil & solvak, 2012; vekkaila, pyhältö, & lonka, 2013). research among undergraduate students, however, suggests that by focusing on strengths, positive emotions and full functioning (bresó, schaufeli, & salanova, 2011; krause & coates, 2008; ouweneel, le blanc, & schaufeli, 2011), a better understanding on doctoral students‟ engagement can be attained. this understanding provides tools for creating increasingly engaging environments for doctoral students (e.g., pontius & harper, 2006). our study aimed at filling the gap in the doctoral education literature by exploring the nature of doctoral students‟ engagement in their thesis work in the domain of behavioural sciences. 1.1 engagement in doctoral work owing to the dual nature of doctoral research, our study draws both on research on work engagement (e.g., schaufeli, martínez, pinto, salanova, & bakker, 2002a; schaufeli, salanova, gonzález-romá, & bakker, 2002b) and on study engagement (e.g., fredricks, blumenfeld, & paris, 2004; reeve, jang, carrell, jeon, & barch, 2004) to examine doctoral student engagement in doctoral work. engagement refers to a student‟s active involvement in a task or an activity at hand (e.g., case 2008; fredricks et al., 2004; reeve et al., 2004). accordingly, doctoral student engagement entails active involvement in the learning opportunities and practices provided by their environments. engagement is characterised by positive, fulfilling experiences including vigour, dedication and absorption (salanova, schaufeli, martínez, & bresó, 2010; schaufeli et al., 2002a, 2002b). vigour refers to high levels of energy and mental resilience while working, the willingness to invest effort in one‟s work, and persistence in the face of difficulties (schaufeli et al., 2002b). dedication, on the other hand, is characterised by a sense of significance, enthusiasm, inspiration, pride and challenge (schaufeli et al., 2002b). being fully concentrated http://www.tandfonline.com/action/dosearch?action=runsearch&type=advanced&result=true&prevsearch=%2bauthorsfield%3a%28mart%c3%adnez%2c+isabel%29 http://www.tandfonline.com/action/dosearch?action=runsearch&type=advanced&result=true&prevsearch=%2bauthorsfield%3a%28mart%c3%adnez%2c+isabel%29 http://www.tandfonline.com/action/dosearch?action=runsearch&type=advanced&result=true&prevsearch=%2bauthorsfield%3a%28bres%c3%b3%2c+edgar%29 j. vekkaila et al. 14 | f l r on and immersed in one‟s work characterises absorption (schaufeli et al., 2002b). absorption is close to the flow experience in which an individual is deeply immersed in an activity that is intrinsically enjoyable (csikszentmihalyi, 1990). there is evidence that engaged doctoral students were likely to feel effective and satisfied with their thesis work, and remained determined when encountering challenges (virtanen & pyhältö, 2012). in contrast, students who suffered from disengagement from their doctoral studies, were likely to feel less satisfied and more likely to give up (vekkaila et al., 2013). moreover, engaged doctoral students have, for instance, been shown to attain better learning outcomes and relationships within their scholarly community (gardner & barnes, 2007). several factors contribute to engagement (e.g., llorens, schaufeli, bakker, & salanova, 2007; reeve et al., 2004; schaufeli & bakker, 2004). for instance, in previous studies on doctoral education good quality supervision, support and constructive feedback (e.g., golde, 2005; hoskins & goldberg, 2005) as well as meaningful interaction within the scholarly community (e.g., gardner, 2007; deem & brehony, 2000; lovitts, 2001; pyhältö, stubb, & lonka, 2009; stubb et al., 2011) have been identified as predictors of doctoral students‟ satisfaction, study persistence and well-being. for instance, weidman and stein (2003) found a link between the number of faculty-student interactions and students‟ involvement in their research projects. moreover, ives and rowley (2005) showed that a constructive supervisory relationship was associated with students‟ progress and satisfaction with their doctoral studies, and hence their involvement in their thesis projects. 1.2 engagement and dynamic interplay between doctoral students and their environments the scholarly community often provides the primary work environment for doctoral students (brew et al., 2011; gardner, 2007; mcalpine & amundsen, 2008; pyhältö et al., 2009). hence, doctoral students‟ learning is highly embedded in the practices of a scholarly community. however, this community itself is a complex, multilayered, nested entity (mcalpine & norton, 2006) that can be defined as a discipline such as „education,‟ as a faculty, or as a specific research group (e.g., austin, 2002; pyhältö, nummenmaa, soini, stubb, & lonka, 2012a; white & nonnamaker, 2008). accordingly, the community provides various arenas and forms for student participation such as interaction with faculty, participation in international conferences, peer collaboration, working in a research group and teaching undergraduate students (brew et al., 2011; pyhältö & keskinen, 2012). further, students‟ involvement in the various arenas such as conducting research work, attending courses and participating in research collaboration may promote their experiences of dedication to and vigour in earning the doctorate as well as absorption in conducting research work. the previous findings on doctoral education imply that the doctoral student engagement is regulated by a complex, dynamic interplay between the student and the environment rather than a single individual or environmental attribute (e.g., golde, 2005; virtanen & pyhältö, 2012; vekkaila, pyhältö, hakkarainen, keskinen, & lonka, 2012; vekkaila et al., 2013). this includes that doctoral students‟ experiences of engagement are constantly constructed and re-constructed in the student-environment interaction. such interaction entails the students‟ prior learning experiences, beliefs, goals, and the practices and culture of the environment. doctoral students‟ perceptions, participation and other practices are mediated by their prior experiences and knowledge that have developed during their undergraduate studies and in their other professional careers or personal lives. the culture and practices of the environment, in turn, affect doctoral students‟ thinking, actions, and engagement. accordingly, the complex doctoral student-learning environment interrelation mediates students‟ engagement in the doctoral process. the dynamic interplay between the learner and learning environment (e.g., lindblom-ylänne & lonka, 2000) contributes to not only whether or not students engage in their studies (e.g., fredricks et al., 2004; leiter & bakker, 2010) but also to the ways in which they engage in their studies. accordingly, the quality of the dynamics between the doctoral student and the environment is likely to contribute to the ways the student engages in doctoral work. dynamics between students and their environment can contribute j. vekkaila et al. 15 | f l r to students‟ sense of relatedness, competence, autonomy (deci & ryan, 2002, 2008) and contribution (eccles, 2008). deci and ryan (2002) have proposed that the experiences of relatedness, competence and autonomy are the prerequisites for individuals‟ personally meaningful actions and experiences (see the selfdetermination theory). the sense of relatedness refers to feeling connected to others, having sense of belonging both with other individuals and with one‟s community, and be integral to and accepted by others (deci & ryan, 2002). the sense of competence, in turn, focuses on feeling effective and confident in one‟s on-going actions within the social environment and experiencing opportunities to express and exercise one‟s capacities (deci & ryan, 2002). when individuals are autonomous they feel as if they are the source of their own actions and behaviour even when those actions are influenced by outside forces (deci & ryan, 2002). that is, their actions are based on their own personal interests and values. furthermore, it is important to feel a sense of contribution when acting in a personally meaningful way (eccles, 2008). thus, the experiences of belonging, competence, autonomy and contribution are necessary in order to promote doctoral students‟ engagement (mason, 2012; virtanen & pyhältö, 2012). for instance, appel and dahlgren (2003) found that doctoral students were inspired in their studies by the opportunities available for intellectual development, feelings of having internal locus of control and academic freedom as a researcher, and chances to make a difference by their doctoral project. in addition, stubb et al. (2011) and pyhältö and keskinen (2012) more recently found that the doctoral students who experienced their scholarly community in a positive way, that is, as empowering, or who perceived themselves as active agents, less often reported lack of interest towards their own studies and considered interrupting their doctoral process less often than those students who had negative experiences or perceived themselves as passive objects. this indicates that doctoral students can be active in certain interaction arenas of a scholarly community whereas in some other communities they may participate infrequently and be more in a role of an observer. this, in turn, is likely to contribute to their engagement in doctoral work. it follows that also students‟ engagement in terms of how agentic (reeve & tseng, 2011) they experience themselves in their doctoral work may vary. at its best a doctoral student‟s peripheral role gradually evolves towards active, relational agency as the student is involved more intensively in the research group‟s shared knowledge creation practices, and develops a sense of ownership of one‟s own doctoral research and identity as a researcher (hakkarainen, hytönen, makkonen, seitamaa-hakkarainen, & white, 2013; hopwood, 2010; pyhältö & keskinen, 2012; pyhältö et al., 2012a; vekkaila et al., 2012). relational agency (edwards, 2005) refers to the capacity of doctoral students to work with other members of their research community in order to better respond to complex research problems (pyhältö & keskinen, 2012). this means that doctoral students are not influenced only by the scholarly community but can, at least to some extent, choose their primary arenas in which to participate and take initiative, direct and re-direct their own activity and learning (pyhältö & keskinen, 2012). therefore by adopting different strategies, the students can actively modify their environment, and hence their opportunities to engage in the scholarly community in question (virtanen & pyhältö, 2012) and, further, in their doctoral work. sometimes students‟ engagement in doctoral work may be inspired mainly by their work-life experience. mäkinen, olkinuora, and lonka (2004) showed that especially in fields such as teacher education, law, and medicine, where the student aimed at professional development rather than abstract, theoretical understanding, the so-called work-life orientation dominated. in previous studies those university students who expressed so-called work-life orientation were interested in professional development and saw their studying as training for a certain profession or vocation (lonka & lindblom-ylänne, 1996; mäkinen et al., 2004; vermunt, 1996). they appreciated directly useful, concrete and applicable knowledge (lonka & lindblom-ylänne, 1996; mäkinen et al., 2004; vermunt, 1996). such orientation on studying was considered to reflect practical interest rather than scientific ambition (lonka & lindblom-ylänne, 1996; mäkinen et al., 2004). brint, cantwell, and hannerman (2008), for instance, found in their study on undergraduate students that the culture of engagement in the arts, humanities and social sciences focused on participation and interest in ideas, whereas the culture of engagement in the natural sciences and engineering focused more on improvement of research skills, collaborative study, and the labour market. in our recent study on natural sciences doctoral students, this was not as straightforward: the students‟ inspiration and engagement in the doctoral work was often due to a strengthened sense of belonging and participation in the various practices j. vekkaila et al. 16 | f l r of their research community (vekkaila et al., 2012). this suggests that students may have different ways of being engaged in their doctoral process and earning the doctoral degree. the present study focused on exploring behavioural sciences students‟ engaging doctoral experiences. 2. the aim of this study this study is a part of a larger national research project on doctoral education in finland that aims to understand the process of phd education (see pyhältö et al., 2009). the present study aimed at gaining a better understanding of doctoral student engagement in thesis work. in our study, the following research questions were addressed: 1. what kinds of experiences of engagement did the doctoral students describe? 2. what were the sources of engagement in doctoral work? 3. were there qualitatively different forms of engagement? 3. method 3.1 doctoral education in finnish context in finland, doctoral studies are heavily focused on conducting thesis research. there is no extensive separate course work required before launching the doctoral research project. in fact, course work from 40 to 80 european credit transfer and accumulation system (ects) credits worth of postgraduate studies depending on the discipline included in doctoral studies are usually individually constructed and based on personal study plans that typically include international conferences and some methodological studies. in behavioural sciences, an article compilation with a summary has become the dominant form of thesis (66%) during recent years (pyhältö, stubb, & tuomainen, 2011). the article compilation is more dominant in psychology, whereas in educational sciences the dominant form is the monograph (a book format). the article compilation consists of three to five internationally refereed journal articles often coauthored with the supervisors and a summary that includes an introduction and a discussion bringing together the separate articles. doctoral supervision is usually based on an apprenticeship, in both research groups and supervisor-student dyads (löfström & pyhältö, 2012). in finland, students can conduct doctoral studies full-time or part-time. the general target duration for full-time studies for the doctorate is four years. however, often the completion time for the doctorate is longer than this. according to a recent survey the average time for completing the degree in behavioural sciences is five to six years (sainio, 2010). however, some sources indicate that the average completion time may be higher ranging from seven to over ten years (pyhältö et al., 2011). this may be explained by the heavy requirements of earning the doctorate. the articles included in the article compilation need to be published in peer reviewed journals, students need to write the summary of them, the thesis need to examined by two or three pre-reviewers, a students need to defend the thesis publicly before the faculty council decides whether to award the doctoral degree. long completion times may also be explained by the nature of finnish doctoral education system, that is, doctoral studies are free for the students, the licence to conduct doctoral studies is valid for life and students can conduct their doctoral studies part-time and have other professional full-time jobs. although the doctoral education is publicly funded, the students have to cover their costs of living, which is typically done through personal grants, project funding or wages earned by working outside the university (pyhältö et al., 2011). doctoral education in finland is more detailed described by the international postgraduate student mirror (2006) and pyhältö et al. (2012a). j. vekkaila et al. 17 | f l r 3.2 participants the participants were 21 behavioural sciences doctoral students (female: 17; male: 4) from a major research-intensive finnish university. all the participants were from the same case community participating in the larger national research project on doctoral education in finland (see pyhältö et al., 2009) and its all doctoral students were invited to participate in the study. participation was voluntary. the case community was chosen because it represented a national and international well-established research group and was considered to be good representative of organisation of doctoral education. eleven of the participants were full-time doctoral students and ten were part-time. six participants were pursuing a monograph, seven a summary of articles; eight participants were unsure of the form their theses would take. all the participants had master‟s degrees, typically in educational sciences and they were in different phases of their doctoral process. according to the participants‟ own estimates, twelve of them were in the beginning of the doctoral process meaning that they were typically launching their research projects, collecting and/or analysing data, or writing their first and/or second article. four were in the middle part of the process that typically included data analysis, and writing the monograph, or writing third and/or fourth article. four of the participants were in the last part of the process that typically meant finalising the monograph or the last articles and the summary of the articles. one of participants had already graduated. all the participants were interviewed on a voluntary basis. 3.3 interviews semi-structured interview (e.g., kvale, 2007) data were collected in 2007–2008. the interviews were designed to investigate the doctoral students‟ experiences of their thesis process and their views of themselves within it (see appendix 1). at the beginning of the interviews, the students were asked some background information questions about their discipline or subject, time spent on their thesis/studies, the phase of the process and time of graduation, as well as the form of the thesis and whether they were working on it full-time or part-time. the interview focused both on the retrospection of previous experiences of the ph.d. process and on the present situation. (stubb, 2012.) the interview was piloted before the actual data collection. in the first stage, it was tested with four doctoral students in behavioural sciences, and minor modifications to the questions were made. then the interview was tested with seven science students and no further modifications were required. all interviews were conducted by a researcher from the authors‟ research group (except one, which was done by a trained research assistant). each interview lasted approximately one hour (ranging from thirty minutes to almost three hours). the interviews were recorded and transcribed. (stubb, 2012.) 3.4 analysis the interview data were qualitatively content analysed (e.g., patton, 1990) by relying on an abductive strategy (e.g., coffey & atkinson, 1996; morgan, 2007). hence, the data observations and prior understanding based on theories were repeatedly assessed in relation to each other in order to acquire the most optimal understanding of the phenomenon (coffey & atkinson, 1996; morgan, 2007), that is, doctoral student engagement, when categorising the data. at the beginning of the first analysis phase, all the text segments in which the doctoral students referred to engaging experiences in terms of their doctoral work were coded into the same hermeneutic category by using a grounded strategy (e.g., harry, sturges, & klingner, 2005; mills, bonner, & francis, 2006). accordingly, all the text segments referring to engaging doctoral experiences from the 21 interviews were grouped together and formed the ground data for further analysis. the unit of analysis included the totality of thought referring to engaging experiences ranging from a sentence to dozen sentences. these text segments included expressions of interest, inspiration, energy, devotion, meaningfulness and positive doctoral thesis related emotions. j. vekkaila et al. 18 | f l r after this, the analysis focused on what the participants experienced, that is, the different qualities of engaging doctoral experiences. data were coded into three exclusive main categories by relying on research on characteristics of engagement introduced in the literature review (e.g., salanova et al., 2010; schaufeli et al., 2002a, 2002b) as follows: (a) dedication including participants‟ experiences where they expressed earning the doctorate, being a doctoral student, and conducting research and studies as personally highly meaningful and significant, and entailing strong devotion and positive emotions such as joy, enthusiasm and inspiration; (b) efficiency including participants‟ experiences of having willingness to invest effort in their research work and studies, strengthened self-images of themselves as researchers and having the effective and energetic drive to conduct doctoral work, and (c) absorption including participants‟ experiences of intensive situations where they experienced being fully concentrated on and engrossed in their research work and studies. the three main categories reflected the main experiences of engagement in doctoral work. the category labelled “efficiency” came close to vigour (e.g., schaufeli et al., 2002a, 2002b), however, the category was named as efficiency because in students‟ descriptions experiences of strengthened self-efficacy beliefs and an energetic drive with the research work were emphasised. at the end of the first phase, the analysis focused on what contributed to students‟ experiences of engagement in their doctoral work. the text segments in the categories representing the main experiences of engagement were coded into four basic categories according to the primary sources of, that is, causes for engagement as described by the participants by relying on deci and ryan‟s (2002, 2008) as well as eccles‟s (2008) works introduced in the literature review: (a) competence including participants‟ descriptions where their experiences of engagement in doctoral work were promoted by development of their academic skills and expertise, learning and developing understanding of the domain and own topic, and gaining insights into their own research; (b) relatedness including participants‟ descriptions where their experiences of engagement in doctoral work were strengthened by having dialogues and collaboration with supervisors, other researchers and peers, as well as participating in and becoming a valued part of a scholarly community; (c) autonomy including participants‟ descriptions where their experiences of engagement in doctoral work were promoted by being in control of their own research work, and following their own interest in their doctoral process, and (d) contribution including participants‟ descriptions where their experiences of engagement in doctoral work were strengthened by producing such significant scientific knowledge that make a difference, and seeing the value of their own research in practice. a visualisation of the first analysis phase is provided in figure 1. the agreement between the two classifiers regarding the independent parallel analysis of 30% (f = 36) of the text segments in relation to the main experiences of engagement was 94% and in relation to the sources of engagement was 97%. interrater reliability measured with cohen‟s kappa (κ) in regard to the main experiences of engagement was 0.91 and in regard to the sources of engagement was 0.95, indicating almost complete agreement. the text segments related to the main experiences and sources of engagement were quantified and the relation between them was analysed with cross-tabulation and χ²-tests. j. vekkaila et al. 19 | f l r figure 1. a visualisation of the first analysis phase. in the second phase of the analysis, a person-oriented analysis strategy was applied. the personoriented analysis involved that the analysis focused on identifying the forms of student‟s engagement in doctoral work. in practice, each participant‟s engaging experiences that were identified at the very beginning of the analysis from the interview data were grouped together, that is, formed own unity, and were separated from the experiences reported by the other participants. at first, the different forms of engagement presented in each participant‟s descriptions were investigated to delineate the initial categories by exploring the patterns, that is, differences and similarities in the main experiences and sources of each individual student‟s engaging experiences. also, each participant‟s engaging experiences were interpreted within the larger interview context. then, the similarities and differences in the main experiences and sources were explored across all participants‟ descriptions of engaging experiences. as a result, the experiences were divided into their own categories based on their differences, following the idea that the experiences presenting a certain form of engagement in one category were mutually similar, while being distinct enough from the other categories. the categories appeared to differ from each other in terms of how the participants expressed: (1) the dynamics between themselves and their scholarly community in the engaging experiences, and (2) the source of inspiration in their doctoral work in the engaging experiences. from the participants‟ descriptions three categories representing the qualitatively different forms of engagement in doctoral work were identified: adaptive form of engagement, agentic form of engagement and work-life inspired form of engagement. in adaptive and work-life inspired forms of engagement the dynamics between the doctoral students and their scholarly community was expressed as being static in nature, that is, providing an arena for adjusting and acquiring knowledge, whereas in the agentic form of engagement the dynamic was expressed as being reciprocal, that is, an arena for dialogue. in the students‟ expressions the source of inspiration in doctoral work in adaptive form of engagement was adapting and conforming to the current conditions and acquiring j. vekkaila et al. 20 | f l r the knowledge and skills that were valued in the scholarly community. in turn, in agentic form of engagement creating new knowledge was more emphasised as source of inspiration in doctoral work, whereas in work-life inspired form of engagement the students highlighted the importance of applying the new knowledge and skills acquired in the scholarly community in order to solve practical problems and contribute to the work-life outside academia. the qualitatively different forms of engagement were also studied in relation to the phase of studies. study phase was determined based on students‟ own evaluation of whether they were at the beginning, middle, or end of their own doctoral process. although in the participants‟ descriptions typically at least two of the forms of engagement were present, one of the forms was emphasised in their descriptions. in the results, we provide quotations of participants‟ descriptions that were translated from finnish into english. 4. results the results suggested that there was a variation in the participants‟ experiences of engagement. the doctoral students‟ descriptions of dedication, efficiency and absorption ranged from experiencing their doctoral work as highly meaningful to having energetic drive while conducting it. moreover, the sources of engagement varied from developing an understanding of one‟s own research into belonging to the scholarly community. the students also described qualitatively different forms of engagement. 4.1 main experiences of engagement in doctoral work the participants emphasised experienced dedication (53%) in their doctoral work (see table 1). for instance, the students perceived earning the doctorate and training as personally meaningful and significant, and described their strong devotion in their doctoral process and interest in their research. they also expressed extremely positive emotions including pleasure, satisfaction and joy. they were also enthusiastic about being doctoral students and pursuing their phds in the training program. for instance, as one of the students described: i like this graduate school because every time we have here a seminar, i leave it with a growing zeal. i think that conducting research is the right work for me. participating in this graduate school and its seminars really promote my excitement and inspiration. (p10) the students also often highlighted a sense of efficiency in conducting their doctoral work (40%). they reported positive, strengthened perceptions of their self-efficacy beliefs as researchers and their clear perceptions of the next steps in their research, and ability to organise and steer their own doctoral process. they were also willing to make efforts for their doctoral work and described having active, efficient and energetic drive when conducting it. as one of the students shared: when i present my work in different seminars and receive feedback . . . it has a practical influence on my work and then i really need to get to work with my research; then i know what i have to get working on next . . . it gives me energy to conduct my research further and i try to find time to conduct it . . . then my research moves forward . . . (p17) the students rarely described experiences of absorption in the doctoral work (7%). in these cases, for example, they described intensive episodes during which they were fully immersed in their work, including data analysis or writing the thesis. they were involved in the doctoral work even to the extent that other activities were brushed aside. as one student described: j. vekkaila et al. 21 | f l r then came this very intensive period . . . i was in the field collecting data every day for several months . . . i was immersed in the data collection for several years, because i found the situation in the field really interesting. (p21) 4.2 sources of engagement in doctoral work the sources that the participants identified as contributing to the engaging experiences varied. however, the engagement was often described in relation to learning and developing as a researcher as well as interacting with other researchers. table 1 shows that the participants emphasised an increased sense of competence (39%) as an important source of their engagement. the students‟ sense of competence often emerged as development of understanding or new academic skills. hence, their engagement often stemmed from learning and development as scholars. these experiences included, for instance, deepening their understanding of research work and theories, creating new knowledge, and developing their thinking and learning about their themes in more profound ways, as well as providing new insights in their research. as one student remarked: i think that finding and learning new knowledge is fun. my supervisor says that i should not read anymore, but when new research is published, i have to read it. i suppose i like to gain new insights and understanding about my research theme. they are really the best experiences in this work. (p15) almost as often, the students highlighted their sense of relatedness (37%) as a significant source of engagement in their doctoral work (table 1). characteristic of the situations in which the students‟ experiences of relatedness were promoted was that they perceived being actively involved in their scholarly community, and having a sense of belonging to it and being valued by others. they also described various participation and interaction arenas including research collaboration, receiving constructive feedback and discussions, and sharing interest and expertise with more experienced researchers, supervisors, and peers on research work in general and especially on their own doctoral research. as one of the students described: usually i become inspired by our seminars and discussions. the first thing that comes to my mind is professor h’s ways of stating concepts. he somehow makes theories clearer and adds new perspectives. i have also participated in a group where we have discussed the doctoral theses of other advanced doctoral students and through those discussions i have had many new ideas . . . i get the feeling that it is wonderful that i am able to do this and it is amazing to be here, that this work is really fun. (p1) sometimes the students described their sense of autonomy (13%) as the source of engagement in their doctoral work. they expressed the significance of being able to conduct such research work that was one of their personal interests, based on their own decisions, were in their own control, and defined on their own terms even though they often worked in research projects with other researchers. as one student commented: that seminar began and there we read the central texts related to the theory together in our graduate school group. it was an amazing time and we were given time and space to think . . . it was really nice time . . . [it was a] time when i did not have to limit myself and had the freedom to do and be. (p16) less often, the students expressed sense of contribution (11%) to be a source of their engagement. when the students described sense of contribution they typically reported the importance of being able to produce original scientific knowledge with significance and develop such understanding of the research themes that would be valued and making a difference especially in the practical work-life outside academia. as one of the students shared: j. vekkaila et al. 22 | f l r it really inspired me that some group with our support would innovate and develop a new way of performing and working and they would begin to apply it in practice. it is inspiring to be involved in those processes. i think that this research is useful and i can have an impact on something larger through this work. (p3) further investigation showed that there was a relation between the main experiences and sources of engagement (χ² =13.42, df =6, p =0.037). the sources of students‟ dedication and efficiency in terms of their doctoral work were typically their strengthened senses of competence and relatedness (table 1). table 1 the main experiences and sources of engagement in doctoral work (based on 120 engaging experiences reported by the participants) experiences of engagement sources of engagement dedication f (%) efficiency f (%) absorption f (%) total f (%) competence 19 (16%) 24 (20%) 4 (3%) 47 (39%) relatedness 22 (18%) 20 (16%) 3 (3%) 45 (37%) autonomy 12 (10%) 2 (2%) 1 (1%) 15 (13%) contribution 11 (9%) 2 (2%) 13 (11%) total 64 (53%) 48 (40%) 8 (7%) 120 (100%) 4.3 qualitatively different forms of engagement in doctoral work our person-oriented analysis showed that the participants‟ descriptions included three qualitatively different forms of engagement (see table 2). in each form the dynamics between the doctoral students and their scholarly community as well as the source of inspiration in doctoral work were expressed differently by the students. the first category was labelled adaptive form of engagement, where the students emphasised their experiences of dedication and efficiency through adapting and adjusting to their scholarly community and its research traditions and practices. such experiences reflected a static, one-directional relation between the students and their scholarly community. the students usually reported their relatedness to their own research community which provided the arena for acquiring knowledge from more experienced researchers, for instance, through supervision and following theoretical discussions. the students expressed adapting and conforming to the current conditions and acquiring the knowledge and skills that were valued in their scholarly community as the significant source of inspiration in doctoral work. such knowledge and skills included, for instance, writing skills and gaining the relevant theoretical understanding. being able to conduct the research according to the community‟s framework and criterion was also important. the students, for instance, described adaptive engagement in relation to their supervision and research as follows: i got a good feeling when i exchanged a few words with my supervisor. then it was all clear how i should continue my work . . . i learned something relevant or gained insights, because this is a new world for me . . . (p1) j. vekkaila et al. 23 | f l r overall this graduate school has been rewarding because it was a new experience to create the research plan but at the same time i could see what others had done and from others’ work i got some hints . . . i made notes and out of that mess i gradually came up with a logical vision and started to lay out my research plan. (p19) such adaptive form of engagement was most often described by students who were at the beginning of their doctoral process. in the second category, agentic form of engagement, the students emphasised their experiences of dedication and efficiency through a dialogical relationship between themselves and their scholarly community. such experiences reflected an active and re-forming interplay between the students and their community. the students also perceived their relatedness to both their own research community and the larger scholarly environment including international conferences that provided an arena for sharing research ideas, receiving constructive feedback and collaboration. the students highlighted creation of new knowledge as the important source of their inspiration in doctoral work. this included, for instance, being able to redefine their own research work in relation to their research community‟s framework, becoming autonomous and work on their own terms, and being able to argue their own point of view when contributing to their scholarly community. for example, the students expressed their agentic engagement in terms of dialogues with others and their own research work as follows: the most rewarding for me are the moments when i can share my thinking with others . . . for instance, i have those experiences where there were interesting discussions and i could present my point of view and we can develop some insights . . . i have found pleasure in those encounters in the field, or with my supervisor, when she can follow my ideas and clarify them, or through some e-mail conversations with a colleague. of course, these experiences require that i must also write something and then share it with others. (p16) at first, i did not know much and i was all at sea about on what theme i should focus my research; it was quite superficial . . . now i have hope . . . i have familiarised with it little by little and now i develop and cherish my own ideas. now i feel that it is my own project, more than before . . . (p12) the agentic form of engagement was typically reported by those students who were either halfway through or at the end of their process. the third category was labelled work-life inspired form of engagement, in which the students emphasised the influence of their professional lives on their dedication to their doctoral work. such experiences reflected three-directional relations between the students, their work-life outside academia and their research community. typically, the research community where they were receiving doctoral training provided the arena for acquiring such theoretical knowledge and research skills that extended their understanding of their research questions evolved from their work-live contexts. the students emphasised applying the new knowledge and skills in order to solve practical problems and contribute to the work-life outside academia as the significant source of inspiration in doctoral work. the students described their worklife inspired engagement, for instance, as follows: these were those moments of insight. i really understand my [professional] work now in a more profound way and can combine concepts that i have not previously realised to be related. i find answers to those questions from practical problems that i have seen in my own work . . . and i have gained a lot from the graduate school seminars where there have been discussions on these ideas . . . now, for instance, i have read a doctoral thesis and then i have gained some new insights into my own data and concepts, and through those concepts i can understand better my data . . . (p4) j. vekkaila et al. 24 | f l r actually the inspiring experiences and moments of joy or inspiration related to doctoral studies arise when i lead the groups involved in the project . . . their own zeal also encouraged me to continue and the idea that my research work could make a difference and support these practices in the future. (p10) such work-life inspired form of engagement was reported by the students in different phases of their doctoral process. table 2 qualitatively different forms of engagement (based on the person-oriented analysis of the participants’ engaging experiences) qualitatively different forms of engagement what kind of dynamic exists between the doctoral students and their scholarly community the source of inspiration in doctoral work adaptive agentic work-life inspired dedication and efficiency through a one-directional relation where the scholarly community provides the arena for the students to adjust and acquire knowledge dedication and efficiency through a dialogical relation between the students and the scholarly community where both the students and the community re-form dedication through a three-directional relation where the scholarly community provides the arena for the students to acquire knowledge to answer questions that have evolved from their work-life outside academia dedication and efficiency through conforming to the current conditions and acquiring the knowledge and skills valued in the scholarly community dedication and efficiency through creating new knowledge in relation to the scholarly community‟s theoretical framework, being able to work on their own terms and develop their own points of view dedication through applying the scholarly community‟s theoretical knowledge and research skills in order to solve practical problems and contribute to the worklife outside academia 5. discussion 5.1 theoretical reflections and implications engaging doctoral experience is rarely explored in both doctoral education and engagement literature. hence, our study provided new insight into doctoral student engagement by breaking down the complexity of engagement by identifying qualitatively different experiences and sources of engagement. results showed that the main experiences of engagement in doctoral work were dedication and efficiency. experiences of absorption were rarely reported. our finding were in line with the previous findings of work engagement research carried out in other work-life contexts and among undergraduate students where engagement is explored in terms of dedication, vigour and absorption (e.g., bresó et al., 2011; krause & coates, 2008; ouweneel et al., 2011; salanova et al., 2010; schaufeli et al., 2002a, 2002b). this implies that previous research on work engagement (e.g., salanova et al., 2010; schaufeli et al., 2002a, 2002b) appeared to provide a functional framework for exploring students‟ engagement in their doctoral work. j. vekkaila et al. 25 | f l r further investigations showed that the primary sources for engaging doctoral experiences were increased sense of competence and relatedness. the students reported sometimes sense of autonomy and contribution as sources for engagement in their doctoral work. this is in line with previous research suggesting that students‟ and workers‟ self-motivation, optimal functioning and psychological well-being are fostered when their senses of relatedness, competence, autonomy (e.g., deci & ryan, 2008; niemic & ryan, 2009; see also mason, 2012; virtanen & pyhältö, 2012) and contribution (eccles, 2008; see also virtanen & pyhältö, 2012) are promoted. however, the findings here clarify further the understanding of different sources of engagement in doctoral work. in our findings the experiences of competence and relatedness were emphasised. this may reflect the development of engagement during the doctoral process. in the present study, the doctoral students‟ dedication and sense of efficiency appeared to be strengthened when they developed their competences as researchers and became more related to their scholarly community. it may be that when students perceive themselves as competent and acknowledged members experiences of having autonomy and making contributions may become more salient ones. in addition, our results confirmed the previous findings suggesting that students‟ feelings of belonging and participation in a scholarly community contribute to their positive experiences, wellbeing as well as satisfaction with and persistence in doctoral studies (deem & brehony, 2000; golde, 2005; hoskins & goldberg, 2005; lovitts, 2001; pyhältö et al., 2009; pyhältö, vekkaila, & keskinen, 2012b; stubb et al., 2011). the results of our study also provided new insights by demonstrating how students‟ experiences of belonging were significant in terms of their engagement in doctoral work. the significance of experienced belonging among the behavioural science doctoral students may have to do with the nature of the research in their discipline. as part of the soft sciences, the behavioural sciences are sometimes characterised by solitary research work in libraries, archives or in the field (lovitts, 2001). one would therefore expect that relatedness would not be as important. in our participants reports the possibilities for experiencing being a valued, acknowledged member of a scholarly community was important. however, some students in these fields may also work in research groups (e.g., austin, 2010), for instance, in archaeology. also different forms of engagement, including adaptive engagement, agentic engagement and worklife inspired engagement, were identified. to our knowledge, qualitatively different forms of engagement have not been previously reported among university students. hence, this study contributed to the literature on doctoral student engagement by opening the nature of engagement at the interfaces of studying and working by shedding light on the dual role of doctoral students as both students and professional researchers. it is possible that the varying forms of engagement reflect the different meanings of doctoral work that were given by the participants (e.g., meyer, shanahan, & laugksch, 2005; stubb, pyhältö, & lonka, 2012a, 2012b). for instance, to some extent our results resembled the different perceptions of doctoral research found by stubb et al. (2012b). in their research the doctoral students perceived research work as 1) “a personal learning process”, 2) a “job to do”, 3) “making a contribution” and 4) “obtaining qualifications and gaining accomplishments”. the first category and the agentic and work-life inspired forms of engagement overlap with one other since in all of them the significance of exploring something that was defined in one‟s own terms or was personally interesting were emphasised by the participants. in turn, the second category and the adaptive form of engagement resemble each other, because, for both of these, the participants highlighted doctoral research as an activity in which they follow the traditions and practices of the scholarly community or its use in fulfilling the community‟s requirements for a doctorate. in addition, in the third category, answering interesting questions that made a difference was viewed as meaningful to the doctoral students, and, hence, has similarities with work-life inspired engagement. however, in work-life inspired engagement, the contribution focused mainly on professional contexts outside academia, whereas in the third category, the contribution focused both on the discipline and society. moreover, the fourth category of “accomplishment” not only included demonstrating one‟s excellent performance, but also the creation of new knowledge and, therefore, has similarities with the agentic form of engagement. however, gaining merit and status were also emphasised in this particular category, but were not expressed by the participants in relation to agentic engagement. hence, it may be that the sources of inspiration in doctoral work at least j. vekkaila et al. 26 | f l r partially reflect the students‟ motives, goals and aspirations related to their phds. furthermore, the meanings of doctoral research given by students and goals for earning the doctorate may affect what kinds of scholarly identities (e.g., pyhältö et al., 2012a) doctoral students construct, for instance, a professionally oriented one, and also their engagement in the doctoral process. if students perceive the meaning of doctoral work to be obtaining qualifications for work-life outside university and construct their identity through their professional careers it is likely to be reflected into their engagement in the doctoral process. then it may be that doctoral experiences that promote the connection between the doctoral work and professional life, and practical meaning and value of doctorate are likely to enhance students‟ engagement in doctoral work. in turn, experiences that do not enable students to make a meaningful connection between the doctorate and their aspirations may reduce their engagement in their doctoral work. moreover, our findings suggested that the qualitatively different forms of engagement were emphasised differently by the participants in different phases of the doctoral process. adaptive engagement was more often described by the students who were at the beginning of their doctoral process, agentic engagement by those students who were either halfway through or at the end of their process. work-life inspired engagement was reported by the students in all phases of the doctoral process. a reason for the adaptive form of engagement was being emphasised at the beginning of the doctoral process maybe that doctoral students‟ active agency and participation in their scholarly communities increases over time as they progress in their thesis process (e.g., hakkarainen et al., 2013; hopwood, 2010; pyhältö & keskinen, 2012). 5.2 methodological reflections and its limitations in this study, semi-structured interview data were collected and qualitative content analysis relying on abductive strategy that combined both grounded and theory-guided analyses (e.g., coffey & atkinson, 1996; harry et al., 2005; kvale, 2007; mills et al., 2006; morgan, 2007; patton, 1990) was used to identify the students‟ experiences of engagement in doctoral work. engagement has typically been investigated by using quantitative methods (e.g., ouweneel et al., 2011; salanova et al., 2010, schaufeli et al., 2002a, 2002b). the strength of our approach was that it allowed us to explore students‟ experiences of engagement in a profound manner and provided insights in the various aspects of engagement in doctoral work. certain challenges are involved in using a retrospective approach (e.g., cox & hassard, 2007). the participants‟ experiences and their overall life situations are often difficult to recall and sum up in a single interview (kvale, 2007). accordingly, the retrospection was likely to have affected the data, including a generalisation of experiences. the retrospective approach and semi-structured interviews also had their advantages (e.g., cox & hassard, 2007). the reflective and process-oriented design gave the participants an opportunity to reflect on their doctoral journey and identify significant experiences in it. this resulted in rich data and ensured that the participants recalled and reported only significant experiences. moreover, we explored the engagement among 21 behavioural sciences doctoral students who were conducting their thesis in one top-level research community. because of the distinctive features of the discipline (e.g., lindblom-ylänne, trigwell, nevgi, & ashwin, 2006; mccune & hounsell, 2005) and the limited sample size, generalising the results to other disciplines and in other countries should be done with caution. however, we have looked at doctoral students‟ experiences in other domains, and, for instance, results here resemble our (vekkaila et al., 2012) recent findings regarding natural sciences students‟ significant engaging and disengaging doctoral experiences. further longitudinal studies are needed to explore the development of engagement (e.g., demerouti, bakker, nachreiner, & schaufeli, 2001) in doctoral work among doctoral students from different domains and countries. this may provide a better understanding, for instance, of whether students experience their engagement in their thesis work differently in various domains and at the different phases of the doctoral process. also, the relation between engagement in the doctoral process, the meanings of earning the doctorate given by the students and development of a scholarly identity is worth of further investigation. j. vekkaila et al. 27 | f l r 5.3 educational implications in terms of developing more engaging learning environments for doctoral students, our findings imply that engagement is not a singular entity; instead it is multidimensional and entails various qualities. doctoral student may experience engagement in their doctoral work in varying ways, and hence it is one matter to be dedicated to doctoral research and another to experience oneself as an efficient researcher or absorption in research activities at hand. dedicated doctoral students are likely to be engaged in their doctoral work by their sense of significance, commitment and positive thesis related emotions. students feeling efficiency, in turn, are likely to express their engagement through their positive self-images as researchers and by their energetic actions, whereas absorption is likely to entail students‟ full concentration and being totally immersed in their study or research activities for certain periods of time. accordingly, the ways to support doctoral student engagement need to be diverse. our results implied that doctoral students‟ engagement in doctoral work can be supported by enhancing their experiences of being competent researchers and integrated into their scholarly community. such experiences can be supported by, for instance, facilitating doctoral students‟ participation in collaborative academic practices. an example of a practice that is likely to promote students‟ engagement is a learning community formed around certain academic activities (zhao & kuh, 2004) which are designed to strengthen students‟ positive self-efficacy beliefs, provide academic challenges, and involve active and collaborative learning techniques, interaction opportunities and social support (bresó et al., 2011; overall, deane, & peterson, 2011; umbach & wawrzynski, 2005). this could be applied in doctoral education both in supervisory meetings and in different academic groups that support, for instance, peer learning, writing processes, dialogues and collaborative problem solving (aitchison & lee, 2006; boud & lee, 2005; lonka, 2003). it is interesting that the doctoral students rarely described experiences of absorption in their doctoral work. a reason for this maybe that the experiences of absorption remain an unidentified or unused resource for supporting students‟ engagement in their doctoral work. absorption resembles the flow experience (csikszentmihalyi, 1990); hence, emerge of such intrinsically enjoyable experience can be fostered by optimising the balance between the challenges of learning tasks and students experiencing competence (e.g., inkinen et al., 2013). for instance, in their recent study on university students inkinen et al. (2013) noted that although positive and active emotions are only one aspect of the complex flow experience, they found that these kinds of emotions occurred when the perceived challenge and required skills were both very high and in balance. the balance may be reached by providing doctoral students the resources they need such as supervision, constructive feedback of their learning and development as a researcher, peer support, and control over their own research. then, when doctoral students experience balance between their resources and the unique challenges set by the doctoral research and intensively work at the edge of their competences they are more likely to experience absorption. moreover, based on our results doctoral students‟ engagement in their doctoral work may be facilitated by shared meaning-making among doctoral students and supervisors regarding their goals for the doctorate and meanings of research work given by both students and supervisors. in practice, this can be supported, for instance, by encouraging supervisors and students to reveal and elaborate on their perceptions in supervisory discussions. such elaborations may provide a tool for and support supervisors and students to construct a shared understanding of the focus of supervision. supervisory discussions on the goals and perceptions of doctoral research are important especially at the beginning of the doctoral process when supervisory relationships are formed and students plan and get started with their doctoral projects. golde (1998), for instance, showed that one of the main reasons for doctoral students leaving their studies during the first year was a mismatch between the students‟ and supervisors‟ goals, expectations and practices. at the same time, there may be both individual and contextual variations. doctoral students, supervisors and other members of a scholarly community face more and less difficult times. there is also the reciprocal, continuously evolving relation between students and their environments in which engagement is constructed. it follows that both the students and scholarly community need to be constantly adjusting. the results of our study can be used both by students themselves for preparing themselves for the doctoral j. vekkaila et al. 28 | f l r process and considering meaningful and active participation strategies, and by supervisors and other doctoral educators for supporting their students‟ engagement in the best possible ways. designing more engaging learning environments for today‟s doctoral students is also an investment for the future academics and other knowledge workers. doctoral students‟ experiences of engagement are likely to have long-lasting effects. for instance, stubb et al. (2012a) demonstrated a relation between doctoral students‟ perceptions of their doctoral project, well-being and engagement. the results showed that participants who perceived their doctoral research as a process (e.g., learning and developing as a researcher) reported less stress, exhaustion, anxiety and lack of interest than students who perceived their research as a product (e.g., career qualification) or both as a process and product. moreover, those students who reported process related-meaning had less frequently considered interrupting their studies than others. accordingly, students‟ experiences of engagement during their doctoral process may function as a basis for their further engagement and well-being. keypoints engaging doctoral experience is rarely explored in both doctoral education and engagement literature. this study provided new insight in doctoral student engagement by breaking down the complexity of engagement by identifying qualitatively different experiences and sources of engagement. this study contributed to the literature on doctoral student engagement by opening the nature of engagement at the interfaces of studying and working by shedding light on the dual role of doctoral students as both students and professional researchers. the results encourage designing such engaging learning environments for doctoral students that promote their experiences of being competent researchers and integrated into their scholarly community. the relation between engagement in the doctoral process, the meanings of earning the doctorate given by the students, and development of a scholarly identity is worth of further investigation. acknowledgements the work has been supported by a grant from the finnish cultural foundation to the first author, grant 2106008 from the university of helsinki, and grant 1121207 from the academy of finland. references aitchison, c., & lee, a. (2006). research writing: problems and pedagogies. teaching in higher education, 11(3), 265–278. doi:10.1080/13562510600680574 alexander, a. (1992). domain knowledge: evolving themes and emerging concerns. educational psychologist, 27(1), 33–51. doi:10.1207/s15326985ep2701_4 austin, a. e. (2002). preparing the next generation of faculty: graduate school as socialization to the academic career. the journal of higher education, 73(1), 94–122. doi:10.1353/jhe.2002.0001 austin, a. e. (2010). expectations and experiences of aspiring and early career academics. in l. mcalpine & g. åkerlind (eds.), becoming an academic. international perspectives (pp. 18–44). united kingdom: palgrave macmillan. j. vekkaila et al. 29 | f l r appel, m., & dahlgren, l. (2003). swedish doctoral students‟ experiences on their journey towards a phd: obstacles and opportunities inside and outside the academic building. scandinavian journal of educational research, 47(1), 89–110. doi:10.1080/0031383032000033380 biglan, a. (1973a). the characteristics of subject matter in different academic areas. journal of applied psychology, 57(3), 195–203. doi:10.1037/h0034701 biglan, a. (1973b). relationship between subject matter characteristics and the structure and output of university departments. journal of applied psychology, 57(3), 204–213. doi:10.1037/h0034699 boud, d., & lee, a. (2005). peer learning as pedagogic discourse for research education. studies in higher education, 30(5), 501–516. doi:10.1080/03075070500249138 boud, d., & tennant, m. (2006). putting doctoral education to work: challenges to academic practice. higher education research & development, 25(3), 293–306. doi:10.1080/07294360600793093 bourner, t., bowden, r., & laing, s. (2001). professional doctorates in england. studies in higher education, 26(1), 65–88. doi:10.1080/03075070020030724 bresó, e., schaufeli, w. b., & salanova, m. (2011). can a self-efficacy-based intervention decrease burnout, increase engagement, and enhance performance? a quasi-experimental study. higher education, 61(4), 339–355. doi:10.1007/s10734-010-9334-6 brew, a., boud, d., & namgung, s. u. (2011). influences on the formation of academics: the role of the doctorate and structured development opportunities. studies in continuing education, 33(1), 51–66. doi:10.1080/0158037x.2010.515575 brint, s., cantwell, a. m., & hannerman, r. a. (2008). two cultures: undergraduate academic engagement. uc berkeley: center for studies in higher education. doi:10.1007/s11162-008-9090-y case, j. (2008). alienation and engagement: development of an alternative theoretical framework for understanding student learning. higher education, 55(3), 321–332. doi:10.1007/s10734-007-9057-5 chiang, k. (2003). learning experiences of doctoral students in uk universities. international journal of sociology and social policy, 23(1/2), 4–32. doi:10.1108/01443330310790444 coffey, a., & atkinson, p. (1996). making sense of qualitative data. complementary research strategies. thousand oaks, ca: sage. cox, j. w., & hassard, j. (2007). ties to the past in organization research: a comparative analysis of retrospective methods. organization, 14(4), 475–497. doi:10.1177/1350508407078049 csikszentmihalyi, m. (1990). flow: the psychology of optimal experience. new york: harper-perennial. deem, r., & brehony, k. j. (2000). doctoral students access to research cultures – are some unequal than others? studies in higher education, 25(2), 149–165. doi:10.1080/713696138 deci, e. l., & ryan, r. m. (2002). an overview of self-determination theory: an organismic-dialectical perspective. in l. deci & r. m. ryan (eds.), handbook of self-determination research (pp. 3–33). rochester: the university of rochester press. deci, e. l., & ryan, r. m. (2008). facilitating optimal motivation and psychological well-being across life‟s domains. canadian psychology, 49(1), 14–23. doi:10.1037/0708-5591.49.1.14 demerouti, e., bakker, a. b., nachreiner, f., & schaufeli, w. b. (2001). the job demands – resources model of burnout. journal of applied psychology, 86(3), 499–512. doi:10.1037/0021-9010.86.3.499 eccles, j. s. (2008). agency and structure in human development. research in human development, 5(4), 231–243. doi:10.1080/15427600802493973 edwards, a. (2005). relational agency: learning to be a resourceful practitioner. international journal of educational research, 43(3), 168–182. doi:10.1016/j.ijer.2006.06.010 fredricks, j. a., blumenfeld, p. c., & paris, a. h. (2004). school engagement: potential of the concept, state of the evidence. review of educational research, 74(1), 59–109. doi:10.3102/00346543074001059 gardner, s. k. (2007). “i heard it through the grapevine”: doctoral student socialization in chemistry and history. higher education, 54(5), 723–740. doi:10.1007/s10734-006-9020-x gardner, s. k., & barnes, b. j. (2007). graduate student involvement: socialization for the professional role. journal of college student development, 48(4), 1–19. doi:10.1353/csd.2007.0036 golde, c. m. (1998). beginning graduate school: explaining first-year doctoral attrition. new directions for higher education, 101, 55–64. doi:10.1002/he.10105 golde, c. m. (2005). the role of the department and discipline in doctoral student attrition: lessons from four departments. the journal of higher education, 76(6), 669–700. doi:10.1353/jhe.2005.0039 http://www.tandfonline.com/doi/abs/10.1080/03075070124819 http://dx.doi.org/10.1080/0158037x.2010.515575 http://web.ebscohost.com/ehost/viewarticle?data=dgjymppp44rp2%2fdv0%2bnjisfk5ie46bnkrquutrgk63nn5kx95uxxjl6oru%2btqk5jszaxurksueiulr9lporweezp33vy3%2b2g59q7sbwmtki0rlfks5zqeezdu33snoj6u9vmgktq33%2b7t8w%2b3%2bs7srattfgzrk8%2b5oxwhd%2fqu37z4uqm4%2b7y&hid=126 http://web.ebscohost.com/ehost/viewarticle?data=dgjymppp44rp2%2fdv0%2bnjisfk5ie46bnkrquutrgk63nn5kx95uxxjl6oru%2btqk5jszaxurksueiulr9lporweezp33vy3%2b2g59q7sbwmtki0rlfks5zqeezdu33snoj6u9vmgktq33%2b7t8w%2b3%2bs7srattfgzrk8%2b5oxwhd%2fqu37z4uqm4%2b7y&hid=126 j. vekkaila et al. 30 | f l r hakkarainen, k., hytönen, k., makkonen, j., seitamaa-hakkarainen, p., & white, h. (2013). interagency, collective creativity, and academic knowledge practices. in a. sannino & v. ellis (eds.), learning and collective creativity: activity-theoretical and socio-cultural studies. london, england: routledge, taylor & francis group. harry, b., sturges, k. m., & klingner, j. k. (2005). mapping the process: an exemplar of process and challenge in grounded theory analysis. educational researcher, 34(2), 3–13. doi:10.3102/0013189x034002003 hopwood, n. (2010). a sociocultural view of doctoral students‟ relationships and agency. studies in continuing education, 32(2), 103–117. doi:10.1080/0158037x.2010.487482 hoskins, c. m., & goldberg, a. d. (2005). doctoral student persistence in counselor education programs: student-program match. counselor education and supervision, 44(3), 175–188. doi:10.1002/j.15566978.2005.tb01745.x hyun, j. k., quinn, b. c., madon, t., & lustig, s. (2006). graduate student mental health: needs assessment and utilization of counseling services. journal of college student development, 47(3), 247–266. doi:10.1353/csd.2006.0030 inkinen, m., lonka, k., hakkarainen, k., muukkonen, h., litmanen, t., & salmela-aro, k. (2013). the interface between core affects and the challenge–skill relationship. journal of happiness studies. an interdisciplinary forum on subjective well-being. doi:10.1007/s10902-013-9455-6 international postgraduate student mirror. (2006). catalonia, finland, ireland and sweden. högskoleverket, swedish national agency for higher education, 29r. retrieved from http://www.ub.edu/depdibuix/ir/0629r-shv_se-catalonia.pdf ives, g., & rowley, g. (2005). supervisor selection or allocation and continuity of supervision: ph.d. students‟ progress and outcomes. studies in higher education, 30(5), 535–555. doi:10.1080/03075070500249161 krause, k-l., & coates, h. (2008). students‟ engagement in first-year university. assessment & evaluation in higher education, 33(5), 493–505. doi:10.1080/02602930701698892 kurtz-costes, b., helmke, a. l., & ülkü-steiner, b. (2006). gender and doctoral studies: the perceptions of phd students in an american university. gender & education, 18(2), 137–155. doi:10.1080/09540250500380513 kvale, s. (2007). doing interviews. london: sage publications. leiter, m. p., & bakker, a. b. (2010). work engagement: an introduction. in a. b. bakker & m. p. leiter (eds.), work engagement: a handbook of essential theory and research (pp. 1–9). london and new york: psychology press. lindblom-ylänne, s., & lonka, k. (2000). interaction between learning environment and expert learning. lifelong learning in europe, 5(2), 90–97. lindblom-ylänne, s., trigwell, k., nevgi, a., & ashwin, p. (2006). how approaches to teaching are affected by discipline and teaching context. studies in higher education, 31(3), 285–298. doi:10.1080/03075070600680539 llorens, s., schaufeli, w., bakker, a., & salanova, m. (2007). does a positive gain spiral of resources, efficacy beliefs and engagement exist? computers in human behavior, 23(1), 825–841. doi:10.1016/j.chb.2004.11.012 lonka, k. (1997). explorations of constructive processes in student learning. a doctoral dissertation. helsinki: university press. lonka, k. (2003). helping doctoral students to finish their theses. in l. björk, g. bräuer, l. rienecker, g. ruhmann, & p. stray jørgensen (eds.), teaching academic writing across europe (pp. 113–131). dordrecht, the netherlands: kluwer university press. lonka, k., & lindblom-ylänne, s. (1996). epistemologies, conceptions of learning, and study practices in medicine and psychology. higher education, 31(1), 5–24. doi:10.1007/bf00129105 lovitts, b. e. (2001). leaving the ivory tower: the causes and consequences of departure from doctoral study. lanham, md: rowman and littlefield. lovitts, b. e., & nelson, g. (2000). the hidden crisis in graduate education: attrition from ph.d. programs. academe, 86(6), 44–50. doi:10.2307/40251951 j. vekkaila et al. 31 | f l r löfström, e., & pyhältö, k. (2012). the supervisory relationship as an arena for ethical problem-solving. education research international. doi:10.1155/2012/961505 mason, m. m. (2012). motivation, satisfaction, and innate psychological needs. international journal of doctoral studies, 7, 259–277. retrieved from http://ijds.org/volume7/ijdsv7p259277mason0345.pdf mcalpine, l., & amundsen, c. (2008). academic communities and developing identity: the doctoral student journey. in p. richards (ed.), global issues in higher education (pp. 57–83), ny: nova publishing. mcalpine, l., & norton, j. (2006). reframing our approach to doctoral programs: an integrative framework for action and research. higher education research & development, 25(1), 3–17. doi:10.1080/07294360500453012 mccune, v., & hounsell, d. (2005). the development of students‟ ways of thinking and practicing in three final-year biology courses. higher education, 49(3), 255–289. doi:10.1007/s10734-004-6666-0 meyer, j. h. f., shanahan, m. p., & laugksch, r. c. (2005). students‟ conceptions of research i: a qualitative and quantitative analysis. scandinavian journal of educational research, 49(3), 225–244. doi:10.1080/00313830500109535 mills, j., bonner, a., & francis, k. (2006). the development of constructivist grounded theory. international journal of qualitative methods, 5(1), 25–35. retrieved from http://ejournals.library.ualberta.ca/index.php/ijqm/article/view/4402/3795 morgan, d. l. (2007). paradigms lost and pragmatism regained: methodological implications of combining qualitative and quantitative methods. journal of mixed methods research, 1(1), 48–76. doi:10.1177/2345678906292462 mäkinen, j., olkinuora, e., & lonka, k. (2004). students at risk: students‟ general study orientations and abandoning/prolonging the course of studies. higher education, 48(2), 173–188. doi:10.1023/b:high.0000034312.79289.ab nettles, m. t., & millet, c. m. (2006). three magic letters: getting to ph.d. baltimore: the john hopkins university press. niemic, c. p., & ryan, r. m. (2009). autonomy, competence, and relatedness in the classroom. applying self-determination theory to educational practice. theory and research in education, 7(2), 133–144. doi:10.1177/1477878509104318 ouweneel, e., le blanc, p. m., & schaufeli, w. b. (2011). flourishing students: a longitudinal study on positive emotions, personal resources, and study engagement. the journal of positive psychology: dedicated to furthering research and promoting good practice, 6(2), 142–153. doi:10.1080/17439760.2011.558847 overall, n. c., deane, k. l., & peterson, e. r. (2011). promoting doctoral students‟ research self-efficacy: combining academic guidance with autonomy support. higher education research & development, 30(6), 791–805. doi:10.1080/07294360.2010.535508 park, c. (2005). new variant phd: the changing nature of the doctorate in the uk. journal of higher education policy and management, 27(2), 189–208. doi:10.1080/13600800500120068 patton, m. q. (1990). qualitative research and evaluation methods (2nd ed.). newbury park, ca: sage publications. pontius, j. l., & harper, s. r. (2006). principles for good practice in graduate and professional student engagement. new directions for student services, 115, 47–58. doi:10.1002/ss.215 pyhältö, k., & keskinen, j. (2012). doctoral students‟ sense of relational agency in their scholarly communities. international journal of higher education, 1(2), 136–149. doi:10.5430/ijhe.v1n2p136 pyhältö, k., nummenmaa, a. r, soini, t., stubb, j., & lonka, k. (2012a). research on scholarly communities and development of scholarly identity in finnish doctoral education. in s. ahola & d. m. hoffman (eds.), higher education research in finland. emerging structures and contemporary issues (pp. 337–357). jyväskylä: jyväskylä university press. pyhältö, k., stubb, j., & lonka, k. (2009). developing scholarly communities as learning environments for doctoral students. international journal for academic development, 14(3), 221–232. doi:10.1080/13601440903106551 j. vekkaila et al. 32 | f l r pyhältö, k., stubb. j., & tuomainen, j. (2011). international evaluation of research and doctoral education at the university of helsinki to the top and out to society. summary report on doctoral students‟ and principal investigators‟ doctoral training experiences. retrieved from http://wiki.helsinki.fi/display/evaluation2011/survey+on+doctoral+training pyhältö, k., vekkaila, j., & keskinen, j. (2012b). exploring the fit between doctoral students‟ and supervisors‟ perceptions of resources and challenges vis-à-vis the doctoral journey. international journal of doctoral studies, 7, 395–414. retrieved from http://ijds.org/volume7/ijdsv7p395414pyhalto383.pdf reeve, j., jang, h., carrell, d., jeon, s., & barch, j. (2004). enhancing students‟ engagement by increasing teachers‟ autonomy support. motivation and emotion, 28(2), 147–169. doi:10.1023/b:moem.0000032312.95499.6f reeve, j., & tseng, c-h. (2011). agency as a fourth aspect of students‟ engagement during learning activities. contemporary educational psychology, 36(4), 257–354. doi:10.1016/j.cedpsych.2011.05.002 sainio, j. (2010). asiantuntijana työmarkkinoille vuosina 2006 ja 2007 tohtorin tutkinnon suorittaneiden työllistyminen ja heidän mielipiteitään tohtorikoulutuksesta [experts for the labour market the employment of doctors who earned their doctoral degree in 2006-2007 and their perceptions of doctoral training]. tampere: kirjapaino hermes oy. retrieved from http://www.aarresaari.net/pdf/asiantuntijana_tyomarkkinoille_netti.pdf salanova, m., schaufeli, w., martínez, i., & bresó, e. (2010). how obstacles and facilitators predict academic performance: the mediating role of study burnout and engagement. anxiety, stress & coping: an international journal, 23(1), 53–70. doi:10.1080/10615800802609965 schaufeli, w. b., & bakker, a. b. (2004). job demands, job resources, and their relationship with burnout and engagement. journal of organizational behavior, 25(3), 293–315. doi:10.1002/job.248 schaufeli, w. b., martínez, i. m., pinto, a. m., salanova, m., & bakker, a. b. (2002a). burnout and engagement in university students. journal of cross-cultural psychology, 33(5), 464–481. doi:10.1177/0022022102033005003 schaufeli, w. b., salanova, m., gonzález-romá, v., & bakker, a. b. (2002b). the measurement of engagement and burnout: a two sample confirmatory factor analytic approach. journal of happiness studies, 3(1), 71–92. doi:10.1023/a:1015630930326 stubb, j. (2012). becoming a scholar: the dynamic interaction between the doctoral student and the scholarly community (doctoral dissertation). university of helsinki, faculty of behavioural sciences, department of teacher education, research report 336. retrieved from http://urn.fi/urn:isbn:978952-10-6867-6 stubb, j., pyhältö, k., & lonka, k. (2011). balancing between inspiration and exhaustion: phd students‟ experienced socio-psychological well-being. studies in continuing education, 33(1), 33–50. doi:10.1080/0158037x.2010.515572 stubb, j., pyhältö, k., & lonka, k. (2012a). the experienced meaning of working with a phd thesis. scandinavian journal of educational research, 56(4), 439–456. doi:10.1080/00313831.2011.599422 stubb, j., pyhältö, k., & lonka, k. (2012b). conceptions of research: the doctoral student experience in three domains. studies in higher education. 1–14. ifirst article. doi:10.1080/03075079.2011.651449 toews, j. a., lockyer, j. m., dobson, d. j. g., simpson, e., brownell, a. k. w., brenneis, f., macpherson, k. m., & cohen, g. s. (1997). analysis of stress levels among medical students, residents, and graduate students at four canadian schools of medicine. academic medicine, 72(11), 997–1002. retrieved from http://journals.lww.com/academicmedicine/abstract/1997/11000/analysis_of_stress_levels_among_m edical_students,.19.aspx turner, g., & mcalpine, l. (2011). doctoral experience as researcher preparation: activities, passion, status. international journal for researcher development, 2(1), 46–60. doi:10.1108/17597511111178014 umbach, p. d., & wawrzynski, m. r. (2005). faculty do matter: the role of college faculty in student learning and engagement. research in higher education, 46(2), 153–184. doi:10.1007/s11162-0041598-1 http://wiki.helsinki.fi/display/evaluation2011/survey+on+doctoral+training http://www.aarresaari.net/pdf/asiantuntijana_tyomarkkinoille_netti.pdf http://www.tandfonline.com/action/dosearch?action=runsearch&type=advanced&result=true&prevsearch=%2bauthorsfield%3a%28mart%c3%adnez%2c+isabel%29 http://www.tandfonline.com/action/dosearch?action=runsearch&type=advanced&result=true&prevsearch=%2bauthorsfield%3a%28bres%c3%b3%2c+edgar%29 http://www.tandfonline.com/loi/gasc20?open=23#vol_23 http://www.tandfonline.com/action/dosearch?action=runsearch&type=advanced&result=true&prevsearch=%2bauthorsfield%3a%28mart%c3%adnez%2c+isabel%29 j. vekkaila et al. 33 | f l r vassil, k., & solvak, m. (2012). when failing is the only option: explaining failure to finish phds in estonia. higher education, 64(4), 503–516. doi:10.1007/s10734-012-9507-6 vekkaila, j., pyhältö, k., hakkarainen, k., keskinen, j., & lonka, k. (2012). doctoral students‟ key learning experiences in the natural sciences. international journal for researcher development, 3(2), 154–183. doi:10.1108/17597511311316991 vekkaila, j., pyhältö, k., & lonka, k. (2013). experiences of disengagement – a study of doctoral students in the behavioral sciences. international journal of doctoral studies, 8, 61–81. retrieved from http://ijds.org/volume8/ijdsv8p061-081vekkaila0402.pdf vermunt, j. (1996). metacognitive, cognitive and affective aspects of learning styles and strategies: a phenomenographic analysis. higher education, 31(1), 25–50. doi:10.1007/bf00129106 virtanen, v., & pyhältö, k. (2012). what engages doctoral students in biosciences in doctoral studies? psychology, 3(12a), 1231–1237. doi:10.4236/psych.2012.312a182 weidman, j. c., & stein, e. l. (2003). socialization of doctoral students to academic norms. research in higher education, 44(6), 641–656. doi:10.1023/a:1026123508335 white, j., & nonnamaker, j. (2008). belonging and mattering. how science doctoral students experience community. naspa journal, 45(3), 350–372. doi:10.2202/1949-6605.1860 zhao, c., & kuh, g. d. (2004). adding value: learning communities and student engagement. research in higher education, 45(2), 115–138. doi:10.1023/b:rihe.0000015692.88534.de appendix 1 doctoral student interview discipline or subject: been as a phd student since: i‟m doing a monograph/collection of articles: i‟m female/male: i‟m doing my thesis full-time/ part-time: phase of my study: 1. how did you become a phd student? what is the topic of your phd work? how did you come up with this topic? does it relate to the work of others in your group? 2. what motivates you to do your phd research? 3. describe in your own words, how has your phd process gone so far? 4. describe some situation, event or episode from your phd studies that has really influenced your own thoughts about doing phd research or something else related to that. what happened? why? what did you think of and how did you feel? 5. at the moment, do you have some question/challenge that you are wondering about? if so, what? why? 6. what is the most enjoyable thing in postgraduate studies? what is the hardest? 7. describe a situation that gave you inspiration. what happened? why do you think it happened? what did you do, think and feel? describe a situation in your phd process that was in some way negative. what happened? why do you think it happened? what did you do, think and feel? 8. what kind of supervision have you gotten in your phd process? what kind of supervision would you hope for? 9. do you get support to your work from somewhere else? what kind of support? would you need something more? 10. describe a situation in your phd process where you felt that your supervisor especially succeeded. what happened and why was that situation meaningful to you? 11. what kind of role do other researchers and phd students have in your process? 12. in your opinion, how should postgraduate education be developed? http://web.ebscohost.com/ehost/viewarticle?data=dgjymppp44rp2%2fdv0%2bnjisfk5ie46bnkrquutrgk63nn5kx95uxxjl6oru%2btqk5jszaxuq6muemylr9lporweezp33vy3%2b2g59q7rbgotvg0qbrqtzzqeezdu33snoj6u9vmgktq33%2b7t8w%2b3%2bs7t7etsem2r7e%2b5oxwhd%2fqu37z4uqm4%2b7y&hid=126 http://web.ebscohost.com/ehost/viewarticle?data=dgjymppp44rp2%2fdv0%2bnjisfk5ie46bnkrquutrgk63nn5kx95uxxjl6oru%2btqk5jszaxuq6muemylr9lporweezp33vy3%2b2g59q7rbgotvg0qbrqtzzqeezdu33snoj6u9vmgktq33%2b7t8w%2b3%2bs7t7etsem2r7e%2b5oxwhd%2fqu37z4uqm4%2b7y&hid=126 http://www.springerlink.com/content/0361-0365/ http://www.springerlink.com/content/0361-0365/ j. vekkaila et al. 34 | f l r 13. what kind of advice would you give to a student who is considering phd studies? why? 14. is there still something you would like to tell? 15. what would you have wished to be asked about? frontline learning research 1 (2014) 1-21 issn 2295-3159 corresponding author: hilde haider, department of psychology, university of cologne, richard-strauss-str. 2 50931 cologne, germany, phone: +49-221-4704719, email: hilde.haider@uni-koeln.de http://dx.doi.org/10.14786/flr.v2i1.37 1 | f l r how we use what we learn in math: an integrative account of the development of commutativity hilde haider a , alexandra eichler a , sonja hansen a , bianca vaterrodt b , robert gaschler c , peter a. frensch b a university of cologne, germany b humboldt-university berlin, germany c university koblenz-landau, germany article received 28 may 2013 / revised 6 january 2014/ accepted 16 january 2014 / available online 27 january 2014 abstract one crucial issue in mathematics development is how children come to spontaneously apply arithmetical principles (e.g. commutativity). according to expertise research, well-integrated conceptual and procedural knowledge is required. here, we report a method composed of two independent tasks that assessed in an unobtrusive manner the spontaneous use of procedural and conceptual knowledge about commutativity. this allowed us to ask (1) in which grade students spontaneously apply this principle in different task formats and (2) in which grade they start to possess an integrated concept of the commutativity. procedural and conceptual knowledge of 8 to 9 year olds (163 second and 180 third graders) as well as 46 adult students was assessed independently and without any hint concerning commutativity. results indicated procedural as well as conceptual knowledge about commutativity for second graders. however, their procedural and conceptual knowledge was unrelated. an integrated relation between the two measures first emerged with some of the third graders and was further strengthened for adult students. keywords: conceptual knowledge; procedural knowledge; commutativity; integrated concept. h. haider et al. 2 | f l r 1. introduction one major skill in mathematics is the acquisition of adaptive expertise. that is, students should be able to deliberately recognize those constraints that allow to apply a certain mathematical principle (e.g., torbeyns, de smedt, ghesquière & verschaffel, 2009; verschaffel, luwel, torbeyns & van dooren, 2009). for instance, in the pisa mathematical literacy test students have to spontaneously apply mathematical principles in order to solve problems in real-world contexts. given the important role of self-guided learning and performance in the development of mathematical abilities and concepts, some recent studies have started to focus on spontaneous recognition of mathematical aspects in natural surroundings (e.g., hannula, & lehtinen, 2005; hannula, lepola, & lehtinen, 2010; mcmullen, hannula-sormunen, & lehtinen, 2011). an important question with regard to adaptive expertise is how students come to recognize that they can use a certain principle in order to facilitate calculation. or to put it in other words, what kind of knowledge underlies the ability to adaptively apply a mathematical principle spontaneously whenever it facilitates calculation? in the research on adaptive expertise, it is widely accepted that this ability is not only based on procedural knowledge (knowing how to apply a certain strategy), but also on conceptual knowledge (knowing when and why a certain principle applies). procedural and conceptual knowledge should be integrated and the resulting knowledge base should be abstract enough to ensure flexibility in knowledge application (e.g., anderson & schunn, 2000; baroody, 2003; gentner & toupin, 1986; haider & frensch, 1996; koedinger & anderson, 1990; star, 2005; verschaffel et al., 2009). as one example of linking concepts and procedures, this research has revealed that conceptual knowledge is important to guide attention to task relevant information in order to solve problems (e.g., baroody & rosu, 2006). an abstract conceptual understanding might also be of particular importance when knowledge has to be transferred from one domain to another (e.g., goldstone & sakamoto, 2003; kaminsky, sloutsky & heckler, 2008). it supports flexible shortcut application when problems to which a principle applies are presented mixed with problems to which the principle does not apply. for instance, siegler and stern (1998) have shown that second graders relied less on inversion-based procedures when inversion problems (a + b – b) were randomly interspersed with control problems (a + b – c) compared to blocked presentation. mixed presentation of inversion and control problems hindered the use of inversion short-cuts. this suggests that younger children do not deliberately recognize the constraints important for applying the inversion principle. rather, they simply seem to know that the strategy applies for a certain class of problems. likely they cannot rely on a well integrated, abstract understanding of the inversion principle (i.e., they have not yet developed adaptive expertise). but, how and when do children develop an integrated representation of basic mathematical principles? the first goal of the current study was to develop a method to unobtrusively measure the spontaneous application of procedural and conceptual knowledge taking the commutativity principle as a test case. the second goal was to shed some light on the development of an abstract and well integrated representation of the commutativity principle. with regard to the second goal, we pursued to different questions: (1) in which grade are students able to spontaneously apply commutativity knowledge in different task formats and (2) in which grade starts performance expressed in the different tasks to correlate with oneanother? for this purpose, we investigated the deliberate use of the commutativity principle in two different situations. the term “deliberate” means that children did not receive any hint about the commutativity principle at all (e.g., torbeyns et al., 2009). in the first test children simply solved addition problems that sometimes allowed for a shortcut based on the commutativity principle (procedural knowledge). the second test was aimed at assessing conceptual knowledge. children were instructed to mark – without solving the problem – those problems that they believed could be solved without calculation. hence, this task required children to realize that the order of addends does not change cardinality. the correlation between these two independent tasks allowed us to gauge how well integrated children’s knowledge was. focusing on just two unobtrusive measures (one procedural and one conceptual), the current work can potentially lay the ground to develop multi-method approaches in the same spirit, safeguarding that multiple testing does not cue participants towards what the test situation is about. h. haider et al. 3 | f l r we focused on the commutativity principle as it is one of the most basic properties in mathematics. it refers to the principle that changing the order of operands in addition and multiplication does not change the end result. it is known as a fundamental property of many binary operations. the commutativity of simple operations, such as the multiplication and addition of numbers are usually acquired throughout elementary school. however, many mathematical proofs also depend on this property. 1.1 development of procedural and conceptual knowledge about commutativity former research in the field of developmental psychology has already shown that children acquire informal knowledge of commutativity as an arithmetic principle long before they enter school (e.g., baroody & gannon, 1984; baroody, ginsburg & waxman, 1983; canobi, reeve & pattison, 1998, 2002; cowan & renton, 1996; resnick, 1992; siegler & jenkins, 1989; sophian, harley & manos martin, 1995). one potential reason for this early development is that at least the core property of commutativity, the orderirrelevance principle, applies to many non-numerical situations. for example, children may experience that some tasks require a certain sequence (e.g., putting on one’s clothes), whereas others do not (e.g., laying the table). already toddlers have many opportunities to learn that order does not affect the end result in some situations, but does in others. order-irrelevance is also a core principle for counting (e.g., gelman, 1990; gelman & gallistel, 1978). learning to count requires children to learn, on the one hand, that the sequence of number words is relevant. on the other hand, the sequence in which the objects are counted is irrelevant. consequently, briars and siegler (1984) found that children need time to understand order-irrelevance in counting. furthermore, counting is the dominant skill through which preschool children learn to map concrete objects to numbers. also, counting is one of the important precursors of addition. through counting, pre-school children can learn order-irrelevance in a numerical manner before entering school. they thus do not only have the chance to understand order-irrelevance in a non-numerical manner. however, even though considerable interest in research on counting and addition principles emerged already in the 1980s and still continues (e.g., baroody, 1984; baroody & gannon, 1984; baroody et al., 1983; briars & siegler, 1984; canobi, reeve & pattison, 1998, 2002, 2003; fuson, 1988; gelman & gallistel, 1978; gelman & meck, 1983; resnick, 1992; sophian & adams, 1987; starkey & gelman, 1982), the central question has not been solved yet: how and when do children acquire integrated knowledge representations, in the sense of true formal arithmetic principles? for instance, geary (2006) stated that it is not clear when children “explicitly understand commutativity as a formal arithmetical principle” (p. 791). on the one hand, the difficulties in answering this question are due to the fact that researchers by no means agree upon the characteristics of procedural or conceptual knowledge that must be given in order to conclude that children possess an abstract mathematical concept (cf., star, 2004). concerning procedural knowledge, most researchers agree that it refers to the ability to apply a certain strategy when performing a mathematical task (e.g., hiebert & lefevre, 1986). conceptual knowledge or metastrategic competences (kuhn, garcia-mila, zohar & andersen, 1995) often is assumed to refer to children’s explicit understanding of a certain principle (i.e., why and when it is allowed to a apply a certain strategy; e.g., baroody, feil & johnson, 2007; hiebert & lefevre, 1986; rittle-johnson, siegler & alibali, 2001). on the other hand, there is no consensus how best to assess procedural and conceptual knowledge. one frequently used approach to measure procedural knowledge is to ask children to solve addition problems and afterwards have them explain their strategies (e.g., baroody & gannon, 1984; baroody, ginsburg & waxman, 1983; bisanz & lefevre, 1992; canobi et al., 1998, 2002, 2003; cowan & renton, 1996). conceptual knowledge in these studies has, for example, been assessed by letting children observe a puppet solving problem pairs (see, e.g. baroody et al., 1983; canobi et al, 1998). if a child on enquiry stated that the puppet could know the answer to the second problem from looking at the previous one, he or she was asked for reasons and eventually prompted for more detailed explanations. this form of assessment implies that children are being informed about the underlying arithmetic principle – at least they are made aware that different efficient strategies are applicable. such procedural and the conceptual knowledge tests might guide h. haider et al. 4 | f l r children’s attention to the task-relevant information. also, it is conceivable that they look at the problems more attentively when they are asked to verbalize their strategies. consequently, conclusions concerning the question whether a child possesses abstract conceptual knowledge may vary depending on the tests that were applied. based on the above-mentioned forms of assessment, the empirical research on commutativity suggests that conceptual and procedural knowledge in this domain are moderately related (e.g., baroody et al., 1983; canobi, 2004; canobi et al., 1998). however, the findings do not allow to exclude that the acquired conceptual knowledge of first or even second graders is still domain-specific rather than akin to an abstract concept representing the formal arithmetic principle of commutativity (e.g., bisanz, watchhorn, piatt & sherman, 2009; geary, hoard, byrd-craven & desoto, 2004; lefevre et al., 2006). therefore, investigating the spontaneous application of commutativity knowledge would complement and broaden this research. in summary, the goal of our study was twofold: first, we aimed to develop a method to unobtrusively test for spontaneous application of procedural and conceptual commutativity knowledge. our second goal was to investigate the degree of integration of this spontaneously expressed procedural and conceptual knowledge of second and third graders. additionally, for means of comparison, we also tested adult students. as described above, knowledge about a mathematical principle like, for instance, the commutativity principle, can be said to represent an integrated or abstract concept in the sense of a true formal mathematic principle when learners are able to apply their knowledge whenever task constraints permit. that is, learners should be able to deliberately recognize task properties allowing them to apply the mathematical principle irrespectively of task context. 2. method 2.1. general method we investigated the commutativity principle with three-element addition problems (i.e., 5+3+7 = ?) 1 . these three-element problems are unfamiliar at least for younger students. since we wanted to investigate whether or not students would recognize the applicability of the commutativity principle without any further information about this principle, we needed less familiar problems. therefore, we accepted that threeelement problems implicitly presuppose knowledge about associativity (e.g., canobi et al., 1998). the three-element problems were always presented in blocks, one problem beneath the other. unbeknownst to the participating students, some problems consisted of identical addends in a different order as the preceding problem, and thus could be solved without calculation (commutative problems, hereafter). students received two different and completely independent task formats. the first task, the arithmetic task, consisted of two blocks. one block contained interspersed commutative problems, the other one did not. participants did not receive any information about the existence of these commutative problems. rather, they were simply asked to solve the two blocks of addition problems as fast and accurately as possible. if students are faster when working on the block that includes three-element commutative problems as compared to the block which does not contain such shortcut options, they can be said to possess procedural knowledge about commutativity. 1 some researchers use the term associativity instead of commutativity when an addition or multiplication problem has more than two addends or factors (geary et al., 2008). other researchers (canobi, et al., 1998) refer to commutativity as the property that problems containing the same terms in a different order have the same answer independent of the number of terms, whereas associativity is the property that problems in which terms are decomposed and recombined in different ways have the same answer [(a + b) + c = a + (b + c)]. h. haider et al. 5 | f l r in the second task, the so-called judgment task, students were instructed to identify those problems which they believed need no calculation. they explicitly were told to refrain from calculating any problems. if students understand that the order of identical addends does not change the cardinality, they should be able to correctly mark the commutative problems. by virtue of this second task type, we were able to assess conceptual knowledge about commutativity without cueing the concept. importantly and in contrast to former experiments (e.g., baroody et al., 1983; canobi, 2005), participants in our experiment did not receive any hint about the existence of commutative problems in either task. that is, they were not instructed to further explain their strategies. the rationale behind this procedure was that any instruction to think about the strategies used to solve the problems might trigger active search for regularities, thereby making it impossible to assess the spontaneously activated concept of commutativity. if students possess an abstract understanding of the commutativity principle, they should be able to recognize and rely on the relevant task characteristics in any task context and without any hint (e.g., bisanz et al., 2009; prather & alibali, 2009). to the extent that children have acquired an abstract concept of commutativity, performance should correlate between both of our two tasks reflecting knowledge about this principle. likely procedural and conceptual knowledge about commutativity becomes iteratively more integrated in the first years of primary school. we should thus find that the relation between procedural and conceptual knowledge is stronger in third graders as compared to second graders (e.g., lefevre et al. 2006). 2.2 participants overall, 163 second graders (79 girls) with a mean age of 8 years 1 month (sd = 7 months), 180 third graders (91 girls) with a mean age of 9 years 1 month (sd = 8 months) participated in the study. as a control condition, we also collected data of 46 students of the university of cologne (37 women) with a mean age of 23.6 years (sd = 5.2). children were recruited from six different elementary schools located in middle socioeconomic status suburbs of cologne. all children had their parents’ or guardians’ permission to participate in the study. 2.3 procedure and materials the study consisted of two parts. in the first part, participants received the arithmetic task: one block with interspersed repetitions of addends in changed order in consecutive problems (commutativity block) and one block without such repetitions (control block). in the second part, participants were administered the judgment task. both tasks were designed as paper-pencil tests and children and adult students were tested in groups of up to 25 participants in a classroom-like setting. we generated three sets of 30 arithmetic problems with three addends between 2 and 9 (e.g., 3 + 6 + 8 = ?; maximum result was 24; 1 as an addend was not included). the problems in all three sets yielded at least approximately the same totals and within a problem each numeral could only occur once. the 30 problems of each of the two blocks were distributed over five pages with six problems on each page. in the commutativity block, each page contained two pairs of commutative problems (i.e., one problem and its repetition with a different order of addends). in the control block no such commutative pairs occurred. instead participants received pairs of control problems which yielded the same results but were composed of different addends. in both blocks, participants were instructed to calculate the problems page by page from top to bottom. the judgment task consisted of overall 30 problems with 10 problems per page. on each page, three pairs were commutative pairs and the remaining four problems were filler problems. the first page was for practice only. participants were instructed to first solve all 10 problems from top to bottom on the page. afterwards they were asked to mark those problems that needed no calculation on that page. in particular, they were told that some of the problems need no calculation and that they should figure out for which of these problems they could have written down the result without calculation. after this practice page, h. haider et al. 6 | f l r participants were instructed to only judge on the next pages whether or not they needed to calculate the result for a problem without actually attempting to solve it. therefore, all problems on pages 2 and 3 were presented without equal sign. instead, there was a circle to the right of each problem and participants were told to mark this circle when they believed they did not need to calculate the result. again, students were instructed to work on the problems from the top to the bottom of each page. table 1 depicts examples of the problems in each of the two arithmetic blocks and the judgment task. table 1 examples of the problems presented on one page in the two arithmetic blocks (commutativity and control block) and the judgment task arithmetic task judgment task commutativity block control block 3 + 5 + 4 = 4 + 9 + 8 = 4 + 8 + 9 = 6 + 2 + 5 = 9 + 7 + 2 = 2 + 7 + 9 = 5 + 3 + 4 = 8 + 9 + 4 = 6 + 7 + 8 = 5 + 2 + 6 = 2 + 7 + 9 = 9 + 4 + 5 = 2 + 7 + 9 9 + 5 + 4 2 + 6 + 5 6 + 5 + 2 8 + 7 + 5 3 + 5 + 6 6 + 5 + 3 2 + 9 + 5 6 + 7 + 9 9 + 6 + 7           problems in bold indicate the commutative pairs of the respective task each of the two arithmetic blocks was administered as a separate booklet, as was the judgment task. students only worked with a pencil and were not allowed to use an eraser. rather, to increase the reliability of the timing measure, they were told to cross out any errors and to write the correct answer right beside the problem. an experimenter instructed all participants in the classroom. the experiment started with six arithmetic practice problems with three addends. the only goal of this phase was to familiarize the children with the task requirements. students were given 2 minutes to solve these six warm-up problems. then, the first of the two arithmetic blocks was presented. approximately half of the children (second graders and third graders) and all adults in the control condition received the commutativity block first, followed by the control block. the remaining participants started with the control block and subsequently received the commutativity block. the time limit was set to 3 minutes per block (1 minute for adult students) with a 1-minute break between blocks. one minute after having finished the second arithmetic block, the judgment task was presented. participants were allowed 2 minutes (adult students again 1 minute) to calculate the problems on the practice page. afterwards they had the same amount of time for marking those problems they believed they could h. haider et al. 7 | f l r have solved without calculation. after the practice phase, the same time limit was applied for the two subsequent pages, so that time did not suffice to calculate the problems and to concurrently mark those problems requiring no calculation. in addition, up to four additional experimenters observing small groups of children (up to six) ensured that they were not calculating the problems. after this last block, all children received some sweets. adult students in the control condition were debriefed about the study. 2.4. design independent variables were grade (second versus third graders) and block type (commutativity versus control block in the arithmetic task). dependent variables in the arithmetic problem blocks were calculation time per problem in each of the two arithmetic blocks, as well as the number of correct results. calculation time was computed separately for each participant and each of the two arithmetic problem blocks by dividing the individual number of completed problems by the total time given for the block in the respective age group (three minutes for second and third graders; 1 minute for adults). for the judgment task, the dependent variables were relative number of hits (correctly identified commutative problems) and false alarms (problems incorrectly identified as commutative problems), as well as the sensitivity index d’ from signal detection theory (i.e., the difference between z-transformed hit rate and false alarms rate). 2.5 split-half reliability in order to check if our measures are reliable, we computed split-half reliabilities for each task type. that is, for each of the two age groups and the adults, we calculated correlations between the two arithmetic blocks (control and commutativity block) and between the second and third pages of the judgment task (the practice page of the judgment task was excluded). table 2 shows the spearman-brown corrected correlation coefficients separately for each age group and each task format (arithmetic task and judgment task). table 2 spearman-brown corrected correlation coefficients for the arithmetic task and the judgment task for all participants and separately for the three age groups (arithmetic task: correlation between the amount of computed problems in the commutativity block and the control block; judgment task: correlation between correct responses on the first and on the second pages of the test) arithmetic task amount of computed commutative problems all participants grade 2 grade 3 adults amount of computed control problems .90 .88 .88 .95 judgment task amount of correct judgments on the first page all participants grade 2 grade 3 adults amount of correct judgments on the second page .82 .83 .78 .86 as can be seen from table 2, the correlation coefficients in each age group ranged between r = .78 and r = .95. thus, the two tasks used to assess participants’ procedural and conceptual knowledge seem to be reliable measures. 3. results second or third graders were excluded from further analyses if they completed less than 16 problems across the two arithmetic problem blocks (i.e., 2 standard deviations below the group means; 15 second graders, 12 third graders, and 1 adult). they were also excluded from further analyses if they solved all 30 h. haider et al. 8 | f l r problems in the control and the commutativity block, as calculation times could not be calculated for these participants (2 second graders, 23 third graders, and 8 adults). this led to 146 second graders, 145 third graders, and 37 adult students in the control condition. the following result section is divided into three parts. we first describe the results for the arithmetic problem blocks. second, we report the performance in the judgment task. lastly, we analyze the relation between these two tasks. 3.1 arithmetic task as a preliminary analysis did not reveal substantial effects of the order of presentation (commutative problem first followed by control problem or vice versa), we collapsed the data for all participants within the groups of second and third graders. table 3 depicts the calculation times for problems in the commutativity and the control blocks per age group. mean calculation times suggest that second and third graders benefitted from the commutative problems whereas adult students did not. table 3 calculation times per task in the commutativity and the control block for each age group. the table holds the means and standard deviations for the different age group in seconds as well as lower and upper limit of the 95-% confidence interval (ci; loftus & masson, 1994) age group commutativity block control block n m (sd) m±95%ci m (sd) m±95%ci grade 2 12.33 (3.74) 12.09 12.66 13.28 (4.39) 12.95 13.60 146 grade 3 9.55 (2.51) 9.34 9.76 9.93 (3.12) 9.72 10.14 145 adults 3.79 (1.04) 3.66 3.92 3.63 (1.04) 3.50 3.76 37 a 2 (age group) x 2 (block type: commutativity block vs. control block) analysis of variance (anova) with calculation time as dependent variable revealed significant main effects of age group (f[1, 289] = 66.23, mse = 20.64, p .23), and of block type (f[1, 289] = 15.94, mse = 4.04, p = .06). the interaction between age group and block type was close to significance (f[1, 289] = 2.98, mse = 4.04, p = .088). planned contrasts revealed that only second graders significantly profited from the commutative problems (second graders: f[1, 289] = 16.32, mse = 4.04, p .06; third graders: f[1,289] = 2.59, mse = 4.04, p = .108). a separate t-test with block type as within-participants variable revealed that the adult control group did not show a significant benefit from commutative problems (t < 1). in addition, we also analyzed the percentage of correct responses in the commutativity and the control blocks. table 4 presents the percentage of correct responses in the two age groups for these two types of problems. as can be seen from table 4, percentage of correct responses was higher for thirdas compared to second graders. accordingly, the 2 (age group) x 2 (block type) anova yielded a significant main effect of age group (f[1, 289] = 3.8, mse = 118.27, p < .05, η² = .01). no other effect was significant. h. haider et al. 9 | f l r table 4 mean percent correct responses in the three age groups for the commutativity and control blocks. also depicted are standard deviations (in parentheses) and the lower and upper limit of the 95-% confidence interval (ci; loftus & masson, 1994) age group commutativity block control block n m (sd) m±95%ci m (sd) m±95%ci grade 2 92.66 (9.85) 91.64 93.68 93.11 (9.13) 92.0894.13 146 grade 3 94.33 (8.62) 91.64 93.68 94.82 (8.31) 94.0495.60 145 adults 96.66 (9.08) 94.8198.51 95.94 (12.55) 94.0997.78 37 overall, the results up to this point show that third graders were faster and less error prone as compared to second graders. furthermore and more importantly, second graders showed a substantial benefit of commutative problems. third graders in tendency also profited from these problems, but for them the effect was not significant. adults, by comparison, did not show such a benefit, probably due to a floor effect based on the simplicity of the problems (for similar results, see robinson & dubé, 2009; robinson & ninowski, 2003). 3.2 judgment task for each student, we individually computed the hit rate, false alarms rate, and the sensitivity index (d’) from signal detection theory 2 . table 5 depicts the means for these dependent measures separately for each of the two age groups and the adult students. as expected, hit rate was higher than false alarms rate in all age groups. accordingly, the sensitivity index d’ differed significantly from chance (all ts > 2.5, ps < .01). this suggests that students were able to correctly identify at least some of the commutative problems. in addition, d’ was substantially higher in thirdas compared to second graders (t[289] = 2.18, p < .05, η² = .02). 2 separately for each student within the respective age groups, we computed z-score of his or her hit and false alarms rate. then, we individually computed the sensitivity index d’ from signal detection theory by subtracting the ztransformed false alarms rate from the z-transformed hit rate. h. haider et al. 10 | f l r table 5 rate of hits, false alarms and d’ in each of the three age groups in the judgment task. standard deviants are given in parentheses. age group judgment task hits false alarms d’ grade 2 .70 (.28) .36 (.33) 1.81 (2.35) grade 3 .81 (.22) .35 (.32) 2.41 (2.32) adults .82 (.25) .11 (.18) 3.92 (2.33) low sensitivity could result from two different sources: the difficulty to identify commutative problems (hit rate) or a tendency to mark other than commutative problems (false alarm rate). therefore, we additionally analyzed the hit and false alarm rates in the two age groups. these analyses revealed that higher sensitivity in grade 3 as compared to grade 2 was mainly due to a higher hit rate. second graders were less able to identify the commutative problems than third graders (t[289] = 3.44, p < .01, ² = .04). the false alarm rate did not differ significantly between these two age groups (t < 1). by contrast, as can be seen from table 5, the higher sensitivity in adults as compared to third graders resulted from a lower false alarm rate in adults, whereas hit rate was almost identical in these two age groups. thus, older participants were better able to discriminate between commutative and control problems than younger participants. this finding from our cross-sectional age-comparison suggests a progress in conceptual knowledge with increasing age (as cohort differences are unlikely). 3.3 relation between procedural and conceptual knowledge the results reported up to this point are somewhat counterintuitive. even though third graders were better able to identify commutative problems in the judgment task, they seemed to rely less on a commutativity-based shortcut during calculation than second graders. in addition, for adults we found no benefit of commutative problems in the arithmetic task. thus, it seems that either the willingness to use more efficient arithmetic strategies or procedural knowledge of commutativity itself decreases (while conceptual knowledge increases with age). the last analyses of the relationship between procedural and conceptual knowledge might help to reconcile this picture. these analyses will answer the research question whether or not participants possess an integrated concept of commutativity. if so, we should find significant positive correlations between the use of the commutative-based shortcut in the arithmetic task (procedural knowledge) and the ability to correctly identify the commutative problems in the judgment task (conceptual knowledge). for procedural knowledge, we used for each participant the average calculation time per problem in the control and the commutativity block of the arithmetic task as well as the difference between these two measures (i.e., savings; with positive values indicating shorter calculation times in the commutativity block). for conceptual knowledge, we used hit rate, false alarms rate, and the sensitivity measure d’. in a first analysis we calculated correlations across second and third graders. second, we calculated correlations within the two age groups and for the adults. table 6 depicts the correlation between procedural and conceptual knowledge. h. haider et al. 11 | f l r table 6 correlation coefficients between procedural and conceptual knowledge depicted separately for all second and third graders as well as for the three age groups. procedural knowledge is indicated by calculation times in seconds for commutative problems, control problems, and in addition for savings. hit rate, false alarms, and d’ indicate conceptual knowledge arithmetic tasks judgment task commutative problems control problems savings n second and third graders hits -0.23 ** -0.17 ** 0.001 291 false alarms -0.09 -0.09 0.02 d’ -0.05 -0.02 0.01 grade 2 hits -0.23 ** -0.18 ** -0.04 146 false alarms -0.13 -0.15 -0.03 d’ -0.09 -0.04 -0.008 grade 3 hits -0.02 0.06 0.11 145 false alarms 0.06 -0.04 0.01 d’ 0.05 0.09 0.06 adults hits -0.23 0.03 0.39 * 37 false alarms -0.15 -0.24 -0.14 d’ -0.11 0.13 0.36 * ** p < .01; * p < .05 as can be seen from table 6, hit rate for the entire group correlated negatively with calculation time for commutative and control problems. that is, the faster second and third graders solved the arithmetic problems, the better they were able to identify commutative problems in the judgment task. savings in solution time due to commutative problems were not related to the ability to identify commutative problems, suggesting that their knowledge about commutativity was not very well integrated. a closer look at the different age groups revealed, however, that adults showed the expected positive correlation between savings and sensitivity. adults who applied the commutativity-based shortcut in the arithmetic blocks were also those who were better able to identify the commutative problems in the judgment task. this correlation suggests that the tested adults do possess an integrated knowledge representation of the commutativity principle. h. haider et al. 12 | f l r in contrast, second and third graders’ procedural and conceptual knowledge were only weakly related at best. as table 6 additionally reveals, second graders’ hit rate correlated negatively with calculation time. again, this correlation suggests that the faster second graders solved the arithmetic problems the better they were able to identify the commutative problems in the judgment task. thus, second graders’ ability to discriminate between commutative and control problems was linked to more general calculation competencies rather than to their procedural knowledge about using the commutativity-based shortcut. third graders, by contrast, did not show any significant correlation between calculation performance and discrimination. overall, these findings suggest that only adults’ spontaneous application of commutativity knowledge is based on an integrated concept of the commutativity principle. in contrast, procedural and conceptual knowledge seem to be only weakly related in second and third graders. note that alternatively, one also could argue that our assessments of procedural and conceptual knowledge are not sufficiently reliable (but, see table 2). in order to further rule out this latter argument and to better understand the missing correlations between savings in the arithmetic task (procedural) and the sensitivity index in the judgment task (conceptual knowledge), we conducted a final fine grained analysis for second and third graders. in the judgment task, false alarms rate of second and third graders was rather high (approximately 40%; table 5) and differed largely between participants in both age groups. presumably, children with a high false alarms rate might have correctly recognized the commutative problems in the judgment task, but at the same time might have believed that also easy to calculate problems (i.e., those with comparatively small addends) needed no calculation. this might have inflated false alarms rate and thus might have reduced the correlations between procedural and conceptual knowledge within second and third graders. following up on these assumptions, we divided the second and third graders into three different groups according to their false alarm rate: children with no false alarms, with up to 50% false alarms, and children with a false alarm rate higher than 50%. table 7 presents the number of participants within these three groups as well as the hit rates separately per grade. table 7 mean hit rates for second and third graders with no (fa = 0), medium (fa ≤ 50%), or high (fa > 50%) false alarms rate in the judgment task false alarm rate no false alarms medium fa-rate high fa-rate hit rate n hit rate n hit rate n grade 2 .81 41 .52 61 .85 44 grade 3 .88 44 .73 60 .84 41 as can be seen from table 7, for second and third graders hit rate was high when either the false alarm rate was low or when the false alarm rate was high (i.e., some children marked only the commutative problems while others marked the commutative and many other problems). this might have caused the overall low correlations between procedural and conceptual knowledge within these two age groups. therefore, we re-analyzed the correlation between hit rates and d’ and arithmetic abilities separately for these three groups within second and third graders. table 8 presents these correlations. in both age groups, only those participants who produced high hit rates without incorrectly marking the filler problems also h. haider et al. 13 | f l r showed substantial correlations. however, second and third graders differed qualitatively with regard to these correlations. table 8 correlations between procedural and conceptual knowledge for second and third graders with no, medium, or high false alarms rate. procedural knowledge is indicated by calculation times in seconds for commutative and control problems as well as savings. hit and false alarms rate (fa) indicate conceptual knowledge grade 2 no false alarms (n = 41) medium fa-rate (n = 61) high fa-rate (n = 44) hits fa hits fa hits fa commutative -.50 ** --.11 .07 -.01 -.01 control -.39 ** --.11 -.10 .02 -.04 savings .03 --.11 -.15 .04 .04 grade 3 no false alarms (n = 44) medium fa-rate (n = 60) high fa-rate (n = 41) hits fa hits fa hits fa commutative -.25 - .04 .12 .04 -.22 control -.01 -.07 .08 -.04 -.18 savings .31 * -.05 -.03 -.02 -.04 ** p < .01; * p < .05 once again, the results suggest that the second graders’ ability to discriminate between commutative and control problems in the judgment task is mainly related to their general calculation skills rather than to their ability to rely on efficient calculation strategies (i.e., the commutativity-based shortcut strategy). by contrast, third graders with high discrimination abilities seem to already possess integrated procedural and conceptual knowledge, starting to form an abstract understanding of commutativity. they use this knowledge, on the one hand, to identify commutative problems in the context of control problems and, on the other hand, to increase efficiency in solving arithmetic problems. 4. discussion with the current study we aimed at presenting an approach to unobtrusively measure the spontaneous usage of procedural and conceptual knowledge of the commutativity principle. apart from providing a basis to develop the method further (see below), the second goal of our study was to investigate h. haider et al. 14 | f l r the relation between procedural and conceptual knowledge about the commutativity principle in second and third graders. for this we asked (1) at which grade the different forms of commutativity knowledge can be detected and (2) at which grade they start to correlate with one-another. overall, our study yielded three main results: first, as expected, third graders showed higher general calculation proficiency (procedural knowledge) and more conceptual knowledge about commutativity than second graders. second, a solution time benefit based on the procedural use of the commutativity principle was only found for second graders. they calculated commutative problems faster than control problems. neither calculation times of third graders nor of adult students reflected significant profit from interspersed commutative problems. third, the correlation between (a) the benefit resulting from a commutativity-based shortcut and (b) conceptual knowledge of commutativity was rather weak in second and third graders. the relation seems to arise in some of the third graders. the link was also present in the control group (adults). the second and third findings merit some further discussion before we come to the theoretical and practical implications. the second finding (i.e., that only second graders’ calculation performance reflected the exploitation of commutativity whereas that of third graders and adults did not) was somewhat surprising. interestingly however, gaschler, vaterrodt, frensch, eichler, and haider (2013) found similar patterns of results with the identical arithmetic task. therefore, we assume that this finding is not due to a sample artefact. nevertheless, it does not fit the general claim that with experience, children become faster and more accurate at solving addition problems and also tend to use more sophisticated strategies, such as order-irrelevant, decomposition, and retrieval strategies (baroody et al., 1983; canobi et al., 1998, 2002; 2003; geary, brown & samaranayake, 1991; goldman, mertz & pellegrino, 1989; resnick, 1992; rittle-johnson & siegler, 1998; siegler, 1987; but see, mcneil, 2007; robinson & dubé, 2009; robinson & ninowski, 2003; torbeyns et al., 2009). it also seems to contradict the results of baroody et al. (1983), which show that approximately 80% of their third graders applied the commutativity-based shortcut to solve arithmetic problems (see also, canobi et al., 2003). one obvious reason for these divergent findings might be that we used three-element addition problems which probably were hard for second graders but (due to the rather small addends) easy for third graders and adults. this may have caused second graders to rely on the more efficient commutativity-based shortcut strategy, whereas third graders and adults were fast in solving the problems anyway, so that they did not consider any gain through using the shortcut strategy. for instance, siegler and araya (2005) mentioned that participants are more likely to adopt solution strategies if they contribute to significant performance advantages. this argument is further supported by the results of gaschler et al. (2013) who found larger benefits when presenting problems with large rather than with small addends. a second reason might be that in former studies (e.g., baroody et al., 1983; canobi et al., 1998; farrington-flint, canobi, wood & faulkner, 2010) students were instructed to explain their strategy immediately after responding. by contrast, our participants received no hint about the existence of commutative problems. while in our study the use of any shortcut strategy was spontaneous, it is possible that the explanation required in the baroody et al.’s study might have triggered students to apply the commutativity-based strategy. for instance, torbeyns et al. (2009) found less strategy application when students could spontaneously apply different strategies during calculation than when they were instructed to do so. in a similar vein, a yet unpublished study from our labs revealed that second and third graders as well as adults substantially benefitted from instruction (compared to a non-instructed group). participants reminded of the commutativity principle and alerted to the fact that commutative problems might occur, showed a larger solution time advantage on commutative problems as compared to control problems. as students of all three age groups relied on the commutativity principle after being instructed accordingly, it seems justified to conclude that (with the exception of second graders) students in our study indeed did not profit much from spontaneously applying the commutativity-based shortcut strategy. concerning our second research question, we found that the second graders’ understanding of commutativity was unrelated to their use of commutativity-based shortcut strategies. first signs of an integrated concept (assessed by the correlation of procedural and conceptual knowledge measures) occurred h. haider et al. 15 | f l r in a small group of third graders and were substantial only for adult students. thus, the integration of procedural and conceptual knowledge seems to increase with age. however, it also suggests that second graders may have used the shortcut strategy without entirely understanding the commutativity principle. this finding seems at odds with the early onset assumption of, for example, baroody and gannon (1984). furthermore, canobi et al. (1998; 2002) had found that second graders’ conceptual (assessed by an explanation task) and procedural knowledge (assessed by solving addition problems) correlated moderately. however, as already discussed concerning baroody et al.’s (1983) findings, canobi (2009; canobi et al., 1998, 2002) assessed conceptual knowledge by asking participants to explain their strategies after they had solved an addition problem. thus, even though canobi et al. used different tasks for assessing procedural and conceptual knowledge, the knowledge assessed by their addition task might have resulted from a mixture of procedural and conceptual competencies (see also, robinson & dubé, 2009). this might have led to a higher correlation between both tests compared to a variant were spontaneous application of procedural and conceptual knowledge is independently assessed. alternatively, one might suspect that our measures of procedural and conceptual knowledge were unreliable. however, this is not likely as we did find satisfying split-half reliabilities for all age groups and both task formats (see, table 2). in addition, our results showed significant correlations (a) for all second graders between the calculation time and percentage of hits and (b) for at least some third graders and all adult students between savings due to the use of commutativity-based shortcuts and hits in the judgment task. therefore, it seems worthwhile to ask for further theoretical causes concerning our third finding. 4.1 theoretical implications at first glance, the results seem to fit with a procedural-first development of commutativity (e.g., baroody et al., 2007; briars & siegler, 1984; siegler & stern, 1998). that is, second graders use the commutativity-based shortcut before having acquired an abstract understanding of the principle. it therefore appears that the development of conceptual knowledge progresses more slowly than that of procedural knowledge – at least as measured in this study and for commutativity (cf. canobi, 2004; canobi et al., 1998). however, in their review about relations between children’s understanding of mathematical concepts and their ability to execute arithmetic procedures, rittle-johnson and siegler (1998) provided ample evidence that with regard to commutativity, children first acquire conceptual knowledge before then applying corresponding strategies. in order to reconcile this conflict, we refer to the iterative model of the development of conceptual and procedural knowledge (e.g., resnick, 1992; rittle-johnson et al., 2001). our findings suggest that second graders possess at least rudimentary conceptual knowledge of the commutativity principle, but their conceptual representation of the commutativity principle is less well integrated (with procedural knowledge) than that of third graders or adult students. this is in line with many findings in the field of mathematic development showing that already second graders possess conceptual knowledge about commutativity (for a review, see rittle-johnson & siegler, 1998). also, our sensitivity index d’ indicated such knowledge. however, the consistent use of this knowledge may still be reduced; that is, their competency to identify the relevant task properties for applying a certain shortcut strategy has not fully developed yet. therefore, they may need a certain external trigger in order to activate their knowledge about commutativity and the corresponding strategies. our instructions for the arithmetic and the judgment task did not provide any such trigger which probably made it rather difficult, particularly for second graders and also for most of the third graders, to realize that they should rely on the commutativity principle in both tasks. consequently, it may be that some participants applied the commutativity-based shortcut strategy to solve the arithmetic problems, but did not use it in the judgment tasks or vice versa. this does not imply that they first learn procedures before they acquire conceptual knowledge. rather, we assume that such a finding mainly reflects that children’s conceptual knowledge is not sufficiently integrated to spontaneously recognize that they could rely on the commutativity principle. in a similar vein, research on expertise (e.g., anderson & schunn, 2000; gentner & toupin, 1986; haider & frensch, 1996; koedinger & anderson, 1990) also shows that wellintegrated and thus abstract conceptual knowledge is required to identify task relevant information in order h. haider et al. 16 | f l r to solve problems and to flexibly transfer knowledge from one task domain to another (see also e.g., sloutsky & fisher, 2008; star & seifert, 2006). to summarize, we assume that the divergent findings concerning the development of an abstract understanding of the commutativity principle reflect the fact that after students have acquired some procedural and conceptual knowledge in this domain, this knowledge needs to be integrated. this integration of knowledge, we suspect, is done in an iterative way which means that procedures are applied, which then refine the conceptual knowledge (rittle-johnson et al., 2001). the conceptual knowledge is then used to guide children’s attention to information which is needed to adaptively apply efficient strategies. 4.2 further improvements of the measurement of spontaneous application of a mathematical principle with the current study, we took a first step to measure spontaneous usage of commutativity knowledge in a non-reactive way. participants worked on the paper-and-pencil tasks in a setting very similar to other tests in the classroom. in the arithmetic blocks, we asked for fast and correct solutions to the arithmetic problems and did not mention that regularities in the task material might be exploited for efficient task processing. we inferred procedural knowledge of the commutativity principle from the performance benefits on material containing identical addends in changed order in consecutive problems (as compared to material that did not contain such pairs of problems). probing for conceptual knowledge, we asked participants to indicate in which cases calculation was not necessary – again without hinting that it might be the commutativity principle that made calculation superfluous. the rationale behind this procedure was that participants who had well integrated knowledge of the commutativity principle should recognize the respective arithmetic problems and consequently should be able to relate it to the task demand (marking problems where calculation was not necessary). age-related changes in hits and false alarms in the judgment task suggested that this was indeed the case. similar indirect approaches to measure knowledge have been developed in order to measure insight (cf. haider & rose, 2007). when investigating insight, it is not feasible either to directly ask participants again and again if they already have discovered the regularity in the task material – without providing them with a strong hint that such a regularity exists. one way to further improve the method would be to include control problems which also feature the same numbers as their predecessor problems in changed order but do not allow to apply the commutativity principle (i.e. subtractions). such interspersed control problems could help to rule out superficial matching strategies (i.e. “same numbers = same result”) that do not capture the essence of commutativity. first explorations in our labs indicate that second graders do not confuse subtraction problems containing the same digits as a preceding addition with genuine commutative problems. furthermore, it would be interesting to implement our instruments within a multi-method approach, administering multiple measurements per person and construct (cf. prather & alibali, 2009). as a first step one would have to estimate to what extend repeated testing of spontaneous usage of a mathematical principle induces participants to recognize and use the principle – and by this spoils the possibility to assess spontaneous usage. paper-and-pencil-based testing in the classroom has the advantage that the test situation is similar to other tests the students take. in a parallel line of research, we have started to employ eyetracking to obtain process measures related to commutativity knowledge (e.g., gaschler et al. 2013; godau, wirth, hansen, haider, & gaschler, in press). for instance, it is possible to quantify the extent to which a child searches for repetitions of addends in subsequent addition problems. however, when they are tested individually with an eyetracking system, children are aware that the measurement is about where they look and how they calculate. class-based testing in computer labs within schools might offer the possibility to obtain process measures while keeping up the character of the assessment as allowing to measure spontaneous application of the principle knowledge. h. haider et al. 17 | f l r 4.3 theoretical conclusions and practical implications recently, prather and alibali (2009; see also, bisanz et al., 2009; schneider & stern, 2010) called for multifaceted knowledge assessment in the context of arithmetic development. our use of the arithmetic and judgment tasks in order to independently assess procedural and conceptual knowledge can be seen as a first step in this direction. complementing earlier work on commutativity knowledge (e.g., baroody & gannon, 1984; canobi et al., 2002), our findings suggest that, when second graders and third graders are not alluded to rely on the commutativity principle, second graders and most of the third graders show but weak signs of spontaneous application of commutativity and interrelation of different forms of commutativity knowledge. this suggests that they do not possess well-integrated knowledge about commutativity in the sense of an abstract formal mathematical principle. our results suggest that even if children use procedures that suggest integrated conceptual knowledge about commutativity, the learning process has by far not reached an endpoint. rather, it still progresses, before leading to a well-integrated, abstract representation of the mathematical principle as with our measures found in adults. as long as children do not possess such an abstract representation, they will not be able to flexibly and adaptively use the commutativity principle in different task contexts. accordingly, we suspect that increasing experience in the field of mathematics is needed in order to better integrate conceptual knowledge about various arithmetic principles. this might explain why transfer of knowledge from one context to another is often found to be rather weak (e.g., frensch & haider, 2008; kaminsky, sloutsky & heckler, 2008; siegler & stern, 1998; sloutsky & fisher, 2008). therefore, helping students to develop well-integrated knowledge concepts should be one of the most important tasks education has to fulfill (see, e.g., geary et al., 2008; prather & alibali, 2009; verschaffel et al. 2009). in more practical terms, if children are taught the commutativity principle in the context of addition, they seem to learn that they can use this principle to avoid unnecessary labor. however, our results suggest that this does not mean that they concurrently acquire an idea of the abstract principle of cardinality. we suspect that many children only acquire a procedure (or a strategy) that they can easily apply for twoelement addition problems. in order to help students to understand the abstract principle of commutativity, it might be worthwhile to activate students’ prior knowledge of this principle, such as the order-irrelevance principle they already use in counting. when introducing the commutativity principle in addition (or multiplication) it might be helpful to tell students that they already have used this principle in other contexts and explain how and why it works in all these different situations. this then might help them to understand the commutativity principle in a more abstract manner and probably also to understand task properties needed to correctly apply this principle. further, it may help to support children in recognizing the consequences of using alternative strategies in order to ensure representational redescription (e.g., baroody & gannon, 1984). keypoints procedural and conceptual commutativity knowledge increase with increasing age. second graders show no signs of an integrated concept of commutativity. first signs of an integrated concept of commutativity emerge in grade three. acknowledgements this research was supported by the german research foundation (dfg; hh1471/12-1). some of the results were presented at the kongress der deutschen gesellschaft für psychologie 2010 in bremen, germany. we thank annette bräutigam, yvonne radermacher, pia blase and ester jung for help with data collection. h. haider et al. 18 | f l r references anderson, j. r. & schunn, c. d. (2000). implications of the act-r learning theory: no magic bullets. in r. glaser (ed.), advances in instructional psychology: educational design and cognitive science (vol. 5, pp. 1-33). mahwah, nj: lawrence erlbaum associates publishers. baroody, a. j. (1984). more precisely defining and measuring the order-irrelevance principle. journal of experimental child psychology, 38, 33-41. doi:10.1016/0022-0965(84)90017-1 baroody, a. j. (2003). the development of adaptive expertise and flexibility: the integration of conceptual and procedural knowledge. in a. baroody & a. dowker (eds.), the development of arithmetic concepts and skills. constructing adaptive expertise (1st ed., pp. 1–34). mahwah, nj: lawrence erlbaum associate publishers. baroody, a. j., feil, y. & johnson, a. r. (2007). an alternative reconceptualization of procedural and conceptual knowledge. journal for research in mathematics education, 38, 115-131. baroody, a. j. & gannon, k. e. (1984). the development of the commutativity principle and economical addition strategies. cognition & instruction, 1, 321-339. doi:10.1207/s1532690xci0103_3 baroody, a. j., ginsburg, h. p. & waxman, b. (1983). children's use of mathematical structure. journal for research in mathematics education, 14, 156-168. doi:10.2307/748379 baroody, a.j., & rosu, l. (2006). adaptive expertise with basic addition and subtraction combinations the number sense view. paper presented at theannual meeting of the american educational research association (april) san francisco, ca. bisanz, j., & lefevre, j. (1992). understanding elementary mathematics. in j. d. campbell (ed.) , the nature and origins of mathematical skills (pp. 113-136). oxford england: north-holland. doi:10.1016/s0166-4115(08)60885-7 bisanz, j., watchorn, r. p. d., piatt, c., & sherman, j. (2009). on “understanding” children's developing use of inversion. mathematical thinking and learning, 11, 10–24. doi:10.1080/10986060802583907 briars, d. & siegler, r. s. (1984). a featural analysis of preschoolers' counting knowledge. developmental psychology, 20, 607-618. doi:10.1037/0012-1649.20.4.607 canobi, k. h. (2004). individual differences in children's addition and subtraction knowledge. cognitive development, 19, 81-93. doi:10.1016/j.cogdev.2003.10.001 canobi, k. h. (2005). children's profiles of addition and subtraction understanding. journal of experimental child psychology, 92, 220-246. doi:10.1016/j.jecp.2005.06.001 canobi, k. h. (2009). concept-procedure interactions in children's addition and subtraction. journal of experimental child psychology, 102(2), 131-149. doi:10.1016/j.jecp.2008.07.008 canobi, k. h., reeve, r. a. & pattison, p. e. (1998). the role of conceptual understanding in children's addition problem solving. developmental psychology, 34, 882-891. doi:10.1080/01443410903473597 canobi, k. h., reeve, r. a. & pattison, p. e. (2002). young children’s understanding of addition concepts. educational psychology, 22, 513-532. doi:10.1080/0144341022000023608 canobi, k. h., reeve, r. a. & pattison, p. e. (2003). patterns of knowledge in children’s addition. developmental psychology, 39, 521–534. doi:10.1037/0012-1649.39.3.521 cowan, r. & renton, m. (1996). do they know what they are doing? children's use of economical addition strategies and knowledge of commutativity. educational psychology, 16, 407-420. doi:10.1080/0144341960160405 farrington-flint, l., canobi, k.h., wood, c. & faulkner, d. (2010). children’s patterns of reasoning about reading and addition concepts. british journal of developmental psychology, 28, 427–448. doi:10.1348/026151009x424222 frensch, p. a. & haider, h. (2008). transfer and expertise: the search for identical elements. in h. l. roediger, iii (ed.), cognitive psychology of memory. vol. [2] of learning and memory: a comprehensive reference (pp. 579-596) oxford: elsevier. doi:10.1016/b978-012370509-9.00177-7 fuson, k. c. (1988). children's counting and concepts of number. new york, ny: springer-verlag publishing. h. haider et al. 19 | f l r gaschler, r., vaterrodt, b., frensch, p. a., eichler, a. & haider, h. (2013). spontaneous usage of different shortcuts based on the commutativity principle. plos one 8(9): e74972. doi:10.1371/journal.pone.0074972 geary, d. c. (2006). development of mathematical understanding. in w. damon, r. m. lerner, & n. eisenberg (eds.), handbook of child psychology: social, emotional, and personality development (vol. 3, pp. 777–810). john wiley and sons. geary, d. c., boykin, a. w., embretson, s., reyna, v., siegler, r., berch, d. b., et al. (2008). report of the task group on learning processes. in national mathematics advisory panel, reports of the task groups and subcommittees (pp. 4-1–4-211). geary, d. c., brown, s. c. & samaranayake, v. a. (1991). cognitive addition: a short longitudinal study of strategy choice and speed-of-processing differences in normal and mathematically disabled children. developmental psychology, 27, 787-797. doi:10.1037/0012-1649.27.5.787 geary, d. c., hoard, m. k., byrd-craven, j. & desoto, m. c. (2004). strategy choices in simple and complex addition: contributions of working memory and counting knowledge for children with mathematical disability. journal of experimental child psychology, 88, 121–151. doi:10.1016/j.jecp.2004.03.002 gelman, r. (1990). first principles organize attention to and learning about relevant data: number and the animate-inanimate distinction as examples. cognitive science, 14, 79-106. gelman, r. & gallistel, c. r. (1978). the child's understanding of number. in. cambridge, ma: harvard university press. gelman, r. & meck, e. (1983). preschoolers' counting: principles before skill. cognition, 13, 343-359. doi:10.1016/0010-0277(83)90014-8 gentner, d. & toupin, c. (1986). systematicity and surface similarity in the development of analogy. cognitive science, 10, 277-300. doi:10.1207/s15516709cog1003_2 godau, c., wirth, m., hansen, s., haider, h., & gaschler, r. (in press). from marbles to numbers estimation influences looking patterns on arithmetic problems. psychology. goldman, s. r., mertz, d. l. & pellegrino, j. w. (1989). individual differences in extended practice functions and solution strategies for basic addition facts. journal of educational psychology, 81, 481-496. doi:10.1037/0022-0663.81.4.481 goldstone, r. l., & sakamoto, y. (2003). the transfer of abstract principles governing complex adaptive systems. cognitive psychology, 46, 414-466. doi:10.1016/s0010-0285(02)00519-4 haider, h. & frensch, p. a. (1996). the role of information reduction in skill acquisition. cognitive psychology, 30, 304-337. doi:10.1006/cogp.1996.0009 haider, h. & rose, m. (2007). how to investigate insight: a proposal. methods, 42, 49–57. doi: 10.1016/j.ymeth.2006.12.004 hannula, m. m., & lehtinen, e. (2005). spontaneous focusing on numerosity and mathematical skills of young children. learning and instruction, 15, 237-256. doi:10.1016/j.learninstruc.2005.04.005 hannula, m. m., lepola, j., & lehtinen, e. (2010). spontaneous focusing on numerosity as a domainspecific predictor of arithmetical skills. journal of experimental child psychology, 107, 394-406. doi:10.1016/j.jecp.2010.06.004 hiebert, j., & lefevre, p. (1986). conceptual and procedural knowledge in mathematics: an introductory analysis. in j. hiebert (ed.), conceptual and procedural knowledge: the case of mathematics (pp. 127). hillsdale, nj: erlbaum. kaminski, j. a., sloutsky, v. m., & heckler, a. f. (2008). learning theory: the advantage of abstract examples in learning math. science, 320, 454–455. koedinger, k. r. & anderson, j. r. (1990). abstract planning and perceptual chunks: elements of expertise in geometry. cognitive science, 14, 511-550. doi:10.1207/s15516709cog1404_2 kuhn, d., garcia-mila, m., zohar, a., & andersen, c. (1995). strategies of knowledge acquisition. society for research in child development monographs, 60 (4), serial no. 245. lefevre, j.-a., smith-chant, b. l., fast, l., skwarchuk, s.-l., sargla, e., arnup, j. s., et al. (2006). what counts as knowing? the development of conceptual and procedural knowledge of counting from kindergarten through grade 2. journal of experimental child psychology, 93, 285-303. doi:10.1016/j.jecp.2005.11.002 h. haider et al. 20 | f l r mcmullen, j.a., hannula-sormunen, m.m., & lehtinen, e. (2011). young children’s spontaneous focusing on quantitative aspects and their verbalizations of their quantitative reasoning. in ubuz, b. (ed.). proceedings of the 35th conference of the international group for the psychology of mathematics education, 3, pp. 217-224. ankara, turkey: pme. mcneil, n. m. (2007). u-shaped development in math: 7-year-olds outperform 9-year-olds on equivalence problems. developmental psychology, 43, 687–695. doi:10.1037/0012-1649.43.3.687 prather, r. w., & alibali, m. w. (2009). the development of arithmetic principle knowledge: how do we know what learners know? developmental review, 29, 221–248. doi:10.1016/j.dr.2009.09.001 resnick, l. b. (1992). from protoquantities to operators: building mathematical competence on a foundation of everyday knowledge. in g. leinhardt, r. t. putnam & r. a. hattrup (eds.), analysis of arithmetic for mathematics teaching (pp. 373-429). hillsdale, nj, uk: lawrence erlbaum associates, inc. rittle-johnson, b. & siegler, r. s. (1998). the relation between conceptual and procedural knowledge in learning mathematics: a review. in c. donlan (ed.), the development of mathematical skills (pp. 75-110)). east sussex, uk: psychology press. rittle-johnson, b., siegler, r. s. & alibali, m. w. (2001). developing conceptual understanding and procedural skill in mathematics: an iterative process. journal of educational psychology, 93, 346362. doi:10.1037/0022-0663.93.2.346 robinson, k. m. & dubé, a. k. (2009). children’s understanding of addition and subtraction concepts. journal of experimental child psychology, 103, 532–545. doi:10.1016/j.jecp.2008.12.002 robinson, k. m., & ninowski, j. e. (2003). adults’s understanding of inversion concepts: how does performance on addition and subtraction inversion problems compare to performance on multiplication and division inversion problems? canadian journal of experimental psychology, 57 321-330. doi:10.1037/h0087435 schneider, m. &. stern, e. (2010). the developmental relations between conceptual and procedural knowledge: a multimethod approach. developmental psychology, 46, 178–192. doi:10.1037/a0016701 siegler, r. s. (1987). the perils of averaging data over strategies: an example from children's addition. journal of experimental psychology: general, 116, 250-264. doi:10.1037/0096-3445.116.3.250 siegler, r. s. & araya, r. (2005). a computational model of conscious and unconscious strategy discovery. in r. v. kail (ed.), advances in child development and behavior (vol. 33) (pp. 1-42). oxford, uk: elsevier. siegler, r. s. & jenkins, e. (1989). how children discover new strategies. hillsdale, nj, uk: lawrence erlbaum associates, inc. siegler, r. s. & stern, e. (1998). conscious and unconscious strategy discoveries: a microgenetic analysis. journal of experimental psychology: general, 127, 377-397. doi:10.1037/0096-3445.127.4.377 sloutsky, v. m. & fisher, a. v. (2008). attentional learning and flexible induction: how mundane mechanisms give rise to smart behaviors. child development, 79, 639-651. doi:10.1111/j.14678624.2008.01148.x sophian, c. & adams, n. (1987). infants' understanding of numerical transformations. british journal of developmental psychology, 5, 257-264. doi:10.1111/j.2044-835x.1987.tb01061.x sophian, c., harley, h. & manos martin, c. s. (1995). relational and representational aspects of early number development. cognition & instruction, 13, 253-268. doi:10.1207/s1532690xci1302_4 star, j.r. (2004, april). the development of flexible procedural knowledge in equation solving. paper presented at the annual meeting of the american educational research association, san diego. star, j. r. (2005). reconceptualizing procedural knowledge. journal for research in mathematics education, 36, 404-411. star, j. r. & seifert, c. (2006). the development of flexibility in equation solving. contemporary educational psychology, 31, 280-300. doi:10.1016/j.cedpsych.2005.08.001 starkey, p. & gelman, r. (1982). the development of addition and subtraction abilities prior to formal schooling in arithmetic. in t. p. carpenter, j. m. moser & t. a. romberg (eds.), addition and subtraction: a cognitive perspective (pp. 99–116). hillsdale, nj: erlbaum. h. haider et al. 21 | f l r torbeyns, j., de smedt, b., ghesquière, p. &verschaffel, l. (2009). acquisition and use of shortcut strategies by traditionally schooled children. educational studies in mathematics, 71(1), 1-17. verschaffel, l., luwel, k., torbeyns, j. & van dooren, w. (2009). conceptualizing, investigating, and enhancing adaptive expertise in elementary mathematics education. european journal of psychology of education, 24(3), 335-359. doi:10.1007/bf03174765 microsoft word lindblom-ylänne et al_publication.docx           frontline learning research vol.3 no. 2 (2015) 47-62 issn 2295-3159   academic procrastinators, strategic delayers and something betwixt and between: an interview study sari lindblom-ylännea1, emmi saariahoa, mikko inkinena, anne-haarala-muhonenb, telle hailikaria afaculty of behavioural sciences, university of helsinki, finland bfaculty of law, university of helsinki, finland article received 1 march 2015 / revised 24 april 2015 / accepted 25 may 2015 / available online 12 june 2015 abstract the study explored university undergraduates’ dilatory behaviour, more precisely, procrastination and strategic delaying. using qualitative interview data, we applied a theory-driven and person-oriented approach to test the theoretical model of klingsieck (2013). the sample consisted of 28 bachelor students whose study pace had been slow during their first university year. three student profiles emerged. the first concerned strategic delay and was represented by motivated students with strong self-efficacy beliefs who had intentionally postponed their studying. the second consisted of students whose delaying was unnecessary in nature; these students had minor self-regulation problems but were still motivated to study. the third profile consisted of procrastinating students who lacked self-regulation skills and had weaker self-efficacy beliefs. the results indicate that dilatory behaviour can vary from strategic delay to dysfunctional procrastination, and that different factors are related to these various types of dilatory behaviour. this study adds to our theoretical understanding of academic procrastination by empirically testing a new theoretical model of procrastination. in addition, the study shows the value of using a qualitative approach in understanding the phenomenon of dilatory behaviour. keywords: academic procrastination; strategic delay; dilatory behaviour; university student                                                                                                                           1  corresponding author: sari lindblom-ylänne, institute of behavioural sciences, university of helsinki, p.o. box 9, 00014 university of helsinki , finland, email: sari.lindblom@helsinki.fi doi: http://dx.doi.org/10.14786/flr.v3i2.154 lindblom-­‐ylänne  et  al       | f l r     48   1. introduction research has shown that academic procrastination is very common among university students: almost all occasionally procrastinate in one or another domain of their studies, and approximately every second student regularly procrastinates (rothblum, solomon & murakami, 1986; steel, 2007). however, research in this area often lacks precision in the definition of procrastination, with the concept being used to describe different types of delay varying from functional to dysfunctional (e.g., klingsieck, 2013; schraw, wadkins & olafson, 2007; steel, 2007). an example of functional procrastination is a situation in which a student studies effectively and attains favourable results under pressure of an approaching deadline (e.g., choi & moran, 2009; chu & choi, 2005; schraw et al., 2007). examples of dysfunctional procrastination are delaying the decided time for beginning study processes, moving scheduled study periods for the future and engaging in study-irrelevant behaviour (e.g., schouwenburg, 1995). the different and even contrasting definitions of procrastination have made it difficult to understand the phenomenon and to follow the research. further, the definitions’ lack of precision also influences the way researchers operationalise these constructs and the analyses they perform. to address this pervasive problem, klingsieck (2013) recently provided an excellent meta-analysis of the different definitions of procrastination and of the trends in procrastination research. she suggests that a clear distinction should be made between procrastination and strategic delay, in other words, between dysfunctional and functional forms of delay. klingsieck (2013) proposed that the following seven parameters be used for this purpose: 1) the delay of an overt or covert act; 2) the act is intended to be started and/or completed; 3) the act is necessary or of personal importance; 4) the delay is voluntary; 5) the delay is unnecessary or irrational; 6) the act is delayed despite being aware of the potential negative consequences; and 7) the delay is accompanied by subjective discomfort or other negative consequences. according to klingsieck (2013), parameters 1 and 2 characterise any form of delay, and 3 and 4 both procrastination and strategic delay. what differentiates procrastination from strategic delay is the nature of the delay itself. the delay in procrastination is unnecessary, irrational and even harmful (5). in strategic delay, a student is confident that the positive consequences will eventually outweigh the potential negative ones, whereas procrastination involves negative consequences and is accompanied by subjective discomfort or other negative consequences (6 and 7). to summarise, “there is no functional form of procrastination, but there is a functional form of delay” (klingsieck 2013, 26), in other words, strategic delay. in the light of klingsieck’s distinction, research that has emphasised the adaptive forms of procrastination (e.g., choi & moran, 2009; chu & choi, 2005; ferrari, johnson & mccown, 1995; schraw, et al., 2007) could be considered as research on strategic delay. choi and colleagues (choi & moran, 2009; chu & choi, 2005) have used the terms ‘passive’ and ‘active’ procrastination. by passive procrastination they refer to postponing tasks “until the last minute because of an inability to make the decision to act in a timely manner” (choi & moran, 2009, 196). their definition of active procrastination fits klingsieck's (2013) definition of strategic delay where students are highly motivated by time pressure, and are able to complete tasks before deadlines and achieve satisfactory outcomes. thus, typical of active procrastinators is a preference for working under pressure (chu & choi, 2005). corkin, lu and lindt (2011) have argued that active procrastination is distinct from procrastination in general, and should be referred to as ‘active delay.’ their definition of active delay is very close to klingsieck's (2013) strategic delay. corkin et al. (2011) also showed active delay to be associated with adaptive self-regulatory processes and academic achievement, and procrastination to be associated with mastery-avoidance goals and a lack of metacognitive strategy. finally, grunschel, partzek and fries (2013a) used the term ‘purposeful delay,’ which can be considered synonymous with strategic delay. on the basis of this distinction between procrastination and strategic delay, procrastination can be defined as “the voluntary delay of an intended and necessary and/or [personally] important activity, despite expecting potential negative consequences that outweigh the positive consequences of the delay” (klingsieck, 2013, 26). klingsieck’s definition extends steel’s (2007, 65) definition of procrastination as “a prevalent and pernicious form of self-regulatory failure.” strategic delay may be defined here as the lindblom-­‐ylänne  et  al       | f l r     49   voluntary delay of an intended and necessary and/or [personally] important activity in which positive consequences are believed to outweigh negative consequences in the long run. from the point of view of individual students, procrastination and strategic delay can be closely intertwined when examining individuals in different study contexts, in different study situations and at different times. therefore, to understand individual dilatory behaviour more deeply, it is important, as klingsieck (2013) suggests, to expand the investigation of procrastination from one specific context and cover longer periods of time and different contexts. furthermore, when exploring the nature of procrastination it is important to take into account the whole procrastination process, from its reasons and contexts to actual procrastination behaviour and its consequences, as in grunschel, partzek and fries (2013b). the present study aims to empirically test the theoretical model of klingsieck (2013), in which strategic delay is separated from procrastination, by using a qualitative theory-driven and person-oriented approach. as a starting point we use the above-mentioned definitions of procrastination and strategic delay. in addition, the study aims to explore motivational, volitional and situational factors related to procrastination and strategic delay. 1.1 motivational, volitional and situational factors that promote procrastination theoretical approaches in procrastination research vary (klingsieck, 2013). the present study focuses on the motivational-volitional and situational dimensions of procrastination. researchers have identified many motivational-volitional and situational factors that accompany procrastination. low intrinsic study motivation, problems in self-regulation, poor time-management and/or organising skills and weak selfefficacy beliefs have been shown to be key factors in leading to procrastination (e.g., grunschel et al., 2013a; lee, 2005; pychyl, morin & salmon, 2000; rebetez, rochat & van der linden, 2015; steel, 2007; strunk, cho, steele & bridges, 2013; tice & baumeister, 1997; wolters, 2003). many studies have shown a link between procrastination and both extrinsic motivation and a lack of motivation (grunschel et al., 2013b; lee, 2005; pychyl et al., 2000; rebetez et al., 2015; tice & baumeister, 1997). lack of self-regulation has also been shown to increase procrastination (e.g., steel, 2007). self-regulation refers to a student’s own active role in his or her learning process (e.g., pintrich, 1995; vermunt & van rijswijk 1988; vermunt & verloop, 1999; zimmermann, 1994). characteristic of self-regulation is monitoring one’s actions, using metacognition, and regulating motivational and emotional states (e.g., pintrich, 1995; zimmerman, 1994). if these important skills are missing or if they are poorly developed, it is more difficult for a student to control his or her cognition, motivation, actions and emotions (pintrich, 1995), and this lack of control can promote procrastination. students who lack self-regulation skills often struggle alone, whereas students with good self-regulation skills are able to seek help for their study-related problems (newman, 1994; pintrich, 2004). klassen, krawschuk and rajani (2008) interestingly showed that low self-efficacy for self-regulation was a stronger predictor of the tendency to procrastinate than poor self-regulation skills or weak self-efficacy beliefs alone, or other motivation variables. according to them, self-efficacy for self-regulation “reflects an individual’s beliefs in his or her capabilities to use a variety of learning strategies, resist distractions, complete schoolwork, and participate in class learning” (p. 918). they showed that “self-efficacy to structure the learning environment […] leads to timely task completion and successful academic achievement” (klassen, et al., 2008, 922). furthermore, problems in self-regulation, together with underachievement and the avoidance of tasks seen as demanding, are characteristic of a self-handicapping strategy (e.g., eerde, 2003; garcia & pintrich, 1994; howell & watson, 2007), which in turn has been shown to be related to procrastination (ferrari & tice, 2000; solomon & rothblum, 1984; steel, 2007). self-handicapping is a cognitive strategy, which concerns avoiding effort, and in this way preventing potential failure from lowering self-esteem. for persons applying a self-handicapping strategy, avoiding effort is a way to make good performance less likely and to protect their sense of self-competence (jones & berglas, 1978). furthermore, typical of such a strategy is maladaptive task-irrelevant behaviour and a preference for external regulation in which responsibility for the lindblom-­‐ylänne  et  al       | f l r     50   learning process is shifted to the teacher (heikkilä & lonka, 2006). in addition, achievement goals are related to procrastination, and can be divided into mastery and performance goals. mastery goals focus on developing new skills, whereas performance goals focus on demonstrating ability and skills (e.g., ames & archer, 1988; elliot & harackiewicz, 1994). strong achievement goals have been shown to reduce procrastination, whereas having weak or no achievement goals increases procrastination (howell & buro, 2009). further, the effect of achievement goals on study process and study success seems to be mediated by students’ self-efficacy beliefs. the higher the students’ perceived self-efficacy is, the higher they set their goals, which in turn leads to better academic achievement (e.g., cheng & chiou, 2010). solomon and rothblum (1984) found that both fear of failure and feeling the task at hand to be disagreeable caused procrastination. in the case of fearing failure, procrastination has been explained by traits such as perfectionism, anxiety and low self-efficacy beliefs regarding one’s skills in organizing and regulating oneself in order to succeed at specific tasks (e.g., bandura, 1997). experiencing a task as disagreeable has been explained by problems in time management. fear of failure, low self-efficacy beliefs, task aversiveness and laziness have all been repeatedly mentioned as factors leading to procrastination (e.g., blunt & pychyl, 2000; eerde, 2003; ferrari & tice, 2000; howell & watson, 2007; rothblum et al., 1986; wolters, 2003). 1.2 aims of the study the present study has two objectives. firstly, we aim to empirically test the theoretical model of klingsieck (2013), in which strategic delay is separated from procrastination, by using a qualitative theorydriven and person-oriented approach. secondly, we aim to clarify on how motivational, volitional and situational factors are related to procrastination and strategic delay. the purpose is to explore long-term dilatory behaviour, i.e., procrastination and strategic delay, among university students during the first study year. because research on academic procrastination has mainly focused on short-term, task-related procrastination, knowledge about long-term dilatory behaviour is scarce. there are, however, a few interesting exceptions. a good example of a long-term research design is wäschle, allgaier, lachner, fink and nückles (2014): long-term procrastination and its relation to self-efficacy among university students were investigated using self-monitoring protocols during an academic term. the present study extends longterm dilatory behaviour to the whole academic year. 2. methodology 2.1 participants the participants comprised bachelor-level humanities and law students (n=28) from the university of helsinki, whose study pace had been slow during their first academic year. these individuals lacked at least a quarter of the credits students at the university are expected to earn each year. students earning less than 46 credits during their first year must submit a report to the university regarding their slow study pace, and create a detailed plan for future studies. the 46-credit limit is also the minimum requirement for receiving government-financed study grants, which are an important student benefit. we applied purposive sampling and invited these students to participate in interviews after their first year of study. two sample groups were assembled. the first consisted of humanities students from the faculty of arts, where graduation times are the longest (more than four years), and the second consisted of students from the faculty of law, where the average graduation times are the shortest (approx.. three and a half years). in addition, the two bachelor curricula are different in nature: the law curriculum is professional, lindblom-­‐ylänne  et  al       | f l r     51   comprised mainly of law studies with few optional courses, whereas humanities students can freely choose their minors and there are fewer compulsory elements. a total of 154 humanities students were enrolled in three bachelor of arts undergraduate programs, of whom 27 (17.5%) had earned less than 46 credits. of these 27 students 17 (63%) volunteered to be interviewed. their mean age was 24 years, ranging from 20 to 36. altogether 27% of these participants were male and 73% female. male students were slightly over-represented in the sample, their proportion of the whole cohort being 23%. a total of 247 law students were enrolled in bachelor of law undergraduate program, of whom 36 students (15%) had earned less than 46 credits. of these 36 students, 11 (31%) volunteered to be interviewed. their mean age was 22.5 years, ranging from 20 to 26. altogether 55% of these participants were male and 45% female. similarly to the humanities sample, male students were over-represented, their proportion of the whole cohort being 43%. in the results section, the humanities students are referred to as h, and law students as l. 2.2 materials research on procrastination has mainly applied quantitative approaches in which the data have been collected through various self-report instruments. to our knowledge, only four studies have applied a qualitative approach in order to explore procrastination or delay in academic studying (grunschel et al., 2013b; klingsieck, grund, schmid & fries, 2013; patrzek, grunschel & fries, 2012; schraw et al., 2007). these studies vary in their definitions of procrastination and in their qualitative methodological designs. furthermore, schraw et al. (2007) questioned whether self-report instruments could capture all of the possible dimensions of procrastination and therefore chose a qualitative approach. we also opted for a qualitative approach, but applied it differently compared to the four above-mentioned qualitative studies. our study’s sample comprised students whose study pace had been slow during their first university year, but we did not particularly ask the students whether, when or how they procrastinated. instead, we used in-depth interviews to explore the study processes as well as the expectations, experiences, interest and motivation of students who had studied slowly during their first year. further, we applied a qualitative person-oriented approach (e.g., vanthurnout, 2011), meaning that we used individual students as units of analysis. interviewing university students after their first study year made it possible to expand the scope of our investigation of procrastination and strategic delay from individual assignments or courses to a whole academic year. the participants volunteered to be interviewed after their first study year. the data collection was approved by the faculties. the students were informed that the results of the study would be used to enhance the program design and the development of the teaching-learning environments. the students gave their informed consent to participate in the study and were told that they could withdraw from it at any time. the interviews concentrated on three broad themes, which were based on previous research on motivational, volitional and situational factors related to dilatory behaviour. the first two explored the motivational and volitional dimensions of dilatory behaviour, whereas the third focused on the situational dimensions, as follows: a) students’ evaluations of themselves as university students, their study aims and future goals, motivation to study, and study success, b) descriptions of the students’ study processes and practices, and c) experiences of the teaching-learning environment. contrary to klingsieck et al. (2013) and grunschel et al. (2013b), we did not specifically ask for the students’ views, explanations or definitions of procrastination. instead, the interviews focused on their aims, study processes, evaluations and experiences of their first study year, and thus ‘circled around’ the lindblom-­‐ylänne  et  al       | f l r     52   phenomenon. by doing so we wanted to ensure that the students’ spontaneous and personal views were heard, and that our questions did not steer the students to explore their first study year specifically from the point of view of procrastination. consequently, our interviews were longer and less structured than those of schraw et al. (2007), klingsieck et al. (2013) and grunschel et al. (2013b). the fourth and fifth authors acted as interviewers; the fourth author interviewed the law students and the fifth the humanities students. the length of the interviews varied from approximately forty minutes to an hour, and the interviews were transcribed verbatim. for each profile, the most typical and representative extracts were selected. the selected extracts were translated into english. due to this translation process, the extracts do not represent authentic spoken english. to ensure the anonymity of the interviewees, the age and gender of the participants are not revealed. all students are referred to as ‘she’. 2.3 procedure we applied a theory-driven approach in which we used klingsieck’s model (2013) as a theoretical basis for the study. we developed the analysis process by using the model of deductive content analysis by elo and kyngäs (2008) as the starting point. all five authors were involved in the analysis process, which consisted of four phases. during phase 1 the data were prepared for the analysis and the unit of analysis was defined. as we applied a person-oriented approach, we used individual students or whole interviews as units of analysis. we also developed a categorisation matrix in which criteria were created for the seven parameters (table 1). klingsieck’s descriptions of the seven parameters of her model were quite short, which presented several obstacles with respect to our theory-driven approach. therefore it was important to create more explicit criteria. most challenging was to devise criteria for parameters 4 and 5, more precisely to define the difference between voluntary and unnecessary delay. defining this difference was particularly important, because according to klingsieck, parameters 3 and 4 represent strategic delay and parameters 5 to 7 procrastination. in addition, ‘unnecessary’ and ‘irrational’ in the description of the parameter 5 seemed very different from each other, which made this parameter quite broad in nature. finally, it was important to define what we meant by ‘act’. we defined it as study activities, processes, assignments and tasks during the first study year, in other words, not as a specific study task at a course level. lindblom-­‐ylänne  et  al       | f l r     53   table 1. the categorisation matrix. criteria for the seven parameters for the theory-driven data analysis. note. additional criteria for phase 3 are presented in italics. in phase 2 the categorisation matrix was used to review all data parameter by parameter. the first and fifth authors independently analysed the interview transcripts of all 28 students by checking each of the seven parameters of klingsieck’s model one by one, moving from the first parameter to the seventh. each interview transcript was therefore analysed independently by these authors in a cycle of seven rounds. in each round the data were coded for correspondence with criteria for each parameter. for example, for klingsieck’s seven parameters of delay (2013). criteria 1) an overt or covert act is delayed. evidence of delay, e.g., a low number of credits, unfinished courses, or not following timetables and not meeting deadlines. 2) the start or the completion of the act is intended. evidence of studies having been started and no evidence of dropping out from the program. 3) the act is necessary or of personal importance. at least one of the following elements: intention to graduate from university, commitment to studying or graduating, interest in studying or motivation to study. 4) the delay is voluntary and not imposed on oneself by external matters. no evidence of external reasons for the delay, such as sickness or family crisis. 5) the delay is unnecessary or irrational. evidence of possibilities to act or choose another way to proceed in studying, such as working fewer hours, spending less time on hobbies, or using more time for studying. no evidence of clear reasons for the delay. 6) the delay is achieved despite being aware of its potential negative consequences. evidence of the awareness of possible negative consequences. awareness is explicitly stated. 7) the delay is accompanied by subjective discomfort or other negative consequences. evidence of subjective discomfort or of negative consequences. discomfort is not necessarily expressed by using only negative words. therefore, jokes and laughing are scrutinized within their contexts. lindblom-­‐ylänne  et  al       | f l r     54   parameter 1 ‘an overt or covert act is delayed’ we coded all data segments in the interviews, which showed evidence of delay of an act, i.e., delay of study activities, processes, assignments or tasks during the first study year. the pieces of evidence were, for example, unfinished courses, not following timetables or not meeting deadlines. the two authors then compared their findings, and were unanimous about all participants having fulfilled the parameters 1 to 3: all students had delayed an overt or covert act (1), and intended to start and/or complete the act (2). it was also clear that the act had been necessary or of personal importance to all students (3). in addition, the two authors were unanimous regarding the voluntary versus involuntary nature of students’ dilatory behaviour (4). despite difficulties in creating criteria for ‘unnecessary’ or ‘irrational’ dilatory behaviour (5), the comparisons concerning this showed no differences between the two authors. however, the authors differed slightly regarding parameters 6 and 7. in some cases, it had been difficult to evaluate whether a student had been aware of the potential negative consequences (6). for phase 3 we clarified the criteria for this parameter so that a student’s awareness needed to be explicitly stated, and thus not inferred by the authors. furthermore, the two authors discussed criteria for ‘subjective discomfort and other negative consequences’ (7). in most of the cases, parameter 7 was easy to evaluate, although a number of problematic instances were found in which students had laughed and/or joked about their dilatory behaviour. thus, subjective discomfort seemed to be sometimes disguised by joking. the criteria for parameter 7 was further specified so that joking or laughing about dilatory behaviour needed to be scrutinized within their contexts, and that individual jokes would not automatically fulfil this parameter. at the end of this phase, all authors agreed on the adjusted criteria. phase 3 concentrated only on parameters 4 to 7, because the previous phase showed that each participant undoubtedly met parameters 1 to 3. the 17 interviews transcripts of the humanities students were independently analysed by the second author, and the 11 interview transcripts of the law students were analysed by the fourth. the analysis results were then compared and discussed between all authors. in phase 4 the student profiles were created, with all authors being involved. this phase concentrated particularly on analysing the individual profiles of students who met the criteria for strategic delay, but not all of the criteria for procrastination. 3. results 3.1 the three dilatory profiles the first aim of the study was to empirically test klingsieck’s model (2013), in which strategic delay and procrastination differ from any kind of delay in that the delay must be voluntary. the analysis revealed that one humanities student did not meet this criterion, because the slow progress of her studies due to external factors, more precisely, by unexpected family crises. altogether ten participants (37%) comprising six humanities and four law students, perfectly fit into klingsieck’s characterization of strategic delay (i.e., parameters 1 to 4). this profile was named strategic delayers. the following extract was very typical of the students in this profile: well, it [slow study pace] was mainly because of my own choices, i mean how to use your time…whether to study or go to work and so on, hobbies as well. they are just my own choices. of course the other school [completing studies at another university] affected, and partly also the fact that some courses took place simultaneously. (student h5) four students met all seven parameters of klingsieck’s (2013) model. these students were aware of the negative consequences of delay. further, their delay was unnecessary or irrational: i don’t know how the others…maybe they just are able to study and make it work, i can’t. in one course i should have completed two essays, but i just couldn’t do the other one at lindblom-­‐ylänne  et  al       | f l r     55   all. i finished one essay and got a poor grade, but that other…it was just, i just could not start. i even talked to the teacher and got good advice on how to start and what points to include, but somehow everything just vanished from my head. now i doubt whether i can ever succeed in writing essays, but i just have to try. (student h13) in addition, all four students expressed subjective discomfort, as the following extract shows: i’m independent and ambitious, but i easily get nervous and lose self-confidence. then i start feeling that i cannot do this, and that i’d just like to leave everything…or is it wise to spend all my time on this and just go crazy. so i dropped out from many courses and felt like a loser. (student h8) interestingly, 13 students’ profiles (48%) could not be determined as representing either strategic delayers or procrastinators. of these, seven students met the first six parameters meaning that these students were aware of the potential negative consequences of their dilatory behaviour, but had not experienced subjective discomfort or other negative outcomes. this was the only aspect separating these students from the four who met all seven parameters. we therefore created a procrastinators profile consisting of two subgroups: procrastinators not expressing subjective discomfort (n=6; 22%) and procrastinators experiencing subjective discomfort (n=4; 15%). this can be seen as following klingsieck’s model, because she mentions that procrastination often entails subjective discomfort or negative consequences, but not always. the subgroup procrastinators not expressing subjective discomfort consisted of one law and five humanities students, and the subgroup procrastinators experiencing subjective discomfort consisted of three humanities and one law student. despite the negative consequences of their unnecessary delay, procrastinators not expressing subjective discomfort did not exhibit anxiety or stress: i’m really bad in doing anything independently. maybe i do some assignments, but i don’t read much of the literature, which i should be able to do in this field. i’m also too lazy to reserve time for that. for example, i feel now so tired that there is no way i’m going to the library even though i should. instead, i go home and do anything else but study. maybe at some point i do some studying – a little [laughs]. (student h6) finally, the remaining seven students met the parameters 1 to 5: their dilatory behaviour had been unnecessary or irrational, which was not characteristic of the strategic delayers’ profile. these students had the possibility to choose another way to proceed in their studying, such as working fewer hours, spending less time on hobbies, or using more time to studying. however, there was no evidence of subjective discomfort and of being aware of potential negative consequences, which were typical of the procrastinators’ profile. these seven students seemed to fall between the strategic delayers and procrastinators, forming another profile we termed unnecessarily delaying students (n=7, 26%). two students represented humanities and five law. the following extract was very typical of them: i’m doing ok and have liked it here [at the faculty of law], but it was quite a surprise how much one should read and study for exams. i’ve had many other things to do, work and hobbies. i kept up my study pace quite nicely during the fall semester, but in the spring i started to slip. i realised too late that i didn’t start early enough and hadn’t read enough for the exams. (student l1) figure 1 summarises how klingsieck’s seven parameters (2013) were met in the humanities and law students’ dilatory profiles. parameters 1 to 3 were met in all profiles. strategic delayers met parameters 1 to 4 and unnecessarily delaying students parameters 1 to 5. the first subgroup of procrastinators, i.e., procrastinators not expressing subjective discomfort, met parameters 1 to 6, and the second subgroup of procrastinators, i.e., procrastinators experiencing subjective discomfort, met all seven parameters. lindblom-­‐ylänne  et  al       | f l r     56      figure 1. the dilatory profiles created on the basis of klingsieck’s (2013) seven parameters (n=28). 3.2 motivational, volitional and situational factors related to strategic delay and procrastination the second aim of the study was to explore how motivational, volitional and situational factors are related to procrastination and strategic delay. next, we explore each dilatory profile in more detail from the point of view of these factors. 3.2.1 strategic delayers the strategic delayers’ evaluations of themselves as students, their study experiences, and their experiences of the teaching-learning environment were positive. their strong volition was expressed by good self-regulation and time management, as shown in the following extract: i think i have good time-management skills, because i have been able to do a lot of sports and work while studying. i think there is a nice balance now and i’m quite happy about what i’m able to do in one week. i try to plan my schedule about a month ahead. (student h2) some students had encountered minor difficulties in self-regulation or time management, but had already sought help or changed their study practices after reflecting upon their life situation and study plans, which also reflects volition: at the beginning of the fall semester i lost my study rhythm, but now i have found a good one. i realised that i couldn’t be successful in studying while both working part-time and doing a lot of sports, so i had to lessen my exercising hours. (student l9) the students in this profile were able to describe their study processes and practices in greater detail than those in the other profiles. furthermore, strategic delayers seemed to be more successful at combining their studies with family and/or working life. all strategic delayers also showed personal interest in studying as well as intrinsic motivation. the following extract is very representative of this profile: 0   1   2   3   4   5   6   7   procras2nators  expressing   subjec2ve  discomfort  (n=4)   procras2nators  not  experiencing   subjec2ve  discomfort  (n=6)   unnecessarily  delaying  students   (n=7)   strategic  delayers  (n=10)   no  strategic  delay  or   procras2na2on  (n=1)   parameters   di la to ry  p ro fi le s   1 act is delayed 2 act isintended to be started / completed 3 act is necessary 4 act is voluntary 5 act is unneccesary 6 delay despite being aware of consequences 7 subjective discomfort lindblom-­‐ylänne  et  al       | f l r     57   i chose this field as a result of my own personal interests. the subject has interested me through my whole school history. it was one of my favourite subjects at school. in addition, i’m devoted to a hobby, which even increases my interest in studying, because it gives me a personal perspective on this field. (student h9) 3.2.2 unnecessarily delaying students unnecessarily delaying students shared the same positive study experiences and high interest and motivation with strategic delayers, but seemed to show weaker volition, as the following typical extract shows: there seems to be all kinds of disrupting factors…i’m sometimes nit in the right mood to study effectively. it’s difficult to describe this feeling, it’s a kind of lack of being able to concentrate…i would say that my biggest problem is to begin reading. after i have started, then it becomes easier, but then something interrupts my studying again. (student l11) typical of unnecessarily delaying students was also a discrepancy between their own study objectives and the actual study practices, in other words an intention-action gap: i’m very good at making plans, but quite poor at executing them. i try to allow myself free time as well, but studying seems to steal it. i try to use my time effectively, but because i work part-time, it’s often difficult to really have enough time for everything. (student h16) while strategic delayers had succeeded through their time management and self-regulation in combining studies, family life and work, unnecessarily delaying students had more difficulty doing so, which resulted in slower study progress. 3.2.3 procrastinators selecting one’s own field of study had not been a clear or easy task for the students in both subgroups of the procrastinators profile. these students had selected their major subjects either on the basis of success in this subject in high school, or because they could not decide upon a better option. they might also have seriously considered other disciplines before finally deciding, as the following extracts indicate: when i realised that i don’t have the skills and ambition [to reach my dream profession], i created a backup plan. i thought of different options, but this field has been my choice for a couple of years now. (student h6, subgroup procrastinators not expressing subjective discomfort) furthermore, most students did not have clear plans for the future, as the following extract shows: well, it's not very convincing [future employment]. everyone asks me what i’m going to become when i graduate, and i cannot answer. maybe something. maybe a rare type of expert, but who could hire someone like that? (student h13, subgroup procrastinators experiencing subjective discomfort) time management was also difficult for all students in this profile: just thinking about a calendar gives me the creeps. so i didn’t have a calendar, because i felt it would control my life. however, i finally gave in very reluctantly, and started to use a mobile calendar. that, however, was actually a good thing, because when it beeps it reminds me of things. maybe i just feel that others have more free time than i have. maybe others are just better at organising. because i’m less focused as a person, my use of time is not efficient, and i realise that. (student h12, subgroup procrastinators not expressing subjective discomfort) lindblom-­‐ylänne  et  al       | f l r     58   procrastinators had also missed lectures and stopped attending some courses. further, their study experiences seemed less positive than those of the first two dilatory profiles. however, the subgroups differed from each other in terms of their study experiences in that procrastinators experiencing subjective discomfort expressed more negative emotions and had less positive experiences than procrastinators not expressing subjective discomfort. clear differences were noted between the two subgroups in self-efficacy beliefs, experienced stress and study motivation. procrastinators experiencing subjective discomfort were less motivated to study than procrastinators not expressing subjective discomfort: the most important reason for my lowered interest was a feeling of strain or burden. now when i think about it logically, there were no clear reasons for that feeling, but what can you do about your feelings? (student h3, subgroup procrastinators experiencing subjective discomfort) furthermore, procrastinators experiencing subjective discomfort seemed to lack the ability to evaluate their own knowledge and skills, whereas procrastinators not expressing subjective discomfort had a more realistic view of themselves as students, describing themselves honestly and in a manner that showed no without anxiety or stress: lazy. aimless. floating. so i’m not the kind of person who has clear long-term objectives. i could be more active. in high school i liked that someone was watching and monitoring me. at university i should also have someone with authority to push me forward. you know, it’s too easy here to think that i will do this the next year, so delaying is not that harmful. i know that some students succeed at being efficient and organized, but i’m not the only one like this. (student h12, subgroup procrastinators not expressing subjective discomfort) the first study year of procrastinators experiencing subjective discomfort had been much more difficult than they had anticipated, and they were puzzled by the problems that had arisen. the vicious circle of procrastination resulted from a combination of lack of regulation skills, high workloads, lack of interest, low self-efficacy beliefs, and exhaustion, as the following typical extract indicates: from primary to upper secondary school i was a really good student, but now i feel that i can’t learn anything about any topic. this depresses me. i don’t have enough time to really learn something, and that feels bad. maybe i’m aiming too high, and when i can’t reach my aims, i get depressed. (student l4, subgroup procrastinators experiencing subjective discomfort) the students were unevenly distributed in the profiles in terms of their discipline, except for the profile strategic delayers, in which about 37% of both humanities and law students belonged. the humanities students were over-represented in both procrastinators subgroups. almost half of the law students belonged in the profile unnecessarily delaying students whereas only 13% of the humanities belonged in this profile. 4. discussion the present study empirically tested klingsieck’s (2013) theoretical model, and provided support for the suggestion to differentiate procrastination from strategic delay. interestingly, our results showed that dilatory behaviour is even more complex than klingsieck suggests. our in-depth qualitative analyses revealed forms of dilatory behaviour lying somewhere between procrastination and strategic delay, that do not meet the criteria for either strategic delay or procrastination: unnecessarily delaying students showed no lindblom-­‐ylänne  et  al       | f l r     59   awareness of potential negative consequences and did not exhibit subjective discomfort, and procrastinators not expressing subjective discomfort did not meet klingsieck’s (2013) seventh criterion of subjective discomfort or other negative consequences. in addition, the results interestingly showed only a small percentage of students meeting all seven criteria of procrastination, making this form of dilatory behaviour the least common in our data. the strategic delayers had good self-regulation and time-management skills as well as self-efficacy for self-regulation, which was shown by klassen et al. (2008) to impede procrastination. these students also exhibited strong achievement goals as well as interest and intrinsic motivation with respect to studying, all of which has been shown by howell and buro (2009) to reduce procrastination. furthermore, they exhibited good metacognitive and reflective skills in evaluating and developing their study processes and practices. these students’ dilatory behaviour was due to their life situations and how they had prioritised their study tasks. unnecessarily delaying students were also motivated to study and showed an interest in their majors, as was the case with strategic delayers. many had quite clear study plans as well. however, typical of these students was an intention-action gap: they often could not execute their study plans, which indicates problems in self-regulation and time management. this is in line with studies showing a relation between self-regulation problems and dilatory behaviour (e.g., corkin et al., 2011; wolters, 2003). both subgroups of the procrastinators profile lacked self-regulation and time-management skills, had low self-efficacy for self-regulation, and interest in their major subject was lower than that of the two other profiles. they also lacked intrinsic motivation and clear goals. however, procrastinators not expressing subjective discomfort seemed more motivated and interested than procrastinators experiencing subjective discomfort. these factors are in line with research showing that low intrinsic study motivation, problems in self-regulation, poor time-management skills and weak self-efficacy beliefs to be key factors in promoting procrastination (e.g., grunschel, et al., 2013a; lee, 2005; pychyl, et al., 2000; steel, 2007; strunk, et al., 2013; tice & baumeister, 1997; wolters, 2003). procrastinating students often mentioned personal characteristics when explaining their slow study pace, which could indicate trait procrastination (schouwenburg, 1995). this is consistent with research that has emphasized the maladaptive nature of procrastination (e.g., pychyl, et al., 2000; schouwenburg, 1995; solomon & rothblum, 1984; steel, 2007). shraw et al. (2007) have reported that some students postpone their studying because of fatigue and burnout. wäschle et al. (2014) showed that low goal achievement decreased perceived self-efficacy and increased procrastination. according to them, this might be due to both an expectation of repeated failure as well as negative emotions, which was also evident in our ‘procrastinating’ students. wäschle et al. (2014, 112) further showed that “instead of increasing their learning and raising cognitive strategy use, these students tended to irrationally postpone their studying.” the procrastinating students in our study also seemed to ‘freeze’ instead of act when confronting study problems, thus indicating a self-handicapping strategy (eerde, 2003; garcia & pintrich, 1994; howell & watson, 2007). other interesting differences were found between the humanities and law students. as mentioned earlier, the average graduation time in humanities is longer than in law. our results were in line with this, because the percentage of humanities students was higher in the procrastinators profile, whereas that of the law students was higher in the unnecessarily delaying students profile. the difference is probably due to situational factors: the law curriculum consists largely of mandatory legal courses and leaves little freedom of choice, whereas humanities students must make their own decisions concerning their minors. thus the results suggest that having the freedom to choose between numerous possibilities may promote procrastination, particularly for students with low self-regulation skills. however, because of the small sample size, we cannot make any generalisations of the relation between study context and the nature of dilatory behaviour. lindblom-­‐ylänne  et  al       | f l r     60   our study has several limitations. the number of participants was quite low and the students represented a narrow range of academic disciplines, i.e., humanities and law. in addition, even though our qualitative approach and means of data collection were chosen deliberately, it is possible that memory distortion may have affected our data. utilising the experience-sampling method developed by csikszentmihalyi, larson and prescott (1977) combined with interviews might have improved the participants’ memories of their first study year. pychyl et al., (2000) have successfully applied this method in procrastination research. another option would have been to complement interviews with other kinds of selfreport methods, such as self-monitoring protocols (wäschle et al., 2014) or learning diaries. however, as all previously mentioned data-collection methods are based on self-reports, it would be important in the future to complement self-report data with evaluations by tutors or study counsellors. it is also possible that errors were made in assigning students to dilatory profiles. however, we tried to minimise the errors by involving all five authors to independently assign students to the dilatory profiles. much of the research to date on procrastination and dilatory behaviour has applied mainly quantitative approaches and has focused on short-term procrastination. our theory-driven, person-oriented approach and our focus on long-term procrastination yielded a profound understanding of factors related to slow study progress among university students. in our view, future research into dilatory behaviour should endeavour either to be qualitative or use mixed-method or multi-method designs. in this way it is possible to better capture the richness and variation in dilatory behaviour. the results of the study can be applied to support the individual study paths of university students. the results imply that students representing different dilatory profiles need different kind of support during their studies. while strategic delayers might do well in a study environment offering ample alternatives and a freedom of choice, for unnecessarily delaying students this kind of an environment can be more harmful. unnecessarily delaying students can benefit from support for developing their self-regulation and timemanagement skills. further, the results indicate that procrastinators would have needed help from the beginning of their university studies, and during the spring term the latest, because the harmful effects procrastination were clearly visible after the first study year. these students’ self-efficacy beliefs had already weakened and they had started to doubt their skills, motivation and interest in their study fields. therefore, it is important to develop such study-counselling practices that can help diagnosing and solving students’ study problems at an early stage. keypoints the study empirically tested klingsieck’s (2013) theoretical model, and provided support for the suggestion to differentiate procrastination from strategic delay. a qualitative approach can better capture the richness of dilatory behaviour. the theory-driven, person-oriented approach and the focus on long-term procrastination yielded a profound understanding of factors related to slow study progress among university students. references ames, c. & archer, j. (1988). achievement goals in the classroom: students’ learning strategies and motivation processes. journal of educational psychology, 80(3), 260-267. bandura, a. (1997). self-efficacy. the exercise of control. new york: w. h. freeman and company. blunt, a., & pychyl, t. (2000). task aversiveness and procrastination: a multi-dimensional approach to task aversiveness across stages of personal projects. personality and individual differences, 28, 153-167. cheng, p.y., & chiou, w.-b. (2010). achievement, attributions, self-efficacy, and goal setting by accounting undergraduates. psychological reports, 106(1), 1-11. lindblom-­‐ylänne  et  al       | f l r     61   choi, j.n., & moran, s. v. (2009). why not procrastinate? development and validation of a new active procrastination scale. the journal of social psychology, 149(2), 195–211. chu, a.h.c., & choi, j. n. (2005). rethinking procrastination: positive effects of “active” procrastination behavior on attitudes and performance. the journal of social psychology, 145(3), 245–264. corkin, d, yu, s., & lindt, s. (2011). comparing active delay and procrastination from a self-regulated learning perspective. learning and individual differences 21, 602–606. doi:10.1016/j.lindif.2011.07.005 csikszentmihalyi, m., larson, r., & prescott, s. (1977). the ecology of adolescent activity and experience. journal of youth and adolescence, 6, 281-294. eerde, w. (2003). a meta-analytically derived nomological network of procrastination. personality and individual differences, 35, 1401-1418. doi: 10.1016/s0191-8869(02)00358-6 elliott, a., & harackiewicz, j. (1994). goal setting, achievement orientation, and intrinsic motivation: a meditational analysis. journal of personality and social psychology, 66(5), 968-980. elo, s. & kyngäs, h. (2008). the qualitative content analysis process. journal of advanced nursing, 62(1), 107-115. doi: 10.1111/j.1365-2648.2007.04569.x. ferrari, j. r., johnson, j. l., & mccown, w. g. (1995). an overview of procrastination. in j. r. ferrari, j. l. johnson, & w. g. mccown (eds.). procrastination and task avoidance. theory, research and treatment (pp. 1-20). new york: plenum press. ferrari, j. r., & tice, d. m. (2000). procrastination as a self-handicap for men and women: a task-avoidance strategy in a laboratory setting. journal of research in personality, 34, 73–83. doi:10.1006/jrpe.1999.2261 garcia, t., & pintrich, p. r. (1994). regulating motivation and cognition in the classroom: the role of self schemas and self-regulatory strategies. in d. h. schunk & b. j. zimmerman (eds.) self-regulation of learning and performance. issues and educational applications (pp. 127–153). new jersey: lawrence erlbaum associates, inc., publishers. grunschel, c., patrzek, j., & fries, s. (2013a). exploring reasons and consequences of academic procrastination: an interview study. european journal of psychology of education, 28(3), 841-861. doi 10.1007/s10212-012-0143-4 grunschel, c., patrzek, j., & fries, s. (2013b). exploring different types of academic delayers: a latent profile analysis. learning and individual differences, 23, 225-233. doi: 10.1016/j.lindif.2012.09.014. heikkilä, a., & lonka k. (2006). studying in higher education: students’ approaches to learning, self regulation, and cognitive strategies. studies in higher education, 31, no. 1, 99–117.     doi: 10.1080/03075070500392433. howell, a., & buro, k. (2009). implicit beliefs, achievement goals, and procrastination: a mediational analysis. learning and individual differences, 19, 151-154. doi:10.1016/j.lindif.2008.08.006 howell, a., & watson, d. (2007). procrastination: associations with achievement goal orientation and learning strategies. personality and individual differences, 43, 267-178. doi:10.1016/j.paid.2006.11.017. jones, e. & berglas, s. (1978). control of attributions about the self through self-handicapping strategies: the appeal of alcohol and the role of underachievement. personality and social psychology bulletin, 4(2), 200-2006. klassen, m., krawchuk, l., & rajani, s. (2008). academic procrastination of undergraduates: low self efficacy to self-regulate predicts higher levels of procrastination. contemporary educational psychology, 33, 915-931. doi:10.1016/j.cedpsych.2007.07.001. klingsieck, k. (2013). procrastination. when good things don’t come to those who wait. european psychologist, 18(1), 24–34. doi: 10.1027/1016-9040/a000138. klingsieck, k., grund, a., schmid, s., & fries, s. (2013) why students procrastinate: a qualitative approach. journal of college student development, 54(4), 397-412. doi: 10.1353/csd.2013.0060 lee, e. (2005). the relationship of motivation and flow experience to academic procrastination in university students. the journal of genetic psychology, 166, 5-14. doi:10.3200/gntp.166.1.5-15. newman, r. s. (1994). adaptive help seeking: a strategy of self-regulated learning. in d. h. schunk d. h., & b. j. zimmerman (eds.) self-regulation of learning and performance. issues and educational applications (pp. 283–301). new jersey: lawrence erlbaum associates, inc., publishers. patrzek, j., grunschel, c., & fries, s. (2012). academic procrastination: the perspective of university lindblom-­‐ylänne  et  al       | f l r     62   counsellors. international journal for the advancement of counselling, 34(3), 185-201. doi: 10.1007/s10447-012-9150-z. pintrich, p. r. (1995) understanding self-regulated learning. in p. r. pintrich (ed.) understanding self regulated learning (pp. 3-12). san francisco: jossey-bass publishers. pintrich, p. r. (2004). a conceptual framework for assessing motivation and self-regulated learning in college students. educational psychology review, 4, 385-408. pychyl, t. a., morin, r. w., & salmon, b. r. (2000). procrastination and the planning fallacy: an examination of the study habits of university students. journal of social behavior and personality, 15, 135–150. rebetez, m., rochat, l., & van der linden, m. (2015). cognitive, emotional, and motivational factors related to procrastination: a cluster analytic approach. personality and individual differences 76, 1–6. http://dx.doi.org/10.1016/j.paid.2014.11.044. rothblum, e. d., solomon, l. j., & murakami, j. (1986). affective, cognitive, and behavioral differences between high and low procrastinators. journal of counseling psychology, 33(4), 387-394. schouwenburg, h. c. (1995). academic procrastination: theoretical notions, measurement, and research. in j. r. ferrari, j. l. johnson, & w. g. mccown (eds.) procrastination and task avoidance. theory, research and treatment (pp. 71–96). new york: plenum press. schraw, g., wadkins, t. & olafson, l. (2007). doing the things we do: a grounded theory of academic procrastination. journal of educational psychology, 99, 12–25. doi: 10.1037/0022-0663.99.1.12. solomon, l. j. & rothblum, e. d. (1984). academic procrastination: frequency and cognitive-behavioral correlates. journal of counseling psychology, 31, 503–509. steel, p. (2007). the nature of procrastination: a meta-analytic and theoretical review of quintessential self regulatory failure. psychological bulletin, 133, 65–94. doi: 10.1037/0033-2909.133.1.65 strunk, k., cho, j., steele, m., & bridges, s. (2013). developing and validation of a 2 x 2 model of time related academic behaviour: procrastination and timely engagement. learning and individual differences, 25, 35-44. doi: 10.1016/j.lindif.2013.02.007. tice, d, m., & baumeister r, f. (1997). longitudinal study of procrastination, performance, stress and health: the costs and benefits of dawdling. psychological science, 8(6), 454–458. vanthurnout, g. (2011). patterns in student learning. exploring a person-oriented and a longitudinal research perspective. antwerpen-apeldoom: garant. doctoral thesis. university of antwerpen, belgium. vermunt, j., d., h., m., & van rijswik, f., a., w., m. (1988). analysis and development of students’ skill in self-regulated learning. higher education 17, 647–682. vermunt, j., d., & verloop n. (1999). congruence and friction between learning and teaching. learning and instruction 9, 257–280. wolters, c. a. (2003). understanding procrastination from a self-regulated learning perspective. journal of educational psychology, 95(1), 179–187. doi: 10.1037/0022-0663.95.1.179. wäschle, k., allgaier, a., lachner, a., fink, s., & nückles, m. (2014). procrastination and self-efficacy: tracing vicious and virtuous circles in self-regulated learning. learning and instruction, 29, 103 114. http://dx.doi.org/10.1016/j.learninstruc.2013.09.005. zimmerman, b. j. (1994). dimensions of academic self-regulation: a conceptual framework for education. in d. h. schunk d. h., & b. j. zimmerman (eds.) self-regulation of learning and performance. issues and educational applications (pp. 3–21). new jersey: lawrence erlbaum associates, inc., publishers. microsoft word fleckenstein_publication.docx             frontline learning research vol. 3 no. 2 (2015) 27 46 issn 2295-3159   corresponding author: johanna fleckenstein, leibniz institute for science and mathematics education, department of educational research, olshausenstr. 62, 24118, kiel, germany. email address: fleckenstein@ipn.uni-kiel.de doi: http://dx.doi.org/10.14786/flr.v3i2.162 what works in school? expert and novice teachers’ beliefs about school effectiveness johanna fleckensteina, friederike zimmermannb, olaf köllera, jens möllerb aleibniz institute for science and mathematics education, germany bkiel university, germany article received 4 april 2015 / revised 4 april 2015 / accepted 27 april 2015 / available online 4 june 2015 abstract in 2009, john hattie first published his extensive metasynthesis concerning determinants of student achievement. it provides an answer to the question: “what works in school?” the present study examines how this question is answered by preand in-service teachers, how their beliefs correspond to the current state of research and whether they differ according to the teachers' level of expertise. thus, it takes on a novel approach as it draws on data from two sources in the field of education -empirical research and teachers’ beliefs -and examines their similarities and differences. the teachers’ beliefs were elicited by asking n = 729 participants to estimate the effect sizes of several determinants of student achievement. those were compared to the empirical effect sizes found by hattie (2009). profile correlations showed that expert teachers’ beliefs are more congruent with current research findings than those of novice teachers. we further examined where expert and novice teachers’ beliefs differ substantially from each other by using confirmatory factor analysis (cfa) and comparing group means in latent variables. our findings suggest that teachers’ beliefs about school effectiveness are related to professional experience: expert teachers showed a stronger overall congruence with empirical evidence, scoring higher in achievement-related variables and lower in variables concerning surfaceand infrastructural conditions of schooling as well as student-internal factors. results are discussed with regard to teacher-education practices that emphasize research findings and challenge existing beliefs of (prospective) teachers. keywords: teacher beliefs; teacher education; professional competence; school effectiveness   fleckenstein  et  al   f   | f l r     28   teachers’ beliefs are often guided by subjective experience rather than by empirical data. thus, it is to be expected that they generally diverge from research findings. this assumption also pertains to the specific case of teachers’ beliefs that manifest in their response to the question: “what works in school?” school effectiveness research has tried to answer this question, the latest attempt being hattie’s (2009, 2012) metasynthesis of factors that influence student achievement. however, little is known of the practitioner’s answer to one and the same question: what factors do (expert and novice) teachers believe to influence student achievement? where do their beliefs differ from research findings substantially? these questions are highly relevant as teachers’ beliefs have been shown to influence teaching and learning. if they differ substantially from results of school effectiveness research, we have reason to assume a negative effect on educational outcomes. for example, if a teacher undervalues a particular teaching method or overvalues surface structural aspects like class size, this could be a serious threat to effective teaching. an investigation into teachers’ beliefs concerning school effectiveness can show us where these discrepancies are and thus inform teacher education practice and classroom instruction. teachers’ beliefs are typically represented as part of a multi-dimensional construct of teachers’ professional competence (baumert & kunter, 2006). they influence teachers’ perceptions and judgments and, consequently, affect their classroom instruction (calderhead, 1996; pajares, 1992). beliefs are subject to change and thus can be expected to differ in expert and novice teachers. in the present study, we are particularly interested in a specific subset of teachers’ beliefs, namely those regarding the effectiveness of schooland education-related factors. in the last few decades, there has been a lot of research concerning the question of what works in school – and what does not (fraser, walberg, welch & hattie, 1987; walberg, 1986; hattie, 2003; wang, haertel, & walberg, 1990, 1993). the most recent and extensive example of this is hattie’s synthesis of meta-analyses (2009, 2012), in which he examines the influence of 138 factors on student achievement. teachers should be familiar with such findings of school effectiveness research in order to make informed decisions and focus on the most effective interventions. for this kind of evidence-based practice, however, teachers not only have to know about such findings but actually to believe that they are true. hence, we asked pre-service (“novice”) and in-service (“expert”) teachers for their beliefs about the efficacy of certain determinants of student achievement. these ratings of our expert and novice teachers were then compared with each other and contrasted with the findings of hattie (2009). in the following we will provide a brief theoretical background on teachers’ beliefs. since the body of research on teachers’ beliefs is quite extensive we concentrate on the following relevant aspects: the theoretical construct of teachers’ beliefs, the influence of teachers’ beliefs on classroom processes and student outcomes, the issue of teachers’ beliefs being guided by subjective experience rather than objective fact, and the general differences in beliefs of novice and expert teachers. subsequently, we locate teachers’ beliefs within a model of teachers’ professional competence, which – in accordance with the prevailing expert(-novice) paradigm – suggests the malleability of all its components. the beliefs examined here focus on determinants of student achievement, therefore, we will also show the most important and recent results of school effectiveness research. as hattie (2009) serves as the basis of our study, particular attention is paid to his comprehensive aggregation of existing meta-analyses (metasynthesis; see zell & krizan, 2014). 1. theoretical background 1.1 teachers’ beliefs according to barcelos (2003), beliefs are types of thoughts which provide a basis for decisions and actions. harvey (1986) described a belief system as “a set of conceptual representations which signify to its   fleckenstein  et  al   f   | f l r     29   holder a reality or a given state of affairs of sufficient validity, truth or trustworthiness to warrant reliance upon it as a guide to personal thought and action” (p. 146). beliefs in general are thought of as psychologically held understandings, premises, or propositions about the world that a person perceives as being veritable (richardson, 1996). teachers’ beliefs can be seen as a substructure of the general belief system. they consist of beliefs that serve as a guide when dealing with schooland instruction-related situations. in educational settings, haney et al. (2003) defined beliefs as “one’s convictions, philosophy, tenets, or opinions about teaching and learning” (p. 367). as such, teachers’ beliefs may include subjective theories about how students learn, what a teacher should or should not do and which instructional strategies work effectively. the last few decades have brought out a substantial body of research on the beliefs of teachers (for comprehensive research reviews see calderhead, 1996; fang, 1996; pajares, 1992; nespor, 1987; richardson, 1996; stuart & thurlow, 2000; verloop et al., 2001; wenden, 1999; woods, 1996; zheng, 2009). teachers’ beliefs influence perception and judgment which in turn guide their actions in the context of school and education (pajares, 1992). prior research has shown that teachers’ beliefs have a critical impact on the way they teach in the classroom, learn how to teach, and perceive educational reforms (m. borg, 2001; allen, 2002; s. borg, 2003; freeman, 2002; yook, 2010). other studies have shown the importance of teachers’ beliefs for student achievement (peterson, fennema, carpenter & loef, 1989; staub & stern, 2002). an especially well-researched issue in this context are teachers’ self-efficacy beliefs. according to tschannen-moran et al. (1998, p. 233), teacher efficacy is “the teacher’s belief in his or her capability to organize and execute courses of action required to successfully accomplish a specific task in a particular context”. such beliefs have been shown to critically influence a teacher’s performance and motivation (bandura, 1997; ross,1998; tschannen-moran & woolfolk hoy, 2001; tschannen-moran, woolfolk hoy & hoy,1998; woolfolk & hoy, 1990; woolfolk, rosoff & hoy, 1990; woolfolk hoy & davis, 2006) as well as his or her students’ achievement in school (bates, latham & kim, 2011; muijs & rejnolds, 2001; ross, 1992, 1998). against this background it becomes evident that the beliefs of teachers are an important issue for teaching and learning. the definitions above show that beliefs are always perceived to be correct by the individual. however, this is not necessarily the case: teachers’ beliefs – as a subgroup of beliefs in general – are especially likely to be flawed in the sense that they contradict empirical evidence. in the following we discuss why this is and in what way it can lead to substantial problems, in particular for novice teachers. teachers’ beliefs are highly subjective, tend to be persistent and develop at a rather early stage in life (lortie, 1975; pajares, 1992). this is partly due to the long-term experience with schools and classrooms during a teacher’s own time as a student, which serves as the starting point of his or her training and career. hattie (2009) calls this relatively stable system of beliefs the “grammar of schooling” (p. 5). it contains tacit and simplified notions of what a good teacher is and how students are supposed to behave (clark, 1988; nespor, 1987). as such, this belief system often diverges from empirical findings and can fatally influence the process of teaching (kunter & pohlmann, 2009). understanding and challenging one’s own beliefs is therefore considered an important aspect of teacher qualification (bromme, 1997; woolfolk-hoy, davis & pape, 2006). however, this does not seem to be an easy task, since not even the confrontation with dissonance (e.g., induced by empirical evidence that contrasts one’s beliefs) necessarily leads to a corresponding change in beliefs (hart, 2004; pajares, 1992). hattie (2009) claims that teachers experience almost everything they do in the classroom to have a positive influence on their students’ learning. the remarkable differences in effectiveness of their efforts are easily overlooked due to the lack of firsthand comparison, since teachers are usually confined to their own classroom. they see that what they do seems to work fine, as (almost) everything a teacher does leads to an increase in students’ achievement. thus, there is always anecdotal evidence for the effectiveness of certain methods, even though the actual extent to which they support students’ learning often differs dramatically. hattie describes this basic principle of teaching as “just leave me alone as i have evidence that what i do enhances learning and achievement” (hattie, 2009, p. 6). admittedly, this is a rather simplified account of   fleckenstein  et  al   f   | f l r     30   teachers’ self-efficacy beliefs. most teachers indeed struggle with the teaching methods they use, try out new things and reflect whether or not they seem to be working. however, they often do this within their own frame of reference, not based on research evidence. the fact that teachers’ beliefs do not necessarily coincide with empirical findings can be problematic if not reflected thoroughly. especially novices can have inadequate notions of what constitutes good teaching. weinstein (1989) found that on average, pre-service teachers overestimate affective and social variables of classroom instruction (such as patience and the ability to relate to children) and underestimate cognitive and academic variables (such as organization and challenging). however, when contrasted with the perspectives of educational policy makers and researchers, in-service teachers’ beliefs seem very similar to those of their less experienced colleagues. while the former two speak of good teaching in terms of outcomes in standardized assessment and direct instruction models (policy makers), as well as ‘masterful teachers’ with a whole set of well-defined professional skills (researchers; e.g. shulman, 1987), the latter two have a notion of ‘good teachers’ that can be described as “warm, caring individuals who enjoy working with children” (weinstein, 1989, p. 59). hogan, rabinowitz, and craven (2003) compared novice and expert teachers and found that student achievement was important for expert teachers, while novice teachers paid more attention to student interest. 1.2 teachers’ professional competence teacher training and professional development are central issues in the international discussion on teacher effectiveness (bauer & prenzel, 2012; cochran-smith & zeichner, 2005; darling-hammond & bransford, 2005). the development of standards in teacher education requires an explicit analysis of challenges teachers face in their everyday professional life. moreover, it demands a specification of competences necessary to master these challenges. the quest for the good teacher is not new; however, the specific facets of teachers’ professional competence and the premise that teachers’ cognitions are modifiable by means of training are direct results of a relatively recent objective in teacher(-education) research: the expert(-novice) paradigm (berliner, 2004; bromme, 1997, 2003; ericsson & lehmann, 1996; ericsson, charness, feltovich & hoffman, 2006). accordingly, this paper is based on the central assumptions that (a) good teachers are experts of learning and teaching, and (b) they achieve this expertise in the form of professional competence through continuous teacher education and professional experience. experts are roughly defined as “individuals who exhibit reproducibly superior performance on representative, authentic tasks in their field” (ericsson, 2006, p. 688). it is assumed that teachers’ expertise or professional competence is acquired throughout preand in-service training as well as by hands-on experience in the classroom (berliner, 2004). teachers’ professional competence is usually represented as a multi-dimensional construct. as such, based on the five core propositions of the national board for professional teaching standards (nbpts), baumert and kunter (2006) proposed a model of teachers’ professional action competence with four non-hierarchical dimensions: (1) specific declarative and procedural knowledge, which further distinguishes between content knowledge (ck), pedagogical knowledge (pk), and pedagogical content knowledge (pck) (shulman, 1986, 1987); (2) professional beliefs, values, subjective theories, normative preferences and objectives; (3) motivational orientations; and (4) meta-cognitive skills and professional self-regulation. in line with the expert(-novice) paradigm we assume that all of these competencies are subject to change throughout a teacher’s professional life. the boundaries between these categories of teachers’ cognition, however, are more or less fuzzy: knowledge – pk and pck in particular – and beliefs are strongly interrelated theoretical constructs, even though they rely on dissimilar epistemological notions, as „belief is based on evaluation and judgment; knowledge is based on objective fact“ (pajares, 1992, p. 313). one and the same response to a pedagogical question can either demonstrate well-founded knowledge or it can be based on subjective belief. the two answers may differ in their epistemological status, though their distinction – philosophically speaking – is a mere social construct: leatham (2006) argues that beliefs (things we just believe) and knowledge (things we   fleckenstein  et  al   f   | f l r     31   more than believe) can be viewed as complementary subsets of the things we believe. in comparison to belief, knowledge is characterized by a higher degree of certainty, for example, by being grounded in empirical evidence. in many empirical studies on teacher beliefs, however, the distinction between knowledge and beliefs is rather blurry. it is very difficult to distinguish whether teachers refer to their knowledge or beliefs when they plan, make decisions, or act in classroom (verloop, van driel, & miejer, 2001). in the present study, we use the concept of teachers’ beliefs to refer to cognitions of teachers that are subjective and normative in nature, while they may or may not coincide with the more objective construct of knowledge. 1.3 school effectiveness research the present study deals with teachers’ beliefs concerning factors of school effectiveness. thus, in the following we briefly summarize the central findings of school effectiveness research. we focus on proximal vs. distal aspects of schooling as there is a broad consensus concerning this dichotomy in the literature. the analyses performed in this paper focus on the broad categories school, teaching and student, so we aim at summarizing the comprehensive literature in school effectiveness research on this rather general level. subsequently, we give a more detailed overview of hattie’s 2009 metasynthesis, concentrating on those variables and categories that were used in our questionnaire to elicit teacher’ beliefs. in the last few decades, increased efforts were made in school effectiveness research to study the importance of a range of determinants on successful schooling. the question of what works in school (and what does not) has been the central issue of a number of meta-analyses and research syntheses (coleman et al., 1966; fraser et al., 1987; hattie, 2003, 2009, 2012; jencks et al., 1973; scheerens & bosker, 1997; seidel & shavelson, 2007; walberg, 1986; wang et al., 1990, 1993). those studies do not always agree on the specific size of an effect, however, there is a general tendency with regards to certain factors of school effectiveness. in general, the majority of these studies suggest that the amount of variance explained by proximal – schoolor classroom-related – variables is considerable, and has a greater influence on student learning than more distal aspects such as school system and educational policy (seidel & shavelson, 2007). the general emphasis on proximal variables was, for example, shown by scheerens and bosker (1997). they combined the results of three meta-analyses as well as a re-analysis of an international data set and found that school-organizational factors (e.g., monitoring/ evaluation, orderly climate), instructional conditions (e.g., opportunity to learn, homework), and aspects of structured teaching (e.g., feedback, cooperative learning) are a better explanation for the differences between the achievement of students than more distal aspects such as resource input factors (e.g., student-teacher ratio, teachers’ salary). moreover, the rank-ordering presented by wang et al. (1993) put student characteristics and classroom practices ahead of design of program and school demographics. they found particularly strong effects for the variables meta-cognition, classroom management and quantity of instruction. fraser et al. (1987) found the highest correlations with performance tests for variables related to student characteristics (especially cognitive ones), learning strategies, and structured or direct teaching. the results also revealed that open teaching and individualization are less powerful factors, at least when the dependent variable is (cognitive) achievement. similar findings were shown by walberg (1986). there have been many attempts to find a comprehensive consensus with regards to the effects of certain factors of school effectiveness. scheerens (2004) presented the effectiveness enhancing conditions of schooling in five review studies (cotton, 1995; levine & lezotte, 1990; purkey & smith, 1983; sammons, hillman & mortimore, 1995; scheerens, 1992): a consensus was reached with respect to many instruction-related factors, such as achievement orientation, high expectations, frequent testing/ monitoring, professional development, and structured or purposeful teaching. hattie’s (2009) synthesis of over 800 meta-analyses was one of the most recent milestones in school effectiveness research: 52,637 individual studies with over 83 million students were used in order to determine the relevance of 138 factors for student achievement. for each of these factors he determined   fleckenstein  et  al   f   | f l r     32   cohen’s d as the averaged effect size. as a convention in the school context, effect sizes of d > .40 are considered substantial, since this would imply greater effects than one year of average schooling (köller, 2012); hattie calls this the zone of desired effects. hence, the point of reference for the effectiveness of an innovation is not d = 0, but d = .40. hattie’s results were largely in line with the prior findings in school effectiveness research as described in the preceding paragraphs. he systematized the individual factors according to six superordinate categories: student, teacher, teaching, curriculum, school, and family. the central results can be summarized as follows: more or less ineffective factors with d < .40 were primarily infrastructural conditions of schooling, such as withinor between-class grouping, finances, and reduction of class size. moreover, aspects of the surface structure of teaching, which is often associated with progressive teaching approaches (e.g. open learning, multi-grade/-age classes, team teaching), did not show to be very effective either. these results may be surprising considering the socio-political discourse on education; however, against the background of modern classroom research they are to be expected: research has shown that successful learning can be better predicted by the deep structure of teaching and learning than by the surface structure (e.g., seidel & shavelson, 2007). the latter can be observed and described without much effort, while the former requires more elaborate assessment. the use of surface-structure learning methods is not beneficial by itself, but only if it affects the level of deep-structure cognitive processing (e.g., by giving constructive feedback or teaching meta-cognitive strategies). in line with prior research on school effectiveness, high effect sizes could also be shown for cognitive and emotional student characteristics (e.g. prior knowledge, motivation) and instructional, achievement-related variables such as direct instruction and high expectations of the teacher. in agreement with prior research, hattie’s findings suggest that more distal factors are less important than proximal factors and that the structural conditions of teaching are less important than the process of teaching itself. the results also highlight the importance of the students’ cognitive and noncognitive prerequisites for learning. 2. the present study in our study we attempted a direct comparison of the results of school effectiveness research (i.e. the effect sizes from hattie’s study) with the beliefs of novice and expert teachers (i.e. their ratings of effect sizes). this was a rather novel approach; however, wang et al. (1993) adopted a similar strategy when they compared the results of 91 meta-analyses with the ratings of experts in education, namely 61 distinguished educational researchers. the correlation they found between expert ratings and meta-analyses was .59 (p < .01). the authors concluded that there is a general agreement between expert ratings and the meta-analyses regarding the effect of different variables on student learning and their relative strength. while wang et al. (1993) examined the judgments of experts in educational research, our study dealt with the beliefs of teachers that are either enrolled in a teacher training program or work as teachers and school administrators. the objective of wang et al. was to build a knowledge base in school effectiveness research: they used three different methods – content analyses, expert ratings, and results from meta-analyses – to quantify the importance and consistency of variables that influence student learning. our objective, on the other hand, was to explicitly address the beliefs of those groups that actually are or are going to be working in the field and directly influence classroom processes. prior research on teachers’ beliefs (in general and especially those of novice teachers) has shown that they tend to be very subjective and are rather unlikely to be guided by empirical evidence, so we had reason to assume that is also the case also for beliefs concerning determinants of student achievement. thus, we expected our preand in-service teachers’ beliefs to differ more strongly from the findings of school effectiveness research than the ratings of expert researchers. with the epistemological question in mind that we raised above, one could also argue that wang et al. (1993) examined knowledge while we examined beliefs.   fleckenstein  et  al   f   | f l r     33   school effectiveness research gives us a good theoretical understanding of what works in school. however, we have reason to assume that teachers’ cognitions are not congruent with these findings, since their beliefs are at risk to be guided by subjective experience and beliefs rather than by empirical data. the present study focused on teachers’ beliefs about which factors determine their students’ achievement. hence, the central questions were: a) what are teachers’ beliefs about the impact of the above-mentioned factors on student achievement, and to what extent do these beliefs diverge from findings of empirical research (i.e., the effect sizes of hattie’s research synthesis)? b) what are the differences in the beliefs of novice and expert teachers on a latent level of meaningful factors of school effectiveness? 3. methods 3.1 sample the sample comprised n = 729 participants (64% female); n = 358 were in-service (“expert”) teachers and n = 371 pre-service (“novice”) teachers. teachers of the first group were in service at different schools in the federal states schleswig-holstein and hamburg, germany. of these participants 53 % were women, the medium age was m = 52.3 years (sd = 8.8) ranging from 28 to 64 years. this subgroup included teachers from different types of schools (primary and secondary).the data from this subsample was collected in the context of professional development lectures for in-service teachers. though attendance was not mandatory, the lectures were open for all teachers from the two states. the attending teachers can be considered true experts as many of them were in leadership positions at their schools (training supervision, school administration, etc.). the pre-service teachers were university students enrolled in the first year of a master of education (m.ed.) at a university in the northern part of germany. to illustrate the background of our sample we briefly outline a typical teacher training program in germany: at most german universities teacher training is composed of a three-year bachelor (b.a./ b.sc.) program and a two-year master (m.ed.) program. it includes the academic study of two scientific disciplines and didactics for the corresponding school subjects. in addition, students take a variety of courses in educational sciences. after university studies, students transfer to the more practical part of teacher training. they train teaching in schools for one to two years before they become proper teachers. our pre-service teachers had completed a bachelor program that prepared them for graduate studies in teacher education. they had done practical training in schools for a period of six weeks in total; however, the bachelor program was clearly focused on the theoretical study of the two scientific disciplines. only a small proportion of the degree was dedicated to introductory classes on educational sciences and didactics. thus, we can assume that their prior knowledge concerning these subjects was not very advanced. the percentage of female participants in this group was 71%. their mean age was m = 24.7 years (sd = 2.5), ranging from 23 to 39 years. the data was collected during a lecture on psychology in education, in which all students enrolled in the m.ed. were required to participate. 3.2 procedure in order to assess teachers’ beliefs about the effectiveness of factors for students’ achievement a questionnaire was developed based on 16 determinants of student learning (see table 2) selected from hattie (2009). criteria for the selection of items were the coverage of a large range of effect sizes (d = .01-.73) and   fleckenstein  et  al   f   | f l r     34   the coverage of the a priori categories school, teaching and student from hattie’s study. these categories were chosen as a focus of our study since there seemed to be the highest consensus about the extent of their impact on student learning among school effectiveness studies. moreover, we selected those variables that we assumed even inexperienced university students would be familiar with, as most of them are also a frequent issue in political and academic discourse. table 1 intervals of effect sizes and their interpretations by hattie (2009) and köller (2012)   range of effect sizes interpretation by hattie (2009) interpretation by köller (2012) d < 0 reverse effects harmful 0 ≤ d < 0.15/0.2 developmental effects not harmful, not helpful 0.15/0.2 ≤ d < 0.4 teacher effects a little helpful 0.4 ≤ d < 0.6 zone of desired effects helpful d ≥ 0.6 very helpful the questionnaire was administered in the context of a lecture on hattie’s study. first of all, participants were introduced to the concept of meta-analysis in general and the design of hattie’s research synthesis in particular. subsequently, they were familiarized with the concept of cohen’s d and its practical implications: the formula d = (mtest – mcontrol) / sdpooled was presented to the teachers and explained in detail with the help of examples. in order to give practical significance to the rather abstract notion of effect size the interpretation of effect sizes as presented in the rightmost column of table 1 was introduced and displayed during the completion of the questionnaire. the participants were asked to estimate the impact of each of the factors on a scale of effect sizes ranging from d = -0.4 to d = 1.0. the precise instruction was: “please estimate the effect of each of the factors below on students’ achievement”, followed by the list of variables. participants were briefly familiarized with the 16 factors, that is, short comments were given on what was meant by the factors. 3.3 statistical analyses first, group means were calculated on the level of the 16 manifest variables and further analyzed by comparing them to the effect sizes of hattie (2009). this was achieved by calculating the correlation coefficient pearson’s r for each person’s rating profile with the distribution of hattie’s effect sizes. the coefficients were transformed by fisher’s z in order to approximate constant variance for all values. the use of fisher’s z transformation is recommended when averaging correlation coefficients as the distribution of r is skewed (silver & dunlap, 1987). the resulting fisher’s z coefficients were aggregated per group (preand in-service teachers) and the resulting means (mz) of the two groups were compared using an independentsample t-test. thus, we could determine the difference in preand in-service teachers in terms of congruence with hattie’s results. second, we examined whether ratings on several individual items could be aggregated on a higher level, that is, in a latent variable model. confirmatory factor analysis (cfa) was carried out in mplus 7 (muthén & muthén, 1998-2012) in order to analyze the underlying latent structure of the data, which then served as a basis for comparisons between preand in-service teachers on the reduced number of meaningful categories on a higher level. the assumed factor structure was based on hattie’s (2009) a priori categorization of the variables. the factor teaching contained variables from hattie’s categories teaching and teacher, the factor school contained variables from hattie’s category school, and the factor student contained variables from hattie’s category student. modification indices indicated, however, that the factor teaching   fleckenstein  et  al   f   | f l r     35   was to be split in two (teaching and achievement). in the final measurement model we specified four latent variables using maximum likelihood estimation (ml). the latent variables were allowed to correlate and some error terms of manifest variables were allowed to covary if considered plausible. missing data (< 5%) were estimated with the help of the full information maximum likelihood (fiml) procedure. subsequently, in order to allow for meaningful comparisons between the subgroups, measurement invariance was tested using a multiple-group modeling approach (meredith & teresi, 2006). for a comparison of latent means across the two groups (prevs. in-service teachers) at least partial scalar invariance is required (byrne, shavelson & muthén, 1989). in multiple group analysis, when the specified model includes a mean structure, both the intercepts and factor loadings of the continuous factor indicators are held equal across groups to specify (scalar) measurement invariance. the intercepts of the factors are fixed at zero in the first group and are free to be estimated in the other groups. thus, differences between the two groups can be determined based on the latent factors. 4. results 4.1 descriptive statistics and item level analysis table 2 shows descriptive statistics for the two groups – preand in-service teachers – on item level as well as hattie’s (2009) research results. the factors are ranked by the size of their effect (d) on student achievement as found by hattie in his metasynthesis. in the following we elaborate on the descriptive results, especially on those variables with the highest and lowest ratings in each group. both groups seemed to believe in the importance of student variables, as they both showed the highest means on the factors motivation and attitude. ranked third by both preand in-service teachers was feedback. multi-grade/age learning (in-service group) and direct instruction (pre-service group), respectively, had the lowest ratings of all 16 factors. pre-service teachers had significantly higher effect sizes for the variables feedback, prior achievement, motivation, attitude, class size, co-/team teaching, within-class grouping, and open learning. in-service teachers’ beliefs showed higher effect sizes for direct instruction, high expectations, and selfconcept. table 2 hattie's effect sizes (d), group means (mgroup), and standard deviations (sd) factors dhattie min-service mpre-service feedback .73 .55 (.23) .62 (.24)a teaching meta-cognitive strategies .69 .53 (.25) .53 (.26) prior achievement .67 .32 (.23) .39 (.24)a professional development .62 .40 (.23) .42 (.25) direct instruction .59 .28 (.23)a .18 (.25) motivation .48 .63 (.24) .73 (.22)a expectations .43 .36 (.25)a .23 (.29) self-concept .43 .55 (.21)a .48 (.24) attitude .36 .56 (.24) .66 (.21)a frequent/effects of testing .34 .34 (.24) .34 (.26)   fleckenstein  et  al   f   | f l r     36   class size .21 .34 (.31) .59 (.30)a co-/team teaching .19 .37 (.27) .45 (.28)a within-class grouping .16 .48 (.27) .61 (.24)a problem-based learning .15 .52 (.23) .52 (.25) multi-grade/age classes .04 .19 (.26) .22 (.27) open learning .01 .29 (.27) .37 (.27)a a superscript characters indicate statistically significant (p < .01) higher mean effect sizes for the respective group (results of two-sample independent t-tests using bonferroni correction to account for the multiple comparisons problem) we determined bivariate correlations for each teacher’s rating profile on the one hand and hattie’s results on the other by calculating pearson’s r for each person. the coefficients were transformed by fisher’s z and aggregated per group. these means were then compared using an independent-sample t-test. for the pre-service teachers the mean (fisher z-transformed) correlation with hattie’s d was mz=.06 (sd =.32), for the in-service teachers it was mz = .23 (sd = .35). these group differences were statistically significant (t[715] = 7.12; p < .001; d = .51), indicating a substantially higher degree of conformity of the experts’ ratings with hattie’s results. 4.2 cfa and multiple-group analysis a priori, for the item pool selected for this study we assumed three latent factors based on hattie’s categorization of indicators: school, teaching/teacher and student (χ2[95]=586.36; cfi=.85; rmsea=.08; tli=.81; srmr=.08). empirically, however, a four-dimensional structure resulted in a better model fit (χ2[92] = 340.85; cfi = .92; rmsea = .06; tli = .90; srmr = .05). the factor teaching was split in two, separating the strongly achievement-focused indicators from more specific instructional teaching behaviors. the improvement in goodness-of-fit indices was substantial (δcfi > .01; δrmsea > .015) (cheung & rensvold, 2002), so we decided on the four-dimensional model (see table 3). residual correlations were allowed for some indicators with substantial covariance that was not explained by the latent factor. all items loaded significantly (p < .001) and almost all items loaded substantially (λ ≥ .4) on one of the latent factors. the only items with factor loadings slightly below the minimum value were prior achievement (λ = .37) and multi-grade/age classes (λ =.37). the former is the only indicator for the factor student that focuses on cognitive rather than motivational aspects of a student’s academic prerequisites. this might explain the low factor loading. the latter may have been a difficult concept for many of the participants as by far not all teachers encounter this instructional challenge throughout their careers. due to high modification indices, residual correlations were allowed for six item pairs (multigrade/age classes with open learning and co-/team teaching, class size with co-/team teaching and withinclass grouping, teaching meta-cognitive strategies with feedback, motivation with attitude). the majority of these modifications were performed within the factor structure. they seemed to be theoretically sound as certain infrastructural conditions of schooling (class size, multi-grade/age classes) are strongly associated with or even demand certain surface-structural aspects of learning (open learning; co-/team teaching; withinclass grouping). teaching meta-cognitive strategies and feedback are both direct and concrete instructional measures of the teacher, motivation and attitude towards subject refer to very similar student-internal constructs (as opposed to the other respective indicators). the allowed covariances were the same for both models (three and four latent factors) that were tested.   fleckenstein  et  al   f   | f l r     37   table 3 standardized factor loadings matrix of the cfa teaching achievement school student 1 feedback 0.62 2 teaching meta-cognitive strategies 0.68 3 professional development 0.70 4 problem-based learning 0.69 5 direct instruction 0.52 6 expectations 0.69 7 frequent/effects of testing 0.48 8 multi-grade/age classes 0.37 9 open learning 0.65 10 class size 0.51 11 co-/team teaching 0.56 12 within-class grouping 0.77 13 motivation 0.67 14 prior achievement 0.37 15 self-concept 0.50 16 attitude towards subject 0.57 in the following we will explain the four dimensions in more detail: the factor school comprises infrastructural conditions of schooling (class size; multi-grade/age classes) and the surface-structure of learning (open learning; co-/team teaching; within-class grouping). the factor teaching contains manifest variables concerning instructional methods (feedback; teaching meta-cognitive strategies; problem-based learning) and the teacher (professional development), while the factor achievement emphasizes achievementfocused and teacher-centered variables (direct instruction; teacher expectations; frequent/effects of testing). student-internal prerequisites (motivation; self-concept; prior achievement; attitude) constitute the factor student. table 4 correlation matrix of latent factor model achievement school student teaching .41** .70** .66** achievement -.16** .17* school .78** **p < .01; *p < .05 table 4 shows bivariate correlations of the four latent variables, which were all statistically significant. all coefficients were positive apart from the one between achievement and school, which showed a negative relationship. the intercorrelations were strongest for teaching/school, teaching/student, and school/student. the factor achievement consistently showed the weakest relationships with all the other   fleckenstein  et  al   f   | f l r     38   factors. these results indicate that, in general, teachers had a tendency to believe in either high or low effect sizes for factors of school effectiveness. however, achievement-related variables seemed to be the exception. table 5 measurement invariance across groups (preand in-service teachers) model parameters constrained χ² df cfi tli rmsea srmr 1 none (configural invariance) 405.22 184 .932 .912 .058 .055 2 fl (metric invariance) 428.09 196 .929 .913 .057 .057 3 fl, il (scalar invariance) 627.50 208 .872 .852 .075 .069 3b fl, il (partial scalar invariance) 448.51 205 .925 .913 .058 .058 note. fl = factor loadings. ii = item intercepts. for model identification in model 1 and 2 (item intercepts freely estimated) latent means were fixed to zero. cfi = comparative fit index. tli = tucker-lewis index. rmsea = root mean square error of approximation. srmr = standardized root mean square residual. the model showed partial scalar invariance (strong measurement invariance; see table 5), which indicated that factor structure and factor loadings as well as item intercepts were equal for both groups. pre and in-service teachers attributed the same meaning to the latent constructs and the levels of the underlying items. thus, the proposed factor model can be assumed to represent the belief structure of pre-service as well as in-service teachers. strong measurement invariance allows for the comparison of latent group means. in the latent mean structure analysis the pre-service teachers were chosen as a reference group so that the difference in means between preand in-service teachers on each construct equals the mean of the nonreference group (in-service teachers). the means of the in-service teachers are as shown in table 6. they valued the achievement-related factor considerably higher, while rating the effects of infra-/surface-structure and student-internal variables lower than pre-service teachers. the group differences concerning the factor teacher/teaching were not significant. table 6 mean group differences in latent variables* factor mdifference p teaching -0.12 ns achievement 0.67 <.001 school -0.56 <.001 student -0.70 <.001 *mpre-service = 0; min-service = mdifference 5. discussion 5.1 summary the present paper deals with expert and novice teachers’ beliefs about school effectiveness. we investigated the differences between (a) teachers’ beliefs versus findings of school effectiveness research (cf. hattie, 2009), and (b) expert versus novice teachers’ beliefs. for this purpose preand in-service teachers   fleckenstein  et  al   f   | f l r     39   were asked to rate the effect sizes of several determinants of student achievement. profile correlations were aggregated and compared in terms of similarity to recent empirical findings (hattie, 2009). we found significant differences between preand in-service teachers, as the latter showed a stronger overall congruence with hattie’s results. subsequently, data were combined by using a four-dimensional cfamodel with the latent factors school, teaching, achievement, and student. partial measurement invariance could be established allowing the comparison of latent factor means of the two groups. in-service teachers showed higher means in achievement-focused variables (e.g., direct instruction, and high expectations) and lower means in variables concerning the infraand surface-structural conditions of schooling (e.g., class size, co-/team teaching, within-class grouping, and open learning) as well as student-internal variables (e.g., prior achievement, motivation, and attitude). the structure of teachers’ beliefs concerning school effectiveness seemed to resemble the a priori categorization of relevant research studies. the assumed categorical structure (school, teaching, and student) used by hattie (2009) to organize his metasynthesis was largely met by the data. an exception was the separation of the factor teaching into teaching and achievement. teachers seemed to distinguish two types of instructional preferences: one that foregrounds the support of students’ learning (feedback, teaching metacognitive strategies) and one that focuses mainly on cognitive achievement (high expectations, frequent/effects of testing). while school effectiveness research has shown that both have a strong positive influence on student achievement (see chapter 1.3), teachers undervalued instructional choices that point directly and explicitly at academic achievement. one key finding of this study was that in-service experience and training seem to be associated with teachers’ beliefs about school effectiveness. the beliefs of experienced teachers were more consistent with empirical results than those of novice teachers. this is in line with the current theoretical expert(-novice) paradigm, according to which teachers develop expertise in the course of their education and career. first of all, the ratings of our in-service teachers suggested that they value a kind of activating, teacher-directed instruction, which is supposed to affect the deep structure of classroom learning more than the pre-service teachers do. second, in comparison to the novices they believed infraand surface-structural variables to be not as relevant. this hierarchy is also emphasized by the results of school effectiveness research as outlined in chapter 1.3. this compliance indicates that expert teachers are more competent in assessing the effects of a range of variables than novice teachers. furthermore, it shows that educational research appears not as far from a teacher’s perceived everyday reality as is often suggested. in turn, we must acknowledge that some of the beliefs concerning influences on student achievement – particularly (but not only) those of pre-service teachers – diverge from empirical findings quite dramatically. similarly to weinstein (1989), we found that affective (e.g. motivation, attitude) and social (e.g. within-class grouping) variables were overestimated, while cognitive variables (e.g. direct instruction, prior achievement) were underestimated, especially by the pre-service teachers, but also – to a lesser extent – by in-service teachers. this insight is quite valuable considering the impact of beliefs on classroom instruction. hence, our results call for a paradigmatic change in the way teachers are trained. in the following we consider the limitations of this study before we conclude with its strengths and practical implications for the field of teacher education. 5.2 limitations one problem with the interpretation of our results was the epistemological status of the information given by the participants: did we actually elicit their beliefs or rather their theoretical knowledge about what should be correct. as we mentioned above, the confrontation with objective data, such as the findings of empirical investigations, does not necessarily lead to a change in individual beliefs: knowing something does not equal believing in it. so the response of the teachers to a certain item may not have shown their beliefs but instead, represent the knowledge they have about school effectiveness research. hence, the epistemological distinction between beliefs and knowledge that we addressed above could not be followed   fleckenstein  et  al   f   | f l r     40   through completely. similarly, the issues of tacit knowledge, its accessibility and the reciprocal relationship of implicit and explicit knowledge are highly relevant topics that go beyond the scope of this study. even if we assume that we are dealing with teachers’ beliefs (not knowledge), we cannot be certain that those beliefs are actually put into practice. in our report of prior research on teachers’ beliefs we pointed out that beliefs are considered to affect classroom processes and, in turn, student outcomes. however, the assumption that teachers apply everything they believe to classroom practice would go too far (beliefbehavior gap; cf. sheeran, 2002). so a general identification of “good beliefs” with “good teacher” is too simplistic. this issue requires further research that examines the relevance of teachers’ beliefs (concerning factors of school effectiveness) for instructional choices and student achievement. another drawback is that in the present study hattie’s findings served as a kind of “quasi-reality”. despite all its merits, as a synthesis of meta-analyses it also poses some methodological difficulties (for a more extensive discussion see terhart, 2011), which limits the interpretability of the discrepancy between “real” and teacher-estimated effect sizes. thus, even though we concentrated on those variables that have shown similar results in other school effectiveness studies, the comparison with hattie’s results is to be interpreted with caution. the same holds for the comparison of preand in-service teachers: as our data was not longitudinal, strictly speaking we cannot interpret the discrepancy between the groups as a development or acquisition of competence. additionally, in order to validate and stabilize the latent factor structure further investigations with a larger number of variables from hattie’s study are needed. the generalizability of our results to other countries is limited for two different reasons: firstly, the intercultural generalizability of hattie’s study is already questionable. it mainly relies on research findings from anglophone countries, not all of which are equally relevant for education in germany and for the beliefs of german teachers. secondly, we also have to take into account that our sampling was restricted to the german education system and, moreover, to one specific teacher training program. beliefs are rather likely to differ in terms of structureand content-related aspects of teacher education in different countries and institutions. furthermore, they are subject to the more general cultural and academic situation in a certain time and place. for example, the acceptance and appreciation of empirical research may differ considerably from country to country and from faculty to faculty. in case of this study, the fact that german teacher training programs generally focus on the study of two scientific disciplines rather than on educational science and field experience might impact the beliefs of teachers. the question of differences in teachers’ beliefs according to differences in their (cultureand program-specific) education and practice would be an interesting topic for future research. our research was restricted to teachers’ beliefs concerning cognitive achievement. unfortunately, hattie’s work does not identify determinants of motivation, attitude, self-concept or other affective variables. a study that examines determinants of affective outcomes to the extent that hattie did this for cognitive outcomes is still missing in educational effectiveness research. thus, there is no basis for research on teachers’ beliefs concerning those variables yet. 5.3 strengths and educational implications especially the beliefs of pre-service teachers differed significantly from the results of hattie’s research synthesis. of course neither hattie’s findings nor the beliefs of expert teachers can be taken as ultimately true or as factual reality. however, they both emphasize similar aspects of schooling: the role of the teacher as an activator (rather than a facilitator), the importance of academic achievement and the comparably little significance of structural conditions. if teacher educators address these issues explicitly and confront their students with their own beliefs as well as with the findings of school effectiveness research, they can help (prospective) teachers to focus on what has yet been shown to work best in school. the fact that pre-service teachers’ beliefs diverged from empirical findings more strongly than those of experienced teachers suggests that learning opportunities in the field do make a difference. the experience   fleckenstein  et  al   f   | f l r     41   that teachers gain in their years of classroom practice seems to affect their judgment in a way that is beneficial for their belief systems. one could argue that the constant feedback they get from their students’ performance in terms of (successful and unsuccessful) interventions helps them challenge and adjust their beliefs where necessary. in-service teachers’ focused on instructional strategies and factors that support the deep structure of learning might thus be a reaction to their (more or less systematic) observation and monitoring of what actually works in their own classroom. pre-service teachers, however, are missing this direct feedback in terms of actual student outcomes as their confrontation with actual classroom situations is very limited. holding on to established beliefs might be a result of them not being challenged by reality in the field. moreover, one could expect novice teachers to be quite overwhelmed by their first tentative efforts in teaching. the demands they have to meet in the classroom are manifold and they need to concentrate on various things at once. in such a state, focusing on the surface structure of learning seems easier than focusing on deep-structural aspects. only with substantial practice and experience, when other processes come to them more naturally and intuitively, teachers get the opportunity to pay attention to those instructional details that have actually been shown to work in school. early and regular work experience in school during teacher training might aid the acquisition of the necessary professionalism, as it presents the opportunity to familiarize pre-service teachers with real classroom situations and their role as a teacher. however, this should be realized only with respect to the current state of research as we have little reason to assume that practical school experiences for pre-service teachers automatically lead to better teaching abilities or a better understanding for the purposes and consequences of teaching. (tabachnik et al., 1979-1980). practical experience in school does not automatically make better teachers. this also applies to their beliefs: studies on short-term work experiences during teacher training has shown the resilience of teachers’ beliefs to change (hascher, 2012; richardson, 1996). in order to avoid this misguided process, these early experiences need to be instructed and accompanied by professional teacher trainers. if planned and exercised carefully, practical teaching experience during teacher training at university can lay a solid foundation for a teacher’s career. the other central finding of this study was the fundamental discrepancy of teachers’ beliefs and empirical evidence from school effectiveness research. to some extent, this might be due to shortcomings in teacher training programs to convey to future teachers the importance of evidence-based practice. the australian educational scholar and administrator michele bruniges puts the lack of data usage in the teaching profession into the following words: “a greek philosopher might suggest that evidence is what is observed, rational and logical; a fundamentalist – what you know is true; a post modernist – what you experience; a lawyer – material which tends to prove or disprove the existence of a fact and that is admissible in court; a clinical scientist – information obtained from observations and/or experiments; and a teacher – what they see and hear” (bruniges, 2005; p. 102). while systematic observation and monitoring of students’ learning processes are very desirable actions to be taken by teachers, “seeing and hearing” should not be the only sources for their professional choices and actions. our study supports the claim that teachers rarely rely on available research evidence. their assessment of what actually works in school rather seems to be guided by subjective experiences that are usually gained in the isolation of their own classrooms. but it would certainly be wrong to lay all the blame on the teachers: what we need is an evidencebased culture of improvement in teaching and learning. in order to achieve this goal, three professions in the field of education need to assume responsibility: researchers, teacher trainers, and teachers themselves. first of all, educational researchers are confronted with the issue of making their findings available to teachers. more often than it is already done, they should break rather abstract studies down to what is of practical relevance for the field. such efforts may counteract aversion to empirical research on the side of the teachers. with his follow-up book “visible learning for teachers”, hattie sets a good example for this kind of transfer. secondly, those who educate and train pre-service teachers need to make sure their students are   fleckenstein  et  al   f   | f l r     42   familiar with relevant research findings, can interpret them appropriately, and have the necessary skills to implement them in school. in addition to assessing students’ knowledge, they should be attentive to their beliefs and make room for critical discussion of empirical versus anecdotal evidence. explicitly addressing the issue of teachers’ beliefs and confronting (future) teachers with cognitive dissonance might support a critical reflection and examination of existing beliefs. last but not least, it is a necessity for teachers themselves to stay in touch with research communities in order to understand current developments and to constantly reflect on their beliefs in comparison with crucial evidence provided by researchers. keypoints teachers’ beliefs diverge from empirical evidence expert teachers’ beliefs diverge from novice teachers’ beliefs expert teachers show more congruence with empirical evidence than novice teachers expert teachers believe in the effectiveness of achievement-related variables novice teachers believe in the effectiveness of structural and student factors references allen, l. q. (2002). teachers’ pedagogical beliefs and the standards for foreign language learning. foreign language annals, 35, 518-529. http://dx.doi.org/10.1111/j.1944-9720.2002.tb02720.x bandura, a. (1997). self-efficacy: the exercise of control. new york: freeman. barcelos, a. m. f. (2003). researching beliefs about sla: a critical review. in p. kalaja and a. m. f. barcelos (eds.), beliefs about sla: new research approaches (pp. 7-33). dordrecht: kluwer academic publishers. http://dx.doi.org/10.1007/978-1-4020-4751-0_1 bates, a. b., latham, n. & kim, j. (2011). linking preservice teachers' mathematics self-efficacy and mathematics teaching efficacy to their mathematical performance. school science and mathematics. 111, 325-333. http://dx.doi.org/10.1111/j.1949-8594.2011.00095.x bauer, j. & prenzel, m. (2012). european teacher training reforms. science, 336, 1642-1643. http://dx.doi.org/10.1126/science.1218387 baumert, j., & kunter, m. (2006). stichwort: professionelle kompetenz von lehrkräften. zeitschrift für erziehungswissenschaft, 9, 469-520. http://dx.doi.org/10.1007/s11618-006-0165-2 berliner, d. (2004). describing the behavior and documenting the accomplishments of expert teachers. bulletin of science, technology and society, 24, 200-214. http://dx.doi.org/10.1177/0270467604265535 borg, m. (2001). teachers’ beliefs. elt journal, 55, 186-188. http://dx.doi.org/10.1093/elt/55.2.186 borg, s. (2003). teacher cognition in language teaching: a review of research on what language teachers think, know, believe, and do. language teaching, 36, 81-109. http://dx.doi.org/10.1017/s0261444803001903 bromme, r. (1997). kompetenzen, funktionen und unterrichtliches handeln des lehrers. in f. e. weinert (ed.), psychologie des unterrichts und der schule. enzyklopaedie der psychologie (vol. 3, pp. 177212). goettingen: hogrefe. bromme, r. (2003). on the limitations of the theory metaphor for the study of teachers' expert knowledge. in m. kompf & p. denicolo (eds.). teacher thinking twenty years on: revisiting persisting problems and advances in education (pp. 283-294). liss, nl: swets & zeitlinger. bruniges, m. (2005). an evidence-based approach to teaching and learning. http://research.acer.edu.au/research_conference_2005/15.   fleckenstein  et  al   f   | f l r     43   byrne, b. m., shavelson, r. j., & muthén, b. (1989). testing for the equivalence of factor covariance and mean structures: the issue of partial measurement invariance. psychological bulletin, 105, 456-466. http://dx.doi.org/10.1037/0033-2909.105.3.456 calderhead, j. (1996). teachers: beliefs and knowledge. in d. berliner & r. calfee (eds.), handbook of educational psychology (pp. 709-725). new york: macmillan. cheung, g. w., & rensvold, r. b. (2002). evaluating goodness-of-fit indexes for testing mi. structural equation modeling, 9, 235-55. http://dx.doi.org/10.1207/s15328007sem0902_5 clark, c. m. (1988). asking the right questions about teacher preparation: contributions of research on teacher thinking. educational researcher, 17, 5-12. http://dx.doi.org/10.3102/0013189x017002005 cochran-smith, m., & zeichner, k. (eds.). (2005). studying teacher education: the report of the aera panel on research and teacher education. mahwah: lawrence erlbaum associates. coleman, j. s., campbell, e. q., hobson, c. f., mcpartland, j., mood, a. m., weifeld, f. d., et al. (1966). equality of educational opportunity. washington, d.c.: u.s. government printing office. cotton, k. (1995). effective schooling practices: a research synthesis. 1995 update. school improvement research series. northwest regional educational laboratory. darling-hammond, l., & bransford, j. (eds.). (2005). preparing teachers for a changing world. what teachers should learn and be able to do. san francisco: jossey-bass. ericsson, k. a. (2006). the influence of experience and deliberate practice on the development of superior expert performance. in k. a. ericsson, n. charness, p. feltovich, and r. r. hoffman, r. r. (eds.). cambridge handbook of expertise and expert performance (pp. 685-706). cambridge, uk: cambridge university press. ericsson, k. a., & lehmann, a. c. (1996). expert and exceptional performance: evidence of maximal adaptations to task constraints. annual review of psychology, 47, 273-305. http://dx.doi.org/10.1146/annurev.psych.47.1.273 ericsson, k. a., charness, n., feltovich, p. j., & hoffman, r. r. (2006). the cambridge handbook of expertise and expert performance. cambridge: cambridge university press. http://dx.doi.org/10.1017/cbo9780511816796 fang, z. (1996). a review of research on teacher beliefs and practices. educational research, 38, 47-65. http://dx.doi.org/10.1080/0013188960380104 fraser, b. j., walberg, h. j., welch, w. w., & hattie, j. a. (1987). syntheses of educational productivity research. international journal of educational research, 11, 147-252. http://dx.doi.org/10.1016/08830355(87)90035-8 freeman, d. (2002). the hidden side of the work: teacher knowledge and learning to teach. language teaching, 35, 1-13. http://dx.doi.org/10.1017/s0261444801001720 haney, j. j., lumpe, a. t., & czerniak, c. m. (2003). constructivist beliefs about the science classroom learning environment: perspectives from teachers, administrators, parents, community members, and students. school science and mathematics, 103, 366-377. http://dx.doi.org/10.1111/j.19498594.2003.tb18122.x hart, l. c. (2004). beliefs and perspectives of first-year, alternative preparation, elementary teachers in urban classrooms. school science & mathematics, 104, 79-88. http://dx.doi.org/10.1111/j.19498594.2004.tb17985.x harvey, o. j. (1986). belief systems and attitudes toward the death penalty and other punishments. journal of psychology, 54, 143-159. http://dx.doi.org/10.1111/j.1467-6494.1986.tb00418.x hascher, t. (2012). lernfeld praktikum – evidenzbasierte entwicklungen in der lehrer/innenbildung. [learning setting student teaching – evidence-based developments in teacher education]. zeitschrift für bildungsforschung, 2, http://dx.doi.org/10.1007/s35834-012-0032-6 hattie, j. (2003). teachers make a difference: what is the research evidence? paper presented at the australian council for educational research annual conference on building teacher quality. hattie, j. (2009). visible learning. a synthesis of over 800 meta-analyses relating to achievement. london & new york: routledge. hattie, j. (2012). visible learning for teachers. london: routledge.   fleckenstein  et  al   f   | f l r     44   hogan, t., rabinowitz, m., & craven, j. a. (2003). representation in teaching: inferences from research of expert and novice teachers. educational psychologist, 38, 235-247. http://dx.doi.org/10.1207/s15326985ep3804_3 jencks, c., smith, m. s., ackland, h., bane, m. j., cohen, d., grintlis, h., et al. (1973). inequality: a reassessment of the effect of family and schooling in america. new york: basic books. köller, o. (2012). what works best in school? hatties befunde zu effekten von schulund unterrichtsvariablen auf schulleistungen. [what works best in school? hattie’s findings concerning school and instructional effectiveness on student achievement]. psychologie in erziehung und unterricht, 59, 72-78. http://dx.doi.org/10.2378/peu2012.art06d kunter, m., & pohlmann, b. (2009). lehrer. in j. möller & e. wild (eds.), einführung in die pädagogische psychologie (pp. 261-282). berlin: springer. http://dx.doi.org/10.1007/978-3-540-88573-3_11 leatham, k. (2006). viewing mathematics teachers’ beliefs as sensible systems. journal of mathematics teacher education, 9, 91-102. http://dx.doi.org/10.1007/s10857-006-9006-8 levine, d.k., & lezotte, l.w. (1990). unusually effective schools: a review and analysis of research and practice. madison, wise: nat. centre for effective schools research and development. lortie, d. (1975). schoolteacher: a sociological study. chicago: university of chicago press. meredith, w., & teresi, j. a. (2006). an essay on measurement and factorial invariance. medical care, 44, 69-77. http://dx.doi.org/10.1097/01.mlr.0000245438.73837.89 muijs, r. d., & rejnolds, d. (2001). teachers' beliefs and behaviors: what really matters. journal of classroom interaction, 37, 3-15. muthén, l.k. and muthén, b.o. (1998-2012). mplus user’s guide. seventh edition. los angeles, ca: muthén & muthén. national board for professional teaching standards (2002). what teachers should know and be able to do. arlington. nespor, j. (1987). the role of beliefs in the practice of teaching. journal of curriculum studies, 19, 317-328. http://dx.doi.org/10.1080/0022027870190403 pajares, m. f. (1992). teachers' beliefs and educational research: cleaning up a messy construct. review of educational research, 62, 307-332. http://dx.doi.org/10.3102/00346543062003307 peterson, p., fennema, e., carpenter, t. p., & loef, m. (1989). teachers' pedagogical content beliefs in mathematics. cognition and instruction, 6, 1-40. http://dx.doi.org/10.1207/s1532690xci0601_1 purkey, s.c., & smith, m.s. (1983). effective schools: a review. the elementary school journal, 83, 427452. http://dx.doi.org/10.1086/461325 richardson, v. (1996). the role of attitudes and beliefs in learning to teach. in j. sikula, t. buttery, & e. guyton (eds.), handbook of research on teacher education (pp. 137-147). new york: macmillan. ross, j. a. (1992). teacher efficacy and the effect of coaching on student achievement. canadian journal of education, 17, 51-65. http://dx.doi.org/10.2307/1495395 ross, j. a. (1998). the antecedents and consequences of teacher efficacy. in j. brophy (ed.), advances in research on teaching, vol. 7 (pp. 49-74). greenwich, ct: jai press. sammons, p., hillman, j., & mortimore, p. (1995). key characteristics of effective schools: a review of school effectiveness research. london: ofsted. scheerens, j. (1992). effective schooling, research, theory and practice. london: cassell. scheerens, j. (2004). review of school and instructional effectiveness research. paper commissioned for the efa global monitoring report 2005, the quality imperative. unesco, 2005/ed/efa/mrt/pi/44. scheerens, j., & bosker, r. j. (1997). the foundations of educational effectiveness. oxford: elsevier science ltd. seidel, t., & shavelson, r. j. (2007). teaching effectiveness research in the past decade: the role of theory and research design in disentangling meta-analysis results. review of educational research, 77, 454499. http://dx.doi.org/10.3102/0034654307310317 sheeran, p. (2002). intention-behavior relations: a conceptual and empirical review. in w. strobe and m. hewstone (eds.) european review of social psychology, vol. 12. chichester: wiley, 1-30. http://dx.doi.org/10.1080/14792772143000003   fleckenstein  et  al   f   | f l r     45   shulman, l. s. (1986). those who understand: knowledge growth in teaching. educational researcher, 15, 4-14. http://dx.doi.org/10.3102/0013189x015002004 shulman, l. s. (1987). knowledge and teaching: foundations of the new reform. harvard educational review, 57, 1-22. silver, n. c., & dunlap, w. p. (1987). averaging correlation coefficients: should fisher’s z transformation be used? journal of applied psychology, 72, 146-148. http://dx.doi.org/10.1037/0021-9010.72.1.146 staub, f. c., & stern, e. (2002). the nature of teachers’ pedagogical content beliefs matters for students’ achievement gains: quasi-experimental evidence. journal of educational psychology, 94, 344-355. http://dx.doi.org/10.1037/0022-0663.94.2.344 stuart, c., & thurlow, d. (2000). making it their own: pre-service teachers’ experiences, beliefs, and classroom practices, journal of teacher education, 51(2), 112-121. http://dx.doi.org/10.1177/002248710005100205 tabachnik, b. r., popkewitz, t., & zeichner, k. (1979-1980). teacher education and the professional perspectives of student teachers. interchange, 80, 12-29. http://dx.doi.org/10.1007/bf01810816 terhart, e. (2011). has john hattie really found the holy grail of research on teaching? an extended review of visible learning. journal of curriculum studies, 43(3), 425-438. http://dx.doi.org/10.1080/00220272.2011.576774 tschannen-moran, m., & woolfolk hoy, a. (2001). teacher efficacy: capturing an elusive concept. teaching and teacher education, 17, 783-805. http://dx.doi.org/10.3102/00346543068002202 tschannen-moran, m., woolfolk hoy, a., & hoy, w. k. (1998). teacher efficacy: its meaning and measure. review of educational research, 68, 202-248. verloop, n., van driel, j., & meijer, p. (2001). teacher knowledge and the knowledge base of teaching. international journal of education research, 35, 441-461. http://dx.doi.org/10.1016/s08830355(02)00003-4 walberg, h. j. (1986). synthesis of research on teaching. in m. c. wittrock (ed.), handbook of research on teaching. new york: macmillan. wang, m. c., haertel, g. d. & walberg, h. j. (1990). what influences learning? a content analysis of review literature. journal of educational research, 84, 30-43. http://dx.doi.org/10.1080/00220671.1990.10885988 wang, m. c., haertel, g. d. & walberg, h. j. (1993). toward a knowledge base for school learning. review of educational research, 63, 249-294. http://dx.doi.org/10.3102/00346543063003249 weinstein, c. s. (1989). teacher education students' preconceptions of teaching. journal of teacher education, 40, 53-60. http://dx.doi.org/10.1177/002248718904000210 wenden, a. l. (1999). an introduction to metacognitive knowledge and beliefs in language learning: beyond the basics. system, 27, 435-441. http://dx.doi.org/10.1016/s0346-251x(99)00043-3 woods, d. (1996). teacher cognition in language teaching: beliefs, decision-making, and classroom practice. cambridge: cambridge university press. woolfolk, a. e., & hoy,w. k. (1990). prospective teachers' sense of efficacy and beliefs about control. journal of educational psychology, 82, 81-91. http://dx.doi.org/10.1037/0022-0663.82.1.81 woolfolk, a. e., rosoff, b., & hoy, w. k. (1990). teachers' sense of efficacy and their beliefs about managing students. teaching and teacher education, 6, 137-148. woolfolk hoy, a., & davis, h. a. (2006). teacher self-efficacy and its influence on the achievement of adolescents. in f. pajares & t. urdan (eds.), self-efficacy of adolescents (pp. 117-137). greenwich, connecticut: information age publishing. http://dx.doi.org/10.1016/0742-051x(90)90031-y woolfolk hoy, a., davis, h., & pape, s. j. (2006). teacher knowledge and beliefs. in p. a. alexander & p. h. winne (eds.), handbook of educational psychology (2 ed., pp. 715-737). mahwah, nj: erlbaum. yook, c. m. (2010). korean teachers' beliefs about english language education and their impacts upon the ministry of education-initiated reforms. applied linguistics and english as a second language dissertations. paper 14. zell, e., & krizan, z. (2014). do people have insight into their abilities? a metasynthesis. perspectives on psychological science, 9, 111-125. http://dx.doi.org/10.1177/1745691613518075   fleckenstein  et  al   f   | f l r     46   zheng, h. (2009). a review of research on pre-service teachers’ beliefs and practices. journal of cambridge studies, 4, 73-81. frontline learning research 2 (2013) 99-101 issn 2295-3159 corresponding author: michael schneider, university of trier, www.educational-psychology.uni-trier.de, m.schneider@uni-trier.de, and peter edelsbrunner, eth zurich, www.ifvll.ethz.ch, peter.edelsburnner@ifv.gess.ethz.ch http://dx.doi.org/10.14786/flr.v1i2.74 99 | f l r modelling for prediction vs. modelling for understanding: commentary on musso et al. (2013) peter edelsbrunner a , michael schneider b a eth zurich, switzerland b university of trier, germany article received 11 september 2013 / accepted 12 december 2013 / available online 20 december 2013 abstract musso et al. (2013) predict students’ academic achievement with high accuracy one year in advance from cognitive and demographic variables, using artificial neural networks (anns). they conclude that anns have high potential for theoretical and practical improvements in learning sciences. anns are powerful statistical modelling tools but they can mainly be used for exploratory modelling. moreover, the output generated from anns cannot be fully translated into a meaningful set of rules because they store information about input-output relations in a complex, distributed, and implicit way. these problems hamper systematic theory-building as well as communication and justification of model predictions in practical contexts. modern-day regression techniques, including (bayesian) structural equation models, have advantages similar to those of anns but without the drawbacks. they are able to handle numerous variables, non-linear effects, multi-way interactions, and incomplete data. thus, researchers in the learning sciences should prefer more theory-driven and parsimonious modelling techniques over anns whenever possible. keywords: artificial neural networks; black box; student achievement; statistical modelling musso, kyndt, cascallar, and dochy (2013) conducted a study in which the statistical modelling technique of artificial neural networks (anns) was used to predict the academic achievement of university students a year in advance. the measures used were attention, working memory, learning strategies, and demographic variables. the results were precise estimations of each student’s achievement tercile after their first year at university. this is an impressive success, demonstrating the usefulness of anns as a statistical modelling tool. p. edelsbrunner and m. schneider 100 | f l r the study has raised an important question of the preferred statistical methods used by researchers in learning sciences. should anns replace conventional statistical methods such as multiple regression, discriminant analysis, and structural equation modelling? – the potential of anns cannot be denied especially as a tool to examine predictive patterns in complex systems. however, musso and colleagues overestimate the ability of anns in their application to the learning sciences. they do not mention shortcomings of anns, while overemphasizing shortcomings of competing conventional methods. anns are limited in at least two important ways. first, the construction of ann models such as those used by musso et al. is highly explorative apart from choosing relevant input and output variables (günther, pigeot, &bammann, 2012; scarborough & somers, 2006). the connection weights, which determine how an ann transforms input into output patterns, are not specified by the researchers or based on theory. they are set to random values and changed gradually by an optimisation algorithm. this process usually involves thousands of iterations until each input pattern leads to the desired output pattern in the training data set. anns, thus, cannot be entirely compared to conventional methods since the latter are aimed at confirming or disconfirming pre-specified relations and interactions. in other words, the research question should determine whether the exploratory nature of anns is adequate, or if a conventional, confirmatory model should be the method of choice. second, connection weights cannot be codified into a coherent set of rules that delineate the process by which anns transform input patterns into output patterns. anns typically have a high number of connections between neurons (e.g., 300 in ann1 by musso et al.). the transformation process of input into output patterns is determined by non-linear, multi-way interactions of these connection weights. recent research has attempted to increase the interpretability of anns, for example with the help of visualizations for complex interactions (e.g., cortez & embrechts, 2013; intrator & intrator, 2001). however, the basic problem of how non-linear interactions between hundreds of variables can be understood and communicated in meaningful terms has not yet been solved, causing anns to be frequently characterised as “black boxes” (cf. benitez, castro, & requena, 1997). while one can assess how well an ann works, it is difficult to comprehensively explain why it performs well or not (scarborough & somers, 2006). to interpret their results, musso and colleagues list an importance parameter for each predictor but these parameters do not explain interaction effects or non-linear relations among the variables. in addition, it is difficult to integrate the results of anns across studies and also generalise from samples to underlying populations due to the lack of output parameters such as standard errors and error probabilities. the explorative and opaque nature of anns impedes theory-developing and limits their practical application. each relation in a statistical model should ideally correspond to a matching relation in an educational or psychological theory that justifies and explains the assumed statistical relation. researchers can compare competing theories and advance assumptions that are not in line with the empirical data by fitting a series of statistical models that differ in theoretically relevant aspects (kaplan, 1990). this is not possible with anns because the input-output relations are implicitly coded and distributed over all connection weights, preventing researchers from being able to map elements of an ann and elements of a theory onto each other (luger, 2009, p. 680). the results obtained from ann models are also of limited use for solving real-life problems. this limitation can be illustrated in a situation where diagnosticians would have to tell certain high school students that despite achieving satisfactory levels in their current academic performances, they cannot be admitted to college because an ann predicts low academic performance in the future. in justifying the results, the diagnosticians would have to admit that they cannot explain how the different predictors statistically combine, nor describe the causal processes that will contribute to the anticipated decrease in the students’ achievement. these limitations are unsatisfactory from diagnostic, educational, and public policymaking perspectives. conventional methods represent more parsimonious and theory-driven alternatives to anns because they use smaller numbers of parameters, which enhances the interpretability of results. like anns, modern regression techniques can account for non-linear relations (bates & watts, 2007) and complex interactions between variables (aiken & west, 1991). structural equation models are built on regression techniques and p. edelsbrunner and m. schneider 101 | f l r allow a simultaneous analysis of numerous variables. these models can be estimated by methods that are robust to missing data and non-normal distributions, account for hierarchical data structures, and identify heterogeneous sub-populations in mixture-models (hoyle, 2012). especially bayesian structural equation models represent a strong advancement in modelling non-linear relations, assessing unspecified relations and handling highly non-normal and hierarchical data (song & lee, 2012). in contrast to anns, these modelling techniques require explicit theoretical assumptions about the relationship of the variables and they allow for explicit tests of these assumptions. this might limit their predictive power compared to anns, but it aids theory-building, hypothesis testing, and the communication of model results in practical applications. keypoints artificial neural networks are powerful statistical tools for pattern recognition and prediction. artificial neural networks transform input patterns into output patterns by non-linear multi-way interactions between simulated neurons that are governed by information that is stored in connection weights in an implicit and distributed way. this “black box” nature of artificial neural networks hampers the systematic testing of theories and the communication of results in practical settings. more conventional regression-type models can also handle non-linear relations, interaction effects, and a high number of variables, correlated errors, missing values, and non-normal distributions. artificial neural network analysis cannot replace conventional statistical methods in the learning sciences but may be applicable in specific cases. references aiken, l. s., & west, s. g. (1991). multiple regression: testing and interpreting interactions. newbury park, ca: sage. bates, d. m., & watts, d. g. (2007). nonlinear regression analysis and its applications (2nd ed.). hoboken, nj: wiley. benitez, j. m., castro, j. l., & requena, i. (1997). are artificial neural networks black boxes? ieee transactions on neural networks, 8, 1156-1164. doi:10.1109/72.623216 cortez, p., & embrechts, m. j. (2013). using sensitivity analysis and visualization techniques to open black box data mining models. information sciences, 225, 1-17. doi:http://dx.doi.org/10.1016/j.ins.2012.10.039 günther, f., pigeot, i., & bammann, k. (2012). artificial neural networks modeling gene-environment interaction. bmc genetics, 13(1), 37. doi:10.1186/1471-2156-13-37 hoyle, r. h. (ed.). (2012). handbook of structural equation modeling. new york: guilford press. intrator, o., & intrator, n. (2001). interpreting neural-network results: a simulation study. computational statistics & data analysis, 37, 373-393. doi:10.1016/s0167-9473(01)00016-0 kaplan, d. (1990). evaluating and modifying covariance structure models: a review and recommendation. multivariate behavioral research, 25, 137-155. doi:10.1207/s15327906mbr2502_1 luger, g. f. (2009). artificial intelligence: structures and strategies for complex problem solving (6th ed.). boston, ma: pearson education. musso, m. f., kyndt, e., cascallar, e. c., & dochy, f. (2013). predicting general academic performance and identifying the differential contribution of participating variables using artificial neural networks. frontline learning research, 1, 42-71. retrieved from http://journals.sfu.ca/flr/index.php/journal/article/view/13 p. edelsbrunner and m. schneider 102 | f l r scarborough, d., & somers, m. j. (2006). neural networks in organizational research: applying pattern recognition to the analysis of organizational behavior (pp. 137-144). washington, dc: american psychological association. song, x. y., & lee, s. y. (2012). basic and advanced bayesian structural equation modeling: with applications in the medical and behavioral sciences. chichester, uk: john wiley & sons. microsoft word dörrenbächer et al_publication.docx ! ! ! ! ! frontline)learning)research)vol.3)no.)4)(2015))14<)36) issn)2295<3159)) ! volition completes the puzzle: development and evaluation of an integrative trait model of self-regulated learning laura dörrenbächer1, franziska perels department of educational sciences, saarland university article received 20 may / revised 25 october / accepted 26 october / available online 15 december abstract most self-regulated learning theories are imbedded within a social-cognitive framework and comprise cognitive, metacognitive and motivational components. nevertheless, these theories partly neglect volition, which is necessary for implementing learning intentions. therefore, the present study is frontline as it aimed to integrate volition within a comprehensive trait model of self-regulated learning (srl) while proposing a new conception of trait volition for learning. a sample of n = 377 college students (70.1% female, mage = 23.36, sdage = 4.12) filled out questionnaires concerning volitional, cognitive, metacognitive, and motivational belief aspects of srl. the results of confirmatory factor analysis speak in favour of integrating the highly interrelated constructs of procrastination, future time perspective, and academic delay of gratification in order to depict volition for srl. moreover, the structural equation modelling results favour a twofold motivational component for srl that comprises both motivational beliefs and volition instead of including volition as a separate component aside from cognitive, metacognitive and motivational belief components. additionally, the comprehensive trait model of srl is related to gpa, which is a first indication of its validity. therefore, the study empirically investigates a new conception of trait volition for learning environments as well as its integration within a comprehensive srl framework. future research should consider the importance of volitional components for srl and could investigate individual differences concerning the modelled components. keywords: self-regulated learning; volition; academic delay of gratification; procrastination; future time perspective !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 1 corresponding author: laura dörrenbächer, department of educational sciences, saarland university, campus building a 4 2, 66123 saarbrücken, germany. phone: +49(0)681 / 30258337, fax: +49(0)681/ 30258341, email: laura.doerrenbaecher@uni-saarland.de doi: http://dx.doi.org/10.14786/flr.v3i4.179 ! dörrenbächer*&*perels* * * | f l r ! ! 15! 1. introduction although there are many different models to explain self-regulated learning (hereinafter referred to as srl), they all characterize the learner as an individual that is self-determined and that actively creates his or her learning process (efklides, 2011). a self-regulated learner therefore has the ability to set goals and to accomplish these goals by monitoring, controlling, and altering his or her behaviour, motivation and cognition adaptively as a response to changing environmental factors (pintrich, 2000; zimmerman, 2000). consistently, most authors agree that srl embraces cognitive, metacognitive as well as motivational components that interact reciprocally (boekaerts, 1999). recent research has shown that srl positively influences academic outcomes in different areas of education (e.g. dignath, büttner, & langfeldt, 2008; kitsantas, winsler, & huie, 2008) and that college students show better performance when self-regulative strategies are used (nandagopal & ericsson, 2012). concerning its conceptual status, srl can be seen as a trait that influences individual learning processes on a general level (e.g. boekaerts, 1999) or as a dynamic state that changes cyclically according to situational demands (e.g. schmitz & wiese, 2006). recently, there have also been models proposed that integrate both conceptual dimensions (e.g. efklides, 2011) because srl can be regarded as an aptitude and an event (winne & perry, 2000). matthews, schwean, campbell, saklofske, and mohamed (2000) concordantly argue that srl has nomothetic (trait) as well as idiographic (state) qualities. as states are influenced by corresponding traits (hong, 1995) and traits help to explain individual differences (hong & o’neil, 2001), the present study aims to develop and evaluate an integrative trait model of srl that could be useful for future research. the trait model of hong & o’neil (2001) is used as a basis as it already has been tested empirically. we investigate an extended model that also includes a cognitive component (boekaerts, 1999) and considers several subcomponents of metacognition and motivation that are important for depicting srl comprehensively. even though the definition of self-regulated students as “metacognitively, motivationally, and behaviourally active participants in their own learning process” (zimmerman, 2008, p. 167) takes into account volitional aspects, srl research largely has underemphasized such abilities that can predict academic achievement as well (duckworth, gendler, & gross, 2014). volition is the capability to inhibit distracting behaviours in order to attain a higher goal (duckworth & seligman, 2006) and helps to protect learning intentions from action tendencies competing with that goal (corno, 2001). as it has been mostly described within action-control theory (heckhausen & kuhl, 1985), several authors demand for adding volitional aspects above and beyond motivational, cognitive, and metacognitive components within a socialcognitive framework when modelling srl (wolters & benzon, 2013; zimmerman, 2011). therefore, the present study examines the conceptual structure of volition comparing two integrative trait models of srl (see figure 1): one model treats volition as a separate component of srl besides cognitive, metacognitive and motivational belief components (corno, 2001), while the second model categorizes it as a motivational subcomponent besides motivational beliefs2 (zimmerman, 2008). in order to validate the models, their relation to gpa of university entrance diploma will be analysed using structural equation modelling. in the context of integrating volition into the srl framework, a new conceptualization of trait volition is presented: we chose procrastination, future time perspective, and academic delay of gratification as these constructs represent three volitional traits that are highly interrelated (e.g. bembenutty & karabenick, 2004; sirois, 2014) and that act as important supporters for srl (e.g. park & sperling, 2012; zimmerman, 2011). altogether, the present study adds to research because it evaluates an extended trait model of srl that takes into account cognitive and metacognitive components as well as motivational beliefs and volition while the position of volition within the srl framework is examined. in this context, a new conceptualization of volition for learning environments is developed and tested. the present study is frontline as it brings together two highly relevant theoretical frameworks for the field of educational psychology whose relation has been neglected for a long time. after presenting the basis for our !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 2 the term motivation will be used to refer to motivational components in general, while motivation beliefs refer to selfbeliefs concerning self-efficacy, goal orientation or intrinsic task value and are distinguished from volition as a second motivational component. dörrenbächer*&*perels* * * | f l r ! ! 16! comprehensive srl trait model, we will describe the research lines of srl and volition in order to point out differences between them. in a next step, we present attempts to integrate both frameworks and depict the conceptualization of trait volition for srl. 1.1 trait-conception of srl traits represent relatively stable characteristics that influence and predict performance across a wide range of tasks (hertzog & nesselroade, 2003). thus, trait srl can be described as a general disposition of students and learners in general (boekaerts & corno, 2005) or as relatively stable tendencies to use srl strategies. individuals therefore will respond relatively consistently to a range of different learning situations: learners with high srl trait values should show more metacognitive skills and be more self-efficacious than should individuals with low values on this trait (hong & o’neil, 2001). the fact that srl is related to other personality traits and achievement motives also speaks in favour of the trait perspective (wolters & hussain, 2014). although the concept of state srl has recently gained importance in literature and stimulated a great deal of research (azevedo, 2014), the examination of trait srl is necessary to understand individual differences in srl and achievement. accordingly, several studies have shown the positive influence of srl on academic outcomes in all areas of education (e.g. dignath, büttner, & langfeldt, 2008; kitsantas, winsler, & huie, 2008) and therefore constitute its relevance and its meaning for lifelong learning (bronson, 2000). in line with this, college students show better test performance when self-regulative strategies are used during test preparation and accomplishment (kitsantas, 2002), and their grade point average (gpa) level differs depending on the use of self-regulative learning strategies (nandagopal & ericsson, 2012). moreover, the examination of srl as a trait can help to explain differential effects of intervention programs (hong & o’neil, 2001). therefore, it can have practical implications for dealing with heterogeneity in learning and in fostering learning competences. the trait srl model of hong & o’neil (2001) can be used as a starting point for the development of an integrative model. they conceptualize srl as a third-order factor that subsumes the two second-order factors metacognition and motivation, which are highly relevant components of srl and which most models have in common (efklides, 2011). the metacognitive factor comprises the subcomponents of planning one’s time and strategy use as well as the construct of self-checking, which is a method to control proceedings and adapt learning behaviour in a goal-oriented way. the conceptualization of the motivational factor, which is represented by effort as well as self-efficacy, indicates a combination of two motivational concepts: whereas self-efficacy refers to more or less unconscious beliefs or attitudes about one’s skills and competences (bandura, 1997), effort comprises processes that include deliberate thoughts or behaviours in order to reach a goal (carver & scheier, 2000). therefore, self-efficacy as a motivational belief mostly influences goalsetting processes, whereas effort is important for the initiation and implementation of an action (corno, 2001). effort therefore can be considered as a volitional component of srl, helping students to focus attention and to deal with distractions of personal and environmental origin (zimmerman, 2011). even though volition can have the role of a mediator between learning intentions and the actual use of learning strategies, models of srl mostly have overlooked volitional components (garcia, mccann, turner, & roska, 1998). although the previously described model (hong & o’neil, 2001) integrates various important components of srl, it has several points of criticism: at first, the motivational belief component is only represented by self-efficacy, although goal orientations, intrinsic motivation, or causal attributions are also crucial motivational properties (zimmerman, 2008). moreover, it simplifies the volitional component by reducing it to effort and therefore neglects other important volitional factors. additionally, the model lacks the integration of a second-order cognitive factor or the use of cognitive learning strategies that represent a further important srl component (e.g. organization, critical thinking; pintrich, 2000). as the metacognitive factor comprises the construct of self-checking, it blends the capabilities of self-recording as an dörrenbächer*&*perels* * * | f l r ! ! 17! observational method and self-evaluation as a judgment of one’s own actions although these are located in different phases of a self-regulated learning cycle (zimmerman, 2000). concluding, hong and o’neil’s model (2001) is a first attempt to integrate several important components of srl and to test this structure empirically but has several shortcomings. motivated by the abovementioned points of criticism, the present study adds to research as a cognitive factor is included, the metacognitive and motivational belief factors are extended by adding several subcomponents and the structural position of the volition factor is examined in more detail. therefore, the next sections aim to theoretically integrate trait volition within a broader framework of srl considering the previously described shortcomings and to propose an extended concept of trait volition with regard to srl. as far as we know, this study is the first to bring together these research lines and to test such an integrative srl trait model empirically. 1.2 integration of volition into srl 1.2.1 linking two frameworks self-regulated learning is described as the ability to set goals that are accomplished by monitoring, controlling, and altering one’s behaviour, motivation and cognition in response to environmental conditions that are continuously changing (zimmerman, 2000). this interaction of personal, behavioural and environmental processes reflects a social-cognitive perspective for describing learning processes (bandura, 1986). feedback loops between personal and environmental factors are assumed and represent their interdependence. accordingly, zimmerman’s process model of self-regulated learning (2000) distinguishes between forethought, performance and self-reflection phases that comprise these interacting factors. the learner adapts his or her thoughts, affects and behaviour cyclically in order to attain a previous set goal. the social-cognitive framework therefore underlines the role of the learner’s context and situation when goals are set as well as of previous performance when expectations are formed (zimmerman & schunk, 2001). volition is defined as the capability to inhibit irrelevant behaviours in order to attain a higher goal (duckworth & seligman, 2006). concerning learning environments, volition helps to protect learning intentions in the presence of competing action tendencies or obstacles (corno, 2001) by the use of action control strategies (kuhl, 2000). therefore, volition is of importance when students have to maintain their concentration and effort in the presence of internal or external distractions because it supports tenacity (zimmerman, 2011). the construct has mostly been described within action-control theory (kuhl, 1984), where it is seen as a mediator between the intention to learn and the actual use of learning strategies (corno, 1993). concerning this framework, goal-directed activities can be divided into two distinct phases (heckhausen & kuhl, 1985): the pre-decisional phase comprises motivational processes that entail the choice of a specific goal as well as the appraisal of the goal’s value and intention formation. the subsequent post-decisional phase involves the implementation of goal-directed behaviours by the use of volitional strategies that help to maintain one’s intentions (corno, 2001; garcia et al., 1998). in this context, the framework describes how motivation and volition are interwoven and hints at the assumption that volition is a part of a broader motivational concept: choice motivation influences goal setting and expectancy-value processes in the pre-decisional phase and therefore represents motivational beliefs like self-efficacy, task value and goal orientations (husman, mccann & crowson, 2000). executive motivation however takes place in the post-decisional phase and affects action implementation and effort maintenance. according to gollwitzer (1996), choice motivation is characterized by a motivational mindset, whereas a volitional mindset comes along with executive motivation. as most srl theories are imbedded within a social-cognitive framework (zimmerman, 2000), motivational beliefs within the predecisional phase (e.g. self-efficacy, goal orientation) have been of particular interest for this line of research. concerning the postdecisional phase of learning processes, srl models focus on the application of cognitive learning strategies and therefore partly neglect volitional strategies necessary to implement one’s intentions (garcia et al., 1998). as garcia et al. (1998) put it, executive motivation has mostly been neglected within srl frameworks and the majority of srl models dörrenbächer*&*perels* * * | f l r ! ! 18! propose that processes of choice motivation “complete the puzzle” (p. 396). hence, there is a lack of studies that investigate the relation of volition to srl. accordingly, several authors speak in favour of integrating volition within models of srl (e.g. duckworth et al.; 2014, zimmerman, 2008) in addition to aspects of cognition, metacognition, and motivation. concordantly, corno (2001) argues that volition complements motivation and that both concepts taken together represent an action disposition. nevertheless, action-control theory assumes a relatively absolute distinction between both concepts as choice motivation is terminated when the pre-actional phase and volitional processes start (heckhausen & kuhl, 1985). a social-cognitive view would be more dynamic, assuming an interaction between choice and executive motivational processes (wolters, 2003b) and would question if volition is separable from traditional motivational measures like goal expectations (zimmerman & schunk, 2001). therefore, a cross-fertilization of the conceptions could be beneficial for both social-cognitive and action-control theory (duckworth et al., 2014). the abovementioned twofold conception of motivation (self-efficacy and effort) in hong and o’neil’s model (2001) is a first attempt to accommodate for this demand. moreover, a recent study found that the construct of selfdiscipline, which is an expression of executive motivation and therefore stems from action-control theory, and srl, which is a social-cognitive construct, were highly interrelated (zimmerman & kitsantas, 2014). nevertheless, srl had a significantly higher predictive validity for gpa than did self-discipline, and a twofactor model solution provided better fit indices in this study sample. one attempt to integrate motivation and volition within a social-cognitive framework came from corno and kanfer (1993) who developed a model that illustrates the role of volition in the context of learning and motivation. besides intrinsic and extrinsic motivation, volitional styles, action control and goalrelated cognitions are integrated within the model. nevertheless, this model has, as far as we know, not been empirically tested yet and it neglects metacognitive components which represent a key factor of selfregulated learning. moreover, by integrating a strict distinction between decision-making and action implementation, the model is relatively rigid and neglects possible interactions between the two components. another research line accounting for volitional aspects of learning that is imbedded within the socialcognitive framework of srl focuses on the regulation of motivation and speaks in favour of subsuming motivational beliefs and volition under the broader term of motivation. the construct of motivation regulation is volitional in nature and comprises actions that support the purposeful initiation and maintenance of goal-directed behaviour (wolters, 2003b). therefore, it is a critical aspect of srl that must coalesce with motivational beliefs and metacognitive processes to ensure success when learning. this importance speaks in favour of integrating it into the broader system of srl (wolters, 1999, 2003b). wolters and benzon (2013) have argued that empirical work concerning the link between regulation of motivation and other dimensions of srl has to be extended, and that it has to be examined how this relation could be modelled at a general level (wolters, 2003b). nevertheless, motivational regulation only affects one category of volitional control and therefore does not depict the whole volitional framework (kuhl, 1985). moreover, although volition seems to support the influence of motivational processes on cognitive effort (garcia et al., 1998), husman et al. (2000) argue that research has not yet examined the relationship of motivational processes and volition. concluding, the existing literature speaks in favour of integrating volition within models of srl. as volition has mostly been neglected in srl research, we will present a new conceptualization of volition for srl by integrating academic delay of gratification, future time perspective and procrastination. 1.2.2 an extended conceptualization of trait volition for srl as with srl, volitional traits can be distinguished from volitional states: it is argued that the use of volitional strategies may be an indicator of a dispositional trait ability to reach goals by controlling distractions (corno, 1994). volitional styles represent dispositional tendencies that influence goal implementation and are relatively stable (corno & kanfer, 1993). kuhl’s (1985) differentiation of actionvs. state-orientation is an example for treating volitional styles as a predisposition that influences action. concordantly, boekaerts and corno (2005) designated those strategies as habits supporting an effective working style, suggesting a trait conception. in order to conceptualize a volitional trait in the framework of srl, we integrate academic delay of gratification, procrastination, and future time perspective. these dörrenbächer*&*perels* * * | f l r ! ! 19! constructs have been selected because they represent volition within learning environments (e.g. steel, 2007), show a relatively high stability (e.g. sirois, 2014), are considered as srl features (e.g. bembenutty & karabenick, 2004), and are highly interrelated (e.g. dewitte & lens, 2000). moreover, several authors have argued for their investigation within an srl framework (bembenutty & karabenick, 2004). academic delay of gratification is defined as postponing proximate, impulse satisfying actions to sustain previously intended actions oriented towards a distant but apparently more valuable academic goal (bembenutty, 2008), and is therefore volitional in nature (bembenutty & karabenick, 2004). depending on the underlying theoretical model, academic delay of gratification can be seen as “a volitional strategy, a cognitive schema, a general disposition or a personal trait” (pintrich, 1999, p. 346). since it is mostly described as the ability to wait for temporarily distant rewards (bembenutty, 2008), the construct is seen as a trait in the present study. moreover, several studies speak in favour of the construct’s relevance for academic achievement (e.g. di benedetto & bembenutty, 2013). it has been shown that this trait is associated with the level of self-regulation as well as the use of srl strategies, and can be embedded within a broader framework of srl (bembenutty, 2008, 2009; bembenutty & karabenick, 2004). contradictory to this construct, procrastination is defined as a deliberate delay of intended actions although this delay in all probability has negative effects on reaching an important goal (steel, 2007). it is therefore the opposite of motivated and volitional behaviour (keller, 2008), and demonstrates a lack of effort regulation (rakes & dunn, 2010) or a volitional breakdown (sirois, 2014). concerning its conceptual status, most authors conclude that procrastination results from a trait-like tendency in behaviour (sirois, 2014), with high stability over a period of ten years (steel, 2007). procrastination is a highly relevant construct in academic learning settings because it is very common among students (schouwenburg, 2004) and negatively related to academic achievement (akinsola, tella, & tella, 2007). for the integration of procrastination within an srl framework, it could be argued that the two concepts may be opposite ends on the same regulatory continuum (dietz, hofer, & fries, 2007; park & sperling, 2012), and that the level of procrastination is negatively related to metacognitive strategies (howell & watson, 2007; wolters, 2003a). the concept of future time perspective is represented by a conceptualization of time that is directed to the future and entails future-oriented beliefs concerning specific life domains (peetsma, schuitema, & van der veen, 2012). the construct supports volition because it has a positive influence on the maintenance of motivation (dewitte & lens, 2000) and intention implementation (de bilde, vansteenkiste, & lens, 2011). it is therefore related to the concept of maintenance self-efficacy (luszczynska & sutton, 2006) that is important for persistence on tasks. future time perspective moreover fosters an inner pressure to achieve goals and is therefore a form of volitional motivation. although it can be seen as a cognitive-motivational concept, the perceived instrumentality of future goals causes a stronger effort in learning (simons, dewitte & lens, 2004), why we consider future time perspective as volitional. similar to academic delay of gratification and procrastination, it is conceptualized as a general predisposition that is stable over time (peetsma et al., 2012; zimbardo & boyd, 1999). students with this cognitive temporal bias tend to show better study outcomes (horstmanshof & zimitat, 2007) and are more conscientious (zimbardo & boyd, 1999). several authors have stated that a time perspective directed towards the future has an important influence on motivational processes and the use of srl strategies (de bilde et al., 2011; miller & brickman, 2004; zimmerman, 2011). several results speak in favour of a high interrelation of these constructs, and therefore support the hypothesis that they represent volition. a meta-analysis of 14 samples showed that students with a future time perspective displayed less procrastination (sirois, 2014), and the fact that procrastinators are unable to postpone gratification expresses low levels of academic delay of gratification and low inhibitory control (tuckman, 1991). moreover, some authors have argued that future time perspective is responsible for the conceptual relationship of procrastination and academic delay of gratification (dewitte & lens, 2000) and that it is necessary to delay gratifications in academic settings (bembenutty & karabenick, 2004). concluding, it seems reasonable to integrate these three constructs in order to depict volition within a srl framework. dörrenbächer*&*perels* * * | f l r ! ! 20! 1.3 purpose of the present study as abovementioned, the present study aims to investigate several theoretical aspects concerning the conceptualization of volition within srl: the overall aim is the development and evaluation of an integrative srl trait model that takes into account the previously addressed shortcomings of the hong and o’neil model (2001). besides the integration of a cognitive factor, the proposed model extends the metacognitive and motivational belief component and includes a volitional factor. accompanying this extension, the comparison of two integrative srl trait models should clarify the role of the volitional factor: one model integrates volition as a separate aspect distinct from cognition, metacognition and motivational beliefs (corno, 2001), whereas the other one specifies the motivational component by integrating motivational beliefs and volition (garcia et al., 1998; wolters, 2003b; zimmerman, 2008; see figure 1). moreover, a new conception of trait volition for srl that integrates academic delay of gratification, procrastination, and future time perspective is presented and evaluated. in order to validate the models, their relation to gpa of university entrance diploma will be analysed using structural equation modelling because srl is associated with academic achievement (dignath et al., 2008). altogether, the present study adds to research by examining the structural relationship of academic delay of gratification, future time perspective and procrastination concertedly within an integrative trait model of srl that encompasses cognitive, metacognitive, motivational belief as well as volitional components. it is frontline as it brings together two highly relevant educational concepts whose relation has been analysed rarely. figure 1. two proposed integrative srl models: model 1 includes volition (vol) as a factor besides cognition (cog), metacognition (meta) and motivational beliefs (mb) while model 2 extends the motivational factor (mot) by including volition besides motivational beliefs. 2. method 2.1 sample and sampling procedure the sample comprised n = 381 undergraduate students from a southwestern german university. as four participants had missing data on all variables, they were excluded from the following analyses. therefore, a total sample of n = 377 (70.1% female) with an age range from 17 to 45 years (m = 23.36, sd = 4.12) was analysed. the mean gpa of university entrance diploma was m = 2.10 (sd = 0.60, ranging from 1 = excellent to 4 = poor), indicating that there is no ceiling effect concerning the sample’s academic achievement level. the students were enrolled in very different fields of study (pre-service teachers of dörrenbächer*&*perels* * * | f l r ! ! 21! different subjects [65.0%], psychology [16.2%], languages and cultural studies [9.8%], natural sciences [2.9%], economics and law [2.4%], informatics [1.3%], other/not specified [2.2%]) and therefore reflected a large portion of available fields of study in germany. the sample comprised students of all phases of their studies (year one: 25.0%, year two: 16.9%. year three: 25.7%, year four: 16.8%, year five or higher: 15.7%), while year of study was no predictor for srl (t(376) = -0.24, p = .81). testing was embedded into the first session of several university courses and students had the chance to win a shopping voucher. participants had to sign an informed consent as participation was voluntary and data were anonymised by codes. at the beginning of the test session, every participant received an informed consent form that explained the purpose of the study and the use of the data gained. by signing the form, participants agreed to this procedure. the data collection was part of a larger project and was conducted with the lecturers’ permission. 2.2 instruments 2.2.1 demographic information and academic performance the first part of the questionnaire recorded demographic information such as gender, age, field of study, and gpa of university entrance diploma (ranging from 1 = excellent to 4 = poor). this university entrance diploma is the result of national school exams that are curricular-based and therefore comparable across different schools and regions. all university students pass this exam in the same class level ensuring a comparable educational level. as gpa of university entrance diploma is used for applicant selection at many universities and has a strong relationship with later university achievement (wedler, troche, & rammsayer, 2008), it is very central in the german educational system. moreover, it is comparable between students of all subjects of study which would not be the case for gpa of subject of study (müller-benedict & tsarouha, 2011). 2.2.2 srl inventory in order to measure students’ srl trait, a questionnaire consisting of 32 items was developed. questionnaires are appropriate to measure traits as they record generalized assessments concerning specific domains and abilities (veenman, 2011). the items were adopted from existing inventories that measure srl (e.g. jerusalem & schwarzer, 1981; pintrich, smith, garcia, & mckeachie, 1991; wild & schiefele, 1994) or have been newly developed in order to optimize several subscales. the inventory measured cognitive, metacognitive, and motivational belief variables, and therefore represents the most popular categories of srl traits (boekaerts, 1999). the items were rated on a four-point likert-type format, ranging from 1 (i don’t agree at all) to 4 (i totally agree). subscales with item examples, cronbach’s alphas, and number of items can be seen in table 1. the factor structure of the three srl components was tested using confirmatory factor analysis. the components (cognition, metacognition, motivational beliefs) each were modelled as latent second-order factors with their subscales as latent first-order factors and items as manifest variables. the results are acceptable for motivational beliefs [χ² (50) = 128.61, p < .01, χ²/df = 2.57, rmsea = 0.065 [0.051 – 0.078], srmr = 0.061, cfi = 0.934], metacognition [χ² (71) = 194.74, p < .01, χ²/df = 2.74, rmsea = 0.068 [0.057 – 0.079], srmr = 0.067, cfi = 0.934], and cognition [χ² (8) = 22.31, p < .01, χ²/df = 2.79, rmsea = 0.069 [0.036 – 0.104], srmr = 0.042, cfi = 0.949]. 2.2.3 volition inventory as our study aimed to integrate three volitional traits (future time perspective, procrastination, and academic delay of gratification), a questionnaire comprising items of these three constructs was developed. the items stem from existing inventories to measure these constructs (academic delay of gratification scale, bembenutty & karabenick, 1998; procrastination scale, tuckman, 1991; future scale of the zimbardo time perspective inventory, zimbardo & boyd, 1999). the items of future time perspective and procrastination were rated on a four-point likert-type format, ranging from 1 (i don’t agree at all) to 4 (i totally agree). the procrastination items register the presence of procrastinating behaviour and therefore dörrenbächer*&*perels* * * | f l r ! ! 22! present a lack of volition. as the items of the academic delay of gratification scale consisted of two action alternatives, a and b, that represent the ends of a “delay of gratification-continuum”, these were rated on a four-point answer format, ranging from 1 (definitely chose a), 2 (rather chose a), 3 (rather chose b), and 4 (definitely chose b). b represents the alternative that reflects the highest ability to delay gratification and therefore the strongest volitional control. all items should be answered with regard to students’ learning behaviour. subscales with item examples, cronbach’s alphas and number of items can be seen in table 1. dörrenbächer*&*perels* * * | f l r ! ! 23! table 1 scales and subscales of the self-regulated learning and volition inventory scale subscale item example cronbach’s alpha (number of items) metacognition planning “i write a time schedule before i start learning.” .88 (5) self-recording “i pay attention to not miss my goal when i’m learning.” .70 (4) self-evaluation “after learning, i check if i’ve reached my goals.” .76 (5) cognition organization “i draw charts or diagrams in order to structure learning materials.” .52 (3) critical thinking “i critically question things i learn.” .70 (3) motivational beliefs self-efficacy “i’m able to find a solution for every problem.” .77 (5) intrinsic motivation “i enjoy learning.” .70 (3) goal orientation “i prefer tasks that are interesting, even if they’re difficult to solve.” .68 (4) volition future time perspective “i even work on difficult and boring tasks, when i know, that they are important for my future.” .74 (4) academic delay of gratification „i would rather a spend time with friends shortly before an exam or b learn each day for the exam and spend less time with friends.” .68 (4) procrastination “if something is too difficult to start with, i postpone the task.” .87 (6) dörrenbächer*&*perels* * * | f l r ! ! 24! 2.3 data analysis in order to test our conception of trait volition and the extended model of trait srl, we used maximum likelihood parameter estimation with mplus7 (muthén & muthén, 2012). using confirmatory factor analysis, the factorial structure and the fit of the proposed models were estimated. the model fit is assessed by evaluating the model’s χ², its rmsea (root mean square error of approximation), srmr (standardized root mean square residual), and cfi (comparative fit index). a good fit is characterized by a non-significant χ² (p > .05). as this test is less reliable with large sample sizes (kline, 2005), one can examine the χ²/df-ratio, which should be below 2:1 to mark an acceptable fit (schermelleh-engel, moosbrugger, & müller, 2003). rmsea and srmr values should be ≤ 0.08 and the cfi as a fit-index has to be > 0.90 to indicate a good fit (kline, 2005). besides these model fit estimations, we assessed the criterion validity by analyzing the models’ relation to the gpa of university entrance diplomas using structural equation modelling (sem). 3. results 3.1 initial data screening table 2 shows the descriptive statistics and the bivariate zero-order correlation matrix of the measured variables. in a first step, the data were screened in order to find outliers and to examine missing data, as well as to assess the linearity and normality of the data. except for the four participants that were excluded from the analyses because they had missing values on all variables, there were no participants with a lot of missing data or outlier values. for all variables of the srl questionnaire as well as the volition inventory and the gpa of university entrance diploma little’s mcar test (little & rubin, 2002) indicated that the missing data in this study occurred completely at random (χ² (35) = 29.69, p = .72). mplus7 uses the fiml-estimator (full information maximum likelihood) to treat missing values, so we did not impute them. moreover, the data violated the assumption of a normal distribution why the mlr estimator—which is robust to non-normality—was used to run the analyses. dörrenbächer*&*perels* * * | f l r ! ! 25! table 2 bivariate correlations among variable as well as their means and standard deviations 1 2 3 4 5 6 7 8 9 10 11 12 1. gpa .16** -.17** -.23** -.08 -.18** -.22** -.16** -.17** -.15** -.13* -.16** 2. procrastination -.74** -.42** -.11* -.29** -.26** -.53** -.49** -.43** -.33** -.22** 3. future time perspective .47** .09 .26** .29** .51** .54** .38** .36** .25** 4. academic delay of gratification -.10 .28** .20** .39** .30** .28** .30** .13* 5. self-efficacy .26** .43** .05 .12* .01 .06 .17** 6. intrinsic motivation .49** .10 .20** .16** .23** .35** 7. goal-orientation .18** .27** .18** .21** .35** 8. planning .50** .44** .37** .12* 9. self-recording .54** .45** .42** 10. self-evaluation .23** .26** 11. organization .28** 12. critical thinking m 2.10 2.60 2.78 3.01 2.81 3.05 3.05 2.44 2.85 2.60 2.77 2.44 sd 0.60 0.72 0.57 0.63 0.54 0.56 0.51 0.75 0.48 0.57 0.61 0.60 note. 313 ≤ n ≤ 377, * p < .05, ** p < .01 dörrenbächer*&*perels* * * | f l r ! ! 26! 3.2 confirmatory factor analysis of trait volition in order to examine the factor structure of the 14 items of the volition inventory, a latent model of volition with three latent first-order factors (procrastination, academic delay of gratification, future time perspective) and one latent second-order factor (volition) using confirmatory factor analyses was tested. the results speak in favour of a good model fit: χ² (75) = 123.12, p < .01, χ²/df = 1.64, rmsea = 0.041 [0.028 – 0.054], srmr = 0.039, cfi = 0.971. figure 2 shows the model of trait volition with standardized factor loadings that are all significant (all p values < .001). the first-order procrastination factor has a negative loading as the procrastination items register the presence of procrastinating behaviour and therefore present a lack of volition. figure 2. trait model of volition for srl with standardized coefficients. vol volition, ftp future time perspective, pro procrastination, adog academic delay of gratification. all factor loadings are significant (p < .001). dörrenbächer*&*perels* * * | f l r ! ! 27! 3.3 testing the hypothesized model in a second step, we tested the comprehensive trait model of srl by including the volitional factor besides cognitive, metacognitive and motivational belief components into the model of srl. as the internal consistencies of the srl inventory show acceptable to satisfying values and the factor structure of the volition inventory was confirmed, the latent factors were estimated by using the respective subscales as observed variables. because we wanted to examine whether volition—as we conceptualize it—is a fourth factor of srl (model 1) or is a part of motivation in addition to motivational beliefs (model 2), we compared two models. table 3 shows the fit-indices of both models, indicating that the model with volition as a subcomponent of motivation fits the data more adequately (model 2). moreover, as the model alternatives were constructed based on theoretical assumptions, we used the bayesian information criterion to compare the two models (burnham & anderson, 2004). model 2 has the lower bic value and therefore seems more appropriate to model the data (geiser, 2011, see table 3). table 3 fit-indices of the compared srl-models model χ² df χ²/df rmsea srmr cfi bic 1 54.13 32 1.69 0.043 [0.022 – 0.062] 0.040 0.980 6287.76 2 45.11 31 1.46 0.035 [0.003 – 0.056] 0.035 0.987 6283.48 3.4 testing the criterion validity of srl for achievement in order to test the relation of our integrative srl model with academic achievement, we included gpa of the university entrance diploma as a manifest variable into the structural model. although this measure is not really predictive because it was obtained in the past, we had to take this gpa instead of current gpa of university subject because it is more comparable between students of different subjects of study (müller-benedict & tsarouha, 2011). additionally, gpa of university entrance diploma is a very central achievement marker in the german education system (wedler et al., 2008). moreover, one could argue that srl as a trait should also have predictive value for past indicators because it should not change that much over time. both models yield a good fit (model 1: χ² (42) = 70.73, p < .01, χ²/df = 1.68, rmsea = 0.043 [0.024 – 0.059], srmr = 0.043, cfi = 0.975, bic = 6960.42; model 2: χ² (41) = 63.75, p < .01, χ²/df = 1.55, rmsea = 0.038 [0.018 – 0.056], srmr = 0.041, cfi = 0.981, bic = 6958.49) with all significant factor loadings (p < .001) and highly significant correlations with gpa (model 1: r = -0.25, p < .001; model 2: r = -0.23, p < .001). although the correlation with gpa is slightly higher for model 1, we depict model 2 in figure 3 as it yields better fit indices. dörrenbächer*&*perels* * * | f l r ! ! 28! figure 3. structural equation model of trait srl and achievement with standardized coefficients. gpa grade point average, srl self-regulated learning, cog cognitive components, ct critical thinking, org organization, meta metacognitive components, seva self-evaluation, srec self-recording, plan planning, mot motivational components, vol volition, adog academic delay of gratification, pro procrastination, ftp future time perspective, mb motivational beliefs, se self-efficacy, im intrinsic motivation, go goal orientation. all factor loadings are significant (p < .001). 4. discussion the present study aimed to test a comprehensive trait srl model that integrates volition besides cognition, metacognition and motivational beliefs. moreover, a new conception of trait volition for learning was tested empirically. the results confirm the hypothesized structure of volition and speak in favour of a twofold motivational component for srl that comprises motivational beliefs as well as volition. furthermore, the trait model of srl is related to gpa of university entrance diploma emphasizing the importance of srl for academic achievement. the first aim of the present study was to examine a new conception of trait volition for learning environments that integrates academic delay of gratification, future time perspective and procrastination. the results suggest that volition within learning environments comprises the competence of postponing available gratification in order to attain important academic goals, and thus is the opposite of delaying an intended action. this motivational regulation is supported by a personal time frame that is directed to the future and focuses on the instrumentality of long-term goals. the present study therefore answers some authors’ call for dörrenbächer*&*perels* * * | f l r ! ! 29! a systematic investigation of the hypothesized correlational framework (bembenutty & karabenick, 2004) and is the first to examine the three constructs concertedly by testing their underlying structure empirically. a shortcoming of the conception is the relatively low factor loading of academic delay of gratification, which could be caused by the items’ answer format. as it differs from that of the procrastination and future time perspective items, an adaption to the same measurement format could improve the factor loading. although the model is not exhausting, it presents an integrative and broad conception that helps to explain and examine trait volition during learning. the second aim of the study was to test a comprehensive trait model of srl by extending the hong and o’neil model (2001) and by incorporating the new conception of trait volition (e.g. duckworth et al., 2014; wolters & benzon, 2013). firstly, the modelling results speak in favour of including a cognitive component that comprises the usage of organization as well as elaborative learning strategies. secondly, results allow for the conclusion that the metacognitive, volitional and motivational belief components should be specified by considering further subcomponents. thirdly, the findings support the assumption of a comprehensive motivational component that comprises motivational beliefs coinciding with volition. this result clarifies the role of volition as a critical part of goal-oriented behaviour: optimal motivational beliefs are only valuable if distractions in the course of action can be handled, especially when tasks take several weeks to be completed (husman et al., 2000). the motivational component of the presented model is therefore in agreement with the differentiation made in action-control theory: the choice of goals during the pre-decisional phase is influenced by self-efficacy, intrinsic value and goal orientation and therefore is named choice motivation. the implementation of planned intentions in the post-decisional phase however is named executive motivation and can only be secured if volition supports motivational beliefs and therefore regulates motivation (zimmerman, 2011). hence, our model integrates executive motivation in addition to choice motivation, which has largely been neglected within self-regulated learning research (garcia et al., 1998). nevertheless, future studies are needed to test sequential aspects of this model. longitudinal and experimental investigations could help to clarify predictive relations between motivational beliefs and volitional actions as they should reflect different segments of a goal-oriented process. although the model with a twofold motivational component yields a good fit, it is striking that the latent second-order motivational variable is represented strongly by the first-order volitional factor and only moderately by the first-order motivational belief factor. one explanation for this result could be that the relationship between motivational beliefs and volition is not linear, but rather curvilinear in type. wolters (2003b) has argued that highly motivated students do not need to make use of volitional strategies, whereas totally unmotivated students cannot summon the willingness to enact such strategies. volition thus could be regarded as mediating the influence of motivational beliefs on the use of srl strategies (gaeta, teruel, & orejudo, 2012; garcia et al., 1998) or as supportive for student’s motivation in general (husman et al., 2000). propositions of action-control theory are in line with the mediator hypothesis: whereas motivational beliefs are central components of the pre-decisional phase that lead to the choice of a goal, volition refers to post-decisional processes that concern the implementation of intentions (zimmerman, 2011). future research could analyse this hypothesis by conducting longitudinal research with a cross-lagged panel design and study the direction of possible effects. an additional shortcoming of our model is the fact that the motivational belief factor shows a weak factor loading for the self-efficacy indicator. this finding is consistent with the notion of some authors that self-efficacy is rather a motivational precursor than a part of motivation (usher & pajares, 2008). moreover, only one item used in the questionnaire referred to self-efficacy for university context, while four items asked participants to rate their general self-efficacy (jerusalem & schwarzer, 1981). as general self-efficacy concerns the handling of global, unspecified problems, it is not necessarily related to academic self-efficacy (topkaya, 2010). thus, it would be more adequate to use items that measure academic self-efficacy, or even self-efficacy for self-regulated learning and the use of learning strategies (schunk & usher, 2011). the third aim of the present study was to validate the srl models by testing their relation to gpa of university entrance diploma. while model 1 showed a slightly higher relation to gpa, model 2 yielded a dörrenbächer*&*perels* * * | f l r ! ! 30! better fit. therefore, we decided to favour the model that integrates volition within the motivational component besides motivational beliefs as this is in accordance with action-control theory. although the relation with gpa is moderate, it is highly significant and therefore emphasizes the importance of srl for academic achievement. an explained variance of about 5% is small but comparable to the results of similar studies (e.g. balkis, duru, & bulus, 2013). the inclusion of interindividual variables, such as intelligence, personality, or attitudes could increase the prediction of gpa. taken together, the results of the present study underline the importance of volition within srl frameworks and confirm the construct’s relevance for academic achievement, especially because students with high values on the examined srl traits show better academic performance. 4.1 limitations although the data speak in favour of integrating volition within a broader trait srl framework, several methodological limitations are present that should be considered when interpreting the results: the achievement marker used in the present study is not optimal as it is retrospective. we chose the criterion of gpa of university entrance diploma since it is very central in the german educational system: it is the result of national exams that are curricular-based and thus comparable across different schools and regions. all university students pass this exam in the same class level ensuring a comparable educational level. moreover, it is used for applicant selection at many universities and has a strong relationship with later university achievement (wedler et al., 2008). therefore, it is comparable between students of all subjects of study. although the analysis corresponds to some kind of retrodiction, it is justifiable because we regard srl as a stable trait that should be related to past indicators. future studies could analyse the predictive validity of the integrative srl model for current gpa of subject of study. in order to obtain reliable results, students should be from the same field of study because grades of different subjects are not comparable in the german college system (müller-benedict & tsarouha, 2011). moreover, they should be in the same phase of their studies in order to obtain comparable experiences with university exams. additionally, future studies should aim at using objective achievement markers stemming from performance tests to validate models like the one proposed in this study. another limitation is the heterogeneity of the sample used: participants studied a wide range of subjects and were in different phases of their studies. this reduces the interindividual comparability concerning study experiences and interest structures. nevertheless, the obtained results using such a heterogeneous sample speak in favour of srl as an important factor for all fields of study. moreover, the sample was highly selective, as college students have the highest school degree available in germany and thus represent the upper ability continuum. future studies should validate srl models on a sample that is more representative for the diversity of our society and the lifelong learners living in it. additionally, the participants were predominantly female. as previous studies have shown, males report higher self-efficacy values (huang, 2013) and show higher values concerning the use of elaborative learning strategies (bembenutty, 2007). nevertheless, females report higher academic delay of gratification (bembenutty, 2007). an exploratory multivariate analysis of variance using the srl and volition subscales as dependent variables and gender as independent variable indicated an effect of gender as females in our study reported significant higher academic delay of gratification and significant lower self-efficacy beliefs than males. consequently, structural analyses could be conducted for both genders separately to investigate if factor structures differentiate between the groups. as mentioned previously, the different answer formats of the instruments used as well as the low reliability for the subscale of organizational learning strategies represent further limitations. future examinations should aim to adjust the instruments concerning their structure in order to make them more comparable and should choose more reliable items to depict organizational learning strategies. dörrenbächer*&*perels* * * | f l r ! ! 31! 4.2 implications and future directions the presented findings have several implications for educational researchers: as the results speak in favour of the integration of academic delay of gratification, future time perspective, and procrastination to conceptualize volition for learning, future research could cross-validate this conception with samples of different age and educational groups. in addition, the findings support the trait srl model with an expanded motivational component that integrates motivational beliefs and volition. this is why future research should include volitional factors into the theoretical research basis along with other motivational constructs. moreover, analyses of the model’s relationship with several variables of interest, such as intelligence, personality traits, or attitudes could result in further important insights. latent profile analyses could support the investigation of the relation between interpersonal differences and the trait model of srl and could help to indentify types of learners that need different types of interventions. moreover, it would be interesting to examine the model’s stability by conducting longitudinal research as stability measurements would militate in favour of the trait concept of srl. questionnaire methods are indispensible for measuring traits because participants have to aggregate their behaviour concerning several situations, which is in accordance with the situation-independent trait perspective. nevertheless, future research could complement self-report scales with qualitative instruments, such as interviews or thinking aloud protocols (veenman, 2011), representing a multimethod approach. subsequently, with regard to the analysis of trait srl, future studies could transfer the model to state level and examine its structure using process measures. as trait and state srl are highly interrelated (hong, 1998), it would be interesting to systematically analyse whether the components of our model are present in all phases of srl using process models as a theoretical basis (e.g. zimmerman, 2000). moreover, state analyses could focus the question which strategies are used to ensure volitional control. altogether, it seems appropriate to assume that cognitive and metacognitive variables as well as motivational beliefs and volition are important for planning, performing, and reflecting upon one’s learning. nevertheless, more research is needed to derive practical implications based on these theoretical findings. longitudinal studies that investigate the stability of the constructs, their reciprocal relations as well as their development and interconnection in earlier stages of life could be helpful. keypoints academic delay of gratification, procrastination and future time perspective can be integrated in order to depict volition for srl. an srl trait model that comprises cognitive, metacognitive and motivational components yields a good fit. volition for learning can be integrated within that model by extending the motivational component and adding volition above and beyond motivational beliefs. the proposed srl trait model is related to gpa which is a first hint of its validity. dörrenbächer*&*perels* * * | f l r ! ! 32! acknowledgments the authors wish to thank the doctoral training programme of saarland university (gradus) for funding the present research. references akinsola, m. k., tella, a., & tella, a. (2007). correlates of academic procrastination and mathematics achievement of university undergraduate students. eurasia journal of mathematics, science & technology education, 3(4), 363-370. azevedo, r. (2014). issues in dealing with sequential and temporal characteristics of self-and sociallyregulated learning. metacognition and learning, 9(2), 217-228. doi:10.1007/s11409-014-9123-1 balkis, m., duru, e., & bulus, m. (2013). analysis of the relation between academic procrastination, academic rational/irrational beliefs, time preferences to study for exams, and academic achievement: a structural model. european journal of psychology of education, 28(3), 825-839.!doi:10.1007/s10212012-0142-5 bandura, a. (1986). social foundations of thought and action: a social cognitive theory. englewood cliffs, nj: prentice hall. bandura, a. (1997). self-efficacy: the exercise of control. new york: w. h. freeman. bembenutty, h. (2007). self-regulation of learning and academic delay of gratification: gender and ethnic differences among college students. journal of advanced academics, 18(4), 586-616. doi:10.4219/jaa-2007-553 bembenutty, h. (2008). academic delay of gratification and expectancy–value. personality and individual differences, 44(1), 193-202.!doi:10.1016/j.paid.2007.07.025 bembenutty, h. (2009). teaching effectiveness, course evaluation, and academic performance: the role of academic delay of gratification. journal of advanced academics, 20(2), 326-355. doi:10.1177/1932202x0902000206 bembenutty, h., & karabenick, s. a. (1998). academic delay of gratification. learning and individual differences, 10(4), 329-346. doi:10.1016/s1041-6080(99)80126-5 bembenutty, h., & karabenick, s. a. (2004). inherent association between academic delay of gratification, future time perspective, and self-regulated learning. educational psychology review, 16(1), 35-57. doi:10.1023/b:edpr.0000012344.34008.5c boekaerts, m. (1999). self-regulated learning: where we are today. international journal of educational research, 31(6), 445-457. doi:10.1016/s0883-0355(99)00014-2 boekaerts, m., & corno, l. (2005). self/regulation in the classroom: a perspective on assessment and intervention. applied psychology, 54(2), 199-231. doi:10.1111/j.1464-0597.2005.00205.x bronson, m. b. (2000). self-regulation in early childhood. nature and nurture. new york: the guilford press. burnham, k. p. & anderson, d. r. (2004). multimodel inference: understanding aic and bic in model selection. sociological methods & research, 33(2), 261-304. doi:10.1177/0049124104268644 carver, c. s., & scheier, m. f. (2000).on the structure of behavioral self-regulation. in m. boekaerts, p. r. pintrich, & m. zeidner (eds.), handbook of self-regulation (pp 42-80). san diego: academic press. corno, l. (1993). the best-laid plans: modern conceptions of volition and educational research. educational researcher, 22(2), 14-22. doi:10.3102/0013189x022002014 corno, l. (1994). student volition and education: outcomes, influences, and practices. in d. h. schunk & b. j. zimmerman (eds.), self-regulation of learning and performance: issues and educational applications (pp. 229-251). hillsdale, nj: lawrence erlbaum associates. dörrenbächer*&*perels* * * | f l r ! ! 33! corno, l. (2001). volitional aspects of self-regulated learning. in b. j. zimmerman & d. h. schunk (eds.), self-regulated learning and academic achievement. theoretical perspectives (pp. 191-226). mahwah, nj: erlbaum. corno, l., & kanfer, r. (1993). the role of volition in learning and performance. in l. darling-hammond (ed.), review of research in education (vol. 21, pp. 301-341). itasca, il: f.e. peacock publishers. de bilde, j., vansteenkiste, m., & lens, w. (2011). understanding the association between future time perspective and self-regulated learning through the lens of self-determination theory. learning and instruction, 21(3), 332-344. doi:10.1016/j.learninstruc.2010.03.002 dewitte, s., & lens, w. (2000). procrastinators lack a broad action perspective. european journal of personality, 14(2), 121-140. doi:10.1002/(sici)1099-0984(200003/04)14:2<121::aidper368>3.0.co;2-# dibenedetto, m. k., & bembenutty, h. (2013). within the pipeline: self-regulated learning, self-efficacy, and socialization among college students in science courses. learning and individual differences, 23, 218-224. doi:10.1016/j.lindif.2012.09.015 dietz, f., hofer, m., & fries, s. (2007). individual values, learning routines and academic procrastination. british journal of educational psychology, 77(4), 893-906. doi:10.1348/000709906x169076 dignath, c., buettner, g., & langfeldt, h. p. (2008). how can primary school students learn self-regulated learning strategies most effectively?: a meta-analysis on self-regulation training programmes. educational research review, 3(2), 101-129. doi:10.1016/j.edurev.2008.02.003 duckworth, a. l., gendler, t. s., & gross, j. j. (2014). self-control in school-age children. educational psychologist, 49(3), 199-217. doi:10.1080/00461520.2014.926225 duckworth, a. l., & seligman, m. e. (2006). self-discipline gives girls the edge: gender in self-discipline, grades, and achievement test scores. journal of educational psychology, 98(1), 198. doi:10.1037/00220663.98.1.198 efklides, a. (2011). interactions of metacognition with motivation and affect in self-regulated learning: the masrl model. educational psychologist, 46(1), 6-25. doi:10.1080/00461520.2011.538645 gaeta, m. l., teruel m. p., & orejudo, s. (2012). motivational, volitional and metacognitive aspects of self regulated learning. electronic journal of research in educational psychology, 10(1), 73-94. garcia, t., mccann, e. j., turner, j. e., & roska, l. (1998). modeling the mediating role of volition in the learning process. contemporary educational psychology, 23(4), 392-418. doi:10.1006/ceps.1998.0982 geiser, c. (2011). datenanalyse mit mplus [data analysis using mplus]. wiesbaden: vs verlag. gollwitzer, p. m. (1996). the volitional benefits of planning. in p. m. gollwitzer, & j. a. barge (eds.), the psychology of action (pp. 287-312). new york: guilford press. heckhausen, h., & kuhl, j. (1985). from wishes to action: the dead ends and short cuts on the long way to action. in m. frese & j. sabini (eds.), goal-directed behavior: the concept of action in psychology (pp. 134-160). hillsdale, nj: lawrence erlbaum associates. hertzog, c., & nesselroade, j. r. (2003). assessing psychological change in adulthood: an overview of methodological issues. psychology and aging, 18(4), 639. doi:10.1037/0882-7974.18.4.639 hong, e. (1995). a structural comparison between state and trait self/regulation models. applied cognitive psychology, 9(4), 333-349.!doi:10.1002/acp.2350090406 hong, e. (1998). differential stability of state and trait self-regulation in academic performance. the journal of educational research, 91(3), 148-159. doi:10.1080/00220679809597536 hong, e., & o'neil jr, h. f. (2001). construct validation of a trait self-regulation model. international journal of psychology, 36(3), 186-194. doi:10.1080/00207590042000146 horstmanshof, l., & zimitat, c. (2007). future time orientation predicts academic engagement among first/ year university students. british journal of educational psychology, 77(3), 703-718. doi:10.1348/000709906x160778 howell, a. j., & watson, d. c. (2007). procrastination: associations with achievement goal orientation and learning strategies. personality and individual differences, 43(1), 167-178. doi:10.1016/j.paid.2006.11.017 dörrenbächer*&*perels* * * | f l r ! ! 34! huang, c. (2013). gender differences in academic self-efficacy: a meta-analysis. european journal of psychology of education, 28(1), 1-35. doi:10.1007/s10212-011-0097-y husman, j., mccann, e., & crowson, h. m. (2000). volitional strategies and future time perspective: embracing the complexity of dynamic interactions. international journal of educational research, 33(7), 777-799. doi:10.1016/s0883-0355(00)00050-1 jerusalem, m., & schwarzer, r. (1981). selbstwirksamkeit [self-efficacy]. in r. schwarzer (eds.), skalen zur befindlichkeit und persönlichkeit [scales to measure mood and personality] (pp. 15-28). berlin: freie universität berlin. keller, j. m. (2008). an integrative theory of motivation, volition, and performance. technology, instruction, cognition, and learning, 6(2), 79-104. kitsantas, a. (2002). test preparation and performance: a self-regulatory analysis. the journal of experimental education, 70(2), 101-113. doi:10.1080/00220970209599501 kitsantas, a., winsler, a., & huie, f. (2008). self-regulation and ability predictors of academic success during college: a predictive validity study. journal of advanced academics, 20(1), 42-68. doi:10.4219/jaa-2008-867 kline, r. b. (2005). principles and practice of structural equation modeling (2nd ed.). new york: guilford. kuhl, j. (1984). volitional aspects of achievement motivation and learned helplessness: toward a comprehensive theory of action control. in b. a. maher & w. b. maher (eds.), progress in experimental personality research (pp. 101-171). orlando: academic press. kuhl, j. (1985). volitional mediators of cognition-behaviour consistency. self-regulatory processes and action vs. state orientation. in j. kuhl & j. beckmann (eds.), action control: from cognition to behavior (pp. 101-128). new york: springer. kuhl, j. (2000). the volitional basis of personality systems interaction theory: applications in learning and treatment contexts. international journal of educational research, 33(7-8), 665-703. doi:10.1016/s0883-0355(00)00045-8 little, r. j. a., & rubin, d. b. (2002). statistical analysis with missing data. hoboken: john wiley & sons. luszczynska, a., & sutton, s. (2006). physical activity after cardiac rehabilitation: evidence that different types of self-efficacy are important in maintainers and relapsers. rehabilitation psychology, 51(4), 314321. doi:10.1037/0090-5550.51.4.314 matthews, g., schwean, v. l., campbell, s. e., saklofske, d. h., & mohamed, a. r. (2000). personality, self-regulation, and adaptation: a cognitive-social framework. in m. boekaerts, p. r. pintrich, m. zeidner, m. boekaerts, p. r. pintrich, m. zeidner (eds.) , handbook of self-regulation (pp. 171-207). san diego, ca, us: academic press. miller, r. b., & brickman, s. j. (2004). a model of future-oriented motivation and self-regulation. educational psychology review, 16(1), 9-33. doi:10.1023/b:edpr.0000012343.96370.39 müller-benedict, v., & tsarouha, e. (2011). können examensnoten verglichen werden? eine analyse von einflüssen des sozialen kontextes auf hochschulprüfungen [are grades in exams comparable to each other? the impact of social context on grading in higher education]. zeitschrift für soziologie, 388-409. muthén, l. k., & muthén, b. o. (2012). mplus: statistical analysis with latent variables (version 7) [computer software]. los angeles: authors. nandagopal, k., & ericsson, k. a. (2012). an expert performance approach to the study of individual differences in self-regulated learning activities in upper-level college students. learning and individual differences, 22, 597–609. doi:10.1016/j.lindif.2011.11.018 park, s. w., & sperling, r. a. (2012). academic procrastinators and their self-regulation. psychology, 3(01), 12. doi:10.4236/psych.2012.31003 peetsma, t., schuitema, j., & van der veen, i. (2012). a longitudinal study on time perspectives: relations with academic delay of gratification and learning environment. japanese psychological research, 54(3), 241-252. doi:10.1111/j.1468-5884.2012.00526.x pintrich, p. r. (1999). taking control of research on volitional control: challenges for future theory and research. learning and individual differences, 11(3), 335-354. doi:10.1016/s1041-6080(99)80007-7 dörrenbächer*&*perels* * * | f l r ! ! 35! pintrich, p. r. (2000). the role of goal orientation in self-regulated learning. in m. boekaerts, p. r. pintrich & m. zeidner (eds.), handbook of self-regulation (pp. 451-502). san diego: academic press. pintrich, p. r., smith, d. a. f., garcia, t., & mckeachie, w. j. (1991). a manual for the use of the motivated strategies for learning questionnaire (mslq). michigan: university of michigan. rakes, g. c., & dunn, k. e. (2010). the impact of online graduate students’ motivation and self-regulation on academic procrastination. journal of interactive online learning, 9(1), 78-93. schermelleh-engel, k., moosbrugger, h., & müller, h. (2003). evaluating the fit of structural equation models: tests of significance and descriptive goodness-of-fit measures. methods of psychological research online, 8(2), 23-74. schmitz, b., & wiese, b. s. (2006). new perspectives for the evaluation of training sessions in self-regulated learning: time-series analyses of diary data. contemporary educational psychology, 31(1), 64-96. doi:10.1016/j.cedpsych.2005.02.002 schouwenburg, h. c. (2004). academic procrastination: theoretical notions, measurement, and research. in h. c. schouwenburg, c. h. lay, t. a. pychyl, & j. r. ferrari (eds.), counseling the procrastinator in academic settings (pp. 3–17). washington, dc: american psychological association. schunk, d., & usher, e. (2011). assessing self-efficacy for self-regulated learning. in b. zimmerman & d. schunk (eds.), handbook of self-regulation of learning and performance (p. 282-297).!new york, ny: routledge. simons, j., dewitte, s., & lens, w. (2004). the role of different types of instrumentality in motivation, study strategies, and performance: know why you learn, so you'll know what you learn!. british journal of educational psychology, 74(3), 343-360. doi:10.1348/0007099041552314 sirois, f. m. (2014). out of sight, out of time? a meta/analytic investigation of procrastination and time perspective. european journal of personality, 28(5), 511-520. doi:10.1002/per.1947 steel, p. (2007). the nature of procrastination: a meta-analytic and theoretical review of quintessential selfregulatory failure. psychological bulletin, 133(1), 65-94. doi:10.1037/0033-2909.133.1.65 topkaya, e. z. (2010). pre-service english language teachers' perceptions of computer self-efficacy and general self-efficacy. turkish online journal of educational technology, 9(1), 143-156. tuckman, b. w. (1991). the development and concurrent validity of the procrastination scale. educational and psychological measurement, 51(2), 473-480. doi:10.1177/0013164491512022 usher, e. l., & pajares, f. (2008). self-efficacy for self-regulated learning: a validation study. educational and psychological measurement, 68(3), 443-463. doi:10.1177/0013164407308475 veenman, m. v. (2011). alternative assessment of strategy use with self-report instruments: a discussion. metacognition and learning, 6(2), 205-211. doi:10.1007/s11409-011-9080-x wedler, b., troche, s., & rammsayer, t. (2008). bericht: studierendenauswahl – eignungsdiagnostischer nutzen von noten aus schule und studium [report: selection of students – diagnostic benefit of grades in school and study]. psychologische rundschau, 59(2), 123-125. doi:10.1026/0033-3042.59.2.123 wild, k., & schiefele, u. (1994). lernstrategien im studium: ergebnisse zur faktorenstruktur und reliabilität eines neuen fragebogens [learning strategies in college: results concerning the factorial structure and reliability of a new questionnaire]. zeitschrift für differentielle und diagnostische psychologie, 15(4), 185-200. winne, p. h., & perry, n. e. (2000). measuring self-regulated learning. in m. boekaerts, p. r. pintrich & m. zeidner (eds.), handbook of self-regulation (pp. 532-568). san diego: academic press. wolters, c. a. (1999). the relation between high school students' motivational regulation and their use of learning strategies, effort, and classroom performance. learning and individual differences, 11(3), 281299. doi:10.1016/s1041-6080(99)80004-1 wolters, c. a. (2003a). understanding procrastination from a self-regulated learning perspective. journal of educational psychology, 95(1), 179. doi:10.1037/0022-0663.95.1.179 wolters, c. a. (2003b). regulation of motivation: evaluating an underemphasized aspect of self-regulated learning. educational psychologist, 38(4), 189-205. doi:10.1207/s15326985ep3804_1 dörrenbächer*&*perels* * * | f l r ! ! 36! wolters, c. a., & benzon, m. b. (2013). assessing and predicting college students’ use of strategies for the self-regulation of motivation. the journal of experimental education, 81(2), 199-221. doi:10.1080/00220973.2012.699901 wolters, c. a., & hussain, m. (2014). investigating grit and its relations with college students’ selfregulated learning and academic achievement. metacognition and learning, 1-19. doi:10.1007/s11409014-9128-9 zimbardo, p. g., & boyd, j. n. (1999). putting time in perspective: a valid, reliable individual-differences metric. journal of personality and social psychology, 77(6), 1271-1288. doi:10.1037/00223514.77.6.1271 zimmerman, b. j. (2000). attaining self-regulation: a social cognitive perspective. in m. boekaerts, p. r. pintrich & m. zeidner (eds.), handbook of self-regulation (pp. 13 – 41). san diego: academic press. zimmerman, b. j. (2008). investigating self-regulation and motivation: historical background, methodological developments, and future prospects. american educational research journal, 45(1), 166–183. doi:10.3102/0002831207312909 zimmerman, b. j. (2011). motivational sources and outcomes of self-regulated learning and performance. in b. j. zimmerman & d. h. schunk. (eds.), handbook of self-regulation of learning and performance (pp. 49-64). new york, ny: routledge. zimmerman, b. j., & kitsantas, a. (2014). comparing students’ self-discipline and self-regulation measures and their prediction of academic achievement. contemporary educational psychology, 39(2), 145-155. doi:10.1016/j.cedpsych.2014.03.004 zimmerman, b. j., & schunk, d. h. (2001). self-regulated learning and academic achievement: theoretical perspectives (2nd ed.). mahwah, nj, us: lawrence erlbaum associates publishers. microsoft word raes_publication.docx           frontline learning research vol.3 no. 2 (2015) 63-89 issn 2295-3159   1  corresponding author: elisabeth raes, dekenstraat 2, 3772, 3000 leuven, belgium, phone: +32 16 32 62 40, e-mail: elisabeth.raes@ppw.kuleuven.be, doi: http://dx.doi.org/10.14786/flr.v3i2.166   turning points during the life of student project teams: a qualitative study elisabeth raesa, eva kyndta, & filip dochya a occupational & organizational psychology and professional learning, university of leuven, belgium article received 24 april 2015 / revised 28 june 2015 / accepted 2 july 2015 / available online 19 august 2015 abstract in this qualitative study a more flexible alternative of conceptualising changes over time in teams is tested within student project teams. the conceptualisation uses turning points during the lifespan of a team to outline team development, based on work by erbert, mearns, & dena (2005). turning points are moments that made a significant difference during the course of the collaboration as a team. in this study, they are tracked by means of team interviews and reflection papers of team members. a method of coding was created to collect all information about the turning points, their causes and consequences. by means of a thorough analysis of these coded data an overview of their nature and their effects on the rest of the team process as perceived by the team members themselves is provided. results show that the development paths of the three teams were differentiated in terms of turning points that occurred and, especially, in the order in which the turning points occurred. however four types of turning points (two at the task level en two at the interpersonal level) were remarkable due to their occurrence in all three project teams. keywords: team development, knowledge work teams, turning points, qualitative raes       | f l r     64   1. introduction in organisations, teamwork is set up to create interdependent collaboration between team members to accomplish a common task. typically, these teams are composed of team members with diverse background, habits and behaviour patterns concerning work and (team) work relationships. team members are confronted with the challenge to efficiently combine their individual knowledge, skills and attitudes into an effective working team in order to finish the tasks they were assigned (salas, sims & burke, 2005). the way team members recognise their diversity and learn to use it as strength is of crucial importance for the ability of the team members to work together effectively to accomplish the team goal. based on previous research, it is well established that when teamwork is characterised by constructs such as trust, psychological safety, cohesion, interdependence, or team efficacy, it tends to be more effective (e.g. raes, boon, kyndt, & dochy, 2015.; edmondson, 1999; jehn, greer, levine & szulanski, 2008; raes, kyndt, decuyper, van den bossche & dochy, 2015). the characteristics described above are in the literature referred to as emergent states, which means that they are seen as ‘…properties of the team that are typically dynamic in nature and vary as a function of team context, inputs, processes and outcomes. team emergent states describe cognitive, motivational and affective states of teams …’ (marks, mathieu, & zaccaro, 2001, p. 357). because emergent states can be considered as both inputs and outcomes of the reality they are part of (jehn, greer, levine, & szulanski, 2008), they play a crucial role in understanding team dynamics and changes in teamwork over time. within the teamwork research tradition, different team development models have been created that describe the occurrence of these emergent states over time (e.g. mcgrath, 1991; tuckman, 1965; wheelan, 2005). team development can be seen as change over time in teams (tuckman, 1965). change in this context is defined as ‘an alteration in the nature of group interaction or performance, in the stage of the group as a whole, or a second-order change in the patterning of group processes’ (arrow et al., 2004, p. 80). alterations can be triggered by influences from inside (e.g. diversity among team members) or outside (e.g. deadline pressure) the team. and alterations are learning activities of the team. these learning activities occur in the form of adaptive, generative or transformational learning efforts from the team members that alter the way they approach collaboration and task work (sessa & london, 2011). the traditional team development models describe the occurrence of emergent states as a fixed process within a prescribed pattern of sequential stages toward a more mature team (raes et al., 2015). based on these models it can be stated that teamwork of teams that are more mature is characterised by the presence of more robust emergent states, such as stronger cohesion and high, more stable, levels of trust between team members (tuckman, 1965; wheelan, 2005). there is evidence for this statement, as the presence of more robust forms of these emergent states in more mature teams is related to the occurrence of several team processes and eventually to team outcomes such as team performance (mathieu, maynard, rapp & gilson, 2008). for example, raes, et al. (2015) found that teams who perceived themselves to be in later phases of team development also report to express more team learning behaviours and that this relationship is mediated by the presence of perceived team psychological safety and group potency. the traditional team development models have been proven useful for team members to better understand what is happening in the group they are functioning in (bonebright, 2010). additionally, when team members are informed about the path of development as outlined in these models, they are bound to develop faster (wheelan, 2009). however, these models conceptualise team development as if all teams experience the same learning curve towards effective functioning. as such, they seem to neglect that emergent states can change unpredictably over the lifetime of a team, caused by the unpredictability of influences from in and outside of the team. different teams can respond differently to the same challenges, translating in different solutions to these challenges, and, consequently, different team paths (sessa & london, 2011). several researchers and practitioners have noted that traditional models are too general and rigid to be applicable to specific teams (poole, 1983; raes et al., 2015; rickard & morger, 2000; tuckman, 1965). for example, the idea of different consecutive stages seems to be problematic in the sense that not all teams follow the different phases in the described order. additionally, empirical confirmation for these models is scarce (kuipers & stoker, 2009). starting from these observations new conceptualisations of group development leave more raes       | f l r     65   and more room for unique team paths and influences from outside of the team (raes et al., 2015). in this study, a more flexible approach towards conceptualisation of team development will be applied to study the changes over time in three temporary project teams. the inspiration for the creation of this approach was found in findings about previous research on team development. in the following paragraph an overview of these findings is given. 1.1 team development in knowledge work teams – state of the art. in their review of team development models raes et al. (2015) selected 15 models that are applicable to knowledge work teams. they categorised this selection of team development models based on their flexibility. flexibility is the degree of predictability by which the team development is conceptualised in the team development model on a dimenson from fixed to random (raes et al., 2015). stated differently, flexibility represents the extent to which team development models take into account the dynamic nature of teamwork. within this review, three facets of flexibility were identified: (1) flexibility of the events described in the model, (2) flexibility of the order of the events and (3) flexibility in co-existence of development paths (raes et al., 2015). an event refers to a phase or critical moment described in the team development models. a phase is a distinct sub period within the lifespan of a team with a specific set of characteristics that shape team interactions during that phase. a critical moment refers to a specific moment of change during the lifespan of a team (raes et al., 2015). because it is not within the scope of this paper to give an exhaustive overview of team development models that apply to knowledge work teams, a selected number of models are used to illustrate the different trends within the literature in the following paragraph (for an exhaustive overview, see raes et al., 2015). additionally, the limitations and advantages of the different presented conceptualisations will be discussed. 1.2 fixed models of team development a collection of traditional team development models, also referred to as linear progressive models (chidambaram & bostrom, 1996), that is applicable to knowledge work teams can be classified as the least flexible type of team development model on three facets. firstly, these models describe a number of fixed events that teams go through. additionally, the order of occurence of these events is determined. lastly, the co-existence of different events in different paths of development (e.g., events on an interpersonal level are inevitable linked with events on the task level) is fixed (raes et al., 2015). the textbook example of these type of models is the linear development model of tuckman (1965) and tuckman and jensen (1977). the model describes the maturion of a team over time in five consecutive phases. the development of the interpersonal path and the path of task behaviours are described parallel to each other and every phase consists of characteristics that are specific for teamwork of a team in that phase. interaction during the forming phase is characterised by orientation. team members discover the boundaries of the social situation by testing behaviours towards and in interaction with each other and use cues from the team leaders’ behaviours as a guidance in this process. at a task level, team members are preoccupied with identification of the task and necessary task behaviours. during the storming phase hostility is the key characteristic of team member interactions. team members express resistance towards each other and the team as a whole in order to protect their individuality. a comparable process occurs in relation to the task: team members oppose against the task as a reaction to the invasion of the task into personal goals. interaction during the norming phase is characterised by acceptance and openess. team members accept personal idiocyncracies of others and create norms that adhere to and respect the differences and similarities between them. task behavior is characterised by an open exchange of information and relevant interpretation related to the task. as a consequence, the stage is set for effective team work. during the performing phase, team members no longer feel the need to establish social relationships because all these issues have been handled during earlier phases. as a consequence, they can approach each other and the task in a pragmatic way. team members are given specific roles that can be employed when and where necessary. this enhances succesfull task raes       | f l r     66   completion. ten years after the publication of the model, tuckman and jensen (1977) reviewed literature on empirical studies that validated the model. based on this review, they added an adjourning stage. during this stage, team members start to alienate from the team and the team task due to approaching termination of the team members working together. 1.3 towards a more flexible approach to study team development more recent but less known models of team development implemented more flexibility in their conceptualisation of team development by means of identifying multiple possible development paths. most of these models were created as a reaction towards the unitary nature of the traditional team development models. researchers implemented different features into their models to allow more flexibility in terms of outlining a development path (raes et al., 2015). morgan, salas, and glickman (1993) created a model that finds its roots in the model of tuckman (1965). however, different features were added to the model to create possible alternative development paths for teams and, as such, to increase the fit of the theoretical model with the actual development path of teams. for example, the different phases are not hypothesised to occur in a strict order, they overlap and can be repeated several times. they also suggest a differential maturation of teamwork and task skills. the model of mcgrath (1991) conceptualises team development by means of three functions or tracks of activities through which teams contribute to the system (production – well-being – member support) and four different modes in which these functions can be found (mode i: inception – mode ii: problem solving – mode iii: conflict resolution – mode iv: execution) . the functions exist independently next to each other and are not hypothesised to go through all the modes in a specific order. these models are more flexible than the linear progressive models on all three dimensions mentioned above (raes et al., 2015). first, the events that teams encounter are not fixed, only suggested as a possible activity that could occur within a function. secondly, the order of events is not fixed and third, the different paths are independent (in the case of mcgrath’s model). however, there are still restrictions in terms of the range of important events that can occur and paths that can determine team development. 1.4 turning points the most flexible conceptualisations of team development consist of a description of critical events during the lifespan of the team based on the judgement of the party that is asked (e.g. team members, team leader or objective external observers) (raes et al., 2015). erbert et al. (2005) introduced one of these conceptualisations, namely turning points. the concept of turning points is borrowed from research in romantic relationships (bolton, 1961). a turning point is operationalised as critical moments in the lifespan of a team that is perceived as a significant change in the course of the collaboration by the team members themselves, both on a task or interpersonal level and both in a negative or positive direction. erbert et al. (2005) identified turning points in teamwork based on interviews with individual team members during which team members were asked to report about the turning points that occurred within the lifespan of their team. additionally, they studied the turning points with more depth in order to get more insight into the concept of change in teams using the dialectical theory. the dialectical theory recognises complexities and contradictions within social interactions and takes this into account when studying the way humans make sense of their everyday experiences. erbert et al. (2005) asked team members to rate the importance of the turning points on six dialectical contradictions, e.g. autonomy vs. dependence. they concluded that most of the identified turning points cluster around one out of seven issues: cohesion, project management, socialization, member change, competence, workload and conflict. additionally, the most important dialectical contradictions for the identified turning points were team vs. individual and competence vs. incompetence. lastly, erbert et al. (2005) asked the team members to score their satisfaction with these turning points and discovered six different paths of satisfaction over time (erbert et al., 2005). however, contrary to other team development models, they did not use this conceptualisation to descriptively outline the development path of the teams they studied over time. raes       | f l r     67   2. present study the main goal of this study is to enhance knowledge about the team development paths of temporary project teams taking into account the dynamic nature of emergent states. the concept of turning points, borrowed from bolton (1961) and erbert et al. (2005), is amplified in order to create a method to outline the development path of teams over time. the conceptualisation of team development in this study is based on the positioning of turning points towards each other over time. this approach towards conceptualisation of team development respects the uniqueness of every individual team because it allows turning points to occur at any moment in time and relative towards the position of other turning points. in further research, this method could be used in order to understand the emergence of team emergent states over time. through further elaboration about the causes and consequences of these turning points more information about the emergence of team emergent states over time will be discovered. the focus is on turning points during the development of temporary project teams. temporary project teams are teams that consist of highly interdependent members that each need to use their own specialised expertise, knowledge and judgement in order to accomplish the team task. these teams are composed with a specific, unique team task for which the team members have a shared responsibility (cohen & bailey, 1997; devine, 2002; tannenbaum, mathieu, salas & cohen, 2012). the teams studied are knowledge work team, because their main focus is on a cognitive task that requires low physical effort (devine, 2002). by means of focussing on one specific type of teams, it is possible to eliminate ‘noise’ in the results that is created by differences between different types of teams (devine, 2002; dochy, gijbels, raes, & kyndt, 2014). as such, enhancing our knowledge about the development of one type of team allows in depth study of specific findings that appear and can give rise to recommendations for this type of team. in this study, attention will be given to (1) the team members’ perceptions about turning points in the lifespan of their team and the position of the turning points towards each other; (2) the characteristics of teamwork during, before and after the identified turning points. using this information, the individual developmental path of three project teams will be recreated. in a next step, a cross-case analysis to look for similarities between the three teams will be conducted. this leads to the following research questions: [1] which turning points can be identified in each of the studied teams as perceived by the team members? a) why are these moments turning points? b) why do they emerge? c) what is their effect on the team? d) where can they be situated in the lifespan of the team? e) on which level (task or interpersonal) are the turning points situated? [2] how are the turning points situated within the life course of the team? [3] can similarities and differences be found in the observed turning points and their positioning over time across the three teams? 3. method 3.1 the teams the subjects of this study are the team members of three student project teams working together on an authentic problem in collaboration with real companies for a period of three months. the projects are part of the project-based course ‘labour pedagogy projects’ organised in the third year of the bachelor program of educational sciences in a large european research-intensive university. more information about the set-up of this course is given in the next paragraph. the teams studied adhere to the characteristics of temporary project teams since they have a one-time assignment that is unique and that requires a new solution. the raes       | f l r     68   team works on an authentic organisational problem within a limited timespan. the skills and knowledge of the each of the individual team members are valuable and useful to accomplish the task. in the following paragraph a short description of the three project teams and their project is given. team a consisted of seven female team members. the authentic assignment they were confronted with was to optimise the welcomeand socialisation procedure for new employees at a college for higher vocational education. team b consisted of seven female team members. the authentic question they were confronted with also came from a college of applied science. they were asked to determine the necessary steps for designing one of the educational programs as more competence based. team c consisted of six female team members and one male team member. the question they were confronted with came from a consultancy company specialised in guidance of change processes. the project team was asked to update a tool for guiding these processes. 3.2 the course ‘labour pedagogy projects’ in the course ‘labour pedagogy projects’ that entails the subject ‘learning and development in professional organisations’ the student teams were confronted with an authentic question presented to them by professional organisations. the project teams worked together for a delineated period of three months at the end of which they had to deliver a tangible product to the professional organisation. the practiceoriented end product(s) had to be accompanied by a collection of relevant scientific literature on the topic of their question. within a well-designed, supportive environment, the main focus was on the autonomy of the team. both on a task and interpersonal level, the team was self-steering and encouraged to create an environment that facilitates their collaboration and output production. the supportive environment was shaped by a number of supporting activities that are part of the standard curriculum of the course. every team was supported by an academic coach and a contact person from the professional organisation. the academic coach provided ad hoc guidance on a task and process level. and the team received the advice and coaching on the task from the professional coach. additionally, more structural activities were set up. on a task level, the team received feedback on the minutes from every team meeting. the team was obliged to visit the professional organisation at the beginning of the project to ensure that the question is clear for both parties. the team had to discuss about the task and progress to a professional coach on a regular basis. at the end of the process, the team organised a closing event with the stakeholders from the professional organisation and the academic coach during which a final presentation of the work and results was given. at the midpoint of the three months, the team was asked to provide an intermediate report on the work progress. additionally, an intermediate evaluation was set up during which the team members were obliged to organise a moment to reflect on their activities both at process level and at task level including the functioning of the group on both tracks. 3.3 data collection data were collected by means of two sources of information: individual reflection papers and team interviews. after the end of the project, the team members of the three student project teams were asked to individually reflect on the development of their team by means of writing a reflection paper. they were asked about turning points that they experienced during the development of their team. afterwards, the team members were gathered for a team interview. these team interviews were conducted by one interviewer and were videotaped for further analysis. the approach to the interview was semi-structured (wengraf, 2001). the interview consisted of two parts. first, the team members were asked to outline the course of their team over time collectively1, both on a task and interpersonal level. after ten minutes of preparations, the graph                                                                                                                           1 the members of the team had to provide an answer as a group. raes       | f l r     69   was presented and the interviewer asked for clarifications. this information was used to position the turning points in time during the second part of the interview. during this second part, team members were asked to pinpoint and elaborate on turning points that occurred during the course of their project. a turning point was defined as a moment during which change occurred. change could imply both progress and regression. team members were asked to indicate the turning points on the graph with blue post-its (see figure 1 – 2 – 3). after ten minutes of preparations, the team members were asked to elaborate on the turning points based on the interview guideline. this two-source data collection was setup to enhance reliability of the study. the information in the reflection papers was used as a control for the information that was given during the interview (golafshani, 2003). first, it was used to see whether similar information was reported right after the collaboration and four months after the collaboration. additionally, it was used to avoid the creation of a fragmented picture due to a lack of safety to speak out between the team members during the team interview. both data sources were compared for substantial differences in information about occurred turning points. as no substantial differences were found the information from both data sources was used for analysis. 3.4 method of analysis data were analysed using the theoretical thematic data analysis technique (braun & clarke, 2006) focussing on themes in the data. in this case the themes are all the facets of information about different aspects of the turning points as described in the research questions (e.g. description, cause, consequences). as such, the coding and analysis of the data were driven by the research questions with the goal to report the reality of the participants’ experiences concerning the reported turning points. a semantic analysis of what has been said and written was conducted (braun & clarke, 2006). to support the use of this technique, a coding procedure was created and executed with the help of the qsr international’s nvivo 10 software. this is a program that facilities the coding of text documents and the additional analysis of the identified codes. in the following paragraph, the coding procedure and the following analysis activities are described. 3.4.1 coding procedure. first, a basic coding schema was set up to systematically collect all the given information about one turning point. because the different teams encountered a different number of turning points, the number of basic coding schemes used was different for each of the three teams. for example, in team a, four turning points were identified. as such, the final coding schema to analyse the turning points of team a used in the nvivo software consisted of four basic coding schemes. basic coding scheme. the codes in this basic scheme were created based on the sub questions under research question 1 of this study. for example, research question one focuses on the different turning points identified by the team members. the code ‘objective description’ was created to code all the information that was given during the interview and in the reflection papers that provides an answer to this question. the codes were fine-tuned by means of different trial versions with the data from team a. for example, the second code ‘subjective description’ was created to code the information that provides an answer to research question 1a, since most of the answers to the question ‘why is this moment a turning point?’ entailed a description of how team members experienced this moment. the different codes and the questions to which they provide an answer are outlined in table 1. raes       | f l r     70   table 1 codes based on research questions code description objective description information about what happened during the turning point subjective description information about how the turning point is experienced by team members consequence consequences of the turning point as perceived by the team members cause what caused the turning point? time situating the turning point in time additionally, the information was also coded based on the topic it was referring to. it was coded ‘task’ when it referred to an aspect that was directly related to the task being performed, it was coded ‘procedure’ when it referred to the way the team members collaborated to perform the task. it was coded ‘interpersonal’ when the code referred to topics that were related to individual team members and their relationships. 3.4.2 abstracting information about the turning points from the coded data within the teams. after coding, a document was generated via the qsr international’s nvivo 10 software that consisted of an overview of all collected text data per code from both the team interviews and the reflection papers. this selection of text data was used to identify themes within these coded data. as such, a thorough answer to the different research questions was formulated based on the collected data. for example, all the coded text material about the cause of one turning point within one team was collected and studied as a whole. based on this collection and codes, an extended description of the different turning points was made. within this text material, an abstraction of the different themes that emerged concerning the causes of the turning points as outlined by the team members was made. this analysis resulted in a narrative description of the different themes that were determined for each individual turning point. 3.4.3 identifying similarities and differences in the turning points across the teams. in the next step, analysis that was done for every code for every individual team as described in the previous paragraph was scanned for recurring themes across the different teams. as such, similarities in the turning points, their causes and consequences across the different teams could be identified. 4. results first, an overview will be given of the information about the turning points that was collected while analysing the coded team interviews and individual reflection papers interview. additionally, the results of the cross case analysis will be given to outline the differences and similarities in the findings over the different investigated teams. raes       | f l r     71   4.1 turning points identified in the three teams for the three teams, a synthesis of the recurring themes in the information about the turning points (description, cause, consequence) that was given by the team members will be presented per turning point. additionally, a visual representation of the chronological position of these turning points within the lifespan of the team will be given. using this set-up, research questions one and two will be addressed for every team individually. it should be noted that – since teams are unique differences in the sort and amount of information that was given occur between the teams. additionally, in this section the turning points are referred to by the name the team members gave them during the interview. team a. figure 1 represents the course of the team a with the turning points positioned over the lifespan of the team. four turning points were identified during the group interview. figure 1. course of development team a turning point 1. the first turning point mentioned by the team members, and nominated as the most important turning point, is the intermediate evaluation. this turning point is situated at the interpersonal level. cause. the team members indicated the fact that they had to reflect on their own and the others performance and the collaboration as a cause for its occurrence. they also appointed great importance to the way they choose to execute this evaluation as an important cause for its occurrence. they feel that they created a safe and stimulating environment by setting up an evaluation method that entailed one-on-one feedback for every team member followed by a group conversation about the functioning of the team. interviewer: ‘could you explain to me why you think this [evaluation] was so effective?’ team member: ‘ i think that, because there was a combination of one-on-one and group, i think that it is really good if one evaluates like that. when there is only in group [evaluation], there are always thinks such as this or that that you would not express towards the person involved and if there is only one-on-one [evaluation], eventually you also have to evaluate in group. i think that both is the best.’ this set-up allowed them to privately discuss what they felt could not be said in front of the whole group. the fact that they expressed both positive and negative points ensured a constructive atmosphere. team member: ‘if it wasn’t for this moment, we wouldn’t have shared negative experiences.’ raes       | f l r     72   consequence. the main consequence is reported on an interpersonal level. an increase in familiarity or interpersonal knowledge between team members (gruenfeld, mannix, williams, & neale, 1996) enhanced feelings of safety. team member 1: ‘it was a bit tricky, all girls together, it could have gotten very bitchy.’ team member 2: ‘yes, but it was not the case, and that was reassuring.’ this allowed the team members to express feedback (positive and negative) towards each other and triggered the presence of more acceptance and understanding of the behaviour of other team members. these shifts had positive repercussions on the process level. team members expressed an increase in the quality of their collaboration: the understanding of the natural division of roles was enhanced due to openness about capacities of the team members. team member: ‘because now i knew other team members thought i was good at this or that, i felt more secure to take up this task, without discussion about it being needed.’ based on this knowledge, some shifts in roles occurred spontaneously. team member: ‘i used to take on the leadership role, but it became clear that other team members were also very good at this, so i let them take the stage more often’ on a task level, their efficiency went up after the turning point. the reason this moment had such significant effect on the teamwork was that it set the stage for openness between team members to discuss frustrations, critiques and accomplishments. turning point 2. the second turning point refers to the moment where one of the team members prepared a presentation in which she presented all her considerations and questions about the state of the art of the project and the future steps in the process towards task achievement. cause. the turning point was caused by chaos at the task level. in the previous meeting, the team members expressed, and agreed about, a shared feeling of lack of clarity about the task, about what the professional organisation expected from them and a lack of information about what other team members were doing. team member: ‘you are working on all kind of things and you are thrown in and after some time you just don’t know [what the next step is] anymore. everybody experienced this feeling, except for [name team member].’ due to this feeling of safety created by the first turning point, the team members were able to express doubts and insecurities concerning the task. this was an important condition for this turning point to emerge. consequence. the consequence was situated at a task level: the team started working more focused. on a process level, team members started collaborating as a team and bundle their resources in a more efficient way (e.g. role division was refined as a team leader was appointed). team member 1: ‘we started working again.’ team member 2: ‘more goal-directed looking for texts, goal-directed reading, not just because i did it that way but yes really for the assignments.’ team member 3: ‘more together again, because before it was kind of individual, everybody did her thing, but now it was very much in the same direction again.’ on an interpersonal level, this moment confirmed the safe climate that was established during the first transition point. raes       | f l r     73   turning point 3. the third transition point indicated by the team members was situated during the period of interviews with employees conducted in the professional organisation to collect information necessary for their task process. cause. the cause of this transition point was threefold. first, by interviewing the employees they came in close contact with the core of the problem they had been working on. the closeness to the authentic work situation made them realise the value of the project and their expertise to work on the problem. secondly, by leading the interviews the team members became more confident about their competences. these two causes triggered the understanding that their project could make a useful difference for the professional organisation. team member: ‘it was the feeling of finally something, yes, that you could lead something yes, that was weird, that you would also think that those people would see you as the one with the expertise, at least a bit.’ the third reason was the strengthening of their interpersonal relations due to getting to know each other in a setting outside the meeting room and the university (cfr. familiarity). consequence. as a consequence of this turning point, team members felt more motivated to work on the project. they felt very proud of themselves as individuals and as a team. team member: ‘i think this period showed us what we were capable of.’ at a task level, the effect was visible trough high productivity during the focus interviews. team member: ‘a drive to continue, as if we can maybe actually mean something useful for this organisation. well, because at the beginning you do think what can we as students change there and then you notice that it might be useful, because those people [employees] also said that it was useful to talk about these things and then you think: aha we can actually contribute something.’ at an interpersonal level, this turning point improved the atmosphere in the team. turning point 4. the last transition point was situated around negative feedback they received from the academic team coach on one of their products. cause. they state that the moment was caused by frustration and stress due to the approaching deadline. however, as the interviewer elaborated on this moment during the team interviews of this study, it became clear that the frustration had a deeper cause: not long before, the coach had encouraged them to take initiative. by providing negative feedback on their work, they experienced a discrepancy in the behaviour of the coach. it is the experience of this discrepancy that disappointed them and made them angry. they felt that the role of the coach had shifted from being a companion in the work process to one of an evaluator of their work. team member: ‘everybody thought: finally we have something, it was something concrete for a change and then it got totally demolished and then yes, it’s a push back, back to reality i guess.’ consequence. the turning point created an ‘us against her’ attitude in the team. at a task level, this attitude became a very strong motivator to prove that they could meet the high standards that were set by the coach. team b – general development. figure 2 represents the course of team b. the turning points identified by the team members are outlined on the timeline. the team members reported seven turning points. they are not discussed in chronological order, because three of the turning points are related and as such described together under the subtitle ‘a cluster of three turning points’ raes       | f l r     74   figure 2. course of development team b turning point 1. the first transition point covers a longer period of time. at the beginning of the project the team members visited different schools to collect information. the number of visits was divided between different sub-teams of two or three members to reduce workload. cause. these sub-teams spent a lot of time together. due to the lengthy personal, one-on-one contact, familiarity between the duo’s and trio’s increased. team member: ‘i think that, because you were alone with somebody for a couple of hours, it caused you to talk about the project and than you would understand better why people think in a specific manner of, yes, what their opinion was on certain things, and that helped also when you were talking to somebody else about the same topic, to tell to them how this person explained it to you.’ consequence. this turning point had a positive effect at the interpersonal level and the process level: team members felt that a base of trust was created which had a positive effect on their collaboration. team member: ‘you get to know each other a little better, you get more insight in how other people think and you, well, it also was a onset to be more open towards each other.’ however, they also report a negative side effect on the task level: sometimes the close friendship distracted them while working on the task. team member 1: ‘yeah, maybe the collaboration went a bit easier or smoother on moments when everything went fine.’ team member 2: ‘yes, but this was also the reason, i think, that it got a bit annoying. because in the library, i know, or when we had to collaborate somewhere, we would just be talking about all kinds of stuff while we were actually supposed to continue working. and, in itself, that was not such a good input for the task.’ turning point 2. the second transition point that is put forward by the team members is a meeting with their coach from the professional organisation. cause. during this meeting they received the necessary information to continue their work and to channel their efforts in the right direction. raes       | f l r     75   team member: ‘before [that moment] we did not really know what to do, we were working without knowing where to go and by talking to these people, a lot became clear for us.’ consequence. as a consequence, at the task level, they could continue working and felt very motivated to do so. team members: ‘this lead us to continue with our work, well yes, we know what we had to do and that could go more fluent for a while.’ cluster of three interrelated turning points 3a – 3b – 3c. from the start of the project, team members experienced frustrations about the other team members. over time, this became stronger and more detrimental to the positive atmosphere in the team. at different points in time this came to the front triggered by specific events. the team members appointed each of these events as turning points. the first turning was appointed as such spontaneously. the second and third turning points were added during the elaboration about the first turning point. for clarity reasons they are here described jointly as three turning points within the same scope. a. during the first weeks of the project, frustrations about how individual team members approached the task work (e.g. the importance team members granted to the task or the amount of time they spent on the task) emerged. cause. these frustrations were not discussed overtly, until one of the team members explicitly expressed hers. this triggered other team members to do the same. team member: ‘we were working for a couple of weeks and there were little frustrations. we did talk about it, but never in the group or to the people we really wanted to say it to or where the message should end up and uhm, at some point, i had the feeling that we really needed to say something about it.’ consequence. at first, team members felt this was a positive event, because it had the potential to set the stage for talking about the issues openly. however, this turning point had a negative longterm outcome in the form of a recurring habit of outing of frustrations at the end of every team meeting. team member: ‘in the beginning, i think it was good that we were able to say it, because it was already present at the surface for a very long time, but in general it had a very negative effect.’ even though the frustrations were always about task or process aspects (e.g. criticising team members on spelling mistakes) some team members perceived them as personal attacks. team member: ‘well yeah, if you keep hearing about the spelling mistakes that you made and the fact that you didn’t re-read the things you wrote, after a while… you start … [expresses frustrations with facial expressions].’ these events had a significant impact on their interpersonal relationship in terms of safety. the team members reported a constant fear of doing something wrong and of the possibility that other team members would call their attention to the mistake. the pressure of the workload strengthened this feeling. in general, these recurring events were associated with negative feelings. b. at a certain point, the team members decided to take measures by implementing a positive attitude towards each other. at the end of every meeting they installed a round of constructive feedback supported through a shared responsibility. for example, when team members expressed a frustration in a negative way, the others reminded them that they had to focus on constructive feedback. raes       | f l r     76   cause. after subsequent meetings of strengthening the habit of expressing negative feedback, some team members were confused about this pattern to such an extent that they felt they could not respond to the repeated attacks in an appropriate way. team member: ‘i think it kept going on until we, during one of the meetings, said that is was important to start focusing on the positive aspects again and on the positive things other people were able to do.’ team member 1: ‘at some point, i said, you guys, we should look back [to what we achieved already]. team member 2: ‘yes, that happened’ team member 3: ‘and then we did that and everybody was very surprised. we were like: yes, actually we can be happy and proud about what we achieved so far.’ consequence. in general, most team members felt that this approach worked, despite a couple of backdrops. but for one of the team members, this perception was not present. c. after the last turning point (cfr. turning point 5) the team members had a very positive attitude. however, the pattern of mutual frustration occurred again very quickly. cause. the attitude of being positive disappeared again due to deadline stress. different aspects of the task did not go as planned. team member: ‘we had to start working again really hard and the frustration just came back.’ consequence. as a consequence, new frustrations arose and were expressed, due to stress and burnout feelings of the team members. for some team members’ trust disappeared due to interpersonal issues. other team members choose to personally opt-out due to the lack of interest or the lack of coping options. team members: ‘yes, it became interpersonal, but also personal in the first place. i mean, it just began because it was too much. you did not see the end of it, you got annoyed from always working, you could not do anything fun, you were always working and in the end you were so annoyed that it is not so easy to let these things [frustrations towards other people] rest. turning point 4. the meeting after the intermediate evaluation is put forward as the fourth turning point of the team. this moment was situated between transition point 3a and 3b (cfr. infra). during their obliged intermediate evaluation the team mainly focused at the task and procedure level and not at the interpersonal level. a little while later, they also set up a moment to focus on this interpersonal level. this moment is the turning point. cause. the academic coach advised the set-up of a new evaluation moment focused on interpersonal relations and functioning of individual team members. the team members also felt that they needed some feedback on how others perceived them and their functioning. consequence. this turning point was perceived as a positive and useful moment. due to the turning point the team members became aware of their work points and this awareness stayed until the end of their project. the primary effect of this turning point was on a (inter)personal level: team members felt more secure. they were more aware of their influence on the team (both positive and negative) and really started tackling their points to improve. team member 1: ‘i think that at that moment it became clear for everybody what their position in the group was or what the others thought about them. more security.’ raes       | f l r     77   team member 2: ‘and you were also more aware of the things you need to work on, or thing others could see as a problem.’ the consequence at a process level was that the team members started to communicate more with each other about decisions and they got back to working (cfr. turning point 3b). however, the effect was not sustainable. due to deadline stress, all the positive effects of this moment disappeared after a while. team member: ‘i have the feeling that we very quickly forgot all our points to improve. over time, we needed to keep our head above water more and more because we needed to finish so many things. because there was so much work, we seemed to have forgotten al those points for improvement.’ turning point 5. the last turning point put forward by the team is situated around the final presentation of their product for the professional organisation. cause. in preparation of this presentation the team members became aware of their competences and on the teams competences. they went to the presentation with an attitude of self-security. team member: ‘there were things that we gave each other feedback at so many levels that we were so convinced about it that we forgot to send it to the coaches for feedback.’ during and after the presentation, the enthusiasm and praise of the different stakeholders (academic coach, professional coach, employers) about there work confirmed this feeling. consequence. this gave the team and the individual team members a serious boost, both at a task and interpersonal level. however, the effect of the boost was of short duration due to the pressure of the deadline that came back very soon. team member: ‘during the train ride after the presentation we were constantly talking about the presentation and about how the employees responded to how we approach the issue and that we really had the feeling that we could make a difference and yes, that gives you a bit of a boost.’ team c – general description. the four turning points of team c are presented chronologically in figure 3. figure 3. course of development team c raes       | f l r     78   turning point 1. the team members referred to the first turning point as the complex moment. team members referred to this as the most important turning point in the lifespan of the team. during the first meetings of the team, a number of team members with a dominant character took the lead. other team members felt threatened by them and started back talking about these issues. at the same time they kept up appearances about the atmosphere in the team during the meetings; by organising informal lunch dates; and by making explicit statements about the importance of being honest to each other for the collaboration, while it was very clear that this was not the case at that point. team member: ‘we were not honest during direct interaction with each other, and on the essential moments during meetings, we were just not honest on a relational level that is. and then, we went to a have lunch together with the idea to do something about this hypocrisy. but actually, that only made it worse. we were really not honest towards each other, drinking chocolate milk together pretending nothing was wrong and assuming it would pas. but that didn’t work. then the crisis came…’ at a certain moment the existing issues were made explicated. the actual turning point is situated during the meeting after this moment, when the team members agreed to openly reflect upon the events. cause. due to open communication team members started to accept the different personalities of different team members. consequence. as a consequence of this turning point, team members became more familiar with each other and learned to understand each other better and what to expect from each other. at a process level, this facilitated collaboration. they agreed that they needed to be much more explicit in terms of their communication and planning. the performance of the task also got a boost due to this event. team members explicitly named this the most important event for their team. team member: ‘it lead to, not only about this situation, but also in general, why aren’t we always straightforward with each other. we should make things more explicit, like what we want to do and just make sure the communication is much more transparent.’ turning point 2. the team members referred to the second turning point as the light point. two team members felt that some parts of the project were very vague. they decided to look at it during an afternoon in the library. they collected all the information and had thorough discussions about it. cause. they figured out the missing links and presented the other team members with their solution to the lack of clarity. team member: ‘the thing was, we did not have three goals as we believed first, but only two main goals and some sub goals and that is the way we started looking at it at that moment. we were talking about it and figuring it out and then we decided to share it with the rest of the group. and actually, what was very vague or difficult suddenly became, well, everybody was like ‘yes… even though we thought that the others could just as well react with confusion to what we were saying. but it turned out that it could just as well have been two other people that did the thinking exercise that we had just done.’ raes       | f l r     79   consequence. due to their effort, the other team members also started to see the whole picture; they all felt that this information made their efforts to accomplish the task more focused and more aligned towards the goal. team member: ‘it felt like we could finally continue’ at an interpersonal level, this strengthened the trust and openness between team members they felt that from that moment on a setback in the task could no longer break the strong interpersonal relationships. turning point 3. the third turning point appointed by the team members is a workshop that was offered to them by the professional organisation. this workshop was set up to reflect on behaviours styles of the team members and to reflect on how the different behaviour styles interrelate between team members. cause. during the workshop, they were confronted with the behaviour styles of the other team members, which differed a lot from theirs, but they were forced to find a fit between different behaviour styles. experiencing this process further clarified why the team had some trouble working together at the beginning of the project and why, after getting to know each other better, they were able to overcome differences and use them as strengths of the team. team member: ‘well, the thing that could be seen as a weakness by one person was also formulated as a strength and then you could also start understanding it as a strength. i experienced that very strongly then. and, well yes, also, we were already collaborating well at that point, but then it kind of got explained why that was. additionally, the academic coach provided a feedback workshop during the same period. the feeling and effect of this workshop was similar to that of the workshop provided by the professional organisation. consequence. as a consequence of the session those individual team members felt more involved and this made it easier for them to speak up in the group. they became closer and even more familiar and the team cohesion became stronger: more knowledge was collected about behaviours of team members, about which behaviour they could expect, how this could be interpreted and what it meant. in this way team members were able to be more thoughtful of what certain behaviours could mean and how they should be understood. as a consequence, roles became clearer, for example team members knew whom to approach when they had a problem. team member: ‘well yes for example, when i thought, well i don’t know how to do this anymore and i am very easily confused. and then i would think: whom can i ask? … someone with a [name of profile] profile, maybe [name of team member]. yes that was kind of what it was like of the workshops.’ turning point 4. the last turning point put forward by the team, is the end of the project. all the references were checked, spelling mistakes, typing mistakes, everything was set in the right colour. team member: ‘it was a great moment, because all the references were checked, all different possible typo’s were corrected and spelling mistakes and then it was finished and i was very happy.’ cause. they finished the project. raes       | f l r     80   team member: ‘it was really finished and that was really like we are very happy that it is finished. and not just finished because the deadline was there, but really finished as it should be finished like it is now.’ consequence. when they handed in their work, they described the feeling of stress fading away. they describe the effect at a task level: the feeling of relief and victory. 4.2 a cross case analysis of team development and turning points in this section, a cross case analysis of the three project teams in terms of the similarities and differences that can be discovered in the turning points and their occurrence over time will be given. 4.2.1. the four turning points that occurred in the three teams. two types of turning points at an interpersonal level and two types of turning points at a task level occurred in the three teams. the first type that was identified in all teams is the turning point that triggered openness about the relational conflicts (team a: turning point 1; team b: turning point 4; team c: turning point 1). the team members of team a and team c explicitly put forward this turning point as the most important event during the lifespan of their team due to its positive (long and short term) effects on familiarity, openness and trust. also the team members of team b mentioned that this was an important moment during their collaboration. the notion of relational conflicts refers to ‘an awareness of interpersonal incompatibilities, which includes affective components such as feeling tension and friction. relationship conflict involves personal issues such as dislike among group members and feelings such as annoyance, frustration and irritation’ (jehn & mannix, 2001, p. 238). the presence of a relational conflict was situated at a covert level, but was reported during the team interviews. the type of turning point discussed here resolved the conflict by openly expressing differences and incompatibilities. the second turning point that was traceable within the development of the three teams at an interpersonal level was the period when team members got to know each other on a more personal level, unrelated to the collaboration itself (team a: turning point 3; team b: turning point 1; team c: turning point 3). this getting to know each other turning point triggered the enhancement of familiarity between team members. it strengthened the interpersonal relationship regardless of the task. the main difference between the first and the second turning point discussed here is that the first one involved open communication about incompatibilities between team members. for example, team members would give each other feedback on strengths and points for improvement of individual contributions to task work; or the different aspects of the collaboration between team members would be thoroughly discussed. in the case of the second turning point, the team members were merely socializing with each other and engaging in casual conversation. the first type of turning point at a task level is the aha – experience (team a: turning point 2; team b: turning point 2 ; team c: turning point 2). the different teams reported a moment during which clarity in terms of the task and the task goal was created by actions of the team members themselves or by an outside stakeholder. this event led to an increase in the performance of the different teams. the last common turning point is labelled achievement (team a: turning point 3; team b: turning point 5; team c: turning point 4) and is characterised by an achievement on the task level, for example positive feedback on work, finishing a product, reaching a goal. their similarities. the short-term effects of these turning points were very similar for the three teams (see table 2). raes       | f l r     81   table 2. overview of the scope of the effects of turning points. task level interpersonal level process level ‘openness about relational conflict’ ‘getting to know each other’ ‘aha – experience’ ‘achievement’ the main observed consequence of the openness about relational conflict moment were enhanced familiarity, trust and mutual understanding of attitudes and behaviours among team members. as a consequence of this event, team members started to understand their own and the other team member’s behaviours and attitudes in the team and this created awareness and understanding of processes and mechanisms that occurred within the collaboration of the team members. this understanding led to a decrease in frustration about differences in approaches to the task and an increase of a safe team climate, openness and trust. as a consequence, team members in the three teams report similar boosts at a process level: facilitation of collaboration and at a task level: focus on the task, productivity, efficiency. however the sustainability of this effect seems to depend on other dynamics at play and the subsequent events during the course of the team. in team b a process conflict about the division of the work and the responsibilities of individual team members for the work (jehn et al. 2008; jehn & rupert, 2001) occurred at the same time (cfr. cluster turning point 3a – 3b – 3c). additionally, the team members expressed the presence of a heavy workload. it seemed that these events eliminated the initial positive effect of the openness about relational conflict turning point. the team members of team b reported that they were able to establish a strong feeling of friendship between team members, but failed to create robust openness and trust between the different team members. the getting to know each other turning point created more familiarity among team members, which also contributed to an enhanced understanding of each other’s behaviour during the collaboration. additionally, the turning point also had repercussions on a mere interpersonal level: namely in terms of enhanced stronger friendship and a basis of trust among team members. the common consequence of the aha-experience is focused on the task level: it created a clear and shared understanding of the task. clearing out the confusion about the common goal facilitated the teams’ focus and goal directedness. due to a clear, shared vision of the mutual goal, the collaboration became more effective and smoother. the achievement turning point mainly enhanced social cohesion between team member and belief of the team members that they are capable of working on the task. their differences. the previous paragraph shows that four types of turning points emerged in the three teams and that these turning points had similar consequences in the three teams. also a lot of difference can be observed concerning these events. firstly, the specific circumstances of the turning points as they occurred within every single team are unique. even though the main scope of the turning point is the same, each one has its own specific setup and details. for example, even though the obliged and planned intermediate evaluation triggered openness about the relational conflicts within two teams (team a and team b), the intermediate evaluation in both teams itself was setup differently. and within team c this turning point occurred in very different circumstances, namely escalation of the interpersonal conflict very early on in their development. secondly, the order of the turning points within the lifespan of the project teams is different (see figure 2, 3 and 4 for an overview of the turning points in chronological order). additionally, the position of a turning point relative to other turning points seemed to have an effect on its impact. for raes       | f l r     82   example, in team b and c an aha-experience was reported very early on in the development of the team. the positive effect that was reported was situated on a task level (clarity concerning the task). in team a this moment occurred later during the lifespan of the team, when openness and trust were already established. the effects of this moment in team a was more extended as it also confirmed the trust climate that was created during the openness about relational conflict moment. even though the same type of turning points seemed to have similar consequences across the three teams, the same type of turning points were triggered by different causes across the different teams. for example, the aha-experience was triggered by an external stakeholder in team b and c, and by a team member (internal stakeholder) in team a. additionally, it is found that some turning points within the development of a team have a double function. the third turning point of team a was both a getting to know each other as well as an achievement turning point. 4.2.2 differences in observed turning points across the teams. the range of turning points in team a and b is supplemented with additional turning points that can not be placed under one of the four types of turning points mentioned above. in team c, the only types of turning points mentioned are the four mentioned above. for both teams, this entailed turning points that had a negative impact on their functioning. 5. conclusion this study was set up to conceptualise team development in three temporary project teams taking into account the dynamic nature of emergent states. the focus of this conceptualisation of team development is on turning points, or significant changes in the lifespan of the team as perceived by the team members, and their positioning towards each other over the lifespan of the team. in the following section a short overview of the findings for the three teams is given. because the interest is in the development of these teams over time, the turning points will be discussed in terms of their effect on team maturity both on an interpersonal and a task level. next, the main take away points of the cross case analysis are given. this overview recaps the main direct conclusions of the study. the team members of team a put forward four turning points as significant events within their development. the first turning point, which consisted of an evaluation of the task work and the team work, enhanced the feeling of being a team between team members and this was seen an important prerequisite for the following turning points and was also strengthened after every other turning point. the second turning point, which entailed a brainstorming session to clarify the future directions on the task path, focused on task work, and mostly strengthened the feeling of competence of the team motivating them towards better achievement. the third turning point, during which the team members had close contact with each other and their target audience, enhances the feeling of being a team and their feelings of competence as a team. the last turning point, which was a negative confrontation with the academic coach, had a strengthening influence on the feeling of being a team and, on after a while also enhanced motivation to deliver a better product. over the whole line, each one of these turning points increased and strengthened the maturity of the team in terms of collaboration and performance. the team members of team b identified five turning points. the first turning point, during which team members got to know each other better by spending time with each other in duo’s, set the stage for the construction of a good team. the second turning point consisted of a meeting with the contact person of the professional organisation to clarify the organisational question and enhanced team competence and performance. a cluster of turning points around interpersonal frustrations emerged. the first two turning points in this cluster, which entailed fostering a positive attitude towards each other, had a short-term good effect on the collaboration of the team members. their long-term effect however, was negative because their intentions to engage in constructive collaboration failed due to high stress levels. one turning point occurred raes       | f l r     83   during which the team members reflected on their collaboration. this had a positive effect on their collaboration. a turning point that was created due to the success of their final presentation strengthened their team cohesion. however, the last turning point of the frustration cluster had a negative effect on the collaboration. as such, the turning points that occurred in the development of this team had an overall positive short-term effect on collaboration and team performance. however, the recurring covert and over frustrations towards each other seemed to have had a detrimental effect on this overall positive effect. even though the team members reported successful and increasing team performance and good interpersonal relationships, the high perceived workload seemed to have had a recurring negative effect on the interpersonal maturity of the team. in team c four turning points were identified. the first one reflected about interpersonal interaction and had a positive effect on the collaboration. during the second turning point, the team members figured out how to approach the task that lay ahead of them. this fostered their collective efficacy. the third turning point entailed reflection on individual characteristics of the team members and how this interplays with the collaboration. this enhanced their trust and facilitated collaboration as a team. lastly, they put forward the moment at which their goals were achieved as a turning point. the turning points that occurred during the lifespan of this team each had an overall increasingly positive effect on the collaboration and team performance over time. by approaching team development through pinpointing important moments of change during the lifespan of a team, a number of conclusions can be made when comparing the teams in the three cases. in this study, three similar teams were investigated and a common ground could be detected in terms of occurring events that shape the development of these teams, for example the occurrence of a relational conflict or a moment of task clarity. as such, four types of turning points were identified that occurred within the three teams: openness about the relational conflict, getting to know each other better, ahamoment and achievement. the consequences of these turning points on the further lifespan of the team also show great similarity across the three teams. however, the triggers for occurrence, the timing of one type of turning point and the order of occurrence of these different turning points are different for the three teams. additionally, the events behind the turning points are also considerably different. for example, in two teams the openness about the relational conflict was triggered by the intermediate evaluation. in the third team it was triggered by an unforeseen circumstance. aside from the common turning points, two of the three teams identified unique turning points that are considered important for the development of the team. the scope and content of these turning points are different and they can thus not be compared to each other. yet all these events had an important impact on the course of the team, as perceived by the team members. 6. discussion in this last section, the findings of this study are elaborated on in more depth: the turning points approach is compared with and contrasted against existing theories. first the similarities and differences with a more popular conceptualisation of team development are discussed. based on this comparison, the strengths of the turning point method will be highlighted and potential opportunities and limitations for the use of this method are outlined. following, the findings in this study in terms of the types of turning points that were found will be discussed in the light of relevant existing team theories. 6.1 turning points: a more flexible way of conceptualising team development. the turning point conceptualisation of team development is compared and contrasted with the linear development model of tuckman (1965). this comparison is made because the tuckman model is traditionally the most used team development model in research and in practice (raes et al., 2015). the raes       | f l r     84   scope of the four types of turning points that could be identified within each of the three teams (e.g., getting acquainted with each other, relational conflict, sudden focus on the task, achieving as a team) can also be detected in different phases of the linear development model of tuckman (1965) (e.g. resp. forming, conflict and performing phase). additionally, two of the three teams (team a and team c) reported a gradually increase in team collaboration and team performance. as such reporting the gradual maturation of teams over time that is outlined in the traditional models. this is not surprising because these models originate from studying the same subject: the process of team development (raes et al., 2015). however, the turning points conceptualisation is suggested here as a more flexible alternative for the linear development model, because in contrast to the former provides an alternative path for development when a team does not pass one of the phases, skips one of the phases, relapse to one of the earlier phases or passes the phases in a different order, whereas the latter does not (bettenhausen & murninghan, 1985). for the teams in the current study, this is specifically problematic in the case of team b. when the development of this team was outlined based on the model of tuckman (1965) the description would stop after the second phase, because the team did not fully pass this storming phase, even though the team members also experienced dynamics that are characteristic for the later phases of development. the development of team a and c could be fitted within the path outlined by the tuckman model, but they also described events and orders of events that not fit the traditional model. this shortcoming of the tuckman model could be explained by stating that it is predictive in nature (bushe & coetze, 2007). as such it prescribes a development path for teams that, when closely followed, results in the highest level of performance possible. however, other researchers approach this model as descriptive when using it to identify the state of development in which a team is situated at a certain moment in time (e.g. miller, 2003; raes et al., 2015; wheelan, 2009). by using it as a descriptive model, these researchers assume that all teams will follow exactly this development path. lack of clarity about the potential use of the tuckman model, and by extension the other existing team development models, encourages misinterpretations and a very rigid view on team development. the strength of the turning point conceptualisation as an alternative for the tuckman model consists in the fact that it is more flexible in terms of descriptive possibilities to outline team development of individual teams. the important events that teams encounter during their lifespan are well captured in the traditional models. however, these models do not allow for the description of other events that play a role that could have an important influence, or do not allow to report about events that occur in a different order then described by the model. teams can encounter every possible event that influences their development and change. this conceptualisation allows other events to play a role in team development than the ones described in a fixed model, including events with a negative outcome. and it also allows events that have an overall influence during the development, such as the cluster of turning points in team b. as such, this approach should thus be considered a descriptive approach. it can be used to describe team development paths taking into account specific dynamics of individual teams, because it allows flexibility on different aspects. this description has potential for both research and practice. using this approach, more in depth knowledge of the dynamics of emergent states over time could be collected by identifying the effect of these turning points on a specific emergent state, such as psychological safety, instead of the explorative approach used in this study. 6.2 contrasting the findings in this study with earlier research on teamwork. teamwork has been studied for many years; as such the findings in this study can be compared with earlier theories studying phenomena that were discovered when studying turning points over the lifespan of temporary project teams. first, as mentioned above erbert et al. (2005) identified several turning points in their study. interestingly, the scope of the turning points identified in their studie and in this study show considerable overlap. most identified turning points clustered around following topics: cohesion, project management, socialisation, member change, competence, workload and conflict. additionally, most turning points were positioned at the dialectical contradictions of team vs. individual and competence vs. incompetence. the raes       | f l r     85   turning points identified in this study can be situated on the described topics, except for project management and member change. the later one is not so relevant for the type of teams studied here due to fixed membership of the studied teams. additionally, the identified turning points were mainly situated on the two dialectical contradictions described above. the turning points at an interpersonal level entailed events during which the team members encountered a tension between the identify of the individual and the identity of the team played an important role: the openness about relational conflicts turning point started with tensions due to diversity between the different team member and created a collective understanding of the influences of the diversity of these individual team members on the collective collaboration. this dialectical contradiction can also be identified within the dynamics of the getting to know each other turning point. it also enhances understanding of the position of individual team members within the whole of the team. the aha-experience and achievement turning point are situated at the competence vs. incompetence contradiction for the obvious research that they increased a feeling of competence among team members, and, in some cases, started from a feeling of incompetence (cfr. turning point 2, team a). also in the case of the turning points that were only observed in one of the three teams at least one of the contradictions can be observed. turning point 5 in team a could be situated at both, as this turning point was about (1) the team against the coach and (2) feeling incompetent due to negative feedback from the coach. the cluster of turning points in team c is clearly situated at the individual vs. team contradiction due to the interpersonal frustration that repeatedly arose. in two of the three teams, the team members spontaneously2 put forward the ‘openness about the relational conflict’turning point as the most important event within the team development. additionally, the effect of this turning point was particularly important for fostering a positive interpersonal context for constructive collaboration. the third team also recognised its importance and short-term positive effect on collaboration. due to the importance team members ascribed to this event and its facilitating effect, we decided to confront the findings of our study with earlier research on conflict and conflict management. this finding is not completely in line with previous research and theory about relational conflict in teamwork. jehn and mannix (2001) state that the occurrence of relational conflict is always undesirable in a team context in terms of performance. in their questionnaire study, they confirmed the hypothesis that highperforming teams experience lower levels of relational conflict compared to low-performing teams. however, based on the findings of the current study, it seems that the occurrence of a relational conflict, defined as team members’ awareness of interpersonal incompatibilities accompanied by an emotion of friction or frustration (jehn & mannix, 2001), is a natural part of the process of the examined teams. moreover, the uncovering of this conflict has led to an increase of social learning, defined as ‘the process though which team members get to know each other better as individuals and learn to interpret each other’s behaviour in the context of personal life and personality’ by jehn and ruppert (2008, p. 128) within each of the three teams. when social learning is present and it enhances familiarity and understanding of each other’s motives for behaviour it can create empathy and facilitates relational interaction with other team members (huckman, staats, & upton, 2009; jehn & rupert, 2008). this has positive influences on the team work because it facilitates collaboration which in turn leads to higher team effectiveness and efficiency (for an overview of beneficial effects, see jehn & rupert, 2008). the team members also reported a positive effect of this turning point on team performance. it seems that sometime at the beginning of the project, the openness about relational conflict is to some extent necessary in order to facilitate future collaboration. this mechanism could be identified thanks to the qualitative nature of the currend study and the focus on studying a longer period of time. de dreu and van vianen (2001) found in their study that avoidance of relational conflict is most beneficial for team functioning and team performance. the current study suggests that it is necessary to uncover the relational conflict in order to make task conflicts effective. the presence of trust and openness fostered by the turning point enhances the quality of constructive task conflict (van den bossche et al.,                                                                                                                           2 this was not a question during the team interview and the interviewer did not ask team members to elaborate on this topic. raes       | f l r     86   2006). overlooking this mechanisms, could also explain why de dreu and weingart, 2003 found that both task and relational conflict are negative for team functioning. they found that a small amount of conflict is good for team performance; however, an increase in intensity of the conflict is detrimental to the functioning of the team. it could be that whenever a conflict, no matter which type of conflict, becomes too intense, a relational aspect of awareness of perceived differences is experienced. however, when relational frictions are solved, task conflict can occur without the emergence of intense negative emotions. in line with this, the statement that a relational conflict is unrelated to the task (de dreu & weingart, 2003; jehn & mannix, 2001) can be refuted for the studied teams: it requires collaboration on the task and task processes for team members to become aware of incompatibilities and for irritations and frictions to emerge. clearing out emotions by overtly discussing relational conflicts seems to have beneficial effects. however, occurrence of this confrontation does not necessarily mean that the collaboration is safeguarded (cfr. team b). the question remains what triggers a good outcome of this event and what triggers a bad outcome both on the long and short term. the three teams reported a stronger influence of interpersonal processes on task processes than the other way round. dysfunctional interpersonal relations had a detrimental effect on the task process. however, when they were going good, this had a positive effect on the task process. the task did not have this kind of effect on the interpersonal relations. this emphasises the importance of the quality of these interpersonal processes. when interpersonal matters are openly discussed, this openness can also prevent escalation of a relational conflict. more research is necessary to tap into the specific timing and characteristics of this turning point. 6.3 context of this study and future research. the subjects of this study were functioning within the context of a course that is part of the curriculum of an educational program. one reason to set up this type of education is to familiarize students with working in teams. within the context of the course, a number of interventions are organised to enhance collaboration and facilitate the task process. this context can be considered as a limitation of the study, because some of the interventions that are part of the course set-up are closely linked to some of the identified turning points. for example, the visit to the client organisation is a mandatory part of the program for the teams. in one of the teams this triggered the aha-experience. another example is the openness of the relational conflicts turning point that was closely related to the intermediate evaluation with two of the three teams. the closing event triggered the achievement turning point in team b. these observations raise the question to what extent the observed turning points occurred as a natural aspect of the development of the teams. given the scope of this study, it is not appropriate to make statements about (the absence of) causal effects of interventions. however, it is a valid observation that these turning points occurred in the three teams with different triggers, in some cases related, and in some cases unrelated, to the organised interventions in context of the course. further research is necessary to untangle the complexity of the dynamics of different influencing factors of the turning points. the method developed in this study to identify changes in the emergent states over time can be used with different purposes in future research. for example, it can be used to specifically focus on the fluctuations of one emergent state over time to understand its emergence more in-depth. additionally, changes in other aspects of teamwork can be measured before and after the occurrence of a turning point. or interventions can be tested to trigger turning points and as such foster collaboration. this can contribute to the entanglement of the complexity of teamwork. raes       | f l r     87   keypoints individual teams report unique team development paths over time, which is not in line with the traditional team development models. the innovative flexible approach towards team development allows unique paths of individual teams and provides opportunities to study generalizability and common grounds. openness about relational conflict is the most important turning point within the development of student project teams. references arrow, h., poole, m. s., henry, k. b., wheelan, s., & moreland, r. (2004). time, change, and development: the temporal perspective on groups. small group research, 35, 37-105. doi:10.1177/1046496403259757 bettenhausen, k., & murnighan, j. k. (1985). the emergence of norms in competitive decision-making groups. administrative science quarterly, 30, 350372. bolton, c. d. (1961). mate selection as the development of a relationship. marriage and familiy living, 23, 234-240. bonebright, d. a., (2010). 40 years of storming: a historical review of tuckman’s model of small group development. human resource development international, 13, 111-120. doi:10.1080/13678861003589099 boon, a., raes, e., kyndt, e., & dochy, f. (2013). team learning beliefs and behaviours in response teams. european journal of training and development, 37, 357–379. doi:10.1108/03090591311319771 braun, v., & clarke, v. (2006). using thematic analysis in psychology. qualitative research in psychology, 3, 77-101. doi:10.1191/1478088706qp063oa. bushe, g. r., & coetzer, g. h. (2007). group development and team effectiveness. using cognitive representations to measure group development and predict task performance and group viability. journal of applied behaviour science, 43, 184-212. doi:10.1177/0021886306298892 cohen, s. g., & bailey, d. e. (1997). what makes teams work: group effectiveness research from the shop floor to the executive suite. journal of management, 23, 239-290. doi:10.1177/014920639702300303 chidambaram, l., & bostrom, r. p. (1996). group development (i): a review and synthesis of development models. group decision and negotiation, 6, 159-187. doi:10.1023/a:1008603328241 erbert, l. a., mearns, g. m., & dena, s. (2005). perceptions of turning points and dialectical interpretations in organizational team development. small group research, 36, 21-58. doi:10.1177/1046496404266774 de dreu, c. k. w., & van vianen, a. e. m. (2001). managing relational conflicts and effectiveness of organizational teams. journal of organizational behavior, 22. doi:10.1002/job.71 de dreu, c. k. w., & weingart, l. r. (2003). task versus relationship conflict, team performance, and team member satisfaction: a meta-analysis. journal of applied psychology, 88. doi:10.1037/00219010.88.4.741 devine, d. j. (2002). a review and integration of classification systems relevant to teams in organizations. group dynamics: theory, research, and practice, 6, 291-310. doi:10.1037//1089-2699.6.4.291 funk, c. a., & kulik, b. w. (2012). happily ever after: toward a theory of late stage group performance. group & organization management, 37, 36-66. doi:10.1177/1059601111426008 dochy, f., gijbels, d., raes, e., & kyndt, e. (2014). team learning in education and professional organisations. in billett s., harteis c., gruber h. (eds.), international handbook of research in professional and practice-based learning. pp: 987-1020 . the netherlands: springer. golafshani, n. (2003). understanding reliability and validity in qualitative research. the qualitative report, 8, 597-607. retrieved from: http://www.nova.edu/ssss/qr/qr84/golafshani.pdf raes       | f l r     88   gruenfeld, d. h., mannix, e. a., williams, k. y., & neale, m. a. (1996). group composition and decision making: how member familiarity and information distribution affect process and performance. organizational behavior and human decision processes, 67, 1 – 15. jehn, k. a., greer, l., levine, s., & szulanski, g. (2008). the effects of conflict types, dimensions, and emergent states on group outcomes. group decision negotiation, 17, 465-495. doi:10.1007/s10726008-9107-0 jehn, k. a., & mannix, e. a. (2001). the dynamic nature of conflict: a longitudinal study of intragroup conflict and group performance. academy of management journal, 44, 238-251. doi:10.2307/3069453 jehn, k. a. & rupert, j. (2008), group fautlines and team learning: how to benefit from different perspectives. in v.i. sessa & m. london (eds.), work group learning: understanding, improving & assessing how groups learn in organizations (pp. 19-148). mahaw, new jersey: lawrence erlbaum associates. kuipers, b. s., & stoker, j. i. (2009). development and performance of self-managing work teams: a theoretical and empirical examination. the international journal of human resource management, 20, 399-419. doi:10.1080/09585190802670797 mathieu, j., maynard, m. t., rapp, t., & gilson, l. (2008). team effectiveness 1997-2007: a review of recent advancements and a glimpse into the future. journal of management, 34, 410-476. doi:10.1177/0149206308316061 marks, m. a., mathieu, j. e., & zaccaro, s. j. (2001). a temporally based framework and taxonomy of team processes. academy of management review, 26, 356-376. doi:10.5465/amr.2001.4845785 mcgrath, j. e. (1991). time, interaction and performance (tip): a theory of groups. small group research, 22, 147-174. doi:10.1177/1046496491222001 miller, d. l. (2003). the stages of group development: a retrospective study of dynamic team processes. canadian journal of administrative sciences, 20, 121-134. doi:10.1111/j.1936-4490.2003.tb00698.x morgan, b. b., salas, e., & glickman, a. s. (1993). an analysis of team evolution and maturation. the journal of general psychology, 130, 277-291. nvivo qualitative data analysis software; qsr international pty ltd. version 10, 2014. poole, m. s. (1983). decision development in small groups, iii: a multiple sequence model of group decision development. communication monographs, 50, 321-341. raes, e., boon, a., kyndt, e., & dochy, f., (2015). team’s anatomy. exploring change in teams (unpublished doctoral dissertation). university of leuven, leuven. raes, e., kyndt, e., decuyper, s., van den bossche, p., & dochy, f. (2015). an exploratory study of group development and team learning. human resource development quarterly. doi:10.1002/hrdq.21201 rickards, t. & moger, s. (2000). creative leadership processes in project team development: an altenative to tuckman's stage model. british journal of management, 11, 273-283. doi:10.1111/1467-8551.00173 salas, e., sims, d. e., & burke, c. s. (2005). is there a “big five” in teamwork? small group research, 36, 555-599. doi:10.1177/1046496405277134 sessa, v. i., london, m., pingor, c., gullu, b., & patel, j. (2011). adaptive, generative, and transformative learning in project teams. team performance management, 17, 146-167. doi:10.1108/13527591111143691 sweet, m., & michaelsen, l. (2007). how group dynamics research can inform the theory and practice of postsecondary small group learning. educational psychology review, 19, 31-47. doi:10.1007/s10648006-9035-y tannenbaum, s. i., mathieu, j. e., salas, e., & cohen, d. (2012). teas are changing: are research and practice evolving fast enough? industrial and organizational psychology, 5, 2-24. doi:10.1111/j.17549434.2011.01396.x tuckman, b. w. (1965). development sequence in small groups. psychological bulletin, 36, 384-399. tuckman, b. w., & jensen, m. a. c. (1977). stages in small group development revisited. group and organizational studies, 2, 419-427. wheelan, s. a. (2005). group processes. a developmental perspective (2nd ed.). boston: pearson education, inc. raes       | f l r     89   wheelan, s.a. (2009). group size, group development, and group productivity. small group research, 40, 247-262, doi:10.1177/1046496408328703 wengraf, t. (2001). qualitative research interviewing. sage publications ltd.: londen microsoft word kospentaris et al_publication.docx           frontline  learning  research  vol.4  no.  1  (2016)  40-­‐57   issn  2295-­‐3159       visual and analytic strategies in geometry george kospentarisa, stella vosniadoub, smaragda kazic, emilian thanoud anational and kapodistrian university of athens, greece bthe flinders university of south australia, australia and national and kapodistrian university of athens, greece cpanteion university of social and political sciences, greece dnational and kapodistrian university of athens, greece article received 13 november / revised 28 december / accepted 19 january / available online 3 march   abstract we argue that there is an increasing reliance on analytic strategies compared to visuospatial strategies, which is related to geometry expertise and not on individual differences in cognitive style. a visual/analytic strategy test (vast) was developed to investigate the use of visuo-spatial and analytic strategies in geometry in 30 mathematics teachers and 134 11th grade students. students’ performance in the vast was also compared to performance in tests of visuo-spatial abilities, of abstract reasoning, and of geometrical knowledge. the results showed high performance of all the participants in the vast items that could be solved by relying on visuo-spatial strategies. however, only the math teachers showed high performance in the vast items that required the application of analytic geometrical strategies. there were high correlations between the students’ performance in the tests of visuo-spatial and abstract reasoning abilities and the vast analytic strategies scale, but the contribution of these tests to the vast analytic performance became statistically insignificant when geometrical knowledge was used as a mediating factor. the implications of this work for the learning and assessment of geometrical knowledge are discussed. keywords: geometry learning and instruction; visual-spatial reasoning; analytic strategies; assessment of geometry corresponding author: george kospentaris, anational and kapodistrian university of athens, dimitsanas str. 20, 11522, athens, greece, email: 25aris@math.uoa.gr doi: http://dx.doi.org/10.14786/flr.v4i1.226 kospentaris  et  al   1. introduction in recent years research has accumulated showing that spatial thinking is central to success in science, technology, engineering, and mathematics, the so-called stem disciplines. spatial thinking is thinking about the location of objects and their relations and requires both visuo-spatial ability – the ability to mentally visualize the rotation of objects – and spatial abstract reasoning – the ability to identify analogical relations amongst patterns (see newcombe, 2010; wai, lubinski, & benbow, 2009). there is convincing evidence that there are important individual differences in spatial thinking and that spatial thinking abilities can predict success in stem disciplines (hegarty & waller, 2006; wai, et al., 2009). particularly impressive are the analyses of large data sets showing that people with high scores on tests of spatial thinking in high school are more interested in science and math, are more likely to get advanced degrees in stem, and are more likely to pursue stem careers (shea, lubinski, & benbow, 2001; wai et al., 2009). this has led to an increase in training studies that aim at improving spatial thinking as a means of improving performance in stem disciplines (sanchez, 2012; uttal et al., 2013). there is little doubt that much of the problem solving done in science, mathematics and engineering requires the use of spatial thinking (kozhevnikov, motes, & hegarty, 2007; stieff, 2007, zazkis, dubinsky, & dautermann, 1996). in euclidean geometry, where figures are the main objects of study, the role of spatial thinking is of utmost importance. visualizing the shapes and their relation is a standard prerequisite for the understanding of geometrical propositions (battista, 2007). although a great deal of this spatial thinking can be achieved using visuo-spatial strategies – i.e., strategies that allow individuals to obtain spatial information from immediate perceptual processes – spatial thinking can also be achieved using analytic strategies, where rules provide access to spatial information without recourse to visual perception and mental animation (stieff, 2007). zazkis et al. (1996) showed that the majority of the university students participating in abstract algebra courses used a combination of visuo-spatial and analytic strategies (see also schwartz & black, 1996; stieff, 2007). the use of visuo-spatial vs. analytic strategies has been predominately examined from an individual differences point of view, as an individual characteristic or a cognitive style (eisenberg & dreyfus, 1991; pitta-pantazi & christou, 2009). this emphasis on individual differences has obscured the fact that reliance on analytic strategies also characterizes the acquisition of expertise in many domains of stem. as expertise is acquired, problem solving increasingly relies on specialized, domain-specific, rule-based, analytic approaches compared to visual, perceptual information and mental rotation. for example, in the domain of organic chemistry expert chemists develop a predilection for analytic strategies to solve chemistry tasks (stieff, 2007). in geometry, reliance on visuo-spatial strategies seems to coincide with the level 1 of the van hiele (1986) theory in geometrical thinking. at later levels geometrical thinking increasingly requires an understanding of the logical systems that geometry represents. in geometry, shapes are represented by a set of properties and their relations and geometrical thinking is characterized by the formal manipulation of a logical system. thus, when geometrical expertise is achieved, geometrical thinking relies increasingly on analytical formal processes based on geometrical knowledge. the purpose of the present research is to develop a task that can differentiate visual from analytic reasoning in geometry -a visual/analytic strategy task (vast) -and to validate it by comparing novices to experts in geometry. in the next section we present a summary review of the literature on geometrical thinking and define and explain our theoretical position with respect to the use of visuo-spatial and analytic strategies. 1.1 geometrical thinking piaget and his collaborators (piaget, inhelder, & szeminska, 1948/1960; piaget & inhelder, 1948/1967) were the first to study the psychological foundations of geometrical thinking and to propose that it develops in four sequential and hierarchical stages1. in subsequent years, van hiele (1986) argued that                                                                                                                           1 there is a great deal of research on spatial development in young children and different theoretical approaches have appeared after piaget’s seminal work (see newcombe & huttenlocher, 2000; spelke & kinzler, 2007) but this is not the focus of the present paper. 41 kospentaris  et  al           42   there are five, qualitatively distinct, hierarchical levels of thought in geometry. in contrast to piaget, van hiele strongly emphasized the crucial role of school instruction in the acquisition of geometrical knowledge (van hiele, 1986, p. 65-66). more recently, houdement and kuzniak (2003) proposed (on theoretical grounds) that the five van hiele levels can be reduced to three kuhnian-like paradigms: geometry 1-natural geometry; geometry iinatural axiomatic geometry, and geometry iii – formalist axiomatic geometry. empirical research so far has failed to confirm the predictions of the van hiele theory that students move through discrete levels of geometrical thought, each characterized by different internal conceptual organization (battista, 2007). it appears instead that students oscillate between different levels of geometric understanding depending on the context and the nature of the problems to be solved. for this reason some researchers have argued that although there might be different levels of geometric thinking as identified by van hiele, these do not represent distinct stages but develop in parallel and without discontinuities between them (clements & battista, 2001; lehrer, jenkins, & osana, 1998). it follows from the above that we need a theoretical framework that can account for the considerable conceptual re-organizations that take place in the process of acquiring and using geometrical knowledge without posing the existence of hierarchical and well-defined distinct stages. for these reason, it is proposed here that it might be fruitful to examine geometrical thinking from a conceptual change point of view, and that the framework theory (ft) approach to conceptual change (vosniadou, 2013; vosniadou & skopeliti, 2014) can serve as an anchor for examining changes in geometrical knowledge after exposure to instruction. the ft belongs to a class of conceptual change approaches known as ‘theory-theory’ (carey, 2009), but also differs from them in important ways. briefly, the ft claims that (a) there are systems of core cognition that bootstrap cognitive development (carey, 2009; spelke & kinzler, 2007), without making strong nativist interpretations of early infant competencies2, and (b) that conceptual development consists of episodes of qualitative change, which, however, are not discontinuous or stage like. rather, conceptual change is seen as a slow and gradual learning process greatly facilitated by sociocultural and educational inputs. according to the ft, the same constructive-type mechanisms that are involved in all learning processes are also involved in conceptual change processes often producing fragmentation and misconceptions, but eventually having the potential to lead to qualitatively different conceptual organizations (vosniadou & skopeliti, 2014). finally, the ft claims that initial systems of thought continue to exist and influence thinking, even after instruction-induced conceptual changes have occurred (shtulman & varcarcel, 2012; vosniadou et al., 2015). seen from this theoretical perspective, it is argued that geometrical knowledge is originally built on two core cognitive systems (spatial and numerical) that rely on visuo-spatial information (newcombe & frick, 2010; spelke, lee, & izard, 2010), but that it gradually develops through systematic instruction to rely on more analytic strategies based on formal geometrical knowledge. in other words, we claim that there is a growing reliance on analytic strategies in geometric thinking with the acquisition of expertise, and that the systematic use of analytic strategies is a product of conceptual changes that take place in the subject-matter area of geometry. we do not claim that visuo-spatial strategies become extinct and that experts rely on analytic strategies only. unlike stage theories, we argue that the initial, visuo-spatial, approach to geometry is not supplanted by the analytic one, but continues to exist and to be used when contextually appropriate. the ability to systematically employ analytic strategies in geometry, however, is a major intellectual achievement and not a matter of individual differences in cognitive style. it is the product of a conceptual change which takes place over many years and which requires the acquisition of new concepts and new forms of geometrical thinking. we believe that many geometry education researchers would agree with this account of the development of geometrical knowledge. geometry is undeniably a formal system and geometric reasoning                                                                                                                                                                                                                                                                                                                                                                                                                     2 various theories are attempting to explain early spatial development including connectionist interpretations and neoconstructivist approaches (see newcombe, uttal, & sauer, in press, for an extensive review). kospentaris  et  al           43   consists of using this formal system to reason about shape and space. according to battista (2007), underlying this formal system is a ‘primitive’ system of visuo-spatial thinking allowing individuals to ‘see’, inspect, and reflect on spatial objects, images, relationships and transformations (p. 843). this ‘primitive’ system is characterized by what he calls ‘perceptual objects’ – i.e., mental entities perceived by an individual when viewing physical objects in the real world, including geometrical diagrams. in contrast, expert geometrical knowledge operates on ‘conceptual objects’ – i.e., abstract, completely idealized and general mental entities based on ‘formal’ categories, which are explicitly circumscribed according to verbally stated, property-based definitions. the difference between a geometric diagram and a figure captures this basic dichotomy: the former is a material entity, a concrete case that imperfectly represents the abstract concept, while the latter is a theoretical, ideal object without any physical properties. similarly, fischbein (1993) argues that experts in geometry form and reason with ‘figural concepts’. a figural concept is controlled by logical rules in the context of an axiomatic system but is also a mental entity, an image with a spatial-figural content, although devoid of any concrete sensorial properties (fischbein, 1993, p. 148). battista’s (2007) and fischbein’s (1993) arguments are consistent with the proposal that there are ontological and representational shifts that take place in the development of geometrical knowledge analogous in some respects to the ontological shifts that take place in learning science (chi, 2008; vosniadou, 2013). for example, a circle, this quite familiar shape, changes from a visual gestalt (figure 1a) and becomes the locus of all plane points characterized with the property that they are equidistant from its center (figure 1b), or, in the conceptual frame of analytic geometry, to an equation (the plane points satisfying x2+y2 =r2, figure 1c). in addition, the theoretical explanations in the domain also change. at the beginning, geometrical propositions are mainly inductive generalizations based on empirical observations and experimentation with perceptual objects and not on proofs and deductive procedures based on accepted axioms and previously proven propositions. figure 1. changes in the representation of the circle. it could be argued that the above arguments would also be acceptable by stage theories, such as the van hiele theory. if this is the case, then what can the ft offer in our theoretical understanding of geometrical expertise? although some stage theories allow for intra-individual differences across tasks (a phenomenon known as decalage –piaget & inhelder, 1948/1967), they nevertheless assume that a) the new forms of thinking that develop in geometry gradually transform and eventually replace the ‘primitive’ visuospatial system with a more advanced system of thought based on analytical, formal knowledge, and b) that this process leads to distinct, qualitatively different stages in students’ thinking. from the perspective of the ft, however, knowledge acquisition does not proceed through hierarchical and well defined distinct stages, but through the gradual assimilation of the new information into the initial, ‘primitive’ system, creating in the process inert knowledge, fragmentation, and misconceptions, many of which are ‘synthetic’ conceptions3                                                                                                                           3 synthetic conceptions are formed when learners assimilate scientific information to their incompatible prior knowledge producing in the process an alternative, erroneous conception, which however has some internal consistency and explanatory value, such as the ‘impetus misconception’ in mechanics (clement, 1982), the ‘molecules in matter’ model in the atomic-molecular theory (wiser & smith, 2013), and the ‘hollow sphere’ model in observational kospentaris  et  al           44   (vosniadou & skopeliti, 2014). although there is an order of acquisition in this conceptual development and some learning progressions can be identified, these cannot be characterized as ‘stages’, both because they cannot be clearly identified as such, and because the ‘primitive’ visuo-spatially-based system – is not eradicated but continues to co-exist with the formal, analytical modes of thought (shtulman & valcarcel, 2012). the purpose of the present research is not to test a full-blown theory of geometrical thinking, but to start in this direction by developing a valid task that can help us distinguish the use of visuo-spatial from analytic strategies in geometry. in the next section the rationale behind the development of the visual/analytic strategy task is described. 1.2 the visual/analytic strategy task (vast). several tasks have been developed over the years to test students’ movement from the visual to the descriptive/analytic van hiele level (e.g., burger & shaughnessy, 1986; gutiérrez & jaime, 1998; lynn & lynch, 2010; usiskin, 1982). the main limitation of such tasks is that they either favour the recall or recognition of definitions of shapes and their properties over their understanding and their application in novel situations, or that they do not require thinking based on more sophisticated, relational properties (battista (2007). in addition to the above, there are several other standardized geometry tests that assess students’ level of geometrical knowledge, such as the california standards geometry test (csgt, 2009). these standardized tests examine mainly the extent to which students can perform school procedures, e.g., to apply a known formula for some computation within a narrow formal context, which often imposes a particular solution method. thus it remains unclear whether the students who succeed in these tests would present the same level of geometry knowledge in situations where the test format would not be similar to the way they have been taught. in the present research a different method to measure visual and analytic reasoning was developed, based on the following considerations: first, we avoided setting our task in the typical geometry textbook style that could suggest deductive requirements and delimit visualization or measuring. second, we did not impose a particular solution method to the solver but rather selected problems that could be solved using either visuo-spatial or analytic strategies, so that we could investigate spontaneous strategy choice. third, we wanted to investigate not only whether students are able to use analytic strategies but also whether they are able to do so in situations that require them to inhibit visual-perceptual information processing and reason instead along formal geometric lines. thus, a task was needed in which the perceptual difficulty of comparing shapes would be intensive and where the use of analytic strategies would lead to conclusions sometimes conflicting with visual-perceptual information. the above theoretical considerations led to the development of the present visual/analytic strategy task (vast). the vast is a verbal/picture verification task. the participants are presented with a geometrical configuration that includes two shapes and are asked to decide whether a verbal statement that states that these shapes are congruent (congruence domain), similar (similarity domain), or occupy the same area (area domain), is true or false. as shown in figure 2, there are four types of configuration conditions in each geometrical domain: (a) the ‘appearance+/reality+’ condition where the two shapes both appear to be and are indeed congruent, similar or area equivalent; (b) the ‘appearance-/reality-’ condition where the two shapes neither appear nor are congruent, similar or area equivalent; (c) the ‘appearance+/reality-’ condition where the two shapes appear to be but are not congruent, similar or area equivalent; and (d) the ‘appearance-/reality+’ condition where the two shapes do not appear to be but are congruent, similar or                                                                                                                                                                                                                                                                                                                                                                                                                     astronomy (vosniadou, 2013; vosniadou & brewer 1994). in geometry, the ‘figural object’ described by fischbein (1993) to be formed from the synthesis of the ‘perceptual’ and ‘conceptual’ objects described by batista (2007) is such a hybrid, synthetic conception. kospentaris  et  al           45   area equivalent. on the top of each configuration there is a verbal statement, such as, for instance, ‘the lengths of the routes are the same’ (see figure 2, upper row). the participants are asked to decide whether this statement is true or false with respect to the geometrical configuration to which it refers. in all three geometrical domains, the conditions (a) and (b) involve items purposely designed to be solved by visual estimation alone and which are consistent with the adoption of either a visual or an analytic geometric strategy (thereafter the vast consistent subscale, or vast-cons). the conditions (c) and (d) are inconsistent with reliance on visual estimation alone and require for their correct solution reliance on geometrical knowledge and the adoption of analytic strategies (thereafter the vast-incons). the geometrical knowledge required involves either measurement and empirical confirmation or euclidean deductive argumentation. more specifically, in the case of the geometrical configurations a1 and a2 in figure 2, the conclusions that the two routes are equal (in a1) and unequal (in a2) can be reliably achieved through visualspatial inspection. they can also be deduced on the basis of known geometrical properties: in a1, the conclusion of equality can be deduced from the congruence of the corresponding line segments, which are the opposite sides of the formed rectangles. in a2, the conclusion that nick’s path is shorter than john’s path can be deduced from the geometrical axiom of triangle inequality – that the hypotenuse is always shorter than the sum of the right angled segments. in the case of the geometrical configuration a3 and a4, however, the conclusions cannot be deduced by using visual strategies, but only through reliance on geometrical knowledge. in a3, in order to deduce the inequality of length line segments, one has to compute the hypotenuses of the formed right triangles and compare the oblique line segments with the vertical or horizontal ones. in a4 by drawing horizontal and vertical lines, the segments forming the zigzag “john’s route” are equal to the corresponding segments forming the direct “nick’s route”, as opposite sides of rectangles. the above rationale applies to all items of the test. figure 2. sample items from the visual/analytic shift test (vast). kospentaris  et  al           46   1.3 questions and hypotheses of the present study our purpose in the present study was to examine if the vast is a reliable and valid test of visual and analytic reasoning in geometry. with respect to reliability, we wanted to find out whether kuder-richardson (k-r) reliability index was acceptable across the different sub-scales (hypothesis 1). with respect to validity, and in view of our theoretical position that the use of visual and analytic strategies is related to geometry expertise, we wanted to find out if the vast would be able to differentiate the performance of experts in geometry from that of novices. for this reason, the vast was administered to a group of mathematics teachers with extensive experience in teaching geometry and to a group of 11th grade students who had been exposed to euclidean geometry teaching. we hypothesized that if the vast is a good test of the use of visual and analytic strategies, then the mathematics teachers should have high scores both in the vast-cons and in the vast-incons because of their expertise in geometry. on the contrary, the students would obtain high scores only in the vast-cons, which can be solved with visuo-spatial strategies and not in the vast-incons, which requires reliance on analytic strategies based on geometrical knowledge and the inhibition of visual estimation (hypothesis 2). the above hypotheses are different from what would be expected assuming that performance in the vast is related only to individual differences in strategy use as opposed to geometry expertise. individual differences in strategy use would not predict systematic differences in the performance of the high school students. rather, some participants should do better in the visuo-spatial items and some others in the analytic items, regardless of their geometry knowledge (hypothesis 3). hypothesis 4 concerned the relation between the vast-incons and geometrical knowledge, as measured by school grades in geometry (gg) and performance in a standardized test of geometrical knowledge (the california standards geometry test csgt). high correlations were predicted between performance in the vast-incons, csgt and gg, because they are all alternative measures of geometrical problem solving and geometrical knowledge. finally, we investigated the relation between vast-icons and two cognitive abilities that comprise spatial thinking: (i) abstract reasoning ability, i.e., the ability to identify patterns, analogical relationships and logical rules, and (ii) visuo-spatial ability – i.e., the ability to mentally visualize the rotation of objects. in view of the well-documented findings in the literature that spatial thinking is strongly related with students’ performance in stem subjects, we hypothesized that performance in the vast-incons should correlate positively with performance in the these two tests of spatial thinking (hypothesis 5). however, in accordance with our theoretical position, namely that it is the acquisition of geometrical knowledge that leads to the use of analytic strategies, we expected that geometry knowledge (as measured by the csgt) would significantly contribute to vast-incons performance, reducing the influence of the spatial thinking factor (hypothesis 6). 2. method 2.1 participants the participants included 30 mathematics teachers (age range 30-55 years, 18 men) and 134 11th grade students (age range 16.4-17.5 years, 71 boys). the mathematics teachers had considerable experience kospentaris  et  al           47   teaching high school geometry. the students were of middle-class backgrounds, came from two different schools, had four different geometry teachers, and were towards the end of a five-year course in geometry4. 2.2 materials the visual analytic shift test (vast) consisted of four items for each geometrical domain counterbalanced across the four conditions described earlier. thus, there were a total of 48 items (4 items for each geometrical domain × 3 domains × 4 conditions), randomly ordered. the california standards geometry test (csgt) consisted of 15 items (5 for each geometrical domain) selected from the overall 96 items of the california standards geometry test sample released in 2009 (http://www.cde.ca.gov/ta/tg/sr/documents/cstrtqgeomapr15.pdf). the selection was made on the basis of the relevance of each item to the national geometry curriculum. the purdue visualization of rotations test (rot) is a test that determines how well one can visualize the rotation of threedimensional objects. it is among the tests of spatial thinking less likely to be contaminated by analytical abilities. to restrict analytical processing, a time limit of 10 minutes for the 20item version of this test was strictly enforced (bodner & guay, 1997). the abstract reasoning test (art) is one of the tests of spatial thinking used by wai et al., (2009). it is a non-verbal measure of fluid intelligence, consisting of 15 items. school grades in geometry were collected for all students participating in the study. 2.3 procedure the vast was administered to the mathematics teachers individually in their school office. the vast and the csgt were administered to the students as a group test during a 45-minute class session. their order of presentation was counterbalanced. the students were instructed to answer the vast and csgt items using whatever method they found suitable. formulas were provided to students individually, if they asked for them. the rot and art were administered to small groups of students in the school computer lab. completion of the electronic tests required approximately 30 minutes. 3. results 3.1 reliability indices since all measures were binominal, the kuder-richardson (k-r) reliability test was applied. the results showed that the reliability of the two subscales was acceptable (vast-cons, kuder-richardson (kr) = .70; vast-incons, kuder-richardson (k-r) = .75). reliability of the rest scales are as follows: csgt (kuder-richardson (k-r)= .76, range of mean percentage performance: 13.33-100.00, mean= 63.82, std= 22.35), rot (kuder-richardson (k-r)= .74, range: 1-20, mean= 8.59, std= 3.81), and art (kuderrichardson (k-r)= .57, range: 5-15, mean= 9.89, std= 2.49).                                                                                                                           4 the students had been taught the basic geometric concepts and methods based on empirical measurements and inductive generalizations in grades 7, 8, and 9. in grades 10 and 11 they were introduced to the procedures of deductive proofs characterizing euclidean geometry. kospentaris  et  al           48   3.2 performance of the experts vs novices in order to examine the effect of the visual vs. analytic component on the participants’ performance, two composite scores were computed: the mean percentage performances in the vast consistent subscale (vast-cons) and in the vast inconsitent subscale (vast-incons). examination of the mean and standard deviation of the performance on the vast-cons, showed that the scale almost reached a ceiling effect (mean percentage= 83.86, see table 1). thus, as was planned, this subscale consisted of easy items that could be successfully solved by the math teachers as well as by the students. given that normality assumptions did not hold for this particular subscale, no parametric tests were applied. table 1 means and std as a function of expertise and item type (vast-cons vs. vast-incons) of the vast participants item type vast-cons vast-incons total mean std mean std mean std teachers 92.000 7.575 78.841 11.685 85.392 7.525 students 81.637 12.454 50.751 15.314 66.175 10.312 total 83.857 12.321 56.770 18.606 70.293 12.566 a t-test for independent samples was applied on vast-incons performance. results showed significant difference between the two groups [t(138)= -9.324, p<.001, mean= 50.75, for the students, and mean= 78.84, for the math teachers]. the math teachers answered correctly almost all of the items in the vast-incons, whereas the 11th graders had considerable difficulty with the vast-incons. in agreement with hypothesis 2, teachers’ and students’ performance was clearly differentiated in the vast-incons, where the performance of students was considerably lower than that of the teachers. in order to examine hypothesis 3, we plotted the students’ individual mean percentage performance in the vast-cons and the vast-incons. figure 3 shows the mean percent score of each participant on the y-axis. as can be seen, it was not the case that some participants performed well in the vast-cons and others in the vast-incons, as would have been predicted by the individual differences/cognitive style hypothesis. with very few exceptions, the items in the vast-cons that could have been solved using visuospatial strategies were much easier for each individual participant than the items in the vast-incons, which required recourse to analytic strategies. it can also be seen, that many students performed well only in the case of the vast-cons. kospentaris  et  al           49   figure 3. individuals’ mean performance in the vast-cons and the vast-incons. the performance of the teachers is shown above the horizontal line. 3.3 relations between the vast-cons, vast-incons csgt, gg rot and art due to the violation of the normality assumption of the vastcons, a spearman’s rho correlation analysis was performed on the mean scores of the vastcons with the vastincons, csgt, gg, rot and art. results showed that vastcons correlated moderately only with rot (rho=.190, p=.047), whereas all other correlations were insignificant (with vastincons, rho=.136, with csgt, rho=.165, with gg, rho=.054, and with art, rho=.176). 3.4 relations between the vast-incons, csgt and gg in order to examine hypothesis 5, a correlation analysis (pearson’s r) was performed on the mean percentage scores on the vast-cons, the vast-incons, rot (cronbach alpha= .74, range: 1-20, mean= 8.59, std= 3.81), and art (cronbach alpha= .57, range: 5-15, mean= 9.89, std= 2.49). as predicted, the results showed statistically significant correlations between performance in the two vast subscales, rot and art (table 2). 3.5 relations between the vast-incons, rot and art in order to examine hypothesis 5, a correlation analysis (pearson’s r) was performed on the mean percentage scores on the vast-incons, rot (cronbach alpha= .74, range: 1-20, mean= 8.59, std= 3.81), and art (cronbach alpha= .57, range: 5-15, mean= 9.89, std= 2.49). as predicted, the results showed kospentaris  et  al           50   statistically significant correlations between performance in the two vast subscales, rot and art (table 2). table 2 correlation between the measures of the study   1 2 3 4 5 1. vast-incons 2. csgt .429** 3. gg .318** .575** 4. rot .313** .357** .249** 5. art .366** .414** .285** .327** **. correlation is significant at the 0.01 level (2-tailed) *. correlation is significant at the 0.05 level (2-tailed) a stepwise regression analysis was applied with the purpose of examining in greater detail the contributions of the above-mentioned measures of this study to vast-incons performance. the measures were inserted in the analysis in the following order: rot, art, and csgt. the order of insertion followed the theoretical rationale of the present study, that is, general visuo-spatial ability (rot) was inserted first, followed by general abstract reasoning ability (art), and, finally, by performance on csgt, which incorporated all the above and included geometrical knowledge. based on our theoretical analysis, performance in csgt should predict performance in the vast-incons best. table 3 results of step-wise regression of rot, art, and csgt on vast-incons model b std. error β t sig. 1 rot 1.258 .432 .295 2.909 .005 2 rot .901 .430 .211 2.096 .039 art 1.772 .581 .307 3.048 .003 3 rot .522 .442 .122 1.179 .242 art 1.223 .603 .212 2.027 .046 csgt .196 .076 .282 2.561 .012 kospentaris  et  al           51   the results (see table 3) showed that all three consecutive models had a good fit. for the first step (rot) [f (1, 89) =8.461, p= .005], for the second step (rot and art) [f (2, 88) = 9.269, p< .001], and for the third step (rot, art and csgt) [f (3, 88) = 8.755, p< .001]. as it can be seen in table 3, when csgt was entered in the model the contribution of rot became non-significant (p= .242) and the contribution of art became marginally significant (p= .046). these results fully confirmed hypothesis 6, indicating that geometrical knowledge, and not general visuo-spatial abilities and abstract spatial reasoning, accounted for the visual/analytic strategy shift as measured by the vast-incons. in order to further validate the above result, a mediation analysis of the patterns of relations was applied on the data (see figure 4), by using amos (version spss21) through bootstrapping (number of bootstrap samples=2000, bias corrected confidence intervals= .95). first, the direct relations between rot and art on the vast-incons were computed. for rot and vast: two-tailed significance p< .045, and for art and vast: two-tailed significance p< .019. thus, the results indicated that both paths were significant. we then tested two models, one with the csgt and the other with geometry grades as mediating variables. when the csgt was added as a mediating variable, the indirect effect (i.e., the mediating path from rot through csgt to vast-incons) was significant (p= .006), and so was the mediating path from art through csgt to vast-incons (p= .001). inspection of the direct effects showed that both relations were completely mediated by csgt (for the path between rot and vast-incons, p= .143 and for the path art to vast-incons, p= .133). the best model that resulted (see figure 4), eliminating only the direct relation from rot to the vast-incons, had an acceptable fit (χ2 (1)= 2.384, p= .123, cfi= . 977, standardized rmr= .04). figure 4. regression weights of the mediation analysis between rot, art, csgt, and vast-incons. when geometry grades were treated as the mediating variable, the indirect effect from rot through geometry grades to vast-incons was significant (p= .042), and so was the mediating path from art through school grades to vast-incons (p= .025). inspection of the direct effects showed that the relation between rot and the vast-icons was completely mediated by geometry grades (for the path between rot and vast-incons p= .086), whereas the relation between art and the vast-incons was partially mediated by grades (p= .048). the final model, after eliminating the direct relation between rot and the vast-incons did not, however, show a good fit [χ2 (1)= 3.431, p= .06, cfi= .942, standardized rmr= .05], since the value of χ2/df exceeded the value of 3, and model’s p value was statistically significant. to conclude and summarize, the results of the mediation analysis confirmed the hypothesis that performance on the vast-incons will be mediated by geometrical knowledge, particularly when kospentaris  et  al           52   geometrical knowledge was measured by csgt. in addition, the model with the best fit still retained the direct relation between analytic reasoning as measured by art and the vast-incons. 4. discussion our main purpose in this study was to develop and validate a task that could distinguish visuo-spatial from analytic reasoning in geometry. as mentioned in the introduction, there have been several attempts so far to develop tasks that capture the change from the visual to the descriptive/analytic level in geometry. these previous attempts were not very successful because they were based on the recall or recognition of definitions of shapes and their properties, did not require thinking based on relational properties, and/or did not require the application of formal geometrical thinking in novel situations. the vast differs substantially from these previous attempts because it avoids the typical geometrystyle problems that impose an analytical solution method and because it consists of tasks that can be solved using either visuo-spatial or analytic strategies. the greatest advantage of vast, compared to previous tests, is that it focuses specifically on the antagonism between the two substantially different types of strategy in geometry and allows us to check individuals’ abilities to spontaneously choose and adequately apply the correct strategy. finally, the vast investigates the ability to use analytic strategies in situations where the visual element plays quite a central role and where visual information processing must be inhibited in favour of formal geometrical thinking. the results of the present study showed that the vast consists of items that have good internal consistency. most importantly the results show that the vast is a valid test because it can differentiate geometry teachers from students and because it correlates highly with other tests of spatial thinking and geometrical knowledge. 4.1 are the differences in vast performance related to geometry expertise? despite small differences in the ease or difficulty of the solution of individual items, the main pattern of results was the same: the items in the vast-cons that could be solved correctly using visuo-spatial strategies were much easier for all participants than the items in the vast-incons which required the use of analytic strategies and the inhibition of the visual element. as predicted (hypothesis 2), the math teachers were able to solve both the vast-cons and the vast-incons. this result suggests that the math teachers have access to both visuo-spatial and analytic strategies. on the contrary, the 11th grade students could easily solve only the vast-cons, but had great difficulty with the vast-incons. this finding indicates that the students relied predominantly on visuospatial strategies and had difficulty in employing analytic strategies when required by the task, further confirming hypothesis 2. the finding that there were systematic differences in the performance of the math teachers and the students -the math teachers performed well in both the vast-cons and the vast-incons, but the 11th grade students were able to perform well only in the vast-cons -supports the argument that the use of analytic strategies is related to the acquisition of geometry expertise. as explained in the introduction, if the differences in vast performance were due to individual differences in strategy use, then we should expect some mathematics teachers and some students to perform well in the vast-cons and others in the vastincons (hypothesis 3). this was not however the case. the items in the vast-incons were systematically more difficult than the items in the vast-cons both for teachers and for students. moreover, the students performed well only in the vast-cons, indicating that the majority of the students were able to successfully apply only simple visuo-spatial strategies. kospentaris  et  al           53   additional evidence in favour of the interpretation that the use of analytic strategies requiring geometrical knowledge in the vast-inconsistent sub-scale comes from the results of the mediation analysis which showed that relations between performance in the rot and the vastinconsistent subscale was completely mediated by performance in the csgt. this means that high performance in the vastinconsistent scale requires not just domain-general analytic skills, but domain-specific geometrical knowledge and ability to use analytic thinking in a geometrical context. the results showed considerable individual differences in the performance of the students in the vast-incons. these differences seem to be related to differences in geometry expertise within the student group. this conclusion can be deduced from the high correlations that were obtained between performance in the vast-incons, performance in the csgt, and gg (hypothesis 4). students’ performance in the vast-cons and assumed ability to apply simple visuo-spatial strategies did not correlate significantly with their gg or with their performance in the csgt, confirming hypothesis 4, namely, that this performance does not necessarily require geometrical knowledge. 4.2 what is the contribution of the cognitive skills involved in spatial thinking in vast performance? the results of the present study showed high correlations between performance in the vast-cons and the vast-incons with both rot and art, confirming the prediction that spatial thinking as measured by tests of visuo-spatial and abstract (spatial) reasoning abilities contributes to students’ performance in both scales of the vast. however, performance in the vast-incons was also highly correlated with performance in the test of geometrical knowledge (the csgt) and geometry grades. most importantly the results of the stepwise regression confirmed that when the csgt performance was inserted in the last step of analysis, the contribution of rot and art (which were statistically significant in the previous steps of the analysis), became nonor marginally significant. furthermore, when a mediation analysis was applied on the data, the previously significant direct relations between performance in the vast-incons and rot were mediated by the students’ geometrical knowledge, as measured by performance in the csgt. since the direct relations between art and vast-incons were not eliminated, however, it might be the case that analytic abilities are directly contributing to performance in the vast-incons. to sum up, the results indicate that the correct use of analytic strategies cannot be explained only on the basis of visuo-spatial abilities and abstract reasoning, i.e., the cognitive skills that comprise spatial thinking, but requires the accumulation of considerable geometrical knowledge. 4.3 implications for a theory of geometrical thinking one of the main reasons we developed the vast was in order to show that the learning of geometry requires significant conceptual changes to take place, and that instruction-induced conceptual changes culminate in the ability to use formal, geometrical knowledge in problem-solving and in the flexible use of visual/spatial and analytic strategies appropriate in the given contexts. the role of instruction here is quite crucial. as fischbein (1993) stressed, “the development of figural concepts generally is not a natural process” (p. 161). the present findings support the hypothesis that the use of analytic strategies in geometry is not a matter of individual differences in cognitive style but a major intellectual achievement, a conceptual change, which requires the acquisition of new forms of geometrical thinking. the application of analytic, formal geometrical knowledge in problem solving does not mean that visuo-spatial geometrical reasoning disappears. the fact that the math teachers could easily solve the items in the vast-cons suggests that they still had access to visual strategies. this issue needs to be investigated further, however, in view of the fact that analytic strategies could be used in the vast-cons also. finally, a great deal more research is also required to investigate the hypotheses of the ft according to which the kospentaris  et  al           54   processes of acquisition of geometrical knowledge are slow and gradual rather than sudden or stage-like, and that these processes can give rise to fragmentation and synthetic models. 4.4 implications for learning and instruction the low performance of the students in the vast-incons after almost five years of instruction in geometry supports the argument that the systematic use of analytic strategies is a major intellectual achievement that requires considerable conceptual changes. it should be added here that the 11th grade students were in their 5th year of geometry instruction, they had completed a two year course in plane eucleadian geometry, which was taught in a formal manner and accompanied by many geometry problems to be solved, and that they had an additional year’s course in analytic geometry designed for students opting for university study in stem subjects. consequently, all the students had been taught the formulae and theorems required to answer correctly all the vast items, both in the consistent and inconsistent sub-scales. thus, we can conclude from the above that it is possible to have acquired a great deal of school-type knowledge in geometry and not yet function at an analytic level in geometric thinking. although geometry instruction focuses almost exclusively on the acquisition of formal geometrical knowledge and the application of analytic strategies, it does not seem to be very successful in transferring to situations different from the narrow school context in which it is taught and in producing the necessary conceptual changes. this situation is similar to what is happening in other stem domains, such as physics or chemistry, where school instruction often fails to produce the necessary conceptual changes (e.g., clement, 1982; disessa, 1982). we hope that the present findings will further sensitize educators of the need to develop instruction that emphasizes not the recall and rigid application of formal definitions and rules but the constructive, dynamic activities of students that can help them understand how formal definitions fit with their visualspatial experiences and representations of geometrical shapes (fischbein, 1993; kilpatrick, hoyles, skovsmose, & valero, 2005; lehrer jenkins & osana, 1998). 4.5 limitations of the present study and future research the present work has a number of limitations and leaves several open questions to be answered by future research. first, the sample of the present study is small when attempting to validate a new measure such as the vast and therefore the results presented in this exploratory study need to be replicated with larger samples and more age groups. in addition, reaction time studies as well as use of qualitative methods, such as interviews, think-aloud protocols and eye-movement tracking could be used to further validate the vast. although performance in the vast differentiates teachers from experts and is related to geometry expertise, the results do not provide information about how analytic reasoning in geometry actually develops. developmental research is needed to further examine the hypotheses of the ft that knowledge acquisition in geometry is a continuous and not a stage-like process, during which fragmentation and synthetic conceptions are formed. future research needs to also investigate the contribution of intellectual abilities and mathematics knowledge to vast performance by administering additional tests, such as a propositional test measuring verbal analytical skills, measures of visuo-spatial and phonological memory, speed of processing and executive function, as well as measures of mathematics abilities. results from these tests could be used as co-variates in order to reduce as much as possible the effect of individual differences. last but not least, the relation between individual differences in spatial thinking and the use of visual and analytic strategies in geometry needs to be investigated further, preferably using longitudinal designs. if, as we claim, geometry expertise (and possibly expertise in other domains of stem) requires the eventual development of analytic strategies, why are individuals who are good in spatial thinking more successful in stem disciplines than those who are not? a possible explanation of this finding may be that individuals kospentaris  et  al           55   who are good in spatial thinking find it easier to do well in geometry early on, before conceptual changes in this domain require the development of analytic strategies based on formal, geometrical knowledge. maybe because of these early successes, these students develop an interest in geometry (or other stem-related disciplines), spend more time studying, and thus eventually undergo the conceptual changes and develop the analytic strategies required for geometrical expertise. this conjecture is consistent with the findings of the present study that spatial thinking as measured by visuo-spatial and abstract (spatial) reasoning tests contributes to success in the vast, but is rendered insignificant for the vast-incons when geometrical knowledge is taken into account. this issue needs to be further researched. keypoints considerable conceptual changes are required to go from visuo-spatial reasoning to analytic strategies in geometry these changes are related to geometry expertise and not to individual differences questions arise about the role of visual-spatial reasoning in geometry expertise acknowledgments the research reported in this paper is supported by a grant from the greek ministry of education, general secretariat for research and technology, molvisedu, thalis – aristotle university of thessaloniki. we would like to thank petros roussos for helpful comments. references battista, m. t. (2007). the development of geometric and spatial thinking. in f. lester (ed.), second handbook of research on mathematics teaching and learning, nctm (pp. 843-908). reston, va: national council of teachers of mathematics. bodner, g.m., & guay, r.b. (1997). the purdue visualization of rotations test. the chemical educator, 2(4), 1-17. burger, w. f., & shaughnessy, j. m. (1986). characterizing the van hiele levels of development in geometry. journal for research in mathematics education, 17, 31-48. csgt california standards test geometry sample (2009). http://www.cde.ca.gov/ta/tg/sr/ documents/cstrtqgeomapr15.pdf carey, s. (2009). the origin of concepts. new york, ny: oxford university press. cheng, y.l., & mix, k.s. (2012). spatial training improves children's mathematics ability. journal of cognition and development (published online, september 19, 2012). http://www.tandfonline.com/doi/pdf/10.1080/15248372.2012.725186 chi, m. t. h., (2008). three types of conceptual change: belief revision, mental model transformation, and categorical shift. in s. vosniadou (ed.), international handbook of research on conceptual change (pp. 61-82). new york, ny: routledge. clement, j. (1982). students' preconceptions in introductory mechanics. the american journal of physics, 50(1), 66-71. clements, d. h., & battista, m. t. (2001). logo and geometry. journal for research in mathematics education monograph.reston, va: national council of teachers of mathematics. disessa, a., (1982). unlearning aristotelian physics: a study of knowledge-based learning. cognitive science, 6(1), 37-75.doi: 10.1207/s15516709cog0601_2 kospentaris  et  al           56   eisenberg, t., & dreyfus, t. (1991). on the reluctance to visualize in mathematics. in w. zimmermann & s. cunningham (eds.), visualization in teaching and learning mathematics (pp. 26-37). washington, dc: maa. evans, j. st. b.t., & stanovich, k. (2013). dual-process theories of higher cognition: advancing the debate. perspectives on psychological science, 8(3), 223–241. doi: 10.1177/1745691612460685 fischbein, e. (1993). the theory of figural concepts. educational studies in mathematics, 24 (2), 139-162. gutiérrez, a., & jaime, a. (1998). on the assessment of the van hiele levels of reasoning. focus on learning problems in mathematics, 20(2-3), 27-46. hegarty, m. (1992). mental animation: inferring motion from static diagrams of mechanical systems. journal of experimental psychology: learning, memory and cognition, 18, 1084–1102. hegarty, m., montello, d. r., richardson, a. e., ishikawa, t., & lovelace, k. (2006). spatial abilities at different scales: individual differences in aptitude-test performance and spatial-layout learning. intelligence, 34, 151-176. hegarty, m., & waller, d. (2006). individual differences in spatial abilities. in p. shah & a. miyake (eds.), handbook of visualspatial thinking. cambridge, ma: cambridge university press. houdement, c., & kuzniak, a. (2003). elementary geometry split into different geometrical paradigms. proceedings of cerme 3, belaria, italy. http://www.dm.unipi.it/~didattica/ cerme3/proceedings/groups/tg7/tg7_houdement_cerme3.pdf kastens, k.a., & ishikawa, t. (2006). spatial thinking in the geosciences and cognitive sciences: a crossdisciplinary look at the intersection of the two fields. in c. a. manduca & d. w. mogk (eds.), earth and mind: how geologists think and learn about the earth (pp. 53-76). boulder, co: geological society of america. kilpatrick, j., hoyles, c., & skovsmose, o. (eds.) in collaboration with valero, p. (2005). meaning in mathematics education, usa: springer. kozhevnikov, m., motes, m., & hegarty, m. (2007). spatial visualization in physics problem solving. cognitive science, 31, 549-579. lehrer, r., jenkins, m., & osana. h. (1998). longitudinal study of children’s reasoning about space and geometry. in r. lehrer & d. chazan (eds.), designing learning environments for developing understanding of geometry and space (pp. 137-167). mahwah, nj: erlbaum. lynn, b. m., & lynch, c. m. (2010). van hiele revisited. mathematics teaching in the middle school, 16(4), 232-238. newcombe, n. s. (2010). picture this: increasing math and science learning by improving spatial thinking. american educator, summer, 29-43. newcombe, n. s., & frick, a. (2010), early education for spatial intelligence: why, what, and how. mind, brain and education, 4(3), 102-111. newcombe, n. s., & huttenlocher j. (2000). making space: the development of spatial representation and reasoning. cambridge, ma: mit press. newcombe, n. s., uttal, d. h. & sauter, m. (in press). spatial development. in p. zelazo (ed), oxford handbook of developmental psychology. new york, ny: oxford university press. piaget, j., & inhleder, b., (1948/1967). the child’s conception of space. london: w.w. norton & company. piaget, j., inhelder, b., & szeminska, a. (1948/1960). the child’s conception of geometry. london: routledge and kegan paul. pitta-pantazi, d., & christou, c. (2009). cognitive styles, dynamic geometry and measurement performance. educational studies in mathematics, 70, 5-26. sanchez, c.a. (2012). enhancing visuospatial performance through video game training to increase learning in visuospatial science domains. psychonomic bulletin & review, 19 (1), 58–65. schwartz, d.l., & black, j.b. (1996). shuttling between depictive models and abstract rules: induction and fallback. cognitive science, 20, 457-497. shea, d. l., lubinski, d., & benbow, c. p. (2001). importance of assessing spatial ability in intellectually talented young adolescents: a 20-year longitudinal study. journal of educational psychology, 93(3), 604–614. kospentaris  et  al           57   shtulman, a., & valcarcel, j. (2012). scientific knowledge suppresses but does not supplant earlier intuitions. cognition, 124, 209–215. spelke, e. s., & kinzler, k. d. (2007). core knowledge. developmental science, 10, 89–96. spelke, e.s., lee, s. a., & izard, v. (2010). beyond core knowledge: natural geometry. cognitive science, 34(5), 863-884 stieff, m. (2007). mental rotation and diagrammatic reasoning in science. learning and instruction, 17, 219234. usiskin, z. (1982). van hiele levels and achievement in secondary school geometry. colombus, oh: eric. uttal, d. h., meadow, n. g., tipton, e., hand, l. l., alden, a.r., warren, c., & newcombe, n.s. (2013). the malleability of spatial skills: a meta-analysis of training studies. psychological bulletin, 139, 352402. van hiele p. m. (1986). structure and insight: a theory of mathematics education. london: academic press inc. vosniadou, s., & brewer, w. f. (1994). mental models of the day/night cycle. cognitive science, 18, 123183. vosniadou, s. (2013). conceptual change in learning and instruction: the framework theory approach. in s. vosniadou (ed.), international handbook of research on conceptual change, 2nd edition (pp. 11-30). new york, ny: routledge. vosniadou, s., & skopeliti, i. (2014). conceptual change from the framework theory side of the fence. science and education, 23(7), 1427-1445. vosniadou, s., pnevmantikos, d., makris, n., ikospentaki, k., lepenioti, d., chountala, a., & kyrianakis, g. (2015). executive functions and conceptual change in science and mathematics learning, 7th annual conference of the cognitive science society, pasadena, ca. https://mindmodeling.org/cogsci2015/papers/0434/paper0434.pdf wai, j., lubinski, d., & benbow, c. p. (2009). spatial ability for stem domains: aligning over 50 years of cumulative psychological knowledge solidifies its importance. journal of educational psychology, 101(4), 817–835. wiser, m., & smith, c. (2013). learning and teaching about matter in the middle school years. how can atomic-molecular theory be meaningfully introduced? in s. vosniadou (ed.), international handbook of research on conceptual change, 2nd edition (pp. 177-194). new york, ny: routledge. zazkis, r., dubinsky, e., & dautermann, j. (1996). coordinating visual and analytic strategies: a study of students' understanding of the group d 4. journal for research in mathematics education, 27(4), 435– 457. frontline learning research 6 (2014) 46-66 issn 2295-3159 corresponding author: fang zhao, department of psychology, university of koblenz-landau, germany, email: zhao@unilandau.de doi: http://dx.doi.org/10.14786/flr.v2i4.98 46 | f l r eye tracking indicators of reading approaches in text-picture comprehension fang zhao a , wolfgang schnotz a , inga wagner a , robert gaschler a a university of koblenz-landau, germany article received 21 march 2014 / revised 4 august 2014 / accepted 21 october 2014 / available online 27 october 2014 abstract despite numerous studies on reading and multimedia comprehension, the usage of text and picture with different reading strategies has rarely become a focus of research. the current study aims to explore whether the usage of text differs from the usage of picture when readers follow different strategies of knowledge acquisition. in a within-subjects design using eye tracking, seventeen secondary school students comprehended blended text and picture materials with three different strategies. (1) initial coherence-formation strategy, which requires students to process text and picture unguidedly. (2) consecutive task-oriented strategy, which requires them to gather information to answer the question (which explains the task) provided after prior experience with text and picture. (3) initial task-oriented strategy, which requires them to comprehend text and picture to solve the task equipped with the prior information about the question from the beginning. eye tracking data showed that text and picture play different roles in these processing conditions. (1) the results are in line with the assumption that text (rather than picture) is more likely used to construct mental models in initial coherence-formation processing of text and picture. (2) students seem to primarily rely on the picture to answer the question after the prior experience with the material with consecutive task-oriented strategy. (3) text and picture are both used heavily when the question is presented first, enabling students to selectively process question-relevant aspects of the material at first contact. keywords: multimedia learning; mental model; eye tracking f. zhao et al 47 | f l r after learning to read in primary school, students are required to use their reading skills for learning from written materials. in secondary school, these materials usually include text and different kinds of pictures, such as maps and diagrams. students therefore need skills for integrating text and picture information in order to build the required knowledge structures (ainsworth, 1999; kintsch & van dijk, 1978). to support students’ learning from text and pictures, we need sophisticated knowledge about the usage of text and pictures under different learning conditions. unfortunately, there is so far not much knowledge available about integrated processing of text and pictures. abundant studies have explored reading strategies in text comprehension (e.g. anderson & pearson, 2002; frase, 1967, 1968; rothkopf, 1964, 1966; rouet, 2006). the aim of reading can fundamentally influence cognitive processing due to the application of different reading strategies (andre, 1979; hamilton, 1985; rickards, 1979). however, there has not been much research about reading strategies for combined comprehension of text and pictures. the current study aims at exploring whether the usage of texts differs from the usage of pictures when readers follow different strategies of knowledge acquisition. first, we will introduce a theoretical framework of text-picture comprehension. second, we will formulate research questions and derive hypotheses about the usage of text and pictures under the condition of different reading strategies. third, we will describe a study that was designed to test these hypotheses. fourth, we present the results in the light of the previously mentioned hypotheses. fifth, we will discuss the empirical findings and analyze their relations to findings from the previous literature. 1. theory 1.1 theories of text-picture comprehension many theories focused on text-picture comprehension (tpc; kress & leeuwen, 1996; rouet & britt, 2011; zwaan, 1998). however, there are mainly three theories specifically targeting formats of mental representations involved in tpc: (1) dual coding theory (paivio, 1986), (2) the cognitive theory of multimedia learning (mayer, 2005) and (3) the integrative model of tpc (schnotz & bannert, 2003). the three theories share the assumption of separate channels for text and picture processing but differ on other issues. the dual coding theory focuses on the referential connections between text and pictures and assumes that people can retrieve information better by using two different channels. the cognitive theory of multimedia learning assumes that people process information through an auditory-verbal channel and a visual-pictorial channel of limited capacity. multimedia learning is assumed to include: (i) selecting relevant words; (ii) selecting relevant images; (iii) organizing the selected words into a verbal mental model; (iv) organizing the selected images into a pictorial mental model; and (v) integrating the verbal model and the pictorial model with prior knowledge into a coherent mental representation. the integrative model of tpc combines the concepts of multiple memory systems, multiple sensory modalities, and two kinds of representations: descriptions (such as natural language and propositional representations) and depictions (such as pictures, visual images and mental models). according to this theory, readers construct only one mental model, which contains verbal and pictorial information. due to the importance of distinguishing descriptive and depictive representations for the analysis of reading strategies, our study is mainly inspired by the integrative model of tpc. 1.2 reading strategies of text-picture comprehension in order to explain differences between reading strategies when learning from texts, rickards and denner (1978) suggested a distinction between general and specific processing. general processing deals with the global thematic coherence of the text, whereas specific processing focuses on unique information required for specific purposes. there seems to be an inherent conflict between these two kinds of processing, f. zhao et al 48 | f l r as rickards and denner found that pre-posed questions can lead to a highly selective processing, but at the expense of global understanding of the text. according to our knowledge, this distinction between two different kinds of processing has not been applied to the combined processing of text and pictures, yet. in line with rickards and denner, we differentiate between general coherence-formation processing and selective task-oriented processing of text and pictures. the two kinds of processing are not meant as a strategy dichotomy. instead, coherenceformation processing and task-oriented processing are strategy components that can be combined. as time and processing resources are limited, however, the two components cannot be both maximized at the same time. instead, they can obtain different emphasis in the process of tpc. thus, there is a continuum with a primarily general coherence-formation processing at the one end and a primarily selective task-oriented processing at the other end. depending on the specific learning situation, different kinds of processing will be combined into a suitable strategy. in the following, we will consider three different learning situations: (a) if a reader receives a text with pictures without a specific goal in mind, he/she will put the emphasis on general coherence-oriented processing. that is, he/she will try to construct coherent mental representations based on the available information. we will refer to this kind of processing as the initial coherence-formation strategy. (b) if a reader has first processed a text with pictures without a specific goal in mind (i.e., applied an initial coherence-formation strategy) and then gets access to the specific questions to be answered, he/she will put emphasis on selective task-oriented processing and focus on task-relevant information. we will refer to this kind of processing as the consecutive task-oriented strategy. (c) if a reader is presented specific questions before receiving a text with pictures to be used for answering these questions, he/she will put more emphasis on selective task-oriented processing. however, although processing is goal-directed from the very beginning on, some coherence-oriented processing is also required because the reader needs some understanding of what the text and the picture are about. we will refer to this kind of processing as the initial task-oriented strategy. 2. research questions and hypotheses the abovementioned strategies refer to general vs. specific information processing (rickards & denner, 1978), without specifying the differential roles that text vs picture might play in these strategies. in order to fill this conceptual gap, we proposed research questions and hypotheses on the usage of text vs. picture taking the described strategies into consideration. we arranged the order of presentation of (a) the question (b) the material containing text and picture, as well as potential prior exposure to the material in that way, that these context factors simulated different processing. this allowed us to compare the relative amount of text vs. picture processing among the three strategies based on eye tracking data. as mentioned above, the initial coherence-formation strategy is comparatively general and coherence-driven. as learners have not been provided with a question to be solved based on the material, yet they might process the material in order to construct a coherent mental model covering the general content of the material. the consecutive task-oriented strategy is fairly specific and goal-driven. it is supposed to come into play when learners are provided with a question they should solve based on the material which they have already processed before. conversely, the initial task-oriented strategy is a combination of coherencedriven and goal-driven processing. in this case, learners are provided with the question to be solved before they get access to the material. they can thus selectively process information that is relevant to the question at first contact. as there are two strategies with a higher proportion of question-oriented processing, we will focus (1) on the comparison between initial coherence-formation processing and consecutive task-oriented processing and (2) on the comparison between initial coherence-formation processing and initial taskoriented processing. our research questions concerned potential differences between the usage of texts and the usage of pictures in the different processing. more specifically, our first research question was: f. zhao et al 49 | f l r (1) does the initial coherence-formation strategy (no question yet) differ from the consecutive taskoriented strategy (question after material) in terms of using texts and pictures? we hypothesized that text processing differs fundamentally from picture processing with the initial coherence-formation strategy and with the consecutive task-oriented strategy. this may be routed in different functions of text and pictures in reading comprehension as well as in the effect of reading strategy. reading through a text is a major activity to make the meaning of the content (schmidt-weigand et al., 2010), whereas pictures mainly serve as an external representation scaffolding the answering of questions (eitel et al., 2013). initial coherence-formation processing guides readers to comprehend without any task, which is more coherence-driven and general. in order to understand the main content, learners probably pay more attention to text with the initial coherence-formation processing than with the consecutive task-oriented processing. in comparison, consecutive task-oriented processing guides readers to solve the task after their prior knowledge construction, which is rather task-oriented and selective. participants might thus pay more attention to pictures with consecutive task-oriented processing than with initial coherence-formation processing in order to solve the question. (2) does initial coherence-formation processing (no question yet) differ from initial task-oriented processing (question before material) in terms of using texts and pictures? we also expected text and pictures to be comprehended differently between the initial coherenceformation strategy and the initial task-oriented strategy. the initial task-oriented strategy combines coherence formation and selective processing at first contact with the material. when readers engage in initial task-oriented processing, they, on the one hand, know the aim from the start and reading is relatively goal-directed and selective. on the other hand, they need to understand the content to search for the relevant information, which makes reading comparatively coherence-driven. in other words, the initial task-oriented strategy is more goal-driven and less coherence-driven than the initial coherence-formation strategy. as readers primarily use text for understanding, we assumed that participants focus more on text with the initial coherence-formation strategy than with the initial task-oriented strategy. as pictures can assist task solving (larkin & simon, 1987), we hypothesized that participants focus more on pictures and less on texts with the initial task-oriented strategy than with the initial coherence-formation strategy. 3. method 3.1 participants seventeen participants from secondary schools in germany were included in this study (m = 13 years, sd = 3.4 years). eleven participants were male and six were female. we used the heller and perleth (2000) intelligence tests on spatial ability and verbal ability. participants were marginally above average of the norm sample in spatial ability (average t of 53.65; sd = 7.39) and verbal ability (average t of 55.59; sd = 9.08). 3.2 materials in a previous pilot study, we had selected 60 text-picture units from textbooks about geography and biology with 288 test questions. the units and questions were tested with 1060 students in grades 5 to 8 according to item-response theory including dif-analyses for gender, grade and school. additionally, we carried out a rational task analysis on questions (schnotz et al., 2011). in order to answer the questions correctly, participants need to process both text and pictures. as we adopted eye-tracking methodology, we used only a subset of the text-picture units for pragmatic reasons. each participant received six text-picture units. the units were selected in a way that the type of image (realistic pictures vs. graphs) and the level of difficulty (easy vs. medium vs. difficult) were balanced throughout the experiment. participants were f. zhao et al 50 | f l r randomly distributed to different task orders to control for sequencing effects. the selected units (see appendix) and their average difficulties (beta-values in terms of item-response theory) were as follows: 1) banana trade: easy level (beta = -0.95) containing 95 words; realistic image in geography. 2) legs of insects: easy level (beta = -0.75) containing 59 words; realistic image in biology. 3) auditory: medium level (beta = 0.10) containing 122 words; graph in biology. 4) pregnancy: medium level (beta = 0.39) containing 136 words; realistic image in biology. 5) map of europe: difficult level (beta = 1.37) containing 136 words; map in geography. 6) savannah: difficult level (beta = 1.47) containing 170 words; graph in geography. 3.3 stimulating specific strategies each participant was instructed to process six text-picture units under different conditions in order to gather eye tracking indicators of text vs. picture processing under three different processing conditions. three units were presented without any information about the task (question) to be solved afterwards. this was expected to stimulate an initial coherence-formation strategy. after this first phase of processing, the task appeared on the screen and participants were asked to solve the task. participants could then re-read the text and re-observe the picture under the guidance of the task. this second phase of processing was expected to stimulate a consecutive task-oriented strategy. the other three units were presented when the participants had the task to be solved already in mind. participants read the question first. after participants had read the task, the corresponding text and pictures appeared on the screen. therefore, participants could explore the text and the pictures under the guidance of the question from the very beginning on. this was expected to stimulate an initial task-oriented strategy. f. zhao et al 51 | f l r figure 1. design overview for the three reading conditions. the timeline indicates the time that text, picture and question appeared on the screen. 3.4 procedure we conducted the experiments (each taking about 45 min) individually in a lab environment with the permission of students’ parents. materials and instructions were in german (native language of the research participants). after being informed about the purpose of the study, participants took the paper-pencil iq tests and watched an instruction video. participants were seated at 60-65 cm distance from the 24-inch monitor of the eye tracker positioned vertically at the eye level. a 5-point calibration was conducted before participants read the material. once the calibration was successful, the experiment would start. the experiment included a warm-up phase and a main test. the aim of the warm-up phase was to give participants the opportunity to get familiar with the eye tracking system and with using keyboard and mouse for turning pages and answering the questions. we used a tobii xl60 eye tracker to record eye movements at the rate of 60 hz. the system can compensate for head movements and thus provides relatively precise data from young participants and a comfortable testing situation. for each unit, the strategy was indicated by a written instruction presented upfront. finally, participants were thanked and rewarded with € 12 for taking part in our experiment. 3.5 indicators of cognitive processing past research has established the basic assumption that cognitive processes influence and are mirrored by eye movement indicators (e.g. gaschler, marewski, & frensch, in press; godau et al., 2014). the positions that a person’s eyes fixate are related to cognitive processing according to the eye-mind hypothesis and the immediacy hypothesis (just & carpenter, 1980). the eye-mind hypothesis assumes that eye movements can reflect information processing. the immediacy hypothesis states that the processing of information is immediate and happens directly after it is perceived. fixation counts and fixation duration, i.e. accumulated fixation counts and the fixation time at a particular area of interest (aoi) are associated with the depth of cognitive processing and spatial distribution of attention (hegarty, 1992; rayner, 1998). visit counts and visit duration (i.e. number of entries into and exits from an aoi and the accumulated duration of visits at an aoi) are also used in detecting reading processing as they mirror the importance of the aoi and the perceived informativeness (jacob & karn, 2003; friedman & liebelt, 1981). time to the first fixation of an aoi (i.e. latency until an aoi is fixated for the first time) indicates readers’ interest in the aoi (goldberg f. zhao et al 52 | f l r & kotval, 1999). transitions between aois (i.e., frequency of the eye movements from one aoi to the other) can mark the integration of contents presented in the aois (johnson & mayer, 2012). 3.6 scoring the aois were drawn manually using tobii studio software. each task had three separated aois: the picture, the text and the question. the average picture aoi for the six materials covered 25.14% of the screen; and the average text aoi covered 27.72%; the average question aoi occupied 18.55%. participants obtained one point when they answered a question correctly. they could obtain a maximum score of six points (i.e., accuracy rate of 100%) and a minimum score of zero points. 4. results participants had an average of 59% accurate answers to the questions (sd = 25%) with initial coherence-formation strategy combined with consecutive task-oriented strategy and an average of 53% accurate answers (sd = 36%) with initial task-oriented strategy, t(16) = 0.65, p = .52, d = 0.19. the average time for comprehending text and pictures and answering the question was 1.14 seconds per word (sd = 0.39 second per word) from a range of 1.71 seconds per word to 0.61 seconds per word. most importantly, the different experimental conditions of processing strategies differed in eye fixations on text vs. picture. considering the variation of reading speed between participants and the limited number of participants, we decided to use proportion measures for capturing the relative weight of text vs. picture (i.e. proportion of fixation counts on text, proportion of fixation time on text, proportion of visit counts on text and proportion of visit duration on text; holmqvist et al., 2011). for example, the proportion of fixation counts on text was calculated by dividing this count by the sum of fixation counts on text and picture. due to the dependency of text and picture, we only show the data for text with these indicators to avoid redundancy. thus, if we refer to high proportion of fixations on text, this implies a low proportion on the picture and vice versa. in order to explore whether the usage of text differs from the usage of picture with different strategies, we conducted two one-way repeated-measures multivariate analysis of variance (manova) with four eye-tracking indicators in three reading conditions. as there were two types of task-driven strategies, two manovas were performed: (1) initial coherence-formation strategy (no question yet) vs. consecutive task-oriented strategy (question after material), and (2) initial coherence-formation strategy vs. initial taskoriented strategy (question before material). the eye-tracking indicators included the proportion of fixation counts and fixation duration on text, the proportion of number of visit and visit duration on text, the time to the first fixation on text and picture and the number of transitions between text and picture. 4.1 initial coherence-formation strategy vs. consecutive task-oriented strategy in order to check whether the initial coherence-formation strategy (no question yet) differs from the consecutive task-oriented strategy (question after material) in fixations on text vs. picture, we performed a manova and univariate analyses (anovas) with the eye tracking indicators listed in table 1. the initial coherence-formation strategy and consecutive task-oriented strategy differed significantly with respect to the relative weight on text (rather than picture) across the eye tracking indicators, f(7, 26) = 14.10, p < .001, ηp 2 = .79. as reported below, separate anovas confirmed the difference between the two strategies for indicators like fixation counts and fixation duration, visit counts and visit duration and time to the first fixation. (1) fixation indicators during text-picture comprehension the anova revealed that the proportion of fixation counts on texts (rather than on picture) was significantly higher with initial coherence-formation processing than with consecutive task-oriented f. zhao et al 53 | f l r processing, f(1, 32) = 56.51, p < .001, ηp 2 = .64. proportion of fixation counts on texts and proportion of accumulated fixation duration on texts were correlated by r = .97. thus, unsurprisingly the proportion of accumulated fixation duration on text was higher with initial coherence-formation strategy than with consecutive task-oriented strategy, f(1, 32) = 70.41, p < .001, ηp 2 = .69. (2) visit indicators during text-picture comprehension participants visited the text aoi (rather than the picture aoi) more often with the initial coherenceformation strategy than with the consecutive task-oriented strategy. the proportion of number of visits on text was higher when participants engaged in initial coherence-formation processing rather than in consecutive task-oriented processing, f(1, 32) = 7.93, p = .008, ηp 2 = .20. participants had a higher proportion of accumulated visit duration on text with initial coherence-formation processing than with consecutive task-oriented processing, f(1, 32) = 14.24, p = .001, ηp 2 = .31. (3) time to first fixation on text and picture participants fixated within a shorter time on texts with the initial coherence-formation strategy than with the consecutive task-oriented strategy, f(1, 32) = 9.43, p = .004, ηp 2 = .23. they also fixated more quickly on pictures with the initial coherence-formation strategy than with the consecutive task-oriented strategy, f(1, 32) = 4.96, p = .033, ηp 2 = .13. pictures were fixated slightly quicker than texts: f(1, 32) = 0.87, p = .358, ηp 2 = .03 for the initial coherence-formation strategy; f(1, 32) < 1, for the consecutive taskoriented strategy. (4) transitions between text and picture the transitions between text and picture did not show a robust difference between initial coherenceformation and consecutive task-oriented processing, f(1, 32) = 2.34, p = .136, ηp 2 = .07. with the consecutive task-oriented strategy, participants had 27% of transitions (sd = 10.6%) between text and picture, 25% of transitions (sd = 10.4%) between text and question and 48% of transitions (sd = 13.7%) between picture and question. in short, participants transferred their eyes slightly more often between text and picture with initial coherence-formation strategy than with consecutive task-oriented strategy. when questions were illustrated, participants mainly transferred their attention between picture and question. table 1 means and standard deviations of eye tracking indicators in different goal-oriented strategies eye tracking indicators initial coherenceformation consecutive task-oriented initial taskoriented m (sd) m (sd) m (sd) % fixation counts on text 80.11 (12.81) 47.87 (12.19) 64.95 (16.4) % fixation time on text 81.27 (12.99) 44.00 (12.73) 66.93(15.72) % visit counts on text 48.29 (9.13) 39.64 (8.79) 46.16 (8.47) % visit time on text 66.72 (20.27) 44.88 (12.61) 69.31 (11.91) time to first fixation on text (sec) 2.11 (1.44) 6.60 (5.85) 2.21 (1.92) time to first fixation on picture (sec) 1.57 (1.92) 5.02 (6.09) 0.87 (2.07) average number of transitions between text and picture 21.59 (13.4) 14.94 (14.47) 28.53 (16.18) f. zhao et al 54 | f l r 4.2 initial coherence-formation strategy vs. initial task-oriented strategy eye tracking indicators also showed that the initial coherence-formation strategy (i.e. general processing without task) differed from the initial task-oriented strategy (i.e. task was presented at the beginning) in terms of text and picture processing. the manova showed the significant effect of goalorientation on eye tracking indicators, f(7, 26) = 3.49, p = .009, ηp 2 = .48. separate anovas yielded effects of strategy condition on fixation indicators but not on visit indicators. (1) fixation indicators during text-picture comprehension there was a significantly higher proportion of fixation counts on text (rather than on picture) with the initial coherence-formation strategy than with the initial task-oriented strategy, f(1, 32) = 9.03, p = .005, ηp 2 = .22. as proportion of counts and of duration of fixations were highly correlated (r = .84), this was mirrored by a similar effect on proportion of fixation duration on text, f(1, 32) = 8.41, p = .007, ηp 2 = .21. (2) visit indicators during text-picture comprehension participants had a similar pattern of results in visiting text and picture for both experimental strategy conditions. no difference was detected for the proportion of visit counts and visit duration on text between the initial coherence-formation strategy and the initial task-oriented strategy, fs(1, 32) < 1. (3) time to first fixation on text and picture we did not find any difference between the experimental strategies for latency of first fixation on text or picture, fs(1, 32) < 1.03. participants fixated the picture marginally sooner than the text with initial task-oriented processing, f(1, 32) = 3.79, p = .06, ηp 2 = .11. (4) transitions between text and picture transitions between text and picture did not differ when participants followed the initial coherenceformation strategy vs. the initial task-oriented strategy, f(1, 32) = 1.48, p = .23, ηp 2 = .04. for the initial taskoriented strategy, participants had 45% (sd = 13%) of transitions between text and picture, 20% (sd = 7.6%) between text and question and 35% (sd = 17.5%) between picture and question. in brief, participants transferred their eyes dominantly between text and picture, secondarily between picture and question and lastly between text and question. 5. discussion the current eye tracking study provides first methodological tools and results to specify the distinction between general and specific processing proposed by rickards and denner (1978) for processing of text vs. picture in mixed material. general processing (i.e., when the question is not known yet) deals with the global thematic coherence of the material. pre-posed questions can lead to a highly selective processing. in order to apply this account to the processing of mixed material (text and picture), the current study examined whether text processing differs from picture processing and whether this difference is moderated by the strategies used by the learner. specifically, we compared (1) initial coherence-formation strategy (no question yet) with consecutive task-oriented strategy (question after material) and (2) initial coherenceformation strategy with initial task-oriented strategy (question before material). a higher emphasis on text relative to picture was expected for the initial coherence-formation strategy. pictures were assumed to lead to higher values in fixation indicators with the consecutive task-oriented processing. we used eye tracking indicators to reveal the processing of text and picture, according to the eyemind theory and the immediacy theory. our results confirmed general differences of text vs. picture processing (mcnamara, 2007). importantly, eye tracking indicators of relative emphasis on text (rather than on picture) differed among the experimentally induced processing strategies. for the initial coherenceformation strategy, participants primarily fixated on text rather than on picture. as fixation count and f. zhao et al 55 | f l r fixation duration have been linked to the depth of cognitive processing and distribution of attention (rayner, 1998), the result suggest the primary usage of text during mental model construction in initial coherenceformation processing. the same is true for visit counts and visit duration. as visit indicators have been linked to the importance and informativeness of the aois (jacob & karn, 2003), participants possibly consider text to be important and informative with initial coherence-formation strategy. participants also needed little time to proceed to text and picture (i.e. time to the first fixation) with initial coherenceformation strategy. this indicator has been linked to participants’ interest on text and picture when reading is general (goldberg & kotval, 1999). in addition, participants had frequent transitions between text and picture with initial coherence-formation strategy. according to the integrative model of tpc, participants may establish their mental model by integrating text and picture. this is consistent with our assumption that participants intensely processed the content to establish an initial mental model with initial coherenceformation strategy. for the consecutive task-oriented strategy, picture was mainly used to scaffold question solving after the initial construction of the mental model. the results from fixation counts and fixation duration were consistent with the idea that pictures are primarily used when participants need to answer the question with consecutive task-oriented processing (cf. hochpöchler et al., 2012). they had a high amount of visit counts and visit duration on pictures with consecutive task-oriented processing. it seems that participants considered pictures important and informative when they were asked to solve tasks after the initial construction of the mental model. besides, data also revealed that participants perceived pictures sooner than texts with both processing strategies. this can be explained by pictures attracting readers’ attention (mayer, 1989; tversky, 2001; winn, 1989). transition data showed that they focused their main attention on picture and question, less attention on text and picture and the least on text and question when they followed consecutive taskoriented strategy. participants seemed to primarily use the picture to scaffold question answering. they paid less attention on text and picture because they have already constructed the initial mental model and they just need to further construct or update this mental model in order to answer the question. they paid the least attention on text and question, which implies that the text also helps participants to answer the question. pictures possibly serve as a tool for question solving with consecutive task-oriented processing and text may be used for building and updating the mental model. these results correspond to the assumption of unequal usage of text and pictures in the integrative model of tpc. according to this model, pictures are mainly processed as an external representation to solve questions. with the initial task-oriented strategy, participants invested a large amount of time on pictures because participants may have used pictures as an external tool to scaffold question answering. they visited more frequently and spent longer time on text than on picture. this might be explained by coherenceformation being supported by the text. in our design, participants need to process both text and picture to get the correct answer. although participants were instructed with questions in initial task-oriented processing, they still needed to understand the text and the picture, thus also constructing a mental model. therefore, text and pictures were both used with initial task-oriented processing. also, time to the first fixation suggested that text and pictures drew participants’ interest with initial task-oriented processing. learners might integrate text and picture to build the initial mental model, as assumed in the integrative model of tpc. similar patterns were shown for transitions between text and pictures with initial task-oriented strategy. participants transferred their attention most frequently between text and picture, less between picture and question and the least between text and question with the initial task-oriented strategy. on the one hand, participants integrated text and picture with initial task-oriented strategy. on the other hand, more transitions between picture and question than between text and question support our assumption that picture is primarily used to solve the question. this result also corresponds to the assumption that text is more likely to be used for general coherence-formation processing and picture is especially used for specific task-oriented processing. summarizing our results, we found that text processing and picture processing differ substantially when learners are exposed to text-and-picture material. the differences are moderated by processing strategies triggered by context factors such as presentation of the question prior to vs. after the first exposure to the text-and-picture material. likely, text is primarily used for coherence-formation processing. it assists f. zhao et al 56 | f l r learners to comprehend the content of the materials, which generate an initial mental model or coherent semantic representation. picture is likely used for task-oriented specific processing. when learners have constructed the initial mental model, picture is mainly used as an external representation to update the mental model and to answer the question. when learners have the tasks beforehand, picture might serve mainly to scaffold the initial mental model construction. future studies should provide more evidence for the link between (a) differences in fixation patterns elicited by different processing strategies and (b) the formation and usage of mental models integrating text and picture information. in conclusion, returning to the hypotheses posed at the beginning of this study, it is now possible to state that text processing differs fundamentally from picture processing and that this difference is moderated by different reading strategies. more specifically, this study suggested that text is mainly used to build the mental model with coherence-formation general strategy. comparatively, picture is more likely to guide readers to solve questions with task-oriented selective strategy. similar to the previous studies on text comprehension, reading strategies also influence the comprehension of text and picture. our findings expand the rickards and denner (1978) account of global vs. selective processing to the domain of mixed (picture and text) materials. the results suggest that eye tracking indicators can play a major role in assessing and scaffolding text and picture comprehension. eye tracking indicators might be used to assess whether the learner is following an approach suitable to the current context factors (i.e. presentation of the question prior to vs. after the first exposure to the text-and-picture material). based on such assessment, interventions guiding visual attention to areas relevant at the current processing stage can involve salient visual cues presented online during task processing (cf. rouinfar et al. 2014). keypoints according to eye tracking indicators, text processing differs from picture processing and this difference is moderated by processing strategies. a high emphasis on text when processing material before the question is known, suggests that text is mainly used to build a mental model in general coherence-formation processing. picture is more likely to guide readers when they need to solve a question with selective taskoriented processing. acknowledgments this study is part of the on-going bite project on text-picture integration, which is funded by the german science foundation (grant no: schn 665/3-1, schn 665/6-1). we appreciate all the help from student participants and student assistants involved in this study. we thank for the cooperation of parents, teachers and principals. we also express our gratitude to dr. loredana mihalca, dr. axel zinkernagel and dr. thorsten rasch for their suggestions during setting up the experiment and interpreting the data. f. zhao et al 57 | f l r appendix six materials displayed on tobii eye tracker (translated from german) [each topic has three levels of questions. they were analysed on item-response theory with a one-parametric logistic model (raschmodel): level 1 (beta = -1.15); level 2 (beta = +0.57); level 3 (beta = +1.39). we also carried out a rational task analysis, which showed the mapping procedures between text and picture: level 1 (1.25); level 2 (1.50); level 3 (3.00)]. 1. banana trade figure a1. banana trade. many people like to eat bananas. they are planted in countries like ecuador, costa rica or columbia and then exported to europe. undoubtedly, this is related to costs. the banana that you see in the picture costs just one euro. this euro includes... (a) salary of the farmers; (b) cost of the fertilizer; (c) cost for transportation to the harbour; (d) profit of the plantation owners; (e) tax for bananas; (f) cost for shipping; (g) profit of the wholesalers; (h) cost for storage; (i) profit of the retailers question level 1 how many cents can a retailer earn from a banana? (3 cent/ 21 cent/ 31 cent/ 7 cent) question level 2 who do people pay the least if they buy a banana? (profit of the wholesalers/ profit of the plantation owners/ profit of the retailers/ salary of the farmers) question level 3 if we compare the farmers, retailers and wholesalers, then... (farmers earn the most, wholesalers earn less and retailers earn the least from a banana/ retailers earn the most, wholesalers earn less and farmers earn the least from a banana/ retailers earn the most, farmers earn less and wholesalers earn the least from a banana/ wholesalers earn the most, farmers earn less and retailers earn the least from a banana) f. zhao et al 58 | f l r 2. legs of insects figure a2. legs of insects. the legs of insects presented in figures a to d have the same structure: hip (orange ), leg ring (brown ), thigh (green ), bar (pink ), and foot (blue ). the legs are primarily organs for movement, which can be used for: running (a), swimming (b) or jumping (c). however, they can also be used for cleaning (d). question level 1 which insect has a cleaning leg? (ant/ honey bee/ water bug/ grasshopper) question level 2 which type of leg has the longest leg ring? (leg for cleaning/ leg for jumping/ leg for running/ leg for swimming) question level 3 does the leg for cleaning compared to the leg for jumping have… (a longer bar but a shorter foot/ a thicker thigh but a thinner leg ring/ a longer foot but a shorter bar/ a shorter foot but a longer leg ring) f. zhao et al 59 | f l r 3. auditory figure a3. auditory. tones and sounds are sound-waves. the faster the sound-vibrations the higher we perceive the sound/tone. the human ear can differentiate sounds/tones with low vibrations (20) and high vibrations (20 000) per second. the number of vibrations per second is called frequency; its unit is hertz (hz). a ten-year-old child is able to hear every sound/tone between the frequencies of 20 and 20000 hertz. this area is called the hearing range, which is displayed in blue in the picture. furthermore, a ten-year-old child is able to produce sounds/tones, for example by speaking, which are between 70 and 1000 hertz. this area is called vocal range, which is displayed in pink the picture. the illustration shows the hearing and vocal ranges of different species. question level 1 which of the following species are able to perceive tones/sounds at 120 hertz? (dog/ cricket/ cat/ bat) question level 2 which of the following species has a vocal range for producing the lowest tone? (human being, 45 year-old/ cat/ dog/ cricket) question level 3 which of the following four species is able to hear tones below 100 hertz as well as produce tones below 1000 hertz and above 1500 hertz? (human being, 10 year-old/ human being, 45 years old/ cricket/ cat) f. zhao et al 60 | f l r 4. pregnancy figure a4. pregnancy. the child develops in the uterus, or uterine wall (1). from the fourth month of gestation it is called a fetus (2). the fetus is nourished by the placenta (3). it is here where exchange between the blood vessels of the fetus (8) and the blood vessels of the mother (9) takes place (look at zoomed area). in the blood vessels, nutrients and oxygen (o2) as well as waste products and carbon dioxide (co2) are exchanged. the fetus is connected to the mother by the umbilical cord (4). the amniotic fluid protects the child. it could be called a protective-pillow for the fetus, because it helps to cushion the fetus from impact. when the mother’s water membrane (7) has broken, the delivery process is initiated and the fetus will move through the cervix (5) during the labour. question level 1 what is the name of the pink area? (placenta/ umbilical cord/ uterine wall/ amniotic fluid) question level 2 which parts do not directly link to each other? (blood vessels of the mother and blood vessels of the child/ cervix and amniotic fluid/ umbilical cord and water membrane/ placenta and uterine wall) question level 3 which path does the blood of the fetus take after getting nutrition and oxygen (o2) from the mother? it flows … (back through the placenta and by umbilical cord to the fetus/ back through the placenta and by the blood vessels of the fetus to the water membrane/ back through the placenta and by blood vessels of the fetus to the uterine wall/ back through the placenta and by blood vessels of the mother directly to the fetus) f. zhao et al 61 | f l r 5. map of europe figure a5. map of europe. the big map shows the continent of europe. actually, europe is not an independent continent. together with asia, it forms the continent “eurasia”. in the south, west and north, the border of europe is clearly defined by the seas. the delineation in the east is more difficult because there are no natural borders. an agreement set the borders at the ural mountains and further south, so parts of russia and kazakhstan belong to both europe and asia. the states of europe are divided into different districts according to economic and geographic features. these subspaces are: northern europe (blue) western europe (dark green) central europe (light green) southern europe (red) eastern europe (yellow) in figure a, one box represents one unit of the european area. in figure b, one box represents one unit of the european population. question level 1 in which subspace are countries located that belong to both europe and asia? f. zhao et al 62 | f l r (northern europe/ southern europe/ middle europe/ eastern europe) question level 2 which subspaces have the same number of area units? (eastern europe and northern europe/ southern europe and south-eastern europe/ middle europe and southern europe/ west europe and middle europe) question level 3 which subspace has one more unit of area, but four fewer units of population, when compared to western europe? (middle europe/ northern europe/ southern europe/ south-eastern europe) f. zhao et al 63 | f l r 6. savannah figure a5. savannah. because there are different rainy seasons, the savannah is differentiated into separate types. the city of enugu is located in the wet savannah; accra and ouagadougou are located in a dry savannah and the city of sinder in a thorn-bush savannah. depending on the amount of rainfall (e.g. ouagadougou 887mm rainfall per year) different food is cultivated and plants exported in each region and city. the months with enough rain for the respective plants to grow are shown in the diagrams by the areas with blue stripes above the red lines for temperature. if the line for rainfall (blue) is above the line for temperature (red), then there is more rainfall than evaporation. each month is represented by a letter in the lower part of the diagrams, for example f = february. the following plants need different amounts of rainfall per year for ideal growth: millet: 180 to 700 mm manioc: 500 to 2000 mm yams: more than 1500 mm peanut: 250 to 700 mm cotton: 700 to 1500 mm question level 1 what is the amount of rainfall per year in accra? (887 mm/ 1661 mm/ 787 mm/ 549 mm) question level 2 which plant can grow well in enugu? (peanut/ yams/ cotton/ millet) f. zhao et al 64 | f l r question level 3 which plants will grow well in the most cities, respective of the type of savannah? (millet/ manioc/ cotton/ peanut) f. zhao et al 65 | f l r references anderson, r.c., & pearson, p.d. (2002). a schema-theoretic view of basic processes in reading comprehension. in p. d. pearson, b. rebecca, m. l. kamil & p. mosenthal (eds.), handbook of reading research (pp. 255-291). mahwah, nj: lawrence erlbaum. ainsworth, s.e., (1999) a functional taxonomy of multiple representations. computers and education, 33(2/3), 131-152. doi:10.1016/s0360-1315(99)000299 andre, t. (1979). does answering higher-level questions while reading facilitate productive learning? review of educational research, 49(2), 280-318. doi: 10.3102/00346543049002280 eitel, a., scheiter, k., schüler, a., nyström, m., & holmqvist, k. (2013). how a picture facilitates the process of learning from text: evidence for scaffolding. learning and instruction, 28, 48-63. doi: 10.1016/j.learninstruc.2013.05.002 frase, l.t. (1967). learning from prose material: length of passage, knowledge of results, and position of questions. journal of educational psychology, 58(5), 266-272. doi: 10.1037/h0025028 frase, l.t. (1968). effect of question location, pacing, and mode upon retention of prose material. journal of educational psychology, 59(4), 244-249. doi: 10.1037/h0025947 friedman, a., & liebelt, l. s. (1981). on the time course of viewing pictures with a view towards remembering. in d. f. fisher, r. a. monty & j. senders (eds.), eye movements: cognition and visual perception (pp. 137-155). hillsdale, nj: lawrence erlbaum associates. gaschler, r., marewski, j. n., & frensch, p. a. (in press). once and for all – how people change strategy to ignore irrelevant information in visual tasks. quarterly journal of experimental psychology. doi: 10.1080/17470218.2014.961933 godau, c., wirth, m., hansen, s., haider, h., & gaschler, r. (2014). from marbles to numbers – estimation influences looking patterns on arithmetic problems. psychology, 5, 127-133. doi: 10.4236/psych.2014.52020 goldberg, j. h., & kotval, x. p. (1999). computer interface evaluation using eye movements: methods and constructs. international journal of industrial ergonomics, 24(6), 631-645. doi: 10.1016/s01698141(98)00068-7 hamilton, r. j. (1985). a framework for the evaluation of the effectiveness of adjunct questions and objectives. review of educational research, 55(1), 47-85. doi: 10.3102/00346543055001047 hegarty, m. (1992). mental animation: inferring motion from static displays of mechanical systems. journal of experimental psychology: learning, memory, and cognition, 18(5), 1084-1102. doi: 10.1037/02787393.18.5.1084 heller, k.a. & perleth, c. (2000). kft 4-12+r. kognitiver fähigkeitstest für 4. bis 12. klassen, revision. göttingen: beltz test gmbh. hochpöchler, u., schnotz, w., rasch, t., ullrich, m., horz, h., mcelvany, n., & baumert, j. (2012). dynamics of mental model construction from text and graphics. european journal of psychology of education, 1(22). doi: 10.1007/s10212-012-0156-z holmqvist, k., nyström, m., andersson, r., dewhurst, r., jarodzka, h., & weijer, j. v. d. (2011). eye tracking: a comprehensive guide to methods and measures. oxford; new york: oxford university press. jacob, r. & karn, k. s. (2003). commentary on section 4. eye tracking in human-computer interaction and usability research: ready to deliver the promises. in j. hyönä, r. radach, & h. deubel (eds), the mind’s eye: cognitive and applied aspects of eye movement research (pp. 573-605). oxford: elsevier science. johnson, c. i., & mayer, r. e. (2012). an eye movement analysis of the spatial contiguity effect in multimedia learning. journal of experimental psychology: applied, 18(2), 178-191. doi: 10.1037/a0026923 just, m. a., & carpenter, p. a. (1980). a theory of reading: from eye fixations to comprehension. psychology review, 87, 329-354. doi: 10.1037/0033-295x.87.4.329 kintsch, w., & van dijk, t. a. (1978). toward a model of text comprehension and production. psychological review, 85(5), 363-394. doi: 10.1037/0033-295x.85.5.363 f. zhao et al 66 | f l r kress, g. r., & leeuwen, v. t. (1996). reading images: the grammar of visual design. london; new york: routledge. larkin, j. h., & simon, h. a. (1987). why a diagram is (sometimes) worth ten thousand words. cogs cognitive science, 11(1), 65-100. doi: 10.1111/j.1551-6708.1987.tb00863 mcnamara, d. s. (ed.). (2007). reading comprehension strategies: theories, interventions, and technologies. mahwah, nj: lawrence erlbaum associates publishers. mayer, r. e. (1989). models for understanding. review of educational research, 59(1), 43-64. mayer, r. e. (2005). cognitive theory of multimedia learning. in r. e. mayer (ed.), the cambridge handbook of multimedia learning (pp. 31-48). cambridge, u.k.; new york: cambridge university press. paivio, a. (1986). mental representations: a dual coding approach. new york: oxford university press. rayner, k. (1998). eye movements in reading and information processing: 20 years of research. psychological bulletin, 124(3), 372-422. doi: 10.1037/0033-2909.124.3.372 rickards, j. p. (1979). adjunct postquestions in text: a critical review of methods and processes. review of educational research, 49(2), 181-196. doi: 10.3102/00346543049002181 rickards, j. p., & denner, p. r. (1978). inserted questions as aids to reading text. instructional science, 7(3), 313-346. doi: 10.1007/bf00120936 rothkopf, e. z. (1964). learning and the educational process; selected papers from the research conference on learning and the educational process, held at stanford university, june 22-july 31, 1964. in j. d. krumboltz (ed.), research conference on learning the educational process (pp. 193-221). chicago: rand mcnally. rothkopf, e.z. (1966). learning from written instructive materials: an exploration of the control of inspection behavior by test-like events. american educational research journal, 3(4), 241-249. doi: 10.3102/00028312003004241 rouet, j.-f. (2006). question answering and document search. in j.-f. rouet (ed.), the skills of document use. from text comprehension to web-based learning (pp. 93-121). mahwah, nj: erlbaum. rouet, j.-f., & britt, m. a. (2011). relevance processes in multiple document comprehension. in m. t. mccrudden, j. p. magliano & g. schraw (eds.), text relevance and learning from text. (pp. 19-52). charlotte, nc, us: iap information age publishing. rouinfar, a., agra e., larson, a. m., rebello, n. s., & loschky, l. c. (2014). linking attentional processes and conceptual problem solving: visual cues facilitate the automaticity of extracting relevant information from diagrams. frontiers in psychology, 5:1094. doi:10.3389/fpsyg.2014.01094 schmidt-weigand, f., kohnert, a., & glowalla, u. (2010). explaining the modality and contiguity effects: new insights from investigating students' viewing behaviour. applied cognitive psychology, 24(2), 226-237. doi: 10.1002/acp.1554 schnotz, w., & bannert, m. (2003). construction and interference in learning from multiple representation. learning and instruction, 13(2003), 141-156. doi: http://dx.doi.org/10.1016/s0959-4752(02)00017-8 schnotz, w., ullrich, m., hochpöchler, u., horz, h., mcelvany, n., schröder, s., & baumert, j. (2011). what makes text-picture integration difficult? a structural and procedural analysis of textbook requirements. ricerche di psicologia, 1, 103-135. doi: 10.1007/s10212-011-0078-1 tversky, b. (2001). spatial schemas in depictions. in m. gattis (ed.), spatial schemas and abstract thought (pp. 79-112). cambridge, mass.: mit press. winn, w. (1989). the role of graphics in training documents: toward an explanatory theory of how they communicate. ieee trans. profess. commun. ieee transactions on professional communication, 32(4), 300-309. doi: 10.1109/47.44544 zwaan, r., radvansky, g., hilliard, a., & curiel, j. (1998). constructing multidimensional situation models during reading. scientific studies of reading, 2(3), 199-220. doi: 10.1207/s1532799xssr0203_2 spe frontline learning research vol.4 no. 4 special issue (2016) 39 47 issn 2295-3159 corresponding author: crina i. damşa, department of education, university of oslo, po box 1092 blindern, 0317 oslo, norway. email: crina.damsa@ils.uio.no. doi: http://dx.doi.org/10.14786/flr.v4i4.208 revisiting learning in higher education—framing notions redefined through an ecological perspective crina damsa, alfredo jornet university of oslo, norway article received 15 september / revised 8 january / accepted 22 february / available online 10 january abstract this article employs an ecological perspective as a means of revisiting the notion of learning, with a particular focus on learning in higher education. learning is reconceptualised as a process entailing mutually constitutive, epistemic, social and affective relations in which knowledge, identity and agency become collective achievements of whole ecosystems. this conceptualisation implies that learning involves a trans-contextual and multimodal process, in which both learners and their social and material environments change. this article examines the implications of an ecological perspective on framing notions central to learning and current educational research, namely (a) knowledge co-construction and epistemic agency, (b) the role of (material) knowledge resources in the learning process and (c) the trans-contextuality that characterises learning in today’s knowledge society. the discussion concludes by identifying prospects that an ecological perspective offers to education and research on learning in higher education. the insights emerging from this reconceptualisation imply changes in the ways we can enhance and analytically account for the transformative potential of education. they also indicate the necessity for further advancing our understanding of learners’ ways of assembling the epistemic spaces necessary to engage in meaningful learning, their agency in this process and their relationship with the (social and material) environment. keywords: higher education, ecological perspectives on learning, co-construction of knowledge, agency, materiality, trans-contextuality, transformative nature of learning mailto:crina.damsa@ils.uio.no http://dx.doi.org/10.14786/flr.v4i4.208 damsa et jornet | f l r 40 1. introduction this article revisits the notion of learning, particularly in the context of higher education, by discussing key ideas derived from an ecological perspective. such a perspective is timely, considering today’s complex epistemological, social and institutional context, in which interdependent links between human subjectivities, collective human cultures and their environments are becoming more visible. learning is no longer viewed as the mastering of a given subject; it involves being knowledgeable across a variety of contexts, with the ability to connect to remote knowledge resources, communities and (work) sites no longer bound to one particular physical context (carvalho & goodyear, 2015; säljö, 2010). an ecological perspective is of particular interest in higher education, in which changing societal contexts and knowledge dynamics are creating new, open-ended and often unexplored opportunities and challenges in helping learners create professional futures. we are particularly interested in the broadening connections emerging between learning settings in higher education and in other contexts, in which curricular crossovers between scholarly knowledge and professional practices or cross-boundary learning arrangements (such as internships) are frequent. more knowledge is needed about the kind of learning opportunities that emerge as students, educators, professionals and other actors enter into new social and material configurations that are essentially uncertain and open-ended (markauskaite & goodyear, 2014; richter et al., 2015).. the aims of this article are to elaborate on the ecological premises that underlie existing sociocultural, situative and sociomaterial approaches and to discuss the implications for learning research and practice. thus, in this article, we do not develop a distinctly new approach, but rather make visible and elaborate on essential premises that are common to these other frameworks but which often remain tacit or underdeveloped. most importantly, an ecological perspective conceives of learning as an irreducible, mutually constitutive set of relationships between individuals and their social and material environments. thoroughly following this conception leads to new insights about how learners and social contexts develop together, and it also challenges the remaining dualisms present in the literature. we begin this article by laying out the basic premises of an ecological perspective. drawing on an empirical study of project-based, collaborative learning in higher education, we then revisit notions important in current educational research, namely knowledge co-construction and agency, knowledge resources and materials and trans-contextuality. in each case, we discuss how a consideration of the ecological premises allows us to reconceptualise these notions, highlighting aspects of learning that are not so visible when these premises are implicit or simply ignored. 2. learning from an ecological perspective ecology—the study of the relationships of organisms among one another and to their environment— is a central subject in biology. it has also been important for some of the most influential theories on learning and education. these include vygotsky’s (2012) cultural-historical theory, dewey and bentley’s 1949/1999 views of knowing as entailing transactional—i.e., mutually transforming—relations between organisms and the environment, gibson’s (1979) ecological psychology and bateson’s (1972) theory of learning and communication. common to these otherwise disparate scholars are two interlocked postulates that contrast with the individualistic and constructivist theories still present in current research on learning in general and in higher education in particular: (a) learning is not a private, internal process, but involves transactions between people and their socio-material environment, in which both people and environments are transformed; and (b) learning involves not only intellectual dimensions, but also practical and affective ones. in learning, the entire person-in-setting is transformed. we elaborate on this viewpoint by highlighting arguments that propound an expansion of the damsa et jornet | f l r 41 sociocultural theory, emphasising its underlying relational ontology and transformative stance. accordingly, human subjectivity, intersubjective (i.e., social) exchange, and material practice and production are irreducible aspects of a three-fold dialectical system (stetsenko, 2008). the individual actively relates to the environment and other individuals, and those relations then come to form part of how a person relates to herself and how she comes to know and develop (vygotsky, 1987). however, at the same time, the individual engages in the production of new material conditions, and thus acts upon and changes the world so that ‘the individual could no longer be understood without his/her cultural means; and the society could no longer be understood without the agency of individuals who use and produce artefacts (engeström, 2001, p. 134). from an ecological perspective, learning involves not only epistemology—how we come to know things—but also, and most fundamentally, ontology (packer & goicoechea, 2000). that is, knowledge, knowing and knowledgeable action are not ontologically separated from human development, but are inherently related to it. viewed from this perspective, learning is not a process whereby stable, unchanging things become known by unchanging individuals. rather, learning comprises changes in the conditions of human life and activity, in which both individuals and environments change. dewey (1938/1997) captured this mutually transforming relation in the principle of the continuity of experience, in which, through experience of the world—and precisely because experience involves material and bodily engagement—the world changes, thus changing the conditions under which new experiences are had. this is a change that involves not only the intellect but the whole person and how one relates to oneself and to others. experience changes not only the way we intellectually know the world but also the way we affectively and perceptually relate to it (roth & jornet, 2014). although these ecological principles, which imply the primacy of the social ecosystem over the individual, might not be new to readers familiar with sociocultural and situative approaches, their implications are still under-developed in the context of educational research and practice (roth, 2015). 3. redefining key framing concepts from an ecological perspective to better understand how focus on the ecological premises described here contributes to reconceptualising learning, we revisit three notions that are important in current educational research and particularly important for higher education: (a) knowledge co-construction and agency, (b) knowledge resources and (c) trans-contextuality. we draw from a case involving groups of computer engineering students enrolled in an undergraduate introductory course in web design and development. the course included bi-weekly lectures in web development (e.g., html5, java programming languages), lab sessions and a four-week collaborative web design and development project as the main course assignment. the student groups in the course were to receive guidance from the teachers and had various knowledge resources at their disposal. the setting is particularly interesting because it illustrates ways in which higher education programmes are attempting to prepare students to enter professional domains and societal contexts. we focus on one specific group mainly because their active and sustained engagement in the collaborative project illustrates both opportunities and challenges associated with the learning processes. our description of the case is based on video-recordings of actual group interactions and group interviews. damsa et jornet | f l r 42 the focus group consisted of four male students with a genuine interest in software development. the group chose to design and develop a website for an external customer. their learning process was characterised both by opportunities and challenges, related to both the learning of new content (i.e., the programming language and its application) and to ways of thinking and working in the field of web development. on the one hand, the students organised themselves effectively and employed varied and unexpected resources, most of them beyond the formal course curriculum (textbook). they organised the project work by dividing tasks and then holding long face-to-face meetings, during which they discussed strategy, searched for resources, integrated individually programmed codes and fixed bugs. they worked iteratively on their software product by developing and refining (paper and digital) mock-ups, following the methods of experienced web developers, which they explored online or by talking to experts. the group used and engaged with some resources provided in the course and with external resources from the web development community (crowd-source online programming platforms). the feedback on the developing product and on project management was mainly provided by the customer. on the other hand, the students experienced difficulties in understanding the complexity of the task’s requirements. this often led to crashing prototypes and put pressure on the group’s interaction. the students indicated that they found the project very interesting, but that the requirements were not specific enough and discussion often went on in circles, without a productive loop. these emerging issues were solved through group discussions, trial and error and by using clues found online and in customer feedback. this unplanned feedback compensated for the relatively little guidance received by the group. although the group’s assessment of the assignment was positive, they received a lower grade than they expected. the students critiqued the complexity of the task, which combined technical and project management challenges, neither of which were made completely clear in the assignment guidelines. in addition, the students considered that identifying and addressing errors before grading could have been facilitated by more sustained guidance during the development process. 3.1 redefining knowledge co-construction and agency the notion of knowledge construction is widely used in educational research (e.g., schellens & valcke, 2006), most often to denote individuals’ formations of mental models and representations. in an attempt to overcome focus on the individual, higher education research has more recently used the notion of co-construction to denote processes that focus on collective participation in learning activities and on transforming the environment (damşa, ludvigsen, & andriessen, 2013; richter et al., 2015;). the case above clearly offers an example of such a co-construction process. the students worked jointly to create a product, and in this context, learning was the result not of an individual but of a social process, joint efforts and the resources involved. however, the notion of co-construction is often associated with a focus on how participants ‘negotiate’ meanings about given practices and topics (e.g., heo, kim, & kim, 2010). here, a triad formed by subject, object and meaning is maintained, in which each of the three elements remains selfcontained and therefore ontologically primary. co-construction thus turns the focus away from individual minds and toward joint group cooperation. however, in doing so, it still retains the ontological primacy of subjects and objects over the social, transformative process. latour (2013) eloquently depicts this limitation: ‘every use of the word construction … opens up an enigma as to the author of the construction’ (p. 158). in the case of the student group described above, every step in the development process opened up new paths of inquiry and development, but it also forced the group to face new problems and choices. by engaging with these problems and activities, they generated new ideas and conceptual artefacts (the programming code), which in turn represented new departure points in their endeavour (damşa & nerland, 2016). an ecological perspective gives primacy to the social, shifting the focus away from subjects and objects and pointing to a ‘a better appreciation of the material flows and currents of sensory awareness within which both ideas and things reciprocally take shape’ (ingold, 2011, p. 10). in the described case, the students damsa et jornet | f l r 43 develop a product together. yet, the product itself is in constant transformation, taking different forms (sketches, drawings and prototypes). the students themselves do not have a clear idea of what they are in the process of ‘co-constructing’, and much of what happens is not planned, but emerges. there is not just jointly knowing, but there is also being uncertain, a condition that nonetheless does not impede the students’ engagement in professional practices for which they have not yet developed expertise. as has been thematised in recent research on transfer taking and the ecological perspective, this participation is possible not because the individual students carry with them already formed understandings; rather, it is because there is an emergent constitutive order that cannot be attributed to the individual mind, but to an unfolding field of action (damşa et al., 2010; jornet, roth, & krange, 2016). here, the learner’s receptivity, affectivity and competence to engage in social relations with others is primary over their individual (intellectual) intentions. in analysing so-called co-construction events, the focus cannot be on either the constructing agents or the constructed objects, because both are constantly changing. this has implications for the notion of agency, which has not received adequate attention in higher education research (ashwin, 2008). some studies have begun to reconceptualise learning-related agency in terms of shared epistemic agency (damşa et al., 2010). these studies have considered the social–relational aspects of learning and how knowledgegenerating processes become more than an individual endeavour. transformative conceptions of agency are also being examined (e.g., kumpulainen, 2013; engeström, sannino, & virkukken, 2014). attention to the ecological premises creates the need to consider how collaborations involve affective and perceptual changes, in which learners are not only intellectual agents (how could they otherwise engage in practices for which they do not yet have the required knowledge?) but also subject to the performative and affective relations in which they engage (roth & jornet, 2014). in the case presented here, the students not only (co)construct but draw from and appropriate cultural resources that are not their own. an ecological approach should be able to account for the role of these resources in ways that do not reify individual (agent, subject)– tool (object, world) dualism. 3.2 redefining knowledge resources and materials traditionally, domain-related knowledge has been ‘translated’ into classroom curricula that emphasise conceptual knowledge and understanding, in which teaching materials are often seen as involving knowledge representations or tools (säljö, 2010). with the growth of ubiquitous information and communication technologies (icts), the range of knowledge resources available for teaching and learning has dramatically expanded. this is visible in the case above, in which the students did not turn to the textbook, but most often relied on online professional programming and/or social platforms. classical literature has it that learning involves a process of interpreting and decoding representations to solve an already given problem. not possibly knowing the domain and its practices in advance (learning these is the goal of the course), the students’ engagement with resources and materials was problemand world-forming. yet, how materials partake in the formation of students’ worlds (i.e., perceptions, knowledge or identity), that is, in processes of ontogenesis, is rarely discussed in connection to learning and research in higher education. an ecological perspective is in line with recent sociocultural conceptualisations that view materials as meaning or sense-making resources (säljö, 2010). according to this view, materials come to form integral part of thinking and doing through processes of sign formation (vygotsky, 1987). it is not, as often is implied, that materials and technology ‘mediate’ between learners and the world, which would maintain cartesian dualism (stetsenko, 2005). rather, materials become entangled with people’s lives and form new organs, which cannot be reduced to either learner (subject) or material (object, tool) (vygotsky, 1989). learners orient towards materials, which organise the participants’ perceptions and actions. at the same time, these actions transform the very materials that shaped them in the first place. accordingly, the ‘things’ of learning—that is, ‘teachers, learning activities and spaces, knowledge representations such as texts, pedagogy, curriculum content, and so forth’ (fenwick et al., 2012, p. 2)—cannot be taken for granted, but damsa et jornet | f l r 44 are seen as ‘themselves effects of heterogeneous relations’ (p. 2). in the case described above, the learning materials are diverse, but they certainly come to form an ecology that both organises and depends upon the organisation of the students’ joint work. by coming into contact with knowledge resources related to professional practice (through delivering to customers, searching for information in professional and crowdfunded fora and using already existing codes), the participants are not so much being mediated to access knowledge as they are developing habits, dispositions and forms of orienting towards the epistemic, digital and physical world of which they already form an integral part (damşa & nerland, 2016). 3.3 trans-contextuality an ecological perspective is relevant for explaining learning as a social and continuous process occurring across contexts and occasions. individualist approaches rely on notions of transfer of knowledge to explain how learners move knowledgeably across contexts (e.g., reed, 2012). boundary crossing has been formulated as an alternative notion to account for how moving across settings involves social and material (rather than only mental and abstract) processes (akkerman & baker, 2011). taking the perspective of the students in our case, however, no explicit boundaries were apparent between their university setting and the professional world, in which they were already participating in several ways (through meetings, online or in contact with customers). indeed, the students’ concerns emerged with respect to both the formulation of the task (an aspect of schooling practice with which they are familiar) and aspects of the professional programming practices that may be said to be beyond the university’s boundary. an ecological perspective challenges the notion of boundary, demanding instead an account that adequately describes the lines of becoming (intellectual, social and relational) that learners and materials together constitute and undergo. by crossing contexts, learners assemble an epistemic space (markauskaite & goodyear, 2014), in which individual and collective goals, needs and epistemological orientations develop, capitalising on teaching and guidance, resources and infrastructure. there are practical and methodological challenges associated with learning trajectories traversing time and space through (digital) technology (carvalho & goodyear, 2015; erstad, 2013). learners have more access to information from a multitude of sources; it is a ‘polyphonic’ (säljö, 2010), multi-contextual world. although it is generally considered beneficial for learning, capitalising on widely available and distributed knowledge, resources and tools is not a straightforward process (orlikowski, 2007). in our exemplary case, challenges and tensions emerged in relation to various aspects of the learning situation: the affordances offered by the state-of-the-art knowledge, practices and technologies, the students’ positioning in relation to the tasks and the domain, and the teaching and assessment practices within the institutional setting. however, from a perspective that takes the ecological premises laid out here seriously, these challenges and tensions cannot be the result of self-contained learners, which inter-act with self-contained resources and self-contained (i.e., contained within boundaries) practices. as higher education continues to develop outreach practices in which schooling goes into professional practice, and vice-versa, we no longer have a crossing of boundaries, but a new line of development within which different materials and subjectivities unfold. 4. concluding remarks: ecology and learning in higher education this contribution elaborates on considerations of learning as a set of mutually constitutive relationships among individual, institutional and societal contexts. the empirical material illustrates how, as a group of students engaged in actual relations in and across knowledge domains (e.g., the school, the professional field of software programming), there is not just co-construction of knowledge but also reconfiguration of their affective and relational orientations. while learning software development and damsa et jornet | f l r 45 programming, the students generated knowledge and re-enacted practices and objects (see stetsenko, 2005) that fed into and transformed their knowledge landscape and their personal horizons. in line with the transformational ontology posited by the ecological perspective, a reconsideration of the role of higher education involves preparing people to not only learn and adapt to existing knowledge, practices and environments but also to actively transform them (stetsenko, 2008). such a perspective has the potential to guide learners, education and society towards a notion of learning that accounts for the fluid elements of the epistemic, social and material-digital environments we are surrounded by and partake in (ingold, 2011; säljö, 2010). within this augmented learning context, the role of formal educational settings, such as higher education, entails more than simply organising learning and helping learners go through authoritative obligatory passage points (callon, 1984). it needs to offer resources to allow students become critical and productive participants, helping them manage their own learning and development trajectories. acknowledging that learning is an achievement of whole (eco-) systems, and not primarily of individuals alone, educational settings should orient not towards individuals but towards transformational potentials. if learning is not about acquiring knowledge but about changing the world, then providing tools and opportunities for that change should become primary. an ecological perspective thus addresses the need to view learning not from a normative perspective—i.e., in terms of the competences we want learners to achieve—but in terms of the life world of the learner, for whom the structures in the world (and the boundaries thereof) are not the same as that of the educator or researcher. the insights emerging from this reconceptualisation indicate the necessity for a more sophisticated and versatile account of the transformative potential of learning and for advancing our understanding of learners’ authoritative positioning and agency in learning (kumpulainen, 2013), their relationship with the (epistemic, social and material) environment and the way they assemble the epistemic space necessary to engage in meaningful and transformative learning. keypoints this article provides an ecological perspective for revisiting premises and notions fundamental to learning and development, in relation to higher education contexts, expands on current sociocultural and sociomaterial theories by proposing learning as a transformative process whereby both the learner and the environment change, and which entails development that is not only intellectual but also social and affective, builds on empirical material from a study of a higher education course aimed at bridging educational and professional contexts and challenges higher education to reconsider the premises for defining learning and to provide the appropriate framework for transformative learning to take place and for remaining dualistic viewpoints to be overcome. references akkerman, s. f., & bakker, a. (2011). boundary crossing and boundary objects. review of educational research, 81(2), 132–169. ashwin, p. (2008). accounting for structure and agency in ‘close-up’ research on teaching, learning and assessment in higher education. international journal of educational research, 47, 151–158. bateson, g. (1972). steps to an ecology of mind: collected essays in anthropology, psychiatry, evolution, and epistemology. chicago, il: university of chicago press. damsa et jornet | f l r 46 callon, m. (1984). some elements of a sociology of translation: domestication of the callps and the fishermen of st brieuc bay. the sociological review, 32, 196–233. carvalho, l., & goodyear, p. (2014). the architecture of productive learning networks. new york, ny: routledge. damşa, c. i., kirschner, p. a., andriessen, j. e. b., erkens, g., & sins, p. h. m. (2010). shared epistemic agency an empirical study of an emergent construct. journal of the learning sciences , 19(2), 143–186. doi:10.1080/10508401003708381 damşa, c., ludvigsen, s., & andriessen, j. (2013). knowledge co-construction – epistemic consensus or relational assent? in m. baker, j. andriessen, & s. jaarvela (eds.), affective learning together: social and emotional dimensions of collaborative learning (pp. 97-119). london, england: routledge academic publishers & taylor and francis group. damşa, c.i., & nerland, m. (2016). student learning through participation in inquiry activities: two cases from teaching and computer engineering education. vocations and learning, doi:10.1007/s12186-0169152-9 dewey, j. (1997). education and experience. new york, ny: touchstone. (original work published 1938) dewey, j., & bentley, a. f. (1999). knowing and the known. in r. handy & e. e. hardwood (eds.), useful procedures of inquiry (pp. 97–209). great barrington, ma: behavioral research council. (original work published 1949) engeström, y. (2010). expansive learning at work: toward an activity theoretical reconceptualization. journal of education and work, 14(1), 133-156. engeström, y., sannino, a., & virkkunen, j. (2014). on the methodological demands of formative interventions. mind, culture and activity, 21(2), 118–128. doi:10.1080/10749039.2014.891868 erstad, o. (2013). digital learning lives: trajectories, literacies, and schooling. new york, bern, berlin, bruxelles, frankfurt am main, oxford, wien, peter lang publishing group. fenwick, t., edwards, r., & sawchuk, p. r. (2012). emerging approaches to educational research: tracing the socio-material. london: routledge. gibson, j. j. (1979). the ecological approach to visual perception. boston, ma: houghton mifflin. heo, h., lim, k. y., & kim, y. (2010). exploratory study on the patterns of online interaction and knowledge co-construction in project-based learning. computers & education, 55, 1383–1392. ingold, t. (2011). redrawing anthropology: materials, movements, lines. aldershot, england: ashgate. ingold, t. (2015). the life of lines. london, england: routledge. jornet, a., roth, w.-m., & krange, i. (2016). a transactional approach to transfer episodes. journal of the learning sciences. doi:10.1080/10508406.2016.1147449 jornet, a., & steier, r. (2015). the matter of space: bodily performances and the emergence of boundary objects during multidisciplinary design meetings. mind, culture, and activity, 22, 129–151. kumpulainen, k. (2013). the legacy of productive disciplinary engagement. international journal of educational research, 64, 215–220. doi:http://dx.doi.org/10.1016/j.ijer.2013.07.006 latour, b. (2013). an inquiry into modes of existence: an anthropology of the moderns. cambridge, ma: harvard university press. markauskaite, l., & goodyear, p. (2014). professional work and knowledge. in s. billett, c. harteis, & h. gruber (eds.), international handbook of research in professional and practice-based learning (pp. 79– 106). dordrecht, netherlands: springer. orlikowski, w. (2007). sociomaterial practices: exploring technology at work. organization studies, 28, 1435–1448. doi:10.1177/0170840607081138 packer, j. m., & goicoechea, j. (2000). sociocultural and constructivist theories of learning: ontology, not just epistemology. educational psychologist, 35(4), 227–241. doi:10.1207/s15326985ep3504_02 reed, s. k. (2012). learning by mapping across situations. journal of the learning sciences, 21, 353–398. richter, c., allert, h., albrecht, j., & ruhl, e. (2015). grappling with the not-yet-known. in o. lindwall, p. häkkinen, t. koschman, p. tchounikine, & s. ludvigsen (eds.), exploring the material conditions of learning: the computer supported collaborative learning (cscl) conference 2015, volume 1 (pp. 284–291). gothenburg, sweden: the international society of the learning sciences. damsa et jornet | f l r 47 roth, w.-m. (2015). the primacy of the social and sociogenesis. integrative psychological and behavioral science. doi:10.1007/s12124-015-9331-5 roth, w.-m., & jornet, a. (2014). toward a theory of experience. science education, 98, 106–126. schellens, t., & valcke, m. (2006). fostering knowledge construction in university students through asynchronous discussion groups. computers & education, 46, 349–370. stetsenko, a. (2005). activity as object-related: resolving the dichotomy of individual and collective planes of activity. mind, culture, and social interaction, 12(1), 70–88. doi:10.1207/s15327884mca1201_6 stetsenko, a. (2008). from relational ontology to transformative activist stance on development and learning: expanding vygotsky’s (chat) project. cultural studies of science education, 3, 471–491. säljö, r. (2010). digital tools and challenges to institutional traditions of learning: technologies, social memory and the performative nature of learning. journal of computer assisted learning, 26, 53–64. doi:10.1111/j.1365-2729.2009.00341.x vygotsky, l. s. (1987). the collected works of l. s. vygotsky: vol. 1. problems of general psychology. new york, ny: plenum. vygotsky, l. s. (1989). concrete human psychology. soviet psychology, 27, 53–77. vygotsky, l. s. (2012). thought and language (rev. ed.). cambridge, ma: mit press. (original work published 1986) frontline learning research vol. 5 no. 3 special issue (2017) 139 154 issn 2295-3159 * corresponding author: a.m. williams, department of health, kinesiology, and recreation, college of health, university of utah, 250 s. 1850 e. rm 200, salt lake city, ut 84112. email: mark.williams@health.utah.edu. doi: http://dx.doi.org/10.14786/flr.v5i3.267 using the ‘expert performance approach’ as a framework for improving understanding of expert learning a. mark williamsa*, bradley fawvera, & nicola j. hodgesb a university of utah, usa b university of british columbia, canada article received 15 august / revised 14 december / accepted 23 march / available online 14 july abstract the expert performance approach, initially proposed by ericsson and smith (1991), is reviewed as a systematic framework for the study of ’expert’ learning. the need to develop representative tasks to capture learning is discussed, as is the need to employ process-tracing measures during acquisition to examine what actually changes during learning. we recommend the use of realistic retention and transfer tests to infer what has been learned, so that the effects of various interventions on learning may be evaluated. a focus on individual differences in learning within groups of expert performers is considered as a way to identify the characteristics of more efficient and effective learners. the identification and study of expert (or good) learners will enhance our understanding of skill acquisition and how this may be promoted using instructional interventions and practice opportunities. although these ideas are predicated on our research on perceptual-cognitive expertise in sport, we argue that they have general merit beyond this domain. the challenge for scientists is to generate new knowledge that helps those involved in developing learners who can acquire and refine skills more efficiently and effectively across professional domains. keywords: perceptual-cognitive training; skill acquisition; representative task-design; process-tracing measures; deliberate practice http://dx.doi.org/10.14786/flr.v5i3.267 williams et al | f l r 140 1. introduction over recent decades there has been significant growth of interest in the study of expert performance, with a specific focus on perceptual-cognitive expertise (for up-to-date reviews, see baker & farrow, 2015; ericsson, hoffman, aaron, & williams, in press; farrow, baker & macmahon, 2013; hodges & williams, 2012). the growth of interest in topics such as anticipation and decision making has spanned multiple domains including sport (williams & abernethy, 2012), medicine (mcrobert, causer, vasiliadus, watterson, & williams, 2013), aviation (kennedy, taylor, reade, & yesavage, 2010), and automobile driving (stahl, donmez, & jamieson, 2016). the typical finding is that experts can be differentiated from less expert or novice counterparts based on a number of domain-specific, perceptual-cognitive skills. these skills include the ability to pick up biological motion information from the movements of others (e.g., abernethy & zawi, 2007), a capacity to identify familiarity or patterns in structured displays (e.g., north, ward, ericsson, & williams, 2011), and a more refined knowledge of likely situational or event probabilities (e.g., ward, ericsson, & williams, 2013). systematic differences have been reported in the gaze behaviours underpinning these perceptual-cognitive skills, with the specific strategies employed being task and context specific (see roca, ford, & williams, 2013). in contrast, evidence pointing towards individual differences in measures of basic visual and cognitive functions (i.e., domain-generic skills) is relatively weak or at best mixed (voss, kramer, basak, prakash, roberts, 2010). the suggestion is that expertise arises through adaptations that are mostly specific to the target domain of expertise (ericsson & kintsch, 1995). in this paper, we consider how research on expertise in sport, particularly related to perceptual-cognitive expertise, could impact how we study expert performance in general and ‘expert’ or good learning more specifically. we argue that a distinction between expert performance and expert learning could provide significant insight into the processes underpinning skill acquisition. although the number of published reports on expertise has grown substantially, there remain significant shortcomings with existing literature. first, while the vast majority of researchers have used the traditional expert-novice paradigm to examine differences as a function of performance, there has been a paucity of published papers focusing on expertise within the context of learning in contrast to expertise in terms of performance on the task itself. there are numerous questions that remain unanswered. is it true that expert performers are always expert learners? what is the relationship between performance and the ability or receptiveness to acquire skills efficiently and effectively? are there individual differences with respect to how skills are acquired and refined that distinguish individuals even within skill groups? what is the relationship between skill acquisition processes, for component skills, and eventual skilled performance/expertise? in this review article, the merits of using the ‘expert performance approach’ as a framework to study expert learning rather than, or as well as, expert performance are highlighted. generally, research into skill development and learning has lacked a systematic framework (including methods and measures) for the study of perceptual-cognitive expertise and, equally importantly, its acquisition. while the expert performance approach is not new, having originally been proposed more than two decades ago (see ericsson & smith, 1991), it does offer a framework to study how skills are learned as well as performed. we acknowledge that this approach puts a focus on the acquisition of new skills, rather than their refinement, with the supposition that intentional learning and factors related to skill learning are likely different to the background corrections (bernstein, 1967/1996) and refinements in skill that might take place at a more implicit/non-conscious or subcortical level. in the following sections, the three stages outlined in the expert performance approach are introduced with some examples of how the framework has been used successfully to evaluate expertise across domains. in the second half of the article, the focus shifts to examining how the expert performance approach may be used to evaluate expert learning rather than expert performance. our intention is that some of the questions, issues and ideas we raise in this paper will encourage those working in the learning sciences to adopt this approach (or at least consider these issues) as they pursue further research in this area. williams et al | f l r 141 2. the expert performance approach the expert performance approach was first presented by ericsson and smith in 1991. the three-stage approach to studying expert performance is highlighted in figure 1, with each stage described in detail below. figure 1. the expert performance approach proposed by ericsson and smith (1991). adapted from williams and ericsson (2005). 2.1 stage 1 – capturing expert performance: what factors discriminate? in the first stage, the aim is to develop a representative task that enables expert performance to be captured in a reliable and objective manner. performance may be captured under controlled conditions in the laboratory or using appropriate measurement systems in the field (williams & ericsson, 2005). the goal is to develop tests that discriminate at an empirical level those who are skilled from those less skilled on the task(s). the stable and reliable aspects of performance are captured in a repeated manner. initially, researchers focused primarily on cognitive tasks representative of performance in domains such as chess or mathematics. more recently, there have been efforts to determine representative tasks for both perceptualcognitive and perceptual-motor skills (williams, ford, eccles, & ward, 2011). for example, in sport, these latter efforts have focused both on capturing performance in the field using player and motion tracking systems or by recreating realistic simulations of performance situations using filmor virtual-reality simulations (williams & abernethy, 2012). figure 2 presents a typical example of a film-based simulation set up designed to capture anticipation and decision making in tennis. the advantage with such methods is that they enable performance on the task to be measured accurately such that groups of participants varying in expertise may be compared under standardised and reproducible test conditions. williams et al | f l r 142 2.2 stage 2 – identifying mechanisms underpinning expert performance: how do experts perform better? in the second stage, the aim is to identify the processes and mechanisms underpinning superior performance. process-tracing measures such as eye movement recording and think-aloud verbal protocols are employed to examine the perceptual-cognitive processes underlying expert performance (e.g. mcrobert, ward, eccles, & williams, 2011; roca, ford, mcrobert, & williams, 2013). neurophysiological techniques, such as transcranial magnetic stimulation (tms), electroencephalography (eeg) and functional magnetic resonance imaging (fmri) can be employed to identify and relate areas of brain activity to performance (e.g., balser et al., 2014; dennis, rowe, williams, & milne, in press; wright, bishop, jackson, & abernethy, 2010, 2011). psycho-physiological measures such as pupil diameter, inter-beat heart rate intervals, electromyography (emg), and galvanic skin response may be recorded to examine changes in mental effort or stress/anxiety as a function of skill on the task (e.g., vater, roca, & williams, 2015). biomechanical measurement tools can also be employed to record kinetic and kinematic variables that give insight into movement planning and execution processes (e.g., müller, brenton, dempsey, harbaugh & reid, 2015; smeeton & williams, 2012). finally, various behavioural manipulations, such as dual-task paradigms, can be used to determine the type of processes engaged during anticipation and decision making (e.g., mulligan, lohse & hodges, 2016a,b) or the cognitive and attentional load associated with a task or skill level (e.g., broadbent, causer, williams, & ford, in press). a combination of different process-tracing measures may be needed to cross-validate findings and to provide a more complete picture of the important processes that mediate expert performance. the challenge is to identify how experts demonstrate superior performance on the task compared to less expert individuals with the overall aim of enhancing conceptual and empirical understanding of performance processes. 2.3 stage 3 – facilitating the acquisition of expert performance: why has expertise developed? the aim in the final stage is to improve understanding of how adaptations occur during the acquisition of expertise and to use the knowledge acquired to develop training interventions that help facilitate the more rapid acquisition of expertise. in regards to understanding how these adaptations occur, the prototypical approach has been to use qualitative and quantitative methods to probe practice histories; such as questionnaires, interviews and practice logs. the deliberate practice theory (ericsson et al., 1993) is usually used as the framework around which empirical questions are generated and related (e.g., see ward, hodges, starkes, & williams, 2007; ford & williams, 2012). preto post-test intervention designs have also been used to study the efficacy of shortto medium-term training interventions on performance and learning. such interventions have focused on determining whether the practice behaviours of experts differ from less skilled individuals and how these behaviours relate to current understanding of best practice methods (e.g., coughlan, ford, mcrobert, & williams, 2014; hodges, edwards, luttin, & bowcock, 2011). other researchers have studied how the perceptual-cognitive skills that underpin anticipation and decision making can be trained using simulationor field-based interventions (e.g., see broadbent, causer, williams, & ford, 2015a; williams, ward, & chapman, 2003; williams, ward, knowles, & smeeton, 2002). the general intention is to develop interventions that facilitate the more rapid and robust acquisition of skills, based on knowledge of what good learners or elite performers do, as well as understanding of how and why this works. williams et al | f l r 143 figure 2. an illustration of a film-based simulation used to evaluate anticipation and decision making in tennis. the participant is required to anticipate the location and type of shot played by the opponent located at the net and to decide on an appropriate response. verbal or motor responses are recorded (both speed and accuracy), with the latter from pressure sensitive mats placed on the floor around the participant. 3. applying the expert performance to the study of expert learners the expert performance approach has been used to study differences between expert and less expert or novice performers. however, the approach has not been used systematically to evaluate how skilled individuals acquire (and refine) skills. there are many factors which are likely to affect the rate of acquisition and retention of skills, including the dispositional characteristics of the performer and exposure to previous learning experiences (e.g., hodges et al., 2014), existing skill sets (e.g., hodges et al., 2011), motivation (wulf, shea, & lewthwaite, 2010), and potentially, individual differences related to intelligence or other factors (e.g., ackerman, 1987). the expert performance approach offers a framework to address issues concerning how experts learn and engage in practice to acquire new skills, characteristics that define good learners with respect to factors such as efficiency in rate of skill acquisition and potentially transferability of skills to new conditions, as well as to assess relationships between current performance and learning. as such, by studying how expert performers learn and what factors distinguish the best learners from the worst, we might learn more about best practice principles for more elite performers and about processes of learning in general that could impact conceptual understanding and applied interventions. 3.1 how can we capture learning and how do we define ‘expert’ and ‘less expert’ learners? learning is differentiated from performance during practice on the grounds that the latter is evaluated through observation of current behaviour, whereas estimates of learning usually involve both a measure of skill retention (to measure the longevity of the change in performance) and an assessment of transfer to probe what has been learned and the robustness of learning (schmidt & lee, 2013). typically, learning is assessed at least 24 hours after the end of practice (i.e., after a period of sleep consolidation, walker et al., 2002), under equitable conditions for all groups (such as in the absence of feedback). thus far, however, limited attention has been paid to the efficiency of learning and how it differs within or between individuals. the general focus has been on testing how various interventions facilitate, or not, learning at the group rather than individual level. in fact, the process of collating mean scores across both trials and individuals and summating these largely eliminates what may be interesting variability in the data, which may be a functional component of learning (davids, bennett, & newell, 2006). williams et al | f l r 144 the rate of learning may be negatively related to long-term retention. for example, in short-term adaptation studies, where individuals practice making novel aiming movements in visually-rotated conditions, there is evidence for both fast and slow adaptation processes, with the slow process being more robust over time and arguably predictive of better learning (smith, ghazizadeh & shadmehr, 2006). this slow process is more implicit in nature, showing little sensitivity to error, in comparison to the fast process that is more explicitly-driven and highly responsive to errors. in a study of sequence learning, where individuals practiced three different sequences in a serially repeating order, there was evidence that slower learners, or the ones who spent more time in what was termed the cognitive phase of skill acquisition, had better retention (wadden, hodges, de asis, neva & boyd, 2017). these data seem to support an efficiencyeffectiveness trade-off in learning, which is underscored by current conceptualizations of practice and learning relating to the need for ‘challenge points’ during practice (guadagnoli & lee, 2004) and the promotion of desirable difficulties (bjork & bjork, 2011). however, there are individuals that show both fast acquisition and good retention and “good” learners do not always adhere to established principles of practice. for example, when learners are allowed to schedule their own practice, a useful method to assess characteristics of good or expert learners, the best learners are not always the ones who adhere to good practice principles (such as high contextual interference between attempts of different skills). as long as practice is self-determined, low levels of switching (i.e., contextual interference) between to-be-acquired skills can produce good learning, without accuracy costs in acquisition (hodges et al., 2011, 2014; keetch & lee, 2007). also, there has been resistance to methods which have a negative impact on performance in the short-term (in order to benefit retention), as these can discourage change, demotivate learners and of course have a cost function in terms of amount or duration of practice (especially pertinent when practice time is limited or safety concerns are at play, see lee & wishart, 2005). systematic investigations of experts acquiring new skills (or studying good learners who are able to circumvent practice time) may help to give us insights into this efficiency-effectiveness relationship and any individual differences that might impact these variables. in order to enhance measurement sensitivity, it is necessary to consider the variability in learning across participants to determine how individuals differ in the amount of learning that has occurred (perhaps in response to varying instructions or demonstrations). such an analysis is important if we are to better understand the subtle differences that may exist between ‘expert’ and ‘less expert’ learners. one approach would be to classify performers based on the amount of learning that has occurred by looking at either absolute retention, change in performance (from preto retention tests) or learning efficiency, with respect to the relationship or ratio between rate of acquisition and retention. with respect to this latter measure, those with the best ratio (i.e., efficient and effective) would be classified as a more ‘expert learner’ (or the best within the group). if a sufficiently large sample is used, more and less expert learning groups may be created using a quartileor median-split approach or alternatively, regression analyses may be used, with all participants included, to examine changes in performance and how these ultimately may be linked to changes in process measures. such an approach has been used previously to stratify participants into high and low performing individuals based on their perceptual-cognitive expertise (e.g., bourne, bourne, bennett, hayes, & williams, 2011; savelsbergh, van der kamp, williams, & ward, 2006; williams, ward, bell-walker, & ford, 2011), yet, thus far, the approach has not been employed to stratify participants based on their proficiency in learning a skill. it is often difficult to assess performance directly in many domains, making it hard to measure the amount of improvement that occurs during learning. when performance can be evaluated using standard units of assessment (e.g., time and distance), it is relatively straightforward to evaluate learning. yet, in domains such as in team sports, the military and emergency room medicine, performance is hard to measure as it is not expressed through a single unit of measurement. in such scenarios, performance is typically made up of several individual components that could interact. as such, particularly when the components are more independent, a sensitive assessment could be provided through measurement of an isolated component of performance. for example, a specific test of decision making could be designed, such that performance on this component can be isolated for assessment in the laboratory. a within-task criterion (i.e., actual score on williams et al | f l r 145 the test) may then be used to identify those who are high or low on decision making, rather than general performance in the domain. an advantage of this method is that the efficiency and effectiveness of learning on a specific component of performance may be identified, allowing identification of participants who are more or less able to learn and improve on that component. 3.2 identifying the processes and mechanisms that mediate expert learning when conducting traditional experimental work on skill learning, the vast majority of researchers have relied almost exclusively on changes in outcome measures of performance. process-tracing measures, although used, have been employed far less frequently. the difficulty in relying on outcome measures is often we have limited understanding of what actually changes during learning and neglect the fact that positive changes in process may not necessarily be (immediately) reflected in changes in outcomes (schmidt & lee, 2013). eysenck and calvo (1995), in their processing efficiency theory, present a conceptual account of the effects of anxiety on performance. the model differentiates between changes in processing efficiency and processing effectiveness. as participants become anxious often there is no immediate change in effectiveness, but there may be an increase in the amount of mental effort or resources that need to be devoted to the task, thereby decreasing performance efficiency. at higher levels of anxiety, the cognitive resources needed for task execution may exceed available resource capacity, leading to declines in both efficiency and effectiveness. we argue that the process of learning may function in a similar way such that the efficiency and effectiveness of learning may not necessarily be highly correlated. a call is therefore made for the use of process-tracing measures in conjunction with outcome measures in skill acquisition research. it could be argued that the collection of process-measures to study learning may be as important as the adoption of measures of retention and transfer proved to be in the motor learning literature in the late 1980s (williams & ericsson, 2007). we need to better differentiate between the process and product of learning and to ascertain the efficiency and effectiveness gains (or trade-offs) that may be obtained through different interventions. visual gaze has been used as an index of attention during performance on a specific task (e.g., anticipation), but thus far, only in a few published reports have gaze behaviours been recorded before and after an intervention or practice phase (for exceptions, see alder, ford, causer, & williams, 2016; breslin, hodges, williams, kramer, & curren, 2007; causer, holmes, & williams, 2011). measures of gaze can help establish whether more stable or efficient patterns of visual search behaviour in some learners (i.e., longer duration fixations on information rich areas of a display) inform what information is being acquired, how improvements in performance are attained, and/or explain “learning” in the absence of performance effects. such an approach is especially important when studying expert learning, where subtleties in processes may be more evident following practice than (positive) changes in behavioural outcomes. moreover, there is a paucity of research where verbal reports have been gathered on a pre-test as well as on subsequent retention and transfer-tests (for an exception, see coughlan et al., 2014). whilst kinematic measures are commonly employed on simple tests of motor skill learning, detailed motion analysis is less common for the acquisition of more complex skills, such as those involved in sport. researchers have tended to focus on changes in motor control by measuring variation in single measures (e.g., range of motion, linear velocity) rather than in global motor coordination patterns (e.g., conjugate cross correlations and angle-angle plots as measures of coordination; see carling, reilly, & williams, 2009). the use of neurophysiological measures such as fmri, tms or eeg are reported more frequently as process measures of skill learning in simple, constrained laboratory-based tasks of motor skill (e.g., key press sequencing or single limb adaptation to novel environments; for reviews see hardwick, rottschy, miall, & eickhoff, 2013; lohse, wadden, boyd, & hodges, 2014). however, the recording of neurophysiological measures on preand retention/transfer-tests following the acquisition of more complex skills has been rare (for an exception, see bezzola, merillat, gaser, & jancke, 2011). the difficulty is that while we have a reasonable understanding of the effectiveness of different instructional interventions at the outcome level, only limited knowledge has been generated in regards to how changes in process-measures williams et al | f l r 146 accompany changes in outcome. an interesting question is the extent to which process-tracing measures (such as visual search) may predict changes in learning across and within individuals. such research would help us better understand the processes underpinning expert learning and how these relate to the development of mechanisms that promote expertise. 3.3 tracing the development of expert learners an extensive body of research now exists focusing on deliberate practice and its contribution to expert performance. the findings remain somewhat controversial with wide ranging views regarding the variance in performance across individuals accounted for by hours accumulated in deliberate practice (see ericsson et al., in press; hambrick, altmann, oswald, meinz, & gobet, 2014; mcnamara, hambrick, & oswald, 2014). perhaps the most interesting issue is that that variability in the amount of hours accumulated across individuals has been largely unexplored. numerous researchers have reported extremely large standard deviations in the number of hours athletes accumulate in different types of practices activities during development, implying that practice (at least given issues in measurement) is not the sole factor in developing expertise (e.g., see ford et al., 2012; hopwood, macmahon, farrow & baker, 2016). the difficulty with the majority of the existing research on deliberate practice is that it tends to be correlational in nature, rather than prospective or intervention-based. the data may largely be describing the social and cultural backgrounds surrounding a particular cohort or domain, rather than highlighting causal factors which promote excellence. if the deliberate practice framework is to continue to make a valuable contribution to the learning sciences, better efforts are needed to link engagement in specific types of practice activities with specific improvements in related components of performance. the need to be able to draw firmer inferences about causality may necessitate moving away from retrospective, historical accounts of practice (i.e., the ‘bean counting’ approach) towards more prospective approaches combining some of the tenants of deliberate practice theory with more traditional, quasi-experimental designs. a published report by coughlan, williams, and ford (2014) nicely illustrates how the tenants of deliberate practice may be studied under controlled conditions involving a traditional learning design, including measures of transfer and retention, and process measures of learning. the need remains to specify not just how much deliberate practice occurs amongst learners, but also how deliberate the learners are during practice itself. self-regulated learners take control of their learning environment and engage in specific learning strategies to improve (zimmerman, 2008). the motivation to develop deeper understanding of key concepts and subject matter lead to long-term retention and a greater ability to transfer skills across domains. similarly, mental toughness, grit and resilience may mediate expert learning as performers must develop new and innovative ways to overcome obstacles and continue to challenge themselves (hodges, ford, hendry, & williams, in press). do expert learners simply respond better to challenges in the learning environment or do they create those challenges to push themselves to new heights? although research using the deliberate practice framework and cross-sectional, expert-novice type comparisons may be criticised, shortcomings equally exist with more ‘traditional’ learning designs that rely on an experimental approach. the majority of researchers who study motor-skill learning have employed novice participants who often have no or very little skill (or interest) in the chosen task. novel tasks are often used, in efforts to equate individuals before practice, which have little applicability to real-world scenarios and relatively short acquisition periods, sometimes not even including a retention and/or transfer test. moreover, experience and skill are often confounded in the choice of participants, such that, for example, participants high in skill are typically highly experienced making it impossible to disassociate these two factors. the absence of research with more skilled performers is notable, particularly involving real-world tasks. no published reports exist using more and less expert learners. what is needed is a refinement in the methods used by the different camps of researchers and a greater emphasis on using mixed-methods, combining retrospective, experimental and prospective designs and employing processes-tracing measures. williams et al | f l r 147 the advantages and disadvantages of retrospective practice history profiling, traditional learning studies and prospective designs are highlighted in table 1. table 1 the advantages and disadvantages associated with using retrospective practice history profiles, traditional experimental approaches to learning as well as prospective designs (adapted from williams & ericsson, 2005). retrospective practice history profiling traditional preto post-test designs prospective designs provides a general description of the types of activities performers have engaged in to become expert enables the validity of specific instructional interventions and practice schedules to be examined under controlled conditions leading to inferences regarding causality can answer questions about why differences exist across individuals and enable causality judgements strong emphasis on identifying the macro rather than micro structure of practice strong emphasis on the micro-structure of practice, rather than macro-level allow subtle measurement of changes in performance with practice that might not be seen in typical behavioural measures limited attempts to identify the specific practice activities engaged in and to link these to changes in specific components of performance overreliance on simplistic and novel laboratory tasks leading to concerns regarding generalizability of findings (to the real-world and to experts) allows tracking of specific activities within a targeted group of individuals few attempts to use the approach in conjunction with experimentally-based, pre-post-test designs short-term interventions with limited follow-up regarding long-term retention of skills assessment of short-term interventions is permitted over long periods of time absence of control groups (matched for age and experience) applicability of findings somewhat dependent on strengths of transfer tests employed, particularly if stressors such as fatigue and anxiety missing prospective designs allow for better reliability and validity in measures/conclusions lack of focus on individual differences and large standard deviations in data often an absence of retention tests, or the use of short retention periods, and inappropriate use of, or absent, transfer tests age can be factored into longitudinal designs, such that comparisons can be made across age groups and across individuals over time not possible to imply causality from such data majority of published reports involve novice participants with limited focus on expert performers and on expert learners prospective designs afford a better focus on changes within individuals and causality statements concerns with validity and reliability of retrospective estimates of practice hours lack of research studying how instructional variables, practice variables and learner’s skill interact, leading to a limited, reductionist approach to exploring skill acquisition multiple measures of process and outcome can be integrated to identify efficiency-effectiveness trade-offs and how these change over time an area of study that has attracted a growing body of research in recent years relates to the use of video and other forms of simulation that enable opportunities for anticipation and decision making to be isolated under controlled conditions in the laboratory, greatly increasing the opportunity for repetition and engagement in deliberate practice (for a recent review, see broadbent et al., 2015). the prototypical williams et al | f l r 148 approach involves filming an action sequence (e.g., the serves of opponents in tennis) from the perspective of the player and then replaying the action, which may be occluded at various time intervals before, at or after ball-racket contact. the task for the participant would be to anticipate where and what type of serve the opponents would employ before deciding on the correct return shot. such an approach would be used for the preand post-test, whereas a transfer test would likely involve serves from opponents not previously observed, and potentially some form of on-court data collection. the intervention period would typically require players to view other servers and to receive instruction related to the pick-up of key postural cues to facilitate anticipation. feedback regarding task performance would be provided (for examples, see abernethy, schorer, jackson, & hagemann, 2012; smeeton, williams, hodges, & ward, 2005; williams et al., 2002). while there remain limitations with existing work on perceptual-cognitive training, notably in regards to the measurement of skill transfer, the paradigm has been used to examine the conceptual underpinnings of a range of issues in the learning sciences, including: practice scheduling (broadbent, causer, williams, & ford, 2015b; hodges et al., 2011); focus of attention (abernethy, schorer, jackson, & hagemann, 2012); the effectiveness of imagery (smeeton, hibbert, stevenson, cumming, & williams, 2015); perceptual cueing (ryu, kim, abernethy, & mann, 2012); and training under pressure (alder et al., 2015). although there have been attempts to look at individual differences in perceptual-cognitive skills and determine how well these predict performance (e.g., mangine et al., 2014), such an approach has not been used to examine how elite athletes improve or respond to practice over a relatively short intervention (for an exception, see faubert, 2013). in terms of differentiating across skill groups, the evidence is pretty mixed regarding the ability of general perceptual-cognitive skills, such as attention or spatial-iq, to distinguish better from worse athletes (for a review, see voss et al., 2010). with respect to learning, this general abilities approach may be informative, especially if combined with sport-specific measures of skill. for example, do the athletes who show better recall or recognition for patterns of play specific to their sport, also respond better to practice manipulations designed to improve perceptual cognitive skill in recognizing deceptive plays for example? whilst there is anecdotal evidence for “good” learners in sports, the reasons for this receptivity have not been explored and as such, studying variables such as current skill level, playing experience, attention or cognitive capacities, might provide insight into variables and characteristics which are most related to continued learning and the overcoming of challenges associated with necessary technique changes (such as in the case of injury or changing demands in the sport). in addition to the relative short acquisition periods that are employed in laboratory based research on motor learning (e.g., 45-60 min, see williams, ward, & chapman, 2003), there is an absence of longitudinal work. consequently, we have limited idea of how skill learning progresses and changes over time or the extent to which the typically observed changes are durable and lasting. we know very little about the factors that differentiate someone who learns these skills effectively and efficiently from those who record smaller changes in performance either over time or as a result of the training intervention. the absence of any longitudinal work also makes it difficult to judge whether these measures of specific task components (e.g., perceptual-cognitive skills) have any predictive utility from a talent identification perspective. if someone scores well on these tests at an early age does this suggest that performance will continue to be high relative to others in an older age grouping? moreover, are there key time windows for the acquisition of these skills and how is this linked to general development in young children? do individuals that demonstrate ‘exceptional talent’ learn skills more efficiently following brief exposure or is the development of perceptual-cognitive skills non-linear, varying from one stage of development to the next, making prediction difficult? williams et al | f l r 149 4. conclusions in this article, the expert performance approach, originally introduced by ericsson and smith (1991) was reviewed as a systematic framework for the study of expert learning. over recent decades, this approach has helped to stimulate research in the area of perceptual-cognitive expertise. in contrast to the growing research on expert performance across numerous domains, as well as significant research focusing on the practice history profiles of experts and novices, there remains a need for research on how expert learners continue to learn new skills and refine existing ones. we advocate a more systematic approach to the study of (motor) learning in general, based on process tracing measures and identification of individual differences related to effectiveness and efficiency. someone who acquires, retains and transfers skill better than another individual may perhaps be categorised as a skilled or more expert learner. the absence of research on the above topic largely arises because the prototypical approach has been to use participants classified based on their current level of performance rather than their expertise in learning (which of course requires some formative assessment). this latter approach necessitates the selection of participants (typically retrospectively) based on the level of performance change observed over time on a particular component of performance. a strong focus on identifying individual differences in learning is essential, rather than relying on group means over practice blocks. the use of process-tracing measures during acquisition will improve understanding of how learning takes place rather than the amount of learning that has occurred, as is the case when relying solely on outcome scores. we suggest that the expert performance approach has considerable potential in offering a framework to identify expert learners across domains and that it offers a systematic, guiding framework for better understanding how experts have learned the skills needed to perform at high levels in their respective domains. a stronger focus on the science of learning is key to enhance knowledge of how skills are acquired and how we then can promote more efficient and effective skill acquisition cross many professional domains. keypoints the expert performance approach presents a systematic framework for examining how expert learners acquire and refine skills across domains. more effort is needed to identify individuals who demonstrate exceptional learning on specific skills (as defined by efficiency and effectiveness in attainment) when compared to norms, if we are to develop more refined methods to accelerate skill learning across domains. a stronger focus is needed on exploring individual differences in learning and on using processtracing measures to evaluate how learning progresses, rather than on an existing overreliance on outcome scores, averaged across individuals. measures of learning efficiency (rate of learning, number of practice trials or days of practice) and effectiveness (i.e., accuracy, consistency, skill quality, speed) are needed to gain a more accurate picture of how experts learn across many domains of professional activity. references abernethy, b., schorer, j., jackson, r. c., & hagemann, n. (2012). perceptual training methods compared: the relative efficacy of different approaches to enhancing sport-specific anticipation. journal of experimental psychology: applied, 18, 143. http://psycnet.apa.org/doi/10.1037/a0028452 abernethy, b., & zawi, k. (2007). pickup of essential kinematics underpins expert perception of movement patterns. journal of motor behavior, 39, 353-367. http://dx.doi.org/10.3200/jmbr.39.5.353-368 http://psycnet.apa.org/doi/10.1037/a0028452 http://dx.doi.org/10.3200/jmbr.39.5.353-368 williams et al | f l r 150 ackerman, p. l. (1987). individual differences in skill learning: an integration of psychometric and information processing perspectives. psychological bulletin, 102, 1, 3. http://dx.doi.org.ezproxy.lib.utah.edu/10.1037/0033-2909.102.1.3 alder, d., ford, p. r. causer, j., & williams, a. m., (2016). the effects of highand low-anxiety training on the anticipation judgements of elite performers. journal of sport & exercise psychology, 38, 93-104. http://dx.doi.org/10.1123/jsep.2015-0145 balser, n., lorey, b., pilgramm, s., stark, r., bischoff, m., zentgraf, k., williams, a. m., & munzert, j. (2014) prediction of human actions: expertise and task-related effects on neural activation of the action observation network. human brain mapping, 35, 4016-4034. doi:10.1002/hbm.22455 baker, j. & farrow, d. (2015). routledge handbook of sport expertise. london: routledge bezzola, l., merillat, s., gaser, c., & jancke, l. (2011). training-induced neural plasticity in golf novices. the journal of neuroscience, 31, 35, 12444-12448. https://doi.org/10.1523/jneurosci.1996-11.2011 bjork, e. l., & bjork, r. a. (2011). making things hard on yourself, but in a good way: creating desirable difficulties to enhance learning. in m. a. gernsbacher, r. w. pew, l. m. hough, & j. r. pomerantz (eds.), psychology and the real world: essays illustrating fundamental contributions to society (pp. 5664). new york: worth publishers. breslin, g., hodges, n.j., williams, a.m., kremmer, j. & curren w. (2006). the role of intraand interlimb relative motion information in modeling a novel motor skill. human movement science, 25, 6, 753-752. http://dx.doi.org/10.1016/j.humov.2006.04.002 broadbent, d. p., causer, j., williams, a. m., & ford, p. r. (2015a). perceptual-cognitive skill training and its transfer to expert performance in the field: future research directions. european journal of sport science, 15, 322-331. http://dx.doi.org/10.1080/17461391.2014.957727 broadbent, d. p., causer, j., ford, p. r., & williams, a. m. (2015b). contextual interference effect on perceptual-cognitive skills training. medicine and science in sports and exercise, 47, 6, 1243-1250. http://dx.doi.org/10.1249/mss.0000000000000530 broadbent, d. p., causer, j., williams, a. m., & ford, p. r. (in press). the role of cognitive effort and error processing in the contextual interference effect during perceptual-cognitive skills training. journal of experimental psychology: human perception & performance. http://psycnet.apa.org/doi/10.1037/xhp0000375 bourne, m., bennett, s.j., hayes, s., & williams, a.m. (2011). the dynamical structure of handball penalty shots as a function of target location. human movement science, 30, 1, 40-5. http://dx.doi.org/10.1016/j.humov.2010.11.001 carling, c., reilly, t.p., & williams, a.m. (2009). the handbook of soccer match analysis. taylor & francis: london. causer, j., holmes, p.s., & williams, a.m. (2011). quiet eye training in elite performers. medicine & science in sport & exercise, 43, 1042-1049. doi: 10.1249/mss.0b013e3182035de6 chiviacowsky, s., & wulf, g. (2002). self-controlled feedback: does it enhance learning because performers get feedback when they need it? research quarterly for exercise and sport, 73, 408-415. http://dx.doi.org/10.1080/02701367.2002.10609040 coughlan, e., williams, a.m., mcrobert, & ford., p. (2014) a novel test of deliberate practice theory: how experts learn. journal of experimental psychology: learning, memory & cognition, 40, 449-458. http://psycnet.apa.org/doi/10.1037/a0034302 davids, k., bennett, s.j., & newell, k.m. (2006). movement system variability. human kinetics: champaign, illinois. dennis, d., rowe, r., williams, a.m., & milne, e. (2017). the role of cortical sensorimotor oscillation in action anticipation. neuroimage, 146, 1102-1114. http://dx.doi.org/10.1016/j.neuroimage.2016.10.022 ericsson, k. a., & kintsch, w. (1995). long-term working memory. psychological review, 102, 211-245. http://psycnet.apa.org/doi/10.1037/0033-295x.102.2.211 ericsson, k. a., & smith, j. (1991). prospects and limits of the empirical study of expertise: an introduction. in k. a. ericsson & j. smith (eds.), towards a general theory of expertise: prospects and limits (pp. http://dx.doi.org.ezproxy.lib.utah.edu/10.1037/0033-2909.102.1.3 http://dx.doi.org/10.1123/jsep.2015-0145 https://doi.org/10.1523/jneurosci.1996-11.2011 http://dx.doi.org/10.1016/j.humov.2006.04.002 http://dx.doi.org/10.1080/17461391.2014.957727 http://dx.doi.org/10.1249/mss.0000000000000530 http://psycnet.apa.org/doi/10.1037/xhp0000375 http://dx.doi.org/10.1016/j.humov.2010.11.001 http://dx.doi.org/10.1080/02701367.2002.10609040 http://psycnet.apa.org/doi/10.1037/a0034302 http://dx.doi.org/10.1016/j.neuroimage.2016.10.022 http://psycnet.apa.org/doi/10.1037/0033-295x.102.2.211 williams et al | f l r 151 1-38). new york: cambridge university press. ericsson, k. a., & williams, a. m. (2007). capturing naturally occurring superior performance in the laboratory: translational research on expert performance. journal of experimental psychology: applied, 13(3), 115-123. http://psycnet.apa.org/doi/10.1037/1076-898x.13.3.115 ericsson, k. a., krampe, r. t., & tesch-römer, c. (1993). the role of deliberate practice in the acquisition of expert performance. psychological review, 100, 363-406. http://psycnet.apa.org/doi/10.1037/0033-295x.100.3.363 ericsson, k. a., hoffman, r., aaron, k., & williams, a. m. (in press) (ed.) the cambridge handbook of expertise (second edition). cambridge university press. eysenck, m. w., & calvo, m. g. (1992). anxiety and performance. the processing efficiency theory. cohgnition and emotion, 6, 409-434. http://dx.doi.org/10.1080/02699939208409696 farrow, d., baker, j., & macmahon, c. (2013). developing sport expertise: researchers and coaches put theory into practice (2nd edition). routledge/taylor and francis. faubert, j. (2013). professional athletes have extraordinary skills for rapidly learning complex and neutral dynamic visual scenes. scientific reports, 3, 1154. doi:10.1038/srep01154 ford, p. r., & williams, a. m. (2012). the developmental activities engaged in by elite youth soccer players who progressed to professional status compared to those who did not. psychology of sport & exercise, 13, 349–352. http://dx.doi.org/10.1016/j.psychsport.2011.09.004 ford, p. r., hodges, n. j., &williams, a. m. (2013). expert sports performance and its development. in beyond “talent or practice?”: the multiple determinants of greatness (edited by b. kaufman), 391414. oxford: oxford university press. ford, p. r., carling, c., garces, m., marques, m., miguel, c., farrant, a., & williams, a.m. (2012). the developmental activities of elite soccer players aged under-16 years from brazil, england, france, ghana, mexico, portugal and sweden. journal of sports sciences, 30, 1653-1663. http://dx.doi.org/10.1080/02640414.2012.701762 hambrick, d. z., altmann, e. m., oswald, f. l., meinz, e. j., & gobet, f. (2014). facing facts about deliberate practice. frontiers in psychology, 5, 751. doi:10.3389/fpsyg.2014.00751 hardwick, r. m., rottschy, c., miall, r. c., & eickhoff, s. b. (2013). a quantitative meta-analysis and review of motor learning in the human brain. neuroimage, 67, 283-297. http://dx.doi.org/10.1016/j.neuroimage.2012.11.020 hodges, n. j., edwards, c., luttin, s., & bowcock, a. (2011). learning from the experts: gaining insights into best practice during the acquisition of three novel motor skills. research quarterly for exercise & sport, 82, 178-187. http://dx.doi.org/10.1080/02701367.2011.10599745 hodges, n. j., lohse, k. r., wilson, a., lim, s. b., & mulligan, d. (2014). exploring the dynamic nature of contextual interference: previous experience affects current practice but not learning. journal of motor behavior, 46(6), 455-467. http://dx.doi.org/10.1080/00222895.2014.947911 hodges, n. j., ford, p. r., hendry, d. j., & williams, a. m. (in press). getting gritty about practice and success: motivational characteristics of great performers. progress in brain research. http://dx.doi.org/10.1016/bs.pbr.2017.02.003 hodges, n. j., & williams, a. m. (2012). skill acquisition in sport: research, theory and practice (second edition). routledge: london. hopwood, m., macmahon, c., farrow, d., & baker, j. (2016). is practice the only determinant of sporting expertise? revisiting starkes (2000). international journal of sport psychology, 47(1), 631-651. keetch, k. m., & lee, t. d. (2007). the effect of self-regulated and experimenter-imposed practice schedules on motor learning for tasks of varying difficulty. research quarterly for exercise and sport, 78(5), 476-486. http://dx.doi.org/10.1080/02701367.2007.10599447 kennedy, q., taylor, j. l., reade, g., & yesavage, m. d. (2010). age and expertise effects in aviation decision making and flight control in a flight simulator. aviation, space and environmental medicine, 81, 489-497. https://doi.org/10.3357/asem.2684.2010 lee, t. d., & wishart, l. r. (2005). motor learning conundrums (and possible solutions). quest, 57(1), 6778. http://dx.doi.org/10.1080/00336297.2005.10491843 http://psycnet.apa.org/doi/10.1037/1076-898x.13.3.115 http://psycnet.apa.org/doi/10.1037/0033-295x.100.3.363 http://dx.doi.org/10.1080/02699939208409696 http://dx.doi.org/10.1016/j.psychsport.2011.09.004 http://dx.doi.org/10.1080/02640414.2012.701762 http://dx.doi.org/10.1016/j.neuroimage.2012.11.020 http://dx.doi.org/10.1080/02701367.2011.10599745 http://dx.doi.org/10.1080/00222895.2014.947911 http://dx.doi.org/10.1016/bs.pbr.2017.02.003 http://dx.doi.org/10.1080/02701367.2007.10599447 https://doi.org/10.3357/asem.2684.2010 http://dx.doi.org/10.1080/00336297.2005.10491843 williams et al | f l r 152 lohse, k. r., wadden, k., boyd, l. a., & hodges, n. j. (2014). motor skill acquisition across short and long time scales: a meta-analysis of neuroimaging data. neuropsychologia, 59, 130-141. http://dx.doi.org/10.1016/j.neuropsychologia.2014.05.001 mangine, g.t., hoffman, j.r., wells, a.j., gonzalez, a.m., townsend, j.r., jajtner, a.r., beyer, k.s., bohner, j.d., pruna, g.j., fragala, m.s., stout, j.r. (2014). visual tracking speed is related to basketball-specific measures of performance in nba players. journal of strength conditioning research, 28 (9), 2406-2414. doi: 10.1519/jsc.0000000000000550 mcnamara., b. n., hambrick, d. z., & oswald, f. c. (2014). deliberate practice and performance in music, games, sports, education, and professions: a meta-analysis. psychological science, 25, 1608-1618. doi:10.1177/0956797614535810 mcrobert, a., causer, j., vasiliadus, j., watterson, l., & williams, a. m. (2013). contextual information influences diagnosis accuracy and decision-making in simulated emergency medicine emergencies. british medical journal: quality & safety, 22, 478-484. http://dx.doi.org/10.1136/bmjqs-2012000972 müller, s., brenton, j., dempsey, a. r., harbaugh, a. g., & reid, c. (2015). individual differences in highly skilled visual perceptual-motor striking skill. attention, perception, & psychophysics, 77(5), 1726-1736. doi:10.3758/s13414-015-0876-7 mulligan, d., lohse, k. r., & hodges, n. j. (2016a). an action-incongruent secondary task modulates prediction accuracy in experienced performers: evidence for motor simulation. psychological research, 80, 496-509. doi:10.1007/s00426-015-0672-y mulligan d., lohse, k.r., & hodges, n.j. (2016b). evidence for dual mechanisms of action prediction dependent on acquired visual-motor experiences. journal of experimental psychology: human perception and performance, 42, 1615-1626. doi: 10.1037/xhp0000241 north, j. s., ward, p., ericsson, a., & williams, a. m. (2011). mechanisms underlying skilled anticipation and recognition in a dynamic and temporally constrained domain. memory, 19(2), 155-168. http://dx.doi.org/10.1080/09658211.2010.541466 roca, a., ford, p. r., mcrobert, a. p., & williams, a. m. (2013). perceptual-cognitive skills and their interaction as a function of task constraints in soccer. journal of sport & exercise psychology, 35, 144155. http://dx.doi.org/10.1123/jsep.35.2.144 ryu, d., kim., s., abernethy, b., & mann, d. l. (2012). guiding attention aids the acquisition of anticipatory skill in novice soccer goalkeepers. research quarterly for exercise & sport, 84, 252-262. http://dx.doi.org/10.1080/02701367.2013.784843 savelsbergh, g.j.p., van der kamp, j., williams, a.m., & ward, p. (2005). anticipation and visual search behavior in expert soccer goalkeepers. a within-group comparison. ergonomics, 48(11-14), 1686-1697. http://dx.doi.org/10.1080/00140130500101346 schmidt, r., & lee, t. (2013). motor learning and performance, 5th edition: from principles to application. champaign, il: human kinetics. smeeton, n. j., & williams, a. m. (2012). the role of movement exaggeration in the anticipation of deceptive soccer penalty kicks. british journal of psychology, 103, 539-555. doi: 10.1111/j.20448295.2011.02092.x smeeton, n. j., williams, a. m., hodges, n. j., & ward, p. (2005). the relative effectiveness of various instructional approaches in developing anticipation skill. journal of experimental psychology: applied, 11, 98–110. http://psycnet.apa.org/doi/10.1037/1076-898x.11.2.98 smeeton, n. j., hibbert, j. r., stevenson, k., cumming, j., & williams, a. m. (2013). can imagery facilitate improvements in anticipation behavior? psychology of sport & exercise, 14, 200-210. http://dx.doi.org/10.1016/j.psychsport.2012.10.008 smith, m. a., ghazizadeh, a., & shadmehr, r. (2006). interacting adaptive processes with different timescales underlie short-term motor learning. plos biol, 4(6), e179. http://dx.doi.org/10.1371/journal.pbio.0040179 stal, p., donmez., b., & jamieson, g. a. (2016). supporting anticipation in driving through attentional and interpretational in-vehicle displays. accident, analysis and prevention, 91, 103-113. http://dx.doi.org/10.1016/j.neuropsychologia.2014.05.001 http://dx.doi.org/10.1136/bmjqs-2012-000972 http://dx.doi.org/10.1136/bmjqs-2012-000972 http://dx.doi.org/10.1080/09658211.2010.541466 http://dx.doi.org/10.1123/jsep.35.2.144 http://dx.doi.org/10.1080/02701367.2013.784843 http://dx.doi.org/10.1080/00140130500101346 http://psycnet.apa.org/doi/10.1037/1076-898x.11.2.98 http://dx.doi.org/10.1016/j.psychsport.2012.10.008 http://dx.doi.org/10.1371/journal.pbio.0040179 williams et al | f l r 153 http://dx.doi.org/10.1016/j.aap.2016.02.030 vater, c., roca, a., & williams, a. m. (2015). effects of anxiety on anticipation and visual search in dynamic, time-constrained situations. sport, exercise, & performance psychology. advance online publication. http://dx.doi.org/10.1037/spy0000056 voss, m. w., kramer, a. f., basak, c., prakash, r. s., & roberts, b. (2010). are expert athletes ‘expert’ in the cognitive laboratory? a meta‐analytic review of cognition and sport expertise. applied cognitive psychology, 24(6), 812-826. doi:10.1002/acp.1588 wadden, k.p., hodges, n.j., de asis, k.l., neva, j.l., & boyd, l.a. (2017). individualized challenge point practice informs motor sequence learning: more time in an early phase of practice benefits later retention. plosone. walker, m. p., brakefield, t., morgan, a., hobson, j. a., & stickgold, r. (2002). practice with sleep makes perfect: sleep-dependent motor skill learning. neuron, 35(1), 205-211. http://dx.doi.org/10.1016/s0896-6273(02)00746-8 ward, p., & williams, a.m. (2003). perceptual and cognitive skill development in soccer: the multidimensional nature of expert performance. journal of sport & exercise psychology, 2(1), 93-111. http://dx.doi.org/10.1123/jsep.25.1.93 ward, p., ericsson, k. a., & williams, a. m. (2013). complex perceptual-cognitive expertise in a simulated task environment. journal of cognitive engineering & decision making, 7(3), 231-254. doi:10.1177/1555343412461254 ward, p., hodges, n.j., williams, a.m., & starkes, j. (2007). the role of deliberate practice in the development of expert performers. high ability studies, 18, 119-153. http://dx.doi.org/10.1080/13598130701709715 wright, m. j., bishop, d., jackson, r. c., & abernethy, b. (2010). functional mri reveals expert-novice differences during sport-related anticipation. neuroreport, 21, 94-98. doi:10.1097/wnr.0b013e328333dff2 wright, m. j., bishop, d., jackson, r. c., & abernethy, b. (2011) cortical fmri activation to opponents’ body kinematics in sport-related anticipation: expert-novice differences with normal and point-light video. neuroscience letters, 500(3), 216-221. http://dx.doi.org/10.1016/j.neulet.2011.06.045 williams, a. m., & abernethy, b. (2012). anticipation and decision-making: skills, methods and measures. in g. tenenbaum, r. c. eklund & a. kamata (eds.), measurement in sport & exercise psychology. (pp. 191-202), champaign, il: human kinetics. williams, a. m., & ericsson, k. a. (2007). perception, cognition, action and skilled performance. journal of motor behavior, 39(5), 338-340. doi:10.3200/jmbr.39.5.338-340 williams, a. m., & ericsson, k. a. (2005). perceptual-cognitive expertise in sport: some considerations when applying the expert performance approach. human movement science, 24(3), 283-307. http://dx.doi.org/10.1016/j.humov.2005.06.002 williams, a.m., ward, p., & chapman, c. (2003). training perceptual skill in field hockey: is there transfer from the laboratory to the field? research quarterly for exercise & sport, 74, 98-104. http://dx.doi.org/10.1080/02701367.2003.10609068 williams, a. m., ford, p. r., eccles, d. w., & ward, p. (2011). perceptual-cognitive expertise in sport and its acquisition: implications for applied cognitive psychology. applied cognitive psychology, 25, 432442. doi:10.1002/acp.1710 williams, a. m., north, j. s., & hope, e. r. (2012). identifying the mechanisms underpinning recognition of structured sequences of action. the quarterly journal of experimental psychology, 65, 1975-1992. http://dx.doi.org/10.1080/17470218.2012.678870 williams, a.m., ward, p., bell-walker, j., & ford, p. (2012). discovering the antecedents of anticipation and decision making skill. british journal of psychology, 103, 393-411. doi: 10.1111/j.2044-8295.2011.02081.x. williams, a. m., ward, p., knowles, j. m., & smeeton, n. j. (2002). anticipation skill in a real-world task: measurement, training, and transfer in tennis. journal of experimental psychology: applied, 8, 259270. http://psycnet.apa.org/doi/10.1037/1076-898x.8.4.259 http://dx.doi.org/10.1016/j.aap.2016.02.030 http://dx.doi.org/10.1037/spy0000056 http://dx.doi.org/10.1016/s0896-6273(02)00746-8 http://dx.doi.org/10.1123/jsep.25.1.93 http://dx.doi.org/10.1080/13598130701709715 http://dx.doi.org/10.1016/j.neulet.2011.06.045 http://dx.doi.org/10.1016/j.humov.2005.06.002 http://dx.doi.org/10.1080/02701367.2003.10609068 http://dx.doi.org/10.1080/17470218.2012.678870 http://psycnet.apa.org/doi/10.1037/1076-898x.8.4.259 williams et al | f l r 154 wulf, g., shea, c., & lewthwaite, r. (2010). motor skill learning and performance: a review of influential factors. medical education, 44, 75-84. doi:10.1111/j.1365-2923.2009.03421.x zimmerman, b.j. (2008). investigating self –regulation and motivation: historical background, methodological developments, and future prospects. american educational research journal, 45, 166-183. doi:10.3102/0002831207312909 frontline learning research 5 (2014) 115-139 issn 2295-3159 corresponding author: kerstin helker, institute for education, rwth aachen university, eilfschornsteinstraße 7, 52056 aachen, germany, email: kerstin.helker@rwth-aachen.de doi: http://dx.doi.org/10.14786/flr.v2i3.99 115 | f l r responsibility in the school context – development and validation of a heuristic framework kerstin helker a , marold wosnitza a,b a rwth aachen university, germany b murdoch university, perth, australia article received 27 march 2014 / revised 25 may 2014 / accepted 23 june 2014 / available online 1 july 2014 abstract existing research has identified feelings of responsibility as having major motivational implications for a person’s actions. a person identifying as being responsible for a certain task will perceive themselves as self-determined and thus invest considerable effort in the task. despite being conceptualised as an individual’s sense of internal obligation, responsibility in everyday contexts is often attributed by and to other people. different perspectives on responsibility may, however, not always overlap, especially in the school context where tasks and liabilities often remain ill-defined. this paper thus presents a framework of responsibility in the school context which assumes teachers, students and parents to share a certain number of microsystems which may (indirectly) influence one another. in order to test the usefulness of the proposed framework, a series of studies were conducted collecting data on teachers’, students’ and parents’ views of their own and one another’s responsibility in the school context. 4339 statements were assigned to categories representing different parts of the framework and reveal its usefulness for describing the complexity of responsibility attributions and its influences in the school context. findings show the framework will be helpful to embrace existing research and develop questions for further research that address central educational issues such as student and teacher motivation, teacher burnout as well as prerequisites for students’ high or low achievement. keywords: teacher responsibility; student responsibility; parent responsibility; school context k. helker & m. wosnitza 116 | f l r in the last decade the extension of demands on schools led to an extensive discussion on the particular competencies and responsibilities of stakeholders in the school context. the challenge is that for most specific responsibilities, due to the complexity and the fact that tasks and liabilities in schools are often ill-defined (fischman, dibara, & gardner, 2006), can often not clearly be attributed to one specific person or group of people. to further complicate this, there is an absence of an agreed-upon definition of the term responsibility that can lead to conflicts between stakeholders perceiving their own and others‟ responsibility differently (lauermann & karabenick, 2011). conflicts are especially likely to occur between teachers, students and parents, all emphasizing different goals and judging their own and others‟ responsibility against the background of their own sphere of experience. a review of the empirical work on responsibility in the school context underlines the complexity of the concept responsibility (author/s). it shows that when teachers, parents and students talk about students‟ learning and achievement they assign the same responsibility differently to each other. it furthermore shows that the context or the cultural setting in which this responsibility attribution takes place plays a significant role. the interplay between the stakeholders‟ attributions of responsibility is still underresearched. thus, this paper aims to examine how teachers, students and parents attribute responsibility to themselves and one another, to disentangle these often implicit and confused responsibility attributions, and represent them in a heuristic framework. 1. responsibility despite being used in a multitude of contexts and sometimes being considered a “core concept of social life” (hamilton, 1978 p. 326), the term responsibility remains unclear (del schalock, 1998; fischman et al., 2006; maulbetsch, 2010). the multitude of perspectives from which responsibility has generally been studied indicates the fluid nature of the concept (lauermann & karabenick, 2011), with perspectives ranging from conceptualising responsibility as a relatively stable disposition of a person (bierhoff, 2000) to the interrelation between personal sense of responsibility and locus of control (guskey, 1981, 1982; rose & medway, 1981a, 1981b). due to this diversity of theoretical perspectives, responsibility in the literature is often conceptualised as a multirelational construct of at least three components, which in each context are engaged differently: somebody is responsible for something under supervision or judgment of some kind of sanctioning instance (auhagen, 1999; auhagen & bierhoff, 2001; bayertz, 1995; grotlüschen, 2008; höffe, 2008; schleißheimer, 1984). this judging instance can take many forms ranging between a court and the internal conscience. one of the most elaborated constructs of responsibility is lenk‟s (lenk, 1992; lenk & maring, 1993) six-component model asking: who (subject of responsibility) is attributed responsibility for what (object of responsibility), in view of whom (addressee) by whom (judging instance) in relation to what (normative) criteria and in what realm (of responsibility or action)? this construct of responsibility was taken up by lauermann and karabenick (2011) who studied the components and theoretical status of teacher responsibility in order to tease out the complexity of its different meanings. one basic aspect of their work was the distinction of responsibility from accountability, with the latter being an explicit, formal attribution of tasks. responsibility was defined as “a sense of internal obligation and commitment to produce or prevent designated outcomes or that these outcomes should have been produced or prevented” (lauermann & karabenick, 2011, p. 135). this definition accommodates the two perspectives implied in responsibility attributions: the retrospective (that something should have happened), which is often linked to questions of fault or guilt (weiner, 1995), and the prospective view (that something should happen), denoting a subject‟s obligations for certain people, things or states (werner, 2006). these prospective responsibilities can furthermore emerge in two different ways. despite some researchers assuming feelings of responsibility to only be the result of a personal disposition (bierhoff, 1995, 2000) or of social attributions (bayertz, 1995) it is generally assumed that responsibility can either result from attributions by other people or a person‟s own sense of obligation (auhagen, 1999; bacon, 1991; kammerl, 2008; kaufmann, 1995). these two perspectives very often overlap, especially when it comes to k. helker & m. wosnitza 117 | f l r rather ill-defined tasks like the teaching profession (fischman et al., 2006). the extensive discussion of teachers‟ professional behaviour and ethics shows this lack of conventional means for defining teacher responsibility. often, teachers only face a broad description of the field of activity and are (sometimes even contractually) attributed the paramount responsibility to define their specific tasks and what they feel responsible for (werner, 2006). based on self-determination theory, this could be considered positive, as it can be assumed that an internal sense of responsibility evokes more positive motivational responses as this person perceives themselves as self-determined which enhances engagement (berkowitz & daniels, 1963; ryan & deci, 2000) and work satisfaction (müller, 2009), whereas people only being attributed responsibility from external instances would have to be controlled for compliance (lauermann & karabenick, 2011). assuming much responsibility in response to broad or non-existing guidelines for action has, however, been hypothesised to cause burnout (fischman et al., 2006). 2. responsibility in school the above indicates the relevance of discussing responsibility in relation to teaching and learning in schools today. by acting as a teacher, a person is, as in any other job, attributed a specific task responsibility whose nature is determined by the specific role this person incorporates (leithwood, edge, & jantzi, 1999). due to teachers‟ tasks and liabilities being rather ill-defined (feiks, 1992; fischman et al., 2006; pätzold, 2008; tenorth, 2004), teachers are left to define what they assume themselves, and others respectively, to be responsible for. students and their parents, in return, can be assumed to also go through the process of defining their own and others‟ responsibilities which again influences teachers (fischman et al., 2006), who according to feiks (1992) are expected to do more than just fulfilling their explicitly set duties. empirical research up to this point, however, seems to have been guided by the role of teachers as the only bearer of responsibility in the classroom (bastian, 1995; del schalock, 1998; eikenbusch, 2009) and has strongly focused its attention on teacher responsibility, linking this research field to aspects like sources of teacher responsibility, contextual influences on perceptions of responsibility (responsibility as a social, situational phenomenon) and limitations of responsible actions. students‟ and parents‟ responsibility mostly served as confinements of teacher responsibility rather than being studied for their own sake. regarding teacher responsibility some studies focused on general objects of teacher responsibility (bourke, 1990), which they found to be centred around preparation and structuring learning materials, while others specifically studied teachers‟ sense of personal responsibility for their students‟ educational outcomes (bracci, 2009; halvorsen, lee, & andrade, 2009; matteucci & gosling, 2004; potvin & papillon, 1992). results showed that teachers were more ready to assume responsibility for their students‟ success than failure with responsible teachers being better prepared and attending more advanced training units while also experiencing more support and encouragement by school administrators. the direct school setting, its socioeconomic background (diamond, randolph, & spillane, 2004), size (lee & loeb, 2000) and perceived family influences (thrupp, mansell, hawksworth, & harold, 2003) but also more remote factors like the cultural influences on teachers‟ perception of their professional identity (barrett, 2005; karakaya, 2004) were found to affect teachers‟ perceptions and perceived limitations of their responsibility. fischman, dibara and gardner (2006), however, found that teachers, compared to other professions, were more ready to take responsibility. they generally counteract the missing of clear instructions and norms of behaviour by steadily focusing their sense of responsibility on their students whom they perceive to be primary addressees of their responsibility. teachers state to meet higher academic, social, emotional and developmental demands of students resulting from problematic environments by expanding their sphere of action and sense of responsibility. this behaviour could be assumed to deprive students and parents of their responsibilities, which to some might be considered a positive development. research has shown that student responsibility is widely understood to be limited to cooperative and social behaviour in the classroom (lewis, 2001), meeting expectations and learning goals (bryan & mclaughlin, 2005) and basically “doing the work” and “obeying the rules” (bacon, 1993). the students in bacon‟s study, being asked about what they thought they were k. helker & m. wosnitza 118 | f l r responsible for, indicated to mostly feel to be held responsible rather than have feelings of personal responsibility which based on ryan and deci (2000) can be assumed to deter students from developing feelings of self-determination. in contrast to these findings, zimmerman and kitsantas (2005) found students in self-regulated learning settings to generally rate their abilities higher and attribute more responsibility to learners than teachers. these students did not limit their responsibility to classroom learning, as was done in other studies, but also felt responsible for contextual factors outside the classroom that might indirectly influence their learning. one major factor in this respect, acknowledged by all three, teachers, students and parents, is how central a student‟s parents are with supporting their child‟s school work (ballard & bates, 2008). despite emphasising the importance of parent involvement, only few teachers state to feel responsible for establishing connections with parents but rather hold them responsible for getting engaged in school matters (ramirez, 1999). such views have, however, been found to also be context-specific as in china (katyal & evers, 2007) as well as turkey (korkmaz, 2007) parents are expected to provide a loving and supportive home in which the child is well cared for and thereby equipped for school, but leave educational matters to professional educators such as teachers. thus, parent involvement in school is neither supported nor expected. up to today, to our knowledge, only one new zealand study has studied all three, teachers‟, students‟ and parents‟ perception of their own and others‟ (retrospective) responsibility. peterson, rubiedavies, elley-brown, widdowson, dixon and irving (2011) found students to view themselves as most responsible for their learning outcomes. students, however, indicate influences by (for example more or less sympathetic) teachers and their parents‟ responsibility for supporting them and providing a stimulating learning environment. while interviewed parents shared this view, teachers emphasised students‟ responsibility for their motivation and success as well as contextual matters of school facilities and resources that might influence learning – and enable teachers to deny their responsibilities and attribute them to others. in sum, a review of the empirical literature regarding teacher, student and parent responsibility showed that these three central agents in schools assume or are attributed specific prospective responsibilities, some of which only become apparent in retrospect (helker & wosnitza, 2014). this entails potential for conflict as it is the nature of things that people can only be made accountable for issues they knew about being responsible for beforehand – what you do not know, you cannot take or be attributed responsibility for (lenk, 1992, p. 10). prior research revealed that most of these three major agents in schools direct their behaviour to those things they personally feel responsible for. own responsibilities are outlined by attributing all remaining tasks to other agents. the perception of one‟s own and others‟ prospective and retrospective responsibility is a highly individual matter (gärtner, 2010) which is influenced by the subjective perception of the importance of specific tasks and of situational factors. as these views seldom are openly addressed and negotiated, different perspectives are likely to not always overlap, which might generate conflicts when conflicting goals are emphasized (lauermann & karabenick, 2011). furthermore, empirical research suggests the importance of the role of context and interactions between agents, as findings show perceptions of own and others‟ responsibility to strongly vary with the (national, economic, social etc.) setting (e.g., katyal & evers, 2007; korkmaz, 2007). the following section will present existing models of context which have in the past been applied to (the analysis of) schools‟ working and learning processes and appear relevant for describing responsibility in school. 3. relevance of context for describing responsibility in school a number of models of context have been put forward in existing literature focusing on the multiple aspects of context (see wosnitza & beltman, 2012 for an overview). wosnitza and beltman (2012), who developed a model for analysing context of specific situations with regard to the level of interaction, perspective (subjective/objective) and content (social/physical/formal). the aspect of level of interaction in this model is, as in most other work relating to context, conceptualised closely along the lines of bronfenbrenner‟s model of the ecological environment. despite covering the aspect of differing perspectives k. helker & m. wosnitza 119 | f l r people may hold on various levels of context, the model focuses on explaining the context of a specific situation rather than describing interrelations between different agents. gurtner, monnard and genoud (2001) applied a model of the school context to explore its impact on students‟ motivation. drawing on the model‟s notion of indirect as well as bidirectional influence between the person and their environment, these authors highlighted that “two students placed in an apparently identical situation may react to it differently since the context in which each one will embed that situation might be quite different.” (p. 191). representing the nature of partnerships and relationships between schools, families and communities, joyce l. epstein‟s (2011) framework of overlapping spheres of influence has become widely acknowledged and applied especially to discussions of questions regarding parental involvement in schools (e.g., galindo & sheldon, 2012; katyal & evers, 2007; lawson, 2003). based on prior research that parents can influence student educational outcomes and achievement (e.g., leichter, 1974; lightfoot, 1978; marjoribanks, 1979), epstein (2011) developed a model of school, family and community partnerships which she applied to the development of research questions as well as strategies for action in improving those partnerships. in this model, school, family and community are represented as three spheres of context which overlap to a certain degree which is determined by external forces and internal actions (e.g., time, backgrounds, actions taken in families and schools). these spheres can have unique and also combined influences on children through the interactions of parents, teachers, students and community partners. taking action for bringing together the different partners and thus enlarging the overlap between these spheres is considered to help identify shared responsibilities of home, school, and community and to increase positive influences on children. also, the degree of overlap obscures boundaries between school and family, so that the influence of one sphere can still be at work while the student is involved in the other (epstein, 2011). in proposing this framework, epstein called for the recognition of shared goals and responsibilities for the socialisation and education of the child (epstein, 2011, p. 26) and for researchers to recognise schools‟, homes‟ and communities‟ simultaneous and cumulative effects on student development and learning (epstein & sheldon, 2006). while epstein strongly focused on what strategies could be applied by schools and educators to establish functioning and reliable partnerships with their students‟ families and communities, other researchers have emphasised the influences between these spheres. christenson (2004) pointed out that different antecedents may result in the same outcome (equifinality) while similar initial conditions may still lead to dissimilar results (multifinality). applying bronfenbrenner‟s model of the ecological environment (1979), she denied the possibility of developing “uniform prescriptions” for involving parents to improve students‟ school performance, as interfaces of home and school may be variably overlapping (p. 87). to sum up, existing models can partially account for attributions of responsibility and how responsible or irresponsible actions of teachers, students and parents influence what happens in the specific or related contexts. nevertheless, to the best of our knowledge, no work has been published which developed a theoretical background for research presenting the different agents in the school context and as context for schools and one another. up to now, the different research perspectives appear isolated, incommensurate and thereby impeding a broad understanding of the phenomenon of responsibility in the school context. thus, in the following, a heuristic framework shall be presented which draws on the models of context already presented in order to comprehensively describing and structure responsibility in the school context. in order to validate this framework, empirical data will be applied to support its different components. 4. towards a framework of responsibility in the school context based on the above, we propose a heuristic framework for representing the origins and impacts of responsibility attributions in school context. teachers, students and parents attribute responsibility to themselves and the other agents in the school context to prevent or produce certain outcomes (prospective) k. helker & m. wosnitza 120 | f l r based on their perception of the respective (professional) roles, context and individual spheres of action. responsibility attributed from external sources does not automatically imply an internal sense of responsibility, because often no comparisons are made between different perspectives which can evoke differences in (the possibility of) retrospective attributions of responsibility. furthermore, attributed responsibilities may not overlap due to different perspectives on the context and the settings in which a specific person is involved. regarding the individual spheres of actions, applying bronfenbrenner‟s (1979) model of the ecological environment to responsibility attributions in the school context, we propose that teachers, students and parents engage in several subcontexts, i.e. microsystems, which sum up to constitute this person‟s mesosystem. when it comes to their school-related activities, teacher, students and parents share a specific number of the microsystems in which they are involved, i.e. interfaces of their mesosystems (general sphere of action). a mathematics teacher, for example shares the microsystem „math lesson‟ with their students while he or she might not be involved in the microsystem „english lesson‟, the students share with somebody else. while teachers‟ and parents‟ mesosystems overlap on parents‟ days, parents, just like any other of the named agents, are involved in many other, not school-related microsystems, none of the other agents is part of as for example home, work or free-time activities. actions that are not located in one of the shared microsystems (exosystems) these interfaces comprise (e.g., events outside school), might, however, have an indirect, yet considerable effect on what happens there. a parent-teacher talk might be a microsystem, the teacher and a student‟s parents share, in which the student is not involved but will certainly be affected by. it thus represents an exosystem for the student. while this example is quite obvious, it could be assumed that many other exosystems influence school microsystems which the agents sharing this microsystem are not aware of. all of the above described levels of context are embedded in the macrocontext that could be the cultural setting or the socioeconomic background of the school or family. in sum, the above assumptions allow for the following conclusions. the school context is understood as consisting of a multitude of microsystems which are determined by the agents involved. due to their bidirectional influence with the environment, in the school context, multiple actors function as the context for one another. thus, certain educational outcomes like students‟ success and failure in school cannot be traced back to one specific incident but result from the interplay of the many microsystems a student is involved in as well as the indirect influences of different exosystems. in conclusion, we assume that when it comes to their responsibility, teachers, students and parents can be and are often attributed not only specific responsibilities but also a general responsibility for certain outcomes which is not directly related to what happens in one specific microsystem but rather all of the microsystems, i.e. the mesosystem, this person, representing this specific role, is involved in (e.g., for a student home, school, meetings with friends etc.). the macrosystem in which all the other systems are embedded may influence the nature of the mesosystem as well as the attributed responsibilities (e.g., karakaya, 2004). parents from a different cultural background may thus assume teachers to be involved in microsystems or responsible for objects that they are not in this culture and would thus deny. the nature of the individual‟s mesosystem and thereby the microsystems and interfaces a person is involved in (i.e. who else is involved in this setting, physical and material nature of the setting), determines the objects this person feels or is held responsible for. drawing on lenk‟s (1992) argument that a person can only be held responsible for such things they are aware of, following duff (1998) we furthermore suggest the view that a person can only be attributed responsibilities which could be fulfilled in microsystems they are part of (e.g., teachers cannot be held responsible for what happens in the student‟s home as they are not involved in this microsystem and to not have control over events). although teachers may indirectly influence these events by their actions at school, they do not have a direct control over the events in a student‟s home. thus, the attribution of prospective responsibility requires a profound understanding of what microsystems an individual‟s sphere of action (mesosystem) comprises and which issues they have the capacity to act on. in addition, retrospective judgments of whether and how attributed responsibilities have been fulfilled are only valid if the respective person knew about and at that time was able to act on them. k. helker & m. wosnitza 121 | f l r 4.1 the heuristic framework and its structural elements the proposed heuristic framework of responsibility in the school context brings together the above considerations. as illustrated in figure 1, the framework most importantly comprises teachers, students and parents as the three central subjects, i.e. bearers, of responsibility in the school context who share specific microsystems which determine to what degree their mesosystems overlap. figure 1. heuristic framework for structuring responsibility in the school context. as illustrated in figure 1, on a conceptual level, the following subjects, i.e. bearers, of responsibility, can be identified: the teacher (t), the particular student (s), his/her parent(s) (p) and his/her classmates (c). the mesosystem of a teacher, illustrated by the circle at the top, comprises a multitude of microsystems in which they take responsibility, some of them shared with other people, others not shared with anyone. furthermore, this mesosystem is partly embedded in the macrocontext (i.e. all microsystems in which the person is involved). besides this, the teacher also shares a number of microsystems with people outside school (indicated by the dotted line) which are part of the wider community in which he or she lives, and thus takes responsibility these contexts. these microsystems can also serve as exosystems to what happens in school microsystems, with the influence and thus indirect responsibility being mediated by the teacher. interface st (student-teacher) comprises all the microsystems in which a specific student interacts with his or her teacher and both may feel or be held responsible. within the macrocontext of school, however, these are not the only microsystems the student and his/her teacher share. as the model focuses its representation on one specific student, another conceptual component of the model are the classmates of the respective student, who as a collective can also be attributed certain responsibilities which have to be differentiated from those attributed to the respective student. thus, in this model, microsystems, which the student and teacher share with the student‟s classmates (e.g., a lesson), are located in interface stc (students-teacher-classmates). the microsystems located here are not only classroom settings but also any incident in which these three actors are involved and can thus be held responsible. correspondingly, microsystems located in interface st may be set in the classroom because teachers might only interact with one student although being in the classroom with the whole group. as these groups of students can exist without the respective student being part of them and thus has to be conceptually differentiated from them, interface sc (student-classmates) comprises those microsystems shared by the student in focus and his or her classmates. this interface does not necessarily have to be embedded in the school context, as students are often friends and meet outside school. also, just k. helker & m. wosnitza 122 | f l r like in the teacher‟s case, the student‟s mesosystem involves microsystems he or she does not share with anyone or with people outside school like in a sports club, which can also be assumed to have an impact on what happens in other microsystems this student is involved in. thus, he or she might feel or be held responsible for objects located there. classmates were not explicitly attributed responsibilities, as they represent the students as a group. besides teachers and students, parents were identified to constitute the third agent in the school context. parents can be assumed to mainly be involved in the context of school when they share microsystems with either their child or their child‟s teacher. thus, interface sp (student-parents) subsumes those microsystems in which the student interacts with his or her parents, e.g. if parents help their child with their homework or ask about school over dinner. there are also microsystems, in which parents, their child and their child‟s teacher interact (interface stp, student-teacher-parents) and thus feel or are held responsible, this area being most strongly addressed by current research on parent involvement. furthermore, there are also microsystems, which parents only share with their child‟s teacher (e.g., parents‟ evenings), which are located in interface tp (teacher-parents). of course, one could argue that parents might also participate in microsystems which their child, his/her teacher and the classmates participate in and thus an interface was missing from this model. we propose, however, that if parents participate in their child‟s classroom (and also as attendants on field trips etc.) there is a change of roles by which parents assume the role of a teacher. these microsystems should thus also be located in the interfaces st and stc. this consideration indicates that responsibilities attributed to teachers, students and parents in the school context can be assumed to result from the specific roles a person is incorporating which invites the attribution of certain tasks and liabilities. furthermore, we hypothesise that the perception of what microsystems a person is involved or not involved in strongly affects what responsibilities are attributed to him or her. 5. validation of the framework suggested for structuring responsibility in the school context empirical studies into the issue of responsibility as were presented above have produced results that can be linked to further explicate some aspects of this framework of responsibility in the school context. no studies, however, have, to our knowledge, yet addressed the issue in its full complexity. therefore, the proposed framework shall be examined along data from several empirical studies in order to examine whether the model is useful to account for questions regarding the issue of responsibility attributions between teachers, students and parents. to meet this goal, the aim of this paper is to examine whether the proposed framework is useful and adequate for structuring responsibility attributions in the school context. furthermore, the study will look at the nature of the interfaces of teachers‟, students‟ and parents‟ mesosystems and examine what responsibilities these agents attribute to themselves and each other. 5.1 method the data presented in the following to support the above presented framework for structuring responsibility in the school context result from a series of studies each exploring the matter of teacher, student and parent responsibility in german secondary education. included studies and specific foci: (1) online survey with students about teacher responsibility, including lauermann & karabenick‟s (2013) teacher responsibility scale k. helker & m. wosnitza 123 | f l r (2) online survey with students about student responsibility, including a newly-developed student responsibility scale along the lines of lauermann & karabenick (2013) (3) online survey with parents about teacher responsibility, including lauermann & karabenick‟s (2013) teacher responsibility scale (4) pen and paper survey with parents of one local school (highest educational track) on shared and individual responsibility of teachers, students and parents (5) pen and paper survey of teachers of the highest educational track on shared and individual responsibility of teachers, students and parents (6) pen and paper survey of teachers of all educational tracks on shared and individual responsibility of teachers, students and parents (7) pen and paper survey of students of the highest educational track on shared and individual responsibility of teachers, students and parents. although every one of these studies had a different focus, they all contained at least one open-ended question each asking participants about their understanding of responsible teacher‟s, student‟s or parent‟s behavior (e.g., “what behavior characterizes a responsible teacher?”). some of the questionnaires contained three questions about all three agents, some of the studies only focused on student responsibility and thus only asked participants to characterize responsible students‟ behavior. the phrasing of the question aimed at catching broad descriptions of perceived teachers‟, students‟ and parents‟ responsibility. participants could name as many responsibilities as they liked for each agent. each of these mentioned responsibilities was later coded and counted separately. statements from all studies were organized into nine groups regarding the perspective (e.g., if respondents were teachers) and focus (e.g., statement about student responsibility) of the statement. table 1 provides an overview of the characteristics of these so-combined groups, characteristics of the sample and the number of statements in the perspective indicated (e.g., first cell: 68 teachers of which 58.8% were female and 42.2% were younger than or 40 years old provided 177 statements about teacher responsibility.). also, total numbers of statements per respondent group and subjects of responsibility are presented. table 1 overview of samples regarding perspective teachers‟ view of… students‟ view of… parents‟ view of… total # of statements teacher responsibility n=68; ♀ 58.8%; age: 42.4% ≤ 40years statements: 177 n=610; ♀ 60.3%: age: m=14.6 sd=2.4 statements: 1475 n=162; ♀ 88.9%; child: ♀ 54.9%; age: m=12.8 sd=2.1 statements: 535 2187 student responsibility n=68; ♀ 58.8%; age: 42.4% ≤ 40years statements: 161 n=279; ♀ 59.3%: age: m=15.1 sd=2.6 statements: 763 n=106; ♀ 85.8%; child: ♀ 57.5%; age: m=12.6 sd=0.8 statements: 364 1288 parent responsibility n=68; ♀ 58.8%; age: 42.4% ≤ 40years statements: 143 n=164; ♀ 58.9%: age: m=15.1 sd=3.0 statements: 405 n=106; ♀ 85.8%; child: ♀ 57.5%; age: m=12.6 sd=0.8 statements: 316 864 total # of statements 481 2643 1215 4339 all 4339 statements regarding teachers‟, students‟ and parents‟ responsibility were analysed using nvivo10 software for qualitative data analysis. all data were coded by a second coder and intercoder agreement was 74.1%. k. helker & m. wosnitza 124 | f l r data were coded into categories representing the six interfaces (see fig.1: st, stc, sc, sp, stp, tp) of these agents‟ mesosystems (i.e. what these people are feeling or being held responsible for in these specific microsystems.). during the coding process it became obvious that teachers, students and parents are often attributed general responsibilities that they are responsible to fulfill in all microsystems they are involved in (e.g. for being honest and trustworthy) that could not be coded into one specific interface. while these statements could have been coded into each of the above categories, as the specific person is stated to always be responsible for this object, the coders decided to code them separately in order to adequately test the framework. thus, general responsibilities being attributed to a person to be fulfilled in all the microsystems in which they are involved, were categorized into three groups representing teachers‟, students‟ and parents‟ mesosystems (i.e., sum of their microsystems). these main categories also included data on the respective person‟s responsibility for interactions with people not included in the framework for reasons of complexity (like colleagues, other parents, friends outside school etc.) in order to learn more about the responsibilities of each of the nine main categories, data in these were in a second step further categorized into sub-categories representing the different objects of responsibility in order to empirically describe the main categories. in some of these main categories, a further sub-division of statement was not necessary, due to the data varying regarding their levels of differentiation and depth. statements regarding the influences between different microsystems and also clear-cut distinctions between different areas of involvement were double-coded in an additional category for further analyses. 6. results of the overall 4339 statements about teacher, student and parent responsibility, 3993 statements could clearly be attributed to one of nine main categories suggested by the proposed framework. the remaining 346 could not be coded for reasons of ambiguity, incomprehensibility or lack of relevance regarding the research topic. the three categories covering students‟, teachers‟, and parents‟ general responsibility include those statements concerning what the specific agent is responsible for in all the microsystems they are involved in. furthermore, six main categories represent responsibilities resulting from the agents‟ shared microsystems, i.e. the interfaces of their mesosystems (see fig 1). results showed that no person was attributed responsibility in contexts in which they were not involved (e.g., parents were not attributed responsibilities that could be categorised in interface stc) and with few exceptions, all people involved were, however, attributed responsibility in the contexts in which they are involved. these exceptions concern interfaces stp and tp, in which students do not attribute any responsibility whatsoever to teachers. comparing the number of statements in the nine main categories, representing mesosystems and interfaces of these, and emphasis of a group‟s statements, with 35.2% of statements by students, the most prominent theme in the students‟ statements was the interface stc, the interaction of the teacher with their students. for the statements by teachers, the main emphasis lies on four categories, namely, students‟ (23.1% of teachers‟ statements) and teachers‟ (19.5%) general responsibilities, interface stc (18.2%) and sp (19.7%). statements by parents focused on students‟ general responsibility (20.1%) as well as interfaces stc (25.0%) and sp (19.3%). just focusing on the interfaces, stc and sp show to be the most frequently mentioned. in a second step, data attributed to these categories were analysed further for emergent themes. this section will start out by describing teachers‟, students‟ and parents‟ general responsibilities, i.e. those things, these three agents feel or are held responsible for in all the school-related microsystems they are involved in. furthermore, microsystems which these agents do not share with others or with people beyond the focus of this approach will be indicated. describing these three main categories will provide an overview of teachers‟, students‟ and parents‟ responsibilities which do not result from their being involved in a specific k. helker & m. wosnitza 125 | f l r microsystem they share with one of the other agents but from their generally being involved in the school context. these results will be followed by the presentation of the objects of responsibility arising from the shared microsystems of teachers, students and parents. the six categories representing the interfaces of teachers‟, students‟ and parents‟ mesosystems all contain responsibilities. the people involved are attributed these responsibilities because they interact with the other person(s) in this context. 6.1 teacher responsibility regarding teachers‟ general responsibility they feel or are expected by others to fulfill in all microsystems they are involved in, 16 categories were identified in the teachers‟, students‟ and parents‟ data. table 2 provides an overview of the categories as well as total numbers of categorized statements and percentages of each respondents‟ (teachers, students and parents) group. table 2 categories, frequencies and percentages of teachers’ general responsibility. categories teachers are generally responsible for… total # of statements teachers’ statements students’ statements parents’ statements …being attentive, empathic, caring and compassionate. 116 (19.66%) 32.94% 15.78% 22.14% …appearing nice and friendly in their interactions with others. 109 (18.47%) 2.35% 25.67% 8.40% ...being honest, reliable and trustworthy. 69 (11.69%) 12.94% 10.70% 13.74% …having pedagogical, methodological and content knowledge (showing teaching competence). 58 (9.83%) 11.76% 8.56% 12.21% …being ready, motivated, willing and trying hard to teach their students. 57 (9.66%) 11.76% 7.75% 13.74% …preparing lessons. 36 (6.10%) 10.59% 5.35% 5.34% …their relations with students, parents and other teachers in general. 32 (5.42%) 4.71% 5.35% 6.11% …being helpful. 31 (5.25%) 8.02% 0.76% …their self-reflection. 18 (3.05%) 3.53% 2.41% 4.58% …being patient. 17 (2.88%) 1.18% 2.94% 3.82% …doing their job and fulfilling their duties. 11 (1.86%) 3.53% 1.34% 2.29% …being well-organised. 10 (1.69%) 2.41% 0.76% …being open. 8 (1.36%) 3.35% 0.53% 2.29% …sticking to the given rules. 8 (1.36%) 2.14% … their cooperation and relations with other teachers. 5 (0.85%) 0.80% 1.53% …getting advanced training. 5 (0.85%) 1.18% 0.27% 2.29% total 590 85 (100%) 374 (100%) 131 (100%) teachers were on this general level attributed responsibilities connected to their fulfilling of their job and its requirements or were attributed responsibility for their relations with other agents (students, parents, colleagues) in the school context and for partially personal qualities. thus, teachers were attributed the responsibility for being attentive, empathic, caring and compassionate (e.g., “showing interest in every student and getting this across to the students” (t#13tr1)), which was the main category in teachers‟ and 1 abbreviation “t#13tr“ means “teacher no. 13 about teacher responsibility”. k. helker & m. wosnitza 126 | f l r parents‟ numbers of statements regarding general teacher responsibility. in the students‟ statements, the teacher responsibility for appearing nice and friendly in their interaction with others was identified as the main theme. further responsibilities regarding their social interactions, were teachers‟ responsibility for being honest, reliable and trustworthy and for being helpful and also being patient in their interactions with others (e.g., “being patient with every student (even if it‟s hard)” (s#141tr)) regarding teachers‟ responsibility for objects more related to their job, teachers were attributed the responsibility for doing their job and fulfilling their duties, having distinctive pedagogical, methodological and content knowledge (e.g., “high social and pedagogical competence” (p#114tr); “ability to teach us a lot” (s#2tr)), sticking to the given rules (e.g., “a responsible teacher has to also meet the rules set for the students (no mobile phone, chewing gum, etc.)” (s#113tr)) and being well-organized (e.g., “not being sloppy/ forgetful” (s#514tr)). furthermore, teachers‟ responsibility for being ready, motivated, willing and trying hard to teach their students (e.g., “love children and his job” (p#111tr)) was identified in the data. independently of others, teachers are also involved in microcontexts they do not share with the named agents, but in which they are attributed the responsibility for preparing lessons and getting advanced training (e.g., “be ready and motivated to get advanced training regarding contents and pedagogical matters” (t#36tr)). the data suggested another microsystem, i.e. teachers‟ cooperation and relations with other teachers (e.g., “cooperate with colleagues” (s#105tr); “not insulting other teachers” (s#512tr)). 6.2 student responsibility altogether, 17 general responsibilities of students were identified. these are represented in table 3 including the numbers of statements overall as well as per group of respondents. these general responsibilities, i.e., objects of student responsibility in all microsystems they are involved in, could be subdivided into 17 sub-categories. table 3 categories, frequencies and percentages of students’ general responsibility. categories students are generally responsible for… total # of statements teachers’ statements students’ statements parents’ statements …being ready, motivated, willing and trying their best to learn and be successful. 157 (22.14%) 21.51% 19.60% 27.06% …monitoring and adapting their own learning progress and study. 89 (12.55%) 9.68% 15.58% 8.26% …having positive relations with other people. 86 (12.13%) 8.60% 11.06% 15.60% …doing their homework. 74 (10.44%) 8.60% 13.07% 6.42% …doing their work and fulfilling their duties. 62 (8.74%) 11.83% 8.04% 8.72% …being honest and reliable. 35 (4.94%) 7.53% 3.52% 6.42% …school matters in general. 30 (4.23%) 3.23% 2.76% 7.34% …being self-reliant and take responsibility. 29 (4.09%) 5.38% 4.27% 3.21% …turning to others when needing help and accepting help given. 26 (3.67%) 4.30% 2.76% 5.05% …sticking to the rules. 22 (3.10%) 4.30% 3.27% 2.29% …their self-reflection. 21 (2.96%) 4.30% 2.51% 3.21% …preparing exams. 19 (2.68%) 1.08% 4.27% 0.46% …accuracy and order. 16 (2.26%) 2.15% 2.26% 2.29% …preparing lessons and revise contents taught. 16 (2.26%) 1.08% 3.52% 0.46% …getting good grades. 11 (1.55%) 1.08% 2.51% 0.00% …being open. 10 (1.41%) 2.15% 0.25% 3.21% …students are responsible to engage in sports 6 (0.85%) 3.23% 0.75% 0.00% k. helker & m. wosnitza 127 | f l r or social clubs outside school. total 709 93(100%) 398 (100%) 218 (100%) all three respondent groups, teachers, students and parents, mentioned students‟ responsibility for being ready, motivated, willing and trying their best to learn and be successful (e.g., “being ambitious to achieve the best possible result” (p#69sr)) most often. students also frequently mentioned their own responsibility to monitor and adapt their learning progress and study (e.g., “revise contents” (s#206sr)) as well as doing their homework. students were attributed the responsibility for doing their work and fulfilling their duties (e.g., “fulfilling even displeasing tasks” (t#1sr)), sticking to the rules, being self-reliant (e.g., “being self-reliant” (s#109sr)) and school matters in general (e.g., “attending school on a regular basis” p#43sr)). furthermore, students are held responsible for having positive relations with other people (e.g., “having a pronounced social behavior” (t#43sr)) and being honest and reliable (e.g., “keeping agreed dates (e.g. for projects)” (s#152sr)). regarding specifically students‟ school work, the responsibility for getting good grades but also for turning to others when needing help and accepting help given appear in the data (e.g., “notifying parents and teachers about problems and accepting help” (s#128sr)). independently of others, students alone are responsible for working and studying at home, which includes completing tasks and preparing for exams, but also monitoring their own learning success and study accordingly (e.g., “recognizing when more has to be done for school” (t#7sr)). apart from school matters, some data suggested students‟ responsibility to further engage in sports or social clubs outside school (e.g., “[a responsible student] shows social commitment in his free time (youth fire fighters, sports club)” (t#44sr)). 6.3 parent responsibility just like all other microsystems parents are involved in, the number of statements attributed to this category was comparatively low with 5.3% of teachers‟, 0.7% of students‟ and 14.5% of parents‟ statements. however, five categories of parents‟ responsibility could be identified as well as differences regarding how strongly these categories were emphasized by each respondent group. table 4 categories, frequencies and percentages of parents’ general responsibility. categories parents are generally responsible for… total # of statements teachers’ statements students’ statements parents’ statements …getting involved and engage in school matters. 31 (44.29%) 70.83% 23.53% 34.48% …being interested. 23 (32.86%) 16.67% 64.71% 27.59% …being objective, diplomatic and able to conciliate. 10 (14.29%) 12.50% 11.76% 17.24% …having a positive attitude towards school. 3 (4.29%) 0.00% 0.00% 10.34% …their interactions with other parents. 3 (4.29%) 0.00% 0.00% 10.34% total 70 24 (100%) 17 (100%) 29 (100%) the data suggested that parents‟ general responsibilities were for getting involved and engage in school (e.g., “cooperation with the school” (t#27pr); “participating in school life” (t#53pr)). this responsibility for getting involved in school was the most dominant one in teachers‟ statements (70.83%) and also strongly emphasized by parents themselves (34.48%) who also stressed parents‟ being responsible for being interested (27.59%), a responsibility being particularly stressed in the students‟ statements (64.71%). k. helker & m. wosnitza 128 | f l r furthermore, parents were stated to have the general responsibility for having a positive attitude towards school (e.g., “have a positive attitude towards school and lessons” (p#10pr)), being objective and diplomatic regarding school matters (e.g., “being objective towards teachers” (p#50pr)) and a specific microsystem that emerged from the data was parents‟ interactions with other parents‟ for which they are also stated to be responsible (e.g., “cooperation and communication with other parents” (p#62pr)). 6.4 interface st – student-teacher to this main category, all statements were assigned which concerned the respective student‟s interactions with their teacher independent of other people. while every teacher has a large number of students, this category represents those responsibilities attributed to teachers in every context they interact with an individual student. this interface is different from interface stc in that it only includes responsibilities that derive from a specific student interacting with a specific teacher. these dyadic interactions may, in fact, be also set in the classroom but do exist irrespective of co-students being involved (as is the case for responsibilities coded in interface stc). thus, one could hypothetically assume these responsibilities to also derive from microsystems that students share with their teachers and classmates (stc) because teachers might only interact with one student although being in the classroom with the whole group. the two categories are, however, conceptually different and thus treated separately here. table 5 provides an overview of teachers‟ responsibilities in their interactions with their students. table 5 categories, frequencies and percentages of teachers’ responsibility in their interactions with their students. categories teachers are responsible for… total # of statements teachers’ statements students’ statements parents’ statements …developing a positive, caring, interested personal relationship with every student. 136 (32.69%) 16.67% 32.25% 37.07% …listening to, helping, counselling students when they have problems inor outside school. 117 (28.13%) 8.33% 30.80% 25.86% …enhancing every single student‟s learning. 87 (20.91%) 33.33% 19.20% 22.41% …treating students with respect. 42 (10.10%) 12.50% 9.78% 10.34% …being a role model. 22 (5.29%) 16.67% 5.43% 2.59% …developing student‟s personality. 8 (1.92%) 12.50% 1.09% 1.72% …the individual student‟s grades. 4 (0.96%) 0.00% 1.45% 0.00% total 416 24 (100%) 276 (100%) 116 (100%) teachers are described by in students‟ and parents‟ statements as most responsible for developing a positive, caring, interested, personal relationship with every student. students furthermore emphasizing teachers‟ responsibility to listen to, help and counsel their student whenever he/she turns to them with problems inor outside school (e.g., “to be responsive to students‟ problems in and outside school.” (t#61tr)). regarding school outcomes, teachers‟ responsibility to enhance every single student‟s learning (e.g., “adapt to different types of learners and try for all of them to have the same chances.” (s#114tr)) lies in the center of teachers‟ statements about their own responsibility. teachers are furthermore described as responsible for being a role model for this student (e.g., “a responsible teacher exemplifies positive behavior (meet deadlines, being organized, on time, fair…)” (t#26tr)) and treating the student with respect. other responsibilities attributed to the main category of teachers‟ general responsibilities were the ones for the student‟s grades (only mentioned by students) and the development of their personality (e.g., “co-education, e.g. what‟s good or not for a student at the moment and in the future” (p#39tr)). of the mentioned, the personal relationship lies more in the center of students‟ statements than do educational matters. k. helker & m. wosnitza 129 | f l r regarding students‟ responsibility in their interaction with the teacher, four categories could be identified in the data. table 6 provides an overview. table 6 categories, frequencies and percentages of students’ responsibility in their interactions with their teacher. categories students are responsible for… total # of statements teachers’ statements students’ statements parents’ statements …treating teacher with respect, accept their authority. 30 (73.17%) 100.00% 65.22% 81.25% …contacting teacher with problems and questions. 6 (14.63%) 17.39% 12.50% …not provoking the teacher. 4 (9.76%) 17.39% …trusting the teacher. 1 (2.44%) 6.25% total 41 2 (100%) 23 (100%) 16 (100%) most statements in all three groups of respondents‟ statements concern students‟ being responsible to treat their teacher with respect and accept his/her authority (e.g., “respect the teacher, even if they haven‟t earned it” (s#105sr)). this may include the sub-category of not provoking teachers, which was only mentioned by students. furthermore, students were held responsible for contacting the teacher with their problems and questions (e.g., “if you have a problem, dare to ask the teachers” (s#47sr)). 6.5 interface stc – teacher-all students interface b, comprising all those microsystems in which the teacher interacts with all students, holds all those statements regarding teaching the specific lesson. within the data categorized into this main category, 14 teacher (table 7) and 7 student responsibilities (table 8) were identified. table 7 categories, frequencies and percentages of teachers’ responsibility in their interactions with all their students. categories teachers are responsible for… total # of statements teachers’ statements students’ statements parents’ statements …teaching good, interesting lessons adapted to their students. 262 (27.29%) 21.57% 28.02% 26.29% …showing authority and be strict – consequent classroom management. 200 (20.83%) 9.80% 22.70% 17.37% …being fair. 129 (13.44%) 15.69% 11.78% 18.31% …contributing to a positive learning atmosphere. 61 (6.35%) 13.73% 6.47% 4.23% …coming to class on time and prepared. 52 (5.42%) 7.84% 5.89% 3.29% …keeping calm and continue lessons. 50 (5.21%) 1.96% 5.60% 4.69% …intervening in students‟ conflicts and work for group‟s team spirit. 43 (4.48%) 0.00% 4.02% 7.04% …motivating their students. 42 (4.38%) 11.76% 2.01% 10.33% …(objective) assessment. 31 (3.23%) 3.92% 3.88% 0.94% …involving all students in lessons. 26 (2.71%) 3.92% 2.30% 3.76% …supervising their students. 24 (2.50%) 3.92% 2.87% 0.94% …caring for the group and being their 17 (1.77%) 1.96% 2.16% 0.47% k. helker & m. wosnitza 130 | f l r students‟ advocate. …teaching the contents required, follow the curriculum. 12 (1.25%) 3.92% 0.86% 1.88% …appropriately using homework. 11 (1.15%) 0.00% 1.44% 0.47% total 960 51(100 %) 696 (100%) 213 (100%) table 8 categories, frequencies and percentages of students’ responsibility in their interactions with their teacher and other students. categories students are responsible for… total # of statements teachers’ statements students’ statements parents’ statements …participating in lessons and pay attention. 143 (48.81%) 35.48% 52.13% 45.95% …coming to class on time and prepared. 47 (16.04%) 22.58% 15.96% 13.51% …contributing to a positive working atmosphere. 38 (12.97%) 0.00% 14.36% 14.86% …respectful relations with classmates and teachers in class. 36 (12.29%) 29.03% 7.45% 17.57% …being interested in the contents taught. 18 (6.14%) 3.23% 5.85% 8.11% …trying to understand and learn the contents taught. 8 (2.73%) 9.68% 2.66% 0.00% …trying to meet the expectations the teacher sets for the class. 3 (1.02%) 0.00% 1.60% 0.00% total 293 31 (100%) 188 (100%) 74 (100%) while both teachers and students are attributed the responsibility in this area to come to class on time and prepared (e.g., “being well-prepared” (t#38tr)) and to create/contribute to a positive working atmosphere in class (e.g., “to try not to disturb the lessons in order not to ruin others learning success” (s#250sr)), other responsibilities differ. regarding teacher responsibility, the category most of all three groups of respondents‟ statements could be assigned to, however, was the responsibility for teaching good and interesting lessons that are adapted to the students (e.g., “adapting lessons to students” (t#7tr)). also, especially by their students, teachers are attributed responsibilities that can be linked to issues of classroom management (showing authority and being strict (e.g., “clear instructions and strict implementation” (t#57tr)). further objects of teacher responsibility were supervising the students, intervene in students‟ conflicts and establish the group‟s team spirit, keeping calm and continuing lessons no matter what happens (e.g., “sometimes just turning a blind eye” (t#43tr))). further personal relations between the teacher and their class like caring for the group and being their students‟ advocate (e.g., “a responsible teacher backs their students” (s#378tr)) as well as being fair. the latter is by some statements linked to teachers‟ responsibility to include all students in lessons (e.g., “a responsible teacher tries to get quiet students out of their shell” (s#145tr)) and (objective) assessment (e.g., “clear, transparent and explained grading” (s#113tr)). regarding these teaching aspects, teachers were also attributed the responsibility to teach the contents required (follow the curriculum) and appropriately setting homework and motivate the students (e.g., “get children excited about learning” (p#103tr)). responsibilities attributed to students in this respect somewhat complemented the above, with the students‟ responsibility for participating and paying attention in lessons (e.g., “show interest in the lesson and actively participate” (p#16sr)) being the most emphasized responsibility in all three groups of k. helker & m. wosnitza 131 | f l r respondents‟ statements. furthermore, students were stated to be responsible for trying to be interested in the contents (e.g., “interest in the contents and other students‟ answers” (t#54sr)), understanding and learn the contents (e.g., “truly trying to understand the contents rather than simply marking lesson time” (t#33sr)) and meeting expectations. statements regarding students‟ responsibility to treat their classmates and teachers in lessons with respect were also included in this category. 6.6 interface sc – student-classmates in contrast to the above, this category includes all statements regarding students‟ responsibility in their interaction with their classmates. this category was conceptually different from stc because, although it can be assumed to address classroom settings in which the teacher is also present, these are responsibilities that derive from students‟ interactions with their co-students, irrespective of whether the teacher is around. only one category was evident, namely students‟ responsibility to have positive, helpand peaceful relations with their classmates (e.g., “not bully or even hurt anyone” (p#59sr); “save others from mean people” (s#46sr)). table 9 categories, frequencies and percentages of students’ responsibility in their interactions with their classmates. categories students are responsible for… total # of statements teachers’ statements students’ statements parents’ statements …their interactions with their classmates. 134 (100.00%) 100.00% 100.00% 100.00% total 134 17 (100%) 88 (100%) 29 (100%) 6.7 interface sp – student-parents this main category subsumes all those statements regarding those responsibilities resulting from students‟ interactions with their parents. no statements focused on students‟ responsibilities in this area. table 10 provides an overview. table 10 categories, frequencies and percentages of parents’ responsibility in their interactions with their child. categories parents are responsible for… total # of statements teachers’ statements students’ statements parents’ statements …supporting and helping their child. 177 (26.50%) 22.47% 27.93% 25.79% …having a positive and caring relationship with their child. 124 (18.56%) 19.10% 17.04% 20.81% …keeping tabs on their child‟s learning and school life. 119 (17.81%) 24.72% 15.08% 19.46% …learning with their child. 57 (8.53%) 4.49% 11.73% 4.98% …educating their child (besides school matters). 52 (7.78%) 14.61% 5.59% 8.60% …motivating their child. 48 (7.19%) 3.37% 6.98% 9.05% …providing the prerequisites for student learning. 38 (5.69%) 5.62% 5.31% 6.33% …giving their child space and freedom. 31 (4.64%) 3.37% 6.42% 2.26% …physical and structural care for their child. 22 (3.29%) 2.25% 3.91% 2.71% total 668 89 (100%) 358 (100%) 221 (100%) k. helker & m. wosnitza 132 | f l r parents in this category were attributed broader responsibilities as for generally supporting and helping them (e.g., “support and encourage” (t#1pr)), a category which about a quarter of each group of respondents statements were assigned to. a quarter of the teachers‟ statements in this overall area also mentioned parents‟ responsibility for keeping the tabs on their child‟s learning and school life (e.g., “say when and how much i have to learn or do my homework” (s#50pr)), which was not equally strongly emphasized by students and parents. further general responsibilities assigned to parents were for having a positive and caring relationship with their child (e.g., “take time for their child” (p#3pr); “take notice of me” (s#15pr)) and giving their child space and freedom (e.g., “no pressuring expectations” (s#101pr)). further responsibilities mentioned could be subdivided into responsibilities concerning the child‟s care and education (educating the child (e.g., “parents have educated their children at home and taught them values and norms.” (t#25pr)); providing physical and structural care (e.g., “send children well-prepared (fed, well-rested, low pressure) to school” (p#71pr))) as well as responsibilities related to the child‟s school work. thus, parents were stated to be responsible to provide the prerequisites for student learning (materials, working atmosphere at home (e.g., “parents are responsible for creating the ideal conditions so that the child can show the best performance in school.” (p#77pr))), motivating their child and learn with their child (e.g., “learn with the child” (p#88pr)). while 668 statements could clearly be categorized as regarding parents‟ responsibility in their interaction with their child, no statement whatsoever could be identified as describing students‟ responsibility in these microsystems they share with their parents. 6.8 interface stp – student-teacher-parents just as in the above interface, there are also no responsibilities attributed to students in the microcontexts they share with their teachers and parents. table 11 categories, frequencies and percentages of teachers’ responsibility in their interactions with the students and their parents. categories teachers are responsible for… total # of statements teachers’ statements students’ statements parents’ statements …their interactions with students and their parents. 5 (100.00%) 100.00% 100.00% total 5 1 (100%) 4 (100%) table 12 categories, frequencies and percentages of parents’ responsibility in their interactions with the students and their parents. categories parents are responsible for… total # of statements teachers’ statements students’ statements parents’ statements …their interactions with their children and their teachers. 14 (100.00%) 100.00% 100.00% total 14 3 (100%) 11 (100%) at the sp interface, exchange between these participants is found in the data to mostly concern students‟ learning difficulties or underachievement. students are not seen to have much responds whereas teachers are stated to be responsible to be “ready to communicate about problems with student and parents” k. helker & m. wosnitza 133 | f l r (p#99tr)) just as parents are (e.g., “in case there is a problem (e.g., underachievement, bullying) search for a solution together with the child and the teacher” (p#81pr)). 6.9 interface tp – teacher-parents this interface comprises all statements regarding responsibilities resulting from teachers‟ and parents‟ interaction. table 13 and 14 provide an overview. table 13 categories, frequencies and percentages of teachers’ responsibility in their interactions with the students’ parents. categories teachers are responsible for… total # of statements teachers’ statements students’ statements parents’ statements …their interactions with their students‟ parents. 27 (100.00%) 100.00% 100.00% total 27 2 (100%) 25 (100%) table 14 categories, frequencies and percentages of parents’ responsibility in their interactions with their child’s teachers and respective numbers of statements. categories parents are responsible for… total # of statements teachers’ statements students’ statements parents’ statements …their interactions with their children‟s teachers. 66 (100.00%) 100.00% 100.00% 100.00% total 66 18 (100%) 14 (1000%) 34 (100%) regarding the interaction of parents with their child‟s teacher, both teachers and parents are attributed the responsibility to communicate and cooperate with the other, contacting one another when there is a problem that needs solving (e.g., “involve parents” (p#5tr); “accessible for parents” (p#95tr); “inform parents at an early stage” (p#62tr); “work with, not against the teachers” (t#5pr); “use parents‟ nights” (t#37pr)). 6.10 exosystems the framework suggested that a microsystem an agent is involved in, may be influenced by exosystems, i.e. microsystems this person is not part of. the data provided some examples of these indirect influences between microsystems. especially parents fulfilling their responsibility at home (i.e., interface sp) were often stated to influence other microsystems: “improve the relations between the teacher and the child by encouraging the child.” (p#12pr); “if their child complains about a teacher or a subject, taking that seriously but not taking the same line but trying to find solutions” (p#70pr). students were stated to be responsible for sometimes acting as their classmates‟ advocate (“standing up for other students‟ needs and issues in the face of teachers and classmates” p#2sr). regarding teachers‟ responsibility, one parent claimed that a teacher was responsible “to not, if they have not managed to get through with their lesson plans, shift the contents as homework to the parents.” (p#159tr). furthermore, one student described how much their teacher‟s interaction with other teachers influenced the microsystems in interface b: “he is responsible for not going into a cover lesson and have other teachers say bad things about the class beforehand but he/she is responsible for judging the group by themselves.” (s#604tr). k. helker & m. wosnitza 134 | f l r 7. discussion aim of this paper was to study in how far the here developed framework is useful to structure responsibility attributions in the school context. to address this issue, data from a series of studies were coded into the postulated structure of the model in order to learn about the usefulness of the model and the nature of its elements. this analysis has provided the necessary starting point for studies of the complexity of responsibility in the school context based on the here proposed framework. overall 84 objects of responsibility in the school context were identified in this study. all of these responsibilities could be allocated to specific spheres of interaction of teachers, students and parents. the results showed that the proposed framework of responsibility in the school context is adequate for structuring responsibilities that teachers, students and parents attribute to themselves and one another in the different microsystems they share in the school context. in accordance with the theoretical considerations, that a person‟s responsibility is determined by the specific role this person incorporates (leithwood et al., 1999), the qualitative data showed that all three agents were only attributed responsibility in those contexts, in which they are actively involved. accordingly, as suggested in the literature (e.g., duff, 1998) no responsibilities for objects in specific microsystems were attributed to a person who is not involved in this microsystem, i.e., no person was assigned responsibility for objects beyond their reach. thus, no parent responsibilities were mentioned in the data for goings-on in specific lessons or classroom settings just as teachers were not attributed responsibilities for students‟ learning at home. this finding may, however, vary between different cultural contexts, school systems and naturally different stages of schooling (e.g., primary vs. high school) and should thus be focused on in further research including more heterogeneous groups in this respect the data revealed, that in some areas, agents involved are not attributed any responsibility as for example in interface sp, which students share with their parents. with students mentioning significantly less responsibilities that could be categorized into this area than teachers and parents did, it can be hypothesized that students do not perceive this area as one in which they are responsible agents. the fact, that also neither teachers nor parents mentioned any student responsibilities here, supports this area‟s minor role in student responsibility. it however constitutes the category into which most of all three agents‟ statements on parents‟ responsibility were coded. regarding parents‟ responsibilities in interactions with teachers (interfaces stp and tp), the data support ramirez‟ (1999) findings that teachers attribute responsibility for parent-teacher interactions more to parents than to themselves. also, the main focus of the statements categorized into one of the shared microsystems sometimes differed between the three groups of respondents. this was most obvious with regard to interface st, which comprises teachers‟ and students‟ responsibilities resulting from their interaction. while a third of teachers‟ statements focus on their responsibility to enhance every single student‟s learning, students and parents emphasise teachers‟ responsibility to develop a positive, caring and interested personal relationship with the students in about a third of their statements. these examples show how tensions may occur between various actors emphasizing different responsibilities or goals (lauermann & karabenick, 2011). the limits of each of the described interfaces were found to be rather clear-cut which becomes specifically apparent between interfaces sc and sp. thus, a teacher is attributed the responsibility for settling conflicts between their students when he or she realizes a problem (“keeping an eye on their students, not letting everything go just using the excuse that students should and could „sort it out by themselves‟” (p#129tr)). students however believe that teachers “should not always intervene. some things have to be left for the students to settle.” (s#262tr). thus, whether teachers fulfill their responsibility for settling conflicts among students depends on whether they think students are able to do it themselves, i.e. whether this matter is part of the microsystem only the students share, or not. this example also shows that responsibilities are not attributed once for all times. this interplay between different microsystems is further supported by the close alignment of teacher and student responsibilities in interfaces st and stc, with teachers being attributed the responsibility for k. helker & m. wosnitza 135 | f l r enhancing every single student‟s learning in interface st, which affects their responsibility for involving all students in lessons and teaching lessons that are adapted to the level of all of their students. the data also reveal the fulfillment of responsibilities in different microsystems to depend on whether responsibilities in others are met. this for example holds true for teachers‟ and students‟ responsibility for coming to class prepared, which naturally subsumes these agents‟ responsibility to work and prepare classes at home, i.e. in another microcontext. in this specific context but also generally, responsibilities identified in this study showed to be far more wide-ranging than was assumed in previous research (e.g., bourke. 1990; lewis, 2001). the finding of data revealing influences between different microsystems is closely connected to these considerations, as the possibility of a person‟s actions to (in-)directly influence other microsystems despite them not being involved in them, raises their responsibility in those microsystems they actively engage in. despite these influences having been discussed in a lot of literature on, in the broadest sense, the (social) contexts of school (e.g. epstein, 2011), the question of shared responsibility has not yet been posed. the data clearly reveal that the general responsibility for advancing student learning, all three agents share, is in fact built up of many different issues different agents are responsible for taking care of in the microsystems they are involved in. everyone is responsible for doing their share in their sphere of action, i.e. their mesosystem. this also holds true if the different microsystems a person is involved in may influence one another by means of this person‟s involvement in them, i.e. serve as exosystems to one another. thus, the data revealed a teacher being held responsible for acting in a specific way (e.g., not letting other teachers influence them regarding their opinion of the students in order to not letting these prejudices influence them in their interactions with these students.) although object and addressee of responsibility are not located in the same microsystem in this case. in regard to student responsibility, parallels between bacon‟s (1993) study of student responsibility and findings in this study can be drawn. students in bacon‟s study mentioned they felt responsible for “doing the job” and “obeying the rules”, i.e. being held responsible rather than feeling personally responsible. being a student thus comes with certain responsibilities which resemble those of an employee: punctuality, having their working material on them, discipline, trying to meet expectations, working diligently and orderly etc. interestingly, most of these responsibilities are alike for teachers and students. besides this employee‟s role, students were also found to have certain responsibilities as a learner, which go beyond the school context and are not closely controlled. as a learner, a student is responsible for showing interest and motivation, trying to understand and learn what is taught in lessons and studying the materials at home. despite these two roles‟ potential overlap, they may also become quite contradictory for students being controlled for fulfilling responsibilities by their parents and teachers. in this respect, it seems obvious that parents‟ responsibility to keep the tabs on students‟ learning conflicts with their responsibility for giving their child some space and freedom. however, this conflict should also be noted as a limitation of this study, as this study has only looked at data from teachers, students and parents involved in secondary schooling. responsibility attributions can be assumed to differ in primary school settings with parents sharing a larger part of their children‟s school lives. microsystems parents share with teachers may thus become more important and also more extended regarding responsibility attributions. comparable influences on responsibility attributions can be expected for the social, cultural and economic context of the school which have been suggested by prior research. after generally studying the usefulness of the framework for structuring responsibility in the school context in this study, these aspects should definitely be addressed in further research in order to identify influences, especially because the analyses presented here only included data from german teachers, students and parents. this study has provided an overview of what responsibilities are attributed to teachers, students and parents on different levels of the school context and how these can be structured with a theoretical framework. this framework will also allow to more closely analyse responsibility attributions within the school context and the influences these might have on other contexts, i.e. microsystems. the results of the preliminary study of teacher, student and parent responsibility, especially the multitude of named shared objects of each agent‟s responsibility, strongly suggest the need for further, also k. helker & m. wosnitza 136 | f l r quantitative research. this research might provide insights into the extent to which teachers, students and parents attribute responsibility for specific objects to themselves and one another. focusing on family-school mesosystem christenson (2004) claims that the interface between home and school may be “strong for some families, weak for others, and non-existent for others.” (p. 87). thus, further research will have to focus on different types of teachers, students and parents showing specific levels and constellations of responsibility for student motivation, learning and achievement. also, the relevance of role-taking has been shown in these preliminary analyses and shall thus be extended with regard to individual roles in a context of other agents as for example suggested by goffman (1959) or other sociological concepts of interactions (e.g., weber, 2005). furthermore, existing literature has suggested studying the role the school setting and other structural factors play with regard to responsibility perception (e.g., lee & loeb, 2000; thrupp et al., 2003). consequently, future research will have to focus on the interplay of teachers‟, students‟ and parents‟ attributions of responsibility and by specifically linking these three agents exploring patterns of responsibility attributions and their influence on student achievement and motivation. keypoints a heuristic framework for structuring responsibility in the school context is developed. the framework is validated by coding teachers‟, students‟ and parents‟ statements on their own and others‟ responsibility into the suggested model. results support the usefulness of the framework and reveal the objects of teachers‟, students‟ and parents‟ responsibility in the school context and influences. acknowledgements the authors would like to thank judith fränken for the help with coding the data and sue beltman for her many invaluable comments on the manuscript. references auhagen, a. e. (1999). die realität der verantwortung [the reality of responsibility]. göttingen: hogrefe. auhagen, a. e., & bierhoff, h.-w. (eds.). (2001). responsibility the many faces of a social phenomenon. london: routledge. bacon, c. s. (1991). being held responsible versus being responsible. the clearing house, 64, 395-398. bacon, c. s. (1993). student responsibility for learning. adolescence, 28(109), 199–212. ballard, k., & bates, a. (2008). making a connection between student achievement, teacher accountability, and quality classroom instruction. qualitative report, 13(4), 560-580. barrett, a. m. (2005). teacher accountability in context: tanzanian primary school teachers' perceptions of local community and education administration. compare, 35(1), 43-61. doi: 10.1080/03057920500033530 bastian, j. (1995). verantwortung: pädagogik zwischen freiheit und verbindlichkeit [responsibility: pedagogy between freedom and obligation]. pädagogik, 47(7-8), 6–10. bayertz, k. (1995). eine kurze geschichte der herkunft der verantwortung [a short history of the origin of responsibility]. in k. bayertz (ed.), verantwortung (pp. 3–71). darmstadt: wiss. buchges. berkowitz, l., & daniels, l. r. (1963). responsibility and dependency. journal of abnormal and social psychology, 66(5), 429–436. doi: 10.1037/h0049250 k. helker & m. wosnitza 137 | f l r bierhoff, h. w. (1995). verantwortungsbereitschaft, verantwortungsabwehr und verantwortungszuschreibung [disposition for responsibility, denial of responsibility and attribution of responsibility]. in k. bayertz (ed.), verantwortung (pp. 217-240). darmstadt: wiss. buchges. bierhoff, h. w. (2000). skala der sozialen verantwortung nach berkowitz und daniels: entwicklung und validierung [the social responsibility scale by berkowitz and daniels: development and validation]. diagnostica, 46(1), 18–28. doi: 10.1026//0012-1924.46.1.18 bourke, s. (1990). responsibility for teaching: some international comparisons of teacher perceptions. international review of education, 36(3), 315-327. doi: 10.1007/bf01876000 bracci, e. (2009). autonomy, responsibility and accountability in the italian school system. critical perspectives on accounting, 20(3), 293-312. doi: 10.1016/j.cpa.2008.09.001 bronfenbrenner, u. (1979). the ecology of human development: experiments by nature and design. cambridge, ma: harvard university press. doi: 10.1525/aa.1981.83.3.02a00220 bryan, l. a., & mclaughlin, h. j. (2005). teaching and learning in rural mexico: a portrait of student responsibility in everyday school life. teaching and teacher education, 21(1), 33-48. doi: 10.1016/j.tate.2004.11.004 christenson, s. l. (2004). the family-school partnership: an opportunity to promote the learning competence of all students. school psychology review, 33(1), 83-104. doi: 10.1521/scpq.18.4.454.26995 del schalock, h. (1998). student progress in learning: teacher responsibility, accountability, and reality. journal of personnel evaluation in education, 12(3), 237–246. doi: 10.1023/a:1008063126448 diamond, j. b., randolph, a., & spillane, j. p. (2004). teachers' expectations and sense of responsibility for student learning: the importance of race, class, and organizational habitus. anthropology & education quarterly, 35(1), 75–98. doi: 10.1525/aeq.2004.35.1.75 duff, r. a. (1998). "responsibility" routledge encyclopedia of philosophy (vol. 8, pp. 290–294). eikenbusch, g. (2009). classroom management für lehrer und für schüler: wege zur gemeinsamen verantwortung für den unterricht [classroom management – for teachers and students: ways to a shared responsibility for teaching]. pädagogik, 61(2), 6–10. epstein, j. l. (2011). school, family, and community partnerships preparing educators and improving schools (2 ed.). boulder, co: westview press. epstein, j. l., & sheldon, s. b. (2006). moving forward: ideas for research on school, family, and community partnerships. in c. f. conrad & r. serlin (eds.), the sage handbook for research in education: engaging ideas and enriching inquiry (pp. 117-137). london: sage. feiks, d. (1992). zur pädagogischen verantwortung des lehrers [about the pedagogical responsibility of teachers]. lehren und lernen, 18(4), 1–20. fischman, w., dibara, j. a., & gardner, h. (2006). creating good education against the odds. cambridge journal of education, 36(3), 383–398. doi: 10.1080/03057640600866007 galindo, c., & sheldon, s. b. (2012). school and home connections and children's kindergarten achievement gains: the mediating role of family involvement. early childhood research quarterly, 27, 90-103. doi: 10.1016/j.ecresq.2011.05.004 gärtner, h. (2010). wie schülerinnen und schüler ihre lernumwelt wahrnehmen: ein vergleich verschiedener maße zur übereinstimmung von schülerwahrnehmungen [how students perceive their learning environment: a comparison of four indices of interrater agreement]. zeitschrift für pädagogische psychologie, 24(2), 111–222. doi: 10.1024/1010-0652/a000009 goffman, e. (1959). the presentation of self in everyday life. new york: doubleday. grotlüschen, a. (2008). verantwortung und verantwortungsabwehr bei der zusammenarbeit mit bildungsfernen schichten [responsibility and denial of responsibility in working with educationally disadvantaged social strata]. in h. pätzold (ed.), verantwortungsdidaktik: zum didaktischen ort der verantwortung in erwachsenenbildung und weiterbildung (pp. 95–112). baltmannsweiler: schneiderverlag hohengehren. gurtner, j.-l., monnard, i., & genoud, p. (2001). towards a multilayer model of context and its impact on motivation. in s. volet & s. järvelä (eds.), motivation in learning contexts: theoretical advances and methodological implications (pp. 189-208). amsterdam: pergamon. k. helker & m. wosnitza 138 | f l r guskey, t. r. (1981). measurement of the responsibility teachers assume for academic successes and failures in the classroom. journal of teacher education, 32(3), 44–51. doi: 10.1177/002248718103200310 guskey, t. r. (1982). differences in teachers' perceptions of personal control of positive versus negative student learning outcomes. contemporary educational psychology, 7(1), 70–80. doi: 10.1016/0361476x(82)90009-1 halvorsen, a.-l., lee, v. e., & andrade, f. h. (2009). a mixed-method study of teachers' attitudes about teaching in urban and low-income schools. urban education, 44(2), 181-224. doi: 10.1177/0042085908318696 hamilton, l. (1978). who is responsible? toward a social psychology of responsibility attribution. social psychology, 41(4), 316-328. doi: 10.2307/3033584 helker, k., & wosnitza, m. (2014). verantwortung im schulkontext – ein systematisches review des empirischen forschungsstandes [responsibility in the school context – a systematic literature review of the state of empirical research]. unterrichtswissenschaft, 42(3), 261 – 279. höffe, o. (2008). verantwortung [responsibility]. in o. höffe (ed.), lexikon der ethik (pp. 326–327). münchen: beck. kammerl, r. (2008). divergente verantwortungszuschreibungen als problemfeld beruflicher ausund weiterbildung [diverging responsibilty attributions as a problem area of vocational education and advanced training]. in h. pätzold (ed.), verantwortungsdidaktik: zum didaktischen ort der verantwortung in erwachsenenbildung und weiterbildung (pp. 31–48). baltmannsweiler: schneiderverlag hohengehren. karakaya, s. (2004). a comparative study: english and turkish teachers' conceptions of their professional responsibility. educational studies, 30(3), 195-216. doi: 10.1080/0305569042000224170 katyal, k. r., & evers, c. w. (2007). parents partners or clients? a reconceptualization of home-school interactions. teaching education, 18(1), 61-76. doi: 10.1080/10476210601151573 kaufmann, f.-x. (1995). risiko, verantwortung und gesellschaftliche komplexität [risk, responsibility and social complexity]. in k. bayertz (ed.), verantwortung (pp. 72–97). darmstadt: wiss. buchges. korkmaz, i. (2007). teachers' opinions about the responsibilities of parents, schools, and teachers in enhancing student learning. education, 127(3), 389-399. lauermann, f., & karabenick, s. (2011). taking teacher responsibility into account(ability): explicating its multiple components and theoretical status. educational psychologist, 46(2), 122-140. doi: 10.1080/00461520.2011.558818 lauermann, f., & karabenick, s. (2013). the meaning and measure of teachers' sense of responsibility for educational outcomes. teaching and teacher education, 30, 13-26. doi: 10.1016/j.tate.2012.10.001 lawson, m. a. (2003). school-family relations in context: parent and teacher perceptions of parent involvement. urban education, 38(1), 77-133. doi: 10.1177/0042085902238687 lee, v. e., & loeb, s. (2000). school size in chicago elementary schools: effects on teachers' attitudes and students' achievement. american educational research journal, 37(1), 3-31. doi: 10.3102/00028312037001003 leichter, h. j. (ed.). (1974). the family as educator. new york: teachers college press. leithwood, k., edge, k., & jantzi, d. (1999). educational accountability: the state of the art. gütersloh: bertelsmann. lenk, h. (1992). zwischen wissenschaft und ethik [between science and ethics]. frankfurt am main: suhrkamp. lenk, h., & maring, m. (1993). verantwortung normatives interpretationskonstrukt und empirische beschreibung [responsibility – normative construct of interpreattion and empirical description]. in l. h. eckensberger (ed.), ethische norm und empirische hypothese (pp. 222–243). frankfurt am main: suhrkamp. lewis, r. (2001). classroom discipline and student responsibility: the students' view. teaching and teacher education, 17(3), 307-319. doi: 10.1016/s0742-051x(00)00059-7 lightfoot, s. l. (1978). worlds apart: relationships between families and schools. new york: basic books. marjoribanks, k. (1979). families and their learning environments an empirical analysis. london: routledge. k. helker & m. wosnitza 139 | f l r matteucci, m. c., & gosling, p. (2004). italian and french teachers faced with pupil's academic failure: the "norm of effort". european journal of psychology of education, 19(2), 147-166. maulbetsch, c. (2010). person und verantwortung zur grundlegung einer pädagogischen handlungstheorie unter dem aspekt der erziehung zur verantwortung im kontext schule [person and responsibility – about laying the basis for a pedagogical theory of cation regarding the aspect of educating for responsibility in the school context]. münster: waxmann. müller, f. (2009). verantwortung für sich selbst übernehmen: arbeitsmotivation und spielräume im berufsalltag [taking responsibility for oneself: work motivation and scopes in the work routine]. pädagogik, 61(10), 32–33. pätzold, h. (2008). vom professionellen umgang mit verantwortung [about the professional handling of responsibility]. in t. rihm (ed.), teilhaben an schule (pp. 253–264). wiesbaden: vs verl. für sozialwiss. peterson, e. r., rubie-davies, c. m., elley-brown, m. j., widdowson, d. a., dixon, r. s., & irving, e. s. (2011). who is to blame? students, teachers and parents views on who is responsible for student achievement. research in education, 86(1), 1-12. potvin, p., & papillon, s. (1992). teacher's sense of responsibility towards student achievement and their attitude. canadian journal of special education, 8(1), 33-42. ramirez, a. y. f. (1999). survey on teachers' attitudes regarding parents and parental involvement. school community journal, 9(2), 21-39. rose, j., & medway, f. (1981a). measurement of teachers' beliefs in their control over student outcome. journal of educational research, 74(3), 185–189. rose, j., & medway, f. (1981b). teacher locus of control, teacher behaviour and student behaviour as determinants of student achievement. journal of educational research, 74(6), 375–381. ryan, r., & deci, e. (2000). self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. american psychologist, 55(1), 68–78. doi: 10.1037/0003-066x.55.1.68 schleißheimer, b. (1984). die verantwortung des erziehers: vorüberlegungen zu einer ethik pädagogischen handelns [the responsibility of the educator: pre-considerations about ethics of pedagogical actions]. vierteljahresschrift für wissenschaftliche pädagogik, 60(1), 1–17. tenorth, h.-e. (2004). lehrerarbeit: strukturprobleme und wandel der anforderungen [teacher work: structural problems and change in demands]. in u. beckmann, h. brandt & h. wagner (eds.), ein neues bild vom lehrerberuf? (pp. 14–25). weinheim: beltz. thrupp, m., mansell, h., hawksworth, l., & harold, b. (2003). "schools can make a difference" but do teachers, heads and governors really agree? oxford review of education, 29(4), 471-484. doi: 10.1080/0305498032000153034 weber, m. (2005). wirtschaft und gesellschaft [economy and society]. frankfurt: zweitausendeins. weiner, b. (1995). judgments of responsibility: a foundation for a theory of social conduct. new york: guilford press. werner, m. (2006). verantwortung [responsibility]. in m. düwell (ed.), handbuch ethik (pp. 541–548). stuttgart, weimar: metzler. wosnitza, m., & beltman, s. (2012). learning and motivation in multiple contexts: the development of a heuristic framework. european journal of psychology of education, 26(2), 177-193. doi: 10.1007/s10212-011-0081-6 zimmerman, b. j., & kitsantas, a. (2005). homework practices and academic achievement: the mediating role of self-efficacy and perceived responsibility beliefs. contemporary educational psychology, 30(4), 397-417. doi: 10.1016/j.cedpsych.2005.05.003 frontline learning research vol.5 no. 3 special issue (2017) 31 42 issn 2295-3159 elizabeth a. krupinski, phd, department of radiology & imaging sciences, emory university 1364 clifton rd ne d107 atlanta, ga30322, united states. email: ekrupin&emory.edu. doi: http://dx.doi.org/10.14786/flr.v5i3.250 receiver operating characteristic (roc) analysis elizabeth a. krupinski emory university, usa article received 13 april / revised 23 march / accepted 23 march / available online 14 july abstract visual expertise covers a broad range of types of studies and methodologies. many studies incorporate some measure(s) of observer performance or how well participants perform on a given task. receiver operating characteristic (roc) analysis is a method commonly used in signal detection tasks (i.e., those in which the observer must decide whether or not a target is present or absent; or must classify a given target as belonging to one category or another), especially those in the medical imaging literature. this frontline paper will review some of the core theoretical underpinnings of roc analysis, provide an overview of how to conduct an roc study, and discuss some of the key variants of roc analysis and their applications. keywords: receiver operating characteristic analysis, roc, observer performance, visual expertise http://dx.doi.org/10.14786/flr.v5i3.250 krupinski | f l r 32 1. introduction visual expertise can be measured or assessed in a variety of ways, but in many cases it is the behavioral outcome related to visual expertise that is of ultimate concern. for example, in medical imaging (e.g., radiology, pathology, telemedicine) a physician must view an image (e.g., x-ray exam, pathology slide, photograph of a skin lesion) and render a diagnostic decision (e.g., tumor present or absent), prognosis (e.g., malignant vs benign) and/or a recommendation for a treatment plan (e.g., additional exams or surgery). in the military, radar screens need to be monitored for approaching targets (e.g., missiles) and decisions made as to whether they are friend or foe and whether to escalate the finding to the next level of action. there are many other situations where these types of visual detection and/or classification tasks take place in real life, and where investigations take place to assess how expertise impacts these decisions and what we can do to improve them. in medicine this is particularly important as when decision errors are made they have direct and significant impacts on patient care and well-being [1-2]. the problem with “real life” is that is often very difficult to determine how well the interpreter is actually performing. feedback is often not provided and when it is it is often separate in time from the actual decision making event – often quite disparate in time, making the feedback less impactful. compounding the problem is that there is often not a single correct answer. thus, the majority of performance assessments are done in the context of research and are often focused on assessing observer performance in the context of comparing a new technique or technology to an existing one. to deal with the complex nature of decision interpretation tasks where it is important to understand and balance the consequences both correct and incorrect decisions, receiver operating characteristic (roc) analysis is a very valuable tool. it is important throughout this review to keep in mind that performance is not a constant. one’s performance on any given task will vary as a function of a number of factors and thus the metrics and principles discussed will reflect those changes and differences. for example, when someone first learns a decision task, their criteria for rendering a decision may be based more on didactic learning that they have engaged in and the exact examples of the task they have encountered to date. as expertise grows and they encounter more and varied examples, the criteria they use in their decisions is likely to change as their knowledge and skill develops. even when one is an expert, changes in the environment, the stimuli and the consequences of the decisions rendered may change, making it necessary to adjust one’s criteria to the new situation. for example, on a chest x-ray image cancer looks like a white spot(s) on the lung and even residents in training quickly learn to detect and diagnose lung cancer. however, in the southwest united states and other desert regions there is a condition known as valley fever (in infection caused by the fungus coccidioidomycosis) that appears in the lungs as a white spot(s). radiologists who move to places like arizona where valley fever is quite common go through a period of criteria adjustment – initially calling nearly everything cancer (high false positive rate) but soon learning to distinguish valley fever from lung cancer as they see more cases and learn the discriminating features. 2. basics of decision making for roc roc analysis was developed in the early 1950s based on principles from signal-detection theory for evaluation of radar operators in the detection of enemy aircraft and missiles [3-4], and additional contributions were thereafter made by researchers in engineering, psychology, and mathematics [5-7]. roc was introduced into medicine in the 1960s by lee lusted [8-11], with significant efforts devoted to gaining a better understanding of decision-making [12-15]. this entrée into medicine was the result of a series of studies in radiology that began soon after world war ii to determine which of four radiographic and fluoroscopic techniques (e.g., radiographic film vs slides) was better for tuberculosis (tb) screening [16-17]. the goal was to find a single imaging technique that would outperform all the others (in other words allow krupinski | f l r 33 the radiologists to reach the correct decision). instead what they found was that the intra-observer and interobserver variation was unexpectedly so high that it was impossible to determine which one was best. this was unexpected as until then it was presumed that given the same image data all radiologists looking at the images would be seeing the same diagnostic features and information, detecting the tb if it was present, and rendering the same diagnostic decision. the idea that everyone may “see” something different in the images, perhaps as a function of experience or expertise, had never been considered. thus, it was necessary to build systems that could generate better images so radiologists’ performance could improve (i.e., reduce observer variability), and develop methods to evaluate these new systems and assess their impact on observer performance. although there are newer methods that allow for more than two decisions in the roc task environment [18-19], roc is traditionally a binary decision task – target/signal (e.g., lesion, disease, missile) present versus target/signal absent, or in the case of classification rather than detection the target/signal belongs to class 1 (e.g., cancer, enemy) or class 2 (e.g., not cancer, friend). for roc analysis, these two conditions must be mutually exclusive. there must also be some sort of “truth” or gold standard for each option. in radiology for example, pathology is often used as the gold standard. in cases where there is no other definitive test or method for determining the truth, panels of experts are often used to establish the gold standard [20-21]. given the truth and the decisions of the observers in the study, a 2x2 table readily summarizes all four possible decisions: true positive (tp) (target present, observer reports as present), false negative (fn) (target present, observer reports as absent), false positive (fp) (target absent, observer reports as present), and true negative (tn) (target absent, observer reports as absent). the tp and tn decisions are correct while the fn and fp decisions are incorrect. 2.1 common performance metrics suppose you have an observer who is an expert skilled at visually detecting a specific poisonous frog in the jungle versus a similar but non-poisonous frog. knowing that she can make the correct decision is important because she is your guide on an expensive eco-jungle tour and if she says a given cute little frog is not poisonous but in reality is, you might reach out to touch it with potentially fatal consequences. in radiology, one of the common sources of litigation is mammography – mammographers either missing potential breast cancers or overcalling benign conditions as malignant causing undo stress and anxiety. real life examples of important decision tasks abound, all of which require careful assessment of correct and incorrect decisions. from the basic 2x2 matrix of 4 decisions described above come some key metrics often used to assess performance in visual expertise and other observer performance studies. the two most commonly used are sensitivity and specificity. sensitivity reflects the ability of the observer to correctly classify the target present stimuli (e.g., x-ray or other images) and is calculated as: sensitivity = tp/(tp + fn) (2.1) specificity reflects the ability of the observer to correctly classify the target absent stimuli and is calculated as: specificity = tn/(tn + fp) (2.2) when you combine these decisions, a measure of accuracy can be obtained as: accuracy = (tp + tn)/ (tp + fn + tn + fp) (2.3) two other common metrics that are used in medicine are positive and negative predictive value: positive predictive value (ppv) = tp/(tp + fp) (2.4) negative predictive value (npv) = tn/(tn + fn) (2.5) krupinski | f l r 34 in general, there is a trade-off between sensitivity and specificity – as you increase one you decrease the other. in other words, if you want to detect more targets (high sensitivity) it often occurs as a result of making more false positives (decreased specificity). why would you want to use sensitivity/specificity versus ppv/npv? basically, the former are independent of the prevalence of targets in the case sample while the latter are not. an example might be useful. suppose you have an observer who is an expert skilled at visually detecting a specific poisonous frog in the jungle versus a similar but non-poisonous frog and her sensitivity is 95% and specificity 80%. in jungle #1 there are 1000 frogs total with a prevalence of 50% poisonous (n = 500). in jungle #2 there are also 1000 frogs total but only 25% are poisonous (n = 250). based on this, our observer has the following performance levels. it can be seen that depending on which jungle the observer study is conducted, the performance even of an expert observer will differ. jungle #1: tp = 475 fn = 25 fp = 100 tn = 400 accuracy = (475 + 400)/(475 + 25 + 100 + 400) = 0.88 or 88% ppv = 475/(475 + 100) = 0.83 or 83% npv = 400/(400 + 25) = 0.94 or 94% jungle #2: tp = 238 fn = 12 fp = 150 tn = 600 accuracy = (238 + 600)/(238 + 12 + 150 + 600) = 0.84 or 84% ppv = 238/(238 + 150) = 0.61 or 61% npv = 600/(600 + 12) = 0.98 or 98% in many cases sensitivity and specificity are more than adequate measures of performance for visual search tasks, but it becomes complicated when the test sets contain cases with a range of difficulty levels. for example, in radiology some bone fractures are very obvious and thus easy to detect but others are quite subtle and can readily be missed. in the frog example, a bright red frog in a green jungle is likely easier to detect than a light green frog in a dark green jungle. in cases that are not obvious the decision as to whether or not the target is actually there becomes less certain and observers may not be willing to give a binary yes/no present/absent decision, but may be more willing to report their decision as a function of confidence, for example reporting a target (or lack of target) as definitely present, probably present or possibly present. in other words, the observer’s decision threshold for reporting can change as a function of many things, including but not limited to the nature of the target, target prevalence, background complexity within which the target is embedded, number and type of targets, and observer experience or expertise. this is where roc analysis becomes useful. 2.2 the roc curve even the visual expert may not perform as one would think without delving further into the nature of the task. for example, in radiology decision thresholds can change within and between observers as a function of the nature of the task and its consequences. in chest ct exam interpretation a radiologist may adopt a very conservative threshold for reporting possible abnormalities in the lungs that could be cancer nodules but are less than 5 mm in size because they know that ct is typically the best imaging exam to do (i.e., no other follow-up imaging options) and obtaining a biopsy on such a small target is very difficult (potentially leading to a pneumothorax or puncture of the lung) and unlikely to yield a specimen large enough to get an accurate biopsy on. instead of reporting it, the radiologist may recommend a 6-month watch period and another ct exam. in mammography however, small lesions are easier to biopsy and there is no risk of puncturing a lung or other vital structure, so mammographers tend to be more liberal in their reporting of potential cancerous findings. this comes at the risk of more false positives but before doing a biopsy additional x-ray images or an ultrasound is often recommended, reducing the impact of a false positive even further. krupinski | f l r 35 roc analysis and the resulting roc curve is a method that captures the relationship (i.e., trade-offs) between sensitivity and specificity as well as the range of decision thresholds that every observer has no matter what their level of expertise and experience. the roc curve (figure 1) is a graphical representation of this relationship, plotting sensitivity (the tp fraction) versus 1-specificity (the fp fraction or 1 – tn/(tn + fp) = fp/(fp + tn)) for all possible decision thresholds. the axes go from 0 to 1 since sensitivity and specificity are typically represented as proportions. the diagonal line is chance performance or guessing and the curves indicate increasingly better performance moving to the upper left corner which represents perfect performance. thus in terms of visual expertise, one would expect that those with more expertise would tend to have curves closer to the upper left corner. the fact that those with more expertise tend to have curves closer to the upper left corner is true, but as noted above much depends on the decision thresholds that the individual observer has for a given task. each point on a given curve represents a specific tp and fp fraction or decision threshold setting. for example, on the “good” curve, the plus sign represents a rather conservative decision maker – not reporting a lot of targets (low sensitivity) but also not making a lot of false positives (high specificity or low 1specificity). the cross sign represents a more liberal decision maker – reporting a lot of targets (high sensitivity) but making a lot of false positives (low specificity or high 1-specificity). referring back to the example of the radiologist moving to arizona and overcalling valley fever as cancer, the figure could just as well represent his/her learning or adjustment period, in which the lower curve is when they first move and overall cancers and the upper curve is their performance after a few months of seeing more exemplars of valley fever so better discrimination is possible. figure 1. example of a typical roc curve plotting sensitivity vs 1-specificity. the diagonal line is chance performance or guessing and the curves indicate increasingly better performance moving to the upper left corner which represents perfect performance. 2.3 how to plot an roc curve it is rare that someone actually plots an roc curve by hand as there are a number of software programs available both in commercial statistical packages and freely downloadable from research web sites. however, it is useful to see how the confidence ratings discussed above lead to the generation of an roc curve. thus, suppose you ran a visual detection experiment and the observer was required to identify whether or not a given image had a target (n = 50) or not (n = 50) and then report confidence in that decision as definite, probable or possible. this yields 6 categories of responses where 1 = absent, definite and 6 = present, definite. you can then create a table (table 1) showing the distribution of responses as a function of truth (whether or not the image actually contains a target). it should be noted that continuous rating scales (e.g., 0 – 100) can be used as well and methods exist for generating operating points (decision thresholds) and plotting these as well [22]. krupinski | f l r 36 table 1 an example of the distribution of confidence scores for a subject in an observer performance study with a 6point confidence scale and images with a target present or absent (truth). truth 1 2 3 4 5 6 present 2 3 2 5 20 18 absent 16 15 10 4 3 2 the sensitivity, specificity and fp fraction can then be determined at each threshold or cutoff point as in table 2. table 2 sensitivity, specificity and fp fraction can then be determined at each threshold or cutoff point. result positive is > sensitivity specificity fp fraction 2 probably absent 0.96 (48/50) 0.32 (16/50) 0.68 3 possibly absent 0.90 (45/50) 0.62 (31/50) 0.38 4 possibly present 0.86 (43/50) 0.82 (41/50) 0.18 5 probably present 0.76 (38/50) 0.90 (45/50) 0.10 6 definitely present 0.36 (18/50) 0.96 (48/50) 0.04 the roc plot can then be generated plotting sensitivity on the y-axis and 1-specificity (fp fraction) on the x-axis. figure 2. roc curve generated from the data in table 2. in terms of fitting an actual curve to the data points in the roc plot, there are a variety of methods. simply “connecting the dots” is the empirically based version but it creates a stepped or jagged plot. a smooth curve reflecting the theoretical “true” curve is much more desirable. there are basically two ways to approach generating the curve – parametric and non-parametric [22-28]. the non-parametric approach does not have any assumptions about the structure of the underlying data distribution and essentially smooths the histograms of the output data for the two classes. the parametric methods do rely on the validity of the underlying distribution assumptions. most researchers prefer the parametric approaches and much of the available software uses these approaches as well. 2.4 interpreting the roc curve there are a few key metrics used to interpret the roc curve and characterize observer performance. the most common one is the area under the curve (auc or az). as noted above the diagonal line in the roc plot represents chance or guessing performance and it clearly divides the roc space into two halves thus representing an auc of 0.5. the top left corner is perfect performance and encompasses all of the area krupinski | f l r 37 below it, thus auc is 1.0. any curve lying between chance and perfect performance will have a value between 0 and 1.0, with better performance having values closer to 1.0. as with generation of the curve itself, there are a variety of methods to calculate auc [22-29] and most programs use one of these methods. partial auc acknowledges that the more traditional auc is often not appropriate, as not all decision thresholds or operating points are equally important [30-31]. in other words, in real life observers may not actually operate at certain thresholds for one reason or another. in medicine for example, a diagnostic test with low specificity (a high false positive rate) may not be clinically acceptable. in this case it may be useful to select a given (acceptable) fp rate, determine its associated sensitivity (tp rate), and then calculate the area under the curve only up to that operating point (i.e., capturing only a part of the total auc). partial auc is very common in the development of computer-aided detection and discrimination algorithms for medical imaging. other metrics used less often are d´, de´, m,  and zk [32-33]. 2.5 comparing roc curves although a single roc curve is common, quite often studies are designed to compare performance, for example between experts and novices on a given visual task. thus there are two curves – one for experts and one for novices – and the question is whether there is a significant difference in performance (auc typically) or not. visually it is not always possible to tell if the difference is significant. this is especially true when the roc curves cross at some point (usually the upper right quadrant/corner) as in figure 3 [3435]. therefore, statistical methods have been developed, some for comparing only two curves and some for comparing multiple curves. again, there are parametric and non-parametric options [22-23, 36-40]. one of the most common methods for comparing multiple observers and multiple cases is the multi-reader multicase method developed by dorfman, berbaum and metz [39]. figure 3. example of 2 roc curves that cross. figure 4 is an example of the output from an mrmc analysis on a study that had 6 observers viewing a series of images in 2 conditions (different computer monitors for displaying medical images). the visual task was to search for subtle fractures in bone x-ray images. readers 1, 3 and 5 were expert radiologists and 2, 4 and 6 were resident trainees. the upper portion provides the auc values for each observer in each condition, followed by the difference between the two conditions. auc values are usually reported in publications out to three decimal points maximum. the lower portion shows the results of the analysis of variance (anova) comparing the aucs in the standard anova output format. the actual output document provides more information than provided here, such as the variance components and errorcovariance estimates, and different ways of treating the various variables (e.g., random readers and random cases, fixed readers and random cases, random readers and fixed cases), but this example shows how many available programs output relevant data comparing roc curves. the analyses can also take into account krupinski | f l r 38 level of expertise by comparing the two groups of observers as described above (in this case there were differences but it did not reach statistical significance). figure 4. an example of the output from an mrmc analysis. 3. other types of roc the discussion of roc analysis up to this point has been about tasks that involve the detection of one target and for the most part assess fps only as they occur in target absent images (again 1 per). in real life however, scenes and other visual stimuli often contain multiple targets and fps can occur in both target present and target absent stimuli. traditional roc analysis also typically does not ask or require the observer to locate the target once it has been detected. thus there is always some question (unless the target appears in a specific location (e.g., center of the display) every time) as to whether the observer actually detected the true target (tp) or called something else in the image (fp) thereby actually missing the true target (fn). the earliest attempt to account for location in roc tasks was the lroc (location roc). in this method, the observer provides a confidence rating that somewhere in the image there is a target, then marks the location of the most suspicious region [41-42]. lroc is an advantage over roc in that it takes location into account, but it still only allows for a single target. to account for multiple targets, free response roc (froc) was developed in which observers mark different locations and provide a confidence rating for each mark [43-46]. the problem with froc is that when the roc is plotted the x-axis, rather than going from 0 to 1 like the traditional rfoc curve, goes from 0 to infinity (based on however many fps are reported). this makes calculating the area under curve quite difficult and the comparison of two curves even more challenging. the alternative froc (afroc) method was developed to address issue, creating a plot that has both axes going from 0 to 1 [47]. the jackknife afroc (jafroc) method was then developed to allow for generalization to the population of readers and cases, in the same way that mrmc roc does [4446]. krupinski | f l r 39 4. other considerations in addition to deciding which type of roc analysis is best (which really depends on the hypotheses, nature of the task, types and number of images and targets), there are two other aspects that are typically important. as already discussed above, the truth or gold standard for cases must be known in advance. with simple psychophysical studies this is quite easy but for real life images (scenes, medical images, industrial images, satellite images) this can be more difficult. other considerations when selecting cases include: how subtle or obvious the targets are, how much and what type of background “noise” is in the image, where the targets are located (random or in specified locations), the size(s) of the targets and how much background image is included, how long the images will remain available for viewing, whether the images can be manipulated (e.g., zoom/pan or window/level) by the observers, and target prevalence as noted above. sample size is the other key issue with respect to setting up an roc study – how many images and observers are required to achieve adequate power once the study is completed. as with any other power calculation, sample size will depend on a number of factors including the metric under consideration (e.g., auc or partial auc) and the design (e.g., repeated measures with the same observers viewing the same images in two or more conditions or different readers viewing the images in different conditions). there are a number of key papers describing methods to calculate sample sizes for various study designs, many of which include representative tables showing sample sizes required for different power estimates [48-51]. some of the available roc programs also include a power calculator and some will provide power in the analysis output. 5. software programs it is not possible to list all of the available software programs for roc analysis as there are always new ones being released. there are however some reliable sites where the more commonly used programs can be found. the medical image perception laboratory web site [52] at the university of iowa and the roc software site at the university of chicago [53] contain the programs developed by that team (dorfman, berbaum, metz, hillis) including the mrmc roc, rocfit, labroc4, corroc, indroc, rocket, labmrmc, proproc, rscore, bigamma, rscore-j and sas programs to perform sample size estimates. software for froc, afroc and jafroc (chakraborty) analyses are available as well [54]. some commercial statistical software also has modules for roc analysis [55-59]. keypoints roc analysis provides metrics of observer performance in visual detection and discrimination tasks area under the curve (auc) is the most commonly used metric of performance common variants of roc analysis allow for multiple targets and location accuracy key study design issues include target characteristics and establishing a gold standard software is available to conduct roc analyses krupinski | f l r 40 references analyse-it. http://analyse-it.com/docs/220/method_evaluation/roc_curve_plot.htm last accessed april 13, 2016. birkelo, c.c., chamberlain, w.e., phelps, p.s. (1947). tuberculosis case finding. a comparison of the effectiveness of various roentgenographic and photofluorographic methods. jama, 133, (6), 359-366. pmid: 20281873 bunch, p.c., hamilton, j.f., sanderson, g.k. simmons, a.h. (1978). a free-response approach to the measurement and characterization of radiographic-observer performance. j appl photogr eng, 4,166– 171. chakraborty, d.p., berbaum, k.s. (2004). observer studies involving detection and localization: modeling, analysis, and validation. med phys, 31 (8), 2313-2330. doi: 10.1118/1.1769352 chakraborty, d.p. (2005). recent advances in observer performance methodology: jackknife free-response roc (jafroc). rad protect dosim, 114 (1), 26-31. doi: 10.1093/rpd/nch512 chakraborty, d.p. (2006). analysis of location specific observer performance data: validated extensions of the jackknife free-response (jafroc) method. acad radiol, 13 (10), 1187-1193. doi: 10.1016/j.acra.2006.06.016 chakraborty, d.p., winter, l.h.l. (1990). free-response methodology: alternate analysis and the new observer-performance experiment. radiol, 174 (3), 873-881. doi: 10.1148/radiology.174.3.2305073 delong, e.r., delong, d.m., clarke-pearson, d.l. (19880. comparing the areas under two or more correlated receiver operating characteristics curves: a non-parametric approach. biometrics, 44 (3), 837-845. pmid: 3203132 dev chakraborty’s froc web site. http://perception.radiology.uiowa.edu/software/receiveroperatingcharacteristicroc/tabid/120/defaul t.aspx last accessed april 13, 2016. dorfmam, d.d., alf, e. (1968). maximum likelihood estimation of parameters of signal detection theory: a direct solution. psychometrika, 33 (1), 117-124. doi: 10.1007/bf02289677 dorfman, d.d., alf, e. (1969). maximum likelihood estimation of parameters of signal detection theory and determination of confidence intervals – rating method data. j math psychol, 6 (3), 487-496. doi: http://dx.doi.org/10.1016/0022-2496(69)90019-4 dorfman, d.d., berbaum, k.s., metz, c.e. (1992). receiver operating characteristic rating analysis: generalization to the population of readers and patients with the jackknife method. invest radiol, 27(9), 723731. pmid: 1399456 dorfman, d.d., berbaum, k.s., metz. c.e., lenth, r.v., hanley, j.a., dagga, h.a. (1997). proper receiver operating characteristic analysis: the bigamma model. acad radiol, 4 (2), 138-149. pmid: 9061087 edwards, d.c. (2013). validation of monte carlo estimates of three-class ideal observer operating points for normal data. acad radiol, 20 (7), 908-914. doi: 10.1016/j.acra.2013.04.002 egan, j.p. (1975). signal detection theory and roc analysis. new york, ny: academic press. faraggi, d., reiser, b. (2002). estimation of the area under the roc curve. stats med, 21 (20), 3093-3106. doi: 10.1002/sim.1228 garland, l.h. (1949). on the scientific evaluation of diagnostic procedures. radiol, 52, (3), 309-328. doi: http://dx.doi.org/10.1148/52.3.309 green, d.m., swets, j.a. (1974). signal detection theory and psychophysics. huntington, ny: krieger publishers. hajian-tilaki, k.o., hanley, j.a., joseph, l., collet, j.p. (1997). a comparison of parametric and nonparametric approaches to roc analysis of quantitative diagnostic tests. med decis making, 17 (1), 94102. doi:10.1177/0272989x9701700111 hanley, j.a. (1988). the robustness of the “binormal” assumptions used in fitting roc curves. med decis making, 8 (3), 197-203. doi: 10.1177/0272989x8800800308 hanley, j.a., mcneil, b.j. (1983). a method for comparing the areas under receiver operating characteristic curves derived from the same cases. radiol, 148 (3), 839-843. doi: 10.1148/radiology.148.3.6878708 http://analyse-it.com/docs/220/method_evaluation/roc_curve_plot.htm https://doi.org/10.1118/1.1769352 http://dx.doi.org/10.1093/rpd/nch512 https://doi.org/10.1016/j.acra.2006.06.016 https://doi.org/10.1148/radiology.174.3.2305073 http://perception.radiology.uiowa.edu/software/receiveroperatingcharacteristicroc/tabid/120/default.aspx http://perception.radiology.uiowa.edu/software/receiveroperatingcharacteristicroc/tabid/120/default.aspx http://dx.doi.org/10.1016/0022-2496%2869%2990019-4 https://doi.org/10.1002/sim.1228 http://dx.doi.org/10.1148/52.3.309 https://doi.org/10.1177/0272989x9701700111 https://doi.org/10.1177/0272989x8800800308 https://doi.org/10.1148/radiology.148.3.6878708 krupinski | f l r 41 hanley, j.a., mcneil, b.j. (1982). the meaning and use of the area under a receiver operating characteristic (roc) curve. radiol, 143 (1), 29-36. doi: http://dx.doi.org/10.1148/radiology.143.1.7063747 institute of medicine. (1999). to err is human: building a safer health care system. washington, dc: national academy press. jiang, y., metz, c.e., nishikawa, r.m. (1996). a receiver operating characteristic partial area index for highly sensitive diagnostic tests. radiol, 201 (3), 745-750. doi: 10.1148/radiology.201.3.8939225 kundel, h.l., polansky, m. (1997). mixture distribution and receiver operating characteristic analysis of bedside chest imaging with screen-film and computed radiography. acad radiol, 4 (1), 1-7. pmid: 904086. lusted, l.b. (1960). logical analysis in roentgen diagnosis. radiol, 74,178-193. doi: http://dx.doi.org/10.1148/74.2.178 lusted, l.b. (1968). introduction to medical decision making. springfield, il: charles c. thomas publishers. lusted, l.b. (1969). perception of the roentgen image: applications of signal detection theory. rad clin n am, 7, 435-459. lusted, l.b. (1971). signal detectability and medical decision making. science, 171, (3977), 1217-1219. doi: 10.1126/science.171.3977.1217 mcclish, d.k. (1989). analyzing a portion of the roc curve. med decis making, 9 (3), 190-195. doi: 10.1177/0272989x8900900307 mcneil, b.j., adelstein, s.j. (1976). determining the value of diagnostic and screening tests. j nuc med, 17, (6), 439-448. pmid:1262961 mcneil, b.j., hanley, j.a. (1984). statistical approaches to the analysis of receiver operating characteristic (roc) curves. med dec making, 4, (2), 137-150. doi:10.1177/0272989x8400400203 mcneil, b.j., keeler, e., adelstein, s.j. (1975). primer on certain elements of medical decision making. ne j med, 293, (5), 211-215. doi: 10.1056/nejm197507312930501 metz, c.e., herman, b.a., shen, j.h. (1998). maximum-likelihood estimation of roc curves from continuously-distributed data. stats med, 17 (9), 1033-1053. pmid: 9612889 metz, c.e., kronman, h.b. (1980). statistical significance tests for binormal roc curves. j math psych, 22 (3), 218-243. doi: http://dx.doi.org/10.1016/0022-2496(80)90020-6 metz, c.e., pan, x. (1999). “proper” binormal roc curves: theory and maximum-likelihood estimation. j math psych, 43 (1), 1-33. doi: 10.1006/jmps.1998.1218 nakas, c.t. (2014). developments in roc surface analysis and assessment of diagnostic markers in threeclass classification problems. revstat – stat j, 12 (1), 43-65. ncss statistical software. http://www.ncss.com/software/ncss/procedures/ last accessed april 13, 2016. obuchowski, n.a. (1997). testing for equivalence of diagnostic tests. am j roentgen, 168 (1), 13-17. doi: 10.2214/ajr.168.1.8976911 obuchowski, n.a. (1994). computing sample size for receiver operating characteristic studies. invest radiol, 29 (2), 238-243. doi:10.2214/ajr.175.3.1750603 obuchowski, n.a. (2000). sample size tables for receiver operating characteristic studies. am j roentgen, 175 (3), 603-608. doi:10.2214/ajr.175.3.1750603 obuchowski, n.a. (2004). how many observers care needed in clinical studies of medical imaging? am j roentgen, 182 (4), 867-869. doi: 10.2214/ajr.182.4.1820867 peterson, w.w., birdsall, t.l., fox, w.c. (1954). the theory of signal detectability. ire prof gp in theory trans pgit, 4, (4), 171-212. doi: 10.1109/tit.1954.1057460 petrick, n., gallas, b.d., samuelson, f.w., wagner, r.f., myers, k.j. (2005). influence of panel size and expert skill on truth panel performance when combining expert ratings. proc spie med imag, 5749, 596286. doi: 10.1117/12.596286 schulman, k.a., kim, j.j. (2000). medical errors: how the us government is addressing the problem. curr control trials cardiovasc med, 1(1), 35-37. doi: 10.1186/cvm-1-1-035 spss statistics. http://www-03.ibm.com/software/products/en/spss-statistics last accessed april 13, 2016. starr, s.j., metz, c.e., lusted, l.b., goodenough, d.j. (1975). visual detection and localization of http://dx.doi.org/10.1148/radiology.143.1.7063747 https://doi.org/10.1148/radiology.201.3.8939225 http://dx.doi.org/10.1148/74.2.178 https://doi.org/10.1177/0272989x8900900307 https://doi.org/10.1177/0272989x8400400203 http://dx.doi.org/10.1016/0022-2496%2880%2990020-6 https://doi.org/10.1006/jmps.1998.1218 http://www.ncss.com/software/ncss/procedures/ https://doi.org/10.2214/ajr.168.1.8976911 https://doi.org/10.2214/ajr.175.3.1750603 https://doi.org/10.2214/ajr.175.3.1750603 https://doi.org/10.1109/tit.1954.1057460 https://dx.doi.org/10.1186%2fcvm-1-1-035 http://www-03.ibm.com/software/products/en/spss-statistics krupinski | f l r 42 radiographic images. radiol, 116 (3), 533-538. doi:10.1148/116.3.533 stata data analysis and statistical software. http://www.stata.com/features/overview/receiver-operatingcharacteristic/ last accessed april 13, 2016 swensson, r.g. (1996). unified measurement of observer performance in detecting and localizing target objects on images. med phys, 23 (10), 1709-1725. doi:10.1118/1.597758 swets, j.a. (1979). roc analysis applied to the evaluation of medical imaging techniques. radiol, 14 (2), 109-121. pmid: 478799 swets, j.a., dawes, r.m., monahan, j. (2000). psychological science can improve diagnostic decisions. psych sci public interest, 1 (1), 1-26. doi: 10.1111/1529-1006.001 swets, j.a., pickett, r.m. (1982). evaluation of diagnostic systems. methods from signal detection theory. new york, ny: academic press. tanner, w.p., swets, j.a. (1954). a decision-making theory of visual detection. psych rev, 61, (6), 401-409. pmid: 13215690 university of iowa medical image perception roc software. http://perception.radiology.uiowa.edu/software/receiveroperatingcharacteristicroc/tabid/120/defaul t.aspx last accessed april 13, 2016. university of chicago roc software. http://metz-roc.uchicago.edu/ last accessed april 13, 2016. wald, a. (1950). statistical decision functions. new york, ny: wiley, inc. zou, k.h., hall, w.j., shapiro, d.e. (1997). smooth non-parametric receiver operating characteristic (roc) curves for continuous diagnostic tests. stats med, 16 (19), 2143-2156. pmid: 9330425 zhou, x.h., obuchowski, n.a., mcclish, d.k. (2002). statistical methods in diagnostic medicine. new york, ny: wiley. medcalc statistical software. https://www.medcalc.org/manual/roc-curves.php last accessed april 13, 2016. https://doi.org/10.1148/116.3.533 http://www.stata.com/features/overview/receiver-operating-characteristic/ http://www.stata.com/features/overview/receiver-operating-characteristic/ https://doi.org/10.1118/1.597758 https://doi.org/10.1111/1529-1006.001 http://perception.radiology.uiowa.edu/software/receiveroperatingcharacteristicroc/tabid/120/default.aspx http://perception.radiology.uiowa.edu/software/receiveroperatingcharacteristicroc/tabid/120/default.aspx http://metz-roc.uchicago.edu/ https://www.medcalc.org/manual/roc-curves.php frontline learning research 5 special issue „learning through networks‟ (2014) 56-71 issn 2295-3159 corresponding author: matthieu.vaessen@ou.nl doi: http://dx.doi.org/10.14786/flr.v2i2.92 56 | f l r networked professional learning: relating the formal and the informal matthieu vaessenᵃ, antoine van den beemtᵇ, maarten de laatᵃ a welten institute, open university, heerlen, the netherlands b eindhoven school of education, eindhoven university of technology, the netherlands article received 17 february 2014 / revised 31 march 2014 / accepted 30 june 2014 / available online 15 july 2014 abstract the increasing complexity of the workplace environment requires teachers and professionals in general to tap into their social networks, inside and outside circles of direct colleagues and collaborators, for finding appropriate knowledge and expertise. this collective process of sharing and constructing knowledge can be considered 'networked learning'. the processes involved are informal and largely invisible to the official framework of the organisation. consequently, a large amount of learning that takes place is unrecognised and the dynamics, impacts and benefits of such networked learning are often overlooked by organisations. this situation brings about tensions between formal and informal processes, which in turn raise issues concerning adequate professional development, professional autonomy and management. it also leads to questions about facilitating the creation and exchange of knowledge and expertise within the existing social networks. taking an interdisciplinary approach, we explore a number of educational and organisational studies. our key questions are: what are the formal and informal mechanisms underlying networked professional learning, related to professional development, autonomy and management? how can networked learning be positioned in the most optimal way? currently, a clear academic understanding of how to optimally align and make use of networked learning is lacking. the goal of our exploratory review is to describe mechanisms that influence the alignment of informal and formal learning of teachers within their workplace: schools. we work towards a theoretical and practical integration of the different chosen fields by means of a framework of mechanisms related to networked learning. keywords: networked learning; teachers; professional development vaessen et al. 57 | f l r 1. introduction entering the 21st century, pervasive communication technologies together with increased attention for situational knowledge have resulted in an emphasis on collaboration and exchange, highlighting the importance of social networks both within organisations and across organisational borders (lieberman, 2000; price, 2013; pugh & prusak, 2013). making good use of social networks has become increasingly important in educational settings, where teachers develop relationships within and outside schools that help them to learn, solve problems, and innovate their teaching (de laat, 2012). access to networks resulting from these informal relationships has become an important aspect of continued professional development. these informal networks help teachers to deal with the increasing complexity of their work. research shows that most of what professionals learn is learnt informally (cross, 2007), which highlights the need for professional autonomy and personal creativity in problem-solving and professional development. furthermore, research shows the need to understand the role and impact of informal social networks on teacher professional development (villegas-reimers, 2003; darling-hammond et al., 2009; boud & hager, 2012; hargreaves & fullan, 2012). 1.1 the value of networked professional learning networked learning is a perspective on social learning that describes how participants learn through communication, exchange and connections. people in a person's network can be seen as a source of knowledge (siemens, 2004). learning in networks can be informal (a chat during a break) or formal (attending a group training), and the networks themselves can be formal (a taskforce) or informal (talking to a student's parent). learning networks often can be of value when we are in need of certain knowledge, especially the „weak ties‟; those people that we know but don‟t interact with very often can have something „new‟ to offer (granovetter, 1973). learning in networks is nothing new, it happens where people interact and gain experience (eraut, 2004), connected to the work context (billet, 2004). professional learning has proven to drive organisational learning and innovation (bessant et al., 2012). addressing complex problems is a forte of the networked learning paradigm (earl & katz, 2007; hodgson, de laat, mcdonnel & ryberg, 2014). 1.1.1 networked learning and professional development in spite of the proven importance of informal networks, professional development of teachers is almost invariably approached in a largely formal manner (darling-hammond & wei, 2009; villegasreimers, 2003). school organisations often think of schooling in terms of hiring an expert, in-house training, or individual training trajectories such as coaching. however, formal trajectories are seldom tailored to the challenges teachers face in daily practice. furthermore, these challenges at work induce teachers to learn informally (billet, 2004). both formal trajectories and informal learning processes are part of the learning of teachers, and professionals in general (billet, 2002; le clus, 2011). unfortunately, this continuous process of workplace learning, where people customarily exchange knowledge with others in their networks, is hardly ever recognised as professional development. as such, informal learning processes are often overlooked by the management and as a consequence do not receive adequate attention. this suboptimal situation (billet, 2004; de laat, 2012) can be remedied by aligning formal and informal learning processes through networked learning. instead of contrasting formal with informal learning, we emphasise the need to develop a hybrid form of learning where both formal and informal learning activities are recognised and promoted (cf. mcguire & gubbins, 2010). this requires a new role from school management, one that expands a culture of learning by creating social learning spaces for professional development (de laat, 2012). growing evidence is available that shows how informal professional development can be given a place within the formal organisational context by establishing learning networks and professional learning communities, such as „communities of practice‟ (cf. wenger, dermott & snyder, 2002). promoting and strengthening these vaessen et al. 58 | f l r informal networks builds on the already existing social structures and networks within and between schools (cf. hodkinson & hodkinson, 2005; de laat, 2012). 1.1.2 networked learning and professional autonomy participation in learning networks is aimed at sharing knowledge and expertise as individuals personally see fit. networked learning, in our view, is aimed at promoting professional autonomy, selfdirectedness and independent decision-making. networked learning opens up the social environment to optimally make use of (new) possibilities to connect to other professionals (cf. de laat, 2012). for networked learning to be effectively integrated into the organisation, a balanced and integrated approach is required (agterberg, 2012). since informal learning through networks is often bottom-up, self-governing, spontaneous and practice-driven, it is not an easy task to combine this with the formal need for control and performance: management and „personnel‟ have different roles and outlooks (fuller & unwin, 2003; hargreaves & fullan, 2012; hodkinson & hodkinson, 2005) as soon as the management gets involved too much, participants in learning networks risk losing their sense of autonomy, the result of which can be loss of motivation (agterberg, 2012; kubiak & bertram, 2010). related to issues of teacher autonomy are teachers‟ influence on management control and leadership (forrester, 2000) as well as teachers‟ participation in planning and innovation (de laat, 2012). 1.1.3 networked learning and management providing autonomy, which allows individuals to interact and develop expertise as they see fit, means lowering formal control (hulsbos, andersen, kessels & wassink, 2012). this brings into view issues of management and leadership, which directly influence the amount of professional autonomy that individuals have in the organisation (bass, 1991; tynjälä, 2013). with greater individual autonomy, thinking, learning and acting independently is increased and people can personally take up responsibility. this requires a different style of leadership, where responsibilities are shared among the members of the organisation: distributed leadership. distributed leadership promotes the sharing of knowledge and increases motivation for work and learning (spillane, 2008). when leaders pay attention to informal factors in the organisation, such as the personal interests of individuals (i.e. „transformative leadership‟) this increases commitment to organisation goals. this can be contrasted with purely transactional leadership, which functions according to standards, performance and rewards, which can engender mediocrity in the organisation (bass, 1991). to create an organisation where the day-to-day complexity is successfully dealt with and different interests are accounted for, where responsibility is shared and where people can grow and together create value and quality, the management needs to shift focus from a traditional centralised role to a position that reflects a deeper insight into the dynamics of the organisation. this entails an integrated view of formal and informal dynamics. directions and strategies can be developed „top-down‟ as well a „bottom-up‟ (groot and homan, 2012). networked learning then involves the organisation as a whole, management as well as teachers. 1.1.4 aim of this study we have argued the importance of informal networked learning and illustrated how this relates to professional development, autonomy and management of informal and formal learning in organisations. however, to date, these areas of research have not been integrated in the scientific literature. theory in the field of teacher professional development is still much under development (mccormick, 2010). findings from the private sector can advance theory and practice in the public sector (binz-scharf, lazer, & mergel, 2011). in this study we examine underpinning mechanisms, using a networked learning perspective, in order to develop a better conceptual understanding and to examine how this facilitates a better alignment of informal and formal learning in organisations. since professional development of teachers is directly related vaessen et al. 59 | f l r to teaching quality (darling-hammond & wei, 2009; villegas-reimers, 2003), we deem it important that this topic receives the attention it deserves. 1.2 the ‘iceberg’ metaphor as background of our study: formal and informal working and learning the formal side of how things are officially organised, and the informal side of how in everyday life people work, learn, experience and give meaning to their work, are two faces of the organisation. the analogy of an iceberg illustrates this point. the visible tip of the iceberg represents the formal organisation, where planned decisions are made and organisational structures are developed in order to divide the work, create order, and provide stability. under the waterline we find the huge mass of the iceberg, largely invisible, informally structured, yet much larger and often at least as influential as the official organisation structures, consisting of everything that is not formal (de caluwe and vermaak, 2003; de laat, 2012). „formal‟ and „informal‟ aspects of working and learning both are part of professional life and play a role at the level of individuals, groups, and organisations. the worlds „above‟ and „under‟ water mutually influence each other: by interacting in networks people create and influence both the formal and the informal organisation. within both formal and informal networks we find aspects of control, autonomy, performance, development and management. actions and procedures can be planned or spontaneous, visible or invisible, controlled or chaotic, under orders or autonomous. both formal procedures and informal influences are crucial for the organisation and its members (brown, collins & duguid, 1989; snowden, 2005). formal and informal mechanisms play a role for all individuals, groups and organisations, be it „above water‟, or „under water‟. 2. research question in this paper we research the mechanisms underlying networked professional learning in order to increase our understanding of how to optimally align networked learning in the school organisation. our key questions are:  what are the formal and informal mechanisms underlying networked professional learning, related to professional development, autonomy and management?  how can networked learning be positioned in the most optimal way? the term mechanism is used here as: the way in which something functions. we first address how networked learning contributes to professional development. then, because a prerequisite for networked learning is the possibility of spontaneous and autonomous action and decisionmaking, we outline how networked learning is related to professional autonomy. lastly, we explore how networked learning is related to issues of management and leadership. literature in these different research areas has until now not been integrated, and we work toward a framework in order to bring these different areas together. we do this by means of analysing formal and informal mechanisms that play a role in networked learning. 3. methodology the studies presented in this exploratory review were identified in several systematic steps. first, searches on the database of ebscohost were applied. we chose this database as it is a multi-disciplinary meta-database that allows to search for articles that covers studies in education and professional development, management and organisational learning. ebscohost includes, amongst others, the databases vaessen et al. 60 | f l r of academic search elite, business source premier, e-journals, psycinfo, and eric. peer reviewed journal articles and international peer reviewed book chapters published between january 1st 2004 and january 1st 2014 were included in the search. the following keywords were used for a boolean search: „networked learning‟ or 'learning networks' and „professional development‟ and „teachers‟ not „online‟. this search resulted in 74 articles. the aim of the literature research was to recognise formal and informal mechanisms underlying networked learning, related to professional development, professional autonomy and management. for this purpose, the abstract, summary and references of all selected sources were studied first, 26 studies were shortlisted and the articles were read, which resulted in a final selection of 22 sources. the other 52 articles were left out of further analysis because they did not discuss networked professional development of teachers, or had a single focus on online learning tools. the snowball method of checking references in the remaining articles resulted in 22 extra references relevant to our aim. in total 44 studies (see appendix 2) were read in depth and provided the basis of our analysis. the result is an overview of formal and informal mechanisms involved in networked professional learning. this overview is then condensed into a conceptual framework. 4. findings first we discuss our findings of how networked learning is related to professional development. after this, we look at networked learning and professional autonomy. then we consider the relation of networked learning and management. we conclude each section with an overview of formal and informal mechanisms regarding networked learning found in the literature. 4.1 networked learning and professional development professional development comprises formal and informal activities related to intellectual, personal and social domains (de laat, schreurs & nijland, 2013), and can be seen as a “non-linear ongoing process rather than as an outcome of linear, one-off training events” (varga-atkins, o‟brien, burton, campbell & qualter, 2008, p.42). furthermore professional development can be regarded as “a flow of acquired knowledge, as well as participation in a learning community” (pahor, škerlavaj & dimovski, 2008). in networked learning communities, knowledge is constructed and developed, rather than being transferred from one person to the next (schultz, 2011). influence from colleagues can be noted as a contributing factor in order to learn and develop, for example, in changing a style of teaching (supovitz, sirinides & may, 2009). it is argued that theory in the field of professional development still has to be developed, insights gained from networked learning could contribute to how and what teachers learn professionally (cf. appleby & hiller, 2012; mccormick, 2010). exchange between individuals happens through formal and informal networks (carmichael, fox, mccormick, procter & honour, 2006) and the flow of knowledge related to professional development occurs both between organisations and within organisations (jones, 2006; seezink, poell & kirschner, 2010) as well as cross-culturally (ryan, kang, mitchell & gaalen, 2009). professional learning activities can be formal (obtaining a diploma or a degree from an institute), or informal (sharing a drink after a conference day). studies comparing effectiveness of professional development programmes have found that collaborative approaches are more effective than individual ones (varga-atkins et al., 2010), for example when teachers together research and evaluate their own practices (bartlett & burton, 2006). baker‐ doyle and yoon, (2011) also found that while teachers personally gather information, it is within and through social networks that this information comes to life as it is shared, interpreted, developed and sustained. professional development can be seen as an ongoing process of becoming where people grow and learn in connection with each other and events in their professional life (boud & hager, 2012; poell & van der krogt, 2013). schools however, have traditionally been formally designed in a way that teachers work individually. “they have rarely been given time together to plan lessons, share instructional practices, assess students, design curriculums, or contribute to administrative or managerial decisions” (darling-hammond & wei, 2009, p.11). increasing possibilities for communication and exchange across organisational boundaries is therefore vaessen et al. 61 | f l r an important aspect of networked learning initiatives, aiming to bring together people in order to exchange and create knowledge to support each other. for example, questions can be explored, new insights can be discussed, or meeting an expert can provide valuable new information. both formal and informal learning opportunities enable teachers to improve their practice (o‟brien, varga-atkins, burton, campbell & qualter, 2008). making social learning processes part of a learning programme can complement or replace formal education such as seminars in situations where this formal education does not address the learning needed. for example, a project was carried out in a primary school setting where teachers, parents and other parties outside of the school studied problems together (angelides, georgiou & kyriakou, 2008). these learning networks, aimed at developing a social learning approach, were found to facilitate experimentation and reflection. the teachers felt strengthened in their profession when being able to collaborate with the outsiders (school advisors or academics) that came to the school (angelides et al., 2008). learning through networks and partnerships within and between schools sustains contextualised knowledge (baumfield & butterworth, 2005). beckett (2012) describes a situation in which school staff operated in a political context focused on targets and performance levels. the school was situated in a poor area, which required adaptation and dealing with complexities. the school staff felt that the governmentimposed recommendations were not reflecting their immediate concerns, and developed a school network including researchers in order to develop understanding about the relation between poverty and children‟s educational experiences. professional learning networks can function as a „learning incubation centre‟ (attard, 2012). participating in a learning network can promote reflective awareness and development through collaborative analysis, for example when participants note that they “started to dig deeper into their experience” (p.199). when what happens in learning networks is of direct relevance to the participants' needs, this can increase participants‟ motivation to engage in the reflective process that the network entails (attard, 2012). the main findings of this section are: professional learning is an ongoing process, rather than something occasional, which naturally happens in formal and informal social structures. furthermore, networked learning is often situated and most effective when it is directly related to the work practices. promoting collaboration through networks has proven to be effective to enhance the learning process. in table 1 we outline the formal and informal mechanisms regarding networked learning and professional development that we have found in this section. table 1 formal and informal mechanisms in networked learning regarding professional development mechanisms 'informal' 'formal' knowledge is constructed knowledge is transferred invisible visible transcending borders within boundaries continuous event-driven demand-driven supply-oriented voluntary under orders 4.2 networked learning and autonomy if teachers are to improve their skills, they must have the possibility to influence their work and the way they learn (cf. villegas-reimers, 2003). learning networks provide individuals with the opportunity to vaessen et al. 62 | f l r learn about topics they personally find of interest to their practice or personal development. in addition to being able to choose what they want to learn, networks also open up the environment by providing links to others outside of the direct working environment (cf. büchel & raub, 2002). the option to personally choose the areas to explore improves a person‟s performance (akkerman, petter & de laat, 2008) because the opportunity to choose brings a feeling of responsibility which increases personal motivation (vargaatkins et al., 2010). research shows that when teachers have more autonomy they are more committed and share more of their practices (hökkä & eteläpelto, 2013; imants, wubbels & vermunt, 2013). trotman (2009) warns for too much pressure to meet formal performance standards, pointing out that one should be careful to ensure that true learning is happening, where professionals are intrinsically motivated because of their own interest. for reflective processes to take place among colleagues, there must be trust, so that mistakes can be discussed openly and learned from (hargreaves et al., 2013). positive school culture and atmosphere for collaboration are thus important contributors to quality of networked professional development (vargaatkins et al., 2010). hodkinson and hodkinson (2005) refer to the notion of an „expansive‟ rather than a „restrictive‟ learning environment where formal learning is combined with an effective approach to informal and networked learning. through networked learning, possibilities for collaboration and personal initiative can be created (hodkinson & hodkinson, 2005). learning networks can function as open platforms where participants can meet and develop issues of their own interest. however, issues surrounding accountability can come up when learning networks are misunderstood and misused, for example when formal leaders take part, disturbing genuineness and exchange, or when financial interests are involved that create pressure (trotman, 2009). group processes of power, role ambiguity, and lack of direction can create complications. when personal responsibility takes the form of accountability toward control from superiors or school inspection, spontaneous learning processes can be impeded (hargreaves et al., 2013). among members of the group a sense of autonomy is created and sustained and in this sense, autonomy does not mean acting alone as an isolated individual (hargreaves et al., 2013; imants et al., 2013). a flat organisation structure and a culture that fosters democracy and participation, allows for easier contact between people and increases the chance that networked learning occurs. in open organisational environments where people freely can use their networks to connect to each other and learn, it is easier to find and contact the right person to learn from. hierarchy and a centralised culture can hinder possibilities to learn from more experienced people (pahor et al., 2008). trust is an important factor when it comes to developing networked learning communities (day & hadfield, 2004; trotman, 2009). penuel et al. (2009) describe how in a school there were more opportunities to learn from colleagues, because the principal and the teachers themselves encouraged sharing and communication. authority structures were more open, and teachers often used their networks to go outside the school for helpful resources. the school showed a pragmatic attitude towards teachers using these networks and resources, rather than one requiring formal approval from superiors. this led to a high level of trust in relationships and a sense of collective responsibility. more openness, generated by trust and social coherence, can lead to more success in implementing change and development (penuel et al., 2009). promoting open collaboration requires trust in order for members to open up, discuss differences, deal with uncertainty and respect individual differences (attard, 2012). hökkä and eteläpelto (2013), studying autonomy and learning of teachers, note three aspects to consider in order to improve continuous professional learning facilitated by networks: teachers often do not identify with their role as active researchers and developers, barriers between groups can hinder collaboration between groups in different fields, and too strongly adhering to one‟s views can limit collaboration, cultural change and organisational learning. hanraets, hulsebosch and de laat (2011) note that networking skills need to be developed over time in order to make better use of the social environment . employing initiative, valuing others with whom you learn, sharing responsibility and building relations or actively looking for connections are not necessarily skills that people have by nature. new skills have to be developed, by getting used to the new networked way of thinking and working (day & hadfield, 2004). vaessen et al. 63 | f l r concluding, an important aim of promoting networked learning is to provide individuals with more professional autonomy by creating an open environment in which people can connect to others to learn. we have seen that a number of mechanisms that play a role here: freedom of choice, commitment, responsibility, accountability, power, control, trust, communicative openness and willingness to share and reflect are all factors that contribute to the professional autonomy of the individual, and to a collaborative atmosphere in the organisation, and the success of networked learning activities. we stipulate that aiming to integrate these informal tendencies with the necessary formal requirements (see table 2) will create a situation with most value for all involved. table 2 formal and informal mechanisms regarding professional autonomy mechanisms 'informal' 'formal' personal choice rules commitment accountability personal interest/development performance standards personal reflection directives communicative openness communicative barriers trust control in what follows we outline how networks and networked learning are related to management, how networked learning is important, and what can be done to promote it. we identify formal and informal mechanisms that are of influence in the context of management and networked learning. 4.3 networked learning and management schools can be seen as examples of ´open practices´ (de laat, schreurs & nijland, 2014), connecting different parties and practices in an open and complex environment as they are directly related with governments, parents and families, companies and other collaborative institutions (darling-hammond & wei, 2009; villegas-reimers, 2003). the importance of networks for the organisation and the way they are embedded within organisational structures have been widely recognised (cf. carmichael et al., 2006 ). knowledge developed in learning networks form a significant part of the „social capital‟ of an organisation (van emmerik, jawahar, schreurs & cuyper, 2011), and learning networks build capacity for change (edwards, 2012). since networked processes comprise a large part of the learning in organisations, it raises the question of how to manage the relations and knowledge involved. by relinquishing some control, managers can provide a creative and productive network environment where organisation members take part out of their own interest, understanding the benefits of having a strong professional network (büchel & raub, 2002). leaders need to „let it happen‟ while at the same time facilitating adequate room for emerging networks and embedding network activities in the organisation (kubiak & bertram, 2010). leadership is not only embedded in formal positions, but emerges from interactions between people and activities that are performed (scribner, sawyer, watson & myers, 2007). in a more open and decentralised authority structure, leadership is less central but distributed over the members in the networks of the organisation (cf. frost, 2008). büchel and raub note the importance of multi-directionality, each member or unit can learn from all the others. responsibility for success lies within all the network members. learning networks can be designed for problem-solving and creating new knowledge, generated by input from all participants. vaessen et al. 64 | f l r although the motivation of the participants is crucial in attaining success, learning networks need to be supported by the management (büchel & raub, 2002; carmichael et al., 2006). promoting learning and change entails that both formal processes and informal processes are considered important and where possible brought into agreement. when the formal and the informal organisation of a school are in harmony, it increases the chance of successful collaboration (penuel et al. 2010). managing responsibilities and allocation of time and resources have found to be of influence to perceptions of the social space on the work floor. the “designed” and “lived” organisations are equally important and influence each other mutually (penuel et al., 2010). in addition to promoting an open culture of learning and exchange in general, organising network activities or setting up networked learning communities can be helpful to promote the exchange of knowledge (moses, skinner, hicks & o‟sullivan, 2009) and to create a more distributed leadership where members of the organisation all can contribute their expertise (baumfield & butterworth, 2005). holmes (2004) describes a networked learning project where collective enquiry was the underlying mechanism that fuelled the activity in the learning networks. in order to be successful, a learning network needs a common purpose which benefits individual needs, fruitful collaboration which promotes commitment, purposeful and relevant network activities, a good facilitator who has sound knowledge and expertise in the given area, and funding (varga-atkins et al., 2010). fostering networked learning communities is most successful when participants have shared goals, such as clearly defined aims and activities, where a balance between short and long-time goals is important, observe kubiak and bertram (2010). in order to promote learning networks, it has shown to be important to respect the natural bottom-up, self-governing culture of learning. since informal learning is often spontaneous and practice-driven, it is not an easy task to combine this with the need for control and performance of „above the waterline‟: management and employees have different roles and outlooks. as soon as the management gets involved too much, learning networks risk losing their sense of autonomy, the result of which can be loss of motivation (agterberg, 2012; kubiak & bertram, 2010). for a network facilitator, his or her task involves creatively working with whatever emerges and take up the role of for example an inspirer, guide, pr-manager or an investigator. in order to work with bottom-up processes the facilitator has to develop a non-directive attitude, and to investigate profoundly the needs and expectations of the participants and use this information to make suggestions for developing the network. also, coaching participants intensively in personal and communication skills and online literacy can be part of the procedures. furthermore it can be necessary to promote networked learning as a recognised strategy for professional development in order for it to be understood and supported by supervisors and managers (hanraets et al. 2011).. school principals are important agents when it comes to implementing learning networks. they can act as gate-keepers, facilitators or as barriers (o‟brien et al., 2008). the way networks are promoted and developed by leaders and co-leaders is highly influential (daly, moolenaar, bolivar & burke, 2009), while the way networks develop can vary from network to network (kubiak, 2009, kubiak & bertram, 2010). some may be more short-lived, others become more mature and individuals and schools might opt in or out according to their individual needs. network leaders, being aware of these particularities and developing appropriate strategies, can prove vital for the healthy development of learning networks (fox, haddock & smith, 2007; kubiak & bertram, 2010; schechter, 2012; varga-atkins et al., 2010). hökkä and eteläpelto (2013) conclude that because the management is crucial in creating openness and the possibility for change, leaders and managers themselves need to reflect on their own identity, since they are the ones implementing strategic decisions and then deal with the emotions of the personnel. concluding; regarding networked learning and organisational leadership, we found a number of mechanisms at play. managerial acknowledgement of informal networks, promoting networked learning, organisational structure, a distributed leadership, open communication patterns, and an organisational culture in favour of collaboration and exchange, not only between direct colleagues but also between different organisational layers, all contribute to an environment that promotes a healthy learning culture that is conducive to both formal learning procedures and informal networked learning (see table 3). vaessen et al. 65 | f l r table 3 formal and informal mechanisms regarding management mechanisms 'informal' 'formal' recognition of informal networks recognition of formal authority structures shared leadership centralised leadership bottom-up decision-making top-down decision-making open organisational structure rigid organisational structure open communication closed communication learning and working together in an inspiring environment is more likely to succeed when the work floor and the management understand each other and respect each others‟ decisions. networked learning facilitates understanding and collaboration in respect to the content of work practices, and also contributes to the formal and informal organisational context. 5. conclusion and discussion in this study we examined underpinning mechanisms regarding networked learning and professional development, autonomy, and management. we used the perspective of networked learning in order to develop a better conceptual understanding and to examine how this facilitates a better alignment of informal and formal learning in organisations. our key questions were:  what are formal and informal mechanisms underlying networked professional learning related to professional development, autonomy and management?  how can networked learning be positioned in the most optimal way? 5.1 formal and informal mechanisms underlying networked professional learning concerning our first question: we analysed the formal and informal mechanisms that we found in each of the sections of the results (see appendix 1) and found three main groups of mechanisms at play: learning mechanisms: what we have seen in the literature indicates that networked learning is a natural activity through which professionals develop their expertise, in addition to participating in formal learning procedures. this form of professional development is a continuous process. networked learning is often directly related to work practices and promoting it has proven to be effective to enhance the learning process. mechanisms regarding autonomy can be considered to be motivational: networked learning provides individuals with the opportunity to connect to others with the same interests, in this way opening up the learning environment to learn what one deems necessary. personal learning and learning initiatives can be promoted through networked learning. issues of trust, freedom of choice, and willingness to share and connect are intrinsically motivated factors that play a role here. this can be contrasted with pressure to perform, obligations to follow rules, and follow strict regulations which, however necessary, creates an external motivational force (cf. ryan & deci, 2000). vaessen et al. 66 | f l r organisational mechanisms: if management acknowledges the value of informal networks, professionals can be encouraged to make use of their informal networks in order for the organisation to adapt to the always changing environment. through networks, organisational structures become more flexible, and open communication can be promoted. in an expansive rather than a restrictive organisational environment, leadership can be seen as a process where responsibilities are distributed and „bottom-up‟ initiatives are encouraged. the management has an important role in creating a conducive and collaborative learning environment by providing opportunities for networked learning activities and structuring the formal organisation accordingly these three groups of mechanisms can be brought together in the following framework against the background of our „iceberg‟ (figure 1). figure 1. three groups of formal and informal mechanisms related to networked learning in school organisations 5.2 how can networked learning be positioned in the most optimal way? our second key question was: how can networked learning be positioned in the most optimal way? as we have argued in the introduction, formal and informal learning procedures in teacher professional development often not are integrated in a satisfactory way. the core mechanisms depicted in the formal-informal framework illustrate how networked learning can be positioned so that formal learning procedures can be augmented, complemented and informed by informal networked learning. already existing informal networks can be made visible and then strengthened by giving them a place in the organisation. for this to happen it is helpful for the networks to develop a learning agenda that is visible to the management (de laat, 2012), and have support from the management (büchel and raub, 2002). for members of networks to be motivated, autonomy, trust and efficacy are important factors in order for networks to be effective (cf. van den beemt, ketelaar & diepstraten, 2014). networking skills need to be developed by both the participants in learning networks and by the management of school organisations in order for networked learning to be most effective. formal regulations and standards are a professional reality, but school leaders, in addition to judging teachers‟ performance through accountability practices, can strive to create an open organisational culture where responsibilities are shared, encourage participation, and promote looking for new ideas outside of the direct working environment in order to create an environment learning formal informal autonomy organisation planned extrinsic motivation restrictive spontaneous intrinsic motivation permissive vaessen et al. 67 | f l r where formal study and informal learning can both have their place. recognising both parts of the „iceberg‟ by understanding the mechanisms at play is helpful in order to understand how to balance and integrate both positions so that professionalism can prosper. 5.3 discussion research in this area raises questions about how, what, why, and when teachers learn. currently we do not know much about the way the different mechanisms that we found in this study influence each other, which, in our view, merits further investigation. developing a „social awareness‟ of learning processes (boud & hager, 2012) can help to develop new metaphors for professional development (cf. de laat, schreurs & nijland, 2013) and open up new avenues of practice and research. findings from this study can be used to advance the theoretical understanding about the alignment of informal and formal professional development (cf. evers et al., 2011; mcguire & gubbins, 2010) and develop an instrument to engage school leaders and teachers in a constructive dialogue, and collect further data. our study has its limitations. by focusing on the interplay of formal and informal processes, we have provided a far from exhaustive overview of the findings in each of the chosen fields related to the subject. however, combining the insights from different areas of research in order to come to a shared framework there is scientific relevance to our study and our findings can be further conceptualised and validated. we would like to add to this the observation that there might not be one specific „optimal situation‟ for (networked) professional development to be effective; different people have different needs and views. organisations can be seen as a „complex responsive process‟ with many unexpected complexities and local realities, and only one-third of change efforts to improve quality in organisations are considered successful (pieterse, caniëls & homan, 2012). we believe that this is where making use of networks can be helpful: to provide open space for communication and learning, where individual differences can exist and prosper. openness, exchange, trust, and communication are relevant to both school leaders and teachers. promoting openness and development in the light of performance pressure, market-oriented reforms, and centrally imposed standards is no easy task. however, to be in control can sometimes mean, within limits, letting go of control. networks flourish by a healthy balance between formalities and informalities. striking this balance can be achieved by aiming both at facts and figures and at shared values and meaning. keypoints networks and networked learning are increasingly impostant for the work of teachers because of the increased complexity of the work professional development entails formal and informal processes. informal processes, that take up a large proportion of the learning process, are often overlooked and consequently do not receive much attention. networked learning can be helpful to integrate informal processes in the formal school context and align formal and informal learning procedures. creating a balance between the personal interest and performance requirements can provide for a healthy level of professional autonomy and increase motivation for working and learning adopting a perspective of networked learning can have implications for management and leadership. leadership and responsibilities can be shared in order to create a more „open‟ organisation. striking a healthy balance between attention for formal and informal processes means paying attention to both facts and figures and shared values and meaning. vaessen et al. 68 | f l r acknowledgements this paper has been supported by skoem, the netherlands. references agterberg, l. c. m. (2012). walking a tightrope: the dynamics of coordinating intra-organizational networks of practice (doctoral dissertation).vu, amsterdam. akkerman, s., petter, c., & laat, m. de. (2008). organising communities-of-practice: facilitating emergence. journal of workplace learning, 20(6), 383–399. doi: 10.1108/13665620810892067 angelides, p., georgiou, r., & kyriakou, k. (2008). the implementation of a collaborative action research programme for developing inclusive practices: social learning in small internal networks. educational action research, 16(4), 557–568. doi: 10.1080/09650790802445742 appleby, y., & hillier, y. (2012). exploring practice–research networks for critical professional learning. studies in continuing education, 34(1), 31-43. doi: 10.1080/0158037x.2011.613374 attard, k. (2012). public reflection within learning communities: an incessant type of professional development. european journal of teacher education, 35(2), 199-211. doi: 10.1080/02619768.2011.643397 baker‐ doyle, k. j., & yoon, s. a. (2011). in search of practitioner‐ based social capital: a social network analysis tool for understanding and facilitating teacher collaboration in a us‐ based stem professional development program. professional development in education, 37(1), 75–93. doi: 10.1080/19415257.2010.494450 bartlett, s., & burton, d. (2006). practitioner research or descriptions of classroom practice? a discussion of teachers investigating their classrooms. educational action research, 14(3), 395–405 doi:10.1080/09650790600847735 bass, b. m. (1991). from transactional to transformational leadership: learning to share the vision. organizational dynamics, 18(3), 19-31. doi: 10.1016/0090-2616(90)90061-s baumfield, v., & butterworth, m. (2005). developing and sustaining professional dialogue about teaching and learning in schools. journal of in-service education, 31(2), 297-312. doi:10.1080/13674580500200280 beckett, l. (2012). “trust the teachers, mother!”: the leading learning project in leeds. improving schools, 15(1), 10–22. doi: 10.1177/1365480211433723 berry, j. (2012). an investigation into teachers’ professional autonomy in england: implications for policy and practice. submitted in partial fulfilment of the requirements of the degree of phd. university of hertfordshire. bessant, j., alexander, a., tsekouras, g., rush, h., & lamming, r. (2012). developing innovation capability through learning networks. journal of economic geography, 12(5), 1087-1112. doi: 10.1093/jeg/lbs026 billett, s. (2004). workplace participatory practices: conceptualising workplaces as learning environments. journal of workplace learning, 16(6), 312-324. doi: 10.1108/13665620410550295 binz-scharf, m. c., lazer, d., & mergel, i. (2011). searching for answers: networks of practice among public administrators. the american review of public administration, 42(2), 202–225. doi: 10.1177/0275074011398956 boud, d., & hager, p. (2012). re-thinking continuing professional development through changing metaphors and location in professional practices. studies in continuing education, 34(1), 17-30. doi: 10.1080/0158037x.2011.608656 brown, j, collins, a; duguid p (1989). situated cognition and the culture of learning. educational researcher, vol. 18, no. 1. (jan. feb., 1989), pp. 32-42. doi: 10.3102/0013189x018001032 büchel, b., and raub, s. (2002). building knowledge-creating value networks. european management journal, 20(6), 587-596. doi: 10.1016/s0263-2373(02)00110-x carmichael, p., fox, a., mccormick, r., procter, r., & honour, l. (2006). teachers’ networks in and out of school. research papers in education, 21(2), 217–234. doi: 10.1080/02671520600615729 vaessen et al. 69 | f l r clus, le, m. (2011). informal learning in the workplace: a review of the literature. australian journal of adult learning, 51(2) 355-373. cross, j (2007). informal learning, rediscovering the natural pathways that inspire innovation and performance. san francisco: pfeiffer. daly, a. j., moolenaar, n. m., bolivar, j. m., & burke, p. (2010). relationships in reform: the role of teachers’ social networks. journal of educational administration, 48(3), 359–391. doi: 10.1108/09578231011041062 day, c., & hadfield, m. (2004). learning through networks: trust, partnerships and the power of action research. educational action research, 12(4), 575–586. doi: 10.1080/09650790400200269 darling-hammond, l., wei, r. c., andree, a., richardson, n., & orphanos, s. (2009). professional learning in the learning profession. washington, dc: national staff development council. de caluwe, l., & vermaak, h. (2003). learning to change: a guide for organization change agents. thousand oaks: sage. de laat, m. (2012). enabling professional development networks: how connected are you? inaugural address. heerlen, open universteit. de laat, m.f., schreurs, b., & nijland, f. (2014). communities of practice and value creation in networks. in. r.f. poell, t. rocco, & g. roth. (eds.). the routledge companion to human resource development. new york: routledge. earl, l., & katz, s. (2007). leadership in networked learning communities: defining the terrain. school leadership & management, 27(3), 239–258. doi: 10.1080/13632430701379503 edwards, f. (2012). learning communities for curriculum change: key factors in an educational change process in new zealand. professional development in education, 38(1), 25-47. doi: 10.1080/19415257.2011.592077 emmerik, h. van, jawahar, i. m., schreurs, b., & cuyper, n. de. (2011). social capital, team efficacy and team potency: the mediating role of team learning behaviors. career development international, 16(1), 82–99. doi: 10.1108/13620431111107829 eraut, m. (2004). informal learning in the workplace. studies in continuing education, 26(2), 247-273. doi: 10.1080/158037042000225245 evers, a. t., kreijns, k., van der heijden, b. i., & gerrichhauzen, j. t. (2011). an organizational and task perspective model aimed at enhancing teachers’ professional development and occupational expertise. human resource development review, 10(2), 151-179. doi: 10.1177/1534484310397852 forrester, g. (2000). professional autonomy versus managerial control: the experience of teachers in an english primary school. international studies in sociology of education, 10(2), 133–151. doi: 10.1080/09620210000200056 fox, a., haddock, j., & smith, t. (2007). a network biography: reflecting on a journey from birth to maturity of a networked learning community. curriculum journal, 18(3), 287–306. doi: 10.1080/09585170701589918 fuller, a., & unwin, l. (2003). learning as apprentices in the contemporary uk workplace: creating and managing expansive and restrictive participation. journal of education and work, 16(4), 407-426. doi: 10.1080/1363908032000093012 granovetter, m. (1973). the strength of weak ties. american journal of sociology, 78, 1360–1380. groot, n., & homan, t. h. (2012). strategising as a complex responsive leadership process. international journal of learning and change, 6(3/4), 156. doi: 10.1504/ijlc.2012.050858 hargreaves, e., berry, r., lai, y. c., leung, p., scott, d., & stobart, g. (2013). teachers’ experiences of autonomy in continuing professional development: teacher learning communities in london and hong kong. teacher development 17(1), 19-34 doi: 10.1080/13664530.2012.748686 hargreaves, a., & fullan, m. (2012). professional capital: transforming teaching in every school. new york: teachers college press. hodgson, v., de laat, m., mcconnell, d., ryberg, th. (eds.) (2014) the design, experience and practice of networked learning. new york: springer hodkinson, h., and hodkinson, p. (2005). improving schoolteachers' workplace learning. research papers in education, 20(2), 109-131. doi: 10.1080/02671520500077921 vaessen et al. 70 | f l r holmes, d. (2004). nuts, bolts, levers and cranks: designing enquiry-based learning in hartlepool. improving schools, 7(2),121–128.doi: 10.1177/1365480204047344 hökkä, p., & eteläpelto, a. (2013). seeking new perspectives on the development of teacher education: a study of the finnish context. journal of teacher education, 65(1), 39–52. doi: 10.1177/0022487113504220 hulsbos, f., anderson, i., kessels, j., & wassink, h. (2012). professionele ruimte en gespreid leiderschap [professional autonomy and distributed leadership]. look report 37, open university, the netherlands. imants, j., wubbels, t., & vermunt, j. d. (2013). teachers’ enactments of workplace conditions and their beliefs and attitudes toward reform. vocations and learning, 6(3), 323–346. doi: 10.1007/s12186013-9098-0 jones, j. (2006) leadership in small schools: supporting the power of collaboration. management in education 20(2) 24-28. doi: 10.1177/089202060602000207 kubiak, c. (2009). working the interface: brokerage and learning networks. educational management administration & leadership, 37(2), 239–256. doi: 10.1177/1741143208100300 kubiak, c., & bertram, j. (2010). facilitating the development of school-based learning networks. journal of educational administration, 48(1), 31–47. doi: 10.1108/09578231011015403 lieberman, a. (2000). networks as learning communities shaping the future of teacher development. journal of teacher education, 51(3), 221-227. doi: 10.1177/0022487100051003010 mccormick, r. (2010). the state of the nation in cpd: a literature review. curriculum journal, 21(4), 395– 412. doi: 10.1080/09585176.2010.529643 mcguire, d., & gubbins, c. (2010). the slow death of formal learning: a polemic. human resource development review, 9(3), 249-265. doi: 10.1177/1534484310371444 moses, a. s., skinner, d. h., hicks, e., & o‟sullivan, p. s. (2009). developing an educator network: the effect of a teaching scholars program in the health professions on networking and productivity. teaching and learning in medicine, 21(3), 175–9. doi: 10.1080/10401330903014095 o‟brien, m., varga-atkins, t., burton, d., campbell, a., & qualter, a. (2008). how are the perceptions of learning networks shaped among school professionals and headteachers at an early stage in their introduction? international review of education, 54(2), 211–242. doi: 10.1007/s11159-0089084-1 pahor, m., škerlavaj, m., & dimovski, v. (2008). evidence for the network perspective on organizational learning, journal of the american society for information science and technology, 59(12), 1985– 1994. doi: 10.1002/asi penuel, w. r., riel, m., joshi, a., pearlman, l., kim, c. m., & frank, k. a. (2010). the alignment of the informal and formal organizational supports for reform: implications for improving teaching in schools. educational administration quarterly, 46(1), 57–95. doi: 10.1177/1094670509353180 penuel, w., riel, m., krause, a., & frank, k. (2009). analyzing teachers' professional interactions in a school as social capital: a social network approach. the teachers college record, 111(1), 124-163. pieterse, j. h., caniëls, m. c., & homan, t. (2012). professional discourses and resistance to change. journal of organizational change management, 25(6), 798-818. doi: 10.1108/09534811211280573 poell, r. f. & van der krogt f.j. (2013). the role of human resource development in organizational change : professional development strategies of employees, managers and hrd practitioners, 1– 20. chapter to be published in international handbook of research in professional and practicebased learning. s. billett, c. gruber (eds.). dordrecht: springer (in press) price, d. (2013). open: how we’ll work, live and learn in the future. kindle edition. pugh, k. and prusak, l. (2013). designing effective knowledge networks. retrieved october, 24, 2013, from http://sloanreview.mit.edu/article/designing-effective-knowledge-networks/. ryan, j., kang, c., mitchell, i., & erickson, g. (2009). china's basic education reform: an account of an international collaborative research and development project. asia pacific journal of education, 29(4), 427-441. doi: 10.1080/02188790903308902 vaessen et al. 71 | f l r ryan, r. m., & deci, e. l. (2000). self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. american psychologist, 55(1), 68. doi: 10.1037/0003066x.55.1.68 scribner, j. p., sawyer, r. k., watson, s. t., & myers, v. l. (2007). teacher teams and distributed leadership: a study of group discourse and collaboration. educational administration quarterly, 43(1), 67–100. doi: 10.1177/0013161x06293631 schultz, k. (2011). beginning with the particular: reimagining professional development as a feminist practice. the new educator, 7(3), 287–302. doi: 10.1080/1547688x.2011.593997 seezink, a., poell, r., & kirschner, p. (2010). soap in practice: learning outcomes of a cross‐ institutional innovation project conducted by teachers, student teachers, and teacher educators. european journal of teacher education, 33(3), 229–243. doi: 10.1080/02619768.2010.490911 schechter, c. (2012). the professional learning community as perceived by israeli school superintendents, principals and teachers. international review of education, 58(6), 717-734. doi: 10.1007/s11159012-9327-z siemens g., (2004) connectivism: a learning theory for the digital age. retrieved from http://www.elearnspace.org/articles/connectivism.htm snowden, d. (2005). from atomism to networks in social systems. the learning organization, 12(6), 552– 562. doi:10.1108/09696470510626757 spillane, j. p. (2005, june). distributed leadership. in the educational forum (vol. 69, no. 2, pp. 143-150). taylor & francis group. doi: 10.1080/00131720508984678 supovitz, j., sirinides, p., & may, h. (2009). how principals and peers influence teaching and learning. educational administration quarterly, 46(1), 31–56. doi: 10.1177/1094670509353043 trotman, d. (2009). networking for educational change: concepts, impediments and opportunities for primary school professional learning communities. professional development in education, 35(3), 341–356. doi: 10.1080/13674580802596626 tynjälä, p. (2013). toward a 3-p model of workplace learning: a literature review. vocations and learning, 6(1), 11-36. doi: 10.1007/s12186-012-9091-z varga-atkins, t., o‟brien, m., burton, d., campbell, a., & qualter, a. (2009). the importance of interplay between school-based and networked professional development: school professionals’ experiences of inter-school collaborations in learning networks. journal of educational change, 11(3), 241–272. doi: 10.1007/s10833-009-9127-9 van den beemt, a., ketelaar, e., & diepstraten, i. (2014). reciprocity in knowledge networks. paper presented at the 2014 networked learning conference, edinburgh, uk. villegas-reimers, e. (2003). teacher professional development: an international review of the literature. paris: international institute for educational planning. wenger, e., mcdermott, r. and snyder, w. m. (2002). cultivating communities of practice. boston, mass: harvard business school press. http://psycnet.apa.org/doi/10.1037/0003-066x.55.1.68 http://psycnet.apa.org/doi/10.1037/0003-066x.55.1.68 http://dx.doi.org/10.1007%2fs12186-012-9091-z microsoft word bempeni & vamvakoussi_publication.docx           frontline learning research vol. 3 no. 1 (2015) 18 – 35 issn 2295-3159   individual differences in students’ knowing and learning about fractions: evidence from an in-depth qualitative study maria bempenia,*, xenia vamvakoussib a university of ioannina, greece buniversity of ioannina, greece article received 22 november 2014 / revised 2 february 2015 / accepted 11 february 2015 / available online 1 april 2015 abstract we present the results of an in-depth qualitative study that examined ninth graders’ conceptual and procedural knowledge of fractions as well as their approach to mathematics learning, in particular fraction learning. we traced individual differences, even extreme, in the way that students combine the two kinds of knowledge. we also provide preliminary evidence indicating that students with strong conceptual fraction knowledge adopt a deep approach to mathematics learning (associated with the intention to understand), whereas students with poor conceptual fraction knowledge adopt a superficial approach (associated with the intention to reproduce). these findings suggest that students differ in the way they reason and learn about fractions in systematic ways and could be used to inform future quantitative studies. keywords: fractions; conceptual/procedural knowledge; individual differences; learning approach                                                                                                                           * corresponding author. current address: trempessinas 32, athens, 12136, greece. e-mail address: mbempeni@gmail.com doi http://dx.doi.org/10.14786/flr.v3i1.132 m. bempeni & x. vamvakoussi           | f l r   19   1. theoretical background the distinction between procedural and conceptual knowledge has elicited considerable research and discussion among researchers in the fields of cognitive-developmental psychology and mathematics education. procedural knowledge is defined as the ability to execute action sequences to solve problems and is usually tied to specific problem types, whereas conceptual knowledge is defined as knowledge of concepts pertaining to a domain and related principles (rittle-johnson and schneider, in press). the relation between the two types of knowledge, particularly with respect to their order of acquisition has elicited considerable discussion, and there is evidence in favour of contradictory views – in the words of rittle-johnson, siegler, and alibali (2001), “concepts-first” and “procedures-first” theories. according to concepts-first theories, children develop (or are born with) conceptual knowledge in a domain and then use this knowledge to select procedures for solving problems. according to procedures-first theories, children learn procedures for solving problems in a domain and later extract domain concepts from repeated experience in solving problems. in the area of mathematics education research, the two types of knowledge (sometimes referred to by other terms) are deemed practically inseparable (gilmore & papadatou-pastou, 2009; hiebert & wearne, 1996). nevertheless, it is assumed that procedural knowledge plays an important role in the development of conceptual understanding (dubinsky, 1991; gray & tall, 1994; sfard, 1991). more specifically, it is suggested that mathematical concepts develop out of related mathematical processes. such accounts share two common background assumptions, namely that there is a single developmental path and that this path is independent of the particular domain considered. rittle-johnson and siegler (1998) challenged the latter providing evidence that the order of acquisition many vary, depending of the domain considered. in any case, the two types of knowledge appear closely related. thus, rittle-johnson et al. (2001) argued for an iterative model, according to which the two types of knowledge develop in a hand-over-hand process and gains in one type of knowledge lead to improvements in the other. this model is supported by empirical evidence and seems to provide an adequate description of the relation between conceptual and procedural knowledge (rittle-johnson & schneider, in press). nevertheless, there is evidence that sometimes the development of one type of knowledge does not necessarily lead to the development of the other. indeed, in the area of fraction learning it has been shown that some students have the ability to perform fraction procedures without exhibiting comparable conceptual understanding or without being able to explain why they are using these procedures (kerslake, 1986; peck & jencks, 1981). on the other hand, resnick (1982) presented evidence showing that some children may exhibit conceptual understanding of principles underlying subtraction without showing procedural fluency. recently, a different explanation for the contradictory findings has been proposed, namely that not enough attention has been paid to the individual differences in the way that students combine the two types of knowledge (gilmore & bryant, 2008; gilmore & papadatou-pastou, 2009; hallett, bryant, & nunes, 2010; hallett, nunes, bryant and thrope, 2012). hallett and colleagues examined the procedural and conceptual fraction knowledge of students at grade 4 and 5 (2010) as well as at grade 6 and 8 (2012). they identified groups of students who had strong (or weak) procedural as well as conceptual knowledge. however, they also consistently traced two substantial groups of students who demonstrated relative strength with one form of knowledge and weakness with the other, with differences between the two types of knowledge becoming less salient with age. these findings challenge the assumption that all children follow a uniform sequence in gaining the two types of knowledge (see also canobi, reeve, & pattison, 2003). in their attempts to explain how such individual differences arise, some researchers appealed to differences in students’ prior knowledge in the domain in question (schneider, rittle johnson, & star, 2011); differences in students’ cognitive profiles (gilmore & bryant, 2008; hallett et al., 2012) and differences in students’ educational experiences (canobi, 2004; gilmore & bryant, 2008; hallett et al., 2012). however, empirical evidence in support of these assumptions is so far lacking. for example, schneider et al. (2011) found no evidence supporting the hypothesis that the correlation between the two kinds of knowledge might m. bempeni & x. vamvakoussi           | f l r   20   vary with different levels of prior knowledge in the area of equation solving. hallett et al. (2012) investigated whether individual differences in procedural and conceptual knowledge of fractions can be explained by differences in students’ general procedural and conceptual ability (measured by standardized tests); they found no such evidence. in addition, hallett et al. (2012) examined the role of school experience, which they measured as school attendance, that is, they investigated whether attending different schools could explain the individual differences in question; they found no such relation. further research, possibly with different measures, is necessary to clarify the role of the above factors in individual differences in procedural and conceptual knowledge, in particular of fractions. we argue that a factor also worth investigating is the individual student’s learning approach to mathematics. in the literature there is an overarching distinction between the deep approach to learning, associated with the individual’s intention to understand; and the surface approach, associated with the individual’s intention to reproduce. there are several ways of characterizing each learning approach, mainly adapted to tertiary education (entwistle & mccune, 2004). stathopoulou and vosniadou (2007) proposed a model, which was tested with secondary students. they included three categories for each learning approach, namely goals, (study) strategy use, and awareness of understanding. a deep approach to learning involves goals of personal making of meaning, deep study strategy use (e.g., integration of ideas), and high awareness of understanding. a superficial approach involves performance goals, superficial strategy use (e.g., rote learning), and low awareness of understanding. using these categories, stathopoulou and vosniadou showed that students with strong conceptual understanding of science concepts adopted a deep approach to science learning, whereas students with poor conceptual understanding adopted a superficial approach. a similar association might be present in the case of mathematics as well. indeed, a student that follows a deep learning approach to mathematics is more likely to pay attention to the concepts and principles in the domain in question, to be aware of conceptual difficulties, and to invest the effort necessary to overcome them. on the contrary, a student with a superficial approach is more likely to focus on memorizing procedures, especially if procedures are emphasized in instruction, as is often the case (moss, 2005). before we formulate our hypotheses, we turn to a methodological issue, namely the difficulty to measure the two types of knowledge validly and independently of each other (e.g., gilmore & bryant, 2006; hiebert & wearne, 1996; rittle-johnson & schneider, in press; schneider & stern, 2010; silver, 1986). the development of a procedural test that would be conceptual free (and vice versa) is a challenging task, since this type of tests may be person, content and context sensitive (haapasalo & kadijevich, 2000; schneider et al., 2011). moreover, for tasks administered in paper-and-pencil tests, it is often impossible to decide how the student actually solved the task. for such reasons, hiebert and wearne (1996) suggested that attention should be also paid to students’ solution strategies (see also faulkenberry, 2013). a distinction between procedural and conceptual strategies (alsawaie, 2011; clarke & roche, 2009; yang, reys, & reys, 2007) is relevant at this point: procedural strategies are related to rules and exact computation algorithms learnt from instruction. conceptual strategies, on the other hand, are diverse, and tailored to the specific problem at hand; they are mostly invented by (some) students themselves that use them flexibly in order to avoid lengthy computations as well as to deal with unfamiliar problems (see also smith, 1995). in this study, we examined ninth graders’ conceptual and procedural fraction knowledge. taking into account the methodological issue mentioned above, we designed a qualitative study in order to also monitor students’ strategies. similarly to hallett et al. (2010, 2012), we hypothesized that there are individual differences in the way students combine the two kinds of knowledge. we were particularly interested in extreme cases, namely students with strong conceptual knowledge and weak procedural knowledge, and vice versa. such cases are theoretically interesting, since they are not compatible with the iterative model (rittlejohnson et al., 2001). moreover, tracing extreme cases at grade 9 would indicate that individual differences may persist, although the general tendency is for them to become less salient with age (hallett et al., 2012). in addition, we examined students’ learning approach to mathematics learning, particularly fraction learning. following stathopoulou and vosniadou (2007), we explored whether students with strong m. bempeni & x. vamvakoussi           | f l r   21   conceptual knowledge adopt a deep learning approach to mathematics, whereas students with weak conceptual knowledge adopt a superficial approach. 2. methodology 2.1 participants the participants were seven greek students at grade nine (three girls), from seven different schools in the area of athens. the selection of the participants was not random. first, based on their school grades, all participants could be characterized as medium level students in mathematics. second, they all had the same mathematics tutor, starting from the last grades of the elementary school. their tutor provided information about their mathematical behaviour. based on this information, we had reasons to expect some variation in their conceptual and procedural knowledge of fractions. we note that by grade seven greek students are taught all the material related to fractions as well as decimals, and are introduced to the term “rational numbers”. we stress that at the moment this study took place the mathematics curriculum as well as the mathematics textbooks, were “traditional’, in the sense that they emphasized general, computation-intensive procedures for dealing with fraction tasks (smith, 1995). consider, for example, that mental calculations and estimation strategies were not among the curricular goals. based on information provided by our participants’ tutor, who had extensive knowledge about their homework assignments as well as their assessment tests on a long-term basis, we had good reasons to believe that instruction relied heavily on the textbooks, at least with respect to what students were expected to do. 2.2 materials we used thirty fraction tasks grouped in four categories (see appendix a). category a included five procedural tasks, that is, tasks that for which a standard procedure was taught at school: four tasks that examined operations with fractions (q.1.1-q.1.4); and one task that required conversion to an equivalent fraction (q.1.5). category b, consisting of eight tasks, targeted on conceptual knowledge. four tasks involved fraction representations (q.1.6-q.1.9); one task required recognizing fraction as a ratio (q.1.10); one item focused in the role of the unit of reference (q.1.11); and two tasks targeted on the understanding of the effect of multiplication and division with fractions (q.1.12, q.1.13). there were no tasks similar to q.1.10-q.1.13 in the textbooks, either at the elementary, or at the secondary level. on the other hand, the area model for the representation of fractions was salient in the elementary school textbooks, but unlike q.1.8., the shape was typically given, already equally partitioned; examples of improper fraction representations were scarce (q.1.9), and there was no task similar to q.1.7. category c consisted of seven comparison (q.1.14-q.1.17, q.1.20-q.1.22) and two ordering tasks (q.1.18, q-1.19). although these tasks could be solved by standard methods taught at school, they could also be solved by a variety of conceptual strategies. finally, the tasks of the category d required deep conceptual understanding or the combination of conceptual understanding and procedural fluency. more specifically, there were two tasks regarding locating fractions on the number-line (q.1.23, q.1.24); one problem that involved an intensive quantity and required the comparison of ratios (q.1.25); one task regarding estimation of a fraction sum (q.1.26); one task that required substituting variables with non-natural numbers (q.1.27); one task that tested the use of the inverse relationship between addition and subtraction, as well as between multiplication and division with fractions m. bempeni & x. vamvakoussi           | f l r   22   (q.1.28); and two tasks targeting the dense ordering of rational numbers (q.1.29, q.1.30). there were no tasks similar to q.1.27-q.1.30 in the mathematics textbooks. locating fractions on the number line was presented at the secondary level (grade 7), albeit not particularly emphasized. the selection and categorization of the tasks was based on relevant literature (e.g., clarke & roche, 2009; hallett et al., 2010, 2012; mcintosh, reys, & reys, 1992; moss & case, 1999; smith, 1995). we note that we included items targeting students’ awareness of the differences between natural and rational numbers (e.g., q.1.12, q.1.13, q.1.29, q.1.30) which is considered an important aspect of conceptual knowledge (vamvakoussi & vosniadou, 2010; mcmullen, laakkonen, hannula-sormumen, & lehtinen, 2014). we also used a considerable number of tasks related to fraction magnitude (e.g., category d tasks, q.1.23, q.1.24) (for the importance of accessing fraction magnitude in students’ developing knowledge see siegler & pyke, 2013). we stress, however, that this categorization was tentative, since we also looked into students’ strategies. this consideration is particularly important for category c tasks, but relevant for all tasks. in addition, we developed twelve items so as to explore students’ learning approach (deep/superficial) to fraction and, more generally, to mathematics learning (see appendix b). the items were presented as scenarios describing a situation that the student had to react to. 2.3 procedure in the first phase of the study each student was asked to solve the fraction tasks, thinking aloud and explaining their answers. no time limit was imposed. in the second phase three participants were selected to participate in an in-depth, semi-structured individual interview about their learning approach to mathematics. because this was a first attempt to explore a potential relation between individual differences in conceptual and procedural fraction knowledge and the individual’s learning approach, we selected one student with strong procedural, but weak conceptual knowledge; one student with strong conceptual but weak procedural knowledge; and one student who combined both procedural and conceptual knowledge. these students were additionally asked to comment on the responses of the first questionnaire (certainty about the solution, awareness of their performance in the tasks). the second interview took place about one week later. each interview lasted about one hour. all interviews were recorded and transcribed. 2.4 data analysis first, we assessed the accuracy of students’ responses in all tasks. second, we examined the strategies used. we categorized a strategy as procedural, if it was based on instructed rules and procedures related to our research tasks. based on mathematics textbooks, as well as information by our participants’ mathematics tutor, we categorized as procedural strategies the standard algorithms for fraction operations; and transformation strategies (smith, 1995), namely converting to equivalent fractions, similar fractions, decimals, or mixed numbers. transformation strategies are relevant to operations as well as comparison, and they were over-emphasized in the textbooks. we also categorized as procedural the instructed method for q.1.25, namely the construction of a 2x2 table placing the like quantities one below the other, and forming and comparing the ratios. regarding the placement of fractions on the number line, the instructed method involved segmenting the unit in the appropriate number of parts. finally, given the salience of the area model for the representation of fractions, particularly in the elementary grades, we reasoned that it had the status of definition for fractions. we thus did not consider that students used a strategy, either conceptual, or procedural in the related tasks (q.1.6–q.1.9). we categorized as conceptual the strategies that were not based on instructed procedures. for comparison tasks, such strategies involved, for example, the use or reference numbers, such as the unit and one half; and also residual thinking, that is, comparing the complementary fractions (alsawaie, 2011; clarke & roche, 2009; smith, 2005; yang et al., 2007). in a more general fashion, we categorized as conceptual m. bempeni & x. vamvakoussi           | f l r   23   strategies the ones that relied on estimation of fraction magnitudes, on spontaneous use of representations, and on spotting and employing the multiplicative relations present in the task at hand (e.g., in q.1.25). we categorized a strategy as conceptual/procedural if it involved conceptual and procedural features, such as adjusting a procedural strategy to deal with a novel task. a prominent example was the use of a transformation strategy, namely converting to equivalent fractions, as a first step to deal with q.1.30, combined with the idea that this process can be repeated infinitely many times. we also note that in certain cases students provided immediate responses that were not based on a specific strategy; rather, they relied on a holistic understanding of the situation at hand. this was the case mainly for tasks targeting the differences between natural numbers and fractions (q.1.12, q.1.13, q.1.29, q.1.30). for example, some students answered immediately that there is no other number between 2/5 and 3/5, directly drawing on their natural number knowledge. we categorized the strategy of relying on natural number knowledge as conceptual. for the second phase of the study, the categories (i.e., goals, strategy use, and awareness of understanding) and the related indicators used by stathopoulou and vosniadou (2007) were our starting point for the analysis. we reviewed all transcripts and coded them when possible. we selected sentences as unit of analysis, but in some cases we used paragraphs so as to obtain a sense of the whole. we looked for utterances that included keywords pertaining to the indicators of each category (e.g., remember, memorize, memory and similar expressions for the indicator ‘‘rote-learning’’ as a superficial strategy use). we placed the sentences in the coding categories according to the initial indicators and developed new indicators when needed. after coding, data that could not be coded were identified and analyzed to determine if they represented a new category. one new category emerged, namely engagement factors, consisting of two subcategories: preferred tasks/strategies (conceptual/procedural), and also motivation (intellectual challenge/coping). in addition, we replaced the category awareness of understanding with the more general category awareness with indicators pertaining to awareness of understanding (high/low) as well as to awareness of the effectiveness of one’s personal study strategies (high/low). the categories are presented in table 5. 3. results of the 1st phase of the study tables 1-4 present how students performed in the tasks of categories a-d, respectively; and the type of strategy (conceptual, procedural, or a combination of both) they used in each task. as shown in tables 1-4, students 1, 2, and 3 were rather successful across all task categories. students 4, 5, and 6 were successful in categories a and c, but not in categories b and d. student 7 failed in category a, but was rather successful in categories b, c, d. we placed the students in three profiles: a) conceptual-procedural (students 1, 2, and 3); b) procedural (students 4, 5, and 6); and c) conceptual (student 7). in the following we present these profiles in more detail. 3.1 conceptual procedural profile the conceptual-procedural students succeeded in all tasks of category a using procedural strategies, that is, standard algorithms (table 1). student 1 and student 3 (hereafter, kosmas) also succeeded in all tasks of category b (table 2). all three students relied heavily on conceptual strategies (reference numbers, residual thinking) to deal with the tasks of category c (table 3). all three performed well in the tasks of this category, with kosmas responding correctly to all tasks. m. bempeni & x. vamvakoussi           | f l r   24   table 1 students’ performance (success, failure) and type of strategy used (conceptual, procedural, or conceptual-procedural) in the tasks of category a student q.1.1 q.1.2 q.1.3 q.1.4 q.1.5 profile 1 s, p s, p s, p s, p s, p conceptual/procedural 2 s, p s, p s, p s, p s, p 3 (kosmas) s, p s, p s, p s, p s, p 4 s, p s, p s, p s, p f, p procedural 5 s, p s, p s, p s, p s, p 6 (stella) s, p s, p s, p s, p s, p 7 (filio) f, p f, p f, p f, p s, p conceptual kosmas was the only student who responded correctly to all tasks of category 4. in general, however, all three students performed well in category d tasks, showing a rather strong conceptual understanding, combined with procedural fluency. a good indicator of their conceptual understanding is their responses to the density tasks (q.1.29, q.1.30), in particular to the first that is the most challenging. student 2 and kosmas provided an impressively sophisticated answer, stating explicitly that there is no such number and explained that, given any number, no matter how small, one can always find a smaller one. student 1, on the other hand, assumed that such a number exists, thus typically his answer is incorrect; however, he stated that this number cannot be found, not even by a computer; and described it as “zero point zero, followed by infinitely many zeroes, and one unit in the end”. these students’ tendency to prefer conceptual over procedural strategies manifested itself in the tasks of category d as well. none of them applied the instructed method to solve q.1.25; instead, they focused on the relations between the quantities involved. in the words of student 2: “stella’s milk tastes sweeter, because george dissolved the double quantity of chocolate in the triple quantity of milk”. the data presented in tables 1-4 show that kosmas was the only one who succeeded in all tasks. moreover, kosmas’s responses were more elaborated than his peers’ in terms of completeness as well as of the explanations he provided. consider, for example, q.1.26 that asked for the estimation of 7/15 and 5/12. all three students noticed that each addend was smaller than 1/2 and concluded that the sum was smaller than the unit. kosmas, however, went farther to notice that “this sum equals the unit minus 0.5/15+1/12. the missing part is close to 0.1; more precisely, a bit bigger than 0.1”. he reached this close estimate of the missing part mainly via mental calculations, writing down some of the intermediate results. m. bempeni & x. vamvakoussi           | f l r   25   table 2 students’ performance (success, failure) and type of strategy used (conceptual, procedural, or conceptual-procedural) in the tasks of category b student q.1.6 q.1.7 q.1.8 q.1.9 q.1.10 q.1.11 q.1.12 q.1.13 profile 1 s s s s s s s, c s, c conceptual/procedural 2 s s s s s f s, c s, c 3 (kosmas) s s s s s s s, c s, c 4 s f f f f f f, c f, c procedural 5 s f f f f f f, c f, c 6 (stella) f f f f f f f, c f, c 7 (filio) s s s s s s s, c s, c conceptual table 3 students’ performance (success, failure) and type of strategy used (conceptual, procedural, or conceptual-procedural) in the tasks of category c student q.1.14 q.1.15 q.1.16 q.1.17 q.1.18 q.1.19 q.1.20 q.1.21 q.1.22 profile 1 s, c s, c s, c s, c s, c s, c f, c s, c s, c conceptual/ procedural 2 s, c s, c s, c s, c s, c f, c/p s, c s, c s, c 3 (kosmas) ssss) s, c s, c s, c s, c s, c s, c s, c s, c s, c 4 s, p s, p s, p s, p s, p s, p s, p s, p s, p procedural 5 s, p s, p s, p s, p s, p s, p s, p s, p s, p 6 (stella) s, p s, p s, p s, p s, p s, p s, p s, p s, p 7 (filio) s, c s, c s, c s, c s, c s, c s, c s, c s, c conceptual m. bempeni & x. vamvakoussi           | f l r   26   table 4 students’ performance (success, failure) and type of strategy used (conceptual, procedural, or conceptual-procedural) in the tasks of category d student q.1.23 q.1.24 q.1.25 q.1.26 q.1.27 q.1.28 q.1.29 q.1.30 profile 1 f, c f, c s, c/p f, c s, c s, c/p f, c s, c/p conceptual/procedural 2 s, c/p s, c/p s, c s, c s, c f, c s, c s, c/p 3 (kosmas) s, c/p s, c/p s, c/p s, c s, c/p s, c/p s, c s, c/p 4 f, p f, p f, p f, c f, p f, p f, c f, c procedural 5 f, p f, p f, p f, c f, p f, p f, c f, c 6 (stella) f, p f, p f, c f, c f, p f, p f, c f, c 7 (filio) s, c s, c s, c/p s, c s, c f, c/p f, c s, c conceptual 3.2 procedural profile as shown in table 1, the students of this profile performed very well in the tasks of category a (table 1). on the contrary, their performance was very law in the tasks of category b (table 2). in particular, student 3 (hereafter, stella) failed in all the tasks of this category. she stated that “the nominator shows how many pieces to take” to justify her answer in q.1.6, and she drew a circle and partitioned it in three unequal parts in q.1.8 (figure 1). none of these students exhibited any understanding of the fundamental principle that the fractional parts of the unit should be equal, as also evidenced by their performance in q.1.7 (table 2). figure 1. stella’s response to q.1.6, q.1.8: representations for the fractions 1/4 and 2/3, respectively. all three students failed to represent the improper fraction 5/3 (q.1.9). figure 2 presents s5 and stella’s attempts to deal with this task. s4 gave no answer to the problem. figure 2. procedural profile: student 5 and stella’s’ attempt to represent the fraction 5/3. m. bempeni & x. vamvakoussi           | f l r   27   in addition, all three students failed in q.1.10, explaining that the denominator shows how many pieces the pizza had, and the nominator how many pieces were eaten. they also failed in q.1.11, since they did not consider that the units of reference might be different. moreover, they all insisted on executing the calculations in q.1.12 and q.1.13. when they were explicitly instructed not to do it, they came up with the rule “multiplication makes bigger, whereas division makes smaller”. all students of this profile were flawless in the tasks of category c, using only procedural strategies. they were, however, very reluctant to try without using paper and pencil, when they were asked to. in case they tried, their responses reflected severe lack of understanding. for example, stella claimed that 123/220 is greater than 6/5 because the numbers 123 and 220 are greater than 6 and 5, respectively. the students of this profile failed in all tasks of category d (table 4). again, they relied heavily on procedural strategies, in particular transformation strategies. for example, they all converted fractions into decimals in q.1.23 and q.1.24. they also attempted to use this strategy or to perform the calculation in the estimation task q.1.26, although they were specifically asked not to. stella, in particular, explicitly stated that it is impossible to solve the task without converting to similar fractions or to decimals first. students 4 and 5 applied the instructed method q.1.25. however, they were not able to interpret the result. consider, for example, the answer and the explanation provided by student 5: “george’ s milk tastes sweeter, because his proportion 600/100=6 is better than stella’s 200/50=4”. on the other hand, stella’s answer indicated that she neglected the multiplicative relations defining the relative quantities that are involved in the situation: “the girl’s quantities are rather small compared to the boy’s. so i believe that her milk tasted sweeter”. these students’ responses to the tasks on dense ordering (q.1.29, q.1.30) were immediate and reflected the idea that fractions (or decimals, in case they had converted them) are discrete, like the natural numbers. stella stated that “there are no other numbers between 2/5 and 3/5, because 3 comes right after 2”. according to stella, one was the smallest positive number, while students 4 and 5 proposed 0.1. 3.3 conceptual profile as mentioned above, there was only one student placed in this profile, namely filio. as shown in table 1, filio failed in all tasks of category a, except for q.1.5, since she was quite competent with equivalent fractions (see also her solution in q1.25 below). on the contrary, she succeeded in all tasks of category b (table 2). she was able to explain adequately her responses. for example, to explain her disagreement with maria in q.1.10, filio said that “i don’t know how many pieces this pizza had. kostas could have eaten 3 pieces, only if the pizza was cut in four”. similarly, in q.1.11, she exclaimed: “where are the pizzas? i need to see the pizzas. are they the same or not?” while dealing with q.1.12 and q.1.13, she explicitly stated that the outcome is not necessarily bigger, just because there is multiplication involved. she tried with several numbers, and eventually came up with a generalization: “when we multiply a number a by a fraction smaller than the unit, the product is smaller than the number a”. filio succeeded in all tasks of category c (table 3) using consistently only conceptual strategies. interestingly, she also succeeded in most of the tasks of category d (table 4). her responses in q.1.23, q.1.24, were based on estimation of the fraction magnitudes and a rough approximation of their location on the numbers line. unlike the students of the conceptual-procedural profile, she didn’t attempt to find the exact locations by partitioning the line segments. quite similar to these students, however, she focused on the relations between the quantities in q.1.25, employed a transformation strategy, and came up with a solution that is not taught at school: “the 50gr of chocolate powder that stella put in 200gr milk is half the quantity that george put in 600gr. so i double the quantities 50/200 and i get 100/400. then, 100 in 400 means more chocolate powder in the milk than 100 in 600! so, stella’s milk tastes sweeter.” similarly to kosmas, filio explicitly stated that there are infinitely many pairs whose product is 3 (q.1.27). moreover, she also stated that there are infinitely many numbers between 2/5 and 3/5 (q.1.30). m. bempeni & x. vamvakoussi           | f l r   28   unlike all other participants, she justified her answer using spontaneously a rather sophisticated representation: “if we locate them on the number-line, there is definitely a gap in between. in this gap, there are infinitely many numbers”. we note that filio explicitly expressed her discomfort with tasks in which she could not avoid using procedures (e.g., category a tasks, q.1.28). we also note that filio was monitoring her performance during the solution process. she explicitly expressed doubt about responses that were actually incorrect; she also revised certain answers herself. for example, when solving q.1.18, she initially answered that the fractions 3/4 and 6/7 are equal, because for both one fractional unit is needed to complete the unit. she revised this answer after locating the two fractions on the number line. 3.4 conclusions the first phase of the study revealed three different student profiles: the conceptual-procedural profile consisted of three students with quite strong conceptual knowledge of fractions, combined with procedural fluency. these students appeared to prefer conceptual strategies over procedural strategies, when this was possible. one of these students, namely kosmas (student 3), was exceptionally strong: not only did he succeed in all tasks, but he also gave the most complete and elaborated answers. the procedural profile consisted of three students who were capable of applying instructed procedures. this capability allowed them to deal very successfully with the tasks that could actually be solved by an instructed procedure. however, these students failed in most tasks that required conceptual knowledge, exhibiting lack of understanding for even the most fundamental fraction ideas. stella, in particular, failed in the simplest conceptual tasks. these students relied heavily on procedural strategies and avoided consistently to try otherwise. when they did try, they typically failed. finally, the conceptual profile consisted of one student, namely filio (student 7). filio consistently avoided applying procedures throughout the interview, and she failed when she had to do it. she nevertheless exhibited a firm understanding of fundamental fraction ideas; and thus she managed to deal quite successfully with many tasks by applying consistently conceptual strategies. thus, in line with recent discussions regarding the relation between conceptual and procedural knowledge of fractions (e.g., hallett et al., 2010, 2012), we found individual differences in the way that students combine the two kinds of knowledge. moreover, we showed that these differences can be extreme – consider, for example, stella and filio. 4. results of the 2nd phase of the study table 5 presents the categories that describe the deep learning approach and the superficial learning approach to mathematics, and their indicators. in the following we present excerpts from transcribed interviews of kosmas, filio, and stella, in order to highlight the similarities and the differences in their learning approaches to mathematics, along these categories. 4.1 goals kosmas and filio repeatedly referred to the importance of learning with understanding in mathematics, which they both juxtaposed with rote learning. for them, learning with understanding meant personal making of meaning. this point is illustrated in the following excerpts, in which kosmas and filio explain how they would help a hypothetical younger student that is challenged by the comparison of fractions: m. bempeni & x. vamvakoussi           | f l r   29   “perhaps i could try to explain fractions as i understand them. he has to find a personal way of thinking though. he could study the rules. in fact, there are two ways: in the case of fractions, the first one is to memorize the rules and apply them. for example, between two fractions with the same numerator 3/5 and 3/7 the bigger is this one with the smaller denominator. alternatively, he would compare the two fractions to the unit, that is, notice that 3/7 is closer to 1 than 3/5. there is a difference: in the second case you have understood exactly what happens with fractions-the first is rote learning. you can reach a conclusion regarding which of two fractions is bigger but you don’t understand why. personally, if i saw these two fractions, i would compare the fractions to the unit so as to check the validity of the rule.” (kosmas, q.2.11) “i would help him understand the concept of fraction. but, you know, everyone has their own way of thinking. mathematics is not rote learning, you have to put your mind to the work. […] i could explain to him how to compare fractions based on the rules, but if he wants be really able to compare fractions, i think that he should understand the concept of fraction. he must understand what fractions are and then he will do well in fractions.” (filio, q.2.11) consider also the following excerpts: “the most important thing is to understand. knowing the rules will also help you, there is no doubt about it. but understanding is the most important thing.” (kosmas, q.2.11) “if i understand the meaning of what i do, then i can solve the exercises.” (filio, q.2.2) table 5 deep vs. superficial learning approach to mathematics: categories and indicators categories sub-categories indicators deep approach superficial approach goals understanding / personal making of meaning focus on what is required /assessed at school study strategies combining theory and practice systematic, long-term time investment memorizing and rehearsing more rehearsing awareness of understanding high low effectiveness of own study strategies high low engagement factors task/strategy preferences conceptual procedural motivation intellectual challenge coping on the contrary, stella repeatedly referred, explicitly or implicitly, to the importance of complying with what is assessed at school and appeared to focus exclusively on the material taught at school. this is summarized nicely in the following excerpt: “what i would advise a younger student is to look at the exercises solved at school, to focus on what is likely to be asked in the exams, and to pay attention to what the teacher has emphasized on.” (stella, q.2.1) m. bempeni & x. vamvakoussi           | f l r   30   4.2 study strategies kosmas and filio both stressed that in order to study efficiently in mathematics one needs to combine studying theory in depth and extensive practice with exercises. they also expressed their conviction that solving unfamiliar problems is important as a study strategy as well as an indicator of understanding. “you have to know the theory very well so as to understand mathematics. if you only solve exercises, your competence is very limited. one has to understand the theory in depth before trying to solve exercises.” (kosmas, q.2.2) “if you give me any problem and i can solve it, then it means i have understood well.” (kosmas, q.2.9) “one should understand the theory very well and practice a lot as well; and solve exercises beyond the ones in the textbook.” (filio, q.2.2) in contrast, stella’s study strategies were limited to memorizing and rehearsing: “studying what is needed for solving the exercises is pretty much sufficient.” (stella, q.2.2) “studying the theory is good, because you have to know some theory to be able to solve the exercises. but i think that it is better to focus on exercises. personally, i look at what we have done at school, so as to remember how the exercises are solved. i solve them again and again, and then i check if they are correct.” (stella, q.2.3) in addition, unlike stella, kosmas and filio appeared to value the hypothetical students’ study strategies in q.2.3, although they both admitted that they don’t study like this. “there is no doubt that this is the appropriate way of studying the theory. […] this is how i should study but, unfortunately, i don’t. that’s why i am not strong in mathematics.” (kosmas, q.2.3) “what she does is just fine. i don’t study like this, but i wish i did.” (filio, q.2.3) moreover, kosmas and filio referred to the importance of investing time on mathematics studying. they distinguished between merely spending time on studying, and studying systematically and in depth. “mathematics is a course that has to do a lot with understanding, so you have to study a lot. you have to start systematically in mathematics from the beginning. gaps are difficult to cover, one needs to dedicate lots of time for both theory and exercises.” (kosmas, q.2.2) “i was preparing for a mathematics test and i spent lots of time, but only during the last two days before the exam. i believe that studying in depth results to success. if you study superficially, you are not prepared appropriately. when we talk about mathematics, you can’t prepare at the last minute. if you do it, you will fail. it is impossible to learn mathematics two days before the exams.” (kosmas, q.2.4) “it’s not only the time spent on studying, it’s also the way you study. […] you may feel well-prepared for a test because you have spent lots of time on solving exercises and fail in the end. for example, what has happened to me is to face unfamiliar problems in a test and fail. in that test, our teacher tested whether we can think for ourselves, so he examined us in different tasks than the ones we had solved in the classroom. […] in order to succeed, you must have understood the concepts and have practiced a lot.” (filio, q.2.4) stella also mentioned time as an important factor of success in mathematics. for stella, however, spending more time on studying meant more rehearsing: m. bempeni & x. vamvakoussi           | f l r   31   “[one of my classmates] is a very good at math. i believe that i am good too, but not exactly at the same level. […] i think he spends more hours studying than i do. […] perhaps he solves the exercises more times than i do.” (stella, q.2.5) 4.3 awareness 4.3.1 awareness of understanding kosmas felt confident that he was able to assess his performance in mathematical tasks in general. in fact, he was very accurate in assessing his performance regarding the fraction tasks. as already mentioned, filio was monitoring her performance in the fraction tasks and corrected several mistakes herself in the process. she also detected practically all the tasks that she had answered incorrectly. in addition, she was aware that she lacked procedural fluency: “i don’t remember rules and procedures regarding fractions. however, if someone reminded me of them, i could apply them.” (filio, q.2.9) filio acknowledged that fractions require “a lot of thinking” and recalled that she was challenged by fractions at the elementary school. interestingly, she mentioned that she managed to grasp the meaning of fractions, by connecting the “school fractions” with the fractions she met at her music courses. (filio, q.2.6) stella, on the other hand, was confident that she had answered pretty much all fraction tasks correctly. she appeared to detect her mistake in q.1.9, and she revised her answer. however, her second attempt was again incorrect, since it was based on the assumption that 5/3 is “a bit bigger than 0.5”. nevertheless, stella believed that she had a firm understanding of fractions in general: “i believe that i understand everything about fractions. i never had any difficulty with fractions. i found them very easy at the elementary school, too. in general, i have never had any problems with mathematics, as far as i can remember.” (stella, q.2.8) 4.3.2 awareness of the effectiveness of own study strategies as mentioned before, kosmas and filio both admitted that they did not follow effective study strategies in mathematics, although they recognized and appreciated them. in addition, they both attributed the fact that they didn’t excel in mathematics to their own way of studying. “[one my classmates] is really strong in mathematics. i am at a considerably lower level. this is because i don’t invest enough time to study seriously in mathematics. [...] often i only solve the exercises that i have as homework and stop there. [...] i could be as strong as my fellow student, provided that i would be determined to study seriously (kosmas, q.2.5) “i could be as good as him [my fellow student]. how? the old-fashioned way: putting time and effort in studying as i should.” (filio, q.2.5) on the contrary, not once did stella question her study strategies: “every time something went wrong, this happened because i was not so careful. […] or i thought i knew the material and that there was no need to look at it again, but in fact i did not remember it well. but in cases that i had studied as i should, i believe that stress was responsible for my failure.” (stella, q.2.4) m. bempeni & x. vamvakoussi           | f l r   32   4.4 engagement factors 4.4.1 task/strategy preferences as already mentioned, during the first phase of the study it was more than obvious that filio resented the tasks that she perceived as procedural. for instance, she grew impatient with q.1.28 and quitted trying, exclaiming “i’ve had enough! i spent too much time on this already. i can’t do it, i won’t do it!”. kosmas, on the other hand, never expressed any discomfort when he had or chose to apply procedural strategies. in spite of this important difference, these students both expressed their preference for conceptual over procedural tasks, when they were explicitly asked to chose: “this is an easy choice! i would choose the second one, because i do not like using methods. i do know, however, that the first one is easier. at any moment you can open your book and remember how it is solved.” (kosmas, q.2.10) “not the first one, for sure. it’s better to think something new, instead of constantly doing the same. i find no meaning in the application or rules and procedures. it is not interesting. it is like rote learning, you know, a method to solve exercises.” (filio, q.2.10) unlike kosmas and filio, stella showed a clear preference for procedural strategies during the first phase of the studies; and she explicitly stated that she would prefer the standard, procedural task in q.2.10. 4.4.2 motivation as it may be evident by their responses to q.2.10, kosmas and filio were motivated by novel and challenging tasks. there were clear such indications about kosmas already in the first phase of the study. for instance, when he first saw q.1.29, his immediate reaction was the following: “the smallest positive number! this is a nice question, isn’t it?” in fact, kosmas was the only participant who chose to deal with the most demanding and unfamiliar tasks first. when asked why, he replied: “i like challenging tasks much more. i find no interest in solving exercises similar to the ones i have met before. the point is to think of something new.” similarly, filio explained her choice of the unfamiliar task in q.2.10 as follows: “when you try to solve an exercise and you finally discover that something that you thought for yourself is correct, you get a very nice feeling.” (filio, q.2.10). on the contrary, stella’s main concern was to stay on the safe side. as may be evident by her responses presented above, she was mainly interested in good school performance. when she explained why she would prefer the “standard”, procedural, task in q.2.10, she indicated that she was minding the possible failure that guided her choice: “i would choose the first one because it involves operations, which i already know. so i would be sure that i can respond correctly. the second one may involve something i don’t know or never met before.” finally, we note that for kosmas and filio learning with understanding, besides being an important goal in mathematics learning, also had a motivational aspect. consider, for example, the following excerpts: “if you are to study mathematics, you should understand what you’re doing. you should find meaning in what you do.” (filio, q.2.2) “[my classmate who excels in mathematics] has a special interest in math, he loves it. he finds meaning in what he does. that’s why he dedicates so many hours to studying.” (filio, q.2.5) m. bempeni & x. vamvakoussi           | f l r   33   both students mentioned that they felt they understood mathematics at the elementary level, but not so at the secondary level. this was due to the fact that procedures are over-emphasized at the secondary level and this appeared to be demotivating for them. “instruction on fractions is based on algorithms and students do not understand the concept of fraction. for example, in the addition of fractions we learn a priori that fractions must have the same denominator without understanding why. something similar happens to mathematics teaching in general. we should understand mathematics deeper and i think that teachers must help us. how? i don’t know.” (kosmas, q.2.7). 4.5 conclusions as evidenced by their interview excerpts, kosmas and filio exhibited similar features along the categories goals, (study) strategy use, awareness, and engagement factors. specifically, they both appeared to value understanding and personal making of meaning in mathematics learning; they were convinced that the study of mathematics requires combining deep understanding of theory as well as extensive practice; systematic and long-term time investment was a key issue for them, as they appeared aware that merely spending time on mathematics studying is not enough to succeed in mathematics. kosmas and filio showed high awareness of understanding in the domain of fractions; they were also highly aware of their limitations as students in mathematics. finally, they showed a clear preference for tasks that require conceptual understanding and present an intellectual challenge, which appeared to be motivating for them. on the contrary, filio differed across all categories. specifically, filio’s goal was to cope successfully with what was required at school; her study strategies were limited to memorizing rules and procedures as well as solving similar or even the same exercises repeatedly; she preferred procedural tasks because she was confident that she would succeed. finally, she showed practically no awareness of her (extremely limited) conceptual understanding of fractions, and no awareness of the limitations of her study strategies. 5. discussion our results support the hypothesis that there are individual differences in the way that students develop conceptual and procedural knowledge of fractions. similarly to hallett et al. (2010, 2012), we identified students who were strong with respect to one type of knowledge, but weak with respect to the other. although the findings of hallet et al. (2012) indicate that such individual differences become less salient with age, we showed that for some students they remain extreme, even at grade 9. consider, for example, stella and filio: it appears that for these students conceptual and procedural knowledge of fractions have not developed in a hand-over-hand process, as predicted by the iterative model (rittle-johnson et al., 2001). in addition, our study provides preliminary evidence indicating that the individual student’s learning approach to mathematics is worth investigating in relation to individual differences in conceptual and procedural knowledge. similarly to stathopoulou and vosniadou (2007), we found that kosmas and filio, who exhibited strong conceptual knowledge of fractions, both valued a deep approach to mathematics learning; whereas stella, who exhibited poor conceptual knowledge of fractions, appeared to follow a superficial approach. this finding cannot, of course, be generalized, given that it comes from a qualitative study, with small sample. moreover, it is based on “extreme” cases of individuals. nevertheless, this qualitative evidence can inform the hypotheses and the design of future quantitative studies. investigating individual differences in conceptual and procedural knowledge is important for understanding mathematical development (canobi, 2004; hallett et al., 2010, 2012). from an educational perspective, however, encouraging the symmetrical development of the two kinds of knowledge is an m. bempeni & x. vamvakoussi           | f l r   34   important goal, since they are both considered essential for students’ mathematical competence (rittlejohnson & schneider, in press). to this end, probably the first step would be to foster learning environments in which both conceptual and procedural knowledge are valued – and also assessed. keypoints there are individual differences, even extreme, in the way students combine conceptual and procedural knowledge of fractions. the individual student’s learning approach to mathematics is a factor worth investigating with respect to individual differences in conceptual and procedural fraction knowledge. references alsawaie o. (2011). number sense-based strategies used by high-achieving sixth grade students who experienced reform textbooks. international journal of science and mathematics education, 10(5), 1071-1097. doi: 10.1007/s10763-011-9315-y canobi, k. h. (2004). individual differences in children's addition and subtraction knowledge. cognitive development, 19, 81-93. doi: 10.1016/j.cogdev.2003.10.001 canobi, k. h., reeve, r. a., & pattison, p. e. (2003). patterns of knowledge in children’ s addition. developmental psychology, 39, 521-534. doi: 10.1037/0012-1649.39.3.521 clarke, d. m., & roche α. (2009). students’ fraction comparison strategies as a window into robust understanding and possible pointers for instruction. educational studies in mathematics, 72(1), 127138. doi: 10.1007/s10649-009-9198-9 dubinsky, e. (1991). reflective abstraction in advanced mathematical thinking. in d. o. tall (ed.) advanced mathematical thinking (pp. 95-123). kluwer: dordrecht, doi: 10.1007/0-306-47203-1_7 entwisle, n. & mccune v. (2004). the conceptual bases of study strategy inventories. educational psychology review, 16, 325-345. doi: 10.1007/s10648-004-0003-0 faulkenberry, t. j. (2013). the conceptual/procedural distinction belongs to strategies, not tasks: a comment on gabriel et al.(2013). frontiers in psychology, 4, 1-2. doi: 10.3389/fpsyg.2013.00820 gilmore, c. k., & bryant, p. (2006). individual differences in children’s understanding of inversion and arithmetical skill. british journal of educational psychology, 76, 309–331. doi: 10.1348/000709905x39125 gilmore, c. k., & bryant, p. (2008). can children construct inverse relations in arithmetic? evidence for individual differences in the development of conceptual understanding and computational skill. british journal of developmental psychology, 26, 301–316. doi: 10.1348/026151007x236007 gilmore, c. k., & papadatou-pastou, m. (2009). patterns of individual differences in conceptual understanding and arithmetical skill: a meta-analysis. mathematical thinking and learning, 11, 25-40. doi: 10.1080/10986060802583923 gray, e., tall. d. (1994). duality, ambiguity, and flexibility: a proceptual view of simple arithmetic. journal for research in mathematics education 25(2), 407-428. hallett, d., nunes, t., & bryant, p. (2010). individual differences in conceptual and procedural knowledge when learning fractions. journal of educational psychology, 102, 395–406. doi: 10.1037/a0017486 haapasalo, l., & kadijevich, d. (2000). two types of mathematical knowledge and their relation. jmd - journal for mathematic-didaktik, 21, 139-157. doi: 10.1007/bf03338914 hallett, d., nunes, t., bryant, p., & thorpe, c. m. (2012). individual differences in conceptual and procedural fraction understanding: the role of abilities and school experience. journal of experimental child psychology, 113, 469-486. doi: 10.1016/j.jecp.2012.07.009 m. bempeni & x. vamvakoussi 35 | f l r     hiebert, j., & wearne, d. (1996). instruction, understanding, and skill in multidigit addition and subtraction. cognition & instruction, 14(3), 251-283. doi: 10.1207/s1532690xci1403_1 kerslake, d. (1986). fractions: children’s strategies and errors: a report of the strategies and errors in secondary mathematics project. windsor, uk: nfer–nelson. mcintosh, a., reys, b. j., & reys, r. e. (1992). a proposed framework for examining basic number sense. for the learning of mathematics, 12, 2-8. mcmullen, j., laakkonen, e., hannula-sormumen m., & lehtinen e. (in press). modeling developmental trajectories of rational number. learning and instruction. doi:10.1016/j.learninstruc.2013.12.004. moss, j., & case, r. (1999). developing children’s understanding of the rational numbers: a new model and an experimental curriculum. journal for research in mathematics education, 30, 122-147. national council of teachers of mathematics. (1989). curriculum and evaluation standards for school mathematics. reston, va: author. peck, d. m., & jencks, s. m. (1981). conceptual issues in the teaching and learning of fractions. journal for research in mathematics education, 12(5), 339-348. resnick, l. b. (1982). syntax and semantics in learning to subtract. in t. p. carpenter, j. m. moser & t. a. romburg (eds.), addition and subtraction: a cognitive perspective (pp. 136–155). hillsdale, nj: erlbaum. rittle-johnson, b., siegler, r. s., & alibali, m. w. (2001). developing conceptual understanding and procedural skill in mathematics: an iterative process. journal of educational psychology, 93, 346-362. doi: 10.1037/0022-0663.93.2.346 rittle-johnson, b., & siegler, r. s. (1998). the relations between conceptual and procedural knowledge in learning mathematics: a review. in c. donlan (ed.), the development of mathematical skills (pp. 75110). east sussex, uk: psychology press. rittle-johnson, & b., schneider, m. (in press). developing conceptual and procedural knowledge of mathematics. in r. kadosh & a. dowker (eds), oxford handbook of numerical cognition. oxford press. schneider, m., & stern, e. (2010). the developmental relations between conceptual and procedural knowledge: a multimethod approach, developmental psychology, 46, 178-192. doi: 10.1037/a0016701 schneider. m., rittle-johnson b, & star j. (2011). relations among conceptual knowledge, procedural knowledge, and procedural flexibility in two samples differing in prior knowledge. journal of developmental psychology, 47, 1525-1538. doi: 10.1037/a0024997 siegler, r, pyke, a. (2013). developmental and individual differences in understanding of fractions. developmental psychology, 49, 1994–2004. doi: 10.1037/a0031200 silver, e. a. (1986). using conceptual and procedural knowledge: a focus on relationships. in j. hiebert and p. lefevre (eds.), conceptual and procedural knowledge: the case of mathematics. new jersey: erlbaum associates. smith, j. (1995). competent reasoning with rational numbers. cognition and instruction, 13(1), 3-50. doi: 10.1207/s1532690xci1301_1 sfard, a. (1991). on the dual nature of mathematical conceptions: reflections on processes and objects as different sides of the same coin. educational studies in mathematics, 22, 1-36. doi: 10.1007/bf00302715 stathopoulou, c., & vosniadou, s. (2007). conceptual change in physics and physics-related epistemological beliefs: a relationship under scrutiny. in s. vosniadou, a. baltas, & x. vamvakoussi (eds.), reframing the conceptual change approach in learning and instruction (pp. 145-164). oxford, uk: elsevier. doi: 10.1016/j.cedpsych.2005.12.002 vamvakoussi x. & vosniadou, s. (2010). how many decimals are there between two fractions? aspects of secondary school students’ reasoning about rational numbers and their notation. cognition & instruction, 28(2), 181-209. doi: 10.1080/07370001003676603 yang. d., c., reys, r., & reys, b. (2007). number sense strategies used by pre-service teachers in taiwan. international journal of science and mathematics education, 7(2), 383-403. doi: 10.1007/s10763-0079124-5 microsoft word geeraerts et al_publication.docx           frontline  learning  research  vol.5  no.  2  (2017)  78  -­‐  98   issn  2295-­‐3159         corresponding author: kendra geeraerts, gratiekapelstraat 10, 2000 antwerp, belgium email: kendra.geeraerts@uantwerpen.be doi: http://dx.doi.org/10.14786/flr.v5i2.293   intergenerational professional relationships in elementary school teams: a social network approach kendra geeraertsa, piet van den bosschea,b, jan vanhoofa, nienke moolenaarc auniversity of antwerp, belgium buniversity of maastricht, the netherlands cuniversity of utrecht, the netherlands article received 20 february / revised 22 august / accepted 22 august / available online 8 september abstract this paper examines the extent to which school team members’ professional relationships are affected by being part of a certain generational cohort. these professional relationships provide opportunities for intergenerational knowledge flows and can therefore be relevant for intergenerational learning. nowadays these topics have gained more attention due to worldwide demographic changes such as increased retirement rates and high levels of teacher dropout. data were gathered through a survey with socio-metric questions among 299 school team members in 15 elementary schools in the netherlands. using social network analysis, in particular p2 modelling, we analysed the effect of being part of a generational cohort on teachers’ likelihood of having professional relationships in networks such as discussing work, asking and providing advice, and collaboration. findings indicate that generational cohorts based on chronological age do matter in the formation of work related ties. these findings also support the importance of focusing on different professional networks since different age dynamics can be at play. our findings also show that school team members of the youngest cohort tend to form intra-generational relationships, whereas older generational cohort members prefer inter-generational relationships. this study is innovative due to its application of social network analysis to investigate intergenerational knowledge flows. keywords: intergenerational learning; school teams; teacher development; social network analysis; p2 modelling geeraerts  et  al       | f l r     79   1. introduction nowadays, the role of knowledge management within schools as an organizational context has received more attention due to its potential to encourage innovative practices and to avoid knowledge loss within school teams (thambi & o'toole, 2012). knowledge loss can occur when workers leave the profession without sharing their knowledge or without turning implicit knowledge into an explicit mode. similar to other countries, schools in the netherlands are confronted with a large outflow of older teachers and a challenge to retain young teachers into the teaching profession. as compared to secondary school teams, elementary schools are often smaller organizations and characterized by a more cohesive organizational culture (johnson, 1990). also, the tasks of elementary school teachers show more similarities than for secondary school teachers. therefore, we assume elementary school teams to be a fruitful context for exchange of knowledge which offers an interesting case to investigate how professional relationships are shaped. facilitating intergenerational learning interactions seems to be a promising way to prevent knowledge loss within organizations (gerpott, lehmann-willenbrock, & voelpel, 2016; ropes, 2013; starks, 2013). intergenerational learning in school teams is mainly conceptualised as an interactive process between teachers of different generations that results in learning from one or both parties (novotný & brücknerová, 2014; ropes, 2011). in this study, we refer to generations of teachers by using the term generational cohort, which in turn refers to being born in the same chronological time period. individuals of generational cohorts are found to possess different kinds of knowledge (gerpott et al., 2016). previous research within the context of school teams showed that teachers’ knowledge varies depending on their generational cohort or level of experience. for instance, young teachers are perceived to possess innovative teaching methods and ict skills, while teachers of the oldest generational cohort are perceived to have excellent classroom management skills and content knowledge (geeraerts, vanhoof, & van den bossche, 2016). simultaneously, classroom management skills are known to be a challenge for beginning teachers (wolff, van den bogert, jarodzka, & boshuizen, 2015). these findings make age diversity and intergenerational learning within school teams relevant. moolenaar (2010) highlighted the importance of interactions between school team members to facilitate knowledge sharing and learning. consequently, intergenerational knowledge sharing can be understood as a socio-constructive process in which interaction plays a facilitating role (geeraerts et al., 2016; novotný & brücknerová, 2014). this implies that school team members must be aware of resources such as information, knowledge and expertise of their colleagues, and make use of their social relationships to access these assets. dynamics of sending and receiving professional relationships between or within generational cohorts have the potential to facilitate knowledge flows within school teams. these flows can be influenced by the fact that different generations of teachers also differ in life and work experiences or feelings. for example, a great number of early career teachers face a practice shock that is accompanied by feelings of uncertainty (pillen, beijaard, & den brok, 2013; stokking, leenders, de jong, & van tartwijk, 2003). in addition, young teachers fear being perceived as incompetent or vulnerable by their more experienced colleagues, also described as feeling evaluated (kelchtermans & ballet, 2002). on the contrary, older teachers or experienced teachers are perceived to have a high level of self-confidence in their profession (geeraerts et al., 2016). consequently, we expect that teachers of different generations show differences in the formation of professional relationships. in this study, we examine the formation of relationships in elementary school teams in the netherlands, focusing on discussing work, asking and providing advice, and collaboration. we label these relationships as professional relationships and question them from an intergenerational perspective, meaning that these relationships can be formed within or across different generations of school team members. hereto, we apply methods from social network analysis (sna) enabling us to investigate social relationships and to look more deeply into interactions among school team members. when investigating professional relationships we focus on sending and receiving relationships. specifically, this study investigates to what extent school team members of different generational cohorts differ in the number of geeraerts  et  al       | f l r     80   professional relationships they send and receive, and secondly, how being part of the same generational cohort affects the likelihood of engaging in professional relationships. in the following, we start with framing our work within the literature of generations, and school team members’ social relationships. the latter part builds on social capital theory and social network theory. 2. theoretical framework 2.1. generational diversity among school team members the concept of generations was first introduced by mannheim (1952) and referred to a generation as a group of individuals who share mutual social and historical events during their lifespan. according to edge (2014) school teams include three generations of teachers and leaders: baby boomers (1946-65), generation x (1966-80) and generation y (1981-2003). in literature, there are a lot of inconsistencies concerning the labels of these generational cohorts and the boundaries used to determine a generational cohort. brücknerová and novotný (2016) relied in their study on intergenerational learning of teachers on teachers’ own perceptions of themselves as a member of a particular generation. most studies used the chronological dimension of age to frame school team members’ generational cohorts (edge, descours, & frayman, 2016; geeraerts et al., 2016). looking at school teams through a generational lens sheds light on the level of age diversity within school teams (brücknerová & novotný, 2016). previous studies in other contexts have already recognized the benefits of age diversity in teams since individuals contribute different kinds of information, knowledge, skills, and expertise to the team (gerpott, lehmann-willenbrock, & voelpel, 2016; williams & o'reilly, 1998). the traditional point of view in which older workers are perceived as experts is questioned nowadays. fuller and unwin (2004) acknowledged that young workers have already developed different kinds of knowledge and skills, and also have higher educational levels than many of their older counterparts with whom they interact in the workplace. this finding brings the importance of the bidirectional character within the learning process under attention. previous research by geeraerts et al. (2016) showed that teachers of different generational cohorts are perceived to possess different kinds of knowledge. whereas young teachers were seen as a knowledge source for innovative teaching methods and ict skills, teachers older than 50 were rather associated with classroom management skills and subject knowledge. many studies have focused on differences between novice teachers and experienced ones (wolff et al., 2015), but studies in which the interactions between both parties are investigated are rather scarce. thus, interactions between teachers of different generations can provide opportunities to learn from each other’s knowledge, especially when these learning processes are characterized by bidirectional interactions instead of unidirectional ones. these findings underline the added value of the formation of relationships across different generations of teachers for the construction and transfer of knowledge, and raises important questions on knowledge management within school teams. 2.2. the social side of learning across generations in this paper, we argue that professional relationships between teachers of different generations may provide opportunities to learn from each other. this notion builds on a more 'social' interpretation on how learning takes place. lee et al. (2004) highlight that learning is a complex concept due to different approaches to learning. the traditional approach to learning describes 'learning as acquisition’ and builds on cognitive psychology and behaviourism (sfard, 1998). this learning as acquisition approach focusses mainly on the individual. a more recent approach refers to 'learning as social participation’ and focusses on learning geeraerts  et  al       | f l r     81   through social relations and participation of individuals within communities of practice (lave & wenger, 1991; sfard, 1998). this implies that learning can be understood as a social, interactive process of coconstruction. learning in the workplace is mainly informal and involves learning from colleagues on the job (eraut, 2004). knowledge acquisition and access to information from others are equally important contributors to learning processes in workplaces (ashton, 2004). therefore, many researchers in the field of workplace learning emphasize the importance of social relationships for informal learning (doornbos, bolhuis, & simons, 2004; eraut, 2004; tynjälä, 2008). also in terms of teacher learning, the social side of learning is emphasised based on the idea that cognition is situated in nature (e.g. kwakman, 2003; lohman, 2000, 2006; meredith, van den noortgate, struyve, gielen, & kyndt, 2017; van waes, van den bossche, moolenaar, de maeyer & van petegem, 2015). a shift from a focus on the individual towards a more social approach contributed to the popularity of social network research methods to investigate relationships of teachers (baker-doyle, 2015). we will now further explore how knowledge sharing among generational cohorts may contribute to intergenerational learning by first zooming into teachers' relationships, and then, into some relevant social network concepts and dynamics. 2.3. school team members’ social relationships the attention for social relationships between teachers is underpinned by research on the importance of social capital for school improvement and instructional reform (spillane, kim, & frank, 2012). social capital within an organization reflects an investment in social relationships through which valuable resources such as knowledge, information and expertise can be accessed, borrowed, or leveraged (daly, 2010; lin, 1999). the general concept of social capital provides a framework to conceptualize how individuals have access to resources (e.g. information, expertise) by the web of social relationships surrounding them, thereby offering (or hindering) opportunities for (intergenerational) learning. borgatti, everett, and johnson (2013) distinguish two main types of relations: relational states and relational events. relational states refer to continuously persistent relationships between individuals (e.g. being a colleague, teacher or friend), whereas relational events refer to discrete events such as interactions (e.g. asking a colleague for advice). the outcomes of interactions are flows, and can contain, for instance, information, knowledge, and expertise (borgatti et al., 2013). we build on earlier work, suggesting that social relationships offer opportunities for knowledge creation, knowledge retention, and knowledge transfer (argote, mcevily, & reagans, 2003). in school teams, collegial relationships have the potential to initiate occasions to learn from each other within or between generational cohorts. informal learning interactions between teachers occur in the form of engaging in dialogue, collaborating, sharing resources such as information, lesson materials, ideas, advice, etc. (baker-doyle, 2015; kwakman, 2003; lohman, 2006). accordingly, knowledge flows are the result of interactions between two teachers through which information is exchanged. in addition, social relationships can be described in terms of the content of the relationship. ibarra (1993, 1995) distinguishes between expressive and instrumental relationships. this distinction also applies to school teams (moolenaar, 2010). whereas expressive relationships do not directly aim at work related issues (e.g. friendship), instrumental relationships do aim to achieve organisational goals. for instance, work related discussions and asking questions are interactions that enable exchange of expertise (gerpott et al., 2016). furthermore, spillane et al. (2012) see advice and information seeking relationships as critical for teachers’ professional development and for knowledge development. also, relationships in terms of teacher collaboration can be seen as an indicator of informal learning within school teams (richter, kunter, klusmann, lüdtke, & baumert, 2011). all of these relationships: discussing work, asking advice, providing advice, and collaboration, provide opportunities for intergenerational learning and can be labelled as instrumental (geeraerts et al., 2016; novotný & brücknerová, 2014). moolenaar (2010) found only partial overlap between these different networks, which highlights the semi-unique character of these networks as a source for exchange of knowledge, expertise, teaching materials, and other resources valuable to teacher geeraerts  et  al       | f l r     82   learning and school performance. this study focuses on instrumental relationships and labels them as professional relationships due to its professional nature and closer link to organisational benefits. taken together, professional relationships among school team members are essential since they provide access to social resources such as knowledge, information, and expertise. school team members can only benefit from these resources when they have access to them through social interactions, and these interactions are facilitated by social relationships. a research method that allows us to investigate the formation of school team members’ relationships is social network analysis. in the following, we discuss degree centrality and network homophily as two important network concepts. 2.4. social network concepts and dynamics 2.4.1. degree centrality social network research has the potential to reveal the underlying network structure in an organization so that more insight in the exchange of resources within an organization can be established (cross, parker, & borgatti, 2002). in this study, networks are represented by school teams in which school team members are the actors or nodes. centrality is a commonly used concept within social network research that focuses on the position of a node or actor within a network (borgatti et al., 2013). the concept of centrality identifies the structural importance of a node, by looking at how many connections or relations one node has to other nodes. within social network research, these relationships are often referred to as ties. since our study aims to investigate relationships, we do not approach our data from the perspective of centrality as such, but rather from the perspective of interactions, measured by degree centrality. degree centrality is a frequently used measure for relationships within networks, and it refers to the number of ties a node has to other nodes. in a directed graph or network, degree centrality has two types: indegree and outdegree (borgatti et al., 2013). indegree centrality involves the number of incoming ties of an actor within a network. it can be seen as a measure of individual popularity (in the case of a positive tie network), since this measure is the number of colleagues by whom the respondent, or school team member, was nominated. outdegree centrality counts the number of outgoing ties of an actor within a network. it involves the number of colleagues nominated by the respondent, or school team member, which suggests a measure of individual activity (borgatti et al., 2013). consequently, degree centrality refers to the individual node level. a normalized inand outdegree score can be interpreted as the percentage of relationships that school team members maintain within the whole network. we focus on the calculation of degree centrality within the networks of discussing work, asking and providing advice, and collaboration, since these networks can be seen as a potential indicator of learning within school teams. moreover, we assume that there might be differences in degree centrality between different generational cohorts. for instance, kelchtermans (2006) mentioned that asking for advice from a colleague might be seen as a request for help, which is accepted for young or beginning teachers but not for experienced ones. this might imply that teachers of the youngest cohort are more likely to form more ties to ask advice than teachers of the older generational cohorts. spillane et al. (2012) see advice and information relationships as critical for teachers’ professional development and for knowledge development. in their study, more experienced teachers were less likely to receive advice and information from other colleagues, as compared to early career teachers (spillane et al., 2012). in terms of teacher collaboration, previous research of richter et al. (2011) showed that young teachers tend to collaborate more frequently than older colleagues. teacher collaboration seems to decrease with age (richter et al., 2011). regarding work related conversations, beginning teachers are more likely to interact with colleagues in order to overcome professional challenges and to exchange teaching ideas, as compared to experienced colleagues (grangeat & gray, 2007). geeraerts  et  al       | f l r     83   2.4.2. network homophily the concept of network homophily, often referred to by the proverbial expression ‘birds of a feather flock together’, captures the idea that individuals are more likely to have ties with others who are similar to themselves on attributes such as age, race, gender, education, and values, than with individuals that are dissimilar to them (feld, 1982; mcpherson, smith-lovin, & cook, 2001). a study of marsden (1988) revealed that the greater the age difference between individuals, the less likely they were to discuss important matters with each other. in particular, the youngest age cohorts tend to have confiding relationships with individuals of their own cohort. whereas degree centrality focused on the individual level, network homophily refers to similarities between individuals and therefore focuses on the dyad level within networks. these dyads can be mutual, asymmetric or null dyads (wasserman & faust, 1994). in mutual dyads, actors within the network choose each other; ties are reciprocated. an asymmetric dyad refers to a one-directional tie in which one actor chooses the other actor, without being reciprocal. a null dyad indicates the absence of a tie between two actors. when dyads occur between actors with similar attributes, in our case being part of the same generational cohort, this can be labelled as homophily. individuals with similar background characteristics are more likely to have mutual experiences which in turn results in shared knowledge (reagans & mcevily, 2003). according to reagans and mcevily (2003) this common knowledge has a positive effect on knowledge transfer. therefore, we expect that teachers of the same generational cohort are more likely to engage in work related or so-called instrumental interactions. this concept of network homophily is also supported by the ideas of social identity theory (tajfel & turner, 1986). this theory follows a similar reasoning, suggesting that individuals have more positive perceptions towards people who are similar to them, compared to people who are dissimilar. this results in categorizations of in (“us”) and out (“them”) groups. within work-group diversity research, williams and o'reilly (1998) referred to this way of categorizing as a social categorization perspective. similarity of age characteristics can be seen as a trigger for these inand out-group categorizations (dencker, joshi, & martocchio, 2007). consequently, teachers might be more likely to form ties with colleagues of the same generational cohort. this implies that resources such as information and knowledge tend to flow within the particular generational cohort. following this reasoning, both outgoing and incoming relationships, also referred to as ties of sender and receiver, occur within a generational cohort rather than between different generational cohorts. 3. research questions we investigate whether belonging to a certain generation affects individual school team members’ likelihood of having relationships in networks through which resources can flow. more specifically, we focus on four instrumental networks, referring to professional relationships in which school team members discuss their work, receive or provide advice, and collaborate, since these relationships can be relevant for learning in school teams (geeraerts et al., 2016; novotný & brücknerová, 2014). therefore, the following research questions are set forward: rq1: to what extent do school team members of different generational cohorts differ in the number of professional relationships (indegree and outdegree; sending and receiving relationships)? rq2: to what extent does being part of the same generational cohort affect the likelihood of engaging in professional relationships within school teams (network homophily)? geeraerts  et  al       | f l r     84   4. methodology 4.1. sample the data from this study was collected at 15 elementary schools in the netherlands. the data collection was part of a larger project on school improvement in the netherlands in which 53 schools participated. all schools were organised as a cluster of catholic schools, supported by a single catholic school board. we selected the subsample of schools based on the following criteria: school team size was 10 or more teachers, and in each school, each generational cohorts was represented by at least 20% of the respondents. this resulted in a final selection of 15 schools, which enabled us to investigate intergenerational relationships in schools where all generational cohorts were sufficiently represented. the sample consisted of all principals and teachers, including instructional coaches (teachers with specialised instructional tasks, such as emotional/behavioral support), since we wanted this selection to be as close as possible to the core teaching team. the sample did not include temporary and replacement teachers. the overall sample contained 284 teachers and 15 school principals (n=299). generational cohorts are based on chronological age. within this sample, three generational cohorts can be distinguished. the 'young' cohort contained 94 educators aged 35 years old or younger. the 'middle' cohort contained 87 educators from 36 to 50 years old. the 'old' cohort consists of 118 teachers older than 50 years. most school members (75%) were female. the sample demographics are summarized in table 1. table 1 sample demographics (n schools= 15, n respondents=299) categories teachers principals total number of school team members percentage generational cohort young cohort (-36 yrs) 93 1 94 31% middle cohort (36 – 50 yrs) 84 3 87 30% old cohort (50+ yrs) 107 11 118 39% gender male 63 11 74 25% female 221 4 225 75% total number 284 15 299 100% 4.2. data collection the survey included questions on job satisfaction, leadership, school team, strategy and policy, processes, citizenship in the classroom and general questions. the section on ‘school team’ contained sociometric questions that questioned the social networks within the school team. network questions used in this study were: geeraerts  et  al       | f l r     85   • discussing work: whom do you turn to in order to discuss your work? • asking advice: whom do you prefer to go to for work related advice? • providing advice: to whom do you give work related advice? • collaboration: with whom do you like to collaborate the most? to answer these socio-metric questions, respondents were provided with a list of their school team members. according to marsden (2011), this list helps respondents to remind the alters in their network and so it also minimizes measurement error. in order to contribute to the premise of anonymous analysis of the data, the alter list contained a letter code for each alter (e.g. jessica thompson = ab). respondents were asked to indicate this letter code by completing the survey instead of mentioning the name and surname of the alter. there was no limitation to the number of colleagues a respondent could indicate as part of his/her network. an example of the visualization of a school team network is displayed in figure 1. figure 1. example of ‘asking advice’ network. 1 4.3. measures our dependent variable is the existence or the absence of a professional relationship between two school team members (a dyad). concretely, for every pair of school team members i and j, a value of 1 represents a relationship between i and j. for instance, i provides advice to j. a value of 0 indicates the absence of a tie between i and j. the mathematical representation of these relationship is an adjacency matrix composed by 0s and 1s (van duijn & vermunt, 2006). 4.3.1. individual level measures involve characteristics of the individual school team members. generational cohort in line with the findings of a study by richter et al.(2011), based on spearman’s correlation we found a high correlation between age and the number of years of teacher experience within education or within the school, r=0.84 and r=0.52 respectively. this implies that these measures are nearly interchangeable and suggests that the number of teachers who enter the teaching profession at later stages in their life is limited. consequently, we only take into account the variable of generational cohort based on chronological age. this measure refers to three age related categories: young, middle and old cohort. in the survey, teachers were asked to indicate their age (1 = 20-25 years, 2 = 26-30 years, 3 = 31-35 years, 4 = 36-40 years,                                                                                                                           1 the youngest cohort is represented by circles, the middle cohort by squares, and oldest cohort by triangles geeraerts  et  al       | f l r     86   5 = 41-45 years, 6 = 46-50 years, 7 = 51-55 years, 8 = older than 55). these age categories were first recoded into three categories that are in line with generational cohorts that can be found in the literature under the labels of generation y, generation x, and the baby boomers generation (young = 20-35 years, middle = 3650 years, old = 51 and older) (edge, 2014; geeraerts et al., 2016; glass, 2007; novotný & brücknerová, 2014). consequently, this categorical variable contains three categories that each cover approximately 15 years. finally, two dummy variables were generated, in which ‘middle’ and ‘old’ were contrasted with the ‘young’ cohort. gender this individual feature is coded in the following way: a value of 0 refers to male school team members, and a value of 1 refers to female school team members. function this measure takes a value of 0 for a teacher position, and a value of 1 for a principal position. 4.3.2. dyadic level measures, also named relationship covariates. generational cohort similarity three dummy variables of a generational cohort are used: youngest cohort, middle cohort, oldest cohort. the ‘absolute difference’ function in the p2 module is used to investigate the likelihood of a relationship when actors are part of a different generational cohort, in other words, not being part of the same generational cohort. 4.4. data analysis in order to respond to rq1, social network properties at the individual level were calculated by using software package ucinet 6.0 (borgatti et al., 2013). normalized degree centrality, both indegree and outdegree, was calculated. these measures can be interpreted as percentages. consequently, these normalized inand outdegree scores have a value from 0 to 100, in which 0 indicates the absence of relationships, and 100 refers to being tied to the entire school team. the variations in the percentages of incoming and outgoing relationships are provided by the standard deviations of the normalized inand out degree measures. in addition, the statistical software program ibm spss 24 was used to measure the effects of individual node characteristics on the network properties as a dependent variable. a one-way analysis of variance (anova) was conducted on each network question. for comparisons of the mean scores among the different generational cohorts, the post hoc test tukey with its significant difference procedure (α=0.05) was used. regarding rq2, we used the p2 package within the social network software stocnet (boer et al., 2006). by using p2 modelling, we investigated dyadic ties as the dependent variable. dyadic level factors focus on similarities and differences. the p2 model is a model for the statistical analysis of directed binary relationship data with actor and/or dyadic covariates (boer et al., 2006; zijlstra & van duijn, 2003). as such, a p2 model is designed to predict the likelihood of the formation of social relationships (e.g. work discussions) between pairs of actors (e.g. teachers and principals) based on individual and dyadic variables (e.g. belonging to a generational cohort (boer et al., 2006; zijlstra & van duijn, 2003). p2 models can be seen as a type of logistic regression model that takes the dependency between relationships from one to another actor into account (lazega & van duijn, 1997). ordinary logistic regression models cannot be used here since the assumption of data independence is violated. p2 modelling specifically focusses on complete, directed networks, which implies that every actor within the network can have ties with all other actors. however, the model can handle (some) missing data (van duijn & vermunt, 2006). the multilevel variant of the p2 model is an extension of the p2 model which can be used for the analysis of multiple networks. parameter estimates of the p2 model geeraerts  et  al       | f l r     87   and the multilevel p2 model derive from the markov chain monte carlo (mcmc) procedures which are integrated in the p2 module of the social network analysis software stocnet (boer et al., 2006; zijlstra, van duijn, & snijders, 2006). the p2 model is designed to compute the likelihood of sending a relationship (cf. out-degree; called sender effect), receiving a relationship (cf. in-degree; called a receiver effect), and the likelihood of engaging in a relationship based on dyadic similarity (cf. homophily; called a reciprocity effect). a positive significant parameter estimate indicates a positive effect of the variable on the likelihood to form a relationship. for example, a positive significant sender effect of gender (male/female) indicates that female teachers have a higher likelihood to send relationships within the network than male teachers. in order to investigate homophily effects, the p2 software constructs dyadic matrices based on the absolute difference between two actors within the network. for instance, a dyad between a school team member of the youngest generational cohort and a school team member of the middle generational cohort represents a relationship between school team members of a different generational cohort. this absolute difference between being part of the youngest and oldest cohort (dummy variable=0) and being part of the middle cohort (dummy variable=1) is 1. in this example, a negative parameter estimate would suggest that a difference in generational cohort is related to a lower likelihood of having relationships. consequently, a negative parameter suggests that relationships between members of the same generational cohort are more likely to occur. as such, a negative reciprocity effect signals the existence of a homophily effect. regarding the significance level of the parameter estimates, the p2 output in stocnet does not directly provide this information. an additional wald test needs to be calculated by dividing the parameter estimate by the standard error of the estimator. when the ratio is smaller than -2 or larger than 2, a significant effect occurs at 0.05 level. 5. findings 5.1. to what extent do generational cohorts differ in the number of professional relationships? the first part of the study investigated if teachers of different generational cohorts show differences in degree centrality measures of the network questions. a one-way analysis of variance (anova) was conducted on each network question. the independent variable contained the generational cohorts and the dependent variable contained indegree centrality and outdegree centrality. 5.1.1. the number of outgoing professional relationships no significant generational cohort differences were found with regard to normalized outdegree centrality measures of the networks, as displayed in table 2. this implies that school team members of different generational cohorts do not significantly differ statistically, with regard to the average number of sending ties within the networks of discussing work, asking advice, providing advice, and collaboration. geeraerts  et  al       | f l r     88   table 2 anova of normalized outdegree (n=15, n=299) young (1) middle (2) old (3) anova main effect m sd m sd m sd f p discussing work 0.32 0.19 0.33 0.22 0.27 0.19 2.448 0.088 asking advice 0.20 0.15 0.17 0.14 0.16 0.13 2.304 0.102 providing advice 0.20 0.19 0.22 0.23 0.19 0.19 0.460 0.632 collaboration 0.27 0.20 0.27 0.24 0.25 0.22 0.362 0.697 5.1.2. the number of incoming professional relationships regarding normalized indegree centrality, the main effect analysis found a statistically significant generational cohort difference in the network question ‘providing advice’ [f(2, 297)=9.003, p=0.000] and the network question ‘collaboration’ [f(2, 297)=9.367, p=0.000]. no significant differences were found for the networks of ‘discussing work’ and ‘asking advice’. descriptives of the mean values for the normalized indegree of the generational cohorts are displayed in table 3. table 3 anova of normalized indegree (n=15, n=299) young (1) middle (2) old (3) anova main effect post hoc tukey m sd m sd m sd f p discussing work “being chosen to discuss work with” 0.32 0.21 0.29 0.17 0.30 0.19 0.530 0.589 asking advice “being asked for advice” 0.16 0.18 0.18 0.16 0.18 0.16 0.633 0.532 providing advice “being provided with advice” 0.24 0.14 0.21 0.12 0.17 0.12 9.003** 0.000 1 > 3 ** collaboration “being chosen to collaborate with” 0.32 0.18 0.25 0.12 0.23 0.14 9.367** 0.000 1 > 3 * 1 > 2 ** * significant at 0.05 level ** significant at 0.01 level the post hoc test tukey with its significant difference procedure (α=0.05) was used for comparisons of the mean scores among the different generational cohorts. regarding the indegree of ‘providing advice’, members of the youngest cohort (m=0.24, sd=0.14) were found to have significantly higher ratings than members of the oldest cohort (m=0.17, sd=0.12). this implies that the youngest school team members geeraerts  et  al       | f l r     89   receive advice from more colleagues than the members of the oldest generational cohort do. whereas young school team members are tied to on average 24% of the network, the oldest cohort is tied to on average 17% of the school team. there were no statistically significant differences found between the other generational cohorts in terms of indegree on giving advice. regarding the indegree of ‘collaboration’, the youngest group of team members (m=0.32, sd=0.18) was found to have significantly higher scores than members of the middle and oldest cohorts (m=0.25, sd=0.12; and m=0.23, sd=0.14, respectively). this implies that the youngest team members are chosen by more colleagues to collaborate with, as compared to the two older generational cohorts. young school team members form ties with on average 32% of their network, in contrast to 25% for the middle cohort and 23% for the oldest cohort. the above described anova analysis and findings provide insight in the average number of outgoing and incoming ties, however, insight in which generational cohort sends ties to and receives ties from which cohort is missing. in addition, the anova did not control for gender and function, which might give a different picture of sending and receiving ties. also, the occurrence of a homophily effect could not be tested by the previous analysis. therefore, we ran the p2 analysis to further our understanding in these dynamics. parameter estimates of the multilevel p2 models for investigating the effect of individual and dyadic level demographics on the likelihood of having relationships within the networks of discussing work, asking advice, providing advice, and collaboration, are presented in table 4. geeraerts  et  al       | f l r     90   table 4 the effect of sender and receiver demographic variables on the likelihood of having relationships within the networks of discussing work, asking advice, providing advice, and collaboration. parameter estimates of the multilevel p2 models. (n=299) network discussing work asking advice providing advice collaboration pe (se) pe (se) pe (se) pe (se) overall effects density -1.98 (0.22) -2.12 (0.15) -2.51 (0.22) -2.17 (0.25) reciprocity 2.63 (0.18) 2.21 (0.17) 2.11 (0.18) 2.21 (0.21) sender covariates middle cohort 0.10 (0.21) -0.36 (0.22) 0.40 (0.23) 0.29 (0.27) old cohort -0.28 (0.18) -0.39 (0.13) 0.47 (0.19) 0.15 (0.21) gender 0.20 (0.15) 0.00 (0.00) 0.25 (0.18) 0.25 (0.21) function -0.09 (0.31) 0.62 (0.33) -0.06 (0.30) 0.86 (0.36) receiver covariates middle cohort -0.16 (0.19) 0.28 (0.12) -0.37 (0.14) -0.44 (0.17) oldest cohort -0.08 (0.16) 0.07 (0.15) -0.60 (0.13) -0.55 (0.14) gender 0.15 (0.17) 0.00 (0.00) 0.24 (0.13) 0.21 (0.14) function 1.35 (0.29) 0.97 (0.21) -0.38 (0.25) -0.17 (0.28) relationship covariates youngest cohort -0.21 (0.06) -0.18 (0.08) -0.15 (0.08) -0.15 (0.07) middle cohort 0.01 (0.07) 0.08 (0.08) 0.03 (0.09) 0.10 (0.08) oldest cohort -0.10 (0.06) -0.08 (0.07) -0.04 (0.08) -0.09 (0.07) random effects sender variance 1.19 (0.15) 0.72 (0.11) 0.54 (0.18) 1.78 (0.22) receiver variance 1.06 (0.13) 1.18 (0.16) 0.36 (0.08) 0.68 (0.11) covariance -0.78 (0.12) -0.62 (0.11) -0.58 (0.11) -0.71 (0.13) note: pe= parameter estimate; se= standard error; bold typeface refers to a significant pe; n=9309 dyadic relations from 299 school team members in 15 elementary schools geeraerts  et  al       | f l r     91   first of all, overall effects show negative density effects and positive reciprocity effects within the four networks. these findings suggest that the networks are overall rather sparse, meaning that the likelihood of having a tie is lower than 50% for the reference group, which are dyads of young male teachers. the positive parameter estimates of reciprocity indicate a tendency of reciprocated ties instead of unidirectional ties throughout the different networks. with regard to the random effects, the positive and significant effects of sender and receiver variance indicate that there is considerable variation among school team members in the amount of ties they send and receive within the four networks. the negative sender-receiver covariance suggests that school team members who report to send more ties have a lower likelihood of receiving ties within their network, when allowing for differences between schools. 5.1.3. the likelihood to send professional relationships looking at the sender covariates, we found no significant effects for the network discussing work. in other words, none of the individual characteristics affected the likelihood of sending ties in a positive or negative way. more specifically, school team members of the middle or oldest cohort did not send more ties than young school team members, female not more than male, and school principals not more than teachers. within the other three networks, results indicated that some of the individual characteristics affected the likelihood of sending relationships. within the network of asking advice, being part of the oldest cohort decreases the likelihood of asking advice. in addition, the oldest cohort tends to send significantly more relationships of providing advice. within the network of collaboration, results reveal that being a school principal increases the likelihood of sending collaboration ties. 5.1.4. the likelihood to receive professional relationships regarding the receiver covariates, significant effects were found in all the four networks. being a school principal increases the likelihood of receiving relationships in discussing work and asking advice networks. also, being part of the middle cohort increases the likelihood of receiving relationships in the network of asking advice. this implies that principals and teachers of the middle cohort are more likely to be sought out for advice. in addition, the network of ‘providing advice’ shows that teachers of the middle and oldest cohort have a lower likelihood to receive advice relationships. the same trend can be found within the network of collaboration. both individual characteristics, being part of the middle and being part of the oldest cohort, decreases the likelihood of receiving collaboration ties. for both networks, providing advice and collaboration, the effect of being part of the oldest cohort is stronger than the middle cohort. this suggests that being part of the oldest cohort decreases the likelihood of the formation of a tie to a greater extent than their colleagues of the middle cohort. gender did not affect any of the sender and receiver relationships in the four instrumental networks. 5.2. to what extent does being part of the same generational cohort affect the likelihood of engaging in professional relationships? regarding the effects of the relationship covariates, homophily effects can be found for the youngest cohort within all the networks. this finding suggests that school team members of the youngest cohort are more likely to form ties with colleagues of the same generational cohort than with colleagues of the middle or the oldest generational cohort. this tendency of homophily occurs at the level of significance for the youngest generational cohort in the networks of discussing work, asking advice, and collaboration. the other generational cohorts, middle cohort and oldest cohort, do not show a significant homophily effect in all the networks. relationships in these cohorts can therefore be described as heterogeneous. we conclude that, in particular, school team members of the youngest cohort tend to form intra-generational ties, whereas older generational cohort members form inter-generational ties. geeraerts  et  al       | f l r     92   6. conclusion and discussion in the present study we have focused on the role of being part of a generational cohort in the formation of professional relationships within elementary school teams in the netherlands. our first research question focused on differences between generational cohorts in terms of professional relationships being sent or received. to answer this question, we included both findings from descriptive anova analysis of degree centrality measures, and from a p2 model based on probability distributions. the latter analysis gave a slightly different picture which we explain by the fact that our p2 model provides a more sophisticated investigation of sender and receiver tendencies within the networks. therefore, we elaborate our conclusions primarily on the basis of this p2 model and indicate how they are in line with the anova analyses. when looking for differences in sending and receiving relationships, we noticed a significant role of generational cohort within the networks of asking advice, providing advice, and collaboration. within the network of discussing work we did not find any significant impact of belonging to a certain generation, neither in the anova and the p2 findings. next, we discuss the findings for each of the respective networks studied. regarding asking and providing advice, our results reveal that the middle cohort can be seen as an important source of advice within the school team. this is the cohort who is asked by most colleagues for advice. on the other hand, we noticed that the oldest cohort sees themselves as an important provider of advice, since this cohort provides more colleagues with advice. it must be noticed that members of the oldest cohort provide advice to colleagues who did not necessarily ask for it. the middle cohort does not certainly provide more advice, but they are asked for it by more colleagues. this finding underlines the importance of the middle cohort as a source of knowledge within the school team. in addition, we found that the youngest cohort is provided with advice by more colleagues, as compared to their older counterparts. relationships within the networks asking advice and providing advice can be interpreted as complementary. when looking at these networks asking and providing advice, it can be questioned to what extent providing advice is the result of asking advice. further research might pay attention to what degree providing advice is a voluntary action or the result of being asked for advice. when looking to the effects of generational cohort within these two advice networks, we also recognize the complementary tendencies. being part of the oldest cohort did decrease the likelihood of sending relationships for asking advice, and did increase the likelihood of sending relationships for providing advice. a second complementary tendency for asking and providing advice was found for the middle cohort. while being part of the middle cohort increased the likelihood of receiving relationships within the network asking advice, it decreased the likelihood of receiving relationships within the network of providing advice. being part of the oldest cohort decreased the likelihood of being provided with advice even more than the middle cohort. the anova results were supporting these findings, revealing that young school team members were more provided with advice than their oldest counterparts. a similar tendency has been observed by spillane et al. (2012), where more experienced colleagues were less likely to receive advice. our findings of both asking and providing advice relationships do relate to the traditional mentor models in which older or experienced teachers serve as knowledge providers to younger ones. within recent developments and ideas on intergenerational learning, knowledge supplies and demands are seen as important for all generational cohorts. consequently, our findings bring the importance of stimulating intergenerational relationships under attention. within the last network, collaboration, we found that being part of the oldest cohort did decrease the likelihood of receiving collaboration ties. a decrease of receiving collaboration relationships was also found for the middle cohort. these findings were also supported by the anova results. the youngest cohort is the most preferred cohort to collaborate with. previous research revealed that older teachers have positive perceptions towards their youngest counterparts in terms of enthusiasm and creativity (geeraerts et al., 2016). these positive perceptions about young teachers might contribute to preference to collaborate with them. we found similarities within the tendencies of sending and receiving relationships between the providing advice network and collaboration network. this raises questions on the existence of overlap geeraerts  et  al       | f l r     93   between both networks. further research might investigate the extent to which instrumental relationships show overlap. whereas gender did not affect the likelihood to send or receive professional relationships, the other control variable function did. principals tend to mention more different colleagues they prefer to collaborate with. in addition, principals were also chosen by more colleagues to discuss work with and to ask advice to. further research might foreground the role of the school principal and investigate, for instance, the network position of the principal within professional school team networks, and, the role of the principal’s age. regarding our first research question we conclude that, for some networks, generational cohorts based on chronological age do matter in the formation of professional relationships. these findings also underline the importance of focussing on different instrumental networks since different age dynamics can be at play, and therefore give support to the findings of moolenaar (2010) to approach different professional networks as unique networks. future research might include expressive relationships in addition to instrumental relationships, for instance, by including friendship relationships. by doing this, the extent to which instrumental relationships are explained by expressive relationships can be investigated. our second research question captured the mechanisms of homophily within elementary school teams. teachers of the youngest cohort in particular seem to form relationships within their own generation for discussing work, asking advice, and collaboration, which is in line with the homophily effects for the young individuals in other contexts found by marsden (1988). from the viewpoint of intergenerational learning and intergenerational knowledge sharing, this is a worrying finding. this suggests that facilitating bidirectional intergenerational relationships is important for practice. further research, for instance by using qualitative methods, may dive deeper into the reasons why young teachers have this tendency. factors such as the level of trust, a safe and respectful climate, as perceptions of being evaluated or hierarchical perceptions of seniority are worth taking into account. in addition, it is worth to investigate how this tendency of homophily relates to early career teacher dropout and challenges such as practice shock or feelings of uncertainty (pillen, beijaard, & den brok, 2013; stokking, leenders, de jong, & van tartwijk, 2003). also, literature on intergenerational relationships often focusses on age bias and generational stereotyping within organizations (e.g. king & bryant, 2017; rupp, vodanovich, & credé, 2006). the effects of age stereotyping in school teams on the formation of professional relationships and intergenerational knowledge exchange might offer interesting starting points for further research. this study has limitations that suggest additional paths for future research. first of all, sending and receiving advice ties are only general indicators of knowledge flows. the content of the advice ties has not been included in this study. given the idea that teachers of different generations are able to provide different kinds of information and knowledge, it might be interesting to investigate content related advice relationships, for instance, whom do you go to for advice on classroom management? this will provide information on which knowledge can be seen as a demand or supply for a certain generational cohort. a second limitation is related to the fact that we did not have information about the amount of advice shared and the frequencies of interactions within the networks. an actor can receive more advice from one colleague than from a number of different colleagues. further research can more explicitly map the strength of teacher relationships across generational cohorts by looking at the frequency, length, and duration of contact (e.g. van waes et al., 2015). also, the relevance of the received advice has not been discussed yet; to what extent is provided advice valuable advice to a teacher? exchange of information or advice is no guarantee for learning to occur. for instance kyndt, vermeire, and cabus (2016) did not find a significant relationship of knowledge acquisition and access to information with informal workplace learning outcomes. this finding underlines the importance of further evaluating the relevance of the advice that is being shared (e.g. van der rijt, van den bossche, van de wiel, de maeyer, gijselaers, & segers, 2013; van der rijt, van den bossche, van de wiel, segers, & gijselaers, 2012). similarly, information on the quality of relationships between generational cohorts would be a valuable contribution for further research. this also opens up the discussion on the ‘unknown side of intergenerational learning’ which refers to the idea that intergenerational learning does not necessarily lead to positive learning outcomes. this is also connected to newly emerging geeraerts  et  al       | f l r     94   social network research on ‘negative ties’. an interesting perspective on knowledge sharing within school teams might be to investigate the opposite, for instance, knowledge hiding. further research might focus on reasons why school team members would be shielding knowledge from their school team. previous research suggested not to use a too narrow approach on ‘generation’, but also to take into account factors such as work tenure and years of job experience (geeraerts et al., 2016; kooij, de lange, jansen, & dikkers, 2008). we did not include years of experience since this variable was too highly correlated with our age variable, which might cause problems of multicollinearity. we would argue that the conceptualisation of generations and operationalization of generations of teachers needs further elaboration. for instance, future research might focus on the relevance of age boundaries of generational cohorts or reveal whether there exists a linear effect of age? the division of generational cohort in this study was based on previous studies within the context of school teams (edge, 2014; geeraerts et al., 2016). in this study, our generational cohorts are diverse. for instance, the youngest cohort includes both inexperienced teachers in their induction phase and teachers with 10 years of experience. this division has not been included in this study, but potentially offers an important perspective to further unravel the complexity of teacher generations. due to the selection of our sample in which we targeted balance in terms of the presence of three generational cohorts within school teams, the generalizability of our study is limited to schools that are characterized by equal distribution of generational cohorts. an interesting path for further studies is to look at schools with different constellations of generational cohorts and examine whether these different constellations paint a similar picture regarding the formation of intergenerational relationships within school teams. would our results remain if a certain generational cohort is absent within a school team? differences in age demographic profiles of school teams might raise interesting questions for further research. also, our analysis technique (multilevel p2 modelling) has certain limitations since this model is restricted to dyadic relationships (spillane et al., 2012). the extent to which relationships between pairs of teachers are influenced by their relationships within the larger structure of the network (e.g. triads) are not taken into account. based on our findings, future studies might hypothesize more complex social network models and for instance use exponential random graph models (ergm) to explore triadic structures in intergenerational relationships. to conclude, we state that being part of a generational cohort based on chronological age does matter within these elementary school teams and that it plays a role in the formation of teachers’ professional relationships. both researchers and practitioners may regard social networks as a valuable concept to contextualise and investigate teacher interactions in order to further understand and support teacher intergenerational learning and teachers’ professional development in general. keypoints being part of a generational cohort affects the formation of relationships within the networks asking and providing advice, and collaboration. different generation dynamics are at play within different networks. young teachers are more likely to be mentioned as a preferred partner to collaborate with. old teachers are less likely to ask advice, while teachers of the middle generation are more being asked for advice. young teachers in particular tend to form relationships within their own cohort. geeraerts  et  al       | f l r     95   acknowledgments we would like to thank prof. dr. marijtje van duijn (university of groningen, the netherlands) for her helpful comments on our p2 model. references argote, l., mcevily, b., & reagans, r. (2003). managing knowledge in organizations: an integrative framework and review of emerging themes. management science, 49(4), 571-582. doi:10.1287/mnsc.49.4.571.14424 ashton, d. n. (2004). the impact of organisational structure and practices on learning in the workplace. international journal of training and development, 8(1), 43-53. doi:10.1111/j.13603736.2004.00195.x baker-doyle, k. (2015). no teacher is an island: how social networks shape teacher quality. in g. k. letendre & a. w. wiseman (eds.), promoting and sustaining a quality teacher workforce (international perspectives on education and society (vol. 27, pp. 367-383). boer, p., huisman, m., snijders, t. a. b., steglich, c., wichers, l. h. y., & zeggelink, e. p. h. (2006). stocnet: an open software system for the advanced statistical analysis of social networks. version 1.7. . groningen: ics/science plus. borgatti, s. p., everett, m. g., & johnson, j. c. (2013). analyzing social networks. london: sage publications ltd. brücknerová, k., & novotný, p. (2016). intergenerational learning among teachers: overt and covert forms of continuing professional development. professional development in education, 1-20. doi:10.1080/19415257.2016.1194876 carolan, b. v. (2014). social network analysis and education. theory, methods & applications. los angeles: sage. choo, c. w. (1998). the knowing organization: how organizations use information to construct the meaning, create knowledge, and make decisions. new york, ny: oxford university press. cross, r., parker, a., & borgatti, s. p. (2002). making invisible work visible: using social network analysis to support strategic collaboration. california management review, 44(2), 25-46. daly, a. j. (2010). social network theory and educational change. cambridge: harvard education press. dencker, j. c., joshi, a., & martocchio, j. j. (2007). employee benefits as context for intergenerational conflict. human resource management review, 17(2), 208-220. doi:10.1016/j.hrmr.2007.04.002 doornbos, a. j., bolhuis, s., & simons, p. r. j. (2004). modeling work-related learning on the basis of intentionality and developmental relatedness: a noneducational perspective. human research development review, 3(3), 250-274. doi:10.1177/1534484304268107 edge, k. (2014). a review of the empirical generations at work research: implications for school leaders and future research. school leadership & management, 34(2), 136-155. doi:10.1080/13632434.2013.869206 edge, k., descours, k., & frayman, k. (2016). generation x school leaders as agents of care: leader and teacher perspectives from toronto, new york city and london. in k. leithwood, j. sun & k. pollock (eds.), how school leaders contribute to student success: springer international publishing. eraut, m. (2004). informal learning in the workplace. studies in continuing education, 26(2), 247-273. doi:10.1080/158037042000225245 feld, s. l. (1982). social structural determinants of similarity among associates. american sociological review, 47(6), 797-801. retrieved from http://www.jstor.org/stable/2095216 fuller, a., & unwin, l. (2004). young people as teachers and learners in the workplace: challenging the novice-expert dichotomy. international journal of training and development, 8(1), 32-42. doi:10.1111/j.1360-3736.2004.00194.x geeraerts  et  al       | f l r     96   geeraerts, k., vanhoof, j., & van den bossche, p. (2016). teachers' perceptions of intergenerational knowledge flows. teaching and teacher education, 56 (may 2016), 150-161. doi: 10.1016/j.tate.2016.01.024 gerpott, f. h., lehmann-willenbrock, n., & voelpel, s. c. (2016). a phase model of intergenerational learning in organizations. academy of management learning & education. doi:10.5465/amle.2015.0185 glass, a. (2007). understanding generational differences for competitive success. industrial and commercial training, 39(2), 98-103. doi:10.1108/00197850710732424 grangeat, m., & gray, p. (2007). factors influencing teachers' professional competence development. journal of vocational education & training, 59(4), 485-501. doi:10.1080/13636820701650943 johnson, j. c. (1990). the primacy and potential of high school departments. in m. w. mclaughlin, j. e. talbert & n. bascia (eds.), the contexts of teaching in secondary schools (pp. 167-184). new york: teachers college press. kelchtermans, g. (2006). teacher collaboration and collegiality as workplace conditions. a review. zeitschrift für pädagogik, 52(2), 220-237. kelchtermans, g., & ballet, k. (2002). the micropolitics of teacher induction. a narrative-biographical study on teacher socialisation. teaching and teacher education, 18, 105-120. king, s. p., & bryant, f. b. (2017). the workplace intergenerational climate scale (wics): a self-report instrument measuring ageism in the workplace. journal of organizational behavior, 38(1), 124-151. doi:10.1002/job.2118 kooij, d., de lange, a., jansen, p., & dikkers, j. (2008). older workers' motivation to continue to work: five meanings of age. a conceptual review. journal of managerial psychology, 23(4), 364-394. doi:10.1108/02683940810869015 kwakman, k. (2003). factors affecting teachers' participation in professional learning activities. teaching and teacher education, 9(2), 149-170. doi:10.1016/s0742-051x(02)00101-4 kyndt, e., vermeire, e., & cabus, s. (2016). informal workplace learning among nurses. organisational learning conditions and personal characteristics that predict learning outcomes. journal of workplace learning, 28(7), 435-450. doi:10.1108/jwl-06-2015-0052 lave, j., & wenger, e. (1991). situated learning: legitimate peripherical participation. cambridge: cambridge university press. lazega, e., & van duijn, m. (1997). position in formal structure, personal characteristics and choices of advisors in a law firm: a logistic regression model for dyadic network data. social networks, 19, 375397. lee, t., fuller, a., ashton, d. n., butler, p., felstead, a., unwin, l., & walters, s. (2004). workplace learning: main themes & perspectives. learning as work research paper 2. university of huddersfield. lin, n. (1999). building a network theory of social capital. connections, 22(1), 28-51. lohman, m. c. (2000). environmental inhibitors to informal learning in the workplace: a case study of public school teachers. adult education quarterly, 50(2), 83-101. doi:10.1177/07417130022086928 lohman, m. c. (2006). factors influencing teachers' engagement in informal learning activities. journal of workplace learning, 18(3), 141-156. doi:10.1108/13665620610654577 mannheim, k. (1952). essays on the sociology of knowledge. london, uk: routledge & kegan paul. marsden, p. v. (1988). homogeneity in confiding relations. social networks, 10, 57-76. marsden, p. v. (2011). survey methods for network data. in j. scott & p. j. carringston (eds.), the sage handbook of social network analysis (pp. 310-388). thousand oaks, ca: sage publications. mcpherson, j. m., smith-lovin, l., & cook, j. m. (2001). birds of a feather: homophily in social networks. annual review of sociology, 27, 415-444. doi: 10.1146/annurev.soc.27.1.415 meredith, c., van den noortgate, w., struyve, c., gielen, s., & kyndt, e. (2017). information seeking in secondary schools: a multilevel network approach. social networks, 50, 35-45. doi:10.1016/j.socnet.2017.03.006 moolenaar, n. m. (2010). ties with potential. nature, antecedents, and consequences of social networks in school teams. (doctoral dissertation), university of amsterdam. geeraerts  et  al       | f l r     97   novotný, p., & brücknerová, k. (2014). intergenerational learning among teachers: an interaction perspective. studia paedagogica, 19(4). doi:10.5817/sp2014-4-3 pillen, m., beijaard, d., & den brok, p. (2013). tensions in beginning teachers' professional identity development, accompanying feelings and coping strategies. european journal of teacher education, 36(3), 240-260. doi:10.1080/02619768.2012.696192 reagans, r., & mcevily, b. (2003). network structure and knowledge transfer: the effects of cohesion and range. administrative science quarterly, 48(2), 240-267. doi:10.2307/3556658 richter, d., kunter, m., klusmann, u., lüdtke, o., & baumert, j. (2011). professional development across the teaching career: teachers' uptake of formal and informal learning opportunities. teaching and teacher education, 27, 116-126. doi: 10.1016/j.tate.2010.07.008 ropes, d. (2011). intergenerational learning in organisations. a research framework. in cedefop (ed.), working and ageing. guidance and counselling for mature learners. luxembourg: publications office of the european union. ropes, d. (2013). intergenerational learning in organizations. european journal of training and development, 37(8), 713-727. doi:10.1108/ejtd-11-2012-0081 rupp, d. e., vodanovich, s. j., & credé, m. (2006). age bias in the workplace: the impact of ageism and causal attributions. journal of applied social psychology, 36(6), 1337-1364. doi:10.1111/j.00219029.2006.00062.x sfard, a. (1998). on two metaphors for learning and the dangers of choosing just one. educational researcher, 27(2), 4-13. doi:10.3102/0013189x027002004 spillane, j. p., kim, c. m., & frank, k. a. (2012). instructional advice and information providing and receiving behavior in elementary schools: exploring tie formation as a building block in social capital development. american educational research journal, 49(6), 1112-1145. doi:10.3102/0002831212459339 starks, a. (2013). the forthcoming generational workforce transition and rethinking organizational knowledge transfer. journal of intergenerational relationships, 11, 223-237. doi:10.1080/15350770.2013.810494 stokking, k., leenders, f., de jong, j., & van tartwijk, j. (2003). from student to teacher: reducing practice shock and early dropout in the teaching profession. european journal of teacher education, 26(3), 329-350. doi:10.1080/0261976032000128175 tajfel, h., & turner, j. c. (1986). the social identity theory of intergroup behavior. in s. worchel & w. g. austin (eds.), psychology of intergroup relations (pp. 7-24). chicago, il: nelson-hall. thambi, m., & o'toole, p. (2012). applying a knowledge management taxonomy to secondary schools. school leadership & management, 32(1), 91-102. doi:10.1080/13632434.2011.642350 tynjälä, p. (2008). perspectives into learning at the workplace. educational research review, 3, 130-154. doi:10.1016/j.edurev.2007.12.001 van duijn, m., & vermunt, j. k. (2006). what is special about social network analysis? methodology, 2(1), 2-6. doi:10.1027/1614-1881.2.1.2 van der rijt, j., van den bossche, p., van de wiel, m. w. j., de maeyer, s., gijselaers, w. h., & segers, m. s. r. (2013). asking for help: a relational perspective on help seeking in the workplace. vocations and learning, 6, 259-279. doi:10.1007/s12186-012-9095-8 van der rijt, j., van den bossche, p., van de wiel, m. w. j., segers, m., & gijselaers, w. h. (2012). the role of individual and organizational characteristics in feedback-seeking behaviour in the initial career stage. human research development international, 15(3), 283-301. doi:10.1080/13678868.2012.689216 van knippenberg, d., de dreu, k. k. w., & homan, a. c. (2004). work group diversity and group performance: an integrative model and research agenda. journal of applied psychology, 89(6), 10081022. doi:10.1037/0021-9010.89.6.1008 van waes, s., van den bossche, p., moolenaar, n. m., de maeyer, s., & van petegem, p. (2015). knowwho? linking faculty's networks to stages of instructional development. higher education, 70(5), 807826. doi:10.1007/s10734-015-9868-8 geeraerts  et  al       | f l r     98   wasserman, s., & faust, k. (1994). social network analysis: methods and applications. new york: cambridge university press. williams, k. y., & o'reilly, c. a. (1998). demography and diversity in organizations: a review of 40 years of research. research in organizational behavior, 20, 77-140. wolff, c. e., van den bogert, n., jarodzka, h., & boshuizen, h. p. a. (2015). keeping an eye on learning: differences between expert and novice teachers' representations of classroom management events. journal of teacher education, 66(1), 68-85. doi:10.1177/0022487114549810 zijlstra, b. j. h., & van duijn, m. a. j. (2003). manual p2. version 2.0.0.7. groningen: iec progamma/university of groningen. zijlstra, b. j. h., van duijn, m. a. j., & snijders, t. a. b. (2006). the multilevel p2 model: a random effects model for the analysis of multiple social networks. methodology, 2(1), 42-47. doi:10.1027/16141881.2.1.42 codepen 5minnameierpub frontline learning research special issue vol 8, no. 5 (2020) 70 91 issn 2295-3159 explaining happy victimizing in adulthood – a cognitive and economic approach gerhard minnameiera agoethe university frankfurt am main, germany article received 1 june 2018 / revised 21 november 2019 / accepted 21 november / available online 1 july 2020 abstract while acknowledging the phenomenon of “happy victimizing” (hv), the classical explanation is questioned and challenged. hv is typically explained by a lack of moral motivation (mm) that is thought to develop in late childhood and adolescence. apart from empirical evidence for widespread hv in adulthood, there are also strong theoretical arguments against the classical explanation. firstly, there are arguments against the coherence of the very concept of mm. secondly, while the classical explanation focuses on internal drivers (in the sense of mm), the one proposed in the present paper focuses on the patterns of interaction. accordingly, hv may depend less on internalised values and individual motivation (whether in terms of moral internalism or moral externalism), and more on the “rules of the game” that are established in social interaction (or not). on this account, hv appears where higher order moral rules are not established and cannot be established, either due the circumstances or due to the unwillingness (or incapability) to play by the rules of these higher order games (where “games” are to be understood in the game-theoretic sense). the ordinary one-shot prisoners’ dilemma is a case in point. it precludes promise-giving as well as other higher order moral regimes, but instead forces the agents into a conflict of interest, where everyone has to mind their own business. moreover, claiming that all players have to pursue their own self-interest, can be understood as moral rule of its own. keywords: happy victimizer phenomenon; moral stages; moral reasoning; moral motivation; moral internalism and externalism; game theory; norms and conventions info corresponding email: minnameier@econ.uni-frankfurt.de doi: https://doi.org/10.14786/flr.v8i5.381 1. introduction it is well known and empirically well established that 4 to 6 years old children have a marked propensity for the so-called happy victimizer pattern (hvp). this is the combination of violating a known moral rule and feeling good about it (nunner-winkler & sodian 1988; arsenio, gold, & adams, 2006). while older children have more often mixed feelings (bad about acting immorally, good about what they get by doing so), the younger ones do not seem to feel any remorse, even though they know and accept the moral rules they violate. in the first place, hvp was discovered using a projective method, where the children had to look at a series of pictures that illustrated a case of immoral behaviour. then they were asked whether the behaviour of the protagonist was ok or not. finally they were asked how he or she felt and why. later on, the study was replicated asking the children directly what they would think and feel, if they were in the protagonist’s shoes (keller, lourenço, malti, & saalbach, 2003). even in these circumstances, about 50 per cent of the participants showed hvp. the proposed explanation was straightforward. since children obviously understand and seem to have internalised the moral norms, they are said to have the moral knowledge necessary for moral action, but lack moral motivation (nunner-winkler & sodian, 1988; nunner-winkler 2007; 2013; arsenio et al., 2006; krettenauer, malti & sokol, 2008). moral motivation has even been thought to be the crucial ingredient that developed only gradually in late childhood and adolescence and prevented many from doing the right thing (nunner-winkler 1999; 2007; malti & krettenauer, 2013). however, what is moral motivation? even though the concept is old, it is still difficult to grasp (brink, 1997; smith, 1994/2005; zangwill, 2003; minnameier, 2010; malti & krettenauer, 2013; wren 2013; heinrichs, minnameier, gutzwiller-helfenfinger, & latzko, 2015; rosati, 2016). furthermore, concerning its development, the proponents of moral motivation remain mostly silent. while nunner-winkler has thought that it develops gradually throughout childhood and adolescence (see above), others think it is there right from the start (malti & krettenauer, 2013, referring to warneken & tomasello, 2009, who have found strong evidence for pro-social behaviour among toddlers). what we know, however, is that hvp has turned out to be salient even among young adults and contrary to the classical explanation, according to which hvp should vanish in late childhood (nunner-winkler 2007; krettenauer, malti & sokol, 2008; minnameier & schmidt, 2013; heinrichs et al. 2015). beyond the narrow frame of research on hvp, there is also huge evidence for happy-go-lucky cheating and other forms of moral victimisation among adults (batson et al., 1999; 2002; ariely, 2012, rustichini, & villeval, 2014). in particular, it is well known that students of economics and business administration act more selfishly in various experiments than students of other subjects (marwell & ames, 1981; frey, 1986; carter & irons, 1991; frank et al., 1993; 1996; frey & meier, 2003; rubinstein 2006). further analyses have mainly focused on whether this points to a self-selection of self-interested individuals into economics-related studies or to an “indoctrination effect”, which means that economics is taught in such a way that makes students more selfish and competitive. the differentiation was introduced by carter and irons (1991) who found evidence in favour of the self-selection hypothesis, based on the ultimatum game. however, selten and ockenfels (1998) as well as frank et al. (1993) found evidence for an indoctrination effect. within this economic body of research, however, it has never been asked how people actually take their decisions and what is really appropriate in specific circumstances. after all, we know that people use moral principles in a situation-specific manner (krebs & denton, 2005; rai & fiske, 2011; minnameier, beck, heinrichs, & parche-kawik, 1999; minnameier & schmidt, 2013). hence, the question arises how economists and non-economists actually take their decisions and what is (more or less) appropriate. by the same token, hvp in general might be explainable as an action pattern that is morally justified (or at least justifiable) under specific conditions. thus, while the phenomena are almost crystal clear, the explanation is not. in section 2.1, i will summarise arguments against the received explanation from moral motivation that i have elaborated elsewhere in detail (minnameier 2010; see also 2012; 2013). on this account, the common understanding of moral motivation has to be turned around almost completely. what i suggest instead in the remainder of section 2 is to replace the false dichotomy of moral cognition and moral motivation by a more comprehensive theory of moral judgement and agency that includes processes of abduction, deduction and induction. this approach converges not only with the integrative account of rai and fiske (2011), but also with a “reason-based theory of rational choice” (dietrich & list, 2013a and b) that allows us to integrate cognitive moral psychology with rational choice theory (section 2.2). as it turns out, however, this theory of moral decision-taking has to be integrated into a game-theoretic view, because moral principles represent more than just personal values (section 2.3). in section 3 a study is presented that contains important evidence in favour of the reason-based approach, and section 4 contains an extensive discussion that also introduces further theoretical ramifications. section 5 concludes. 2. moral motivation and beyond 2.1. the problem of moral motivation the idea of moral motivation has different sources. for instance, kant needed it to explain why an individual ought not only to follow moral principles, but to engage in moral reasoning and develop moral principles in the first place (ameriks, 2006). this is why kant says that “(t)here is nothing it is possible to think of anywhere in the world, or indeed anything at all outside it, that can be held to be good without limitation, excepting only a good will” (2002/1785, p. 9 [4.393] ). in modern philosophy and psychology, the concept has been used to explain hvp-like behaviour (in terms of a lack of moral motivation). philippa foot (1972) kicked off the modern debate in philosophy observing that we can very well be indifferent to morality without being irrational (zangwill, 2003). and in moral psychology it was augusto blasi (1984) and james rest (1984) who held against kohlberg that moral reasons did not motivate moral action directly, but needed to be sided by judgements of responsibility (blasi) or moral motivation (rest). both strands of reasoning, the one in moral philosophy and the one in moral psychology, are directed against a view traditionally ascribed to socrates. socrates is said to have claimed that knowing what morality (or virtue) demands motivates virtuous action and that acting otherwise indicates the agent’s ignorance about the moral course of action (see e.g., brickhouse & smith, 2010, esp. chap. 3). advocates of moral externalism hold that one can very well know what the moral course of action would be, yet not be motivated and therefore act in a different way. moreover, they do not interpret this as a mere case of weakness of the will, in which agents would act against their own intentions and possibly be racked by remorse as a consequence. their view is that a moral judgement has to combine with a desire, where both are related only contingently. therefore, on the externalist account, moral judgement does not motivate by itself, but needs to be seconded by moral motivation (see e.g., rosati, 2016). this is precisely the way in which james rest, in the psychological camp, has defined moral motivation, i.e., to “select among competing value outcomes of ideals the one to act on; deciding whether or not to try to fulfil one’s moral ideal” (1984, 27). and he differentiates it from willpower, which means “to execute and implement what one intends to do” (ibid.). this is, by and large, the current state of affairs in moral psychology concerning the concept of moral motivation (see e.g. thoma & bebeau, 2013; nunner-winkler, 2013). against this view, i have argued elsewhere that this conceptual framework is inconsistent for two reasons (see minnameier, 2010; 2013). the first is that rest’s definition of moral motivation implies a judgment: if individuals are not driven against their will (which would indicate lack of willpower rather than lack of moral motivation), but select freely among competing value outcomes, this kind of selection requires a decision based on some criterion. in other words, moral motivation would not be independent from moral judgement, but would have to involve some kind of moral judgement resulting in an intention that motivates action. in fact, deciding whether to go for some personal benefit or to dispense with it for some other-regarding motive is even the paradigmatic case of a moral problem as such. therefore, the classical definition of moral motivation seems ill-conceived? the second point pertains to what makes moral motivation moral. how could moral motivation possibly be distinguished from other kinds of motivation, if not for an underlying reason or so? i see no other way to explain the moral aspect of moral motivation than to refer to some underlying reason. hence, moral motivation would have to be derived from the conviction of being morally obliged in some way. 2.2. inferential moral reasoning and a reason-based theory of rational choice on the cognitive view, moral motivation in the sense of rest’s third component is not just eliminated, but rather replaced by a specific kind of judgement within a more comprehensive notion of moral judgement. this conception allows for three different parts of moral reasoning which play distinctive roles in the formation of a moral intention. in particular, this broader notion of moral reasoning comprises abduction, deduction, and induction, which in turn are based on the charles s. peirce’s pragmatist theory of inferential reasoning (see minnameier, 2004; 2017). according to this approach, both the adoption and the application of moral principles are mediated by three characteristic inferences called “abduction”, “deduction” and “induction”. any kind of development is triggered by a negative, i.e., disconfirming, induction, which means that a certain principle that one previously adhered to fails to do the job in the given situation. for instance, if a child uses the golden rule (“do unto others as you would have others do unto you”) and now faces a decision at school about where to go for a class outing, this rule might not suffice. for even if everybody followed this principle, no unitary decision might be obtained. hence, instead of trying to square diverging individual interests, a decision has to be taken at the level of the group. majority voting would be a solution for this kind of problem, and this solution allows us to transcend individual interests in order to determine what is best for the group as a whole. moreover, group decisions call for loyalty on the part of the outvoted members. the shift onto the higher stage can be understood as being mediated, firstly, by a negative induction which leads to the insight that the golden rule fails to solve the problem at hand. secondly, a new principle has to be invented, which is captured by an abductive inference. thirdly, deduction tells us what follows, if the principle is applied to the situation at hand, in particular what kind of action would have to be taken. fourth, and finally, induction leads to the adoption (or rejection) of that principle as a guideline for the present situation, but equally for all situations of the same type. these inferential processes can also be assumed to mediate moral action, where moral principles are not invented ab ovo in a genuine developmental transformation, but merely activated in relevant situations (see minnameier, 2013). in those cases, the inferential reasoning may take an explicit or implicit (habitual) form, especially in well-rehearsed situations, for which habitual action schemes have been established. in the latter case, we may immediately know how to react without engaging in any (further) reasoning, but would have to be able to explicate our reasons on demand. the inferential approach meshes perfectly with a reason-based theory of rational choice (rbt) (dietrich & list, 2013a and b; minnameier 2016a). dietrich and list strongly criticise the behaviourist approach that is still prominent in economics today (see e.g. gul & pesendorfer, 2008) and which relies on the concept of “revealed preferences” (samuelson, 1938; 1948). this means the preferences are revealed from choice data and rationality assumptions like completeness and transitivity, just to name the most important. if you choose apples, where you could have had pears for the same price, this reveals that you prefer apples over pears. however, our “real” preferences are usually of a deeper psychologic nature, so that we might buy a bunch of red roses to impress a beloved person, not because we find them particularly attractive and worth the price for what they are as such. therefore, dietrich and list formulate a rational choice theory in which motivating reasons (i.e., the underlying drivers which need not necessarily be made conscious) play a central role. we could also call them the “fundamental preferences” that express what we are really striving for, as opposed to the concrete or instrumental preferences, like the one for red roses in the example. the latter are formed based on the former and the specific restrictions that hold in a given situation (like time, money, and so forth). if we take moral principles as fundamental preferences or “reasons” in terms of rbt, then the situation-specific adaptation of morality is theoretically tractable within a rational-choice-theoretic context. the situational cues provide information about the positive and negative restrictions (affordances and constraints), which we would have to understand as “beliefs”, because what counts for the explanation of behaviour is how the agent views the situation, not an objective account of it. the fundamental preferences and the beliefs determine the concrete preferences that we could also call “intentions” in the context of morality. this in turn, would determine a rational action. however, people might fail to act according to their own intentions, especially for lack of willpower, which allows us to include “irrational choice” as well. in the inferential context, we can say, accordingly, that a moral problem perceived as a feature of the situation abductively leads to a moral principle that captures this problem. from this principle and the situational premises we can deductively derive one or several moral courses of action, which just follow from applying the principle to the situation. the last, inductive, step is to determine whether this solution of the moral problem can be appropriately implemented. sometimes, the price one has to pay, or the risk, might be too high, so that we might legitimately discard a certain course of action at this stage. however, if there are no such problems, we commit ourselves and so a moral intention is formed. this would then have to be carried out on pain of irrationality. finally, i think this reconstruction is valid also for intuitive moral agency, because if we do not assume any kind of moral reason or principle to underlie the act, we could never distinguish self-interested action from moral action. therefore, even an intuitive helping behaviour that ought to be characterised as “moral” must imply a moral point of view (e.g., that the other person is in need of help and therefore should be helped). consequently, some kind of moral cognition would always have to be abducted, even in non-deliberative moral decision making (see minnameier, 2016b; 2017; hermkes, 2016). 2.3. morality in the game-theoretical context rational choice theory branches out into decision theory and game theory (see e.g., binmore, 2009, p. 25). the former relates to the choices a rational agent takes in a certain environment (“games against nature”), whereas the latter concerns the interaction with one or more other rational agents. on the reason-based account, moral agency is modelled within a decision-theoretic framework, and moral values are treated as personal values (fundamental preferences) according to which the individual wants to live. however, from such a point of view, moral judgements are (involuntarily) reduced to prudential judgements, and so is moral action reduced to prudential action, because whatever an agent does, it is always explained in terms of maximising utility with respect to satisfying fundamental (but still personal) preferences (whether we call them morals, values or virtues). this raises the question of what morality really is. this question is very topical also in the context of prosocial behaviour among toddlers and apes (for reviews see paulus, 2014; killen & smetana, 2015). it appears as if they have a sense of morality or justice. however, it may be a mistake to infer directly from prosociality to morality. first of all, agents may have altruistic or other-regarding orientations, but these would none the less be their own orientations. they might be explained in terms of vicarious experience caused mirror-neuron activity. however, in whichever way these orientations are explained, the hard problem remains. the hard problem is that stretching morality to include mere prosociality would make the distinction between morality and prudence obsolete or impossible. conversely, if (first person) prosociality is distinguished from (third person) morality, the morality requires a third-person perspective (or meta-perspective) from which the differentiated perspectives of self and other(s) are looked upon and coordinated in some way. this idea of a hard problem can even be taken further. we typically conceive morality in terms of (other-regarding) personal values. people who have internalised such values are regarded as highly moral people in the sense that they are intrinsically motivated to act morally. conversely, those who are only extrinsically motivated appear to use morality instrumentally and are therefore considered not to be morally motivated. in other words, they are not moral agents, but only do as if they were moral for some selfish reasons. again, they seem to be guided by a prudential rationality rather than a moral rationality. however, what appears as truly moral on this account, still runs into the hard problem, because if agents follow their personal values and try to maximise utility in this respect, they will always have to follow a prudential rationality, whatever the specific content of their values and however other-regarding they may be. this hard problem remains, because morality is reduced to a “decision-theoretic” problem rather than a “game-theoretic” one. however, morality is basically a social project, not a question of the good life for the individual, and hence not just a question of personal values. if moral principles are to make any sense, they do not only address the agent as an individual, but all agents involved in a moral problem. moral rules are social rules, and these rules have to be accepted and heeded by the set of agents to whom the rules apply. a further common misunderstanding or misconception is this: if we think of morality in the game-theoretic sense, we often treat moral issues in terms of a zero-sum game. a zero-sum game is one in which a fixed sum of money or amount of goods is to be allocated. the overall sum does neither increase nor decrease. in this very sense we think that rich people should transfer some of their riches to the poor, thereby leaving the total unchanged. on this account, those who give suffer a loss (for moral reasons), and the needy enjoy the benefit. however, apart from zero-sum games we have two other types, i.e., coordination games and cooperation games. coordination games are less problematic, because coordination is always beneficial for every agent. for instance, in some countries people drive on the right-hand side, in some they drive on the left-hand side. none the less there is no reason, say, for those from the continent to drive on the left while in the united kingdom, and those from the uk would have no incentive to drive on the right while on the continent (unless they wanted to commit suicide). where we drive is a question of “conventions”, because one just has to follow a uniform rule, and there seems to be no deeper reason, at least no moral one, why we should decide to drive on the right or on the left. thus, conventions are the solution concept for coordination games, and i think domain theorists like turiel (2002) and nucci (2008) could benefit from consulting game theory to distinguish conventional problems from moral problems. cooperation games are different. the prisoners’ dilemma is a classic example. in this type of situation, “cooperation” would benefit everyone (like in a coordination game), but it does not constitute a stable equilibrium in the sense of a nash equilibrium. in the prisoners’ dilemma the agents have an incentive to defect, which takes them to a social trap in the form of a pareto-inferior nash equilibrium. their only way to overcome this deadlock is to invent a so-called institution, i.e., a rule backed by suitable sanctioning mechanisms that ensures cooperation. if the prisoners’ dilemma were not as restrictive as it is (where the agents are not allowed to communicate with each other), they would perhaps agree not to confess to the judge and threaten each other with the punishment of comrades outside the prison or with their own fierce revenge in case the other might defect. however, such an institution is tantamount to a moral rule, in particular that of a mutual promise or contract. and the potential sanctions may not only be negative, but also positive (respect for one’s reliability and a mutual willingness to cooperate in the future). hence, moral principles can be straightforwardly reconstructed as solution concepts for so-called cooperation problems? i think this hits the point, and i also think that both the decision-theoretic interpretation as well as the understanding of morality in terms of a zero-sum logic are severely flawed. these misconceptions are at the heart of fatal errors about moral motivation and the related normative questions (see e.g. minnameier, 2013, and the discussion of empirical results below). 3. empirical evidence of hvp in the light of the reason-based approach 3.1. review of a study on morality in the economic context here, i would like to summarise the data and results of a recent study on how economists and non-economists choose to act in different framings of the famous prisoners’ dilemma (minnameier, heinrichs, & kirschbaum, 2016), which also yields important insights about how prevalent hvp is among different groups and in different contexts, and on how it might be brought about. this is essential for the crucial question of whether those who show the pattern are really motivated immorally or not, and whether the self-interested behaviour is to be morally condemned and educationally tackled or is rather acceptable or even desirable as an adequate situation-specific response to the constraints the individual faces. as mentioned in the introduction, there is clear evidence that students of economics and business administration behave more selfishly than others, at least in certain situations. and by the same token we may assume that hvp is more prevalent among these students. however, not so much is known about the motivation of either self-interested or other-regarding behaviour. moral psychologists tend to ask why people act selfishly and fail to follow moral rules. economists wonder about other-regarding behaviour which they consider irrational. thus, there seems to be a need for clarification. the prisoners’ dilemma is a case in point. it has been used to measure moral motivation (see nunner-winkler, 2007), where cooperation indicates high moral motivation and defection low moral motivation. on this view, however, the moral course of action is the exact opposite of rational action in terms of an economic analysis. conversely, defection, which is the dominant strategy and thus rational from an economic point of view, is thought to indicate a lack of moral motivation. hence, you could either be “rational” or “moral”, but not both. luckily, the reason-based approach sketched above allows us to reconcile morality and rationality. based on it, we may assume that even though economists might generally be more self-interested than students of other subject matters (especially following the self-selection hypothesis), not all those who take self-interested decisions may do it for lack of moral motivation, but may simply be more realistic and prudent in their decision-making. to test this assumption, we have carried out a study using the prisoners’ dilemma (pd) in the two framings already discussed above, i.e., we called it “wall street game” in one condition and “community game” in the other. additionally, we let participants express their feelings and reasons. our hypotheses in this context were the following (we use “economists” as a short cut for “students of economics and business administration”, and “b&e education” for “business and economics education”): compared with other students, economists (a) defect to a greater extent and (b) are less vulnerable to the framing (since they know they are in a pd and that defection is the dominant strategy). accordingly, we expect (a) more happy victimizers than happy moralists among economists and (b) fewer happy victimizers than happy moralists among the students of other subjects. we expect that at least in part these differences do not indicate differences in moral judgement and/or moral motivation, but merely different views of the situational constraints. since we have also students of b&e education in our sample, we expect them to score at intermediate levels as compared with the other two groups. 3.2. sample and design the sample consists of 481 undergraduates from two german universities (frankfurt am main and bamberg). 54 percent of the sample are bachelor students of economics and business administration (“economists”), 14 percent are bachelor students of b&e education, and 32 percent are teacher students with no economics-related subject (henceforth “teacher students”). 39 percent of our participants are male, 61 percent are female. the average age is 22.4 years. the participants were presented a classical pd in two different frames. for one half it was called “wall street game” (wg), for the other half it was called “community game” (cg) (see liberman et al., 2004; ellingsen et al., 2012). they were randomly assigned to one of the two conditions. of the 481 participants, 227 were in the cg-condition, 254 in the wg-condition (χ^2 = 1.516, p = 0.218). the description and instructions were exactly the same in both conditions with the only exception that the respective name of the game appeared three times in the instructions. they had to decide between two options, a and b, and were given the following information: if both you and the other person choose a you both get €50. if both you and the other person choose b you both get €20. if you choose a and the other person chooses b you get €5 and the other person gets €80. if you choose b and the other person chooses a you get €80 and the other person gets €5. after having chosen one of the two possible strategies – which are commonly called “cooperate” (a) and “defect” (b) – they had to explain, first, why they decided the way they did. after this, they had to rate how they feel about the decision on a four-point likert-scale (good, rather good, rather bad, bad) and, finally, they had to explain their feelings. the data on the participants’ emotions allowed us to code the answers in the happy-victimizer framework (arsenio, gold & adams, 2006; nunner-winkler, 2007; 2013), where good or rather good feelings indicate “happiness” and bad or rather bad feelings indicate “unhappiness”. concerning the choice options, cooperation codes for “moralists” and defection for “victimizers”. thus we classify the answers as happy (hv) or unhappy victimizers (uv) and happy (hm) or unhappy moralists (um). 3.3 results as for the framing in terms of wg and cg, no framing effect can be identified in the total sample (χ^2 = 0.112, df = 1, n = 481, p = 0.783). the same is true for economists (χ^2 = 0.007, df = 1, p = 1) and b&e education (χ^2 = 0.122, df = 1, p = 0.804), but a rather strong one among the non-economist teacher students (χ^2 = 5.374, df = 1, p = 0.029), with 80 percent cooperation in the cg-condition and only 63 percent in the wg-condition (see table 1). economists only cooperate by 45 percent (cg) and 44 percent (wg) respectively. this is not surprising, since economists know the prisoners’ dilemma and have identified it as such. apart from the framing, however, the defection rate of economists is much higher, in general, than that of non-economists. students of b&e education score at intermediate levels. table 1 framing, cooperation and defection in the pd as table 1 also reveals, the study programme has a significant effect on the decisions taken (χ^2 = 24.799, df = 2, n = 476, p = 0.000). this effect remains, if we control for gender (male: χ^2 = 10.988, df = 2, n = 187, p = 0.002 (one-tailed); female: χ^2 = 13.208, df = 2, n = 289, p < 0.001, one-tailed). thus, hypotheses 1a and 1b are both strongly confirmed. and concerning hypotheses 4 we can state that students of b&e education are at an intermediate level with respect to the proportion of cooperation and defection. perhaps even more importantly, we see a major difference in terms of agency (hv/uv/hm/um), with 55 percent hvs among economists, but only 31 percent among non-economists. conversely, 67 percent of the non-economists are hms, whereas only 36 percent of the economists are hms (see figure 1; χ^2 = 26.094, df = 6, n = 337, p = 0.000, based on fisher’s exact test). if we control for gender, results are still statistically significant (male: χ^2 = 11.016, df = 6, n = 130, p = 0.044 (one-tailed); female: χ^2 = 12.052, df = 6, n = 204, p = 0.026, one-tailed). hypothesis 2 is confirmed by these data, and so is hypothesis 4. figure 1: the proportions of hv, uv, hm, and um. another important result relates to the reasons the participants give for their choices and feelings. we classified them in terms of neo-kohlbergian stages, that are explained below in the discussion. three such types can be distinguished (see figure 2): 1. those who ignore or fail to see the conflict of interest inherent in the pd and just do what they think is best for all, i.e., to cooperate (stage 1c). 2. those who see the conflict of interest and argue that they just have to pursue theirs, which is to defect (stage 2a). 3. those who see the conflict of interest and wish to overcome it. they decide to cooperate, but also express that they would be ready to defect, should their partner be unwilling to cooperate (stage 2c). figure 2: types of moral reasoning. with respect to the third type, there is no difference between the groups. therefore, the overall difference in moral orientations is down to a trade-off between the other two types. economists have a comparatively strong tendency to argue that as an agent in this game they have to pursue their personal interest, whereas teacher students show a strong orientation towards what would be collectively rational (χ^2 = 29.548, df = 4, n = 338, p = 0.000). if we control for gender, however, we only get significant results for female participants (χ^2 = 17.653, df = 4, n = 205, p < 0.001; male: χ^2 = 7.648, df = 4, n = 130, p = 0.053. one-tailed). among economists, 60.2 percent are type 2, and 16.1 percent are type 1. conversely, among teacher students we find 32.2 percent type 2 and 44.8 percent type 1 participants. the fact that there are no differences with respect to the more sophisticated third type is at least an indication that the participants might not differ in their moral preferences, but instead in their relevant beliefs. the strong defective orientation, coupled with type-2-reasoning that economists reveal, may be due to the fact that they are highly aware of the situational constraint the pd poses. conversely, the teacher students seem to be oblivious to this very fact. they seem to focus on the common interest, ignoring that the pd models conflicting interests. in other words, the common interest should be to overcome this problem of conflicting interests, which is precisely what type-3-participants aim at. however, type-1-participants are unable (or unwilling) to see this basic problem. in the context of the differentiation between preferences and restrictions this means that they, as a matter of fact, conceive of quite different restrictions than participants of type 2 or type 3. inasmuch as the differences in agency derive from differences in beliefs rather than differences in basic moral preferences, the different groups might differ neither in their moral judgement competence nor in what might be called their moral motivation, but in their comprehension of situational constraints. a further analysis, illustrated in figure 3 supports this view. here we only look at hvs and the reasons they give for their choices and feelings. some of them explicitly refer to the situational constraints and state that cooperation would have been better, but in the situation as it was, they just had to choose the other option. they clearly would have preferred to cooperate. they are not unhappy, however, because they think they have taken the right decision. in quite some participants we could identify this kind of reasoning, and we call these “strategic moralists”, because they are morally motivated, but also focus on the prudential aspect of what kind of morality can be implemented in the present situation. figure 3: strategic moralists and happy victimizers in the strict sense. again, we see more strategic moralists among economists than among teacher students. this does not mean that economists are, generally speaking, not more selfish than others. hypothesis 3 only states that the stronger tendency towards an hv-like agency on the part of the economists is not only an effect of a higher level of selfishness, but in part down to different beliefs. this is confirmed by the data. furthermore, students of b&e education score on intermediate levels also in this respect. hypothesis 4 is thus confirmed throughout. figure 3: strategic moralists and happy victimizers in the strict sense. again, we see more strategic moralists among economists than among teacher students. this does not mean that economists are, generally speaking, not more selfish than others. hypothesis 3 only states that the stronger tendency towards an hv-like agency on the part of the economists is not only an effect of a higher level of selfishness, but in part down to different beliefs. this is confirmed by the data. furthermore, students of b&e education score on intermediate levels also in this respect. hypothesis 4 is thus confirmed throughout. 4. discussion: moral principles as social institutions 4.1. morality in the prisoners’ dilemma above, i have made a claim for a game-theoretic understanding of morality, where moral rules have to be rules of a game that the agents play. if they are to be rules in the game-theoretic sense, however, they have to be self-enforcing (binmore, 2010). that is, following moral rules must be somehow pay – especially in “moral currencies” like respect, reputation and the like – and violating them must translate into costs, so that not complying with the rules is clearly irrational (at least for those who understand the rules). put in other words, a simple appeal to moral rules and precincts is worth nothing, if those rules cannot secure compliance by themselves (and if this fault is not compensated by some other rules). this is the case in the (one-shot) pd, where a morality of contract or the golden rule cannot be implemented, simply because these moral rules are either impossible (striking a deal) or not enforceable (golden rule). the golden rule is not enforceable, because no signals of approval or disproval, or even of the appropriateness of the golden rule, can be exchanged, which means that this “moral currency” is invalid in this case. this explains why such moral orientations are usually crowded out very quickly (for the remaining levels of cooperation see footnote 9). however, this is not the whole story. if we take moral rules as rules of the game, we can make even more sense of the strategies that individuals choose in the game. first of all, even where individuals defect, they may still follow a moral rule, the rule that everyone has interests to pursue. this is one variant of kohlberg’s stage 2, which concerns the “awareness that each person has interests to pursue and that these may conflict” (colby & kohlberg, 1987, p. 26). defectors, whom we have classified as hvs, do not betray morality, but follow a specific morality that acknowledges the dignity of each person’s individual interest and that these interests may conflict. this is the “type 2” morality specified above. in such cases of conflict of interest (and in the absence of any means to mediate between these conflicting interests), agents are morally justified to pursue their interests, but also have to accept others doing so. this morality is employed by both hvs and sms, where the difference is that sms explicitly state that a higher order morality would not work. secondly, “type 1” morality seems naïve, because it ignores the conflict of interest and rather assumes that each individual’s interest in the other person’s well-being is strong enough to preclude defection. this is a kind of morality that would work among closely affiliated agents (e.g., family members and friends are typically supposed to care for each other and dispense with a personal profit or benefit to support the other). their moral appeals for solidarity would therefore fall flat. thirdly, “type 3” morality is employed by those who see a chance to coordinate with others, but this higher form of morality is almost certain to be crowded out, since individuals frequently express they would switch to defection, if their partners fail to cooperate. fourthly, the situation would be different, if we allowed for changes in the game. for instance, the pd could be played repeatedly. in this case, a tit-for-tat strategy is possible. if player 2 understands the intention of player 1’s choices, cooperation can easily emerge in this repeated interaction. player 1’s cooperation in the first round can be rewarded by player 2’s cooperation in the second round, and vice versa. conversely, player 1’s defection in one round can be punished by player 2’s defection in the following round. every cooperative choice can be understood as a promise to cooperate in the future, provided that the other player follows suit. hence, under these circumstances the morality of promise-giving can be implemented und generate benefits for both 4.2. the economics of morality in the game-theoretic context finally, we can generalise the kind of moral functioning just explained with respect to the pd (see also minnameier, 2018). the pd is one example of what game-theorists call a “cooperation game”. perhaps counterintuitively, cooperation games are the situations in which cooperation typically fails if no institution is established, because the pareto-efficient point is not a nash-equilibrium. this formal structure of cooperation games is illustrated in figure 4. figure 4: the basic structure of cooperation games. the simplest version of such a game is the situation that thomas hobbes models in the leviathan: based on the “law of nature” according to which “every man has right to everything” (1651/2001, p. 65 [chap. 15, §2]) a “war of every one against every one“ (1651/2001, p. 59 [chap. 14, §4]) is thought to ensue. this marks the very beginning of morality, where, e.g., children get into conflict about some resources like food or toys and they compete in trying to appropriate things. what they have to learn is to mutually respect property or rights of use (e.g. that the one who had it first has the right to use a certain item). however, before they learn this, they have to experience the social trap (i.e., the inefficient nash equilibrium) they reach in the bellum omnium contra omnes, for this is the problem that calls for innovation. at the same time, this problem allows them to take the perspective of the other individual who has the opposite point of view. both “players” understand that when they win something, the other one loses it (win-lose), and vice versa (lose-win). and eventually they learn that their conjoint activity produces a lose-lose-result. establishing the moral norm allows them to move into the win-win-zone. since it applies to a cooperation game, a moral norm – as an “institution” – has to go with the possibility of sanctions. in the simplest case of morality in narrow social relationships, the agents sympathise with each other, want to make and keep friends and to win each other’s affections. therefore, signs of fondness and attachment, like smiles and so on, function as positive sanctions, whereas repudiation and anger function as negative sanctions. if the sanctions work this way in a social relationship, the payoff matrix is changed accordingly (see figure 5). figure 5: payoff matrix with premiums (+3) and discounts (-3) for sanctions. if the institution, i.e., the moral rule, is understood by both players, defection is discounted by the costs for spoiling the personal relationship and incurring dislike. conversely, cooperation is enhanced by the returns in affection currency. the most important aspect, however, of this change is that the whole game changes decisively. it is no longer a cooperation game, but a so-called coordination game in which the pareto-efficient combination is also a nash-equilibrium. the main result of this analysis is, therefore, that moral rules allow us to turn cooperation games into coordination games, and in this very sense they become self-enforcing, as long as the players understand them and as long as the sanctions work. the latter explains, why morality frequently vanishes in situations characterised by anonymity and social distance (see e.g., dana, weber, & kuang, 2007; andreoni & bernheim, 2009). 4.3 moral stages and appropriate sanctions the economic rationale just sketched allows us to construct ever more complex kinds of social interaction and cooperation. this entails a succession of cooperation games that have to be turned into coordination games with the help of moral principles as the regulating institutions. another property of this succession of stages is their dialectic order. if mutual acceptance of property rights marks the beginning of moral reasoning, this means that the perspectives and legitimate claims of different individuals are treated equally, each in their own right. however, these perspectives are kept independent of each other, and the morality of property rights secures and underpins this independence. this principle ceases to fit, if individuals differ in their property rights. if one individual owns a certain item desired by the other(s), a sharing norm is what we need. therefore, we usually claim that one ought to share, at least with friends or with those who are dear to us. the sharing norm relates these perspective reciprocally to each other. however, there is a simple form of sharing, where those who have something have to give to others and have to share in equal parts, usually. and there is an advanced form of sharing, where relevant inter-individual differences (of deservingness in terms of need, effort, and so on) have to be taken into account. in this latter case, equal sharing is considered unjust, and a norm of equitable sharing has to be put in place. this basic triad of types of morality follows from a dialectical approach advocated by the mature piaget in one of his last works (piaget & garcia, 1989), where he calls them “intra”, “inter”, and “trans”, respectively. empirical evidence is available today, e.g., from experiments with young children between three and six years of age (paulus & moore, 2014). another tenet of piaget’s final version of the mind’s architecture is that the “trans”-stage forms a complex unity which however can break up into a set of trans-stages as a further morally relevant aspect enters the scene that cannot be integrated into the present trans-stage. and so a higher level with a new stage triad opens up. in this very sense, the early morality, which is based on sympathy and altruism, gives way to new and higher form of morality relating to the cases where helping someone else is futile and goes against one’s own interest. in early forms of morality, helping others is never in contradiction with one’s own interest, i.e., one enjoys helping others and caring for them. however, if you ought to care for a competitor in business or in sport, or where someone, say, wants you to go for a walk together, when you yourself prefer to stay at home, this marks a true clash of interests. according to kohlberg, the morality of conflicting interests is captured by his stage 2 (colby & kohlberg, 1987). and since kohlberg had also invented the idea of sub-stages “a” and “b” for every stage – that he later changed into “types” 11 – i label the first triad as stages 1a, 1b, and 1c, and the succeeding triad stage 2a, 2b, and 2c, and so on. in this so-called neo-kohlbergian framework we find 9 stages altogether (see minnameier, 2014). they cannot be explained in detail here, but the first three of them are briefly described in table 2, together with the sanctions that work in its respective context. 12 table 2 moral principles and sanctioning potentials for stages 1 to 3 since stage 1 is based on sympathy, sanctions are imposed in terms of affection (positive) and dislike (negative). this is the code in which moral discourse takes place at this stage (or the “moral currency” of stage 1). if moral discourse in terms of this stage – or any other stage – fails because of unwillingness or inability, one can always back out of “the game”, shift to a lower stage and thus play a different game. for instance, if others simply wouldn’t stop taking things away from me or using and possibly spoiling my property illicitly, i will have to fight back in some way. at stage 1 this always relates to the people with whom one wants or has to keep up (like parents, mates and so on). thus, the extended punishment at stage 1a means to shift to stage 0, i.e., to defend oneself by taking vengeance (this is the hobbesian state of nature). similarly, at stage 1b one can ultimately stop sharing, so that everybody sticks with what they have (which is the principle of stage 1a), and at stage 1c one can always revert to strict reciprocity rather than an equitable or caring interaction. at stage 2 the legitimacy of interests and the conflicts of interest that arise are the core problem. this stage applies to situations in which the agents are mutually disinterested, either with respect to the person of the other or with respect to a specific activity. thus, while siblings are generally quite interested in each other and are involved in a close relationship, they may none the less have diverging interests in terms of leisure time activities, and then it is legitimate for them to go their own sweet ways, as it were. accordingly, a conflict may arise, when one, e.g., wants to listen to loud music while the other has to prepare for an exam (the same of course applies to students sharing a flat). if no ways can be found in which each one can pursue their interests without encroaching upon the others’, one has to go separate ways. this is meant by “suspension” or “separation” as the extended punishment at stage 2a, where the agents are relegated to contexts, where they do not interfere anymore with each other and only deal with those with whom one gets on well and wants to affiliate. stages 2b and 2c provide forms of coordinating in conflicts of interest. mutual promises at stage 2b agents with different interests to strike deals so that everybody’s interest is furthered. any ordinary deal is an example. the pd is a situation in which the agents would like to strike such a deal to overcome the dilemma. but the rules of this game preclude this. therefore, agents have to revert to stage 2a and simply pursue their own interest (by defecting). in real life this can be a form of punishing those who go back on their promises. in the pd it is morally just to act selfishly, because the conflict of interest cannot be resolved. stage 2b requires that both parties have something to trade. stage 2c goes beyond this, because here one makes a contract with oneself, determining what you would have yourself do, if you were the other person. however, this requires trustful relationships and that the favours you offer are paid back in case the roles were reversed some other time when you are in need of assistance or so. in case of violation of this principle one could revert to 2b and demand immediate compensation for any service to be given. at stage 3a individual interests are merged into group interests. in a way, stage 2c implies that one tries to please everybody and coordinate diverging individual interests in this way. however, there are situations in which this is impossible, e.g., if one works for a company where pleasing customers of suppliers too much means to spoil the company’s business. at this point it is important to think in terms of social units like companies, departments, families and clans, or peer groups and teams (in sport or at work). stage 3a relates to these social units and the roles one takes on in these contexts. role-related reputation is what one can gain, and disrepute may be the price one has to pay, if one fails to fulfil one’s tasks. in the limiting case, one is excluded from the group – either literally, e.g., by being laid off from a firm, or in the sense that group cohesion breaks apart. in the latter case, one can still treat each other with respect and rely on each other in terms of stage 2c, but not in terms of role-related duties and commitments. stage 3b concerns inter-group relationships, with customs and generalised expectations as moral principles. examples are fairness rules in sport, social “conventions” like what kinds of behaviour are expected from superiors and subordinates, honest practices in commercial relations (e.g., whether a handshake is an obligation or not). where there are diverging views among honest people – i.e., “honest” in the sense of stage 3b – about what is decent and what not, an authority will have to decide, typically the leader or leading authority of group, which can be a political community, a company or any other kind of organisation. that there is a leader or an authority who knows (or must know) what is right and may legitimately decide, is the core of stage 3c. at this stage, the verdict of an authority is believed to be sacrosanct. at this point i end the description of moral stages. however, this is not to be mistaken as the endpoint. an authority as just described has to be legitimate, and what follows the acceptance of some kind of leadership is the question how we can determine whether certain laws, rules and forms of government are legitimate or not. this is the proper field of ethics that we would then enter. altogether, the neo-kohlbergian taxonomy of stages comprises 9 stages (minnameier, 2000; 2001; 2005). finally, turning back to the three stages just described and to the way the sanctions work, i would like to point out that there are always two kinds of punishment available (see table 2): one is punishment within an institution, which relates to the proper meaning of the moral principle and reclaims that it be heeded. if the other does not play according to these rules, one can still revert to a lower stage and consequently play another moral game at this lower level. this precludes self-exploitation in situations where certain moral rules cannot be implemented for some reason. 5. conclusion hvp does not seem to be a real problem. when people illegitimately act in selfish ways, well-functioning social systems should be able to counter such behaviour by appropriate sanctions. and if these sanctions fail to work effectively, we have to adapt our tools, in particular by playing different (lower-stage) moral games with other kinds of sanctions (as explained in section 5.3). conversely, however, we will not solve such problems by trying to increase moral motivation in the sense criticised in this article (least of all among moral transgressors). there are situations, where “happy victimizing” actually seems to be an appropriate behavioural orientation, in particular in strictly competitive situations, as when applying for a job or competing over a large order in business. in such situations, it is morally mandatory to pursue one’s self-interest (as long as one is playing fair). the same is true for situations in which contracts are either not fulfilled or not feasible (like e. g., in the pd). acknowledgements i wish to thank two anonymous reviewers who have read the paper very thoroughly and carefully. they have helped me to improve it and correct remaining errors. keypoints the classical (morally externalistic) explanation based on moral motivation has been criticised and confronted with an internalistic alternative that incorporates a theory of inferential reasoning and a reason-based theory of rational choice. the reason-based approach and moral functioning can also be interpreted in terms of game theory, where moral principles (or preferences) become preferences for games. empirical evidence supports the reason-based approach, which is then extended to an overall theory of moral principles as institutions in the institutional-economic sense. one main outcome of this analysis is that morality requires positive and negative sanctioning mechanisms that must be operative, if specific types of morality are to be upheld in specific contexts. this new approach also revives kohlbergian moral theory and leads to a neo-kohlbergian theory of moral reasoning and moral functioning. footnotes 1 as a world-wide citation standard, the two numbers refer to the volume and the page in the famous “akademie-textausgabe”. 2 brickhouse and smith show that, contrary to the received view, socratic moral psychology is not naïvely cognitivistic, but more in line with what it is claimed in the present contribution. 3 this is an induction, because it is not just a formal deductive judgement based on deductive premises, but determines a belief about what is feasible and effective in real world interactions. furthermore, this belief includes the insight that the golden rule not only fails in the present situation, but would equally fail in all situations of the same kind as the present one. in this sense, any induction, even if it actually only addresses a single situation, is always generalizable, in principle. and this applies to positive (confirming) as well as negative (disconfirming) inductions. 4 note that already in gary s. becker’s economic approach to human behaviour (1976; 1993), preferences are conceived as very fundamental. however, he uses rational choice theory to explain any kind of (stable) human behaviour, i.e., he uses the principle of rationality as an explanatory tool, so that clearly nothing remains as irrational, since what is irrational on this account, is simply not rationalisable. weakness of will, however, is part of an explanation, and then this weakness is interpreted as a constraint on the agent’s actions, which rationalises even choices that run counter to the agent’s intentions. 5 we have a total of 587 participants in the study. however, 78 of them study something other than the three groups discussed here of have failed to specify the study programme. 28 have not completed the questionnaire and were therefore excluded from the analysis. 6 while the proportion of male and female participants is fairly equilibrated for economists (50/50 percent) and b&e education (42/58 percent), it is not for teacher students (21/79 percent). this study does not focus on the gender aspect. none the less, we control for gender in the subsequent analyses of moral orientations. in terms of age the differences are small (economists: 22.42 years; b&e education: 24.83; teacher students: 21.21 years) but significant (according to the kruskal-wallis-test, since variances are heterogeneous). 7 in the analyses on gender differences, only 476 cases are included, because the remaining five have failed to indicate their gender. 8 since both kinds of reasoning complemented each other, we have taken them together as expressing their moral judgement. 9 if one player does not understand the rules, the game is not played (at least not with this player), because a game implies that the players know the rules. if they don’t, they may still play a game, but a different one. 10 of course, self-signalling is always possible, and this may be the mechanism that makes some people cooperate, even after having gained experience with the pd (ledyard, 1995). some take this as a strong moral self (blasi, 1984; 1995; bergman, 2002; 2004; krettenauer, 2013). however, since this is tantamount to self-exploitation, one is certainly not morally obliged to cooperate in the face of (a high risk of) others defecting, especially since such defection cannot be penalised in any way. hence, whether one should engage in one-tailes cooperation is a question of personal value and prudential reasoning, but not a moral issue in the strict sense, even though we often associate it with morality. 11 kohlberg first introduced these forms as „sub-stages“ (see e.g. 1984), but later treated them as mere „types“, because he noticed anomalies in the developmental sequence (colby & kohlberg, 1987). from the point of view of the neo-kohlbergian taxonomy, however, these anomalies are integrated and therefore constitute no systematic problem anymore. 12 neo-kohlbergian stages roughly conform to kohlberg’s original stages (at least with respect to stages 1 through 5). however, there are a few important differences with respect to particular substages. for instance, the “golden rule” integrates conflicts of interests and identified as stage 2c in the neo-kohlbergian framework, while kohlberg takes it as a form of stage 3. 13 kohlberg associated the golden rule with stage 3. however, here he seems mistaken, since the golden rule applies to balance individual interests, whereas – at least in the neo-kohlbergian framework – stage 3 is based on social units to which individual interests are merged. in this sense any stage 3 morality differs sharply from the golden rule, which, however, is integrated in this higher order of hierarchical complexity. 14some separate strictly between moral issues and conventional issues (most prominently turiel, 1983; 2002; nucci, 2008). with respect to the systematic differentiation between “norms” and “conventions” in game theory (see above), i fully endorse this. however, social conventions can also function as norms in the sense and in the contexts discussed here. for further discussions of conventions functioning as norms see bicchieri (2006) and sugden (2010, where the latter explicitly discusses turiel’s and nucci’s approach). references ameriks, k. (2006). kant and motivational externalism. in h. f. klemme, m kühn & d. schönecker (eds.), moralische motivation: kant und die alternativen (pp. 3-22), hamburg: meiner. andreoni j., & bernheim, d. b. (2009). social image and the 50–50 norm: a theoretical and experimental analysis of audience effects. econometrica, 77, 1607–1636. https://doi.org/10.3982/ecta7384 ariely, d. (2012). the (honest) truth about dishonesty: how we lie to everyone – especially ourselves. new york: harper collins. arsenio, w. f., gold, j., & adams, e. (2006). children’s conceptions and displays of moral emotion. in: m. killen/j. g. smetana (eds.). handbook of moral development (pp. 581-609). mahwah, nj: erlbaum. batson, c. d., thomson, e. r., & chen, h. (2002). moral hypocrisy: addressing some alternatives. journal of personality and social psychology, 83, 330-339. https://doi.org/10.1037/0022-3514.83.2.330 batson, c. d., thomson, e. r., seuferling, g., whitney, h, & strongman, j. a. (1999). moral hypocrisy: appearing moral to oneself without being so. journal of personality and social psychology, 77, 525-537. https://doi.org/10.1037/0022-3514.77.3.525 becker, g. s. (1976). the economic approach to human behavior. chicago, il: university of chicago press. becker, g. s. (1993). nobel lecture: the economic way of looking at behavior. journal of political economy, 101, 385–409. bergman, r. (2002). why be moral? a conceptual model from developmental psychology. human development, 45, 104-124. bergman, r. (2004). identity as motivation: toward a theory of the moral self. in d. k. lapsley & d. narvaez (eds.), moral development, self, and identity (pp. 21-46), mahwah, nj: lawrence erlbaum associates. bicchieri, c. (2006). the grammar of society: the nature and dynamics of social norms. cambridge, uk: cambridge university press. binmore, k. (2009). rational decisions. princeton: princeton university press. binmore, k. (2010). game theory and institutions. journal of comparative economics, 38, 245-252. https://doi.org/10.1016/j.jce.2010.07.003 blasi, a. (1984). moral identity: its role in moral functioning. in w. m. kurtinez & j. l. gewirtz (eds.), morality, moral behavior, and moral development (pp. 128-139). new york: wiley. blasi, a. (1995). moral understanding and the moral personality: the process of moral integration. in w. m. kurtinez & j. l. gewirtz (eds.), moral development: an introduction (pp. 229-253). boston: allyn and bacon. brickhouse, t. c., & smith, n. d. (2010 ). socratic moral psychology. cambridge: cambridge university press. brink, d. o. (1997). moral motivation. ethics, 108, 4-32. carter, j. r., & irons, m. (1991). are economists different, and if so, why? journal of economic perspectives, 5(2), 171–177. colby, a., & kohlberg, l. (1987). the measurement of moral judgment, vol. i: theoretical foundations and research validation. cambridge, ma: cambridge university press. dana, j., weber, r. a., & kuang, j. x. (2007). exploiting moral wiggle room: experiments demonstrating an illusory preference for fairness. economic theory, 33, 67–80. doi: 10.1007/s00199-006-0153-z dietrich, f., & list, d. (2013a). a reason-based theory of rational choice. noûs, 47, 104-134. https://doi.org/10.1111/j.1468-0068.2011.00840.x dietrich, f., & list, d. (2013b). where do preferences come from? international journal of game theory, 42, 613–637. doi: 10.1007/s00182-012-0333-y foot, p. (1972). morality as a system of hypothetical imperatives. philosophical review, 81, 305-316. doi: 10.2307/2184328 frank, r. h., gilovich, t., & regan, d. t. (1993). does studying economics inhibit cooperation? journal of economic perspectives, 7 (2), 159-171. frank, r. h., gilovich, t., & regan, d. t. (1996). do economists make bad citizens? journal of economic perspectives, 10(1), 187-192. doi: 10.1257/jep.10.1.187 frey, b. s. (1986). economists favour the price system. who else does? kyklos, 39(4), 537–563. https://doi.org/10.1111/j.1467-6435.1986.tb00677.x frey, b. s., & meier, s. (2003). are political economists selfish and indoctrinated? evidence from a natural experiment. economic inquiry, 41, 448-462. https://doi.org/10.1093/ei/cbg020 gul, f., & pesendorfer, w. (2008). the case for mindless economics. in a. caplin & a. schotter (eds.), t he foundations of positive and normative economics: a handbook (pp. 3-39). oxford: oxford university press. heinrichs, k., minnameier, g., gutzwiller-helfenfinger, e., & latzko, b. (2015). „don’t worry, be happy“? – das happy-victimizer-phänomen im berufsund wirtschaftspädagogischen kontext. zeitschrift für berufs und wirtschaftspädagogik, 111, 31-55 . hermkes, r. (2016). perception, abduction, and tacit inference. in l. magnani & c. casadio (eds.), model-based reasoning in science and technology – logical, epistemological, and cognitive issues (pp. 399-418). heidelberg: springer. hobbes, t. (1651/2001). leviathan. south bend, in: infomotions. kant, i. (2002/1785). groundwork for the metaphysics of morals (ed. and transl. by a. w. wood). new haven, ct: yale university press. keller, m., lourenço, o., malti, t., & saalbach, h. (2003). the multifaceted phenomenon of „happy victimizers“: a cross-cultural comparison of moral emotions, british journal of developmental psychology, 21, 1-18. doi: 10.1348/026151003321164582 killen, m., & smetana, j. g. (2015). origins and development of morality. in m. e. lamb (ed.), handbook of child psychology and developmental science, vol. 3 (7th ed.; pp. 701-749). ny: wiley-blackwell. https://doi.org/10.1002/9781118963418.childpsy317 kohlberg, l. (1984). essays on moral development, vol. 2: the psychology of moral development . san francisco, ca: harper & row. krebs, d. l., & denton, k. (2005). toward a more pragmatic approach to morality: a critical evaluation of kohlberg’s model, psychological review, 112, 629-649. https://doi.org/10.1037/0033-295x.112.3.629 krettenauer, t. (2013). moral motivation, responsibility and the development of the moral self. in f. oser, k. heinrichs & t. lovat (eds.), handbook of moral motivation: theories, models, applications. (pp. 215-228). rotterdam: sense. krettenauer, t., malti, t., & sokol, b. w. (2008). the development of moral emotion expectancies and the happy victimizer phenomenon: a critical review of theory and application, european journal of developmental science, 2, 221-235. doi: 10.3233/dev-2008-2303 ledyard, j. (1995). public goods: a survey of experimental research. in j. kagel & a. roth (eds.), handbook of experimental economics (pp. 253–279). princeton: princeton university press. malti, t., & krettenauer, t. (2013). the relation of moral emotion attributions to prosocial and antisocial behavior: a meta-analysis. child development, 84, 397-412. doi: 10.1111/j.1467-8624.2012.01851.x marwell, g., & ames, r. (1981). economists free ride, does anyone else? journal of public economics, 15(3), 295–310. minnameier, g. (2000). strukturgenese moralischen denkens eine rekonstruktion der piagetschen entwicklungslogik und ihre moraltheoretischen folgen . münster: waxmann. minnameier, g. (2001). a new stairway to moral heaven – a systematic reconstruction of stages of moral thinking based on a piagetian 'logic' of cognitive development. journal of moral education , 30, 317-337. https://doi.org/10.1080/03057240120094823 minnameier, g. (2005). developmental progress in ancient greek ethics. european journal of developmental psychology, 2, 71-99. https://doi.org/10.1080/17405620444000274a minnameier, g. (2010). the problem of moral motivation and the happy victimizer phenomenon – killing two birds with one stone . new directions for child and adolescent development, 129, 55-75 . https://doi.org/10.1002/cd.275 minnameier, g. (2012). a cognitive approach to the ‘happy victimiser’. journal of moral education, 41, 491-508. https://doi.org/10.1080/03057240.2012.700893 minnameier, g. (2013). deontic and responsibility judgments: an inferential analysis. in f. oser, k. heinrichs & t. lovat (eds.), handbook of moral motivation: theories, models, applications. (pp. 69-82). rotterdam: sense. minnameier, g. (2014). moral aspects of professions and professional practice. in s. billet, c. harteis & h. gruber (eds.), international handbook of research in professional and practice-based learning (pp. 57-77). berlin: springer. minnameier, g. (2016a). rationalität und moralität – zum systematischen ort der moral im kontext von präferenzen und restriktionen. zeitschrift für wirtschaftsund unternehmens¬ethik, 17, 259-285. minnameier, g. (2016b). abduction, selection, and selective abduction. in l. magnani & c. casadio (eds.), model-based reasoning in science and technology – logical, epistemological, and cognitive issues (pp. 309-318). heidelberg: springer. minnameier, g. (2017). forms of abduction and an inferential taxonomy. in l. magnani & t. bertolotti (eds.),springer handbook of model-based reasoning (pp. 175-195). berlin: springer. minnameier, g. (2018). reconciling morality and rationality – positive learning in the moral domain. in o. zlatkin-troitschanskaia, g. wittum & a. dengel (eds.), positive learning in the age of information (plato) a blessing or a curse? (pp. 347-361). wiesbaden: springer vs. minnameier, g., & schmidt, s. (2013). situational moral adjustment and the happy victimizer. european journal of developmental psychology, 10, 253-268. doi: 10.1080/17405629.2013.765797 minnameier, g., beck, k., heinrichs, k., & parche-kawik, k. (1999). homogeneity of moral judgement? apprentices solving business conflicts. journal of moral education , 28, 429-443. https://doi.org/10.1080/030572499102990 minnameier, g., heinrichs, k., & kirschbaum, f. (2016). sozialkompetenz als moralkompetenz – theoretische und empirische analysen. zeitschrift für berufsund wirtschaftspädagogik, 112, 636-666. nucci, l. (2008). social cognitive domain theory and moral education. in l. nucci, & d. narvaez (eds.), handbook of moral development and character education (pp. 291–309). oxford: routledge. nunner-winkler, g. (1999). development of moral understanding and moral motivation. in f. e. weinert & w. schneider (eds.), individual development from 3 to 12 (pp. 253–292). cambridge, uk: cambridge university press. nunner-winkler, g. (2007). development of moral motivation from childhood to early adulthood. journal of moral education, 36, 399-414. https://doi.org/10.1080/03057240701687970 nunner-winkler, g. (2013). moral motivation and the happy victimizer phenomenon. in f. oser, k. heinrichs & t. lovat (eds.), handbook of moral motivation: theories, models, applications. (pp. 267-288). rotterdam: sense. nunner-winkler, g., & sodian, b. (1988). children’s understanding of moral emotions, child development, 59, 1323-1338. doi: 10.2307/1130495 paulus, m. (2014). the emergence of prosocial behavior: why do infants and toddlers help, comfort, and share? child development perspectives, 8, 77-81. https://doi.org/10.1111/cdep.12066 paulus, m., & moore, c. (2014). the development of recipient-dependent sharing behaviour and sharing expectations in preschool children. developmental psychology, 50, 914-921. https://doi.org/10.1037/a0034169 piaget, j., & garcia, r. (1989). psychogenesis and the history of science. new york: columbia university press. rai, t. s., & fiske, a. p. (2011). moral psychology is relationship regulation: moral motives for unity, hierarchy, equality, and proportionality. psychological review, 118, 57-75. https://doi.org/10.1037/a0021867 rest, j. r. (1984). the major components of morality. in w. m. kurtinez & j. l. gewirtz (eds.), morality, moral behavior, and moral development (pp. 24-38). new york: wiley. rosati, c. s. (2016). moral motivation. in e. n. zalta (ed.), the stanford encyclopedia of philosophy, url = . rubinstein, a. (2006). a sceptic’s comment on the study of economics. economic journal, 116, c1-c9. https://doi.org/10.1111/j.1468-0297.2006.01071.x rustichini, a., & villeval, m. c. (2014). moral hypocrisy, power, and social preferences. journal of economic behavior & organization, 107, 10-24. https://doi.org/10.1016/j.jebo.2014.08.002 samuelson, p. a. (1938). a note on the pure theory of consumer’s behaviour. economica, 5, 61–71. doi: 10.2307/2548836 samuelson, p. a. (1948). consumption theory in terms of revealed preference. economica, 15, 243–253. 10.2307/2549561 selten, r., & ockenfels, a. (1998). an experimental solidarity game. journal of economic behavior & organization 34 (4), 517-539. https://doi.org/10.1016/s0167-2681(97)00107-8 smith, m. (1994/2005). the moral problem. oxford: blackwell. sugden, r. (2010). is there a distinction between morality and convention? in m. baurmann, g. brennan, r. e. goodin & n. southwood (eds.), norms and values: the role of social norms as instruments of value realization (pp. 47-65). baden-baden: nomos. thoma, s. j., & bebeau, m. j. (2013). moral motivation and the four component model. in f. oser, k. heinrichs & t. lovat (eds.), handbook of moral motivation: theories, models, applications. (pp. 49-68). rotterdam: sense. turiel, e. (1983). the development of social knowledge: morality and convention. new york: cambridge university press. turiel, e. (2002). the culture of morality: social development, context, and conflict. new york: cambridge universtiy press. warneken, f., & tomasello, m. (2009). the roots of human altruism. british journal of psychology, 100, 455–471. https://doi.org/10.1348/000712608x379061 zangwill, n. (2003). externalist moral motivation. american philosophical quarterly, 40, 143-154. mittelmeier et al publication frontline learning research vol.6 no. 2 (2017) 20 38 issn 2295-3159 ‘a double-edged sword. this is powerful but it could be used destructively’: perspectives of early career education researchers on learning analytics jenna mittelmeiera, rebecca l. edwardsb, sarah k. davisb, quan nguyen a, victoria l. murphyaa, leonie brummercc, bart rientiesaa athe open university, uk buniversity of victoria, canada cuniversity of groningen, the netherlands article received 2 march / revised 8 june / accepted 3 august/ available online 17 august abstract learning analytics has been increasingly outlined as a powerful tool for measuring, analysing, and predicting learning experiences and behaviours. the rising use of learning analytics means that many educational researchers now require new ranges of technical analytical skills to contribute to an increasingly data-heavy field. however, it has been argued that educational data scientists are a ‘scarce breed’ (buckingham shum et al., 2013) and that more resources are needed to support the next generation of early career researchers in the education field. at the same time, little is known about how early career education researchers feel towards learning analytics and whether it is important to their current and future research practices. using a thematic analysis of a participatory learning analytics workshop discussions with 25 early career education researchers, we outline in this article their ambitions, challenges and anxieties towards learning analytics. in doing so, we have provided a roadmap for how the learning analytics field might evolve and practical implications for supporting early career researchers’ development. keywords: learning analytics, educational research, early career researchers info corresponding author mail: jenna.mittelmeier@open.ac.uk doi: 10.14786/flr.v6i2.348 1. introduction there is increased awareness in education that using learning analytics to measure students’ and teachers’ attitudes, behaviours, and cognition provides potential opportunities to enhance and enrich learning and teaching (agudo-peregrina, iglesias-pradas, conde-gonzález, & hernández-garcía, 2014; dawson & siemens, 2014; tempelaar, rienties, & giesbers, 2015; van leeuwen, janssen, erkens, & brekelmans, 2015; winne, 2017). furthermore, fine-grained learning analytics data can provide opportunities to test, validate, or build educational theories (malmberg, järvelä, & järvenoja, 2017; winne, 2017). despite these potentials, large-scale uptake of learning analytics has so far been rather limited (ferguson et al., 2016; papamitsiou & economides, 2016). this is reflective of wider scepticism about whether learning analytics-supported teaching and, more generally, technology-enhanced learning, can be as effective as ‘traditional’ methods (kirschner & erkens, 2013; selwyn, 2015, 2016). in this regard, a group of prominent researchers in learning analytics and educational data mining recently argued that educational data scientists were a ‘scarce breed’(buckingham shum et al., 2013). they further noted that critical reflection was needed on how the educational research community could train, support, and prepare the next generation of early career education researchers to work across traditional and new boundaries of research (buckingham shum et al., 2013). however, in light of this suggestion, questions remain around how early career education researchers reflect and act upon an increasingly complex, data-heavy research environment. to answer this, we conducted an in-depth, participatory workshop with 25 early career education researchers at the earli jure 2017 conference, which collaboratively and critically evaluated the affordances and limitations of learning analytics. in particular, we outline through a thematic analysis of workshop discussions the ambitions, challenges, and anxieties of early career researchers towards learning analytics, and how and why they might (not) use learning analytics methods in their own research now and in the future. 2. learning analytics: six main research topics learning analytics can be defined as ‘the measurement, collection, analysis, and reporting of data about learners and their contexts for the purpose of understanding and optimising learning and the environments in which it occurs’ (lak, 2011). recent literature in the learning analytics field addresses a large range of trends and challenges in this emerging research field (ferguson et al., 2016; lang, siemens, wise, & gašević, 2017; papamitsiou & economides, 2016; selwyn, 2015; siemens, 2013). for example, in the recent handbook of learning analytics, lang et al. (2017) identified 27 topics of interest. given the limited time available for the workshop, we aimed to narrow this work to outline overarching topics of importance that represented cutting-edge debates in the learning analytics field. to aid in this process, we cross-referenced this list with two recent meta-analyses of learning analytics (ferguson et al., 2016; papamitsiou & economides, 2016), which have provided in-depth reviews of current findings, debates in the field, and limitations of recent work. to select topics that were timely to the field, we considered common themes presented at the recent learning analytics and knowledge conferences (lak), which all workshop facilitators had previously attended. altogether, these various sources of evidence and expertise were combined to narrow down and identify six main research topics in learning analytics to focus upon in this workshop, which are outlined in table 1. table 1 six topics of discussion during earli jure learning analytics workshop in the next sections, a description and justification for using these six topics in our workshop are provided, followed by a description of our research questions for this study. 2.1 learning analytics measurements the opportunity to capture and measure rich and fine-grained data from authentic learning environments in real-time has been one of the unique advantages of learning analytics research. for example, in a large-scale study amongst 141 courses, rienties and toetenel (2016) found that the way teachers designed courses significantly predicted how 111,000 students were studying online, whereby follow-up training with 80+ teachers helped to improve the learning design. however, since the data are often collected authentically (versus in controlled lab settings) and automatically (as opposed to collection first-hand by researchers), questions about how to compute meaningful metrics remain. for example, the typical log files that record user activities on a browser lack insight into many factors, such as who is behind the screen, whether they are active or taking a break, or whether they are focused on task or distracted (kovanovic, gašević, dawson, joksimovic, & baker, 2016; kovanovic et al., 2015). without explicit engagement with the underlying assumptions behind these data, researchers could overestimate or underestimate results (gašević, dawson, & siemens, 2015). therefore, it is crucial for early career education researchers to be aware of both the strengths and limitations of these measurements to adopt learning analytics in meaningful manners. 2.2 ethics several ethical debates have emerged around the practical application of learning analytics techniques. for example, questions remain around transparency of how and why data is collected and used (slade & prinsloo, 2013), what information is missing from automatically captured data (ruppert et al., 2015), what biases and subjectivities are inherent to techniques used for analysis (boyd & crawford, 2012), and whether educational institutions have an obligation to utilise learning analytics to support students (prinsloo & slade, 2017a). perhaps most prominently, there are criticisms around data ownership and students’ consent for their data to be collected. whether data are owned by universities, students, or even researchers has large implications for the ethicality of learning analytics (drachsler & greller, 2016). furthermore, generic consent is often collected at enrolment, but, as a wide variety of research projects may potentially use students’ data in new and various ways, it is difficult to argue that this is informed consent (slade & prinsloo, 2013). as the field struggles to answer these questions, it is critical that early career education researchers engage in the debates surrounding ethical learning analytics. 2.3 learning design in recent years, there has been increasing interest in the connection between learning design and learning analytics. in his lak16 keynote speech, kirschner (2016) critically reflected on the lack of pedagogical context in learning analytics research, which could prevent researchers from asking the right questions, selecting the right metrics, or interpreting results in meaningful manners (gašević et al., 2015). in this way, learning design is an emerging field that aims to develop a ‘descriptive framework for teaching and learning activities (“educational notation”), and to explore how this framework can assist educators to share and adopt great teaching ideas’ (dalziel, 2016). by capturing and visualising the design of learning activities, learning design could provide pedagogical context to support interpreting and translating learning analytics findings into tangible interventions (lockyer et al., 2013; persico & pozzi, 2015). a gradual accumulation of empirical evidence has reinforced the importance of aligning learning analytics with learning design, as the way teachers design for learning has been found to be significantly related to students’ engagement, pass rates, and satisfaction (nguyen et al., 2017; rienties & toetenel, 2016). in this way, learning design offers opportunities for early career education researchers to engage with educational theories and pedagogies that underpin their analysis. 2.4 learning dispositions early learning analytics research often focused on predictive models based on extracting data from digital platforms (such as learning management systems) and institutional records of student information. while these studies (e.g., agudo-peregrina et al., 2014) provided valuable foundations for understanding learning analytics’ potential, they relied primarily on demographics, grades, and trace data, providing a relatively simplistic story of students’ experiences and behaviours. a reliance on such narratives without more complex insights may lead to difficulties in designing pedagogically-informed interventions (tempelaar et al., 2015; tempelaar et al., 2017). in response, buckingham shum and deakin crick (2012) proposed a ‘dispositional learning analytics’ approach that combines behavioural data (e.g., that which is mined from learning management systems) with more complex data about the learner (e.g., their values, attitudes, and dispositions measured through self-reported surveys) (malmberg et al., 2017; tempelaar et al., 2015). for example, recent research in blended courses of mathematics and statistics (tempelaar et al., 2017) found that linking learning dispositions data with behavioural learning analytics data can lead to actionable feedback. as the learning analytics field matures, therefore, it becomes more necessary for early career education researchers to understand and apply more complex learner data for use in tandem with learning analytics. 2.5 combining learning analytics with other methods in light of the issues outlined above, combining learning analytics with other methods offers opportunities for both early career researchers and the wider field. for example, it has been argued that combining more classically automated and quantitative learning analytics data with qualitative methods (interviews, focus groups, think-aloud protocols, etc.) can develop insight into the how and why factors impacting quantitative findings (chatti et al., 2012). similarly, merceron, blikstein, and siemens (2016) described learning as ‘multimodal’ and involving many simultaneous mental and physical processes, noting maturity in learning analytics research that incorporates multiple methods to capture this. some potential additional methods include natural language processing, eye tracking, social network analysis (such as through surveys), visualisation techniques, discourse analysis, or emotional measurements (suthers & verbert, 2013). combining learning analytics with these techniques reflects the diversity of the learning analytics field as a whole, which crosses multiple academic disciplines (dawson, gasevic, siemens, & joksimovic, 2014). this diversification of learning analytics research has implications for early career education researchers as they develop skills and competencies to contribute to the field. 2.6 personalisation versus generalization learning analytics applications can be continuously adapted and, to some extent, personalise the learning environment for students (ifenthaler & widanapathirana, 2014). however, personalisation relies on categorising students based on behaviours and traits (buckingham shum & deakin crick, 2012), and some have noted that there are grey areas between ‘categorising’ and ‘stereotyping’ (prinsloo & slade, 2017b; scholes, 2016; slade & prinsloo, 2013). for instance, learning analytics could lead to erroneous assumptions about future behaviours and performance (scholes, 2016). issues also remain around the bias of researchers or teachers, whose views towards certain students or categories of students might influence the interpretation of findings. therefore, it is important for early career education researchers to be explicitly cognisant of the fine line between personalisation and generalisation to develop sustainable algorithms and interpretations that limit researcher bias. 3. method 3.1 purpose and research questions current discussions in learning analytics research, as outlined by the six topics above, depict a complex and diverse emerging field. as such, the next generation of early career education researchers represent an important voice in understanding how learning analytics research will develop and progress (buckingham shum et al., 2013). in this study, we aimed to understand how these early career researchers make sense of learning analytics in light of these ongoing debates and discussions in the field. using an interactive workshop structured around these six learning analytics topics amongst 25 early career education researchers in the education field, we explored the following research questions: 1. what are the main strengths and limitations of learning analytics according to early education career researchers? 2. in what ways do early career education researchers embed learning analytics approaches into their own research practices? in answering these questions, we have outlined an in-depth account of how early career education researchers approach learning analytics techniques, providing insight into potential pathways forward for the field. 3.2 procedure and setting this study describes the results of an invited workshop at the earli jure 2017 conference in tampere, finland, which was entitled ‘three different perspectives on why you need learning analytics and educational data-mining’. earli (european association for research on learning and instruction) is an international organisation for researchers in the education field, and jure (junior researchers of earli) focusses specifically on early career researchers in education. early career researcher in this context is defined as master’s students, phd students, and those within two years of receiving their phd. more information about the conference and its aims are available at: http://www.earli-jure2017.org/ the workshop was facilitated by three researchers from the open university (ou), including a phd student, a postdoctoral researcher, and a professor. the workshop was 90 minutes long and consisted of opening and closing sessions, combined with breakaway small group discussions for each of the six topics outlined in the introductory sections of this article (see table 1). participants were free to choose which small group topic they wished to discuss and not all participants discussed all topics. on average, each participant attended two small group discussions. the overall aim of the workshop was to encourage critical assessment and debate on the merits and drawbacks of learning analytics approaches in participants’ own institutional and cultural contexts. table 2 describes the workshop schedule adopted. table 2 workshop schedule outline the workshop was highly participatory with opportunities for attendees to take part in both small and large group discussions. a pilot was conducted two weeks before the earli jure workshop with 12 early career researchers at the open university to test and fine-tune the workshop design, as well as the final selection of the six topics. the role of the facilitators was mostly limited to moderating and encouraging discussion, as the aim was to offer collaborative experiences between early career education researchers. the workshop facilitators only provided a one-minute introduction of each of the six topics outlined in table 1 before breaking into small groups and contributed to discussions only to facilitate conversation between participants (i.e., ‘what are your thoughts on this topic?’ or ‘how does this relate to your own work?’). the full slides used by the facilitators during the workshop is available online (see: https://bit.ly/2mep4u0). this design allowed early career researchers to contribute their own voices and experiences, as well as reflect upon their own practices. in terms of the six learning analytics topics (table 1), our aim was to encourage discussion around current debates on these topics while creating an environment that allowed participants to be critical about their relevance and contribute their own opinions. the workshop design further allowed participants to draw upon their own needs as researchers, including what their practices had in common with learning analytics in the wider field, as well as what is unique about their own experiences. during the full group discussion at the start of the workshop (activity 2 in table 2), 19 of the 25 participants contributed their opinion about the strengths and limitations of learning analytics. in the small group discussions (activity 3 in table 2), each of the 25 participants made at least one comment about their chosen small group discussion topic. altogether, all participants were active during the workshop and most provided substantial discourse about their opinions and experiences. 3.3 participants workshop participants were recruited via the earli jure conference registration process and 25 conference attendees joined the session. overall, there was a wide range of participants, including phd students, postdoctoral researchers, those from non-academic sectors (e.g., non-profit or regulatory bodies), and policymakers. in terms of geographical spread, participants were from belgium, canada, china, germany, finland, the netherlands, pakistan, spain, singapore, vietnam, the united kingdom, and the united states. although all participants were conducting education-related research, they also came from a variety of disciplines (e.g., educational technology, computer science, learning sciences, primary education, secondary education, higher education, educational psychology, etc.). following the workshop, four participants collaborated with the research team as co-researchers in this study and contributed to the analysis and writing. the research team was a diverse mix of early career researchers from canada, the netherlands, uk, usa, and vietnam, who conducted research in canada, the netherlands, and the uk. this collaboration ensured the credibility and authenticity of the findings and contributed to the validity by ‘building the participant’s view into the study’ (creswell & miller, 2000, p. 128). 3.4 instruments 3.4.1 online surveys an online survey tool (pollev.com) was used to probe six questions to participants during the workshop. these questions are outlined in table 3. in total, 17 to 20 participants completed each of the questions at the start of the workshop (questions 1-4), 22 participants completed the topic ranking activity (question 5), and 13 participants completed end-of-workshop satisfaction questions (questions 6-7). descriptive statistics of our findings from these questions are compiled in the results section to provide a general understanding of participants’ backgrounds and views on learning analytics and its applicability to their own research. table 3 survey questions asked of participants 3.4.2 transcripts of discussions both large and small group discussions were recorded during the workshop and transcribed verbatim. the workshop facilitators also wrote discussion points and perspectives from participants on flipcharts, which served as a secondary data source for qualitative analysis. in total, eight transcripts and three documents were used to inform our findings. although workshop participants were split into small groups to discuss different topics during activity 3 (table 2), we were interested in understanding the common and overarching themes that emerged across the six discussion topics. as such, we used a thematic analysis method, as described by lichtman (2013), to summarise and group common themes across all the data available from this workshop, including whole group discussions, small group discussions, workshop notes, and facilitator notes and reflections. for the first stage of the thematic analysis, each member of the research team provided an initial summary of key emergent themes for one of the transcripts or data sources, which were then each reviewed and confirmed by two additional members of the research team for accuracy and reliability. in the next stage, the themes from each data source were then combined to develop an overarching coding book of common themes across all data sources (as outlined in appendix 1). this two-stage analysis provided an understanding of participants’ sentiments both within and across the whole and small group discussions. afterwards, notes and reflections on these themes were compared between all authors to confirm and validate findings, counter any disagreements, and further develop the narrative of our findings. 3.4.3 post-workshop reflections directly after the workshop, the three workshop facilitators and a selection of workshop participants met to reflect upon the discussions and the main themes that emerged within and across the six discussion topics. this conversation was recorded and transcribed to inform the data analysis described in section 3.3.2. 4. results the descriptive statistics from the online survey provided an overview of participants’ backgrounds and views towards learning analytics. the first question asked on a 1-5 scale the degree of understanding participants felt they had about learning analytics (1 = no idea and 5 = expert). to this, only 15% (n = 3) considered themselves experts, while the majority (80%, n = 16) rated their understanding was somewhere in the middle (between scales 2 – 4). we next asked whether participants felt that learning analytics was relevant to their own research on a 1-5 scale (1 = irrelevant, 5 = extremely relevant) and the majority (80%, n = 16) agreed (scored 3 and above). most participants (83%, n = 17) also felt that learning analytics are of importance for their universities (scored 3 and above). overall, despite a large variation in terms of the level of familiarity with learning analytics, the majority of early career education researchers indicated that learning analytics is relevant to their work and institutions. 4.1 rq 1. main strengths and limitations of learning analytics in the opening large-group discussion, participants outlined a number of key strengths and weaknesses of learning analytics, which is summarised in table 4. perspectives were provided from 19 of the 25 participants during this portion of the workshop, indicating high participation across the full sample. table 4 main strengths and limitations from the opening session at the jure workshop in terms of strengths, participants indicated first and foremost that learning analytics data can provide insights into actual learner behaviours. six whole-group discussion participants, in that way, argued that learning analytics data is more ‘objective.’ for example, one participant commented, ‘it’s not relying on self-report.’ of course, whether learning analytics data is actually objective can be debated (mirriahi & vigentini, 2017; prinsloo & slade, 2017b). this point was further articulated by one participant, who suggested ‘perceived objectivity…because you have to do something with the data to get something out of it.’ in addition, the power to visualise data and provide immediate feedback to both teachers and students was frequently noted by several participants, as highlighted by recent research (charleer, klerkx, duval, de laet, & verbert, 2016; nguyen et al., 2017). furthermore, two additional participants noted that learning analytics can measure ‘a lot of things at the same time,’ while most educational researchers primarily have to be selective in terms of the instruments and constructs that can be collected or measured. another potential strength highlighted by two participants was the opportunity to give direct, personalised feedback to learners, or so-called formative learning analytics (sharples et al., 2016). finally, one participant noted that learning analytics research may provide cross-institutional collaboration opportunities to test and validate large and small educational theories (winne, 2017). altogether, it was apparent that early career participants were engaged with the potential power of learning analytics and felt that it had value for the future of educational research, as indicated by over half of the comments provided during the whole group discussion. in terms of limitations (also summarised in table 4), there were fears from four participants that parameter-driven analytics might lead to ‘false data’, which potentially could lead to wrong conclusions or inappropriate interventions based on an incomplete picture of learning. additionally, concerns were voiced by five participants around the potential for learning analytics to seem like a ‘black box’ (garcía et al., 2012; kovanovic et al., 2015). four participants mentioned that learning analytics required an intimidating level of statistical techniques and technical competencies, which had implications for their skills development pathways and overall perception of the field amongst peers. relatedly, another comment outlined potential risks of cognitive overload for teachers and students when making sense of data. ethics was also a frequent cause for concern. in particular, four participants were worried about ‘who owns the data’ students, teachers, institutions, governments, companies? there were some additional concerns around informed consent and whether learners were aware that their data was being collected (slade & prinsloo, 2013). at the same time, they noted difficulties in that some countries or institutions limited or even forbade measuring and monitoring students’ and teachers’ activities (prinsloo & slade, 2017b; slade & prinsloo, 2013), which was detrimental to their research. a final but crucial limitation outlined by the final comment that links with all points above was the risk of labelling (i.e., ‘the oversimplification of student characteristics’), whereby students (or teachers) may be identified to fit in a particular category of users based upon a mix of objective and subjective indicators and algorithms. altogether, nearly half of the discussion points made by early career participants in the whole-group discussion appeared to be engaged with the current debates in the field and felt that creative solutions were needed to both accomplish their personal research goals and drive the field forward. 4.2. rq 2. embedding learning analytics into research practice in the online survey, we asked workshop participants to rank their personal interest in the six identified learning analytics topics. the subjects were collectively ranked in the following order: (1) learning analytics measurement, (2) learning design, (3) combining learning analytics with other methods, (4) generalisation versus personalisation, (5) ethics, and (6) learning dispositions. this seemed to initially indicate that early career education researchers were more strongly engaged with issues pertaining to methods and methodologies over those of more theoretical importance. one explanation could be that many of the students were phd students who were interested in discussing and framing their research projects and methodologies. in the small groups, discussions centred on the chosen topic with a focus on how participants were currently integrating learning analytics into their own research practices. potential future uses of learning analytics to enhance the participants’ research projects were also explored in the dialogue in light of the small group topic. although discussions varied across the small groups and it was not possible to elicit whether sentiments were shared across all participants, our thematic analysis identified several common themes across all six small group discussions in this workshop. altogether, the majority of early career education researchers in this workshop: 1. used learning analytics to support student success through self-regulated learning approaches 2. felt that learning analytics were on the cutting edge of educational research, which posed unique challenges for their own research and the wider field 3. believed that approaches of combining learning analytics data with other methods strengthened learning analytics and aimed to incorporate multiple methods into their own research 4. contemplated theory-driven learning analytics versus learning analytics-driven theory 5. aimed to contextualise their data to the realities of students’ lives and experiences 6. recognised a multitude of ethical ramifications for their own learning analytics research definitions of each theme and its subthemes are provided in appendix 1. the following sections provide a detailed narrative of each theme derived from our analysis. 4.2.1 theme 1: using learning analytics to support student success by taking a self-regulated learning approach ‘there is a lot of use of demographic data in la and i like what phil winne said in the handbook of learning analytics…that he thinks the focus should be on things learners can change about themselves.’ in all six of the small discussion topic groups, participants suggested that combining learning and learner data provides opportunities to support student success, both for students who are at risk of failure and students who are objectively successful but could achieve more. in four of the six small groups, this discussion was framed through the perspective of researching things that students have agency over (e.g., behaviours that they can change) rather than static demographic variables. by focusing on these types of variables, several participants noted that it might be possible to provide just-in-time and just-enough feedback to students so that they can strategically adjust their learning approaches. particularly in topic 4, 5, and 6 discussions, participants suggested that it is critical for students to take an active role and learn to effectively regulate their own learning, thereby not use learning analytics as a ‘crutch.’ in this way, participants in these groups argued that it is important to ‘sell’ learning analytics as beneficial to students, not just advisors and teachers. 4.2.2 theme 2: learning analytics are on the cutting edge of educational research and this poses unique challenges ‘for right now, it's just me and i have to find everyone. so, i'm building it from the ground up and i have to find people at my institution with these skills.’ in the whole group discussion and our workshop survey, participants noted there are wide variations in levels of familiarity and comfort with learning analytics techniques at their institutions. in five of the six small groups, participants elaborated that, despite this inconsistency, they personally viewed learning analytics as a powerful method because it brings together researchers from diverse disciplines in a way that traditional educational research methods have not. however, participants have found it challenging to establish (or find) needed expertise, work with complex data, learn and use advanced analysis techniques, and access needed technologies. this notion was highlighted in the whole-group discussion, the workshop survey, and in all six of the small group discussions. across the various discussions, there was a sense that these challenges stemmed from the newness of the field and participants often likened using learning analytics to embarking alone on an adventure for which you are not prepared. 4.2.3 theme 3: combining learning analytics methods with other methods ‘i ran in a lot of trouble of, like, getting a lot of data but not really feel how to interpret it. because you only see the patterns, but you don’t know why or how … i’m really thinking about mixing it, with the log data analysis but then also interviewing them … you usually need somebody’s explanation about why they did what they did.’ many participants in both the whole and small group discussions consistently noted the value of combining learning analytics with more traditional education research methodologies. although this was particularly prevalent in topic 5 (combining learning analytics with other methods) and topic 4 (learning dispositions), the value of using other methods in combination with the more automated and quantitative learning analytics data was outlined in three other small group discussions. when discussing their own work throughout the workshop, comments from 15 of the 25 participants indicated that early career education researchers were either already multiple methods or hoped to combine methods in the future. among other methods, participants referenced social network analysis surveys, interviews, document analysis, eye tracking technologies, and psychometric questionnaires. in general, participants aimed to combine learning analytics with other methods because they believed that this would provide a stronger evidence base, mitigate biases, and add the ‘how-and-why’ or intent to their analysis and findings. however, several participants in the topic 5 discussion expressed concerns that they may not be able to access the data or participant numbers they would require to use multiple methods. 4.2.4 theme 4: contemplating theory-driven learning analytics versus learning analytics-driven theory ‘…the theory often seduces you to see something in the data which might not entirely be there. it’s still quite tricky…’ there was some debate in the whole group discussions (four comments) and across the small group discussions (twelve comments) on whether learning analytics should be theory-driven or data-driven. discussed by nearly half of the participants throughout the workshop, there were conflicting views between them about whether theory should inform data analysis choices or whether the data should be explored to develop or inform new theories. four participants described that focusing on existing theories was limiting, as the data should ‘speak for itself’. in this way, there were suggestions that theories should be developed from the analysis and findings because existing theories might not match students’ measurable behaviours. at the same time, five other participants gravitated towards more theory-driven analytics to understand what variables to consider and to bring the ‘why’ factor into their findings. this second perspective was more in line with more established experts in the field (gašević et al., 2015), who have argued that learning analytics researchers must explicitly engage with the underlying assumptions of using learning analytics data and the educational theories underpinning their research. 4.2.5 theme 5: aiming to contextualise their data to the realities of students’ lives and experiences ‘sometimes it’s not the amount of time they spent, but the sequence of time they spent on something. or you have to know the time when they have an exam, so you need to take into account the context of the course.’ throughout all workshop activities, participants aimed to contextualise the data they collected in light of students’ realities and experiences. this was an explicit focus in the topic 6 small group discussion, but these sentiments were brought up in the remaining five small groups as well. in all groups, participants suggested that data is easier to understand in context and that findings do not always translate between contexts. this reflects positively on the future of the field, as nearly all participants recognised a need for complex, well-rounded research that goes beyond simple click data to capture the nuances of student experiences. at the same time, several participants, particularly in the topic 2, 3, and 5 discussions, noted limitations in their own access to data to accomplish this, which ultimately may skew results. for example, two participants commented that data may only be available from learning management systems, while students are also using another platform for their learning. early career education researchers may also not have access to participants for interviews or other studies. therefore, the ability to access data from multiple platforms and gain access to participants were key areas of concern for several participants. 4.2.6 theme 6: recognising a multitude of ethical ramifications ‘it’s about the argumentation of why you focus on the women rather than the men, maybe. for example, if the beneficial effects are larger for women than for men, then i would say, yeah, maybe. but, then again....’ across the whole group and all six small group discussions, most participants were keenly aware that strong ethics policies are needed when using learning analytics. although a key focus in the topic 2 discussion, participants in all small groups were concerned with protecting privacy and individual rights and frequently suggested that it might not be ethical to track everything that one might be capable of tracking. in the topic 6 discussion, participants indicated that learning analytics, particularly sharing learning analytics data back to teachers and students, might lead to unintended consequences. in the whole group discussion, one participant went so far as to call learning analytics a ‘double-edged sword,’ pointing out that learn analytics are ‘powerful, but could be used destructively’. in this way, participants throughout the workshop argued that learners, educators, researchers, and institutions might not be fully prepared to wield learning analytics in a useful and productive way. however, several participants, particularly those in topic 1, 2 and 5, did point out that the use of learning analytics has ethical benefits. for example, these participants felt that learning analytics provides a way to track and support all students to reduce bias and inequalities. altogether, there was an engagement across the workshop by most participants with ethical implications of learning analytics in line with current debates in the field (gašević, dawson, & jovanovic, 2016; prinsloo & slade, 2017b; slade & prinsloo, 2013). 5. discussion and practical implications previously, researchers have outlined a need for preparing early career education researchers for an increasingly data-heavy education sector (buckingham shum et al., 2013). in response, our thematic analysis of the earli jure learning analytics workshop activities have provided a detailed and nuanced account of early career education researchers’ feelings towards learning analytics (rq1) and the role it plays in their own research practices (rq2). our findings overall indicated that the early career participants were engaged with current debates and discussions in the field, were interested in developing creative solutions to overcome perceived issues and desired support for building their skills and expertise on the topic. workshop participants were broadly in agreement about several themes, including the potential of learning analytics to support learning, the value of combining learning analytics with other methods, the need to contextualise learning analytics to learners’ own lives, and the need to engage with ethical implications. these views are in line with experts in the field and indicate that most early career education researchers using learning analytics are in tune with the wider community of researchers (ferguson et al., 2016; papamitsiou & economides, 2016). at the same time, there were inconsistencies between participants in their views around issues such as the connection between theory and data and whether quantitative learning analytics data is ‘objective’, which was often in contrast to recent theoretical developments in the learning analytics field (dawson & siemens, 2014; gašević et al., 2015; kovanovic et al., 2015; lockyer et al., 2013). this suggests the need for early career education researchers to engage more with learning analytics theory in addition to its methodological and analysis affordances. the sentiments expressed during the workshop appeared to be consistent across participants and could potentially be generalisable to other emerging learning analytics researchers, particularly considering the wide diversity of researchers in attendance. however, it is important to note that most participants were already using or considering using learning analytics, and more diverse views or scepticism would likely be present among those who do not use or have more limited knowledge of learning analytics. drawing on the findings we have described, this study reveals four important areas for supporting early career education researchers in the learning analytics field: 1) providing adequate statistical and technical training, 2) building on their curiosities about the bigger picture of learning analytics research, 3) supporting access to different kinds of data, and 4) developing and fostering engagement with ethical issues. using a combination of large and small group interactions in this study, we revealed that early career researchers echo current debates frequently discussed by learning analytics experts, but are also concerned about how they can contribute their own perspectives to this emerging field. providing adequate statistical and technical training learning analytics research requires a certain amount of statistical and computational understanding (lang et al., 2017) and current training for education researchers does not always adequately prepare them to conduct these analyses confidently. if early career education researchers want to learn the skills needed to conduct learning analytics research, they often have to pursue this on their own time with little direction or guidance on what they should be learning. this was made evident in our study, whereby early career education researchers modestly rated their knowledge of learning analytics and expressed concerns and anxieties related to their own statistical and technical skills. the comments also link with recent work, which has highlighted poor data literacy as a key barrier to learning analytics adoption (ferguson et al., 2014). therefore, resources are needed to support early career education researchers in building capacities to interpret, analyse, and evaluate big (and small) data. collaborating on learning research projects with other experienced learning analytics researchers could also help demystify the processes used and provide opportunities for contributions from early career researchers. building on early career researchers’ curiosities about the bigger picture of learning analytics research in this study, early career education researchers showed concerns about learning analytics research presenting an incomplete picture of students’ learning. specifically, they were concerned about the oversimplification of students’ characteristics and using demographic data collected about students to inform interventions. participants recognised learning analytics might provide only a snapshot of students’ behaviour and performance. additionally, they wondered how data collected could be used to inform theories of learning (and vice versa) and how variables represent different aspects of learning. these echo common concerns and theoretical arguments that are being developed by experts in the wider learning analytics field (dawson & siemens, 2014; gašević et al., 2015; kovanovic et al., 2015; lockyer et al., 2013). as such, phd and postdoctoral research projects represent a valuable resource for pushing the boundaries of the learning analytics field. supervisors and line managers should, therefore, build on the curiosities of early career education researchers by supporting theoretical engagement, creative research ideas, and multi-method and multi-disciplinary approaches, in addition to technical skill development and automated learning analytics methods. supporting access to different kinds of data along with curiosities about the ‘bigger picture,’ early career education researchers in our study frequently noted the value of combining learning analytics with other methods. nearly all participants noted an interest and curiosity in experimenting with new approaches to answer complex research questions. by combining data from different methods, including qualitative methods, participants understood they could develop a more complete picture of learning than by only collecting log file data, which was in line with others in the field (chatti et al., 2012; suthers & verbert, 2013). in this regard, combining multiple methods can help bridge the gap between generalisation and personalisation of visualisations and feedback to be used by students and teachers. however, there was a recognition of challenges in accessing and incorporating different types of data and the expense of using new technologies with large groups of students. therefore, more support is needed to ensure that early career researchers have resources (time, access, training, etc.) to incorporate diverse methods into their research. developing and fostering ethical and theoretical engagements to effectively incorporate learning analytics into their own research, early career education researchers need to be well-versed in current debates around ethics (drachsler & greller, 2016; prinsloo & slade, 2017b; slade & prinsloo, 2013). the early career researchers in this workshop appeared to be aware of these issues, but often were unsure about how to address them. participants often felt they were small cogs in large institutional machines, with little power to influence or contribute to debates around ethics. although the wider field is attempting to address these issues through a growing number of ethical frameworks and privacy guidelines (gašević et al., 2016), early career researchers had major questions about how to collect data unbiasedly, while still capturing complex learning processes. therefore, an explicit engagement with developing early career education researchers’ knowledge of ethical frameworks and theories in learning analytics is vital. at the same time, the knowledges and experiences of early career researchers are important voices for institutions to consider when developing ethical frameworks. 5.1 limitations through an in-depth analysis of an earli jure conference workshop, this study has provided a nuanced look at the thoughts and experiences of early career education researchers in regards to learning analytics. we believe our study is the first of its kind to engage with how early career education researchers perceive and use learning analytics, providing insight into how the next generation of researchers will move this emerging field forward. in doing so, however, we note several limitations. first, it is recognised there was a potential for sample bias in relation to who volunteered to attend the workshop. it is likely that those who attended the workshop were already interested in and using learning analytics and that more diverse views could be elicited from early career researchers who have opted not to engage with the field. future research, therefore, should explore our research questions with a broader population of early career education researchers in different contexts to identify barriers experienced by those with little to no experience or knowledge of learning analytics. second, we note that the audio recordings for small group discussions were made in a busy, noisy room and, as such, were not always complete. some participants’ statements may have been omitted or unusable due to the quality of the recording. at the same time, discussion topics varied across the six small groups and it was not possible to elicit thoughts on some topics or themes from all participants. however, all main points from the small-group discussions were summarised by the facilitator, and all data were analysed and cross-examined by the seven early career researchers. 6. conclusion in this study, early career education researchers were aware of the strengths and limitations of learning analytics research and were grappling with previously identified issues in the field, including ethical and privacy issues, institutional barriers, concerns over atheoretical approaches, and developing complex stories about learning processes. this shows an awareness of field issues, as well as a familiarity with the literature and current research. the findings of this study also revealed how early career education researchers are struggling to situate themselves within the field and gain the complex skills necessary to appropriately embed learning analytics approaches into their own practices. these issues are essential considerations for the earli community and wider field of educational research, as it is clear that more resources are needed to support and develop the valuable expertise needed for early career researchers to contribute to the growing field of learning analytics. keypoints early career researchers in this study reflected positively on using learning analytics in their research to support and understand students’ learning processes early career researchers recognised a variety of benefits and challenges of using learning analytics approaches in their research, which were frequently in line with the theorisation of experts in the field most early career researchers favoured mixed methods approaches by combining learning analytics data with other quantitative and qualitative methods early career researchers debated the role of theory in learning analytics and, in particular, whether data should support theory or theory should support data common barriers to using learning analytics for early career researchers included access to data, technical skills in processing and analysing data, and contextualising findings to students’ lives and experiences references agudo-peregrina, á. f., iglesias-pradas, s., conde-gonzález, m. á., & hernández-garcía, á. (2014). can we predict success from log data in vles? classification of interactions for learning analytics and their relation with performance in vle-supported f2f and online learning. computers in human behavior, 31, 542-550. doi: 10.1016/j.chb.2013.05.031 boyd, d., & crawford, k. (2012). critical questions for big data. information, communication & society, 15(5), 662-679. doi: 10.1080/1369118x.2012.678878 buckingham shum, s., & deakin crick, r. (2012). learning dispositions and transferable competencies: pedagogy, modelling and learning analytics. paper presented at the learning analytics and knowledge (lak), vancouver, canada. buckingham shum, s., hawksey, m., baker, r., jeffery, n., behrens, j. t., & pea, r. (2013). educational data scientists: a scarce breed. paper presented at the learning analytics and knowledge (lak), leuven, belgium. charleer, s., klerkx, j., duval, e., de laet, t., & verbert, k. (2016). creating effective learning analytics dashboards: lessons learnt. paper presented at the european conference on technology enhanced learning (ec-tel) lyon, france. chatti, m. a., dyckhoff, a. l., schroeder, u., & thüs, h. (2012). a reference model for learning analytics. international journal of technology enhanced learning, 4(5-6), 318-331. doi: 10.1504/ijtel.2012.051815 creswell, j. w., & miller, d. l. (2000). determining validity in qualitative inquiry. theory into practice, 39(3), 124-130. doi: 10.1207/s15430421tip3903_2 dalziel, j. (2016). learning design: conceptualizing a framework for teaching and learning online . london, uk: routledge. dawson, s., gasevic, d., siemens, g., & joksimovic, s. (2014). current state and future trends: a citation network analysis of the learning analytics field . paper presented at the proceedings of the fourth international conference on learning analytics and knowledge, indianapolis, indiana, usa. dawson, s., & siemens, g. (2014). analytics to literacies: the development of a learning analytics framework for multiliteracies assessment. international review of research in open and distributed learning, 15 (4). doi: 10.19173/irrodl.v15i4.1878 drachsler, h., & greller, w. (2016). privacy and analytics: it's a delicate issue a checklist for trusted learning analytics . paper presented at the learning analytics and knowledge (lak), edinburgh, united kingdom. ferguson, r., brasher, a., cooper, a., hillaire, g., mittelmeier, j., rienties, b., . . . vuorikari, r. (2016). research evidence of the use of learning analytics: implications for education policy. in r. vuorikari & j. castano-munoz (eds.), a european framework for action on learning analytics. luxembourg: joint research centre science for policy report. ferguson, r., clow, d., macfadyen, l., essa, a., dawson, s., & alexander, s. (2014). setting learning analytics in context: overcoming the barriers to large-scale adoption . paper presented at the learning analytics and knowledge (lak), indianapolis, indiana, usa. garcía, r. m. c., pardo, a., kloos, c. d., niemann, k., scheffel, m., & wolpers, m. (2012). peeking into the black box: visualising learning activities. international journal of technology enhanced learning, 4(1-2), 99-120. doi: 10.1504/ijtel.2012.048313 gašević, d., dawson, s., & jovanovic, j. (2016). ethics and privacy as enablers of learning analytics. journal of learning analytics, 3 (1), 4. doi: 10.18608/jla.2016.31.1 gašević, d., dawson, s., & siemens, g. (2015). let’s not forget: learning analytics are about learning. techtrends, 59(1), 64-71. doi: 10.1007/s11528-014-0822-x ifenthaler, d., & widanapathirana, c. (2014). development and validation of a learning analytics framework: two case studies using support vector machines. technology, knowledge and learning, 19 (1), 221-240. doi: 10.1007/s10758-014-9226-4 kirschner, p. (2016). learning analytics: utopia or dystopia paper presented at the learning analytics and knowledge (lak) keynote. http://lak16.solaresearch.org/?events=keynote-slides-available kirschner, p., & erkens, g. (2013). toward a framework for cscl research. educational psychologist, 48(1), 1-8. doi: 10.1080/00461520.2012.750227 kovanovic, v., gašević, d., dawson, s., joksimovic, s., & baker, r. (2016). does time-on-task estimation matter? implications on validity of learning analytics findings. 2016, 2(3), 81-110. doi: 10.18608/jla.2015.23.6 kovanovic, v., gašević, d., dawson, s., joksimović, s., baker, r. s., & hatala, m. (2015). penetrating the black box of time-on-task estimation. paper presented at the learning analytics and knowledge (lak), poughkeepsie, new york, usa. lak. (2011). call for papers: 1st annual conference on learning analytics and knowledge. lang, c., siemens, g., wise, a. f., & gašević, d. (2017). handbook of learning analytics: society for learning analytics research. lichtman, m. (2013). qualitative research in education: a user's guide (3 ed.). thousand oaks, ca: sage. lockyer, l., heathcote, e., & dawson, s. (2013). informing pedagogical action: aligning learning analytics with learning design. american behavioral scientist, 57(10), 1439-1459. doi: 10.1177/0002764213479367 malmberg, j., järvelä, s., & järvenoja, h. (2017). capturing temporal and sequential patterns of self-, co-, and socially shared regulation in the context of collaborative learning. contemporary educational psychology, 49, 160-174. doi: 10.1016/j.cedpsych.2017.01.009 merceron, a., blikstein, p., & siemens, g. (2016). learning analytics: from big data to meaningful data. journal of learning analytics, 2 (3), 5. doi: 10.18608/jla.2015.23.2 mirriahi, n., & vigentini, l. (2017). analytics of learner video use. in c. lang, g. siemens, a. f. wise, & d. gasevic (eds.), handbook of learning analytics (pp. 49-57): society for learning analytics research. nguyen, q., rienties, b., toetenel, l., ferguson, r., & whitelock, d. (2017). examining the designs of computer-based assessment and its impact on student engagement, satisfaction, and pass rates. computers in human behavior, 76, 703-714. doi: 10.1016/j.chb.2017.03.028 papamitsiou, z., & economides, a. (2016). learning analytics for smart learning environments: a meta-analysis of empirical research results from 2009 to 2015. in j. m. spector, b. b. lockee, & d. m. childress (eds.), learning, design, and technology: an international compendium of theory, research, practice, and policy (pp. 1-23). london, uk: springer. persico, d., & pozzi, f. (2015). informing learning design with learning analytics to improve teacher inquiry. british journal of educational technology, 46(2), 230-248. doi: 10.1111/bjet.12207 prinsloo, p., & slade, s. (2017a). an elephant in the learning analytics room: the obligation to act. paper presented at the learning analytics & knowledge (lak), vancouver, canada. prinsloo, p., & slade, s. (2017b). ethics and learning analytics: charting the (un)charted. in c. lang, g. siemens, a. f. wise, & d. gasevic (eds.), handbook of learning analytics (pp. 49-57): society for learning analytics research. rienties, b., & toetenel, l. (2016). the impact of learning design on student behaviour, satisfaction and performance: a cross-institutional comparison across 151 modules. computers in human behavior, 60, 333-341. doi: 10.1016/j.chb.2016.02.074 ruppert, e., harvey, p., lury, c., mackenzie, a., mcnally, r., baker, s. a., & lewis, c. (2015). socialising big data: from concept to practice. working paper series (no. 138): centre for research on socio-cultural change. scholes, v. (2016). the ethics of using learning analytics to categorize students on risk. educational technology research and development, 64(5), 939-955. doi: 10.1007/s11423-016-9458-1 selwyn, n. (2015). data entry: towards the critical study of digital data and education. learning, media and technology, 40(1), 64-82. doi: 10.1080/17439884.2014.921628 selwyn, n. (2016). education and technology: key issues and debates. london: bloomsbury publishing. sharples, m., de roock, r., ferguson, r., gaved, m., herodotou, c., koh, e., . . . wong, l. h. (2016). innovating pedagogy 2016: open university innovation report 5. milton keynes: the open university. siemens, g. (2013). learning analytics: the emergence of a discipline. american behavioral scientist, 57(10), 1380-1400. doi: 10.1177/0002764213498851 slade, s., & prinsloo, p. (2013). learning analytics: ethical issues and dilemmas. american behavioral scientist, 57(10), 1510-1529. doi: 10.1177/0002764213479366 suthers, d., & verbert, k. (2013). learning analytics as a "middle space". paper presented at the proceedings of the third international conference on learning analytics and knowledge, leuven, belgium. tempelaar, d. t., rienties, b., & giesbers, b. (2015). in search for the most informative data for feedback generation: learning analytics in a data-rich context. computers in human behavior, 47, 157-167. doi: 10.1016/j.chb.2014.05.038 tempelaar, d. t., rienties, b., mittelmeier, j., & nguyen, q. (2017). student profiling in a dispositional learning analytics application using formative assessment. computers in human behavior, 78, 408-420. doi: 10.1016/j.chb.2017.08.010 van leeuwen, a., janssen, j., erkens, g., & brekelmans, m. (2015). teacher regulation of cognitive activities during student collaboration: effects of learning analytics. computers & education, 90, 80-94. doi: 10.1016/j.compedu.2015.09.006 winne, p. h. (2017). leveraging big data to help each learner upgrade learning and accelerate learning science. teachers college record, 119(13), 1-24. appendix 1: thematic analysis codes hämäläinen et al publication frontline learning research vol.6 no.3 (2018) 204 227 issn 2295-3159 it’s not only what you say, but how you say it: investigating the potential of prosodic analysis as a method to study teacher’s talk raija hämäläinena, bram de weverbteija waaramaac, anne-maria laukkanenc, joni lämsäa a university of jyväskylä, finland b ghent university, belgium c university of tampere, finland article received 13 may 2018 / revised 22 october/ accepted 28 november / available online 19 december abstract in this study, we introduce new insights into prosodic analyses as an emerging method to study what happens in classrooms interactions. we claim that the prosodic aspects (features of speech such as intonation, volume and pace) of talk are important, but under-represented in the learning sciences. these prosodic aspects may be used to complement, intensify or even reverse the linguistic content of speech. thus far, most research on classrooms has focused on the content (what is said) rather than on understanding the meaning of the prosodic features (how it is said) of talk. in this study, we introduce prosodic analyses as a method to study classroom discussions. our exploratory experiment focuses on the prosodic perspective of teacher’s talk to shed light on classrooms interactions. we present a case in which we align prosodic features with the content of teacher's talk during a nine-week physics course. this article shows that prosodic analyses may have added value for research on learning and professional development. namely, we illustrate that acting in an authentic classroom setting might trigger specific prosodic aspects in teacher's talk. we further found indications that the teacher applied different voice prosody regarding certain patterns of classroom talk. for the future, we suggest that a combination of content and prosodic analysis is a promising tool for gaining new insights into classroom interactions. keywords: teacher’s talk; prosodic analyse info corresponding author mail raija.h.hamalainen@jyu.fi doi: https://doi.org/10.14786/flr.v6i3.372 1. introduction multiple methods and techniques are required to understand what happens in classrooms, and while many researchers have investigated the content of talk – for example, with content analysis (de wever et al., 2006; hämäläinen & de wever, 2013) and discourse or conversation analysis (mercer & dawes, 2014; warwick, vrikki, vermunt, mercer, & van halem, 2016) – few researchers have attempted to understand the prosodic features of talk (gweon et al., 2013). the present study methodologically bridges the gap between two research domains to advance research on classroom discussions via the analysis of prosodic features. this analysis focuses on elements such as intonation and pitch on teacher’s talk in an authentic classroom context from a sociocultural perspective. 1.1 teachers’ talk in the classroom: educational dialogues and teacher monologues according to mercer, dawes, and staarman (2009), authentic classroom situations typically involve non-dialectic teacher monologues and educational dialogues. in non-dialectic situations, typically only the teacher talks (teacher monologue). in classroom contexts, in addition to educational dialogues, non-dialectic situations may be necessary and an intriguing way to stimulate learning. on the other hand, nowadays, the role of the teacher is changing from only providing knowledge to also supporting students’ knowledge construction activities. during educational dialogue – also referred to as productive classroom talk, which has an analogous meaning (muhonen et al., 2017) – both students and teachers talk. according to mercer (1995), classrooms set up ‘sceneries’ of educational dialogues where, ideally, teachers and their students will collaboratively discuss the topic about which they are learning. in these educational dialogues, teachers engage their students in discussions that include a series of questions and answers. for sociocultural research, the creation of meaning is inherently an intrapersonal process, and ways of thinking are embedded in particular ways of using language (littleton & mercer, 2013). dialogue may thus be said to be more than ‘just talk’ (o’connor & michaels, 2007). dialogue is talk that is productive and, therefore, should be the central interest of analysis (grounded on the work of vygotsky, 1987). according to muhonen et al. (2017), previous studies on learning have indicated that the quality of teacher-student dialogue is associated with the growth of students’ understanding (alexander, 2001; lemke, 1990; mortimer & scott, 2003). as a direct result, the field particularly needs to understand teacher talk during educational dialogues that creates opportunities to promote learning (see also mortimer & scott, 2003). educational dialogues are influenced first by differential power relations between teachers and students (lemke, 1990) and second by differential knowledge relations, for example by being either the ‘primary’ knower (typically a teacher) or a ‘secondary’ knower (typically a student) regarding the topic under discussion (berry, 1981). the teacher also plays a special role in guaranteeing that students benefit from classroom activities (nassaji & wells, 2000). some scholars have argued that in (science) classrooms, educational dialogues are likely to follow triadic dialogue patterns (lemke, 1990), especially during whole-class discussions (lemke, 1990; mehan, 1978; mortimer & scott, 2003; salloum & boujaoude, 2017). for example, according to nassaji and wells (2000), educational dialogues typically proceed along an initiation-response-feedback (i-r-f) pattern, which includes three phases. first, the teacher initiates a question (usually with a known answer); second, one or more students respond to that question; third, the teacher evaluates the answers, provides feedback and may or may not ask for follow-up questions or activities. wells (1993) further highlights that the students’ responses are a crucial element of i-r-f, since without such responses, there is no exchange (dialogue). sinclair and coulthard (1975) have also called the third stage (feedback) ‘follow-up’ and further define three types of follow-up acts: (1) accept or reject, (2) evaluate and (3) comment, which includes exemplifying, expanding, and justifying. 1.2 acoustic speech research from the research area of acoustic speech and voice research, it is known that prosodic features affect speech (later referred to as talk) perception and discussion. by prosodic features, we mean vocal characteristics like pitch variation and stress pattern, pausing, tempo, mean pitch and loudness, and vocal quality. prosody refers to the intentional or unintentional use of these characteristics to convey the meaning of an utterance (see the links in the appendix for more information on prosody). for example, prosody may signal one’s psychophysiological activity level and emotions. changes in pitch and/or loudness and tempo may reflect changes in activity or arousal level (vilkman & manninen, 1986; laukkanen et al., 1997; waaramaa et al., 2010; waaramaa et al., 2014). activity or arousal level can be low, moderate, or high. typically, a positive emotion of joy and a negative emotion of anger have a high arousal level, while tenderness and sadness have a low arousal level. a high arousal level is typically expressed by high pitch and loudness and a firmer (more pressed, tense) voice quality (laukkanen et al., 1997; waaramaa et al., 2010). in a low arousal level, the mean pitch and loudness are lower, and the voice quality is less firm (softer, laxer). the valence of the emotion (i.e. whether it is positive, negative, or neutral) may be conveyed by a complex combination of features including voice timbre (e.g. a brighter voice is associated with a more positive emotion than a darker voice, possibly because smiling makes the voice timbre brighter) (laukkanen et al., 1997; waaramaa et al., 2006). 1.3 analysing the prosody of teacher talk from the methodological perspective to understand what happens in classroom, we have to take into account that perceptions of paralinguistic and nonverbal characteristics of talk are, to a large extent, subconscious (see e.g. zald, 2003). these subconscious characteristics influence discussion processes, as according to brazil (1978), intonation in discourse illustrates the interaction. various studies have shown the effect of prosodic aspects on the listeners’ opinion of the speaker and perceptions of his/her personality (e.g. addington, 1968; lukkarila et al., 2012; scherer, 1972; zellner keller, 2004). prosodic (or suprasegmental) features intensify what is said or add meaning to the segmentals (phonemes). additionally, several studies have indicated that prosodic features can even reverse the meaning of a message (see e.g. laver, 1991; lehiste, 1970; scherer & giles, 1979) and some studies have shown that when linguistic-semantic content and prosodic-paralinguistic content are contradictory, the latter usually wins (lyons, 1977). therefore, we need a better understanding on what kind of role prosodic aspects play in teacher’s talk. divergent intonation patterns are also known to be used to praise and encourage students, to minimise the embarrassing effect of a wrong answer and to open and close discussions (hellermann, 2003). furthermore, intonation may be used to mark text cohesion (e.g. halliday & hasan, 1976), which improves speech comprehension. studies have also shown that a teacher’s dysphonic voice quality (e.g. irregularities in the sound signal, perceived as vocal fry or hoarseness) negatively affects students’ comprehension of instruction (see imhof et al., 2014; lyberg åhlander et al., 2014; rogerson & dodd, 2005). on the other hand, the classroom activity is also affected by environmental factors, like the size of the student group and classroom acoustics. for instance, a noisy environment requires a higher volume. consequently, a teacher may unintentionally modify other prosodic aspects, like intonation or voice quality, which may restrict the natural use of prosodic variation in conveying content or even provoke contradictory connotations. this may influence teacher’s talk and is another reason to investigate possibilities and limits of prosodic analysis as an approach for gaining insights into classroom talk. 1.4 aims the motivation for this article is that the content features of talk might not be sufficient for describing and understanding the true nature of classroom activities. therefore, methodological development is needed. this article aims to identify novel methods for studying teacher talk, combining both content and prosodic perspectives in this analysis. we concretize the methodological approach in light of two selected physics lessons. special attention will be paid to the applicability and restrictions of the selected methods for analysing the features of talk. to illuminate the method, we advanced two research questions: (rq1): how were the prosodic features of teacher talk influenced by the contextual factors of the authentic classroom, such as noisy classroom conditions? (rq2): how did the teacher’s use of prosody vary between different kinds of talk patterns? 2. method 2.1 context and data the present work is an exploratory case study based on data obtained in an authentic classroom setting. twenty-seven seventh-grade students and a teacher worked in a computer-supported inquiry science classroom during a nine-week physics course (27 hours of teaching and studying). one researcher observed the lessons and took ethnographic field notes during classroom observations (derry et al., 2010). based on these observation notes, we selected two lessons that included different types of teacher’s talk (see, next section) for future analysis. the two respective 45-minute videos of physics lessons from which both the teacher’s and students’ dialogue were transcribed, served as data for the present study. the teacher played a central role in planning the practical organisation of the project, and she was not given any specific instructions regarding her role as a teacher in the project. the teacher was fully responsible for implementing the instructional design without interference from the researchers. one video camera and three audio recording systems taped lessons. 2.2 analysing teacher’s talk in this study, we are interested in teacher talk in classrooms as a medium for pedagogy (e.g. kumpulainen et al., 2010). first, our analysis focused on how the prosodic features of teacher talk were influenced by the contextual factors of an authentic classroom (rq1). this investigation is necessary because most prosodic research has been conducted in settings in which participants use their ‘natural’ voice, while teachers may use their voice differently in classroom conditions, as the classroom is a specific condition with rather high levels of noise due to students. second, we sought whether (and how) a teacher’s use of prosody varied between different kinds of talk patterns (educational dialogues and teacher monologues, rq2). in-depth qualitative analysis and descriptive statistics were used to analyse and interpret the teacher’s talk. the identified 110 episodes of educational dialogue (n=42) and teacher monologues (n=68) were analysed. frequency counts and illustrative qualitative analyses were combined to explore the teacher talk in detail. the data analysis of classroom talk was grounded in educational dialogues and teacher monologues; it was also adjusted to sociocultural discourse analysis (mercer & dawes, 2014; niemi, 2016). the talk was analysed sequentially, which means that each utterance in a selected sequence is understood and viewed in relation to the previous utterance in the ongoing discussion. according to linell (1998), analytical descriptions are thus oriented towards the dialectical achievements of the participants. we identified key episodes related to educational dialogues based on triadic dialogue (an i-r-f pattern, lemke, 1990) and teacher monologues. as nassaji and wells (2000) have noted, however, this basic structure of triadic dialogue can be used for many purposes, particularly because the nature of the feedback the teacher is providing may vary. in our analysis, we focussed on these variations of teacher’s feedback. we based this analysis on sinclair and coulthard’s (1975) follow-up acts – accept/reject, evaluate and comment – however, we further developed the analysis of follow-up moves. we argue that sinclair and coulthard’s (1975) follow-up moves may not fully account for the influence of feedback variations that are present in the current inquiry-based science classrooms. teacher talk has radically changed since 1975, especially in the context of inquiry-influenced science classrooms. while the role of the teacher was once that of knowledge provider, classroom talk today is more based on shared discussions in which teachers try to trigger and support their students’ knowledge construction activities. for example, whether a teacher accepts or rejects a student’s response will make a difference. we thus broadened the analysis of the i-r-f pattern and labelled follow-up moves as cumulative, promotive and disputational i-r-f patterns (see table 1 for more details) (see also an analysis of student-student collaboration: exploratory, cumulative and disputational talk, mercer & wegerif, 1999). table 1 the talk patterns used for coding transcribed data to form a meaningful unit of analysis, typically several utterances, both from the teacher and student, needed to be combined. consequently, there were cases in which sequentially analysed utterances had characteristics of two or more classes from table 1. in these situations, the coding of units of analysis was based on the teacher’s talk. in addition to three educational dialogue patterns, there were episodes of teacher monologues when only/mostly the teacher spoke. during teacher monologues, the teacher, for example, gave physics-related information without dialogic discussion with students. she also organised the groups to get them to work more effectively. these patterns of talk are referred to as teacher presentations and group organising, respectively. finally, there was ‘other’ talk. all these classes with their descriptions are presented in table 1. the coding was done with software for qualitative data analysis (atlas.ti). this coding (six classes, see table 1) allowed an examination of educational dialogues and teacher monologues, which enabled talk episodes to be identified as patterns for frequency counts. from the coded data, the durations of the units of analysis were determined as measured in seconds. subsequently, based on the coded frequency counts of teacher talk patterns were selected for the intonation analysis. we did not conduct an intonation analysis for ‘other’ talk because this group was so heterogeneous in its contexts, for example the teacher discussed with her colleague (without the students hearing it) a student who probably ran away from school. in summary, in our analysis we first identified talk episodes, such as (1) cumulative, (2) promotive, and (3) disputational i-r-f patterns in the educational dialogues, as well as (4) teacher presentation and (5) group organising in the teacher monologues. after that, we investigated how the prosody related to these patterns of talk was characterised and whether variations could be identified. the educational dialogues and teacher monologues took place in finnish, and the researchers translated the excerpts presented in this article. pseudonyms were used to report the results. our analysis revealed typical patterns of how the teacher used her voice regarding different classroom situations. we aimed to select representative excerpts of classroom talk. there are several reasons for selecting these specific excerpts. first, in line with how the prosodic features of teacher talk were influenced by contextual factors (rq1), we selected excerpts that illustrate prosodic challenges that emerged when teachers acted in authentic classroom settings and that may be a useful starting point for future studies. additionally, to illustrate how the teacher’s use of prosody vary between different kinds of talk patterns (rq2), excerpts were selected in accordance with the theoretical perspective regarding educational dialogues and teacher monologues. due to space limitations, only some examples of the talk are illustrated in detail. however, we do not claim that the episodes presented here are necessarily typical of the larger sample. rather, they were chosen in view of our aim to illustrate the method developed and show that prosodic analyses may be an interesting venue for further research. however, to increase reliability, the data excerpts and the analyses have been actively discussed within our research group. critical comments and joint analysis efforts have contributed to strengthening the validity of the empirical analysis. 2.3 acoustic analysis in the prosodic approach to our data analysis, we focused on pitch and intonation. additionally, voice quality was addressed, as it may change involuntarily as a consequence of environmental challenges. for readers who are unfamiliar with the measures within prosody research, we briefly explain the terminology in detail in the next section. readers who are familiar with this terminology can move directly to the section ‘analyses in the present study’. 2.3.1. introduction to acoustic analysis human voice production can be divided into three main parts: power source, vibrator and filter. airflow from the lungs provides the power source for vocal fold vibration. the vocal tract (space from the vocal folds to the lips and nostrils) acts as a filter, thereby colouring and amplifying the sound produced by vocal fold vibration. through articulation, we modify the vocal tract to produce speech. speech consists of linguistic content and paralinguistic cues (prosodic, suprasegmental). linguistic content refers to the words used, and paralinguistic cues are the way the words are expressed and what other sounds or modifications are included (laughter, crying, smiling). paralinguistic content contains prosodic elements. prosodic features are said to be suprasegmental, as they are properties of speech units larger than the individual segment. it is necessary to distinguish between the personal, background characteristics that belong to an individual’s voice (for example, one’s habits influencing the pitch range) and the independently variable prosodic features that are used contrastively to communicate meaning (for example, the use of changes in pitch to distinguish questions from statements). we can alter our vocal production in many ways. the human voice, like sound in general, has three basic characteristics: pitch (how high or low the sound is), loudness (how soft or loud it is) and quality (timbre/colour, i.e. whether the sound is dark or bright, or sounds tense or lax). in addition, sounds have duration (how long or short the signal is). prosodic characteristics consist of the manipulation of these aspects. thus, they include the average pitch and loudness used by the speaker, as well as variations in pitch and loudness during sentences. variations in pitch during a sentence are called intonation. variations of loudness during a word or a sentence are used to stress (highlight) the important parts of the message. the manipulations of temporal aspects in talk include altering the duration of sounds and talk tempo and talk rhythm, consisting of temporal aspects and pausing. besides these characteristics, we can change our vocal quality: we can speak in a tenser way or a laxer way, or we can use various types of vocal fold vibration: chest voice, falsetto or vocal fry. furthermore, prosodic aspects are used to complement, intensify or even contradict one’s conversational content. prosodic aspects also convey (even subconsciously or involuntarily) information about our psychophysiological state, with aspects like mood, emotion and attitude, as well as our physical status (age, gender, health etc.). voices can be studied acoustically. the characteristic of sound that is perceived as pitch originates from the fundamental frequency (f0). this, in turn, in human voice production, corresponds to the number of vocal fold vibrations per second. it is measured in hertz (hz); one vibration per second is 1 hz. the faster the vocal folds vibrate, that is the more vibrations they produce per second, the higher a pitch we hear. the average f0 of a male speaking voice is about 120 hz and a female voice about 200 hz; however, the frequencies (i.e. the pitch use) may be somewhat language and culture dependent (pépiot, 2013). below, in figure 1, we present an example of the pitch curve, or f0 curve, made from a sentence read by a female speaker: ‘to my chagrin (in finnish: ‘harmikseni’), i did not find the basket there any longer.’: in this example, the mean f0 of the low-pitched finnish female speaker is 165 hz, which corresponds to the e note (e3) on a musical scale. the highest peak (f0 maximum) at the beginning of the sentence is 308 hz (dis1 or dis4), and the lowest f0 at the end of the sentence is 122 hz (h or h2). thus, the pitch variation range in this sentence is approximately 16 semitones (a semitone, i.e. a half step, is the smallest interval in western tonal music). in general, peaks on an f0 curve are associated with sentence stress (emphasis). figure 1. an example of pitch curve, or the f0 curve, and analysis. x-axis: time (in seconds), y-axis: fundamental frequency (f0), pitch in hertz (hz, 1 hz is the inverse of the duration of one vocal fold vibration, thus telling how many vibrations fit in one second of time). the high peak that is clearly seen at the beginning results from sentence stress placed on the word ‘harmikseni’ (in english: ‘to my chagrin’). the main acoustic correlate of perceived loudness is the sound pressure level (spl), which is most often measured in decibels (db). voice quality, in turn, can be acoustically studied, for example, with spectrum analysis. there, the sound is divided into components (harmonics). these components are present simultaneously. how strong these components are in relation to each other affects how the voice sounds, that is, whether the voice’s timbre is bright or dark. in a bright voice, there are stronger components in the high-frequency range compared to a darker voice. the overall tilt of the spectrum tells how the voice is produced. the tilt is steeper in a soft or breathy voice (see figure 2). figure 2. an illustration of two examples of long-term average spectra (ltas) from two [a:] vowel samples from the same female speaker. the solid line describes ltas from a pressed voice, while the dotted line shows a breathy voice. the spectrum tilts more steeply in the sample with a breathy voice. this means that the harmonic energy declines faster as a function of frequency. the curves drawn on top of the spectrum (green for breathy and red for pressed) show this phenomenon. the higher the peaks are in the spectrum, the more energy there is in the expression, and the louder the voice sounds. 2.3.2. analyses in the present study in the present study, a fundamental frequency (f0) analysis was conducted to give physical correlates of pitch and intonation for sentences that were classified to represent the five teacher-talk patterns mentioned in table 1: cumulative, promotive, and disputational i-r-f patterns; teacher presentations; and group organising talk. the f0 analysis was performed using praat software (boersma & weenink, 2006, version 6.0.21). we studied the f0 curve both qualitatively and quantitatively. in the latter approach, we measured the mean, range and standard deviation (sd) of f0. the sd of f0 illustrates f0 variation in intonation more reliably than the f0 range, as the latter can be affected by unintentional voice quality-related matters (like the use of vocal fry with a very low f0) or random errors in the automatic f0 analysis. voice quality was illustrated through ltas and praat analysis. 3. results 3.1 prosodic challenges when teachers act in authentic classroom settings (rq1) our methodological approach offers possibilities to show prosodic challenges that emerged when teachers acted in authentic classroom setting. we found that when the teacher interacted with students in authentic (meaning often rather noisy) classroom conditions, she used a loud voice, which led to a heightened pitch. figure 3 illustrates this phenomenon in terms of an f0 curve. as we can see in figure 3, there is a part of teacher presentation, interrupted by a question from a student, followed by the teacher’s answer to the question. teacher: speaking of next week’s exam that we thought about today… student: so, when it will be next week? (in finnish: ’eli millon se on ens’ viikolla?’) teacher: on wednesday of next week. here, the teacher’s mean f0 is 296 hz (d1) in the beginning, which represents her general way of speaking to the whole class. at the end part of the curve, the mean f0 is 220 hz (a), as she answers a student’s question seemingly using her natural conversational volume. thus, at the beginning she raised her pitch by circa 5 semitones to speak to all the students. this may also restrict the habitual livelier use of pitch variation, as can be seen by comparing the beginning part of the f0 curve in figure 3 with the f0 curves seen in the other figures. in figure 4, we illustrate this exchange in terms of voice quality. the black line shows a spectrum of the part: ‘speaking of next week’s exam that we thought about today…’ in this part, the teacher speaks loudly, and her perceptual vocal quality is pressed (see e.g. kankare et al., 2012; waaramaa & kankare, 2013; waaramaa et al., 2014). the grey line shows the spectrum of the part: ‘on wednesday of next week.’ in this part, the teacher’s voice sounds more relaxed. thus, in pressed phonation, there is more sound energy at the higher frequency range than in the ordinary phonation type. the use of a pressed vocal quality poses more biomechanical load on the vocal folds than the use of an ordinary, relaxed voice. additionally, this example illustrates that a teacher seemed unintentionally modify prosodic aspects, which may restrict the natural use of prosodic variation in conveying content or even provoke contradictory connotations. for instance, a teacher may involuntarily sound angry when trying to get his/her voice heard over background noise. this, in turn, may influence how teacher’s talk is interpreted in classroom. figure 3. an example of how the teacher raises her pitch by circa 5 semitones (left part of this figure) when raising her voice to speak to the whole group of students (compared to when she is speaking to a single student, right part of the figure). figure 4. two long-term average spectra of teacher talk. y-axis: mean sound energy in db, x-axis: frequency in hz. the black line represents a pressed and loud voice (addressing the complete group of students), while the grey line represents a relaxed voice (talking to an individual student). 3.2 differences in a teacher’s use of prosody (rq2) in table 2, we summarise the f0 characteristics in the five patterns of talk studied (see table 1 for a description of the patterns). even though based on this small sample size, it is impossible to claim direct correspondences between the patterns of talk and intonation patterns, some typical features could be identified based on this exploratory case study. in general, cumulative i-r-f patterns seemed to use a moderate mean pitch level and a moderate pitch variation (see table 2), with frequent word stresses that were realised using the same pitch pattern (see figure 5). promotive i-r-f patterns showed a low mean pitch, a wide pitch range and large emphatic sentence stresses (high peaks in the f0 curve; see figure 6). disputational i-r-f patterns seemed to use moderate mean pitch level with large pitch variation and strong emphatic stresses (see figures 7 and 8). teacher presentation resulted in a high mean pitch level with small pitch variation (figure 9), while group organising was characterised by a relatively high mean pitch level with moderate pitch variation (figure 10). in the following sub-sections, our methodological approach is exemplified with empirical examples. we demonstrate how language was manifested in various talk patterns and what kinds of prosodic features were typical for each type of talk pattern. for each type of talk listed in tables 1 and 2, we first introduce a representative example of a talk episode, followed by a representative figure (graph of a prosodic phase) illustrating how the teacher’s intonation is used and varies. within the selected representative sentences, bold words refer to stressed words in the sentence. table 2. teacher’s f0 (pitch) variation and range in hz and in semitones. 3.2.1. teacher’s talk prosody in educational dialogues cumulative i-r-f patterns were rare and emerged nine times (8.2%) with a total duration of 333 seconds (12.2%). they involved speakers in pleasant, uncritical exchanges that built towards a common understanding through accumulated repetition and confirmation. from a prosodic perspective, cumulative i-r-f patterns were associated with a relatively narrow pitch range and frequent word stresses that were realised using the same pitch pattern. typically, the intonation pattern repeated itself, without extreme deviations from the mean pitch level. the following excerpt 1 of cumulative i-r-f pattern shows how the teacher cumulated dialogue. first, she conformed: ‘well, now, it’s here’, and she asked what material a jar of jam is made from. the student pekka responded. this was followed by a new question from the teacher and another response from pekka. subsequently, the teacher inquired what happens to materials when they are heated, and she accepted a trivial answer from joel, a student: ‘the material gets warmer’ (i.e. constructing positively). then, the teacher tries to remind the students of a video in which the same phenomenon was shown. in this way, the common knowledge is constructed further. the teacher repeats her student elvira’s answer that, when matter becomes warmer, it expands (conformation and repetition). figure 5 is selected from excerpt 1 to show a typical example of the flow of the cumulative i-r-f pattern. in this example, the mean f0 is 232 hz, and the sd is 46 hz, that is, seven semitones, and the total f0 range is 112–335 hz. furthermore, the graphic lines show that the intonation pattern repeats itself, without extreme deviations from the mean pitch level. between the vertical lines is a sentence: ‘expands (in finnish: ‘laajenee’), when matter becomes warmer (in finnish: ‘lämpenee’), it expands (in finnish: ‘laajenee’). now, when the jar is made of glass, the lid is made of metal. does it expand in the same way?’ excerpt 1: an example of cumulative i-r-f patterns teacher: well, now, it’s here. think about a jar; what material is a jar usually made of? a jar of jam. pekka: glass. teacher: and what material is the lid of the jar, typically? pekka: metal. teacher: think about glass and metal; when you put the jar under hot water, what happens to it? what happens to matter when it becomes warmer? when getting warmer, matter, what…? joel: it becomes warmer. teacher: becomes warmer, but at the same time …? quite in the beginning we made those, there was … you have the kind of a video there, with the hole and the bullet, and the bullet is heated. elvira: expands. teacher: expands, when matter becomes warmer, it expands. now, when the jar is made of glass, the lid is made of metal. does it expand in the same way? liisa: the lid gets larger. teacher: yes, the lid gets larger than the jar, so then we get it open. figure 5. an example of a cumulative i-r-f pattern. during promotive i-r-f patterns, the teacher engaged constructively with students’ ideas, trying to trigger productive collaboration (to trigger exploratory talk, see, mercer & wegerif, 1999). promotive i-r-f patterns were applied fairly actively (n=20, 18.2%; total duration of 11.4 minutes, 25%). the prosodic analysis illuminated that promotive i-r-f patterns were associated with a wider pitch range than cumulative i-r-f patterns and larger emphatic sentence stresses than the cumulative i-r-f patterns presented previously. the following excerpt 2 is a typical example of a discussion between the teacher and her students. the teacher asks her students to think of the three materials and infer which expands differently from the others. the teacher urges her students to consider a phenomenon and converse about it, and she offers an immediate verbal response to her students’ reactions. as we can see, the teacher engages students to actively encounter the phenomenon: ‘now, think of these three materials. can you infer which one of these expands differently from the others?’ when juuso’s response is correct, the teacher becomes excited and immediately gives him positive feedback: ‘correct! well done!’ the teacher continues to explore the issue and asks her students to make a hypothesis regarding what happens to metals when they are heated. in this case, the teacher tries to engage her students to consider and justify their answers (aiming to promote students’ exploratory talk). she also offers alternative hypotheses: ‘well, what happens to metals when they exp … become warmer? do they shrink, stretch or not change?’ after ilkka replies correctly, the teacher again provides positive feedback. figure 6 highlights the discussion between the teacher and a student. for figure 6, the mean f0 is 199 hz, and sd 55hz, i.e. 9.8 semitones, and the total range is 98–310 hz. we can see that promotive i-r-f patterns showed a wider pitch range and larger emphatic sentence stresses than the cumulative i-r-f patterns presented previously. furthermore, an emotional state of excitement involves a high arousal level that is seen in high f0 peaks in the intonation curve: ‘correct! well…’ (in finnish: ‘aivan! hyvin…’) the two high peaks in the pitch curve at the end of the sample reflect stressed words, expressing the teacher’s excitement when getting a correct answer from her student: ‘correct! well done!’ (in finnish: ‘aivan! hyvin pää(telty)!’) the last syllables in the parentheses are expressed in a whisper. the excerpt also shows how the intonation curve drops in a question in finnish (the first half of the picture). excerpt 2: an example of promotive i-r-f patterns teacher: if you think of these three materials, could you infer which one of those expands differently from the others? juuso: well, water. teacher: correct! well concluded! juuso: do i win a prize? teacher: nope. well, what happens to metals when they exp … become warmer? do they shrink, stretch or not change? ilkka: they stretch. teacher: good! the first page is completed. figure 6. an example of promotive i-r-f -pattern. finally, disputational i-r-f patterns were characterised by disagreements, by teacher disagreeing and showing critical approach to student response(s) and by short, often confrontational, interchanges from the students. it emerged 13 times (11.8%) with a total duration of 346 seconds (12.7%). figures 7 and 8 below illustrate that during disputational i-r-f -patterns, the pitch variation seemed to be the greatest and sentence stresses strongest. in the following excerpt 3, we can see how the teacher gets frustrated when trying to motivate the students to think about the insulation properties of a thermos bottle. the teacher asks many times if the students could explain why the inner surface of the thermos bottle is glossy without giving them sufficient resources. there is also evidence that the students are not listening actively to the teacher, as markus asks regarding the glossy surface: ‘on the inside, you mean?’ even though the teacher has consistently been talking about the inner surface of the thermos bottle. finally, when markus responds to the question, the teacher disagrees with him: ‘that isn’t enough, markus, that it keeps the drink warm.’ in practice, the teacher demands more information from markus. of the interactive patterns, pitch variation seems to be largest and sentence stresses strongest for disputational i-r-f -patterns. for figure 7, the mean f0 is 230 hz, and sd 66 hz, i.e. 10 semitones, and the total range is 84–346 hz. in figure 7, we can also notice that the sentence stress is on the word ‘enough’ (in finnish: ‘riitä’), which is indicated by the highest peak in the intonation curve: ‘that isn’t enough (in finnish: ‘toi ei riitä...’), markus, that it keeps the drink warm.’ when comparing figure 7 to figure 6, high f0 peaks can be seen in both figures reflecting a high arousal level in the teacher’s talk. however, in figure 6, a positive emotion of excitement was expressed and in figure 7 a negative emotion, perhaps frustration considering the content of the teacher’s talk. thus, the emotional valence (positive, negative or neutral) cannot be directly concluded from the acoustic cues of the arousal level e.g. related to intonation (bänziger & scherer 2005). figure 8 below illuminates how this disputational i-r-f pattern continues with similar significant prosody variation. here, the teacher is still not happy with the students’ study process, and she is still demanding more from the students. in the episode represented in figure 8, the teacher’s mean f0 was 210 hz, the sd was 46 hz, i.e. ca 8 semitones, and the range was 122–317 hz. thus, in the continuation of disputational i-r-f pattern the pitch variation continues to be great and sentence stresses strong: ‘it is kind of a fact (in finnish: ‘fakta’) why a vacuum bottle is used (in finnish: ‘käytetään’), but you should now explain why the glossy (in finnish: ‘kiiltävä’) interior is helpful there.’ the very low f0 values in the total range reflect the use of vocal fry phonation, especially in sentence endings. excerpt 3: an example of disputational i-r-f -patterns jan: a thermos bottle is a bottle that is heat insulated from its environment as much as possible. teacher: yes. does it explain why the inner surface is made glossy? it is one of the many ways by which it is made a good insulation, the structure of the whole bottle, but… (off-topic discussions between the students) teacher: yes, but why does the glossy surface keep it warm unlike a red surface, for example? jan: i don’t know. teacher: and you cannot find any explanation, can you, if you go and study the material there? [miscellaneous noise] teacher: hey – now! you were supposed to think now of the glossy inner surface of a thermos bottle. what might be [i know] the reason? well? markus: on the inside, you mean? teacher: uhum, on the inside, yes. have you ever looked into a thermos bottle? [miscellaneous noise, the teacher’s comments to one student.] think; search for information. it can be found there in the heat transfer mechanisms section. teacher: you can’t search for anything, can you? markus: me? yes, watch out; it’s coming soon…keeps warm…. teacher: that isn’t enough, markus, that it keeps the drink warm. it’s like a fact why people use a thermos bottle, but you should now explain why the glossy surface helps there. so, because what...? markus: i don’t know. teacher: well, it cannot be necessarily found from wikipedia now. [miscellaneous noise] teacher: it is kind of a fact why a vacuum bottle is used, but you should now explain why the glossy interior is helpful there. figure 7. an example of disputational i-r-f -pattern figure 8. another example of disputational i-r-f -pattern, a continuation from figure 7. 3.2.2. teacher’s talk prosody in teacher monologues the teacher’s talk initiated by a student’s question or her talk at the beginning of different sections of the lesson was labelled as teacher presentation. the teacher gave general instructions to the students, so they could start working on their tasks, the teacher interrupted the students’ working so they could begin reviewing the correct answers to the problems, or the teacher gave a short lecture about the theory when students faced challenges while solving problem; these are typical examples of when talk classified as teacher presentation emerged. in general, it can be said that when there was a need for straightforward instruction (mercer, 1995) the teacher’s talk had characteristics of teacher presentation. overall, teacher presentation was a common type of talk, comprising one fourth of the total duration of different talk patterns (n=27, 24.5%; total duration 11.5 min, 25.3 %). as an example, in the following excerpt 4 the teacher does not provide physics-related information, but rather general instruction to the whole group about the lesson plan of the day. she also gives general feedback to the students about combining information from different sources when taking their previous exam. there are no attempts (e.g. questions) to invite students to take part in this discussion, so the conversation can be referred to as teacher monologue. figure 9 displays an example of teacher presentation talk: ‘there were some minor difficulties in the answers (in finnish: ‘pieniä ongelmia vastauksessa’) to the exam last week (in finnish: ‘viime viikon’). you were not able to combine (in finnish: ‘osannu yhdistellä’) pieces of information …’ the mean pitch of the teacher’s talk is relatively high (290hz), and there are no wide changes in f0 during intonation. the range of f0 variation comprises frequencies from 193 to 374 hz, and the sd is 37 hz, i.e. about 4.4 semitones. excerpt 4: an example of teacher presentation teacher: there were some minor difficulties in the answers to the exam last week. you were not able to combine pieces of information (or) find it, so i think that we will practice a little for the exam. there are similar types of questions to those that will be on the exam, so let’s take a look (at these). first, go through (the problems) yourself or with a pair or a group, and think about how you would answer. together, try to find what kind of answer would be good when you have to combine pieces of information from many resources now. figure 9. an example of teacher presentation talk. in addition to teacher presentation, there were 31 units of analysis (28.2 %) belonging to group organising, but the total duration (9.0 min, 19.8 %) of those units was not that high, mainly because excerpts were usually short remarks and comments from the teacher to the students relating to their behaviour and studying methods. group organising, like teacher presentation, was characterised by teacher monologue. in general, there was a need for group organising at regular intervals when students solved the problems themselves (6.5 min, 25.4 %, cf. teacher presentation: 3.9 min, 15.4 %). when the section of the lesson was more teacher-led by nature (the students started to go through the correct answers with the teacher), the teacher had to guide the students less frequently to concentrate on the teaching (2.5 min, 12.5 %, cf. teacher presentation: 7.6 min, 38.0 %). in the following excerpt 5, the teacher explains to the students that the problems they are going to solve are a rehearsal for the exam. when a student points out that he did not get a problem sheet, the teacher says, with a twinkle in her eyes, that some of the students may have taken more than one (same) problem sheet, as there were not enough papers in the stack. even though there are utterances both from the teacher and the student, there is no typical triadic dialogue visible, and the conversation can be referred to as teacher monologue without true collaboration between the participants. the main motive of the teacher was to get the problem sheets for everyone so the students can start revising physics-related issues for the exam. a situation given below illustrates the use of prosody during group organisation. in figure 10, we can see how the teacher says that the exercise ‘is a rehearsal (in finnish: ‘harjoittelua’) for the exam’, and a student responds that ‘i didn’t get one’ (in finnish: ‘mä en saanu.’) then, the teacher answers that ‘there should be (more) in the stack (in finnish: ‘siin pinos’), so perhaps somebody took more than one. see, the most hard-working (in finnish: ‘ahkerimmat’) students do two (papers).’ as we can see from figure 10, the teacher’s mean f0 was 248 hz, the absolute minimum 56 hz representing vocal fry (creaky sound), 175 hz was the lowest f0 without vocal fry, and the maximum was 419 hz. thus, the total frequency range was 15 semitones, i.e. 1 ¼ octaves. the sd of f0, which reflects the intonation range more reliably, was 55 hz, i.e. approximately 8 semitones. excerpt 5: an example of group organising teacher: this is a rehearsal for the exam. jenna: i didn’t get (one). teacher: there should be (more) in the stack, so perhaps somebody took more than one. see, the most hard-working students do two (papers). figure 10. an example of group organising. 4.discussion the present study is a first step towards developing a new method of analysing classroom interaction. we investigate the potential of prosodic analyses of teacher talk. our argument is that analysis of prosodic features has been underrepresented when classroom talk is analysed. we claim that in addition to analysing the content of talk (focusing on what is said), analysing the prosodic features of talk (focusing on how something is said, thus considering elements as intonation, volume, and pace) is also important. in some cases, it might be as important as – or even be more important than –what is actually said. for example, a teacher asking a student, ‘what do you think about this?’ may be simply inquiring for a student’s opinion. however, the same question – in exactly the same wording – can be conveyed in such way (by changing intonation and stressing other words) that a student really feels involved in the discussion process and appreciates being invited to share his/her ideas. at the same time, exactly the same words can be pronounced in such a way that the student feels threatened and reprimanded for not listening attentively. in this article, the methodological development grounds on a notion that voices can be studied acoustically. we focused on teacher’s talk in an authentic classroom, and two research questions were addressed in relation to the general aim of investigating the potential of prosodic analysis. knowing that a classroom is far from a laboratory setting, we investigated how the prosodic features of teacher talk were influenced by the contextual factors of the authentic classroom (i.e. an often quite noisy environment) (rq1). with our methodological approach, we were able to illustrate some specific prosodic challenges. our findings showed that when the teacher acted in the authentic classroom setting, she often used her voice in a different way. the results show that when addressing the complete classroom, her voice was more raised, resulting in a more pressed voice (indicated by a higher pitch) than in other occasions, such as talking to the student in a one-to-one way or to small group of students when guiding them. in the latter situation, the voice was more relaxed and thus closer to her natural voice. we argue that how teachers use their voice may have an influence on teachers’ health, teacher-student interaction, and classroom climate. firstly, the risk of vocal fatigue increases when using a pressed voice (e.g. kankare et al., 2012). secondly, a pressed voice quality is related to the expression of anger (e.g. laukkanen et al., 1997; waaramaa et al., 2010; 2014). therefore, speaking in a large and noisy classroom may lead to involuntary and misleading prosodic characteristics. these characteristics can be interpreted as shouting in anger, which may affect negatively on teacher-student interaction and classroom climate (see, 3.1). this may be disconcerting as a high-quality teacher–student interaction and a supportive classroom climate is one possible protective factor against the negative impacts of learning (kiuru et. al., 2012). the second research question focused on studying how the teacher’s use of prosody varies between different kinds of talk patterns (rq2). therefore, we first identified talk episodes, such as (1) cumulative, (2) promotive, and (3) disputational i-r-f patterns in the educational dialogues, and (4) teacher presentation and (5) group organising in the teacher monologues. next, we checked how the prosody related to these patterns of talk was characterised and whether differences could be identified. we found that cumulative i-r-f patterns seemed to use less pitch variation and word stress patterns were often repeated here. on the opposite, a wide pitch range and clear emphatic sentence stresses with large f0 jumps characterised disputational as well as promotive i-r-f patterns. thus, the intonation pattern in cumulative i-r-f patterns seems to reflect continuation, while the strong emphatic stresses with relatively wide pitch intervals marks a contrast e.g. between the student’s answer and the teacher’s instruction. in promotive i-r-f patterns, some strong accents with high pitch peaks were used to acknowledge correct answers and to give support. regarding teacher monologues, teacher presentation used a high mean pitch, a narrower pitch variation and a more pressed voice quality, while group organising was characterised by a relatively high mean pitch level with moderate pitch variation. in sum, by combining the prosodic and content characteristics of teacher’s talk, we were able to identify initial variations in how the teacher used her voice in diverse educational dialogues and teacher monologues. 4.1 limitations and critical issues the strength of this study is that, along with studying the content of the talk, it pays attention to the potential offered by the prosodic perspective of teacher’s talk that has rarely been explored to date. however, there are several limitations and critical issues to consider as this study was an initial attempt to illustrate how the teacher’s intonation varies depending on the situation. first, this study is exploratory in nature, and although we were able to show that different prosodic characteristics are somehow related to distinctive patterns of talk content-wise, additional explorative and hypothesis-testing research is needed to analyse this relationship more specifically. second, as this case study like case studies in general is based on a small sample, all limitations thereof should be duly considered. moreover, there are three more limitations to our study that are related to the use of prosodic analysis in general in this type of research settings. the third limitation concerns the use of acoustic speech methodology in an authentic classroom setting and comprises three problematic aspects: (1) although the technology is available, it is not necessarily easy to get the hardware needed (especially when using it on a larger scale) and to use this hardware to capture voices in classrooms without compromising the authenticity of the setting; (2) the teachers may tend to use their voice in different ways depending on the specific conditions within the classroom; and (3) authentic classroom conditions may hamper the quality of audio recordings and thus limit the usability of the method. the fourth and fifth limitations also pertain to acoustic speech research in general, both of which may make it more difficult to establish a catalogue of normative data on classroom interactions. the fourth limitation is that the interpretation of how voice is used might be culturally bound (waaramaa, 2014; waaramaa & leisiö, 2013), and this aspect was not considered in this study. on a general level, this limitation might also make it more difficult to compare findings from different classrooms around the world. the fifth limitation is that language specificity might form another barrier for the comparability of the research findings. specifically, when studying collaboration, language is always the central aspect under investigation, and with regard to prosodic analysis, specific features of different languages may have specific characteristics (see, method section). therefore, we briefly discuss the specifics of intonation patterns and features of finnish language (compared to other languages) in the remainder of this paragraph. in finnish, sentence stress and intonation do not serve linguistic purposes to the same extent as e.g. in swedish or english, as finnish takes advantage of enclitics. the intonation curve is typically declining in statements, and a relatively smooth and high intonation pattern is used to express continuation (aaltonen & wiik, 1979). a pitch rise in sentence endings has been regarded as untypical for finnish language, even though lately it has become a characteristic of teenagers’ talk (routarinne, 2003 a and b; härkönen, 2016). in general, finnish talk has been described as characterised by soft phonation, a low mean pitch, small intervals in intonation and a relatively ‘tame’ expression of emotions (hakulinen, 1979). on the other hand, despite the differences between the languages, prosody has been analysed in other fields, e.g. the therapist–patient dialogues (e.g. leszcz, 2017). from this research, we know that differences between the languages can be considered and dealt with. in the present study, we investigated the intonation pattern of a teacher’s talk and our findings can be considered to be in line with earlier results e.g. reported by o’connor and arnold (1973) for english language. however, even though we can take differences between languages into account, it could be interesting to explore the value of prosodic analysis in view of analysing teacher talk in different languages. 4.2 directions for future research in this section we put forward many opportunities that this new method may bring, and we relate this to some elements to be further developed and explored in future research. first of all, we see the potential for developing our methodological approach towards (semi-)automatic analysis of audio and video data. in a first phase, a possible methodological application could be to identify interesting discussion phases based on the prosodic features, which can then be further analysed and interpreted by educational researchers. based on the results of this explorative study, we are optimistic that using prosodic analysis in such a semi-automatic way is a likable future application, and even if it means that the identified phases still need to be interpreted by researchers, this is a promising venue, as often researchers have an enormous amount of data, and thus being able to use prosodic analyses to pre-process and reduce this amount of data for manual coding would be a useful application. in a second phase, future research could focus on investigating whether it is possible to move to fully automatic analyses of talk, based on prosodic analyses. a second opportunity and direction for further research is to broaden the scope to also analyse students’ talk. while our study focused on teacher’s talk, we suggest that next methodological step should be taken by combining prosodic and content analysis to study students’ talk, and more specifically student–student dialogues that are happening as a part of collaborative learning. in this respect, prosodic analyses could be applied to identify different types of talk or collaboration on the one hand, while on the other hand it could be used to capture students’ (and also teachers’) emotions. earlier research in the field of therapist–patient dialogues has shown that in addition to verbal, non-verbal, para-verbal, implicit and explicit communication, prosodic analyses are useful in capturing and predicting emotions (leszcz, 2017). the role of emotions has often been underestimated and could be of great importance (isohätälä et al., 2017). what is particularly important is the consideration of how, when and why students’ emotions arise and how they shape interaction (student-student/teacher-student) and affect students’ dedication towards collaboration and learning. this may be associated what and how is said in the classroom context (e.g. our results about teacher involuntary sounding angry). recently, positive activating emotions have been shown to be related to good academic success (postareff et al., 2017), so being able to capture and analyse emotions through prosodic analyses – and in a next phase do this ad hoc, on the fly and provide teachers with this information through a learning analytics powered dashboard – could be very interesting and valuable future application. related to this, a third opportunity that we put forward is the development of new digital tools to support teachers (see also, harteis, 2018). automatic prosodic analyses could be an interesting feature to inform teachers of students’ collaborative discussions, e.g. by signalling group processes to the teachers. if students’ voices could be interpreted on the fly, data from these analyses could be used to create process indicators in an automatic way. by adding information from automated prosodic analysis, existing tools could be extended. as an example, we can think of how a lantern device (dillenbourg et al., 2011) could be fed by prosodic data. the lantern device of dillenbourg and colleagues (2011) is a small device with leds that is controlled by students to allow them to indicate which exercises or phase of a collaborative activity they are working on (i.e. by changing the colour of the lamp), and if they have questions for the instructor, it allows them to signal this to the instructor (i.e. by making the lantern blink). the blinking rate is furthermore increased over time, allowing the instructors to see how long students have been waiting for them (for more details, we refer to dillenbourg et al., 2011). the goal of the devices was to provide the instructors with some awareness of the teams’ behaviour. in their implementation, students controlled the tool themselves, but in future extensions, based on automated on-the-fly analyses of the prosodic features of students’ collaborative discussions, the tool could provide additional useful information about collaboration processes for instructors. finally, a fourth opportunity is allied to teacher training and teachers’ professional development. being able to capture, interpret and understand students’ emotions on the fly while engaged in technology-enhanced collaborative learning may be helpful for teachers’ professional development. this is also related to the question of how emotional valence — whether positive, neutral, or negative — can be derived from the teacher’s talk. typically, vocal emotions are studied first for their arousal level and second for their valence. in the present investigation, we concentrated on arousal level, displayed by intonation curves. in our future research, we will scrutinise teacher’s vocal expression of valence, how the teacher uses his/her voice to convey emotions related to the content of the talk, e.g. when encouraging the students, when expressing contentment or disappointment, and how valence expressed is associated with teacher’s talk. in this respect, research needs to focus on triangulating data resources. so far, there is research available focusing on physiological measures of emotions (e.g. with the smart rings, e.g. http://www.moodmetric.com) or heart rate variability measures (see e.g. https://www.firstbeat.com/en/) and self-report measures of emotions (see oksanen & hämäläinen 2010; castellar et al., 2014). we argue that an application could be to add prosodic analyses as a method in combination with these methods, as another source to triangulate from. to conclude, there is a current trend of exploring more advanced methods to capture social, cognitive, and emotional features of classroom talk, as these novel approaches are needed to meet the analytical challenges of making sense of the processes of learning and instruction (damsa & ludvigsen, 2016). the present exploratory study can in this view be seen as one contribution. we showed that acting in an authentic classroom setting might trigger specific prosodic aspects in the teacher's talk. additionally, we were able to identify differences in how the teacher used her voice and relate those to diverse educational talk patterns. we believe that prosodic analyses may be one novel approach that allows us to understand learning and instruction processes better. keypoints multiple methods and techniques are required to understand what happens in classrooms prosodic aspects (features of speech such as intonation, volume, and pace) of talk are under-represented in the field of the learning sciences we introduce prosodic analyses as a method to study teacher talk in classroom we showed that the teachers’ prosody varied depending on different patterns of talk that were identified based on the content. this article shows that prosodic analyses may have an added value for research on learning and professional development acknowledgments this work was supported by the academy of finland under grant numbers 292466 and 318095 [the multidisciplinary research on learning and teaching profile of jyu] and by the emil aaltonen foundation and the finnish cultural foundation. references aaltonen, o., & wiik, k. (1979). (1979). suomen jatkuvuuden intonaatiosta. in p. hurme. (eds.) jyväskylän yliopiston suomen kielen ja viestinnän laitoksen julkaisuja, 18 1. fonetiikan päivät (the first finnish phonetics symposium), (pp. 23-33). alexander, r. j. (2001). culture and pedagogy: international comparisons in primary education (pp. 391-528). oxford: blackwell. addington, d. w. (1968). the relationship of selected vocal characteristics to personality perception.speech monographs, 35(4), 492-503. berry, m. (1981). systemic linguistics and discourse analysis: a multi-layered approach to exchange structure. studies in discourse analysis, 1, 20-145. bänziger, t., & scherer, k. r. (2005). the role of intonation in emotional expressions.speech communication, 46(3), 252-267. boersma, p., & weenink, d. (2006). praat: doing phonetics by computer. brazil, d. c. (1978). discourse intonation ii, discourse analysis monographs ii(1st ed.). birmingham: university of birmingham, english language research. castellar, e. n., oksanen, k., & van looy, j. (2014). (2014). assessing game experience: heart rate variability, in-game behavior and self-report measures. in anonymous (eds.) 2014 sixth international workshop on quality of multimedia experience (qomex) (pp. 292-296). de wever, b., schellens, t., valcke, m., & van keer, h. (2006). content analysis schemes to analyze transcripts of online asynchronous discussion groups: a review.computers & education, 46(1), 6-28. derry, s. j., pea, r. d., barron, b., engle, r. a., erickson, f., goldman, r., . . . sherin, b. l. (2010). conducting video research in the learning sciences: guidance on selection, analysis, technology, and ethics. journal of the learning sciences, 19(1), 3-53. hakulinen, l. (1979). suomen kielen rakenne ja kehitys(4th ed.). helsinki, finland: otava. halliday, m. a. k., & hasan, r. (1976). cohesion in english. london: longman. harteis, c. (2018). machines, change, work: an educational view on the digitalization of work. in c. harteis (hrsg.), the impact of digitalization in the workplace – an educational view (s. 1-10). dordrecht: springer. hämäläinen, r., & de wever, b. (2013). vocational education approach: new tel settings—new prospects for teachers’ instructional activities? international journal of computer-supported collaborative learning, 8 (3), 271-291. härkönen, r. (2016). tilanteen vaikutus 14-vuotiaiden puheen akustisiin ja perkeptuaalisiin piirteisiin (acoustical and perceptual analysis of the situational effect on 14 year-olds' speech) hellermann, j. (2003). the interactive work of prosody in the irf exchange: teacher repetition in feedback moves.language in society, 32(1), 79-104. imhof, m., välikoski, t., laukkanen, a., & orlob, k. (2014). cognition and interpersonal communication: the effect of voice quality on information processing and person perception. studies in communication sciences, 14(1), 37-44. isohätälä, j., järvenoja, h., & järvelä, s. (2017). socially shared regulation of learning and participation in social interaction in collaborative learning. international journal of educational research, 81, 11-24. kankare, e., laukkanen, a., ilomäki, i., miettinen, a., & pylkkänen, t. (2012). electroglottographic contact quotient in different phonation types using different amplitude threshold levels. logopedics phoniatrics vocology, 37(3), 127-132. kiuru, n., poikkeus, a. m., lerkkanen, m. k., pakarinen, e., siekkinen, m., ahonen, t., & nurmi, j. e. (2012). teacher-perceived supportive classroom climate protects against detrimental impact of reading disability risk on peer rejection. learning and instruction, 22(5), 331-339. kumpulainen, k., & lipponen, l. (2010). productive interaction as agentic participation in dialogic enquiry. in k. littleton, & c. howe (eds.), educational dialogues: understanding and promoting productive interaction (pp. 48-63). london: routledge. laukkanen, a., vilkman, e., alku, p., & oksanen, h. (1997). on the perception of emotions in speech: the role of voice quality. logopedics phoniatrics vocology, 36(3), 465-475. laver, j. (1991). voice quality and indexical information. in j. laver (ed.), the gift of speech. papers in the analysis of speech and voice(pp. 147-161). edinburgh: edinburgh university press. lehiste, i. (1970). suprasegmentals. cambridge, massachusetts: mit press. lemke, j. l. (1990). talking science: language, learning, and values.new jersey, usa: ablex publishing corporation. leszcz, m. (2017). how understanding attachment enhances group therapist effectiveness.international journal of group psychotherapy, 67(2), 280-287. linell, p. (1998). approaching dialogue : talk, interaction and contexts in dialogical perspectives . amsterdam ; philadelphia, pa: j. benjamins pub. co. littleton, k., & mercer, n. (2013). interthinking: putting talk to work. london: routledge. lukkarila, p., laukkanen, a., & palo, p. (2012). influence of the intentional voice quality on the impression of female speaker. logopedics phoniatrics vocology, 37(4), 158-166. lyberg-åhlander, v., haake, m., brännström, j., schötz, s., & sahlén, b. (2015). does the speaker's voice quality influence children's performance on a language comprehension test? international journal of speech-language pathology, 17(1), 63-73. lyons, j. (1977). semantics. cambridge, uk: cambridge university press. mehan, h. (1978). structuring school structure .harvard educational review, 48(1), 32-64. mercer, n. (1995). the guided construction of knowledge : talk amongst teachers and learners . clevedon, avon, england: multilingual matters. mercer, n., & dawes, l. (2014). the study of talk between teachers and students, from the 1970s until the 2010s. oxford review of education, 40(4), 430-445. mercer, n., dawes, l., & staarman, j. k. (2009). dialogic teaching in the primary science classroom.language and education, 23(4), 353-369. mercer, n., wegerif, r., & dawes, l. (1999). children's talk and the development of reasoning in the classroom. british educational research journal, 25(1), 95-111. mortimer, e., & scott, p. (2003). meaning making in secondary science classrooms (1st ed.). berksire, england: open university press. muhonen, h., rasku-puttonen, h., pakarinen, e., poikkeus, a., & lerkkanen, m. (2017). knowledge-building patterns in educational dialogue. international journal of educational research, 81, 25-37. nassaji, h., & wells, g. (2000). what's the use of 'triadic dialogue'?: an investigation of teacher-student interaction. applied linguistics, 21(3), 376-406. niemi, k. (2016). moral beings and becomings: children's moral practices in classroom peer interaction o’connor, c., & michaels, s. (2007). when is dialogue ‘dialogic’? human development, 50(5), 275-285. o'connor, j. d., & arnold, g. f. (1973). intonation in colloquial english (2nd ed.). london: longman. oksanen, k., & hämäläinen, r. (2010). (2010). using psychophysiological methods in the research of collaborative learning games. in anonymous (eds.) international symposium on collaborative learning and argumentation (icla 2010), pépiot, e. (2013). voice, speech and gender: male-female acoustic differences and cross-language variation in english and french speakers. xvèmes rencontres jeunes chercheurs de l’ed 268, hal id: halshs-00764811 postareff, l., mattsson, m., lindblom-ylänne, s., & hailikari, t. (2017). the complex relationship between emotions, approaches to learning, study success and study progress during the transition to university. higher education, 73(3), 441-457. rogerson, j., & dodd, b. (2005). is there an effect of dysphonic teachers' voices on children's processing of spoken language? journal of voice, 19(1), 47-60. routarinne, s. (2003a). parenteesit ja nouseva sävelkulku keskustelun kielioppiin.virittäjä, 107(3), 398. routarinne, s. (2003b). tytöt äänessä. parenteesit ja nouseva sävelkulku kertojan vuorovaikutuskeinoina . helsinki, finland: suomalaisen kirjallisuuden seura. salloum, s., & boujaoude, s. (2017). the use of triadic dialogue in the science classroom: a teacher negotiating conceptual learning with teaching to the test.research in science education, scherer, k. r. (1972). judging personality from voice: a cross-cultural approach to an old issue in interpersonal perception. journal of personality, 40(2), 191-210. scherer, k. r., & giles, h. (1979). social markers in speech. cambridge, uk: cambridge university press. sinclair, j. m., & coulthard, r. m. (1975). towards an analysis of discourse: the english used by teachers and pupils (1st ed.). london: oxford university press. vilkman, e., & manninen, o. (1986). changes in prosodic features of speech due to environmental factors .speech communication, 5(3-4), 331-345. vygotsky, l. s. (1987). thinking and speech . in r. w. rieber, & a. s. carton (eds.), (1st ed., ). new york: plenum. waaramaa, t. (2014). perception of emotional nonsense sentences in china, egypt, estonia, finland, russia, sweden, and the usa. logopedics, phoniatrics, vocology, 40(3), 129-135. waaramaa, t., alku, p., & laukkanen, a. (2006). the role of f3 in the vocal expression of emotions.logopedics phoniatrics vocology, 31 (4), 153-156. waaramaa, t., laukkanen, a., airas, m., & alku, p. (2010). perception of emotional valences and activity levels from vowel segments of continuous speech.journal of voice, 24(1), 30-38. waaramaa, t., palo, p., & kankare, e. (2014). emotions in freely varying and mono-pitched vowels, acoustic and egg analyses. logopedics phoniatrics vocology, 40(4), 156-170. waaramaa, t., & kankare, e. (2013). acoustic and egg analyses of emotional utterances.logopedics phoniatrics vocology, 38(1), 11-18. waaramaa, t., & leisiö, t. (2013). perception of emotionally loaded vocal expressions and its connection to responses to music. a cross-cultural investigation: estonia, finland, sweden, russia, and the usa.frontiers in psychology, 4, 344. warwick, p., vrikki, m., vermunt, j. d., mercer, n., & van halem, n. (2016). connecting observations of student and teacher learning: an examination of dialogic processes in lesson study discussions in mathematics.zdm, 48(4), 555-569. zald, d. h. (2003). the human amygdala and the emotional evaluation of sensory stimuli (review).brain research reviews, 41, 88-123. zellner keller, b. (2004). prosodic styles and personality styles: are the two interrelated? in anonymous (eds.) proceedings of the 2nd international conference on speech prosody – sp2004 international conference on speech prosody – sp2004, (pp. 383-386). appendix db: http://science.howstuffworks.com/question124.htm. dec 12th 2016. frequency: https://www.merriamwebster.com/dictionary/frequency#medicaldictionary. dec 12th 2016. ? fundamental frequency (f0 and pitch): https://www.researchgate.net/post/what_is_pitch_or_pitch_frequency_of_a_speech_signal. dec 12th 2016.? hertz (hz): https://www.merriam-webster.com/dictionary/hertz#medicaldictionary. dec 12th 2016. loudness: http://hyperphysics.phy-astr.gsu.edu/hbase/sound/loud.html. dec 12th 2016.? pitch: https://www.merriam-webster.com/dictionary/pitch#medicaldictionary. dec 12th 2016. prosody: https://www.merriam-webster.com/dictionary/prosody. dec 12th 2016; ? http://grammar.about.com/od/pq/g/prosodyterm.htm. dec 14th 2016. semi-tone: https://www.merriam-webster.com/dictionary/semi-tone. dec 12th 2016. codepen gutzwillerlatzko frontline learning research special issue vol.8 no.5 (2020) 47 69 issn 2295-3159 happy victimizing in emerging adulthood: reconstruction of a developmental phenomenon? eveline gutzwiller-helfenfinger a & brigitte latzkob auniversity of duisburg-essen, germany buniversity of leipzig, germany article received 1 june 2018 / revised 7 november / accepted 21 november 2019 / available online 1 july 2020 abstract this study contributes to a developmental approach focusing on emotions as being of key significance in explaining the happy victimizer pattern (hv pattern) among adults. based on findings from our own research on moral emotions within the happy victimizer paradigm, we claim that a purely cognitive approach to explain the hv is overly narrow. instead, we argue that emotion attributions serve as a source for moral motivation. by identifying new dimensions (i.e., deontic judgment; own action choice; self-constructed emotion attributions) to explain the complexity of moral functioning in emerging adulthood, the current study contributes to a theoretical and methodological framework that integrates both cognitive and emotional processes to bridge the gap between moral thought, emotion, and action with the aim of fostering moral learning across the lifespan. keywords: happy victimizer phenomenon; adult moral development; moral emotions; developmentally appropriate assessment info corrseponding author email: eveline.gutzwiller-helfenfinger@uni-due.de doi: https.www.doi.org/10.14786/flr.v8i5.382 1. introduction the happy victimizer phenomenon (hvp ) denotes the empirical finding that children aged four to seven ascribe positive emotions like satisfaction or happiness to a rule transgressor although they know that a moral rule was broken (arsenio & kramer, 1992; arsenio & lover, 1995; nunner-winkler & sodian, 1988; nunner-winkler, 1999, 2012). in contrast, older children ascribe negative emotions like guilt or shame to the rule transgressor. the classical explanation according to nunner-winkler and sodian (1988) states that moral cognitions (making and justifying judgments) develop earlier than moral emotions. the attribution of positive emotions to a rule transgressor is interpreted as indicating a lack of moral motivation (nunner-winkler & sodian, 1988), the lack of moral motivation representing a lack of readiness to act on moral commitments (thorkildsen, 2013). based on the assumption that moral emotions like guilt can be seen as showing that the self not only knows a moral rule, but also feels committed to it (gibbard, 2002; malti, 2010; malti, gummerum, keller, & buchmann, 2009), the hvp can also be interpreted as a lack of moral commitment accompanying the absence of negative (i.e., moral) emotion attributions. owing to the well-established finding that by age eight or nine a shift towards negative emotion attributions can be observed, the hvp is also understood as representing a developmental transition, with the attribution of positive emotions being replaced by the attribution of negative emotions in the course of moral development (arsenio et al., 2006; lagattuta, 2005; krettenauer, malti, & sokol, 2008). according to this understanding, the hvp is restricted to preand early school age. findings from recent studies, however, challenge this position, as identical reasoning patterns (i.e., judging a transgression as wrong while attributing positive emotions to the transgressor) can be found for adolescents and adults as well (e.g., heinrichs, minnameier, gutzwiller-helfenfinger, & latzko, 2015; krettenauer, asendorpf, & nunner-winkler, 2013; krettenauer & eichler, 2006; nunner-winkler, 2007). therefore, the question arises how the occurrence of these reasoning patterns in adolescence and adulthood can be explained.2 researchers have offered different explanations for the occurrence of the hv pattern in adolescence and adulthood. heinrichs and colleagues (2015) for example argue that specific factors of a given moral conflict and its associated context might influence individuals’ respective evaluation (situation-specifity). minnameier and schmidt (2013) conceptualise the hv pattern in adolescence and adulthood as a particular moral judgment structure triggered by situation-specific factors and used by individuals to adjust to the requirements of the situation (adjustment-focused). from a perspective of the development of moral emotions, we raise the question whether the classical interpretation of the hv as a lack of moral motivation (nunner-winkler & sodian, 1988; nunner-winkler, 2007; 2013a) also applies to adolescence and adulthood. a second issue to be resolved is whether the hvp actually does disappear in the course of development. this is a relevant question: if the hvp represents a developmental transition (e.g., krettenauer et al., 2008) which is experienced by all children (i.e., a normative transition), then we need longitudinal research to indicate whether adults showing the hv pattern are displaying delayed or even dysfunctional moral development. first longitudinal findings have not yielded a clear picture (krettenauer et al., 2013). however, the basic question whether the judgment-emotion-attribution-justification patterns found in adults really represent the hvp (as documented for children) can already be approached by cross-sectional research. the aim of our study is to investigate whether the hvp or similar judgment-emotion attribution-justification patterns can be found in adults, and if so, how these patterns might be explained. more specifically, the question is whether the explanation used with children, that is, a lack moral motivation, also applies to adults. to address these aims, we also need to consider measurement issues related to the age-sensitive assessment of moral rule knowledge and moral emotions. 1.1 assessing the happy victimizer phenomenon in childhood in the classical study by nunner-winkler and sodian (1988), the hvp as such was studied systematically for the first time, although the term “happy victimizer” was not yet used. the authors sought to replicate and extend findings from a previous study by barden, zelko, duncan and masters (1980) who had observed that children aged four to five ascribed mainly positive emotions (mostly happiness) to a protagonist whose theft passed undetected, whereas children aged nine to ten and twelve to thirteen attributed negative emotions, especially fear and sadness. nunner-winkler and sodian (1988) sought to more deeply investigate children’s understanding of moral emotions: “if the developmental trend observed by barden et al. (1980) is a stable phenomenon, this may point to an important change in children's conceptions of the determinants of emotion between the preschool and the elementary school years” (nunner-winkler & sodian, 1988, p. 1324). they devised and implemented a series of three experiments to pursue this issue. the first experiment served as a test of the generality of the expected emotion attribution patterns and was intended also to offer first explanations for these patterns. children aged 4, 6 and 8 were given two emotion attribution and one moral judgment task, all of them including pictures so assist children’s understanding. the emotion attribution tasks included two parallel stories about a protagonist taking or not taking sweets or chestnuts from another child in that child’s absence (the term “stealing” was not used). each story was therefore available in a moral (not taking) and an immoral version (taking the sweets/chestnuts). counterbalancing was used in that children were presented only one version of each story, that is, either the moral version of the first and the immoral version of the second story, or vice versa. for each story, children were first told that the protagonist considered taking the sweets/chestnuts and then asked whether the protagonist was allowed to do so. this question was used to control for children’s rule understanding, namely, whether they knew that one is not allowed to take another’s belongings. afterwards, they were asked to tell how the protagonist felt. although the experimenter did not ask children to justify their emotion attributions, most of them did so. accordingly, justifications were included in analyses. in the subsequent moral judgment task, children were told a story about two children, each of whom had stolen a toy car from a friend. the experimenter asked whether it was right for them to have taken the car or not, and why. then, pictures of both protagonists were shown, with one having a happy face (because s/he now has the beautiful car) and the other having a sad face (because s/he is sorry for taking the car). children were then asked whether the happy or the sad child was worse, or whether both were the same and asked to justify their judgment. in the second experiment, the potential influence of the salience of morality in a given context on children’s emotion attributions was addressed in a sample of children aged 4 to 5. salience of morality was manipulated along the dimensions of (a) tangibility of profit of an immoral action (achieving possession of a desired object vs. managing to annoy another child); and (b) severity of transgression (telling a lie in story 1 vs. physically harming another child in story 2). story 1 (telling a lie) was acted out using two puppets. story 2 (harming) was narrated and accompanied by coloured drawings. half of the children were assigned to the tangibleand non-tangible profit conditions, respectively, and stories were told in counterbalanced order. understanding of the story was checked directly after introducing the rule transgression. after each story, first the test question (“what do you think? how does [protagonist] feel now?” “why?”) and then the control question tapping rule understanding (“what do you think about what [protagonist] did: was it right or was it not right?” “why?”) were asked. after finishing both stories children were asked whether the protagonist who sent another in the wrong direction or the protagonist who pushed another child from the swing was worse or whether they were both the same. experiment 3 investigated whether the attribution of positive emotions to a wrongdoer was limited to instances of intentional harm in a sample of 4-to-5-year-olds. four contrasting conditions were defined: intentional harm by ill-motivated actor; unintentional harm by ill-motivated actor; unintentional harm by neutrally motivated actor; and bystander witnessing someone being hurt. the harm done always referred to physical injury in the context of children playing together. four story frames were constructed, and each frame was used to create four stories representing the four contrasting conditions. movable coloured figures were used to enact the stories. each child was presented four different stories, one for each of the experimental conditions (including counterbalancing and randomisation of story frames and experimental conditions, respectively). after each story, first emotion attributions (“what do you think? how does [protagonist] feel now?” “why?”) and then moral judgments (“was [protagonist] bad or was she [he] not bad?” “why?”) were elicited. if emotions were not attributed spontaneously, the experimenter asked: “do you think [protagonist] is happy or do you think [protagonist] is sad?” “why?”. after eliciting emotion attributions and moral judgments, a control question referring to intentionality of harm was asked: “did [protagonist] hurt [victim] intentionally or did he [she] hurt him [her] not intentionally?”. for the bystander condition, the control question was: “did [protagonist] hurt [victim] or did he [she] watch [victim] being hurt?” to sum up, slightly different methods (stimulus materials, nature and functions of the questions asked, and sequence of the questions) were used in each experiment. in experiment 1, moral emotion attributions and moral judgment were assessed separately. in the emotion attribution task (coming before the moral judgment task) children’s moral rule understanding was determined before eliciting emotion attributions. in experiment 2, children’s emotion attributions were elicited before determining their moral rule understanding. moral judgment was assessed last, after finishing the two stories involved, but not in a separate task. in experiment 3, children’s emotion attributions were elicited before moral judgments. again, the latter were not assessed in a separate task. also, control questions were used for intentionality or bystander perspective. especially the nature and function as well as the sequence of the questions used pose a challenge when it comes to deconstructing the hv. interestingly, nunner-winkler and sodian (1988) never used the term “happy victimizer” for the phenomenon they explored but mentioned a “happy wrongdoer”3. in subsequent research, a more standardised assessment procedure was developed, leading to the following prototype: presenting the story; assessment of rule understanding; introducing the transgression; asking for a moral judgment and its justification; and eliciting the attribution of an emotion to the perpetrator and the justification of that emotion. variations included for example asking control questions ensuring understanding of the story (e.g., gasser & keller, 2009), probing for deserved punishment (e.g., smetana, toth, cicchetti, bruce, kane, & daddis, 1999), severity of transgression (e.g., smetana et al., 1999), evaluation of interpersonal consequences (e.g., malti & keller, 2009), ruleand authority independence (e.g., malti, gasser, & gutzwiller-helfenfinger, 2010), or including the perspective of self-as-perpetrator in addition to other-as-perpetrator (e.g., keller, lourenço, malti, & saalbach, 2003). often, following nunner-winkler and sodian (1988), additional materials like pictures (e.g., keller, 2006), cartoons (e.g., malti & keller, 2009) or figurines and dolls (e.g., woolgar, steele, steele, yabsley, & fonagy, 2001) were used to illustrate or even enact the story vignettes and support children’s understanding. for further discussions of methodological variations see for example krettenauer, malti and sokol (2008). some variation can also be found for the assessment of moral judgments and emotion attributions in particular. with respect to morally judging the transgression (sometimes also termed understanding of rule validity to emphasise the necessity of knowing a rule and its validity before being able to judge its transgression as wrong), various, slightly differing probes were used, like for example “is it right what x (victimizer) did? why/why not?” (e.g., keller et al., 2003; malti et al., 2009); “is it right to do what the victimizer did? why/why not?” (malti & keller, 2009); “is it right or wrong to do x? why?” (e.g., gasser & keller, 2009); or “is it okay or not okay for the child to do x? why? ”. action choice and its moral evaluation were used in one out of four vignettes in the study by malti & keller (2009): “how does the protagonist decide in this situation? why?” “is this the right decision or not? why?” thus, whereas nunner-winkler and sodian (1988) used the concept of “being allowed”, other researchers focused also on the “rightness” of an act, the dichotomies of “right vs. wrong” or “okay vs. not okay”, or, very rarely, asked for an action decision and its justification. with respect to eliciting emotion attributions (also called emotion expectancies to specify that emotions attributed represent what children expect others to feel in a given situation), again, slightly differing probes were utilised. regarding attributing emotions to the transgressor, probes like “how does x (victimizer) feel at the end of the story? why does s/he feel this way?” (e.g., keller et al., 2003); “how does the victimizer feel? why?” (e.g., malti & keller, 2009); or “how do you think this child will feel after s/he (x)es?” why? (malti et al., 2010) were used. when attributing emotions to the self as transgressor, probes included “how would you feel if you had done that? why would you feel that way?” (e.g., malti et al., 2009) or “how would you feel if you did x? ‘why would you feel that way?” (e.g., gasser & keller, 2009). in some studies, attribution of victims’ emotions was also elicited (e.g., gasser, malti, & gutzwiller-helfenfinger, 2012) using simple probes like “how does the victim feel? (why?)”. sometimes, children could freely attribute emotions (e.g., gutzwiller-helfenfinger, gasser, & malti, 2010), whereas in some studies they were either presented with response scales including a variety of emotions (e.g., gasser et al., 2012) or depicting gradations between “good” and “bad” (in some cases accompanied by schematically drawn faces, e.g., in the study by arsenio & kramer, 1992). to sum up, while nunner-winkler and sodian used a general probe about the way the protagonist feels, introduced by “what do you think…?”, other researchers used also more direct probes not stressing participants’ thinking, specifications of the point in time (at the end of the story, after s/he xes) or – in the case of self-as-perpetrator – formulations including conditionals (“would”), sometimes including also pre-defined response scales. as indicated above there is no doubt that happy victimizer patterns can be found in adolescents and adults. a large part of the studies addressing the hv either in adolescence or adulthood from a developmental psychological perspective have focused on its relationship with antisocial or aggressive behaviour (for a review see malti & krettenauer, 2013). empirical findings indicating that immoral conduct (e.g. breaking rules, aggressive behaviour) is, in part, related to a lack of moral emotions (malti & krettenauer, 2013) underline the assumption that moral emotions have a great impact on regulating social interaction. moral emotions are considered to be the key elements of socio-emotional competences, because they help children and adolescents (and even adults) to anticipate the outcomes of socio-moral events and adjust their social interaction accordingly (arsenio, gold, & adams, 2006; malti, 2010). within this context moral emotions have the power to regulate social interaction in the sense of providing the motivation to do the good and avoid doing the bad (kroll & egan 2004). therefore, investigating the hv, particularly the question how moral emotions impact adolescents’ and adults’ social behaviour is of great significance in explaining the causes of adaptive and maladaptive behaviour. krettenauer et al. (2013) reported a significant relationship between moral emotion attributions at the ages of 18 and 23 with antisocial conduct at age 23. additionally, they found that moral emotion attributions at the ages of 18 and 23 predicted antisocial conduct at age 23, both directly and indirectly. a study by perren and gutzwiller-helfenfinger (2012) showed that moral emotion attributions (a lack of remorse) predicted both traditional and cyberbullying in adolescents aged 12-19. moreover, a vast body of research on moral disengagement in children, adolescents and adults has consistently shown that the selective activation of moral distancing processes enables individuals to feel indifference and even happiness when harming others’ welfare (for an overview, see bandura, 2016). finally, a set of studies on cheater’s high indicated that although individuals predicted that they would feel guilty and experience increased negative affect after acting unethically, those individuals who actually did cheat consistently experienced more positive feelings than those who did not (ruedy, moore, gino, & schweitzer, 2013). however, two major issues are still unresolved, one of them of a conceptual and one of them of a methodological nature. first, conceptually, most of these studies have used the hv as an explanatory variable, shedding light on the role moral motivation (to be exact: emotion attributions) plays in explaining negative behaviour in children and youth. only few studies have investigated the hv in adolescence or adulthood as a phenomenon in its own right (nunner-winkler, 2007; minnameier, 2012) and sought to explain it. accordingly, the question why hv patterns can be found in children, adolescents, and adults still remains unresolved, as remains also the question whether the childhood phenomenon is identical with the hv patterns occurring in later life. a purely cognitive-structural explanation, postulating that the hvp can be reconstructed as a specific moral judgment structure, that is, level 2b reasoning (cf. minnameier, 2012; minnameier & schmidt, 2013), is not sufficient and falls short because both the phenomenon and its explanation are assessed on the basis of one and the same judgment-emotion attribution-justification pattern. more precisely, the pattern is used both to “diagnose” the hvp and to identify the respective moral judgment structure, leading to circular reasoning. presently we do not know whether the specific moral judgment structure (2b) can also be found in persons not displaying the hvp. only if the judgment structure cannot be found in persons not displaying the hvp can we conclude that this specific structure is found only within the hvp and might thus offer a potential explanation thereof. both the transitional and the dysfunctional hypotheses, however, are in line with the assumption that developmental changes in moral motivation cannot be explained solely as caused by changes in moral judgment (krettenauer & montada, 2005). instead, moral rule knowledge and self-evaluating moral emotions are increasingly coordinated (krettenauer & montada, 2005). according to blasi (1993, 1999) and damon and hart (1988) these processes are based on the development of a moral self or a moral identity, respectively. one’s moral identity shows itself to the extent that moral notions, such as being fair, just, and good are important to one’s self-understanding (blasi, 1984). or, conversely, when the self is not constructed or defined with reference to moral categories and shows no commitment to moral values, one does not have a moral identity (see lapsley & narvaez, 2004, p.192). taken together this position points towards a “weak” moral self as a further explanation supporting the dysfunctional hypothesis (cf. krettenauer, 2012) and raises the question whether hv is associated with a lower commitment to moral values. a second issue relates to the measurement of hv across different age groups. we have to keep in mind that investigating a developmental phenomenon necessitates the implementation of developmentally appropriate, that is, sensitive, assessment methods. to learn more about the hv across different age groups, we must ensure that we adequately measure the underlying moral mechanisms. only then can we ascertain whether the phenomenon (as a judgment-emotion attribution-justification pattern) actually can be found in adults, that is, in analogy to findings for children. thus, we have to analyse the conceptual foundations of the phenomenon as well as its assessment across the various age groups. accordingly, we will first analyse the hvp, that is, dissect it into its constituent parts on the basis of the measurement method used in the original studies on children. this will enable us to critically discuss existing findings on the hv pattern in adults. afterwards, we will transfer the operationalisation of the individual components as used for investigating children to adulthood, that is, develop an operationalisation of these components that represents a developmentally appropriate assessment suited to study the phenomenon in adults. this new measurement approach will then be used to investigate a) whether and how the phenomenon (judging a transgression as wrong while attributing positive emotions to the transgressor) manifests itself in young adults, addressing the reconstruction of the phenomenon; and b) whether the moral reasoning structures found actually represent the hvp or whether they reproduce a specific judgment-emotion attribution-justification pattern which resembles the hvp on the surface but means something different on the conceptual level. the core research questions we pursue are: what patterns of moral judgments, emotion attributions, and associated justifications do we find in young adults? can the hvp be reconstructed for adults on the basis of our findings? additionally, we wanted to explore the potential relationship between participants’ patterns of moral judgments, emotion attributions, and associated justifications and their commitment to moral values. we hypothesised that participants displaying hv reasoning patterns would show a lower level of commitment to moral values than participants not displaying hv reasoning patterns. 2. method 2.1 participants 285 pre-service teachers enrolled in a teacher education programme at the university of leipzig (germany) participated in the study. participants’ age ranged from 19 to 40 years (mage=21,74, sd=2.59), with 98.2% of the sample being 29 years or younger. 69% of participants were female. this corresponds well with the overall gender balance of the german pre-service-teacher-population. 60% of participants were enrolled in a secondary i programme, whereas 40% attended a special needs education programme focusing on socio-emotional development. most of the students studied humanities (21%) and languages (25%), only a few studied subjects related to natural sciences (5%). study participation took place in the context of a lecture on developmental psychology and was voluntary and anonymous. 2.2 instruments and procedure following krettenauer and eichler’s (2006) argumentation relating to a potential social desirability bias in adolescents’ and emerging adults’ responses in an individual interview setting, we decided to provide a half-standardised paper-and-pencil questionnaire to provide a more anonymous setting in which participants feel free to share their (written) reflections about rule transgressions and the emotions involved with the researcher. the questionnaire consisted of three parts: in the first part, participants were asked to work through two vignettes, each reflecting a moral norm conflicting with personal desires in order to assess their moral judgements, emotion attributions, and their respective justifications. in the second part, participants’ moral values were assessed. in the third part, participants provided some general sociodemographic information. 2.2.1 moral judgments, emotion attributions, and justifications in line with the traditional happy victimizer paradigm three vignettes describing the following moral rule transgressions were used: keeping excess change money (10 euros) after buying a new bicycle light (gutzwiller-helfenfinger & perren, 2015; 2016) (this vignette was called the “change money” vignette; see also study 2 in heinrichs, gutzwiller-helfenfinger, latzko, minnameier, & döring, this issue); breaking one‘s promise to wait for a prior customer while selling the motorbike to someone offering a better price (see döbert & nunner-winkler, 1983; see also study 1 in heinrichs, gutzwiller-helfenfinger et al., this issue) (this vignette was called the “the motorbike” vignette); lying to a potential customer to prevent him/her from employing another company (minnameier & schmidt, 2013) (this vignette was called the “lying to a customer” vignette). the “change money” vignette involved a passive moral temptation, where a protagonist has no intention to transgress and only realises that s/he might do so as a result of specific circumstances (gutzwiller-helfenfinger & perren, 2015; 2016; heinrichs, minnameier, gutzwiller-helfenfinger, & latzko, 2015), while the other scenarios involved proactive transgressions. each participant worked through a combination of two vignettes. the combinations were as follows: change money-the motorbike, lying to a customer-change money, the motorbike-lying to a customer. female participants received a female version (jana, maria, petra) whereas male participants received corresponding male versions (jan, mark, peter). the order of vignettes per combination was counterbalanced. the english version of the vignettes can be found in the appendix. to provide developmentally appropriate assessment of adults’ moral competencies in the context of the vignettes, we used an extended measurement approach regarding moral rule understanding and moral emotions. moral rule understanding was not only assessed by asking participants to judge the transgression, but, to gain insights into participants’ initial construction of the situation, by asking them to make a deontic judgment and to justify it. to give more room to potential complexity and richness of emotion attributions, participants had to indicate first whether they ascribed positive, negative, or mixed (i.e., both positive and negative) emotions to the protagonist and then to specify (i.e., construct) the emotion(s) attributed and justify them. the exact procedure was as follows: for each vignette, participants were first asked to make a deontic judgment: they had to indicate what the protagonist should do by ticking the appropriate box (transgress, not transgress) and to justify their judgment. for example, in the “change money” vignette, they had to indicate whether jana should keep the money or give it back, and write down the reason for doing so in their own words. afterwards, the vignette was continued by saying that the protagonist had transgressed the moral rule. in the “change money” vignette, this was expressed as follows: “let us suppose that jana kept the money”. participants then had to judge the rule transgression (classical happy victimizer judgment): they indicated whether the behaviour was “wrong” or “right” by marking the appropriate answer; afterwards, they had to tick one out of three boxes marked “good”, “bad”, and “mixed” to indicate how the protagonist felt after the transgression. additionally, they were asked to specify the exact emotion(s) in their own words and to write in their own words why the protagonist felt that way. finally, participants had to make a self-judgment. they had to indicate what they themselves would do in the given situation, that is, transgress or not transgress, by ticking the appropriate box (e.g., keep or give back the money); justify this decision in their own words; attribute emotions to themselves by ticking “good”, bad”, or “mixed”; specify the exact emotion(s) in their own words; and justify the emotion attribution(s) in their own words. 2.2.2 moral values the second part of the questionnaire assessed the values participants were committed to by using the ideal self values scale (pratt, hunsberger, pancer, & alisat, 2003). the scale includes the following twelve values: “polite and courteous”, “trustworthy”, “good citizen”, “honest/truthful”, “ambitious/hard-working”, “be open and communicate”, “careful/cautious”, “independent”, “kind and caring”, “fair and just”, “loyal”, “integrity”. six of these values belong explicitly to the moral domain (“trustworthy”, “good citizen”, “honest/truthful”, “kind and caring”, “fair and just”, “integrity” and represent a general index of commitment to a moral valuing self (campbell, 2004; pratt et al., 2003). participants had to indicate how important the twelve values were for their own life on a 6-point-likert-scale (0=unimportant; 6=important). afterwards, they had to pick and range the three most important values. 2.3 analyses emotion attributions included both a general attribution (good, mixed, bad) as well as an emotion specification (i.e., construction of the respective emotion/s) and related justification of the emotion specification. emotion specifications were categorised separately from justifications. if specifications of “good” or “bad” included more than one emotion, the most concrete and/or most complex was used, following the classification of emotions according to harris (2008). an example for “bad” was “tense, anxious, unwell”. in this case, anxiety was used because it was the most concrete emotion. an example for “good” was “good, proud”. here, pride was coded because pride was both the most concrete and the most complex emotion. an overview of positive and negative emotion specifications categorised according to their complexity is given in the results section (see also tables 5 and 6). in line with study 2 in heinrichs, gutzwiller-helfenfinger et al. (this issue), justifications of judgments and emotion attributions were content analysed using categories from research within the happy victimizer paradigm. as no new inductive categories were found for the change money vignette in relation to the categories identified in study 2 in heinrichs, gutzwiller-helfenfinger et al. (this issue), the existing categories were summarised into the following category groups: morality (i.e., referring to moral principles), empathy towards the victim, consideration of consequences, law and order, hedonism, blaming the victim, and affective distancing (i.e., stating that the protagonists’ emotions cannot be inferred). inter-rater reliability including two independent raters (10% of scenarios) was high (percentage of perfect agreement = 96.8 %, cohen’s kappa κ = .81). inter-rater reliability for emotion specifications was perfect (percentage of perfect agreement = 100%; cohen’s kappa κ = 1.0). data from the ideal self values scale were analysed for internal consistency. for both the moral (6 items) and the non-moral (6 items) subscales, cronbach’s alpha was calculated. both subscales had only moderate internal consistency (.50 for the moral and .47 for the nonmoral subscales, respectively). accordingly, the six moral values were used as single-item measures of the respective values. 3. results in order to reconstruct the happy victimizer phenomenon, we combined its basic and constitutive elements in a step-by-step procedure. in a first step, to refer back to the original hvp, we included only data from the classical happy victimizer condition referring to the evaluation of the rule transgression. accordingly, we identified participants who judged the transgression as wrong while attributing positive or mixed emotions to the perpetrator. table 1 shows the distribution of participants across the categories of “pure happy victimizer”, “mixed happy victimizer” (attributing mixed emotions) and “no happy victimizer” for the first vignette. (to include all participants the distribution is shown across all vignettes given in the first situation.) as can be seen, almost no one judged the respective transgression as wrong while attributing purely positive emotions (2.2%), indicating that we found only very few participants displaying the pure happy victimizer. however, almost half of participants judged the transgression as wrong and attributed mixed, that is, positive and negative emotions (42.5%). table 1 distribution of participants in the classical happy victimizer condition across vignettes for the first vignette presented if we break this down for the individual vignettes, we see that the distribution of participants across the three hv categories differs, with the “lying to a customer” vignette having the highest number of participants belonging in the no hv category (table 2). table 2 distribution of participants in the classical happy victimizer condition across situations 1 and 2 for each vignette additionally, we took a closer look at the data from the 104 participants categorised as no hv in the classical victimizer condition4 in the “change money” vignette. 50 (48.1%) of those (104) participants actually said that it was okay for jana to keep the money. 14 participants attributed positive, 3 attributed negative, and the remaining 33 attributed mixed emotions. similar distributions were found for the no hv category the “the motorbike” and the “lying to a customer” vignette table 3 adding the deontic judgment to the classical happy victimizer condition across vignettes for situation 1 in a second step, we added the data from the deontic judgment condition to see what the distribution of hv patterns as displayed in table 1 would look like. this meant that we now considered whether participants had also spontaneously said in the deontic judgment that the protagonist should not transgress (i.e., give back the money in the “change money” vignette), representing a more appropriate assessment of moral rule understanding in adults. thus, in the “pure hv” category we now had those participants who had initially (deontic judgment) said that the protagonist should not transgress and who afterwards – when the transgression had been introduced 5 – said that the transgression was not okay but had attributed positive emotions to the protagonist who had transgressed. only 2.2% of participants actually belonged in that category. however, the “mixed hv” category increased, with 64.5% of participants initially saying that the protagonist should not transgress and who afterwards judged the transgression as wrong but still attributed mixed (i.e., positive and negative) emotions to the protagonist. the “no hv” category, accordingly, had shrunk to 33.3% (see table 3). a chi square test revealed a significant change of the distribution of the different hv patterns (“pure”, “mixed”, “no”) by judgment condition, that is, classical vs. deontic, χ2(4,279) = 317. 38, p >.001. in a third step, we focused on the justifications participants gave for the positive emotions they attributed to the perpetrator in the classical hv condition. due to the vignette-effect reported above, we decided to perform these in-depth analyses for individual vignettes. we selected the “change money” vignette because the situation depicted there (getting too much change) was closest to participants’ everyday life-experience. table 4 justifications of positive emotions in the change money vignette for the pure and mixed hv categories a these instances were formulated in the negative, for example, having no empathy for the victim. according to the classical findings by nunner-winkler and sodian (1988), positive emotions should be justified by hedonistic reasons. as most participants attributing positive emotions actually attributed mixed emotions, that is, positive and negative emotions (see above), we again included both participants showing the “pure” and participants showing the “mixed” hv pattern. hedonistic reasons were the predominant category used to justify positive emotions (80.9%; see table 4). still, most of the other justification categories were also used in this vignette, though rather infrequently. table 5 specification and justification of emotion attribution “good” the analysis of the specifications and justifications of positive emotions in the “change money” vignette for the pure and mixed hv categories revealed that one and the same emotion attributed, that is “good”, assumed different meanings. to illustrate this finding, three different specifications and justifications of the response category “good” are shown in table 5. thus, “good” meant happiness in response 190 whereas in 275 it expresses the rejection of any concern for the shop assistant. in response 2, “good” assumed the meaning of “feeling comfortable”. to find out whether the various specifications of “good” represented different levels of emotional complexity, we summarised and categorised them on the basis of harris’ (2008) taxonomy of emotions (see also pons, harris, & de rosnay, 2004) (see table 6). there is agreement among experts that emotions run at different levels of complexity, for example basic/primary, secondary, and tertiary level emotions (parrott, 2001). all seven levels of complexity were found, with the majority of specifications covering levels 1 and 2, the lowest two levels of complexity. table 6 emotion specification of the emotion attribution “good” the same analysis was performed for negative emotions. table 7 displays the summarised emotion specifications for “bad” categorised according to harris (2008). here, seven out of ten levels of complexity were covered. while a large part of specifications covered the lowest three levels, a substantial portion (32.9%) ranged on level six referring to bad conscience and guilt, indicating more complexity and differentiation for specifications of “bad”. table 7 emotion specification of the emotion attribution “bad” in a fourth and last step we explored the potential relationship between participants’ hv status assessed in the classical hv condition in the “change money” vignette and the moral values they identified as relevant to themselves. due to the low internal consistency of the moral subscale, separate univariate anovas were performed for each item (value), that is, the degree to which that value was considered important to the self (0=unimportant to 6=important). as cell size was <5 for several cells, the hv categorisation was collapsed into no hv and hv. only for honest/truthful was a significant difference found: participants categorised as hv ascribed more importance of honest/truthful to themselves than participants categorised as no hv (µhv=5.17, sdhv=.89; n=47; µnohv=4.58, sdnohv=1.25, n=98; f[1,98=6.94, p=.01, eta2=.07]. 4. discussion the present study had two aims. first, we wanted to investigate whether the happy victimizer phenomenon can be found in emerging adults, and if so, whether the explanation used with children, that is, a lack moral motivation, also applies to emerging adults. to achieve this, we reconstructed the hv in a step-by-step analytic procedure based on written data from 285 pre-service teachers working through a set of hypothetical vignettes. as our second aim, we wanted to explore the potential relationship between participants’ patterns of moral judgments, emotion attributions, and associated justifications and their commitment to moral values. our stepwise reconstruction of the hv using a developmentally appropriate measurement approach in our sample of emerging adults yielded a number of noteworthy findings. results from our first step indicated that in vignette 1 in the classical hv condition virtually no one (2.2% of participants) judged the transgression as wrong while attributing positive emotions to the perpetrator (i.e., displayed the classical hv reasoning pattern). thus, the classical hv phenomenon hardly ever emerged in our emerging adult sample. however, two fifths were identified as falling into the “mixed” hv category, attributing both positive and negative emotions to the perpetrator while judging the transgression as wrong, confirming earlier findings involving adult samples (heinrichs et al., 2015). hence, this result implies that the classical hv phenomenon can only insufficiently be used to characterise (emerging) adults’ moral functioning in the context of hypothetical vignettes. still, the relatively high percentage of response patterns falling into the mixed hv category indicates that the happy victimizer research paradigm is relevant for the study of moral functioning in emerging adulthood. however, it is necessary to use a measurement approach going beyond the classical assessment procedure to capture the potential complexities of (emerging) adults’ moral functioning. this point will be elaborated on in more detail in the subsequent sections. analysing the distribution of patterns in the individual vignettes we found that, while the classical hv pattern was low for all three vignettes, the proportion of participants showing the mixed hv pattern differed across vignettes. it seems that the specific vignette contexts contributed to the interpretation of the respective rule transgressions in the situations depicted. there is substantial earlier research showing that both the situations and contexts involved and the specific moral principles underlying hypothetical vignettes influence the way they are interpreted and judged (cf. nunner-winkler, 2013b). in this regard, our results also confirm earlier findings on the context and situation specifity of moral judgments (krettenauer & johnston, 2011). additionally, analyses indicated that for all three vignettes about half of those participants falling into the “no hv” group said that it was okay for the perpetrator to transgress, for example, to keep the change money in the “change money” vignette. within the classical hv paradigm involving (young) children this would be seen as indicating insufficient rule knowledge or rule understanding (e.g., nunner-winkler & sodian, 1988). in a sample of emerging adults, it would be absurd to think that this is actually the case. moreover, participants said that the rule should be transgressed, for example, that jana should keep the money, in the deontic judgment condition which came before the transgression was introduced in the classical hv condition. similar findings, that is, participants saying that it is okay for the perpetrator to transgress were reported by heinrichs et al. (2015) and by heinrichs, gutzwiller-helfenfinger et al. (this issue) for studies 2, 3, and 4. in the case of study 2, the sample consisted of 14-year-olds, indicating that such transgression-friendly judgments can already be found in adolescence. that these transgression-friendly judgments were found also in the deontic judgment and in the self judgment conditions in studies 2, 3, and 4 in heinrichs, gutzwiller-helfenfinger et al. (this issue) further indicates that we cannot assume that this result is due to a methodological artefact relating to the use of the classical hv condition, where a transgression is stated as “fait accompli”. results from our second step show that, across vignettes, when more deeply examining emerging adult participants’ rule knowledge by adding the judgment made in the deontic judgment condition, the proportion of participants identified as belonging in the mixed hv category significantly increased, whereas the no hv category shrank. no change was found for the “pure” hv category. thus, in the “mixed” hv category we now had participants initially saying that the protagonist should not transgress and who afterwards judged the transgression as wrong but still attributed mixed (i.e., positive and negative) emotions to the protagonist. accordingly, we can say that insufficient rule knowledge or rule understanding very probably does not lie at the heart of the “pure” and “mixed” hv reasoning patterns. the third analytical step, addressing the justifications participants gave for the positive emotions they attributed to the perpetrator in the classical hv condition in the “change money” vignette, yielded that participants in the “pure” and “mixed” hv categories predominantly mentioned hedonistic reasons. this is in line with the classical finding by nunner-winkler & sodian (1988) as well as subsequent research. nevertheless, in almost 20% of the cases additional justification categories emerged, indicating that in our emerging adult sample the classical pattern of positive emotion attributions as justified by hedonistic reasons is not the only pattern included in participants’ constructions. it seems that those participants’ socio-moral meaning making moved beyond the classical pattern, suggesting that they were able to construct alternative interpretations why a protagonist feels good after breaking a moral rule. of course, the nature of the vignette may have played a vital role in stimulating these interpretations. in the “change money” vignette, no pro-active, planned rule transgression occurs, no negative duty is violated. instead, the protagonist is thrown into the situation, that is, tempted not to give back some money s/he mistakenly receives. this is reflected in the responses of participants classified as displaying one of the hv patterns in the classical hv condition. for example, id 275 argued that jana feels good about keeping the money “because it is not her fault, her own advantage counts for more than the shop assistant’s stupidity”. justifications as these reflect specific strategies of moral disengagement (cf. bandura, 2016) which make it possible for an individual to feel good after breaking a moral rule by cognitive reconstruction of the situation: the behaviour or its consequences are reconstructed as less harmful, the individual’s responsibility is denied or weakened, or the victim is blamed or denigrated (e.g., gutzwiller-helfenfinger, 2015a). further, when analysing specifications and justifications of positive (good”) and negative (“bad”) emotion selections, we found that a broad range of meanings was associated with those specifications. for example, “good” could mean happiness, feeling unconcerned, or feeling comfortable. this differentiation of meanings was further confirmed when we categorised the specifications according to complexity after harris (2008). in the case of positive emotions, all seven levels of complexity were found, with the vast majority of specifications (94.2%) covering the lowest two levels. a different picture emerged for negative emotions. while almost all levels were used (eight out of ten), and while the majority of specifications belonged to the lowest three levels (59.2%), still 39.4% of specifications related to guilt and mention of a bad conscience, the sixth level of complexity. it seems that for negative emotions, guilty feelings are both salient and relevant in participants’ constructions of the situation. thus, the “change money” vignette, despite its context of a passive moral temptation (no intention to transgress) and its relation to a positive duty (which has a weaker moral appeal than negative duties, e.g., belliotti, 1981; see study 2 in heinrichs, gutzwiller-helfenfinger et al., this issue), still triggers also guilty feelings in participants’ evaluation of the transgression. on the methodological level, our results suggest that it is important to have (emerging adult) participants construct (specify) their emotion attributions, instead of offering a pre-selection without asking for any specification, in order to more closely examine their moral functioning. moreover, emerging adults’ use of also more complex levels of emotions going along with the use of justifications other than hedonism provides evidence that emerging adults’ moral meaning-making is differentiated, complex, and in some cases even sophisticated. this can be seen as an indication of complex reasoning processes taking place when emerging adults evaluated this hypothetical moral vignette. and although the results of their reasoning find expression in written form only, still these written answers are sufficient to reflect the processual nature of participants’ responding. thus, what might look similar for children and emerging adults when considering the surface (judging the transgression as wrong while attributing positive emotions) carries differential depths of understanding. accordingly, while the hv framework is still relevant in studying emerging adults’ moral functioning, our results indicate that the patterns found mean something different from the phenomenon as identified in young children. consequently, we vote for using different terms in order to mark this developmentally relevant difference. we suggest to reserve the term “happy victimizer phenomenon” to studies with preschool and young schoolchildren. based on the findings in this study as well as those from study 2 in heinrichs, gutwziller-helfenfinger et al. (this issue) using a similar measurement approach we suggest to use the term “happy victimizer pattern”. the second aim of our study was to explore the potential relationship between participants’ patterns of moral judgments, emotion attributions, and associated justifications and their commitment to moral values, the latter representing a moral self. here, results were modest. the poor measurement properties of the ideal values scale (pratt et al., 2003) as determined from our data did not make it possible to create a moral subscale. it is possible that effects of culture may have influenced the interpretation of these items, as the scale was developed in a us context while we used them in a sample of german pre-service teachers. this calls for a validation of the scale for the european and more specifically, for the german context. accordingly, analyses could only be performed on the level of individual items. from the six moral items, only the item “honest/truthful” yielded a significant effect: participants categorised as displaying hv patterns (mixed or pure) saw this value as more important for their own than those categorised as no hv. however, collapsing hv categories due to the small cell size of the pure hv category is not really satisfactory, because the two categories carry different meanings. in the pure category, only positive emotions were attributed, whereas participants in the mixed category attributed also negative emotions, reflecting their inner struggle to make meaning of the situation. that those participants displaying hv patterns assigned more personal relevance to honesty/truthfulness implies that there is a discrepancy between the abstract importance they assign to this value and the concrete evaluation of a situation where this value becomes relevant. it is possible that – due to the more open nature of the passive moral temptation, these participants were not sufficiently aware that the situation was morally relevant, and that an orientation towards honesty/truthfulness might guide the interpretation of the situation by attributing guilty feelings (i.e., negative emotions) to the rule transgressor. hence, this would raise the issue of moral sensitivity (rest, narvaez, bebeau, and thoma, 1999) and its relation to a moral self (blasi, 1993). in any case, our findings underline the claim by krettenauer et al. (2008) arguing that it is necessary to explore the link between moral emotions expectancies and the development of the moral self empirically. it is probable that moral emotion expectancies are intimately linked to the development of the moral self. to our knowledge, hardly any empirical research has analysed this relationship directly (e.g., see krettenauer, campbell, & hertz, 2013), so the ideas presented here remain largely theoretical. to sum up, our findings suggest that moral emotions play an important role in emerging adults’ evaluations of morally relevant situations and moral rule transgressions. by identifying new dimensions (i.e., deontic judgment; own action choice; self-constructed emotion attributions) to explain the complexity of moral functioning in emerging adulthood the current studies contribute to a theoretical (and methodological framework) that integrates both cognitive and emotional processes to bridge the gap between moral thought, emotion, and action (malti & latzko, 2010). second, our results imply that also in emerging adulthood emotions play a central role when it comes to explaining the complexity of moral functioning. a purely cognitive-structural approach as suggested by minnameier (e.g., 2012) is not sufficient to explain our findings. very basically, even if the pattern of judging the transgression as wrong while attributing positive emotions is constructed as reflecting a specific substage of moral judgment and thus seen as basically a cognitive phenomenon, the occurrence of mixed emotions and the associated mixed hv pattern cannot be grasped by this approach. the construction or attribution of also negative emotions, together with positive emotions, is not envisaged there. finally, we cannot say that emerging adults displayed the classical hv phenomenon, which in itself would be an indicator that either the developmental transition had not been made or that participants would display a dysfunctional morality. however, many of them showed complex mixed hv reasoning patterns, at least in the context of passive moral temptations, suggesting that emerging adults’ moral functioning often includes internal struggles and ambivalence when thinking about moral issues. we need more research, also longitudinal, to explore the morality of emerging adults and potential developmental trajectories across adulthood. 4.1 limitations and outlook there are several limitations to our study. first, we used self-report data to assess participants’ moral functioning. especially regarding the ideal self values scale, it is possible that, despite the anonymous setting, participants’ answers may have been influenced by social desirability. however, participants’ answers in the hypothetical vignettes included many socially undesirable instances relating to the breach of moral rules in connection with showing no indications of guilt. yet for further studies it would be important to include a measure of social desirability to rule out this possibility and thereby strengthen the internal validity of our measurement. second, including a convenience sample of pre-service teachers implies that we can generalise our findings only to a certain extent. although our results confirm earlier findings, still the data were collected in a rather homogeneous, well-educated university sample. accordingly, at this point we cannot say whether the hv patterns can also be found in the general population of (german) emerging adults. for future studies, we need to include more diverse samples and also assess other relevant personal characteristics like for example socio-economic status, level of education, or migration background in order to more deeply explore (emerging) adult moral functioning and potential mediating or moderating factors. this also means that it is necessary to include further potentially relevant variables associated with moral functioning like empathy, social perspective-taking, moral sensitivity, or interpersonal problem solving. moreover, we need to use also behavioural measures, for example in the context of experimental settings (malti & latzko, 2017) to bridge the gap between moral functioning in hypothetical contexts and actual moral behaviour. our findings have practical implications relating to moral learning and development. by using passive moral temptations, we stimulated participants to explore the boundaries of morality, that is, situations where right and wrong are not as clear-cut as for example in situations where a negative duty like stealing, lying, etc. is violated. this resulted in a surprisingly high proportion of participants showing mixed hv reasoning patterns, indicating that our scenarios stimulated participants to think more deeply about the issues raised in the vignettes. accordingly, such materials may be well suited for use with children, adolescents, and emerging adults to stimulate their moral growth. in line with the kohlbergian tradition, we state that it is not the direction of a moral judgment or evaluation per se but the reasoning process which both reflects and stimulates moral growth. we assume that when it comes to the important and often neglected issue of vertical moral development (cf. schuster, 2001), that is, learning to transfer moral principles and reasoning to other domains, contexts, and situations, scenarios including passive moral temptations may be especially fruitful. thus, while we do not offer a «plus one» stimulation in the vygotskyan sense (blatt & kohlberg, 2006) that is, a stimulation based on a higher developmental level, we argue that passive moral temptations may offer a «plus horizon», that is, a vertical, horizon-extending stimulation. as everyday moral situations involve a high level of variation, for example lying to a stranger vs. lying to a friend or stealing out of hunger vs. stealing just for fun, it is important that individuals of all ages are offered multiple opportunities to practice their reasoning skills in a variety of educational settings. these may include well-established approaches like conflict discussions, role-play, creative writing, and so on. the important issue is that shades and gradations of meaning can be explored, reflected upon, and experienced (cf. gutzwiller-helfenfinger, 2015b). keypoints a clear distinction needs to be made between the happy victimizer phenomenon as relating to young children’s, and the happy victimizer pattern as relating to emerging adults’ (and adolescents’) moral functioning, respectively. despite some similarities on the surface level, the respective reasoning (i.e., judgment-justification-emotion attribution-justification) patterns are developmentally distinct. exploring the link between moral emotion attributions and the commitment to moral values is a promising pathway for further research on (adult) moral functioning. the assessment of (emerging) adults’ moral reasoning tapping both cognitive and emotional components necessitates the use of developmentally appropriate or sensitive assessment approaches. our findings emphasise the situation specifity of moral emotion attributions. further research is needed to explore the potential of educating moral emotions footnotes 1 we use hvp as an abbreviation to refer to the happy victimizer phenomenon. 2we use the term emerging adulthood to refer to our own sample and refer to adolescence, adulthood, and young adulthood when using the terminology of the research cited. 3 it seems that arsenio & kramer (1992) were the first to use this term. 4in the present paper, the term “condition” refers to the specific form of assessment. for example, the deontic judgment condition refers to the assessment of the deontic judgment. 5participants had to make a deontic judgment before the transgression was introduced and afterwards had to judge the transgression. references arsenio, w. f., gold, j., & adams, e. (2006). children’s conceptions and displays for moral emotions. in m. killen & j. g. smetana (eds.), handbook of moral development (pp. 581–609). mahwah, nj: erlbaum. arsenio, w. f., & kramer, r. (1992). victimizers and their victims: children’s conceptions of the mixed emotional consequences of moral transgressions. child development, 63(4), 915–927. doi: 10.2307/1131243 arsenio, w. f., & lover, a. (1995). children’s conceptions of sociomoral affect: happy victimizers, mixed emotions, and other expectancies. in m. killen & d. hart (eds.), morality in everyday life: developmental perspectives (pp. 87–128). cambridge: cambridge university press. barden, c. r., zelko, f. a., duncan, w. s., & masters, j. c. (1980). children’s consensual knowledge about the experiential determinants of emotion. journal of personality and social psychology, 39(5), 968–976. https://doi.org/10.1037/0022-3514.39.5.968 blasi, a. (1984). moral identity: its role in moral functioning. in w. kurtines, & j. gewirtz (eds.), morality, moral behavior and moral development (pp. 128 –139). new york: wiley. blasi, a. (1993). the development of identity: some implications for moral functioning. in g. g. noam & t. e. wren (eds.), the moral self (pp. 99 – 122). cambridge, ma: mit press. blasi, a. (1999). emotions and moral motivation. journal for the theory of social behaviour, 29(1), 1-19. doi:10.1111/1468-5914.00088 bandura, a. (2016). moral disengagement: how people do harm and live with themselves. new york: worth publishers. blatt, m. & kohlberg, l. (2006). the effects of classroom moral discussion upon children's level of moral judgment. journal of moral education, 4(2), 129-161. doi: 10.1080/0305724750040207 campbell, k. m. (2004). moral identity, youth engagement, and discussions with parents and peers . st. catharines, ontario, canada: brock university. damon, w., & hart, d. (1988). self-understanding in childhood and adolescence. cambridge: cambridge university press. döbert, r. & nunner-winkler, g. (1983). moralisches urteilsniveau und verlässlichkeit. die familie als lernumwelt für kognitive und motivationale aspekte des moralischen bewusstseins in der adoleszenz. in g. lind, h.h. hartmann, & r. wakenhut (hrsg.), moralisches urteil und soziale umwelt [moral judgment and social environment] (s. 95-122). weinheim und basel: beltz verlag. gasser, l., & keller, m. (2009). are the competent the morally good? perspective taking and moral motivation of children involved in bullying. social development, 18, 798–816. doi: 10.1111/j.1467-9507.2008.00516.x gasser, l., malti, t., & gutzwiller-helfenfinger, e. (2012) aggressive and nonaggressive children's moral judgments and moral emotion attributions in situations involving retaliation and unprovoked aggression. the journal of genetic psychology, 173(4), 417-439. doi: 10.1080/00221325.2011.614650 gibbard, a. (2002). normative and recognitional concepts. philosophy and phenomenological research, 64(1), 151-167. doi: 10.1111/j.1933-1592.2002.tb00148.x gutzwiller-helfenfinger, e. (2015a). moral disengagement and aggression. comments on the special issue. merrill palmer quarterly, 61(1), 192-211. doi: 10.13110/merrpalmquar1982.61.1.0192 gutzwiller-helfenfinger, e. (2015b). die wirkung von erweitertem rollenspiel auf soziale perspektivenübernahme und antisoziales verhalten. in t. malti, & s. perren (hrsg),soziale kompetenz bei kindern und jugendlichen [social competence in children and adolescents] (s. 244-261; 2. überarb. und erw. aufl.). stuttgart: kohlhammer. gutzwiller-helfenfinger, e., & perren, s. (2015). adolescents’ evaluations of passive moral temptations – relations to bully-victim problems. paper presented in the symposium morality and bully-victim problems (chairs: l. kollérova & d. strohmeier). 17th european conference on developmental psychology (ecdp), braga (portugal), september 8-12, 2015. gutzwiller-helfenfinger, e., & perren, s. (2016). the relationship between adolescents’ bully-victim problems and their use of mechanisms of moral disengagement in the context of passive moral temptations. paper presented in the symposium moral disengagement in the production of aggression (chairs: k. runions & e. gutzwiller-helfenfinger). 22nd world meeting of the international society for research on aggression (isra), sydney (australia), july 19-23, 2016. gutzwiller-helfenfinger, e., gasser, l., & malti, t. (2010). moral emotions and moral judgments in children’s narratives: comparing real-life and hypothetical transgressions. new directions for child and adolescent development, 129, 11–31. doi: 10.1002/cd.273 harris, p.l. (2008). children’s understanding of emotions. in m. lewis, j. haviland-jones, & l. f. barrett (eds.), handbook of emotions (3rd ed., pp. 320-331). new york: the guilford press. heinrichs, k., minnameier, g., gutzwiller-helfenfinger, e., & latzko, b. (2015). „don’t worry, be happy?" das happy-victimizer-phänomen im berufsund wirtschaftspädagogischen kontext. zeitschrift für berufs-und wirtschaftspädagogik, 111(1), 31–55. keller, i. (2006). moralische motivation bei achtjährigen kindern. zusammenhänge von moralischer motivation mit sozialverhalten und der beliebtheit unter gleichaltrigen (forschungsbericht nr. 3 aus der reihe z-proso). zürich: pädagogisches institut der universität zürich. keller, m., lourenço, o., malti, t., & saalbach, h. (2003). the multifaceted phenomenon of ‘happy victimizers’: a cross-cultural comparison of moral emotions. british journal of developmental psychology, 21 , 1–18. doi: 10.1348/026151003321164582 krettenauer, t. (2012). linking moral emotion attributions with behavior: why “(un)happy victimizers” and “(un)happy moralists” act the way they feel. new directions for youth development, 136, 59-74. doi: 10.1002/yd krettenauer, t., asendorpf, j. b., & nunner-winkler, g. (2013). moral emotion attributions and personality traits as long-term predictors of antisocial conduct in early adulthood: findings from a 20-year longitudinal study. international journal of behavioral development, 37, 192–201. doi: 10.1177/0165025412472409 krettenauer, t., campbell, s., & hertz, s. (2013). moral emotions and the development of the moral self in childhood. european journal of developmental psychology, 10(2), 159-173. doi: 10.1080/17405629.2012.762750 krettenauer, t., & eichler, d. (2006). adolescents' self-attributed emotions following a moral transgression: relations with delinquency, confidence in moral judgment, and age. british journal of developmental psychology, 24, 489–506. doi: 10.1348/026151005x50825 krettenauer, t. & johnston, (2011). positively versus negatively charged moral emotion expectancies in adolescence: the role of situational context and the developing moral self. british journal of developmental psychology 29(3), 475-488. doi: 10.1348/026151010x508083 krettenauer, t., malti, t., & sokol, b. (2008). the development of moral emotion expectancies and the happy victimizer phenomenon: a critical review of theory and application. european journal of developmental science, 2, 221–235. doi: 10.3233/dev-2008-2303 krettenauer, t. & montada, l. (2005). entwicklung von moral und verantwortlichkeit. in j. b. asendorpf (hrsg.), soziale, emotionale und persönlichkeitsentwicklung [social, emotional, and personality development] (s. 141-189). göttingen: hogrefe. kroll, j., & egan, e. (2004). psychiatry, moral worry, and the moral emotions. journal of psychiatric practice, 10(6), 352-360. doi: 10.1097/00131746-200411000-00003 lagattuta, k. h. (2005). when you shouldn’t do what you want to do: young children’s understanding of desires, rules, and emotions. child development, 76(3), 713–733. doi: 10.1111/j.1467-8624.2005.00873.x lapsley, d. k., & narvaez, d. (2004). a social-cognitive approach to the moral personality. in, d. k. lapsley & d. narvaez (eds.), moral development, self and identity (pp. 189 212). mahwah, n. j.: erlbaum. malti, t., gasser, l., & buchmann, m. (2009). aggressive and prosocial children’s emotion attributions and moral reasoning. aggressive behavior, 35, 90–102. doi: 10.1002/ab.20289 malti, t., gasser, l., & gutzwiller-helfenfinger, e. (2010). children’s interpretive understanding, moral judgments, and emotion attributions: relations to social behaviour. british journal of developmental psychology, 28, 275–292. doi: 10.1348/026151009x403838 malti, t., & keller, m. (2009). the relation of elementary-school children’s externalizing behaviour to emotion attributions, evaluations of consequences, and moral reasoning. european journal of developmental psychology, 6(5), 592–614. malti, t., keller, m., gummerum, m., & buchmann, m. (2009). children’s moral motivation, sympathy, and prosocial behavior. child development, 80(2), 442–460. malti, t., & krettenauer, t. (2013). the relation of moral emotion attributions to prosocial and antisocial behavior: a meta-analysis. child development, 84(2), 397–412. doi: 10.1111/j.1467-8624.2012.01851.x malti, t., & latzko, b. (2010). children's moral emotions and moral cognition: towards an integrative perspective. new directions for child and adolescent development, 129, 1-10. malti, t., & latzko, b. (2017). moral emotions. in j. stein (ed.), reference module on neuroscience and biobehavioral psychology. oxford, uk: elsevier. doi: 10.1016/b978-0-12-809324-5.06491-9 minnameier, g. (2012). a cognitive approach to the ‘happy victimiser’. journal of moral education, 41, 491-508. doi: 10.1080/03057240.2012.700893 minnameier, g., & schmidt, s. (2013). situational moral adjustment and the happy victimizer. european journal of developmental psychology, 10, 253-268. doi: 10.1080/17405629.2013.765797 nunner-winkler, g. (2007). development of moral motivation from childhood to early adulthood. journal of moral education, 36(4), 399–414. https://doi.org/10.1080/03057240701687970 nunner-winkler, g. (2013a). moralische entwicklung. in m. stamm & d. edelmann (eds.), handbuch frühkindliche bildungsforschung [handbook of early childhood educational research] (pp. 653–665). wiesbaden: springer fachmedien. https://doi.org/10.1007/978-3-531-19066-2 nunner-winkler, g. (2013b). moral motivation and the happy victimizer phenomenon. in k. heinrichs, f. oser & t. lovat (eds.), handbook of moral motivation. theories, models, applications (pp. 267-288). rotterdam: sense publishers. nunner-winkler, g., & sodian, b. (1988). children’s understanding of moral emotions. child development, 59(5), 1323–1338. doi: 10.2307/1130495 parrott, w. (2001). emotions in social psychology. philadelphia: psychology press. perren, s., & gutzwiller-helfenfinger, e. (2012). cyberbullying and traditional bullying in adolescence: differential roles of moral disengagement, moral emotions, and moral values. european journal of developmental psychology, 9(2), 195-209. doi: 10.1080/17405629.2011.643168 pons, f., harris, p.l., & de rosnay (2004 ). emotion comprehension between 3 and 11 years: developmental periods and hierarchical organization . european journal of developmental psychology, 1(2), 127–152. doi: 10.1080/17405620344000022 pratt, m. w., hunsberger, b., pancer, s. m., & alisat, s. (2003). a longitudinal analysis of personal values socialization: correlates of a moral self-ideal in late adolescence. social development, 12, 563–585. doi:10.1111/1467-9507.00249 ruedy, n. e., moore, c., gino, f., & schweitzer, m. e. (2013). the cheater’s high: the unexpected affective benefits of unethical behavior. journal of personality and social psychology, 105(4), 531. schuster, p. (2001). von der theorie zur praxis – wege zur unterrichtspraktischen umsetzung des ansatzes von kohlberg, in w. edelstein, f. oser and p. schuster, eds., moralische erziehung in der schule. entwicklungspsychologie und pädagogische praxi s [moral development in school. developmental psychology and educational practice] (pp. 177-212). weinheim und basel:beltz. smetana, j. g., toth, s. l., cicchetti, d., bruce, j., kane, p., & daddis, c. (1999). maltreated and nonmaltreated preschoolers’ conceptions of hypothetical and actual moral transgressions. developmental psychology, 35, 269-281. doi: 10.1037/0012-1649.35.1.269 thorkildsen, t. a. (2013). motivation as the readiness to act on moral commitments. in k. heinrichs, f. oser & t. lovat (eds.), handbook of moral motivation. theories, models, applications (pp. 83-96). rotterdam: sense publishers. woolgar, m., steele, h., steele, m., yabsley, s., & fonagy, p. (2001). children's play narrative responses to hypothetical dilemmas and their awareness of moral emotions. british journal of developmental psychology, 19(1), 115-128. appendix hypothetical scenarios used change money jana uses her bike every day to go to school. she urgently needs a new tail light. she buys a suitable tail light in a bike shop nearby. she chooses one that costs 32.euros. jana pays cash with a 50-euro bill and, when leaving the shop, notices that the shop assistant gave her 10.euros too much in change. what should jana do (keep the money / return the money)? why? please justify your choice. suppose jana keeps the money. is it okay for jana to keep the money (okay / not okay)? how does she feel (good / mixed feelings [both good and bad], bad)? please name jana’s feelings precisely. please justify your assessment. what would you do in this situation (keep the money / return the money)? how would you feel (good / mixed feelings [both good and bad], bad)? please name your feelings precisely. please justify your assessment. the motorbike peter offers his motorbike for sale. he wants to sell it for 800.euros. a young man is interested in the bike. he beats peter down to 700.euros. the two men come to an agreement. however, the young man does not have enough cash on him. but he promises to be back with the money in half an hour. peter says: “agreed, i will wait for you.” a short time afterwards, though, another customer joins peter. he is prepared to pay the 800.euros in cash right on the spot. what should peter do (sell the motorbike to the new customer / wait for the first customer)? why? please justify your choice. suppose peter sells the motorbike to the new customer. is it okay for peter to sell the motorbike to the new customer (okay / not okay)? how does he feel (good / mixed feelings [both good and bad], bad)? please name peter’s feelings precisely. please justify your assessment. what would you do in this situation (sell the motorbike to the new customer / wait for the first customer)? how would you feel (good / mixed feelings [both good and bad], bad)? please name your feelings precisely. please justify your assessment. lying to a customer maria has founded an enterprise in an innovative technology sector. the enterprise is in a critical start-up phase. maria struggles with financial straits and a fluctuating order situation. she has just overcome a slack season. now the business is running smoothly again. she receives an order that needs to be processed at very short notice. maria knows already that she will not be able to meet the deadline and will have to stave off the customer. moreover, when attending a start-up workshop, she happened to learn about a rival company that would be able to process the order both more speedily and reliably. she ponders whether to inform the customer about the rival company or whether to keep the order in her company. what should maria do (inform the customer / not inform the customer)? why? please justify your choice. suppose maria does not inform the customer. is it okay for maria not to inform the customer (okay / not okay)? how does she feel (good / mixed feelings [both good and bad], bad)? please name maria’s feelings precisely. please justify your assessment. what would you do in this situation (inform the customer / not inform the customer)? how would you feel (good / mixed feelings [both good and bad], bad)? please name your feelings precisely. please justify your assessment. frontline learning research vol. 5 no. 3 special issue (2017) 55 65 issn 2295-3159 corresponding author: ass. prof. adam szulewski bsc, md, frcpc, mhpe, dep. of emergency medicine, queen’s university, empire 3, kingston general hospital, 76 stuart street, kingston, ontario, k7l 2v7,canada, aszulewski@qmed.ca doi: http://dx.doi.org/10.14786/flr.v5i3.256 pupillometry as a tool to study expertise in medicine adam szulewskia, danielle keltonb, daniel howesc adepartment of emergency medicine, queen’s university, canada bfaculty of medicine, queen’s university, canada cdepartments of emergency medicine and critical care medicine, queen’s university, canada article received 2 may / revised 9 november / accepted 23 march / available online 14 july abstract background pupillometry has been studied as a physiological marker for quantifying cognitive load since the early 1960s. it has been established that small changes in pupillary size can provide an index of the cognitive load of an individual as he/she performs a mental task. the utility of pupillometry as a measure of expertise is less well established, although recent research in the fields of education, medicine and psychology indicates that differences in pupillary size during domain-specific tasks allows differentiation between experts and novices in appropriately designed experiments. purpose the goal of this review is to explore the existing body of evidence for the use of pupillometry as a measure of expertise and to identify its strengths and constraints within the context of expertise research in the medical sciences. results pupillometry is a robust metric that allows researchers to better understand cognitive load in medical practitioners with varying levels of expertise. in medical expertise research, it has been used to study surgeons, anesthetists and emergency physicians. its strengths include its ability to provide quantitative and objective outputs, to be measured unobtrusively with new technology and to be precisely computed as cognitive load changes over the course of completion of a task. constraints associated with this methodology include its potential inaccuracy with changes in ambient light and pupillary accommodation as well as the need for relatively expensive equipment. conclusion with recent technological advances, pupillometry has become a simple and robust method for quantifying physiological changes attributable to cognitive load and is increasingly being utilized in medical education. it can be used as a reliable marker of cognitive load and has been shown to differentiate levels of expertise in medical practitioners. keywords: pupillometry; cognitive load; expertise; medical education http://dx.doi.org/10.14786/flr.v5i3.256 szulewski et al | f l r 56 1. background the measurement of human cognitive load has been of interest to researchers for decades. knowing how intensely a person is thinking has implications beyond knowing what that person is thinking about. this is particularly relevant in the context of professional domains (like medicine) where critical and cognitively loading decisions often need to be made with limited time and in the context of other competing priorities. this “intensity” of thinking, which is related to cognitive load, functions within the constraints of a limited working memory. working memory is a key executive function. executive functions are a group of mental processes that are required when an individual has to pay attention, and when it would be considered inappropriate, insufficient or impossible to rely on instinct or to respond automatically (burgess & simons, 2005). in addition to working memory, executive functioning involves two other core activities: inhibition (self-control, selective attention, cognitive inhibition) and cognitive flexibility (which is closely related to creativity). these three core activities are combined in different ways to build higher order executive functions such as reasoning, problem solving and planning (diamond, 2013). working memory, which is responsible for the manipulation of stored information (or our ability to “think”), is generally thought to be limited (paas, tuovinen, tabbers, & van gerven, 2003). ongoing work in this area now suggests that with experience, experts are able to expand working memory capacity by having developed methods for storage and retrieval of domain-specific information in long-term memory – so called long-term working memory (ericsson & kintsch, 1995). this is accomplished, in part, by pattern recognition and schema development, and a resultant relative decrease in the cognitive load that a problem or situation imposes as an individual becomes more expert-like (szulewski, roth, & howes, 2015). functionally, cognitive load can be thought of as the mental capacity that is allocated to performing a task (paas et al., 2003). it is thought to be comprised of three components: intrinsic cognitive load (icl), extraneous cognitive load (ecl) and germane cognitive load (gcl) (young, van merrienboer, durning, & ten cate, 2014). icl is a function of expertise and task complexity, while ecl is related to suboptimal information presentation conditions. gcl refers to the working memory resources that are dedicated to processing icl, and thus to learning (sweller, 2010). in general, researchers measure cognitive load using psychometric scales, physiological variables and secondary task methodology (paas et al., 2003). briefly, psychometric scales gather subjective data from participant self-reports after task completion. physiological variables use task-evoked pupillary responses (teprs) or pupillometry, heart rate variability, galvanic skin response (among others) as surrogate markers of cognitive load. secondary task methodology relies on participants’ performance on a secondary task (that requires sustained attention, like detecting an auditory signal) and uses this information to glean the level of cognitive load imposed by the primary task. each of these techniques has its own strengths and limitations; but in general, each is thought to provide data about total (or measurable) cognitive load. the contribution of intrinsic, extraneous and germane cognitive load to each of these measurement techniques remains to be elucidated (leppink, paas, van gog, van der vleuten, & van merriënboer, 2014). this review focuses on one particular physiological method of measuring cognitive load – pupillometry, which is the study of changes in pupil size. we will first examine the technique of pupillometry as a surrogate marker for cognitive load in non-medical domains and then we will focus the discussion on pupillometry research in medicine and how this relates to the development of expertise. finally, constraints of the technique will be discussed. szulewski et al | f l r 57 2. pupil physiology dilation and constriction of the human pupil is necessary for day-to-day visual tasks. dilation (mydriasis) of the pupil is accomplished by the contraction of the iris dilator (radial) muscle, which is controlled by the sympathetic nervous system. constriction (miosis) of the pupil occurs when the iris sphincter (circular) muscle contracts, which is controlled by the parasympathetic nervous system. two commonly tested reflexes in clinical medicine are the light reflex and the accommodation reflex. during the light reflex, the pupil dilates in low luminance environments and constricts in high luminance environments. in the accommodation reflex, as an individual changes visual focus from a distant object to a closer object, the pupil constricts, and vice-versa (lang, 2015). in addition to these clinically measurable and commonly discussed reflexes, pupils also change in size as a result of non-visual stimuli. this was first described in detail by hess and polt (1960) where it was shown that pupil size varied when participants viewed particular images (for example, sexually suggestive ones). follow-up studies by this group and others further demonstrated that pupil size could be used to measure cognitive load (or mental effort). physiologically, it is thought that pupil size changes with cognitive loading as a result of pathways that originate in the locus coeruleus, which is a major norepinephrine source in the brain (laeng, sirois, & gredebäck, 2012). in fact, locus coeruleus activity has been shown to be very closely related to sympathetic activity and changes in pupil size (aston-jones & cohen, 2005). these pupillary responses are spontaneous and very difficult to control voluntarily. the voluntary dilation of a subject’s pupils is only possible indirectly if the subject imagines a situation (e.g. selfinduced sexual imagery) where his/her pupils would normally dilate (whipple, ogden, & komisaruk, 1992). this would be particularly difficult, if not impossible, to systematically do while simultaneously performing other cognitively loading tasks. this makes the technique robust. although other autonomic measurements like heart rate and skin resistance have also been found to provide similar information regarding sympathetic activity (and thus cognitive loading), pupillometry has been found to yield the most consistent and readily analysable results (kahneman, tursky, shapiro, & crider, 1969). 3. pupillometry as a measure of cognitive load in a seminal article, hess and polt (1964) found there to be a strong correlation between difficulty of arithmetic problems posed to participants and the magnitude of the increase in their pupil sizes. further, they observed that after a question was asked, participants’ pupils showed a gradual increase in diameter, reached a maximum size just prior to reporting an answer, and then reverted back to their original diameter shortly thereafter. in another article, beatty and kahneman (1966) built upon these original experiments and were able to confirm two phases in the pupillary response to cognitive processing. first, they noted a loading phase with dilation corresponding to information gathering and an unloading phase where the pupil constricted as answers were verbalized by the participants. based on the results from these studies as well as others, it became generally accepted that changes in pupil size reflect changes in cognitive processing load during task performance and provide information about processing resources. specifically, more difficult cognitive tasks were found to cause both an increase in the amplitude and the latency of pupillary dilation (beatty, 1982). these early experiments were carried out with relatively onerous experimental processes that involved developing large quantities of photographs taken by cameras in precisely controlled environments and then manually measuring pupil size with a ruler. this made large-scale experiments impractical. modern technology has allowed researchers to electronically collect pupil size data with stationary as well as mobile devices, obviating the need for time-consuming manual measurement and allowing for less stringent experimental environments. some of the previously described studies that used arithmetic problems have szulewski et al | f l r 58 now been replicated with the new technology, showing similar results. figure 1 is taken from one of these studies that used a mobile eye-tracker to capture participant pupil size at a rate of 30hz during arithmetic problem solving. as was first shown in the original experiments, the new technology also demonstrated that pupil size increased with increasing problem difficulty and changed predictably with phases of information gathering and delivery of responses (szulewski, fernando, baylis, & howes, 2014). figure 1. “difficult questions resulted in peak dilation of 11.8% compared to baseline whereas “easy” questions resulted in peak dilation of 5.0% compared to baseline (p = 0.005). time 0 to 3 seconds serves as baseline (3 seconds prior to question presentation); time 3 to 8 seconds corresponds to the time that the question was on the screen; time 8 to 11 seconds corresponds to the 3 seconds after the question was removed and the black dot appeared. [from szulewski, a., fernando, s. m., baylis, j., & howes, d. (2014). increasing pupil size is associated with increasing cognitive processing demands: a pilot study using a mobile eye-tracking device. open journal of emergency medicine, 2014. reprinted with permission]. the ease of use and precision of the newer technology has expanded the role of pupillometry to more theoretical realms. in addition to reliably demonstrating increased cognitive load with increasing question difficulty, pupillometry data have also shown that the modality of information presentation has cognitive loading effects. using a remote eye tracker, klingner, tversky, and hanrahan (2011) showed that cognitive load is higher for the same tasks when they are presented orally as opposed to visually. these experiments underscore the precision and expanded applications of this technique. other groups of researchers have also investigated the ability to measure cognitive load in novel environments using pupillometry. one such experiment by palinko, kun, shyrokov, and heeman (2010) investigated measuring mean pupil diameter change in drivers as they operated a simulated vehicle while they were involved in simultaneous spoken dialogues. pupil diameter changed as expected and the authors concluded that pupillometry was better in quantifying small changes in cognitive load in the simulator compared with other measures like lane position and steering wheel angle. results from studies like this one suggest that pupillometry can be reliably used in more true-to-life situations in addition to well controlled laboratory settings. importantly, during the driving simulator experiment, luminance varied only ± 5% in the simulated experimental environment which likely minimized the contribution of the light reflex to pupillary szulewski et al | f l r 59 changes and allowed for a relatively clean signal. changes in luminance become more of an experimental issue in real-world environments where background luminance varies to greater degrees. on the whole, these studies seem to suggest that the construct being measured with pupillometry is cognitive load, although there is research to suggest that other factors (e.g. emotion, fatigue, age, pain and certain drugs) also contribute to change in pupil size (holmqvist et al., 2011). validity evidence for the use of pupillometry to measure cognitive load specifically was recently described by szulewski, gegenfurtner, howes, sivilotti, and van merriënboer (2016). in this study, pupillometric measurements of cognitive load were compared to psychometric measurements of cognitive load across different question types, question difficulty and experience levels in a testing environment. based on the predictability of the results and the strong correlation of the measurement instruments, the authors concluded that there is validity evidence to use either psychometric or pupillometric measurements to measure cognitive load in traditional testing environments. 4. pupillometry research in medicine given the promising results of pupillary analysis in experimental settings and the increasing availability of the new technology, researchers have started to expand its use into other domains, including medicine. medicine is a particularly interesting field in which to study cognitive load given its inherent characteristics where physicians regularly make high stakes decisions, often under considerable external pressures (including time and stress). these characteristics are emphasized during non-routine emergency situations. one study that investigated critical incidents in the operating room examined anaesthesiology trainees’ pupil sizes, among other physiological responses as surrogate markers of cognitive workload (schulz et al., 2011). participants’ pupil sizes were found to increase as the severity of a critical incident increased. although this pattern held true within scenarios, the authors found that there was no difference between sessions or individuals. this was thought to be due to individual pupil variations as well as external factors like lighting. these issues raise the concern that external factors can skew pupillometric results and make it difficult to interpret the data reliably in real-world environments where luminance is not adequately controlled. all real-life physician-patient clinical encounters would, as a result, be affected. the main issue in these scenarios involves the light reflex which is capable of causing pupil diameter changes of up to 120% from baseline, which is far greater than the changes of up to 20% that can be attributed to cognitive processing demands. (holmqvist et al., 2011; laeng et al., 2012). in an effort to mitigate the confounding effects of the light reflex, investigators often try to control for luminance during their experiments. zheng, jiang, and atkins (2015) did just this and were able to confirm that pupil responses behaved as expected with changing sub-task difficulty in a simulated laparascopic surgical experiment. in a related study, the same group noted that the rate of change of pupil size was better than pupil diameter in assessing mental workload of the simulated laparscopic task (jiang, zheng, tien, & atkins, 2013). to address some of these issues, new techniques (like the index of cognitive activity) have been designed to separate out the light reflex from pupil changes secondary to cognitive workload by measuring abrupt discontinuities in the pupil size signal (marshall, 2002). this index of cognitive activity has been utilized in the objective assessment of surgical skill where pupil size (along with other eye and pupillary metrics) was used to objectively classify non-expert from expert surgeons in environments that were uncontrolled for luminance including a simulator as well as a live operating room (richstone et al., 2010). szulewski et al | f l r 60 5. pupillometry and expertise performance on tests has universally been utilized to measure the construct of ability, intelligence, competence or expertise in a domain. despite its wide use, test-taking is known to have many limitations as a surrogate marker to measure these constructs. applications of pupillometry have allowed researchers to delve deeper into this area than simply examining performance. this is particularly interesting when considering cognitive processing as subjects answer questions correctly. based on the traditional view of assessment in test-taking, two individuals who get the same score on a test are thought to have equal domainspecific skill (or ability or expertise). the reality is more nuanced. figure 2 is taken from a study by ahern and beatty (1979) which shows the cognitive processing demands (as measured by pupillometry) of participants with both high and low intelligence as defined by scholastic aptitude test scores as they were faced with arithmetic problems (and answered these problems correctly). based on traditional assessment modalities, both groups of individuals would be assessed equally for having answered correctly. a closer analysis, however, revealed that the group with “lower intelligence” had greater increases in pupillary dilation than the “higher intelligence” group at all question difficulty levels. essentially, the group with the lower intelligence had to “think harder” to achieve the same correct response as the group with higher intelligence. figure 2. averaged task-evoked pupillary responses for correctly solved problems at three levels of difficulty for subjects in the high and low groups of psychometrically measured intelligence. at all difficulty levels, larger pupillary responses are observed for subjects in the low group. [from ahern, s., & beatty, j. szulewski et al | f l r 61 (1979). pupillary responses during information processing vary with scholastic aptitude test scores. science, 205(4412), 1289-1292. reprinted with permission from aaas.] moving away from intelligence defined by standardized testing, a study by szulewski et al. (2015) found similar results in novices and trained physicians as they answered clinically-based multiple choice questions. the participant groups in this study were divided not by intelligence, but by clinical experience. those with more clinical experience (the trained physician group) had smaller changes in pupil diameter as they answered the questions compared to the more novice group when both groups answered correctly (see figure 3). in another study, tien et al. (2015) found that junior surgeons had greater pupil sizes than expert surgeons during open inguinal hernia repair. both of these studies (which divided physician participant groups based on experience level) emphasize that those with less experience expend more cognitive load than those with more experience when they perform domain-specific tasks, even when the measured outcome is the same. although it is reasonable to assume that these observed differences are due to different experience levels, one might argue that there may be other confounding factors between the groups that could skew the results. this is a potential issue for any cross-sectional study. a study by richstone et al. (2010) suggests that it is in fact experience/expertise that is responsible for the pupillometric changes between groups, as opposed to another confounder. as part of their study, they examined one non-expert surgeon three times over the course of 18 months both in simulated and live surgical environments. during this longitudinal analysis, they found that it became increasingly difficult to differentiate this non-expert from the expert surgeon group as his pupil metrics became more expert-like over time with increased training and experience. this finding suggests that the differences between groups of participants are in fact due to skill or expertise, as opposed to another confounding factor. overall, these studies suggest that there is empirical evidence that those with more domain-specific experience exhibit a certain cognitive efficiency as they perform tasks associated with their training and experience that their novice counterparts have not yet developed. figure 3. results of an analysis of correctly answered clinical multiple-choice questions. the increase in the pupil diameter of novices was significantly greater than that of trained physicians (p < 0.001). [from szulewski, a., roth, n., & howes, d. (2015). the use of task-evoked pupillary response as an objective measure of cognitive load in novices and trained physicians: a new tool for the assessment of expertise. academic medicine, 90(7), 981-987. reprinted with permission from aamc.] szulewski et al | f l r 62 it is debatable whether differences in cognitive efficiency are relevant during a test where a student is asked a sequence of questions and he/she generally focuses all of his/her working memory onto the question at hand before moving onto the next one. moreover, it is debatable whether an assessor would even want to know this information. arguably, however, this cognitive efficiency in the “more intelligent” or “more skilled” or “more experienced” or “more expert-like” group becomes relevant in complex situations with competing priorities, as the less cognitively strained individual will have a greater proportion of his/her working memory available for other cognitively demanding executive functions. one such area where competing priorities often coexist and where cognitive efficiency might be beneficial is clinical medicine. during medical emergencies in particular, a physician team leader is cognitively tasked not only with making appropriate medical decisions but also employing crisis resource management techniques (like leadership skills, situational awareness, communication skills and resource utilization) to optimize patient care (hicks, bandiera, & denny, 2008). it logically follows that cognitive efficiency in medical decision-making will more readily allow the physician leader to perform these simultaneous crisis resource management tasks to a higher level given the real constraints of human working memory. anecdotally, cognitive efficiency seems to evolve with experience. the “anatomy” of working memory is thought to change with the development of expertise and it is likely that certain clinical tasks cognitively load experts and novices in different ways (szulewski et al., 2015). this evolution of the thinking process is tied to expertise development. 6. typical experimental conditions for pupillometry studies as outlined in this review, researchers have successfully used pupillometry as a cognitive load measurement tool in numerous experimental conditions. these range from relatively simple experiments where pupillometric data are gathered as participants are presented with written or verbal questions and are tasked to solve problems in fields including arithmetic and language, among others. other experiments involve the use of different stimuli including photographs or even simulated driving environments. in medical applications, pupillometry has similarly been used in various settings including test-taking as well as more high fidelity environments like simulation and actual physician-patient clinical encounters. the task instructions provided to participants are equally variable and range from solving provided problems to performing operations in live surgical environments. 7. constraints of pupillometry though it is clear that pupillometry provides useful information about both visual as well as nonvisual stimuli, the technology has a number of constraints. until relatively recently, accurate pupillometry studies required cumbersome experimental environments and tedious data collection and analysis. although some of these issues have been addressed with new technology, the cost of this technology poses new financial barriers for certain researchers on smaller budgets. this is especially relevant for those part-time researchers who may want to incorporate pupillometry into their professional and teaching duties, like academic physicians. this reality suggests that, for the time being, given the costs, pupillometry research is more likely to occur at a theory-building level. as a result, some of its potential benefits in adjusting task difficulty for an individual learner and individualizing and optimizing education will remain elusive until the technology becomes cheaper and more readily available for teachers. szulewski et al | f l r 63 another significant constraint of the technology relates to the accommodation and light reflexes. although pupillometry provides consistent and fairly easy-to-interpret data in experimental conditions of constant ambient light and focus distance, data output in real-world conditions is suboptimal. as previously discussed in this review, the index of cognitive activity has been designed in an effort to overcome some of these obstacles. although this technique allows for extraction of valuable information from large data sets with changing ambient light, the results are more coarse and provide less precise and detailed information about shorter-term cognitive changes that might be relevant in studying precise comparisons between groups performing shorter tasks (klingner, kumar, & hanrahan, 2008). in addition, because this metric is a commercial product and its algorithm is not made publicly available, it cannot be replicated nor adequately studied. as a result, it is of limited benefit to researchers. another consideration in pupillometry research is participant age. older individuals generally have pupils that are smaller and are more restricted in their ability to dilate compared to younger people (holmqvist et al., 2011; piquado, isaacowitz, & wingfield, 2010). since many studies compare cognitive load between novices (who are usually younger) to experts (who are generally older), this might lead to confounding, as a smaller pupil diameter change may be due to a combination of increased age as well as decreased cognitive load. cognitive load researchers should be aware of this issue and either control for participant age (where possible) or correct for it. correction measures include expressing pupil size changes relative to a baseline measurement and/or age-adjusting for pupil size and reactivity based on participants’ pupil responsiveness to a range of experimental light stimuli (piquado et al., 2010). finally, the accuracy of pupillometric measurement is dependent to some degree on gaze position, with greater systematic error occurring when the eye is looking away from the eye-tracker’s camera (brisson et al., 2013). different eye-tracking devices attempt to correct for this error, but the accuracy of pupillometric measures suffers from variable quality under these conditions. this is especially relevant for researchers studying cognitive load where the participant’s gaze may move away from centre. 8. conclusion pupillometry is a robust and reliable method for studying cognitive load. since its inception as a scientific field in the 1960’s, it has evolved greatly. the development of new technology to measure pupil size that can electronically gather pupil data at high rates has led to the increased use of pupillometry in diverse fields. despite the inherent constraints of the technique including interference by luminance and its cost, pupillometry remains a promising metric for researchers to utilize in the study of cognitive load. it can provide insights into the human thinking process that would otherwise be unobservable. it has a particularly promising role in the field of medicine and in the study of physician expertise development. utilizing pupillometry to better understand and optimize physician cognitive load (and overload) is clinically relevant and has the potential to directly impact medical education and ultimately patient care. keypoints pupillometry is a robust method of quantifying cognitive load. otherwise unobservable insights into cognitive processes can be gleaned with the use of pupillometry. pupillometry research in medicine is contributing to a better understanding of expertise development across medical domains. szulewski et al | f l r 64 despite its benefits, pupillometry data in real-world applications suffers in quality as a result of the light and accommodation reflexes. references ahern, s., & beatty, j. (1979). pupillary responses during information processing vary with scholastic aptitude test scores. science, 205(4412), 1289-1292. aston-jones, g., & cohen, j. d. (2005). an integrative theory of locus coeruleus-norepinephrine function: adaptive gain and optimal performance. annu. rev. neurosci., 28, 403-450. doi: 10.1146/annurev.neuro.28.061604.135709 beatty, j. (1982). task-evoked pupillary responses, processing load, and the structure of processing resources. psychological bulletin, 91(2), 276. beatty, j., & kahneman, d. (1966). pupillary changes in two memory tasks. psychonomic science, 5(10), 371-372. brisson, j., mainville, m., mailloux, d., beaulieu, c., serres, j., & sirois, s. (2013). pupil diameter measurement errors as a function of gaze direction in corneal reflection eyetrackers. behavior research methods, 45(4), 1322-1331. doi:10.3758/s13428-013-0327-0 burgess, p. w., & simons, j. s. (2005). theories of frontal lobe executive function: clinical application. in p. w. halligan & d. t. wade (eds.), effecitveness of rehabiliation for cognitive deficits (pp. 211231). new york: oxford univeristy press. diamond, a. (2013). executive functions. annual review of psychology, 64, 135-168. doi:10.1146/annurevpsych-113011-143750 ericsson, k. a., & kintsch, w. (1995). long-term working memory. psychological review, 102(2), 211. hess, e. h., & polt, j. m. (1960). pupil size as related to interest value of visual stimuli. science, 132(3423), 349-350. hess, e. h., & polt, j. m. (1964). pupil size in relation to mental activity during simple problem-solving. science, 143(3611), 1190-1192. hicks, c. m., bandiera, g. w., & denny, c. j. (2008). building a simulation‐ based crisis resource management course for emergency medicine, phase 1: results from an interdisciplinary needs assessment survey. academic emergency medicine, 15(11), 1136-1143. doi: 10.1111/j.15532712.2008.00185.x holmqvist, k., nyström, m., andersson, r., dewhurst, r., jarodzka, h., & van de weijer, j. (2011). eye tracking: a comprehensive guide to methods and measures: oup oxford. jiang, x., zheng, b., tien, g., & atkins, m. (2013). pupil response to precision in surgical task execution. studies in health technology and informatics, 184, 210. kahneman, d., tursky, b., shapiro, d., & crider, a. (1969). pupillary, heart rate, and skin resistance changes during a mental task. journal of experimental psychology, 79(1p1), 164. klingner, j., kumar, r., & hanrahan, p. (2008). measuring the task-evoked pupillary response with a remote eye tracker. paper presented at the proceedings of the 2008 symposium on eye tracking research & applications. klingner, j., tversky, b., & hanrahan, p. (2011). effects of visual and verbal presentation on cognitive load in vigilance, memory, and arithmetic tasks. psychophysiology, 48(3), 323-332. doi: 10.1111/j.14698986.2010.01069.x laeng, b., sirois, s., & gredebäck, g. (2012). pupillometry a window to the preconscious? perspectives on psychological science, 7(1), 18-27. doi: 10.1177/1745691611427305 lang, g. k. (2015). ophthalmology: thieme. leppink, j., paas, f., van gog, t., van der vleuten, c. p., & van merriënboer, j. j. (2014). effects of pairs of problems and examples on task performance and different types of cognitive load. learning and instruction, 30, 32-42. szulewski et al | f l r 65 marshall, s. p. (2002). the index of cognitive activity: measuring cognitive workload. paper presented at the human factors and power plants, 2002. proceedings of the 2002 ieee 7th conference on. paas, f., tuovinen, j. e., tabbers, h., & van gerven, p. w. (2003). cognitive load measurement as a means to advance cognitive load theory. educational psychologist, 38(1), 63-71. doi: 10.1207/s15326985ep3801_8 palinko, o., kun, a. l., shyrokov, a., & heeman, p. (2010). estimating cognitive load using remote eye tracking in a driving simulator. paper presented at the proceedings of the 2010 symposium on eyetracking research & applications. piquado, t., isaacowitz, d., & wingfield, a. (2010). pupillometry as a measure of cognitive effort in younger and older adults. psychophysiology, 47(3), 560-569. doi:10.1111/j.1469-8986.2009.00947.x richstone, l., schwartz, m. j., seideman, c., cadeddu, j., marshall, s., & kavoussi, l. r. (2010). eye metrics as an objective assessment of surgical skill. annals of surgery, 252(1), 177-182. doi: 10.1097/sla.0b013e3181e464fb schulz, c., schneider, e., fritz, l., vockeroth, j., hapfelmeier, a., wasmaier, m., . . . schneider, g. (2011). eye tracking for assessment of workload: a pilot study in an anaesthesia simulator environment. british journal of anaesthesia, 106(1), 44-50. doi: 10.1093/bja/aeq307 sweller, j. (2010). element interactivity and intrinsic, extraneous, and germane cognitive load. educational psychology review, 22(2), 123-138. doi: 10.1007/s10648-010-9128-5 szulewski, a., fernando, s. m., baylis, j., & howes, d. (2014). increasing pupil size is associated with increasing cognitive processing demands: a pilot study using a mobile eye-tracking device. open journal of emergency medicine, 2014. doi: 10.4236/ojem.2014.21002 szulewski, a., gegenfurtner, a., howes, d. w., sivilotti, m. l. a., & van merriënboer, j. j. g. (2016). measuring physician cognitive load: validity evidence for a physiologic and a psychometric tool. advances in health sciences education, 1-18. doi: 10.1007/s10459-016-9725-2 szulewski, a., roth, n., & howes, d. (2015). the use of task-evoked pupillary response as an objective measure of cognitive load in novices and trained physicians: a new tool for the assessment of expertise. academic medicine, 90(7), 981-987. doi: 10.1097/acm.0000000000000677 tien, t., pucher, p. h., sodergren, m. h., sriskandarajah, k., yang, g.-z., & darzi, a. (2015). differences in gaze behaviour of expert and junior surgeons performing open inguinal hernia repair. surgical endoscopy, 29(2), 405-413. doi: 10.1007/s00464-014-3683-7 whipple, b., ogden, g., & komisaruk, b. r. (1992). physiological correlates of imagery-induced orgasm in women. archives of sexual behavior, 21(2), 121-133. young, j. q., van merrienboer, j., durning, s., & ten cate, o. (2014). cognitive load theory: implications for medical education: amee guide no. 86. medical teacher, 36(5), 371-384. doi: 10.3109/0142159x.2014.889290 zheng, b., jiang, x., & atkins, m. s. (2015). detection of changes in surgical difficulty: evidence from pupil responses. surgical innovation, 22(6), 629-635. doi: 10.1177/1553350615573582 frontline learning research 5 special issue ‗learning through networks‘ (2014) 15-37 issn 2295-3159 corresponding author: kaisa hytönen, department of education, 20014 university of turku, finland, sakahy@utu.fi doi: http://dx.doi.org/10.14786/flr.v2i2.90 15 | f l r cognitively central actors and their personal networks in an energy efficiency training program kaisa hytönen a , tuire palonen a , kai hakkarainen a a university of turku, finland article received 17 february 2014 / revised 31 march 2014 / accepted 29 may 2014 / available online 15 july 2014 abstract this article aims to examine cognitively central actors and their personal networks in the emerging field of energy efficiency. cognitively central actors are frequently sought for professional advice by other actors and, therefore, they are positioned in the middle of a social network. they often are important knowledge resources, especially in emerging fields where standard knowledge exchange mechanisms are weak. by adopting a personal network approach, we identified the cognitively central participants of a one-year energy efficiency training program, studied the structure and heterogeneity of their personal networks and determined which features were relevant to achieving these cognitively central positions. at the end of the training, the social networking questionnaire was sent to 74 course participants. semi-structured interviews were conducted for the six mostcentral actors, whose personal networks were larger than those of the other participants. these six actors differed from each other in many respects; there did not appear to be a single explanation for why these persons achieved their central positions. in conclusion, we propose that becoming a cognitively central actor is an intricate process. it cannot be explained only, for instance, by actors’ educational backgrounds, the level of their previous energy efficiency knowledge or their field of know-how. to understand this phenomenon, we must examine which organizations such people come from and how their expert profiles, which are related to their fields and competences, fit into the wider context of energy efficiency. more research is needed to determine whether the results are only typical of emerging fields. keywords: personal networks; cognitive centrality; advice seeking; social network analysis; emerging fields; energy efficiency training programme hytönen et al. 16 | f l r 1. introduction in rapidly changing and complex environments and their associated emergent knowledge-laden global problems and challenges, professionals must share their knowledge and expertise (hakkarainen, palonen, paavola, & lehtinen, 2004) rather than rely on mere individual competencies. this study focuses on examining key experts who have crucial roles in adaptively coping with novel challenges and changing professional requirements emerging from swiftly transforming professional fields. the key experts are often considered to be exceptionally valuable networking partners and collaborators because they have strategic knowledge and competence as well as in-depth meta-level vision regarding a transforming multi-professional field. their knowledge and competence is likely to be seen valuable by colleagues because they are deliberately building personal networks to interconnect heterogeneous social resources, expertise and knowhow and reaching beyond their immediate peers and bridging professional fields, thereby changing the ecology of their professional learning. as a consequence, the key experts are most often sought for advice and assistance by those struggling with novel professional challenges. personal networking connections with key professionals and the expert cultures they represent are important in updating the expertise and skills needed for responding to the professional challenges of future working lives, especially in turbulent environments (lehtinen, hakkarainen, & palonen, in press). professionals must be able to solve unforeseen complex problems and to share knowledge and competences, often breaking the boundaries of traditional disciplines. energy efficiency is one of the rapidly developing fields that has emerged through the intersection of several professional domains. therefore, there does not appear to be one unified system to direct professional activity, and the standard knowledge exchange mechanisms are weak. cooperation between professionals from diverse fields, who master varying bodies of expertise and pursue divergent professional tasks and projects, plays an important role in energy efficiency work. extensive professional experience alone does not automatically guarantee a central professional position; deliberate and sustained efforts to work at the edge of competence and cultivate expertise play a critical role as well (bereiter & scardamalia, 1993). although developing efficient energy usage practices and meeting global and national standards and directives regarding energy efficiency are some of the most important challenges of the 21st century, there are no established educational methods and practices for cultivating associated expertise in finland. therefore, efforts to create multifaceted personal expert networks and informal learning seem to play a significant role in professional development and updating expertise (see the similar situation regarding magicians‘ expert networks in rissanen, palonen, pitkänen, kuhn and hakkarainen, 2013). in our previous study (hytönen, palonen, lehtinen, & hakkarainen, 2014), we examined whether a training model that we call academic apprenticeship education initiated in finland in 2009, could help increase professional networking ties among participants. the study revealed that this energy efficiency training program, organized for actors who were already working on expert-level tasks, did not effectively support comprehensive networking or the creation of a knowledge exchange forum among the participants. however, there were some key professionals who were able to create valuable personal networking connections and contribute to professional collaboration during the training. this paper focuses on them. 1.1 conceptual background in complex and changing professional environments, targeted knowledge or competence is not always easily found or verified. in order to acquire new knowledge and appropriate novel professional practices as well as find required professional help and advice, professionals have to deliberately build and extend their personal networks (see pataraia, margaryan, falconer, & littlejohn, 2013). resources obtained through personal networks can benefit professional development by providing access to networking partners and associated professional support and opportunities for informal learning. in order to obtain new knowledge, many key experts have to rely on their personal social networks, reaching beyond the boundaries of their workplace organizations rather relying merely on traditional institutional resources (nardi, whittaker, & schwarz, 2000). however, to benefit from personal professional learning networks, workers must have cultivated networking competencies in terms of having the capability of finding and creating hytönen et al. 17 | f l r useful connections, as well as maintaining and activating these connections when needed (gruber, lehtinen, palonen, & degner, 2008; rajagopal, joosten-ten brinke, van bruggen, & sloep, 2012). the factors influencing the choices involved in building, maintaining and activating personal professional networks are related to (a) the trajectories of an actor‘s personal professional interests and needs, (b) the features of the contacts, such as the like-mindedness, benevolence and the potential learning and collaboration value of the relationship, and (c) the characteristics of the work environment (rajagopal, joosten-ten brinke, van bruggen, & sloep, 2012). according to the homophily principle, people often interact and create strong ties with those who have similar characteristics to themselves (kleinbaum, stuart, & tushman, 2013; mcpherson, smith-lovin, & cook, 2001; reagans, 2011). it follows that networks are often homogeneous in nature; people are more likely to create contacts with others who share the same gender, age, educational level, professional group and structural position. therefore, homophily often impacts the information people receive from their personal social networks, the attitudes they form and the interactions they experience (lozares, verd, cruz, & barranco, 2013; mcpherson, smith-lovin, & cook, 2001). professionals functioning in such networks often share a great deal of their knowledge and practices, immediately understanding each other (wenger, 1998). homogeneous professional networks do not, however, provide an adequate way of coping with the challenges involved in profound transformations of professional practices extending across multiple fields, such as in the case of energy efficiency work; personal networks are rich repositories of professional knowledge if they involve people with heterogeneously distributed knowledge and expertise and, thus, provide access to the resources embedded in these social relations (lin, 2001). cultivating strategic competence in complex and extended professional fields, such as energy efficiency, appears to require deliberate efforts of creating networking connections across the boundaries of several fields of professional activity (akkerman, admiraal, simons, & niessen, 2006). such efforts of crossing boundaries between professional cultures are likely to characterize networking activities of key experts, allowing them to mediate knowledge across the borders of different cultures and environments and bridge various fields of expertise with one another. key persons are positioned in the middle of the communication structure and therefore have access to extended pools of knowledge and diverse sources of information. in the literature, actors with strategic networking positions mediating, translating and transmitting knowledge and good practices and creating connections between diverse people between different cultures (meyer, 2010) are referred to as knowledge brokers (sverrisson, 2001), gatekeepers (morrison, 2008), stakeholders (krueger, page, hubacek, smith, & hiscock, 2012; svendsen & laberge, 2005), stars (borgatti, mehra, brass, & labianca, 2009) and hubs (barabasi, 2002). sverrisson (2001) has distinguished between three approaches to knowledge brokering. networking brokerage refers to connecting people, knowledge orientated brokerage relates to translating concepts and theories across disciplines that are critical to applying knowledge in complex projects, and organizational or technological brokerage involves facilitating novelty and innovation. overall, people in the middle of the social network often disseminate knowledge culture by sharing information with people around them and between workplace organizations and their surrounding environments, and by building bridges among people and between bodies of knowledge (burt, 1999). to cope with the challenges of rapidly transforming environments of professional activity, key experts have to cultivate practices of adaptive expertise (hatano & inagaki, 1986). such practices involve the cultivation of competency in successfully dealing with challenging, novel and unanticipated professional problems instead of clinging to old routines. adaptive experts are those who deliberately invest resources released by accumulating experience in new learning and seek challenges that assist and elicit their learning and the development of expertise. toward that end, many participants create deliberately novel networking connections and engage in inspiring encounters with heterogeneous networking partners. the creation of versatile networking connections and sustained sharing of professional expertise elicits the development of relational expertise, which is understood as the capability to productively tailor and fine-tune personal expertise to create joint or shared competence within communities and organized groups of experts and professionals (edwards, 2010). people working in the emerging fields often come from different working sectors and representing various fields of know-how when combining different fields of expertise appears to hytönen et al. 18 | f l r be important (see mieg, 2006). relational expertise recognizes the importance of resources provided by the different actors and the relevance of generating mutual understanding and shared goals over the borders of different fields of expertise, enabling collaboration (edwards, 2010). one way of assessing key experts‘ positions within a social network is the number of networking partners seeking their advice. advice networks are comprised of relations through which participants share resources, supporting the completion of their assignments (sparrowe, liden, wayne, & kraimer, 2001). who people contact when needing knowledge and advice and the reasons for seeking advice from these people has been studied (creswick & westbrook, 2010; nebus, 2006), as has the kind of knowledge sought in advice networks (cross, 2004; cross, borgatti, & parker 2001). motivations for asking for professional advice from someone seem to be related to the relevance and value of their information, the level of interpersonal trust (levin & cross, 2004), the advice seeker‘s perceptions of the knowledge source‘s expertise and credibility, accessibility, the expectations on how the contact will respond, and the assessed value and costs of seeking advice (nebus, 2006). investigations have revealed that information and advice relationships cultivated by people provide several types of knowledge, such as answers to know-what, knowhow and know-who questions as well as meta-knowledge concerning where information needed for answering these questions may be found. in addition, knowledge received from advice networks might help to think differently about problems faced as well as validate and legitimize solutions and plans made (see cross, 2004). attainment of a central networking position in advice networks is often related to personal characteristics, such as an in-depth professional commitment, motivational engagement (aalbers, doflsma, & koppius, 2013), a high level of professional performance (sparrowe, liden, wayne, & kraimer, 2001) and transformational leadership (bono & anderson, 2005). in this study, we adopt a personal (egocentric) network approach to identify and examine key experts in an energy efficiency training program whose professional knowledge the other course participants frequently sought to share. we call such key experts, whose cognitive achievements are shared by their professional peers, cognitively central actors. the concept of cognitive centrality is derived from studies on group decision making in a social network framework (kameda, ohtsubo, & takezawa, 1997). kameda, ohtsubo and takezawa (1997) suggested that the more knowledge and competence a person shares with the other group members, the more cognitively central position he or she has in the group. cognitively central group members who contribute intensively in collective problem-solving efforts are more influential in decision-making situations than peripheral members (stasser, abele, & vaughan parsons, 2012). here, the concept of cognitively central actors will be used to refer to course participants who were positioned in the middle of the social network, have valuable, extended and heterogeneous networking connections, and, therefore, provide other participants with new and relevant knowledge, competences and assistance more often than others (see kameda, ohtsubo, & takezawa, 1997; palonen, hakkarainen, talvitie, & lehtinen, 2004). in many cases, they appeared to have a high level of relational expertise in terms of having metaknowledge regarding the social distribution of relevant knowledge across professional networks (i.e., knowing who knows what in a professional network). traditionally, it is thought that persons who are often sought for professional and work-related advice are more knowledgeable and have more expertise than others. however, this is not necessarily the case in the emerging fields of complex professional activity where expertise needed for solving emerging problems is radically distributed or may not exist to begin with. more symmetric advancement of heterogeneously distributed knowledge (scardamalia, 2002) by different participants may characterize such situations. under such conditions, participants having a comprehensive vision of the future of their field as well as a high level of discernment, that is, a capability of assessing knowledge relationally in context (facer, 2011), may become cognitively central participants. this paper examines more closely who the cognitively central participants are in the field of energy efficiency, and it attempts to determine the possible reasons or personal features for achieving this kind of important networking position in the emerging field. we aim to understand why certain key persons are contacted and asked for knowledge and advice more often than others. personal networks offer illustrative ways to examine knowledge exchanges and communication in complicated environments by enabling the integration of individual and community level attributes. therefore, they enable the analysis of the properties of the one person ―owning‖ the network (ego) and the properties of people belonging to his or her network hytönen et al. 19 | f l r (alters), as well as the attributes of ego-alter ties and alter-alter ties (hakkarainen, palonen, paavola, & lehtinen, 2004). as a unit of analysis, personal networks, supplemented by other techniques, enabled us to look at network connections from different angles and across several levels and thus achieve a more accurate picture of multi-faceted and complicated social structures (fuhse & mützel, 2011). in all, we shall examine cognitively central actors‘ personal, social and organizational features relevant to achieving a central, strategic position among the participants in the energy efficiency training. 2. the aim of the study the purpose of this study is to examine the personal networks of those key energy efficiency professionals who are often sought for professional information and advice by other actors working in the field, in other words, the cognitively central actors. our specific focus is analysing how knowledge and competence sharing regarding energy efficiency issues was organized around particular persons and whether there were some features explaining why certain persons achieved a cognitively central position. the study was carried out in the context of a year-long energy efficiency training program. our hypotheses are as follows: 1) at the overall network level, cognitively central participants can be identified using an advice size indicator, that is, from whom the participants ask advice regarding their energy efficiency related problems. 2) a) at the ego-alter level, the structure of the cognitively central participants‘ personal networks differs from that of other course participants‘ so that their networks are bigger, denser and they have more broker capacity, that is, they connect the members in their personal networks. b) the central participants‘ personal networks are expected to be diverse in relation to their members‘ genders, university divisions, working sectors, educational backgrounds, previous experience-based knowledge of energy efficiency and the fields of their know-how. 3) at the ego level, the cognitively central participants have certain features that explain their prominent networking position. such features can be expected to relate to their personal attributes and affiliations. 3. methods 3.1 energy efficiency training program this study was conducted in the context of the one-year academic apprenticeship education program in the field of energy efficiency (hytönen, palonen, lehtinen, & hakkarainen, 2014). it was a pilot educational program organized for the first time in finland in 2010–2011. the energy efficiency training aimed to support the cultivation of energy efficiency expertise in the public and private sectors, promote professional networking between the actors in the field and encourage the sharing of good professional practices. three technical universities organized the training collaboratively: universities a (n = 29) and b (n = 28) organized education mainly for actors working in the public sector, and university c (n = 30) organized education for actors working in the private sector. fourteen participants working in the private sector participated in the educational training organized by universities a and b because there were not enough spaces for all willing private-sector participants at university c. altogether, 74 of 87 course participants completed the training program; 13 participants dropped out for various reasons. hytönen et al. 20 | f l r the energy efficiency training was based on real-life working practices and included theoretical studies and workplace learning. the theoretical studies were organized into seven contact days, including lectures, small group work and discussions. the first three and the last contact days were organized jointly for all course participants, but the remaining three days were organized separately for the public and private sector actors. the three separated contact days involved themes that were relevant especially to either the public or the private sector. the timespan and practices for organizing the contact days were the same for all three universities. about 70–80% of the active time in the training program was assumed to take place in the participants‘ workplaces, where the participants conducted a developmental study project. the developmental study project aimed to support participants‘ professional development as well as the development of the workplaces‘ energy efficiency practices. on the last contact day, each participant presented his or her developmental study project. networking between the participants was supported by small group work. in each university, the course participants were organized into five small groups of five to six members according to their places of residence. in addition to small group work taking place during the contact days, the small group members were advised to meet at least three times during the training to discuss their developmental study projects and provide peer support. in addition, the course participants were encouraged to use the virtual learning environments provided by each university to support open discussion and knowledge exchange. however, the small group meetings and the use of the virtual learning environments were not controlled by any means. furthermore, each participant was assigned an academic expert advisor on behalf of the universities and a workplace supervisor from his or her workplace organization. their role was to provide professional support for participants in their developmental study projects and the process of workplace learning. the practices of the energy efficiency training are presented in further detail in hytönen, palonen, lehtinen and hakkarainen (2014). 3.2 participants at the overall level of analysis, all course participants were asked to participate in this study. participation was voluntary, and the energy efficiency training was independent of this research. the participants were engineers, architects and other professionals with a mastersor bachelors-level education and varied lengths of experience in professional practices related to energy efficiency. at the ego-alter level of analysis, the participants were the 40 members (alters) of the central participants‘ personal networks in the context of the energy efficiency training. personal networks included only other course participants; the academic expert advisors, the workplace supervisors and other colleagues were not investigated. twentyfour of the alters were male and 16 were female. fifteen of the alters participated in the education organized by university a, 12 participated in the education organized by university b and 13 participated in the education organized by university c. more detailed information regarding the alters is provided in the results section. at the ego level of analysis, the participants in the study were six cognitively central actors from the energy efficiency training who were identified from all the course participants by analysing adviceseeking in the first section of the analysis. they are described in more detail in the results section. 3.3 social network methods network data were collected by administering an online social networking questionnaire to all 74 course participants (males, 50; females, 24) at the end of the training, out of whom 52 responded; the response rate was 70%. we also collected networking data in a similar way in the beginning of the training, but this study is based only on the latter data. the results concerning the changes in networking ties during the training are reported elsewhere (hytönen, palonen, lehtinen, & hakkarainen, 2014). the networking questionnaire involved a list of the names of all course participants, and in relation to one another, the respondents were asked to assess the following: 1) from whom they sought advice regarding energy efficiency and 2) with whom they collaborated in terms of energy efficiency activity. to measure the strength of the networking relations, the respondents were asked to rate each of these items on a valued scale of 0 (no connection), 1 (a connection) or 2 (a strong connection). hytönen et al. 21 | f l r a social network analysis (sna) was conducted via ucinet 6 (borgatti, everett, & freeman, 2002). we examined both the advice-seeking network, that is, how the participants sought energy efficiency information from one another, and the collaboration network, that is, how the participants collaborated with one another regarding energy efficiency issues. sna was conducted at the overall network level and the egoalter level. the different levels of analysis provided complementary dimensions for examining the cognitively central participants‘ networking. regarding the overall network, multidimensional scaling (mds) and advice size variables were used. in relation to the ego-alter level, the structure of connections between ego and alters was examined using different networking methods. information about the features of alters was collected by a networking questionnaire that was developed according to earlier studies (palonen, 2003). at the overall network level of analysis, both the advice-seeking and collaboration networks were examined. from these two, the advice-seeking network was used to identify the cognitively central participants in the training because it is asymmetric in nature and does not require reciprocal networking connections. therefore, it functions well as an indicator of a person‘s cognitive centrality (palonen, hakkarainen, talvitie, & lehtinen, 2004; sparrowe, liden, wayne, & kraimer, 2001). the cognitive centrality of the course participants was examined by calculating the centrality value (advice size), which indicates the amount of information that a person provides to the other members of the network. this was done using freeman‘s in-degree measurement, which revealed how many course participants sought energy efficiency advice from the actor in question, that is, the number of incoming networking linkages based on peer evaluation. the analysis indicated how significant a role an actor‘s expertise played in the social network and thereby allowed one to identify the cognitively central actors among the participants. the analysis was conducted for the dichotomized network, so the frequency of communication was not analysed. further, the network cohesion for the overall advice-seeking network was analysed via a density measure that characterized the number of existing networking ties in relation to all possible ties. to illustrate the structure of the overall network of all course participants and the structural position of the cognitively central participants, the advice-seeking and collaboration networks were visualized using the spindel visualization tool (see www.spindel.fi) using the participants‘ network distances, which were provided by mds techniques. at the ego-alter level, to deepen the analysis, the structure and heterogeneity of the central participants‘ personal networks were examined. the advice-seeking and collaboration networks were merged for the following analyses by summing them up, and the merged network was dichotomized (cut point 0). the egocentric network was used as the unit of analysis. the structure of the central participants‘ personal networks was analysed by size, density and a brokering index. size indicates the number of alters the ego is directly connected to; central members are expected to have a high number of contacts. density was calculated among the central participants‘ network members; the number of ties was divided by the number of pairs multiplied by 100. a high density in the alter network indicates a low brokering or mediating role for a given ego. on the other hand, a low density indicates that the ego‘s position in the alter network is crucial. the brokering index is the number of times an ego lies on the shortest path between two alters. it is a parallel indicator for knowledge mediating. an undirected type of ego neighbourhood was used, meaning that all actors connected to and from an ego were considered (borgatti, everett, & freeman, 2002). the mannwhitney u-test was used to analyse whether the structure of the central participants‘ personal networks differed from the structure of all other course participants‘ personal networks. the heterogeneity of the central participants‘ personal networks was analysed by comparing the various properties among alters, as well as the properties between the egos and alters. first, we identified the alters by examining the egos‘ neighborhood in advice-seeking and collaboration. second, we classified all participants in terms of the university they belonged to, educational background, working sector, gender, level of previous experience-based knowledge in energy efficiency and field of know-how. the estimation of the alters‘ previous energy efficiency knowledge was based on their self-reports. the central participants‘ personal networks were visualized using cytoscape. the advice-seeking and collaboration networks were merged for the visualizations. hytönen et al. 22 | f l r 3.4 semi-structured interviews and qualitative content analysis semi-structured interviews were conducted with all the cognitively central actors to complement the social networking data at the ego level of analysis. the interviews were carried out to examine the features of the cognitively central participants and the possible reasons they achieved a central networking position among energy efficiency workers. data collection was carried out in two phases. four of the six central participants identified were interviewed, both in the beginning and at the end of the training, as a part of broader data collection. after we conducted sna and identified the cognitively central participants, we complemented the interviews by asking them to assess the possible reasons for their central networking positions. at this stage, the two remaining central participants were interviewed as well. the interview themes addressed the participants‘ educational backgrounds, work experiences, current work assignments and professional roles in relation to energy efficiency; their reasons for attending the training; their views on the energy efficiency field; their networking with the other course participants and other energy efficiency professionals, future prospects of developing energy efficiency expertise and their own opinions regarding the possible reasons for their cognitive centrality. the interviews were audio recorded and transcribed by the first author. qualitative content analysis was performed using atlas.ti 6.2. the analysis was conducted by identifying expressions related to the themes of adaptive expertise, relational expertise, disseminating knowledge culture and knowledge brokering. content was identified and clustered independently by two researchers. 4. results 4.1 identifying the cognitively central participants at the overall network level at the overall network level, we identified the cognitively central actors of the energy efficiency training program. the density analysis for the overall network revealed that 5% (sd = 21.8) of all potential networking linkages were present in the advice-seeking network. all course participants‘ cognitive centrality was analysed via freeman‘s in-degree measure in the advice-seeking network. the measure is based on peer evaluation, and it reveals how many course participants have selected the actor in question as an information source. the cognitively central participants were selected on the basis of their high in-degree value, i.e., a minimum of seven linkages, as compared to the average for all course participants (m = 3.7; sd = 2.0) in the advice-seeking network (see table 1). we selected six actors (a20, a26, b2, b21, c17 and c23) who were most often sought advice by their peers. there were two central actors from each university. multidimensional scaling (figure 1) representing the overall network of all course participants revealed that the central actors from the public sector universities (a20, a26, b2 and b21) were located in the middle of the network, indicating that they were in close connection with participants from both public sector universities (see the video of figure 1). central participant c17 from university c appeared to be connected mainly with the other private sector participants. however, the other central participant from the private sector (c23) was positioned between the private and public sector universities. overall, the course participants from universities a and b were clustered more closely than the participants from university c. hytönen et al. 23 | f l r figure 1. overall network. the mds figure is based on collaboration ties, whereas lines reflect advice ties. the figure, visualized using spindel tools (www.spindel.fi), reveals how the central participants were positioned in the network of all course participants. the colour code in the graphs represents the university that the person comes from: red, university a; green, university b; blue, university c. the central actors are indicated by the large nodes and personal numbers. click here to start the video. 4.2 central participants’ personal networks at the ego-alter level at the ego-alter level of ties, we examined the structure and heterogeneity of the central participants‘ personal networks. the structure of the personal networks was assessed using the ego networks‘ basic measures, which are reported in table 1. two of the central participants, a26 and b21, did not respond to the networking questionnaire. thus, their measures are based only on information provided by other course participants. we used a mann whitney u-test to analyse whether the structure of the central participants‘ personal networks differed from the structure of all other course participants‘ personal networks. it appeared that there was a statistically significant difference in relation to the size (z = -3.368; p = .001), which was self-evident, and density (z = -2.009; p = .045) of the personal networks, as well as the brokering index (z = -3.275; p = .001). to conclude, in addition to the fact that central members were most often asked for advice (that was the defining criterion), they had larger networks that were relatively sparse, indicating their own mediation role, which was also shown by the broker indicator. a20 had an especially large network, in which her own contribution was important and her brokering role was essential. https://www.youtube.com/watch?v=hbajsn_i3ty&feature=youtu.be hytönen et al. 24 | f l r table 1. in-degree and ego network measures in-degree measures size density (%) broker a20 9 19 8 158 a26* 10 10 13 39 b2 7 11 23 42.5 b21* 9 9 8 33 c17 7 9 39 22 c23 8 11 22 43 m 8.3 11.5 18.8 56.3 sd 3.8 11.9 50.5 measures for all other course participants (does not include central participants‘ measures) m 3.7 5.8 38.7 16.0 sd 2.0 3.7 29.7 29.7 * a26 and b21 did not respond to the networking questionnaire, and, therefore, their measures are based only on information provided by other course participants. the heterogeneity of the central participants‘ personal networks was examined by analysing the network alters‘ university divisions, working sectors, educational backgrounds, genders, previous experience-based knowledge of energy efficiency (self-reported) and field of know-how. in table 2 (see appendix 1), we have provided the frequencies of alters belonging to each central participant‘s personal network, indicating the heterogeneity of the networks. figure 2. network members‘ working sector. the colour code represents the working sector of the participants: green, public sector; blue, private sector. the large spheres represent the cognitively central participants and the small ones represent their network alters. for every participant, we have provided a personal number and a code identifying the university. the six central participants‘ personal networks are merged for the visualization. alter-alter ties are not represented in the figure. hytönen et al. 25 | f l r figure 2 is a visualization of the central participants‘ personal networks in terms of their alters‘ working sectors. all personal networks have been merged into the same figure. the results indicate that the heterogeneity of the central participants‘ networks varied in terms of alters‘ home universities and working sectors, that is, whether they came from the public or private sector (see table 2 and figure 2). a20‘s and c23‘s personal networks were the most heterogeneous in this respect; they included rather even amounts of actors from both the public and private sectors and from all three universities. b2, who worked in the private sector but participated in the public sector education, had contacts with only the participants from the public sector universities. obviously, participating in home university activities had more influence than the working sector as such. figure 3. network members‘ educational background. the colour code represents the educational background of the participants: orange, engineer; blue, architect; green, other; white, information missing. the large spheres represent the cognitively central participants and the small ones represent their network alters. for every participant, we have provided a personal number and a code identifying the university. the six central participants‘ personal networks are merged for the visualization. alter-alter ties are not represented in the figure. furthermore, the public sector actors (a20, a26, b2 and b21) had varied educational backgrounds (see figure 3 and table 2), as did their alters, whereas the personal networks of c17 and c23, who worked in the private sector, had low levels of variety in this respect. with one exception, they included only engineers. the personal networks of a20, b2, b21 and c23 were rather heterogeneous in respect to their alters‘ know-how, whereas a26‘s and c17‘s personal networks were more homogeneous; in the personal network of a26, there were many alters doing either land use planning or construction planning in the public sector, and the majority of c17‘s alters were industrial planners in the private sector (see figure 4 and table 2). hytönen et al. 26 | f l r figure 4. network members‘ filed of know-how. the shape and colour code represent the field of know-how of the participants: circle: red, land use planning; orange, construction planning; brown, environmental surveillance; violet, other; white, information missing. square: blue, planning for industry; green, consultant/surveillance/planning; white, information missing. the large spheres represent the cognitively central participants and the small ones represent their network alters. for every participant, we have provided a personal number and a code identifying the university. the six central participants‘ personal networks are merged for the visualization. alter-alter ties are not represented in the figure. figure 5 reveals that in the personal networks of a20, a26, b21 and c23, there were nearly the same number of female and male alters (see also table 2). in the networks of b2 and c17, there were more participants from their own gender group. it appeared that in the private sector, the central participants‘ personal networks were more male-oriented. this could be explained by the fact that in the context of this particular energy efficiency training, males worked in the private sector more often than females. hytönen et al. 27 | f l r figure 5. network members‘ gender distribution. the colour code represents the genders of the participants: blue, male; red, female. the large spheres represent the cognitively central participants and the small ones represent their network alters. for every participant, we have provided a personal number and a code identifying the university. the six central participants‘ personal networks were merged for the visualization. alter-alter ties are not represented in the figure. figure 6 visualizes the central participants‘ personal networks in terms of their alters‘ previous selfreported, experience-based energy efficiency knowledge. the figure reveals that two of the central participants (a26 and b2) had little or no previous knowledge of energy efficiency (see also table 2). obviously, their central networking position is explained by something else. overall, in each central participant‘s personal network, there were alters with varying amounts of previous energy efficiency knowledge. in this respect, the personal networks of inexperienced participants did not differ from those of experienced participants. hytönen et al. 28 | f l r figure 6. network members‘ previous experience-based professional knowledge of energy efficiency. the colour code represents the level of participants‘ previous knowledge of energy efficiency: green, strong; blue, some; red, minor or none; white, information missing. the large spheres represent the cognitively central participants and the small ones represent their network alters. for every participant, we have provided a personal number and a code identifying the university. the six central participants‘ personal networks are merged for the visualization. alter-alter ties are not represented in the figure. 4.3 ego level: features for achieving the cognitively central position at the ego level of analysis, using the interview data, we examined who the cognitively central actors were and which features were relevant to achieving a central position in more detail. the cognitively central actors differed from one another in terms of age, educational background and the length of work experience (see table 3). in addition, they had different levels of previous energy efficiency-related knowledge in terms of their job description. the interviews revealed that there was not one common explanation as to why these six participants achieved cognitively central positions. instead, various features were emphasized. it is obvious that a central position was not achieved based only on the strength of personal characteristics but also on the basis of what kind of information the other participants were requesting from the cognitively central actors. therefore, the features relevant to achieving the central position are related to the nature of the central participants‘ expertise, their knowledge brokering roles or positions between various fields or cultures, the nature of their employers and their own attitudes towards energy efficiency. in addition, they appeared to be interested in pursuing careers in the energy efficiency field. central participant a20, ―a knowledge-sharing representative of an important organization‖, had strong and wide-ranging working experience in energy efficiency in both the public and private sectors. in her current workplace, a significant public organization, she worked as an energy efficiency expert. as the organization‘s ―internal help‖, she was responsible for ensuring that energy efficiency was taken into account in the organization‘s approaches and decisions, and she advised fellow workers on energy efficiency issues. a20 appeared to have versatile professional connections that supported her daily work, and she emphasized the importance of professional collaboration. her employer functioned as a forerunner in developing and implementing energy-efficient practices and operational models in the public sector. she hytönen et al. 29 | f l r considered it crucial to openly discuss and share the newest knowledge and experiences among actors working with energy efficiency issues in order to promote the development of the energy efficiency field, energy efficiency consciousness and good operational practices: ―i‘ve pretty openly adopted the orientation that i‘m just going to talk and give those ideas.‖ in her experience, it is important to freely discuss both successful and unsuccessful undertakings because this benefits the development of the entire domain. a20 herself assessed that her open attitude towards sharing all types of energy efficiency knowledge was the most important reason for her cognitively central position. by performing research, a20 deliberately aimed to expand her own know-how regarding energy efficiency as well as to produce new information. she highlighted the fact that even though plenty of theoretical and technical energy efficiency knowledge and expertise exists, it is important to produce more practical knowledge and real-life examples to help steer the work of actors working with energy efficiency issues. the challenge is also to produce intelligible energy efficiency knowledge for common people: ―about 80 percent of the others [populace] don‘t understand anything about basic facts if you don‘t translate them into images, and they don‘t need to, because i don‘t understand anything about basic medication. it‘s the doctor who tells me what i have to eat to cope with those symptoms.‖ table 3. background information for the central participants age gender education work experience (years) job description in relation to energy efficiency a20 35–39 female engineer 11–15 a central part of the job description a26 55–59 female architect 36–40 in the background b2 40–44 female engineer 11–15 in the background b21 30–34 male m.sc. 1–5 about half of the job description c17 25–29 male engineer 1–5 a central part of the job description c23 30–34 female m.sc. 1–5 a central part of the job description central participant a26, ―an experienced worker and ‗missionary‘‖, was an architect by training and, like a20, worked at a remarkable organization in the public sector. she did not have any actual experience in energy efficiency issues before participating in the energy efficiency training but did have a great deal of work experience in her own field. by participating in the energy efficiency training and other available education, a26 aimed to become a kind of ―internal energy efficiency consultant‖ in her employing organization: ―it‘s like i have this kind of a model currently in my mind, or that‘s developed, about how i can first get this workplace community educated about taking the importance of energy efficiency into consideration.‖ in this way, she wished to be able to raise the awareness of energy efficiency practices and deliver them to her employer; she described herself as ―a kind of a missionary‖, though she reported holding a peripheral position in her workplace, without any support from her superintendent. central participant b2, ―a gatekeeper for electrical engineering‖, worked in a small private company, although she participated in education that was organized mainly for the public sector actors. she had strong technical know-how related to electrical engineering. as an electrician, she worked with assignments that were not directly related to energy efficiency, and her previous energy efficiency knowledge was minor. however, she highlighted the fact that awareness of energy efficiency matters is increasing in electrical engineering because of changing legislation and the increasing demands of customers; in the future, designs will have to be sustainable in the long term and not ―only such easy fixes‖. b2 emphasized that in electrical engineering, actors are ―contemplating their navels‖ too much instead of collaborating with other domains. participating in the training and networking with the other course hytönen et al. 30 | f l r participants widened b2‘s own professional viewpoint and convinced her of the importance of networking and collaboration across the borders of professional fields: ―we often considered, together with the planners, before the basic elements of a construction project, how could energy efficiencies be defined and such, even before the building is on the table.‖ the field of electrical engineering was unfamiliar to the most of the other course participants, and therefore b2 herself was able to provide them with a new kind of knowledge, presenting a novel perspective on energy efficiency. central participant b21, ―a liaison and eco-man‖, was a m.sc. by training. his job description and know-how comprised mainly of eco-efficiency, thus including many aspects of energy efficiency: ―i am some sort of eco-man, so in a sense, when situations emerge in which i have to take a position on climate or energy issues, then i‘m involved in such projects.‖ in his workplace, b21 functioned as a coordinator and knowledge mediator between land-use-planning actors and environmental authorities regarding issues related to energy efficiency: ―it is just this kind of role of ‗combiner‘, because of course i don‘t know about energy issues as much an engineer from city energy [name changed]. on the other hand, he doesn‘t know anything about land-use planning. still, i‘m not such a great land use designer either, so we have several architects, but then again, they don‘t necessarily know anything about energy efficiency.‖ overall, b21 emphasized that cross-administrative and versatile professional network connections are important in dealing with daily assignments. his employer was a significant public organization that functioned as an example for smaller municipalities. b21 emphasized that, as a large organization, it has better resources with which to develop energy-efficient operational models than smaller municipalities: ―we have really been able to do the kind of development work that not many municipalities can afford or even have time for maybe, so in that sense, we‘ve got a pioneering role.‖ therefore, b21 had profitable energy efficiency related knowledge and advice that he could share with the other course participants working with similar questions. central participant c17, ―an adaptive expert in the industrial sector‖, worked in a private company. although he had only three years of working experience, he had developed strong expertise in energy efficiency issues in a particular industrial field in which his daily work assignments were directly related. c17 had acquired his current energy efficiency knowledge through a few years of purposeful and deliberate efforts toward self-development, and further, he aimed to achieve a comprehensive understanding of all kinds of energy efficiency matters: ―since i started working in this company, i‘ve tried to find an extensive vision for the industrial air pressure systems, their energy efficiency and industrial energy efficiency in general.‖ c17 highlighted the extreme importance of increasing the awareness of efficient energy usage in the industry so that energy efficient behaviour will become a natural and axiomatic part of daily routines, instead of being ―a mandatory chore‖. overall, c17 wished for more openness and interaction between those actors dealing with energy efficiency issues in order to promote the diffusion of good ideas and, more generally, the development of the entire energy efficiency field: ―my overall opinion, outside of this training in general, is a desire to pursue openness and open communication, like exchanging ideas and not holding back information.‖ he aimed to promote this himself by sharing new energy efficiency knowledge, information and perspectives with his colleagues, as well as to customers and other actors in the industry. he had a mission of ―starting, so to speak, to declare our message to our customers and collaborators and other possible parties‖. central participant c23, ―a bridge between the public and private sector‖, had a degree in environmental technology. therefore, she had a different educational background than the majority of her colleagues and other course participants, who were mainly engineers, and a less technical perspective on energy efficiency. she had become acquainted with energy efficiency matters in her current workplace, a private company; her job description included consultancy and planning related to various energy efficiency issues and projects. c23‘s clients were actors and organizations from both the private and public sectors, and therefore, she had gained wide-ranging knowledge and experience in various kinds of energy efficiency issues that could be exploited in industrial and public sector assignments. because of her professional position in the intersection of these two sectors, many participants already knew her before the training: ―i work on both the municipal and the industrial sides, which is probably why, in my training, the people on the municipal and industrial sides knew me. i was probably in the middle there.‖ c23 emphasized that networking and collaboration are required in the diverse energy efficiency field; she stated that it is hytönen et al. 31 | f l r important to have a network of professionals with various kinds of know-how to consult when help and advice are needed—a kind of meta-knowing about who-knows-who-knows-what (borgatti & cross, 2003): ―nobody can be an expert in everything, so it‘s good to know about people who know about some issues and to be able to create such [connections] if you end up working on some projects for customers.‖ to sum up, becoming a cognitively central actor is an intricate process that cannot be reduced to personal characteristics. it is related to the organizations that the actors represent, the expert profiles or competences that they have and how these complement the wider context. cognitive centrality is therefore not only an individual-level capacity. 5. discussion in this study, we relied on the personal network approach to examine which features were relevant to achieving a cognitively central networking and knowledge sharing position in the academic apprenticeship education program in the field of energy efficiency. in emerging fields such as energy efficiency, where standard knowledge exchange mechanisms are still weak, cognitively central members, whose professional knowledge is frequently sought by other actors, are expected to be very important knowledge resources for other members in the network in terms of mediating knowledge and creating connections between different professional cultures. the analysis revealed that the six most central participants differed from each other in many respects, including the length of their work experience, educational background, how much they were involved in energy efficiency and what kind of organizations they came from. whatever the reason, these participants were asked for energy efficiency-related information more often than the other participants, and their knowledge mediating role in energy efficiency issues was essential. thus, the results revealed that there was not a single shared feature that can explain why certain participants became more cognitively central than their peers. according to the homophily principle, people tend to interact more frequently with those who have similar characteristics to themselves, such as those with similar educational levels or members of a joint professional group (mcpherson, smith-lovin, & cook, 2001). the present analysis of the central participants‘ personal networks, in contrast, revealed that many of the networks were rather heterogeneous in nature, including a rich variety of people with different educational and working backgrounds, as well as professional and energy efficiency-related experiences. in particular, the personal network of central participant a20, who had the most important knowledge sharing position in the training, was outstandingly heterogeneous in nature. such heterogeneous resources are obviously needed for coping with a continuously changing environment. even though our previous study (hytönen, palonen, lehtinen, & hakkarainen, 2014) indicated that the energy efficiency training did not support participants in comprehensive networking, the creation of an occupational knowledge-exchange forum and the use of one another‘s complementary expertise on a large scale, the results of this study revealed that some course participants were able to find valuable new connections with people who had novel perspectives on energy efficiency and to cross the boundaries of their immediate professional fields (akkerman, admiraal, simons, & niessen, 2006). apparently, the cognitively central actors possessed knowledge that other course participants found usable, even though they did not necessarily represent the same professional context or culture (see edwards, 2010). cognitive centrality is obviously not related only to personal attributes, such as a high level of professional experience, previous energy efficiency-related knowledge or personal characteristics. it is also related to social contexts, for instance, the nature of the operational environments and employing organizations that the participants represented. in addition, the results indicated that the participants‘ forms of expertise and competence were relationally and contextually assessed (mieg, 2006); their fields of knowhow were not necessarily energy efficiency, but they had strategic and special knowledge in some particular area, such as electrical engineering, that was found useful by other course participants. in addition, the participants representing significant public sector organizations appeared to possess advanced and trustworthy knowledge that was valued by the other participants and that they needed in their own professional contexts (levin & cross, 2004). hytönen et al. 32 | f l r in advice-seeking networks, help is often asked for from persons presumed to be the most knowledgeable and having the strongest experience in the issue in question (nebus, 2006). cumulative individual experience is expected to increase individual proficiency (reagans, argote, & brooks, 2005). however, this investigation revealed that it is not only lengthy professional experience or strong expertise in energy efficiency that make a person cognitively central. other factors such as personal enthusiasm or energy efficiency awareness were, in some cases, more important than strong professional competency or an extensive experience in the field. it seemed to us that young workers with rather limited working experience may quickly acquire relatively strong expertise and become cognitively central knowledge mediating professionals if they deliberately attempt to increase their expertise and succeed in reaching considerable professional capability (ericsson, 2006; hatano & inagaki, 1986). this can be the case especially in emerging fields, in which there are no strong established paradigms and working cultures and where good operational practices are still developing. recent changes in the working world highlight the importance of multi-professional collaboration (edwards, 2010) and a role as a boundary-spanning knowledge broker for professionals (johri, 2008). in addition to mediating knowledge, the key persons acting as knowledge brokers often produce a new kind of brokered knowledge that has been assembled based on knowledge collected from different cultures (meyer, 2010). one essential reason for achieving a cognitively central networking position in the energy efficiency training program was obviously bridging the gaps between various professional cultures and working environments, that is, those between the public and private sectors and between disciplines. in these positions, the cognitively central participants were able to process, build and even create new energy efficiency knowledge to be utilized in novel situations and tasks. it appears to us that the three brokering roles introduced by sverrisson (2001) were present at least in some forms in the central participants personal networks; they obviously connected people working with the energy efficiency issues (networking brokerage); created and translated concepts, theories and new knowledge of energy efficiency (knowledge oriented brokerage); and facilitated innovations and good operational practices and new operational models (brokerage of organizational or technological novelties) in and between the public and private sector organizations. the results indicated that the knowledge mediating role of the central participants was important both in the energy efficiency training and in their larger working environments in terms of aiming to increase awareness of energy efficiency practices and disseminating them to their workplaces. in addition to efforts towards purposeful and continuous self-development (ericsson, 2006; hatano & inagaki, 1986), some cognitively central participants showed a strong willingness to promote the overall development of the energy efficiency field by systematically creating and sharing knowledge and working for the diffusion of good energy efficiency practices. in this, the importance of socially shared professional goals appeared to have essential role (edwards, 2010). in emerging fields, there is often a lack of a stable knowledge base and formal education, as is the case in the field of energy efficiency in finland, and therefore, professional learning takes place through informal and incidental learning (watkins, marsick, & fernández de álava, 2014; palonen, lehtinen, & boshuizen, 2014). finally, it is presumably not possible to determine all possible reasons why someone is a hub for communication. the interviews revealed that informal networking connections and collaboration had important roles in professional activities and development. one example of this was found in the context of participants‘ joint discussions related to everyday energy-efficient practices, such as cooling gardens in the summertime. informal and incidental learning happens without much external facilitation and often occurs unsystematically, and it is therefore difficult to elicit and understand from an outside perspective. 5.1 limitations and further steps one of the limitations of this study was that two of the central participants did not respond to the networking questionnaire. therefore, their data were based only on information given by the other course participants, and we were not able to examine those relationships that they themselves may have had with others, that is, outgoing linkages. in addition, only a limited number of the course participants were interviewed, and, therefore, more research is needed to generalize the results. however, this study demonstrates the potential value of the personal network approach in the study of professional knowledge hytönen et al. 33 | f l r exchange in complex environments. sna provided a useful multi-level approach for determining the cognitively central actors possessing strategic competence in multi-professional fields, studying their role in professional networking and knowledge exchange and examining both the social context and the characteristics of individual actors in these processes. personal networks are often studied via egocentric network interviews in which the participants (egos) are asked to list the alters belonging to their personal networks and to evaluate the relationship between themselves and the alters as well as between each individual pair of alters. in this study, we used the overall network data to study the cognitively central participants‘ personal networks. this approach allowed us to use ties incoming from other course participants to estimate cognitive centrality and to analyse the structure of the personal networks (mccarty & govindaramanujam, 2005) and to visualize the networks on both the overall (sociocentric) and personal levels (see mccarty, molina, aguilar, & rota, 2007). our study contributes to professional learning research by elaborating the concept of cognitive centrality and widening its use outside a small group research. this approach is useful, especially for extension studies. future studies should examine in detail what kind of advice is sought from the cognitively central participants and how it is related to the nature of their expertise. in addition, more research is needed to better understand the phenomenon of cognitive centrality and to discover whether the results found are typical for emerging fields but not generalizable to other contexts. keypoints methods of analysing personal social networks provide a functional unit of analysis for studying the personal and social features of knowledge exchange in complex environments. this study introduces a concept of cognitive centrality and the diverse reasons behind this phenomenon. the article explains how cognitive central actors can be identified and how the flow of advice is centralized in the context of professional networks. this study addresses professional development in an emerging field (i.e., energy efficiency) where the knowledge base is not yet stable or consolidated. the paper focuses on learning processes in the context between working life and higher education institutions and explicates the features that are essential there. acknowledgments research has been funded by futurex project that is part of european social fund programme and finnish ministry of education and culture (asko-project). we would like to thank otto and antti seitamaa for translating transcribed interviews from finnish to english. references aalbers, r., dolfsma, w., & koppius, o. (2013). individual connectedness in innovation networks: on the role of individual motivation. research policy, 42, 624–634. doi: 10.1016/j.respol.2012.10.007. akkerman, s., admiraal, w., simons, r. j., & niessen, t. (2006). considering diversity: multivoicedness in international academic collaboration. culture & psychology, 12, 461–485. doi: 10.1177/1354067x06069947. barabasi, l-l. (2002). linked: how everything is connected to everything else and what it means for business, science, and everyday life. cambridge, ma: perseus publishing. hytönen et al. 34 | f l r bereiter, c., & scardamalia, m. (1993). surpassing ourselves: an inquiry into the nature and implications of expertise. chicago, il: open court. bono, j. e., & anderson, m. h. (2005). the advice and influence networks of transformational leaders. journal of applied psychology, 90, 1306–1314. doi: 10.1037/0021-9010.90.6.1306. borgatti, s. p., & cross, r. (2003). a relational view of information seeking and learning in social networks. management science, 49, 432–445. borgatti, s. p., everett, m. g., & freeman, l. c. (2002). ucinet 6 for windows. harvard, ma: analytic technologies. borgatti, s. p., mehra, a., brass, j. p., & labianca, g. (2009). network analysis in the social sciences. science, 323, 892–895. doi: 10.1126/science.1165821. burt, r. s. (1999). entrepreneurs, distrust, and third parties: a strategic look at the dark side of dense networks. in l. l. thompson, j. m. levine, & d. m. messick (eds.), shared cognition in organizations: the management of knowledge (pp. 213–243). mahwah, nj: erlbaum. creswick, n., & westbrook, j. i. (2010). social network analysis of medication advice-seeking interactions among staff in an australian hospital. international journal of medical informatics, 79, e116–e125. doi:10.1016/j.ijmedinf.2008.08.005. cross, r. (2004). more than an answer: information relationships for actionable knowledge. organization science, 15, 446–462. doi: 10.1287/orsc.1040.0075. cross, r., borgatti, s. p., & parker, a. (2001). beyond answers: dimensions of the advice network. social networks, 23, 215–235. edwards, a. (2010). being an expert professional practitioner. london: springer. ericsson, k. a. (2006). the influence of experience and deliberate practise on the development of superior expert performance. in k. a. ericsson, n. charness, p. j. feltovich, & r. r. hoffman (eds.), the cambridge handbook of expertise and expert performance (pp. 683–703). cambridge: cambridge university press. fuhse, j., & mützel, s. (2011). tackling connections, structure, and meaning in networks: quantitative and qualitative methods in sociological network research. quality & quantity, 45, 1067–1089. doi: 10.1007/s11135-011-9492-3. gruber, h., lehtinen, e., palonen, t., & degner, s. (2008). persons in shadow: assessing the social context of high ability. psychology science quarterly, 50, 237–258. hatano, g., & inagaki, k. (1986). two courses of expertise. in h. a. h. stevenson, & k. hakuta (eds.), child development and education in japan (pp. 262–272). new york: freeman. hakkarainen, k., palonen, t., paavola, s., & lehtinen, e. (2004). communities of networked expertise. professional and educational perspectives. amsterdam: elsevier. hytönen, k., palonen, t., lehtinen, e., & hakkarainen, k. (2014). does academic apprenticeship increase networking ties among participants? a case study of an energy efficiency training program. higher education. doi: 10.1007/s10734-014-9754-9. johri, a. (2008). boundary spanning knowledge broker: an emerging role in global engineering firms. proceedings from 38th asee/ieee frontiers in education conference. saratoga springs, ny. kameda, t., ohtsubo, y., & takezawa, m. (1997). centrality in sociocognitive networks and social influence: an illustration in a group decision-making context. journal of personality and social psychology, 73, 296–309. doi: 10.1037/0022-3514.73.2.296. kleinbaum, a. m., stuart, t. e., & tushman, m. l. (2013). discretion within constraint: homophily and structure in a formal organization. organization science, 24, 1316–1336. doi: 10.1287/orsc.1120.0804. krueger, t., page, t., hubacek, k., smith, l., & hiscock, k. (2012). the role of expert opinion in environmental modelling. environmental modelling & software, 36, 4–18. doi: 10.1016/j.envsoft.2012.01.011. lehtinen, e., hakkarainen, k., & palonen, t. (in press). understanding learning for the professions: how theories of learning explain coping with rapid change. in s. billett, c. harteis, & h. gruber. (eds.), international handbook of research in professional and practice-based learning. dordrecht: springer. levin, d., & cross, r. (2004). the strength of weak ties you can trust: the mediating role of trust in effective knowledge transfer. management science, 50, 1477–1490. doi:10.1287/mnsc.1030.0136. hytönen et al. 35 | f l r lin, n. (2001). social capital. a theory of social structure and action. cambridge: cambridge university press. lozares, c., verd, j. m., cruz, i., & barranco, o. (2013). homophily and heterophily in personal networks. from mutual acquaintance to relationship intensity. quality & quantity. doi: 10.1007/s11135-0139915-4. mccarty, c., & govindaramanujam, s. (2005). a modified elicitation of personal networks using dynamic visualization. connections, 26, 61–69. mccarty, c, molina, j. l., aguilar, c., & rota, l. (2007). a comparison of social network mapping and personal network visualization. field methods, 19, 145–162. doi: 10.1177/1525822x06298592. mcpherson, m., smith-lovin, l., & cook, j. (2001). birds of a feather: homophily in social networks. annual review of sociology, 27, 415–444. meyer, m. (2010). the rise of knowledge broker. science communication, 32, 118–127. doi: 10.1177/1075547009359797. mieg, h. a. (2006). social and sociological factors in the development of expertise. in k. a. ericsson, n. charness, p. feltovich, & r. hoffman (eds.), the cambridge handbook of expertise and expert performance (pp. 743–760). cambridge: cambridge university press. morrison, a. (2008). gatekeepers of knowledge within industrial districts: who they are, how they interact. regional studies, 42, 817–835. nardi, b. a., whittaker, s., & schwartz, h. (2000). it‘s not what you know, it‘s who you know: work in the information age. first monday, 5. nebus, j. (2006). building collegial information networks: a theory of advice network generation. the academy of management review, 31, 615–637. doi: 10.2307/20159232. palonen, t. (2003). shared knowledge and the web of relationships. turku: painosalama. palonen, t., hakkarainen, k., talvitie, j,. & lehtinen, e. (2004). network ties, cognitive centrality, and team interaction within a telecommunication company. in h. p. a. boshuizen, r. bromme, & h. gruber (eds.), professional learning: gaps and transitions on the way from novice to expert (pp. 271– 294). dordrecht: kluwer academic publisher. palonen, t., lehtinen, e,. & boshuizen, h. p. a. (2014). how expertise is created in emerging professional fields. in s. billett, t. halttunen, & m. koivisto (eds.), promoting, assessing, recognizing and certifying lifelong learning: international perspectives and practices (pp. 131–149). dordrecht: springer. pataraia, n. margaryan, a., falconer, i., & littlejohn, a. (2013). how and what do academics learn through their personal networks. journal of further and higher education. doi: 10.1080/0309877x.2013.831041. rajagopal, k., joosten-ten brinke, d., van bruggen, j., & sloep, p. (2012). understanding personal learning networks: their structure, content and the networking skills needed to optimally use them. first monday, 17, 1–12. reagans, r. (2011). close encounters: analyzing how social similarity and propinquity contribute to strong network connections. organization science, 22, 835–849. reagans, r., argote, l., & brooks, d. (2005). individual experience and working together: predicting learning rates from knowing who knows what and knowing how to work together. management science, 51, 869–881. doi: 10.1287/mnsc.1050.0366. rissanen, o., palonen, t., pitkänen, p., kuhn, g., & hakkarainen, k. (2013). personal social networks and the cultivation of expertise in magic: an interview study. vocations and learning, 6, 347–356. doi: 10.1007/s12186-013-9099-z. scardamalia, m. (2002). collective cognitive responsibility for the advancement of knowledge. in b. smith (eds.), liberal education in a knowledge society (pp. 67–98). chicago: open court. sparrowe, r. t., liden, r. c., wayne, s. j., & kraimer, m. l. (2001). social networks and the performance of individuals and groups. academy of management journal, 44, 316–325. doi: 10.2307/3069458. stasser, g., abele, s., & vaughan parsons, s. (2012). information flow and influence in collective choice. group processes and intergroup relations, 15, 619–635. doi: 10.1177/1368430212453631. svendsen, a. c., & laberge, m. (2005). convening stakeholder networks: a new way of thinking, being and engaging. journal of corporate citizenship, 19, 91–104. hytönen et al. 36 | f l r sverrisson, á. (2001). translation networks, knowledge brokers and novelty construction: pragmatic environmentalism in sweden. acta sociologica, 44, 313–327. doi: 10.1177/000169930104400403. watkins, k. e., marsick, v. j., & frenández de álava, m. (2014). evaluating informal learning in the workplace. in t. halttunen, m. koivisto, & s. billett (eds.), promoting, assessing, recognizing and certifying lifelong learning (pp. 59–77). dordrecht: springer. wenger, e. (1998). communities of practice: learning, meaning, and identity. cambridge: cambridge university press. hytönen et al. 37 | f l r appendix 1 table 2. heterogeneity of personal networks university working sector education gender previous energy efficiency knowledge a b c public private engineer architect other not known m f strong some minor not known a20 10 6 3 12 7 11 3 2 3 10 9 7 4 5 3 a26 7 3 0 10 0 5 4 1 0 4 6 3 1 6 0 b2 7 4 0 10 1 6 2 2 1 2 9 5 1 4 1 b21 4 5 0 7 2 6 2 0 1 5 4 3 1 4 1 c17 0 0 9 0 9 7 0 0 2 8 1 3 2 1 3 c23 2 2 7 4 7 8 0 1 2 7 4 3 2 4 2 field of know-how public sector private sector land use planning construction planning environmental surveillance other not known industrial planning consulting, surveillance, planning not known a20 5 3 0 3 1 0 5 2 a26 5 5* 1 1 0 0 0 0 b2 5 1* 1 3 1 0 1 0 b21 3 3* 1 0 1 0 2 0 c17 0 0 0 0 0 5 2 2 c23 1 1 2 0 0 4 1 2 a the number in each column indicates how many alters the central participants have in their personal network in relation to specific indicators (university, working sector, educational, gender, previous energy efficiency knowledge and field of know-how). b * for land use planning and construction planning the expertise areas are overlapping and there are 2 persons that have been added to both columns. microsoft word möller et al_publication.docx           frontline learning research vol.4 no. 2 special issue (2016) 1 – 11 issn 2295-3159 __________________________ corresponding author: jens möller, kiel university, institute of psychology, department of educational psychology, olshausenstraße 75, 24118 kiel, germany. email address: jmoeller@psychologie.uni-kiel.de  doi:   http://dx.doi.org/10.14786/flr.v4i2.169   the generalized internal/external frame of reference model: an extension to dimensional comparison theory jens möllera, hanno müller-kalthoffa, friederike helma, nicole nagya, herb w. marshb akiel university, germany baustralian catholic university, king saud university article received 6 may / revised 24 july / accepted 3 november / available online 20 january abstract the dimensional comparison theory (dct) focuses on the effects of internal, dimensional comparisons (e.g., “how good am i in math compared to english?”) on academic selfconcepts with widespread consequences for students’ self-evaluation, motivation, and behavioral choices. dct is based on the internal/external frame of reference model (i/e model) which integrates dimensional and external, social comparisons (e.g., “how good am i in math compared to my classmates?”). this article presents an extension, the generalized i/e model, which describes effects of dimensional and social comparisons in various areas. firstly, it proposes that such comparisons are carried out not only within the academic area but also within other areas. secondly, it proposes effects of social and dimensional comparison for other variables besides self-concepts, i.e. for motivational constructs, learning behaviors, or personality characteristics. the present article closes with an examination and discussion of the contributions of the dct by applying standards of good theories to it. keywords: dimensional comparison; social comparison; self-concept; domain-specificity   möller  et  al   f     | f l r     2   1. introduction  to the dimensional comparison theory and the i/e model this paper deals with a recently developed theory in the field of educational psychology, the dimensional comparison theory (dct; möller & marsh, 2013). first we will present the central ideas of the dct and the empirical support for its assumptions with regard to the antecedents of dimensional comparisons and the psychological processes carried out while dimensionally comparing aspects of different domains. then we will present a recent extension derived from the dct, the generalized internal/external frame of reference model (gi/e model). whereas the internal/external frame of reference model (i/e model; marsh, 1986) deals with the relations between math and verbal achievements and self-concepts and proposes positive effects from math and verbal achievements to corresponding self-concepts and negative effects on non-corresponding selfconcepts, its generalization allows the application of the relations and effects described therein to other domains as well. the dct and the gi/e model both are presented with reference to their motivational implications. in the discussion, dct’s gains with regard to the criteria of good theories developed by van lange (2013) will be summarized. 1.1 the dimensional comparison theory like social comparison theory (festinger, 1954) or the temporal comparison theory (albert, 1977), the dimensional comparison theory (möller & marsh, 2013) details the cognitive process of evaluating a certain target by comparing it to a certain standard. this cognitive process comprises four stages: the selection of a certain target for evaluation, the selection of a certain standard, the comparison of the target with the standard, and finally the evaluation of the target (biernat & eidelman, 2007; mussweiler, 2003). whereas social comparisons use information on others as the standard (festinger, 1954), and temporal comparisons use prior information on oneself as the standard (albert, 1977), dimensional comparisons use information on other attributes of the same person as a standard (möller & köller, 2001a, b). more precisely, dimensional comparisons (like temporal comparisons) are intra-individual comparisons, for example comparing one’s own mathematical achievement with one’s own verbal achievement affecting both mathematical and verbal selfevaluations. in the dct, dimensional comparisons are defined as taking place when people compare their achievement in one domain (the target domain) with their achievement in another domain (the standard domain). most of the research on dct is quantitative in nature and data come from field studies. however, there are some experimental studies clearly demonstrating effects of dimensional comparisons on self-concepts. in möller and köller (2001a) as well as in pohlmann and möller (2009), participants received dimensional comparison information indicating that their performance on the domain a is ranked worse (better) than their performance on the domain b. participants felt better about their performance in their better-off domain and worse in their worse-off domain, controlled for the presence of social comparison information (see also strickhouser & zell, 2015). research on dimensional comparison has also shown that these comparisons happen in everyday life situations. for example, in a diary study möller and husemann (2006) examined spontaneous dimensional comparisons using qualitative data. their participants were told to note and describe any dimensional comparison that came to their mind during a period of 14 days. university students (study 1) and high school students (study 2) recorded an average of more than six dimensional comparisons during the two weeks, clearly supporting the assumption that dimensional comparisons occur in everyday life. students were asked to mark which domain served as target and which domain served as the comparison standard. results showed that academic matters were most commonly used as target domains (i.e., “we were given our school reports and i compared my grade in religion with my grade in mathematics”). personal relationships with friends, partners, and family as well as the general well-being, physical appearance, and personality characteristics were also used   möller  et  al   f     | f l r     3   as targets (“although i am not that thin, i am not touchy”), yet less frequently. additionally, target and standard often belonged to the same domain. it was also shown that people often carry out dimensional comparisons when they are motivated to enhance themselves or to improve their mood. particularly in situations of failure, upward dimensional comparisons with a better-off standard (möller & husemann, 2006) serve compensational needs: when i fail at math, it is more pleasurable to concentrate on my verbal abilities. if self-enhancement is the major motivation for dimensional comparison, it is beneficial to use a better-off domain as a comparison standard. such an upward comparison often leads to a higher self-esteem and a more positive mood state in the better-off domain (despite some costs in self-concepts in the worse-off domain). the gains in the better-off domain following downward dimensional comparison (from this perspective) should be stronger than the losses in the worse-off domain following upward dimensional comparison (from this perspective): the net effect of dimensional comparison on self-evaluations empirically seems to be slightly positive (pohlmann & möller, 2009). in the diary study by möller and husemann (2006), upward dimensional comparisons were more frequent than downward comparisons. in study 1, the majority of dimensional comparisons were upward (70.5%). participants reported 20.9 % downward comparisons and 8.7% horizontal comparisons. in study 2, 52.3% of all comparisons were upward, 34.9% downward, and 12.8% horizontal. however, the need for selfenhancement is not the only motivation for dimensional comparisons. following möller, helm, müller-kalthoff, nagy, & marsh (2015), motivations for domain-specific self-evaluation and for self-improvement may also lead to dimensional comparisons. for example, someone trying to self-evaluate how verbally talented he or she is may choose his/her math ability as a comparison standard even if choosing math as a (worse-off) standard may not be beneficial to the actual self. dickhäuser, reuter, and hilling (2005) and nagy et al. (2006) showed that the probability of choosing a particular course in school is positively affected by high achievement in corresponding subjects and negatively affected by high achievement in non-corresponding subjects (i.e., influenced by dimensional comparisons). imagine a student who has to decide whether he/she wants to concentrate on language or on science courses. one criterion for his/her decision will be his/her achievement in these domains; he/she may ask him/herself: “am i better at science or in language arts?”. then again, selfimprovement motivation may also trigger dimensional comparisons: to become better in math, a student might analyze his/her motivation and learning behavior in his/her better-off verbal subjects and transfer them to math. here, the comparison might lead to a behavioral assimilation. one might think that verbal and math achievements both are based on motivation, intelligence, and adequate learning behavior so that the more positive achievement in verbal subjects could be transferred to math. this would lead to a more positive learning behavior in math as well. in addition, some hints on the antecedents of dimensional comparisons could be found: dimensional comparisons (like social comparisons) were shown to be triggered by motivational needs and/or by external forces (möller et al., 2015). the dct is inspired by the i/e model (marsh, 1986, figure 1), which posits the joint operation of both social comparisons and dimensional comparisons to construct domain-specific academic self-concepts. students conduct social comparisons by comparing their achievement with the achievement of their classmates (external frame of reference). for example, if a student’s verbal achievement is lower than that of his/her classmates, likely his/her verbal self-concept will also be lower. in addition, students conduct dimensional comparisons by comparing their achievement in a given subject with their achievements in another subject (internal frame of reference). for example, if a student’s verbal achievement is lower than his/her math achievement, his/her verbal self-concept will suffer and his/her math self-concept will benefit from dimensional comparisons. möller, pohlmann, köller, and marsh (2009) meta-analyzed 69 studies with n = 125,308 students on the relations between academic achievements and self-concepts (see figure 1). the average correlation between math and verbal achievements was strongly positive (r = .67), and much higher than the average correlation between math and verbal self-concepts (r = .10), indicating a strong domain-specificity of academic self   möller  et  al   f     | f l r     4   concepts. moreover, the effects of external comparisons, i.e. the effects of math achievement on math selfconcept (β = .61) and of verbal achievement on verbal self-concept (β =.49), were substantial and positive (see the horizontal paths in figure 1). however, the effects of dimensional comparisons, i.e. the effects from verbal achievement to mathematical self-concept (β = −.27) and of mathematics achievement on verbal self-concept (β = −.21), were negative (see the cross-dimensional paths in figure 1). integrating the results leads to the central assumption of the i/e model: the strong positive correlation between subjects-specific achievements does not lead to strong positive correlations between subject-specific self-concepts. the reason for this is the negative effect of dimensional comparisons between subject-specific achievements. the results of the meta-analysis indicate the effects of social and dimensional comparisons described in the classic i/e model to be valid for different achievement measures (grades as well as standardized achievement test scores), for different grades, gender groups, and countries (möller et al., 2009). despite the various studies supporting the i/e model, the actual psychological processes behind dimensional comparisons remain rather unexplored. a central assumption of the dct is that dimensional comparison effects are moderated by the perceived similarity of the compared school subjects. according to the dct, different school subjects form a similarity continuum (marsh, byrne, & shavelson, 1988; marsh, lüdtke et al., 2015) that explains the different outcomes of dimensional comparisons. the dct predicts that for dissimilar subjects (so-called far comparisons) like math and english, dimensional comparisons lead to contrast effects, whereas smaller contrast effects or even assimilation effects result between similar subjects (so-called near comparisons). for example, möller, streblow, pohlmann, and köller (2006) found positive path coefficients from achievements to non-corresponding self-concepts between relatively similar subjects like english and german or math and physics. möller, streblow, and pohlmann (2006) asked students directly for their belief in a negative interdependence of math and verbal abilities, that is, whether they thought of math and verbal abilities as negatively correlated or not. stronger beliefs in a negative interdependence of math and verbal ability were accompanied by more negative path coefficients from grades in one subject to academic selfconcepts in the other subject. if students considered abilities in two subjects to be positively correlated, the impact of dimensional comparisons even showed a positive assimilation effect. therefore, similarity perceptions regarding different school subjects seem to be composed to a great amount of interdependence beliefs students hold about underlying abilities. we propose that the similarity of school subjects influence dimensional comparisons in a manner that is described for social comparisons by the selective accessibility model (sam) designed by mussweiler (2003). according to sam, the comparison of a certain target to a given standard is influenced by perceptions of the general similarity of the target and the standard. we assume that when two subjects like math and english are selected as target and standard for dimensional comparison, the comparison process is driven by dissimilarity assumptions. when two more similar subjects like math and physics are dimensionally compared, the comparison process is driven by similarity assumptions instead. the similarity assumption will make commonalities between math and physics more accessible, which will result in lower contrast or even assimilation effects in self-concepts. the dissimilar perception will make differences between subjects more accessible and lead to contrast effects as described in the original i/e model.   möller  et  al   f     | f l r     5   figure 1. the i/e model: results of a meta-analytic path-analysis on the relations between math and verbal achievement and math and verbal self-concept (from möller et al., 2009). mach = math achievement; mself = math self-concept; vach = verbal achievement; vself = verbal self-concept. 1.2 the generalized internal/external frame of reference model (gi/e model) whereas the i/e model is originally restricted to math and verbal achievements and math and verbal selfconcepts, its logic is extended in the dct to a variety of other variables. therefore, we introduce a generalized i/e model (see figure 2) which may serve as a kind of a guide to look for more i/e like relations between independent and dependent variables. in this model, a person carries out social and dimensional comparisons. for example a student compares his/her own standing or perception in a certain domain with someone else’s standing, and as a result the student is able to form an opinion on his/her own standing in that particular domain. the student also carries out a dimensional comparison when comparing his/her perception of aspects of a particular domain a with his/her perception of aspects of a particular domain b coming to a conclusion about his/her standing in domain a in comparison to his/her standing domain b. both comparisons may have consequences for any kind of domain-specific thought and learning behavior. if a student perceives him/herself as being better than most of his/her classmates in a domain, the effects of social comparisons are positive for self-evaluations. however, the effects of dimensional comparisons are often negative for a particular domain. if someone perceives that he/she is better in sports than in music he/she might neglect his/her musical activities because he/she prefers sports. in extension to the original i/e model, the gi/e model allows an integration of each domain-specific aspect that students tend to compare externally and internally as a predictor or an independent variable. it also integrates the consequences of self-evaluation, motivation, and learning behavior as criteria or dependent variables. at the moment our assumptions on the generalizability of the i/e model and effects to different constructs are rather hypothetical. however there are already some empirical studies that provide preliminary evidence of the validity of our extensions. we will present the description of these studies with regard to the question whether they extended the i/e model on the side of the predictors (i.e., examining the effects of different independent variables on the academic self-concept), on the side of the criteria (i.e., examining the effects of academic achievements on different dependent variables) or on both (i.e., examining i/e-like effects in completely different domains).   möller  et  al   f     | f l r     6   changing predictors. most recent extensions on the side of the predictors integrate more or other school subjects than merely the native language and math (e.g. chiu, 2012; marsh, lüdtke et al., 2015; jansen, schroeders, lüdtke & marsh, 2015; möller, streblow, pohlmann, & köller, 2006; nagy, trautwein, baumert, köller, & garrett, 2006), i.e., multiple academic subjects (native and foreign language, history, biology, physics, and math). as already outlined, the application of the i/e model to two similar subjects does not typically lead to contrast effects in subject-specific self-concepts. it rather leads to no significant effect from subject-specific achievement to self-concept in the other subject or even to assimilative effects, i.e., positive effects from achievement to self-concept in the other subject. a study by marsh, lüdtke et al. (2015) offers an illustration of the distinction of between-domain comparisons and within-domain comparisons, showing significant contrast effects for so-called far comparisons (between dissimilar subjects) and significantly less contrast or even assimilation effects for so-called near comparisons (for similar subjects). jansen et al. (2015) analyzed dimensional comparison effects for five domains and found support for the hypotheses which derived from the dct. both contrast and assimilation effects can result from dimensional comparisons: mathematics, physics, and chemistry showed contrast effects to german self-concept, whereas more assimilative effects were found from achievements in the three subjects to mathematics, physics, and chemistry self-concepts. furthermore, tietjens, möller, and pohlmann (2005) successfully replicated the i/e model using sports achievement as predictors. performance in track and field negatively affected the self-concept in swimming and basketball and swimming performance negatively influenced the self-concept in soccer (see also chanal, sarrazin, guay, & boiché, 2009). in the diary study mentioned above (möller & husemann, 2006), participants were found to compare a vast variety of different domains intra-individually, like personality characteristics and physical attractiveness. figure 2. the generalized i/e model. an extension of the i/e model to other domains and consequences. changing criteria. some studies in the tradition of the i/e model used math and verbal achievements as comparison target and standard, but then used different dependent variables analyzing effects on variables other   möller  et  al   f     | f l r     7   than self-concepts, i.e. self-regulated learning (miller, 2000), emotions (goetz, frenzel, hall, & pekrun, 2008), intrinsic motivation (marsh, abduljabbar, parker, morin, abdelfattah, nagengast, möller, & abu-hilal, 2015), and interest (schurtz, pfost, nagengast, & artelt, 2014). with regard to motivation we refer to the corpus of motivation terms relevant to academic achievement murphy and alexander (2000) discussed. they differentiated between self-schema (including self-efficacy and attribution), interest (situational and individual), intrinsic and extrinsic motivation, and goal orientation (including learning, performance, and work avoidance goals). our general answer to the question which constructs fit into i/e relations as criteria is based on the domain-specificity of the motivational constructs: whereas self-schema, interest, and intrinsic/extrinsic motivation in most research studies are domain-specific constructs, goal orientation is often conceptualized as more domain-general (bong, 2013). the more a motivational construct is specific to a domain, the more we will assume effects of dimensional comparisons. motivational constructs that are less specific to a domain are rarely able to be the subject of dimensional comparisons. there is one important exception to this rule: in the möller et al. (2009) meta-analysis, the most important moderator was the type of self-schema measure. when measures of self-efficacy served as selfconcept indicators that included the target in the item, analyses revealed larger relations between math and verbal self-evaluations and the i/e model did not fit the data well. the most important difference between such self-efficacy beliefs and self-concept with regard to comparison processes may be that self-efficacy beliefs, when including the target in the item beliefs, are much more driven by former experiences with similar types of tasks. so far, there is a lack of studies that analyze the effects of dimensional comparisons on other criteria than students’ self-reported self-concepts. although there is already evidence for i/e like relations between students’ verbal and math grades and other-ratings of students’ ability beliefs (e.g., dickhäuser, 2005; möller, 2005), it would be interesting to analyze the effects of social and dimensional comparisons using a broader variety of domain-specific criteria (observed and self-reported) such as time spent on homework, teacher-ratings of students’ classroom behavior, or students’ perceptions of instructional quality of classes (see arens & möller, 2016). in one of the first studies directly referring to the gi/e model, arens and möller (2016) asked students for their grades in math and german and for their perceptions of the learning environment in each their math and german classes from two perspectives: first, the students were asked to rate their relationships to their math teacher and their german teacher. second, the students were asked to judge the perceived quality of the instruction they received in math and german classes. analyses revealed positive paths from grades to corresponding student-teacher relationships and qualities of instruction. more importantly for the dct, negative paths occurred from grades to non-corresponding perceptions of the quality of instruction and positive paths occurred from grades to non-corresponding student-teacher relationships, indicating dimensional comparison processes. changing predictors and criteria. very few studies replaced both independent and dependent variables. however dietrich, dicke, kracke, & noack (2015) found dimensional comparison effects when analyzing crossdomain relations of teacher support and motivation: higher levels of perceived teacher support in one subject were negatively related to students’ intrinsic value and effort in another subject. möller and savyon (2003) analyzed dimensional comparison effects between intelligence and honesty. they gave their participants success or failure feedback on anagram tasks. people in the failure condition rated themselves as being more honest than did students who received positive feedback in the anagram task, indicating that processes of dimensional comparison were carried out between intelligence and honesty. möller and marsh (2013) suggested a transfer of the i/e model to basic personality characteristics. they viewed the “big two” personality dimensions agency (competence) and communion (warmth; abele &   möller  et  al   f     | f l r     8   wojciszke, 2014) as ideal candidates for an extension of the dct since both dimensions are independent of each other in self-perception, as are math and verbal self-concepts. a first re-analysis of the data of abele, rohe, and hauke (2013, merged from studies i and ii) revealed some support for a “big two i/e model”. helm et al. (under review) revealed typical i/e patterns between other-rated agency and communion as predictors and selfrated agency and communion as criteria, i.e. positive effects on corresponding self-ratings and negative effects on non-corresponding self-ratings. such extensions to personality variables like agency and communion widen the opportunities delivered by the generalized i/e model. to sum up, the generalized i/e model may serve as a matrix for other juxtapositions of domaincharacteristics with consequences for domain-specific beliefs and learning behaviors. the successful transfer of our assumptions to agency and communion may serve as an example for the research possibilities that arise from these extensions, within and outside of learning and motivation research. 1.3 discussion the aim of the present article was to give an overview on a new theory in motivation research, the dimensional comparison theory, as well as to devise a critical extension to the core assumption of the theory. namely, we introduced a generalized i/e model assuming external and in particular internal, dimensional comparisons to be rather general comparison processes not only limited to the formation of academic selfconcepts alone, but to apply to evaluations of different constructs as well. the gi/e model may enable future research to go beyond the relations between verbal and math self-concepts and apply the rationale underlying external and dimensional comparisons in different fields and disciplines as well. in conclusion, we would like to evaluate the dct in regard to its usefulness in terms of a good theory. we will try to evaluate dct from our (subjective) perspective. according to van lange (2013) truth, abstraction, progress, and applicability as standards (tapas) may serve as ideals for theories in psychology. in the following section, we would like to apply tapas to dct: truth. the ideal of truth is met when a theory allows hypothesizing testable relations between the constructs that the theory deals with. although according to popper (1959) truth remains an unreachable ideal, empirical studies allow researchers, who are testing hypotheses derived from theories, to evaluate what aspects of a theory are more or less adequate descriptions of the data. with regard to motivation and learning research, such a theory has to describe or explain data on motivational constructs, e.g. relations between two motivational constructs or between motivation and achievement. in the case of the dct, we have shown that there is strong evidence (a) for the occurrence of dimensional comparisons inside and outside of schools and (b) for crossdomain effects between achievements and academic self-concepts. the empirical support is smaller with regard to other motivational constructs. abstraction. the second ideal of abstraction asks theories to go beyond single empirical studies, generalize findings, and verbalize relations between constructs. in our opinion, the dct is abstract enough while still exposing the causal relations between aspects of domains and corresponding self-evaluations as well as noncorresponding self-evaluations of these aspects, and grounding them on psychological principles. the dct overcomes the limitations of the i/e model, which is a more descriptive approach.     progress. thirdly, ideal theories strive for progress. they should add some new insights to the prior knowledge or provide a new and intriguing perspective on a phenomenon. the dct emphasizes a comparison process rather neglected in research outside of educational psychological self-concept research. a first visible sign of progress may be that research on the dct is published in a major social psychology journal (strickhouser & zell, 2015). from our perspective and concerning the dct, it is essential to generalize the i/e model and to initiate more (and more experimental) research on the psychological phenomena associated with   möller  et  al   f     | f l r     9   dimensional comparisons. derived from dct, the gi/e model allows the application of the relations and effects described in the i/e model to other domains and person characteristics as well. this would offer a new theoretical perspective on the formation of self-concepts and on the emergence of motivational tendencies in different areas of life. applicability. fourthly, a theory should be applicable to real-world concerns, an ideal that is in particular important in an applied field of research like motivation and learning research. the dct does not only claim to be applicable to the well-known and practically important relations between academic achievement and academic self-concept, but has the potential of explaining several real world phenomena that may benefit from considering dimensional comparisons as an underlying psychological process. the dct is applicable in each situation, which forces people to choose between alternatives. in particular, when students have to choose between courses, select careers and academic majors, dimensional comparisons will be at work. to sum up, the gi/e model was introduced overcoming the i/e model’s limitations with regard to math and verbal affairs. first empirical support was presented underlining the fruitfulness of the model to initiate further research activities. the gi/e model is a consequence of the formulation of the dct and enriches the theories’ purpose. further research is needed that adds new domains as comparison targets and standards as well as new motivational and learning outcomes as consequences of external and internal comparisons. keypoints the internal/external frame of reference model (i/e model) describes the relations between math and verbal achievements and self-concepts. the dimensional comparison theory (dct) focuses on the internal frame of reference and extends the i/e model. instead of being limited to math and verbal achievements, it deals with other domains as well. instead of being limited to math and verbal self-concepts, it deals with motivational constructs, learning behaviors, and personality characteristics. therefore, a generalized i/e model is proposed that allows the application of previous findings and will initiate further research. dct is discussed by applying standards of good theories to it. references abele, a., rohe, m., & hauke, n. (2013). does my friend see me like i do? friendship quality and self-other agreement on the fundamental dimensions of social judgment. unpublished manuscript, university of erlangen-nuremburg, germany. abele, a. e., & wojciszke, b. (2014). communal and agentic content. a dual perspective model. advances in experimental social psychology, 50, 198–255. arens, k. & möller, j. (2016). dimensional comparisons: effects on perceived learning environments. learning and instruction, 42, 22-30. doi: 10.1016/j.learninstruc.2015.11.001 albert, s. (1977). temporal comparison theory. psychological review, 84, 485–503. doi: http://dx.doi.org/10.1037/0033-295x.84.6.485   möller  et  al   f     | f l r     10   biernat, m., & eidelman, s. (2007). standards. in a. w. kruglanski and e. t. higgins (eds.), social psychology: handbook of basic principles, volume 2 (pp. 308–333). new york: the guilford press. bong, m. (2013). self-efficacy. in j. hattie, e. m. anderman (eds.), international guide to student achievement (pp. 64-66). new york, ny, us: routledge/taylor & francis group. chanal, j. p., sarrazin, p. g., guay, f., & boiché, j. (2009). verbal, mathematics, and physical education selfconcepts and achievements: an extension and test of the internal/external frame of reference model. psychology of sport and exercise, 10, 61–66. doi: http://dx.doi.org/10.1016/j.psychsport.2008.06.008 chiu, m.-s. (2012). the internal/external frame of reference model, big-fish-little-pond effect, and combined model for mathematics and science. journal of educational psychology, 104, 87–107. doi: http://dx.doi.org/10.1037/a0025734 dickhäuser, o. (2005). teachers' inferences about students' self-concepts – the role of dimensional comparison. learning and instruction, 15(3), 225-235. doi: http://dx.doi.org/10.1016/j.learninstruc.2005.04.004 dickhäuser, o., reuter, m., & hilling, c. (2005). coursework selection: a frame of reference approach using structural equation modelling. british journal of educational psychology, 75, 673–688. doi: http://dx.doi.org/10.1348/000709905x37181 dietrich, j., dicke, a. l., kracke, b., & noack, p. (2015). teacher support and its influence on students' intrinsic value and effort: dimensional comparison effects across subjects. learning and instruction, online first. doi: http://dx.doi.org/10.1016/j.learninstruc.2015.05.007 festinger, l. (1954). a theory of social comparison processes. human relations, 7, 117–140. doi: http://dx.doi.org/10.1177/001872675400700202 goetz, t., frenzel, c. a., hall, n. c., & pekrun, r. (2008). antecedents of academic emotions: testing the internal/external frame of reference model for academic enjoyment. contemporary educational psychology, 33, 9–33. doi: http://dx.doi.org/10.1016/j.cedpsych.2006.12.002 helm, f., müller-kalthoff, h., & nagy, n., & möller, j. (under review). dimensional comparison effects on academic self-concepts: similarity of school subjects. university of kiel, germany. jansen, m., schroeders, u., lüdtke, o., & marsh, h. w. (online first, 2015). contrast and assimilation effects of dimensional comparisons in five subjects: an extension of the i/e model. journal of educational psychology. doi: http://dx.doi.org/10.1037/edu0000021 marsh, h. w. (1986). verbal and math self-concepts: an internal/external frame of reference model. american educational research journal, 23, 129–149. doi: http://dx.doi.org/10.3102/00028312023001129 marsh, h. w., abduljabbar, a. s., parker, p. d., morin, a. s., abdelfattah, f., nagengast, b., möller, j., & abuhilal, m. m. (2015). the internal/external frame of reference model of self-concept and achievement relations: age-cohort and cross-cultural differences. american educational research journal, 52(1), 168202. doi: http://dx.doi.org/10.3102/0002831214549453 marsh, h. w., byrne, b. m., & shavelson, r. j. (1988). a multifaceted academic self-concept: its hierarchical structure and its relation to academic achievement. journal of educational psychology, 80, 366–380. doi: http://dx.doi.org/10.1037/0022-0663.80.3.366 marsh, h. w., lüdtke, o., nagengast, b., trautwein, u., abduljabbar, a. s., abdelfattah, f., & jansen, m. (2015). dimensional comparison theory: paradoxical relations between self-beliefs and achievements in multiple domains. learning and instruction, 35, 16-32. doi: http://dx.doi.org/10.1016/j.learninstruc.2014.08.005 miller, j. w. (2000). exploring the source of self-regulated learning: the influence of internal and external comparisons. journal of instructional psychology, 26, 47–52. möller, j. (2005). paradoxical effects of praise and criticism: social, dimensional and temporal comparisons. british journal of educational psychology, 75(2), 275-295. doi: http://dx.doi.org/10.1348/000709904x24744   möller  et  al   f     | f l r     11   möller, j., helm, f., müller-kalthoff, h., nagy, n., & marsh, h. w. (2015). dimensional comparisons: theoretical assumptions and empirical results. in: j. wright (ed.), international encyclopedia of the social and behavioral sciences, 2nd ed. (pp. 430-436). elsevier: oxford, gb. möller, j., & husemann, n. (2006). internal comparisons in everyday life. journal of educational psychology, 98, 342–353. doi: http://dx.doi.org/10.1037/0022-0663.98.2.342 möller, j., & köller, o. (2001a). dimensional comparisons: an experimental approach to the internal/external frame of reference model. journal of educational psychology, 93, 826–835. doi: http://dx.doi.org/10.1037/0022-0663.93.4.826 möller, j., & köller, o. (2001b). frame of reference effects following the announcement of exam results. contemporary educational psychology, 26, 277-287. doi: http://dx.doi.org/10.1006/ceps.2000.1055 möller, j. & marsh, h. w. (2013). dimensional comparison theory. psychological review, 120, 544-560. doi: http://dx.doi.org/10.1037/a0032459 möller, j., pohlmann, b., köller, o., & marsh, h. w. (2009). a meta-analytic path analysis of the internal/external frame of reference model of academic achievement and academic self-concept. review of educational research, 79, 1129–1167. doi: http://dx.doi.org/10.3102/0034654309337522 möller, j., & savyon, k. (2003). not very smart thus moral: dimensional comparisons between academic selfconcept and honesty. social psychology of education, 6, 95–106. doi: http://dx.doi.org/10.1023/a:1023247910033 möller, j., streblow, l., & pohlmann, b. (2006). the belief in a negative interdependence of math and verbal abilities as determinant of academic self-concepts. british journal of educational psychology, 76, 57–70. doi: http://dx.doi.org/10.1348/000709905x37451 möller, j., streblow, l., pohlmann, b., & köller, o. (2006). an extension to the internal/external frame of reference model to two verbal and numerical domains. european journal of psychology of education, 21, 467–487. doi: http://dx.doi.org/10.1007/bf03173515 murphy, p. k., & alexander, p. a. (2000). a motivated exploration of motivation terminology. contemporary educational psychology, 25(1), 3-53. doi: http://dx.doi.org/10.1006/ceps.1999.1019 mussweiler, t. (2003). comparison processes in social judgment: mechanisms and consequences. psychological review, 110, 472–489. doi: http://dx.doi.org/10.1037/0033-295x.110.3.472 nagy, g., trautwein, u., baumert, j., köller, o., & garrett, j. (2006). gender and course selection in upper secondary education: effects of academic self-concept and intrinsic value. educational research and evaluation, 12, 323–345. doi: http://dx.doi.org/10.1080/13803610600765687 pohlmann, b., & möller, j. (2009). on the benefits of dimensional comparisons. journal of educational psychology, 101, 248–258. doi: http://dx.doi.org/10.1037/a0013151 popper, k. r. (1959). the logic of scientific discovery. new york, ny: harper. (original work published as logik der forschung, 1935) schurtz, i. m., pfost, m., nagengast, b., & artelt, c. (2014). impact of social and dimensional comparisons on student's mathematical and english subject-interest at the beginning of secondary school. learning and instruction, 34, 32-41. doi:http://dx.doi.org/10.1016/j.learninstruc.2014.08.001 strickhouser, j. e. & zell, e. (2015, online first). self-evaluative effects of dimensional and social comparison, journal of experimental social psychology. doi: http://dx.doi.org/10.1016/j.jesp.2015.03.001 tietjens, m., möller, j., & pohlmann, b. (2005). leistung und selbstkonzept in verschiedenen sportarten [achievement and self-concept in various sports]. zeitschrift für sportpsychologie, 12, 135–143. doi: http://dx.doi.org/10.1026/1612-5010.12.4.135 van lange, p. m. (2013). what we should expect from theories in social psychology: truth, abstraction, progress, and applicability as standards (tapas). personality and social psychology review, 17(1), 40-55. doi: http://dx.doi.org/10.1177/1088868312453088 salmeron et al frontline learning research vol.6 no. 3 (2018) 105 122 issn 2295-3159 using eye-tracking to assess sourcing during multiple document reading: a critical analysis ladislao salmeróna, laura gilaivar bråten b adepartment of developmental and educational psychology, university of valencia, spain bdepartment of education, university of oslo, norway article received 14 may 2018/ revised 17 october/ accepted 28 november/ available online 7 december abstract during the last 15 years, there have been some efforts to extend the use of eye-tracking to researching reading in complex contexts, such as the reading of multiple documents. the research community involved in this extension has been interested in higher-order comprehension processes occurring in complex reading contexts, such as sourcing, defined as the processes of attending to, representing, evaluating, and using available or accessible information about the sources of textual content. in this article, we argue that extending eye-tracking research to investigate more complex reading contexts has been made without critically reflecting on its validity in those contexts. specifically, because eye-tracking captures automatic as well as conscious processes, it is currently an open question how consistently eye-tracking captures the strategic sourcing processes that take place during multiple document reading, in particular when using real documents that include salient source information that may attract bottom-up fixations. in contrast, subjective methods, such as interviews, mainly target conscious processes, and may therefore be a more valid and generalizable measure of strategic sourcing activities. we compared sourcing indicators based on eye-tracking measures to sourcing indicated by a post-reading interview. results suggested that current eye-tracking indices of sourcing are not universally valid measures, and that simpler methods, such as asking readers whether they paid attention to source information, may be more suited to assess strategic sourcing during multiple document reading. keywords: eye-tracking; interview; reading comprehension; multiple documents; sourcing. info mail corresponding author ladislao.salmeron@uv.es doi: https://doi.org/10.14786/flr.v6i3.368 1. introduction during the last four decades, reading researchers have extensively used eye-tracking measures to understand reading comprehension processes as they unfold during actual reading (for reviews, see leinenger & rayner, 2017; rayner, 1998, 2009). this approach is based on the mind-eye hypothesis (just & carpenter, 1980), which assumes that readers fixate words as long as they are being processed. therefore, eye-movements may be used as an indicator of what and how readers process when interacting with text. a major advantage of eye-tracking compared to subjective measures, such as verbal reports or interviews, is that this is an objective measure that captures subtle and automatic processes during reading, processes that may not be conscious to the reader. our current understanding of reading processes owes much to the use of eye-tracking methodology, but the use of eye-tracking also comes with certain limitations. to be able to map specific eye-movements (e.g., regressions) onto specific reading processes (e.g., interpreting anaphoric references), researchers have developed highly controlled experimental stimuli. this approach has resulted in a research corpus on reading comprehension with an overrepresentation of simple text materials, such as sentences, paragraphs, or, at best, brief narratives (jarodzka & brand‐gruwel, 2017). however, during the last 15 years, there have also been some efforts to extend the use of eye-tracking to researching reading in more complex contexts, such as the reading of long expository texts (e.g., catrysse, gijbels, donche, de maeyer, lesterhuis, & van den bossche, 2018; hyönä, lorch jr., & kaakinen, 2002; kaakinen, hyönä, & keenan, 2002), digital hypertexts (e.g., salmerón, baccino, cañas, madrid, & fajardo, 2009; salmerón, naumann, garcía, & fajardo, 2017), or literary pieces (e.g., jacobs, 2016). the research community involved in this extension has been interested in higher-order comprehension processes occurring in complex reading contexts, such as the selection of information, identification of text relevance, and integration of textual information. in this article, we will argue that extending eye-tracking research to investigate more complex reading contexts has been made without critically reflecting on the validity of this approach when used in those contexts. specifically, we will focus on the recent use of eye-tracking to study sourcing during the reading of multiple documents. finally, we will present an empirical study to test the applicability of current eye-tracking indices to detect sourcing. in the context of reading multiple documents, sourcing refers to processes of attending to, representing, evaluating, and using available or accessible information about the sources of textual content, for example about the author or publisher of different texts (bråten & braasch, 2018). when people read multiple documents about complex, controversial issues of which they have limited prior knowledge, such as climate change, they will have great difficulties determining the accuracy or truth value of information directly through “first-hand evaluation” (stadtler & bromme, 2014). in such complex reading contexts, they may therefore profitably resort to the strategy of evaluating information indirectly in light of the features of the sources (termed “second-hand evaluation” by stadtler and bromme). this is consistent with the discrepancy-induced source comprehension (d-isc) model (braasch & bråten, 2017; braasch, rouet, vibert, & britt, 2012), which describes processes occurring when readers try to understand conflicting information presented by different sources. according to d-isc, readers may find it difficult, if not impossible, to construct a coherent mental representation of an issue if conflicting information is presented across multiple sources. in such situations, readers will therefore strategically shift cognitive resources towards constructing a mental representation of the issue that includes source information (e.g., about the document authors) as organizational elements (braasch & bråten, 2017). in other words, conflicting information across multiple sources becomes an impetus for strategic sourcing processes and such processes, in turn, allow readers to construct a meaningful interpretation of the issue despite existing conflicts. because eye-tracking presumably captures automatic as well as conscious processes (mccormick, 1997), it is currently an open question how accurately and consistently eye-tracking captures the strategic sourcing processes described above, in particular compared to subjective methods that mainly target conscious processes, such as interviews. in their recent review of the threats to validity in eye-movements research, orquin and holmqvist (2018) identified the distinction between bottom-up and top-down fixations as a major potential problem. specifically, a stimulus can be fixated due to top-down processes, which indicates that readers strategically allocate their attention with the purpose of extracting information. for example, when reading a sentence, readers focus their attention sequentially on words to extract meaning. but a stimulus can also be fixated due to bottom-up processes, for example, because it is visually salient. such fixations do not represent strategic actions, and are not implemented with the goal of extracting meaning from the stimulus. bottom-up fixations may not be common while reading simple text materials typically used in laboratory settings, such as a paragraph written in uniform font size. this is because other information from the documents is eliminated to avoid distractions in such settings. accordingly, the mind-eye hypothesis (just & carpenter, 1980) can accurately describe the knowledge extraction processes associated with reading in such settings, because most fixations are driven by top-down processes. however, bottom-up fixations may be frequent, and thus pose a threat to validity, when the texts are presented as real documents, for example as a real newspaper. in such a case, a fixation on the nameplate on the front-page of the newspaper (e.g. “the new york times”) may be triggered by readers’ attempts to identify and evaluate the document source (i.e., top-down), or simply occur because the nameplate was a salient feature of the document (i.e., bottom-up). therefore, interpretations of such fixations as purely strategic processes will be highly problematic. during the last decade, several authors have attempted to use eye-tracking to measure sourcing in multiple document reading situations (braasch et al., 2012; brand‐gruwel, kammerer, van meeuwen, & van gog, 2017; gerjets, kammerer, & werner, 2011; kammerer, kalbfell, & gerjets, 2016; mason, pluchino, & ariasi, 2014; van strien, kammerer, brand-gruwel, & boshuizen, 2016). these researchers have used different indices derived from fixations on source information available on pages, typically document logos (e.g., the name of a company), ‘about us’ information, or names of authors. continuous indices include first and second pass fixation duration on source names in headlines or logos (mason et al., 2014) and total fixation time on logos (gerjets et al., 2011; van strien et al., 2016). non-continuous indices include percentage of page logos fixated for more than 100 msec (brand‐gruwel et al., 2017). one set of studies have used eye-tracking measures to test the hypothesis that readers tend to consider sources in trying to resolve inconsistencies across documents, in accordance with the d-isc model (braasch & bråten, 2017; braasch et al., 2012). for example, authors’ affiliations could be used to interpret the extent to which they are knowledgeable about the issue and likely to provide unbiased information about it. braasch et al. (2012, exp 1) had participants read brief texts in which two authors provided either consistent or inconsistent accounts of the same event. when reading inconsistent accounts, participants were found to refixate the source area including information about author affiliation more often than when reading consistent accounts. more recently, kammerer et al. (2016) replicated this effect in a multiple document scenario. in two studies, participants read two webpages about a controversial health-related issue that provided either consistent or inconsistent information about the issue. eye-tracking data were used to measure the time spent on reading the ‘about us’ information when encountering consistent versus inconsistent information across webpages. supporting the d-isc model, participants spent more time reading the source information when reading inconsistent compared to consistent information. however, another set of studies have used eye-tracking indices of sourcing in a more exploratory (rather than experimental) manner, that is, without a clear connection to theoretical models of sourcing, such as the d-isc model (braasch & bråten, 2017). thus, none of these studies investigated sourcing during specific reading episodes, such as when encountering conflicting information, but rather tried to capture sourcing through eye-tracking from beginning to end. results reported in those studies provide examples that question the validity of eye-tracking measures to accurately assess readers’ sourcing activities. for example, participants explicitly told to assess the trustworthiness of webpages have not been found to spend more time inspecting source information from the pages (i.e. page logos) than those not told to do so (gerjets et al., 2011). moreover, a range of studies within multiple document literacy has indicated that readers with higher topic knowledge typically display more sourcing than do readers with lower topic knowledge (braasch, bråten, strømsø, & anmarkrud, 2014; bråten, strømsø, & salmerón, 2011; rouet, britt, mason, & perfetti, 1996; rouet, favart, britt, & perfetti, 1997; salmerón, kammerer, & garcía-carrión, 2013; stang lund, bråten, brandmo, brante, & strømsø, in press; strømsø, bråten, & britt, 2010). these studies have used various methodologies, such as expert-novice comparisons or think aloud protocols, to capture the effects of prior topic knowledge on sourcing measured offline (e.g., by means of argumentative essays of source memory). however, attempts to replicate this effect using eye-tracking indicators of sourcing have been unsuccessful. for example, brand-gruwel et al. (2017) reported no difference between experts and novices in the percentage of webpages on which source information was fixated, and van strien et al. (2016) found no difference between students with more or less prior topic knowledge in this regard. finally, researchers within multiple document literacy have identified a positive relationship between sourcing and the comprehension of multiple conflicting documents (anmarkrud, bråten, & strømsø, 2014; barzilai & eshet-alkalai, 2015; barzilai, tzadok, eshet-alkalai, 2015; bråten, strømsø, & britt, 2009; goldman, braasch, wiley, graesser, & brodowinska, 2012; salmerón, gil, & bråten, 2018; strømsø et al., 2010; wiley et al., 2009). to the best of our knowledge, only one study has directly tested the relationship between sourcing measured via eye-tracking and students’ performance. in that study, van strien et al. (2016) found a weak correlation between total fixation times on logos and number of ideas included in a post-reading summary, but only for documents that provided information that was in line with participants’ own attitudes towards the topic (but not for documents that contradicted participants’ own attitudes towards the topic). in sum, existing evidence suggests that current eye-tracking indices of sourcing are not universally valid measures, and that they may be especially problematic when they are not tied to experimental manipulations but, rather, used in exploratory studies. in the present study, we built on the d-isc model in assuming that sourcing represents conscious processing in the service of constructing coherent mental representations (braasch & bråten, 2017), and that it, therefore, will be more adequately measured by means of a subjective measure, such as an interview, than by means of eye-tracking. therefore, we expect that sourcing as measured by what could be considered gold standard measures (i.e., source citation in essays and memory for sources) would be predicted by participants’ reporting of sourcing activities in a post-reading interview, but not by different eye-tracking indicators of attention to source information during reading. 2. method 2.1. participants forty-three undergraduates from a public university in eastern spain participated for extra course credit. the sample included 33 females and 10 males, who ranged in age from 17 to 32 years and had an overall mean age of 19.9 (sd = 2.7). all participants were caucasians and native spanish speakers. participants were assigned randomly to conditions in which they read either real or print-out versions of multiple documents (see sections 2.2.3 and 2.3 below). 2.2. instruments and materials 2.2.1. documents participants read four authentic documents containing relevant information about the topic of climate change. the first document was a 526-word excerpt from a textbook that provided basic knowledge about the greenhouse effect and described current controversies concerning climate change. the second document was a 363-word editorial from a major spanish newspaper focused on the negative consequences of climate change. the third document was a 296-word blog entry written by an expert in climatology that discussed the potential benefits of increases in co2. the last document was a 286-word article from a popular science magazine arguing that efforts to prevent global warming, although necessary, will have great economic costs. in addition to information about the sources of the documents themselves, all the documents cited at least two other sources (i.e., contained embedded sources) but none of the documents referred to each other. more specifically, the textbook excerpt included two embedded sources, while the newspaper editorial included four, the blog entry three, and the popular science article two embedded sources. the number of embedded sources was identical in the print-out and real document conditions. a more detailed description of these four documents can be found in salmerón et al. (2018). documents were presented in two conditions that varied in their design: real documents or print-out copies of them (see below). such manipulation was originally used to test the document as entities hypothesis (britt, rouet, & braasch, 2013), which suggests that readers’ integration of information across documents is likely to increase when source characteristics are more salient. in the context of the current manuscript, this manipulation served to compare the validity of eye-movements indices when reading complex real documents as compared to visually more uniform print-out versions, as well as to test the generalizability of the results to different types of texts. in the real documents condition, participants read the textbook excerpt in a printed textbook, the newspaper editorial in a printed newspaper, the blog entry on a tablet, and the popular science article in a printed popular science magazine. thus, in this condition, the documents differed with regard to a number of physical properties, such as size, weight, shape, and color, and participants had to adapt their body posture and their hands to read each text. when presented with the textbook, the newspaper, and the popular science magazine, participants were instructed to look for post-it notes that indicated the start and the end of the text to be read. in the print-out condition, participants read print-out versions of the same documents, with each text printed in black on a separate a4 sheet of paper, using times new roman, font size 12, and line space 1.5. exactly the same source information that was available about the documents in the real document condition (i.e., author, document type, publisher, and date) was presented at the beginning of the texts, and the same embedded sources were included in the print-our versions of the documents. 2.2.2. writing task after reading the documents, all participants wrote an argument essay given the following written instruction: the texts that you just read present different perspectives on climate change. please write a report where you describe and evaluate the different perspectives, taking the arguments and evidence presented in the texts into consideration. use approximately half a page for your essay. all essays were coded to indicate students’ sourcing. in doing so, we identified all the segments that included specific references to source information from one or more of the documents as source citations. a segment was coded as a source citation when it referred to either accurate source feature information about the documents themselves, for example to document author, document type, or publication (e.g., “as described in the textbook”, “the text from the newspaper explains”, “the article published in the newspaper el país”), or to embedded sources (e.g., “at the end of the 19th century arrhenius prophesied”, “according to the ipcc”). two independent judges independently scored a random selection of 10% of the essays. the cohen’s kappa agreement was .84 for the coding of students’ source citations, which indicates an almost perfect agreement according to landis and koch (1977). all disagreements were resolved in thorough discussion between the two raters, and one of them scored the remaining essays with respect to source citations using the same coding system. 2.2.3. source memory task participants were requested to write down all the information they remembered about the documents they had just read, including the author, the document type, the publisher, and the date. for each document, we coded three source features: the author or publisher (i.e., the surname of the document’s author or the name of the document’s publisher), the document type (i.e., book or textbook/editorial, newspaper, or journalistic article/blog or webpage/magazine, science magazine, or science article), and the year of publication. the reason author and publisher were not scored separately is that the source information provided for the textbook excerpt and the newspaper editorial did not mention the author of the documents. participants received one point for each source feature that they reported correctly (maximum score = 12 points). a random sample of 10% of the source memory statements were independently scored by two raters, yielding adequate reliability (cohen’s kappa = .79). all disagreements were resolved in thorough discussion between the two raters, and one of them scored the remaining source memory statements according to the same coding system. 2.2.4. interview following completion of the source memory task, the first or second author interviewed participants about their reading of the documents. during the interview, participants were asked questions about their use of specific strategies while reading the documents and their perceptions of the potential benefits of using each strategy. for the purpose of this study, only responses to a question concerning participants’ attention to source information were analyzed. specifically, they were asked: while reading the texts, did you ever pay attention to the source of the document, such as the author or publisher ...? if needed, the interviewer provided more examples to ensure that the student had understood the question correctly. participants’ responses to this question were scored 1 when they stated that they had paid attention to source information during reading and 0 when they said they had paid no attention to sources. 2.2.5. apparatus students’ eye-movements during the reading task were assessed by an eye-tracker mounted in a pair of glasses (smi eye tracking glasses), with a sampling rate of 30 hz. the selection of this apparatus, over more robust ones, was due to the fact that the stimuli were provided in paper, and therefore the glasses were the only available eye tracking equipment to use. the smi eye tracking glasses use an event detection algorithm that identifies saccades and blinks, and classifies the remaining samples as visual intake or fixation. fixations shorter than 50 msec are excluded. a pre-processing algorithm corrects for head movements. in the print-out condition, the area of interest (aoi) corresponding to source information was a short text that described the type of document, the name of the publisher, the author’s name, and the date. for example, the popular science magazine included the following source information: article from the journal scientific american, by john matson, journalist. number 76, 2nd trimester 2014. this text was always located in the upper right part of the print-out (figure 1, left). in the real document condition, source information was located in different aois and included the logo of the publisher, the author’s name, and the date (figure 1, right). for example, source information for the popular science magazine included the name of the journal and the date on the front page, and the author’s name and occupation below the title on the text page. we computed two types of indices of sourcing based on eye-movement to the source aois: continuous and non-continuous. continuous indices included a) first pass dwell time that started the first time students entered an aoi and ended when they shifted to a different aoi (in milliseconds), b) number of fixations of a first pass dwell episode, c) second pass dwell time corresponding to all subsequent revisits to that source’s aoi, and d) number of fixations of all second pass dwell episodes. non-continuous indices were based on the number of sources (n = 4) fixated for a particular time period. as there are no clear guidelines to determine a particular threshold, we explored four different time thresholds. specifically, we computed the total number of sources fixated for at least 100, 500, and 1000 milliseconds (cf. brand‐gruwel et al., 2017; kammerer & gerjets, 2014). bromme, stadtler, and scharrer (2018) have argued that source information is not only meta-textual, as some relevant aspects, such as document genre, can be derived from the text itself. for example, science magazines tend to use neutral voice, while blogs tend to be written in first-person voice. in an attempt to isolate potential effects of text processing on source identification, we included total text reading time and total number of text fixations as covariates in our analyses. figure 1. examples of the newspaper editorial presented in the print-out (top) and real document (bottom) conditions. aois for the content are represented by solid line squares, while aois for source information are indicated by dash line squares. note that in the real document condition the source information (nameplate and date) was located on the front-page and the content on an inside page. 2.3. procedure participants were tested individually in sessions lasting approximately 90 minutes. on arrival, they were sequentially assigned to either the real document condition (n = 20) or the print-out condition (n = 23). first, they answered a questionnaire on demographics and completed prior knowledge and interest measures (not reported in this study, but see salmerón et al., 2018). after having completed these measures, they were presented with the documents and orally given the following instruction: you are now going to read four texts that present different perspectives on climate change. after reading the texts, you are going to write a report that discusses the different perspectives based on the arguments and evidence presented in the texts. there is no time limit, and you can read and reread the texts in the order you prefer. please note that you will not have the texts available while writing your report. the presentation order for the four documents was randomized for each participant in the real document condition, and each participant in the print-out condition was presented with the documents in the same random order as the preceding participant in the real document condition. in both conditions, the four documents were piled on a table in random order and all were available to participants during the entire reading task. calibration was set by using an a4 sheet of paper with three points (two in the upper corners, and one in the middle of the bottom of the page) that students held as if they were reading it. after participants had finished reading the documents, they completed the writing task, a text comprehension measure (also not reported in this study), and the source memory task in that order on a laptop using a word-processing application. finally, participants were interviewed about their strategy use when reading the documents. 3. results first, we describe the different sourcing indices that we explored by means of eye tracking and compare them across the two conditions (i.e., real vs. print-out versions of documents). next, we analyze the extent to which sourcing based on both continuous and discontinuous eye-tracking data, as well as sourcing reported in the interview, were related to students’ offline sourcing (i.e., source citation and source memory). finally, we triangulate the different sourcing indices extracted from eye-tracking and interview data. 3.1. descriptive analyses across conditions, sourcing as indicated by both continuous and non-continuous eye-tracking data took longer time and was more frequent in the print-out than in the real document condition (table 1). this pattern may simply be explained by the fact that in the print-out version, the amount of source information was greater (as the type of document was made explicit) and more salient (located together rather than distributed across several pages). in contrast, the number of participants who said in the interview that they had attended to source information during reading was similar in the two conditions. table 1 descriptive statistics for all the measures by condition (real document and print-out versions). 1 squared values were used to correct the positive skewness. 3.2. sourcing as measured by continuous eye-movement indices to test the validity of the continuous eye-movement indices of sourcing, we performed partial correlations comparing those values with offline sourcing as indicated by source citations in the essays and source memory, after controlling for the effect of total text reading time (for the indices based on dwell time) and total number of fixations on the text (for the indices based on fixation number). in the real documents condition, there were no significant correlations between the eye-tracking measures and indicators of offline sourcing (table 2, first four columns). in the print-out condition, however, second-pass measures correlated positively with source citations, but not with source memory (table 2, last four columns). table 2 partial correlations between sourcing as measured by continuous eye-movement indices and offline indicators of sourcing by condition (real documents and print-out versions), controlling for the effect of total text reading time (for the indices based on dwell time) or total number of fixations on the text (for the indices based on fixation number). note: * p < .05 3.3. sourcing as measured by discontinuous eye-movement indices we first compared the number of sources fixated by participants (maximum 4 sources), as a function of fixation threshold, by condition (real documents and print-out versions). the number of sources fixated varied considerable between thresholds, and across conditions (table 3). chi square analyses comparing condition indicated that there was no significant difference for the 100 msec threshold, 2(3)= 3.93, p = .27, but there were significant differences for the 500 and 1000 msec thresholds (2(4)= 13.47, p < .01 and 2(4)= 12.44, p = .01, respectively). for the higher thresholds, participants in the real document conditions fixated less sources than in the print-out version. table 3 number of sources fixated by participants, as a function of fixation threshold, by condition (real documents and print-out versions) (in percentage of participants). next, to test the validity of the discontinuous eye-movement indices of sourcing, we computed partial correlations between the discontinuous indices of sourcing and the offline sourcing measures, controlling for the effect of total text reading time. in the real documents condition, there were no correlations between the eye-tracking measures and offline sourcing (table 4, left columns). in the print-out condition, however, there were positive relations between number of sources fixated for at least 100, 500, and 1000 msec and source citations (table 4, right columns). memory for sources correlated only with the index using the 100 msec threshold. table 4 partial correlations between sourcing as measured by discontinuous eye-movements indices and offline indicators of sourcing by condition (real documents and print-out versions), controlling for the effect of total text reading time. note: * p < .05 3.4. sourcing as measured in the interview to assess the predictive power of sourcing as measured in the interview, we computed a series of anovas with condition (real vs. print-out versions of documents) and reported sourcing (yes or no) as independent variables, and the two offline sourcing measures as dependent variables. there were main effects of reported sourcing on the offline sourcing measures, with students reportedly having attended to sources while reading scoring higher on the offline measures (see table 5, right column). none of the interactions between condition and reported sourcing were significant. in other words, the effects of reported sourcing were independent of whether students read real or print-out versions of the documents. table 5 anovas of reported sourcing (as measured in the interview) and condition (real documents and print-out versions) on offline indicators of sourcing. 3.5. triangulation of the sourcing measures finally, we compared sourcing as indicated by the different eye-movement indices and the interview. specifically, we ran a set of anovas with reported sourcing as independent variable and eye-movement indices as dependent variables (see table 6). overall, there were no major differences in sourcing as indicated by the eye-movements between students who reportedly had or had not paid attention to sources while reading. thus, regardless of what they reported about sourcing in the interview, students apparently looked at source information to some extent. the only significant difference emerged for the number of first pass fixations on source information, which increased for students who reportedly paid attention to sources while reading. this effect was especially pronounced in the print-out versions of the documents (table 6, right column). in this condition, the trend was observed for other eye-tracking indicators, with participants who reportedly paid attention to sources seemingly looking longer at and fixating more on source information. however, such differences did not reach statistical significance, probably due to a lack of statistical power. table 6 anovas of reported sourcing (as measured in the interview) and condition (real documents and print-out versions) on two sets of dependent variables (continuous and discontinuous eye-movement indices of sourcing). 4. discussion and conclusions results from our study converge with existing evidence suggesting that current eye-tracking indices of sourcing have several limitations (brand‐gruwel et al., 2017; gerjets et al., 2011; mason et al., 2014; van strien et al., 2016), especially when they are not tied to experimental manipulations performed to test specific hypotheses based on a theoretical model of sourcing, such as the d-isc (braasch & bråten, 2017). our results also suggest that simpler methods, such as asking readers whether they paid attention to source information, actually may be better suited to assess strategic sourcing during multiple document reading. specifically, sourcing measured by means of the interview predicted two different sourcing measures, source citations and source memory, independently of the type of reading materials (real documents or print-out versions). relationships between sourcing measured by eye-movement indices and offline sourcing were more complex, however, as they depended on the type of reading materials and the type of indices used. thus, neither continuous nor discontinuous eye movement indices predicted offline sourcing when reading real documents. this probably reflects the fact that fixations to source information in real documents is a combination of bottom-up and top-down processes, which makes such indicators less valid (orquin & holmqvist, 2018). yet, for print-out versions of the documents, late eye movement continuous indices (i.e., second dwell time and number of second pass fixations), which usually reflect more strategic processing (rayner, 2009), correlated positively with source citations, but not source memory. when using dichotomous eye movement indices sourcing as indicated by a 100 msec threshold positively correlated with both source memory and source citations, and higher thresholds (i.e., 500 and 1000 msec) positively correlated with source citations. of note is that in these analyses, potential effects due to the time participants devoted to reading the actual text were controlled for. the fact that participants accurately reported their sourcing activities in the interview support the idea that attending to and processing source information while reading multiple documents represent strategic and conscious activities, which is consistent with the d-isc model recently elaborated by braasch and bråten (2017). presumably, readers resort to this strategy in an effort to create coherent, meaningful mental representations from diverse documents on the same topic, for which they may need to include source information as organizational elements (e.g., to qualify claims according to source or better understand conflicts among sources) (see also van den broek & kendeou, 2015). as students are aware of such strategic sourcing activities in the service of meaning-making, post-task interviews seem suited to capture the use of sourcing in multiple document contexts. in comparison, measures based on eye-movements may incorporate a mixture of automatic and strategic processes, which make such indices less valid when used during the reading of more complex and diverse reading materials. the mixture of automatic and strategic processes in eye-tracking data seemed to be particularly the case when participants read real documents, which included more visually salient information not present in the print-out versions. for example, the nameplate of the newspaper “el país” occupied a big portion of the upper part of the front page of the real document, while this source feature in the print-out version had the similar size as other textual information. overall, the results challenge the current use of eye-tracking indices solely based on fixations at source information (e.g., logos or ‘about us’ information). we suggest that future work could increase the validity of such measures in at least two different ways. first, although most students directly attend to source features for short periods of time, they may also evaluate sources by focusing on the quality of the arguments in the texts. that is, readers will not necessarily have to look at source information to judge the credible of a document. accordingly, a combination of indices of source and text processing may provide a more valid picture of students’ sourcing activities. second, as indicated by our literature review, another way to improve the validity of eye-movement indices of sourcing in such challenging reading task contexts is to identify specific reading episodes in which sourcing may be expected based on theoretical assumptions, such as those forwarded by the d-isc model (braasch et al., 2012; kammerer et al., 2016). sourcing could then be analyzed after the critical episode takes place, for example, after students finish reading a conflicting claim likely to involve a break in situational coherence. presumably, this will reduce the probability that non-strategic, automatic processing is captured by the eye-tracking indicators. in the absence of such theoretical assumptions, when the goal of the study is to explore sourcing from beginning to end during reading, subjective methods such as interviews seem preferable when assessing strategic sourcing. this study does not come without limitations, of course. among them are the relatively small sample size, which made statistical power lower than desirable, the particular eye-tracker that we used, which may have influenced the accuracy level of the data, and the complexity of the reading materials that we presented. moreover, it cannot be entirely ruled out that participants’ reports of sourcing were influenced by their performance on the preceding source memory task. we could go on describing further potential limitations and qualifications. at the same time, given the sample, the apparatus, the reading materials, and the operationalizations that we did use, we believe our results still merit attention and hope they will provide impetus to more careful consideration of eye movement indices as valid measures of sourcing in future research. keypoints use of eye-movements to investigate complex reading requires critical reflection on the validity of the measures eye-movement indicators of students’ sourcing do not always predict offline sourcing interview data seem to be a more valid indicator of strategic sourcing in multiple document contexts acknowledgments funding for the research reported in this article was provided by grant edu2014-59422 from the spanish secretaría general de universidades to ladislao salmerón and by grant 237981/h20 from the research council of norway to ivar bråten. references anmarkrud, ø., bråten, i., & strømsø, h.i. (2014). multiple-documents literacy: strategic processing, source awareness, and argumentation when reading multiple conflicting documents. learning and individual differences, 30, 64-76. barzilai, s., & eshet-alkalai, y. (2015). the role of epistemic perspectives in comprehension of multiple author viewpoints . learning and instruction, 36, 86-103. barzilai, s., tzadok, e., & eshet-alkalai, y. (2015). sourcing while reading divergent expert accounts: pathways from views of knowing to written argumentation. instructional science, 43, 737-766. brand‐gruwel, s., kammerer, y., van meeuwen, l., & van gog, t. (2017). source evaluation of domain experts and novices during web search. journal of computer assisted learning, 33, 234-251. braasch, j.l.g., & bråten, i. (2017). the discrepancy-induced source comprehension (d-isc) model: basic assumptions and preliminary evidence. educational psychologist, 52, 167-181. braasch, j.l.g., bråten, i., strømsø, h.i., & anmarkrud, ø. (2014). incremental theories of intelligence predict multiple document comprehension. learning and individual differences, 31, 11-20. braasch, j. l.g., rouet, j. f., vibert, n., & britt, m. a. (2012). readers’ use of source information in text comprehension. memory & cognition, 40, 450-465. bråten, i., & braasch, j.l.g. (2018). the role of conflict in multiple source use. in j.l.g. braasch, i. bråten, & m.t. mccrudden, m.t. (eds.), handbook of multiple source use (pp. 184-201). new york: routledge. bråten, i., strømsø, h.i., & britt, m.a. (2009). trust matters: examining the role of source evaluation in students’ construction of meaning within and across multiple texts. reading research quarterly, 44, 6-28. bråten, i., strømsø, h.i., & salmerón, l. (2011). trust and mistrust when students read multiple information sources about climate change. learning and instruction, 21, 180-192. britt, m. a., rouet, j. f., & braasch, j. l. g. (2013). documents as entities: extending the situation model theory of comprehension. in m. a. britt, s. r. goldman, & j. f. rouet (eds.), reading: from words to multiple texts (pp. 160–179). new york: routledge. bromme, r., stadtler, m., & scharrer, l. (2018). the provenance of certainty: multiple source use and the public engagement with science. in j. l. g. braasch, i. bråten, & m. t. mccrudden (eds.), handbook of multiple source use (pp. 269-284). new york, ny: routledge. catrysse, l., gijbels, d., donche, v., de maeyer, s., lesterhuis, m., & van den bossche, p. (2018). how are learning strategies reflected in the eyes? combining results from self‐reports and eye‐tracking. british journal of educational psychology, 88, 118-137. gerjets, p., kammerer, y., & werner, b. (2011). measuring spontaneous and instructed evaluation processes during web search: integrating concurrent thinking-aloud protocols and eye-tracking data. learning and instruction, 21, 220-231. goldman, s.r., braasch, j.l.g., wiley, j., graesser, a.c., & brodowinska, k. (2012). comprehending and learning from internet sources: processing patterns of better and poorer learners. reading research quarterly, 47, 356-381. hyönä, j., lorch jr., r. f., & kaakinen, j. k. (2002). individual differences in reading to summarize expository text: evidence from eye fixation patterns. journal of educational psychology, 94, 44-55. jacobs, a. m. (2016). the scientific study of literary experience. scientific study of literature, 5, 139-170. jarodzka, h., & brand‐gruwel, s. (2017). tracking the reading eye: towards a model of real‐world reading. journal of computer assisted learning, 33, 193-201. just, m. a., & carpenter, p. a. (1980). a theory of reading: from eye fixations to comprehension. psychological review, 87, 329-354. kaakinen, j. k., hyönä, j., & keenan, j. m. (2002). perspective effects on online text processing. discourse processes, 33, 159-173. kammerer, y., & gerjets, p. (2014). the role of search result position and source trustworthiness in the selection of web search results when using a list or a grid interface. international journal of human computer interaction, 30, 177–191. kammerer, y., kalbfell, e., & gerjets, p. (2016). is this information source commercially biased? how contradictions between web pages stimulate the consideration of source information. discourse processes, 53, 430-456. landis, j. r., & koch, g. g. (1977). the measurement of observer agreement for categorical data. biometrics, 33, 159-174. leinenger, m., & rayner, k. (2017). what we know about skilled, beginning, and older readers from monitoring their eye movements. j. a. león & i. escudero (eds.), reading comprehension in educational settings (pp. 1-17). amsterdam: john benjamins. mason, l., pluchino, p., & ariasi, n. (2014). reading information about a scientific phenomenon on webpages varying for reliability: an eye-movement analysis. educational technology research and development, 62, 663-685. mccormick, p. a. (1997). orienting attention without awareness. journal of experimental psychology: human perception and performance, 23, 168-180. orquin, j. l., & holmqvist, k. (2018). threats to the validity of eye-movement research in psychology. behavior research methods, 50, 1645-1656. rayner, k. (1998). eye movements in reading and information processing: 20 years of research. psychological bulletin, 124, 372-422. rayner, k. (2009). the 35th sir frederick bartlett lecture: eye movements and attention in reading, scene perception, and visual search. quarterly journal of experimental psychology, 62, 1457-1506. rouet, j.f., britt, m.a., mason, r.a., & perfetti, c.a. (1996). using multiple sources of evidence to reason about history. journal of educational psychology, 88, 478–493. rouet, j.f., favart, m., britt, m.a., & perfetti, c.a. (1997). studying and using multiple documents in history: effects of discipline expertise. cognition and instruction, 15, 85-106. salmerón, l., baccino, t., cañas, j.j., madrid, r. i., & fajardo, i. (2009). do graphical overviews facilitate or hinder comprehension in hypertext? computers & education, 53, 1308-1319. salmerón, l., gil, l., & bråten, i. (2018). effects of reading real versus print-out versions of multiple documents on students’ sourcing and integrated understanding. contemporary educational psychology, 52, 25-35. salmerón, l., kammerer, y., & garcía-carrión, p. (2013). searching the web for conflicting topics: page and user factors. computers in human behavior, 29, 2161–2171. salmerón, l., naumann, j., garcía, v., & fajardo, i. (2017). scanning and deep processing of information in hypertext: an eye-tracking and cued retrospective think-aloud study. journal of computer assisted learning, 33, 222–233. stadtler, m., & bromme, r. (2014). the content-source integration model: a taxonomic description of how readers comprehend conflicting scientific information. in d.n. rapp & j.l.g. braasch (eds.), processing inaccurate information: theoretical and applied perspectives from cognitive science and the educational sciences (pp. 379-402). cambridge, ma: the mit press. stang lund, e., bråten, i., brandmo, c., brante, e.w., & strømsø, h.i. (in press). direct and indirect effects of textual and individual factors on source-content integration when reading about a socio-scientific issue. reading and writing: an interdisciplinary journal. strømsø, h.i., bråten, i., & britt, m.a. (2010). reading multiple texts about climate change: the relationship between memory for sources and text comprehension. learning and instruction, 20, 192-204. van den broek, p., & kendeou, p. (2015). building coherence in web-based and other non-traditional reading environments: cognitive opportunities and challenges. in r.j. spiro, m. deschryver, m.s. hagerman, p.m. morsink, & p. thompson (eds.), reading at a crossroads? disjunctures and continuities in current conceptions and practices (pp. 104-114). new york: routledge. van strien, j. l., kammerer, y., brand-gruwel, s., & boshuizen, h. p. (2016). how attitude strength biases information processing and evaluation on the web. computers in human behavior, 60, 245-252. wiley, j., goldman, s.r., graesser, a.c., sanchez, c.a., ash, i.k., & hemmerich, j.a. (2009). source evaluation, comprehension, and learning in internet science inquiry tasks. american educational research journal, 46, 1060-1106. microsoft word sjöblom et al_publication.docx           frontline  learning  research  vol.4  no.  1  (2016)  17-­‐39   issn  2295-­‐3159       corresponding author: kirsi sjöblom, research group of educational psychology, department of teacher education, faculty of behavioural sciences, p.b. 9 (siltavuorenpenger 5), 00014 university of helsinki, finland. e-mail: kirsi.sjoblom@helsinki.fi doi: http://dx.doi.org/10.14786/flr.v4i1.217 does physical environment contribute to basic psychological needs? a self-determination theory perspective on learning in the chemistry laboratory kirsi sjöblom, kaisu mälkki, niclas sandström, kirsti lonka university of helsinki, finland article received 1 / october / revised 28 december / accepted 8 january / available online 10 february   abstract the role of motivation and emotions in learning has been extensively studied in recent years; however, research on the role of the physical environment still remains scarce. this study examined the role of the physical environment in the learning process from the perspective of basic psychological needs. although self-determination theory stresses the role of the social and cultural environment, as yet the role of the physical environment has been unexplored. the study focused on beginning chemistry university students’ (n=21) experiences in a chemistry laboratory. the data consisted of focus-group interviews and self-report questionnaires. the results indicate that the physical environment can support or thwart the fulfillment of the basic psychological needs. the usability and functionality of spaces and tools contributed to not just the fluency of the intellectual activity but also to the related emotional experience of oneself acting in a particular environment. the physical environment was a source of procedural facilitation: it complemented and challenged the students’ existing skills, contributing to their experiences of autonomy and competence. the everyday successes or struggles in the laboratory built on the students’ developing professional identity as well as their sense of belonging to the professional community. this study demonstrates that the design and functionality of the physical environment has a significant role in users’ intellectual and emotional functioning. it is essential to utilize psychological and pedagogical knowledge when designing or renovating work and learning environments in order to fully make use of the potential of physical environments as part of human performance. keywords: self-determination theory; basic psychological needs; physical environment; learning environment; indoor environment; usability sjöblom  et  al       | f l r     18   1. introduction in recent years, the broadening field of research on the role of motivation and emotions in learning has produced important new information on how to optimally arrange the study environment (see e.g. csíkszentmihályi, 2014; dweck, 2006; heikkilä & lonka, 2006; heikkilä, lonka, nieminen & niemivirta, 2012; hidi & renninger, 2006; job, walton, bernecker & dweck, 2015; lindblom-ylänne & lonka, 2000; mälkki, 2010; ryan & deci, 2009; seligman, ernst, gillham, reivich & linkins, 2009; tuominen-soini, salmela-aro & niemivirta, 2008). strikingly, even though knowledge on the study environment, and especially its social attributes, is vast, knowledge on how the physical environment is related to psychological and pedagogical phenomena as yet remains scarce (sandström, sjöblom, mälkki & lonka, 2013; beard 2009, 2012; lansdale, parkin, austin & baguley, 2011; lonka, 2012; woolner, hall, higgins, mccaughey & wall, 2007). intellectual and emotional functioning is always nested in the physical environment, even when working in virtual learning environments. however, most of the research on physical environment has traditionally focused on minimizing its negative effects on health or determining how individuals interact with the environment on a perceptual level (see e.g. alfonsi, capolongo & buffoli, 2014; evans, bullinger & hygge, 1998; parsons & hartig, 2000; ulrich, 1981), rather than on unveiling the role of the physical environment with regard to cognitive and emotional functioning. this study examines the role of the physical environment in supporting learning and basic psychological needs. previous research has indicated that the physical environment is far from irrelevant with regard to intellectual functioning: the design and functionality of the physical environment contribute to physically distributed intelligence (norman, 1993), stress over safety issues and the cognitive capacity available for higher intellectual functioning such as learning (sandström, sjöblom, mälkki & lonka, 2013). being organized in a given way, the physical space also conveys assumptions and ideologies (beard, 2012; beard & price, 2010) e.g. on the activity taking place and, as such, tunes the users into different mental modes and roles (mälkki, sjöblom & lonka, 2014). thus, similarly to the social environment, the physical environment can be seen as either facilitating learning and well-being or posing a challenge to them. moreover, of particular interest is the emotional experience related to the activity taking place in a given physical space. this experience may likely bear meaning in the process of forming a relation to the place and, more broadly, of developing one’s identity as a professional in a given field. in modern-day society people spend most of their time in indoor environments, and new multidisciplinary information is needed on how to design these spaces to best support the activity expected to take place in them. both human resources and physical spaces are valuable and costly resources: typically around 90% of business operating costs consist of direct or indirect staff costs (alker et al., 2015), and as to physical spaces, expensive indoor environments need to be used efficiently. at the same time, the industry policy of most western societies prioritizes innovation. we need to acquire further knowledge on how to facilitate the thriving of the human potential by creating fruitful grounds for it. when designing physical learning spaces, it is essential to not only take into account the most fundamental needs of the students, but also to gain understanding on the relations between the physical surroundings and the more refined psychological processes. these issues are focal in both learning environments and environments dedicated to other purposes, such as work or recreation. finally, it is not quite enough to focus on the design and functionality of physical space and tools as such. the use of available premises and equipment is essentially determined by the social practices applied in them; for instance, technology advances learning only through transformed social practices (hakkarainen, 2009; paavola, lipponen & hakkarainen, 2004). thus, although in this article we examine the role of the physical environment in the fulfillment of basic psychological needs, we do not assume that it is only a matter of a relation between the individual and the physical environment. rather, we approach the theme from the perspective that the users’ experience of the physical environment is mediated by social practices and culturally shared meanings. in a broader sense, we are approaching the intriguing interplay between the human and the material, as well as the intellectual and the emotional. sjöblom  et  al       | f l r     19   2. theoretical framework 2.1 basic psychological needs in this study we approach questions of learning and well-being with regard to the physical learning environment from the perspective of basic psychological needs as laid out by the self-determination theory developed by deci and ryan (1985, 2000, 2008; ryan & deci 2000, 2009). this is a macro-theory of human motivation, personality development and well-being that focuses especially on volitional behavior and the surrounding conditions that support it (ryan, 2009). the theory views all human beings as inherently self-determined, actively evolving organisms, with a natural aspiration for continuous psychological development and growth. however, in order to these propensities to be actualized, the satisfaction of basic psychological needs must be sufficiently supported. according to the theory, the social and cultural environment can support the satisfaction of basic psychological needs and the self-determined behavior to varying degrees. thus, the process of growth is essentially seen to take place in relation to the surrounding conditions that, for their part, contribute to the individuals’ possibilities to embrace their full, natural potential. aligned with this emphasis, it is also relevant to study in more detail how the physical environment may, for its part, contribute to the interaction between the individual and the environment and the fulfilment of basic psychological needs (e. deci, personal communication with the first author, october 28, 2014). self-determination theory is currently one of the most prevalent and utilized theories on motivation. in the decades following the formal introduction of the theory in the 1980s, research on the theory has dramatically increased. consequently, the theory has been subject to criticisms and suggestions for further development as well. a common criticism of the theory is its cultural applicability, posing that the core features of the theory, such as the need for autonomy, are mainly descriptive of a western individual, rather than of people raised in and surrounded by more collectivist cultures (e.g. iyengar & devoe, 2003; markus & kitayama, 1991). however, further research has verified that psychological needs are equally imperative with regard to psychological well-being in both individualistic and collectivist cultures (e.g. chirkov, ryan, kim & kaplan, 2003; ryan & deci, 2006). the formal framework of self-determination theory consists of five mini-theories (ryan, 2009). this study focuses on the mini-theory of basic psychological needs. the theory states that all people, universally and regardless of their age or gender, share the same basic psychological needs, namely the needs for autonomy, competence and relatedness. these needs are seen to be central prerequisites with regard to healthy human functioning. autonomy refers to perceiving oneself as the origin or source for one’s own behavior (deci & ryan, 1985; ryan & connell, 1989; ryan & deci, 2002, 2006), competence refers to a felt sense of confidence and effectance in one’s own actions (ryan & deci, 2002), and relatedness refers to feeling connected and having a sense of belonging with regard to both other individuals and with one’s community (baumeister & leary, 1995; ryan, 1995; ryan & deci, 2002). in order to function effectively and to be psychologically healthy, these needs must be sufficiently satisfied (deci & ryan, 2008). more specifically, the satisfaction of basic psychological needs is relative to the activity and functioning pursued; needs may be seen to specify necessary nutriments with regard to healthy development and vitality as well as constructive and creative outputs (deci & ryan, 2002). thus, rather than being a goal in itself, the satisfaction of basic psychological needs is seen to facilitate intrinsic motivation, learning and well-being (niemiec & ryan, 2009; ryan & deci, 2009) as well as eudaimonic happiness (ryan, huta & deci, 2008). the theory of basic psychological needs is widely studied empirically, including in the context of learning in higher education (see e.g. black & deci, 2000). in particular, the need for autonomy and the possibilities to support it have acquired much needed attention in the context of learning and instruction (see sjöblom  et  al       | f l r     20   e.g. jang, reeve & deci, 2010; niemiec & ryan, 2009; soenens, sierens, vansteenkiste, goossens & dochy, 2012; vansteenkiste et al., 2012). however, the research has predominantly focused on the social aspects of the learning environment, such as the interaction between the students and the teacher, while the research on the psychological needs of an individual with regard to the physical environment has been extremely scarce (see e.g. gay, 2008; gay, saunders, & dowda, 2011; rutten, boen & seghers, 2012). 2.2 the learning environment similarly to the research on the basic psychological needs, research on learning environments has mainly focused on the social learning environment while the physical learning environment has for the most part been ignored. for example, lave and wenger’s idea of legitimate peripheral participation (1991) places high importance on social engagements that provide the proper context for learning to take place. by participating in the activities of an expert community, a novice is gradually able to assimilate the professional practices and become part of the community. these kinds of views stress the role of the social learning environment in the development of professional abilities, yet neglect the physical environments in which the social activity takes place. empirical research on physical environments, on the other hand, has traditionally focused on factors related to physical health or discomfort (e.g. küller & lindsten, 1992; winterbottom & wilkins, 2009). knowledge on how the physical environment, i.e. physical spaces, tools and equipment, is related to psychological and pedagogical phenomena is still rare (lansdale, parkin, austin & baguley, 2011; lonka, 2012; woolner, hall, higgins, mccaughey & wall, 2007). while the importance of individual characteristics and the social environment should not be underestimated (e.g. perry, turner & meyer, 2006), the role of the physical environment in the learning process calls for more rigorous attention in the field of learning research. more knowledge is needed on how the physical environment can support learning, wellbeing, engagement and commitment. research on learning environments has shown that the physical environment conveys assumptions (beard, 2012; beard & price, 2010) and activates students’ previous assumptions regarding similar environments (mälkki, sjöblom & lonka, 2014). the assumptions conveyed by the physical environment may involve underlying conceptions on the learning process and the roles of the participants: an auditorium implies a different positioning and division of roles than a classroom where the desks are organized in groups and the teacher has no central position but is instead moving around the classroom on a chair. this demonstrates how the physical space itself tunes the students into different mental modes and roles. the arrangement of physical space in ways that the participants are not used to may as such turn into a disorienting dilemma, challenging existing conceptions and ways of thinking and possibly triggering reflection (mälkki, sjöblom & lonka, 2014). thus, the space or equipment cannot be seen as a separate entity, detached from the present culture. rather, social practices are embedded in the physical arrangements (hakkarainen, 2009) and also have an impact on how the physical environment is perceived and experienced by the users. along with the idea of socially and physically distributed cognition (hakkarainen, palonen, paavola & lehtinen, 2004; hutchins, 2000, 2006), physical environments also vary with regard to the degree they facilitate the activity that is expected to take place in them. for example, the space may be equipped with modern technology and devices that assist the learning process, which makes the learning process markedly different from one that is carried out without any needed assistance, such as calculators, to begin with. the very fact that learners are able to choose a suitable environment for different learning tasks is helpful with regard to completing the tasks. a concrete example of this might be having to work on a group assignment in a silent library hall or endeavoring to understand new theoretical material in a noisy hallway. in fact, the physical environment consists of affordances that may, at best, facilitate the development of new skills, help people overcome the limitations of their own capabilities and make them feel like active agents; or in contrast, the lack of needed affordances may pose a significant challenge to carrying out the sjöblom  et  al       | f l r     21   expected activities, handicapping the cognitive functioning in the space and making people feel incapable of performing the expected tasks (sandström, sjöblom, mälkki & lonka, 2013; norman, 1993; sandström, eriksson, lonka & nenonen, 2015). thus, the physical environment for its part offers a varying degree of procedural facilitation (bereiter & scardamalia, 1987) of the aspired activity. if for instance students lack enough space for their work or constantly have to worry about unclear safety issues, these issues inevitably take a toll on the cognitive resources available for learning (sandström, sjöblom, mälkki & lonka, 2013; see also sandström, ketonen & lonka, 2014). thus, a dysfunctional environment may be handicapping with regard to intellectual activity at the most basic level. consequently, we postulate that the design and the functionality of the physical environment play a role in the students’ experiences related to the basic psychological needs. 2.3 the context of the study: exploring the basic psychological needs in a chemistry laboratory learning environment in our study we focus on beginning university chemistry students’ learning, in particular on their experiences during laboratory work, in order to unveil the dynamics between physical environment and basic psychological needs. chemistry as a study context offers an intriguing and relevant terrain for researching this interplay. namely, the physical laboratory environment, which includes not only desks and chairs but also the diverse and complex laboratory instrumentation, is especially focal in learning chemistry. focusing on the experiences of first-year students is fruitful from the perspective of their emerging sense of relatedness to the professional field. furthermore, sense of autonomy and competence are expected to develop in a study context, which, similarly to a working environment, represents a performance-oriented environment. in addition, the aforementioned topics may be particularly present in the students’ experiences in the beginning stage of their studies. as argued earlier in the text, current research on basic psychological needs in study contexts has focused especially on students’ sense of autonomy. this is a central question as a study context has traditionally been an environment where the action to a large extent is guided by the teacher, while at the same time, the students have the need to develop their sense of autonomy and competence in the field. this need for a constructive friction between the students’ existing capabilities and an appropriate amount of guidance provided by the teacher has also been addressed by vermunt and verloop (1999). in our view, it is important to look more closely at the emerging sense of relatedness with regard to the study community and the physical premises, and more generally, to the professional field. in the chemistry context this may have particular importance: for example, in finland many students discontinue their chemistry studies after a year or two. for some of these students, this may be due to a transfer to pursue studies in the faculty of medicine, where the chemistry studies serve as a platform to develop the abilities needed to be accepted into that faculty. however, this is not the case for all of the students who drop out of their chemistry studies. in order to increase understanding on student experiences in the chemistry learning context, it is particularly interesting to examine the role of the physical learning environment from the viewpoint of psychological needs and the support the physical environment could offer for learning. what are the most central characteristics of the physical environment that contribute to emerging experiences of competence, autonomy and relatedness? if we are able to consider the psychological needs of the students when designing learning spaces, we can create a fruitful ground for thriving, productive students and, at best, further current understanding on how to design leading university campuses (lonka, 2012; nenonen, kärnä, junnonen, tähtinen & sandström, 2015). sjöblom  et  al       | f l r     22   3. the aims of the study this study explored the role of the physical environment with regard to learning from the perspective of basic psychological needs. the research questions were as follows: 1. what is the role of the physical environment in the experience of the basic psychological needs? 2. what is the role of the physical environment in the learning process from the perspective of the basic psychological needs? aligned with the theory of basic psychological needs, we postulated that the satisfaction of these needs is not a goal as such, but rather a facilitator with regard to productivity and well-being. consequently, it was relevant to study the experience of the basic psychological needs in relation to the activity pursued, that is, learning. we hypothesized that if the physical environment contributes to the experience of basic psychological needs, this may have a mediating effect on the process of learning; by supporting the fulfillment of basic psychological needs, the physical environment may facilitate learning and study engagement. in addition, we aimed at furthering the interactional perspective of the theory of basic psychological needs by considering the role of the physical environment in facilitating or posing a challenge to the fulfillment of the core needs. rather than focusing on individual experiences regarding the core needs, our emphasis was on exploring the dynamics of the phenomenon on a more theoretical level. in order to capture the diversity and depth of the student’s experiences regarding this fairly new research topic, the study approached the relations between the physical environment and basic psychological needs with qualitative methodology. while much of the research on motivation is based on self-report questionnaires in order to measure individuals’ views and beliefs, classroom observations and interviews can provide a richer depiction of situated motivation (wigfield, cambria & eccles, 2012). 4. method 4.1 participants the participants of the study were beginning-stage chemistry students (n=21, representing both genders) from a finnish university. the participants were selected based on their willingness to participate as well as the appropriate timing of their current laboratory project; in other words, participation in the interview and selection for a particular focus group also depended on whether they were able to leave their laboratory work for an hour to complete the interview. 4.2 materials the data consists of focus group interviews and questionnaires that were completed by each participant individually before entering the interview. the questionnaire served as an orientation to the interview, whereas the qualitative analysis is based on the material from the focus group interviews. the questionnaire included both open-ended and multiple choice questions. the themes of the questionnaire focused on helpful and challenging aspects of the physical environment with regard to learning as well as typical study-related use of physical spaces, equipment and technological devices: a) sources of interest and engagement in the laboratory work (open-ended), b) sources of challenge and difficulty in the laboratory work (open-ended), c) typical study-related use of technological tools in learning (multiple choice questions assessing the frequency of the use on a scale 1-6; e.g. smartphone, laptop), sjöblom  et  al       | f l r     23   d) typical study-related use of spaces in learning (multiple choice questions assessing the frequency of the use on a scale 1-6; e.g. library, hallways, cafeterias, home), e) concrete tools, equipment or other aspects of the laboratory work that are experienced as particularly well-functioning or engaging (open-ended), f) concrete tools, equipment or other aspects of the laboratory work that are experienced as particularly cumbersome or counterproductive with regard to learning (open-ended), g) other comments and suggestions with regard to the physical learning environment (openended). the interview elaborated on the same questions with the group. 4.3 procedures 4.3.1 interviews semi-structured focus-group interviews in groups of three to four students were collaboratively carried out by two of the authors. the interviews were conducted contextually in the middle of a laboratory work session. the students completed the questionnaires in the actual laboratory space, an organic chemistry laboratory, and the interviews were carried out in an adjoining room in order to ensure privacy and focused environment. by having the students complete the questionnaire individually before entering the interview, we aimed at giving the students the space to reflect on the topics based on their own experience and perspective first, and the views could then be elaborated further in the group. the interviews followed an interpretivist approach (scott & usher, 1999; williams, 2000), aiming at "making sense of actor's actions and language within their 'natural' setting" (williams, 2000). the method of the interview was designed to leave space for the participants to freely discuss themes that they experienced as important. as the topic of the research is rather new, the structure of the questionnaire and the interview had to be open enough not to restrict the participants but to genuinely leave space for unexpected material and directions, regardless of the preconceptions or hypotheses of the researchers. moreover, the phenomena and the related experiences are such that a clearly articulated view from the students is hardly expected; rather, the data had to be approached in a holistic way to seek understanding on the phenomena. thus, the questions were formed rather open so that the interview and the discussion in the groups could develop the topics further. as a result, the interview data brought about a rich milieu of aspects of the students’ experiences, beyond the expected themes and hypotheses. 4.3.2 analysis the interviews were transcribed verbatim, and the transcriptions were then analyzed by the authors iteratively with the help of the atlas ti program. repeated stages of individual and collaborative analysis were conducted to find central categories and patterns in the reported experiences. the initial stage of the analysis and classification was data driven in order to capture unforeseen observations and patterns in the data. when the researchers gathered to discuss the initial results of the first round of analysis that each had conducted individually, it was noted that despite the differences in the conceptualizations and terminologies of the classifications among the researchers, many of the central themes and categories fell into the dimensions of basic psychological needs. the results of the first round of analysis supported the theory of basic psychological needs as a relevant theoretical approach through which to frame the findings and acquire further understanding on the role of the physical environment in the learning process. the following rounds of analysis focused on elaborating specifically on this approach with continued iterative individual and collaborative rounds. sjöblom  et  al       | f l r     24   indeed, the theory of basic psychological needs was not yet our framework when collecting data. we aimed at more generally unveiling the role of the physical environment in the learning process. along with the data-driven analyses and initial findings, we started seeing the relevance of further rounds of analysis from the perspective of basic psychological needs. consequently, the analysis is by no means exhaustive with regard to the relation between basic psychological needs and the physical environment but rather an opening for research on the topic. this study was a deepening reanalysis of previously analyzed data (sandström, sjöblom, mälkki & lonka, 2013) on the chemistry laboratory as a physical learning environment. the previous study shed light on the role of the physical space in the learning process: the physical space may contain guidance implemented in it, and the physical space and its usability contribute to the students' sense of safety, which in turn is crucial when students are expected to engage in demanding cognitive activities. however, early in the initial phases of analysis, it seemed that in addition to the aforementioned findings, the data also offered intriguing perspectives on the dynamics between the physical environment, learning and experiences of oneself as a learner in that given environment, which deserved a deepening reanalysis. the authors represented different fields of expertise, namely educational psychology, clinical psychology, adult education and linguistics. the analysis aimed at utilizing and building on the diversity of the scholarly backgrounds of the researchers to explore different approaches to the phenomena as well as reach understanding on the core features presented in the data. as with the participants of the study, we aimed at both capturing the individual approaches and views as well as elaborating them further by combining the views and abilities of the whole group collaboratively. most of the work on the study was carried out collaboratively, with the team of researchers working on the material and writing the text in the same physical space, which added value to the depth of the analysis, as opposed to each researcher separately adding their own share of expertise to the study (hakkarainen, palonen, paavola & lehtinen, 2004). moreover, the researchers altered and modified the physical spaces in which they were working during the research process. choosing a suitable physical space with the required technological tools to accommodate a given work assignment, for example, a collaborative writing session, brought further understanding on the role of the physical space in the work process itself. the approach to the current study was abductive by nature; by utilizing the theory in the analysis of the data, we aimed at a deeper understanding of the phenomenon as well as at furthering the theory. our main focus was on the dynamics between the physical environment and the experiences of the learner rather than on a purely deductive approach driven by an emphasis on testing the theory. from a methodological point of view, our aim was not to cover all possible variations of the interplay between the physical environment and the psychological needs in the context of chemistry studies. rather, our study was aimed at serving as an opening for research on previously unmapped ground. even though the sample size can be seen as a limitation of the study and a broader sample could have been advantageous, from a theoretical point of view (see mälkki, 2012) the data were rich and offered relevant material for an exploratory analysis on the dynamics of the topic. 5. results in the following sections we will focus on how the three core needs, autonomy, competence and relatedness (ryan & deci, 2002), manifest in relation to the physical environment and the learning context. as our approach stretches the theory of psychological needs out of its usual sphere of application, we employed an abductive approach to be open to dynamics of the phenomenon that are not readily conceptualized in self-determination theory. for analytical clarity, we will in the following first examine each dimension individually, and secondly we will discuss how these dimensions are intertwined in the data. sjöblom  et  al       | f l r     25   5.1 autonomy 5.1.1 physically mediated guidance and the use of modern technological devices in supporting students’ sense of autonomy within the context of learning and instruction, the issue of autonomy is often regarded to predominantly concern the balance between the control over one’s work and the received guidance, which is usually seen as socially mediated. students need sufficient guidance and should not be “abandoned,” but the teacher should not regulate or perform on behalf of the students the tasks and challenges that they already master, thus disturbing the sense of autonomy experienced by the students. as for the laboratory as a physical entity, guidance may be seen not merely as socially mediated but also as physically mediated (sandström, sjöblom, mälkki & lonka, 2013; hutchins, 2006); information may be embedded in the physical space itself. for instance, different tags and signs can be seen as affordances (norman, 1993) that assist individual information processing. they help people overcome the boundaries of their intellectual capacities. thus, the workspace itself can be seen as cognitively structuring, also with regard to the clarity of close surroundings such as desks. architecturally, the space itself may also communicate information, which is the case, for instance, when signs in a hallway are not needed to locate the corridor to the restrooms. on the other hand, a lack of needed information or tools provided by the physical environment can reduce one’s prerequisites for performing various tasks, either practical or intellectual, in the space. not only does this happen factually, but this may also challenge the experiences of one’s own ability and autonomy, at worst bringing about a sense of inability due to a dysfunctional environment. in this sense, properties of the physical surroundings become incorporated as capabilities of the individual. the information embedded in the physical environment may also reduce the need to seek instructions for tasks on the very basic level of functioning, such as finding the appropriate equipment to perform a given task. in contrast, if a student is not capable of navigating independently in the space without constantly asking for information on the most basic level, this can be harmful not only for the process of learning but also for the sense of autonomy experienced by the student. in fact, the guidance provided by the physical space itself may be seen as more supportive of the autonomy of the students as they take on a more active role when searching for the needed information from the physical environment, as opposed to being socially given the information that the teacher assumes that they need. they are “the origin or source for one’s own behavior” (deci & ryan, 1985; ryan & deci, 2002), and the more they can autonomously direct their study-related behavior in meaningful ways, the more they themselves are in control of the learning process. for these purposes, the physical environment may provide not only information and guidance, but also tools for searching for the requisite information. for example, the students described their frequent use of modern technology, such as smartphones, tablets and laptops, in searching for relevant information. the use of modern technologies was experienced as handy and quick in comparison to searching for the information from the library. it also seemed that the use of modern technology was at times more supportive of the students’ sense of autonomy as it reduced the need to lean on the teacher as a source of information within the laboratory space. however, some students found that the physical space did not accommodate the use of modern tools as well as they would have hoped. the students reported that workspaces crowded with chemical equipment often did not leave room for laptops even though they would have been an important part of the study process. while independent search for information requires self-directedness, it also changes some of the social aspects of having to ask for additional information. instead of presenting his or her imperfections, the student can independently approach the question and, optimally, succeed in solving it. at best, this may foster the student’s sense of autonomy. in addition, providing information in excess through various physical modalities is hardly a risk, whereas with socially mediated guidance this can often be a challenge: sjöblom  et  al       | f l r     26   at times the assistant may come and do the thing for you, and it would be nicer to get to do it yourself, just to take the instructions and try to get something out of it. sometimes when you’ve wanted help, verbally or such, then the assistant has come and put together that instrument there and taken care of it. the independent work to me too is great, really... at least for me, even though group work is okay and nice but if the other person gets things faster and better, then i’m just like, the other person says well go find this and i do and i’m getting nothing about anything, --so then you have to take responsibility for your own work and understanding too. indeed, in light of this data, the role of socially mediated guidance in learning was as underlined as it was dilemmatic. by socially mediated guidance we refer to the support that the student receives in the learning process either from teachers or from fellow students. while the students appreciated the space and freedom to process things themselves and be independently responsible for their progress in the chemistry tasks, they felt a strong need for reassurance that they are progressing in the right direction. many students emphasized the importance of receiving social confirmation and affirmation for their assumptions either from their peers or from the teacher. 5.1.2 the volitional nature of the study activities finally, when asked about the meaning of the physical environment in their studies, throughout the data many of the students mentioned how being able to practice in the actual laboratory brought a sense of meaning and purpose to their studies. the activities performed in the laboratory demonstrated why they were there in the first place, what they would be doing in the future, and why they should proceed and advance in their studies: here the students are doing their work and the assistants are only there to see that nothing particular is happening. in the future if you’re working in the laboratory, there will probably be no one telling you to “do this, do this”. instead, you have to use your own head when you’re working there, and here you get to practice that. for me too, with this instrumentation that i’ve never got to use before, it is a fine feeling of ‘hey this is how it works’; there are levers and tubes and glass and all kinds of things gathered there. it is awfully great to get to use things that you never have before. and overall, the engagement of the laboratory work, that feeling when you’ve actually succeeded, you have that aspirin weighed, measured, everything checked – that feeling: yes i’ve accomplished something today! even though it’s nothing bigger than some ten grams of aspirin, still. as was clearly manifested in the students’ reports, the physical spaces and tools enabled study processes that were highly valued by the students and appeared to strengthen not only the sense of autonomy but also competence and relatedness to their professional community. 5.2 competence 5.2.1 the importance of practical conditions on intellectual and emotional functioning: ergonomics, usability and the fluency of the activity in the physical environment in a study context the need to be able to perform and accomplish tasks is accentuated. a predominant feature of a chemistry laboratory as a study context is that it involves concrete activities with physical equipment and tools. when asked about helpful and challenging aspects of the physical learning environment, the students brought up the importance of ergonomics in the laboratory settings. they sjöblom  et  al       | f l r     27   mentioned how their work can be significantly disturbed by challenging external conditions, for example, when they have to work in unergonomic positions. this was evident in the experience of a student who reported having at times to do his laboratory work “in a highly confined space in a fetus-like position.” with that example in mind, one may recognize how the physical environment may have a fundamental effect in hindering or disturbing the student in applying his competence to the task at hand.  the questions of usability were equally important regarding modern tools such as technological devices and software. if the prerequisites for accomplishing a task are not taken care of and the environment does not provide the needed procedural facilitation, the student cannot experience himor herself as competent in the given physical environment. in consequence, these kinds of external factors may lower the internal sense of competence; the functionality of the physical environment may not only have an enabling role with regard to the concrete activity, but there is an essentially emotional component to this as well. equipment and tools, traditional or modern, a chair or a smartphone, may hinder one’s experienced competence but also elevate it and take it to the next level. 5.2.2 the physical environment and tools: tangible indications of competence and sources of engagement on the other hand, the students frequently brought up that proper and well-functioning practical tools offered them a concrete indication of competence and accomplishment as well as a source of engagement in the learning process: i do like it that with the kind of proper practical tools one can practice making real things, that it’s not just all on the pages of the books, that it motivates and in my opinion grows that confidence, hey i could do this, hey this resulted in such a good yield. for me, that inspires me to go forward in the studies. it appeared that at best, the equipment offered stimulus for a positive, reinforcing cycle when the student was able to master, put together and utilize equipment initially experienced as strange and intimidating due to its complexity and sophistication: to me, successful reactions or syntheses help me greatly [to engage in learning]. and special and new equipment too, that you get to familiarize yourself a little with, you wonder what to do with them and they look completely strange, and you have absolutely no clue what to do with them. and then someone clears that up for you and you’re like “aah okay!” it is so nice! … especially when putting together the distillation apparatus for the first time, it was absolutely horrible, and such an awful chaos! but now that you’ve done a fair amount of that, you’re, well… it is wonderful to notice that it doesn’t take 15 minutes of agonizing anymore, you take the right instruments almost automatically. in these cases the elements of the physical environment that used to communicate strangeness became familiar and meaningful. instead of communicating difficulty and incapability, they offered a sense of mastery as well as an indication of progress in the learning process. more broadly, the mere observation that with time and practice the student could navigate and function in an environment that in the beginning had been fairly demanding may also be seen as a positive indicator of the learning process and the development of competence in the context. this kind of feedback on one’s abilities, stemming from the mundane concrete doings in the laboratory and involving both the cognitive and the emotional dimension, may be seen to be functional in nature as it is not given by someone else but emerges through the experience of success in a practical task. 5.2.3 the challenges of competent functioning in the complex physical environment: providing cognitive structuring and procedural facilitation in the space itself in order to successfully function in the laboratory environment, the students need to not only have the appropriate theoretical grounding and understanding of the phenomena, but they also need to familiarize sjöblom  et  al       | f l r     28   themselves with the social practices of applying the information in practice in a given field. this is not straightforward as the shared practices are often in the form of an expert’s silent information, which can be best assimilated by participating in the actual procedures and operations, or by becoming part of the professional community. this, however, may be challenging as the time spent in the actual laboratory setting is limited. many students reported feeling that they were expected to be more competent in laboratory work than their actual level of competence was. the laboratory environment was highly complex and demanding for them to begin with. for example, the students mentioned that watching a security video once does not necessarily mean that they have assimilated the information and would be able to take the crucial points into account when working in the laboratory setting. this, for many students, resulted in recurrent uncertainty and pondering over safety issues: in theory you do know these things since you’ve studied the course on safe work in the laboratory, but then when you come to strange circumstances like these, it may happen that that part of your brain is not working, and you’re like, there’s all the rest of the hustle and bustle and the poisons there. in terms of the theory of flow (csíkszentmihályi, 1988; see also inkinen et al., 2013), if the challenges of the task are considerably higher than the students’ abilities to respond to them, the students are at risk of experiencing predominantly worry and anxiety, which does not facilitate their learning or wellbeing. if students are frequently experiencing failure and inability with regard to the expectations rather than meeting the expectations and noticing progress in their learning, the students’ sense of competence in the given physical environment may be hindered. the more complex the activity and the environment, the more cognitive structuring is needed. as mentioned earlier in the text, by physical means this scaffolding can be provided e.g. by adding tags, signs and information boards as well as paying attention to the overall clarity of the physical environment. 5.3 relatedness within research on learning, relatedness has mainly been studied in relation to a given social community, such as a professional community, instructors or peer students. although feelings of relatedness may not be connected to the mere physical surroundings, we considered it important to study the role of the physical environment from the viewpoint of an emerging sense of relatedness to a professional community. more specifically, based on the analysis, it appeared that the students referred to the role of the physical environment as part of their experiences of belonging to given physical premises or the lack of belonging. therefore, in the following we will also use the notion of belonging when approaching questions of relatedness with regard to the physical environment. while the dimensions of autonomy and competence were particularly central in the interview data, experiences of relatedness were less prevalent, which seems to be an important observation since within the field of chemistry there seems to be a challenge with regard to students' commitment to their studies and to the professional field. as the students described their relation to the physical space, two central themes became relevant. firstly, the students perceived different kinds of study activities to belong to different physical surroundings and associated a certain value to them. secondly, as elaborated earlier, how the physical environment accommodates the activities expected to be performed in it has importance with regard to the emerging sense of competence. consequently, the study activity, be it fluent or laborious, contributes to how easy or difficult it is for the students to proceed and succeed in their study tasks and influences how the students view themselves when working in that particular environment. in a broader sense, this bears relevance to their developing sense of relatedness to the professional field. in the field of chemistry, the laboratory surroundings are an especially central if not inseparable feature of the work itself, and therefore chemistry sjöblom  et  al       | f l r     29   offers an intriguing terrain for studying the role of the physical environment with regard to a broader formation of relatedness. in the following we will discuss each of these points in more detail. 5.3.1 from hallways to lecture rooms: spaces of status, ownership and functionality with regard to physical space and the sense of relatedness, for the students it was important to have certain physical spaces as anchors for their activities so that they could repeatedly utilize certain spaces instead of floating around without a “home” for their activities. when asked about their preferred study environments, the students seemed to experience most ownership and belonging with regard to spaces where the activity is not instructed but rather informal, such as the tables and chairs in the hallways, libraries, the student union room and, obviously, home, that is, spaces which the students were able to enter and use on their own and where the role of teacher was not as predominant. in fact, the spaces in which the students seemed to experience belonging often were also such that supported the students’ sense of autonomy, both with regard to being able to choose and enter the space rather freely and self-directedly, as well as to the nature of the activity taking place in the space. indeed, just the very fact of being able to choose between differentiable and flexible spaces in order to best accommodate the given study activity may be seen as supporting students’ sense of autonomy and their active role in guiding their own learning process. while most of the spaces utilized by the students were not officially designated for any specific task, the students nevertheless seemed to have a clear vision regarding which spaces they would use for which study activity. for informal tasks, such as group work, the students reported choosing mainly informal environments, such as university hallways, cafeterias or public transportation. the faculty library or classrooms, instead, were perceived as natural venues for pursuing more serious and ambitious studying. further, the students associated a certain value with certain study environments. some students seemed to regard as “proper learning” those study activities that were situated in formal learning environments. for pursuing “serious study activities,” students reported choosing formal study environments, such as the faculty library. the laboratory environment, clearly being a formal learning environment, was regarded as an environment where serious study activity and “proper learning” takes place. in contrast, the activities conducted in informal environments, such as group assignments, were not described as worthy and official, albeit that these study activities may be highly essential in the process of learning. in fact, based on the students’ reports, collaborative study activities were not recognized as learning as clearly as individual work, either when instructed by a teacher or accomplished alone. specifically with regard to the laboratory environment and the related sense of belonging, some of the students indicated that in their experiences the laboratory space is not a space that belongs to them in the first place. rather, many students perceived themselves as visitors in this space that is occupied by others, such as teachers, more advanced students and the researchers who are its main users. 5.3.2 welcoming, functional and dysfunctional spaces: allowing users to be human the issue of belonging may also be seen as related to how the space communicates with work and tasks. thus questions of usability become relevant: the dysfunctionality or impracticality of the environment does not support experiences of one or one’s work belonging in the given space. the space can be seen as inviting or welcoming in relation to the individual’s own functioning; for instance, how the space is designed to meet the ergonomic needs of users builds experiences of fluency vs. laboriousness: maybe the most important thing in interior design would be functionality, as you have people of different sizes, the adjustability of the surroundings, so that the work would be ergonomic. if you have to reach something from high above, that you would have some tool or a strategy, whatever it is, so that you can reach things from above safely. at times when you are taking those poisons from somewhere terribly high, me too, a small person, it is a bit like, will it come down and will my hand slip… sjöblom  et  al       | f l r     30   if the environment is predominantly uncomfortable and performing tasks in it is cumbersome, this does not enhance the experience of being capable or, more broadly, belonging to function in that space: sometimes you kind of know what you’re doing or what you’d like to do, but somehow you can’t as the instrument…or the practicalities don’t always work. there is no space or there are too many flies in the ointment to be able to do a simple thing. if you’re working in a fume cupboard with acid solution, you need ph paper, if i do that i first pass three chairs, three buddies, i only get to the hallway there. then i walk past devices where there are possibly people working so i have to dodge them too, and then i get to the assistants’ room where there are three other people asking them something. i stretch there and i take the ph paper... at worst there are so many things in the way, to sum it up, there are many switchbacks there. how the space communicates with the student’s needs or expectations may also stem from the way the student is able and allowed to individually customize the space and the facilities according to his or her own preferences, thus bringing about a personal touch with regard to the given physical surroundings. for instance, the student should be able to adjust the equipment to meet his or her ergonomic needs or to customize the environment to adapt to personal work habits as opposed to being forced to work in a space occupied by another person who has completely opposite habits. for example, the students had varying preferences regarding the need for clarity versus stimuli from the proximal surroundings and differed as to what point they started to feel the need to clear the space or wash the glassware. whereas some students wished to have all their equipment immediately available and within reach, other students experienced this kind of abundance as overwhelming and chaotic, disturbing both their cognitive processing and conduction of practical tasks. specifically related to chemistry laboratory work, an important issue is also how the environment allows the students to be human. that is to say, at the beginning of studies it is natural to make mistakes and break glassware or other equipment by accident. the students described the importance of the policy in the faculty regulations on whether the students need to pay none of the expenses, part of them or all of them, as this influences their confidence to practice the work that they do not yet master. in a broader sense, these kinds of background factors may also have an impact on the students’ perception of how effortless it is to be working in the space and whether it is meant for their work and incompleteness in the first place. 5.3.3 the challenges of forming a relationship with a space and place: esthetics and uninviting spaces in addition to the various subtler indications of the students’ experienced relatedness either to the physical space that they inhabit, their peers or the field in general, the data also included indications of spaces experienced as actually uninviting. some students described experiences of unpleasantness or repulsiveness, such as a space being esthetically so unsightly that it may actually have an alienating influence on the user: one student described as a freshman coming to the study premises full of enthusiasm, but considered changing the major because of the highly uninviting physical surroundings. in this case the student was never in close enough proximity to form a personal relationship with the physical study environment, as this was actually prevented by the strong initial sensation of the facilities as non-welcoming and uninviting. thus the comfort, coziness and even the materials of the physical space are not irrelevant in the process of forming a relationship to the space and place. as another example, many students mentioned the relevance of the colors in the physical environment. they were hoping for fresh, calming colors, as opposed to mirthless or exceedingly bright colors that were felt to be jarring and almost obtrusive in the study environment. while the possible lack of esthetic beauty or the experience of distaste may not, as such, prevent the experience of belonging to the physical surroundings within the study environment, it certainly does not improve the situation. the aforementioned aspects of pleasantness may be seen to point to matters that might, for their part, create beneficial circumstances for the experience of belonging to emerge. sjöblom  et  al       | f l r     31   5.4 conclusions on the intertwinedness of the basic psychological needs within the context of chemistry studies: the physical environment as a gateway to a professional community and practices above we have considered the needs for autonomy, competence and relatedness as separate dimensions. these dimensions, however, are not detached from each other; we have held to this division for analytic purposes. rather, as is implicit in the analysis above, the dimensions of autonomy, relatedness and competence are essentially intertwined. in the following, we will specifically explicate this intertwinedness in the studied chemistry context. within the light of the basic psychological needs, what at first came across in the students’ reports was their relation to autonomy. namely, they appeared to emphasize the need for self-directedness already in their first year of studies. this may derive from the fact that the laboratory as a space offered them a direct connection to their possible future job in the laboratory, and thus they were constantly mirroring their everyday laboratory chores to the expectations of the profession: an independent role in a laboratory, possibly working alone or as the only chemist on the premises. with this vision in their minds, they desired to form a similarly self-driven and independent work ethos already at the beginning of their studies. as the profession of a chemist can be seen not only as an academic profession but also as handicraftmanship, the relation between the future profession and the novice stage courses is much closer than in many other academic fields in which the first years of studies are often mainly filled with theoretical courses. the laboratory environment represents a physical professional environment that the student is able to enter at an early stage of studies and, with practice, to increasingly master. indeed, the students often seemed to experience that the work in the laboratory bridged the gap between the rookie and professional stages: by accomplishing their concrete study tasks in the laboratory, they were doing similar tasks as professionals, which served as a gateway to the professional practices of chemists. this advance in study practices can also be seen as progress in terms of legitimate peripheral participation (lave & wenger, 1991); as the students are admitted to participate in procedures in a given professional context, they become involved in the professional community and culture and its shared social practices and are able to proceed from the fringe areas of professional abilities towards more internalized and well-established professional practices and expertise. from the viewpoint of basic psychological needs, an environment that supports feelings of efficacy as well as a connection with those who convey it is most likely to promote internal motivation (ryan, 2009). as the students experience the laboratory environment as closely representing their future workplace and mirror their actions to their future role as a professional, it is particularly important to pay attention to how the initial experiences of working as a chemist in a laboratory setting are built. here the design of a functional and pedagogically purposeful environment becomes central. to conclude, based on the results, we suggest that the experience of a given physical space builds through the activities performed in that space. the functionality and usability of the space and tools are highly important as they contribute to the fluency of the activity taking place, which builds the students’ view of themselves acting in that given environment. when a student experiences the space as a place that involves equipment and functions that he or she can master and perceives himor herself as someone successfully functioning in that environment, he or she is more likely to experience belonging to that environment and context. thus, how the physical environment manages to accommodate the most mundane everyday activities may, for the users, build on a broader experience of relatedness. this may be of importance when building a professional identity and creating a sense of belonging to a professional community. thus, by providing sufficient or even optimal premises for study activities, the physical environment may facilitate this process to varying degrees. sjöblom  et  al       | f l r     32   conclusion examples from the data practical implications physically mediated guidance and the use of modern technological devices may support students’ sense of autonomy and competence. students were hoping for clear, well-structured spaces, where the basic-level information may be implemented in the space, or the students can acquire it with the help of technological devices, in order to enable them to navigate and function in the space in a selfdirected manner. socially mediated guidance was regarded as important in confirming one’s assumptions, in a facilitating rather that instructing manner. it is important to distinguish between physically and socially mediated guidance and their purposeful roles. physically mediated guidance should be more widely acknowledged and utilized in communicating information on a basic level, such as where to find needed equipment or dispose of substances, whereas social guidance is needed in the more complex cognitive processing. the physical environment may complement the students’ existing competence and offer procedural facilitation for their learning processes. the chemistry laboratory as a new and complex working environment seemed to be highly challenging, if not intimidating for the students at first. however, if the students were able to successfully enter and learn to master the equipment and the space, it offered them fruitful and highly engaging learning experiences. students should be provided with suitable spaces and tools as well as sufficient guidance in using them in order to ensure the scaffolding of the learning processes by both physical and social means. the more complex the activity and the environment, the more cognitive structuring is needed. being able to utilize diverse learning environments in a selfdirected manner may support students’ sense of autonomy in directing and regulating their own learning process. the students associated certain study activities as well as a certain value, status and ownership to different learning environments. formal learning environments, such as lecture halls, libraries and laboratories, as well as the formal and focused learning activities occurring in them, were often perceived as more “proper” than the informal and collaborative learning environments and activities, even though the latter were experienced as crucial in the learning process. flexible, diverse and freely accessible spaces should be available for students in order to accommodate the variety of study activities as well as support students’ sense of autonomy and relatedness. informal environments may promote more sense of belonging and ownership in novice students; the possibility to act in a professional work environment may bridge the gap between the rookie and professional stages and also bring a sense of meaning and purpose to the studies. the functionality of the physical environment contributes to the cognitive processes of the users as well as to the related emotional experience of oneself acting in the given environment. consequently, the physical environment may be instrumental in the development of the students’ sense of relatedness to the professional community. for the students the laboratory strongly represented their future work environment as chemists, and the experiences occurring in it were frequently mirrored to their future professional identity. the functionality of the physical environment and the fluency of the activity appeared to contribute to students’ sense of belonging to the professional context. special attention should be paid to the functionality of the physical environment as well as the fluency of short periods of practical work, as the experience of a physical environment builds through the activity performed in the environment. table 1 summary of results   sjöblom  et  al       | f l r     33   6. discussion in this study we analyzed the role of the physical environment in learning and well-being from the viewpoint of self-determination theory and basic psychological needs. the physical environment may support not only learning and well-being, but also autonomy, competence and relatedness with regard to the learning environment and the professional field. in the following we will elaborate on the broader theoretical and practical implications of the results. the physical space and tools can be seen as facilitating or posing a challenge to study activities and cognitive functioning by various means. the physical environment not only influences the cognitive learning process but inevitably gives rise to an emotional experience, as well. for instance, if the physical environment poses a challenge to study activities, and because of this the students constantly feel incompetent in the learning context, this experience builds on their views of themselves acting in that particular environment, and consequently, they may be less likely to frequently and willingly approach the same environment in the future. moreover, in order to reduce unnecessary anxiety over tough challenges with regard to their existing abilities, as well as to provide optimal grounds for learning to occur, it would be important to complement the students’ existing competence by offering procedural facilitation and support in both the physical and social environment. the emotional experience resulting from the concrete activities taking place in the physical space can support committing to that particular working environment, as well as the broader context related to it, such as the professional community of chemists. recent pedagogical research has emphasized the emotional components of the learning process, such as interest and engagement (see e.g. csíkszentmihályi, 2014; heikkilä, niemivirta, nieminen & lonka, 2011; hidi & renninger, 2006; inkinen et al., 2013; lonka 2012; lonka & ketonen 2012), as opposed to more traditional views concerning merely cognitive aspects of learning. furthermore, engagement in learning has been approached through conceptualizing cyclical stages in the learning process and defining optimal practices. we want to shed light on the role of the physical environment in the learning process: how the physical environment may support or hinder learning practices, and how that, in turn, contributes to the emotional experience and sense of commitment or the lack of it. this broadened viewpoint involving the role of the physical environment in learning may be utilized in envisioning a more holistic approach to engaging learning. the functionality and usability of the space and the equipment, the guidance implemented in the space as well as other support available (peers, teacher) all play key roles in the learning process. from the viewpoint of self-determination theory, physical environment represents a novel context for the application of the theory. based on this study, similarly to social and cultural environment, physical environment can also support or thwart the fulfillment of the basic psychological needs. furthermore, this study raises theoretical questions concerning the role of the three basic psychological needs as well as their interrelations in different contexts. while the fulfillment of all three needs is essential, within the light of these data it strongly appeared that in a study context, perhaps similar to other contexts that are highly demanding in relation to existing abilities, the dimension of competence seemed to be very central, if not a prerequisite, for experiences of autonomy or belonging to emerge. for example, it is challenging for students to develop a sense of belonging to the professional community if they mostly feel incapable of performing basic tasks and thus find themselves incompetent in the field in general. while the developers of the theory strongly emphasize the importance of all three needs as well as the synergy between them, depending on the nature of the activity, relatedness, for instance, may at times be less central to intrinsic motivation than autonomy and competence (deci & ryan, 2000). on the other hand, in other occasions, such as with children or adolescents who are at risk of dropping out of school, it may be most crucial to support the experience of relatedness (e. deci, personal communication with the first author, october 28, 2014). furthermore, it has been acknowledged that the dimensions of competence, autonomy and relatedness are strongly interrelated, and for instance, an autonomy-supporting atmosphere will assist in promoting relatedness and competence as well (deci & ryan, 1987; wolters & gonzalez, 2008). acknowledging these previously researched viewpoints, we wish to both emphasize the importance of sjöblom  et  al       | f l r     34   promoting the fulfillment of all three basic needs in the learning context as well as further examine their interrelations and prerequisites. in our view the three dimensions may not in all contexts be equally interrelated and in identical interaction with each other. instead, they may be interdependent or sequential depending on the context. this context-driven analysis of the underlying dynamics may be an intriguing terrain for further research on the theory. what is of particular interest in the field of higher education is how the basic psychological needs interact with vital study-related phenomena such as the commitment to studies and the development of professional identity, and how to best take this into account when designing learning processes. this study demonstrates the importance of the physical environment for intellectual as well as emotional functioning. the intellectual functioning of an individual is always nested in a given physical environment, even when the work is carried out in a virtual environment. in fact, it may be that the impact of the physical environment on psychological functioning is often highly underestimated. with regard to future research it would be intriguing to untangle the effects that different physical space solutions have on human functioning. it is likely to make a difference whether one is working in a familiar workspace or in increasingly common open-plan multispace offices, not only with regard to ergonomics but also with regard to experiences of belonging or recovery. for instance, experiences of ownership and relatedness or beneficial, uplifting and inspiring mental modes can be supported by various means in both stable and mobile offices. in addition to the focal social aspects such as the shared culture of the community, some physically mediated options might include customizing the physical space with personal items but also utilizing modern and mobile technological means, such as customized technological tools or screen savers. moreover, the bodily dimensions of office environments beyond ergonomics offer an intriguing aspect to the psychophysical experience. for instance, the possibilities that the spaces or furniture offer for varied bodily postures and physical movement all contribute not just to physical health but also to the psychological experience and functioning. to conclude, it is essential to utilize psychological and pedagogical knowledge when designing work and learning environments. by considering the interplay between the material world and human functioning, we can create fruitful ground for thriving users and develop novel design for leading university campuses and other indoor environments. keypoints similarly to social and cultural environment, physical environment can also support or thwart the fulfillment of the basic psychological needs. learning and wellbeing can be facilitated by developing physical environments that support the basic psychological needs. the physical environment contributes to the cognitive functioning of the users as well as to the related emotional experience of oneself acting in the given environment. for example, a wellstructured physical environment may offer physically mediated guidance, cognitive structuring and procedural facilitation for the students’ learning processes. it may complement the students’ existing competence and scaffold the students’ sense of control in situations where the challenge of the task is experienced as high. physical spaces and tools should be utilized in offering students functional feedback, engaging learning experiences and gateways to practicing their future profession. in order to support the basic psychological needs as well as help the students to regulate their own learning process the students should be provided with suitable spaces and tools as well as sufficient guidance and autonomy in using them. special attention should be paid to the functionality of the physical environment, as the experience of a physical environment builds through the activity performed in the environment. sjöblom  et  al       | f l r     35   the results provide both theoretical and practical value in understanding the role of the physical environment as part of human functioning and serve as an opening to a previously unexplored ground. by bringing together the theoretical approaches of socially and physically distributed intelligence and research on motivation, this study demonstrates the importance of the physical environment for intellectual as well as emotional functioning. the intellectual functioning of an individual is always nested in a given physical environment, even when the work is carried out in a virtual environment. utilizing psychological and pedagogical knowledge is essential when designing or renovating work and learning environments in order to fully make use of the potential of physical environments as part of human performance. acknowledgements this study was funded by the tekes (the finnish funding agency for technology and innovation) rym indoor environment project (project number 462054), the academy of finland project mind the gap (project number 1265528) as well as personal grants from finnish cultural foundation (1st and 3rd autor) and alfred kordelin foundation (2nd author). references alfonsi, e, capolongo, s. & buffoli, m. (2014). evidence based design and healthcare: an unconventional approach to hospital design. annali di igiene : medicina preventiva e di comunità, 26, 137–143. doi:10.7416/ai.2014.1968 alker, j., malanca, m., pottage, c., o’brien, r., akhras, d., ambrose, b., …wong, j. (2015). health, wellbeing and productivity in offices. world green building council, http://www.worldgbc.org/activities/health-wellbeing-productivity-offices/research. baumeister, r. f. & leary, m. r. (2000). the need to belong: desire for interpersonal attachments as a fundamental human motivation. in higgins, e. t. & kruglanski, a. w. (eds.), motivational science: social and personality perspectives (pp. 24–49). new york, ny: psychology press. beard, c. (2009). space to learn: the development and evolution of new learning environments in higher education. in buswell, j. & becket, n. (eds.), enhancing student centred learning in business and management, hospitality, leisure, sport, and tourism. oxford, uk: threshold press. beard, c. (2012). spatial ecologies: learning and working environments that change people and organisations. in alexandra, k. & price, i. (eds.), managing organisational ecologies (pp.69–80). new york, ny: routledge. beard, c. & price, i. (2010). space, conversations and place: lessons and questions from organisational development. international journal of facility management, 1. bereiter, c. & scardamalia, m. (1987). the psychology of written composition. hillsdale, nj: lawrence erlbaum associates inc. black, a. e. & deci, e. l. (2000). the effects of instructors’ autonomy support and students’ autonomous motivation on learning organic chemistry: a self-determination theory perspective. science education, 84, 740–756. doi:10.1002/1098-237x(200011)84:6<740::aid-sce4>3.0.co;2-3 chirkov, v., ryan, r. m., kim, y., & kaplan, u. (2003). differentiating autonomy from individualism and independence: a self-determination theory perspective on internalization of cultural orientations and well-being. journal of personality and social psychology, 84, 97–110. doi:10.1037/0022-3514.84.1.97 csíkszentmihályi, m. (1988). optimal experience: psychological studies of flow in consciousness. new york, ny: cambridge university press. csíkszentmihályi, m. (2014). applications of flow in human development and education: the collected works of mihály csíkszentmihályi. new york, ny: springer science + business media. sjöblom  et  al       | f l r     36   deci, e. l. & ryan, r. m. (1985). intrinsic motivation and self-determination in human behavior. new york, ny: plenum. deci, e. l. & ryan, r. m. (1987). the support of autonomy and the control of behavior. journal of personality and social psychology, 53, 1024–1037. deci, e. l. & ryan, r. m. (2000). the “what” and “why” of goal pursuits: human needs and the selfdetermination of behavior. psychological inquiry, 11, 227–268. doi:10.1207/s15327965pli1104_01 deci, e. l. & ryan, r. m. (2002). self-determination research: reflections and future directions. in deci, e. l. & ryan, r. m. (eds.), handbook of self-determination research (pp. 431–442) rochester, ny: the university of rochester press. deci, e. l. & ryan, r. m. (2008). self-determination theory: a macrotheory of human motivation, development, and health. canadian psychology, 49, 182–185. doi:10.1037/a0012801 dweck, c. s. (2006). mindset: the new psychology of success. new york, ny: random house. evans, g. w., bullinger, m. & hygge, s. (1998). chronic noise exposure and physiological response: a prospective study of children living under environmental stress. psychological science, 9, 75–77. doi:10.1111/1467-9280.00014 gay, j. l. (2008). testing self-determination theory and the roles of the social and physical environments in an adult beginning exerciser population. columbia, sc: university of south carolina. gay, j. l., saunders, r. p. & dowda, m. (2011). the relationship of physical activity and the built environment within the context of self-determination theory. annals of behavioral medicine, 42, 188– 196. doi:10.1007/s12160-011-9292-y hakkarainen, k. (2009). a knowledge-practice perspective on technology-mediated learning. computersupported collaborative learning, 4, 213–231. doi:10.1007/s11412-009-9064-x hakkarainen, k., palonen, t., paavola, s. & lehtinen, e. (2004). communities of networked expertise: professional and educational perspectives. bingley, uk: emerald group publishing limited. heikkilä, a., & lonka, k. (2006). studying in higher education: students’ approaches to learning, selfregulation, and cognitive strategies. studies in higher education, 31, 99–117. doi:10.1080/03075070500392433 heikkilä, a., lonka, k., nieminen, j. & niemivirta, m. (2012). relations between teacher students' approaches to learning, cognitive and attributional strategies, well-being, and study success. higher education, 64, 455–471. doi:10.1007/s10734-012-9504-9 heikkilä, a., niemivirta, m., nieminen, j. & lonka, k. (2011). interrelations among university students’ approaches to learning, regulation of learning, and cognitive and attributional strategies: a person oriented approach. higher education, 61, 513–529. doi:10.1007/s10734-010-9346-2 hidi, s., & renninger, k.a. (2006). the four-phase model of interest development. educational psychologist, 41, 111–127. doi:10.1207/s15326985ep4102_4 hutchins, e. (2000). distributed cognition. international encyclopedia of the social and behavioral sciences, 2068–2072. hutchins, e. (2006). the distributed cognition perspective on human interaction. in enfield, n. j. & levinson, s. c. (eds.), roots of human sociality: culture, cognition and interaction (pp. 375–398). inkinen, m., lonka, k., hakkarainen, k., muukkonen, h., litmanen, t. & salmela-aro, k. (2013). the interface between core affects and the challenge-skill relationship. journal of happiness studies, 15, 891–913. doi:10.1007/s10902-013-9455-6 iyengar, s. s., & devoe, s. e. (2003). rethinking the value of choice: considering cultural mediators of intrinsic motivation. in murphy-berman, v. & berman, j. j. (eds.), cross-cultural differences in perspectives on self (pp. 146–191). lincoln, ne: university of nebraska press. jang, h., reeve, j. & deci, e. l. (2010). engaging students in learning activities: it is not autonomy support or structure but autonomy support and structure. journal of educational psychology, 102, 588–600. doi:10.1037/a0019682 job, v., walton, g. m., bernecker, k. & dweck, c. s. (2015). implicit theories about willpower predict selfregulation and grades in everyday life. journal of personality and social psychology, 108, 637–647. doi:10.1037/pspp0000014 sjöblom  et  al       | f l r     37   küller, r. & lindsten, c. (1992). health and behavior of children in classrooms with and without windows. journal of environmental psychology, 12, 305–317. doi:10.1016/s0272-4944(05)80079-9 lansdale, m., parkin, j., austin, s. & baguley, t. (2011). designing for interaction in research environments: a case study. journal of environmental psychology, 31, 407–420. doi:10.1016/j.jenvp.2011.05.006 lave, j. & wenger, e. (1991). situated learning: legitimate peripheral participation. new york, ny: cambridge university press. lindblom-ylänne, s. & lonka, k. (2000). interaction between learning environment and expert learning. lifelong learning in europe, 5, 90–97. lonka, k. (2012). engaging learning environments for the future: the 2012 elizabeth w. stone lecture. in gwyer, r., stubbings, r. &walton, g. (eds.), the road to information literacy: librarians as facilitators of learning (pp. 15–30). berlin, germany: de gruyter. lonka, k. & ketonen, e. (2012). how to make a lecture course an engaging learning experience? studies for the learning society, 2, 63–74. doi:10.2478/v10240-012-0006-1 markus, h. r., & kitayama, s. (1991). culture and the self: implications for cognition, emotion, and motivation. psychological review, 92, 224–253. doi:10.1037/0033-295x.98.2.224 mälkki, k. (2010). building on mezirow’s theory of transformative learning: theorizing the challenges to reflection. journal of transformative education, 8, 42–62. doi:10.1177/1541344611403315 mälkki, k. (2012). rethinking disorienting dilemmas within real-life crises: the role of reflection in negotiating emotionally chaotic experiences. adult education quarterly, 62, 207–229. doi:10.1177/0741713611402047 mälkki, k., sjöblom, k. & lonka, k. (2014). transformation of the physical space and transformation of the subject. in nicolaides, a. & holt, d. (eds.), spaces of transformation and transformation of space: proceedings of the xi international transformative learning conference (pp. 550–556). new york, ny: teachers college, columbia university. nenonen, s., kärnä, s., junnonen, j.-m., tähtinen, s. & sandström, n. (eds.) (2015). how to co-create campus. tampere, finland: suomen yliopistokiinteistöt oy, tampere juvenes print. (in finnish) niemiec, c. p. & ryan, r. m. (2009). autonomy, competence, and relatedness in the classroom applying self-determination theory to educational practice. theory and research in education, 7, 133–144. doi:10.1177/1477878509104318 norman, d. a. (1993). things that make us smart: defending human attributes in the age of the machine. cambridge, ma: perseus books. paavola, s., lipponen, l. & hakkarainen, k. (2004). models of innovative knowledge communities and three metaphors of learning. review of educational research, 74, 557–576. doi:10.3102/00346543074004557 parsons, r. & hartig, t. (2000). environmental psychophysiology. in cacioppo, j. t., tassinary, l. g. & berntson, g. (eds.), handbook of psychophysiology (pp. 815–846). new york, ny: cambridge university press. perry, n. e., turner, j. c. & meyer, d. k. (2006). classrooms as contexts for motivating learning. in alexander, p. a. & winne, p. h. (eds.), handbook of educational psychology. mahwah, nj: lawrence erlbaum associates publishers. rutten, c., boen, f. & seghers, j. (2012). how school social and physical environments relate to autonomous motivation in physical education: the mediating role of need satisfaction. journal of teaching in physical education, 31: 216–230. ryan, r. m. (1995). psychological needs and the facilitation of integrative processes. journal of personality, 63, 397–427. ryan, r. m. (2009). self-determination theory and wellbeing. wellbeing in developing countries research review, 1. ryan, r. m. & connell, j. p. (1989). perceived locus of causality and internalization: examining reasons for acting in two domains. journal of personality and social psychology, 57, 749–761. doi:10.1037/00223514.57.5.749 sjöblom  et  al       | f l r     38   ryan, r. m. & deci, e. l. (2000). self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. american psychologist, 55, 68–78. doi:10.1037/0003066x.55.1.68 ryan, r. m. & deci, e. l. (2002). overview of self-determination theory: an organismic dialectical perspective. in e. l. deci & r. m. ryan (eds.), handbook of self-determination research (pp. 3–33) rochester, ny: university of rochester press. ryan, r. m. & deci, e. l. (2006). self-regulation and the problem of human autonomy: does psychology need choice, self-determination, and will? journal of personality, 74, 1557–1586. doi:10.1111/j.14676494.2006.00420.x ryan, r. m. & deci, e. l. (2009). promoting self-determined school engagement: motivation, learning and well-being. in wentzel, k. r. & wigfield, a. (eds.), handbook of motivation at school (pp. 171–195). new york, ny: routledge. ryan, r. m., huta, v. & deci, e. l. (2008). living well: a self-determination theory perspective on eudaimonia. journal of happiness studies, 9, 139–170. doi:10.1007/s10902-006-9023-4 sandström, n., eriksson, r., lonka, k. & nenonen, s. (accepted for publication 2015). usability and affordances for inquiry-based learning in a blended learning environment. facilities. sandström, n., ketonen, e. and lonka, k. (2014). the experience of laboratory learning – how do chemistry students perceive their learning environment? european journal of social and behavioural sciences, 11, 1612–1625. doi:10.15405/ejsbs.144 sandström, n., sjöblom, k., mälkki, k. & lonka, k. (2013) the role of physical, social and mental space in chemistry students’ learning. european journal of social and behavioural sciences, 6, 1134–1139. doi:10.15405/ejsbs.90 scott, d. & usher, r. (1999). researching education: data, methods and theory in educational enquiry. london, uk: continuum. seligman, m., ernst, r., gillham, j., reivich, k. & linkins, m. (2009). positive education: positive psychology and classroom interventions. oxford review of education, 35, 293–311. doi:10.1080/03054980902934563 soenens, b., sierens, e., vansteenkiste, m., goossens, l., & dochy, f. (2012). psychologically controlling teaching: examining outcomes, antecedents, and mediators. journal of educational psychology, 104, 108–120. doi:10.1037/a0025742 tuominen-soini, h., salmela-aro, k. & niemivirta, m. (2008). achievement goal orientations and subjective well-being: a person-centred analysis. learning and instruction, 18, 251–266. doi:10.1016/j.learninstruc.2007.05.003 ulrich, r. s. (1981). natural versus urban scenes: some psychophysiological effects. environment and behavior, 13, 523–556. doi:10.1177/0013916581135001 vansteenkiste, m., sierens, e., goossens, l., soenens, b., dochy, f., mouratidis, a., aelterman, n., haerens, l. & beyers, w. (2012). identifying configurations of perceived teacher autonomy support and structure: associations with self-regulated learning, motivation and problem behavior. learning and instruction, 22, 431–439. doi:10.1016/j.learninstruc.2012.04.002 vermunt, j. d. & verloop, n. (1999). congruence and friction between learning and teaching. learning and instruction, 9, 257–280. doi:10.1016/s0959-4752(98)00028-0 wigfield, a., cambria, j. & eccles, j. s. (2012). motivation in education. in ryan, r. m. (ed.), the oxford handbook of human motivation. new york, ny: oxford university press. williams, m. (2000). interpretivism and generalisation. sociology, 34, 209–224. doi:10.1177/s0038038500000146 winterbottom, m. & wilkins, a. (2009). lighting and discomfort in the classroom. journal of environmental psychology, 29, 63–75. doi:10.1016/j.jenvp.2008.11.007 wolters, c. a. & gonzalez, a-l. (2008). classroom climate and motivation: a step toward integration. in maehr, m. l., karabenick, s. a. & urdan, t. c. (eds.), advances in motivation and achievement social psychological perspectives. bingley, uk: emerald group publishing limited. sjöblom  et  al       | f l r     39   woolner, p., hall, e., higgins, s., mccaughey, c., & wall, k. (2007). a sound foundation? what we know about the impact of environments on learning and the implications for building schools for the future. oxford review of education, 33, 47–70. doi:10.1080/03054980601094693 smets et struyven publication frontline learning research vol.6 no. 2 (2018) 66 80 issn 2295-3159 aligning with complexity: system-theoretical principles for research on differentiated instruction wouter smetsa, katrien struyvenb c akarel de grote university college, belgium b vrije universiteit brussel, belgium c uhasselt, belgium article received 23 november 2017 / revised 17 july/ accepted 31 august/ available online 1 october abstract differentiated instruction is a teaching philosophy and practice that deals with responding appropriately to student heterogeneity. in order to gain deep understanding of this complex concept, research methodology is challenged to use appropriate data collection and data analysis. the aim of this paper is to reflect on how system theory may be used as ontological and epistemic grounding for research on differentiated instruction. three challenges for this research are presented: to focus on the interplay between the individual and complex collective behaviour; to acknowledge the external influences in research design; and to describe patterns of non-linear causality and emergence. three design principles for research on differentiated instruction are presented to address these challenges: organic design, interaction and reflectivity. by using these principles, we believe research on differentiated instruction would be aligned with the theoretical foundations of the concept. keywords: differentiated instruction – system theory – complexity – emergence – nestedness – forest-tree perspective info. corresponding author mail: wouter.smets@kdg.be doi: https://doi.org/10.14786/flr.v6i2.340 1 the methodological need for a forest-tree perspective in an increasingly diverse world the call for teachers to provide instruction that caters for different learning needs sounds ever clearer. teachers are expected to design instruction that takes diversity among students seriously (schleicher, 2013). research over the last years has taken a lot of effort to find and study new ways of teaching that respond to diversity in heterogeneous classes (gay, 2002; ware, 2006). a central challenge for teaching in heterogeneous classes is to adapt teaching strategies which are designed, organised and assessed at the classroom level, taking into account the apparent student diversity. theoretically this is described by bronfenbrenner’s (1977) analytical stance that calls for a naturalistic perspective on psychological research. bronfenbrenner argues that, in order to understand the complexity of education, phenomena must be studied from diverse perspectives. his ecology of human development is defined as: “the scientific study of the progressive, mutual accommodation, throughout the life span, between a growing human organism and the changing immediate environments in which it lives, as this process is affected by relations obtaining within and between these immediate settings, as well as the larger social contexts, both formal and informal, in which the settings are embedded.” (p. 514). bronfenbrenner uses different systemic levels to describe interactions between individuals and their surroundings. he defines the microsystem as a “complex of relations between the developing person and […] the immediate setting containing that person” (p. 514). phenomena related to students’ individual experiences are thus described as the individual level within a learning ecosystem, whereas phenomena related to dynamics among (groups of) students are described as the micro-level of the learning ecosystem. also interactions between a teacher and students are situated at the micro-systemic level. other aspects which may influence learning (such as school culture or leadership) are described as the mesoor macro-level. scholarly research is challenged by bronfenbrenner’s approach to take different perspectives in order to study the complexity of the phenomena: the study of the phenomena with an exclusive focus on one systemic level would thus be seen as reductionist. a metaphor proposed by jacobson and kapur (2012) describes the scope of the challenge to teach in heterogeneous classes. they suggested that, in order to increase our understanding on learning environments, not solely individual phenomena or phenomena at a collective level must be studied, but rather the combination of both. metaphorically, they suggest, not solely a forest perspective must be taken, nor solely a tree perspective, but rather a forest-tree perspective. with regard to teaching in heterogeneous classes, jacobson and kapur’s (2012) metaphor stresses the vital role of distinguishing individual perspectives of particular students from the collective behaviour within the microsystem. hence, also within the microsystem, it should not be assumed that all students act and react similarly to changing conditions. the forest-perspective addresses the microsystem of a particular class. meanwhile, a focus on the learning of individual students would be seen as a tree-perspective. figure 1: the forest-tree perspective: multiple individuals within a learning ecosystem in this reflective study the concept of differentiated instruction takes a central place, as this concept tends to merge the perspective of teaching to heterogeneous groups, the micro-level of the learning ecosystem, with the perspective of the learning of heterogeneous individuals. it essentially takes heterogeneity of classes as a given fact and, therefore, assumes that each teaching process needs to be not only based on the targeted learning goals, but also on the apparent student diversity. notwithstanding the great potential of differentiated instruction for the sciences of teaching and learning, we believe that research on it is faced with important challenges. a key methodological challenge for educational science is how to design research which merges the forest and tree perspectives, to a forest-tree perspective. at present this forest-tree perspective is largely absent in research on differentiated instruction. thus, bronfenbrenner’s naturalistic approach remains a challenge for this research. much research on differentiated instruction is currently being conducted. in the next section we provide more details of the construct. two main research focuses may be discerned in research on the topic: first some studies focus on the role of teachers in a differentiated classroom. these studies, as a consequence, do not address the role of students in a differentiated classroom, nor do they focus on teacher-student interactions (de neve, devos, & tuytens, 2015). other studies on differentiated instruction focus on the learning outcomes of (groups of) students for a particular type of strategy (e.g. chen, yang, & hsiao, 2016; van klaveren, vonk, & cornelisz, 2017). we do not contest the added value of such a perspective on differentiated instruction, however it is argued in this study, that both these approaches are reductionist. the epistemic assumptions that underpin such approaches focus on parts of the teaching process instead of documenting interactions within the microsystem throughout the teaching process. it is argued in this study that the central idea of differentiated instruction lies in the responsive act(s) of a teacher which links the chosen teaching strategy to the given heterogeneity of the group. in order to grasp the full complexity of this idea, we argue for system-theoretical epistemology and hence for the use of research designs which are methodologically aligned with it. insofar as system-theoretical concepts are already in use in various fields of educational research, we discuss their usefulness ontologically, epistemologically and methodologically to underpin research on differentiated instruction. to do so, three design principles are proposed at the end of this study: organic design, interaction and reflectivity. in what follows, more details are provided on the concept of differentiated instruction in order to be able to reflect on the methodological challenges the concept poses to research. 2. differentiated instruction given the fundamental characteristics of differentiated instruction, it may be noticed that the concept entails difficulties for scholars trying to study it. differentiated instruction has been proposed as a teaching philosophy and practice that intends to maximise learning outcomes of all students in a class by responding to students’ different learning needs, namely their readiness level, interest or learning profile (tomlinson, 2000). the added value of tomlinson’s approach is that it merges a microsystem-perspective of heterogeneous classes with the perspective of individual students with particular characteristics. hence, she implicitly uses the aforementioned forest-tree perspective: teachers are supposed to teach a heterogeneous class (the forest), however their instructional design is supposed to be adaptive, and thus based on students’ individual perspectives (the tree). in doing so, differentiated instruction encompasses other strategies that focus on student heterogeneity, such as cultural responsive teaching (gay, 2002) or inclusive education (schumm & vaughn, 1995) that focus on specific types of student heterogeneity. tomlinson accepts heterogeneity as a given fact without detailing the origins of it. differentiated instruction rests on the constant responsiveness of the teaching based on these perpetual changing characteristics. tomlinson’s work at the beginning of the 21st century (tomlinson, 2000, 2001; tomlinson et al., 2003) pioneered instructional design to a broad range of student differences. a lot of enthusiasm has arisen for her practice-oriented publications. the idea is not to address any target-specific type of diversity such as learning disabilities or students at risk of academic failure, rather differentiated instruction intends to foster learning by giving the proper attention to heterogeneity in the broad sense. students’ readiness levels, interests and learning profiles are the three large, overlapping and constantly changing categories of heterogeneity used to adapt instruction. as differentiated instruction stresses the role of tailoring instruction to students’ characteristics, tomlinson’s ideas were initially picked up primarily by scholars in the activity theory tradition (shabani, khatib, & ebadi, 2010; wass & golding, 2014). yet now, the concept of differentiated instruction is also commonly applied in more cognitive scholarly approaches which focus, for instance, on the learning effect of differentiated instruction (e.g. deunk, doolaard, smale-jacobse, & bosker, 2015; prast, weijer-bergsma, kroesbergen, & luit, 2015). following tomlinson, teachers must carefully consider which teaching strategy is appropriate at which particular moment for a particular group of students. multiple teaching strategies are used, based on flexible grouping, to build learning paths respecting the unique group composition. a cyclical approach to teaching is proposed in which it is the teacher’s responsibility to engage in ongoing assessment and to adapt instructional design based on that. it is a ‘key principle that assessment and instruction are inseparable’ for differentiated teaching (tomlinson, 2000, p. 20). while practising differentiated instruction, teachers use the output of a learning sequence as the starting point of a subsequent one. dependent on students’ readiness, teachers build on prior knowledge, or on previously acquired strategies and schemes. also, teachers may respond to differences in students’ interests or learning profiles. this relationship between input and output is a fundamental characteristic of the instructional design of differentiated instruction. it marks the cyclical – responsive – character of the approach. further, we elaborate on how this cyclical character of differentiated instruction challenges research design. dweck’s (2008) growth mindset theory is often proposed as an essential characteristic that guides teaching in a differentiated classroom (coubergs, struyven, vanthournout, & engels, 2017). this theory stresses the importance of an incremental theory of intelligence or, in other words, the idea of the potential growth of students’ talents in order to maximise learning outcomes (rattan, savani, chugh, & dweck, 2015). teachers responding to the learning needs of their students will need a growth mindset in order to fulfil the learning potential of all students. teaching in a differentiated classroom seems to be closely tied to such a growth mindset, as a teacher needs to see and develop the student’s growth potential. moreover, also students’ growth mindsets need to be developed, in order to fulfil their full learning potential. in summary, differentiated instruction stands for responding to students’ heterogeneity based on a cyclical process of formative assessment and observation. the principles of growth mindset are used to guide teachers’ practice in a differentiated classroom. building on these intertwined characteristics we believe that scholarly study of the concept cannot solely rely upon classic reductionist empirical epistemology. in the following paragraph, we detail insights of systems theory, which provides a useful conceptual framework for ground research on differentiated instruction. further we elaborate how systems theory can be used to build the needed ontological and epistemic foundations to study the concept of differentiated instruction. 3 systems theory and complexity systems theory has its roots in physics and environmental sciences (von bertalanffy, 1968). for decades it has taken effort to use concepts and metaphors to describe comparable patterns across different scientific fields (luhmann, 2013 ). essential for systems theory is the notion of open systems, which stands opposed to closed systems. open systems have interactions with external surroundings. often open systems are described as complex in the sense that the properties of a system as a whole cannot necessary be derived from the properties of individual components within a system (prigogine, 1980). by conceptualising a classroom as a learning ecosystem (see §1) we are able to understand better the complexity in education (bakker & montesano montessori, 2016). the ‘open’ character of this systemic approach lies in acknowledging its interactions between the microsystem and other systemic levels. systems theory has built a reputation for helping understanding counterintuitive phenomena. its ambition is to fully acknowledge the complexity of phenomena. by doing so it stands in opposition to more reductionist scientific approaches which aim at more specific insights (sawyer, 2002). morrison (2008) stresses it is vital to see complexity theory as a collection of ideas, metaphors and concepts to describe and not to prescribe educational phenomena. thus, with this reluctance for the prescriptive ambitions, systems theory seems to stand in a postmodernist ontological and epistemological tradition. it may therefore be criticised as relativist (morrison, 2008). we agree that some system-theoretical notions (e.g. emergence or nonlinearity, see further) make insights into causal relationships hard to achieve and, as a consequence, limit the prescriptive ambitions of educational science. however, describing patterns of change may in itself be a sufficient added value to gain deeper understanding of teaching and learning in a differentiated classroom, and thus to legitimise a system-theoretical perspective. table 1 systems theory compared to modernistic approaches jacobson, kapur, and reimann (2016) proposed a framework that conceptualises the role of complexity in the learning sciences: the complex systems conceptual framework of learning (cscfl). this framework intends to reframe the traditional situated versus cognitive debate among educational scholars. for decades scholars have discussed ontological and epistemological issues on how learning processes must be interpreted (derry & steinkuehler, 2003). the discussion on the primacy of the cognitive (anderson, reder, & simon, 1996) or situated perspectives (engeström, 2001; greeno, 1997) has so far not yielded a sustainable consensus (jacobson et al., 2016). the cscfl adds a new perspective to this debate based on systems theory: it intends to harmonise both views, building on notions of complexity science. two central domains of the framework are: (1) complex collective behaviour in systems, and (2) behaviour of individual agents in systems. each of these domains is characterised by concepts which illustrate the complex character of learning. complex collective behaviour of agents or elements within a system follows the idea of self-organisation or emergence. this means that dynamics within and between systems are sensitive to initial conditions and are nonlinear. the notion of emergence is pivotal for systems theory, it is used to describe patterns of change. sawyer described it as an ‘attempt to bridge the micro-macro divide’ (2005, p. 210). using this concept of emergence, systems theory describes counterintuitive phenomena. resnick (1996) famously referred to the emergence of traffic jams or termite constructions. “strong emergence presents a direct challenge to determinism (the idea that given one set of circumstances there is only one logical outcome). with strong emergence, what emerges is always radically novel” (osberg & biesta, 2007, p. 34). it gives an insight into how non-linear patterns of interaction influence the relationships among individual agents of systems and how complex collective behaviour emerges out of it. these patterns are then fed by positive or negative feedback loops which result in these sometimes unexpected outcomes. essentially, jacobson et al. (2016) describe learning as not ontologically determined something that is rather as emergent. understanding learning and teaching through this prism evidently challenges research methodology. the idea of nestedness (burns & knox, 2011) is used to describe interactions between individual agents within systems and other systemic agents at, for instance, the mesoor macro-level. interactions within a classroom or interactions with external influences are, therefore, crucial to grasp why dynamics may be different among systems. it refers to the common idea of the contextual nature of learning (greeno, 1997). however, jacobson et al. (2016) use the term nestedness, which is more common in systems theory across many disciplines. an essential consequence of system theory is that teaching and learning must be understood as interaction between elements or agents in a system. this organic view sees the interactions in a class as fundamentally related. this contrasts with a mechanistic world view in which all elements of a system can be understood separately. a major question is therefore whether elements of the teaching process should be studied separately or not. following systems theory, agents within a system are in continuous interaction with the system in which they are active. it is therefore crucial to see the role of complex collective behaviour and of individual agents in systems. if learning must be interpreted as a complex phenomenon for which these characteristics are genuinely valid, this poses a tremendous challenge for our concept of the role of a teacher in it. davis and sumara (2007) described the shift from a mechanical view on teaching to a more organic one. damsa and jornet (2016) comparably argued to reframe learning as ‘collective achievements of whole ecosystems’ (p. 39). some distinctive properties, compared with mechanical management, are that knowledge in organic systems is said to be structured anywhere in the system, compared to top-down knowledge structures. communicative relations are horizontal rather than vertical in organic systems. and individual tasks within organic systems are said to be continuously adjusted and refined, compared to mechanical structures where tasks are specialised and differentiated. resnick (1996) argued for a decentralised concept of social institutions in order to account for self-organisation and emergence. ‘from the perspective of complexity multidimensional relationships and dynamic interactions among agents and elements, rather than predictable linear effects, are responsible for patterns and phenomena’ (cochran-smith, ell, ludlow, grudnoff, & aitken, 2014, p.5). davis and sumara (2007) comparably argue to reposition the role of teachers and teaching: teaching is not to be understood any more by what a teacher does or intends. tomlinson’s ideas on the responsiveness of teachers and reliance on formative assessment as grounds for differentiated instruction align with this repositioning of teaching and learning. differentiated instruction is here not a linear type of instructional design initiated by the teacher. the responsiveness of differentiated instruction is characterised by a decentralised concept of teaching. it accounts for the nestedness of learning and for patterns of emergence at a level of complex collective behaviour. moreover, it accounts for interactions within and across (open) systems. 4 system-theoretical grounding for research on differentiated instruction the concept of differentiated instruction is multifaceted. it invites teachers to adopt a growth mindset. it involves formative assessment to gather data on student heterogeneity and it is essentially characterised by the act of responding to these differences. system theory provides a useful theoretical framework to understand more deeply the complexity of differentiated instruction. in particular, we believe the notions of nestedness and emergence (combined with non-linearity) are of particular value for the study of differentiated instruction. in the following paragraph we link the fundamental properties of the concept of differentiated instruction to systems theory. we argue that, to fully acknowledge the complexity of differentiated instruction, empirical data are needed which are grounded on systems theory. in this section, three methodological challenges are presented which could increase our understanding of teaching in a differentiated classroom. although the challenges partially overlap, we present them in three different sections. 4.1 focus on the interplay between the individual and complex collective behaviour generalisation is often thought to be one of the main quality criteria in educational sciences (hammersley, 1997). with regard to teaching effectiveness, many educational scholars tend to generalise the validity of their ideas (cohen, manion, & morrison, 2007). cochran-smith et al. (2014) have noticed that these claims of generalisation in the educational sciences are challenged by complexity science. building on this claim, we add that, in a differentiated classroom generalised claims cannot account for all deviant profiles of individuality. therefore, scholarly research on differentiated instruction is challenged to describe the interplay between individual behaviour and complex collective behaviour. an exclusive focus on one type of agent of the differentiated classroom does not permit study of the interplay between all systemic agents. empirical data with an exclusive focus on the individual level (the role of teachers or students in a differentiated classroom) or on the microsystem-level (instructional design) may have an important added value for the debate on differentiated instruction. however, they do not permit a comprehensive insight into the complexity of it. to do so, the interplay between individual agents and complex collective behaviour within systems also needs to be studied. for differentiated instruction, this idea would imply the systematic study of the responsive act of teaching in a heterogeneous class, meaning analysing the impact of it at both studentand teacher-level. such studies would add substantially to our understanding of the responses within a differentiated teaching process. to monitor both the individual impact on students and the collective behaviour of a group of students are therefore important challenges for research on differentiated instruction. rarely, however, do studies seek to understand the link between phenomena at an individual level and the management of collective behaviour, which essentially is the teaching process in a differentiated class. as differentiated instruction essentially intends to maximise learning opportunities for all students in a class by taking their individual characteristics into account, a randomly composed research sample that undergoes a homogenised treatment, or that aims at reaching a common goal, is exactly the opposite of what the idea of differentiated instruction is. differentiated instruction essentially intends to take into account, not only the average student in a class (the microsystem-level), but also acknowledges deviant or changing profiles and characteristics of individual students (the individual perspective). this core characteristic of differentiated instruction makes scholarly research about it standing at odds with classic randomised research designs which aim at generalised research conclusions. methodologically these assertions bring us to plea for empirical data that inform on the interactions between agents in a learning system. with its responsive approach, differentiated instruction stands in an interactionist tradition (mead, 1934). building on blumer (1973), a long research tradition has focused on studying interactions in detail in order to understand the relationship between the social and the individual. classic interactionist methodology that documents 1-on-1 interactions is now critiqued for not (fully) accounting for the complexity of the interactions (e.g. sawyer, 2005). from a systems perspective, interactions between all systemic agents must be documented in order to gain deep insight into the interplay between the individual and complex collective behaviour. with regard to differentiated instruction, such an approach should lead to documenting, at least, the interactions among students and describing teacher-student interactions in a differentiated classroom. 4.2 interdependence with other systems the actual heterogeneity of a class changes throughout the year. students who drop out may influence opportunities for collaborative learning. moreover, newly arrived students may lack sufficient prior knowledge in order to participate in learning activities. it also occurs that less visible changes in the class influence the actual learning process when motivation, or other personal characteristics, change as a result of external influences. if a student experiences anything interesting or emotional in his or her personal life (e.g. a trip to a foreign country, the death of relative, an unusual encounter) this could, in a differentiated classroom, be a relevant take-off point for instruction. this nestedness of differentiated instruction is described by tomlinson as follows: “teachers who care about their students as individuals accept the difficult task of trying to identify the interests students bring to the classroom with them.” (tomlinson, 2001, p. 53). empirical methodology that intends to control data collection and data analysis cannot account for this nestedness. classic research designs make an effort to control the variables they study, and hence to wipe out the effects of external influences (sansone, morf, & panter, 2004). from a mechanistic point of view, these influences are apt to be avoided or neglected. by controlling for external bias, external validity of research conclusions may be increased (tipton, 2013). it needs to be questioned whether the idea of (semi-)controlled research conditions can result in external validity towards situations where external variables will have considerable influence. it is exactly for the lack of external validity of experimental research that bronfenbrenner’s naturalistic approach (1977) to research was grounded. a systemic perspective on data collection intends to account for externals instead of neglecting them. it could be assumed that research conclusions have stronger external validity when external influences are not ignored but seen as part of the complex reality. in system-theoretical terminology this would be described as accounting for the nestedness of systemic patterns. the philosophy of differentiated instruction implies that influences at the mesoor macro-level on the teaching process are taken into account. therefore it would be useful to build on the tradition that argues for the relevance of this stance. cultural-historical activity theory argues that this addresses the challenges and possibilities of inter-organisational learning (engeström, 2001). moreover a multi-systemic approach intends to stress cross-boundary relationships of agents in educational systems (akkerman & van eijck, 2013; bronkhorst & akkerman, 2016). in order to obtain a deep understanding of differentiated instruction, it must be targeted to describe concisely the nestedness of teaching and learning in a differentiated classroom. methodological choices for empirical research on the matter need to be informed by this nestedness. building on this argument some scholars ask for increased attention for the researcher’s role of reflexive methodology (alvesson & sköldberg, 2009) to acknowledge this nestedness. 4.3 non-linearity and emergence the character of differentiated instruction implies that teachers respond to student diversity. the cyclical process of teaching, learning, adapting teaching and further learning is essential to it. if teaching is not (only) to be seen as an activity with linear causal consequences, but as an agent (the teacher, the student) within a system that adds to the emergence of complex collective behaviour (such as learning and interactions between learners), then research is challenged to study the dynamics of the relationship between these two perspectives. these patterns are non-linear, they are cyclical in the sense that feedback mechanisms are at work. to illustrate the importance of non-linearity for differentiated instruction, we elaborate here on the role of mindset theory. if, indeed, a growth mindset is a central concept that facilitates the application of differentiated instruction both for teachers and students, then the influence of this factor should always be taken into account for studies on differentiated instruction. many studies have described spectacular results based on growth mindset interventions (blackwell, trzesniewski, & dweck, 2007; dweck, 2015; rattan et al., 2015). linear mechanistic research interventions have studied whether a growth mindset affects learning by isolating the growth mindset-factor from other motivational components in the learning process. in a systems-theoretical perspective on research methodology, however, this factor may not be isolated from other factors. a systemic approach will, in consequence, not study whether there is an effect of a growth mindset, but describe how growth mindset affects learning. it could, for instance, be hypothesised that feedback mechanisms between a growth mindset and goalsetting or tenacity of students result in the emergence of learning. differently stated: from a systems-theoretical perspective this influence of growth mindset on other aspects of the teaching process must not be isolated, rather must it be studied how feedback-mechanisms influence the outcomes of the learning process. if research is designed to be static or linear it is restricted in its scope to describe patterns of emergence between systemic agents. systems theory, however, contrasts with linear-causal thinking, given the assumption that, as a result of feedback mechanisms, the outcome of processes are thought to be unpredictable (brown, 2016; osberg & biesta, 2007). therefore, also patterns of non-linear causality need to be studied in order to understand fully the complexity of differentiated instruction. to incorporate feedback mechanisms in research designs means to challenge research to study patterns of non-linearity and emergence. as long as a mechanistic view on teaching and learning is adopted, this cyclical approach on learning does not necessarily pose problems for research design. with its focus on ongoing assessment and adaptive instructional design, differentiated instruction holds an iterative view on teaching. it acknowledges differences in learning pace between students and, consequently, adjusts the teaching process for students depending on the pace at which learning occurs. the use of formative assessment is seen as fundamental to document students’ learning needs and, hence, optimal learning chances are provided for all students in the classroom (coubergs et al., 2017; hall, 2006; tomlinson, 2015). building on ideas of non-linearity and emergence, this cyclical approach on learning would provide important new insights for scholarly research on differentiated instruction. classic planned experimental design-based research is difficult to align with these concepts of non-linearity or emergence. data collection that opens up for emergence is needed to mirror the complexity of learning processes in a differentiated classroom. as long (2001) states: “intervention is an on-going transformational process that is constantly re-shaped by its own internal organizational and political dynamic and by the specific conditions it encounters or (…) creates” (p. 27). this implies that data collection goes further than describing linear patterns of change, but includes more complex patterns of change within educational systems. describing these mechanisms at work is a major challenge for research on differentiated instruction. 5. design principles for research on differentiated instruction generalised knowledge on the micro-level of a classroom stands at odds with the concept of differentiated instruction. moreover, ideas of emergence and nestedness provide fundamental challenges for research designs on differentiated instruction. in this section we propose three design principles for this research that aim at aligning research design with the philosophy and practice of differentiated instruction such as proposed by tomlinson. examples of existing empirical research are added to illustrate these principles. although these examples all refer to learning in heterogeneous settings, not all of them explicitly refer to the construct of differentiated instruction such as proposed by tomlinson. principle 1: organic design. understanding the complexity of teaching in a differentiated classroom implies a holistic focus on the interaction between agents and components of systems instead of mechanical design which isolates particular agents or components. this means that what happens at the level of these agents or components is not necessarily seen as representative on a higher (meso-)level. only a holistic analysis can bring about the necessary understanding on teaching and learning in a differentiated classroom. jacobson et al. (2016) claim the concept of emergence must necessarily be considered when reflecting on causal relationships with regard to teaching and learning. moreover, certainly with regard to differentiated instruction, processes of emergence must stand central in empirical data collection. interventions that open up for non-linear patterns of change are needed for this purpose. applied to research on differentiated instruction this would mean, for instance, studying feedback mechanisms at work within a classroom related to a growth mindset theory. the following example illustrates how non-linear interventions could be designed in order to study these types of patterns. jafari and hashim (2012) described the use of advance organisers in order to improve english foreign language listening skills. an experimental intervention was designed to document students’ learning progress. advanced organisers were administered for a treatment group of students. however, depending on their actual learning progress, the strategy was differentiated. the monitoring of the learning progress of subgroups of students (higher or lower performing) permitted them to assess the strategy at this subgroup level. repeated formative assessment was used to document students’ learning progress and, hence, the further development of the intervention. in addition to monitoring students’ learning progress, this study also gathered qualitative data on the affective outcomes of the chosen strategy. again these data were related to the subgroups of students’ achievement levels. this type of intervention design represents closely the instructional design as it would be applied in a differentiated classroom. through extending the intervention for students who needed more practice or more extended direct teacher instruction, this study permitted them to gain insight into the structure of the learning process of diverse types of students within a group. referring to jacobson and kapur (2012) metaphor of the forest-tree perspective, we believe that this type of study approaches the idea of merging both perspectives on teaching and learning in a differentiated classroom. the teachers set a targeted goal for a heterogeneous group of students. however, patterns of change – learning – are monitored at the level of subgroups of students in order to gain insight into how learning emerges at the level of these subgroups. the organic nature of this intervention lies in the fact that it acknowledges diverse needs of students in its data collection (cognitive and affective, high and low performing). unfortunately, no data were provided in this study on how dynamics among students added to the emergence of learning at an individual or collective level. evidently, data collection that provides more detailed insight of learning at the individual level of students would come even closer to this system-theoretical principle. principle 2: interaction. studies on differentiated teaching must aim at matching the perspective of heterogeneous groups with learning at the level of the individuals within it. as a consequence, interactions between these levels must be monitored. responsiveness being one of the main characteristics of differentiated instruction, this element must necessarily occupy a central position in research design. this means that the students’ individual and collective characteristics are used as a basis for teaching and that the teachers’ response to these depends on formative assessment of students. students’ initial characteristics are pre-assessed and their progress is monitored using formative assessment. understanding how responsiveness of teaching is related to students’ individual learning is therefore a major challenge for empirical research. a study of martin-beltran, guzman, and chen (2017) describes how teachers differentiate discourse in order to foster collaboration between linguistically diverse students. this study is a typical example of interaction in the sense that it studies the interaction between teachers and their students. it draws upon system-theoretical principles in the sense that it intends to describe the complexity of discourse that teachers use in order to cater for diversity in their classes. these patterns of interaction are essential to understand how differentiated instruction materialises into everyday practice. recently the study of interactions within learning systems has attracted a lot of attention due to research design in which interactive software allows the documentation in detail of the learning processes of students. jacobson, kapur, so, and lee (2011) describe, for instance, how systems of hypermedia learning environments work when different types of scaffolding are provided. they collect data through interactive software. they argue how performance on problem solving transfer tasks is determined by the different types of scaffolding provided. building on the interactions between software, individual students and the scaffolds provided, systemic patterns could be described. principle 3: reflexivity. studies on differentiated instruction must acknowledge the interdependence of systems by adopting reflexivity. a more reflexive attitude of researchers is needed in order to achieve more transparency with regard to diverse external or internal dynamics that lie out of the control of researchers (tracy, 2010). building on the idea that control of all external influences is not achievable, it is our suggestion to increase reflexivity about conditions that lie out of control. the idea of proposing a research design in which all necessary factors are controlled seems unachievable with regard to teaching in a differentiated classroom. therefore, instead of controlling all potential disturbing variables, a researcher’s reflexive stance is needed to account for the systems’ interdependence. the notion of reflexivity encompasses different sorts of reflections on how the choices of a particular research design influence its results. alvesson and sköldberg (2009) suggest that this notion should be used not only for reflexivity on the choices made with regard to the systematics of data collection and techniques of procedures of data analysis but also propose to reflect on the interpretative and political-ideological character of research. with regard to differentiated instruction where the responsive character of teaching always implies teachers and, hence, also educational researchers to make difficult choices, we believe reflexivity to be the most credible option to foster transparency in educational research. a study by pilten (2016) comes close to what would be meant with the concept of reflexivity. it documents the experiences with the implementation of differentiated reading instruction of seventeen turkish elementary school teachers. their experiences are limited and their implementation of differentiated instruction is reluctant. although participants in this study often see a potential advantage of the idea of differentiated instruction, most of them classify the use of differentiated instruction as impracticable and thus hard to implement in practice. interestingly the authors of this study chose to reflect on the validity and reliability of their findings in the method section of their study. by doing so, they openly reflect on the extent to which their findings are credible. the phenomenological approach they use allows the authors to dig deep into the complexity of the participants’ teaching practice. building on the aforementioned ideas on open systems, it appears that the implementation of differentiated instruction by essence always relies on dynamics between open systems at the micro-level and other systemic levels. in this case, the implementation of differentiated instruction by the participants of this study could be mediated by external factors. this is why reflection on the validity is desirable. the act of reflecting on the way in which controlled conditions have been achieved and reflecting on potential inter-systemic relations should stand at the heart of methodological sections of studies on differentiated instruction. in the aforementioned example of pilten’s (2016) study, we believe that reflections on, for instance, growth mindset could have been an added value to strengthen further the reflexivity component of this study. 6. limitations we have sought to retheorise empirical research on differentiated instruction, drawing upon system-theoretical epistemic and ontological positions. a major critique on systems theory is that no consensus exists (yet) on the conceptualisation of some of its central concepts. according to fenwick, complexity science remains “slippery, heterogeneous and contested” (2012, p.110 ). most importantly, we notice a certain ambiguity in descriptions of how non-linearity and emergence are related to each other. in addition to this we believe that some of the concepts that are commonly used in systems theory, are sometimes differently conceptualised in more traditional educational approaches. we have built on the cscfl which provides terminology that is accessible for scholars in both cognitive and situated learning traditions. our choice to draw on the classic systems-theoretical terminology of this framework does not imply a positioning in favour of, or against, conceptualisation as situated learning theory or any other research tradition. we want to broaden further and refine research on differentiated instruction, but not by disputing any approach. however, by showing complexity, we argue for the added value of a system-theoretical stance. finally, it may be noticed that we have used a human-centred interpretation of systems theory. as systems theory originates from physics, the consequences of it cannot be strictly focused on human beings. we believe our choice to interpret differentiated instruction with a dominant human interactionist focus, may be argued referring to the existing literature of tomlinson et al. (2003). it could, however, be worthwhile to reinterpret differentiated instruction by tracing more clearly its socio-materiality. research on differentiated instruction has a tendency to couple learning and teaching with a strictly human-centered ontology. fenwick (2012) argues however, against the tendency to focus on human learning figures: “complexity science urges a re-focusing on the relations that produce things, not the things themselves”, (2010, p.111) several scholars have treated material conditions in which differentiated instruction is enacted (gaitas & martins, 2017; keuning et al., 2017). they see them as fostering or inhibiting teaching practice. future research could determine to which extent these material conditions are actually shaping the nature of differentiated teaching and learning. 7. conclusion the concept of differentiated instruction describes a philosophy and an approach to teaching to adapt to diversity in heterogeneous classroom settings (tomlinson, 2015). the complexity of the concept challenges scholarly research on it: it seeks to practice a responsive approach to teaching in which a variety of differences among students are addressed. a range of strategies is used for flexible grouping of students. moreover a growth mindset is adopted in order to maximise learning of all students. system-theoretical insights are needed to describe concisely and to understand deeply, teaching and learning in a differentiated classroom. the notions of non-linearity and emergence, and the concept of nestedness challenge scholarly study on differentiated instruction to broaden and refine research methodology. they help understanding the complex interplay between individual and collective behavior in a differentiated classroom. moreover, they provide insight in the role of interdependence with other systems that mediate learning in a differentiated classroom. methodology that draws upon the description of human interactions, or that includes interventions that open up for non-linearity or emergence, may be used to underpin empirical research on differentiated instruction. moreover, scholars need to reflect on conditions that lie out of their control during data collection. based on these ideas of systems theory, we suggest three design principles for research on differentiated instruction: organic design, interactions and reflectivity. organic design could apply a holistic focus to differentiated instruction. focus on interactions could draw attention to the role of responsivity of the construct. reflexivity is needed in order to account for conditions that lie out of control of studies that focus on differentiated instruction. keypoints differentiated instruction is a complex teaching concept that needs research aligned with this complexity educational research on it is challenged to use theoretical foundations that align with this complexity, both ontological and epistemological three methodological design principles are proposed to align scholarly research on differentiated instruction with the notions of non-linearity and emergence, and nestedness these principles are: organic design, interaction and reflectivity references akkerman, s. f., & van eijck, m. (2013). re-theorising the student dialogically across and between boundaries of multiple communities. itish educational research journal, 39(1), 60-72. doi:10.1080/01411926.2011.613454 alvesson, m., & sköldberg, k. (2009). reflexive methodology. new vistas for qualitative research (2nd ed.). london: sage. anderson, j., r., reder, l., m. , & simon, h., a. (1996). situated learning and education. educational researcher, 25(4), 5-11. doi:10.3102/0013189x025004005 bakker, c., & montesano montessori, m. (2016). complexity in education. from horror to passion. rotterdam: sense. blackwell, l. s., trzesniewski, k. h., & dweck, c. s. (2007). implicit theories of intelligence predict achievement across an adolescent transition: a longitudinal study and an intervention. child development, 78(1), 246-263. doi:10.1111/j.1467-8624.2007.00995.x blumer, h. (1973). a note on symbolic interactionism. american sociological review, 38(6). bronkhorst, l. h., & akkerman, s. f. (2016). at the boundary of school: continuity and discontinuity in learning across contexts. educational research review, 19, 18-35. doi: https://doi.org/10.1016/j.edurev.2016.04.001 bown, b. (2016). a systems thinking perspective on change processes in a teacher professional development programme. journal of education, 66, 37-64. burns, a., & knox, j. (2011). classrooms as complex adaptive systems: a relational model. teaching english as a second or foreign language, 15(1). chen, s. c., yang, s. j. h., & hsiao, c. c. (2016). exploring student perceptions, learning outcome and gender differences in a flipped mathematics course. itish journal of educational technology, 47(6), 1096-1112. doi:10.1111/bjet.12278 cochran-smith, m., ell, f., ludlow, l., grudnoff, l., & aitken, g. (2014). the challenge and promise of complexity theory for teacher education research. teachers college record, 116(5). doi:10.1007/s10833-012-9183-4 cohen, l., manion, l., & morrison, k. (2007). research methods in education. oxford: routledge coubergs, c., struyven, k., vanthournout, g., & engels, n. (2017). measuring teachers’ perceptions about differentiated instruction: the di-quest instrument and model. studies in educational evaluation, 53, 41-54. doi:10.1016/j.stueduc.2017.02.004 damsa, c., & jornet, a. (2016). revisiting learning in higher education—framing notions redefined through an ecological perspective. frontline learning research, 4(4), 39-47. davis, b., & sumara, d. (2007). complexity science and education: reconceptualizing the teacher’s role in learning. interchange, 38(1), 53-67. doi:10.1007/s10780-007-9012-5 de neve, d., devos, g., & tuytens, m. (2015). the importance of job resources and self-efficacy for beginning teachers' professional learning in differentiated instruction. teaching and teacher education, 47, 30-41. doi:10.1016/j.tate.2014.12.003 derry, s. j., & steinkuehler, c. a. (2003). cognitive and situative theories of learning and instruction. in l. nadel (ed.), encyclopedia of cognitive science (pp. 800–805). london nature. deunk, m., doolaard, s., smale-jacobse, a., & bosker, r. (2015). differentiation within and across classrooms: a systematic review of studies into the cognitive effects of differentiation practices . retrieved from groningen: dweck, c. s. (2008). mindset: the new psychology of succes. new york: ballantine. dweck, c. s. (2015). growth mindet. itish journal of educational psychology, 85(2), 242-245. doi:10.1111/bjep.12072 engeström, y. (2001). expansive learning at work: toward an activity theoretical reconceptualization. journal of education and work, 14(1), 133-156. doi:10.1080/13639080020028747 fenwick, t. (2012). tracing the socio-material: emerging approaches to theory and research in adult education. in t. fenwick, r. edwards, & p. sawchuk (eds.), emerging approaches to educational research. london: routledge. gaitas, s., & martins, m. a. (2017). teacher perceived difficulty in implementing differentiated instructional strategies in primary school. international journal of inclusive education, 21(5), 544-556. doi:10.1080/13603116.2016.1223180 gay, g. (2002). preparing for culturally responsive teaching. journal of teacher education, 53(2), 106-116. greeno, j. g. (1997). on claims that answer the wrong questions. educational researcher, 26(1), 5-17. doi: http://doi.org/10.3102/0013189x026001005 hall, t. s., nicole; meyer, anne. (2006). differentiated instruction and implications for udl implementation [press release] hammersley, m. (1997). educational research and teaching: a response to david hargreaves' tta lecture. itish educational research journal,, 23(2), 141-161. jacobson, m., & kapur, m. (2012). learning environments as emergent phenomena: theoretical and methodological implications of complexity. in d. h. jonassen & l. s. (eds.), theoretical foundations of learning environments (2nd ed.). london: routledge. jacobson, m., kapur, m., & reimann, p. (2016). conceptualizing debates in learning and educational research: toward a complex systems conceptual framework of learning. educational psychologist, 51 (2), 210-218. doi:10.1080/00461520.2016.1166963 jacobson, m., kapur, m., so, h., & lee, j. (2011). the ontologies of complexity and learning about complex systems. instructional science, 39(5), 763-783. doi:10.1007/s11251-010-9147-0 jafari, k., & hashim, f. (2012). the effects of using advance organizers on improving efl learners' listening comprehension: a mixed method study. system, 40 (2), 270-281. doi:10.1016/j.system.2012.04.009 keuning, t., geel, m. v., frèrejean, j., merriënboer, j. v., dolmans, d., & visscher, a. j. (2017). differentiëren bij rekenen: een cognitieve taakanalyse van het denken en handelen van basisschoolleerkrachten. pedagogische studiën, 94, 160-181. long, n. (2001). development sociology: actor perspectives. london: routledge. luhmann, n. (2013 ). introduction to systems theory. cam idge: polity press. martin-beltran, m., guzman, n. l., & chen, p. j. j. (2017). "let's think about it together:' how teachers differentiate discourse to mediate collaboration among linguistically diverse students. language awareness, 26(1), 41-58. doi:10.1080/09658416.2016.1278221 mead, g. (1934). mind, self, and society. chicago: university of chicago press. morrison, k. (2008). educational philosophy and the challenge of complexity theory. educational philosophy and theory, 40(1), 19-34. doi: https://doi.org/10.1111/j.1469-5812.2007.00394.x osberg, d., & biesta, g. j. j. (2007). beyond presence: epistemological and pedagogical implications of ‘strong’ emergence. interchange, 38(1), 31-51. doi:10.1007/s10780-007-9014-3 pilten, g. (2016). a phenomenological study of teacher perceptions of the applicability of differentiated reading instruction designs in turkey. educational sciences-theory & practice, 16(4), 1419-1451. doi:10.12738/estp.2016.4.0011 prast, e. j., weijer-bergsma, e. v. d., kroesbergen, e. h., & luit, j. e. h. v. (2015). readiness-based differentiation in primary school mathematics: expert recommendations and teacher self-assessment. frontline learning research, 3(2), 90-116. prigogine, i. (1980). from being to becoming: time and complexity in the physical sciences. new york: w h freeman & co . rattan, a., savani, k., chugh, d., & dweck, c. s. (2015). leveraging mindsets to promote academic achievement: policy recommendations. perspectives on psychological science, 10(6), 721-726. doi:10.1177/1745691615599383 resnick, m. (1996). beyond the centralized mindset. journal of the learning sciences, 5(1), 1-22. sansone, c., morf, c. c., & panter, a. t. (2004). the sage handbook of methods in social psychology. london: sage. sawyer, k. (2002). emergence in psychology: lessons from the history of non-reductionist science. human development, 45, 2-28. sawyer, k. (2005). social emergence: societies as complex systems. cam idge: cam idge university press. schleicher, a. e. (2013). preparing teachers and developing school leaders for the 21st century. lessons from around the world . retrieved from paris: schumm, j. s., & vaughn, s. (1995). getting ready for inclusion: is the stage set? learning disabilities research and practice, 10 (3), 169-179. shabani, k., khatib, m., & ebadi, s. (2010). vygotsky's zone of proximal development: instructional implications and teachers' professional development. english language teaching, 3(4), 237-248. tipton, e. (2013). improving generalizations from experiments using propensity score subclassification: assumptions, properties, and contexts. journal of educational and behavioral statistics, 38 (3), 239-266. doi:10.3102/1076998612441947 tomlinson, c. a. (2000). the differentiated classroom: responding to the needs of all learners . alexandria: association for supervision and curriculum development. tomlinson, c. a. (2001). differentiating instruction in mixed-ability classrooms (2nd ed.). alexandria: association for supervision and curriculum development. tomlinson, c. a. (2015). teaching for excellence in academically diverse classrooms. society, 52(3), 203-209. doi:10.1007/s12115-015-9888-0 tomlinson, c. a., ighton, c., hertberg, h., callahan, c. m., moon, t. r., imijoin, k., . . . reynolds, t. (2003). differentiating instruction in response to student readiness, interest, and learning profile in academically diverse classrooms: a review of literature. journal for the education of the gifted, 27(2-3), 119-145. tracy, s. (2010). qualitative quality: eight “big-tent” criteria for excellent qualitative research. qualitative inquiry, 16(10), 837-852. van klaveren, c., vonk, s., & cornelisz, i. (2017). the effect of adaptive versus static practicing on student learning evidence from a randomized field experiment. economics of education review, 58 , 175-187. doi: https://doi.org/10.1016/j.econedurev.2017.04.003 von bertalanffy, l. (1968). organismic psychology and systems theory. worcester: clark university press. ware, f. (2006). warm demander pedagogy culturally responsive teaching that supports a culture of achievement for african american students. urban education, 41(4), 427-456. doi:10.1177/0042085906289710 wass, r., & golding, c. (2014). sharpening a tool for teaching: the zone of proximal development. teaching in higher education, 19 (6), 671-684. frontline learning research vol. 5 no. 3 special issue (2017) 94 122 issn 2295-3159 corresponding author information: margje w. j. van de wiel, dep. of work and social psychology, faculty of psychology and neuroscience, maastricht university, p.o. box 616, 6200 md maastricht, the netherlands. email: m.vandewiel@maastrichtuniversity.nl, phone: +31-43-3882171, fax: +31-43-3884211. doi: http://dx.doi.org/10.14786/flr.v5i3.257 examining expertise using interviews and verbal protocols margje w. j. van de wiel maastricht university, the netherlands article received 4 may / revised 2 march / accepted 23 march / available online 14 july abstract to understand expertise and expertise development, interactions between knowledge, cognitive processing and task characteristics must be examined in people at different levels of training, experience, and performance. interviewing is widely used in the initial exploration of domain expertise. work and cognitive task analysis chart the knowledge, skills, and strategies experts employ to perform effectively in representative tasks. interviewing may also shed light on the learning processes involved in acquiring and maintaining expertise and the way experts deal with critical incidents. interviews may focus on specific tasks, events, scenarios, and examples, but they do not directly tap the representations involved in task performance. methods that collect verbal protocols during and immediately after task performance better probe the ongoing processes in representing problems and accomplishing tasks. this article provides practical guidelines and examples to help researchers to prepare, conduct, analyse, and report expertise studies using interviews and verbal protocols that are derived from thinking aloud, dialogues or group discussions, free recall, explanation, and retrospective reports. in a multi-method approach, these methods and other techniques need to be combined to fully grasp the nature of expertise. this article shows how the cognitive processes in data collection constrain data quality and highlights how research questions guide the development of coding schemes that enable meaningful interpretation of the rich data obtained. it focuses on professional expertise and provides examples from medicine including visual tasks. this comprehensive review of qualitative research methods aims to contribute to the advancement of expertise. keywords: expertise, interviews, verbal protocols, cognitive processing, analysis mailto:m.vandewiel@maastrichtuniversity.nl http://dx.doi.org/10.14786/flr.v5i3.257 van de wiel | f l r 95 1. introduction for this special issue on “methodologies for studying visual expertise”, the present article will discuss the qualitative research methods of interviews and verbal protocols to examine expertise and expertise development. this article aims to guide students, practitioners, and researchers new to the field of expertise research, or these types of qualitative research, when and how to use these methods to answer their research questions. starting from a theoretical framework of expertise and cognitive processing in task performance, this article provides practical guidelines so that researchers can prepare, conduct, analyse, and report expertise studies using interviews and verbal protocols. the rationale behind the methods, as well as their strengths and weaknesses, are explained to understand how procedures should be designed to maximise the quality of the data. although these guidelines for research are applicable to all expertise domains, the focus here is on professional expertise. most examples will be drawn from medical expertise research, as it has a long-standing tradition of using diverse methods of verbal protocols, and includes various areas of visual expertise. this comprehensive overview of qualitative research methods contributes to the literature by showing why and how the methods can best be used to deliver valid, high-quality verbal data when examining expertise. the literature is reviewed from an analytical and practical perspective to connect different research traditions that shed light on the nature and origins of expertise. expertise research may add to the advancement of any domain, as careful analysis of task characteristics, performance, and underlying knowledge and cognitive processing is at the heart of improving current practices. this review aims to provide researchers and practitioners who want to embark on this endeavour with fundamental insights into theory and methods that help to further develop the level of expertise in their domain of interest. the article is organised into five further sections. first, expertise is defined in terms of outcomes, underlying knowledge and processes, and their interaction with task and domain characteristics. second, the steps in preparing expertise research are discussed, starting with the research questions, familiarisation with the domain of expertise and the specific tasks at hand, the selection of experts and other participants, and the main criteria for choosing between interviews or verbal protocols. third, the interview method is described and placed within the context of research on work and expertise. fourth, the characteristics of interview and verbal protocol methods are described in light of the cognitive processing involved in task performance and data collection. moreover, five methods used to gather verbal protocols to reveal expert task performance are discussed and illustrated in more detail. finally, a conclusion is provided that summarises the main issues to be considered when designing studies that examine expertise using the qualitative research methods of interviews and verbal protocols. 2. expertise two dominant perspectives on expertise can be distinguished in the literature. the expertperformance approach (ericsson, 1996, 2004, 2015; ericsson & smith, 1991) characterises expertise as the capability to demonstrate reproducible superior performance on representative tasks in a specific domain. the highest expertise level is achieved when individuals are able to go beyond mastery and contribute their creative ideas and innovations to the task at hand. although years of practice and experience are needed to become an expert, skilled performance and experience alone are not enough. routine behaviour and full automaticity should be counteracted by gaining high-level control of performance that allows further improvements to be made. in the expert-novice research approach (chi, glaser, & farr, 1988; chi, 2006a), expertise has been characterised by differences in performance and underlying knowledge between groups with increasing levels of experience in a particular domain. experts have a large and well-developed knowledge-base, that is tuned to the tasks performed and the problems encountered, and allows fast and accurate performance in routine situations. in more complex situations, they can apply their knowledge flexibly when trying to understand the situation and decide upon further actions. van de wiel | f l r 96 the obvious similarity in these characterisations of expertise is that they both emphasise routine, automatic versus controlled, deliberate performance that is adapted to the task at hand. both approaches explain how the development of knowledge and skills underlies expert performance (feltovich, prietula, & ericsson, 2006). simply said, expertise is the result of activating the right knowledge at the right time (anderson, 1996). experts have developed rich and coherent knowledge structures that allow immediate access to the relevant knowledge, strategies, skills, and control mechanisms. domain-specific task performance is mediated by evolving representations of the task and problem they attend to. this enables experts to perform effectively and efficiently, coordinating automatic thoughts and actions with deliberate thinking. problem representations guide them in selectively focusing on relevant information and features that novices are not aware of. moreover, they help experts to carefully monitor and adapt their performance in an ongoing process. figure 1 illustrates how both incoming information and the experts’ knowledge in long-term memory continuously interact to determine the content of working memory in task performance. the capacity of working memory is enhanced by retrieval cues that directly access the relevant parts of experts’ knowledge in long-term working memory (ericsson & kintsch, 1995), enabling them to coordinate thoughts and actions in cognitive processing. the evolving mental representations in task performance reflect the content of working memory. experts update their knowledge and skills by means of study, practice and experience. they enhance learning from their experiences by seeking feedback and reflecting upon their performance to find weak aspects in processes and outcomes that might be improved. expertise development is a gradual process in which the knowledge and skills needed to plan, monitor and evaluate performance are refined during practice. this requires the motivation to improve performance and invest effort in deliberate practice (ericsson, krampe, & tesch-röme, 1993; ericsson & pool, 2016). figure 1. information processing in task although both approaches define expertise in relative terms and focus on tasks representative of the domain, one important difference between them relates to the standards of performance. whereas the expertperformance approach (ericsson, 1996, 2015) focuses on top-performance that can be objectively measured, the expert-novice approach (chi et al., 1988, 2006a) is more pragmatic in comparing novices and students with intermediate levels of training to experienced performers within a particular domain. in professional domains, such as medicine, auditing, law, teaching, software engineering and psychotherapy, it is not as easy to objectively measure performance as it is in domains, such as sports and games, in which clear outcomes (e.g., time, points gained) are available. the experience of the professional and the presence of professional criteria, such as degrees, licenses, memberships of professional organisations, prizes, and teaching experience usually work well to identify experts (evetts, mieg, & felt, 2006; hoffman, shadbolt, burton, & klein, 1995; mieg, 2006). there is a notable absence of a ‘gold standard’ of professional performance that is based on a validated objective outcome measure (ericsson, 2004; weiss & shanteau, 2003; shanteau, weiss, thomas, & pounds, 2001) and one best solution or approach to a problem may not even exist (tracey, wambold, lichtenberg, & goodyear, 2014). in medicine, for example, physicians have to make decisions van de wiel | f l r 97 under conditions of uncertainty when they encounter more complex and rare patient problems. in many other professions, experts are confronted with uncertainty and new situations that require performance on the edge of what they may accomplish based on their knowledge and skills (klein, 2008; salas & klein, 2001). experts, furthermore, play an important role in the advancement of their domain and in setting (new) standards for performance (boshuizen & van de wiel, 2014; ericsson, 2009, 2015; evetts et al, 2006; lesgold, 2000). in summary, how expert performance can best be defined depends on several factors including the domain, the tasks, and the type of problems to be solved. the accumulated body of knowledge and skills available in a domain constrains the level of performance that can be acquired by individuals. shanteau (1992) found in his analysis that performance is better in structured domains in which incoming information is static, problems are predictable, conditions are similar, tasks are repetitive, and objective analysis, feedback and decision aids are available. in these structured domains, individuals have more chances to learn and improve their performance as compared to less structured domains, which do not share these task characteristics. kahneman and klein (2009) discuss how intuitive expertise, i.e., automatic accurate judgment, can only be developed in high-validity domains in which the environment is predictable via the recognition of a set of cues. if professionals are given adequate opportunity to practice, they can learn the causal structure and/or the statistical regularities that enable this recognition process. classical examples of ill-structured or low-validity domains include wine tasting, stock broking, and clinical psychology, all domains in which judgment is inconsistent. ericsson (2014, 2015; ericsson & pool, 2016) argues that professional performance can only be improved by searching for and identifying reproducible superior outcomes, which can then be used to guide deliberate practice. libraries of problem situations with known outcomes and simulators enable intensifying practice with immediate feedback for problems that are uncommon or have high-stakes in real practice. research on expertise contributes to the development of a domain and the performance levels that can be achieved by professionals on essential tasks. 3. examining expertise when examining expertise, first the research question needs to be clearly formulated: “what do you want to know about expertise?” as expertise is based on the development of a well-organised body of knowledge that determines the processes and strategies used in task performance, the most obvious questions are “what knowledge and skills underlie expert performance and how do they develop?”. the research question can focus on the representations of the problem to be solved, or the task to be performed and how these differ between novices and experts (e.g., chi, feltovich, & glaser, 1981; van de wiel, boshuizen, & schmidt, 2000), or change as a result of practice (boshuizen, van de wiel, & schmidt, 2012). but it may also focus on the knowledge and strategies used in problem solving and in performing a task (e.g,, boshuizen & van de wiel, 1999; diemers, van de wiel, scherpbier, baarveld, & dolmans, 2015; gilhooly et al., 1997; lesgold et al., 1988; kok et al., 2015). it can focus on the learning processes and how teaching and instruction can help novices to become experts (e.g., chi, bassok, lewis, reimann, & glaser, 1989; kok, de bruin, robben, & van merriënboer, 2013). it may also focus on the activities experts engage in to learn from their experience, further develop and maintain their expertise (e.g., ericsson et al., 1993; van de wiel, van den bossche, janssen, & jossberger, 2011), and the self-regulations skills they apply to plan, control and evaluate their performance. finally, it may be important to investigate the ways in which the knowledge of experts can fall short by focusing on biases and near-errors and how these might be overcome (e.g., chi, 2006a; hashem, chi, & friedman, 2003; elstein & schwarz, 2002). the general themes addressed by these research questions will require further specification depending on previous research, the domain of expertise, and the research interests. in the professional domain of medicine, there is a long tradition of expertise research that started with the seminal work of elstein, shulman, & sprafka in 1978. while searching for general problem solving van de wiel | f l r 98 skills, studies in the early years consistently showed that experts and novices used the same strategy of generating and testing hypotheses in diagnostic problem solving, but that experts generated diagnostic hypotheses faster and more accurately (elstein et al., 1978; norman, eva, brooks, & hamstra, 2006; neufeld, norman, feightner, & barrows, 1981). the accuracy of physicians’ diagnoses, however, was found to be case-specific and tied to the domain of clinical experience (elstein et al., 1978). research was then directed at uncovering the nature and organisation of knowledge underlying physicians’ performance in interaction with the patient cases diagnosed. results have shown that their large and well-developed knowledge base enables physicians to automatically retrieve the relevant knowledge in routine cases, as well as to analytically process cases that are difficult or evoke a sense of alarm (elstein & schwarz, 2002; stolper et al., 2010). elucidating the nature and acquisition of medical expertise is still an active research field that contributes to safe patient care and medical education. studying visual expertise in medicine is a rapidly growing field, as exemplified by this special issue. imaging techniques are important diagnostic tools that develop quickly and require complex knowledge and skills that need to be learned and assessed (gegenfurtner, siewiorek, lehtinen, & säljö, 2012). while examining images, bottom-up and top-down processes continuously interact and determine whether significant features are recognised and correctly interpreted. their knowledge ultimately determines whether experts see what can be detected, and understand what they see. a broad array of research questions is open to investigation in this field. having established what you want to know about expertise, the second question that needs to be answered in expertise research is “who are the experts?” as expertise manifests itself in the context of the tasks that experts engage in, this question is intricately intertwined with the question: “what tasks are critical to the domain?” depending on the type of performance outcomes available, it may be more or less straightforward to identify experts that consistently show superior performance on representative tasks. a thorough familiarisation with the domain which focuses on the tasks performed is needed to find out what may characterise expert behaviour. if objective outcome measures do not exist, professional criteria that provide social recognition, such as experience, degrees, licenses, job titles, status, and prizes, might be used to define experts (evetts, mieg, & felt, 2006; hoffman et al., 1995; mieg, 2006), as well as peer judgments that ask professionals to identify the best performers in their field, or those whom they would go to for advice (ericsson, 2006a; kahneman & klein, 2002; shanteau, 2002). other groups of participants must be included to examine in what way experts differ from those who are less experienced within the domain, or those who have worked as long in the profession but are considered to have less expertise. to study the development of expertise, groups with different levels of experience in the domain (ranging from naïve, novice, intermediate, and advanced to expert) are compared to each other. a group of trainees can also be followed on their developmental path towards becoming a professional. another crucial question to be addressed in preparing expertise research is “what research method(s) are most suitable to investigate the research questions in this field?”. in this article, the focus is on qualitative research methods of interviews and verbal protocols as they play a crucial role in uncovering the characteristics and origins of expertise in any domain. interviewing is a very straightforward way to initially explore a specific domain of expertise. it can be used to gather information on the relevant tasks undertaken, the knowledge and skills needed to perform these tasks and solve problems, the learning processes involved in education and continuous development, and the pitfalls associated with expertise that need to be dealt with. interviews deliver verbal protocols as data resulting from the answers to the interview questions. however, to better capture the cognitive processing of experts in task performance, verbal protocols must be gathered that are directly related to the task-specific processes. expertise researchers have, therefore, developed methods that study, in addition to behavior and outcomes of representative task performance, the thinking processes involved. they do so by probing the experts’ underlying representations, knowledge, and reasoning (chi, 2006b, 1997; ericsson & simon, 1980, 1993; feltovich et al. 2006; hoffman et al., 1995). these methods provide insight into the content of working memory during, or immediately after, task performance (see figure 1). two common methods used to assess online thinking by gathering verbal data during task performance are thinking aloud and discourse analysis of dialogues and group discussions. three common methods used to gather verbal protocols after task processing are free recall protocols, explanation van de wiel | f l r 99 protocols, and retrospective reports. table 1 provides an overview of the qualitative research methods discussed in the present article, and how they are related to both the domain and the task when examining expertise. table 1 overview of qualitative methods used to examine expertise in relation to the domain and the task exploration of domain expertise cognitive processing related to a specific task during task performance after task performance interviews thinking aloud free recall focus groups dialogues and group discussions explanation retrospective reports to guarantee the quality of the data obtained by interviews and verbal protocols in expertise research, careful preparation is required. the most critical steps in preparing studies using these qualitative research methods are summarised in table 2. in addition to the steps outlined above, how the protocols will be coded, and how the study will be communicated to the participants, are two important considerations for both interviews and verbal protocols. in relation to interviews, the emphasis is on preparing the interaction with the interviewee, whereas for verbal protocols the emphasis is on selecting and designing tasks that may differentiate expert and novice behaviour. these tasks should reflect the same goal-directed processing as required in real-world tasks. in the following sections, the methods are discussed to provide practical guidelines for developing, conducting, and analysing expertise studies that deliver valid, high-quality verbal data, which can then be reported in a transparent way. in addition, the strengths and weaknesses of these particular methods are highlighted and compared, and specific issues related to expertise research are explained and illustrated. table 2 steps in preparing studies using interviews and verbal protocols to examine expertise formulating the research question familiarisation with the domain and tasks determining the experts and participants interviews developing the interview guide preparing for the role of interviewer verbal protocols selecting and analysing the task developing the task materials and instructions developing coding schemes for analysing the verbal protocols framing the communication to participants van de wiel | f l r 100 4. interviews interviewing is one of the most common methods used to gather information about a given topic. it is a very natural process of inquiry that is used in everyday communication; just think about how often we engage in asking questions and receiving answers. the key to all good interviews is to clearly ask what you want to know, and to make sure that you receive the answer that allows you to know what you want to know. this sentence describes interviewing in a nutshell, highlighting the importance of the research questions, the interview guide, and the role of the interviewer in asking questions and evaluating answers. in relation to expertise research, it is important that this characterisation of interviewing also shows a realistic perspective on research. it assumes that we can come to understand a topic by obtaining relevant information in an objective way by interviewing a representative sample of participants (emans, 2004; king & horrocks, 2010). the interviewer wants to reveal what interviewees know, do, think, feel, believe, intend to do, want, or need, and assumes that the interviewee can communicate this during the interview. the basic processes involved in interviewing for research purposes are well explained by emans (2004) and summarised in figure 2. the goal is to reveal the cognitions of the interviewee, as related to a certain topic, i.e., the interviewee’s mental processes and the products of these processes, usually in the form of information, knowledge, thoughts, feelings and ideas about the topic. the task of the interviewer is to create a situation and ask questions that motivate an interviewee to connect to these cognitions and verbalise them in a reliable manner. the interviewer must carefully listen to the answers provided, and check whether these answers include the information necessary to answer the research questions. if not, further questions need to be asked. to obtain data for subsequent analysis, the whole interview needs to be recorded and transcribed verbatim, thus without any interpretation of the data. figure 2. processes in interviewing (adapted from emans, 2004). in the context of work, interviewing is the most common method used to interact with experienced practitioners as subject-matter experts to gather information about all kind of aspects of the work they are engaged in and the vocabulary they use. in job analysis, cognitive task analysis, and knowledge elicitation, interviews are used to yield primary insights into the tasks experts perform, the knowledge and skills underlying their performance, and the conditions that shape their performance. job analysis focuses on work activities, worker attributes, and/or work context and is used to inform human resource management practices, such as personnel selection, training, and performance management (bartram, 2008; sanchez & levine, 2012). job analysis is also a first step in job (re)design, workplace and equipment design and organising team work and provides an overview to determine what tasks need to be further scrutinised by task analysis or cognitive task analysis (chipman, schraagen & shalin, 2000; dubois & shalin, 2000). detailed analysis of tasks in terms of goals, actions, and thought processes helps to articulate what is not directly observable but which may be expressed by experts when they are sufficiently guided. in knowledge engineering, knowledge elicitation techniques are specifically designed for this purpose in order to develop expert systems and knowledge management systems (hoffman et al., 1995; hoffman & lintern, 2006; van de wiel | f l r 101 shadbolt & smart, 2015). interviews in different forms have been applied in a wide variety of professional domains to elicit knowledge from experts. the unstructured interview is often employed in an initial exploratory phase in which investigators familiarise themselves with the domain in an informal setting. in later phases, more structured interviews can scaffold the knowledge elicitation process by focusing on specific events, such as critical incidents in the critical incident technique (flanagan, 1954) and critical decisions made in unusual and challenging cases in the critical decision method (hoffman & lintern, 2006; shadbolt & smart, 2015). experts can also be asked to respond to an evolving scenario or to specific probe questions in order to systematically unravel their task representations (shadbolt & smart, 2015). the literature on job analysis, cognitive task analysis, and knowledge elicitation also emphasises the need for research methods to be combined to achieve a thorough understanding of the job and the tasks under study. as complex tasks often require teamwork, interviewing teams in addition to individual experts can provide further insight into how experts from different disciplines work together, build a shared understanding of tasks and situations, and coordinate and distribute tasks amongst themselves (salas, rosen, burke, goodwin, & fiore, 2006). the interview method provides relatively quick access to information from several teams within a domain, as compared to observation in the field or simulation of task performance. furthermore, as jobs, tasks, equipment, and work roles are not stable but rather continuously developing, experts in the field par excellence may provide valuable perspectives on future developments, and innovations, and insights into novel problems and how to deal with them (boshuizen & van de wiel, 2014; lesgold, 2000; sanchez & levine, 2012). in fact, experts shape the advancement of their field, and may share their ideas about these developments, and how to support them by individual and organisational learning, in future-oriented interviews (bartram, 2008; evetts et al, 2006; lesgold, 2000). whereas less structured, informal interviews are very suitable for an initial exploration of a domain or topic, semi-structured interviews are most useful in expertise research as they result in more objective data that also allow for comparisons to be made between different expertise groups. these interviews provide enough guidance to structure the conversation, but also enable the acquisition of meaningful information, as the interviewer may interact with the interviewee after an initial open question (emans, 2004). focus groups that investigate the opinions and experiences of people while they are interacting in groups are valuable tools to examine expertise for both exploratory and comparative purposes. in focus groups, initial questions guide the interview by inviting participants to share their views. in preparing interviews and focus groups for expertise research, the steps outlined in table 2 must be kept in mind. in the following sections, the development of the interview guide, the preparation for the tasks of interviewer, and the analysis of the data are described and illustrated. although face-to-face interviews are most commonly used, and taken as a starting point in the descriptions, the guidelines are largely applicable to interviews conducted in groups and by telephone or skype (deakin & wakefield, 2014; emans, 2004). in a subsequent section, the focus group method is explained in more detail. 4.1 the interview guide an interview guide is a script that helps to ensure that the interview is conducted in a standardised way. it consists of an introduction to the study, the body of the interview outlining the main questions and possible follow-up questions, the transitions between questions, and a conclusion (emans, 2004; king & horrocks, 2010; skopec, 1986). in the recruitment phase, participants already receive information about the study that may influence their contributions, and thus, the quality of the data gathered. therefore, it is good practice to compose this information as the first step of the interview guide. for researchers new to the field of the interview method table 3 provides an overview of the topics advised to be addressed in the interview guide. van de wiel | f l r 102 table 3 general outline of the interview guide information sheet provide context and purpose of the research; invite the participant; explain the procedure and expectations; explain the costs and benefits for participants; emphasise that participation is voluntary and may be withdrawn at any time; explain how data will be used; ask if there are any remaining questions; provide details of the research team and contact information (check compliance to ethical standards) introduction of the interview introduce yourself as interviewer; explain the purpose of research; show appreciation for participation; give the information letter and ask for informed consent; give a brief outline of the interview topics and indicate the time involved; clarify the roles and expectations of both interviewer and interviewee; ask permission to record the interview; ask if there are any remaining questions body of the interview introduction theme 1 question 1 + (probes) + follow-up questions question 2 + (probes) + follow-up questions question 3 + (probes) + follow-up questions et cetera introduction theme 2 question 1 + (probes) + follow-up questions question 2 + (probes) + follow-up questions et cetera conclusion of the interview indicate the end of the interview; summarise the main points; thank the interviewee for the valuable contribution; reiterate how the data will be used; provide a debriefing of the study; offer to provide a summary of the results after data processing has taken place; ask if there are any questions or comments developing the questions for the interview is an iterative process that is based on the research goals and the specific research questions (emans, 2004). as described above, the goal is to have a clear idea of what you want to know, and then to ask questions that elicit this information from the interviewees. in interview studies, it is also important to make the need for information explicit by defining the variables to be examined. the choice of variables should reflect a thorough understanding of the domain of study and, whenever possible, should be grounded in theory and based on previous research. formulating the possible outcomes of these variables assists in the phrasing of the interview questions and the analysis of the data. in fact, the possible outcomes are the type of answers you would expect to receive from the participants in the interview and, thus, provide clear guidelines for formulating the interview questions. an example of a study on the development of medical expertise in professional practice may illustrate this approach. in this study, our exploratory research questions were: “how do physicians learn in, from, and for their daily work, and how deliberate is this learning process?” (van de wiel et al., 2011). in addition, we wanted to examine differences in workplace learning between three groups of physicians. if we had asked our research questions directly, we would have cued the participants’ connotations about learning. this may have led to the physicians limiting their answers based on their interpretation of the concept, and focusing their attention too much on deliberate processes. based on research on workplace learning, deliberate practice, and selfregulated learning, we identified several relevant work-related activities from which physicians could learn: problem solving, consultation of colleagues, having differences of opinion, explaining to others, seeking and receiving feedback, evaluating performance, professional development activities, and participation in research. these work-related activities allowed us to formulate more specific research questions that helped us to decide what variables to focus on in the interviews. table 4 shows how specific research questions, van de wiel | f l r 103 variables, possible outcomes, and interview questions are related. for example, the specific research question formulated for the work-related activity of problem solving concerns what physicians can learn from the problems they encounter. these problems provide a chance to learn because they need to be solved, or dealt with as part of the job. the two most important variables to examine, therefore, are the types of problems encountered (i.e., what topics can physicians learn about?) and how they solve these problems (i.e., in what way do they learn?). some answers can already be anticipated and listed as possible outcomes in order to guide the formulation of the interview questions on this topic. in our study, an additional question asked for an example of a specific problem and how it was solved in order to illustrate and corroborate the previous answers. a subsequent research question that we addressed was what physicians can learn from solving problems by consulting colleagues. the variables of interest follow-up on the problem solving strategies previously mentioned by the interviewees. as we expected the consultation of colleagues to be an important strategy, we first allowed participants to mention it spontaneously, and then elaborated on the topic in the next interview question. in this study, we started with general questions and moved on to more specific learning activities. moreover, we were careful to introduce the study using the term professional development and not learning, as interview questions prime the way in which participants think and answer. table 4 relationships between specific research questions, variables, possible outcomes, and interview questions. examples are taken from a study on physicians’ learning in the workplace (van de wiel, et al., 2011) specific research questions variables possible outcomes interview questions what can physicians learn from the problems encountered in their practice? types of problems encountered problems in diagnosing patients problems in deciding on the treatment problems in interacting with patients 1. what problems do you usually encounter when diagnosing and treating patients? types of strategies in problem solving asking colleagues for advice searching the literature 1a. what strategies do you use to solve these problems? (if needed can you explain why?) example 1 example 2 et cetera 1b. can you give an example of a problem and how you solved it? what can physicians learn from consulting colleagues? kind of situations in which advice is asked lack of knowledge being uncertain 2. in what situations do you choose to ask a colleague for advice? frequency of advice seeking daily weekly monthly 2a. can you indicate how often you ask for advice? (ask to specify) reasons for advice seeking et cetera direct need for patient care learning 2b. why do you ask for advice? (probe for purpose/goal) the type of variables tapped into by the interview questions constrains the answers that can be expected. with closed questions, answers are restricted to a closed set, such as years of experience, age, or van de wiel | f l r 104 frequency (e.g., question 2a in table 4). in bipolar questions, this set is restricted to the answers yes or no. these questions are mostly asked in order to get a clear answer that can be followed up. in open questions, the interviewees are invited to share their perspective on a topic (e.g., questions 1, 1a, 1b, 2, and 2b in table 4). these questions must be unambiguously formulated, ask about one specific topic at a time, and should not be leading, i.e., should not suggest a particular answer option. interviewers must be neutral and not introduce bias by showing their own opinions, ideas, or feelings, or by suggesting answer options by providing examples or disclosing expectations. the goal is to obtain valid objective data and these types of suggestive questions may influence the thought processes of the interviewee. in answering open questions, interviewees must be encouraged to bring forward what is most prominent in their mind from their own frame of reference. the research team must critically review the interview guide and pilot the interviews with representatives of the target group. this will optimise the informational value of the data collected and prevent the use of restrictive and leading questions. 4.2 the tasks of the interviewer interviewing is a complex task in which the interviewer has to obtain the required information as efficiently as possible. the interviewer must create an interview situation in which the interviewee feels encouraged to speak up, but at the same time, the interviewer must remain in control by asking questions, evaluating the answers, and probing for more elaborate or meaningful answers if necessary (emans, 2004; skopec, 1986). this means that the interviewer must build up a positive relationship with the interviewee to make him or her feel at ease, and be well-prepared to guide the interviewee through the questions in an unobtrusive manner. while it is vital to standardise the situation over different interviews to obtain objective and comparable data, a natural conversation will usually occur if the interviewer is clear, kind, and genuinely interested in a professional way. the interviewer observes the interviewee and listens carefully to the answers by keeping the final goal in mind: gathering valid, complete, relevant and clear answers that reflect the interviewee’s cognitions and provide the information needed for the study. the interaction with the interviewee is steered by verbal and non-verbal probing. nondirective probes serve to fuel the conversation by encouraging interviewees to continue without interrupting their line of thought. effective nondirective probes include: silence accompanied by non-verbal signs of attention; neutral phrases, such as “um hmm”, “oh”, “yes”, “interesting”; rephrasing part of the answer; reflection of feelings that cannot be neglected by showing understanding; making brief summaries of what has been said to check the main points; and general elaborations such as, “can you tell me more?” and, “could you explain further?”. direct probes more actively intervene in the flow of the conversation and are used to focus attention on specific topics. common directive probes include: elaborations that ask for specifications; clarification questions when answers given are imprecise or not well understood; repetition of a question when it appears that the interviewee has not understood the question or avoids answering it; and confrontation when answers seem inconsistent. the interviewer must take the lead and coordinate questioning and probing with appropriate non-verbal behaviour in terms of eye-contact, facial expressions, gestures, and position, to obtain the information required and maintain a positive atmosphere throughout the interview. 4.3 analysing and reporting the data analysing interviews can be relatively straightforward, if during the translation of the specific research questions into variables and interview questions in the preparation phase (see table 4), the anticipated answers were matched to the intended results (emans, 2004). a good match guarantees the internal validity of the study (neuendorf, 2002). the analysis starts with transcribing the interviews verbatim and proceeds by completing the list of possible answers anticipated while preparing the interview with the answers actually given by the interviewees. the more open the questions, the more diverse the answers can be. the more limited the possible answer set, the easier it is to create an overview of the results. for closed questions asking for numerical data, such as work experience, age, and frequency, the results can be van de wiel | f l r 105 processed as they are in quantitative studies, and groups can be compared. for the open questions, researchers need to categorise the answers per variable using content, thematic, or template analysis (neuendorf, 2002; king & horrocks, 2010; brooks, mccluskey, turley, & king, 2015). categorisation starts by developing a coding scheme indicating all emergent answer categories per variable. coding is usually an ongoing and iterative process. a team of researchers reads and codes the transcripts and discusses these codes as new themes and subthemes in the answers emerge and are agreed upon. it is vital to code all relevant parts of participants’ answers. irrelevant answers, i.e., those that do not relate to the research questions, may be categorised under a separate code. this helps the researchers to check those answers and be open-minded to the possibility of finding unexpected emergent themes in the analysis, related to questions answered throughout the interview. the coding can be done manually or can be facilitated by software programs for qualitative analysis. after the coding of all transcripts is complete, the frequencies per code can be listed to indicate the most common answers to each question, as well as the exceptional ones. a good overview shows the main themes and subthemes and helps to uncover patterns in the data as well as any differences between participant groups. in essence, the analysis and reporting of interview data is a process of data reduction and summarisation. illustrative quotes from participants’ answers enrich reporting by giving insight in how participants typically expressed themselves. the description so far may be very abstract, but an example will show that the approach is actually very pragmatic. in the study we conducted about physicians’ workplace learning (van de wiel et al., 2011), we started by categorising the answers per interview question and later grouped them per variable and per specific research question (see table 4) in order to meaningfully report the data. for example, regarding the interview question about what problems participants encountered when diagnosing and treating patients (question 1 in table 4), we found that most participants encountered problems that could be categorised as problems with diagnosis, choosing diagnostic tools and treatments, interaction with patients, and practical organisational issues. participants gave specific examples of some problems, and these were summarised. the ways in which they solved these problems were also categorised and summarised. the answers to these questions were often intermingled, and just as the interviewer had to make sure that all topics were addressed, the researchers had to combine all answers in the analysis. as theory and previous research provided the basis for our specific research questions and variables, we chose to report the data by presenting these as themes and subthemes in the results section. our intention was to summarise what had been said by the participants and to indicate to what extent they concurred and differed on themes and subthemes. in our reporting of results, we also referred to characteristic quotes. analysis was an iterative process in which two coders consecutively categorised sets of data. these categories were reviewed and critically discussed by the research team. in a follow-up study we wanted to quantify to what extent the physicians were deliberately engaged in the work-related learning activities and relate this to other variables (van de wiel, & van den bossche, 2013). in a second step, we therefore analysed the data used in the van de wiel et al. (2011) study from this perspective. the analysis approach taken in this study illustrates how rich verbal interview data are, and how these data can be analysed in different ways depending on the specific research questions. the themes reported in the results of the qualitative analysis of the first study (i.e., work-related learning activities in medical practice) were used as variables to be coded to explore the extent of deliberate engagement in these activities in the second study. in an iterative process, three researchers coded a subset of the interviews until they could reliably distinguish three levels of deliberate practice for each variable: (0) not engaged in learning activity, (1) engaged in learning activities inherent to the job, such as solving a problem, and (2) engaged in deliberate practice as indicated by showing greater motivation and effort for learning to improve competence. clear definitions of codes guided the continuation of coding by one researcher, who consulted the others when in doubt. two themes that pervaded the entire interview were added as variables: reflection on diagnosis and treatment and planning learning activities. the categorisations were reported in a table to allow comparison between the medical residents and the experienced physicians participating in the study. the table displayed the frequencies of both groups’ learning activities, representing the ten variables at each of the three levels of deliberate practice. the outcomes were described in the text and illustrated by quotes. van de wiel | f l r 106 4.4 focus groups focus groups are the most common method used to investigate the opinions and experiences of carefully selected groups of people with regard to all kind of topics and across a wide range of fields (morgan, 1996; krueger & casey, 2015; stalmeijer, mcnaughton, & van mook, 2014). in relation to expertise research, this method is a valuable tool that can be used to gain insight into how different people experience a task, situation, or phenomenon. it can be used both to explore a topic and to compare different groups. only a few initial questions are needed to open up the discussion and encourage participants to share their view on the topic at hand. the advantage of this group interview method is that as people interact, they discuss and analyse the topic from different perspectives, ask each other questions, and may refine their views. a moderator guides the discussion and, as in individual interviews, is responsible for making the participants feel at ease while at the same time probing them to specify their contributions in order to obtain relevant data. a co-moderator usually assists in this process. the discussions are transcribed and then analysed in a bottom-up way by identifying themes and subthemes throughout the text that are relevant to the specific research questions. the researchers also look for relationships between these themes. at least two coders analyse the data independently and then critically review the coding scheme until they reach agreement. the researchers write a summary of the findings that may then be sent to participants to check whether they have suggestions for adjustments that would better represent their discussion. the summary also pinpoints issues that need further clarification and can be brought up at the next focus group. usually 34 rounds of focus group discussions with 5-8 participants in the target group are needed. extra groups are added as long as new information emerges in the discussions, i.e., until data saturation has been achieved. a synthesis of all themes and subthemes coded in the different groups is the basis for reporting the data. quotes can illustrate characteristic utterances. the focus group method has been frequently applied in medicine, often for purposes of medical education (stalmeijer et al., 2014), for example to gain insight into how to improve study arrangements (e.g., de leng, dolmans, van de wiel, muijtjens, & van der vleuten, 2007; van de wiel, schaper, scherpbier, van der vleuten, & boshuizen, 1999), and how the transition from medical school to clinical practice is perceived and can be supported (prince, van de wiel, van der vleuten, boshuizen, & scherpbier, 2004). in medical expertise research, this method has been used to understand how general practitioners approach the diagnostic task and what role non-analytical reasoning plays in their diagnostic process (stolper et al., 2009). 5. verbal protocols verbal protocols add another dimension to the examination of expertise as they deliver verbal data in relation to cognitive processing either during or directly after task performance. interviews are limited in that they provide self-reports by eliciting cognitions about task performance and expert behaviour in a general way. this may induce participants to interpret their own processes from an evaluative perspective and lead to reconstruction and generalisation of their memories of specific task performance (ericsson & simon, 1980; 1993; van someren, barnard, & sandberg, 1994). to minimise these effects, verbal protocol methods aim to capture the processes in performing representative domain tasks by tapping the content of working memory during or immediately after task processing (see figure 1). this diminishes participants’ opportunity to theorise and rationalise what they do and keeps the time delay in processing and verbalising to a minimum. however, requesting participants to verbalise their thoughts while they engage in a task may interfere with natural processing. interference may also occur when participants know beforehand that they will be asked to report back on their task performance. it is necessary, therefore, that studies using verbal protocol methods are carefully designed to capture the natural task processes and problem representations. the formulation of task instructions and selection of tasks and problem situations are critical to ensure that expert behaviour can be demonstrated, and contrasted to novice behavior, in goal-directed, realistic task performance. table 5 van de wiel | f l r 107 shows how the different methods discussed in this article score on the most important criteria in determining whether verbalisation of cognitions impacts the validity of the verbal data gathered. table 5 advantages and disadvantages of the qualitative methods of interviews and verbal protocols in relation to task processing criteria that may impact the validity of verbal data method specific task performance interference with task processing inducing interpretation reflection of working memory interviews +/ focus groups +/ during task think aloud + +/+ dialogues and group discussions + + probing + + +/+/ after task free recall + + explanation + + retrospective reports + + probing + +/+/ two methods that lie at the intersection of interviews and verbal protocols probe participants’ cognitive processing by means of questions during or after specific task performance. as depicted in table 5, one advantage of these probing methods is that cognitions can be examined in direct relation to the task at hand, yielding more specific and precise information than interviews. the disadvantage of probing during task performance is that questioning interrupts participants’ thoughts and actions, altering their normal task processing. this may even encourage participants to adopt an interpretative mind-set (ericsson & simon, 1980, 1993; van someren et al, 1994). probing after task performance overlaps with interviews that focus on specific tasks, events, scenarios, and examples as used in knowledge elicitation techniques (hoffman et al., 1995; hoffman & lintern, 2006; shadbolt & smart, 2015), as well as with some types of explanation protocols and retrospective reports. this method can deliver very valuable information regarding the research questions. as discussed in the section on interviews, this is dependent upon the way the questions are phrased and embedded within the interview guide. in expertise research, data gathered by interviews and verbal protocols may complement each other as participants’ overall cognitions about their expertise domain can be combined with assessments of task-specific processing and outcomes. as expertise is domain and task specific, the selection of participants, tasks, and particular problems to solve are crucial steps in setting up verbal protocol studies that examine expertise (see table 2). both task characteristics and the experts’ knowledge and experience determine the cognitive processes and evolving representations in task performance (see figure 1). if a representative domain task has been chosen, e.g., diagnosis in medicine, the next step is to decide upon the problems to be solved and the presentation format. in medicine, patient cases that reflect a consultation with a physician can, for example, be summarised in a brief description and might be supplemented with information from the patients’ record to mimic the situation in real practice. most important here is that the experimental situation captures the essentials of the task and the problem under investigation. in order to identify characteristics of expert behaviour, cognitive processing, and outcomes, as well as differences with other groups of participants, problems may be presented at various levels of difficulty. task materials and conditions may also be manipulated to van de wiel | f l r 108 investigate the effects of changing normal processing. in routine problems, experts are expected to automatically activate the right knowledge, but in more difficult problems and under special conditions, they will coordinate automatic thoughts with analytical thinking. the amount of problem information presented and the time in which the information becomes available also impact cognitive processing: the more information and time involved, the more coordination is required, and the more elaborate and deliberate thinking will be. factors such as these that are related to the nature of the task performed are also emphasised in the cognitive continuum theory (custers, 2013; hamm, 1988; hammond, hamm, grassia, & pearson, 1987). this theory situates most thinking somewhere in between intuition and analysis, which are conceptualised as two ends of a continuum of cognitive processing. where on the continuum thinking falls, depends on specific task characteristics. this theory provides a framework for analysing tasks in relation to processing requirements. verbalisation theory, in addition, provides a framework for analysing in what way verbalisation of the content in working memory influences task processing (ericsson & simon, 1980, 1993). if information and knowledge are already represented in a verbal format, they only have to be vocalised. if, however, they are represented in a visual or motor modality, encoding of the representations into verbal format is required. this demands extra cognitive processing that may alter the way in which the task normally proceeds. in preparing verbal protocol studies, both the task and problem characteristics need to be well thought out, and piloted to ascertain that the research questions can be answered. these characteristics may also be manipulated to test specific hypotheses. when combining verbal protocol methods in one study, researchers need to be careful that instructions given and procedures followed do not influence subsequent task processing. in conducting verbal protocol studies, task performance must be monitored and verbalisations need to be recorded and transcribed verbatim (chi, 1997; ericsson & simon, 1980, 1993; van someren et al, 1994). the outcome measures of the task comprise the dependent variables that are to be used to assess differences in expertise between groups or improvement over time. these may be complemented with quantitative processing measures, such as time used to solve or to explain a problem. the verbal protocols provide rich data on the underlying cognitions in task performance that need to be coded and interpreted to answer the research questions. analysing the data from a clear perspective enhances the acquisition of valuable, objective information and may reduce the workload. a coding scheme depicting the variables (e.g., knowledge used) and the related coding categories (e.g., biomedical and clinical knowledge), including definitions and examples of utterances per category, needs to be developed to guide data analysis (chi, 1997; van someren et al., 1994). the coding scheme can be based on theory, previous research, and/or cognitive task analysis or can be developed in a bottom up way (chi, 1997; hsieh & shannon, 2005; van someren et al., 1994). an important issue to decide upon when developing the coding scheme is the unit of analysis used, as this may vary from a word, single unit of information or proposition, to a clause, reasoning chain, or turn in a discussion, and is to be guided by the research goals. particularly in theory-driven research, it is good practice to develop the coding scheme using a set of pilot protocols and to test the hypotheses on another sample of protocols. in more exploratory research, the development of coding schemes may lead to theory-building and hypotheses generation. during analysis, the audioor video-recordings can be listened to or watched, if it is necessary to improve understanding and interpretation. the coding of verbal protocols is an iterative process, in which coders must seek agreement in order to obtain reliable data. although the analysis of verbal protocols is qualitatively in nature, the data may be quantitatively described by tallying the number of utterances per coding category (chi, 1997; krippendorf, 2012; neuendorf, 2002). such quantitative descriptions help to create an overview, reduce subjectivity in interpretation, and find and report meaningful patterns in the data, while protocol fragments show examples of the coded categories in relation to each variable. recapitulating the steps to be taken in preparing expertise research using qualitative methods (see table 2) clearly shows how important it is that the study design is guided by the research questions and by the strengths and weaknesses of the methods employed. understanding the ways in which various methods affect cognitive task processing in data collection helps to establish what interview or verbal protocol method to use. in designing verbal protocol studies, the next critical step is to choose task characteristics and van de wiel | f l r 109 requirements in a manner that can reveal expert behaviour under conditions reflecting or manipulating the essence of the task. this step is closely tied to the development of a coding scheme in which it is anticipated how each variable of interest can be measured in the verbal protocols. the final step of communicating the study to participants must comply with general research guidelines, as described in the outline of the interview guide (see table 3). in the following sections, the five verbal protocol methods of thinking aloud, dialogues or group discussions, free recall, explanation, and retrospective reports will be discussed. drawing from research on medical and visual expertise, examples will be provided of specific aspects of the design and analysis of each method. 5.1 think-aloud protocols the think-aloud method has been frequently used to investigate the knowledge and processes in task performance and is well-documented in the literature (chi, 1997; ericsson, 2006b; ericsson & simon, 1980, 1993; hassebrock & prietula, 1992; van someren, et al., 1994; shadbolt & smart, 2015). participants are asked to say everything that comes to their mind while they engage in a task. the rationale behind this method is that the verbalised thoughts reflect the evolving mental representations in working memory during task performance (see figure 1). verbalisation of thoughts does not change the sequence of actions, and usually does not disturb the cognitive processes engaged in, but merely slows down these processes. in some tasks, however, verbal encoding and vocalisation of information interferes with natural task performance. this is because the increased load on working memory makes it difficult to keep up with the flow of information that needs to be attended to, i.e., cognitive processing cannot be slowed down to successfully accomplish the task. verbalising is easiest when thoughts are already verbally represented, and requires extra cognitive effort in visual and motor tasks. in highly skilled and expert performance, cognitive processes are largely automated and think-aloud protocols will only reveal those thoughts that consciously come to mind. this method, then, shows what knowledge is activated, which parts of cognitive processing are automated, and when deliberate, analytical thinking is involved in specific groups, tasks and problem situations. natural thinking in cognitive tasks can be easily disturbed by task instructions. it is therefore important that these instructions are formulated in such a way that participants are not tempted to explain or justify what they do to the experimenter. practicing the think-aloud procedure with participants maximises the chance of obtaining valid data. when verbal materials are presented, these need to be read out loud to facilitate expressions of thoughts during information intake and signal the cues that trigger these thoughts. the data can be systematically analysed and reported in many different ways, depending on the research goals. in medicine, the think-aloud procedure has, for the most part, been applied in diagnostic tasks to reveal the knowledge and reasoning involved. hassebrock & prietula (1992) made a detailed analysis of diagnostic reasoning, focusing on knowledge states, conceptual operations, and lines of reasoning to explore different types of diagnosis, the cognitive activities engaged in (e.g., data examination, data explanation, hypothesis evaluation, meta-reasoning), and the links between patient cues, pathophysiological conditions and hypotheses, respectively. the coding scheme they used was very elaborate allowing for precise statements on the knowledge representations and the (causal) lines of reasoning in diagnosing cases of congenital heart disease. these statements could be compared with expert models of reasoning on the topic. it is a good illustration of how cognitive task analysis may guide the interpretation of qualitative data. however, although it is tempting to carry out this type of detailed analysis when data are so rich, it might be more practical to focus on some of the main inferences made, as in a study conducted by gilhooly, et al. (1997). they asked participants to diagnose eight ecg traces while thinking aloud. a computer program first listed all of the technical terms used by participants and an expert categorised these words into three coding categories. subsequently, the program counted the number of words indicating trace characteristics, clinical inferences and biomedical inferences. this method enabled the researchers to compare the knowledge used in visual diagnosis between different expertise groups and across ecg traces varying in difficulty, as well as between the think-aloud protocols and explanation protocols that were collected a week later. in this special issue, helle (2017) discusses the relationship between eye tracking data and verbal data. van de wiel | f l r 110 5.2 discourse analysis of dialogues and group discussions recording the natural discourses of collaborating experts or of students and teachers in education is a valuable method that can be used to gain insight into online group decision making, problem solving, and learning (chi, 1997; salas et al., 2006). it shows what issues participants attend to, how they regulate their discussions and task processes, how they interact, what kinds of knowledge and strategies they use, what they might learn, and what they can improve. the main task of researchers is to select a representative sample of meetings in which participants engage in knowledge sharing as part of their work or training, audioor video-record these meetings, transcribe them verbatim, and then analyse what has been said. participants might notice that they are being observed in the very beginning, but as soon as they start their tasks, they will proceed as usual. this method delivers very rich data, and reducing the data to manageable proportions in coding is a challenge, and a process that will be guided by the research goals. analysing and describing the data at different levels of detail, for example, in terms of the sequence of actions taken in discussing a patient, the type of patient problems discussed, and the content of a sample of these discussions, helps to create overview and pinpoint the issues of interest. patel and colleagues investigated team problem solving and decision making in the complex environment of hospital intensive care units. in their research, work domain analysis and communication patterns between attending physicians, residents, nurses, and consulted specialists provided an invaluable framework that could be used to analyse the content of the individual contributions and understand the processes in context (patel & arocha, 2001; patel, kaufman, & magder, 1996). the researchers particularly focused on discussions that took place during morning rounds. during these rounds the team discusses the patients on the ward, evaluating each patient in detail and planning future actions. they enriched their analysis by examining complementary data from patient charts, recordings of morning lectures, and interviews with participants. they segmented the protocols into episodes that distinguished between subsequent phases in the discussions, and further segmented these phases into thematic idea units. categorisation of these units classified the contributions at four knowledge levels: (1) observations of patient signs, (2) findings referring to clinically significant clusters of observations, (3) facets referring to pathophysiological states or broad categories of disease, and (4) diagnoses referring to clinical conclusions. they also categorised three types of decisions at the level of: findings, actions to be taken in patient management, and assessments of the overall state. the results showed how contributions were distributed over the different participants, how content changed over three consecutive days of patient care, how interactions, content, and reasoning differed in two types of intensive care units, and what episodes were particularly useful for expertise development. the data were both qualitatively and quantitatively described. this research in the tradition of naturalistic decision making (klein, 2008) provides a good example of how expertise can be examined when it is distributed over multiple agents in a complex, real-life, dynamic environment. two other studies using this method have analysed the discussions in learning situations to examine the application of biomedical and clinical knowledge in problem-based learning tutorials with real-patients (diemers, van de wiel, scherpbier, heineman, & dolmans, 2011), and the topics discussed in tutorial dialogues on diagnostic reasoning of trainees and their supervisors in general practice (stolper et al., 2015). in the study conducted by diemers and colleagues, a purposive sample of tutorial group discussions was divided into a preparation and a reporting phase. based on a technique of proposition analysis for medical protocols (patel & groen, 1986), all transcripts were segmented into small meaningful information units or propositions. propositions connect two concepts by a qualifier, such as “a pseudo polyp is characteristic of ulcerative colitis”. guided by previous research and pilot interviews, we coded the propositions as patient information, formal clinical knowledge, biomedical knowledge, informal clinical knowledge, procedural information, and other information, and also indicated if the proposition was put forward by the tutor or a student. in this way, we could compare both the number of propositions per coding category, and the number of propositions contributed by tutors and students in both phases. to analyse the function of biomedical knowledge, we categorised what the biomedical lines of reasoning in the protocols explained, and found that van de wiel | f l r 111 they mostly linked underlying mechanisms of disease to clinical features of patients, as intended by the educational format. in the study conducted by stolper and colleagues, a representative sample of tutorial dialogues was taken and segmented based on turns in the conversation and also on content changes. the coding proceeded in a bottom-up and iterative way, but was informed by the researchers knowledge of diagnostic reasoning and research goals, which helped to characterise the topics of discussion. as trainees usually presented several patient cases for discussion, we differentiated between a reporting and an analysis phase. segments were double-coded to indicate the contributions of trainees and supervisors. the number of words per code was counted so that we could examine to what extent the different topics were discussed and by whom. in line with the research questions, the data for all main coding categories were reported in tables, and specified in more detail for the categories of diagnostic reasoning and gut feelings. in addition, the tutorial dialogues, diagnostic reasoning, and the way in which gut feelings featured in the dialogues were described in qualitative terms. in both studies, the methods of analysis yielded rich but precise data that could be used to quantitatively compare the attention paid by the participant groups to the coding categories and the variables of interest, and qualitatively describe, interpret and illustrate these variables. 5.3 free recall protocols free recall is a classical method used to study problem representation underlying task performance (chi, 2006b; chi et al., 1988; feltovich et al., 2006). the most well-known example of a free recall study is the work on chess expertise of de groot (1946/1978), who asked participants to think of a move in an actual game position before they had to recall the position. chess masters selected the best moves and recalled more chess pieces than other players. this was explained as being the result of their better overall conception of the problem. masters could grasp the problem at a high level in a very short time showing their expert knowledge of chess. the memory paradigm used, i.e., asking the chess players to recall briefly presented chess positions, has since been adopted to chart the problem representations and underlying knowledge structures by reviewing how individual chess pieces are chunked together in memory (chase & simon, 1973). the rationale behind the free recall measure is that the content of working memory is retrieved immediately after task performance (see figure 1). therefore, matching the goals of processing in the experimental situation to the real task is a prerequisite consideration in the development of task instructions (ericsson & smith, 1991; ericsson, patel & kintsch, 2000). this is underlined by levels of processing theory showing that meaningful processing gives the best results due to the richer connections in memory (craik, 2002). the use of the free recall method in medical expertise research clearly shows that results are influenced by task processing conditions. in contrast to the superiority of expert memory for meaningful materials found in a wide variety of domains (chi, 2006b; feltovich et al., 2006), in medicine, findings have been less straightforward (norman, et al., 2006). in most circumstances, experienced physicians appear to represent patient cases in a more condensed way than advanced students, if they process them for diagnosis (schmidt & boshuizen, 1993; de bruin, van de wiel, rikers, & schmidt, 2005). a sequence of recall studies manipulating case materials and task instructions suggests that clinical case processing of experts using patient descriptions is rather robust. however, memorisation instructions, perceiving the task as a memory task, or instructions for elaborate processing may enhance their recall (de bruin, et al., 2005; van de wiel, schmidt, & boshuizen, 1998; van de wiel, ploegh, boshuizen, & schmidt, 2005; wimmers, schmidt, verkoeijen, & van de wiel, 2005). if lab data were processed without further patient information and under elaborate problem formulation conditions, experts outperformed students in recalling the lab data after this analytical diagnostic task (norman et al., 1989; wimmers et al., 2005). if medical students knew in advance that they would be asked to recall a patient case (intentional recall condition) they recalled more case information than if they did not know this (incidental recall condition) (van de wiel et al., 2005). in conclusion, the method has some drawbacks that are hard to control in experimental research, if only because more than one patient case should be presented. a lesson learned is that diagnostic tasks should be presented in a realistic way by aligning processing goals and providing the patient information that would be available van de wiel | f l r 112 in practice. for research in visual expertise this means that both the image and the information physicians have before interpreting the image should be presented (hatala, norman, & brooks, 1999; kulatungamoruzi, brooks, & norman, 2004). to capture the nature of expertise, the information provided in patient case descriptions needs to be presented in a standard order and phrased as it is communicated in practice, either in the words of patients or of colleagues, and interpretation must be left to the participants. the manipulation of task performance conditions can be an important strategy to further uncover expert knowledge structures. a good example is constraining the time in which participants have to process case materials, as it might be expected that this will have a lower impact on experts than on less advanced participants (schmidt & boshuizen, 1993; van de wiel et al., 1998). in addition, the method of free recall should be supplemented with other verbal protocols in order to corroborate findings. representations of patient cases, for example, might be investigated by asking physicians how they would summarise or characterise a case or describe what information was critical for their diagnosis. in visual domains, recall protocols of the images presented can be obtained by asking participants to describe what they saw, and this can be supplemented with instructions to indicate the relevant features on the image, or draw the image (e.g., lesgold et al., 1988; gilhooly et al., 1997). in this way, the representation of features recognised can be separated from the interpretation of the pattern of features in diagnosis. analysis of free recall protocols, characteristic summaries, or protocols with critical cues can proceed in a straightforward way using proposition analysis. the number of propositions in the protocol that match the propositions in the case materials is counted. the number of summaries, i.e., inferences referring to more than one case proposition, may be counted separately to show to what extent participants represent the case information at a higher interpretative level. an example of a summary encompassing four propositions is “auscultation reveals mitral valve insufficiency”, which summarises the more detailed information, “auscultation reveals a holosystolic murmur at the apex radiating towards the axilla” (van de wiel et al., 1998). the summaries can be provided at different levels of detail, varying from interpretation of lab data to the encapsulation of data into pathophysiological mechanisms or diagnostic labels. moreover, the order of the information recalled may also reveal how the problem and underlying knowledge are represented in memory (e.g., boshuizen & claessen, 1985). the research questions will determine what to focus on in designing the study and analysing the data. 5.4 explanation protocols in medical expertise research, the post-hoc explanation method was introduced by feltovich and barrows (1984), and further developed by patel and groen (1986) to gain insight into the knowledge used in the diagnostic process and the causal lines of reasoning. participants are asked to provide a pathophysiological explanation of the signs and symptoms in a patient case that they have processed for diagnosis. just as in free recall, it is assumed that the knowledge activated in case processing will be retrieved when providing the explanation (see figure 1). although, in general, this seems to be the case for expert processing, explanations may be influenced by the knowledge available to understand the case materials. as the knowledge of novices and intermediates, for example, is fragmented and less coherently organised, they tend to elaborate on their knowledge when explaining the recalled case features (boshuizen & schmidt, 1992). it is clear, however, that explanations of patient cases can reveal the knowledge participants use in linking case information, pathophysiological mechanisms, and disease. explanations allow for analysis of both the content and the structure of knowledge, i.e., how knowledge elements are related (chi, 1997). the explanation method is easy to use, but may be labour-intensive to analyse. based on two studies in which we applied this method to compare the knowledge structures of medical students with those of experienced physicians (van de wiel et al., 2000), and to examine the knowledge development of medical students over the period of their course (diemers, et al., 2015), a detailed account of how to proceed with analysis and reporting are provided. for these research goals of comparing knowledge between groups and over time, the first step in analysis is to develop a model explanation. a model explanation links the signs van de wiel | f l r 113 and symptoms in a case to the diagnosis via the most important biomedical and clinical concepts explicating the disease processes in a network representation much like a concept map. the model explanation represents a causal model of a disease and must be developed in close collaboration with experts. concept mapping is a particularly useful tool to elicit experts’ knowledge for this purpose (hoffman & lintern, 2006; shadbolt & smart, 2015). the participants’ written explanation protocols are then translated into networks of linked concepts and compared to the model explanation network. the network representations enable an assessment of explanation quality in terms of correctness of both concepts and links. in our studies, the variables we coded in the protocols included the total number of concepts used in explanations, specified as the number of model concepts, alternative concepts, detailed concepts, and wrong concepts, as well as the total number of links, specified as model links, alternative links, detailed links, wrong links and shortcuts in reasoning. as measures of quality, we used the percentage of model concepts (relative to the total number of concepts used in the protocols) and the percentage of model links (relative to the total number of links used in the protocols). in the diemers et al. study, we also counted the number of biomedical and clinical concepts used. the outcomes of the variables were depicted in graphics or tables to support interpretation. the method provides detailed insight into the quality of knowledge structures, (causal) lines of reasoning, and the nature of the knowledge used. moreover, it allows meaningful comparisons between groups and conditions. the combination of variables allows consistent patterns to be found in the data. explanation protocols provide a rich source of clear examples of expert and flawed knowledge and reasoning. two studies asking for explanations in the visual domain demonstrate the use of a keyword technique to characterise the type of utterances made by participants (gilhooly et al., 1997; jaarsma, jarodzka, nap, van merriënboer, & boshuizen , 2014). gilhooly and colleagues were interested in whether cardiologists had more biomedical knowledge available to explain ecg traces than less experienced participants. participants could look at the the ecg traces while providing the explanations. a computer program analysed how many words in the protocols indicated trace characteristics, clinical inferences, and biomedical inferences (similar to the analysis of the ecg think-aloud protocols that were collected a week before the explanations). this procedure resulted in very clear data which revealed the expected expertise effect. in the study conducted by jaarsma and colleagues, participants were asked to give an explanation for their diagnosis after they had seen a microscopic image for two seconds. two coders categorised the words in the combined protocols of 20 explanations. some of these categories were based on previous research and others emerged from the data. the number of words in each coding category characterised patterns of reasoning processes and the types of knowledge used by each of the three expertise groups. these examples show that explanation protocols can be reliably used to examine expertise differences in the visual domain, and that indicative words can be used as practical units of analysis. providing explanations can be a task in its own right (chi, 1997), as participants can for example be asked to explain a concept. this use of explanation protocols overlaps with knowledge elicitation techniques. explanations collected in medical expertise research can be analysed from a psychological perspective to reveal differences between expertise groups in terms of the content and organisation of their knowledge. for example, in a study on the explanation of clinical concepts, we analysed the elaborateness, quality, and fluency with which explanations were provided (van de wiel, boshuizen, schmidt, & schaper, 1999). the researchers told participants that they were interested in what they knew about certain concepts, and asked them to explain 20 concepts to the experimenter in approximately 2 minutes. the instructions were carefully crafted to ensure that participants communicated all they knew during the full time slot, indicated when they were not sure, and also explained how they could recognise a particular concept in patients. throughout the experiment, the experimenter guided the elicitation process according to this procedure in order to increase data quality. for each concept, a model explanation was constructed. this highlighted model concepts and referred to a definition of the concept, the major causes and clinical consequences, and the essential pathophysiological mechanisms of disease. the model explanations were based on medical literature and checked by medical specialists. the participants’ transcribed explanation protocols were segmented into meaningful information units that were then coded based on content. elaborateness of the protocols was measured by counting the total number of medical concepts used and specified per category: definition, van de wiel | f l r 114 cause, clinical knowledge, pathophysiological knowledge, and therapeutic knowledge. quality of the explanations was measured by comparing the explanation protocols to the model explanation revealing the number of model concepts and imprecise expressions used, and the number of clinical concepts that were unknown. accessibility of knowledge was operationalised by the fluency with which participants provided their explanations and measured by the number of times they abruptly changed the subject, used thinking pauses, stumbled, thought aloud, or referred to their lack of knowledge. coding was very precise and laborious and, as a result, reliable. the results could be clearly depicted in a table and gave good insights into the availability and accessibility of knowledge in the three expertise groups. the data were also interpreted in a qualitative manner for medical education purposes to illustrate major misconceptions in concepts that were weakly explained. this method may be particularly useful to chart knowledge and misunderstandings in complex domains, including visual expertise (e.g., the feature description test as used by kok et al., 2013), and to effectively design instruction. if explanations are requested in a written format, the method is more feasible and shows availability of knowledge but not the fluency by which it is accessed. the method can also be used for assessment of and feedback to students. 5.5 retrospective reports the questions asked to participants in retrospective reports of problem solving are similar to those that can be asked in interviews. examples of such questions are: “how did you solve the problem?” or “what did you think while problem solving?”. the important difference is that the delay between problem solving and answering these questions is minimised in retrospective reports that are generated immediately after task performance. this delay can undermine the validity of the answers because people tend to present their thought processes as being more coherent and intelligent than they are, and reconstruct their memories of the problem-solving processes based on the outcomes (ericsson & simon, 1980, 1993; van someren et al., 1994). the effect of delay may also exist in retrospective reports if the task lasts longer than 10 seconds, as, after that time, the sequence of thoughts is no longer readily available in working memory (ericsson & simon, 1993). another threat to the validity of the data is if, during questioning, researchers probe participants to give post-hoc rationalisations of what they did, e.g., by asking: “can you explain how you proceeded in solving the problem?”, “what general approach did you use in problem solving?” or “why did you solve the problem in this way?” if, for example, participants did not use a general, structured approach, but rather solved the problem by trial and error, they may feel embarrassed to say so. however, when interviews are well-prepared and conducted, the interviewer can tap into relevant knowledge and strategies the participants may have used by emphasising that there are no right or wrong answers, encouraging participants to think aloud, and probing. as demonstrated by knowledge elicitation practices, researchers can scaffold participants in a collaborative process to articulate what they know, even if this has not been articulated before (hoffman et al., 1995; hoffman & lintern, 2006). a good example of such a method is cued or stimulated recall, in which participants are walked through the task by watching and/or listening to a recording of the problem-solving process while expressing what they remember of their thoughts at specific points. this method may be particularly useful if concurrent thinking aloud is not possible because the incoming information that needs to be attended to is presented faster than it can be verbalised. the recording then provides cues in working memory that trigger retrieval of the cognitive processes and mental representations involved in task performance (see figure 1). for this reason, the method might be well suited to examine processing and interpretation of visual data. to obtain valid data, retrospective questions or reports should be gathered in representative samples of participants and validated by the results of other methods, such as thinking aloud. to analyse the data, researchers can develop a coding scheme (just as in interviews or think-aloud protocols) and report the data in both qualitative and quantitative ways. the following two examples will show how the method of stimulated recall can be effectively applied to gather additional data on participant’s thoughts while accomplishing a task. in a study on problem analysis in a tutorial group, the recorded discussion was presented to each individual group member immediately after the meeting to elicit their memories of their thinking during the discussion (de grave, van de wiel | f l r 115 boshuizen, & schmidt, 1996). participants could stop the videotape at any time to recall what they had in mind while the others were talking. the research goal was to investigate whether problem-based learning leads to conceptual change when students participate in small group discussions. the data showed that this stimulated recall procedure provided further insight into the prior knowledge invoked, the changes in reasoning based on the group’s theory building process, and the metacognitive reflections engaged in, as well as participants’ thoughts about the group process. a coding scheme guided analysis of both the group interaction and the stimulated recall protocols, enabling comparisons between the number of clauses in each coding category across the two types of protocols. a temporal analysis of the categories of theory building and meta-reasoning showed how these thinking processes interacted and how conceptual change came about. this was illustrated with the stimulated recall protocol of one student. this example shows that stimulated recall can be a very informative means of elucidating covert thinking processes in group discussions. in the other example, retrospective reports of electrical circuit problem solving were elicited by presenting the participant’s eye fixations and mouse-keyboard operations on the computer screen in the original task (van gog, paas, van merriënboer, witte, 2005). this may be a very good way of capturing the cognitive processing in perceptual-motor tasks. the goal of the study was to compare the results collected when the different methods of thinking aloud, retrospective reporting, and cued retrospective reporting were used. participants were asked to tell what they were thinking during the task while watching the record of their eye movements and actions. unfortunately, the record was replayed at the same speed as in task performance, and participants could not stop the recording to share their thoughts. this was probably the reason why cued retrospective reporting did not deliver more information than thinking aloud for three of four coding categories. task performance can slow down during thinking aloud, but this was not possible during the cued retrospective reporting procedure used in this study. in conclusion, the examples show that stimulated recall may provide valuable information about the knowledge and reasoning involved in task performance when it is well-designed and implemented. 6. conclusion the methods of interviewing and collecting verbal protocols provide rich data to examine expertise and expertise development from different perspectives in all kinds of domains. interviews are needed in the exploratory phase of research to gather information on the tasks and the problem situations to be investigated. they can also be used to obtain objective data to compare different expertise groups on targeted topics. in verbal protocol studies, knowledge, mental representations, and reasoning are examined in direct relation to representative tasks in which experts should outperform less advanced or experienced participants. the protocols collected as a result of thinking aloud, dialogues, and group discussions tap into the concurrent cognitive processes of experts and students during task performance, whereas verbal protocols of free recall, explanations, and retrospective reports are gathered after the task has been performed. all methods have their advantages and disadvantages in terms of the validity of the data obtained. to grasp the full nature of domain expertise, different methods should be applied in order to complement one another. sometimes, different verbal protocols can be gathered in the same study, for example, thinking aloud while interpreting patient information and making a diagnosis, incidental recall of the last case presented, and explanation of the case materials a week later (e.g., gilhooly et al., 1997). the selection of tasks and case materials is crucial and should be well-aligned with practices in real-life settings in order to capture expertise. varying task difficulty is an important manipulation that can be used to investigate in what situations, and for which participant groups, automatic processing falls short and deliberate thinking is involved. the interactions between knowledge, cognitive processing, and task characteristics are key to the understanding of expertise. interviews and cognitive task analysis play an important role in identifying the relevant interaction patterns, before research using verbal protocols can be designed. the formulation of interview questions and task instructions need careful attention in order to safeguard data validity when examining the underlying cognitions of expertise. analysis of interviews and verbal protocols is usually van de wiel | f l r 116 labour intensive, and reducing analysis to manageable proportions is best guided by the specific research goals. quantifying the qualitative variables identified in developing the coding schemes provides an overview that helps the researchers to find, interpret, and report main patterns in the data. all methods described in this article may contribute to further developing different domains of expertise. they can also be used to examine visual expertise, as professionals are usually able to communicate what they see, what conclusions they come to, and what they plan to do when collaborating with colleagues and teaching students. keypoints the research questions, available theory, and previous findings are the starting point for coding schemes used to analyse interviews and verbal protocols. interviews can be used to obtain relevant information in an objective way to compare groups with different levels of expertise. interviews provide indispensable insight in the tasks and materials to be selected for research gathering verbal protocols. as expertise emerges from expert knowledge in interaction with the task and the problems to be solved, the selection of participants and materials as well as the design of task instructions are critical in preparing verbal protocol studies. different methods of research, such as interviews and verbal protocols, need to be combined to fully grasp expertise in complex domains, such as visual expertise. acknowledgments i am very grateful for the valuable feedback i received from my colleagues, els boshuizen and fleurie nievelstein, and two anonymous reviewers on earlier drafts of this article. references anderson, j. r. (1996). act: a simple theory of complex cognition. american psychologist, 51(4), 355. doi:10.1037/0003-066x.51.4.355 bartram, d. (2008). work profiling and job analysis. in n. chmiel (ed.), an introduction to work and organizational psychology: a european perspective (2nd ed.) (pp. 3-28). oxford, uk: blackwell. brooks, j., mccluskey, s., turley, e., & king, n. (2015). the utility of template analysis in qualitative psychology research. qualitative research in psychology, 12(2), 202-222. doi:10.1080/14780887.2014.955224. boshuizen, h. p. a., & van de wiel, m. w. j. (1998). multiple representations in medicine: how students struggle with it. in m. w. van someren, p. reimann, h. p. a. boshuizen, & t. de jong (eds.). learning with multiple representations. amsterdam, the netherlands: elsevier. boshuizen, h. p. a., van de wiel, m. w. j., & schmidt, h. g. (2012). what and how advanced medical students learn from reasoning through multiple cases. instructional science, 40(5), 755-768. doi:10.1007/s11251-012-9211-z. boshuizen, h. p. a., & van de wiel, m. w. j. (2014). expertise development through schooling and work. in a. littlejohn, & a. margaryan (eds.), technology-enhanced professional learning: processes, practices and tools (pp. 71-84). new york, ny: taylor & francis/routledge. http://psycnet.apa.org/doi/10.1037/0003-066x.51.4.355 http://dx.doi.org/10.1007/s11251-012-9211-z van de wiel | f l r 117 chase, w. g., & simon, h. a. (1973). perception in chess. cognitive psychology, 4(1), 55-81. chi, m. t. h., (1997). quantifying qualitative analyses of verbal data: a practical guide. the journal of the learning sciences, 6(3), 271-315. doi:10.1207/s15327809jls0603_1 chi, m. t. h. (2006a). two approaches to the study of experts’ characteristics. in k. a. ericsson, n. charness, p. j. feltovich, & r. r. hoffman (eds.), the cambridge handbook of expertise and expert performance (pp. 21-30). new york, ny: cambridge university press. chi, m. t. h. (2006b). laboratory methods for assessing experts’ and novices’ knowledge. in k. a. ericsson, n. charness, p. j. feltovich, & r. r. hoffman (eds.), the cambridge handbook of expertise and expert performance (pp. 167-184). new york, ny: cambridge university press. chi, m. t. h., bassok, m., lewis, m. w., reimann, p., & glaser, r. (1989). self-explanations: how students study and use examples in learning to solve problems. cognitive science, 13(2), 145-182. doi:10.1016/0364-0213(89)90002-5 chi, m. t. h., feltovich, p. j., & glaser, r. (1981). categorization and representation of physics problems by experts and novices, cognitive science, 5, 121-152. doi:10.1207/s15516709cog0502_2 chi, m. t. h., glaser r., & farr, m. j. (1988). the nature of expertise. hillsdale, nj: lawrence erlbaum. chipman, s. f., schraagen, j. m., & shalin, v. l. (2000). introduction to cognitive task analysis. in j. m. schraagen, s. f. chipman, & v. l. shalin (eds.), cognitive task analysis (pp. 3-23). mahwah, nj: lawrence erlbaum. claessen, h. f. a., & boshuizen, h. p. a. (1985). recall of medical information by students and doctors. medical education, 19(1), 61-67. doi:10.1111/j.1365-2923.1985.tb01140.x craik, f. i. (2002). levels of processing: past, present... and future? memory, 10(5-6), 305-318. doi:10.1080/09658210244000135 custers, e. j. (2013). medical education and cognitive continuum theory: an alternative perspective on medical problem solving and clinical reasoning. academic medicine, 88(8), 1074-1080. doi:10.1097/acm.0b013e31829a3b10 deakin, h., & wakefield, k. (2014). skype interviewing: reflections of two phd researchers. qualitative research, 14(5), 603-616. doi:10.1177/1468794113488126 de bruin, a. b. h., van de wiel, m. w. j.,. rikers, r. m. j. p, & schmidt, h. g. (2005). examining the stability of experts’ clinical case processing: an experimental manipulation. instructional science, 33, 251-270. doi:10.1007/s11251-005-3598-8 de grave, w. s., boshuizen, h. p. a., & schmidt, h. g. (1996). problem based learning: cognitive and metacognitive processes during problem analysis. instructional science, 24(5), 321-341. doi:10.1007/bf00118111 de groot, a. d. (1978). thought and choice in chess. (2nd ed). the hague, the netherlands: mouton. (original work published in 1946) de leng, b., dolmans, d., van de wiel, m. w. j., muijtjens, a., & van der vleuten, c. (2007). how video cases should be used as authentic stimuli in problem-based medical education. medical education, 41, 181-188. doi:10.1111/j.1365-2929.2006.02671.x diemers, a. d., van de wiel, m. w. j., scherpbier, a. j., baarveld, f., & dolmans, d. h. (2015). diagnostic reasoning and underlying knowledge of students with preclinical patient contacts in pbl. medical education, 49(12), 1229-1238. doi:10.1111/medu.12886. diemers, a. d., van de wiel, m. w. j, heineman, e., scherpbier, a. j. j. a., & dolmans d. h. j. m. (2011). pre-clinical patient contacts and the application of biomedical and clinical knowledge. medical education, 45, 280-288. doi:10.1111/j.1365-2923.2010.03861.x. dubois, d., & shalin, v. l. (2000). describing job expertise using cognitively oriented task analyses (cota). in j. m. schraagen, s. f. chipman, & v. l. shalin (eds.), cognitive task analysis (pp. 4155). mahwah, nj: lawrence erlbaum. elstein, a. s., & schwarz, a. (2002). clinical problem solving and diagnostic decision making: selective review of the cognitive literature. british medical journal, 324, 729-732. elstein, a. s., shulman, l. s., & sprafka, s. a. (1978). medical problem solving: an analysis of clinical reasoning. cambridge, ma: harvard university press. http://dx.doi.org/10.1016/0364-0213(89)90002-5 van de wiel | f l r 118 emans, b. (2004). interviewing: theory, techniques and training. groningen, the netherlands: woltersnoordhoff. ericsson, k. a. (1996). the acquisition of expert performance: an introduction to some of the issues. in k. a. ericsson (ed.), the road to excellence: the acquisition of expert performance in the arts and sciences, sports and games. mahwah, nj: lawrence erlbaum. ericsson, k. a. (2004). deliberate practice and the acquisition and maintenance of expert performance in medicine and related domains. academic medicine, 79(10 suppl), s70-81. doi:00001888-20041000100022 ericsson, k. a. (2006a). an introduction to the cambridge handbook of expertise and expert performance: its development, organization, and content. in k. a. ericsson, n. charness, p. j. feltovich, & r. r. hoffman (eds.), the cambridge handbook of expertise and expert performance (pp. 3-19). new york, ny: cambridge university press. ericsson, k. a. (2006b). protocol analysis and expert thought: concurrent verbalisations of thinking during experts’ performance on representative tasks. in k. a. ericsson, n. charness, p. j. feltovich, & r. r. hoffman (eds.), the cambridge handbook of expertise and expert performance (pp. 223-241). new york, ny: cambridge university press. ericsson, k. a. (2009). development of professional expertise: toward measurement of expert performance and design of optimal learning environments. new york, ny: cambridge university press. ericsson, k. a. (2015). acquisition and maintenance of medical expertise: a perspective from the expertperformance approach with deliberate practice. academic medicine, 90(11), 1471-1486. doi:10.1097/acm.0000000000000939 ericsson, k. a. (2014). how to gain the benefits of the expert performance approach in domains where the correctness of decisions are not readily available: a reply to weiss and shanteau. applied cognitive psychology, 28(4), 458-463. doi:10.1002/acp.3029 ericsson, k. a., & kintsch, w. (1995). long-term working memory. psychological review, 102(2), 211245. doi:10.1037/0033-295x.102.2.211 ericsson, k. a., krampe, r. t., & tesch-römer, c. (1993). the role of deliberate practice in the acquisition of expert performance. psychological review, 100(3), 363-406. doi:10.1037/0033-295x.100.3.363 ericsson, k. a., patel, v., & kintsch, w. (2000). how experts' adaptations to representative task demands account for the expertise effect in memory recall: comment on vicente and wang (1998). psychological review, 107(3), 578-592. doi:10.1037/0033-295x.107.3.578 ericsson, k. a., & r. poole (2016). peak: secrets from the new science of expertise. london, uk: the bodley head. ericsson, k. a., & simon, h. a. (1980). verbal reports as data. psychological review, 87(3), 215-251. doi:10.1037/0033-295x.87.3.215 ericsson, k. a., & simon, h. a. (1993). protocol analysis: verbal reports as data. cambridge, ma: mit press. ericsson, k. a., & smith, j. (1991). prospects and limits of the empirical study of expertise: an introduction. in k. a. ericsson, & j. smith (eds.), toward a general theory of expertise: prospects and limits (pp. 138). cambridge: cambridge university press. evetts, j., mieg, h. a., & felt, u. (2006). professionalization, scientific expertise, and elitsm: a sociological perspective. in k. a. ericsson, n. charness, p. j. feltovich, & r. r. hoffman (eds.), the cambridge handbook of expertise and expert performance (pp. 105-123). new york, ny: cambridge university press. feltovich, p. j., & barrows, h. s. (1984). issues of generality in medical problem solving. in h. g. schmidt, & m. l. de volder (eds.), tutorials in problem-based learning. new directions in training for the health professions, (pp. 128-142). assen/maastricht: van gorcum. feltovich, p. j., prietula, m. j., & ericsson, a. (2006). studies of expertise from psychological perspectives. in k. a. ericsson, n. charness, p. j. feltovich, & r. r. hoffman (eds.), the cambridge handbook of expertise and expert performance (pp. 41-67). new york, ny: cambridge university press. http://psycnet.apa.org/doi/10.1037/0033-295x.102.2.211 http://psycnet.apa.org/doi/10.1037/0033-295x.107.3.578 http://dx.doi.org/10.1037/0033-295x.87.3.215 http://dx.doi.org/10.1037/0033-295x.87.3.215 van de wiel | f l r 119 flanagan, j. c. (1954). the critical incident technique. psychological bulletin, 51(4), 327. doi:10.1037/h0061470 gegenfurtner, a., siewiorek, a., lehtinen, e., & säljö, r. (2013). assessing the quality of expertise differences in the comprehension of medical visualizations. vocations and learning, 6(1), 37-54. doi:10.1007/s12186-012-9088-7. gilhooly, k. j., mcgeorge, p., hunter, j., rawles, j. m., kirby, i. k., green, c., & wynn, v. (1997). biomedical knowledge in diagnostic thinking: the case of electrocardiogram (ecg) interpretation, european journal of cognitive psychology, 9(2), 199-223. doi:10.1080/713752555. hamm, r. m. (1988). clinical intuition and clinical analysis: expertise and the cognitive continuum. in j. dowie, & a. elstein (eds.), professional judgment: a reader in clinical decision making, (pp.78-105). cambridge, ma: cambridge university press. hammond, k. r., hamm, r. m., grassia, j., & pearson, t. (1987). direct comparison of the efficacy of intuitive and analytical cognition in expert judgment. transactions on systems, man and cybernetics, ieee, 17(5), 753-770. doi:10.1109/tsmc.1987.6499282 hashem, a., chi, m. t., & friedman, c. p. (2003). medical errors as a result of specialization. journal of biomedical informatics, 36(1), 61-69. doi:10.1016/s1532-0464(03)00057-1 hassebrock, f., & prietula, m. j. (1992). a protocol-based coding scheme for the analysis of medical reasoning. international journal of man-machine studies, 37(5), 613-652. doi:10.1016/00207373(92)90026-h. hatala, r., norman, g. r., & brooks, l. r. (1999). impact of a clinical scenario on accuracy of electrocardiogram interpretation. journal of general internal medicine, 14(2), 126-129. doi:10.1111/j.1525-1497.1999.tb00008.x helle, l. (2017). prospects and pitfalls in combining eye-tracking data and verbal reports. frontline learning research, 5(3), 1-12. doi:10.14786/flr.v5i3.254 hoffman, r. r., & lintern, g. (2006). eliciting and representing the knowledge of experts. in k. a. ericsson, n. charness, p. j. feltovich, & r. r. hoffman (eds.), the cambridge handbook of expertise and expert performance (pp. 203-222). new york, ny: cambridge university press. hoffman, r. r., shadbolt, n. r., burton, a. m., & klein, g. (1995). eliciting knowledge from experts: a methodological analysis. organizational behavior and human decision processes, 62(2), 129-158. hsieh, h. f., & shannon, s. e. (2005). three approaches to qualitative content analysis. qualitative health research, 15(9), 1277-1288. doi:10.1177/1049732305276687 jaarsma, t., jarodzka, h., nap, m., merrienboer, j. j. g., & boshuizen, h. p. a. (2014). expertise under the microscope: processing histopathological slides. medical education, 48(3), 292-300. doi:10.1111/medu.12385. kahneman, d., & klein, g. (2009). conditions for intuitive expertise: a failure to disagree. american psychologist, 64(6), 515-526. doi:10.1037/a0016755 king, n., & horrocks, c. (2010). interviews in qualitative research. london, uk: sage. klein, g. (2008). naturalistic decision making. human factors, 50(3), 456-460. doi:10.1518/001872008x288385. kok, e. m., jarodzka, h., de bruin, a. b. h., binamir, h. a. n., robben, s. g. f., & van merriënboer, j. j. g. (2015). systematic viewing in radiology: seeing more, missing less?. advances in health sciences education, 1-17. doi:10.1007/s10459-015-9624-y. kok, e. m., de bruin, a. b. h., robben, s. g. f., & van merriënboer, j. j. g. (2013). learning radiological appearances of diseases: does comparison help? learning and instruction, 23, 90-97. doi:10.1016/j.learninstruc.2012.07.004. krippendorff, k. (2012). content analysis: an introduction to its methodology (3rd ed.). newsbury park, ca: sage. krueger, r. a., & casey, m. a. (2015). focus groups: a practical guide for applied research (5th edition). thousand oaks, ca: sage. http://dx.doi.org/10.1109/tsmc.1987.6499282 http://dx.doi.org/10.1016/s1532-0464(03)00057-1 http://dx.doi.org/10.1016/0020-7373%2892%2990026-h http://dx.doi.org/10.1016/0020-7373%2892%2990026-h http://dx.doi.org/10.1016/j.learninstruc.2012.07.004 van de wiel | f l r 120 kulatunga-moruzi, c., brooks, l. r., & norman, g. r. (2004). using comprehensive feature lists to bias medical diagnosis. journal of experimental psychology: learning, memory, and cognition, 30(3), 563572. doi:10.1037/0278-7393.30.3.563 lesgold, a., rubinson, h., feltovich, p., glaser, r., klopfer, d., wang, y., et al. (1988). expertise in a complex skill: diagnosing x-ray pictures. in m. t. h. chi, r. glaser, & m. j. farr (eds.), the nature of expertise (pp. 311-342). hillsdale, nj: lawrence erlbaum. lesgold, a. (2000). on the future of cognitive task analysis. in j. m. schraagen, s. f. chipman, & v. l. shalin (eds.), cognitive task analysis (pp. 451-465). mahwah, nj: lawrence erlbaum. mieg, h. a. (2006). social and sociological factors in the development of expertise. in k. a. ericsson, n. charness, p. j. feltovich, & r. r. hoffman (eds.), the cambridge handbook of expertise and expert performance (pp. 743-760). new york, ny: cambridge university press. morgan, d. l. (1996). focus groups as qualitative research (2nd ed.). thousand oaks, ca: sage. neuendorf, k. a. (2002). the content analysis guidebook. thousand oaks, ca: sage. neufeld, v. r., norman, g. r., feightner, j. w., & barrows, h. s. (1981). clinical problem-solving by medical students: a cross-sectional and longitudinal analysis. medical education, 15(5), 315-322. doi:10.1111/j.1365-2923.1981.tb02495.x norman, g. r., brooks, l. r., & allen, s. w. (1989). recall by expert medical practitioners and novices as a record of processing attention. journal of experimental psychology: learning, memory, and cognition, 15(6), 1166-1174. doi:10.1037/0278-7393.15.6.1166 norman, g. r., eva, k., brooks, l., & hamstra, s. (2006). expertise in medicine and surgery. in k. a. ericsson, n. charness, p. j. feltovich, & r. r. hoffman (eds.), the cambridge handbook of expertise and expert performance (pp. 339-354). new york, ny: cambridge university press. patel, v. l., & arocha, j. f. (2001). the nature of constraints on collaborative decision making in health care settings. in e. salas, & klein, g. (eds.), linking expertise and naturalistic decision making (pp. 383-405). mahwah, nj: lawrence erlbaum. patel, v. l., & groen, g. j. (1986). knowledge based solution strategies in medical reasoning. cognitive science, 10(1), 91-116. doi:10.1207/s15516709cog1001_4 patel, v. l., kaufman, d. r., & magder, s. a. (1996). the acquisition of medical expertise in complex dynamic environments. in k .a. ericcson (ed.), the road to excellence: the acquisition of expert performance in the arts and sciences, sports and games, (pp. 127-165). mahwah, nj: lawrence erlbaum. prince, k. j. a. h., van de wiel, m. w. j., van der vleuten, c. p. m., boshuizen, h. p. a., & scherpbier, a. j. j. a. (2004). junior doctors' opinions about the transition from medical school to clinical practice: a change of environment. education for health, 17(3), 323-331. doi:10.1080/13576280400002510 salas, e., & klein, g. a. (eds .) (2001). linking expertise and naturalistic decision making. mahwah, nj: lawrence erlbaum. salas, e., rosen, m. a., burke, c. s., goodwin, g. f., & fiore, s. m. (2006). the making of a dream team: when expert teams do best. in k. a. ericsson, n. charness, p. j. feltovich, & r. r. hoffman (eds.), the cambridge handbook of expertise and expert performance (pp. 439-453). cambridge, uk: cambridge university press. sanchez, j. i., & levine, e. l. (2012). the rise and fall of job analysis and the future of work analysis. annual review of psychology, 63, 397-425. doi: 10.1146/annurev-psych-120710-100401. schmidt, h. g., & boshuizen, h. p. a. (1993). on the origin of intermediate effects in clinical case recall. memory and cognition, 21, 338 351. doi:10.3758/bf03208266 shanteau, j. (1992). competence in experts: the role of task characteristics. organizational behavior and human decision processes, 53(2), 252-266. shanteau, j., weiss, d. j., thomas, r. p., & pounds, j. c. (2002). performance-based assessment of expertise: how to decide if someone is an expert or not. european journal of operational research, 136(2), 253-263. skopec, e. w. (1986). situational interviewing. prospects heights, ii: waveland press. van de wiel | f l r 121 stalmeijer, r. e., mcnaughton, n., & van mook, w. n. k. a. (2014). using focus groups in medical education research: amee guide no. 91. medical teacher, 36(11), 923-939. doi:10.3109/0142159x.2014.917165 stolper, e., van bokhoven, m., houben, p., van royen, p., van de wiel, m. w. j., van der weijden, t., & dinant, g. j. (2009). the diagnostic role of gut feelings in general practice. a focus group study of the concept and its determinants. bmc family practice, 10(1), 17. doi:10.1186/1471-2296-10-17. stolper, e., van de wiel, m. w. j., hendriks, r. h. m., van royen, p., van bokhoven, m., van der weijden, t., & dinant, g. j. (2015). how do gut feelings feature in tutorial dialogues on diagnostic reasoning in gp traineeship? advances in health sciences education, 20, 499-513. doi:10.1007/s10459-014-9543-3. stolper, e., van de wiel, m. w. j., van bokhoven, m., van royen, p., van der weijden, t., & dinant, g. j. (2011). gut feelings as a third track in general practitioners’ diagnostic reasoning. journal of general internal medicine, 26, 197-203. doi:10.1007/s11606-010-1524-5. shadbolt, n. r., & smart, p. r. (2015) knowledge elicitation. in j. r. wilson, & s. sharples (eds.), evaluation of human work (4th ed.). boca raton, fl: crc press. tracey, t. j., wampold, b. e., lichtenberg, j. w., & goodyear, r. k. (2014). expertise in psychotherapy: an elusive goal? american psychologist, 69(3), 218229. doi:10.1037/a0035099 van de wiel, m. w. j., boshuizen, h. p. a., & schmidt, h. g. (2000). knowledge restructuring in expertise development: evidence from pathophysiological representations of clinical cases by students and physicians. european journal of cognitive psychology, 12(3), 323-355. doi:10.1080/09541440050114543 van de wiel, m. w. j., boshuizen, h. p. a., schmidt, h. g. & schaper, n. c. (1999). the explanation of medical concepts by expert physicians, clerks and advanced students. teaching and learning in medicine, 11(3), 153-163. doi:10.1207/s15328015tl110306 van de wiel, m. w. j., ploegh, k., boshuizen, h. p. a., & schmidt, h. g. (2005). the influence of diagnosis and memorization instructions on clinical case processing by students and physicians. paper presented at the annual meeting of the american educational research association 2005. montreal, canada, april 11-15. van de wiel, m. w. j., schaper, n. c., scherpbier, a. j. j. a., van der vleuten, c. p. m., & boshuizen, h. p. a. (1999). students' experiences with real patient tutorials in a problem-based curriculum. teaching and learning in medicine, 11(1), 12-20. doi:10.1207/s15328015tlm1101_5 van de wiel, m. w. j., & schmidt, h. g., boshuizen, h. p. a. (1998). a failure to reproduce the intermediate effect in clinical case recall. academic medicine, 73(8), 894-900. van de wiel, m. w. j., & van den bossche, p. (2013). deliberate practice in medicine: the motivation to engage in work-related learning and its contribution to expertise. vocations and learning, 6(1), 135158. doi:10.1007/s12186-012-9085-x. van de wiel, m. w. j., van den bossche, p., janssen, s., & jossberger, h. (2011). exploring deliberate practice in medicine: how do physicians learn in the workplace? advances in health sciences education. 16(1), 81-95. doi:10.1007/s10459-010-9246-3. van de wiel, m. w. j., van den bossche, p., & koopmans, r. p. (2011). deliberate practice, the high road to expertise: k.a. ericsson. in dochy, f., gijbels, d., segers, m., & van den bossche, p. (eds.), theories of learning for the workplace: building blocks for training and professional development programs (pp. 1-16). london, uk: routledge. van gog, t., paas, f., van merriënboer, j. j., & witte, p. (2005). uncovering the problem-solving process: cued retrospective reporting versus concurrent and retrospective reporting. journal of experimental psychology: applied, 11(4), 237-244. doi:10.1037/1076-898x.11.4.237 van someren, m. v., barnard, y. f., & sandberg, j. a. (1994). the think aloud method: a practical approach to modelling cognitive processes. london, uk: academic press. weiss, d. j., & shanteau, j. (2003). empirical assessment of expertise. human factors, 45(1), 104-116. http://psycnet.apa.org/doi/10.1037/a0035099 van de wiel | f l r 122 wimmers, p. f., schmidt, h. g., verkoeijen, p. p. j. l., & van de wiel, m. w. j. (2005). inducing expertise effects in clinical case recall through the manipulation of processing. medical education, 39, 949-957. doi:10.1111/j.1365-2929.2005.02250.x frontline learning research vol.5 no. 2 (2017) 60 77 issn 2295-3159 corresponding author: jannis bosch, university of potsdam, inclusive education – research methods and diagnostics, karl-liebknecht-str. 24-25, 14476 potsdam, germany. e-mail: jbosch@uni-potsdam.de, phone: +49-331977-2462. doi: http://dx.doi.org/10.14786/flr.v5i2.292 contrast and assimilation effects on task interest in an academic learning task jannis bosch & jürgen wilbert university of potsdam, germany article received 14 february / revised 18 may / accepted 2 august / available online 7 august abstract information on social comparison is one of the major factors used to evaluate academic achievement. the presence of big-fish-little-pond (bflp) and basking-in-reflectedglory (birg) effects of academic achievement on the self-concept have been extensively researched in various observational studies. recent research suggests that these effects can also be transferred to motivational variables such as task interest. this paper uses an experimental paradigm to take a closer look at the mechanisms expected to be behind bflp and birg effects: contrast and assimilation effects of task performance. the analyses are based on n = 129 primary education students who completed a computer-based learning task. during this task, participants received social comparative feedback that was experimentally manipulated based on 2x2 conditions: social position (high vs. low) and peer performance (high vs. low). task interest was measured both before the start of the learning task and after it was finished. results indicate a positive influence of high social position and high peer performance on the development of task interest from preto posttest compared to the low social position and peer performance conditions. further, the hypothesized positive association between self-concept, task interest and performance in the academic learning task could also be shown. keywords: social comparison; feedback; interest; contrast effect; assimilation effect mailto:jbosch@uni-potsdam.de http://dx.doi.org/10.14786/flr.v5i2.292 bosch et wilbert | f l r 61 1. theory a lot of time during childhood, adolescence, and early adulthood is spent in institutionalized learning situations, such as school, vocational training, or university. usually, these institutions are group-based and therefore social in nature. while these settings enable the use of group-based learning activities that can have several positive effects on learning achievement and attitudes towards learning (springer, stanne, & donovan, 1999), they also facilitate the use of social comparison information in the evaluation of students' learning progress or achievement. hence, information on social comparison is one of the major factors used to evaluate academic achievement in many educational contexts, both formally (e.g., grades) and informally (e.g., contact with peers). while social comparison can have desirable effects, it can also have considerable downsides, such as negative affect and lack of self-confidence, especially for lower-achieving students: “in classrooms characterized by frequent grades and public evaluation, students become focused on their ability and the distribution of ability in the classroom group. many students not only come to believe that they lack ability but this perception also becomes shared among peers” (ames, 1992, p. 264). this becomes even more relevant considering the increasing heterogeneity of academic learning groups that can be observed lately, due to the implementation of the “convention of the rights of persons with disabilities” in 2008 and the corresponding move to an inclusive school system in many countries. processes of social comparison not only present a risk for low-achieving students to develop a negative self-concept, but can also influence academic achievement itself. studies have shown that motivational (i.e. interest; möller, retelsdorf, köller, & marsh, 2011; schiefele, krapp, & winteler, 1992) and self-related emotional (i.e. self-concept; guay, marsh, & boivin, 2003; huang, 2011; marsh, trautwein, lüdtke, köller, & baumert, 2005; möller et al., 2011; valentine, dubois, & cooper, 2004) variables play a very important role in the academic learning process. further, the relation between academic achievement and both interest and self-concept seems to be bidirectional, as research reveals that low academic achievement in turn leads to a negative academic self-concept and hampers the development of interest in an academic topic (denissen, zarrett, & eccles, 2007; marsh et al., 2005), which can potentially lead to a vicious circle of negative self-affect, diminished interest and negative performance feedback in academic subjects. hence, the establishment of motivating and emotionally stabilizing evaluation practices should also be a major goal to ensure successful institutionalized learning for every student. to be able to do that, we need to analyse how the use of social comparison data in academic evaluation variously influences the interest and self-concept of students with heterogeneous prerequisites. although there is ample evidence of the impact socially comparative feedback has on the development of an individual's selfconcept, the role of interest and motivation has been considerably less well scrutinized. that said, theoretical models, as well as empirical evidence, propose a tight connection between self-concept and interest in an academic context (e.g. fryer, 2015). both constructs are domain-specific, and the associations between them always correspond to a specific domain (i.e. math self-concept is associated with math interest). obviously, self-concept and interest are two distinctive constructs nonetheless: interest, on one hand, is the motivational orientation of a person toward a certain object, activity, or area of knowledge that is independent of any direct external rewards and can therefore be described as domain-specific intrinsic motivation (schiefele, 1992). self-concept, on the other hand, consists of a person’s perceptions of himor herself. these perceptions are formed based on self-evaluation as well as the evaluation of significant others. that means that self-concept is self-ascribed competence in a specific domain (marsh & shavelson, 1985; shavelson, hubner, & stanton, 1976). the self-determination theory (deci & ryan, 1985), one of the most popular motivational theories, has identified feelings of competence in a specific domain as crucial for the establishment of intrinsic motivation in that same domain, suggesting a possible explanation for empirical connections between selfconcept and interest found in several studies (e.g. fryer, 2015; schurtz, pfost, nagengast, & artelt, 2014). taking the feeling of competence into focus as a bridge between self-concept and interest, empirical evidence for the effects of social comparison on self-concept will be presented in the following section. according to the internal/external frame of reference model by marsh (i/e-model; 1986) development of the self-concept is based on comparison of perceived competence in a specific domain with both internal (i.e. perceived competence in other domains) and external (i.e. perceived competence of other persons in that same bosch et wilbert | f l r 62 domain) frames of reference. while the former can be expected to have an overall neutral effect on the sum of self-concepts in different domains (e.g. a person perceives himself as good in math but bad in arts), the latter can have predominantly positive or negative effects on the self-concept in various domains. hence, social comparison processes are a major factor for the development of a positive self-concept, especially when the evaluative atmosphere of the typical classroom makes it hard for students not to compare themselves with each other (dijkstra, kuyper, van der werf, buunk, & van der zee, 2008). as stated earlier, social comparison processes can be detrimental to the development of a positive self-concept in low-achieving students. the results of a study by dickhäuser and galfe (2004) suggest that the direction of comparison (i.e. comparison with a better or worse performing peer) has an influence on the development of the self-concept. in their study, students were asked to whom they compared themselves after their math test scores were announced. additionally, math self-concept was assessed before and after the math test. the results showed a negative effect of upward social comparison (i.e. comparison with a better performing peer) on math self-concept, while downward social comparison (i.e. comparison with a worse performing peer) tended to produce more positive effects. hence, the predominance of upward comparison targets might be one factor responsible for the previously stated negative effects of social comparison on self-concept development in low-achieving students. this notion is further supported by two studies that investigated the effects of ability grouping on self-concept: the first study showed that ability grouping leads to decreased academic self-concept for those in higher ability groups compared to those in lower ability groups (mulkey, catsambis, steelman, & crain, 2005), while the second study found a similar effect of lowered academic self-concept for academically handicapped special school students who were integrated into mainstream schooling (strang, smith, & rogers, 1978). therefore, the influence of social comparison on self-evaluation, and its effects on self-concept, cannot be properly understood without considering the composition of the reference group. the intimate nature of the process connecting self-evaluation and the attributes of social comparison groups have been outlined in the inclusion/exclusion model (iem; bless & schwarz, 2010; schwarz & bless, 1992). according to the iem, evaluation of a particular stimulus is based on the stimulus itself and a standard of comparison. because humans can never consider all potentially relevant information, both the stimulus itself and the standard of comparison are defined by the subset of information that is most accessible at the time of judgment. the model assumes two possible effects of the standard of comparison on the evaluation of the target: a contrast effect, i.e. the use of the standard of comparison as a reference against which to judge any characteristic of the target, and an assimilation effect, i.e. the transfer of characteristics of the standard of comparison to the target. the iem covers both situations in which the target is superordinate to the standard of comparison, i.e. a group is judged in relation to an individual that is part of the group, and situations in which the target is subordinate to the standard of comparison, i.e. a member of the group is judged in relation to the group. if we apply the iem to the process of self-evaluation, the target (e.g., a student) is part of a superordinate category (e.g., the class or school). while there are various sources of information that can be used for self-evaluation, information about other students in one’s class (or school) is very salient and readily available in school contexts and will therefore most likely play a major role in school-based self-evaluations. depending on which information is actually used to evaluate the self, both contrast effects, i.e. the usage of one’s class (or school) as a reference against which to judge any of one's own characteristics, or assimilation effects, i.e. the transfer of characteristics of the students of one’s class (or school) to oneself, can occur (see bless & schwarz, 2010 for a detailed discussion of influencing factors). 1.1 social comparison effects in heterogeneous classrooms the concept of assimilation and contrast effects supposedly lies at the core of one of the most popular phenomena in educational research. based on their observation that students attending schools with a low average socio-economic status (ses) had higher academic self-concepts than equally performing students attending schools with a higher average ses, marsh and parker (1984) came up with the big-fish-little-pond (bflp) and basking-in-reflected-glory (birg) effects. the bflp effect is the transfer of a contrast effect to an institutionalized learning context; it states that self-evaluation of performance is based on the average bosch et wilbert | f l r 63 performance of other students in the same school, thus resulting in a higher academic self-concept for a student in a lower-performing school compared to an equally performing student in a higher-performing school. the birg effect is the transfer of an assimilation effect to an institutionalized learning context and has contrary expectations: a student in a high-performing school will have a higher academic self-concept than an equally performing student in a lower-performing school, because both students identify with their reference group and adapt their academic self-concept based on their group affiliation. however, because of a relatively stronger bflp effect, the net result of the reference group's strength in terms of academic self-concept is still expected to be negative (marsh, 1987). since then, the bflp effect on academic self-concept has been confirmed in school children from different countries (marsh et al., 2005; schurtz et al., 2014; wouters, de fraine, colpin, van damme, & verschueren, 2012), as well as in larger studies covering schools from various culturally and economically diverse countries (chiu, 2012; marsh & hau, 2003; seaton, marsh, & craven, 2009; wang, 2015). similar bflp effects could also be found in young adults (jonkmann, becker, marsh, lüdtke, & trautwein, 2012). further, several studies have also shown empirical evidence for the existence of an albeit usually smaller birg effect in children and young adults (huguet et al., 2009; marsh, kong, & hau, 2000; preckel & brüll, 2010; trautwein, lüdtke, marsh, & nagy, 2009). while the bflp effect has mainly been investigated regarding academic self-concept, there are also several studies showing that the underlying processes may additionally affect other self-related constructs, as well as emotional variables. more specifically, these include bflp effects on self-related constructs such as self-esteem (jansen, scherer, & schroeders, 2015) and self-efficacy (marsh, trautwein, lüdtke, & köller, 2008), on general motivational variables such as educational and occupational aspirations (marsh, 1991; marsh & o’mara, 2010; nagengast & marsh, 2011), on control expectations and strategies (marsh et al., 2008), and on emotional variables such as test anxiety (goetz, preckel, zeidner, & schleyer, 2008; zeidner & schleyer, 1999) and emotionality (goetz et al., 2008). additionally, a similar bflp effect could also be shown in regard to task interest (köller, schnabel, & baumert, 2000; schurtz et al., 2014; trautwein, lüdtke, marsh, köller, & baumert, 2006), possibly reflecting the close connection between self-concept and interest. trautwein and colleagues (2006) additionally tested for a birg effect on task interest. although they did not find evidence for a birg effect on task interest, the failure to reproduce the original birg effect on self-concept renders those results somewhat fragile. even though the presented results provide indirect evidence for the existence of contrast and assimilation effects in institutionalized education contexts, the non-experimental nature of the described studies still leaves some questions in relation to the assumed modes of action unanswered. in order to be able to investigate the direct influence of social comparative performance feedback (i.e. contrast and assimilation effects) on task interest, experimental designs explicitly manipulating the performance feedback are necessary. to the author’s knowledge, only one study used such an experimental design to show that contrast effects of performance feedback have a direct influence on task interest (pohlmann & möller, 2006). a previous experimental study by wilbert, grosche, and gerdes (2010) could also show that social comparison feedback tends to increase task-specific motivation in high-achieving students while undermining low-achieving students’ motivation. however, they did not control for the fact that highachieving students usually receive positive feedback, while low-achieving students receive negative feedback. 1.2 research question and hypotheses in this paper, we want to further investigate the assumed modes of action behind the bflp and birg effects: the direct influence contrast and assimilation effects in social comparative performance feedback have on students' interest in an academic learning task. the only study directly investigating the influence of social comparative performance feedback on task interest could find evidence for a contrast effect (pohlmann & möller, 2006). they did not, however, test for potential assimilation effects. therefore, we want to investigate the directs effects of social comparative feedback on task interest to shed more light on the modes of action expected to be behind the popular bflp and birg effects. bosch et wilbert | f l r 64 hence, the research question is as follows: how do contrast and assimilation effects of social comparative performance feedback directly influence task interest in an academic learning task? based on the empirical results and theoretical considerations presented above, we expect both contrast and assimilation effects to influence interest in an academic learning task. that is, we expect a higher social position within a reference group (compared to a lower social position), as well as a higher peer performance (compared to lower peer performance) to be associated with higher interest in an academic learning task. in order to achieve a clear differentiation between the quality of feedback and actual performance, and therefore to increase the internal validity, feedback related to both social position and peer performance was experimentally manipulated in this study. that means the performance feedback every participant received was independent of actual performance. to directly investigate potential contrast effects the percentile rank of performance feedback given was experimentally manipulated (high and low social position conditions). this was done to simulate the respective performance feedback high and low performing learners usually receive. further, potential assimilation effects were investigated by experimentally manipulating the percentage of points the reference group scored (high and low peer performance conditions). additionally, we want to replicate previously presented positive associations between self-concept, interest and achievement in an academic learning task. finally, the potential influences of contrast and assimilation effects on task enjoyment, a further variable strongly associated with task-interest, are explored. 2. method 2.1 participants participants were 129 first year pedagogy students at university of potsdam studying to become elementary school teachers. seven participants were excluded from further analyses due to missing data or because they expressed doubt about the authenticity of the performance feedback, leaving us with 122 participants. each participant was randomly assigned to one of 2x2 experimental conditions: 31 participants were in the high social position / high peer performance (sp+/pp+) condition, 30 participants in the low social position / high peer performance (sp-/pp+) condition, 32 participants were in the high social position / low peer performance (sp+/pp-) condition, and 29 participants were in the low social position / low peer performance (sp-/pp-) condition. 110 participants were female, 11 were male, and 1 did not report a gender. the age of the participants ranged from 18 to 37 (m = 22.19, sd = 4.29). all participants attended a lecture on inclusive education and were recruited during the class. 2.2 learning task the learning task we used in this study was presented to the participants as the “flag game.” participants were told that the game consisted of two phases: a learning and a performance phase. during the learning phase, participants were presented with a map of the entire african continent and the country outlines for every country. for each of the 30 learning trials, the outlines of one country were highlighted, accompanied by the corresponding national flag. each combination was presented for up to 10 seconds. participants could shorten the presentation of each pair by pressing the space key. after completion of the learning phase, the performance phase was started. during the performance phase, the same 30 items were presented. this time, however, five different country outlines were highlighted with every flag, that is, four distractors in addition to the correct country. participants were instructed to choose the correct country as quickly as possible by clicking on the corresponding country name. there was no time limit in the performance phase. after the first run, feedback was given depending on the respective feedback condition. then, a second run comprising of bosch et wilbert | f l r 65 another learning and performance phase with the same 30 items was started. after the second run, no feedback was given. to avoid sequencing effects between conditions, trial sequences were pseudo-randomized separately for each phase. distractor items were taken from the item pool, balanced and pseudo-randomized separately for each performance phase. hence, every participant was presented with the exact same trial sequences and distractor items. figure 1. feedback slides for all four experimental conditions (feedback slides were originally in german and were translated for this article). the slides on the left were presented to the respective low social position conditions (top side: low social position/low peer performance, bottom side: low social position/high peer performance), while the slides on the right were presented to the respective high social position conditions (top side: high social position/low peer performance, bottom side: high social position/high peer performance). 2.3 experimental feedback conditions during instruction, participants were told that they would receive information about their performance compared to the last 50 students of the same degree course. however, feedback did not reflect their real performance, but was given based on the respective feedback condition. to test for contrast and assimilation effects on task interest, we manipulated both social position (sp; high vs. low) and peer performance (pp; high vs. low) in a 2x2 design. participants were told their performance rating could range from 0 to 100%, based on both correctness and the speed of their responses. response speed was included to make it harder for the participants to predict their own performance. performance feedback was shown graphically, with 50 stick figures in a graph representing the scores of the reference group and a highlighted stick figure representing the participants’ own score (see figure 1). scores for the reference group were approximately normally distributed and ranged between 0 and 50% in the low peer performance conditions. in the high peer performance conditions, the reference group score was simply raised by 30% for every reference group member, resulting in values between 30% and 80%. based on the respective social position condition, the participant’s individual bosch et wilbert | f l r 66 score was placed either at the top or the bottom of the distribution. this resulted in individual scores of 79% (high social position / high peer performance), 49% (high social position / low peer performance and low social position / high peer performance), or 19% (low social position / low peer performance). 2.4 measures 2.4.1 self-concept the self-concept scale consists of four items: two positive (e.g., “usually i have no trouble with learning tasks”) and two negative (e.g., “dealing with learning tasks isn’t one of my strengths”) statements about the participant’s self-rated ability to resolve learning tasks. participants rated the items on a likert-type scale ranging from 1 (strongly disagree) to 7 (strongly agree). negative items were reversed and a mean selfconcept score was calculated. cronbach’s alpha of the scale was α = .84. 2.4.2 task interest (preand post-test) the interest scale consists of eight items: four items concerning interest in the task itself (e.g., “i like learning tasks such as this one”) and four items concerning interest in the content of the task (e.g., “i’m interested in country flags”). the items were adapted versions of an interest scale from a german motivation questionnaire (rheinberg, vollmeyer, & burns, 2001). each item was rated on a likert-type scale ranging from 1 (strongly disagree) to 7 (strongly agree). pre-test interest had a cronbach’s alpha of .87, while post-test interest had a cronbach’s alpha of α = .90. 2.4.3 task enjoyment task enjoyment was measured by means of a single item at the end of the study (“did you enjoy the study?”). the item was rated on a likert-type scale from 1 (definitely no) to 4 (definitely yes). 2.4.4 task performance (preand post-test) pre-test performance was the proportion of correctly answered multiple-choice items in the first performance phase, while post-test performance was the percentage of correctly answered items in the second performance phase. both performance phases consisted of the same 30 items. 2.5 procedure participants were tested in groups of four to six persons in our learning laboratory. they were welcomed and informed about the upcoming procedure. written informed consent was obtained from every participant prior to testing. each participant was placed in front of a separate computer. computers were properly shielded to avoid participants seeing each other’s screens. firstly, the self-concept questionnaire was administered. during the computer-based instruction, participants were informed about the upcoming learning task and how feedback about their performance would be provided. after completion of the instruction, a questionnaire concerning task interest was administered. when all participants finished filling out the questionnaire, the computer-based learning task was begun. afterwards, the task interest questionnaire was administered a second time to investigate any potential changes. in the end, participants answered two questions concerning whether they were motivated to participate in similar studies and an open question that functioned as a manipulation check. the true nature of the study and the experimental manipulation was then explained to all participants. bosch et wilbert | f l r 67 2.6 statistical analyses to test the previously presented hypotheses that there are contrast and assimilation effects on domainspecific interest, we followed the suggestion by everitt and hothorn (2011, p. 232) and used linear mixed models with fixed and random effects. this was done because regular linear regression is unable to deal with data from repeated measures. maximum likelihood (ml) estimation was used for all models. based on the suggestions by nakagawa and schielzeth (2013), marginal and conditional r² were reported in order to determine variance explained by fixed factors and by both fixed and random factors, respectively. model 1 contained the subject as a random intercept factor and preand post-test interest as level 1 variables. contrasts for the measurement time factor were coded 0 for the preand 1 for the post-test measurements. for model 2, we added the level 2 variable social position and the interaction with measurement time. then, peer performance, as well as interactions with measurement time and social position were added for model 3. finally, we also added the standardized (m = 0, sd = 1) self-concept and interaction terms to the other predictors in order to investigate possible moderating effects in model 4. the contrasts were effect-coded as 1 for low social position or peer performance and +1 for high social position or peer performance, respectively. to determine the model best fitting our data, likelihood ratio (lr) tests comparing each model to the previous one were calculated. this was done to test whether the subsequent addition of predictors significantly increased variance explained. to investigate the possible effects of experimental conditions on learning task performance, corresponding models were also calculated with task performance as the dependant variable. because task enjoyment was only retrieved at a single point in time at the end of the study, generalized least squares (gls) models with ml estimation and only one measurement per subject were calculated to investigate bflp effects on task enjoyment. in order to determine explained variance in the gls models, likelihood ratio r² was calculated as suggested by magee (1990). this time, an intercept-only model with task enjoyment as the dependent variable was used as the base model 1. again, social position was included in model 2, while peer performance and the interaction between social position and peer performance were added to model 3 and standardized self-concept and corresponding interactions to model 4. once again, lr tests were calculated to compare the four models. all calculations were carried out using the statistical software r with nlme and mumin packages. 3. results table 1 shows intercorrelations and descriptive statistics for all observed variables. as can be seen, self-concept is positively associated with preand post-test interest. associations between self-concept and preand post-test performance was also positive, but only significant in the latter case. further, pre-test interest showed a significant positive association to preand post-test performance. comparable associations could be found between post-test interest and preand post-test performance. hence, results support the hypothesized associations between interest, self-concept and performance. bosch et wilbert | f l r 68 table 1 descriptive statistics and intercorrelations for all observed variables selfconcept interest pre-test interest post-test interest diff. perf. pre-test perf. post-test perf. diff. enjoy self-concept 1 interest pre-test .52** 1 interest post-test .39** .87** 1 interest diff. -.16 -.07 .44** 1 perf. pre-test .13 .23* .31** .20* 1 perf. post-test .23* .20* .25* .14 .74** 1 perf. diff. .15 -.02 -.06 -.08 -.30** .42** 1 enjoy .09 .24* .29** .15 .01 .04 .04 1 m 4.39 3.81 3.88 0.07 .51 .71 .20 3.40 sd 1.03 1.16 1.28 0.64 1.15 1.16 1.11 0.65 n 122 122 122 122 122 122 122 121 note. ** p < .01, * p < .05; interest diff. = difference between interest preand post-test, perf. pre-test = performance pre-test, perf. post-test = performance post-test, perf. diff. = difference between performance preand post-test, enjoy = enjoyment table 2 descriptive statistics for all observed variables for each experimental condition sp + pp + sp – pp + sp + pp – sp – pp – self-concept pre-test 4.47 (1.04) 4.21 (0.95) 4.31 (1.16) 4.59 (0.96) interest pre-test 4.02 (1.14) 3.44 (1.13) 3.91 (1.11) 3.85 (1.23) post-test 4.39 (1.12) 3.43 (1.33) 3.94 (1.17) 3.73 (1.39) difference 0.37 (0.64) 0.01 (0.69) 0.03 (0.57) 0.12 (0.58) performance pre-test 0.53 (0.16) 0.49 (0.15) 0.52 (0.16) 0.50 (0.12) post-test 0.75 (0.16) 0.71 (0.16) 0.70 (0.17) 0.69 (0.12) difference 0.21 (0.12) 0.22 (0.12) 0.18 (0.11) 0.19 (0.10) enjoyment post-test 3.61 (0.50) 3.24 (0.69) 3.44 (0.72) 3.28 (0.65) note. sp = social position; pp = peer performance; performance variables reflect percentage of correct answers bosch et wilbert | f l r 69 table 2 shows the descriptive results separately for each experimental condition. one-factorial anovas with the experimental condition as independent variable showed no significant differences between groups for self-concept (f3, 118 = 0.78, p = .51), pre-test interest (f3, 118 = 1.45, p = .23) or pre-test performance (f3, 118 = 0.53, p = .67). hence, there were no significant pre-test differences between the groups in relevant study variables. 3.1 contrast and assimilation effects on task interest table 3 shows the model fit and parameter estimates for all models predicting task interest. model 2 including measurement time, social position and the interaction between social position and measurement time showed significantly better model fit compared to model 1 only containing measurement time (lr = 9.79, p < .01). inclusion of peer performance and interactions (model 3) did increase the model fit, albeit not significantly (lr = 7.72, p = .10). addition of the self-concept and interactions to all previous predictors (model 4) further increased the model fit (lr = 38.29, p < .001). as can be seen in model 1, measurement time did not significantly predict task interest, showing that there were no differences in task interest between measurement times across all experimental conditions. model 2 contained the hypothesized interaction between social position and measurement time, reflecting a more positive development of task interest in the high social position conditions, compared to the low social position conditions. model 3 further contained the hypothesized interaction between measurement time and peer performance as a significant predictor, once again reflecting a more positive interest development in the high peer performance compared to the low peer performance conditions. model 4 additionally included self-concept as a significant predictor, showing a strong association between self-concept and task interest. the interaction between self-concept and measurement time, however, was small and did not reach significance. self-concept also did not moderate interactions between social position and measurement time or between peer performance and measurement time. hence, even though task interest did not change from preto post-test measure across the whole sample, there were interactions between social position and measurement time as well as between peer position and measurement time, reflecting a more positive development of task interest in the high social position and high peer performance conditions compared to the low social position and low peer performance conditions, respectively. additionally, self-concept was associated with pre-test task interest, but did not influence the effect of experimental conditions on the development of task interest. therefore, our results support the hypothesized influence of contrast and assimilation effects on task interest. 3.2 contrast and assimilation effects on task performance table 4 shows the model fit and parameter estimates for all models predicting task performance. neither model 2 containing social position and the interaction between social position and measurement time (lr = 1.29, p = .52), nor model 3 additionally containing peer performance and corresponding interactions (lr = 2.91, p = .57), nor model 4 further containing self-concept and corresponding interactions (lr = 10.13, p = .26) showed significantly better model fit then the model containing only repeated measurements. model 1 did contain a significant predictor for measurement time, showing that our participants performed significantly better in the second compared to the first run. these differences in task performance were influenced by neither social position nor peer performance. bosch et wilbert | f l r 70 table 3 fixed effects for mixed models predicting interest (n = 122) unstandardized estimate b (se) parameter model 1 model 2 model 3 model 4 intercept 3.81** (0.111) 3.80** (0.109) 3.81** (0.109) 3.81** (0.099) level 1 mt 0.07 (0.058) 0.07 (0.057) 0.07 (0.056) 0.07 (0.057) level 2 sp 0.16 (0.109) 0.16 (0.109) 0.14 (0.099) pp 0.08 (0.109) 0.05 (0.099) sc 0.58** (0.098) interactions mt x sp 0.13* (0.057) 0.13* (0.056) 0.13* (0.057) mt x pp 0.12* (0.056) 0.12* (0.057) mt x sc 0.09 (0.056) sp x pp 0.13 (0.109) 0.05 (0.099) sp x sc 0.02 (0.098) pp x sc 0.14 (0.098) mt x sp x pp 0.06 (0.056) 0.07 (0.057) mt x sp x sc 0.04 (0.056) mt x pp x sc 0.06 (0.056) sp x pp x sc 0.01 (0.098) mt x sp x pp x sc 0.01 (0.056) model indices loglik 311.13 306.24 302.38 283.24 df 4 6 10 18 r²m .00 .04 .06 .27 r²c .86 .87 .88 .88 note. ** p < .01, * p < .05; mt = measurement time (0/1), sp = social position (-1/+1), pp = peer performance (-1/+1), sc = self-concept (centered), loglik = logarithmized likelihood of the model, df = degrees of freedom, r²m = marginal r², r²c = conditional r² bosch et wilbert | f l r 71 table 4 fixed effects for mixed models predicting task performance (n = 122) unstandardized estimate b (se) parameter model 1 model 2 model 3 model 4 intercept 0.51** (0.013) 0.51** (0.014) 0.51** (0.014) 0.61** (0.014) level 1 mt 0.20** (0.010) 0.20** (0.010) 0.20** (0.010) 0.10** (0.010) level 2 sp 0.02 (0.014) 0.02 (0.014) 0.02 (0.014) pp 0.01 (0.014) 0.00 (0.014) sc 0.02 (0.014) interactions mt x sp 0.00 (0.010) 0.00 (0.010) 0.00 (0.010) mt x pp 0.01 (0.010) 0.01 (0.010) mt x sc 0.02 (0.010) sp x pp 0.01 (0.014) 0.00 (0.014) sp x sc 0.02 (0.014) pp x sc 0.01 (0.014) mt x sp x pp 0.00 (0.010) 0.00 (0.010) mt x sp x sc 0.00 (0.010) mt x pp x sc 0.00 (0.005) sp x pp x sc 0.01 (0.014) mt x sp x pp x sc 0.00 (0.010) model indices loglik 161.70 162.34 163.80 168.86 df 4 6 10 18 r²m .31 .31 .32 .35 r²c .82 .82 .82 .82 note. ** p < .01, * p < .05; mt = measurement time (0/1), sp = social position (-1/+1), pp = peer performance (-1/+1), sc = self-concept (centered), loglik = logarithmized likelihood of the model, df = degrees of freedom, r²m = marginal r², r²c = conditional r² 3.3 contrast and assimilation effects on task enjoyment table 5 shows the model fit and parameter estimates for all four models predicting task enjoyment. model 2 containing the social position showed significantly better model fit then the intercept-only model 1 (lr = 5.15, p < .05), while inclusion of peer performance and the interaction between social position and peer performance in model 3 did not further increase the model fit (lr = 1.25, p = .54). inclusion of self-concept (model 4) also did not significantly increase model fit (lr = 2.63, p = .62). model 2 included social position as a significant predictor, while in model 3 neither peer performance nor the interaction between social position and peer performance significantly predicted task enjoyment. in model 4, self-concept was also not a significant predictor. bosch et wilbert | f l r 72 hence, while the social position condition did influence task enjoyment in the hypothesized direction (i.e. feedback of a high social position led to higher task enjoyment compared to feedback of a low social position), there was no effect of the peer performance condition on task enjoyment. table 5 fixed effects for mixed models predicting task enjoyment (n = 121) unstandardized estimate b (se) parameter model 1 model 2 model 3 model 4 intercept 3.40** (0.059) 3.39** (0.058) 3.39** (0.059) 3.40** (0.060) level 2 sp 0.13* (0.058) 0.13* (0.059) 0.12* (0.060) pp 0.04 (0.059) 0.04 (0.060) sc 0.05 (0.059) sp x pp 0.05 (0.059) 0.04 (0.060) sp x sc 0.03 (0.059) pp x sc 0.04 (0.059) sp x pp x sc 0.06 (0.059) model indices loglik 119.37 116.80 116.17 114.86 df 3 4 6 10 r²lr .00 .04 .05 .07 note. ** p < .01, * p < .05; sp = social position (-1/+1), pp = peer performance (-1/+1), sc = self-concept (centered), loglik = logarithmized likelihood of the model, df = degrees of freedom, r²lr = likelihood ratio r² 4. discussion in this study, we tested whether contrast and assimilation effects on task interest, which are thought to be the basis of popular bflp and birg effects, could be replicated in an experimental paradigm. as expected, the feedback of both a high social position and high peer performance had a positive influence on the change in our participants’ interest over the course of this study (compared to low social position and peer performance, respectively). when considering the sample as a whole, there was only a minor, non-significant increase in task interest between preand post-test. however, when taking a closer look at group differences in interest development, a pattern matching the hypothesized additive effects of contrast and assimilation effects could be found. while feedback of high social position / high peer performance showed a clear increase, task interest remained virtually unchanged in both intermediate groups (i.e. feedback of high social position / low peer performance and feedback of low social position / high peer performance). further, feedback of low social position / low peer performance even caused a decrease in task interest. hence, this study provides additional experimental evidence for the hypothesis that contrast and assimilation effects of performance (as described in the iem by schwarz and bless, 1992), and subsequent positive or negative performance feedback are likely to be at the core of bflp and birg effects in task interest. bosch et wilbert | f l r 73 however, when it comes to task interest social comparative feedback is a zero-sum game. not only are there winners, but also a rather consistent group whose interest is systematically undermined by continuing processes of negative social comparative evaluation, which could lead to further decline of academic achievement (möller et al., 2011; schiefele, 1992). this tends to be less of a problem in relatively homogeneous groups, where shortcomings in prior knowledge can more easily be overcome by increased effort. in rather heterogeneous groups, however, the influence of effort on academic evaluation is much lower, resulting in continuing negative social-comparative evaluation and therefore diminishing academic interest in slower learners. hence, the focus on grades as the most popular form of formal feedback in institutionalized learning settings should be re-evaluated in order to facilitate effective learning for all students. butler (1992), for example, could show that task interest for sixth graders differed depending on whether the instruction focused on task mastery or social comparison as well as on their self-rated performance. subjects who rated their performance highly showed higher interest in the social comparison condition, while subjects with low self-rated performance showed higher interest in the mastery condition. the results from harks, rakoczy, hattie, besser, and klieme (2014) point in a similar direction: process-oriented feedback was perceived as more useful by ninth graders than grade-oriented feedback and also had a positive effect on the development of mathematics achievement and interest. feedback of a high social position did additionally lead to the hypothesized increase in task enjoyment compared to a low social position, suggesting a similar contrast effect of task performance on task enjoyment. peer performance did not influence task enjoyment in the hypothesized direction, however, not supporting our hypothesis of an assimilation effect in task enjoyment. hence, an effect of social comparative performance feedback on task enjoyment similar to the one on task interest could only be partially confirmed. further, we explored associations between self-concept, task interest and task performance and potential reference group effects on task interest. we were able to confirm the hypothesized positive associations between self-concept, interest and performance in an academic learning task. hence, a positive self-concept usually coincided with higher interest as well as performance in our learning task. 4.1 limitations the experimental paradigm of this study was chosen to increase internal validity and further help to shed light on the exact mechanisms behind the contrast and assimilation effects of academic achievement. the downside is that the learning task in this study was strictly linear and lacked several dimensions of classroom learning (e.g. cooperation, teacher-student interaction, etc.). therefore, experimental studies with learning situations more closely resembling those of classroom learning are necessary in order to gain a better understanding of the influence of contrast and assimilation effects on interest in actual school settings. especially the operationalization of the two different peer performance conditions focuses on criterial differences in test score and lacks important aspects of the birg effect (e.g. being chosen into a highly valued group; highly visible selection process) that have been suggested (marsh, 1987; marsh, kong, & hau 2000). further, pre-/post-differences in interest between the two conditions with individual scores of 49 % (i.e. high social position / low peer performance and low social position / high peer performance) were relatively small, possibly hinting at effects from the criterial score on interest that could not be controlled within the design of this study. additionally, self-concept was only measured once at the beginning of the study. the addition of a post-learning-task measurement of self-concept would have enabled a deeper investigation of potential interactions between self-concept and task interest. bosch et wilbert | f l r 74 keypoints self-concept, interest and performance show hypothesized positive associations in an academic learning task both high social position and high criterial peer performance could be shown to positively influence the development of task interest in an experimental paradigm social comparative feedback seems to promote interest only for those receiving feedback above average, while having contrary effects for those receiving feedback below average references ames, c. (1992). classrooms: goals, structures, and student motivation. journal of educational psychology, 84(3), 261–271. https://doi.org/10.1037/0022-0663.84.3.261 bless, h., & schwarz, n. (2010). mental construal and the emergence of assimilation and contrast effects. advances in experimental social psychology, 42, 319–373. https://doi.org/10.1016/s00652601(10)42006-7 butler, r. (1992). what young people want to know when: effects of mastery and ability goals on interest in different kinds of social comparisons. journal of personality and social psychology, 62(6), 934– 943. https://doi.org/10.1037/0022-3514.62.6.934 chiu, m.-s. (2012). the internal/external frame of reference model, big-fish-little-pond effect, and combined model for mathematics and science. journal of educational psychology, 104(1), 87–107. https://doi.org/10.1037/a0025734 deci, e. l., & ryan, r. m. (1985). intrinsic motivation and self-determination in human behavior. new york: plenum. denissen, j. j. a., zarrett, n. r., & eccles, j. s. (2007). i like to do it, i’m able, and i know i am: longitudinal couplings between domain-specific achievement, self-concept, and interest. child development, 78(2), 430–447. https://doi.org/10.1111/j.1467-8624.2007.01007.x dickhäuser, o., & galfe, e. (2004). besser als ..., schlechter als... leistungsbezogene vergleichsprozesse in der grundschule. zeitschrift für entwicklungspsychologie und pädagogische psychologie, 36(1). dijkstra, p., kuyper, h., van der werf, g., buunk, a. p., & van der zee, y. g. (2008). social comparison in the classroom: a review. review of educational research, 78(4), 828–879. https://doi.org/10.3102/0034654308321210 everitt, b., & hothorn, t. (2011). an introduction to applied multivariate analysis with r. new york, ny: springer new york. https://doi.org/10.1007/978-1-4419-9650-3 fryer, l. k. (2015). predicting self-concept, interest and achievement for first-year students: the seeds of lifelong learning. learning and individual differences, 38, 107–114. https://doi.org/10.1016/j.lindif.2015.01.007 goetz, t., preckel, f., zeidner, m., & schleyer, e. (2008). big fish in big ponds: a multilevel analysis of test anxiety and achievement in special gifted classes. anxiety, stress & coping, 21(2), 185–198. https://doi.org/10.1080/10615800701628827 guay, f., marsh, h. w., & boivin, m. (2003). academic self-concept and academic achievement: developmental perspectives on their causal ordering. journal of educational psychology, 95(1), 124–136. https://doi.org/10.1037/0022-0663.95.1.124 harks, b., rakoczy, k., hattie, j., besser, m., & klieme, e. (2014). the effects of feedback on achievement, interest and self-evaluation: the role of feedback’s perceived usefulness. educational psychology, 34(3), 269–290. https://doi.org/10.1080/01443410.2013.785384 huang, c. (2011). self-concept and academic achievement: a meta-analysis of longitudinal relations. journal of school psychology, 49(5), 505–528. https://doi.org/10.1016/j.jsp.2011.07.001 bosch et wilbert | f l r 75 huguet, p., dumas, f., marsh, h., régner, i., wheeler, l., suls, j., … nezlek, j. (2009). clarifying the role of social comparison in the big-fish–little-pond effect (bflpe): an integrative study. journal of personality and social psychology, 97(1), 156–170. https://doi.org/10.1037/a0015558 jansen, m., scherer, r., & schroeders, u. (2015). students’ self-concept and self-efficacy in the sciences: differential relations to antecedents and educational outcomes. contemporary educational psychology, 41, 13–24. https://doi.org/10.1016/j.cedpsych.2014.11.002 jonkmann, k., becker, m., marsh, h. w., lüdtke, o., & trautwein, u. (2012). personality traits moderate the big-fish–little-pond effect of academic self-concept. learning and individual differences, 22(6), 736–746. https://doi.org/10.1016/j.lindif.2012.07.020 köller, o., schnabel, k. u., & baumert, j. (2000). der einfluß der leistungsstärke von schulen auf das fachspezifische selbstkonzept der begabung und das interesse. zeitschrift für entwicklungspsychologie und pädagogische psychologie, 32(2), 70–80. https://doi.org/10.1026//0049-8637.32.2.70 magee, l. (1990). r 2 measures based on wald and likelihood ratio joint significance tests. the american statistician, 44(3), 250. https://doi.org/10.2307/2685352 marsh, h. w. (1986). verbal and math self-concepts: an internal/external frame of reference model. american educational research journal, 23(1), 129149. https://doi.org/10.3102/00028312023001129 marsh, h. w. (1987). the big-fish-little-pond effect on academic self-concept. journal of educational psychology, 79(3), 280–295. https://doi.org/10.1037/0022-0663.79.3.280 marsh, h. w. (1991). failure of high-ability high schools to deliver academic benefits commensurate with their students’ ability levels. american educational research journal, 28(2), 445. https://doi.org/10.2307/1162948 marsh, h. w., & hau, k.-t. (2003). big-fish--little-pond effect on academic self-concept: a cross-cultural (26-country) test of the negative effects of academically selective schools. american psychologist, 58(5), 364–376. https://doi.org/10.1037/0003-066x.58.5.364 marsh, h. w., kong, c.-k., & hau, k.-t. (2000). longitudinal multilevel models of the big-fish-little-pond effect on academic self-concept: counterbalancing contrast and reflected-glory effects in hong kong schools. journal of personality and social psychology, 78(2), 337–349. https://doi.org/10.1037//0022-3514.78.2.337 marsh, h. w., & o’mara, a. j. (2010). long-term total negative effects of school-average ability on diverse educational outcomes: direct and indirect effects of the big-fish-little-pond effect. zeitschrift für pädagogische psychologie, 24(1), 51–72. https://doi.org/10.1024/1010-0652/a000004 marsh, h. w., & parker, j. w. (1984). determinants of student self-concept: is it better to be a relatively large fish in a small pond even if you don’t learn to swim as well? journal of personality and social psychology, 47(1), 213–231. https://doi.org/10.1037/0022-3514.47.1.213 marsh, h. w., & shavelson, r. (1985). self-concept: its multifaceted, hierarchical structure. educational psychologist, 20(3), 107–123. https://doi.org/10.1207/s15326985ep2003_1 marsh, h. w., trautwein, u., lüdtke, o., & köller, o. (2008). social comparison and big-fish-little-pond effects on self-concept and other self-belief constructs: role of generalized and specific others. journal of educational psychology, 100(3), 510–524. https://doi.org/10.1037/0022-0663.100.3.510 marsh, h. w., trautwein, u., lüdtke, o., köller, o., & baumert, j. (2005). academic self-concept, interest, grades, and standardized test scores: reciprocal effects models of causal ordering. child development, 76(2), 397–416. https://doi.org/10.1111/j.1467-8624.2005.00853.x möller, j., retelsdorf, j., köller, o., & marsh, h. w. (2011). the reciprocal internal/external frame of reference model: an integration of models of relations between academic achievement and selfconcept. american educational research journal, 48(6), 1315–1346. https://doi.org/10.3102/0002831211419649 mulkey, l. m., catsambis, s., steelman, l. c., & crain, r. l. (2005). the long-term effects of ability grouping in mathematics: a national investigation. social psychology of education, 8(2), 137–177. https://doi.org/10.1007/s11218-005-4014-6 bosch et wilbert | f l r 76 nagengast, b., & marsh, h. w. (2011). the negative effect of school-average ability on science self-concept in the uk, the uk countries and the world: the big-fish-little-pond-effect for pisa 2006. educational psychology, 31(5), 629–656. https://doi.org/10.1080/01443410.2011.586416 nakagawa, s., & schielzeth, h. (2013). a general and simple method for obtaining r2 from generalized linear mixed-effects models. methods in ecology and evolution, 4(2), 133–142. https://doi.org/10.1111/j.2041-210x.2012.00261.x pohlmann, b., & möller, j. (2006). vergleichseffekte auf kognitive, affektive und motivationale variablen. zeitschrift für entwicklungspsychologie und pädagogische psychologie, 38(2), 79–87. https://doi.org/10.1026/0049-8637.38.2.79 preckel, f., & brüll, m. (2010). the benefit of being a big fish in a big pond: contrast and assimilation effects on academic self-concept. learning and individual differences, 20(5), 522–531. https://doi.org/10.1016/j.lindif.2009.12.007 rheinberg, f., vollmeyer, r., & burns, b. d. (2001). fam: ein fragebogen zur erfassung aktuller motivation in lernund leistungssituationen. diagnostica, 47(2), 57–66. https://doi.org/10.1026//0012-1924.47.2.57 schiefele, u. (1992). topic interest and leveles of text comprehension. in k. a. renninger, s. hidi, & a. krapp (eds.), the role of interest in learning and development (pp. 151–182). hillsdale, n.j: l. erlbaum associates. schiefele, u., krapp, a., & winteler, a. (1992). interest as a predictor of academic achievement: a metaanalysis of research. in k. a. renninger, s. hidi, & a. krapp (eds.), the role of interest in learning and development (pp. 183–212). hillsdale, n.j: l. erlbaum associates. schurtz, i. m., pfost, m., nagengast, b., & artelt, c. (2014). impact of social and dimensional comparisons on student’s mathematical and english subject-interest at the beginning of secondary school. learning and instruction, 34, 32–41. https://doi.org/10.1016/j.learninstruc.2014.08.001 schwarz, n., & bless, h. (1992). constructing reality and its alternatives: assimilation and contrast effects in social judgment. in l. l. martin & a. tesser (eds.), the construction of social judgments (pp. 217–245). hillsdale, n.j: lawrence erlbaum associates. seaton, m., marsh, h. w., & craven, r. g. (2009). earning its place as a pan-human theory: universality of the big-fish-little-pond effect across 41 culturally and economically diverse countries. journal of educational psychology, 101(2), 403–419. https://doi.org/10.1037/a0013838 shavelson, r. j., hubner, j. j., & stanton, g. c. (1976). self-concept: validation of construct interpretations. review of educational research, 46(3), 407. https://doi.org/10.2307/1170010 springer, l., stanne, m. e., & donovan, s. s. (1999). effects of small-group learning on undergraduates in science, mathematics, engineering, and technology: a meta-analysis. review of educational research, 69(1), 21–51. https://doi.org/10.3102/00346543069001021 strang, l., smith, m. d., & rogers, c. m. (1978). social comparison, multiple reference groups, and the self-concepts of academically handicapped children before and after mainstreaming. journal of educational psychology, 70(4), 487–497. trautwein, u., lüdtke, o., marsh, h. w., köller, o., & baumert, j. (2006). tracking, grading, and student motivation: using group composition and status to predict self-concept and interest in ninth-grade mathematics. journal of educational psychology, 98(4), 788–806. https://doi.org/10.1037/00220663.98.4.788 trautwein, u., lüdtke, o., marsh, h. w., & nagy, g. (2009). within-school social comparison: how students perceive the standing of their class predicts academic self-concept. journal of educational psychology, 101(4), 853–866. https://doi.org/10.1037/a0016306 valentine, j. c., dubois, d. l., & cooper, h. (2004). the relation between self-beliefs and academic achievement: a meta-analytic review. educational psychologist, 39(2), 111–133. https://doi.org/10.1207/s15326985ep3902_3 wang, z. (2015). examining big-fish-little-pond-effects across 49 countries: a multilevel latent variable modelling approach. educational psychology, 35(2), 228–251. https://doi.org/10.1080/01443410.2013.827155 bosch et wilbert | f l r 77 wilbert, j., grosche, m., & gerdes, h. (2010). effects of evaluative feedback on rate of learning and task motivation: an analogue experiment. learning and disabilities: a contemporary journal, 8(2), 43–52. wouters, s., de fraine, b., colpin, h., van damme, j., & verschueren, k. (2012). the effect of track changes on the development of academic self-concept in high school: a dynamic test of the bigfish–little-pond effect. journal of educational psychology, 104(3), 793–805. https://doi.org/10.1037/a0027732 zeidner, m., & schleyer, e. j. (1999). the big-fish–little-pond effect for academic self-concept, test anxiety, and school grades in gifted children. contemporary educational psychology, 24(4), 305– 329. https://doi.org/10.1006/ceps.1998.0985 microsoft word osterhaus et al_publication.docx ! ! ! ! ! frontline)learning)research)vol.3)no.)4)(2015))56)=)94) issn)2295=3159)) children’s understanding of experimental contrast and experimental control: an inventory for primary school christopher osterhausa, susanne koerbera, beate sodianb afreiburg university of education, department of psychology, germany bludwig-maximilians-university munich, department of psychology, germany article received 18 october / revised 29 november / accepted 4 december / available online 19 january abstract experimentation skills are a central component of scientific thinking, and many studies have investigated whether and when primary-school children develop adequate experimentation strategies. however, the answers to these questions vary substantially depending on the type of task that is used: while discovery tasks, which require children to engage in unguided experimentation, typically do not reveal systematic skills in primary school, choice tasks suggest an early use of adequate experimentation strategies. to acquire a more accurate description of primary-school experimentation, this article proposes a novel multiple-select paper-and-pencil inventory that measures children’s understanding of experimental design. the two reported studies investigated the psychometric properties of this instrument and addressed the development of primaryschool experimentation. study 1 assessed the validity of the item format by comparing 2 items and an interview measure in a sample of 71 thirdand fourth-graders (9and 10-year-olds), while study 2 investigated the reliability and the convergent validity of the inventory by administering it to 411 second-, thirdand fourth-graders (8-, 9and 10-year-olds) and by comparing children’s performance in the 11-item scale to 2 conventional experimentation tasks. the obtained results demonstrate the reliability and validity of the inventory and suggest that a solid understanding of experimental design first emerges at the end of primary school. keywords: experimentation skills; control-of-variable strategy (cvs); scientific thinking; primary (elementary) school; paper-and-pencil assessment corresponding author: christopher osterhaus, department of psychology, freiburg university of education, kunzenweg 21, 79117 freiburg, germany. phone: +49(0)761 / 682-164, fax: +49(0)761 / 682 98164, email: osterhaus@ph-freiburg.de. doi: http://dx.doi.org/10.14786/flr.v3i4.220 osterhaus)et )al ) ) ) ) ) ) 57! ! 1. introduction experimentation skills constitute a fundamental component of scientific thinking, and many developmental research studies have investigated children’s acquisition of adequate experimentation strategies (e.g., case, 1974; inhelder & piaget, 1958; kuhn & phelps, 1982; siegler & liebert, 1975; for a review see zimmerman, 2007). two central research questions have been whether and when children begin to master the so-called control-of-variables strategy (cvs). this strategy, which is also referred to as the “vary-one-thing-at-a-time” strategy (tschirgi, 1980), requires informative experiments to contrast a single (focal) variable while keeping all other (non-focal) variables constant. most studies of experimentation skills probe children’s use of cvs in two different task settings: (1) discovery tasks with unrestricted variable configurations (production of cvs) and (2) choice tasks with restricted response options (choice of cvs). discovery tasks typically require children to explore the causal relations between different candidate causes and an outcome over a set of experimentation trials (e.g., kuhn et al., 1995; schauble, 1990). in each of these trials, mature reasoners form hypotheses about the system of independent variables (i.e., which variables are causal and which are non-causal), and based on these, they use cvs to isolate a single variable, which they contrast in an experiment that controls all non-focal variables (controlled-contrastive experiment). by updating their hypotheses and repeating this procedure over multiple trials, mature reasoners arrive at a final theory of the causal standing of each variable. discovery tasks hence involve reasoners in multiple phases of constructing scientific knowledge (i.e., reiterative formation of hypotheses, experimentation and data interpretation), and therefore they offer a high ecological validity and are particularly well suited for microgenetic studies that investigate strategy change in detail by repeatedly using the same task with a high density of observations in the short period of time when change is assumed to occur. choice tasks, in contrast, typically present children with a specific and directed hypothesis regarding one of the candidate causes. in addition, they include restricted answer options that represent distinct variable configurations from which children are allowed to choose (multiple choice [mc]; e.g., croker & buchanan, 2011; koerber, mayer, osterhaus, schwippert, & sodian, 2015; mayer, sodian, koerber, & schwippert, 2014; tschirgi, 1980). mature reasoners understand that they need to compare conditions, and their command of cvs is demonstrated by selecting the answer option in which only the focal variable is varied. this requirement of a single choice makes choice tasks easier to administer and hence well suited for large-scale, paper-and-pencil-based assessments, which are important research tools for investigating the large interindividual differences that already exist in primary-school experimentation (bullock, sodian, & koerber, 2009). despite their common measurement focus on cvs and their same basic task requirement (i.e., understanding that conditions need to be compared and an appropriate comparison needs to be made), discovery and choice tasks reveal a substantially different picture of primary-school children’s experimentation skills. while discovery tasks tend to reveal that only a small number of students produces experiments that are designed in accordance with cvs (e.g., 17% initially used cvs in a sample of fifthand sixth-graders; schauble, 1990), choice tasks reveal that most primary-school children prefer controlled-contrastive experiments over confounded experiments when they are allowed to choose between different experimental designs (e.g., 54% correct cvs choices by fourthgraders in koerber, mayer, et al., 2015; around 60% in bullock & ziegler, 1999). although differences between production and choice tasks (i.e., between open-answer and closed-response items) are a common empirical finding for knowledge scales (e.g., nehm & schonfeld, 2008), the large discrepancies between the two types of cvs tasks suggest that differences in performance are not solely attributable to the increased probability of correct guessing in choice tasks. performance differences between discovery and choice tasks might be attributable to their specific and discrepant task demands and solution strategies (see table 1). while discovery tasks osterhaus)et )al ) ) ) ) ) ) 58! ! require the repeated generation of hypotheses, which is an ability that may develop later than do experimentation skills (piekny & maehler, 2013) and which requires extensive memory and processing capacities, cvs choice tasks might not assess children’s understanding of experimentation because they can be solved by lower-level heuristics, such as varying a single variable while keeping the others constant, which does not need to be motivated by children’s full understanding of experimentation. other evidence, such as children’s justifications, is therefore required before concluding whether children fully understand the rationale for cvs. table 1 student performance in and characteristics of discovery tasks, choice tasks and understanding of experimental design (unex). discovery tasks choice tasks unex description iterative experimentation and use of cvs to uncover causal effects single choice of a (controlledcontrastive) experimental comparison identification of design errors in (controlled-) contrastive experiments task format open; often computerized mc; sometimes other formats that do not favour guessing (e.g., bullock & ziegler, 1999) ms use of cvs 17% initial use in at ages 11 and 12 (schauble, 1990) 8–68% (croker & buchanan, 2011), 54% (koerber, mayer, et al., 2015), and 60% (bullock & ziegler, 1999) all at age 10 29% at age 10 (cf. study 2) validity moderate guessing is not a problem low guessing is a problem (especially for mc items) high guessing is not a problem* task requires nonessential skills: tracking of variables (memory and processing skills), hypothesis generation may be solved by lower-level heuristics understanding features of experimental contrast and experimental control required reliability high performance estimates stable across problems low performance estimates differ across tasks and contexts high performance estimates stable across problems* assessment time-consuming rapid; allows for large-scale assessment rapid; allows for largescale assessment note. *characteristics investigated in the two studies presented in this article; cvs = control-ofvariables strategy; mc = multiple choice; ms = multiple select. a further criticism of both discovery and choice tasks is that studies of scientific thinking— with their narrow focus on cvs—have not explored the broader context of children’s understanding osterhaus)et )al ) ) ) ) ) ) 59! ! of the experimental method. in addition to cvs, a basic feature of experimentation is that planned comparisons are necessary to test for differences between conditions, as are randomization and an understanding of local control, which is necessary to reduce variation due to extraneous factors and which goes beyond the control of non-focal variables. bullock et al. (2009) developed an interview that investigates children’s metaconceptual understanding of experimental design (unex) in participants aged 12–22 years. this instrument asks children to review a set of fictitious experiments that contain diverse design errors that violate the principles of local control or experimental contrast. children need to recognize that ill-designed controlled-contrastive experiments do not provide an adequate test of hypothesis because variation due to extraneous factors is not reduced or non-focal variables are not controlled for (violation of cvs); they also must understand that hypothesis testing is impossible in ill-designed contrastive experiments because a single observation is made (either due to the missing variation of the focal variable, or the lack of an initial measurement in a pre-post design) and the principle of experimental contrast therefore is ignored. this latter experiment type provides no information about the control of nonfocal variables, and the ‘violated’ design feature is simply the principle of experimental contrast, which should be easier to understand than ‘experimental control’. primary-school children’s unex has not been investigated previously. however, bullock and her colleagues found that sixth-graders (the youngest age group interviewed with such an instrument) solved around 50% of tasks correctly, by identifying the experimental design error and justifying their opinion about whether the experiment was appropriately designed. it therefore appears that even very young children can show unex if they are questioned in an age-appropriate format that offers contextual support, such as providing a graphic representation of fictitious experiments and offering answer alternatives in a closed-response format. the present studies investigated whether a paper-and-pencil version of unex can provide valid and reliable measurements in primary-school children, in order to determine whether early abilities can be identified in this age group. our inventory of primary-school unex uses 11 closedresponse, multiple-select (ms) items (examples are provided in figure 1, and appendix 1 provides the full item set). for each item the children have to answer whether or not they consider a fictitious experiment to be well designed, and whether or not they agree with each of three separate justifications concerning the good or bad quality of the experiment (which includes one that identifies the design error in question). more conventional mc items provide a single choice, whereas the three separate decisions required by our ms procedure have two important advantages: (1) the probability of correct guessing is substantially reduced (i.e., 12.5% instead of 33% in the case of three statements), and (2) the ms format investigates potential inconsistencies in children’s understanding of experimentation. mc items only make it possible to conclude that children consider their chosen answer to be superior to the nonselected alternatives, whereas ms items give additional information about their view on all response options (i.e., although children may recognize the design error, they may still hold naïve beliefs regarding the production of effects). the correct and incorrect answer options used in the present studies are based on children’s answers to open items in prestudies and they draw on a conceptual-development model of scientific thinking (koerber, mayer, et al., 2015). specifically, each item includes naïve, intermediate and advanced-level answers. naïve-level answers reveal no understanding of hypothesis testing, instead referring to the production of an effect (cf. answer option 3 in figure 1), intermediate-level answers demonstrate a first understanding of the necessity of hypothesis testing and the existence of relevant design features (e.g., testing and sample size are important; cf. answer option 2), and advanced-level answers identify the specific design error and recognize that it restricts the information that can be drawn from the experiment (cf. answer option 1). osterhaus)et )al ) ) ) ) ) ) 60! ! figure 1. sample item. while research has shown that children’s (dis-)agreement with conceptually different levels on ms items matches the levels found in an interview conducted after exposure to the ms answer options (ms item then open interview [i-after]; cf. koerber, osterhaus, & sodian, 2015), little is known about the relation between children’s ms choices and the beliefs they hold before being presented with the ms item (open interview then ms item [i-before]). an instrument’s reliability and validity depends on a close relation between initial beliefs and the levels identified in the ms item. the present work therefore first investigated (in study 1) whether the ms item format is valid, resulting in ascriptions of levels that are significantly related with those levels found in i-before, and then (in study 2) addressed whether the inventory results in a reliable scale with satisfactory content and convergent validities. because primary-school children’s unex has not been investigated previously, study 2 also addressed the abilities and development of primary-school children. 2. study 1 study 1 investigated the validity of the ms item format by comparing children’s performance in two interview measures, which were conducted before (i-before) and after (i-after) the presentation of the closed-response answer options. study 1 also investigated the relation between the novel ms and conventional mc formats, as well as the relation between mc and the interview, both before and after the presentation of the ms task (i.e., i-before and i-after), to obtain a better understanding of the relative performances of diverse testing formats. 2.1 methods 2.1.1 participants the 71 included primary-school children comprised 56 third-graders (mean age 8 years, 8 months, sd=5 months; 31 girls) and 15 fourth-graders (mean age 10 years, 1 month; sd=4 months; 17 the hospital on planet mola, many molans are sick. in the hospital on planet “mola”, a scientist wants to find out whether the molans recover faster if they are allowed to receive visitors. he performs an experiment: 20 molans who have a heat disease are allowed to receive visitors for 2 hours per day during 2 weeks. 2 visiting hours 20 molans who have a broken antenna are not allowed to receive visitors during these 2 weeks. no visiting hours he compares both groups and finds out: after 2 weeks, all molans with heat disease have recovered. molans with a broken antenna, who have not received visitors, are still sick after 2 weeks. the scientist is convinced: “whether or not the molans recover fast depends on the visiting hours.” was this a good experiment? � yes � no 18 on planet “mola” susan, lisa, and vera think about whether this was a good experiment. who is right and who is not? is right is not right 1. susi says: “it was not a good experiment because the scientist should have compared molans with the same disease.” � � 2. lisa says: “it was a good experiment because he investigated many molans.” � � 3. vera says: “it was a good experiment because 20 molans have recovered from their disease.” � � which of the three girls has the best answer? no.______ osterhaus)et )al ) ) ) ) ) ) 61! ! 7 girls). children were recruited from three predominantly middle-class schools in germany. parental informed consent was obtained for all children. 2.1.2. materials the children were presented with and interviewed about (see procedure) two items from our inventory: one contrastive and one controlled-contrastive experiment. both experiments were set up in an artificial context in order to reduce interferences of children’s content knowledge on design evaluations. item 1 (contrastive experiment) presented the children with a story about a scientist who wants to test the hypothesis that yellow fertilizer increases plant growth significantly more than blue fertilizer. the scientist administers yellow fertilizer to 100 plants. he observes that all plants got bigger blossoms and therefore concludes that the yellow fertilizer works better than the blue one. children were asked whether or not this was a good experiment, and evaluated three explanations for their opinion (ms) on three hypothesized levels: (1) an explanation based on the production of effects (naïve level; “it was a good experiment because all plants that received the yellow fertilizer became bigger”), (2) an explanation that contained a design feature that was not the crucial to the experiment’s validity (intermediate level; “it was a good experiment because he tested the yellow fertilizer on many plants”) and (c) the correct explanation that identified the design error in question (advanced level; “it was not a good experiment because he only tested the yellow fertilizer”). after evaluating each of the three levels (ms), the children also indicated which of the three explanations they considered the best (mc). item 2 (controlled-contrastive experiment) presented the children with a story about two grandmothers who use lake or river water to water their plants. while those plants watered with lake water grow well, plants that are watered with river water are withering. children had to evaluate whether the experimental data was sufficient to support the hypothesis that lake water causes plants to grow well, or whether more information would be needed (i.e., control of non-focal variables). analogously to item 1, the three explanations reflected the distinct levels; however, while the naïve level as was the case for item 1 referred to the production of effects and the advanced level identified the design error in question, the intermediate level for this item included a reference to a potential causal mechanism rather than to a non-crucial design feature (i.e., “lake water contains more minerals than lake water. therefore, lake water is better for plants”). 2.1.3 procedure both items were read out loud to the children in a one-on-one interview. children marked their answers in their own booklets. interviews were conducted at two points during the presentation of the item: (1) after the children’s initial design evaluation and before presenting the ms answer options (ibefore; “why was the experiment good/bad?”), and (2) after presenting the answer options and the children choosing the best one (i-after; “why did you consider this answer to be the best one?”). both instances included follow-up questions such as “why did this [the reason children named] make the experiment a good/bad one?” or “would you have done anything differently?” osterhaus)et )al ) ) ) ) ) ) 62! ! table 2 coding of ms and interviews levels children agreed to naïve intermediate advanced coding of ms (final level) naïve x ---- naïve x x -- naïve x x x naïve x --x naïve ------ intermediate --x -- intermediate --x x advanced ----x coding of interviews sample answers naïve “because the fertilizer made the plants look more beautiful” (production of effect) intermediate “because he [the scientist] tried it on so many plants” (reference to non-relevant design feature) “because lake water is often dirty, whereas not much rubbish is thrown into river water” (reference to mechanism) advanced “because he [the scientist] did not try both fertilizers” (correct identification of design error) 2.1.4 transcription and coding of children’s answers interviews were audiotaped, transcribed verbatim and coded by two independent raters (see table 2 for coding examples and the ms coding). using a strict criterion, the lowest level which children agreed to was taken as the final level in the coding of the ms item (e.g., if a child accepted the naïve and intermediate levels simultaneously, the ms item was coded as naïve). the interrater kappa reliability values for i-before and i-after were .94 and .88, respectively, for item 1, and .73 and .90 for item 2. since some children gave invalid answers (especially on i-before; e.g., they refused to answer or only gave answers that were irrelevant to the question), some of the subsequent analyses involved a smaller sample. 2.2. results and discussion 2.2.1 core performance core performance data for items 1 and 2 (table 3) and a wilcoxon signed-rank test of i-before revealed that, as expected, more children recognized the design error in the contrastive (mdn=1) than the controlled-contrastive (mdn=0) experiment, z=2.92, p<.01, r=.30. osterhaus)et )al ) ) ) ) ) ) 63! ! 2.2.2 comparison of the different formats to investigate whether meaningful relations between the different formats in the two items existed (i.e., whether there was a high agreement in the levels assigned), we assessed spearman correlations (rho) and, in addition, computed wilcoxon signed-ranks statistics (z) when nonsignificant correlations suggested low convergence, which may have been a result of overor underestimation in one of the formats. for item 1 (table 4), correlations between formats were significant for all comparisons, except for i-before and mc, where the correlation did not reach significance. however, as a wilcoxon signed-ranks test revealed, there was no significant difference in difficulty of mc or i-before, and none of the formats led to a systematic overor underestimation. rather, 24% of all children performed better on mc than on i-before while 27% were assigned a higher level in i-before than in mc (cf. table 4). for item 2 (table 5), only the correlation between i-before and the ms task was significant. a wilcoxon signed-ranks test revealed that the mc question (mdn=1) significantly overestimated performance with respect to i-before. similarly, ms was significantly more difficult than mc, just as i-before was more difficult than i-after, suggesting that presenting answer options significantly decreased the difficulty of identifying the design error in this second item with a more complex experimental design. table 3 core performance data for items 1 and 2 (contrastive and controlled-contrastive experiment) item / format naïve level intermediate level advanced level total item 1 i-before 15 (33) 11 (24) 19 (42) 45 (100) ms 62 (87) 1 (1) 8 (11) 71 (100) mc 23 (32) 14 (20) 34 (48) 71 (100) i-after 18 (26) 29 (41) 23 (33) 70 (100) item 2 i-before 25 (51) 16 (33) 8 (16) 49 (100) ms 53 (75) 13 (18) 5 (7) 71 (100) mc 10 (14) 26 (37) 35 (49) 71 (100) i-after 21 (30) 26 (37) 23 (33) 70 (100) notes. data are n (%) values. i-before = open interview then ms item; i-after = ms item then open interview. osterhaus)et )al ) ) ) ) ) ) 64# # table 4 convergence between different item formats for item 1 (contrastive experiment) notes. data are n (%) values. int. = intermediate; adv. = advanced; z = wilcoxon signed-rank test. overand underestimation of the second comparison (b) relative to the first (a). rho and z are computed for children with complete observations on both measures for each test pair (i.e., children who did not give a valid answer on a measure that was part of the respective test pair were excluded from the respective analysis). * p<.05. ** p<.01. table 5 convergence between different item formats for item 2 (controlled-contrastive experiment) notes. data are n (%) values. ** p<.01. *** p<.001 . 1 r=.29. 2 r=.39. 3 r=.53 same level (a=b) overestimation (b>a) underestimation (ba) underestimation (b60%), and lower frequencies of intermediate and advanced conceptions (all <20% and <25%, respectively). 3.2.2 scale analysis the reliability of the scale was determined by fitting a partial credit model (masters, 1982) to the children’s responses. this model assumes that developmental progression is unidirectional and that categories (i.e., naïve, intermediate and advanced) are hierarchical, reflecting the assumption of our conceptual-development model. all but one item (item u5, a contrastive experiment) had a good fit to the model (i.e., 0.85> infit mean-square statistic [mnsq] <1.15; see table 6). removing this poorly fitting item yielded a scale with an expected a posteriori estimate based on a plausible values (eap/pv) reliability of .82 (weighted-likelihood estimator person separation reliability=.55, cronbach’s α=.85; scale mean=5.35, sd=5.85, minimum=0, maximum=22). all items satisfied the partial credit model’s requirements of increasing point-biserial correlations and ability estimates per category. however, none of our items fulfilled a third criterion (ordered delta parameters). it is debated in the literature as to whether this is a mandatory requirement for model fit to hold (adams, wu, & wilson, 2012), but the unordered delta parameters are consistent with children’s low frequency of choosing the intermediate level. we therefore dichotomized the data by collapsing the naïve and intermediate levels. a rasch model fitted to the resulting binary data revealed a good fit for all items (0.85< infit mnsq <1.15) and an eap/pv reliability of .71 (weightedlikelihood estimator person separation reliability=.62, cronbach’s α=.72). all subsequent analyses are based on these binary data. ! ! ! ! ! ! ! 67! ! table 6 percentage of answers on the naïve, intermediate and advanced level for the 11 unex items of study 2 overall, and the item difficulty, discrimination and item fit. notes. difficulty, discrimination and infit mean-square statistic (mnsq) are based on an analysis of the partial credit model; large negative item difficulties indicate easy items, while large positive values indicate difficult ones; infit mnsq should be between .85 and 1.15 for an item to show a good fit to the partial credit model; # indicates an item with a poor fit according to infit mnsq. item numbering applies to study 2 and the full item set given in the appendix. 3.2.3 developmental patterns and interindividual differences a univariate analysis of variance revealed a significant main effect in unex for grade, f(2,402)=21.75, p<.001, partial η2=.10. interestingly, while fourth-graders recognized significantly more design errors than did third-graders, f(1,402)=24.14, p<.001, partial η2=.06 (see figure 2), there was no difference between grades 2 and 3, f(1,402)=1.89, p>.05. there were also no differences between boys and girls, f(1,402)=3.18, p>.05. 3.2.4 content validity an explanatory item response model (de boeck & wilson, 2004) with a fixed person effect (age) revealed that controlled-contrastive experiments were, as hypothesized, more difficult than contrastive experiments (see table 7 for model fit, table 8 for parameter estimates). figure 2. percentage of correct answers for contrastive and controlled-contrastive experimental designs per grade. error bars indicate 95% confidence intervals. level item experiment type naïve (0) inter. (1) adv. (2) difficulty discrimination infit mnsq u1 contrastive 62.4 19.9 17.7 –.14 .57 1.10 u2 contrastive 67.8 9.3 22.9 –.32 .66 1.03 u3 contrastive 74.2 7.6 18.2 .20 .69 0.92 u4 contrastive 73.2 4.4 22.4 –.14 .66 1.03 u5 contrastive 65.2 10.4 24.3 –.48 .80 0.74# u6 contrastive 73.2 5.6 21.2 .01 .45 0.96 u7 controlled-contrastive 73.7 15.1 11.2 .56 .43 1.04 u8 controlled-contrastive 72.7 4.6 22.7 –.15 .63 1.05 u9 controlled-contrastive 72.3 11.5 16.2 .69 .49 0.90 u10 controlled-contrastive 68.1 12.9 19.9 –.04 .69 0.91 u11 controlled-contrastive 74.0 5.6 20.44 –.06 .72 0.89 ! ! ! ! ! ! ! 68! ! table 7 model comparisons for the explanatory item response model notes. m0 = reference model (1-parameter logistic model); aic = akaike information criterion; bic = bayesian information criterion; -2ll = deviance; df = degrees of freedom; lr test = likelihood ratio test. *** p<.001. table 8 coefficients for m2 3.2.5 convergent validity one-third (130) of the children chose cvs for both the cars and slopes tasks; however, only 55 children (14%) chose cvs consistently across the 2 items. the cars task was performed correctly by 20%, 33% and 44% of second-, thirdand fourth-graders, respectively; the corresponding rates for the slopes task were 18%, 35% and 46%. the overall performance was thus slightly worse than that reported by bullock and ziegler (1999), where approx. 40% and 60% of thirdand fourth-graders, respectively, correctly solved a cvs choice task. these differences are probably due to differences in sociodemographic characteristics between the previous urban sample and our more rural sample (cf. koerber, mayer, et al., 2015). interestingly, whereas the change-all strategy (vary all—including non-focal—variables) was the third most frequently chosen strategy for the cars task (20%), it was the least popular strategy for the slopes task (4%). in contrast to the cars task, where low performance resulted mostly from children disregarding the necessity of experimental control, errors in the slopes task were primarily due to children choosing an incorrect focal variable (i.e., size or position of the marble instead of slope steepness). while both these errors reflect unsuccessful coordination of hypothesis and evidence (cf. kuhn, 2011), they suggest that the content domain of the specific task and children’s knowledge thereof influence the expression of this interference between children’s hypotheses and their construction of evidence. a binomial regression revealed that performance in the cars task was significantly predicted by controlled-contrastive unex, χ2(1)=4.67, p=.03. specifically, children with a more profound controlledcontrastive unex were more likely to use cvs than any other strategy, β=.13, t(1)=4.62, p=.03, odds ratio=1.13. however, contrastive-unex experiments did not predict strategy choice in the cars task, β=.02, t(1)=.11, p=.74. in contrast, the use of cvs in the slopes task was predicted by contrastive unex, χ2(1)=8.16, p=.004. specifically, children with a high contrastive unex showed an increased use of cvs, β=.13, t(1)=8.02, p=.005, odds ratio=1.14, while there was no positive effect of controlled-contrastive model effects (fixed) effects (random) aic bic -2ll df lr test m0 (1pl) --intercept 3493.3 3506.1 3489.3 m1 experiment type intercept 3483.3 3502.6 3477.4 2 11.87*** m2 experiment type + age intercept 3465.4 3491.0 3457.4 1 20.01*** model b se p intercept –7.20 1.14 < .001 age .54 .12 < .001 controlled-contrastive (reference: contrastive) –.33 .01 < .001 ! ! ! ! ! ! ! 69! ! unex, β=–.08, t(1)=2.03, p=.15. this finding is consistent with the descriptive data for the slopes task, which suggest that few children choose an experiment in which all non-focal variables are varied. these results might be due to children only contrasting non-focal variables when their individual influences are unknown to them and they want to find out about them in a single test (cf. schauble, 1990). while 68% and 56% of the children with mastery in the controlled-contrastive unex (≥four items correctly solved) solved the cars and slopes tasks correctly, 70% and 68% of the incompetent children ( 0.95, ifi > 0.95, srmr < 0.08, and rmsea < .06 to indicate acceptable model fit. table 4 explained total variance of factors in the transfer contexts study and work 3. results the study aimed to estimate the extent to which peer support (hypothesis 1), supervisor support (hypothesis 2), openness to change (hypothesis 4), and feedback (hypothesis 5) influenced pre-training intention to transfer in the two transfer contexts study and work. table 5 presents the means, standard deviations, reliability estimates, and intercorrelations among all factors. table 5 correlation matrix of all variables the five-factor model yielded an acceptable fit in both transfer contexts. table 6 presents the psychometric properties. in the transfer context study, the x2 was 99.43 (df = 80), cfi = 0.99, ifi = 0.99, srmr = 0.04, and rmsea = 0.03 (90 % ci = 0.00, 0.05). in the transfer context work, the x2 was 206.39 (df = 80), cfi = 0.96, ifi = 0.96, srmr = 0.04, and rmsea = 0.07 (90 % ci = 0.06, 0.08). these estimates suggest acceptable model fit that was slightly better for the transfer context study compared to work. table 6 goodness-of-fit indices of the structural models in the transfer contexts study and work figures 2 and 3 present the model parameter estimates of the structural relations among factors for the transfer contexts study and work, respectively. in the transfer context study, intention to transfer was positively predicted by feedback (β = 0.37, p < 0.01), supervisor support (β = 0.31, p < 0.01), and openness to change (β = 0.10, p < 0.05); the relationship between peer support and intention to transfer (β = 0.04) was statistically non-significant. in the transfer context work, intention to transfer was predicted by supervisor support (β = 0.31, p < 0.01) and feedback (β = 0.19, p < 0.01); the relationship of intention to transfer with peer support (β = -0.05) and openness to change (β = -0.04) were statistically non-significant. figure 2. measurement and structural model parameter estimates of transfer context: study figure 3. measurement and structural model parameter estimates of transfer context: work a comparison between the two transfer contexts study and work indicates that the model parameter estimates between the independent and dependent variables were higher in the study context. table 7 presents the differences between beta coefficients for all variables. the highest difference emerged for feedback (study context: β = 0.37, work context: β = 0.19, δ = 0.18) followed by openness to change (study context: β = 0.10, work context: β = -0.04, δ = 0.14). these analyses tend to indicate the benefits of examining multiple transfer contexts when estimating organizational predictors of intention to transfer. table 7 comparison of beta coefficients between the two transfer contexts study and work 4. discussion the goal of this study was to complement previous research on the intention to transfer new learning by investigating the students' pre-course perception of the influence of five support-related variables on their intention to transfer learning to both their study and their work context. in educational settings transfer of learning is typically measured during a test after the training. the first finding of our study is that already before the actual training, in this case a course in generic information literacy competences, five of the eight variables used for both the study and work context were considered significant for the students’ intention to transfer competences they were about to learn during the course. from the literature we learn that transfer is a dynamic process with a temporal dimension, influenced by a multitude of variables, not only during and after, but also already before an intervention (baldwin & ford, 1988; broad, 2005, burke & hutchins, 2008; gegenfurtner, veermans, festner, & gruber, 2009; grossman & salas, 2011; holton & baldwin, 2003). this study confirms this influence in the first pre-training stage of the transfer process. instructional designers might take this into account when framing and presenting a training. (baldwin & magjuka, 1991). preceding a training, for example when designing the course programme but also later on during the course, they might communicate to the students the importance of the variables that appear to be significant predictors of intention to transfer in the students' study and work context and how they are integrated into the course. follow-up studies will investigate the other temporal dimensions. a second finding from the model parameter estimates shows that there is a difference between the beta coefficients for the two transfer contexts study and work. this is especially relevant in situations where it involves education in so-called generic competences like information literacy that can be, or are meant to be applied in various contexts, be it study, work or private life. this finding tends to indicate that, for educational designers, it is beneficial to gain insight into the conditions of, and main actors within the intended transfer contexts. one option would be to involve the input of former students when designing and framing a specific course, for example by using course evaluation reports. a third finding refers to the variable 'opportunity to use' that was dropped from this study. in previous reviews and studies opportunity appeared to be one of the strongest predictors of transfer. in our exploratory factor analysis it loaded onto one factor together with intention to transfer. correlation analysis also showed a relatively strong positive relationship. this might be difficult to explain when looking at the items that were used for both constructs. the ones used for intention referred to the respondents' plans or efforts to apply new learning, while items for opportunity to use specifically referred to workload and the availability of sufficient facilities and opportunities to apply new learning. follow-up interviews with respondents might have shed more light on why they have interpreted both variables in the same manner. looking at the differences between the two transfer contexts in more detail we notice that in the study context feedback (0.37, hypothesis 5), supervisor support (0.31, hypothesis 2), and openness to change (0.10, hypothesis 4) appear to be significant moderate predictors of the students' intention to transfer learning. this is not surprising as especially in an educational context, where students are gaining new competencies, feedback and the support by supervisors, in this case lecturers, is considered essential to, and even a precondition for the proper application of new learning. communicating, for example in the course guide, that both are important aspects of the course might boost the students' pre-course self-confidence and subsequently their intention to apply their newly acquired competences. another aspect within the study context that is significant for the students' intention to transfer, although with a significantly lower beta coefficient, is openness to change. this might be understandable as behavioural change is part and parcel of educational settings and openness to these changes might therefore be taken for granted by the students. the specific educational setting in this research, distance education at the open university with little or no direct contact between students, might explain why peer support isn't seen as significant for the intention to transfer setting. in the students' work context supervisor support (0.31, hypothesis 2) and feedback (0.19, hypothesis 5) appear to be significant for the students' intention to apply new learning. one explanation might be that doing a master programme beside work is a challenging, time and energy consuming activity that will not only improve the learners' individual competences but will also be beneficial for the quality of the organization at large. students therefore might expect that when applying these new competences this will be acknowledged and facilitated by the organization in the person of the supervisors. when it involves an in-company training instructional designers might pay special attention to the organizational support that is available or might have to be organized in order to enhance transfer of learning. as educational designers generally don't have any influence on the supportive conditions within the students' work context they might integrate discussions about these aspects in the course, asking students how relevant feedback and supervisor support are for them, if and how both are available in their organization and in case it is lacking, what steps might be taken to organize it. these discussions might also give an indication of what kind of support is expected. extending the current research by interviewing respondents might also give an explanation of why support of colleagues and an open, innovative team or organizational climate are not relevant for their application of new learning. 4.1. theoretical implications the theoretical relevance of this study is that it adds a novel dimension to the conceptual development of transfer studies. a multicontextual perspective on transfer is currently absent from the training literature, as are studies that systematically compare transfer to more than one context within one sample. we expect that studies addressing this multicontextuality of transfer will continue to confirm or challenge previous monocontextual research on the importance of environmental predictors in specific settings. 4.2 practical implications the practical value of this study lies in the field of educational design. the results confirm that transfer of learning isn't only happening after an intervention. it is a longitudinal process in which various aspects may influence the learners' intention to transfer new learning not only after but already before the actual intervention. furthermore, from a lifelong learning perspective and especially when generic competences are involved, it is important for educational designers not only to focus on the educational setting of an intervention but also to pay attention to aspects of current or prospective transfer contexts, be it study, work or private life, that might be relevant for the transfer process. in order to enhance multicontextual transfer, learners could be involved in the design process for their knowledge of the transfer context. 4.3 limitations and directions for future research one limitation of this study is the use of self-report surveys. although there are valid arguments against using self-report like social desirability or common method bias, some specific situations might prevent additional external measurements. this can be the case when it involves a construct like intention that is difficult to estimate with behavioural measures, or when it involves more autonomously working respondents. previous studies on the other hand indicate that respondents are very well capable of reporting on their own transfer process (chiaburu & tekleab, 2005; devos, dumay, bonami, bates, & holton, 2007; facteau et al., 1995; velada et al., 2007) and that they themselves know best which variables they consider relevant for their intention to transfer. we have attempted to limit undesirable bias by communicating that data was collected anonymously and by offering the opportunity to answer the surveys electronically and therefore in private. for future research however data triangulation using additional measurement instruments would be advisable. another limitation is the setting in which this study has taken place namely adult learners in a distance education context. distance learning creates specific conditions that may have implications for some transfer related variables like peer support. furthermore, according to heery (1996, p. 5) non-traditional students like distance learners "often show an unusual degree of motivation and commitment”. we therefore recommend future studies to be carried out in other learning contexts, for example at regular universities of applied sciences where learning is also directed towards transfer of new learning to study and work contexts. finally, the course and also the integrated survey were mandatory. various studies recommend voluntary enrolment in order to enhance motivation and transfer (gegenfurtner, könings, kosmajac, & gebhardt, 2016) while others argue that mandatory participation enhances transfer as it expresses the importance of the course to the organization (baldwin & magjuka, 1991). future research might investigate the implications of these differences in more detail. future steps in this research will investigate the influence of learner characteristics and intervention variables on the intention to transfer learning to multiple contexts, as well as their longitudinal development over time, measured directly after and three months after the course. despite the extensive body of knowledge on transfer of learning, including on the intention to transfer, studies typically have focussed on transfer within one specific context. a multicontextual focus however is in place when it concerns transfer of generic competences that can be, or are intended to be applied in more than one context. this has been confirmed by the results of the present study. furthermore, in contexts where it’s difficult to measure actual transfer, intention can function as a valuable precursor to and predictor of transfer. the present study indicates that variables and items used in the study offer a valuable contribution to the design of a practical instrument to measure the influence of a selection of key factors on the transfer process. this instrument in turn will help instructional designers to create educational interventions that enhance the transfer of learning. keypoints the present study extends previous research on transfer of learning by introducing a multicontextual perspective when designing training of generic competences that are meant to be applied in more than one context, for example study, work or daily life. we hypothesised positive relationships of five organizational variables on students' pre-training intention to transfer generic competences from the prospective training to both their study and work context. participants were 303 students at an open university starting a course in information literacy. data was collected by means of pre-course self-reports and analysed using structural equation modelling. before starting the course supervisor support and feedback were the strongest predictors of intention to transfer new learning in both the study and the work context. our study confirmed the presumption that transfer of learning is a process that already starts before an actual training. references aguinis, h., & kraiger, k. (2009). benefits of training and development for individuals and teams, organizations, and society. annual review of psychology, 60, 451-474.doi:10.1146/annurev.psych.60.110707.163505 ajzen, i. (1991). the theory of planned behavior. organizational behavior and human decision processes, 50(2), 179-211. doi:10.1016/0749-5978(91)90020-t ajzen, i., & madden, t. j. (1986). prediction of goal-directed behavior: attitudes, intentions, and perceived behavioural control. journal of experimental social psychology, 22(5), 453-474. doi:10.1016/0022-1031(86)90045-4 al-eisa, a.s., furayyan, m.a., & alhemoud, a.m. (2009). an empirical examination of the effects of self-efficacy, supervisor support and motivation to learn on transfer intention. management decision, 47 (8), 1221–1244. doi:10.1108/00251740910984514 allison, p. d. (2003). missing data techniques for structural equation modeling. journal of abnormal psychology, 112(4), 545-557. doi:10.1037/0021-843x.112.4.545 alvero, a. m., bucklin, b. r., & austin, j. (2001). an objective review of the effectiveness and essential characteristics of performance feedback in organizational settings (1985-1998). journal of organizational behavior management, 21(1), 3-29. doi:10.1300/j075v21n01_02 baldwin, t. t., & ford, j. k. (1988). transfer of tranining: a review and directions for future research. personnel psychology, 41(1), 63-105. doi:10.1111/j.1744-6570.1988.tb00632.x baldwin, t. t., & magjuka, r. j. (1991). organizational training and signals of importance: linking pretraining perceptions to intentions to transfer. human resource development quarterly, 2(1), 25-36. pp. 80–96. doi:10.1002/hrdq.3920020106 bandura, a. (1977). social learning theory (englewood cliffs, nj: prentice hall). bates, r. a., holton lll, e. f., & hatala, j. p. (2012). a revised learning transfer system inventory: factorial replication and validation. human resource development international, 15(5), 549-569. doi:10.1080/13678868.2012.726872 blume, b. d., ford, j. k., baldwin, t. t., & huang, j. l. (2010). transfer of training: a meta-analytic review. journal of management, 36(4), 1065-1105. doi:10.1177/0149206309352880 brand-gruwel, s., wopereis, i., & vermetten, y. (2005). information problem solving by experts and novices: analysis of a complex cognitive skill. computers in human behaviour, 21(3), 487-508. doi:10.1016/j.chb.2004.10.005 brand-gruwel, s., wopereis, i., & walraven, a. (2009). a descriptive model of information problem solving while using internet. computers & education, 53(4), 1207-1217. doi:10.1016/j.compedu .2009.06.004 brinkerhoff, r. o., & montesino, m. u. (1995). partnerships for training transfer: lessons from a corporate study. human resource development quarterly, 6(3), 263-274. doi:10.1002/hrdq.3920060305 broad, m. l., & newstrom, j. w. (1992). transfer of training: action packed strategies to ensure high payoff from training investments. reading, ma: addison-wesley. broad, m. l. (2005). beyond transfer of training: engaging systems to improve performance . john wiley & sons. burke, l. a., & hutchins, h. m. (2007). training transfer: an integrative literature review. human resource development review, 6(3), 263-296. doi:10.1177/1534484307303035 cheng, e. w. l., & hampson, i. (2008). transfer of training: a review and new insights. international journal of management review, 10(4), 327-341. doi:10.1111/j.1468-2370.2007.00230.x chiaburu, d. s., van dam, k., & hutchins, h. m. (2010). social support in the workplace and training transfer: a longitudinal analysis.international journal of selection and assessment, 18(2), 187–200. doi:10.1111/j.1468-2389.2010.00500.x. chiaburu, d. s., & marinova, s. v. (2005). what predicts skill transfer? an exploratory study of goal orientation, training self‐efficacy and organizational supports. international journal of training and development, 9(2), 110-123. doi:10.1111/j.1468-2419.2005.00225.x chiaburu, d. s., & tekleab, a. g. (2005). individual and contextual influences on multiple dimensions of training effectiveness. journal of european industrial training, 29(8), 604-626. doi:10.1108/03090590510627085 clarke, n. (2002). job/work environment factors influencing training transfer within a human service agency: some indicative support for baldwin and ford’s transfer climate construct. international journal of training and development, 6(3), 146–162. doi:10.1111/1468-2419.00156 cohen, s., underwood, l. g., & gottlieb, b. h. (eds.). (2000). social support measurement and intervention. oxford: oxford university press. colquitt, j. a., lepine, j. a., & noe, r. a. (2000). toward an integrative theory of training motivation: a meta-analytic path analysis of 20 years of research. journal of applied psychology, 85(5), 678–707. doi:10.1037//0021-9010.85.5.678 cromwell, s. e., & kolb, j. a. (2004). an examination of work-environment support factors affecting transfer of supervisory skills training to the workplace. human resource development quarterly, 15(4), 449–471. doi:10.1002/hrdq.1115 den ouden, m. d. (1992). transfer na bedrijfsopleidingen: een veldonderzoek naar de rol van voornemens, sociale normen, beheersing en sociale steun bij opleidingstransfer. doctoral dissertation. amsterdam: thesis publishers. de rijdt, c., stes, a., van der vleuten, c., & dochy, f. (2013). influencing variables and moderators of transfer of learning to the workplace within the area of staff development in higher education: research review. educational research review, 8(0), 48-74. doi:10.1016/j.edurev.2012.05.007 devos, c., dumay, x., bonami, m., bates, r., & holton, e. (2007). the learning transfer system inventory (ltsi) translated into french: internal structure and predictive validity. international journal of training and development, 11(3), 181-199. doi:10.1111/j.1468-2419.2007.00280.x egan, t. m., yang, b., & bartlett, k. r. (2004). the effects of organizational learning culture and job satisfaction on motivation to transfer learning and turnover intention. human resource development quarterly, 15(3), 279-301. doi:10.1002/hrdq.1104 facteau, j. d., dobbins, g. h., russell, j. e., ladd, r. t., & kudisch, j. d. (1995). the influence of general perceptions of the training environment on pretraining motivation and perceived training transfer. journal of management, 21(1), 1-25. doi:10.1177/014920639502100101 ford, j.k. (1997). advances in training research and practice: an historical perspective. in ford, j., kozlowski, s., kraiger, k., salas, e., & teachout, m. (eds), improving training effectiveness in work organizations (pp. 1-18) mahwah, nj: lawrence erlbaum associates ford, j. k., quinones, m. a., sego, d. j., & sorra, j. (1992). factors affecting the opportunity to perform trained tasks on the job. personnel psychology, 45(3), 511-527. doi:10.1111 /j.17446570.1992.tb00858.x ford, j.k., & weissbein, d. (1997). transfer of training: an updated review. performance improvement quarterly, 10(2) , 22–41. doi:10.1111/j.1937-8327.1997.tb00047.x ford, j. k., yelon, s. l., & billington, a. q. (2011). how much is transferred from training to the job? the 10% delusion as a catalyst for thinking about transfer. performance improvement quarterly, 24(2), 7-24. doi:10.1002/piq.20108 foxon, m. j. (1994). a process approach to transfer of training: part 2; using action planning to facilitate the transfer of training. australian journal of education and technology, 10(1), 1-18. doi:10.14742/ajet.2080 gabelica, c., van den bossche, p., segers, m., & gijselaers, w. (2012). feedback, a powerful lever in teams: a review. educational research review, 7(2), 123-144. doi:10.1016/j.edurev.2011.11.003 gegenfurtner, a. (2011). motivation and transfer in professional training: a meta-analysis of the moderating effects of knowledge type, instruction, and assessment conditions. educational research review, 6 (3), 153–168. doi:10.1016/j.edurev.2011.04.001 gegenfurtner, a. (2013). dimensions of motivation to transfer: a longitudinal analysis of their influence on retention, transfer, and attitude change. vocations and learning, 6(2), 187-205. doi:10.1007/s12186 -012-9084-y gegenfurtner, a., könings, k. d., kosmajac, n., & gebhardt, m. (2016). voluntary or mandatory training participation as a moderator in the relationship between goal orientations and transfer of training. international journal of training and development, 20(4), 290-301. doi:10.1111/ijtd.12089 gegenfurtner, a., veermans, k., festner, d., & gruber, h. (2009). integrative literature review: motivation to transfer training: an integrative literature review. human resource development review, 8(3), 403-423. doi:10.1177/1534484309335970 gilpin-jackson, y., & bushe, g. r. (2007). leadership development training transfer: a case study of posttraining determinants. the journal of management development, 26(10), 980-1004. doi:10.1108/02621710710833423 govaerts, n., & dochy, f. (2014). disentangling the role of the supervisor in transfer of training. educational research review, 12, 77-93. doi:10.1016/j.edurev.2014.05.002 grossman, r., & salas, e. (2011). the transfer of training: what really matters. international journal of training and development, 15(2), 103-120. doi:10.1111/j.1468-2419.2011.00373.x haskell, e. (2001). transfer of learning: cognition, instruction, and reasoning. san diego, usa: academic press. heery, m. (1996). academic library services to non-traditional students. library management, 17(5), 3-13. doi:10.1108/01435129610119584 hoekstra, m. r. (1999). gedragsbeïnvloeding door cursussen. een studie naar de effecten van persoons-, cursusen omgevingskenmerke. [influencing behaviour through training programmes: a study of the effects of personal, training programme and environmental characteristics]. amsterdam, the netherlands: vrije universiteit holton iii, e. f. (2005). holton's evaluation model: new evidence and construct elaborations. advances in developing human resources, 7(1), 37-54. doi:10.1177/1523422304272080 holton, iii, e. f., & baldwin, t. t. (2003). making transfer happen. in holton iii, e. f., & baldwin, t. t. (eds.) improving learning transfer in organizations (pp. 3-15). san francisco, ca: jossy-bass. holton iii, e. f., bates, r. a., & ruona, w. e. (2000). development of a generalized learning transfer system inventory. human resource development quarterly, 11(4), 333-360. doi: 10.1002/1532-1096(200024)11:4<333::aid-hrdq2>3.0.co;2-p holton iii, e. f., chen, h. c., & naquin, s. a. (2003). an examination of learning transfer system characteristics across organizational settings. human resource development quarterly, 14(4), 459–482. doi:10.1002/hrdq.1079 homklin, t., takahashi, y., & techakanont, k. (2014). the influence of social and organizational support on transfer of training: evidence from thailand, international journal of training and development, 18(2), 116-131. doi:10.1111/ijtd.12031 horner, m. t. (2010). toward an understanding of when and why situational constraints influence performance . doctoral dissertation, texas a & m university. hu, l. t., & bentler, p. m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives. structural equation modeling, 6(1), 1-55. doi:10.1080/10705519909540118 hutchins, h. m. and burke, l. a. (2007), identifying trainers' knowledge of training transfer research findings – closing the gap between research and practice. international journal of training and development, 11 (4), 236–264. doi:10.1111/j.1468-2419.2007.00288.x hutchins, h. m., nimon, k., bates, r., & holton iii, e. (2013). can the ltsi predict transfer performance? testing intent to transfer as a proximal transfer of training outcome. international journal of selection and assessment, 21(3), 251-263. doi:10.1111/ijsa.12035 katsioloudes, v. (2015). supervisory ratings as a measure of training transfer: testing the predictive validity of the learning transfer system inventory . [doctoral dissertation]. kline, r. b. (2015). principles and practice of structural equation modeling (4th ed.). new york: guilford. knowles, m., holton iii, e., & swanson, r. (1998). the adult learner (5th ed.). houston, tx: gulf publishing. kontoghiorghes, c. (2004). reconceptualizing the learning transfer conceptual framework: empirical validation of a new systemic model.international journal of training and development, 8(3), 210-221. doi:10.1111/j.1360-3736.2004.00209.x lim, d. h., & johnson, s. d. (2002). trainee perceptions of factors that influence learning transfer. international journal of training and development, 6(1), 36-48. doi:10.1111/1468-2419.00148 ling, g. j., & yusof, h. m. (2017). a review of the linkage between supervisory support and training transfer. sains humanika, 9(1-3). doi:10.11113/sh.v9n1-3.1140 massenberg, a. c., schulte, e. m., & kauffeld, s. (2017). never too early: learning transfer system factors affecting motivation to transfer before and after training programs. human resource development quarterly, 28(1), 55-85. doi:10.1002/hrdq.21256 massenberg, a. c., spurk, d., & kauffeld, s. (2015). social support at the workplace, motivation to transfer and training transfer: a multilevel indirect effects model. international journal of training and development, 19(3), 161-178. doi:10.1111/ijtd.12054 mathieu, j. e., martineau, j. w., & tannenbaum, s. i. (1993). individual and situational influences on the development of self-efficacy: implications for training effectiveness. personnel psychology, 46(1), 125–147. doi:10.1111/j.1744-6570.1993.tb00870.x mathieu, j.e., tannenbaum, s.i., & salas, e. (1992). influences of individual and situational characteristics on measures of training effectiveness. academy of management journal, 35(4), 828–847. doi:10.2307/256317 merriam, s. b., & leahy, b. (2005). learning transfer: a review of the research in adult education and training. paace journal of lifelong learning, 14(1), 1-24. naquin, s. s., & baldwin, t. t. (2003). managing transfer before learning begins: the transfer-ready learner. in e. f. holton, iii & t. t. baldwin (eds), improving learning transfer in organizations (pp. 80–96). san francisco, usa: jossey-bass. ng, k. h. (2013). the influence of supervisory and peer support on the transfer of training. studies in business & economics 8(3), 82-97. nijman, d. j. (2004). supporting transfer of training: effects of the supervisor [doctoral dissertation].(phd), enschede, the netherlands: universiteit twente. retrieved from http://doc.utwente.nl/76049/ nijman, d., nijhof, w., wognum, a., & veldkamp, b. (2006). exploring differential effects of supervisor support on transfer of training. journal of european industrial training, 30(7), 529–549. doi:10 .1108/03090590610704394 noe, r. a. (1986). trainees' attributes and attitudes: neglected influences on training effectiveness. academy of management journal, 11(4), 736-749. doi:10.5465/amr.1986.4283922 oxford dictionaries, british & world english, retrieved february 18, 2018 from https://en.oxforddictionaries.com/definition/feedback pidd, k. (2004). the impact of workplace support and identity on training transfer: a case study of drug and alcohol safety training in australia. international journal on training and development, 8(4), 274–288. doi:10.1111/j.1360-3736.2004.00214.x reece, g. j. (2007). critical thinking and cognitive transfer: implications for the development of online information literacy tutorials. research strategies, 20, 482-493. doi:10.1016/j.resstr.2006.12.018 reinhold, s., gegenfurtner, a., & lewalter, d. (2018). social support and motivation to transfer as predictors of training transfer: testing full and partial mediation using meta-analytic structural equation modeling. international journal of training and development, 22(1). doi:10.111/ijtd.12115 de rijdt, c., stes, a., vleuten, c. v. d., & dochy, f. (2013). influencing variables and moderators of transfer of learning to the workplace within the area of staff development in higher education: research review. educational research review, 8, 48-74. doi:10.1016/j.edurev.2012.05.007 rouiller, j. z., & goldstein, i. l. (1993). the relationship between organizational transfer climate and positive transfer of training. human resource development quarterly, 4(4), 377-390. doi:10.1002/hrdq.3920040408 russ-eft, d. (2002). a typology of training design and work environment factors affecting workplace learning and transfer. human resource development review, 1(1), 45-65. doi:10.1177/1534484302011003 salas, e., & cannon-bowers, j. a. (2001). the science of training: a decade of progress. annual review of psychology, 52, 471-499. doi:10.1146/annurev.psych.52.1.471 seyler, d. l., holton iii, e. f., bates, r. a., burnett, m. f., & carvalho, m. a. (1998). factors affecting motivation to transfer training. international journal of training and development, 2(1), 2-16. doi:10.1111/1468-2419.00031 sheeran, p. (2002). intention-behavior relations: a conceptual and empirical review. in w. stroebe, & m. hewstone (eds.), european review of social psychology, vol. 12 (pp. 1–36). london: wiley. simons, p. r. j. (1999). transfer of learning: paradoxes for learners. international journal of educational research, 31(7), 577-589. doi:10.1016/s0883-0355(99)00025-7 smith-jentsch, k. a., salas, e., & brannick, m. t. (2001). to transfer or not to transfer? investigating the combined effects of trainee characteristics, team leader support, and team climate. journal of applied psychology, 86(2), 279-292. doi:10.1037/0021-9010.86.2.279 taylor, p. j., russ-eft, d. f., & chan, d. w. l. (2005). a meta-analytic review of behavior modeling training, journal of applied psychology, 90(4), 692–709. doi:10.1037/0021-9010 .90.4.692 thayer, p. w., & teachout, m. s. (1995). a climate for transfer model. brooks air force base, texas. thurstone, l. l. (1947). multiple-factor analysis. chicago: the university of chicago press. tonhäuser, c., & büker, l. (2016). determinants of transfer of training: a comprehensive literature review.international journal for research in vocational education and training, 3(2), 127-165. doi:10.13152/ijrvet.3.2.4 tracey, j. b., hinkin, t. r., tannenbaum, s., & mathieu, j. e. (2001). the influence of individual characteristics and the work environment on varying levels of training outcomes. human resources development quarterly, 12(1), 5–23. doi:10.1002/1532-1096(200101/02)12:1<5::aid-hrdq2>3.0.co;2-j tracey, j. b., tannenbaum, s. i., & kavanagh, m.j. (1995). applying trained skills on the job: the importance of the work environment. journal of applied psychology, 80(2), 239–252. doi:10.1037/0021-9010.80.2.239 triandis, h. c. (1980). values, attitudes, and interpersonal behavior. in h. e. howe & m. page (eds.), nebraska symposium of motivation, 27 (pp. 195–259). lincoln, ne: university of nebraska press. tziner, a., haccoun, r. r., & kadish, a. (1991). personal and situational characteristics influencing the effectiveness of transfer of training improvement strategies, journal of occupational psychology, 64(2), 167–77. doi:10.1111/j.2044-8325.1991.tb00551.x van den bossche, p., segers, m., & jansen, n. (2010). transfer of training: the role of feedback in supportive social networks, international journal of training & development, 14 (2), 81–94. doi:10.1111/j.1468-2419.2010.00343.x van der klink, m., gielen, e., & nauta, c. (2001). supervisory support as a major condition to enhance transfer. international journal of training and development, 5(1), 52-63. doi: 10.1111/1468-2419.00121 van merriënboer, j. j. g., & kirschner, p. a. (2018). ten steps to complex learning: a systematic approach to four-component instructional design . new york, ny: routledge. velada, r., caetano, a., michel, j. w., lyons, b. d., & kavanagh, m. j. (2007). the effects of training design, individual characteristics and work environment on transfer of training. international journal of training & development, 11(4), 282-294. doi:10.1111/j.1468-2419.2007.00286.x wexley k. n., & latham g. p (1991). developing and training human resources in organizations. new york: harper-collins. wopereis, i., frerejean, j., & brand-gruwel, s. (2015). information problem solving instruction in higher education: a case study on instructional design. communications in computer and information science, 552, 293-302. doi:10.1007/978-3-319-28197-1_30 wopereis, i., frerejean, j., & brand-gruwel, s. (2016). teacher perspectives on whole-task information literacy instruction. communications in computer and information science, 676, 678-687). doi: 10.1007/978-3-319-52162-6_66 frontline learning research vol.6 no. 1 (2018) 19-30 issn 2295-3159 corresponding author: nathalie roland, psychological sciences research institute, université catholique de louvain, place du cardinal mercier 10, 1348 louvain-la-neuve, belgium. email: nathalie.roland@uclouvain.be doi : 10.14786/flr.v6i1.327 field-identification iat predicts students’ academic persistence over and above theory of planned behavior constructs nathalie rolanda, adrien mieropa, mariane frenaya, olivier corneillea auniversité catholique de louvain, belgium article received 8 september 2017 / article revised 3 november / accepted 20 march / available online april 30 2018 abstract ajzen and dasgupta (2015) recently invited complementing theory of planned behavior (tpb) measures with measures borrowed from implicit cognition research. in this study, we examined for the first time such combination, and we did so to predict academic persistence. specifically, 169 first-year college students answered a tpb questionnaire and completed a field-identification implicit association test (iat). the iat measure largely predicted academic persistence 6 months later over and above tpb constructs, including behavioral intention. we discuss interpretations of this finding and its relevance to educational research. keywords: theory of planned behavior; implicit association test; academic persistence; college students; field-identification mailto:nathalie.roland@uclouvain.be roland et al 2 | f l r 1. introduction dropout in first year at university affects about 25% of students, and comes with major social, organizational and economic costs (finnie & qiu, 2008; national center for education statistics, 2016; organization for economic cooperation and development, 2013; schmitz & frenay, 2013). over the last 30 years, education research has significantly advanced our understanding of the drivers of students’ persistence and drop out. historically, researchers have focused on the socio-demographic characteristics of students (e.g., ethnicity, parent income, parent’s third-level education, sex, educational background) to understand academic persistence (otero, rivas, & rivera, 2007; pascarella & terenzini, 2005; ratelle, larose, guay, & senécal, 2005; vermandele, dupriez, maroy, & van campenhoudt, 2012). in parallel, researchers started exploring the role of motivational variables (e.g., expectancy-value, intention, self-efficacy, control) and educational variables (e.g., institutional experiences, academic and social integration, social pressure) (braxton, hirschy, & mcclendon, 2004; cabrera, castaneda, nora, & hengstler, 1992; eccles & wigfield, 2002; nora, cabrera, hagdorn, & pascallera, 1996; pritchard & wilson, 2003; schmitz, frenay, neuville, boudrenghien, wertz, noël, & eccles, 2010; tinto, 2006). in this context, the theory of planned behavior (or tpb; fishbein & ajzen, 2010) was recently considered, with the aim of covering under a single theoretical umbrella the most widely studied determinants of academic persistence (davis, ajzen, saunders, & williams, 2002; houme, 2010; roland, frenay, & boudrenghien, 2016a). in tpb research, people’s behavior is thought be best predicted by their intention to perform this behavior. in turn, people’s intentions are determined by their attitudes, perceived norms and perceived control. those are all ultimately based on their beliefs that the behavior is likely to be under their control and to serve their best interests. this rational behavioral model relies on direct (i.e., self-reported) measures of tpb constructs, such as self-reported intentions (fishbein & ajzen, 2010). the tpb provided a very integrative understanding of academic persistence (davis et al., 2002; houme, 2010; roland et al., 2016a). recently, however, prominent tpb and implicit cognition researchers have called for studies that include both direct and indirect measures for predicting behavior. for instance, ajzen and dasgupta (2015) noted: “complementing the reasoned action approach, a great deal of research in recent years has focused on implicit cognitions and their effects on behavior. the general theorizing behind this line of work is the proposition that dormant beliefs, attitudes, intentions, and other constructs of this kind can be activated while still remaining below conscious awareness, and that these implicit reactions can have observable effects on judgments and actions” (ajzen & dasgupta, 2015, p. 136) in the present study, we examined how the latter combination between explicit and implicit cognition measures may help to better predict academic persistence. we did so by collecting both tpb constructs measures and a field-identification implicit association test (iat) measure. the iat has been extensively used in implicit cognition research (hofmann, gschwendner, nosek, & schmitt, 2005; rothermund & wentura, 2004). it makes it possible to assess spontaneous associations between one’s self-concept and semantic attributes (greenwald et al., 2002). the use of an iat has been shown to improve the prediction of behaviors in many areas (e.g., mental and physical health, employment, job performance, stereotypes), but to our knowledge it has never been used in the context of academic persistence. and, perhaps even more important, to our knowledge, the iat and the tpb have never been jointly considered in predicting behaviors. hence, the innovation of the current research is twofold: examining how an iat measure contributes to predicting academic persistence, and examining the degree to which it may do so over and above more analytical (and self-reported) tpb measures. in doing so, this research also contributes to strengthening ties between two areas of psychology that are rarely related to each other, namely educational psychology and social psychology (both in its explicit and implicit cognition dimensions). at first sight, everything opposes the tpb and the iat approaches, since the first involves detailed and deliberate measures that address explicit cognition processes whereas the second involves holistic and spontaneous measures thought to tackle implicit cognition processes. the use of both direct and indirect measures can, however, help better predict a variety of behaviors, such as intergroup behavior (e.g., roland et al 3 | f l r greenwald, mcghee, & schwartz, 1998; ottaway, hayden, & oakes, 2001), job performance (e.g., srivastava & banaji, 2011), and alcohol and drug use (e.g., chassin, presson, sherman, seo, & macy, 2008; wiers, van woerden, smulders, & de jong, 2002). furthermore, social behaviors characterized by relatively high levels of involvement, such as political voting (e.g., arcuri, catselli, galdi, zogmaister, & amadori, 2008), participation in rallies (e.g., zerhouni, rougier, & muller, 2016), or even suicidal behavior (e.g., nock et al., 2010) are often best predicted when using both direct and indirect measures. we reasoned that the predictive contribution of indirect measures could also be observed in the case of academic persistence. this is consistent with evidence suggesting that for some students vocational decisions are based on a deliberate decision-making style, while for other students the decision is more intuitive (arroba, 1977; gati, landman, davidovitch, asulin-peretz, & gadassi, 2010). in fact, gati and colleagues (2010) proposed that, for the same person, the vocational decision might be based on different decision-making styles. on the basis of these elements, it seemed important to examine the role played by more spontaneous and intuitive process in academic persistence. a field-identification iat seems particularly suited in this regard. 2. method 2.1. participants 169 first-year belgian college students in psychology agreed to participate in our research and were assured of the confidentiality of their responses. 86.4 % of the participants were female and 13.6 % were male. this ratio is representative of students in psychology at the hosting university. the mean age of the participants was 19.72 years (sd = 1.02). all participants were french speakers. among these participants, 34.3 % dropped out whereas 66.7% persisted in psychology studies. note that no observation was excluded from our sample. this sample size provided adequate statistical power (1 – β = .8) with a type i error of .05 to detect an effect size as small as cohen’s f² = .046 (which consists of a small-to-medium effect size that is suited for multiple regressions design according to cohen (1988)). 2.2. procedure in april 2016, students enrolled in first year psychology -since september 2015were asked to complete an iat and a tpb questionnaire. completion of both the iat and the tpb questionnaire was confidential and participants collaborated on a voluntary basis. the study was carried out in accordance with the ethical standards of our institution. those who agreed to participate in this research took the iat in a computer room and completed a tpb questionnaire directly afterwards. the total duration of task completion was 15 minutes on average. in september 2016, with permission from both the academic authorities and the participants, we obtained information about participants’ registration for the next year in their academic program (i.e., persistence measure). 2.3. measurement implicit association test. the iat (greenwald et al., 1998) is a computer test based on participants’ response times in classifying stimuli. it measures spontaneous associations between concepts, in this case selfconcept and psychology. more specifically, the iat consisted of 7 consecutive blocks in which participants had to quickly categorize words appearing on the screen along two dimensions by using two keys (“e” and “i”). the words used to represent the self-concept came from the literature (i.e., nock et al., 2010; nosek, banaji, & greenwald, 2002) and the words used to represent psychology were selected on the basis of a pilot study (see table 1) (bellezza, greenwald, & banaji, 1986). the pilot study was conducted on a separate sample roland et al 4 | f l r of 159 participants who responded to a solicitation on a popular social network. respondents came from different backgrounds (15% were professionals in various fields of psychology). the pilot study allowed selecting seven words highly associated with psychology (m = 35.63, sd = 11.76) and seven words highly associated with other professional and curricula domains (m = -38.46, sd = 12.75) (f(1, 15) = 2762.6, p < .001, η² = .95). table 1 words used in the implicit association test in french (english) self-related non-self-related psychology-related non-psychology-related mon (mine) je (i) mes (mines) ma (mine) mienne (mine) leur (their) il (he) ses (his) sa (her) eux (them) burnout (burnout) dépression (depression) phobie (phobia) névrose (nevrosis) emotion (emotion) freud (freud) thérapeute (therapist) archéologie (archeology) comptabilité (accounting) sol (ground) eprouvette (test tube) pesticide (pesticide) bâtiment (building) agricole (agricultural) in the first block, students had to categorize five words related to the self (e.g., “me”, “mine”) and five words related to others (e.g., “he”, “his”) by using the “e” key of the keyboard if it was self-related, and by using the “i” key if it was related to others. participants were instructed to be as quick as possible while trying to make as few mistakes as possible. each word was presented twice, resulting in 20 trials. in the second block, students had to categorize seven words related to psychology (e.g., “emotion”, “therapist”) with the “e” key and seven words related to other disciplines (e.g., “accounting”, “archeology”) with the “i” key in 20 trials (each word was presented at least once). in the third block and fourth blocks, both selfand psychologyrelated words were to be responded to with the “e” key and the non-self and non-psychology related words with the “i” key. these third and fourth “congruent” blocks consisted in 20 and in 40 trials, respectively. in a fifth block, students were asked to reverse the key mapping for the words related to psychology (by using the “i” key) or not psychology (by using the “e” key) in 20 trials. in the sixth and seventh “incongruent blocks”, students had to categorize words using the “e” key both for self-related and non-psychology related words and the “i” key” for the no-self related and the psychology-related words, respectively for 20 and 40 trials. students were randomly assigned to a condition in which congruent blocks preceded the incongruent blocks or viceversa (reversing blocks 2, 3 and 4 with blocks 5, 6 and 7). as a result, depending on the experimental block, students had to use the same key or not to categorize the words as related to the self and to psychology. the rationale is that if they associate themselves strongly with psychology, categorizing self-related words along with psychology-related words should be facilitated. likewise, categorizing self-related words with nonpsychology-related words using the same key should interfere (greenwald et al., 1998 ; hofmann et al., 2005). the relative strength between “self” and “psychology” was computed for each participant by calculating a d score [-2 ; +2] (nosek, bar-anan, sriram, axt, & greenwald, 2014). according to nosek and colleagues (2014), response times exceeding 10.000 ms and/or students with more than 10% of their response time below 300 ms had to be excluded from the sample. indeed, nosek and colleagues (2014) found that including very fast responses (<300 ms) or on the contrary slow responses (>10.000 ms) disrupted roland et al 5 | f l r psychometric properties enough to warrant excluding them. based on these rules, none of the participants of this study had to be excluded from the sample. to calculate the d score, the average response time of the congruent blocks was subtracted from the average response time of the incongruent blocks and this score was then divided by the standard deviation of response times per participant. the d score thus represents the level of spontaneous identification of the self with psychology, a positive d score meaning a stronger association between “self” and “psychology” and a negative d score meaning a stronger association between “self” and “not psychology”. persistence. to measure persistence, we used the student’s registration in the same field of studies in the next academic year (nora et al., 1996; robbins et al., 2004). students who continued in the same field after their first academic year were coded “1”; students who did not continue in the same field were coded “0”. to measure the constructs of the theory of planned behavior, we used scales adapted from fishbein and ajzen (2010). most items on the self-report questionnaire were rated on 5-point likert-type scales (generally with 1 = strongly disagree and 5 = strongly agree). the exceptions are presented below. factorial analyses were performed and revealed one factor by scale, except for perceived behavioral control where three factors were found. intention. students’ intention to persist was assessed using three items (α = .88) (“i intend to stay registered in psychology studies next year”). attitude. the evaluation of attitude was obtained by means of a set of evaluative semantic differential scales. the statement “staying enrolled in psychology studies next year will be…” was rated on five bipolar adjective scales (pleasant-unpleasant; positive-negative; useful-useless; good-bad; important-not important) taken from osgood, suci and tannenbaum (1957) (α = .89). injunctive norms. four items were used to measure injunctive norms (α = .71). one item was “my relatives believe that i should not stay enrolled in psychology studies next year”. descriptive norms1. six items were used to assess descriptive norms concerning their father and mother (α = .77). one item was “during her studies, my mother remained enrolled in the studies she had initiated”. perceived behavioral control. ten items were created for assessing perceived behavioral control, following the recommendations of fishbein and ajzen (2010). factorial analysis revealed three factors. the first factor was composed of four self-efficacy items (α = .85) (“i am sure that i will be able to stay enrolled in my psychology studies next year”). the second factor consisted of three items about the control a person has over the decision to drop out (α = .86) (“i will decide to enroll in new studies next year”). finally, the third factor was about control a person has over the decision to persist and was made up of two items (r = .75) (“i’ll decide to stay enrolled in psychology next year”). 3. results 3.1. preliminary analyses correlational analyses and summary statistics of each study’s variables were performed (table 2). the scores corresponding to skewness and kurtosis were found to be within the normal values (hae-young, 2013). correlational analyses showed outcomes consistent with relationships postulated in the tpb (for more details on tpb applied to academic persistence, see houme, 2010; roland et al., 2016a). it is interesting to note that the iat was significantly correlated with persistence, and with persistence only. performance on the iat was compared, using a t test for independent samples, between students who 1 this scale, although part of the tpb, was not taken into account in all the analysis since only students with at least one graduate parent could answer this question. this therefore excluded many of the participants. roland et al 6 | f l r persisted and students who gave up. results revealed that students who persisted initially had a stronger implicit association between self and psychology (m = .82, sd = .37) than those who eventually decided to drop out six months later (m = .47, sd = .40) (t(167) = 5.76 ; p < .001). table 2 correlations between the tpb constructs and the iat m sd skew. kurt. 1. 2. 3. 4. 5. 6. 7. 8. 9. 1. persistence .65 .48 -.67 -1.57 1 .41** .19* .07 .19** .01 .20** .02 .11 2. iat .70 .42 -.32 1.41 1 -.02 -.03 -.05 -.01 .01 -.13 -.08 3. intention 4.68 .71 -2.52 3.4 1 .63** .36** -.12 .50** .17* -.09 4. attitude 4.56 .49 -1.22 1.01 1 .37** -.08 .43** .18* -.07 5.injunctive norms 4.24 .68 -.91 .23 1 -.02 .33** .13 -.04 6.descriptive norms 4.04 1.25 -.61 -.51 1 .52 .32 .05 7.self-efficacy beliefs 4.28 .65 -.79 -.17 1 .20** .11 8.control on persistence 4.85 .43 -2.48 4.01 1 .08 9.control on dropout 3.21 1.49 -.30 -1.35 1 ** = p < .01; * = p < .05 3.2. main analyses of critical interest to the present research is whether the field-identification iat would predict persistence over and above tpb measures. to this end, we ran a logistic regression analysis that regressed persistence on tpb and iat measures (see table 3). logistic regression is the appropriate regression analysis to conduct when the dependent variable is dichotomous. this analysis showed that self-reported intention and the iat measure predicted persistence in this joint model2. of importance, intention was less predictive of persistence than the iat measure: the odds of persisting were much higher when the iat score was high (exp b = 18.25) than when intention was high (exp b = 2.28). 2 that other constructs of tpb do not predict persistence in this model is not surprising since they are supposed to impact persistence through intention. preliminary analyses showed that when intention is not introduced into the regression, self-efficacy predicts persistence, which is consistent with research (barry & finney, 2009; vuong, brown-welry, & tracs, 2010; wright, jenkins-guarnieri, & merdock, 2012). table 1 also shows simple correlations between those constructs, intentions and persistence, which are fully consistent with tpb. roland et al 7 | f l r table 3 logistic regression analysis predicting academic persistence b (se) exp b constant -5.18 (2.74) .01 attitude -.88 (.52) .42 injunctive norms .56 (.31) 1.74 self-efficacy beliefs .48 (.35) 1.62 control on persistence .05 (.48) 1.05 control on dropout -.15 (.14) .86 intention .83 (.38) 2.28* iat 2.90 (.59) 18.25*** r² (cox & snell) = .26 ; r² (nagelkerke) = .36 ; model 2(df) = 51.07(7)*** * p < .05 ; ***p < .001 4. discussion the general goal of this research was to follow a recent invitation to examine the joint contribution of tpb and implicit cognition measures for predicting social behaviors (ajzen & dasgupta, 2015). we did so by combining a tpb measure and a field-identification iat measure to predict academic persistence. in doing so, we combined for the first time tpb and iat measures to predict a behavior. as an additional asset, we also introduced for the first time a measure borrowed from implicit cognition in academic persistence research, the latter of which has extensively relied on self-reported measures, so far (e.g., cabrera et al., 1992; eccles & wigfield, 2002; pascarella & terenzini, 2005; roland et al., 2016a; schmitz et al., 2010; tinto, 2006). the measure that best illustrates this deliberate approach is students’ intention to persist (braxton, bray, & berger, 2000; dadeppo, 2009; schmitz & frenay, 2013). in contrast, the use of more indirect measures has never been examined in research on academic persistence, although it was examined in research on vocational choice (arroba, 1977; gati et al., 2010). we found that the field-identification iat measure (i) strongly predicts persistence (i) strongly predicts it over and above comprehensive tpb measures, and (iii) predicts, by a large margin, persistence more strongly than tpb measures do, when collected six months ahead of students’ actual persistence decisions. these findings may be interpreted in different ways. one interpretation is that unconscious determinants of people’s behavior operate independently of people’s conscious beliefs and deliberate intentions. at first sight, this interpretation is consistent with our finding that the iat measure was uncorrelated with tpb measures and that the association between the iat measure and persistence was not mediated by intentions (as iat and intentions were not associated with each other). we do not subscribe to this latter interpretation. as a matter of fact, recent research shows that constructs measured by the iat can be formed through fully deliberate learning processes (e.g. gast & de houwer, 2012; kurdi & banaji, 2017; van dessel, de houwer, gast, & smith, 2015) and that people are able to consciously introspect their iat score (hahn, judd, hirsch, & blair, 2014). more generally, that indirect measures reflect the operation of either independent roland et al 8 | f l r learning pathways or behavioral expression pathways has been questioned lately (for a recent discussion, see corneille & stahl, 2018) instead, the current findings may suggest that students’ intentions are less stable than their identification to the field. as time goes by, students may experience situations that lead them to revise their beliefs and so update their attitudes, perceived norms, sense of control, and ultimately their intention to persist. this is in line with research showing that temporal distance weakens the intention-to-behavior association (mceachan, conner, taylor, & lawton, 2011; sheeran, orbell, & trafimow, 1999). this second interpretation suggests that the predictive advantage of the iat resides in the higher stability of the construct it tackled (i.e., identification to the field). for instance, in the first months of their studies, many students are disappointed that they have to attend very general courses (e.g., physiology and statistics) instead of more specialized psychology courses (neuville, frenay, & bourgeois, 2007; roland, frenay, & boudrenghien, 2016b). those students may thus revise their intention to persist in a program that does not fulfill their expectations, while still feeling attracted to psychology in general (roland et al., 2016b). field-identification may be less sensitive to the latter disappointment and motivate students to persist in their studies. if this interpretation is correct, one may speculate that for students more advanced in their curriculum, the predictive gap between intention and iat measures decreases (as the courses become more specialized from year to year). one may also predict that, within the first year, the iat-intention predictive gap decreases over the year, and also that the iat measure is more stable than the intention measure. finally, one may predict this iat-intention predictive gap to be smaller in academic fields where students’ expectations about curricula are possibly more realistic (e.g., mathematics, engineering). the hypothesis that the iat measure was more stable than the direct measure of intention should, however, be confirmed in a longitudinal study. in any case, when it comes to the question of which measure may be favored for diagnostic purposes, the current analysis clearly supports the field-identification iat, at least for the student population considered here, one that suffers from massive dropout rates in the first year (romainville & michaut, 2012). this conclusion is in line with research showing that the link between intention and persistence is highly variable and typically weak (bers & smith, 1991; cabrera et al., 1992; pascarella, duby, & iverson, 1983; sandler, 2000). note that an interesting question for future research is whether a deliberate field-identification measure may not serve the same purpose, for instance using a self-reported pictorial self-categorization measure (schubert & otten, 2002). answering this question would be of paramount interest for both implicit cognition theorization and academic guidance. one possibility is that iat and deliberate identification measures complement each other. alternatively, a field-identification iat may outperform, as this measure is likely to be less contaminated by social demands and introspective effects. finally, another interesting finding of the current research is that a precise and deliberate behavior can be better predicted by an iat measure than by several tpb questions. two comments have to be made here. first, this finding may seem inconsistent with the compatibility principle inherent in the tpb, which states that a specific behavior is best predicted by a diversity of precise questions (at least when they are separated by a significant time delay). again, however, it should be noted that tpb measures were probably collected here too early to secure their maximum predictive value. second, the fact that the iat measure predicted a behavior as deliberate as persisting in an academic curriculum may seem problematic for research suggesting doubledissociations in measure and behavior types such that deliberate and conscious behaviors would be best predicted by self-reports whereas more automatic behaviors would be best predicted by indirect associative measures as the iat (e.g., friese, hofmann, & schmitt, 2009; hofmann et al., 2005). however, numerous studies, along with the present one, have reported findings inconsistent with the dissociative view. for instance, iat measures have been shown to successfully predict behaviors as deliberate as political votes (arcuri et al., 2008) or suicide attempts (nock et al., 2010). the present results should be interpreted in the light of several limitations, which call for future research. first, this study was conducted only with psychology students. as discussed above, it is important to replicate this study with students from other faculties, and also from other universities. to achieve this, it would be necessary to conduct pilot studies to identify stimuli relevant to each field of study (greenwald et al., 1998; nosek, banaji, & greenwald, 2010). second, the stimuli used in the iat tested here focused mostly roland et al 9 | f l r on clinical psychology. these stimuli stem from the pilot study we carried out and may therefore be considered as valid. although the resulting iat predicted overall academic persistence in psychology, it is possible that some students, who were attracted to other fields of psychology (e.g., work psychology), felt less identified with these stimuli and that for these students the iat was less predictive of their specific future persistence. despite its limitations, however, this study provides a first test of the role of implicit cognition measures in predicting academic persistence, and more generally in complementing tpb measures. neither of these objectives was empirically addressed so far. indeed, direct measures are by very far the dominant norm in educational research. this more generally points to the interest of bringing different fields of research together; in the present case, educational and social psychology. keypoints academic persistence is predicted by both direct and indirect measures. a field-identification implicit association test strongly predicts students’ academic persistence six months ahead of decision, over and above theory of planned behavior measures. the interpretations of this finding and its relevance to educational research are discussed. references ajzen, i., & dasgupta, n. (2015). explicit and implicit beliefs, attitudes, and intentions: the role of conscious and unconscious processes in human behavior. in p. haggard, & b. eitam (eds.), the sense of agency (pp. 115-144). new york: oxford university press. arcuri, l., castelli, l., galdi, s., zogmaister, c., & amadori, a. (2008). predicting the vote: implicit attitudes as predictors of the future behavior of decided and undecided voters. political psychology, 29(3), 369-387. doi: 10.1111/j.1467-9221.2008.00635.x arroba, t. (1977). styles of decision making and their use: an empirical study. british journal of guidance and counselling, 5, 149-158. doi: 10.1080/03069887708258110 barry, c. l., & finney, s. j. (2009). can we feel confident in how we measure college confidence? a psychometric investigation of the college self-efficacy inventory. measurement and evaluation in counseling and development, 43(3), 197-222. doi: 10.1177/0748175609344095 bellezza, f. s., greenwald, a. g., & banaji, m. r. (1986). words high and low in pleasantness as rated by male and female college students. behavior research methods, instruments, and computers, 18(3), 299-303. doi: 10.3758/bf03204403 bers, t. h., & smith k. e. (1991). persistence of community college students: the influence of student intent and academic and social integration. research in higher education, 32(5), 529-556. doi: 10.1007/bf00992627 braxton, j.-m., bray, n. j., & berger, j.-b. (2000). faculty teaching skills and their influence on the college student departure process. journal of college student development, 41(2), 215-227. braxton, j.-m., hirschy, a. s., & mcclendon, s. a. (2004). understanding and reducing college student departure. san francisco: jossey-bass. cabrera, a. f., castaneda, m. b., nora, a., & hengstler, d. (1992). the convergence between two theories of college persistence. journal of higher education, 63(2), 143-164. doi: 10.2307/1982157 chassin, l., presson, c., seo, d., sherman, s. j., macy, j., wirth, r. j., & curran, p. (2008). multiple trajectories of cigarette smoking and the intergenerational transmission of smoking: a multigenerational, longitudinal study of a midwestern community sample. health psychology, 27, 819-828. doi: 10.1037/0278-6133.27.6.819 cohen, j. (1988). statistical power analysis for the behavioral sciences (2nd ed). hillsdale, nj: erlbaum. roland et al 10 | f l r corneille, o., & stahl, c. (in press). associative attitude learning. getting closer to evidence and how it relates to attitude models. personality and social psychology review. doi: 10.1177/1088868318763261 dadeppo, l. (2009). integration factors related to the academic success and intent to persist of college students with learning disabilities. learning disabilities research & practice, 24(3), 122-131. doi: 10.1111/j.1540-5826.2009.00286.x davis, l. e., ajzen, i., saunders, j., & williams, t. (2002). the decision of african american students to complete high school: an application of the theory of planned behavior. journal of educational psychology, 94(4), 810-819. doi: 10.1037//0022-0663.94.4.810 eccles, j. s., & wigfield, a. (2002). motivational beliefs, values, and goals. annual review of psychology, 53, 109-132. doi:10.1146/annurev.psych.53.100901.135153 finnie, r., & qiu, h. (2008). the patterns of persistence in post-secondary education in canada: evidence from the yits-b dataset. toronto, on: educational policy institute. fishbein, m., & ajzen, i. (2010). predicting and changing behavior: the reasoned action approach. new york: psychology press (taylor & francis). friese, m., hofmann, w., & schmitt, m. (2008). when and why do implicit measures predict behaviour ? empirical evidence for the moderating role of opportunity, motivation, and process reliance. european journal of social psychology, 19(1), 285-338. doi: 10.1080/10463280802556958 gast, a., & de houwer, j. (2012). evaluative conditioning without directly experienced pairings of the conditioned and the unconditioned stimuli. the quarterly journal of experimental psychology, 65(9), 1657-1674. doi: 10.1080/17470218.2012.665061 gati, i., landman, s., davidovitch, s., asulin-peretz, l., & gadassi, r. (2010). from career decisionmaking styles to career decision-making profiles : a multidimensional approach. journal of vocational behavior, 76(2), 277-291. doi: 10.1016/j.jvb.2009.11.001 greenwald, a. g., banaji, m. r., rudman, l. a., farnham, s. d., nosek, b. a., & mellott, d. s. (2002). a unified theory of implicit attitudes, stereotypes, self-esteem, and self-concept. psychological review, 109(1), 3-25. doi: 10.1037//0033-295x.109.1.3 greenwald, a. g., mcghee, d. e., & schwartz, j. l. k. (1998). measuring individual differences in implicit cognition : the implicit association test. journal of personality and social psychology, 74(6), 1464-1480. doi: 10.1037/0022-3514.74.6.1464 hae-young, k. (2013). statistical notes for clinical researchers: assessing normal distribution using skewness and kurtosis. restor dent endod, 38(1), 52-54. doi: 10.5395/rde.2013.38.1.52 hahn, a., judd, c. m., hirsh, h. k., & blair, i. v. (2014). awareness of implicit attitudes. journal of experimental psychology: general, 143(3), 1369-1392. doi: 10.1037/a0035028 hofmann, w., gschwendner, t., nosek, b. a., & schmitt, m. (2005). what moderates implicit-explicit consistency ? european review of social psychology, 16(1), 335-390. doi: 10.1080/10463280500443228 houme, k. p. (2009). application de la théorie du comportement planifié pour prédire la persévérance des étudiants en sciences naturelles de l’université de lomé (togo) (doctoral dissertation). retrieved from http://www.theses.ulaval.ca/2009/26821/26821.pdf mceachan, r. r. c., conner, m., taylor, n. j., & lawton, r. j. (2011). prospective prediction of healthrelated behaviors with the theory of planned behavior: a meta-analysis. health psychology review, 5(2), 97-144. doi:10.1080/17437199.2010.521684 kurdi, b., & banaji, m. r. (2017). repeated evaluative pairings and evaluative statements: how effectively do they shift implicit attitudes? journal of experimental psychology: general, 146(2), 194–213. http://doi.org/10.1037/xge0000239 national center for education statistics [nces] (2016). persistence and attainment of 2011-12 firsttime postsecondary students after 3 years (research report n°2016-401). washington: department of education, office of educational research and improvement. neuville, s., frenay, m., & bourgeois, e. (2007). task value, self-efficacy and goal orientations: impact on self-regulated learning. psychologica belgica, 47(1-2), 95-117. doi : 10.5334/pb-47-1-95 https://dx.doi.org/10.5395%2frde.2013.38.1.52 http://doi.org/10.1080/10463280500443228 http://doi.org/10.1080/10463280500443228 http://www.theses.ulaval.ca/2009/26821/26821.pdf roland et al 11 | f l r nock, m. k., park, j. m., finn, c. t., deliberto, t. l., dour, h. j., & banaji, m. r. (2010). measuring the “suicidal mind:” implicit cognition predicts suicidal behavior. psychological science, 21(4), 511517. doi: 10.1177/0956797610364762 nora, a., cabrera, a., hagdorn, l. s., & pascarella, e. (1996). differential impacts of academic and social experiences on college-related behavior outcomes across different ethnic and gender groups at four-year institutions. research in higher education, 37(4), 427-451. doi: 10.1007/bf01730109 nosek, b. a., banaji, m. r., & greenwald, a. g. (2002). math = male, me = female, therefore math = me. journal of personality and social psychology, 83(1), 44-59. doi: 10.1037//0022-3514.83.1.44 nosek, b. a., banaji, m. r., & greenwald, a. g. (2010). project implicit. retrieved from https://implicit.harvard.edu/implicit/ nosek, b. a., bar-anan, y., sriram, n., axt, j., & greenwald, a. g. (2014). understanding and using the brief implicit association test: recommended scoring procedures. plos one 9(12), 1-31. doi:10.1371/journal.pone.0110938 organization for economic cooperation and development [oecd] (2013). regards sur l’éducation. retrieved from http://www.oecd.org/fr/education/rse-indicateurs.htm osgood, c. e., suci, g. j., & tannenbaum, p. h. (1957). the measurement of meaning. urbana: university of illinois press. otero, r., rivas, o., & rivera, r. (2007). predicting persistence of hispanic students in their first year of college. journal of hispanic higher education, 6(2), 163-173. doi:10.1177/1538192706298993 ottaway, s. a., hayden, d. c., & oakes, m. a. (2001). implicit attitudes and racism: effects of word familiarity and frequency on the implicit association test. social cognition, 19(2), 97-144. doi: 10.1521/soco.19.2.97.20706 pascarella, e., duby, p., & iverson, b. (1983). a text and reconceptualization of a theoretical model of college withdrawal in a commuter institution setting. sociology of education, 56(2), 88-100. doi: 10.2307/2112657 pascarella, e., & terenzini, p. (2005). how college affects students: a third decade of research (2nd ed.). hoboken, nj: john wiley & sons. pritchard, m. e., & wilson, g. s. (2003). using emotional and social factors to predict student success. journal of college student development, 44(1), 18-28. doi: 10.1353/csd.2003.0008 ratelle, c. f., larose, s., guay, f., & senécal, c. (2005). perceptions of parental involvement and support as predictors of college students’persistence in a science curriculum. journal of family psychology, 19(2), 286-293. doi :10.1037/0893-3200.19.2.286 robbins, s. b., lauver, k., le, h., davis, d., langley, r., & carlstrom, a. (2004). do psychosocial and study skill factors predict college outcomes? a meta-analysis. psychological bulletin, 130(2), 261288. doi:10.1037/0033-2909.130.2.261 roland, n., frenay m., & boudrenghien, g. (2016a). understanding academic persistence through the theory of planned behaviour: normative factors under investigation. journal of college student retention: research, theory & practice. doi: 10.1177/1521025116656632 roland, n., frenay, m., & boudrenghien, g. (2016b). towards a better understanding of academic persistence among freshmen: a qualitative approach. journal of education and training studies, 4(12), 175-188. doi: 10.11114/jets.v4i12.1904 romainville, m., & michaut, c. (2012), réussite, échec et abandon dans l’enseignement supérieur. bruxelles: de boeck. rothermund, k., & wentura, d. (2004). underlying processes in the implicit association test (iat): dissociating salience from associations. journal of experimental psychology: general, 133(2), 139165. doi: 10.1037/0096-3445.133.2.139 sandler, m. e. (2000). career decision-making self-efficacy, perceived stress, and an integrated model of student persistence: a structural model of finances, attitudes, behavior, and career development. research in higher education, 41(5), 537-580. doi: 10.1023/a:1007032525530 schmitz, j., & frenay, m. (2013). la persévérance en première année à l’université : rôle des expériences en classe, de l’intégration sociale et de l’ajustement émotionnel. in s. neuville, m. frenay, b. noël, https://implicit.harvard.edu/implicit/ http://www.oecd.org/fr/education/rse-indicateurs.htm https://doi.org/10.1353/csd.2003.0008 roland et al 12 | f l r & v. wertz (eds.), persévérer et réussir à l'université. louvain-la-neuve: presses universitaires de louvain. schmitz, j., frenay, m., neuville, s., boudrenghien, g., wertz, v., noël, b., & eccles, j. (2010). étude de trois facteurs clés pour comprendre la persévérance à l'université. revue française de pédagogie, 172, 43-61. doi : 10.4000/rfp.2217 schubert, t. w., & otten, s. (2002). overlap of self, ingroup, and outgroup: pictorial measures of selfcategorization. self and identity, 1(4), 353-376. doi: 10.1080/152988602760328012 sheeran, p., orbell, s., & trafimow, d. (1999). does the temporal stability of behavioural intentions moderate intention-behavior and past behavior-future behavior relations? personality and social psychology bulletin, 25(6), 724-734. doi:10.1177/0146167299025006007 srivastava, s. b., & banaji, m. r. (2011). culture, cognition, and collaborative networks in organizations. american sociological review, 76(2), 207-233. doi: 10.1177/0003122411399390 tinto, v. (2006). research and practice of student retention: what next? journal of college student retention, 8(1), 1-19. doi: 10.2190/c0c4-eft9-eg7w-pwp4 van dessel, p., de houwer, j., gast, a., & smith, c. t. (2015). instruction-based approach-avoidance effects. changing stimulus evaluation via the mere instruction to approach or avoid stimuli. experimental psychology, 62, 161-169. doi: 10.1027/1618-3169/a000282 vermandele, c., dupriez, v., maroy, c., & van campenhoudt, m. (2012). réussir à l’université : l'influence persistante du capital culturel de la famille. les cahiers de recherche en éducation et formation, n°87. vuong, m., brown-welty, s., & tracz, s. (2010). the effects of self-efficacy on academic success of first-generation college sophomore students. journal of college student development, 51(1), 50-64. doi : 10.1353/csd.0.0109 wiers, r. w., van woerden, n., smulders, f. t. y., & de jong, p. j. (2002). implicit and explicit alcohol-related cognitions in heavy and light drinkers. journal of abnormal psychology, 111(4), 648-658. doi: 10.1037/0021-843x.111.4.648 wright, s. l., jenkins-guarnieri, m. a., & murdock, j. l. (2012). career development among first-year college students: college self-efficacy, student persistence, and academic success. journal of career development, 40(4), 292-310. doi: 10.1177/0894845312455509 zerhouni, o., rougier, m., & muller, d. (2016). “who (really) is charlie?” french cities with lower implicit prejudice toward arabs demonstrated larger participation rates in charlie hebdo rallies. international review of social psychology, 29(1), 69-76. doi: 10.5334/irsp.50 http://doi.org/10.5334/irsp.50 microsoft word malmberg et al_publication.docx           frontline  learning  research  vol.4  no.  5  (2016)  62  -­‐  82   issn  2295-­‐3159       contact information: lars-erik malmberg, department of education, university of oxford, 15, norham gardens, ox2 6py, oxford, uk. email: lars-erik.malmberg@education.ox.ac.uk. doi: http://dx.doi.org/10.14786/flr.v4i5.227   within-students variability in learning experiences, and teachers' perceptions of students' task-focus lars-erik malmberga, wee h. t. lima, asko tolvanenb and jari-erik nurmib a university of oxford, united kingdom b university of jyväskylä, finland article received 16 november / revised 7 july / accepted 7 september / available online 11 january   abstract in order to advance our understanding of educational processes, we present a tutorial of intraindividual variability. an adaptive educational process is characterised by stable (less variability), and a maladaptive process is characterised by instable (more variability) learning experiences from one learning situation to the next. we outline step by step how we specify a multilevel structural equation model of state, trait and individual differences in intraindividual variability constructs, which can be appropriately fitted to intraindividual data (e.g., time-points nested in persons, intensive longitudinal data). in total 285 primary school students’ (years 5 and 6) completed the learning experience questionnaire using handheld computers, on average 13.6 learning episodes during one week (sd = 4.6; range = 5-29; nepisodes = 3,433). we defined mean squared successive differences (mssd) for each manifest indicator of task difficulty, competence evaluation and intrinsic motivation. we also demonstrate how to specify multivariate models for investigating convergent validity of the variability constructs. overall, our study provides support for intraindividual variability as a construct in its own right, which has the potential to provide novel insight into students’ learning processes. keywords: intraindividual variability; multilevel structural equation model (msem); learning experience; ecological momentary assessment malmberg  et  al         | f l r 63     1. introduction there is a growing interest in the study of students’ learning processes using diary and real-time data (schmitz, 2006). these micro-longitudinal studies expand our knowledge about learning processes beyond what we can learn from single time-point cross-sectional studies, in at least three ways. first, there is considerable variation in students’ learning experiences, e.g., their engagement, beliefs, motivation, emotions, and performance from one situation to another (i.e., intraindividual variation), more so than there is variation between students (i.e., interpersonal variation; schmitz & skinner, 1993). second, situationspecific learning experiences vary as a function of contextual features, such as perceived autonomy support (tsai, kunter, lüdtke, & trautwein, 2008), and extrinsic motivation (malmberg, pakarinen, vasalampi, & nurmi, 2015). this means that situation specific opportunities and constraints, such as provision of support and levels of expectation, form an integral part of students’ learning experiences. third, students’ individual characteristics can moderate the relationship between experiences. compared with relatively lower achievers, higher achievers had more stable control beliefs and perceived task ease from one situation to the next (musher-eizenman, nesselroade, & schmitz, 2002), and exerted more effort when confronted with difficult tasks (malmberg, walls, martin, little, & lim, 2013). what we know less about is the intraindividual variability in students’ learning experiences from one situation to the next. whilst intraindividual variation captures the differences between individuals’ experiences above or below their own average experience (i.e., an “individual standard deviation” of own “ups” and “downs”), intraindividual variability, inconsistency, or instability refers to the magnitude of short-term fluctuations in the order of the ups and downs from one time-point to the next (e.g., jahng, wood, & trull, 2008; kernis, grannemann, & barclay, 1989). this magnitude of intraindividual variability is larger when the shifts between highs and lows are more abrupt, occur more often, and the swings go from one extreme to the other. in the present study we go beyond previous real-time studies of students’ learning experiences in two ways. first, we propose a methodology for specifying a within-person variability construct alongside state and trait constructs, using state-of-the-art multilevel structural equation models (msem). the msem allows us to model latent constructs net of measurement error at two (or more) levels of data. second, we include teacher perceived student task-focus as an indicator of convergent validity of students’ intraindividual variability. accumulated research shows that students, who in the eyes of their teacher are generally task-focused, are: intrinsically motivated, deploy task-focused rather than task-avoidant behavioural engagement, exert effort, seek help when they need it, and persist when they encounter difficulties (eccles, wigfield, & schiefele, 1998; nurmi, hirvonen & aunola, 2008, zimmerman, 2000). it would be important to know whether students who teachers regard as task-focused, are also more stable in their learning experiences, i.e., less variability in students’ perceptions of task difficulty, competence beliefs and intrinsic motivation from one learning situation to the next. to this end we provide a brief overview of intraindividual research in education, task-focused learning, a didactical example of the mean squared sequential difference (mssd) index of intraindividual variation, and an msem specification. 1.1. intraindividual resesarch to education there appears to be a surge in intraindividual research in education. since the seminal diary studies by schmitz and skinner (1993) and musher-eizenman et al. (2002), an up-swing in the number of publications has been seen, for example schmitz and wiese (2006), and tsai et al., (2008). recent studies have used experience sampling of students’ academic emotions (goetz, frenzel, stoeger & hall, 2010), coping with boredom (nett, goetz & hall, 2011) and metacognitive strategies (nett, goetz, hall & frenzel, 2012); ecological momentary assessment studies of effort exertion, competence beliefs and task difficulty (malmberg, walls et al., 2013); and contextual activity sampling of university students’ challenge, competence and emotions (inkinen et al., 2013). data in these studies were collected at multiple time-points in their natural settings, as close in time as possible to events, thus reducing retrospection bias (wilhelm, perrez, & pawlik, 2012). the importance of the intraindividual perspective on learning experiences is threefold. these studies pave the way for understanding, first, learning processes as they occur in real-time; second, individual differences in such learning processes, and third, how teachers might differentially malmberg  et  al         | f l r 64     support individual students. taken together, an intraindividual approach to learning can help us understand both learning processes and the ways in which teachers can support these (schmitz, 2006). 1.2. intraindividual variation and variability in the research fields of personality and psychiatry, affect instability is characteristic of personality disorders (jahng et al., 2008; trull et al., 2008), with particular focus on negative mood (eid & langeheine, 2003), affect (eid & diener, 1999), mood and job satisfaction (ilies & judge, 2002), affect and mood instability (jahng et al., 2008), short-term fluctuations in self-esteem (kernis et al., 1989), and mood variability (mcconville & cooper, 1997). expanding into other fields, recent studies of intraindividual variability include secure attachment (la guardia, ryan, couchman, & deci, 2000), temperament (hooker, nesselroade, nesselroade, & lerner, 1987), perceived control (eizenman, nesselroade, featherman, & rowe, 1997), and coping (roesch et al., 2010). a range of techniques have been suggested for aggregating measures of within-person variability (for a review, see jahng et al., 2008): the intraindividual standard deviation (or variance), first-order autocorrelation coefficients r, and the mean square successive difference (mssd; von neumann, kent, bellinson, & hart, 1941). while the intraindividual standard deviation is intuitively appealing, it does not capture the frequency of change (larsen, 1987). the mssd calculates an aggregate that takes the sequential order of the events into account (equation 1). mssd =   ! !!!   (x! +  1 −  x!)!!!! !!!   (1), where xi + 1 is the lagged value of xi. the squared difference between xi + 1 and xi assures that the magnitude of the successive differences is captured. there are n-1 observations in the dataset (see appendix 1). in a didactic simulation shown in figure 1 we exemplify the conduct of the mean (m), standard deviation (sd), the mean square successive difference (mssd), and the autocorrelation (r), in three scenarios (panels a, b and c). for a similar simulation see jahng et al. (2008). when we observe the raw data in panel a (figure 1) we find that the m and sd are the same as in panel c, in which the data has been rank-ordered in descending order. the m and sd in panel b, in which each data-point has been multiplied by two, are the same as multiplying the m and sd of those in panel a by two. while the sd indeed captures variation, it is not sufficient for capturing the magnitude of variation. the stability over time captured by the autocorrelation r remains the same in panels a and b, demonstrating that r does not capture the magnitude of change either. the autocorrelation coefficient r is different in panel c demonstrating that the order of events matter. finally mssd differs in all three panels demonstrating that it is both sensitive to magnitude (panel b) and order of change (panel c). in the present study we use mssd for investigating intraindividual variability. in previous studies, a range of models for investigating lagged associations have been specified, including time-series and spectral analysis (larsen, 1987; ram et al., 2005), the mixed-effects location scale model (li & hedeker, 2012), generalized multilevel model (jahng et al., 2008), and mixture distribution models (eid & langeheine, 2003). however, these models do not correct for measurement error in constructs. to do so, we calculated mssd for each indicator of our latent constructs and modelled these using multilevel structural equation models (msem). although time-series typically requires longer stretches of time-points, the mssd method is suggested to be robust also for shorter time-series e.g., a number of time-points during each day (ebner-priemer, eid, kleindienst, stabenow, & trull, 2009). malmberg  et  al         | f l r 65     figure 1. three example time-series and indices of intraindividual variability (cf. jahng, et al., 2008). note: panel a represents one sample student for whom 29 situation reports were observed for intrinsic motivation (1 = low motivation, 4 = high motivation). panel b represents each numerical value in panel a multiplied by 2, so the scale now spans 2 to 8. panel c represents the raw data from panel a but now rank-ordered in descending order. m = mean, sd = standard deviation, mssd = the mean square successive difference, and r = autocorrelation. 1.3. research questions and hypotheses a) what is the structural validity of the state, trait and intraindividual variability constructs? b) what is the association between trait and intraindividual variability constructs? c) how do trait and intraindividual variability constructs of students’ learning experiences converge with teacher-reported task-focus? hypothesis 1: we expected convergence between teacher-reports of students' task-focus (nurmi et al., 2008), higher level of task-focus positively and moderately associated with trait-levels of each construct, and negatively associated with variability constructs (i.e., higher task-focus less variability, lower task-focus more variability). 2. method 2.1. sample and procedure in total, 353 students in 16 classrooms in 11 schools participated in the learning every lesson (lel) study (for details see malmberg, woolgar, & martin, 2013; malmberg, walls et al., 2013; malmberg et al., 2015), with informed parental or guardian consent. the study was carried out in two quite diverse areas in southeast england, uk. students were asked to complete the electronic learning experience questionnaire (leq) for personal digital assistant (pda) at the end of each learning episode or at least once per lesson. teachers or teaching assistants were asked to complete a brief one-page report of each student they taught. teaching arrangements differed across the classes. in half of the classrooms, one teacher reported on all his or her students; in four classrooms two teachers reported on the students; in two malmberg  et  al         | f l r 66     classrooms, there was a mix of students with one or two teacher reports; and in another two classrooms, two or three teachers reported. in order to investigate the correspondence between students’ and teachers’ views of the students, in the final study sample we included all observations for which both teacher and student reports for any given student were available. the intraclass correlation for teacher-reported task-focus was ricc = .08 between classrooms and ricc = .08 between teachers (malmberg et al., 2015). in order to not burden the models with additional hierarchical levels, teacher reports were aggregated for each student, weighted for the number of experiences with each teacher. however, for the purpose of aggregating mssd-indices of the lagged relationships between the time-points, we carried out analyses for those students who had at least five timepoints of data available (roughly the possible number of reports per day). there were 285 students who reported on 3,433 learning episodes: on average 13.6 learning episodes (sd = 4.6; range = 5-29) combined with 434 teacher reports (139 students had one teacher report, 143 had two reports and 3 had three reports). of these there were 126 boys (44.2%) and 159 girls (55.8%), 104 were in year 5 (36.5%) and 181 in year 6 (63.5%). they were 10.5 years old on average (sd = 0.64). 2.2. student-reported measures students’ learning experiences were measured using the validated leq (reliability, structural and external validity), covering sources of motivation, learning behaviour, competence evaluation and affect (malmberg, woolgar, & martin, 2013). 2.2.1. task difficulty students completed a single item measuring task difficulty: “the learning task i was doing was”, on a four-point scale (1= very easy, 4= very hard). 2.2.2. competence evaluation students responded to two items indicating competence evaluation (mα = .70; sdα = .18): “how well were you doing at this task” on a five-point scale (1 = poorly, 5 = very well), and “how much did you understand” on a four-point scale (1 = all of it, 4 = none of it; reverse-coded). 2.2.3 intrinsic motivation students were asked “why were you doing this task?” and responded to three items measuring intrinsic motivation: “i enjoyed it”, “i chose to do it”, and “i was interested in it”. when we split the data by day and learning experience, the average internal consistency was mα = .85 (sdα = .09). 2.3. teacher-reported measures teachers reported on each student’s task-focused characteristics and behaviour. 2.3.1. task-focus we used teacher-reports of each student’s task-focus in school in general. task-focus was measured with six items modified from the observer-rating scale of achievement strategies (osas; nurmi, & aunola, 1998), and the behavioural strategy rating scale ii (bsr-ii; aunola, nurmi, parrila, & onatsuarvilommi, 2000; zhang, nurmi, kiuru, lerkkanen, & aunola, 2011). teachers were asked to think about each student’s behaviour and work habits in class, and respond on five-point scales (0 = not at all, 1 = rarely, 2 = sometimes, 3 = often, 4 = very often), to what extent each of the six statements characterise the way each student typically behaves in learning situations. half of the items were positively worded (indicating taskfocus): “actively attempts to solve even difficult tasks”, “demonstrates initiative and persistence in activities and tasks”, and “tries hard to finish even difficult tasks”. the three negatively worded items (indicating taskmalmberg  et  al         | f l r 67     avoidance) were: “has a tendency to find something else to do, instead of focusing on the task at hand”, “gives up easily”, and “loses focus if a task or activity is not going well” (α = .88). we specified the construct so that higher values indicated more focus on tasks. task-focus was strongly and positively related to academic performance (malmberg et al., 2015). 2.4. analytic procedures we specified multilevel structural equation models (msem) in mplus (muthén & muthén, 2012). at the within level we specified a latent state construct ξw1 using x1 to x3 as indicators (see figure 2). at the between level, we specified a correspondence between level trait construct ξb1, equating factor loadings across the levels for metric invariance between the state and trait constructs (morin, marsh, nagengast, & scalas, 2014). we then specified a second between-level construct, which captures interindividual differences in intraindividual variability, ξb2 using k indicators. figure 2. msem of statetrait and variability constructs note: indicators are raw data of time-points (t) nested in students (i). circles above (at the between level, e.g., x1b) and below (at the within level e.g., x1w) the indicators depict latent constructs of decomposed betweenand within-level indicators respectively. there is one within-level latent construct (ξw1) and two between-level constructs (ξb1 and ξb2), with factor loadings (λ, one-headed arrows) linking constructs to level-specific indicators. variances of latent constructs are indicated in double headed arrows (ψ). residuals of indicators are also depicted with double headed arrows (ε), at the within-level measurement error. the mean-structure (triangle with 1 inside) is estimated at the between-level (i.e., cluster-intercepts, τ). in the dataset we created lagged variables (xkt+1) of each indicator (xkt) for each student. this gave 285 additional lines of data, one for each participant in our data-matrix, giving a total of nti = 3,718 lines of malmberg  et  al         | f l r 68     data (see appendix 1). we then, in mplus, defined intraindividual squared deviations (xkt+1 xkt)2 which we used as indicators (see appendix 2). the scalar of the mssd equation, ! !!! , was not necessary to apply as there are n-1 number of successive differences for each participant. calculating the average of the successive differences is to divide the sum of the squared successive differences by n-1. we specified msems, presented in figures 3-5, for each construct using one (difficulty), two (competence), and three indicators (intrinsic motivation) for each latent construct separately (models 1-3). we then illustrated how to specify three multivariate models, presented in figure 6 for investigating convergence between trait and variability constructs, and between variability and task-focus (models 4-6). we inspected indices of convergence (association of higher magnitude where expected) and divergence (lack of association where expected; campbell & fiske, 1959). model fit was assessed by inspecting cut-offs for goodness of fit indices: ≤.06 for good model fit using the root mean square error of approximation (rmsea) and the standardized root mean square residual for the within (srmrw) and the between level (srmrb), and ≥.90 for acceptable and ≥.95 for good model fit for the comparative fit index (cfi; browne & cudeck, 1993). assuming mar we treated missing data (4.8% of the missing data-points, in the dataset with the non-lagged variables) using the default fiml algorithm in mplus (muthén & muthén, 2012). we used the robust maximum likelihood estimator (mlr) which corrects standard errors for non-normality. 3. results in order to test structural validity of the state, trait, and variability-constructs of each learning experience, we present a univariate msem specified with one manifest indicator (task-difficulty), two indicators (competence evaluation), and three indicators (intrinsic motivation). to investigate the association between trait and variability-constructs we report on the correlation between these latent constructs. 3.1. univariate models as shown in fig 3, we illustrate how to specify our proposed model using a single item indicator. to identify this model we fixed a number of parameters: all factor loadings (at 1), and residuals (at 0). the pooled within-level variance was ψw1= 0.88 and between ψb1 = 0.26, showing that 22.5% of the variance of task difficulty resided at the between level. we note that the variance of the variability construct, ψb2 = 2.05, was larger than the variance of the trait construct. the association between trait-task-difficulty and variability in task-difficulty was ρ = 0.46, that is the more difficult tasks appeared on average during the week, the more variability in task-difficulty (i.e., larger ups and downs in task difficulty during the week). malmberg  et  al         | f l r 69     figure 3. multilevel structural equation model of latent state, trait and intraindividual variability of task difficulty. note: manifest indicators are diff = task-difficulty as shown in fig 4, we illustrate how to specify our proposed model using two indicators. the pooled within-level variance was ψw1 = 0.32 and between ψb1 = 0.13, showing that 29.5% of the variance of competence beliefs resided at the between level. we note that the variance of the variability construct, ψb2 = 0.66, was larger than the variance of the trait construct. the association between trait intrinsic motivation and variability in intrinsic motivation was ρ = -0.72, that is the more competent students thought they were on average during the week, the less variable they thought their competences were during the week (i.e., smaller ups and downs in competence belief during the week). malmberg  et  al         | f l r 70     figure 4. multilevel structural equation model of latent state, trait and intraindividual variability of competence belief. note: manifest indicators are well = how well?, und = understanding as shown in fig 5, we illustrate how to specify our proposed model using three indicators. the pooled within-level variance was ψw1= 0.51 and between ψb1 = 0.32, showing that 38.7% of the variance of competence evaluation resided at the between level. we note that the variance of the variability construct, ψb2 = 1.26, was larger than the variance of the trait construct. the association between trait-task-difficulty and variability in task-difficulty was ρ = -0.35, that is the more intrinsically motivated students thought they were on average during the week, the less their motivation fluctuated during the week (i.e., smaller ups and downs in intrinsic motivation during the week). malmberg  et  al         | f l r 71     figure 5. multilevel structural equation model of latent state, trait and intraindividual variability of intrinsic motivation. note: manifest indicators are enj = enjoyment, int = interest, and cho = choice. 3.2. multivariate models in models 4 to 6 we present three possible models for investigating convergent validity of the variability constructs (see fig 6). in model 4 we specified three state constructs at the within-level, with three corresponding trait constructs at the between-level, and three variability constructs. the structural parameters of interest are the associations between the trait-construct and variability-construct of task difficulty (ρ = 0.46), competence belief (ρ = -0.64) and intrinsic motivation (ρ = -0.35) respectively. these associations were similar to the ones found in the separate univariate models. in model 5 we specified associations between the three trait-constructs and teacher-reported student task-focus, and between variability constructs and task-focus. more task-focused students, on average during the week, found tasks easier (ρ = -0.19), felt more successful (ρ = 0.38), and were more intrinsically motivated (ρ = 0.25). more task-focused students found tasks of more equal difficulty (ρ = -0.26), fluctuated less in their competence beliefs (ρ = -0.35), but were not significantly less variable in their intrinsic motivation. malmberg  et  al         | f l r 72     in model 6 we specified a higher-order construct of variability. this means that the factor loadings of the higher-order construct explain the associations between the latent constructs. students who were more variable overall (i.e., a higher value on the higher-order variability construct) found, on average during the week, task more difficult (ρ = 0.50), their competence lower (ρ = -0.62) and were less intrinsically motivated (ρ = -0.32). figure 6. multivariate models of state, trait and variability constructs. note: only structural parts of the models shown for clarity. estimates (standardized correlations) are from mplus 7.4 (muthén & muthén, 2012). malmberg  et  al         | f l r 73     4. discussion in order to advance our understanding of students’ learning processes in real time, we investigated intraindividual variability in students’ learning experiences, and convergence between intraindividual variability and teacher-reported task-focus. inclusion of such intraindividual variability construct(s) in process models of learning experiences would expand current modelling practices in the field. up to now, models include: (1) decomposition of learning experiences into within (time-points) and between (students) components, (2) random (moderator) effects of perceptions of the context on learning experiences, and (3) fixed and moderation effects of personal characteristics on learning experiences. as the mssd captures both magnitude and order of events (von neumann et al., 1941; jahng et al., 2008), we created a dataset with lagged variables and specified aggregated variables to use in msems. we specified three latent constructs: a state-construct at the within-level, and a trait and an intraindividual variability construct at the between-level. in response to our first research question regarding the structural validity of state-trait and variability dimensions of learning experiences, we found support for the specificity of the intraindividual variability dimension. importantly, this suggests that intraindividual variability in learning experiences such as motivation, adaptive behaviours, and competence evaluations capture an important dimension of students’ experiences of learning, in addition to variability of constructs in other fields of research, e.g., affect instability (trull et al., 2008). in response to our second research question, we confirmed the hypothesis that trait and variability dimensions of learning experiences converged with teacher-reported task focus. importantly, teacher reports of higher task-focus were related to more adaptive learning experiences on average during the week (i.e., less difficulty, feeling more competent, experiencing higher intrinsic motivation), and to less variability in these same learning experiences (i.e., a smaller magnitude in momentary fluctuations from one learning episode to the next). 4.1. state, trait and individual differences in intraindividual variability we found three sources of support for the distinction between state, trait and variability dimensions of the same construct. first, msem of each learning experience construct in turn (models 1-3) suggested that state, trait and intraindividual variability are separable constructs. importantly, this expands existing two-level models in which states and traits have been modeled at the within and between levels respectively (e.g., roesch et al., 2010). second, inspection of associations between trait and intraindividual variability constructs at the between level suggested convergence, that is traits and variability dimensions of each construct were moderately to strongly associated |ρ| = .35 to .64. third, model 6 suggested that intraindividual variability could be specified as a higher-order construct. taken together, our msem using aggregates of mssds of each indicator was deemed successful for portraying the variability dimension. going beyond previous studies (malmberg et al., 2013; schmitz & skinner, 1993) which have shown that there is more variance within (i.e., intraindividual) than between students (i.e., interindividual difference), we suggest it is possible to retrieve systematic variance of intraindividual experiences by specifying intraindividual variability dimensions of constructs. importantly, this demonstrates that there are systematic individual differences in how students vary within themselves. there are at least two research contexts in which it could be useful to implement the specification of such an intraindividual variability construct in its own right. first, it could be possible to design intervention studies with bursts of collections of intensive longitudinal data (schmitz, 2015; walls, barta, stawski, collyer, & hofer, 2011). if measures of intraindividual variability could be obtained at both preand post-tests, it would be possible to investigate treatment effects geared towards decreasing such intraindividual variability. second, if reports of studentteacher interaction were possible to collect alongside collection of students’ self-reported learning experiences, it would be possible to specify models in which teacher sensitivity to disengagement or off-task behaviour might alleviate intraindividual variability. with regard to questionnaire design, we suggest it could be possible to create at least two types of psychometric measures, for researchers who do not aspire to measure processes by collecting intensive malmberg  et  al         | f l r 74     longitudinal data. first, while previous self-report measures of emotional self-concept (i.e., emotional stability, meaning the perception of feeling calm, emotionally stable and worried; e.g., marsh, 1989) and stability of self-esteem (rosenberg, 1965), have indeed focused on intraindividual variability as a trait, the findings from our present study suggest that it could be possible to create a wider range of “variability-as-atrait” constructs. second, while present observation instruments of students’ engagement and task-focus are designed to capture trait aspects (zhang et al., 2011), future instruments could focus on variability of such observations. 4.2. intraindividual variability and teacher perceptions of task-focus task-focus, as a psychometric construct, is operationalized as a trait-level of students’ adaptive work habits in classrooms, the extent to which each individual student attempts difficult tasks, persists and stays focused on these (nurmi, & aunola, 1998; aunola et al., 2000; zhang et al., 2011). model 5 showed that teacher-reported student task-focus was associated with less difficulty, stronger sense of competence and more intrinsic motivation, in line with hypothesis 1. in a previous study, lower achievers were found to withdraw effort when confronted with a difficult task, while higher achievers exerted more effort (malmberg, walls, et al., 2013). consistent with the idea that students who experience learning as inherently interesting, pay attention and focus on their task at hand in that learning situation (nurmi et al., 2008), teacher rated task-focus was also related with higher levels of intrinsic motivation. this does not mean that teachers can "see" students' motivation as such (lee & reeve, 2012), but rather intrinsic motivation manifested as energized behaviour (sheldon & elliot, 1998). a higher level of task-focus was also related to students’ trait competence evaluation. competence belief also varied from one learning situation to another. future studies should investigate to what extent this is linked with experiences of particular school subjects or particular teachers. taken together, the current and the previous study hint at the importance to further investigate stability and variability in the ways teachers support and place demands on students, that are optimal for students’ learning over time. there were three important associations between task-focus and variability. a higher level of taskfocus was related, first, to less variability in difficultly and competence beliefs, and, second, to a higher level of the higher-order variability-construct. importantly, it appears that variability could be construed as a trait dimension in itself. the variability-construct might play an important role in models of self-regulation (boekaerts & corno, 2005), in which “top-down” self-regulation is typical of students who steer their learning processes by setting goals for enhancing their knowledge by sustaining motivation, rather than being obstructed by situational demands and setbacks typical of “bottom-up” self-regulation. an important future research task would be to investigate intraindividual variability in relation to self-set learning goals (or the lack of such goals). there are two important implications of students’ variability in learning experiences for the different ways in which teachers can support different students. first, the variability in itself demonstrates that all students have ups and downs, less task-focused students more so than more task-focused students. it would be important for teachers to capture the “ups”, particularly of less task-focused students. teachers want to capitalize on students’ “ups” as these would be teachable moments (hamre & pianta, 2005; pianta & hamre, 2009). second, the variability in students’ learning experiences is inherently linked to experiences of the learning context. from the teachers’ point of view there lurks a danger in them classifying students as “engaged” or “disengaged” as all students have their ups and downs. this means that “engaged” students have their disengaged moments. these are moments when teachers can redirect students. it also means that “disengaged” students have their engaged moments. these moments are the ones to capitalize on for learning; the others for redirection. it would be important also for teacher educators to focus on the meaning of intraindividual variability, allowing prospective teachers to focus on changes in student behaviours and actions in real time. in future studies, it would be important to combine studies of students’ intraindividual variability and measure of teacher support. teachers in classrooms can promote students’ learning processes by supporting malmberg  et  al         | f l r 75     their autonomy (reeve, jang, carrell, jeon, & barch, 2004), being involved with students and structuring the learning contents (skinner & belmont, 1993), providing task-contingent praise (deci, koestner, & ryan, 1999), providing feedback directed at reducing discrepancies between current understanding and learning goals (hattie & timperley, 2007), and tailoring goals to individual learners (hattie, biggs, & purdie, 1996; butler & winne, 1995; pianta, belsky, vandergrift, hours, & morrison, 2008), in order to enhance motivation, effort, and selection of optimally difficult tasks. 4.3. limitations there were three limitations of the present study. first, the lagged variables we created spanned from 5 to 29 observations, which is shorter than the typical time-series model. however, for the purpose of calculating the mssd the number of observations might be sufficient (ebner-priemer et al., 2009). future studies of students’ learning processes should aim at collecting more repeated measures for the purpose of creating longer time-series. such a demand needs to be carefully weighed against the risks of response fatigue of participants. second, the model we applied assumes equivalent duration between each time-lag (e.g., jahng et al., 2008). thus our models do not account for unequal number of responses per day, unequal lags between each subsequent reports, and downtime when not at school. the current findings would need to be replicated with models more suitable for unequally spaced lagged data. third, the empirical data stem from a particular age-group in a particular sociocultural context, england, so replications in other contexts and age-groups would be valuable to carry out. 4.4. conclusions variability in intraindividual learning experiences captures the abruptness, frequency, order, and magnitude of students’ “ups and downs” in their engagement during a week at school. teacher perceived student task-focus converged with both trait-levels and variability in task difficulty, competence beliefs and intrinsic motivation. intraindividual variability formed a higher-order construct. overall, our study provides support for intraindividual variability as a construct in its own right, which has the potential to provide novel insight into students’ learning processes. keypoints we investigated intraindividual variability of primary school students' experience of learning. we used the mean square successive differences (mssd) as an index of magnitude of variability. we specified latent state, trait, and intraindividual variability constructs using multilevel structural equation models (msem). higher teacher-reported-task-focus was related to less variability in learning experiences. intraindividual variability is an important educational construct in its own right. acknowledgments the learning every lesson (lel) study was funded by the john fell foundation, and carried out during the first author’s research councils uk (rcuk) fellowship 2007-12. malmberg  et  al         | f l r 76     references aunola, k., nurmi, j.-e., parrila, r., & onatsu-arvilommi, t. (2000). behavioral strategy relating scale ii. unpublished measurement instrument jyväskylä: university of jyväskylä, finland. boekaerts, m., & corno, l. (2005). self-regulation in the classroom: a perspective on assessment and intervention. applied psychology: an international review, 54, 199-231. doi: 10.1111/j.14640597.2005.00205.x browne, m. w., & cudeck, r. (1993). alternative ways of assessing model fit. in k. a. bollen, & j. s. long (eds.), testing structural equation models (pp. 136-162). beverly hills, ca: sage. butler, d. l., & winne, p. h. (1995). feedback and self-regulated learning: a theoretical synthesis. review of educational research, 65, 245-281. doi:10.3102/00346543065003245 campbell, d. t., & fiske, d. w. (1959). convergent and discriminant validation by the multitraitmultimethod matrix. psychological bulletin, 56, 81-105. doi:10.1037/h0046016 deci, e. l., koestner, r., & ryan, r. m. (1999). a meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. psychological bulletin, 125, 627–668. doi:10.1037/0033-2909.125.6.627 ebner-priemer, u. w., eid, m., kleindienst, n., stabenow, s., & trull, t. j. (2009). analytic strategies for understanding affective (in)stability and other dynamic processes in psychopathology. journal of abnormal psychology, 118, 195–202. doi:10.1037/a0014868 eccles, j. s., midgley, c., wigfield, a., buchanan, c. m., reuman, d., flanagan, c., & maciver, d. (1993). development during adolescence: the impact of stage-environment fit on young adolescents’ experiences in schools and in families. american psychologist, 48, 90-101. doi:10.1037/0003066x.48.2.90 eccles, j. s., wigfield, a., & schiefele, u. (1998). motivation to succeed. in w. damon & n. eisenberg, (eds.), handbook of child psychology, 5th ed.: vol 3. social, emotional, and personality development (pp. 1017-1095). hoboken, nj.: john wiley & sons. eid, m., & diener, e. (1999). intraindividual variability in affect: reliability, validity, and personality correlates. journal of personality and social psychology, 76, 662-676. doi:10.1037/00223514.76.4.662 eid, m., & langeheine, r. (2003). separating stable from variable individuals in longitudinal studies by mixture distribution models. measurement: interdisciplinary research and perspectives, 1, 179-206. doi:10.1207/s15366359mea0103_01 eizenman, d. r., nesselroade, j. r., featherman, d. l., & rowe, j.w. (1997). intraindividual variability in perceived control in an older sample: the macarthur successful aging studies. psychology and aging, 12, 489-502. doi:10.1037/0882-7974.12.3.489 georgiou, g., manolitsis, g., nurmi, j.-e., & parrila, r. (2010). does task-focused versus task-avoidance behavior matter for literacy development in an orthographically consistent language? contemporary educational psychology, 35, 1−10. doi:10.1016/j.cedpsych.2009.07.001 gill, p., & r. remedios, r. (2013). how should researchers in education operationalise on-task behaviours? cambridge journal of education, 43, 199-222. doi:10.1080/0305764x.2013.767878 goetz, t., frenzel, a. c., stoeger, h., & hall, n. c. (2010). antecedents of everyday positive emotions: an experience sampling analysis. motivation and emotion, 34, 49-62. doi:10.1007/s11031-009-9152-2 hamre, b. k., & pianta, r. c. (2005). can instructional and emotional support in the first-grade classroom make a difference for children at risk of school failure? child development, 76, 949-967. doi:10.1111/j.1467-8624.2005.00889.x hattie, j., biggs, j., & purdie, n. (1996). effects of learning skills interventions on student learning: a metaanalysis. review of educational research, 66, 99-136. doi:10.3102/00346543066002099 hattie, j., & timperley, h. (2007). the power of feedback. review of educational research, 77, 81-112. doi:10.3102/003465430298487 hirvonen, r., tolvanen, a., aunola, k., & nurmi, j.-e. (2012). the developmental dynamics of taskavoidant behavior and math performance in kindergarten and elementary school. learning and individual differences,22, 715-723. doi: 10.1016/j.lindif.2012.05.014 malmberg  et  al         | f l r 77     hooker, k., nesselroade, d. w., nesselroade, j. r., & lerner, r. m. (1987). the structure of intraindividual temperament in the context of mother-child dyads: p-technique factor analyses of short-term change. developmental psychology, 23, 332-346. doi:10.1037/0012-1649.23.3.332 ilies, r., & judge, t. a. (2002). understanding the dynamic relationships among personality, mood, and job satisfaction: a field experience sampling study. organizational behavior and human decision processes, 89, 1119–1139. doi:10.1016/s0749-5978(02)00018-3 inkinen, m., lonka, k., hakkarainen, k., muukkonen, h., litmanen, t., & salmela-aro, k. (2013). the interface between core affects and the challenge-skill relationship. journal of happiness studies, 15, 891-913. doi:10.1007/s10902-013-9455-6 jahng, s., wood, p. k., & trull, t. j. (2008). analysis of affective instability in ecological momentary assessment: indices using successive difference and group comparison via multilevel modeling. psychological methods, 13, 354–375. doi:10.1037/a0014173 kernis, m. h., grannemann, b. d., & barclay, l. c. (1989). stability and level of self-esteem as predictors of anger arousal and hostility. journal of personality and social psychology, 56, 1013-1022. doi:10.1037/0022-3514.56.6.1013 la guardia, j. g., ryan, r. m., couchman, c. e., & deci, e. l. (2000). within-person variation in security of attachment: a self-determination theory perspective on attachment, need fulfillment, and well-being. journal of personality and social psychology, 79, 367-384. doi:10.1037/0022-3514.79.3.367 larsen, r. j. (1987). the stability of mood variability: a spectral analytic approach to daily mood assessments. journal of personality and social psychology, 52, 1195-1204. doi:10.1037/00223514.52.6.1195 lee, w., & reeve, j. (2012). teachers’ estimates of their students’ motivation and engagement: being in synch with students. educational psychology, 32, 727–747. doi:10.1080/01443410.2012.732385 li, x., & hedeker, d. (2012). a three-level mixed-effects location scale model with an application to ecological momentary assessment data. statistics in medicine, 31, 3192-3210. doi:10.1002/sim.5393 malmberg, l.-e., walls, t., martin, a. j., little, t. d., & lim, w. h. t. (2013). primary school students' learning experiences of, and self-beliefs about competence, effort, and difficulty: random effects models. learning and individual differences, 28, 54–65. doi:10.1016/j.lindif.2013.09.007 malmberg, l.-e., woolgar, c., & martin, a. (2013). quality of measurement of the learning experience questionnaire for personal digital assistants. international journal of quantitative research in education, 1, 275-296. doi:10.1504/ijqre.2013.057689 malmberg, l.-e., pakarinen, e., vasalampi, k., & nurmi, j-e. (2015). students’ school performance, taskfocus, and situation-specific motivation. learning and instruction, 39, 158-167. doi:10.1016/j.learninstruc.2015.05.005 marsh, h. w. (1989). age and sex effects in multiple dimensions of self-concept: preadolescence to early adulthood. journal of educational psychology, 81, 417-430. doi:10.1037/0022-0663.81.3.417 mcconville, c., & cooper, c. (1997). the temporal stability of mood variability. personality and individual differences, 23, 161-164. doi:10.1016/s0191-8869(97)00013-5 morin, a. j. s., marsh, h. w., nagengast, b., & scalas, l. f. (2014). doubly latent multilevel analyses of classroom climate: an illustration. the journal of experimental education, 82, 143-167. doi:10.1080/00220973.2013.769412 musher-eizenman, d. r., nesselroade, j. r., & schmitz, b. (2002). perceived control and academic performance: a comparison of highand low-performing children on within-person change patterns. international journal of behavioral development, 26, 540–547. doi:10.1080/01650250143000517 muthén, l. k. & muthén, b. o. (2012). mplus statistical analysis with latent variables: user’s guide (version 7). los angeles, ca: muthén & muthén. nett, u. e., goetz, t., & hall, n. c. (2011). coping with boredom in school: an experience sampling perspective. contemporary educational psychology, 36,49–59. doi:10.1016/j.cedpsych.2010.10.003 nett, u. e., goetz, t., hall, n. c., & frenzel, a. c. (2012). metacognitive strategies and test performance: an experience sampling analysis of students’ learning behavior. education research international. article id 958319, 16 pages. [online] http://kops.ub.uni-konstanz.de/handle/urn:nbn:de:bsz:352-206102 (accessed 12 april 2013). malmberg  et  al         | f l r 78     nurmi, j.-e., & aunola, k. (1998). observer rating scale of achievement strategies (osas). unpublished measurement instrument. jyväskylä: university of jyväskylä, finland. nurmi, j.-e., hirvonen, r., & aunola, k. (2008). motivation and achievement beliefs in elementary school: a holistic approach using longitudinal data. unterrichtswissenschaft, 36(3), 237–254. doi:10.3262/uw0803237 pianta, r. c., & hamre, b. k. (2009). conceptualization, measurement, and improvement of classroom processes: standardized observation can leverage capacity. educational researcher, 38, 109-119. doi:10.3102/0013189x09332374 pianta, r. c., belsky, j., vandergrift, n., hours, r., & morrison, f. (2008). classroom effects on children’s achievement trajectories in elementary school. american educational research journal, 45, 365-397. doi:10.3102/0002831207308230 ram, n., chow, s.-m., bowles, r. p., wang, l., grimm, k., fujita, f., & nesselroade, j. r. (2005). examining interindividual differences in cyclicity of pleasant and unpleasant affect using spectral analysis and item response modeling. psychometrika, 70, 773-790. doi:10.1007/s11336-001-1270-5 reeve, j., jang, h., carrell, d., jeon, s., & barch, j. (2004). enhancing students’ engagement by increasing teachers’ autonomy support. motivation and emotion, 28, 147-169. doi:10.1023/b:moem.0000032312.95499.6f roesch, s. c., aldridge, a. a., stocking, s. n., villodas, f., leung, q., bartley, c. e. & black, l. j. (2010). multilevel factor analysis and structural equation modeling of daily diary coping data: modeling trait and state variation. multivariate behavioral research, 45, 767-789. doi:10.1080/00273171.2010.519276 rosenberg, m. (1965). society and the adolescent self-image. princeton, n j: princeton university press. schmitz, b. (2006). advantages of studying processes in educational research. learning and instruction, 16, 433-449. doi:10.1016/j.learninstruc.2006.09.004 schmitz, b. (2015). the study of learning processes using time-series analyses [video file]. available from http://www.education.ox.ac.uk/network-on-intrapersonal-research-in-education-nire/seminar1/bernhard-schmitz/ schmitz, b., & skinner, e. (1993). perceived control, effort, and academic performance: interindividual, intraindividual, and multivariate time-series analyses. journal of personality and social psychology, 64, 1010-1028. doi:10.1037/0022-3514.64.6.1010 schmitz, b. & wiese, b. s. (2006). new perspectives for the evaluation of training sessions in self-regulated learning: time-series analyses of diary data. contemporary educational psychology, 31, 64 – 96. doi:10.1016/j.cedpsych.2005.02.002 sheldon, k. m. & elliot, a. j. (1998). not all personal goals are personal: comparing autonomous and controlled reasons for goals as predictors of effort and attainment. personality and social psychology bulletin, 24, 546-557. doi:10.1177/0146167298245010 skinner, e. a., & belmont, m. j. (1993). motivation in the classroom: reciprocal effects of teacher behavior and student engagement across the school year. journal of educational psychology, 85, 571-581. doi:10.1037/0022-0663.85.4.571 trull, t. j., solhan, m. b., tragesser, s. l., jahng, s., wood, p. k., piasecki, t. m., & watson, d. (2008). affective instability: measuring a core feature of borderline personality disorder with ecological momentary assessment. journal of abnormal psychology, 117, 647-661. doi:10.1037/a0012532 tsai, y-m., kunter, m., lüdtke, o. & trautwein, u. (2008). day-to-day variation in competence beliefs: how autonomy support predicts young adolescents’ felt competence in h. w., marsh, r. g., craven, & d. m., mcinerney (eds.), self-processes, learning, and enabling human potential: dynamic new approaches (pp. 119-143). charlotte, nc:information age publishing.. von neumann, j., kent, r. h., bellinson, h. r., & hart, b. i. (1941). the mean square successive difference. the annals of mathematical statistics, 12, 153–162. doi:10.1214/aoms/1177731746 walls, t. a., barta, w. d., stawski, r. s., collyer, c, & hofer, s. m. (2011). time-scale dependent longitudinal designs. in b. laursen, t. d. little, & n. card (eds.), handbook of developmental research methods (pp. 46-64). new york: guilford press. malmberg  et  al         | f l r 79     wilhelm, p., perrez, m., & pawlik, k. (2012). conducting research in daily life. in m. r. mehl, and t. s. conner, (eds.), handbook of research methods for studying daily life (pp. 62-68). new york: guilford press. zhang, x., nurmi, j-e., kiuru, n., lerkkanen, m.-k., & aunola, k. (2011). a teacher-report measure of children's task-avoidant behavior: a validation study of the behavioral strategy rating scale. learning and individual differences, 21, 690–698. doi:10.1016/j.lindif.2011.09.007 zimmerman, b. j. (2000). self-efficacy: an essential motive to learn. contemporary educational psychology, 25, 82–91. doi:10.1006/ceps.1999.1016 malmberg  et  al         | f l r 80     appendix 1. hand-calculation of the mean square successive difference (mssd) this hand-calculation presents the values of panel a in figure 1. the columns below represent: timepoint = 28 time-points for one individual. note that the numeric values of intrt+1 are replicated in the intrt column, only one row below each corresponding intrt+1 value. this gives a 29th time-point; intrt+1 = intrinsic motivation at time-point t+1; intrt = intrinsic motivation at time-point t. this variable was created by lagging the intrt+1 variable one time-point step. time-point t becomes a predictor of time-point t+1; δ = difference between intrt+1 and intrt; δ2 = squared difference between intrt+1 and intrt; l = missing data created by the lag. time-point intrt+1 intrt δ δ2 1 4.0 l l l 2 3.5 4.0 -0.5 0.25 3 2.0 3.5 -1.5 2.25 4 2.5 2.0 0.5 0.25 5 4.0 2.5 1.5 2.25 6 2.0 4.0 -2.0 4.00 7 2.0 2.0 0.0 0.00 8 2.0 2.0 0.0 0.00 9 2.0 2.0 0.0 0.00 10 2.5 2.0 0.5 0.25 11 1.0 2.5 -1.5 2.25 12 2.0 1.0 1.0 1.00 13 2.5 2.0 0.5 0.25 14 2.5 2.5 0.0 0.00 15 2.5 2.5 0.0 0.00 16 2.0 2.5 -0.5 0.25 17 1.0 2.0 -1.0 1.00 18 2.5 1.0 1.5 2.25 19 2.0 2.5 -0.5 0.25 20 1.5 2.0 -0.5 0.25 21 1.5 1.5 0.0 0.00 22 2.5 1.5 1.0 1.00 23 1.5 2.5 -1.0 1.00 24 1.0 1.5 -0.5 0.25 25 2.0 1.0 1.0 1.00 26 1.0 2.0 -1.0 1.00 27 1.0 1.0 0.0 0.00 28 1.0 1.0 0.0 0.00 (29) l 1.0 l l m 2.05 2.05 -0.11 0.78 sd 0.83 0.83 0.89 1.01 r(t+1,t)     0.36 the mssd is the average of δ2 using n-1 (28-1=27) as denominator. r(t+1,t) is the autocorrelation between intrt+1 and intrt. malmberg  et  al         | f l r 81     appendix 2. mplus code for single construct model (intrinsic motivation) title: mssd 28 june 2016 ; data: file is "c:\variability.txt" ; variable: names are studid sequence lag_n diff_t1 diff_t0 well_t1 well_t0 und_t1 und_t0 enj_t1 enj_t0 int_t1 int_t0 cho_t1 cho_t0 focus1 focus2 focus3 avoid1 avoid2 avoid3 ; !enj=enjoyment, int=interest, cho=choice, t1 = time t+1, t0 = time t usevar = enj_t1 int_t1 cho_t1 enj_var int_var cho_var ; !include three observed and three defined variables missing all (-9) ; between enj_var int_var cho_var ; !defined variables are at level 2 cluster = studid; !clustering by student define: enj_va = (enj_t1 enj_t0)**2 ; !squared difference of enjoyment int_va = (int_t1 int_t0)**2 ; !average squared difference of enjoyment cho_va = (cho_t1 cho_t0)**2 ; !squared difference of interest enj_var = cluster_mean (enj_va) ; !average squared difference of interest int_var = cluster_mean (int_va) ; !squared difference of choice cho_var = cluster_mean (cho_va) ; !average squared difference of choice center (grandmean) enj_t1 int_t1 cho_t1 enj_var int_var cho_var ; !grand mean centre level 2 indicators analysis: type = twolevel ; model: %within% w_intr by enj_t1 (a) int_t1 (b) cho_t1 (c) ; ! w_ = state-construct (within-level) ! factor loadings of within and between indicators are equated between level w_intr (var_w) ; !estimate variance, and use for calculating new parameter %between% b_intr by enj_t1 (a) int_t1 (b) cho_t1 (c) ; ! b_intr = trait-construct intr_var by enj_var int_var cho_var ; ! variability construct b_intr with intr_var ; b_intr (var_b) ; malmberg  et  al         | f l r 82     !estimate variance, and use for calculating new parameter int_t1*.05 (br1) ; ! estimate error variance model constraint: br1 > 0 ; !br2 > 0 ; new(var_comp); var_comp = var_b / (var_b + var_w) ; ! calculate intraclass correlation of latent constructs output: stand sampstat tech1 tech2 ;   donker et al publication frontline learning research vol.6 no. 3 (2018) 162 185 issn 2295-3159 a quantitative exploration of two teachers with contrasting emotions: intra-individual process analyses of physiology and interpersonal behavior monika h. donkera, tamara van goga m. tim mainharda adepartment of education, utrecht university, the netherlands article received 14 may 2018 / revised 13 september / accepted 20 september / available online 19 december abstract although the association between teacher-student relations, teacher emotions, and burnout has been proven on a general level, we do not know the exact processes underlying these associations. recently there has been a call for intra-individual process measures that assess what happens from moment-to-moment in class in order to better understand inter-individual differences in emotions and burnout between teachers. this paper explored the use of process measures of teachers’ heart rate and their interpersonal behavior during teaching. our aim was to illustrate different ways of analyzing and combining physiological and observational time-series data and to explore their potential for understanding between-teacher differences. in this illustration, we focused on two teachers who represented contrasting cases in terms of their self-reported teaching-related emotions (i.e., anxiety and relaxation) and burnout. we discuss both univariate process analyses (i.e., trend, autocorrelation, stability) as well as state-of-the-art multivariate process analyses (i.e., cross-correlations, dynamic structural equation modeling). results illustrate how the two teachers differed in the nature of their physiological responses, their interpersonal behavior, and the association between these two process measures over time. along implications and suggestions for further research, it is discussed how the process-based, dynamic assessment of physiology and interpersonal behavior may ultimately help to understand differences in more general teaching-related emotions and burnout. keywords: teacher emotions; process analyses; physiology; heart rate; interpersonal teacher behavior info. corresponding author mail m.h.donker@uu.nl doi: https://doi.org/10.14786/flr.v6i3.372 1. introduction a currently much debated question in educational research is how actual classroom processes are associated with more general outcomes, such as teacher emotions and burnout (e.g., frenzel, 2014). although most theories describe emotions as dynamic intra-individual processes (frijda, 1986; scherer, 2009), many empirical studies only used inter-individual (i.e., between-person) comparisons at one moment in time to investigate them. on a general inter-individual level, the interpersonal relationship between teacher and students has been found to be an important factor in the emergence of negative emotions and teacher burnout (chang, 2009; hoglund, klingle, & hosan, 2015; roorda, koomen, spilt, & oort, 2011; spilt, koomen, & thijs, 2011; van droogenbroeck, spruyt, & vanroelen, 2014). however, these findings at an inter-individual level might not generalize to the intra-individual level (fisher, medaglia, & jeronimus, 2018; murayama et al., 2017), while interventions with teachers most likely need to focus on what happens from moment-to-moment in class (pennings & mainhard, 2016; van vondel, steenbeek, van dijk, & van geert, 2017). the current paper illustrates the use of two process measures capturing teaching as it occurs in terms of a teachers’ physiology and interpersonal behavior. both physiology and interpersonal behavior have been described as theoretical components of emotional processes (scherer, 2009). new methodological developments now enable us to study interpersonal and affective variables as processes at the intra-individual level during real-class teaching. more specifically, the current study explores the value of observational, quantitative time-series data of interpersonal teacher behavior coded from moment-to-moment (lizdek, sadler, woody, ethier, & malet, 2012; sadler, ethier, gunn, duong, & woody, 2009), combined with continuous heart rate data as a physiological proxy for teachers’ affective processes during teaching (cf. de geus, willemsen, klaver, & van doornen, 1995). the aim of the current paper was to illustrate the promises and challenges of such state-of-the-art methods and their analysis by an in-depth quantitative examination of two contrasting cases: two teachers with pronounced differences in their self-reported habitual discrete emotions and burnout levels. 1.1 measuring emotions emotions are an integral part of the classroom setting. many studies in academic settings have focused on discrete emotions of students (such as enjoyment, anxiety, and boredom), and their relation with motivation and achievement (e.g., goetz, frenzel, pekrun, hall, & lüdtke, 2007; pekrun, goetz, titz, & perry, 2002). currently, research on teacher emotions is becoming more and more prominent. teacher emotions have been associated with teacher wellbeing (spilt et al., 2011), the quality of teacher-student interaction (hagenauer, hascher, & volet, 2015), as well as with student outcomes in the affective domain (e.g., enjoyment; frenzel, goetz, lüdtke, pekrun, & sutton, 2009; mainhard, oudman, hornstra, bosker, & goetz, 2018) and the cognitive domain (e.g., report card grades; reyes, brackett, rivers, white, & salovey, 2012). measuring a person’s emotions is challenging (mauss & robinson, 2009), and emotion researchers have used a variety of definitions, approaches, and instruments. kleinginna and kleinginna (1981) compiled a widely used definition based on 102 different definitions used in research on emotions: “emotion is a complex set of interactions among subjective and objective factors, mediated by neural/hormonal systems, which can (a) give rise to affective experiences such as feelings of arousal, pleasure/displeasure; (b) generate cognitive processes such as emotionally relevant perceptual effects, appraisals, labeling processes; (c) activate widespread physiological adjustments to the arousing conditions; and (d) lead to behavior that is often, but not always, expressive, goal-directed, and adaptive’’ (p. 355). scherer (1999) emphasized that emotions emerge based on an individual’s appraisal of events. however, appraisals are only one component in the multi-componential perspective of sutton and wheatley (2003), who define “appraisal, subjective experience, physiological change, emotional expression, and action tendencies’’ (p. 329) as the components of emotions. in line with key theorists (e.g. frijda, 1986; lazarus, 1991), most studies on emotions in academic settings, and on teacher emotions specifically, have adopted this multi-componential view. these studies often measured the affective, cognitive, motivational, and physiological components of discrete emotions with questionnaires (e.g., frenzel et al., 2016; pekrun, goetz, frenzel, barchfeld, & perry, 2011). questionnaires are helpful tools when investigating discrete emotions during salient episodes (becker, goetz, morger, & ranellucci, 2014; goetz et al., 2015), but retrospective bias may occur when assessing emotions after the situation of interest has already ended (wilhelm & grossman, 2010). in addition, these retrospective questionnaires, that are administered once and typically at the end of a day (or longer period), do not capture fluctuations in teachers’ emotions (sutton & wheatley, 2003), whereas current emotion researchers emphasize the value of measuring emotion as a constantly changing process within individuals (goetz et al., 2015; hollenstein, 2015; kuppens, 2015; scherer, 2009). indeed, recent findings indicate that important affective differences between people may not lie in mean levels of emotions, but in how people’s emotions are changing from moment-to-moment (becker, keller, goetz, frenzel, & taxer, 2015; hamaker, 2012). several indicators can be used to assess these affective dynamics, such as variability (i.e., deviation from an individual’s average; wichers, wigman, & myin-germeys, 2015), and inflexibility or inertia (i.e., resistance to change/temporal dependency; hollenstein, 2015; trull, lane, koval, & ebner-priemer, 2015). for example, and emphasizing the need to study emotions as processes, increased variability and higher levels of inertia have been related to less successful coping and emotion regulation (bonanno & burton, 2013), more depressive symptoms (koval, pe, meers, & kuppens, 2013), and low psychological well-being (houben, van den noortgate, & kuppens, 2015). ¬a recent development in the field of emotion research that taps into these continuous, ongoing aspects of emotions, is the experience sampling methodology (esm; csikszentmihalyi & larson, 1987; scollon, kim-prieto, & diener, 2009). this method enables recording of (discrete) feelings and activities, by repeatedly asking people to answer a small set of questions in situ. however, repeatedly reporting on emotions might influence subsequent emotions and patterns of emotions over time (kassam & mendes, 2013; scollon et al., 2009; wilhelm & grossman, 2010), and places a heavy demand on participants (scollon et al., 2009). although esm has successfully been applied in studies on teacher emotions, it also interrupts the ongoing lesson and therefore endangers the authenticity of the measurement (becker et al., 2015). in addition, although esm tackles the issue of retrospective biases, self-report data is sensitive to other biases, such as social desirability (hagenauer et al., 2015). to overcome these biases, it has been argued that research on teacher emotions may benefit from using physiological measures (frenzel, becker-kurz, pekrun, & goetz, 2015). 1.2 physiological measures to assess emotion physiology refers to the study of the structure and function of bodily systems. psychophysiology, more specifically, is concerned with linking the bodily system to the experience, feelings, and behavior of people (cacioppo, tassinary, & berntson, 2017). while it is evident that many physiological responses are associated with emotions, there is no consensus in the field of psychophysiology on the exact linkage between physiological signals and discrete emotions; physiological measures are typically indirect (cacioppo & tassinary, 1990; gratton & fabiani, 2016). there have been efforts to tap into discrete emotions by combining several physiological signals (wilhelm, pfaltz, & grossman, 2006; for a review, see kreibig, 2010), and although some patterns emerged, overall, clear physiological distinctions between emotions are hard to establish. because of this current state of affairs, in the present paper, we conceptualized a physiological reaction (after correction for physical activity) as an affective response. during the last decades, technical innovations have produced several devices that are suitable for ambulatory assessment of physiological data (wilhelm & grossman, 2010). this enables researchers to measure physiological activation as a proxy for affective responses within authentic contexts like classrooms, thereby increasing the ecological validity of the measurements (trull & ebner-priemer, 2013; wilson, 1992). until now, ambulatory devices have mostly been used to compare physiological signals in an aggregated way across two or more conditions (e.g., kreibig, gendolla, & scherer, 2012) and to compare individual mean levels (e.g., pieper, brosschot, van der leeden, & thayer, 2010). in the present study however, our interest lies in moment-to-moment changes in physiological activation. two bodily systems that are regularly examined in psychophysiological research are the autonomic nervous system (ans) and the hypothalamic-pituitary-adrenal (hpa) axis. the hypothalamic-pituitary-adrenal axis is a slowly reacting system, that comes into action when the stressor is large or endures for a longer time (sapolsky, romero, & munck, 2000). the autonomic nervous system is the system that reacts very fast to changes and stressors in the environment (nederhof, marceau, shirtcliff, hastings, & oldehinkel, 2014), and is therefore more suited for research on emotion dynamics at smaller timescales (i.e., minutes or seconds). the autonomic nervous system consists of the sympathetic nervous system (sns; ‘fight-or-flight’ system) and the parasympathetic nervous system (pns; ‘rest-and-digest’ system). in reaction to a stressor, the parasympathetic nervous system puts its actions on hold to support the sympathetic nervous system in the stress response. the sympathetic and parasympathetic nervous system work together to reach an equilibrium during and after a stressful situation (nederhof et al., 2014). there are several indices of autonomic nervous system activation, such as cardiovascular activity (blood circulatory system) and electrodermal activity (sweat glands; mauss & robinson, 2009; wilhelm, grossman, & müller, 2012). electrodermal activity (eda) captures the flow of electricity through the skin of a person. sweat production increases under stress, resulting in increased conductance of the skin (sharma & gedeon, 2012). however, measures of electrodermal activity become less precise when hand palms are too wet and when there are environmental temperature changes (boucsein et al., 2012), or when an emotional trigger is relatively short (oliveira-silva & gonçalves, 2011). the cardiovascular system has been featured in many studies, because it is one of the strongest human biosignals and sensitive to almost all emotions (myrtek, 2004) besides various other physical and psychological processes (berntson, quigley, norman, & lozano, 2016). several cardiovascular measures can be used, such as heart rate (hr), heart rate variability (hrv), blood pressure (bp), and pre-ejection period (pep; mauss & robinson, 2009). heart rate is one of the most widely used biosignals (kreibig, 2010), encompasses both sympathetic and parasympathetic activation (nederhof et al., 2014), and can be measured from moment-to-moment non-invasively (van dijk et al., 2013). therefore, the focus of the present study was on continuous measurement of heart rate as a proxy for an affective response. heart rate is usually converted to beats per minute (bpm), but can also be expressed as heart period, representing the time in milliseconds between two heartbeats (berntson et al., 2016). when applying heart rate measures in a larger sample, it is advised to use heart period, because this measure is less affected by inter-individual variability in baseline heart rate (see berntson, cacioppo, & quigley, 1995; porges & byrne, 1992). heart rate variability (hrv) is another often used measure (e.g., dishman et al., 2000; koval, ogrinz, et al., 2013), and represents the variability of heart period over time. some drawbacks that limit the potential of heart rate variability for betweenand intra-individual comparisons are that it is a purely parasympathetic measure (mauss & robinson, 2009), that it needs to be summarized over longer periods of time, and that it is influenced to a relatively large extent by heart rate, age, and respiration (berntson et al., 2016; myrtek, 2004). further, blood pressure has often been used in ambulatory monitoring studies, but also not on a moment-to-moment basis (holt-lunstad, uchino, smith, olson-cerny, & nealey-moore, 2003; vrijkotte, van doornen, & de geus, 2000). there have been some efforts to measure blood pressure noninvasively (friedman, christie, sargent, & weaver, 2004), but to the best of our knowledge, these methods are only applicable in lab settings due to the large size of the devices. finally, pre-ejection period is an index of sympathetic nervous system activity only, and is very sensitive to changes in posture or physical activity (vrijkotte, van doornen, & de geus, 2004). this limits its use for moment-to-moment measures, because movement cannot be restricted in most real-life activities like teaching. heart rate is of course always related to physical activity and postural changes (e.g., in order to provide muscles with oxygen), which needs to be taken into account in real-life situations (houtveen & de geus, 2009; myrtek, 2004). according to rohrbaugh (2016), physiological changes cannot be interpreted without information on the demands placed on the bodily system by physical activity. therefore, some researchers restricted their analyses to periods with comparable posture and levels of physical activity, which is of course not preferable when studying dynamic processes in real-life settings with many potential changes in posture and movement. myrtek (2004) describes the additional heart rate (ahr) method, which mathematically controls for movement; an increase in heart rate is continuously compared to the heart rate in a pre-defined preceding period, while controlling for physical activity. in the present study, an affective response is assumed, when the heart rate rises to a larger extent than can be accounted for by physical activity (carroll, phillips, & balanos, 2009; myrtek, 2004). we thereby view an increase in heart rate (controlled for physical activity) as indicating a change in the environment that is emotionally meaningful to the teacher (cf. mauss & robinson, 2009). 1.3 interpersonal behavior as a context for affective responses emotions, including physiological responses, are intertwined with the social context in which they are experienced (fischer & van kleef, 2010). people are continually adapting their emotions based on the interactions and emotional displays of people around them (butler, 2015; mainhard et al., 2018; van kleef, 2009). because of this dynamic, in which emotions are affected by and simultaneously affect social interaction and behavior (haines et al., 2016; keltner & haidt, 1999), we investigated teachers’ affective responses in combination with their interpersonal behavior in the classroom context. interpersonal theory is one of the most basic approaches to describing the interpersonal meaning of behavior exhibited in the vicinity of others (horowitz & strack, 2010). interpersonal theory posits that two general dimensions underlie all behavior of people in interaction with others: agency (i.e., dominance, power or social influence) and communion (i.e., friendliness, affection or warmth). these dimensions are at the same time necessary and sufficient to exhaustively capture interpersonal processes, that is, all behavior displayed in the presence of others can be described as a specific combination of these two dimensions (for empirical evidence, see (sadler et al., 2009; wubbels, brekelmans, den brok, & van tartwijk, 2006). the interpersonal circle for teachers (ipc-t; see figure 1; wubbels et al., 2012) is an adaptation of interpersonal theory to the classroom setting.  figure 1. the interpersonal circle for teachers (ipc-t; wubbels et al., 2012). the ipc-t can be used to assess both selfand student-perceptions of teachers’ general agency and communion in class via questionnaires (i.e., questionnaire on teacher interaction, qti; wubbels, créton, & hooymayers, 1985). research has shown that teacher behavior characterized by relatively high levels of agency and communion, that is, directing and helpful behavior, is preferred by both teachers and students, and is associated with more pleasurable student emotions (mainhard et al., 2018) and better cognitive outcomes (roorda, jak, zee, oort, & koomen, 2017; wubbels, brekelmans, mainhard, den brok, & van tartwijk, 2016). sadler et al. (2009) developed a computer-joystick method to observe interpersonal behavior continuously over time: continuous assessment of interpersonal dynamics (caid). with this method, teachers’ agency and communion can be coded by external observers as it unfolds over time (pennings, brekelmans, et al., 2014). research on these moment-to-moment processes has shown that teachers’ interpersonal behavior is not constant but dynamic, and more variability in teacher interpersonal behavior tends to occur in classrooms with poorer social climates (mainhard, pennings, wubbels, & brekelmans, 2012; pennings et al., 2017). the caid method is described in more detail in section 2.3.2. 1.4 the present study the aim of the present study was to showcase different approaches to the analysis of physiological and behavioral time-series data in the context of teacher-emotion research. we illustrate both univariate (i.e., separate for each time series) and multivariate (i.e., combining time series) process analyses to investigate the dynamics of teachers’ physiological responses and interpersonal behavior during teaching. combining continuous physiological measures with continuous coding of interpersonal behavior has the potential to advance our understanding of why teachers vary in their teaching-related emotions. to get an idea of the potential value of process measures for the study of teacher emotion in connection to work-related stress and well-being, we selected two teachers representing contrasting cases (i.e., similar interpersonal student perceptions, but diverging emotion and burnout scores). the teachers’ physiological activation was measured as a proxy for their affective response during a regular classroom lesson, which was also filmed to code teachers’ interpersonal behavior. 2. method 2.1 participants we present an in-depth investigation of process measures of physiology (i.e., heart rate) and interpersonal behavior of two dutch secondary school teachers from a pilot study (n = 10) that was part of the dynamics of emotional processes in teachers (depth) project. teacher a was a younger female teacher of social studies with 5 years of teaching experience and teacher b was a somewhat older male science teacher with 20 years of experience. both classes were of approximately equal size (about 25 students) and consisted of an approximately equal number of boys and girls with a mean age of about 14 years. the selected teachers differed in terms of reported levels of teaching-related discrete emotions (i.e., anxiety and relaxation) and work-related burnout, whereas their general interpersonal teaching style exhibited in class was rather similar according to student questionnaires. an overview of the measures used to characterize the two teachers is provided in table 1. the dissimilar combination of interpersonal and emotional variables was chosen to be able to explore the predictive potential of intra-individual processes and the potential link with teachers’ emotions reported after the lesson.   table 1 description of the selected teachers in terms of emotion, burnout, and interpersonal style a dutch translation and adaptation of the achievement emotions questionnaire (aeq: pekrun et al., 2011) and teacher emotions scales (tes; frenzel et al., 2016). means are based on the current sample (n=10). b utrecht burnout scale (schaufeli & van dierendonck, 2000). population mean scores are based on a sample of teachers in secondary education (n=603). for depersonalization, the norm scores are different for males/females. c questionnaire on teacher interaction (wubbels et al., 2006, 1985). population mean scores are based on a sample of 1668 students (mainhard et al., 2018). 2.2 procedure teachers selected the, according to them, most challenging group of students they currently taught and a lesson with a considerable amount of teacher-student interaction (i.e., no extended seat work). an online questionnaire was administered before the selected lesson, assessing teachers’ personal ideal levels of agency and communion in class, self-efficacy, general teaching-related emotions, burnout, and work engagement. just before the lesson, a heart rate measurement device was attached to the teacher’s chest. lessons were filmed from the back of the classroom, focused on the teacher. during the last ten minutes of the lesson, both the teacher and the students filled out a paper-and-pencil questionnaire focused on their emotions during the lesson. the general set-up of this study was approved by the ethics committee of the faculty of social and behavioral sciences of utrecht university (fetc16-074). all teachers and students included in this study provided their active written informed consent. students that did not agree to participate were either not videotaped or their faces were made indistinct, and their questionnaire data was not used in the analyses. 2.3 measures 2.3.1 heart rate teacher’s cardiovascular and physical activity was measured continuously with the vu university – ambulatory monitoring system (vu-ams; www.vu-ams.nl). seven electrodes were placed on the teacher’s chest to minimize the influence of movement (porges & byrne, 1992) and to get a more reliable estimate of heart rate than with wristbands based on pulse plethysmography (ppg; rohrbaugh, 2016; schäfer & vagedes, 2013). the electrocardiogram (ecg) data was processed in three stages: signal enhancement, data reduction, and statistical analysis (following the guidelines of gratton & fabiani, 2016). to enhance the signal, we checked the raw data for artefacts (i.e., bad ecg quality) and outliers (i.e., anomalous interbeat intervals, wrongly placed r-peaks) with the help of the automated detection of the vu-ams software. data correction or deletion was not necessary for the fragments in the current study. we controlled for the influence of physical activity in line with myrtek’s (2004) additional heart rate approach. a challenge in our case was that physical activity was related not only to heart rate, but also to teachers’ interpersonal behavior (e.g., interpersonal agency or dominant teacher behavior was associated with relatively more movement). thus, the part of the physical activity that showed unique overlap with the heart rate signal was filtered out by 1) regressing physical activity on interpersonal behavior, and 2) regressing heart rate on the residual physical activity. the resulting residual heart rate values were used in the process analyses. in the data reduction stage, one of the main challenges of dynamic process analyses (as compared to averaging over situations or experimental conditions) is choosing a time interval to represent the data (hamaker & wichers, 2017). we followed gratton and fabiani (2016) in selecting a time period that balanced the richness of the data, which we wanted to maximize, and the amount of noise or measurement error, which we wanted to minimize (porges & byrne, 1992). in our case, we used 5 second-averages, as this time interval has at least 4-5 heart beats for each person, but still reflected the small-scale changes we were interested in and which would be lost when applying larger time intervals (e.g., there is no change if data is aggregated into a one-minute aggregate; see figure 2). figure 2. heart rate at different levels of aggregation. 2.3.2 interpersonal teacher behavior two trained coders independently rated teacher’s interpersonal behavior continuously according to the caid approach (lizdek et al., 2012; pennings, brekelmans, et al., 2014; sadler et al., 2009). this approach enables observers to take into account both verbal and non-verbal behavior with a clear interpersonal meaning (markey, lowmaster, & eichler, 2010). while watching the classroom video on one side of a computer screen, teacher’s interpersonal behavior was assessed from moment-to-moment by moving with a joystick apparatus over the interpersonal circle (see figure 1) on the other side of the screen. that is, the observers followed the teacher’s behavior and positioned it on the orthogonal agency and communion axes simultaneously. the position of the joystick on the axes represents the type of behavior, whereas the distance from the origin embodies the intensity. an increase in agency was registered, for example, when teachers were directing the conversation vs. following what students demanded. examples of behaviors that were coded as an increase in communion were smiling or supporting students, and a decrease would have been recorded when teachers were not responding to students or made ‘cold’ comments (ross et al., 2017). the software saves separate values for teacher agency and communion every half-second, on a scale from minus 1000 to 1000. intra-class correlations (consistency) between both coders’ ratings were icc(2,2)= .95 for agency and icc(2,2) = .55 for communion, which can be interpreted as respectively very strong and moderate agreement (cicchetti, 1994; koo & li, 2016; lebreton & senter, 2008). for the analyses, the ratings of both coders were averaged to account for idiosyncratic observations. following thijs, koomen, roorda, and ten hagen (2011), and to match the heart rate data, 5-second averages of agency and communion were used in the analyses. we calculated the 5-second average as an aggregate of the ten underlying half-second codes instead of using only one code per five seconds. with this approach, the averages were controlled for possible outliers that could otherwise have influenced the results. the correlation between the resulting values in both approaches were all above .97 (i.e., for both teachers, and for both agency and communion). 2.4 data-analysis we focused our analysis on the first 12 minutes of the lesson, because the lesson start is at the same time very important and very demanding regarding teacher-student interaction (pennings et al., 2017; van tartwijk, brekelmans, wubbels, fisher, & fraser, 1998). first, we examined univariate characteristics of the heart rate, agency, and communion time series: a) linear, quadratic, and cubic trend models, which indicate whether teacher’s heart rate (or agency, or communion) inor decreased continuously during the lesson and whether it deor accelerated at some points, b) autocorrelation, which indicates to what extent heart rate values carry over to the next moment in time (i.e., temporal dependency), and c) stability, which combines the temporal dependency with the amplitude of heart rate changes. these measures provided not only a description of the time series but are also potential predictors of macro-level outcomes. for example, houben et al. (2015) found in their meta-analysis that higher autocorrelation and lower stability in emotions were associated with lower psychological wellbeing. second, the time series of each teacher were analyzed simultaneously in a multivariate fashion to investigate how they were related to each other over time with a cross-correlation analysis (e.g., the extent to which teacher agency was on average related to heart rate), and by using dynamic structural equation modeling (dsem; asparouhov, hamaker, & muthén, 2017) to get insight in sequential or quasi-causal relations between the three time series within a teacher (e.g., whether a decrease in teacher communion was followed by an increase in agency, heart rate, or both). these multivariate process analyses of intra-individual dynamics provide information about the coupling of interpersonal and affective processes as they unfold over time. the nature of this coupling, again, represents a potential predictor of macro-level outcomes such as general feelings of burnout (cf. hamaker, asparouhov, brose, schmiedek, & muthén, 2018). one issue in time-series analyses is the handling of trends. in our case we chose not to remove trends, but rather analyze them, as they may capture valuable and essential information, and removing them might blur the interpretation of the results (boker, rotondo, xu, & king, 2002; but also see warner, 1998; wu, huang, long, & peng, 2007 for a discussion of detrending time-series data). data and syntax for this article can be found at https://osf.io/p9ak2/. 3. results 3.1 case description figure 3 presents the time series of heart rate, agency, and communion for teacher a and b. each time series consisted of 144 data points (12 minutes x 5-second intervals). in line with the general student perceptions based on the qti, and as could be expected based on research relating general interpersonal student perceptions with coded interpersonal behavior (pennings, van tartwijk, et al., 2014), the mean level of observed interpersonal behavior, for teacher a and b respectively, was similar for both the agency (m = 422, sd = 328; m = 569, sd = 276) and the communion dimension (m = 380, sd = 166; m = 336, sd = 185). for both teachers, the lesson started with approximately five minutes (300s) where students entered the classroom, and the teacher was talking to a smaller group of students, while other students were talking to each other. after these five minutes, teacher a started the lesson with an overview of the lesson, and discussed a previous written assessment including a whole-class review of the test. also, teacher b started with a general review of a test, instructing students with lower scores, followed by an introduction to a group assignment (i.e., to set up a small physics experiment) that students worked on during the remainder of the lesson. the time-series data (see figure 3) and large sd values showed that the two teachers were variable in their heart rate as well as their interpersonal behavior over time during the lesson start, which indicated that a detailed process analyses could potentially add information beyond examining mean-level associations. figure 3. time series for heart rate, agency, and communion. 3.2 univariate process analyses 3.2.1 trend the trend gives an indication of the overall directionality of variables. we explored the existence of linear, quadratic, and cubic trends in the three time series. table 2 shows that the time series for agency and communion for both teachers were best described by the cubic trend model (i.e., linear + quadratic + cubic term), which corresponds to previous findings (pennings et al., 2017). halfway the 12 minutes (approximately 300s), there was an increase in agency for both teachers, reflecting the moment where they took the lead in opening the lesson, and their agency-level stabilized at about 7.5 minutes (450s). regarding communion, teacher a first became friendlier, but then her friendliness decreased while her agency increased (i.e., she became more imposing and strict). for teacher b the general trend was less pronounced, as the lower amount of variance explained by the trends indicated. heart rate showed a significant linear + quadratic trend for teacher a. her heart rate first increased for about eight minutes (480s), and then leveled off slightly. this may indicate an affective response in particular to opening the lesson. the heart rate of teacher b showed a somewhat reversed pattern. after a relatively higher starting level, there was a linear decreasing trend for teacher b, which may be a sign of more effective emotion regulation, that is, the parasympathetic part of the autonomic nervous system might have become more active, resulting in a lower heart rate (gross, 1998; sack, hopper, & lamprecht, 2004). table 2 the amount of explained variance (r2) per time series by adding a linear, quadratic, and cubic trend 3.2.2 autocorrelation autocorrelations indicate how well one state in a time series can be predicted by its previous state, this is also referred to as carry-over effect, inertia, or temporal dependency (houben et al., 2015). a high autocorrelation value on affective variables is usually interpreted as low reactivity to changes in the environment and less recovery, and this has been related to lower psychological wellbeing and psychopathology (houben et al., 2015; wichers et al., 2015). for interpersonal behavior, a high autocorrelation appears to be associated with more positive teacher-student relations (i.e., predictable behavior and few abrupt changes; mainhard et al., 2012; pennings, brekelmans, et al., 2014). in the present study, the autocorrelations for teacher a and b respectively were .47 vs. .27 for heart rate, .99 vs. .98 for agency, and .92 vs. .94 for communion. thus, the autocorrelations indicated that there was a higher carry-over effect of heart rate for teacher a. following hollenstein (2015), this can be seen as teacher a being less flexible in her affective response and potentially less adaptive to specific classroom situations, or being less able to regulate her affective response (e.g., less cognitive reappraisal; gross & john, 2003). for both teachers, there was a large carry-over effect for interpersonal behavior, which indicated that both teachers were quite predictable in their behavior (i.e., not many sudden changes from friendly to unfriendly or from dominant to submissive behavior). we checked whether the high autocorrelations for agency and communion could be explained by the small time frame between two data points (i.e., whether interpersonal behavior might take more than five seconds to change substantially) by testing larger time intervals and by using a lag of two time points. exploratory analyses with larger time intervals (i.e., 10 and 30 seconds) showed that interpersonal behavior still had a relatively large autocorrelation. the autocorrelation became negative when leaving out the first time lag and using a time lag of two, which might indicate a cyclical pattern. however, as research has shown that trends and cyclical patterns are individual teacher characteristics (pennings et al., 2017), and in the absence of strong methodological guidelines (hamaker & wichers, 2017), we did not want to draw conclusions based on the data of only two teachers, but decided to follow previous studies in choosing a first-order time lag and a relatively small time interval (dormann & griffin, 2015). moreover, although both teachers’ level of interpersonal behavior was highly predictable from the value five seconds before, this does not mean that there were no meaningful fluctuations in agency and communion during the lesson start (as is also evident from the graphs presented in figure 3).3.2.3 stability stability combines the temporal dependency with the variability of time series. this measure illustrates not only whether participants are changing (i.e., auto-correlation), but also takes into account how large these changes are (e.g., standard deviation). in the teaching context, small changes from moment-to-moment might be seen as adaptive; however, large changes (instability) may make the teachers’ behavior unpredictable for students and may primarily indicate changes towards more negative teacher behavior (mainhard et al., 2012). also, affective instability has been related to psychological maladjustment (koval, ogrinz, et al., 2013). the mean square successive difference (mssd; jahng, wood, & trull, 2008) is a stability measure which captures this variability between measurements and is not influenced by trends in the data (hamaker, ceulemans, grasman, & tuerlinckx, 2015). as the name suggests, mssd is the average of the squared difference between moment t-1 and t. lower values indicate more stability. the mssd, for teacher a and b respectively, was 152 vs. 44 for heart rate, 1483 vs. 1454 for agency, and 3517 vs. 3051 for communion. these results showed that teacher a was somewhat less stable in all variables compared to teacher b, in particular regarding heart rate, which might be related to her self-reported overall more negative feelings about the lesson. although teacher a was less stable in her heart rate (i.e., higher mssd), there was a larger carry-over effect (i.e., high autocorrelation) at the same time. in this case, the mssd might have been influenced mainly by the higher variability (i.e., standard deviation) of the raw heart rate changes in teacher a (wang, hamaker, & bergeman, 2012). another interpretation might be that teacher a was more rigid in her affective responses, but that changes in affective response (although less frequent) were relatively large and demanding to regulate, for example, when she had to discipline students. such a pattern would correspond with teacher a reporting less relaxation and more anxiety after the lesson, and might thus be a risk factor for psychological problems (hollenstein, 2015; houben et al., 2015; koval, ogrinz, et al., 2013; wichers et al., 2015). 3.3 multivariate process analyses by combining the time series, we explored teachers’ heart rate in the context of specific interpersonal behavior, that is, we examined the intertwinedness of a teacher’s interpersonal behavior and affective response. figure 4 visually combines teacher heart rate with a teacher’s behavioral trajectory projected on the ipc-t. the color of the dots represents the teacher’s heart rate at the moment at which this specific combination of agency and communion was observed. overall, both teachers showed comparable behavioral trajectories (i.e., positive levels of agency and communion), which corresponded to their similar general interpersonal style as perceived by their students (cf. pennings, van tartwijk, et al., 2014). figure 4 shows that teacher a had, based on her individual range, a relatively high heart rate during the entire lesson start (i.e., more dark/red dots overall), and more specifically when showing directing behavior (i.e., combining high agency and moderately high communion). in contrast, teacher b had a relatively low heart rate while displaying directing behavior. the trajectory of teacher b indicated one specific situation, where this teacher had to correct a student, became less friendly (negative communion) and had a relatively higher heart rate.   figure 4. visual combination of continuous assessment of interpersonal dynamics with heart rate values. darker colors indicate higher heart rate values, and the darkest (resp. lightest) dot represents an individual’s maximum (resp. minimum) observed heart rate value. 3.3.1 cross-correlation when combining time series in analyses, the cross-correlation represents the overall coordination between time series over the entire time frame. as with regular correlation coefficients, this value can be interpreted as a combination of the strength and direction of the concurrent relationship between two variables. for both teachers, the cross-correlations indicated a negative association between agency and communion, but more strongly so for teacher a (see table 3). thus, relatively higher levels in agency tended to go together with relatively lower levels in communion (and vice versa). this can be interpreted as a potential interpersonal pitfall, as combining high agency with relatively high communion can be understood as more productive in class (i.e., directing instead of imposing teacher behavior; aelterman, vansteenkiste, van den berghe, de meyer, & haerens, 2014; mainhard et al., 2018). interestingly, for both teachers, changes in affective responses were more strongly related to changes in agency than communion. however, this association was different in nature for the two teachers. for teacher a, the association between agency and affective response was positive (but statistically non-significant), whereas this association was negative for teacher b. thus, teacher b tended to have an affective response (i.e., relatively higher heart rate) in situations characterized by relatively lower teacher agency. again, these patterns may represent potential predictors that help to understand different emotional outcomes for teacher a and b. table 3 cross-correlations between teachers’ affective response (i.e., heart rate), agency, and communion (n=144 time points) 3.3.2 dynamic structural equation modeling cross-correlations describe the nature of the concurrent association between two time series (i.e., at the same point in time). however, cross-correlation analysis does not take into account relationships over time and (the in our case large) autocorrelations, which may lead to overestimation of the cross-correlations (see hamaker et al., 2018; sadler et al., 2009). dynamic structural equation modeling (dsem; asparouhov et al., 2017) is a new technique that gives insight in cross-lagged relations between variables, that is, how one variable is associated with other variables at the previous time point, while controlling for their own value at the previous time point. the dsem results for teacher a and b are presented in figure 5. these indicated a positive sequential relation between heart rate at t-1 and agency at t for teacher a. thus, her affective response may have resulted from anticipating situations where she had to become more agentic in class. also, agency at t-1 had a negative association with communion at t for this teacher. this corresponds to the negative association between agency and communion we discussed above (i.e., moving towards imposing or strict behavior); but in addition the cross-lagged effect suggested that less friendly behavior of teacher a (lower communion) tended to be preceded by higher agency and not the other way around. in contrast, for teacher b, agency was negatively associated with heart rate at the next time point. this teacher seems to be less aroused after taking a position with relatively more agency (e.g., he may feel more at ease in situations where he has the lead in class). this might be related to his level of experience in teaching, as beginning teachers often have more difficulty with this type of classroom management (aldrup, klusmann, & lüdtke, 2017; brekelmans, wubbels, & van tartwijk, 2005). figure 5. dynamic structural equation model (dsem; asparouhov et al., 2017). the subscript t refers to a time point and the subscript t-1 refers to the previous time point of the time series. the reported autocorrelations are slightly different from section 3.2.2 due to the addition of the cross-lagged associations. only statically significant cross-lagged effects are presented (n=144 time points per teacher). hr = heart rate; ag = agency; com = communion. 4. discussion teacher-student relations have been associated with teacher outcomes, such as emotions and burnout, on a general level (chang, 2009; roorda et al., 2011; spilt et al., 2011). to better understand how these associations emerge from the classroom situation, and in order to design individualized interventions, there has been a recent call for intra-individual process analysis in educational research (murayama et al., 2017; pennings & mainhard, 2016; van vondel et al., 2017). we therefore explored the value of measuring and analyzing teachers’ interpersonal behavior and physiological processes as two important mechanisms in the emergence of more habitual emotions. the present study showed that in-depth intra-individual analyses of physiology and interpersonal behavior uncovered potentially meaningful teacher-specific characteristics that may go unnoticed when using, for example, more general mean-level indicators. the two teachers examined here showed marked differences in intra-individual tendencies, which may be related to divergent self-reported emotions and burnout, despite their similar habitual interpersonal styles in class as perceived by their students and comparable mean levels of observed interpersonal behavior. the use of physiological measures enabled us to get some insight into teachers’ affective responses without disrupting the teaching process (mauss & robinson, 2009) and to reduce issues with social desirability, retrospective bias, and high cognitive load (becker et al., 2015; goetz et al., 2015; scollon et al., 2009; wilhelm & grossman, 2010). moreover, we found that heart rate measures discriminated between both teachers, even when their interpersonal behavior during the lesson start was relatively similar. especially combining physiological data with teachers’ interpersonal behavior on a moment-to-moment level, which is directly in line with prominent theories on human emotional functioning (frijda, kuipers, & ter schure, 1989; scherer, 2009), enabled us to estimate teacher-specific links between physiology and interpersonal behavior (i.e., cross-correlations or cross-lagged effects). this link may represent teachers’ individual action tendencies (i.e., their habitual reaction pattern to stimuli in the classroom environment), which is one important component in the process leading to emotional outcomes (frijda et al., 1989; scherer, 2009). moreover, it might not be the situation itself that leads to emotional outcomes, but rather the individual appraisal of the situation (chang, 2013; frenzel, 2014; moors, ellsworth, scherer, & frijda, 2013). this might explain why similar interpersonal behavior could result in different emotions for the teachers in our study. for example, whether the situation was in line with the teachers’ lesson goals, or whether the teacher felt able to cope with the situation, is likely to have influenced their physiological (affective) and behavioral response (kreibig et al., 2012). it might be beneficial for teachers and teacher educators to become aware of such (often tacit) appraisal processes and action tendencies, for example through video feedback coaching (e.g., van vondel et al., 2017). beyond investigating more general patterns of within-lesson physiological and interpersonal processes, the richness of the collected data in the present study would also allow pinpointing the specific classroom situations that trigger these particular affective responses. heart rate could be used to select moments during the lesson (e.g., selecting episodes with a particularly high heart rate) and teachers could be asked to reflect on the causes and consequences of their behavior (e.g., their appraisals and expectations). this could be a first step towards identifying for example less functional appraisals or emotion regulation strategies (gross & john, 2003). 4.1 limitations and future directions we used a quantitative contrasting cases approach with two teachers differing in their general level of emotions and burnout to gauge the value of intra-individual process measures for explaining inter-individual differences between teachers. case studies are quite frequent in the emerging field of intensive longitudinal data due to the large time investment needed to collect, code, and analyze the data. however, we agree that the preeminent value of physiological and moment-to-moment data lies in their explanatory value of higher-level outcomes in a larger sample (hamaker & wichers, 2017; martin et al., 2015). these multilevel models are ultimately needed to check our current speculation about the relation between moment-to-moment dynamics of physiology and interpersonal behavior and general-level teacher emotional outcomes. in larger samples it would be important to investigate also how other teacher characteristics that were contrasting in our two teachers, such as gender or experience, are related to both intra-individual dynamics as well as between-person outcomes. dsem provides the opportunity to model these relations in an integrated analytical framework, i.e., between time-series variables (e.g., autocorrelations and cross-lagged associations) and higher-level predictor or outcome variables (e.g., self-reported emotions), by modelling random slopes that allow for individual differences in the time-series analysis (asparouhov et al., 2017; hamaker et al., 2018). although recruiting teachers for this type of intensive longitudinal studies is challenging, a relatively large sample (n>50) might be needed to test these multilevel models (schultzberg & muthén, 2018). it should also be taken into account that the dsem outcomes only hold for the specific time interval tested, and that there might be, as in most statistical analyses, omitted variables that may change the model (hamaker et al., 2018). future research should therefore also include other time lags or implement continuous time models (cf. deboeck & preacher, 2015). there are several other avenues for future research. in the current study, we chose to use heart rate as an indicator of teachers’ affective response, because of its reactivity and robustness in terms of signal quality in real-life situations. other physiological measures (such as electrodermal activity or heart rate variability) might be worthwhile to explore as well, as there might be different relations with interpersonal behavior and emotions (kreibig, 2010). we used the additional heart rate approach to control the heart rate signal for physical activity before the analyses. however, in a multilevel framework, physical activity could be added directly to the model. also triangulation with other unobtrusive measures such as automatic face detection might be a possible avenue for further research on teachers’ emotional processes (e.g., bosch et al., 2015). furthermore, combining several time-series descriptors and cross-lagged patterns might help to understand their interrelatedness (e.g., by using latent class/mixture modeling; flunger et al., 2015; van den bergh & vermunt, 2017). finally, teachers’ interpersonal behavior is only one of the two constituent parts that make up classroom interaction. in future research, it would be interesting to combine physiological measures with observation of students’ interpersonal behavior, to see whether, for example, refraining from unfriendly reactions to hostile student behavior is as challenging for teachers as is often claimed (de jong, van tartwijk, verloop, veldman, & wubbels, 2012; pennings et al., 2017; sadler et al., 2009). 4.2 conclusion to conclude, we want to emphasize that moment-to-moment behavioral dynamics as well as physiological data have a role to play in gaining understanding of the complex interpersonal and affective processes underlying teaching and emotions. based on our exploration of two contrasting cases, we believe that differences between teachers in their intra-individual physiological and behavioral processes may reflect the working of teachers’ action tendencies and appraisal processes and that these might be related to teacher emotions and burnout symptoms. future research may need to combine different physiological indicators and both interand intra-individual analyses to sketch a more complete picture of emotionally relevant processes in the classroom (cf. mauss & robinson, 2009). in the end, this might help to advance teachers’ interpersonal behavior and affective responses during the lesson, and thereby improve both teachers’ job satisfaction (chang, 2013; hakanen, bakker, & schaufeli, 2006) and student outcomes (arens & morin, 2016; hoglund et al., 2015). keypoints teacher-student relations, teacher emotions, and burnout are related on a general level, but information is missing on underlying classroom processes continuous assessment of interpersonal dynamics is a method to code teachers’ interpersonal behavior during teaching from moment-to-moment heart rate is a non-invasive and objective physiological measure, and can be used as a continuous indicator of teachers’ affective responses teachers differed in their intra-individual physiological and behavioral processes, potentially indicating individual action tendencies and appraisals time-series analyses have the potential to strengthen research on the link between classroom social processes, teacher emotions, and burnout acknowledgments this research was funded by the netherlands initiative for education research (nro/proo grant 405-14-300-039), which resides under the netherlands organization for scientific research (nwo). the authors would like to thank the teachers and their classes that participated in the study, and the students for their assistance with data collection and coding. references aelterman, n., vansteenkiste, m., van den berghe, l., de meyer, j., & haerens, l. (2014). fostering a need-supportive teaching style: intervention effects on physical education teachers’ beliefs and teaching behaviors. journal of sport and exercise psychology, 36, 595–609. https://doi.org/10.1123/jsep.2013-0229 aldrup, k., klusmann, u., & lüdtke, o. (2017). does basic need satisfaction mediate the link between stress exposure and well-being? a diary study among beginning teachers. learning and instruction, 50, 21–30. https://doi.org/10.1016/j.learninstruc.2016.11.005 arens, a. k., & morin, a. j. s. (2016). relations between teachers’ emotional exhaustion and students’ educational outcomes. journal of educational psychology, 108, 800–813. https://doi.org/10.1037/edu0000105 asparouhov, t., hamaker, e. l., & muthén, b. (2017). dynamic structural equation models.structural equation modeling: a multidisciplinary journal, 0, 1–30. https://doi.org/10.1080/10705511.2017.1406803 becker, e. s., goetz, t., morger, v., & ranellucci, j. (2014). the importance of teachers’ emotions and instructional behavior for their students’ emotions an experience sampling analysis. teaching and teacher education, 43, 15–26. https://doi.org/10.1016/j.tate.2014.05.002 becker, e. s., keller, m. m., goetz, t., frenzel, a. c., & taxer, j. l. (2015). antecedents of teachers’ emotions in the classroom: an intraindividual approach. frontiers in psychology, 6, 1–12. https://doi.org/10.3389/fpsyg.2015.00635 berntson, g. g., cacioppo, j. t., & quigley, k. s. (1995). the metrics of cardiac chronotropism: biometric perspectives. psychophysiology , 32, 162–171. https://doi.org/10.1111/j.1469-8986.1995.tb03308.x berntson, g. g., quigley, k. s., norman, g. j., & lozano, d. l. (2016). cardiovascular psychophysiology. in j. t. cacioppo, l. g. tassinary, & g. g. berntson (eds.), handbook of psychophysiology(4th ed., pp. 183–216). cambridge, england: cambridge university press. boker, s. m., rotondo, j. l., xu, m., & king, k. (2002). windowed cross-correlation and peak picking for the analysis of variability in the association between behavioral time series. psychological methods, 7, 338–355. https://doi.org/10.1037//1082-989x.7.3.338 bonanno, g. a., & burton, c. l. (2013). regulatory flexibility: an individual differences perspective on coping and emotion regulation. perspectives on psychological science, 8, 591–612. https://doi.org/10.1177/1745691613504116 bosch, n., d’mello, s., baker, r., ocumpaugh, j., shute, v., ventura, m., … zhao, w. (2015). automatic detection of learning-centered affective states in the wild. in proceedings of the 20th international conference on intelligent user interfaces (pp. 379–388). https://doi.org/10.1145/2678025.2701397 boucsein, w., fowles, d. c., grimnes, s., ben-shakhar, g., roth, w. t., dawson, m. e., & filion, d. l. (2012). publication recommendations for electrodermal measurements. psychophysiology, 49, 1017–1034. https://doi.org/10.1111/j.1469-8986.2012.01384.x brekelmans, m., wubbels, t., & van tartwijk, j. (2005). teacher-student relationships across the teaching career. international journal of educational research,43, 55–71. https://doi.org/10.1016/j.ijer.2006.03.006 butler, e. a. (2015). interpersonal affect dynamics: it takes two (and time) to tango. emotion review, 7, 336–341. https://doi.org/10.1177/1754073915590622 cacioppo, j. t., & tassinary, l. g. (1990). inferring psychological significance from physiological signals. american psychologist, 45, 16–28. https://doi.org/10.1037/0003-066x.45.1.16 cacioppo, j. t., tassinary, l. g., & berntson, g. g. (2017). strong inference in psychophysiological science. in j. t. cacioppo, l. g. tassinary, & g. g. berntson (eds.), handbook of psychophysiology(4th ed., pp. 3–15). cambridge, england: cambridge university press. carroll, d., phillips, a. c., & balanos, g. m. (2009). metabolically exaggerated cardiac reactions to acute psychological stress revisited. psychophysiology, 46, 270–275. https://doi.org/10.1111/j.1469-8986.2008.00762.x chang, m. l. (2009). an appraisal perspective of teacher burnout: examining the emotional work of teachers. educational psychology review, 21, 193–218. https://doi.org/10.1007/s10648-009-9106-y chang, m. l. (2013). toward a theoretical model to understand teacher emotions and teacher burnout in the context of student misbehavior: appraisal, regulation and coping. motivation and emotion, 37, 799–817. https://doi.org/10.1007/s11031-012-9335-0 cicchetti, d. v. (1994). guidelines, criteria, and rules of thumb for evaluating normed and standardized assessment instruments in psychology. psychological assessment,6, 284–290. https://doi.org/10.1037/1040-3590.6.4.284 csikszentmihalyi, m., & larson, r. (1987). validity and reliability of the experience-sampling method. journal of nervous and mental disease, 175, 526–536. https://doi.org/10.1007/978-94-017-9088-8 de geus, e. j. c., willemsen, g. h. m., klaver, c. h. a. m., & van doornen, l. j. p. (1995). ambulatory measurement of respiratory sinus arrhythmia and respiration rate. biological psychology, 41, 205–227. https://doi.org/10.1016/0301-0511(95)05137-6 de jong, r. j., van tartwijk, j., verloop, n., veldman, i., & wubbels, t. (2012). teachers’ expectations of teacher–student interaction: complementary and distinctive expectancy patterns. teaching and teacher education, 28, 948–956. https://doi.org/10.1016/j.tate.2012.04.009 deboeck, p. r., & preacher, k. j. (2016). no need to be discrete: a method for continuous time mediation analysis.structural equation modeling: a multidisciplinary journal, 23, 61–75. https://doi.org/10.1080/10705511.2014.973960 dishman, r. k., nakamura, y., garcia, m. e., thompson, r. w., dunn, a. l., & blair, s. n. (2000). heart rate variability, trait anxiety, and perceived stress among physically fit men and women. international journal of psychophysiology,37, 121–133. https://doi.org/10.1016/s0167-8760(00)00085-4 dormann, c., & griffin, m. a. (2015). optimal time lags in panel studies. psychological methods, 20, 489–505. https://doi.org/10.1037/met0000041 fischer, a. h., & van kleef, g. a. (2010). where have all the people gone? a plea for including social interaction in emotion research. emotion review, 2, 208–211. https://doi.org/10.1177/1754073910361980 fisher, a. j., medaglia, j. d., & jeronimus, b. f. (2018). lack of group-to-individual generalizability is a threat to human subjects research. proceedings of the national academy of sciences, 201711978. https://doi.org/10.1073/pnas.1711978115 flunger, b., trautwein, u., nagengast, b., lüdtke, o., niggli, a., & schnyder, i. (2015). the janus-faced nature of time spent on homework: using latent profile analyses to predict academic achievement over a school year. learning and instruction, 39, 97–106. https://doi.org/10.1016/j.learninstruc.2015.05.008 frenzel, a. c. (2014). teacher emotions. in e. a. linnenbrink-garcia & r. pekrun (eds.), international handbook of emotions in education (pp. 494–519). new york, ny: routledge. frenzel, a. c., becker-kurz, b., pekrun, r., & goetz, t. (2015). teaching this class drives me nuts! examining the person and context specificity of teacher emotions. plos one, 10(6), 1–15. https://doi.org/10.1371/journal.pone.0129630 frenzel, a. c., goetz, t., lüdtke, o., pekrun, r., & sutton, r. e. (2009). emotional transmission in the classroom: exploring the relationship between teacher and student enjoyment. journal of educational psychology, 101, 705–716. https://doi.org/10.1037/a0014695 frenzel, a. c., pekrun, r., goetz, t., daniels, l. m., durksen, t. l., becker-kurz, b., & klassen, r. (2016). measuring enjoyment, anger, and anxiety during teaching: the teacher emotions scales (tes). contemporary educational psychology, 46, 148–163. https://doi.org/10.1016/j.cedpsych.2016.05.003 friedman, b. h., christie, i. c., sargent, s. l., & weaver, j. b. (2004). self-reported sensitivity to continuous noninvasive blood pressure monitoring via the radial artery. journal of psychosomatic research, 57, 119–121. https://doi.org/10.1016/s0022-3999(03)00597-x frijda, n. h. (1986).the emotions. cambridge, england: cambridge university press. frijda, n. h., kuipers, p., & ter schure, e. (1989). relations among emotion, appraisal, and emotional action readiness. journal of personality and social psychology,57, 212–228. https://doi.org/10.1037/0022-3514.57.2.212 goetz, t., becker, e. s., bieg, m., keller, m. m., frenzel, a. c., & hall, n. c. (2015). the glass half empty: how emotional exhaustion affects the state-trait discrepancy in self-reports of teaching emotions. plos one, 10(9), 1–14. https://doi.org/10.1371/journal.pone.0137441 goetz, t., frenzel, a. c., pekrun, r., hall, n. c., & lüdtke, o. (2007). betweenand within-domain relations of students’ academic emotions. journal of educational psychology, 99, 715–733. https://doi.org/10.1037/0022-0663.99.4.715 gratton, g., & fabiani, m. (2016). biosignal processing in psychophysiology: principles and current developments. in j. t. cacioppo, l. g. tassinary, & g. g. berntson (eds.), handbook of psychophysiology(4th ed., pp. 628–661). cambridge, england: cambridge university press. gross, j. j. (1998). antecedentand response-focused emotion regulation: divergent consequences for experience, expression, and physiology. journal of personality and social psychology, 74, 224–237. https://doi.org/10.1037/0022-3514.74.1.224 gross, j. j., & john, o. p. (2003). individual differences in two emotion regulation processes: implications for affect, relationships, and well-being. journal of personality and social psychology, 85, 348–362. https://doi.org/10.1037/0022-3514.85.2.348 hagenauer, g., hascher, t., & volet, s. e. (2015). teacher emotions in the classroom: associations with students’ engagement, classroom discipline and the interpersonal teacher-student relationship. european journal of psychology of education, 30, 385–403. https://doi.org/10.1007/s10212-015-0250-0 haines, s. j., gleeson, j., kuppens, p., hollenstein, t., ciarrochi, j., labuschagne, i., … koval, p. (2016). the wisdom to know the difference: strategy-situation fit in emotion regulation in daily life is associated with well-being. psychological science, 27, 1651–1659. https://doi.org/10.1177/0956797616669086 hakanen, j. j., bakker, a. b., & schaufeli, w. b. (2006). burnout and work engagement among teachers. journal of school psychology, 43, 495–513. https://doi.org/10.1016/j.jsp.2005.11.001 hamaker, e. l. (2012). why researchers should think “within-person”. a paradigmatic rationale. in m. r. mehl & t. s. conner (eds.), handbook of research methods for studying daily life(pp. 43–61). new york, ny: guilford publications. hamaker, e. l., asparouhov, t., brose, a., schmiedek, f., & muthén, b. (2018). at the frontiers of modeling intensive longitudinal data: dynamic structural equation models for the affective measurements from the cogito study. multivariate behavioral research, 3171, 1–22. https://doi.org/10.1080/00273171.2018.1446819 hamaker, e. l., ceulemans, e., grasman, r. p. p. p., & tuerlinckx, f. (2015). modeling affect dynamics: state of the art and future challenges. emotion review,7, 316–322. https://doi.org/10.1177/1754073915590619 hamaker, e. l., & wichers, m. (2017). no time like the present: discovering the hidden dynamics in intensive longitudinal data. current directions in psychological science, 26, 10–15. https://doi.org/10.1177/0963721416666518 hoglund, w. l. g., klingle, k. e., & hosan, n. e. (2015). classroom risks and resources: teacher burnout, classroom quality and children’s adjustment in high needs elementary schools. journal of school psychology, 53, 337–357. https://doi.org/10.1016/j.jsp.2015.06.002 hollenstein, t. (2015). this time, it’s real: affective flexibility, time scales, feedback loops, and the regulation of emotion. emotion review, 7, 308–315. https://doi.org/10.1177/1754073915590621 holt-lunstad, j., uchino, b. n., smith, t. w., olson-cerny, c., & nealey-moore, j. b. (2003). social relationships and ambulatory blood pressure: structural and qualitative predictors of cardiovascular function during everyday social interactions. health psychology, 22, 388–397. https://doi.org/10.1037/0278-6133.22.4.388 horowitz, l. m., & strack, s. (2010). handbook of interpersonal psychology: theory, research, assessment, and therapeutic interventions . hoboken, new jersey: john wiley & sons. houben, m., van den noortgate, w., & kuppens, p. (2015). the relation between short-term emotion dynamics and psychological well-being: a meta-analysis. psychological bulletin, 141, 901–930. https://doi.org/10.1037/a0038822 houtveen, j. h., & de geus, e. j. c. (2009). noninvasive psychophysiological ambulatory recordings: study design and data analysis strategies. european psychologist,14, 132–141. https://doi.org/10.1027/1016-9040.14.2.132 jahng, s., wood, p. k., & trull, t. j. (2008). analysis of affective instability in ecological momentary assessment: indices using successive difference and group comparison via multilevel modeling. psychological methods, 13, 354–375. https://doi.org/10.1037/a0014173 kassam, k. s., & mendes, w. b. (2013). the effects of measuring emotion: physiological reactions to emotional situations depend on whether someone is asking. plos one, 8(6), 1–8. https://doi.org/10.1371/journal.pone.0064959 keltner, d., & haidt, j. (1999). social functions of emotions at four levels of analysis. cognition & emotion, 13, 505–522. https://doi.org/10.1080/026999399379168 kleinginna, p. r., & kleinginna, a. m. (1981). a categorized list of motivation definitions, with a suggestion for a consensual definition. motivation and emotion, 5, 263–291. https://doi.org/10.1007/bf00993889 koo, t. k., & li, m. y. (2016). a guideline of selecting and reporting intraclass correlation coefficients for reliability research. journal of chiropractic medicine,15, 155–163. https://doi.org/10.1016/j.jcm.2016.02.012 koval, p., ogrinz, b., kuppens, p., van den bergh, o., tuerlinckx, f., & sütterlin, s. (2013). affective instability in daily life is predicted by resting heart rate variability. plos one, 8 (11), 1–10. https://doi.org/10.1371/journal.pone.0081536 koval, p., pe, m. l., meers, k., & kuppens, p. (2013). affect dynamics in relation to depressive symptoms: variable, unstable or inert? emotion, 13, 1132–1141. https://doi.org/10.1037/a0033579 kreibig, s. d. (2010). autonomic nervous system activity in emotion: a review. biological psychology, 84, 394–421. https://doi.org/10.1016/j.biopsycho.2010.03.010 kreibig, s. d., gendolla, g. h. e., & scherer, k. r. (2012). goal relevance and goal conduciveness appraisals lead to differential autonomic reactivity in emotional responding to performance feedback. biological psychology, 91, 365–375. https://doi.org/10.1016/j.biopsycho.2012.08.007 kuppens, p. (2015). it’s about time: a special section on affect dynamics. emotion review, 7, 297–300. https://doi.org/10.1177/1754073915590947 lazarus, r. s. (1991). emotion and adaptation. new york: oxford university press. lebreton, j. m., & senter, j. l. (2008). answers to 20 questions about interrater reliability and interrater agreement. organizational research methods, 11, 815–852. https://doi.org/10.1177/1094428106296642 lizdek, i., sadler, p., woody, e., ethier, n., & malet, g. (2012). capturing the stream of behavior: a computer-joystick method for coding interpersonal behavior continuously over time. social science computer review, 30, 513–521. https://doi.org/10.1177/0894439312436487 mainhard, m. t., oudman, s., hornstra, l., bosker, r. j., & goetz, t. (2018). student emotions in class: the relative importance of teachers and their interpersonal relations with students. learning and instruction, 53, 109–119. https://doi.org/10.1016/j.learninstruc.2017.07.011 mainhard, m. t., pennings, h. j. m., wubbels, t., & brekelmans, m. (2012). mapping control and affiliation in teacher-student interaction with state space grids. teaching and teacher education, 28, 1027–1037. https://doi.org/10.1016/j.tate.2012.04.008 markey, p., lowmaster, s., & eichler, w. (2010). a real-time assessment of interpersonal complementarity. personal relationships, 17, 13–25. https://doi.org/10.1111/j.1475-6811.2010.01249.x martin, a. j., papworth, b., ginns, p., malmberg, l.-e., collie, r. j., & calvo, r. a. (2015). real-time motivation and engagement during a month at school: every moment of every day for every student matters. learning and individual differences, 38, 26–35. https://doi.org/10.1016/j.lindif.2015.01.014 mauss, i. b., & robinson, m. d. (2009). measures of emotion: a review. cognition & emotion, 23, 209–237. https://doi.org/10.1080/02699930802204677 moors, a., ellsworth, p. c., scherer, k. r., & frijda, n. h. (2013). appraisal theories of emotion: state of the art and future development. emotion review,5, 119–124. https://doi.org/10.1177/1754073912468165 murayama, k., goetz, t., malmberg, l.-e., pekrun, r., tanaka, a., & martin, a. j. (2017). within-person analysis in educational psychology: importance and illustrations. british journal of educational psychology monograph series ii: psychological aspects of education --current trends: the role of competence beliefs in teaching and learning , 71–87. myrtek, m. (2004). heart and emotion. ambulatory monitoring studies in everyday life. göttingen, germany: hogrefe & huber publishers. nederhof, e., marceau, k., shirtcliff, e. a., hastings, p. d., & oldehinkel, a. j. (2014). autonomic and adrenocortical interactions predict mental health in late adolescence: the trails study. journal of abnormal child psychology, 43, 1–15. https://doi.org/10.1007/s10802-014-9958-6 oliveira-silva, p., & gonçalves, ó. f. (2011). responding empathically: a question of heart, not a question of skin. applied psychophysiology biofeedback, 36, 201–207. https://doi.org/10.1007/s10484-011-9161-2 pekrun, r., goetz, t., frenzel, a. c., barchfeld, p., & perry, r. p. (2011). measuring emotions in students’ learning and performance: the achievement emotions questionnaire (aeq). contemporary educational psychology, 36, 36–48. https://doi.org/10.1016/j.cedpsych.2010.10.002 pekrun, r., goetz, t., titz, w., & perry, r. p. (2002). academic emotions in students’ self-regulated learning and achievement: a program of qualitative and quantitative research. educational psychologist, 37, 95–105. https://doi.org/10.1207/s15326985ep3702 pennings, h. j. m., brekelmans, m., sadler, p., claessens, l. c. a., van der want, a. c., & tartwijk, j. van. (2017). interpersonal adaptation in teacher-student interaction. learning and instruction, 55, 41–57. https://doi.org/10.1016/j.learninstruc.2017.09.005 pennings, h. j. m., brekelmans, m., wubbels, t., van der want, a. c., claessens, l. c. a., & van tartwijk, j. (2014). a nonlinear dynamical systems approach to real-time teacher behavior: differences between teachers. nonlinear dynamics, psychology, and life sciences, 18, 23–45. pennings, h. j. m., & mainhard, m. t. (2016). analyzing teacher-student interactions with state space grids. in m. koopmans & d. stamovlasis (eds.), complex dynamical systems in education: concepts, methods, and applications (pp. 233–271). cham, switzerland: springer. https://doi.org/10.1007/978-3-319-27577-2 pennings, h. j. m., van tartwijk, j., wubbels, t., claessens, l. c. a., van der want, a. c., & brekelmans, m. (2014). real-time teacher–student interactions: a dynamic systems approach. teaching and teacher education, 37, 183–193. https://doi.org/10.1016/j.tate.2013.07.016 pieper, s., brosschot, j. f., van der leeden, r., & thayer, j. f. (2010). prolonged cardiac effects of momentary assessed stressful events and worry episodes. psychosomatic medicine, 72, 570–577. https://doi.org/10.1097/psy.0b013e3181dbc0e9 porges, s. w., & byrne, e. a. (1992). research methods for measurement of heart rate and respiration. biological psychology, 34, 93–130. https://doi.org/10.1016/0301-0511(92)90012-j reyes, m. r., brackett, m. a., rivers, s. e., white, m., & salovey, p. (2012). classroom emotional climate, student engagement, and academic achievement. journal of educational psychology, 104, 700–712. https://doi.org/10.1037/a0027268 rohrbaugh, j. w. (2016). ambulatory and non-contact recording methods. in j. t. cacioppo, l. g. tassinary, & g. g. berntson (eds.), handbook of psychophysiology(4th ed., pp. 300–338). cambridge, england: cambridge university press. roorda, d. l., jak, s., zee, m., oort, f. j., & koomen, h. m. y. (2017). affective teacher–student relationships and students’ engagement and achievement: a meta-analytic update and test of the mediating role of engagement. school psychology review, 46, 239–261. https://doi.org/10.17105/spr-2017-0035.v46-3 roorda, d. l., koomen, h. m. y., spilt, j. l., & oort, f. j. (2011). the influence of affective teacher-student relationships on students’ school engagement and achievement: a meta-analytic approach. review of educational research, 81, 493–529. https://doi.org/10.3102/0034654311421793 ross, j. m., girard, j. m., wright, a. g. c., beeney, j. e., scott, l. n., hallquist, m. n., … pilkonis, p. a. (2017). momentary patterns of covariation between specific affects and interpersonal behavior: linking relationship science and personality assessment. psychological assessment, 29, 123–134. https://doi.org/10.1037/pas0000338 sack, m., hopper, j. w., & lamprecht, f. (2004). low respiratory sinus arrhythmia and prolonged psychophysiological arousal in posttraumatic stress disorder: heart rate dynamics and individual differences in arousal regulation. biological psychiatry, 55, 284–290. https://doi.org/10.1016/s0006-3223(03)00677-2 sadler, p., ethier, n., gunn, g. r., duong, d., & woody, e. (2009). are we on the same wavelength? interpersonal complementarity as shared cyclical patterns during interactions. journal of personality and social psychology, 97, 1005–1020. https://doi.org/10.1037/a0016232 sapolsky, r. m., romero, l. m., & munck, a. u. (2000). how do glucocorticoids influence stress responses? integrative permissive, suppresive, stimulatory, and preparative actions. endocrine reviews, 21, 55–89. https://doi.org/10.1210/er.21.1.55 schäfer, a., & vagedes, j. (2013). how accurate is pulse rate variability as an estimate of heart rate variability?: a review on studies comparing photoplethysmographic technology with an electrocardiogram. international journal of cardiology,166, 15–29. https://doi.org/10.1016/j.ijcard.2012.03.119 schaufeli, w. b., & van dierendonck, d. (2000). utrechtse burnout scale (ubos). testhandleiding. [utrecht burnout scale. test manual]. amsterdam, the netherlands: harcourt test services. scherer, k. r. (1999). appraisal theory. in t. dalgleish & m. j. power (eds.), handbook of cognition and emotion(pp. 637–663). chichester, uk: john wiley & sons. https://doi.org/10.1002/0470013494.ch30 scherer, k. r. (2009). the dynamic architecture of emotion: evidence for the component process model. cognition & emotion, 23, 1307–1351. https://doi.org/10.1080/02699930902928969 schultzberg, m., & muthén, b. (2018). number of subjects and time points needed for multilevel time-series analysis: a simulation study of dynamic structural equation modeling. structural equation modeling , 25, 495–515. https://doi.org/10.1080/10705511.2017.1392862 scollon, c. n., kim-prieto, c., & diener, e. (2009). experience sampling: promises and pitfalls, strenghts and weaknesses. in e. diener (ed.), assessing well-being: the collected works of ed diener(pp. 157–180). social indicators research series. https://doi.org/10.1007/978-90-481-2354-4 sharma, n., & gedeon, t. (2012). objective measures, sensors and computational techniques for stress recognition and classification: a survey. computer methods and programs in biomedicine, 108 , 1287–1301. https://doi.org/10.1016/j.cmpb.2012.07.003 spilt, j. l., koomen, h. m. y., & thijs, j. t. (2011). teacher wellbeing: the importance of teacher-student relationships. educational psychology review, 23, 457–477. https://doi.org/10.1007/s10648-011-9170-y sutton, r. e., & wheatley, k. f. (2003). teachers’ emotion and teaching: a review of the literature and directions for future research. educational psychology review,15, 327–358. https://doi.org/10.1023/a:1026131715856 thijs, j., koomen, h., roorda, d., & ten hagen, j. (2011). explaining teacher–student interactions in early childhood: an interpersonal theoretical approach. journal of applied developmental psychology, 32, 34–43. https://doi.org/10.1016/j.appdev.2010.10.002 trull, t. j., & ebner-priemer, u. (2013). ambulatory assessment. annual review of clinical psychology, 9, 151–176. https://doi.org/10.1146/annurev-clinpsy-050212-185510 trull, t. j., lane, s. p., koval, p., & ebner-priemer, u. w. (2015). affective dynamics in psychopathology. emotion review, 7, 355–361. https://doi.org/10.1177/1754073915590617 van den bergh, m., & vermunt, j. k. (2017). building latent class growth trees. structural equation modeling, 25(3), 1–12. https://doi.org/10.1080/10705511.2017.1389610 van dijk, a. e., van lien, r., van eijsden, m., gemke, r. j. b. j., vrijkotte, t. g. m., & de geus, e. j. (2013). measuring cardiac autonomic nervous system (ans) activity in children. journal of visualized experiments, (74), 1–6. https://doi.org/10.3791/50073 van droogenbroeck, f., spruyt, b., & vanroelen, c. (2014). burnout among senior teachers : investigating the role of workload and interpersonal relationships at work. teaching and teacher education, 43, 99–109. https://doi.org/10.1016/j.tate.2014.07.005 van kleef, g. a. (2009). how emotions regulate social life. current directions in psychological science, 18, 184–189. https://doi.org/10.1111/j.1467-8721.2009.01633.x van tartwijk, j., brekelmans, m., wubbels, t., fisher, d. l., & fraser, b. j. (1998). students’ perceptions of teacher interpersonal style: the front of the classroom as the teacher’s stage. teaching and teacher education, 14, 607–617. https://doi.org/10.1016/s0742-051x(98)00011-0 van vondel, s., steenbeek, h., van dijk, m., & van geert, p. (2017). ask, don’t tell; a complex dynamic systems approach to improving science education by focusing on the co-construction of scientific understanding. teaching and teacher education, 63, 243–253. https://doi.org/10.1016/j.tate.2016.12.012 vrijkotte, t. g. m., van doornen, l. j. p., & de geus, e. j. c. (2000). effects of work stress on ambulatory blood pressure, heart rate, and heart rate variability. hypertension,35, 880–886. https://doi.org/10.1161/01.hyp.35.4.880 vrijkotte, t. g. m., van doornen, l. j. p., & de geus, e. j. c. (2004). overcommitment to work is associated with changes in cardiac sympathetic regulation. psychosomatic medicine, 66, 656–663. https://doi.org/10.1097/01.psy.0000138283.65547.78 wang, l. p., hamaker, e., & bergeman, c. s. (2012). investigating inter-individual differences in short-term intra-individual variability. psychological methods, 17, 567–581. https://doi.org/10.1037/a0029317 warner, r. m. (1998). spectral analysis of time-series data. new york, ny: the guilford press. wichers, m., wigman, j. t. w., & myin-germeys, i. (2015). micro-level affect dynamics in psychopathology viewed from complex dynamical system theory.emotion review, 7, 362–367. https://doi.org/10.1177/1754073915590623 wilhelm, f. h., & grossman, p. (2010). emotions beyond the laboratory: theoretical fundaments, study design, and analytic strategies for advanced ambulatory assessment. biological psychology, 84, 552–569. https://doi.org/10.1016/j.biopsycho.2010.01.017 wilhelm, f. h., grossman, p., & müller, m. i. (2012). bridging the gap between the laboratory and the real world: integrative ambulatory psychophysiology. in m. r. mehl & t. s. conner (eds.), handbook of research methods for studying daily life.(pp. 210–234). new york, ny: the guilford press. wilhelm, f. h., pfaltz, m. c., & grossman, p. (2006). continuous electronic data capture of physiology, behavior and experience in real life: towards ecological momentary assessment of emotion. interacting with computers, 18, 171–186. https://doi.org/10.1016/j.intcom.2005.07.001 wilson, g. f. (1992). applied use of cardiac and respiration measures: practical considerations and precautions. biological psychology, 34, 163–178. https://doi.org/10.1016/0301-0511(92)90014-l wu, z., huang, n. e., long, s. r., & peng, c.-k. (2007). on the trend, detrending, and variability of nonlinear and nonstationary time series. proceedings of the national academy of sciences, 104, 14889–14894. https://doi.org/10.1073/pnas.0701020104 wubbels, t., brekelmans, m., den brok, p., levy, j., mainhard, m. t., & van tartwijk, j. (2012). let’s make things better. in t. wubbels, p. den brok, j. van tartwijk, & j. levy (eds.), interpersonal relationships in education: an overview of contemporary research (pp. 225–250). rotterdam, the netherlands: sense publishers. wubbels, t., brekelmans, m., den brok, p., & van tartwijk, j. (2006). an interpersonal perspective on classroom management in secondary classrooms in the netherlands. in c. evertson & c. weinstein (eds.), handbook of classroom management: research, practice, and contemporary issues (pp. 1161–1191). mahwah, nj: lawrence erlbaum associates. wubbels, t., brekelmans, m., mainhard, m. t., den brok, p., & van tartwijk, j. (2016). teacher-student relationships and student achievement. in k. r. wentzel & g. b. ramani (eds.), handbook of social influence on social-emotional, motivation, and cognitive outcomes in school contexts (pp. 127–142). new york, ny: routledge. wubbels, t., créton, h. a., & hooymayers, h. p. (1985). discipline problems of beginning teachers, interactional teacher behaviour mapped out. abstracted in resources in education, 20, 12, p. 153. eric document ed260040. microsoft word vokatis & zhang_publication.docx           frontline  learning  research  vol.4  no.  1  (2016)  58  -­‐  77   issn  2295-­‐3159       corresponding author: barbara vokatis, department of elementary education and reading, state university of new york at oneonta, ravine parkway 108, oneonta, ny 13820, usa, email: barbara.vokatis@oneonta.edu doi: http://dx.doi.org/10.14786/flr.v4i1.223 the professional identity of three innovative teachers engaging in sustained knowledge building using technology barbara vokatisa, jianwei zhangb astate university of new york at oneonta, united states bstate university of new york at albany, united states article received 30 october / revised 15 february / accepted 16 february / available online 14 april   abstract diffusing inquiry-based pedagogy in schools for deep and lasting change requires teacher transformation and capacity building. this study characterizes the professional identity of three elementary school teachers who have productively engaged in inquirybased classroom practice using knowledge building pedagogy and knowledge forum, a collaborative online environment. grounded theory analysis of teacher interviews, supplemented with field observations, highlights five distinctive features of the teachers’ identity: (a) teachers as professional knowledge builders to explore new visions of teaching for continual improvement of knowledge building; (b) teachers as co-learners to form symmetrical relationships with students so they can take on the highest level of responsibility; (c) teachers as problem-solvers and barrier-breakers holding a proactive stance toward the contexts of practice; (d) teachers as members of a professional community that encourages collaboration, innovation, and continual improvement; and (e) an empowering relationship with the principal who supports teacher innovation and collaboration. keywords: teacher identity; knowledge building; inquiry learning; problem-centred pedagogy; technology vokatis  &  zhang       | f l r     59   1. introduction current education reforms require teachers’ capacity to incorporate authentic inquiry practices by which students construct powerful explanations and designs to address authentic problems (national research council, 2012). while various professional development resources are initiated to help teachers understand inquiry-based teaching strategies and technologies, the knowledge of teaching methods does not suffice to become a responsive and thoughtful inquiry-based educator (fairbanks, duffy, faircloth, ye, levin, rohr, & stein, 2010). the complexity of teaching demands that teachers develop adaptive expertise (bransford, darling-hammond, & lepage, 2005) and capacity to engage in wise and collaborative improvisation in response to students’ evolving thinking and changing needs (little, lampert, graziani, borko, clark, & wong, 2007; sawyer, 2004). therefore, good inquiry-based teaching cannot be reduced to prescriptive techniques, but comes from the identity of the teacher as a whole person (palmer, 1997). this aspect of teaching gives educators a sense of their selfhood as dynamically connected with their students and subject areas, allowing them to have a clear vision of themselves and what is important for them to accomplish with children (duffy, 2005), and to have the soul and agency to overcome obstacles as they constantly look for ways to respond to the needs and thinking of students (fairbanks et al., 2010). the goal of this study is to provide a detailed account of the professional identity of three elementary school teachers who have been working persistently and productively with knowledge building pedagogy and technology, one of the most influential computer-supported collaborative inquiry programs to cultivate creative knowledge work among students (scardamalia & bereiter, 2006). our conceptual framework guiding this study is two-fold, focusing on teacher identity and knowledge building, respectively. 1.1 conceptualizing teacher identity the concept of teacher identity refers to how teachers identify themselves as teachers, including who they are as professionals, and who they strive and are empowered to become in a constant process of reflecting on their practices and experiences. teacher identity is not a static entity; a teacher constantly constructs and develops a reflective sense of self through looking into his or her practice and life of teaching, as a mirror (palmer, 1997). teachers teach who they are (clandinin & huber, 2005; palmer, 1997); a teacher’s identity is associated with his or her distinctive set of practices (gee, 2001), such as inquiry-based teaching. in this sense, teacher identity is intertwined with teacher practices (enyedy, goldberg, & welsh, 2005). teachers’ professional identity arises out of their various types of teaching practices across contexts in which they construct holistic views of themselves in relation to students, colleagues, professional purposes, and circumstances of teaching (beijaard, meijer, & verloop, 2004; dillabough, 1999; olsen, 2008). in this sense, teacher identity differs from teachers’ specific practices and functional roles: their roles are associated with specific jobs and skills of teaching while teacher identity is a more personal entity that indicates how one identifies himself or herself as a teacher (mayer, 1999). our conceptualization of identity draws heavily on the work of gee (2001) and other researchers who highlight a number of important characteristics of teacher identity (beijaard et al., 2004; connelly & clandinin, 2000; rodgers & scott, 2008; sfard & prusak, 2005). first, teacher identity is a constantly undergoing process in which a person interprets and reinterprets oneself as a certain kind of person and is also recognized as a certain kind of person in a particular context (gee, 2001). it is not limited to answering the question “who am i at this moment?”, but also entails answering the question: “who do i want to become?” (beijaard et al., 2004). thus, teachers need to constantly explore and reflect on who they are as professionals based on their experiences (antonek, mccormick, & donato, 1997; brooke, 1994) and actively look for new ways to define their professional work to approach important educational issues (coldron & smith, 1999). they frame and develop who they are through reflective story telling about what they strive for and do as teachers: “stories to live by” that are shaped by the past and project into their ongoing lives and works (connelly & clandinin, 2000). vokatis  &  zhang       | f l r     60   as its second feature, teacher identity is shaped by multiple contexts of teaching practices (beijaard et al., 2004; rodgers & scott, 2008). contexts entail larger socio-cultural-historical processes that influence teachers’ identity (varelas, house, & wenzel, 2005), personal histories that alter teachers’ beliefs and values, the culture of the institution, including the history of the institution, and values held by its administrators and other members. through reflecting on their practice and identity, teachers “become more in tune with their sense of self and with a deep understanding of how this self fits into a larger context which involves others” (beauchamp & thomas, 2009, p. 182). therefore, teacher identity involves sub-identities that are reflected in their relationships with peer teachers, students, administrators, and other members of their school communities (beijaard et al., 2004). examining a teacher’s identity requires understanding how the teacher forms certain relationships with his or her students, peer teachers, and school administrators and positions himself or herself toward the context such as the curriculum, school policy, and physical school environment. social relationships are crucial to identity, because to have an identity one must be recognized as a particular “kind of person” by others (gee, 2001). how a teacher identifies himself or herself stems from the nature of social interactions the teacher has with his or her peers and others (dillabough, 1999). within multiple contexts, a teacher forms multiple relationships that bring forth multiple aspects of himself or herself (gee & crawford, 1998; rodgers & scott, 2008). teacher identity is co-constructed “through engagement with others in cultural practice” (smagorinsky, cook, moore, jackson, & fry, 2004, p. 21). a teacher’s identity influences how he or she negotiates his or her role in relation to administration, curriculum and students, and these relationships further influence the teacher’s identity (enyedy et al., 2005). in an autographic study, brooke (1994) also found that becoming a professional teacher involves interacting with others’ views. the conflict between one’s own images of teaching and peers’ expectations of what makes a professional teacher may lead to deep reflection on identity (volkmann & anderson, 1998). a final, critical feature of teachers’ identity pertains to their agency and voice to shape their own professional paths (rodger & scott, 2008). teachers’ professional identity develops as a result of the negotiation between given factors, such as existing social structures and policies, and teachers’ agencydriven participation and engagement with educational resources and ideas (coldron & smith, 1999). agency is the empowerment to act (holland, lachicotte, skinner, & caine, 1998). a teacher with agency not only knows how to act within the existing world of education but also acts upon and remakes the world in line with his or her vision. such agency results from a teacher’s realization of his or her identity (beauchamp & thomas, 2009; parkinson, 2008) and further drives his or her continual efforts to explore and form new identifies as he or she goes beyond current classroom practices. in this study, we investigated the identity of inquiry-based teachers who have engaged in knowledge building pedagogy and technology as a school-wide innovation implemented over a decade. school-based professional development has been provided to support their ongoing reflection on their knowledge building practices in classrooms as well as what it means to be inquiry-based, knowledge building teachers. this study examines how the teachers reshape their identity through their long-term engagement in knowledge building practices. in line with the above conceptualization, we analyse their reflective stories about what they strive for through their unique set of practices (gee, 2001), supported by their relationships with their students, colleagues, and the principal. 1.2 teacher identity in inquiry-based, knowledge building communities with schools increasingly incorporating problem-centred, inquiry-based pedagogy to develop student productive knowledge and higher-order competencies (e.g. creativity, collaboration, and other 21st century competencies), research on teacher learning and development needs to understand the professional identity of teachers who are dedicated to implementing and sustaining inquiry-based learning in their practice. the literature on inquiry-based, collaborative learning suggests various new roles to be played by the teacher: a designer, facilitator, mentor, modeller of authentic inquiry processes, and a partner or covokatis  &  zhang       | f l r     61   learner who co-engages in the inquiry processes with his or her students, valuing students as collaborative contributors while fostering their ownership and agency (belland, glazewski, & richardson, 2008; brush & saye, 2000; crawford, 2000; hjalmarson & diefes-dux, 2008; hmelo-silver & burrows, 2006; lunn & solomon, 2000; mills, 2014; tabak & baumgartner, 2004; zhang & sun, 2011; zhang, hong, scardamalia, & morley, 2011). however, existing research on teacher learning to support inquiry-based classroom innovation primarily focuses on teacher knowledge and practices (see fishman, davis, & chan, 2014 for a review), with scarce research efforts centring on the professional identity of teachers who are innovative, persistent, and productive in implementing inquiry-based learning (enyedy et al., 2005). teacher identity aligned with inquiry-based pedagogy allows teachers to pursue and persist in implementing adaptive, responsive teaching (duffy, 2005; fairbanks et al., 2010) and continually inventing new and more productive practices (davis, 2006). therefore, in order to better understand teacher identity in support of inquiry-based pedagogy, the present study examines the professional identity of teachers who are deeply engaged in continuous classroom innovation using knowledge building pedagogy and technology (scardamalia & bereiter, 2006). the knowledge building pedagogy belongs to the larger family of problem-centered, inquiry-based learning programs, with a particular focus on designing inquiry following authentic knowledge creation processes (scardamalia & bereiter, 2006). beyond project-based inquiry that requires students to address pre-defined problems and tasks, knowledge building pedagogy approaches inquiry as progressive problem solving achieved by sustained collective discourse: students identify new and deeper problems as old ones are addressed, driving sustained advancement of collective understandings. students work as a knowledge building community to engage in such sustained inquiry and discourse and collectively advance the “state of the art” of their community’s collective knowledge (scardamalia & bereiter, 2006). they identify deepening problems of understanding, develop and contribute ideas to a public space, engage in collaborative discourse and experimentation, and use a wide variety of resources to advance their ideas. a networked knowledge building environment—knowledge forum, formerly known as csile (computer-supported intentional learning environment)—has been developed to support knowledge building discourse and processes (see scardamalia & bereiter, 2006). knowledge forum provides a collective knowledge space that gives student ideas a public, permanent representation. students contribute diverse ideas to ongoing conversations and collectively advance the ideas through constructive criticisms, mutual build-on, and progressive problem solving, with new and deeper challenges identified as their understanding is advanced (bereiter, 2002). specifically, students record ideas in views (workspaces). these workspaces correspond with their focal goals. students write notes in these views in order to contribute their ideas, data, and related information using text and graphics. knowledge forum has supportive features for knowledge-building discourse that allows students to co-author, build on, and annotate notes. students can also create reference links with citations to existing notes, as well as add keywords and create rise-above notes to summarize and advance their discussions (scardamalia, 2004). knowledge forum scaffolds additionally support both individual contributions and learning as well as collaboration, turning over to students higher-level knowledge processes. customizable scaffolds are designed to support various knowledge processes, such as using the sentence starters “my theory,” “i need to understand,” “this theory cannot explain,” “a better theory,” and “putting our knowledge together” to support theory development (scardamalia, 2004). knowledge building is a dynamic, social activity system in which students interact with diverse people and ideas to advance their collective knowledge. the social and cognitive complexity of this process requires a principle-based, adaptive approach to classroom design and practice, which differs from procedure-based inquiry designs that require students and their teacher to work on pre-defined project tasks following pre-scripted procedures and timelines (zhang et al., 2011). knowledge building in classrooms is guided by a set of 12 knowledge building principles, including epistemic agency, real ideas and authentic problems, continual idea improvement, collective responsibility for community knowledge, knowledge building discourse, and constructive use of authoritative sources (scardamalia, 2002). epistemic agency allows students to set goals for learning, initiate and sustain knowledge advancement, and engage in higherlevel knowledge work normally left to the teacher. the principle of real ideas and authentic problems allows vokatis  &  zhang       | f l r     62   students to identify problems that stem from their curiosity and efforts to understand the world. the principle of continual idea improvement treats ideas as ever improvable, not simply rejected or accepted. the principle of collective responsibility for community knowledge places a responsibility on all participants for contributing to community goals and advancing community knowledge, not only individual learning. the principle of knowledge building discourse asks students to engage in discursive practices whose goals are not only to share, but also to transform and advance knowledge. the idea behind the principle of constructive use of authoritative sources is accessing and critically evaluating sources of information. students use these sources to support and refine their ideas, not just to find “the answer.” the 12 principles and corresponding support in knowledge forum create affordances for knowledge building in a community. in a knowledge building initiative focusing on a deep curriculum area(s), teachers and their students co-construct goals of inquiry based on progressive questions from students, and design knowledge building activities in light of the principles. resources and supports are in place for teachers to share lesson examples and reflect on knowledge building processes using real-time analytic data about idea contributions and social interactions. teachers have utilized the automated feedback generated by the automated tools for the purpose of facilitating reflection and improving practice. this principle-based approach gives teachers a high-level ownership over classroom practice and innovation, so they can continually improve and adapt classroom designs and pedagogical understandings to enable increasingly productive knowledge building experiences among their students (chan, 2011; zhang et al., 2011). this study examines the professional identity of a group of teachers from an elementary school that has been implementing knowledge building pedagogy and technology for more than a decade. our previous analysis based on rich data collection over eight years demonstrated the teachers’ continual improvement of inquiry practice as reflected in student active and collaborative engagement in knowledge building (zhang et al., 2011). enabling such sustained innovation and improvement, the teachers formed into a professional knowledge building community themselves to discuss advances and challenges, co-design and test classroom designs, reflect on their practice based on data collected, and continually deepen their understanding of knowledge building principles that inform new possibilities of improvement (zhang et al., 2011). the purpose of this study is to investigate the professional identity of these teachers in the context of knowledge building classrooms. our research question asks: what characterizes the professional identity of these teachers who are dedicated to and capable of sustained innovation using knowledge building pedagogy and technology? 2. method this case study was conducted as a part of a larger research initiative to examine the enactment of knowledge building as a principle-based innovation at an elementary school over a decade: the dr. eric jackman institute of child study laboratory school in toronto (zhang et al., 2011). the school was established in 1926, partly inspired by the work of john dewey. it enrols approximately 200 students from nursery (pre-k), junior kindergarten, senior kindergarten, to grade 6, with 22 students on average per class. most families come from a middle class background and pay a tuition fee. as a laboratory school, jackman ics has been involved in initiating and disseminating new ideas related to improving education. it makes daily contributions to teacher training, providing internship opportunities for graduate students in the programs of child development and education. knowledge building pedagogy and csile/knowledge forum were first introduced in 1994, tested by a few classrooms between 1996-2000, and adopted across the entire school since 2000. for the larger study, we analysed the knowledge building initiatives facilitated by the teachers over eight years focused on core scientific themes as well as social topics. the analysis of student online discourse in knowledge forum demonstrated increasing levels of collaborative knowledge advancement vokatis  &  zhang       | f l r     63   associated with years of teachers’ experience. teacher interviews, reflection journals, and on-site observations helped to elaborate the teachers’ efforts and school conditions. qualitative analysis revealed that the teachers continually and collaboratively worked on improving their practice and deepening their understanding of knowledge building pedagogy. through experimenting with new ideas and openly sharing them with other teachers, they developed adaptive expertise (crawford, schlager, toyama, riel, & vahey, 2005; zhang et al., 2011). the present study re-analysed the interviews with three teachers to understand their professional identity. these teachers were chosen because they represented different grade levels, had the most extensive experience with knowledge building pedagogy at their school, and were often requested to provide mentoring support to other teachers from the international network of knowledge building communities. the teachers included raphael, a male teacher teaching grades 4-6; zanna, a female teacher teaching grades 2-3; and cadence, a female teacher teaching kindergarten (these are pseudonyms). all the teachers were middle-aged and had over five years of experience in teaching. each interview took approximately 30 minutes, focusing on how the teachers approached and improved their classroom practices to better support knowledge building. example interview questions included: how do you see your role as a teacher? what are the three most important qualities you would like to develop in your students? what are the major things you do to develop these qualities? what have been your three most important improvements in your teaching in the past years? in what way do you see your colleagues/principal as supportive of your efforts for seeking innovation in teaching? the analysis of identity was mainly based on the interviews in which the teachers reflect on who they are and what they strive for. the analysis was contextualized by observational records of the teachers’ classroom practices, which were systematically analysed for the larger project. the interviews were fully transcribed, and then analysed using a grounded theory approach (strauss & corbin, 1998). the grounded theory approach suits the subject of this research because, although teacher professional identity has been investigated in the literature in light of theories of identity, inquiry-based teachers with high levels of innovativeness have never been examined for this purpose. we also argue that incorporating already existing ideas from literature regarding what constitutes teacher identity not only does not “compromise methodological ‘purity’” (dunne, 2011, p. 113), but “can actually enhance rigor” (p. 113). the literature review provides a clear rationale for the study and a specific research approach (coyne & cowley, 2006; mcghee, marland, & atkinson, 2007). secondly, it is helpful in contextualizing the study (mccann & clark, 2003), provides the researcher with an aim (urquhart, 2007), and makes known how the phenomenon has been researched (denzin, 2002; mcmenamin, 2006). thirdly, it is helpful in developing ‘sensitising concepts’ (coffey & atkinson, 1996; mccann & clark, 2003) and promoting “clarity in thinking about concepts and possible theory development” (henwood & pidgeon, 2006, p. 350). following procedures of grounded theory analysis, the first author first read and re-read the interview transcriptions, and created open codes that reflect specific features of the teachers’ identity. these codes were then categorized into primary themes that include subthemes to capture prominent features of the teachers’ professional identity in connection with their knowledge building practice. the creation of the themes was informed by teacher identity and knowledge building as our two-fold conceptual framework and by the features of teacher identity reviewed in the beginning of this article while remaining open to possible new aspects of teacher identity in the contexts of knowledge building. the two authors then co-reviewed the open codes and initial themes and subthemes and discussed any disagreements. for example, several raw codes such as classroom as a community of researchers, teacher as an authentic co-learner, and faith in students developed into the following subtheme: “teacher as a co-learner: sustaining students-driven inquiry through symmetrical teacher-student interactions in which the teacher is not an intellectual authority but a co-learner.” this subtheme became then the core aspect of the following theme: “teachers as co-learners: forming symmetrical relationships with students so they can take on the highest level of responsibility for learning and knowledge advancement.” in grounded theory analysis, the degree of agreement between researchers is not as important as “the content of disagreements and the insights that discussion can provide for refining coding frames” (barbour, 2001). in addition, we employed a reflexive approach to the analysis to ensure reliability and validity (barry, vokatis  &  zhang       | f l r     64   britten, barber, bradley, & stevenson, 1999). schwandt (1997) specifies reflexivity as two-fold. the first aspect involves being part of the setting, context, and a phenomenon being researched. the second is “[a] process of self reflection of one’s biases, theoretical predispositions, preferences and so forth” (p. 135). we used reflexivity “to move us outward to achieve an expansion of understanding” (barry, britten, barber, bradley, & stevenson, 1999, p. 30). we were reflexive not only by keeping reflexive diaries and recording analytic decisions in memos, but also by being reflexive about every decision we made (mason, 1996). in addition, as a team, we engaged in group reflexivity, making sure that there was a dialogue between our individual reflexivities and our group reflexivity (barry et al., 1999). in negotiation of our ideas, we developed a dialectic that improved our thinking. that is, by sharing and negotiating our thinking and differences, we thought through our positions and justified them, and if an argument could not be justified, it became apparent that it was weak (barry et al., 1999). the themes and subthemes were then refined and further validated through relating and comparing the themes, checking data against the themes, and triangulating the identified themes with data from the teachers’ journals and field observations. the refined themes and subthemes are elaborated in results. 3. results the data analysis identified five overarching themes—each involving a number of subthemes—that characterize the professional identity of the knowledge building teachers. these themes are summarized in table 1 and elaborated below. table 1 themes and subthemes that characterize the identity of the knowledge building teachers theme subtheme teachers as professional knowledge builders to explore new visions of teaching: viewing teaching as ever improvable to open new possibilities for student knowledge building and development a vision of teaching for lifelong learning and whole child development (e.g. intellectual curiosity, creative problem solving, caring and collective responsibility, open-mindedness) beyond curriculum coverage; a strong belief that pedagogical knowledge and practice need to be continually built and refined to foster increasingly productive knowledge building; an adaptive, open approach to teaching so new classroom arrangements, procedures, and technologies are continually tested and flexibly adapted and integrated in the service of knowledge building and inquiry. teachers as co-learners: forming symmetrical relationships with students so they can take on the highest level of responsibility for learning and knowledge advancement teacher as a co-learner: sustaining students-driven inquiry through symmetrical teacher-student interactions in which the teacher is not an intellectual authority but a co-learner; student agency: honouring students as research team members who are responsible for proposing goals and ideas for research, building and assessing theories, and designing experiments and other activities; collective engagement: respecting and engaging each student as a contributive member of a knowledge building community; students-driven discourse: striving for spontaneous, idea-centred conversations co-improvised by all community members in both faceto-face and online environment—with the teacher as one of them. vokatis  &  zhang       | f l r     65   teachers as problem-solvers and barrierbreakers: holding a proactive stance toward the contexts of practice to address challenges, constraints, and barriers for continual improvement a commitment to developing context-adaptive strategies to make knowledge building possible and effective across age groups and classroom settings; a barrier-breaking attitude to address practical challenges such as time limit and technical problems through flexible and integrated arrangements. teachers as members of a professional community that encourages collaboration, innovation, and continual improvement: building collaborative relationships with colleagues to share, discuss, design, and reflect on innovative classroom practices a shared focus on continual improvement and invention in teaching beyond routine procedures; conviction that improvement is achieved through collaborative efforts; continual, professional knowledge building discourse that supports collaborative problem solving and ideation in teaching; boldness to share and reflect on both successes and failures; confidence in accepting risk-taking as inevitable in experimenting with new approaches; a hybrid identity that integrates practice with research for continual improvement of teaching. an empowering relationship with the principal: perceiving the principal as both a leader and a professional colleague who supports teacher innovation and collaboration democratic, supportive, and professionally centred relationship with the principal as crucial in all teachers’ undertakings; empowerment to innovate resulted from the relationship in which collaborative experimentation and risk-taking are valued and encouraged. 3.1 teachers as professional knowledge builders to explore new visions of teaching: viewing teaching as ever improvable to open new possibilities for student knowledge building and development as a crucial aspect of who they are and mean to achieve, the teachers’ comments in the interviews reveal a perception of themselves as professional knowledge builders who are committed to explore innovative visions of teaching to open new possibilities for student knowledge building and development. first, the focal teachers see themselves as teachers for whole child development and lifelong learning, not just to cover the curriculum. specifically, they are committed to developing crucial qualities that go beyond curriculum content coverage, such as curiosity, intellectual thinking skills, creative problem solving, social caring, collective responsibility, and open-mindedness. they stress that these qualities are important for student development and further needed for students to engage in productive knowledge building. developing curiosity is one of the most important qualities that constitute who they are as young children’s educators. raphael (teaching grades 4-6) recognizes the limitation of the curriculum in stimulating children’s natural curiosity that drives inquiry learning; thus, in his classrooms, he particularly encourages students to ask deeper and deeper questions and engage in curiosity-driven inquiry. similarly, cadence and zanna (grades 2-3) mention the importance of exploring questions that are asked by children and of their interest. specifically, they elaborate that this way of teaching, contextualized and situated in children’s lives, is aligned with children’s needs and curiosity, leading to active engagement and deep exploration, which is the essence of inquiry. as zanna says, she wants to instil “a love of learning so that when they come to school they don’t see the work of school as being just for school but they see that it is important for their life.” at the same time, zanna underscores that this attitude to teaching is a part of who she always was: “i really tried to hold onto it in the public school system, but to come back here and find everyone like-minded it has just brought me right back to the way children learn best, developmentally appropriate practice.” curiosity, according to the teachers, as a tenet that stimulates desires to investigate something in depth, is vital in problem solving and nurturing an inquiring mind. teaching how to solve problems also starts very early. it is in kindergarten where children learn that problems can be solved with their efforts and vokatis  &  zhang       | f l r     66   that there are specific words and strategies that can help them. when problems are brought to the group, through conversation children find out how to solve them. cadence, the kindergarten teacher, stresses, “we talk a lot. we bring problems of understanding, social issues to the group... no matter what kind of problem it is, you can solve something piece by piece... you can bring it to the community and see that you aren't just addressing the small problem that came up or that conflict, you are actually figuring out how to figure out all conflicts.” zanna adds, “we have a class meeting every friday where we sit in a circle and we talk about the week and one thing that kids can do is to raise up a problem they have.” at the same time, raphael underscores that major knowledge advancements may not be achieved all the time as he would hope, but at least he observes “the desire to go deeper” among his students as a result of their engagement in the inquiry activities. developing independence in thinking is another important part of what the teachers mean to achieve. cadence mentions developing independence and open-mindedness in thinking very early in children’s education: “i want them to be knowing that they can act independently; they don't need to have a teacher there, guiding them in the whole way, and telling them what they're doing is right or wrong.” for zanna, developing independence is also important: “i want the kids to rely on each other so they don't feel they need to come to me. i don't want to be centre of the class.” raphael stresses that he wants to instill in children a disposition of “seeking the information as opposed to children thinking that things have to come to them either from a teacher or a book.” aligned with their commitment to growing care, curiosity, and independent thinking in their students, the teachers focus their role on creating a community of knowledge builders who share collective responsibility. the teachers stress collective responsibility that allows students to keep moving forward in their common pursuit of knowledge and hold each other accountable without the teacher stepping in all the time. cadence mentions that her kindergarteners already learn that being a part of a community entails certain behaviours, such as respecting everybody as a member of the community and valuing everyone’s ideas. zanna adds, “i want them to take responsibility for their actions, not to blame others or to say ‘oh i didn’t do it’ but to make good choices and when they don’t make good choices to admit to it.” raphael gives a specific example of such collective responsibility, “…the children will say ‘you’ve been researching for two days and you haven’t written anything on the database yet, we need to know what you’ve done. we’ve given you this time and you have to give us back some information’ and so that completely changes the nature of the community it is very responsible and works as a unit and where they are leading it themselves.” to explore and achieve their vision of teaching, the teachers embark on a sustained, reflective journey to explore and build new knowledge about their profession. in the interviews, the teachers comment that they constantly rediscover what it means to be a knowledge building teacher. their understanding of knowledge building pedagogy has been constantly evolving over time. in this process, they are all learners. raphael made a comment that encapsulates this critical belief in rethinking, refinement, and improvement: “but five years from now, we can come back and say: ‘i don't know what i was talking about then, and this feels like knowledge building!’ so there is constant improvement. just like ideas are improvable, the process of knowledge building is improvable.... you're constantly going deeper in what this means... none of us is the learned. we're all learners.” cadence adds: “i never try to think that worked really well, i'm going to do the same thing again. i always look for ways to improve my practice.” the teachers’ comments on new advances they have made in their classrooms demonstrate their adaptive, open approach to teaching by which new classroom arrangements, procedures, and technologies are continually tested, flexibly adapted, and integrated to serve and strengthen knowledge building and inquiry. what these teachers have learned and experienced in this process further strengthens their identity as knowledge building teachers. for zanna, continual testing of adaptive approaches is vital to her professional identity. she has experimented with various strategies to support student knowledge building discourse both in face-to-face interactions and in knowledge forum and found out that giving children the opportunity of recording their ideas in the online space creates the possibility to revisit the ideas later for further vokatis  &  zhang       | f l r     67   exploration. she expresses her thoughts in this regard, “...because we have the software, the questions live there and they do get answered.” reflecting on his experimentations in the classroom, raphael comments on the changes he has made to develop dynamic collaboration structures for knowledge building over three years. he began with collaboration in fixed small-groups in the first year, evolved to collaboration in interacting groups, and eventually to opportunistic collaboration among students based on emergent needs without fixed smallgroups. analysis of the online discourse showed that increased connectivity and productivity among his students resulted from these changes (see zhang, scardamalia, reeve, & messina, 2009; zhang & messina, 2010). raphael has also changed the way classroom conversations are organized; instead of scheduling them in advance, he decided to allow the conversations to emerge naturally, when children felt that they had something important to discuss. similarly, the use of technology in his classroom has gone from teacherdirected tasks to trusting children and allowing them to decide if an idea or question is suitable and important for discussion, online or face-to-face. as all the teachers stress, such efforts to continually improve and innovate teaching is an important characteristic of who they are. as new notions and strategies of teaching continue to develop, the teachers further redefine their relationship with students in the classroom. 3.2 teachers as co-learners: forming symmetrical relationships with students so they can take on the highest level of responsibility for learning and knowledge advancement the focal teachers identify themselves as co-learners who honour students as research team members for collective knowledge building. the teachers describe themselves as authentic members of the classroom community who co-engage in the knowledge building and problem solving processes with students. raphael stresses his position: “we are a community of researchers in the classroom…i try to do things that i don’t know the answer to so that it becomes an authentic kb [knowledge building] experience for me as well, so that i can say to the students ‘i’m not exactly sure. let’s find out.’” zanna adds, “i don't want to be centre of the class. i want to be another member of the community.” cadence also underscores this distinctive relationship that she creates in her classroom, saying: “you're not always the intellectual authority.” perceiving themselves as co-learners leads to a more symmetrical relationship with students. they honour students as research team members who have the agency and capability of proposing goals and topics for research, building and evaluating theories, designing experiments, and forming collaborative groups. raphael underscores this point with enthusiasm: “imagine if a child feels that from the very beginning they could add by connecting things in an interesting way… they might be adding a new perspective, a new theory.” the teachers trust that children can take on high-level responsibility in the classroom to generate deepening questions and ideas. they communicate this trust to their students during activities and encourage children to ask questions that would direct and deepen their collective inquiry. as zanna notes: “children could put their ideas in the pocket if you want to talk about. so the students have more agency in what's happening in the [knowledge building] talks. it's not me deciding, but they identify ‘we want to talk about this,’ ‘we want to put the view up on the wall.’ whatever they want to do.” the symmetrical relationship with students is reflected in students-driven, open-ended discourse in the classroom and on knowledge forum, which are not pre-scripted by the teacher but co-improvised by all community members—with the teacher as one of them. the teachers comment on the importance of respecting diverse ideas. they model their respect of student ideas in the classroom and further create a community that respects diverse voices from all members. the diverse ideas are treated as the driving force to deepen classroom discussions. cadence underscores her openness to follow children’s deepening questions and ideas and let them explore these questions for sustained inquiry and discourse. she describes how children’s interest in how trees breathe led to a three-month investigation: vokatis  &  zhang       | f l r     68   it was the very first day of school. i thought it would be interesting to do a study of trees. … and i tried to think where it might go…every year in the fall, [students] bring in different colours of leaves, they look at the shapes...i think i would probably be talking about leaves and colours and maybe get to the cells... so the very first day, i started asking kids what they knew about trees. and as they told me about different parts of trees, i drew on a piece of chart paper. so someone said branches...twigs...and then a child said: "lungs." and i just stopped… it's such a clear way that puts me in an interesting position. so i said: "where would i put the lungs?" and she said: "i don't know. they have to breath, don't they? they're alive." and for the next months, we looked into how trees breathe. that's how it caught children's interests in the class. ... and it was amazing to notice that you don't have to have these arbitrary barriers, that you can study so many things: do literacy and drama, deep thinking, and specific experiments... so for me it was a huge moment as a teacher to realize just how much you can blast open the possibilities of depth and time. in the above example, when one of the students mentioned lungs as a part of a tree, the teacher did not just say that trees do not have lungs but treated this as a real and meaningful idea, which has the potential to stimulate investigations of how trees breathe and live. through acknowledging the student’s idea and asking a question “where would i put the lungs?” the teacher helped the students to recognize an authentic problem, leading to a deep inquiry beyond the teacher’s imagination. in addition to their trust in student agency and capability to generate questions and ideas for deep inquiry, the teachers further encourage students to take on high-level responsibilities that are usually enacted by the teacher in traditional classrooms. these include giving input to high-level decisions about what needs to be studied, through what activities, who will do what, when, and how the online discussion space should be structured and used. the teachers all stress that it is important to encourage children to propose specific problems for discussions as opposed to following teacher-set topics and schedules. the example coming from the classroom of raphael is particularly striking. at the beginning, he simply planned and scheduled knowledge building talks—a structure co-developed by the teachers to facilitate interactive discourse focusing on advancement of ideas beyond information sharing (see zhang et al., 2011). then, he hung pockets in the classroom to encourage students to drop a note when they have important problems or knowledge advances to talk about. through this and other changes, the knowledge building talks in his classroom have become much more spontaneous and organic, with continual improvement of ideas as the focus. in our classroom observations, we captured chunks of metacognitive discourse embedded on the ongoing classroom dialogues, which focus on issues such as: are we making progress? what are the areas that need more research? what kinds of information should be recorded in knowledge forum? student input to these questions leads to collective decisions about how the community should focus and refine their knowledge building work in the next phase. the focal teachers see knowledge forum as an enabler for the shift of high-level responsibility to students. for example, zanna values the use of knowledge forum to support sustained discourse and further make students’ questions and progress visible for reflection. the software provides a space where “the questions live there and they do get answered.” she expresses that before the use of knowledge forum, it was hard to trace which questions were answered. with knowledge forum’s scaffolds, marking questions using “i need to understand” and theories using “my theory” and “a better theory,” both teachers and their students can trace progress in addressing progressive questions and find areas that need deeper contributions, assisting collective decision making about unfolding directions and deeper actions. while all three teachers emphasize the importance of a symmetrical relationship with their students, their personal styles vary. zanna prefers to be a quiet speaker in the classroom. raphael emphasizes that it is ok to intervene in a classroom discussion actively when needed, such as to recall and model the rules of contribution. vokatis  &  zhang       | f l r     69   3.3 teachers as problem-solvers and barrier-breakers: holding a proactive stance toward the contexts of practice to address challenges, constraints, and barriers for continual improvement developing innovative practices in line with their visions of teaching require the teachers to face and address a range of challenges and barriers resulting from the contexts, such as time limit and school schedule, subject area limitations, age differences, and technology malfunction. instead of being defeated by the challenges and barriers, the teachers become active problem-solvers and barrier-breakers. while facing various challenges, the teachers have developed adaptive strategies to make knowledge building possible and effective across student age groups and classrooms, and they speak about these efforts as a part of who they are. as they strive to engage students of all ages in knowledge building, they need to develop adaptive ways of addressing the challenge of developmental differences. for younger students, the teachers especially focus on modelling tenets that are foundational for knowledge building, such as developing respect for others and different perspectives in order to adapt knowledge building principles to younger students. cadence provides the support when she models how children should respectfully converse with each other, “i try to model a lot of my expectations for the children. so when we sit on the carpet, i don't sit on the chair...i think it's an important thing for me because i'm at their level... i hope i'm showing what kinds of comments, what kinds of questions have value for the whole group, and that, again, every voice needs to be heard.” such modelling does not need to be as intensive for older students.   but raphael underscores that while he strives to be just a member of the community, he never forgets about his modeling role, “...we are still modelling for children.” another significant challenge the teachers have encountered is how to foster deep inquiry in different subject areas within typical time constraints. the teachers comment on a strategy they have developed to integrate different subjects into a sustained knowledge building initiative that addresses core contents of all the areas for integrated understanding. one of the most striking examples comes from cadence who integrated several subjects under one big topic: trees and how they breathe. she says, “...you can study so many things: do literacy and drama, and deep thinking, and specific experiments, every kinds of learning we want the children to do, you can actually do as one topic, because if it's a good topic, like trees and how they breathe, it is so rich, there're so many directions you can go.” efficiently integrating different subjects allowed the teachers to reallocate the time needed for each subject while further fostering the connected understandings among their students. the intensive use of knowledge forum and other technology tools also requires the teachers to solve emergent problems related to technology use. the teachers comment on their constant experimentations to find meaningful ways to use technology for knowledge building and address issues of unproductive technology use. raphael shared a story that a few of his students once refused to write on knowledge forum. by sitting down to listen to the students’ concerns, he realized that the problem resulted from his procedural use of knowledge forum: students were assigned to write online based on a preset schedule when they might not have deep ideas to contribute. raphael made a change to encourage students to use the technology only when it is necessary, focusing on contributing important ideas instead of simple facts from books. doing so helped to increase student engagement. he then reflected on what this struggle taught him, “so we have to really be careful of how we use the technology, that is not for the sake of technology. it has to be for the sake of knowledge building.” in terms of technological reliability, the teachers also need to learn to solve various technical problems themselves (e.g. internet connection, forgetting passwords) due to the lack of a full time technology support specialist. they treat this challenge as an opportunity to model to children how problems can be solved and create alternative arrangements when technology does not work. raphael stresses his persistent and proactive approach to solving problems with a “strong stomach.” he says: “it puts you in a role where you have to be happy all the time with technology, and that's a lot of work. the children are watching you. they could give up easily, because the frustrations sometimes are huge. so we need always to be able to be flexible... it’s about saying ‘oh ok that’s not working, let’s do this over here....” vokatis  &  zhang       | f l r     70   even at this laboratory school, one that has a supportive context for innovative classroom practices, the teachers experience challenges and struggles that they have to address in order to effectively implement knowledge building in their classrooms. working collaboratively as a community helps them to share and address the challenges with mutual social support. 3.4 teachers as members of a professional community that encourages collaboration, innovation, and continual improvement: building collaborative relationships with colleagues to share, discuss, design, and reflect on innovative classroom practices teachers’ innovative collaboration with colleagues and considering themselves as not only teachers but also researchers is another major aspect of who they are as professionals. the teachers see themselves and their teaching as a part of a professional community of teachers they work with. as zanna says, “...everyone here is so interested in their teaching and improving it.” furthermore, all teachers underscore that forming such a professional group is crucial to innovation and improvement of their teaching. raphael comments: “it creates an environment where you say something and even by just talking about it you are improving your understanding.” cadence adds, “anytime i have an idea, a question and i want to connect with another class or another teacher you pretty much have people who are willing to go ahead and do it.” moreover, conducting professionally oriented discourse at weekly knowledge building meetings, which supports their collaborative problem solving and formation of new ideas, shapes who they are as professionals. at the meetings, they exchange their classroom designs, insights, and challenges, ask questions, and continually develop better understanding and strategies for deeper and more productive knowledge building. raphael stresses, “…each one of these [meetings] has completely changed me, my practice, you can, we all know, you can create a kb [knowledge building] environment and not be a knowledge-builder yourself. and to truly understand you need to be immersed in a kb experience yourself.” they also stress that this community, which strives for excellence in teaching, is open to sharing both successes and failures. such freedom and boldness in terms of talking about both successes and failures comes from the teachers’ shared belief that risk-taking is inevitable in experimenting with new approaches that lead to the improvement of teaching practice. in this community, as zanna underscores, “there is not that sort of pretending that everything is going great. people bring their problems up and admit when things aren’t going well.” these teachers also identify themselves as both teachers and researchers and stress that researching their own practice, as well as working with other researchers, is a critical component of good teaching that promotes innovation, refinement, and change. raphael elaborates on this important connection between teaching and researching, “the researcher part informs the teaching and the teaching informs the researcher part of me.” zanna sums up, “for me research goes along really well with teaching ... good teachers constantly reflect on their teaching and think about how to improve it... it is just a natural part of good teaching.” 3.5 an empowering relationship with the principal: perceiving the principal as both a leader and a professional colleague who supports teacher innovation and collaboration for continual improvement the focal teachers develop a supportive relationship with the principal who is professionally instead of administratively oriented. through weekly knowledge building meetings and informal ongoing interactions, the teachers share with the principal and other colleagues their teaching expertise, ideas and designs, understanding of children’s development and needs, and vision of innovative teaching. the principal participates in the professional dialogues and gives her input. cadence comments on this sharing, vokatis  &  zhang       | f l r     71   “if you have an idea, you present it to her [the principal] and it makes sense to her, she is going to back you up. she may have questions about it and ask you to think about it in a slightly different way that has more value, but she is really going to support it.” zanna underscores that she gains support from the principal to sustain her innovative practices, saying: “she is a fabulous leader and it allows me to teach the way i want to teach, be innovative, and reflect on my practice.” raphael expresses the essence of this relationship, “we are incredibly empowered. we are given a lot of support, but with that comes a huge amount of responsibility as well.” 4. discussion the present study sought to illuminate the professional identity of three teachers who have been working persistently and productively with knowledge building pedagogy and technology. while the existing literature on teachers in inquiry-based settings focuses on investigating teacher practices and strategies to facilitate collaborative inquiry and the interplay with teacher knowledge, beliefs, and goals (see fishman et al., 2014 for a review), this study is the first to examine the new professional identity of teachers who have engaged in sustained knowledge building pedagogy and classroom innovation for multiple years. our interview data captured their reflective story telling about what they do and mean to achieve as teachers (connelly & clandinin, 2000). as the review of literature suggests, teachers constantly construct and refine their reflective sense of self through looking into their practice of teaching (antonek, mccormick, & donato, 1997; brooke, 1994; palmer, 1997). through their long-term engagement in knowledge building pedagogy and technology, as a distinctive set of practices (gee, 2001), the focal teachers in this study develop new understandings of who they are and what it means to be inquiry-based knowledge building teachers. specifically, the analysis elaborates five important distinctive facets of the teachers’ identity that fits into the larger context of practice (beauchamp & thomas, 2009) involving their students, colleagues, and principal. first, beyond routine implementers of teaching, the teachers are professional knowledge builders who explore new visions and possibilities of teaching and test new and adaptive teaching designs for continual improvement. their visions of teaching concentrate on whole child development, including intellectual curiosity, creative problem solving, caring, respect, collective responsibility, and openmindedness. they see knowledge building pedagogy and technology as supporting their visions. since there are no given classroom procedures for achieving these high-order learning outcomes, the teachers have to work as professional knowledge builders to develop and improve specific designs in light of principles of knowledge building. in this undergoing process (gee, 2001), they are empowered to engage in constant interpretation and reinterpretation (beijaard et al., 2004) of who they are as knowledge building teachers. they are deeply aware that who they are as teachers constantly changes because of their strong sense that they need to become educators who help children to engage in increasingly productive knowledge building. this mindset allows the teachers to develop an adaptive approach that demonstrates itself in readiness and openness to change existing classroom arrangements and processes and test new and improved strategies, including new ways to use technology. their identity as vision-directed professional knowledge builders is consistent with and supported by their knowledge building practices, which require high-level dynamics and adaptation in classroom work. the teachers reflect on who they are and what they do in the context of knowledge building pedagogy, which in itself demands that teachers build knowledge about their pedagogy and develop it. a related aspect of the teachers’ identity has to do with how they position themselves in relation to the contextual challenges and constraints of their work (beijaard et al., 2004, gee, 2001; rodgers & scott, 2008; varelas et al., 2005). the literature suggests that teachers often see obstacles and contextual constraints as preventing them from innovation and change. the teachers in this study actively identify and vokatis  &  zhang       | f l r     72   address challenges instead of avoiding them. they identify themselves as problem solvers and barrierbreakers who continuously develop adaptive strategies to make knowledge building possible and productive. they approach obstacles in a proactive way so they can solve the problems with their colleagues and students, transform obstacles into innovative ideas and opportunities, and implement and improve knowledge building under new conditions. such a proactive stance is an important part of the teachers’ identity, allowing them to resolve dilemmas and make decisions to sustain student-centred inquiry (see also, enyedy et al., 2005). another important characteristic of the professional identity that the focal teachers display is reflected in their social relationships with others (beijaard et al., 2004; gee, 2001; sfard & prusak, 2005). through forming relationships with students, peer teachers, and administrators, the teachers brought forth multiple identities (gee, 2001), or aspects of oneself (rodgers & scott, 2008), that are connected to “their performances in society” (gee, 2001, p. 99). first, the teachers’ relationships with students constitute their identity. their understanding of themselves as a certain kind of professionals deeply involves students (beauchamp & thomas, 2009) and forming certain kinds of relationships with them. different from traditional authoritative roles, they identify themselves as co-learners with their students in a community of knowledge builders. this results in a symmetrical relationship with students in which students assume highlevel agency for continually evolving knowledge building. this relationship is aligned with the way the teachers approach classroom discussions in both face-to-face and online settings through knowledge forum, not as teacher-planned conversations but students-driven, spontaneous, and co-improvised conversations driven by students’ authentic questions and ideas. such a symmetrical relationship has been evidenced to some extent in research on inquiry-oriented teachers; however, this symmetry was either not always sustained (enyedy et al., 2005) or was incidental (crawford, 2000; tabak & baumgartner, 2004). yet another essential attribute of teachers’ identity is deeply intertwined with the kind of relationships they form with other teachers and with their principal. their sense of belonging to a professional community, as the teachers express in their reflective narratives, substantially strengthens their bold vision of innovative and adventurous teaching (cohen, 1989). in the collaborative team, devoted to inquiry and improvement of teaching, the teachers also display a hybrid identity (bereiter, 2002) by describing themselves not only as teachers but also researchers who research their own practice, in cooperation with other teachers and researchers, to continually advance their pedagogical insights and strategies. increasingly effective knowledge building practice is what results from this multidimensional identity that involves co-developing better understanding and designs of classroom practices, supporting each other to solve problems and take risks, sharing successes and failures based on formal and informal data collection, and challenging one another in an atmosphere of mutual respect and sharing. supporting their exploration and improvement of classroom practices to facilitate knowledge building, the teachers further develop a democratic and professionally oriented relationship with their principal. underpinning this relationship is a mutual understanding that continual innovation and experimentation are necessary for educational improvement and that teaching needs to be coupled with research. these teachers treat the principal both as a leader who is devoted to the school and as an educator who can always share ideas and expertise and engage in professional conversation with teachers. this type of relationship results in teachers’ seeing themselves as empowered to pursue teaching according to their vision, take risks to experiment with innovative approaches, and collaborate and share with their principal and other colleagues about advances and challenges. this kind of democratic relationship with the principal and its direct connection to how teachers perceive and identify themselves has never been evidenced in literature on teacher identity. these various aspects of teacher identity appear to be deeply connected to one another, depicting a coherent image of the teacher’s self in the contexts of inquiry-based, knowledge building classrooms. the teachers’ identity as vision-driven professional knowledge builders who continually improve classroom practice is supported by their role as problem-solvers and barrier-breakers to address contextual challenges and constraints and by their relationships with their students, peers, and principal. they co-construct their vokatis  &  zhang       | f l r     73   identity through engagement with their students, peers, and principal in transformative cultural practices (smagorinky et al., 2004), which focus on collaborative knowledge building. through co-engaging with their students in knowledge building and reflecting on such experiences, they notice and are impressed by the deep ideas and active thinking of their students, which further reinforce their trust in student potential and agency and help the teachers to envision new possibilities to further engage students’ responsibility through improved classroom designs. through ongoing dialogue at weekly meetings that focus on knowledge building progress, strategies, and challenges, the teachers support and acknowledge one another as professional knowledge builders, problem solvers, and co-learners. these identities are further empowered by the democratic relationship with the principal who facilitates a supportive school culture for sustained innovation (zhang et al., 2011). with these important characteristics of who they are and what they strive for, the teachers are able to be persistent and productive in implementing and improving knowledge building practices and addressing various challenges on an ongoing basis. at the point of this study, the teachers were still searching for effective ways to implement knowledge building in mathematics, drawing on a set of strategies tested. 5. implications in conclusion, this study of the three teachers who have been engaging in knowledge building pedagogy for continual innovation contributes to understanding the new professional identity of inquirybased teachers in computer-supported collaborative classrooms. the teachers’ identity is multifaceted, as vision-driven professional knowledge builders, problem solvers, co-learners with students, and innovative collaborators with colleagues. such identity is co-constructed through sustained engagement in the pedagogical practice of knowledge building that both the teachers and administrators value as beneficial for children’ development as well as the constant improvement of teaching in spite of challenges. it is further shaped and sustained through the symmetrical relationship with students, innovative collaboration with other teachers, and the democratic relationship with the principal. teacher development initiatives to support authentic inquiry practices and educational innovations need to nurture the new aspects of teacher identity identified in this study. the best form of professional development is probably to create collaborative, professional knowledge building communities among teachers in which such important new identities are valued and enacted, as featured in this study. in the field of computer-supported collaborative learning, researchers are developing innovative efforts to create reflective communities and professional networks among teachers so they can develop the capacity to implement collaborative knowledge building among their students (chan, 2011; laferrière, breuleux, allaire, hamel, law, et al., in press). a primary focus of such communities is on engaging teachers in collaborative sharing of pedagogical understandings and co-creation of classroom designs (voogt, laferrière, breuleux, itow, hickey, & mckenney, 2015). in light of the findings of this study, researchers may additionally test systematic efforts to help teachers reflect on and transform their professional identity, including their vision of teaching, stance toward classroom practice, and relationships with their students, colleagues, and administrators. such identity-focused reflection may be designed using a narrative approach (sfard & prusak, 2005) to engage teachers in collaborative story telling about who they are now and who they hope to become professionally, always in relation to their context that includes both the practice and the types of relationships they build with others. comparing the stories among different teachers, including their principal, and reflecting on the stories in relation to the principles of collaborative inquiry and knowledge building may create valuable opportunities for teachers to transform their identity and practices. we are interested to test this possibility in our future studies. as a potential shortcoming of this study, the findings were generated based on the analysis of a small sample of teachers who have been implementing knowledge building as a specific model of inquiry-based vokatis  &  zhang       | f l r     74   pedagogy. the results are limited to understanding teacher identity in the contexts of sustained, open-ended inquiry in which the teachers are not charged to follow a set of tasks and procedural steps designed by researchers and curriculum developers, but to design, improvise, and deepen the inquiry process as it unfolds, based on interactive input from students. future research needs to look at other innovative groups of teachers, at different stages of their career, to see similarities and differences in how they understand and perform their identities, and conduct deeper analyses of teacher identity performed in their classroom practices. keypoints bringing computer-supported collaborative knowledge building into classrooms requires new professional identities of teachers. the new professional identities involve teachers as vision-driven professional knowledge builders; as problem-solvers to address contextual challenges; as co-learners with students; and as innovative collaborators with colleagues. teacher development efforts to support authentic inquiry and knowledge building need to nurture these new aspects of teacher identity. the best form of professional development is probably to create collaborative, professional knowledge building communities among teachers in which such important new identities are valued, enacted, and reflected upon. acknowledgments this research was supported by the national science foundation (iis #1441479). any opinions expressed in this paper are those of the authors and do not necessarily reflect the views of the national science foundation. part of the findings has been presented at the international conference on computer supported collaborative learning (cscl 2015, gothenburg, sweden). the authors would like to thank the teachers, principal, and students of the dr. eric jackman institute of child study of the university of toronto for the insights, accomplishments, and research opportunities enabled by their work. references antonek, j. l., mccormick, d. e., & donato, r. (1997). the student teacher portfolio as autobiography: developing a professional identity. the modern language journal, 81(1), 15–27. doi: 10.2307/329158 barbour, r. (2001). checklists for improving rigour in qualitative research: a case of the tail wagging the dog? british medical journal, 322, 1115–1117. doi: http://dx.doi.org/10.1136/bmj.322.7294.1115 barry, c. a., britten, n., barber, n., bradley, c., & stevenson, f. (1999). using reflexivity to optimize teamwork in qualitative research. qualitative health research, (9)1, 26–44. beauchamp, c., & thomas, l. (2009). understanding teacher identity: an overview of issues in the literature and implications for teacher education. cambridge journal of education, 39(2), 175–189. doi:10.1080/03057640902902252 beijaard, d., meijer, p. c., & verloop, n. (2004). reconsidering research on teachers' professional identity. teaching & teacher education: an international journal of research and studies, 20(2), 107–128. doi:10.1016/j.tate.2003.07.001 vokatis  &  zhang       | f l r     75   belland, b. r., glazewski, k. d., & richardson, j. c. (2008). a scaffolding framework to support the construction of evidence-based arguments among middle school students. educational technology research and development, 56(4), 401–422. doi:10.1007/s11423-007-9074-1 bereiter, c. (2002). education and mind in the knowledge age. mahwah, nj: erlbaum. bransford, j., darling-hammond, l., & lepage, p. (2005). introduction. in l. darling-hammond & j. bransford (eds.), preparing teachers for a changing world: what teachers should learn and be able to do (pp. 1–39). san francisco, ca: jossey-bass. brooke, g. e. (1994). my personal journey toward professionalism. young children, 49(6), 69–71. brush, t., & saye, j. (2000). implementation and evaluation of a student-centered learning unit: a case study. educational technology research and development, 48(3),79–100. chan, c. k. k. (2011). bridging research and practice: implementing and sustaining knowledge building in hong kong classrooms. international journal of computer-supported collaborative learning, 6(2), 147–186. doi:10.1007/s11412-011-9121-0 clandinin, d. j. & huber, m. (2005). shifting stories to live by: interweaving the personal and the professional in teachers’ lives. in d. beijaard, p. meijer, g. morine-dershimer, & h. tillema (eds.), teacher professional development in changing conditions (pp. 43–61). dordrecht: springer. coffey, a., & aktinson, p. (1996). making sense of qualitative data: complementary research strategies. thousand oaks, ca: sage. cohen, d. k. (1989). teaching practice: plus que ca change....in p. w. jackson (ed.), contributing to educational change: perspectives on research and practice (pp. 27–84). berkeley, ca: mccutchan. coldron, j., & smith, r. (1999). active location in teachers’ construction of their professional identities. journal of curriculum studies, 31(6), 711–726. doi:10.1080/002202799182954 connelly, f. m., & clandinin, d. j. (2000). shaping a professional identity: stories of education practice. london, on: althouse press. coyne, i., & cowley, s. (2006). using grounded theory to research parent participation. journal of research in nursing, 11(6), 501–515. doi: 10.1177/1744987106065831 crawford, b. a. (2000). embracing the essence of inquiry: new roles for science teachers. journal of research in science teaching, 37, 916–937. doi: 10.1002/1098-2736(200011)37:9<916::aidtea4>3.0.co;2-2 crawford, v. m., schlager, m., toyama, y., riel, m., & vahey, p. (2005, april). characterizing adaptive expertise in science teaching: report on a laboratory study of teacher reasoning. paper presented at the annual meeting of the american educational research association, montreal, canada. davis, e. (2006). characterizing productive reflection among preservice elementary teachers: seeing what matters. teaching and teacher education, 22, 281–301. doi:10.1016/j.tate.2005.11.005 denzin, n. k. (2002). the interpretive process. in m. huberman & m. b. miles (eds.), the qualitative researcher’s companion (pp. 340–368). thousand oaks, ca: sage. dillabough, j. a. (1999). gender politics and conceptions of the modern teacher: women, identity and professionalism. british journal of sociology of education, 20(3), 373–394. duffy, g. (2005). metacognition and the development of reading teachers. in c. block, s. israel, k. kinnucan-welsch, & k. bauserman (eds.), metacognition and literacy learning (pp. 299–314). mahwah, nj: lawrence erlbaum. dunne, c. (2011). the place of the literature review in grounded theory research. international journal of social research methodology, 14(2), 111–124. doi: 10.1080/13645579.2010.494930 enyedy, n., goldberg, j, & welsh, k. m. (2005). complex dilemmas of identity and practice. science education, 90(1), 68–93. doi 10.1002/sce.20096 fairbanks, c. m., duffy, g. g., faircloth, b. s., ye, h., levin, b., rohr, j., & stein, c. (2010). beyond knowledge: exploring why some teachers are more thoughtfully adaptive than others. journal of teacher education, 61(1/2), 161–171. doi: 10.1177/0022487109347874 fishman, b. j., davis, e.a., & chan, c. k.k. (2014). a learning sciences perspective on teacher learning research. in r. k. sawyer (ed.), cambridge handbook of the learning sciences (2nd ed., pp.750–769). new york: cambridge university press. vokatis  &  zhang       | f l r     76   gee, j. p. (2001). identity as an analytic lens for research in education. in w. g. secada (ed.), review of research in education, vol. 25 (pp. 99–125). washington, dc: american educational research association. gee, j., & crawford, v. (1998). two kinds of teenagers: language, identity, and social class. in d. alvermann, k. hinchman, d. moore, s. phelps, & d. waff (eds.), reconceptualizing the literacies in adolescents’ lives (pp. 225–245). mahwah, nj: erlbaum. henwood, k., & pidgeon, n. (2006). grounded theory. in g. m. breakwell, s. hammond, c. fife-shaw, & j. a. smith (eds.), research methods in psychology (3rd ed., pp. 342–365). thousand oaks, ca: sage. hjalmarson, m. a., & diefes-dux, h. (2008). teacher as designer: a framework for teacher analysis of mathematical model-eliciting activities. interdisciplinary journal of problem-based learning, 2(1), 57–78. doi:10.7771/1541-5015.1051 hmelo-silver, c. e., & burrows, h. s. (2006). goals and strategies of a problem-based learning facilitator. interdisciplinary journal of problem-based learning, 1(1), 21–39. doi:10.7771/1541-5015.1004 holland, d., lachicotte, w., skinner, d., & caine, c. (1998). identity and agency in cultural worlds. cambridge, ma: harvard university press. laferrière, t., breuleux, a., allaire, s., hamel, c., law, n., montané, m., hernandez, o., turcotte, s., & scardamalia, m. (in press). the knowledge building international project (kbip): scaling up professional development for effective uses of collaborative technologies. in c.-k. looi & l. w. teh (eds.), scaling educational innovations. new york: springer. little, j. w., lampert, m., graziani, f., borko, h., clark, k. k., & wong, n. (2007, april). conceptualizing and investigating the practice of facilitation in content-oriented teacher professional development. symposium conducted at the annual meeting of the american educational research association, chicago. lunn, s., & solomon, j. (2000). primary teachers' thinking about the english national curriculum for science: autobiographies, warrants, and autonomy. journal of research in science teaching, 37(10), 1043–1056. doi: 10.1002/1098-2736(200012)37:10<1043::aid-tea2>3.0.co;2-s mason, j. (1996). qualitative researching. london: sage ltd. mayer, d. (1999) building teaching identities: implications for pre-service teacher education. paper presented to the australian association for research in education, melbourne. mccann, t., & clark, e. (2003). grounded theory in nursing research: part 1 – methodology. nurse researcher, 11(2), 7–18. mcghee, g., marland, g. r., & atkinson, j. (2007). grounded theory research: literature reviewing and reflexivity. journal of advanced nursing, 60(3), 334–342. doi: 10.1111/j.1365-2648.2007.04436.x mcmenamin, i. (2006). process and text: teaching students to review the literature. ps: political science and politics, 39(1), 133–135. mills, h. (2014). learning for real. portsmouth, nh: heinemann. national research council (2012). a framework for k-12 science education: practices, crosscutting concepts, and core ideas. washington, dc: the national academies press. olsen, b. (2008). introducing teacher identity and this volume. teacher education quarterly, 35(3), 3–6. palmer, p. j. (1997). the courage to teach: exploring the inner landscape of a teacher’s life. san francisco, ca: jossey-bass publishers. parkison, p. (2008). space for performing teacher identity: through the lens of kafka and hegel. teachers and teaching: theory and practice, 14(1), 51–60. doi:10.1080/13540600701837640 rodgers, c. r., & scott, k. h. (2008). development of the personal self and professional identity in learning to teach. in m. cochran-smith & s. feiman-nemser (eds.), handbook of research in teacher education (pp. 732–755). mahway, nj: lawrence. earlbaum. schwandt, t. (1997). qualitative inquiry: a dictionary of terms. thousand oaks, ca: sage. sfard, a., & prusak, a. (2005). telling identities: in search of an analytic tool for investigating learning as a culturally shaped activity. educational researcher, 34(4), 14–22. doi: 10.3102/0013189x034004014 sawyer, r. (2004). creative teaching: collaborative improvisation. educational leadership, 33, 12–20. doi: 10.3102/0013189x033002012 vokatis  &  zhang       | f l r     77   scardamalia, m., & bereiter, c. (2006). knowledge building: theory, pedagogy, and technology. in r. k. sawyer (ed.), cambridge handbook of the learning sciences (pp. 97–115). new york: cambridge university press. smagorinsky, p., cook, l. s., moore, c., jackson, a.y., & fry, p. g. (2004). tensions in learning to teach: accommodations and the development of a teaching identity. journal of teacher education, 55(1), 8– 24. doi: 10.1177/0022487103260067 strauss, a., & corbin, j. (1998). basics of qualitative research: techniques and procedures for developing grounded theory (2nd ed.). newbury park, ca: sage. tabak, i., & baumgartner, e. (2004). the teacher as partner: exploring participant structures, symmetry, and identity work in scaffolding. cognition and instruction, 22(4), 393–429. urquhart, c. (2007). the evolving nature of grounded theory method: the case of the information systems discipline. in a. bryant & k. charmaz (eds.), the sage handbook of grounded theory (pp. 339–360). london: sage. varelas, m., house, r., & wenzel, s. (2005). beginning teachers immersed into science: scientist and science teacher identities. science education, 89(3), 492–516. doi: 10.1002/sce.20047 volkmann, m. j., & anderson, m. a. (1998). creating professional identity: dilemmas and metaphors of a first-year chemistry teacher. science education, 82(3), 293–310. voogt, j., laferrie`re, t., breuleux, a., itow, r. c., hickey, d. t., & mckenney, s. (2015). collaborative design as a form of professional development. instructional science, 43, 259–282. doi: 10.1002/(sici)1098-237x(199806)82:3<293::aid-sce1>3.0.co;2-7 zhang, j., hong, h.-y., scardamalia, m., teo, c. l., & morley, e. a. (2011). sustaining knowledge building as a principle-based innovation at an elementary school. journal of the learning sciences, 20(2), 262–307. doi: 10.1080/10508406.2011.528317 zhang, j., & messina, r. (2010). collaborative productivity as self-sustaining processes in a grade 4 knowledge building community. in k. gomez, j. radinsky, & l. lyons (eds.), proceedings of the 9th international conference of the learning sciences (pp. 49-56). chicago, il: international society of the learning sciences. zhang, j., scardamalia, m., reeve, r., & messina, r. (2009). designs for collective cognitive responsibility in knowledge building communities. journal of the learning sciences, 18, 7–44. doi:10.1080/10508400802581676 zhang, j, & sun, y. (2011). reading for idea advancement in a grade 4 knowledge building community. instructional science, 39(4), 429–452. doi: 10.1007/s11251-010-9135-4 microsoft word meredith et al_publication.docx           frontline  learning  research  vol.5  no.  2  (2017)  24  -­‐  35   issn  2295-­‐3159       the measurement of collaborative culture in secondary schools: an informal subgroup approach chloé mereditha*, nienke m. moolenaar b, charlotte struyve a, machteld vandecandelaere a, sarah gielen a & eva kyndt a a ku leuven, belgium b utrecht university, the netherlands article received 28 november / revised 17 january / accepted 21 january / available online 28 march abstract research on teacher collaboration underlines the importance of a collaborative culture for teachers’ functioning. however, while scholars usually regard collaborative culture as a school team characteristic, this study argues that subgroups may be more meaningful units of analysis to conceptualize and assess teachers’ perceptions of collaborative culture. based on the assumption that collaborative culture is developed, expressed, and maintained in frequent work-related interactions, this study hypothesizes that collaborative culture is not homogenously spread over the school but rather varies between informal subgroups. data from 760 flemish teachers were examined using social network analysis and consensus analyses. the results provided evidence that perceptions on collaborative culture are more homogeneous within informal subgroups that are characterized by frequent interactions than the entire school team. this finding stresses the importance of assessing the meaningful unit of analysis for collective-level and socially-constructed concepts, such as collaborative culture. moreover, the benefits and potential of a social network approach to identify (socially stable) subunits within the school team are illustrated. keywords: collaborative culture, informal subgroups, social network analysis, secondary schools                                                                                                                           * corresponding author: chloé meredith, faculty of psychology and educational sciences, ku leuven. dekenstraat 2 – bus 3772, 3000 leuven, belgium. chloe.meredith@kuleuven.be doi: http://dx.doi.org/10.14786/flr.v5i2.283 meredith  et  al       | f l r     25   1. introduction several studies have described the benefits of having a collaborative culture in schools. as a result, teachers have repeatedly been advised to move away from traditional norms of isolation and individuality and move towards greater collaboration (louis & kruse, 1995; marks & louis, 1997). teachers are encouraged to discuss problems, offer different viewpoints, leading to constructive collaboration and consensus (barczak, lassk, & mulki, 2010). in line with research in various organizations, hargreaves (1994) and leithwood, leonard, and sharatt (1998) indicated that cultures with characteristics expressed in terms of collegiality and collaboration generally are those types that promote feelings of professional involvement, efficacy and satisfaction. collaborative culture can be regarded as a part of the organizational culture (peterson & beard, 2004, sveiby & simons, 2002). organizational culture has been described as the shared, and often taken for granted, values, norms and practices within an organization (schein, 2010). in previous studies, collaborative culture has often been conceptualized and assessed as a characteristic of the entire organization, in this case the school (e.g., strahan, 2003; waldron & mcleskey, 2010). however, questions can be asked whether the school is always the most meaningful unit to conceptualize and assess collaborative culture. a precondition for and significant characteristic of organizational culture is that values, norms and practices are shared by a significant portion of organizational members (yang, 2007). these shared values, norms, and practices are developed, expressed and maintained in the frequent communication of organizational members (dumay, 2009). however, within a school, especially in larger schools, interactions are often concentrated in subgroups (firestone & pennell, 1993; frank, 1995). as a result, subgroup members may develop their own norms, values and practices concerning collaboration, resulting in subcultures within the school (soeters, 1988). if we want to investigate collaborative culture as a working condition that affects teachers, it is crucial to identify the relevant unit of analysis. in order to make correct inferences, this unit of analysis needs to reflect the actual collaborative culture one works in. this study therefore focuses on the conceptualization and assessment of collaborative culture and investigates whether subgroups are more meaningful units of analysis than the entire school team. as culture is assumed to result from frequent work-related interactions among teachers, subgroups are identified by means of social network analysis. in this way, this study not only contributes to the assessment of socially-constructed concepts, such as collaborative culture, but also offers a different, innovative perspective on the identification of (socially stable) subunits within the school. 2. conceptual framework 2.1. conceptualizing collaborative culture collaborative culture can be defined as the shared values, norms and practices on the matter of teamwork and communication. flores (2004) referred to collaborative cultures in schools as the working relationships, which are spontaneous, voluntary, evolutionary, and development-oriented, wherein the stance of working together becomes part of the personality of the school. collaboration can then be regarded as a process in which teachers come together to discuss, share knowledge, coach each other, reflect on common experiences and build the curriculum together (lieberman, 1990, 1995). by doing this, teachers create a culture of acceptance for mutual support and collaboration (louis, marks, & kruse, 1996), wherein a norm of collegiality becomes a part of the working stance (little, 1982; nias, 2005). this ‘sharedness’, or homogeneity, of values, norms, and practices is developed, expressed, and performed in the day-to-day interactions of individuals working in the same context, facing the same challenges and goals (dumay, 2009; harris, 1994; van maanen & schein, 1979). however, in most schools, frequent interactions with all school team members are practically impossible and as a result, culture is meredith  et  al       | f l r     26   dispersed over different parts (sackmann, 1992). this dispersion within an organization does not necessarily imply that a culture is heterogeneous, but rather that there are subgroups who have their own culture (adkins & caldwell, 2004). previous studies already provided evidence that members of subgroups share beliefs and values, and exhibit similar actions based on these beliefs and values (ashforth & mael, 1989; harris, 1994). moreover, social psychologists and sociologists indicated that individuals perceive and evaluate their organization based on what happens in their close social neighborhood, meaning the people with whom they engage in frequent interactions (lock & crawford, 1999; sackmann, 1992). this is important, especially because culture is manifested in and often measured by the perceptions of individuals (hofstede, 1998). in other words, teachers that frequently interact, will not only develop and maintain shared values, norms and practices concerning collaboration, they will also perceive and evaluate collaboration in the school based on what happens in their social neighborhood. previous studies have already indicated that subgroups can be found within the school team (frank, 1996). frank (1995) argued that subgroups can be regarded as the crucial link between individuals and organizations, and may therefore be more meaningful units to conceptualize and measure collaborative culture. 2.2. assessing collaborative culture: an informal subgroup approach previous studies often focused on formal subgroups to determine the boundaries of subunits within the school (e.g., busher & blease, 2000; visscher & witziers, 2004). this approach is based on the assumption that interactions among teachers follow the boundaries of this formal structure. however, along with the recent rise of interest in social networks in organizations, studies have started focusing on informal subgroups that are characterized by actual social interactions, and do not necessarily follow the same constellations as their formal counterparts (frank, penuel, & krause, 2015). however, based on the assumption that culture is constructed and evaluated in frequent work-related interactions, it is crucial to identify subgroups that are characterized by frequent interactions. therefore, this study adopts an informal subgroup approach in order to conceptualize and measure collaborative culture. an informal subgroup approach takes the work-related social network of members into account to subdivide the school in subgroups (frank, 1995, 1996). as a result, these informal subgroups can be regarded as stable social units within the broader organization (frank & yasumoto, 1998). we therefor expect that these subgroups may be more meaningful units of analysis to assess collaborative culture, in comparison to the whole school team. to empirically substantiate this assumption, the following research hypothesis is tested: perceptions of collaborative culture are more homogenous within informal subgroups compared to on the entire school team’s aggregated perception on collaborative culture. 3. method 3.1. participants data for this study were collected in secondary schools in flanders (northern part of belgium with approximately 6.500.000 inhabitants), using an online survey. the “school team questionnaire” was part of the liso-project (dutch acronym for student careers in secondary education) and was administered in 20 secondary school teams. in this study, a ‘school team’ consisted out of all teachers who actually worked in the same secondary school campus (= in the same geographical unit). all teachers of the 20 secondary schools received an invitation to fill out the online questionnaire that addressed the working conditions in the school. however, we could only include secondary schools that had a response rate of at least 75%, which is the threshold to reliably identify informal subgroups using social network analysis (borgatti, carley, & krackhardt, 2006; kossinets, 2006). in 13 schools, a response rate of 75% or more was reached, resulting in meredith  et  al       | f l r     27   a dataset of 760 teachers. respondents in our dataset were on average around 40.65 years old (sd = 10.62), had 11.65 years of experience (sd = 9.68) in their current school and 64% of them were female. 3.2. instrument and analytic strategy 3.2.1. determining informal subgroups in order to determine informal subgroups within the secondary school, we included a sociometric question in the survey. as we wanted to grasp the day-to-day work-related interactions of teachers, we decided to focus on the information network and asked the following question: “whom do you go to for class-related information? (for instance, for teaching materials and methods, learning, content and classroom management)”. a roster with all teachers’ names was presented and participants could indicate to whom of their colleagues they went to and on what frequency basis (presented in 8 categories going from ‘once a year’ to ‘daily’). teachers could indicate an unlimited number of colleagues. in order to identify informal subgroups, we used a social network approach. social network analysis makes it possible to investigate the patterns of interactions between actors (scott, 1991). it focuses on a set of individuals and the relationships connecting them (wasserman & faust, 1994). several authors indicated that social network analysis helps to unravel social phenomena that can partly be explained by social structure (e.g., wellman & frank, 2001). based on the patterns and frequency of work-related interactions, informal subgroups can be identified. in this study, the approach of frank (1996) was used to determine nonoverlapping cohesive subgroups. a non-overlapping approach, meaning that teachers could only be member of one subgroup, was selected, as an overlapping approach makes it difficult to establish “an inside and an outside” of a social unit (abbott, 1995, p. 872). in order to identify these non-overlapping cohesive subgroups, the kliquefinder software was used (frank, 1996). the procedure implemented in kliquefinder is based on the exponential random graph model (ergm) framework. the algorithm within this software identifies subgroups by iteratively reassigning actors in order to see if the probability that two actors interact increases if they are members of the same subgroup. actors are assigned to a cohesive subgroup once the maximum probability of having a tie with other subgroup members is reached (for more information about the software and algorithm, see frank, 1995). 3.2.2. evaluating homogeneity of collaborative culture in order to test the hypothesis that perceptions on collaborative culture are more homogenous within the informal subgroup than within the entire school team, we included a scale to assess collaborative culture. schein (1985) indicated that actual practices are the most visible and tangible aspects of organizational culture, and underlying assumptions, values and norms come to surface in these practices. we selected six items of the scale of leonard (2002) addressing the perceptions of teachers on collaborative culture. we only included the items that addressed the actual collaboration within the school, such as ‘in my school, teacher collaboration is strong’ and ‘in my school, teaching is a team activity, rather than an individual activity’. the cronbach’s alpha of this subscale was satisfactory (α = .85). based on a confirmatory factor analysis, the fit of the scale was found acceptable (cfi = .954, tli = .968, srmr = .042, rmsea = .04). homogeneity of perceptions was operationalized by combining two complementary approaches to assess within-group agreement, namely 1) a consensus-based approach, using the intraclass correlation coefficient (icc(1)) as a measure of homogeneity; and 2) a disconsensus-based approach, calculating the average deviation index (admj) as a measure of lack of homogeneity. at the same time, these measures give information on the aggregation of individual perceptions to the proposed level of analysis (raudenbush & bryk, 2002). meredith  et  al       | f l r     28   a) a consensus-based approach to homogeneity of perceptions. first, we assessed the icc(1). this icc(1) reflects the agreement between any pair of individuals within the same group (mcgraw & wong, 1996). the outcome is the portion of total variance in a variable that can be explained by membership in a group (raudenbush & bryk, 2002). in other words, a low icc(1) indicates that only a small proportion of the variance can be explained by membership of the informal subgroup, and variance within the subgroup is high. the icc(1) is interpreted as an effect size with values of .01, .10 and .25 respectively indicating a small, median and large effect (bliese, 2000; raudenbush & bryk, 2002). the advantage of using the icc(1) to reflect subgroup homogeneity of perceptions is that it corrects for group size and provides group-level properties that are not biased by either group size or the number of groups in the sample (bliese, 2000). b) a disconsensus-based approach to homogeneity of perceptions. admj is the degree of within-group disconsensus. it reflects the within-group variability, so high admj scores indicate within-group dispersion. the calculation of it consists out of two steps. first the average deviation for each scale item (j) is computed. second, the average deviation for the six items of the scale is calculated. this measure provides an alternative for the highly criticized within-group agreement (rwg) of james, demaree, and wolf (1984), because it does not require an a priori specification of a null response range and it takes the metric of the original response scale into account (gonzález-romá, peiró, & tordera, 2002). both measures were calculated for informal subgroups and school teams to compare the homogeneity of perceptions on collaborative culture in subgroups and schools. these measures reflect to what extent teachers perceive collaborative culture in a similar/dissimilar way. a high degree of correspondence is needed to draw inferences about organizational working conditions, such as culture, and to aggregate perceptual scores to the level of the organization or organizational unit (james & jones, 1974; schneider, 1975). aggregation is interesting because, and only justified when, the mean score goes beyond individual perceptions, and reflects an actual organizational condition (payne, fineman, & wall, 1976). in other words, in order to investigate collaborative culture as a characteristic of the organization or organizational subgroup, perceptions need to reflect sufficient agreement. 4. results 4.1. the identification of informal subgroups based on the analyses in kliquefinder, the 13 schools were found to be consisting of 136 informal subgroups. on average, there were 11.33 subgroups within a school team (min = 5, max = 25, sd = 6.15), containing an average of 6.31 members (sd = 1.99). within these subgroups, the average age was 40.67 years (sd = 6.13), average experience was 11.78 years (sd = 5.36), and averagely 62.35 % (sd = .33) of the subgroup members were female. for all schools, cluster p-values were below .01, indicating that teachers were significantly more likely to interact frequently with colleagues that were part of the same identified subgroup. 4.2. cultural consensus within schools and informal subgroups the results of homogeneity of collaborative culture indicated that the intraclass correlation for the subgroup level (icc(1) = .21) is almost twice as high as for the school level (icc(1) = .11), reflecting lower variance within the informal subgroup than within the school team. this means that agreement is higher meredith  et  al       | f l r     29   within the informal subgroup. within-group disconsensus results showed that within-group dispersion was lower on the subgroup level (admj = .66) in comparison to the school team level (admj = .77). these results provide support for our research hypothesis, namely that perceptions on collaborative culture are more homogenous within the informal subgroup in comparison to the entire school team. in other words, the perceptions on collaborative culture are more similar among teachers within the same informal subgroup compared to teachers the same school, even when correcting for group size. as a high degree of correspondence is necessary to make inferences about organizational features, the informal subgroup seems a more appropriate unit to conceptualize and assess collaborative culture. 5. discussion over the last decades, the importance of a collaborative culture within organizations, such as secondary schools, has been advocated by several researchers. in many cases, collaborative culture is regarded as a school feature, with the whole school team being the meaningful unit of analysis. however, based on the assumption that culture is developed, maintained and evaluated in the day-to-day communication among organizational members, this study hypothesized that the subgroup might be a more meaningful unit of analysis for both the conceptualization and measurement of collaborative culture. to test this hypothesis, a social network approach was adopted to identify stable social units that are characterized by frequent work-related interactions among their members. thereafter, we compared the homogeneity of perceptions of collaborative culture in these informal subgroups and the entire school team. in what follows, several theoretical and methodological implications are discussed and limitations and suggestions for future research are indicated. 5.1. theoretical and methodological implications 5.1.1. collaborative culture as an informal subgroup characteristic the main goal of this study was to address the conceptualization and measurement of collaborative culture in secondary schools. our findings indicated that the perceptions on the school’s collaborative culture are more homogenous within the informal subgroup than the entire school team. in other words, members of an informal subgroup evaluate collaborative culture more similarly than all members of the entire school team. linking these findings to the conceptual framework, two important, interrelated conclusions can be drawn. first, the fact that there is more homogeneity within subgroups suggests that there are, to some extent, collaborative subcultures within secondary school teams. second, our findings provide evidence for earlier conclusions that teachers perceive and evaluate the organization based on what happens in their close social neighborhood (van maele, moolenaar, & daly, 2015). based on this, it can be concluded that informal subgroups are more meaningful units of analysis to conceptualize and assess collaborative culture, in comparison to the entire school team. in secondary schools, it is often impossible to communicate, and consequently collaborate, with all other school team members. this conclusion stresses the importance of identifying the appropriate level of assessment, especially for socially-constructed concepts. this conclusion is not solely relevant for the concept of collaborative culture and the context of secondary schools. the finding that culture is indeed developed, maintained and evaluated in informal subgroups, characterized by frequent interactions, could also be relevant for other collective-level concepts, such as other types of organizational culture, collective efficacy or organizational trust. moreover, the reasoning behind the importance of (informal) subgroups might also applicable in other organizations. research in organizations should therefore be aware that collective-level concepts are, to some extent, meredith  et  al       | f l r     30   socially constructed and, as a result, differ between subgroups within the organization. in line with frank (1995), we argue that subgroups can be regarded as the crucial link between individuals and organizations. 5.1.2. using social network analysis to identify informal subgroups in order to identify subgroups, this study adopted a social network approach. this approach made it possible to identify subgroups that are characterized by frequent informal work-related interactions, and can therefore be considered as stable social units (frank & zhao, 2005). while previous studies relied on the boundaries of formal structure (e.g., departments or teams) to identify subgroups (e.g., ball & lacy, 1984; hargreaves & hargreaves, 2006; siskin, 1991), this study used the work-related information seeking network of teachers. this approach has not only the advantage that it captures meaningful subunits that are characterized by actual interactions. it provides the possibility to apply a uniform approach for all schools in the sample, even if they differ in formal structure or the meaning that is given to it. for instance, two of the schools in our sample indicated that they did not have the usual ‘subject department structure’. four schools indicated that next to subject department structure, they had other formal subunits, such as grade-level teams or working groups. finally, three schools mentioned that the subject department structure was a purely formal constellation that did not reflect the actual interactions and subunits within the school. a social network approach helps to overcome the issues of identifying the (most) meaningful subunit. the interest in social networks has increased exponentially during the last three decades (lazega & snijders, 2016). based on the assumption that ‘relationships matter’, several researchers already gained thorough understanding of the structure and content of teachers’ professional relationships (e.g., coburn & russel, 2008; daly, moolenaar, bolivar, & burke, 2010; moolenaar, sleegers, & daly, 2012). this study provided further insight into the social structure of secondary schools and contributed to the understanding on how social-structural and cultural aspects of secondary schools are intertwined. moreover, the results showed that research focusing on the school team features benefits from a social network approach, as it provides the possibility to identify stable social units within the broader organization. 5.2. limitations and suggestions for future research first, several limitations can be formulated on the identification of informal subgroups. based on theoretical ideas and empirical findings, this study adopted an informal subgroup approach to conceptualize and measure collaborative culture. to identify our informal subgroups, data on the information seeking network of teachers were used. the selection of the criterion ‘seeking information’ was based on the reasoning that interactions that are characterized by ‘seeking work-related information’ are one of the most general work-related interactions. as we were interested in subgroups that were characterized by day-to-day work-related interactions, this criterion seemed to fit our purpose best. in advance, we checked whether this criterion did not already reflect collaborative culture itself and calculated the correlation between subgroup density, namely the extent to which subgroups members are connected by information ties, and perceptions on collaborative culture in the subgroup. our results showed that both measures were not significantly related (r = .13, p > .05). this indicated that subgroups with more information seeking interactions are not necessarily perceived as having a higher collaborative culture. in other words, the criterion of ‘seeking information’ (and the determination of subgroups) is not related to the perception of collaborative culture itself. it seems that more profound interactions are necessary to establish and measure collaborative by means of interactions. future research could include different types of work-related interactions to further look into the identification of informal subgroups and the measurement of collaborative culture by means of social networks. also, to identify our informal subgroups; the non-overlapping cohesive subgroup approach of frank (1995, 1996) was adopted. we chose this approach as it has been applied and validated in a wide array of network studies (e.g., foster-fishman, berkowitz, lounsbury, jacobson, & allen, 2001; frank & yasumoto, 1998; mcfarland, 2001). however, methodologists have developed and employed various other meredith  et  al       | f l r     31   techniques to identify cohesive subgroups based on the extent of interaction between actors (borgatti & everett, 2000). future research could explore and compare other ways of attributing members to subgroups. finally, we treated these subgroups as autonomous entities with their own norms values and beliefs. however, future research could also address the spillover effects of interactions between subgroups, taking the larger social structure of the school into account. the inclusion of both the subgroup and school as units of analysis seems promising to capture important working conditions that affect the functioning of teachers. second, limitations concerning the measurement of collaboration can be identified. the scale is aimed at measuring a school related concepts and the items of the scale are formulated assuming that the school is the reference group. the approach adopted in this study can be justified by the assumption that teachers’ perceptions are based on what happens in their social neighborhood. however, in future research, a more accurate measurement of collaborative culture could use the, formal or informal, subgroup as reference group. this could make it possible to distinguish between school collaborative culture and subgroup collaborative culture. for instance, the research of adkins and caldwell (2004) made a difference between subgroup culture and the organizational culture. moreover, as schein (1985) indicated that actual practices are the most visible and tangible aspects of organizational culture, it would be interesting to measure collaborative culture by means of observing the actual collaborative practices that take place in the school. this would not only lead to higher response rates, but also reduce potential response biases, such as social desirability bias, which can be present in self-report instruments such as online questionnaires (krumpal, 2013). third, it would be interesting to include other features of the subgroup to investigate the presence of collaborative culture. for instance, the research of barczak, lassk, and mulki (2010) found that team trust impacts collaborative culture. in addition, also the consequences of collaborative culture could be further researched, as well as other types of organizational culture. fourth, as in most survey-based research, our data collection was characterized by incomplete data. missing data can be a significant problem in social network analysis because, as noted, individuals are considered interdependent from social ties in their social context (kossinets, 2006). in order to achieve high response rates in each school, several strategies were adopted to convince teachers to fill out the questionnaire. first, teachers were informed about the questionnaire by means of mails, posters and brochures. in this communication, it was emphasized that all responses would be anonymized, that each response was crucial to get more insight in teacher careers and the working conditions in schools, and that general conclusions would be fed back to policy agencies. second, response rates were followed up and several reminders were sent out. finally, school leaders were motivated to convince their team to fill out the questionnaire. if the school team achieved a response rate of more than 90%, an anonymized feedback report was provided. as only schools with a response rate of 75% or more were included, missing data was limited. in studies adopting a classic statistical approach, and where standard sampling is used to draw a representative sample from a population, there are special techniques available to correct parameter estimates for imperfect response rates (rubin & little, 2002). however, in the case of social network analysis, these kind of treatments are questionable, and although methods for dealing with such issues have been proposed, researchers agree that, at this point in time, list wise deletion is the most appropriate approach (robins & alexander, 2004). finally, our results showed that socially constructed working conditions, more specifically collective-level concepts that are developed and sustained in collegial interactions, cannot simple be regarded as organizational-level concepts. through interactions, members of subgroups develop and maintain shared values, norms and practices. at the same time, this social neighborhood defines how teachers perceive the school and its working conditions. future research therefore needs to pay attention to working conditions that may potentially differ between subgroups within the school. when collective-level concepts are at stake, it is crucial to identify the meaningful level of analysis. only in this way, correct interferences about the importance of these working conditions can be made. the main innovation of this paper is the adoption of a social network approach to identify meaningful subunits and measure collaborative meredith  et  al       | f l r     32   culture. we provided an example of how social network analysis can provide more insight and methodological advancements in educational research. keypoints an informal subgroup approach was adopted to determine meaningful subunits within secondary school teams informal subgroups, characterized by frequent work-related interactions, are identified by means of social network analysis perceptions of collaborative culture are more homogenous within informal subgroups compared to on the entire school team’s aggregated perception on collaborative culture. the importance of identifying the meaningful unit of analysis for collective-level and sociallyconstructed concepts is stressed acknowledgments this study was conducted within the framework of the policy research center on study and school careers and was financed by the flemish ministry of education (belgium). the conclusions of the study do not necessarily reflect the views (of (and do not commit) the financing body. references abbott, a. (1995). things of boundaries. social research, 62(4), 857–882. adkins, b., & caldwell, d. (2004). firm or subgroup culture: where does fitting in matter most? journal of organizational behavior, 25(8), 969-978. doi: 10.1002/job.291 ashforth, b. e., & mael, f. (1989). social identity theory and the organization. academy of management review, 14(1), 20-39. doi: 10.5465/amr.1989.4278999 ball, s. j. & lacy, c. (1984) subject disciplines as the opportunity for group action: a measure critique of subject subcultures. in a. hargreaves & p. woods (eds.), classrooms and staffrooms: the sociology of teachers and teaching (pp. 232-244). milton keynes, uk: open university press. barczak, g., lassk, f., & mulki, j. (2010). antecedents of team creativity: an examination of team emotional intelligence, team trust and collaborative culture. creativity and innovation management, 4, 332-345. doi: 10.1111/j.1467-8691.2010.00574.x bliese, p. d. (2000). within-group agreement, non-independence, and reliability: implications for data aggregation and analysis. in k.j. klein & s.w. kozlowski (eds.), multilevel theory, research, and methods in organizations (pp. 349-381). san francisco, ca: jossey-bass. borgatti, s. p., carley, k. m., & krackhardt, d. (2006). on the robustness of centrality measures under conditions of imperfect data. social networks, 28, 124-136. doi: 10.1016/j.socnet.2005.05.001 borgatti, s. p., & everett, m. g. (2000). models of core/periphery structures. social networks, 21(4), 375395. doi: 10.1016/s0378-8733(99)00019-2 busher, h., & blease, d. (2000). growing collegial cultures in subject departments in secondary schools: working with science staff. school leadership & management, 20(1), 99-112. doi: 10.1080/13632430068905 coburn, c. e., & russell, j. l. (2008). district policy and teachers’ social networks. educational evaluation and policy analysis, 30, 203-235. doi: 10.3102/0162373708321829 meredith  et  al       | f l r     33   cohen, s. g., & bailey, d. e. (1997). what makes teams work: group effectiveness research from the shop floor to the executive suite. journal of management, 23(3), 239-290. doi: 10.1016/s01492063(97)90034-9 daly, a. j., moolenaar, n. m., bolivar, j. m., & burke, p. (2010). relationships in reform: the role of teachers' social networks. journal of educational administration, 48, 359-391. doi: 10.1108/09578231011041062 dumay, x. (2009). origins and consequences of schools’ organizational culture for student achievement. educational administration quarterly, 45(4), 523-555. doi: 10.1177/0013161x09335873 firestone, w. a., & pennell, j. r. (1993). teacher commitment, working conditions, and differential incentive policies. review of educational research, 63(4), 489-525. doi: 10.3102/00346543063004489 flores, m. a. (2004). the impact of school culture and leadership on new teachers' learning in the workplace. international journal of leadership in education, 7, 297-318. doi: 10.1080/1360312042000226918 foster-fishman, p. g., berkowitz, s. l., lounsbury, d. w., jacobson, s., & allen, n. a. (2001). building collaborative capacity in community coalitions: a review and integrative framework. american journal of community psychology, 29(2), 241-261. doi: 10.1023/a:1010378613583 frank, k. a. (1995). identifying cohesive subgroups. social networks, 17, 27-56. doi: 10.1016/03788733(94)00247-8 frank, k. a. (1996). mapping interactions within and between cohesive subgroups. social networks, 18, 93119. doi: 10.1016/0378-8733(95)00257-x frank, k. a., penuel, w. r., & krause, a. (2015). what is a “good” social network for policy implementation? the flow of know-­‐‑how for organizational change. journal of policy analysis and management, 34(2), 378-402. doi: 10.1002/pam.21817 frank, k. a., & yasumoto, j. y. (1998). linking action to social structure within a system: social capital within and between subgroups. american journal of sociology, 104, 642-686. doi: 10.1086/210083 frank, k. a., & zhao, y. (2005). subgroups as meso-level entities in the social organization of schools. in l.v. hedges & b. schneider (eds.), social organization of schooling. new york, ny: sage publications. gonzález-romá, v., peiró, j. m., & tordera, n. (2002). an examination of the antecedents and moderator influences of climate strength. journal of applied psychology, 87(3), 465-473. doi: 10.1037//00219010.87.3.465 hargreaves, a. (1994). changing teachers, changing times: teachers' work and culture in the postmodern age. new york, ny: teachers’ college press. hargreaves, d. h., & hargreaves, d. (2006). social relations in a secondary school. london, uk: routledge & kegan paul. doi: 10.4324/9780203001837 harris, s. g. (1994). organizational culture and individual sensemaking: a schema-based perspective. organization science, 5(3), 309-321. doi: 10.1287/orsc.5.3.309 hofstede, g. (1998). attitudes, values and organizational culture: disentangling the concepts. organization studies, 19(3), 477-493. doi: 10.1177/017084069801900305 james, l. r., demaree, r. g., & wolf, g. (1984). estimating within-group interrater reliability with and without response bias. journal of applied psychology, 69(1), 85-98. doi: 10.1037/0021-9010.69.1.85 james, l. r., & jones, a. p. (1974). organizational climate: a review of theory and research. psychological bulletin, 81(12), 1096-1112. doi: 10.1037/h0037511 kossinets, g. (2006). effects of missing data in social networks. social networks, 28, 247-268. doi: 10.1016/j.socnet.2005.07.002 krumpal, i. (2013). determinants of social desirability bias in sensitive surveys: a literature review. quality & quantity, 47(4), 2025-2047. doi: 10.1007/s11135-011-9640-9 lazega, e., & snijders, t.a.b. (2016). multilevel network analysis for the social sciences. cham, ch: springer. doi: 10.1007/978-3-319-24520-1 leithwood, k., leonard, l., & sharratt, l. (1998). conditions fostering organizational learning in schools. educational administration quarterly, 34, 243-276. doi: 10.1177/0013161x98034002005 meredith  et  al       | f l r     34   leonard, l. j. (2002). schools as professional communities: addressing the collaborative challenge. iejll: international electronic journal for leadership in learning, 6(17). lieberman, a. (1990). schools as collaborative cultures: creating the future now. bristol, uk: the falmer press. lieberman, a. (1995). practices that support teacher development: transforming conceptions of professional learning. phi delta kappan, 76, 591-596. little, j. w. (1982). norms of collegiality and experimentation: workplace conditions of school success. american educational research journal, 19, 325-340. doi: 10.3102/00028312019003325 lock, p., & crawford, j. (1999). the relationship between commitment and organizational culture, subculture, leadership style and job satisfaction in organizational change and development. leadership and organizational development journal, 20, 365-373. doi: 10.1108/01437739910302524 louis, k. s., & kruse, s. d. (1995). professionalism and community: perspectives on reforming urban schools. thousand oaks, ca: sage publications ltd. louis, k. s., marks, h. m., & kruse, s. (1996). teachers’ professional community in restructuring schools. american educational research journal, 33, 757-798. doi: 10.3102/00028312033004757 marks, h. m., & louis, k. s. (1997). does teacher empowerment affect the classroom? the implications of teacher empowerment for instructional practice and student academic performance. educational evaluation and policy analysis, 19, 245-275. doi: 10.3102/01623737019003245 mcfarland, d. a. (2001). student resistance: how the formal and informal organization of classrooms facilitate everyday forms of student defiance. american journal of sociology, 107(3), 612-678. doi: 10.1086/338779 mcgraw, k. o., & wong, s. p. (1996). forming inferences about some intraclass correlation coefficients. psychological methods, 1(1), 30-46. doi: 10.1037/1082-989x.1.4.390 moolenaar, n. m., sleegers, p. j., & daly, a. j. (2012). teaming up: linking collaboration networks, collective efficacy, and student achievement. teaching and teacher education, 28, 251-262. doi: 10.1016/j.tate.2011.10.001 nias, j. (2005). why teachers need their colleagues: a developmental perspective. in d. hopkins (ed.), the practice and theory of school improvement (pp. 223-237). dordrecht, nl: springer. doi: 10.1007/14020-4452-6 payne, r. l., fineman, s., & wall, t. d. (1976). organizational climate and job satisfaction: a conceptual synthesis. organizational behavior and human performance, 16(1), 45-62. doi: 10.1016/00305073(76)90006-4 peterson, t. o., & beard, j. w. (2004). workspace technology's impact on individual privacy and team interaction. team performance management: an international journal, 10(7/8), 163-172. doi: 10.1108/13527590410569887 raudenbush, s. w., & bryk, a. s. (2002). hierarchical linear models: applications and data analysis methods (vol. 1). thousand oaks, ca: sage. robins, g., & alexander, m. (2004). small worlds among interlocking directors: network structure and distance in bipartite graphs. computational & mathematical organization theory, 10(1), 69-94. doi: 10.1023/b:cmot.0000032580.12184.c0 rubin, d. b., & little, r. j. (2002). statistical analysis with missing data. new york, ny: john wiley & sons. sackmann, s. a. (1992). culture and subcultures: an analysis of organizational knowledge. administrative science quarterly, 37(1), 140-161. doi: 10.2307/2393536 schein, e. h. (1985). defining organizational culture. classics of organization theory, 3, 490-502. schein, e. h. (2010). organizational culture and leadership (vol. 2). san francisco, ca: john wiley & sons. schneider, b. (1975). organizational climate: individual preferences and organizational realities revisited. journal of applied psychology, 60(4), 459-465. doi: 10.1037/h0076919 scott, j. (1991) social network analysis: a handbook. london, uk: sage. doi: 10.4135/9781446294413 meredith  et  al       | f l r     35   siskin, l. s. (1991). departments as different worlds: subject subcultures in secondary schools. educational administration quarterly, 27, 134-160. doi: 10.1177/0013161x91027002003 soeters, j. (1988). organisatiecultuur: inhoud, betekenis en veranderbaarheid [organizational culture: content, meaning and changeability]. in j.j. swanink (ed.), werken met organisatiecultuur: de harde gevolgen van de zachte factor [working with organizational culture: the harsh effects of the soft factor] (pp. 15-27). vlaardingen: nederlands studie centrum. strahan, d. (2003). promoting a collaborative professional culture in three elementary schools that have beaten the odds. the elementary school journal, 104(2), 127-146. doi: 10.1086/499746 sveiby, k. e., & simons, r. (2002). collaborative climate and effectiveness of knowledge work-an empirical study. journal of knowledge management, 6, 420-433. doi: 10.1108/13673270210450388 van maanen, j., & schein, e. h. (1979). towards a theory of organizational socialization. in b. m. staw (ed.), research in organizational behavior, vol. 1 (pp. 209-264). greenwich, ct: jai press. van maele, d., moolenaar, n.m., & daly, a.j. (2015). all for one and one for all: a social network perspective on the effects of social influence on teacher trust. in m. dipaola & w.k. hoy (eds.) (pp. 171-196), leadership and school quality. greenwich, ct: information age publishing. visscher, a., & witziers, b. (2004). subject departments as professional communities? british educational research journal, 30(6), 785-800. doi: 10.1080/0141192042000279503 waldron, n. l., & mcleskey, j. (2010). establishing a collaborative school culture through comprehensive school reform. journal of educational and psychological consultation, 20(1), 58-74. doi: 10.1080/10474410903535364 wasserman, s., & faust, k. (1994). social network analysis: methods and applications (vol. 8). cambridge, uk: cambridge university press. doi: 10.1017/cbo9780511815478 wellman, b., & frank, k. (2001). network capital in a multilevel world: getting support from personal communities. in n. lin, r.s. burt & k. cook (eds.), social capital: theory and research (pp. 233273), new york, ny: aldine de gruyter hawthorne. yang, j. t. (2007). knowledge sharing: investigating appropriate leadership roles and collaborative culture. tourism management, 28, 530-543. doi: 10.1016/j.tourman.2006.08.006 puurtinen publication frontline learning research vol 6 no. 3 special issues (2018) 148-161 issn 2295-3159 learning on the job: rethinks and realizations about eye tracking in music-reading studies marjaana puurtinen a a university of turku, finland article received 23 may 2018 / revised 31 august / accepted 18 september / available online 19 december abstract the application of new methods and measures in domains with few methodological traditions of that kind often presents researchers with a challenge; they may have to take up the task of developing their understanding of the phenomenon while, at the same time, creating the practices for its study. for us, the method was eye tracking, and the topic, music reading. one key characteristic of music reading is that the music reader’s gaze moves slightly ahead of the current point of performance. this gap allows the performer to prepare for the upcoming motoric responses. in this paper, we present our 10-year-long path, describing the steps we have taken while studying this “looking ahead” in music reading. we will point out how we have, after both advances as well as setbacks, come to change our views on how best to explain the various components affecting this specific act and how it is best measured. finally, we discuss some of the lessons we have learned, hoping in this way to provide practical suggestions for others who plan to take up methods from other domains and use them in novel ones. keywords: expertise; eye tracking; eye-hand span; eye-time span; music reading corresponding author email: marjaana.puurtinen@utu.fi doi: https://doi.org/10.14786/flr.v6i3.38 1. introduction similar to all researchers who plan to apply a method more often used for studying other topics than the one they are wishing to target, we in our team have had to learn about expertise in music reading in parallel with developing the eye-tracking methodology for its study. our initial educated guesses about how to go about this were based on one team member’s prior experience with eye tracking and on all team members’ much better comprehension of western music notation than of the cognitive side of things (due to our background in music and academic training in other disciplines than psychology). as often occurs in research, some of our hunches were better than others, and our views about both what to study and how to study it have therefore changed considerably throughout the 10 years we have conducted our work. it may be that the interest in process measures in educational sciences puts more and more project leaders, research coordinators, doctoral candidates and their supervisors in similar situations to where we were some time ago. to be sure, many will be much faster in learning their topics than we have been. still, for anyone intending to take up eye tracking or any other similarly complicated method and apply it to a little explored topic, we suggest that it is good to prepare for similar kinds of rethinks and realizations to what we have faced and, in advance, already be thinking of both practical and emotional ways to cope with them. we begin our history by giving some background about eye tracking in the music domain (based on our current understanding of the matter), and then we briefly describe some of our more or less successful data collections and attempts to report our findings from 2008 onward with respect to the study of “looking ahead” during music reading. for the sake of presentation, we write about our sub-studies as steps 1, 2 and 3, although the whole process consisted of several overlapping stages. finally, based on our history, we state some of the lessons we have learned from the perspective of taking up new research methods in little explored fields. 1.1why use eye tracking in the music domain? with eye tracking, we can accurately record the durations and locations of fixations, that is the short moments (typically around 200–400 ms) when we target our gaze to a certain location in our visual array and process information from it. during one fixation, we only see a narrow area accurately (for instance one word of this text), and in order to collect more visual information, we must quickly move our eyes to fixate upon another spot (for instance the next word). the perceptual span is the area from which useful visual information is obtained. for instance, when we fixate on this word, a few words to the left and right of it fit within the perceptual span but are blurred. this blurred visual information helps us, however, to see where the next word begins, and we can then target our next fixation on that word. our eye movements are so rapid that we get the impression that our eyes do not perform these skips and jumps, and we may also fixate upon targets (for instance reread a word) without noticing it ourselves. (for details, see rayner, 2009; holmqvist et al., 2015.) we therefore need eye trackers to record and show us how we are looking at a painting, reading a newspaper, checking the traffic while driving or otherwise examining our surroundings. eye tracking has a long research tradition in psychological studies about text reading. in these studies, eye-movement data is considered to give insight into the cognitive processes related to the reading and comprehending of words and longer texts (see rayner, 1998; 2009). the method has also been used to study visual expertise in various domains (jarodzka, holmqvist, & gruber, 2017), of which chess and medicine are the ones typically mentioned first. in expertise research, the goal is to make experts’ automated visual practices visible and, similarly to text-reading research, to search for indicators of the underlying cognitive processes. often this is done by comparing the visual processing of experts to that of novices or intermediates while they perform a domain-specific task. music reading would seem like a natural part of studies about visual expertise. it is a visual motor task rehearsed and widely used by a large group of amateur musicians and used professionally (and on very high level of expertise) by some, while there are still many almost illiterate with respect to western music notation. music reading also has one feature which points especially to the usefulness of the eye-tracking method; while reading, the performer produces a continuous motor response, as she is constantly “acting out” the stimulus. this gives the researcher constant verification of the quality of the visual processing as well as the possibility of synchronizing, throughout a recording, the “intake” of visual information with a motor output. this is unlike the process with many other visual tasks. (see puurtinen, 2018.) at this point, we need to specify what we mean by “music reading”. to us, it is a task where someone reads music notation and executes the symbols in one way or another. the motor part may consist only of tapping rhythms, singing or playing an instrument (see video 1). we consider it important to make a distinction between performance tasks and tasks where music notation is only read and not performed in any way. to highlight the difference, we have come to call the latter silent (music) reading (penttinen, huovinen, & ylitalo, 2013). as an example, think of a piano teacher taking a pile of sheet music in her hands in order to go through the sheets and select which piece to play with a student. the teacher is probably not reading all of the note symbols to the extent that they could be performed (cf. the performer in video 1) but is scanning the music quickly in order to see whether one piece would seem to be of suitable difficulty. in this type of a task, there are no temporal restrictions imposed upon the reading (see section 1.2, below), nor any need to plan for motor responses. thus, compared to a performance task, we believe that the silent-reading task presents the readers with very different goals for going through the material and, accordingly, with very different cognitive demands (see penttinen, et al., 2013; puurtinen, 2018). however, due to the lack of a need to plan for a motor response, silent-reading tasks could be used as a reference when comparing music reading with the processing of other types of domain-specific symbol systems. for those interested in the visual processing of note symbols with no performance requirements, we suggest that they turn to studies which have applied a non-performance design (e.g., waters & underwood, 1998; burman & booth, 2009; drai-zerbib & baccino, 2013; penttinen, et al., 2013). video 1. the author’s eye movements during the piano performance of a familiar children’s song, “mary had a little lamb”. a metronome (the click heard in the background) is set at 60 beats per minute and provides the temporal framework. recorded with a tobii tx300 eye tracker and a yamaha electric piano. overall, then, the use of the eye-tracking method seems well suited for music-reading studies. still, for some reason, such studies are quite few, and the field lacks a coherent research tradition (for reviews, see madell & hébert, 2008; puurtinen, 2018). we can, however, say that in general, experts in music reading are faster at encoding note symbols than are novices and intermediate performers, and therefore, they can read the music with shorter fixation durations. this expertise effect has been noticed both in silent-reading tasks (waters & underwood, 1998; penttinen, et al., 2013) and during-performance tasks. considering the latter, experts’ advanced skill has appeared both when more and less experienced musicians’ performances differ in terms of performance tempo (truitt, clifton, pollatsek, & rayner, 1997; gilman & underwood, 2003; penttinen & huovinen, 2011) but also when all performances are alike in a temporal sense and in performance quality (penttinen, huovinen, & ylitalo, 2015). the features of the music notation, too, play a role in the reading, though this side of the process has only rarely been studied and when it has been, often only on a very general level (madell & hébert, 2008; puurtinen, 2018). 1.2 music reading is all about the use of time the uniqueness of music reading as a visual motor task is in the fact that it is temporally regulated. time can be thought to affect the reading on two levels. first, the overall tempo affects the amount of absolute time a performer has for inspecting and performing the visual elements on the score. for example, if one performer uses 10 seconds to perform the melody in video 1, and another one plays the same melody in 5 seconds, it is obvious that the latter performer has less time to study the written symbols and plan for the motor responses. secondly, the course of reading is also regulated, because the relative durations of individual notes are signalled by the symbolic system itself. in figure 1, in measure 2, the performer has to perform more (two eighth notes) during the first beat and less (only one quarter note) during the second beat. naturally, the performer may choose to ignore the durations of individual symbols and stop to correct errors, but that will lead to a “staggering” performance, where the flow of the music is disrupted. this is exactly what beginning musicians do when they are incapable of interpreting the symbols and performing them in a given tempo (e.g. drake & palmer, 2000). all of this is unlike the situation in text reading, where readers can return to difficult sections or words and where, in fact, targeted rereading may even result in better text comprehension and be considered a preferable strategy (e.g. hyönä, lorch, & kaakinen, 2002). another complexity of western music notation is the fact that the same symbols contain information about both the relative duration of each symbol and the pitch height of each (e.g. which key to press on a piano keyboard), and the reader has to process both features in order to perform accurately. figure 1 demonstrates the relative durations of notes in “mary had a little lamb” (the first line) along with the whole melody, with pitch heights included (the second line). (note that this is a highly simplified example of western music notation; the amount of information on musical scores can vary from this type of simple information to detailed information about simultaneously performed notes, performance tempos and expressive interpretation.) figure 1. above, the rhythms of “mary had a little lamb”. in western music notation, the score is typically divided into “measures” marked by vertical bar lines. in this example, in measure 2, the performer first performs two eighth notes during one quarter beat (the beat marked in red) and then one quarter note which lasts for the whole quarter beat (the beat marked in blue). the correct timing of the notes (which land either on a beat onset or between two beat onsets) is of importance, as it secures the steady flow of the music. below is the full song, the melody represented through the different pitch heights of the note heads. thus, musicians have to make sense of the complex symbolic system and select what to target their gaze at, as well as when and for how long, in order to perform successfully while operating within the given temporal framework. we know that musicians manage this by maintaining their gaze slightly ahead of the current point of performance, and with the help of this buffer (typically of perhaps around 1–2 seconds [furneaux & land, 1999; penttinen, et al., 2015; huovinen, ylitalo, & puurtinen, 2018; rosemann, altenmüller, & fahle, 2016; see video 1]), the performer prepares for the upcoming motoric responses. the gaze may typically tend to remain very close to the performed notes (truitt et al., 1997; penttinen, et al., 2015), but instead of a steady “looking ahead”, at least for skilled music readers, the reading may also consist (mainly or in part) of rapid back and forth eye movements (goolsby, 1994; cf. penttinen, et al., 2015). overall, this “looking ahead” behaviour, often called the eye-hand span (madell & hébert, 2008; holmqvist et al., 2015) should work as an indicator of music-reading-related cognitive processes, since it seems to vary due to performers’ expertise as well as stimulus features. however, the interplay of all the factors involved is still not properly understood, and the field is methodologically scattered to the extent that only the most elementary features of the “looking ahead” can be regarded to be established ones. (for a recent review of findings thus far, see huovinen et al. [2018].) before 2008, the year of our first data collection, published works on this topic were indeed very few, and we therefore had to start considerably near the beginning. 2. our eye-tracking investigations about the “looking ahead” in music reading 2.1 step 1: focus only on gaze targets (penttinen & huovinen, 2009; 2011) 2.1.1. study summary we began our work in 2008–2009 by recording non-professional pianists’ sight-reading of simple melodies prior to, during and after nine months’ training. our aim was to track down potential eye-movement indicators of early skill development, especially relating to the reading of larger melodic intervals. our author team then consisted of a phd candidate in education (who was also an active musician) along with a musicologist, who were supported by a group of colleagues from a department of teacher education. forty-nine education majors took part in the experiment, of whom 15 novices and 15 amateur musicians’ data was included in the final analyses. the data collection was organized alongside a compulsory music course (with piano playing as one of the course topics), and this allowed us to follow the novices’ development during a real-life learning task. our four experimental melodies consisted of quarter notes (see figure 2) played only with the right hand, with a metronome signalling the onset of each quarter note at a relatively slow pace (60 beats per minute, where each beat thus lasted for 1 s). the melodies were stepwise apart from two larger melodic intervals (see the “skips” at two of the bar lines in figure 2). figure 2. one of the four piano performance tasks applied in penttinen and huovinen (2009; 2011). in this melody, each of the right-hand fingers could be put on one key (the thumb on the first note), and the performer did not have to move her hand. for the most part, even for the novices, this eliminated the need to look at the fingers; with difficult pieces, pianists often need to divide their visual attention between the score and the keyboard to ease motor coordination. we built a set-up where music was presented from the screen of a 50 hz tobii 1750 eye tracker, and an electric piano, attached to a laptop with sequencer software, was positioned in front of the tracker. after careful consideration, no chin rest was used, to allow the performers as natural a playing position as possible. we assumed that although during our simple tasks the performers did not need to look at their fingers to any great extent (apart from prior to performing when placing the hand on the keyboard), preventing that altogether with a chin rest might have had larger effects on the eye-movement data than the rare looks toward the keyboard. previous music-reading studies have not been very precise about how participants in them were trained for the study protocols (sometimes there is no indication that any training took place), but we also took care in preparing a practice trial with exactly the same kind of protocol (down to every instruction slide appearing on the screen) than what was then applied in the actual recording. this proved a useful procedure, and we have applied that practice ever since. we attempted to launch both the eye-tracking and sequencer recording simultaneously from a third computer, but though the system worked in pilots, inconsistent lags started to appear during the actual data collection. thus, we could not synchronize the performance and eye-movement data after all, and our analyses had to be limited to the allocation of fixation time without information about where the performance was ongoing at each fixation. we reported indicators of novices’ skill development in three respects. first, after training, the novices performed the large intervals with better timing and accuracy. second, together with increasing performance accuracy, the novices began to allocate more fixation time to the last two notes of each measure (similar to the amateurs, who did so from the first measurement session onward). third, we noticed that the novices, with training, perhaps began to identify the latter notes in large intervals more quickly. this was suggested by the gradual shortening of the average first-pass fixation durations for those symbols. there were, however, some matters which suggested that some of the results should be treated with caution, and we therefore tried to corroborate them later on. after two revisions, the manuscript was accepted for publication (penttinen & huovinen, 2011). 2.1.2. rethinks and realizations after step 1 this was the first music-reading study we designed, and it certainly had some drawbacks, but perhaps with beginners’ luck, some benefits as well. to begin with the latter, due to some of our background in text-reading studies, we decided to apply in our music-reading studies the eye-tracking measures suggested by hyönä, lorch and rinck (2003) as suitable common ones for text-reading research. we only later fully realized the need for music-reading research, too, to be much more consistent in the use of the measures; we continued with these measures in our later work. another benefit was the choice of very simple melodies; we realized that the set-up should perhaps be kept simple, but the reasons were not solely due to our less sophisticated understanding of the eye-tracking method and the kind of stimuli with which it was most usefully applied. we also simply considered what kinds of melodies our novices could learn to play during the nine-month-long training. however, we did come to think (to some extent, at least) of the fact that due to the multidimensional nature of western music notation (see section 1.1), we should be able to distinguish whether any eye-movement effects were due to the interpretation or planning of the rhythm or melodic features of a certain melody. by simplifying the rhythm in these tasks, we could narrow down our interpretations and suggest that our novices’ eye-movement effects were indeed indicators of the interpreting and planning of the melodic features in the melodies, rather than being about reading and planning the execution of the rhythmic patterns. this has not been the case in many music-reading studies (see, puurtinen, 2018). however, we were also faced with several challenges, both while collecting the data and especially when analysing it. first, the synchronization of performance and eye-tracking data failed, as described above. for this reason, we could only speculate about what had caused shorter or longer fixation times and, importantly, when particular fixations took place. we had no idea whether our novices and amateurs read in a steady manner approximately one second ahead of the performed notes or whether they tended to do longer advance inspections and then return toward the point of performance, as was proposed as one possible pattern for skilled readers (goolsby, 1994). when discussing the findings, we could only offer alternative hypotheses about the “looking ahead” in this specific task. the fixation durations suggested that not all note symbols were treated equally, but we did not have the full explanation of why. and in the analysis of note-specific fixation times, we also applied the kruskal-wallis analyses with dependent data in one part of the analyses, which, we know now, omitted the within-subject dependence. furthermore, we presented the four melodies in the same order to all participants and did not counterbalance them. thus, although the mean values for first-pass fixation times for selected note symbols decreased for the novices during training, the statistical testing should be interpreted with caution and retested with a better study design. we tried to address these matters in our follow-up studies. 2.2 step 2: adding the hand (penttinen, huovinen, & ylitalo, 2012; 2015) 2.2.1. study summary in 2008–2009, when we collected the data set described above, we also invited the same participants to perform another task, to play a familiar children’s song, “mary had a little lamb”, a few times. some versions of the melody contained a one-measure-long “mistake” or notes which were “wrong”, and our idea was to trace the eye-movement indicators of coping with such surprising misprints in the music. here, too, we applied a metronome, which forced the performers to solve the problem of playing something against their expectations at a given tempo. our attempt with the 49 participants failed (for details, see section 2.2.2, below), and we ended up using only five experienced musicians’ data as a pilot study (reported as study 1 in penttinen, et al. [2012]). we developed the setting further and collected a new data set in 2011. at this point, we had also included a student of statistics in our team. due to many novices’ difficulties in performing the melody accurately and in the given tempo in our first try, this time we only asked skilled performers (amateur and professional-level pianists) to perform the same song and two of the variations from the previous attempt. an electric piano was placed in front of a 300-hz tobii tx300 eye tracker, and the performances were recorded into a separate computer. the metronome was provided by the second computer, which recorded the performances in midi format. both computers were operated by the first author, who coordinated the change of slides on the eye tracker monitor so that they occurred at metronome beats. this procedure provided both the eye tracking and midi data streams with simultaneous events, and we created a laborious way to synchronize the data streams based on these markers (for details, see penttinen, et al., 2015). this enabled us to calculate what we called the eye-hand span, that is, how much the gaze was ahead of the onset of a performed note (see figure 3). figure 3. how we calculated the eye-hand span in penttinen and colleagues (2015). in this melody, two quarter beats (beat 1 in red, beat 2 in blue) fit into one measure. the performer plays all of the note symbols. when performing the first note symbol of measure three (in red), the performer’s gaze is targeted near the note in the second beat area of the measure (in blue). thus, at this point, the eye-hand span is, roughly put, one quarter beat. in the tempo of 60 beats per minute, this equals one second. we also calculated what we called gaze activity, namely how many quarter beat areas were fixated on when performing the notes of one quarter beat (i.e. notes fitting under a red or a blue line). this was typically only two beat areas. we observed that professional performers applied longer eye-hand spans more often than amateurs did. also, during one-second intervals, professional pianists fixated on more of the music than amateur performers did. we called this finding an increase in their gaze activity. however, these between-group differences disappeared in the face of melodic deviations, suggesting that performing against expectations inhibited the professional performers’ reading, to some extent. we were also able to compare in detail the “looking ahead” during the performance of different rhythm symbols and obtained some information about the reading of quarter notes and the more rapidly performed eighth notes (see figure 1). for the analyses, we used tand chi-square tests and were fairly content with our protocols for data collection and analyses — for a while. our first full-paper submission was not successful, but the second one was (penttinen, huovinen, & ylitalo, unpublished manuscript; penttinen, et al., 2015). 2.2.2. rethinks and realizations after step 2 as described above, our first attempt (in 2008-2009) with this particular research design failed. first of all, there was a lack of synchronization of the performance and eye-movement data. secondly, the task was too difficult for the most novice participants, and this reduced the data set considerably. erroneous performances were so much unlike each other (with different types of errors appearing in different parts of the melodies) that they could not be pooled. thirdly, to present the melody in a big enough font size, we placed it on two rows (first four measures on row 1, last four on row 2). thus, the performers’ gaze moved from the end of the first row to the beginning of the second, and in such a short reading task, all additional and accidental fixations landing here and there during these sweeps from one line to the next made the eye-movement data even more complicated to handle. finally, one of the three variations we used had the “mistake” in the second-to-last measure. however, during music reading, the endings differ from other parts of the reading, since there is nothing more to look in advance, and thus, the “extra” time gained is just spent on the final measures. thus, this variation did not work out similarly to the other two variations. after we addressed these deficiencies in our second data collection in 2011, we began to get a bit closer to what we were after. skilled performers’ adjustment of the eye-hand span and gaze activity, when something unexpected had to be performed, suggested that the “looking ahead” is indeed involved in the planning of motor responses and that it is affected both by the performers’ expertise and by stimulus characteristics. this time, we had also been able to synchronize the performance and eye-movement data with what we thought to be sufficient accuracy, and we had found measures which had brought forth something new about the “looking ahead”. however, even though we were content with this improved data set and thought it could indeed contribute to our understanding about the “looking ahead”, the first manuscript we submitted was rejected and only later accepted by another journal. this was, at the latest, the moment when we realized that it really is hard to find publication forums for this type of work (apart from music education, which we thought to be one relevant field). as we framed our work both within educational sciences, where music is by no means a mainstream topic, and within cognitive musicology, where in 2011, eye-tracking was still a relatively unknown method, and as none of us were psychologists and therefore not accustomed to write to that audience, we were caught somewhere in between these fields. (at this point, we need to thank all reviewers who took up the job of reviewing our work; some of them have stated that they were not experts in music or eye-tracking, and we do appreciate the effort they put into their reviews.) in addition, we later came to be more critical of our analytical approaches; they were still only a series of separate (piecemeal) comparisons. we started to think that if several factors (e.g. performer characteristics, even the minutest features of the stimulus, and most likely also the performance tempo) did co-influence the “looking-ahead” in music reading, as now seemed to be the case, we should study them in a way which would separate the overriding factors from those which interacted with other factors. we had also applied only one tempo in these studies. (some of these realizations were reached while already working on our next data sets [see 2.3, below]). although this performance task, with its authentic melody and violation of expectations, differed from the simple performance task in step 1 and was designed to address very different kinds of research questions, in retrospect, the studies seem to have shared more common features in a methodological sense than we had previously considered. 2.3 step 3: replacing the hand with metrical time (huovinen, et al., 2018) 2.3.1. study summary step 3 had its origins in the data collection in 2008–2009. we first attempted to publish the 2008–2009 data as a separate paper which focused not on novices’ reading but on the reading of the two different kinds of experimental melodies (ones which had a large interval in the middle of a measure and ones where the “skip” was across a bar line) during correct performances. however, even after submitting a revised version, our work was rejected, and we realized that we did need the synchronized performance data in order to fully explain our findings (huovinen & penttinen, unpublished manuscript; huovinen, penttinen, & ylitalo, unpublished manuscript). thus, in 2011, we asked the same skilled participants who had performed the “mary had a little lamb” task to also perform another task. we modified the 2008–2009 set-up (see 2.1, above) by adding eight more melodies to the set, including four with no large intervals, or “skips”, in them. we had realized that we needed these stepwise melodies so that we would have a baseline with which to compare the melodies with the large intervals. we also asked the performers to play half of the melodies in a slower tempo and half in a faster one. the tobii tx300 eye tracker was used, and all was considered to go smoothly; our skilled performers did the tasks with very good temporal accuracy and few mistakes, and though the work was laborious, we were able to synchronize the performance and eye-movement data with a similar technique as in the task in step 2. our previous measure for the eye-hand span was calculated from the point of performance onward (see figure 3). this was the traditional approach in other prior studies, too. however, we now noticed that in fact, this measure describes more those moments in the performance when the performer can look ahead (huovinen, et al., 2018), that is, moments when performing a symbol is easy enough to allow the reading of symbols further ahead. in our case, that was not the issue of interest (especially when the melodies were almost overly simple to perform). since we had simple melodies with “more difficult” notes in them at the large intervals, our question was whether these notes were fixated upon, or “looked ahead”, as early on as possible and whether the visual processing was therefore affected by such minute features of the written music. this we did not know beforehand. thus, we came to realize that the eye-hand span measure should be calculated exactly the other way around from how we had done it; if we wanted to examine what the performer needed to look at in advance, we needed to take the first-pass fixation upon a specific target as the starting point and calculate the “span” backward to where the metrical time was at that particular moment (see, figure 4). in practice, this would tell us when a performer attempts to fixate a symbol as early on as possible. we therefore differentiated between two measures, the traditional eye-hand spanperformance and the new eye-hand spanfixation. figure 4. how we calculated the eye-hand spanfixation, or eye-time span, in huovinen and colleagues (2018). (this is only for illustration; please note that those studies actually applied stimuli resembling that in figure 2.) the red arrow marks the steady progress of metrical time. the eye-time span was calculated from the first fixation to a note symbol (in this example, from the last note in measure three). by regarding the metrical time as a continuous axis, we could calculate where the metrical time was at the onset of a fixation. in this example, the eye-time span is 1.5 beats or, in a tempo of 60 bpm, 1500 ms (i.e. the metrical time is ongoing 1500 ms behind the fixated note). the keypress (or “hand”) information was only used to estimate the correctness of the performance as well to synchronize the metrical time with the eye-movement recording; after that, all calculations were based on the timestamps for fixation onset with respect to the thus abstracted metrical time. we considered the adding of the eye-hand spanfixation measure a significant advance in our thinking. in terms of data analyses, we also turned to generalized estimation equations and began to statistically model our data instead of listing a series of pairwise comparisons. this shift in analytical thinking later greatly affected how we designed our studies. at this point, we thought we had finally come to understand the “looking ahead” from all its angles and what kinds of questions should be answered with what kind of approach. (later, we noticed that in domains other than music reading, the eye-hand span can also be calculated as the distance between the first fixation on a target and its subsequent performance [holmqvist et al., 2015; huovinen, et al., 2018].) however, our measures and findings were far from easy to explain or present, and our first manuscript, too, was not accepted for publication (huovinen, ylitalo, penttinen, & penttinen, unpublished manuscript). in 2015, we decided to collect one more data set in order to corroborate our previous findings. this data collection was enabled by large project funding, which also gave the three authors a new support team, with representatives from the fields of education and musicology. in what was titled experiment 2 (data from 2011 being experiment 1), we had only professional pianists performing longer melodies which had more difficult larger intervals, again in the same two tempos as in the 2011 study. this time we used the 60-hz tobii t60xl eye tracker. the idea was to make the large intervals more salient and thus make any potential “looking ahead” they may cause more profound. from this perspective, it was only the eye-hand spanfixation measure that could answer this question. thus, we left out the comparison of the two eye-hand measures and moved on to explain the effects caused by the now ever more salient large intervals only with the eye-hand spanfixation measure, and we gave it a new more appropriate name, the eye-time span (since the hand was not involved in the calculations). again, we thought that we were on to something; we were focusing on one characteristic of the music notation and how it affected the “looking ahead”, having also developed a measure which could be used for its study. the feedback for this was not favourable, however, and the manuscript, in which we reported experiments 1 and 2 together, was rejected (huovinen, ylitalo, & puurtinen, unpublished manuscript). after rethinking again the nature of the phenomenon under study, and the measures, we prepared yet a new manuscript, this time including a more detailed description of the eye-hand span and eye-time span, their measurement, and use. in it we reported that besides the general effects of performance tempo and musical expertise on the eye-time span, the local melodic complexity (i.e. the larger intervals) could elicit longer advance inspections; however, these inspections might not land on the “complex” notes themselves but on the ones preceding them (huovinen, et al., 2018). 2.3.2. rethinks and realizations after step 3 all of step 3 indeed took a while, and those years included a number of disappointments and moments when we had to admit that what we had thought to be a good idea was not. nevertheless, despite the emotional struggle, in retrospect, we know that we needed all of that earlier work. during each stage (be it data collection or the writing of a rejected manuscript), we realized and came to think of something which then helped us with the next study design and analyses. we also improved the formulation of research questions and are now better able to estimate what kinds of evidence eye tracking can provide us with in this context. as a result, we can now argue for slightly more straightforward explanations of our findings and can see more clearly which way we or some other team should turn next. however, the results we reported by huovinen and colleagues (2018) are still not the full answer to all of our questions. if, in suitable circumstances, as the data suggests, the pianist’s gaze is drawn toward salient features, but fixations do not necessarily land exactly on them, this suggests that the perceptual span (see 1.1, above) has come more clearly into play. in some cases, then, the salient features observed (albeit in a blurry form) within the perceptual span may attract a visual response so early that the performer can not target a fixation upon that exact note; if she did, the gaze might move too far away from where the music is ongoing. this would perhaps make it difficult to keep all the notes in between in short-term memory and perform them correctly. whether this undershooting is involuntary due to temporal restrictions, or deliberate because the procedure is enough for fitting the problematic symbols into the area of accurate vision, we can not be sure. however, perhaps the visual information which is visible from that fixation point is still enough to allow the performer to get ready for a correct motor response. in order to search for the full answer to this question, we would have to conduct studies that apply the gaze-contingent methodology, that is, where a limited part of the notation is shown to the reader at a time (rayner, 1998). the method has apparently been applied only in two studies about the size of the perceptual and eye-hand spans during music reading (gilman & underwood, 2003; truitt et al., 1997), but not with the most controlled music stimuli—and we do claim, based on all our research, that the stimulus features should be included as part of the analyses and should not be thought of only as part of the study protocol (puurtinen, 2018). we have, to this day, opted for a “natural” layout of the music, showing all of the music during our recordings, but after step 3 it now seems that even more experimental set-ups may be needed in order to further explain the interplay between the eye-time span and the perceptual span during music reading. another matter we have come to recognize is that both the eye-hand span and the eye-time span we have calculated have one drawback: we have only calculated them at certain points during the performances (although in step 2, the gaze activity measure described to some extent what happened between two quarter beats). the eye-time span is actually a continuously changing parameter, similar, for instance, to pupil size. thus, to truly understand the “looking ahead”, the study of this aspect should not be limited in future only to within-subject or between-subject comparisons at certain notes. instead, we should also develop methods which allow the examining of the “looking ahead” as a process, a characteristic of a full performance, varying according to the factors which appear to affect it. 3. lessons learned as the examples above perhaps demonstrate, our study of the “looking ahead” in music reading has been a long and, on occasion, also tiring endeavour. we have came to change a number of things in our methodologies, such as the skill level of our study participants (starting with novices and amateurs and finally recording performances of professionals only), our measures (developing one, the eye-time span, that we thought will be the answer to our questions but which actually lead to even more open ones), as well as our analytical approach (from piece-meal comparisons to statistical modelling). already these examples should demonstrate the methodological twists and turns of our work. to be sure, each step we have taken, from planning an experiment to the final published report of the findings, has made us less the novices we initially were. still, we are aware that the path toward expertise in this methodology and domain will require the taking of many more similar steps. hopefully, this history of our choices and our reasons for moving away from our earlier choices are of benefit to others planning to use new methods and measures and perhaps even use them in domains where there is little methodological support available. we will now attempt to summarize some of the more general lessons we have learned for the benefit of others who may find themselves in a similar situation to where we have been (and still are), using our own experiences as an example. first, many of us who are trained in something other than the computer sciences might not be in our comfort zone when we first start to apply research methods which require the use of technology we are not used to. for us, for instance, the synchronizing of the two recording onsets in step 1 was done with remote support from a manufacturer. we were able to build the system ourselves, but when it failed, there was nothing we could do about it. when learning the use of new equipment, and with limited skills in solving problems related to it, we recommend that researchers create a plan b in their research design; they should ensure that their research questions can be answered, at least in part, even when something goes wrong on the technical side. in our case, the comparison of novices and amateurs helped us to make use of the data, and we also included a task of silent music reading where the synchronization was not needed (penttinen, et al., 2013). however, more insight into the participants’ reading would have been gained had the technical side been more successful. second, when starting to work with a new method and/or new types of stimuli, an exploratory pilot with a complex and perhaps realistic task is a possible starting point for creating research hypotheses (see, puurtinen, 2018), but to get beyond descriptive data, it may be useful for researchers to control their stimuli (or task) as much as possible. this might not answer the most interesting of research questions, but enables one to test the possibilities of the method. for instance, in prior work about eye movements in music reading, the role of the applied music was thought of more as a part of the study protocol, something “read” or “played”, although it now seems that even the minutest details of the music (whether there is one quarter note or two eighth notes to be played, or a slightly larger interval among an otherwise stepwise melody) are already reflected in the eye-movement process. thus, following this line of thought, before knowing how a complex task is actually represented in the measure of the researchers’ choice, they should start with simple set-ups. the reporting of the findings might be challenging (we have heard a number of times how our tasks are not “musical” at all, since they are so simple), but in the end, one can argue with more certainty what one’s findings might mean and what their cause might be. third, for others beginning their work with new methods, we suggest the use of measures from other domains (or previous studies, if available). importantly, pay attention to the fact that measures which different research teams consider the same may, in fact, be calculated quite differently. with unique measures and their definitions, one’s results might be very hard to align with others’ findings. in music-reading studies, there is great variability in the measures and in how they are calculated (puurtinen, 2018); the eye-hand span, with its many ways of operationalization, is a very good (or bad) example (huovinen, et al., 2018). in our work, partly because they were familiar to us, we began from the very beginning with some measures applied in text-reading research, and we have been content with that early choice we made. in developing areas for research, we should aim at methodological consistency to the extent that it is possible and should provide data and reports which can later be used in meta-analyses. to be sure, measures also need to be developed further and new ones tested; still, existing measures provide a good starting point. fourth, after conducting some studies with similar protocols, test the protocols. this is something we should also do. for instance, our team members’ own background in music made us think of using a metronome in order to secure the temporal similarity of the performances. in practice, we had an external “click” marking the onset of certain beats. we only later noticed that this was not a standard protocol, up until that point, at least. the external temporal control means that the performer herself does not have to use any of her cognitive capacity to maintain the tempo (something that is challenging, especially for novice and intermediate performers), but still, the possible effects caused by the external click should also be tested. it may be that in our data, there are main effects caused either by the strain of following the external click or by the lack of strain due to the relaxation of the need to maintain the tempo without any help. similarly, with other methods, there may be protocols which first seem sensible and which are ones we become accustomed to, but those could also raise new research questions when further thought through. fifth, if possible, create or join a team, and make it something more than just a group of colleagues. in research areas which reach out to several background theories or methodologies in particular, no one can be the expert on all of the relevant topics, and it is difficult for one person to reach all the audiences which might be interested in and whose input would be beneficial for the work. also, potential disappointments and difficulties are much easier to handle with people whom the researchers trust and who support them — and the celebrations are also much more fun like that. thinking of our step 3, one may stop to think whether, alone, they would have done the three data collections, written and submitted the several manuscripts and revisions which were rejected, or kept one subproject going for the ten years it took to get one piece of the work published. we now think it is a good thing that we made it (though we immediately saw the next steps someone should definitely take). of course, in long-term projects and among people who work with something as personal as their intellectual skills, small or large conflicts necessarily arise at one point or another — but with good enough emotional bonding and similar goals, these do not break the team and just become a part of life. to conclude, we have here presented how we have tried to develop our methodologies for the study of “looking ahead” in music reading, openly listing our unpublished work and what we now know were mistakes or simply bad ideas. overall, as will probably also be the case for many others, we do not think our path has simply been either a success or a failure. instead, it has been a back-and-forth movement (cf. goolsby, 1994), and we are glad to have the chance to report also the backward steps. hopefully, other researchers, too, can share those parts of their projects which are faded out of polished publications and presentations — those are, after all, the ones we could really learn from. keypoints the application of a method in a research domain where there is little methodological tradition of that kind presents the researcher with a challenge; likely, there is a need to develop the understanding of the phenomenon and the methodology in parallel. by using our own 10-year-long work as an example, we demonstrate the advances and setbacks we have faced in our eye-tracking studies about music reading. we discuss some of the lessons we have learned in order to provide practical suggestions to others planning to take up a new method or apply one in a novel domain. acknowledgments this work was supported by the academy of finland (project number 275929). the author would like to thank erkki huovinen and anna-kaisa ylitalo for collaboration in the presented work, the members of the “reading music: eye movements and the development of expertise” consortium for their support, the study participants involved in the data collections described in this paper, and the anonymous reviewers for their feedback and suggestions. references burman, d. d., & booth, j. r. (2009). music rehearsal increases the perceptual span for notation. music perception: an interdisciplinaryjournal, 26(4), 303–320. doi: 10.1525/mp.2009.26.4.303 drai-zerbib, v. & baccino, t. (2013). the effect of expertise in music reading: cross-modal competence. journal of eye movement research, 6(5), 1–30. doi: 10.16910/jemr.6.5.5 drake, c., & palmer, c. (2000). skill acquisition in music performance: relations between planning and temporal control. cognition, 74(1), 1–32. doi: 10.1016/s0010-0277(99)00061-x furneaux, s., & land, m. f. (1999). the effects of skill on the eye-hand span during musical sight-reading. proceedings of the royal society of london, series b, 266 (1436), 2435–2440.doi: 10.1098/rspb.1999.0943 gilman, e., & underwood, g. (2003). restricting the field of view to investigate the perceptual span of pianists. visual cognition, 10(2), 201–232. doi: 10.1080/713756679 goolsby, t. w. (1994). eye movement in music reading: effects of reading ability, notational complexity, and encounters. music perception, 12(1), 77–96. doi: 10.2307/40285756 holmqvist, k., nyström, m., andersson, r., dewhurst, r., jarodzka, h., & van de weijer, j. (2015). eye tracking: a comprehensive guide to methods and measures. oxford, uk: oxford university press huovinen, e., & penttinen, m.(unpublished manuscript). the allocation of fixation time in simple sight-reading tasks. huovinen, e., penttinen, m., & ylitalo, a.-k.(unpublished manuscript). the visual processing of melodic group boundaries: an eye-movement study. huovinen, e., ylitalo, a.-k., penttinen, m., & penttinen, a.(unpublished manuscript). the where and when of sight reading: effects of performance tempo and musical structure on eye movements. huovinen, e., ylitalo, a.-k., & puurtinen, m.(unpublished manuscript).the eye-time span: structural salience and looking ahead in simple sight-reading tasks. huovinen, e., ylitalo, a.-k., & puurtinen, m. (2018).early attraction in temporally controlled sight reading of music. journal of eye movement research, 11(2), 1-30. doi: 10.16910/jemr.11.2.3 hyönä, j., lorch, r. f., jr., & kaakinen, j. k. (2002). individual differences in reading to summarize expository text: evidence from eye fixation patterns. journal of educational psychology, 94(1), 44–55. doi:10.1037//0022-0663.94.1.44. hyönä, j., lorch, r. f. jr, & rinck, m. (2003). eye movement measures to study global text processing. in j. hyönä, r. radach, & h. deubel (eds.), the mind’s eye: cognitive and applied aspects of eye movement research (pp. 313–334). amsterdam: elsevier science. jarodzka, h., holmqvist, k., & gruber, h. (2017). eye tracking in educational science: theoretical frameworks and research agendas. journal of eye movement research, 10(1):3, 1–18. doi: 10.16910/jemr.10.1.3 madell, j., & hébert, s. (2008). eye movements and music reading: where do we look next? music perception, 26(2), 157–170. doi: 10.1525/mp.2008.26.2.157 penttinen, m., & huovinen, e. (2009). the effects of melodic grouping and meter on eye movements during simple sight-reading tasks. in j. louhivuori, t. eerola, s. saarikallio, t. himberg, & p. eerola (eds.), proceedings of the 7thtriennial conference of european society for the cognitive sciences of music , jyväskylä, finland. (pp. 416-424). available at https://jyx.jyu.fi/dspace/handle/123456789/20910 penttinen, m., & huovinen, e. (2011). the early development of sight-reading skills in adulthood: a study of eye movements. journal of research in music education, 59(2), 196-220. doi: 10.1177/0022429411405339 penttinen, m., huovinen, e., & ylitalo, a.-k. (unpublished manuscript). eye-movement effects of melodic deviations in temporally controlled sight reading. penttinen, m., huovinen, e., & ylitalo, a.-k. (2012).unexpected melodic events during music reading: exploring the eye-movement approach. in e. cambouropoulos, c. tsougras, p. mavromatis, & k. pastiadis (eds.), proceedings of the 12thinternational conference on music perception and cognition and the 8thconference of the european society for the cognitive sciences of music, thessaloniki, greece. (pp. 792-798). available at http://icmpc-escom2012.web.auth.gr/sites/default/files/papers/792_proc.pdf penttinen, m., huovinen, e., & ylitalo, a.-k. (2013).silent music reading: amateur musicians’ visual processing and descriptive skill. musicæ scientiæ, 17(2), 198-216. doi: 10.1177/1029864912474288 penttinen, m., huovinen, e., & ylitalo, a.-k. (2015).reading ahead: adult music students’ eye movements in temporally controlled performances of a children’s song.international journal of music education: research, 33(1) , 36-50. doi:10.1177/0255761413515813 puurtinen, m. (2018). eye on music reading: a methodological review of studies from 1994 to 2017. journal of eye movement research, 11 (2), 1-16. doi: 10.16910/jemr.11.2.2 rayner, k. (1998). eye movement in reading and information processing: 20 years of research. psychological bulletin, 124(3), 372–422. doi: 10.1037/0033 rayner, k. (2009). eye movements and attention in reading, scene perception, and visual search. the quarterly journal of experimental psychology, 62(8), 1457–1506. doi: 10.1080/17470210902816461 rosemann, s., altenmüller, e., & fahle, m. (2016).the art of sight-reading: influence of practice, playing tempo, complexity and cognitive skills on the eye-hand span in pianists. psychology of music, 44(4), 658–673. doi: 10.1177/0305735615585398 truitt, f. e., clifton, c., pollatsek, a., & rayner, k. (1997). the perceptual span and the eye-hand span in sight-reading music. visual cognition, 4(2), 143–161. doi: 10.1080/713756756 waters, a., & underwood, g. (1998). eye movements in a simple music reading task: a study of experts and novice musicians. psychology of music, 26(1), 46–60. doi: 10.1177/0305735698261005 microsoft word chorney_publication.docx             frontline  learning  research  vol.5  no.  1  (2017)  43  -­‐  57   issn  2295-­‐3159       contact information: sean chorney, 8888 university drive, v5a 1s6, burnaby, canada, email: sean_chorney@sfu.ca doi: http://dx.doi.org/10.14786/flr.v5i1.229   re-animating the mathematical concept: a materialist look at students practicing mathematics with digital technology sean chorney simon fraser university, canada article received 27 november / revised 30 september / accepted 26 october / available online 15 february abstract this paper proposes a philosophical approach to the mathematical engagement involving students and a digital tool. this philosophical proposal aligns with other theories of learning that have been implemented in mathematics education but rearticulates some metaphors so as to promote insight and ideas to further support continued investigations into the learning of mathematics. in particular, this philosophical proposal takes seriously the notion that a priori to activity, there are no objects which in turn challenge the notions of intention, affordance and/or representation. to exemplify this perspective, two episodes of grade nine students using a dynamic geometry software are analysed to elaborate how mathematics can be seen to emerge from working with a tool. keywords: post-humanism; materialism; mathematics education; theories of learning   chorney       | f l r     44   1. introduction with the proliferation of digital tools in the modern age, the implementation of these tools into the mathematics classrooms is becoming ubiquitous. these tools that include graphing calculators, ipads, iphones are implemented for the purpose of improving mathematical learning. since technology is changing so fast and new applications are being created everyday, different learning theories are necessary to accommodate or adapt to these new technologies. indeed, it is ultimately a philosophical question as to how a material object, a technological/mathematical tool for example, can become part of, contribute to, and help develop mathematical thinking. however, in addressing this philosophical challenge, this paper moves away from the “learner” as central and asserts that mathematics emerges amid the dynamic relations of humans and materials. the highlighting of materials partially aligns with actor network theorists since actor network scholars value the contributions of material objects in the production of knowledge. however, this study notes that materials are variable and often boundary-less and as such the perspective taken in this paper takes seriously the positioning exemplified by ingold’s (2011) provocation that there are no materials a priori to an activity. this paper explores a profound philosophical shift in terms of where it locates the "thinking" in a mathematical activity. this shift has been called post-humanist (barad, 2007) as well as new materialist (de freitas & sinclair, 2014). it is common in teaching to describe knowing in a concrete and human-centric way. “she knows the quadratic formula” is a simple phrase that conveys the ownership of a proposition. however the implication in such statements place the human as central knower and agent. in education studies, the student is often central, and although this seems appropriate within an educational system with a mandate to educate the individual, new theories are pointing to alternate conceptions that may offer productive new ways of understanding tool-based mathematical activity. i adopt the term post-humanist (barad, 2007) to challenge the isolating tendency or reductionist approach to see the human as the central actor. in post-humanist practice, the human is considered to be just one of many “actors” involved in that practice. other actors can include such things as tools, social influences, concepts or even a task itself. i draw on post-humanism not to deny the intangible, experiential aspect of mathematical practice but instead to implicate the dynamic and significant effect of the artefact on practice, and, in this study, how this implicates the emergence of the mathematical. that is, the material aspect of working with concrete tools is more than a matter of mastering a tool to exploit its affordances. a student and a tool might instead be seen as being on more equal footing, so that the tool is not simply subjugated to the all-knowing or all-deciding human. in this way post-humanism refers to the idea that mind and matter are not ontologically distinct. new materialism aligns with the ontological monism of spinoza and the historical materialism of marx, adopting the important aspects of the political, the ontological and the epistemological into mathematics education. the term “new”, however, in “new materialism” moves into a world of agency and animation and argues for mobility and action as necessary components to generative development. when distinctions of human agency and matter are challenged, particularly in education, different ways of seeing emerge in how the use of tools can generate interesting ways of thinking about learning. indeed, similar to learning theories, many current methods privilege the human actor (e.g. in transcripts, in modes of describing events in terms of being actor-driven). while there are numerous contributory frameworks that address the questions that emerge in an analysis of the relational engagement of a human working with a material, i suggest this study differs from many approaches by adopting both a post-humanist and a new materialist perspective. throughout this paper i draw on how these perspectives come together in one framing. in part, the post-humanism draws attention away from human and the new materialist approach looks more to the material. it is the combination of both approaches that i suggest differs from previous work in mathematics education. i suggest that to create turbulence in how one “reads” common learning situations is the first step towards new ways of seeing. i do not present insights based on the intrinsic nature of learning but rather a   chorney       | f l r     45   metaphoric re-description (hess, 1980) that in and of itself can lead to different ways of seeing. the question addressed in this paper is not about informing educators about the way things are. the purpose of introducing a post-humanist and new materialist approach is for the production of potential possibilities and to explore alternative ways of understanding different framings of mathematical activity. following from this re-conception of mathematical practice, the underlying question that emerges in this paper asks what insightful understandings and interventions emerge from a post-humanist/new materialist approach. in the first section, i discuss current approaches to theorizing tool-use in mathematics education. i then contrast these with a post-humanist, new materialist perspective on tool-use, which challenges the human-centric view of activity and argues for a process ontology. i draw upon the work of andrew pickering and articulate his resistance/accommodation model as the way in which scientists and mathematicians create new machines and/or ideas. i also expand on tim ingold’s work and his argument that tools are not distinct entities and should be seen as narratives—that is, how they may contribute to function. in this study i present two narratives that support a storied knowledge of actual engagement1. finally, i briefly describe the work of karen barad, who offers the powerful notion of intra-action in her new materialist account of the nature of scientific concepts, in which concepts and tools, for example, are seen as a single entity (a relation) rather than as two interacting things (or relata). i suggest that each of these scholars has identified particular and distinctive approaches to tools that are relevant to the study of tool-use in the mathematics classroom. i finish with the analysis of two episodes of grade nine students using a dynamic geometry software in terms of the framework described in this paper. 2. tools in mathematics the theoretical foundation of tool-use in mathematics education continues to be problematic (waltz, 2006). at one extreme, mathematics is valued as a mental discipline where concrete tools are dismissed as being a mere aid to learning, but not as constitutive of the knowing or of the mathematics (see balacheff, 1988). another approach is that tools can be considered an essential element to mathematical practice—a position that rotman (2008) encapsulates nicely when he says that mathematics has been, and will continue to be, involved in a two-way, co-evolutionary relationship with machines. here, the implication is that there would be no mathematics without mathematical tools, and vice versa. in mathematics education research, some of the more common frameworks for addressing the challenges of accounting for how tools are brought into and included in mathematical activity are instrumental genesis and semiotic mediation. for example, instrumental genesis suggests the instrumentalizing of an artefact develops over time as people use them for particular purposes, thus transforming them into tools. this approach distinguishes the instrumetalization of a tool as part artefact, part cognitive schema (artigue, 2002). the individual’s mental schemes together with the artefact’s inherent potential is what makes the artefact an instrument. the artefact is a material object, but an instrument is a psychological construct. verilon and rabardel (1995) define this process as one of appropriation, of making the tool one’s own. instrumental genesis does not focus solely on individuals with tools; instead, it also incorporates socio-cultural issues, including institutional meaning, class norms, and teacher’s expectations. according to ruthven (2002), these social factors are integral to the activity of instrumentation. another well-developed framework for thinking about tools within mathematics education is that of semiotic mediation (bartolini-bussi & mariotti, 2008). the theory of semiotic mediation has been developed specifically within the field of mathematics education with a focus on analysing the semiotic potential of a                                                                                                                           1   ingold seems to use “narrative” and “storied knowledge” interchangeably. however, in the mathematics education literature, story and narrative often have different and distinct meanings (see dietiker, 2013)     chorney       | f l r     46   tool—that is, its potential for linking personal meanings with mathematical ones—and then studying how it can be realized by teachers in classroom interactions. this theory is based on the work of vygotsky and has been significant in socio-cultural and –historical understanding of thinking. in each of these frameworks, the tool plays an important role and influences the material situation of learning. another framework that has recently emerged is proposed by nemirovsky et al. (2013). nemirovsky et al. argue against the dualism of tool-mediated expression and mathematical understanding (and its related dualism of body and mind). in their non-dualist approach, they offer an approach to mathematical thinking that involves the temporal and developing entwining of perception and motor skills, and call the “interpenetration” of perception and motor skills “fluency”, the development of which constitutes mathematical thinking. a mathematical instrument for nemirovsky et al. is material, semiotic and has a set of embodied practices. they contrast their use of the term “instrument” with that of the theory of instrumental genesis, where an instrument is a mixed entity, part artifact (material) and part schema (mental). they see mathematical thinking as the bodily experience of developing fluency on the mathematical instrument, stating, “mathematical activity is constituted by bodily activity” and, correspondingly, mathematical learning as the “transformations in learners’ engagement in mathematical practices” (p. 376). the theory of enactivism based on systems theory (eg. varela, thompson, & rosch, 1991) has also been implemented in mathematics education and goes further in challenging boundaries of human body and environment. goodchild (2014) describes enactivism “as active processes that occur directly through the interaction between the cognizing subject and the environment, rather than as a construction of representations of the environment by the cognizing subject” (page 210). these theories have contributed insightful approaches and encourage further exploration to enrich perspectives and provide continued generative ways of thinking about education. the approach outlined in this paper aligns in many way with the previously mentioned frameworks. i suggest however that this study deviates from these approaches in two significant ways. i draw upon ingold’s notion that there are no objects a priori to activity. i refer to roth (2011) to support this difference, he notes that “kant, piaget, the constructivists, and the embodiment/enactivist theorists all presuppose a subject…” (p. 225). ingold’s provocation goes one step further than roth’s observation by denying any a priori distinctions of individuated entities. this has profound implications for what mathematics is since it cannot be found “in” the cognizing subject. this is the second deviation. mathematical concepts are framed as material. in the following section, i describe three different theoretical approaches that challenge the positioning of humans in activity and also in their conceptualizing of tools and of concepts. 3. frameworks to move away from subject object positions in this section, i look at andrew pickering, who draws attention to the temporal aspect of practice and gives materials a voice. i also look at tim ingold, an anthropologist who proposes that all “things” are snapshots of processes, including tools. i also draw on barad for her approach to concepts as physical apparatus. finally, i move to de freitas and sinclair’s elaboration of barad’s materialism in a mathematics educational framework. these perspectives each contribute a certain way of looking at data, and i suggest they lead to pressing questions and insights that support a unique way of seeing mathematics emerge from tool use. 4. material agency   chorney       | f l r     47   pickering (1995) offers the most accessible approach to the interactions of tool and human arguing for a back and forth model. he is concerned with the advancement and progression of scientific practice (including mathematics) in activity. in one of his case studies, he articulates how donald glaser attempts to accumulate data on strange particles. pickering argues that in glaser’s experimentation obstacles were a natural part of interacting with materials. consequently, glaser’s final result was very different from his initial intentions. pickering uses this example to emphasize that materials influence our activities. thus, pickering presents an analysis of scientific advancement in action in which material agency, as he calls it, is attributed to the natural phenomena with which scientists interact. his model can be summed up simply: materials can be understood as having agency when their structures, make-up and design restrict the subject within a context of activity. pickering then argues that scientists will accommodate their actions to overcome obstacles, and identifies accommodation as a response to these obstacles. in this way accommodation can be seen as a readjustment of action. i suggest that material agency has its relevance in mathematics specifically when working with tools. i draw on pickering’s constructs of resistance and accommodation as a way of approaching the phenomenon of an individual interacting with a material, specifically, material tools. for pickering, the construct of resistance can be seen as an obstacle to performing an action. on the other hand, accommodation is the response to resistance, usually in the form of a readjustment of an act. pickering’s construct of resistance is the important construct and may need some elaboration. for pickering, the construct of resistance can be seen as an obstacle to performing an action. it is important to understand in his model that resistance is not a human action but a material action. that is, resistance is a material obstruction to a goal or an intention. but resistance, according to pickering, goes beyond the simple notion of impeding, it is for him both a micro and a macro construct in that it is not only a “challenge” to overcome but also a way that materials, those outside the bounds of social norms and subjective interpretation, determine action. resistance, for example, can be thought of as a “voice” of an entity—a voice that is not audible but is expressed as a dynamic, in the moment, emergence of form. in this way, resistance can be thought of as an expression of the material's agency, giving it a dynamic, active “voice”. given that accommodation, or readjustment, on the part of the human depends upon the resistance of the material, and drives the activity in a fundamental way, it highlights that movement does not reside solely in human initiation. although a quick interpretation of this model seems to support a binary divide between the human and the tool, as well as centralizing the construct of accommodation in the human, this reading does not necessarily commit us to an anthropocentric point of view. in fact, one of the goals of pickering is to move away from human centric activity and he does this by offering the construct of material agency. his model is not about expressing a truth statement (lyotard, 1984) of what is occurring, it is rather a shift of attention that can help illuminate different ways of seeing. 5. becoming as opposed to being the importance of activity and interaction identified in the previous section, leads me to identify a process ontology as a significant part in developing theoretical frames in approaching mathematical activity. i draw from alfred north whitehead (1929/1978) who privileges process as the ontological realty of the world. for any entity to be identified, it is important to consider that there was a preceding activity that brought it to its current “state”. ingold (2011) extends whitehead by describing all “things” as processes, and while “being” implies a state, “becoming” is a process. humans become as they unfold within the weave of the world. he writes, “to move, to know, and to describe are not separate operations that follow one another in series, but rather parallel facets of the same process” (p. xii) implicating not only the practices of a researcher but also the process of learning. in movement then, a person elicits a knowing which is not a   chorney       | f l r     48   “property of knowing” but a “practice of knowing” and where “knowledge is perpetually under construction within the field of relations" (p. 159) in particular material contexts. ingold also values the material experience before the mental act. in the western world, epistemology is seen in terms of ideas and images, but for ingold, meanings of things come from our embodied experience. ingold states, “practical activity brings incorporeal minds into contact with a material world” (p. 21) and such a process defines the actors. the person is seen as a manifestation of a process of becoming, of continuous creation. this process ontology moves us from a noun-oriented understanding of things to a more verb-oriented approach where all “things” are always in motion. the significance of tools is not in their distinct demarcation, but rather in the role these features play—their dynamism—in relations with human actors. in this study’s analysis, i will mobilize the notion that meaning emerges only in relations. for ingold, the metaphor of lines is very valuable for distinguishing the process of becoming from the state of being. he contrasts the notion of direct lines, which imply transport or a passage from one place to another, with that of wayfaring, which can be thought of as improvisational movement that carves out a path from a starting point. while direct lines require the implementation of intentions, mental images or models in order to get from one place to another (for example, from one van hiele level to the next), wayfaring is first and foremost about the going or the moving. ingold uses the example of looking at art in a gallery to illuminate the difference between direct lines and wayfaring by arguing that looking at art is not a “shuttling back and forth between radically opposed and mutually exclusive domains of mind and world […] but rather to bind mind and world in an ongoing movement” (p. 178). as can be seen in this quotation, ingold is trying to reinscribe thinking within the temporal world of lived experience where there are no station stops to mark new thoughts. if we take looking at art, in ingold’s description, to be like using mathematics tools, then we can see the latter as being less about the development of mental schemes or intentional, efficient deployment, and more about a fused, mind-world movement. in this framing, it becomes important to attend to the largely spontaneous movements of the wayfarer. this can be hard to do in post-hoc assessments of what has changed during mathematical activity, when the direct lines metaphor encourages a view of there having been transport from state a to state b. it is within this framework that ingold (2011) presents a perspective that does not see “things” but adopts emergence in activity as a way of “seeing” reality. that is, he argues one makes a conscious choice to bypass an a priori reflection and offers the notion of narrative to describe the process of movement and knowing. it may seem common sense to identify objects when we look at the world, identifying this and that as objects ready-to-use or as things that exist. however, if we are to take seriously a process ontology, identifying objects can misinform. things “exist” and are “present” if we choose to draw a boundary around them and speak and/or think of them differently than the “other things” around it, but this, of course, is less an ontological reality than it is a choice of distinction or classification. according to linguist, deutscher (2010), language affects how one thinks and perceives the world. the act of nominalization that is common in represetationalist theories of mind can affect how one sees. english is a noun-based language and its use can suggest that it is things or beings that act. ingold’s metaphor is important in this regard because the line does not connect the subject and object like it might do in pickering’s framing or in actor network theory (latour, 1987). instead the line travels in a wavy pattern back and forth between the “subject” and “object” but the line only moves toward or away. the purpose of this metaphor is that it helps shift attention away from the objects so that theorizing of individual things are curtailed because according to a non-dualist ontology nothing acts alone. it is not just an attempt to break down the ontological divide between being and things but also to shift our gaze of analysis to one of function. ingold elaborates his argument against identifying objects as things in and of themselves by suggesting that tools are not determined by their names but by their “storied past”. he uses his own experience of sawing wood as an example to illustrate how the saw as well as himself are drawn into use and become tool and sawer together in activity. he argues that it is not an example of a human using an appropriated tool. the functionality of a tool, if one can describe it as such, is not a result its form or its design alone but is based on its history of use. a distinguishing aspect of ingold’s articulation of tool as   chorney       | f l r     49   narrative is his moving away from an outline of form. an outline, for ingold, is to make a distinction between inner and outer; it is to establish a boundary and, for ingold, this boundary is the delineation of a closure in which movement is restricted. it is within all these considerations that ingold challenges the notion of affordance originally put forth by gibson. according to ingold, gibson’s notion of affordance is an elusive quality, one that gibson has trouble reconciling. gibson states, “but, actually, an affordance is neither an objective property nor a subjective property; or it is both if you like.” (in ingold, 2011). ingold notes that gibson cannot quite commit to whether an affordance exists in the object or in the relation of use. ingold troubles the notion of affordance arguing that an affordance cannot exist prior to activity. if, for example, a child is placed in a room with a bunch of toys, one might argue that they will see toys and play with them but this only shows that a child can listen to instruction and imitate an approach to a world that has already been organized by way of language and interaction. for ingold, these toys would not be toys until they are “toyed” with, and only then the playing and the engagement that emerges is what “possibly” makes them toys, not because they are called toys in the first place. meta-level descriptions of students working with tools may seem similar in certain ways. for example, the students drew a circle with a compass. however, in more nuanced and detailed observation the narratives of each student might be quite different: some students had their compass slip, one drew an ellipse, another drew part of the circle off their paper. but even in these examples, there are problems because if we accept the notion that there are no things preceding activity, we should not begin the narrative with a student nor a compass. we should begin with a narrative of movement. this approach engages descriptions that do not define a trajectory-like approach to an outcome (eg. drawing a circle) but engages us in a description of a wayfarer, one who becomes and does in the moment, where an endpoint is not conceived (partly because there is no endpoint). 6. concept as material: mobilizing the concept in this section i go further with relations in activity and challenge the boundaries of things. barad (2007), a philosopher, feminist and physicist, adopts a post-humanist perspective in analyzing neils bohr’s quantum physics theories emerging out of early 20th century. she seeks to address bohr’s philosophy that concepts depend upon arrangement of apparatus. bohr’s contemplations of the two-slit experiment and the resulting wave-particle duality of light motivates barad. she notes while much of the community debated the nature of light, bohr observed that it was the “actual experiments that displayed the “dual” nature of matter and light” (p. 105). by looking at this example, barad concludes “the nature of observed phenomenon changes with corresponding changes in apparatus” (p. 106). concepts are not ideational, they depend upon physical apparatus. apparatus is not something that sits on a shelf waiting to serve a particular purpose, barad argues that science experiments have shown throughout history that apparatus are constituted through practice and are often rearticulated in new reworkings; thus, the apparatus is usually thought to merely aid in identifying the concept but is more appropriately thought of as creating them. similar to ingold, barad argues apparatus does not “pre-exist the experiment but rather emerges from it” (p. 142). barad aligning with ingold argues “objects are not already there; they emerge through specific practices” (p. 157). to move away from object and subject, barad uses the term intra-action as an alternative to interaction, positing that the prefix “intra-” is more ontologically sound than “inter-” given that things do not come to act together as individual, fixed parts but rather become together depending on their relation. building upon barad’s work, de freitas and sinclair (2014) describe the inseparability of concept and matter. they argue that the performative boundary making in activity animates the mathematical. this perspective is much more temporal and fluid, challenging the traditional approach of “acquiring” mathematical ideas. instead, mathematics becomes more an experience than reflection. de freitas and sinclair, drawing from barad and extending to mathematics via the work of the philosopher of mathematics gilles châtelet (2000),   chorney       | f l r     50   term their approach “inclusive materialism” indicating that their “theory of matter [...] resists the binary divide between human agency and inert passive matter” (p. 39). revitalizing materials often considered to be passive or inert challenges the idea of concepts, animating them in terms of materiality. de freitas and sinclair theorize mathematics concepts as engaging both the logical and ontological and argue the fusing and couplings of speech, movement and material items including the body, is the mathematics. in this dynamic relation where mathematics is animated by material tools, practice is affected. as an example, de freitas and sinclair (2014) draw on speech in a learning environment as an example of becoming in the moment. they argue that speech is not a medium to reflect what the thinker has already articulated in the mind, but instead organically flows in real time, as context, gestures and other factors are all active, changing and influencing. to ascribe speech a meaning of communication and/or representation relies too heavily on the meaning of words and to the intention of the speaker, but instead it should be seen as one of the many material factors that participate in a material assemblage2 of learning. these writers focus on activity, not as a process of acting based on what is already known, but as becoming and learning in movement. the tool and student, together, and what they can do, together, is what can be referred to as coupling (de freitas & sinclair, 2014) or fusing (barad, 2007). i refer to a hyphenated combination “student-tool” to specify this single entity. as such, the focus now is what the new entity can do as opposed to concern over what individuated objects bring to each other. mathematics, tools and students become a single focus, with a singular agency of production. the inclusive materialism, outlined by de freitas and sinclair describe the practice of mathematics as a material engagement of student, movement, tool. the mathematical concept is the intra-action, a human subject with the tool. the concept, therefore, partakes of the physical world; the concept is material. although a classroom can be seen as divided up discretely (ie. projectors, students), all movements, shifts, transformations are continuous and universes are being made known. there is neither origin nor final cause. there is only a weave of becoming. 7. towards post-human methods to address a method of observing students intra-acting with a digital tool is challenging when both ingold and barad argue there are no things a priori to activity. it is important to adopt a more accessible way to engage in an analysis and i draw on pickering who provides a methodology that identifies subject and object and offers an accessible way to engage in language about things. for example, one can identify resistance by looking for obstacles. such a task is not easy from a nonanthropocentric point of view since obstacles are typically framed in relation to what the student is trying to do. nonetheless, i suggest that it is always possible to counter the anthropocentric perspective by taking into consideration that the student must respond to the resistance—which is to say the material agency—of the tool. one can identify an obstacle by watching for sudden changes in action and attempting to decipher whether there was a significant reason for doing so. for example, if a student is using a compass and the graphite piece breaks, and the student subsequently stops using the compass, one can infer the breaking of the tool was an obstacle. to take another example, if in using a ruler to draw a line, the ruler slides slightly so that the angle of the line changes and the student stops drawing the line, then one can identify the sliding ruler as an obstacle. in these examples, the student is not able to perform an action—his or her motions or gestures are restricted. this noticing of obstacles offers a way of identifying the student’s active, in-themoment experiences in mathematical practice.                                                                                                                           2  assemblage is a notion introduced by deleuze and guattari, 1980, meaning an emergent unity joining together heterogeneous bodies in a “consistency.”   chorney       | f l r     51   while pickering’s model will be used as a way of engaging in observation the other theoretical lenses, materialism (drawing on barad as well as de freitas and sinclair) and a process-ontology (drawing on whitehead and ingold) will be mobilized as well. a process approach attends to how particular aspects of students and the tool come together and their role in performing an action together. one can pair a process approach with a materialist view noting that it is important not to assign any aspect of movement or action to a hidden phenomenon such as thinking, but rather only to that which is observable. process attends less to nouns and more to verbs; it is the relation and action between fused people-and-things that are the fundamental building blocks of reality (whitehead, 1978). 8. site of data collection data collection for this study took place at a high school in western canada involving a grade nine class of approximately thirty students. the grade nine class was chosen because the mathematics curriculum for this grade includes a comprehensive geometric component where students work with rotational and line symmetries, reflections, polygons and circle geometry. the students had studied basic polygons and their properties in previous grades. therefore, the pedagogical purpose of the activity was for the students to experience some common polygons as a form of review and renewal. the episode outlined in this paper occurred in a computer lab that the students had not visited before and where students were introduced for the first time to the geometer’s sketchpad (gsp) (jackiw, 2001). they were requested to construct a triangle and a square using gsp. consider the following two episodes: 8.1 episode 1: is this a triangle? two grade nine students, calvin and jonas, constructed a triangle by using the segment tool and also by constructing an interior. they then dragged the triangle all over the screen and when the triangle was half off the screen, one of the boys posed the question, “is this a triangle?” before addressing this interesting question, i will first look at the process of movement involving the body, the tool and the triangle and subsequently identify resistance within that movement. a b c figure.1. calvin’s triangle being translated around the screen figures 1 a, b, and c, are snapshots from a mathematical activity, in which calvin moved the triangle around the screen employing different motions, in different directions and at different speeds. these images, when seen as static hide the previous activity as well as in-between and future movements. in terms of their temporal existence, it might seem like they are static representations of the tool’s affordance; that the student was the agent and the tool passively allowed the student’s predetermined actions.   chorney       | f l r     52   if movement occurs before thought, as ingold insists, then the hand-action that glides the mouse across the desk, while consequently dragging the triangle’s vertex on the screen, fuses human and tool (hand and mouse) in a continuous flow. this flow of movement saw the triangle move upward, downward, to the left, and line up with the edge of the screen. human-centric theorists might describe this event in terms of calvin’s use of a “random” or “wandering” drag mode (arzarello et al., 2002), but this ascribes all the agency and intention to calvin (who may be wanting to see simply what happens or does not know what else to do). in contrast, through a movement lens, the mind is not observing and leading, it is participating, in parallel with the tools (and the size, as well as friction, of the desk). the movements translate the knowing and learning; the knowing is the translated triangle. after a short period of time, the movement translated into the question “is this a triangle?” to describe the question as separate from the boy’s activity is to miss the process of his movements as well as create a binary divide between thought and action. it theorizes a dual role for the student—one as the mover and the other as the observer—but only one role for the tool: allowing the student to act upon and question it. ingold (2011) reminds us that moving and knowing are the same process; similarly, roth and radford (2011) insist that thinking is not externalized but becomes. it is a challenge to shift from a paradigm that sees the student dragging the shape around the screen and attributing the intention, “oh i can drag the triangle off the screen”, to a paradigm that sees the thinking in the hand-mouse gesture. while calvin is limited in perception by the material tools, including a bordered screen, and in movement, by the shape of his hand and the desk, he nonetheless “feels” the ability of “hiding”, “cutting”, and “dragging off” part of the triangle. doing justice to the movement of a wayfarer requires honouring it as it unfolds, rather than seeing it as a path of transport towards pre-determined ends or as a result of a conscious, deliberate decision. in this episode, the movement seems to give rise to a kind of resistance, or an obstacle, which does not appear until the translating motion of the shape is limited by the edge of the screen. in other words, the edge of the screen only becomes interesting under motion, without it, it may not ever be noticed. when its boundary is approached, reached and then transgressed, it begins to intra-act with the moving hand-mouse, and becomes a participant in occasioning new activity. in this case, the transgression gives rise to a triangle that seems visually truncated at the end of the dragging, while retaining an integrity in time that warrants the name “triangle”. the question, “is this a triangle?” highlights an accommodation-like action. the resistance of the edge of the screen, which “hides” part of the triangle, and the question that emerges becomes part of a story, part of an action. a truncated triangle drawn on a whiteboard would probably not be considered a triangle, especially if the action of the drawing and/or gesturing hand, which may be extending off the edge is ignored. however, in the gesture-triangle-edge intra-action that produces the shape that used to be a triangle, the resistance cannot just be seen as being imposed by the edge of the screen, but also in the question itself. indeed, if it was a triangle, it should still be a triangle, and it would also be a triangle if the edge of the window were enlarged. and surely the gesturing hand is still holding the invisible vertex in place, even when moving it around so as to affect the visible sides of the triangle? 8.2 episode 2: the almost-square students were asked to make a square but often their shape did not hold under dragging, the students tried other methods to create more robust squares. what emerged from this process is what i term an almostsquare (figure 2). they may well have looked liked squares before, but the visible measurements exposed the second decimal place difference in the unequal lengths.   chorney       | f l r     53   figure.2. an almost-square before using the measure tool, the dragging hand was entangled with the watchful eye, as the quadrilateral changed on the screen. instead of focusing on the student’s intention and the tool’s response, the notion of assemblage draws our attention to the relations among the tool, the student and the concept. once the measuring tool comes into play, as well as the measurements lastingly being visible on the screen, the activity changed. the students began to drag the vertices of the square in order to adjust the measurements so that the side lengths would all match. but as soon as they dragged one vertex to match two sides, a third side would change in length and therefore no longer match the other sides. in response to this, a certain student pair then dragged the vertex associated with the distorted side ever so slightly, in an attempt to make all four sides match, only to find in movement that once again this led to a different kind of disfigurement. this back-and-forth dragging continued for a while. now the almost-square emerged out of the actions of the dragging hand, which changed the measurements on the screen, which provoke further dragging. the relation changed as the segments themselves faded away to give primacy to the numbers, which now dictate the extent to which the shape could be considered a square. the measuring became part of (and set in motion) an activity of adjustment, which can be seen as an accommodation that gave rise to a new assemblage. the resistance manifests itself through the unequal measurements, which nonetheless have the potential to be made equal (or so it may seem), which are related to the known properties of the square. the dragging of the vertex to try to “fix” the almost-square can be seen as an accommodation made by the hand and prompted by the resistance encountered in movement. the properties gain new meanings in that the previous way in which the sides could merely look the same now gets replaced with them all having to be equal in measure, one to the other. it is crucial here to recognize that the dragging is not controlled by the student, but spurred on by the changing measurements. for it is only when the measurements indicate a need for change did the vertex get dragged. accommodation in this episode is not simply overcoming the stubbornness of the tool in an intentional change of strategy; rather, it is about a changing assemblage in which new relations come to the fore, leaving previously central “actors” (like the segments) to recede while new ones (the measurements) emerge. throughout, the hand, tool, eye are in motion, in continual reconfiguration. the on-going struggle for equilibrium is initiated and promoted by this assemblage of technology and student. the process of how the almost-square came to be involves multiple steps. the software’s response, the movement of the hand–mouse, the pointing and the student’s readings of the measurements, all contribute to the movement from which the almost-square emerges. from a process perspective, the almostsquare only exists in the activity of movement; it is dependent upon the intra-action with the student in terms of the initial construction and subsequent “adjustments” and “fine-tunings”. if left alone, it might merely be termed a quadrilateral, but in ongoing movement, it remains an almost-square. this is because in a process approach, we read it not from its shape at any particular point in time but through the process in which it is engaged—which, in this case, is asymptotically becoming a square. therefore, within a process ontology, the almost-square does not exist as a noun, but as movement in a becoming state. 8.3 episode 1: resistance as relational while pickering’s constructs of resistance/accommodation are centered on the tool (which resists) and the user (who accommodates), my reading of the translation of the triangle showed that the resistance   chorney       | f l r     54   (the edge of the screen) can be seen as highlighting a certain relation—a calvin–dragging relation that encounters an unexpected event that would not exist if each were taken in isolation. in other words, the resistance in this case was a result of activity between the student–tool and not the student alone, nor the tool alone. in addition, the resistance was less about avoiding or overcoming a problem than it was about dwelling in the problem itself. as the triangle was dragged around the screen and eventually partly off the screen, the student–tool experience can be viewed as having provided an opportunity for the student–tool to experience something unexpected, which ended up flushing out a question—is this a triangle?—that reconfigured the relations among the triangle, the dragging and calvin. this prompts a question about when resistance might be seen as productive in a learning environment. the teacher could have insisted that the students restrict their dragging to the contours of the screen, in which case the boundary between triangle and not-triangle would have remained unchallenged. but this is the very boundary that the lakatosian practice of concept formation through monster-adjusting and monster-barring seeks to identify and stress. is the fact that a tool is involved in the process somehow less mathematical? or might we come to see calvin’s question as a legitimate one that perturbs boundaries imposed by the visible and the static in school geometry? in which case, accommodation is not to be seen as a singular response that fixes a transgression, but rather as a setting off of possibilities for engagement. 8.4 episode 2: from resistance as back-and-forth to resistance as narrative in contrast to episode 1, the resistance that was identified in the almost-square episode did not seem to be a result of the fused student–tool, but of fusing itself. that is, the resistance aligned more closely with pickering’s model of a back-and-forth between the student and the tool. this leads to the question of whether there are different levels or layers of resistance. while in episode 1 intra-action seemed to be in the student– tool relation, the identification of resistance in this episode highlighted the back-and-forth between tool and student potentially supporting a perspective that distinguishes between the tool and student. so the question emerges of how to deal with resistance in this kind of example without assuming an a priori individuation of relata. perhaps the back-and-forth interpretation is not as relevant in the context of instant feedback where the measurements change as the dragging changes. if both are changing simultaneously, then the temporal ordering implied by a back-and-forth interpretation is misleading. it encourages a reading in which the student makes an intentional choice based on the outcome of the dragging, which reasserts the primacy of the student over the tool. the back-and-forth activity does not entail a metaphor of transport, but one of wayfaring. it is not that student is moving from an almost-square to a square as might be seen in a transport but that in movement, in the back and forthing, the almost-square emerges. but what might be gained from focusing on the relation rather than the relata? from a baradian point of view, we are forced to reckon with the impossibility of separating the relata, which, in this case, encourages us to attend to the dragging/measuring process. this leads to different questions: instead of asking whether the student will ever succeed in making a square or whether the tool’s measurements will always prevent a square from happening, we ask what new meanings (about “square”) emerge? or, from an ingoldian perspective, we ask what new narratives emerged. indeed, ingold helps with this relata-versus-relation tension by drawing on the construct of things-as-narrative. from this perspective, there is no student, nor is there a tool, before the back-and-forth activity. the student–tool is, in and of itself, a process according to ingold. to identify movement within the student–tool entity is natural through an ingoldian process lens.   chorney       | f l r     55   9. conclusions the post-human starting point of this paper invites us to move away from the tendency to locate learning and thinking solely within the mind of the learner. because of the tendency to place the learner at the centre of thinking and learning, and to look for changes in how the learner talks, acts and moves, it can be challenging to study mathematical learning situations from a post-humanist point of view. by drawing on some of the main constructs offered in the literature, such as resistance/accommodation, assemblage, intraaction, process and tools as narrative, none of which have been explicitly designed to account for educational research, i have attempted to find out what new insights such post-humanist constructs can provide and what new questions they prompt. de freitas & sinclair (2014) attempt to mobilise the notion of assemblage by referring, for example, to the student-tool-concept as a whole. they also suggest ways both of attending to and of analysing data that re-materialise language and attempt to include broader material dimensions of activity, instead of simply attending to spoken words or to gestures—highlighting other sounds and rhythms that shape the ongoing activity. the construct of resistance/accommodation was fruitful in identifying turning points in the two episodes, but also potentially invited a human-centric view in which the tool ends up being subordinate to the human. the constructs of intra-action and assemblage were then brought in as starting points for overcoming a human-centric perspective. these new constructs drew attention to the changing nature of relations and, in turn, to the evolving assemblages involved in each episode. however, it is challenging to describe mathematical learning in terms of assemblages. we may be able to see difference, but know very little about whether the difference is mathematically relevant. the shift to the physical reconfigures what it means to do mathematics. as opposed to a theorizing of how the mind is learning, responding or interiorizing, this shift draws back to the “in-the-moment”, embodied experience of tooling, rather than using a tool—much as we talk about walking rather than using our legs. considering the material and temporal experiences is to study mathematics education in a posthuman and materialist way, but one, that as discussed, does have challenges. to think about mathematics in the classroom involving “tooling” rather than using tools, requires avoiding the tendency and tradition of isolating nouns, like “student”, “tool” and “concept” and locating meaning in their entanglement. this is particularly difficult when wishing to discuss particular aspects of an assemblage, as the tendency is then to detach that aspect or part from the whole. this approach demands that when discussing any part of an assemblage, the focus should be on the part’s role in the intra-activity, which is to say, how it acts as a verb and how it contributes to the assemblage’s movement. a process ontology, the actions of student as wayfarer, and especially the notion of tool as narrative, draw attention to movement and temporality. movement and temporality were vital in interpreting the exercise involving the triangle and the exercise involving the almost-square; the triangle and the almostsquare would not have emerged—would not have meaning—if only the student or only the tool were considered as isolated and inert nouns. it was the doing—the fused nouns participating in the temporal verb of acting—that was the mathematical practice. to look at any single movement of transport from one static point to another, within the process of the student–tool assemblage changes the meaning entirely: one might miss the triangle entirely if, for example, it was only seen when it existed partly-off screen; one might only see a quadrilateral (noun) instead of the in-action almost-square. in analysing mathematical practice through a relational and verb-oriented process ontology, in seeing the student–tool as a line of becoming, the significance of the embodied experience and the material world are animated, consequentially changing the understanding of mathematics teaching and learning. this study was motivated and mobilized by a reconfiguration of perspective that challenged presupposed objects and their subsequent interactions. this study did not specifically address the age old philosophical quandary of how material and mental combine but helped generate a shift to addressing how the fusing of student-tool reaches out into the larger classroom environment with questions like “is this a triangle?” or “is this a well-constructed square?” these questions emerge from a paradigm of intra-action   chorney       | f l r     56   and not solely a mental contemplation. i suggest there is significant difference in these two perspectives. in this study’s approach, mathematics is less a label for a specific type of practice than it is a name for a specific kind of coupling. references artigue, m. (2002). learning mathematics in cas environment: the genesis of a reflection about instrumentation and the dialectics between technical and conceptual work. international journal of computers for mathematical learning, 7:245-274. arzarello, f., olivero, f., paola, d., & robutti, o. (2002). a cognitive analysis of dragging practises in cabri environments. zentralblatt für didaktik der mathematik, 34(3), 66-72. balacheff, n. (1988). aspects of proof in pupils’ practice of school mathematics. in d. pimm (ed.), mathematics, teachers and children. p. 216-235. london: hodder & st. barad, karen. (2007). meeting the universe halfway: quantum physics and the entanglement of matter and meaning. durham, n.c.: duke university press. bartolini bussi, m. g., & mariotti, m. a. (2008). semiotic mediation in the mathematics classroom: artifacts and signs after a vygotskian perspective. in l. d. english (ed.), bussi, m. b., jones, g. a., lesh, r. a., sriraman, b. (assoc. eds.), handbook of international research in mathematics education. new york: routledge. châtelet, g. (2000). figuring space: philosophy, mathematics and physics. dordrecht, the netherlands: kluwer. de freitas, e. & sinclair, n. (2014). mathematics and the body: material entanglements in the classroom. cambridge university press. deutscher, g. (2010). through the language glass: why the world looks different in other languages. henry holt and co. new york. dietiker, l. (2013). mathematical texts as narrative: rethinking curriculum. for the learning of mathematics, 33(3), 14-19. goodchild, s. (2014). enactivist theories. in s. lerman (ed.), encyclopedia of mathematics education. heidelberg: springer. hesse, m. (1980). revolutions and reconstructions in the philosophy of science. bloomington: indiana university press. ingold, t. (2011). being alive: essays on movement, knowledge and description. london: routledge. latour, bruno. (1987). science in action: how to follow scientists through society. cambridge, ma: harvard university. jackiw, n. (2001). the geometer’s sketchpad. emeryville, ca: key curriculum press. lyotard, jean-françois. (1984). the postmodern condition: a report on knowledge. trans. geoff bennington and brian massumi. minneapolis: university of minnesota. nemirovsky, r., kelton, m. l. & rhodehamel, b. (2013). playing mathematical instruments: emerging perceptuomotor integration with an interactive mathematics exhibit. journal for research in mathematics education, 44(2), 372–415. pickering, a. (1995). the mangle of practice: time, agency, and science. chicago: the university of chicago press. roth, wolff-michael. (2011). passibility: at the limits of the constructivist metaphor. springer, netherlands. roth. w-m., & radford, l. (2011). a cultural historical perspective on mathematics teaching and learning. sense publishers. rotman, b. (2008). becoming beside ourselves: the alphabet, ghosts, and distributed human being. duke university press, london. ruthven, k. (2002). instrumenting mathematical activity: reflections on key studies of the educational use of computer algebra systems. international journal of computers for mathematical learning, 7:275-291.   chorney       | f l r     57   varela, f.j. & thompson, e. & rosch, e. (1991). the embodied mind: cognitive science and human experience (first mit press paperback edition, 1993) cambridge, massachusetts, london, england: the mit press. verillon, p. & rabardel, p. (1995). cognition and artifacts: a contribution to the study of thought in relation to instrumented activity. european journal of psychology in education, 9(3): 77-101. waltz, s. (2006). nonhumans unbound: actor-network theory and the reconsideration of “things”. in educational foundations, summer-fall. whitehead, a. n. (1978). process and reality. new york: the free press. frontline learning research vol. 5 no. 3 special issue (2017) 14 30 issn 2295-3159 neural correlates of visual perceptual expertise: evidence from cognitive neuroscience using functional neuroimaging andreas gegenfurtnera1, ellen m. kokb, koos van geelb, anique b. h. de bruinb, bettina sorgerc a institut für qualität und weiterbildung, technische hochschule deggendorf, germany b school of health professions education, maastricht university, the netherlands c department of cognitive neuroscience, maastricht university, the netherlands article received 7 may / revised 23 march / accepted 23 march / available online 14 july abstract functional neuroimaging is a useful approach to study the neural correlates of visual perceptual expertise. the purpose of this paper is to review the functional-neuroimaging methods that have been implemented in previous research in this context. first, we will discuss research questions typically addressed in visual expertise research. second, we will describe which kinds of stimuli are employed and which functional-neuroimaging techniques are implemented in this kind of research, with a special focus on electroencephalography (eeg) and functional magnetic resonance imaging (fmri). third, we will summarize the outcomes of recent studies that addressed the neural correlates of visual expertise and will particularly focus on studies that examined the neural correlates of visual expertise in medical image diagnosis. finally, the review closes with a discussion of the benefits, caveats, and future directions of cognitiveneuroscience research for studying visual expertise. keywords: perceptual expertise; eeg; fmri; n170; ffa. 1 corresponding author: andreas gegenfurtner, institut für qualität und weiterbildung, technische hochschule deggendorf, dieter-görlitz-platz 1, 94469 deggendorf, germany. email: andreas.gegenfurtner@th-deg.de doi: http://dx.doi.org/10.14786/flr.v5i3.259 http://dx.doi.org/10.14786/flr.v5i3.259 gegenfurtner et al | f l r 15 1. introduction expertise can be defined as maximal adaptations to task constraints (ericsson & lehmann, 1996; gruber, jansen, marienhagen, & altenmueller, 2010) which can take many forms, including, among others, motor expertise, memory expertise, or perceptual expertise (ericsson & lehmann, 1996). perceptual expertise can be further categorized as visual, auditory, tactile, olfactory, vestibular, or gustatory expertise. visual expertise is evident, for example, when bird experts classify a passing little bird as an oriole or a cardinal (tanaka & curran, 2001) or when clinicians diagnose digitized slides of human tissue as pathologically normal or abnormal (helle et al., 2011). assuming that individual differences in visual perceptual expertise should be reflected in differences in the brain, the following question arises: can we reliably measure/objectify neural correlates of visual expertise with currently available functionalneuroimaging methods and therewith explain inter-individual behavioral differences with respect to visual perceptual expertise? in line with the overall goal of this special issue to introduce and discuss methodological approaches in visual expertise research (gegenfurtner & van merriënboer, 2017), the purpose of the present methodological review is to reflect on the promises and pitfalls of cognitive-neuroscience methods in the study of visual perceptual expertise. while the review can offer input for discussions among scholars experienced in conducting neuroscientific studies, the manuscript is mainly written to inform scholars who are unfamiliar with the methodological repertoire of functional neuroimaging and its use in expertise research. in this review, we will particularly address expertise in medical image diagnosis, which can be defined as the inspection and interpretation of a visual representation of the human anatomy or its functions (gegenfurtner, kok, van geel, de bruin, jarodzka, szulewski, & van merriënboer, 2017); but because this body of research is still limited and in its infancy, we will extend our review to other content domains with the aim of offering a more useful overview of current methodological decisions in the visual perceptual expertise literature. there are already several systematic reviews available on the neural aspects of visual perceptual expertise (for example, richler & gauthier, 2014, for face perception or gegenfurtner, siewiorek, lehtinen, & säljö, 2013, for medical image diagnosis). the present review has a particular emphasis on implementing cognitive-neuroscience (especially functional-neuroimaging) methods on visual perceptual expertise, and will follow four steps. first, we will start with a short discussion of typically addressed research questions. second, we will describe which kinds of stimuli are employed and which functionalneuroimaging methods are used, with a special focus on the frequently implemented methods eeg and fmri. third, we will summarize the outcomes of studies that addressed the neural correlates of visual perceptual expertise. and finally, we close this review with a discussion of the benefits, caveats, and future directions of cognitive-neuroscience research for studying visual expertise. 2. research questions research on visual perceptual expertise has focused on a wide range of different research questions. these research questions can be clustered in three distinct types: contrastive, developmental, and conditional. naturally, research questions strongly correspond with the research design. for example, contrastive research questions ask how participants of different levels of expertise vary in different neural measures. in a classic study, haller and radue (2005) were interested in examining “neuronal activations during processing of radiologic and non-radiologic images by experienced radiologists and non-radiologist subjects by using event-related functional magnetic resonance (mr) imaging” (p. 983). this is a representative example for the first type of research questions (contrastive research questions). a second type of research questions, developmental research questions, asks how participants neurally adapt to visual perceptual training. these studies typically employ a paradigm implementing a training of inexperienced participants over the course of several weeks. for example, gauthier and colleagues (1998) were interested in examining if increased experience with so-called ‘greebles’, artificially created stimuli (see figure 1), gegenfurtner et al | f l r 16 would yield to an increase of fmri activation in a particular brain region, the so-called ‘fusiform face area’ (outcomes of this study presented and discussed below). finally, the third type of research questions, conditional research questions, addresses the extent to which expertise effects – be they contrastive or developmental – are contingent on task conditions such as the duration of stimulus presentation or different manipulations of the presented stimuli. for example, bilalić and colleagues (2016) were interested in unravelling if expertise effects are moderated by the orientation of the presented stimulus, in their case, xray films showed either in a normal, upright position or in an inverted position (rotated by 180°). comparability across studies depends on the used research question and design. when designing a cognitiveneuroscience study, one can follow a single research question or several research questions even from different research-question types (see above). typically, conditional research questions that address the moderating effect of stimulus or task conditions are often combined with contrastive or developmental research designs. 3. methodology in visual perceptual expertise research in this section, we review established methods implemented in cognitive-neuroscience studies in the field of visual perceptual expertise. we first describe frequently used artificial and naturalistic stimuli. we then look at the methodology of fmri and eeg, in particular, on what kind of information can be derived from fmri and eeg signals, and we also outline other, less frequently used techniques in cognitive neuroscience. figure 1. examples of smoothie, spikie, and cubie objects (open access from op de beeck et al., 2006). 3.1. stimuli 3.1.1. artificial stimuli artificially created stimuli are objects that have no common reference in the real world. this is a deliberate choice to avoid any confounding effects that may be induced from familiarity with the object. gegenfurtner et al | f l r 17 several groups of artificial stimuli have been introduced. some of those used more frequently are ‘smoothies’, ‘spikies’, and ‘cubies’, and, perhaps most prominently, greebles (mentioned above). smoothies, spikies, and cubies are matlab-generated classes of objects that “were designed to have different shape properties and to seem novel (i.e., they did not immediately suggest associations with everyday object categories” (op de beeck, baker, dicarlo, & kanwisher, 2006, p. 13025). figure 1 shows example smoothies, spikies, and cubies. these artificial stimuli were created with variations of different dimensions, so that participants need to process more than one location of the object to attain high rates of discrimination. greebles are objects specifically constrained to be similar to faces along several dimensions. figure 2 shows example greebles. greebles are photo-realistically rendered, three-dimensional, computergenerated objects that all share a common configuration. as gauthier, williams, tarr, and tanaka (1998) explain: “each greeble is made up of a vertically-oriented ‘body’ with four protruding ‘appendages’, from top to bottom, two ‘boges’ a ‘quiff’ and a ‘dunth’” (p. 2402). greebles come in two different genders (called “glip” and “plok”) and five families (called “galli”, “osmit”, “radok”, “samar”, and “tasio”). greebles have been used in a range of studies using both fmri and eeg. figure 2. examples of greeble objects in their two genders and five families (open source from: https://commons.wikimedia.org/wiki/category:greeble). 3.1.2. real-world stimuli in opposition to these artificial stimuli that share little to no resemblance with naturally occurring objects, researchers also use real-world stimuli. these stimuli are called “real-world” to indicate that these objects are not researcher-generated. associated with real-world stimuli is the assumption that there are realworld experts that have developed visual skills related to these objects (shen, mack, & palmeri, 2014), so these material are used in an attempt to create ecologically valid domain-specific tasks. real-world stimuli can be classified as faces and non-face objects. first, photographs of faces (or of parts of faces) are extensively used as stimuli in visual perceptual expertise research because we have so much exposure to faces that this makes us all experts in face recognition (bentin, allison, puce, perez, & mccarthy, 1996; richler & gauthier, 2014). second, photographs of non-face objects includes cars (gauthier, skudlarski, gore, & anderson, 2000), different animal species such as birds (tanaka, curran, & sheinberg, 2005) or dogs (tanaka & curran, 2005), and also letters such as japanese (maurer, zevin, & mccandliss, 2008) and chinese characters (fan, chen, zhang, qi, jin, wang, et al., 2015; qi, wang, hao, zhu, he, & luo, 2016). researchers also use representations of chess positions (bilalić, langner, ulrich, & grodd, 2011) and medical images (haller & radue, 2005). in many studies, these real-world stimuli are presented either in gegenfurtner et al | f l r 18 original form or in inverted, rotated, or otherwise artificially distorted. the rationale behind these artificial manipulations is to complicate and change pattern recognition for expert participants. typically, real-world stimuli in these studies are static and two-dimensional. if we assume that the comprehension of visualizations is moderated by variations in dimensionality and dynamics (for a meta-analysis testing this assumption, see gegenfurtner, lehtinen, & säljö, 2011), then it seems surprising that the literature on the neural correlates of real-world visual perceptual expertise has not yet systematically compared how the brain processes of experts and novices differ when they view static vs. dynamic stimuli or two-dimensional vs. three-dimensional visualizations. 3.2. apparatus while viewing different kinds of stimuli, participants’ neural correlates can be measured with cognitive-neuroscience techniques. measuring these neural correlates is contingent on the study interests and research questions. typically, if researchers are interested in the temporal aspects of image processing, they use electroencephalography (eeg). conversely, if researchers are interested in the spatial aspects of image processing, they use functional magnetic resonance imaging (fmri). in addition to eeg and fmri, there are also several other measurement techniques, including magnetoencephalography (meg), positron emission tomography (pet), and functional near-infrared spectroscopy (fnirs). offering detailed descriptions of each of these techniques is beyond the scope of this review. ward (2006) and squire and colleagues (2013) offer easy-to-understand introductions. but it is informative here to briefly describe the two most frequently used techniques to illustrate how they work and what they measure. these are eeg and fmri. 3.2.1 eeg neurons communicate through electrical signals transmitted along axons and dendrites. when populations of neurons that are oriented in parallel are synchronously active, their electrical signals can be measured with electrodes placed on the scalp. electroencephalography (eeg) records and amplifies these electrical signals over time. when we perceive a picture, particular populations of neurons in our brain respond to this picture. this response is measurable as a change in voltage at the scalp before, while and after seeing the picture. if we average the recorded eeg signal across many trials, random brain activity that is unrelated to the neural processing of the picture is cancelled out. the relevant (stimulus-related) signal is preserved and called the ‘event-related potential’ (erp). when recording eeg from participants while they looked at pictures of faces, bentin and colleagues (1996) found a negative event-related potential (n) that reached its maximum at approximately 172 ms (n170) after picture onset. since this pioneering study, the n170 has become a widely studied eeg component in cognitive neuroscience. eeg measures have a high temporal resolution and are therefore time-sensitive. thus, eeg can be especially used to investigate temporal patterns of brain activity. however, eeg has a relatively low spatial resolution meaning that the localization of the eeg signal source (i.e., the location of the specific neuronal populations evoking the electrical brain activity) cannot be ascertained with high precision. 3.2.2 fmri fmri indirectly measures neural activity through its vascular response: following (e.g., visual) stimulation, neuronal activity in particular (e.g., visual) brain regions increases which results in enhanced local oxygen consumption. neuronal tissue gets new oxygen from the oxygenated hemoglobin in the blood. within a few seconds, the blood flow and the concentration of oxygenated hemoglobin in the blood increases in the particular brain region. this increase is called the hemodynamic response. since oxygenated and deoxygenated hemoglobin have different magnetic properties, the hemodynamic (or the blood oxygenation level-dependent, bold) response can be imaged using fmri. note that the hemodynamic response is considerably delayed and expanded which puts some constraints when designing fmri experiments. compared to eeg, the temporal resolution of fmri is rather low (one data point is normally obtained within 1-2s). however, the spatial resolution is considerably higher (in the mm3 range) meaning that fmri can gegenfurtner et al | f l r 19 provide specific information about the origin of the brain signal and therewith information about which part of the brain is involved in a particular activity (e.g., visual perception). note that fmri only measures relative (and not absolute) changes of the oxygenation level and that fmri visualizations are actually representations of statistical differences of the fmri signal across different experimental conditions. in summary, eeg has a very high temporal resolution and is therefore a suited method to investigate timing of brain activity. fmri has a much higher spatial resolution than eeg and is an appropriate method to indicate which brain regions are involved in a particular (e.g., perceptual) task. 4. results of visual expertise research this section presents the outcomes of studies addressing neural correlates of visual perceptual expertise. how does the development of expertise change temporal and spatial aspects of information processing? first, we summarize the findings of eeg research. second, we review the fmri findings. and finally, in a special section, we zoom in on the relatively new field of cognitive-neuroscience research applied to medical image diagnosis. 4.1. eeg research: the n170 using erps based on eeg measurements, cognitive-neuroscience research has provided strong support for the idea that a particular early erp component, namely the n170 (introduced above), plays a significant role when participants process photographs and pictures of faces (bentin et al., 1996; for a metaanalysis of this research, see hinojosa, mercado, & carretié, 2015). interestingly, it could be demonstrated that patients who suffer from face blindness, also called prosopagnosia (the inability to recognize faces), did not show this larger magnitude of the n170 component when processing faces (for reviews, see richler & gauthier, 2014; towler, fisher, & eimer, 2017). this body of evidence on face processing has inspired research on perceptual expertise because of the assumption, in part, that all humans are ‘face experts’. if the n170 was such a stable neurophysiological marker in face perception, would the enhanced n170 also reflect expert processing of other familiar, domain-specific objects? a pioneering study by tanaka and curran (2001) confirmed this hypothesis. eeg was recorded while participants viewed photographs of cars or birds. approximately 164 ms after stimulus onset, participants who were car experts showed a larger n170 component for cars compared to birds, and participants who were bird experts showed a larger n170 component for birds compared to cars. tanaka and curran (2001) carefully controlled for stimulus artefacts including image properties and task instruction, and also for group effects in that the same participants viewed photos of cars and birds and were thus expert and novice in different trials of the experiment. in summary, this study revealed that visual perceptual expertise is associated with an enhanced n170 component and therewith with very early stages of visual information processing. in recent years, this effect has been replicated with both car (gauthier & curby, 2005; scott, tanaka, sheinberg, & curran, 2008) and bird stimuli (scott, tanaka, sheinberg, & curran, 2006; tanaka et al., 2005). research also disclosed the expertise effect on n170 using artificial stimuli including blobs (curran, tanaka, & weiskopf, 2002) and greebles (rossion, gauthier, goffaux, tarr, & crommelinck, 2002; rossion, kung, & tarr, 2004) and with non-object letter symbols, including japanese (maurer et al., 2008) and chinese characters (fan et al., 2015; qi et al., 2016). in summary, eeg studies suggest that, similar to face perception (hinojosa et al., 2015; richler & gauthier, 2014; towler et al., 2017), visual expertise modifies the temporal aspects of information processing related with an enhanced n170 component for trained or domain-specific objects. gegenfurtner et al | f l r 20 4.2. fmri research: revealing the functional role of the ffa in 1997, kanwisher and colleagues located a brain region in the fusiform gyrus that is strongly activated when humans view faces. this region was called the fusiform face area (ffa). two years later, in 1999, gauthier, tarr, anderson, skudlarski, and gore demonstrated that the ffa is not only activated when viewing faces but also indicates the level of expertise with artificial objects (in this case greebles). the assumption was that the selectivity of ffa reflects a more generalized form of visual perceptual expertise that is not intrinsically specific or restricted to processing face stimuli (tarr & gauthier, 2000). this assumption was confirmed with bird (gauthier et al., 2000), car (gauthier et al., 2000), and artificial stimuli (gauthier et al., 1999). however, an early criticism of these studies was that these stimuli were similar to faces: indeed, parts of greebles evoke resemblance to faces, birds have faces, and also cars, at least in threequarter frontal views, resemble faces (kanwisher, 2000; grill-spector, knouf, & kanwisher, 2004). the conclusion was, thus, that ffa activation was more likely the result of face similarity than object expertise. to minimize the effect of faces, xu (2005) used side view photographs of birds and cars, and reported that visual perceptual expertise was still associated with ffa activation. since then, a rich plethora of fmri studies supported gauthier’s initial assumption that visual expertise in object perception was associated with activation in the ffa independent of face similarity (bilalić et al., 2011; bukach, gauthier, & tarr, 2006; palmeri & gauthier, 2004; righi, tarr, & kingon, 2013; but see bartlett, boggan, & krawczyk, 2013, for a study that did not find differences between experts and novices in ffa activation in a chess task. in that study, artificially inverted and distorted chess stimuli were used, so it is a matter of debate if these stimuli were suitable to trace chess expertise). in recent years, the discussion around face selectivity tended to be replaced with a more recent discussion whether ffa was the only region relevant for processing familiar objects or whether visual perceptual expertise was associated with the interaction between different brain regions (e.g., bilalić, langner, campitelli, turella, & grodd, 2015; harel, kravitz, & baker, 2013; wong & wong, 2014). in short, there seems to be broad consensus in the field that the processing of objects involves more than just ffa. specifically, wong and wong (2014, p. 308) explain that “perceptual expertise researchers have been considering the interaction between perceptual and cognitive processing as an important component in understanding perceptual expertise for different objects. it is therefore unnecessary to create the debate between the so-called “perceptual view” and “interactive view” of expert object recognition, as the interaction between perceptual and cognitive processing has been well accommodated in perceptual expertise research.” overall, studies using fmri demonstrated that when experts view domainspecific stimuli, the ffa and other brain regions are activated. the precise location of these “other” brain regions and their particular interaction patterns with ffa, however, are still under investigation. gegenfurtner et al | f l r 21 table 1 studies examining neural correlates of expertise differences with medical images as stimuli first author (year) participants stimulus task (apparatus) main findings bilalić (2016) 16 radiologists, 15 students upright or inverted chest x-ray films, photographs of faces, rooms, and tools viewing task: 1-back task (fmri) expertise effect in ffa fiorio (2010) 8 clinicians, 10 students photographs and videos of healthy and dystonic writing movement decision task: judgment if and to what extent writing was dystonic (tms) corticospinal activation in students, but not in experts haller (2005) 12 radiologists, 12 laypersons original and manipulated radiologic images detection task: finding manipula tions on the image (fmri) increased activa-tion in temporal and frontal gyri in radiologists harley (2009) 7 radiologists, 6 4th-yr residents, 7 1st-yr residents normal and abnormal chest x-ray films detection task: finding nodules (fmri) positive correlation of ffa activity with expertise melo (2011) 25 radiologists chest x-ray films that included lesions, animals, or letters detection task: finding lesions, animals, or letters (fmri) activation in left inferior frontal sulcus and posterior cingulate cortex ribas (2013) 29 radiologists veterinary x-ray films decision task: choose between four diagnosis options (eeg) positive correlation of expertise with electrode activity c4, f3, f8, oz, t6 4.3. visual expertise in medical image diagnosis gegenfurtner, siewiorek, lehtinen, and säljö (2013) reviewed the literature on visual expertise in relation to medical image diagnosis and identified three of 21 studies that examined neural correlates of expert-novice differences when inspecting medical visualizations (haller & radue, 2005; fiorio et al., 2010; harley et al., 2009). since the review of gegenfurtner et al. (2013), two additional studies were published that addressed the neural basis of visual perceptual expertise in medicine (bilalić et al., 2016; ribas et al., 2013). we briefly review these studies here, together with an additional paper (melo et al., 2011) that examined the neural correlates of radiologists’ diagnoses. although melo et al. (2011) did not analyse the effect of expertise, the study is relevant in the current context as it discusses the involvement of brain regions when deriving diagnoses from medical visualizations. table 1 offers an overview of the six studies. bilalić and colleagues (2016) asked radiologists and medical students to indicate if the current stimulus they were seeing was the same as the previous one. stimuli were chest x-ray films that were either presented in upright position or rotated by 180° (inverted), as well as stimuli including photographs of faces, rooms, and tools. the findings suggest that the ffa of radiologists compared to medical students was more sensitive in differentiating upright or rotated x-ray films from the photographs showing rooms and tools. bilalić et al. (2016) conclude that the ffa activation was likely associated with the level of participant expertise effect. also harley and colleagues (2009) found a positive correlation between ffa activation and the visual expertise of radiologists. however, harley et al. (2009) also reported that activity in the right ffa did not differ between radiologists and first-year residents looking at radiological images. haller and radue (2005) presented radiologic images (computer tomography scans, magnetic resonance images, and ultrasound pictures) that were either original or manipulated to radiologists and non-radiologists. the participants were asked to decide if the presented images were original or manipulated. the group of gegenfurtner et al | f l r 22 radiologists showed significantly stronger activation than the group of non-radiologists in the bilateral middle and inferior temporal gyrus, bilateral medial and middle frontal gyrus, and left superior and inferior frontal gyrus—regions that are allegedly associated with visual attention and memory retrieval (wager & smith, 2003). haller and radue’s (2005) study is interesting because it is the first to indicate that different brain regions interact when experts visually process medical images. the findings of melo et al. (2011) and ribas et al. (2013) further support this notion. particularly, using eeg, ribas and colleagues (2013) report that participations with higher levels of expertise had more electric activity compared to participants with lower levels of expertise. fiorio et al. (2010) used transcranial magnetic stimulation (tms) to examine how participants differed when viewing photographs and short video sequences of healthy and dystonic writing. briefly, in tms, a magnetic field generator is placed in close proximity to the head of a participant in order to evoke electric currents in brain areas (for introductions to tms, see walsh & cowey, 2000; ward, 2006). the authors showed that “observation of pathological actions differently modulates the viewer’s motor resonant system, depending on previous knowledge, visual expertise, and ability to recognize sub-optimal movement kinematics” (p. 698). fiorio and colleagues (2010) used dynamic stimuli, which is still rare in the field of cognitive-neuroscience methods applied to medical diagnosis. on the basis of the studies reviewed here, it seems safe to conclude that expertise in medical diagnosis cannot be located in and isolated to a single brain area but, instead, expertise seems to be associated with changes in activation in a multitude of neural regions as a function of experience, amount of training, and knowledge structures. we should note, however, that this interpretation is contingent on the level of task complexity in the original studies. it seems likely that more brain regions are activated when the task is more complex, while many studies employ simplified versions of the task of medical image diagnosis. the six studies reviewed in table 1 differ in their task complexity. the complexity of the employed tasks was categorized following the four-level model of task complexity in the comprehension of visualizations (gegenfurtner et al., 2011) shown in table 2. this model defines task complexity on the basis of contextual demands that differ as a function of the number of desired outcomes, the multiplicity of paths to attain desired outcomes, and the coordinative complexity of informational cues in the task material while moving toward task completion. the reviewed studies include one viewing task, in which participants had to say if they had just seen the same image (bilalić et al., 2016); three detection tasks, in which participants were asked to search for an abnormality or specific target within the image (haller & radue, 2005; harley et al., 2009; melo et al., 2011); and two decision tasks, in which the participants had to choose among a given set of options (fiorio et al., 2010; ribas et al., 2013). the study by fioro et al. (2010) is the only one using tms and the study by ribas et al. (2013) the only one using eeg; thus, findings from these two studies cannot easily be compared to the other four studies using fmri. somewhat surprisingly, to date, no study has asked participants to produce a full diagnosis from a presented visual material, perhaps because tasks inside a magnetic resonance scanner are kept deliberately simple and diagnostic problem-solving tasks would be too complex; even a simple blink causes severe artifacts in electroencephalograms. typing is practically impossible, and speaking might be hard to record due to the noise made by the scanner. furthermore, if the task is too complex in comparison to the control task, this might lead to differences in brain activation that are so widespread that it might no longer be able to meaningfully interpret them. this explains perhaps the scarcity of cognitive-neuroscience studies in medical image diagnosis relative to its wide application in the visual perceptual expertise literature. gegenfurtner et al | f l r 23 table 2 four-level model of task complexity in the comprehension of visualizations (adapted from gegenfurtner, lehtinen, & säljö, 2011) task type multiplicity of solution paths number of desired outcomes coordinative complexity example viewing task low low low looking at medical images detection task low low high searching for lung nodules decision task low high high deciding between given options problem-solving task high high high generating a diagnosis if we compare the studies using medical images as experimental stimuli (reviewed in table 1) with the wider visual perceptual expertise literature and their findings of ffa activation and the enhanced n170 component, it is evident that the expertise effect in ffa was partially confirmed with medical images (bilalić et al., 2016; harley et al., 2009). an increased n170 component has not yet been systematically addressed. it is very encouraging that, since our review some years ago (gegenfurtner et al., 2013), more and more cognitive-neuroscience studies using fmri or eeg emerge that address medical image diagnosis. we do expect that future research will proliferate in this area in an attempt to replicate ffa activation and n170 enhancement as neural correlates of visual expertise in the medical domain. these studies will help us understand how medical expertise changes temporal and spatial brain activation patterns associated with the diagnosis of medical visualizations. 5. discussion after reviewing typical research questions and experimental stimuli, describing eeg and fmri, and reporting the current state of the neural correlates of visual perceptual expertise, this section will now elaborate on the advantages and limitations of cognitive-neuroscience research in the current context. what are the benefits of using fmri and eeg? what are caveats of these methodologies? and what are directions for future research that originate from this review? 5.1. benefits the benefits of applying cognitive-neuroscience methods in research on visual perceptual expertise relate to the extension of behavioural research, high temporal and spatial sensitivity, and high levels of control. we elaborate on each of these benefits in turn. first, cognitive neuroscience can extend behavioural research. in particular, cognitive neuroscience affords different units and levels of analysis; these, in turn, make visible some of the neural correlates underlying cognitive processes that would not be accessible with behavioural measures (ansari, de smedt, & grabner, 2012). framing this triangulation, stern and schneider (2010) introduced the metaphor of a digital road map: with cognitive neuroscience, researchers can zoom in to the neural levels of cognition and perception and examine processes inside the human brain. if researchers are interested in these processes, eeg and fmri offer measures that can unveil neural activation as the basis for observable expertise differences (gegenfurtner et al., 2013; gruber et al., 2010). gegenfurtner et al | f l r 24 another advantage of many cognitive-neuroscience methods is their very high temporal and spatial resolution. more precisely, eeg takes measures in the range of milliseconds. thus, if we are interested in the temporal aspects of expert performance, then time-sensitive eeg is a very suited method, especially if measured in parallel with pupillometry (szulewski, gegenfurtner, howes, sivilotti, & van merriënboer, 2017) and eye tracking (holmqvist, nyström, anderson, dewhurst, jarodzka, & van de weijer, 2011; jarodzka, jaarsma, & boshuizen, 2015). in contrast, fmri has a unique capability in locating very precisely the brain regions that are active, e.g., when participants of varying levels of expertise interpret complex images. if research aims to uncover where and when neural activity occurs during expert performance, then eeg and fmri (at best in combination) are two very powerful, non-invasive methodological tools. finally, because of the extremely high sensitivity of eeg (temporal) and fmri (spatial), experiments in cognitive neuroscience are typically very controlled. these levels of control afford high levels of external validity (ansari et al., 2012; de smedt, 2014). because researchers invest a considerable amount of time and energy in securing experimental control, including a strict selection of participants (for example: only righthanded people) and carefully filtered stimulus materials (exemplarily reflected in the huge effort of creating greebles), findings from eeg and fmri often result in stable, generalizable inferences. these generalizable inferences can inform researchers when developing theories of visual expertise. because cognitive-neuroscience methods have high levels of temporal and spatial sensitivity, as well as experimental control, neural correlates of visual expertise can be used in theory testing and development (bilalić et al., 2015). neuroscientific findings thus have the potential to inform expertise research in two ways. first, they can be used to test the predictive validity of existing models and theories, for example on how expertise develops in novices (kok, de bruin, robben, & van merriënboer, 2012; van geel, kok, dijkstra, robben, & van merriënboer, 2017), intermediates (boshuizen & schmidt, 1992; ericsson & lehmann, 1992), and experts (gegenfurtner, 2013; gegenfurtner, nivala, lehtinen, & säljö, 2009). second, they can be used to develop novel theories to account for expertise differences revealed by methods of cognitive-neurosciences; differences that would have remained unobservable with behavioral methods alone (bilalić et al., 2015). 5.2. caveats no method comes without limitations. powerful and elegant as cognitive neuroscience may appear, its methodology also includes different costs that can compromise the available evidence. caveats include the temporal and spatial resolution, ecological validity, a reductive bias, and limited implications for educational practice. first, and perhaps surprisingly, the extreme sensitivity of eeg and fmri measures, positive on one side, introduces of course a number of limitations to the experimental setup. for example, in eeg research, already the slightest motions like a blink or moving the nostrils creates severe data artifacts. research projects thus often lose a considerable amount of data because participants had not been motionless enough while their neural activity was recorded (ansari et al., 2012; de smedt, 2014). this is particularly detrimental if we consider the financial costs of data collection and if we consider that cognitive neuroscience often works with small sample sizes. to cover for this data loss, even higher and stricter experimental controls are developed, implemented, and employed. the fact that eeg and fmri are so sensitive to small motions causing artefacts means that the ecological validity is easily compromised: the sensitivity to motion restricts the possibilities for ecologically valid experiments. this is related to compromises in the ecological validity of cognitive neuroscience experiments. experts are not typically motionless or work inside the tube of a 3-tesla mri scanner. it is thus a matter of debate to which extent cognitive neuroscience can reflect the complexity of processes and practices associated with real-world expertise. do we force experts to act in too artificial ways? can we capture how experts diagnose a patient case when we show them a chest x-ray in a rotated, blurred, or otherwise distorted mode, for the duration of only a few seconds? the high level of experimental control, that is clearly a benefit of cognitive neuroscience, comes at the same time with limitations to ecological validity. the limited gegenfurtner et al | f l r 25 ecological validity is also associated with tasks that are typically used. instead of complex problem-solving tasks that would reflect medical diagnosis, many of the reviewed studies employed lower levels of task complexity (gegenfurtner et al., 2011). furthermore, study participants are asked to complete these tasks repeatedly in longer sessions to get readable signals, which can further compromise ecological validity. this is in line with de smedt’s (2014) observation that tasks used in neuroscience “need to be very elementary, because the larger the number of cognitive processes in a particular task, the more difficult it will be to disentangle these cognitive processes physiologically.” cognitive neuroscience is interested in the neural correlates of behavior, cognition, emotion etc. (squire et al., 2013; stern & schneider, 2010; ward, 2006). epistemologically, from a neuroscience perspective, visual perceptual expertise tends to be reduced to changes in electrical activity or blood flow. while this can render fascinating findings, other important ingredients of expert performance are ignored. certainly, all research is reductionist (lehtinen, 2012; säljö, 2009). one must make decisions what to measure because we simply cannot account for all relevant aspects in a single study, as interesting these aspects may be (damşa et al., 2017). focusing on neural levels does not imply that we uncover the basis of human learning. one could easily argue that the basis is the social context within which we are situated (gegenfurtner & szulewski, 2017; säljö, 2009). in describing this reductive bias, lehtinen (2012) notes: “because of the impressive technical development of brain research during the last two decades (…) many neuroscientists have quite a strong tendency towards downwards reductionism (emphasis in the original). this reductionism stems from the idea that research registering brain processes with complex technical tools finally opens up a real scientific approach to learning research.” cognitive neuroscientists are well aware that eeg and fmri are just two among the many other methods of learning research (e.g., de smedt, 2014). this reductive bias is not only associated with limitations in how expertise is measured and methodically approximated; it also signals limitations in how expertise and performance are theorized and conceptually framed (lehtinen, 2012; säljö, 2009; siewiorek & gegenfurtner, 2010). more specifically, theories of visual expertise that are exclusively grounded on neuroscientific evidence risk to de-emphasize other facets of how expert performance is enacted and displayed in real-world activities and practices (gibson, 1986; goodwin, 1994; see also de bruin, 2017; gegenfurtner et al., 2017). this risk is of course inherent in all mono-method designs, largely because single method studies capture a limited number of units of analysis. conversely, the combination of approaches in mixed-method or multi-method designs allow for the triangulation of units of analyses, which can inform a theory of visual perceptual expertise that encompasses different analytic levels beyond what is evident from single method approaches like eeg, tms, or fmri. while the benefits of bridging methods of expertise research are clear, methods are always part of a scientific community. these communities have agency as political actors and “defend” their methods against the influences of concurrent academic realms (al lily, foland, stoloff, gogus, erguvan, awshar, et al., 2017), so it will be an interesting observation to see if and to what extent expertise researchers will (continue to) embrace methodological triangulations and combine cognitive neurosciences with other method approaches in their studies for the purpose of theory development. related to that is a false belief that findings from eeg and fmri would be directly applicable and informative for re-designing learning environments and curricula. educational neuroscientists work hard to deemphasize the hopes that many practitioners have when they read that finally, once we understand the brain, we understand how learning, expertise, and education “work”. ansari and colleagues (2012) write that “the most obvious question a teacher may ask is, ‘how will i be able to apply this knowledge?’ there is, in our view, no reason to expect that neuroimaging research, will determine directly how teaching should take place. this is considered by many ‘a bridge too far’”. thus, cognitive or educational neuroscience may have a very limited impact on educational practices. does this mean we should not conduct this kind of research? certainly not; but we should, perhaps, rethink our expectations about what neuroscience measures can do (ansari et al., 2012; de smedt, 2014; lehtinen, 2012; säljö, 2009; stern & schneider, 2010). gegenfurtner et al | f l r 26 5.3. directions for future research examining the neural correlates of visual expertise is a fascinating endeavour. this review has identified a small, still limited number of studies that examined how visual perceptual expertise in medical image diagnosis correlates with eeg, fmri, and tms measures. what are directions for future research that follow from this review? first, all but one of the reviewed studies in table 1 used static pictures. only fiorio and colleagues (2010) used video sequences as stimuli. we thus recommend exploring and testing if and to what extent neuroscience-based visual perceptual expertise research can use dynamic stimuli. second, future research can make more use of eeg as well as other neuroscience approaches such as tms or meg to study the neural correlates of visual expertise in medical image diagnosis. another possible direction for future research is the combination of neuroscience methods with other online measures of expertise, including eye tracking and pupillometry (holmqvist et al., 2011; gegenfurtner & seppänen, 2013; kok et al., 2012; szulewski et al., 2017) if the constraints of different temporal scales can be accommodated for. such combinations would be interesting theoretically as a means to inquire how eye movements and neural activity correlate in expert diagnostic reasoning. fourth, implications of cognitive-neurosciences for education and training need to be explored. to what extent can clinical practitioners and medical educators benefit from neuroscientific measures? this is a question that applies to the field of medical image perception more generally and is not exclusive to cognitive-neurosciences; for example, also eye tracking used to be criticized for not being relevant enough to medical education and training, but has demonstrated its benefits in the form of eye movement modeling examples (jarodzka, balslev, holmqvist, nyström, scheiter, gerjets, et al., 2012; seppänen & gegenfurtner, 2012). it remains to be seen in future research if, and how, a similar approach can be developed for functional imaging. we should note, however, that cognitive-neurosciences are useful methods in addition to instructional design studies: while design studies reveal what works, neural correlates can indicate why it works (gegenfurtner et al., 2013; kok, van geel, van merriënboer, & robben, 2017). finally, eeg and fmri are measures into the temporal and spatial configurations of visual perceptual expertise. these measures should be incorporated into existing theory frameworks of visual perceptual expertise to advance our conceptual understanding of how experts, intermediates, or novices comprehend medical visualizations. 6. conclusion as noted at the outset, if we assume that individual differences in visual expertise are reflected in differences in the brain, then cognitive neuroscience methods can be used to examine the neural correlates of the experts’ visual skills. these methods can complement other methodologies interested in how experts in medical disciplines form their diagnoses. this review summarized research on visual perceptual expertise and described which research questions were typically asked, which stimuli and functional neuroimaging methods were frequently used, and how experts and novices differ in their neural representations (e.g., with respect to activation within the ffa). we also outlined some of the benefits, limitations, and future directions of cognitive-neuroscience research as they apply to the comprehension of (medical) visualizations. this methodological review closes with the hope that interested researchers, who are perhaps yet inexperienced with cognitive neuroscience, will find this paper a useful introduction into the neural correlates of visual expertise. keypoints cognitive neuroscience can uncover the neural correlates of visual perceptual expertise electroencephalography can reveal temporal adaptations of expertise (e.g., n170) gegenfurtner et al | f l r 27 functional magnetic resonance imaging can reveal spatial adaptations of expertise (e.g., ffa) cognitive neuroscience examining expertise in medical image diagnosis is promising but still in its infancy eeg and fmri can complement and extend each other as wells as other methodologies in expertise research references al lily, a., foland, d., stoloff, d., gogus, a., erguvan, i. d., awshar, m. t., et al. (2017). academic domains as political battlegrounds: a global enquiry by 99 academics in the fields of education and technology. information development. doi:10.1177/0266666916646415 ansari, d., de smedt, b., & grabner, r. (2012). neuroeducation a critical overview of an emerging field. neuroethics, 5, 105-117. doi:10.1007/s12152-011-9119-3 bartlett, j., boggan, a. l., & krawczyk, d. c. (2013). expertise and processing distorted structure in chess. frontiers in human neuroscience, 7, 825. doi:10.3389/fnhum.2013.00825 bentin, s., allison, t., puce, a., perez, e., & mccarthy, g. (1996). electrophysiological studies of face perception in humans. journal of cognitive neuroscience, 8, 551-565. doi:10.1162/jocn.1996.8.6.551 bilalić, m., grottenthaler, t., nägele, t., & lindig, t. (2016). the faces in radiologic images: fusiform face area supports radiological expertise. cerebral cortex, 26, 1004-1014. doi:10.1093/cercor/bhu272 bilalić, m., langner, r., campitelli, g., turella, l., & grodd, w. (2015). editorial: neural implementation of expertise. frontiers in human neuroscience. doi:10.3389/fnhum.2015.00545 bilalić, m., langner, r., ulrich, r., & grodd, w. (2011). many faces of expertise: fusiform face area in chess experts and novices. journal of neuroscience, 31, 10206-10214. doi:10.1523/jneurosci.5727 -10.2011 boshuizen, h. p. a., & schmidt, h. g. (1992). on the role of biomedical knowledge in clinical reasoning by experts, intermediates, and novices. cognitive science, 16, 153-184. doi:10.1207/s15516709cog1602_1 bukach, c. m., gauthier, i., & tarr, m. j. (2006). beyond faces and modularity: the power of an expertise framework. trends in cognitive science, 10, 159-166. doi:10.1016/j.tics.2006.02.004 curran, t., tanaka, j. w., & weiskopf, d. m. (2002). an electrophysiological comparison of visual categorization and recognition memory. cognitive, affective, and behavioral neuroscience, 2, 1-18. doi:10.3758/cabn.2.1.1 damşa, c. i., froehlich, d. e., & gegenfurtner, a. (2017). reflections on empirical and methodological accounts of agency at work. in m. goller & s. paloniemi (eds.), agency at work: an agentic perspective on professional learning and development. new york: springer. de bruin, a. b. h. (2017). the potential of neuroscience for health sciences education: towards convergence of evidence and resisting seductive allure. advances in health sciences education. doi:10.1007/s10459016-9733-2 de smedt, b. (2014). advances in the use of neuroscience methods in research on learning and instruction. frontline learning research, 6, 7-14. doi:10.14786/flr.v2i4.115 fan, c., chen, s., zhang, l., qi, z., jin, y., wang, q., et al. (2015). n170 changes reflect competition between faces and identifiable characters during early visual processing. neuroimage, 110, 32-38. doi:10.1016/j.neuroimage.2015.01.047 fiorio, m., cesari, p., bresciani, m. c., & tinazzi, m. (2010). expertise with pathological actions modulates a viewer’s motor system. neuroscience, 167, 691-699. doi:10.1016/j.neuroscience.2010.02.010. gauthier, i., & curby, k. m. (2005). a perceptual traffic jam on highway n170. current directions in psychological science, 14, 30-33. doi:10.1111/j.0963-7214.2005.00329.x gauthier, i., skudlarski, p., gore, j. c., anderson, a. w. (2000). expertise for cars and birds recruits brain areas involved in face recognition. nature neuroscience, 3, 191-197. doi:10.1038/72140 gegenfurtner et al | f l r 28 gauthier, i., tarr, m. j., anderson, a. w., skudlarski, p., & gore, j. c. (1999). activation of the middle fusiform ’face area’ increases with expertise in recognizing novel objects. nature neurosciences, 2, 568-573. doi:10.1038/9224 gauthier, i., williams, p., tarr, m. j., & tanaka, j. (1998). training ’greeble’ experts: a framework for studying expert object recognition processes. vision research, 38, 2401-2428. doi:10.1016/s00426989(97)00442-2 gegenfurtner, a. (2013). transitions of expertise. in j. seifried & e. wuttke (eds.), transitions in vocational education (pp. 305-319). opladen: budrich. gegenfurtner, a., kok, e., van geel, k., de bruin, a., jarodzka, h., szulewski, a., & van merriënboer, j. j. g. (2017). the challenges of studying visual expertise in medical image diagnosis. medical education, 51, 97-104. doi:10.1111/medu.13205 gegenfurtner, a., lehtinen, e., & säljö, r. (2011). expertise differences in the comprehension of visualizations: a meta-analysis of eye-tracking research in professional domains. educational psychology review, 23, 523-552. doi:10.1007/s10648-011-9174-7 gegenfurtner, a., nivala, m., säljö, r., & lehtinen, e. (2009). capturing individual and institutional change: exploring horizontal versus vertical transitions in technology-rich environments. in u. cress, v. dimitrova, & m. specht (eds.), learning in the synergy of multiple disciplines. lecture notes in computer science (pp. 676-681). berlin: springer. doi:10.1007/978-3-642-04636-0_67 gegenfurtner, a., & seppänen m. (2013). transfer of expertise: an eye-tracking and think-aloud study using dynamic medical visualizations. computers & education, 63, 393-403. doi:10.1016/j.compedu.2012.12.021 gegenfurtner, a., siewiorek, a., lehtinen, e., & säljö, r. (2013). assessing the quality of expertise differences in the comprehension of medical visualizations. vocations and learning, 6, 37-54. doi: 10.1007/s12186-012-9088-7 gegenfurtner, a., & szulewski, a. (2016). visual expertise and the quiet eye in sports – comment on vickers. current issues in sport science, 1, 108. doi:10.15203/ciss_2016.108 gegenfurtner, a., & van merriënboer, j. j. g. (2017). methodologies for studying visual expertise. frontline learning research. gibson, j. (1986). the ecological approach to visual perception. new york: psychology press. goodwin, c. (1994). professional vision. american anthropologist, 96, 606-633. doi:10.1525/aa.1994.96.3. 02a00100 grill-spector, k., knouf, n., & kanwisher, n. (2004) the fusiform face area subserves face perception, not generic within-category identification. nature neuroscience, 7, 555-562. doi:10.1038/nn1224 gruber, h., jansen, p., marienhagen, j., & altenmüller, e. (2010). adaptations during the acquisition of expertise. talent development & excellence, 2, 3-15. haller, s., & radue, e. w. (2005). what is different about a radiologist’s brain? radiology, 236, 983-989. doi:10.1148/radiol.2363041370 harel, a., kravitz d., & baker c. i. (2013). beyond perceptual expertise: revisiting the neural substrates of expert object recognition. frontiers in human neuroscience, 7, 885. doi:10.3389/fnhum.2013.00885 harley, e. m., pope, w. b., villablanca, j. p., mumford, j., suh, r., mazziotta, j. c., et al. (2009). engagement of fusiform cortex and disengagement of lateral occipital cortex in the acquisition of radiological expertise. cerebral cortex, 19, 2746-2754. doi:10.1093/cercor/bhp051 helle, l., nivala, m., kronqvist, p., gegenfurtner, a., björk, p., & säljö, r. (2011). traditional microscopy instruction versus process-oriented virtual microscopy instruction: a naturalistic experiment with control group. diagnostic pathology, 6, s81-s89. doi:10.1186/1746-1596-6-s1-s8 hinojosa, j. a., mercado, f., & carretié (2015). n170 sensitivity to facial expression: a meta-analysis. neuroscience & biobehavioral reviews, 55, 498-509. doi:10.1016/j.neubiorev.2015.06.002 holmqvist, k., nyström, n., andersson, r., dewhurst, r., jarodzka, h., & van de weijer, j. (2011). eye tracking: a comprehensive guide to methods and measures. oxford: oxford university press. gegenfurtner et al | f l r 29 jarodzka, h., balslev, t., holmqvist, k., nyström, m., scheiter, k., gerjets, p., et al. (2012). conveying clinical reasoning based on visual observation via eye-movement modelling examples. instructional science, 40, 813-827. doi:10.1007/s11251-012-9218-5 jarodzka, h., jaarsma, t., & boshuizen, h. p. a. (2015). in my mind: how situation awareness can facilitate expert performance and foster learning. medical education, 49, 854-856. doi:10.1111/medu.12791 kanwisher, n. (2000). domain specificity in face perception. nature neuroscience, 3, 759-763. doi: 10.1038/77664 kanwisher, n., mcdermott, j., & chun, m. m. (1997). the fusiform face area: a module in human extrastriate cortex specialized for face perception. journal of neuroscience, 17, 4302-4311. kok, e. m., de bruin, a. b. h., robben, s. g. f., & van merriënboer, j. j. g. (2012). looking in the same manner but seeing it differently: bottom-up and expertise effects in radiology. applied cognitive psychology, 26, 854-862. doi:10.1002/acp.2886 kok, e. m., van geel, k., van merriënboer, j. j. g., & robben, s. g. f. (2017). what we do and do not know about teaching medical image interpretation. frontiers in psychology, 8, 309. doi:10.3389/fpsyg.2017.00309 lehtinen, e. (2012). learning of complex competences: on the need to coordinate multiple theoretical perspectives. in a. koskensalo, j. smeds, r. de cillia, & á. huguet (eds.), language: competencies change contact (pp. 13-27). berlin: lit. maurer, u., zevin, j. d., & mccandliss, b. d. (2008) left-lateralized n170 effects of visual expertise in reading: evidence from japanese syllabic and logographic scripts. journal of cognitive neuroscience 20, 1878-1891. doi:10.1162/ jocn.2008.20125 melo, m., scarpin, d. j., amaro, e., passos, r. b., sato, j. r., friston, k. j., et al. (2011). how doctors generate diagnostic hypotheses: a study of radiological diagnosis with functional magnetic resonance imaging. plos one, 6, e28752. doi:10.1371/journal.pone.0028752. nishimura, m., & maurer, d. (2008). the effect of categorisation on sensitivity to second-order relations in novel objects. perception, 37, 584-601. doi:10.1068/p5740 op de beeck, h. p., baker, c. i., dicarlo, j. j., & kanwisher, n. g. (2006). discrimination training alters object representations in human extrastriate cortex. journal of neuroscience, 26, 13025-13036. doi:10.1523/jneurosci.2481-06.2006 qi, z., wang, x., hao, s., zhu, c., he, w., & luo, w. (2016). correlations of electrophysiological measurements with identification levels of ancient chinese characters. plos one, 11, e0151133. doi: 10.1371/journal.pone.0151133 palmeri, t. j., & gauthier, i. (2004). visual object understanding. nature reviews neuroscience, 5, 291-303. doi:10.1038/nrn1364 ribas, l. m., rocha, f. t., siqueira ortega, n. r., freitas de rocha, a., & massad, e. (2013). brain activity and medical diagnosis: an eeg study. bmc neuroscience, 14, 109. doi:10.1186/1471-2202-14-109 richler, j. j., & gauthier, i. (2014). a meta-analysis and review of holistic face processing. psychological bulletin, 140, 1281-1302. doi:10.1037/a0037004 righi, g. r., tarr, m., & kingon, a. (2013). category-selective recruitment of the fusiform gyrus with chess. in j. staszewski (ed.), expertise and skill acquisition: the impact of william g. chase (pp. 261280). new york: taylor & francis. rossion, b., gauthier, i., goffaux, v., tarr, m. j., & crommelinck, m. (2002). expertise training with novel objects leads to left-lateralized facelike electrophysiological responses. psychological science, 13, 250257. doi:10.1111/1467-9280.00446 rossion, b., kung, c.-c., & tarr, m. j. (2004). visual expertise with nonface objects leads to competition with the early perceptual processing of faces in the human occipitotemporal cortex. proceedings of the national academy of sciences, 101, 14521-14526. doi:10.1073/pnas.0405613101 säljö, r. (2009). learning, theories of learning, and units of analysis in research. educational psychologist, 44, 202-208. doi:10.1080/00461520903029030 gegenfurtner et al | f l r 30 scott, l. s., tanaka, j. w., sheinberg, d. l., & curran, t. (2006). a reevaluation of electrophysiological correlates of expert object processing. journal of cognitive neuroscience, 18, 1453-1465. doi:10.1162/jocn.2006.18.9.1453 scott, l. s., tanaka, j. w., sheinberg, d. l., & curran, t. (2008). the role of category learning in the acquisition and retention of perceptual expertise: a behavioral and neurophysiological study. brain research, 1210, 204-215. doi:10.1016/j.brainres.2008.02.054 seppänen, m., & gegenfurtner, a. (2012). seeing through a teacher’s eyes improves students’ imaging interpretation. medical education, 46, 1113-1114. doi:10.1111/medu.12041 sergent, j., ohta, s., & macdonald, b. (1992). functional neuroanatomy of face and object processing. a positron emission tomography study. brain, 115, 15-36. doi:10.1093/brain/115.1.15 shen, j., mack, m. l., & palmeri, t. j. (2014). studying real-world perceptual expertise. frontiers in psychology, 5, 857. doi:10.3389/fpsyg.2014.00857 siewiorek, a., & gegenfurtner, a. (2010). leading to win: the influence of leadership style on team performance during a computer game training. in k. gomez, l. lyons, & j. radinsky (eds.), learning in the disciplines (vol. 1, pp. 524-531). chicago, il: international society of the learning sciences. squire, l. r., berg, d., bloom, f. e., du lac, s., ghosh, a., & spitzer, n. c. (2013). fundamental neuroscience (4th ed.). oxford: academic press. stern, e., & schneider, m. (2010). a digital road map analogy of the relationship between neuroscience and educational research. zdm the international journal on mathematics education, 42, 511-514. doi:10.1007/s11858-010-0278-1 szulewski, a., gegenfurtner, a., howes, d., sivilotti, m., & van merriënboer, j. j. g. (2017). measuring physician cognitive load: validity evidence for a physiologic and a psychometric tool. advances in health sciences education. doi:10.1007/s10459-016-9725-2 tanaka, j. w., & curran, t. (2001). a neural basis for expert object recognition. psychological science, 12, 43-47. doi:10.1111/1467-9280.00308 tarr, m. j., & gauthier, i. (2000). ffa: a flexible fusiform area for subordinate-level visual processing automatized by expertise. nature neuroscience, 3, 764-769. doi:10.1038/77666 towler, j., fisher, k., & eimer, m. (2017). the cognitive and neural basis of developmental prosopagnosia. the quarterly journal of experimental psychology, 72, 316-344. doi:10.1080/17470218.2016.1165263 van geel, k., kok, e. m., dijkstra, j., robben, s. g. f., & van merriënboer, j. j. g. (2017). teaching systematic viewing to final-year medical students improves systematicity but not coverage or detection of radiologic abnormalities. journal of the american college of radiology, 14, 235-241. doi:10.1016/j.jacr.2016.10.001 walsh, v., & cowey, a. (2000). transcranial magnetic stimulation and cognitive neuroscience. nature reviews neuroscience, 1, 73-80. ward, j. (2006). the student's guide to cognitive neuroscience. new york: psychology press. wong, a. c.-n., palmeri, t. j., & gauthier, i. (2009). conditions for facelike expertise with objects. becoming a ziggerin expert—but which type? psychological science, 20, 1108-117. doi:10.1111/j.1467-9280.2009.02430.x wong, a. c.-n., & wong, y. k. (2014). interaction between perceptual and cognitive processing well acknowledged in perceptual expertise research. frontiers in human neuroscience, 8, 308. doi:10. 3389/fnhum.2014.00308 xu, y. (2005). revisiting the role of the fusiform face area in visual expertise. cerebral cortex, 15, 12341242. doi:10.1093/cercor/bhi006 microsoft word hagenauer et al_publication.docx ! ! ! ! frontline)learning)research)vol.4)no.)3)(2016))44)<)74) issn)2295<3159)) ! university teachers’ perceptions of appropriate emotion display and high-quality teacher-student relationship: similarities and differences across cultural-educational contexts gerda hagenauera1, michaela gläser-zikudab, & simone e. voletc auniversity of bern, switzerland buniversity of erlangen-nuremberg, germany cmurdoch university, australia article received 27 january / revised 8 april / accepted 27 april / available online 12 may ! abstract research on teachers’ emotion display and the quality of the teacher-student relationship in higher education is increasingly significant in the context of rapidly developing internationalization in higher education, with scholars (and students) moving across countries for research and teaching. however, there is little theoretically grounded empirical research in this area, and the different research strands remain relatively unconnected. the present study aimed to address this gap. psychological, educational and cross-cultural theories were brought together to investigate the interplay of emotion display and the quality of the teacher-student relationship from a teachers’ perspective and across “cultural-educational” contexts. given that social interaction, and the mores and norms associated with emotion display are often culturally underpinned, this study explored how university teachers in two so-called “individualistic” countries with different educational systems displayed positive and negative emotions in their teaching and what they perceived as an ideal teacher-student relationship. australian (n = 15) and german (n = 9) university teachers in teacher education were interviewed. the study revealed that while both groups viewed the open expression of positive emotions as integral to teaching, and negative emotions to be controlled based on their understanding of professionalism, significant group differences were also found. while the australian teacher educators reported higher and more intense expression of positive emotions, their german counterparts reported more open anger display. subtle yet noteworthy differences in the tsr quality between the two groups of teachers emerged. the findings of this study have implications for research and practice in international higher education. keywords: teacher emotions, emotion display, higher education, teacher-student relationship, cross-cultural comparison, internationalization of higher education !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 1 corresponding author: gerda hagenauer; university of bern, institute of educational science, fabrikstrasse 8, 3012 bern, switzerland; email: gerda.hagenauer@edu.unibe.ch doi: http://dx.doi.org/10.14786/flr.v4i3.236 hagenauer(et (al ( ( ( | f l r ! ! 45! 1. introduction emotions in teaching have received increased attention in educational research in the school context (newberry, gallant, & riley, 2013; schutz & zembylas, 2011, schutz, 2014). teachers’ experience of emotions and the communication of these emotions are expected to have a significant impact on the quality of their teaching practice and the socio-emotional climate in the classroom (jennings & greenberg, 2009). while the literature on teacher emotions at school steadily increases, affective factors in teaching-learningprocesses remain largely neglected in the higher education (he) literature (beard, clegg, & smith, 2004; quinlan, 2016), in particular teacher emotions (moore & kuol, 2007). however, recent empirical research has shown that teaching is also experienced emotionally in he (hagenauer & volet, 2014a; postareff & lindblom-ylänne, 2011), and is related to quality indicators of teaching, e.g. student-centered teaching (trigwell, 2012), that affect students’ learning (e.g., students’ engagement; zhang & zhang, 2013). besides emotional aspects, the teacher-student relationship (= tsr) in he is also a highly neglected field of research, in particular from a he teachers’ perspective (hagenauer & volet, 2014c; walker & gleaves, 2016). some research has shown that emotions and relationships are strongly intertwined (e.g. parkinson, fischer, & manstead, 2005): if someone seeks to understand people’s emotions in interactions, the quality of the relationship has to be considered as it contributes to the quality of the emotion evoked. in turn, the way emotions are communicated (= emotion display) contributes to the development of relationships (boiger & mesquita, 2012). thus, research that combines the two research strands – emotions / emotion display and the tsr – is warranted in order to better understand he teacher-student interactions that form the basis of quality teaching and learning processes. the research presented here follows an earlier study on australian university teachers’ emotions and emotion display (hagenauer & volet, 2014a, b). the present research aimed to increase understanding in this area through a cross-cultural perspective, employing interview accounts of australian and german he teachers, and by bringing in the aspect of the tsr. based on psychological research on cross-cultural differences in emotions combined with a social-psychological lens on emotions (parkinson, fischer, & manstead, 2005), we examined whether display modes of emotions differ between teachers in an australian and german he context, if differences in the quality of the tsr can be found between countries and how emotion display and the tsr are linked. perceived differences in relationship quality may signal the presence of differences in display modes. the focus is on a particular group of teachers in he, namely teacher educators, who fulfill a special function as they do not only teach content in their respective subject, but are also expected to model teaching behaviour to their students (lunenberg, korthagen, & swennen, 2007). thus, the way university teacher educators relate to their students and how they display their emotions, does not only affect the teaching-learning-environment but also serves as a model for future school-teaching practices of pre-service teacher students. the present study is innovative in three aspects: firstly, in a field of research with scarce empirical evidence so far, it brings together in one study psychological and educational theoretical strands of research, which have traditionally been researched separately, namely research on teacher emotions, emotion display and the tsr in he. the cross-cultural perspective adds another dimension, which is important to explore in the context of the internationalization of he. secondly, it adopts the concept of “cultural-educational context” (volet, 2001), and thus moves beyond the frequently made cultural comparisons between collectivistic and individualistic countries. thirdly, it uses a qualitative approach to investigate differences across cultural-education contexts. this enables an in-depth exploration of cultural-educational practices, which differs methodologically from the typically quantitative driven approaches used in cross-cultural emotion research (e.g. safdar et al., 2009). 1.1 emotion and emotion display in higher education from an appraisal theoretical perspective (ellsworth & scherer, 2003) it is assumed that emotions arise from the cognitive evaluation of a situation (e.g., a teacher judges the learning behaviour of a particular hagenauer(et (al ( ( ( | f l r ! ! 46! student). as evaluations of situations vary across people, the same situation can trigger different emotions in different people. according to this approach, emotions only develop if a situation is of relevance for people (= primary appraisal, lazarus, 1999); otherwise people remain emotionally “untouched”. further appraisal cognitions, e.g. in terms of controllability or goal attainment, determine the quality of the respective emotion. appraisals also influence whether people decide to suppress or to show the particular emotion. the ability to display emotions appropriately is a competence that can be linked to teachers’ overall emotion regulation competence (gross, 2002). according to gross (2010, p. 497) emotion regulation “refers to how we try to influence which emotions we have, when we have them, and how we experience and express these emotions.” emotion expression or emotion display are terms that describe the same phenomenon of how emotions are communicated, which is a constituent part of emotion regulation besides the internal regulation of emotion (e.g. down-regulating the intensity and duration of anger). the (appropriate) expression of emotions in he is discussed in various ways in the literature. in an overview, gates (2000) identified studies in which authors have argued that a neutral teaching and learning environment may be best for students’ learning in he classrooms, which suggests emotion suppression. however, the majority of literature advocates for authenticity, which allows for teachers’ emotion display in a controlled manner (e.g. cranton & carusetta, 2004). an authentic display of emotions fulfills a relevant function in establishing genuine and caring relationships with students (yuu, 2010; see also fischer & manstead, 2010 discussing the social functions of emotions in general), and is also significant for maintaining teachers’ health (zhang & zhu, 2008). research also suggests that emotion suppression can have adverse cognitive implications (gross, 2010), which in turn may also impact teaching quality while cognitive resources and attention are focused on the process of emotion suppression. the research strand on “emotional labour” discusses the requirement to mask emotions on the job (hochschild, 1983). employees aiming at successful fulfillment of work tasks are expected to follow particular occupational emotion display rules, which usually means masking (overly intense) negative feelings and acting in an emotionally positive manner. if students in he are regarded as customers, this also applies to teachers as department employees with teaching duties. according to fischer, manstead, evers, timmers and valk (2004), the display of desired emotions in the job does not necessarily cause negative side effects, presupposing that appropriate regulation of emotions is accepted as a part of teachers’ role-identity. however, negative consequences may result if emotional dissonance occurs, since the expressed emotions do not coincide with one’s identity. if teachers experience emotional labour over a sustained period, it could result in negative consequences, such as decreased satisfaction in the job, or burnout symptoms (e.g. emotional exhaustion) (barber, grawitch, carson, & tsouloupas, 2010; zhang & zhu, 2008). emotional labour can also arise from the role of a teacher, a job or function that includes moral elements (chen & kristjansson, 2011), such as pastoral care (isenbarger & zembylas, 2006; oplatka, 2007; yuu, 2010) or being role models. from this perspective, it is pertinent that teachers as educators are in control of their emotions, which requires a degree of emotional labour. concluding, lack of competence in appropriate communication of emotion can not only damage the wellbeing of he teachers but also endangers the development of positive relationships with students. consequently, emotions and the display of emotions contribute to the quality of interactions (boiger & mesquita, 2012). at university, interactions between teachers and students occur formally in courses and informally on campus. these interactions ultimately lead to the establishment of relationships, which are multilayered as they are built on a professional level (= working relationship) and on an interpersonal level (= closeness, affiliation) (hagenauer & volet, 2014c; nias, 1989). richardson and radloff (2014) have highlighted the significance of teacher-student-interactions for he students’ positive experiences, however they also cautioned that the frequency of direct interactions steadily decreases due to changes in the he context (e.g. increases in student-staff ratio or online learning; as found in the australian context). this development is alarming, given the fact, that fulfillment of the basic need for belongingness or relatedness (baumeister & leary, 1995; deci & ryan, 2002) is likely to be relevant for both – he teachers (e.g. in terms of their workplace satisfaction) and he students (e.g. in terms of their study commitment). hagenauer(et (al ( ( ( | f l r ! ! 47! as empirical evidence on the role of emotions in teacher-student interaction in he is rather scarce, we draw from research on schoolteachers in order to better frame this phenomenon. from this research we know that interactions with students are frequently emotionally laden and that student-teacher-interactions are the most prevalent source of teacher emotions (for a review, see sutton & wheatley, 2003). recently, hagenauer, hascher and volet (2015) have shown that “closeness” (as a result of many positive studentteacher interactions) to students (e.g., liking students; knowing them personally to some degree) predicted teachers’ emotions experienced when teaching in the classroom. the effect was particularly strong for the experience of joy. more generally, the model on teacher emotions developed by frenzel (2014) suggests, that teachers’ perceptions regarding the degree of goal attainment in particular teaching and/or interaction situations with single students, groups of students or classrooms, determines the quality of the emotion evoked. frenzel’s model relies on an appraisal-theoretical approach to emotions and shows overlaps with the control-value theory of achievement emotions, which is primarily applied in research on students’ emotions (pekrun, 2006). according to that model, a teacher might experience anger if students are unengaged in class, on the expectation that participation is linked to high achievement goals, performance and motivation. furthermore, it is reasonable to assume that teachers and students’ emotions are dynamically linked. students are expected to react sensitively to their teachers’ emotions, and reciprocally teachers would typically notice the emotional state of their students (garner, 2010). this was demonstrated in a study by frenzel, goetz, luedtke, pekrun and sutton (2009), which revealed that teacher’s joy supported schoolstudents’ enjoyment in learning. in the he context, titsworth, mckenna, mazer and quinlan (2013) detected a link between teacher immediacy, a quality indicator of a positive nonverbal teacher-student interaction, and positive student emotions. the phenomenon of emotions that spread from person to person is termed “emotional contagion” (e.g. hatfield, bensman, thornton, & rapson, 2014). thus, the quality of teacher-student interactions in the classroom is highly relevant for the concrete emotional experiences of the actors involved. however, empirical evidence on emotions and emotion display in teacher-student interactions in he, and more specifically in teacher education, is lacking. with regard to teacher education, some research has focused on the emotions of student teacher in the teaching practicum (e.g., hascher & hagenauer, 2016; pillen, beijaard & den brok, 2013; timostsuk & ugaste, 2012) or school-based teachers (hastings, 2008); but the emotional aspects of the university-based part of teacher education are overlooked. 1.2 cultural aspects of emotion display and relationships cross-cultural comparisons in emotion research are not new (e.g. eid & diener, 2001). according to mesquita (2007), emotions are “culturally situated” (p. 410) and are not just an individual phenomenon. emotions can be described as socio-cultural phenomena since most human emotions are evoked in social situations of interaction (boiger & mesquita, 2012). many studies of emotions across cultures have focused on the expression of emotions in social situations, which are underpinned by so-called display rules. for safdar et al. (2009, p.1), display rules “influence the emotional expression of people from any culture depending on what that particular culture has characterized as an acceptable of unacceptable expression of emotion”. besides personality traits that affect emotion display (e.g. people with high scores on extraversion display emotions more intensively; matsumoto, 2006), culture has been claimed to be an influencing factor on emotion display. in countries categorized as “individualistic” (hofstede & hofstede, 2005), where a strong independent self is highly valued, different modes of emotion display can be observed in comparison to countries labeled “collectivistic”, where an interdependent self is encouraged (markus & kitayama, 1991). despite evidence that the display of (negative) emotions is regarded as the right of the individual in “individualistic” countries, from a collectivistic point of view, emotions tend to be controlled in favor of the enhancement of positive relationships and harmony (safdar et al., 2009). furthermore, in countries labeled “individualistic” people have been found to value high intensive positive emotions (e.g. excitement), whereas people in more hagenauer(et (al ( ( ( | f l r ! ! 48! collectivistic cultures strive for the experience of low(er) intensive positive emotions (e.g. feeling calm) (tsai, knutson, & fung, 2006). based on hofstede and hofstede’s research (2005), western countries, such as the u.s., australia and many european countries, are considered “individualistic” countries, and east asian countries as “collectivistic”. however, such a broad-brush categorization has been widely criticized, given that there is great variability within cultural dimensions, including in terms of emotion display (schwarz & ros, 1995). consequently, there is a need for extending our understanding of emotion display rules within, and not only between, individualistic and collectivistic countries, which has been largely unexplored to date (see also koopman-holm & matsumoto, 2011, comparing the u.s. with germany). furthermore, in he, teachers act in a professional setting that has its own display rules. while emotions must be more controlled in workplace settings in comparison to private settings (moran, diefendorff, & greguras, 2013), another dimension requiring attention in cross-cultural emotion display research is that of context, and the roles people play and relationships they form in particular contexts. therefore in addition to potential variability of emotion display within so called “individualistic” and “collectivistic” countries, the context and situation should be considered (see also volet, 2001). an actor will vary in emotion display depending on the current role, as mother/father or university lecturer for example, and will also be impacted by the particular relationship quality. as safdar et al. (2009) observed, the display mode of emotion is usually contingent on the interaction partner, this interdependency having consequences for the relationship, due to the reciprocal influence of emotion display and relationship-quality (boiger & mesquita, 2012; eid & diener, 2001; fischer & manstead, 2010). concluding, similarly to the display of emotions, the quality of the tsr is influenced by cultural and institutional (organizational) norms, functioning as cultural guidance modeling “ideas, meanings, and practices of how to be a person and how to relate to others” (boiger & mesquita, 2012, p. 224). 1.3 the present study: aims and relevance the present study explored how australian and german university teachers, in so-called “individualistic” countries, displayed their positive and negative emotions when teaching at university and how they perceived a positive tsr. the review article of quinlan (2016) on he teaching and learning underpins the importance of relationships for emotions in interactions and vice versa, which stresses the need for more empirical research on that issue. furthermore, this study aimed to deepen understanding about university teachers’ emotion display through a cross-cultural lens. this research is timely given the pace of he’s internationalization (altbach & knight, 2007) in both research and teaching (e.g., shimmi, 2014). understanding the cultural specifics of both home and host country appears imperative for cultural adaptation, and the quality of teaching practice from an international perspective. more specifically, the following research questions were addressed: a) what do german and australian teacher educators perceive as appropriate emotion display in terms of positive and negative emotions when teaching and interacting with their students? what are similarities and differences in their perceptions? b) how do german and australian teacher educators construe the quality of the tsr, and the “ideal tsr”? what are similarities and differences in their views? c) how do german and australian teacher educators’ modes of emotion display and views of the quality of the tsr quality interrelate? d) how do particular background variables (e.g. position at university, background as a school teacher) contribute to explain the mode of emotion display and views of the quality of the tsr among australian and german teacher educators? the choice of samples within the individualistic cluster was based on convenience, and represent the authors’ respective, personal cultural background and familiarity with specific cultural-educational he contexts. however, this choice also addresses concerns that cross-cultural comparisons should go beyond the frequently made broad discrimination of collectivistic versus individualistic countries. indeed, the two hagenauer(et (al ( ( ( | f l r ! ! 49! individualistic countries chosen for the present study score differently on hofstede and hofstede’s (2005) degree of individualism. based on their “individualism index” (p. 78), australia ranks higher on individualism (2 out of 74 countries/regions) compared to germany (ranked 18). consequently, it was reasonable to anticipate some variability in display practices as well as perceptions of quality tsr, which could be interpreted in regard to their respective cultural-educational contexts (cultural and educational dimensions being confounded). 2. method 2.1 the participants fifteen australian (6 male, 9 female) and nine german (5 male, 4 female) teacher educators from two public universities in australia and one public university in germany participated voluntarily in the study. participants for the study were approached informally by one of the authors from the same culturaleducational context. selection criteria aiming at achieving representativeness in terms of relevant demographic characteristics pertaining to the population of teacher educators were: (1) at least two years of teaching experience (as there is evidence that the emotional experiences of beginning teachers are of particular quality, e.g. ria, sève, saury, theureau, & durand, 2003); (2) teaching in different subject areas across teacher education, for example, introductory courses in educational psychology, school pedagogics, mathematics and science education, civic education, and literacy education (as different subject areas may attract different styles of “communication”); (3) holding different positions in the university, ranging from full professor, to (full-time; part-time) lecturer, to phd-student with teaching duties (as particular positions in the university system are known to determine he teachers’ duties, and in turn the relative importance they give to teaching, frequency of interactions with students etc.). this purposive sampling strategy captured the typical heterogeneity of the population of teacher educators in these two countries, which is necessary for exploring the phenomenon in its breadth. the australian sample comprised twelve lecturers in teacher education (e.g., associate lecturer, senior lecturer, lecturer; most of them at post-doctoral level), and three associate professors in education with broader teaching and research responsibilities. in contrast, the german sample consisted of five full professors, one full-time lecturer, one post-doctoral fellow, and two phd students with teaching duties. the cultural-educational background of participants differed across countries the australian sample was culturally more diverse than the german sample, which is typical of the profile of german and australian teacher educators in general. none of the german teacher educators had any personal or professional experience of another cultural-educational setting, whereas five australian teachers came from another country. however, these teachers had already some years of teaching experience in australia before the interview was conducted. to protect the anonymity of the australian participants, no details pertaining to their specific cultural background can be provided. 2.2 the context: teacher education in germany and in australia teacher education in germany is structured in two main phases. phase 1 covers predominantly academic studies at a university (in general for 6 to 10 semesters), including some phases of school practice. however, most of the practical preparation is provided in a second phase, taking place in special, generally small, institutions operated by state governments and known as studienseminare. the second phase typically lasts 24 months (könig & blömeke, 2013). the present study took place in the first phase of teacher education at a university in thuringia. no selection procedures (numerus clausus) for entry into teacher education are applied at that university. thus, all students who had successfully completed their abitur, the hagenauer(et (al ( ( ( | f l r ! ! 50! secondary school completion diploma, were qualified to commence teacher education. students at that university studied teacher education to practice in high-track secondary schools, so called gymnasien or realschulen. typically, most of these students have an academic family background and are not from a low ses or migrant family background. teacher education in australia is structured differently, depending on whether it prepares students to become primary school teachers, or secondary school teachers specializing in the teaching of particular subjects, e.g. mathematics, languages, science. in the two universities where the study was conducted, students taught by the respondents were predominantly future primary school teachers. primary teacher education, for students coming straight out of high school (the majority), is usually completed in a four-year period. academic subjects are typically interspersed with practical preparation during the first three years with a strong practicum component in the last year. entry into primary teacher education is based on academic results in high school, or completion of other tertiary study, but entry levels tend to be lower than in other fields of university study. a substantial proportion of future primary school teachers, taught by the respondents at these two universities, would have been from a low ses, and possibly migrant family background, although primary school teachers in australia are not a culturally diverse professional group, which contrasts somehow with the overall diverse population. 2.3 interviews and procedure the first author conducted semi-structured face-to-face individual interviews with the teachers in each country. the style of the interviews was open, informal and conversational. university ethics approvals and informed consent from participants were obtained prior to participation. most of the interviews were conducted in the interviewee’s offices on campus, and a few in the staff room, if the teachers preferred this more informal context. interview duration ranged from 35 to 75 minutes; most interviews were 45-55 minutes. interviews were digitally recorded and transcribed verbatim. in the australian context, the interviews were in english; in the german context they were in german. a german speaking person proficient in english translated the quotes used in the present article. for crosschecking, an english speaking person with sound german language skills did the same, leading to english transcriptions of the german interview accounts. a semi-structured interview guideline framed the basic three themes that were addressed during the interviews: (1) emotions when teaching; (2) emotion regulation (internal regulation and emotion display) and (3) the tsr at university. in this paper, we refer to the interview accounts on theme 2 – focusing on the display of emotions – and theme 3 – the tsr. emotional and relationship issues can be addressed directly or indirectly in the interviews. in our study, and consistent with ethical standards in qualitative research, the purpose of the research was communicated to the participants right from the beginning and participants were asked direct questions pertaining to the main focus of the research. thus, the participants were fully aware that the interview would focus on emotions and relationships with students in he. in other words, there was no artificially masking the central topics through indirect questions. furthermore, we expected these experienced teacher educators would have reflected on their teaching and would feel ready to talk about their emotions and relationships with students. in terms of emotion display, the leading question was “do you show and express your feelings while teaching and interacting with students or do you also hide them sometimes?” followed by various probes that included descriptions of concrete interaction situations that were experienced by the teachers in teaching situations, focusing on teaching in small-groups settings, such as seminars or workshops (up to about 30 students), and teaching first-year students. the teachers also talked about concrete emotions, e.g. joy or anger. if these accounts included information about the display of emotions, they were coded within the category “emotion display” as well. regarding the tsr, teachers were asked to describe the “ideal” tsr at university from their perspective. again, probes were used to elicit further elaboration on the tsr. if descriptions of emotional interaction situations contained details pertaining to the tsr, these accounts were coded within the category of tsr as well. hagenauer(et (al ( ( ( | f l r ! ! 51! 2.4 data analysis the interview material was coded based on a category system in orientation to a deductive-inductive qualitative content analysis (mayring, 2000; gläser-zikuda & mayring, 2003). the analysis involved several steps. step 1: first, a theoretically based category system was developed to code the content of the interviews. then, the transcripts were read several times, and all text passages that could be allocated to the broad category of “emotion display” (subcategories: display of positive emotions versus display of negative emotions) or “the ideal tsr” were electronically coded using the maxqda software (“structuring content analysis” according to mayring, 2000). this step was done by the first author only, as due to the direct question format the extraction of these interview accounts was a very clear coding process with little room for interpretation. the three main codes – illustrated in table 1 – were derived deductively from the theory and the main research questions. to ensure an objective and reliable coding procedure, coding rules were formulated and anchor examples identified in the interviews. table 1 the coding scheme code number of accounts code description and anchor example displaying positive emotions aus: 46 ger: 28 this code is used when teachers talk about how to display positive emotions. example: so, okay, the positive ones are easy to handle. just join it, just share the fun. displaying negative emotions aus: 80 ger: 60 this code is used when teachers talk about how to display negative emotions. example: uhm, probably. i am sure i do. i am a bit of an open book. so, i think, you know, i don’t ... i don’t hide my feelings or even though i try to ... as i’ve said i am not gonna show that i am angry. the quality of the tsr aus: 69 ger: 24 this code is used for teachers’ answers on the question pertaining to “ideal tsr”. no distinction is made between the professional and interpersonal tsr at this coding step, as statements on both dimensions are frequently intertwined. example: i am careful, because i don’t want it to seem unprofessional. but yeah, i also think it’s important for them to see that ... it’s a whole person. step 2: after that, and in order to get greater insight into each individual teacher and his/her perception in terms of negative and positive emotion display and the tsr, a summary for each interviewee was prepared based on the extracted interview accounts resulting from the first coding step (“summarizing content analysis”; mayring, 2000). these summaries provided a concise overview of each case (or he teacher educator). after that, the original interview accounts and the summaries formed the basis for coding each case according to the relevant categories. a separate coding scheme for negative emotion display, positive emotion display and the tsr was applied to each case as the unit of analysis. [1] in terms of the display of negative emotions (= anger), three aspects were coded: a) did the teacher perceive the direct communication of anger as appropriate or not? (1= yes; 2 = no; 3 = ambivalent) b) did the teacher think that the communication of anger has to be controlled? (1 = yes; 2 = no; 3 = ambivalent) c) how does the teacher communicate his/her anger to the students? (1 = i-messages; 2 = group-messages; 3 = transfer messages; 4 = argumentative confrontation; 5 = provocative confrontation; 6!= sarcasm, irony; 7 = raising the voice; 8 = threat / hagenauer(et (al ( ( ( | f l r ! ! 52! classroom relegation; 9 = one-on-one contact after the lesson; 10 = i-messages (using less intensive words) while the codes 1 and 2 were derived deductively, the codes applied in (3) (= communication of anger) were inductively developed from the data. [2] in terms of the display of positive emotions, two aspects were coded: a) did the teacher perceive the direct communication of positive emotions as appropriate or not? (1=!yes; 2 = no; 3 = ambivalent). as not many accounts revealed information on the control of positive emotions, the code “emotion control” was not applied for the display of positive emotions. b) how does the teacher communicate his/her positive emotions? (1= positive feedback (neutral); 2 = praising students; 3 = communicating positive emotions intensively (verbally); 4 = hugging students; 5 = displaying enthusiasm; 6 = sharing humour). again, code 1 was derived deductively from theory, and the concrete emotional communication inductively from the data. [3] in terms of the ideal tsr two codes were applied: a) does the teacher perceive the professional aspect of the tsr as relevant? (1 = yes, 2 = no, 3 = ambivalent) b) does the teacher perceive the interpersonal aspect of the tsr as relevant? (1 = yes, 2 = no, 3 = ambivalent) both codes were derived deductively according to the conceptualization of the tsr (hagenauer & volet, 2014c). as an example, the coding scheme for the display of positive emotions is illustrated in table 2. hagenauer(et (al ( ( ( | f l r ! ! 53! table 2 the coding scheme for the display of positive emotions code definition anchor example direct expression of positive emotions yes (1) “yes” is coded, if teachers say that they communicate positive emotions to students. no (2) “no” is coded, if teachers do not communicate their positive emotions to students. ambivalent (3) “ambivalent” is coded if teachers basically agree that showing positive emotions is possible, but are hesitant about it at the same time. way of communicating emotions giving positive feedback (1) the teacher communicates positive emotions neutrally by giving contentfocused feedback. i give feedback to the students at the end of the session how i perceived the session. how i perceived the progress of the course. (i5, germany) praising students (2) the teacher communicates positive emotions (e.g., satisfaction) by praising the students. praise incorporates some kind of emotionality in the feedback. well, from my perspective praising students is very important. (i6, germany) expressing positive emotions intensively (verbally) (3) the teacher communicates positive emotions intensively verbally. the communication is more intensively compared to code 2. i would equally say "i am so happy for you." you know, if somebody gets a job or if somebody gets an award or something else. (i4, australia) hugging students (intense physical reaction) (4) the teacher hugs the students. i would hug students and students would hug me that ... that ... not all the time but i wouldn't hold back from doing that kind of thing. (i4, australia) displaying enthusiasm (5) the teacher shows enthusiasm evoked by the content / subject. i get excited about things. and i'll say: oh, guess what, guys! look at this! check this out! everybody come over! (i13, australia) sharing humour (6) the teacher shares humour in the classroom. having a laugh with the group. that's important. (i15, australia) interrater-reliability was calculated for the second coding step. interview accounts from eight interviewees (four from germany, four from australia), which represents about a third of the whole data set, were randomly selected and coded independently by two of the authors. both are fluent in german and english, which was critical as the german interviews were not translated to english. the result of the double-coding procedure is presented in table 3. percentage agreement was 92.60 % (54 coding options; 4 disagreements). in order to account for randomly reached agreement, the corrected cohen’s kappa was calculated (brenan & prediger, 1981). it lied at .85, which is satisfactory (bortz & döring, 2006). all disagreements were discussed between the two coders (including also going back to the whole interview incorporating any information that would help to clarify the respective code) until agreement was reached. after that, the first author went back to the data and validated the codes of the other interviews taking into account the aspects that had been discussed between the two coders. hagenauer(et (al ( ( ( | f l r ! ! 54! table 3 interrater agreement (selection of 8/24 teacher educators, 4 german and 4 australian) ger1 ger2 ger3 ger4 aus1 aus2 aus3 aus4 display of positive emotions x x x x x x x x display of negative emotions x x x x x x x x interpersonal teacher-student relationship x (x) x x x x x x emotion control (display of negative emotions) x x x x x x x x communicating positive emotions (1) x x (x) x x (x) x x communicating positive emotions (2) x x x (x) x x communicating positive emotions (3) x communicating negative emotions (1) x x x x x x communicating negative emotions (2) x note. x = agreement; (x) = disagreement; empty field = n.a., if for example, an interviewee only mentioned one way of communicating positive emotions. step 3: the final step complemented the comparison of emotion display modes and the quality of the tsr across the two countries based on the interview accounts extracted in coding step 1 and the case summaries resulting from coding step 2. regularly revisiting the original transcripts was undertaken at step 3 in order to confirm the accuracy of inferred conclusions. step 3 resulted in a case overview pertaining to the three main categories, namely, “display of negative emotions”, “display of positive emotions” and the “ideal tsr”.!! ! 3. findings as aforementioned in the introduction, emotion display and relationships are reciprocally entwined. the findings are structured around the four research questions, (1) starting by german and australian teacher educators’ perceptions of appropriate emotion display in their interactions with students, (2) followed by their views of the quality of the tsr at university and the “ideal tsr”, (3) the relationship between their reported modes of emotion display and views of the quality of the tsr, (4) and the examination of background variables that may contribute to explain modes of emotion display and views of the quality of the tsr across samples. 3.1 modes of emotion display (res q 1) 3.1.1 modes of positive emotion display across interviews, teacher educators stressed issues concerning the display of positive emotions less frequently than the display of negative emotions. for most, it was clear that positive emotions evoked in the classroom are shared easily, although the majority of german teacher educators appeared more reluctant in expressing positive emotions in an (emotionally) intense and direct manner than their australian counterparts. australian teacher educators frequently expressed strong positive feelings about students, as illustrated by emotionally laden expressions, such as “being thrilled about” or “getting very excited in the classroom”. in contrast, german teacher educators reported expressing positive emotions less directly, mainly communicating their satisfaction or joy by praising the achievement of particular students or a group. when probed about displaying positive feelings in the classroom, german teacher educators mostly tended to respond in a way similar to this example: hagenauer(et (al ( ( ( | f l r ! ! 55! i: “and how about the positive emotions? do you express them as well?” a: „yes, it is very similar. at the end of each course i give feedback how i experienced it. (i5, male, germany) thus, in the german context positive emotions were typically reflected in teacher’s feedback or praise, and commonly communicated in a relatively neutral manner. alternatively, the australian teachers reported more emotion-laden interactions in expressing praise. for example, so i have sent an announcement to everyone saying, i am really thrilled and proud of the feedback you are giving on your (anonymized; virtual platform). (i8, male, australia) a female australian teacher also mentioned that she would hug students if she felt deep joy, e.g. due to the success of a particular student. further, some australian accounts revealed how emotions transfer between teachers and students (e.g. a student feels happy following success; this happiness transfers to the teacher). taken together, it is clear that expressing positive emotions was widely regarded as appropriate and relevant (e.g., in terms of fostering students motivation) in both cultural-educational contexts. however, the actual mode of expression differed somewhat across contexts. this pertained mainly to the student-teacher interaction (e.g. how to praise students) and less to the teacher-subject-interaction (e.g., how to respond emotionally on subject matter). australian and german teacher educators expressed relaying enthusiasm for their taught subject similarly, but a little more pronounced in the australian context, as illustrated by the following quote: i mean, i get excited about things. very enthusiastic about the science. and i’ll say: oh guess what, guys! look at this! check this out! everybody come over! (imitates excitement in a classroom). now, is that emotional? yes! i get very enthused about the science or you know, that type of thing. it’s really cool stuff. but as far as reacting to students, i try not to be way up or way down. (i13, female, australia) finally, what was perceived as appropriate display of emotions varied among teachers from the same country, indicating that the expression of emotions may be influenced not only by the cultural-educational context but also by individual characteristics, such as personality or position at university. for example, contrary to most of her colleagues, one german teacher reported teaching in a very emotional manner, as she had experienced that kind of enthusiasm as a relevant antecedent of learner’s motivation during her former work as a schoolteacher. 3.1.2 modes of negative emotion display in terms of negative emotions, annoyance dominated in both the german and australian accounts. therefore the following analysis focused on teachers’ annoyance and anger (for an overview on the range of negative and positive emotions typically expressed by teacher educators, see hagenauer & volet, 2014a). in the present study, we focused on the most frequently mentioned negative behaviours that evoked teachers’ anger or annoyance (e.g. student disengagement or classroom disturbance). comparing case summaries revealed a marked difference in the accounts of the german and australian teachers. while nearly all the australian teachers shared the opinion that negative emotions should be suppressed for professional and role-modeling reasons, german teachers were less reluctant in expressing their anger directly to students. but both groups shared the belief that the display of anger should be controlled according to professional standards. more concretely, most australian teacher educators advocated for the need to suppress negative feelings in class. if classroom disturbances occurred, both groups claimed that the teacher should intervene calmly if the disturbance affected the learning of the others but would strategically ignore the disturbance if the classroom learning process were not endangered. they usually would not express negative emotions directly, but would talk to the particular student(s) on a one-on-one basis in or after the course to address the problematic behavior: hagenauer(et (al ( ( ( | f l r ! ! 56! but if they talk to each other, i would mention that. yeah, i would say: i would expect you, you know, you are listening at this point. i try and do that sometimes in a private way rather than publicly. (i14, male, australia) in contrast, the german teacher educators reported a variety of reactions when they experienced anger, which were made visible to students more pronounced relative to their australian counterparts. for example, one german female teacher reported that she experienced anger if students submitted assignments the evening before the seminar. upon probing how she expressed her anger, she replied: i tell them, it is not ok, if they don’t get the timing right and i have to pay for it. it’s not acceptable for me. i don’t want that and then i also justify why i don’t want it that way. (i9, female, germany) only one of the nine german teacher educators said he suppressed his anger completely, which he traced back to his personality and his difficulty in coping with conflict. the other german teacher educators reported using mainly verbal strategies in such circumstances, such as sending i-messages addressing problematic behaviour and how it affected them (“i am annoyed that…”), or the group (“do you think it is fair to your group members…?”), or asking the students how they would cope with such situations in their own classroom (bringing in the professional perspective). i-messages were also a popular method to deal with disturbances from the australian teachers’ perspective. if perceived as necessary, problems were also addressed verbally; however experienced negative emotions were frequently addressed more cautiously, by substituting with a less-intense emotion display. for example, one teacher reported dealing with anger and disappointment in class contributions: i mean, if you’re disappointed, if you are not happy with something, you have got to tell them. and i’ll tell them. if i say: look, i am not, i am not pleased. nobody seems to be contributing to this discussion today. we need to contribute to discussions. this is how we learn. we need to talk about things. so, you know, please get involved.” (i9, male, australia) when students expressed opposing views, or displayed lack of openness, which caused anger in some of the german teachers’ accounts, the german teachers frequently dealt with their annoyance through starting an argumentative and sometimes provocative discussion in class. some teachers also reported raising their voice or using sarcasm as verbal reactions in situations they found annoying. one teacher mentioned that he was willing to react directly by removing the problematic student from the classroom if the student did not respond favorably after a few prompts. some reactions to anger that were mentioned by the german teacher educators would be perceived as problematic in the australian context. for example, the display of anger was interpreted as a means of maintaining a productive professional working relationship from some german teachers’ perspective, since not displaying anger, and not intervening in difficult classroom situations would be regarded in that context as displaying a lack of professionalism. it should be noted, however, that such views were mainly expressed by less experienced teachers. in contrast, one australian teacher stated that he “would lose face” if he showed his annoyance directly in class. another australian teacher mentioned that he did not want to risk the progress of the work by letting negative emotions interfere while another female australian teacher said she would not want to risk the positive relationships with students by reacting angrily. figure 1 and 2 and table 4 outline the findings pertaining to the overall display and communication of positive and negative emotions. each teacher educator is listed as a case in the table and is represented within one column. in the figures the percentage of cases within each group of teacher educators (german, australian) is provided. the calculation of the percentage made it possible to compare the responses of german and australian teacher educators directly as within each group – although based on a different hagenauer(et (al ( ( ( | f l r ! ! 57! frequency they finally summed up to 100 %. thus, in the australian sample a teacher educator accounts for about 7 %; in the german sample for about 11 %. as can be seen, german and australian teacher educators shared many ways of displaying their emotions in the classroom, but there were also noticeable differences. german teacher educators frequently reported direct display of anger but this was not the case among the australians. in contrast, the direct display of positive emotions was slightly more pronounced within the australian compared to the german sample. the german teacher educators also reported higher intensity of direct anger display and also a higher variety of possible responses. in particular, german teacher educators reported using various forms of verbal reactions when facing difficult student behaviour, while their australian counterparts tended to avoid direct confrontation and preferred talking to their students after class one-on-one. verbal and rational reactions, termed as “positive feedback”, also seemed to be the german teachers’ preferred way of displaying positive emotions, while the emotional aspect of feedback/praise came through emotionally more intensely in the australian teachers’ accounts. figure 1: ways of communicating positive emotion by german and australian teacher educators figure 2: ways of communicating anger by german and australian teacher educators 56 33 11 22 22 7 14 21 50 43 0 10 20 30 40 50 60 70 positive feedback praise intensive emotional feedback enthusiasm humour p er ce nt ag e of c as es (w ith in th e cu lt. -e du c. c lu st er ) german teacher educators (n = 9) australian teacher educators (n = 14) 50 50 14 14 14 7 7 0 0 0 0 0 56 33 22 22 44 22 11 22 0 10 20 30 40 50 60 70 p er ce nt ag e of c as es (w ith in th e cu lt. -e du c. c lu st er ) australian teacher educators (n = 14) german teacher educators (n = 9) hagenauer(et (al ( ( ( | f l r ! ! 58! table 4 display of positive and negative emotions: case overview comparing german and australian teacher educators’ perceptions german sample australian sample g1 g2 g3 g4 g5 g6 g7 g8 g9 a1 a2 a3 a4 a5 a6 a7 a8 a9 a10 a11 a12 a13 a14 a15 f f m m m m f m f m* f* f* f* f f m f m f f m f* m m gender is it appropriate to display positive emotions and anger (as a frequent negative emotion) in he teaching? (x = yes; (x) = ambivalent; empty field = no) x (x) x x x x x (x) (x) (x) x x x n.a. n.a. x x x x x x x x x pos.e. x x x x x x x x x n.a. x x x (x) anger is it necessary to control the display of negative emotions? (x = yes; (x) = ambivalent) x x x x (x) x x x x x x x x n.a. x x x x x x x x x x control the display of positive emotions: reactions (x = mentioned in the interview) x x x x x x pof x x x x x pra x x x x ier/h x x x x x x x x x ent x x (x) x x x x x x hum the display of negative emotions: reactions (x = mentioned in the interview) x x x x x x x ooo x x x x x x x lie x x x x x x x i-m x x x x x g-m x x x x t-m x x x ac x x x x x pc x x sar x rai x x thr note. pos.e. = positive emotions; pof = positive feedback; pra = praise; ier/h = intensive emotional reaction (verbally) + hugging students; ent = enthusiasm; hum = humour; ooo = oneon-one contact after the lesson; lie = using less intensive emotion words; i-m = i-messages; g-m = group messages; t-m = transfer messages; ac = argumentative confrontation; pc = provocative argumentation; sar = sarcasm/irony; rai = raising the voice; thr = threat, classroom relegation; n.a. = not applicable (no information on that aspect in the interview): *coming from another cultural background; the arrow represents an increasing intensity of the respective emotional reaction; e.g. in terms of anger: talking one-on-one privately after the lesson (ooo) is less intensive than directly addressing the problematic behaviour in class (e.g. through i-messages; lie; i-m); in terms of positive emotions: giving neutral feedback (pof) is less intensive than praising students (pra) hagenauer(et (al ( ( ( | f l r ! ! 59! 3.2 quality of the teacher-student relationship (res q 2) according to hagenauer and volet (2014c) the quality of the tsr can be described in terms of support (or professional) dimension or in terms of affective (or interpersonal) dimension. our findings are presented in relation to these two dimensions, starting with the teachers’ views on the quality of the professional tsr. 3.2.1 the professional teacher-student relationship by comparing the accounts and reflections of the australian and german teacher educators, it is evident that both groups regarded the tsr as predominantly a professional one, with particular boundaries that must not be overstepped, but with “room to personalize” this relationship, as stated by one australian female teacher educator (i5). however, there were also differences between the two groups of teachers, particularly with regard to the amount of formality versus informality of the interactions, and the amount of (interpersonal) care expressed within this relationship. most of the german teacher educators described their role within the professional tsr mainly as being a moderator, a generator of ideas and inspiration or a specialist in terms of the teaching content, who designs effective learning environments in which students are required to contribute actively. this professional tsr understanding is illustrated below: i would like to be an instigator. i want to challenge them, they should think about stuff, care about things, which i think are relevant, yes, be an inspiration, a person who challenges, you know…the person, who asks good questions and starts new thinking processes. i’m not a guru. i can see that with some colleagues and that’s scary to me, you know, to have something like a fan club. i’m not the head-teacher. i’m more the person who asks questions. an instigator. (i9, female, germany) in this professional working-relationship mutual appreciation and respect are important components from the teacher educator’s perspective and a well-adjusted give-and-take basis is expected (e.g., in terms of engagement). a male professor called this kind of relationship “mutual receptiveness and openness” (i3). furthermore, many german teacher educators also expressed approachability as a relevant dimension of the relationship, equalizing approachability mostly with openness for content-related questions of students. informal contact between students and teachers apart from the regular course setting and the official officehours was rare. the following quote gives an example of a male teacher’s perception on approachability, which also addresses the idea that a certain distance in the tsr might not only be a need of teachers but also of students: the direct connection with students, if they [the students] want it, it’s not a problem for me. in lectures i say to them: if you have questions or if you need anything else you can come and talk to me during office hours, or they can have an additional appointment. it’s all possible. if i have the time, i will give it to them. but it has to be in a, you know, professional setting. well, it needs to stay connected to the topic. (i5, male, germany) in another interview a female university teacher-professor stressed that office hours must be kept, the dilemma being that good teaching is frequently unrewarded, which affects the amount of effort invested into teaching. ultimately, this also impacts on the frequency and intensity of teacher-student interactions: i plan 1.5 hours for the office hour, most of the time i need 2, i use a watch for it. well, they would like to be looked after for half an hour. that’s not possible. after 15 minutes they have hagenauer(et (al ( ( ( | f l r ! ! 60! to leave …at the latest. they have to ask precisely. it’s strange, but it feels that university teaching stops me from my work. […] umm, and it’s not that i don’t think that teaching is not important, it just feels, like doing something that nobody sees or doesn’t count. that’s the frustrating thing. (i9, female, germany) 3.2.2 the interpersonal teacher-student relationship australian teacher educators emphasized mutual respect and appreciation as relevant characteristics of the tsr, but they considered its quality in terms of informality in interactions. thus, within the interpersonal dimension of the tsr some sort of “closeness” is coming into play, as the following account reveals: i think closeness... and caring is quite important. we routinely here, as you noticed, we are not status-bound. i introduce myself to my students as xy (first name). i say: just call me xy (first name). and i want them to see me as someone who is here to help them, not someone who has an authority-status.[…] so, we have an ethos in tune. i think we've always had that. (i7, male, australia) in contrast, “closeness” was viewed skeptically by most of the german teacher educators: i would say, relationships between students and teachers or lecturers should always be at a professional level. of course there are sympathies. there are aversions. they are okay and legitimate. but they musnt’t disturb the sequence [of the course]. (i5, male, germany) as expected, these “relational” differences underpinned teacher-student interactions. in contrast to germany, it was found that australian students could call their university teachers by their first name, and many teachers reported an open-door policy, so students could approach them whenever they wanted to. it appeared also not unusual in the australian context for teachers and students to share personal information, sometimes during the course but also anywhere on campus. a male teacher discussed his willingness in establishing interpersonal relationships: i try and take an interest in them personally ...with their jobs and their families and the rest of their lives. and i try and make it clear that i am not just a teacher of [subject area]. i do other things as well. (i7, male, australia) furthermore, for the australian teachers, approachability implied that students could approach them when they had content-related questions, but also if they were dealing with personal issues that interfered with their study. aspects of personal care that appeared to come into play were more visible in the interview accounts of the australian teacher educators. this might be partly related to the fact that many students in the two australian universities in the present study were from a lower socioeconomic background. during the interviews, a number of teacher educators expressed worry or concern about the study success of these students and expressed a willingness to listen to their problems, for example, by granting extensions in terms of submitting assignments and spending extra time one-on-one to discuss open questions. and they have ...you know, there are lots of personal issues, particularly in a lower socioeconomic area ...[…] i generally find, that if someone has an issue and if you manage to build that rapport and that relationship they are happy to come and talk to me about it and to say, you know, that they're struggling because of being at the doctor last week and having heart tests and that they're, you know, are so stressed and not knowing what the results are hagenauer(et (al ( ( ( | f l r ! ! 61! and things like that...that, you know, i am happy to sort of say to them: well, fine. this assignment is due in friday night. monday is fine. (i8, female, australia) interestingly, when probed for worry or concern, many of the german teacher educators said that these emotions did not play any important role in their teaching practice. but in a workshop or seminar, whether they achieve or not, i’m not too worried, because ultimately, they are adults and i can’t do everything for them. (i2, female, germany) not only the background of students, but also the system more generally appeared to contribute to the explicit caring attitude of teacher educators in australia. as retention rate and learning outcomes of students are considered by funding bodies as important indicators of the “efficacy” of university teaching, these teachers felt somehow obligated to maximize students’ academic success, which sometimes created friction: it is hard, yeah. in australia very much the emphasis is to try and help them to pass. and that's something i am not used to…because sometimes i think some students really need to fail. (laughs). but you try and help them as best as you can. so, if they are struggling in language you'll offer them support in their language. but sometimes you think: really, this person shouldn't be teaching! (laughs) well i just, i sense that we are maybe a little bit softer than other countries. (i12, male, australia) based on the informality of the tsr and the amount of care invested in this relationship, the interpersonal aspect of the tsr emerged not only as stronger in the australian sample, but often as an explicit goal in teaching (e.g. building a rapport with students). the informality and caring attitude, highlighted by australian teacher educators, was also occasionally questioned, but mainly by teachers from another cultural background, suggesting that such informality may have to be learned by new teachers: here are more informal kinds of relationships. but it doesn't mean...being informal does not mean that there is no distinction between workshop leader and students. i find it very hard to balance. it's sometimes informal. okay, we are like equal, you know. we are like friends. […] so i adapted to it and, yeah, i still need to, i am still learning. i feel, because it's a long drawn thing having to find a nice balance or effective balance. i don't have to be nice but i need to be effective. (i1, male, australia) but not only the informality constitutes an aspect that new teacher educators may have to adapt to in the australian context. the same may apply to the adoption of an explicit “caring attitude” as observed by non-native teacher educator, who described her local colleagues as follows: most of the lecturers here would have been teachers at some stage. so, most of them would come with that caring attitude. you know, wanting to establish good relationships, wanting to have, you know, like the best possible environment, where they can teach and their students can learn. (i2, female, australia) taken together, these findings show that australian and german teacher educators share a similar belief about the necessity to form professional relationships with students at university. however, due to differences in how interactions are realized – in particular due to higher informality and more pronounced caring attitude in australia – the interpersonal tsr seemed to be closer in the australian sample, while the professional and more formal working relationship dominated in nearly all of the accounts of the german teachers (see table 5). hagenauer(et (al ( ( ( | f l r ! ! 62! table 5 case overview: the ideal teacher-student relationship (tsr): professional and/or interpersonal? german sample australian sample 1 2 3 4 5 6 7 8 9 1* 2* 3* 4* 5 6 7 8 9 10 11 12 13* 14 15 gender f f m m m m f m f m f f f f f m f m f f m f m m prof. tsr x x x! x! x! x! x! x! x! x! x! x! x! x! x! x! x! x! x! x! x! x! x! x! interp. tsr (x) x x x x x x x (x) x *coming from another cultural background 3.3 relationship between modes of emotion display, quality of the teacher-student relationship and cultural-educational background of teacher educators (res q 3) in a final step, three target factors were interrelated: (1) mode of emotion display; (2) quality of tsr and (3) cultural-educational background of teacher educators (see table 6). in order to explain the relevance of the tsr on the display reactions of emotion, the differences in display modes within the group of german teacher educators and within the group of australian teacher educators was of interest (= controlling for cultural-educational background). looking at the australian sample, it became clear that the teacher educators, whose tsr was more pronounced on an interpersonal level, displayed their positive emotions with higher intensity (marked in bold font, see table 6); this was the same for the german teacher educator who also formed strong interpersonal relations with her students. thus, for both cultural-educational groups the intense communication of positive emotions was strongly connected to the interpersonal tsr. in regard to the communication of anger, the australian teacher educators communicated anger similarly, regardless of whether they had formed a strong interpersonal tsr or not. however, a slight difference appeared: two of the teacher educators, who formed the tsr also on an interpersonal level, report using “moral” strategies when communicating their anger (we-messages; transfer messages). none of the australian colleagues who focused less on the interpersonal tsr reported these strategies. having a closer look at the german teacher educator with a more pronounced affective bond to her students, it became apparent that she used lessintensive anger reactions than some of her colleagues who adopted a greater distance and a less explicit caring attitude to their students. in a second step, a comparison within the respective tsr group was made, in order to explore whether cultural-educational differences were still noticeable when accounting for the quality of the tsr. within the group of teachers who formed strong interpersonal tsr, the german teacher educator showed similar reactions in terms of communicating positive emotions; but her display of negative emotions through i-messages was more directly compared to the main reactions of the australian counterparts, who preferred mainly one-on-one contacts after the course/workshop. within the group of teachers who formed the tsr mainly on a professional level, the obvious difference in the direct communication of anger and positive emotions between australian and german teacher educators persisted: on a professional level, the german teacher educators communicated their positive emotions more neutrally, and their anger more directly than their australian counterparts. thus, the results highlight two main points: first, an emphasis on the interpersonal tsr mainly goes along with the display of high intense positive emotions and less intense negative emotions. second, differences in emotion display between australian and german teacher educators were found also within the respective tsr grouping (professional + interpersonal, professional only), which points to the importance of cultural-educational background as an influencing factor on teacher educators’ emotion display that goes above and beyond the quality of the tsr. hagenauer(et (al ( ( ( | f l r ! ! 63! table 6 case overview: display of positive and negative emotions and quality of the tsr, taking teacher educators’ cultural-educational background into account professional relationship + interpersonal relationship professional relationship only g7 a3 a4 a6 a7 a8 a11 a12 a13 a14 g1 g2 g3 g4 g5 g6 g8 g9 a1 a2 a9 a10 a15 f f* f* f m f f m f* m f f m m m m m f m* f* m f m is it appropriate to display positive emotions and anger (as a frequent negative emotion) in he teaching? (x = yes; (x) = ambivalent; empty field = no) x x x n.a. x x x x x x x (x) x x x x (x) (x) (x) x x x x pos.e. x x (x) x x x x x x x x x x anger is it necessary to control the display of negative emotions? (x = yes; (x) = ambivalent) x x x x x x x x x x x x x x (x) x x x x x x x x control the display of positive emotions: reactions (x = mentioned in the interview) x x x x x x pof x x x x x pra x x x x ier/h x x x x x x x x x ent x x x x x x x x hum the display of negative emotions: reactions (x = mentioned in the interview) x x x x x x x ooo x x x x x x x x lie x x x x x x x i-m x x x x x g-m x x x x t-m x x x ac x x x x x pc x x sar x rai x x thr note. light shading = german teacher educators; dark shading =australian teacher educators; pos.e. = positive emotions; pof = positive feedback; pra = praise; ier/h = intensive emotional reaction (verbally) + hugging students; ent = enthusiasm; hum = humour; ooo = one-on-one contact after the lesson; lie = using less intensive emotion words; i-m = i-messages; g-m = group messages; t-m = transfer messages; ac = argumentative confrontation; pc = provocative argumentation; sar = sarcasm/irony; rai = raising the voice; thr = threat, classroom relegation; n.a. = not applicable (no information on that aspect in the interview); *coming from another cultural background a5 is not presented in the overview as she explicitly stated that she “does not take the emotional aspects” which made it impossible to code the main categories on emotion display. hagenauer(et (al ( ( ( | f l r ! ! 64! 3.4 factors contributing to explain modes of emotion display and the quality of the teacherstudent relationship in german and australian teacher educators: intensive and deviant case analyses (res q 4) finally, it was of interest to explore the extent to which some central background variables of these teacher educators may contribute to explain the way they communicated their emotions in the classroom and their views of the quality of the tsr. this exploration was done through the close analysis of deviant and intensive cases. in the german sample, the deviant case was g7, the teacher who reported strong interpersonal relationships with her students, whereas all other german teacher educators were allocated to the group “professional tsr only”. two german teachers (g2 and g5) who expressed rather extreme negative emotions in comparison to their colleagues were potential candidates for the intensive case for the german context. in order to control for possible gender effects, the female teacher (g2) was chosen for comparison. in the australian sample the deviant case was a15, the teacher who expressed the most distant relationship to students and who emphasized “professionalism” the most, and the intensive case was a4, the teacher who showed the closest relationship to students and adopted the most intensive way of communicating her positive emotions. as background criteria the following aspects were used: position at university, age, experience in he teaching, experience in school teaching, amount of out-of-class interactions, dedication for he teaching (in teacher education) and having research duties beyond teaching duties. the results of the comparisons are depicted in figures 3 and 4. for both samples, the identification that a teacher reveals from teaching in he and in teacher education in particular generated the most salient difference in how interactions with students were implemented (= dedication he teaching). both teachers who described an intense display of positive emotions and “close(r)” interpersonal relationships to students reported that they loved teaching and were very keen on teaching in teacher education, and enjoying the sharing of their own experiences as school teachers to student teachers. g7 could dedicate a lot of her time doing this, as she was a full-time lecturer; thus the position matched her interests as the following account illustrates: that's a big part of my professional understanding, and afterwards i go out of these classes and i’m really ... feel like being on cloud nine. then, i feel really good, then i know: yeah, that's my job, that’s what i wanna do. you know when i can inspire them for what they aspire to be. wow, this really just warms my heart. so yes, that's, um, that's really my idealism that i bring into this job. you know, i really really like speaking about teaching, i like teaching a lot, but i also like to talk about school lessons, and i like to pique the students’ interest. (i7, germany ) the dedication to teaching in teacher education was not as visible in the accounts of those two teachers who displayed more distant relationships and a less intense way of communicating positive emotions to their students. that might be traced back to several factors: this german teacher did not have a school teaching background herself and was quite inexperienced in he teachings, which was accompanied by some problems in classroom management and thus feelings of insecurity. in the australian sample, position at university might have contributed to differences. i15 had a background as a schoolteacher, but he was a sessional lecturer (= an external lecturer) who taught at university only once a week. he taught different courses on that day and thus, the amount of informal interactions he could have with students was limited. in addition, he expressed that he was not sure how long he would teach at university which suggested low(er) identification as a teacher educator at university. hagenauer(et (al ( ( ( | f l r ! ! 65! ! figure 3: case comparison: german sample ! figure 4. case comparison: australian sample hagenauer(et (al ( ( ( | f l r ! ! 66! 4. discussion the present study explored the emotion display and the quality of the tsr at university from a teachers’ perspective, with particular attention to cultural-educational differences between a sample of teachers from australia and germany. the findings revealed many similarities but also differences between the australian and german teacher educators, in their emotion display and in the quality of the tsr. bringing these two strands of findings together, and drawing on the literature on emotions in social relations (boiger & mesquita, 2012; parkinson et al., 2005), it can be concluded that not only cultural aspects but also the quality of the relationship formed between teachers and students within their respective cultural he setting, is likely to impact on what is perceived by university teachers as appropriate ways of displaying emotions. although a predominantly egalitarian tsr in the professional dimension prevailed in both groups, in which considering the instructor as facilitator supporting critical thinking and student centered learning processes (in contrast to a predominantly hierarchical tsr, which may be more common in the chinese he; see zhang & zhang, 2013), the quality of the interpersonal dimension of the tsr appeared to differ between the two cultural-educational contexts. interpersonal relations between teachers and students were relatively pronounced in the australian setting, while in the german sample the professional working relationship between students and teachers was emphasized, resulting in higher formality in student-teacher interactions. the importance given to caring in the interpersonal tsr by the australian teacher educators might be partly explained in relation to the higher proportion of their students being from low ses in comparison to the german context. however, the formality versus informality in the interactions that also substantially contributed to the interpersonal tsr may be culturally informed and less dependent on student characteristics. future studies will need to account for possible moderator variables, such as the students’ background, when investigating teacher-student interactions and relationships. as argued by safdar et al., (2009), relationship quality typically affects the way emotions are communicated, as the “closeness” of the interaction partner is regarded as a significant influencing factor of emotion display. typically, people express emotions with higher intensity when relationships are close (fischer & manstead, 2010). this was also evident in our study, but only for the positive emotions: teacher educators, who formed more pronounced interpersonal relationships with their students, displayed their positive emotions more intensely, while they communicated the negative ones in a relatively reserved way. on the other hand, if teachers formed relationships mainly on a professional level, their positive emotions were communicated less intensely. although all teacher educators in this study believed in the necessity to control emotions in professional settings, the way emotion display was acted out in practice differed in the two culturaleducational contexts. while the majority of the australian teachers communicated positive emotions immediately but avoided the direct and immediate display of anger, their german counterparts were less direct in the display of positive emotions but more immediate in the communication of anger. furthermore, most of the german teacher educators did not perceive the controlled expression of anger as a threat for the professional tsr but instead regarded it as a necessary component of their function as a teacher. some german teacher educators explicitly stated that it is not about “liking them”, but about learning as much as possible from their courses. these views are consistent with yuu (2010), who argued that expression of negative emotions in an appropriate way does not risk relationships but bridges “the psychological and emotional distance between teachers and students” (p. 76). in contrast, some australian teacher educators stressed their effort not to display strong negative emotions in order to maintain positive relationships with students and not to risk authority. in a recent study, bartram (2015) found that sometimes students take advantage of close and caring relationships with he teachers, a behavior he coined “affective strategizing” (p. 68). this suggests that a caring attitude of he teachers may not necessarily be positive, as it comes with individual costs for the teacher (e.g. a higher workload) and with students capitalizing on this attitude (for a conceptualization on the caring he teacher, see walker & gleaves, 2016). interestingly, the less pronounced acceptance of anger expression in the australian sample appears consistent with countries that score higher on hofstede and hofstede’s (2005) collectivistic dimension, hagenauer(et (al ( ( ( | f l r ! ! 67! whereas the display of negative emotions is regarded as the right of the individual in more individualistic countries. however, the conclusion that australian he teacher educators might display more collectivistic tendencies than their german counterparts cannot be made based on the present data. in fact, australia scores higher on hofstede and hofstede’s individualism index than germany. in order to clarify that to some degree unexpected finding, future research will need to explore the extent to which different degrees of individualistic and collectivistic tendencies can be found within so-called “individualistic or collectivist” cultures / countries and how these differences are connected to teacher motivation, emotion and behaviour. we would argue that “individualism-index” dimensions are not sensitive enough to interpret differences in interactions and relationships, and that the educational context and the norms, values and rules within this context provide more promising explanatory elements. the combination of cultural and contextual elements was captured in the term “cultural-educational context” (volet, 2001) that we have used throughout the article. apart from differences emerging between the accounts of the australian and german teacher educators, it is important to also note their shared beliefs. for example, beliefs regarding the necessity to control the display of negative emotions by simultaneously maintaining a relatively high degree of authenticity (an explicit goal in individualistic countries, safdar et al., 2009), were shared by australian and german teacher educators alike, which supports previous findings. for example, mendzheritskaya and hansen (2013) (see also mendzheritskaya, hansen & horz, 2015) found that german and russian he lecturers showed a greater authenticity in displaying positive emotions, which was also the case in the present sample, while negative emotions were displayed with lower expressivity based on the belief that negative emotions must be controlled in a professional setting. another aspect shared by many teachers in both countries, in particular in australia, was the importance of displaying enthusiasm for the teaching subject in class. this is reminiscent of what neumann (2006) referred to as “passionate thought” inspired by scholarly work in the respective subject area, and which should, according to neumann, be transferred to students through an enthusiastic teaching practice in order to cultivate their motivation for the subject. finally, a high variety of emotion display modes within each cultural-educational context were detected, which could be traced back to different factors on the individual level, such as personality (matsumoto, 2006). this supports the findings resulting from a meta-analysis done by van hemert, poortinga, and van de vijver (2007) who have found several moderators that reduced the cross-cultural variance in emotion variables. a particular relevant factor influencing the display mode of emotions and the quality of the tsr seems to be the dedication for teaching in general and teaching in teacher education specifically. as the comparative case-analysis has shown, teachers who identified strongly with the teaching profession, reported closer relationships with their students and also expressed more intensive positive emotions. furthermore, position at university and prior teaching experience (in he and at school) also appeared to contribute to the quality of the tsr and the display of emotions. 5. conclusion: study limitations and implications the results of the present study have significant implications particularly for research in the internationalization of he. while there is extensive literature on the experiences of international students in he, these experiences are under-explored for he faculty (bedenlier & zawacki-richter, 2015; shimmi, 2014). our interview accounts have revealed different modes of forming relationships with students and displaying emotions in teaching, depending on the cultural-educational context. there was also evidence that new teachers from a cultural background different from the local context can feel insecure about how to behave in an unfamiliar classroom environment. as reported in volet and jones’ (2012) review of literature, scholars teaching in another country may need support to adapt to new and culturally different teaching environments. many strategies may support their adaptation. for example, in their first semester, an hagenauer(et (al ( ( ( | f l r ! ! 68! international and a national scholar could conduct joint courses, which would allow the scholars to share how students and teachers typically interact in the respective countries. in doing so, the international scholar could develop a sense of the “academic habitus” that prevails in the host country. reflection groups that consist of international teachers and at least one national teacher meeting regularly to exchange and discuss experiences may be useful. darwin and palmer (2009) introduced the idea of “mentoring circles”, which they present as a form of group mentoring aimed at staff development in a he context. mentoring circles could be appropriate environments to assist international scholars to adapt to their new work environment. more generally, luna and cullen (1995) describe mentoring as a key principle in empowering faculty, with special strategies for mentoring minority faculty (e.g., scholars with a different cultural background). helping new academics from a different cultural-educational background to adapt to local teaching practices is important but it is particularly critical in the field of teacher education. this is because teacher educators do not only teach but are also expected to model teaching behaviour to prospective teachers. therefore they should be aware, not only of appropriate teaching practices in the he field, but also of those expected in the school context. for example, they should become familiar about the feedback practices in the new country (how is it expected to communicate satisfaction with a students’ achievement, whether in an emotionally intense way or in a more reserved way?), and about the degree of informality or formality in teacher-student interactions at school. furthermore, as bridging educational theory with educational practice (garcia-aracil, 2012) is an important contributing factor to study satisfaction, teacher educators coming from another cultural-educational background should gain in-depth understanding of the educational practices of the new country. more generally, fostering socio-emotional competence (jennings & greenberg, 2009) in teacher educators, for example by analyzing classroom instructional and interaction practices through videos (for schoolteachers, see i.e. seidel et al., 2011), appears an important avenue in the professionalizing of teacher educators. the finding of differences in the perceived-as-appropriate display rules of emotions in educational settings, as well as the quality of the tsr, suggests that teaching strategies cannot be considered as techniques that can simply be transferred to different cultural-educational settings. international educators need to become aware of the frequently unexpressed and implicit cultural specifics of teaching in he classrooms. awareness of cultural specifics of their own familiar educational environments as well as those of the new environment, would not only promote teaching quality from the outset but also teachers’ and students’ satisfaction in the teaching learning process. diener and diswas-diener’s (2008) research has found that satisfaction with relationships is the most important predictor for life-satisfaction. developing competence in a range of cultural appropriate display rules can, thus, be regarded as an important factor for maintaining teachers’ wellbeing (woods, 2010) in international education settings at home or abroad (for a critical discussion on the increased emphasis on emotions and emotional well-being in education, see for example ecclestone & hayes, 2009; ecclestone, 2012). many directions for future research emerge from the present study. methodologically, there is a need for more situated approaches to explore teacher emotions in the classroom. for example, systematic (video-) observations in concrete teaching situations, combined with subsequent stimulated recall interviews, would generate rich insights on the cultural-specifics of teacher-student interactions across countries, as well as culturally diverse classroom settings. video-observations could be used together with teacher questionnaires (e.g. using the “display rule assessment inventory” adapted to the he context; see mendzheritskaya and hansen 2013), following a mixed-methods approach (creswell & plano-clark, 2011). the present study relied on self-reports of teachers, which may have drawn some responses that were affected by social desirability bias to some extent. furthermore our study addressed emotions and relationships directly. other research – for example the study by postareff and lindblom-ylänne (2011) – did not explore directly he teachers’ emotions, but issues related to emotions nevertheless emerged from the answers given to other interview questions. both ways of exploring emotions are possible avenues in emotion research and the appropriateness of the chosen approach depends on the research aim and the sample (e.g in terms of their willingness and ability to reflect educational practices including the own emotional and mental processes). future studies should combine such subjective measures with more objective measures, regardless of hagenauer(et (al ( ( ( | f l r ! ! 69! whether emotions are assessed through direct or indirect questioning in the subjective-coined part of the study. it also seems necessary to develop quantitative measure that would make it possible to assess the multidimensionality of the quality of the tsr in he (hagenauer & volet, 2014c). this could involve the development of coding schemes to record systematic observations of teacher-student interactions; coding schemes that would allow inferences to be drawn on the quality of the tsr. furthermore, correlational studies should follow as well as longitudinal studies exploring how display modes might affect the establishment of relationships and vice versa. in terms of limitations, the present study was conducted with teacher educators, which suggests caution in generalizing to university teachers in other domains. as argued by alheit (2009), teaching practices vary across domains, departments, and institutions due to differences in the academic habitus. the impact of moderating variables should therefore be examined systematically in future studies, for example, the teaching domain (e.g. soft versus hard sciences), teachers’ role and status in the university (professor, lecturer, phd-student with teaching duties, casual) (richardson & radloff, 2014), the amount and quality of prior teaching experience (meanwell & kleiner, 2014), and the value placed on teaching at the respective university. the latter might also influence teachers’ identification as a teacher and his/her motivation for teaching, which might in turn impact on the emotional labour required in teaching (visser-wijnveen, stes, & van petegem, 2014). by taking these factors into account, variability in emotion display and the relationship quality in different cultural-educational contexts could be better understood. on an interpersonal level, gender, personality (e.g., the big five or people’s idiocentric vs. allocentric tendencies) as well as ethnicity and cultural background should be accounted for. in terms of cultural background, decuir-gunby and williams-johnson (2014) discussed people’s difficulties to read the expression of emotions of culturally distant others. this is an important issue if international students are taught in he classrooms or if teachers teach in a foreign or ethnically unfamiliar context. in addition, by applying an autobiographical approach it would be possible to explore how the past experiences or the “personal history”, as coined by day and leitch (2001), could contribute to our understanding of emotion display and tsr-building processes in he settings. finally, some cultures may be more heterogeneous than others. in our study, the australian sample comprised a few immigrants, while the german teacher educators were all local. furthermore, it is well known that equating country with culture is problematic (e.g. matsumoto et al., 1998; volet & jones, 2012), therefore there should be caution in the interpretation of the findings, keeping in mind that educational practices are constantly changing in increasingly global environments, and most importantly also reflect individual preferences. concluding, as observed by van hemert et al. (2007), studies on cultural issues in he teaching should go beyond frequently made comparisons between individualistic and collectivistic cultures or countries. this study revealed a high variation of cultural-educational practices within two socalled “individualistic” countries (hofstede & hofstede, 2005). this calls for a reconceptualization of crosscultural research on the relationship between emotion display and the tsr in he teaching contexts. keypoints insight into university teachers’ perceptions of the characteristics of quality tsr and appropriate emotion display is essential in the context of the fast developing internationalization of higher education. this study revealed major qualitative differences in the display of positive and negative emotions in two distinct cultural-educational contexts. the study also unveiled differences across cultural-educational contexts in the importance given by teachers to keeping professional and formal working relationship with students or alternatively showing informality and caring in the interpersonal tsr. hagenauer(et (al ( ( ( | f l r ! ! 70! international scholars and teachers, like mobility students, need support to adapt to unfamiliar cultural-educational environments. references alheit, p. (2009). die symbolische macht des wissens. exklusionsmechanismen des universitären habitus [the symbolic power of knowledge. mechanism of exclusion of the academic habitus]. presentation held at the university of heidelberg (3.6.2009). altbach, p. g., & knight, j. (2007). the internationalization of higher education: motivations and realities. journal of studies in international education, 11 (3), 290-305. doi: 10.1177/1028315307303542 barber, l. k., grawitch, m. j., carson, r. l., & tsouloupas, c. n. (2010.) costs and benefits of supportive versus disciplinary emotion regulation strategies in teachers. stress and health, 27, e173-e187. doi: 10.1002/smi.1357. bartram, b. (2015). emotion as a student resource in higher education. british journal of educational studies, 63 (1), 67-84. doi: 0.1080/00071005.2014.980222 baumeister, r. f., & leary, m. r. (1995). the need to belong: desire for interpersonal attachments as a fundamental human motivation. psychological bulletin, 117, 497-529. doi: 10.1037/00332909.117.3.497 beard, c., clegg, s., & smith, k. (2007). acknowledging the affective in higher education. british educational research journal, 33 (2), 235-252. doi: 10.1080/01411920701208415 bedenlier, s., & zawacki-richter, o. (2015). internationalization of higher education and the impacts on academic faculty members. research in comparative and international education, 10 (2), 185-201. online first: doi: 10.1177/1745499915571707 boiger, m., & mesquita, b. (2012). the construction of emotion in interactions, relationships, and cultures. emotion review, 4 (3), 221-229. doi: 10.1177/1754073912439765 brennan, r. l., & prediger, d. j. (1981). coefficient kappa: some uses, misuses, and alternatives. educational and psychological measurement, 41, 687-699. doi: 10.1177/001316448104100307 bortz, j., & döring, n. (2006). forschungsund untersuchungsplanung. heidelberg: springer. chen, y.-h., & kristjansson, k. (2011). private feelings, public expression: professional jealousy and the moral practice of teaching. journal of moral education, 40 (3), 349-358. doi: 10.1080/03057240.2011.596336 cranton, p., & carusetta, e. (2004). perspectives on authenticity in teaching. adult education quarterly, 55 (5), 5-22. doi: 10.1177/0741713604268894 creswell, j. w., & plano clark, v. l. (2011). designing and conducting mixed methods research (2nd ed.). los angeles: sage. darwin, a., & palmer, e. (2009). mentoring circles in higher education. higher education research & development, 28 (2), 125-136. doi: 10.1080/07294360902725017 day, c., & leitch, r. (2001). teachers’ and teacher educators’ lives: the role of emotion. teaching and teacher education, 17, 403-415. doi: 10.1016/s0742-051x(01)00003-8 decuir-gunby, j., & williams-johnson, m. r. (2014). the influence of culture on emotions. implications for education. in r. pekrun, & l. linnenbrink-garcia (eds.), international handbook of emotions in education (pp. 539-557). new york: routledge. deci, e. l., & ryan, r. m. (eds.). (2002). handbook of self-determination research. rochester: university of rochester press. diener, e., & biswas-diener, r. (2008). happiness: unlocking the mysteries of psychological wealth. malden, ma: blackwell publishing. ecclestone, k. (2012). from emotional to psychological well-being to character education: challenging policy discourses of behavioural science and “vulnerability”. research papers in education, 27 (4), 463-480. doi: 10.1080/02671522.2012.690241 hagenauer(et (al ( ( ( | f l r ! ! 71! ecclestone, k., & hayes, d. (2009). changing the subject: the educational implications of developing emotional well-being. oxford review of education, 35 (3), 371-389. doi: 10.1080/03054980902934662 eid, m., & diener, e. (2001). norms for experiencing emotions in different cultures: inter-and intranational differences. journal of personality and social psychology, 81 (5), 869-885. doi: 10.1037//00223514.81.5.869 ellsworth, p. c., & scherer, k. r. (2003). appraisal processes in emotion. in r. j. davidson, k. r. scherer, & h. hill goldsmith (eds.), handbook of affective sciences (pp. 572-595). oxford: university press. fischer, a. h., & manstead, a. s. r. (2010). social functions of emotions. in m. lewis, j. m. havilandjones, & l. feldman barrett (eds.), handbook of emotions (3rd ed.) (pp. 456-468). new york: guilford press. fischer, a. h., manstead, a. s. r., evers, c., timmers, m., & valk, g. (2004). motives and norms underlying emotion regulation. in p. phillippot & r. s. feldman (eds.), the regulation of emotion (pp. 187-210). mahwah, nj: lawrence erlbaum. frenzel, a. (2014). teacher emotions. in r. pekrun, & l. linnenbrink-garcia (eds.), international handbook of emotions in education (pp. 494-519). new york: routledge. frenzel, a. c., goetz, t., luedtke, o., pekrun, r., & sutton, r. (2009). emotional transmission in the classroom: exploring the relationship between teacher and student enjoyment. journal of educational psychology, 101, 705-716. doi: 10.1037/a0014695 garcia-aracil, a. (2012). a comparative analysis of study satisfaction among young european higher education graduates. irish educational studies, 31, 223-243. doi: 10.1080/03323315.2012.660605 garner, p. w. (2010). emotional competence and its influences on teaching and learning. educational psychology review, 22, 297-321. doi: 0.1007/s10648-010-9129-4 gates, g. s. (2000). the socialization of feelings in undergraduate education: a study of emotional management. college student journal, 34 (4), 485-504. gläser-zikuda, m. & mayring, ph. (2003). a qualitative approach to learning emotions at school. in ph. mayring, & ch. v. rhöneck (eds.), learning emotions. the influence of affective factors on classroom learning (pp. 103-126). berlin: peter lang. gross, j. j. (2010). emotion regulation. in m. lewis, j. m. haviland-jones, & l. feldman barrett (eds.), handbook of emotions (3rd ed.) (pp. 497-512). new york: guilford press. gross, j. j. (2002). emotion regulation: affective, cognitive, and social consequences. psychophysiology, 39 (3), 281-291. doi: 10.1017.s0048577201393198 hagenauer, g., hascher, t., & volet, s. e. (2015). teacher emotions in the classroom: associations with students’ engagement, discipline in the classroom and the interpersonal teacher-student relationship. european journal of psychology of education, 30 (4), 385-403. doi: 10.1007/s10212-015-0250-0 hagenauer, g., & volet, s. e. (2014a). “i don't think i could, you know, just teach without any emotion”: exploring the nature and origin of university teachers’ emotions. research papers in education, 29 (2), 240-262. doi: 10.1080/02671522.2012.754929 hagenauer, g., & volet, s. e. (2014b). “i don’t hide my feelings, even though i try to”: insight into teacher educator emotion display. australian educational researcher, 41 (3), 261-281. doi: 10.1007/s13384013-0129-5 hagenauer, g., & volet, s. e. (2014c). student-teacher relationship at university: an important yet underresearched field. oxford review of education, 40 (3), 370-388. doi: 10.1080/03054985.2014.921613 hascher, t., & hagenauer, g. (2016). openness to theory and its importance for pre-service teachers’ selfefficacy, emotions, and classroom behaviour in the teaching practicum. international journal of educational research, 77, 15-25. doi: 10.1016/j.ijer.2016.02.003 hastings, w. (2008). i felt so guilty: emotions and subjectivity in school-based teacher education. teachers and teaching: theory and practice, 14 (5-6), 497-513. doi: 10.1080/13540600802583655 hatfield, e., bensman, l., thornton, p. d., & rapson, r. l. (2014). new perspectives on emotional contagion: a review of classic and recent research on facial mimicry and contagion. interpersona, 8 (2), 159-179. doi:10.5964/ijpr.v8i2.162 hagenauer(et (al ( ( ( | f l r ! ! 72! hochschild, a. r. (1983). the managed heart: commercialization of human feelings. berkeley: university of california press. hofstede, g., & hofstede, g. j. (2005). cultures and organization: software of the mind. london: mcgrawhill. isenbarger, l., & zembylas, m. (2006). the emotional labour of caring in teaching. teaching and teacher education, 22, 120-134. doi: 10.1016/j.tate.2005.07.002 jennings, p. a., & greenberg, m. t. (2009). the prosocial classroom: teacher social and emotional competence in relation to student and classroom outcomes. review of educational research, 79, 491-525. doi: 10.3102/0034654308325693 könig, j., & blömeke, s. (2013). germany. in j. schwille, l. ingvarson, & holdgreve-resendez (eds.), teds-m encyclopedia. a guide to teacher education context, structure, and quality assurance in 17 countries. retrieved from http://www.iea.nl/fileadmin/user_upload/publications/electronic_versions/teds-m_encyclopedia.pdf koopman-holm, b., & matsumoto, d. (2011). values and display rules for specific emotions. journal of cross-cultural psychology, 42, 355-371. doi: 10.1177/0022022110362753 lazarus, r. s. (1999). stress and emotion. a new synthesis. london: free association books. luna, g., & cullen, d. l. (1995). empowering the faculty. mentoring redirected and renewed. ashe-eric higher education report no. 3. washington, d. c.: the george washington university, graduate school of education and human development. lunenberg, m., korthagen, f., & swennen, a. (2007). the teacher educator as role model. teaching and teacher education, 23, 586-601. doi: 10.1016/j.tate.2006.11.001 markus, h. r., & kitayama, s. (1991). culture and the self: implications for cognition, emotion, and motivation. psychological review, 98, 224-233. doi: 10.1037/0033295x98.2.224 matsumoto, d. (2006). are cultural differences in emotion regulation mediated by personality traits? journal of cross-cultural psychology, 37, 421-437. doi: 10.1177/0022022106288478 matsumoto, d., takeuchi, s., andayani, s., kouznetsova, n., & krupp. d. (1998). the contribution of individualism vs. collectivism to cross-national differences in display rules. asian journal of social psychology, 1, 147-165. doi: 10.1111/1467-839x.00010 mayring, p. (2000). qualitative content analysis. forum qualitative social research, 1 (2). retrieved from http://www.qualitative-research.net/index.php/fqs/article/view/1089/2385 meanwell, e., & kleiner, s. (2014). the emotional experience of first-time teaching: reflections from graduate instructors, 1997-2006. teaching sociology, 42 (1), 17-27. doi: 10.1177/0092055x13508377 mendzheritskaya, j., & hansen, m. (2013, august). shall i show my anger? display rules in lecturer-student interaction in germany and russia. paper presented at the 15th biennal earli conference. munich, germany. mendzheritskaya, j., hansen, m., & horz, h. (2015). emotional display rules at universities in russia and germany. russian psychological journal, 12 (4). mesquita, b. (2007). emotions are culturally situated. social science information, 46, 410-415. doi: 10.1177/05390184070460030107 moore, s., & kuol, n. (2007). matters of the heart: exploring the emotional dimensions of educational experience in recollected accounts of excellent teaching. international journal of academic development, 12 (2), 87-98. doi: 10.1080/13601440701604872 moran, c. m., diefendorff, j. m., & greguras, g. j. (2013). understanding emotional display rules at work and outside of work: the effects of country and gender. motivation and emotion, 37, 323-334. doi: 10.1007/s11031-012-9301-x neumann, a. (2006). professing passion: emotion in scholarship of professors at research universities. american educational research journal, 43 (3), 381-424. retrieved from http://www.jstor.org/stable/4121764 newberry, m., gallant, a., & riley, p. (eds.). (2013). emotion and school: understanding how the hidden curriculum influences relationships, leadership, teaching, and learning. uk: emerald. nias, j. (1989). primary teacher talking. a study of teaching as work. london: routledge. hagenauer(et (al ( ( ( | f l r ! ! 73! oplatka, i. (2007). managing emotions in teaching: towards an understanding of emotion displays and caring as nonprescribed role elements. teachers college record, 109 (6), 1374-1400. parkinson, b., fischer, a. h., & manstead, a. s. r. (2005). emotion in social relations. new york: psychology press. pekrun, r. (2006). the control-value theory of achievement emotions: assumptions, corollaries, and implications for educational research and practice. educational psychology review, 18, 315-341. doi: 10.1007/s10648-006-9029-9 pillen, m., beijaard, d., & den brok, p. (2013). professional identity tensions of beginning teachers. teachers and teaching: theory and practice, 19 (6), 660-678. doi: 10.1080/13540602.2013.827455 postareff, l., & lindblom-ylänne, s. (2011). emotions and confidence within teaching in higher education. studies in higher education, 36, 799-813. doi: 10.1080/03075079.2010.483279 quinlan, k. m. (2016). how emotions matters in four key relationships in teaching and learning in higher education. college teaching. online first: doi: 10.1080/87567555.2015.1088818 ria, l., sève, c., saury, j., theureau, j., & durand, m. (2003). beginning teachers’ situated emotions: a study of first classroom experiences. journal of education for teaching, 29 (3), 219-233. doi: 10.1080/0260747032000120114 richardson, s., & radloff. a. (2014). allies in learning: critical insights into the importance of staff-student interactions in university education. teaching in higher education, 19 (6), 603-615. doi: 10.1080/13562517.2014.901960 safdar, s., friedlmeier, w., matsumoto, d., yoo, s. h., kwantes, c., kakai, h., & shigemasu, e. (2009). variations in emotional display rules within and across cultures: a comparison between canada, usa, and japan. canadian journal of behavioural science, 41 (1), 1-10. doi: 10.1037/a0014387 schutz, p. a. (2014). inquiry on teachers’ emotion. educational psychologist, 49, 1-12. doi: 10.1080/00461520.2013.864955 schutz, p. a., & zembylas, m. (eds.). (2011). advances in teacher emotion research. the impact on teachers’ lives. heidelberg: springer. schwarz, s. h., & ros, m. (1995). values in the west: a theoretical and empirical challenge to the individualism-collectivism cultural dimension. world psychology, 1, 99-122. seidel, t., stürmer, k., blomberg, g., kobarg, m., & schwindt, k. (2011). teacher learning from analysis of videotaped classroom situations: does it make a difference whether teachers observe their own teaching or that or others? teaching and teacher education, 27, 259-267. doi: 10.106/j.tate.2010.08.009 shimmi, y. (2014). experiences of japanese visiting scholars in the united states: an exploration of transition. phd, boston college, 2014. retrieved from http://hdl.handle.net/2345/3783 sutton, r. e., & wheatley, k. f. (2003). teachers’ emotions and teaching: a review of the literature and directions for future research. educational psychology review, 15, 327-358. doi: 10.1023/a:1026131715856 titsworth, s., mckenna, t. p., mazer, j. p., & quinlan, m. m. (2013). the bright side of emotion in the classroom: do teachers’ behaviors predict students’ enjoyment, hope, and pride? communication education, 62 (2), 191-209. doi: 10.1080/03634523.2013.763997 timostsuk, i., & ugaste, a. (2012). the role of emotions in student teachers’ professional identity. european journal of teacher education, 35 (4), 421-433. doi: 10.1080/02619768.2012.662637 trigwell, k. (2012). relations between teachers’ emotions in teaching and their approaches to teaching in higher education. instructional science, 40, 607-621. doi: 10.1007/s1125-011-9192-3 tsai, j. l., knutson, b., & fung, h. h. (2006). cultural variation in affect valuation. journal of personality and social psychology, 90, 288-307. doi: 10.1037/0022-3514.90.2.288 van hemert, d. a., poortinga, y. h., & van de vijver, f. j. r. (2007). emotion and culture: a metaanalysis. cognition and emotion, 21 (5), 913-943. doi: 10.1080/02699930701339293 visser-wijnveen, g. j., stes, a., & van petegem, p. (2014). clustering teachers’ motivations for teaching. teaching in higher education, 19 (6), 644-656. doi: 10.1080/13562517.2014.901953 hagenauer(et (al ( ( ( | f l r ! ! 74! volet, s. (2001). understanding learning and motivation in context: a multi-dimensional and multi-level cognitive-situative perspective. in s. järvelä, & s. volet (eds.)., motivation in learning contexts: theoretical advances and methodological implications (pp. 57-82). elmsford, ny: pergamon. volet, s. e., & jones, c. (2012). cultural transitions in higher education: individual adaptation, transformation and engagement. in s. karabenick, & t. urdan (eds.), advances in achievement and motivation series (vol. 17): transitions across schools and cultures (pp. 241-284). bingley, uk: emerald. walker, c., & gleaves, a. (2016). constructing the caring higher education teacher: a theoretical framework. teaching and teacher education, 54, 65-76. doi: 10.1016/j.tate.2015.11.013 woods, c. (2010). employee wellbeing in the higher education workplace: a role for emotion scholarship. higher education, 60, 171-185. doi: 10.1007/s10734-009-9293-y yuu, k. (2010). expressing emotions in teaching: inducement, suppression, and disclosure as caring profession. educational studies in japan. international yearbook, 5, 63-75. zhang, q., & zhang, j. (2013). instructors’ positive emotions: effects on student engagement and critical thinking in u.s. and chinese classrooms. communication education, 62 (4). 395-411. doi: 10.1080/03634523.2013.828842 zhang, q., & zhu, w. (2008). exploring emotion in teaching: emotional labor, burnout, and satisfaction in chinese higher education. communication education, 57 (1), 105-133. doi: 10.1080/03634520701586310 frontline learning research vol. 5 no 3 special issue (2017) 155-166 issn 2295-3159 the interplay between methodologies, tasks and visualisation formats in the study of visual expertise jean-michel boucheix1 *lead-cnrs, university of bourgogne franche-comté, dijon, france 1. introduction as underlined in the issue by andreas gegenfurtner & jeroen merriënboer (2017, this issue), "expertise can be defined as maximal adaptations to task constraints" (ericsson & lehmann, 1996; gruber, jansen, marienhagen, & altenmueller, 2010). this definition implicitly emphasizes the multimodal dimension of expertise, both cognitive and perceptual. the study of expertise, notably visual expertise (or expert cognition in visual tasks), has a long research tradition, and several important discoveries and theoretical advances have been made and developed within the framework of ericsson’s ‘expert performance approach’ (see williams, fawver, & hodges, this issue). however, the recent refinement and widespread use of technologies such as eye tracking, and the development of new methods such as direct brain investigations (fmri, fnirs etc.) have shed new light on "visual expertise" within the broad perception-cognition interplay system. this special issue on “methodologies for studying visual expertise” offers a unique opportunity to discover the diversity and complementarities of the methods and techniques used to analyse expertise and its development. in contrast to conventional special issues, it does not comprise empirical studies, but a series of research reviews presenting and analysing different methods or frameworks. thus, the first paper offers a sort of methodological guide using ericsson & smith’s (1991) "expert performance approach". this is followed by three papers analysing the use of eye tracking in visual expertise models, and a paper reviewing the use of methods such as eeg and fmri to investigate the neural correlates of visual perceptual expertise. a paper on receiver operating characteristic (roc) analysis brings new insight to the well-known signal detection methods used in control and visual inspection tasks in industry. finally, three papers review methods based on verbal reports and protocols used to investigate the conceptual and meta-cognitive levels of visual expertise. in the following sections, we will (i) describe the methodological contribution of each paper according to five main themes or criteria (see table 1), and (ii) discuss the benefits, limitations, and particularly the opportunities offered by new developments in each methodology and their possible combinations. 1e-mail : jean-michel.boucheix@u-bourgogne.fr doi : http://dx.doi.org/10.14786/flr.v5i3.311 mailto:jean-michel.boucheix@u-bourgogne.fr commentary boucheix 156 | f l r 2. methodological contributions as shown in table 1, the methodological contributions and results differ for each criterion. table 1 summary of each methodological review authors and method investigated visual expertise level process level/ models tasks: -goals, measures materials-stimuli "artificial"/"realistic" static vs. dynamic interactivity, 2/3d results groups of interest williams, fawver, hodges expert performance approach multiple levels of performance (capture, process, learning) long-term working memory (ericsson) detailed measures of learning efficiency process tracing measures visual anticipation decision making (sports) realistic, "quasiecological" scenes film-based simulation: static and dynamic visual stimuli promise of virtual reality (embodiment), 2d -transfer of expertise to several domains experts, novices, learners, how experts refine skills -individual differences krupinski roc analysis based on tds visual detection /discrimination less on process -target detection and decision sensitivity, specificity, accuracy area under the curve metric "realistic" (or modified) medical images: radiology industry (visual inspection of objects) static images (2d) measures of sensitivity, specificity and accuracy mostly experts and novices gegenfurtner kok, van geel, de bruin, sorger functional neuro imaging neural correlates of visual perceptual expertise: eeg fmri -target detection detection tasks decision making (manipulated images) viewing judgments artificial (greebles) and realistic stimuli medical domain: radiology mostly static single images (2d) no interactivity eeg: temporal adaptation of expertise: n170 fmri: spatial adaptation of expertise: ffa (multiple brain areas: cortico-spinal activation, temporal and frontal gyri, left inferior frontal sulcus, posterior cingulate cortex) experts novices fox, faulknerjones eye tracking eye movements -top-down /bottom up dynamics: perceptualcognitive diagnosis focal disease detection gaze fixations, saccades, scan paths. errors realistic medical stimuli radiology target search, recognition diagnosis mostly static single images (2d) no interactivity hand-eye tasks in procedural skills: laparoscopic tasks experts: fewer fixations, less time on aois search-related tasks, top-down, less feed-back needed. holistic/analytic phases: gestalt guided/hybrid search models unexpected paths priority map, attentional priority -assistive technology inattentional blindness missing rare events maintaining eyegaze skills evaluation mostly experts novices (learners) szulewski, kelton, howes cognitive load cl quantification medical diagnosis decision making from visual scenes. medical scenes visual laparoscopic tasks cl quantification pupil dilation (light condition limits) experts novices learners commentary boucheix 157 | f l r pupillometry eye tracking mostly static (some dynamic) 2d, no interactivity litchfield, donovan -fpmwp -flash-preview moving window paradigm holistic model effect of initial glimpse on subsequent search behaviour interplay of holistic/analytic processes target detection focal tumour detection decision making manipulation of the fpmwp variables/measures. everyday visual scenes medicine: radiology images static 2d images no interactivity inconsistent effects first preview impairs experts’ performance experts novices helle combination eye-tracking/ verbal reports speech production visual search conscious and meta-cognitive level, performance levels -verbal reports: thinking aloud, retrospective thinking aloud medical image interpretation radiology dermatology mostly static -2d effectiveness: verbalization time limits <10sec. experts (novices) van de wiel interviews verbal protocols focus groups structured interviews conscious and (meta)-cognitive planning task representation size anticipation free recall explaining reports coding verbal reports medical images and scenes medical tasks: problem solving reasoning insight into the tasks and materials to be selected guide mostly experts ivarsson ethno methodology embodiment of visual expertise visual expertise grounded in body actions communication of visual analysis tasks, diagnosis explanation, gestures video-tapes of experts radiology (thoracic) images mostly static (reference to dynamic images) embodiment results: effect of visual expertise or of explanatory task? experts (few) the next section discusses the specific contributions of the methodologies. 3. lessons and opportunities of each methodology 3.1. different levels of visual and cognitive processes two main levels of visual expertise are explored. first, fine-grained visual-cognitive processes and performance are addressed in the papers on eye tracking by fox & faulkner-jones, litchfield & donovan, and krupinski; in the paper concerning functional neuro-imaging by gegenfurtner, kok, van geel, de bruin, & sorger; and in the paper about pupillometry by szulewski, kelton & howes. of particular interest is their analysis and description of the temporal course of the interplay between holistic and more analytic processing of the visual stimuli with their neural correlates. secondly, the papers on verbal reports by van de wiel and by ivarsson introduce different top-down processing levels, which seem to be more conscious, and (meta)cognitive levels. these include task representation, planning activities involving dimensions such as the extent of anticipation, and the heuristic level used to plan the visual search. the relationship between top-down and bottom-up processes is addressed in all the papers, but particularly by helle who describes a methodology that is well-suited to this question, involving a combination of verbal reports and eye tracking. commentary boucheix 158 | f l r the two neuro-imaging methods described by gegenfurtner, kok, van geel, de bruin, & sorger involve different processing levels: temporal for eeg (n170), and spatial for fmri (the face fusiform area (ffa) and other brain regions). in szulewski, kelton, & howes’ investigation of pupillometry, the authors describe how changes in pupil diameter across time reflect cognitive load levels and the relationship between cognitive load and expertise level. an important aspect concerns differentiating between the specific cognitive processes involved, and whether they are global or local. the authors refer to sweller's distinction between icl (intrinsic cognitive load), ecl (extraneous cognitive load), and gcl (germane cognitive load). is it possible to vary icl, ecl and gcl using the same type of task and the same visual environment stimuli? however, it is more difficult to assess the specific involvement of different cognitive processes, such as information storage, manipulation of information, executive control, etc. it would be interesting to manipulate variables to test pupil diameter reactions when participants perform different types of task (memory, storage, cognitive control) involving different kinds or levels of information processing. for example, in the developing area of neuroergonomics, durantin, gagnon, tremblay, & dehais (2014) measured the cognitive load level of easy vs. difficult flight tasks for aircraft pilots using fnirs (prefrontal dorso-lateral area) and pupil dilation, with similar results to those shown in figure 1 of szulewski, kelton, & howes’s (this issue). finally, in the visual search expertise models described by fox & jones (e.g. guided search model, priority map and hybrid search), measures such as saccades and scan path could be used to study chunking mechanisms. as shown in table 1, different tasks are used for each methodology, yielding apparently varied results. 3.2. task goals, task design and outcomes. as shown in table 1, in most of the methods reviewed in the nine papers, the results of previous research, and the nature of the cognitive and visual search processes involved, differ significantly and are sometimes inconsistent, even when the same visual material is used. with medical images, such as chest x-rays, or in dermatology, the results of the experiments about the performance of experts and novices seem to differ for simple target detection (nodules), diagnosis, categorization, and decision making. in their review of studies of functional neuro-imaging methods, gegenfurtner, kok, van geel, de bruin, & sorger found clear visual expertise effects in the use of eeg (n170) and fmri (fusiform face area ffa activity). however, they also suggest (in their table 1) that the brain regions that are activated depends not only on the method (x-ray, fmri etc.), stimuli (medical vs. non-medical) and expertise, but also on the type of task (e.g. viewing, detection, decision-making, reasoning, image manipulation) and task complexity. for example, the activated regions may not be associated with object recognition (e.g. ffa) but with attention and memory. as stated by the authors, "cognitive neuroscience examining expertise in medical image diagnosis is promising but still in its infancy", and further research is urgently needed in this area. more generally, the question is raised about how tasks can be customized in order to assess specific processing mechanisms, without moving too far from the ecological process: in other words, the art of simplifying the task without changing it too much. in their review of eye-tracking methods, fox & faulkner-jones also found an effect of task type in top-down and bottom-up dynamics, when detecting perturbations or deviance from expectations. similarly, task differences may also raise questions about the initial holistic phase of image processing in nodine & kundel’s (1987) model. another crucial aspect of the effect of task design on performance and cognitive processes in visual expertise is illustrated in litchfield & donovan’s review of the flash-preview moving window paradigm (fpmw). as described by the authors, in the fpmw paradigm (taken from castelhano & henderson, 2007), "observers are briefly shown a preview of the upcoming search scene and are subsequently asked to search for a particular target object whilst their peripheral vision is restricted to a gaze-contingent moving window". in the framework of the holistic model of image processing, a positive flash preview effect of the whole image commentary boucheix 159 | f l r (250ms) has been shown in detection tasks in experts, but not, or much less, in novices. according to litchfield and donovan, “it is within this initial glimpse of the scene –gistthat subsequent eye movements are guided (castelhano & henderson, 2007) with parafoveal and peripheral vision playing an important role in the early comprehension of the gist of a scene and in the detection of targets (henderson, pollatsek, & rayner, 1989)”. however, unexpectedly, experiments conducted with experts and novices using the fpmw method (litchfield & donovan, 2016; see also donovan & litchfield, 2013) found that experts were impaired at identifying target abnormalities (lung nodules, brain tumours, bone fractures) if shown a preview. as stated by the authors, “it is still unclear whether it is an expertise-specific advantage in global processing that specifically contributes to search guidance". these results of the fpmw method involving a particular visual activity raise questions and concerns about the holistic model. we can hypothesize that a potential negative effect of flash previews could also arise in the case of dynamic pictures, such as animations or videos: with this dynamic material, it could be necessary to stop earlier global processing in order to find the boundaries of temporal processes (e.g. meteorology maps, lowe, 1999). this aspect will be discussed further below. in the methods based on verbal reports, van de wiel examined expertise using interviews and verbal protocols and found that the type of interview (the method) and the task determine the data that is collected and the results of the investigation. finally, regarding the ethno-methodology approach, ivarsson followed a group of radiologists and radio physicists who were told to find, discuss, and formulate issues or problems ensuing from the implementation of new radiographic imaging technology. using their visual expertise, they had to communicate their analysis of the image, not only verbally but also with gestures: this has been termed the “enacted production of radiological reasoning”. analysis of the data suggested that visual expertise was, partly at least, due to embodied practice and ability. however, and in relation with a task effect, we also suggest that the embodiment effects that are described as the main finding and purpose of the paper could be partially produced by the communication task, which involved gestures and bodily demonstrations. this effect does not necessarily mean that a visual radiological task (e.g. tumour detection) in everyday practice involves this type of embodiment. it is possible that the experimental task forced them to embody their analysis in order to explain and to be understood. 3.2.1. making relations and drawing inferences many of the studies reviewed in the papers focus on visual search, and apart from the reviews and studies of verbal reports and the global expertise framework (helle, van de wiel, ivarsson, and williams, fawver & hodges), the task (with performance measures) involves the detection of targets or abnormalities. however, visual tasks involving more complex diagnoses or interpretation require making relations between two or more components of the visual scene, or between targets situated at more or less distant locations (see schwan, 2013). in future research with eye-tracking and functional brain imaging methods, it would be very interesting to analyse how experts and novices process or construct such relations. this would be particularly relevant for developing learning tasks for students who are at a stage between novices and experts. regarding this aspect of visual processing, litchfield and donovan, in their review of fpmw studies, make an interesting remark about the re-direction of eye gaze according to time-locked acquisition of new information during the visual processing of medical images. this raises a crucial question about dynamic chunking during these redirections. how does the new information interact with previously acquired knowledge and re-direct or drive the new gaze on the visual scene? processing relations efficiently could be even more important to interpret dynamic visual scenes. the visual format of images is the subject of the next section. 3.3. format of the visual material and interactive features table 1 shows that in most previous research on visual expertise, especially in the medical area (except in william, fawver & hodges’ paper) the format of the visual scenes is limited to conventional 2d images, often presented as a single image. however, with technological advances, images can now be dynamic: animation (realistic or virtual), dynamic scenes, videos etc. medical images, such as ct images, can also be dynamic. for example, "on line" heart rate graphs can be presented dynamically on screen; ultrasound or commentary boucheix 160 | f l r scanner images show the opening and closing of heart valves, and long series of images can be scrolled down and synchronized by the user. in their reviews, gegenfurtner, kok, van geel, de bruin & sorger, and fox & faulkner-jones observe that further research is needed on dynamic stimuli. gegenfurtner et al. observed that it is surprising that the literature on the neural correlates of real-world visual perceptual expertise has not yet systematically compared how the brain processes of experts and novices differ when they view static vs. dynamic stimuli or twodimensional vs. three-dimensional images. fox & faulkner-jones also report that several studies involved the use of visual stimuli that require manipulation by the viewer, such as three-dimensional images and dynamic video recordings. while the literature on dynamic image processing seems to be limited, the situation is very different in the field of learning and education, where a vast amount of research has been conducted in the last 20 years (see for example, lowe, schnotz, 2008; bétrancourt, 2005; mayer, 2005, 2014; van gog & schieter, 2010; höffler & leutner, 2007; boucheix & lowe, 2010, 2013; lowe & boucheix, 2008, 2011, 2016; jarodzka, scheiter, gerjets & van gog, 2010, amongst many others). bétrancourt and tversky (2000) defined computer animation with its notions of dynamic stimuli and transience as follows: "computer animation refers to any application which generates a series of frames, so that each frame appears as an alteration of the previous one, and where the sequence of frames is determined either by the designer or the user" (p. 313). thus, when presented with transient information, viewers (experts, novices or learners) must simultaneously store, process and link both previous and current information. thus, the distinction between static and dynamic images should not be seen as an adjunct to conventional static visual scenes, but as different in nature, with temporal and transient information added to the spatial dimension. with static visual stimuli, the perceptual and cognitive processing unit is the object (or part of it), with its shape and spatial location, but with dynamic stimuli, the perceptual and cognitive processing unit is the event with its associated behaviour (lowe & boucheix, 2008, 2016; kurby & zacks, 2008). lowe & boucheix (2008) proposed an animation processing model of explanatory dynamic visualizations in learners (the apm) based on studies that highlighted temporal processing (see also boucheix & lowe, 2016). fig 1. figure 1. summary of the main phases of the animation processing model (lowe & boucheix, 2008. commentary boucheix 161 | f l r according to the five-phase hierarchical animation processing model, learning from animation is a cumulative process for building dynamic mental models in which events play a crucial role. overall, this learning process can be divided into two broad types of activity: decomposition (apm phase 1) and composition (apm phases 2-5). a distinction is thus made between (i) analytic processing in which the learner must initially decompose the animation’s continuous flux of information into the discrete event units (entities plus their associated behaviours) that provide the raw material for mental model building, and (ii) synthetic processing in which this raw material is cumulatively and iteratively composed into the higher order knowledge structures that comprise a mental model of the target subject matter. the present commentary on the nine papers is not the place to present a theoretical model. however, it illustrates potential changes, especially with learners in the educational area (mostly novices) when using dynamic presentations (see also the work by lowe on expert meteorologists, 1999). as observed by fox & faulkner-jones on page 6: "we now have the tools available to examine the effects of expertise upon the dynamic diagnostic process. the authors have examined pathology trainees at the early and late stages of training, and found support for the use of specific dynamic digital platforms in the acquisition of “expert” patterns of gaze (fox et al., 2016)". in previous research with single static images, the main focus was on targets such as components or parts of objects. studies on dynamic stimuli emphasized the regional level of processing with the idea of macroor micro-dynamic chunks (lowe & boucheix, 2008; boucheix, lowe, putri & groff, 2013). it would be interesting to introduce a regional level in the flash-preview moving window paradigm presented by litchfield & donovan, in the framework of the global-focal search model (nodine & kundel, 1987). the relevance of the regional level clearly depends on the type and content of the images. more generally, it would be interesting to test the expertise effect on dynamic stimuli processing using different methodologies. what would be the effect of dynamic stimuli on sensitivity, specificity and accuracy in the roc analysis method (krupinski), on eye-tracking measures (fox & faulkner-jones), on eye tracking with verbal reports (helle), and on the neuro-correlates of visual expertise? what about the brain regions and neural circuits involved in processing dynamic stimuli? finally, dynamic visualization may have interactive features such as user control; for example, the sequence of frames can be determined either by the designer or by the user. for example, an obstetriciangynaecologist carrying out an ultrasound of the foetus of a pregnant woman, or a radiologist scrolling through a series of frames of a brain scan, can control the sequence of images in order to analyse them. what is the effect of expertise on the cognitive control strategies used to visualize the series of frames? the question is also raised of the interactions between top-down and bottom-up processing during the temporal sequence. previous research in the educational and learning area (school children, students and also professionals) found that interactive features hindered rather than fostered the performance of learners ("what matters is what you see, not whether you interact", keehner, hegarty, cohen, khooshabeh & montelleo, 2008; bétrancourt, 2005; boucheix, 2008). this interactive aspect of user control is also raised in the case of 3d objects, as highlighted in fox & faulkner-jones’ review of eye-tracking methods and in gegenfurtner et al.’s review of neuro-imaging techniques. previous research on 3d-object processing has also revealed difficulties for learners. however, short training sessions seem to lead to significant performance improvement (cohen & hegarty, 2007, 2011). 3.4. groups of interest: learners, learning and individual differences 3.4.1. learners and learning taking the expert performance approach, based on ericsson & kintsch’s (1995) long-term working memory theory, williams, fawver, & hodges underline the need to address the issue of facilitating the acquisition of expert performance: "in contrast to the growing research on expert performance across numerous domains, as well as significant research focusing on practice profiles of experts and novices, there remains a need for research on how expert learners continue to learn new skills and refine existing skills." commentary boucheix 162 | f l r it is also important to examine how visual expertise develops. gegenfurtner et al. identified three research approaches: contrastive (expert vs. novices), developmental (training effect on neural activations), and conditional. however, as shown in table 1 (last column on the right), in many of the methodological reviews, it seems that most studies focus on experts and/or novices and not on learners at different learning stages. longitudinal studies on the development of visual expertise could thus offer relevant insight into expertise development: how and when the cognitive processes involved in the visual search of images change, how eye movement patterns and scan paths change, how neural activations change, what develops and appears or disappears in learners' verbal reports at different professional stages. as observed by williams et al., one of the problems when assessing learning effects is the need for control groups (matched with experimental groups). joint research with researchers in the educational field should thus be encouraged. as shown by fox & faulkner-jones with regard to eye tracking, the top-down and bottom-up dynamics at different stages of learning could be very relevant in the study of expertise. for example, what is the learning process in the changes mentioned by litchfield & donovan: "visual search in medical image perception begins as: search & detect–recognize–decide, but with experience, this develops into: recognize & detect–search–decide”. the holistic visual search model, well suited to experts in the medical area, may not be entirely appropriate for learning complex animated graphics, where novices’ initial “glimpses” concern the most perceptually salient information, which may also be the least relevant (see lowe, 1999, with meteorologists, or boucheix & lowe, 2006, 2010). table 1 also indicates that, in most of the methodological reviews, only a few studies have investigated or focused on individual differences in experts. 3.4.2. individual differences in their review, williams et al. argue that there is a need for "more process-oriented studies about changes during learning that explain or address differences or changes in outcomes". in their view, deliberate practice is not the only factor explaining expertise development. there are vast differences in the amount of practice carried out by experts who have the same performance level. methodologies that record performance (objective measures) combined with verbal reports (more subjective measures) could help understand the course and development of the acquired expertise. similarly, gegenfurtner et al. wonder whether it is possible to "reliably measure/objectify neural correlates of visual expertise with currently available functional neuroimaging methods and therewith explain inter-individual behavioral differences" with respect to visual perceptual expertise. 3.4.3. spatial abilities and visual expertise (in the medical area) in none of the reviews (to the best of my understanding) are visual-spatial abilities considered to be a potential moderator of expertise, whatever the methodology used. however, a plethora of previous studies in this field showed that spatial abilities measured with standardized tests involving mental rotation of objects, spatial perspective taking, paper-folding tasks etc. had a significant effect on graphic and visual scene processing (for a review, see hegarty, 2010; hegarty, de leeuw, & bonura, 2008; höffler & leutner, 2010; boucheix & schneider, 2009; boucheix & chevrey, 2014). investigating the relationship between spatial abilities and the development of expertise could provide fruitful information to further our understanding of professional development and training. furthermore, all the standardized ability tests use static pictures of objects (e.g., metzler figures). with the development of dynamic visual presentations, it could be relevant to use dynamic instead of static spatial ability tests (porte, boucheix & bétrancourt, 2016). commentary boucheix 163 | f l r 4. conclusion: limitations of methodologies for studying visual expertise and the prospect of combined and synchronized methods to conclude this commentary we will briefly remark on two interrelated aspects: the limitations of the methodologies and the interest of using and combining several methodologies in the same research. 4.1. limitations of methodologies for studying visual expertise in each of the papers, the limitations of each methodology are indicated, as illustrated below. gegenfurtner, kok, van geel, de bruin & sorger discuss the spatial limitations of eeg, and the temporal limitations of fmri. limitations regarding the distance from the scalp to the deep structures in the brain have been described for fnirs. until recently, designing more ecological visual situations has been a challenge in neuro-imaging. by contrast, williams et al. describe how more ecological situations have been created in the expert-performance approach, such as film-based simulation to evaluate anticipation and decision-making in tennis (see figure 2, in the paper by williams et al., this issue). fox & faulkner-jones and litchfield & donovan underline the well-known limitation of eye-tracking techniques, namely that only focal vision is recorded, and not peripheral vision. however, the role of peripheral vision can be analysed using the flash-preview moving window paradigm. szulewski, kelton & howes describe how pupillometry requires strict lighting conditions. helle, van de wiel, and also to some extent ivarsson, discuss the limitations of interviews and verbal protocols to assess tacit, implicit, and more or less "unconscious" aspects of expertise. finally, at a more general level, gegenfurtner et al. highlight several possible caveats about the neuroscientific methods in addition to the limitations of spatial and temporal resolution, including ecological validity, reductive experimental bias and limited implications in the field of education. to some extent these limitations apply to many methodologies. 4.2. the prospect of combined and synchronized methods in a recent review, gegenfurtner, kok, van geel, de bruin, jarodzka, szulewski & van merriënboer, (2017) underlined "the need to allow novel and unconventional combinations of methods to aggregate data material that transcends single paradigmatic boundaries". as shown by helle’s analysis of verbal protocol and eye tracking, the use of complementary and synchronized methodologies or techniques could be useful when investigating process-oriented visual expertise. however, the decision to use two (or more) methodologies should not be guided only by the mere idea of collecting more data, but by much more specific goal of research. for example, and as suggested in gegenfurtner et al.’s review in this issue, the combination of pupillometry and eeg could provide interesting data and insight into the correlation between eye movements and neural activity. similarly, recent research on cognitive load in the domain of neuro-ergonomics (e.g. aircraft pilots) combined eye tracking (with pupillometry) and fnirs to understand the potential correlation between eye movements and pre-frontal (dorsolateral) activation during difficult piloting tasks (see durantin, gagnon, tremblay, & dehais, 2014; dehais, causse & cegarra, 2017). using a technique called eye fixation-related potentials (efrps), thierry baccino (see rämä & baccino, 2010) combined eye movement and eeg measures and showed that this technique could be a useful tool to study the temporal dynamics of visual perception and the processes underlying object identification. two other methodologies could complement the ones analysed in this special issue on visual expertise. the first is motion capture, which uses 3d motion sensors to record gestures and body movements, possibly in conjunction with other measures, such as eye movements recorded with mobile eye-tracking glasses. this could be useful to study new aspects of visual tasks involving gestures and hand-eye coordination (see figure 2 in williams et al.’s paper, showing a film-based simulation used to evaluate anticipation and decision-making in tennis). similarly, in arts, the processes underlying drawing or painting tasks could be analysed using a https://www.ncbi.nlm.nih.gov/pubmed/?term=r%c3%a4m%c3%a4%20p%5bauthor%5d&cauthor=true&cauthor_uid=20939937 https://www.ncbi.nlm.nih.gov/pubmed/?term=baccino%20t%5bauthor%5d&cauthor=true&cauthor_uid=20939937 commentary boucheix 164 | f l r combination of eye tracking and motion capture (perdreau & cavanagh et al., 2013). the same coupling of methods could be used in the medical domain, for example in surgery or radiology. the second methodology concerns the recording of physiological measures such as heart rate and skin conductance. these could be particularly useful when studying the emotional aspects of decision-making and/or "intensive" cognition related to cognitive load. similar methods have already been used in the field of car driving, for example in simulation tasks or to study the visual expertise of car drivers (pépin, jallais, fort, navarro & gabaude, 2017). two more general concluding remarks can be made. firstly, all the papers in this special issue contribute to a better understanding of the dynamics between top-down and bottom-up processes in visual expertise and of how perceptual expertise interacts with cognitive expertise. this has important practical implications. for example, in the field of education research, signalling and cueing effects in multimedia presentations have been extensively studied using eye tracking (see for example, jarodozka, de köning, boucheix). one of the main results of these studies is that learners fixate the signalled or cued information in the right place and at the right time, indicating that the signalling effect works (see also, mayer, 2014). however, this is not a guarantee that the information has been understood at a deep level. the second point concerns another practical issue of the research on the methodologies used to study visual expertise. in gegenfurtner et al.’s review of neuro-scientific methods, they observe that "it is a false belief that findings from eeg and fmri would be directly applicable and informative for re-designing learning environments and curricula". indeed, given the level of "reduction" in neuroscience experiments, we need to be very cautious about their practical implications. however, the continuous improvement of cognitive and neuro-scientific technologies in the future will give increasing opportunities to gain in-depth understanding of the processes underlying visual expertise. the nine articles in this special issue contribute to this progress. references berney, s., & bétrancourt, m. (2016). does animation enhance learning? a meta-analysis. computers & education, 101, 150-167. bétrancourt, m. (2005). the animation and interactivity principles in multimedia learning. in r. mayer (ed.). the cambridge handbook of multimedia learning (pp. 287-296). cambridge, uk: cambridge university press. boucheix, j.-m. (2008). young learners’ control of technical animations. in r. k. lowe, & w. schnotz (eds.). learning with animation: research implications for design (pp. 208-234). cambridge, uk: cambridge university press. boucheix, j.-m., & lowe, r. k. (2010). an eye tracking comparison of external pointing cues and internal continuous cues in learning with complex animations. learning and instruction, 20, 123-135. boucheix, j.-m., lowe, r. k., putri, d. k., & groff, j. (2013). cueing animations: dynamic signalling aids information extraction and comprehension. learning and instruction, 25, 71-84. http://dx.doi.org/10.1016/j.learninstruc.2012.11.005. boucheix, j.-m., & schneider, e. (2009). static and animated presentations in learning dynamic mechanical systems. learning and instruction, 19 (2), 112–127. boucheix, j.m. & chevret, m. (2014). alternative strategies in processing 3d objects diagrams: static, animated and interactive presentation of a mental rotation test in a cued retrospective study. in tim dwyer, helen purchase aidan delaney, diagrams 2014, lnai, lecture notes in artificial intelligence (eds.). diagrammatic representation and inference (pp. 138-145). heidelberg: springer-verlag. castelhano, m. s., & henderson, j. m. (2007). initial scene representations facilitate eye movement guidance in visual search. journal of experimental psychology: human perception and performance, 33, 753763. cohen, c.a. & hegarty, m (2007). individual differences in use of external visualizations to perform an internal visualization task. applied cognitive psychology, 21, 701-711 dehais, causse, m., & cegarra, j. (2017). « neuroergonomics: measuring the human operator’s brain in ecological settings ». le travail humain 2017/1 (vol. 80), p. 1-6. doi 10.3917/th.801.0001 http://leadserv.u-bourgogne.fr/fr/membres/jean-michel-boucheix commentary boucheix 165 | f l r donovan, t., & litchfield, d. (2013). looking for cancer: expertise related differences in searching and decision making. applied cognitive psychology, 27, 43-49. durantin, d., gagnon, j.f., tremblay, s. & dehais, f. (2014). using near infrared spectroscopy and heart rate variability to detect mental overload. behavioural brain research, de koning, b., tabbers, h. k., rikers, r. m. j. p., & paas, f. (2010). improved effectiveness of cueing by self-explanations when learning from a complex animation. applied cognitive psychology, 25, 183-194. de koning, b. b., tabbers, h. k., rikers, r. m. j. p., & paas, f. (2007). attention cueing as a means to enhance learning from an animation. applied cognitive psychology, 21, 731-746. ericsson, k. a., & kintsch, w. (1995). long-term working memory. psychological review, 102, 211245. ericsson, k. a., & smith, j. (1991). prospects and limits of the empirical study of expertise: an introduction. in k. a. ericsson & j. smith (eds.), towards a general theory of expertise: prospects and limits (pp. 1-38). new york: cambridge university press. fox, s.e., faulkner-jones, b.e. (2107). eye-tracking in the study of visual expertise: methodology and approaches in medicine. frontline learning research. this issue. gegenfurtner, a., kok, e., van geel,k., de bruin, a., & sorger, b. (2017). neural correlates of visual perceptual expertise: evidence from cognitive neuroscience using functional neuroimaging. frontline learning research. this issue. gegenfurtner, a., kok, e., van geel, k., de bruin, a., jarodszka; h., szulewski , a., & van merriënboer, j. j.g. (2017). the challenges of studying visual expertise in medical image diagnosis. medical education, 51, 97-104. gegenfurtner, a., & van merriënboer, j. j.g. (2017). methodologies for studying visual expertise. frontline learning research. this issue. gruber, h., jansen, p., marienhagen, j., & altenmüller, e. (2010). adaptations during the acquisition of expertise. talent development & excellence, 2, 3-15. hegarty, m. (2010). components of spatial intelligence. in b. h. ross (ed.). psychology of learning and motivations, volume 52 (pp. 265-297). amsterdam: elsevier. hegarty, m., de leeuw, k., & bonura, b. (2008).what do spatial ability tests really measure? in: proceedings of the 49th meeting of the psychonomic society chicago, il, november. hegarty, m., waller, d. (2004). a dissociation between mental rotation and perspective-taking spatial abilities. intelligence, 32, 175-191 helle, l. (2017). prospects and pitfalls in combining eye-tracking data and verbal reports. frontline learning research. this issue. henderson, j. m., pollatsek, a., & rayner, k. (1989). covert visual attention and extrafoveal information use during object identification. perception & psychophysics, 45, 196-208. höffler, t. n., & leutner, d. (2007). instructional animation versus static pictures: a meta-analysis. learning and instruction, 17, 722-738. http://dx.doi.org/10.1016/j.learninstruc.2007.09.013. höffler, t. n. (2010). spatial ability: its influence on learning with visualizations-a meta-analytic review. educational psychology review, 22, 245–269. ivarsson, j. (2017). visual expertise as embodied practice. frontline learning research. this issue. jarodzka, h., scheiter, k., gerjets, p., & van gog, t. (2010). in the eye of the beholder: how experts and novices interpret dynamic stimuli. learning and instruction, 2, 146-154. keehner, m., hegarty, m., cohen, c., khooshabeh, p., montello, d. (2008). spatial reasoning with external visualizations: what matters is what you see, not whether you interact. cognitive science, 32, 10991132. krupinski, e.a. (2017). receiver operating characteristic (roc) analysis. frontline learning research. this issue. kundel, h. l., nodine, c. f., & toto, l. c. (1991). searching for lung nodules. the guidance of visual scanning. investigative radiology, 26, 777-781. kurby, a., & zacks, j. (2007). segmentation in the perception and memory of events. trends in cognitive science, 12, 72-79. http://dx.doi.org/10.1016/ j.tics.2007.11.004 https://scholar.google.fr/scholar?oi=bibs&cluster=13130480308526016709&btni=1&hl=en https://scholar.google.fr/scholar?oi=bibs&cluster=13130480308526016709&btni=1&hl=en commentary boucheix 166 | f l r litchfield, d., & donovan ,t. (2017). the flash-preview moving window paradigm: unpacking visual expertise one glimpse at a time. frontline learning research. this issue. lowe, r.k., & boucheix, j.m. (2016). principled animation design improves comprehension of complex dynamics. learning and instruction, 45, 72-84. http://dx.doi.org/10.1016/j.learninstruc.2016.06.005 lowe, r. k. (1999). extracting information from an animation during complex visual learning. european journal of psychology of education, 14, 225-244. lowe, r. k. (2003). animation and learning: selective processing of information in dynamic graphics. learning and instruction, 13, 157-176. lowe, r. k., & schnotz, w. (2008). learning with animation: research implications for design. new york: cambridge university press.nodine, c. f., & kundel, h. l. (1987). the cognitive side of visual search in radiology. in j. k., o’regan, a. levy-schoen (eds). eye movements: from physiology to cognition (pp. 573–582). amsterdam: elsevier. lowe, r. k., & boucheix, j.-m. (2008). learning from animated diagrams: how are mental models built? in g. stapleton, j. howse, & j. lee (eds.). theory and applications of diagrams (pp. 266-281). berlin: springer. lowe, r. k., & boucheix, j.-m. (2011). cueing complex animation: does direction of attention foster learning processes? learning and instruction, 21, 650-663. mayer, r. e. (2009; 2014 2d edition). multimedia learning. new-york: cambridge university press. nodine, c. f., & kundel, h. l. (1987). the cognitive side of visual search in radiology. in j. k., o’regan, a. levy-schoen (eds). eye movements: from physiology to cognition (pp. 573–582). amsterdam: elsevier. pépin, g., jallais, c., fort, a., moreau, f., navarro, j. & gabaude, c. (20017). « towards real-time detection of cognitive effort in driving: contribution of a cardiac measurement », le travail humain, 2017/1 (vol. 80), p. 51-72. doi 10.3917/th.801.0051. perdreau, f. & cavanagh, p. (2013). the artist's advantage: better integration of objects information across eye movements. i-perception, 4, 380-395. porte, l., boucheix, j.m., bétrancourt, m., lowe, r.k., ceddia, m., & bard, p. (2016). dynamic spatial abilities and learning from animation. in désiron, j., berney, s., bétrancourt, m., & tabbers, h. (eds.). earli sig 2 conference: comprehension of text and graphics, learning from text and graphics in a world of diversity (pp. 145-148). geneva, july 11-13: university of geneva. rämä p., & baccino t., (2010). eye fixation-related potentials (efrps) during object identification. vis neurosci. 2010 nov;27(5-6):187-92. doi: 10.1017/s0952523810000283. epub 2010 oct 13. scheiter, k. (2014). the learner control principle in multimedia learning. in r. e. mayer (ed.). the cambridge hanbook of multimedia learning (pp. 487-512). new-york: cambridge university press. schnotz, w., & lowe, r. k. (2008). a unified view of learning from animated and static graphics. in r. k. lowe, & w. schnotz (eds.). learning with animation: research implications for design (pp. 304-356). new york: cambridge university press. schwan, s. (2013). the art of simplifying events, in a.p. shimamura (ed.). psychocinematics. exploring cognition at movies. (pp. 214-226). new-york: oxford university press. szulewski, a., kelton, d., howes, d. (2017). pupillometry as a tool to study expertise in medicine frontline learning research. this issue. van gog, t., & scheiter, k. (2010). eye tracking as a tool to study and enhance multimedia learning. learning and instruction, 20, 95-99. http://dx.doi.org/ 10.1016/j.learninstruc.2009.02.009. van de wiel, m. (2017). examining expertise using interviews and verbal protocols. frontline learning research. this issue. williams, a.m., fawver, b., hodges, n.j. (2017). using the ‘expert performance approach’ as a framework for examining and enhancing skill learning: improving understanding of expert learning. frontline learning research. this issue. http://leadserv.u-bourgogne.fr/fr/membres/jean-michel-boucheix http://leadserv.u-bourgogne.fr/fr/membres/jean-michel-boucheix https://www.ncbi.nlm.nih.gov/pubmed/?term=r%c3%a4m%c3%a4%20p%5bauthor%5d&cauthor=true&cauthor_uid=20939937 https://www.ncbi.nlm.nih.gov/pubmed/?term=baccino%20t%5bauthor%5d&cauthor=true&cauthor_uid=20939937 https://www.ncbi.nlm.nih.gov/pubmed/20939937 https://www.ncbi.nlm.nih.gov/pubmed/20939937 microsoft word koerber et al_publication.docx             frontline  learning  research  vol.5  no.  1  (2017)  76  -­‐  84   issn  2295-­‐3159       corresponding author: susanne koerber, department of psychology, freiburg university of education, kunzenweg 21, 79117 freiburg, germany. phone: +49(0)761 / 682-278, email: susanne.koerber@ph-freiburg.de. doi: http://dx.doi.org/10.14786/flr.v5i1.265   diagrams support revision of prior belief in primary-school children susanne koerbera, christopher osterhausa, beate sodianb afreiburg university of education, department of psychology, germany bludwig-maximilians-university munich, department of psychology, germany article received 4 july / revised 21 november / accepted 29 november / available online 28 february abstract the reluctance of children to revise their prior beliefs is a prominent phenomenon in the reasoning literature. one way to facilitate belief change is offering explanations, and this study examined whether highlighting (counter)evidence with diagrams leads to belief revision to the same extent. altogether 134 preschoolers and second-graders (5and 7year-olds, respectively) were presented with either counterintuitive data or explanations, both refuting a strong commonly held belief concerning the relation between two variables (e.g. eating carrots improves vision). in the explanation condition, we presented children with an explanatory underlying mechanism for the unexpected causal relation (e.g. spinach and carrots contain the same amount of vitamin a, with both improving vision). in the diagram condition, children were presented with empirical data displayed in a bar graph (non-covariation), which also disconfirmed the initial belief. in both age groups and both conditions we found significant numbers of belief revision with high certainty ratings concerning the new belief. belief change was more pronounced in second-graders, who in addition showed significantly more changes in the diagram condition than in the explanation condition. these findings suggest that the perceptual saliency of (counter)evidence helps children to correctly evaluate hypotheses, which supports changes in their prior belief. keywords: belief revision, hypothesis evaluation, primary-school children, diagram, explanation   koerber  et  al       | f l r     77   1. introduction the prominent role of prior knowledge in the evaluation of evidence is a well-documented phenomenon in the reasoning literature (e.g., croker & buchanan, 2011; kuhn, amsel, & o’laughlin, 1988). especially when evidence is inconsistent with their favored hypotheses or beliefs, children often refuse to update their initial beliefs; rather, they distort their interpretation of the data so that these are in accordance with their initially held belief (klaczynski, 2001; kuhn et al., 1988). the distortion typically entails that children ignore data that are inconsistent with their initial belief, that they selectively attend to only those parts of the data that are consistent with their initial belief, or that they even misinterpret the data (amsel & brock, 1996; chinn & brewer, 1993; chinn & malhotra, 2002; kuhn, et al., 1988; masnick, klahr, & knowles, 2016). a common explanation for this phenomenon has been proposed by kuhn (e.g., kuhn, 2010). kuhn suggests that young children do not understand that theory (beliefs) and evidence (data) are two epistemologically distinct categories. specifically, she argues that young children do not understand that theories need to be fully backed up by evidence, which would result in their selective interpretation of data. in a well-known study, kuhn and her colleagues (1988) asked children to interpret data that suggested a causal relation between a set of variables (e.g., different kinds of food, being healthy or not). specifically, sixthand ninth graders, as well as adults, were presented with fictitious data that, so were told, was obtained in a boarding school, in which different groups of students consumed four different kinds of food or beverage (diet coke or regular coke, baked potato or fries, oranges or apples, and special k or granola). some of the students in the boarding school felt sick after lunch, others were healthy. two variables contradicted the prior belief of the participants and two variables were in accordance with the participants’ beliefs concerning the impact of the different food items on the health of the students when children were asked whether a specific kind of food caused the students’ sickness, kuhn and colleagues found that only the older children and the adults used evidence-based reasoning (i.e., they related their answer to the covariation data). in contrast, most of the younger children did not attend to the evidence at all or they did not interpret it correctly and instead they kept their initial belief. although this finding seems to suggest some severe deficits in data interpretation skills and an insufficient differentiation between hypothesis and evidence in young children, more recent research shows that children as young as six-year olds possess a basic understanding of the hypothesis-evidence relation (e.g. ruffman, perner, olson & doherty, 1993). for instance, koerber, sodian, thoermer, and nett (2005) found that children as young as five-year-olds are able to successfully interpret simple patterns of covariation data, without any distortion of prior beliefs, when data are not overly complex and when the prior belief is not too strong. in their study, koerber and colleagues presented children with a hypothesis held by a story character and a set of covariation data (perfect and imperfect covariations, noncovariation) that contradicted the protagonist’s hypothesis. most five-year-olds successfully attributed a belief revision to the protagonist when the relation presented in the data was straightforward (perfect covariation), showing that they are able to successfully interpret simple patterns of data without any distortions and to incorporate this new evidence into their theories (see also piekny & maehler, 2013; van der graaf, segers, & verhoeven, 2016, for a replication of these findings). these confirmatory findings of early data interpretation skills are in line with a growing literature on early preschool and primary school scientific thinking, which shows that already young children possess a basic understanding of the distinction between hypothesis and evidence (mayer, koerber, sodian & schwippert, 2014; sandoval, sodian, koerber, & wong, 2014). the discovery of increasingly mature scientific thinking skills in this young age group suggests that children, in principle, should be able to use data and evidence to update and revise their initially held beliefs. but then why do studies find such weak performance in belief revision tasks in young children? we argue that one of the main reasons why evidence evaluation often fails to promote belief revision in young children is that many studies use evidence that is either too complex and/or is not salient enough, especially so for young children. asking children to interpret data about the effects of multiple variables in a single design, as for instance was done in kuhn et al. (1988), places heavy demands on children’s general   koerber  et  al       | f l r     78   information-processing capacities and it demands a sequential processing of information. in addition the information typically only enters via a single perceptual route (i.e., data are presented only verbally without graphical depiction). chinn and malhotra (2002) found that indeed children find it difficult to make correct, undistorted observations of (counter)evidence when the data are not salient. data with little salience increased children’s reluctance to change their prior belief in the face of new evidence, showing that the salience of the evidence is an important facilitator of successful belief revision. one possible way to increase salience of data lies in the presentation of evidence in simple bar graphs. bar graphs make possible a salient and meaningful representation of covariation data and even preschoolers can successfully interpret these graphs (koerber & sodian, 2008). their positive effects are, such as those of many well-designed visual displays, mostly attributed to the following three characteristics (e.g. hegarty, 2011; kosslyn, 1989, 1994, 2006; larkin & simon, 1987): first, bar graphs serve as external storage for information and thus they reduce memory capacity. second, the relations between two variables are spatially organized in bar graphs and they can be perceived at a glance (i.e., different pieces of information and relations between variables can be perceived simultaneously and thus bar graphs allow for a more efficient processing of information). and third, complex cognitive processes can be “offloaded” on perception in bar graphs. taken together, bar graphs thus assist information processing by offering an additional perceptual route (in addition to the verbal route) and, in addition, their two-dimensional display of a relation between two variables allows using analogies to space and spatial relations in order to make inferences on non-spatial content domains. graphs thus offer the viewer information (e.g., a linear trend) in a direct way which does not require that this information is inferred or computed from numerical or verbal data. recent research has shown that the positive effects of bar graphs even hold for children as young as six-year-olds. moreover, prior research showed that kindergarteners can successfully read off causal relations from simple bar graphs (koerber & sodian, 2008). the present study investigates whether successful belief revision can be induced in young children when evidence is presented in a salient way (i.e., in a bar graph) that requires reduced information processing. specifically, we compare this way of inducing belief revision to a highly effective means of revising initially held beliefs, which is providing children with explanations and causal mechanisms that link the two variables. this approach has been taken by koslowski (1996, 2012, see also koslowski, marasia, chelenza, & dublin, 2008), who argues that children often do not give up their initially held beliefs in favor of contrary evidence because of their strong subjective causal theories. these causal theories comprise not only information about the statistical association (covariation) between two variables, but they also entail beliefs about the underlying causal mechanism that connects the two variables. according to this account, covariation evidence alone cannot sufficiently challenge the initially held beliefs when no adequate, novel explanations about the underlying causal mechanism is offered simultaneously. in a study involving sixth and ninthgraders as well as college students, koslowski (1996) presented participants with the results of a fictitious study that investigated whether two kinds of food (sweets with low or high concentration of sugar, and milk with low or high concentration of fat) influenced whether or not children could easily fall asleep. all participants initially believed that sugar but not fat had a significant influence on children’s ability to find sleep. participants were assigned to two conditions: in one condition, participants received only covariation evidence that contradicted their initially held belief (“covariation-only-condition”); in the other condition, participants received not only covariation evidence, but also they were given an explanation concerning the mechanism that may link fat to sleep (“covariation-and-mechanism condition”). as hypothesized, receiving an additional explanation led to significant more belief revision than did the covariation-evidence alone. the explanation condition that was included in the present study therefore presented, analogously to koslowski (1996), children with an explanation that accounted for the occurrence of the counterevidence and a mechanism that linked the two variables. the diagram condition in turn, presented empirical data about covariation in a salient way, depicted in bar graphs. participants were preschool and primary-school children, who are at an age when conceptual change occurs over a wide range of knowledge areas. we hypothesized that children in this young age group would change their prior belief when presented with evidence in a   koerber  et  al       | f l r     79   salient way. in addition, we hypothesized that belief revision would, in the diagram condition, not be inferior to belief revision in the explanation condition. 2. methods a 2 (condition: explanation vs. diagram) by 2 (age group: preschoolers vs. second-graders) betweensubjects design was employed to investigate children’s belief revision in four tasks. 2.1. participants participants were 134 children, among them 54 preschoolers (m = 5,6 years, sd = 5 months, nexplanation = 28, ndiagram = 26) and 80 second-graders (m = 7,6 years, sd = 5 months, nexplanation = 40, ndiagram = 40). the 64 girls and 70 boys were sampled from 10 middle-class preschools and schools in a proximity to a large city in southern germany. preschools and schools supplied parents with information material about the study, and parents decided whether or not their children would be allowed to participate in the study. in addition to parental informed consent, child assent was obtained for all participants. participants from both age groups were semi-randomly assigned to one of the two conditions to ensure equal group sizes. 2.2. material and procedure four tasks were used to test children’s belief revision. the contexts used in these tasks were contexts in which children typically hold strong, naïve beliefs about the covariation and causal relation between two variables. these were: (1) eating carrots (but not spinach) improves vision; (2) drinking milk (but not mineral water) increases bone density; (3) eating gummi bears (but not mustard) makes you fat; and (4) eating chocolate (but not bananas) makes you happy. for each context, the children were first asked about their initial belief concerning the relation between the variables (e.g. “does eating carrots or eating spinach improve vision, or do they both equally contribute to good vision?”). children’s answers were visually displayed by placing small plastic cards (e.g., eyes and carrots) in front of them. this was done so that children would remember their initial belief. also, children indicated how confident they were about their initial belief (0, 1, 2). in addition, children were presented with a story protagonist (robbie or anna, depending on the gender of the child which was the same for the protagonist) who held the same initial belief as the child. this was done in order to account for potential differences between children’s own belief revision and the revision they ascribe to another person. depending on children’s initial belief, the following counterevidence was used to induce belief revision in children: in the carrot/spinach context (task 1), children who believed that carrots or spinach improved vision were presented with counterevidence that showed that the effect of the two variables is equally strong (noncovariation). children who initially believed that both factors are equally associated with good vision were presented with counterevidence that showed that carrots improve vision (see table 1 for an overview of which counterevidence was presented in response to varying initial beliefs). because the prior literature revealed that the type of covariation evidence (e.g., perfect covariation or noncovariation) influences preschoolers’ interpretation of data (e.g., koerber et al., 2005, piekny & mähler, 2013), we included counterevidence in the form of noncovariation as well as in the form of perfect covariation in order to account for potential effects.   koerber  et  al       | f l r     80   table 1 design of the four tasks context child said counterevidence (factor a) counterevidence (factor b) counterevidence (a/b same effect) carrot/spinach (task 1) carrot (a) x spinach (b) x doesn’t matter (a/b) x milk/water (task 2) milk (a) x water (b) x doesn’t matter (a/b) x gummi bears/mustard (task 3) gummi bears (a) x mustard (b) x doesn’t matter (a/b) x chocolate/bananas (task 4) chocolate (a) x bananas (b) x doesn’t matter (a/b) x in the diagram condition, counterevidence was presented visually in bar graphs (for an example, see figure 1); in the explanation condition, experimenters presented the children with an explanation about a mechanism that supported the opposite of children’s initial belief. in the carrots/spinach context, for instance, children who initially believed that only carrots would improve vision heard an explanation that maintained that spinach and carrots contain equal amounts of vitamin a, which is the mechanism that improves vision. after the evidence was presented, experimenters reminded children of their initial belief. subsequently, they asked the children about their present belief (including the strength of their confidence in this belief) as well as about the present belief of the story protagonist (third person). children were interviewed individually by two research assistants, who were extensively trained and who each interviewed equal amounts of children in each condition and age group (i.e., there was no interaction between experimenters and condition). to ensure that children in the diagram condition were able to interpret the bar graphs, a short introduction was given in which the experimenters explained how to read a bar graph. figure 1. example of a bar graph (perfect covariation).   koerber  et  al       | f l r     81   3. results preanalyses revealed that children’s confidence in their initial belief [low (=0), moderate (=1) or high (=2)] was high before treatment (m = 1.59; sd =.44 and m = 1.63, sd = .38, respectively, for the preschoolers and second-graders) with no significant difference between the age groups, t(132) =-.547, p>.05. this clearly indicates that all children held strong initial beliefs. figure 2 shows the mean percent of belief revision across all four tasks. in line with our hypothesis, more than 60% of all children changed their initial belief in light of the salient evidence or the explanations. in the explanation condition, between 63% and 71% of all children changed their initial belief (solid lines) and they ascribed a belief revision to the story protagonist (dashed lines) in the light of the counterevidence. in the diagram condition, belief revision was more pronounced in both measures for the second-graders (84% and 86%) than for the preschoolers (59% and 67%). the difference between the diagram condition and the explanation condition was significant in second grade, where children significantly more often changed their own initial belief, t(78)=2.68, p <.01, and significantly more often ascribed a belief revision to the story protagonist, t(78) = 3.44, p=.001, in the diagram condition than in the explanation condition. in the diagram condition, there was, in addition, a significant difference between the two age groups, with five-year-olds changing their initial belief significantly less often than eight-year-olds, t(64) =-4.67, p<.001, and also they ascribed significantly less often a belief revision to the story protagonist than did eight-year-olds, t(64) =– 2.63, p=.01. while a 2 (age) by 2 (condition) analysis of variance of children’s mean number of belief revisions ascribed to the story protagonist revealed no significant main effects for age (f(1, 129) = 1.282, ns) or condition (f(1, 129) = 2.590, ns) a significant interaction, f(1, 129) = 5.631, p < .05, η2 = .042 was revealed. figure 2. mean number of belief revisions categorized by age and condition 4. discussion can successful belief revision be induced in young children when evidence is presented in a salient way? the present study found that salient evidence (diagrams) is as effective to challenge and revise young children’s initial beliefs as explanations. in contrast to the findings of a large number of prior studies (e.g., amsel & brock, 1996; chinn & brewer, 1993), our findings suggest that even preschoolers are able to revise an initial belief when they are presented with counterevidence, be it an explanation about a mechanism or covariation evidence in a salient way. second-graders showed even more belief revision when presented 0%   10%   20%   30%   40%   50%   60%   70%   80%   90%   100%   explanafon     diagram     preschoolers  (self)   preschoolers   (other)   second-­‐graders   (self)   second  graders   (other)       koerber  et  al       | f l r     82   with salient counterevidence (bar graph) than when receiving an explanation, suggesting that this form of evidence presentation may be especially beneficial in this age group. in accordance with our hypothesis, bar graphs helped children to successfully interpret the data and to revise their initial beliefs. bar graphs appear to be such a helpful tool in belief revision because they depict the relation between two or more variables in a straightforward manner; in addition, they enable students to grasp complex data patterns that depict the relations between multiple variables at a single glance, and they help to reduce memory and information processing demands. this way, bar graphs make the counterevidence more salient, and by means of this saliency they lead to substantially higher rates of belief revision than have been found in prior studies (e.g., amsel & brock, 1996). interestingly, our second graders performed even better in the diagram condition than in the explanation condition. this is an interesting finding because research from koslowski (1996) has shown highly beneficial effects of explanations in belief revision. a possible explanation for this finding (the high amount of belief revision in the diagram condition) is that children may have come up with self-generated explanations (legare, 2012). these self-generated explanations may have supported and strengthened the effect of the diagrams. legare showed that twoto six-year-olds produce substantially more often selfgenerated explanations when they are faced with counterevidence than when they interpret evidence that is consistent with their initial beliefs. importantly, our participants did not seem to distort the data in the diagram condition. this is most likely due to the salience of the counterevidence that we presented in the bar graphs, which makes it difficult to ignore or neglect aspects of the data that do not align with children’s initial beliefs. visually perceiving counterevidence thus might have triggered children to generate their own, new explanations for the unexpected relation rather than distorting the data. and it is reasonable to assume that these self-generated explanations are even more beneficial for inducing belief revision than those generated by another person. given this interpretation, it might be that the better performance of the second graders in the diagram condition cannot only be attributed to their perception of the visual information but also to their deeper processing of the information. future studies therefore need to disentangle the impact of the visual salience of the counterevidence and the depth of processing of the information (e.g., the generation of selfexplanations). to this end, “thinking aloud” protocols in which participants are explicitly asked to come up with explanations may be especially helpful. if the salient presentation of evidence and diagrams indeed provoke self-generated explanations, then it is impossible to strictly isolate the effects of the presentation of the diagram and the explanations. thus, in line with koslowski (1996, 2012), we did not regard evidence and explanation as a dichotomy, but rather we suggest that future studies should investigate the way in which the salient presentation of evidences leads to the generation of theory-based explanations. adding a third, mixed condition, in which the participants receive counterevidence and explanations, should then be equally facilitating for belief revision as the diagrams alone. children’s own belief revision was contrasted in this study with their ascription of belief revision to another person. although we found convergent findings in these two measures, the difference between the diagram and the explanation condition was more pronounced in second grade when asking children about another person’s belief revision. the belief revision of a third person has been the dependent measure in many other studies (e.g., koerber et al, 2005; piekny & mähler, 2013), which is typically done because it ensures that the (counter)evidence is evaluated from a more abstract and distant level. the general trends in our data were however similar for our two measures (own belief revision and belief revision ascribed to a third person) so that it seems reasonable to assume that these two measures are closely related, leading to the same conclusions regarding the influence of saliency on evidence interpretation. in sum, our findings suggest that the salient presentation of counterevidence leads to belief revision in children as young as five-year-olds. although explanations are important for children’s belief revision, explanations do not need to be provided externally (e.g., by an adult); instead, our data show that presenting evidence may suffice when it is presented in a way that is salient and that may elicit children’s selfgeneration of explanations. if this interpretation of our findings holds and children indeed generate   koerber  et  al       | f l r     83   explanations while processing salient evidence, this finding will be of high relevance not only for theorists, but also for teachers and practitioners in schools. first, it underlines the important role diagrams can play for illustrating concepts. second, it also reveals their usefulness for eliciting and promoting conceptual change (see also koerber, 2003). therefore, diagrams should play a more important role in school in general and specifically in science and math curricula. keypoints preschoolers are able to revise a prior belief when presented with counterevidence perceptual saliency of evidence (given in a diagram) helps children to evaluate hypotheses and revise beliefs the role of evidence and explanations for supporting hypothesis evaluation should not be viewed as dichotomous. acknowledgements the study was conducted as part of the project “development of understanding graphs and diagrams in preschool age and elementary-school age” and it was funded by the german research council (dfg) (ko 2276/3-1).we would like to thank daniela huber and susanne mikschl for their assistance in data collection and all children, parents, and teachers who supported this study. references amsel, e., & brock, s. (1996). the development of evidence evaluation skills. cognitive development, 11, 523-550. http://dx.doi.org/10.1016/s0885-2014(96)90016-7 chinn, c. a., & brewer, w. f. (1993). the role of anomalous data in knowledge acquisition: a theoretical framework and implications for science instruction. review of educational research, 63, 1-49. chinn, c. a., & malhotra, b. a. (2002). children's responses to anomalous scientific data: how is conceptual change impeded? journal of educational psychology, 94, 327-343. http://dx.doi.org/10.1037/0022-0663.94.2.327 croker, s., & buchanan, h. (2011). scientific reasoning in a real-­‐‑world context: the effect of prior belief and outcome on children's hypothesis-­‐‑testing strategies. british journal of developmental psychology, 29(3), 409-424. doi: 10.1348/026151010x496906 hegarty, m. (2011). the cognitive science of visual-spatial displays: implications for design. topics in cognitive science, 3(3), 446–474. doi: 10.1111/j.1756-8765.2011.01150.x klaczynski, p. a. (2001). analytic and heuristic processing influences on adolescent reasoning and decision-­‐‑ making. child development, 72, 844-861. doi: 10.1111/1467-8624.00319 koerber, s. (2003). der einfluss externer repräsentationsformen auf proportionales denken im grundschulalter. [the influence of external forms of representation on proportional reasoning in elementary school age.] hamburg: verlag dr. kovac. koerber, s., sodian, b., thoermer, c., & nett, u. (2005). scientific reasoning in young children. preschoolers’ ability to evaluate covariation evidence. swiss journal of psychology, 64, 141-152. http://dx.doi.org/10.1024/1421-0185.64.3.141 koerber, s., & sodian, b. (2008). preschool children’s ability to visually represent relations. developmental science, 11, 390-395. doi: 10.1111/j.1467-7687.2008.00683.x   koerber  et  al       | f l r     84   kosslyn, s. m. (1989). understanding charts & graphs. applied cognitive psychology, 3, 185–226. doi: 10.1002/acp.2350030302 kosslyn, s. m. (1994). elements of graph design. new york: w. h. freeman. kosslyn, s. m. (2006). graph design for the eye and mind. new york: oxford university press koslowski, b. (1996). theory and evidence: the development of scientific reasoning. mit press. koslowski, b. (2012). scientific reasoning: explanation, confirmation bias and scientific practice. in g. feist & m. gorman (eds.), handbook of the psychology of science and technology (pp 151-192). dordrecht: springer. koslowski, b., marasia, j., chelenza, m., & dublin, r. (2008). information becomes evidence when an explanation can incorporate it into a causal framework. cognitive development, 23(4), 472-487. http://dx.doi.org/10.1016/j.cogdev.2008.09.007 kuhn, d. (2010). what is scientific thinking and how does it develop? in u. goswami (ed.). handbook of childhood cognitive development (2nd ed., pp. 472-534). oxford, england. blackwell. kuhn, d., amsel, e., & o'loughlin, m. (1988). the development of scientific reasoning skills. orlando, ca: academic. larkin, j., & simon, h. (1987). why a diagram is (sometimes) worth ten thousand words. cognitive science, 11, 65–99. doi: 10.1111/j.1551-6708.1987.tb00863.x legare, c. h. (2012). exploring explanation: explaining inconsistent evidence informs exploratory, hypothesis-­‐‑testing behavior in young children. child development, 83, 173-185. doi: 10.1111/j.14678624.2011.01691.x masnick, a. m., klahr, d., & knowles, e. r. (2016). data-driven belief revision in children and adults. journal of cognition and development, (just-accepted). http://dx.doi.org/10.1080/15248372.2016.1168824 mayer, d., sodian, b., koerber, s., & schwippert, k. (2014). scientific reasoning in elementary school children: assessment and relations with cognitive abilities. learning and instruction, 29, 43-55. http://dx.doi.org/10.1016/j.learninstruc.2013.07.005 piekny, j., & maehler, c. (2013). scientific reasoning in early and middle childhood: the development of domain-­‐‑general evidence evaluation, experimentation, and hypothesis generation skills. british journal of developmental psychology, 31(2), 153-179. doi: 10.1111/j.2044-835x.2012.02082.x ruffman, t., perner, j., olson, d. r., & doherty, m. (1993). reflecting on scientific thinking: children's understanding of the hypothesis-­‐‑evidence relation. child development, 64, 1617-1636. doi: 10.1111/j.1467-8624.1993.tb04203.x sandoval, w. a., sodian, b., koerber, s., & wong, j. (2014). developing children's early competencies to engage with science. educational psychologist, 49, 139-152. http://dx.doi.org/10.1080/00461520.2014.917589 van der graaf, j., segers, e., & verhoeven, l. (2016). scientific reasoning in kindergarten: cognitive factors in experimentation and evidence evaluation. learning and individual differences, 49, 190-200. http://dx.doi.org/10.1016/j.lindif.2016.06.006 microsoft word mckeown & hmelo-silver_publication.docx             frontline  learning  research  vol.4  no.  4  special  issue  (2016)  56  -­‐  58   issn  2295-­‐3159       doi: http://dx.doi.org/10.14786/flr.v4i4.318 commentary: expanding conceptualizations for the study of learning: implications for theory and research jessica m. mckeown & cindy e. hmelo-silver indiana university, usa the articles in this special issue provide examples of how conceptions of studying learning need to expand to incorporate new environments and concerns in education. as technology advances the tools available to learners and instructors in synchronous and asynchronous learning environments, theory must meet the new challenges of conceptualizing these spaces and pedagogical approaches that allow for new ways of organizing learning opportunities and supporting learners. the depiction of chronotope, the ecological perspective, and social practice theory convey the need to consider new conceptualizations. considerations of agency and emotion expand the bounds of learning in ways that are important for the educational futures that we can imagine. these exciting advances have created a need for thinking about learning in a way that reframes existing conceptualizations of how to appropriately analyze learning activities. this collection of articles takes a stridently dynamic view of learning that goes beyond static cognition and recognizes the importance of learner agency. some perspectives are more or less emergent in terms of their conceptions of learning and context. although including context is not new for sociocultural theories, being specific about aspects of context is. in particular, most of these articles are concerned with the role of time and space and the relevant interactions among them. new definitions of learning must consider it as a system that looks across what might appear to be natural boundaries of time, space, and context; cognition and emotion; human and artifact: with individual agents, co-constructing groups and institutions (e.g., engle & conant, 2002; lemke, 2000). as in other systems, what emerges (i.e., what knowledge is socially constructed) is different from what might be predicted from those bounded-appearing elements. the chronotope framework presented by ritella, ligorio, and hakkarainen is a conceptual/analytical tool centered on the premise of interdependency between issues of space and time. education research tends to treat these factors as backdrops where learning occurs instead of focusing on how they interact and present opportunities and constraints that may not occur otherwise. here, learning is a system of complex interactions with agency at different levels, that may not always be predictable, and are defined by the timespace where the learner is situated. these lines of thinking are similar to engle, lam meyer, & nix’s (2015) ideas about how positioning learners’ agency and expansively framing their ideas is a key for learning and transfer. conceptually, ritella et al. argue their framework could be used to design instruction while considering how space-time is used, especially in technology-rich learning environments. for example, when we consider how mobile computer-supported collaborative learning environments are used by different collaborators across different spaces and times. we echo ritella et als. heed to consider how teachers can orchestrate learning in such complex settings. commentary  mckeown  &  hmelo-­‐silver       57 | f l r     similar to this study, penuel also considers the move across settings and learner agency. examining pathways of interest in science, technology, engineering, and mathematics fields, penuel and colleagues adopt social practice theory, which allows for the conceptualization of structures of practice and the emergent constraints and possibilities that exist for the learner(s). longitudinal analysis of an individual path can show us where our assumptions about education reform and programs are falling short and where they are working. this article examines how individual agency is expressed, but also attends to emergent properties of the environment/learner interaction. further considering agency, damsa and jornet discuss the study of higher education through the lens of the ecological perspective, which holds that learning is a set of dialectical relationships between the individual, others, and the environment. learning is portrayed as a cognitive, emotional, and practical system, where the self, society, and tools necessarily impact the transformation of the others. however, each component of the system cannot be understood or analyzed without consideration of the others, as is typical of systems. issues of co-construction and agency, knowledge resources, and trans-contextuality are highlighted in a case study of a class of computer engineering undergraduates as they interact while completing a project for a customer. the authors argue that the field of higher education study needs to consider the transformative nature of learning. two articles in this issue illustrate the challenges of integrating two lesser-studied aspects of cognition, imagination and emotion, into science and history respectively. moving toward a perspective focusing on learner agency, the work detailed next highlights the dynamic and socially negotiated processes that support learner agency. the role of imagination in science education is a relatively new field of study, and as such has been lacking a framework for analysis. building on the work of zittoun & gillespie (2016) and their depiction of the distal and proximal loops of imagination, hilppö and colleagues introduce us to the expansive and refining dynamics of imagination. they argue that these patterns of imagination allow learners to break the constraints of space and time that are often imposed by limitations of classroom settings. the expansive imagination allows shifts in meaning-making between proximal and distal times and spaces, and for pushing the boundaries of what is realistic, and possible, while refining imagination encompasses again the boundaries of reality to imagine what could happen plausibly. benefits of tapping into imagination in science education include agency, and collective imagination, as imagination is situated and shared. goldberg and schwarz take a socio-emotional approach to learning, specifically how students come to reason about historical sources. they compare three approaches to history learning on hot topics (critical inquiry, empathetic dual narrative, and patriotic apologetic teaching). these hot topics are places where emotion, identity, and agency come to the forefront, rather than creating obstacles to disciplinary discussions. the authors provided examples from four episodes, representative of different potential effects of teaching approaches on relations of emotions and history. different teaching approaches lead to different ways that learners invoke identity and emotion—some of these approaches can support perspective taking whereas others can help make learners aware of biases and/ or impede productive engagement. considering multiple perspectives and biases is important for helping learners develop emotional responses that can be productive. learning environment designs that bring out identity-related emotions can support more reflective engagement with historical thinking. it is important to consider the ways in which merging cognition and emotion can support and challenge disciplinary thinking practices. while these focus on history as a discipline, these considerations are also applicable to considering other content, such as socioscientific issues or professional education that might explicitly focus on emotionally laden content (hmelosilver, jung, lajoie, lu, yu, wismeman, & chan, 2016). in order to broaden our consideration of learning as emotionally dialogic, we need to consider the broader contextual and disciplinary spaces in which it is useful to consider this expansive perspective. these articles were provocative in their refined conceptualizations of learning and the aspects that each article brought to the foreground. the frameworks here focused on the conceptual and analytic tools that each of the authors argued will inform education research in the 21st century. to move this research commentary  mckeown  &  hmelo-­‐silver       58 | f l r     forward, and taking the notion of design-based research seriously, the next step is application of these frameworks in designing for learning. this is particularly timely given the rapid changes in technology that break down many boundaries between time and space, but also require learner agency. new designs for learning must take into account a notion of learning as a systemic process with expansive boundaries; one that is dynamically unfolding, with some emergent aspects that can provide new opportunities. but with these opportunities come challenges in how to orchestrate and facilitate learning in complex environments. the designs are themselves tests of these new conceptualizations (e.g., mckenney & reeves, 2013). further research into these topics needs to consider the role of the teachers, designers, and learners in dealing with the complexity of new learning spaces, as individuals influence and are also affected by components of the emergent systems in which they exist. references engle, r. a., & conant, f. r. (2002). guiding principles for fostering productive disciplinary engagement: explaining an emergent argument in a community of learners classroom. cognition and instruction, 20, 399-484. engle, r. a., lam, d. p., meyer, x. s., & nix, s. e. (2012). how does expansive framing promote transfer? several proposed explanations and a research agenda for investigating them. educational psychologist, 47(3), 215-231. doi: 10.1080/00461520.2012.695678 gillespie, a., & zittoun, t. (2016). the gift of a rock: a case study in the emergence and dissolution of meaning. nothingness: philosophical insights to psychology. new brunswick, nj: transaction publishers. hmelo-silver, c. e., jung, j., lajoie, s. p., lu, j., yu, y., wiseman, j., chan, l-k. (2016). video as context and conduit for problem-based learning. in s. m. bridges, s. m., l-k chan, & c. e. hmelo-silver. (eds.), educational technologies in medical and health sciences education (pp.57-77). new york: springer. lemke, j. l. (2000). across the scales of time: artifacts, activities, and meanings in ecosocial systems. mind, culture, and activity, 7(4), 273-290. mckenney, s., & reeves, t. c. (2013). conducting educational design research. routledge. sandoval, w. a. (2014). conjecture mapping: an approach to systematic educational design research. journal of the learning sciences, 23, 18-36. doi: 10.1080/10508406.2013.778204 microsoft word dignath_publication.docx           frontline  learning  research  vol.4  no.  5  (2016)  83  -­‐  105     issn  2295-­‐3159       contact information: charlotte dignath-van ewijk, goethe university frankfurt, theodor-w.-adorno-platz 6, 60323 frankfurt am main, email: dignath@psych.uni-frankfurt.de doi: http://dx.doi.org/10.14786/flr.v4i5.247 which components of teacher competence determine whether teachers enhance self-regulated learning? predicting teachers’ self-reported promotion of self-regulated learning by means of teacher beliefs, knowledge, and self-efficacy charlotte dignath-van ewijk goethe university frankfurt, germany article received 10 march / revised 16 august / accepted 9 october / available online 18 january abstract in this study, the predictive value of three aspects of teacher beliefs regarding teachers’ promotion of self-regulated learning (srl) is modelled by means of structural equation modelling. these include teacher beliefs on (1) instructing srl, (2) regarding their own self-efficacy towards the promotion of srl, and (3) their epistemological beliefs regarding learning. 173 primary school teachers participated in the study. path analysis revealed that teachers’ beliefs on instructing srl, along with their self-efficacy beliefs regarding the promotion of srl, were predicting teachers’ promotion of srl mostly positively. the results offer new insights into teacher beliefs and how they account for self-reported teacher practice regarding the promotion of srl. this study is particularly innovative as it is the first study in the field of teachers and srl to investigate teacher beliefs and teacher selfefficacy as potential determinants of teachers’ promotion of srl in the classroom. these results can serve to construct a model of teachers’ promotion of srl, as well as provide ideas on how to help teachers supporting srl. this study is frontline as it appears that no other research has been published on teachers’ beliefs, in particular self-efficacy beliefs, towards promoting srl and how these are related to teachers’ promotion of srl. most research on teachers and srl has so far focused on training teachers, but no model of the professional competence of teachers in this area exists until now. keywords: self-regulated learning; teacher beliefs; teacher self-efficacy dignath  et  al       | f l r     84   1. theoretical background a large amount of research on the impact of self-regulated learning (srl), as well as on factors that determine the use of srl and how srl can be fostered in learners, has been examined in the past. however, there is still a lack of research on what determines teachers’ promotion of srl. when looking at the literature, one can imagine the amount of research in the area of srl as being similar to the shape of a funnel: the numerous studies on the impact that srl has on a learner’s achievement and motivation have drawn the interest of many researchers to investigate, as a next step, how srl in learners can be improved. several meta-analyses on how to promote srl have summarized the considerable number of training studies in the field of srl (see e.g., hattie, biggs & purdie, 1996; dignath, büttner & langfeldt, 2008). when looking further at how teachers can enhance their students’ srl, the amount of research decreases substantially (for a literature review on the role of the teacher in promoting srl, see moos & ringdal, 2012). and finally, only few studies have explicitly searched for reasons why some teachers do, while others do not promote srl in their classrooms, in particular with regard to their instruction of srl strategies (e.g., chatzistamatiou, dermitzaki & bagiatis, 2014; dignath-van ewijk, dickhäuser & büttner, 2012). many studies in the past showed that individual factors, such as teaching self-efficacy, value beliefs, etc., are related and/or affect teachers' reports regarding several aspects of instructional quality, as well as covering aspects of supporting students’ autonomy while learning (e.g., kunter, tsai, klusmann, brunner, krauss & baumert, 2008). however, knowing more about what determines whether teachers promote srl would be helpful to generate ideas on what teachers should develop in order to enhance srl in their students and how teachers can be supported during the process of teacher training (see kramarski, desoete, bannert, narciss & perry, 2013). to fill this gap, we investigated several potential predictors in order to learn what determines teachers’ self-reported promotion of srl. the aim of the study was to investigate the impact of potential determinants for teachers’ promotion of srl, including teacher beliefs, teacher self-efficacy, and teacher knowledge, on teachers’ self-reported promotion of srl in the classroom in order to find out which teacher characteristics should be addressed when training teachers in srl. teachers had to fill out questionnaires regarding their educational beliefs, self-efficacy beliefs, and their knowledge on how to support students’ srl, and they also had to rate their promotion of srl while they taught. the teacher characteristics which were assessed should be placed in a general model of teaching competence. 1.1 promotion of self-regulated learning zimmerman (2000) defines self-regulated learners as learners who set themselves goals, plan their actions to pursue these goals, monitor their learning, and finally evaluate their learning process. in terms of a feedback loop, the result of this evaluation influences and regulates the following learning process. when looking at how teachers can foster srl in students, paris and paris (2001) describe two different ways: directly, by providing students with knowledge and skills on how to self-regulate (teaching them strategies in terms of informed training (brown, campione & day, 1981), and indirectly, by arranging the learning environment in a constructivist way so that students can and have to self-regulate their learning (e.g., by offering choices to students and providing them with situations in which they can take over responsibility for their learning) (pressley, harris & marks 1992). although some literature can be found on how to promote srl among students using specific interventions (see e.g., dignath, 2009; paris & paris, 2001; perry & vandekamp, 2000; perry, phillips & dowler, 2004; pressley et al., 1992), only little research has been published so far about teachers’ actual promotion of srl in the classroom. the few observational studies that have been conducted in order to register in how far teachers foster srl among their students have concluded that teachers spend only little time on strategy instruction, even if they often design learning environments which require self-regulation (see e.g., bolhuis & voeten, 2001; dignath et al., 2013; hamman, berthelot, saia & crowley, 2000; spruce & bol, 2014; moely et al., 1992). the outcomes of those studies raise the question, why teachers do not dignath  et  al       | f l r     85   invest more in preparing their students for self-regulation. do teachers not support the idea of self-regulated and strategic learning (teacher beliefs), or do they not know how to support it (teacher knowledge)? and how are both, beliefs and knowledge, related? 1.2 teacher beliefs and teacher knowledge when looking at teacher beliefs, a clear labelling of the constructs of beliefs and knowledge seems necessary. while beliefs are supposed to be more affective in this distinction, knowledge is supposed to have the higher epistemic status by being more justifiable when compared with beliefs (fenstermacher, 1984). although the terms have been used interchangeably in teacher education literature (hofer & pintrich, 1997), we draw on the terminology used by pajares (1992) when reviewing the literature of teacher beliefs and teacher knowledge: “belief systems, unlike knowledge systems, do not require general or group consensus regarding the validity and appropriateness of their beliefs. individual beliefs do not even require internal consistency within the belief system. this nonconsensuality implies that belief systems are by their very nature disputable, more inflexible, and less dynamic than knowledge systems.” (pajares, 1992, p. 311). beliefs are supposed to include value commitments, epistemological beliefs, subjective theories about learning, and goals. furthermore, motivational orientations cover teachers’ self-referred cognitions – in particular locus of control and self-efficacy – as well as their intrinsic motivation (baumert & kunter, 2006, 2013). 1.3 epistemological beliefs beliefs on fostering srl could be influenced by more general beliefs on the nature of learning and knowing: epistemological beliefs. epistemological beliefs refer to the assumptions an individual has about the origin, nature and structure of knowledge and knowing (schraw, crippen & hartley, 2006). hofer and pintrich (1997) identified three lines of research on epistemological beliefs. the first line of research has dealt with developmental models of beliefs about knowing and knowledge, building on the initial work of perry (1970). the second line of research has focused on the consequences of epistemological assumptions on the way people think and reason. the last and most recent line of research has been concerned with the structure of a system of beliefs and how these beliefs influence comprehension and academic achievement (e.g., schommer, 1990; schommeraikins, 2004). schommer has developed a taxonomy of beliefs consisting of four more or less independent dimensions: innate ability (learning is not changeable, but rather a fixed ability), quick learning (content is either learned quickly or not at all), simple knowledge (most important information is rather simple), and certain knowledge (information does not change over time). these four dimensions were extracted by factor analysis with each factor seen as a continuum. the factor innate ability is probably of central importance for the promotion of srl in our study and will therefore be described later on in more detail compared with the other three factors of which detailed descriptions can be found elsewhere (e.g., hofer & pintrich, 1997). several researchers have found significant relationships between epistemological beliefs and srl (see e.g., pieschl, stahl & bromme, 2008 or schommer, 1990 for epistemological beliefs in general; see e.g., bendixen & hartley, 2003 for innate ability). in the way that this relationship can be found between learners’ epistemological beliefs and their self-regulation, one can assume that such a relationship exists between teachers’ epistemological beliefs and their promotion of self-regulation. if the assumption that learning and knowing is something that a learner cannot change or influence, then the learner does not feel capable to self-regulate his or her learning. if a teacher assumes that learning and knowing is not changeable, why would he or she try to teach self-regulation of learning? in a meta-synthesis, hattie (2013) synthesized the results of several meta-analyses on the effect that teacher beliefs have on their effectiveness. he found a strong effect of teachers’ epistemological beliefs related to their conceptions of teaching and learning (hattie, 2013). several studies on the association of teachers’ epistemological beliefs and their instruction dignath  et  al       | f l r     86   point into this direction as teachers’ epistemological beliefs affect their curricular and pedagogical decisions (for an overview, see schraw et al., 2006). however, results are sometimes unclear and other studies found the impact of epistemological beliefs on teaching competence to be mixed (e.g., creemers et al., 2013; shraw & olafson, 2003; sosu & gray, 2012). bell (2006) examined the effects of srl and epistemological beliefs on academic achievement while holding constant the effects of self-efficacy and prior college academic achievement. he found students’ prior academic achievement and their expectancy to be the only significant predictors for academic achievement. however, he argues that students’ self-regulation, as well as their epistemological beliefs, probably influence students’ expectancy, which in turn influences academic achievement, but he did not show any own predictive value in the multiple regression analyses that he had run (bell, 2006). we therefore decided to analyze the impact of epistemological beliefs in another way so that we can account for indirect effects on our outcome variable, as epistemological beliefs might not directly influence teachers’ selfreported promotion of srl, but only indirectly via more specific teacher beliefs. 1.4 beliefs on the promotion of self-regulated learning to our knowledge, only few studies thus far have been conducted in order to investigate the relationship between teachers’ educational beliefs and teachers’ promotion of srl. lombaerts and colleagues have addressed the question of the relationship between flemish primary school teachers’ beliefs and their self-reported teaching practice by developing questionnaires to investigate teacher beliefs on promoting srl (lombaerts, engels, van braak & athanasou, 2009), as well as on teachers’ realization of self-regulation during their teaching (lombaerts, engels & athanasou, 2007). they found teacher beliefs about srl in primary school to be a significant predictor for teachers’ self-reported recognition of srl. these teacher beliefs were predicted significantly by beliefs about teacher-level influence on srl, but not by beliefs about school context influence on srl. thereby, teacher beliefs about teacher-level influence on srl cover aspects like e.g., beliefs on instructional pedagogy, and on innovations in teaching and learning (sample item: personal insight into how to support self-regulated learning does influence the introduction of selfregulated learning in my classroom.), while beliefs about school-level influence on srl include beliefs on the stimulation of srl by the school as an organization, collaboration of teachers as part of the school culture, or curriculum changes (sample item: the level of involvement of our team in the school development plan influences the introduction of self-regulated learning in my classroom.) (see lombaerts et al., 2009). moreover, another significant predictor for teachers’ self-reported recognition was teacher-level satisfaction, which was in turn predicted by school context satisfaction, itself not having any direct predictive value for self-reported teacher behavior. the main conclusion of lombaerts and his colleagues (2009) is that beliefs about teacher-level influence on srl predict teachers’ self-reported behavior directly, but beliefs about school context influence on srl do not. vandevelde, vandenbussche and van keer (2012) conducted a study with flemish primary school teachers that revealed that teachers who report developmental educational beliefs reported to implement srl more often than teachers holding transmissive educational beliefs. teachers’ implementation of srl was assessed using the scale that lombaerts and colleagues had developed and used in their study (lombaerts et al., 2007). dignath-van ewijk & van der werf (2012) examined the predictive value of dutch primary school teachers’ educational beliefs and their knowledge about promoting srl for their realization of supporting srl in the classroom. their results showed that teacher beliefs on srl predict teachers’ self-reported implementation of srl, contrary to teachers’ general educational beliefs, or to teachers’ knowledge on promoting srl. moreover, they found that teachers were more positive towards the realization of a constructivist learning environment in general than towards the instruction of srl strategies. spruce and bol (2014) examined the relationship between (1) teacher beliefs and (2) knowledge and (3) the observed classroom practice of ten american primary and middle school teachers in a qualitative case dignath  et  al       | f l r     87   study. they found teachers often deviated within these three constructs: out of the ten teachers that they observed in classrooms, teachers with high knowledge regarding the promotion of srl and rather positive beliefs reached only low observation scores when promoting srl; teachers having a positive attitude towards srl and possessing high knowledge regarding the promotion of srl was not reflected in their promotion of srl, which the researchers observed (spruce & bol, 2014). these studies have investigated the predictive value of teacher beliefs and, regarding teacher knowledge, the promotion of srl on behalf of teachers. however, none of the studies have considered the impact of teachers’ self-efficacy on implementing srl in the classroom. 1.5 self-efficacy beliefs when looking at motivational orientations, self-efficacy and locus of control play an important role in determining teacher behavior (see e.g., baumert & kunter, 2006, 2013; kunter, 2013). the feeling of control over the behavior in question arises from self-efficacy theory (bandura, 1977). according to bandura (1977), self-efficacy beliefs represent a judgment of one’s capabilities to reach a certain goal. with regard to teacher efficacy, tschannen-moran and woolfolk (2001) define teachers’ efficacy beliefs as an assessment of “his or her capabilities to bring about desired outcomes of student engagement and learning, even among those students who may be difficult or unmotivated” (tschannen-moran & woolfolk, 2001, p. 783). bandura (1986) proposed that people with high self-efficacy beliefs are more likely to perform better and show a higher frustration tolerance than people with low self-efficacy beliefs. self-efficacy beliefs would, therefore, play an important role in one’s motivation to initiate an effort, the persistence of effort, and the way in which one deals with setbacks (skaalvik & skaalvik, 2010). four factors are supposed to determine a person's self-efficacy: (1) personal experience of success or failure (also: enactive attainment) whose interpretation is closely related to people’s beliefs and values, (2) vicarious experience (also: modeling) which is influencing the knowledge of what is “right” and “wrong”, (3) social persuasion in terms of encouragement or discouragement, as well as (4) the perception of physiological reactions (bandura, 1977). these four factors also affect teachers’ beliefs and knowledge which will, in turn, act as determinants for teachers’ self-efficacy. these sources of efficacy beliefs are considered to provide the basis for one’s task analysis and appraisal of one’s personal competence. the resulting judgment of the match of task requirements and personal ability creates the teachers’ efficacy belief (tschannen-moran, hoy & woolfolkhoy, 1998). how is the self-efficacy of teachers linked to teacher beliefs, knowledge, and behavior? a large amount of research on teacher efficacy has evolved within the last decades (see klassen, tze, betts & gordon, 2011 for a review). research has shown the relationship between self-efficacy and teacher behavior (see e.g., guo, piasta, justice & kaderavek, 2010; holzberger, phillip & kunter, 2013; tschannen-moran & woolfolk hoy, 2001). studies have found e.g., a significant relationship between teachers’ self-efficacy and teacher beliefs towards instructional innovation (e.g., ghaith & yaghi, 1997; guskey, 1988) and instructional strategies (e.g., tschannen-moran & johnson, 2011; wertheim & leyser, 2002; swars, 2005), and a relationship between teachers’ self-efficacy and a control orientation against teacher control of students (woolfolk & hoy, 1990). with regard to the promotion of srl, chatzistamatiou et al. (2014) found that teachers' self-reported strategies to enhance students' srl in mathematics were predicted by their own selfefficacy beliefs. perry, hutchinson, and thauberger (2008) asked teachers about their instruction of strategies and found teachers to be positive towards the idea of fostering self-regulation of their students, but they did not know how to do this. beside the pure knowledge of doing so, the teachers’ self-efficacy on feeling competent enough to handle this might play a role. dignath  et  al       | f l r     88   1.6 hypotheses in this study, we wanted to investigate determinants of teachers’ promotion of srl when looking at teacher beliefs and teacher knowledge. although also other variables on the institutional level or the teacher level can have an impact on teachers’ promotion of srl, we focused on variables on the teacher level in our investigation in order to start researching potential determinants on a micro-level first. the relationships between the above mentioned concepts, on which we base the theoretical model, can be derived from the theory as follows (see figure 1 for a graphical illustration on the predicted relationships): a) firstly, general teacher beliefs are assumed to predict more specific teacher beliefs (pajares, 1992). therefore, epistemological beliefs about learning as a fixed or changeable ability (in terms of rather general teacher beliefs) are assumed to predict teachers’ beliefs on the promotion of srl (in terms of rather specific beliefs). b) second, knowledge is assumed to be predicted by one’s beliefs. in how far teachers learn new knowledge and how new information will be integrated into their existing knowledge base, will be predicted by the beliefs that teachers already report (ertmer, 2005). thus, teacher beliefs are assumed to predict teacher knowledge on the promotion of srl. c) third, most researchers agree on the strong impact that teacher beliefs have on teaching behavior. most empirical evidence suggests that teacher beliefs have a stronger impact on teaching behavior than does teacher knowledge (see e.g., the reviews of kagan, 1992, and pajares, 1992). we therefore assume that teacher beliefs predict teacher knowledge, and that both teacher beliefs and teacher knowledge predict self-reported teacher practice, with teacher beliefs being a stronger predictor than teacher knowledge. d) whether an intention is carried out as a certain behavior depends not only on one’s beliefs, but also on one’s perceived behavioral control (ajzen, 1991; bandura, 1977). the self-efficacy of teachers to promote srl is thus supposed to predict teachers’ self-reported promotion of srl, next to teacher beliefs and teacher knowledge. e) finally, self-efficacy is supposed to be determined by several factors (see point described earlier in this section), of which the interpretation is influenced by one’s beliefs and knowledge (bandura, 1977). we therefore assume that teacher beliefs and teacher knowledge will predict the self-efficacy of teachers. 2. methods 2.1 sample primary school teachers from southern germany had been contacted via email and telephone to ask if they could complete a questionnaire on teachers’ knowledge and beliefs towards srl. one hundred and seventy-three primary school teachers participated in the study, which equates to a response rate of 41%. the data from 140 teachers was complete and could be included in the analyses. teachers were mainly female (87.1%), and ranged in age from 22 to 64 years (m= 39.6 years, sd=12.63) and had on average 15.06 years of teaching experience (sd=12.91). dignath  et  al       | f l r     89   2.2 procedure in order to find out whether teacher beliefs and knowledge predict teachers’ self-reported promotion of srl, a path analysis was conducted. as we assumed beliefs and self-efficacy to predict teachers’ selfreported practice, self-reported teacher behavior was regressed on beliefs regarding the promotion of srl, as well as epistemological beliefs, on knowledge, and on self-efficacy within the path analysis. furthermore, epistemological beliefs as more general beliefs on learning were assumed to predict teacher beliefs on the promotion of srl, as those are more specific. figure 1 shows the model that was specified based on the previously described theories. figure 1. expected model of predictors for teachers’ self-reported promotion of self-regulated learning the questionnaires were administered to all teachers within a time period of two weeks. most schools received the questionnaires personally; only schools which were situated more than 40 km away received the questionnaires by post. one week after administration, the completed questionnaires were collected from the schools. the questionnaire took approximately 30 minutes to complete. teachers were told about the purpose of the study. questions from the questionnaire were to be answered with regard to teaching in grade 4 of primary school1 to allow for comparisons among teachers. 2.3 instruments the teacher questionnaire consisted of four self-reporting scales and one open question regarding teacher knowledge. teachers first had to answer the open question that aimed at assessing their knowledge on promoting srl before possibly being influenced by the content of the items in the questionnaires. next, teachers had to answer questions using the scales measuring teacher beliefs on srl, teacher self-efficacy beliefs, as well as epistemological beliefs. finally, the questionnaire contained a scale assessing teachers’ self-reported promotion of srl. the order of the scales was chosen in a way that minimizes influences from answering one scale to the next.                                                                                                                           1 in the german school system, students enter primary school at the age of six (1st grade) and leave primary school after 4th grade. epistemological beliefs (“innate”) epistemological beliefs (“changeable”) beliefs on promotion of srl self-efficacy for the promotion of srl knowledge on the promotion of srl self-reported promotion of srl dignath  et  al       | f l r     90   2.3.1 teacher epistemological beliefs teachers’ epistemological beliefs regarding learning were measured using the subscale for fixed ability from a german translation of the epistemological questionnaire by schommer (1990) (schiefele, moschner & husstegge, 2002). innate, or fixed ability, draws on the concept of individuals’ implicit theories of intelligence by dweck and legget (1988) who found individuals to differ in their view on intelligence as being a fixed versus a malleable entity. although schommer (1990) treated the five items assessing innate, or fixed ability, as being one scale, we decided to include two subscales: one measuring the belief whether learning skills are innate (sample item: “there exists an innate talent, which determines how quickly one can learn.”), and one measuring the belief whether learning behavior is changeable (sample item: “the ability to learn can hardly be influenced through practice.”). in former research from schommer, the subset of items that measured ability to learn is innate, was assumed to be part of the factor fixed ability; however, it had not consistently loaded on this factor, but on the factor quick learning (see hofer & pintrich, 1997). schommer (1990) had developed the questionnaire to assess the epistemological beliefs of learners. in our study, we assessed the epistemological beliefs of teachers who completed the questionnaire in the first instance from the perspective of a teacher and not of a learner. for teachers, it makes sense by nature that learning is changeable; otherwise, their entire profession would come into question. however, there is variance in teachers’ beliefs regarding fixed ability that we wanted to be able to capture. epistemological beliefs were therefore assessed by means of two subscales: the belief whether learning skills are innate, and the belief whether learning behavior is changeable. the teachers rated each item on a sixpoint likert scale ranging from not true at all to completely true. internal consistency for the subscale innate was .57 and for the subscale changeable, .62. although .57 is a low coefficient by most reliability standards, results from other studies using the measurement of epistemological constructs have presented reliability coefficients in this range (see e.g., hofer, 2000; schraw et al., 2002). since the problem to be expected from relatively low internal consistency is low statistical power, only the subscale changeable, which consisted of three items and still had an acceptable reliability of α = .62, was kept in the analysis. 2.3.2 teacher beliefs on srl teacher beliefs on instructing srl were assessed with a german version of the self-regulated learning teacher belief scale (lombaerts et al., 2009). the twelve items from this scale focused on teachers’ beliefs about supporting primary school students’ self-regulation of learning using different measures (sample item: “the instruction of learning strategies leads to students being better in evaluating their learning.”) and had to be rated on a five-point likert scale. the internal consistency (cronbach’s α) for the german version of this scale in our sample was .80. 2.3.3 teacher self-efficacy beliefs to measure teachers’ beliefs regarding their own self-efficacy with regard to the promotion of srl, the teacher-self-efficacy scale by schmitz und schwarzer (2000), consisting of ten items, had been tailored to the context of srl. item formulations had been adapted in order to assess specific teacher self-efficacy beliefs concerning the promotion of srl (example: “i know that i manage to help even to the most difficult students to learn the learning content.” [original item] was changed to “i know that i manage to help even to the most difficult students to learn how to learn the learning content themselves.” [adapted item]). for this adapted scale, we found an internal consistency of α = .81 in this sample. 2.3.4 teacher knowledge teacher knowledge was assessed by means of an open-ended question developed by lonka, joram, and bryson (1996), asking teachers about the best way to enhance the learning behavior of students, teaching them learning to learn and why. dignath-van ewijk & van der werf (2012) have developed a coding scheme for this question in order to code the answers of teachers according to models of direct and indirect ways of fostering srl. the teachers’ written answers were transcribed and coded for data analyses according to this dignath  et  al       | f l r     91   coding scheme using two different coders. the coding was based on a theoretical foundation of promoting srl which states that teachers should both provide students with autonomy to regulate their own learning themselves (indirect way of promoting srl) and to teach them srl-strategies on how to deal with such an autonomy (direct way of promoting srl). in terms of scaffolding, the more students are used to selfregulation, the less the teacher has to structure himor herself (so he or she can shift from a more direct way of promoting srl, to a more indirect way). providing students with srl-strategies without enabling them to self-regulate their learning by means of a learning environment that offers the students opportunities to take over responsibility, is not supposed to foster students’ srl, as they learn in a theoretical way about selfregulation but do not have the chance to practice it. the same applies the other way around: creating a learning environment that provides students with autonomous learning opportunities, but that doesn’t teach them how to handle such autonomy, does not foster srl, as many students will not be able to take over the responsibility for their learning when they lack srl-strategies. to this effect, teachers should do both: provide students with learning environments that allow them self-regulation and provide them with strategies to handle these learning environments more effectively (see dignath, 2009; paris & paris, 2001). we therefore coded teacher responses to whether teachers did not show any knowledge of promoting srl (coded as 0: no answer or answer that does not fit to the question, e.g.: “caring for a nice atmosphere in which the pupils feel comfortable.”), only partial knowledge so either mentioning the composition of an autonomous learning environment (coded as 1: example item: “cooperative learning, as learning with peers leads to discussions and that leads to a better understanding instead of the teacher saying how it has to be and then it is like that.”, or the instruction of srl-strategies (coded as 1: “analysing together with the pupils how they learn and make them get to know also other strategies/ways of learning.”), or whether they showed full knowledge (mentioning both, characteristics of an autonomous learning environment, as well as strategy instruction (coded as 2: “on the one hand it is important for pupils to be allowed to work on their own ideas. on the other hand, for some pupils with a greater need for structure, it can also be good to first get familiar with certain learning processes and to first learn how to do that.”). teacher answers for the composition of a learning environment that supports self-regulation were coded according to the coding scheme of an observation instrument developed by dignath-van ewijk et al. (2013) in order to assess teachers’ promotion of srl in the classroom. this coding scheme included categories for teaching formats like cooperative learning, discovery learning, providing students with choices, situated learning, and problem-based learning. teacher answers concerning the instruction of srl-strategies were sorted according to the same coding scheme containing strategy categories like metacognitive strategies (planning, monitoring, evaluating, reflecting), cognitive strategies (organization, elaboration, problem solving), and motivation strategies (e.g., attribution, feedback seeking, cooperation, etc.). for a more precise description of the coding scheme, see dignath-van ewijk et al., 2013 and dignath-van ewijk & van der werf, 2012). teachers were allowed to give more than one answer to that question. this is even desirable as teachers would have to mention at least one direct and one indirect way of fostering srl in order to reach the highest rating. interrater reliability computed with cohen’s kappa revealed an agreement of 𝜅 = .83. 2.3.5 teachers’ self-reported promotion of srl a german version of the self-regulated learning inventory for teachers (lombaerts et al., 2007) was used to assess teachers’ self-reported promotion of srl in their classroom. the questionnaire of lombaerts et al. (2007) consists of 24 items, capturing srl during the three phases of zimmerman’s (2000) model of srl: a planning phase, an action phase, and a self-reflection phase (sample item: “my pupils work on tasks that require them to plan their work themselves towards a deadline.”). we included 16 items from this questionnaire into our questionnaire as some items did not match our interest in teachers’ promotion of srl, but we additionally included five items regarding the direct instruction of srl strategies (sample item: “i ask my students to work independently without explicitly discussing learning strategies beforehand.”) developed by dignath-van ewijk & van der werf (2012). all items had to be rated on a six-point likert scale ranging from never to always. the reliability estimate (cronbach’s α) for the overall scale in our research was .90. dignath  et  al       | f l r     92   table 1 overview of instruments used construct instrument number of items example item teacher beliefs on instructing srl german version of the self-regulated learning teacher belief scale (lombaerts, engels, van braak & athanasou, 2009) 12 the instruction of learning strategies can be realized in primary school. teachers’ epistemological beliefs regarding learning german translation of the epistemological questionnaire by schommer 1990 (schiefele, moschner & husstegge, 2002) 5 there exists an innate talent which determines how quickly one can learn. teachers’ beliefs regarding their own self-efficacy teacher-self-efficacy scale by schmitz und schwarzer (2000) 10 i know that i can teach learning strategies even in difficult situations. teachers’ self-reported promotion of srl german version of the self-regulated learning inventory for teachers (lombaerts, engels & athanasou, 2007) 21 i teach my students how to plan one’s learning. teachers’ knowledge on how to promote srl open-ended question developed by dignath-van ewijk & van der werf (2012) based on lonka et al. (1996) 1 according to you: what is the best way to teach students learning to learn? why? table 2 means, standard deviations and cronbach alpha reliabilities construct m sd cronbach’s α teacher beliefs on instructing srl 2.74 .42 .80 teachers’ epistemological beliefs regarding learning subscale “innate”: 3.54 subscale “changeable”: 1.25 .84 .72 .57 .62 teachers’ beliefs regarding their own selfefficacy 1.78 .41 .81 teachers’ knowledge on how to promote srl cohen’s kappa: 83% teachers’ self-reported promotion of srl 2.58 .59 .90 dignath  et  al       | f l r     93   2.4 path analysis we used path modelling as an extension of multiple regression in order to test both direct and indirect (mediator) effects on the dependent variable. for the model proposed from the data in this study, we included indicators of teacher beliefs (including epistemological beliefs, beliefs on srl, and self-efficacy beliefs) and teacher knowledge, as well as teachers’ self-reported promotion of srl as an outcome measure that is predicted by the other variables in the model. since we did not want to estimate any latent variables, including an explicit estimation of measurement errors, we chose path analysis instead of structural equation modelling using stata. path analysis is based on multivariate regression modelling which allows for more than one dependent variable and the analysis of mediator effects, using one regression equation per endogenous variable derived from the path diagram. the relative strength of the postulated effect is presented in terms of path coefficients, most commonly indicated by the beta coefficient. teachers’ self-reported promotion of srl, as the criterion measure, is considered to be under the influence of all other variables in the model, either directly or when mediated through other variables. 3. results 3.1 test for multivariate normality we conducted the doornik-hansen test for multivariate normality (doornik & hansen, 2008) for the five variables included in the model in order to show that the data was normally distributed (chi2=9.34, p=.50). 3.2 scale intercorrelations table 3 provides an overview of correlations among the scales. results were largely consistent with the hypothesized predictions that can be found in figure 1. scores measuring epistemological beliefs concerning the belief of ability as being changeable correlated negatively with teacher beliefs on srl (r=-.19), teacher self-efficacy beliefs (r=--.21), and self-reported promotion of srl (r=--.21). epistemological beliefs scores did not correlate with teacher knowledge, which was not consistent with our predictions. teacher beliefs on srl correlated positively with teacher knowledge (r=.37), teacher self-efficacy beliefs (r=.43), and teachers’ self-reported promotion of srl (r=.49). self-efficacy beliefs further correlated positively with teacher knowledge (r=.33) and teachers’ selfreported promotion of srl (r=.56). finally, teacher knowledge also correlated with teachers’ self-reported promotion of srl (r=.36). dignath  et  al       | f l r     94   table 3 overview of correlations among the scales teacher knowledge teacher beliefs on srl teacher selfefficacy teachers’ self-reported promotion of srl teachers’ epistemological beliefs “changeable” teacher knowledge 1.00 teacher beliefs on srl 0.37*** 1.00 teacher selfefficacy 0.33** 0.43*** 1.00 teachers’ selfreported promotion of srl 0.36*** 0.49*** 0.56*** 1.00 teachers’ epistemological beliefs “changeable” 0.09 -0.19* -0.21* -0.20* 1.00 note: *correlation is significant at the .05 level, one-tailed. **correlation is significant at the .01 level, onetailed. ***correlation is significant at the .001 level, one-tailed. intercorrelations among teacher knowledge, teacher beliefs, teacher self-efficacy, self-reported teacher behavior, and teachers’ epistemological beliefs (n=140) 3.3 path modelling regression analyses path regression analyses were then conducted according to the hypotheses specified earlier in order to predict teachers’ reported promotion of srl directly through teachers’ knowledge, beliefs and selfefficacy towards the promotion of srl, as well as indirectly through teacher beliefs. theoretically, we assumed that all variables would predict teachers’ self-reported promotion of srl; however, not all of them would have a direct effect. compared to the fully saturated model, we therefore omitted the direct paths between epistemological beliefs and self-efficacy of promoting srl, and between epistemological beliefs and self-reported srl. we calculated direct, indirect, and total effects for all endogenous variables by applying sewall wright’s multiplication rule as the path tracing rule: each structural equation is multiplied by its predetermined variables (wright, 1934). we used stata 13 to estimate the path model and tested the model fit using maximum likelihood estimation. the likelihood ratio, chi square, and corresponding p value, rmsea and cfi, were calculated to assess the model fit. the initial path model contained five predictors and had an r2=.38, indicating that the model explains 38% of the variance in teachers’ self-reported promotion of srl. chi square indicated that the model could be improved in order to fit the data accordingly compared to the fully saturated model: χ2 (2, n = 140) = 8.06, p = 0.02. in a next step, we therefore added the two omitted paths to the model. the path from epistemological beliefs to self-efficacy in promoting srl was significant (p = 0.01), while the path from epistemological beliefs to self-reported srl was not significant (p = 0.17). finally, we added the former path to the model. this model did not fit the data significantly worse than the fully saturated model: χ2 (1, n = 140) = 1.85, p = 0.17). compared to the baseline model without any predictors, the revised model was highly significant (likelihood ratio chi square = 144.76; p < .001). as additional measures of fit, the rmsea and cfi were calculated and both indicated an acceptable fit (rmsea = .076; p(rmsea<.05) = dignath  et  al       | f l r     95   -.11* .25; cfi = .994). as the evidence on the investigated constructs would allow for several theoretical assumptions, we tested five different models and compared their fit indices, as summarized in table 4: table 4 goodness-of-fit-indices for the compared models model 1 model 2 model 3 model 4 model 5 beliefs predicts knowledge and self-efficacy; and self-efficacy is predicted by knowledge and by beliefs beliefs predicts knowledge and self-efficacy beliefs, knowledge, and self-efficacy do not predict each other self-efficacy is predicted by knowledge and beliefs knowledge predicts beliefs and self-efficacy chi2 = 1.853 rmsea = 0.076 chi2 = 15.960 rmsea = 0.219 chi2 = 36.996 rmsea = 0.279 chi2 = 21.473 rmsea = 0.258 chi2 = 28.084 rmsea = 0.239 only for the model in which beliefs predicts knowledge and self-efficacy, while knowledge also predicts self-efficacy, the goodness-of-fit indices are acceptable. we therefore rejected the four alternative models and worked with the model presented in figure 2. figure 2. path model of predictors for teachers’ self-reported promotion of self-regulated learning including estimates (beta weights) of direct and indirect effects and explained variances (adjusted r2 coefficients). *p<.05; **p < .01; ***p < .001; tp < .10. based on the theoretical assumptions, we expected all variables to predict teachers’ selfreported promotion of srl, though not all variables were assumed to have a direct effect on the promotion of srl. the revised model indicates that teachers’ self-reported promotion of srl is predicted by the groups of variables: teacher beliefs, teacher knowledge, and teachers’ self-efficacy towards promoting srl. we found -.11* .10t .54*** .30*** .15*** .68*** epistemological beliefs (“changeable”) beliefs on promotion of srl self-efficacy for the promotion of srl knowledge on the promotion of srl self-reported promotion of srl .40** .18* dignath  et  al       | f l r     96   direct effects on teachers’ self-reported promotion of srl only for teacher beliefs on fostering srl and for teachers’ self-efficacy towards promoting srl.. teachers’ self-efficacy towards fostering srl had the largest direct effect (ß=.36, p<.001), followed by teacher beliefs (ß =.28, p<.001). we found no significant path from teacher knowledge on the promotion of srl to teachers’ self-reported promotion of srl (ß =.13, p=.07). no direct effect had been assumed from the teachers’ epistemological beliefs on teachers’ self-reported promotion of srl. however, teachers’ epistemological beliefs were found to have an indirect effect on teachers’ self-reported behavior in the classroom (see table 5). teachers’ self-efficacy was strongly predicted by teachers’ beliefs (ß =.31, p<.001) and teachers’ knowledge on promoting srl (ß=.29, p<.001), and negatively by teachers’ epistemological beliefs (ß =-.18, p<.05) (teachers who assumed learning behavior to be changeable reported a higher self-efficacy than teachers who assumed learning to be a fixed ability). we found teacher beliefs on the promotion of srl to be a highly significant predictor for teacher knowledge in this field (ß =.36, p<.001), but also teachers’ epistemological beliefs on the changeability of learning ability played a significant role (ß =.16, p<.05), with teachers believing in learning as a fixed ability showing more knowledge on the promotion of self-regulation strategies than teachers who didn’t assume learning to be a fixed ability. finally, teacher beliefs on the promotion of srl were predicted negatively by teachers’ epistemological beliefs in a way that teachers who believed in learning ability as not being fixed, were more positive towards the promotion of srl (ß =-.18, p<.05). table 5 reports regression coefficients of the path regression analyses for direct, indirect, and total effects. table 5 direct, indirect, and total effects for coefficients of the revised path model direct effect indirect effect total effect b (se) b (se) b (se) beliefs promotion of srl ß epistemological beliefs -.11 (.05)* 0 (no path) -.11 (.05)* knowledge promotion of srl ß beliefs promotion of srl epistemological beliefs .68 (.15)*** .18 (.09)* 0 (no path) -.07 (.04)* .68 (.15)*** .11 (.09) specific self-efficacy beliefs ß beliefs promotion of srl knowledge promotion of srl epistemological beliefs .30 (.07)*** .15 (.04)*** -.10 (04)* .10 (.02)*** 0 (no path) -.02 (.02) .40 (.08)*** .15 (.04)*** .12 (.05)** promotion of srl ß beliefs promotion of srl knowledge promotion of srl specific self-efficacy beliefs epistemological beliefs .40 (.11)*** .10 (.05)t .54 (.11)*** 0 (no path) .28 (.05)*** .08 (.02)*** 0 (no path) -.10 (.05)* .68 (.12)*** .18 (.06)** .54 (.11)*** -.10 (.05)* note: ***p < .001; **p < .01; *p < .05; tp < .10. dignath  et  al       | f l r     97   4. discussion 4.1 contribution of the study contrary to the field of srl, which has been covered by a large amount of research in the meantime, teachers’ promotion of srl has been a rather neglected field in this research area, leaving many questions open about how teachers can improve students’ self-regulation and how they can be supported. this study contributes to the field of srl in adding to what we already know about why teachers do or do not promote srl during their teaching. as research in this area has shown, teachers do not promote strategy instruction very often during their teaching (see e.g., bolhuis & voeten, 2001; dignath-van ewijk et al., 2013; hamman et al., 2000; lombaerts et al., 2007; spruce & bol, 2014; vandevelde et al., 2012). only few studies present results about teacher characteristics that have an impact on teachers’ support for srl during their teaching (dignath-van ewijk & van der werf, 2012; lombaerts et al., 2009; spruce & bol, 2014; vandevelde et al., 2012), and if they do, they only cover single aspects such as teachers’ educational beliefs or their knowledge. in order to build a model of teachers’ promotion of srl, we need to know more about the teacher variables that predict teachers’ promotion of srl and how they might interact with each other. moreover, most studies have investigated single aspects of the promotion of srl without drawing on a model of teachers’ professional competence. the present results provide us with a further insight into the predictors of teachers’ enhancement of srl embedded in a framework on teacher competence. 4.2 summary in this study we investigated the predictive impact of primary school teachers’ epistemological beliefs, their beliefs and their knowledge on the promotion of srl, and teachers’ self-efficacy beliefs as determinants of teachers’ self-reported promotion of srl in the classroom. as the results show, teachers’ self-efficacy towards promoting srl has the strongest direct predictive value on self-reported teacher behavior, while teacher beliefs have a strong direct and indirect impact via teacher self-efficacy and teacher knowledge. teacher knowledge has a direct, as well as an indirect, effect via teacher self-efficacy. teachers’ epistemological beliefs towards the changeability of learning ability shows a direct effect on teacher beliefs towards the promotion of srl, and it shows a direct effect on teachers’ self-efficacy beliefs in a way that teachers, who do not assume learning to be changeable, show lower self-efficacy and less positive beliefs towards the promotion of srl. moreover, we found a direct effect of teachers’ epistemological beliefs on teacher knowledge with teachers, who assume learning not to be changeable, showing more knowledge regarding the promotion of srl. the results offer new insights into teacher beliefs and how they might account for teacher (selfreported) behavior regarding the promotion of srl. the more teachers feel capable of instructing selfregulation strategies, and to manage self-regulating students, the more they report to promote more srl when they teach. furthermore, the more teachers report to believe that students can benefit from srl, the more they report to support their students by supplying them with self-regulation strategies, and by offering learning situations that allow for self-regulation. finally, the more teachers have knowledge on how to foster srl, the more they report to show supporting teaching behavior with regard to srl. dignath  et  al       | f l r     98   4.3 conclusions the following conclusions can be drawn with regard to our hypotheses: 4.3.1 hypothesis 1: epistemological beliefs about learning as fixed or changeable ability (in terms of rather general teacher beliefs) are assumed to influence teachers’ beliefs on the promotion of srl(as these are more specific). the more teachers think of learning as not being changeable, the less positive they are towards the promotion of srl. this result is not surprising: why would a teacher, who does not believe in the nature of learning as something to change and to develop, support the idea of students taking over responsibility for their own learning. as a fixed ability, learning to self-regulate can hardly make sense. it also supports the findings of pajares (1992) that more general beliefs predict more specific beliefs. 4.3.2 hypothesis 2: teacher beliefs are assumed to affect teacher knowledge on the promotion of srl. first, we found that teachers, who think of learning as not being changeable, report more knowledge than teachers who do not. at first sight, this result seems to be counterintuitive, but when looking at the operationalization of teacher knowledge, the finding makes more sense. teacher answers that included only one of the two aspects of fostering srl – 1. creation of a learning environment that allows for students’ selfregulation or 2. instruction of strategies that help students to handle a self-regulatory learning environment effectively – were ranked lower than teacher answers that included both aspects. however, most teachers, who included only one aspect, focused on the autonomous learning environment, and not on strategy instruction. this result replicates the results of former studies showing that most teachers, who support srl in school, particularly associate this with allowing students to have more freedom, but rarely providing them with the necessary strategies to deal with this autonomy (see e.g., bolhuis & voeten, 2001; de kock, sleegers & voeten, 2005; dignath-van ewijk et al., 2013). the results found here might indicate that teachers who think of learning as not being changeable, cannot integrate the idea of providing students with autonomy as a mean to support their self-regulation (lower ranked teacher answer) into their concept of knowledge. if, at all, they might accept the idea of giving students more autonomy by additionally teaching them first how to regulate this autonomy (higher ranked teacher answer). secondly, as assumed, the results showed that teachers, who are positive towards the promotion of srl, also share the idea of providing students with strategy knowledge, as well as with autonomy. 4.3.3 hypothesis 3: teacher beliefs and teacher knowledge have direct effects on self-reported teacher behavior, with teacher beliefs being a stronger predictor than teacher knowledge. teachers who know about the importance of creating a learning environment that allow students to self-regulate, and who are aware of the importance of providing students with the necessary strategies to deal with more autonomy, report to also implement these factors in their teaching. furthermore, our results confirm former evidence on the significance of teacher beliefs and teacher knowledge for self-reported teaching behavior. although both can play a role for teachers’ self-reported practice in the classroom, teacher beliefs seem to have a larger impact than does their knowledge (knowledge was not significant on the 5% level): on the one hand, by directly and strongly influencing teachers’ self-reported practice, and on the other hand, by influencing teacher knowledge and teacher self-efficacy which again predict teachers’ self-reported practice as well. this conforms with former evidence of teacher knowledge and teacher beliefs in general (see for an overview e.g., pajares, 1992) and more specifically for the promotion of srl, as well as with the results of spruce and bol (2014). although their findings of ten teachers did not deliver quantitative results, their descriptives showed that those teachers with the highest scores in classroom observations (i.e. teacher behavior) also reached higher scores for teacher beliefs, but not for teacher knowledge. looking at it the other way around, teachers with the highest scores on teacher beliefs also achieved higher observation scores, while teachers with the highest scores on teacher knowledge achieved only low observation scores. one could therefore assume that teacher beliefs would also be a better predictor for teacher behavior than teacher knowledge (spruce & bol, 2014). dignath  et  al       | f l r     99   teachers’ epistemological beliefs were not found to directly predict teachers’ self-reported practice. this is also in line with former research in which no close alignment between teachers’ epistemological world views and their teaching practice could be found (olafson & schraw, 2006), although the results in this area are mixed (see creemers et al., 2013; shraw & olafson, 2002; sosu & gray, 2012). ravindran, greene and debacker (2005) studied the relationships between several epistemological beliefs and the meaningful cognitive engagement of preservice teachers. although they could find positive relationships between most epistemological beliefs and cognitive engagement, the factor innate ability did not turn out to directly predict the cognitive engagement of preservice teachers (ravindran et al., 2005). 4.3.4 hypothesis 4: teachers’ self-efficacy to promote srl predicts teachers’ self-reported promotion of srl, next to teacher beliefs and teacher knowledge. the amount in which teachers feel competent enough to foster their students’ self-regulation depends on their beliefs, as well as on their knowledge. this result is also in line with the results of chatzistamatiou et al. (2014) who found teachers’ self-efficacy beliefs to predict teachers’ instruction of srl in mathematics. it is also in line with the teachers’ answers that perry et al. (2008) found, showing that teachers might have a positive attitude towards srl, but that they just do not feel able to support their students with their selfregulation. moreover, self-efficacy seems to play an even bigger role rather than just being another component in the model, as self-efficacy has the strongest predictive value among teacher beliefs and teacher knowledge. 4.3.5 hypothesis 5: teacher beliefs and teacher knowledge predict teachers’ self-efficacy. as expected, we found that the more positive teachers are towards the promotion of srl in primary school classrooms, and the more they know about supporting their students’ self-regulation, the more competent they feel with handling a learning environment conducive to self-regulation. as the goodness-offit indices had suggested, our initial model showed options to be improved in order to fit the data accordingly. we therefore refitted the model by including another path from teachers’ epistemological beliefs to teachers’ self-efficacy. initially, we did not assume teachers’ epistemological beliefs to predict teacher self-efficacy towards promoting srl directly, but rather only through teachers’ beliefs specifically towards the promotion of srl. however, this added path does make sense theoretically when considering teachers’ self-efficacy as beliefs as well. the self-efficacy beliefs of teachers can be considered as specific beliefs, as well, that are supposed to be influenced by more general beliefs (pajares, 1992). it is therefore theoretically reproducible that there is a direct path from teachers’ beliefs of learning as changeable to their self-efficacy beliefs about feeling able to promote srl in their classrooms. 4.4 limitations as for all research on srl or its promotion by teachers, which is carried out by means of self-report, this limitation also applies to this study: teachers might have tried to present themselves in a socially desired way, or, in other words, more positive (or negative) towards the promotion of srl then they actually are. moreover, the questionnaire data implies the risk that teachers might not understand or misunderstand certain items. finally, having asked teachers questions retrospectively, teacher answers might be incorrect due to problems with recall. all variables had been assessed by means of self-reporting. this problem seems to be the smallest with regards to teacher knowledge, since knowledge can somehow be assessed more objectively than can be beliefs. the problem is biggest for teacher behavior which was assessed as the teachers’ self-reported practice, since teachers might have answered in a most socially desirable way here. classroom observations would be more objective and could add to the reliability of the data in order to judge the teachers’ promotion of srl (see e.g., dignath-van ewijk et al., 2013). furthermore, by solely relying on the teachers’ self-report, all answers have been assessed within the same sample. high intercorrelations between all variables could be attributed to teachers following the same answering pattern for all questionnaires. by including an external perspective, e.g. of the students or external classroom observers, this problem could be resolved. moreover, we have limited this investigation to potential determinants on the dignath  et  al       | f l r     100   teacher level. yet, research by lombaerts et al. had shown that variables on the school level and on the student level could also have an impact on teachers’ promotion of srl (lombaerts et al., 2007). future research should also include these aspects in order to broaden the picture. particularly with regards to adaptive teaching, teachers’ reactions to student variables might play an important role. finally, the generalizability of our results is limited to the limited sample size. as participation in the study was voluntary, it might be that only very motivated teachers or teachers who have been interested in srl already agreed to participate. our sample would then not be representative. yet, the results can provide interesting first insights into the interrelation of determinants of the promotion of srl which should be investigated further with a larger sample. 4.5 implications for future research 4.5.1 implications for intervention research on srl research on the role of teachers in the promotion of srl is most notably intervention research. however, interventions to help teachers promote srl in their classrooms mainly focus on the instruction of what srl is and how teachers can create learning environments to foster srl. first, when developing and evaluating interventions, researchers should also include information about a potential inheritability of learning abilities in order to correct misconceptions of teachers. second, as the results show, teachers’ attitudes towards how beneficial srl is for their students, their epistemological beliefs on whether the learning behavior of their students is innate or not, as well as teachers’ own self-efficacy to promote srl, should be addressed in order to integrate not only cognitive, but also motivational aspects in teacher training. 4.5.2 implications for research on teacher knowledge and srl concerning theoretical contributions to the research of teachers’ promotion of srl, the question of what teacher knowledge on the promotion of srl implies and how it relates to epistemological beliefs has to be investigated further. in this study, teacher knowledge on promoting srl was defined as a two-folded construct, including the instruction of strategies plus the design of the learning environment (see dignathvan ewijk et al., 2013; paris & paris, 2001). more detailed analyses of these two aspects and their relation to teachers’ beliefs and behavior are needed to understand the result found in this study. wilson and bai (2010) investigated the relationship between teachers’ metacognitive knowledge and their pedagogical understandings of what is necessary for the teaching of metacognition. they found that teachers’ understanding of metacognition was related to their ideas of instructional strategies for srl (wilson & bai, 2010). based on their results, it would be interesting to further investigate in how far teachers’ understanding of srl and metacognition influences their knowledge on how to promote srl among their students as two steps in their knowledge of promoting srl: 1. the understanding of srl could serve as prior knowledge, and 2. the understanding of teaching srl would then be based on this prior knowledge. future research should take this first step of knowledge on srl into account when investigating further determinants of teachers’ promotion of srl. 4.5.3 implications for research on teacher self-efficacy and srl former research has shown the negative consequences of low self-efficacy in teachers, e.g. in terms of teacher burnout (skaalvik & skaalvik, 2007) or lower instructional quality (tschannen-moran & johnson, 2011; wertheim & leyser, 2002; swars, 2005). low teacher self-efficacy can have a negative impact on what teachers dare to try out in their classrooms. on the other hand, holzberger et al. (2013) investigated not only the effect of self-efficacy on instructional quality, but, in a longitudinal design, also the impact of instructional quality on teacher self-efficacy in the following school year. although their results also supported the results of former research on the effect of self-efficacy on instructional quality, they additionally found a reciprocal effect (holzberger et al., 2013). their findings suggest that self-efficacy and teacher behavior have rather a reciprocal relationship than just a one-sided one. as holzberger et al. (2013) dignath  et  al       | f l r     101   argue, the experience that teachers make during teaching is used as feedback from the students about the teachers’ instructional success. this feedback can serve in terms of one of bandura’s four factors that determine a person’s self-efficacy as described earlier: enactive attainment (bandura, 1986). for our study, this implies that it would be interesting to investigate in how far teachers’ promotion of srl can predict teachers’ self-efficacy. such a reciprocal relationship between teachers’ self-efficacy and their promotion of srl would not only be interesting for the theoretical understanding of how these aspects of teacher competence are related, but it would also have implications for research on teacher training. future research should investigate in how far students’ reactions on teachers’ behavior could serve as a feedback for the teacher on his or her success in promoting srl among the students, and in how far this feedback can serve to enhance teachers’ self-efficacy for continuing to foster srl. when implementing innovative teaching in schools, the self-efficacy of teachers should be taken into account in order to succeed with the implementation. this implies close cooperation with teachers, including getting to know what teachers find doable and finding out how to support them. 4.5.4 implications for research in other contexts finally, it would be interesting to conduct a similar study with high school teachers in order to compare the results found here with teachers from different contexts and different subject matters. we know from meta-analysis that, for the training of srl, different training characteristics are more or less effective for primary versus secondary school students and for different school subjects (dignath et al., 2008). therefore, one can assume that, for the promotion of srl, teachers might have made different experiences with different groups of students or within different subjects that affect their beliefs. 4.5.5 implications for research on building a model on teachers’ promotion of srl when looking at the literature on srl, one finds only the lack of a clear model on teaching srl. although there are contributions on how srl should be promoted (e.g., paris & paris, 2001; pressley et al., 1992; for an overview, see dignath, 2009), no specific model of teacher competence to foster srl exists yet. future research should connect the findings of studies on determinants of teachers’ promotion of srl (e.g., dignath-van ewijk & van der werf et al., 2012; lombaert et al., 2009; spruce & bol, 2014; vandevelde et al., 2012) and merge them with findings on how to assess teachers’ behavior in the classroom with regard to fostering srl by means of more sophisticated methods than self-reporting (see e.g., dignath-van ewijk et al., 2013 for observation methods on promoting srl) in order to collect more information on teacher competence in effectively promoting srl. keypoints this study reports research on determinants of teachers’ self-reported promotion of self-regulated learning an area which remains a gap in research on self-regulated learning. the investigation includes teacher beliefs on (1) teachers’ self-reported practice on instructing srl, (2) teachers’ self-efficacy with regard to promoting srl, and (3) teachers’ epistemological beliefs regarding learning. path analyses were conducted and reveal new insights into constructing a model of teachers’ selfreported promotion of srl in the classroom. a large teacher sample participated in the study in order to provide representative classroom data. the study is innovative as there is no research yet regarding teacher beliefs and teacher selfefficacy predicting teachers’ self-reported practice of enhancing srl. dignath  et  al         | f l r 102     acknowledgments the author thanks the student assistants cornelia haaß, melanie scheuermann, and katharina uhrig for their help with the data collection, as well as oliver dickhäuser for his comments on an earlier draft of this article. references ajzen, i. (1991). the theory of planned behavior. organizational behavior and human decision processes, 50, 179–211. bandura, a. (1977). self-efficacy: toward a unifying theory of behavioral change. psychological review, 84, 191. bandura (1986). social foundations of thought and action: a social cognitive theory. englewood cliffs, nj: prentice-hall. baumert, j. & kunter, m. (2013). the coactiv model of teachers’ professional competence. in m. kunter, j. baumert, w. blum, u. klusmann, s. krauss & m. neubrand (eds.), cognitive activation in the mathematics classroom and professional competence of teachers (pp. 25-48). new york: springer us. baumert, j. & kunter, m. (2006). stichwort: professionelle kompetenz von lehrkräften [keyword: professional competence of teachers. zeitschrift für erziehungswissenschaft, 9, 469-520. bell, p. d. (2006). can factors related to self-regulated learning and epistemological beliefs predict learning achievement in undergraduate asynchronous web-based courses? perspectives in health information management/ahima, american health information management association, 3, 7. bendixen, l. d. & hartley, k. (2003). successful learning with hypermedia: the role of epistemological beliefs and metacognitive awareness. journal of educational computing research, 28, 15-30. bolhuis, s. & voeten, m. j. m. (2001). toward self-directed learning in secondary schools: what do teachers do? teaching and teacher education, 17, 837–855. brown, a. l., campione, j. c. & day, j. d. (1981). learning to learn: on training students to learn from texts. educational researcher, 10, 14-21. canrinus, e. t., helms-lorenz, m., beijaard, d., buitink, j. & hofman, a. (2012). self-efficacy, job satisfaction, motivation and commitment: exploring the relationships between indicators of teachers’ professional identity. european journal of psychology of education, 27, 115-132. chan, w. y., lau, s., nie, y., lim, s. & hogan, d. (2008). organizational and personal predictors of teacher commitment: the mediating role of teacher efficacy and identification with school. american educational research journal, 45, 597–630. chatzistamatiou, m., dermitzaki, i. & bagiatis, v. (2014). self-regulatory teaching in mathematics: relations to teachers' motivation, affect and professional commitment. european journal of psychology of education, 29, 295-310. de kock, a., sleegers, p. & voeten, m. j. m. (2005). new learning and choices of secondary school teachers when arranging learning environments. teaching and teacher education, 21, 799–816. dignath-van ewijk, c., dickhäuser, o. & büttner, g. (2013). assessing how teachers enhance self-regulated learning a multi-perspective approach. journal of cognitive education and psychology, special issue on self-regulated learning, 21, 338-358. dignath-van ewijk, c. & van der werf, g. (2012). what teachers think about self-regulated learning: an investigation of teacher beliefs about enhancing students' self-regulation and how they predict teacher behavior. education research international, doi:10.1155/2012/741713. dignath  et  al         | f l r 103     dignath, c. (2009). different aspects of the promotion of self-regulated learning: a multi-method investigation on the instruction of self-regulated learning at primary and secondary school. dissertation universität frankfurt. dignath, c., büttner, g. & langfeldt, h.-p. (2008). how can primary school students acquire self-regulated learning most efficiently? a meta-analysis on interventions that aim at fostering self-regulation. educational research review, 3, 101-129. doornik, j. a. & hansen, h. (2008). an omnibus test for univariate and multivariate normality. oxford bulletin of economics and statistics, 70, supplement, 927-939. dweck, c. s. & leggett, e. l. (1988). a social-cognitive approach to motivation and personality. psychological review, 95, 256. ertmer, p. a. (2005). teacher pedagogical beliefs: the final frontier in our quest for technology integration? educational technology research and development, 53, 25–39. fenstermacher, g. d. (1994). the knower and the known: the nature of knowledge in research on teaching. review of research in education, 20, 3-56. fives, h. (2003, april). what is teacher efficacy and how does it relate to teachers’ knowledge? a theoretical review. american educational research association annual conference, chicago. ghaith, g. & yaghi, h. (1997). relationships among experience, teacher efficacy, and attitudes towards the implementation of instructional innovation. teaching and teacher education, 13, 451-458. guo, y., piasta, s. b., justice, l. m. & kaderavek, j. m. (2010). relations among preschool teachers’ selfefficacy, classroom quality, and children’s language and literacy gains. teaching and teacher education, 26, 1094-1103. guskey, t. r. (1988). teacher efficacy, self-concept, and attitudes toward the implementation of instructional innovation. teaching and teacher education, 4, 63-69. hamman, d., berthelot, j., saia, j. & crowley, e. (2000). teachers' coaching of learning and its relation to students' strategic learning. journal of educational psychology, 92, 342. hattie, j. (2013). visible learning: a synthesis of over 800 meta-analyses relating to achievement. london: routledge. hattie, j., biggs, j. & purdie, n. (1996). effects of learning skills interventions on student learning: a metaanalysis. review of educational research, 66, 99-136. hofer, b. k. (2000). dimensionality and disciplinary differences in personal epistemology. contemporary educational psychology, 25, 378-405. hofer, b. k., & pintrich, p. r. (1997). the development of epistemological theories: beliefs about knowledge and knowing and their relation to learning. review of educational research, 67, 88-140. holzberger, d., philipp, a. & kunter, m. (2013). how teachers’ self-efficacy is related to instructional quality: a longitudinal analysis. journal of educational psychology, 105, 774-786. kagan, d. m. (1992). implication of research on teacher belief. educational psychologist, 27, 65–90. klassen, r. m., tze, v. m., betts, s. m. & gordon, k. a. (2011). teacher efficacy research 1998–2009: signs of progress or unfulfilled promise?. educational psychology review, 23, 21-43. kramarski, b., desoete, a., bannert, m., narciss, s. & perry, n. (2013). new perspectives on integrating self-regulated learning at school. education research international, 2013, article id 498214. kramarski, b. &. michalsky, t. (2009). investigating preservice teachers’ professional growth in selfregulated learning environments. journal of educational psychology, 101, 161–175. kunter, m. (2013). motivation as an aspect of professional competence: research findings on teacher enthusiasm. in m. kunter, j. baumert, w. blum, u. klusmann, s. krauss & m. neubrand (eds.), cognitive activation in the mathematics classroom and professional competence of teachers (pp. 273289). new york: springer us. kunter, m., tsai, y. m., klusmann, u., brunner, m., krauss, s. & baumert, j. (2008). enjoying teaching: enthusiasm and instructional behaviors of secondary school mathematics teachers. learning and instruction, 18, 468-482. lombaerts, k., engels, n. & athanasou, j. a. (2007). development and validation of the self-regulated learning inventory for teachers. perspectives in education, 25, 29–47. dignath  et  al         | f l r 104     lombaerts, k., engels, n., van braak, j. & athanasou, j. a. (2009). development of the self-regulated learning teacher belief scale. european journal of psychology of education, 1, 79-96. lonka, k., joram, e. & bryson, m. (1996). “conceptions of learning and knowledge: does training make a difference?” contemporary educational psychology, 21, 240–260. moely, b. e., hart, s. s., leal, l., santulli, k. a., rao, n., johnson, t. & hamilton, l. b. (1992). the teacher's role in facilitating memory and study strategy development in the elementary school classroom. child development, 63, 653-672. moos, d. c., & ringdal, a. (2012). self-regulated learning in the classroom: a literature review on the teacher’s role. education research international, 2012, article id 423284. muis, k. r. (2007). the role of epistemic beliefs in self-regulated learning. educational psychologist, 42, 173-190. olafson, l. & schraw, g. (2006). teachers’ beliefs and practices within and across domains. international journal of educational research, 45, 71-84. pajares, m. f. (1992). teachers’ beliefs and educational research: cleaning up a messy construct. review of educational research, 62, 307–332. paris, s. g. & paris, a. h. (2001). classroom applications of research on self-regulated learning. educational psychologist, 36, 89-101. perry, n. e., hutchinson, l. & thauberger, c. (2008). talking about teaching self-regulated learning: scaffolding student teachers’ development and use of practices that promote self-regulated learning. international journal of educational research, 47, 97–108. perry, n. e., phillips, l. & dowler, j. (2004). examining features of tasks and their potential to promote self-regulated learning. the teachers college record, 106, 1854-1878. perry, n. e. & vandekamp, k. j. (2000). creating classroom contexts that support young children's development of self-regulated learning. international journal of educational research, 33, 821-843. perry, w. j. r. (1970). forms of intellectual and ethical development in the college years: a scheme. new york: holt, rinehart and winston. pieschl, s., stahl, e. & bromme, r. (2008). epistemological beliefs and self-regulated learning with hypertext. metacognition and learning, 3, 17-37. pressley, m., harris, k. r. & marks, m. b. (1992). but good strategy instructors are constructivists!, educational psychology review, 4, 3–31. ravindran, b., greene, b. a. & debacker, t. k. (2005). predicting preservice teachers’ cognitive engagement with goals and epistemological beliefs. the journal of educational research, 98, 222-232. schiefele, u., moschner, b. & husstegge, r. (2002). skalenhandbuch smile-projekt. [scale handbook of the smile project] bielefeld: university of bielefeld. schmitz, g. s. & schwarzer, r. (2000). selbstwirksamkeitserwartung von lehrern: längsschnittbefunde mit einem neuen instrument. [perceived self-efficacy of teachers: longitudinal findings with a new instrument] zeitschrift für pädagogische psychologie, 14, 12-25. schommer-aikins, m. (2004). explaining the epistemological belief system: introducing the embedded systemic model and coordinated research approach. educational psychologist, 39, 19-29. schommer, m. (1990). effects of beliefs about the nature of knowledge on comprehension. journal of educational psychology, 82, 498-504. schraw, g., crippen, k. j. & hartley, k. (2006). promoting self-regulation in science education: metacognition as part of a broader perspective on learning. research in science education, 36, 111139. schraw, g. & olafson, l. (2003). teachers' epistemological world views and educational practices. journal of cognitive education and psychology, 3, 178-235. shulman, l. s. (1986). those who understand: knowledge growth in teaching. educational researcher, 15, 4-14. sinatra, g. m. & kardash, c. m. (2004). teacher candidates’ epistemological beliefs, dispositions, and views on teaching as persuasion. contemporary educational psychology, 29, 483-498. dignath  et  al         | f l r 105     skaalvik, e. m. & skaalvik, s. (2010). teacher self-efficacy and teacher burnout: a study of relations. teaching and teacher education, 26, 1059-1069. skaalvik, e. m. & skaalvik, s. (2007). dimensions of teacher self-efficacy and relations with strain factors, perceived collective teacher efficacy, and teacher burnout. journal of educational psychology, 99, 611. sosu, e. m. & gray, d. s. (2012). investigating change in epistemic beliefs: an evaluation of the impact of student teachers’ beliefs on instructional preference and teaching competence. international journal of educational research, 53, 80-92. spruce, r. & bol, l. (2014). teacher beliefs, knowledge, and practice of self-regulated learning. metacognition and learning, 1-33. swars, s. l. (2005). examining perceptions of mathematics teaching effectiveness among elementary preservice teachers with differing levels of mathematics teacher efficacy. journal of instructional psychology, 32, 139-147. tillema, h. h. (1995). changing the professional knowledge and beliefs of teachers: a training study. learning and instruction, 5, 291–318. tschannen-moran, m. & johnson, d. (2011). exploring literacy teachers’ self-efficacy beliefs: potential sources at play. teaching and teacher education, 27, 751-761. tschannen-moran, m. & woolfolk hoy, a. (2001). teacher efficacy: capturing an elusive construct. teaching and teacher education, 17, 783-805. tschannen-moran, m., hoy, a. w. & hoy, w. k. (1998). teacher efficacy: its meaning and measure. review of educational research, 68, 202-248. vandevelde, s., vandenbussche, l. & van keer, h. (2012). stimulating self-regulated learning in primary education: encouraging versus hampering factors for teachers. procedia-social and behavioral sciences, 69, 1562-1571. weinert, f. e. (2001). concept of competence: a conceptual clarification. in rychen, d. s. & salganik., l. h. (eds.), defining and selecting key competencies (pp. 45-65). ashland, oh: hogrefe. wertheim, c. & leyser, y. (2002). efficacy beliefs, background variables, and differentiated instruction of israeli prospective teachers. the journal of educational research, 96, 54-63. wilson, n. s. & bai, h. (2010). the relationships and impact of teachers’ metacognitive knowledge and pedagogical understandings of metacognition. metacognition and learning, 5, 269-288. wright, s. (1934). the method of path coefficients. the annals of mathematical statistics, 5, 161-215. woolfolk, a. e. & hoy, w. k. (1990). prospective teachers’ sense of efficacy and beliefs about control. journal of educational psychology, 82, 81-91. yadav, a. & koehler, m. (2007). the role of epistemological beliefs in preservice teachers’ interpretation of video cases of early-grade literacy instruction. journal of technology and teacher education, 15, 335361. zimmerman, b. j. (2000). attaining self-regulation: a social cognitive perspective. in m. boekaerts, p. r. pintrich & m. zeidner (eds.), handbook of self-regulation (pp. 13-39). san diego, ca, us: academic press. microsoft word laine et al publications.docx frontline learning research vol.5 no. 4 (2017) 42 – 60 issn 2295-3159 corresponding author: erkka laine, department of teacher education, faculty of education, 20014 university of turku, finland. e-mail: erkka.laine@utu.fi doi: https://doi.org/10.14786/flr.v5i4.306 generation of student interest in an inquiry-based mobile learning environment erkka laine, marjaana veermans, aleksi lahti, koen veermansa auniversity of turku, finland article received 5 may / revised 16 july / accepted 9 october / available online 15 november abstract a declining trend in adolescents’ interest in science learning and attitudes towards science-related careers has been reported during recent years. there has been a call for more motivating learning environments that inspire students to develop interest towards science. this study examines students’ interest development in stem subjects in an ecologically valid setting during one school year and how features of the learning environment affect students’ generation of interest. in a quasi-experimental study design, one class of 7th grade (aged 12 to 13 years) students (n = 18) studied in an inquirybased mobile learning environment that had a special emphasis on integrated curriculum. interest variables were measured three times and focus group interviews were held twice during the school year. from a group of 113 students studying in an ordinary learning setting, a propensity score-matched control group of 18 students was selected based on general self-efficacy, intrinsic goal orientation, interest in technology, and web-user self-efficacy. results from the quantitative analyses revealed only minor differences between the two groups. results from the qualitative analyses indicate that students found the new environment to be interest generating, thus ascribing to the general idea and aim of the new environment, but also that the implementation was in many cases far from ideal, indicating that much of its potential was unrealized. keywords: interest; science learning; inquiry learning; mobile learning; mixed method. laine et al 42 | f l r 1. introduction researchers have reported a continuous worldwide decline in student interest in science learning over the last two decades (e.g. rocard et al., 2007). among others, large-scale quantitative studies by the global science forum (organisation for economic co-operation and development, 2006) and the international study center (martin, mullis, foy & hooper, 2016) indicate a declining trend in adolescents’ interest in science learning and attitudes to science-related careers. while the general trend suggests that students start to lose interest in science-related subjects in upper secondary school, there is considerable variation across science domains. reasons for the observed decline also appear diverse—for example, students’ interest may depend on the quality and type of instruction offered in schools; the psychological demands of adolescent life may take priority; or the student’s ideal self-concept may be at odds with their perception of science learning (krapp & prenzel, 2011). also students’ individual characteristics, such as self-efficacy and goal orientation, have been found to influence interest (glynn, bryan & brickman, 2015; tapola, veermans & niemivirta, 2013). it has been suggested that one factor in the decline in interest in science, technology, engineering, and mathematics (stem) is that these subjects still tend to be taught using traditional teacher-led pedagogies (rocard, 2007). by way of response, several pedagogical approaches have been developed in an attempt to foster self-regulated learning, active agency, and learning by doing. among these, inquiry-based learning has been found to be positively related to academic achievement. several meta-analyses and research syntheses have reported higher overall mean effect sizes for inquiry-based science teaching and learning when compared to more traditional types of instruction (furtak, seidel, iverson & briggs 2012; alfieri, brooks, aldrich & tenenbaum, 2011;. minner, levy & century, 2010). this has also proven to be the case for digital learning environments, where students are offered the necessary support for inquiry (de jong, 2006). though there have been examples of interventional and case studies (e.g. knogler, harackiewicz, gegenfurtner & lewalter, 2015; renninger et al., 2014; tapola et al., 2013;) on the relationship between inquiry learning and students’ interest development these studies have either concentrated on short-term change in students’ situational interest, and/or lacked the ecological validity that would benefit the development of everyday classroom practises. studying interest development in an ecologically valid setting over an extended time frame would allow the examination of the benefits and possible disadvantages of these types of pedagogical approaches as they unfold in real life classroom situations (e.g., glynn et al., 2015). the purpose of the present study was to examine the development of lower secondary school students’ interest in stem domains using an inquiry-based mobile learning environment as compared to an ordinary learning environment, employing a quasi-experimental design. in addition to quantitative measurements of interest and individual characteristics, group interviews were carried out among the experimental group students. the design of the experimental learning environment was inspired by the new finnish curriculum which was under development by the time of this study. the core ideas of the new curriculum are student initiated inquiry activities, collaboration, and integration of different school subjects into one meaningful unit (finnish national board of education, 2014). 2. theoretical framework 2.1. interest according to the four-phase model of interest development (hidi & renninger, 2006; renninger & hidi, 2011), the concept of interest is a multidimensional construct consisting of affect, knowledge, and laine et al 43 | f l r value components that are related to a certain subject, situation or activity. the four developmental phases of the model, namely triggered situational interest, maintained situational interest, emerging individual interest, and well-developed individual interest, differ from each other on how much weight the affect, knowledge, and value components get. situational interest refers to a state of focused attention and affective reaction that appears in a moment and may or may not persist over time. it does not necessarily depend on the individual’s prior knowledge of the subject but is caused and supported by certain external stimuli within the situation (hidi & renninger, 2006). the affective component of interest predominates in the early stages of interest development, as attention is initially triggered by an affective response, and prior knowledge may not yet exist. individual interest, then, can be viewed as a more permanent positive orientation toward an activity or a domain that has value for that person; here, it is less that the activity triggers the engagement than that the person’s prior knowledge and experiences of value and enjoyment in that context determine their interest (ainley, 2012). in this long-term orientation, the person may autonomously engage and re-engage with the content despite significant efforts required and occasional setbacks. at this stage, the learner experiences not only positive feelings but also a sense of greater knowledge and value of the content in question. the more developed the individual interest, the more an individual will participate in self-regulated learning activities such as seeking answers to their own curiosity questions (hidi & renninger, 2006). well-developed individual interest is based strongly on cognitive factors such as well-founded and meaningful goal structures (hidi & renninger, 2006; krapp & prenzel, 2011). interest can also be viewed from the perspective of classroom learning, and especially the various domains of different school subjects. according to fryer, ainley & thompson (2016), domain interest can be defined as students’ interest to a certain body of knowledge, such as interest in mathematics, and it resembles conceptually individual interest because of its connection to knowledge and value components of interest. contrary to the university students that fryer et al. (2016) were studying, primary and secondary school children usually have less freedom to choose what domains to study. this poses challenges to teachers because of the various phases of interest development that the students are in. in order for students’ interest to develop towards a particular domain, the activities and tasks used in the teaching should help to trigger interest in the more novice students and offer enough challenge and autonomy for the more advanced ones. previous research literature has found interest to have several positive outcomes in learning, such as increased attention, use of learning strategies, and goal setting (for reviews, see renninger & hidi, 2011; potvin & hasni, 2014). 2.2 individual characteristics and interest students differ from each other not only in terms of interest development, but also on other individual characteristics. in this study we controlled for two such characteristics, namely self-efficacy, and intrinsic goal orientation, because these have been found to have a relationship with interest. self-efficacy is defined as the individual’s appraisal of one’s own capabilities to carry out a task along with the confidence in one’s skills to complete that task (pintrich, smith, garcia & mckeachie, 1991) it has been found to have at least a moderately sized correlation to interest, and in addition this correlation has been found to be stronger in mathematics and science than in any other subjects (rottinghaus, larson & borgen, 2003; bong, lee & woo, 2015). reasons for this strong relationship may be found in the hierarchical nature of mathematic and scientific knowledge, inflexible characteristics of mathematical instruction, and students’ perceptions of mathematics and science as difficult subjects. the relationship is hypothesized to be reciprocal, where on the one hand student’s level of self-efficacy in a particular domain may influence whether or not he or she generates interest towards it while on the other hand, interest may evoke feelings of competence and selflaine et al 44 | f l r efficacy (bong et al., 2015). intrinsic goal orientation, or mastery goal orientation (for terminology, see ames, 1992; vansteenkiste, lens & deci, 2006), by definition is the individual’s perception of oneself to be participating in a task for the sake of challenge, curiosity, or mastery. in educational settings, a student with and intrinsic goal orientation views an academic task to possess not only instrumental, but also inherent value (pintrich et al., 1991). it has been found to have a positive connection to interest (harackiewicz, barron, linnenbrink-garcia & tauer, 2008; tapola et al., 2013). this relationship is also thought to be reciprocal, so that students’ initial interest in the subject predicts their adoption of mastery goals, and the adoption of mastery goals deepens their interest by motivating them to be more task-focused (harackiewicz et al., 2008). 2.3. generation of interest in inquiry-based learning environments generation of interest refers to the ways in which a learner develops an interest in the activity at hand. on a practical level the challenge is usually twofold: how to trigger the interest of those students who aren’t already drawn to the subject, and at the same time help already interested students to maintain their interest. from the teacher’s perspective, it is essential to know how to design teaching so that it triggers the interest of uninterested students, and supports the development of individual interest for the others. in particular, experiences of novelty, challenge, physical activity, and social involvement have been identified as factors contributing to interest generation (renninger & hidi, 2011; palmer, 2009). also personalizing the context of the instructions to match the students’ individual interests, using problem-based learning scenarios, and helping students find meaning and value in their studies through utility-value interventions have been found helpful in supporting students’ motivation (harackiewicz, smith & priniski, 2016). known sources of individual interest generation in science include seeking opportunities to learn through engagement with parents and peers, initiating one’s own projects, finding settings for both formal and informal learning, and pursuing a variety of information resources (crowley, barron, knutson & martin, 2015). beyond addressing the issue by designing new curriculum content and presentations, students should also be offered early opportunities to experience science as this supports both triggered situational interest and its maintenance and further development (ainley & ainley, 2015). as a one pedagogical method that utilizes some of these principles, inquiry-based, or inquiry learning has been found to relate positively to students’ achievement and interest in science learning (renninger et al., 2014). inquiry learning is a process in which the learner plays an active part in discovering new relationships within a certain phenomenon. the process involves formulating hypotheses and testing their validity through experimentation or observation (pedaste, mäeots, leijen & sarapuu, 2012). for such a process to be efficient, it can be divided into phases, each with separate characteristics and subgoals. in one form or another, these phases commonly include orientation to the current problem, conceptualization of key ideas related to the problem, investigation of the problem through data collection and analysis, interpretation of findings, and discussion of the results and the process (pedaste, et al., 2015). inquiry learning provides opportunities to link learning to students’ own interests, and when mobile devices are embedded in inquiry learning environments, learning can occur ubiquitously, focusing on aspects of students’ everyday experience. the rapid development and uptake of advanced digital devices is both promising and challenging for the design of new learning environments that will interest and engage students by enabling them to choose between different learning activities and tasks. in a study conducted among taiwanese high school students lai, hwang, liang & tsai (2016) found that especially meaningful and relevant information from multiple sources predicted students’ willingness to engage in inquiry learning. it seems that mobile and inquiry learning share some commonalities that may benefit one another. mobile laine et al 45 | f l r learning’s ability to remain detached from spatial and material boundaries offers possibilities for inquiry to expand beyond the traditional classroom environment. inquiry-based approaches have been found to positively influence students’ interest development in different types of science education settings. in their study renninger et al. (2014) found that an inquirybased intervention that was designed to help at-risk middle school students in their science studies had a positive impact on students with varying levels of interest. those students who had already more developed interest at the beginning of the workshop were able to dig deeper into the subject and approach it more individually and systematically. on the other hand, those students with less developed interest could start off less systematically but gradually develop more complex thinking that would also transfer from one session to another. this is one of the possible benefits of inquiry-based learning since it allows more individuality in terms of study pace. 2.4. context of the study the new learning environment was designed to utilize the principles of inquiry-based learning, with a particular emphasis on integrated curriculum. the new finnish curriculum that was under development at the time of this study relied on a conception of the learner as an active agent, setting individual goals and solving problems, independently and in collaboration with others. students’ self-reflection of experiences and emotions, the joy of learning, and creativity are expected to expand their sphere of interests. supporting students’ growth to humanity and ethically responsible citizenship, as well as lifelong learning are seen as central goals of education (finnish national board of education, 2014). to implement the requirements of the new national curriculum, the participating school had designed a new pedagogical approach that would combine different school subjects, especially in the science domain, in various learning projects. the school had received outside funding for a three year project and had selected one 7th grade class of 18 students as the pilot group. this study reports the first year of that project. the students received personal tablet computers as their primary tool for school work. the tablets contained different types of learning software, such as e-books and e-notebooks and applications for mind mapping, instant messaging, and image and video editing. in addition, all of the subject teachers had an opportunity to participate in professional training courses to facilitate implementation of pedagogical practices in the new learning environment. teaching in the new learning environment relied on principles of inquiry-based learning and integrated curriculum. the approach adopted in the present study can be characterised as multidisciplinary, as the learning projects and inquiry activities revolved around one overarching theme which would then be covered in different subjects (drake & burns, 2004). in practice, this meant that the students were planning and conducting experiments in learning projects combining elements from different school subjects and involving work outside the classroom. depending on the requirements of each task, the students worked individually, in groups, or in pairs. students used the tablet computers to document experiments, to retrieve information, to work simultaneously on a task, and to present their findings to other students in the classroom, as well as for homework and exam rehearsals. one example of such activity was a crosscurricular, water-themed project, in which students would conduct field experiments by collecting water samples, analysing them in biology and chemistry classes, taking photographs of the surrounding environment in arts class, and writing reports on the quality of the water in language class. laine et al 46 | f l r 3. aims of the study in examining how lower secondary school students’ domain-specific interest is developed during the first year in an inquiry-based mobile learning environment, external factors such as how features of that environment contribute to students’ interest were of particular interest. a quasi-experimental study design was used to compare the new learning environment with a regular learning environment. the research questions were as follows. 1.1 how does domain-specific interest among students in the digital learning group develop during the school year? 1.2 does this differ from development among control group students? 2.1 what aspects of the new learning environment help or hinder students’ interest generation? 2.2 what components of the experience generated student interest? to answer questions 1.1 and 1.2 quantitative data were collected by means of interest questionnaires. for a more detailed view, questions 2.1 and 2.2 were addressed by qualitative means, using focus group interviews with the experimental group. 4. method 4.1. participants the study participants were 131 7th grade students (69 girls and 62 boys, aged 12–13 years) from seven classes at one lower secondary school in southern finland. one class (n = 18) had been chosen by the school to undergo classroom teaching using only digital learning tools, and these students formed the experimental group in a quasi-experimental research design. the remaining students (n = 113) continued to receive more traditional teaching using materials such as books, notebooks, and handouts. from these 113 students 18 propensity score-matched students were selected as the control group for the experimental digital learning group. the use of propensity score matching creates a control for students’ initial motivational profiles at the beginning of the school year, and ensures that each student from the digital learning group has a counterpart with a similar profile at the beginning of the school year. the defining variables used for matching were general self-efficacy, intrinsic goal orientation, interest in technology, and web-user selfefficacy. 4.2. procedure to establish students’ initial level of interest and motivation, a self-report questionnaire was administered to students at the beginning of the school year (time 1). after this initial measurement, day-today classroom work went on uninterrupted, and teachers and students were free to implement the learning environment as they saw fit. similar sets of questionnaires were administered on two further occasions during the school year: four months after the start of the school year (time 2), and nine months after the start, at the end of the school year (time 3). the questionnaires included measures of individual interest, motivational beliefs, and self-efficacy, which are described in more detail in the next section. laine et al 47 | f l r 4.3 measures 4.3.1 domain-specific individual interest an instrument from tapola et al. (2013) was used to assess students’ domain-specific individual interest in the science domain. the instrument measured interest in three subjects (mathematics, chemistry and physics, and biology) for a single item on a five-point scale, ranging from 1 (not at all interested) to 5 (very interested). an example item would be “how interested are you in mathematics”. interest in technology and interest in collaboration were measured for five items on five-point scales ranging from 1 (not at all interested) to 5 (very interested), with two reversed items in each. example items for both respectively would be “i find working with technology interesting”, and “working together with other students is interesting” 4.3.2 individual characteristics students’ intrinsic goal orientation and self-efficacy for learning and performance were measured using selected scales from the motivated strategies for learning questionnaire (pintrich et al., 1991). intrinsic goal orientation (e.g. in a class like this, i prefer course material that really challenges me so i can learn new things) was measured for four items, and self-efficacy for learning and performance (e.g. i'm certain i can understand the most difficult material presented in the readings for this course) was measured for eight items on a seven-point scale. in addition, students’ web-user self-efficacy was measured using a modified version of the wuse scale (eachus, cassidy & hogg, 2006). the instrument consisted of 14 items, measured on a five-point scale from 1 (strongly disagree) to 5 (strongly agree) (e.g. i would never try to download files from the internet, that would be too complicated). 4.4. focus group interviews to gain a better insight into students’ experiences of the new learning environment, two interview sessions were arranged during the school year. in the first session, the students in the experimental group were interviewed in groups of four or five. the first interview session was conducted in the middle of the autumn semester, two months after the start of the school year; this time point was chosen because the students would by then have been sufficiently familiar with the environment to have formed an opinion about it. the time period was considered short enough to enable recall of aspects of the learning environment perceived as interest-generating, and long enough for differences in interest to emerge. the second interview session was conducted at the end of the school year. in all, four groups of students were interviewed during the autumn semester and three during the spring semester. depending on how actively students participated in the discussion, interviews varied in duration between 20 and 34 minutes. the interviews were designed on the principles of stimulated recall, first asking students to describe an event or topic that they remembered from their studies during the past semester. follow-up questions would then explore what had been of particular interest, and why, with additional emphasis on crosscurricular and inquiry learning elements in the students’ responses. students were also asked to identify elements that they felt had not worked, and to propose how the learning environment could be further developed. 4.5. analyses to answer the first research question, concerning how domain-specific interest develops during the school year and whether there are differences between the digital and traditional learning groups, the laine et al 48 | f l r quantitative data were analysed in two phases. in the first phase, a repeated measures anova was conducted to establish how mean levels of domain-specific interest developed in the digital learning group. this was followed by paired sample t-tests to identify statistically significant changes in those variables based on the anova results. a similar procedure was used to detect mean level differences between the levels of interest of the digital learning group and the propensity score-matched traditional learning group. following repeated measures anovas, statistically significant differences were compared separately for each time point, using independent samples t-tests. there was no indication that the data would not be normally distributed and for that reason the analyses proceeded as planned. to analyse the interview data a twofold analysing scheme was created. the first main category of the scheme, named ‘generation of interest’, sought to clarify what aspects of the new learning environment had helped to generate student interest. this category focused on the learning environment’s pedagogical design and technical features. the second main category of the analysing scheme was called ‘components of interest’ and focused on students’ experiences of studying in the new environment. its aim was to determine which components had contributed to interest generation. in both of these categories, the analyses were data-driven; this meant that there were no ready-made subcategories for the students’ answers (see appendix a for more information on the coding scheme, including categories and subcategories and excerpts from the data). using content analysis, the coding scheme addressed 1) generation of interest (issues that students identified as interesting) and 2) components of interest (why certain issues were perceived to be interesting). each main category was then further and separately analysed, grouping expressions into subcategories according to similarity of content. as each interview session was conducted in groups of four or five students, one expression may consist of several students participating in the conversation at the same time. an expression would be counted as being started when some of the students brought it up first time during the interview, and would end as the conversation flowed to another topic. in all, the generation of interest category attracted five subcategories, and the components of interest category attracted four. in addition, two categories were formed for critical remarks and developmental ideas. the analysis was conducted by the first author, and the second author coded half of the data to assess intercoder agreement (90%). 5. results 5.1. digital learning group: domain specific interest development during the school year item analyses confirmed that all scales were reliable (alphas varying between .72 and .93). to find out how the experimental group’s interest developed during the school year, a repeated measures anova was conducted for all five subject-specific interest measures for time points 1, 2, and 3. the tests showed statistically significant differences between interest in physics and chemistry (f(2, 15.098) = 7.612, p = .002) and interest in biology (f(2, 16.987) = 4.957, p = .013). the mean level development of the experimental group’s subject-specific interest is shown in figure 1. paired sample t-tests showed a significant decrease in experimental group students’ interest in physics and chemistry from time point 1 (m = 4.17, sd = .79) to time point 2 (m = 3.48, sd = .91); t(18) = 2.99, p = .008, cohen’s d = 0.83. during the same time period, students’ interest in biology increased from time point 1 (m = 2.72, sd = .90) to time point 2 (m = 3.38, sd = .90); t(18) = 2.60, p = .019, cohen’s d = 0.73. during the spring semester, between time point 2 and time point 3, no significant changes in interest levels were detected. laine et al 49 | f l r figure 1. digital learning group’s subject-specific interest during the school year. 5.2. differences in interest development between the experimental and control groups a repeated measures anova was conducted to determine whether interest variables, motivational variables, and school grades differed significantly between the two groups. the results show a significant effect of classroom condition on interest in physics and chemistry (f(2, 68) = 6.35, p = .003). as shown in figure 2, the experimental group’s interest in physics and chemistry declined between time point 1 (m = 4.17, sd = 0.79) and time point 2 (m = 3.48, sd = 0.91) but stabilized at time point 3 (m = 3.36, sd = 1.08). the control group’s interest in physics and chemistry increased between time point 1 (m = 3.07, sd = 1.10) to time point 2 (m = 3.46, sd = 0.98) but decreased again between time point 2 and time point 3 (m = 3.02, sd = 0.91). an independent samples t-test revealed that although the two groups were matched on other variables at the start of the school year at time point 1, there was a significant difference in individual interest in physics and chemistry between the experimental group (m = 4.17, sd = 0.79) and the control group (m = 3.07, sd = 1.10) , t(34) = 3.449, p =.002. the experimental group showed a higher interest in physics and chemistry at the beginning of the school year as compared to the control group but not at the other time points. 1 2 3 4 5 time 1 time 2 time 3 m ea n le ve l i nt er es t technology collaboration mathematics physics and chemistry biology laine et al 50 | f l r figure 2. experimental and control groups’ interest in physics and chemistry during the school year. in summary, there was no evidence of statistically significant differences between the groups in interest in technology, interest in collaboration, interest in mathematics, or interest in biology. similarly, no differences were found in the development of students’ intrinsic goal orientation, general self-efficacy, or web-user self-efficacy. the only significant difference between the two groups related to interest in physics and chemistry. the experimental group’s interest in physics and chemistry was higher at the beginning of the school year in september but declined to average level over the following four months. in contrast, the control group were a little less interested in physics and chemistry at the start of the school year but seemed to gain interest during the autumn and reached the same level as the experimental group at the midterm measurement point. however, during the spring term, the experimental group seemed better able to maintain their interest, experiencing only a slight decrease in mean value over the four months, while the control group regressed to the below-average level they had reported at the beginning of the school year some nine months earlier. 5.3. student interest generation: digital learning group during the time 1 interviews, experimental group students identified several features in the learning environment that they found to be interest-generating; these were integration, illustrativeness, learning tools, hands-on activities, ability to use the digital material outside classroom, and versatility. versatile use of learning materials was mentioned most often as interest generating feature of the environment. activities performed outside of classroom were the second often mentioned interest generating feature. table 1 shows the number of mentions for each sub-class of interest generation, an example of sub-classes, as well as interpretation. during the time 2, the students mentioned less interest generating features as seen in table 1. 1 2 3 4 5 time 1 time 2 time 3 m ea n le ve l i nt er es t digital environment ordinary environment laine et al 51 | f l r table 1 generation of interest, an example passage from the interviews and its interpretation generation of interest number of mentions for each sub-class of interest generation example an example passage from the interviews which may include conversation between several students and which counts as one expression interpretation integration time 1 = 2; time 2 = 1 student 1: well, the chemistry test or the experiment thing, where we collected water from nearby lakes. we take water samples from there and then we analyse them in chemistry class. student2: and in biology. in biology, we also do research work, and doesn’t writing the report involve something related to finnish language? during the time 1 and time 2 interviews, students mentioned integrative elements that they found interesting when participating in projects that combined different school subjects. for example, students referred to a science project in which they had collected water samples, analysed them, and wrote reports about their findings. illustrativeness time 1 = 4; time 2 = 0 student 1: especially on those occasions when we record videos to book creator. in the chemistry class, we usually record it (the experiment), and then you can attach the video there. student 2: and then you can watch it from there when you are revising for an exam. how it was done and so on; what went wrong and what could have been done better. as an example from time 1 interviews, the tablet computer was found to have proved useful when taking notes during chemistry class experiments. in the time 1 interviews, students mentioned features of the learning environment that they had found especially illustrative. however, during time 2 interviews, students did not mention any examples of illustrative features. hands-on activities time 1 = 3; time 2 = 0 student 1: everything like water hardness, ph-value, and nitrate content. then, nutrient concentrations; you put some substance in there that changes the colour; like, the sample’s colour changes, and then we can conclude if there is something. during time 1 interviews, the students mentioned that the hands-on work in the chemistry classes had been interesting. this included taking water samples from nearby lakes that were then analysed in the chemistry laboratory. outside of classroom time 1 = 5; time 2 = 3 student 1: well, there’s this home economics diary. student 2: you take a picture of the food you prepared, and then you write what was done today and what was the day’s topic. student 3: and then we can …there’s in both time 1 and 2 interviews, students mentioned that they would sometimes use the tablet computer to engage in learning activities outside the classroom. for example, in home economics, they would document their household work, which the teacher would take into account in the course evaluation. laine et al 52 | f l r this separate section where we can write all the foods that we have prepared at home and take a picture, so then the teacher can see what we have done at home. versatility time 1= 6; time 2 = 5 student 1: well, at least there are videos; in the hardcover books, there are none of those. student 2: and there are multiple choice questions. student 1: as i was saying, there are videos and a lot more pictures than in the hardcover books. student 2: and you can immediately check if the multiple-choice questions were correct. it shows you that. students also mentioned several features of the learning environment that related to the versatile use of the learning material. for example, when discussing the differences between e-books and hardcover books, they noted that e-books included more interactive features. 5.4. components of interest generation: digital learning group after identifying things they had found interesting during their studies, students in the experimental group were then asked to elaborate on why these things were interesting. in the analysis, these were categorised as components of interest. the students’ answers included features such as ease of use, working with the environment as nice or fun, being able to collaborate with other students, and having autonomy over their own learning. most often the students mentioned that the fact that the environment was easy to use was generating their interest. table 2 shows the number of mentions for each sub-classes of components of interest, examples of each sub-category, as well as interpretation. table 2 shows also number of mentions of critical remarks and developmental ideas that were asked from the students at the end of the time 2 interview. table 2 components of interest, an example passage from the interviews and its interpretation components of interest generation number of mentions for each sub-classes of components of interest example an example passage from the interviews which may include conversation between several students and which counts as one expression interpretation easy to use time 1 = 7; time 2 = 6 student 1: there are so many things in this one device that it’s completely different. you can just take this in your hand and do everything at once. with books, you take one book and do stuff there, and then you take another one and do stuff there, and… at both interview times, students felt that the mobile learning environment had made it easier for them to study. in the time point 1 interviews especially, the fact that all the materials they needed were on one device attracted positive remarks. laine et al 53 | f l r nice/fun time 1 = 6; time 2 = 4 student 1: you can do all kinds of fascinating things with it, like molecules and stuff. student 2: we haven’t used it that much yet, only in one lesson. student 1: yeah, in one lesson, but it’s very nice because you can get all these fascinating looking things. in general, students referred to the new learning environment in words and phrases conveying positive emotions. working with the material had felt nice, fun, and exciting, especially in time 1 interviews, when they were still relatively new to the learning environment. for example, in chemistry, they had used modelling software to produce different chemical bonds. collaboration time 1 = 4; time 2 = 6 student 1: it [the software] is a bit like a notebook, and the teacher writes the titles and groups there. and then, let’s say i would be lisa. i would write about a certain topic in religion, and then another group would write about some other topic, and so on. then we all read the texts produced, and the teacher also has access to them. so, we have to read what everybody else has written there. student 2: and everyone can do it with their own tablet. we can write about a topic together in the group by simultaneously using our own tablets. one of the themes that emerged from the interviews was collaboration. the students described situations in which they had worked together as a group, using the mobile learning environment to produce and share material instantly with their peers. autonomy time 1 = 4; time 2 = 6 student 1: but it’s not as if someone is breathing down your neck all the time and watching what you’re doing or that all the words are correct. we are usually just given instructions on what to do—like do exercises or write something down. student 2: yeah, or go somewhere on the internet. at both time points, the students raised issues about their autonomy as learners. in time point 2 interviews especially, they mentioned classroom activities where the teacher had allowed them to choose from different approaches and working methods. critical remarks time 2 = 16 ideas for development time 2 = 11 student 1: well, mathematics. student 2: it is a bit far-fetched how you are supposed to use it there. basically it works pretty fine, but at least personally i don’t feel like i’m learning a lot with it. student 1: and in principle if you would like to have all the (mathematical) symbols you would have to buy a keyboard. like the kind you have in a computer. in the time 2 interviews, the students were also asked to freely criticise features and practices in the environment that they felt were not working properly, and to suggest improvements for future development. their critical remarks related to poor quality e-books (especially in mathematics), problems with ipad functioning and wireless connection, the lack of a good application for taking laine et al 54 | f l r exams, teachers’ reluctance to take full advantage of the digital tools, and concern about losing hand-eye coordination because of working only digitally. 6. discussion this study investigated how students’ interest in the stem domain developed over a one year period in an inquiry-based mobile learning environment, and whether their developmental patterns differed from those in the more traditional learning group. in addition, we hoped that by interviewing the students about their experiences we might acquire more in-depth information about the components that contribute to longterm interest generation and development in such environments. the statistical analyses showed that the digital learning group’s students’ domain specific interests in science and mathematics remained at an average level throughout the school year while their interest in technology and collaboration remained high. these results seem to indicate that students found the new learning environment interest generating. this aligns with previous research suggesting that inquiry-based learning environments trigger and sustain students’ interest in stem (renninger et al., 2014). however, perhaps the most surprising result was that the digital learning group’s students’ interest in physics and chemistry declined significantly during the first half of the school year. this finding seems to be in conflict with some of the interview results, which highlight how the new learning environment was seen as interest supporting by the students. that decline was also the only observed statistical difference between the two groups with the digital learning group’s developmental pattern differing from that of the control group. while the experimental group’s interest declined slightly from high to above average during the autumn semester, the control group’s interest increased to similar level during the same period. this raises the question of whether the new learning environment had been implemented to make the best use of the principles of inquiry learning and to take into account the pedagogical demands of mobile learning. previous research has found that the manner of implementation in the inquiry context is a key element of inquirybased learning interventions (renninger et al., 2014). digital environments do not in themselves necessarily suffice to generate interest; instead, a more comprehensive design approach is needed. it seems, for instance, that a number of questions should be addressed before implementation, including how to arrange instructional support, when to allow collaboration between students, and how best to take advantage of the mobility of the learning environment. furthermore it would be beneficial to apply a broader view of the learning context when mobile learning designs are researched and implemented in practice. frameworks such as integrative learning design framework (bannan, 2016) could be used to systematically integrate analysis, design and development processes that would serve the needs of both pedagogical development, and producing an effective learning innovation. the interviews revealed that, in general, the students found the mobile learning environment to be interest-generating, with particular reference to the following specific features: 1) the possibility of integrating different school subjects; 2) the ease of use of the learning software; 3) the ability to take learning outside the classroom; and 4) the increased possibilities for learning collaboratively. students referred to all four features on both occasions which indicates that there was no sign of a novelty effect wearing off during the one year period. laine et al 55 | f l r the new curriculum was designed to emphasise joint teaching of different school subjects with the aim of increasing student interest and engagement in studying. it was therefore very positive to see that students were referring to the possibility of integrating different school subjects as interest generating. therefore it can be argued that while the new curriculum was especially emphasizing integration the new learning environment had from this point of view at least partly reached its goal. the way that students were referring to the second feature, ease of use was also promising. the analyses of the interviews revealed that students placed considerable emphasis on how the digital tools work and the scope they offer for individual exploration. already in the first interviews, students described several items of learning software they had learned to use in school, including e-books, web-based learning resources, mind maps, and blogs. in some learning activities or projects, students were allowed to choose how to complete the work, using their creativity to modify the end product to their liking and making it more personally relevant by retrieving information from different sources. they also expressed a belief that learning to use a wide range of digital tools in the environment prepared them for success in their future working life; hence they had developed value beliefs towards the technology used in the environment. this indicates that at least some of the students had already progressed towards more individual interest, since its underlying psychological processes are closely related to interpreting and valuing contextual features in relation to potential activities (ainley, 2012). use of mobile learning environments may offer students more autonomy and freedom of expression. some students emphasised the importance of being able to modify their digital notebooks as they saw fit. some had also used the tablet to create their own animated videos and to undertake other creative work in their free time. this seems to have helped them to adopt the tablet as an everyday instrument for learning, increasing their level of engagement with school work. mobile devices offer the possibility of a more personalized approach to learning and stretch the boundaries of the physical classroom. planning the balance of instructional structure and student autonomy is an important part of every learning process, and combining mobile learning with inquiry lends further emphasis to this aspect. the critical comments made by the students regarding their day-to-day schoolwork were indicating that this was currently lacking in some cases. previous literature has found that un-assisted discovery, or inquiry learning is not in any way an optimal solution for educative purposes (see e.g. alfieri et al., 2011) and scaffolds for self-regulation as well as students’ own reflection are needed (pedaste et al., 2012). in the future, technology will probably offer even more personalized learning paths, and this will further increase the demand for teachers to develop their competences and also re-consider their role in the learning process. the mobile environment also made it possible to take learning outside the classroom, most notably in the case of the multidisciplinary water-related learning project. using the tablet computer, students were able to take notes and document the water samples they had collected. they could then work collaboratively on writing the final report, sharing it with other members of their group before displaying the results to a wider audience. although these features have in many cases already formed part of inquiry science instruction (see minner et al., 2010), mobile technology helped to make the process more effortless and versatile. although this study did not directly address the issue of whether interest influences learning outcomes, the interviews raised some interesting notions. some students felt that the activities completed in the mobile learning environment helped them to memorise the material better than when reading from books. their reflections seemed to indicate their ability to use metacognitive skills to assess their learning strategies, and suggested their preference for the new environment. laine et al 56 | f l r 6.1 critical remarks on the learning environment students also made some critical remarks and identified issues related to the learning environment that need to be resolved. these remarks can be divided into two broad categories: 1) pedagogical shortcomings in implementing the environment and 2) technological challenges. in relation to pedagogy, students noted that teachers were not always up to speed with the new technology sometimes the students even needed to guide them on technical issues. this also gave rise to students’ criticism of teaching methods in some of the lessons where according to the students, the teacher had merely replicated ordinary classes with books, notebooks and handouts, utilizing the tablet computer and its applications in a minimal way. these implementation problems may also explain why the statistical analysis did not reveal greater differences in interest levels over time or between the two groups. this criticism brings out the challenges that new mobile learning environments present for teaching. as traxler & kukulska-hulme (2016) point out, in mobile learning the teacher’s role and responsibilities shift from being the facilitator of educational artefacts, towards being one who evaluates and collects suitable content, organizes it and makes educational opportunities possible. in order to succeed in this task the teachers need sufficient support on how to prepare for this change. this underlines the importance of in-service teacher training programs that would offer teachers the possibilities to get an update on their pedagogical skills and knowledge. attempting to use ‘ordinary’ pedagogy in a mobile environment is likely to hinder student engagement in learning and to place an additional burden on the teacher. this kind of criticism also suggests that students are quite well aware of the technological aspects of the learning environment and are also able to reflect on how they learn in different settings. in the future students could be more involved already in the design process when these types of inquiry learning activities are planned. in this way, assignments could focus more on students’ own areas of interests and needs, and challenging or even negative features of their experiences could be addressed and resolved during the process. this would also support students’ selfregulation development which is a crucial element in these types of learning environments. previous findings state that while the mobile learning environment offers tools for ubiquitous learning, students need skills and knowledge of self-regulated learning if they are to realise all the environment’s possibilities (sha, looi, chen & zhang, 2012). another criticism was directed at the technology used. sometimes, the hardware caused problems, such as loss of tablet power, failure to connect to the internet, or the laboriousness of writing longer texts without an external keyboard. additionally, the educational software was either too limited—for instance, allowing only multiple-choice questions—or was still under development, which meant that some features were not yet working. teachers therefore need to be aware of the attributes of each software or application and how to apply these in their teaching. although essential in day-to-day classroom practice, this is not an easy task, and without familiarizing themselves, teachers will be unable to decide the best ways to apply the technology. 6.2 limitations the present study has some limitations that need to be addressed. the first of these concerns the relatively small sample size on which the quantitative analyses were based. because the study was conducted in ecological valid settings, the researchers could not determine the number of participants; instead, the school principal decided who was to be assigned to the mobile learning class. while this may have created bias in the sample, this was controlled by creating a propensity score-matched control group from a pool of 113 students. the propensity scoring matched each of the experimental group’s students with a matching student from the control group, based on their motivational profiles at the start of the school year. a second limitation relates to the measures used for quantitative data collection. students’ interest in stem was measured in terms of school subjects. although single-item instruments have been extensively laine et al 57 | f l r used to measure interest in previous research, it can be argued that using single-item instruments measuring only interest in the subject as a whole was too broad. it would be useful in future studies to consider how to reliably measure interest in its different developmental phases (renninger & hidi, 2011). another measure-related issue is that interest in physics and chemistry was assessed by one combined item and may therefore have caused some reliability issues in terms of the item’s internal consistency. although the calculated reliability estimates were at a sufficient level, some students completing the questionnaire may actually have been estimating their interest in physics while others were concentrating on chemistry. in addition, there was no measure of interest in geography, which is usually included in the stem domain and is taught as a separate subject in the finnish school system. 6.3 recommendations for future research previous research has shown that individual interest affects which features of the learning environment help to trigger students’ situational interest (renninger & hidi, 2011). an essential question for future research, then, is how best to combine measures of situational interest and long-term individual interest, and what other personal and motivational variables should be taken into account in this process. digital devices such as tablet computers facilitate more effortless and less intrusive data collection that can tap into students’ everyday experiences at school. another recommendation relates to the actual day-to-day classroom work in the course of this study. without observational data, it is impossible to determine exactly how much of the teaching in these new learning environments actually relies on inquiry-based approaches. based on the interview data of this study, it seems that at least partly the day-to-day teaching was relying neither on inquiry nor taking the full advantage of the mobile learning aspect of the learning environment. in the future, it would be interesting to add an observational element to data collection to determine, for example, how students’ inquiry skills develop during an inquiry learning activity, and whether or not pedagogies evolve during that time. in addition a measure for learning outcomes could be added. a final recommendation for future research would be to explore how to design mobile learning environments to properly enhance inquiry learning. mobile learning devices offer freedom from temporal and spatial boundaries, and digital evaluation tools may in the future provide personalised information about students’ progress and level of development. keypoints there was clear evidence that the inquiry-based mobile learning environment offered promise in terms of supporting students’ interest in learning science but this did not come without a price. more research is needed to determine the best practices of pedagogical implementation. the pedagogy used during lessons should be integrated with the educational technology, so that they complement each other. this makes it possible to support students’ interest development in the long term. students could be more involved in the implementation process by having their say about the design of the environment. students would benefit from increased autonomy in terms of selfregulation and engagement and teachers would receive valuable feedback on the functionality of the environment. teachers should be offered the possibility to attend in-service training programs that help them to update their pedagogies, and also constant support should be available. laine et al 58 | f l r 7. acknowledgements this study received funding from the finnish cultural foundation and turku university foundation. references ainley, m. (2012). students’ interest and engagement in classroom activities. in s. l. christenson, a. l. reschly, & c. wylie (eds.), handbook of research on student engagement (pp. 283-302). doi:10.1007/978-1-4614-2018-7_13 ainley, m., & ainley, j. (2015). early science learning experiences: triggered and maintained interest. in k. a. renninger, m. nieswandt, s. hidi (eds.), interest in mathematics and science learning (pp. 1733). washington: aera. doi:10.3102/978-0-935302-42-4 alfieri, l., brooks, p. j., aldrich, n. j., & tenenbaum, h. r. (2011). does discovery-based instruction enhance learning? journal of educational psychology, 103(1), 1-18. doi:10.1037/a0021017. ames, c. (1992). classrooms: goals, structures, and student motivation. journal of educational psychology, 84(3) 261-271. doi:10.1037/0022-0663.84.3.261 bannan, b. (2016). analysing context for mobile augmented reality prototypes in education. in j. traxler, & a. kukulska-hulme (eds.), mobile learning: the next generation. (pp. 115-139). london: routledge. doi:10.4324/9780203076095 bong, m., lee, s. k., & woo, y-k. (2015). the roles of interest and self-efficacy in the decision to pursue mathematics and science. in k. a. renninger, m. nieswandt, s. hidi (eds.), interest in mathematics and science learning (pp. 189-202). washington: aera. doi:10.3102/978-0-935302-42-4 crowley, k., barron, b.j., knutson, k., & martin, c. (2015). interest and the development of pathways to science. in k. a. renninger, m. nieswandt, s. hidi (eds.), interest in mathematics and science learning (pp. 297-314). washington: aera. doi:10.3102/978-0-935302-42-4 de jong, t. (2006). computer simulations: technological advances in inquiry learning. science, 312, 532533. doi:10.1126/science.1127750 drake, s. m., & burns, r. c. (2004). meeting standards through integrated curriculum. alexandria, va: association for supervision and curriculum development. eachus, p., & cassidy, s. (2006). development of the web users self-efficacy scale (wuse), issues in informing science and information technology journal, 3, 199-209. retrieved from: https://www.informingscience.org/publications/883?search=development%20of%20the%20web%20 users%20self-efficacy%20scale finnish national board of education (fnbe). (2014). national core curriculum for basic education 2014. helsinki: national board of education. retrieved from http://www.oph.fi/ops2016 [in finnish]. fryer, l. k., ainley, m., & thompson, a. (2016). modelling the links between students' interest in a domain, the tasks they experience and their interest in a course: isn't interest what university is all about?. learning and individual differences, 50, 157-165. doi: 10.1016/j.lindif.2016.08.011 furtak, e.m, seidel, t., iverson, h., briggs, d. (2012). experimental and quasi-experimental studies of inquiry-based science teaching: a meta-analysis. review of educational research, 82(3), 300-329. doi: 10.3102/0034654312457206 glynn, s. m., bryan, r. r., brickman, p., & armstrong, n. (2015). intrinsic motivation, self-efficacy, and interest in science. in k. a. renninger, m. nieswandt, s. hidi (eds.), interest in mathematics and science learning (pp. 189-202). washington: aera. doi:10.3102/978-0-935302-42-4 harackiewicz, j. m., durik, a. m., barron, k. e., linnenbrink-garcia, l., & tauer, j. m. (2008). the role of achievement goals in the development of interest: reciprocal relations between achievement goals, laine et al 59 | f l r interest, and performance. journal of educational psychology, 100(1), 105-122. doi: 10.1037/00220663.100.1.105 harackiewicz, j. m., smith, j. l., priniski, s. j. (2016). the importance of promoting interest in education. policy insights from the behavioral and brain sciences, 3(2), 220-227. doi: 10.1177/2372732216655542 hidi, s. & renninger, a. (2006). the four-phase model of interest development. educational psychologist, 41(2), 111-127. doi:10.1207/s15326985ep4102_4 knogler, m., harackiewicz, j. m., gegenfurtner, a., & lewalter, d. (2015). how situational is situational interest? investigating the longitudinal structure of situational interest. contemporary educational psychology, 43, 39-50. doi:10.1016/j.cedpsych.2015.08.004 krapp, a., & prenzel, m. (2011). research on interest in science: theories, methods, and findings. international journal of science education, 33, 27-50. doi:10.1080/09500693.2010.518645 lai, c. l., hwang, g. j., liang, j. c., & tsai, c.-c. (2016). differences between mobile learning environmental preferences of high school teachers and students in taiwan: a structural equation model analysis. educational technology research and development, 64(3), 533-554. doi:10.1007/s11423016-9432-y martin, m. o., mullis, i. v. s., foy, p., & hooper, m. (2016). timss 2015 international results in science. chestnut hill, ma: timss & pirls international study center, boston college. retrieved from: http://timssandpirls.bc.edu/timss2015/international-results/ minner, d.d., levy, a.j. & century, j. (2010). inquiry-based science instruction what is it and does it matter? results from a research synthesis years 1984 to 2002. j. research in science teaching, 47, 474–496. doi:10.1002/tea.20347 organisation for economic co-operation and development. (2006). evolution of student interest in science and technology studies. policy report. global science forum. retrieved from: www.oecd.org/science/sci-tech/ palmer, d. h. (2009). student interest generated during an inquiry skills lesson. journal of research in science teaching, 46, 147-165. doi:10.1002/tea.20263 pedaste, m., mäeots, m., leijen, ä., & sarapuu, s. (2012). improving students' inquiry skills through reflection and self-regulation scaffolds. technology, instruction, cognition and learning, 9, 81–95 pedaste, m., mäeots, m., siiman l. a., de jong, t., van riesen, s. a. n., kamp, e. t., manoli, c. c., zacharia, z. c. & tsourlidaki, e. (2015). phases of inquiry-based learning: definitions and the inquiry cycle. educational research review, 14, 47–61. doi: 10.1016/j.edurev.2015.02.003 pintrich, p., smith, d., garcia, t. & mckeachie, w. (1991). a manual for the use of the motivated strategies for learning questionnaire (mslq). ann arbor, mi: national center for research to improve postsecondary teaching and learning. retrieved from: https://archive.org/details/eric_ed338122 potvin, p., & hasni, a. (2014). interest, motivation and attitude towards science and technology at k-12 levels: a systematic review of 12 years of educational research. studies in science education, 50(1), 85129. doi:10.1080/03057267.2014.881626 renninger, k. a., & hidi, s. (2011). revisiting the conceptualization, measurement, and generation of interest. educational psychologist, 46, 168–184. doi:10.1080/00461520.2011.587723 doi: 10.1080/03057267.2014.881626 renninger, k. a., austin, l., bachrach, j. e., chau, a., emmerson, m.s., king, b. r., riley, k. r., stevens, s. j. (2014). going beyond whoa! that’s cool! achieving science interest and learning with the ican intervention. in s. karabenick & t. urdan (eds.), motivation-based learning interventions: advances in motivation and achievement series (vol. 18, 107–140). doi:10.1108/s0749-742320140000018003 rocard, m., csermely, p., jorde, d., lenzen, d., walberg-henriksson, h., & hemmo, v. (2007). science education now: a renewed pedagogy for the future of europe. retrieved from: http://ec.europa.eu/research/swafs/index.cfm?pg=library&lib=science_edu laine et al 60 | f l r rottinghaus, j. p., larson, l. m., & borgen, f. h. (2003) the relation of self-efficacy and interests: a metaanalysis of 60 samples. journal of vocational behavior, 62, 221-236. doi:10.1016/s00018791(02)00039-8 sha, l., looi, c.-k., chen, w. and zhang, b.h. (2012), understanding mobile learning from the perspective of self-regulated learning. journal of computer assisted learning, 28, 366–378. doi:10.1111/j.13652729.2011.00461.x tapola, a., veermans, m., & niemivirta, m. (2013). predictors and outcomes of situational interest during a science learning task. instructional science, 41(6), 1047–1064. doi:10.1007/s11251-013-9273-6 traxler, j., & kukulska-hulme, a. (2016). introduction to the next generation of mobile learning. in j. traxler, & a. kukulska-hulme (eds.), mobile learning: the next generation. (pp. 1-10). london: routledge. doi:10.4324/9780203076095 vansteenkiste, m., lens, w., & deci, e. (2006) intrinsic versus extrinsic goal contents in selfdetermination theory: another look at the quality of academic motivation. educational psychologist 41(1), 19-31. doi:10.1207/s15326985ep4101_4 microsoft word schindler et al publication.docx frontline learning research vol. 5 no. 4 (2017) 76-88 issn 2295-3159 mood moderates the effect of self-generation during learning julia schindler1, tobias richter1, & carolin eyßer2 1university of würzburg, germany 2university of regensburg, germany article received 09 march 2017 / revised 16 november / accepted 21 november / available online 18 december abstract generating information, compared to reading, improves learning and enhances long-term retention of the learned content. this so-called generation effect has been demonstrated repeatedly for recall and recognition of single words. however, before adopting generating as a learning strategy in educational contexts, conditions moderating the effect need to be identified. this study investigated the impact of positive and negative mood states on the generation effect with short expository texts. according to the dual-force framework (fiedler, nickel, asbeck, & pagel, 2003), positive mood should facilitate generation by enhancing creative knowledge-based top-down processing (assimilation). negative mood, however, should facilitate learning in the read-condition by enhancing critical stimulus-driven bottom-up processing (accommodation). in contrast to our expectations, we found no general generation effect but an overall learning advantage of read compared to generated texts. however, a significant interaction of learning condition and mood indicates that learners in a better mood recall generated texts better than learners in a more negative mood, whereas no mood effect was found when the texts were read. the results of the present study partially support the predictions of the dual-force framework and are discussed in the context of recent theoretical approaches to the generation effect. keywords: generation effect, learning with expository texts, mood states info corresponding author email: julia.schindler@uni-wuerzburg.de doi: http://dx.doi.org/10.14786/flr.v5i4.296 schindler et al 77 | f l r 1. introduction a common assumption is that learning is most effective when it is easy. however, research suggests that under certain conditions learning is more effective when learners intentionally make it more difficult by, for example, distributing learning sessions, interleaving topics and tasks, testing learned content, and generating knowledge (bjork & bjork, 2011). the advantage of generated compared to read information in memory tasks (generation effect) has been investigated extensively (for a meta-analysis, see bertsch, pesta, wiscott, & mcdaniel, 2007). in the classical generation paradigm (e.g., mcdaniel, waddil, & einstein, 1988; slamecka & graf, 1978) participants read an associated word pair (purr – cat) or they complete a fragment of the target word (purr – c_ _). in a subsequent learning test, memory for generated target words is better than for read words. findings like this suggest that using generation in every-day learning situations might be beneficial. however, before adopting generation as a learning strategy in educational contexts, conditions moderating the generation effect such as learners’ cognitive abilities, motivation, or mood states need further clarification. research on the impact of mood on information processing has demonstrated that positive and negative mood states serve a regulatory function in terms of processing depth, processing capacity, and processing strategies (for an overview see bless & fiedler, 2005; fiedler, 2001). fiedler assumes that these mood states trigger one of two possible learning settings that should affect generation differently. the learning settings and their consequences for the efficiency of generation are described in fiedler, nickel, asbeck, and pagel’s (2003) dual-force framework. the framework is based on fiedler’s assumptions that (1) cognitive processes usually contain the two components of conservation of perceived information (accommodation) and the generation of new information based on internal knowledge structures (assimilation) and (2) positive and negative mood states trigger two different learning settings that encourage either assimilation or accommodation. positive mood is assumed to trigger an appetitive learning setting that encourages exploration and creative and elaborative top-down processes. in other words, positive mood should encourage assimilative processes as required in generating activities. negative mood, however, is assumed to trigger an aversive learning setting that encourages critical stimulus-driven processes. thus, negative mood should encourage accommodative bottom-up processes as primarily required in reading (see fiedler, 2001). fiedler et al. (2003) conducted three experiments to test some of the framework’s implications, but they primarily addressed mood-congruency effects in word learning (words congruent with a learner’s mood are recalled better than non-congruent words). hence, their results provide only little insight in the predicted mood-generation interaction with neutral verbal material such as expository texts, which are usually used in educational contexts. in an earlier study, fiedler, lachnit, fay, and krug (1992, exp. 4) found that positive mood compared to neutral mood enhanced the generation effect for neutral word pairs as predicted by the dualforce framework, but a negative mood state was not induced. the aim of the present study was to further investigate the impact of positive and negative mood states on the generation effect with more complex and naturalistic learning material. 2. the present study slamecka and graf’s (1978) fragment-completion paradigm was adopted for short naturalistic texts from a psychology text book (mazur, 2006). participants read either complete or generated fragmented definitions containing a concept and its description (e.g., spontanerholung = das wie_erau_tre_en einer zu_or ge_ösc_te_ rea_tio_, nachdem lä_ge_e ze_t ohne wei_ere kon_itio_ie_un_sdu_ch_än_e ver_tri_he_ ist/ spontaneous recovery = recurrence of an extinguished reaction, after some time has passed without further conditioning) after receiving a positive or negative mood induction. we used short definitions instead of one longer text to be able to generalize possible findings across texts while keeping the learning phase after the schindler et al 78 | f l r mood induction as short as possible. also, short definitions such as those used in the present study are typical every-day learning contents in school and university. hence, we assume that these minimalistic texts are suitable to test the dual-force framework for naturalistic learning material in a first attempt. fragment completion was used because it has been found to induce the generation effect quite constantly in several studies with target words as learning material (see the meta-analysis by bertsch et al., 2007) and to improve memory for fairy tales and narratives (e.g., einstein, mcdaniel, bowers, & stevens, 1984; mcdaniel, 1984; mcdaniel & kerwin, 1987). we expected learners to use the sentence context and, if available, prior knowledge to infer the fragmented words in the generation condition, thereby encouraging the construction of an elaborated mental representation of the texts. we expected to find a generation effect (better recall for generated compared to read definitions). moreover, based on the dual-force framework, positive mood was expected to enhance recall in the generation condition, whereas negative mood was expected to enhance recall in the read condition. 3. methodology participants were 55 undergraduates (48 psychology students) from the university of kassel and one non-student participant (all native speakers of german). ten participants were male and 46 were female with a mean age of 21.96 (sd = 3.46; min = 18; max = 36). all participants provided written informed consent. in a first step, participants’ prior knowledge on the to-be-learned topic was assessed via a computerized 20-item classification task. they were presented with 10 concepts from the field of learning and behavior (e.g., post-reinforcement pause) and with 10 pseudo-concepts (e.g., comparative reinforcement). participants were asked to indicate for each concept whether it belongs to the field of learning and behavior or not. positive mood (n = 28) or negative mood (n = 28) was induced via the computerized mood-induction procedure by robinson, grillon, and sahakian (2012). participants listened to happy or sad music via headphones. at the same time, they were presented with 60 emotionally charged statements from velten (1968) (e.g., happy: i often feel great and motivated; sad: i often feel sad and depressed) and were asked to read them as referring to themselves. after the mood induction, participants read or generated 32 psychological definitions (varied within subjects) in two 15 min sessions with 16 definitions each (paper and pencil format). the order of learning condition (generate-read vs. read-generate) was counterbalanced across participants. after the second learning session, participants completed a sociodemographic questionnaire (e.g., age, sex, native language, highest level of education). in the subsequent test phase, participants were asked to provide the description for each concept. at the end of the experiment, participants in the negative mood-induction group received an additional positive mood induction, and participants in the positive mood-induction group received an additional neutral mood induction. self-reported mood was assessed via three visual analogue scales (how happy are you?/how sad are you?/how depressed are you?) ranging from 1 to 100 at six measuring times during the experiment (time 1: before the first mood induction; time 2: after the first mood induction and prior to learning phase 1; time 3: prior to learning phase 2; time 4: prior to the test phase; time 5: after the test phase; time 6: after the final mood induction). schindler et al 79 | f l r 4. results self-reported mood measures from the sad and depressed-mood scales were inverted so that higher values indicated a better mood. participants’ answers in the learning test were coded by two independent raters. the maximum scores ranged from 2 to 3 points. for each definition, the proportion of accurate recall was calculated (i.e., achieved score divided by maximum score). considering that inter-rater reliability was very high for proportion of accurate recall (icc(just) = .95), we combined the ratings into one measure of recall accuracy (ranging from 0 to 1 with 1 indicating maximum score achievement). for mood measures (manipulation check) and rated accuracy of recall, we estimated linear mixed models (lmm, baayen, davidson, & bates, 2008). all models were estimated and tested with the software packages lme4 (bates et al., 2014) and lmertest for r (kutznetsova, brockhoff, & christensen, 2014). all significance tests were based on a type-i error probability of .05. one-tailed tests have been used for all directional hypothesis. mood measures (manipulation check). six linear mixed models were estimated for each mood measure (happy, sad, depressed) as dependent variables to derive the simple main effects for mood induction at the six times of measurement during the experiment. mood induction was included as a contrast-coded fixed effect (-1 = negative mood, 1 = positive mood) and measuring time in the form of five dummy-coded fixed effects (with time 1 as the reference category). intercepts were allowed to vary randomly between participants (random effects of participants). as expected, for the happy mood measure, the simple main effects for mood induction were significant at time 2 (b = 23.04; t(44) = 9.50; p <.001), time 3 (b = 7.13; t(44) = 2.94; p <.01), time 4 (b = 8.70; t(44) = 3.59; p <.001) and time 5 (b = 7.39; t(44) = 3.05; p <.01) indicating that participants’ self-reported mood was significantly better in the positive than in the negative mood-induction group at all measuring times between the first and the final mood induction. at time 6, participants self-reported mood was significantly lower in the positive mood-induction group after the final neutral mood induction compared to the negative mood-induction group, which had just received a final positive mood induction (b = -8.70; t(44) = -3.59; p <.001). for the inverted sad mood measure, the simple main effects for mood induction were significant at time 2 (b = 18.79; t(44) = 6.78; p <.001), time 3 (b = 8.63; t(44) = 3.11; p <.01), time 4 (b = 10.00; t(44) = 3.61; p <.001), and time 5 (b = 7.38; t(44) = 2.66; p <.01), and for the inverted depressed mood measure at time 2 (b = 18.86; t(44) = 5.98; p <.001), time 3 (b = 7.59; t(44) = 2.41; p <.05), time 4 (b = 10.36; t(44) = 3.28; p <.01) and time 5 (b = 5.63; t(44) = 1.78; p <.05; one-tailed). both measures indicate that participants’ self-reported mood was better in the positive than in the negative mood-induction group between time 2 and time 5. figure 1 illustrates the differences in self-reported mood between both mood-induction groups at the six times of measurement for the three mood measures. given the significant differences between both moodinduction groups at time 2 to time 5 for all three measures, we consider the mood induction to have been successful. schindler et al 80 | f l r a) b) c) figure 1. mood measures for positive and negative mood-induction groups: differences between moodinduction groups estimated at six measuring times for happy mood (a), inverted sad mood (b), and inverted depressed mood (c). rated recall accuracy. a linear mixed model was estimated for rated recall accuracy as dependent variable. learning condition was included as a contrast-coded fixed effect (-1 = read, 1 = generated). mood measures were combined for the three mood scales to include self-reported mood as a single fixed-effect predictor into the model. the learners’ mood prior to learning phase 1 (time 2) was used as a predictor for recall of definitions from learning phase 1, and mood prior to learning phase 2 (time 3) was used as a predictor for definitions from learning phase 2 (centered around mean mood at time 1 prior to first mood induction). in order to test whether it is adequate to integrate the three mood items at times 2 und 3 into an integrated mood scale, we used confirmatory factor analysis (software package lavaan for r, rosseel, 2012) with maximum likelihood to estimate and test a measurement model with one latent variable (mood) for each measuring time. the factor loadings for the same items at the two measuring times were set equal. in addition, the errors of pairs of the same items at the two times of measurement were allowed to covary, with the covariances of the three pairs of items set equal. this measurement model showed an acceptable model fit, c2= 13.44 (n = 56, df = 9), p = 0.078, cfi = .98, tli = .96, rmsea = .094 (90% confidence interval: .00 .19). based on these results, it seems adequate to form an integrated mood scale. to control for effects of prior knowledge, participants’ proportion of correct responses in the priorknowledge test was included as a grand-mean centered fixed effect into the model. we expected better test performance for learners with high compared to low prior knowledge. to account for the different retention intervals of stimuli presented in learning phases 1 and 2, learning phase was included as another contrastcoded fixed effect (-1 = learning phase 1, 1 = learning phase 2). we expected better recall for definitions from learning phase 2 than from learning phase 1 due to the longer retention interval for definitions from learning phase 1. finally, interaction terms between predictor variables were included in the model. the intercept was allowed to vary randomly between participants and definitions (random effects of participants and items). linear mixed model analysis revealed a significant main effect for learning phase (b = 0.02; t(40) = 2.45; p <.01, one-tailed) indicating that definitions from learning phase 2 were recalled better than definitions from learning phase 1. moreover, the analysis revealed a significant main effect for prior knowledge (b = 0.07; t(40) = 2.11; p <.05, one-tailed), which was further characterized by a significant two-way interaction of learning condition and prior knowledge (see below). in contrast to our expectations, we found no generation effect but a learning advantage for read compared to generated definitions (b = -0.04; t(40) = -5.09; p <.001), which was moderated by prior knowledge (b = 0.02; t(40) = 2.42; p <.05) and mood (b = 0.001; t(40)=1.85; p schindler et al 81 | f l r <.05, one-tailed). the simple slope of prior knowledge was significant for generated but not for read definitions. the higher the learners’ prior knowledge, the better they recalled generated definitions (figure 2). the effect of mood was also significant for generated but not for read definitions. as expected, learners in a better mood recalled generated definitions better than learners in a more negative mood (figure 3). figure 2. estimated recall accuracy (proportions) for generated and read definitions: simple slopes for prior knowledge and differences between learning conditions estimated at three different levels of prior knowledge. figure 3. estimated recall accuracy (proportions) for generated and read definitions: simple slopes for mood and differences between learning conditions estimated at three different levels of mood. 5. discussion the aim of the present study was to test the predictions derived from fiedler et al.’s (2003) dual-force framework that positive mood benefits learning with generated texts, whereas negative mood benefits learning with read material. according to the framework, positive mood triggers an appetitive learning setting that encourages assimilative, i.e. creative and elaborate, top-down processes, which are central to generation. consistent with this assumption, we found that learners in a better mood recalled generated definitions better than learners in a more negative mood. the finding that learners with higher prior knowledge recalled generated (but not read) definitions better than learners with lower prior knowledge further supports fiedler et al.’s assumption that generation requires assimilative knowledge-based elaboration. schindler et al 82 | f l r moreover, the dual-force framework assumes that negative mood triggers an aversive learning setting that encourages accommodative bottom-up processes, which are central to reading. however, we did not find negative mood to improve learning in the read condition. a possible explanation might be that most participants (fortunately) seem to have had a rather positive prevailing mood. even after the participants’ mood in the negative mood-induction group decreased significantly at times 2 and 3 under the prevailing positive mood level, it increased to a more neutral mood level at time 3. thus, the negative mood induction, although successful and persistent, might have been too weak to ensure a constantly aversive learning setting. another, more likely explanation for the lack of negative mood influence, however, is that text comprehension (in contrast to single word reading) always requires certain knowledge-driven assimilative processes such as drawing inferences or establishing relations of coherence in order to establish a rich and coherent mental model of the text (graesser, millis, & zwaan, 1997; van dijk & kintsch, 1983). consequently, it is not surprising that reading coherent texts benefits less from a negative mood state than reading single words. given the different cognitive processes involved in word and text reading, differential implications should be derived from the dual-force framework for learning with words and more naturalistic texts. an unexpected finding was the absence of an overall generation effect. in contrast to extant research demonstrating the beneficial effect of fragment completion on text memory (einstein et al., 1984; mcdaniel, 1984), learners in our study recalled read definitions better than generated definitions. a likely explanation is provided by mcdaniel and butler’s (2010) contextual framework. they assume that learning difficulties such as generation are only desirable when they stimulate cognitive processing that is not stimulated by the material itself. they further assume that descriptive texts primarily encourage concept-specific processing at the word and proposition level (for a detailed explanation see mcdaniel & einstein, 2005). therefore, letter deletion, which is also assumed to encourage concept-specific processing, benefits memory for descriptive texts less than a task that stimulates relational processing (between propositions) such as reordering scrambled sentences (mcdaniel, einstein, dunay, & cobb, 1986). according to the contextual framework, a relational processing task might have evoked the expected generation effect with the learning material used in our study. the results of the present study need to be interpreted with its limitations in mind. first, we used a rather specific sample (primarily psychology students). the dual-force framework, however, focuses on basal learning processes and, therefore, does not differentiate between different learner groups. thus, we assume that our sample was as good as any to start with. moreover, we used just one type of generation task and text material. replicating the present results with different samples, generation tasks, and learning material would further corroborate our findings and conclusions on the moderating role of mood states for generation. this seems even more important given the absence of the expected overall generation effect. considering that learning outcome is always the result of a complex interaction of learner characteristics, learning material, generation task, and criterial task requirements (contextual framework, mcdaniel & butler, 2010), farreaching conclusions can only be drawn with caution from a single study. however, our study was the first attempt to test the dual-force framework for naturalistic texts and, thus, can be seen as a starting point for future research. despite these limitations, our findings have important implications for developing activities in applied educational contexts. given that learners in a better mood benefit more from generation than learners in a bad mood, students who are in a bad mood momentarily or who suffer from constantly bad mood states such as depression or anxiety disorders might be at a disadvantage when it comes to generating activities in classroom settings. in contrast, our findings suggest that games or exercises that enhance the students’ mood might be used systematically to enhance the benefits students gain from generation. schindler et al 83 | f l r keypoints generation effect learning with expository texts mood states acknowledgments the research presented in this article was supported by the federal state of hessen and its loewe research initiative desirable difficulties in learning (loewe: landes-offensive zur entwicklung wissenschaftlichökonomischer exzellenz [state offensive for the development of scientific and economic excellence]). we thank our student assistants for assisting in coding the participants’ answers. the publication of this article was funded by the german research foundation (dfg) and the university of würzburg in the funding programme open access publishing. references baayen, r. h., davidson, d. j., & bates, d. m. (2008). mixed effects modeling with crossed random effects for subjects and items. journal of memory and language, 59, 390–412. doi:10.1016/j.jml.2007.12.005 bates, d., maechler, m., bolker, b., walker, s., christensen, r. h. b., & sigmann, h. (2014). lme4: linear mixed-effects models using eigen and s4 [software]. r-package version 1.1-6. retrieved from: http://cran.r-project.org/package=lme4 bertsch, s., pesta, b. j., wiscott, r., & mcdaniel, m. a. (2007). the generation effect: a meta-analytic review. memory & cognition, 35, 201–210. doi:10.3758/bf03193441 bjork, e. l., & bjork, r. a. (2011). making things hard on yourself, but in a good way: creating desirable difficulties to enhance learning. in m. a. gernsbacher, r. w. pew, l. m. hough, & j. r. pomerantz (eds.), psychology and the real world: essays illustrating fundamental contributions to society (pp. 56– 64). new york: worth. bless, h., & fiedler, k. (2005). mood and the regulation of information processing and behavior. in j. p. forgas (ed.), affect in social thinking and behavior (pp. 65–84). new york: psychology press. einstein, g. o., mcdaniel, m. a., bowers, c. a., & stevens, d. t. (1984). memory for prose: the influence of relational and proposition-specific processing. journal of experimental psychology: learning, memory, and cognition, 10, 133–143. doi:10.1037/0278-7393.10.1.133 fiedler, k. (2001). affective states trigger processes of assimilation and accommodation. in l. l. martin & g. l. clore (eds.), theories of mood and cognition: a user’s guidebook (pp. 85–98). mahwah, nj: erlbaum. fiedler, k., lachnit, h., fay, d., & krug, c. (1992). mobilization of cognitive resources and the generation effect. the quarterly journal of experimental psychology, section a, 45, 149–171. doi:10.1080/14640749208401320 fiedler, k., nickel, s., asbeck, j., & pagel, u. (2003). mood and the generation effect. cognition and emotion, 17, 585–608. doi:10.1080/02699930302301 schindler et al 84 | f l r graesser, a. c., millis, k. k., & zwaan, r. a. (1997). discourse comprehension. annual review of psychology, 48, 163–189. doi:10.1146/annurev.psych.48.1.163 kuznetsova, a., brockhoff, p. b., & christensen, r. h. b. (2014). lmertest: tests for random and fixed effects for linear mixed effect models (lmer objects of lme4 package). r-package version 2.06. retrieved from: http://cran.r-project.org/web/packages/lmertest/index.html mazur, j. e. (2006). lernen und verhalten [learning and behavior] (6th ed.). münchen, germany: pearson studium. mcdaniel, m. a. (1984). the role of elaborative and schema processes in story memory. memory & cognition, 12, 46–51. doi:10.3758/bf03196996 mcdaniel, m. a., & butler, a. c. (2010). a contextual framework for understanding when difficulties are desirable. in a. s. benjamin (ed.), successful remembering and successful forgetting: a festschrift in honor of robert a. bjork (pp. 175-198). new york: psychology press. mcdaniel, m. a., & einstein, g. o. (2005). material appropriate difficulty: a framework for determining when difficulty is desirable for improving learning. in a. f. healy (ed.), experimental cognitive psychology and its applications (pp. 73–85). washington, dc: american psychological association. mcdaniel, m. a., einstein, g. o., dunay, p. k., & cobb, r. e. (1986). encoding difficulty and memory: toward a unifying theory. journal of memory and language, 25, 645–656. doi:10.1016/0749596x(86)90041-0 mcdaniel, m. a., & kerwin, m. l. (1987). long-term prose retention: is an organizational schema sufficient? discourse processes, 10, 237–252. doi:10.1080/01638538709544674 mcdaniel, m. a., waddil, p. j., & einstein, g. o. (1988). a contextual account of the generation effect: a three-factor theory. journal of memory and language, 27, 521–536. doi:10.1016/0749-596x(88)90023x robinson, o., grillon, c., & sahakian, b. (2012). the mood induction task: a standardized, computerized laboratory procedure for altering mood state in humans. protocol exchange, 10. doi:10.1038/protex.2012.007 rosseel, y. (2012). lavaan: an r package for structural equation modeling. journal of statistical software, 48, 1–36. retrieved in february 2017 from: http://www.jstatsoft.org/v48/i02/ slamecka, n. j., & graf, p. (1978). the generation effect: delineation of a phenomenon. journal of experimental psychology: human learning and memory, 4, 592–604. doi:10.1037/0278-7393.4.6.592 van dijk, t. a., & kintsch, w. (1983). strategies of discourse comprehension. new york: academic press. velten, e. (1968). a laboratory task for induction of mood states. behavioural research and therapy, 6, 473– 482. doi:10.1016/0005-7967(68)90028-4 schindler et al 85 | f l r appendix table 1. stimulus material: concepts and descriptions in the read and the generate condition no. concept description in the read condition description in the generate condition 1 abergläubisches verhalten superstitious behavior verhalten, das auftritt, weil ihm zu einem früheren zeitpunkt zufällig oder versehentlich ein verstärker folgte. behavior that occurs because it was by chance or accidentally followed by a reinforcer before. ve_hal_en, das auf_ri_t, weil ihm zu einem frü_ere_ zei_pun_t zu_äl_ig oder ver_ehen_lic_ ein ver_tär_er fol_te. 2 bestrafung typ i punishment type i verhaltensreduktions-methode, bei der einem bestimmten verhalten ein aversiver reiz folgt. procedure for reducing behavior in which a specific behavior is followed by an aversive stimulus. ver_alte_sre_uktio_s-met_ode, bei der einem bes_imm_en ver_alte_ ein ave_sive_ r_iz fol_t. 3 diskrimination discrimination das lernen, auf einen stimulus, nicht aber auf einen anderen, ähnlichen, zu reagieren. learning to respond to a specific stimulus but not to a similar one. das ler_en, auf einen sti_ulu_, nicht aber auf einen an_eren, äh_li_hen, zu rea_ieren. 4 diskriminativer stimulus discriminative stimulus beim operanten konditionieren ein stimulus, der anzeigt, ob eine reaktion zur verstärkung führt. a stimulus that indicates if a reaction will be reinforced in operant conditioning. beim ope_an_en ko_di_io_ieren ein sti_ulu_, der an_eigt, ob eine rea_tio_ zur ver_tär_ung füh_t. 5 fester intervallplan fixed interval plan verstärkungsplan, in dem die erste reaktion nach einem festen zeitintervall verstärkt wird. reinforcement plan in which the first reaction is reinforced after a fixed period of time. ve_stä_ku_gsp_an, in dem die ers_e rea_tio_ nach einem fe_te_ zeiti_ter_all ver_tär_t wird. 6 generalisierung generalization übertragung einer gelernten reaktion von einem stimulus auf einen anderen, ihm ähnlichen. transfer of a conditioned response to one stimulus to another similar one. übe_t_agun_ einer gele_nte_ reak_ion von einem s_imu_us auf einen an_eren, ihm äh_lic_en. 7 gesetz des effekts law of effect reaktionen, denen angenehme oder befriedigende reize folgen, werden verstärkt und zukünftig öfter stattfinden. reactions followed by pleasant or satisfying stimuli are reinforced and will occur more often afterwards. rea_tio_en, denen an_ene_me oder be_riedi_ende rei_e fol_en, werden ver_tär_t und zu_ün_tig öf_er sta_t_inden. 8 habituation habituation nachlassen der stärke einer reflexartigen reaktion nach wiederholter präsentation des auslösenden stimulus. the intensity of a reflexive response to a specific stimulus decreases when the stimulus is presented repeatedly. nac_la_sen der s_är_e einer re_lexa_ti_en rea_tio_ nach wie_er_ol_er prä_enta_ion des au_lö_en_en s_imu_us. 9 konditionierter inhibitor conditioned inhibitor ein konditionierter stimulus, der eine konditionierte reaktion abschwächt oder verhindert. a conditioned stimulus that attenuates or prevents a conditioned response. ein ko_di_ionie_ter sti_ulu_, der eine ko_di_ionie_te rea_tio_ ab_ch_äc_t oder ver_inde_t. 10 konditionierter verstärker ein ursprünglich neutraler stimulus kann eine reaktion durch wiederholte paarung mit einem primären verstärker stärken. ein ur_prü_g_ic_ neu_rale_ s_imu_us kann eine reak_io_ durch wie_er_ol_e schindler et al 86 | f l r conditioned reinforcer an originally neutral stimulus reinforces a reaction by being repeatedly paired with a primary reinforcer. paa_un_ mit einem pri_äre_ ver_tär_er stä_ken. 11 kontext-interferenz context interference situationsmerkmale, die das lernen einer neuen aufgabe erschweren, aber langfristig zu einer besseren leistung führen können. situational characteristics that make learning of a new task more difficult but enhance performance in the long run. si_uatio_s_erk_ale, die das le_nen einer ne_en au_ga_e er_c_were_, aber la_g_ri_tig zu einer be_se_en lei_tun_ füh_en können. 12 kontra-imitation contra-imitation wenn jemand das gegenteil des verhaltens ausführt, das ein modell vorgemacht hat. if someone shows the opposite of the behavior which was demonstrated by a model. wenn jemand das ge_en_eil des ve_ha_te_s au_füh_t, das ein mo_ell vor_ema_ht hat. 13 lerntransfer learning transfer im motorischen lernen die auswirkung der erfahrung mit einer aufgabe auf die performanz bei einer anderen aufgabe. in motor learning the impact of experience with a specific task on performing a different task. im mo_oris_he_ le_nen die au_wi_kun_ der er_ah_un_ mit einer auf_abe auf die pe_fo_ma_z bei einer an_ere_ au_ga_e. 14 löschungsresistenz resistance to extinction der grad, in dem eine reaktion anhält, wenn sie nicht länger verstärkt wird. the extant that a reaction lasts when it is no longer reinforced. der g_ad, in dem eine rea_tio_ an_äl_, wenn sie nicht län_e_ ver_tär_t wird. 15 nachverstärkungspause post-reinforcement pause eine reaktionspause, die bei festen quotenplänen in der regel nach jedem verstärker eintritt. a reaction pause that generally occurs after the presentation of each reinforcer in the context of a fixed ratio plan. eine rea_tio_spau_e, die bei fes_en quo_en_lä_en in der re_el nach je_em ve_s_ärke_ ei_tri_t. 16 negative bestrafung negative punishment verfahren zur verhaltensreduktion, bei dem ein erwünschter reiz beseitigt oder entzogen wird, wenn das verhalten auftritt. procedure for reducing behavior in which an appetitive stimulus is removed or detracted after a specific behavior occurred. ve_fah_en zur ver_al_ens_edu_tion, bei dem ein er_üns_h_er re_z besei_ig_ oder ent_oge_ wird, wenn das ver_al_en auf_ri_t. 17 negative verstärkung negative reinforcement methode der verhaltensstärkung, bei der ein aversiver reiz beseitigt oder entzogen wird, wenn das verhalten auftritt. procedure for reinforcing behavior in which an aversive stimulus is removed or detracted after a specific behavior occurred. me_ho_e der ve_hal_ens_tär_un_, bei der ein ave_si_er r_iz besei_ig_ oder ent_oge_ wird, wenn das ver_al_en auf_ri_t. 18 nichtkontingente verstärkung non-contingent reinforcement die verabreichung von verstärkern zu zufälligen zeitpunkten, unabhängig vom verhalten. random presentation of reinforcers irrespective of the behavior. die ve_ab_eic_ung von ve_s_ärke_n zu zu_äl_ige_ zei_pu_k_en, una_hän_ig vom ver_al_en. 19 primärer verstärker primary reinforcer ein verstärker, der naturgemäß jede reaktion verstärkt, auf die er folgt. a reinforcer that naturally reinforces each reaction that directly preceded the reinforcer. ein ver_tär_e_, der natu_ge_äß je_e rea_tio_ ver_tär_t, auf die er fo_g_. 20 proaktive interferenz wenn zuvor erlernter stoff das erlernen neuer inhalte beeinträchtigt. wenn zu_or erle_n_er sto_f das er_erne_ ne_er inha_te beei_trä_hti_t. schindler et al 87 | f l r proactive interference when previously learned information hinders the learning of new contents. 21 retroaktive interferenz retroactive interference wenn die präsentation von neuem material die erinnerung an früher gelerntes beeinträchtigt. when learning new information hinders the memory of previously learned information. wenn die prä_en_atio_ von ne_em ma_eria_ die eri_ne_un_ an frü_e_ ge_ernte_ beei_trä_hti_t. 22 sättigung saturation eine methode zur verhaltensreduktion, bei der der verstärker in solchen mengen gegeben wird, dass er seine wirksamkeit verliert. procedure for reducing behavior in which the reinforcer is presented repeatedly until its effectiveness decreases. eine me_hode zur ve_hal_ens_edu_tio_, bei der der ve_s_är_er in so_che_ me_gen ge_ebe_ wird, dass er seine wi_k_am_eit ver_iert. 23 selbstverstärkung self-reinforcement eine verhaltensmodifikations-technik, bei der das individuum seine eigenen verstärker für angemessenes verhalten liefert. technique for modifying behavior in which the individual provides reinforcers for appropriate behavior on his/her own. eine ver_al_ens_odi_ikatio_s-tec_ni_, bei der das in_ivi_uum seine eige_en ve_stä_ke_ für ange_esse_es ve_ha_te_ lie_ert. 24 shaping shaping eine methode zum erlernen einer neuen verhaltensweise, bei der die zunehmenden annäherungen an das erwünschte verhalten verstärkt werden. procedure for learning a new behavior in which each behavior is reinforced that increasingly approaches the target behavior. eine me_ho_e zum erle_ne_ einer ne_en ver_al_ens_eise, bei der die zu_eh_ende_ an_ähe_un_en an das er_ünsc_te ve_ha_te_ ve_stä_k_ werden. 25 spontanerholung spontaneous recovery das wiederauftreten einer zuvor gelöschten reaktion, nachdem längere zeit ohne weitere konditionierungsdurchgänge verstrichen ist. recurrence of an extinguished reaction, after some time has passed without further conditioning. das wie_erau_tre_en einer zu_or ge_ösc_te_ rea_tio_, nachdem lä_ge_e ze_t ohne wei_ere kon_itio_ie_un_sdu_ch_än_e ver_tri_he_ ist. 26 teilnehmende modellierung participating modeling eine art des modelllernens, bei der der lernende das verhalten des modells bei jedem behandlungsschritt nachahmt. kind of model learning in which the learner imitates the behavior of a model in each step of the treatment. eine a_t des mo_el_le_ne_s, bei der der le_ne_de das ve_hal_en des mo_el_s bei je_em beha_dlu_gs_ch_itt na_hah_t. 27 transsituationalitätsprinzip principle of transsituationality ein stimulus, der in einer situation als verstärker fungiert, wird auch in anderen situationen als verstärker dienen. a stimulus that serves as reinforcer in one situation will also serve as a reinforcer in different situations. ein s_imu_us, der in einer si_uatio_ als ve_s_är_er fu_gie_t, wird auch in an_ere_ si_uatio_en als ve_s_ä_ke_ die_en. 28 tropismus tropism eine angeborene bewegung eines ganzen organismus als reaktion auf einen spezifischen stimulus. a hereditary movement of a whole organism as a response to a specific stimulus. eine an_ebo_ene be_egun_ eines gan_en or_ani_mu_ als rea_tio_ auf einen spe_ifi_che_ sti_ulu_. 29 überlernen overlearning die fortgesetzte übung einer reaktion, nachdem die leistung scheinbar perfekt ist. continuing with exercise after the performance is apparently perfect. die fo_tge_et_te übu_g einer rea_tio_, nachdem die leis_un_ scheinbar pe_fe_t ist. 30 variabler quotenplan ein verstärkungsplan, bei dem ein verstärker nach einer variablen und nicht vorhersehbaren zahl von reaktionen verabreicht wird. ein ve_stä_kun_s_lan, bei dem ein ve_stä_ke_ nach einer va_iab_en und nicht vor_erse_ba_en za_l von rea_tio_en ve_ab_eich_ wird. schindler et al 88 | f l r variable ratio plan a reinforcement plan in which a reinforcer is presented after a variable and not predictable number of reactions. 31 vermeidung avoidance eine art der negativen verstärkung, bei der durch eine bestimmte reaktion ein aversiver reiz vermieden wird. kind of negative reinforcement in which an aversive stimulus is avoided by a certain reaction. eine a_t der ne_ati_en ve_stä_kun_, bei der durch eine bes_im_te rea_tio_ ein ave_si_er re_z ver_iede_ wird. 32 verteilte übung distributed practice ein übungsverfahren, in dem sich sehr kurze übungsphasen mit ruhezeiten abwechseln. exercise technique in which very brief phases of exercise alternate with pauses. ein übu_g_ve_fah_en, in dem sich sehr kur_e übu_g_p_ase_ mit ru_ezei_en ab_ech_el_. notes. all definitions (concepts and descriptions) were taken from a psychology text book on learning and behavior by mazur (2006) and were slightly modified for the purpose of this study. microsoft word jarodzka et boshuizen_publication.docx frontline learning research vol. 5 no. 3 special issue (2017) 167183 issn 2295-3159 unboxing the black box of visual expertise in medicine halszka jarodzka1,a,b & henny p.a. boshuizena,c awelten institute – research centre for learning, teaching and technology, open university of the netherlands beye tracking laboratory, lund university, sweden cschool of education, university of turku, finland abstract visual expertise in medicine has been a subject of research since many decades. interestingly, it has been investigated from two little related fields, namely the field that focused mainly on the visual search aspects whilst ignoring higher-level cognitive processes involved in medical expertise, and the field that mainly focused on these higher-level cognitive processes largely ignoring the relevant visual aspects. consequently, both research lines have traditionally used different methodologies. recently, this gap is being increasingly closed and this special issue presents methods to investigate visual expertise in medicine from both research lines, namely those investigating vision (eye tracking, pupillometry, flash preview moving window paradigm), verbalisations, brain activity, and performance measures (roc analysis, gesture coding, expert performance approach). we discuss the benefits and drawbacks of each method and suggest directions for future research that could help to unbox the black box of visual expertise in medicine. keywords: expertise; visual expertise; cognition; medicine 1 corresponding author: halszka jarodzka, valkenburgerweg 177, 6419 at heerlen, the netherlands, email: halszka.jarodzka@ou.nl. doi: http://dx.doi.org/10.14786/flr.v5i3.322 jarodzka et boshuizen 168 | f l r 1. visual expertise in medicine as a research field expertise is known to be highly domain-specific (chi, 2006). hence, generalizations from findings cannot be made across different domains and often enough not even across tasks. thus, the organizers of this special issue made a very important decision to focus on medical visual expertise. in this way, we can draw concrete conclusions from each contributing article to unbox the nature of visual expertise in medicine. the topic of visual expertise in medicine has received increasing attention over the past years (gegenfurtner, lehtinen, & säljö, 2011; kok & jarodzka, 2016; kok & jarodzka, 2017; norman, coblentz, brooks, & babcook, 1992; reingold & sheridan, 2011; van der gijp et al., 2016). this comes as no surprise as most medical tasks require some sort of visual input, in the form of a medical image, a tissue sample, or the patient himor herself. as a result, different theoretical models were constructed that describe different aspects of medical expertise. holistic models (for recent descriptions, see kundel, nodine, conant, & weinstein, 2007; nodine & mello-thoms, 2010) describe how experts visually search abnormalities on medical images. another group of theories focus more on cognitive aspects of expertise and concretely describe the development of expertise (boshuizen & schmidt, 1992; feltovich, johnson, moller, & swanson, 1984; norman, young, & brooks, 2007). decades of research related to this group of models have provided us with a thorough idea of how expertise is constituted, however, most often only taking the cognitive component into consideration (e.g., ecg studies gilhooly et al., 1997). lesgold et al. (1988) tried decades ago to marry these two research lines into one unified model. unfortunately, this model was not further developed ever since. recently we have combined the model of lesgold et al. (1988) with more recent cognitive expertise models (boshuizen & schmidt, 2008a) into one model presented in figure 1 (jarodzka, boshuizen, & kirschner, 2012). the fact that these types of models emerged from rather independent research fields, resulted in different types of research methodologies they use. for instance, while visual search models were often studied with eye tracking data and roc abnormality detection rates, cognitive models were often studied with different forms of verbal data. however, over the past few years these differentiations do not hold any more and both research lines learn from each other. the current special issue presents these methodologies, irrespective in which research line they were originally used. this provides the ground towards a more unified theoretical model of visual expertise in medicine which will ultimately unbox the black box of visual expertise in medicine. jarodzka et boshuizen 169 | f l r ` figure 1. theoretical model combining the classical model of lesgold et al (1988) with recent cognitive models of medical expertise (boshuizen & schmidt, 2008) as published in jarodzka, boshuizen, & kirschner (2012). 2. methods presented in the current special issue and the concepts they address this special issue brings together diverse methods that all aim at “unboxing the black box” of medical expertise from different angles. we chose to structure them according to the concepts they are mainly focusing at (figure 2). our structure corresponds to the structuring as gegenfurtner and van merriënboer (2017) use it in their introduction to the current special issue into activation (brain activity), detection (vision), inference (verbalisation), and practice (observations). jarodzka et boshuizen 170 | f l r figure 2. methods to capture diverse concepts related to visual expertise presented in the current special issue. 2.1 vision – the sensory input these articles present methods to measure the sensory input of the medical specialist. how can measuring sensory input help us understanding (visual) expertise? efficient information-processing is a central part of expertise and its development (reingold & sheridan, 2011). on the one hand, an expert is able to detect subtle cues and interpret them within a certain context, but s/he also can detect patterns within seeming random elements. on the other hand, our information-processing system is not a passive recipient, but rather in active search for meaningful information. the theory of boshuizen and schmidt (2008b) gives indications on how this process unfolds: information elements that enter the cognitive systems of novices and intermediates activate nodes within large knowledge networks – depending on their experience, this happens more or less efficient. experts, on the other hand, also begin with a passive reception of information elements. these, however, instantly activate one or several illness-scripts, which in turn, guide the further active search for information (jarodzka, boshuizen, et al., 2012). hence, the passive intake or the active search of (visual) information elements has the potential to reveal crucial aspects of a person’s expertise. the current special issue, provides three manuscripts that discuss methodology to do address this aspect of expertise. 2.1.1 eye tracking the idea that medical experts “see more” than untrained individuals do, appears immediately, when you see a medical expert “reading” from an x-ray or a mammogram. hence, obviously medical expertise was already very early investigated with eye tracking (for a comprehensive overview, see reingold & sheridan, 2011). eye tracking (holmqvist et al., 2011) entails (1) the apparatus that measures the motion of the eye balls in relation to a stimulus, (2) the software that allows to relate parameters derived from the eye movements to certain parts of the stimulus (in time or space), and (3) the researchers, who interpret these findings within a theoretical framework. the clear benefits of this methodology are that it unobtrusively captures unconscious processes directly; it captures all relevant visual input to working memory. on the jarodzka et boshuizen 171 | f l r other hand, eye tracking data are ambiguous, task-dependent, idiosyncractic, and last but not least challenging in data collection and analysis. the article by fox and faulkner-jones (2017) provides a brief historical overview of eye tracking. for a broader view, we would like to refer the reader to the informative and entertaining book by wade and tatler (2005) on the history of eye tracking. we applaud fox and faulkner-jones for their excellent analysis of the different medical tasks and how they should be differently approached by means of eye tracking. we fully agree with them, in particular, as it is well-known how task-dependent expertise is and how broad the field of medicine is, at the same time. we hope, that this detailed analysis will forgo overly generalized statements, such as one these authors surprisingly made themselves “eye-tracking studies across medical specialties have suggested that more experienced physicians require fewer fixations, and less time spent on areas of interest, […] than novices.” (p.3). such statements are not only too reductionistic to reveal interesting insights into the nature of expertise, but even worse: they are often enough simply wrong. in many medical areas, we find exactly the opposite to be true, namely that experts are looking longer at relevant areas of interest (e.g., balslev et al., 2012). that does not mean that studies finding the one or the other were wrong; it means that these findings cannot be generalized, but depend on the exact task and the stimuli that were used. what many of the here reported eye tracking studies in medical expertise “suffer” from, is, that they report eye tracking measures, that are too basic to allow drawing conclusions on the nature of expertise. one example are the findings reported from fox, law, and faulkner-jones (2016) that trainees make more eye movements than experts. by itself, this statement comes down to a simple time-on-task difference, that does not make use of the potentials that eye tracking as a methodology offers. one example of how to make more concrete statements from eye tracking more concrete is a study by kok, de bruin, robben, and van merriënboer (2012). these authors have investigated how experts, in comparison to medical residents and students, visually explored focal and diffused diseases on chest x-rays. amongst others, they calculated a measure that captured, how broadly an image was scrutinized, by calculating the global/local ratio of saccades. in this way, the authors showed that images containing a diffuse disease (i.e., a disease that is spread all over the lungs and cannot be brought down to one location) were visually examined in a broader way, by inducing a higher global/local ration. similarly, in one of our own studies (jaarsma, jarodzka, nap, van merriёnboer, & boshuizen, 2014), we have investigated how expert pathologists, pathology residents, and medical students visually examine pathological slides. amongst other, we found that experts and residents diagnosed the slides equally well. however, the way they processed the slides differed severely: while experts looked immediately to the relevant location and scrutinized it with long fixations, they had afterwards time to explore the rest of the slide – with short fixations – for other potential abnormalities. residents on the other hand, took quite some time to detect the relevant location, and they examined it up to the end of the trial to verify their diagnosis. hence, experts had capacities left over in this task, while residents were at their maximum. this could have made a difference for more complex cases, for instance, with several different diseases in one case. hence, the potentials of eye tracking can be explored much more by going beyond the ‘standard’ eye tracking measures and looking more concretely into the characteristics of the task and the stimulus at hand. another important aspect that fox and faulkner-jones (2017) point out is the lack in research on 3d and dynamic medical images. we fully agree with that, but would like to point towards research not jarodzka et boshuizen 172 | f l r mentioned by these authors, i.e., by bertram, helle, kaakinen, and svedström (2013) on ct images and our own research on interactive digital pathology slides (jaarsma et al., 2016; jaarsma, jarodzka, nap, van merriënboer, & boshuizen, 2015; jaarsma et al., 2014) and on patient-video cases (balslev et al., 2012). moreover, the authors mention on several occasions the potential eye tracking has for medical education. we agree. on that note, we would like to point towards the idea of using eye movements of experts in instructional videos (van gog, jarodzka, scheiter, gerjets, & paas, 2009) and its successful application to medicine (e.g., jarodzka, balslev, et al., 2012). on a final note, it is important to mention that fox and faulkner-jones (2017) have discussed many visual search models, but did not make the connection to current models on medical expertise and its development (jarodzka, boshuizen, et al., 2012). however, as we already argue elsewhere (kok & jarodzka, 2016), this connection is crucial for theoretical development. 2.1.2 pupillometry szulewski, kelton, and howes (2017) describe pupillometry as a method to capture cognitive load in relation to visual expertise in medicine. this idea is very persuasive as it would allow to unobtrusively measure online cognitive processes during medical task performance. pupillometry is actually the use of one very specific measure derived from eye tracking equipment, namely the size of the pupil and how it changes over time. we have to keep in mind though, that the main purpose of changes in the size of the pupil, is the adaptation of the eyes’ photoreceptors to the lighting conditions to allow for optimal vision (davson, 2012). this is a reflex that everyone can easily observe: look close to a mirror in a brightly lighted room. now close one of your eyes off with your hand. open this eye after a moment and observe how your two eyes differ: the eye that was open all the time – and thus exposed to light – has a small pupil. the other eye, the one that was exposed to relative darkness has a larger pupil. this, however, changes very quickly and you can observe how your pupil shrinks to adapt your sight to the changed lightning conditions. research has shown that the size of the pupil may also change if lighting conditions are stable. it may vary according to the interest a participant shows in a stimulus (hess & polt, 1960), their emotional state (vanderhasselt, remue, ng, & de raedt, 2014), the musical chill they experience (laeng, eidet, sulutvedt, & panksepp, 2016) and many other exciting concepts (loewenfeld, 1999). this measure can even be a valid indicator for certain diseases, such as parkinson’s (wang, mcinnis, brien, pari, & munoz, 2016). on the other hand, pupillometry is a rather coarse measure that often cannot provide specific predictions (failure in predicting sexual orientation: savin-williams, cash, mccormack, & rieger, 2016). in any case, it is important to keep in mind, that all these changes in pupil size are subtle and can be easily overruled by a ray of light falling onto the eye. it is therefore important, to keep meticulously equal lightning conditions for both eyes over the entire experiment. this not only holds for the laboratory room, but even the stimulus presentation screen luminosity has to be kept stable. these are conditions that can be realized in fundamental laboratory research, but are difficult to realize in applied research, such as medical education. hence, we will often have to wonder whether the pupil size changes were due to the emotional state of the participants, their mental effort etc. or rather the inevitable changes in the lightning falling onto their eyes in this particular stimulus or their position in the recording room. szulewski et al. (2017) also discuss these and other severe drawbacks of using pupillometry in reallife settings, still they come to an optimistic conclusion that this method would have “a particularly promising role in the field of medicine and in the study of physician expertise development”. we would rather suggest to address these methodological problems by triangulating pupillometry with other mental jarodzka et boshuizen 173 | f l r effort measures that are less vulnerable in real-life settings, for instance, questionnaires (e.g., paas, 1992), dual-task paradigms (brünken, steinbacher, plass, & leutner, 2002) or other less vulnerable physiological data, such as skin resistance (e.g., nourbakhsh, wang, chen, & calvo, 2012). 2.1.3 the flash-preview moving-window (fpmw) paradigm the manuscript by litchfield and donovan (2017) presents the ‘moving window paradigm’ to investigate visual expertise. mcconkie and rayner (1975) introduced this paradigm to investigate the socalled perceptual span in reading. based on the current fixation within a word, a few characters to the left and to the right are masked to investigate to which extent this influences reading. the underlying idea is, that we move our eyes in one direction when reading a text (often from left to right or vice versa, depending on the language), and that the amount of information we can take in towards this direction, without fixating it, increases with increasing expertise in reading. as the visual processing of a written text is so clearly predefined, we know exactly how our eyes will move on a line. unsurprisingly, mcconkie and rayner (and many more afterwards confirmed and specified this) showed that our perceptual span in reading is skewed to the right and indeed depends on our expertise. this is a method widely used and well-established in reading research. when looking at medical pictures, however, there is no such clearly predefined gazing direction as in reading (e.g., line by line, from left to right). hence, the methodological set-up is more complicated. typically, in such ‘scene perception’ settings, researchers use a gausian blurring technique to capture the size of the perceptual span (i.e., the image is blurred apart from a certain area around the current fixation point). the “flash preview moving window paradigm” (litchfield & donovan, 2017) shows an alternative solution. it combines a method of providing participants only with a short glimpse of an image (‘flash preview’: kundel & nodine, 1975) and the above described ‘moving window paradigm’. the really clever aspect about using this method to investigate visual expertise in medicine, is, that it allows estimating (1) to which extend the pre-activation of a schema based on visual input influences, how a consecutive visual search is carried out, and (2) the exact size of the perceptual span in relation to expertise level. the latter shows that – at least in other domains – the size of the visual span increases with higher levels of expertise. more interesting is the first point, though: to which extent, does an initial schema activation guide the actual search for information relevant to this schema? litchfield and donovan could not find strong empirical evidence for this idea (based on the model of kundel, nodine, & toto, 1991). this comes as no surprise, when consulting cognitive theories on medical expertise (as summarized in jarodzka, boshuizen, et al., 2012). these theories assume that medical experts activate a set of schemata that are – partly –instantiated, tested and often discarded. hence, from these theories we would assume that this is rather an ongoing process than a strictly serial one. what would be very interesting for future research, is to take these cognitive theories on medical expertise to inform fpmw research in medicine. considering how a medical expert would pursue in the real world aligns well with these cognitive theories (jarodzka, boshuizen, et al., 2012) and would be very interesting in at least two ways. first, in a real-life situation the expert would first receive several background information of the patient, which would activate an illnessscript (more complex than a schema, see also figure 1). this activated illness-script would already guide the expert in his or her subsequent visual search on the medical image. interestingly, litchfield and donovan (2017) did something in this direction within their third experiment, by showing the target word to the participant right before the flash-preview. it would be very informative for further theoretical developments to extend this approach, by using realistic patient data. second, in real-life, the task of the medical expert is jarodzka et boshuizen 174 | f l r not to simply state the presence or location of a target, but goes far beyond: providing a diagnosis, requesting further examinations of the patient, and finally suggesting a treatment. including these aspects into fpmw experiments, would allow seeing not only a potential influence on the visual search itself, but also on its accompanying cognitive processes. 2.2 verbalisations – the working memory output from a cognitivistic perspective, medical expertise development has been mainly investigated using verbal methods; even for those domains, that heavily draw on visual skill (e.g., radiology lesgold et al., 1988 and ecg interpretation gilhooly et al., 1997). at the same time, the visual challenges of those fields were not much in focus within this paradigm, as it was felt, that verbal methods lacked the acuity to discern and untangle perceptual processes. van de wiel’s article (2017) does a great job in showing how verbal methods can be used in showing reasoning lines and knowledge application by medical experts, intermediates, and novices. it also shows that the validity of verbal methods depends a lot on the vocabulary mastered by the participants, the skill of the researcher to identify references to visual qualities (helle (2017) refers to ericsson’s (2006) “non-verbal thoughts”, captured by briefs labels and referents), and on the successfulness of separating perception of features of the image (if necessary by means of additional methods such as pointing or drawing) from the interpretation of patterns of features in the protocols. the usefulness of verbal methods, stand-alone or in combination with eye-tracking, also depends on a couple of other things: the kind of visual information involved, and – not surprisingly – the research question. even when we constrain ourselves to (stacks of) pictures resulting from medical imaging techniques (e.g., eeg, ecg, x-ray images, microscopic pathological slides), the differences in the amount of information embedded in such pictures are huge. a basic ecg consists of one wiggly line that should show a repetitious pattern with several features such as a p, r and t-top and associated q and s-inflections with their specific amplitude and latency. absence of these features, and irregularities of the pattern may have clinical meaning. importantly, students learn the interpretation of these visual presentations as a combination of features of the visual appearance, the associated vocabulary, and the biomedical and clinical interpretation. though visualising complex phenomena, an ecg is a simple line, though more advanced equipment can visualise several measurements simultaneously (up to twelve for complex diagnoses). a similar analysis applies to eeg, though the number of channels recorded is much higher. the potential amount of information in x-rays pictures, fmri-s, pet-scan and microscopic images is much higher than in line graphs. they are (stacks of) 2d, multi-coloured or grayscale pictures that may maximally vary pixel by pixel, independent of the colour or the greyness of an adjacent pixels. a final difference is related to the nature of certain diseases that may be focal or diffuse. images of local conditions may show isolated, discernible lesions that can be pointed at, but other disease processes only show themselves in the qualities of the image (e.g., ‘cloudy’ or ‘milky’; see kok et al., 2012). these differences in visual qualities of the domain under investigation have implications for vocabulary building, and thus for the usefulness of verbal reports generated by participants of different expertise levels (see van de wiel, 2017), and for foveal detection, and thus for the usefulness of eye-tracking (see helle, 2017). it is almost a platitude to state that the research question affects the investigative data collection and analysis methods to be used. yet, the articles by van de wiel and by helle forget to problematize the assumption that feature extraction should be differentiated from pattern recognition and interpretation (van jarodzka et boshuizen 175 | f l r de wiel), and whether deep (meaning) coding of think-aloud protocols is better than superficial coding that stays close to the exact verbalisations. there are strong indications that visually detecting and interpreting relevant visual features among irrelevant ones, goes hand in hand with interpretation, and is guided by a person’s expectations. many interpretative choices in visual information processing seem to take place in the early and non-analytic phase (norman et al., 2007). on the other hand, recognition of even obvious features (for instance, a clear-cut jaundice) is a gradual process that is interlaced with the developing hypotheses about what might be wrong. manipulation of these hypotheses can lead to the non-recognition of such features (brooks, leblanc, & norman, 2000; leblanc, brooks, & norman, 2002; leblanc, dore, norman, & brooks, 2004; leblanc, norman, & brooks, 2001). the alleged superficiality of our own analyses of verbal protocols turned out to reveal aspects of perception and cognition (jaarsma et al., 2014) that we have never become aware of in earlier research that made use of deep, semantic analyses (see for instance, boshuizen & schmidt, 1992). lack of a professional vocabulary in novice groups seems to be associated with a lack of a repertoire of perceptual features in that domain thus hampering perception and interpretation. stated differently, what might be interpreted as ‘poor’ protocols, may be a veridical representation of the perceptual and cognitive skills of the participant. a decision pro or con one or the other interpretation cannot be made just on procedural features of the way the research method was applied, but requires the theorybased assessment of expertise level, image quality and research question. 2.3 brain activity – neural activity and blood flow measuring the neural activity of the brain, gives us the most ‘pure’ look inside the brain. it is very persuasive to believe that one day we will be able to observe the brain of experts while they are performing a task of their domain and observe how their brilliance unfolds under our very eyes. however, we are not quite there, yet, and the question is whether we ever will – or ever need to be. as fascinating as fancy new visualizations of activities in the brain might be, we have to keep in mind what they represent: increased blood flow in some regions of the brain in comparison to others (fmri: huettel, song, & mccarthy, 2008) or increased neural activity somewhere in the brain (eeg: niedermeyer & da silva, 2005). thus, we can observe and record some physiological processes either specific over time (eeg) or space (fmri). what we cannot tell, though, is which thoughts these activities represent. we need to keep that in mind, when estimating the insights we can derive from such techniques. imagine lying on a narrow stretcher, while your head is tied within a small cage-like object. you are asked not to move and are left alone in the room. then the stretcher begins to move into a tube which makes strange noises. first you might feel very scared (many people do!) and later on coming close to fall asleep (also quite common). this is what it feels like when you participate in an fmri scan. from this we can immediately tell, that this is an extremely fundamental laboratory study. the reason for this extremely restricted position for the participant is that even the slightest movements (even blinks and eye movements) can cause neural activity that was not induced by the experimental intervention. that is also the reason for the very many repetitive trials the experimenter has to run on one participant (to filter this noise out). the situations are similar for other measures of neural activity. when talking about visual expertise research, this seems to be an extremely reductionistic approach. this has to do with two issues: first, expertise is a result of decades of deliberate practice and only measures under representative circumstances (e.g., ericsson & lehmann, 1996). obviously, the scenario described jarodzka et boshuizen 176 | f l r above does not represent any form of visual expertise in nowadays medical expertise. a serious problem is that expertise is extremely domainand task-specific and thus evolves only under the very specific circumstances of the very task. drawing conclusions on expertise from pseudo-expertise tasks (note that studying artificial objects even for several sessions does not qualify for the definition of real expertise), will never be possible. second, when measuring neural activity or blood flow, do we really do justice to the nature of expertise in its complexity? for these and other more pragmatical reasons (costs of conducting such research), most of the research presented in the article by gegenfurtner, kok, van geel, de bruin, and sorger (2017) is actually not related to medical expertise. actually, only three of the presented studies investigated medical expertise as they involved real, medical tasks (fiorio, cesari, bresciani, & tinazzi, 2010; melo et al., 2011; ribas, rocha, siqueira ortega, freitas de rocha, & massad, 2013). interestingly, first studies show activation differences in relation to medical expertise – at least for certain stimuli (hruska et al., 2016). the challenge remains to understand what these found differences actually mean in terms of expertise development. hence, as interesting as these studies are, it is difficult to draw concrete conclusions from them already and clearly far more research is needed. so, do we argue that such medical expertise research is pointless? on the contrary! but we need to be very careful, what it can be used for. we argue that such research on neurological processes can inform fundamental research on memory and attention, which in turn can inform cognitive science, which forms a basis for (medical) expertise research. on that route, it could even eventually reach educational research. what is now urgently needed to make this information flow possible are solid theoretical models that allow for these connections between these research fields. 2.4 observations – behavioural performance one key step in expertise research is to estimate whether participants are ‘real’ experts by checking whether their performance systematically exceeds the one of individuals with less expertise (ericsson & lehmann, 1996; ericsson & smith, 1991). this is easier for some domains than it is for others (e.g., chess expertise can be clearly defined by the elo system). for visual expertise in medicine performance estimates are not trivial as krupinski (2017) describes in her article. in this section, we review three articles that capture very different aspects of performance related to visual expertise in medicine. 2.4.1 roc analysis krupinski (2017) describes the well-established method of receiver operating characteristics (roc) analysis to tackle this issue. this analysis method allows scrutinizing the ability of medical specialists or lay persons to detect one abnormality in a medical image in a very detailed manner. this detailed statistical analysis allows for clear interpretations of the findings. this article provides a very comprehensive and concrete ‘hands-on’ on how to conduct this specific methodology. such an article is of extreme practical value for other researchers. unfortunately, such publications are still rare, even though more of these would be needed. this holds even true for the current special issue; the article by krupinski (2017) has the highest practical value for other researchers interested in the topic of visual expertise in medicine – and not only! many other domains of visual expertise could benefit from this approach, too. as long as the task can be boiled down to a binary, exclusive decision. jarodzka et boshuizen 177 | f l r this is also where the drawbacks of the roc method begin. even though, it is very well-established as a sophisticated statistical method, it is only applicable for a very limited type of task, namely the binary, exclusive decision. however, visual expertise in medicine goes far beyond that, as the author mentions already in the very beginning of her article. already the detection itself goes beyond this simplified present vs. not-present decision: the medical specialist needs to know where the abnormality is located, whether other ones are present as well, etc. more recent types of roc analyses take these issues into account (lroc, jafroc, etc.) and have been used for many years already. although it is important to understand roc first, similar articles to this one on more up-to-date versions of roc would be very important. but still, visual expertise in medicine goes beyond the mere detection of abnormalities. it is only the very first step and does not capture the entirety of medical expertise performance (e.g., the following diagnostic or treatment decision). the next question is thus, whether it is possible to extend these methods to other forms of performance as well. on a final note, the author explains how the specific roc curve and its interpretation depend on the observer’s background and experience. however, she does not draw the connection to existing and wellestablished theories on medical expertise and its development (for a summarized model of them, see jarodzka, boshuizen, et al., 2012). for this method to be able to further contribute to more general (medical) expertise research, this connection to existing cognitive theories on medical expertise (development) is urgently needed and should be the next step in research line. 2.4.2 gesturing the study by ivarsson (2017) investigates a different observable aspect of medical (visual) expertise, namely gestures. as described above, verbalising visual process is a rather difficult endeavour for which gestures can be really helpful. the article by ivarsson shows very concrete examples that depict exactly this. what is important to know is that this study – in contrast to all other articles in this special issue – does not deal with the diagnosis of an individual expert. instead, it acknowledges that a lot of medical practice is carried out in groups. this fact already makes this article a unique contribution to the current special issue. ivarsson investigates a communication situation between several medical professionals. in such a scenario, communication comes into play, of which non-verbal aspects are a crucial part of, especially in the medical profession as this article shows. it might be interesting to know that even though the author mainly addresses movements of the limbs, there seems to be another body part involved as well: examples [21] and [22] indicate that the professional was not only gesturing with her hands, but also guiding the listeners’ attention with her gaze. this is a well-studied and important phenomenon, and following the gazes of others strongly guides our attention (anstis, mayhew, & morley, 1969) in particular in conversations (argyle & cook, 1976; mansfield, farroni, & johnson, 2003). hence, gaze guidance could be explicitly included in such an analysis. an important question within this research line is what the exact function of the gesture is. in the case described in this article, the main function of the gestures used is to establish a common ground between the professionals in a discourse. but is this really the sole purpose of these gestures? are the merely communicative or are they an inherent part of a schema? research in eye tracking has shown, that people often make eye movements even when these are not providing any information, such as when looking at a blank screen or being in the dark (foulsham et al., 2012; johansson & johansson, 2014). in these cases, participants automatically move their eyes when remembering a prior encoded scene. if they are forced not jarodzka et boshuizen 178 | f l r to move their eyes, their recoding performance significantly drops. hence, these eye movements (also a form of body motion) have a functional role in restoring long term memory content. couldn’t the same be true – at least in part – for gestures? ivarsson (2017) chose a rather atypical task for these experts. the situation is very specific and the episode rather short. so, what can we learn from that? the purpose of this endeavour can only be hypothesis building and further research needs to follow to test these hypotheses. these might, for instance, investigate whether the here found gestures are typically used by these professionals? for instance, the ones indicating the representations of digital manipulations [7-9]. is there a gesture ‘language’ that professional use? a different way to apply such methodology would be to investigate individual professionals. for instance, future research could investigate gestures that are part of the clinical routine (e.g., surgery, but also a radiologist turning the x-ray upside down or holding it in a particular angle) or are used as preparations for the clinical routine (a phenomenon that often can be observed in sports). these analyses could be triangulated with other data. in our own research, we have also investigated the interplay between hand movements as navigations within a pathological digital slide, eye movements on this slide as well as the verbalization about this examination (jaarsma et al., 2016; jaarsma et al., 2015; jaarsma et al., 2014). this approach was also already investigated in a more natural setting with mobile eye tracking, although in the non-medical task of tea making (tatler et al., 2013). such a triangulated analysis of eye-hand coordination could be a very meaningful addition to the analysis of gestures. 2.4.3 the expert performance approach williams, fawver, and hodges (2017) provide an excellent overview of methodologies of research on expertise. they describe three steps, namely (1) developing representative tasks that elicit systematic performance differences in individuals of different levels of expertise, (2) process-tracing techniques to study processes underlying expert performance and (3) identifying individual characteristics in training or learning that lead to the expert level. we would like to add to this list the importance of a thorough definition and description of different expertise levels, which is described very concretely in the medical domain by boshuizen and schmidt (2008b). what williams and colleagues (2017) then mainly focus on, is studying the learning towards expertise. in principle, this could be done from two perspectives. on the one hand, one can identify successful methods that improve performance towards visual expertise. one example of an instructional method to train aspects of visual expertise are eye movement modelling examples (van gog et al., 2009). these are instructional videos showing how an expert model approaches a task. therefore, the model verbally explains the steps taken in this task. moreover, the attentional focus of the expert (based on his or her eye movements) is overlaid on this video. we have already successfully applied this method in the medical domain (jarodzka, balslev, et al., 2012). we showed that diagnosing patient-video cases improved after such a training not only in terms of performance, but also on the visual processes. williams and colleagues (2017) favour another approach: the investigation of individual learning trajectories to identify ‘good’ and ‘poor’ learners. based on such an analysis they argue, not only instruction could be improved, but also the identification of future experts might be possible. a long-standing research line of self-regulated learning in non-medical professions could be very informative for such a research (kicken, brand-gruwel, van merriënboer, & slot, 2009). in any case, williams et al. – and we fully agree with them – call for two important issues: (a) more longitudinal studies to truly understand the development of visual expertise in jarodzka et boshuizen 179 | f l r medicine, as well as (b) a process-tracing approach to identify relevant (and probably heterogeneous) processes underlying this development. 3. lessons learned and future avenues this special issue presented different methods to investigate different aspects of visual expertise in medicine emerging from different research lines within this field. to reach the aim of unboxing the black box of visual expertise in medicine, we argue that future research in this line, should consider the following points: • methodological triangulation: we saw that each of the methods presented in this special issue have unique potentials, but also severe drawbacks. to counterbalance these drawbacks, we need to use more methodological triangulation of different methods when conducting studies within visual expertise in medicine (see also: gegenfurtner et al., 2016; kok & jarodzka, 2016). • we need to systematically discuss more on the challenges we face with new methodologies and detailed process measures. unfortunately, traditional empirical articles leave hardly room to do this. we thus plead for a forum where such issues could be identified, evaluated and solutions to them agreed upon. • several articles in this special issue have shown the importance of more interdisciplinary research that combine different research fields, such as medical image perception and scene processing. • finally, as uttered many times throughout this discussion, we need more solid theoretical models that allow to form bridges between these different methodologies presented in this special issue (see also: gegenfurtner et al., 2016; kok & jarodzka, 2016). references anstis, s. m., mayhew, j. w., & morley, t. (1969). the perception of where a face or television portrait is looking. american journal of psychology, 82, 474-489. argyle, m., & cook, m. (1976). gaze and mutual gaze. oxford, england: cambridge university press. balslev, t., jarodzka, h., holmqvist, k., de grave, w. s., muijtjens, a., eika, b., . . . scherpbier, a. (2012). visual expertise in paediatric neurology. european journal of paediatric neurology, 16 (2), 161166. doi: 10.1016/j.ejpn.2011.07.004 bertram, r., helle, l., kaakinen, j. k., & svedström, e. (2013). the effect of expertise on eye movement behaviour in medical image perception. plos one, 8(6), 1-15. doi: 10.1371/journal.pone.0066169 boshuizen, h. p. a., & schmidt, h. g. (1992). on the role of biomedical knowledge in clinical reasoning by experts, intermediates, and novices. cognitive science, 16(2), 153-184. doi: 10.1207/s15516709cog1602_1 boshuizen, h. p. a., & schmidt, h. g. (2008a). the development of clinical reasoning expertise. in j. higgs, m. jones, s. loftus & n. christensen (eds.), clinical reasoning in the health professions (pp. 113122). oxford: butterworth heinemann / elsevier. boshuizen, h. p. a., & schmidt, h. g. (2008b). the development of clinical reasoning expertise: implications for teaching. in j. higgs, m. jones, s. loftus & n. christensen (eds.), clinical reasoning in the health professions (vol. 3rd rev. ed. ed., pp. 131-121). oxford, uk: butterworthheinemann/elsevier. jarodzka et boshuizen 180 | f l r brooks, l. r., leblanc, v. r., & norman, g. r. (2000). on the difficulty of noticing obvious features in patient appearance. psychological science, 11, 112-117. doi: 10.1111/1467-9280.00225 brünken, r., steinbacher, s., plass, j., & leutner, d. (2002). assessment of cognitive load in multimedia learning using dual-task methodology. experimental psychology, 49, 109-119. doi: 10.1027//16183169.49.2.109 chi, m. t. h. (2006). two approaches to the study of experts' characteristics. in k. a. ericsson, n. charness, r. r. hoffman & p. feltovich (eds.), the cambridge handbook of expertise and expert performance (pp. 21-30). cambridge: cambridge university press. davson, h. (2012). physiology of the eye: elsevier. ericsson, k. a. (2006). protocol analysis and expert thought: concurrent verbalizations of thinking during experts’ performance on representative tasks. in k. a. ericsson, n. charness, p. j. feltovich & r. r. hoffman (eds.), the cambridge handbook of expertise and expert performance (pp. 223-241). cambridge: cambridge university press. ericsson, k. a., & lehmann, a. c. (1996). expert and exceptional performance: evidence of maximal adaption to task constraints. annual reviews in psychology, 47(1), 273-305. doi: 10.1146/annurev.psych.47.1.273 ericsson, k. a., & smith, j. (1991). prospects and limits in the empirical study of expertise. in k. a. ericsson & j. smith (eds.), towards a general theory of expertise: prospects and limits (pp. 1-38). cambridge, ma: cambridge university press. feltovich, p. j., johnson, p. e., moller, j. h., & swanson, d. b. (1984). lcs: the role and development of medical knowledge in diagnostic expertise. in w. j. clancey & e. h. shortliffe (eds.), readings in medical artificial inteligence: the first decade (pp. 275-319). reading, ma: addison-wesley publishing company. fiorio, m., cesari, p., bresciani, m. c., & tinazzi, m. (2010). expertise with pathological actions modulates a viewer’s motor system. neuroscience, 167, 691-699. doi: 10.1016/j.neuroscience.2010.02.010 foulsham, t., dewhurst, r., nyström, m., jarodzka, h., johansson, r., underwood, g., & holmqvist, k. (2012). comparing scanpaths during scene encoding and recognition: a multidimensional approach. journal of eye movement research, 5(5), 1-12. doi: 10.16910/jemr.5.4.3 fox, s. e., & faulkner-jones, b. e. (2017). eye-tracking in the study of visual expertise: methodology and approaches in medicine. frontline learning research, 5(3), 43 54. doi: 10.14786/flr.v5i3.258 fox, s. e., law, c. c., & faulkner-jones, b. e. (2016). quantitative gaze assessment of a dual ”side by side” viewer versus a single whole slide image viewer for pathology education. journal of pathology informatics. gegenfurtner, a., kok, e., van geel, k., de bruin, a. b. h., jarodzka, h., szulewski, a., & v., v. m. j. j. g. (2016). the challenges of studying visual expertise in medical image diagnosis. medical education, 51(1), 97-104. doi: 10.1111/medu.13205 gegenfurtner, a., kok, e., van geel, k., de bruin, a. b. h., & sorger, b. (2017). neural correlates of visual perceptual expertise: evidence from cognitive neuroscience using functional neuroimaging. frontline learning research, 5(3), 14 30. doi: 10.14786/flr.v5i3.259 gegenfurtner, a., lehtinen, e., & säljö, r. (2011). expertise differences in the comprehension of visualizations: a meta-analysis of eye-tracking research in professional domains. educational psychology review, 23(4), 523-552. doi: 10.1007/s10648-011-9174-7 gegenfurtner, a., & van merriënboer, j. (2017). methodologies for studying visual expertise. frontline learning research, 5(3), 1-13. doi: 10.14786/flr.v5i3.361 gilhooly, k. j., mcgeorge, p., hunter, j., rawles, j. m., kirby, i. k., green, c., & wynn, v. (1997). biomedical knowledge in diagnostic thinking: the case of electrocardiogram (ecg) interpretation. european journal of cognitive psychology, 9(2), 199-223. doi: 10.1080/713752555 helle, l. (2017). prospects and pitfalls in combining eye-tracking data and verbal reports. frontline learning research, 5(3), 81 93. doi: 10.14786/flr.v5i3.254 jarodzka et boshuizen 181 | f l r hess, e. h., & polt, j. m. (1960). pupil size as related to interest value of visual stimuli. science, 132(3423), 349-350. doi: 10.1126/science.132.3423.349 holmqvist, k., nyström, m., andersson, r., dewhurst, r., jarodzka, h., & van de weijer, j. (2011). eye tracking: a comprehensive guide to methods and measures. oxford: oxford university press. hruska, p., hecker, k. g., coderre, s., mclaughlin, k., cortese, f., doig, c., . . . krigolson, o. (2016). hemispheric activation differences in novice and expert clinicians during clinical decision making. advances in health science education, 21(5), 921–933. doi: 10.1007/s10459-017-9787-9 huettel, s., song, a., & mccarthy, g. (2008). functional magnetic resonance imaging. sunderland, ma: sinauer associates. ivarsson, j. (2017). visual expertise as embodied practice. frontline learning research, 5(3), 123 138. doi: 10.14786/flr.v5i3.253 jaarsma, t., boshuizen, h. p. a., jarodzka, h., nap, m., verboon, p., & van merriënboer, j. j. g. (2016). tracks to a medical diagnosis: expertise differences in visual problem solving. applied cognitive psychology, 30(3), 314-322. doi: 10.1002/acp.3201 jaarsma, t., jarodzka, h., nap, m., van merriënboer, j. j. g., & boshuizen, h. p. a. (2015). expertise in clinical pathology: bridging the gap. advances in health sciences education, 20(4), 1089-1106. doi: 10.1007/s10459-015-9589-x jaarsma, t., jarodzka, h., nap, m., van merriёnboer, j. j. g., & boshuizen, h. p. a. (2014). expertise differences under the microscope: processing histopathological slides. medical education, 48(3), 292-300. doi: 10.1111/medu.12385 jarodzka, h., balslev, t., holmqvist, k., nyström, m., scheiter, k., gerjets, p., & eika, b. (2012). conveying clinical reasoning based on visual observation via eye-movement modelling examples. instructional science, 40(5), 813-827. doi: 10.1007/s11251-012-9218-5 jarodzka, h., boshuizen, h. p. a., & kirschner, p. a. (2012). cognitive skills in medical diagnosis and intervention. in p. lanzer (ed.), catheter-based cardiovascular interventions; knowledge-based approach (pp. 69-86). berlin, germany: springer. johansson, r., & johansson, m. (2014). look here: eye movements play a funtional role in memory retrieval. psychological science, 25, 236-242. doi: 10.1177/0956797613498260 kicken, w., brand-gruwel, s., van merriënboer, j. j. g., & slot, w. (2009). the effects of portfolio-based advice on the development of self-directed learning skills in secondary vocational education. educational technology, research and development, 57, 439-460. doi: 10.1007/s11423-009-9111-3 kok, e. m., de bruin, a. b. h., robben, s., g. f., & van merriënboer, j. (2012). looking in the same manner but seeing it differently: bottom-up and expertise effects in radiology. applied cognitive psychology, 26(6), 854-862. doi: doi: 10.1002/acp.2886 kok, e. m., & jarodzka, h. (2016). before your very eyes: the value and limitations of eye tracking in medical education. medical education. doi: 10.1111/medu.13066 kok, e. m., & jarodzka, h. (2017). beyond your very eyes: eye movements are necessary, not sufficient. medical education. doi: 10.1111/medu.13384 krupinski, e. (2017). receiver operating characteristic (roc) analysis. frontline learning research, 5(3), 31 42. doi: 10.14786/flr.v5i3.250 kundel, h. l., & nodine, c. (1975). interpreting chest radiographs without visual search. radiology, 116, 527-532. kundel, h. l., nodine, c. f., conant, e. f., & weinstein, s. p. (2007). holistic component of image perception in mammogram interpretation: gaze-tracking study1. radiology, 242(2), 396-402. doi: 10.1148/radiol.2422051997 kundel, h. l., nodine, c. f., & toto, l. (1991). searching for lung nodules. the guidance of visual scanning. investigative radiology, 26(9), 777-781. laeng, b., eidet, l. m., sulutvedt, u., & panksepp, j. (2016). music chills: the eye pupil as a mirror to music’s soul. consciousness and cognition, 44, 161-178. doi: 10.1016/j.concog.2016.07.009 jarodzka et boshuizen 182 | f l r leblanc, v. r., brooks, l. r., & norman, g. r. (2002). believing is seeing: the influence of a diagnostic hypothesis on the interpretation of clinical features. academic medicine, 77(10), 67-69. leblanc, v. r., dore, k., norman, g. r., & brooks, l. r. (2004). limiting the playing field: does restricting the number of possible diagnoses reduce errors due to diagnosis-specific feature identification? medical education, 38, 17 24. doi: 10.1111/j.1365-2923.2004.01730.x leblanc, v. r., norman, g. r., & brooks, l. r. (2001). effect of a diagnostic suggestion on diagnostic accuracy and identification of clinical features. academic medicine, 76(10), 18-20. lesgold, a., rubinson, h., feltovich, p., glaser, r., klopfer, d., & wang, y. (1988). expertise in a complex skill: diagnosing x-ray pictures. in m. t. h. chi, r. glaser & m. farr (eds.), the nature of expertise (pp. 311-342). hillsdale, nj: erlbaum. litchfield, d., & donovan, t. (2017). the flash-preview moving window paradigm: unpacking visual expertise one glimpse at a time. frontline learning research, 5(3), 66 80. doi: 10.14786/flr.v5i3.269 loewenfeld, i. (1999). the pupil: anatomy, physiology, and clinical applications. oxford: butterworthheinemann. mansfield, e., farroni, t., & johnson, m. (2003). does gaze perception facilitate overt orienting? visual cognition, 10(1), 7-14. doi: 10.1080/713756671 mcconkie, g. w., & rayner, k. (1975). the span of the effective stimulus during a fixation in reading. perception & psychophysics, 17, 578-586. melo, m., scarpin, d. j., amaro, e., passos, r. b., sato, j. r., & friston, k. j. (2011). how doctors generate diagnostic hypotheses: a study of radiological diagnosis with functional magnetic resonance imaging. plos one, 6, e28752. doi: 10.1371/journal.pone.0028752 niedermeyer, e., & da silva, f. (2005). electroencephalography. lippincott williams & wilkins. nodine, c. f., & mello-thoms, c. (2010). the role of expertise in radiologic image perception. in e. samei & e. krupinski (eds.), medical image perception and techniques (pp. 139-156). cambridge: cambridge university press. norman, g. r., coblentz, c. l., brooks, l. r., & babcook, c. j. (1992). expertise in visual diagnosis: a review of the literature. academic medicine, 67(10), 78-83. norman, g. r., young, m., & brooks, l. (2007). non-analytical models of clinical reasoning: the role of experience. medical education, 41, 1140-1145. doi: 10.1111/j.1365-2923.2007.02914.x nourbakhsh, n., wang, y., chen, f., & calvo, r. a. (2012). using galvanic skin response for cognitive load measurement in arithmetic and reading tasks. paper presented at the proceedings of the 24th australian computer-human interaction conference. paas, f. (1992). training strategies for attaining transfer of problem-solving skill in statistics: a cognitiveload approach. journal of educational psychology, 84, 429-434. doi: 10.1037//0022-0663.84.4.429 reingold, e. m., & sheridan, h. (2011). eye movements and visual expertise in chess and medicine. in s. p. liversedge, i. d. gilchrist & s. everling (eds.), oxford handbook of eye movements (pp. 523-550). oxford: oxford university press. ribas, l. m., rocha, f. t., siqueira ortega, n. r., freitas de rocha, a., & massad, e. (2013). brain activity and medical diagnosis: an eeg study. bmc neuroscience, 14, 109. doi: 10.1186/1471-2202-14-109 savin-williams, r. c., cash, b. m., mccormack, m., & rieger, g. (2016). gay, mostly gay, or bisexual leaning gay? an exploratory study distinguishing gay sexual orientations among young men. archives of sexual behavior, 1-8. doi: 10.1007/s10508-016-0848-6 szulewski, a., kelton, d., & howes, d. (2017). pupillometry as a tool to study expertise in medicine. frontline learning research, 5(3), 55 65. doi: 10.14786/flr.v5i3.256 tatler, b. w., hirose, y., finnegan, s. k., pievilainen, r., kirtley, c., & kennedy, a. (2013). priorities for selection and representation in natural tasks. philosophical transactions of the royal society b, 368(1628), 1-9. doi: 10.1098/rstb.2013.0066 jarodzka et boshuizen 183 | f l r van de wiel, m. (2017). examining expertise using interviews and verbal protocols. frontline learning research, 5(3), 94 122. doi: 10.14786/flr.v5i3.257 van der gijp, a., ravesloot, c., jarodzka, h., van der schaaf, m., van der schaaf, i., van schaik, j., & ten cate, t. j. (2016). how visual search relates to visual diagnostic performance: a narrative systematic review of eye-tracking research in radiology. advances in health sciences education, 1-23. doi: 10.1007/s10459-016-9698-1 van gog, t., jarodzka, h., scheiter, k., gerjets, p., & paas, f. (2009). attention guidance during example study via the model‘s eye movements. computers in human behavior, 25(3), 785-791. doi: 10.1016/j.chb.2009.02.007 vanderhasselt, m.-a., remue, j., ng, k. k., & de raedt, r. (2014). the interplay between the anticipation and subsequent online processing of emotional stimuli as measured by pupillary dilatation: the role of cognitive reappraisal. frontiers in psychology, 5, 207. doi: 10.3389/fpsyg.2014.00207 wade, n. j., & tatler, b. (2005). 'the moving tablet of the eye': the origins of modern eye movement research. oxford: oxford university press. wang, c.-a., mcinnis, h., brien, d. c., pari, g., & munoz, d. p. (2016). disruption of pupil size modulation correlates with voluntary motor preparation deficits in parkinson's disease. neuropsychologia, 80, 176-184. doi: 10.1016/j.neuropsychologia.2015.11.019 williams, a. m., fawver, b., & hodges, n. j. (2017). using the ‘expert performance approach’ as a framework for improving understanding of expert learning. frontline learning research, 5(3), 139 154. doi: 10.14786/flr.v5i3.267 frontline learning research vol.6 no 1 (2018) 54-76 issn 2295-3159 corresponding author: cornelis j. de brabander, department of educational studies, leiden university, the netherlands, email address: brabander@fsw.leidenuniv.nl doi: http://dx.doi.org/10.14786/flr.v6i1.342 testing a unified model of task-specific motivation: how teachers appraise three professional development activities cornelis j. de brabandera,b,c & folke j. glastraa a department of educational studies, leiden university, the netherlands b query informatisering, voorhout, the netherlands c welten institute, open university of the netherlands article received 1 january / revised 20 february / accepted 20 february / available online june 21 abstract this article tests the tenability of a unified model of task-specific motivation (umtm). the umtm integrates task-specific components from several theories of motivation. core of the model are four interacting but relatively independent types of valences. affective and cognitive valences represent feelings while doing an activity and thoughts about the value of its consequences respectively; both affective and cognitive valences can be positive and negative, hence calling for approach and avoidance motivation respectively. the interaction between these four types of valences results in a valence appraisal that influences readiness for action. task-specific antecedents, autonomy, feasibility, social relatedness and subjective norm, influence valences. 441 primary school teachers provided judgments of all components of the model except social relatedness for three imaginary professional learning activities. the three activities were framed as a school board decided, a team decided and a personally decided learning activity. structural equation modelling showed that for each activity a separate model was needed. how valences influenced readiness for action was specific to each activity. in the board and team decided activities, for instance, readiness for action appeared to be based predominantly on cognitive valences, while in the personally decided activity affective and cognitive valences showed a more balanced contribution. regarding task-specific antecedents, however, the picture was less clear. nevertheless, the umtm proved to offer rich possibilities for the explanation of complex motivational phenomena and promises a significant reduction of the superabundance of theories that encumbers motivation research. keywords: motivation; intrinsic motivation; extrinsic motivation; subjective task value; professional development mailto:brabander@fsw.leidenuniv.nl de brabander et glastra 54 | f l r 1. introduction this article explores the structural relations between the components of a unified model of taskspecific motivation (umtm). in order to allow this exploration we collected appraisals from primary school teachers about imaginary professional learning activities. in the field of educational psychology, and not only there, we may notice a vast proliferation of theories of motivation (boekaerts, van nuland, & martens, 2010; schunk, pintrich, & meece, 2008). these theories appear to conflict in different ways. the umtm (figure 1) is an attempt to integrate existing theories on task-specific motivation and is explained in depth by de brabander & martens (2014). the umtm focuses on task-specific aspects of motivation, that is, motivation for rather specific action options that are open to the actor. the model does not refer to motivation for broad fields of action like sports or mathematics, but is applicable to rather specific acts like reading job relevant books or papers. furthermore, the model focuses on motivation and consequently disregards actual action or the feedback loops from actual action. its intention is to provide a picture of the components that are necessary to describe the motivation of a person for a specific activity at a specific point in time. the constituent parts of the model stemmed from or were suggested by several motivation theories, like the self-determination theory (deci & ryan, 2000; ryan & deci, 2000a, 2000b), the flow theory of csikszentmihalyi (1990), the person-object theory of interest proposed by krapp (2002; 2005), and several expectancy*value theories, namely the social-cognitive theory of bandura (1977; 1986; 1992; 1997; 2001; schunk & pajares, 2010), the expectancy*value theory of achievement motivation (wigfield & eccles, 2000) and the theory of planned behavior of ajzen and fishbein (ajzen, 1991; ajzen & fishbein, 2008). the reasons for the selection of these theories and the exclusion of others are clarified by de brabander and martens (2014). 2. the unified model of task-specific motivation for a full explanation of the unified model of task-specific motivation we refer to de brabander and martens (2014). the account of the model given here is a slightly adapted version. by and large, however, the analysis and arguments of the original authors still apply. 2.1. affective and cognitive valences motivation is defined as readiness for action. it is a certain level of willingness to do an activity. readiness for action is influenced by a valence appraisal in which affective and cognitive valences are combined on a common scale. the valence appraisal represents the overall attractiveness of an activity. affective and cognitive valences are produced by separate and relatively independent systems of behaviour regulation. affective valences are produced by an affective system of behaviour regulation; cognitive valences are produced by a cognitive behavioural regulation system (krapp, 2002, 2005; epstein, 1994). affective valences are defined as feelings that originate from undertaking an activity (cf. self-determination theory, flow theory). affective valences represent the levels of positive and negative feelings one experiences when doing an activity. they emanate from an automatic (unavoidable) and mechanical, that is unintentional, response to an action ‘object’ the person apperceives. any action that comes to mind as an action option immediately brings about feelings about that action. the important characteristic of affective valences is not that they have no reason or function, but that their reason or function is not necessarily known and often is not known. cognitive valences are defined as the articulation and valuation of consequences of performing an activity (cf. expectancy*value theories). cognitive valences are explicit and brought about by active reflection of the prospective actor. any activity has multiple consequences, intended and unintended, each of which will have a separate cognitive valence. the quality of cognitive valences obviously depends on the actor’s competence to articulate and value consequences of an activity. de brabander et glastra 55 | f l r cognitive valences can be broken up in different parts depending on who acquires the cognized outcome value. evidently, the first one that profits from an activity is always the acting person himself or herself. what other components of cognitive valence need to be distinguished depends on the specific context. in the context of education with teachers as actors, the student and the school as an institution are natural candidates. in other contexts these non-personal parts of cognitive valences will have different labels. affective valences may also have different parts, because some parts or aspects of an activity may be pleasurable, while others may be less pleasurable or even unpleasurable, but affective valences have no non-personal component. relevant are only the feelings that the person self experiences about an activity. feelings of other persons (as attributed by the actor) can only enter the motivational play in the form of non-personal cognitive valences. an important field of conflict between different motivation theories is the relation between affective and cognitive valences. in self-determination theory affective and cognitive valence are oppositional endpoints on a single dimension: from completely extrinsic to completely intrinsic. expectancy*value theories blur the conceptual distinction and regard both types as additive: intrinsic value is just another type of subjective value. the umtm follows krapp (2002, 2005) by defining affective valences and cognitive valences as two theoretically independent categories of motives that emanate from two interacting, but distinct systems of behaviour regulation. what is pleasurable to do, in general is also recognized as profitable and vice versa, but discrepancies between the two are also possible (‘i know it is good for me, but i absolutely don’t feel like it’ or ‘i have not the faintest idea what it is good for, but it certainly is fun to do) and it is the merit of selfdetermination theory to have demonstrated the possible adverse effects of such discrepancies in educational figure 1. a unified model of task-specific motivation. adapted from de brabander & martens (2014). de brabander et glastra 56 | f l r contexts on performance quality, self-esteem, and general well-being (ryan & deci, 2000a). this view of affective and cognitive valences as discrete but related implies that they are sometimes additive and sometimes opposed, and that readiness for action consequentially depends on the specific configuration of all valences involved. 2.2. positive and negative valences the distinction between affective and cognitive valences is combined with the distinction between approach and avoidance motivation (elliot, 2006; elliot & church, 1997). this distinction is related to the possibility of valences to be positive or negative. positive valences call for an action that promises to realize them. negative valences call for refraining from an action or for initializing a counteraction in order to prevent their realization. thus, the quantity of valences is not one-dimensional, but needs to be split up in two relatively independent dimensions: positive valences, running from neutral to highly positive, and negative valences, running from neutral to highly negative. combining the distinction between positive and negative valences with the distinction between affective and cognitive valences allows for four relatively independent motivational components. affective and cognitive valences, both consisting of positive and negative valences, combine in a resultant value, which we call valence appraisal, which in turn determines readiness for action. such a conception is very much compatible with the notion of multiple consequences (including costs!) combining into the cognitive valence of a course of action. affective valences also comprise multiple aspects, some of which are pleasurable whereas others may be unpleasurable. this fourfold conceptualization promises to provide more and better possibilities for the explanation of complex motivational phenomena. 2.3. task-specific antecedents the umtm identifies autonomy, feasibility, relatedness, and, adopted from the theory of planned behaviour, subjective norm as important task-specific antecedents of valences. autonomy refers to the origin of the action: the self or an internal or external source that is experienced as foreign. feasibility that replaces the concept of competence for reasons that will become clear below concerns the question whether it is possible to complete an activity successfully. relatedness refers to the level of belongingness between the people that participate in the action context. subjective norm entails the inclination to abide by the task-related norms of others in the action context. there is a multitude of possible relations between these antecedents and action valences, but for the moment, it is assumed that they all contribute to all kind of valences. that is to say, that a higher level of task-specific autonomy, feasibility, relatedness and subjective norm can influence task-specific motivation, because they can affect the level of affective and cognitive, positive and negative valences. it is important to emphasize that the model is limited to task-specific aspects of these antecedents. the umtm focuses, for instance, on the autonomy one has in a specific activity, not on the level of autonomy one experiences in general in a certain context. evidently, that is not to say that the latter is unimportant, but it is assumed that in the end its influence is actualized in a task-specific implementation. a particular feature in the model is the distinction between the person and his or her context, which leads to a partition of some components in the model. the distinction between personal and non-personal cognitive valences was introduced above already. in addition also autonomy and feasibility are divided in a personal and a contextual facet. the personal facet of autonomy is labelled sense of personal autonomy and is defined as the extent to which the person experiences himself or herself as the originating force that drives performance of an activity. the distinction here is whether the self feels driving or driven. the contextual part of autonomy is labelled perceived freedom of action and entails the liberty a person perceives in the action context to decide independently about choice and execution of action alternatives (cf. reeve, nix, & hamm, 2003). of course, both aspects are related. their fundamental distinctiveness, however, is disclosed by the possibility that one can still experience the self as the driving force in a situation where one has no freedom of action. that possibility is also the reason why it is assumed that sense of personal autonomy is the shaping de brabander et glastra 57 | f l r factor and that in a specific context perceived freedom of action is important to the extent it contributes to sense of personal autonomy. notice again that we are dealing here only with task-specific factors. the role of a more general freedom of action (for instance at the level of the school in general) might very well be different. furthermore, it was speculated that sense of personal autonomy is differently related to different types of valences. sense of personal autonomy affects cognitive valences, but sense of personal autonomy and affective valence are presumably reciprocally related. if i experience myself as origin of my actions, it is very likely that i also experience pleasure during performance, and the other way around: if i experience pleasure performing an activity, it is very likely that i have a feeling of personal autonomy. in concrete situations sense of personal autonomy and affective valence act like communicating vessels. feasibility has also two sides. the personal side of feasibility is sense of personal competence and refers to the estimate of personal capacities, knowledge and skills that are available to complete an activity at an acceptable level of performance. the contextual part, an estimate of circumstances that hinder or support performance of a specific activity, is called perceived external support (de brabander, rozendaal, & martens, 2009; imants & de brabander, 1996; smit, de brabander, & martens, 2014). both are needed for successful performance, which is why a feasibility appraisal was conceived as the result of both components. it is assumed, however, that sense of personal competence in general will supply the higher contribution of the two, because people can overcome hindrances and obstacles by stepping up their effort, for instance. theoretically it would be possible to apply the distinction between person and context also to relatedness, with on one hand the level of relatedness one perceives between the members of an action context in general as the contextual aspect, and on the other the level of relatedness one personally feels with the members of an action context as the personal aspect. however, as long as a distinctive role of the contextual part is unlikely, the model makes only room for the personal part: sense of personal relatedness. there is one more aspect that makes relatedness different from the other conditional factors, not theoretically, but practically. certainly in educational settings, the group of people that populates the action context is relatively stable across tasks. in practice then, sense of personal relatedness is less task-specific than autonomy and feasibility. 2.4. direct or mediated effects different theories also disagree about the paths of influence of different antecedents. the socialcognitive theory posits that efficacy expectations shape outcome expectations, which in turn determine motivation. (bandura, 1977; 1986; 2001; schunk & pajares, 2010) the theory of planned behaviour, however, claims that perceived behavioural control and subjective norm both directly influence intention formation (ajzen, 1991). de brabander & martens (2014), introducing their unified model of task-specific motivation, favoured the mediated option. they found it difficult to understand, for instance, why the sheer expectation of a successful performance by itself would increase readiness for action. it seemed more logical that a higher feasibility would lead to a more positive and a less negative valence appraisal and that this valence appraisal would determine readiness for action. their first empirical exploration of the model, however, suggested that task-specific antecedents could exert their influence along all kinds of paths, direct and mediated by valences (de brabander & martens, 2018). in that study, however, no data on readiness for action were available and instead reports of actual participation were used as a proxy. they found that in addition to effects that were mediated by valences, task-specific antecedents could also have direct effects on readiness for action. we assume, therefore, that to some extent and in some cases, people may gain readiness for action simply because they feel being the origin of their action, because they see good opportunities for a successful performance, because they experience positive relations to other people in the action context, and/or because they want to abide by normative beliefs about that action of significant people in the action context. 2.5. time of appraisal de brabander et glastra 58 | f l r developing their unified model of task-specific motivation de brabander and martens (2014) used a future time perspective. they concentrated on activities in the near future. as a result they defined all components of the model as expectancies. however, this future time perspective was only a methodical aid to ease the development of the model. appraisal of all components of the model can be made at different points in time, not only in foresight, but also during performance of an activity and in hindsight. point of time will certainly make a difference. it will be easier to anticipate the consequences of an activity during performance of an activity, because a more precise estimate of its outcomes is possible. assessment of outcomes after finishing an activity requires no anticipating competence. but in all cases, also in the latter, people will vary in what consequences they perceive and how they value them. it is conceivable that the value of outcomes is revealed to the person even many years after completion of an activity (e.g. high school graduation). therefore, we adapted in the foregoing the definition of all components of the model, where necessary, to enable different times of appraisal by changing “expectancy” into “appraisal”. 2.6. research questions the current study was one of the very first attempts to acquire empirical support for the unified model of task-specific motivation. as described above this model aims to identify important constructs that are needed to understand the level of motivation and the factors that contribute to this motivation. however, what level of readiness for action a person reaches and how it is formed depends on the valences that actually play a role and on the specific configuration of task-specific antecedents, which in turn depend among others on the specific task at hand. in the current investigation, therefore, we were not in a position to formulate strict hypotheses, but relied instead on two guiding questions. the first question was to what the extent empirical relations between the measures of the components of the umtm are in accordance with model expectations. the second question was to what extent different activities require different models as specific instantiations of the general model. in the current investigation we tested the unified model of task-specific motivation in the context of motivation for professional learning activities of teachers in primary education. professional learning activities are seen as an important instrument for the improvement of the quality of education. this belief is shared between the dutch ministry of education, school boards and teachers (research voor beleid, 2011). although in this paper we did not have an interest in teacher professional development per se, it is a relevant field from the perspective of motivation theory. the reasons for this can be found in three debates with regard to teachers and their professional development. we will briefly discuss debates about the professional identity of teachers, educational leadership, and the design of teacher professional development. with regard to the professional identity of teachers, swann, mcintyre, pell, hargreaves, and cunningham (2010) point out that teacher autonomy has become a controversial issue lately. in recent education policies, governments in many countries are seeking to replace traditional notions of professional autonomy among teachers by a modern professional identity based on value added performance, efficiency, accountability, teamwork and the willingness to implement change (ball, 2003; hardy & lingard, 2008). against the background of current educational reform, the question is raised whether teachers are to be conceived as real professionals in the technical sense of the word or as executive educational workers; should their professional autonomy in educational matters be trusted or should their performance be more or less scripted and closely monitored (apple & jungck, 1990; ball, 2003; ballet & kelchtermans, 2009; clark, livingstone, & smaller, 2012; gleeson & husbands, 2003; milner, 2013); should their professional development be determined by themselves or should it be structured by governmental educational reform agenda’s (borko, elliott & uchiyama, 2002; starkey et al., 2009). in 2012, the dutch ministry of education concluded with some satisfaction that teachers’ professional development choices reflected its educational reform agenda to a considerable extent (ministerie van ocw, 2012, pp. 18-20). for teachers in canada, clark, livingstone, and smaller (2012) have shown that teachers experienced increasing constraints in determining the course of their professional development. policy-driven, compulsory professional development programs met with fierce criticism among these teachers (clark, livingstone, & smaller, 2012, p. 93; also locke, de brabander et glastra 59 | f l r vulliamy, webb, & hill, 2005). a dutch survey among primary and secondary teachers (onderwijscoöperatie, 2016) reported that decisions about teacher participation in professional development were made by principals or school boards in 48% of the cases, in 1% by the teachers, and in 36% the decision was made jointly by teachers and school boards or principals. closely linked to the issue of teacher professional identity is the debate on educational leadership. decisions about teacher learning issue from the leadership structures and cultures of schools in which both the discretionary powers of school boards or principals, and the ‘professional space’ for teacher decisions are shaped (bakkenes, 1996). silins, mulford, and zarins (2002) have argued that leadership in schools can be organized on several organizational levels. according to them, the presence of leadership at each organizational level, i.e. school board or principal, the school team, and the individual teachers, is an important condition for organizational learning and professional development. the literature about distributed and teacher leadership takes the team level of leadership as its focal point (devos, tuytens & hulpia, 2013; hulpia, devos & van keer, 2011). the assumption here is that the impact that teachers have, both individually and collectively as a team on school decisions will have a beneficial influence on their commitment to school improvement and to professional development. hargreaves and dawe (1990) made an important proviso to such expectations. they made a distinction between teacher collaborative cultures where teacher define their own professional development and innovation goals as a community, and ‘contrived collegiality’ where teacher collaboration is administratively enforced in the context of learning to implement skills and programs developed elsewhere. a third debate concerns the design of teachers’ professional development. little (1993, p.142) refers to a dominant training and guidance model of professional development of teachers with ‘packaged knowledge and standardized programs’ (cf. hardy, 2010). they are provided by a market for formal training programs and bought off the shelf by schools accountable for their investment in professional development to school boards and inspectorates. in these programs, teachers are positioned as more or less passive consumers. little (1993) states that the fit of these programs with learning needs of teachers is often problematic. she goes on to describe alternatives to the standardized professional development format: teacher and school collaboratives and partnerships, which give central importance to teaching contexts and teacher experiences (penuel, fishman, yamaguchi, and gallagher, 2007). these alternative programs demand active and collaborative construction of knowledge from teachers. from these three debates, we derived a dimension that relates to teacher autonomy in the sense that it identifies the source of decision-making on participation in professional learning activities. principals or school boards as decision makers impact the professional autonomy of teachers directly since they imply as a rule mandatory participation for the teachers. penuel, fishman, yamaguchi, and gallagher (2007) have pointed out that this will in many cases result in incoherence between professional development programs and individual and situation dependent learning goals of teachers. individual teachers as decision makers will tend to reinforce the notion of professional autonomy among teachers, and may imply better chances of coherence between professional development and individual and situational learning needs, although they may also bring a risk of narrowing down the scope of learning (little, 1993). teams as decision makers about teachers’ professional development, the intermediate level of decision making between school board and individual decisions, may involve different levels of collective professional autonomy, dependent upon whether they are instances of teacher collaborative cultures or issue from ‘contrived collegiality’. hence, in our research we discern three sources of decision-making. the decision to take part in professional learning activities can be taken by the school board or the principal, the team of teachers, or the teacher personally. to sum up, when teachers are confronted with three possible professional learning activities, differing in how they were decided upon, namely by the school board, the team of teachers, or the teacher personally, (1) to what extent is it possible to model the motivational data on these activities according to the umtm, and as the umtm is not a deterministic model but a model in which differences between activities determine the de brabander et glastra 60 | f l r relations between different components, (2) do the different activities lead to different models and in what respect? 3. method 3.1. sample the sample was a convenience sample of 441 teachers from 54 primary schools in the netherlands, 377 female and 63 male (1 missing value). mean age was 42.17 years (s.d. = 12.09). the age distribution, however, had two modes, at 28 and 56 years approximately. mean number of years of experience was 16.92 (s.d. = 11.25), again with a bi-modal distribution with modes at 6 and 34 years of experience. of the 402 teachers who reported their appointment size (mean = 30.42 hours, s.d. = 9.22), 121 (30%) had a full-time appointment (40 hours). relatively more common part-time appointment sizes were 20 hours (10.2%), 24 hours (8%) and 32 hours (10.4%). 3.2. variables all components of the umtm (except sense of personal relatedness) were translated in a single question with a bipolar seven-point scale (table 1). sense of personal relatedness was discarded because to measure it as a task-specific aspect would require the subject to report on the relations with people involved in a specific activity, and who would be involved in that imaginary activity was by definition not known in advance. the items on positive and negative cognitive valences were differentiated into three items according to the recipient of these valences: the actor self (p-pcv respectively p-ncv), the students (s-pcv respectively sncv), or the school as an institute (i-pcv respectively i-ncv). this was accomplished by interposing between the two poles three scales that were labelled “for me personally”, “for the student”, and “for the school”. a short explanation of their meaning preceded the items on affective and cognitive valences. this set of questions was administered three times, once for each of the three professional learning activities. teachers were asked how they would appreciate these imaginary activities. each activity was described as a professional learning activity aimed at improving a specific instructional problem, which would take place in the very near future. activities differed in how they were decided upon. the first activity was framed as a local school board decision, the second as a team decision, and the third as a personal decision. the three activities were introduced as follows: “based on its accountability program, the school board decided to focus on improvement of technical reading and reading comprehension. the board has recruited a third party to organize a course for all teachers of the schools for which it is responsible. your school received an announcement that the next seminar will be on technical reading and reading comprehension.” “after careful consultation the teachers assembly determined a number of professional development priorities. the topic of the forthcoming seminar will be one of these priorities: performance-oriented practice.” “year after year you noticed that some students have great difficulty to master fraction calculus. you want to participate in a seminar about a new approach that promises to offer a better support for these students. your school provides the budget you need to follow this course.” 3.3. analysis the relations between the components of our theoretical model were analysed with structural equation modelling. in line with our research questions we used an analysis process that is called ‘model generation’ de brabander et glastra 61 | f l r (kline, 2011, p. 8). model generation involves alteration of an initial theoretical model, in our case derived from the umtm, based on residual covariances, lagrange multiplier test, and wald tests until a model is found that theoretically makes sense, is reasonably parsimonious, and has an acceptably close correspondence to the data. kline also notes that this probably is the most common type of use of structural equation modeling. but, more importantly, it fitted our theoretical purpose better than a strict confirmatory approach, because we expected the three types of activities to be different specific instances of a general model. what level of readiness for action a person eventually reaches depends very much on the valences that actually play a role and on the specific configuration of task-specific antecedents, which in turn both depend on the specific task at hand. the umtm implies a distinction between direct and mediated effects. to determine their existence we relied on the so-called test of joint significance, which is superior to using bias-corrected bootstrapping and at the same time far more efficient (leth-steensen & gallito, 2016). this test means that a mediation effect simply is significant when all paths contributing to the mediated effect are significant. developing different models for different tasks would eventually lead to models in which constructs were represented by single items. we did not choose to employ multiple items for each construct to spare teachers the task of completing a lengthy questionnaire that would have to be administered three times anyway. obviously, with single item measures reliability cannot be determined, but reliability is implied by systematic covariation with other measures. a second problem of single item measures relates to the measurement level. technically, a seven points scale provides data on an ordered categorical level. however, using 5 categories or more approximate continuous data and using robust maximum likelihood estimation, which we already needed because many of our variables did not satisfy normality assumptions, further reduces the bias level (finney & distefano, 2013). all analyses were implemented with eqs (bentler, 2008). de brabander et glastra 62 | f l r table 1 item formulations for task-specific components sense of personal autonomy (spa) i would have the feeling that i participated in such an activity … only because i had to only because i myself wanted to. perceived freedom of action (pfa) such an activity would offer … very much very little … opportunities for free choice. perceived external support (pes) i find the facilities and circumstances in our school … very frustrating very conducive … to completing such an activity successfully. sense of personal competence (spc) i myself feel … very competent not competent at all … to complete such an activity successfully. subjective norm (snc) i think that colleagues whom i feel connected to, would assess my participation in such an activity … not positively very positively. positive affective valences (pav) during preparation and execution of such an activity i would have … very often seldom or never … a positive feeling. negative affective valences (nav) during preparation and execution of such an activity i would have … seldom or never very often … a negative feeling. positive cognitive valences (p-pcv, spcv, i-pcv) considering the positive consequences, such an activity would be … not or barely profitable very profitable. negative cognitive valences (p-ncv, sncv, i-ncv) to my estimate the costs and any unwanted consequences would be … very consequential negligible. readiness for action (rfa) if such an activity was about to take place, i would be willing to invest … very little very much … effort in it. de brabander et glastra 63 | f l r table 2 distribution characteristics of task-specific components n mean sd skewness kurtosis value std. error value std. error board activity spa 438 3.84 1.423 .059 .117 -.586 .233 pfa 436 3.36 1.296 .102 .117 -.532 .233 pes 436 4.97 1.187 -.556 .117 .150 .233 spc 437 5.51 1.304 -1.260 .117 1.442 .233 snc 437 4.35 1.340 -.344 .117 -.344 .233 pav 439 4.89 1.100 -.490 .117 -.035 .233 nav 439 3.33 1.238 .252 .117 -.702 .233 p-pcv 438 5.23 1.112 -.982 .117 1.293 .233 s-pcv 437 5.34 1.130 -1.059 .117 1.402 .233 i-pcv 437 5.36 1.076 -.952 .117 1.186 .233 p-ncv 438 3.59 1.459 .201 .117 -.784 .233 s-ncv 436 3.00 1.477 .464 .117 -.596 .233 i-ncv 433 3.51 1.386 .191 .117 -.688 .234 rfa 435 4.90 1.162 -.814 .117 .666 .234 team activity spa 440 4.65 1.339 -.453 .116 -.310 .232 pfa 439 4.18 1.430 -.168 .117 -.728 .233 pes 441 5.06 1.187 -.639 .116 .117 .232 spc 441 5.41 1.431 -1.125 .116 .563 .232 snc 437 4.90 1.223 -.625 .117 -.016 .233 pav 439 5.02 1.222 -.859 .117 .289 .233 nav 439 3.01 1.194 .776 .117 .295 .233 p-pcv 440 5.42 1.066 -1.083 .116 1.831 .232 s-pcv 441 5.54 1.002 -.845 .116 1.538 .232 i-pcv 441 5.58 .990 -1.001 .116 2.094 .232 p-ncv 437 3.32 1.426 .457 .117 -.556 .233 s-ncv 436 2.84 1.409 .639 .117 -.377 .233 i-ncv 436 3.22 1.349 .438 .117 -.558 .233 rfa 435 5.19 1.148 -1.063 .117 1.582 .234 personal activity spa 440 6.24 1.030 -2.020 .116 5.600 .232 pfa 437 5.11 1.737 -.696 .117 -.639 .233 pes 439 5.22 1.331 -.773 .117 .079 .233 spc 439 5.90 1.317 -1.888 .117 3.522 .233 snc 437 5.30 1.366 -.971 .117 .333 .233 pav 438 5.70 1.312 -1.549 .117 2.426 .233 nav 438 2.48 1.335 1.316 .117 1.356 .233 p-pcv 438 6.10 .983 -1.412 .117 2.652 .233 s-pcv 436 5.92 .934 -1.171 .117 2.199 .233 i-pcv 438 5.69 1.022 -1.051 .117 1.938 .233 p-ncv 439 2.77 1.471 .895 .117 .094 .233 s-ncv 438 2.52 1.391 .973 .117 .444 .233 i-ncv 436 2.89 1.378 .647 .117 -.207 .233 rfa 439 5.89 1.056 -1.271 .117 2.082 .233 de brabander et glastra 64 | f l r 4. results 4.1. descriptive statistics in table 2 the distribution characteristics of the task-specific components were collected (cf. also table 1). higher scores represent a higher level of the variable involved. figure 2. correlograms of task-specific components in three activities. legend: color shades and pie charts indicate the correlation level. blue shades and clockwise pie charts represent positive correlations, red shades and counter-clockwise pie charts represent negative correlations. de brabander et glastra 65 | f l r 4.2. differences between activities the distribution characteristics of the task-specific components (table 2) revealed already some differences between the three professional learning activities. and although the pattern of correlations between task-specific components showed a substantial agreement (figure 2), the fit of a confirmatory factor analysis model with 14 factors under the assumption that there are essentially no differences between activities, was highly unacceptable: χ2 (n=396) = 3446.529, df = 728, cfi = .605, nnfi = .533, rmsea = .097 90% ci: .094 .100. from this analysis we concluded that it was indeed necessary to develop a separate model for each activity. we were not able, however, to corroborate this conclusion with a multitrait-multimeasure model in which each task-specific component was allowed to load both on a task-specific component and on an activity: the estimation phase took 332 iterations to alleviate non-positive definiteness and another 396 iterations to reach convergence and still produced illogical parameter estimates. figure 3. structural regression model of the task-specific components for the board decided professional development activity. thus, each professional development activity was analysed separately. therefore, the majority of components in the model was measured by one item, and consequently was represented by manifest variables. in each case we started with the relations as suggested by the umtm. the endogenous variable in this model was readiness for action (rfa). according to the umtm, readiness for action is influenced by a valence appraisal (va). as we did not have direct measurements of this valence appraisal, we modelled the covariance between valences and readiness for action as a latent variable with valences as indicators. however in each type of activity it proved to be possible to combine cognitive valences for the school (institute, i-pcv and incv) and for the student (s-pcv and s-ncv) in two latent variables: non-personal positive cognitive valences (np-pcv) and non-personal negative cognitive valences (np-ncv). thus six types of valences were combined in de brabander et glastra 66 | f l r a latent factor: positive and negative affective valences (pav and nav), personal positive cognitive valences (ppcv), non-personal positive cognitive valences (np-pcv), personal negative cognitive valences (p-ncv), and non-personal negative cognitive valences (np-ncv). we hypothesized that these six types of valences, in turn, were influenced by task-specific antecedents. sense of personal autonomy influenced all valences and was influenced by perceived freedom of action (pfa). in the umtm sense of personal competence (spc) and perceived external support (pes) combine into a feasibility appraisal, however, based on their low correlation in our dataset and considering that we had no direct measurement of feasibility appraisal, we treated them as separate components. the last task-specific antecedent that was expected to influence valences was subjective norm of colleagues (snc). in each analysis, we removed or added paths based on residual covariances, lagrange multiplier tests and wald tests to attain three models as specialized instances of the general model. thus we arrived at the three models that are pictured in figure 3, 4, and 5. to ease the readability of these figures, the multitude of paths between task-specific antecedents and the different valences was presented in embedded tables. the fit measures of all three models were satisfactory and are reported in the figures. the results of these analyses are presented under three headings: the influence of valences on readiness for action, the influence of task-specific antecedents on valences, and the direct influence of task-specific antecedents on readiness for action. the correlated error terms that were left free for estimation in order to improve model fit are reproduced in the figures. figure 4. structural regression model of the task-specific components for the team decided professional development activity. de brabander et glastra 67 | f l r 4.3. influence of valences on readiness for action for two activities, the board activity and the personal activity, one valence appraisal factor was not enough to capture the covariance between valences and readiness for action (rfa) sufficiently. to account for residual covariances between valences and readiness for action we had to assume a second valence appraisal factor. in the model for the team activity, however, it was possible to develop a model with one valence appraisal factor with a more or less equal fit compared to a model with two valence appraisal factors. we chose the latter based on its comparability to the other two and its interpretability. though this part of all three models had thus the same components, the path coefficients were quite different. in the model for the board activity the first valence appraisal factor (va1) pav had the highest loading, followed by nav with a negative loading. the positive cognitive valence components had smaller loadings and negative cognitive valence components had the smallest loadings on va1. in the second valence appraisal factor (va2) only cognitive valences were combined with loadings that are more or less of the same level as the loadings on va1. however, the path coefficient of va2 to rfa (.57) was substantially higher that the path coefficient of va1 to rfa (.23). this was also the case in the model for the team activity (.60 vs .23). however, the loading of pav on va1 was much lower. here, the negative loading of nav was the highest, and the loadings of cognitive valences on va1 were at the same (smaller) level. again for this activity, the loadings of cognitive valences on va2 were very modest. in the model for the personal activity the path coefficients were very different. differences in loadings on va1 between affective and cognitive valences were smaller and all loadings were rather substantial. and here, only the personal cognitive valences had loadings on va2 and they were relatively small. furthermore, the path coefficients of va1 and va2 to rfa were equal. de brabander et glastra 68 | f l r figure 5. structural regression model of the task-specific components for the personally decided professional development activity. 4.4. influence of task-specific antecedents on valences sense of personal autonomy (spa) influenced all valences in all activities with a relatively high path coefficient on personal positive cognitive valence (p-pcv). rather surprisingly there were hardly any influences of sense of personal competence (spc) on valences in the three activities. subjective norm of colleagues (snc) had influence on four valence components in the board activity, on all six in the team activity, but only on two in the personal activity, namely the positive cognitive valences, and both at a hardly meaningful level. perceived freedom of action (pfa) contributed to spa in all activities. in the team activity and the personal activity pfa contributed also to spc. furthermore, in the board activity and the team activity pfa had also a small influence on pav. in none of the activities perceived external support (pes) influenced any of the valences. in all three activities snc influenced pfa, pes and spa. 4.5. direct influence of task-specific antecedents on readiness for action all activities revealed a substantial direct effect of spa on rfa, especially in the personal activity. there was in none of the activities a direct effect of spc on rfa. snc showed a modest direct contribution to rfa in the board and the team activity, but not in the personal activity. de brabander et glastra 69 | f l r 5. discussion the results of our analyses showed that cognitive valences dominated readiness for action in the board decided and the team decided activity, but that the influence of affective and cognitive valences was more balanced in the personally decided activity. in the latter case the first valence appraisal factor was a balanced combination of affective and cognitive valences, whether positive or negative; a small part of the valences, namely the personal cognitive valences combined in the second valence appraisal factor. still these two factors of valences contributed equally to readiness for action. these results imply that in the board decided and the team decided activities readiness for action was predominantly based on the estimated value of the expected consequences of such an activity. in the personally decided activity on the other hand affective and cognitive valences were more balanced and made a stronger contribution to readiness for action, but on top of that, personal cognitive valences made a special and substantial contribution. this means that the feelings that the subjects expected to accompany the personally decided activity, and the estimated values of its consequences contributed equally and substantially to readiness for action, and that the subjects’ readiness for action was heightened considerably to the extent they expected valuable outcomes for themselves personally. the board decided activity scored a lower mean value on readiness for action than the team decided activity, that in turn scored lower than the personally decided activity, but all mean values were relatively high. with respect to the task-specific antecedents of valences, the picture was less clear. only sense of personal autonomy had a clear influence on all types of valences: a higher sense of personal autonomy led to more positive and less negative valences. the influence of perceived freedom of action was largely mediated by sense of personal autonomy, though, except in the personally decided activity, there was a small contribution to positive affective valences: perceiving more freedom of action slightly raised expectations of positive feelings during these activities. sense of personal competence, however, did barely influence the valences. and its pattern of influences was rather incomprehensible. only in the case of positive affective valences we might discern a pattern: the influence of sense of personal competence was lowest in the board decided activity, somewhat higher in the team decided activity and rose to a meaningful level in the personally decided activity. if we could order the activities according to ownership, the board decided activity would be on the low end and the personally decided activity on the high end. thus a bold guess would be that sense of personal competence mattered to the level of positive feelings the teacher expected to the extent he or she felt ‘owner’ of an activity. however, the results of this study on the relation between valences and feasibility as an immediate task-specific antecedent did not support the umtm. in all theories of motivation competence plays an important role in motivation and that role is well established by research. maybe the skewedness of the scores on sense of personal competence posed a limit on its relations with other variables, but this cannot be the full explanation: for instance, its skewedness was highest in the personally decided activity, where the relation with positive affective valence was also highest. perceived external support played a very minor role in the board and team decided activity. possibly, the proposed activities were simply too imaginary to allow an estimate of the influence of perceived external support on their expected performance. perceived freedom of action, however, meaningfully influenced sense of personal competence in the team decided and the personally decided activity: if i am not free to decide, i can’t do it. subjective norm had its widest influence on valences in the team decided activity. the board decided activity came second and its role in the personally decided activity was hardly worth mentioning. furthermore, in all three activities subjective norm contributed positively to perceived freedom of action, sense of personal autonomy, and perceived external support. in all activities, we witnessed also some direct influences on readiness for action from the task-specific antecedents. most notably sense of personal autonomy contributed to readiness for action in all three activities, in the board decided activity, slightly higher in the team decided activity and highest in the personally decided activity. subjective norm showed also some direct effects, but only in the board and the team decided activities. de brabander et glastra 70 | f l r all in all we concluded that the differentiation in different types of valences, which is a hallmark of the umtm, was strongly supported by the different configuration of valences that we found in different activities. each activity provoked a different set of affective and cognitive, and positive and negative valences that carried different weights for the resulting readiness for action. the path coefficients of the valences in different activities aptly demonstrated the intricate interaction between different types of valences. with respect to the relation between task-specific antecedents and valences we needed a more balanced account. sense of personal autonomy influenced all valences as expected. however, sense of personal competence largely failed to influence valences: sense of personal competence was relevant in some cases to some valences and generally at a very low level. a possible explanation could be the imaginary character of the activities. what exactly would be the requirements of the proposed activities possibly was not clear yet. on the other hand, these activities were not uncommon, so that competence expectations could have been taken from experience. subjective norm was relevant to the valences, provided that the activity had a social character, as was the case in the board decided and the team decided activity, but less so in the personally decided activity. de brabander and martens (2018) similarly witnessed this conditional role of subjective norm. taking also its direct influence on readiness for action in two of the three activities into account, we may conclude that subjective norm proved to be an important addition to the umtm. autonomy appears to be a special factor in the way teachers frame their professional live. sense of personal autonomy influenced all valences in every activity. and in all activities, though in somewhat varying degrees, there was a direct effect of sense of personal autonomy on readiness for action. furthermore, perceived freedom of action was related to sense of personal competence in the two activities, where the teacher was not completely free. originally (de brabander & martens, 2014) the assumption prevailed that the influence of task-specific antecedents is mediated by valences. thus, subjective norm, for instance, would affect readiness for action because, when everybody involved supports an activity, doing that activity would be more pleasurable and would lead to more and better outcomes. the current study and the study by de brabander & martens (2018) leave no other option than to revise this position. the direct effects of task-specific components on readiness for action (in figure 1 represented by dashed lines) have to be acknowledged as real possibilities. and these possibilities have to be extended to include also autonomy. sense of personal autonomy has a direct effect on readiness for action. the high strength of this effect could have been the consequence of the type of framing we used in this study: the locus of decision-making. an open question remains how to interpret these direct effects. the easiest case may be subjective norm. when important colleagues are in favour of an activity, that alone may be enough to make one willing to adopt that activity: people need not bother to consider any consequences or to pay attention to whether it is pleasurable or not; they are willing to get into that activity, because they are inclined to live up to expectations of valued colleagues. it becomes more difficult to imagine what happens when feasibility directly influences readiness for action. we did not find such an effect in this study, but it was found by de brabander & martens (2018). such an effect would involve doing an activity simply because you can do it. it might be the case, however, that valences are taken more for granted to the extent one feels competent at a certain activity. the case that is most difficult to understand is represented by the direct effect of sense of personal autonomy: doing an activity for the sake of being the driving force in that activity. in the current study this effect may have been exaggerated because of the type of framing used, but in other cases (de brabander & martens, 2018) this effect was also apparent. as was stipulated before, this study was one of the first attempts to collect empirical evidence on the umtm. in these attempts we had no intention to develop sharp hypotheses in order to discriminate between alternative theories. it is nevertheless worthwhile to reflect on the implications of the results of our analyses for the different theories that contributed to the umtm. an obvious first issue here is the relation between affective and cognitive valences. self-determination theory assumes an opposition between intrinsic and extrinsic motivation: motivation is optimal when it is intrinsic, but as soon as consequences of an activity play a motivating role, motivation is less optimal. extrinsic motivation is further differentiated into different types, de brabander et glastra 71 | f l r where some types of extrinsic motivation (identification and integration), which together with intrinsic motivation are called autonomous motivation, are less detrimental than other types (external regulation and introjection), which are called controlled motivation (ryan & deci, 2000a; 2000b). the umtm assumes a conceptual independence both between cognitive and affective, and between positive and negative valences. all structural models in the current study showed that positive cognitive and affective valences were positively related. this was, obviously, not a very surprising outcome as research based on expectancy*value theory consistently found a positive relation between intrinsic value and other types of subjective value (eccles, wigfield, harold, & blumenfeld, 1993; eccles & wigfield, 1995; wigfield & cambria, 2010). our results are no exception in that respect. once again, we may conclude that the adoption of a universal opposition between intrinsic and extrinsic motivation is not warranted. additional support for this conclusion comes from impeccable sources, namely meta-analyses of the effect of verbal rewards (deci, koestner, & ryan, 1999; 2001). if extrinsic versus intrinsic motivation is a one-dimensional contrast, we would expect autonomous types of extrinsic motivation to have a less impeding effect than controlled types of extrinsic motivation. however, these meta-analyses showed that verbal rewards (positive feedback) that are experienced as autonomy supportive (i.e. an autonomous type of extrinsic motivation) actually enhanced motivation substantially, that is, were more motivating than intrinsic motivation on its own (cf. de brabander & martens, 2014). furthermore, by asking in the current study also about negative valences, we showed that there are always also negative valences that are negatively related to positive valences, implying in the first place that cognitive valences can be opposed to affective valences. secondly, we found in our results not only signs of possible oppositions between cognitive and affective valences, but also between affective valences and between cognitive valences. an activity does have both consequences that are valued positively and consequences that are valued negatively, which is understandable when we realize that profits come only at the expense of effort, and that activities have multiple consequences, some valued positively and others negatively. our results showed, furthermore, that activities also may spawn both positive and negative feelings, in different quantities. an activity is an assembly composed of different aspects, some of which may be more or less pleasant and others more or less unpleasant. thus, this differentiation of conceptually independent positive and negative, and affective and cognitive valences is more plausible than the conception of an opposition of intrinsic and extrinsic motives, irrespective of the further differentiation of the latter. we do regard the oppositional view of intrinsic and extrinsic motivation as a premature coagulation of theoretical thoughts about motivation in the early days of self-determination theory. confronted with the widespread use of extrinsic rewards, unrelated to the inherent goals of learning activities, and the conspicuous differences with activities apparently based on the joy of performing, self-determination theorists adopted an oppositional view that was generalized to all types of extrinsic motivation. it is about time to revise this view, which we do regard to affect only a minor part of the rich body of knowledge that self-determination theory has to offer, including the valuable distinction between different types of extrinsic motivation. expectancy*value theories construe an additive relation between different types of subjective value (de brabander & martens, 2014). the reflection on our results led us to conjecture that not this additivity per se is the key issue after all the combination of positive and negative, affective and cognitive valences in a resultant valence appraisal also represents some sort of weighted compound but instead the fact that this additive relation is based on a failure of expectancy*value theories to recognize the conceptual difference between affect and cognition. in the expectancy*value theory of achievement motivation (wigfield & eccles, 2000) intrinsic value, attainment value, and utility value are seen as different types of subjective value, which gives them implicitly a certain level of empirical independence: for instance, for some people behavioural choices might be determined more by intrinsic value, whereas for others attainment value might be more important. but fundamentally, all values belong to the same category of subjective values. and although it is nowhere explicitly defined as such, this would imply that the expectancy*value theory of achievement motivation views intrinsic value actually as a cognitive appraisal of the pleasure that is connected to a specific activity. that is a different view than the conceptualization of affective valences, which, to play their role in behavioural regulation, may, but need not be conscious. the theory of planned behaviour is not explicit about the differences between affect and cognition, but in one study on leisure activities ajzen (1991, p. 200) refers to empirical differences between “affective” and “evaluative” judgments of activities. however, despite their de brabander et glastra 72 | f l r convergent and discriminant validity, the distinction was not judged relevant, let alone conceptually scrutinized, as it did not improve the prediction of leisure intentions. our results obviously do not offer any definite arguments for a choice between the umtm and expectancy*value theories, because the methods we used were comparable. we tried to elicit affective valences by asking our subjects to reflect on the feelings that would accompany different activities, just the same as in studies based on expectancy*value theories. by definition, therefore, we have missed feelings that failed to reach consciousness and were nevertheless influential. this difference between implicit and explicit valences will be an important matter of concern in future research. the second point for reflection relates to the distinction between approach and avoidance motivation. both self-determination theory and expectancy*value theories do not examine potential differences between positive and negative motivators. in self-determination theory (ryan & deci, 2000a; 2000b) external regulation, which is one of the four types of extrinsic motivation, involves motivation based on external rewards and punishments. as such it is the only type of motivation that mentions positive and negative incentives explicitly, but without attention for any conceptual difference. introjection is another type of controlled motivation that is based on approval from self or others. a guileless reader might be inclined to see disapproval as implicitly included in approval, but then again there is no attention for their difference. other types of extrinsic motivation, identified regulation and integrated regulation, do not provide room for negative incentives. expectancy*value theories do not distinguish conceptually between positive and negative values. the expectancy*value theory of achievement motivation admittedly defines costs, but as just another type of subjective value (wigfield & eccles, 2000), which subsequently is more or less forgotten as a topic for research (flake et al., 2015; wigfield & cambria, 2010; for an exception see battle and wigfield, 2003). according to the theory of planned behavior consequences can be beneficial or detrimental but their differences are not conceptually designated (ajzen, 1991). our study did, obviously, not lead to a sharp verdict on this matter, but the differential role of negative valences in our analyses suggest that a neglect of the conceptual difference between approach and avoidance motivation is not justified. although the difference between approach and avoidance motivation is convincingly substantiated empirically (elliot, 2006; carver, 2006), it is for future research to establish its relevance for the distinction between positive and negative valences in the umtm. all results of our studies confirm that the umtm is not a deterministic model. it is a framework that describes different components that may be important for motivated action in concrete cases. it is a heuristic model that is useful to describe motivation in specific situations. which components play a role and how important they are depends very much on the specific activity, its specific context and the specific characteristics of the actor. therefore, we also need many more studies of the umtm involving a wide variety of activities and of actors to explore the impact of this variation on the relative importance of different components. for instance, we referred above to the conspicuous role of autonomy-related aspects in the umtm. an interesting question is to what extent these autonomy-related phenomena are typical of the population of teachers. the role of autonomy in our data reminded us of the topic of teacher isolation that received much attention in the past (bakkenes, 1996; bakkenes, de brabander, & imants, 1999; little, 1990; lortie, 1975). we need many more distinctive examples to be able to separate the peculiarities from the generalities. however, the categorization of valences in the umtm as either affective or cognitive, and either positive or negative, and the differentiation of cognitive valences in personal and non-personal valences has proven to offer better possibilities for the explanation of motivational phenomena than either a onedimensional opposition between intrinsic and more extrinsic types of motivation or a simple addition of different types of subjective values. de brabander et glastra 73 | f l r keypoints intrinsic and extrinsic motivation, respectively subjective value are redefined as affective and cognitive valences valences can be positive (approach motivation) or negative (avoidance motivation) these four types of valences appear to be relatively independent the key sources of valences, autonomy, feasibility, relatedness, and subjective norm, are not equally important in all concrete action contexts in the context of professional development of teachers, autonomy and subjective norm proved affect readiness for action directly references ajzen, i. (1991). the theory of planned behavior. organizational and human decision processes, 50, 179211. doi: 10.1016/0749-5978(91)90020-t ajzen, i., & fishbein, m. (2008). scaling and testing multiplicative combinations in the expectancy–value model of attitudes. journal of applied social psychology, 38, 2222-2247. doi: 10.1111/j.15591816.2008.00389.x apple, m. w., & jungck, s. (1990). "you don't have to be a teacher to teach this unit": teaching, technology and gender in the classroom. american educational research journal 27, 227-251. bakkenes, i. (1996). professional isolation of primary school teachers; a task-specific approach. doctoral dissertation. leiden: dswo press, leiden university. bakkenes, i., de brabander, c. j., & imants, j. g. m. (1999). teacher isolation and communication network analysis in primary schools. educational administration quarterly, 35, 166-202. doi: 10.1177/00131619921968518 ball, s. j. (2003). the teacher’s soul and the terrors of performativity. journal of education policy, 18, 215228. doi: 10.1080/0268093022000043065 ballet, k., & kelchtermans, g. (2008). workload and willingness to change: disentangling the experience of intensification. journal of curriculum studies, 40, 47-67. doi: 10.1080/00220270701516463 bandura, a. (1977). self-efficacy: toward a unifying theory of behavioral change. psychological review, 84, 191-215. doi: 10.1037//0033-295x.84.2.191 bandura, a. (1986). social foundations of thought and action. englewood cliffs, nj: prentice-hall. bandura, a. (1992). on rectifying the comparative anatomy of perceived control: comments on ‘cognates of personal control’. applied and preventive psychology, 1, 121-126. doi: 10.1016/s09621849(05)80153-2 bandura, a. (1997). self-efficacy: the exercise of control. new york: w. h. freeman & co. bandura, a. (2001). social cognitive theory: an agentic perspective. annual review of psychology, 52, 126. doi: 10.1146/annurev.psych.52.1.1 battle, a., & wigfield, a. (2003). college women’s value orientations toward family, career, and graduate school. journal of vocational behavior, 62, 56–75. doi:10.1016/s0001-8791(02)00037-4 bentler, p. m. (2008). eqs 6 structural equations program manual. encino, ca: multivariate software, inc. de brabander et glastra 74 | f l r boekaerts, m., van nuland, h.j.c., & martens, r.l. (2010). perspectives on motivation: what mechanism energise students’ behaviour in the classroom. in k. littleton, c. wood, & j. kleine staarman (eds.), international handbook of psychology in education (pp. 535-568). bingley, uk: emerald group publishing limited. borko, h., elliott, r., and uchiyama, k. (2002). professional development: a key to kentucky's educational reform effort. teaching and teacher education, 18, 969-987. doi: 10.1016/s0742-051x(02)00054-9 carver, c. s. (2006). approach, avoidance, and the self-regulation of affect and action. motivation and emotion, 30, 105-110. doi: 10.1007/s11031-006-9044-7 clark, r., livingstone, d. w., & smaller, h., eds. (2012). teacher learning and power in the knowledge society. rotterdam, netherlands: sense. csikszentmihalyi, m. (1990). flow: the psychology of optimal experience. new york: harper & row. deci, e. l., koestner, r., & ryan, r. m. (1999). a meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. psychological bulletin, 125, 627–668. doi: 10.1037/00332909.125.6.627 deci, e. l., koestner, r., & ryan, r. m. (2001). extrinsic rewards and intrinsic motivation in education: reconsidered once again. review of educational research, 71, 1–27. doi: 10.3102/00346543071001001 deci, e. l., & ryan, r. m. (2000). the “what” and “why” of goal pursuits: human needs and the selfdetermination of behaviour. psychological inquiry, 11, 227-268. 10.1207/s15327965pli1104_01 de brabander, c. j., & martens, r. l. (2014). towards a unified theory of task-specific motivation. educational research review, 11, 27-44. doi: 10.1016/j.edurev.2013.11.001 de brabander, c. j., & martens, r. l. (2018). empirical exploration of a unified model of task-specific motivation. psychology, 9, 540-560. doi: 10.4236/psych.2018.94033 de brabander, c. j., rozendaal, j. s., & martens, r. l. (2009). investigating efficacy expectancy as criterion for comparison of teacherversus student-regulated learning in higher education. learning environments research, 12, 191-207. doi: 10.1007/s10984-009-9062-y devos, g, tuytens, m, & hulpia, h. (2014). teachers’ organizational commitment: examining the mediating effects of distributed leadership. the american journal of education, 20, 205-231. doi: 10.1086/674370 eccles, j., & wigfield, a. (1995). in the mind of the actor: the structure of adolescents’ achievement task values and expectancy-related beliefs. personality and social psychology bulletin, 21, 215-225. doi: 10.1177/0146167295213003 eccles, j., wigfield, a., harold, r. d., & blumenfeld, p. (1993). age and gender differences in children’s selfand task perceptions during elementary school. child development, 64, 830-847. doi: 10.2307/1131221 elliot, a. j. (2006). the hierarchical model of approach-avoidance motivation. motivation and emotion, 30, 111-116. doi: 10.1007/s11031-006-9028-7 elliot, a. j., & church, m. a. (1997). a hierarchical model of approach and avoidance achievement motivation. journal of personality and social psychology, 72, 218-232. doi: 10.1037/00223514.72.1.218 elliot, a. j., & covington, m. v. (2001). approach and avoidance motivation. educational psychology review, 13, 73-92. doi: 10.1023/a:1009009018235 epstein, s. (1994). integration of the cognitive and the psychodynamic unconscious. american psychologist, 49, 709-724. doi: 10.1037//0003-066x.49.8.709 de brabander et glastra 75 | f l r finney, s. j., & distefano, c. (2013). nonnormal and categorical data in structural equation modeling. in: g. r. hancock & r. o. mueller (eds.), structural equation modeling: a second course, 2nd edition (pp. 439-492). charlotte, nc: information age publishing. flake, j. k., barron, k. e., hulleman, c., mccoach, b. d., & welsh, m. e. (2015). measuring cost: the forgotten component of expectancy-value theory. contemporary educational psychology, 41, 232-244. doi: 10.1016/j.cedpsych.2015.03.002 gleeson, d., & husbands, c. (2003). modernizing schooling through performance management: a critical appraisal. journal of education policy, 18, 499-511. doi: 10.1080/0268093032000124866 hardy, i. (2010) critiquing teacher professional development: teacher learning within the field of teachers' work. critical studies in education, 51, 71-84. doi: 10.1080/17508480903450232 hardy, i. & lingard, b. (2008). teacher professional development as an effect of policy and practice: a bourdieuian analysis. journal of education policy, 23, 63-80. doi: 10.1080/02680930701754096 hargreaves, a, & dawe, r. (1990). paths of professional development: contrived collegiality, collaborative culture, and the case of peer coaching. teaching and teacher education, 6, 227-241. doi: 10.1016/0742051x(90)90015-w hulpia, h, devos, g, & van keer, h. (2011). the relation between school leadership from a distributed perspective and teachers' organizational commitment: examining the source of the leadership function. educational administration quarterly, 47, 728-771. doi: 10.1177/0013161x11402065 imants, j. g. m., & de brabander, c. j. (1996). teachers' and principals' sense of efficacy in elementary schools. teaching and teacher education, 12, 179-195. doi: 10.1016/0742-051x(95)00053-m kline, r. b. (2011). principles and practice of structural equation modeling: third edition. new york: the guilford press. krapp, a. (2002). structural and dynamic aspects of interest development: theoretical considerations from an ontogenetic perspective. learning and instruction, 12, 383-409. doi: 10.1016/s09594752(01)00011-1 krapp, a. (2005). basic needs and the development of interest and intrinsic motivational orientations. learning and instruction, 15, 381-395. doi: 10.1016/j.learninstruc.2005.07.007 leth-steensen, c., & gallitto, e. (2016). testing mediation in structural equation modeling: the effectiveness of the test of joint significance. educational and psychological measurement, 76, 339351. doi: 10.1177/0013164415593777 little, j. (1990). the persistence of privacy: autonomy and initiative in teachers’ professional relations. teachers college record, 91, 509-536. little, j. w. (1993). teachers' professional development in a climate of educational reform. educational evaluation and policy analysis, 15, 129-151. doi: 10.2307/1164418 locke, t, vulliamy, g, webb, r, & hill, m. (2005). being a ‘professional’ primary school teacher at the beginning of the 21st century: a comparative analysis of primary teacher professionalism in new zealand and england. journal of education policy, 20, 555-581. doi: 10.1080/02680930500221784 lortie, d. (1975). schoolteacher: a sociological study. chicago, il: university of chicago press. milner, h.r. (2013). policy reforms and de-professionalization of teaching. boulder, co: national education policy center. retrieved from http://nepc.colorado.edu/publication/policy-reformsdeprofessionalization. ministerie van ocw. (2012). nota werken in het onderwijs 2012. [memorandum working in the field of education] the hague, the netherlands: ministerie van ocw. http://nepc.colorado.edu/publication/policy-reforms-deprofessionalization http://nepc.colorado.edu/publication/policy-reforms-deprofessionalization de brabander et glastra 76 | f l r onderwijscoöperatie (2016). de staat van de leraar 2016. [the state of the teacher 2016] retreived from http://www.fvov.nl/wp-content/uploads/2016/04/w-20160413-staat_van_de_leraar.pdf penuel, w. r., fishman, b. j., yamaguchi, r., & gallagher, l. p. (2007). what makes professional development effective? strategies that foster curriculum implementation. american educational research journal, 44, 921–958. doi: 10.3102/0002831207308221 reeve, j., nix, g., & hamm, d. (2003). testing models of the experience of self-determination in intrinsic motivation and the conundrum of choice. journal of educational psychology, 95, 375-392. doi: 10.1037/0022-0663.95.2.375 research voor beleid. (2011). tussenmeting convenant leerkracht 2011 [intermediate assessment teacher covenant 2011]. zoetermeer, the netherlands: research voor beleid. ryan, r. m., & deci, e. l. (2000a). self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. the american psychologist, 55, 68-78. doi: 10.1037/0003066x.55.1.68 ryan, r. m., & deci, e. l. (2000b). intrinsic and extrinsic motivations: classic definitions and new directions. contemporary educational psychology, 25, 54-67. doi: 10.1006/ceps.1999.1020 silins, h. c., mulford, w. r., & zarins, s. (2002). organizational learning and school change. educational administration quarterly, 38, 613-642. doi: 10.1177/0013161x02239641 schunk, d. h., & pajares, f. (2010). self-efficacy beliefs. in p. peterson, e. baker, & b. mcbarry, eds., international encyclopedia of education (third edition), pp. 668-672. oxford: elsevier. doi: 10.1016/b978-0-08-044894-7.00620-5 schunk, d. h., pintrich, p. r., & meece, j. l. (2008). motivation in education. theory, research, and applications. upper saddle river, new jersey: pearson education, inc. smit, k., de brabander, c. j., & martens, r. l. (2014). student-centred and teacher-centred learning environment in pre-vocational secondary education: psychological needs, and motivation. scandinavian journal of educational research, 58, 695-712. doi: 10.1080/00313831.2013.821090 starkey, l., yates, a., meyer, l. h., hall, c., taylor, m., stevens, s., & toia, r. (2009). professional development design: embedding educational reform in new zealand. teaching and teacher education, 25, 181-189. doi: 10.1016/j.tate.2008.08.007 swann, m., mcintyre, d., pell t., hargreaves, l. & cunningham, m. (2010). teachers' conceptions of teacher professionalism in england in 2003 and 2006. british educational research journal, 36, 549 571. doi: 10.1080/01411920903018083 wigfield, a. & cambria, j. (2010). expectancy-value theory: retrospective and prospective. in s. a. karabenick & t. c. urdan, eds., the decade ahead: theoretical perspectives on motivation and achievement. advances in motivation and achievement, volume 16, part a, pp 35-70. bingley, uk: emerald group publishing. doi: 10.1108/s0749-7423(2010)000016a005 wigfield, a., & eccles, j. s. (2000). expectancy–value theory of achievement motivation. contemporary educational psychology, 25, 68-81. doi: 10.1006/ceps.1999.1015 2.2. positive and negative valences 2.3. task-specific antecedents 2.4. direct or mediated effects brom et al publication frontline learning research vol. 7 no 3 (2019) 64 90 issn 2295-3159 it’s better to enjoy learning than playing: motivational effects of an educational live action role-playing game cyril broma, viktor dobrovolný a, b filip děchtěrenkob, c tereza stárková a, b edita bromová a, d afaculty of mathematics and physics, charles university, czech republic b faculty of arts, charles university, czech republic c institute of psychology, the czech academy of sciences, czech republic b film and tv school, academy of performing arts in prague, czech republic article received 17 february 2019/ revised 20 june/ accepted 19 july / available online 16 august abstract game-based learning is supposed to motivate learners. however, to what degree does motivation driven by interest in playing an instructional game affect learning outcomes compared to motivation driven by interest in the very learning process? this is not known. in this study with a unique design and intervention, young adults (n = 128; a heterogeneous sample) learned how to control an electro-mechanical device in a 40-minute-long learning session integrated into a 2-hour-long educational live action role-playing game (edu-larp). edu-larps are supposedly engaging games where players take part in team role-playing by physically enacting characters in a fictional universe. in our edu-larp, players had to understand how the to-be-learned device worked in order to win the game. departing from typical game-based learning research, learningand playing-related variables were assessed for each learner separately (i.e., a within-subject design). affective-motivational factors related to playing (rather than learning) predicted learning outcomes in a positive, but considerably weaker, way compared to learning-related, affective-motivational factors. developed interest in larp-like games was primarily related to enjoying the game rather than better learning outcomes; whereas, developed interest in the instructional domain was primarily related to enjoyment of learning and better learning outcomes. overall, autonomous motivation to play was connected to higher learning outcomes, but this connection was weak. keywords: game-based learning; autonomous motivation; learning-outcomes; developed interest; live action role-playing games; edu-larps info corresponding author brom@ksvi.mff.cuni.cz doi: 10.14786/flr.v7i3.459 1. introduction game-based learning experiences are widely supposed to boost the autonomous motivation of learners. autonomous motivation refers to doing an activity for its own sake or for its perceived personal importance (vansteenkiste et al., 2009). higher autonomous motivation is supposed to increase cognitive engagement and thereby facilitate learning (e.g., moreno, 2005; see also cordova & lepper, 1996; grolnick & ryan, 1987). however, to what degree does motivation driven by interest in playing an educational game impact learning outcomes? is the influence of this motivation on learning outcomes higher, lower or the same compared to how much motivation driven by interest in the instructional topic impacts learning? this research question is addressed in this study. participants played a 2-hour-long educational live action role-playing game (edu-larp) with a sci-fi plot, that included a 40-minute-long learning session organically integrated into the middle of game play. edu-larps are supposedly engaging game-based learning experiences that emphasize team role-play and an element of physical enactment (bowman, 2014; bowman & standiford, 2015; hyltoft, 2010; montola, 2008; vanek & peterson, 2016). to win the game, players had to learn (in the integrated learning session) how to control an electro-mechanical device. the device was fictitious, having meaning only in the game context (to minimize the influence of prior knowledge). from an educational perspective, learning how the device worked represented general mental models acquisition in science, technology, and engineering contexts. participants included young adults having varying degrees of developed interest in sci-fi larp-like games and developed interest in electro-mechanics/ict. developed interest refers to a relatively enduring predisposition to re-engage with particular types of content over time (hidi & renninger, 2006). the edu-larp was designed to appeal primarily to those with the former developed interest. the learning session and the device were designed to engage primarily those with the latter developed interest. the reason for this was to create two distinct drives of autonomous motivation: one being interest in playing and the other interest in the instructional domain. enjoyable instructional experiences may not only increase autonomous motivation; they can also cause greater cognitive load (e.g., mayer, 2014; rey, 2012; um et al., 2012). cognitive load refers to the amount of mental activity imposed by the educational experience on working memory (sweller, ayres, & kalyuga, 2011). high cognitive load, especially if triggered by too complex or distractive elements in the learning environment, has opposite effects on learning processes than does motivation. it can overwhelm limited working memory resources and thereby hamper learning (e.g., sweller, ayres, & kalyuga, 2011; see also um et al., 2012). in this study, we measured autonomous motivation as well as cognitive load variables. we measured them twice during the game: with respect to game play (i.e., before the learning session) and with respect to the learning (i.e., after the learning session). we contrasted how motivation to play versus motivation to learn, and game-engendered versus learning-engendered cognitive load, influenced learning outcomes (i.e., within-subject design). learning outcomes were measured both after the edu-larp and a month later. we also examined whether developed interest in the instructional domain, and in the game, affected the two autonomous motivations, two cognitive loads, and learning outcomes. finally, we explored whether the (possible) relationship between the two developed interests and learning outcomes was mediated by affective-motivational or cognitive load variables. 2. study background 2.1 games and learning there are many definitions of games. in this paper, we understand game in the terms of jull’s definition (2003): as a rule-based system with variable, quantifiable outcomes, in which actors (i.e., players) have the possibility to “attach” themselves intrinsically to different outcomes of the playing process. at the same time, they can influence the game state while working toward the goal (with or without real-life consequences). players must make an effort to influence the game state. one of the key drivers for exerting such effort is autonomous motivation to play the game. various types of learning exist; for instance, skill training, learning of facts, or mental models acquisition. in this study, we focus on mental models acquisition, as understood in constructivist frameworks (e.g., mayer, 2009). mental models are knowledge structures that represent processes and/or systems and that enable drawing inferences about the processes/systems. these structures are built in learners’ minds within the context of knowledge structures previously acquired. the process is not automatic: learners have to exert mental effort to do so. one of the key premises of game-based learning is that motivation to play the game can positively influence learning processes and thereby enhance learning outcomes. how exactly motivation to play should do this (and whether it can actually do this, and if so, how strongly) has rarely been examined in game-based learning literature. this is discussed in detail in the next section. for example, when can effort invested into playing (driven by autonomous motivation to play) be transferred to effort invested into learning? this is not known. the present study examines one possible mechanism through which motivation derived from playing can enhance learning, as outlined in section 2.4. 2.2 game-based learning and edu-larps the game-based learning field has been dominated by digital games. however, other approaches, such as edu-larps, are also becoming popular (see bowman, 2014). as said above, the motivational potential of game-based learning has been generally assumed, but evidence substantiating it is limited. for example, digital games slightly enhance learning compared to traditional instructional approaches (e.g., meta-analyzed in clark, tanner-smith, & killingsworth, 2016; wouters, van nimwegen, van oostendorp, & van der spek, 2013), but the extent to which they do so through affective-motivational factors is unclear. many studies included in the meta-analyses have not researched affective-motivational factors (see sitzmann, 2011; wouters et al., 2013). the affective-motivational dimension has thus been omitted from most meta-analyses. wouters and colleagues, who did examine the effects of games in this dimension, reported that instructional games are more motivational compared to traditional types of education, but the difference only approached significance. 1 narrative reviews of game-based learning literature (e.g., bowman, 2014; boyle et al., 2016; jabbar & felicia, 2015) also did not provide information about how much game-derived motivation influences learning outcomes. claims about the possible influence of game-derived motivation on learning outcomes are substantiated by experimental studies that have examined the effects of individual game design elements (see, e.g., clark et al., 2016; wouters & oostendorp, 2017). some of these elements have been shown to elevate affective-motivational factors as well as enhance learning outcomes. these include personalization and choice (cordova & lepper, 1996), intrinsic integration of the learning content with game mechanics (habgood & ainsworth, 2011), and team role-playing activities (brom, šisler, slussareff, selmbacherová, & hlávka, 2016). however, correlational studies examining the strengths of associations between game-derived affective-motivational factors and learning outcomes have offered a mixed picture (e.g., a negative influence: iten & petko, 2014; a positive influence: sabourin & lester, 2014). the reason behind this ambiguity could be that the motivational–learning correlations have been confounded by contextual factors, most notably, by different levels of cognitive load caused by different game designs. even in comparative studies, experimental and control conditions could differ in the levels of cognitive load (often uncontrolled) imposed upon learners. to take the next step in answering the question on the strength of game-derived motivation’s learning impact, it would be useful if motivational effects triggered by an educational game were contrasted with effects triggered by an unquestioned, robust, “baseline” motivational factor while all participants were undergoing the same intervention. this would make levels of cognitive load dependent on differences between individuals rather than between contextual factors. what should this robust “baseline” be? high interest in a learning domain, and/or in an instructional topic, is straightforwardly implied by theories of motivation and interest in enhancing learning processes (e.g., see eccles & wigfield, 2002; hidi & renninger, 2006; keller, 2010). the positive effects of this interest on motivation to learn and learning outcomes have been repeatedly demonstrated (e.g., brom et al., 2017; fulmer, d’mello, strain & graesser, 2015; schiefele & krapp, 1996; schiefele, 1999). therefore, we used motivation driven by interest in the learning domain as our “baseline”. we contrasted effects of the “baseline” motivation on learning outcomes with the effects of motivation to play the game. the latter motivation was triggered using one of the approaches having been previously shown to be instrumental in doing this: team role-play activities in an edu-larp. edu-larps derive mostly from leisure time social role-playing games (bowman, 2014). social role-playing games emphasize players’ interaction, which distinguishes them from single-player role-playing games. social role-playing games are organized around a modeled scenario/narrative. in edu-larps, players enact this scenario physically, unlike in table-top and digital role-playing games. edu-larps present an old educational approach (kot, 2012; vanek & peterson, 2016), which recently saw a surge in interest (see bowman, 2014). they have been utilized for a range of curricular objectives and they are supposed to have various instructional benefits (bowman, 2014; hyltoft, 2008). for instance, they seem to be useful when a model scenario is to be re-enacted: such as for skills-training (e.g., hayden, smiley, alexander, kardong-edgren, & jeffries, 2014) or understanding complex, socio-historical relationships (e.g., brom et al., 2016; mochocki, 2014). motivational potential is yet another possible advantage of edu-larps (e.g., bowman & standiford, 2015; vanek & peterson, 2016). as in the case of digital game-based learning, researchers are still developing empirical evidence on edu-larps’ effectiveness (bowman, 2014). the key evidence substantiating the claim about the motivational potential of edu-larps comes from the experimental study with a large sample by brom and colleagues (2016), who demonstrated that learning outcomes and affective-motivational variables were enhanced when high school students learned through an edu-larp compared to a discussion-based control without the role-playing game element. the affective-motivational variables partially mediated the game’s positive effect on learning outcomes. supplementary evidence for the motivational potential of edu-larps comes from anecdotal reports (see bowman, 2014; bowman & standiford, 2015). 2.3 individual differences in developed interest different learners are interested in different things. learners’ interests may influence how much they will be motivated to study a particular topic or learn using a game-based approach. for example, not all learners are equally motivated by edu-larps. some learners do not want to get involved in role-play (vanek & peterson, 2016); especially those with low prior experience with this game format (mochocki, 2014). learners with applied study backgrounds found an edu-larp approach more appealing than learners with social sciences backgrounds (brummel et al., 2010). whereas some learners enjoyed an edu-larp experience, others were stressed by the necessity to interact socially during the game (brom et al., 2014). on a theoretical level, these ideas can be embraced using the notion of well-developed individual interest (called developed interest here for brevity). according to the four phase model of interest development (hidi & renninger, 2006), this is the most enduring form of interest: a relatively stable predisposition to re-engage with particular types of content. two developed interests are important here. the first one is interest in sci-fi larps and similar games and game-like experiences (hereafter also called gamer scores). the second one is developed interest in the instructional domain/topic, i.e., ict/electro-physics (hereafter also called techie scores). 2.4 autonomous motivation and self determination theory how can gamer scores and techie scores be theoretically connected to motivations to play and/or to learn? how can these motivations be theoretically linked to learning outcomes? (we note that we do not focus here on leveraging general school motivation, but on motivation and learning outcomes related to a particular learning experience.) a useful way of organizing different forms of motivation provides a framework, which differentiates between autonomous and controlled motivations (ryan et al., 2006; vansteenkiste et al., 2009). this framework is based on self-determination theory (deci & ryan, 1985). autonomous motivation, which is of present interest, is characterized by an internal perceived locus of causality (decharms, 1968; ryan & deci, 2000): learners perceive this motivation as originating within themselves. (controlled motivation is characterized by an external perceived locus of causality.) autonomous motivation is the desired type of motivation (deci & ryan, 2008), because it is linked to several advantages: including better learning outcomes (e.g., cordova & lepper, 1996; grolnick & ryan, 1987; vansteenkiste et al., 2005; see also schiefele, 1999; vansteenkiste et al., 2009). we share this assumption here (general prediction). within self-determination theory, autonomous motivation has two subcomponents (vansteenkiste et al., 2009; see also ryan & deci, 2000): intrinsic motivation that refers to doing an activity for its own sake (because it is inherently enjoyable) and identified regulation, a form of extrinsic motivation that refers to doing an activity because of its perceived importance. self-determination theory maintains that autonomous motivation is fostered when learning environments facilitate learner satisfaction of needs for autonomy, competence, and relatedness (deci & ryan, 1985; vansteenkiste et al., 2009). as concerns intrinsic motivation to learn, the need for competence would be satisfied more often among persons with a high developed interest in the learning domain (i.e., techies) than for those with a low developed interest therein (i.e., non-techies). the reason is that techies would typically feel more competent in solving the learning task than non-techies. therefore, self-determination theory predicts that techie scores will positively relate to intrinsic motivation for learning how an electro-mechanical device works (prediction sdt1a; sdt = self-determination theory). consequently, the techie scores will positively relate to learning outcomes (prediction sdt1b). based on self-determination theory, intrinsic motivation to play would be hampered when a game undermines one of the above needs: it will be lower for those who find it hard to play an edu-larp (competence) or who feel uncomfortable during role-playing/social interaction (relatedness). intrinsic motivation to play will thus positively relate to the gamer score (prediction sdt2a). as concerns identified regulation to learn, the following applies: high autonomous motivation to play, presumably more prevalent among gamers, can be transferred to identified regulation to learn, because players motivated to play may invest more into learning (they feel it is important as part of the game). consequently, the gamer score should relate positively to learning gains (prediction sdt2b). however, the gamer scores → learning outcomes link may be weaker than the techie scores → learning outcomes link, because it is not guaranteed that motivation to play would project to identified regulation to learn for all participants (prediction sdt3). 2.5 distraction and cognitive load theory cognitive load theory (sweller, ayres, & kalyuga, 2011) is a theoretical framework based on a model of human cognitive architecture, which enables educational designers to construct instructionally efficient learning environments. its key assumptions are that working memory has limited capacity and duration; whereas, long-term memory serves for permanent storage with unlimited capacity and duration. during learning, incoming information is first represented in working memory and eventually integrated with pre-existing knowledge structures in long-term memory. these knowledge structures also organize temporary representations in working memory: more advanced structures enable the representation of incoming information using fewer information elements. following its recent adjustment (kalyuga, 2011), cognitive load theory posits two types of working memory load: intrinsic and extraneous. learners must allocate cognitive resources to deal with these loads. if they fail to do so, or if total load overwhelms working memory resources, learning is hampered. intrinsic load is imposed on learners by the complexity of the learning task. intrinsic load is essential for comprehending the learning message: dealing with it results in learning. in determining the task’s complexity, one has to consider learner’s prior knowledge (i.e., knowledge structures in long-term memory prior to learning). what is complex for a novice may not be complex for an expert. in this study, learners learn to operate a fictitious electro-mechanical device. prior knowledge of the device is null for everyone, but techies will possess high-quality knowledge structures concerning certain general electro-mechanical concepts; e.g., “electrical signal”. these structures will enable them to represent information in their working memory more efficiently. within cognitive load theory, this means techies will have lower intrinsic load (prediction clt1a; clt = cognitive load theory). therefore, it is less likely that techies’ working memory would be overloaded compared to non-techies. consequently, in agreement with prediction sdt1b, techie scores will be positively related to learning outcomes (prediction clt1b). extraneous load is caused by the processing of sub-optimally designed features of learning environments. it should be minimized, because accommodating it depletes cognitive resources that could otherwise aid in dealing with intrinsic load. extraneous load can arise from two sources within a game-based learning environment: a) from the sub-optimal design of the instructional content embedded in the game and b) from game play as such. those with high techie scores will likely cope better with possible sub-optimal design, which adds weight to prediction clt1b. as concerns the game-related source of extraneous load, a portion of players’ cognitive resources will be devoted to thinking about playing the game. these thoughts will deflect learners’ attention away from learning. game-related, but learning-irrelevant, thoughts are likely to be amplified for non-gamers, who do not yet have well-developed game-related schemata/skills. therefore, game-engendered extraneous load will be negatively related to gamer scores (prediction clt2a). because it may cause cognitive overload and hamper learning, gamer scores will relate positively to learning outcomes (prediction clt2b). 3. this study – overview and hypotheses this study examines how much autonomous motivation driven by developed interest in an educational game influences learning outcomes compared to motivation driven by developed interest in the instructional domain (i.e., within-subject design). autonomous motivation is referred to hereafter as motivation for the sake of brevity; the first motivation is referred to as motivation to play and the second one as motivation to learn. the two developed interests are called gamer and techie scores, respectively. the study also investigates how these two scores affect motivation to play, motivation to learn, overall game enjoyment, difficulty of game play (as a proxy variable to game-engendered extraneous cognitive load), cognitive loads engendered by the learning experience, and learning outcomes. directional hypotheses: h1: the techie scores will relate • (h1a) positively to motivation to learn (based on prediction sdt1a); • (h1b) negatively to cognitive loads engendered during learning (prediction clt1a); • (h1c) positively to learning outcomes (predictions sdt1b and clt1b). h2: the gamer scores will relate • (h2a) positively to motivation to play and also overall game enjoyment (prediction sdt2a); • (h2b) negatively to difficulty in playing the game (prediction clt2a); • (h2c) positively to learning outcomes (prediction sdt2b and clt2b). the directional hypotheses concerning influences of the techie and gamer scores are summarized in table 1. there relationships for the remaining pairs of variables (i.e., in columns and rows from table 1) will be explored (exploratory goals e1, e2). with respect to the influences of both motivations on learning outcomes, the following hypotheses will be examined (table 2): h3 • (h3a) motivation to play will be positively related to learning outcomes (prediction sdt2b and general prediction); • (h3b) motivation to learn will be positively related to learning outcomes (general prediction). the link between cognitive loads and learning outcomes will be explored (exploratory goal e3). we put forward these mediation hypotheses (table 3): h4: the relationship between techie scores and learning outcomes will be mediated • (h4a) positively by motivation to learn (prediction sdt1a and sdt1b); • (h4b) negatively by cognitive loads evoked during learning (prediction clt1a and clt1b). h5: the relationship between gamer scores and learning outcomes will be mediated • (h5a) positively by motivation to play and also overall game enjoyment (prediction sdt2a and sdt2b); • (h5b) negatively by difficulty in playing the game (prediction clt2a and clt2b). finally, we will explore (e4) whether the relationships between a) gamer scores and learning outcomes and b) motivation to play and learning outcomes is weaker compared to complementary relationships between (a) techie scores or (b) motivation to learn and learning outcomes (cf. prediction sdt3). table 1 hypotheses/exploratory goals related to techie and gamer scores a indexed by proxy variables. note: + positive relationship expected; – negative relationship expected; ? no expectation. table 2 hypotheses/ exploratory goals related to how variables assessed in situ predict learning outcomes aindexed by proxy variables. note: + positive relationship expected; ? no expectation. table 3 hypotheses/exploratory goals related to mediation aindexed by proxy variables. note: expected relationship: + positive; – negative. 4. methods 4.1 participants participants were recruited from university pools and via the facebook pages of larp/sci-fi communities and using short-term job advertisement servers. we emphasized that we were seeking both novice and seasoned larp players. participants received financial compensation (400 czk, ~15 eur). prospective participants completed an online questionnaire, which provided demographic data (mage = 24.7; sdage = 3.72) and data on prior larp-related experience. we invited selected participants for one of 11 game runs, such that people with both low and high prior larp-related experience and with diverse study/employment backgrounds (see suppl. mat. a) participated in every run (10-13 participants per run). ultimately, 128 participants were included in the analysis. two additional participants were excluded due to health issues. seventeen participants were excluded from the analyses of delayed learning outcomes data because they did not attend delayed testing session. the sample was heterogeneous with respect to two techie and gamer scores (figure 1). these two variables also did not correlate (r = .06); i.e., we recruited techie gamers, techie non-gamers, non-techie gamers, and non-techie non-gamers. figure 1. participants’ techie and gamer scores. 4.2 materials edu-larp. we designed the larp as a 2-hour-long game that can teach a scientific topic by means of an embedded learning session. the larp was a sci-fi space opera with a plot created by a seasoned larp script writer. the story started in the midst of a journey on a generation spaceship. the players took on roles of technical school students therein. certain events triggered a mutiny, during which the fighting parties damaged input cables to a device for controlling correction thrusters (i.e., sideways-pointed motors used to make small corrective movements when a spaceship is already in space). this left the ship on a collision course with an asteroid. only one access route to the device remained passable: it led from the technical school via an escape corridor. the players had to learn how the device works (figure 2, 3), locate it, set it to manual control, and tweak its cables to avoid hitting the asteroid and save the ship’s population. the complexity of roles was determined during pilot experiments. the roles were less complex than in a typical larp for seasoned players, but still relatively complex for novice players (i.e., a compromise). figure 2. a: model of the device available during the learning session. b: rewired cabling. c: the actual device. d: schematic drawing of device. figure 3. a: players discussing game options. b: players interacting during the learning session. c: players rewiring cabling on a model. in the middle of the game, the teacher (game master) initiated a 40-minute-long learning session, during which he taught the students how to operate the device (figure 3b, c). the plot was designed so that all players were motivated to learn this (see suppl. mat. b for details). the players worked with functional models of the device. they did not have access to these models outside the session, and the teacher never commented on the device’s functioning, except during this session. a clear beginning and end point for the session also enabled the teacher to administer questionnaires during the game at well-defined moments. toward the end of the game, the players gained access to the actual device. they had 9 minutes to tweak it in order to correct the ship’s course. teaching session. this was a 40-minute-long frontal lecture with slides interspersed with teacher-guided hands-on-practice segments. during the lecture, players stayed in their game roles the whole time. each player was given the device’s model and schematic drawing (figure 2a, 2d). this was the first time they saw the model/device. to solve the final game task and win the game, players had to understand how the device works, i.e., acquire its mental model (mere superficial memorization was insufficient). the model and the learning task. the key part of the model/device, the so-called control calculator a) controlled whether the ship’s course correction could be made given the current status of the correction thrusters and b) converted the course correction given in angular units to the force of the thrusters. input wiring to the calculator relayed a signal from the deck carrying the correction command and status reports from the thrusters. output wiring carried signals to the thrusters with force data and a command (either to perform the maneuver or test if it was possible to carry out the maneuver). the device featured a manual control which could partially override commands from the deck. with a sufficient level of understanding, one could re-wire the cabling (figure 2b, 3c); by-passing certain calculations and/or changing the partial override mode to full override mode (see suppl. mat. c for further details). it was necessary to do these steps in solving the final task. the actual device and the final task. the device was similar to the model (figure 2c): but larger. it had a similar layout, but different graphics. it contained a timer that counted down the time to the ship’s final maneuver. during this time (9 minutes), the players, as a group, had to tweak the cabling and set the ship’s new course. this was a near transfer task with respect to the tasks assigned in the learning session. 4.3 measurements we faced a challenge because instruments assessing relevant constructs in the context of edu-larps were lacking. also, the instruments needed to be short; especially those to be used during the game. therefore, we adjusted several instruments from neighboring research fields, established face validity, and fine-tuned the questions during pilot experiments. all questions and scales are detailed in suppl. mat. d. 4.3.1 participant variables demographic data. when expressing interest in participating in this research, participants reported online their gender, age, prior larp-related experience, study background, and possible employment. prior larp-related experience was measured using six questions we developed, which assessed participants’ experience with various types of larp-like games. techie scores. this variable should reflect participants’ developed domain interest. such interest was, in the present case, related to developed interests in ict and electro-physics. it was also related to participants’ study types, as study type can generally be assumed to reflect the person’s interests. therefore, the techie score was computed as a weighted sum of these two interests and a score assigned based on the participant’s study type (see suppl. mat. e for the exact equation). developed interest in electro-physics (4 items; α = .87) and in ict (4 items; α = .89) was assessed based on work done by renninger and schofield (2014). a score for the participant’s study type was assigned based on a rubric detailed in suppl. mat. e. ict and electro-physics developed interests were strongly inter-correlated (r = .68) and moderately-to-strongly correlated with the study type score (r = .42, .44). gamer scores. this variable should reflect participants’ developed interest in the type of games exemplified by our intervention; i.e., a sci-fi larp. it was computed as a weighted sum of participants’ prior larp-related experience (as this reflects voluntary experience with larp-like games) and developed sci-fi interest (α = .73; see suppl. mat. e for details). prior larp-related experience was reported online when participants expressed interest in our research (see above). sci-fi developed interest was assessed similarly to the ict/electro-physics developed interest (i.e., based on renninger & schofield, 2014) (5 items; α = .95). 4.3.2 dependent variables autonomous motivation: positive affect, flow, and learning enjoyment. motivations to play and to learn were examined separately. because it would be difficult for learners to distinguish between these two motivations after the game ended, we assessed them in situ: at appropriate moments during the game play but without disrupting the play (i.e., using “gamified” questionnaires). we measured motivation to play through two proxy variables: flow and generalized positive affect (referred to hereafter as positive affect). we measured motivation to learn through three proxy variables: flow, positive affect, and learning enjoyment. positive affect is related to various positively-valenced, activating feelings (e.g., excitation, activity). we measured it using positive and negative affect schedule (i.e., panas; watson et al., 1988) (α = .82 – .93).2 flow refers to pleasant absorption of an activity that one takes part in (csikszentmihalyi, 1975). we measured it using three items from the flow short scale (rheinberg et al., 2003) (α = .85 – .89). learning enjoyment was assessed using two questions from the interest/enjoyment subscale of the intrinsic motivation inventory (mcauley et al., 1989) (r = .88). only subsets of questions were used for brevity. learning-engendered cognitive load. in addition to learning-related motivation, we measured learning-engendered intrinsic load (2 items; r = .82) and extraneous load (3 items; α = .78) using questions adopted from the questionnaires by leppink and colleagues (2014) and naismith and colleagues (2015). overall game difficulty, overall game enjoyment. as a proxy variable to game-engendered extraneous load, overall game difficulty was measured using three items we created (α = .84). overall game enjoyment was assessed with 10 questions we created (α = .89); based on motivation/enjoyment items from other questionnaires (e.g., mcauley et al., 1989; schraw et al., 1995). items were tailored for the specifics of larps. retention test. it had one question: “draw a diagram of the device for controlling the ship’s correction thrusters showing all elements of the device”. a point was awarded for correctly drawing an element, for correctly positioning it with respect to other elements, for correctly naming it, and for correctly drawing a cable crossing (scale: 0 – 80). two independent raters scored the answers with a nearly perfect agreement (immediate: weighted cohen’s κ = .995; delayed: κ = .999; cohen, 1986). transfer test. we developed two complementary versions of the transfer test (for immediate versus delayed testing; counterbalanced across participants). one version had four and the other five open-ended questions, e.g., “imagine that the spaceship does not have three correction thrusters, but four instead. what changes would you have to make to the device controlling the correction thrusters in order for the device to function with four thrusters?”. participants were awarded 1 point for each correct solution or 0.25 or 0.5 points for a partially-correct solution (scales: 0 – 26 and 0 – 27, respectively). two raters scored the answers with a substantial agreement (immediate: κ = .977; delayed: κ = .968). raters’ scores were averaged for the subsequent analysis. prior to averaging, transfer test scores were z-transformed for each version of the test to obtain comparable values. 4.4 procedure participants received general larp rules, a description of the setting, and brief descriptions of all roles in advance. they selected online which roles they would prefer to play. upon arrival, participants filled in the initial questionnaire (see figure 4 for the experimental schedule). afterwards, a warm-up period started: the larp rules were recapitulated and participants could decorate the rooms with supplied thematic set pieces. participants were then assigned roles (based on the preferences they had previously expressed), read their descriptions, and introduced their roles to fellow players. figure 4. experimental schedule. afterwards, the game started. the teaching session started about an hour into the game. before it began, the teacher distributed the game-related flow and panas questionnaires (with a cover story about the ship’s supreme inspectorate evaluating the quality of his teaching). after the lecture, the teacher distributed the questionnaire that yielded post-learning flow, positive affect, enjoyment, and intrinsic/extraneous cognitive load data. after roughly 20 minutes, players eventually found access to the escape corridor and located the device, wherein they solved the final game task. participants then filled in the final questionnaire, which primarily yielded overall game enjoyment and difficulty assessments. afterwards, retention and transfer tests were distributed. finally, a game debriefing and an interview were organized. about three weeks after the larp, participants arrived for the delayed testing session. they were given retention and transfer tests and developed interest questionnaires. 4.5 data treatment post-learning positive affect and flow tapped at-the-moment experienced affective-motivational states. these states could be influenced not only by learning, but also by game play that preceded the learning session (this was not the case of enjoyment/cognitive load scales, because the respective questions referred to the learning session as such, see suppl. mat. d). we wanted to use, as proxies to motivation to learn, positive affect/flow-derived variables that satisfied two requirements: they were a) related to pre-/post-learning change in positive affect/flow and b) independent of pre-learning positive affect/flow. therefore, learning-related positive affect/flow were computed as pre-/post-learning positive affect/flow residual differences (by regressing the post-learning positive affect/flow on the pre-learning positive affect/flow). 5. results the means and averages of the dependent variables are included in suppl. mat. a. the correlation matrix is also provided therein. 5.1 techie scores we used multiple linear regressions with two independent factors: gamer and techie score. when entered together with gamer scores into the models, techie scores modestly predicted motivation to learn (table 4), so hypothesis 1a was supported. techie scores strongly predicted cognitive load induced by learning (in the negative direction) and learning outcomes, so hypotheses h1b and h1c were also supported. as concerns exploratory goals e1a and e1b, techie scores were unrelated to game-induced flow and positive affect. however, they were modestly related to game enjoyment. this can be explained by the fact that game enjoyment was measured after the larp ended, so it was arguably also influenced by liking the learning session (unlike game-induced flow/positive affect measured before the learning session started). techie scores were unrelated to perceived game difficulty. we conclude that techies liked the learning session more (compared to non-techies), it was easier for them to learn, and they learned better. however, playing a larp was not more motivating for them. 5.2 gamer scores gamer scores strongly predicted motivation to play, overall game enjoyment, and (in the negative direction) game difficulty (table 4). hypotheses h2a and h2b were thus supported. gamer scores were modestly related to learning gains (except for immediate retention). hypothesis h2c was thus partially supported. as concerns exploratory goals e2a and e2b, gamer scores were unrelated to motivation to learn and learning-engendered cognitive loads. we conclude that participants with higher gamer scores were relatively more motivated to play the larp and playing was easier for them. they also learned slightly better. however, gamer scores were not showed to be connected to motivation to learn and cognitive load. table 4 standardized beta coefficients for a multiple regression model (predictors: techie/gamer scores) note: pa = positive affect. hypotheses: ✓ supported; [✓] partially supported. relationships found: ○ no; + positive; – negative. a pre-post residuum. *p < .05 **p < .01 ***p < .001 5.3 affective-motivational–learning relationship which in situ measured variables predicted learning outcomes? the effects of the following variables were investigated: game-induced positive affect/flow, learning-induced positive affect/flow, learning enjoyment, intrinsic/extraneous load. to facilitate interpretation, we reduced the number of variables using exploratory factor analysis. the kaiser-meyer-olkin adequacy was .65 (above the recommended value .6) and bartlett’s test of sphericity was significant (p < .001). factors were extracted using the ordinary least squares method with varimax rotation. both scree plot analysis and parallel analysis suggested the presence of three factors. these factors corresponded to our three umbrella constructs: motivation to play, motivation to learn, and cognitive load; and the factors were thus labeled so. in our main analysis, we computed regression models with the newly-created factors (all three factors where entered together into each model)3. both motivation factors significantly predicted transfer; however, only motivation to learn predicted also retention (table 5). hypothesis h3b was thus supported and h3a was supported only as concerns transfer. table 5 hypotheses, predictions, and exploratory goals related to the motivation–learning link note: hypothesis: ✓ supported; ( ✓ ) partially supported; x not supported. relationship found: – negative. athe factor to which game-induced positive affect/flow primarily load. bthe factor to which learning enjoyment/positive affect/flow primarily load. cthe factor to which intrinsic/extraneous load primarily load. as concerns exploratory goal e3, cognitive load predicted well all learning outcome variables in the negative direction. as concerns exploratory goal e4, gamer scores as well as motivation to play were weaker predictors of learning outcomes compared to techie scores and motivation to learn (and cognitive load). however, they still played certain roles (tables 4, 5). we conclude that participants with higher motivation to learn and/or with lower cognitive load learned better than those with lower motivation and/or higher cognitive load. also, those motivated to play the game performed somewhat better than those less motivated to play, but only on transfer test tasks. 5.4 mediation analysis we used the package mediation (tingley et al., 2014) for causal mediation analysis. we computed estimates for indirect effect using quasi-bayesian monte carlo simulations (n = 10,000) (preacher & hayes, 2004). the relationship between techie scores and learning outcomes was mediated both by motivation to learn and cognitive load engendered during learning (table 6). therefore, hypotheses h4a and h4b were supported. the relationship between gamer scores and transfer was mediated by overall game enjoyment and, for immediate transfer, marginally mediated by overall game difficulty. no other variable was confirmed as a mediator (p > .106). hypothesis h5a was thus supported only with respect to transfer and overall game enjoyment, and hypothesis h5b with respect to immediate transfer. table 6 mediation analysis note: ci = confidence interval. hypothesis: ✓supported; [✓] partially supported; x not supported. †p < .10 *p < .05 **p < .01 athe fa factors 5.5 interview data we inspected negative evaluations of the larp and whether motivation to learn transformed to identified regulation to play, as gauged by participants. to summarize these results, we split participants into four groups based on median split of transfer (average value of immediate and delayed transfer test scores; i.e., high vs. low transfer) and motivation to play (average value of z-scores from game-induced flow and positive affect; i.e., high vs. low motivation). sample statements from the subgroups’ participants are shown in table 7. generally, qualitative data showed: i) higher learning outcomes were not necessarily connected to excitation from the game; ii) some learners (around 20% of the sample) did not like the approach; iii) some participants (around 15%) claimed the game had a substantial positive motivational effect on their learning. table 7 sample statements during interviews and participant interests aactual range. btwenty-seven participants could not be classified as hard vs. social science students or had partly missing data. 5.6 supplementary results for control purposes, we also conducted a supplementary study, in which we tested how much participants would learn from the same 40-minute-long learning session as embedded in the larp: but outside the game. we recruited 48 participants to match a selected sub-sample from the main study (i.e., a quasi-experimental comparison without randomization). participants were matched based on their study backgrounds (see suppl. mat. g for further details). at the beginning of the learning session, these non-game participants were told that they should imagine they are students on a generation ship (to contextualize the device). the same learning outcomes and autonomous motivation measures were used. positive affect of non-larp learners measured immediately after the narrative introduction, i.e., immediately before the learning session, was significantly lower compared to positive affect of the matched larp participants measured in the game, before the learning session (d = 0.45). this demonstrates a medium effect size motivational advantage of the edu-larp (see suppl. mat. g for descriptive data and analyses). motivation to learn (i.e., measured immediately after the learning session ended) was comparable for both groups of learners. non-larp learners exhibited a faster decline of transfer learning outcomes (immediate – after three weeks) compared to the matched larp participants (d = 0.67). no other significant between-group difference as concerns learning outcomes was found (d = –0.16 – 0.20) (suppl. mat. g). this means that larp players forgot less in terms of conceptual knowledge (moderate-to-strong effect). 6. discussion we investigated how much motivation driven by interest in playing an educational game impacts learning outcomes compared to motivation driven by interest in the instructional domain (while controlling for levels of cognitive load). motivation to play was shown to be related to learning outcomes, but its influence was dwarfed by the effects of the natural motivation to learn the given topic exhibited primarily by participants with developed domain interest. specifically, developed domain interest was clearly related to autonomous motivation to learn and learning-evoked cognitive load, which were clearly related to learning outcomes. however, autonomous motivation to play, exhibited primarily by participants with developed interest in sci-fi larp-like games, was only slightly related to transfer learning outcomes (and unrelated to retention). also, motivation to play did not noticeably mediate the relationship between interest in sci-fi larp-like games and learning outcomes. this pattern of results – that of a relatively strong effect of domain interest on learning outcomes and weaker, sometimes even negative, effect of interestingness of framing the educational message on learning outcomes – is generally consistent with findings from the fields of multimedia learning (e.g., rey, 2012) and hypermedia learning (e.g., moos & marroquin, 2010). this study has demonstrated this pattern in the context of game-based learning, for which large motivational benefits have been envisioned: it is better to enjoy learning than playing. 6.1 contributions 6.1.1 theoretical contributions from the self-determination theory perspective, it is noteworthy that we distinguished between autonomous motivations to learn versus to play, which is rare in game-based learning literature. in agreement with self-determination theory, our results demonstrated that motivation to learn was a better predictor of learning outcomes compared to motivation to play; yet the study also provided provisional evidence suggesting that intrinsic motivation to play can transform to identified regulation to learn. future studies should explore in more detail how motivation derived from playing transfers to identified regulation to learn, as this is one of the key ways how motivation to play can influence learning processes within game-based learning. other ways how motivation to play can impact learning processes should also be considered and examined (e.g., through enhanced self-efficacy). from the cognitive load theory perspective, it was equally important that we distinguished between game-evoked extraneous load and learning-evoked loads. in agreement with cognitive load theory, the results confirmed learning-evoked loads as mediators of the effect of developed domain interest on learning outcomes. complementary meditational analysis showed that perceived game difficulty (a proxy variable to game-engendered extraneous load) only tended to mediate the effect of developed interest in sci-fi larp-like games on immediate transfer. either the distraction from the game was not a big deal for participants, or our measurement did not assess the game-engendered cognitive load well. validated methods for measuring this construct are needed, because the extent to which different types of games or game attributes evoke extraneous cognitive load is a pressing issue. with respect to the four-phase model of interest development (hidi & renninger, 2006), our results corroborated the idea that developed interest in a learning domain is connected to enhanced learning in (at least) two different ways: cognitive and affective-motivational ones. this interest was linked to both lower cognitive load (presumably due to pre-existing, task-fitting schemata in the long-term memory) and higher values of motivation variables. both cognitive load and motivation also mediated the effect of domain interest on learning outcomes. 6.1.2 practical contributions our results showed that edu-larps may enhance learning through positive affective-motivational factors for some learners, but this educational method is not unanimously liked. this study and some prior research (brummel et al., 2010; mochocki, 2014; vanek & peterson, 2016) thus indicate that edu-larps are accepted differently by different learner-types. acceptance-/suitability-related problems should be addressed in future research as well as in applying this educational method in practice (cf. mochocki, 2014; vanek & peterson, 2016). practitioners should also keep in mind that larps take a long time to prepare and complete. on a more positive note, a single edu-larp can have multiple educational objectives at the same time (bowman, 2014). 6.1.3 methodological contributions this work capitalized on the fact that game-related versus learning-related developed interest, affective-motivational, and cognitive load variables were assessed separately. we suggest considering this approach in future game-based learning research, as this can help elucidate complex roles different components of interest, motivation, and cognitive load play in learning processes. 6.2 limitations no work is without limitations. the key thorny issue is the nearly-absolute lack of validated measurements for edu-larp contexts (and beyond, as detailed below). first, developed instruments, such as egameflow (fu et al., 2009), are very long. a construct needs to be assessed with a few questions within a game, and, should learning tests be administered, also after the game (to avoid fatigue). second there is an on-going discussion about how to measure cognitive load and distinguish between different types of load (e.g., leppink et al., 2014; naismith et al., 2015). cognitive load instruments that would be validated in the same way as, for example, panas are lacking. finally, there is no agreed-upon method for measuring developed interests (renninger & pozos-brewer, 2015). we were more satisfied with the gamer score variable, because it was more internally consistent compared to the techie score variable (despite all three subcomponents of the latter variable were theoretically related to developed domain interest). we believe that the pattern of our data is clear enough and consistent with underlying theories to warrant our interpretations of the findings. nevertheless, valid instruments would be useful in future. in an ideal world, this study would have had a control, non-larp, condition and participants would have been randomly assigned to the larp and non-larp conditions. this would enable the contrasting of learning within the game to learning outside the game, whilst the content and the method (i.e., frontal lecture with hands-on practice) would be the same. unfortunately, true randomization is rarely possible in research using edu-larps due to practical and ethical reasons (further detailed in suppl. mat. g). we have focused here on the within-subject comparison part (done within what would be an experimental condition) and have drawn conclusions from this part. we also recruited a quasi-experimental control group, but data on comparing performance of larp learners to non-larp learners should be treated cautiously because of the lack of randomization. 7. conclusions this study contributes to game-based learning literature, but also to the research base on productive enhancement of interest and motivation in academic contexts. its key message is that learners’ developed domain interest and motivation to learn a particular topic contribute more toward enhancing learning outcomes than supposedly appealing, game-based augmentations of the educational message. our results can be directly generalized probably only to games that have a similar level of complexity as our larp. however, there is growing, parallel evidence suggesting that the message above is quite general and concerns many different types of learning environments and materials, such as textbooks or hypermedia. supplementary materials supplementary materials include: supplementary data and analyses (suppl. mat. a, f), detailed description of the edu-larp and procedure (suppl. mat. b), description of the function of the experimental device (suppl. mat. c), questionnaire items (suppl. mat. d), description of developed interest variables (suppl. mat. e), the supplementary study (suppl. mat. g). keypoints educational live action role-playing games (edu-larps) are supposed to enhance learning outcomes by motivating learners. in this study, learners played a 2-hour-long edu-larp with an integrated learning session. we asked to what degree does motivation driven by interest in playing the edu-larp affect learning outcomes compared to learning-driven motivation. learning-driven motivation (rather than playing-driven motivation) predicted learning outcomes; the effects of the latter were positive, but small. developed topic interest (rather than interest in larps) predicted learning outcomes; the effects of the latter were still positive, but small. acknowledgments we thank research assistants who helped to conduct the experiments, most notably: t. zoulová, n. frollová, k. koppová, n. střádalová, and p. šustová. we thank suzanne hidi and k. ann renninger for discussing methods for measuring interest with us. we also thank sarah lynne bowman for discussing this project with us and for commenting on early versions of this manuscript. we thank lucie filipenská for making the videos and all the actors: anna kratochvílová, anežka rusevová, pavol smolárik, and luboš veselý; including actors in pilot videos: břetislav dufek, tereza “tess” kovanicová, and jan kovanic. we also thank david obdržálek, jan hrach, tomáš “jethro” pokorný, and jana stárková for making the devices, and brmlab, a community-run hackerspace in prague, for assistance. this study was primarily funded by czech grant science foundation (ga čr), project nr. 15-14715s. work of f. d. was supported by rvo 68081740 by the czech academy of sciences. table of footnotes 1 the p-value, unreported in the paper, is .076 (pieter wouters; email dating from 16 dec 2013). 2 some questions administered in this study were not analyzed and reported here. examples include negative affect (from panas), which was out of present scope (see suppl. mat. a for descriptive data), and various manipulation check questions (e.g., on initial interest). 3 see suppl. mat. a for factor loadings and supplementary analyzes. references bowman, s. l. (2014). educational live action role-playing games: a secondary literature review. in: the wyrd con companion book (pp. 112-131): wyrd con. bowman, s. l., & standiford, a. (2015). educational larp in the middle school classroom: a mixed method case study. international journal of role-playing, 5, 4-25. boyle, e. a., hainey, t., connolly, t. m., gray, g., earp, j., ott, m., ... & pereira, j. (2016). an update to the systematic literature review of empirical evidence of the impacts and outcomes of computer games and serious games. computers & education, 94, 178-192. doi: 10.1016/j.compedu.2015.11.003 brom, c., buchtová, m., šisler, v., děchtěrenko, f., palme, r., & glenk, l. m. (2014). flow, social interaction anxiety and salivary cortisol responses in serious games: a quasi-experimental study. computers & education, 79, 69-100. doi: 10.1016/j.compedu.2014.07.001 brom, c., šisler, v., slussareff, m., selmbacherová, t., & hlávka, z. (2016). you like it, you learn it: affectivity and learning in competitive social role play gaming. international journal of computer-supported collaborative learning, 11 (3), 313-348. doi: 10.1007/s11412-016-9237-3 brom, c., děchtěrenko, f., frollová, n., stárková, t., bromová, e., & d’mello, s. k. (2017). enjoyment or involvement? affective-motivational mediation during learning from a complex computerized simulation. computers & education, 114, 236-254. doi: 10.1016/j.compedu.2017.07.001 brummel, b. j., gunsalus, c., anderson, k. l., & loui, m. c. (2010). development of role-play scenarios for teaching responsible conduct of research. science and engineering ethics, 16(3), 573-589. doi: 10.1007/s11948-010-9221-7 clark, d. b., tanner-smith, e. e., & killingsworth, s. s. (2016). digital games, design, and learning a systematic review and meta-analysis. review of educational research, 86(1), 79-122. doi: 10.3102/0034654315582065 cordova, d. i., & lepper, m. r. (1996). intrinsic motivation and the process of learning: beneficial effects of contextualization, personalization, and choice. journal of educational psychology, 88 (4), 715-730. doi: 10.1037/0022-0663.88.4.715 csikszentmihalyi, m. (1975). beyond boredom and anxiety: jossey–bass, san francisco, ca. deci, e. l., & ryan, r. m. (1985). intrinsic motivation and self-determination in human behavior. new york: plenum. deci, e. l., & ryan, r. m. (2008). facilitating optimal motivation and psychological well-being across life's domains. canadian psychology, 49(1), 14-23. doi: 10.1037/0708-5591.49.1.14 decharms, r. (1968). personal causation: the internal affective determinants of behavior . new york: academic press. eccles, j. s., & wigfield, a. (2002). motivational beliefs, values, and goals. annual review of psychology, 53(1), 109-132. doi: 10.1146/annurev.psych.53.100901.135153 fu, f.-l., su, r.-c., & yu, s.-c. (2009). egameflow: a scale to measure learners’ enjoyment of e-learning games. computers & education, 52(1), 101-112. doi: 10.1016/j.compedu.2008.07.004 fulmer, s. m., d'mello, s. k., strain, a., & graesser, a. c. (2015). interest-based text preference moderates the effect of text difficulty on engagement and learning. contemporary educational psychology, 41, 98-110. doi: 10.1016/j.cedpsych.2014.12.005 grolnick, w. s., & ryan, r. m. (1987). autonomy in children's learning: an experimental and individual difference investigation. journal of personality and social psychology, 52(5), 890-898. doi: 10.1037/0022-3514.52.5.890 habgood, m. j., & ainsworth, s. e. (2011). motivating children to learn effectively: exploring the value of intrinsic integration in educational games. journal of learning sciences, 20(2), 169-206. doi: 10.1080/10508406.2010.508029 hayden, j. k., smiley, r. a., alexander, m., kardong-edgren, s., & jeffries, p. r. (2014). the ncsbn national simulation study: a longitudinal, randomized, controlled study replacing clinical hours with simulation in prelicensure nursing education. journal of nursing regulation, supplement 5(2), s1-s64. doi: 10.1016/j.ecns.2012.07.070 hidi, s., & renninger, k. a. (2006). the four-phase model of interest development. educational psychologist, 41(2), 111-127. doi: 10.1207/s15326985ep4102_4 hyltoft, m. (2008). the role-players’ school: østerskov efterskole playground worlds: creating and evaluating experiences of role-playing games (pp. 12-25): ropecon ry. hyltoft, m. (2010). four reasons why edu-larp works larp: einblicke (pp. 43-57): zauberfeder verlag. iten, n., & petko, d. (2014). learning with serious games: is fun playing the game a predictor of learning success? british journal of educational technology, 47(1), 151-163. doi: 10.1111/bjet.12226 jabbar, a. i., & felicia, p. (2015). gameplay engagement and learning in game-based learning: a systematic review. review of educational research, 85(4), 740-779. doi: 10.3102/0034654315577210 juul, j. (2003). the game, the player, the world: looking for a heart of gameness. plurais-revista multidisciplinar, 1(2), 248-270. kapp, k. m. (2014). do not use games for “stealth learning”. retrieved from http://karlkapp.com/do-not-use-games-for-stealth-learning/ (acessed 15-06-2019) kalyuga, s. (2011). cognitive load theory: how many types of load does it really need? educational psychology review, 23(1), 1-19. doi: 10.1007/s10648-010-9150-7 keller, j. m. (2010). motivational design for learning and performance: the arcs model approach: new york: springer. kot, y. i. (2012). educational larp: topics for consideration. in: wyrd con companion book (pp. 118-127): wyrd con. leppink, j., paas, f., van gog, t., van der vleuten, c. p., & van merrienboer, j. j. (2014). effects of pairs of problems and examples on task performance and different types of cognitive load. learning and instruction, 30, 32-42. doi: 10.1016/j.learninstruc.2013.12.001 mayer, r. e. (2014). incorporating motivation into multimedia learning. learning and instruction, 29, 171-173. doi: 10.1016/j.learninstruc.2013.04.003 mcauley, e., duncan, t., & tammen, v. v. (1989). psychometric properties of the intrinsic motivation inventory in a competitive sport setting: a confirmatory factor analysis. research quarterly for exercise and sport, 60(1), 48-58. doi: 10.1080/02701367.1989.10607413 mochocki, m. (2014). larping the past: research report on high-school edu-larp. in s. l. bowman (ed.), the wyrd con companion book (pp. 132-149): wyrd con. montola, m. (2008). the invisible rules of role-playing: the social framework of role-playing process. international journal of role-playing, 1(1), 22-36. moos, d. c., & marroquin, e. (2010). multimedia, hypermedia, and hypertext: motivation considered and reconsidered. computers in human behavior, 26(3), 265-276. doi: 10.1016/j.chb.2009.11.004 naismith, l. m., cheung, j. j., ringsted, c., & cavalcanti, r. b. (2015). limitations of subjective cognitive load measures in simulation-based procedural training. medical education, 49(8), 805-814. doi: 10.1111/medu.12732 pekrun, r. (2006). the control-value theory of achievement emotions: assumptions, corollaries, and implications for educational research and practice. educational psychology review, 18(4), 315-341. doi: 10.1007/s10648-006-9029-9 renninger, k. a., & pozos-brewer, r. k. (2015). interest, psychology of. in: international encyclopedia of the social & behavioral sciences, 2nd ed. (pp. 378-385). oxford: elsevier. renninger, k. a., & schofield, l. s. (2014). assessing stem interest as a developmental motivational variable. poster presented as part of a structured poster session. in: current approaches to interest measurement. philadelphia, pa: american educational research association. rey, g. d. (2012). a review of research and a meta-analysis of the seductive detail effect. educational research review, 7(3), 216-237. doi: 10.1016/j.edurev.2012.05.003 rheinberg, f., vollmeyer, r., & burns, b. d. (2001). fam: ein fragebogen zur erfassung aktueller motivation in lern-und leistungssituationen [in german]. diagnostica, 47, 57-66. rheinberg, f., vollmeyer, r., & engeser, s. (2003). die erfassung des flow-erlebens [in german]. in: diagnostik von motivation und selbstkonzept (pp. 261-279): hogrefe. ryan, r. m., & deci, e. l. (2000). intrinsic and extrinsic motivations: classic definitions and new directions. contemporary educational psychology, 25(1), 54-67. doi: 10.1006/ceps.1999.1020 ryan, r. m., rigby, c. s., & przybylski, a. (2006). the motivational pull of video games: a self-determination theory approach. motivation and emotion, 30(4), 344-360. doi: 10.1007/s11031-006-9051-8 sabourin, j. l., & lester, j. c. (2014). affect and engagement in game-based learning environments. ieee transactions on affective computing, 5(1), 45-56. doi: 10.1109/t-affc.2013.27 schiefele, u. (1999). interest and learning from text. scientific studies of reading, 3(3), 257-279. doi: 10.1207/s1532799xssr0303_4 schiefele, u., & krapp, a. (1996). topic interest and free recall of expository text. learning and individual differences, 8(2), 141-160. doi: 10.1016/s1041-6080(96)90030-8 schraw, g., bruning, r., & svoboda, c. (1995). sources of situational interest. journal of literacy research, 27(1), 1-17. doi: 10.1080/10862969509547866 sharp, l. a. (2012). stealth learning: unexpected learning opportunities through games. journal of instructional research, 1, 42-48. doi: 10.9743/jir.2013.6 sitzmann, t. (2011). a meta‐analytic examination of the instructional effectiveness of computer‐based simulation games. personnel psychology, 64(2), 489-528. doi: 10.1111/j.1744-6570.2011.01190.x sweller, j., ayres, p., & kalyuga, s. (2011). cognitive load theory. new york: springer. tingley, d., yamamoto, t., hirose, k., keele, l., & imai, k. (2014). mediation: r package for causal mediation analysis. journal of statistical software, 59(5). doi: 10.18637/jss.v059.i05 um, e. r., plass, j. l., hayward, e. o., & homer, b. d. (2012). emotional design in multimedia learning. journal of educational psychology, 104(2), 485-498. doi: 10.1037/a0026609 vanek, a., & peterson, a. (2016). live action role-playing (larp): insight into an underutilized educational tool. in: learning, education and games. volume two: bringing games into educational contexts (pp. 219-240): etc press. vansteenkiste, m., sierens, e., soenens, b., luyckx, k., & lens, w. (2009). motivational profiles from a self-determination perspective: the quality of motivation matters. journal of educational psychology, 101(3), 671-688. doi: 10.1037/a0015083 watson, d., clark, l. a., & tellegen, a. (1988). development and validation of brief measures of positive and negative affect: the panas scales. journal of personality and social psychology, 54(6), 1063-1070. doi: 10.1037/022-3514.54.6.1063 wouters, p., van nimwegen, c., van oostendorp, h., & van der spek, e. d. (2013). a meta-analysis of the cognitive and motivational effects of serious games. journal of educational psychology, 105(2), 249-265. doi: 10.1037/a0031311 wouters, p., & van oostendorp, h. (2017). overview of instructional techniques to facilitate learning and motivation of serious games. in instructional techniques to facilitate learning and motivation of serious games (pp. 1-16): springer. järvenoja et al publication frontline learning research vol.6 no. 3 (2018) 85 104 issn 2295-3159 capturing motivation and emotion regulation during a learning process hanna järvenojaa, sanna järveläapiia näykkia, jonna malmberga, kristiina kurkia, arttu mykkänena, tiina törmänena, jaana isohätäläa alearning & educational technology research unit, university or oulu, finland article received 14 may 2018/ revised 12 september/ accepted 20 october/ available online 7 december abstract this paper describes our research approach in which we have focused on situational and contextual variations in motivation and emotion regulation to better understand its role, appearance and function in collaborative learning situations. we have used research designs that employ process-oriented measures combined with subjective interpretations to capture motivation and emotion regulation. analysing on-line process data poses several challenges such as variation in the granularity of different data sources, problems that emerge due to the complexity of contextual and situational factors in ecologically-valid learning situations or, currently, challenges in the use of multiple data channels and their analyses. in this paper, we present three claims underlying our research, particularly the motivation and emotions and their regulation in learning. the claims are as follows: (1) motivation and emotion regulation is situation and context specific, (2) motivation and emotion regulation is influenced by multi-layered nature of motivation and (3) motivation and emotion regulation is intertwined with other processes of learning and can be captured from their temporal manifestation. we present an example from our empirical study to discuss how these claims have led us to employ multiple process-oriented methods that include both subjective and objective data sources, including different combinations of situation-specific self-reports, video and physiological data. we then describe opportunities and challenges involved in the empirical studies. keywords: self-regulated learning; emotion regulation; motivation regulation; multiple data channels info mail corresponding author hanna.jarvenoja@oulu.fi doi: https://doi.org/10.14786/flr.v6i3.369 1. introduction the role of emotions and motivation in fostering learning and achievement has been acknowledged in the field of learning sciences (linnenbrink-garcia, & pekrun, 2011; pekrun, goetz, titz, & perry, 2002). theories on motivation and emotions in learning cover multiple concepts and theoretical models explaining the relationship between learners’ beliefs and feelings in learning (pekrun, 2016). motivation theories are particularly involved in describing the reasons why and how learners pay attention to, concentrate on, invest effort in and persist during their academic learning (schutz & pekrun, 2007; volet & järvelä, 2001). still, theories of motivation and emotion have been criticised from not translating self-evidently into classroom practices (boekaerts & corno, 2005; dignath, buettner, langfeldt, & goethe, 2008; pintrich, 2003). too often, motivational constructs are well recognised in research but do not have sufficient practical implications in to real learning contexts. this may stem from research approaches that study certain motivational components in isolation from the actual learning process, limiting the complex and still unclear issue of the dynamic relationships between context and motivation to a background variable (volet & järvelä, 2001; volet & vauras, 2013). many conventional self-reports of students’ emotions and motivation, for example, measure students’ appraisals and perceptions of their emotional or motivational experiences, but do not explore how motivation and emotions are situated and realised in the learning context or how they fluctuate between situations or over time (järvelä, salonen, & lepola, 2001; paris & turner, 2012; winne & perry 2000). we argue that one way to make research on motivation and emotion more effective is to develop research designs and methods that capture the dynamics of students’ emotions and motivational factors during the learning process in ecologically-valid, authentic, learning contexts (pekrun, 2016). accordingly, we engage a process-oriented perspective on studying motivation and emotion regulation as a part of self-regulated learning (srl). srl theory provides us with a theoretical framework that allows us to bring the situational and contextual variation of learners’ motivation and emotions into the focus by targeting their actualized regulation of motivation and emotions (boekaerts & pekrun, 2015; pintrich, 2000; zimmerman & schunk, 2011; winne & hadwin, 1998). it emphasises motivation and emotions as essential features of the learners’ commitment in a learning situation and offers a way to conceptually and empirically grasp the multiple layers of motivation and emotions that are realized in the regulation actions in-situ (järvenoja, järvelä, & malmberg, 2015; volet, vauras, & salonen, 2009). as for engaging in emotion regulation in the learning situation, learners regulate their affect and emotional experience to ensure emotionally solid (social) ground for completing academic tasks (boekaerts & pekrun, 2015; pekrun, 2016). emotion regulation is required, for example, to reduce negative affective responses in a socio-emotionally challenging situation or harmful effects of emotional experiences for learning and academic performance. regulation of motivation is composed of purposeful and appropriate strategic activities through which individuals or group members in coordination initiate, control, and supplement their willingness to maintain the learning process and achieve learning goals (boekaerts & pekrun, 2015). motivation regulation can be directed, for example, at strengthening or redirecting interest, motivational goals or self-efficacy beliefs (wolters & benzon, 2013). finally, we have built our research approach on empirical evidence that emphasises that group members’ motivation and emotions and their effects on learning cannot be thoroughly comprehended without considering the social context in which they occur (dowson & mcinerney, 2003; hickey & mccaslin, 2001; volet & järvelä, 2001). all learning includes social dimensions and is inherently social, but the meaning of the social context and interaction is particularly essential in collaborative learning contexts (baker, 2015). collaborative learning has become an essential 21st century skill (griffin, mcgaw, & care, 2012; sawyer, 2014); its benefits for learning have been emphasised by many researchers (miyake, 1986; roschelle & teasley, 1995; webb, troper, & fall, 1995), and collaborative learning has also become an increasingly valued teaching and learning practise in schools. motivation and emotion regulation in collaborative learning has been characterized as a fundamental part of effective collaborative learning interactions as a part of an increasing interest in defining the regulation mechanism of the groups (hadwin, järvelä, & miller, 2018). lately, the research has become increasingly interested in exploring how motivation and emotion regulation function and fluctuate during collaborative learning situations, and for example, when and how group members share the regulatory responsibilities (järvenoja, järvelä, & malmberg, 2017). however, compared to progress in research on groups’ cognitive regulation mechanism, studies focusing on motivation and emotion regulation in collaborative learning situations are still scarce. in this paper, we introduce our research approach by presenting three claims that address the investigation of motivation and emotion regulation in-situ, describe our methodological orientation and justify the choices of implemented methods. while the three claims overlap somewhat, each claim aims to underlie certain aspects of the research approach. the first claim, motivation and emotion regulation is situation and context specific, grounds our choice of investigating motivation and emotion regulation in authentic learning contexts. the second claim, motivation and emotion regulation is influenced by the multi-layered nature of motivation, gives reasons why it is not enough to focus solely on traceable regulation activities but to complement process data on regulation activities with other measures that capture different components of individual beliefs and interpretations about their motivation and emotions. finally, with the third claim, motivation and emotion regulation is intertwined with other processes of learning and can be captured from its temporal manifestation, we introduce our recent endeavours to explore the possibilities of employing data on group members’ reactions in tracing meaningful situations in terms of motivation and emotion regulation. in relation to the claims, we illustrate with an empirical case example our methodological decisions and highlight the possible advantages and challenges. altogether, the three claims illustrate how the implementation of process analysis does not derive from a simple desire to adopt a certain choice of method(s) but a research agenda that influences the choices made, from the research design to the analysis. 2. claim: motivation and emotion regulation is situation and context specific the first claim, motivation and emotion regulation is situation and context specific, has directed our choices for research contexts towards ecologically valid learning settings where the students collaborate as a part of their regular studying (järvenoja et al., 2015; volet & järvelä, 2001). studies that are carried out in these situations where students have an authentic need to overcome possible challenges in their collaboration (f. kirschner, paas, & kirschner, 2011) enable to us to analyse motivation and emotion regulation in a context that is not isolated but includes all the situational and contextual features that affect to the activation of motivation and emotion regulation. following a situated perspective on learning (greeno, 2005), a learner enters into the learning situation with unique, socio-historically founded personal motivational beliefs and emotional experiences. however, these structures are not static but contextually and situationally changing (volet & järvelä, 2001). each situation forms a unique composition of different factors that together form the circumstances for learning and regulation. situational and contextual factors such as social interaction and the nature of the task, mediate the need for motivation and emotion regulation and, correspondingly, through motivation and emotion regulation the beliefs and experiences can be actively changed or modified in the situation (isohätälä, järvenoja, & järvelä, 2017; kurki, järvenoja, järvelä, & mykkänen, 2017; mykkänen, perry, & järvelä, 2017; whitebread et al., 2009). therefore, regulation of motivation and emotions is socially situated, involving a dynamic interplay between learners, tasks, teachers, peers and parents and is bound up with the context (hadwin et al., 2017; järvenoja, et al., 2015). this has led us to carry out studies where development of situational challenges creates the need for motivation and emotion regulation as well as the actual manifestation of such a regulation is followed from collaborative groups’ learning process. in these case studies, the main source of data has typically been video recorded group activities. for example, the study by järvenoja, näykki, törmänen, and järvelä (2018) implemented moment-by-moment video analysis with 30-second time segments as a unit of analysis to locate the situation-specific challenges and related group-level emotion regulation strategies during higher education students’ (n = 62; mean age 23 years) mathematics’ class. the course was part of the students’ regular study program and the course design involved collaborative learning tasks that were videotaped. the results from 87 hours of video recordings showed that a wide range of challenges (f = 1301 events) emerged in groups, covering challenges with cognitive, motivational and emotional issues, and different sociallyand contextually-oriented challenges. the study also followed how emotion regulation was enacted during these different challenging situations and revealed that emotion regulation was not bound merely to one type of challenge situation, but all of the challenging situations could potentially trigger group-level regulation. again, the type of regulation (co vs. socially-shared) was interconnected with the experienced challenge. co-regulation was used, for example, to increase awareness of the emotional aspects whereas socially-shared emotion regulation was more typically used when social reinforcement was manifested. overall, the above example illustrates how emotion and motivation are bound to the context in which it takes place. particularly, the triggers that activated group members’ coand socially-shared regulation emerge from the current situation and were bound to different contextual features. methodologically speaking, how regulation emerged in the group situations and how regulation interacted in relation to the experienced challenges could not be found if the collaborative learning process was not analysed with micro-level video analyses. studying motivation and emotion regulation in detail as a part of the learning process is essential for recognising and understanding the phenomenon and for operationalising the different indicators of motivation and emotion regulation from the process data. the clear advantage of this type of micro-level video analysis is that it allows the possibility of following how the regulatory behaviours, contextual conditions for motivated learning both shape and are shaped by the situational circumstances of a given time and place. this is important for understanding the mechanisms of emotion and motivation regulation in collaborative learning. detailed video data analysis illustrates the changing composition of interactions, context and activities that give rise to self-regulation, co-regulation and shared regulation of motivation and emotions in the situation (hadwin et al., 2017). process data from these unique situations helps to study motivation, emotions and their regulation ‘in action’ in different context and situations. the pitfall is that contextualized, detailed analysis of video data does not allow for the exploration of the phenomena in a more generalisable manner. case analysis provides understanding and provokes new research questions but for examining, for example, temporal traces and patterns of motivation and emotion regulation, larger data sets are needed (malmberg, järvenoja, & järvelä, 2013). our challenge has also been that operationalisation of different components of analysis, such as a ‘challenging situation’, is not specific enough to be transferred to the systematic analysis of larger data sets as it would be extremely time-consuming (see also wise & schwarz, 2017). for example, in the above described example, the analysis phases included several rounds of categorizing the 87 hours of video data corpus to first reliably locate the challenges and the related regulation. regardless of a growing number of developed methods, such as video recording systems and different state-of-the-art technology tools that make it possible to capture classroom interaction, it has still been difficult to process video data systematically when emotion and motivation regulation is traced from group members’ interactions (azevedo, moos, johnson, & chauncey, 2010). this is particularly challenging as the results tend to suggest repeatedly that while regulation of motivation and emotion is critical for collaborative groups, their occurrence during collaboration is rare and situation specific (järvenoja et al., 2017; sobocinski, malmberg, & järvelä, 2017). this poses a challenge for analysis; even with relatively large video data, the frequency of the coded episodes remains low, limiting considerably the possibilities for analysis and generalisation. however, it also relates to the main premise of our first claim; the triggers for groups to activate emotion and motivation regulation are context-bound, encompassing a situated combination of different features. much more evidence covering a critical mass of process data is required to confirm the role and meaning of emotion and motivation for regulated learning progress. finally, the analysis of motivation and emotion regulation solely from video data is limited to visible reactions and verbalised expressions, leaving out motivational and emotional processes that are silent and hence, invisible to traditional video observation. even though this type of research design provides valuable instruments and methods for contextand task-specific measures to explore regulatory actions as they occur in real time, it is not possible to reach individual internal reactions, beliefs and interpretations of the situation solely with process data. the pitfall is that the researcher relies on his/her interpretations of the phenomena on the verbal and non-verbal behaviours of the learners, but does not know the actual reasons behind them. this leads us to the next claim. our second claim, motivation and emotion regulation is influenced by the multi-layered nature of motivation, refers to the multiple ways and levels in which motivation and emotions function in learning. we started the first claim by arguing that motivation and emotions in the situation are dependent on both learners’ prior socio-historical experiences and social contexts (elliot et al., 2016; levy, kaplan, & patrick, 2004; nolen, horn, & ward, 2015). the multiple emotional and motivational factors, namely appraisals, expectancies, values, beliefs and goals, together form the (pre)conditions for motivation and emotion regulation (pekrun, 2016). the learning situations are built on these conditions, indicating that learner approaches and decision-making processes are personalised by prior individual and group experiences over time and events (hadwin, järvelä, & miller, 2017). even though ‘objective’ process methods, (e.g., video records or log traces that accumulate from technology-based learning environments) can provide rich accounts of learners’ and groups’ actions and visible reactions in the moment, understanding reasons behind these actions require ‘subjective data’ (e.g., questionnaires, diaries, interviews) (berger & karabenick, 2016; fincham et al., 2018; järvelä, malmberg, sobocinski, haataja, & kirschner, 2018; malmberg et al., 2013). one of our first efforts to develop measures to capture individual beliefs in a specific learning situation was the development of an instrument, called adaptive instrument for regulation of emotions (aire) (järvenoja, volet, & järvelä, 2013), which is a self-report instrument designed to access students’ experiences of individual and socially-shared regulation of emotions repeatedly in a specific group learning activity. the instrument is comprised of four interrelated sections, each with a different focus. these four components identify (1) personal goals, (2) the socio-emotional challenges experienced while studying in a collaborative situation, (3) group members’ own evaluation of their individual and group-level attempts to regulate the immediate emotions evoked by the challenges, and (4) reflections on personal goal attainment in the learning situation and how the collaborative group work contributed to it. as each group members’ response to the aire is individual, it is sensitive to students’ unique experiences. the aire instrument enables analyses that compare the coherence between individual group members’ appraisals of the reasons for socio-emotional challenges within specific learning situations, their personal goals and satisfaction with collaborative learning experience. consequently, it captures the regulatory process of the whole group. furthermore, when group members’ responses to different components of the aire instrument are compared, it is possible to form a group-level interpretation of the situation (järvenoja & järvelä, 2009). this affords the possibility of moving the focus when needed to a group level and to provide an estimation of the groups’ joint understanding of the experienced socio-emotional challenges and motivation and emotion regulation. in its entirety and choice of analytical possibilities, aire is an example of subjective self-report instruments that support our preliminary assumptions of ecologically valid, situation-specific and process-oriented measures. while self-report measures have been criticised as being too stable and trait oriented (winne & perry, 2000), the need for self-reports cannot be neglected since they access the self-perceptions of individual learners (bandura, 2011; mccardle & hadwin, 2015). rather, we emphasise that process-oriented analysis should not be construed as a substitute for learners’ subjective appraisals but if the form of self-reports is purposefully considered, they can become an approach that reaches beyond the observable or measurable reactions and behaviours and provides the possibility of focusing on learner experiences (koskey, karabenick, woolley, bonney, & dever, 2010; nolen, 2006). to study motivation and emotion regulation as a multi-layered phenomenon and as a part of the learning process, however, presumes going beyond recognising and capturing the subjective interpretations and explanations or focusing solely on the process data. as we have argued above, the process in which the effects of motivational and emotional conditions actualise into products of learning and motivation regulation are composed of both individual beliefs and processes taking place in the situation (bakhtiar, webster, & hadwin, 2018). subjective data (e.g., repeated and contextualised self-reports) can reveal students’ intentions to learn and the type of beliefs they have about themselves as learners (mccardle & hadwin, 2015). conversely, objective data (e.g., video and log data, eye movements and physiological responses) provide continuous information about behavioural and mental indicators such as confusion and increasing effort or attention, which are almost impossible to capture otherwise. when the two types of data sources are combined, we can both trace occurrence of emotion and motivation regulation and explain the conditions and products it is bound to (bakhtiar et al., 2018). näykki, järvelä, kirchner, & järvenoja. (2014), for example, combined video-observation data with video-stimulated recall interview data to follow and explain how higher education student collaboration led to a severe socio-emotional conflict. the group dynamics and task characteristics were depicted through process analysis of video data that revealed emerging motivational (e.g., task commitment problems) and emotional (e.g., frustration due to overruling interaction) challenges. the personal reasons behind the conflict was then revealed in the video-stimulated interviews where group members’ provided individual, subjective explanations for the experienced challenges. the combination of two data sources explained why one of the groups failed in regulation of their emotions and motivation, while the other groups managed to regulate the challenging situations. in our ongoing study, we have moved forward in specific operationalisation of the situational indicators of challenges that include triggers for activating emotion and motivation regulation. we have collected data from 12-year old primary school students (n = 41), as they collaborated in three-member groups to complete a design task to construct a model of an energy-efficient house making use of solar energy. apart from process data comprised of video-taped working and physiological data (see claim 3 for more details), both general level and situated self-reports were collected. general level self-reports captured two different motivational constructs: individual students’ self-efficacy beliefs (usher & pajares, 2008) and interest towards science (cleary, 2006) and self-reported beliefs regarding students’ self-regulation skills (cleary, 2006). situated self-reports measured individual group members’ perceived valence of emotional state with the emotion awareness tool (ema) before and after the group task. at the end of the collaborative task, students also evaluated their satisfaction towards the collaborative work of their group. in the part of the data analysis that makes use of the subjective self-report data to explain the actualised motivation and emotion regulation captured from video data, we are currently focusing on the self-reported valence measured before and after the collaborative learning session and its relation to actualised regulation in collaborative learning episodes having a negative group-level valence (kurki, järvenoja, törmänen, & bakhtiar, 2018). while self-reported ema data provides us the individual accounts of students’ self-perceived emotional valence in the situation, the analysis of video data extends the analysis on the process level, exploring the characteristics of the group process in terms of socio-emotional interactions, their valences, and actualised regulation of motivation and emotions. the analysis of the process data began by depicting the groups’ socio-emotional interaction episodes from the video data corpus. emotional valence and related emotion and motivation regulation were then coded in relation to these episodes. the video data was processed using observer xt software. the valence coding protocol categorization is presented in table 1 as an example of our typical video coding protocol. the table explicitly states how individual-level emotional indicators are translated into a group level emotional valence in the coding. it also provides examples of motivation and emotion regulation that may emerge in relation to these episodes. table 1 an example of coding categories for emotional valence after operationalising and coding the different components, the analysis proceeded to compare negative socio-emotional episodes and related group-level regulation with individual student reports on their emotional valence before and after the task. the statistical correlations between the variables were calculated using the estimates of spearman’s correlation coefficient, and the preliminary results indicate that students’ negative emotional states before the task did not increase negative interactions, but actually promoted regulation of motivation and emotion during the negative interactions (kurki et al., 2018). no clear connection was found between the actualised regulation and the valence. however, the self-perceived valence increased in general after the task. the next step is to add a new layer by adding the general self-report data to the analysis. the advantages of combining subjective interpretations with process data analysis lie in the possibility of tracing the dual relationship between individual beliefs and actualised regulation at the group level. in the above example, this meant that by collecting the student evaluations before and after the learning task, we were able to first investigate how emotional valence was reflected in the regulation during the collaborative learning task, and second, how regulation may have modified the individual student’s emotional valence after the group task. when the data on more general-level beliefs is added to the analysis, it will be possible to analyse how subjective beliefs about motivation and emotion in learning shape the students’ preconditions to the collaborative learning and thus, to explore how these beliefs shape the motivation and emotion regulation and, vice versa, how the regulation processes shape students’ future motivational beliefs and emotional experiences. furthermore, combining these multiple layers provides the possibility of exploring in more detail what characterizes the socio-emotional interactions and their coand socially-shared regulation (järvelä; järvenoja, & veermans, 2008; hadwin et al., 2017). it is an advantage from both theoretical and methodological perspectives to combine multiple data sources that capture different dimensions of motivation and emotions instead of relying on a sole source of data (ochoa, 2017). when individual subjective meaning is emphasised as the conditions upon which the contextual interpretations are made, data on these subjective appraisals cannot be neglected. different data sources provide possibilities for data triangulation (azevedo, 2015). this is essential particularly when the theoretical definition and empirical evidence is still evolving. however, the implementation of multiple data sources that are operationally distinct from each other does not come without pitfalls. when merging subjective interpretation data with process data analysis, we first deal with the challenges of different types of data sources (winne & perry, 2000). however, the main challenge we face is consolidation of data gathered with different methods. we are constantly struggling, for example, to find meaningful correlations with individual self-report data, learning outcomes and process data from student interaction (e.g., winne & jamieson-noel, 2002). while this could be an indication of inaccurate selection of certain instruments or data analysis protocol, it could also derive from differences in granularity size of different data sources or the unit of the analysis. typically, the number of participants in the studies that collect process data from real-life collaborative learning is low, which restricts the possibility of self-report data, particularly when we want to engage group-level analysis when the total number of participants is further divided into the number of the collaborative groups. the amount of equipment, space, student groups, etc. that the process data collection in authentic collaborative learning tasks requires is also a bottleneck. while collecting process data from multiple sessions cumulates the data mass, and partly also the situated self-report data mass, the total number of participants remains relatively low as was seen in above example. to summarise, multi-layered process data collection sets challenges for a) combining different data sources, b) combining and selecting instruments that would be comparable and theoretically valid with each other and process data, c) recruiting a sufficient number of participants and finally, d) having an infrastructure that enables the collection of process data in ecologically valid settings. 4. claim: motivation and emotion regulation is intertwined with other processes of learning and can be captured from their temporal manifestation. the third claim, motivation and emotion regulation is intertwined with other processes of learning and can be captured from their temporal manifestation, refers to our latest empirical endeavours to find ways to more systematically explore the different sub-processes that can trigger or indicate the emergence of motivation and emotion regulation during collaborative learning. while we know that coand socially-shared regulation appears at low incidence that are interconnected with situation (järvenoja et al, 2018), to progress with research, we are destined to explore ways of locating and analysing motivation and emotion regulation more systematically from the learning process, in parallel, embedded and interdependent with other motivational and emotional processes and in relation to cognitive manifestation of regulated learning (hadwin et al., 2017; järvelä & hadwin, 2013; järvelä, järvenoja, malmberg, isohätälä, & sobocinski, 2016). in the majority of studies implementing multiple process data, the operationalisation of regulatory interactions and behaviours are limited to ‘cognitive’ episodes alone. for example, malmberg, järvenoja and järvelä (2017) explored the sequential patterns of self-, co-, and socially-shared regulation of learning in the context of collaborative learning but focused solely on cognitive and metacognitive processes. while this body of research has progressed well, providing understanding, for example, of temporal manifestation and patterns of regulated learning in individual and group contexts (e.g., molenaar & chiu, 2014; sonnenberg & bannert, 2015; zheng & yu, 2016), it still has often undermined the salience of motivational and affective processes and beliefs (hadwin et al., 2017). process-oriented methodology that is not limited only to observations, could offer possibilities for systematic analysis that integrates motivation and emotion regulation with the temporal analysis regulated learning more expansively. currently, we are progressing in our analysis of different motivational and emotional indicators that contribute to the process of regulated learning and provide evidence on how motivation and emotion regulation plays a role in group interaction during collaborative learning (järvenoja et al., 2015; malmberg et al., 2017; malmberg, järvelä, holappa, haataja, siipo, & huang, 2018). at the moment, we are exploring how the research on motivation and emotion regulation in collaborative learning could benefit from implementation of physiological measures. combining physiological measures with situated, process-oriented approaches is an uninvestigated area that has a potential to advance previous research regarding the influence of motivation and emotion regulation on learning and achievement. however, we argue that accurate inferences about motivation and emotion regulation require objective data to be carefully contextualised by subjective data about learners’ intents and beliefs in the same relative moment. for example, ahmed, van der werf, and minnaert (2010) combined multi-method qualitative methods and a physiological reactivity measure to investigate students’ emotional experiences in the classroom and henriques, paiva, and antunes (2013) measured electrodermal activity to access emotional patterns occurring in affective interactions. physiological measures have also been used to complement video and log data to filter, organise and classify data into meaningful episodes according to the criteria derived from the theoretical justification and research focus (appelhans & luecken, 2006; mcrae et al. 2012). in relation to the regulation processes in collaborative learning, haataja, malmberg, and järvelä (2018) studied physiological concordance (pc) of group members’ electrodermal activity and observed regulation episodes, and found a weak positive connection between them. how to relate individual physiological reactions to the regulation of emotion or motivation during collaborative learning, however, is still to be explored. therefore, we next present an example from our current, on-going study to explore the possibilities of implementing physiological data in recognising and tracing emotion and motivation regulation in the course of collaborative learning. in the following example, we are measuring collaborative group members’ electrodermal activity from the skin to depict temporal variations in their physiological arousal (see also ahonen, cowley, hellas, & puolamäki, 2018; gillies et al., 2016). to contextualise the physiological data, we relate it with video data codings of valence in socio-emotional episodes (see example in claim 2), which considered as potential indications of situations in which group members’ joint coand socially-shared emotion and motivation regulation can emerge. ultimately, we aim for more systematic evidence of how motivation and emotion regulation is affecting the temporal progress of collaborative learning. the example derives from the same data set presented under claim 2 (kurki, et al., 2018). to reveal the complexity of motivation and emotion regulation in collaborative learning, we collected two types of process data, 360° video data and electrodermal activity (eda) gathered with empatica e4 wristbands (garbarino, lai, bender, picard, & tognetti, 2015). the analysis started by processing the data from two different channels separately. the video data processing was founded on the data coding presented in the claim 2. the initial coding included indications of different learning phases (brainstorming, planning and building), locating socio-emotional interaction episodes (järvenoja et al., 2017; linnenbrink-garcia, rogat, & koskey, 2011), defining the groups’ emotional valence during the socio-emotional interaction episodes, and, finally, searching for regulation of motivation and emotion in relation to these episodes (järvelä, et al., 2016). with eda data, we focused on the phasic skin conductance response (scr) peaks as the former research had indicated that scr peaks are strongly associated with emotional responses that are related to significant external stimulus. hence, scr peaks are considered more reactive to variations in experimental conditions than tonic skin conductance level (scl) (christopoulos, uy, & yap, 2016; dawson, schell, & filion, 2007). after collecting the physiological data, it was first imported to empatica e4 connect software. the physiological eda data was then processed using python software and microsoft excel. three groups had to be left out of the analysis in this initial data processing phase due to poor quality data of some of the group members. the baseline was computed using a third-order low-pass filter and scr peaks were detected using the minimum value of 0.05 µs between the baseline and peak (boucsein, 2012; dawson, schell, & filion, 2007). figure 1 shows an example of one group member’s raw eda data visualized in empatica e4 connect software (a) and processed eda data with detected scr peaks (b) which was used later in the analysis. figure 1. an example of the visualisation of one group member’s raw eda data extracted from empatica e4 connect software (a) and how the raw data was processed on the individual level to track scr peaks (b). after processing both video data and eda data independently, the analysis proceeded to group-level analysis, which combined coded video data and physiological data. already in the data collection phase, the video cameras and empatica wristbands were synchronized, which enabled us to find timely commensurable indicators of the groups’ emotional state from two different data channels. in this phase of the analysis, both video and eda data were segmented in 30-second segments to make the data commensurable for further analysis. the 30-second time frame was chosen based on the preliminary socio-emotional episode coding of three videos, which defined the mean duration of coded episodes to 24.6 seconds. the segmentation proceeded by re-coding the whole video data corpus in terms of group members’ socio-emotional interaction into the 30-second segments. this coding was a dichotomous yes/no coding based on group members’ visible emotional expressions (table 1). the segments that included no emotional expressions or expressions only from one group member were considered as neutral in terms of the groups’ emotional state and were not analysed further. the segments including socio-emotional expressions were categorised according the valence of groups’ emotional state (positive, negative, mixed, unclear) and group-level emotion regulation, namely coand socially-shared regulation and was performed for 40% of the coded videos using cohen’s kappa statistics. substantial agreement was reached for both socio-emotional segment (κ = 0.693) and valence (κ = 0.723) coding. eda data was further processed by determining the frequency of each group members’ scr peaks in each 30-second segment. when both data sets were processed independently, they were synchronised based on the timestamps. to explore the possible association between the two sets of data as a whole, the relation between eda segments and video observation segments was explored with chi-square statistics. for this purpose, the data on scr peaks was reorganised into group level by categorizing each segment based on how many group members (0, 1, 2, 3) were having scr peaks in the segment. significant associations between observed segments and eda segments were further explored with significant z scores from adjusted residuals with alpha levels 0.05 (z < 1.96) and 0.001 (z < 2.58). as the chi-square test showed a significant relation between socio-emotional expressions observed and number of group inter-rater reliability analysis members with scr peaks in the segment (χ2 (3) = 27.106, p < 0.000), the analysis was continued by exploring this relationship in more detail within groups. altogether, 95 segments in which all group members are having scr peaks and group-level valence were found. figure 2 presents a three-member case group example of how the video data codings and eda data were combined (törmänen, järvenoja, kurki, devai, & järvelä, 2018). in figure 2, the frequency of each group member’s (jenny, alan and miia) individual scr peaks in the situation is compared with the group-level valence coding to illustrate the associations between the observed emotional reactions and physiological reactions. the segments in which all three group members had scr peaks as well as emotional expressions in the video are highlighted with vertical, coloured bars. as can be seen from the figure 2, most of these situations occur at the end of the task, indicating that most of the emotionally activating sequences in this case group were located to the end of the collaborative learning activity. figure 2. example of the analysis that combines data from two data channels, namely (1) valence (red, green, yellow, and blue dots) detected from the video data and (2) the number of skin conductance response (scr) peaks (blue line) from the eda data. to detect emotionally activating situations on a group level, segments in which all group members are having scr peaks and also expressing emotions in the video are located from the learning session (coloured segments) (törmänen et al., 2018). next, we progressed into combining the motivation and emotion regulation coding into the analysis. our assumption was that activating motivation and emotion regulation at the group level requires a trigger, for example a socio-emotionally challenging situation that invites the group members to regulate emotions and maintain motivation (hadwin et al., 2017). that is why we will explore how the sequences that were located by combining scr peaks with group level valence in the previous analysis phase are related with the coand socially-shared regulation episodes coded from the video data. from 95 emotionally activating segments, 28% included either coor socially-shared regulation. this is significantly more than the occurrence of coand socially shared motivation and emotion regulation in the whole data set. these preliminary results indicate that when physiological reactions, namely peaks in scr and observable emotional reactions in a group occur at the same time, they together may indicate a presence of a situational trigger that invites the group members to activate motivation and emotion regulation. however, further in-depth exploration needs to be conducted for the located segments to find the common situational characteristics, regardless whether they are cognitive, motivational or emotional, that determine the trigger features of these situations and consequently relate them with the coded motivation and emotion regulation. it should be noticed as well that a significant amount of motivation and emotion regulation was not located with this analytical approach. in the future, we need to explore whether some of these episodes would be better tracked by using some other features of motivation, emotions or physiological data. from the eda data, moments of physiological concordance (pc), for example, could be explored in relation to the video coding (see e.g., haataja, et al., 2018). the type of analysis presented above has a clear advantage in that it combines data sources that derive from the same situation and target the same process. however, when with eda data analysis, the interpretation is not as straightforward as with video data. in addition, the challenge in adding the physiological data is that it adds a new data channel and consequently new layer of analysis to an already complex combination of data sources and analysis dimensions. therefore, careful consideration is needed when combining data with different foci (i.e., individualor group-level data) and granularities (i.e., millisecond versus minutes of episodes as a unit of analysis) in a theoretically meaningful and empirically valid analysis. these analyses also face challenges with the granularity of the basic unit of the analysis. for example, the sampling rate of the empatica e4 wristband is 4hz (that is, four times per second) (garbarino et al., 2015), whereas a meaningful time span for video data coding covers sequences that span from a few seconds to episodes lasting several minutes. in the example above, we presented how we managed to match the two originally very different time scales together in a timely manner. we are still struggling, however, with how to match motivation and emotion regulation with this analysis. the main challenge seems to lie in the fact that within the group, the emotion and motivation regulation is not activated every time the valence coding and scr peaks indicate potential for it. in addition, regulation does not fall in the same time sequence frequency with the two other dimensions; the scr peaks and valence are more instant reactions while regulation requires awareness that is created through (metacognitive) monitoring, which is not instant. accordingly, our data seems to indicate that there can be a lag between the emotional reactions and activated regulation. in addition, the temporal dependence of the time series data and physical movement, should be taken into consideration. finally, it is possible we have not yet found the right indicators of the triggers that activate regulation and which act as mediators between physiological reactions and the actualised regulation. the physiological data in general, and electrodermal activity particularly, challenge us to consider what the meaningful physiological indicators regarding groups’ emotional experiences, learning and interaction are in the first place. that is, physiological reactions are just reactions and are meaningless for possible regulation without individual motivational and emotional interpretations of the situation. eda, for example, is a signal related to the intensity of arousal, meaning that it does not provide information on the quality of the valence of students’ emotional states, not to mention regulation (eiser, 1986). therefore, eda is a measure that provides quantitative properties, while the determination of qualitative properties requires measures of subjective interpretations and measures that provide a meaningful context to explain the reactions captured with the physiological measures (boucsein, 2012; immordino-yang, christodoulou, pekrun, & linnenbrink-garcia, 2014). thus, we still need to explore in more detail the theoretically meaningful associations between different layers of the multiple data before this analysis can truly add to our understanding of the role of motivation and emotion regulation in learning. when this is accomplished, this type of methodological design will enable extension to open in real life learning settings. 5. conclusion through the three claims and related examples presented in this paper, we described our research approach, which emphasises the role of motivation and emotion regulation in the regulated learning process (ben-eliyahu & linnenbrink-garcia, 2013; duffy et al., 2015; järvelä et al., 2016; kwon, liu, & johnson, 2014; rogat & adams-wiggins, 2015). the three claims highlighted a particular viewpoint to the approach: a requisite to study motivation and emotion regulation as situated in the learning context, a need to acknowledge both the process in which regulation is actualised as well as individuals’ subjective beliefs and appraisals of these processes, and finally, a possibility to understand and capture motivation and emotion regulation by tracking related indicators from learning process. in relation to each claim, we discussed the advantages and challenges we have faced regarding to that claim. most of the advantages and the challenges we have faced are not related only to one of the claims. instead, when putting the research in action, the different claims become intertwined and the pros and cons of our approach derive from this entity. this was realized in the empirical data examples presented in relation to claims 2 and 3. the examples derived from a single data collection, and together they showcased our endeavour to capture motivation and emotion regulation by implementing multiple and complementary methods. together, the examples illustrated particularly, how we have utilized the concept of valence as a theoretical construct to indicate potential for motivation and emotion regulation to emerge. valence proved to be useful as it was possible to measure or locate it from data deriving from different sources. hence, valence mediated the combination of different data channels of subjective interpretations, observable (inter)actions and physiological reactions. it was also theoretically connectable with motivation and emotion regulation depicted from video as well as with physiological arousal collected with the empatica wristbands. the example in relation to claim 3 showed that by combining valence and arousal levels within a group, we were able to increase the probability of finding episodes of motivation and emotion regulation to some extent. to conclude, by adding to our analysis of motivation and emotion regulation a well-defined construct that better determined the nature of socio-emotional interaction, we were also better able to utilize the different data channels in the comprehensive analysis. while the process is still on-going, focusing on constructs, which can be clearly operationalised, appears promising to find the systematic way to locate and trace motivation and emotion regulation in action. however, we are still facing significant challenges with the large and complex data sets that are context specific (azevedo, 2015; bannert, reimann, & sonnenberg, 2014; gašević, dawson, & siemens, 2015). it is evident that a single operationalised indicator, in our case valence, is not solely enough to indicate the triggers that activate groups’ shared motivation and emotion regulation. it is presumably that we need more than one indicator—a set of operationalised indicators—to be able to efficiently employ the potential of physiological reactions in tracing more comprehensively the situations where motivation and emotion regulation emerges. the complex designs produce a significant amount of multiple data, but ‘more data’ does not inevitably lead to ‘deeper data’ with meaningful results (reimann, markauskaite, & bannert, 2014). although the data is rich and allows opportunities to reach motivation and emotion regulation in parallel with, for example, the analysis of other regulatory processes, these operationalised indicators are essential to manage with the complexity and richness of the multiple data channels each producing raw data from different processes and levels. in addition, the examples presented here still concern a relatively small data set in which the learning process is shaped by the complexities of different contextual and situational aspects encountered in real-life learning situations. therefore, it suffers from the inability to differentiate between the more and less meaningful indicators in tracing motivation and emotion regulation. baker, hershkovitz, rossi, goldstein, and gowda (2013) pointed out that a major limitation of multi-layered data analysis, including video observations, has difficulty in scaling large amounts of data or large numbers of students. tracing different motivational and emotional variables ‘moment to moment’ has the potential to reveal the antecedent and consequences of motivation and emotion regulation. this is needed to explain (successful) emotion and motivation regulation, but the current limitation is that with the amount of data we are able to process at the moment, it is difficult to interpret the relations and meaningfulness of these moment-to-moment indicators for regulation and further, to groups’ learning. new analysis methods, such as educational data mining or time-series methods could be useful in the future to be able to search for dependencies between triggers and occurrence of motivation or emotions within the group also from larger sets of data. we conclude that capturing motivation and emotion regulation means knowing something about learners’ internal perceptions and intent and relating that with the (inter)actions in the current situation (wise & schwartz, 2017). social interactions, sequences and patterns need to be further contextualised in larger episodes of activity with attention to individual and collective goals, plans and reflection to delineate metacognitively-driven regulatory processes versus extemporaneous patterns of interaction (järvelä et al., 2017). we argue that even with the cost of future errors and mistakes, this endeavour of unlocking the regulated learning process through the analysis of data gained from multiple data channels has the potential to produce knowledge regarding groups’ motivated learning and regulation. we further argue for the importance of drawing upon multiple analytical methods and invite scholars to multidisciplinary collaboration (e.g., data scientists) to examine the data and multiple modes of motivation and emotion regulation as they operate and support (or hinder) learning and collaboration. these studies can take important steps towards overcoming the dichotomy between modes of regulation, such as cognition, metacognition, motivation and emotion and instead examine the interplay between them. keypoints research approach to study motivation and emotion regulation in learning process is introduced. motivation and emotion regulation is situation and context specific. motivation and emotion regulation is influenced by multi-layered nature of motivation. motivation and emotion regulation is intertwined with other processes of learning and can be captured from their temporal manifestation. multiple data channels can help to capture the motivation and emotion regulation from the process data. acknowledgments this work was supported by the finnish academy [grant number 297686; 24302574; 316129]. references ahmed, w., van der werf, g., & minnaert, a. (2010). emotional experiences of students in the classroom: a multimethod qualitative study.european psychologist, 15(2), 142-151. http://dx.doi.org/10.1027/1016-9040/a000014 ahonen, l., cowley, b. u., hellas, a., & puolamäki, k. (2018). biosignals reflect pair-dynamics in collaborative work: eda and ecg study of pair-programming in a classroom environment. scientific reports ,8(1), 1–16. https://doi.org/10.1038/s41598-018-21518-3 appelhans, b. m., & luecken, l. j. (2006). heart rate variability as an index of regulated emotional responding.review of general psychology, 10(3), 229-240. http://dx.doi.org/10.1037/1089-2680.10.3.229 azevedo, r., moos, d. c., johnson, a.m., & chauncey, a. d. (2010). measuring cognitive and metacognitive regulatory processes during hypermedia learning: issues and challenges . educational psychologist 45(4), 210-223. https://doi.org/10.1080/00461520.2010.515934 azevedo, r. (2015). defining and measuring engagement and learning in science: conceptual, theoretical, methodological, and analytical issues.educational psychologist, 50(1), 84–94. https://doi.org/10.1080/00461520.2015.1004069 baker, m. j. (2015). collaboration in collaborative learning. interaction studies, 16, 451–473. https://doi.org/10.1075/is.16.3.05bak baker, r. s., hershkovitz, a., rossi, l. m., goldstein, a. b., & gowda, s. m. (2013). predicting robust learning with the visual form of the moment-by-moment learning curve. journal of the learning sciences,22(4), 639–666. https://doi.org/10.1080/10508406.2013.836653 bakhtiar, a., webster, e. a., & hadwin, a. f. (2018). regulation and socio-emotional interactions in a positive and a negative group climate. metacognition and learning, 13(1), 57–90. https://doi.org/10.1007/s11409-017-9178-x bandura, a. (2011). on the functional properties of perceived self-efficacy revisited. journal of management, 38(1), 9–44. https://doi.org/10.1177/0149206311410606 bannert, m., reimann, p., & sonnenberg, c. (2014). process mining techniques for analysing patterns and strategies in students’ self-regulated learning. metacognition and learning, 9 (2), 161–185. https://doi.org/10.1007/s11409-013-9107-6 basil-shachar, j., hod, y., & ben-zvi, d. (2015). the emergence of norms in a technology-enhanced learning community. in o. lindwall, p. koschman, t. tchounikine, & s. ludvigsen (eds.), exploring the material conditions of learning: the computer supported collaborative learning conference (cscl ), volume ii (pp. 807-808). gothenburg, sweden: the international society of the learning sciences. ben-eliyahu, a., & linnenbrink-garcia, l. (2013). extending self-regulated learning to include self-regulated emotion strategies.motivation and emotion, 37(3), 558–573. https://doi.org/10.1007/s11031-012-9332-3 berger, j. l., & karabenick, s. a. (2011). motivation and students’ use of learning strategies: evidence of unidirectional effects in mathematics classrooms. learning and instruction, 21(3), 416-428.boekaerts, m., & corno, l. (2005). self regulation in the classroom: a perspective on assessment and intervention.applied psychology, 54(2), 199–231. https://doi.org/10.1111/j.1464-0597.2005.00205.x boekaerts, m., & pekrun, r. (2015). emotions and emotion regulation in academic settings. in l. corno & e. m. anderman (eds). handbook of educational psychology (3rded.) (pp. 76-90). routledge boucsein, w. (2012). electrodermal activity.springer science & business media., 1–8. https://doi.org/10.1007/978-1-4614-1126-0 christopoulos, g. i., uy, m. a., & yap, w. j. (2016). the body and the brain: measuring skin conductance responses to understand the emotional experience. organizational research methods, 1–27. https://doi.org/10.1177/1094428116681073 cleary, t. j. (2006). the development and validation of the self-regulation strategy inventory—self-report. journal of school psychology, 44(4), 307–322. https://doi.org/https://doi.org/10.1016/j.jsp.2006.05.002 dawson, m. e., schell, a. m., & filion, d. m. (2007). the electrodermal system. in j. cacioppo, l. g. tassinary, & g. g. berntson (eds.), handbook of psychophysiology(3rd ed., pp. 159–181). cambridge: cambridge university press. dignath, c., buettner, g., langfeldt, h.-p., & goethe, j. w. (2008). how can primary school students learn self-regulated learning strategies most effectively? educational research review, 3(2), 101–129. https://doi.org/10.1016/j.edurev.2008.02.003 dowson, m., & mcinerney, d. m. (2003). what do students say about their motivational goals? towards a more complex and dynamic perspective on student motivation. contemporary educational psychology,28(1), 91–113. https://doi.org/10.1016/s0361-476x(02)00010-3 duffy, m. c., azevedo, r., sun, n. z., griscom, s. e., stead, v., crelinsten, l., … lachapelle, k. (2015). team regulation in a simulated medical emergency: an in-depth analysis of cognitive, metacognitive, and affective processes. instructional science, 43(3), 401–426. https://doi.org/10.1007/s11251-014-9333-6 eiser, j. r. (1986). social psychology: attitudes, cognition and social behaviour. cambridge university press. elliot, a. j., aldhobaiban, n., kobeisy, a., murayama, k., gocłowska, m. a., lichtenfeld, s., & khayat, a. (2016). linking social interdependence preferences to achievement goal adoption. learning and individual differences, 50, 291-295. fincham, o. e., gasevic, d. v., jovanovic, j. m., & pardo, a. (2018). from study tactics to learning strategies: an analytical method for extracting interpretable representations. ieee transactions on learning technologies. garbarino, m., lai, m., bender, d., picard, r. w., & tognetti, s. (2015). empatica e3 a wearable wireless multi-sensor device for real-time computerized biofeedback and data acquisition. proceedings of the 2014 4th international conference on wireless mobile communication and healthcare “transforming healthcare through innovations in mobile and wireless technologies”, mobihealth 2014 , 39–42. https://doi.org/10.1109/mobihealth.2014.7015904 gašević, d., dawson, s., & siemens, g. (2015). let’s not forget: learning analytics are about learning. techtrends, 59(1), 64–71. gillies, r. m., carroll, a., cunnington, r., rafter, m., palghat, k., bednark, j., & bourgeois, a. (2016). multimodal representations during an inquiry problem solving activity in a year 6 science class: a case study investigating cooperation, physiological arousal and belief states. australian journal of education,0(0), 117. https://doi.org/10.1177/0004944116650701 greeno, g. j. (2005). learning in activity. in r. k. sawyer (ed.), the cambridge handbook of the learning sciences(pp. 79–96). cambridge: cambridge university press. https://doi.org/doi: 10.1017/cbo9780511816833.007 griffin, p., mcgaw, b., & care, e. (2012). the changing role of education and schools. in p. griffin, b. mcgaw, & e. care (eds.), assessment and teaching of 21st century skills(pp. 1-16). dordrecht, germany: springer science+business media b.v. http://dx.doi.org/10.1007/978-94-007-2324-5_2 haataja, e., malmberg, j., & järvelä, s. (2018). monitoring in collaborative learning: co-occurrence of observed behavior and physiological synchrony explored. computers in human behavior, 87(june), 337–347. https://doi.org/10.1016/j.chb.2018.06.007 hadwin, a. f., järvelä, s., & miller, m. (2017). self-regulation, co-regulation and shared regulation in collaborative learning environments. in d. schunk & j. greene (eds.), handbook of self-regulation of learning and performance(2nd ed.). new york, ny: routledge. henriques r., paiva a., & antunes c. (2013) on the need of new methods to mine electrodermal activity in emotion-centered studies. in: cao l., zeng y., symeonidis a.l., gorodetsky v.i., yu p.s., singh m.p. (eds) agents and data mining interaction. admi 2012. lecture notes in computer science, vol 7607. springer, berlin, heidelberg hickey, d. t., & mccaslin, m. (2001). a comparative, sociocultural analysis of context and motivation. in s. volet & s. järvelä (eds.), advances in learning and instruction series. motivation in learning contexts: theoretical advances and methodological implications (pp. 33-55). elmsford, ny: pergamon press. immordino-yang, m. h., christodoulou, j. a., pekrun, r., & linnenbrink-garcia, l. (2014). neuroscientific contributions to emotion measurement in educational contexts. in handbook of emotions in education(pp. 607–624). new york: routledge. isohätälä, j., järvenoja, h., & järvelä, s. (2017). socially shared regulation of learning and participation in social interaction in collaborative learning.international journal of educational research, 81, 11–24. https://doi.org/10.1016/j.ijer.2016.10.006 järvelä, s., & hadwin, a. f. (2013). new frontiers: regulating learning in cscl. educational psychologist, 48(1), 25–39. https://doi.org/10.1080/00461520.2012.748006 järvelä, s., järvenoja, h., & veermans, m. (2008). understanding the dynamics of motivation in socially shared learning. international journal of educational research, 47, 122–135. järvelä, s., järvenoja, h., malmberg, j., isohätälä, j., & sobocinski, m. (2016). how do types of interaction and phases of self-regulated learning set a stage for collaborative engagement?learning and instruction, 43, 39–51. https://doi.org/10.1016/j.learninstruc.2016.01.005 järvelä, s., malmberg, j., sobocinski, m., haataja, e., & kirschner, p. (2018). what multimodal data tell about self regulated learning process? submitted. järvelä, s., salonen, p., & lepola, j. (2001). dynamic assessment as a key to understanding student motivation in a classroom context. in pintrich, p., & maehr, m. (eds.), advances in motivation and achievement (pp.207–240). amsterdam: jai press. järvenoja, h., & järvelä, s. (2009). emotion control in collaborative learning situations do students regulate emotions evoked from social challenges? british journal of educational psychology, 79, 463–481. doi:10.1348/000709909x402811 järvenoja, h., järvelä, s., & malmberg, j. (2015). understanding regulated learning in situative and contextual frameworks.educational psychologist, 50(3), 204–219. https://doi.org/10.1080/00461520.2015.1075400 järvenoja, h., järvelä, s., & malmberg, j. (2017). supporting groups’ emotion and motivation regulation during collaborative learning.learning and instruction, (november), 0–1. https://doi.org/10.1016/j.learninstruc.2017.11.004 järvenoja, h., näykki, p., törmänen, t. & järvelä, s. (2018). emotional regulation in collaborative learning: when do higher education students activate group level regulation in the face of challenges? submitted. järvenoja, h., volet, s., & järvelä, s. (2013). regulation of emotions in socially challenging learning situations: an instrument to measure the adaptive and social nature of the regulation process.educational psychology, 33, 31–58. https://doi.org/10.1080/01443410.2012.742334 kirschner, f., paas, f., & kirschner, p. a. (2011). task complexity and collaborative learning efficiency: the collective working-memory effect. applied cognitive psychology, 25, 615–624. koskey, k. l. k., karabenick, s. a., woolley, m. e., bonney, c. r., & dever, b. v. (2010). cognitive validity of students’ self-reports of classroom mastery goal structure: what students are thinking and why it matters. contemporary educational psychology, 35(4), 254-263. http://dx.doi.org/10.1016/j.cedpsych.2010.05.004 kreibig, s. d. (2010). autonomic nervous system activity in emotion: a review. biological psychology, 84(3), 394–421. https://doi.org/10.1016/j.biopsycho.2010.03.010 kurki, k., järvenoja, h., järvelä, s., & mykkänen, a. (2017). young children’s use of emotion and behaviour regulation strategies in socio-emotionally challenging day-care situations.early childhood research quarterly, 41(june 2016), 50–62. https://doi.org/10.1016/j.ecresq.2017.06.002 kurki, k., järvenoja, h., törmänen, t. & bakhtiar, a. (2018). socio-emotional interaction in collaborative learning: combining individual emotional experiences and group level regulation. paper presented at international conference on motivation, aarhus. kwon, k., liu, y. h., & johnson, l. p. (2014). group regulation and social-emotional interactions observed in computer supported collaborative learning: comparison between good vs poor collaborators.computers and education, 78, 185–200. https://doi.org/10.1016/j.compedu.2014.06.004 levy, i., kaplan, a., & patrick, h. (2004). early adolescents' achievement goals, social status, and attitudes towards cooperation with peers. social psychology of education, 7(2), 127-159. linnenbrink-garcia, l., & pekrun, r. (2011). students’ emotions and academic engagement: introduction to the special issue.contemporary educational psychology, 36(1), 1–3. https://doi.org/10.1016/j.cedpsych.2010.11.004 linnenbrink-garcia, l., rogat, t. k., & koskey, k. l. k. (2011). affect and engagement during small group instruction.contemporary educational psychology, 36(1), 13–24. https://doi.org/10.1016/j.cedpsych.2010.09.001 malmberg, j., järvelä, s., holappa, j., haataja, e., & siipo, a, huang, x. (2018).going beyond what is visible –what physiological measures can reveal about regulated learning in the context of collaborative learning? computers and human behavior. malmberg, j., järvelä, s., & järvenoja, h. (2017). capturing temporal and sequential patterns of self-, co-, and socially shared regulation in the context of collaborative learning.contemporary educational psychology, 49, 160–174. https://doi.org/10.1016/j.cedpsych.2017.01.009 malmberg, j., järvenoja, h., & järvelä, s (2013). patterns in elementary school students´ strategic actions in varying learning situations. instructional science. doi 10.1007/s11251-012-9262-1 mccardle, l., & hadwin, a. (2015). using multiple, contextualized data sources to measure learners’ perceptions of their self-regulated learning. metacognition and learning(vol. 10). https://doi.org/10.1007/s11409-014-9132-0 mcrae, k., gross, j. j., weber, j., robertson, e. r., sokol-hessner, p., ray, r. d., … ochsner, k. n. (2012). the development of emotion regulation: an fmri study of cognitive reappraisal in children, adolescents and young adults. social cognitive and affective neuroscience, 7 (1), 11–22. http://doi.org/10.1093/scan/nsr093 miyake, n. (1986). constructive interaction and the iterative process of understanding. cognitive science, 10, 151–177. doi: 10.1207/s15516709cog1002_2 molenaar, i., chiu, m. m. (2014). dissecting sequences of regulation and cognition: statistical discourse analysis of primary school children’s collaborative learning. metacognition and learning, 9, 137-160. doi:10.1007/s11409-013-9105-8 mykkänen, a., perry, n., & järvelä, s. (2017). finnish students’ reasons for their achievement in classroom activities: focus on features that support self-regulated learning. education 3-13, 45 (1), 1–16. https://doi.org/10.1080/03004279.2015.1025802 näykki, p., isohätälä, j., järvelä, s., pöysä-tarhonen, j., & häkkinen, p. (2017). facilitating socio-cognitive and socio-emotional monitoring in collaborative learning with a regulation macro script – an exploratory study. international journal of computer-supported collaborative learning ,12(3), 251–279. https://doi.org/10.1007/s11412-017-9259-5 näykki, p., järvelä, s., kirschner, p. a., & järvenoja, h. (2014). socio-emotional conflict in collaborative learning-a process-oriented case study in a higher education context.international journal of educational research, 68, 1–14. https://doi.org/10.1016/j.ijer.2014.07.001 nolen, s. b. (2006). validity in assessing self-regulated learning: a comment on perry & winne. educational psychology review,18(3), 229-232. https://doi.org/10.1007/s10648-006-9016-1 nolen, s. b., horn, i. s., & ward, c. j. (2015). situating motivation.educational psychologist, 50(3), 234–247. https://doi.org/10.1080/00461520.2015.1075399 ochoa, x. (2017). multimodal learning analytics. in c. lang, g. siemens, a. wise, & d. gasevic (eds.), handbook of learning analytics (pp. 129–141). solar. society for learning analytics research. paris, s. g., & turner, j. c. (2012). situated motivation. in student motivation, cognition, and learning (pp. 229-254). routledge.pekrun, r. (2016). academic emotions. in k. r, wentzel, d. b. miele (eds.), handbook of motivation at school(pp.120-146). new york, ny: routledge. pekrun, r., goetz, t., titz, w., & perry, r. p. (2002). academic emotions in students’ self-regulated learning and achievement: a program of qualitative and quantitative research. educational psychologist,37(2), 91–106. https://doi.org/10.1207/s15326985ep3702 pintrich, p. r. (2000). the role of goal orientation in self-regulated learning. in m. boekaerts. p.r. pintrich & m. zeidner (eds.), handbook of self-regulation, (pp. 451–502) san diego, ca: academic press. pintrich, p. r. (2003). multiple goals and multiple pathways in the development of motivation and selfregulated learning. in l. smith, c. rogers, & p. tomlinson (eds.), development and motivation: joint perspectives (pp. 137–155). leicester: british psychological society. reimann, p., markauskaite, l. and bannert, m. (2014), e‐research and learning theory. british journal of educational technology, 45(3), 528–540. https://doi.org/10.1111/bjet.12146 roschelle, j., & teasley, s. (1995). the construction of shared knowledge in collaborative problem solving. in c. e. o'malley (ed.), computer-supported collaborative learning (pp. 69–97). heidelberg: springer-verlag. rogat, t. k., & adams-wiggins, k. r. (2015). interrelation between regulatory and socioemotional processes within collaborative groups characterized by facilitative and directive other-regulation.computers in human behavior, 52, 589–600. https://doi.org/10.1016/j.chb.2015.01.026 russell, j. a., & barrett, l. f. (1999). core affect, prototypical emotional episodes, and other things called emotion: dissecting the elephant. journal of personality and social psychology,76(5), 805–819. https://doi.org/10.1037/0022-3514.76.5.805 sawyer, k. (2014). introduction: the new science of learning. in r. k. sawyer, (ed.), the cambridge handbook of the learning sciences(2nd ed., pp. 1-18). new york, ny: cambridge university press. schutz, p. a., & pekrun, r. (2007). introduction to emotion in education. in p. a. schutz & r. pekrun (eds.), emotion in education(pp. 3–10). burlington: elsevier inc. sobocinski, m., malmberg, j., & järvelä, s. (2017). exploring temporal sequences of regulatory phases and associated interactions in lowand high-challenge collaborative learning sessions.metacognition and learning(vol. 12). https://doi.org/10.1007/s11409-016-9167-5 sonnenberg, c., & bannert, m. (2015). discovering the effects of metacognitive prompts on the sequential structure of srl-process using process mining techniques. journal of learning analytics, 2 (1), 72–100. törmänen, t., järvenoja, h., kurki, k., devai, r., & järvelä, s. (2018, august). students’ emotional valence and physiological arousal during collaborative learning – a case study . poster presented at 16th international conference on motivation 2018, aarhus, denmark. usher, e. l., & pajares, f. (2008). sources of self-efficacy in school: critical review of the literature and future directions. review of educational research, 78(4), 751–796. https://doi.org/10.3102/0034654308321456 volet, s., & järvelä, s. (eds.). (2001). advances in learning and instruction series. motivation in learning contexts: theoretical advances and methodological implications. elmsford, ny: pergamon press. volet, s., & vauras, m. (2013). interpersonal regulation of learning and motivation: methodological advances . new york: routledge. volet, s., vauras, m., & salonen, p. (2009). self-and social regulation in learning contexts: an integrative perspective. educational psychologist, 44(4), 215-226.webb, n. m., troper, j. d., & fall, r. (1995). constructive activity and learning in collaborative small groups. journal of educational psychology, 87(3), 406–423. doi: 10.1037/0022-0663.87.3.406 whitebread, d., coltman, p., pasternak, d. p., sangster, c., grau, v., bingham, s., … demetriou, d. (2009). the development of two observational tools for assessing metacognition and self-regulated learning in young children. metacognition and learning, 4, 63–85. https://doi.org/10.1007/s11409-008-9033-1 winne, p. h., & hadwin, a. f. (1998). studying as self-regulated learning. in d. j. hacker, j. dunlosky, & a. graesser (eds.), metacognition in educational theory and practice(pp. 277–304). hillsdale, nj. winne, p.h., & jamieson-noel, d.l. (2002). exploring students’ calibration of selfreports about study tactics and achievement. contemporary educational psychology , 27 , 551–572. winne, p.h., & perry, n.e. (2000). measuring self-regulated learning. in p. pintrich, m. boekaerts & m. zeidner (eds.), handbook of self-regulation(pp. 531-566). orlando, fl: academic press. wise, a. f., & schwarz, b. b. (2017). visions of cscl: eight provocations for the future of the field. international journal of computer-supported collaborative learning , 12(4), 423-467. wolters, c. a., & benzon, m. b. (2013). assessing and predicting college students’ use of strategies for the self-regulation of motivation. the journal of experimental education, 81(2), 199-221.zheng, l., & yu, j. (2016). exploring the behavioral patterns of co-regulation in mobile computer-supported collaborative learning. smart learning environments, 3(1), 1-20. 10.1186/s40561-016-0024-4 zimmerman, b., & schunk, d. h. (eds.). (2011). handbook of self-regulation of learning and performance. new york, ny: taylor & francis frontline learning research vol. 5 no. 3 special issue (2017) 43 54 issn 2295-3159 corresponding author: sharon e. fox, m.d., ph.d., dep. of pathology and laboratory medicine, southeast louisiana veterans healthcare system, new orleans email: sharon.fox4@va.gov doi: http://dx.doi.org/10.14786/flr.v5i3.258 eye-tracking in the study of visual expertise: methodology and approaches in medicine sharon e. foxa,b & beverly e. faulkner-jonesa abeth israel deaconess medical center, boston, ma, usa; harvard medical school, boston, ma, usa blsu health sciences center, new orleans, la, usa article received 7 may / revised 23 march / accepted 24 march / available online 14 july abstract eye-tracking is the measurement of eye motions and point of gaze of a viewer. advances in this technology have been essential to our understanding of many forms of visual learning, including the development of visual expertise. in recent years, these studies have been extended to the medical professions, where eye-tracking technology has helped us to understand acquired visual expertise, as well as the importance of visual training in various medical specialties. medical decision-making involves a complex interplay between knowledge and sensory information, and the study of eye-movements can reveal the mechanisms involved in acquiring the visual component of these skills. eye-tracking studies have even been extended to develop computational models of procedures for “expert” skill assessment, and to eliminate potential sources of error in image-based diagnostics. this review will examine the current eye-tracking frontier for the study of visual expertise, with specific application to medical professions. keywords: eye-tracking; visual expertise; digital imaging http://dx.doi.org/10.14786/flr.v5i3.258 fox et al | f l r 44 1. introduction eye-movements and their meaning have long been the subject of scientific study, and recent advances in technology have allowed for the study of eye gaze – or eye-tracking – during the acquisition of complex forms of visual expertise. in this review of eye-tracking methodology, we will examine the way in which the science of acquired visual expertise, as well as the development of eye-tracking technology, have allowed us to better understand visual training in medicine. training and expertise in medicine involves not only acquisition of knowledge, but also the integration of sensory information during the process of diagnosis and disease management. in many instances, that sensory information is visual, and the study of eye-movements can reveal not only the cognitive processes behind medical expertise, but also the mechanisms involved in acquiring these skills. furthermore, studies of expert gaze patterns can help us to understand common perceptual pitfalls, and to develop technologies that may assist in training medical professionals and eliminate sources of error. 2. history & eye-tracking mechanisms eye movements represent one of the most frequent sensorimotor activities in humans. large scanning movements, or saccades, typically occur 3-4 times per second (holmqvist, 2011). the most frequently reported eye gaze metric, however, represents the relative “pause” between saccades, known as fixations. individual fixations generally last approximately 200-300 milliseconds, with the fovea of the eye remaining relatively still along a point of gaze (holmqvist, 2011). in the early 1970s, eye-tracking techniques were advanced by using video-based techniques, wherein recorded features of reflections of light from the eye could be systematically tracked. one option was to scan for the lack of reflectance from the pupil (“dark-pupil” tracking), although low contrast between the pupil and dark-brown irises led to suboptimal results. if the eye is lit from the front, the light will bounce off the back of the lens and appear very bright (“bright-pupil” tracking). this bright circle can then be more reliably detected than the dark-pupil technique (merchant, morrissette, & porterfield, 1974; rayner, 1978). in addition to the ability to capture the motion of the eyes, a second and important component of eye-tracking for the study of vision is the ability to capture the visual stimulus in such a way that gaze can be determined. most early eye-tracking studies involved significant restraint of the head for this purpose. an important innovation in this area was the development of eye-tracking systems that measured multiple features of the eye in order to infer position relative to a visual stimulus (holmqvist, 2011). optical properties of the eye, such as corneal reflection and pupil location, vary differently under conditions of head versus eye movement, and their recordings can be used to solve for the actual gaze point of the viewer (for additional discussion, (wolfe, evans, drew, aizenman, & josephs, 2015)). by utilizing non-visible light, such as near-infrared light sources, such eye-trackers can be made less noticeable, and easier to use in a variety of lighting conditions. such systems are particularly useful in the study of natural learning environments, and allow results to be generalized to real world situations. the balance between obtaining a high-precision record of an observer's point-of-regard and allowing natural headand body-movements is where much of the technological advancement in eye-tracking has arisen over the past twenty years, in addition to "scene cameras" that allow eye-tracking data to be superimposed on the naturalistic point-of-view of the participant during movement (browatzki, bulthoff, & chuang, 2014; huette, winter, matlock, & spivey, 2012; johnson, liu, thomas, & spencer, 2007). computer embedded and table-mounted remote optical eye-trackers have also been developed that allow some natural head movement while sitting in front of a computer screen for two-dimensional stimulus presentation. such advances have placed fewer constraints on experimental subjects, such that modern researchers are able to record the eye movements of freely moving subjects carrying out everyday tasks. fox et al | f l r 45 these features, combined with improvement in user interface, have allowed eye-tracking to be used in the study of visual training, and the application of visual training within the field of medicine. it is also possible that future advances in this technology could allow eye-tracking devices to assist visual processes in medical education and decision-making. 3. eye-tracking applied to medical expertise 3.1 introduction eye-tracking provides a potential means for understanding the nature and acquisition of visual expertise as it relates to medical knowledge. as a method of assessing gaze patterns during task execution, eye-tracking has been used to understand which parts of an image are important to medical decision-making, and also the process which experts adopt to analyze these images. to date, the majority of the published eyetracking studies involving medical expertise are naturally in the fields requiring extensive visual training, such as radiology (drew, evans, vo, jacobson, & wolfe, 2013; kundel, nodine, krupinski, & mellothoms, 2008; g. tourassi, voisin, paquit, & krupinski, 2013; g. d. tourassi, mazurowski, harrawood, & krupinski, 2010; wolfe et al., 2015). these studies often utilize static images, which are relatively easy to employ in eye-tracking experimental designs and analyses, but which don’t always translate to the visual environments of other clinical specialties. in recent years, additional studies have appeared in relation to fields such as dermatology, surgery, and anatomic pathology (ahmidi, ishii, fichtinger, gallia, & hager, 2012; bombari, mora, schaefer, mast, & lehr, 2012; brunye et al., 2014; fox, law, & faulkner-jones, 2017; e. krupinski, chao, hofmann-wellenhof, morrison, & curiel-lewandrowski, 2014; e. a. krupinski, graham, & weinstein, 2013; e. a. krupinski et al., 2006; law, atkins, lomax, & wilson, 2003; tiersma, peters, mooij, & fleuren, 2003). the ability to use dynamic stimuli in eye-tracking research (ahmidi, ishii, et al., 2012; drew, vo, olwal, et al., 2013; fox et al., 2017; mallett et al., 2014), as well as the use of wearable and increasingly portable devices has expanded the opportunities in which eye-tracking can be utilized to understand the gain of visual expertise in clinical settings. in the study of medical education, most experimental designs involve two or three participants groups, each at a different level of medical training (ahmidi, ishii, et al., 2012; e. a. krupinski et al., 2013; mallett et al., 2014; phillips et al., 2013). analyses vary widely, however, with some reporting qualitative information, such as regions of an image with the greatest or least attention, and others using quantitative analyses of fixations in relation to defined areas of interest (aois) (ahmidi, ishii, et al., 2012; brown et al., 2014; fox et al., 2017; e. a. krupinski et al., 2013; mallett et al., 2014; phillips et al., 2013). a recent, and notable study of the use of eye-tracking in medical education involved the use of expert eye-tracking data to train observational techniques (h. jarodska et al., 2012), and will be further discussed in subsequent sections. overall, eye-tracking studies across medical specialties have suggested that more experienced physicians require fewer fixations, and less time spent on areas of interest, while performing at a higher rate of accuracy than novices. the visual patterns identified by eye-tracking experiments, however, depend in part on the type of visual expertise acquired, and the integration of medical knowledge in that process. in order to understand how eye-tracking methodology may be applied to the study of visual expertise in medicine, we must first understand the differences between visual tasks in a variety of clinical scenarios. 3.2 search-related expertise radiology is the most extensively studied field of medicine in relation to visual expertise, and it is not surprising that a significant literature involving eye-tracking has arisen in relation to this specialty (drew, evans, et al., 2013; kundel et al., 2008; g. tourassi et al., 2013; g. d. tourassi et al., 2010; wolfe et al., fox et al | f l r 46 2015). the visual expertise learned by radiologists is an example of a search-related task, in which the radiologist identifies visual “targets” in an image containing both expected and “distractor” elements (wen et al., 2016; wolfe et al., 2015). screening tests, both in radiology and pathology, are generally performed using a visual search model (stewart et al., 2007; wolfe, 1995). in this visual task, search is required because everything in the visual field cannot be identified and processed simultaneously. object recognition is limited to one, or a small number of objects at one time (wolfe, 2012b; wolfe et al., 2015). attention may appear random, but is often guided by multiple cognitive mechanisms. at the most basic level, exploratory visual gaze is directed towards items with “bottom-up” salience (braun, 1994; wolfe & horowitz, 2004). these items do not depend on the purpose of the search, and therefore are not the result of acquired visual expertise. by contrast, gaze patterns of expert radiologists during search tasks may be less affected by bottom-up salience when these elements are known distracters from a visual target (wolfe et al., 2015). bottom-up salience of a visual item is determined by basic features, such as color, size, contrast, movement and luminosity (braun, 1994; wolfe & horowitz, 2004; wolfe, horowitz, kenner, hyle, & vasan, 2004). wolfe et al. (2015) classifies these attributes as “pre-attentive,” because they do not require a specific goal or form of expert attention to bias patterns of gaze. through medical training, including knowledge acquisition and an understanding of the visual properties of targets, “top-down,” or user-guided visual search develops (draschkow, wolfe, & vo, 2014; wolfe, 1994). at this level of visual processing, the radiologist guides attention toward a mental representation of the features and potential location of a target (wen et al., 2016). top-down processing is also employed at the point of diagnosis, when the visual item is matched to both a mental depiction of the target, and general medical knowledge (drew, evans, et al., 2013; drew, vo, olwal, et al., 2013). wolfe et al. (1994) summarize this process in the “guided search model,” in which bottom-up attention to distracters is at least partially suppressed by the top-down effects of visual expertise. the expert’s knowledge provides scene guidance to relevant parts of the image, which combines with these effects to create a mental representation of the likely location of targets (wen et al., 2016). this course of attentional gaze may be modulated by systematic training algorithms for visual inspection (for example, a trained pattern of attention to each potential diagnostic target within a chest radiograph), as well as clinical information (draschkow et al., 2014; drew, evans, et al., 2013; wolfe et al., 2015; wolfe et al., 2007). eye-tracking has been utilized to better understand the development of a “priority map,” within a radiological image, and how this priority map may evolve during the diagnostic process. it is proposed that the priority map must change as the radiologist’s eyes move about the image as the salience of items will change with their distance from the present fixation (wen et al., 2016). interestingly, the development of visual expertise with training in radiology seems to indicate a move towards efficiency, and away from repeated attention to “non-priority” regions of the diagnostic image (g. tourassi et al., 2013; wolfe et al., 2015). search for multiple visual targets, held in memory, is known as “hybrid search” (wolfe, 2012a). attention is not guided as effectively in this form of search, which may play a role in the training process of visual expertise in medicine (eckstein, 2011). this effect can be highlighted as it relates to simulation studies in medical education. in a study of senior nursing students administering medications in a clinical simulation setting, 40% administered a contraindicated medication to a patient with a known allergy (b. amster et al., 2015). eye-tracking data in this educational study was used to determine whether students administered a contraindicated medication due to a knowledge deficit, or because the information required was not visualized. in this case, the necessary information was visualized by all students, and the deficit in knowledge related to pharmacology could be corrected. in addition, acquired knowledge through experience in the field, as well as risk-aversion in certain types of medical cases, may lead to altered salience depending upon “value.” the learned value of visual stimuli significantly affects attentional priority (laurent, hall, anderson, & yantis, 2015; sali, anderson, & yantis, 2014). in basic eye-tracking experiments involving manipulation of object low level properties, participants quickly learned to search for objects of a property that would produce a valuable reward (laurent et al., 2015; sali et al., 2014). thus, if a radiologist is “rewarded,” for example, for finding cancerous lesions as opposed to incidental pathology, these objects may affect salience maps. fox et al | f l r 47 in contrast to the study of errors due to a knowledge deficit (amster et al., 2015), eye-tracking has also been utilized to identify the errors which can occur during visual search guided by expertise. in a version of the “invisible gorilla” study drew et al. utilized an eye-tracker to follow the eye movements of radiologists as they searched for lung nodules in a serial stack of ct images. on the last case, a gorilla was inserted into the lung, and the majority of radiologists failed to notice this (drew, vo, & wolfe, 2013). eyetracking was able to demonstrate that this was not because the radiologists were negligent – they did, in fact, attend to the region of the gorilla – but their expert attention to expected visual targets, acquired through medical training, led to an “inattentional blindness” to the gorilla (drew, evans, et al., 2013; drew, vo, & wolfe, 2013). this particular study illustrates the way in which eye-tracking can be used to study the visual process of “missing” an obvious abnormality as compared to the presence of disease. wolfe and van wert (2010) have also discussed the importance of prevalence effects in designing appropriate eye-tracking experiments for understanding medical expertise. screening tasks such as screening for breast cancer on a mammogram, or cervical cancer on a pap smear, represent an important class of search with low prevalence of visual targets within the population screened (wolfe & van wert, 2010). studies of vigilance and low-prevalence search tasks have shown that rare events are missed more often than common ones (wolfe et al., 2007; wolfe & van wert, 2010). wolfe and van wert (2010) explain that when targets are rare, observers are more likely to reject an ambiguous target. by contrast, participants are much more likely to label an ambiguous item as a target when prevalence is high. this effect is attributed to an unconscious decision rule that is changing (wolfe & van wert, 2010). observers also become faster to declare themselves to be finished with an image under low-prevalence conditions, but forcing them to slow down does not make observers less likely to reject low-prevalence targets (wolfe, 2012b). importantly, evans et al. (2013) found the same effect to be true for both experts and novices. in an experiment involving manipulation of prevalence in the setting of mammography, radiologists had a false-negative rate of 12 % in the setting of high prevalence, and 30% at low prevalence (evans, birdwell, & wolfe, 2013). similar results were found with cytologists reading cervical cancer screening slides (evans, tambouret, evered, wilbur, & wolfe, 2011). in all cases, rare targets were missed more often. this has important implications for the generalizability of eye-tracking experiments related to visual expertise, as it suggests that prevalence in the experimental setting may affect visual performance in a way that is not seen in clinical practice. in some cases, prevalence in an experimental setting may actually offer the opportunity for increased training (jarodska et al., 2012), and the subsequent improvement of visual patterns. while some studies have demonstrated an acquisition of a pattern of “expert” gaze with focused training on gaze patterns (jardoska et al., 2012; e.a. henneman et al., 2014), it would be highly informative to examine the effects of such training over a prolonged period of time to ascertain whether this represents a permanent shift due to the use of eyetracking as an educational tool. furthermore, the patterns of vision seen in the search-related expertise employed for medical screening may differ significantly from the examination of images from which a pathologic diagnosis is expected. 3.3 “gestault” or holistic expertise we have discussed several studies in which eye-tracking has elucidated search-related features of visual expertise, however some forms of visual expertise acquired in medicine involve a visual categorization or “gestault” assessment. dermatologists, for example, are often required to diagnose a clearly visible skin lesion or rash from clinical appearance. similarly, pathologists are frequently asked to render a diagnosis from a distinct image of a lesion or cell type. while search-related expertise is certainly employed in both of these fields, a significant proportion of visual expertise is devoted to recognition and identification, rather than locating the target. kundel (2008) noted that visual search was not necessary for all radiologic images, and he incorporated the idea of holistic visual processing in radiologic diagnosis into eyetracking experiments. the ability to interpret complex visual information in a short period of time is common to all humans, most notably in the form of face perception (bukach, gauthier, & tarr, 2006; rossion, collins, goffaux, & curran, 2007). it is likely that this technique is employed by medical experts fox et al | f l r 48 who have significant experience with the visual characteristics of a diagnostic entity. experts in these fields may describe recognizing an image like one does an acquaintance, which implies visual processing that may be similar to face processing – a visual task long studied with eye-tracking methodology (gauthier & nelson, 2001; pascalis et al., 2005; vanderwert et al., 2015; wagner, hirsch, vogel-farley, redcay, & nelson, 2013). the eye-tracking methods used to study holistic visual processing during medical training differ from those employed for search-related expertise. one study of holistic processing asked if mammographers could look at a bilateral mammogram for less than a second and determine if the woman should be called back (evans, georgian-smith, tambouret, birdwell, & wolfe, 2013). the technique of a brief exposure is often useful in identifying holistic processing, as well as the most important areas of interest for the rapid interpretation of images by experts (evans, georgian-smith, et al., 2013; e. krupinski et al., 2014; kundel et al., 2008). with increasing medical expertise, several studies have also shown specific changes in eye movement patterns (drew, vo, olwal, et al., 2013; fox et al., 2017; e. a. krupinski et al., 2013). characteristically, trainees make more eye movements when evaluating an image than do experts, and those eye movements cover more of the area of the image. this development of visual efficiency is also seen in the development of human face processing from infancy to adulthood, suggesting another link between these two forms of holistic visual processing (fox et al., 2017; gauthier & nelson, 2001; pascalis et al., 2005). we have noted in preliminary work that while this form of visual efficiency seems to be naturally acquired over time, more efficient patterns can develop as a result of visual tools designed to enhance the educational process (fox et al., 2017). this is similar to the finding that training for efficiency of gaze, as well as accuracy in the assessment of infant seizures, could be enhanced through the use of image visualizations based upon expert eye-tracking data (jarodska et al., 2012). 3.4 hand-eye and procedural expertise in recent years, the importance of visual as well as tactile expertise in medical procedures has been recognized, and several eye-tracking studies have examined the type of visual expertise acquired in the training of medical procedural skills (ahmidi, ishii, et al., 2012; law et al., 2003). one such study compared the eye movements utilized by expert and novice surgeons performing a laparoscopic procedure in a computer-based simulator (law et al., 2003). the results of this study showed that novices needed more visual feedback of the tool position to complete the task than did experts. in addition, the experts tended to maintain eye gaze on the surgical target while manipulating surgical instruments, whereas novices were more varied in eye-hand coordination, and often tracked the surgical tool rather than the surgical target (law et al., 2003). the development of robotic and laparoscopic surgical instruments has also allowed for the acquisition of kinematic data related to common procedures, and skill level can be evaluated with these measures in conjunction with eye-tracking data (ahmidi, ishii, et al., 2012). data from robotic surgical systems show that hidden markov models (hmm) can facilitate the recognition of surgical skill level (ahmidi, hager, ishii, gallia, & ishii, 2012; ahmidi, ishii, et al., 2012; ahmidi et al., 2015). methods for the assessment of surgical skill utilizing eye-tracking as a quantitative measure have also been developed (ahmidi, ishii, et al., 2012). in the experiment by ahmidi et al. (2012), sinus procedures were performed by experts and novices on cadavers with the use of an endoscope and a visualization screen. a 50hz remote eye-tracker was utilized in this case, with alignment to video data collected through the endoscope. the results of hmm generated from both kinematic and eye-tracking data reveal that eye-gaze does contain expertise-related structures, and the addition of this data to kinematic information improves models of skill expertise by 13.2% for expert and 5.3% for novice levels (ahmidi, ishii, et al., 2012). models combining both measures can reportedly quantify a surgeon’s skill level on a specific procedure with an accuracy of 82.5% (ahmidi, ishii, et al., 2012). in the field of medical learning, several studies involving all levels of medical professionals have used individual eye-tracking data for both training, and debriefing after simulations (jarodska et al., 2012; henneman et al., 2014). as a debriefing strategy for medical simulations, eye-tracking can offer useful fox et al | f l r 49 information about errors that cannot be readily observed or verbalized. further, allowing medical trainees to observe expert scanpaths derived from controlled eye-tracking has been proven to be more effective than verbal didactics describing the methods of visual assessment. the scanpaths of trainees in these procedural settings can sometimes be used to assess best practices in medicine, as well as assure competency. this was performed as a follow-up to the study of errors in medication administration (marquard et al, 2011; amster, 2015), to analyze patterns of gaze most associated with identification errors among nurses. nurses who recognized errors were more likely to focus on one piece of identification information at one time, comparing medication labels and the corresponding patient information on an id badge in sequence, as opposed to reading either the badge or medication bottle in full before changing the point of fixation. eyetracking has also been a useful technique for the study of real world disruptions in the medical training environment. for example, in simulated emergency room settings, medical professionals who visually engaged an interruption during a visual task were more likely to commit an error (marquard et al., 2011). taken together, these studies of gaze patterns in procedural settings suggest a role for eye-tracking as both a training and assessment tool in medical education, and the necessity of a realistic clinical environment for the application of trained visual patterns of gaze. with the continuing trend towards quantitative assessment of procedural skill in medical training, these early studies suggest a role for eye-tracking methodology in the study and evaluation of visual training as one component of technical proficiency. 3.5 three-dimensional and dynamic visual stimuli several of the studies mentioned have necessitated the use of visual stimuli that require manipulation by the viewer, such as three-dimensional images and dynamic video recordings. several recent studies provide evidence that radiologists examining ct images in three-dimensions developed visual patterns that involved maintaining little movement in one dimension while scanning across the plane of the other two dimensions (drew, vo, olwal, et al., 2013; wen et al., 2016). two distinct techniques within this pattern could be clustered, but no significant superiority of one technique over another was demonstrated. recent advances in software for analyzing dynamic scenes or areas of interest allow for greater opportunity to investigate expertise with these types of medical images in the clinical environment (holmqvist, 2011; mallett et al., 2014; phillips et al., 2013). in the field of pathology, for example, advances in digital imaging have allowed for the development of three-dimensional models (jeong et al., 2010; ward, rosen, law, rosen, & faulkner-jones, 2015), as well as whole-slide images, which more closely replicate the experience of microscopy (fallon, wilbur, & prasad, 2010; fox et al., 2017). while some educational tasks involve the use of static images, the diagnostic process of anatomic pathology almost always involves movement of the slide image, as well as adjustment of magnification. while several studies of eye-tracking have focused on specific features of digital pathology images (bombari et al., 2012; fox et al., 2017; e. krupinski et al., 2014; e. a. krupinski et al., 2013; e. a. krupinski et al., 2006; tiersma et al., 2003), we now have the tools available to examine the effects of expertise upon the dynamic diagnostic process. the authors have examined pathology trainees at the early and late stages of training, and found support for the use of specific dynamic digital platforms in the acquisition of “expert” patterns of gaze (fox et al., 2017). as mentioned previously, it may be possible to use specific viewing platforms to teach a pattern of evaluation that replicates that of medical experts (fox et al., 2017; jarodska et al., 2012). this could potentially involve the observation of expert gaze during a procedure or visual diagnostic process, the direction of trainee gaze through focused visualization guided by this data (jardoska et al., 2012), or the development of training viewers designed to improve gaze direction and efficiency. fox et al | f l r 50 4. constraints of eye-tracking experiments while there are many important questions related to the study of visual expertise in medicine, a major constraint on these studies is always participant number and available time. medical professionals are a limited participant resource, and they often do not have the time to participate in experimental settings. furthermore, there is a limit to how one can manipulate clinical practice for research purposes, and the generalizability of experimental findings to real-world workflow. for this reason, most studies involving eye-tracking are designed to utilize only a small number of participants (often 5-15 per group), as well as fewer visual exemplars to allow for time efficiency. the invention of increasingly portable eye-trackers that can be integrated with a variety of cameras or image interfaces allows for greater accessibility, and perhaps studies which can occur within real clinical settings. in a study of operating room technicians utilizing circulation machines, eye-tracking revealed that expert technicians visually fixated upon a larger number of critical sources of information during the operative procedure as compared to novices (y. tomizawa et al., 2012). with knowledge of these differences, it may be possible to incorporate eye-tracking as a component of self-assessment in early medical practice. a key feature in this form of training would be the instruction of optimal gaze – either in a didactic, or simulation environment – followed by evaluation in applicable realworld scenarios. the collection of expert gaze patterns in the clinical environment, as well as the assessment of trainees in real applications, will allow for greater understanding of the generalizability of the results derived from research settings. furthermore, the feedback of scanpath and fixation data to individual medical trainees may prove useful as not only a research, but an educational tool. in two separate studies, comparisons of standard verbal debriefing of medical trainees following a simulated patient encounter, debriefing involving videos of their scanpaths, and debriefing involving both verbal and scanpath information, revealed that addition of eye-tracking data most significantly improves subsequent performance (jarodska et al., 2012; henneman et al., 2014). it is likely that the increasing ease of use of eye-tracking tools will allow investigators to recruit a greater number of “expert” participants, and thereby provide accurate and generalizable data to the medical community. 5. conclusion visual expertise is an important component of medical learning, and eye-tracking is one method by which we can better understand this skill and its acquisition in relation to the clinical workflow. studies to date have divided visual expertise among categories of search, image categorization, and procedural skill, with significant overlap of these tasks between fields of medicine. with the continued development of cheaper, faster, and more ergonomic eye-tracking devices, it is likely that we will have many opportunities to study the increasing number of professional tasks requiring visual expertise, and to utilize these results for improved medical training and quality of care. the use of eye-tracking in the training of visual medical expertise, as well as self-evaluation, has the potential to impact overall competency. visualization of the eye movements of expert clinicians may provide insights into a diagnostic process, or the means to avoid medical errors. these tools can enhance the traditional processes of learning through lectures, simulations, or observational sessions. specifically, the process of observing the scanpath or directed gaze of a medical expert during a visual task can improve trainee performance, while eye-tracking data from trainees can provide a method of feedback and self-assessment to students. finally, eye-tracking research may allow us to understand the complex patterns of gaze that underlie diagnostic reasoning, and provide further insight into additional learning methods that improve upon clinical expertise. fox et al | f l r 51 keypoints eye-tracking technology has evolved from one-dimensional photographic techniques, to noninvasive and increasingly portable methods that can be used in a variety of medical setting. visual expertise in medicine is acquired in conjunction with clinical knowledge, and can be characterized as search-related, holistic, or in association with kinematic skills. eye-tracking can assist in the assessment of expertise, as well as address human errors in visually based medical decision-making. acknowledgments we would like to give special thanks to dr. charles nelson, and dr. jeremy wolfe for their expertise and guidance in writing this manuscript. in addition, we would like to give thanks to nih support for dr. beverly faulkner-jones (nih: nibib sbir grant 2r44eb013518-02a1) and to dr. sharon fox (nih: nibib 3r25ns070682-04s1). references ahmidi, n., hager, g. d., ishii, l., gallia, g. l., & ishii, m. (2012). robotic path planning for surgeon skill evaluation in minimally-invasive sinus surgery. med image comput comput assist interv, 15(pt 1), 471-478. doi: 10.1007/978-3-642-33415-3_58 ahmidi, n., ishii, m., fichtinger, g., gallia, g. l., & hager, g. d. (2012). an objective and automated method for assessing surgical skill in endoscopic sinus surgery using eye-tracking and tool-motion data. int forum allergy rhinol, 2(6), 507-515. doi:10.1002/alr.21053 ahmidi, n., poddar, p., jones, j. d., vedula, s. s., ishii, l., hager, g. d., & ishii, m. (2015). automated objective surgical skill assessment in the operating room from unstructured tool motion in septoplasty. int j comput assist radiol surg, 10(6), 981-991. doi:10.1007/s11548-015-1194-1 amster b., marquard j., henneman e., fisher d. (2015). using an eye tracker during medication administration to identify gaps in nursing students' contextual knowledge: an observational study. nurse educ, 40:83–86. doi: 10.1097/nne.0000000000000097 bombari, d., mora, b., schaefer, s. c., mast, f. w., & lehr, h. a. (2012). what was i thinking? eyetracking experiments underscore the bias that architecture exerts on nuclear grading in prostate cancer. plos one, 7(5), e38023. doi:10.1371/journal.pone.0038023 braun, j. (1994). visual search among items of different salience: removal of visual attention mimics a lesion in extrastriate area v4. j neurosci, 14(2), 554-567. browatzki, b., bulthoff, h. h., & chuang, l. l. (2014). a comparison of geometricand regression-based mobile gaze-tracking. front hum neurosci, 8, 200. doi:10.3389/fnhum.2014.00200 brown, p. j., marquard, j. l., amster, b., romoser, m., friderici, j., goff, s., & fisher, d. (2014). what do physicians read (and ignore) in electronic progress notes? appl clin inform, 5(2), 430-444. doi:10.4338/aci-2014-01-ra-0003 brunye, t. t., carney, p. a., allison, k. h., shapiro, l. g., weaver, d. l., & elmore, j. g. (2014). eye movements as an index of pathologist visual expertise: a pilot study. plos one, 9(8), e103447. doi:10.1371/journal.pone.0103447 bukach, c. m., gauthier, i., & tarr, m. j. (2006). beyond faces and modularity: the power of an expertise framework. trends cogn sci, 10(4), 159-166. doi:10.1016/j.tics.2006.02.004 crane, h. d., & steele, c. m. (1985). generation-v dual-purkinje-image eyetracker. appl opt, 24(4), 527. https://doi.org/10.1097/nne.0000000000000097 fox et al | f l r 52 dodge, r., & cline, t. s. (1901). the angle velociy of eye movements. psychological review, 8(2), 145157. doi:http://dx.doi.org/10.1037/h0076100 draschkow, d., wolfe, j. m., & vo, m. l. (2014). seek and you shall remember: scene semantics interact with visual search to build better memories. j vis, 14(8), 10. doi:10.1167/14.8.10 drew, t., evans, k., vo, m. l., jacobson, f. l., & wolfe, j. m. (2013). informatics in radiology: what can you see in a single glance and how might this guide visual search in medical images? radiographics, 33(1), 263-274. doi:10.1148/rg.331125023 drew, t., vo, m. l., olwal, a., jacobson, f., seltzer, s. e., & wolfe, j. m. (2013). scanners and drillers: characterizing expert visual search through volumetric images. j vis, 13(10). doi:10.1167/13.10.3 drew, t., vo, m. l., & wolfe, j. m. (2013). the invisible gorilla strikes again: sustained inattentional blindness in expert observers. psychol sci, 24(9), 1848-1853. doi:10.1177/0956797613479386 eckstein, m. p. (2011). visual search: a retrospective. j vis, 11(5). doi:10.1167/11.5.14 evans, k. k., birdwell, r. l., & wolfe, j. m. (2013). if you don't find it often, you often don't find it: why some cancers are missed in breast cancer screening. plos one, 8(5), e64366. doi:10.1371/journal.pone.0064366 evans, k. k., georgian-smith, d., tambouret, r., birdwell, r. l., & wolfe, j. m. (2013). the gist of the abnormal: above-chance medical decision making in the blink of an eye. psychon bull rev, 20(6), 1170-1175. doi:10.3758/s13423-013-0459-3 evans, k. k., tambouret, r. h., evered, a., wilbur, d. c., & wolfe, j. m. (2011). prevalence of abnormalities influences cytologists' error rates in screening for cervical cancer. arch pathol lab med, 135(12), 1557-1560. doi:10.5858/arpa.2010-0739-oa fallon, m. a., wilbur, d. c., & prasad, m. (2010). ovarian frozen section diagnosis: use of whole-slide imaging shows excellent correlation between virtual slide and original interpretations in a large series of cases. arch pathol lab med, 134(7), 1020-1023. doi:10.1043/2009-0320-oa.1 fox, s. e., law, c. c., & faulkner-jones, b. e. (2017). quantitative gaze assessment of a dual ”side by side” viewer versus a single whole slide image viewer for pathology education. manuscript submitted for publication. gauthier, i., & nelson, c. a. (2001). the development of face expertise. curr opin neurobiol, 11(2), 219224. henneman, e.a., cunningham h., fisher d.l., et al. (2014) eye tracking as a debriefing mechanism in the simulated setting improves patient safety practices. dimens crit care nurs, 33:129–135. doi: 10.1097/dcc.0000000000000041. holmqvist, k. (2011). eye tracking : a comprehensive guide to methods and measures. oxford ; new york: oxford university press. huette, s., winter, b., matlock, t., & spivey, m. (2012). processing motion implied in language: eyemovement differences during aspect comprehension. cogn process, 13 suppl 1, s193-197. doi:10.1007/s10339-012-0476-6 jarodzka, h., balslev, t., holmqvist, k., nyström, m., scheiter, k., gerjets, p., & eika, b. (2012). conveying clinical reasoning based on visual observation via eye-movement modelling examples. instructional science, 40(5), 813-827. doi: 10.1007/s11251-012-9218-5 jeong, w. k., schneider, j., turney, s. g., faulkner-jones, b. e., meyer, d., westermann, r., . . . pfister, h. (2010). interactive histology of large-scale biomedical image stacks. ieee trans vis comput graph, 16(6), 1386-1395. doi:10.1109/tvcg.2010.168 johnson, j. s., liu, l., thomas, g., & spencer, j. p. (2007). calibration algorithm for eyetracking with unrestricted head movement. behav res methods, 39(1), 123-132. doi: 10.3758/bf03192850 krupinski, e., chao, j., hofmann-wellenhof, r., morrison, l., & curiel-lewandrowski. (2014). understanding visual search patterns of dermatologists assessing pigmented skin lesions before and after online training. j digit imaging, 27, 779-785. doi:10.1007/s10278-014-9712-1 krupinski, e. a., graham, a. r., & weinstein, r. s. (2013). characterizing the development of visual search expertise in pathology residents viewing whole slide images. hum pathol, 44(3), 357-364. doi:10.1016/j.humpath.2012.05.024 http://dx.doi.org/10.1037/h0076100 fox et al | f l r 53 krupinski, e. a., tillack, a. a., richter, l., henderson, j. t., bhattacharyya, a. k., scott, k. m., . . . weinstein, r. s. (2006). eye-movement study and human performance using telepathology virtual slides: implications for medical education and differences with experience. hum pathol, 37(12), 15431556. doi:10.1016/j.humpath.2006.08.024 kundel, h. l., nodine, c. f., krupinski, e. a., & mello-thoms, c. (2008). using gaze-tracking data and mixture distribution analysis to support a holistic model for the detection of cancers on mammograms. acad radiol, 15(7), 881-886. doi:10.1016/j.acra.2008.01.023 laurent, p. a., hall, m. g., anderson, b. a., & yantis, s. (2015). valuable orientations capture attention. vis cogn, 23(1-2), 133-146. doi:10.1080/13506285.2014.965242 law, b., atkins, m. s., lomax, a. j., & wilson, j. g. (2003). eye trackers in a virtual laparoscopic training environment. stud health technol inform, 94, 184-186. mallett, s., phillips, p., fanshawe, t. r., helbren, e., boone, d., gale, a., . . . halligan, s. (2014). tracking eye gaze during interpretation of endoluminal three-dimensional ct colonography: visual perception of experienced and inexperienced readers. radiology, 273(3), 783-792. doi:10.1148/radiol.14132896 marquard j.l., henneman p.l., he z., jo j., fisher d.l., henneman e.a. (2011). nurses' behaviors and visual scanning patterns may reduce patient identification errors. j exp psychol appl, 17:247–256. doi: http://dx.doi.org/10.1037/a0025261 merchant, j., morrissette, r., & porterfield, j. l. (1974). remote measurement of eye direction allowing subject motion over one cubic foot of space. ieee trans biomed eng, 21(4), 309-317. doi:10.1109/tbme.1974.324318 pascalis, o., scott, l. s., kelly, d. j., shannon, r. w., nicholson, e., coleman, m., & nelson, c. a. (2005). plasticity of face processing in infancy. proc natl acad sci u s a, 102(14), 5297-5300. doi:10.1073/pnas.0406627102 phillips, p., boone, d., mallett, s., taylor, s. a., altman, d. g., manning, d., . . . halligan, s. (2013). method for tracking eye gaze during interpretation of endoluminal 3d ct colonography: technical description and proposed metrics for analysis. radiology, 267(3), 924-931. doi:10.1148/radiol.12120062 rayner, k. (1978). eye movements in reading and information processing. psychol bull, 85(3), 618-660. rossion, b., collins, d., goffaux, v., & curran, t. (2007). long-term expertise with artificial objects increases visual competition with early face categorization processes. j cogn neurosci, 19(3), 543-555. doi:10.1162/jocn.2007.19.3.543 sali, a. w., anderson, b. a., & yantis, s. (2014). the role of reward prediction in the control of attention. j exp psychol hum percept perform, 40(4), 1654-1664. doi:10.1037/a0037267 stewart, j., 3rd, miyazaki, k., bevans-wilkins, k., ye, c., kurtycz, d. f., & selvaggi, s. m. (2007). virtual microscopy for cytology proficiency testing: are we there yet? cancer, 111(4), 203-209. doi:10.1002/cncr.22766 tiersma, e. s., peters, a. a., mooij, h. a., & fleuren, g. j. (2003). visualising scanning patterns of pathologists in the grading of cervical intraepithelial neoplasia. j clin pathol, 56(9), 677-680. doi: http://dx.doi.org/10.1136/jcp.56.9.677 tomizawa y., aoki h., suzuki s., matayoshi t., yozu r. (2012). eye-tracking analysis of skilled performance in clinical extracorporeal circulation. j artif organs, 15:146–157. doi: 10.1007/s10047012-0630-z tourassi, g., voisin, s., paquit, v., & krupinski, e. (2013). investigating the link between radiologists' gaze, diagnostic decision, and image content. j am med inform assoc, 20(6), 1067-1075. doi:10.1136/amiajnl-2012-001503 tourassi, g. d., mazurowski, m. a., harrawood, b. p., & krupinski, e. a. (2010). exploring the potential of context-sensitive cade in screening mammography. med phys, 37(11), 5728-5736. doi:10.1118/1.3501882 vanderwert, r. e., westerlund, a., montoya, l., mccormick, s. a., miguel, h. o., & nelson, c. a. (2015). looking to the eyes influences the processing of emotion on face-sensitive event-related potentials in 7month-old infants. dev neurobiol, 75(10), 1154-1163. doi:10.1002/dneu.22204 fox et al | f l r 54 wagner, j. b., hirsch, s. b., vogel-farley, v. k., redcay, e., & nelson, c. a. (2013). eye-tracking, autonomic, and electrophysiological correlates of emotional face processing in adolescents with autism spectrum disorder. j autism dev disord, 43(1), 188-199. doi:10.1007/s10803-012-1565-1 ward, a., rosen, d. m., law, c. c., rosen, s., & faulkner-jones, b. e. (2015). oxalate nephropathy: a three-dimensional view. kidney int, 88(4), 919. doi:10.1038/ki.2015.31 wen, g., aizenman, a., drew, t., wolfe, j. m., haygood, t. m., & markey, m. k. (2016). computational assessment of visual search strategies in volumetric medical images. j med imaging (bellingham), 3(1), 015501. doi:10.1117/1.jmi.3.1.015501 wolfe, j. m. (1994). guided search 2.0 a revised model of visual search. psychon bull rev, 1(2), 202-238. doi:10.3758/bf03200774 wolfe, j. m. (1995). the pertinence of research on visual search to radiologic practice. acad radiol, 2(1), 74-78. wolfe, j. m. (2012a). saved by a log: how do humans perform hybrid visual and memory search? psychol sci, 23(7), 698-703. doi:10.1177/0956797612443968 wolfe, j. m. (2012b). when do i quit? the search termination problem in visual search. nebr symp motiv, 59, 183-208. doi: 10.1007/978-1-4614-4794-8_8 wolfe, j. m., evans, k. k., drew, t., aizenman, a., & josephs, e. (2015). how do radiologists use the human search engine? radiat prot dosimetry. doi:10.1093/rpd/ncv501 wolfe, j. m., & horowitz, t. s. (2004). what attributes guide the deployment of visual attention and how do they do it? nat rev neurosci, 5(6), 495-501. doi:10.1038/nrn1411 wolfe, j. m., horowitz, t. s., kenner, n., hyle, m., & vasan, n. (2004). how fast can you change your mind? the speed of top-down guidance in visual search. vision res, 44(12), 1411-1426. doi:10.1016/j.visres.2003.11.024 wolfe, j. m., horowitz, t. s., van wert, m. j., kenner, n. m., place, s. s., & kibbi, n. (2007). low target prevalence is a stubborn source of errors in visual search tasks. j exp psychol gen, 136(4), 623-638. doi:10.1037/0096-3445.136.4.623 wolfe, j. m., & van wert, m. j. (2010). varying target prevalence reveals two dissociable decision criteria in visual search. curr biol, 20(2), 121-124. doi:10.1016/j.cub.2009.11.066 frontline learning research vol.4 no. 4 special issue (2016) 1 6 issn 2295-3159 doi: http://dx.doi.org/10.14786/flr.v4i4.320 expanding conceptualizations for the study of learning antti rajalaa, giuseppe ritellaa, kristiina kumpulainena, & louise wilkinsonb auniversity of helsinki, finland bsyracuse university, usa 1. the background and need for this special issue this special issue is dedicated to expanding conceptualizations for the study of learning in contemporary education. ongoing social changes in the private, public, and economic spheres create new and various demands for learning and education. as the articles of this special issue propose, reasoning, critical thinking, imagination, and managing emotions in dealing with controversial issues have become increasingly important learning requirements in the pursuit of interests toward learning across diverse settings. also, the way in which people learn to take part in such practices in contemporary contexts, both in formal education and everyday life, are shifting. for instance, novel kinds of digital tools create continuously evolving spaces for learning that transform social interactions and learning practices across contexts and time (ritella, ligorio, & hakkarainen, 2016). overall, ongoing changes in society, and the learning requirements they entail, challenge research communities to reconsider how to understand and advance learning in diverse settings. it is also increasingly recognised that there is a need to apply and further develop conceptual and methodological frameworks that are able to account for the complexity of learning in contemporary societal conditions (kumpulainen & erstad, 2016). the five articles included here address several important and under-researched topics in contemporary learning and education. each article also proposes and elaborates on potential conceptual frameworks for expanding conceptualizations for the study of learning in the 21st century. before introducing the articles, we will describe the impetus for the publication of the special issue. after that we describe each of the contributions including their specific research topics and conceptual frameworks. we then outline and discuss some cross-cutting themes emerging from our reading of the articles and conclude by pointing out some key arguments made by the commentators to further the ongoing dialogue and research in the field. http://dx.doi.org/10.14786/flr.v4i4.320 introduction 2 | f l r 2. the impetus for publishing the special issue the impetus for this publication was a symposium, “evolving theoretical frameworks for studying collaboration in diverse 21st century learning contexts,” presented at the meeting of the earli special interest groups 10, 21, and 25 in padova, italy on august 27‒30, 2014. the conference, “open spaces for interaction and learning diversities,” was sponsored by earli and the university of padova. the original call for proposals emphasized the need to address the challenges that global movements and cross-cultural communication continue to pose for learning and education. the meeting was the joint effort of three special interest groups: sig 10, which represents researchers who study aspects of the field of social interaction in learning and instruction; sig 21 with an emphasis on learning and teaching in culturally diverse settings; and sig 25 with a focus on educational theory. in particular, the business meeting of sig 25 (educational theory) encouraged us to publish a collection that would explicitly address the issue of reconceptualizing learning. subsequent to the earli symposium, we decided to propose a special issue of frontline learning research that would capture and extend the discussion. additional articles were commissioned and this special issue represents the culmination of this process. as guest editors, it was our aspiration from the beginning to both extend and catalyze further discussions about the future of research on learning and education in changing societies. 3. descriptions of the individual articles and commentaries the special issue consists of five articles and two commentaries. the first article by tsafrir goldberg and baruch schwarz asserts that scant attention has been paid to the role of emotions in students’ reasoning. contemporary society is also characterized by tensions and social conflicts related to cultural and religious diversity, which carry interpretations of historical facts that are heavily loaded with emotions. goldberg and schwarz address these tensions in history education by introducing a framework for studying the role of emotions and identity for deliberative argumentation. the article provides an overall conceptual framework for the centrality of emotion in cognitive processes such as reasoning and argumentation, and it reports the results of an analysis of students’ responses to three different approaches to teaching history: (1) a conventional authoritative approach involving patriotic apologetic teaching, (2) an empathetic dual-narrative approach involving nonjudgmental listening to collective narratives and identifying with emotions and values, and (3) a critical disciplinary inquiry approach involving the critical analysis and synthesis of conflicting sources. the article reports a study of peer discussions between jewish and arab students regarding the 1948 “war of independence.” the approaches resulted in different processes and outcomes where learners invoked identity, emotion, and perspective-taking; students became aware of their own biases, which limited, or eliminated in some cases, productive discussions. goldberg and schwarz propose that teachers consider ways of engaging students’ positive and negative emotions and construct authentic learning activities; in this way, students may harness their emotions and engage in critical classroom discussions of emotionally-charged topics in content areas such as history. the article by jaakko hilppö, antti rajala, tania zittoun, kristiina kumpulainen, and lasse lipponen addresses the role of imagination in learning, an aspect of the cognitive process that has received scant attention in prior research. the article proposes an overall conceptual framework of the centrality of imagination in learning, and it reports the findings from a case analysis of finnish primary-school students in a science classroom. imagination, from their point of view, is characterized by a partial and temporary separation from the immediate, proximal experience of the social and material world into distal experiences, which are ultimately connected back to the present experience. the analysis of the single case illustrates several aspects of imagination as identified by their conceptual framework. these include encouraging students to break the constraints of time and space that are often superimposed on the teaching process in introduction 3 | f l r classrooms both in europe and in the u.s. their discussion illustrates how students make sense of scientific phenomena and how that sensemaking expands both in time and in space. throughout this process, the students’ thinking becomes more refined and differentiated through what they reference as “loops of imagination,” referring to the back-and-forth movement between proximal and distal experiences. finally, they acknowledge that while the analysis of the single case is instructive, further research needs to be conducted to determine how the conceptual framework applies to varied instances of imagination. the article by william penuel, daniela digiacomo, katie van horne, and ben kirshner argues that although equity-oriented efforts of expanding young people’s access and learning opportunities in science, technology, engineering, and math (stem) are laudable, to date, too little attention has been paid to gaining a more nuanced understanding what it entails to develop and sustain an interest-driven stem-learning pathway across settings and over time among diverse youth. to overcome this limitation, the authors propose a social-practice theory as a prominent theoretical lens to guide the study of learning pathways across varied contexts and over time. they contend that using social-practice theory in the analysis of youth learning is potentially transformative as it unpacks potential leverage points for transforming systems to enable broader participation in stem. in order to justify their argument, the authors apply social-practice theory to interpret the learning pathway of one adolescent, jerome, who they followed as part of a longitudinal study of interest-related stem learning. in their analysis of jerome, the authors demonstrate how he pursued diverse concerns and became aware of new possibilities for action as he moved across different settings of practice and learned to adjust his contributions to the flow of ongoing activity to fit demands and structures of local institutions. the analysis also powerfully illustrates how institutional structures of practice framed the choices jerome made about his participation, learning, and becoming in relation to stem. the article by crina damsa and alfredo jornet shows how an ecological perspective can be used to redefine key concepts of research on learning in higher education. in the article, learning is conceptualized as an achievement of whole ecosystems and involves mutually transformative transactions between people and their sociomaterial environments. although the article builds on existing sociocultural, situative, and sociomaterial approaches to learning, it contributes new insights into conceptualizing learning by discussing the implications of the ecological premises underlying these approaches. by analyzing a video-based case study of an undergraduate course in web design and development, the article shows how ‟co-construction” in a collaborative student group can be characterized as an unfolding field of action in which intellectual agency intertwines with affective and performative relations. they also demonstrate how ‟knowledge resources and materials” form an ecology that becomes inseparably entangled with the students’ activities. finally, damsa and jornet discuss the limitations of the notions of transfer and boundary crossing in accounting for learning in new higher-education contexts where, they argue, students’ activity takes place within a ‟trans-contextual” ecology of learning in which no clear boundary exists between university and professional settings. the authors suggest practical implications for reorganizing higher education to address transformative potentials emerging in the students’ activities. lastly, giuseppe ritella, beatrice ligorio, and kai hakkarainen introduce chronotope as a conceptual tool to examine if and how the organization of space and time might affect learning processes. some previous research, the authors argue, demonstrates that spatial and temporal relations are important for learning, but that our knowledge on this topic is still limited. a problem concerning the examination of space-time is that – as mentioned by wegerif in his commentary in this issue – its dynamics are often implicit, going beyond verbally articulated understanding, and thus, it is challenging to grasp its effects on learning. the authors’ claim is that the emergence of chronotope as a scientific concept might help us to verbalize – and scientifically investigate – what is usually implicit, allowing reflection and dialogue on a dimension of learning that is often taken for granted but that seems to exert a silent influence on how we learn. in this sense, the aim of the article is not to conceptualize education exclusively in terms of spacetime, but to suggest a theoretically founded way to examine how the variation of spatial and temporal relations might affect learning processes. three main features of chronotope are presented and discussed to explore its value as a scientific concept: the examination of the potential interdependency between space and introduction 4 | f l r time; the focus on the social negotiation of space-time; and the coordinated examination of material and discursive processes involved in the negotiation of space-time frames. using some examples from their own empirical investigations and from the existing literature, the authors discuss how we can gain further insights about learning processes by examining the space-time relations of learning by using chronotope as analytic lens. 4. cross-cutting themes and emphases a closer reading of the conceptualizations presented in the articles revealed three main themes, which we believe are central for the study of learning and educational practice in changing societies. below, we briefly introduce these themes and discuss how the articles addressed them. 4.1 theme 1: expanding time-space contexts of learning first, some of the articles in this collection challenge the static framing of time and space that often underpins the research of learning. these studies make visible the implicit spatial and temporal infrastructure on which learning and education rely and which is, in turn, being shaped by the processes of learning and education. common to these studies is that they conceptualize space-time contexts as dynamically intertwined with the activities of the people who are being studied. in this respect, ritella et al. introduce the concept of chronotope to create a nuanced theoretical account of how the digitalization of education reshapes the time and space relations of learning activities. they argue that the new technological innovations and ongoing educational reforms transform the spatial and temporal organization of learning in ways that have profound implications for the study of learning. for instance, they use the example of the ‟flipped learning” approach to exemplify how a new pedagogical approach combined with novel digital technology transforms the learning process by changing where and when school learning takes place. penuel et al. use the concept of learning pathway to account for learning as movement across settings of practice and over time while people pursue their interests. hilppö et al. show how primary school students use their imaginations to explore temporally and spatially distant phenomena relevant to the science topic they are studying. 4.2 theme 2: agency-driven learning some of the contributions in this collection develop concepts that help to examine student agency in learning. agency is an emerging research topic in the learning sciences that accounts for acting upon and transforming activities and life circumstances (rajala, kumpulainen, & martin, 2016). the focus on agency permits approaching the culture-learning interfaces in terms of both enculturation and transformation and balances the over-emphasis on the collective and reproductive dimensions of learning (kumpulainen & renshaw, 2007). damsa and jornet seek to redefine agency in knowledge creation in a way that avoids the dualism of the agent and the material world. in the ecological perspective, people are construed both as active in transforming their circumstances and as passive in being subject to the performative and affective relations in which they engage. ritella et al. echo the idea of agency as transformation by arguing that the time-space relations of an activity are amenable to change through the actions of the participants. they illustrate their argument by discussing a study of student teachers who managed the time-space contexts of their activities by arranging their bodies in space, searching for resources in the environment, and exploring physical and virtual space. introduction 5 | f l r some of the articles address agency indirectly. the article by penuel et al. conceptualizes interestdriven science learning as involving agentic learners who direct their learning pathways across a range of settings. the article also discusses how structures of practice constrain agency and limit access in some settings. the article by hilppö et al. examines imagination, which can be considered a prerequisite for agency that allows people to distance themselves from the immediate constraints of action and imagine alternatives to the present circumstances (emirbayer & mische, 1998). 4.3 theme 3: new directions for the study of non-cognitive dimensions of learning the dominant discourse on learning still concerns standardized measurement of cognitive learning outcomes. the articles of this collection challenge this discourse and make room for imagination, interest, and emotion as well as discussion of values in learning and education. imagination is an under-researched topic that merits further attention. building on the pioneering work in cultural psychology of zittoun and gillespie (2016), hilppö et al. develop a framework for researching processes of imagination in science learning. this framework accounts for the back-and-forth movement of imagination between proximal and distal experiences in classroom discourse. penuel et al. highlight interest as a driving force in learning across contexts. in their conceptualization, interest is not seen as confined within an individual but as emergent in practice and shaped by the available resources and opportunities. goldberg and schwarz argue that emotions are often considered an obstacle for deliberation and reasoning in classrooms. however, they develop a framework for supporting engagement with emotions to promote critical and productive engagement with history topics. their article also makes an important contribution in starting to theorize how controversial and politically charged topics can be addressed in classroom situations. thus, their article contributes to a recent discussion in the learning sciences that pays more attention than before to the socio-political contexts of learning (politics of learning writing collective, 2017). this discussion can be considered a partial response to gert biesta’s (2010) critique that the dominant focus on learning in educational research has obscured the the value dimension and promoted a view of education as a technical matter of effectiveness and efficiency. 5. conclusion while addressing important research topics in contemporary education, the articles of this special issue introduce several potential frameworks for expanding conceptualizations for the study of learning in the 21st century. the research foci and the conceptual framings that they discuss also point out the need for research communities in learning and education not only to engage in empirical research in novel settings, but also to reflect upon the theoretical frameworks that explicitly or implicitly inform the research on learning. it is important to unpack the often-implicit assumptions, values, and educational purposes that underlie the theoretical frameworks and to consider the implications for what it means to learn in the 21st century. in his commentary, rupert wegerif applauds the authors of this collection for making their theoretical assumptions visible and bringing them under scrutiny. at the same time, he points out quite rightly that there are always theoretical assumptions involved in research on learning that foreground some relevant educational phenomena and make others more difficult to consider. he cautions that new conceptualizations of learning are not valuable in themselves, and he challenges us to think what is gained if we look at things using the conceptual frameworks proposed in the articles and why we should invest our energy specifically in these conceptual frameworks and not others. introduction 6 | f l r the other commentators, jessica mckeown and cindy hmelo-silver, suggest that the proposed conceptual frameworks can be useful in inspiring new designs for promoting emergent learning that is often valued in contemporary educational settings. as a way forward, both of the commentaries suggest putting the conceptual frameworks advanced in the articles to rigorous empirical tests, for example, through experimental studies and design research. as guest editors, we agree that the worth of the conceptual and theoretical frameworks presented and discussed will be determined by their potential to inform further empirical and interventionist research. yet, we also underscore that there is no straightforward way to determine “what works” in educational research. learning is a normative concept that involves both analytical reasoning concerning the theoretical assumptions that lead educational research and value judgements concerning the purposes of education at large (biesta, 2010). thus, we urge further research and political discussion on how learning is conceptualized in the research community and in the larger society, in attempts to document and assess the value of educational interventions and programs in scientifically sound manners. references biesta, g. (2010). good education in an age of measurement: ethics, politics, democracy. boulder/london: paradigm. publishers. emirbayer, m., & mische, a. (1998). what is agency? american journal of sociology, 103(4), 962–1023. kumpulainen, k., & erstad, o. (2016). (re)searching learning across contexts: conceptual, methodological and empirical explorations. international journal of educational research. kumpulainen, k., & renshaw, p. (2007). cultures of learning. international journal of educational research, 46(3), 109–115. politics of learning writing collective. (2017). the learning sciences in a new era of us nationalism. cognition & instruction, 35(2). rajala, a., martin, j., & kumpulainen, k. (2016). agency and learning: researching agency in educational interactions. learning, culture and social interaction, (10), 1–3. ritella, g., ligorio, m. b., & hakkarainen, k. (2016). the role of context in a collaborative problem-solving task during professional development. technology, pedagogy and education, 25(3), 395–412. egloff et al publication frontline learning research vol.7 no. 1 (2019) 1 22 issn 2295-3159 students´ reading ability moderates the effects of teachers´ beliefs on students´ reading progress frank egloffa, natalie försteraelmar souvigniera auniversity of münster, germany article received 13 october 2017 / revised revised 6 september 2018/ accepted 25 october / available online 16 january abstract teachers’ beliefs about teaching have been found to affect students’ learning growth. the aim of this study was to investigate effects of teachers’ constructivist and direct-transmissive beliefs on learners’ reading progress and whether these effects are influenced by students’ ability. we measured constructivist and direct-transmissive beliefs of 29 teachers and the progress in reading fluency and reading comprehension of their students (n = 568) at eight points of measurement over one school year. results of three-level latent growth curve modeling revealed that only teachers’ global, but not reading specific constructivist beliefs, were generally positively related to learners’ progress in reading fluency. beliefs about teaching had no general effect on growth in reading comprehension, but the relation between constructivist beliefs and students’ progress in reading comprehension was affected by students’ prior skills. teachers with stronger constructivist beliefs effected higher learning growth for high ability compared to low ability learners within their classrooms. no effects were found for direct-transmissive beliefs. this study adds a more differentiated view to findings concerning the effects of teacher beliefs by showing that effects vary depending on the skill under study (fluency vs. comprehension), and that effects of teacher beliefs may depend on students’ ability. keywords: teacher beliefs, beliefs about teaching, reading comprehension, reading fluency info corresponding author email egloff@uni-muenster.de doi: 10.14786/flr.v7i1.336 1. introduction teachers’ beliefs about teaching have been found to influence teachers’ behaviour and thereby affect student learning (pajares, 1992; peterson, fennema, carpenter, & loef, 1989). two theoretically and empirically distinguishable beliefs about teaching are constructivist and direct-transmissive beliefs. a teacher with high constructivist beliefs for example is convinced that students play an active part in their learning and that they should and will develop their own problem-solving strategies. in contrast, a teacher who holds a direct-transmissive view about teaching believes that a teacher should guide students´ learning process. most of the research on teacher beliefs has shown that teachers’ high constructivist and low direct-transmissive beliefs are positively related to higher learning progress (dubberke, kunter, mcelvany, brunner, & baumert, 2008; peterson et al., 1989; souvignier & mokhlesgerami, 2005; staub & stern, 2002). nevertheless, in a sample of particularly low achieving students, contrary results were found (behrmann & souvignier, 2013). thus, similar to the concept of child x instruction interactions (e.g. connor, morrison, & petrella, 2004), which assumes that effects of instruction depend on the fit to students’ abilities, also effects of teachers’ beliefs might vary depending on learners’ initial skills. therefore, the goals of our study were to contribute to the research on effects of teachers’ beliefs by studying—as yet under-investigated¬—students in primary school in the domain of reading and—more importantly—to investigate whether these effects are influenced by prior abilities of individual students or the respective classroom. 1.1 teachers’ beliefs about teaching beliefs are subjective evaluations on whether a specific proposition is true (e.g., pajares, 1992; richardson, 2003). they can be distinguished from knowledge, which is rather based on logical argumentation, fact and thus, expert consensus. in contrast, beliefs are non-consensual because they are rather built on personal emotional experiences (behrmann & souvignier, 2013; nespor, 1987; pajares, 1992; richardson, 2003). teachers’ beliefs are assumed to be very important as they affect what happens in the classroom (e.g., buehl & beck, 2015; dubberke et al., 2008; kagan, 1992; peterson et al., 1989; staub & stern, 2002). they refer to issues that are relevant to teachers’ profession such as their own teaching effectiveness, the nature of knowledge, and how students should be taught (e.g., pajares, 1992). hence, beliefs about teaching, in general, cover all aspects of the spectrum on quality of education (fives, lacatena, & gerard, 2015). within the category of beliefs about teaching, constructivist and direct-transmissive beliefs are most apparent. case studies using analyses of teacher talk during shared planning time (gill & hoffman, 2009) and other qualitative methods mostly revealed constructivist and direct-transmissive beliefs among teachers (see fives et al., 2015; kleickmann, 2007). furthermore, it was demonstrated that operationalisations of constructivist and direct-transmissive beliefs predicted teaching behaviour as well as students’ learning outcomes (e.g., dubberke et al., 2008; staub & stern, 2002). constructivist views on learning and teaching can be related to cognitive constructivist learning theories (see savery & duffy, 1995). according to this perspective, teachers believe that learners play an active role in the process of studying. the underlying idea of learning is that students actively integrate new information into their existing knowledge. to foster an active and autonomous studying process, the teachers’ task is to provide learners with meaningful learning environments (staub & stern, 2002). in contrast, direct-transmissive views on learning and teaching are related to behavioural-associationist learning theories (see resnick & hall, 1998). according to this view, teachers explicitly instruct and guide the students through new contents. thereby, teachers pre-structure the topics and monitor learners’ progress, which leads to a more passive role of the students in the process of knowledge building (staub & stern, 2002). 1.2 effects of teachers’ beliefs about teaching on students’ progress teachers’ beliefs affect perception, information processing, judgement, decision making, and the way of teaching (e.g., buehl & beck, 2015; dubberke et al., 2008; kagan, 1992; pajares, 1992; peterson et al., 1989). hattie (2012) concluded that next to commitment, “teachers’ beliefs (…) are the greatest influence on student achievement over which we have some control” (p. 22). to date, several studies have examined the influence of teachers’ constructivist or direct-transmissive beliefs, showing either positive or negative effects on students’ progress. nevertheless, the number of such studies is limited. peterson et al. (1989) investigated a sample of 39 first-grade math teachers and found constructivist views to be positively related to learners’ mathematical word problem solving. staub and stern (2002) surveyed 27 teachers and similarly demonstrated that constructivist beliefs were positively linked to students’ growth in mathematical word problem solving in second and third grade. studying effects of teachers’ transmissive beliefs in a sample of 155 ninth and tenth-grade school teachers, dubberke et al. (2008) found a negative relation between strong transmissive beliefs and student achievement in mathematics. in reading, only two studies have been conducted, yet. in line with studies in mathematics, data from souvignier and mokhlesgerami (2005) revealed a positive relation between constructivist beliefs and learning growth in reading strategy knowledge for fifth and sixth-grade students from the highest school track. behrmann and souvignier (2013) studied a sample of particularly low performing sixth and seventh graders. in contrast to prior research, they found advantages of direct-transmissive beliefs concerning learners’ declarative and procedural reading strategy knowledge but not their reading fluency. summing up the findings in the literature, positive relations between teachers’ constructivist beliefs and students’ learning growth have been found in mathematics and reading within groups of average to high performing students, whereas direct-transmissive beliefs seem to be negatively related to student learning. this might be explained by higher cognitive activation and motivational advantages of teaching methods that are in line with constructivist beliefs (e.g., dubberke et al., 2008; savery & duffy, 1995; staub & stern, 2002). nevertheless, the opposite pattern of findings in a sample of particularly low achieving students (behrmann & souvignier, 2013) raises the question of whether these effects depend on students’ initial ability. following the concept of child x instruction interactions, connor et al. (2004) found that high achieving students benefitted from more self-regulated phases of reading, while lower performing students needed more pre-structured teaching. given that beliefs affect teachers’ decision-making process and their teaching, the same interactional effects as those between child x instructions might exist for teacher beliefs. thus, a teacher with high constructivist beliefs who provides much child-managed instruction meets the needs of high but not those of low achieving students. in contrast, a teacher who uses many phases of teacher-managed instruction due to his or her high direct transmissive beliefs, might affect higher growth for the low but not the high achieving students in the classroom. the short review of studies that analysed the relation between teachers’ beliefs and learning progress reveals that the majority of studies have been conducted in the domain of mathematics. only two studies (behrmann & souvignier, 2013; souvignier & mokhlesgerami, 2005) investigated effects of teachers’ beliefs on growth in students’ reading related skills. in these studies, reading achievement was assessed specifically according to reading fluency and knowledge of reading strategies. effects on reading comprehension, however, have not been studied yet. therefore, the goal of our study was to broaden the empirical basis concerning effects of teachers’ beliefs in the domain of reading, differentiating between the two key constructs of reading fluency and reading comprehension (national institute of child health and human development [nichd], 2000). 1.3 reading fluency and reading comprehension reading fluency consists of word recognition accuracy and reading speed (samuels, 1979). it is based on the automation of word recognition (e.g., nichd, 2000) and therefore relies on large amounts of student-driven decoding practice (rasinski, reutzel, chard, & linan-thompson, 2011). students’ reading fluency increases from grade to grade with a decreasing growth rate over time (parrila, aunola, leskinen, nurmi, & kirby, 2005; tilstra, mcmaster, van den broek, & kendeou, 2009). reading comprehension in contrast, is a process of constructing a subjective representation of textual information (kintsch, 1998). concerning this skill, two sub-processes can be distinguished. one of them results in a local semantic representation of information explicitly inherent to the text. to build this so-called textbase, the reader has to infer meaning from connecting information of words and sentences of smaller units of a text. the other sub-process connects the semantic information of the textbase with prior knowledge to build a meaningful macrostructure representation of the text, which is called situation model. this skill is based on elaboration, organization and metacognitive processes and thus, cannot be automatised. cromley and azevedo (2007) showed that background knowledge about the content of a text, strategies and vocabulary are important predictors to knowledge-based inferences as well as to reading comprehension itself. longitudinal studies revealed mixed results about growth trajectories. studies of parrila et al. (2005) as well as of farnia and geva (2013) showed that reading comprehension growth decreases over time. nevertheless, studies of tilstra et al. (2009) as well as of nation, cocksey, taylor, and bishop (2010) revealed that growth rates do not decrease over time, but rather follow a linear growth pattern. from the delineation of these two different reading skills, it becomes obvious that they consist of entirely different cognitive processes. while reading fluency is based on automatised word recognition, reading comprehension is based on prior knowledge, metacognitive processes and consciously made inferences. consequently, successful instructional designs to support these skills vary concerning the amount of teacher and student guided activities. reading interventions designed to support reading fluency like paired reading (topping, 1987) and repeated reading (samuels, 1979) encourage students to play an active role in the learning process. these methods especially rely on extended practice so that there is only little need for direct instruction. reading programs to foster reading comprehension (e.g., brown & pressley, 1994; souvignier & mokhlesgerami, 2006; paris, cross, & lipson, 1984), however, demand an active role by the teacher who is supposed to directly instruct and model cognitive and meta-cognitive strategies for reading by a relatively large part. thus, teacher guided instruction may be necessary especially for lower performing students to successfully develop their reading comprehension. 1.4 assessment of teachers’ beliefs about teaching from a learning theory perspective, constructivist and direct-transmissive beliefs are contradictory (behrmann & souvignier, 2013; oecd, 2009; staub & stern, 2002). thus, the assessment of beliefs has often been conceptualized on a constructivist-transmissive continuum, using only one single scale (e.g., peterson et al., 1989; souvignier & mokhlesgerami, 2005; staub & stern, 2002). dubberke et al. (2008), however, used one scale to assess direct-transmissive beliefs only. likewise, behrmann and souvignier (2013) used separate constructivist and direct-transmissive scales in addition to a constructivist-transmissive continuum scale. this alternative approach follows the argumentation that teachers may hold even potentially contradictive perspectives on effective teaching (see fives et al., 2015). it was found that teachers can endorse constructivist and direct-transmissive views at the same time (organisation for economic co-operation and development [oecd], 2009; fives et al., 2015; snider & roehl, 2007). teachers’ beliefs about teaching may be inconsistent because of a vast range of different teaching situations that teachers have to face and to consider (behrmann & souvignier, 2013; fives et al., 2015). this is supported by means of factor analytical results, which show that constructivist as well as direct-transmissive beliefs each build their own factors (e.g., bunting, 1985; woolley, benjamin, & woolley, 2004). thus, conceptualizing and measuring constructivist and direct-transmissive views on teaching with separate scales may be more appropriate (behrmann & souvignier, 2013; buehl & beck, 2015; woolley et al., 2004). nevertheless, strong negative correlations between measures of constructivist and direct-transmissive scales indicate that teachers generally tend to favour one of the two orientations (see behrmann & souvignier, 2013). another important aspect of beliefs about teaching is their specificity with respect to a certain content. beliefs can be conceptualized as content specific (i.e. related to reading instruction) as well as content general. content specific beliefs about teaching may be especially important because they may more precisely apply to content specific teaching situations (peterson et al. 1989; staub & stern, 2002). reading specific beliefs may thus, especially affect students’ reading competence growth. nevertheless, global beliefs about teaching may also have an important impact on students’ learning of reading skills because they may affect teaching on a more general level in most teaching situations. given that reading is not limited to a specific subject like maths, it seems reasonable to assess teachers’ beliefs both in a content specific and in a content general way. 1.5 assessment of student progress to analyse the effects of teachers’ beliefs on learners’ progress, usually data from longitudinal designs with preand posttests on student achievement have been used. difference scores from two measures as an indicator for learning progress, however, have been criticized with respect to limited reliability (willett, 1989). assessing change with multiple points of measurement creates advantages over pre-post measures, because it boosts the reliability of the growth rate estimates (e.g., singer & willett, 2003; willett, 1989). willett (1989) demonstrated that every additional point of assessment helps to deflate standard errors and concluded that “with sufficient waves added, the influence of fallible measurement rapidly dwindles to zero” (p. 598). furthermore, speer and greenbaum (1995) demonstrated that growth modeling based on multiple points of assessment is more sensitive to change than methods based on pre-post measures. in this study, we enhance the reliability of the assessment of growth by modeling learning progress across eight points of measurement over one school year. 1.6 research questions our study addressed three research questions: first, we were interested in general effects of teachers’ beliefs on students’ progress in reading. consistent with previous studies, we expected positive effects from constructivist beliefs and negative effects for direct-transmissive views on reading fluency and comprehension. second, we wanted to investigate if students’ initial achievement moderates the effects of teachers’ beliefs on students’ reading progress. given the different findings in reading with positive effects of either constructivist or direct-transmissive beliefs for high and low achieving students, respectively, we anticipated that constructivist beliefs might be supportive for students with higher reading skills, whereas direct-transmissive beliefs might be more suitable for students with lower reading skills. third, given that moderating effects might not only become apparent at the individual level but also in the entire classroom, we also studied whether effects of teachers’ beliefs on classrooms’ reading growth were moderated by the average prior ability of the classroom. in concordance with the second hypothesis, we expected that constructivist beliefs are supportive for classrooms with higher initial reading ability and direct-transmissive beliefs to be more helpful for classrooms with lower initial reading ability. regarding each of the hypotheses, we expected the same effects for growth in reading fluency and reading comprehension. 2. methodology 2.1 participants and procedure teachers from a previous reading intervention study who voluntarily decided to implement learning progress assessment in the school year 2012-13 were asked to participate in this study. out of 47 teachers, 29 teachers (83% female) agreed to participate. on average they were about 48 years old (m = 47.90 years, sd = 10.99) and had a teaching experience of approximately 22 years (m = 22.45 years, sd = 12.05). the student sample consisted of 568 fourth graders (49% female; 17% with a migration background) from 29 classrooms in 18 german schools. at the first point of measurement, students were approximately 10 years old (m = 9.73 years, sd = 0.48). teachers’ beliefs were assessed with a questionnaire at the beginning of the school year before repeatedly assessing students’ reading skills. participation was voluntary. neither teachers, nor students received incentives for participation. 2.2 measures 2.2.1 teacher beliefs following the procedure by behrmann and souvignier (2013), teachers rated their beliefs on three different scales. the constructivist orientation scale (cos) measures constructivist beliefs with regard to reading instruction. an example for an item of the cos is: “in order to learn how to competently handle texts, it is helpful to let students discuss their own text approaches.” internal consistency of this scale was acceptable (cronbach’s α = .75). the global orientation scale (gos) quantifies global constructivist beliefs. an example for an item of the gos is: “curricular activities should primarily focus on students’ practical learning experiences.” internal consistency of this scale was also acceptable (cronbach’s α = .72). the direct-transmissive orientation scale (dos) measures direct-transmissive beliefs specifically referring to reading instruction. an example for an item of the dos is: “most students are unable to discover reading strategies on their own, and therefore need explicit instruction.” for this scale, internal consistency proved to be good with cronbach’s α = .82. the cos, gos, and dos measures consist of six items each with a four-point likert scale, ranging from (1) i strongly disagree to (4) i strongly agree (see behrmann & souvignier, 2013). a content general direct-transmissive scale was not provided by behrmann and souvignier (2013) and thus not used in this study. 2.2.2 reading progress students’ progress in reading fluency and reading comprehension was assessed over one school year using an internet-based tool for learning progress assessment (see förster & souvignier, 2014). at intervals of three weeks, students individually completed one of eight equivalent reading tests during self-study periods at school. each test took about 10 min on average. in each of the eight tests, learners first completed a maze task in which every seventh word of a text was deleted. students were instructed to replace the 24 gaps as quickly as possible by choosing the correct word among three choices. no time limit was given to assure that all learners had read the complete text. in addition to the number of correct replacements, we also recorded the time needed to complete the maze task. given the need to simultaneously recognize words and construct meaning from text to select the gaps, this test format is in accordance with definitions of reading fluency (e.g., samuels, 1979). we measured reading fluency in the current study as the number of correctly selected words within 1 min. after completing the maze task, students answered 16 comprehension questions that referred to the text from the maze. while answering the questions, the complete and correct text was visible. learners were required to choose the correct answer from four choices. following models of text comprehension (e.g., kintsch, 1998), half of the questions addressed text-based information and thus asked for information that was explicitly contained in the text. the other eight questions assessed the construction of a situation model by requiring students to make inferences from the given information. no time constraints were given to complete the task. we used the number of correct answers as the reading comprehension measure. eight tests were applied over the whole assessment period. four of the tests were based on non-fictional texts about animals and the other four texts were based on fictional detective stories. fictional and non-fictional texts were performed alternately. prior research has documented the psychometric quality of the reading tests (see souvignier, förster, & salaschek, 2014). internal consistencies were found to be high with cronbach’s α ranging from .86 to .89. in addition, correlations to standardized paper-pencil tests measuring reading fluency (r = .60 to .66) and reading comprehension (r = .63 to .65) revealed satisfying criterion validity. the tests demonstrated that they are sensitive to student improvement with significant reading growth over the eight points of measurement. 2.3 data analysis we removed outliers that were two standard deviations below the average of an individuals’ points of measurement because of selective distortion of the data due to guessing, inattention, or failure to make a decision. we also removed outliers that were two standard deviations above the individual average on reading fluency, because fast guesses likely increased these measures. in total, 1.6% of the data were excluded for reading fluency and 0.7% for reading comprehension. data coverage at any point of measurement was continuously higher than 90% with the highest rates of data coverage at the last point of measurement. in total, 6.6% of the reading fluency and 4.4% of the reading comprehension values were missing. we used full information maximum likelihood (fiml) to account for missing data, which has shown to be particularly useful for structural equation modeling (enders & bandalos, 2001). with this procedure, all existing data are used to estimate model parameters. data were nested in three levels. points of measurement (level 1) were nested in students (level 2) which were nested in classrooms (level 3). thus, we applied a three-level latent growth curve model using mplus 8 (muthén & muthén, 2017). with this analysis, prior competence can be modeled on the level of individual students as well as on the level of classrooms. thereby we accounted for the nested structure of the data and prevented underestimation of standard errors (bryk & raudenbush, 1988). given that linear and non-linear reading growth rates have been reported in the literature (parrila et al., 2005; tilstra et al., 2009), we considered linear, quadratic, and free-loading models for the most suitable curve estimation of our data (bollen & curran, 2006; duncan, duncan, & strycker, 2006). we rejected the free-loading model, because its growth factor can only be interpreted as a measure for progress when the determined shape is close to linearity (duncan et al., 2006). after scrutinizing the mean scores of the eight measurement points (see table 4), we only compared fit indices between linear and quadratic models. a linear growth factor is the most legitimate measure of progress for these models (see bollen & curran, 2006). compared to a linear model, a quadratic model additionally has a quadratic growth factor, which is an indicator for an acceleration or deceleration trend in growth over time. its linear factor thereby represents the growth rate at the first measurement point (bollen & curran, 2006). to specify a quadratic model with a comparative linear measure of learning progress over the whole period of the study, we fixed the variance of the quadratic growth factors to zero at the individual and classroom levels. thus, the acceleration or deceleration trend in growth over time was constrained to be equal across all classrooms. at the individual student level, the quadratic factor mean is zero, because lower level (i.e. student level) scores are deviations from means of higher levels (i.e. classroom level) in multilevel modeling. with no variation and a mean of zero, the quadratic growth factor on the individual level remained unspecified. this resulted in a quadratic model with linear growth factors at the individual and classroom level, which could be used as comparable measures of growth over the whole period of measurement. next, solutions of the above described linear and quadratic growth models were compared by using akaike information criterion (aic) and bayesian information criterion (bic). fit indices for reading fluency and reading comprehension suggested an advantage for the quadratic model (see table 1). the classroom level quadratic factor was negative for reading fluency (p < .001) and for reading comprehension (p < .001) indicating decelerated growth over time. based on this result, the quadratic model was selected for the reading fluency and comprehension data. table 1 fit indices of linear and quadratic three-level latent growth curve models without covariates teachers’ beliefs and the interactions of teachers’ beliefs with prior abilities on student and class level were stepwise included into a baseline model resulting in three additional models. all models are shown in table 2 and coefficients are described in table 3. given the three different measures of teachers’ beliefs and the two reading outcomes, we ran 18 models in total. teacher belief data were centred at the grand mean before they were added as covariates to the baseline model in a stepwise procedure (models 1-3, table 6 & 7). this procedure allows for unbiased estimates of higher-level interaction effects (enders & tofighi, 2007). in model 1, only a test for a main effect of teachers’ beliefs on the classrooms’ average learning progress (γ101) was conducted by including the third-level predictor teacher belief (tbj). model 2 additionally tested a cross-level interaction effect (γ201) to analyse if students’ initial achievement moderates effects of teachers’ beliefs on students’ learning gains. thereby, students’ initial reading achievement (r0ij) and progress (r1ij) are determined in relation to their own classrooms’ average initial achievement (ß00j) and progress (ß10j). to determine the cross-level interaction effect in this way, a random effect of prior skill levels on learning growth was required for the student level (ß11j). thus, a regression of individual deviations from class level learning growth (r1ij) on individual deviations from class level prior ability levels (r0ij) was modelled. the parameters r0ij and r1ij thereby indicate the group mean centred individual deviations from the classrooms prior ability and learning progress, respectively. having group mean centred lower level predictors (as r0ij in our case) is crucial to estimate cross-level interaction effects because then, estimates are not biased by interaction, which may be potentially present on class level (enders & tofighi, 2007). model 3 furthermore tested if the classrooms’ initial skills moderated the effect of teachers’ beliefs on the classrooms’ learning growth (γ103). in addition, a test for effects of initial competences on learning progress at the classroom level (γ102) was added to model 3 to allow unbiased estimations of the classroom level interaction effect. table 2 overview of the three-level latent growth curve models table 3 description of coefficients of model 3 3. results 3.1 descriptive statistics and correlations table 4 shows the means, standard deviations, and intercorrelations of reading fluency and reading comprehension data at all points of measurement. all measures of reading ability were highly correlated. table 5 presents the descriptive statistics of the belief scales. moderate to strong positive and negative correlations were found between the cos, gos, and dos measures as expected. a mean score of m = 3.36 for the cos and m = 3.53 for the gos on a 1 to 4 scale indicated that teachers on average agreed with the reading specific and global statements of the constructivist orientation scales. the mean score for the dos was in the middle of the scale (m = 2.48). table 4 intercorrelations, means, and standard deviations of reading fluency and reading comprehension at eight points of measurement (n = 568) table 5 intercorrelations, means, standard deviations of the cos (n = 29), gos (n = 29) and dos (n = 29) scales 3.2 analyses of longitudinal data as shown in table 6, the quadratic three-level latent growth curve model for reading fluency without covariates revealed that students reached an average of 3.63 correctly selected gaps per min at the beginning of fourth grade (γ000). the linear growth (γ100) was 0.39 gaps every three weeks with a moderate deceleration trend indicated by a negative quadratic factor (γ300 = -0.03). we found substantial within-class variance in students’ reading fluency at the beginning of the school year (r0ij = 1.93, p <.01) but no significant variation in linear growth over the course of the school year (r1ij = 0.01, p = 0.15). the same pattern was observed at the classroom level. initial abilities significantly differed between classrooms (u00j = 0.13, p < 0.01), whereas linear growth in reading fluency did not differ between classrooms (u10j = 0.002, p = 0.17). students answered on average 9.79 questions correctly on the reading comprehension test (γ000) at the beginning of fourth grade (see table 5) and had a linear improvement rate of 0.30 answers (γ100) every three weeks on average with a moderate deceleration trend indicated by a negative quadratic factor (γ300 = -0.03). similar to the results found for reading fluency, reading comprehension significantly differed between students from the same classroom at the beginning of the school year (r0ij = 4.91, p <.01), but variation in linear learning growth between individual students of the same classroom was not significant (r1ij = 0.004, p = 0.66). the opposite pattern was found at the classroom level. although no significant differences were found for prior reading comprehension (u00j = 0.66, p = 0.11), different classrooms showed different linear growth in reading comprehension over the school year (u10j = 0.01, p <.01). despite finding some non-significant variances at the student and classroom levels, we analysed the effects of teachers’ beliefs, because adding beliefs as covariates may increase testing power by reducing error variance in the dependent variable (aberson, 2010). table 6 parameters for the three-level latent growth curve baseline models for reading fluency and reading comprehension 3.3 main effects of teacher beliefs results of the three-level latent growth curve models revealed that teachers’ global constructivist beliefs were positively related to students’ progress in reading fluency (see table 7, model 1 & 2) . no significant effects were found for the reading specific cos scale and direct-transmissive beliefs on reading fluency growth. the results for reading comprehension showed that none of the teacher belief scales was significantly related to students’ learning growth (see table 8, model 1 & 2). table 7 effects of teacher beliefs on progress in reading fluency 3.4 interaction of teacher beliefs with student ability we also analysed whether students’ prior ability (relative to the average classroom ability) moderated the effects of teachers’ beliefs on students’ deviation from the average growth of the classroom. results for reading fluency indicate that students’ initial skills did not affect the relation between teachers’ beliefs and individual learning growth (see table 7, model 2 & 3). as hypothesized, however, a significant cross-level interaction was found for reading comprehension (see table 8, model 2 & 3). the effect of teachers’ constructivist beliefs on students’ reading growth was positively moderated by students’ prior skills. hence, teachers with higher constructivist beliefs affected higher growth in reading for students with higher prior ability compared to students with lower ability within their classrooms. we found no interaction for direct-transmissive beliefs (see table 8, model 2 & 3). table 8 effects of teacher beliefs on progress in reading comprehension 3.5 interaction of teacher beliefs with classroom ability we additionally tested whether the prior average reading skills of the classroom moderated the effect of teachers’ beliefs on growth in reading fluency and reading comprehension on the classrooms level. the results show that the classrooms’ prior competences did not moderate the effect of teachers’ beliefs on classrooms’ learning progress (see table 7 & 8, model 3, respectively). in addition, classrooms’ initial reading fluency and reading comprehension ability had no general effect on learning. 4. discussion in this study, we used multilevel latent growth curve modeling to investigate effects of teachers’ constructivist and direct-transmissive beliefs on students’ progress in reading fluency and reading comprehension. moreover, we examined whether effects of teachers’ beliefs depended on prior reading ability. we found that teachers’ global constructivist beliefs had a general positive effect on students’ reading fluency but not on their reading comprehension progress. no significant relations were found between reading specific constructivist beliefs or direct-transmissive beliefs and student growth in reading fluency and reading comprehension. as hypothesized, we found an interaction of teacher beliefs and prior abilities. high achieving students in contrast to low achieving students benefited in their reading comprehension growth from a teacher who holds high constructivist beliefs. the positive effect of teachers’ global constructivist beliefs on reading fluency was unaffected by prior abilities. thus, effects of teacher beliefs seem to depend on the skill under study and interact with students’ prior ability, which we discuss in the following section. 4.1 interpretation of results our finding that general constructivist beliefs of teachers were positively related to students’ growth in reading fluency is in line with our hypotheses. the same effect, however, was expected but not found for reading comprehension and no main effects of reading specific teacher beliefs were found. also, whether or not prior abilities moderated effects of teacher beliefs seems to depend on the respective skill. so how can we explain this pattern of results? given the different findings for reading fluency and reading comprehension, one starting point is to reflect on the specific reading skills and effective ways of teaching these skills. while reading fluency is characterized by the automation of word recognition, reading comprehension requires to intentionally apply reading strategies to construct a representation of the situation model and to connect new information to prior knowledge. consequently, effective reading fluency instruction aims to automatize word recognition, for example by instructing students to (repeatedly) read text passages aloud (e.g., repeated or paired reading; topping, 2006). these decoding practices are mainly student-driven without a particular need of teacher instruction and thus easily match with high constructivist beliefs. effective instruction of reading comprehension, in contrast, is characterized by the explicit instruction of reading strategies by the teacher (souvignier & mokhlesgerami, 2006). this contradicts constructivist views of teaching after which students should and will develop their own problem-solving strategies. actually, as indicated by our finding that prior student ability moderated effects of global and specific constructivist beliefs on students’ progress in reading comprehension, it seems that high achieving students indeed tend to develop their own effective reading strategies and thus profit when a teacher with high constructivist beliefs teaches them. low achieving students, however, might be overstrained to self-regulate their reading comprehension without explicit instruction of strategies. following this argumentation, we would expect that teachers with high constructivist beliefs positively affect the development of skills that require student-driven practices to automatize processes (e.g. word recognition or basic mathematical skills) for all students independent of their prior ability. if, however, the skill is not characterized by high automation but requires strategic behaviour, teachers who provide much student-managed but less teacher-managed instruction due to their high constructivist beliefs might positively affect learning for the high but not the low achieving students. this assumption is in line with findings on child x instruction interactions by connor and colleagues, who showed that low achieving students benefit from teacher-managed instruction but high achieving students benefit from child-managed instruction (e.g. connor et al., 2011; 2004). the positive effects of high constructivist beliefs on student learning have been ascribed to higher cognitive activation and motivational advantages of teaching methods used by teachers with high constructivist beliefs (e.g., savery & duffy, 1995; staub & stern, 2002). these positive motivational advantages, however, will likely not occur if students feel overstrained by the task. we assume that the teaching methods of teachers with high constructivist beliefs will probably be more child-managed and less teacher-managed, and that those methods will likely overstrain low achieving students, who need explicit guidance to build up self-regulated strategic reading behaviour. regarding reading fluency, in contrast, most fourth-grade students will be able to cope with the task of just reading a text, which might enhance their automation of word recognition but not their ability to understand texts. this would explain why we find general positive effects of constructivist beliefs on growth in reading fluency but not in reading comprehension. reading specific views about teaching had no effects on learning growth, which is in contrast to our hypotheses as well as to findings showing that content specific constructivist beliefs about teaching can effect students’ learning (e.g., staub & stern, 2002; souvignier & mokhlesgerami, 2005). nevertheless, results point to expected directions. effects of reading specific beliefs on learning growth may be weaker, because compared to global beliefs, they may be rather limited to lessons specifically dedicated to foster reading skills. we found no effects for direct-transmissive beliefs. thus, although variance in direct transmissive beliefs was highest, this variance did not explain differences in reading fluency or reading comprehension growth. most prior studies have assessed teacher beliefs using a single continuum. the only study investigating teacher beliefs on separate scales in reading is very specific as the teachers applied a strategy-based reading program (text detectives; behrmann & souvignier, 2013). this pattern of results suggests that differences in direct-transmissive beliefs in contrast to constructivist beliefs might be less influential on teaching behaviour and thus do not explain differences in student learning. finally, it should be noted that the interactions with prior ability were found on the student but not the classroom level. an explanation may be that significant variance in prior abilities was found between students of the same classrooms, while the average reading skills of the classrooms in this study seemed to be similar (see table 6). 4.2 limitations a limitation of this study is the small variance in student and classroom level reading growth, despite the representative range of classrooms. the limited variation between classrooms’ growth in reading may have masked effects of teacher beliefs. investigating interaction effects in a sample with a higher variance within and between classrooms would be desirable. moreover, when interpreting our results, one should consider that the belief measures might be affected by social desirability regarding constructivist beliefs leading to ceiling effects and low variances in the cos and gos scales. assuming that social desirability affected our measures it is likely that effects of teachers’ beliefs on students’ reading growth were rather masked than increased by potentially inflated standard errors. given that standard errors are similar for effects of constructivist and direct-transmissive beliefs (see tables 7 & 8), however, we assume that the impact of social desirability may be rather negligible. unfortunately, the construct validity of the three belief scales could not be confirmed using confirmatory factor analysis due to the limited teacher sample. nevertheless, the moderate to strong correlations between the three scales indicate that they are partly independent of each other. in the introduction, we stated that content general as well as reading specific beliefs about teaching are relevant. nevertheless, similar to the study conducted by behrmann and souvgnier (2013), our study included a global constructivist, but not a global direct-transmissive belief scale. further studies should investigate effects of both global constructivist and direct-transmissive beliefs to fully discriminate between content-specific and global beliefs. analysing effects of teachers’ beliefs on students’ learning is largely based on the assumption that beliefs and classroom behaviour of teachers are closely connected (e.g., buehl & beck, 2015). in research designs without classroom observations, as in our study, the variability of teachers’ instructional activity remains an open issue. conversely, by not observing teacher behaviour we ensured unimpaired business-as-usual instruction and thus high ecological validity of the study. regarding the external validity of this study, it should be considered that the results and conclusions of this paper are based on reading skill data from a sample of fourth grade classrooms only. consequently, generalization about different grades and content is limited, especially as our results indicate that effects of teachers’ beliefs may depend on the specific skill under study. the teachers of this study had access to the results of their students’ reading tests, which is a feature of the learning progress assessment tool (förster & souvignier, 2014). thus, in addition to their personal impressions, they had another objective information about the development of their students. we do not assume that the availability of more student information alone is responsible for the effects. moreover, quantity and quality of additional information was the same for all teachers. given that effects of both teacher beliefs and learning progress assessments are assumed to be mediated by instructional behaviour, future studies should investigate the interplay of teacher beliefs, learning progress assessment and instructional behaviour. for example, according to constructivist views, prior knowledge is considered to be particularly important for the learning process (savery & duffy, 1995; staub & stern, 2002). thus, teachers with high constructivist views might be more receptive to additional assessment information. 4.3 conclusion our study complements existing research in a number of ways. first, we investigated effects of teacher beliefs in the domain of reading and–up to our knowledge–provide the first results for effects on reading comprehension. second, we analysed under which conditions which teacher beliefs positively affect students’ reading progress. the analysis of interactions between student ability and different teacher beliefs on both student and classroom level adds to our knowledge of the interplay between teacher and student variables during the learning process and provides a novel perspective to child x instruction interactions. third, we assessed reading fluency and reading comprehension progress with eight measurements across the school year, thereby ensuring the reliable assessment of student progress. this work partly confirmed general effects of teachers’ beliefs about teaching. global constructivist, but not reading specific beliefs about teaching had an impact on students’ reading fluency. no general effects of teachers’ beliefs on reading comprehension progress were found but teachers with stronger constructivist beliefs affected higher learning growth in reading comprehension for students with higher prior ability compared to lower performing students within the classrooms. a similar interaction was not found for reading fluency indicating that effects of teachers’ beliefs on growth in reading fluency is unaffected by prior skills. these skill-specific findings for effects of teachers’ beliefs and their interaction with students’ ability might be explained by differences regarding the optimal instruction of these skills that correspond more or less to constructivist views of teaching. our study thus adds to our understanding of the conditions under which constructivist teacher beliefs are positively associated with student learning. keypoints we investigated effects of teachers’ constructivist and direct-transmissive beliefs on students’ reading fluency and reading comprehension progress. according to child x instruction interactions, we hypothesized that prior abilities moderate the effects of teachers’ beliefs on student learning. we assessed constructivist and direct-transmissive beliefs on separate scales and modelled reading progress across eight points of measurement over one school year. global constructivist beliefs were positively related to reading fluency only. no other main effects were found. effects of teachers’ constructivist beliefs on students’ growth in reading comprehension depended on students’ prior abilities and was higher for high compared to low achieving students. references aberson c. l. (2010). applied power analysis for the behavioral sciences. new york: routledge. behrmann, l., & souvignier, e. (2013). pedagogical content beliefs about reading instruction and their relation to gains in student achievement. european journal of psychology of education, 28 (3), 1023–1044. doi:10.1007/s10212-012-0152-3 bollen, k. a., & curran, p. j. (2006). latent curve models. a structural equation perspective. hoboken, nj: john wiley & sons, inc. brown, r., & pressley, m. (1994). self-regulated reading and getting meaning from text: the transactional strategies instructional model and its ongoing validation. in d. schunk & b. zimmerman (eds.), self-regulation of learning and performance: issues and educational applications (pp. 155–180). hillsdale, nj: erlbaum. bryk, a. s., & raudenbush, s. w. (1988). toward a more appropriate conceptualization of research on school effects: a three-level hierarchical linear model. american journal of education, 97 (1), 65–108. doi:10.1086/443913 buehl, m. b., & beck, j. s. (2015).the relationship between teachers’ beliefs and teachers’ practices. in h. fives, & m. g. gill (eds.), international handbook of research on teachers’ beliefs(pp. 66–84). new york: routledge. bunting, c. e. (1985). dimensionality of teacher educational beliefs: a validation-study. the journal of experimental education, 53 (4), 188–192. doi:10.1080/00220973.1985.10806380 connor, c. m., morrison, f. j., fishman, b., giuliani, s., luck, m., underwood, p. s., bayraktar, a., crowe e. c., & schatschneider, c. (2011). testing the impact of child characteristics x instruction interactions on third graders’ reading comprehension by differentiating literacy instruction. reading research quarterly, 46(3), 189-221. doi: 10.1598/rrq.46.3.1 connor, c. m., morrison, f. j., & petrella, j. n. (2004). effective reading comprehension instruction: examining child by instruction interactions.journal of educational psychology, 96(4), 682–698. doi:10 .1037/0022-0663.96.4.682 cromley, j. g., & azevedo, r. (2007). testing and refining the direct and inferential mediation model of reading comprehension. journal of educational psychology, 99(2), 311–325. doi:10.1037 /0022-0663.99.2.311 dubberke, t., kunter, m., mcelvany, n., brunner, m., & baumert, j. (2008). lerntheoretische überzeugungen von mathematiklehrkräften: einflüsse auf die unterrichtsgestaltung und den lernerfolg von schülerinnen und schülern[mathematics teachers’ beliefs: their impact on instructional quality and student achievement]. zeitschrift für pädagogische psychologie [german journal of educational psychology], 22 (34), 193–206. doi:10.1024/1010-0652.22.34.193 duncan, t. e., duncan, s. c., & strycker, l. a. (2006). an introduction to latent variable growth curve modeling: concepts, issues, and application. mahwah, nj: lawrence erlbaum associates. enders, c. k., & bandalos, d. l. (2001). the relative performance of full information maximum likelihood estimation for missing data in structural equation models. structural equation modeling, 8 (3), 430–457. doi:10.1207/s15328007sem0803_5 enders, c. k., & tofighi, d. (2007). centering predictor variables in cross-sectional multilevel models: a new look at an old issue. psychological methods, 12(2), 121. doi:10.1037/1082-989x.12.2.121 farnia, f., & geva, e. (2013). growth and predictors of change in english language learners' reading comprehension. journal of research in reading, 36(4), 389–421. doi:10.1111/jrir.12003 fives, h., lacatena, n., & gerard, c. (2015). teachers’ beliefs about teaching (and learning). in m. g. gill, & h. fives (eds.), international handbook of research on teachers’ beliefs (pp. 249–265). new york: routledge. förster, n., & souvignier, e. (2014). learning progress assessment and goal setting: effects on reading achievement, reading motivation and reading self-concept. learning and instruction, 32, 91–100. doi:10.1016/j.learninstruc.2014.02.002 gill, m. g., & hoffman, b. (2009). shared planning time: a novel context for studying teachers’ discourse and beliefs about learning and instruction. teacherscollege record, 111(5) , 1242–1273. retrieved from http://www.tcrecord.org/content.asp?contentid =15241 hattie, j. (2012). visible learning for teachers: maximizing impact on learning. london: routledge. kagan, d. m. (1992). implication of research on teacher belief. educational psychologist, 27(1), 65–90. doi:10.1207/s15326985ep2701_6 kintsch, w. (1998). comprehension: a paradigm for cognition. new york: cambridge university press. kleickmann, t. (2007). zusammenhänge fachspezifischer vorstellungen von grundschullehrkräften zum lehren und lernen mit fortschritten von schülerinnen und schülern im konzeptuellen naturwissenschaftlichen verständnis [coherences of elementary teachers ’ subject-related beliefs on teaching and learning with students ’progresses in scientific comprehension](doctoral dissertation). retrieved from https://miami.uni-muenster.de/record/642aa4ce-7149-4cdb-a938-c37f3c64cbe2 muthén, l. k., & muthén b. o. (2017). mplus user’s guide. eighth edition.los angeles: muthén & muthén. nation, k., cocksey, j., taylor, j. s., & bishop, d. v. (2010). a longitudinal investigation of early reading and language skills in children with poor reading comprehension. journal of child psychology and psychiatry, 51(9), 1031–1039. doi:10.1111/j.1469-7610.2010.02254.x national institute of child health and human development (2000). teaching children to read: an evidence-based assessment of the scientific research literature on reading and its implications for reading instruction. washington, dc: u.s. government printing office. nespor, j. (1987). the role of beliefs in the practice of teaching. journal of curriculum studies, 19(4), 317-328. doi:10.1080/0022027870190403 organisation for economic co-operation and development (2009).creating effective teaching and learning environments: first results from talis, oecd. paris, france: oecd. retrieved from http://www.oecd.org/education/school /creatingeffectiveteachingandlearningenvironmentsfirstresultsfromtalis.html pajares, m. f. (1992). teachers’ beliefs and educational research: cleaning up a messy construct. review of educational research, 62(3), 307–332. doi:10.3102 /00346543062003307 paris, s. g., cross, d. r., & lipson, m. y. (1984). informed strategies for learning: a program to improve children’s reading awareness and comprehension. journal of educational psychology, 76, 1239–1252. doi:10.1037/0022–0663.76.6.1239 parrila, r., aunola, k., leskinen, e., nurmi, j. e., & kirby, j. r. (2005). development of individual differences in reading: results from longitudinal studies in english and finnish. journal of educational psychology, 97(3), 299–319. doi:10.1037/0022-0663.97.3.299 peterson, p. l., fennema, e., carpenter, t. p., & loef, m. (1989). teacher's pedagogical content beliefs in mathematics. cognition and instruction, 6(1), 1–40. doi:10.1207/s1532690xci0601_1 rasinski, t. v., reutzel, d. r., chard, d., & linan-thompson, s. (2011). reading fluency. in m. l. kamil, p. d. pearson, birr moija, e., & p. p. afflerbach (eds.), handbook of reading research(vol. iv) (pp. 286–319). new york, nj: routledge. resnick, l. b., & hall, m. w. (1998). learning organizations for sustainable education reform. daedalus, 127(4), 89–118. retrieved from http://www.jstor.org/stable /20027524?seq=1#page_scan_tab_contents richardson, v. (2003). preservice teachers’ beliefs. in j. raths, & a. c. mcaninch (eds.), teacher beliefs and classroom performance: the impact of teacher education, volume 6: advances in teacher education (pp. 1–22). greenwich, ct: information age. samuels, s. j. (1979). the method of repeated readings. the reading teacher, 32(4), 403–408. retrieved from http://www.jstor.org/stable/20194790?seq=1#page_scan_tab _contents savery, j. r., & duffy, t. m. (1995). problem based learning: an instructional model and its constructivist framework. educational technology, 35(5), 31–38. retrieved from http://citeseerx.ist.psu.edu/viewdoc /summary? singer, j. d., & willett, j. b. (2003). applied longitudinal data analysis: modeling change and event occurrence . new york: oxford university press. snider, v. e., & roehl, r. (2007). teachers’beliefs about pedagogy and related issues. psychology in the schools, 44(8), 873–886. doi:10.1002/pits.20272 souvignier, e., förster, n., & salaschek, m. (2014). quop: ein ansatz internetbasierter lernverlaufsdiagnostik mit testkonzepten für lesen und mathematik [quop: an internet based approach with testing concepts for reading and mathematics]. in m. hasselhorn, w. schneider, & u. trautwein (eds.), lernverlaufsdiagnostik [learning progress assessment](pp. 239–256). goettingen: hogrefe. souvignier, e., & mokhlesgerami, j. (2005). implementation eines programms zur vermittlung von lesestrategien im deutschunterricht [moving strategy-oriented reading instruction into the classroom: the role of the teacher]. zeitschrift für pädagogische psychologie [german journal of educational psychology], 19 (4), 249–261. doi:10.1024/1010-0652.19.4.249 souvignier, e., & mokhlesgerami, j. (2006). using self-regulation as a framework for implementing strategy-instruction to foster reading comprehension. learning & instruction, 16, 57-71. doi:10.10167j.learninstruc.2005.12.006 speer, d. c., & greenbaum, p. e. (1995). five methods for computing significant individual client change and improvement rates: support for an individual growth curve approach. journal of consulting and clinical psychology, 63(6), 1044–1048. doi:10.1037/0022-006x.63.6.1044. staub, f. c., & stern, e. (2002). the nature of teachers’ pedagogical content beliefs matters for students' achievement gains: quasi-experimental evidence from elementary mathematics. journal of educational psychology, 94(2), 344–355. doi:10.1037/0022-0663.94.2.344 tilstra, j., mcmaster, k., van den broek, p., kendeou, p., & rapp, d. (2009). simple but complex: components of the simple view of reading across grade levels. journal of research in reading, 32(4), 383–401. doi:10.1111/j.1467-9817.2009.01401.x topping, k. (1987). paired reading: a powerful technique for parent use. the reading teacher, 40(7), 608–614. retrieved from http://www.jstor.org/stable/20199562 topping, k. j. (2006). building reading fluency: cognitive, behavioral, and socioemotional factors and the role of peer-mediated learning. in s. j. samuels, & a. e. farstrup (eds.), what research has to say about fluency instruction.(pp. 106-129). newark, de: international reading association. willett, j. b. (1989). some results on reliability for the longitudinal measurement of change: implications for the design of studies of individual growth. educational and psychological measurement, 49(3), 587–602. doi:10.1177/001316448904900309 woolley, s. l., benjamin, w.-j. j., & woolley, a. w. (2004). construct validity of a self-report measure of teacher beliefs related to constructivist and traditional approaches to teaching and learning. educational and psychological measurement, 64(2), 319–331. doi:10.1177/0013164403261189 microsoft word schwaighofer et al_publication.docx           frontline  learning  research  vol.5  no.  1  (2017)  58  -­‐  75   issn  2295-­‐3159       executive functions in the context of complex learning: malleable moderators? matthias schwaighofer1, markus bühner, frank fischer university of munich, germany article received 19 august / revised 4 january / accepted 9 january / available online 15 february abstract executive functions are crucial for complex learning in addition to prior knowledge. in this article, we argue that executive functions can moderate the effectiveness of instructional approaches that vary with respect to the demand on these functions. in addition, we suggest that engagement in complex activity contexts rather than specific cognitive training paradigms may enhance executive functions and yield practically relevant transfer effects to other cognitive abilities. we develop several hypotheses and principles for how to improve executive functions in these contexts. for future research, we suggest to systematically investigate the moderating role of executive functions in learning environments with varying degrees of instructional support and varying context characteristics. we identify potential factors influencing the improvement of executive functions to be considered in a systematic research program. keywords: executive functions; complex learning; moderation; training                                                                                                                           1corresponding author: matthias schwaighofer, department of psychology, ludwig-maximilians-universität münchen, leopoldstraße 13, 80802 munich, germany. email: m.schwaighofer@psy.lmu.de doi: http://dx.doi.org/10.14786/flr.v5i1.268 schwaighofer  et  al       | f l r     59   1. problem in recent research on learning and instruction, learners’ prior knowledge and working memory capacity are considered crucial prerequisites for complex learning. it is widely accepted that prior knowledge is variable. working memory is usually considered to be limited and constant apart from developmental and intra-individual variation (e.g., due to task demands; trezise & reeve, 2014). moreover, the interplay of prior knowledge and working memory is typically considered to be the main determinant of learning in complex learning environments. two recent patterns of findings in other areas of research may challenge this apparent consensus. first, working memory is but one from a set of basic cognitive functions called executive functions (andersson, 2010; miyake & friedman, 2012). there is evidence that other executive functions are highly relevant for complex learning as well. second, the view that working memory and other executive functions cannot be altered substantially, except through developmental changes and apart from intra-individual variations, has been challenged by recent studies. in this article we link these two bodies of research and delineate the impact of an innovative framework for research on learning and instruction. 2. executive functions in the context of complex learning 2. 1 definition of executive functions   executive functions are basic cognitive functions (andersson, 2010) and as such are typically conceived as effective across domains (friedman & miyake, 2017). executive functions receive attention due to their strong relation to self-control (also referred to as self-regulation) (miyake & friedman, 2012). self-control is considered important for numerous outcomes relevant to daily life, including learning (e.g., blair, ursache, greenberg, vernon-feagans, & the family life project investigators, 2015; mischel et al., 2011). the literature on executive function typically focusses on working memory, shifting, and inhibition as core executive functions (for an overview see miyake & friedman, 2016). working memory is used to store and manipulate information temporarily (baddeley, allen, & hitch, 2011). working memory processes typically involve retrieval and transformation of information from long-term memory and from the environment. these processes are strongly dependent on the capacity of working memory (ecker, lewandowsky, & oberauer, 2014; ecker, lewandowsky, oberauer, & chee, 2010). the capacity of working memory is limited to about four chunks (collections of concepts with strong associations), on average (cowan, 2001). individual capacity limits can differ widely, ranging from 2 to 6 chunks (cowan, 2001). individual differences in working memory capacity may be explained by three mechanisms, which include maintenance of information in short-term storage, retrieving information from long-term memory, and attentional control (shipstead, lindsey, marshall, & engle, 2014). working memory capacity seems to be crucial for complex learning (e.g., sweller, 2011). however, capacity is not the only aspect of working memory that may be relevant for complex learning. updating is sometimes also considered to be an important aspect of working memory. updating involves the continuous monitoring and adding or removal of content in working memory (miyake, friedman, emerson, witzki, & howerter, 2000). based on three experiments, ecker et al. (2014) conclude that the removal of irrelevant information from working memory is the crucial part of substitution, hence also for working memory updating. retrieval and transformation of information from long-term memory as two further updating processes have found to be strongly related to working memory capacity (ecker et al., 2014; 2010). converging evidence indicates that updating and working memory capacity are both strongly related to each other (e.g., wilhelm, hildebrandt, & oberauer, 2013; schmiedek, hildebrandt, lövdén, wilhelm, schwaighofer  et  al       | f l r     60   & lindenberger, 2009; schmiedek, lövdén, & lindenberger, 2014). in a nutshell, updating processes seem to be largely related to working memory capacity that is considered to be crucial for complex learning (see section 2.2.). therefore, for research on complex learning we suggest putting an emphasis on working memory capacity. this emphasis does not however, imply that updating-specific processes, like the removal of items from working memory, are irrelevant for complex learning. shifting, also denoted as attention switching, refers to the ability to switch flexibly between multiple representations and strategies with changing task demands (miyake et al., 2000). inhibition is the ability to overcome dominant or prepotent responses (miyake & friedman, 2012), which enables an individual to avoid acting on the first reaction that comes to mind. inhibition may be involved in working memory and shifting processes. for example, the intent to store only information of a given text in working memory presumably requires inhibition of interference from other sources of information (e.g., thoughts unrelated to the text). switching attention between different parts of a text requires inhibition of previous parts of the text. for relatively simple tasks, studies have shown that there is no specific variance left for inhibition controlling for a common factor extracted from working memory and shifting (e.g., miyake & friedman, 2012). in addition, for several approaches to measuring inhibition the construct was deemed to be multi-dimensional (e.g., friedman & miyake, 2004; krumm et al., 2009). despite these findings, inhibition seems to be important for many processes requiring working memory and shifting and may explain variance in the context of complex learning. thus, we consider the role of the executive function of inhibition for complex learning in our framework. in the remainder of the article, the term executive functions refers to working memory, shifting and inhibition whereas the term cognitive functions also includes cognitive abilities or skills such as fluid intelligence and long-term memory. 2.2 relevance of executive functions for complex learning a main instructional goal with complex learning is to simultaneously facilitate conceptual learning in a domain and the development of skills. complex learning typically provides the learner with non-trivial activities related to inquiry, design or problem solving in real domains. learners’ appropriate engagements in these activities require the coordinated application of domain concepts as well as strategies or skills (see also van merriënboer & kirschner, 2012, p. 2). in the following, we argue that the executive functions of working memory, shifting, and inhibition are important for complex learning. in general, the executive function of working memory seems to be crucial for learning because it provides a mental workspace in which information can be held while carrying out other activities (gathercole & alloway, 2007). working memory capacity is considered to be important for complex learning within cognitive load theory (e.g., sweller, 2011). according to cognitive load theory, the role of working memory capacity is connected to the element interactivity of learning material. elements are information to be learned or information that has been learned (e.g., a concept; sweller, 2010). collections of interacting elements can therefore be conceived as chunks. learning material with many interacting elements, which cannot be learned in isolation, is high in element interactivity and induces a high load on one’s central executive function, namely working memory. interacting elements that have to be simultaneously processed in working memory may be unrelated to learning relevant information. however, the interacting elements in relevant learning material can be crucial to fostering understanding (sweller, 2010). thus, we consider complex learning material as being high in element interactivity. working memory capacity is assumed to play an important role in dealing with element interactivity in fostering learning. connecting elements to units (i.e., chunks) can reduce the number of elements required for complex learning that have to be held in working memory. whether elements can be grouped to units depends on the amount of prior knowledge. learners with high prior knowledge in a domain hold schemas in long-term memory which can be represented as single elements in working memory and therefore induce low working memory load (sweller, 2011). in contrast, for learners with low levels of prior knowledge, complex problem solving can induce a high load in working memory (schnotz & kürschner, 2007). thus, not enough working memory capacity can be devoted to making sense of meaningful material (sweller, 2011) and small learning schwaighofer  et  al       | f l r     61   gains may result (e.g., sweller, mawer, & howe, 1982). therefore, in the framework of cognitive load theory, instructional methods can be improved with respect to their effectiveness and efficiency if they take into account the individual limits of working memory capacity (sweller, 2011). individual differences in working memory capacity have been linked to the variability in several complex learning outcomes. for example, working memory capacity has been shown to be associated with math achievement (peng, namkung, barnes, & sun, 2016), reading comprehension (daneman & merikle, 1996), accomplishment in chemistry (e.g., tsaparlis, 2005), and problem solving (e.g., bühner, kröner, & ziegler, 2008). beyond these associations with complex learning outcomes, working memory capacity is moderately correlated with fluid intelligence (e.g., ackermann, beier, & boyle, 2002; redick, unsworth, kelly, & engle, 2012). compared with research on working memory, the interest in the relevance of shifting, as an executive function for cognitive achievements, has grown only over the last years. research suggests that shifting is relevant for tasks that require switching between different aspects of a problem (e.g., agostino, johnson, & pascual-leone, 2010; blair, knipe, & gamson, 2008). for example, solving mathematical problems involving multiple equations may require switching attention between equations or parts of them. meta-analyses indicate that shifting is associated with reading and math achievements as complex learning outcomes (yeniad, malda, mesman, van ijzendoorn, & pieper, 2013). however, the meta-analysis by yeniad et al. (2013) did not control for other cognitive functions that could explain the associations between shifting and mathematics achievement (bull & lee, 2014). in addition, the studies included in the meta-analysis by yeniad et al. (2013) were conducted on children’s executive functions, which is of note because there is evidence that measures of executive functions do not clearly separate all executive functions prior to the age of 15 years (lee, bull, & ho, 2013). the role of shifting for academic achievements in adulthood is not yet clear. some indications could be taken from research showing a relationship of common executive functioning assessed by one measure in early childhood and a shifting factor in adolescence (friedman, miyake, robinson, & hewitt, 2011). in addition, some studies demonstrate that working memory can already be separated from shifting in late childhood and executive functions show considerable stability from adolescence into early adulthood (friedman et al., 2016). given the stability of shifting and its putative relevance for academic achievements in childhood, we submit that shifting may also be relevant for complex learning of older children, adolescents and adults. with respect to the executive function of inhibition, several primary studies show the association of inhibition and mathematics achievement (e.g., espy, mcdiarmid, cwik, stalets, hamby, & senn, 2004; lan, legare, ponitz, li, & morrison, 2011). results of a meta-analysis by jacob and parkinson (2015) indicate that inhibition is related to math and reading achievements in addition to shifting and working memory. notably, at least for mathematics achievement, it remains unclear whether the role of inhibition depends on updating. several studies found only updating to be associated with math achievement and inhibition was only related to math achievement when there was no control for updating (for an overview, see bull & lee, 2014). however, there are also arguments that consider inhibition as the common factor of executive functions (miyake & friedman, 2012), which could explain associations between working memory capacity and shifting on the one hand and cognitive achievements on the other hand. in sum, individual differences in working memory, shifting, and inhibition are important for explaining variance in various cognitive achievements related to complex learning. as will be discussed in the next section, interactions among these individual differences and different instructional approaches for complex learning have hardly been investigated systematically. 2.3 aptitude treatment interactions in the context of complex learning research on aptitude treatment interactions (atis) (e.g., snow & lohman, 1984) suggests that the effectiveness of instructional approaches depends on learning prerequisites. the prerequisite that is typically investigated in the context of complex learning is prior knowledge (e.g., kester & kirschner, 2012). this is not surprising given that the acquisition of domain-specific schemas can be considered a main goal in complex schwaighofer  et  al       | f l r     62   learning environments (kalyuga & singh, 2015) and prior knowledge is regarded as a strong predictor of learning (e.g., dochy, segers, & buehl, 1999). the moderating influence of prior knowledge on the effectiveness of various instructional supports has been demonstrated frequently. a general finding of this research is that the benefits of instructional support decrease with increasing prior knowledge. this pattern of results has been termed the expertise reversal effect (kalyuga, 2007) and has been found for several instructional approaches, including instructional support using animations, simulations, and hypermedia-based materials (e.g., kalyuga, 2013; kalyuga, rikers, & paas, 2012). executive functions have seldom been measured as potential moderators in research on atis despite their potential relevance for complex learning. perhaps most surprisingly, the executive function of working memory capacity has rarely been assessed as a potential moderator using objective, reliable, and valid measures. de jong (2010) lists some studies as exceptions, but in these as well as more recent studies (e.g., lehmann, goussios, & seufert, 2016; park, korbach, & brünken, 2015), working memory was measured by using only one single task. hence, task-specific influences are not eliminated (schwaighofer, bühner, & fischer, 2016). furthermore, in some of these studies, working memory capacity has been included as a control variable and not as a moderator variable (e.g., berends & van lieshout, 2009; park et al., 2015). working memory capacity might moderate the effects of different kinds of instructional support aimed at reducing working memory load. instructional support through worked examples is assumed to be effective in reducing load in working memory (e.g., renkl, 2014). the benefit of worked examples might be especially pronounced for learners with low working memory capacity (van gog & rummel, 2010). yet, shifting might also moderate the effectiveness of different kinds of instructional support. the moderating role of shifting may depend on context characteristics such as task complexity and their demand for shifting. unsupported problem solving, which requires switching between different sources of information (e.g. within a problem or between problem and additional information), may be overwhelming for learners with low shifting ability. these learners may benefit from instructional support that reduces the demands on shifting (e.g. by integrating information from different sources). in contrast, learners with high shifting ability may not benefit from such instructional support because they can deal with high shifting demands. in a recent study, schwaighofer et al. (2016) investigated the moderating role of the executive functions of working memory capacity and shifting as well as fluid intelligence for the effectiveness of worked examples versus problem solving. compared to problem solving, worked examples had lower benefits the higher the shifting was. presumably, problem-solving puts a high demand on shifting ability. inhibition may moderate the effectiveness of various instructional approaches, especially when these approaches differ in the need for inhibiting responses or resisting interference (e.g., from material irrelevant to learning). furthermore, inhibition may influence the potential moderating role of working memory and shifting due to the involvement of inhibition in working memory and shifting processes. for example, shifting among different information relevant to solving a problem may involve inhibiting irrelevant information. the inhibition of information unrelated to problem-solving may also free up working memory capacity to deal with relevant information (see also trezise & reeve, 2014). learners with high inhibition ability may have few difficulties dealing with problems with high demands on this executive function. therefore, these learners may benefit less from instructional support that lowers these demands. besides these executive functions, other cognitive functions may moderate the effectiveness of different kinds of instructional support. one further potential moderator is fluid intelligence, defined as the “ability to reason and to solve new problems” (könig, bühner, & mürling, 2005, p. 245). fluid intelligence is not as fundamental when compared to executive functions and may be influenced by these functions (diamond, 2013). the conceptualization of intelligence is based on implicit theories of researchers and their assumptions about different methods to test it (oberauer, schulze, wilhelm, & süß, 2005). nonetheless, fluid intelligence is worth considering in the context of complex learning along with executive functions for several reasons. reasoning which information seems to be necessary to solve a novel, complex problem may be a cognitive process unique to fluid intelligence. specifically, there is empirical evidence that fluid intelligence is especially important for solving novel and complex tasks (e.g., primi, ferrão, & almeida, 2010). learners with low fluid intelligence schwaighofer  et  al       | f l r     63   may have severe difficulties figuring out which information is necessary to solve a novel problem. these learners may benefit from instructional support that reduces the demands on fluid intelligence (e.g. by providing information relevant to solve a problem at hand). learners with high fluid intelligence do not need this instructional support because they can more easily use their reasoning capacity to determine which information is necessary to solve a given problem. in line with this argumentation, schwaighofer et al. (2016) found that the benefits of worked examples compared to problem solving were lower the higher fluid intelligence was. in that study, the moderating role of fluid intelligence was independent of working memory capacity and shifting. in summary, both executive functions as well as fluid intelligence are likely to moderate the effects of instruction on complex learning outcomes. these moderation effects could occur independent of prior knowledge or in addition to the moderating role of prior knowledge. however, these factors have hardly been studied systematically and in their interplay in research on complex learning. moreover, there are methodological concerns with respect to atis in the context of complex learning. in several analyses of atis, only global moderation effects of treatments on outcomes have been tested. sometimes only high or low values in an aptitude such as prior knowledge or working memory capacity were compared (e.g., kalyuga & sweller, exp. 3, 2004; seufert, schütze, & brünken, 2009). the main problem of these comparisons and of testing only global moderation effects is that it is not known for which regions of values of a moderator the independent variable has a significant effect on the dependent variable. for example, the regions of values in prior knowledge for which the presence of worked examples has a significant effect on knowledge at post-test is unknown when only high or low values in prior knowledge are compared. the moderating role of executive functions may additionally depend on context characteristics such as task complexity or time constraints. complex tasks can vary in their demand on cognitive functions, and certain instructional support may reduce these demands. for example, shifting demands in a problem-solving condition may be reduced in the presence of worked examples (schwaighofer et al., 2016). in addition to task complexity, other context characteristics such as time constraints may play a role for atis. in the study by schwaighofer et al. (2016) the effectiveness of worked examples was not moderated by working memory capacity. one plausible reason could be that learners did not have to work under time pressure. working memory overload may arise only under time pressure and when offloading (e.g., by using notepads) is prevented (de jong, 2010). thus, the moderating role of executive functions for the effects of instructional approaches on learning outcomes may depend on the context, such as whether learners have to work on tasks under time pressure. in summary, we can say that executive functions have been shown to play an important role for complex learning. specific moderation processes that may involve the interplay of executive functions have not yet been studied systematically in the context of instruction. in addition, instructional research, including measures of executive functions, has so far not fully used the methodological possibilities to objectively, reliably, and validly measure the constructs and explored their possible local moderation effects. implicitly, the instructional research reported in the preceding sections assumes stability of executive functions. however, in a different area of research, several intervention studies have shown that executive functions seem to be malleable to a certain extent (melby-lervåg, redick, & hulme, 2016). we submit that this research has relevance for instructional research as well. in the next section, we outline research on the malleability of executive functions, focusing on approaches that may improve these functions. 2.4 malleability of executive functions 2.4.1 developmental and further influences on executive functions current theorizing and findings indicate that executive functions develop early in childhood, improve until adolescence, and then decline with age. in addition to age, factors such as stress, sadness, and loneliness can have deleterious effects on executive functions (for an overview, see diamond, 2013). schwaighofer  et  al       | f l r     64   importantly, some age-independent factors can lead to variation in executive functioning within small and longer periods of time. for instance, experiencing frustration in the classroom can lead to impaired executive functioning, especially the inhibition of prepotent responses, thus generating high intra-individual variation in this functioning (pnevmatikos & trikkaliotis, 2013). also, higher intrinsic motivation is associated with higher executive functions (sosic-vasic, keis, lau, spitzer, & streb, 2015). negative emotions such as anxiety may impair executive functions (eysenck et al., 2007) but the extent of impairment may also depend on initial levels of executive functioning (working memory) and vary within short periods of time (trezise & reeve, 2014). perceived strain may also impair executive functions temporarily (e.g., liston, mcewen, & casey, 2009); this holds for different age groups (tun, miller-martinez, lachman, & seeman, 2013). 2.4.2 approaches to improve executive functions subsequently, we review existing evidence on the malleability of executive functions in what has been called narrow activity contexts (schwaighofer, fischer, & bühner, 2015) such as computerized training of working memory. these contexts involve material low in complexity such as colored squares. we then turn to conjectures concerned with the improvement of executive functions in more complex activity contexts. recently, a wealth of studies has investigated the malleability of executive functions by applying different training methods. the most popular approach involves computerized training of cognitive functions, particularly working memory training. the aim of working memory training and other approaches to train executive functions is to improve the targeted cognitive function and hence yield transfer effects to other cognitive abilities or skills. although some primary studies yielded promising results (e.g., jaeggi, buschkuehl, jonides, & perrig, 2008), recent meta-analyses cast doubts on the transfer of these effects to other cognitive abilities such as mathematical abilities (melby-lervåg et al., 2016; schwaighofer et al., 2015). in support of these doubts, recent meta-analytic evidence (karbach & verhaeghen, 2014; melby-lervåg & hulme, 2016) suggests that training of executive functions yields no substantial transfer effects to cognitive abilities that were not the focus of the training. this seems to hold, at least when training groups are compared with active control groups (e.g. groups completing physical exercises; karbach & verhaeghen, 2014; melby-lervåg & hulme, 2016). one reason for the lack of practically relevant transfer effects for computerized training of executive functions may lie in the nature of the narrow activity contexts. within these narrow contexts, tasks only tap one or more executive functions in isolation. however, complex-learning tasks may require the well-coordinated interplay of several cognitive functions. for example, mathematical problem solving may tax the executive functions of working memory, inhibition, and shifting and may require their coordination. solving a complex mathematical equation may involve keeping in mind relevant elements in the equation and additional information from texts. this constellation taxes working memory. switching among various information sources to solve the equation might put a high demand on the shifting function. resisting irrelevant information or distracting stimuli (e.g., a smartphone buzzing) taxes inhibition. thus, working memory, shifting, and inhibition may have to operate in concert when solving the mathematical equations. executive functions usually have to be used together and their interplay presumably has to be coordinated. for example, keeping in mind information from additional text while switching among information of the equation and resisting interference from a buzzing smartphone involves working memory capacity, shifting, inhibition, and presumably the coordinated interplay of these functions. the involvement of executive functions and their interplay in complex learning activities may depend on task demands. for example, borella and ribaupierre (2014) found that working memory predicts text comprehension when participants were allowed to look at the pertaining text. when participants were not allowed to look at the pertaining text, inhibition (resistance to distractor interference) predicts text comprehension in addition to working memory. the authors argue that preventing irrelevant information from entering working memory is particularly important when learners have to cope with information in addition to a text comprehension task (borella & ribaupierre, 2014). executive functions appear to be substantially correlated but also have a unique variance (friedman & miyake, 2017). this unique variance of the different executive functions may explain why working memory and shifting can be simultaneously and uniquely related to complex academic achievements such as reading abilities (e.g., jacobson et al., 2016). activities such as solving schwaighofer  et  al       | f l r     65   mathematical problems or reading texts are often embedded in more complex contexts such as classroom instruction at schools. complex activity contexts can be designed such that activities within these contexts improve executive functions and yield transfer effects to academic achievements (see diamond & ling, 2016). the demands on executive functions may depend on the problems or tasks used in complex activity contexts. one could hypothesize that the demands may be reduced or enhanced by varying the degree of instructional support. evidence in support of this hypothesis would be the moderation effects executive functions have in learning environments with varying degrees of instructional support and demand on these functions. the moderating role would suggest that varying degrees of instructional support put different demands on executive functions. therefore, a possible enhancement of executive functions can depend on the degree of instructional support. this has been labeled the “detrimental instruction hypothesis” (schwaighofer et al., 2015), which states that instructional support that reduces the demands on executive functions can be detrimental with respect to improvements of these functions and transfer effects. this hypothesis is related to the assumption that executive functions have to be continuously challenged on a high level (i.e., tasks have to adapt to consistently guarantee a high demand on executive functions; schwaighofer et al., 2015). complementary to this assumption, in their comprehensive review of the literature on training of executive functions, diamond and ling (2016) note that to avoid frustrating or boring the learner, the challenge should not be too high or too low. in addition, the longer executive functions are challenged, the higher the likelihood of training effects. training of executive functions probably should not pause too long because improvements in executive functions usually decline with increasing time lags between training sessions (diamond & ling, 2016). a further speculation concerning the improvement of executive functions in complex activity contexts has been labeled the “cognitive flexibility hypothesis”. according to this hypothesis, executive functions have to be challenged in different contexts to yield transfer of improved executive functioning to other tasks (schwaighofer et al., 2015). the cognitive flexibility hypothesis is based on research suggesting that learning in multiple contexts can foster transfer (e.g., spiro, coulson, feltovich, & anderson, 1988). in multiple complex activity contexts, executive functions might be trained with respect to their coordinated interplay in order to adapt to different complex tasks. thus, executive functions and their interplay may be trained to adapt to different demands. as a result, executive functions may not only improve, but also the chance to yield transfer effects to other complex tasks requiring executive functions and their interplay might increase. in contrast to isolated computerized trainings of executive functions, complex tasks provide the additional opportunity to acquire relevant knowledge of a certain domain. moreover, two principles for the training of executive functions postulated by diamond and ling (2016) deserve attention. first, these authors argue that people with the poorest executive functions seem to gain most from training programs. this principle is based on findings from various intervention studies, ranging from working memory training (e.g., holmes et al., 2010) to physical training (kramer & erickson, 2007). of particular interest is the study by blair and raver (2014), in which a comprehensive curriculum (tools of the mind), shown to be effective in an earlier study (diamond, barnett, thomas, & munro, 2007), was implemented in kindergarten classes for two years. in this curriculum, activities designed to improve executive functions are intertwined with curricular learning activities. teachers are trained to organize and manage instruction in order to improve children’s self-regulation skills through focused interactions with their classmates. because executive functions seem to be important for self-regulation (e.g., blair & ursache, 2011), activities intended to foster selfregulation skills may improve executive functions. the tools of the mind curriculum also includes individualized, differentiated scaffolding depending on the mastery of specific skills. in addition, the children meet with teachers on a weekly basis to reflect on their learning and to develop metacognitive strategies. in contrast to the study by diamond et al. (2007), the study by blair and raver (2014) included baseline measures of executive functions to rule out effects due to pre-existing differences in executive functions. children in the treatment group showed improvements in executive functions and outcomes such as reading and math. the effect sizes for several outcomes, such as measures of executive functions and reasoning were highest for low-income children assumed to be disadvantaged with respect to the development of their executive functions (blair & raver, 2014). according to the second principle postulated by schwaighofer  et  al       | f l r     66   diamond and ling (2016), effects of training are especially pronounced for measures that push the limits of participants’ executive functions. in a study by diamond et al. (2007), the benefits of the tools of the mind curriculum were largest for the executive function tasks with the highest demands. a possible extension of the principle proposed by diamond and ling (2016) refers to the nature of tasks used to assess transfer effects following improvements in executive functions. it is possible that transfer effects are also larger for transfer tasks that put high demands on executive functions. support for this assumption comes from research indicating that executive functions better predict higher than lower level academic skills (e.g., jacobson et al., 2016). tasks used to measure higher academic skills may better capture individual differences in executive functions than transfer effects on lower academic skills. therefore, these tasks could also be more appropriate for measuring transfer effects following improvements of executive functions. the principles and hypotheses mentioned above can be seen as complementary. we assume that the more facilitative principles are considered, the more effective the training of executive functions will be. also, the opportunity to transfer effects to other cognitive abilities (e.g., the ability to solve complex academic tasks in a domain) will increase the more principles are considered. 3. future directions for research on complex learning and executive functions 3.1 investigating the moderating role of executive functions in learning environments with varying demands on these functions in studies on complex learning, the moderating role of learning prerequisites has mainly been analyzed for prior knowledge. different levels of prior knowledge have been linked to varying demands on working memory. working memory demands are lower for persons with high prior knowledge due to a reduced number of elements which have to be stored in working memory or a reduced level of element interactivity, respectively (sweller, 2011). as a consequence, working memory may not moderate the effectiveness of varying degrees of instructional support when learners already dispose of high prior knowledge. within the context of complex learning, interactions between prior knowledge and working memory have been considered to be crucial aspects of the human cognitive architecture. however, research on executive functions indicates that shifting and inhibition can also be relevant for complex learning and potentially moderate the effectiveness of instructional approaches. in addition to working memory, the moderating influence of shifting and inhibition may also depend on prior knowledge. high prior knowledge learners may be better able to quickly identify relevant information to solve a problem. thus, these learners do not have to frequently shift among potentially relevant and irrelevant information to solve a problem. furthermore, learners with high prior knowledge do not have to inhibit a large amount of irrelevant information. nonetheless, for learners with high prior knowledge the demands on executive functions could be high under certain conditions, which have to be systematically explored. executive functions may moderate the effectiveness of instructional approaches varying in the demands on executive functions. accordingly, instructional approaches could be effectively adapted to consider different levels of executive functions. for example, presenting information relevant for solving a problem together with the problem may lower the requirement to integrate spatially separated information. therefore, the demand on working memory may be reduced. executive functions have been included in only a few studies and have not been measured on a latent variable level. reliable and valid measures of executive functions, however, have recently been established (e.g., friedman et al., 2016; redick, broadway et al., 2012). thus, we propose that executive functions receive more attention in the context of complex learning. moreover, fluid intelligence is worthwhile to consider in addition to executive functions because it is presumably important for solving novel complex problems. notably, we do not propose considering the moderating role of executive functions or other cognitive abilities such as fluid intelligence instead of prior knowledge. rather, we schwaighofer  et  al       | f l r     67   suggest systematically addressing the conditions for interactions between prior knowledge and executive functions with respect to their moderating role for the effects of varying instructional approaches. more specifically, we suggest that future research should consider the following three recommendations. a) exploring the conditions for moderation effects in the context of complex learning. to systematically address the potential moderating role of prior knowledge and executive functions, we propose exploring the conditions for moderation effects of one or several cognitive prerequisites. examples of these conditions or context characteristics may include the complexity of tasks or time constraints.  for instance, in a hypothetical scenario learners may have to frequently switch their attention between information from different sources (e.g., different websites or texts) and task instruction to solve a challenging problem. in order to better focus on one source of information after a shift of attention, learners have to inhibit irrelevant information from previous sources (e.g., interesting details of a text that are not relevant for a problem at hand). under these conditions or context characteristics, learners with low shifting ability and/or low inhibition ability in particular could be overstrained. instructional approaches as worked examples may lower the demands on shifting ability and inhibition. providing solution steps (parts of worked examples) for complex tasks may reduce the necessity to switch attention between information from different sources and therefore, reduce shifting demands. moreover, fewer necessary shifts among relevant and irrelevant information can also reduce the demand for inhibiting irrelevant information. learners with low shifting and inhibition ability would benefit from this instructional support. thus, we hypothesize that the executive functions of shifting and inhibition may moderate the effects of the instructional approach (i.e., presence of worked examples) on learning outcomes. it is important to note that the moderating role of executive functions presumably depends on the differences between instructional approaches in their demands on these functions. shifting demands may be high in conditions with low and high degrees of instructional support if both conditions require frequent switching between several information sources to meet task demands. the high degree of instructional support may not significantly reduce the demand on shifting, and learners with low shifting ability will presumably not benefit from the instructional support. thus, shifting ability may not moderate the effect of the degree of instructional support on learning outcomes under these circumstances. accordingly, we hypothesize that executive functions moderate the effects of instructional approaches on learning outcomes only if these approaches significantly differ in their demands on executive functions. it would be a worthwhile task to review different instructional approaches like learning-by-design, inquiry learning, problem-based learning, direct instruction, and worked examples with respect to their typical demands on the different executive functions and their interplay. knowledge on moderation effects can be helpful for selecting or designing particular instructional approaches. a moderating role of executive functions for the effectiveness of different instructional approaches suggests that these approaches differ with respect to the demands on executive functions. approaches with high demand on executive functions may only be beneficial for learners with high executive functions. instructional approaches might be adapted to meet learners’ executive functions or other potential moderators, such as prior knowledge and fluid intelligence. one aim of adapting instructional approaches would be to foster domain knowledge by providing adequate support. a second aim is connected to the training of executive functions by establishing and maintaining an optimal level of demand on executive functions. we elaborate on this second aim in section 3.2. b) measuring executive functions on a latent variable level. in studies investigating interactions between working memory and instructional approaches, working memory has predominantly been measured by using single tasks. the executive function of shifting has rarely been measured in research on atis. we propose to measure executive functions on the latent variable level using several tasks for each cognitive function to minimize task-specific variance. shifting as well as working memory capacity and inhibition can be measured with three computerized tasks for each construct (e.g., friedman et al., 2016; redick, broadway, et al., 2012). in the case of working memory, even short versions allow assessment of working memory capacity on a latent variable level within about 25 minutes (oswald, mcabee, redick, & hambrick, 2015). in these working memory tasks, participants have to complete series of process-storage schwaighofer  et  al       | f l r     68   sequences (e.g., solve a simple mathematical equation and store a letter) and afterwards recall the stimuli (e.g., the letters) (oswald et al., 2015). in typical shifting tasks, participants have to flexibly switch between two tasks, for instance by indicating the shape or the color of a stimulus. a cue presented before the stimulus indicates which task to perform. shifting can be measured as the difference in mean reaction times between correct switch and no-switch trials. in no-switch trials, the same task has to be performed on consecutive tasks (i.e., the same cue appears consecutively). in switch trials, the task changes from one trial to the subsequent one (i.e., the cue changes) (friedman et al., 2011). notably, recent research proposes applying alternative scoring methods to measure shifting. one of these alternative methods involves calculating bin scores. bin scores consider the speed of switching between switch and no-switch trials and accuracy. traditional difference scores do not consider the accuracy when switching between two tasks (draheim, hicks, & engle, 2016). thus, difference scores only reflect the speed of switching between two tasks but do not account for the number of errors or correct shifts respectively. a frequently used task to measure inhibition is the antisaccade task (e.g., friedman et al., 2016). this task requires participants to correctly identify a target stimulus by inhibiting the tendency to look at a cue previously presented at a different location as the target stimulus. the dependent variable for the antisaccade task is typically the proportion of correct responses (friedman et al., 2008). every task for measuring a single executive function involves specific content material. hence, there is systematic variance due to the task context. to alleviate the influence of the task context, several tasks should be used to measure executive functions on a latent variable level (see also schwaighofer et al., 2016; miyake & friedman, 2012). c) applying more appropriate statistical analyses. finally, moderator analyses in the context of complex learning may benefit from recent methodological advances. analyses of moderated moderation effects are feasible with recent analysis techniques (hayes, 2013). these tools also allow, for example, the application of the so-called johnson-neyman technique, which makes it possible to analyze the effects of an independent variable on a dependent variable for different ranges of values of one or two continuous moderators. hence, local moderations can be identified with no loss of information compared to categorizing continuous moderators. notably, categorical moderators and covariates can also be included in moderation analyses. for example, the shared variance of fluid intelligence and working memory can be controlled for. thus, the unique contributions of single moderators can be identified in the complex pattern of effects for several cognitive functions. the moderating role of executive functions could provide information about the conditions under which complex learning activities put lower or higher demands on executive functions. this information is connected to the idea of training executive functions in complex activity contexts over longer periods of time. complex learning activities with high demands on executive functions and their coordinated interplay may be more likely to improve these functions. we propose considering the development of executive functions in complex activity contexts in addition to fostering knowledge acquisition. executive functions may improve in complex activity contexts and therefore acquiring domain knowledge may not be the sole aim for complex learning. 3.2 promising routes for training of executive functions in complex activity contexts cognitive abilities for solving complex academic tasks likely require not only unique contributions of executive functions but also the coordinated interplay of these functions. we submit the hypothesis that training the coordinated interplay of executive functions rather than training isolated functions facilitates transfer. the coordinated interplay of executive functions may be best accomplished in real academic or informal learning environments, including complex tasks (e.g., solving statistical problems) that are optimized with respect to their demands on executive functions when necessary. in contrast, isolated computerized trainings with a limited set of stimuli may be less appropriate for addressing the large number of factors potentially involved in improving executive functions (e.g., coordinated interplay of executive functions in several contexts; or social support). in addition, the aim of acquiring relevant domain knowledge is not usually included in these individualized and narrow training programs. thus, complex activity contexts are promising in that they are intended to help approach the dual goal of advancing domain knowledge and executive functions at the same time. schwaighofer  et  al       | f l r     69   we suggest systematically addressing six issues for enhancing executive functions in complex learning activities. a) investigating the coordinated interplay of executive functions. the hypothesis that the coordinated interplay of executive functions is crucial for improving these cognitive functions should be tested empirically. the overarching rationale behind this line of research has to be further elaborated. isolated trainings of single components of the cognitive system (i.e., training of specific executive functions) with material low in complexity did not show transfer effects to other cognitive abilities (e.g., schwaighofer et al., 2015; melby-lervåg & hulme, 2016). one explanation for the missing transfer effects could be that executive functions rarely operate in isolation when learners are faced with complex learning tasks. if complex learning tasks put a consistently high demand on executive functions and especially their coordinated interplay, these functions and their interplay may adapt to high task demands. the adaptation may result in the improvement of single executive functions and importantly the interplay of executive functions. furthermore, cognitive abilities or skills related to complex learning activities (e.g., reading abilities, learning about physiological, metabolic processes or handling chemical equations) may also improve if they put a similar demand on executive functions operating in parallel. in line with these suggestions, the implementation of training of executive functions within complex activity contexts such as specific learning environments in schools has yielded promising results with respect to the improvement of executive functions and transfer effects to other cognitive abilities (e.g., blair & raver, 2014; diamond et al., 2007; raver et al., 2011). research on training of executive functions in complex activity contexts can benefit from studies investigating the moderating role of executive functions in learning environments with varying degrees of instructional support and context characteristics. the moderating role of executive functions suggests that the demand on these functions may depend on the degree of instructional support and/or context characteristics. one example is the moderating role of shifting for the effects of worked examples versus problem solving. the moderating role suggests that the demand on shifting is high when solving complex problems, which require frequent shifts among relevant and irrelevant information due to a rather low degree of instructional support (schwaighofer et al., 2016). thus, solving complex problems requiring frequent shifts among relevant and irrelevant information could offer a way for improving shifting, especially for learners who are challenged to the limits of their cognitive functions. if shifting, working memory, and inhibition moderate the effects of different learning environments on learning outcomes, their coordinated interplay may be important under particular conditions. therefore, the moderating role of executive functions in complex activity contexts may inform on the demands of instructional approaches on executive functions and their interplay. scenarios with a high demand on executive functions and their interplay can serve as fruitful starting points for improving executive functions in the long run. b) implementing adaptive instructional support. the degree of instructional support may be lowered when executive functions improve over time to put a consistently high demand on executive functions. the degree of instructional support can be adaptively reduced depending on the knowledge of learners, such as by fading solution steps of worked examples (e.g., salden, aleven, renkl, & schwonke, 2009). however, procedures for fading instructional support depending on the level of executive functions (rather than domain knowledge) have yet to be developed. learning environments with low versus high demands on executive functions and their interplay can serve as starting points for determining whether putting different demands on executive functions over longer periods of time influences these functions and their coordinated interplay. the effects of these learning environments on executive functions and their interplay should be considered in addition to effects on knowledge acquisition. further research could compare the effects of instructional approaches adapted to improve executive functions with effects of instructional approaches adapted to foster knowledge acquisition. c) testing the cognitive flexibility hypothesis in the context of executive functioning. the cognitive flexibility hypothesis submits that challenging executive functions in various contexts likely yields transfer effects of improved executive functions to other tasks (schwaighofer et al., 2015). one way schwaighofer  et  al       | f l r     70   of testing this hypothesis would be a comparison of executive function training in various complex contexts with different content materials to training in one context with material similar in content. the various contexts and materials should put different demands on executive functions and their coordinated interplay. the instructional environments would have to establish consistently high demands. this suggestion is in line with the detrimental instructional support hypothesis and the explanations regarding the adaptation of the degree of instructional support detailed above. d) testing the superiority of training effects for learners with poor executive functions. future studies could also address the principle that learners with poor executive functions gain most from training of these functions. many instructional support measures designed to support learners with low prior knowledge appear to have helped them more than learners with higher expertise (e.g., kalyuga, 2013). an open question is whether different levels of prior knowledge and executive functions influence the effects of approaches to improve executive functions in complex activity contexts. these presumably local moderation effects can be tested by applying the johnson-neyman technique (hayes, 2013). e) pushing the limits: new forms of assessment of executive functions and their coordinated interplay. effects of instructional approaches on executive functions should be measured with tasks that push the limits of participants’ executive functions and/or require the interplay of executive functions. in addition to single measures of executive functions that are considered to be independent of domain-specific knowledge (e.g., complex span tasks that measure working memory capacity; see redick, broadway, et al., 2012), more knowledge-rich transfer measures should be used. knowledge-rich transfer measures include problems from specific domains that require the interplay of executive functions and domain knowledge to develop an appropriate solution. f) setting the stage through indirect factors influencing executive functions. finally, complementary to proximal factors that may influence executive functions in the long run, some rather distal variables seem to have an influence on executive functions and thus should not be neglected. diamond and ling (2016) propose that the most successful interventions for improving executive functions will be those that train these functions directly and support them indirectly as well. indirect support involves minimizing factors that appear to decrease executive functions such as stress or sadness and support factors that appear to enhance executive functions such as joy and social support (diamond & ling, 2016). the long-term impact of factors influencing executive functions such as frustration (pnevmatikos & trikkaliotis, 2013) and motivation (sosic-vasic et al., 2015) need to be further addressed in future research. 4. conclusion to conclude, recent instructional approaches typically focus on the acquisition of academically relevant knowledge with complex learning tasks. executive functions are important learning prerequisites and need to be more systematically investigated as to their moderating effect on knowledge acquisition under different instructional support conditions. assuming malleability of executive functions as well, instructional conditions may affect their development, at least in the long run. it is important to know more about the conditions under which instructional approaches influence the development of executive functions. conditions for optimizing knowledge acquisition (e.g., moderate cognitive load) may be detrimental to the development of executive functions because the demand on these functions is low. it is not yet known whether and how learning environments can be designed that foster domain learning and enhance executive functions at the same time. schwaighofer  et  al       | f l r     71   keypoints within the frameworks of complex learning and aptitude treatment interaction, we argue that executive functions and other cognitive functions (particularly fluid intelligence) may moderate the effects of various instructional approaches in interaction with prior knowledge. we suggest that the moderating role of cognitive prerequisites may depend on the degree of instructional support and context characteristics. based on research on training of executive functions, we propose that complex activity contexts have the potential to improve executive functions and yield practically relevant transfer effects to other cognitive abilities. this article proposes several lines for future research on the interaction of cognitive prerequisites and various instructional approaches as well as the training of executive functions in complex activity contexts. references ackerman, p. l., beier, m. e., & boyle, m. o. (2002). individual differences in working memory within a nomological network of cognitive and perceptual speed abilities. journal of experimental psychology. general, 131(4), 567–589. doi: 10.1037/0096-3445.131.4.567 agostino, a., johnson, j., & pascual-leone, j. (2010). executive functions underlying multiplicative reasoning: problem type matters. journal of experimental child psychology, 105(4), 286–305. doi: 10.1016/j.jecp.2009.09.006 andersson, u. (2010). skill development in different components of arithmetic and basic cognitive functions: findings from a 3-year longitudinal study of children with different types of learning difficulties. journal of educational psychology, 102(1), 115–134. doi:10.1037/a0016838 baddeley, a. d., allen, r. j., & hitch, g. j. (2011). binding in visual working memory: the role of the episodic buffer. neuropsychologia, 49(6), 1393–1400. doi: 10.1016/j.neuropsychologia.2010.12.042 berends, i. e., & van lieshout, e. c. d. m. (2009). the effect of illustrations in arithmetic problem-solving: effects of increased cognitive load. learning and instruction, 19, 345–353. doi: 10.1016/j.learninstruc.2008.06.012 blair, c., knipe, h., & gamson, d. (2008). is there a role for executive functions in the development of mathematics ability? mind, brain, and education, 2(2), 80–89. doi: 10.1111/j.1751-228x.2008.00036.x blair, c., & raver, c. c. (2014). closing the achievement gap through modification of neurocognitive and neuroendocrine function: results from a cluster randomized controlled trial of an innovative approach to the education of children in kindergarten. plos one, 9(11), e112393. doi: 10.1371/journal.pone.0112393 blair, c., & ursache, a. (2011). a bidirectional theory of executive functions and self-regulation. in k. vohs, & r. baumeister (eds.), handbook of self-regulation. (2 ed., pp. 300-320). new york: guilford press. blair, c., ursache, a., greenberg, m., vernon-feagans, l., & the family life project investigators (2015). multiple aspects of self-regulation uniquely predict mathematics but not letter–word knowledge in the early elementary grades. developmental psychology, 51(4), 459–472. doi: 10.1037/a0038813 borella, e., & de ribaupierre, a. (2014). the role of working memory, inhibition, and processing speed in text comprehension in children. learning and individual differences, 34, 86–92. doi: 10.1016/j.lindif.2014.05.001 bühner, m., kröner, s., & ziegler, m. (2008). working memory, visual–spatial-intelligence and their relationship to problem-solving. intelligence, 36(6), 672–680. doi:10.1016/j.intell.2008.03.008 bull, r., & lee, k. (2014). executive functioning and mathematics achievement. child development perspectives, 8(1), 36–41. doi: 10.1111/cdep.12059 cowan, n. (2001). the magical number 4 in short-term memory: a reconsideration of mental storage capacity. schwaighofer  et  al       | f l r     72   behavioral and brain sciences, 24(1), 87–185. doi:10.1017/s0140525x01003922 daneman, m., & merikle, p. m. (1996). working memory and language comprehension: a meta-analysis. psychonomic bulletin & review, 3(4), 422–433. doi: 10.3758/bf03214546 de jong, t. (2010). cognitive load theory, educational research, and instructional design: some food for thought. instructional science, 38, 105–134. doi: 10.1007/s11251-009-9110-0 diamond, a. (2013). executive functions. annual review of psychology, 64(1), 135–168. doi: 10.1146/annurevpsych-113011-143750 diamond, a., barnett, w. s., thomas, j., & munro, s. (2007). preschool program improves cognitive control. science, 318 (5855), 1387-1388. doi: 10.1126/science.1151148 diamond, a., & ling, d. s. (2016). conclusions about interventions, programs, and approaches for improving executive functions that appear justified and those that, despite much hype, do not. developmental cognitive neuroscience, 18, 34–48. doi: 10.1016/j.dcn.2015.11.005 dochy, f., segers, m., & buehl, m. m. (1999). the relation between assessment practices and outcomes of studies: the case of research on prior knowledge. review of educational research, 69(2), 145–186. doi:10.3102/00346543069002145 draheim, c., hicks, k. l., & engle, r. w. (2016). combining reaction time and accuracy: the relationship between working memory capacity and task switching as a case example. perspectives on psychological science, 11(1), 133–155. doi: 10.1177/1745691615596990 ecker, u. k. h., lewandowsky, s., & oberauer, k. (2014). removal of information from working memory: a specific updating process. journal of memory and language, 74, 77–90. doi: 10.1016/j.jml.2013.09.003 ecker, u. k. h., lewandowsky, s., oberauer, k., & chee, a. e. h. (2010). the components of working memory updating: an experimental decomposition and individual differences. journal of experimental psychology: learning, memory, and cognition, 36(1), 170–189. doi: 10.1037/a0017891 espy, k. a., mcdiarmid, m. m., cwik, m. f., stalets, m. m., hamby, a., & senn, t. e. (2004). the contribution of executive functions to emergent mathematic skills in preschool children. developmental neuropsychology, 26, 465-486. doi: 10.1207/s15326942dn2601_6 eysenck, m. w., derakshan, n., santos, r., & calvo, m. g. (2007). anxiety and cognitive performance: attentional control theory. emotion, 7(2), 336–353. doi: 10.1037/1528-3542.7.2.336 friedman, n. p., & miyake, a. (2004). the relations among inhibition and interference control functions: a latent-variable analysis. journal of experimental psychology. general, 133(1), 101–135. doi:10.1037/00963445.133.1.101 friedman, n. p., & miyake, a. (2017). unity and diversity of executive functions: individual differences as a window on cognitive structure. cortex, 86, 186–204. doi: 10.1016/j.cortex.2016.04.023 friedman, n. p., miyake, a., altamirano, l. j., corley, r. p., young, s. e., rhea, s. a., & hewitt, j. k. (2016). stability and change in executive function abilities from late adolescence to early adulthood: a longitudinal twin study. developmental psychology, 52(2), 326–340. doi: 10.1037/dev0000075 friedman, n. p., miyake, a., robinson, j. l., & hewitt, j. k. (2011). developmental trajectories in toddlers’ self-restraint predict individual differences in executive functions 14 years later: a behavioral genetic analysis. developmental psychology, 47(5), 1410–1430. doi: 10.1037/a0023750 friedman, n. p., miyake, a., young, s. e., defries, j. c., corley, r. p., & hewitt, j. k. (2008). individual differences in executive functions are almost entirely genetic in origin. journal of experimental psychology: general, 137(2), 201–225. doi: 10.1037/0096-3445.137.2.201 gathercole, s. e., & alloway, t. p. (2007). understanding working memory. a classroom guide. london, uk: harcourt assessment. retrieved from http://www.york.ac.uk/res/wml/classroom%20guide.pdf hayes, a. f. (2013). introduction to mediation, moderation, and conditional process analysis: a regressionbased approach. new york: guilford press. holmes, j., gathercole, s. e., place, m., dunning, d. l., hilton, k. a., & elliott, j. g. (2010). working memory deficits can be overcome: impacts of training and medication on working memory in children with adhd. applied cognitive psychology, 24(6), 827–836. doi: 10.1002/acp.1589 jacob, r., & parkinson, j. (2015). the potential for school-based interventions that target executive function to improve academic achievement: a review. review of educational research, 85(4), 512–552. doi: schwaighofer  et  al       | f l r     73   10.3102/0034654314561338 jacobson, l. a., koriakin, t., lipkin, p., boada, r., frijters, j. c., lovett, m. w., … mahone, e. m. (2016). executive functions contribute uniquely to reading competence in minority youth. journal of learning disabilities. doi: 10.1177/0022219415618501 jaeggi, s. m., buschkuehl, m., jonides, j., & perrig, w. j. (2008). improving fluid intelligence with training on working memory. proceedings of the national academy of sciences, 105(19), 6829–6833. kalyuga, s. (2007). expertise reversal effect and its implications for learner-tailored instruction. educational psychology review, 19, 509–539. doi: 10.1007/s10648-007-9054-3 kalyuga, s. (2013). effects of learner prior knowledge and working memory limitations on multimedia learning. procedia social and behavioral sciences, 83, 25–29. doi:10.1016/j.sbspro.2013.06.005 kalyuga, s., rikers, r., & paas, f. (2012). educational implications of expertise reversal effects in learning and performance of complex cognitive and sensorimotor skills. educational psychology review, 24(2), 313– 337. doi: 10.1007/s10648-012-9195-x kalyuga, s., & singh, a.-m. (2015). rethinking the boundaries of cognitive load theory in complex learning. educational psychology review. doi: 10.1007/s10648-015-9352-0 kalyuga, s., & sweller, j. (2004). measuring knowledge to optimize cognitive load factors during instruction. journal of educational psychology, 96(3), 558–568. doi: 10.1037/0022-0663.96.3.558 karbach, j., & verhaeghen, p. (2014). making working memory work: a meta-analysis of executive-control and working memory training in older adults. psychological science, 25(11), 2027–2037. doi: 10.1177/0956797614548725 kester, l., & kirschner, p. a. (2012). cognitive tasks and learning. in n. seel (ed.), encyclopedia of the sciences of learning (pp. 619-622). new york: springer. könig, c. j., bühner, m., & mürling, g. (2005). working memory, fluid intelligence, and attention are predictors of multitasking performance, but polychronicity and extraversion are not. human performance, 18(3), 243–266. doi:10.1207/s15327043hup1803_3 kramer, a. f., & erickson, k. i. (2007). capitalizing on cortical plasticity: influence of physical activity on cognition and brain function. trends in cognitive sciences, 11(8), 342–348. doi: 10.1016/j.tics.2007.06.009 krumm, s., schmidt-atzert, l., bühner, m., ziegler, m., michalczyk, k., & arrow, k. (2009). storage and nonstorage components of working memory predicting reasoning: a simultaneous examination of a wide range of ability factors. intelligence, 37(4), 347–364. doi:10.1016/j.intell.2009.02.003 lan, x., legare, c., ponitz, c. c., li, s., & morrison, f. j. (2011). investigating the links between the subcomponents of executive function and academic achievement: a cross-cultural analysis of chinese and american preschoolers. journal of experimental child psychology, 108, 677-692. doi:10.1016/j.jecp.2010.11.001 lee, k., bull, r., & ho, r. m. h. (2013). developmental changes in executive functioning. child development, 84, 1933–1953. doi: 10.1111/cdev.12096 lehmann, j., goussios, c., & seufert, t. (2016). working memory capacity and disfluency effect: an aptitudetreatment-interaction study. metacognition and learning, 11(1), 89–105. doi: 10.1007/s11409-015-9149-z liston, c., mcewen, b. s., & casey, b. j. (2009). psychosocial stress reversibly disrupts prefrontal processing and attentional control. proceedings of the national academy of sciences, 106(3), 912–917. doi: 10.1073/pnas.0807041106 melby-lervåg, m., & hulme, c. (2016). there is no convincing evidence that working memory training is effective: a reply to au et al. (2014) and karbach and verhaeghen (2014). psychonomic bulletin & review, 23(1), 324–330. doi: 10.3758/s13423-015-0862-z melby-lervåg, m., redick, t. s., & hulme, c. (2016). working memory training does not improve performance on measures of intelligence or other measures of "far transfer": evidence from a meta-analytic review. perspectives on psychological science, 11 (4), 512-534. doi: 10.1177/1745691616635612 mischel, w., ayduk, o., berman, m. g., casey, b. j., gotlib, i. h., jonides, j., … shoda, y. (2011). “willpower” over the life span: decomposing self-regulation. social cognitive and affective neuroscience, 6(2), 252–256. doi:10.1093/scan/nsq081 miyake, a., & friedman, n. p. (2012). the nature and organization of individual differences in executive schwaighofer  et  al       | f l r     74   functions: four general conclusions. current directions in psychological science, 21(1), 8–14. doi:10.1177/0963721411429458 miyake, a., friedman, n. p., emerson, m. j., witzki, a. h., howerter, a., & wager, t. d. (2000). the unity and diversity of executive functions and their contributions to complex “frontal lobe” tasks: a latent variable analysis. cognitive psychology, 41, 49–100. doi: 10.1006/ cogp.1999.0734 oberauer, k., schulze, r., wilhelm, o., & süß, h. m. (2005). working memory and intelligence—their correlation and their relation: comment on ackerman, beier, and boyle (2005). psychological bulletin, 131, 61–65. doi: 10.1037/0033-2909.131 oswald, f. l., mcabee, s. t., redick, t. s., & hambrick, d. z. (2015). the development of a short domaingeneral measure of working memory capacity. behavior research methods, 47(4), 1343–1355. doi: 10.3758/s13428-014-0543-2 park, b., korbach, a., & brünken, r. (2015). do learner characteristics moderate the seductive-details-effect? a cognitive-load-study using eye-tracking. journal of educational technology & society, 18(4), 24–36. peng, p., namkung, j., barnes, m., & sun, c. (2016). a meta-analysis of mathematics and working memory: moderating effects of working memory domain, type of mathematics skill, and sample characteristics. journal of educational psychology, 108(4), 455–473. doi: 10.1037/edu0000079 pnevmatikos, d., & trikkaliotis, i. (2013). intraindividual differences in executive functions during childhood: the role of emotions. journal of experimental child psychology, 115(2), 245–261. doi: 10.1016/j.jecp.2013.01.010 primi, r., ferrão, m. e., & almeida, l. s. (2010). fluid intelligence as a predictor of learning: a longitudinal multilevel approach applied to math. learning and individual differences, 20(5), 446–451. doi:10.1016/j.lindif.2010.05.001 redick, t. s., broadway, j. m., meier, m. e., kuriakose, p. s., unsworth, n., kane, m. j., & engle, r. w. (2012). measuring working memory capacity with automated complex span tasks. european journal of psychological assessment, 28, 164–171. doi: 10.1027/ 1015 5759/a000123 redick, t. s., unsworth, n., kelly, a. j., & engle, r. w. (2012). faster, smarter? working memory capacity and perceptual speed in relation to fluid intelligence. journal of cognitive psychology, 24, 844–854. doi: 10.1080/20445911.2012.704359 renkl, a. (2014). toward an instructionally oriented theory of example-based learning. cognitive science, 38(1), 1–37. doi:10.1111/cogs.12086 salden, r. j., aleven, v. a., renkl, a., & schwonke, r. (2009). worked examples and tutored problem solving: redundant or synergistic forms of support? topics in cognitive science, 1(1), 203–213. schmiedek, f., hildebrandt, a., lövdén, m., wilhelm, o., & lindenberger, u. (2009). complex span versus updating tasks of working memory: the gap is not that deep. journal of experimental psychology: learning, memory, and cognition, 35(4), 1089–1096. doi: 10.1037/a0015730 schmiedek, f., lövdén, m., & lindenberger, u. (2014). a task is a task is a task: putting complex span, n-back, and other working memory indicators in psychometric context. frontiers in psychology, 5. doi: 10.3389/fpsyg.2014.01475 schnotz, w., & kürschner, c. (2007). a reconsideration of cognitive load theory. educational psychology review, 19(4), 469–508. doi: 10.1007/s10648-007-9053-4 schwaighofer, m., bühner, m., & fischer, f. (2016). executive functions as moderators of the worked example effect: when shifting is more important than working memory capacity. journal of educational psychology, 108(7), 982–1000. https://doi.org/10.1037/edu0000115 schwaighofer, m., fischer, f., & bühner, m. (2015). does working memory training transfer? a meta-analysis including training conditions as moderators. educational psychologist, 50(2), 138–166. doi: 10.1080/00461520.2015.1036274 seufert, t., schütze, m., & brünken, r. (2009). memory characteristics and modality in multimedia learning: an aptitude–treatment–interaction study. learning and instruction, 19(1), 28–42. http://doi: 10.1016/j.learninstruc.2008.01.002 shipstead, z., lindsey, d. r. b., marshall, r. l., & engle, r. w. (2014). the mechanisms of working memory capacity: primary memory, secondary memory, and attention control. journal of memory and language, schwaighofer  et  al       | f l r     75   72, 116–141. doi: 10.1016/j.jml.2014.01.004 snow, r. e., & lohman, d. (1984). toward a theory of cognitive aptitude for learning from instruction. journal of educational psychology, 76, 347–376. sosic-vasic, z., keis, o., lau, m., spitzer, m., & streb, j. (2015). the impact of motivation and teachers’autonomy support on children’s executive functions. frontiers in psychology, 6. doi: 10.3389/fpsyg.2015.00146 spiro, r. j., coulson, r. l., feltovich, p. j., & anderson, d. k. (1988). cognitive flexibility theory: advanced knowledge acquisition in illstructured domains. in proceedings of the tenth annual conference of the cognitive science society (pp. 375–383). hillsdale, nj: erlbaum. sweller, j. (2010). element interactivity and intrinsic, extraneous, and germane cognitive load. educational psychology review, 22(2), 123–138. doi: 10.1007/s10648-010-9128-5 sweller, j. (2011). cognitive load theory. in j. mestre, & b. ross (eds.), the psychology of learning and motivation: cognition in education (vol. 55, pp. 37 – 76). oxford: academic press. sweller, j., mawer, r. f., & howe, w. (1982). consequences of history-cued and means–end strategies in problem solving. american journal of psychology, 95, 455–483. doi: 10.2307/1422136 trezise, k., & reeve, r. a. (2014). cognition-emotion interactions: patterns of change and implications for math problem solving. frontiers in psychology, 5. doi: 10.3389/fpsyg.2014.00840 tsaparlis, g. (2005). non-algorithmic quantitative problem solving in university physical chemistry: a correlation study of the role of selective cognitive factors. research in science & technological education, 23, 125–148. doi: 10.1080/02635140500266369 tun, p. a., miller-martinez, d., lachman, m. e., & seeman, t. (2013). social strain and executive function across the lifespan: the dark (and light) sides of social engagement. aging, neuropsychology, and cognition, 20(3), 320–338. doi: 10.1080/13825585.2012.707173 van gog, t., & rummel, n. (2010). example-based learning: integrating cognitive and social-cognitive research perspectives. educational psychology review, 22(2), 155–174. doi:10.1007/s10648-010-9134-7 van merriënboer, j. j. g., & kirschner, p. a. (2012). ten steps to complex learning: a systematic approach to four-component instructional design (2nd ed.). new york: routledge. wilhelm, o., hildebrandt, a., & oberauer, k. (2013). what is working memory capacity, and how can we measure it? frontiers in psychology, 4. doi: 10.3389/fpsyg.2013.00433 yeniad, n., malda, m., mesman, j., van ijzendoorn, m. h., & pieper, s. (2013). shifting ability predicts math and reading performance in children: a meta-analytical study. learning and individual differences, 23, 1– 9. doi:10.1016/j.lindif.2012.10.004 atteveldt et al publication frontline learning research vol.6 no. 3 (2018) 186 203 issn 2295-3159 neuroimaging of learning and development: improving ecological validity nienke van atteveldtabc, marlieke t.r. van kesterenabc, barbara braamsabc, lydia krabbendamabc avrije universiteit amsterdam, the netherlands binstitute learn!, vrije universiteit amsterdam, the netherlands c institute for brain and behavior amsterdam (ibba), the netherlands article received 13 may 2018/ revised 22 october/ accepted 28 november/ available online 19 december abstract modern neuroscience research, including neuroimaging techniques such as functional magnetic resonance imaging (fmri), has provided valuable insights that advanced our understanding of brain development and learning processes significantly. however, there is a lively discussion about whether and how these insights can be meaningful to the educational practice. one of the main challenges is the low ecological validity of neuroimaging studies, making it hard to translate neuroimaging findings to real-life learning situations. here, we describe four approaches that increase the ecological validity of neuroimaging experiments: using more naturalistic stimuli and tasks, moving the research to more naturalistic settings by using portable neuroimaging devices, combining tightly controlled lab-based neuroimaging measurements with real-life variables and follow-up field studies, and including stakeholders from the practice at all stages of the research. we illustrate these approaches with examples and explain how these directions of research optimize the benefits of neuroimaging techniques to study learning and development. this paper provides a frontline overview of methodological approaches that can be used for future neuroimaging studies to increase their ecological validity and thereby their relevance and applicability to the learning practice. keywords: ecological validity; neuroimaging, educational neuroscience; learning; brain development info corresponding author mail n.m.van.atteveldt@vu.nl doi https://doi.org/10.14786/flr.v6i3.366 1. introduction modern neuroscience research, including non-invasive neuroimaging techniques such as functional magnetic resonance imaging (fmri) or electro-encephalography (eeg), provides exciting new possibilities for investigating the living human brain. for example, cognitive neuroimaging research has yielded many important insights into the neurobiological basis of learning and development (e.g. casey, tottenham, liston, & durston, 2005; crone & elzinga, 2015). this has boosted interdisciplinary research endeavors exploring the educational relevance of these insights, by crossing boundaries between cognitive neuroscience, educational science and educational practice, giving rise to a new research field termed either educational neuroscience or mind, brain and education (ansari, coch, & de smedt, 2011). such endeavors are challenging, as applying insights gained using neuroimaging techniques is accompanied by several limitations and uncertainties (schleim & roiser, 2009; van atteveldt, van aalderen-smeets, jacobi, & ruigrok, 2014). a major limitation is the low ecological validity of the performed experiments. two important factors compromising the ecological validity of neuroimaging studies are 1) the highly controlled stimuli and tasks used and 2) the artificial and isolated environment in which the studies take place (the setting). during a typical neuroimaging experiment, participants perform a cognitive task aimed at isolating one specific skill or cognitive process, typically consisting of highly controlled, simplified stimuli. because of the noisy nature of neuroimaging signals, many repetitions of these simplistic stimuli are needed in the cognitive task. the setting of neuroimaging studies is far removed from a naturalistic learning situation. an mri lab for example, is a sterile, artificial environment that can be intimidating especially for children. participants lie on their back in a narrow scanner bore, with their heads fixated and are alone when they are being scanned. because of the many repetitions of stimuli needed to obtain reliable signals, fmri sessions are long and can be exhausting. this setting is incomparable to daily-life learning environments such as a classroom or the workplace where social dynamics and the learning situation are much more complex. this discrepancy introduces challenges in translating basic insights generated during neuroimaging lab experiments to real-life brain function (kingstone, smilek & eastwood, 2008). fortunately, recent advancements increasingly enable approaches to overcome the constraints that cause the low ecological validity (matusz, dikker, huth & perrodin, 2018). the current paper focuses on how to overcome the challenges yielded by the artificial and highly controlled lab setting of standard neuroimaging research, by outlining methodological approaches to improve the ecological validity of stimuli, tasks and the measurement context (setting). limitations and possibilities of neuroimaging research at a more technical or physiological level are discussed extensively elsewhere (logothetis, 2008; poldrack & farah, 2015; schleim & roiser, 2009), and are included here only when they put constraints on ecological validity. there is currently a lively debate ongoing on the merits of educational neuroscience (bowers 2016, howard-jones et al 2016, gabrieli 2016). here, we specifically address the challenges that arise from low ecological validity of neuroimaging studies rather than the challenges and added value of educational neuroscience in general. the working definition of educational neuroscience used in the current paper is: cognitive neuroimaging studies into learning and development, bringing indirectly and directly useful insights to improve learning and teaching in the practice. in the following, we will begin with providing a brief overview of the most widely used neuroimaging techniques and their strengths and weaknesses. next, we will illustrate what neuroimaging of learning and development can bring to the educational practice, either indirectly by informing teachers and influencing their beliefs and attitudes, or more directly by actually using a brain measure for educational purposes. it should be noted that some factors limiting ecological validity apply to laboratory experiments more generally, i.e., apply also to behavioral experiments rather than specifically to neuroimaging, such as the controlled stimulus and task designs to tap into isolated cognitive processes. however, there are several limitations that are either more extreme or unique to neuroimaging experiments. these more general versus specific limitations are explained in more detail in section 2.4 and in figure 1. next, we will discuss different approaches one can take to improve the ecological validity of neuroimaging, starting with ways to make stimuli and tasks more naturalistic in section 3, such as using videos or virtual reality inside the (f)mri scanner or staging real-life social interactions during lab experiments. in section 4 we focus on the setting, by discussing developments in portable neuroimaging devices (e.g. eeg, or functional near-infrared spectroscopy or fnirs), which enable measuring brain activity in real classrooms or other real-life learning settings. the third approach we discuss is to combine tightly controlled research in the neuroimaging lab with real-life variables (e.g., grades, risk behaviour), and with follow-up field studies (section 5). this approach exploits the technical strengths and data quality of e.g. fmri, but still relating it to real-life learning or behavior within the same study. the final approach we discuss is to include stakeholders from the practice at all stages of the research. this is important to further improve the connection between neuroimaging studies and real-life learning, because teachers have the day-to-day experience with the chaotic and irregular teaching practice that researchers lack, and to ensure realistic expectations on what neuroimaging can and cannot offer as well as to prevent neuromyths (howard-jones, 2014). 2. neuroimaging of learning and development: insights for education studying the human brain in neuroimaging labs has gained many neurobiological insights that changed the way we think about development and learning. such insights can be useful to education either indirectly or more directly. an example of indirect use is that a better understanding of brain plasticity may influence teacher’s beliefs about their student’s learning potential (dubinsky, roehrig, & varma, 2013). in section 2.2, we will highlight two influential examples of such more indirectly useful insights: adolescent brain development and the neurobiology of memory. whether neuroimaging measures or insights can also be used more directly is a matter of debate (bowers, 2016; gabrieli, 2016; howard-jones et al., 2016). as mentioned above, the scope of the current paper is not to discuss the added value of educational neuroscience in general, but to illustrate how improving ecological validity of neuroimaging studies may strengthen its educational relevance. in section 2.3, we will illustrate one potentially promising example of the more direct use of neuroimaging: using a brain measure for its predictive value, also termed neuro-prediction. if adding a brain measure predicts a certain learning or intervention outcome better than behavioral tests alone, it can provide valuable information. 2.1. basic overview neuroimaging techniques the term neuroimaging is used for modern techniques that are able to measure signals from the brains of living (human) participants. a distinction can be made between anatomical and functional techniques. anatomical techniques make static images of structural details of the brain; anatomical details such as volume or thickness of different tissues such as gray vs. white matter in the case of structural mri, water diffusion orientation indicating myelinated axons and their integrity in the case of diffusion tensor imaging (dti). although insights from anatomical techniques can be relevant for educational issues, the link to learning and behavior is always indirect, as they only yield static images and no information about dynamics of brain activity. functional techniques do capture such dynamics, by measuring time-series of brain activity while the participant is doing a task. this task is typically designed to mimic (one specific aspect of) daily-life behavior, but as we are restricted to controlled, simplistic stimuli and responses by button presses allowing only one or multiple fingers to move, often it is only loosely related to naturalistic behavior. the most commonly used functional neuroimaging techniques are fmri and eeg. fmri detects changes in blood oxygenation levels as a measure of neural activation, eeg measures electrical signals emitted by the brain. in table 1, an overview of the practical set-up, nature of measurement, and the strengths and weaknesses of both techniques is provided. other functional techniques are magnetoencephalography (meg), which is very similar to eeg but instead measures magnetic fields produced by the electrical activity of the brain, and functional near-infrared spectroscopy (fnirs) which is more similar to fmri as it also detects changes in blood oxygen level, but using light in the near infrared spectrum instead of radio frequency magnetic pulses. table 1 overview of the practical set-up, nature of the signal, advantages and disadvantages of the most widely used functional neuroimaging techniques fmri and eeg 2.2. indirectly influential insights from the neuroimaging lab one of the major insights in brain development provided by neuroimaging is the protracted development of the adolescent brain. it has long been thought that the brain was mostly developed around adolescence. although there were some indications of protracted development in certain brain regions in the late 1970’s (huttenlocher, 1979), it wasn’t until the rise of mri that adolescents’ brain development could be studied on a large scale in healthy, living adolescents. moreover, as longitudinal use is possible, mri enabled studying adolescents while they go through different stages of development, thereby revealing developmental trajectories rather than only cross-sectional age comparisons (horga, kaur, & peterson, 2014). a couple of landmark studies in the late ’90 and early 2000’s have shown protracted development of certain brain regions during adolescence (giedd et al., 1999; lenroot et al., 2007; paus et al., 1999; sowell, thompson, holmes, jernigan, & toga, 1999; sowell et al., 2004; thompson et al., 2000). some regions, such as temporal regions related to social cognition, show developmental changes in grey matter up to at least age 25 (gogtay et al., 2004). white matter connections between brain areas show non-linear developmental increases in adolescence (bava et al., 2010; giorgio et al., 2010; lebel & beaulieu, 2011). the use of mri techniques opened up the possibility of relating brain developmental to behavior. it has been shown that integrity of white matter pathways between the prefrontal cortex and striatum are positively related to increased patience in a delay discounting task (peper et al., 2013; van den bos, rodriguez, schweitzer, & mcclure, 2014). the integrity of this pathway increases over adolescence and is negatively related to impulsivity in delay discounting behavior (van den bos et al., 2014). together, these studies show that behavioral changes in adolescence are related to developmental changes in neural architecture, and occur until a much later age than previously thought. this knowledge has important implications, especially for secondary and higher education, as brains of adolescents may work more differently from adults than we used to think. over the last couple of years, insights about adolescent brain development has become common knowledge among parents and educators (choudhury, mckinney, & merten, 2012). a better understanding of protracted adolescent brain development may stimulate teachers to use certain teaching practices such as providing more guidance with planning abilities (dekker & jolles, 2015) or simply to be more patient and understanding (hook & farah, 2013). it should be noted that it remains to be established to which degree brain development is a consequence of natural development and what the influence of experience is. this causal direction problem has consequences for, for instance, effectiveness of interventions. some studies have aimed to disentangle the effects of natural development and experience. one example is a school cut-off design in which development in children who attend school is compared to children of the same age who did not yet attend school. children who attended school showed higher improvement in executive function and the rate of improvement was positively correlated with increased activity in the right posterior parietal cortex, an area important for sustained attention (brod, bunge, & shing, 2017), indicating that brain development is at least partly driven by experience. another domain in which neuroimaging research has significantly increased our understanding of learning processes is that of memory research. research into the neurobiology of memory encompasses many different types of memory, such as working memory (i.e., a memory that is shortly kept in mind, such as when reiterating a phone number; baddeley, 2003) and long-term memories (i.e., memories that are stored for a longer period to be retrieved again at a later point in time; jeneson & squire, 2012). with regard to working memory, understanding how we keep information in mind and assessing possible limits of working memory ability can be highly valuable for education (prince & gifford, 2016) because working memory ability is presumed to be malleable (klingberg, 2010) and keeping track of information load can help teachers to rightly dose knowledge. however, while working memory is indeed often found to be trainable, generalization of this enhancement to other cognitive abilities is highly debated (melby-lervåg, redick, & hulme, 2016), and therefore also transfer to learning in the classroom is uncertain. with regard to long-term memory, neuroimaging research has added interesting new evidence supporting the role of prior knowledge in learning. prior knowledge is known to strongly benefit new learning (bartlett, 1932) and as such it has been suggested to facilitate mnemonic processing in the brain (van kesteren, ruiter, fernandez, & henson, 2012). the interplay between the medial temporal lobe (mtl), with the hippocampus at its core, and the medial prefrontal cortex (mpfc) is suggested to be crucial in this process. the mpfc is suggested to help integrate new memories that fit with prior knowledge directly into a prior knowledge network (or a schema), bypassing the hippocampus/mtl system, which is proposed to link to episodic details of memories (teyler & discenna, 1985). this former, integrative process is proposed to yield more semantic, less detailed memories directly after learning. while episodically detailed memories might be more intense, these integrated memories are suggested to be more durable (dudai, karni, & born, 2015). moreover, they could strengthen an existing prior knowledge network such that future learning is also facilitated, which is highly important in educational settings. when understanding more about the interplay between these memory systems, teachers and students can tailor learning by optimally using their prior knowledge networks in order to either store more detailed or more integrated memories. 2.3. how to use lab-generated insights more directly: the example of neuro-prediction neuroimaging can be useful to the educational practice more directly if individual brain measures can be used to predict learning outcomes (peters & crone, 2017) or intervention effectiveness (gabrieli, 2016). several studies demonstrate that including a neural measure to behavioral tests, this improved the prediction of an educational or intervention outcome beyond using behavioral tests alone. for example, hoeft and colleagues showed that adding fmri and dti measures to behavioral tests improved the prediction of how well dyslexic readers compensated for reading difficulty over a period of 2,5 years (hoeft et al., 2011). this may help identifying which children would profit most from extra intervention. another example of the predictive value of neuroimaging is a longitudinal dti study linking white matter development and reading skill development (yeatman, dougherty, ben-shachar, & wandell, 2012). this study showed that reading skill is predicted by how mature the relevant fiber tracts (the left arcuate and left inferior longitudinal fasciculus) are initially. interestingly, the more mature the connections at the first measurement, the lower the reading scores. this study points out that when is the optimal timing to start learning to read may differ between children. studies that improve prediction of an intervention or educational outcome when a neuroimaging measure is added to behavioral tests are promising, but one extra step is still needed, similar to what has been proposed as much needed extra research to test the value of using neuroimaging markers in psychiatry (mayberg, 2014). this step entails performing randomized controlled trials (rct’s) to compare the effectiveness of intervention when the decision between intervention types is based on behavioral tests alone, compared to when the neuroimaging measure is additionally used for this decision. 2.4. general versus specific limitations to ecological validity (stimuli, task, setting) the examples above highlight the potential educational relevance of neuroimaging research into learning and development. as mentioned before, actual use of such insights is limited by the low ecological validity of neuroimaging experiments. most experimental cognitive research, including behavioral studies, has the limitation that stimuli and tasks are tightly controlled and therefore not naturalistic. however, several constraints are either more extreme for, or unique to neuroimaging studies. when measuring signals from the brain there are extra confounding factors to take into account. one of the most pressing issues in neuroimaging is noise in the data due to head movement, eye or muscle movement, physiological measures like heart rate and breathing, and lapses in attention or fatigue due to lengthy experiments in confined spaces (logothetis, 2008). these noise sources have implications for the stimulus and task design: they can be accounted for by adding many (preferably > 20 per condition) highly controlling experimental trials. moreover, experimental trials should optimally be short (i.e., in the range of a few seconds) to really capture brain signals related to one specific process (amaro & barker, 2006; ruiter, van kesteren, & fernandez, 2012). stimuli are best presented visually (at least for fmri because of auditory scanner noise) and should be kept simple and easy to process to extract one isolated cognitive process. for example, in a memory experiment it is better to show participants simple pictures than lengthy sentences because sentences require longer processing and eye movements, activating other neural processes. when accounting for all these factors, there is mostly only a very artificial and simplified version of real-life situations left. in addition, in fmri studies, since it is a costly technique both financially and in time investment for participants and researchers, participants are often scanned one single time. it is currently largely unclear how variable neural measures are and how this variability relates to variability in behavior. also, the results provided by neuroimaging are less straightforward to interpret. hit rates or reaction times resulting from behavioral studies are relatively clear measures, but it is not as clear how the indirect hemodynamic changes measured with fmri, or the voltage change measured at a limited number of scalp locations with eeg, should be interpreted. figure 1. general versus more specific limitations in ecological validity in lab-based experiments in addition to the specific constraints to stimulus and task design, the setting of neuroimaging studies also poses extra challenges for ecological validity. for example, participants have to either lie in a tight fmri-scanner with their heads fixated, or sit behind a computer in a shielded room with an eeg-cap on their heads, while moving as little as possible. the big mri machine and narrow bore can be intimidating and in rare cases even lead to claustrophobic experiences. in developmental mri, there are several approaches to minimize such risks and make children comfortable, such as practicing in a mock scanner (raschle et al., 2009). in sum, the setting leads neuroimaging experiments to be very different from what happens in the educational practice, where free viewing and moving influence learning and attention, and social interactions are constantly in play. meg and fnirs have similar restrictions, although mobility of the equipment varies (mehta & parasaruman, 2013), for example, fnirs even allows participants to walk around (pinti et al, 2018). in figure 1, an overview is provided of the challenges to ecological validity that apply to experimental cognitive research more generally, and which apply more severely or uniquely to neuroimaging. 3. approach 1: improving the ecological validity with more naturalistic stimuli and tasks fortunately, there are ways to circumvent or minimize the limiting factors outlined above. here, we highlight a few of these approaches. for example, one can use stimuli that are often learned in relatively controlled educational settings, such as in vocabulary learning or other simple associations and investigate underlying neural processes that relate to language learning and forming long-term memories (goswami, 2006). another approach is to approximate educational stimuli in an imaging experiment (van kesteren, rijpkema, ruiter, morris, & fernandez, 2014) or run an additional behavioral experiment in which more educationally relevant stimuli are used in a more educationally relevant setting (van kesteren et al., 2018). this can then be compared to the behavioral results of a more strictly controlled neuroimaging study. when behavioral results are comparable for these experiments, a direct relationship between the neural processes and behavioral outcomes in the educationally relevant setup can be inferred. furthermore, more ecologically valid stimuli can be used and analyzed with novel, sophisticated analysis methods, such as video stimuli for memory processes (van kesteren, fernandez, norris, & hermans, 2010; aly et al., 2018; chen et al., 2016) or games for insight processes (milivojevic, vicente-grabovetsky, & doeller, 2015). rather than strongly controlling and isolating stimulus trials and contrasting conditions that only differ in one respect, these analyses use network connectivity measures or inter-subject correlations during free viewing of natural stimuli (hasson & honey, 2012), providing the possibility to consider processes happening during longer stretches of time. this approach thus allows to ask different questions that are e.g. related to inter-subject variability or neural processes that require a more extended time period, and better resembles everyday learning than when simple stimuli are presented in isolation. in addition, new analysis techniques also include machine learning algorithms to decode brain activation patterns automatically with high accuracy (poldrack & farah, 2015). this offers interesting opportunities for understanding how the brain represents real-life information. for example, a recent study successfully decoded fmri response patterns to tell which real-life sounds participants were listening to (santoro et al., 2017). developments in virtual reality (vr) technology for neuroimaging experiments also hold big promises for improving ecological validity of stimuli. as vr enables simulating naturalistic events, brain activity can be measured during learning situations that can closely resemble real-life situations, e.g. including social interactions. vr technology therefore combines a high degree of control with a high degree of ecological validity (bohil, alicea, & biocca, 2011). to use individual neural measures as a predictor (see section 2.3), it is crucial to have reliable neural measures on the individual level. however, the most used current neuroimaging techniques use averaged group data to show patterns of neural activation. recent studies have started to investigate neural measures in individual participants thereby increasing specificity and detail of these measures. since the signal-to-noise ratio in neuroimaging data is low, the same individual needs to be scanned multiple times to acquire sufficient data to characterize individual patterns of activation. it has been shown that scanning the same individual multiple times yields reliable data for this individual (braga & buckner, 2017; gordon et al., 2017; poldrack, 2017). future work should further parse out how these individual neural patterns are related to behavior. more naturalistic stimuli and social interactions are also of high importance for studying social processes. human interaction in real life is complex and flexible. however, most research investigating social cognition has used static stimuli such as pictures of faces or descriptions of persons. these limitations were mainly due to technological limitations. due to recent advances in neuroimaging technology and analysis it is now possible to test more complex social interactions in the mri environment. one of these advances is using naturalistic videos to test which brain areas are selectively responsive to social stimuli (iacoboni et al., 2004; wagner, kelley, haxby, & heatherton, 2016), as also described above. other examples are real-time tracking of the updating of social reputation as a function of social interaction (tavares et al., 2015) and staging real social interactions while being in the scanner, e.g. by having the participant in the scanner interact with persons in his or her social network (fareri, niznikiewicz, lee, & delgado, 2012), or generating a real social touch experience by a person present in the scanner room (gazzola et al., 2012). real social interactions can now be investigated by measuring neural activation of multiple participants simultaneously while they are interacting. this so called ‘hyperscanning’ has opened up new avenues of research and will provide more insight into the complex nature of social interaction (see babiloni & astolfi, 2014, for a review). hyperscanning experiments can be situated in the lab, and developments in portable neuroimaging technology now also enables recording brain activity from multiple participants in a real-life learning setting such as a classroom (bhattacharya, 2017), as will be showcased in section 4. 4. approach 2: improving the ecological validity of the setting by using portable devices as mentioned before, a major limitation of neuroimaging techniques is their immobility. different techniques differ in their degree of immobility, and eeg and fnirs are relatively most mobile (mehta & parasuraman, 2013). this makes it possible to record eeg or fnirs activity during real-life learning settings, such as university lectures. a recent study measured eeg while university students had to respond to visual target stimuli presented during real lectures (ko, komarov, hairston, jung, & lin, 2017), and found a correlation between a student’s sustained attention (i.e., target detection performance) and their eeg spectra. although possible, the setup of regular (stationary) eeg equipment outside the lab is not straightforward. recent developments in portable brain technology have started a new direction of increasing the ecological validity of neuroimaging, as portable equipment drastically facilitates taking the research out of the lab and into working classrooms and other real-life settings. furthermore, portable technologies enable research in infants and young children, and in areas of the world where access to advanced research labs is limited. the main techniques in which portable variants are available are eeg and fnirs, but recently, a mobile version of meg has also been developed (boto et al., 2018). the increasing availability, quality and user-friendly nature of portable devices such as eeg improves the possibilities of measuring brain activity in real-life learning settings enormously. for example, it allows recording brain activity of multiple students in a classroom simultaneously (‘hyperscanning’ as described in section 3), while they engage in natural learning activities and social interactions. in a pioneering study, dikker and colleagues simultaneously measured eeg activity from 12 high school students and their teacher, while they took part in different classroom activities (dikker et al., 2017). the authors were able to establish a link between brain-to-brain synchrony and classroom engagement (dikker et al., 2017), and in a later study, also with learning outcomes (bevilacqua et al., 2018). although the data quality of portable devices does not reach the same level as when using lab equipment, for example due to fewer electrodes and movement artifacts, this direction of portable neuroscience has interesting potential, especially when combined with lab-based neuroimaging to generate complementary high-quality data. interesting future directions are investigating the role of interbrain synchrony during real-life joint attention or collaborative learning (bhattacharya, 2017). in addition to eeg, portable fnirs also opens many opportunities for studying brain function in freely moving participants. an advantage of fnirs is that it is relatively insensitive to motion and muscular artifacts (balardin et al., 2017), making this technique particularly suitable for research in infants and young children, and allows free movement of participants including walking (pinti et al., 2018). 5. approach 3: linking lab-based neuroimaging to real-life behavior 5.1. relating neuroimaging measures to separately measured real-life variables one way to relate neural measures to real-life behavior is by measuring both separately and subsequently linking them. for example, links between the educational practice and brain processes can be made through the relationship between cognition-related brain activity and course grades (van kesteren et al., 2014). the neural measures that link to real-life behavior such as school grades can vary in similarity to the behavior of interest and can range from neural activity during performance of highly controlled tasks, focused on one process, to more complex and ecologically valid tasks. an example of a highly controlled task is a coin flipping game in which participants were instructed to choose heads or tails, if the computer randomly selected the same side of the coin, participants won money (braams et al, 2015). this game results in strong nucleus accumbens activation when winning compared to losing money. nucleus accumbens activation on the task shows a positive correlation with alcohol consumption, a form of real-life risk-taking behavior. however, the real-life processes that give rise to risk-taking behavior are more complex than sensitivity to reward. a more ecologically valid task is the balloon analogue risk task (bart). in this task participants can press a button to inflate a balloon. for each puff of air they receive a small reward. participants can choose to stop and receive the money they earned on that round, or the balloon pops and they lose the money on that round. behavior on the bart shows relationships with real-life behavior such as alcohol use, smoking and drug use (lejuez, aklin, zvolensky, & pedulla, 2003). groups of adolescents using alcohol, marijuana or both show differential recruitment of neural regions while playing the bart compared to adolescents who abstain from using alcohol or marijuana (claus et al., 2017). these studies show that current neuroimaging measures are related to real-life behavior. however, the links are often weak and it is unclear how much variance in real-life behavior is explained exactly by the neural measures. to improve the match between what we measure with neuroimaging and the real-life measure, we can either move the neuroimaging paradigms towards higher ecological validity (e.g. stimulus material as described in section 3, or tasks as described above), or we can improve the real-life measure. in many neuroimaging studies, real-life behavior is measured sparingly, by use of questionnaires or self-report behavior. a much more sensitive approach is the experience sampling method (esm), a structured diary technique allowing for the systematic and repeated sampling of an individual’s mood, behaviour and experiences over time in relation to context (csikszentmihalyi & larson, 1987). esm measures behavior in the real world without the biases associated with reflection and retrospection. these momentary experiences may indeed have different associations with neural patterns of activation, compared to retrospective reports. during scanner-based social rejection, activation in areas associated with negative affect and pain processing was associated with momentary experiences of social distress in daily life social interactions, whereas activation in areas involved in memory and self-referential processing was associated with retrospective reports of social distress (eisenberger, gable, & lieberman, 2007). using structured diary techniques, studies have also revealed longitudinal relationships between social experiences and brain activation. spending more time with friends during adolescence was associated with reduced neural sensitivity to peer rejection two years later (masten, telzer, fuligni, lieberman, & eisenberger, 2010); and low peer support assessed across two years prior to scanning was associated with increased neural sensitivity to peer conflict (telzer, fuligni, lieberman, miernicki, & galván, 2014). another approach to a more detailed assessment of real-life behaviour is the analysis of social networks. social network data are based on reports of groups of individuals about their mutual connections. this provides a better picture of the complexity of social groups than self-reported connections. a recent study combining structural mri with social network analysis found that network size based on the social ties identified by a subject's social connections (indegree network), significantly correlated with a larger regional volume of the orbitofrontal cortex, dorsomedial prefrontal cortex and lingual gyrus. in contrast, no associations between brain structural volume and an individual’s self-reported ties (outdegree network) were found (kwak, joo, youm, & chey, 2018). using logged friendship data from facebook, another study found that connectivity in brain regions associated with pain and mentalizing during social exclusion was modulated by the density of individuals’ social networks (schmälzle et al., 2017). 5.2. using field studies to test real-life validity of hypotheses generated by lab research the results of lab-based neuroimaging studies can be used to generate hypotheses regarding more practical issues, which can be subsequently tested using field research, e.g. by conducting rct’s and intervention effect studies. to come back to the example of the insights into the protracted brain development during adolescence (section 2.2), we mentioned that this insight may have interesting implications for education. however, it is not possible to know what these implications are exactly, based on research in the artificial and highly controlled lab setting. as the brain regions that show protracted development are linked to executive functions and self-regulation, a working hypothesis could be that adolescents may need guidance in self-regulatory behavior to do well at school (i.e., resist distractions, plan school work ahead). this hypothesis can be tested in real life, by designing an intervention based on the lab research and investigating its effectiveness in a field rct. other interventions targeting children’s abilities to self-regulate their behavior and emotions at school have shown positive results (eisenberg, valiente, & eggum, 2010), and these can be fine-tuned by neuroimaging insights. another example is that of memory consolidation and sleep. it has been shown, both behaviorally and using neuroimaging, that sleep is important for memory consolidation (mcgaugh, 2000), a process by which memories are suggested to be strengthened to protect them from forgetting. more specifically, memories are suggested to become semanticized through consolidation (dudai et al., 2015; lewis & durrant, 2011). this means they lose episodic details and integrate with related memories to form a factual knowledge network. however, also here, it is not straightforward what lab findings on the role of sleep in memory consolidation mean for real-life learning. as also raised by sigman and colleagues (sigman, peña, goldin, & ribeiro, 2014), different questions about how to use these insights in the practice should be tested in field studies, such as whether introducing naps at school after learning something improves learning outcomes compared to using this time for extra study. this same research group recently conducted such a field study, demonstrating that post-class naps in school indeed enhance information retention (cabral et al., 2018). testing the potential educational value of lab-based neuroimaging research by additional field studies is as yet uncommon, and described as ‘a nearly unexploited research frontier’ (sigman et al., 2014). potential reasons for the lack of integrated lab-field studies so far are the extra time investment required and the lack of best practices to integrate methodology and standards from both approaches (howard-jones, 2009). however, to test what lab-generated insights into learning and development really imply for the educational practice, future research should make more efforts to integrate lab and field research. 6. approach 4: include stakeholders from the practice at all stages of the research in addition to making stimuli, task and setting more naturalistic, the ecological validity of neuroimaging research on learning and development may further be improved by early and active inclusion of teachers and other stakeholders in the research. for example, involving teacher’s experience and perspectives when designing research may help creating more realistic research designs and selecting stimulus materials and contexts that more closely resemble real-life learning materials and circumstances. another important aim of improving collaboration between researchers and teachers is to prevent misconceptions and unrealistic expectations (howard-jones, 2014), by improving translation accuracy (shonkoff & bales, 2011). the call for more systematic interaction between cognitive neuroscientists and stakeholders from the practice is by no means new. examples of successful teacher-researcher collaborations are abundant (mccandliss, kalchman & bryant, 2003) and the dialogue between teachers and researchers has been put forward as one of the most important building blocks for educational neuroscience research (ansari, coch & de smedt, 2011). one way to increase interactions between teachers and researchers is including more neuroscience and research methodology in teacher training curricula (ansari, coch & de smedt, 2011). as mentioned before, it can be very beneficial to involve teachers and other stakeholders also in the scientific research itself. a systematic approach to do this is the responsible research and innovation (rri) framework (owen, macnaghten & stilgoe, 2012), which aims to align research with values and needs of society. the rri framework has different dimensions, such as active and early inclusion of diverse stakeholders during the entire research process, anticipation of possible future impacts, reflecting on the diverse underlying values and purposes and on whether a certain development or application is desirable, and being responsive by changing the research adaptively. stilgoe and colleagues (2013) offer techniques to implement each of these dimensions into research. implementing this framework to educational neuroscience could be a promising direction, as was recently also suggested for research on electrical brain stimulation for educational purposes (schuijer, de jong, kupper, & van atteveldt, 2017). 7. conclusions one of the main challenges causing the limited applicability of neuroimaging research on learning and development to the educational practice is the low ecological validity of lab settings, making it hard to translate neuroimaging findings to real-life learning situations. in this paper, we described four methodological approaches that increase ecological validity of neuroimaging experiments: using more naturalistic stimuli and tasks, doing neuroimaging research in more naturalistic settings using portable devices, combining tightly controlled lab-based neuroimaging with real-life variables and follow-up field studies, and including stakeholders from the practice at all stages of the research. these approaches can be taken as guidelines for future neuroimaging studies to improve their ecological validity and to maximize their educational value. stepping away from a high level of control is becoming increasingly possible because of improved sensitivity and mobility of neuroimaging techniques, incremental insights in designs and tasks and advancements in data-analysis tools, to improve the connection between behaviors during neuroimaging and natural, real-life activities. in addition to leading to more meaningful results, this progress may also enable addressing a wider array of research questions to which cognitive neuroscience could add relevant insights. it has been argued that neuroscientific data apply mostly to low-level behaviors such as single cognitive constructs (willingham & loyd, 2007) but now, for example by measuring brain activity of groups of students using portable eeg devices, this scope may be widened to include more complex behaviors, such as at the classroom level. keypoints keypoint 1: a main challenge of applying neuroimaging research to the educational practice is the low ecological validity of lab-based neuroimaging research; keypoint 2: the first approach to increase ecological validity of neuroimaging is by using more naturalistic stimuli and tasks; keypoint 3: the second approach is to increase ecological validity of the setting, by using portable neuroimaging devices allowing measuring brain activity in real-life learning situations; keypoint 4: the third approach to increase ecological validity of neuroimaging is to link lab-based neuroimaging measures to separately measured real-life behaviour; keypoint 5: the fourth approach is involving stakeholders (teachers, students) in designing and interpreting neuroimaging research to increase the educational value of neuroimaging. acknowledgments n.v.a. and l.k. were supported by grants from the european research council (erc, starting grant #716736 to n.v.a. and consolidator grant #648082 to l.k.). m.v.k. was supported by a marie curie individual fellowship of the eu horizon2020 framework program (grant #704506). b.b. was supported by a veni grant from the netherlands organization for scientific research (nwo) (nwo veni #451-7-008). references references aly, m., chen, j., turk-browne, n. b., & hasson, u. (2018). learning naturalistic temporal structure in the posterior medial network. journal of cognitive neuroscience 30(9): 1345-1365. amaro, e., jr., & barker, g. j. (2006). study design in fmri: basic principles. brain and cognition, 60(3), 220-232. doi:10.1016/j.bandc.2005.11.009 ansari, d., coch, d., & de smedt, b. (2011). connecting education and cognitive neuroscience: where will the journey take us? educational philosophy and theory, 43(1), 37-42. babiloni, f., & astolfi, l. (2014). social neuroscience and hyperscanning techniques: past, present and future. neuroscience & biobehavioral reviews, 44, 76-93. baddeley, a. (2003). working memory: looking back and looking forward. nat rev neurosci, 4(10), 829-839. doi:10.1038/nrn1201 balardin, j. b., zimeo morais, g. a., furucho, r. a., trambaiolli, l., vanzella, p., biazoli, c., & sato, j. r. (2017). imaging brain function with functional near-infrared spectroscopy in unconstrained environments. frontiers in human neuroscience, 11(258). doi:10.3389/fnhum.2017.00258 bartlett, f. c. (1932). remembering: a study in experimental and social psychology. cambridge, [england]: university press. bava, s., thayer, r., jacobus, j., ward, m., jernigan, t. l., & tapert, s. f. (2010). longitudinal characterization of white matter maturation during adolescence. brain res, 1327, 38-46. doi:10.1016/j.brainres.2010.02.066 bevilacqua, d., davidesco, i., wan, l., oostrik, m., chaloner, k., rowland, j., ding, m., poeppel, d., & dikker, s. (2018). brain-to-brain synchrony and learning outcomes vary by student–teacher dynamics: evidence from a real-world classroom electroencephalography study. journal of cognitive neuroscience: 1-11. bhattacharya, j. (2017). cognitive neuroscience: synchronizing brains in the classroom. current biology, 27(9), r346-r348. bohil, c. j., alicea, b., & biocca, f. a. (2011). virtual reality in neuroscience research and therapy. nature reviews neuroscience, 12, 752. doi:10.1038/nrn3122 boto, e., holmes, n., leggett, j., roberts, g., shah, v., meyer, s. s., . . . brookes, m. j. (2018). moving magnetoencephalography towards real-world applications with a wearable system. nature. doi:10.1038/nature26147 bowers, j. s. (2016). the practical and principled problems with educational neuroscience. psychological review, 123(5), 600. braga, r. m., & buckner, r. l. (2017). parallel interdigitated distributed networks within the individual estimated by intrinsic functional connectivity. neuron, 95(2), 457-471. e455. brod, g., bunge, s. a., & shing, y. l. (2017). does one year of schooling improve children’s cognitive control and alter associated brain activation?. psychological science, 28(7), 967-978. cabral, t., mota, n. b., fraga, l., copelli, m., mcdaniel, m. a., & ribeiro, s. (2018). post-class naps boost declarative learning in a naturalistic school setting. npj science of learning 3(1): 14. casey, b. j., tottenham, n., liston, c., & durston, s. (2005). imaging the developing brain: what have we learned about cognitive development? trends in cognitive sciences, 9(3), 104-110. chen, j., honey, c. j., simony, e., arcaro, m. j., norman, k. a., & hasson, u. (2016). accessing real-life episodic information from minutes versus hours earlier modulates hippocampal and high-order cortical dynamics. cerebral cortex 26(8): 3428-3441 choudhury, s., mckinney, k. a., & merten, m. (2012). rebelling against the brain: public engagement with the ‘neurological adolescent’. social science & medicine, 74(4), 565-573. claus, e. d., feldstein ewing, s. w., magnan, r. e., montanaro, e., hutchison, k. e., & bryan, a. d. (2017). neural mechanisms of risky decision making in adolescents reporting frequent alcohol and/or marijuana use. brain imaging behav. doi:10.1007/s11682-017-9723-x crone, e. a., & elzinga, b. m. (2015). changing brains: how longitudinal functional magnetic resonance imaging studies can inform us about cognitive and social‐affective growth trajectories. wiley interdisciplinary reviews: cognitive science, 6(1), 53-63. csikszentmihalyi, m., & larson, r. (1987). validity and reliability of the experience-sampling method. the journal of nervous and mental disease, 175(9), 526-536. dekker, s., & jolles, j. (2015). teaching about “brain and learning” in high school biology classes: effects on teachers’ knowledge and students’ theory of intelligence. frontiers in psychology, 6, 1848. doi:doi: 10.3389/fpsyg.2015.01848 dikker, s., wan, l., davidesco, i., kaggen, l., oostrik, m., mcclintock, j., . . . ding, m. (2017). brain-to-brain synchrony tracks real-world dynamic group interactions in the classroom. current biology, 27(9), 1375-1380. dubinsky, j. m., roehrig, g., & varma, s. (2013). infusing neuroscience into teacher professional development. educational researcher, 42(6), 317-329. dudai, y., karni, a., & born, j. (2015). the consolidation and transformation of memory. neuron, 88(1), 20-32. doi:10.1016/j.neuron.2015.09.004 eisenberg, n., valiente, c., & eggum, n. d. (2010). self-regulation and school readiness. early education and development, 21(5), 681-698. eisenberger, n. i., gable, s. l., & lieberman, m. d. (2007). functional magnetic resonance imaging responses relate to differences in real-world social experience. emotion, 7(4), 745. fareri, d. s., niznikiewicz, m. a., lee, v. k., & delgado, m. r. (2012). social network modulation of reward-related signals. journal of neuroscience, 32(26), 9045-9052. gabrieli, j. d. (2016). the promise of educational neuroscience: comment on bowers (2016). psychological review, 123(5):613-9. doi: 10.1037/rev0000034. gazzola, v., spezio, m. l., etzel, j. a., castelli, f., adolphs, r., & keysers, c. (2012). primary somatosensory cortex discriminates affective significance in social touch. proceedings of the national academy of sciences, 109(25), e1657-e1666. giedd, j. n., blumenthal, j., jeffries, n. o., castellanos, f. x., liu, h., zijdenbos, a., . . . rapoport, j. l. (1999). brain development during childhood and adolescence: a longitudinal mri study. nat neurosci, 2(10), 861-863. doi:10.1038/13158 giorgio, a., watkins, k. e., chadwick, m., james, s., winmill, l., douaud, g., de stefano, n., matthews, p. m., smith, s.m., johansen-berg, h., james, a. c. (2010). longitudinal changes in grey and white matter during adolescence. neuroimage, 49(1), 94-103. doi:10.1016/j.neuroimage.2009.08.003 gogtay, n., giedd, j. n., lusk, l., hayashi, k. m., greenstein, d., vaituzis, a. c., . . . thompson, p. m. (2004). dynamic mapping of human cortical development during childhood through early adulthood. proc natl acad sci u s a, 101(21), 8174-8179. doi:10.1073/pnas.0402680101 gordon, e. m., laumann, t. o., gilmore, a. w., newbold, d. j., greene, d. j., berg, j. j., . . . sun, h. (2017). precision functional mapping of individual human brains. neuron, 95(4), 791-807. e797. goswami, u. (2006). neuroscience and education: from research to practice? nat rev neurosci, 7(5), 406-411. doi:10.1038/nrn1907. hasson, u., & honey, c. j. (2012). future trends in neuroimaging: neural processes as expressed within real-life contexts. neuroimage, 62(2), 1272-1278. hoeft, f., mccandliss, b. d., black, j. m., gantman, a., zakerani, n., hulme, c., . . . reiss, a. l. (2011). neural systems predicting long-term outcome in dyslexia. proceedings of the national academy of sciences, 108(1), 361-366. hook, c. j., & farah, m. j. (2013). neuroscience for educators: what are they seeking, and what are they finding? neuroethics, 6(2), 331-341. horga, g., kaur, t., & peterson, b. s. (2014). annual research review: current limitations and future directions in mri studies of child‐and adult‐onset developmental psychopathologies. journal of child psychology and psychiatry, 55(6), 659-680. howard-jones, p. a. (2009). introducing neuroeducational research: neuroscience, education and the brain from contexts to practice: routledge. howard-jones, p. a. (2011). from brain scan to lesson plan. psychologist, 24(2), 110-113. howard-jones, p. a. (2014). neuroscience and education: myths and messages. nature reviews neuroscience. howard-jones, p. a., varma, s., ansari, d., butterworth, b., de smedt, b., goswami, u., . . . thomas, m. s. (2016). the principles and practices of educational neuroscience: comment on bowers (2016). huttenlocher, p. r. (1979). synaptic density in human frontal cortex-developmental changes and effects of aging. brain res, 163(2), 195-205. iacoboni, m., lieberman, m. d., knowlton, b. j., molnar-szakacs, i., moritz, m., throop, c. j., & fiske, a. p. (2004). watching social interactions produces dorsomedial prefrontal and medial parietal bold fmri signal increases compared to a resting baseline. neuroimage, 21(3), 1167-1173. jeneson, a., & squire, l. r. (2012). working memory, long-term memory, and medial temporal lobe function. learning and memory, 19(1), 15-25. doi: 10.1101/lm.024018.111 kingstone, a., smilek, d. & eastwood, j. d. cognitive ethology: a new approach for studying human cognition. br. j. psychol. 99, 317–340 (2008). klingberg, t. (2010). training and plasticity of working memory. trends in cognitive sciences, 14(7), 317-324. ko, l.-w., komarov, o., hairston, w. d., jung, t.-p., & lin, c.-t. (2017). sustained attention in real classroom settings: an eeg study. frontiers in human neuroscience, 11(388). doi:10.3389/fnhum.2017.00388 kwak, s., joo, w.-t., youm, y., & chey, j. (2018). social brain volume is associated with in-degree social network size among older adults. proc. r. soc. b, 285(1871), 20172708. lebel, c., & beaulieu, c. (2011). longitudinal development of human brain wiring continues from childhood into adulthood. j neurosci, 31(30), 10937-10947. doi:10.1523/jneurosci.5302-10.2011 lejuez, c. w., aklin, w. m., zvolensky, m. j., & pedulla, c. m. (2003). evaluation of the balloon analogue risk task (bart) as a predictor of adolescent real-world risk-taking behaviours. j adolesc, 26(4), 475-479. doi:10.1016/s0140-1971(03)00036-8 lenroot, r. k., gogtay, n., greenstein, d. k., wells, e. m., wallace, g. l., clasen, l. s., . . . giedd, j. n. (2007). sexual dimorphism of brain developmental trajectories during childhood and adolescence. neuroimage, 36(4), 1065-1073. doi:10.1016/j.neuroimage.2007.03.053 lewis, p. a., & durrant, s. j. (2011). overlapping memory replay during sleep builds cognitive schemata. trends cogn sci, 15(8), 343-351. doi: 10.1016/j.tics.2011.06.004. logothetis, n. k. (2008). what we can do and what we cannot do with fmri. nature, 453, 869-878. masten, c. l., telzer, e. h., fuligni, a. j., lieberman, m. d., & eisenberger, n. i. (2010). time spent with friends in adolescence relates to less neural sensitivity to later peer rejection. social cognitive and affective neuroscience, 7(1), 106-114. matusz, p. j., dikker, s. huth, a. g., perrodin, c. (2018) are we ready for real-world neuroscience? journal of cognitive neuroscience 0(0): 1-12. mayberg, h. s. (2014). neuroimaging and psychiatry: the long road from bench to bedside. hastings center report, 44(s2). mccandliss, b. d., kalchman, m., & bryant, p. (2003). design experiments and laboratory approaches to learning: steps toward collaborative exchange. educational researcher, 32(1), 14–16. mcgaugh, j. l. (2000). memory a century of consolidation. science, 287(5451), 248-251. mehta, r. k., & parasuraman, r. (2013). neuroergonomics: a review of applications to physical and cognitive work. frontiers in human neuroscience, 7, 889. melby-lervåg, m., redick, t. s., & hulme, c. (2016). working memory training does not improve performance on measures of intelligence or other measures of “far transfer” evidence from a meta-analytic review. perspectives on psychological science, 11(4), 512-534. milivojevic, b., vicente-grabovetsky, a., & doeller, c. f. (2015). insight reconfigures hippocampal-prefrontal memories. current biology, 25(7), 821-830. doi:10.1016/j.cub.2015.01.033 owen, r., macnaghten, p., & stilgoe, j. (2012). responsible research and innovation: from science in society to science for society, with society. science and public policy, 39(6), 751-760. paus, t., zijdenbos, a., worsley, k., collins, d. l., blumenthal, j., giedd, j. n., rapoport, j. l., evans, a. c. (1999). structural maturation of neural pathways in children and adolescents: in vivo study. science, 283(5409), 1908-1911. peper, j. s., mandl, r. c., braams, b. r., de water, e., heijboer, a. c., koolschijn, p. c., & crone, e. a. (2013). delay discounting and frontostriatal fiber tracts: a combined dti and mtr study on impulsive choices in healthy young adults. cereb cortex, 23(7), 1695-1702. doi:10.1093/cercor/bhs163 peters, s., & crone, e. (2017). increased striatal activity in adolescence benefits learning. nature communications, 8: 1983. doi: 10.1038/s41467-017-02174-z pinti, p., tachtsidis, i., hamilton, a., hirsch, j., aichelburg, c., gilbert, s., & burgess, paul w. (2018). the present and future use of functional near-infrared spectroscopy (fnirs) for cognitive neuroscience. annals of the new york academy of sciences 0(0). https://doi.org/10.1111/nyas.13948 poldrack, r. a. (2017). precision neuroscience: dense sampling of individual brains. neuron, 95(4), 727-729. poldrack, r. a., & farah, m. j. (2015). progress and challenges in probing the human brain. nature, 526(7573), 371. prince, p., & gifford, k. (2016). working memory goes to school. appl neuropsychol child, 5(3), 194-201. doi:10.1080/21622965.2016.1167502 raschle, n. m., lee, m., buechler, r., christodoulou, j. a., chang, m., vakil, m., stering, p.l., & gaab, n. (2009). making mr imaging child's play-pediatric neuroimaging protocol, guidelines and procedure. journal of visualized experiments: jove, (29). ruiter, d. j., van kesteren, m. t., & fernandez, g. (2012). how to achieve synergy between medical education and cognitive neuroscience? an exercise on prior knowledge in understanding. adv health sci educ theory pract, 17(2), 225-240. santoro, r., moerel, m., de martino, f., valente, g., ugurbil, k., yacoub, e., & formisano, e. (2017). reconstructing the spectrotemporal modulations of real-life sounds from fmri response patterns. proceedings of the national academy of sciences, 114(18), 4799-4804. doi:10.1073/pnas.1617622114 schleim, s., & roiser, j. p. (2009). fmri in translation: the challenges facing real-world applications. frontiers in human neuroscience, 3, 63. schmälzle, r., o’donnell, m. b., garcia, j. o., cascio, c. n., bayer, j., bassett, d. s., . . . falk, e. b. (2017). brain connectivity dynamics during social interaction reflect social network structure. proceedings of the national academy of sciences, 114(20), 5153-5158. schuijer, j. w., de jong, i. m., kupper, f., & van atteveldt, n. m. (2017). transcranial electrical stimulation to enhance cognitive performance of healthy minors: a complex governance challenge. frontiers in human neuroscience, 11, 142. shonkoff, j. p., & bales, s. n. (2011). science does not speak for itself: translating child development research for the public and its policymakers. child development, 82(1), 17–32. http://doi.org/10.1111/j.1467-8624.2010.01538.x sigman, m., peña, m., goldin, a., & ribeiro, s. (2014). neuroscience and education: prime time to build the bridge. nature neuroscience, 17(4), 497-502. sowell, e. r., thompson, p. m., holmes, c. j., jernigan, t. l., & toga, a. w. (1999). in vivo evidence for post-adolescent brain maturation in frontal and striatal regions. nat neurosci, 2(10), 859-861. doi:10.1038/13154 sowell, e. r., thompson, p. m., leonard, c. m., welcome, s. e., kan, e., & toga, a. w. (2004). longitudinal mapping of cortical thickness and brain growth in normal children. j neurosci, 24(38), 8223-8231. doi:10.1523/jneurosci.1798-04.2004 stilgoe, j., owen, r., & macnaghten, p. (2013). developing a framework for responsible innovation. research policy, 42(9), 1568-1580. tavares, r. m., mendelsohn, a., grossman, y., williams, c. h., shapiro, m., trope, y., & schiller, d. (2015). a map for social navigation in the human brain. neuron, 87(1), 231-243. telzer, e. h., fuligni, a. j., lieberman, m. d., miernicki, m. e., & galván, a. (2014). the quality of adolescents’ peer relationships modulates neural sensitivity to risk taking. social cognitive and affective neuroscience, 10(3), 389-398. teyler, t. j., & discenna, p. (1985). the role of hippocampus in memory: a hypothesis. neuroscience and biobehavioral reviews, 9(3), 377-389. thompson, p. m., giedd, j. n., woods, r. p., macdonald, d., evans, a. c., & toga, a. w. (2000). growth patterns in the developing brain detected by using continuum mechanical tensor maps. nature, 404(6774), 190-193. doi:10.1038/35004593 van atteveldt, n., van aalderen-smeets, s. i., jacobi, c., & ruigrok, n. (2014). media reporting of neuroscience depends on timing, topic and newspaper type. plos one, 9(8), e104780. van den bos, w., rodriguez, c. a., schweitzer, j. b., & mcclure, s. m. (2014). connectivity strength of dissociable striatal tracts predict individual differences in temporal discounting. j neurosci, 34(31), 10298-10310. doi:10.1523/jneurosci.4105-13.2014 van kesteren, m. t., fernandez, g., norris, d. g., & hermans, e. j. (2010). persistent schema-dependent hippocampal-neocortical connectivity during memory encoding and postencoding rest in humans. proceedings of the national academy of sciences of the united states of america, 107(16), 7550-7555. van kesteren, m. t. r., krabbendam, l., meeter, m. (2018). integrating educational knowledge: reactivation of prior knowledge during educational learning enhances memory integration. npj science of learning 3(1): 11. van kesteren, m. t., rijpkema, m., ruiter, d. j., morris, r. g., & fernandez, g. (2014). building on prior knowledge: schema-dependent encoding processes relate to academic performance. journal of cognitive neuroscience, 26(10), 2250-2261. doi:10.1162/jocn_a_00630 van kesteren, m. t., ruiter, d. j., fernandez, g., & henson, r. n. (2012). how schema and novelty augment memory formation. trends in neurosciences, 35(4), 211-219. wagner, d. d., kelley, w. m., haxby, j. v., & heatherton, t. f. (2016). the dorsal medial prefrontal cortex responds preferentially to social interactions during natural viewing. journal of neuroscience, 36(26), 6917-6925. willingham, d. t., & lloyd, j. w. (2007). how educational theories can use neuroscientific data. mind, brain, and education 1(3), 140–149. yeatman, j. d., dougherty, r. f., ben-shachar, m., & wandell, b. a. (2012). development of white matter and reading skills. proceedings of the national academy of sciences, 109(44), e3045-e3053. etelapelto et al publication frontline learning research vol.6 no. 3 (2018) 6 36 issn 2295-3159 a multi-componential methodology for exploring emotions in learning: using self-reports, behaviour registration, and physiological indicators as complementary data anneli eteläpeltoa, virpi-liisa kykyria, markku penttonen a, päivi hökkä, a, susanna paloniemi a, katja vähäsantanen a, tuomas eteläpelto b, vesa lappalainen b a faculty of education and psychology, university of jyväskylä, finland b faculty of information technology, university of jyväskylä, finland article received 22 may 2018 / revised 30 august / accepted 1 october/ available online 7 december abstract studies on emotions in learning are often based on interviews conducted after the learning. these do not capture the multi-componential nature of emotions, nor how emotions are related to the processes of learning. we see emotions as dimensional, multi-componential responses to personally meaningful events and situations. in this methodologically advanced pilot study we developed a multi-componential methodology, capable of providing complementary information on emotions in professional learning. for this purpose, we used a within-subject design applied to a single individual, with a focus on emotions during professional learning. within a laboratory setting, the subject was shown personally meaningful video extracts from a learning situation in which she had previously participated. the data were gathered through (i) self-reports of emotions via the emotion circle (ec) online assessment tool, (ii) measures of autonomic nervous system (ans) activity obtained via electrodermal activity (eda) and heart rate variability (hrv), (iii) behavioural registration of facial expression and gaze, and (iv) the stimulated recall interview (sri). self-reports of emotions via ec, and also the emotion-driven sri, were found to be productive, not only in detailing and explaining emotions experienced during the viewing of the videos, but also in bringing about reflective learning and novel insights. eda and hrv provided complementary information on the subject’s ans activity during the learning process. we present conclusions and future challenges in applying a multi-componential methodology to research emotions within professional learning. keywords: learning process; emotions; multimethod measuring; on-line self-reports; autonomic nervous system info. corresponding author mail: anneli.etelapelto@jyu.fi doi: https://doi.org/10.14786/flr.v6i3.379 acknowledgment professor anneli eteläpelto’s highly-meritorious professional career in the field of adult education at university of jyväskylä (finland) will continue in her new role as an emerita. we greatly appreciate her years of dedication to research, and especially all the inspiring ideas she has introduced over the years! among other roles, she has served the academic community as the coordinator of earli sig14 “learning and professional development” and hosted the sig meeting in 2008. further, she is renowned for her repeated successes in securing prestigious grants from the academy of finland and elsewhere. anyone who has had been fortunate enough to collaborate with anneli has probably found her an inspiring colleague and, sometimes, surprising in her contributions. in her new role as a professor emerita, she will undoubtedly continue to contribute actively to the discourse on learning and professional development. this special issue describes a starting point for new direction of research, which she has actively advocated, and we hope she will remain involved in these developments also in the future to see the vision for which she advocated fully realised. all the best and warm regards on behalf of all your colleagues! christian, raija, erno, and stephen 1. introduction there is fairly convincing evidence on the vital role of emotions in learning. emotions are related to motivational processes, self-efficacy, and active engagement, each of which has a salient role in productive learning in schooling contexts (pekrun, elliot & maier, 2006). in research on adult learning, positive emotions have been found to broaden the scope of perception, whereas anxiety and fear have been connected to a narrowing of the perception and curiosity necessary for active and agentic learning (fredrickson & branigan, 2002; hökkä, vähäsantanen, paloniemi & eteläpelto, 2017; perry, 2006; storbeck & maswood, 2015; sung & yih, 2015). in team-based learning, social and self-conscious emotions such as compassion, love, shame, anxiety, and anger have been found to influence how team members see each other, and how they perceive the future of the team (homan, van kleef & sanchez-burks, 2015). furthermore, it has been shown that in work organizations emotions critically influence the work-related learning manifested in job performance, motivation, creativity, decision-making, turnover, psychological wellbeing, teamwork, and leadership (barsade & gibson, 2007). despite convincing findings on the vital role of emotions in human learning, there has been a lack of research on the role of emotions in professional learning settings. these are characterized by processes of active influencing and developing, and by the negotiation of professional identity (eteläpelto, vähäsantanen, hökkä & paloniemi, 2013; vähäsantanen, räikkönen, paloniemi, hökkä & eteläpelto, 2018). recent studies have emphasized that professional learning occurs, in particular, via collaboration within social interaction, plus experimentation and reflection regarding one’s professional mission and practices (zwart et al., 2015). moreover, to achieve useful outcomes, social learning processes should be designed in such a manner that professionals have sufficient resources to shape their own professional identity and work (eteläpelto, vähäsantanen, hökkä & paloniemi, 2014). this touches on the notion that if meaningful learning is to be achieved, it is crucial to acknowledge and promote agency, and negotiations on professional identity (philpott & oates, 2017). here we emphasize reflection and insight concerning one’s own ways of thinking and interacting as important means of professional learning, after particular events. this kind of learning could manifest itself, for example, in insights into one’s ways of thinking and acting, plus the reasons for these, with greater overall self-awareness emerging as a result (vall et al., 2018). note that we have here adopted the term ‘professional learning’ as denoting the individual as an active and reflective participant, that is, someone who is responsible for learning and for constructing change at a personal level within a given context (labone & long, 2016; eteläpelto, 2017; vähäsantanen, paloniemi, hökkä & eteläpelto, 2017a). recent studies have shown that processes of group-based identity learning are particularly imbued with strong emotions, and that these emotions strongly influence actual learning outcomes, manifested as renegotiated work identities (vähäsantanen, hökkä, paloniemi & eteläpelto, 2017b; vähäsantanen, hökkä, paloniemi, herranen & eteläpelto, 2017). however, the findings on emotions are based on interview data collected after the learning experiences. thus, they do not take into account the multi-componential nature of the emotions (involving subjective experience, the autonomic nervous system, and behavioural changes), or how these are related to the processes of learning. a truly multi-componential understanding of emotions would appear to require a multi-componential approach to the measuring of emotions. the question then arises of how to overcome the methodological challenges of multi-componential measurement, and how to apply methodologically advanced tools to investigate emotions and related learning. in fact, there are now unobtrusive, technologically advanced tools such as face reading and gaze analysis (e.g. azevedo et al., 2013, 2016; zembylas & schultz, 2016) which, together with psychophysiological and self-assessment methods, can be used in the simultaneous collection of multiple data from different subsystems of emotions. however, in such simultaneous measurements, one must address questions of complementarity, interchangeability, validity, and reliability. in addition, there are challenges regarding the differing time windows of the various modalities. in addition to evidence on the relationships between learning and emotion, there has long been discussion on how emotions trigger or inhibit human behaviour (hommel, moors, sander & deonna, 2017), and how the emergence of emotions is closely connected to autonomic nervous system (ans) activity (levenson, 2014; mauss et al., 2005). recent discussion on emotion has been critical of mind-body dualism, pointing to evidence on how closely these interact, especially in the affective domain. it has long been recognized that ans has a central role in emotions. as levenson (2014, p. 100) has noted, ‘when it comes to emotions, all roads lead to ans’. the ans functions via two opposite but interacting regulation systems, i.e. the sympathetic and the parasympathetic nervous system. in experiences of fear and anxiety, the sympathetic nervous system produces the ‘fight or flight’ response, whereas the parasympathetic nervous system is responsible for calming our body and mind (‘rest and digest’). the ans is thus closely allied to the human experience of emotions, and the two branches of the ans are closely intertwined with behavioural responses, including changes in facial expressions and the active focusing of attention. it should be noted that the multi-componential nature of emotions has been accepted in a wide range of theoretical models of emotions (kreibig, 2010; mauss & robinson, 2009; gendolla, 2017). since early research on emotions, there has been a consensual understanding on emotions as comprising the following subsystems: (i) subjective experiences, (ii) psychophysiological responses emerging from the unconscious functioning of the ans, and (iii) behavioural and action-related manifestations. although there is agreement on the multi-componential nature of emotions, there is no agreement on the relations between the components of subjective experience, the ans, and behavioural changes. cognitive theories of emotions assume that there is a top-down relationship between levels in the subject’s cognitive assessment of emotions, and that this influences the responses of the ans and behavioural phenomena (e.g. lazarus, 1991). in contrast with this, evolutionarist-functionalist theories view emotions as organizing the activity of the ans and other physiological systems (levenson, 2014). the relationships between the different subsystems are often addressed in terms of coherence, referring to the coordination, or association, of a person’s experiential, behavioural, and physiological responses as an emotion unfolds over time (mauss et al., 2005). despite active theoretical discussion of such coherence, and of the nature of relationships between the three response systems, there is a lack of empirical evidence concerning how emotions organize activity within the ans, and between the ans and other response systems, including facial expression and subjective experience (levenson, 2014). nevertheless, empirical research has revealed that subsystem coherence and synchronization vary between different emotions (kreibig, 2010; levenson, 2014). it has also been found that different emotions (such as fear and joy) may activate different patterns of ans activity. however, research does not agree on the patterning, coherence, and specificity of the different subsystems (levenson, 2014). as indicated above, agreement on the multi-componential nature of emotions implies multi-method measurement of emotions. however, disagreement on the direction of influence between mind, behaviour, and ans responses brings challenges in interpreting the relationships between data derived from different subsystems. this implies a need for a pilot study which would address the relevant methodological challenges. we therefore sought to construct a multimethod research setting within which we could collect information simultaneously from the different subsystems of emotions operating within professional learning. in this article we consider the kinds of information that appear to emerge from such a methodology. our aim is to form some tentative conclusions on how the methods used may provide complementary data, allowing us to capture different aspects of emotions, plus their role within work-related learning processes. the study reported here was based on an understanding of emotions as multi-componential phenomena, manifested simultaneously in subjective experience and in the unconscious autonomic nervous system, and also in behaviour (with nonverbal behaviour operating somewhere in between these, being partly conscious and partly unconscious). the study aimed to capture emotions simultaneously on three levels, namely (i) subjective experience, (ii) ans, and (iii) behaviour. the methods can be summarized as follows (see also figure 2, section 5.1): (i) video-recorded episodes of previous learning situations were presented to the subject here referred to as ‘lisa’ (pseudonym) in a laboratory setting. her subjective experiences of these were elicited using an on-line application called emotion circle (ec), developed for the self-assessment of emotions concurrently with watching videos (see section 5.2.1). the videos were of episodes from a training programme in which lisa had participated previously. the episodes were selected on the basis that lisa had perceived them as personally meaningful, and as containing highly significant events for her learning within the program. the concurrent assessments of emotions (plus their connections to learning) were elaborated using the stimulated recall interview (sri) method (kagan, krathwohl & miller, 1963; kykyri et al., 2017). the sri interview took place immediately after lisa had given her ec assessment. (ii) ans activity was measured using one of the most reliable indicators of arousal emerging from sympathetic nervous system activity, namely electrodermal activity (eda), as measured via skin conductance (sc). in addition, heart rate (hr) and breathing were registered as indicators of ans activity involving the parasympathetic nervous system, manifested as heart rate variability (hrv). (iii) behavioural indicators, which are partly regulated by the autonomous nervous system, include facial expressions. these were captured through non-intrusive video-recordings, capable of being analysed via the computational methods of facereader (facereader version 6.1., 2015). in addition, the focus and direction of visual attention were registered using gaze recording (begaze version 3.7. manual, 2016). in this single case study, a within-subject design was used. a within-subject design is recommended by researchers investigating concurrent responses and the coherence of multiple sub-systems of emotions (levinson, 2014; mauss & robinson, 2009; kreibig, 2010). our purpose was to pilot the research procedure and multi-componential methodology required when one is investigating concurrent measures applicable to changing emotions, especially those emotions that occur within professional learning processes. in the best case, this would give us insights into the subjective experience of emotion – observed concurrently with ans functioning, and the behavioural manifestations of emotions – plus indications on how the measures may provide information on emotions in learning. we anticipated that the study conducted on ‘lisa’ would contribute to an understanding of the kinds of complementarity that can occur between the indicators used and the three different systems under study (subjective experience, the ans, and behaviour). the following section specifies what is known about emotions in terms of (i) the sub-components of emotion, (ii) dimensional vs. discrete models, and (iii) how emotions are related to action. based on this specification, we present our understanding and definition of emotions, necessary for addressing the validity of the methods used. 2. the conceptualization of emotion from darwin (1872/1965) onward, researchers on emotions (ekman, 1992; lazarus, 1991; levenson, 1994) have argued that emotions involve coordinated changes across experiential, behavioural, and physiological response systems (mauss et al., 2005). tomkins (1962) has suggested that emotions are sets of organized responses that (when activated simultaneously) are capable of simultaneously capturing widely distributed organs (the face, the heart, and the endocrines) and imposing on them a specific pattern of correlated responses. most, but not all, of the theorists above have taken a functional perspective, proposing that by imposing coherence across response systems, emotions facilitate the organism’s response to environmental demands, and prepare the organism for a set of diverse actions (levenson, 1994, 2003; mauss et al., 2005). for many theorists, a defining feature of emotion is response system coherence (mauss et al., 2005). this refers to the coordination, association, concordance, or organization of the response tendencies pertaining to a person’s experiential, behavioural, and physiological responses as the emotion unfolds over time (ekman, 1992; lazarus, 1991; levenson, 1994; scherer, 2009; tomkins, 1962). the notion that response coherence is a core feature of emotion suggests two corollaries. first of all, response coherence should increase as the intensity of emotion increases. weak emotions may provoke little coordination of the response systems, whereas strong emotions may provoke greater coordination. secondly, different emotions should be associated with different patterns of experiential, behavioural, and physiological response, tailored to meet the demands of different situations (mauss et al., 2005). for instance, amusement might be associated with facial displays of amusement, increased body activity, and a commensurate pattern of increased cardiovascular and electrodermal responding. by contrast, sadness might be associated with facial displays of sadness, decreased body activity, and a commensurate pattern of decreased cardiovascular and electrodermal response (mauss & robinson, 2009). nevertheless, in a review by mauss and robinson (2009), the researchers observed that, in contrast with the theoretically assumed coherence of the response systems, the empirical findings have been mixed. in fact, psychophysiologists have long emphasized the weak correlations between experiential and physiological response systems, and even between various measures within the physiological response system. along similar lines, more recent studies have found relatively modest correlations between experiential, behavioural, and physiological measures in the context of specific emotional states such as fear (mauss et al., 2005). in recent discussions on emotions, there has been a debate between discrete and dimensional models of emotions. discrete models have assumed that different emotions are associated with discrete and invariant patterns of response within each response system. by contrast, dimensional models assume that measures of emotional response reflect dimensions rather than discrete states. dimensional perspectives argue that emotional states are organized by underlying factors such as valence and arousal (e.g. barrett & russell, 1999). discrete emotion perspectives, by contrast, suggest that each emotion (e.g. anger, sadness, joy) has unique experiential, physiological, and behavioural correlates. having conducted a review, mauss and robinson (2009) concluded that research tends to support the dimensional perspective. thus, the recent consensual, componential model of emotions conceptualizes emotions as dimensional phenomena comprising experiential, physiological, and behavioural responses to personally meaningful stimuli (mauss & robinson, 2009; russell, 2005). in such a model (see figure 1), an emotional response begins with appraisal of the personal significance of an event (lazarus, 1991), which in turn gives rise to an emotional response involving subjective experience, physiology, and behaviour (frijda, 1986). figure 1. a consensual component model of emotional response (modified from mauss & robinson, 2009) a dimensional model of emotions is also present in the circumplex model of emotion (russell, 2005), which we have used as a framework in developing the emotion circle (see section 5.2.1). this is a tool aimed at the on-line assessment and self-reporting of emotions (eteläpelto et al., 2017). the circumplex model proposes that all affective states arise from cognitive interpretations of core neural sensations, which are the product of two independent neurophysiological systems. in the circumplex model, the vertical dimension depicts the level of arousal, whereas the horizontal dimension depicts the valence (pleasure vs. displeasure) of emotions. posner, russell and peterson (2005) argue that the circumplex model of affect is consistent with many recent findings from behavioural, cognitive neuroscience, neuroimaging, and developmental studies of affect. over many years there has been discussion on whether and how emotions influence action, including active learning. in a summary of recent research, hommel, moors, sander, and deonna (2017) conclude that there is little doubt that emotions influence action. emotions can influence the motivation process, and thus action, by fulfilling at least three functions. first of all, the emotions experienced can function as strong need-like motivational states. secondly, anticipated emotions can function as incentives, and justify action. thirdly, emotions can give information on progress in goal pursuit, permitting behaviour calibration (hommel et al., 2017). although the various theories assume a causal relation between emotion and action, they nevertheless disagree on the direction of this causal relation. feeling theories, such that of james (1884), see action as partly the cause of emotion. by contrast, other theories take emotion to be the cause of action. unsurprisingly, theories that see actions as part of emotions expect the processes involved in generating emotions to play a role in action as well (hommel et al., 2017). in the present article, emotions are understood as dimensional and multi-componential (i.e. experiential, physiological, and behavioural) responses to a personally meaningful antecedent event or situation, causing changes in the quality of subjective feeling, expressive behaviour, and physiological activation (kreibig, 2010; levenson, 2014; mauss & robinson, 2009). the multi-componential and dimensional understanding of emotions does not imply coherence between the different subsystems of emotion. neither does it specify the causal relationships that may exist between changes in the three components of emotions. however, coherence might increase as the intensity of the emotion increases. a dimensional understanding means that major variation in the subsystems of emotions may occur in terms of intensity and valence, without separate emotions necessarily having a specific fingerprint within or across different subsystems. furthermore, we assume that emotions are closely intertwined with the active processes of learning, even if the direction of influence or the specific causal links cannot be hypothesized. a lack of coherence between the subsystems of emotion can be manifested for several reasons. for example, subjective experiences of emotions can change without changes in autonomic nervous system responses, and ans changes may occur without concomitant changes in subjective experience. the reasons for this could derive, for example, from the lack of a subject’s faithful report on the emotional states experienced (kreibig, 2010). alternatively, couplings between the subsystems might be absent in self-reports of emotion because of a subliminal stimulus, unconscious emotions, or emotion regulation (kreibig, 2010; öhman, carlson, lundqvist & ingvar, 2007). what does this multi-component and dimensional understanding of emotion imply for the measurement of emotions? mauss and robinson (2009) suggest that the lack of strong convergence among multiple measures of emotion implies that the construct of ‘emotion’ cannot be captured with any one measure considered in isolation. they conclude that the more measures of emotion that are obtained, and the better tailored they are to the particular context and research question, the more likely it is that a researcher will learn from a particular study. since different measures of emotion appear to be sensitive to different dimensional aspects of an emotional state (with eda sensitive to arousal, facial expression sensitive to valence, etc.), they can be expected to act in a complementary way. to sum up, a multiple-component, dimensional understanding of emotions (with not much coherence between the response systems) implies first of all that we need to measure all three response systems. thus, we need to use multiple methods in such way that they cover all three subsystems (subjective experience, behaviour, and the ans). secondly, the dimensional understanding implies that we need to focus on the dimensions of both intensity and valence. thirdly, our interest in investigating emotions in connection with learning processes implies that we need to measure how emotions elicit learning processes and outcomes – and vice versa, i.e. how learning processes elicit emotions. the following section addresses the strengths and limitations of different measures as they encompass subjective experience, the ans, and behavioural subsystems of emotion. 3. the strengths and limitations of self-reporting, physiological, and behavioural indicators of emotions in learning the self-reporting of emotions has several clear strengths. pekrun (2016) suggests that self-reports can render more differentiated assessments of emotions than any other current method. self-reports are especially important for a nuanced description of emotions and thoughts. from a practical perspective, self-reporting is a more economical means of data collection than (for example) observations (pekrun, 2016). however, self-reports also have clear disadvantages, since they are limited to those emotional responses that a participant is aware of, i.e. to emotions that are inside her/his conscious mind. thus, they do not capture unconscious emotional processes (for unconscious emotions see winkelman & berridge, 2004). in addition, if self-reports are collected retrospectively, memory-bias is common. kreibig (2010) suggests that self-reports of emotions are likely to be more valid to the extent that they relate to currently experienced emotions. even in this case, however, there are concerns that not all individuals are aware of and/or capable of reporting on their momentary emotional states (mauss & robertson, 2009). in addition, there are differences between cultures in the understanding of terms relating to emotions, as well as between individuals sharing the same cultural background. an additional threat to validity emerges from the human tendency towards impression-management, which is always present in the social context of data collection (azevedo et al., 2016). the bias pertaining to impression-management and the displaying of emotions might actually be particularly strong among professionals such as teachers and leaders, whose professional work competences comprise the skills of displaying and regulating their own emotions (e.g. kreibig, 2010). nevertheless, even if it were easy to change how one reports one’s emotions for the purposes of impression-management, it may be much more difficult to change one’s level of physical activation in efforts at coping and impression-management (azevedo et al., 2016). from this point of view, physiological and behavioural indicators are needed to increase validity and reliability in the measuring of emotions. all this would imply that self-report methods should not be used alone, and should be supplemented with psychophysiological and behavioural indicators. this would also accord with an understanding of emotions as taking place in three subsystems which do not cohere, such that the measures of one cannot substitute for another. azevedo et al. (2013, 2016) have suggested the use of electrodermal activity (eda), which has long been used as a marker of sympathetic nervous system activity. in addition, researchers addressing emotions in connection with learning have suggested the use of facial expressions and eye-tracking to detect, identify, and classify affective states during learning (azevedo et al., 2013, 2016). eda refers to changes in the skin’s ability to conduct electricity. it can be measured by well-established and non-intrusive techniques that provide unique information on emotional arousal, increased cognitive work load, and task-engagement. overall, the higher the conductance level rises, the more elevated the subject’s emotional arousal becomes (huhtamäki et al., 2017; laitila et al., 2018). it has been suggested that if eda is a good predictor of cognitive load, task difficulty, and task engagement, it might also be used for predicting boredom, disengagement, and frustration in learning (azevedo et al., 2013, 2016). a feature of eda is that it is highly sensitive to small changes in low states of anxiety (boucsein, 2012; hugdahl, 1995). in addition, eda is responsive to behavioural inhibition and to defensive strategies. thus, eda increases if thoughts are suppressed or if emotional expressions are inhibited. nevertheless, an increase in eda is not equivalent to the occurrence of emotions with negative valence or stress; in fact, eda is also responsive to increased activation within the body in the case of happy excitement (benedeck & kaernback, 2010; karvonen, 2017). the limitations involved in using eda are often positioned around the difficulty of extracting information from the ‘raw’ eda signal. in addition, there can be significant daily variation in participants’ eda response, with eda levels increasing linearly throughout the day, and even from one season to another. from such considerations, azevedo et al. (2013, 2016) recommend including eda, but not limiting oneself to it, in efforts to measure ans activity. a central technique for studying ans emotion responses has been heart rate variability (hrv), which refers to a variety of methods for assessing the beat-to-beat change in the heart over time (quintana & heathers, 2014). hrv describes the variation between consecutive inter-beat intervals. both sympathetic and parasympathetic branches of the ans are involved in the regulation of heart rate (hr). sympathetic nervous system (sns) activity increases hr and decreases hrv, whereas parasympathetic nervous system (pns) activity decreases hr and increases hrv (brentson at al., 1997; tarvainen et al., 2018). hrv is commonly used as a tool in assessing cardiac autonomic regulation, and it has been used in a range of wellbeing applications to evaluate the functioning and balance of the ans (tarvainen et al., 2018). it has been suggested that hrv can be used to investigate the relationships between autonomic regulation and interpersonal interaction (porges, 2001). increased hrv has been found to indicate a feeling of safety. in group-based learning contexts, such a feeling has been connected to having space for self-reflection and the emergence of new ideas. in contrast, reduced hrv has been observed in disorders characterized by poor social cognition and emotion regulation (quintana & heathers, 2014). despite this, there are severe limitations concerning the use of hrv. because of the large inter-individual variation in hrv, it has been suggested as more appropriate to a within-subject design. in addition, hrv is affected by respiratory depth and frequency, with both breathing and blood pressure regulation having their own directly mediated relationships with hrv. thus, social-emotional tasks that induce changes in respiratory time variables and/or depth may indirectly influence hrv. in addition, a number of studies have shown continuous focused attention (and thus tasks with increased attentional demands) to reduce hrv, primarily because of changes in respiratory depth and frequency. this creates further difficulties for interpretation (quintana & heather, 2014; quintana, alvares & heathers, 2016). the starting point for hrv analysis is the electrocardiogram (ecg) recording, from which the hrv time series can be extracted. in the formulation of hrv time series, the fundamental issue is determination of the heart beat period (tarvainen et al., 2018). in addition to the ans measures, researchers addressing emotions connected to learning have suggested the use of facial expressions and eye-tracking. these might allow researchers to detect, identify, and classify affective states during learning (azevedo et al., 2013, 2016). in fact, the use of facial expression is a fundamental process-oriented approach in the detection and classification of affective states (azevedo et al., 2016; ekman, 1992). in everyday social interaction, facial expression has a clear communicative function, but it also functions as an automatic and observable indicator of expressed emotion. recently, data collection on facial expression has been conducted through software linked to a video data stream of the learner’s facial expressions. in addition, there are now many comprehensive, widely supported methods to objectively describe the facial expressions of emotions, based on ekman’s facial action coding system (facs). these analyse basic emotions described as (for example) enjoyment, fear, anger, sadness, disgust, and they indicate how the emotions evolve over time. commercial applications for facial expression recognition software have been developed to automate the coding process involved. despite the use of automated software to register the duration, fluctuation, transition, and dynamics of affective states, azevedo et al. (2016) emphasize that several disadvantages should be considered. first of all, the systems rely heavily on the quality of the video stream. there must not be shadows on the faces. moreover, eyeglasses can make it difficult to measure some emotions, and can increase the likelihood of interpretation as a neutral gesture. in addition, postures and body movement are limited, the software can only identify one face at a time, and the results are restricted to sets of specific predefined emotions. despite this, azevedo et al. (2013; 2016) suggest that the benefits of automatic facial expression recognition software outweigh the limitations. it permits non-intrusive, reasonably valid, and concurrent measurements, which can be integrated with other process data channels. thus, the method has clear potential significance. eye tracking is commonly used in learning research. the data can provide detailed description of the areas and focus of the subject’s interest (azevedo et al., 2016; hautala, loberg, hietanen, nummenmaa & astikainen, 2016). eye-tracking data can be used to measure the learner’s fixations (where?), fixation durations (how long?), saccades (eye movements from one fixation to another), plus the order and the gaze patterns (multiple fixations) that occur during learning. from these data, we can gain insights into what learners pay visual attention to during learning, and this can be both indicative and predictive of emotions (bondareva, conati & feyzi-behnagh, 2013). nevertheless, despite the strengths of eye-tracking data for research on emotion and learning, eye tracking alone cannot definitely indicate what elicited the emotion; nor can one determine what the effect of this emotion may be on subsequent activity. thus, eye-tracking data must be supplemented by other data collection methods if one is to capture the subject’s sense-making processes, as they take place in the reciprocal relationship between learning and emotions. the stimulated recall interview (sri) method has been increasingly used as the video-recording of learning situations has become more popular (kagan, krathwohl & miller, 1963). especially in group learning contexts, when one can organize subsequent viewing of the group situation and of one’s ways of acting in the situation, there is the cognitive capacity and space to gain new insights, via reflection and re-evaluation of one’s own actions (huhtamäki et al., 2017; lyle, 2003; vall et al., 2018). in teacher education, the video-assisted procedure has long been used to increase self-reflection, and thus possibilities to learn about oneself as a teacher (fuller & manning, 1973). these goals (involving an increase in self-reflection, and an understanding of oneself as a leader) were in fact the learning goals in the leadership coaching program addressed in this study. hence, sri was expected to be a useful method for our purposes. we expected that viewing the videotaped learning situations in connection with the perceived emotions would provide space for reflection on the original learning situation (at a first viewing) with insights into the emotions experienced (at a second viewing). 4. research questions and aims of the study the research questions can be specified as follows: 1. what kinds of information do concurrent self-reports, indicators of ans, and behavioural measures of emotion provide, in terms of understanding a person’s emotional responses within professional learning? 2. to what extent does the emotion-driven stimulated recall interview (sri) promote reflection and hence learning, related to the original learning situation? the prime aim of the study was to form tentative conclusions on how self-reports, indicators of ans, behavioural measures of emotion, and the emotion-driven stimulated recall interview (sri) may provide complementary information on emotions in professional learning. a further aim was to elaborate the potentials and challenges of a multi-componential methodology for researching emotions in professional learning. 5. methods 5.1. design, procedure, and stimulus material the multi-method measurement procedure described below was designed for ultimate use in wider data collection. in the present case we applied the procedure and data collection setting to a single subject. a practical purpose in this pilot study was thus to elaborate what might need to be changed or developed for wider data collection. the ethical committee of jyväskylä university evaluated the design and approved the study. all the participants mentioned in the article, and all participants at other times, gave their informed consent regarding participation. figure 2 gives a general description of the procedure (including the stimulus material and data collection methods) used in this study. figure 2 shows that in session i, the subject viewed four successive video-recorded episodes (a1, b1, b2, a2) and assessed these via the on-line emotion circle tool, described in detail in section 5.2.1. in session ii, which followed immediately after session i, an emotion-driven stimulated recall interview (sri) was conducted while the subject again viewed the same four episodes (a1, b1, b2, a2), which were shown together with the saved ec data from her emotion assessments in session i. psychophysiological data were collected via measurement of eda, hrv, and respiration. electrodes measuring eda were attached to the left hand, hr electrodes to the chest, and the respiration belt around the chest (see figure 2). behavioural data from eye movement were gathered via the begaze version 3.7 (smi) recorder, which was located at the lower edge of the display. the face reader 6.1. (noldus) video camera (which was used for recording the subject’s facial expressions) was located behind the screen (see figure 2). figure 2 procedures used in this study sessions i and ii took place under laboratory conditions in april 2018. in contrast, the selected episodes (a1, b1, b2, a2), which were used in the laboratory settings as stimulus material, were selected from the original learning settings. these had taken place five years earlier (in 2012–2013.) our subject, lisa, had participated in this earlier group-based training platform (constituting a leadership coaching program). the program was constructed to cultivate (i) participants’ professional identities, (ii) their ways of managing social relationships within the work community, and (iii) their professional communication (hökkä, vähäsantanen, paloniemi & eteläpelto, 2017; vähäsantanen, paloniemi, hökkä & eteläpelto, 2017). the program had twelve workshops in all. it covered a period of eleven months, with one day per month allotted to it. the selected episodes b1 and b2 were those which lisa had mentioned as representing the most important and meaningful learning situations for her, personally, within the program. episodes a1 and a2 were selected as representing neutral episodes. these were episodes which lisa had not mentioned at all. all of them came from the seventh workshop. the procedure of selecting and ordering the episodes to be viewed was constructed according to the abba model. this meant that at the beginning and at the end of the session there were emotionally neutral episodes. the episodes lisa had selected as most influential in terms of her learning (b1 and b2) were placed in the middle of the temporal continuum. the episodes differed in length, being chosen in such a way as to comprise coherent authentic learning episodes, bearing in mind that they would otherwise have been too fragmented for the subject to understand and evaluate. under laboratory condition, the four episodes were played consecutively (total duration = 17 min. 23 s.). the contents of the episodes were as follows: episode a1 (2 min.) depicted an informal get-together situation before the formal start of the training. in this episode, group members came to the seminar room in which the group learning session would take place. about half of the participants, plus the trainer, were already standing there face to face. they were chatting and talking informally with each other. one by one, the remaining participants walked into the room. some of them went directly to sit down in chairs placed in the form of a half circle. there was mix of voices, thus it was impossible to differentiate individual voices, although one could see the people who entered the room. episode b1 (9 min. 23 s.) depicted a situation in which some of the participants (who were close to each other as colleagues) talked about their present feelings using symbolic object working. this took place in such a way that the participants, plus the trainer, (13 persons altogether) were sitting in chairs which were in the form of a circle. in the middle of the circle there was an empty chair. this symbolized an empty place in which each participant could place an imagined object, one that best represented a central issue in their life at that moment. at the start of episode (b1), a participant, ‘bertha’ started to describe how busy she was in her work at that moment, and how she had tried to organize some weekend trip in order to recover a little; however this had merely caused more stress and internal conflict in her mind. the trainer put some questions to bertha concerning the issues that would emerge if there was space for something other than work. after this, another person (‘caroline’, a close colleague of the previous speaker) started to talk. she described very similar feelings of having too much work, with very stressful feelings connected, for example, to the salary negotiations taking place in her work organization. caroline continued to talk about serious issues regarding her health and wellbeing. while talking about these stress-induced health issues, and how they were connected also to family issues, caroline burst into tears. after this she expressed surprise about the strong emotions connected to her situation. immediately after the start of the weeping, lisa stood up. she picked up a tissue from a side table and gave it to the crying person. one by one, two other close colleagues of caroline started to cry (indicating emotional contagion). the remaining participants looked very serious. the trainer put further questions concerning how far caroline had listened to her own mind, and how she could take more time to take care of her own health and wellbeing. the situation ended with the talk of a third person, who was also crying. at the end of the episode, that person made a joke. this caused people to laugh, and feel more relieved. episode b 2 (4 min.) depicted a situation exemplifying how to work with difficult cases. the work had started with a preliminary task in which each participant had called to mind, from their personal work history a difficult colleague. they had described the situation in writing before the session and sent it to the trainer. in this episode, participants first sat face to face, in the form of a circle. the subject of the present study, ‘lisa’, role-played her difficult case using a drama method. she took the role of the difficult person, changing her voice and way of talking, to resemble that of the difficult person. after this role-play, in the next part of the episode, the participants turned their chairs 180 degrees, so that they were no longer face to face. from this ‘turning chair’ position, lisa presented what she was actually thinking, and what she would have liked to say (as her authentic self) to the difficult person. she was then saying aloud what she would have wanted to say to the imagined difficult person, and what she actually could not have said in real life. after this, the chairs were once again turned to the face-to-face position. lisa now explained to the whole group the kind of history she has had with the difficult case. at the beginning of each piece, the trainer always put a question to be answered. in presenting her reasons regarding the difficulty of the person, lisa also described how she had at last become empowered to set limits to the negative and destructive behaviour of the difficult person. episode a2 (2 min.) depicted a pair-discussion session of the whole group. the participants were sitting in chairs, which were in the form of a circle. they actively discussed with the person next to them. there was mix of voices; hence, one could not differentiate individual voices. however, one could see smiling faces, loud laughter, and the participants’ active concentration during discussion with their pairs. 5.2 data collection 5.2.1. self-reports via the emotion circle (ec) self-reports were collected from lisa via the emotion circle (ec) on-line assessment tool. this was developed for concurrent assessment of the quality and intensity of emotions. here it should be noted that there has been a lack of valid, user-friendly tools to capture changing emotions during learning processes. in our research project, we developed the ec on-line application for the self-assessment and reporting of individual shifting emotions within professional learning settings (eteläpelto et al., 2017). ec utilizes a colourful graphic interface containing 12 written emotion words (figure 3). these are presented, in line with the circumplex model of emotions (russell, 2005). the circumplex model has been designed as a single-item scale providing a quick means of assessing affect along the dimensions of pleasure – displeasure (valence) and arousal – sleepiness. pleasure is considered to be the bipolar opposite of displeasure, and the subjective feeling of arousal to be the bipolar opposite of sleepiness. these dimensions are further considered to be orthogonal (i.e. independent) and thus conceptually separated (russell, weiss & mendelsohn, 1989). the single-item scale based on these two dimensions is envisaged as an instrument that will be short and easy to complete in assessing the subjective experience of continuously and rapidly fluctuating emotions in repeated-measures design (russell, weiss, & mendelsohn, 1989). one can anticipate that multiple-item checklists or questionnaires would be too time-consuming and distracting for the purpose of reporting continuously changing emotions within the learning process. in the circumplex model of emotions, valence is described along the x axis so that on the left there are unpleasant (negative) and on the right pleasant (positive) emotions. arousal is described along the y axis so that on the upper segment (containing plus values) there are emotion constructs characterized by high arousal (‘hot’ emotions) whereas on the lower segment (containing minus values) there are emotion constructs characterized by low arousal (‘cold’ emotions). in accordance with the circumplex model of emotions, different emotion words are placed in the emotion circle (ec) (see figure 3). this means that on the upper right quadrant of the circle, there are emotions characterized by pleasant activation and activated pleasure, such as excitement, surprise, and joy. on the lower right quadrant there are emotions characterized by deactivated pleasure and pleasant deactivation (safety, compassion, courage). on the upper left quadrant there are emotions characterized by unpleasant activation and by activated displeasure (irritation, anxiety, and fear). on the lower left quadrant there are emotions characterized by deactivated displeasure and by unpleasant deactivation (shame, frustration, and sadness). the intensity of emotions is depicted in the ec via degrees of colour saturation. in the middle of the ec the colours are lighter, denoting less intensity of emotion. as one moves to the circumference, the colours become more saturated, denoting more intense emotions. figure 3. the display of the emotion circle (ec). figure 3 shows the ec display. the emotion constructs used in the emotion circle were selected on the basis of interviews conducted with the 11 participants of the leadership coaching program. within the interviews, the interviewees were first presented with an open question concerning their perceived emotions during the program. they were then shown a list of 28 emotion constructs that were expected to represent the most common emotions felt during the program. these emotions were based on the final interviews conducted at the end of the program in 2013 (hökkä, vähäsantanen, paloniemi & eteläpelto, 2017). in selecting the relevant emotion constructs, we also utilized prior studies on emotions in leaders’ identity negotiation (e.g. winkler, 2018). self-reports have been criticized on the grounds of the subjects’ tendency to report more positive emotions because of a social preference for these. because of this, ec included roughly the same number of positive (ec right side) and negative (ec left side) emotionally-related words. it was anticipated that this would help to counteract any tendency towards the reporting of positive rather than negative emotions. ec automatically saves the process data (time and object of clicking) from the subject’s assessments. this assessment video (in the present case video recording 1) can be shown to the subject immediately after the assessment. the ec application makes it possible to collect subjects’ self-assessments of their situation-specific emotions, including also data on the quality and intensity of the emotions, plus their dynamic continuity. it transforms and displays the process in such a way that it can be synchronized with other (physiological and behavioural) process data, collected at the same time. the ec also makes it possible to show the subjects’ recorded assessments together with the video-recording of the situation which the subjects have assessed. this is needed for stimulated recall. in using ec, we seek to avoid the memory bias of retrospective interviews. when we used the ec, the subject, lisa, was first given general instructions on it, with opportunities also to practise the use of it. she was asked to click on the emotion word or words which in each situation represented her subjective experience of the emotion. in order to guarantee that she had properly understood the use of ec, she was asked to imagine some emotion, then to click on the ec accordingly. after lisa had confirmed that the use of ec was easy for her, we played the recorded episodes in the order a1, b1, b2, a2, using ec to collect her self-reports, including the nature and intensity of her emotions. 5.2.2. behavioural data we obtained behavioural and expressive data from lisa using automatic gaze recordings (via begaze) and video-recording of gestures (via noldus face reader). the methods used are unobtrusive. they include a gaze recorder, located at the lower edge of the display. a video camera is used to collect data on facial expression (face reader data). this is located behind the display, with an additional light placed on the upper side of the display to prevent shadows on the faces. 5.2. 3. autonomic nervous system recordings autonomic nervous system (ans) recordings were taken from lisa during the video viewing and assessment situation. this further continued during the stimulated recall session which followed the viewing and assessment. during the ans recording sessions the following signals were recorded using the quickamp amplifier and data acquisition system (brain products, gilching, germany): the electrocardiogram (ecg) was recorded with two ag/agci electrodes (ambu neuroline 710, ballerup, denmark), attached above and below the heart, with a similar ground electrode attached over the stomach. electrodermal activity (eda) during the session was recorded via two skin conductance (sc) electrodes (el507, biopac systems, california, usa) on lisa’s non-dominant palm, below the first and fourth digits. the palm was chosen as the location, because in piloting and in previous research (karvonen, 2017), it was found that there is less measurement error from hand movements when the electrodes are in that area, as compared to fingertips. sc was determined using 0.5 v constant voltage (gsr sensor, brain products, gilching, germany). respiration during the session was registered via a fabric belt (respiratory effort sensor, spes medica, genoa, italy). this was fastened on top of lisa’s clothes, on the lower chest area. however, respiration was not analysed in this study, because the data quality was found to be inadequate (due to the belt having been too loosely attached, as discovered after the session). eda and respiration were amplified in dc mode, but the ecg was 0.5 hz high-pass filtered. signals were acquired with a sampling frequency of 1000 hz, using a data acquisition program (brainvision recorder, brain products, gilching, germany). a custom-made marker unit was used to synchronize the ans measures to the video. 5.2.4. the stimulated recall interview (sri) after lisa’s assessment of her emotions during viewing of the video-recorded learning episodes, a videoand ec-assisted stimulated recall interview (see sri; kagan, krathwohl & miller, 1963) was conducted (immediately after the assessment). in this interview, the video episodes were shown to her, along with her emotion assessments given via ec. she was encouraged to share her thoughts, feelings, and reasons at any time while watching the videos, including her assessments of the emotions that she had experienced while watching the videos, and during the original situation as she recalled it. when she started to speak, the video was stopped to give her time to explain and describe her thoughts. to assist in this phase, lisa was given general instructions, in the form of questions to consider, as follows: what thoughts, feelings, or bodily sensations did you have while watching the video and assessing your emotion connected to it? we assumed that naturally, she could have forgotten many of the feelings she had experienced during the original learning sessions (which had taken place five years previously), and that she might simply describe the thoughts, feelings, and sensations evoked by watching the session videos. nevertheless, it seemed reasonable to suppose that she might be able to recollect some intense emotions from the original learning sessions. for this reason, she was further asked to specify whether she had had a particular thought or feeling in the original learning session, or whether that thought or feeling emerged only now in the assessment session. the same ans measures were recorded during the sri as during the emotion assessment session, and the interview was recorded with a video camera. after the stimulated recall interview, lisa was further asked to comment on the user-friendliness of the emotion circle, and of the data collection sessions as a whole. 5.3 analysis of the data so far, there have not been many analytical (and especially statistical) techniques which would be relevant for analysing data from single-case research, characterized by the rapid and randomly determined alternation of conditions or processes. in their review of existing analytical techniques in single-case design, manlov and onghena (2017) found visual analysis, constituting the classical way of analysing single-case data, to be the most frequently applied technique. it has been suggested that in most cases, visual analysis is sufficient to demonstrate evidence of a relationship between conditions and outcome variables (kratochwill et al., 2013). a baseline phase is necessary to represent the control condition and to provide a clear basis for comparison. visual analysis allows comparisons between and within phases, indicating levels, trends, and variability, with data also on overlaps, the immediacy of the effect, and the consistency of patterns. visual analysis can be used to suggest a functional relation between conditions and the target behaviour, and it can indicate the most salient features of the data. visual data can be used as an initial step in the analysis, and it can be regarded as complementary to statistical analysis in making sense of the data obtained (manlov & onghena, 2017). in this study we used visual analysis to compare the levels, trends, and variability of the ans data on eda and hrv, within and between the four time segments (episodes a1, b1, b2, and a2), referring also to the pause segment, which was used to give a baseline value for the comparisons. visual analysis was also used to carry out a general assessment of the data patterns in eda and hrv. statistical analysis was used as complementary to visual analysis in the analyses of eda data. from the eda data, statistically significant (p<.05) values were calculated. in addition, visual analysis was applied in the descriptive analysis of the concurrent variation of different data modalities. this was done in respect of the ans data and the on-line self-report data on emotions collected via the emotion circle. in the visual analysis addressing the complementarity of ans data and self-reports, a figure was constructed using the time stamps written on the excel file. this indicated the exact time points and the specific emotion assessed with ec while the subject reported her emotion. in this visual description, emotions with a positive valence (excitement, surprise, joy, courage, compassion, safety) were placed on the upper side of the horizontal time axis, while emotions with a negative valence (sadness, frustration, shame, fear, anxiety, and irritation) were placed under the time axis. specific colours connected to the different emotions in ec were also presented in the visual description (figure 5.). from the ec data we calculated the absolute and percentual frequencies of clicks on different emotion words within each episode. data from the subject’s eda were analysed with the ledalab program (version 3.4.6) written in matlab (benedek & kaernbach, 2010; www.ledalab.de). before the analysis, the sampling rate was reduced to 10 hz, which was high enough to represent rapid changes in sc related to sns activation. the rapid components of sc were extracted as skin conductance responses (scrs) and written to an excel file. the scrs were normalized by computing the average and standard deviation of the session, and calculating z-scores. values larger than 2.0 were considered to be statistically significant at the p<.05 level (given that 5% of the values have that property; they can thus be considered to represent statistically significant sns activation). the rationale behind the analysis of the statistically significant eda peaks is based on the assumption that eda can track rapid and unconscious changes in sympathetic nervous system (sns) activity in very brief time windows. an increase in sns activity is related to the increased physiological arousal that accompanies most emotions and also preparation for action (boucsein, 2012; kreibig, 2010). in particular, rapid changes in eda (measured as skin conductance responses, and indicated by increased sweating, especially in palms, fingers, and feet) are thought to be a direct measure of the phasic neuronal activity of the sns (benedek & kaernbach, 2010). laitila et al. (2018) have demonstrated the added value gained from analysing the eda responses that occur in the social interaction of a couple therapy session, as a means to detect important moments of change at individual and interpersonal level. in our subjectively meaningful social learning sessions, we were also interested in detecting critical moments of change that had made the selected episodes b1 and b2 personally meaningful for lisa’s learning within the leadership coaching program. in line with the analysis of laitila et al. (2018), we expected that those values of eda which represented rare (p>.05) high peaks would reveal exceptionally high unconscious emotional arousal of the subject, and thus point to critical moments of learning. in the analyses of lisa’s eda, we compared the numbers of statistically significant high peaks between episodes of about the same length. the baseline value during the pause was used in the comparison. nevertheless, more important than mere detection of moments of change is determination of what has led to these moments, thus placing the focus on the learning process itself (laitila et al., 2018). this implies that if one is to make sense of eda data in terms of learning, they need to be complemented with other kinds of data. in the present study, ans data from eda (indicating rapid peaks in the activity of the sns) were complemented with hrv data, indicating the activity of the parasympathetic nervous system (pns). in contrast with the sns responses manifested in eda, the pns responds much more slowly, possibly over minutes rather than seconds. thus, during fairly short episodes (from 2 min to 9 minutes) there could be instances of overlapping from one episode to another. this is a feature that needs to be considered in interpreting hrv trends and changes. hrv was analysed in the time domain using the kubios hrv premium program (version 3.1; www.kubios.com). first of all, r peaks were detected in the ecg to determine the intervals between consecutive heart beats. possible artefacts were removed automatically. rr intervals were determined, and root mean squares of successive rr interval differences were determined for 60-second windows, starting from the beginning of the session, and covering the whole session in 10-second steps. the hrv was not normalized (in order to keep the relevant values transparent in milliseconds). face reader is based on the circumplex model of emotion. the model describes emotions in a two-dimensional circular space, containing arousal on the vertical axis and valence on the horizontal axis. the centre of the circle represents a neutral valence and a medium level of arousal. the circumplex model of facereader version 6.1. is based on the model described by russell (1980). the valence in face reader indicates whether the emotional status of the subject is positive or negative. ‘happiness’ is the only positive emotion, while ‘sadness’, ‘anger’, ‘fear’, and ‘disgust’ are considered to be negative emotions. ‘surprise‘ can be either positive or negative. the valence is calculated as the intensity of ‘happiness’ minus the intensity of the negative emotion with the highest intensity. for instance, if the intensity of ‘happiness’ is 0.8 and the intensities of ‘sadness’, ‘anger’, ‘fear’, and ‘disgust’ are 0.1, 0.0, 0.05, and 0.05 respectively, then the valence is 0.7 (see face reader version 6.1. reference manual, 2015, pp. 80–81). the focus of visual attention > was obtained via begaze, together with the subject’s assessments conducted with the on-line emotion circle. these process data were analysed using a process analysis, conducted via the video recordings, as depicted in the attached video depicting episode a2 (see video 1). the process analysis derived from the video recordings demonstrates how the focus of visual attention always preceded the subject’s selection of the emotion words she clicked. this information can be used in further analyses of the subject’s selection process, as well as in analyses of the usability of the emotion circle. in this study, the data were used to demonstrate the complementarity of the two relevant data modalities (comprising the more objective behavioural gaze data, and the subjective self-assessment data derived via the emotion circle). video 1. assessment process (2 min 3 s) as depicted for the process analysis with the assessments (red square) conducted via the emotion circle, and via the focus of gaze (blue circle). data from episode a2 (mp 4 format file). for the purposes of detailed process analysis, data derived from different methods can be further transformed via the open broadcaster software, changing the data into the mp4 format. this format can be moved to the observer xt12 for simultaneous display. in this way different data sets can be scrutinized simultaneously, and observed moment by moment on the screen. the sri interview data were transcribed verbatim, amounting to 3.5 pages (a4, 1.5 line space). the transcribed data were read and re-read by the second and first authors. we first identified verbal expressions of emotions, plus the explanations given for these. secondly, we identified self-reflective or other reflective contents of the utterances. finally, we focused on displays of insight and novel ideas which were not mentioned in the video-recorded episodes. 6. findings the findings are presented here in line with the research questions. we first describe the kinds of information available from concurrent self-reports, obtained via the emotion circle, the ans indicators, and the behavioural measures of emotion. these shed light on the components of emotions and learning processes, and the emotional responses that occur in professional learning (6.1). thereafter (6.2) we present findings based on the emotion-driven stimulated recall interview (sri), with comments on how these were connected to the learning situations in question. 6.1. information elicited via concurrent self-reports, ans indicators, and behaviour here we first describe findings from lisa’s self-reports on emotions, derived via the on-line emotion circle (6.1.1). we then address the facial expressions obtained via the automatic face reader; these functioned as complementary data to verify the self-reported emotions (6.1.2). ans data concerning electrodermal activity (eda), and heart rate variability (hrv), are presented in 6.1.3. 6.1.1. self-reporting of emotions via the on-line emotion circle (ec) the analysis of lisa’s self-reporting of emotions via the on-line emotion circle (done while watching the selected episodes of videotaped learning sessions) showed that during the four episodes (17 min. 23 s. in total), she clicked on emotions 134 times. table 1 shows that the clicks covered the whole time period fairly evenly. as expected, the quality of the emotions was different between the four episodes. emotionally neutral episodes (a1 and a2) were placed at the beginning and end of the session, with b1 and b2 being placed centrally. these were the episodes arousing the strongest emotions, and also the episodes which lisa had selected as most influential for her learning. regarding the contents of the emotions reported via ec, table 1 shows that only one of the given emotion words (irritation) was not reported at all. the most used emotion word was compassion. lisa would have liked to add to the ec the experience of feeling guilt, especially in episode b1. table 1 shows that emotion words with positive valence (used 94 times) were used more than twice as much as those with negative valence (used 40 times). table 1 further shows that the quality of lisa’s emotions was quite different between the four episodes. episodes a1 and a2 were fairly positive overall, since all the reported emotions exhibited positive valence (characterized by intrinsic pleasantness). as opposed to this, episode b1 included strong negative and unpleasant emotions, such as anxiety, sadness, and fear. these were accompanied by a high degree of compassion. by contrast, episode b2 was assessed as fairly positive in terms of emotions such as courage, surprise, joy, and excitement. in addition, episode b2 was reported as displaying safety and as including also compassion. table 1 number (absolute frequencies and percentages) of emotion words used in lisa’s assessment of four successive video-recorded episodes (a1, b1, b2, a2) via the on-line emotion circle (ec) while watching the videos, the perceived emotions were assessed by clicking on the emotion words given in the emotion circle (session i). however, this is a somewhat limited way of expressing one’s emotional experience. nor, taken by itself, does it indicate or give information on the reasons for the specific reported emotions. because of this, the ec assessments were further elaborated (in session ii) using the stimulated recall interview (sri) method (6.2). 6.1.2. face reader and valence of emotions differences between the episodes in terms of the valence of emotions were further analysed with the face reader (see figure 4). it confirmed the findings based on lisa’s self-reports via ec. episode b2 (figure 4 right) was characterized by positive valence, whereas episode b1 (figure 4 left) was full of very low and even negative valence. it should be noted that the glasses worn by the subject might influence error in counting gestures; thus, the findings showed a large amount of neutral emotion. empty spaces in the graph indicate the face reader software not being able to identify lisa’s gestures because she was temporarily facing away from the video. figure 4. screen captures of the valence of emotions within episodes b1 (left) and b2 (right). 6.1.3. ans responses: eda and hrv figure 5 shows the standardized eda values and hrv values calculated for the four subsequent episodes (a1, b1, b2, a2) in session i. in addition, in the middle of episode b1, there was a pause. this was not planned, and was due to technical problems. in fact, the sound of the video suddenly disappeared six minutes from the start of the viewing of the video, and four minutes from the start of episode b1, during the assessment of lisa’s emotions via ec. the technician then tried to work out the reason for the problem and make technical adjustments. this took eight minutes. during this time lisa could do nothing but sit and wait for the adjustments to be completed. in our analysis of the ans data, we realized that this technical failure provided valuable baseline data concerning lisa’s eda and hrv measures. at the start of the pause there was a rapid decrease in eda. this remained low over the next six minutes. just before the end of the pause there were some new attempts to recover the sound, with confusion and discussion concerning the functioning of the system. lisa participated in these discussions, as indicated by some increase in eda at the start of the pause. however, as figure 5 shows, during the pause, her eda values were at the lowest level for the entire data gathered during the session. the low eda values during the pause could be connected to her passive (non-agentic) motivational state (being unable to do anything). this is in accordance with kreibig’s (2010) suggestion concerning low eda: that it is characteristic of a passive state of action, and thus with deactivation of the sympathetic nervous system. if the eda levels of the pause situation are compared to the eda levels during episodes a1 and a2 (which were selected to represent emotionally neutral episodes) one can detect a clear difference between the pause and the ‘neutral’ episodes. especially in episode a1 (which came at the beginning of session i), the eda levels were relatively high, with two statistically significant (p<.05) peaks within this episode. these peaks apparently manifested lisa’s pleasure at seeing other participants in the video, coming into the room in which the training had taken place. episode a1 came at the start of the task of viewing and assessing emotions while viewing. hence, this represented a novel way of working, one that might set an additional load (manifested as an increased eda level). this is evident if one compares the eda levels between episodes a1 and a2. the latter was at the end of assessment session i, i.e. at a point when lisa was already familiar with the task. in the case of episode a2, she also knew that this was the last part of session i. she might therefore have felt more relieved than at the start of the session. nevertheless, if we focus (figure 5) on the eda levels during the personally meaningful episodes b1 and b2, we can see that there were many statistically significant high peaks in b1 (actually two before the pause, and 16 after the pause). in addition, the frequency of these peaks was fairly dense, and the frequency increased in the course of the episode b1. as described above, in lisa’s self-report concerning her emotions, episode b1 was described mostly in words with negative valence, such as anxiety, sadness, and fear. here, she reported a high level of compassion with the person who was talking about her health problems. in the sri, lisa said that she felt as if she were the body of that person. thus, the high eda values here appear to be connected with the subjective experience of high negative stress (kreibig, 2010). nevertheless, the high eda levels were not connected merely to stress with negative valence. figure 5 further shows that high eda peaks were present also in episode b2, which was subjectively assessed as having fairly positive valence (i.e. in terms of self-reports). episode b2 was described as including the subjective experience of courage, surprise, joy, and excitement. lisa also reported feelings of safety and compassion (see figure 5). this indicates that high eda is not connected merely to emotions with negative valence, and can be related also to emotions with positive valence. in general, lisa’s eda levels here seemed to be related to active cognitive and affective work while viewing episodes that were personally highly important to her, and to assessing her emotions while watching these. these situations, depicted in episodes b1 and b2, were also those which she had selected as highly meaningful for herself in terms of her perceived learning outcomes within the leadership coaching program. the eda peaks, which were connected to watching these episodes, would thus indicate critical incidents in the learning processes, but also a high intensity of emotion. this is in accordance with prior research on eda and its connections with the level of arousal, as manifested within active learning processes, and also in states of intense emotion (kreibig, 2010). figure 5 also shows the values and changes in hrv in the course of the initial assessment session (i). as compared to eda, which mainly indicates the activity of the more rapid sympathetic nervous system, hrv indicates the (generally more slowly fluctuating) activity of the parasympathetic nervous system. hrv was in the present case calculated in 60 s time windows, which were moved forward in 10-second steps to cover the whole session. the calculation of hrv as the root mean square of successive heart beat differences (rmssd) is not usually performed for periods shorter than 60 seconds (usually including 60–100 beats). hence, computationally, the changes in the hrv could not be as rapid as in eda, in which ten samples were used to represent phasic skin conductance responses in a second. an increase in hrv is a general marker of relaxing (stein, bosner, kleiger & conger, 1994). nevertheless, there is also some natural variation in the hrv, and it is thus not directly connected to the events that take place in a given situation. figure 5 shows that there was slowly changing variation in the hrv within session i. at the beginning of the session, during episode a1, the hrv was very high, thus functioning as a marker of a relaxed body. however, at the end of episode a1, the hrv dramatically decreased. it increased again at the start of the next episode, b1. thereafter, the hrv again decreased during the two eda peaks before the pause. at the beginning of the pause, when there was confusion arising from the loss of the sound for technical reasons, the hrv again steeply decreased. after the start of this unexpected technical failure, the hrv then increased during the next six minutes of the pause time. this level of hrv provided lisa’s individual baseline in the passive situation, in which deactivation of the sympathetic nervous system was indicated by low eda levels. after the pause, when episode b1 (characterized by negative valence) continued, the hrv tended to decrease throughout the episode, although the decrease was not linear. especially at the end of episode b1, when there were many high and dense eda peaks, the hrv was clearly decreasing. figure 5. self-assessments of emotions via ec, eda, and hrv during session i. the decrease in the hrv continued further in episode b2 which was characterized by positive valence, but also by high eda peaks. in the subjective self-assessment of episode b2, lisa gave many reports of surprise together with courage (see figure 5). however, towards the end of session b2, her hrv started to increase. here it should be noted that episode b2 was much shorter (4 min) than episode b1 (9 min 23 sec). since activation of parasympathetic nervous system combined with relaxing and calming down (indicated by the increase in hrv) takes place fairly slowly, this might have an influence on the delay in the hrv increase when hrv was activated within episode b2. next subchapter addresses to what extent does the emotion-driven stimulated recall interview (sri) promote reflection and hence learning, related to the original learning situation. 6.2. findings from the stimulated recall interview (sri) the assessments given via emotion circle (ec) together with the video-recorded learning episodes were used as the stimulus material for the sri. in the sri, we asked lisa to explain her thoughts, feelings, and bodily sensations, plus her reasons for her assessments via ec. with this procedure (sri) we mainly aimed to elicit the connections between emotions and the video-recorded learning situations. our findings, based on a qualitative content analysis of the transcribed interview data, showed that the emotion-driven sri produced very rich and heterogeneous descriptions, comprising self-reflection, other-related reflection, new emotion words, plus comments concerning the sri method. these descriptions illustrated reasons for specific emotions, and the connections between the reported emotions and the details of the learning situations and processes. in addition, the sri produced new insights concerning the meanings given by lisa to her own past behaviour, as well as the change in lisa’s behaviour during the learning assignment. table 2 provides a summary of the findings concerning the four episodes (a1, b1, b2, a2), including also a description of the emotional tone during the sri. a1 in the sri for this episode, the emotions lisa mentioned most frequently were ‘surprise’ and ‘joy’. she commented on the reasons behind these emotions in a neutral and calm manner. she had felt surprised and happy at seeing many familiar individuals arriving at the training venue. there were no intense emotional expressions or reflective comments relating to this episode. b1 in the sri for this episode, lisa provided nuanced descriptions, involving: (i) reasons for specific emotions, (ii) new emotions (i.e. emotions which she had not selected in the ec). she also gave reflective comments on the interactions while she viewed them. the comments involved (i) self-reflection on her own activities and bodily experiences, (ii) reflection on the activities of the other participants and the trainer, (iii) interpretative comments concerning other participants’ behaviour, (iv) reflections on the group atmosphere. in this episode, one participant was talking about her health problems. while watching the video, lisa re-lived strong emotions, i.e. ‘anxiety’, ‘sadness’, and ‘fear’. she described her re-lived emotions (and the self-reflection connected to these) as follows: …now i’m getting very anxious and i feel compassion, because this is my group…and i’m responsible for the fact that they have worked too hard, and i feel compassion and guilt, which is not included as an option in the ec. i swing between anxiety and compassion, but guilt is the strongest feeling, since i truly see that people are working too hard, but i wonder how…i haven’t realized that [sighing] so that it feels really bad…that i can feel i’ve failed… and i can feel sadness… lisa also described embodied experiences, in comments such as ‘surely, at that point my electrodermal activity was at a high level’, ‘i tried to take a deep breath in that situation’, or ‘there i definitely have tears in my eyes’. lisa suggested the addition of a new emotion word, ‘guilt’, which was not available in the ec. in addition, within the sri, she named the emotions ‘irritation’, and ‘joy’, which she had not selected in the ec assessment phase, even though these emotion words had been options in the ec. in the sri for b1, lisa also reflected on the behaviour of the group, as follows: ….then i look at the entire group, the way everybody is sort of frozen, or fortunately, it provides space [for the emotional expression] and in this sense there is safety in the group, so that people can empathize with each other and feel relieved... [smiling] b2: this episode consisted of working with ‘a difficult case’ using a drama method. in the related sri, lisa provided diverse descriptions, including (i) reasons for the specific emotions, (ii) self-reflection and self-analysis, (iii) reflection on the group, and (iv) a new insight concerning the reasons why she herself perceived the case as so difficult. the corresponding sri started with lisa’s self-reflection and self-analysis connected to the selected ec emotion of ‘surprise’, and to how this surprise emerged from her role-playing of a difficult character. she described this as follows: …so i’m wondering and i’m surprised, wondering if that’s me, i’m so bad at role-playing, but ok, i’m surprised that it’s as if oh my god…i have that difficult person in my mind all the time the one i’m role-playing…but also i’m astonished, about whether it’s me that’s speaking there or whether it’s my role character….actually i’m a bit ashamed about whether i’m so bad at role-playing but since i really am bad at role-playing… in the citation above, there is also comment on her own behaviour and on issues which are bothering her. these troubling thoughts emerged especially in relation to the other group members, and what they might think about her own role-playing of a difficult person. this further created self-doubt concerning what she had said in the group. she expressed the troubled feelings, including reflection on the group, as follows: …and then i look at those group members wondering what they’re thinking because at that time [referring to the original group learning situation] i couldn’t see it while i was concentrating on pretending to be that role character, so what i was saying … i’ve really got such a that.. uh uh that that somehow i’ve spoken inappropriately… after this, lisa spoke in a way that implied balancing between the feeling of being troubled and the emotion of being courageous (the emotion she had chosen most in the ec). she vacillated between a feeling of having been courageous, and a feeling of being ashamed of her courage. this can be seen as an attempt on her part to seek different interpretations of herself. there seems to be a struggle between, on the one hand, finding her own possible space, within which she can allow herself to be courageous, and on the other hand, social embarrassment, to the point of shame. the reflections here seem to involve tacit (previously unverbalized or unrecognized) emotion. however, within the sri, her talk took a positive direction, ending in laughter. in this way, the balance swung to a positive feeling, and to the conclusion that in fact she had not humiliated herself. after this, a novel insight emerged on the reasons why lisa had perceived the difficult case as so difficult. this insight emerged together with the reflective talk on her emotions, plus the reasons for these. in the sri (while watching the situation again in the video and elaborating her emotion of being courageous), she pondered on her behaviour as follows: … i’m surprised at how courageous i am [in this learning assignment]. um, i was pondering a lot …about whether i was going to dare to say it… that until then i was, well, that she was once my teacher…. when i started my studies in the department… and now [in this learning assignment] i could at last dare to tell her … even though she had been an authority figure to me… well okay i was really courageous… i still keep wondering how and why i was so bold to as say to her that here we are all in the same problematic situation, so why, i guess she had put herself above all the others, and she was a university lecturer when i first started to work in the department. well this has been for me i guess a kind of empowering moment and a really big issue, that i was able to tell her what i was thinking, since i have had a tendency to try to please everyone, and especially her, so that [earlier, when we worked together] i did not, i didn’t dare [laughs]… slightly later in the sri, lisa confirmed her new insight (which only emerged during the sri) as follows: ...now i look at this from a distance, this is how i have been acting… well maybe there’s some new insight about why this person was for me, well since she had previously been a great authority for me then that’s why it was such a big thing that after twenty or thirty years i had the courage to say to her you can’t act like that, tell people they can’t come to a group if they haven’t done their doctoral dissertation, just for that reason, come and make other people depressed, say you people are stupid [laughs] yeah this was something to remember… lisa also confirmed that this new idea, and the emotions attached to it, were merely those that emerged in the present situation (ec plus sri). she could remember that in the original learning setting (five years previously) she had role-played that person, but she did not remember the emotions she experienced at that time. a2: the sri for this episode produced utterances that were descriptive and fairly neutral – but also moderately positive – concerning the group atmosphere. they encompassed feelings of being safe, joyful, and surprised. in addition, she presented descriptive comments concerning individual participants in the group. table 2 summary of findings from the stimulated recall interview (sri) concerning the four episodes a1, b1, b2 and a2. 7. conclusions regarding complementarity and further challenges in using multiple methods our aim in this pilot study was to gain a preliminary understanding of how self-reports, indicators of ans, behavioural measures of emotion, and the emotion-driven stimulated recall interview (sri) may provide complementary information on the function of emotions in professional learning. in conducting the study, we wished also to elaborate the potentials and challenges of the multi-componential methodology we applied, in terms of researching emotions in professional learning. our interest in investigating emotions in relation to learning processes implies a need for a bi-directional perspective, involving how emotions elicit learning processes and outcomes, and how learning processes elicit emotions. in future discussion on emotions in learning we also need to consider recent discussion and disagreement between, on the one hand, scholars who think that emotions are universal and similar over different times and cultures (ekman, 1992), and on the other hand, those who see emotions as historically and culturally determined, and thus learned in socio-cultural contexts (barrett, 2006). the methods and tools used to which measure emotions from facial expressions (such as the facereader used in this study) based on ekman’s (1992) idea of basic emotions and the possibilities for universal measurement of them via facial movements. however, this may be misleading, bearing in mind the criticisms presented against the conception of universality in facial movements as indicators of emotions. human capabilities for emotion regulation, and individual differences in emotional intelligence (manifested as the ability to display emotions), can be expected to influence the presentation of emotions. there is, in fact, considerable evidence concerning the learning of emotions in cultural contexts, and this applies also to the learning and use of emotion words (barrett, 2006). the on-line self-reports (given via ec) and the sris based on these emerged as productive, not only in reporting and explaining one’s emotions during the learning process, but also in terms of promoting new reflective learning. this new learning was evidenced in what our subject, lisa, said in her sris, and especially in her elaborations of the emotions experienced while viewing the episodes. lisa’s self-reports via the ec, and behavioural data derived from her facial expressions, indicated that the selected episodes b1 and b2 were both connected with intense emotions. they also indicated a difference in valence, with the first of these (b1) being characterized by negative valence, and the second (b2) by positive valence. this was evident on the basis of her self-reports, and was validated by behavioural data derived from her facial expressions. the findings here include considerable measurement error, due to our subject’s need to use glasses. nevertheless, they are in line with previous studies showing high agreement between facial expressions and self-report data (harley et al., 2015). in the sri, lisa elaborated on her reasons for her emotion choices in the ec. while viewing the video clips, she could explain what had provoked the emotions. in the episode imbued with emotions of negative valence (b1), lisa reported compassion, anxiety, sadness, and fear. the next episode (b2), which was characterized by emotions with positive valence (surprise, courage, joy, compassion, excitement, and safety) seemed to produce new self-reflective ideas on the previous power relations operating between herself and a particular difficult person, and on why the relationship had been so difficult for her. this finding is in line with previous suggestions regarding the sri method as a means of promoting self-reflection (vall et al., 2018). however, the special feature in the present study was the use of the sri to focus specifically on emotions. the emotion-driven sri seemed to produce – in conjunction with reflections by lisa on herself, other members, the group interaction, and the reasons for her specific emotions – important insights concerning the causes underlying her ‘difficult case’. such novel learning, involving new insights, seems to support the productive nature of emotion-driven reflection on personally meaningful learning episodes. while strong emotions may exist as (partly unconscious) rapid events at particular moments, the elaboration of these moments afterwards, i.e. from a distance, appears to provide a productive basis for identity learning. the video-recorded episodes from personally meaningful learning settings unfolded first from distance in watching and assessing these via ec, and then slowly in the sri. they thus provided options to re-live and re-analyse the connections between the situation in which the emotions arose and the emotional responses that followed (see figure 1). all in all, this implies that within a methodology of presenting stimulus material, it is highly productive to select personally meaningful episodes for further elaboration, and to elaborate these from the perspective of emotions. the self-report method used in this study (combining the on-line ec reporting of emotions with the sri focusing on personally meaningful learning episodes) provided a powerful learning setting. it brought about deepened self-reflection, and novel insights, notably in the episode characterized by emotions with positive valence. overall, we would suggest that a combination of the on-line reporting of emotions with the emotion-driven sri is a promising methodology for investigating how emotions are connected to learning. in addition, such a combination can promote reflective insights that are of value in learning about one’s own identity. there is thus notable pedagogical potential in the sri, which appeared to bring about self-reflection, other-oriented reflection, insight, and learning. this finding resonates with the studies by vall et al. (2018) and huhtamäki et al. (2017), who noticed that the video-assisted sri stimulated reflection and insight, while facilitating therapeutic processing in couple therapy clients. during the learning process in session 1, the psychophysiological indicators eda and hrv were found to provide different and complementary information on the subject’s autonomic nervous system (ans) activity. the eda, which indicates the activity level of the sympathetic nervous system, was found to be at a high level during the viewing and assessing of the videotaped learning episodes. the lowest eda occurred during the pause, i.e. during a passive situation when there was nothing to do. comparison of the eda responses between the watched episodes showed that the highest eda peaks occurred during the viewing of personally significant episodes b1 and b2. in episode b1 (characterized by emotions with negative valence), there were many significantly high peaks, and the peaks were close to each other (i.e. dense). nevertheless, high eda peaks were also present in the other personally significant episode, i.e. the one characterized by positive emotions (b2). in this study, eda could be used to distinguish the personally meaningful episodes b1 and b2 from the passive and more neutral episodes a1 and a2. however, eda did not distinguish between the emotional valences (positive vs. negative) of the situation. for this purpose, we need other complementary methods, such as self-reports. another ans measure used here was heart rate variability (hrv), which is a measure of the activity of the parasympathetic nervous system (pns). in the present study, hrv increased (indicating high activity in the pns) thus pointing to processes of calming and relaxing during the (unintended) pause, as well as during emotionally neutral episodes a1 and a2. in contrast, hrv was found to decrease (thus indicating deactivation of the pns) in the emotionally intense episodes b1 and b2. one very interesting aspect was the decrease in hrv also during the transition from episode b1 (imbued with negative emotions) to episode b2 (imbued with positive emotions). however, the findings here should be treated with caution, given that this is a single-subject case study, and also that sensitivity to breathing (which is a feature of hrv) could have affected the hrv observations. in the present study we were unable to obtain breathing data due to a technical failure. hence, for future collection and analysis of hrv data it will be necessary to improve the reliability of the relevant equipment. the behavioural data (via face reader, gaze) were used merely to increase the reliability of the self-reporting data. in the face expression analysis we utilized – in line with our dimensional understanding of emotions – only the indicator of valence. as regards gaze data, in the present study these data were collected but not further analysed. however, the collection of such data would have utility in showing the focus of attention. in the future, we intend to use gaze data in addressing further the usability of ec. our efforts to construct measurement procedures for use under laboratory conditions produced many insights regarding technical details. these covered the proper installation of devices (such as, in the present case, the breathing belt which did not give proper data because it was too loosely tightened). furthermore, from the unplanned loss of the sound in the video display, and from the ans data collected during the pause, we learned that we actually require such a pause situation at the end of data collection sessions, to obtain the baseline for the subject’s ans data. one central challenge in this kind of multimethod measurement is the synchronisation of different devices. if there are problems with this, analysis of the data becomes very laborious. furthermore, in analysing ans data there are different time windows in eda and hrv. eda, which measures the responses of rapid sympathetic response systems, responds quickly (e.g. in a fight-flight situation). this means that the eda response occurs in just a few seconds from the stimulus situation or event. by contrast, the hrv measures the activity of the (slowly responding) parasympathetic nervous system. this means that in counting the hrv, the time window cannot be less than one minute. hence, the temporal connection to a specific event or situation remains less exact with hrv as compared to eda. for the future development of the emotion circle (ec), we gained much information from the emotion-driven stimulated recall interview (sri). it provided practical information concerning the usability and selection of emotion words. for the future, we need to test the optimal number of emotion words, and also how to take into account the saturation of colours in ec. in addition, there is a need to test a range of pictorial or iconic ways of displaying emotions in the ec. in the future, if our purpose is to measure simultaneously all three components of emotions (subjective experience, ans, and behaviour) within the processes of learning, we shall need a multidisciplinary team. we thus need to recognize that setting up this kind of multi-method measuring system will require a multidisciplinary group of researchers, comprising experts in educational sciences, psychology, psycho-physiology, and information technology. in addition, there will be a need for different kinds of inter-professional practical support and technical services. for future research on emotions in learning, we would suggest a focus on the continuities of emotional processes (in terms of the bodily activity of the ans, plus valence, and the intensity of the experienced emotions). the aim will be to create conditions that are optimal for researching and promoting supportive emotions in professional learning. keypoints this study developed a multi-componential methodology to measure emotions in learning. an on–line assessment tool, emotion circle (ec), was developed for the self-reporting of emotions during learning. the multimethod research design provided complementary information on the experiential, physiological, and behavioural components of emotions. the emotion-driven sri revealed connections between emotions and learning, and was productive in bringing about reflective learning and novel insights. acknowledgments the authors are grateful to the reviewers of the manuscript, and to donald adamson who polished the language of the article. they also wish to thank lauri viljanto and petri kinnunen for technical assistance, and hanna liljapelto who drew the figures. this work was supported by the academy of finland [grant number 288925 the role of emotions in agentic learning at work]. references azevedo, r., harley, j., trevors, g., duffy, m., feyzi-behnagh, m., & bouchet, f. et al. (2013). using trace data to examine the complex roles of cognitive, metacognitive, and emotional self-regulatory processes during learning with multi-agent systems . in r. azevedo & v. aleven(eds.), international handbook of metacognition and learning technologies (pp. 427-449). new york: springer. azevedo, r., taub, m., mudrick, farnsworth, & martin, s. (2016). interdisciplinary research methods used to investigate emotions with advanced learning technologies. in m. zembylas & p. schutz (eds.), methodological advances in research on emotion and education (pp. 231-243). dordrecht: springer. barsade, d., & gibson, s. g. (2007). why does affect matter in organizations? academy of management perspectives, 21, 36–59. barrett, l. f. (2006). are emotions natural kinds? perspectives on psychological science, 1, 28-58. barrett, l. f. & russell, j. a. (1999). the structure of current affect. controversies and emerging consensus. current directions in psychological science, 8,10–14. begaze version 3.7. (2016). manual.smi sensomotoric instruments. benedeck, m. & kaernback, c. (2010). a continuous measure of phasic electrodermal activity. journal of neuroscience methods, 190, 80-91. bondareva, d., conati , c., & feyzi-behnagh , r. (2013). inferring learning from gaze data during interaction with an environment to support self-regulated learning . international conference on artificial intelligence in education(pp. 229-239). springer. boucsein, w. (2012).electrodermal activity. newyork: springer. brentson, g. g., bigger, jr., eckberg, d. l., grossman, p., kaufman, p. g., malik, m., nagaraja, n. h., porges, s. w., saul, j. p., stone, p. h. & van den molen, m. w. (1997). heart rate variability: origins, methods, and interpretive caveats. psychophysiology, 34, 623-648. darwin, c.r. (1872/1965). the expression of emotions in man and animals. london, uk: john murray. dimaggio, g., lysaker, p., carcione, a., nicolo, g., & semerari, a. (2008). know yourself and you shall know the other to a certain extent: multiple paths of influence of self-reflection on mindreading. consciousness and cognition, 17, 778-789. ekman, p. (1992). argument for basic emotions. cognition and emotion, 6(3-4), 169-200. eteläpelto, a. (2017). emerging conceptualisations on professional agency and learning . in m. goller, & s. paloniemi (eds.), agency at work: an agentic perspective on professional learning and development (pp. 183-201). springer: cham. eteläpelto, a., hökkä, p., paloniemi, s., vähäsantanen, k., lappalainen, v., eteläpelto, t., & niskanen, k. (2017). developing an on-line application, emotion circle (ec) for the self-assessment of emotions in agentic learning at work. in l. g. chova, a. l. martínez, & i. c. torres (eds.), iceri 2017 proceedings:10th international conference of education, research and innovation. seville, spain (pp. 7763-7769). eteläpelto, a., vähäsantanen, k., hökkä, p. & paloniemi, s. (2013). what is agency? conceptualizing professional agency at work. educational research review, 10,45-65. http://authors.elsevier.com/sd/article/s1747938x13000274 eteläpelto, a., vähäsantanen, k., hökkä, p. & paloniemi, s. (2014). identity and agency in professional learning. in s. billett, c. harteis and h. gruber (eds.), international handbook of research in professional and practice-based learning volume 2 (pp. 645-672). dordrecht: springer. facereader version 6.1. (2015). reference manual. wageningen, nl: noldus information technology. fredrickson, b. l., & branigan, c. (2002). positive emotions broaden the scope of attention and thought-action repertoires. cognition and emotion,16, 313-332. frijda, n. h. (1986). the emotions. cambridge, uk: cambridge university press. fuller, f. f., & manning, b. a. (1973). self-confrontation reviewed: a video tape. a case study. journal of counseling psychology, 10(3), 237-243. gendolla, g. h. e. (2017). do emotions influence action? of course, they are hypo-phenomena of motivation. emotion review, 9(4), 348-350. harley, j.m., bouchet, f., hussain, m.s., azevedo, r. & calvo, r. (2015). a multi-componential analysis of emotions during complex learning with an intelligent multi-agent system. computers in human behavior, 48,615-625. hautala, j., loberg, o., hietanen, j. k., nummenmaa, l., & astikainen, p. (2016). effects of conversation content on viewing dyadic conversations. journal of eye movement research, 9(7), 1-12. hökkä, p., vähäsantanen, k., paloniemi, s., & eteläpelto, a. (2017). the reciprocal relationship between emotions and agency in the workplace . in m. goller, & s. paloniemi (eds.), agency at work: an agentic perspective on professional learning and development (pp. 161-181). springer: cham. doi:10.1007/978-3-319-60943-0_9 homan, a.c., van kleef, g.a. & sanchez-burks, j. (2015). team members’ emotional displays as indicators of team functioning. cognition and emotion, 30(1), 134-149. hommel, b., moors, a., sander, d., & deonna, j. (2017). emotion meets action: towards an integration of research and theory. emotion review, 9(4), 295-298. huhtamäki, h., lehtinen, r., kykyri, v-l., penttonen, m., karvonen, a., kaartinen, j., & seikkula, j. (2017). oivaltamisen hetket pariterapia-asiakkaiden jälkihaastatteluissa.[moments of insight in the stimulated recall interview of couple therapy clients]. perheterapia,33(4), 36-50. hugdahl, k. (1995). psychophysiology: the mind-body perspective. cambridge, ma: harvard university press. james, w. (1884). what is emotion? mind, 9,188-205. kagan, n., krathwohl, d. r., & miller, r. (1963). stimulated recall in therapy using video tape: a case study. journal of counseling psychology, 10,237-243. karvonen, a. (2017). sympathetic nervous system synchrony between participants of couple therapy. jyväskylä studies in education, psychology and social research599. university of jyväskylä. jyväskylä university printing house, jyväskylä, finland. kratochwill, t.r., hitchcock, j.h., horner, r.h., levin, j.r., odom, s.l., rindskopf, d.m., et al. (2013). single-case intervention research design standards. remedial and special education, 34,26-38. kreibig, s. d. (2010). autonomic nervous system activity in emotion: a review. biological psychology, 84(3), 394–421. kykyri, v.-l., karvonen, a., wahlström, j., kaartinen, j., penttonen, m., & seikkula, j. (2017). soft prosody and embodied attunement in therapeutic interaction: a multimethod case study of a moment of change. journal of constructivist psychology, 30(3), 211-234. doi:10.1080/10720537.2016.1183538 labone, e., & long, j. (2016). features of effective professional learning: a case study of the implementation of a system-based professional learning model. professional development in education, 42(1), 54–77. laitila, a., vall, b., penttonen, m., karvonen, a., kykyri, v-l., tsatsishvili, v., kaartinen, j. & seikkula, j. (2018). the added value of studying embodied responses in couple therapy research: a case study. family process 57,1-13. lazarus, r. s. (1991). emotion and adaptation. new york: oxford university press. levenson, r. w. (1994). human emotion: a functional view. in p. ekman & r. j. davidson (eds.), the nature of emotion: fundamental questions(pp. 123-126). new york: oxford university press. levenson, r. w. (2003). autonomic specifity and emotion. in r. j. davidson, k. r. scherer & h. h. goldschmith (eds.), handbook of affective sciences(pp. 212-224). new york: oxford university press. levenson, r. w. (2014). autonomic nervous system and emotion. emotion review, 6(2), 100-112. lyle, j. (2003). stimulated recall: a report on its use in naturalistic research. british educational research journal, 29(6), 861-878. manolov, r. & onghena, p. (2017). analyzing data from single-case alternating treatment designs. psychological methods, 16 (march), 1-25. mauss, i. b., levenson, r. w., mccarter, l., wilhelm, f. h., & gross, j. j. (2005). the tie that binds? coherence among emotion experience, behaviour and physiology. emotion, 5, 175-190. mauss, i. b., & robinson, m. d. (2009). measures of emotion: a review. cognition and emotion, 23,209-237. pekrun, r. (2016). using self-report to assess emotions in education. in m. zembylas, m. & p. a. schultz (eds.), methodological advances in research on emotion and education. springer international publishing. pekrun, r., elliot, a. j., & maier, m. a. (2006). achievement goals and discrete achievement emotions: a theoretical model and prospective test. journal of educational psychology, 98(3), 583-597. perry, d.p. (2006). fear and learning. trauma-related factors in the adult education process . new directions for adult and continuing education, 110, 21-27. philpott, c., & oates, c. (2017). teacher agency and professional learning communities: what can learning rounds in scotland teach us? professional development in education,43(3), 318–333. porges, s.w. (2001). the polyvagal theory: phylogenetic substrates of a social nervous system. international journal of psychophysiology, 42, 123-146. posner, j., russell, j. a. & peterson, b. s. (2005). the circumplex model of affect: an integrative approach to affective neuroscience, cognitive development, and psychopathology. development and psychopathology, 17(3), 715-734. quintana, d. s. & heathers, j. a. j. (2014). considerations in the assessment of heart rate variability in biobehavioral research. frontiers in psychology, 5, 805. doi: 10.3389/fpsyg.2014.00805 quintana, d. s., alvares, g. a. & heathers, j. a. j. (2016). guidelines for reporting articles on psychiatry and heart rate variability (graph): recommendations to advance research communication. transl. psychiatry, 6, 803. russell,j. a.(1980). a circumplex model of affect. journal of personality and social psychology,39, 1161–1178. russell, j. a. (1989). affect grid: a single-item scale of pleasure and arousal. journal of personality and social psychology,57 (3), 493-502. russell, j. a. (2003). core affect and the psychological construct of emotion. psychological review, 110(1), 145-172. russell, j. a. (2005). emotion in human consciousness is built on core affect. journal of consciousness studies,12(8-10), 26-42. russell, j. a., barrett, l. f. (1999). core affect, prototypical emotional episodes, and other things called emotion: dissecting the elephant. journal of personality and social psychology, 76(5), 805-819. scherer, k. r. (2009). the dynamic architecture of emotion: evidence for the component process model. cognition and emotion, 23,1307-1351. stein, p. k., bosner, m. s., kleiger, r. e., & conger, b. m. (1994). heart rate variability: a measure of cardiac autonomic tone. american heart journal, 127,1376–1381. storbeck, j. & maswood, r. (2015). happiness increases verbal and spatial working memory capacity where sadness does not: emotion, working memory and executive control. cognition and emotion, 30, 925–938. sung, b., & yih, j. (2015). does interest broaden or narrow attentional scope? cognition and emotion, 30(8), 1485-1494. tarvainen, m., lipponen, j., niskanen, j.-p., & ranta-aho, p.o. (2018). kubios hrv(ver. 3.1.) user’s guide. kubios oy (limited company) /www.kubios.com/ tomkins, s. s. (1962). affect, imagery, consciousness. vol. 1. the positive affects. new york: springer. vähäsantanen, k., paloniemi, s., hökkä, p., & eteläpelto, a. (2017). agentic perspective on fostering work-related learning. studies in continuing education, 39(3), 251–267. http://www.tandfonline.com/doi/full/10.1080/0158037x.2017.1310097 vähäsantanen, k., hökkä, p., paloniemi, s., herranen, s., & eteläpelto, a. (2017). professional learning and agency in an identity coaching programme. professional development in education, 43(4), 514-536. http://www.tandfonline.com/eprint/fyumgibsbiigezx5mxzp/full vähäsantanen, k., paloniemi, s., hökkä, p., & eteläpelto, a. (2017) . an agency-promoting learning arena for developing shared work practices . in m. goller, & s. paloniemi (eds.), agency at work: an agentic perspective on professional learning and development (pp. 351–371).springer: cham. doi:10.1007/978-3-319-60943-0_18 vähäsantanen, k., räikkönen, e., paloniemi, s. hökkä, p., & eteläpelto, a. (2018). a novel instrument to measure the multidimensional structure of professional agency. vocations and learning, https://doi.org/10.1007/s12186-018-9210-6 vall, b., laitila, a., borcsa, m., kykyri, v.l., karvonen, a., kaartinen, j., penttonen, m., & seikkula, j. (2018). stimulated recall interviews: how can the research interview contribute to new therapeutic practices? revista argentina de clinical psicologica, xxvii, 2, 274-293, doi: 10.24205/03276716.2018.1068 winkelman, p., & berridge, k. c. (2004). unconscious emotion. current directions in psychological science, 13, 120-123. winkler, i. (2018). identity work and emotions: a review. international journal of management reviews, 20,120–133. öhman, a., carlsson, k., lundqvist, d., & ingvar, m. (2007). on the unconscious subcortical origin of human fear. physiology and behavior, 92, 180-185. zembylas, m., & schultz, p. a. (2016). (eds .), methodological advances in research on emotion and education. springer international publishing. zwart, r.c., korthagen, f.a.j, & attema-noordewiever, s., (2015). a strength-based approach to teacher development. professional development in education, 41(3), 579–566. harteis et al publication frontline learning research vol.6 no. 2 (2018) 57 -71 issn 2295-3159 do we betray errors beforehand? the use of eye tracking, automated face recognition and computer algorithms to analyse learning from errors christian scharingera a leibniz-institut für wissensmedien, germany article received 14 may 2018 / revised 25 september/ accepted 26 september/ available online 7 december abstract during the last decade the combined recording of eye-tracking data and electroencephalographic (eeg) data has led to the methodology of fixation-related potentials analysis (frp). this methodology has been increasingly and successfully used to study eeg correlates in the time domain (i.e., event-related potentials, erps) of cognitive processing in free viewing situations like text reading or natural scene perception. basically, fixation-onset serves as time-locking event for epoching and analysing the eeg data. in this article the methodology of fixation-related frequency band power analysis (frbp) is proposed and conceptually outlined to study cognitive load and affective variations in learners during free viewing situations of multimedia learning materials (i.e., combinations of textual and pictorial elements). the eeg alpha frequency band power at parietal electrodes may serve as a valid measure of cognitive load, whereas the frontal alpha asymmetry may serve as a measure of affective variations. i will introduce and motivate the measures and the methodology and discuss methodological challenges and potential ways to overcome them. the methodology is frontline for learning research, first, as to date the eeg has been rarely used to study design effects of multimedia learning materials and second, as fixation-related eeg data analysis has rarely been done focussing on the frequency domain (i.e., frbp). despite methodological challenges still to be solved, frbp may provide a more in-depth picture of cognitive processing during multimedia learning compared to eye-tracking data or eeg data in isolation and thus may help clarifying effects of multimedia design decisions. keywords: eeg; eye-tracking; fixation-related eeg data analysis; eeg alpha frequency band power; multimedia info corresponding author mail: c.scharinger@iwm-tuebingen.de . doi: https://doi.org/10.14786/flr.v6i3.373 1. introduction there is general agreement in instructional psychology that an adequate design of multimedia learning material (i.e., combinations of text and picture) is crucial for learning success (e.g., mayer, 2009). this is because the design of multimedia learning material can alter the amount of additional, extraneous cognitive load (cl) imposed on the learner, either, in case of "good" design by freeing working memory resources, or, in case of "bad" design, by depleting working memory resources, potentially leading to an overload-situation hampering successful learning (see theoretical accounts like the cognitive load theory, sweller, van merrienboer, & paas, 1998; or the cognitive theory of multimedia learning, mayer, 2009). various multimedia design principles have been described that may alter learners' extraneous cl (mayer & fiorella, 2016; mayer & moreno, 2003). however, it still remains a matter of research of how exactly cl is influenced by certain multimedia elements and multimedia design decisions. for example, how exactly adding decorative pictures to learning materials (i.e., texts) influence cognitive (and affective) processing and consequently alter learning outcomes is still a matter of debate (for a comprehensive review see rey, 2012). such decorative pictures (or graphical elements) adjacent to textual information in multimedia learning materials that are only loosely content-related (e.g., a picture of a lightning stroke adjacent to a text describing the meteorological formation of thunderstorms) have been termed pictorial seductive details (harp & mayer, 1998). pictorial seductive details have resulted in mixed effects on learning outcomes, ranging from beneficial effects (schneider, nebel, & rey, 2016), to no-effects (park & lim, 2007), and even detrimental effects (harp & mayer, 1998; mayer & fiorella, 2016). while the beneficial effects might be explained by affective processes with the pictures positively altering the learners' motivational state (knörzer, brünken, & park, 2016; lenzner, schnotz, & müller, 2013; magner, schwonke, aleven, popescu, & renkl, 2014; schneider et al., 2016), detrimental effects might be explained by increased extraneous cl due to the additional, yet irrelevant pictorial information (mayer & fiorella, 2016; mayer & moreno, 2003), and due to effects of distraction away from and interruption of the learning process (i.e., schema construction; schneider, dyrna, meier, beege, & rey, 2017). clearly, in order to unravel potential reasons for the disparate effects of pictorial seductive details, adequate process measures (i.e., online measures) are necessary to better understand the effects of multimedia design elements on cognitive and affective processing. the electroencephalogram (eeg) and more precisely the eeg alpha and theta frequency band power might serve as such adequate, promising process measures. especially, the methodology of fixation-related eeg frequency band power analysis (frbp) might allow studying the eeg data in ecological valid task settings, that is, in free viewing situations of multimedia learning material, and may allow identifying which multimedia elements (i.e., text or picture) alter cl to what extent. in the next section, i will introduce the eeg alpha (and theta) frequency band power as potential and promising measures of cl and affective processing. i will then describe the frbp methodology on a conceptual level and address open methodological challenges as well as possible approaches to overcome them. note that the main purpose of the current article is to conceptually propose the frbp analysis as a promising, new methodological account for multimedia research by providing an overview of the eeg frequency band power as a measure of cl and affective processing and by making readers aware of potentials as well as weaknesses of the frbp methodology, yet without discussing practicalities and single methodological challenges in depth. the article may nevertheless serve as a valid and helpful primer for future research using the frbp methodology in the context of multimedia materials. 2. eeg frequency band power as a measure of cognitive load and affective processing identifying adequate process measures still is an important and general matter of debate in instructional psychology (brünken, seufert, & paas, 2010; paas, tuovinen, tabbers, & van gerven, 2003). traditionally, most research focused on outcome measures like learning success, retention, or transfer of the learned knowledge. subjective rating scales are used to assess the learners' invested effort during learning (e.g., hart & staveland, 1988; klepsch, schmitz, & seufert, 2017; paas, 1992) with effort sought to be directly related to cl and hence working memory load (schnotz & kürschner, 2007). one general drawback of subjective rating scales is that they only allow assessing cl in hindsight and for longer time periods. whether participants report an averaged impression of cl or certain load-peaks remains elusive (schmeck, opfermann, van gog, paas, & leutner, 2015). more importantly, motivational factors might confound the subjective ratings of cl (schnotz et al., 2009). consequently, the use of objective, online process measures of cl during learning have been proposed (antonenko, paas, grabner, & van gog, 2010; brünken, plass, & leutner, 2003; korbach, brünken, & park, 2017; paas et al., 2003). for example, performance in a parallel secondary task can be used to assess the current cl of the primary task (brünken et al., 2003; park & brünken, 2015). however, the secondary task may unintentionally interfere with the primary task. in contrast, physiological measures like pupil dilation or the eeg, and, more specifically, the eeg alpha (8 – 13 hz) frequency band power at parietal electrodes and the theta (4 – 6 hz) frequency band power at frontal electrodes may serve as valid measures of cl overcoming the aforementioned limitations (antonenko et al., 2010; beatty & lucero-wagoner, 2000). in addition, the frontal alpha asymmetry (faa; smith, reznik, stewart, & allen, 2017) might serve as an index of affective effects in multimedia task materials. the human eeg is typically recorded via several (e.g., 32 or 64) electrodes that are placed on the scalp at different positions consistently defined by the (extended) international 10-20 system (jasper, 1958) using electrode caps with predefined slots for the electrodes (see figure 1 for an exemplary electrode layout). after applying electrode gel to reduce the electrical impedance between the electrodes and the skin, the raw eeg can be measured. the eeg is the amplified recording of the small electrical currents generated within the brain by the summed post-synaptic electrical potentials of large neuronal assemblies consisting of several millions of neurons (pyramidal cells of the cortex) that are oriented in parallel and synchronously active (for reviews see cohen, 2017; jackson & bolger, 2014; olejniczak, 2006). the typical sampling rate of an eeg recording device is 500 or 1000 hz (i.e., one measurement point every two or each millisecond, respectively). the eeg thus reflects the oscillatory activity of specific neuronal populations with excellent time resolution, yet the spatial resolution is rather low (in the range of centimeter; olejniczak, 2006). figure 1. schematic head (nose up) with exemplary electrode layout. highlighted are representative electrode positions for measuring cl (blue color) or affective processing (red color). note. electrodes outside the head are due to the projection of 3d locations onto 2d space. after some data preprocessing steps (e.g., filtering, artefact removal; for guidelines see picton, 2000), the recorded eeg data can be analyzed in the time domain and/or in the frequency domain. for analyses in the time domain, parts of the eeg (i.e., epochs) time-locked to certain events (e.g., stimulus-onsets) are averaged across several trials (i.e., repetitions of a specific event) to increase the signal-to-noise ratio. this procedure results in so-called event-related potential (erp) curves, that is, positive and negative deflections (i.e., components) that have been linked quite specifically to certain cognitive processes (for a comprehensive review see münte, urbach, düzel, & kutas, 2000). for analyses in the frequency domain, the spectrum is calculated for the eeg epochs of interest (e.g., using fast-fourier transforms, fft, or wavelet analysis; cohen, 2014). this results in power values for the different frequencies contained in the eeg signal (and also phase information, which however is beyond the scope of the current article; the interested reader may refer to bastiaansen, mazaheri, & jensen, 2012; cohen, 2014). the more neurons are synchronously active (i.e., 'firing') at a specific frequency in response to an event, the higher is the measured power (i.e., amplitude to the square) at this specific frequency. therefore, an increase in frequency band power after an event is generally termed event-related synchronization (ers), whereas a decrease in frequency band power is termed event-related desynchronization (erd; pfurtscheller & lopes da silva, 1999). the amount of change in erd or ers in relation to an event can be expressed as percentage of change using the erd/ers%-formula given in pfurtscheller & lopes da silva (1999; see also antonenko et al., 2010). five frequency bands have traditionally been differentiated in the eeg reflecting functionally different oscillatory neuronal activity: slow oscillatory activity in the delta (< 3 hz) range, oscillatory activity in the theta (4 – 6 hz), alpha (8 – 13 hz), and beta (13 – 24 hz) range, and fast oscillatory activity in the gamma (> 40 hz) range (for reviews see bastiaansen et al., 2012; krause, 2003). in the context of multimedia research oscillatory activity (and hence, power) in the theta and alpha frequency band is of specific interest as will be detailed below. slow (delta) and fast (gamma) oscillatory activity are rather prone to artefacts (e.g., slow drifts or muscle activity) requiring highly controlled lab settings and rather simple, basic tasks and task materials (see, e.g., bastiaansen et al., 2012), thus being rather inadequate for multimedia research. oscillatory activity in the beta band has mainly been attributed to reflect activity of the motor cortex (pfurtscheller, zalaudek, & neuper, 1998), but might also be interesting for studying cognitive processes (engel & fries, 2010). importantly, the eeg alpha and theta frequency band power might be used as reliable process measures of cl in multimedia research. it has been consistently observed that the eeg alpha frequency band power at parietal electrodes decreases for increasing cl whereas the frontal-central eeg theta frequency band power increases for increasing cl (gevins & smith, 2000; kretzschmar et al., 2013; palomäki, kivikangas, alafuzoff, hakala, & krause, 2012; pesonen, hämäläinen, & krause, 2007; scharinger, soutschek, schubert, & gerjets, 2015; 2017). functionally, the alpha erd (i.e., the decreasing alpha power) has been interpreted to index cortical activity related to attentional and memory processing (klimesch, 1999; krause, 2003). according to klimesch (1999) the alpha frequency band might be functionally divided further in an upper alpha (10 – 13 hz) frequency band that might mainly be related to (semantic) memory processing and a lower alpha (8 – 10 hz) frequency band that might mainly be related to attentional processes. neural activity in the eeg theta frequency band has been associated predominantly with processes of working memory and cognitive control (itthipuripat, wessel, & aron, 2013; nigbur, ivanova, & stürmer, 2011; sauseng et al., 2002; sauseng, griesmayr, freunberger, & klimesch, 2010). note however, that the functional relationship between eeg frequency band power and specific cognitive processes is not as clear and well established as the relationship between certain erp components of the eeg and cognitive processes (krause, 2003). however, in contrast to erps that require a highly controlled, artificial task environment, especially eeg alpha frequency band power has been shown to reliably index cl also in complex task materials like those of instructional psychology (antonenko & niederhauser, 2010; gerlic & jausovec, 2001; scharinger, kammerer, & gerjets, 2015; 2016). importantly, changes in the eeg frequency band power index fluctuations of cl with high temporal acuity. furthermore, apart from indexing cl the eeg alpha frequency band power might also be used to assess emotional and motivational aspects of stimulus-processing. the frontal alpha asymmetry (faa) may serve as such a measure. the faa is a relational measure of the eeg alpha (8 – 13 hz) frequency band power over the left-frontal hemisphere compared to the right-frontal hemisphere, commonly calculated as the difference between corresponding electrodes (e.g., f4 minus f3; smith, reznik, stewart, & allen, 2017). the faa has its origin in cortical activity of the prefrontal cortex. the prefrontal cortex plays an important role in affective processing and emotion regulation showing specific left and right hemispheric lateralization effects for emotions and affective stimulus content (demaree, everhart, youngstrom, & harrison, 2005; dixon, thiruchselvam, todd, & christoff, 2017). although still matter of debate, greater left than right hemispheric cortical activity has been associated with approach motivation and hence predominantly emotionally positively connoted stimuli, whereas greater right hemispheric cortical activity has been associated with withdrawal motivation and hence predominantly emotionally negatively connoted stimuli (ahern & schwartz, 1985; harmon-jones, gable, & peterson, 2010). most eeg studies so far assessed the faa during rest conditions after emotion induction procedures (e.g., coan & allen, 2004). yet, recent research indicated that the faa might be also used during task performance when emotional stimuli are presented for a brief period of time (schöne, schomberg, gruber, & quirin, 2016; weinreich, stephani, & schubert, 2016). thus, potentially, the faa might be used to study affective aspects of multimedia. whether the faa can be used in such a way for complex multimedia materials has however to be studied further. the methodology of fixation-related eeg data analysis may allow analyzing the eeg frequency band power in free viewing situations of multimedia task materials when certain elements (i.e., areas of interest, aois) are fixated and thus may allow to assess how cl (or in case of the faa, affective processing) is altered by specific multimedia elements like pictorial seductive details. in the next section, i will give a brief overview on the methodology, then conceptually pointing out methodological challenges that one has to be aware of. 3. fixation-related eeg data analysis eye-tracking is increasingly used in research on instructional multimedia materials to study underlying cognitive processes during learning (e.g., eitel, scheiter, & schüler, 2012; hyönä, 2010; jarodzka, holmqvist, & gruber, 2017; mayer, 2010; schüler, 2017; van gog & scheiter, 2010; for a recent overview on the literature see alemdag & cagiltay, 2018). eye-tracking data shows the movement of the eyes, basically differentiating between saccades (i.e., the eyes quickly moving) to elements of the visual scenery and fixations (i.e., the eyes practically at rest) on elements of the visual scenery (for a comprehensive discussion of different eye-tracking patterns see holmqvist & andersson, 2017). it is generally agreed on, that during saccades the visual information intake is not possible (e.g., kok & jarodzka, 2017a), whereas it is possible (and most of the time takes place) during fixations. it has been shown that the fixation patterns vary depending on the task materials and the given task (e.g., in picture viewing, yarbus, 1967, or in reading, strukelj & niehorster, 2018; for reviews see kowler, 2011; rayner, 1998; 2009), indicating a plausible link between fixations and cognitive processing (just & carpenter, 1976; 1980). note however, this link might not always be straightforward. visual attention might be slightly ahead of what is fixated (depending on the task, up to 250 ms; deubel, 2008). in addition, due to peripheral viewing especially in reading more elements than those fixated at might be processed (baccino & manunta, 2005; dimigen, kliegl, & sommer, 2012; rayner, 2009). on the contrary, what is fixated at might not always be (consciously) processed. this might be the case during periods of mind-wandering (foulsham, farley, & kingstone, 2013), or if visual elements are not task relevant (as indicated in change blindness paradigms; e.g., triesch, ballard, hayhoe, & sullivan, 2003). nevertheless, fixations have been used as a valid proxy reflecting the individuals' structure of the information intake (and procesing) in free viewing situations and hence as triggers for defining epochs in time for which the eeg data can be analysed. in free viewing or free reading situations no specific stimulus-onset exists. however, the fixation-onset can be used as the event for which the eeg data is time-locked, epoched, and analysed to (depending on the concrete research question, some studies also used saccade-onset as time-locking event; cf. dimigen, sommer, hohlfeld, jacobs, & kliegl, 2011). most studies so far have studied the eeg in the time domain, that is, fixation-related potentials (frps), for example in reading research (dimigen, et al., 2011; frey, lemaire, vercueil, & guérin-dugué, 2018; henderson, luke, schmidt, & richards, 2013; hutzler et al., 2007; kliegl, dambacher, dimigen, & sommer, 2014; kornrumpf, dimigen, & sommer, 2017; kornrumpf, niefind, sommer, & dimigen, 2016; léger et al., 2014; niefind & dimigen, 2016; weiss, knakker, & vidnyánszky, 2016), natural scene perception (giannini, alexander, nikolaev, & van leeuwen, 2018; simola, le fevre, torniainen, & baccino, 2015; simola, torniainen, moisala, kivikangas, & krause, 2013), visual search (brouwer, hogervorst, oudejans, ries, & touryan, 2017; kamienkowski, ison, quiroga, & sigman, 2012; kaunitz et al., 2014; ries, touryan, ahrens, & connolly, 2016; winslow et al., 2010), decision making (frey et al., 2013), and human-computer interaction (léger et al., 2014). despite methodological challenges (see below) these studies concurrently report frp-effects comparable to classical erp-effects (e.g., n400-like effects for word predictability; dimigen et al., 2011), thus validating the basic feasibility and meaningfulness of fixation-related eeg data analyses. only few studies so far have analysed the eeg data in the frequency domain using frbp. for example, scharinger and colleagues (scharinger et al., 2015, experiment 1) used frbp to study cl during hypertext-like reading, comparing several parts of a text that defined two different areas of interest (aois): aois of parts of the text where participants simply read and aois of parts of the text where participants additionally had to perform hyperlink-like selection processes. as hypothesized, the cl was higher when participants had to perform hyperlink-like selection processes in addition to purely text reading, indicated by decreased eeg alpha frequency band power. interestingly, this result was also confirmed by the pupil dilation data, with the pupil showing a larger diameter (i.e., higher cl) for parts of the text with hyperlink-like selection processes as compared to purely text reading. in a second experiment (scharinger et al., 2015, experiment 2) the results could be replicated using classical response-locked eeg data analysis instead of fixation-related eeg data analysis, thus underlining the validity of the fixation-related eeg data analysis methodology. another study by scharinger and colleagues (scharinger et al., 2016) used frbp to study cl (as indexed by the parietal eeg alpha frequency band power) in free viewing and evaluating of search engine result pages. this study indicated that a perfect hit (i.e., a semantically and lexically matching search result for a specific search query) resulted already during initial fixations in decreased eeg alpha frequency band power as compared to search results that are no, or no perfect matches for a given search query. this has been interpreted as indicating that the best hit is recognized early in time and potentially then more thoroughly processed (resulting in increased cl) as compared to other, less fitting search results. finally, vignali and colleagues (vignali, himmelstoss, hawelka, richlan, & hutzler, 2016) used frbp to study semantic violations in sentences in free reading situations. they observed decreased lower-beta band (13-18 hz) power for fixations of semantically unrelated words as compared to semantically related words, also indicating higher cl. to sum up, studies so far indicated the principal feasibility and validity of the methodology of fixation-related eeg data analysis. combining eye-tracking and the eeg in research on instructional multimedia materials seems to be promising for two reasons. first, it allows to study the eeg during learning with multimedia task materials in task settings of high ecological validity (i.e., 'realistic' task materials, with texts and pictures presented simultaneously on one screen). without this methodology, text and pictures would have to be presented separately in time (i.e., in rather artificial sequences) to create stimulus-onsets for the eeg data analysis. second, the eeg data might be used as an additional measure for triangulating the meaning of observed eye-movement patterns as has been proposed for verbal data (kok & jarodzka, 2017b). for example, the eeg might help differentiating whether longer fixations might indicate increased cl (könig et al., 2016; reichle & reingold, 2013) or individuals' expertise (bertram, helle, kaakinen, & svedström, 2013; reingold & sheridan, 2011), or it might help differentiating whether learners are still working on a (difficult) task or mind-wandering (foulsham et al., 2013). yet, some methodological challenges remain that one has to be aware of. 4. methodological challenges of fixation-related eeg data analysis and potential solutions there are some challenges that one has to be aware of when analysing fixation-related eeg data. first, the eeg and eye-tracking data has to be synchronized. this can be done by regularly sending triggers (e.g., each second) during data recording to both recording devices (i.e., the eeg and the eye-tracker). based on these triggers the two data streams can then be synchronized offline using for example the toolbox eeglab (delorme & makeig, 2004) with the eye-eeg plugin (dimigen et al., 2011). the synchronisation of the eeg and eye-tracking data includes matching the sampling rates of both data streams (which is for the eeg typically at 500 hz or 1000 hz and for current remote eye-tracking devices at 120 hz or 250 hz), either by upor down-sampling. once the eeg data and the eye-tracking data are synchronized there remain several challenges that have to be taken into consideration. these include the correction of eye-movement artefacts, dealing with overlapping eeg data segments due to different and rather short fixation durations, and the selection of an adequate baseline (for a comprehensive discussion of these challenges and potential solutions the interested reader may refer to baccino, 2011; dimigen et al., 2011; nikolaev et al., 2016; the purpose of the current article is to make readers aware of potentials as well as weaknesses of the methodology, yet without discussing single methodological challenges in depth). eye-movements (i.e., saccades and blinks) alter the electric fields around the eyes and consequently confound the raw eeg, especially at frontal electrodes (iwasaki et al., 2005). thus, eye-movement artefacts may either mask the eeg correlates of interest or, worse, may induce a systematic error. for example, in multimedia the eye-movement patterns vary between viewing of textual and pictorial elements. thus, when interested in fixation-related eeg data for multimedia, that is, when comparing eeg data for text reading and picture viewing, the eeg may be confounded by different eye-movement patterns. however, several methodologies exist to clean the raw eeg data from eye-movement artefacts (for reviews see croft & barry, 2000; islam, rastegarnia, & yang, 2016). for example, independent component analysis (ica) can be used to reliably identify and correct for eye-movement artefacts (chaumon, bishop, & busch, 2015; delorme, sejnowski, & makeig, 2007; jung et al., 2000; zhou & gotman, 2009). noteworthy, especially in the context of fixation-related eeg data analysis the use of ica has been shown to result in adequately cleaned eeg, outperforming other methodologies like standard regression-based data cleaning (henderson et al., 2013; hutzler et al., 2007). while a variety of methodologies exists for cleaning eeg data from eye-movement artefacts, the varying and rather short length of fixations is another challenge for fixation-related eeg analysis that one has to be aware of. for example, during reading typical fixations last on average between 200-250 ms (e.g., dimigen et al., 2011). this is problematic when analysing the eeg data in the time domain as later components in the frp of a current fixation (e.g., the n400) might be overlapped (i.e., confounded) by early components of the frp of the following fixation. in frp analysis very short fixations are therefore excluded from analysis (e.g., fixations < 80 ms; frey et al., 2018). several statistical methods (e.g., regression-based models) have been proposed to deal with the potentially confounding effect of overlapping eeg data segments (baccino, 2011; dimigen et al., 2011; frey et al., 2018; nikolaev, pannasch, ito, & belopolsky, 2014). it has also been proposed to compare only those data sequences of comparable fixation-patterns between task conditions of interest to minimize potentially confounding effects due to different fixation durations and hence differently overlapping eeg segments (nikolaev et al., 2016). while this proceeding might be feasible for task conditions that are quite comparable (e.g., both within the domain of reading), for multimedia task materials with texts and pictures it might be impossible to adequately match the fixations used for analysis due to the different fixation patterns for reading and picture viewing. moreover, in frbp analysis overlapping eeg data epochs due to short fixation durations might frequently occur. this is because, the length of the eeg data epoch defines the possible frequency resolution of the calculated spectrum. for example, in the theta frequency range (4 – 6 hz) one oscillation lasts at minimum 250 ms. thus, when interested in the theta frequency band power an analysis window of at least 250 ms would be necessary. commonly, an analysis window including several oscillations of the specific frequency of interest is recommended for calculating the spectrum (i.e., an analysis window of at least 500 ms length). consequently, one has to be aware of the constraint that frbp analysis is seldom suitable for analysing the eeg data for single fixations (unless their duration is long enough). yet for research on multimedia task materials interested in cl this constraint might not be of too much relevance, as the aois of interest (i.e., parts of the text versus the pictures) might generally comprise several fixations that one could summarize, resulting in eeg data epochs long enough for analysis (see figure 2). importantly, one still would be able to differentiate between first and later visits of an aoi. nevertheless, a comparison of the eeg in classical sequential stimulus presentation paradigms (e.g., in multimedia research by presenting text and pictures sequentially) with fixation-related paradigms might be necessary to further validate the reliability of the frbp analysis when used in new experimental paradigms or for new, complex multimedia task materials. figure 2. left part: schematic illustration of the problem of overlapping eeg data segments (marked by the lightning symbol) in case of short fixations and a frbp analysis based on single fixations (i.e., using very small aois). right part: schematic illustration of eeg data segments aligned to larger aois which would be typically used in multimedia research. longer eeg segments could be used without overlapping. note. the eeg and eye-tracking data has been artificially combined and does not show data of a real person. another, however easily addressable challenge of overlapping eeg segments in fixation-related eeg data analysis are specific eeg correlates at the very beginning of a stimulus (e.g., when the text for reading is shown on the screen for the first time). to avoid stimulus-presentation associated event-related eeg correlates potentially masking fixation-related eeg correlates, the first 700 ms of stimulus presentation (i.e., the first few fixations) should be excluded from eeg data analysis (dimigen et al., 2011). finally, defining an adequate baseline for eeg data analysis is not trivial in free viewing situations (dimigen et al., 2011; nikolaev et al., 2016). in classical eeg data analysis typically a pre-stimulus baseline is used for baseline correction in order to reduce slow currency drifts (e.g., due to fatigue) in the eeg epochs used for analysis (picton, 2000). as in fixation-related eeg data analysis the pre-fixation baseline is contaminated by saccadic activity, it has been proposed to alternatively use a short time-frame directly after fixation-onset as baseline (e.g., baccino, 2011) or to use a global, pre-stimulus baseline (i.e., a baseline at the beginning of the task; nikolaev et al., 2016). to date, there is no clear recommendation of what time interval is best suited as baseline in fixation-related eeg data analysis. choosing an adequate baseline may largely depend on the concrete task design. potentially, in frbp it might also be possible to report absolute power values (i.e., to avoid using a baseline) if the experimental conditions (i.e., the aois for which the eeg data is analyzed and compared) are fully permutated with respect to spatial and temporal positions (i.e., when spatial or timing issues can be excluded as potential confounds of the data). 5. conclusions despite the remaining challenges of using the methodology of fixation-related eeg data analysis, the methodology clearly is at the frontline of learning research. eeg (alpha and theta) frequency band power may help gaining deeper insight in the cognitive processing of multimedia elements (i.e., text and picture and resulting cl or affective effects). thus, frbp may substantially contribute to a better understanding of multimedia design effects like pictorial seductive details and consequently may foster better instructional design. keypoints eeg alpha frequency band power allows assessing cognitive load (and potentially affective effects) during learning with multimedia task materials. the methodology of fixation-related frequency band power analysis (frbp) allows studying the eeg frequency band power in free viewing situations. frbp thus allows comparing cognitive load when different multimedia elements are fixated (e.g., text versus picture). the methodology may thus provide a deeper understanding of multimedia design effects like the pictorial seductive detail effect. references ahern, g. l., & schwartz, g. e. (1985). differential lateralization for positive and negative emotion in the human brain: eeg spectral analysis. neuropsychologia, 23(6), 745–755. https://doi.org/10.1016/0028-3932(85)90081-8 alemdag, e., & cagiltay, k. (2018). a systematic review of eye tracking research on multimedia learning. computers & education, 125(july), 413–428. https://doi.org/10.1016/j.compedu.2018.06.023 antonenko, p., & niederhauser, d. s. (2010). the influence of leads on cognitive load and learning in a hypertext environment. computers in human behavior, 26(2), 140–150. https://doi.org/10.1016/j.chb.2009.10.014 antonenko, p., paas, f., grabner, r., & van gog, t. (2010). using electroencephalography to measure cognitive load. educational psychology review, 22(4), 425–438. https://doi.org/10.1007/s10648-010-9130-y baccino, t. (2011). eye movements and concurrent event-related potentials’: eye fixation-related potential investigations in reading. in s. liversedge, i. gilchrist, & s. everling (eds.), oxford handbook of eye movements(pp. 857–870). oxford, uk: oxford university press. baccino, t., & manunta, y. (2005). eye-fixation-related potentials: insight into parafoveal processing. journal of psychophysiology, 19(3), 204–215. https://doi.org/10.1027/0269-8803.19.3.204 bastiaansen, m., mazaheri, a., & jensen, o. (2012). beyond erps: oscillatory neuronal dynamics. in s. j. luck & e. s. kappenman (eds.), the oxford handbook of event-related potential components(pp. 31–50). oxford, uk: oxford university press. beatty, j., & lucero-wagoner, b. (2000). the pupillary system. in j. t. cacioppo, l. g. tassinary, & g. berndtson (eds.), handbook of psychophysiology(2nd ed., pp. 142–162). cambridge, uk: cambridge university press. bertram, r., helle, l., kaakinen, j. k., & svedström, e. (2013). the effect of expertise on eye movement behaviour in medical image perception. plos one, 8(6), e66169. https://doi.org/10.1371/journal.pone.0066169 brouwer, a.-m., hogervorst, m. a., oudejans, b., ries, a. j., & touryan, j. (2017). eeg and eye tracking signatures of target encoding during structured visual search. frontiers in human neuroscience, 11, 1–11. https://doi.org/10.3389/fnhum.2017.00264 brünken, r., plass, j. l., & leutner, d. (2003). direct measurement of cognitive load in multimedia learning. educational psychologist, 38(1), 53–61. https://doi.org/10.1207/s15326985ep3801_7 brünken, r., seufert, t., & paas, f. (2010). measuring cognitive load. in j. l. plass, r. moreno, & r. brünken (eds.), cognitive load theory(pp. 181–202). cambridge: cambridge university press. https://doi.org/10.1007/978-1-4419-8126-4_6 chaumon, m., bishop, d. v. m., & busch, n. a. (2015). a practical guide to the selection of independent components of the electroencephalogram for artifact correction. journal of neuroscience methods, 250 , 47–63. https://doi.org/10.1016/j.jneumeth.2015.02.025 coan, j. a, & allen, j. j. b. (2004). frontal eeg asymmetry as a moderator and mediator of emotion. biological psychology, 67(1–2), 7–49. https://doi.org/10.1016/j.biopsycho.2004.03.002 cohen, m. x. (2014). analyzing neural time series data: theory and practice.cambridge, ma: mit press. https://doi.org/10.1007/s13398-014-0173-7.2 cohen, m. x. (2017). where does eeg come from and what does it mean? trends in neurosciences, 40(4), 208–218. https://doi.org/10.1016/j.tins.2017.02.004 croft, r. j., & barry, r. j. (2000). removal of ocular artifact from the eeg: a review. neurophysiologie clinique/clinical neurophysiology, 30 (1), 5–19. https://doi.org/10.1016/s0987-7053(00)00055-1 delorme, a., & makeig, s. (2004). eeglab: an open source toolbox for analysis of single-trial eeg dynamics including independent component analysis. journal of neuroscience methods, 134(1), 9–21. https://doi.org/10.1016/j.jneumeth.2003.10.009 delorme, a., sejnowski, t., & makeig, s. (2007). enhanced detection of artifacts in eeg data using higher-order statistics and independent component analysis. neuroimage,34(4), 1443–1449. https://doi.org/10.1016/j.neuroimage.2006.11.004 demaree, h. a., everhart, d. e., youngstrom, e. a., & harrison, d. w. (2005). brain lateralization of emotional processing: historical roots and a future incorporating "dominance". behavioral and cognitive neuroscience reviews, 4(1), 3–20. https://doi.org/10.1177/1534582305276837 deubel, h. (2008). the time course of presaccadic attention shifts. psychological research,72(6), 630–640. https://doi.org/10.1007/s00426-008-0165-3 deubel, h., & schneider, w. x. (1996). saccade target selection and object recognition: evidence for a common attentional mechanism. vision research, 36(12), 1827–1837. https://doi.org/10.1016/0042-6989(95)00294-4 dimigen, o., kliegl, r., & sommer, w. (2012). trans-saccadic parafoveal preview benefits in fluent reading: a study with fixation-related brain potentials. neuroimage, 62(1), 381–393. https://doi.org/10.1016/j.neuroimage.2012.04.006 dimigen, o., sommer, w., hohlfeld, a., jacobs, a. m., & kliegl, r. (2011). coregistration of eye movements and eeg in natural reading: analyses and review. journal of experimental psychology. general, 140(4), 552–572. https://doi.org/10.1037/a0023885 dixon, m. l., thiruchselvam, r., todd, r., & christoff, k. (2017). emotion and the prefrontal cortex: an integrative review. psychological bulletin, [epub], 1–61. https://doi.org/10.1037/bul0000096 eitel, a., scheiter, k., & schüler, a. (2012). the time course of information extraction from instructional diagrams. perceptual and motor skills, 115(3), 677–701. https://doi.org/10.2466/22.23.pms.115.6.677-701 engel, a. k., & fries, p. (2010). beta-band oscillations signalling the status quo? current opinion in neurobiology, 20(2), 156–165. https://doi.org/10.1016/j.conb.2010.02.015 foulsham, t., farley, j., & kingstone, a. (2013). mind wandering in sentence reading: decoupling the link between mind and eye. canadian journal of experimental psychology/revue canadienne de psychologie expérimentale , 67(1), 51–59. https://doi.org/10.1037/a0030217 frey, a., ionescu, g., lemaire, b., lópez-orozco, f., baccino, t., & guérin-dugué, a. (2013). decision-making in information seeking on texts: an eye-fixation-related potentials investigation. frontiers in systems neuroscience, 7(39). https://doi.org/10.3389/fnsys.2013.00039 frey, a., lemaire, b., vercueil, l., & guérin-dugué, a. (2018). an eye fixation-related potential study in two reading tasks: reading to memorize and reading to make a decision. brain topography, 1–21. https://doi.org/10.1007/s10548-018-0629-8 gerlic, i., & jausovec, n. (2001). differences in eeg power and coherence measures related to the type of presentation: text versus multimedia. journal of educational computing research, 25 (2), 177–195. http://dx.doi.org/10.2190/ydwy-u3fj-4ly4-lynd gevins, a., & smith, m. e. (2000). neurophysiological measures of working memory and individual differences in cognitive ability and cognitive style. cerebral cortex, 10(9), 829–839. https://doi.org/10.1093/cercor/10.9.829 giannini, m., alexander, d. m., nikolaev, a. r., & van leeuwen, c. (2018). large-scale traveling waves in eeg activity following eye movement. brain topography, 1–15. https://doi.org/10.1007/s10548-018-0622-2 harmon-jones, e., gable, p. a., & peterson, c. k. (2010). the role of asymmetric frontal cortical activity in emotion-related phenomena: a review and update. biological psychology, 84(3), 451–462. https://doi.org/10.1016/j.biopsycho.2009.08.010 harp, s. f., & mayer, r. e. (1998). how seductive details do their damage: a theory of cognitive interest in science learning. journal of educational psychology,90(3), 414–434. https://doi.org/10.1037/0022-0663.90.3.414 hart, s. g., & staveland, l. e. (1988). development of nasa-tlx (task load index): results of empirical and theoretical research. in p. a. hancock & n. meshkati (eds.), human mental workload(pp. 139–183). amsterdam, nl. https://doi.org/10.1016/s0166-4115(08)62386-9 henderson, j. m., luke, s. g., schmidt, j., & richards, j. e. (2013). co-registration of eye movements and event-related potentials in connected-text paragraph reading. frontiers in systems neuroscience, 7(28). https://doi.org/10.3389/fnsys.2013.00028 holmqvist, k., & andersson, r. (2017). eye tracking: a comprehensive guide to methods, paradigms, and measures . lund, sweden: lund eye-tracking research institute. hutzler, f., braun, m., võ, m. l.-h., engl, v., hofmann, m., dambacher, m., leder, h., & jacobs, a. m. (2007). welcome to the real world: validating fixation-related brain potentials for ecologically valid settings. brain research, 1172, 124–9. https://doi.org/10.1016/j.brainres.2007.07.025 hyönä, j. (2010). the use of eye movements in the study of multimedia learning. learning and instruction, 20(2), 172–176. https://doi.org/10.1016/j.learninstruc.2009.02.013 islam, m. k., rastegarnia, a., & yang, z. (2016). methods for artifact detection and removal from scalp eeg: a review. neurophysiologie clinique/clinical neurophysiology, 46 (4–5), 287–305. https://doi.org/10.1016/j.neucli.2016.07.002 itthipuripat, s., wessel, j. r., & aron, a. r. (2013). frontal theta is a signature of successful working memory manipulation. experimental brain research, 224(2), 255–262. https://doi.org/10.1007/s00221-012-3305-3 iwasaki, m., kellinghaus, c., alexopoulos, a. v., burgess, r. c., kumar, a. n., han, y. h., lüders, h. o., & leigh, r. j. (2005). effects of eyelid closure, blinks, and eye movements on the electroencephalogram. clinical neurophysiology, 116(4), 878–885. https://doi.org/10.1016/j.clinph.2004.11.001 jackson, a. f., & bolger, d. j. (2014). the neurophysiological bases of eeg and eeg measurement: a review for the rest of us. psychophysiology, 51(11), 1061–1071. https://doi.org/10.1111/psyp.12283 jarodzka, h., holmqvist, k., & gruber, h. (2017). eye tracking in educational science : theoretical frameworks and research agendas. journal of eye movement research, 10(1), 1–18. https://doi.org/10.16910/jemr.10.1.3 jasper, h. h. (1958). the ten-twenty electrode system of the international federation. electroencephalography and clinical neurophysiology, 10, 371–375. jung, t., makeig, s., humphries, c., lee, t., mckeown, m. j., iragui, i., & sejnowski, t. j. (2000). removing electroencephalographic aretfacts by blind source seperation. psychophysiology,37(2), 163–178. https://doi.org/10.1111/1469-8986.3720163 just, m. a., & carpenter, p. a. (1976). eye fixations and cognitive processes. cognitive psychology, 8(4), 441–480. https://doi.org/10.1016/0010-0285(76)90015-3 just, m. a., & carpenter, p. a. (1980). a theory of reading: from eye fixations ot comprehension. psychological review, 87(4), 329–354. https://doi.org/10.1037/0033-295x.87.4.329 kamienkowski, j. e., ison, m. j., quiroga, r. q., & sigman, m. (2012). fixation-related potentials in visual search: a combined eeg and eye tracking study. journal of vision, 12(7), 4–4. https://doi.org/10.1167/12.7.4 kaunitz, l. n., kamienkowski, j. e., varatharajah, a., sigman, m., quiroga, r. q., & ison, m. j. (2014). looking for a face in the crowd: fixation-related potentials in an eye-movement visual search task. neuroimage, 89, 297–305. https://doi.org/10.1016/j.neuroimage.2013.12.006 klepsch, m., schmitz, f., & seufert, t. (2017). development and validation of two instruments measuring intrinsic, extraneous, and germane cognitive load. frontiers in psychology, 8, 1–18. https://doi.org/10.3389/fpsyg.2017.01997 kliegl, r., dambacher, m., dimigen, o., & sommer, w. (2014). oculomotor control, brain potentials, and timelines of word recognition during natural reading. in current trends in eye tracking research(pp. 141–155). cham: springer international publishing. https://doi.org/10.1007/978-3-319-02868-2_10 klimesch, w. (1999). eeg alpha and theta oscillations reflect cognitive and memory performance: a review and analysis. brain research reviews, 29(2–3), 169–195. https://doi.org/10.1016/s0165-0173(98)00056-3 knörzer, l., brünken, r., & park, b. (2016). facilitators or suppressors: effects of experimentally induced emotions on multimedia learning. learning and instruction, 44, 97–107. https://doi.org/10.1016/j.learninstruc.2016.04.002 könig, p., wilming, n., kietzmann, t. c., ossandon, j. p., onat, s., ehinger, b., gameiro r. r., & kaspar, k. (2016). eye movements as a window to cognitive processes. journal of eye movement research, 9(5), 1–16. https://doi.org/10.16910/jemr.9.5.3 kok, e. m., & jarodzka, h. (2017a). before your very eyes: the value and limitations of eye tracking in medical education. medical education, 51(1), 114–122. https://doi.org/10.1111/medu.13066 kok, e. m., & jarodzka, h. (2017b). beyond your very eyes: eye movements are necessary, not sufficient. medical education, 51(11), 1190–1190. https://doi.org/10.1111/medu.13384 korbach, a., brünken, r., & park, b. (2017). measurement of cognitive load in multimedia learning: a comparison of different objective measures. instructional science. https://doi.org/10.1007/s11251-017-9413-5 kornrumpf, b., dimigen, o., & sommer, w. (2017). lateralization of posterior alpha eeg reflects the distribution of spatial attention during saccadic reading. psychophysiology, 54(6), 809–823. https://doi.org/10.1111/psyp.12849 kornrumpf, b., niefind, f., sommer, w., & dimigen, o. (2016). neural correlates of word recognition: a systematic comparison of natural reading and rapid serial visual presentation. journal of cognitive neuroscience, 28(9), 1374–1391. https://doi.org/10.1162/jocn_a_00977 kowler, e. (2011). eye movements: the past 25 years. vision research, 51(13), 1457–1483. https://doi.org/10.1016/j.visres.2010.12.014 krause, c. m. (2003). brain electric oscillations and cognitive processes. in k. hugdahl (ed.), neuropsychology and cognition. experimental methods in neuropsychology (21st ed., pp. 111–130). boston, ma: kluwer academic publishers group. kretzschmar, f., pleimling, d., hosemann, j., füssel, s., bornkessel-schlesewsky, i., & schlesewsky, m. (2013). subjective impressions do not mirror online reading effort: concurrent eeg-eyetracking evidence from the reading of books and digital media. plos one, 8(2), e56178. https://doi.org/10.1371/journal.pone.0056178 léger, p. m., titah, r., sénecal, s., fredette, m., courtemanche, f., labonte-lemoyne, él., & de guinea, a. o. (2014). precision is in the eye of the beholder: application of eye fixation-related potentials to information systems research. journal of the association for information systems, 15, 651–678. http://dx.doi.org/10.17705/1jais.00376 lenzner, a., schnotz, w., & müller, a. (2013). the role of decorative pictures in learning. instructional science, 41(5), 811–831. https://doi.org/10.1007/s11251-012-9256-z magner, u. i. e., schwonke, r., aleven, v., popescu, o., & renkl, a. (2014). triggering situational interest by decorative illustrations both fosters and hinders learning in computer-based learning environments. learning and instruction,29, 141–152. https://doi.org/10.1016/j.learninstruc.2012.07.002 mayer, r. e. (2009). multimedia learning(2nd ed.). new york, ny: cambridge university press. mayer, r. e. (2010). unique contributions of eye-tracking research to the study of learning with graphics. learning and instruction, 20(2), 167–171. https://doi.org/10.1016/j.learninstruc.2009.02.012 mayer, r. e., & fiorella, l. (2016). principles for reducing extraneous processing in multimedia learning: coherence, signaling, redundancy, spatial contiguity, and temporal contiguity principles. in r. mayer (ed.), the cambridge handbook of multimedia learning(pp. 279–315). cambridge: cambridge university press. https://doi.org/10.1017/cbo9781139547369.015 mayer, r. e., & moreno, r. (2003). nine ways to reduce cognitive load in multimedia learning.educational psychologist, 38(1), 43–52. https://doi.org/10.1207/s15326985ep3801_6 münte, t. f., urbach, t. p., düzel, e., & kutas, m. (2000). event-related brain potentials in the study of human cognition and neuropsychology. handbook of neuropsychology, 1, 1–97. niefind, f., & dimigen, o. (2016). dissociating parafoveal preview benefit and parafovea-on-fovea effects during reading: a combined eye tracking and eeg study. psychophysiology, 53(12), 1784–1798. https://doi.org/10.1111/psyp.12765 nigbur, r., ivanova, g., & stürmer, b. (2011). theta power as a marker for cognitive interference. clinical neurophysiology, 122 (11), 2185–2194. https://doi.org/10.1016/j.clinph.2011.03.030 nikolaev, a. r., meghanathan, r. n., & van leeuwen, c. (2016). combining eeg and eye movement recording in free viewing: pitfalls and possibilities. brain and cognition, 107, 55–83. https://doi.org/10.1016/j.bandc.2016.06.004 nikolaev, a. r., pannasch, s., ito, j., & belopolsky, a. v. (2014). eye movement-related brain activity during perceptual and cognitive processing. frontiers in systems neuroscience(8). https://doi.org/10.3389/fnsys.2014.00062 olejniczak, p. (2006). neurophysiologic basis of eeg. journal of clinical neurophysiology,23(3), 186–189. https://doi.org/10.1097/01.wnp.0000220079.61973.6c paas, f. g. (1992). training strategies for attaining transfer of problem-solving skill in statistics: a cognitive-load approach. journal of educational psychology,84(4), 429–434. https://doi.org/10.1037/0022-0663.84.4.429 paas, f., tuovinen, j. e., tabbers, h., & van gerven, p. w. m. (2003). cognitive load measurement as a means to advance cognitive load theory. educational psychologist, 38(1), 63–71. https://doi.org/10.1207/s15326985ep3801_8 palomäki, j., kivikangas, m., alafuzoff, a., hakala, t., & krause, c. m. (2012). brain oscillatory 4-35 hz eeg responses during an n-back task with complex visual stimuli. neuroscience letters, 516 (1), 141–145. https://doi.org/10.1016/j.neulet.2012.03.076 park, b., & brünken, r. (2015). the rhythm method: a new method for measuring cognitive load. an experimental dual-task study. applied cognitive psychology, 29(2), 232–243. https://doi.org/10.1002/acp.3100 park, s., & lim, j. (2007). promoting positive emotion in multimedia learning using visual illustrations sanghoon park and jung lim. journal of educational multimedia & hypermedia, 16 (2), 141–162. pesonen, m., hämäläinen, h., & krause, c. m. (2007). brain oscillatory 4-30 hz responses during a visual n-back memory task with varying memory load. brain research, 1138, 171–177. https://doi.org/10.1016/j.brainres.2006.12.076 pfurtscheller, g., & lopes da silva, f. h. (1999). event-related eeg/meg synchronization and desynchronization: basic principles. clinical neurophysiology, 110(11), 1842–1857. https://doi.org/10.1016/s1388-2457(99)00141-8 pfurtscheller, g., zalaudek, k., & neuper, c. (1998). event-related beta synchronization after wrist, finger and thumb movement. electroencephalography and clinical neurophysiology, 109 (2), 154–160. https://doi.org/10.1016/s0924-980x(97)00070-2 picton, t. w. (2000). guidelines for using human event-related potentials to study cognition: recording standards and publication criteria. psychophysiology, 37(2), 127–152. https://doi.org/10.1111/1469-8986.3720127 rayner, k. (1998). eye movements in reading and information processing: 20 years of research. psychological bulletin, 124(3), 372–422. https://doi.org/10.1037/0033-2909.124.3.372 rayner, k. (2009).eye movements and attention in reading, scene perception, and visual search. quarterly journal of experimental psychology (vol. 62). https://doi.org/10.1080/17470210902816461 reichle, e. d., & reingold, e. m. (2013). neurophysiological constraints on the eye-mind link. frontiers in human neuroscience, 7(july), 1–6. https://doi.org/10.3389/fnhum.2013.00361 reingold, e. m., & sheridan, h. (2011). eye movements and visual expertise in chess and medicine. in s. p. liversedge, i. d. gilchrist, & s. everling (eds.), the oxford handbook of eye movements(pp. 528–550). oxford university press. https://doi.org/10.1093/oxfordhb/9780199539789.013.0029 rey, g. d. (2012). a review of research and a meta-analysis of the seductive detail effect. educational research review, 7 (3), 216–237. https://doi.org/10.1016/j.edurev.2012.05.003 ries, a. j., touryan, j., ahrens, b., & connolly, p. (2016). the impact of task demands on fixation-related brain potentials during guided search. plos one, 11(6), e0157260. https://doi.org/10.1371/journal.pone.0157260 sauseng, p., griesmayr, b., freunberger, r., & klimesch, w. (2010). control mechanisms in working memory: a possible function of eeg theta oscillations. neuroscience and biobehavioral reviews, 34 (7), 1015–1022. https://doi.org/10.1016/j.neubiorev.2009.12.006 sauseng, p., klimesch, w., gruber, w., doppelmayr, m., stadler, w., & schabus, m. (2002). the interplay between theta and alpha oscillations in the human electroencephalogram reflects the transfer of information between memory systems. neuroscience letters, 324(2), 121–124. https://doi.org/10.1016/s0304-3940(02)00225-2 scharinger, c., kammerer, y., & gerjets, p. (2015). pupil dilation and eeg alpha frequency band power reveal load on executive functions for link-selection processes during text reading. plos one, 10(6), e0130608. https://doi.org/10.1371/journal.pone.0130608 scharinger, c., kammerer, y., & gerjets, p. (2016). fixation-related eeg frequency band power analysis: a promising neuro-cognitive methodology to evaluate the matching-quality of web search results? in c. stephanidis (ed.), hci international 2016 posters’ extended abstracts, part i (pp. 245–250). cham, switzerland: springer international publishing. https://doi.org/10.1007/978-3-319-40548-3_41 scharinger, c., soutschek, a., schubert, t., & gerjets, p. (2015). when flanker meets the n-back: what eeg and pupil dilation data reveal about the interplay between the two central-executive working memory functions inhibition and updating. psychophysiology, 52(10), 1293–1304. https://doi.org/10.1111/psyp.12500 scharinger, c., soutschek, a., schubert, t., & gerjets, p. (2017). comparison of the working memory load in n-back and working memory span tasks by means of eeg frequency band power and p300 amplitude. frontiers in human neuroscience, 11(6). https://doi.org/10.3389/fnhum.2017.00006 schmeck, a., opfermann, m., van gog, t., paas, f., & leutner, d. (2015). measuring cognitive load with subjective rating scales during problem solving: differences between immediate and delayed ratings. instructional science,43(1), 93–114. https://doi.org/10.1007/s11251-014-9328-3 schneider, s., dyrna, j., meier, l., beege, m., & rey, g. d. (2017). how affective charge and text–picture connectedness moderate the impact of decorative pictures on multimedia learning. journal of educational psychology. https://doi.org/10.1037/edu0000209 schneider, s., nebel, s., & rey, g. d. (2016). decorative pictures and emotional design in multimedia learning. learning and instruction, 44(march), 65–73. https://doi.org/10.1016/j.learninstruc.2016.03.002 schnotz, w., fries, s., & horz, h. (2009). motivational aspects of cognitive load theory. in m. wosnitza, s. a. karabenick, a. efklides, & p. nenniger (eds.), contemporary motivation research. from global to local perspectives (pp. 69–96). göttingen, germany: hogrefe & huber. schnotz, w., & kürschner, c. (2007). a reconsideration of cognitive load theory. educational psychology review, 19(4), 469–508. https://doi.org/10.1007/s10648-007-9053-4 schöne, b., schomberg, j., gruber, t., & quirin, m. (2016). event-related frontal alpha asymmetries: electrophysiological correlates of approach motivation. experimental brain research, 234(2), 559–567. https://doi.org/10.1007/s00221-015-4483-6 schüler, a. (2017). investigating gaze behavior during processing of inconsistent text-picture information: evidence for text-picture integration. learning and instruction, 49, 218–231. https://doi.org/10.1016/j.learninstruc.2017.03.001 simola, j., le fevre, k., torniainen, j., & baccino, t. (2015). affective processing in natural scene viewing: valence and arousal interactions in eye-fixation-related potentials. neuroimage, 106, 21–33. https://doi.org/10.1016/j.neuroimage.2014.11.030 simola, j., torniainen, j., moisala, m., kivikangas, m., & krause, c. m. (2013). eye movement related brain responses to emotional scenes during free viewing. frontiers in systems neuroscience, 7(41). https://doi.org/10.3389/fnsys.2013.00041 smith, e. e., reznik, s. j., stewart, j. l., & allen, j. j. b. (2017). assessing and conceptualizing frontal eeg asymmetry: an updated primer on recording, processing, analyzing, and interpreting frontal alpha asymmetry. international journal of psychophysiology, 111, 98–114. https://doi.org/10.1016/j.ijpsycho.2016.11.005 strukelj, a., & niehorster, d. c. (2018). one page of text: eye movements during regular and thorough reading, skimming, and spell checking. journal of eye movement research, 11(1), 1–22. https://doi.org/10.16910/jemr.11.1.1 sweller, j., van merrienboer, j. j. g., & paas, f. g. w. c. (1998). cognitive architecture and instructional design. educational psychology review, 10(3), 251–296. https://doi.org/10.1023/a:1022193728205 triesch, j., ballard, d. h., hayhoe, m. m., & sullivan, b. t. (2003). what you see is what you need. journal of vision, 3(1), 86–94. https://doi.org/10.1167/3.1.9 van gog, t., & scheiter, k. (2010). eye tracking as a tool to study and enhance multimedia learning. learning and instruction,20 (2), 95–99. https://doi.org/10.1016/j.learninstruc.2009.02.009 vignali, l., himmelstoss, n. a., hawelka, s., richlan, f., & hutzler, f. (2016). oscillatory brain dynamics during sentence reading: a fixation-related spectral perturbation analysis. frontiers in human neuroscience, 10, 1–13. https://doi.org/10.3389/fnhum.2016.00191 weinreich, a., stephani, t., & schubert, t. (2016). emotion effects within frontal alpha oscillation in a picture oddball paradigm. international journal of psychophysiology, 110, 200–206. https://doi.org/10.1016/j.ijpsycho.2016.07.517 weiss, b., knakker, b., & vidnyánszky, z. (2016). visual processing during natural reading. scientific reports, 6, 1–16. https://doi.org/10.1038/srep26902 winslow, b., carpenter, a., flint, j., wang, x., tomasetti, d., johnston, m., & hale, k. (2010). combining eeg and eye tracking: using fixation-locked potentials in visual search. journal of eye movement research, 6(4), 1–11. https://doi.org/10.16910/jemr.6.4.5 yarbus, a. l. (1967). eye movements and vision. new york, ny, usa: plenum press. https://doi.org/10.1016/0028-3932(68)90012-2 zhou, w., & gotman, j. (2009). automatic removal of eye movement artifacts from the eeg using ica and the dipole model. progress in natural science, 19(9), 1165–1170. https://doi.org/10.1016/j.pnsc.2008.11.013 microsoft word hilppö et al_publication.docx             frontline  learning  research  vol.4  no.  4  special  issue  (2016)  20  -­‐  29   issn  2295-­‐3159       corresponding author: jaakko hilppö, northwestern university, school of education and social policy, walter annenberg hall, 2120 campus drive, evanston, illinois 60208, united states. email: jaakko.hilppo@northwestern.edu doi: http://dx.doi.org/10.14786/flr.v4i4.213 interactive dynamics of imagination in a science classroom jaakko hilppöa, antti rajalab, tania zittounc, kristiina kumpulainenb, lasse lipponenb anorthwestern university buniversity of helsinki buniversity of neuchâtel article received 22 september / revised 20 september / accepted 21 september / available online 19 december   abstract in this paper, we introduce a conceptual framework for researching the dynamics of imagination in science classroom interactions. while educational interest in imagination has recently increased, prior research has not adequately accounted for how imagination is realized in and through classroom interactions, nor has it created a framework for its empirical investigation. drawing on a theory of imagination situated in cultural psychology (zittoun et al., 2013; zittoun & gillespie, 2016), we propose such a framework. we illustrate our framework with a telling case (mitchell, 1984) of imagination from a finnish primary science classroom community. our illustration focuses on the dynamics of imagination as it unfolds in classroom interactions and how qualitatively distinct loops of imagination are formed. in specific, we show how the students’ meaning making expands in time and space and can become more refined and differentiated through loops of imagination and their dynamics. in all, our paper argues that imagination is a constitutive element of science learning. our proposed conceptual framework provides potential avenues for further empirical research on the dynamics of imagination in science learning and teaching. keywords: imagination; science learning; classroom interaction hilppö  et  al       | f l r     21   1. introduction imagination is increasingly identified as an important aspect of learning (nemirovsky, rasmussen, sweeney & wawro, 2012; pelaprat & cole, 2011; zittoun & gillespie, 2016; fleer, 2015). for example, interest in imagination is evident in recent research on learning as knowledge creation (damsa & jornet, this issue), and on playful and creative learning (connery, john-steiner, & marjanovich-shane, 2010; wegerif, 2007). while imagination is closely related to creativity, we consider the outcomes of imagination as creative only when they are socially acknowledged as such (glăveanu, gillespie, & valsiner 2015). the power of imagination lies in its potential to enrich the way people experience and interact with their worlds. in imagination, actions can be disconnected from their usual consequences and phenomena not available in the immediate proximal experience can be conjured. imagination can thus involve experimenting with new ideas without real-life repercussions or constraints, and as such imagination is a fundamental aspect of learning. furthermore, it is a crucial part of science-in-the-making, that is, the actual work of scientists (latour & woolgar, 1986). despite this growing interest, imagination in science learning, has only recently become the focus of conceptual and empirical research. for example, fleer (2015) showed how imagination, emotions and concept formation are united in children’s interaction with their lifeworlds. van eijck & roth (2013), in turn, showed that in science classrooms and textbooks, scientific endeavors are predominantly imagined as epic stories depicting scientists as heroes. while these studies have started to unpack how imagination interacts with other aspects of learning and instruction, the dynamic characteristics of imagination as part of science education are not yet conceptualized. such a lack, in our opinion, undermines attempts to further understand the importance of imagination in learning. in this paper, we build on and extend a theory of imagination situated in cultural psychology (zittoun & gillespie, 2016) to assemble a conceptual framework for researching how students make meaning of science through imagination. we also illustrate and enrich this framework by analyzing a telling case (mitchell, 1984) of the dynamics of imagination in a finnish primary science classroom community. lastly, we shortly discuss the contribution of our framework for empirical research on imagination in learning. we also discuss the strengths and limitations of the framework as well as propose avenues for future research on the topic. 2. assembling a conceptual framework: on loops of imagination... drawing upon a cultural psychological theory of imagination proposed by one of us (tania zittoun) and colleagues (zittoun & gillespie, 2016; see also, vygotsky, 1931), we define imagination in this study as "the process of creating experiences that escape the immediate setting, which allow exploring the past or future, present possibilities or even impossibilities. imagination feeds on a wide range of experiences people have of, or through the cultural world, through diverse senses, now combined, organized and integrated in new forms" (zittoun & gillespie, 2016, p. 2). zittoun & gillespie (2016) continue by theorizing that imagination involves a partial and temporary decoupling from the here-and-now of a particular situation. this partial decoupling is triggered by an event, like being bored in class or struck by an inspirational piece of art, that disrupts one’s ongoing engagement. during this decoupling one’s experience is expanded into semiotically mediated and more distal experiences. a work of art, like a book, might evoke imagining other times and worlds and thus lets the reader enrich her experience by imagining elements not present otherwise. the partial separation of experience ends in the eventual reconnection back to the here-and-now. this process of back-and-forth movement between proximal and distal experiences negotiated in interaction constitutes what zittoun and gillespie (2016) call loops of imagination. hilppö  et  al       | f l r     22   loops of imagination take different context-specific forms. they can be characterized along the dimensions of temporality, plausibility and generalization (zittoun & cerchia, 2013). for example, discussing futuristic societal utopias during civic lessons, reading about ancient egypt or daydreaming about the upcoming summer holiday represent different temporal characteristics, plausibilities, and generalizations as loops of imagination. the temporal dimension designates the temporal “aboutness” of imagining; discussing ancient egypt involves imagining the past and anticipation involves imagining the future. these loops can pertain to issues that are either plausible or improbable in a given social setting and that can vary from concrete personal experiences to generalized abstractions. in this paper, we suggest that the loops of imagination also include a spatial dimension that has not yet been conceptualized in connection to imagination loops. this spatial dimension is evident in the above examples; imagining about future societies, the ancient egypt or the summer holiday all entail a spatial displacement from the here-and-now to an imagined distal space. what zittoun’s and gillespie’s work highlights is the ubiquity of imagination in human activities. rather than being something that people do only at certain times, imagination is a constitutive element of even the most mundane situations (see also pelaprat & cole, 2011). more fundamentally, from a cultural psychological perspective the emergence and use of cultural tools in human activities – namely cultural mediation – has relied on and spurred our capacity to imagine. moreover, although single persons can imagine, their imagination is nonetheless constituted by social and cultural means and their sometimes conflicting perspectives (wertsch, 1991; bakhtin, 1981). 3. ….and their dynamics in science education in the context of formal science education students are often required to imagine phenomena not available to them in the immediate here-and-now. the students are also asked to explore different explanations to phenomena in the natural world and judge their plausibility (varelas & pappas, 2006; kumpulainen et al., 2003), all of which require imagination. in other words, when learning about particular phenomena, students (like scientist) revisit and revise how they imagine these phenomena. while zittoun’s and gillespie’s theory of imagination is sensitive to the process nature and context specificity of imagination, it does not account for this type of dynamics of imagination. they do, however, provides us with initial direction. building on vygotsky (1931) and werner & kaplan (1963), zittoun and gillespie point out that imagination can expand and become more refined over time. for example, after first imagining going to the beach, one can start populating this image with different beach activities or imagine a different destination. on a longer time scale, with the progressive mastery of specific culturally available means, a child’s version of his dream home can become more differentiated when she re-designs it as a professional architect and also expand when she applies her design ideas into a new design domain, like cars. these dynamics are relevant aspects of imagination and thus a conceptual framework of imagination should address them. in this paper, we propose that imagination in science education holds two distinct but intertwined dynamics, expansive and refining. in our proposed conceptual framework the expansive dynamics of imagination accounts for movement in meaning-making between proximal and more distal times and spaces. the expansive dynamics of imagination permits a classroom community to explore topics that extend in space and time beyond their immediately experienced worlds (e.g., addressing phenomena ranging from microscopic to planetary spaces and timescales). this dynamic is thus primarily related to the temporal and spatial dimensions of the loops of imagination. it also relates to the dimension of generality because it involves moving between particular lived experiences and more abstract and general descriptions of objects and events. in turn, the refining dynamics of imagination accounts for how a classroom community develops a progressively more refined understanding of science topics under discussion. this dynamic makes possible hilppö  et  al       | f l r     23   for the classroom community to settle on what they consider as an acceptable explanation. in this regard, the refining dynamics is primarily related to the plausibility dimension of the loops of imagination. thus, this dynamic connects to an important goal of formal science learning: the development of plausible interpretations and explanations that hold across contexts. table 1 summarizes the proposed conceptual framework. table 1 dynamics of imagination in science learning dynamics of imagination definition associated dimensions of imagination loops expansive dynamics • accounts for shifts in meaning-making between proximal and more distal times and spaces. • permits a classroom community to explore topics that extend in space and time beyond their immediately experienced worlds. temporality, spatiality, generality refining dynamics • accounts for how a classroom community develops a progressively refining understanding of science topics under discussion • permits a classroom community to distinguish what is accepted as plausible explanation plausibility we will next illustrate and enrich the proposed conceptual framework by analyzing empirical data consisting of video-recorded social interactions in a science classroom. like others, we consider imagination as a socially mediated process (zittoun & gillespie, 2016; fleer, 2015; pelaprat & cole, 2011). this requires that our analytical approach must be sensitive to how imagination is realized in and through social interaction between participants. the microgenetic analysis we employ to our telling case abides to a dialogical methodology (linell, 1998; jordan & henderson, 1995) and allows us to study imagination as a process and outcome of joint action. our approach resembles those used in research of creative classroom interactions, in which creativity is conceptualized as a social and distributed process (e.g., wegerif, 2007; sawyer, 2004). 4. enriching the conceptual model by analyzing illustrative examples in this section, we analyze a student-led classroom discussion that occurred during a science project. in the discussion the students explore topics which, due to their unfamiliarity, extend beyond their immediately experienced worlds and thus mobilizes the student’s imaginations. the class in question is a culturally and socio-economically heterogeneous third-grade classroom community of eighteen students (910 years old) and their teacher (one of us, antti rajala) from the metropolitan area of helsinki, finland. the teacher’s pedagogical thinking was influenced by sociocultural and dialogic approaches, mainly the hilppö  et  al       | f l r     24   thinking together project (dawes, mercer, & wegerif, 2000), which he applied in his classroom to promote constructive ways of using language as a social mode of thinking (see e.g., rajala, 2016). the discussion analyzed in this paper began when one of the students, maija, asked a question concerning the origin of stones (february 8, 2008). the class was discussing a pile of stones that the students had encountered in the forest. maija requested for the floor and raised a new question (line 134 maija: but tell me that where have the stones come from?). the teacher designated maija as the chair for what became a lively and extended discussion on this topic. maija’s wondering acts as a trigger for a loop of imagination. through this loop, the classroom community decouple from their immediately experienced world to explore distal times and spaces to explain the origin of stones. as a result, a rich array of new meanings are developed, deployed and refined in the ensuing discussion. yet, no conclusion is reached for maija’s question and the topic is postponed until a later, unspecified occasion. the topic resurfaces and a new imagination loop regarding the origin of stones is triggered after more than a month later when maija reiterates her question. this happens while the class was discussing a film on moon formation that they had seen in a science center (march 19, 2008). next, we will show the formation of these imagination loops and illuminate expansive and refining dynamics that loops of imagination can involve. through these two dynamics, the students’ meaning making expands in time and space and becomes more refined and differentiated. 4.1. expansive dynamics of imagination in the expansive dynamics of imagination the students’ meaning making shifts from proximal to increasingly distal times and spaces when the classroom community explores different explanations for the origin of stones in their discussion. when discussion began the students’ explanations were embedded in a spatio-temporal frame that concerned objects and events within the confines of the earth as an existing planet. within this frame the explanations did not extend to the issue of earth’s formation or events or objects beyond its boundaries. the following example illustrates this time-space frame that we named as a geological time-space frame. 4.1.1. example 1. geological time-space frame 149 maija: can i say, if they have come from sand so, like, do stones, like, grow? 150 kimmo: no 151 (unidentified): no 152: teacher: give out turns at talk 153 kimmo: i think they have come, like, there has been some rocks and then the rocks have at some point collapsed. they might have come from there. 154 maija: where have ro 155 kimmo: well there was that rock over there and it might have collapsed 156 maija: but where have the rocks then come from? 157 teacher: maija is now the chair. maija can talk when she wants to and designate speaking turns. 158 kimmo: well, it has come from earth. the example begins when maija reiterates a previous answer to her question on the origin of stones (stones can come from sand) and asks for confirmation. kimmo and others refute this explanation, and kimmo proposes that stones originate from collapsing rocks. kimmo then concretizes this image in terms of hilppö  et  al       | f l r     25   personally experienced time and space by referring to the specific rocks that the class saw during their field trip (line 155). maija provokes kimmo to explain how the rocks have formed in the first place. through kimmo’s response the student’s meaning making begins to expand in time and space. maija’s question (line 156) represents a recurrent social dynamic in the discussion that serves as means for expanding the scope of meaning making. in the next example, we will see how the students’ meaning making further expands to address more distal times and spaces within what we named as a planetary time-space frame. in this frame earth is positioned in the wider solar system, and issues of its formation and interactions with extra-terrestrial objects, such as meteors, are addressed. we enter the interaction when saara asks a question that pushes the class to think beyond their intermediate conclusion that sand is stone. 4.1.2. example 2. from geological to planetary time-space frame 172 saara: so, like, sand is stone. so, like, where has sand come from? -- 180 kimmo: i think it has come from (unclear) earth. like, there has been in the beginning a fiery, a sun or some sort of a fireball and then it has hardened and then 181 (kimmos ideas are laughed at) 182 timo: a fireball (laughing) 183 kimmo: i mean a big one 184 (timo mimics a fireball and ridicules kimmo’s idea) -- 194 oliver: so i’m not like believing that there suddenly comes this fireball, that there comes this fireball or like or like (laughter) then all of a sudden the sand appears. triggered by saara’s question, kimmo first explains that sand has come from earth (line 180). next, he moves to talk about the formation of the planet earth by evoking an image of a blazing celestial body which then solidifies. kimmo’s new explanation expands the class’s meaning making in time and space for the first time to a planetary scale. here, the laughter, timo’s sneering, and most evidently, oliver’s response (line 194) all acknowledge kimmo’s effort in creating the new distal time-space frame. yet, at the same time timo, oliver and the others are resisting the expansion and this tension, stemming from the demand for plausibility of the explanations, becomes a source for further dynamics in the discussion. in all, in both lessons through the loops of imagination the students’ meaning making expands in time and space. in the first lesson, the imagination primarily moves within a geological time-space frame but is expanded to include a planetary time-space frame. imagination also occasionally recouples with personally experienced times and spaces when the students make references to events and objects within the scope of their lifeworlds, such as kimmo’s reference to the nearby rocks (example 1). in contrast, planetary timespace frame is the dominant frame of the second lesson. this shift in the dominant time-space frame is primarily due to the symbolic mediation by the film about moon formation that the class is discussing. during the second lesson the meaning making further expands beyond the planetary time-space frame. this happens when the teacher explains the big bang theory and how it accounts for the emergence of matter in the universe. in response, the class also discusses alternative religious explanations of earth’s creation. maija’s question about the origin of god (line 145 maija: who made god?) further illustrates the role of continuing questions in the expansive dynamics of imagination and how different symbolic materials feed that process. hilppö  et  al       | f l r     26   4.2. refining dynamics of imagination through the loops of imagination meaning making in classroom interaction can also become more refined, as illustrated in the following example. the example begins when oliver continues to ridicule kimmo’s incipient explanation that sand emerged from a ball of fire when earth was formed. 4.2.1. example 3. refining the meaning making 240 oliver: so, like, how come, like, if a piece gets detached from earth, then how come it can, like, suddenly start to come back as a fireball (hussein and timo laugh) so‚ like, the earth circulates the pieces. 241 kimmo: nooo (laughing) 242 timo: are there holes in the earth 243 kimmo: no, if you’d dig long enough into the earth 244 esa: there would be lava (in a silly voice) 245 kimmo: yeah, i meant just that, in the beginning there was a (hussein: it would burn) lava or a ball 246 timo: (overlaps with kimmo) no first, no. there’s the mantle of earth oliver first expresses his doubts about kimmo’s explanation. hussein and timo support oliver’s ridiculing by laughing. kimmo also laughs and refutes oliver’s interpretation of his explanation. timo then asks whether there are holes in earth. kimmo defends his explanation and proposes a thought experiment of digging a deep hole through earth. while the students argue and counter-argue with each other and evaluate various lines of reasoning that are put forth, their meaning making around the formation of earth becomes progressively more refined. the students use their background knowledge to build a more differentiated image of the structure of earth, and kimmo uses this image in an attempt to convince the other students. in particular, he elaborates his explanation that the fireball that is indicated in his explanation is of similar to hot lava deep under the earth’s surface. the students’ deployment of technical language, such as ‘lava’ and ‘mantle of earth’, constitutes effective means whereby their imagination is refined. figure 1 shows a rough overview of both the expansive and refining dynamics of imagination during the two lessons. the expansive dynamics can be seen, for example, in turns 170-194 when the discussion shifts into a new time-space frame. in turn, the refining mode can be seen in the continuous block of speaking turns in the planetary time-space frame between turns 240-246. figure 1. the expansion of meaning making in time and space during lessons february 8 and march 19 and also the location of examples 1-3 in the interaction. hilppö  et  al       | f l r     27   5. discussion although imagination in science learning has been previously conceptualized as a situated and shared process (e.g., fleer, 2015), prior research has not provided for a framework for studying the dynamics of imagination in classroom interactions. in this paper we have presented such a framework by drawing on and extending a theory of imagination situated in cultural psychology (zittoun & gillespie, 2016). we have also illustrated the proposed our framework by analyzing a telling case of imagination in science education. as demonstrated by our telling case, the expansive and refining dynamics foregrounded imagination in the joint meaning making of the classroom. the expansive dynamics of imagination highlighted the dynamics of the loops in relation to their spatial and temporal dimensions. here our new suggestion, the spatial dimension, appeared to sit well with zittoun’s and gillespie’s theory as it provided a way to address the movement of imagination in more detail (cf. zittoun & cerchia, 2013). our framework also highlighted loop dynamics in relation to the dimension of generality when the students repeatedly moved between general descriptions of more distal objects and events and their particular experiences in their joint meaning making. the refining dynamics illuminated the plausibility dimension of imagination. that is, our examples made evident that the students themselves demanded plausible explanations and explicit reasoning from each other based on what they know about the topic or what is common sense to them. in particular, our conceptual framework highlights the formation of a dialogical space (wegerif, 2007) that revolves around the tension between plausibility (what is) and playful exploration of ideas (what could be). this dialogical space created a generative dynamic that pushed joint meaning making to expand and become more refined and differentiated. the ensuing discussion presented the students a fruitful context for defining and learning to use field-specific, differentiated vocabulary. this observation also raises more fundamental questions about what knowing, or learning to know something means in science classroom practices. specifically, learning to know about the origin of stones, or any other phenomena, seems to emerge out of a tensional process where an array of different imagined explanations is ordered in relation to their plausibility (law, 1998). in other words, the importance of imagination for science education is not just in how it makes possible for the students to think about the studied phenomena as they are, but also as they are not, and making the difference between these two. our conceptual framework highlights this process. in all, our conceptual framework illuminates the dynamical characteristics of imagination in science classroom interactions. it also contributes to research on science learning more generally (e.g., varelas & pappas, 2006; kumpulainen et al., 2003), by showing how students – through imagination – can create and evaluate contrasting explanations, even seemingly implausible ones. more specifically, our conceptual framework illuminated what triggered the loops of imagination and the different forms these loops took furthermore, our conceptual framework helped to unpack the social work needed for imagination to take place in classroom interactions. while a single case analysis does not allow us to evaluate how our conceptual framework applies to more varied instances of imagination, we feel that the work presented here offers grounds for further research on the topic. in particular, it generates new relevant research foci and empirical research questions for science education and learning. future research could, for example, investigate the role of imagination in students’ construction of explanations and how teachers and the educational context can support this process. knowing what triggers, sustains and/or hampers as well as personal imagination in educational settings is also of great importance. it would also be interesting to investigate how imagination in science learning is mediated by material tools and artifacts or how, if at all, the dynamics of imagination differ in the different phases of the learning process. future research may also want to address the ‘dark side’ of imagination, that is, how imagination can limit the ways in which people interact with their world(s). hilppö  et  al       | f l r     28   keypoints imagination is a central aspect of science learning and science education through loops of imagination students’ meaning making expands in time and space and becomes progressively more refined and differentiated the introduced conceptual framework for the study imagination opens up new research foci for science education and learning acknowledgments we wish to thank the students for sharing their classroom with us. we would also like to thank the ella and georg ehrnrooth foundation and the jenny and antti wihuri foundation for financially supporting rajala’s contribution to this paper as well as the seduce doctoral school (university of helsinki) for supporting hilppö’s contribution to this paper. lastly, rajala and hilppö would like to thank the lobby sofa of the days hotel coventry for its material and emotional support. references bakhtin, m. m. (1981). the dialogic imagination. four essays by m. m. bakhtin. holquist, m. (ed.) austin, tx: university of texas press. connery, m. c., john-steiner, v., & marjanovic-shane, a. (2010). vygotsky and creativity: a culturalhistorical approach to play, meaning making, and the arts. new york: peter lang. damsa, c., & jornet, a. (this issue). revisiting learning in higher education—framing notions redefined through an ecological perspective. frontline learning research. dawes, l., mercer, n., & wegerif, r. (2000). thinking together: a programme of activities for developing thinking skills at ks2. uk: questions. fleer, m. (2015). imagination and its contributions to learning in science. in m. fleer & n. pramling (eds.) a cultural-historical study of children learning science (pp. 39-57). netherlands: springer. glăveanu, v., gillespie, a., & valsiner, j. (2015). (eds.) rethinking creativity: contributions from social and cultural psychology. london: routledge. kamberelis, g., & wehunt, m. d. (2012). hybrid discourse practice and science learning. cultural studies of science education, 7(3), 505-534. doi:10.1007/s11422-012-9395-1 kumpulainen, k., vasama, s., & kangassalo, m. (2003). the intertextuality of children's explanations in a technology-enriched early years science classroom. international journal of educational research, 39(8), 793-805.doi:10.1016/j.ijer.2004.11.002 jordan, b., & henderson, a. (1995). interaction analysis: foundations and practice. the journal of the learning sciences, 4(1), 39–103.doi: 10.1207/s15327809jls0401_2 latour, b., & woolgar, s. (1986). laboratory life: the social construction of scientific facts. princeton: princeton university press. law, j. (1998). after meta-narrative: on knowing in tension. in r. chia, in the realm of organisation: essays for robert cooper (pp. 90–111). london: routledge. linell, p. (1998). approaching dialogue: talk, interaction and contexts in dialogical perspectives. amsterdam: john benjamins. mcmillan, m. (1904). education through the imagination. london, s. sonnenschein & co., lim. mitchell, c. j. (1984). typicality and the case study. in r.f. ellens (ed.), ethnographic research: a guide to general conduct (pp. 238–241). new york, ny: academic press. hilppö  et  al       | f l r     29   nemirovsky, r., rasmussen, c., sweeney, g., & wawro, m. (2012). when the classroom floor becomes the complex plane: addition and multiplication as ways of bodily navigation. journal of the learning sciences, 21(2), 287-323.doi: 0.1080/10508406.2011.611445 pelaprat, e., & cole, m. (2011). “minding the gap”: imagination, creativity and human cognition. integrative psychological and behavioral science, 45(4), 397-418.doi: doi:10.1007/s12124-011-9176-5 rajala, a. (2016). toward an agency-centered pedagogy: a teacher's journey of expanding the context of school learning. (doctoral dissertation) university of helsinki. van eijck, m., & roth, w. m. (2013). imagination of science in education: from epics to novelization (vol. 7). new york: springer. varelas, m., & pappas, c. c. (2006). intertextuality in read-alouds of integrated science-literacy units in urban primary classrooms: opportunities for the development of thought and language. cognition and instruction, 24(2), 211-259.doi: 10.1207/s1532690xci2402_2 vygotsky, l. s. (1931). pedologija podrostka. moscow-leningrad: uchebno-pedagogicheskoe izdatel’stvo. (title in english: paedology of the adolescent). vygotsky, l. s. (2004). imagination and creativity in childhood. journal of russian & east european psychology, 42(1), 7-97. warren, b., ballenger, c., ogonowski, m., rosebery, a. & hudicourt-barnes, j. (2001). rethinking diversity in learning science: the logic of everyday sensemaking. journal of research in science teaching, 38, 529-552.doi: 10.1002/tea.1017 wegerif, r. (2007). dialogic education and technology: expanding the space of learning (vol. 7). new york: springer. werner, h., & kaplan, b. (1963). symbol formation: an organismic developmental approach to language and the expression of thought. ny: john wiley. wertsch, j. (1991). voices of the mind. a sociocultural approach to mediated mind. cambridge, ma. harvard university press. zittoun, t., & cerchia, f. (2013). imagination as expansion of experience. integrative psychological and behavioral science, 47(3), 305-324. doi:10.1007/s12124-013-9234-2 zittoun, t., & gillespie, a. (2016). imagination in human and cultural development. london: routledge. frontline learning research vol. 5 no. 3 special issue (2017) 81 93 issn 2295-3159 *corresponding author: laura helle, department of teacher education, university of turku, assistentinkatu 5, 20014 turun yliopisto, finland. lhelle@utu.fi doi: http://dx.doi.org/10.14786/flr.v5i3.254 prospects and pitfalls in combining eye-tracking data and verbal reports laura helle university of turku, finland article received 30 april / revised 30 january / accepted 23 march / available online 14 july abstract it is intuitively appealing to try to combine eye-tracking data and verbal reports when investigating medical image interpretation. however, before collecting such data, important decisions must be made, including exactly when and how to collect the verbal reports. the purpose of this methodological article is to reflect on the pros and cons of different solutions and to offer some guidelines to investigators. we start by exploring the ontology of vision and speech production and the epistemology of eye movements to grasp what fixations and verbal reports actually reflect. we are also interested in the major constraints of the two systems. second, we elaborate on two dominant investigational approaches to verbal accounts: concurrent think-aloud and chi’s explanations. later, we move on to other approaches. third, we present and critically evaluate studies from the literature on medical image interpretation, specifically ones that have sought to contrast or integrate eye-movement data and verbal reports. fourth, we conclude with some practical guidelines and suggestions for further research. keywords: eye tracking; gaze tracking; verbal reports; think-aloud; medical images; clinical reasoning mailto:lhelle@utu.fi http://dx.doi.org/10.14786/flr.v5i3.254 helle | f l r 82 1. introduction the study of medical expertise in visual domains, such as radiology and dermatology, is firmly rooted in two distinct investigational approaches, both of which serve certain purposes: (a) the study of visual search or perception using eye-tracking methods (e.g., berbaum et al., 1998; kundel, nodine, & carmody, 1978; krupinski et al., 2006; rubin et al., 2014) and (b) the study of clinical reasoning, usually employing verbal reports (azevedo, faremo, & lajoie, 2007; lesgold, feltovich, glaser, & wang, 1981; morita et al., 2008; van der gijp et al., 2015). before the advent of commercial eye trackers, verbal reports were basically the only way to gain insight into diagnostic reasoning. even today, some type of verbal report is needed because one cannot deduce from, for example, dwell times, whether a viewer actually “sees” a lesion (berbaum, franken, dorfman, caldwell, & krupinski, 2000). a crucial part of the perceptual process is assigning meaning to what one sees (nodine & kundel, 1987). as for the value of eye tracking, krupinski (2006) argued that eye tracking may be useful for developing individual eye movement profiles and for understanding the difference in performance between novices and experts. in addition, they can be useful for developing new visual search strategies. although studies following both lines of investigation have shown important insights, one can question whether either of the approaches alone is sufficient enough to answer important research questions. it is hard to see how medical image perception investigators are meeting the expectations of modeling, for instance, search strategies by relying on eye movement metrics alone. it also hard to see how process models can be justified based on only one source of data. as for the protocol analysts, it is odd that, for example, van der gijp et al. (2014) conceptualized the interpretation of radiological images as a process of perception, analysis, and synthesis but methodologically relied on concurrent think-aloud techniques without the use of eye tracking. in fact, it has been argued in the context of occupational psychology that complex cognitive work tasks should be studied by integrating various sources of information, including eye movement data, when appropriate, with verbal reports (patrick & james, 2004; gegenfurtner et al., 2017). patrick and james (2004) stressed that process tracing involves four stages, and important decisions have to be made in each stage. the stages are the following: (1) collection of data, (2) transcription, integration, and segmentation of the data into a time-lined account, (3) coding, and (4) further analysis of the data from stage 3 and representation of the data. in the data collection stage, one of the most critical decisions involves the timing of data collection, because verbal accounts can be collected concurrently with task performance or retrospectively. as for the transcription phase of verbal reports, the authors present the integrated actions of a person in a single table. alternatively, one could think of either a data matrix containing a time-lined account of actions that is obtained through eye-tracking software or a set of time-stamped, transcribed videos. stage 3 involves coding of the transcribed data either based on theoretical categories or done in a bottom-up fashion. the authors stressed that when categories are derived from a bottom-up approach, independent raters should refine categories iteratively, with some form of reliability being reported. in stage 4, the analyst filters or expands the data using the newly acquired codes from stage 3, and subjects the data to further analysis, whereby certain aspects of cognition are made more salient. the authors stress two points: (a) a minimum level of further analysis is whether a worker’s response or solution is correct; (b) there is a need to capture and represent at a global level a person’s reasoning during a scenario in relation to changes in the task and work situation. the purpose of this methodological article is to reflect on the pros and cons of different solutions and to offer some guidelines for investigators. it is stressed that this endeavor stretches the frontiers of the field because gegenfurtner, siewiorek, lehtinen, and säljö (2013) reported in their systematic review that combining eye tracking and verbal reports remains unexplored. also, this article is not a literature review. to review articles other than the one by gegenfurtner et al. (2017), see al-moteri, symmons, plummer, & cooper (2017); blondon, wipfli, and lovis (2015); and van der gijp et al. (2016). we start by exploring the ontology of vision and speech production and the epistemology of eye movements to grasp the major constraints involved in visual processing and speech production. second, we elaborate on two dominant helle | f l r 83 investigational approaches to verbal accounts and introduce alternative approaches. third, we present and evaluate studies from the literature on medical image interpretation that have sought to contrast or integrate eye movement data and verbal reports. fourth, we conclude with practical solutions and some suggestions for further research. 2. nature of the visual system: what do fixations actually reflect? the visual system is the part of the nervous system that allows organisms to see. it interprets information from the environment to build a representation of the surrounding world. the visual system has the complex task of reconstructing a three-dimensional world from a two-dimensional retinal representation of that world. the performance of the visual system in a constantly-changing visual environment is remarkable. the price, however, is that approximately one-third of the cortex is needed to process visual information (vanni, 2004). then, how does the visual system operate? information from the eyes flows into the brain through the optic nerve. information from the right visual field travels to the left optic tract. information from the left visual field travels to the right optic tract. each optic tract terminates in the lgn in the thalamus. the region that receives information directly from the lgn is called the v1. the macakee ape has over 30 cortical regions, and it is estimated that humans have approximately as many (vanni, 2004). these areas are connected to each other by an intricate wiring containing both feedforward and feedback connections (vanni, 2004). according to the ventral-dorsal model introduced by goodale and milner (1992), information flows in two directions from the primary visual cortex: (a) to the posterior parietal cortex through the dorsal stream and (b) to the inferotemporal cortex through the ventral stream. the dorsal pathway has been characterized as the action stream, a pathway concerned with converting visual inputs into motor outputs, whereas the ventral pathway provides a visual perception of objects and events in the world. goodale (1998) stressed, however, that even a simple action such as picking up a cup a coffee requires activity in both pathways. there are several bottlenecks in the visual system that stem from constraints in anatomy, attention, and working memory. first, high visual acuity is limited to the fovea, a spot on the retina. the fovea is employed for accurate vision in the direction where it is pointed. visual acuity decreases dramatically in the parafoveal area and periphery. in eye-movement research, it is possible to capture the target of foveal inspection through fixations. second, object recognition is limited by capacity and often attention-demanding because one cannot recognize multiple objects with more than one feature simultaneously (such as a letter t containing green and purple). object recognition requires more than 100 ms per item, which refers to processing time instead of presentation time (wolfe, võ, evans, & greene, 2011). third, there is a limit to focusing and shifting one’s attention: people tend to move their eyes between two to four times per second when reading and conducting most visual search tasks (salthouse & ellis, 1980). the gaze, however, can be trained to make the best out of the few fixations: the novice’s gaze is often drawn by salient, bottom-up features, whereas experts more often focus on top-down, task-relevant features, as evidenced by bertram, helle, kaakinen, and svedström (2013). (see also wolfe, evans, drew, aizenman, & josephs, 2015). fourth, although information flows into the system incessantly, working-memory capacity is limited to approximately four “chunks,” or combinations of items, at a time (cowan, 2010). in addition, information in one’s working memory is lost quickly: according to ericsson (2006), for tasks with response latencies of 5– 10 seconds, people are able to recall their sequences of thoughts quite accurately. for the main part, the human brain processes low-level information patterns in the environment automatically (vanni & heikkinen, 2015). studies adhering to a flash-view paradigm (i.e., presenting images to participants for 20–250 ms) have shown that people can partially infer a scene without even fixating on the scene (e.g., kirschner & thorpe, 2006). only a small part of the information reaching the cortex is helle | f l r 84 processed further, with storage capacity representing yet another filter. thus, visual information processing reaching awareness is only the tip of the iceberg. people use fixations to purposively sample information from their surroundings to reconstruct a representation of the surrounding world. however, studies adhering to a flash-view paradigm have shown that people can partially infer a scene without even fixating on the scene. thus, there appears to be two visual pathways, coined a selective pathway involving purposive sampling and a nonselective pathway by wolfe et al. (2011). to answer the question in the title, the sequence of fixations can be seen as reflecting a visual search through the selective pathway (i.e., attentional guidance). 3. speech production and verbal reports: what do verbal reports actually reflect? 3.1 speech production how people produce and why they produce speech is usually taken for granted. speech has many social and cultural functions, such as signifying group identity, social grooming, settling disputes, teaching, and entertainment. naturally, the function of each act of speech shapes speech production, which is a rather complex process. according to levelt (1989, pp. 4–14), speech production involves four stages (originally conceptualized as “processing components”) that depend heavily on “knowledge stores”: (a) conceptualization (i.e., preverbal message generation relying on situational knowledge and content knowledge); (b) formulation, including grammatical and phonological encoding relying on lexical knowledge; (c) articulation (i.e., execution of the phonetic plan by three sets of muscles involving up to 100 different muscles) resulting in overt speech; and (d) self-monitoring (i.e., the normal components of normal language comprehension relying on lexical knowledge). interestingly, the model includes the notion of inner speech, which is the product of the second phase. the model does not include writing as an alternation to articulation, but speech can be encoded into the visual or tactile form in addition to the auditory form. as a result, people manage to produce two to three words per second as a part of fluent conversation, and overtly naming a clear picture of an object can be initiated within 600 ms after the appearance of the picture (levelt, roelofs, & meyer, 1999). in fact, the generation of inner speech may be somewhat ahead of articulation. to cope with asynchrony, it is necessary for the phonetic plan to be stored. the storage mechanism is referred to as the articulatory buffer. it is important to note that these actions tax the speaker’s information-processing capacity, including working memory; in addition, speech production is delayed compared to recognition by the visual system. (recall that object recognition requires 100 ms processing time.) 3.2 two dominant approaches to verbal reports in a research context, verbal reports are heavily shaped by the context in which they are produced, and verbal reports serve various functions in different research traditions. this can be highlighted by comparing two dominant approaches to verbal reports: ericsson’s protocol analysis and chi’s explanations. ericsson and simon’s “verbal reports as data” (1980), with over 13,800 google citations and 1,619 web of science citations, appears to be the most influential piece of work on verbal reports. based on google scholar, ericsson has been the most active author on verbal reports over the last 30 years. second to ericsson and simon’s article is an article by micheline chi: “quantifying qualitative analysis of verbal data: a practical guide.” this paper has over 1,490 google citations and over 480 web of science citations. these two approaches are also frequently used in the context of medical image interpretation. therefore, ericsson’s protocol analysis and chi’s explanations deserve sections of their own. helle | f l r 85 according to ericsson (2006, p. 227), the central assumption of protocol analysis is that it is possible to instruct people “to verbalize their thoughts in a manner that does not alter the sequence and content of thoughts mediating the completion of a task and therefore should reflect immediately available information during thinking”. using levelt’s terminology, ericsson is after “inner speech”. in other words, the purpose is to elicit concurrent, nonreactive reports of thinking to understand expert reasoning and performance. according to the expert performance approach, the best way to obtain valid and complete traces of expert thought is to strive to produce laboratory conditions that capture “the essence of expertise,” where participants perform tasks that are representative of the studied phenomenon and where verbalizations directly reflect the participants’ spontaneous thoughts that are generated while completing the task. the instructions can be as follows (ericsson & simon, 1993, p. 376): “in this experiment, we are interested in what you say to yourself as you perform some tasks that we give you. in order to do this we will ask you to talk aloud as you work on the problems. what i mean by talk aloud is that i want you to say out loud everything that you say to yourself silently. just act as if you are alone in the room speaking to yourself. if you are silent for any length of time i will remind you to keep talking aloud.” in contrast, the goal of chi’s explanations (1997) is to figure out what a learner knows based on what a learner says or does and how that knowledge influences the learner’s reasons. chi avoided giving detailed instructions on how verbal reports should be elicited. instead, she gave detailed instructions on the analysis of such reports. she stressed that one must first determine “what” the learner said (e.g., a set of propositions or concepts). however, after that, to determine the overall structure of knowledge representations, one must assess the relations between the set. for example, a learner with naïve conceptions can hold pieces of unrelated knowledge, or a learner’s knowledge set can be theory-like, meaning that the reasoning can be captured by a few principles. according to chi (1997), the method of coding and analysing verbal data consists of the following eight steps: 1. reducing or sampling the protocols 2. segmenting the reduced or sampled protocols (sometimes optional) 3. developing a coding scheme or formalism 4. operationalizing evidence in the coded protocols 5. depicting the mapped formalism (optional) 6. seeking a pattern in the mapped formalism 7. interpreting the patterns 8. repeating the entire process, perhaps adopting a different grain size (optional). according to chi (1997), there are five key differences between ericsson’s protocol analysis and her verbal analysis. first, there is a clear juxtaposition in the way the verbal reports are collected. ericsson and simon (1993) underlined that research participants are simply verbalizing the information they attend to while generating an answer to a problem instead of describing, explaining, justifying, or rationalizing their actions. second, there is a difference in focus. ericsson and simon (1994) were concerned with tapping the online process of problem solving or decision making, whereas chi was interested in capturing the participants’ knowledge representations. she even argued that the goal of protocol analysis is to test the a priori model rather than to uncover what the participants are actually doing. the third difference has to do with analytical procedures and workloads. according to chi (1997), in protocol analysis, coming up with the ideal template, which requires a cognitive task analysis, represents the majority of the workload. in contrast, in verbal analysis, the referents are unknown; in self-explanation data, one must determine what the participant is talking about (e.g., an inference, plan, or inquiry). fourth, the method of validation or testing is different for the two methods. in ericsson’s protocol analysis, the sequence of verbal utterances is simply compared to the ideal template. the validation of the protocol analysis is “the degree of match” between these two. in the verbal analysis method, validation is achieved by using statistical testing. for example, qualitatively different knowledge representations of different groups of participants can be checked against helle | f l r 86 the answers to some subject-specific questions. ericsson (2006) pointed out that task analysis can be applied to the analysis of think-aloud protocols. however, he added that it is also possible to examine the convergent validity established by different types of data, including reaction times, error rates, patterns of brain activation, and sequences of eye movements. ericsson’s protocol analysis has several advantages. an obvious advantage is that ericsson provided detailed instructions on how to collect data using the method. the other advantage is that ericsson and simon (1993) provided a wealth of evidence indicating that the method is not reactive (i.e., it does not alter the course of cognitive processing). the main disadvantages are the following: (a) as morita et al. (2008) noted, medical image interpretation involves an implicit process that is difficult to verbalize; (b) thinkingaloud generally slows down performance, which may disrupt the execution of dynamic tasks in particular; (c) in certain tasks, it has been shown to alter accuracy (russo, johnson, & stephens, 1989). the disadvantage of prompting for explanations in the middle of an activity is that it has been repeatedly shown to affect behaviour in multiple ways (ericsson & simon, 1993). a more recent study exploring visual search behaviour on different sets of web pages showed that prompting for explanations not only prolonged the task, but also led to more general distributed visual behaviour and the issuing of more commands to navigate within and between the web pages. in addition, mental workload increased (hertzum, hansen, & anderson, 2009). another disadvantage of chi’s explanations is the lack of clear instructions on how to collect “explanations.” if explanations are required, some form of retrospective reporting should be seriously considered. 3.3 retrospective reporting in fact, people can provide quite accurate retrospective reports for short tasks that take 5–10 seconds (ericsson, 2006). the instruction can be as follows: “can you please tell me what you were thinking during problem solving?” (van gog, paas, van merriënboer, & witte, 2005). in the context of medical image interpretation, it is also common to ask the participants to report on the findings and final diagnosis either orally or in writing. as patrick and james (2004) pointed out, ideally, verbal reports should be collected immediately after task completion while the participant’s short-term memory still holds relevant information. according to the authors, when there is a need to rely on the participant’s long-term memory, some type of retrieval cues should be designed. in the case of medical images, which require more than 5–10 seconds to interpret, showing the participants a dynamic presentation of their eye movements (and keyboard movements when applicable) would seem to be a viable cuing solution. there have been some noteworthy efforts to compare concurrent think-aloud with retrospective reports and cued reporting. in the context of troubleshooting electric circuits, van gog, paas, van merriënboer and witte (2005) conjectured that the methods would extract different types of information regarding process tracing. the authors did not report the duration of the troubleshooting tasks, but it seems safe to assume that the task durations exceeded the critical limit of 5–10 seconds. thus, it was not surprising that the concurrent think-aloud and cued retrospective reporting involving showing the participants their eye movements and keyboard strokes resulted in more information than retrospective reporting. the remarkable finding was that the cued retrospective method resulted in less theoretical meaning-making verbalizations (why utterances), whereas the concurrent method resulted in less metacognitive utterances. thus, in addition to the context and population, one needs to consider carefully the type of information one is seeking. the advantage of retrospective reporting is that the method of verbalization does not interfere with task completion. if the task is of a very short duration (under 10 seconds), accurate reports of thinking processes can be expected. the advantage of cued retrospective reporting is that it does not interfere with task performance. however, it may be that the cues are not sufficient enough to recover all task-related information. helle | f l r 87 it is worth noting that morita et al. (2008) showed that it can be worthwhile to triangulate verbal reports obtained through different methods. their results indicated that experts use more conceptual words in thinking-aloud through a visual task, but they use more perceptual words when compared to novices in the writing of the report. the interpretation of the finding was that the development of expertise is based on an ability to build connections between percepts and concepts. 4. critical examination of studies combining eye tracking and verbal reports in the context of medical image perception although there are many arguments for promoting the combination of eye tracking and verbal reports, combining eye tracking and verbal reports is easier said than done, as can be seen from the following studies employing concurrent think-aloud. concurrent think-aloud attempts to capture nonreactive verbal reports of thinking (ericsson, 2006). the notion of “nonreactivity” means that the execution of the primary task is not affected, except for the fact that it may be prolonged. the participants are asked to perform a task while uttering briefly what spontaneously comes to mind. in other words, it aims to “vocalize inner speech.” ericsson emphasized numerous times that participants should be talking to themselves, not explaining what they are doing or why because it has been repeatedly shown that the act of explaining can seriously interfere with the task the investigators are trying to model. the first efforts to triangulate different sources of data obtained from different studies date back to the year 2000. berbaum et al. (2000) were interested in conducting a congenially designed laboratory experiment to determine if satisfaction of search is because of recognition error or because of decision error by two different methods (eye tracking versus protocol analysis). the design involved inserting artificial lesions in an image to see if it decreases the detection of native lesions, indicating satisfaction of search (sos). an earlier study employing eye tracking had indicated that inserting artificial lesions to certain images decreased the reporting of native lesions on those images. in the new experiment, berbaum et al. (2000) discovered two important things: first, the think-aloud condition served to eliminate the satisfaction of the search effect. second, the two methods provided contradictory results: the eye-tracking study suggested sos was because of decision error, whereas the think-aloud study suggested that sos was because of recognition error. the authors concluded that protocol analysis is limited in its ability to differentiate between search error and recognition error. on the other hand, there are perils in assuming that a lesion has been recognized based on dwell time alone. thus, it was hard to reconcile the fact that the two studies produced contradictory findings. also, the fact that the think-aloud procedure affected performance on the primary task casts doubts on the integrity of the entire study: it is difficult to argue that the think-aloud procedure was nonreactive. we speculate that the reactivity was because of the instructions given; the observers were instructed to use a finger to point to where they were looking at and to verbalize the structures and the features they were looking for. these early efforts highlight the difficulties investigators can experience in applying concurrent think-aloud and method triangulation. in fact, the first fundamental issue to consider is whether the concurrent think-aloud condition interferes with the primary task. to our knowledge, only a single study has been conducted on this issue in the context of medical image interpretation. littlefair, brennan, reed, williams, and pietrzyk (2012) explored whether the think-aloud condition affects pulmonary node detection; they did this using a withinsubjects design with seven participants, two viewing sessions with a “wash-out period” separating them, and a set of 30 two-dimensional radiographs. half of the radiographs contained a single artificial nodule, and the rest were non-nodular. the participants were informed that the radiographs may contain a single nodule. no time limit was set for viewing. performance was evaluated in terms of sensitivity, specificity, and roc measures, including multicase multireader roc auc analysis. in addition, the participants’ eye movements were tracked to compare, for example, fixations of areas of interest and time to fixate on areas of interest. helle | f l r 88 results indicated that only half of the nodules ended up correctly localized, indicating an absence of ceiling effects. there were no differences in performance under the two conditions, with the exception of confidence ratings (in the ta condition, the subjects were less confident) and task duration. the latter result was statistically significant. the results are well-aligned with ericsson and simon’s (1993) theoretical account. concurrent think-aloud has also been used, rather surprisingly, in the context of dynamic stimuli. an alternative approach situated in the context of fish locomotion can be seen in jarodzka, scheiter, gerjets, and van gog (2010). balslev et al. (2012) used think-aloud in the context of viewing films depicting infants with seizures and conditions resembling seizures. balslev et al. (2012) had their participants (medical students, residents, and experts) think-aloud while diagnosing the infant seizures presented in the short films, which lasted anywhere from 26–49 seconds. the films were looped and repeated until the observer wished to stop viewing. not surprisingly, the experts scored higher in diagnostic accuracy and spent relatively more time viewing task-relevant features. a content analysis of the verbal accounts revealed that experts engaged more, in relative terms, exploring the material and spent more time building and evaluating hypotheses. this pattern, in turn, explained why the experts returned to the areas of interest. this study showed how the combined use of eye movements and verbal reports can lead to a better understanding of medical image interpretation. finally, li, pelz, alm, and haake (2012) attempted to integrate eye movement information completely with the concurrent verbalizations of a group of dermatologists who differed in their level of training; the groups were asked to observe 42 two-dimensional dermatological images. subsequently, they developed a hierarchical probabilistic framework to extract unique and common eye movement patterns among multiple subjects within each expertise group. the idea was to map specific eye movement patterns to certain cognitions, such as identifying the primary morphology. although the study is a remarkably ambitious endeavor to integrate eye tracking and concurrent verbalizations, as a process-tracing study, the work suffers from the implementation of the concurrent verbalizations. the novices were requested to provide a detailed description of the materials “as if describing to their doctors over the phone.” the medical professionals were instructed to examine and describe the findings to students “as if teaching.” (ibid., p. 395) these are clear violations of the principle of focusing on inner thought, and asking these questions may have seriously interfered with the primary task of image interpretation. therefore, the mapping solution presented may have limited value. retrospective verbalization is a suitable option for very short tasks where there is simply not enough time to verbalize. jaarsma, jarodzka, nap, van merriënboer, and boshuizen (2014) applied a heavily time-constrained research design. the authors presented two-dimensional microscopic images of colon tissue to a group of clinical pathologists, pathology residents, and medical students. the viewing of the images was constrained to 2 seconds. the participants’ eye movements were registered, along with their post hoc verbal accounts of what they had seen. (the authors did not use the expression “retrospective think-aloud”; instead, they referred to post hoc verbalizations.) the investigators analyzed the two sources of data separately. the verbal accounts were analyzed through an elaborate content analysis. the most interesting findings related to the differences between the clinical pathologists and the residents: in their search, the clinical pathologists tended to rely on what they had already seen, further studying the image for other abnormalities, whereas the residents tended to double-check their initial findings. in their post hoc verbalizations, the clinical pathologists focused on the typicality of the tissue, whereas the residents concentrated on naming pathologies. this study showed that important insights can be gleaned by combining eye movements and a form of retrospective verbal reports. interestingly, not a single study could be found where the authors reported collecting the data by using chi’s explanations as the method. instead, going through the literature revealed several studies where the investigators collected the data using concurrent think-aloud and then referring to chi (1997) in the analysis phase (azevedo et al., 2007; van der gijp et al., 2014; van der gijp et al., 2015). it is not unusual to find studies using concurrent verbalizations, which end up gathering explanations (e.g., li et al., 2012; van der gijp et al., 2015). helle | f l r 89 5. conclusions the integration of eye tracking and verbal reports is intuitively appealing, and as illustrated in this methodological article, interesting insights can be gleaned by adopting a mixed-methods approach. however, before collecting such data, important decisions must be made. the first critical decision is timing, that is whether to collect concurrent or retrospective data. in the case of concurrent think-aloud, a decision must be made whether to follow ericsson and simon’s or chi’s advice. in the case of retrospective reports, one must decide whether to play back the eye movements to the observer to aid the retrieval of information from longterm memory. based on this methodological analysis, some methodological issues appear to be solved, whereas others require further investigation. we argue that two issues are solved: first, retrospective reporting without cuing is suitable for perceptual tasks of a very short duration (<10 seconds). the advantages are the following: (a) one can be certain that verbalization does not interfere with the primary task; (b) the observer’s verbalization is not constrained to the speed of the visual system; and (c) it is safe to assume that information is still available in short-term memory. second, there exists a compelling body of literature indicating that for the purposes of process tracing, ericsson’s nonreactive method is superior to the idea of soliciting direct explanations from the observers during task execution because soliciting explanations tends to interfere with the primary task. it is hard to see what the purpose of process tracing would be if the research method results in a substantial change in the primary activity. other issues remain underexplored. first, when tasks are longer than 10 seconds, the pros and cons of ericsson’s concurrent think-aloud versus cued retrospective reporting need to be weighed against each other. in the study by van gog et al. (2005) in the context of troubleshooting electric circuits, the concurrent think-aloud condition produced more theoretical expressions, which were to some extent lost in the stimulated recall condition. we emphasize that this issue has not been explored in the context of medical image interpretation. more fundamentally, there is a need for more studies to be conducted to show that perceptual tasks, such as viewing an x-ray, are not affected by the think-aloud condition. the study by russo et al. (1989) showed that even when experimenters stick meticulously to ericsson and simon’s instructions, the think-aloud condition may affect task performance. russo et al. (1989, p. 758) concluded that “protocol validity should be based on an empirical check rather than theory-based assurances”. what is proposed is a research agenda with two goals: (a) to further explore if think-aloud affects performance in a range of image interpretation tasks; (b) to compare the type of information obtained by concurrent think-aloud and cued retrospective reporting with different types of material (two-dimensional, volumetric, video). it would also be useful to include observers with varying levels of experience. keypoints before attempting to combine eye-tracking data and verbal accounts, important decisions must be made regarding the timing of the verbalizations and possible cuing. ericsson’s concurrent think-aloud is deemed superior to eliciting explanations from the observers during task performance. retrospective think-aloud is suitable for tasks of a very short duration (<10 seconds). a research agenda is proposed for investigating the methodological issues that remain unsolved. helle | f l r 90 acknowledgements the author wishes to thank dr. raymond bertram particularly for advice on writing about speech production. in addition, she is grateful to the two anonymous reviewers for providing exceptionally insightful and constructive feedback on the first version of the manuscript. references al-moteri, m. o., symmons, m., plummer, v., & cooper, s. (2017). eye tracking to investigate cue processing in medical decision-making: a scoping review. computers in human behavior, 66, 52-66. doi.10.1016/j.chb.2016.09.022 azevedo, r., faremo, s., & lajoie, s. p. (2007, january). expert-novice differences in mammogram interpretation. proceedings of the cognitive science society, 29(29). balslev, t., jarodzka, h., holmqvist, k., de grave, w., muijtjens, a. m., eika, b., ... scherpbier, a. j. (2012). visual expertise in paediatric neurology. european journal of paediatric neurology, 16(2), 161-166. doi.10.1016/j.ejpn.2011.07.004 berbaum, k. s., franken, e. a., dorfman, d. d., caldwell, r. t., & krupinski, e. a. (2000). role of faulty decision making in the satisfaction of search effect in chest radiography. academic radiology, 7(12), 1098-1106. doi.10.1016/s1076-6332(00)80063-x berbaum, k. s., franken, e. a., dorfman, d. d., miller, e. m., caldwell, r. t., kuehn, d. m., & berbaum, m. l. (1998). role of faulty visual search in the satisfaction of search effect in chest radiography. academic radiology, 5(1), 9-19. doi.10.1016/s1076-6332(98)80006-8 bertram, r., helle, l., kaakinen, j. k., & svedström, e. (2013). the effect of expertise on eye movement behaviour in medical image perception. plos one, 8(6), e66169. doi.10.1371/journal.pone.0066169 blondon, k., wipfli, r., & lovis, c. (2015). use of eye-tracking technology in clinical reasoning: a systematic review. studies in health technology and informatics, 210, 90-94. doi.10.3233/978-161499-512-8-90 chi, m. t. (1997). quantifying qualitative analyses of verbal data: a practical guide. the journal of the learning sciences, 6(3), 271-315. doi.10.1207/s15327809jls0603_1 cowan, n. (2010). the magical mystery four how is working memory capacity limited and why. current directions in psychological science, 19(1), 51-57. doi.10.1177/0963721409359277 ericsson, k. a. (2006). protocol analysis and expert thought: concurrent verbalizations of thinking during experts’ performance on representative tasks. in k. a. ericsson, n. charness, p. j. feltovich, & r. r. hoffman (eds.), the cambridge handbook of expertise and expert performance (pp. 223-241). cambridge, ma: cambridge university press. ericsson, k. a., & simon, h. a. (1980). verbal reports as data. psychological review, 87(3), 215. ericsson, k. a., & simon, h. a. (1993). protocol analysis. verbal reports as data. cambridge, ma: mit press. gegenfurtner, a., kok, e., geel, k., bruin, a., jarodzka, h., szulewski, a., & merriënboer, j. j. (2017). the challenges of studying visual expertise in medical image diagnosis. medical education, 51(1), 97-104. doi.10.1111/medu.13205 gegenfurtner, a., siewiorek, a., lehtinen, e., & säljö, r. (2013). assessing the quality of expertise differences in the comprehension of medical visualizations. vocations and learning, 6(1), 37-54. doi.10.1007/s12186-012-9088-7 goodale, m. a. (1998). visuomotor control: where does vision end and action begin? current biology, 8(14), r489-r491. doi.10.1016/s0960-9822(98)70314-8 goodale, m. a., & milner, a. d. (1992). separate visual pathways for perception and action. trends in neurosciences, 15(1), 20-25. helle | f l r 91 hertzum, m., hansen, k. d., & andersen, h. h. (2009). scrutinising usability evaluation: does thinking aloud affect behaviour and mental workload? behaviour & information technology, 28(2), 165-181. doi.10.1080/01449290701773842 jaarsma, t., jarodzka, h., nap, m., merrienboer, j. j., & boshuizen, h. (2014). expertise under the microscope: processing histopathological slides. medical education, 48(3), 292-300. doi.10.1111/medu.12385 jarodzka, h., scheiter, k., gerjets, p., & van gog, t. (2010). in the eyes of the beholder: how experts and novices interpret dynamic stimuli. learning and instruction, 20(2), 146-154. doi.10.1016/j.learninstruc.2009.02.019 kundel, h. l., nodine, c. f., & carmody, d. (1978). visual scanning, pattern recognition and decision making in pulmonary nodule detection. investigative radiology, 13(3), 175-181. kirchner, h., & thorpe, s. j. (2006). ultra-rapid object detection with saccadic eye movements: visual processing speed revisited. vision research, 46(11), 1762-1776. doi.10.1016/j.visres.2005.10.002 krupinski, e. a., tillack, a. a., richter, l., henderson, j. t., bhattacharyya, a. k., scott, k. m., … weinstein, r. s. (2006). eye-movement study and human performance using telepathology virtual slides. implications for medical education and differences with experience. human pathology, 37(12), 1543-1556. doi.10.1016/j.humpath.2006.08.024 lesgold, a. m., feltovich, p. j., glaser, r., & wang, y. (1981). the acquisition of perceptual diagnostic skill in radiology (no. lrdc-81/pds-1). pittsburgh university learning research and development center. levelt, w. j. m. (1989). speaking. from intention to articulation. cambridge, ma: mit press. levelt, w. j., roelofs, a., & meyer, a. s. (1999). a theory of lexical access in speech production. behavioral and brain sciences, 22(1), 1-38. li, r., pelz, j., shi, p., alm, c. o., & haake, a. r. (2012, march). learning eye movement patterns for characterization of perceptual expertise. in proceedings of the symposium on eye tracking research and applications (pp. 393-396). acm. littlefair, s., brennan, p., reed, w., williams, m., & pietrzyk, m. w. (2012, february). does the thinking aloud condition affect the search for pulmonary nodules? in spie medical imaging (pp. 83181a83181a). bellingham, wa: international society for optics and photonics. morita, j., miwa, k., kitasaka, t., mori, k., suenaga, y., iwano, s., ... ishigaki, t. (2008). interactions of perceptual and conceptual processing: expertise in medical image diagnosis. international journal of human-computer studies, 66(5), 370-390. doi.10.1016/j.ijhcs.2007.11.004 nodine, c. f., & kundel, h. l. (1987). using eye movements to study visual search and to improve tumor detection. radiographics, 7(6), 1241-1250. patrick, j., & james, n. (2004). process tracing of complex cognitive work tasks. journal of occupational and organizational psychology, 77(2), 259-280. rubin, g. d., roos, j. e., tall, m., harrawood, b., bag, s., ly, d. l., ... choudhury, r. k. (2014). characterizing search, recognition, and decision in the detection of lung nodules on ct scans: elucidation with eye tracking. radiology, 274(1), 276-286. doi.10.1148/radiol.14132918 russo, j. e., johnson, e. j., & stephens, d. l. (1989). the validity of verbal protocols. memory & cognition, 17(6), 759-769. salthouse, t. a., & ellis, c. l. (1980). determinants of eye-fixation duration. the american journal of psychology, 207-234. van der gijp, a., van der schaaf, m. f., van der schaaf, i. c., huige, j. c. b. m., ravesloot, c. j., van schaik, j. p. j., & ten cate, t. j. (2014). interpretation of radiological images: towards a framework of knowledge and skills. advances in health sciences education, 19(4), 565-580. doi.10.1007/s10459013-9488-y van der gijp, a., ravesloot, c. j., van der schaaf, m. f., van der schaaf, i. c., huige, j. c., vincken, k. l., ... van schaik, j. p. (2015). volumetric and two-dimensional image interpretation show different cognitive processes in learners. academic radiology, 22(5), 632-639. doi.10.1016/j.ejrad.2014.12.015 helle | f l r 92 van der gijp, a., ravesloot, c. j., jarodzka, h., van der schaaf, m. f., van der schaaf, i. c., van schaik, j. p. j., & ten cate, t. j. (2016). how visual search relates to visual diagnostic performance: a narrative systematic review of eye-tracking research in radiology. advances in health sciences education, doi. 10.1007/s/10459-016-9698-1 van gog, t., paas, f., van merriënboer, j. j., & witte, p. (2005). uncovering the problem-solving process: cued retrospective reporting versus concurrent and retrospective reporting. journal of experimental psychology: applied, 11(4), 237. doi.10.1037/1076-898x.11.4.237 vanni, s. (2004). näkötiedon käsittely aivokuoressa [processing of visual data in the cerebral cortex]. duodecim, 120, 2653-2662. vanni, s., & heikkinen, h. (2015). onko aivoissamme käyttämätöntä kapasiteettia? [is there unused capacity in our brain?] duodecim, 131, 1644-1649. wolfe, j. m., evans, k. k., drew, t., aizenman, a., & josephs, e. (2015). how do radiologists use the human search engine? radiation protection dosimetry, 501. doi.10.1093/rpd/ncv501 wolfe, j. m., võ, m. l. h., evans, k. k., & greene, m. r. (2011). visual search in scenes involves selective and nonselective pathways. trends in cognitive sciences, 15(2), 77-84. doi.10.1016/j.tics.2010.12.001 microsoft word soto et al publication.docx frontline learning research vol. 6 no. 1 (2018) 31-52 issn 2295-3159 corresponding author: christian soto, department of spanish, university of concepción, victor lamas 1290, concepción, chile, 4030000. e-mail: christiansoto@udec.cl doi: 10.14786/flr.v6i1.328 a deeper understanding of metacomprehension in reading: development of a new multidimensional tool christian sotoa, antonio p. gutierrez de blumeb, rodrigo asúnc, matthew jacovina, and claudio vásquezd auniversity of concepción, chile; b georgia southern university, united states; cuniversity of chile; dautonomous university of chile article received 11 september 2017 / article revised 23 february / accepted 4 april / available online may abstract the purpose of this research endeavor was to develop and validate a new measurement tool predicated on previous research to assess learners’ metacomprehension during reading. in two separate studies with chilean undergraduate students (n = 923), we demonstrate the versatility and utility of our proposed metacomprehension inventory (mi). in study 1, we provide empirical support for the psychometric soundness and construct validity of the mi. in study 2, we provide evidence of the measurement invariance of the mi between males and females. results of study 1 revealed the hypothesized factor structure of the mi is sound, with high factor loadings, excellent model fit, and moderate-to-strong inter-factor correlations. study 2 results indicated that the mi is interpreted similarly by both males and females, as factor loadings were largely statistically identical across the two groups. we discuss implications of our proposed mi for theory and applied research. keywords: metacomprehension; metacognition; reading strategies; factor analysis; validity soto et al 32 | f l r 1. introduction many researchers propose that reading performance is improved when effective metacognitive strategies are implemented (beck & mckeown, 1998; beck & mckeown, 2006; beck & mckeown, 2009; paris & jacobs, 1984; sandora, beck, & mckeown, 1999), such as selecting strategies from one’s repertoire, effectively executing said strategies, and knowing when and why certain strategies do or do not apply, given task demands. this includes both the reader's ability to understand and to apply the necessary strategies during reading. the concept of metacognition has been studied since the 1970s, led by the work of flavell and wellman (1977), who conducted metamemory studies. metacognition studies were then extended to the field of reading and to applied educational work (flavell, 1979; garner, 1987, brown, armbruster, & baker, 1986; paris & paris, 2001, baker, 2002; pressley & block, 2002; block & pressley, 2007; hacker, dunlosky & graesser, 2009, azevedo & aleven, 2013), such as how to improve self-regulated learning skills in the classroom or while studying at home during learning episodes. a crucial challenge for studying metacognition is how to measure its components. several questions arise from this challenge, particularly when studying how metacognitive knowledge, skills, and strategies influence reading comprehension. when considering metacomprehension¸ researchers must consider questions such as how metacognition influences a reader's understanding of a text, how knowledge of metacognition relates to and differs from the enactment of metacognitive strategies, and which methodologies are most appropriate. in this paper, we describe research related to these questions and propose a new metacomprehension inventory that attempts to use what has been learned about metacognition and reading to build a thorough representation of readers' metacognitive knowledge and their conscious use of strategies. the relationship between metacognition and cognition for reading processes historically has not been very clear. a primary complexity has been determining the relative weight and importance of each component of the metacognitive process (jacobs & paris, 1987). a second complexity is that readers must have some awareness of their metacognitive processes in order for it to be accurately measured (jiménez, puente, alvarado, & arrebillaga, 2009). a related third complexity concerns the methods to measure some of the specific processes of metacognition (jiménez et al., 2009), as each method has its limitations and advantages, and is aimed at examining specific aspects of metacognition (pressley & afflerbach, 1995). finally, a fourth complexity has to do with the empirical evidence on the relations between metacognition and reading comprehension, which underpins the question of which methods are most appropriate for studying these processes together. the different methods to assess metacomprehension give indications of how the different components of metacognition and reading may be linked. findings that link reading comprehension training with metacognitive measures are also useful for understanding their relations. methods such as collecting performance/confidence judgments or asking readers to think aloud have been very useful in resolving some of these complexities and ambiguities (bol & hacker, 2001; bol, hacker, o’shea, & allen, 2005; dunlosky, rawson, & mcdonald, 2002; dunlosky, griffin, thiede & wiley, 2005; dunlosky & lipko, 2007; hacker, bol, & bahbahani, 2008; thiede, griffin, wiley, & redford, 2009; gutierrez & schraw, 2015). using inventories, however, has been one of the most common methods, largely because of their practicality (e.g., jacobs & paris, 1987; swanson, 1990; pintrich & de groot, 1990; schmith, 1990, schraw & dennison, 1994; mokhtari & reichard, 2002). these inventories are normally short and easy to administer during different stages of reading. a single thorough metacomprehension inventory could be an excellent tool by itself or when combined with other methods to delve deeper in exploring the relation between metacognition and reading comprehension. for this reason, it is vital for inventories to evolve as new perspectives appear in the field. accordingly, a number of tools have been developed across different languages (see table 1). soto et al 33 | f l r table 1 metacognitive inventories used in psychological research name citation index of reading awareness (ira) jacobs & paris, 1987 metacognitive questionnaire (mq) swanson, 1990 motivated strategies for learning questionnaire pintrich & de groot, 1990 metacomprehension strategy index (msi) schmith, 1990 metacognitive awareness inventory (mai) schraw & dennison, 1994 reading strategy use (rsu) pereira-laird & deane, 1997 metacognitive awaraness of reading strategies inventory (marsi) mokhtari & reichard, 2002 cuestionario de metacomprensión lectora a peronard, crespo, velásquez, &viramonte, 2002 escala de conciencia lectora. (escola) b jiménez, v., puente, a., alvarado, j., & arrebillaga, l. (2009) escala de evaluación de la autorregulación del aprendizaje a partir de textos (aratex-r) c núñez, amieiro, alvarez, garcía & dobarro, 2015 revised metacomprehension scale (rmcs) zabrucky, moore, lin, & cummings, 2015 note. a reading metacomprehension inventory; b reading awareness scale; c evaluation of self-regulation of learning from texts scale. most of these inventories or questionnaires explore people's knowledge (declarative, procedural, conditional) or control/ regulation of the cognitive processes (planning, monitoring, debugging, information management, and evaluation) that are important while reading. questions usually focus on planning, monitoring and evaluation, or ask about strategies and regulatory processes that influence comprehension. some inventories are focused specifically on reading, whereas others are designed to be more general measures of metacognition across multiple domains. in the following paragraphs, we briefly describe some of these instruments. one of the most popular inventories is the metacognitive awareness inventory (mai; schraw & dennison, 1994), which includes 52 questions that assess separate knowledge of cognition (declarative, procedural, and conditional knowledge) and regulation of cognition (planning, monitoring, debugging, information management, and evaluation) factors which are purportedly domain general. sample items include “i try to use strategies that have worked in the past” (procedural knowledge); “i reevaluate my assumptions when i get confused” (debugging); and “i ask myself if i have considered all options after i solve a problem” (information management), answered using a 5-point likert scale. schraw and dennison (1994) reported—in two separate experiments—the mai to have a stable and consistent two-factor structure. another influential inventory is the metacognitive awareness of reading strategies inventory (marsi). marsi is a reading awareness scale, measuring metacomprehension skills in students from 6th through 12th grade. it consists of 30 items and uses a 5-point likert scale. marsi assesses three different soto et al 34 | f l r dimensions of metacomprehension: global reading strategies, problem solving strategies and support reading strategies. according to the authors, the first factor (global reading strategies) contains 13 items about readers’ intentional strategies for analyzing a text at a global level, “setting the stage for the reading act” the second factor, problem-solving strategies, contains 8 items about readers’ strategies for repairing problematic comprehension, particularly when a text is challenging. the third factor, support reading strategies, contains 9 items about strategies that readers employ that use outside reference materials, notes, or consulting others to check their understanding. sample items include “i skim the text first by noting characteristics like length and organization” (global reading strategies); “i try to get back on track when i lose concentration” (problem solving strategies); “i discuss what i read with others to check my understanding” (support reading strategies). escala de conciencia lectora (escola; i.e., reading awareness scale) is a reading awareness scale, measuring metacomprehension skills in spanish speaking students between 8 and 13 years of age. it consists of 56 questions with three possible answers for each. escola assesses three different dimensions of metacomprehension: planning, monitoring and evaluation. each question has a correct answer, and an answer that gives partial credit. thus, the measure is designed to capture students’ metacognitive competency. planning questions measure readers’ knowledge of how to select the most appropriate reading strategies to achieve their reading goal. monitoring questions measure students’ ability to adjust attention and effort during the reading task. evaluation questions measure students’ awareness about whether they appropriately understood the text. sample items (translated by one of the authors of this paper) include “before you start reading, what do you do to help in the reading process? a) i do not make any plans, just start reading [0 points], b) i consider why i'm going to read [2 points], c) i choose a comfortable place to read [1 point]” (planning); “if you are reading a book and find a paragraph difficult to understand, what do you do? a) i stop to think about the problem and how to fix it [2 points], b) i do not keep reading because i cannot solve the problem [0 points], c) i continue to read to see if the meaning is clarified later [1 point]” (monitoring); “in carrying out the activity of reading: a) i think it is useful to assess whether i understood what was written [2 points], b) i think that the evaluation is good but needs to be made by an older person [1 point], c) i do not think that after reading assessment is no longer useful [0 points]” (evaluation). escola's monitoring dimension is similar to the concept of regulation during comprehension, whereas the evaluation dimension is similar to the self-assessment of comprehension. escola has been used and validated in spanish-speaking populations in both spain and argentina, and it has demonstrated sound psychometric properties (jiménez, puente, alvarado, & arrebillaga, 2009; puente, jiménez, & alvarado, 2009). each inventory has a specific focus depending on the emphasis and framework. some help to identify a specific component of the metacomprehension process, emphasizing a person’s awareness of these processes. for example, the mai considers very interesting distinctions between different learning processes, specifically between knowledge and control, as well as some interesting sub-mechanisms of control like debugging. however, this inventory largely focuses on learning in general and not specifically about reading comprehension. marsi, on the other hand, proposes a particular conceptualization of the metacognitive strategies involved in reading, including global reading strategies, problem solving strategies and support reading strategies. these concepts, however, do not consider distinct time points during reading (e.g., reflection during and after reading), and some of the strategies are not commonly used by students. escola differentiates between the different time points during reading comprehension (planning, monitoring and evaluation) but the questions do not clearly situate a reader in terms of his/her current metacognitive knowledge or use of strategies. for example, in one planning question that asks what tasks a reader completes before starting to read, the possible responses are that the reader “does not make any plans,” “[chooses] a comfortable place to read,” and “i consider why i'm going to read,” and these answers receive 1, 2, and 3 points respectively. although it is clear that the third response is the most relevant to planning, it is less clear how choosing a comfortable place to read relates to planning. likert-scale type questions are therefore the preferred choice due to the nature of the variables. soto et al 35 | f l r we agree with schraw and dennison (1994) who, through the mai, make a distinction between the knowledge of cognition as an initial process, and the regulation of cognition as a subsequent process. we also agree with making distinctions between different times during the reading comprehension process. however, there is some confusion between evaluation and regulation processes, and typically their distinction is not deeply considered in metacognitive inventories, despite its importance in understanding the metacognition of reading comprehension in adult populations. this becomes an increasingly important distinction when considering how readers evaluate and regulate during different times or in different situations during the reading comprehension process. 1.1. critical concepts for a new inventory researchers have described monitoring as including different processes. according to hacker and his colleagues (hacker, 1998; keener & hacker, 2012), monitoring has often been discussed as including both the processes of evaluation and regulation. from this perspective, readers’ monitoring would be said to be successful only when they, for example, both noticed that they did not understand some part of the text and also deployed cognitive effort to remedy their understanding (e.g., by rereading the section). an alternative view of monitoring is that there is a clear distinction between the processes of monitoring (e.g., evaluating comprehension), and regulation (e.g., doing something to fix comprehension deficits) (boekaerts, 1999; schraw & dennison, 1994; schraw & moshman, 1995). we adopt this latter view, as it is more useful in analyzing metacomprehension and its influence on comprehension. that is, failures can occur either at the monitoring or regulation stages, and both are interesting for the study of metacomprehension. however, studies of regulation and monitoring not only have different methods and sources of research but would reflect different dimensions of the metacomprehension process. in fact, a failure to detect inconsistencies might not necessarily indicate failures in monitoring understanding, but in this case the reader might be monitoring for purposes unrelated to detecting errors (hacker et al., 1994). there is a close relation between monitoring and regulation, because the regulation is implemented after a reader’s preliminary assessment of his/her understanding. the regulatory process acts as the reader takes action to repair or improve his/her understanding, such as by rereading a part of the text that generated confusion. in short, regulation includes adjustment operations during the comprehension process. evaluation can happen in different stages of reading, both during reading and after a cycle of reading. therefore, an effective inventory should consider these different time points. on the other hand, regulation is a dynamic process that depends on evaluation, but may be implemented for different reasons, such as fixing faulty comprehension or deepening understanding. because traditionally studies on regulation have focused on error detection when reading materials contain inconsistencies, there has been little consideration of regulation as a mechanism to understand ideas more deeply during reading. following from this point of view, there should regularly be situations in which readers decide (consciously or subconsciously) to improve their mental representation using different strategies, even when there is no inconsistency in comprehension. as we describe next, this perspective can be incorporated in a metacomprehension inventory that can assess each one of these components separately, giving a new perspective to both theoretical and applied research, enriching the extant research between these metacognitive components and the reading comprehension measures using this type of tool (azevedo, 2009). 1.2. metacomprehension inventory (mi) the metacomprehension inventory (mi) described in this paper attempts to combine several of the strengths of the previously described inventories, while minimizing the impact of their weaknesses. its goal is to tap into readers’ evaluative and regulatory processes at different time points during reading. throughout, the items focus on specific strategies that readers may be employing. this allows the inventory soto et al 36 | f l r to probe readers’ conscious, strategic behaviors and does not rely on more general questions that might simply ask if readers use strategies at all. further, the inventory uses the distinction made by schraw and dennison (1994) between knowledge about cognition and the control/regulation of cognition. within knowledge of cognition, we consider the three sub-processes described by schraw and dennison: declarative knowledge about a reader’s personal qualities and strategies in general, procedural knowledge about how to use strategies, and conditional knowledge about when and for what purpose to use strategies given task demands. below, we outline the different dimensions, sub-dimensions and components of the mi, that correspond to different stages of the reading process, and are measured with different items. • knowledge about cognition (kac). this dimension considers three components, which we assume to be strongly correlated, so that they will not form independent dimensions: a) declarative knowledge: these items refer to readers’ knowledge about strategies (“i know which might be the characteristics of a good reader”), and about their knowledge of their personal reading skills (“i know the strong and weak points of my reading skills”); b) procedural knowledge: these items refer to readers’ knowledge of how to employ strategies for improved comprehension (“i know how to deal with a text to make it easier to understand for me”); conditional knowledge: these items refer to readers’ knowledge of when it is appropriate to employ comprehension strategies (“i know how to overcome a difficulty when i have problems understanding a text”). control of cognition (coc): this dimension is more complex and includes three sub-dimensions which we propose exhibit some degree of statistical independence from one other: • control of cognition-planning (coc-p). these items refer to activities that occur prior to beginning a reading task such as “i question myself about the topic before starting reading.” and “when i prepare myself to read, i organize my time and reading activities to finish the task on time.” these items relate to control processes that a reader undergoes before beginning the task insofar as they ask whether readers strategically plan their time based on their estimates of task difficulty. • control of cognition-evaluation (coc-e). these items refer to readers’ tendency and ability to examine their reading understanding. we assume that this ability subsumes two temporal sub-dimensions: o control of cognition-evaluation during reading (coc-edr). these items refer to readers’ tendency and ability to examine their understanding during reading (e.g., “while i’m reading i can determine how much i’m understanding”). o control of cognition-evaluation after reading (coc-ear). these items refer to a readers’ tendency and ability to examine their understanding after completing a phase of reading (e.g., “when i finish reading a text i can know which part was more confusing to me”). these questions do not make it explicit what is meant by “after reading,” and so each reader may develop his/her own interpretation (i.e., it is receptive to individual differences). the goal of these items is to gauge reflection as an anterior process as opposed to reflection in the moment (i.e., “during reading”) or posterior. • control of cognition-regulation (coc-r). these items refer to readers’ tendency and ability to regulate their reading understanding. we propose that this ability also subsumes two sub-dimensions: o control of cognition-regulation after problematic understanding (cocrpu). these items refer to readers’ tendency and ability to engage processes and activities soto et al 37 | f l r to repair their understanding when they are confused, feel challenged, or notice a discrepancy (e.g., “when i find some text information strange i stop and read the paragraph more than once”). o control of cognition-regulation to deepen comprehension (coc-rdc): these items refer to readers’ tendency and ability to use reading strategies to improve their comprehension on a regular basis in an attempt to enhance their understanding (e.g., “when i read, i try to explain the text to myself using my own words”). these items explicitly do not mention any difficulties in understanding, and are meant to capture readers’ strategic regulatory behaviors that are used normally when reading. considering readers’ metacognitive knowledge and control processes separately, and at different time points, is crucial for informing interventions and building a theoretical understanding of metacomprehension. as we stated earlier, failures can occur for different reasons and at different stages, and pinpointing the nature and timing of these failures is critical for repairing readers’ comprehension skills and determining appropriate remedial strategies. theoretically, a complete model of metacomprehension should describe when strategic processes are most important; that is, a model must predict when automatic monitoring processes are unlikely to be successful or are potentially not being enacted at all. thus, our proposed inventory, which involves proposing a fourth-order factorial model, is intended to be combined with behavioral data to inform such a model. 1.3. research on gender differences in metacognitive monitoring research has shown that male and female learners may experience metacognitive monitoring in distinct ways (e.g., ackerman, nocera & bargh, 2010; denham et al., 2012; klassen & chiu, 2010). however, extant research on this topic has been inconclusive. ackerman and associates (2010), for instance, found that gender was predictive of lower self-rated driving ability such that females were underconfident regarding their confidence in performance judgments in their driving ability, and that this effect remained even after controlling for baseline driving ability. research on math achievement revealed that males not only exhibited higher achievement than females but that females were underconfident in their math achievement, and that this math performance miscalibration was more pronounced among females (klassen & chiu, 2010; özsoy, 2012; sheldrake, mujtaba, & reiss, 2014). on the other hand, research by nietfeld, shores, and hoffman (2014) did not uncover a significant gender effect on metacognitive monitoring bias or accuracy, as males and females rated confidence similarly and exhibited near similar accuracy. thus, to better disentangle this gender effect, we investigated whether measurement invariance exists between males and females on reading metacomprehension as an index of metacognitive monitoring. 1.4. situating the present research endeavor predicated on the literature review we surveyed, we adopted a two-study approach. in study 1, our main objective was to evaluate the factorial structure of the mi. because our model was based on substantive theoretical claims, we proposed an a priori hypothesized model and employed confirmatory factor analysis (cfa) techniques to assess the validity of the model. thus, our research question for this study was: do the observed data support our proposed latent variable model of metacomprehension in reading among a sample of chilean undergraduates? hypothesis 1: we predicted that the observed data would support our hypothesized conceptualization of latent factors involved in reading metacomprehennsion in adults, as shown in figure 1. we expected this model to fit the observed data exceptionally well. soto et al 38 | f l r figure 1. hypothesized factor structure of the metacomprehension inventory. in study 2 we were interested in exploring the invariance of the model between males and females. hence, our research question in this study was: does the factor structure of the metacomprehension inventory remain consistent among a sample of male and female undergraduate students? hypothesis 2: even though research shows that males and females at times vary in mean scores on metacomprehension scales, we predicted that the factor structure of the mi would remain invariant among males and females. 2. methodology 2.1 participants the participants for both studies were 923 undergraduate students of a chilean university located in the city of talca (autonomous university of chile), whose students are first generation university students. they were selected using a simple random sampling method from a population of 6,525 students. sampling error was 3%, with a 95% confidence level assuming maximum variance. the sample comprised 373 males, 545 females, and 5 who opted not to report gender. the average age was 22 years old (sd = 2.83), and participants were enrolled mainly in the disciplines of health, education, social sciences and business (see table 2). soto et al 39 | f l r table 2 participants by major department n % health sciences 300 32.5% education 245 26.5% social sciences 156 16.9% business administration 100 10.8% law 56 6.1% architecture & construction 35 3.8% engineering 31 3.4% total 923 100% 2.2 instruments we used a questionnaire with sociodemographic questions (i.e., gender, age, year of study and career of study) as well as the 42 likert scale items of the mi. the mi was developed following the fourth-order factor structure shown in figure 1. each endogenous dimension or sub-dimension was measured on a 5-point likert scale. all kac items were answered using a response format from strongly disagree to strongly agree. in contrast, for the coc dimension we used response format from never to always. 2.3. procedure the survey was administered face-to-face in december 2015, after obtaining approval of the directors of the different departments, institutional review board (irb) approval, and informed consent of all the participants. no student refused to participate in the study, each knowing that participation was voluntary. students took about 20 minutes to complete the questionnaire. 2.4. data analysis two studies were conducted: in the first one we used confirmatory factor analysis (cfa) procedures and t-student tests to validate the mi scale by construct and convergent validity methods, whereas the second study assessed the measurement invariance between male and female students by means of the confirmatory factor model validated in the first study using multigroup cfa. soto et al 40 | f l r 3. results 3.1. descriptive analysis of items table 3 shows descriptive statistics for each item, which allows an evaluation of some of its psychometric properties. since all items are written such that greater values indicate a higher level of metacomprehension, no reverse coding was necessary and all means can be interpreted consistently. table 3 means, sd and skewness of the metacomprehension inventory items (range of response 1 to 5) item mean sd skewness item mean sd skewness 1 4.0 0.8 -1.1 22 4.0 0.8 -0.3 2 4.1 0.7 -1.1 23 3.7 0.9 -0.3 3 3.7 0.9 -0.5 24 3.7 0.8 -0.3 4 3.6 0.9 -0.4 25 3.7 0.8 -0.4 5 3.6 0.8 -0.5 26 3.4 1.0 -0.3 6 3.7 0.9 -0.6 27 3.9 0.8 -0.5 7 4.1 0.7 -0.9 28 4.1 0.8 -0.4 8 4.2 0.8 -0.8 29 4.2 0.8 -0.9 9 2.9 1.1 0.1 30 3.9 0.9 -0.5 10 3.6 0.8 -0.3 31 3.9 0.9 -0.5 11 3.9 0.9 -0.5 32 4.0 0.8 -0.5 12 3.6 0.9 -0.3 33 4.0 0.8 -0.3 13 3.2 1.0 -0.2 34 3.9 0.8 -0.3 14 3.3 1.0 -0.2 35 4.1 0.8 -0.4 15 4.0 0.8 -0.4 36 4.0 0.9 -0.7 16 3.4 1.0 -0.3 37 4.0 0.9 -0.8 17 3.6 0.8 -0.2 38 3.8 0.9 -0.4 18 3.5 0.9 -0.4 39 3.9 0.9 -0.5 19 3.4 0.9 -0.4 40 3.6 1.1 -0.4 20 4.1 0.8 -0.6 41 3.6 1.0 -0.4 21 3.8 0.9 -0.5 42 3.9 0.8 -0.4 total 3.8 0.9 -0.5 soto et al 41 | f l r as seen in table 3, participants generally declared relatively high levels of metacomprehension (m = 3.8). the only exception was item 9 (i question myself about the topic before starting reading), which had answers slightly under the midpoint of the range of responses (med. = 3.0). most of items obtained mild negative asymmetry and low kurtosis (only items 1, 2 and 7 showed asymmetry less -1 or kurtosis greater than 1). finally, the standard deviations of the answers demonstrate that all the items were able to discriminate between the subjects. information presented in table 3 does not justify removing any item for low capacity of discrimination, ceiling or floor effects, or very high asymmetry (i.e., non-normal distributions). the foregoing is confirmed by the high significant correlations between the items and the corrected total score, which fluctuated from 0.33 to 0.66. 3.2. study 1: validation of the metacomprehension inventory 3.2.1. construct validation to assess the validity of the measurement model from which the items comprising the mi were generated, we evaluated several cfa models. if the theoretical model fits the observed data, this fact can be interpreted as evidence of construct validity of the instrument (messick, 1995). we detected that 1.2% of the test responses were missing, so, in order to verify that the missing data pattern was missing completely at random (mcar), little’s mcar χ2 statistics (little & rubin, 1989; schaeffer & graham, 2002) were estimated. a significant χ2 (i.e., p > .05) would suggest that the pattern of missing data is not mcar (i.e., missing not at random [mnar]), which poses a problem for interpretation of results because they may be biased due to systematic differences in non-responses. the result (χ2= 1708.4, df=1695, p=.40), suggests that the missing pattern in the data was mcar, and thus, working only with complete answers will not produce bias in results. the cfas were conducted on the polychoric correlation matrix of responses using unweighted least squares (uls) estimation, because the use of both allow for correcting the bias in factor loadings exhibited by the classic factor analysis based on pearson correlations when analyzing ordinal items (asún, rdz-navarro, & alvarado, 2016; forero, maydeu-olivares and gallardo-pujol, 2009). statistical analysis was conducted using mplus 6.11 (muthén and muthén, 2011), which allows a pairwise deletion of missing responses. the fourth-order factorial model shown in figure 1 had a good fit to the data (χ2 (n=923, df=430)=1018.04, p=.001; cfi=.919; tli=.914, rmsea=.057 (ci90%=.052, .061); srmr=.053), and factor loading sizes ranging from .417 to .793. however, the coc-edr and coc-ear factors shared 95% of their variance, while coc-rpu and coc-rdu shared 90%, which does not justify considering them independent dimensions. hence, we opted, based on this evidence, to combine these factors. goodness of fit of a third-order model was evaluated, merging coc-edr with coc-ear items, and coc-rpu with coc-rdu items. the fit of that model was good (χ2 (n=923, df=417)=988.77, p=.001; cfi=.916; tli=.912; rmsea=.057 (ci90%=.053, .060); srmr=.051), and very similar to the first. also, factor loadings were high, ranging from .406 to .776. consequently, we decided to retain this more parsimonious model in lieu of the previous more saturated model. additionally, with the aim of producing a shorter measure (and since a simpler model was retained), the items with the poorest fit, as evidenced by factor loadings to each factor, were eliminated. the elimination process was further supported by the collinearity of our initially proposed factor structure. the result of this process was a third-order model with 23 items, shown in the figure 2. the goodness of fit of this model was better than the previous (χ2 (n=923, df=389)=899.01, p=.001; cfi=.947; tli=.940; rmsea=.060 (ci90%=.057, .063); srmr=.049), and factor loadings were higher (ranging from .445 to .791). soto et al 42 | f l r figure 2. final model of the metacomprehension inventory. figure 2 shows the high factor loadings of the items with their respective factors and the high relations between higher-order factors and their subordinate factors. these results provide evidence that the shorter instrument allows for global and specific scores for each latent dimension within the hierarchy of factors. thus, we retained this best fitting factor structure as our final model. 3.2.2. convergent validation to provide some evidence of the validity of the above model, we evaluated it for significant differences by gender, as there is some evidence that males and females have different levels of metacognitive monitoring (ackerman, et al., 2010; klaasen & chiu, 2010; gutierrez & price, 2017; sharma & bewes, 2011), especially in countries or cultures with social norms and expectations different for men and women (e.g., bembenutty, 2007; bussey & bandura, 1999). given that our sample belongs to a latin american country (chile), where these different social norms still exist, we sought to examine the levels of metacomprehension between both genders. soto et al 43 | f l r our analyses revealed that, compared to males, females reported greater levels of metacomprehension in all factors of the inventory: mi (t(916) = 3.99; p<.001), kac (t(916) = 3.61; p<.001), coc (t(916) = 4.13; p<0.001), coc-p (t(916) = 3.65; p<0.001), coc-e (t(916) = 4.12; p<0.001) and coc-r (t(916) = 4.10; p<0.001). we believe these results are evidence of the validity of the scoring of the instrument. 3.2.3. reliability of the measurement instrument due to the ordinal nature of the items used, the ordinal alpha coefficient (mcdonald, 1985) was employed as an indicator of the reliability (i.e., internal consistency) of the hierarchical factor structure of mi. table 4 shows the results. table 4 reliability (ordinal alpha) of the total instrument, its dimensions and subdimensions third-order second-order first-order reliability metacomprehension inventory (mi) .925 knowledge about cognition (kac) .760 control of cognition (coc) .919 control of cognition: planning (coc-p) .761 control of cognition: evaluation (coc-e) .878 control of cognition: regulation (coc-r) .821 n = 923 it is evident, from data in table 4, that the mi shows high levels of reliability in both the global score and scores at all levels of the factor hierarchy. 3.3. study 2: invariance of the confirmatory model in the second study, we sought to examine the measurement invariance of the validated instrument between male and female students to ascertain whether the factor structure remained consistent between the two groups, despite of their different levels of metacomprehension revealed in section 3.2.2. 3.3.1. data analysis due to the complexity of testing model invariance using statistics for ordinal data, in this study we assumed that responses were obtained on an interval scale and used maximum likelihood procedures to evaluate the models. all data were screened for univariate and multivariate outliers according to the procedures outlined by tabachnick and fidell (2013) using the international business machine (ibm) statistical package for the soto et al 44 | f l r social sciences (spss) statistics 22. no extreme outliers that would otherwise undermine the trustworthiness of the data were detected. the missing values analysis demonstrated that 20 cases (5.3%) in the male group and 23 cases (4.2%) in the female group had missing data. in order to verify that the missing data pattern was missing completely at random (mcar), little's mcar χ2 statistics (little & rubin, 1989; schaeffer & graham, 2002) were requested. the result of this test for the present data was non-significant for both groups (p=.53, male group and p=.18 for the female group), suggesting that the missing pattern in the data was mcar. thus, data analysis proceeded with 875 complete cases (353 for the male group and 522 for the female group). furthermore, data were tested for univariate and multivariate assumptions, including multivariate normality, multicollinearity, and reproducibility of the correlation matrix via residual analysis using eqs 6.1, in order to proceed with the multi-group confirmatory factor analysis (cfa). all assumptions were met, and thus, data analysis proceeded without making any adjustments to the data. multi-group cfa was performed to evaluate the invariance of path coefficients among male and female groups using eqs 6.1. first, a fully constrained, fully-saturated baseline model was established for both groups to examine the feasibility of the hypothesized cfa model presented in figure 2 by specifying the direct paths and by imposing equality constraints on all path coefficients and covariances. subsequently, exploratory model trimming (wald test for dropping parameters) and model building (lagrange multiplier [lm] test for adding parameters) procedures were interpreted in an effort to improve overall model fit of the baseline model. next, equality constraints were individually removed for each parameter (i.e., freely estimated) that reached statistical significance at the p<.05 level using the multivariate lm χ2 univariate increment test for releasing equality constraints. this procedure was repeated until no further parameters’ lm χ2 univariate increment reached statistical significance. this model was then deemed the final model. releasing equality constraints for any given parameter indicates that the parameter in question differs statistically significantly across the male and female groups. finally, the δχ2 (chi-square difference) test was conducted to compare the null (i.e., fully-constrained, fully-saturated) model and the final model (i.e., released equality constraints). 3.3.2. results the baseline model for both groups with equality constraints imposed on all path coefficients and covariances (figure 2) was adequately fitting to the observed data, χ2 (447, n=875) = 777.23, p<.05, tli=.92, cfi=.94, ifi=.94. srmr=.05, rmsea=.04 (ci90%=.03, .05). none of the model building or model trimming statistics was warranted based on theoretical considerations, and hence, this was deemed the final model for both groups. the final model, with one statistically significant equality constraint removed, fit the observed data reasonably well, χ2 (446, n=875) = 760.67, p<.05, tli=.93, cfi=.94, ifi=.95. srmr=.05, rmsea=.04 (ci90%=.03, .04). the correlations between all of the metacomprehension dimensions and subdimensions were statistically significant, but did not differ significantly among the groups. as is evident, δχ2 test results between the fully-constrained baseline model and the final model with one freed equality constraint was not statistically significant, p=.89. therefore, one can conclude that the factor structure of the mi remained consistent between males and females. the only statistically significant difference between the groups was in the path coefficient between the control of cognition-regulation factor (coc-r) and its item, i ask myself if what i am reading is related to what i already know about the content of the text, which was somewhat stronger among females (λ=.79) than males (λ =.70). nevertheless, all other equality constraints, and hence, path coefficients among the groups, remained invariant. this supports our hypothesis regarding the invariance of the mi across males and females. soto et al 45 | f l r 4. general discussion the purpose of study 1 was to evaluate the validity of an innovative, more comprehensive framework for measuring metacomprehension reading strategies among a robust sample of chilean undergraduate students (i.e., the metacomprehension inventory). we hypothesized that this alternative factor structure would provide a more complete and accurate representation of the latent multidimensionality of metacomprehension in reading when compared to previous conceptualizations. our final third-order model (figure 2) with five subordinate factors subsumed by a global metacomprehension factor demonstrated good fit to the observed data, with reasonable fit indices, low residual statistics, and factor loadings within acceptable range and in the expected theoretical direction. given previous measures of metacomprehension (e.g., marsi [mokhtari & reichard, 2002]; mq [swanson, 1990]; mai [schraw & dennison, 1994]; escola [jiménez et al., 2009]), the statistical evidence we provide demonstrates the more comprehensive and complete nature of our model when compared to previous models because our model assesses more fully the multidimensionality of metacomprehension in reading. researchers who study the latent dimensionality of measures have consistently argued that a more complete evaluation of the psychometric multidimensionality of measures that purport to assess a latent construct or set of constructs is essential for drawing more valid inferences and conclusions regarding psychological phenomena, albeit they do not always agree on how to statistically accomplish this (e.g., chemolli & gagné, 2014; guay, ratelle, roy & litalien, 2010; morin, arens, & marsh, 2016; morin, arens, tran, & caci, 2016). nevertheless, these researchers agree that not capturing relevant psychometric multidimensional variance by specifying simpler latent factor structures may bias goodness of fit indices, residual-based statistics, and overestimate factor correlations, possibly leading to a higher likelihood of unnecessarily inflating multicollinearity diagnostics. this leads to the potential of researchers only partially evaluating theoretical frameworks by omitting factors that are (artificially) highly correlated with other factors within the framework, and thus, providing incomplete evidence of the viability of such frameworks. our proposed model in study 1 mitigates these situations by proposing a more complex rather than simple model, as in previous measures of metacomprehension we surveyed. in sum, we believe that our mi will provide a more comprehensive psychometric multidimensionality of metacomprehension in reading than previous attempts. in study 2 we examined the measurement invariance of our proposed mi framework, as prior research has shown that metacomprehension is at times moderated by gender (e.g., ackerman et al., 2010; klaasen & chiu, 2010; gutierrez & price, 2017; sharma & bewes, 2011). even though previous research has demonstrated that self-reported metacomprehension mean scores sometimes vary among males and females, we expected the factor structure of the mi to remain invariant between the groups. even though data of study 1 found that there were statistically significant differences when comparing the mean scale scores between males and females, with females consistently reporting higher mean scale scores, supporting the line of research on gender differences in self-reported metacomprehension (e.g., ackerman et al., 2010; klaasen & chiu, 2010; gutierrez & price, 2017; sharma & bewes, 2011), study 2 findings support the view that the mi factor structure is consistent between males and females. only one parameter in the factor structure of the mi was significantly different between males and females, with females showing a higher standardized path coefficient in the factor loading of the item, i ask myself if what i am reading is related to what i already know about the content of the text, within the coc-r factor, in line with study 1 results. however, the remaining parameter estimates remained invariant between males and females. the fact that the factor structure of the mi remained mostly consistent between this sample of chilean adolescent males and females is encouraging because it does not bias results for either males or females, in spite of research that shows the effects of cultural and social expectations on self-reports of metacomprehension between the two groups. soto et al 46 | f l r 4.1. implications for theory, research, and educational practice theoretically this new inventory considers different metacomprehension components as with extant research, but it also adds a new perspective to approximating the mechanisms underlying metacomprehension. it deeply explores a special focus on evaluation (monitoring) and regulation. different researchers have considered the importance of evaluation in reading comprehension (redford, thiede, wiley, & griffin, 2012), concluding that a calibrated evaluation involves awareness about when/where in the text the reader is getting a better or weaker level of understanding. however, the regulation process has enjoyed less attention in the literature, presumably because the regulation in reading comprehension represents a long tradition associated with the contradiction paradigm, and the nature of those phenomena are more complicated to assess with precision, or because some researchers prefer working with normal texts (azevedo, 2009). however, that situation does not necessarily mean we must stop our efforts to generate new, improved approximations about the particular mechanisms underlying those phenomena, and particularly between the relation between evaluation and regulation (keener & hacker, 2012). this inventory has the potential to offer a new, more specific theoretical perspective because it not only distinguishes between evaluation during and after reading, but also conceptualizes self-regulation more comprehensively than previous attempts. an important contribution is that it considers regulation after difficulties in understanding, where some process is activated to evaluate and implement adjustment to repair the incoherence in the mental representation and other mechanisms within regulation to deepen comprehension. this latter process does not stem from a problem with confusion or a discrepancy, but instead, according to a typical reading process individuals sometimes employ when regulating their reading comprehension to acquire a better mental representation of the text. in fact, under this perspective to use reading comprehension strategies during reading to improve current comprehension could be considered a regulation behavior, regardless of whether the use of strategies is conscious, subconscious, controlled or automatized, while it is generated from the evaluation state to trigger the repair. another contribution of this work is that it better highlights the dynamic relation between cognition and metacognition, and how the relation between those two levels is more flexible than has been previously assumed, albeit a deeper discussion of this point is beyond the scope of this paper. regarding different ways of regulation, the mi considers two different mechanisms about correcting errors in reading, capturing a more complete understanding about the metacomprehension process when using the inventory. 4.2. avenues for future research despite the fact that our studies used a robust sample size, they may not necessarily generalize to other populations or samples of adult learners from other cultures. future research should verify our findings using children and adults from other cultures to ascertain whether our mi factor structure and invariance generalizes to other samples, populations, developmental stages, and cultures. moreover, it would be worthwhile for researchers to explore the utility of the mi in actual classroom settings. finally, future research should examine the relation of the mi to previous measures of reading comprehension and investigate whether the mi’s more comprehensive nature predicts additional variance not attributable to previous measures. 4.3. methodological reflections and limitations no research endeavor involving human beings is ever without limitations. even though the mi represents a more robust psychometric multidimensionality attempt with respect to metacomprehension, with strong support for our hypothesized higher-order factor structure, it still relies on self-report responses from participants. as has been consistently demonstrated, individuals may not be the best raters of their own attitudes, beliefs, and perceptions due to such phenomena as the social desirability bias. moreover, our research design was cross-sectional and correlational in nature, thereby limiting the inferences and soto et al 47 | f l r conclusions we can draw from these data. a longitudinal design may have permitted us to investigate how enduring these perceptions are among this sample of chilean learners. in addition, we acknowledge the common method variance dilemma that may bias correlations and effect sizes due to our single method, single rater approach. in spite of these limitations, our combined studies contribute to a deeper understanding of the complex dynamics involved in metacomprehension of reading. 4.4. conclusion study 1 provided support for our more comprehensive framework for conceptualizing metacomprehension than prior research. our mi measure is also compact, and thus, it can be readily employed with students in authentic, ecologically valid learning environments such as classrooms. information gleaned from the mi could assist classroom teachers to uncover specific deficits in learners’ reading comprehension strategy use. this would then enable teachers to develop individualized reading interventions that could benefit poor reading comprehenders in particular. study 2 showed that, with the exception of one path coefficient that was stronger among females than males, all other parameters in the model were invariant between males and females. this is important information to know because the mi could be applied equally as well to male and female learners without the need to make adjustments. finally, this inventory demonstrated sound psychometric characteristics and conceptual foundations that lead us to conclude that it is an excellent resource for researchers and teachers. acknowledgments we thank the autonomous university of chile at talca, for providing access to students for this research endeavor. a special thanks to the centro de estudios y gestión social (ceges) for making the fieldwork involved in this research possible. finally, we acknowledge the contribution of the it16i10044 project: technology for the improvement of reading comprehension in students of the chilean school system (tecnología para el mejoramiento de la comprensión lectora en estudiantes del sistema escolar chileno; fondef, conicyt). appendix: metacomprehension inventory item spanish (original) version english translation 1 conozco los puntos fuertes y débiles de mis habilidades lectoras i know the strong and weak points of my reading skills 2 sé qué tipo de textos podrían ser más difíciles para mí i know which topics are more complex to read for me 3 sé cómo enfrentar un texto para que se me haga más fácil de entender i know how to deal with a text to make it easier to understand for me 4 conozco las estrategias necesarias para leer mejor i know the strategies needed to read better 5 sé cómo sobreponerme cuando tengo una dificultad al comprender un texto i know how to overcome a difficulty when i have problems understanding a text soto et al 48 | f l r 6 sé cuáles podrían ser las características de un buen lector i know which might be the characteristics of a good reader 7 sé los temas que para mí podrían ser más complejos de leer i know which subject-topics might be more difficult to read for me 8 cuando algo se me hace difícil a veces vuelvo atrás en la lectura when something is difficult to understand, sometimes i turn back into my reading 9 me hago preguntas sobre el tema antes de empezar a leer i question myself about the topic before starting reading 10 después de leer un texto puedo saber con precisión el nivel de comprensión que he alcanzado when finishing my readings i know precisely which level of understanding i achieved 11 durante la lectura me pregunto si voy entendiendo bien o no while i am reading i wonder whether i’m understanding the text correctly or not 12 pienso en lo que realmente necesito comprender antes de empezar una tarea de lectura i think about what i really need to understand before starting a reading task 13 me propongo objetivos específicos antes de empezar una tarea de lectura i define specific objectives for myself before starting a reading task 14 cuando estoy leyendo me pregunto constantemente si estoy alcanzando mis metas de lectura when i'm reading, constantly i ask myself whether i'm reaching my reading goals or not 15 cuando estoy leyendo, de vez en cuando pienso si estoy entendiendo lo que estoy leyendo when i'm reading, occasionally i think whether i'm understanding what i'm reading 16 cuando me dispongo a leer organizo el tiempo y las actividades de lectura para poder terminar la tarea a tiempo when i prepare myself to read i organize my time and reading activities to finish the task on time 17 a medida que voy leyendo puedo ser preciso/a para determinar cuánto voy comprendiendo while i’m reading i can determine how much i’m understanding 18 cuando respondo un test de comprensión lectora sé cómo me ha ido when i respond to a reading comprehension test i know how good or bad was my performance 19 pienso en las distintas maneras de abordar la lectura de un texto y escojo la mejor i think about the different ways to accomplish a reading task and choose the best approach 20 cuando no logro entender una parte de un texto trato de ir más despacio para entender when i don’t understand a part of a text i try to slow down my reading to understand it soto et al 49 | f l r 21 trato de vincular diferentes partes del texto para que tenga más sentido i try to link different parts of the text to make more sense 22 durante la lectura sé si estoy comprendiendo bien, regular o mal during my reading i know whether i am understanding well, not so well or bad 23 cuando termino una lectura puedo saber si logré las metas que me había propuesto when i finish my reading i can tell myself whether i achieved my goals or not 24 cuando estoy confundido/a me pregunto si lo que estaba suponiendo acerca del texto era correcto o no when i feel confused i wonder whether my assumptions about text were correct or not 25 cuando leo trato de estar consciente del nivel de comprensión que voy alcanzando when i read i try to be aware of the level of understanding i’m reaching 26 organizo el tiempo de lectura para lograr mejor mis objetivos i organize my reading time for better achievement of my goals 27 cuando estoy leyendo puedo determinar si alguna parte del texto está siendo más fácil o más difícil para mí when i'm reading i can determine whether part of the text is being easier or harder to understand 28 después de leer un texto puedo determinar qué tan complejo ha sido para mí after reading a text i can determine how complex it has been for me 29 cuando me resulta extraña la información del texto me detengo y la leo más de una vez when i find some text information strange i stop and read the paragraph more than once 30 puedo pensar en ejemplos para poder entender mejor la información del texto i can think about examples to better understand text information 31 leo cuidadosamente los enunciados antes de empezar con la lectura i read carefully the instructions before starting my reading 32 cuando estoy leyendo puedo darme cuenta si el texto es más o menos complejo para mí when i'm reading i can tell whether the text is more or less complex for me 33 cuando termino de leer puedo evaluar si he comprendido bien o no when i finished reading i can evaluate whether i understood the text well or not 34 luego de leer un texto puedo evaluar si interpreté correctamente el texto after reading a text i can assess whether i interpreted the text correctly 35 cuando termino de leer un texto puedo saber qué parte fue más confusa para mí when i finish reading a text i can know which part was more confusing to me 36 cuando el texto se me hace difícil de comprender aumento mi atención y esfuerzo when a text is difficult to understand i increase my attention and effort 37 cuando una palabra es extraña para mi trato de entender por el contexto de la lectura when a word is strange to me i try to understand it by the context soto et al 50 | f l r 38 si no estoy comprendiendo una parte, a veces sigo con la lectura en busca de clarificación if i'm not understanding a part of a text, sometimes i keep reading seeking for clarification 39 cuando leo trato de ir explicándome el texto con mis propias palabras when i read i try to explain the text to myself using my own words 40 mientras leo trato de imaginar lo que vendrá a continuación en el texto while i read i try to imagine what come next in the text 41 busco hacerme preguntas para darle más sentido al contenido del texto i seek inquire myself about what i read to give more meaning to the text content 42 me pregunto si lo que estoy leyendo está relacionado con lo que ya sé del contenido del texto i ask myself whether what i am reading is related to what i know about the text content note: in bold, item retain in final metacognitive inventory. keypoints the metacomprehension inventory (mi) exhibited sound psychometric properties. the mi is an effective tool for measuring reading comprehension among adults. the mi parameter estimates were invariant across males and females. references ackerman, j., nocera, c., & bargh, j. (2010). incidental haptic sensations influence social judgments and decisions. science, 328, 1712-1715. doi: 10.1126/science.1189993 asún, r. a., rdz-navarro, k., & alvarado, j. m. (2016).developing multidimensional likert scales using item factor analysis: the case of four-point items. sociological methods and research, 45(1), 109133. doi: 10.1177/0049124114566716 azevedo, r., & aleven, v. (2013). international handbook of metacognition and learning technologies. new york: springer. azevedo, r. (2009). theoretical, methodological and analytical challenges in the research on metacognition and self-regulation: a commentary. metacognition & learning, 4, 87-95. doi: 10.1007/s11409-0099035-7 baker, l. (2002). metacognition in comprehension instruction. in block, c. c. & pressley, m. (eds.), comprehension instruction: research-based best practices (pp. 77-95). new york: the guil-ford press. beck, i.l. & mckeown, m.g. (1998). comprehension: the sine qua non of reading. in s. patton & m. holmes (eds.), the keys to literacy (pp. 40-52). washington, dc: council for basic education. beck, i.l. & mckeown, m.g. (2006). improving comprehension with questioning the author: a fresh and expanded view of a powerful approach. new york, n.y: scholastic. beck, i.l. & mckeown, m.g. (2009). the role of metacognition in understanding and supporting reading comprehension. in d. j. hacker, j. dunlosky, & a. c. graesser (eds.), handbook of metacognition soto et al 51 | f l r in education. the educational psychology series (pp. 7-25). mahwah, nj: lawrence erlbaum associates. bembenutty, h. (2007). self-regulation of learning and academic delay of gratification: gender and ethnic differences among college students. journal of advanced academics, 18(4), 586-616. block, c. c., & pressley, m. (2007) best practices in teaching comprehension. in l.b. gambrell, l. m. morrow, & m. pressley (eds.), best practices in literacy instruction (pp. 220–242). new york: guilford. boekaerts, m. (1999) self-regulated learning: where we are today. international journal of educational research 31(6), 445-457. doi: 10.1016/s0883-0355(99)00014-2 bol, l., & hacker, d. j. (2001). a comparison of the effects of practice tests and traditional review on performance and calibration. the journal of experimental education, 69(2), 133-151. doi: 10.1080/00220970109600653 bol, l., hacker, d. j., o’shea, p., & allen, d. (2005). the influence of overt practice, achievement level, and explanatory style on calibration accuracy and performance. the journal of experimental education, 73(4), 269-290. doi: 10.3200/jexe.73.4.269-290 brown, a. l., armbruster, b. b. & baker, l. (1986). the role of metacognition in reading and studying. in orasanu, j. (ed.), reading comprehension: from research to practice. hillsdale, nj: lawrence erlbaum. bussey, k., & bandura, a. (1999). social cognitive theory of gender development and differentiation. psychological review, 106(4), 676-713. chemolli, e. & gagné, m. (2014). evidence against the continuum structure underlying motivation measures derived from self-determination theory. psychological assessment, 26(2), 575-585. doi: 10.1037/a0036212 denham s. a., warren-khot h. k., bassett h. h., wyatt t., perna a. (2012). factor structure of selfregulation in preschoolers: testing models of a field-based assessment for predicting early school readiness. journal of experimental child psychology, 111(3), 386-404. doi: 10.1016/j.jecp.2011.10.002 dunlosky, j., rawson, k.a., & mcdonald (2002). influence of practice tests on the accuracy of predicting memory performance for paired associates, sentences, and text material. in t.j. perfect & b.l. schwartz (eds.), applied metacognition (pp. 68-92). new york, ny: cambridge university press. dunlosky, j., griffin, t., thiede, k. & wiley, j. (2005). understanding the delayed keyword effect on metacomprehension accuracy. journal of experimental psychology: learning, memory & cognition, 31, 1267-1280. doi: 10.1037/0278-7393.31.6.1267 dunlosky, j., & lipko, a.r. (2007). metacomprehension: a brief history and how to improve its accuracy. current directions in psychological science, 16(4), 228-232. doi: 10.1111/j.1467-8721.2007.00509.x flavell, j. (1979). metacognition and cognitive monitoring: a new area of cognitive–developmental inquiry. american psychologist, 34(10), 906-911. flavell, j. h., & wellman, h. m. (1977). metamemory. in r. v. kail y j. w. hagen (eds.), perspectives on the development of memory and cognition. hillsdale, n.j.: erlbaum. forero, maydeu-olivares and gallardo-pujol, (2009). factor analysis with ordinal indicators: a monte carlo study dwls and uls estimation. structural equation modeling: a multidisciplinary journal, 16, 625-641. doi: 10.1080/10705510903203573 guay, f., ratelle, c., roy, a., & litalien, d. (2010). academic self-concept, autonomous academic motivation, and academic achievement: mediating and additive effects. learning and individual differences, 20 (6), 644-653. doi: 10.1016/j.lindif.2010.08.001 gutierrez, a. p., & price, a. f. (2017). calibration between undergraduate students’ prediction of and actual performance: the role of gender and performance attributions. the journal of experimental education,85, 486-500. doi: 10.1080/00220973.2016.1180278 soto et al 52 | f l r gutierrez, a. p., & schraw, g. (2015). effects of strategy training and incentives on students’ performance, confidence, and calibration. the journal of experimental education, 83(3) 386-404. doi: 10.1080/2331186x.2017.1314652 hacker, d., plumb, c., butterfield, e., quathamer, d., & heineken, e. (1994). text revision: detection and corrections of errors. journal of educational psychology, 86, 65-78. hacker, d. (1998). definitions and empirical foundations. in d. hacker, j. dunlosky, & a. graesser (eds.), metacognition in educational theory and practice. mahwah, nj: erlbaum. hacker, d., bol, l., & bahbahani, k. (2008). explaining calibration in classroom context. the effects of incentives, reflection, and attributional style. metacognition and learning, 3, 101-121. doi: 10.1080/2331186x.2017.1314652 hacker d. j., dunlosky j., graesser a. c. (2009), handbook of metacognition in education. new york: routledge. jacobs, j. e., and paris, s. g. (1987). children’s metacognition about reading: issues in definition, measurement, and instruction. educational psychologist. 22 (3&4) 255–278. jiménez, v., puente, a., alvarado, j., & arrebillaga, l. (2009) medición de estrategias metacognitivas mediante la escala de conciencia lectora: escola. electronic journal of research in educational psychology,7(2), 779-804. keener, m. c., & hacker d. j. (2012). comprehension monitoring. in n. m. norbert (ed.), encyclopedia of the sciences of learning. new york, ny: springer. klaasen, r., & chiu, m. (2010). effects on teachers' self-efficacy and job satisfaction: teacher gender, years of experience, and job stress. journal of educational psychology, 102(3), 741-756. doi: 10.1037/a0019237 little, r.j. & rubin, d.b. (1989). the analysis of social science data with missing values. sociological methods and research, 18 (2&3), 292-326. mcdonald, r.p. (1985). factor analysis and related methods. hillsdale nj: erlbaum. mokhtari, k., & reichard, c. (2002). assessing students’ metacognitive awareness of reading strategies. journal of educational psychology, 94(2), 249-259. doi: 10.1037/0022-0663.94.2.249 morin, a., arens, k., & marsh, h. (2016). a bifactor exploratory structural equation modeling framework for the identification of distinct sources of construct-relevant psychometric multidimensionality. structural equation modeling: a multidisciplinary journal 23(1), 116-139. doi: 10.1080/10705511.2014.961800 morin, a., arens, k., tran, a. & caci, h. (2016) exploring sources of construct-relevant multidimensionality in psychiatric measurement: a tutorial and illustration using the composite scale of morningness. international journal of methods in psychiatric research 25(4), 277-288. doi: 10.1002/mpr.1485 muthen, l. k., & muthen, b. o. (1998–2011). mplus user’s guide (6th ed.). los angeles, ca: muthen & muthen. nietfeld, j., shores, l., & hoffman, k. (2014). self-regulation and gender within a game-based learning environment, journal of educational psychology, 106(4), 961-973. doi:10.1037/a0037116 özsoy, g. (2012). investigation of fifth grade students’ mathematical calibration skills. educational sciences: theory and practice, 12(2), 1190–1194. paris, s. g. & jacobs, j. e. (1984). the benefits of informed strategies for learning: a program to improve children's reading awareness and comprehension. journal of educational psychology, 76 (6), 12391252. paris, s.g. & paris, a.h. (2001). classroom applications of research on self-regulated learning. educational psychologist, 36(2), 89-101. doi: 10.1207/s15326985ep3602_4 pintrich p. & de groot, e. (1990). motivational and self-regulated learning components of classroom academic performance. journal of educational psychology 82 (1), 33-40. pressley, m., & afflerbach, p. (1995). verbal protocols of reading: the nature of constructively responsive reading. hillsdale nj: lawrence erlbaum associates inc. soto et al 53 | f l r pressley, m., & block, c. c. (2002). summing up: what comprehension instruction could be. in c. c. block & m. pressley (eds.), comprehension instruction: research-based best practices. new york: guilford press. puente, a., jiménez, v. & alvarado, j.m. (2009) escala de conciencia lectora (escola). evaluación e intervención psicoeducativa de procesos y variables metacognitivas durante la lectura. madrid: eos. redford, j., thiede, k., wiley, j. & griffin, t. (2012).concept mapping improves metacomprehension accuracy among 7th graders. learning and instruction, 22 (4), 262-270. doi: doi:10.1016/j.learninstruc.2011.10.007 sandora, c., beck, i., & mckeown, m. (1999). a comparison of two discussion strategies on students' comprehension and interpretation of complex literature. journal of reading psychology, 20 (3), 177212. schaeffer, j. & graham, j. (2002). missing data: our view of the state of the art. psychological methods, 7 (2), 147-177. schmidt, r. (1990). the role of consciousness in second language learning. applied linguistics, 11 (2), 129-158. schraw, g., & dennison, r. s. (1994). assessing metacognitive awareness. contemporary educational psychology, 19 (4), 460-475. schraw, g., & moshman, d. (1995). metacognitive theories. educational psychology review, 7 (4), 351371. sharma, m. d & bewes, j. (2011). self-monitoring: confidence, academic achievement and gender differences in physics. journal of learning design, 4, (3), 1-13. sheldrake, r., mujtaba, t., & reiss, m. (2014). calibration of self-evaluations of mathematical ability for students in england aged 13 and 15, and their intentions to study non-compulsory mathematics after age 16. international journal of educational research, 64, 49-61. doi: 10.1016/j.ijer.2013.10.008 swanson, h. l. (1990). influence of metacognitive knowledge and aptitude on problem solving. journal of educational psychology, 82 (2), 306-314. tabachnick, b. g., & fidell, l. s. (2013). using multivariate statistics (6th ed.). boston: pearson. thiede, k., griffin, t., wiley, j. & redford, j. (2009). metacognitive monitoring during and after reading. in j. dunlosky, a. graesser & j. hacker (eds.), handbook of metacognition in education. new york: routledge. microsoft word warner et al_publication.docx ! ! ! ! ! frontline!learning!research!vol.5!no.!2!(2017)!1!;!23! issn!2295;3159!! ! corresponding author: greta j. warner, university of potsdam, work and organizational psychology, karlliebknecht-str. 24-25, 14476 potsdam, germany. e-mail: warner@uni-potsdam.de, phone: +49-331-977-2336, fax: +49-331-977-2091. doi: http://dx.doi.org/10.14786/flr.v5i2.272 ! relations among personal initiative and the development of reading strategy knowledge and reading comprehension greta j. warner, doris fay, & nadine spörer university of potsdam, germany article received 28 september / revised 21 december / accepted 27 december / available online 28 february abstract reading comprehension is a self-regulated activity that depends on the proactive effort of the reader. therefore, the authors studied the effects of personal initiative (pi) on the development of reading comprehension, mediated by reading strategy knowledge. structural equation modelling was applied to a longitudinal study with two data waves separated by two years. at time 1, the participants (n = 1,102) were either in third or fourth grade. at time 2, third graders had moved to grade five, and fourth graders had moved to grade six (n = 1,009). at both grade levels, pi explained unique variance in reading strategy knowledge and reading comprehension at time 2. moreover, from fourth to sixth grade the effect of pi on reading comprehension was mediated by reading strategy knowledge. no mediation was observed from third to fifth grade. these findings emphasize the relevance of pi in the development of reading strategy knowledge and reading comprehension. they further reveal that the hypothesized mediation process does not unfold until sixth grade. keywords: personal initiative; self-regulation; reading comprehension; strategy knowledge; proactivity warner&et &al & & & | f l r ! ! 2! 1. introduction reading comprehension is understood as a multifaceted and constructive process in which basic information processes and higher cognitive skills must be coordinated and self-regulated (e.g., cain, oakhill, & bryant, 2004). readers must identify word meanings, actively construct text meanings, and embed text information into previous knowledge (kintsch, 1998). the orchestration of these microand macroprocesses requires that students have a clear goal in mind, plan ahead, and stay persistent when comprehension is challenged (e.g., duke & pearson, 2002). therefore, they need to be highly proactive, take responsibility, and demonstrate a good deal of personal initiative (e.g., dermitzaki, andreou, & paraskeva, 2008). personal initiative (pi) defines the behavioural tendency to show self-starting, proactive, and persistent behaviour (see frese & fay, 2001; wollny, fay, & urbach, 2016). in present theories of selfregulated learning this proactive behaviour is understood as an essential antecedent of successful learning processes (e.g., zimmerman, 2002). thus, it seems likely that pi has an impact on students’ academic performance (wollny et al., 2016). yet, research in this area is only just beginning to develop and until now, no study has focused on the impact of pi on one of the most important academic abilities and predictors of future academic performance, i.e., reading comprehension. to date, most empirical studies have focused on the relevance of basic cognitive self-regulatory skills in the development of reading comprehension. for example, there is substantial evidence that executive functions predict reading comprehension in the long-term (e.g., birgisdóttir, gestsdóttir, & thorsdóttir, 2015). these and similar self-regulatory skills are relevant because they reflect a child’s ability to respond to classroom requirements (e.g., follow instructions; mcclelland & cameron, 2011). in addition to these elementary self-regulatory abilities, we propose that pi adds to the literature by capturing the self-starting, proactive, and persistent nature of reading comprehension as emphasized by both researchers and educators (e.g., horner & shwery, 2002; pressley & afflerbach, 1995). pi should, for example, demonstrate a particularly high relation to active reading behaviours, like goal-setting and active problem-solving (e.g., pressley & afflerbach, 1995). however, there seems to be a lack of research about how these proactive behaviours impact the early development of reading comprehension during later school years at elementary school. as of third or fourth grade, students start switching from rather basic word reading to a more global text understanding (e.g., duke & carlisle, 2011). their ability to reflect about their thinking starts developing (i.e., metacognition). this builds the foundation for strategy use, planning, and problem-solving behaviours as involved in pi (e.g., janke & hasselhorn, 2008; schneider, 2010). thus, both reading comprehension and pi likely show a rapid development during that time making it a critical timeframe for studying their relations. the first aim of this study is therefore to explore whether pi represents a predictor of reading comprehension from third to fifth and fourth to sixth grade. in addition, research so far has paid little attention to the mediational mechanisms that explain a positive relation between domain-unspecific self-regulatory constructs and a specific academic ability, such as reading comprehension (e.g., best, miller, & jones, 2009). yet, it is emphasized that general selfregulatory tendencies do not directly, but rather indirectly affect academic performance via more specific learning-oriented actions (e.g., boekaerts, 1999; borkowski, chan, & muthukrishna, 2000). therefore, it is necessary to develop a better understanding of how more general and domain-specific self-regulatory processes interact to influence academic performance. in this study, we assume the effects of pi on reading comprehension to be mediated by self-regulatory skills gained during the reading process. within this context, adequate reading strategy knowledge constitutes one crucial self-regulatory skill, which has consistently demonstrated positive associations with reading comprehension (van kraayenoord, 2010). hence, a second aim of this study is to investigate a mediational process in which pi affects reading comprehension via the development of reading strategy knowledge. warner&et &al & & & | f l r ! ! 3! demonstrating such a mediational process and thereby a positive impact of pi on reading comprehension would have important theoretical and educational implications. first, it would emphasize the relevance of proactivity in the development of reading comprehension. second, it would offer new insights into the relations between domain-unspecific self-regulatory tendencies and reading-specific selfregulatory actions. and third, it would suggest new directions in educational practices and encourage more opportunities of autonomy, choice, and control for students. 1.1 personal initiative and reading comprehension initially developed within the work context (frese & fay, 2001; frese, fay, hilburger, leng, & tag, 1997), the construct of pi represents the behavioural tendency to display self-starting, proactive, and persistent change-oriented behaviour (frese & fay, 2001). self-starting means setting goals that are not explicitly given and go beyond what is expected. proactive means anticipating and acting towards future problems instead of waiting until problems arise. and, persistent means pursuing goals despite potential setbacks or other barriers. in other words, pi represents the opposite of “doing what one is told to do, giving up in the face of difficulties, and reacting to environmental demands” (fay & frese, 2001, p. 97). thus, pi describes behavioural components that have been put forward as core qualities of successful learners, such as the guidance by personal goals, systematic planning, and high perseverance (zimmerman, 2002). with this definition, pi is distinct from the concept of engagement because engagement includes adaptive behaviours that are not proactive in nature (e.g., doing homework, attending school; jimerson, campos, & greif, 2003). likewise, pi is distinct from motivational variables (e.g., self-efficacy, goal orientation), which relate to a broad range of learning behaviours that are not necessarily proactive (e.g., engagement; caraway, tucker, reinke, & hall, 2003). finally, pi is distinct from self-regulated learning because it defines a general tendency rather than a process of learning-related activities (wollny et al., 2016; zimmerman, 2002). based on previous studies, we assume that successful reading comprehension requires all facets of pi. pressley and afflerbach (1995) studied think-aloud protocols of reading and found that good comprehenders were self-starting because they approached texts with a specific purpose in mind, set goals before and during reading, and constantly compared them with their reading progress. good comprehenders were also proactive because they scanned the text, generated hypotheses, anticipated potential problems, and planned how to read the text. finally, they were also persistent, tried new reading strategies, flexibly applied strategies, and discarded inefficient ones. quantitative studies underpin these results by showing that reading comprehension is selfdirected and self-reliant: skilled reading comprehenders display more deliberate reading strategies and rely less on the support of other people than weak comprehenders (e.g., botsas & padeliadu, 2003; denton et al., 2015). moreover, skilled comprehenders help themselves more actively: they use more problem-solving strategies and a larger variety of reading strategies when texts become difficult (cantrell & carter, 2009; kletzien, 1991). and, interventions that seek to enable students to become self-regulated and proactive were found to be more effective in promoting reading comprehension than teaching cognitive reading strategies alone (e.g., dignath, buettner, & langfeldt, 2008; schünemann, spörer, & brunstein, 2013). taken together, these results suggest that the level of reading achievement depends on how actively, self-reliant, and persistently students regulate their own reading process. accordingly, pi was found to predict active reading behaviour and individual word reading trajectories at elementary school beyond the effects of intrinsic motivation (warner, fay, schiefele, stutz, & wollny, 2016; wollny, 2015). moreover, reading achievement has been found to be affected by students’ task-focused and engaged behaviour (e.g., diperna, volpe, & elliott, 2005; hirvonen, georgiou, lerkkanen, aunola, & nurmi, 2010). yet, direct evidence pertaining to a positive relation between pi and children’s reading comprehension development is missing so far. warner&et &al & & & | f l r ! ! 4! 1.2 the mediating role of reading strategy knowledge reading strategy knowledge refers to the knowledge about what reading strategies can be applied, how strategies can be used, and under which conditions they should be used (paris, lipson, & wixson, 1983). with this, reading strategy knowledge represents a specific type of metacognitive knowledge (the knowledge about one’s own and others’ thought processes; pintrich, 2002). according to boekaert’s theory of self-regulated learning (1999), the development of metacognitive knowledge represents an intermediate process between more general self-regulatory tendencies and the actual use of cognitive strategies. thus, the development of reading strategy knowledge should be affected by a child’s general tendency to display pi. moreover, reading strategy knowledge should represent a prerequisite for the efficient use of reading strategies and thereby also for successful reading comprehension. reading strategy knowledge develops when students practice reading strategies on their own, take an active approach in their learning process, and choose moderately difficult texts (borkowski et al., 2000; gourgey, 1998; mckeachie, 1987). pi should provide an advantage within this process. because children who score higher on pi are self-starting, we propose that they show spontaneous use of reading strategies even when not instructed to do so. because they are proactive, we suggest that they realize the long-term benefits of reading strategies and thus are more motivated in acquiring new strategies. moreover, when faced with reading failure, they likely stay persistent and apply a range of problemsolving strategies. and finally, because they are learning-oriented (wollny et al., 2016), we assume that they choose challenging texts that require the use of reading strategies. the development of higher strategy knowledge in turn is thought to promote the development of reading comprehension (borkowski et al., 2000). this assumption rests upon the theoretical considerations and empirical findings that strategy knowledge results in higher frequency and efficiency of strategy usage, which, finally, results in higher reading comprehension (borkowski et al., 2000; kolićvehovec, zubković, & pahljina-reinić, 2014). for example, van kraayenoord and schneider (1999) found that reading strategy knowledge predicted reading comprehension in third and fourth grade students. thus, we assume that pi is a critical factor contributing to the independent use of strategies and thereby also to the acquisition of more reading strategy knowledge. reading strategy knowledge in turn defines a relevant predictor of reading comprehension. these theoretical suggestions and empirical findings imply that reading strategy knowledge should act as a mediator between pi and the development of reading comprehension. 1.3 the present study reading comprehension is commonly described as the result of a highly active process which is self-starting, proactive, and persistent in nature (e.g., pressley & afflerbach, 1995). however, when reviewing the literature two shortcomings become apparent. first, many studies addressed the role of rather basic self-regulatory skills, whereas constructs with an explicit focus on the proactive aspects of self-regulation have received less attention. and second, although a growing body of research has demonstrated positive relations between reading comprehension and domain-unspecific self-regulatory measures, little research exists on the mediating processes of these relations. the present study therefore investigated the following research questions: 1) does pi positively predict the development in reading comprehension? 2) does pi contribute to reading comprehension through the development of reading strategy knowledge? based on the literature reviewed, we hypothesized that pi contributes to students’ development in reading comprehension over time (hypothesis 1). as a mediating process we expected positive indirect contributions of pi to reading comprehension that are mediated by reading strategy knowledge (hypothesis 2). we used a large-scale longitudinal study with two data collection points to examine these proposed associations. the hypothesized model is displayed in figure 1. warner&et &al & & & | f l r ! ! 5! the model included measures of pi and reading comprehension at time 1 and time 2. we assessed pi also at a second time point because this controls for spurious relations and leads to more accurate estimates of the direct and indirect effects (e.g., kenny, 1975; little, preacher, selig, & card, 2007). we examined reading strategy knowledge only at time 2 because previous evidence indicates that children’s reading strategy use is underdeveloped at early grade levels (e.g., kolić-vehovec & bajsanski, 2001; myers & paris, 1978) and therefore their self-reported strategic behaviour demonstrates low or none relations with reading performance (e.g., wernke, 2012). because reading comprehension is thought to affect the development of reading strategy knowledge (e.g., borkowski et al., 2000), we also assumed an effect from time 1 reading comprehension to time 2 reading strategy knowledge. we took effects of grade level into account and examined the proposed mediational process in two groups of elementary students. students of group 1 attended the third grade at time 1 and the fifth grade at time 2. students of group 2 attended the fourth grade at time 1 and the sixth grade at time 2. comparing these two groups provides useful information for future interventions in revealing whether the proposed developmental mechanisms can be generalized across grades or whether they differ by grade level because of developmental dynamics and progress in classroom instruction. figure 1. hypothesized structural equation model. 2. method 2.1 sample and procedure the present research is part of a larger project on intrapersonal risk and protective factors in the development of children and adolescents (pier study). because children were studied on a range of protective and risk factors, only limited time was available for each measure. a school-based sampling t1 personal initiative t2 personal initiative t1 reading comprehension t2 reading comprehension t2 strategy knowledge warner&et &al & & & | f l r ! ! 6! method was used to recruit children and their parents. first, 110 schools from different socioeconomic backgrounds were asked to take part in the study. after schools agreed to participate, parents of first-, second-, and third-grade students were informed about the aims of the study and were asked to participate. this resulted in 1657 children who took part at the first measurement wave (time 0) of the project (580 families refused). the present study focused on the second (time 1) and third measurement wave (time 2) and examined two groups of elementary students and their parents or legal guardians. group 1 consisted of students who attended third grade at time 1 and fifth grade at time 2. this group included 546 students from 66 classes at time 1 (mage = 9.09, sd = 0.49; 52% girls). group 2 consisted of students who attended fourth grade at time 1 and sixth grade at time 2. this group comprised 556 students from 56 classes at time 1 (mage = 10.05, sd = 0.45; 51% girls). at the first wave (time 0), we did not use any measures of reading comprehension because students had not yet achieved this skill level. however, we obtained one of the control measures at this point in time (see below). the time interval between time 0 and time 1 was approximately nine months (mt1−t0 = 8.9 months, sd = 1.7 months) and between time 1 and 2 approximately two years (mt2−t1 = 23.5 months, sd = 1.6 months). students came from 36 different schools in the federal state of brandenburg (germany). at each wave, trained university students tested the school students individually in two separate one-hour sessions. at about the same time, parents were requested to fill in a questionnaire that included measures capturing their child’s behaviour and personality. 2.2 measures 2.2.1. personal initiative at time 1 and 2 pi was assessed by means of an 8-item parent questionnaire (wollny et al., 2016). the questionnaire includes four subscales, each measuring one aspect of pi: self-starting (αt1 = .57, αt2 = .67; e.g., “my child independently searches for new tasks.”), proactive (αt1 = .82, αt2 = .81; e.g., “my child actively approaches problems.”), persistent (αt1 = .72, αt2 = .74; e.g., “if my child has independently set itself a learning goal, then it persistently pursues it.”), and extra-role behaviour (αt1 = .86, αt2 = .88; e.g., “my child takes initiative, even if others do not.”). answers were given by the mother or father of the child on a five-point scale ranging from 1 (does not apply at all) to 5 (fully applies). the internal consistency was high (αt1 = .88; αt2 = .90). the questionnaire was identical to that constructed by wollny et al. (2016). 2.2.2. reading comprehension at time 1 as an indicator for previous reading comprehension, we used the reading comprehension section from the elfe test (lenhard & schneider, 2006). due to time constraints and concerns about training effects, we split the original test into two halves and applied only one of these halves at time 1 (see stutz, schaffner, & schiefele, 2016). students had three and a half minutes to read six short narrative text passages followed by one or two multiple-choice questions. to correctly choose the right answer, students had to find isolated information from the text, make interferences, and build anaphoric relations. at any time, the text passages were available to students. the internal consistency was adequate (α = .79). 2.2.3. reading comprehension at time 2 to measure reading comprehension at time 2, students read an expository text “die tiefsee” (the deep sea) from the german reading test lesen 6-7 (bäuerlein, lenhard, & schneider, 2012). due to time constraints, we shortened the original text by 42% of its length (from 550 to 320 words) and reduced the original 17 multiple-choice items to 10 items. we primarily excluded items on word-, sentence-, and text knowledge in favour of retaining items that required the building of anaphoric inferences and a warner&et &al & & & | f l r ! ! 7! mental model of the text (e.g., “what title would match this text?”). at any time, students were allowed to look back into the text. the internal consistency of the shortened version was acceptable (α =.70). 2.2.4. reading strategy knowledge at time 2 to measure reading strategy knowledge, we developed a self-report questionnaire based on selected strategy questionnaires by wernke, wagener, anschuetz, and moschner (2011), wild and schiefele (1994), pintrich (1991), and lompscher (1995). items of these questionnaires were rephrased in the way that they measure strategy knowledge and not the actual use of strategies. moreover, we took care only to select items which we judged to be helpful for the reading comprehension test at time 2. the resulting nine items assessed knowledge about elaboration strategies (4 items, e.g., “are you familiar with trying to visualize objects, situations, or characters that occur in the text?”), combined organization/elaboration strategies (4 items, e.g., “are you familiar with trying to identify the most relevant information of a text section?”), and control strategies (1 item, e.g., “are you familiar with going back in the text and looking at and reading again some text passages and sections?”). elaboration strategies help to embed new information into prior knowledge, organization strategies facilitate the structuring of new information, and control strategies support the monitoring of comprehension (weinstein & mayer, 1986). answers were given on a three-point scale ranging from no, i’ve never heard before, to yes, i’ve heard before, and yes, i know how to do it. to prevent priming effects, reading comprehension was measured at session 1, whereas strategy knowledge was assessed at session 2 (approximately five days later). we applied exploratory factor analysis with oblique rotation (geomin) in mplus to explore the factor structure of the questionnaire. as extraction method we used the model fit and the theoretical plausibility of factor solutions. the questionnaire demonstrated a unidimensional structure, which was supported by confirmatory factor analysis, χ2(27) = 123.01, p < .001, cfi = .966, rmsea = .059. ordinal alpha was also adequate (α = .86; gadermann, guhn, & zumbo, 2012). to test the validity of the measure, we included one item that presented a non-existing reading strategy, i.e., the “green-paper strategy” (“are you familiar with working on a text by using the green-paper strategy?”). while some students reported to have heard of this strategy (n = 110), only 2.7% (n = 30) of the total sample reported to know how to apply this strategy. we see this as some indication of a valid assessment of strategy knowledge. finally, compared to sixth grade students (n = 60) a higher proportion of fifth grade students (n = 80) reported to have heard from or to know this strategy, χ2 (2) = 7.80, p = .020. 2.2.5. control variables reading performance and strategic skills draw to some extent on an individual’s basic cognitive skills, gender, and socioeconomic status (e.g., artelt, naumann, & schneider, 2010; bowey, 2005; chiu & mcbride-chang, 2006). thus, to ensure that effects of pi were robust against the influence of these variables, we controlled for gender, the number of books at home (as a measure of socioeconomic status), and a measure of cognitive ability assessing different aspects of nonverbal intelligence (digital symbol substitution test; petermann & petermann, 2008). parents rated the number of books at home at time 0. the response scale ranged from 1 (0 – 10 books) to 6 (more than 500 books). 2.3 analytical strategy 2.3.1. structural equation analysis in a first step, we screened all central variables for outliers, but did not identify values deviating more than three standard deviations from the mean. an exception was the measure of reading strategy knowledge which showed eight small outliers. we considered these outliers as realistic and largely due to the limited variance of reading strategy knowledge overall. due to these results, we decided to use robust warner&et &al & & & | f l r ! ! 8! maximum likelihood (mlr; used in confirmatory factor analyses of pi and reading comprehension) and weighted-least square parameter estimates (wlsmv; used in all models including ordinal variables) to control for outliers and nonnormality (brown, 2006; muthén & muthén, 1998-2012). hypotheses were tested using structural equation modelling by means of mplus 7.3 (muthén & muthén, 1998-2012). in line with the conceptualization of the construct and preliminary factor-analytic work (wollny et al., 2016), pi was modelled as a second-order factor with first-order factors each representing one subscale of the questionnaire. reading strategy knowledge was modelled as a one-factor model. because the reading comprehension measures comprised too many items for a stable factor solution, we used the odd-even parceling method to obtain a model of reading comprehension at time 1 and at time 2 that is more parsimonious and less affected by random error (e.g., matsunaga, 2008). confirmatory factor analyses yielded adequate model fits for all measurement models. since the students were nested in school classes, we computed intraclass correlation coefficients (icc) for pi, reading strategy knowledge, and reading comprehension with school class as cluster variable in order to check the level of non-independence. iccs for reading comprehension were .17 and .14 for time 1 and 2, respectively, whereas pi and reading strategy knowledge showed smaller iccs (between .03 and .04). thus, a substantial amount of variance in reading comprehension could be explained by the class membership of students. to take this non-independence of observations into account, we controlled for the nested nature of the data by using the mplus syntax type = complex. reading strategy knowledge and the number of books per household were measured with an ordinal answering scale. therefore, we used a means and variances adjusted weighted least squares estimator (wlsmv) in all models that included these variables. models were evaluated as acceptable if they met the following cut-off criteria: the comparative fit index (cfi) and tucker–lewis index (tli) above .95, the root-mean-square-error of approximation (rmsea) below .08, and the standardized rootmean-square residual (srmr) below .05 (hu & bentler, 1999; schermelleh-engel, moosbrugger, & müller, 2003). we also report the χ2-statistic, which, however, was not used as criterion due to its oversensitivity to large sample sizes (schermelleh-engel et al., 2003). 2.3.2. measurement invariance prior to testing differences between measurement waves/groups, the measurement models have to be tested for invariance. models must demonstrate at least equivalent factor loadings across time/groups (metric invariance; vandenberg & lance, 2000). we therefore analysed measurement invariance across time for pi (because it was assessed by the same questionnaire at time 1 and time 2). we did not test longitudinal measurement invariance of reading comprehension, because different measures of reading comprehension were used at time 1 and 2. we further tested measurement invariance across groups for pi, reading strategy knowledge, and reading comprehension. in doing so, we applied a stepwise procedure. first, we fixed the factor structure (configural invariance), then factor loadings (metric invariance), then intercepts (scalar invariance), and then residual variances (strict invariance) to be equal across time points/groups (vandenberg & lance, 2000). the level of measurement invariance was identified by comparing the model fit at every step of invariance to that of the preceding step of invariance. to this end, we tested measures of reading comprehension at time 1 and time 2 together in an integrated model. this provided the obligatory degrees of freedom for testing all steps of measurement invariance. measures of pi and reading strategy knowledge were tested separately for measurement invariance. as indicators for substantial reduction of model fit, we used the cut-off rules suggested for cfi, rmsea, and srmr by chen (2007; δcfi ≥ -.010, supplemented by δrmsea ≥ .015 or δsrmr ≥ -.010). 2.3.3. missing data at time 2, 1,009 students participated again (dropout rate = 6.3%). moreover, 892 parents participated at time 0, 803 participated at time 1, and 709 at time 2 (dropout rate t0-t1 = 10.0%; t1-t2 = warner&et &al & & & | f l r ! ! 9! 11.7%). note that the only measure obtained at time 0 was the number of books at home rated by parents. beyond these participation and attrition rates, all variables showed less than 4% missing values. if possible, we used the full information maximum likelihood (fiml) to account for missing data. however, when models included ordinal variables (i.e., reading strategy knowledge), we used the wlsmv algorithm as the default estimator of mplus. if control variables are included, the wlsmv estimator handles missing values by a mixture of fiml and pair-wise procedure (newsom, 2015). 3. results 3.1 descriptive results and correlations table 1 presents means and standard deviations of continuous variables for the total sample and separated by group. we conducted several two-tailed t-tests to examine potential group effects on the continuous variables. as could be expected, students of group 2 (older students) had a significantly higher mean level of reading comprehension at time 1 and 2 (table 1). somewhat unexpected, older students showed higher cognitive ability than younger students, although the scores were normed by age (t values). mean levels of pi did not differ between groups, which is consistent with previous findings on older students (study 1 in wollny et al., 2016). because socioeconomic status and reading strategy knowledge were measured on an ordinal scale, we report their frequencies instead. parents reported to have zero to ten books (0.6%), 11 to 25 books (2.5%), 26 to 100 books (21.1%), 101 to 200 books (21.5%), 201 to 500 books (34.4%), or more than 500 books (18.4%) in their household. students reported on average that they do not know any reading strategy (5%), that they have heard of any reading strategy (34.4%), or that they know how to apply any strategy (60.5%). groups did not differ in socioeconomic status, χ2 (5) = 4.01, p = .557. however, as could be expected, students in group 2 reported higher reading strategy knowledge than students in group 1, χ2 (17) = 38.28, p = .002. table 2 presents bivariate correlations between all study variables and control variables. we also controlled for group and computed partial correlations, which were similar to bivariate correlations (table 2). 3.2 measurement invariance the results of testing measurement invariance across time points (time 1 vs. time 2) and groups (group 1 vs. group 2) are displayed in the appendix (table a.1). with respect to measurement invariance across time, we found that the model of pi was strictly invariant. with respect to differences in measurement models across groups, the measurement models of reading strategy knowledge and of reading comprehension proved strict measurement invariance. in case of pi, the second-order factor model did not provide the necessary degrees of freedom to test configural invariance in a conventional sense. however, van de schoot, lugtig, and hox (2012) proposed that configural measurement invariance is also given if the model is found to be valid independently in each group. due to the fact that pi showed an adequate and comparable model fit in both groups (group 1: χ2(111) = 209.82, p < .001, cfi = .964, rmsea = .045; srmr = .050; group 2: χ2(111) = 260.22, p < .001, cfi = .955, rmsea = .055, srmr = .049), we concluded that pi was configural invariant across groups. we then proceeded with the regular way of testing measurement invariance and found that fixing the intercepts to be equal produced an intolerable deterioration of model fit (δcfi ≥ -.010, supplemented by δrmsea ≥ .015 or δsrmr ≥ .010; chen, 2007). as a consequence, with respect to groups, we could only establish metric invariance for pi. warner&et &al & & & | f l r ! ! 10! table 1 mean values, standard deviations, and group differences of continuous variables total sample group 1 group 2 two-tailed t-test effect size variable n range m sd m sd m sd t (df) d t1 cognitive ability 1098 27 80 54.63 8.95 54.09 8.95 55.16 8.91 -1.98* (1096) .12 t1 personal initiative 803 1 5 3.05 0.77 3.02 0.77 3.08 0.77 -1.00 (801) .06 t2 personal initiative 709 1 5 3.10 0.82 3.05 0.80 3.14 0.83 -1.50 (707) .11 t1 reading comprehension 1102 0 10 6.14 2.52 5.41 2.48 6.86 2.35 -9.95*** (1100) .60 t2 reading comprehension 1003 0 10 4.93 2.43 4.50 2.37 5.35 2.42 -5.62*** (1001) .35 note. this table does not include descriptive statistics of socioeconomic status and reading strategy knowledge because these variables were measured on an ordinal scale (instead, see frequencies and chi-squared tests reported in the descriptive results section). group 1 refers to students in 3rd grade at time 1 and 5th grade at time 2. group 2 refers to students in 4th grade at time 1 and 6th grade at time 2. *p < .05. **p < .01. ***p < .001. table 2 bivariate and partial correlations among study variables and control measures (total sample) variable 1 2a 3b 4 5 6 7b 8 9 group --------- t1 gendera .03 -.07* -.20*** -.09* -.17*** -.04 -.07* .06* t0 sesb .00 .07* -.06 .14*** .14*** .07* .27*** .33*** t1 cognitive ability .06* -.19*** .06 -.19*** .17*** .07* .33*** .15*** t1 personal initiative .04 -.09* .14** .19*** -.67*** .16*** .21*** .22*** t2 personal initiative .06 -.17*** .14** .17*** .67*** -.21*** .28*** .25*** t2 strategy knowledgeb .12*** -.04 .07* .08* .17*** .22*** -.20*** .16*** t1 reading comprehension .29*** -.06 .26*** .33*** .21*** .28*** .22*** -.53*** t2 reading comprehension .18*** .07* .32*** .16*** .22*** .26*** .18*** .55*** - note. n = 626-1099. ses = socioeconomic status. bivariate correlations are shown below the diagonal, partial correlations controlled for group are shown above the diagonal. afemale = 1, male = 2. bwe computed spearman rank-based correlations for ordinal variables. *p < .05. **p < .01. ***p < .001. warner&et &al & & & | f l r ! ! 11! 3.3 group effects in structural equation modelling after establishing measurement invariance, the hypothesized model (see figure 1) was tested for moderating effects by group. to this end, we used a multiple group approach with grade level as a group factor. this group-specific mediation model revealed an acceptable model fit. in the next step, we fixed all directed paths to be equal across groups and compared the resulting model fit with the unconstrained model. this yielded a non-significant χ2 difference between the models, δχ2 (22) = 17.17, p = .532. however, this test of overall differences has low sensitivity to differences in single paths. we therefore tested every single directed path for significance and found significant group differences in the paths from time 1 reading comprehension to time 2 reading comprehension (δχ² = 11.76***), time 1 reading comprehension to time 2 pi (δχ² = 12.75***), and time 2 strategy knowledge to time 2 reading comprehension (δχ² = 5.03*). as can be seen in figure 2, the last two of these paths reached significance in group 2 but not in group 1. moreover, the impact of time 1 reading comprehension on time 2 reading comprehension was found to be higher in group 1 than in group 2. finally, we also found a range of group differences with respect to effects by control variables (see table a.2 in the appendix). together, these differences indicate that the group status – in other words, whether students were in the lower or higher grade – had an essential impact on the estimated structural relations and the hypothesized mediation model. thus, the two groups differed considerably in their developmental dynamics. as a result, we freed all paths that varied significantly across groups. the resulting model exhibited a good model fit with the data, χ2(945) = 1072.23, p = .002, cfi = .972, rmsea = .016. 3.4 relations among personal initiative, reading strategy knowledge, and reading comprehension after the establishment of a parsimonious model that took group effects into account, we tested the hypotheses. as depicted in figure 2, the autoregressive effects of pi and the effect of time 1 reading comprehension on time 2 reading comprehension were high in both groups. in terms of our hypothesized mediation model, we found that time 1 pi positively predicted time 2 reading strategy knowledge in both groups. however, only in the group from fourth to sixth grade reading strategy knowledge significantly contributed to time 2 reading comprehension. thus, we could only test the hypothesized mediation process in group 2. using the monte carlo method (hayes & scharkow, 2013), we found a significant indirect effect of pi on reading comprehension mediated by reading strategy knowledge in group 2, β = .02, p < .05, 95% ci [.009, .036]. furthermore, time 1 pi directly contributed to time 2 reading comprehension in both groups. thus, the observed indirect effect from fourth to sixth grade indicated partial mediation. our model also included a number of paths that were not part of our mediation hypothesis. as we had expected, time 1 reading comprehension positively predicted reading strategy knowledge at time 2 in both groups. by testing bidirectional relations, we further observed that time 1 reading comprehension significantly predicted time 2 pi in group 2. 4. discussion previous research has put little emphasis on the role of pi in the prediction of reading comprehension. this study aimed at closing this gap in literature by examining the role of pi in predicting reading comprehension in two groups of elementary students and by investigating the underlying mechanism of this proposed relation. in particular, we hypothesized a mediational process in which pi is positively related to reading comprehension via the development of reading strategy knowledge. in support of hypothesis 1, we found that pi directly contributed to the development of reading strategy knowledge and warner&et &al & & & | f l r ! ! 12! reading comprehension from third to fifth and fourth to sixth grade. in terms of the proposed mediation (hypothesis 2), the results differed by grade level. figure 2. multiple group structural equation model. only standardized path coefficients are displayed. the first coefficient refers to group 1 and the second coefficient refers to group 2. group 1 refers to students in 3rd grade at time 1 and 5th grade at time 2. group 2 refers to students in 4th grade at time 1 and 6th grade at time 2. directed paths were constrained to be equal except the paths from t2 reading strategy knowledge to t2 reading comprehension, t1 reading comprehension to t2 pi, and t1 cognitive ability to t2 pi. for clarity, residual variances, indicators, and effects of control variables (see appendix, table a.2) are not shown here. *p < .05. **p < .01. ***p < .001. 4.1 personal initiative and reading strategy knowledge it is one of our noteworthy results that pi predicted reading strategy knowledge consistently among two groups of students even though we controlled for previous reading comprehension and for several covariates. this finding provides rare evidence for a positive longitudinal relation between a domainunspecific self-regulatory construct, i.e., pi, and the development of specific metacognitive knowledge, i.e., reading strategy knowledge. it further complements the view of self-regulated learning as a primarily domain-specific construct (e.g., massey, 2009) and emphasizes the interconnectedness of domain-unspecific with domain-specific self-regulatory behaviour. moreover, scholars repeatedly emphasized that metacognitive knowledge relies considerably on direct instructions through teaching (e.g., borkowski et al., 2000; pintrich, 2002). in agreement with this, class membership represented an important determinant of reading comprehension in the present sample. therefore, in the analyses, we took into account that students were nested in classes. even when taking these systematic differences based on instruction into account, pi remained a significant predictor.!thus, although t2 reading comprehension t2 strategy knowledge t1 personal initiative t2 personal initiative t1 reading comprehension r 2 = .50***/ .61*** r 2 = .60***/ .64*** r 2 = .09***/ .09*** .07/ .23* .14/ -.14 .73***/ .72*** .65***/ .62*** .20**/ .21*** .17***/ .19*** .05/ .24*** -.00/ .11* ..21***/ .22*** .11*/ .11* warner&et &al & & & | f l r ! ! 13! the instruction of reading strategy knowledge is firmly anchored in german curricula, the present findings suggest that the acquisition of appropriate reading strategy knowledge depends to some extent also on the proactive efforts of an individual. this by no means indicates that school instruction is dispensable in the acquisition of reading strategy knowledge. instead, it highlights that students might differ in how much they apply, expand, and transfer this knowledge on their own initiative. 4.2 reading strategy knowledge and reading comprehension when testing the mediation hypothesis, the results only provided partial support for the proposed link between reading strategy knowledge and reading comprehension. with regard to sixth grade students, we found a positive effect of reading strategy knowledge on reading comprehension. in fifth grade students, however, reading strategy knowledge failed to show any substantial association with reading comprehension. two reasons could account for this finding. first, it is possible that fifth grade students might have had difficulties in transferring their reading strategy knowledge into efficient reading activities. this is referred to as utilization deficiency. this deficiency describes a developmental stage, in which students correctly know how to use reading strategies but have not automated them to such an extent that their reading comprehension benefits notably from their application (miller, 1990). according to the curriculum of the federal state of brandenburg, only from fifth grade onwards the instruction of reading strategies is systematically extended to a larger range of school subjects (ministery of education of the federal state of brandenburg, 2004). thus, fifth grade students might have just been beginning to extend their reading strategy knowledge to other school subjects and might simply have had less practice in applying reading strategies than sixth grade students. second, we found that the measure of reading strategy knowledge was less valid in the sample of fifth grade students than in the sample of sixth grade students. by asking about their knowledge of a nonexisting reading strategy (see 2.2.4.), we observed that fifth grade students more frequently reported to be familiar with this strategy than sixth grade students. this is in accordance with findings that younger children’s metacognitive self-reports are prone to overestimation (e.g., ferreira, simão, & da silva, 2015). however, regardless of the underlying mechanism (i.e., utilization deficiency or overestimation), our findings complement results of previous studies in which reading strategy knowledge seemed to be more highly related with reading performance at higher grade levels (e.g., kolić-vehovec et al., 2014; van kraayenoord & schneider, 1999). 4.3 personal initiative and reading comprehension the hypothesized indirect effect from pi on reading comprehension via reading strategy knowledge was supported for more advanced students (i.e., students who moved from fourth to sixth grade from time 1 to 2), but not for the less advanced group of students (i.e., students who moved from third to fifth grade). thus, with regard to older elementary students, a general tendency to show pi seems to lead to the development of reading strategy knowledge, which in turn promotes reading development. these findings highlight the potential of pi in promoting learning processes (such as reading development), as emphasized in theories of self-regulated learning (e.g., zimmerman, 2002). they further represent a step forward in supporting theories that propose a complex interconnection between a person’s relatively stable selfregulatory tendencies (e.g., pi), domain-specific strategic behaviour (e.g., development of reading strategy knowledge), and domain-specific academic success (e.g., reading comprehension; e.g., boekaerts, 1999; borkowski et al., 2000; efklides, 2011). it is important to note, however, that the effect size of the observed mediation was small. moreover, in both groups we also found a direct contribution of pi on reading comprehension that was not mediated by reading strategy knowledge. this might be due to methodological reasons. for example, students might have reported that they know a certain strategy, but might have failed to apply the strategy during the actual warner&et &al & & & | f l r ! ! 14! reading task (bjorklund & douglas, 1997; flavell, 1970). moreover, variance of strategy knowledge was comparatively low (see table 1) and the questionnaire did not assess the entire range of potentially useful reading strategies (see mokhtari & reichard, 2002). thus, students might have used alternative reading strategies, which were simply not assessed in the present study and were thus not reported by the students. however, the direct effect of pi on reading comprehension might also indicate that other mediating mechanisms exist beyond the use of and knowledge about reading strategies. for example, students with a high level of pi tend to be learning-oriented (wollny et al., 2016) and therefore might simply read more, prefer to read a larger variety of text genres, be more engaged in reading lessons, and choose more challenging text material than students lower on pi. finally, there was also a reversed long-term effect of reading comprehension on pi in the group of sixth grade students. this effect was not surprising considering that adequate reading performance promotes the development of academic self-concepts (e.g., chapman, tunmer, & prochnow, 2000) and might intensify a potential big fish little pond effect. thus, good readers likely develop strong internal control and selfefficacy beliefs, which in turn relate to higher scores of pi in children (wollny et al., 2016). because the developmental shift from “learning to read” to “reading to learn” happens around fourth grade (rupley & willson, 1996), we assume that the mutually supportive mechanisms between pi and reading comprehension take some time to unfold. this might explain why we found this reversed effect only in the group of sixth grade students. 4.4 limitations and future research there are some limitations to bear in mind when interpreting the results from the present study. first of all, we investigated the hypothesized mediational process by using only two measurement points. thus, the observed associations between reading strategy knowledge and reading comprehension were only crosssectional. additional research is therefore necessary to replicate our findings within a three-wave longitudinal study. second, because this study was conducted within the scope of a larger longitudinal project, we had to use relatively short measures of reading strategy knowledge and reading comprehension. as a result, the variance of our measure of reading strategy knowledge was low, which may have reduced the magnitude of the observed mediational relations. future studies might benefit from using measures that cover a larger variety of reading strategies (e.g., mokhtari & reichard, 2002; wernke et al., 2011). this requires also the use of more extended assessments of reading comprehension, such as the use of longer texts and different text types. because pi is a domain-unspecific construct, it might show similar effects across text types. on the other hand, the effects of pi might also be contingent on students’ preferences for particular text types (e.g., fiction vs. non-fiction) or might be particularly helpful in reading texts that require many proactive efforts (e.g., hypertexts). third, the measures of pi and reading comprehension demonstrated a high stability over time leaving little variance to be explained by predictors. the control of stabilities therefore likely reduced the magnitude of observed effects of pi and reading strategy knowledge on the development of reading comprehension (as well as of reading comprehension on the development of pi). however, although the observed effect sizes appeared to be small, they are not trivial in light of the conservative longitudinal approach of this study (e.g., adachi & willoughby, 2015). fourth, this study focused entirely on intrapersonal predictors of reading comprehension leaving the question open as to how predictors beyond individual characteristics, such as the quality of instruction or the level of classroom performance, affect these intrapersonal processes. finally, we examined pi by using parent ratings. parents observe their child’s pi in a range of contexts that are not limited to the academic learning environment. thus, parent ratings of pi represent a broad measure of pi, which may have had a diminishing effect on the relations found in the present study. warner&et &al & & & | f l r ! ! 15! therefore, in future studies, researchers might consider using self-reports and observational measures of pi and combining them within a multi-method approach. 5. conclusions results from the present study are unique in that they document the contribution of students’ pi on the development of both reading strategy knowledge and reading comprehension. these findings underline the potential relevance of pi as a contributor to academic success. it should be noted, however, that we found only limited support for the proposed mediation. this limitation aside, the results from the present study suggest that efforts to foster pi might have positive long-term effects on reading comprehension. in adults, emerging evidence indicates that pi is amenable to training and can be increased by higher job autonomy and job control (e.g., glaub, frese, fischer, & hoppe, 2014; li, fay, frese, harms, & gao, 2014). likewise, pi relates to self-efficacy and internal control orientations in children (wollny et al., 2016). therefore, teaching practices should be aimed at encouraging students’ sense of responsibility for their own learning progress. this can be achieved by providing open-ended and student-centered learning environments (e.g., paris & paris, 2001; turner, 1995). thus, students should be encouraged to choose their own learning tasks, to seek solutions on their own, and to develop and express their own learning goals (e.g., duke & pearson, 2002; stefanou, perencevich, dicintio, & turner, 2004). keypoints developing reading comprehension requires personal initiative. we used longitudinal structural equation modelling with data on 1,102 elementary students. personal initiative predicted reading strategy knowledge and reading comprehension. strategy knowledge mediated the effects of personal initiative from grade 4 to 6. acknowledgements this research was funded by the german research foundation [dfg; grant number 1668/1]. references adachi, p., & willoughby, t. (2015). interpreting effect sizes when controlling for stability effects in longitudinal autoregressive models: implications for psychological science. european journal of developmental psychology, 12, 116-128. doi:10.1080/17405629.2014.963549 artelt, c., naumann, j., & schneider, w. (2010). lesemotivation und lernstrategien [reading motivation and reading strategies]. in e. klieme, c. artelt, j. hartig, n. jude, o. köller, m. prenzel, w. schneider & p. stanat (eds.), pisa 2009. bilanz nach einem jahrzehnt (pp. 73-112). münster, germany: waxmann. bäuerlein, k., lenhard, w., & schneider, w. (2012). lesetestbatterie für die klassenstufen 6-7 [reading test battery for 6thand 7th-graders]. göttingen, germany: hogrefe. best, j. r., miller, p. h., & jones, l. l. (2009). executive functions after age 5: changes and correlates. developmental review, 29, 180-200. doi: 10.1016/j.dr.2009.05.002 warner&et &al & & & | f l r ! ! 16! birgisdóttir, f., gestsdóttir, s., & thorsdóttir, f. (2015). the role of behavioral self-regulation in learning to read: a 2-year longitudinal study of icelandic preschool children. early education and development, 26, 807-828. doi: 10.1080/10409289.2015.1003505 bjorklund, d. f., & douglas, r. n. (1997). the development of memory strategies. in n. cowan & n. cowan (eds.), the development of memory in childhood (pp. 201-246). hove, england: psychology press/erlbaum (uk) taylor & francis. boekaerts, m. (1999). self-regulated learning: where we are today. international journal of educational research, 31, 445-457. doi: 10.1016/s0883-0355(99)00014-2 borkowski, j. g., chan, l. k., & muthukrishna, n. (2000). a process-oriented model of metacognition: links between motivation and executive functioning. in g. schraw & j. impara (eds.), issues in the measurement of metacognition (pp. 1-41). lincoln, nb, us: university of nebraska press. botsas, g., & padeliadu, s. (2003). goal orientation and reading comprehension strategy use among students with and without reading difficulties. international journal of educational research, 39, 477-495. doi: 10.1016/j.ijer.2004.06.010 bowey, j. a. (2005). predicting individual differences in learning to read. in m. j. snowling, c. hulme, m. j. snowling & c. hulme (eds.), the science of reading: a handbook. (pp. 155-172). malden, ma, us: blackwell publishing. brown, t. a. (2006). confirmatory factor analysis for applied research. new york, ny, us: guilford press. cain, k., oakhill, j., & bryant, p. (2004). children's reading comprehension ability: concurrent prediction by working memory, verbal ability, and component skills. journal of educational psychology, 96, 3142. doi: 10.1037/0022-0663.96.1.31 cantrell, s. c., & carter, j. c. (2009). relationships among learner characteristics and adolescents' perceptions about reading strategy use. reading psychology, 30, 195-224. doi: 10.1080/02702710802275397 caraway, k., tucker, c. m., reinke, w. m., & hall, c. (2003). self-efficacy, goal orientation and fear of failure as predictors of school engagement in high school students. psychology in the schools, 40, 417427. doi: 10.1002/pits.10092 chapman, j. w., tunmer, w. e., & prochnow, j. e. (2000). early reading-related skills and performance, reading self-concept, and the development of academic self-concept: a longitudinal study. journal of educational psychology, 92, 703-708. doi: 10.1037/0022-0663.92.4.703 chen, f. f. (2007). sensitivity of goodness of fit indexes to lack of measurement invariance. structural equation modeling, 14, 464-504. doi: 10.1080/10705510701301834 chiu, m. m., & mcbride-chang, c. (2006). gender, context, and reading: a comparison of students in 43 countries. scientific studies of reading, 10, 331-362. doi: 10.1207/s1532799xssr1004_1 denton, c. a., wolters, c. a., york, m. j., swanson, e., kulesz, p. a., & francis, d. j. (2015). adolescents' use of reading comprehension strategies: differences related to reading proficiency, grade level, and gender. learning and individual differences, 37, 81-95. doi: 10.1016/j.lindif.2014.11.016 dermitzaki, i., andreou, g., & paraskeva, v. (2008). high and low reading comprehension achievers' strategic behaviors and their relation to performance in a reading comprehension situation. reading psychology, 29, 471-492. doi: 10.1080/02702710802168519 dignath, c., buettner, g., & langfeldt, h.-p. (2008). how can primary school students learn self-regulated learning strategies most effectively?: a meta-analysis on self-regulation training programmes. educational research review, 3, 101-129. doi: 10.1016/j.edurev.2008.02.003 diperna, j. c., volpe, r. j., & elliott, s. n. (2005). a model of academic enablers and mathematics achievement in the elementary grades. journal of school psychology, 43, 379-392. doi:10.1016/j.jsp.2005.09.002 duke, n. k., & carlisle, j. (2011). the development of comprehension. in m. l. kamil, p. d. pearson, e. b. moje & p. p. afflerbach (eds.), handbook of reading research (vol. 4, pp. 199-228). new york, ny, us: routledge. warner&et &al & & & | f l r ! ! 17! duke, n. k., & pearson, p. (2002). effective practices for developing reading comprehension. in a. e. farstrup & s. j. samuels (eds.), what research has to say about reading instruction (3rd ed., pp. 205242). newark, de, us: international reading association. efklides, a. (2011). interactions of metacognition with motivation and affect in self-regulated learning: the masrl model. educational psychologist, 46, 6-25. doi: 10.1080/00461520.2011.538645 fay, d., & frese, m. (2001). the concept of personal initiative: an overview of validity studies. human performance, 14, 97-124. doi: 10.1207/s15327043hup1401_06 ferreira, p. c., simão, a. m. v., & da silva, a. l. (2015). the unidimensionality and overestimation of metacognitive awareness in children: validating the catom. anales de psicología, 31, 890-900. doi: 10.6018/analesps.31.3.184221 flavell, j. h. (1970). developmental studies of mediated memory. in h. w. reese & l. p. lipsitt (eds.), advances in child development and behavior (vol. 5, pp. 181-211). new york, ny, us: academic press. frese, m., & fay, d. (2001). personal initiative: an active performance concept for work in the 21st century. in b. m. staw & r. i. sutton (eds.), research in organizational behavior (vol. 23, pp. 133-187). san diego, ca, us: elsevier academic press. frese, m., fay, d., hilburger, t., leng, k., & tag, a. (1997). the concept of personal initiative: operationalization, reliability and validity of two german samples. journal of occupational and organizational psychology, 70, 139-161. doi: 10.1111/j.2044-8325.1997.tb00639.x gadermann, a. m., guhn, m., & zumbo, b. d. (2012). estimating ordinal reliability for likert-type and ordinal item response data: a conceptual, empirical, and practical guide. practical assessment, research & evaluation, 17, 1-13. glaub, m. e., frese, m., fischer, s., & hoppe, m. (2014). increasing personal initiative in small business managers or owners leads to entrepreneurial success: a theory-based controlled randomized field intervention for evidence-based management. academy of management learning & education, 13, 354-379. doi: 10.5465/amle.2013.0234 gourgey, a. f. (1998). metacognition in basic skills instruction. instructional science, 26, 81-96. doi: 10.1023/a:1003092414893 hayes, a. f., & scharkow, m. (2013). the relative trustworthiness of inferential tests of the indirect effect in statistical mediation analysis: does method really matter? psychological science, 24, 1918-1927. doi: 10.1177/0956797613480187 hirvonen, r., georgiou, g. k., lerkkanen, m.-k., aunola, k., & nurmi, j.-e. (2010). task-focused behaviour and literacy development: a reciprocal relationship. journal of research in reading, 33, 302-319. doi:10.1111/j.1467-9817.2009.01415.x horner, s. l., & shwery, c. s. (2002). becoming an engaged, self-regulated reader. theory into practice, 41, 102-109. doi: 10.1207/s15430421tip4102_6 hu, l.-t., & bentler, p. m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives. structural equation modeling: a multidisciplinary journal, 6, 1-55. doi: 10.1080/10705519909540118 janke, b., & hasselhorn, m. (2008). frühes schulalter [early school age]. in m. hasselhorn & r. k. silbereisen (eds.), theorie und forschung, enzyklopädie der psychologie, serie entwicklungspsychologie (vol. 7, pp. 413-447). göttingen: hogrefe. jimerson, s. r., campos, e., & greif, j. l. (2003). toward an understanding of definitions and measures of school engagement and related terms. california school psychologist, 8, 7-27. doi: 10.1007/bf03340893 kenny, d. a. (1975). cross-lagged panel correlation: a test for spuriousness. psychological bulletin, 82, 887-903. doi: 10.1037/0033-2909.82.6.887 kintsch, w. (1998). comprehension: a paradigm for cognition. new york, ny, us: cambridge university press. kletzien, s. b. (1991). strategy use by good and poor comprehenders reading expository text of differing levels. reading research quarterly, 26, 67-86. doi: 10.2307/747732 warner&et &al & & & | f l r ! ! 18! kolić-vehovec, s., zubković, b. r., & pahljina-reinić, r. (2014). development of metacognitive knowledge of reading strategies and attitudes toward reading in early adolescence: the effect on reading comprehension. psihologijske teme, 23, 77-98. doi: 159.946.4.072-053.6 kolić-vehovec, s., & bajsanski, i. (2001). children's metacognition as predictor of reading comprehension at different developmental levels. in g. shiel & u. ni dhalaigh (eds.), other ways of seeing: diversity in language and literacy. dublin, ireland: readin association of ireland. lenhard, w., & schneider, w. (2006). leseverständnistest für erstbis sechstklässler [reading comprehension test for 1stto 6th-graders]. göttingen, germany: hogrefe. li, w.-d., fay, d., frese, m., harms, p. d., & gao, x. y. (2014). reciprocal relationship between proactive personality and work characteristics: a latent change score approach. journal of applied psychology, 99, 948-965. doi: 10.1037/a0036169 little, t. d., preacher, k. j., selig, j. p., & card, n. a. (2007). new developments in latent variable panel analyses of longitudinal data. international journal of behavioral development, 31, 357-365. doi: 10.1177/0165025407077757 lompscher, j. (1995). erfassung von lernstrategien mittels fragebogen [questionnaire for the assessment of learning strategies]. lern-und lehrforschung, 10, 80-136. massey, d. d. (2009). self-regulated comprehension. in s. e. israel & g. g. duffy (eds.), handbook of research on reading comprehension (pp. 389-399). new york, ny, us: routledge. matsunaga, m. (2008). item parceling in structural equation modeling: a primer. communication methods and measures, 2, 260-293. doi:10.1080/19312450802458935 mcclelland, m. m., & cameron, c. e. (2011). self!regulation and academic achievement in elementary school children. new directions for child and adolescent development, 2011, 29-44. doi: 10.1002/cd.302. mckeachie, w. j. (1987). cognitive skills and their transfer: discussion. international journal of educational research, 11, 707-712. doi: 10.1016/0883-0355(87)90010-3 miller, p. h. (1990). the development of strategies of selective attention. in d. f. bjorklund & d. f. bjorklund (eds.), children's strategies: contemporary views of cognitive development. (pp. 157-184). hillsdale, england: lawrence erlbaum associates, inc. ministery of education of the federal state of brandenburg. (2004). rahmenlehrplan grundschule deutsch [curriculum framework for german at elementary school]. berlin, germany: wissenschaft und technik verlag. mokhtari, k., & reichard, c. a. (2002). assessing students' metacognitive awareness of reading strategies. journal of educational psychology, 94, 249-259. doi: 10.1037/0022-0663.94.2.249 muthén, l. k., & muthén, b. o. (1998-2012). mplus user's guide. (7 ed.). los angeles, ca, us: muthén & muthén. myers, m., & paris, s. g. (1978). children's metacognitive knowledge about reading. journal of educational psychology, 70, 680-690. doi:10.1037/0022-0663.70.5.680 newsom, j. t. (2015). longitudinal structural equation modeling: a comprehensive introduction. new york, ny, us: routledge/taylor & francis group. paris, s. g., hamilton, e. e., israel, s., & duffy, g. (2009). the development of children’s reading comprehension. in s. e. israel & g. g. duffy (eds.), handbook of research on reading comprehension (pp. 32-53). new york, ny, us: routledge. paris, s. g., lipson, m. y., & wixson, k. k. (1983). becoming a strategic reader. contemporary educational psychology, 8, 293-316. doi: 10.1016/0361-476x(83)90018-8 paris, s. g., & paris, a. h. (2001). classroom applications of research on self-regulated learning. educational psychologist, 36, 89-101. doi: 10.1207/s15326985ep3602_4 petermann, f., & petermann, u. (2008). hamburg-wechsler-intelligenztest für kinder iv [wechsler intelligence scale for children german version]. bern, switzerland: huber. pintrich, p. r. (1991). a manual for the use of the motivated strategies for learning questionnaire (mslq). ann arbor, mi, us: school of education, university of michigan. warner&et &al & & & | f l r ! ! 19! pintrich, p. r. (2002). the role of metacognitive knowledge in learning, teaching, and assessing. theory into practice, 41, 219-225. doi: 10.1207/s15430421tip4104_3 pressley, m., & afflerbach, p. (1995). verbal protocols of reading: the nature of constructively responsive reading. hillsdale, england: lawrence erlbaum associates. rupley, w., & willson, v. (1996). content, domain, and word knowledge: relationship to comprehension of narrative and expository text. reading and writing, 8, 419-432. doi: 10.1007/bf00404003 schermelleh-engel, k., moosbrugger, h., & müller, h. (2003). evaluating the fit of structural equation models: tests of significance and descriptive goodness-of-fit measures. methods of psychological research, 8, 23-74. schneider, w. (2010). metacognition and memory development in childhood and adolescence. in h. s. waters & w. schneider (eds.), metacognition, strategy use, and instruction (pp. 54-81). new york, ny, us: guilford press. schünemann, n., spörer, n., & brunstein, j. c. (2013). integrating self-regulation in whole-class reciprocal teaching: a moderator–mediator analysis of incremental effects on fifth graders’ reading comprehension. contemporary educational psychology, 38, 289-305. doi: 10.1016/j.cedpsych.2013.06.002 stefanou, c. r., perencevich, k. c., dicintio, m., & turner, j. c. (2004). supporting autonomy in the classroom: ways teachers encourage student decision making and ownership. educational psychologist, 39, 97-110. doi: 10.1207/s15326985ep3902_2 stutz, f., schaffner, e., & schiefele, u. (2016). relations among reading motivation, reading amount, and reading comprehension in the early elementary grades. learning and individual differences, 45, 101113. doi: 10.1016/j.lindif.2015.11.022 turner, j. c. (1995). the influence of classroom contexts on young children's motivation for literacy. reading research quarterly, 30, 410-441. doi: 10.2307/747624 van de schoot, r., lugtig, p., & hox, j. (2012). a checklist for testing measurement invariance. european journal of developmental psychology, 9, 486-492. doi: 10.1080/17405629.2012.686740 van kraayenoord, c. e. (2010). the role of metacognition in reading comprehension. in h. trolldenier, w. lenhard & p. marx (eds.), brennpunkte der gedächntisforschung. entwicklungsund pädagogischpsychologische perspektiven (pp. 277-304). göttingen, germany: hogrefe. van kraayenoord, c. e., & schneider, w. (1999). reading achievement, metacognition, reading self-concept and interest: a study of german students in grades 3 and 4. european journal of psychology of education, 14, 305-324. doi: 10.1007/bf03173117 vandenberg, r. j., & lance, c. e. (2000). a review and synthesis of the measurement invariance literature: suggestions, practices, and recommendations for organizational research. organizational research methods, 3, 4-69. doi: 10.1177/109442810031002 warner, g. j., fay, d., schiefele, u., stutz, f., & wollny, a. (2016). being proactive when reading: academic personal initiative as a predictor of word comprehension development. manuscript submitted for publication. weinstein, c. e., & mayer, r. e. (1986). the teaching of learning strategies. in w. m. c. (ed.), handbook of research on teaching (pp. 315-327). new york, ny, us: macmillan. wernke, s., wagener, u., anschuetz, a., & moschner, b. (2011). assessing cognitive and metacognitive learning strategies in school children: construct validity and arising questions. international journal of research and review, 6, 19-38. wernke, s. (2012). handlungsnahe erfassung von lernstrategien mit fragebögen – eine empirische untersuchung mit kindern im grundschulalter [assessing learning strategies with questionnaires an empirical investigation with elementary school children]. oldenburg, germany: uni oldenburg. wild, k.-p., & schiefele, u. (1994). lernstrategien im studium: ergebnisse zur faktorenstruktur und reliabilität eines neuen fragebogens [learning strategies of university students: factor structure and reliability of a new questionnaire]. zeitschrift für differentielle und diagnostische psychologie, 15, 185-200. warner&et &al & & & | f l r ! ! 20! wollny, a. (2015). eigeninitiative in der kindheit und ihre bedeutung für die entwicklung der lesekompetenz [personal initiative in childhood and its importance for the development of reading compentence] (unpublished doctoral dissertation). university of potsdam, potsdam, germany. wollny, a., fay, d., & urbach, t. (2016). personal initiative in middle childhood. learning and individual differences, 49, 59-73. doi: 10.1016/j.lindif.2016.05.004 zimmerman, b. j. (2002). becoming a self-regulated learner: an overview. theory into practice, 41, 64-70. doi: 10.1207/s15430421tip4102_2 warner&et &al & & & | f l r ! ! 21! appendix table a.1 fitness of measurement models to test measurement invariance across time and groups construct χ2 (df) cfi tli rmsea srmr comparison δχ2 (df) δcfi δrmsea δsrmr personal initiative across time model 1: configural 256.222*** (85) .971 .959 .048 .035 ---- model 2: metric 266.836*** (92) .970 .961 .046 .037 model 1 vs. model 2 6.66 (7) -0.001 -0.002 0.002 model 3: scalar 312.954*** (99) .964 .956 .050 .039 model 2 vs. model 3 48.30*** (7) -0.006 0.004 0.002 model 4: strict 329.192*** (111) .963 .960 .047 .041 model 3 vs. model 4 16.30 (12) -0.001 -0.003 0.003 across groups model 1: configural ---------- model 2: metric 473.509*** (229) .959 .957 .049 .052 model 1 vs. model 2 ---- model 3: scalar 675.932*** (236) .927 .925 .065 .108 model 2 vs. model 3 190.43*** (7) -0.032 0.016 0.056 model 4: strict 683.373*** (248) .927 .930 .063 .108 model 3 vs. model 4 9.68 (12) 0.000 -0.002 0.000 strategy knowledgea model 1: configural 139.060*** (54) .969 .959 .056 ------ model 2: metric 129.243*** (62) .976 .972 .046 -model 1 vs. model 2 7.09 (8) 0.007 -0.010 - model 3: scalar 141.130*** (70) .974 .973 .045 -model 2 vs. model 3 6.56 (8) -0.002 -0.001 - model 4: strict 146.545*** (79) .975 .978 .041 -model 3 vs. model 4 9.97 (9) 0.001 -0.004 - (continued)! warner&et &al & & & | f l r ! ! 22! table a.1 (continued) construct χ2 (df) cfi tli rmsea srmr comparison δχ2 (df) δcfi δrmsea δsrmr reading comprehension model 1: configural 0.224 (2) 1.000 1.007 .000 .002 ----- model 2: metric 5.616 (6) 1.000 1.001 .000 .025 model 1 vs. model 2 5.04 (4) 0.000 0.000 0.023 model 3: scalar 6.496 (8) 1.000 1.002 .000 .029 model 2 vs. model 3 1.06 (2) 0.000 0.000 0.004 model 4: strict 15.781 (12) .997 .997 .024 .051 model 3 vs. model 4 9.71* (4) -0.003 0.024 0.022 note. written in bold are values of δrmsea, δcfi and δsrmr that were above the cut-off values recommended by chen (2007). higher level of measurement invariance is rejected if δcfi ≥ -.010, supplemented by δrmsea ≥ .015 or δsrmr ≥ -.010. awhen using the wlsmv estimator, mplus does not provide values for srmr, and chi-square differences can only be computed with the difftest option. *p < .05, **p < .01. ***p < .001. warner&et &al & & & | f l r ! ! 23! table a.2 estimated effects of control variables group 1 group 2 group differences path (from …) β 95% ci β 95% ci δχ² gendera ! t2 personal initiative -.11** -0.181 -0.044 -.11** -0.174 -0.042 3.62 t1 cognitive ability ! t2 personal initiativeb -.05 -0.042 0.148 -.13* -0.234 -0.015 12.77*** ses ! t2 personal initiativeb .02 -0.090 0.133 -.00 -0.098 0.091 4.20* gendera ! t2 strategy knowledge -.02 -0.081 0.039 -.02 -0.090 0.043 2.27 t1 cognitive ability ! t2 strategy knowledgeb -.02 -0.125 0.095 -.07 -0.187 0.043 4.17* ses ! t2 strategy knowledge -.02 -0.089 0.049 -.02 -0.104 0.057 2.25 gendera ! t2 reading comprehensionb .08 -0.021 0.181 .15** 0.045 0.253 7.18** t1 cognitive ability ! t2 reading comprehensionb -.11* -0.210 -0.011 -.06 -0.170 0.053 11.89*** ses ! t2 reading comprehensionb .16* 0.032 0.290 .21*** 0.099 0.310 12.64*** gendera ! t1 personal initiative -.10** -0.172 -0.034 -.10** -0.171 -0.030 2.57 ses ! t1 personal initiative .15*** 0.080 0.220 .15*** 0.077 0.229 1.36 gendera ! t1 reading comprehensionb -.13* -0.228 -0.022 -.04 -0.143 0.072 4.41* ses ! t1 reading comprehensionb .28*** 0.186 0.381 .34*** 0.267 0.421 4.90* note. ses = socioeconomic status. afemale = 1, male = 2. bpath was freely estimated because of a significant δχ² difference between groups. all other paths were fixed to equality. *p < .05. **p < .01. ***p < .001. endedijk et hoogeboom et al frontline learning research vol.6 no. 3 (2018) 123 147 issn 2295-3159 using sensor technology to capture the structure and content of team interactions in medical emergency teams during stressful moments maaike endedijk*a, marcella hoogeboom*a, marleen groeniera, stijn de laata, jolien van sasa a university of twente, the netherlands *both authors contributed equally to this work article received 13 april 2018 / revised 11 novemner/ accepted 23 november/ available online 7 december abstract in healthcare, action teams are carrying out complex medical procedures in intense and unpredictable situations to save lives. previous research has shown that efficient communication, high-quality coordination, and coping with stress are particularly essential for high performance. however, precisely and objectively capturing these team interactions during stressful moments remains a challenge. in this study, we used a multimodal design to capture the structure and content of team interactions of medical teams at moments of high arousal during a simulated crisis situation. sociometric badges were used to measure the structure of team interactions, including speaking time, overlapping speech and conversational imbalance. video coding was used to reveal the content of the team interactions. furthermore, the empatica e4 was used to unobtrusively measure the team leader’s skin conductance to identify moments of high arousal. in total, 21 four-person teams of technical medicine students in the netherlands were monitored in a simulation environment while they diagnosed and managed a patient with cardiac arrest. outcomes of this exploratory study revealed that more effective teams showed greater conversational imbalance than less effective teams, but during moments of high arousal the opposite was found. also, a number of differences were found for the content of team interaction. combining sensor technology with traditional measures can enhance our understanding of the complex interaction processes underlying effective team performance, but technological advances together with more knowledge about the simultaneous application of these methods are needed to tap into the full potential of wearable sensor technology in team research. keywords: : team interaction; video observation; skin conductance; sociometric badges; medical simulation; action teams. info. mail corresponding authors: a.m.g.m.hoogeboom@utwente.nl and m.d.endedijk@utwente.nl . doi: doi: https://doi.org/10.14786/flr.v6i3.353 1. introduction teams are ubiquitous in organizations. since the end of the 20th century the focus of work in organizations has shifted from the individual employee to employees as part of a team (kozlowski & ilgen, 2006). alongside this shift we have seen an increase in research focusing on the collaborative processes and outcomes of different types of teams (vangrieken, boon, dochy, & kyndt, 2017). one specific form of teams is an action team, “…where members with specialized skills must improvise and coordinate their actions in intense, unpredictable situations” (edmonson, 2003, p. 1421). in other words, it is the task of action teams to quickly establish effective coordination and communication in unexpected situations. as negative outcomes of these team processes can be detrimental for human safety (e.g., in medical teams or aviation teams), it is of utmost importance that these teams are trained well in performing these complex team interaction skills in a realistic environment. simulation rooms provide excellent, risk-free opportunities to practice both technical and team interaction skills in realistic scenarios that allow team members to experience how they will perform during stressful moments (kneebone, nestel, vincent, & darzi, 2007). not only research but also practice can benefit from exploring team interaction processes during such stressful moments during scenario-based training (entin & serfaty, 1999; lei, waller, hagen, & kaplan, 2016). traditionally, debriefing sessions with expert debriefers are used to provide feedback in simulation-based learning (fanning & gaba, 2007). during these sessions, expert debriefers select specific, observable events in the scenario and stimulate trainees to reflect on their behavior and decisions. the use of expert debriefers to provide feedback is not only very costly and timeand labor-intensive, but research has also shown that how debriefers facilitate the debriefing sessions is highly variable (tannenbaum & cerasoli, 2013). in addition, the events they select are not necessarily the moments team members experience as stressful. we argue that a next step is needed to move from traditional human-based observation methods to methods that allow for more objective and timely identification of effective team interaction processes during stressful moments. wearable sensors have opened up a new world of research possibilities to detect body signals and analyze speech from team interactions, providing insights into how people respond and interact without interfering with their natural work processes (fischer & järvelä, 2014). for example, sociometric badges (e.g., olguin et al., 2009; pentland, 2012) are sensors that are worn around the neck, similar to typical id-badges, and are able to detect various features of social interaction. the empatica e4-wristband [empatica ins, cambridge, usa] (garbarino, lai, tognetti, picard, & bender, 2014) is able to measure skin conductance (also known as electrodermal activity: boucsein, 2012) in an unobtrusive way. this can be used as an indicator for identifying moments of high arousal (boucheix, 2017; christopoulos, uy, & yap, 2016), which typically reflect high levels of distress in the context of medical action teams during a crisis situation (hunziker, johansson, et al., 2011). when combined, these sensors enable detailed exploration of social interaction in teams during moments of high arousal in action teams. however, a lot is still unknown about how to use and combine these sensors with more traditional measures to detect effective team interaction processes during moments of high arousal, especially for action teams. therefore, in line with the purpose of this special issue, the goal of this study is to clarify the methodological approach, added value and pitfalls of using and combining different sensor technologies in combination with video observation. our study was performed in a medical simulation room for advanced life support (als) training. als is a complex emergency situation following cardiac arrest of a patient and is characterized by “extreme time pressures, diagnostic uncertainty, and rapidly evolving situations” (doumouras, keshet, nathens, ahmed, & hicks, 2012, p. 274; hunziker, laschinger, et al., 2011). research has shown that human factors such as efficient team communication, coordination, and stress especially affect als efficiency and performance (fernandez castelao, russo, riethmüller, & boos, 2013; hunziker, laschinger, et al., 2011). some have estimated that poor non-technical skills can contribute to 64 to 83% of critical incidents in a medical context or crisis situation, for example, in anesthesia (arnstein, 1997). in this study, we combine traditional video observation methodology to systematically analyze the content of team interaction behaviors with innovative sensor technology to explore the structure of team interaction processes in more detail (sociometric badges) and identify moments of high arousal (empatica e4). the outcomes of this study demonstrate how using a combination of different sensors and traditional measures provides a means to get a rich picture of complex team interactions during moments of high arousal. these insights advance our knowledge of how to use and combine sensor technology and how this information can be used for optimization of simulation environments to help prospective and current medical professionals to improve not only their medical skills, but also their team interaction skills. 2. theoretical framework 2.1 simulation-based medical education simulation in medical education is used as an educational technique to improve health outcomes, dating back to the 17th century (cooke, irby, & o’brien, 2010; mcgaghie, issenberg, petrusa, & scalese, 2009). the number of medical simulation settings has expanded with the development of complex technologies. such advanced technologies enable simulations that closely resemble reality, especially when combining them with high-fidelity scenarios built around events that potentially could have serious consequences for the patient (dias & neto, 2016; grenvik, schafer, devita, & rogers, 2004). skills acquired in a well-designed medical simulation environment show better transfer to improved real-life patient care compared to traditional, on-the-job medical training (mcgaghie, issenberg, cohen, barsuk, & wayne, 2011). simulation is especially useful in training under conditions of uncertainty, ambiguity and rapid situation changes with potentially severe consequences for patient safety (satish & streufert, 2002). high-fidelity simulation-based learning is nowadays the standard for training als teams who must diagnose and manage a patient in cardiac arrest (sahu & lata, 2010). diagnosing and managing a patient in cardiac arrest requires immediate medical intervention and efficient teamwork; otherwise, a patient might not survive (hunziker, johansson, et al., 2011). successful resuscitation depends on the integrated application of technical skills, such as intubation, chest compressions, and clinical reasoning, and non-technical skills related to working in a team, such as communication, decision-making, and leadership (hunziker, tschan, semmer, howell, & marsch, 2010). effective and efficient teamwork in a resuscitation scenario requires a sequence of actions that is performed in the correct way and at the right time (hunziker, johansson, et al., 2011). 2.2 team interaction a fundamental aspect that can lead to high team performance in a crisis situation, such as a cardiac arrest, are the emergent, interactive processes between medical personnel (e.g., hunziker, johansson, et al., 2011). team interaction is defined as a series of ongoing behavioral processes and actions that occur over time (lei et al., 2016; stachowski, kaplan, & waller, 2009). for decades, team researchers have been advocating to capture the dynamic nature of such team interactions, as opposed to using static measures (marks, zaccaro, & mathieu, 2000). there are several ways to quantify verbal team interactions. some researchers focus on the content of the interaction, for example, by using observation schema that identify what team members say, such as the number of agreements, suggestions, or opinions in team conversations (e.g., atwal & caldwell, 2005; hoogeboom & wilderom, 2015). others focus on the structure of the conversation, such as which team member is speaking, the number of interruptions or the degree of turn-taking, regardless of the content (kim, mcfee, olguin olguin, waber, & pentland, 2012; koudenburg, postmes, & gordijn, 2017; pugliese, nicholson, & bezemer, 2015). for example, when providing teams real-time feedback on conversational balance (i.e. over-participators were stimulated to decrease their participation), this influenced their decision making for either better or worse (dimicco, hollenbach, pandolfo, & bender, 2007). in the current study, both the content and structure of team interactions will be taken into account. previous studies have already identified some content and structure related characteristics of effective team interaction in action teams which are dealing with crisis situations. for example, using video observation and coding, previous studies on the structure of interactions of airline teams have shown that effective teams displayed less complex, more homogenous interaction processes (kanki, folk, & irwin, 1991; zijlstra, waller, & phillips, 2012). other related research on the structure of interactions in nuclear power plant control room crews showed that effective interactions during crises consisted of fewer actors and less back-and-forth communication (stachowski et al., 2009). hence, shorter, less complex, and less reciprocal team interaction, indicating somewhat scripted or standardized forms of team interaction, seem to be more effective in action teams. regarding the content of interactions, kolbe et al. (2014) showed that effective medical team members more frequently spoke up and aided assistance after implicit action coordination (i.e. team monitoring). such task-related helping interactions after team monitoring behavior seemed vital for high performance. for the team leader, communicating clear goals and a clear task distribution has proven to reduce the emotional reactions by team members, leading to an increase in performance in stressful situations (zaccaro et al., 2001; andersen, jensen, lippert, & østergaard, 2010; marsch et al., 2004). moreover, research in emergency command-and-control teams has shown the importance of leader structuring behavior, such as clarifying and summarizing (van der haar et al., 2017). during resuscitation, the use of closed-loop-communication is advocated to avoid errors (fernandez castelao et al., 2013). this clear, structured, and standardized form of communication (brindley & reynolds, 2011; härgestam, lindkvist, brulin, jacobsson, & hultin, 2013) consists of an initial message (call-out) by the team leader, which should be confirmed or acknowledged by the receiver (check back) and confirmed once again by the team leader (closing the loop) (davis et al., 2017; härgestam et al., 2013; jacobsson, hargestam, hultin, & brulin, 2012; schmutz, hoffmann, heimberg, & manser, 2015). hence, in general, various studies from different fields suggest that effective and less effective action teams differ in both the content and structure of their team interaction. to enhance the learning opportunities in simulation environments, it is important to delineate the effective team interaction processes that are required for high team performance (goldman, 2014). however, only a small but growing number of studies have investigated in detail how teams interact in the daily context of their work and how this contributes to achieving their goals (humphrey & aime, 2015). nowadays, technical and methodological advances allow us to capture team interactions more precisely (molenaar, 2014). one of the available, but to date less frequently used, devices is the sociometric badge (kim et al., 2012): a wearable that includes several types of technology, namely bluetooth, an infrared sensor, an accelerometer and a microphone. either when used in isolation or when combined with other sensors, the badges allow for fine-grained analysis of the structure of verbal team interactions. for example, when combined with sensors that capture electrodermal activity, the structure of team interactions in higharousal moments can be compared to moments of low arousal. 2.3 moments of high arousal: definitions and effects identifying moments of high arousal during a medical simulation session can inform us about what happens during moments when a person is not able to process the mental effort or cognitive load required or when the perceived demands of the environment exceed a person’s ability to cope with these demands (berntson & cacioppo, 2000; stemmler, 2004; boucheix, 2017; lazarus & folkman, 1984). for example, in the context of medical simulations, higher levels of arousal may ensue when time pressure forces an als team member to act quickly. it should be noted that moments of high arousal can also automatically occur due to positive events, such as positive workplace interaction or excitement (heaphy & dutton, 2008). hence, high arousal is not only attributable to distress, but also to excitement (russell, 1980). in other words, when physiological measures of arousal are used, information is obtained about the intensity of physiological arousal, but not about the psychological state or valence (e.g. distress or excitement) associated with it (e.g., akinola, 2010). however, even though no general inferences about valence can be drawn from the intensity of arousal (e.g., boucsein, 2012; larsen, diener, & lucas, 2002), a previous study on self-reported emotions during resuscitation performance has shown that negative emotions (stress or overload) are significantly higher during resuscitation, while positive emotions were highest before resuscitation, decreased during resuscitation, and increased again when the simulated patient was awake again (hunziker, laschinger, et al., 2011). therefore, in our study we interpret high levels of arousal as an indicator of feelings of distress. heart rate and skin conductance, physiological responses of the body, are examples of markers of autonomic activity of the nervous system, and concomitants of arousal (akinola, 2010; benedek & kaernbach, 2010). particularly during social interaction, skin conductance has been found to be the most sensitive indictor of emotional responsiveness or arousal, as opposed to the other physiological markers (marci, ham, moran, & orr, 2007). this physiological measure captures the intensity of emotions during interactions with others (akinola, 2010; figner & murphy, 2011). skin conductance is defined as variations in the eccrine sweat glands (i.e., sweat glands which are present in all bodily parts, with the highest density in the palms and soles (boucsein, 2012): in response to sweat secretion from the skin (e.g., benedek & kaernbach, 2010). the popularity of using skin conductance measures is due to its direct relation with a stimulus or emotional response, which enables us to capture moments of high arousal, as long as the temperature in the environment is kept constant (boucsein, 2012; lang, bradley, & cuthbert, 1998). previous studies have examined the effect of stress on als performance, but with conflicting results. on the one hand, hunziker, laschinger et al. (2011) and hunziker, semmer et al. (2012) found that perceived stress during early resuscitation negatively influenced als performance. in the latter study, also physiological measures as indicators for stress were used, but no association with team performance was found, possible due to the fact that the team members were engaged in physical activity what distorted the physiological measurements. in a similar vein, the study of sandroni, fenici, et al. (2005) did not find a direct relation between physiological stress measures and individual performance during the als scenario as measured by a written multiple-choice test. other studies which examined the relation between physiological measures of stress and performance in other medical simulation settings found a u-shaped association (also known as the yerkes-dodson law, cf. cohen, 2011): positive relationships between arousal and performance have been reported, whereas extreme levels of arousal were detrimental for performance (e.g., keitel et al., 2011; wetzel et al., 2010). to understand these conflicting findings, we suggest that exploring more in-depth how members interact during moments of high arousal is an interesting endeavor, as this might better explain performance than solely the level of arousal. 3 the present study in this paper, we focus on how we can use the combination of sociometric data, physiological data, and video data to explore if more effective teams – compared to less effective teams alter the content and structure of their team interactions during moments of high arousal. we expect that sensor technology will have added value in addition to self-report measures to objectively identify moments of high arousal and analyze the structure of the team interactions. as these measures are relatively new to the field of social science and specifically educational science, this exploratory study sets out to uncover what the added value, pitfalls and hurdles are when using and combining different sensor technologies with more traditional measures in order to better understand team interactions. to better understand the added value of sensor technology, we have translated the aim of this study into the following research question: how do more versus less effective medical action teams differ in content and structure of team interactions during scenario-based training, and how do their interactions differ during moments of high arousal versus outside these moments of high arousal? we will answer the research question by step by step testing the differences in content and structure of team interactions between 1) moments of high arousal and non-high arousal moments (without taking effectiveness into account); 2) more and less effective teams (without taking arousal into account); and 3) the combination of both: differences between high arousal versus non-high arousal moments separately for more and less effective teams. 4 methodology 4.1 participants and design all 95 first-year master’s students (comprising 24 teams) who enrolled in the course ‘advanced life support (als)’ in the master study program ‘technical medicine’ at university of twente were invited to participate in the study. ninety-two students gave written informed consent to participate in the study. the three students who did not give consent and their team members were excluded from the study, resulting in a data set of 21 four-person teams and one three-person team. to avoid an unequal situation for this three-person team, one person from another team was added to this team during the assessment. to ensure comparability across teams, this three-person team was left out from the analysis (n = 84). on average, the students were 22.4 years old (sd = 1.1) and 44% were male. during the als course, students learned to diagnose and manage a patient with cardiac arrest and perform cardiopulmonary resuscitation (cpr) on an advanced human patient simulator in scenarios of varying complexity. a multimethod design was adopted which included four different sources of data: (1) video coding of the content of team interactions, (2) sociometric measurement to capture the structure of team interactions, (3) the empatica e4-wristband to capture skin conductance/arousal, and (4) teacher surveys to assess team effectiveness. 4.2 procedure prior to data collection, the study was approved by the ethical committee of the university as well as by the teachers involved in the als course. during the first introductory lecture of the course, the students were informed about the study. the data were collected during the final assessment of the course. all teams were assessed on the same day. two rooms were used for the als simulation scenarios, both had a regulated temperature with 0.1-degree celsius temperature tolerance. the temperature for both rooms was kept constant at 20.5 degrees celsius. before the start of the scenario, all four students of each team were randomly assigned to one of the four fixed team roles: team leader, responsible for task distribution, monitoring team performance, creating an overview of the situation, and patient handover at the end of the scenario; medication nurse, responsible for drug administrations and connecting devices to the patient; and two cpr administrators, responsible for chest compressions and airway management. all students had practiced each role at least one time during the course, so they knew what was expected from them at the assessment. in the als context the physical activities that are required by team members in the role of medication nurse and cpr administrator (such as intubation or chest compressions) influence their physiological arousal (i.e., producing distorted or biased physiological measurement: berntson & cacioppo, 2000; stemmler, 2004). the team leader in the als situation is less physically active, so we decided to measure physiological arousal of the team leader to identify moments of high arousal. the researchers distributed a sociometric badge to all team members and secured the empatica e4-wristband on the non-dominant wrist of the student in the role of team leader. as hunziker, laschinger, et al. (2011) described in their paper on stress and team performance during a simulated resuscitation, simulated advanced life support (als) scenarios usually follow a specific pattern. similar to the hunziker study, the als scenarios started with a short briefing about the patient’s history by one of the teachers (maximum 90 seconds), immediately followed by a resuscitation period during which cpr is performed. the scenarios ended with a handover of the patient to another team or specialist. however, contrary to the hunziker study, in the majority of the scenarios in our study the patient was still in a critical condition when the scenario was completed; the patient could breathe on its own, but did not have a steady pulse or sinus rhythm. given that the patient still was in a critical condition, the patient handover at the end is a crucial phase in the scenario. the duration of the scenarios was on average 22.00 minutes (sd = 4.83). previous studies have shown that team behaviors and interactions change as the scenario progresses, with a stronger focus on leadership and coordination skills at the beginning of a scenario (tschan et al., 2006, 2014), and agreement on a shared diagnosis and treatment plan, as well as accurate handover at the end of the scenario. especially in assessment settings the end is perceived as a stressful part of the scenario (sandroni et al., 2005). therefore, although we recorded the complete scenarios, for the analysis we focused on the beginning and end of the scenario, which we defined as the first 16.7% minutes (1st time window) and last 16.7% minutes (2nd time window). the duration of these 16.7%-windows was on average 3 minutes and 40 seconds (sd = 49.6 s) and varied from 2 minutes and 24 seconds for the shortest video to 4 minutes and 52 seconds for the longest video. 4.3 instruments 4.3.1 sociometric badges sociometric badges were used to assess the structure of team interactions (kim et al., 2012). sociometric badges are wearables distributed by humanyze (a spin-out of the mit media lab) that measure proximity, body movement and speech features using respectively bluetooth, infrared, an accelerometer and microphones. data were uploaded from the badges using sociometric datalab research edition (version 3.1.3029) and subsequently exported to excel files using the ‘structured meeting’ setting which disregards the bluetooth and infrared data in order to provide a dataset in which each team member is assumed to be in each other’s proximity during the entire session, which was the case in our setting. also, the setting ‘noisy environment’ was used, which filters out additional environment noise, such as the beeping of the heart monitor or the speech from teachers not wearing a badge. the exported microphone data shows per second whether a participant was talking or silent. from this data, the following four metrics were calculated at the team level: 1) the proportion of the time that one or more of the team members was speaking (i.e., proportion of speaking time), 2) the proportion of the speaking time when at least two team members were speaking at the same time (i.e., proportion of overlapping speech), and 3) the distribution of how each team member contributed to the overall speaking time (i.e., conversational imbalance in speech). the conversational imbalance was calculated as the standard deviation of the amount of time that each team member spoke (regardless of whether another member was speaking at the same time), corrected for the duration of the relevant time window. the higher this standard deviation, the greater the variation in speech among members, that is, the greater the conversational imbalance. in addition, also 4) the proportion of speaking time of the team leader was calculated, measured as the proportion of the time the team leader was speaking, regardless of whether another member was speaking at the same time. all measures were calculated using r (r core team, 2014). 4.3.2 empatica e4-wristband the empatica e4 (hereafter referred to as e4), a relatively unobtrusive wristband, including 8 mm silver-plated electrodes, was used for the skin conductance recording. continuous measures of physiological arousal at a sample rate of 4 hz were collected using this device. the wristband was placed on the wrist of the team leader’s non-dominant hand. continuous decomposition analysis was used to get the relevant phasic electrodermal activity parameters : skin conductance responses (i.e., the number of peaks for certain periods of time which represents the fast-varying phasic component of skin conductance; the amplitude threshold for the extraction of skin conductance responses was .01 micro siemens (µs)) and amplitude of the skin conductance responses (i.e., the height of a single skin conductance response). using this analysis, the classical trough-to-peak parameters such as number of skin conductance responses and amplitude of the skin conductance responses can be obtained (benedek & kaernbach, 2010). to detect the moments of high arousal, skin conductance amplitudes for each individual team leader were computed (hamaker, 2012; wiesenfeld, whitman, & malatesta, 1984). information about peak amplitude of individual skin conductance responses can be used to detect moments of high arousal (bach, flandin, friston, & dolan, 2009). the highest skin conductance amplitude in the beginning and end of the scenarios was determined for each team leader. on the basis of the highest skin conductance amplitude a single segment of 30 seconds was selected both in the beginning and the ending of the scenario. this resulted in two 30-second segments for each team: one in the beginning (first 16.7%) and one in the end (last 16.7%) of the scenario . the 30-second segment was based on the thin-slices theory of social interaction (curhan & pentland, 2007). because several studies have reported an average delay of one to four seconds between a stimulus and a skin conductance response (e.g., dawson et al., 2007; weis & herbert, 2017) we used a segment starting 5 seconds before and 25 seconds after the detected highest amplitude. the highest amplitude peak (in micro siemens, µs) of each team leader in both the beginning (m = 2.34, sd = 2.67) and end (m = 3.63, sd = 3.26) was compared with the mean amplitude of the beginning (m = 1.65, sd = 2.01) and end (m = 2.76, sd = 2.65). this indicated that both in the beginning (t (15) = 3.19, p < .01) and in the end (t (15) = 4.67, p < .01) the highest amplitude peak was significantly higher than the mean amplitude in that segment. 4.3.3 video the content of the team interactions was measured via video recordings. the video cameras were ceiling-mounted, fixed cameras that minimized obtrusiveness and reactivity of the team members. the content of the team interactions was coded for the first and last 16.7% minutes of each scenario. the codebook was specifically designed to capture behaviors that occur frequently in action teams and is rooted in earlier theorizing on teams and action teams (lei et al., 2016; stachowski et al., 2009; zijlstra et al., 2012). additionally, on basis of the theory described in the theoretical framework, two items were added in order to code closed-loop-communication, more specifically: (1) check-back (by a team member), and (2) closing the loop (by the team leader); see table 1 for an overview of definitions and examples of the coded behaviors. a distinction was made between behaviors of the team leader, and of the other team members (followers). exhaustive coding was applied, meaning that all behaviors were coded and a non-observable category was used for behaviors that were not understandable or not relevant. the “observer xt” software program (noldus, trienes, hendriksen, jansen, & jansen, 2000; spiers, 2004) was used to code the videos. one of the coders parsed all sessions, that is: segmented the videos in speaker utterances (klonek, burba, kauffeld, & quera, 2016), using as a unit of analysis “a sentence or part of a compound sentence that can be regarded as meaningful in itself, regardless of the meaning of the coding categories” (strijbos, martens, prins, & jochems, 2006, p. 37). subsequently, all segmented scenarios were systematically coded by two independent, trained coders, who were not informed about the moments of high arousal. overall, an inter-rater agreement of 83.1% (cohen’s kappa = .80; cohen, 1960) was established. after coding the videos, the behavioral codes “not observable”, “external communication”, and the infrequent behaviors “laugh” and “apologies”, were merged into one category “external / other behavior”. the frequencies of the behaviors of a team are highly influenced by the total duration of that video. therefore, all coded behaviors were standardized according to the shortest video using the following formula: standardized frequency of a certain behavior of team x = coded frequency of the behavior of team x * (duration of the shortest video / duration of video team x). this resulted in a time-standardized behavior which enabled direct comparisons of the frequencies of the team members’ behaviors across the different teams. in addition, to enable comparison of the frequencies between the 30-second segment and the rest of the time window, the percentages of every leader and follower behavior were calculated relatively to all of the leader or follower behavior in the relevant time interval. table 1 examples of coded video behaviors. note. tl = team leader; f = follower. * after coding, these behaviors were merged into the category external / other behavior. 4.3.4 teacher ratings of performance to assess team effectiveness, four team effectiveness items by gibson, cooper, and conger (2009) were used: “this team is a consistently well performing team”, “this team is effective”, “this team makes few mistakes”, and “this team delivers high quality work”. the items were directly scored after the scenario on a likert scale from 1 (strongly disagree) to 7 (strongly agree) by two teachers. each teacher scored twelve teams. the internal consistency of the scale was high: cronbach’s alpha was .97. the teams were categorized as ‘more’ or ‘less’ effective on the basis of a median split, which was 5.75. the 10 less effective teams had a mean score of 4.28 (sd= .98), as rated by the teachers (on a scale of 1 to 7); the 12 more effective teams, including those with the median score, scored on average 6.23 (sd = .52) on team effectiveness. 4.4 analysis 4.4.1 selection of the data and dealing with missing data all recorded data were checked for missing values. video recordings and skin conductance data of all 21 teams were successfully obtained. mechanical issues with five sociometric badges resulted in the loss of the data of 20 participants among 13 different teams. this resulted in a total of 18 teams from which sociometric data of the team leader was intact, and nine teams from which complete data of the sociometric badges was available of all team members. the demographic data from these subsamples was comparable with the reported data from the 21 teams; no significant differences were found. 4.4.2 synchronization of the data to combine the data from the e4, sociometric badges and video observations, the data had to be synchronized. coded video behaviors, sociometric data and skin conductance amplitude data were synchronized on the basis of a mutual timeline. the internal clocks in the e4, sociometric badges, and video recording devices were used for synchronization purposes employing custom python code (represented by unix time: number of seconds from 1-1-1970 in coordinated universal time: utc). the data sources could only be synchronized if the clock time of the e4 biosensor and the timestamp of the video recording device were exactly aligned. on the basis of the python code, the clock times of the e4, video, and sociometric badges could not be matched for five teams. possibly, the clock time of the e4 was not synchronized (e.g., with a computer or laptop that had an accurate clock time); therefore, differences in the clock times of the video and sociometric badge on the one hand and the e4 on the other hand might be present. however, it should be noted that it is difficult to pinpoint the exact cause of the synchronization issue. due to these synchronization issues and the earlier mentioned malfunctioning of five sociometric badges, a total of 16 teams were available for further analysis of the video data combined with the skin conductance amplitude data; 13 teams from which the sociometric data of the team leader could be matched to the skin conductance measures and video data, and seven teams from which the sociometric data was available of all team members and could be linked to the skin conductance measures and video data. 4.4.3 analysis of relations between variables to answer the first research question of whether content and structure of team interactions were different for moments of high arousal versus non-high arousal, a series of repeated measures manovas were conducted. for each time window (beginning and end), a separate repeated measures manova was conducted for the dependent variables describing the content of team interaction (team leader and follower behavior) and the variables describing the structure (proportion speaking time, proportion overlapping speech, conversational imbalance). in addition, a dependent sample t-test was conducted to test for differences in proportion speaking time of the team leader. this last variable could not be grouped into one of the manovas due to a different sample size. in all of these analyses the level of arousal (high versus non-high) was used as the within-subject variable. the second research question about the differences between more and less effective teams was also answered with a series of manovas and independent sample t-tests with the same dependent variables, but this time the level of effectiveness (more versus less) was used as the between-subject variable. to answer the third research question, the data was split into two groups (more versus less effective teams) to explore differences between those two groups in team interaction during moments of high arousal and non-high arousal. due to the low sample size, we had to refrain from conducting manovas as this resulted in not enough residual degrees of freedom. therefore, separate independent t-tests were conducted to test for differences in structure and content of team interaction. 5. results 5.1 comparison of structure and content of team interactions between high arousal moments and rest of the time window the results of the comparison of structure and content of team interactions between high arousal moments and rest of the time window is displayed in table 2. as can be seen in table 2, team members were speaking for approximately half of the time, and during the end of the scenario almost two third of the time. the team leader was speaking for roughly a quarter of the time. the repeated measures manovas and dependent t-test showed no differences for the structure of team interactions on team level, both for the first (f(12, 1) = .54, p = .667) and second time window (f(12, 1) = .22, p = .878). also, no differences were found in proportion of speaking time of the leader (t (12) = .41, p = .686 first time window, t (12) = 1.08, p = .303 second time window). regarding the content of the team interaction, in the beginning of the scenario, team leader communication is characterized by many commands, while in the end of the scenario, when teams have to reach a diagnosis and have to handover the patient, there is more external communication. the repeated measures manova showed a significant effect of arousal on the content of team interaction, f(10, 6) = 17.72, p = .001 (first time window) and f(10, 6) = 4.30, p = .044 (second time window). further inspection of the univariate anovas showed significant differences between high arousal moments and the rest of the time window in the first time window for the team leader behaviors questioning (f(1, 15) = 45.52, p < .001) and opinion (f(1, 15) = 8.32, p = .011). this demonstrated that, on average, team leaders showed relatively less questioning (m = 0.0%) and opinion (m = 0.0%) behavior during moments of high arousal than during the rest of the first time window (m = 3.3% and m = 0.8%). for the second time window, the univariate anovas showed significant differences for suggestion and opinion, with f(1, 15) = 10.12, p = .006 and f(1, 15) = 6.96, p = .019 respectively. this indicated that, on average, team leaders showed relatively less suggesting (m = 2.8%) and opinion (m = 0.0%) behavior during moments of high arousal compared to the rest of the second time window (m = 8.1% and m = 1.9%). table 2 comparison in structure and content of team interactions between moments of high arousal and outside these moments. a standardized frequencies (see method section). * p < .05. **p < .01. 5.2 comparison in structure and content of team interactions between more and less effective teams to determine whether the structure and content of interactions differed between more and less effective teams, manovas (see table 3) were conducted to compare differences for the beginning (first time window) and the end (second time window). for both time windows, the manovas did not show significant effects of team effectiveness on the structure of team interactions (f(5, 1) = 2.67, p = .220 and f(5, 1) = 8.94, p = .052 respectively). however, as the analysis of the second time window was approaching significance, we further inspected the outcomes of the univariate anovas, which showed significant differences in the conversational imbalance between the more and less effective teams. both in the beginning (f(1, 5) = 11.94, p = .018) and in the end (f(1, 5) = 11.17, p = .020), the more effective teams showed greater imbalance (respectively m = 0.12 and m = 0.13) than the less effective teams (m = 0.06 for both time windows). in other words, in more effective teams one person was more dominant in terms of speaking time, while in less effective teams the team members contributed more equally; see table 3. for content of team interactions two manovas were conducted that showed no significant effect of team effectiveness on content of team interactions for both time windows (f(10, 5) = .58, p = .785 and f(10, 5) = .52, p = .822 for the first and second time window respectively). separate univariate anovas confirmed no significant effects table 3. comparison in structure and content of team interactions between more and less effective teams. a standardized frequencies (see method section). bn = 10 for more effective teams and n = 6 for less effective teams. cn = 3 for more effective teams and n = 4 for less effective teams. dn = 9 for more effective teams and n = 4 for less effective teams. * p < .05. **p < .01. 5.3 comparison between moments of high arousal and rest of the time window separately for more and less effective teams in addition to the overall differences between more and less effective teams, we were also interested to see whether more and less effective teams were different in how they changed their content and structure of the team interaction between moments of high arousal and outside these moments. therefore, separately for the less and more effective teams, we tested what the differences were between the 30-second segment and the rest of the corresponding time window using dependent t-tests. on average, more effective teams (see table 4a) were less imbalanced during moments of high arousal (m = .09) than during the rest of the second time window (m = .14) (t (2) = 5.6, p = .030). regarding the content of team interaction, in the more effective teams, team leader behavior was characterized by more commands, both during moments of high arousal (54%), and in the rest of the time window (m = 30.7%). during the moment of high arousal, team leaders showed relatively less confirmation (m = 1.7%), less questioning (m = 0.0%), less opinion (m = 0.0%) and less closing the loop (m = 3.1%) than during the rest of the first time window (respectively m = 9.5%, m = 3.0% and m = 0.8% and m = 7.2%; t (9) = -4.77, p = .001, t (9) = -5.17, p = .001 and t (9) = -2.33, p = .045 and t(9) = -4.08, p = .020). in the end of the scenario, the moment of high arousal was characterized by relatively less summary behavior (m = 1.4%) when compared to the rest of the second time window (m = 5.3%), t (9) = -3.36, p = .008. also, for less effective teams (see table 4b) some differences were found regarding the structure and content of team interaction during the high arousal moments, compared to rest of the corresponding time window. contrary to the more effective teams, we found that less effective teams, on average, were more imbalanced during moments of high arousal (m = .11) than during the rest of the window (m = .06), t (3) = 4.1, p = .026); yet, this difference was only significant for the first time window. moreover, in the less effective teams, team leader behavior was characterized by a high percentage of commands (m = 32.0% and m = 30.2%). in addition, in the beginning of the scenario, team leaders of less effective teams showed relatively less questioning (m = 0.0%) and inquiry (m = 0.0%) during the moments of high arousal compared to the rest of the first time window (respectively m = 3.7% and m = 3.3%; t (5) = -4.19, p = .009 and t (5) = -2.92, p = .033). in the end of the scenario, during the moment of high arousal also relatively less inquiry behavior (m = 0.0%) and less suggesting behavior (m = 3.3%) was exhibited compared to the rest of the second time window (respectively m = 5.9% and m = 10.4%; t (5) = -3.40, p = .019, and t (5) = -3.21, p = .024). in addition, a difference in follower behavior was found; in less effective teams; the followers showed relatively less check back behavior during moments of high arousal than during the rest of the time window (m = 15.6% vs. m = 25.4%, t (5) = 2.59, p = .049). table 4a. comparison in structure and content of team interactions between moments of high arousal and outside these moments within more effective teams. astandardised frequencies (see method section). *p < .05. **p < .01. table 4b. comparison in structure and content of team interactions between moments of high arousal and outside these moments within less effective teams. astandardised frequencies (see method section). *p < .05. **p < .01 6. discussion in our study, we combined video data and ratings of team effectiveness with skin conductance measures (empatica e4-wristband), and measures of speech features (sociometric badges) to analyze the structure and content of team interactions in medical action teams. the context of our study, a medical simulation room, was highly relevant for the aim of our study as this enabled us to closely observe medical action teams during a simulated crisis situation (als for a patient with cardiac arrest). by studying the differences between more and less effective teams in the content and structure of their interactions, both during moments of high arousal and outside these moments, we were not only able to add to the understanding of effective team interactions during crisis situations, but also of the added value of combining these various sensors with traditional data. in the following paragraphs, we discuss (1) the insights from the study on structure and content of team interactions of action teams during stressful moments, (2) the added value and challenges of combining information from various sensors to contribute to a more fine-grained understanding of complex team interactions, (3) the limitations of our study and (4) future directions for research using sensor technology to study team interactions. 6.1 discussion of findings understanding team interactions in als teams is crucial because als teams “need to be organized in such a way that the individual skills of the team members can be used efficiently and effectively” (cooper & wakelam, 1999; p. 27). each team member needs a clear understanding of “how decisions are made within the group; what resources are needed and how they are to be utilized; how leadership is exercised; and how staff new to the situation are integrated into the group.” (cooper & wakelam, 1999; p. 27). we analyzed the structure and content of team interactions of medical action teams during a simulated crisis situation and compared team interaction at moments of high arousal with team interaction outside these high arousal moments; as well as differences in team interaction between more and less effective teams. first, when we compared team interaction during moments of high arousal with team interaction outside these moments without taking effectiveness of teams into account, we found differences in the content of the team interaction only. in this specific context, it turned out that the team leader gave no opinions during moments of high arousal, both at the beginning and end of the scenario. moreover, the team leader asked no questions (in the beginning of the scenario) and made less suggestions (in the end). this is in line with previous research stating that during moments of crises fast coordination and clear decision making is needed (tschan et al., 2006). in other words, there is no room for behavior that needs further interpretation or clarification from team members. however, no significant differences were found that indicated which behavior was more frequently present during the moments of high arousal. second, we inspected the differences between more and less effective teams without distinguishing between the moments of high arousal and the rest of the corresponding time windows. overall, no main effects of effectiveness were found on the content and structure of team interactions. of course, we have to consider that we had a very small sample size and thus too low power to detect small differences. the only difference that we found was in the conversational imbalance; more effective teams showed greater imbalance than the less effective teams, both at the beginning and the end of the scenario. in other words, in more effective teams, one person was more dominant in terms of speaking time, while in less effective teams the speaking time was more equally distributed. although equal member contribution is generally seen as positive for high team effectiveness as all opinions can be taken into account (dimicco et al., 2007) this is different for medical action teams during cardiac arrest where swift decisions and clear and effective coordination of the team leader are necessary (andersen et al., 2010; marsch et al., 2004). our study points into the same direction, namely that a greater contribution of one person (greater imbalance) contributes to team performance. however, in extreme contexts effective leaders are receptive to the input of team members (hannah, uhl-bien, avolio, & cavarretta, 2009) and team members of effective medical teams more frequently speak up and provide help (kolbe et al., 2014). therefore, dominancy of the team leader should not be interpreted as no room for other team members to contribute. finally, we analyzed separately for more and less effective teams if there were differences in structure and content of team interaction when comparing moments of high arousal with the remaining time in the begin and end of the scenario. this revealed an interesting regarding the structure of the team interaction: when we did not split up between moments of high arousal and the rest of the corresponding time window, we saw more conversational imbalance for more effective teams. however, after splitting up between moments of high arousal and the remaining time of the window, it became clear that both more and less effective teams show differences in conversational imbalance during moments of high arousal, but in opposite direction. the four less effective teams showed more imbalance, i.e. greater dominancy of one person during moments of high arousal, and were more balanced outside these moments (in the beginning of the scenario). on the contrary, in the three more effective teams team members contributed more equally during the high arousal moment (in the end of the scenario), and they were more imbalanced in their contributions during the rest of the time window. an interesting hypothesis for future research is to test whether a certain degree of conversational imbalance indeed contributes to higher team performance in action teams, while more equal contributions are needed when the team leader gets into a state of high arousal. to better understand these outcomes, it is worth to have a look at the differences in the content of interaction. first, table 4a showed that during moments of high arousal more than half of the team leader communication of the more effective teams consisted of commands. in addition, significant differences were found in other behaviors that occurred less frequently during the moments of high arousal: conformation, questioning, opinion, and closing the loop. for less effective teams (table 4b), it seems that team leader behavior during moments of high arousal was more similar to their behavior outside these moments. differences in team leader behavior in less effective teams were only found in the behaviors inquiry and questioning. in addition, the less effective teams showed differences in follower behavior, namely less check back behavior, which is a crucial step in closed-loop-communication to confirm an initial message from the team leader and to avoid mistakes and misunderstandings (davis et al., 2017; härgestam et al., 2013; jacobsson et al., 2012; schmutz et al., 2015). it could thus be the case that the finding that the team leader is more dominant during stressful moments in less effective teams, is related to less check back behavior from their followers: the team leader has to repeat or rephrase often because the followers do not confirm (check back) his or her commands which leads to a relatively greater contribution of the team leader. again, we have to take the low sample size into account here while interpreting the findings, but what our results in general show is that a more fine-grained analysis of high arousal moments of a scenario enhances our understanding of what is effective team interaction during stressful moments. in addition, simultaneously exploring the structure and the content of the interactions can provide a more in-depth understanding of what constitutes effective team behavior. 6.2 added value and challenges of using sensor technology to gain insights in team interactions 6.2.1 added value and challenges of using the empatica e4-wristband in our study, we used the e4 to detect of moments of high physiological arousal that would not be observable with other methods or measures (boucheix, 2017). unveiling the highest level of physiological activation is shown here to lead to advanced insights, as the structure and content of team interaction that was displayed during these moments somewhat differed from what teams displayed outside these high arousal segments. verbal reports or perceptual recall from participants about when they experienced higher stress or arousal are often distorted and do not accurately reflect the actual stressful moment (cf. hunziker, laschinger, et al., 2011). it should be noted that our results have to be interpreted with caution. first, when collecting data in the field (and not in a controlled laboratory environment) increases or fluctuations in skin conductance might be caused not only by higher mental effort, but also by general arousal or body movements (berntson & cacioppo, 2000; cacioppo & tassinary, 1990; stemmler, 2004). this means that the assumed link between physiological arousal and psychological, behavioral or interaction processes requires careful interpretation (akinola, 2010). as mentioned in the theoretical framework, the e4 captures the level of physiological states of arousal, but does not distinguish between excitement and distress. although other studies using self-reported measures of stress (hunziker, laschinger, et al., 2011) indicate that during als especially the negative emotions are high and positive emotions are low (which is why we interpreted in our study high arousal as distress), additional validation measures would be preferable. also, in an als context only the team leader’s skin conductance can be validly measured, as other members would produce distorted or biased physiological measurement because of their activity concerning the compressions and administration of drugs. in order to understand how this affects the team as a whole, more research is needed to study the physiological concordance: to what extent are the physiological processes aligned and how does that influence the dynamics within the team (marci et al., 2007)? finally, even when we interpret arousal as stress, the trigger for higher levels of arousal is still unknown. in the beginning of the scenarios high arousal might be caused by the need to respond quickly to an overload of information, the diagnostic uncertainty and rapidly evolving situations (doumouras et al., 2012; hunziker, johansson, et al., 2011), while at the end of the scenario the team leaders might experience higher arousal because they have to handover the patient and decide on a final diagnosis. both triggers could give rise to the same amount of arousal, but might result in different kinds of team leader behavior. 6.2.2 added value and challenges of using sociometric badges the sociometric badges provided further insight into team interaction processes during these high-arousal moments that adds to the behavioral aspects shown by the video coding. our results showcase how the differences found between the behavior of the more and less effective teams are enriched by the information from the sociometric badges: a difference in their conversational balance became visible. this highlights the importance of also exploring the structure of the verbal interactions, besides looking at the content of the interactions. the sociometric badges have been proven to be reliable and accurate enough to study high-level team interactions, such as participation in conversations and total speaking time (chen, & miller, 2017). however, in our study, we experienced that the hardware in the sociometric badge or e4 might not always function well, resulting in missing data. using equipment from a specific manufacturer also means that you are dependent on a commercial party whenever the equipment is malfunctioning. despite extensive testing, it appeared that the hardware was not functioning properly and technical support was lacking as the researchers support platform for the sociometric badges was discontinued. especially when collecting data from teams with the sociometric badges, the malfunctioning of one badge obstructs the computation of all team interaction dynamics (e.g., one participant could have a great influence on conversational balance). in this study, the failure of two badges resulted in the loss of data for almost half of the teams: all badges were used continuously during the day, meaning that each badge was used approximately five times that day. as downloading the data takes a substantial amount of time, it was not possible to do this during the experiment. as a consequence, it was discovered only afterwards that two badges malfunctioned, resulting in the loss of data for ten participants. in addition, in this specific study the interpretation of the sociometric measures is still on a superficial level: we do have information about the proportion of overlapping speaking time, but we do not know the specifics. for example, we might know that participants spoke more at the same time (overlap), but not whether this consisted of a lot of brief interruptions (e.g., a confirmatory 'yes') or fewer long interruptions (which might be experienced as more disrupting than brief confirmations). more advanced algorithms could result in additional measures that provide more insight into the effects we found, for example, in relation to the length and number of interruptions. 6.2.3 added value and challenges of combining sensor technology with traditional measures although this triangulation of data sources to capture team dynamics in simulation environments results in rich data, there are several challenges that should be noted when adopting such a research design. first, a great benefit of combining video, sociometric and physiological data is that it offers continuous measures of behavior, interaction and physiological intensity on a temporal scale. using this sensory triangulation enables the use of physiological arousal as a process-tracing method (figner & murphy, 2011). this means that it can provide information about behavioral processes, such as decision making, because such physiological data can be measured and collected continuously. at the same time, we experienced the difficulty of synchronizing the data. by using multiple sensors in combination with traditional measures, the risk that one device is malfunctioning or does not align with one of the other devices is substantial, resulting in a much smaller number of teams due to missing data. for future research, we recommend to always have a back-up procedure for this. for both the e4 and the sociometric badges it is possible to use behavioral markers: you push a button on the device and a time-stamp is saved to the data. when you do this in front of the camera you can always synchronize the video with the sensor device. for our study, it was not possible to include this additional step to the procedure, as we were collecting data in an assessment situation in which we already asked a lot from the participants. although our study shows the potential of combining sensor technologies and the added value for team learning research, further research is necessary to validate and ground these methods. each of the team interaction measures that are used in this study, whether observational, self-reported or technology-based, has its limitations. conflicting information about team interaction from these measures needs to be explained and sources of discrepancies need to be understood. this would require not only validation studies, but also transfer studies from the simulation environment to the field. 6.3 limitations and future research next to the problems and limitations that were related to the technology that was adopted, our study design had some limitations that we also have to acknowledge. to begin with, the small sample size (n = 22 teams) and even smaller sample size for the sociometric data resulted in limitations regarding the statistical power of the analyses and generalizability of our findings. moreover, due to insufficient residual degrees of freedom, we were not able to perform manovas for the third research question, resulting in a series of independent t-tests and thus an increased risk of type i errors. consequently, this research has a more exploratory character, which is why we interpreted the results with caution. in the future, we recommend to conduct similar studies with a bigger sample size. in addition to the low sample size, the observed frequency of the video coded behaviors was sometimes very low. this was due to the fact that we decided to zoom in on the beginning and end of the scenario instead of engaging in the time-consuming process of coding the whole scenario. within these time windows of three to four minutes, some of the behaviors were hardly present and when we further zoomed in on the 30-second segments of high arousal, behaviors became even more infrequent. one could also question how many different behaviors one can display in only a 30-second segment. we therefore recommend for future studies to study multiple 30-second segments and longer time windows to allow for more fair comparisons. furthermore, it is known that stress is a complex phenomenon which is difficult to measure (boucsein, 2012). in the present study we chose to include a physiological measure. as described above, these outcomes should be carefully interpreted, as eustress and distress produce similar physiological results. in addition, another limitation of our study is that in the context in which students were assessed it was not possible to obtain a baseline measure. therefore, we could not compare teams on the team leader’s level of arousal; only within person comparisons could be made where peaks in individual skin conductance data were identified in order to pinpoint moments of relative high arousal of a team leader. therefore, future research is advised to measure skin conductance for longer periods and to obtain a baseline measurement to improve the quality of results. in addition, depending on the context of the study, it might be worthwhile for future research to explore the option of measuring the physiological data on the glabrous palmar or plantar surfaces, as this is more reliable and valid (but also more invasive) (boucsein, 2012). studies are available that provide insight into the differences between wristbands compared to palmar measures of skin conductance (e.g., van lier et al., 2017). despite all experienced hurdles and limitations, our study strengthened our idea that wearable sensor technology has the potential to advance insights in team research. sensor technology has the potential to provide objective and unobtrusive measures of complex behavioral and physiological processes as teams do not have to be interrupted while performing their task, which would disrupt their processes. in our study, the participants indicated that wearing the sociometric badges and e4 did not distract them from performing their tasks. wearable sensor measures of interactions are especially promising when linked to psychological or team level constructs such as leadership emergence (chaffin, heidl, hollenbeck, howe, voorhees, & calatone, 2017) and in studying continuous streams of longitudinal data (mathieu, hollenbeck, van knippenberg, & ilgen, 2017). physiological measures of arousal can help to more objectively select stressful moments, which is highly relevant when studying action teams during crisis situations. however, applying these new methods is less straightforward and brings more challenges than is often suggested by manufacturers. first, in order to optimally use these methods, it is important to familiarize with this type of big data, which also adds computational complexity that requires specific expertise in data cleaning and dealing with noise (van keulen, kaminski, matheia, & katoen, 2018). second, in order to apply sensors in a specific situation, many pre-studies are needed to test the reliability of the measures (cf. chaffin, heidl, hollenbeck, howe, voorhees, & calatone, 2017; de laat, endedijk, ufkes, van keulen, & de vries, 2017). ultimately, if one manages to extract meaningful data from the sensors, a final question is how to integrate these data with more traditional measures. as fielding (2012) concludes in his analysis of how methods – including technological data can be mixed, integrating data sources is an innovation in itself and should be treated as such. in other words, although time saving is often advocated as one of the benefits of using sensor technology, we still have a long road to go before this will become reality. 7 conclusion effective team interaction is vital in medical situations and in medical learning and education. a combination of video-observational, sociometric and physiological data can enhance our understanding of the complex behavioral and interaction processes underlying effective team performance and provides alternative learning methods that can be used in the design of trainings and education of medical professionals. technological advances together with the availability of more knowledge about the simultaneous application of such methods are needed to use the full potential of wearable sensor technology in team research and overcome current teething troubles. outcomes of this and future studies might enable future medical professionals to better understand what is required at stressful moments. using these results in the training and during debriefing sessions can potentially optimize team interactions of future medical professionals and enhance the quality of medical care. keypoints more effective teams show greater conversational imbalance than less effective teams, but not during moments of high arousal. both team leaders and followers show some changes in the content of the team interaction during moments of high arousal. more knowledge about the simultaneous application of wearable sensor technologies is needed to use the full potential of these methods in team research. information from sensor technology can in the future be used during debriefing sessions to improve medical training and simulations references akinola, m. (2010). measuring the pulse of an organization: integrating physiological measures into the organizational scholar's toolbox. research in organizational behavior, 30, 203-223. doi:10.1016/j.riob.2010.09.003 andersen, p. o., jensen, m. k., lippert, a., & østergaard, d. (2010). identifying non-technical skills and barriers for improvement of teamwork in cardiac arrest teams. resuscitation, 81(6), 695-702. doi:https://doi.org/10.1016/j.resuscitation.2010.01.024 arnstein, f. (1997). catalogue of human error. british journal of anaesthesia, 79(5), 645-656. doi:10.1093/bja/79.5.645 atwal, a., & caldwell, k. (2005). do all health and social care professionals interact equally: a study of interactions in multidisciplinary teams in the united kingdom. scandinavian journal of caring sciences, 19(3), 268-273. doi:10.1111/j.1471-6712.2005.00338.x bach, d. r., flandin, g., friston, k. j., dolan, r. j. (2009). time-series analysis for rapid event-related skin conductance responses . journal of neuroscience methods, 184(2), 224-234. doi:10.1016/j.jneumeth.2009.08.005. bach dr, flandin g, friston kj, dolan rj. time-series analysis for rapid event-related skin conductance responses. journal of neuroscience methods. 2009;184(2):224-234. doi:10.1016/j.jneumeth.2009.08.005. benedek, m., & kaernbach, c. (2010). a continuous measure of phasic electrodermal activity. journal of neuroscience methods, 190(1), 80-91. doi:10.1016/j.jneumeth.2010.04.028 berntson, g. g., cacioppo, j. t. (2000). from homeostasis to allodynamic regulation. in cacioppo, j. t., tassinary, l. g., berntson, g. (eds.), handbook of psychophysiology (2nd ed., pp. 459–481). new york: cambridge university press. boucheix, j.-m. (2017). the interplay between methodologies, tasks and visualisation formats in the study of visual expertise. frontline learning research, 5(3), 155-166. doi:10.14786/flr.v5i3.311 boucsein, w. (2012). electrodermal activity (2nd. ed.). new york, ny: springer science & business media. brindley, p. g., & reynolds, s. f. (2011). improving verbal communication in critical care medicine. journal of critical care, 26, 155-159. doi:10.1016/j.jcrc.2011.03.004 cacioppo, j. t., & tassinary, l. g. (1990). inferring psychological significance from physiological signals. american psychologist, 45 (1), 16-28. doi:10.1037/0003-066x.45.1.16 chaffin, d., heidl, r., hollenbeck, j. r., howe, m., yu, a., voorhees, c., & calantone, r. (2017). the promise and perils of wearable sensors in organizational research. organizational research methods, 20(1), 3-31. chen, h. e., & miller, s. r. (2017). can wearable sensors be used to capture engineering design team interactions? an investigation into the reliability of sociometric badges. asme 2017 international design engineering technical conferences and computers and information in engineering conference. 7 . doi:10.1115/detc2017-68183. christopoulos, g. i., uy, m. a., & yap, w. j. (2016). the body and the brain: measuring skin conductance responses to understand the emotional experience. organizational research methods, 1094428116681073. doi:10.1177/1094428116681073 cohen, j. (1960). a coefficient of agreement for nominal scales. educational and psychological measurement, 20(1), 37-46. doi:10.1177/001316446002000104 cohen, r. a. (2011). yerkes–dodson law. in encyclopedia of clinical neuropsychology (pp. 2737-2738). springer, new york, ny. cooke, m., irby, d. m., & o'brien, b. c. (2010). educating physicians: a call for reform of medical school and residency . san francisco, ca: jossey-bass. cooper, s., & wakelam, a. (1999). leadership of resuscitation teams: ‘lighthouse leadership’. resuscitation, 42(1), 27-45. doi:10.1016/s0300-9572(99)00080-5 curhan, j. r., & pentland, a. (2007). thin slices of negotiation: predicting outcomes from conversational dynamics within the first 5 minutes. journal of applied psychology, 92(3), 802-811. doi:10.1037/0021-9010.92.3.802 davis, w. a., jones, s., crowell-kuhnberg, a. m., o’keeffe, d., boyle, k. m., klainer, s. b., … yule, s. (2017). operative team communication during simulated emergencies: too busy to respond? surgery, 161(5), 1348-1356. doi:http://doi.org/10.1016/j.surg.2016.09.027 dawson, m. e., schell, a. m., & filion, d. l. (2007). the electrodermal system. in j. t. cacioppo, l. g. tassinary, & g. g. berntson (eds.), handbook of psychophysiology (3rd ed., pp. 159-181). new york: cambridge university press. de laat, s., endedijk, m. d., ufkes, e. g., van keulen, m., & de vries, r. (2017, 24 november). real-time measures of social interaction as predictors for team effectiveness . paper presented at the waop conference 2017, nijmegen, the netherlands. dias, r. d., & neto, a. s. (2016). stress levels during emergency care: a comparison between reality and simulated scenarios. journal of critical care, 33, 8-13. doi:10.1016/j.jcrc.2016.02.010 dimicco, j. m., hollenbach, k. j., pandolfo, a., & bender, w. (2007). the impact of increased awareness while face-to-face. human-computer interaction, 22(1), 47-96 doumouras, a. g., keshet, i., nathens, a. b., ahmed, n., & hicks, c. m. (2012). a crisis of faith? a review of simulation in teaching team-based, crisis management skills to surgical trainees. journal of surgical education, 69(3), 274-281. doi:10.1016/j.jsurg.2011.11.004 edmondson, a. c. (2003). speaking up in the operating room: how team leaders promote learning in interdisciplinary action teams. journal of management studies, 40(6), 1419-1452. doi: 10.1111/1467-6486.00386 entin, e. e., & serfaty, d. (1999). adaptive team coordination. human factors, 41(2), 312-325. doi:10.1518/001872099779591196 fanning, r. m., & gaba, d. m. (2007). the role of debriefing in simulation-based learning. simulation in healthcare, 2(2), 115-125. doi:10.1097/sih.0b013e3180315539 fernandez castelao, e., russo, s. g., riethmüller, m., & boos, m. (2013). effects of team coordination during cardiopulmonary resuscitation: a systematic review of the literature. journal of critical care, 28(4), 504-521. doi:10.1016/j.jcrc.2013.01.005 fielding, n. g. (2012). triangulation and mixed methods designs. journal of mixed methods research, 6(2), 124-136. doi:10.1177/1558689812437101 figner, b., & murphy, r. o. (2011). using skin conductance in judgment and decision making research. in m. schulte-mecklenbeck, a. kuehberger, & r. ranyard (eds.), a handbook of process tracing methods for decision research: a critical review and user's guide (pp. 163-184). new york: psychology press. fischer, f., & järvelä, s. (2014). methodological advances in research on learning and instruction. frontline learning research, 2(4), 1-6. garbarino, m., lai, m., tognetti, s., picard, r. w., & bender, d. (2014). empatica e3 a wearable wireless multi-sensor device for real-time computerized biofeedback and data acquisition. in wireless mobile communication and healthcare (mobihealth), 2014 eai 4th international conference on (pp. 39–42). athens, greece: ieee. gibson, c. b., cooper, c. d., & conger, j. a. (2009). do you see what we see? the complex effects of perceptual distance between leaders and teams. journal of applied psychology, 94(1), 62-76. doi:10.1108/dlo.2009.08123ead.009 goldman, s. r. (2014). perspectives on learning: methodologies for exploring learning processes and outcomes. frontline learning research, 2(4), 46-55. grenvik, a., schaefer, j. j., devita, m. a., & rogers, p. (2004). new aspects on critical care medicine training. current opinion in critical care, 10(4), 233-237. doi:10.1097/01.ccx.0000132654.52131.32 hamaker, e. l. (2012). why researchers should think “within-person”: a paradigmatic rationale. in m. r. mehl & t. s. conner (eds.),handbook of research methods for studying daily life (pp. 43-61). new york, ny: guilford publications. hannah, s. t., uhl-bien, m., avolio, b. j., & cavarretta, f. l. (2009). a framework for examining leadership in extreme contexts. the leadership quarterly, 20(6), 897-919. härgestam, m., lindkvist, m., brulin, c., jacobsson, m., & hultin, m. (2013). communication in interdisciplinary teams: exploring closed-loop communication during in situ trauma team training. bmj open, 3 (10). heaphy, e. d., & dutton, j. e. (2008). positive social interactions and the human body at work: linking organizations and physiology. the academy of management review, 33(1), 137-162. hoogeboom, a. m. g. m., & wilderom, c. p. m. (2015). effective leader behaviors in regularly held staff meetings: surveyed vs. videotaped and video-coded observations. in j. a. allen & n. lehmann-willenbrock & s. g. rogelberg (eds.), the cambridge handbook of meeting science (pp. 381-412). cambridge handbooks in psychology. cambridge university press. http://dx.doi.org/10.1017/cbo9781107589735.017. humphrey, s. e., & aime, f. (2014). team microdynamics: toward an organizing approach to teamwork. the academy of management annals, 8(1), 443–503. doi:10.1080/19416520.2014.904140 doi:10.1080/19416520.2014.904140 hunziker, s., johansson, a. c., tschan, f., semmer, n. k., rock, l., howell, m. d., & marsch, s. (2011). teamwork and leadership in cardiopulmonary resuscitation. journal of the american college of cardiology, 57(24), 2381-2388. doi:10.1016/j.jacc.2011.03.017 hunziker, s., laschinger, l., portmann-schwarz, s., semmer, n. k., tschan, f., & marsch, s. (2011). perceived stress and team performance during a simulated resuscitation. intensive care medicine, 37(9), 1473-1479. doi:10.1007/s00134-011-2277-2 hunziker, s., semmer, n. k., tschan, f., schuetz, p., mueller, b., & marsch, s. (2012). dynamics and association of different acute stress markers with performance during a simulated resuscitation. resuscitation, 83(5), 572-578. doi:10.1016/j.resuscitation.2011.11.013 hunziker, s., tschan, f., semmer, n., howell, m., & marsch, s. (2010). human factors in resuscitation: lessons learned from simulator studies. journal of emergencies, trauma and shock, 3(4), 389-394. doi:10.4103/0974-2700.70764 jacobsson, m., hargestam, m., hultin, m., & brulin, c. (2012). flexible knowledge repertoires: communication by leaders in trauma teams. scandinavian journal of trauma, resuscitation and emergency medicine, 20 (1), 44. doi:10.1186/1757-7241-20-44 kanki, b. g., folk, v. g., & irwin, c. m. (1991). communication variations and aircrew performance. the international journal of aviation psychology, 1(2), 149-162. doi:10.1207/s15327108ijap0102_5 keitel, a., ringleb, m., schwartges, i., weik, u., picker, o., stockhorst, u., et al. (2011). endocrine and psychological stress responses in a simulated emergency situation. psychoneuroendocrino, 36(1), 98–108. kim, t., mcfee, e., olguin, d. o., waber, b., & pentland, a. (2012). sociometric badges: using sensor technology to capture new forms of collaboration. journal of organizational behavior, 33(3), 412-427. doi:10.1002/job.1776 klonek, f. e., burba, m., kauffeld, s., & quera, v. (2016). group interactions and time: using sequential analysis to study group dynamics in project meetings. group dynamics, 20(3). 209-222. kneebone, r. l., nestel, d., vincent, c., & darzi, a. (2007). complexity, risk and simulation in learning procedural skills. medical education, 41, 808-814. kolbe, m., grote, g., waller, m. j., wacker, j., grande, b., burtscher, m. j., & spahn, d. r. (2014). monitoring and talking to the room: autochthonous coordination patterns in team interaction and performance. journal of applied psychology, 99(6), 1254-1267. doi:10.1037/a0037877 koudenburg, n., postmes, t., & gordijn, e. h. (2017). beyond content of conversation: the role of conversational form in the emergence and regulation of social structure. personality and social psychology review, 21(1), 50-71. doi:10.1177/1088868315626022 kozlowski, s. w., & ilgen, d. r. (2006). enhancing the effectiveness of work groups and teams. psychological science in the public interest, 7(3), 77-124. doi: 10.1111/j.1529-1006.2006.00030.x lang, p. j., bradley, m. m., & cuthbert, b. n. (1998). emotion, motivation, and anxiety: brain mechanisms and psychophysiology. biological psychiatry, 44(12), 1248-1263. doi:10.1016/s0006-3223(98)00275-3 larsen, r. j., diener, e., & lucas, r. e. (2002). emotion models, measures, and individual differences. in r. g. lord, r. j. klimoski, & r. kanfer (eds.), emotions in the workplace (pp. 64–106). san francisco: jossey-bass. lazarus, r. s., & folkman, s. (1984). coping and adaptation. in w. d. gentry (ed.), the handbook of behavioral medicine (pp. 282-325). new york: guilford. lei, z., waller, m. j., hagen, j., & kaplan, s. (2016). team adaptiveness in dynamic contexts: contextualizing the roles of interaction patterns and in-process planning. group & organization management, 41(4), 491-525. doi:10.1177/1059601115615246 marci, c. d, ham, j., moran, e., & orr, s. p. (2007). physiologic correlates of perceived therapist empathy and social-emotional process during psychotherapy. the journal of nervous and mental disease, 195(2), 103-111. doi:10.1097/01.nmd.0000253731.71025.fc marks, m. a., zaccaro, s. j., & mathieu, j. e. (2000). performance implications of leader briefings and team-interaction training for team adaptation to novel environments. journal of applied psychology, 85(6), 971-986. doi:10.1037/0021-9010.85.6.971 marsch, s. c., müller, c., marquardt, k., conrad, g., tschan, f., & hunziker, p. r. (2004). human factors affect the quality of cardiopulmonary resuscitation in simulated cardiac arrests. resuscitation, 60(1), 51-56. mathieu, j. e., hollenbeck, j. r., van knippenberg, d., & ilgen, d. r. (2017). a century of work teams in the journal of applied psychology. journal of applied psychology, 102(3), 452. mcgaghie, w. c., issenberg, s. b., cohen, m. e. r., barsuk, j. h., & wayne, d. b. (2011). does simulation-based medical education with deliberate practice yield better results than traditional clinical education? a meta-analytic comparative review of the evidence. academic medicine: journal of the association of american medical colleges, 86 (6), 706-711. doi:10.1097/acm.0b013e318217e119 mcgaghie, w. c., issenberg, s. b., petrusa, e. r., & scalese, r. j. (2010). a critical review of simulation‐based medical education research: 2003–2009. medical education, 44(1), 50-63. doi:10.1111/j.1365-2923.2009.03547.x molenaar, i. (2014). advances in temporal analysis in learning and instruction. frontline learning research, 2(4), 15-24. noldus, l. p., trienes, r. j., hendriksen, a. h., jansen, h., & jansen, r. g. (2000). the observer video-pro: new software for the collection, management, and presentation of time-structured data from videotapes and digital media files. behavior research methods, instruments, & computers, 32(1), 197-206. doi:10.3758/bf03200802 olguín, d. o., waber, b. n., kim, t., mohan, a., ara, k., & pentland, a. (2009). sensible organizations: technology and methodology for automatically measuring organizational behavior. ieee transactions on systems, man, and cybernetics, part b (cybernetics), 39 (1), 43-55. doi:10.1109/tsmcb.2008.2006638 pentland, a. (2012). the new science of building great teams. harvard business review, 90(4), 60-69. pugliese, a., nicholson, g., & bezemer, p. j. (2015). an observational analysis of the impact of board dynamics and directors' participation on perceived board effectiveness. british journal of management, 26 (1), 1-25. doi:10.1111/1467-8551.12074 r core team. (2014). r: a language and environment for statistical computing. vienna, austria: r foundation for statistical computing. retrieved from http://www.r-project.org/ russell, j. a. (1980). a circumplex model of affect. journal of personality and social psychology, 39(6), 1161. sahu, s., & lata, i. (2010). simulation in resuscitation teaching and training, an evidence based practice review. journal of emergencies, trauma and shock, 3(4), 378-384. doi:10.4103/0974-2700.70758 sandroni, c., fenici, p., cavallaro. f., bocci, m. g., scapigliati, a., & antonelli, m. (2005). haemodynamic effects of mental stress during cardiac arrest simulation testing on advanced life support courses. resuscitation, 66(1), 39-44. satish, u. & streufert, s. (2002). value of a cognitive simulation in medicine: towards optimizing decision making performance of healthcare personnel. quality and safety in health care, 11(2): 163-167. schmutz, j., hoffmann, f., heimberg, e., & manser, t. (2015). effective coordination in medical emergency teams: the moderating role of task type. european journal of work and organizational psychology, 24(5), 761-776. doi:10.1080/1359432x.2015.1018184 spiers, j. a. (2004). tech tips: using video management/analysis technology in qualitative research. international journal of qualitative methods, 3(1), 57-61. doi:10.1177/160940690400300106 stachowski, a. a., kaplan, s. a., & waller, m. j. (2009). the benefits of flexible team interaction during crises. journal of applied psychology, 94(6), 1536-1543. doi:10.1037/a0016903 stemmler, g. 2004. physiological processes during emotion. in p. philippot, & r. s. feldman, (eds.), regulation of emotion (pp. 33–70). mahwah, nj: erlbaum. strijbos, j.-w., martens, r. l., prins, f. j., & jochems, w. m. g. (2006). content analysis: what are they talking about? computer & education, 46, 29-48. tannenbaum, s.i., cerasoli, c. p. (2013).do team and individual debriefs enhance performance? a meta-analysis. human factors, 55 (1): 231–245. tschan, f., jenni, n., semmer, n. k., hunziker, s., marsch, s. u., & kolbe, m. (2014). leadership in different resuscitation situations. trends in anaesthesia and critical care, 4(1), 32-36. tschan, f., semmer, n. k., gautschi, d., hunziker, s., spychiger, m., & marsch, s. u. (2006). leader to recovery: group performance and coordinative activities in medical emergency driven groups. human performance, 19(3), 277-304. van der haar, s., koeslag-kreunen, m., euwe, e., & segers, m. (2017). team leader structuring for team effectiveness and team learning in command-and-control rooms. small group research, 48(2), 215-248. vangrieken, k., boon, a., dochy, f., & kyndt, e. (2017). group, team, or something in between? conceptualising and measuring team entitativity. frontline learning research, 5(4), 1-41. doi:10.14786/flr.v5i4.297 van keulen, m., kaminski, m., matheja, c., katoen, j.-p. (2018). rule-based conditioning of probabilistic data. in: proceedings of the 12th international conference on scalable uncertainty management (sum 2018), 3-5 october 2018, milan, italy. springer. van lier h.g. et al. (2017) design decisions for a real time, alcohol craving study using physioand psychological measures. in: de vries p., oinas-kukkonen h., siemons l., beerlage-de jong n., van gemert-pijnen l. (eds) persuasive technology: development and implementation of personalized technologies to change attitudes and behaviors. persuasive 2017. lecture notes in computer science, vol. 10171. springer, cham. weis, p. p., & herbert, c. (2017). bodily reactions to emotional words referring to own versus other people’s emotions. frontiers in psychology, 8, 1277. doi:10.3389/fpsyg.2017.01277 wetzel, c.m., black, s.a., hanna, g.b., athanasiou, t., kneebone, r.l., nestel, d., et al. (2010). the effects of stress and coping on surgical performance during simulations. ann surg., 251(1), 171–6. wiesenfeld, a. r., whitman, p. b., & malatesta, c. z. (1984). individual differences among adult women in sensitivity to infants: evidence in support of an empathy concept. journal of personality and social psychology, 46(1), 118-124. doi:10.1037/0022-3514.46.1.118 zaccaro, s. j., rittman, a. l., & marks, m. a. (2001). team leadership. the leadership quarterly, 12(4), 451-483. doi:http://doi.org/10.1016/s1048-9843(01)00093-5 zijlstra, f. r., waller, m. j., & phillips, s. i. (2012). setting the tone: early interaction patterns in swift-starting teams as a predictor of effectiveness. european journal of work and organizational psychology, 21(5), 749-777. doi:10.1080/1359432x.2012.690399 fraenken et woznitza publication frontline learning research vol.7 no. 1 (2019) 43 50 issn 2295-3159 students’ objects of pride in a learner-focused school setting: an exploratory study judith fraenkena, marold wosnitzab arwth aachen university, germany bmurdoch university, perth, australia article received 28 june 2018/ revised 5 december/ accepted 19 december / available online 31 january abstract in the past decades, schools have become more autonomous and open learning environments. it therefore seems increasingly important for educational research to also consider contextual influences by including autonomous learning settings in its investigations. studying the positive activating emotion of pride seems useful to learn more about the effects of this schooling as pride results from exactly those aspects promoted by autonomous learning: self-evaluation, reflection, self-responsibility and attribution. moreover, pride becomes relevant for a deeper understanding of students’ learning and achievement as pride promotes the desire to repeat already performed achievements in the future. regarding the growing support of individual learning in schools, the present study investigates objects of pride of students attending a school that promotes autonomous, non-competitive, individualized and cooperative learning. students of this school plan their timetables and learning process individually and document it in learning logbooks in which they furthermore can state once a week what they are proud of. in total, 1063 pride statements from 134 students were collected from the learning logbooks. a complementary study, collecting students’ pride statements detached from the learning logbooks, identified 254 pride statements. results show that the pride focus of students at the examined school is learning-oriented. the findings indicate that the specific learning setting of the examined school provides specific school-based pride triggers and thus promotes the learning-oriented pride focus of the students. this paper shall serve as a basis for further research on students’ pride and objects of pride and its potential effects on motivation, achievement and school life. keywords: students’ objects of pride; school setting; student-centred teaching info corresponding author: email: judith.fraenken@rwth-aachen.de doi: 10.14786/flr.v7i1.387 1. introduction today, education systems all over the world are facing massive changes in schooling with schools becoming more autonomous and open learning environments being introduced in the past decades (eurydice, 2008). it therefore seems increasingly important for educational research to also consider contextual influences (e.g.,urdan & schoenfelder, 2006) by including different, non-traditional learning settings in its investigations. the overall aim of autonomous school settings is not only to individualise learning but have students take responsibility for their own learning, reflect and evaluate their own learning processes, feel that learning outcomes are their own success and set individual direction-giving goals for future learning processes (e.g., assor, 2012; madjar & assor, 2013; niemiec & ryan, 2009). in this context, studying students’ pride seems useful to learn more about the effects of this schooling. this positive activating emotion precisely results from those aspects promoted by autonomous school settings: pride can be seen as the consequence of a successful evaluation of a specific event or object for which one feels responsible (lagattuta & thompson, 2007; lewis, 2016). furthermore, for the elicitation of pride, one’s own actions and outcomes of these actions have to be reflected and attributed to internal factors (hart & matsuba, 2007; kornilaki & chlouverakis, 2004; tracy, robins, & lagattuta, 2005; tracy, shariff, & cheng, 2010). this process is based on a person’s self-concept which includes complex cognitive processes such as self-perception and self-evaluation, which is why pride is defined as a self-conscious emotion (lewis, 2016; tracy & robins, 2007). at the same time, the respective person’s society’s standards, rules and goals (srg’s) serve as reference point for the evaluation of their own actions and determination of success (lewis, 2016). consequently, students’ sense of responsibility, their ability of self-evaluation and reflection are pivotal for the elicitation of pride, which is why a school promoting those aspects should be examined in this research. moreover, the emotion of pride is relevant for students’ learning as pride is described as an emotion that results from personal achievement and further promotes the desire to repeat or even outdo this achievement in the future (fredrickson, 2001; lewis, 2016). pride is furthermore considered an incentive to persevere on a task despite initial costs (williams & desteno, 2008), which is why studies in this field additionally become relevant for a deeper understanding of students’ learning and achievement and the promotion of students’ sense of pride. while previous research has mainly focused on the impact of pride on achievement and motivation as well as the attributions of success or the correlations of pride to various aspects such as goal-regulation, self-control or achievement values (e.g.,buechner, pekrun, & lichtenfeld, 2016; carver, sinclair, & johnson, 2010; oades-sese, matthews, & lewis, 2014; pekrun, goetz, frenzel, barchfeld, & perry, 2011; weidman, tracy, & elliot, 2016; williams & desteno, 2008), the object of pride, i.e. what a person is proud of, has been left somewhat disregarded. an exploratory qualitative study aimed to categorize different domains and emphases of students’ pride in the school context of a traditional teacher-centred german comprehensive school (fraenken & wosnitza, 2018). five main categories could be found namely learning in school (aspects that are directly related to learning in school), social aspects (social aspects that do not have to be directly related to learning in school), activities besides (performance at) school (aspects that are established outside the classroom), me (aspects that are related to one’s own person and not specific actions), and persons and animals (aspects that are related to other people or animals). results indicated that students’ pride regarding learning in school appeared to be more achievement-oriented and less learning-oriented in the traditional teacher centred school (fraenken & wosnitza, 2018). this could be hypothesised to result from the school’s competitive learning setting with clearly defined, standardised goals which promotes achievement-focused goals (self-brown & mathews, 2003) and consequently achievement-focused pride after those goals are being achieved. in the present study, the objects of pride of students from a learner-focused and autonomy promoting school setting are being explored. according to self-brown and mathews (2003), in a non-competitive setting where learning goals are defined and evaluated individually, students’ goals are learning-oriented. concerning the empirically verified positive correlations between achievement goals and pride (pekrun, elliot, & maier, 2006, 2009), it can be assumed that students’ objects of pride are learning-oriented in the examined school where students set their own goals and define their own learning process. the finding that learner-centred and autonomy-supporting structures apparently have the strongest positive relation to students’ mastery goals compared to performance goals (ames, 1992; meece, 2003) strengthens this assumption. with a gender perspective, significant differences can be expected regarding students’ pride in learning in school as boys perform worse than girls in most performance areas of german schools (mößle & lohmann, 2014) and therefore seem to have less reason to be proud at that area. confirming this, female students from a traditional teacher-centred school made significantly more pride statements about learning in school than male students (fraenken & wosnitza, 2018). as women are rated as more communal and socially committed than men (brosi, spörrle, welpe, & heilman, 2016), it can be expected that female students make more statements about social aspects than male students. this was also found in a traditional school, where female students made significantly more pride statements about social aspects than male students (fraenken & wosnitza, 2018). furthermore, tracy and beall (2011) found out that men’s pride is considered an attractive expression whereas women’s pride is one of the least attractive. this could lead to female students revealing overall less pride than male students. by encouraging and supporting children’s feelings of pride in their academic success, their perception of being responsible for their own success is being promoted (thompson, 1991). knowing about the range of objects of pride could help detecting different kinds of potential pride triggers in order to encourage students’ pride. this study is a reaction to the emerging change in schools towards student-centred approaches. the overall aim is to investigate the objects and emphases of pride of students attending one progressive school that promotes autonomous, individualized, cooperative and non-competitive learning. it is expected that the students focus their pride on their learning process and progress and less on their achievement outcomes. the results are thought to lay a foundation for future research on pride and possible connections of objects of pride with motivation, learning and achievement. 2. methodology in order to examine the objects of pride of students in a learner-focused school setting, an exploratory qualitative approach was chosen. data were collected at a german comprehensive school with a student-centred approach to teaching. the five major subjects (german, english, maths, natural sciences, social sciences) are provided as topic-related modules with different degrees of difficulty, on which students work individually in their own time and speed. minor subjects are taught in topicand project-related workshops. the school supplies subject-specific classrooms in which students of all grades work individually or cooperatively on their current modules. teachers act as advisors and tutors while students also support each other. students’ learning process and progress is planned and documented individually by themselves in so-called “learning logbooks”, discussed in individual weekly meetings with a tutor, and monitored by exams which the students take as soon as they have completed a module and state to feel ready. in the learning logbooks, students can voluntarily state their personal goal of the week and additionally what they are proud of on a weekly basis. for this, students complete the phrase “i am proud of…” without having to focus on the school sector. these weekly statements of what students are proud of form the database of this study. the sample consisted of 134 students (school years 5-8, 57% female). pride statements were collected from the learning logbooks of one school year and were separated into single statements (n=1063). based on deductive categorisation according to qualitative content analysis (mayring, 2010, 2015), the existing category system of students’ pride in a traditional school (fraenken & wosnitza, 2018) was used during the coding process. all statements were coded into the category system of students’ pride. two different researchers coded 56.44% of all statements (interrater agreement κ=.91). complementary study: the logbook entries are also visible for the students’ teachers and parents and consequently not anonymous. in order to exclude a methodological effect, a complementary study was conducted to collect students’ pride statements detached from the learning logbooks. for two weeks, 110 students of the same school (school years 5-8, 46% female) wrote down their pride statements anonymously. phrasing was adopted from the learning logbooks “i am proud of…”. 254 single statements were identified. 3. results analyses of data showed a wide range of students’ objects of pride (overview of the number of statements in table 1). with 72.53% of all learning-logbook pride statements, students focused their pride on learning in school. the main emphasis of the pride statements within the category learning in school lied with 73.15% on the statements referring to learning process and progress. within the category learning process and progress various new subcategories, which did not appear in the existing category system of the traditional school, could be found which represented 71.45% of all statements within that category. for example, 43.97% of the pride statements were referring to modules and subjects (“that i could almost finish the math module” (#9-20 )), 22.87% to taking tests or exams (“that i can write my german test this week” (#18-24)) and 4.61% to achieving the personal goal of the week (“that i have reached my goal of the week” (#11-2)). most of the remaining statements (21.45%) regarding learning process and progress stated getting a lot of work done (“because i’ve accomplished so much” (#26-37)). the remaining statements named active participation in the classroom (5.67%), homework or learning at home (1.42%) and understanding something (0.35%). furthermore, within the main category learning in school, 26.85% of the statements focused on achievement outcomes and results, e.g., grades (“that i have a b+ in the spanish exam” (#27-14)) or praise (“that i was praised in class” (#16-7)). social aspects included statements about e.g., social behaviour (“that i deal well with my friends” (#43-14)) or classroom discipline (“because i haven’t broken any rules in three days” (#37-1)), which represented another newly found subcategory. activities besides (performance at) school meant for instance hobbies and spare time (“that i had a good, fun birthday” (#48-19)). the category me included e.g., statements about the own personality and characteristics (“that i am honest with my assessments” (#9-3)) whereas persons and animals referred to others, e.g., teacher (“i am proud of the teachers because they are always nice and helpful” (#41-33)). table 1 main categories of students’ pride in a student-centred school results of the complementary study revealed that the focus also lied on learning in school (46.85% of all statements) but compared to the logbook entries, students made significantly less statements about learning in school anonymously [χ²(1, n=1317)=61.707, p=.000]. within that main category, students focused on learning process and progress (73.11% of all statements within the category learning in school). the remaining 26.89% referred to achievement outcomes and results. in total, 22.04% of the statements were related to activities besides (performance at) school which are significantly more statements than within the logbooks [χ²(1, n=1317)=86.638, p=.000]. of all the remaining statements, 4.72% related to social aspects, 2.76% to me and 6.3% to persons and animals. overall, 3.94% of the statements announced not to be proud of anything and 13.39% of all statements could not be coded (rest). 4. discussion besides pointing out a wide range of pride triggers, the results indicate a connection between the students’ objects of pride and their autonomous and individual way of learning and operating in the examined school. as expected, the students’ pride focused on their learning (represented by the category “learning process and progress” and less on their performance (represented by the category “achievement outcomes and results”). one reason for the large number of statements about their learning process and progress could be the fact that the students can set their own realistic goals by determining and conducting their own learning process self-responsibly. as one has to feel responsible for success to feel pride (lewis, 2016), students have to feel personally responsible for their learning process and progress to consequently be proud of it which is being promoted by the school. this result underlines the importance of students’ perceived responsibilities in the achievement context and the associated pride focus. the responsibility for the students’ own learning process was also being reflected by newly found subcategories which showed that the specific learning setting of the examined school provides specific pride triggers. by providing modules with different degrees of difficulty to be chosen by the students who work on them in their own time and speed (subcategory modules and subjects), the school enables the students’ high prospective and, thus, retrospective responsibility for their personal learning process and progress and therefore a school-specific pride trigger. the subcategory taking tests or exams also shows the students’ high responsibility for their learning process and progress as they could only be proud of being able to write a test soon because they individually set the date for their exams based on their own ability and learning progress. helker and wosnitza (2016) found that students’ sense of responsibility for their own learning process and achievement correlates with their sense of competence and autonomy. this puts emphasis on the assumption that the students’ high level of autonomy promotes their sense of responsibility and therefore their sense of pride. the learning logbooks serve as a foundation for students’ self-reflection and consequently promote their ability of self-evaluation. by filling the logbooks with learning contents, learning goals and achievements or defaults, students reflect their learning process and progress every day and thus create the basis for the ability of attributing success to themselves and consequently feeling pride. by doing so, students do not evaluate their learning process independently of others but with regard to the school’s standards, rules and goals (srgs) which is essential for the feeling of pride (lewis, 2016). keeping the logbooks promotes the awareness and adaption of the school’s srg’s and therefore the elicitation of students’ pride focus on their learning process and progress. whereas keeping learning logbooks appears to promote the students’ sense of responsibility and self-reflection, it must also be considered that the logbooks, as research object, could have an impact on students’ pride statements as they are accessible to teachers and parents. even though, the results of the anonymous complementary study reveal that the students’ focus still lied on learning in school, only less prominent compared to the findings in the logbooks. by contrast, students expressed more pride about activities besides (performance at) school than they did within their learning logbooks. apparently, although the students were not demanded to focus their pride statements on the school sector, they instinctively did so when filling out the logbooks. the very low number of statements about activities besides (performance at) school in the learning logbooks could therefore be explained by social desirability. with regard to learning in school, students of the complementary study still focused on learning process and progress which confirms and reinforces the considerations made above. contrary to our expectations, no gender differences could be found regarding main or sub categories. with regard to the examined students’ focus on their learning process and progress instead of their achievement and outcomes, it is, however, plausible that potential gender differences in performance did not have a huge impact on their sense of pride. furthermore, due to the separately coordinated learning schedules and the attention to the individual, gender stereotypes apparently play a minor role than in a more competitive school with direct comparisons between students’ approaches and achievement. school year differences, however, could be found. the continuously decreased focus on achievement outcomes and results from school year to school year as well as the simultaneously occurred increasing focus on learning process and progress gives reason to assume that the students internalize the school’s srg’s (lewis, 2016) over time and therefore adapt their pride focus on learning continuously. as the elicitation of pride requires complex cognitive processes which are involving over time (hart & matsuba, 2007; lewis, 2016; tracy & robins, 2007), the results could reflect this involvement as the students seem to learn over time that they are responsible for their own learning process and can consequently be proud of it. in summary, the examined school itself offers school-related, learning oriented pride triggers and additionally promotes the students’ ability to perform complex cognitive processes in order to feel pride. in combination with previous research on students’ pride in a traditional school (fraenken & wosnitza, 2018) it can be assumed that different learning settings provide different potential objects of pride and promote different pride focuses. however, the present study is explorative and only related to one particular school. building on this study, future research should include various school settings and a greater number of schools in order to further investigate the assumed connection between students’ objects of pride and their school environment. additionally, the results of the present study should be used to investigate students’ degree of autonomy, responsibility and ability of self-reflection and its impact on their pride and by implication their motivation and achievement. as the results of the study indicate a connection between the students’ objects of pride and their autonomous and self-responsible learning, this should be further explored and verified. furthermore, the impact of students’ objects of pride on their learning process, achievement and motivation should be explored in order to find out if and how the objects of pride themselves matter in the achievement context. acknowledgements our sincere appreciation is extended to kerstin helker for her invaluable advice and her comments on the manuscript. keypoints an autonomous school setting may trigger students’ pride in making them feel responsible for their own learning. a learner-focused school setting promotes learning-oriented rather than outcome-oriented pride. the longer students attend an autonomous school setting, the more they tend to feel proud of their learning process and progress. references ames, c. (1992). classrooms: goals, structures, and student motivation. journal of educational psychology, 84(3), 261–271. doi:10.1037//0022-0663.84.3.261 assor, a. (2012). allowing choice and nurturing an inner compass: educational practices supporting students’ need for autonomy. in s. l. christenson, a. l. reschly, & c. wylie (eds.), the handbook of research on student engagement(pp. 421-439). new york: springer science. brosi, p., spörrle, m., welpe, i. m., & heilman, m. e. (2016). expressing pride: effects on perceived agency, communality, and stereotype-based gender disparities. journal of applied psychology , no pagination specified. doi:10.1037/apl0000122 buechner, v. l., pekrun, r., & lichtenfeld, s. (2016). the achievement pride scales (aps). european journal of psychological assessment, 1-12. doi:10.1027/1015-5759/a000325 carver, c. s., sinclair, s., & johnson, s. l. (2010). authentic and hubristic pride: differential relations to aspects of goal regulation, affect, and self-control. journal of research in personality, 44 (6), 698-703. doi:10.1016/j.jrp.2010.09.004 eurydice. (2008). levels of autonomy and responsibilities of teachers in europe. fraenken, j., & wosnitza, m. (2018). stolz im schulalltag worauf sind schülerinnen und schüler stolz? [pride in everyday school life what are students proud of?]. in g. hagenauer & t. hascher (eds.), emotionen und emotionsregulierung in der schule und hochschule(pp. 15-28). münster: waxmann. fredrickson, b. l. (2001). the role of positive emotions in positive psychology: the broaden-and-build theory of positive emotions. the american psychologist, 56(3), 218-226. doi:10.1037//0003-066x.56.3.218 hart, d., & matsuba, m. k. (2007). the development of pride and moral life. in j. l. tracy, r. w. robins, & j. p. tangney (eds.), the self-conscious emotions: theory and research(pp. 114-133). new york, ny, us: guilford press. helker, k., & wosnitza, m. (2016). the interplay of students’ and parents’ responsibility judgements in the school context and their associations with student motivation and achievement. international journal of educational research, 76, 34-49. doi:10.1016/j.ijer.2016.01.001 kornilaki, e. n., & chlouverakis, g. (2004). the situational antecedents of pride and happiness: developmental and domain differences. british journal of developmental psychology, 22(4), 605-619. doi:10.1348/0261510042378245 lagattuta, k. h., & thompson, r. a. (2007). the development of self-conscious emotions: cognitive processes and social influences. in j. l. tracy, r. w. robins, & j. p. tangney (eds.), the self-conscious emotions: theory and research(pp. 91-113). new york, ny, us: guilford press. lewis, m. (2016). self-conscious emotions. embarrassment, pride, shame, guilt, and hubris. in l. feldman barrett, m. lewis, & j. m. haviland-jones (eds.), handbook of emotions(4th ed., pp. 792-814). new york: the guilford press. madjar, n., & assor, a. (2013). two types of perceived control over learning: perceived efficacy and perceived autonomy. in j. hattie & e. m. anderman (eds.), international guide to student achievement(pp. 439-441). new york: routledge. mayring, p. (2010). qualitative inhaltsanalyse [qualitative content analysis]. in g. mey & k. mruck (eds.), handbuch qualitative forschung in der psychologie(pp. 601-613). wiesbaden: vs verlag für sozialwissenschaften. mayring, p. (2015).qualitative inhaltsanalyse: grundlagen und techniken [qualitative content analysis: basics and techniques] (12th ed.). weinheim: beltz meece, j. l. (2003). applying learner-centered principles to middle school education. theory into practice, 42(2), 109-116. doi:10.1207/s15430421tip4202_4 mößle, t., & lohmann, a. (2014). entwicklung akademischer leistungen im geschlechtervergleich [development of academic performance in gender comparison]. in t. mößle, c. pfeiffer, & d. baier (eds.), die krise der jungen. phänomenbeschreibung und erklärungsansätze (pp. 19-27). baden-baden: nomos. niemiec, c. p., & ryan, r. m. (2009). autonomy, competence, and relatedness in the classroom: applying self-determination theory to educational practice. school field, 7(2), 133-144. doi:10.1177/1477878509104318 oades-sese, g. v., matthews, t. a., & lewis, m. (2014). shame and pride and their effects on student achievement. in r. pekrun & l. linnenbrink-garcia (eds.), international handbook of emotions in education(pp. 246-264). new york: routledge. pekrun, r., elliot, a. j., & maier, m. a. (2006). achievement goals and discrete achievement emotions: a theoretical model and prospective test. the journal of educational psychology, 98(3), 583-597. doi:10.1037/0022-0663.98.3.583 pekrun, r., elliot, a. j., & maier, m. a. (2009). achievement goals and achievement emotions: testing a model of their joint relations with academic performance. journal of educational psychology, 101(1), 115-135. doi:10.1037/a0013383 pekrun, r., goetz, t., frenzel, a. c., barchfeld, p., & perry, r. p. (2011). measuring emotions in students’ learning and performance: the achievement emotions questionnaire (aeq). contemporary educational psychology, 36(1), 36-48. doi:10.1016/j.cedpsych.2010.10.002 self-brown, s. r., & mathews, s. (2003). effects of classroom structure on student achievement goal orientation. the journal of educational research, 97(2), 106-112. doi:10.1080/00220670309597513 thompson, r. (1991). emotional regulation and emotional development. educational psychology review, 3(4), 269-307. doi:10.1007/bf01319934 tracy, j. l., & beall, a. t. (2011). happy guys finish last: the impact of emotion expressions on sexual attraction.emotion, 11(6), 1379-1387. doi:10.1037/a0022902 tracy, j. l., & robins, r. w. (2007). the self in self-conscious emotions. a cognitive appraisal approach. in j. l. tracy, r. w. robins, & j. p. tangney (eds.), the self-conscious emotions: theory and research(pp. 3-20). new york: the guilford press. tracy, j. l., robins, r. w., & lagattuta, k. h. (2005). can children recognize pride? emotion, 5(3), 251-257. doi:10.1037/1528-3542.5.3.251 tracy, j. l., shariff, a. f., & cheng, j. t. (2010). a naturalist’s view of pride. emotion review, 2(2), 163-177. doi:10.1177/1754073909354627 urdan, t., & schoenfelder, e. (2006). classroom effects on student motivation: goal structures, social relationships, and competence beliefs. journal of school psychology, 44(5), 331-349. doi:10.1016/j.jsp.2006.04.003 weidman, a. c., tracy, j. l., & elliot, a. j. (2016). the benefits of following your pride: authentic pride promotes achievement. journal of personality, 84(5), 607-622. doi:10.1111/jopy.12184 williams, l. a., & desteno, d. (2008). pride and perseverance: the motivational role of pride. j pers soc psychol, 94 (6), 1007-1017. doi:10.1037/0022-3514.94.6.1007 microsoft word kjellström et al_publication.docx ! ! ! ! ! frontline!learning!research!vol.4!no.!5!(2016)!1! 0.55) in majority reflect the more sophisticated levels of epistemological thinking of the model of van rossum and hamer (2010, 2013), i.e. di3 (‘different solutions are illustrated from different perspectives’), bt3 (‘focus on application on what is learnt’), un5 (‘realise how different perspectives influence what you see and understand’), un4 (‘see the underlying assumptions and how these lead to particular conclusions’) and un3 (‘to see connections within a larger context’). figure 1. unidimensional factor structure of the teaching and learning questionnaire the relevance of conducting a confirmatory factor analysis prior to the rasch analysis lies in the assessment of the unidimensionality assumption. this is one of the available strategies, i.e. one can verify if the data fits a unidimensional model using confirmatory factor analysis and then proceed to the rasch analysis. kjellström et al ! | f l r ! ! 15! the rating scale model results indicated an adequate fit of the items to the unidimensional model (see table 5), with mean infit meansquare of 0.97 (min = 0.69, max = 1.30). these values are within the range considered productive for measurement (wright, linacre, gustafson & martin-lof, 1994). the location parameter of table 5 is basically the item difficulty (i.e. the resistance to endorsing a rating scale response category), while the threshold are the points where the category curves intersect. so, threshold 1 is the point in the latent variable where the category 1 (‘least important’, value = 0) intersects with category 2 (value = 1; figure 1). this means that people with ability or proficiency in the latent variable equal to threshold 1 (-0.41) will have equal probability of choosing item sb1 to reflect category 1 or category 2. table 5 item infit, location and thresholds of items in the edtlq item infit meansquare location threshold 1 threshold 2 threshold 3 threshold 4 sb1 1.06 0.82 -0.41 0.14 1.17 2.37 sb2 0.97 1.44 0.21 0.76 1.79 2.99 sb3 0.92 0.89 -0.34 0.22 1.24 2.44 sb4 0.80 0.49 -0.74 -0.18 0.84 2.04 sb5 1.17 1.35 0.12 0.67 1.70 2.90 sb6 1.21 1.29 0.06 0.62 1.65 2.85 di1 1.03 1.32 0.09 0.65 1.68 2.88 di2 1.30 2.48 1.25 1.81 2.83 4.03 di3 0.82 0.49 -0.74 -0.19 0.84 2.04 di4 1.07 0.82 -0.41 0.15 1.17 2.38 di5 0.89 0.99 -0.24 0.31 1.34 2.54 di6 1.20 2.29 1.05 1.61 2.64 3.84 ap1 1.04 2.60 1.37 1.93 2.95 4.15 ap2 0.86 0.87 -0.36 0.19 1.22 2.42 ap3 1.02 1.23 0.00 0.55 1.58 2.78 ap4 1.05 -0.65 -1.88 -1.32 -0.29 0.91 ap5 1.03 0.92 -0.32 0.24 1.27 2.47 ap6 1.07 1.15 -0.08 0.47 1.50 2.70 un1 1.08 1.10 -0.13 0.42 1.45 2.65 un2 1.23 3.87 2.64 3.20 4.22 5.42 kjellström et al ! | f l r ! ! 16! un3 0.76 0.19 -1.04 -0.48 0.54 1.75 un4 0.74 0.46 -0.77 -0.22 0.81 2.01 un5 0.76 0.41 -0.82 -0.26 0.77 1.97 un6 0.91 3.11 1.88 2.44 3.46 4.66 re1 0.94 0.90 -0.33 0.22 1.25 2.45 re2 0.72 0.54 -0.69 -0.13 0.90 2.10 re3 1.07 0.81 -0.42 0.14 1.16 2.36 re4 0.80 1.76 0.53 1.08 2.11 3.31 re5 0.67 1.42 0.19 0.75 1.77 2.97 re6 0.93 0.57 -0.66 -0.10 0.92 2.12 bt1 1.21 1.70 0.47 1.03 2.05 3.26 bt2 1.24 0.16 -1.07 -0.52 0.51 1.71 bt3 0.67 1.32 0.09 0.65 1.67 2.87 bt4 1.06 1.41 0.18 0.74 1.76 2.96 bt5 0.80 1.54 0.31 0.86 1.89 3.09 bt6 0.96 2.26 1.03 1.58 2.61 3.81 another indicator of the instrument quality is the person separation reliability and the items separation reliability. both have the same interpretation as the reliability calculated using cronbach's alpha: the closer to one, the greater the reliability of the measurement. with regard to the interpretation this means the closer the value to 1, the better the pattern of people’s responses, or items endorsements fits the measurement structure (hibbard, collins, baker & mahoney, 2009). in other words, the separation reliability of people indicates how sure we can be that a person with an estimated ability (in this case perhaps it is better to use affinity to the item) of 2 has, indeed, a greater ability (i.e. affinity to the item) than another person who has an estimated ability (i.e. affinity to the item) of 1. similarly, the items separation reliability indicates the confidence that an estimated item difficulty 2 has, indeed, a greater difficulty (i.e. lower endorsement level) than another item with estimated difficulty of 1. both are calculated using a relationship between the standard error’s variance and the mean square error (mse): !"#$%&!!"#$%$&'()!!"#$%&$#$'( = !"# !!!!!"#$%#&%!!""#" − (!"!!!) !"# !!!!!"#$%#&%!!""#" !"#$%!!"#$%$&'()!!"#$%&$#$'( = !"# !!!!!"#$%#&%!!""#" − (!"#!!) !"# !!!!!"#$%#&%!!""#" the separation reliability of the edtlq items was 0.99 and the separation reliability of persons was 0.87. kjellström et al ! | f l r ! ! 17! figure 2. category characteristic curves for items (statements) from the question “a good study book”. category one = black line; category two = red line; category three = green line; category four = blue line; category five = light blue line. figure 2 shows the category characteristic curves for each item (statement) of the “a good study book” question. the hypothesized ordering of the categories matches the empirical evidence that the first category (value = 1, “least important”) is less difficult to endorse than the second category (value = 2), which is less difficult to endorse than the third category (value = 3), and so on, until the fifth category, the most difficult to endorse (value = 5, “most important”; see also table 2). each category, in every item, presents a greater probability of being endorsed than the other categories in at least a small range of the latent variable. thus, there was no issue related to the ordering of the scales. the same scenario was found in every item of the questionnaire, as can be verified in table 5. this result means that the response categories of the questionnaire (ranging from “least important” to “most important”) are empirically supported, and do not need to be altered. figure 3 shows the parameter distribution of all the items in the edtlq. in this figure items ap4 is the easiest to endorse, followed by bt2, un3 and a small cluster of un5, un4, di3, sb4, re2 and re6. theoretically these items would mostly reflect epistemological level 4 and 5, implying that these levels of epistemological thinking regarding application, good teaching, understanding and discussion correspond closest to what the respondents in the current study feel is important in learning and teaching. at the other end of the scale, the most difficult to endorse statements reflect both the least and the most sophisticated ways of thinking, including the items ! un2, un6 and ap1 reflecting understanding and application focusing on recall and answering exam questions correctly, and ! di2, di6 and bt6 reflecting discussion aimed at hearing different views and the teacher supplying answers and covering the content in the course. kjellström et al ! | f l r ! ! 18! figure 3. person-item map, all items of the edtlq this seems to point towards a preference of the respondent group (teachers and staff) as a whole for epistemological thinking at level 4 and to a lesser degree 5, and a rejection of both the most sophisticated level 6, as well as the least sophisticated levels 1 and 2, matching the expected pattern under the consistency hypothesis. to examine in more detail if this pattern is consistent for all the scales, each scale is plotted and discussed separately below. first we will discuss four of the six scales that seem to behave in a way that confirms the consistency hypothesis where items reflecting the extremes of epistemic thinking are relatively difficult to endorse. these are the scale regarding discussion (di) – discussed in more detail below – and the scales regarding application (ap) and understanding (un) and good teaching (bt). figure 4 shows the parameter distribution of the items regarding discussions during a course. considering only the discussion during course items, item di3 (“different solutions are illustrated from multiple perspectives”) was the least difficult to endorse, while item di2 (“the teacher answers students’ questions”) was the most difficult to endorse. table 6 links the discussion in course items ordered by the difficulty parameter of the rasch model to the respective epistemological development level of the items, as per van rossum and hamer’s theory. in this scale the items reflecting the least sophisticated levels of thinking (di1, di6 and di2) are on average the most difficult to endorse, reflecting the pattern found for all the items together. kjellström et al ! | f l r ! ! 19! figure 4. person-item map, subset of the discussion items (di) table 6 discussion in course items by order of endorsement and epistemological level (van rossum & hamer, 2010) item code item parameter item endorsement order item epistemological level item di3 0.49 1 5 different solutions are illustrated from multiple perspectives. di4 0.82 2 5 students and teachers learn together and from each other. di5 0.99 3 4 opinions are based on and/or supported by evidence. di1 1.32 4 2 the same students don’t dominate. di6 2.29 5 3 you hear all the different views that people have. di2 2.48 6 2 the teacher answers students’ questions. a similar pattern can be seen for the items regarding being able to apply knowledge (figure 5). items ap4, ap2 and ap5 (reflecting levels 4 and 5) are the least difficult to endorse, while those items reflecting the least sophisticated levels (ap1, ap3; level 2 and 3) and ap6 reflecting the most sophisticated epistemic thinking level were the most difficult to endorse, again reflecting the pattern expected under the consistency hypothesis. the third scale that displays the response pattern expected under the consistency hypothesis is the scale with items regarding the “to understand something means to” issue (figure 6). here we see again that the items reflecting higher mid-level epistemic thinking (un3, un5 and un4, reflecting respectively level 4 and level 5 twice, see table 2) are relatively easy to endorse, whilst those reflecting the less sophisticated kjellström et al ! | f l r ! ! 20! epistemological positions (un1, un6 and un2, reflecting respectively level 3, 2 and 1) are rejected, i.e. difficult to endorse. figure 5. person-item map, subset of the application items (ap) figure 6. person-item map, subset of the understanding items (un) finally, figure 7 shows the parameter distribution of the fourth scale that seems to behave as expected – the items regarding the issue “the best teaching ought to …”. item bt2 (“inspire the students so kjellström et al ! | f l r ! ! 21! that they are motivated to learn”) was the least difficult to endorse, while item bt6 (“make sure that the whole content of the course is covered”) was the most difficult to endorse. examining the theoretically linked level of epistemic thinking to the level of endorsement (table 7), again we see that those items reflecting both the most sophisticated as well as the least sophisticated levels of van rossum and hamer’s model are difficult to endorse, a pattern that is as expected given the rank ordering of the items in combination with the consistency hypothesis. figure 7. person-item map, subset of the best teaching items (bt) table 7 best teaching items by order of endorsement and epistemological level (van rossum & hamer, 2010) item code item parameter item endorsement order item epistemological level item bt2 0.16 1 4 inspire the students so that they are motivated to learn. bt3 1.32 2 3 focus on application of what is learnt. bt4 1.41 3 5 make students question their current view on the world. bt5 1.54 4 5 help students realize the limitations of current knowledge and understandings. bt1 1.70 5 2 focus on conveying the most important facts and knowledge as clearly as possible bt6 2.26 6 1 make sure that the whole content of the course is covered. kjellström et al ! | f l r ! ! 22! the two remaining scales, reflecting the views on a good study book (figure 8 and table 8) and responsibility for learning (figure 9 and table 9) both do not seem to fit the expectation of the consistency hypothesis. figure 8. person-item map, subset of the good textbook items (sb) table 8 good textbook items by order of endorsement and epistemological level (van rossum & hamer, 2013) item code item parameter item endorsement order item epistemological level item sb4 0.49 1 4 evokes a critical attitude and invites reflection. sb1 0.82 2 2 is well structured, gives the main points and the consensus story of the subject. sb3 0.89 3 5 presents alternative conclusions, shows different interpretations within a subject matter or clarifies different perspectives on the content. sb6 1.29 4 5 makes students think about the fundamental question of knowledge: "how can we know this?" sb5 1.35 5 4 provokes and challenges the students’ pre-conceptions sb2 1.44 6 3 provides many up-to-date examples and examples from practice. kjellström et al ! | f l r ! ! 23! in examining the good study book items, figure 8 shows that item sb4 (“evokes a critical attitude and invites reflection”) was the least difficult to endorse, while item sb2 (“provides many up-to-date examples and examples from practice”) was the most difficult to endorse. examining the level of endorsement by theoretically expected level of epistemic thinking (table 8) the parameters seem to indicate a clustering inconsistent with the consistency hypothesis. level 4 thinking represents what the respondent group feels is most important in a good textbook the closest, closely followed by a preference for clear structure (level 2) and clarification of alternative perspectives (level 5). study books that challenge current thinking and require reflection on the nature of knowledge seem to be more difficult to endorse, although books with many examples from practice (level 3) also seem to relatively unimportant to the teachers and researchers in this respondent group. the items regarding the nature of a good study book (van rossum & hamer, 2013) are based on student responses only, and have not been similarly confirmed in multiple studies as views reflected in the previous four scales on discussion (di), application (ap), understanding (un) and good teaching (bt). the final scale is shown in figure 9 which gives the parameter distribution of the items regarding the issue “for students to take responsibility for learning means that students…”. item re2 (“have an active interest in the course and be motivated to learn”) was the least difficult to endorse, while item re4 (“follow their own interests in searching for knowledge”) was the most difficult to endorse. table 9 shows the responsibility items ordered by the difficulty parameter of the rasch model and the respective epistemological development level of the items, as per van rossum and hamer’s theory. item re2 and re6, reflecting “an active interest and motivation to learn” and “fitting what is learned into prior knowledge” are easiest to endorse for the respondent group, whilst items that refer to personal reflection, planning and interests seem to be difficult to endorse and thus seem less important for the teachers and researchers in the respondent group. the response pattern of this scale also does not fit with the expectation under the consistency hypothesis. figure 9. person-item map, subset of the responsibility for learning items (re) kjellström et al ! | f l r ! ! 24! table 9 responsibility items by order of endorsement and assumed epistemological level item code item parameter item endorsement order item level* item re2 0.54 1 3 have an active interest in the course and be motivated to learn. re6 0.57 2 4 systematically reflect on what they have learned and how it fits in what they already know. re3 0.81 3 5 take responsibility for one’s own personal development. re1 0.90 4 2 come to classes well prepared, take notes, review the notes afterwards, and complete assigned tasks. re5 1.42 5 4 are aware of their strengths, monitoring their progress and making adjustments when necessary. re4 1.76 6 6 follow their own interests in searching for knowledge. *) please note: the order of these items are not empirically derived but based developmental research literature and work on responsibility issues (kjellström, 2005; kjellström & ross, 2011) for the overall picture we return to figure 3 which shows the person-item map including all items, sorted by relative endorsement difficulty, we can construct a similar table for all the items together. the figure shows a level wide range of endorsement with infit measures indicating good fit to the unidimensional rasch model. the distribution of the levels of the statement is presented in table 10, and shows that when all statements are taken together, levels 4 and 5 are the easiest to endorse and levels 1 and 2 the most difficult. this response pattern implies that the preferred level of epistemic thinking for this group of teachers is constructivist (level 4 and 5) and teaching by telling and focusing on recall and reproduction is actively rejected. table 10 all statements by order of endorsement, based upon figure 6 item level statement ap4 4 be able to solve real life problems by combining knowledge and skills in new ways. bt2 4 inspire the students so that they are motivated to learn. un3 4 see the connections within a larger context. un5 5 realize how different perspectives influence what you see and understand. un4 5 see the underlying assumptions and how these lead to particular conclusions. di3 5 different solutions are illustrated from multiple perspectives. sb4 4 evokes a critical attitude and invites reflection. re2 3 have an active interest in the course and be motivated to learn. re6 4 systematically reflect on what they have learned and how it fits in what they already know. re3 5 take responsibility for one’s own personal development. sb1 2 is well structured, gives the main points and the consensus story of the subject. kjellström et al ! | f l r ! ! 25! di4 5 students and teachers learn together and from each other. ap2 5 look at a problem from multiple perspectives. sb3 5 presents alternative conclusions, shows different interpretations within a subject matter or clarifies different perspectives on the content. re2 3 have an active interest in the course and be motivated to learn. ap5 4 support your own point of view on an issue using evidence and facts. di5 4 opinions are based on and/or supported by evidence. un1 3 be able to apply and use it properly. ap6 6 use your knowledge to make the world a better place for those around you (society, human kind and/or nature). ap3 3 be able to solve familiar problems using what you have learnt. sb6 5 makes students think about the fundamental question of knowledge: "how can we know this?" bt3 3 focus on application of what is learnt. di1 2 the same students don’t dominate. sb5 4 provokes and challenges the students’ pre-conceptions bt4 5 make students question their current view on the world. re5 4 are aware of their strengths, monitoring their progress and making adjustments when necessary. sb2 3 provides many up-to-date examples and examples from practice. bt5 5 help students realize the limitations of current knowledge and understandings. bt1 2 focus on conveying the most important facts and knowledge as clearly as possible re4 6 follow their own interest in searching for knowledge bt6 1 make sure that the whole content of the course is covered. di6 3 you hear all the different views that people have. di2 2 the teacher answers students’ questions. ap1 2 pass one’s exams. un6 2 be able to answer test or exam questions correctly un2 1 you know something by heart 5. discussion summarizing the results of the analysis of psychometric properties of the six scales comprising the edtlq, the confirmatory factor analysis supports the assumption that the six scales together represent a single latent developmental dimension underlying epistemological development that explains the variability in the items, with sufficient infit statistics for all the items. if the assumption was incorrect, the analysis may have resulted in up to six dimensions, one for each issue covered in a scale of the edtlq. as this was not the case, the first research question can be answered positively. examining the factor parameters, with items representing the least sophisticated way of thinking clustering at one end of the dimension and those reflecting more sophisticated epistemological thinking at the other end, the results further seem to imply that kjellström et al ! | f l r ! ! 26! the underlying dimension is somewhat similar to the surface – deep level dimension discussed in earlier works regarding learning strategies (e.g. entwistle & ramsden, 1983; van rossum & schenk, 1984). the rated statements comprising the edtlq have satisfactory levels of separation, meaning that the items are characterized by clearly different levels of endorsement and therefore separate the respondents to an acceptable degree, answering the second research question positively. the edtlq seems to be a satisfactory tool to measure epistemological views on teaching and learning. the analysis of the category characteristic curves for all of the items in the edtlq indicates that there are no response categories that completely overlap, nor are category curves in an unanticipated sequence, indicating that it is not necessary to change the number or sequence of the response categories per item. the design decision to offer a fivepoint rating scale ranging from “least important” to “important” offers clearly separated response categories. this means that the third research question of this study results in a positive answer as well. the at first glance puzzling finding from a unidimensional epistemological development perspective, is the fact that the response group of university staff as a whole endorse relatively sophisticated ways of thinking, namely the items reflecting epistemological development levels within the constructivist learning paradigm (levels 4 and 5, see van rossum & hamer, 2010). this result indicates that for this response group, the items reflecting constructivist learning reflect their views on what is important in learning and teaching the best. as discussed in the introduction, a hierarchical inclusive model would predict that items reflecting less sophisticated levels of thinking would be easier to endorse than those reflecting sophisticated ways of thinking, and items reflecting the most sophisticated levels of thinking being endorse by the fewest of the response group. however, contrary to expectations, in this study it is the items that reflect the least sophisticated levels of thinking that are rejected, i.e. are not endorsed, by the larger majority of the response group. whilst this outcome seems to indicate that the underlying model of epistemological development is incorrect, and so implies that answer to the fourth research question may be negative, there is reason to believe that the response instructions – to indicate how well the statement reflected an important aspect of learning and teaching – led to a different response pattern. as introduced before, kegan (1994) discusses the consistency hypothesis and states that when circumstances prevent respondents to express their highest level of epistemic beliefs, the pressure felt to “regress” is experienced as unhappiness because “we do not feel ‘like ourselves’” (kegan, 1994, p. 372). something similar is discussed by van rossum & hamer (2010) when they describe what they refer to as disenchantment and its counterpart nostalgia. disenchantment refers to the feeling of disillusionment that students experience when they are exposed to a teaching-learning environment characterized by significantly less epistemological sophistication than their own. this can lead to rebellion or despair (e.g. yerrick, pedersen & arnason, 1998; lindblom-ylänne & lonka, 1999). nostalgia refers to the wistful hankering to a more traditional learning-teaching environment expressed by students when they are over asked and required to function at an epistemic level too far beyond their understanding (van rossum & hamer, 2010, pp 415 426). considering the consistency hypothesis, and the request to rate items to reflect the most important level of epistemological development in learning and teaching, it is then in fact not surprising to see that respondents reject items reflecting less sophisticated ways of thinking. indeed, in van rossum and hamer (2010) many examples are given of students expressing their active rejection of a way of knowing or perception of learning that they feel they have outgrown. something similar may be taking place in the mind of teachers. however, this does mean that the edtlq results of a group reflect the epistemological sophistication level that the majority of the respondent group feel is important to learning and teaching, and when used in a way similar to here does not in fact reflect the epistemological development of a single respondent. to establish what individual respondents feel is important to learning and teaching, reflecting the different preferred levels of epistemological development present in a classroom or lecture hall, will require additional study and analysis, and perhaps different response instructions to those completing the questionnaire. examining the six scales of the edtlq introduced here, which reflect six different contexts in which an underlying epistemic development may express itself, it seems that four of these five scales behave as can be expected under the consistency hypothesis. these are the scales referring to discussion, application, understanding and good teaching. the scales for views on a good study book and the responsibility for learning scale do not seem to fit the expected pattern. kjellström et al ! | f l r ! ! 27! in the case of the good study book scale, there are two possible explanations that need to be considered and require further study. firstly, the underlying empirical data for this scale is relatively recent and the analysis of the original data has not been confirmed by further data. secondly, this scale may provide some insight into the extent to which the teacher responses in this study may be affected by discourse and social desirability. in the pedagogical literature aimed at teachers, significant attention is given to describing constructivist interpretations of application, understanding, the role of discussion in arriving at new knowledge and perspectives and characteristics, providing a rich vein of discourse potentially affecting teacher responses. to the current authors knowledge there is very little of similar discourse available defining characteristics of a good study book. if so, the mixed pattern may reflect the teachers’ epistemic thinking more faithfully than the other scales. regarding the responsibility scale, again there are at least two reasonable explanations. firstly, the scale was developed by the use developmental responsibility research from partly another domain (namely responsibility for health) which may illustrate how difficult it is to transfer ideas in one domain to another, and that there may be domain specific patterns (fischer & pruyne, 2003). secondly, and perhaps more importantly, the responsibility items were allocated developmental levels using a model to which they have no empirical link. hence, the items were interpreted by two of the authors and were allocated a developmental level that may not reflect the actual position on the assumed epistemological development trajectory. without empirical evidence linking the responsibility items to any of the other scales, it is less clear which level to assign as previous research has shown that a particular statement may be attractive to people different stages of development (ross, 2008; loevinger & hy, 1994). remains then the issue that the respondent group of teachers seems to prefer items reflecting relatively sophisticated levels of epistemological development that is usually not observed to this degree and is quite sophisticated even for teachers in higher education (for a review of literature see van rossum & hamer, 2010 chapter 5, 2012). whilst this seems to indicate that the teachers in this study are characterized by a relatively sophisticated view on learning and teaching, there is another possible explanation: the interpretation of the instructions for rating the statements. where the example used to illustrate how respondents were to rate the statements referred to a personal view (‘i know i have done a good job when...’) the items of the five scales were formulated in more general way (‘discussion during a course are good when …’). this means that we cannot be sure the instruction was interpreted by the teachers as a request to rate each statement in importance to their personal view on learning, and therefore that their ratings reflect their personal preference. it is entirely possible that the results reflect ‘more socially constructed discourse and less personal … understanding’ (van rossum & hamer, 2010, p. 449). in other words, qualitative responses of teachers are often thick with socially acceptable discourse that at times obscures their personal understanding or preference of the teaching-learning environment (säljö, 1994, 1997). the pilot of the edtlq was performed in sweden, where the official written curricula focus on skills such as discussing, analyzing and synthesizing rather than on facts and learning by heart. the request to rate statements with regard to their importance to learning and teaching in general may well have activated a preference for items reflecting these written and therefore socially acceptable curricula, above a preference for items reflecting their own beliefs or the reality of the taught curricula. if this latter interpretation is correct, reformulating the instructions towards an importance to a respondent’s own teaching practice may result in a different outcome. based on theory and earlier experimental findings (van rossum & hamer, 2012, 2013), the findings for the teachers as a group may move slightly towards levels 3 and 4. in addition, the results for students would reflect a significantly less sophisticated preference. as students are less well versed in the academic discourse with regard to teaching, the ambiguous instructions may not have the same effect, which means that the preferred items for the freshmen students in particular would differ considerably with a significant shift towards items reflecting levels 2 and perhaps 3. although the results of this study are promising, further analyses as well as improvements to the survey, in particular to the rating instruction, and additional data collection are necessary to definitively answer the final research question in this study. kjellström et al ! | f l r ! ! 28! 5.1 methodological reflections the current study indicates that the unidimensional rasch model fits the set of scales regarding epistemological preference for a group of university staff. a similar analysis needs to be undertaken for the two groups of students that participated, freshmen and senior (third year undergraduate) students separately, as these two groups may not interpret the scale items in the same way as the teachers and researchers who participated in this study. to establish if the edtlq is a viable tool to measure individual epistemological development, further analysis is required to identify specific response patterns for individual participant or participant groups. to clarify if the results reflect the personal preference in teaching in learning it is necessary to clarify the instructions regarding the rating of the statements. whilst the categories ranging from least important to most important work well, the instruction above each set of items should include a reference to this personal preference, e.g. “please rate each statement from 1 (least important) to 5 (most important) to your own personal views regarding an ideal learning and teaching environment.” remains the issue of the potential social desirable bias in the teacher responses. when surveying teachers in particular, using inventory statements it will remain difficult to disentangle the use of social acceptable discourse from personal understanding, and as such it is recommended to supplement the edtlq with other data collection methods that provide more opportunity for teachers to structure their own response and so provide a more clear window into their personal views. an interesting finding is that the response group of university teachers seems to be attracted to the more advanced statements, and we probably would not have the same results with high school teachers who may prefer items that reflect levels 2 and 3 (e.g. brickhouse, 1989, 1990; martens, 1992; maor and taylor, 1995; hashweh, 1996) which fits the experience in the adult development research field, but hitherto has not been shown statistically. in this sense this result is the most interesting as it enhances the knowledge base within adult development regarding how the statements are endorsed within each item and the whole collection of items. whilst the main finding, that the items that are easiest to endorse within each scale reflect the midrange levels of epistemological sophistication of the six stage developmental model of learning-teaching conceptions developed by van rossum and hamer (see table 10) seems contradictory for readers not familiar with adult development theories, both the consistency hypothesis (kegan, 1994) and the experience in adult development of people endorsing values at their own stage and those slightly more advanced would, is, given the respondent group, exactly what one might expect. on a more detailed level, there may be issues with individual items or scales. in table 10 we see that level 4-5 statements are most easily endorsed. examining the individual items more closely, it is interesting to note in this table is that among the level 5 statements some seem easy to endorse and others more difficult. three statements that are all about taken perspectives are easily endorsed: “realize how different perspectives influence what you see and understand”, “see the underlying assumptions and how these lead to particular conclusions”, and “different solutions are illustrated form multiple perspectives”. two level 5 statements that require a more critical awareness and questioning, are more difficult do endorse: “make students questions current view of the world” and “help students realize the limits of current knowledge and understanding”. a possible explanation would be that the easier to endorse items can both be endorsed by respondents viewing learning and teaching from a multiplistic perspective (l3) where all opinions are equally valid (perry, 1970), and by respondents viewing learning and teaching from a relativist perspective (l5), where opinions can differ in validity and in range of supporting evidence, but judging opinions is less acceptable, as in this way of thinking the focus is not on uncovering the most valid opinion, but on empathically understanding the point of view of others (van rossum & hamer , 2010). this means that the easily endorsed items that describe multiple perspective taking need amending, to make them more clearly either level 3 or level 5. further, there are some individual items that seem to not fit the model. in particular this is the case for the items regarding the responsibility for learning in particular. as is indicated above, these items were not developed using empirical data from van rossum and hamer (2010). these items seem to reflect the swedish written curriculum and as such are probably part of the socially acceptable discourse for the respondents in this study. in addition, some items are currently assigned to the lowest level where they apply, e.g. re5 “are aware of their strengths, monitoring their progress and making adjustments when necessary” kjellström et al ! | f l r ! ! 29! reflecting meta-cognition and self-monitoring, whilst these items could well be endorsed by respondents with more sophisticated levels of thinking. finally, there could be cultural differences in understanding the items. van rossum and hamer collected their data from dutch students. although in their 2010 review, van rossum and hamer explicitly address cultural differences with regard to their model, the understanding of individual items could of course differ. such cultural differences could include the afore mentioned emphasis in the written curricula on skills such as discussing, analyzing and synthesizing rather than focusing on facts and recall, but perhaps there are other less obvious cultural influences that will only become manifest if the edtlq is piloted in other countries and with other respondent groups. an obvious choice given the countries represented by the authors would be an english language pilot, and perhaps a translation into portuguese. in addition, respondent groups could include teachers and students in secondary education. keypoints the edtlq measures a single latent developmental dimension underlying epistemological development the response alternatives showed satisfactory levels of separation which means that the alternatives that are rated do not need to be altered the edtlq successfully includes items reflecting rare complex epistemic thinking not captured in similar questionnaires before as analysed here, the edtlq results reflect the dominant epistemic position of the response group. in this study, the response group (he teachers) seems to prefer items reflecting relatively sophisticated levels of epistemological development that is usually not observed even in higher education, which has implications for construction and instructions of questionnaires. the edtlq has a potential to be used as an apt tool for teachers to monitor the development of students as well as to offer professional development opportunities to the teachers acknowledgments we would like to thank kristian stålne, university of malmö, sweden for his valuable contribution to constructing the items and collecting the data. further, we would like to thank the reviewers of flr for their insightful comments which helped to make our paper more concise and informative. references andrich, d. (1978). a rating formulation for ordered response categories. psychometrika 43(4), 561–574. andrich, d. (1988). rasch models for measurement. newbury park, ca: sage. andrich, d. (2004). controversy and the rasch model: a characteristic of incompatible paradigms? medical care, 42(1), 7 – 16. andrich, d. (2011). rating scales and rasch measurement. expert review pharmacoeconomics outcomes reserach, 11(5), 571–585. doi: 10.1586/erp.11.59 arum, r., & j. roksa, (2011). academically adrift: limited learning on college campuses. chicago, city, 2 letter state: university of chicago press. barzilai, s., & eshet-alkalai, y. (2015). the role of epistemic perspectives in comprehension of multiple author viewpoints. learning and instruction, 36, 86 – 103. doi: 10.1016/j.learninstruc.2014.12.003 kjellström et al ! | f l r ! ! 30! barzilai, s., & weinstock, m. (2015). measuring epistemic thinking within and across topics: a scenariobased approach, contemporary educational psychology, 42, 141-158. doi: 10.1016/j.cedpsych.2015.06.006 baxter magolda, m.b. (1992). knowing and reasoning in college. san francisco, ca: jossey-bass. baxter magolda, m.b. (2001). making their own way. sterling, va: stylus publishing. belenky, m.f., clinchy, b.m., goldberger, n.r., & tarule, j.m. (1986 ed. 1997). women’s ways of knowing: the development of self, voice and mind. new york: basic books. bentler, p.m. (1990). comparative fit indexes in structural models. psychological bulletin, 107(2), 238-246. doi: dx.doi.org/10.1037/0033-2909.107.2.238 bentler, p.m., & bonett, d.g. (1980). significance tests and g.oodness of fit in the analysis of covariance structures. psychological bulletin, 88, 588–606. doi: 10.1037/0033-2909.88.3.588 bond, t. g., & fox, c. m. (2001). applying the rasch model: fundamental measurement in the human sciences. mahwah, nj us: lawrence erlbaum associates publishers. doi: 10.1111/j.17453984.2003.tb01103.x. bråten, i. and strømsø, h.i. (2004). epistemological beliefs and implicit theories of intelligence as predictors of achievement goals. contemporary educational psychology, 371-388. doi: 10.1016/j.cedpsych.2003.10.001 brickhouse, n.w. (1989). the teaching of the philosophy of science in secondary classrooms: case studies of teachers personal theories. international journal of science education, 11(4), 437-449. doi: 10.1080/0950069890110408. brickhouse, n.w. (1990). teachers’ beliefs about the nature of science and their relationship to classroom practice. journal of teacher education, 41(3), 53-62. doi: 10.1177/002248719004100307. browne, m.w., & cudeck, r. (1993). alternative ways of assessing model fit. in: bollen, k. a., long, j. s. (ed.) testing structural equation models (pp.136 -162). newbury park: sage. commons, m.l., & ross, s. n. (2008a) special issue: postformal thought and hierarchical complexity. world future: journal of general evolution, 64(5-7): p. 297-562. doi: 10.1080/02604020802301121. commons, m.l., & ross, s. n. (2008b), what postformal thought is, and why it matters. world future: journal of general evolution, 64(5-7): p. 321-329. doi: 10.1080/02604020802301139. commons, m.l. (1989) adult development. vol. 1, comparisons and applications of developmental models. new york: praeger. xii. commons, m.l. (1990) adult development. vol. 2, models and methods in the study of adolescent and adult thought. new york: praeger. dawson, t.l. (2006), the meaning and measurement of conceptual development in adulthood. in c.h. hoare handbook of adult development and learning, oxford university press: oxford. p. 433-454. doi: 10.1558/jate.v6i1.98 dawson, t., goodheart, e., draney, k., wilson, m., & commons, m. (2010). concrete, abstract, formal, and systematic operations as observed in a 'piagetian' balance-beam task series. journal of applied measurement, 11(1), 11-23. dawson, t. l., xie, y., & wilson, m. (2003). domain-general and domain-specific developmental assessments: do they measure the same thing? cognitive development, 18, 61-78. doi: 10.1016/s0885-2014(02)00162-4. demick, j., & andreoletti, c. (2003) handbook of adult development. the springer series in adult development and aging. 2003, new york: springer. demetriou, a., & kyriakides, l. (2006). the functional and developmental organization of cognitive developmental sequences. british journal of educational psychology, 76(2), 209-242. doi: 10.1348/000709905x43256 embretson, s.e., & reise, s. p. (2000). item response theory for psychologists. london: erlbaum. entwistle, n.j., & ramsden, p. (1983). understanding student learning. new york, ny: nichols. epskamp, s. (2014). semplot: path diagrams and visual analysis of various sem packages' output. r package version 1.0.1. http://cran.r-project.org/package=semplot fischer, k.w., & pruyne, e, (2003) reflective thinking in adulthood: emergence, development, and variation. in j. demick & c. andreoletti handbook of adult development, kluwer academic/plenum: new york. p. 169-198. glas, c.a. (2007). multivariate and mixture distribution rasch models. new york: springer-verlag. kjellström et al ! | f l r ! ! 31! gow, l., & d. kember, d. (1990). does higher education promote independent learning? higher education, 19, p. 307-322. doi: 10.1007/bf00133895. golino, h.f., gomes, c.m.a., commons, m.l., & miller, p. (2014). the construction and validation of a developmental test for stage identification: two exploratory studies. behavioral development bulletin, 19, 37-54. doi: 10.1037/h0100589 hambleton, r. k. (2000). response to hays et al and mchorney and cohen: emergence of item response modeling in instrument development and data analysis. medical care, 38(9), ii-60. hamer, r., & van rossum, e.j. (2010). de zes talen in het onderwijs: communicatie en miscommunicatie [the six languages in education: communication and miscommunication]. onderwijsvernieuwing, 17 (november). hamer, r., & van rossum, e.j. (2016). students’ conception of understanding and its assessment. in e. cano & g. ion (eds.), innovative practices for higher education assessment and measurement. (pp. 141-162). heshey, pa: igi global. hasweh, m.z. (1996). effects of science teachers’ epistemological beliefs in teaching. journal of research in science teaching, 33(1). 47-63. hibbard, j., collins, p., mahoney, e., & baker, l. (2010). the development and testing of a measure assessing clinician beliefs about patient self-management. health expectations: an international journal of public participation in health care & health policy, 13(1), 65 72. doi: 10.1111/j.13697625.2009.00571.x. hoare, c.h., (2006). handbook of adult development and learning. oxford, uk: oxford university press. xviii, 579 p. hornik, k. (2014). the r faq. retrieved from http://cran.r-project.org/doc/faq/r-faq.html hu, l.t., & bentler, p.m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives. structural equation modeling, 6(1), 1-55. doi: 10.1080/10705519909540118. irving, p.w., & sayre, e.c. (2013). upper-level physics students’ conceptions of understanding. paper presented at 2012 physics education research conference. in aip conference proceedings (vol. 1513, pp. 98-201). doi:10.1063/1.4789686. kegan, r. (1994). in over our heads: the mental demands of modern life. cambridge, ma: harvard university press. king, p.m., & kitchener, k.s. (1994). developing reflective judgment. san fransisco, ca: jossey-bass publishers. king, p.m., & kitchener, k.s. (2004). reflective judgment: theory and research on the development of epistemic assumptions through adulthood. educational psychologist, 39(1), p. 5-18. doi: 10.1207/s15326985ep3901_2. kjellström, s. (2005). responsibility health and the individual [ansvar, hälsa och människan – en studie av idéer om individens ansvar för sin hälsa]. linköping studies in arts and science, nr 318. linköping university. kjellström, s. & s.n. ross (2011). older persons’ reasoning about responsibility for health: variations and predictions”, international journal of aging and human development, 73(2), 99-124. kjellström, s. & sjölander, p. (2014). the level of development of nursing assistants’ value system predicts their views on paternalistic care and personal autonomy. international journal of ageing and later life, 9(1), 35-68. doi: 10.3384/ijal.1652-8670.14243 kohlberg, l. (1984). essays on moral development: vol. 2. the psychology of moral development san francisco: harper & row. kuhn, d. (1991). the skills of argument. cambridge, uk: cambridge university press: linacre j. m. (2002). what do infit and outfit, mean-square and standardized mean? rasch measurement transactions, 16 (2), 878. lindblom-ylänne, s., & lonka, k. (1999). individual ways of interacting with the learning environment – are they related to study success? learning and instruction, 9, 1-18. doi: 10.1016/s09594752(98)00025-5. linder, c.j. (1992). is teacher-reflected epistemology a source of conceptual difficulty in physics? international journal of science education, 14(1), 111-121. doi: 10.1080/0950069920140110. kjellström et al ! | f l r ! ! 32! loevinger, j., & blasi, a. (1976). ego development: conceptions and theories (1. ed.). san francisco: jossey-bass. loevinger, j., & hy, l. x. (1996). measuring ego development (2. ed.). mahwah, n.j.: l. erlbaum associates.francisco: jossey-bass. mair p., & hatzinger, r. (2007a). extended rasch modeling: the erm package for the application of irt models in r. journal of statistical software, 20(9), 1–20. mair p., & hatzinger, r. (2007b). cml based estimation of extended rasch models with the erm package in r. psychology science, 49, 26–43. mair, p., hatzinger, r., & maier, m.j. (2014). erm: extended rasch modeling. r package version 0.15-4. http://erm.r-forge.r-project.org/. maor, d. and taylor, p.c. (1995). teacher epistemology and scientific inquiry in computerized classroom environments. journal of research in science teaching, 32(8). 839-354. doi: 10.1002/tea.3660320807. martens, m.l. (1992). inhibitors to implementing a problem-solving approach to teaching elementary science: case study of a teacher in change. school science and mathematics, 92(3). 150-156. mishra, p., & kereluik, k. (2011). what 21st century learning? a review and a synthesis. in c. d. maddux, m. j.koehler, p. mishra, & c. owens (eds.), proceedings of society for information technology & teacher education international conference 2011 (pp. 3301–3312). chesapeake, va: aace. newble, d.i., & clarke, r.m. (1986). the approaches to learning of students in a traditional and an innovative problem-based medical school, medical education,20, 267-273. doi: 0.1111/j.13652923.1986.tb01365.x. panayides, p., robinson, c., & tymms, p. (2010). the assessment revolution that has passed england by: rasch measurement. british educational research journal, 36(4), 611 626. doi: 10.1080/01411920903018182. paulsen, m.b. and feldman, k.a. (2005). the conditional and interaction effects of epistemological beliefs on the self-regulated learning of college students: motivational strategies. research in higher education, 46(7), 731-768. doi: 10.1007/s11162-004-6224-8. perkins, d. (1993). teaching for understanding. american educator: the professional journal of the american federation of teachers, 17 (3), 28-35. document available online http://www.exploratorium.edu/ifi/resources/workshops/teachingforunderstanding.html, 24 july 2011. perry, w.g. (1970). forms of intellectual and ethical development in the college years: a scheme. new york, ny: holt, rinehart & winston. piaget, j. (1954). the construction of reality in the child. new york: basic books. pornprasertmanit, s., miller, p., schoemann, a., & rosseel, y. (2014). semtools: useful tools for structural equation modeling.. r package version 0.4-6. http://cran.r-project.org/package=semtools qian, g. and alvermann, d.e. (2000). relationship between epistemological beliefs and conceptual change lea.rning. reading &writing quarterly, 16, 59-74. doi: 10.1080/105735600278060. qian, g. and pan, j. (2002). a comparison of epistemological beliefs and learning from science text between american and chinese high school students. in. b.k. hofer and p.r. pintrich (eds.). personal epistemology. (pp. 365-385). mahwah, new jersey u.s.a.: lawrence erlbaum ass. r core team (2014). r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. url http://www.r-project.org/. rasch, g. (1960). probabilistic models for some intelligence and attainment tests. danish institute for educational research, copenhagen. raykov, t. (2012). scale construction and development using structural equation modeling. in r. h. hoyle (ed.), handbook of structural equation modeling (pp. 472-494). new york: guilford. richardson, j.t.e. (2012). erik jan van rossum and rebecca hamer – the meaning of learning and knowing. book review. higher education, 64 (5), pp 735-738. doi: 10.1007/s10734-012-9518-3. ross, s. n. (2008). postformal (mis)communications. world future: journal of general evolution, 64 (5-7), 530-535. doi: 10.1080/02604020802303952. rosseel, y. (2012). lavaan: an r package for structural equation modeling. journal of statistical software, 48(2), 1-36. url http://www.jstatsoft.org/v48/i02/. säljö, r. (1994). minding action; conceiving of the world versus participating in cultural practices. nordisk pedagogik, 14(2), 71-80. kjellström et al ! | f l r ! ! 33! säljö, r. (1997). talk as data and practice – a critical look at phenomenographic inquiry and the appeal to experience. higher education research & development, 16, 173-190. doi: 10.1080/0729436970160205. schommer, m. (1990). effects of beliefs about the nature of knowledge on comprehension. journal of educational psychology, 82(3), 498-504. doi: 10.1037/0022-0663.82.3.498. schommer, m. (1994). an emerging conceptualization of epistemological beliefs and their role in learning. in garner, r and alexander, p.a. (eds.). beliefs about text and instruction with text (pp. 2540). hillsdale: lawrence erlbaum associates. schommer, m. (1998). the influence of age and education on epistemological beliefs. british journal of educational psychology, 68, 551-561. doi: 10.1111/j.2044-8279.1998.tb01311.x. schreiber, j.b. and shinn, d. (2003). epistemological beliefs of community college students and their learning processes. community college journal of research and practice, 27, 699-709. doi: 10.1080/713838244. sjölander, p., n. lindström, a. ericsson, s. kjellström, 2014, a pattern recognition method for disclosing different levels of value system from questionnaire data, behavioral development bulletin, 19(3), 114127. doi: 10.1037/h0100596. strijbos, j., engels, n. & struyven, k. (2015). criteria and standards of generic competences at bachelor degree level: a review study. educational research review, 14, 18-32). doi:10.1016/j.edurev.2015.01.001. van rossum, e.j., & hamer, r.n. (2010). the meaning of learning and knowing. rotterdam, netherlands: sense publishers. van rossum, e.j., & hamer r., (2012). analysing higher education teachers’ learning–teaching conceptions with a model of student thinking. paper presented at the aera conference, vancouver, canada. van rossum, e.j., & hamer, r. (2013). the relationship between students’ conception of good teaching and their views on a good textbook. paper presented at the earli conference, munich, germany. van rossum, e.j., & schenk, s.m. (1984). the relationship between learning conception, study strategy and learning outcome. british journal of educational psychology 54, 73-83. doi:10.1111/j.20448279.1984.tb00846.x. vermunt, j.d. (1996). metacognitive, cognitive and affective aspects of learning styles and strategies: a phenomenographic analysis. higher education, 31, 25-50. doi: 10.1007/bf00129106. voogt, j, erstad, o., dede, c., & mishra p. (2013) challenges to learning and schooling in the digital networked world of the 21st century. journal of computer assisted learning, 29, 403–413, doi: 10.1111/jcal.12029. west, e.j., (2004) perry's legacy: models of epistemological development. journal of adult development, 112, 61-70. doi: 10.1023/b:jade.0000024540.12150.69. wright, b. d., linacre, j. m., gustafson, j. e., & martin-lof, p. (1994). reasonable mean-square fit values. rasch measurement transactions, 8(3), 370. yerrick, r.k., pedersen, j.e., & arnason, j. (1998). “we're just spectators”: a case study of science teaching, epistemology, and classroom management, science education, 82(6), 619-648. doi: 10.1002/(sici)1098-237x(199811)82:6<619::aid-sce1>3.0.co;2-k. kuisma et nokelainen frontline learning research vol.6 no. 2 (2018) 1 19 issn 2295-3159 effects of progressive inquiry on cognitive and affective learning outcomes in adolescents’ geography education merja kuismaa, petri nokelainenb auniversity of tampere, finland btampere university of technology, finland article received 19 juni 2017/ revised 5 november 2017 / accepted 16 july 2018 / available online 30 july abstract adolescents need skills to acquire information and compare, analyze, transform, and experiment with knowledge. however, little research has been conducted on the content and pedagogical practices that are necessary to achieve these skills. this article seeks to contribute to this discussion because geography enables the attainment of the so-called higher-order thinking skills, and the progressive inquiry model provides suitable pedagogical practices. this study provides empirical evidence on the effects of the progressive inquiry teaching method and learning model on cognitive and affective learning outcomes. this paper focuses on learning outcomes among 253 finnish middle and upper secondary school students. this comparison between different developmental stages reveals the effects of the teaching and learning method in question. the results indicate that the progressive inquiry method improves cognitive learning results at both educational levels in the context of geography education. the research provides evidence that older students benefit more from the learning model. additionally, the self-regulated learning skills that the students possess at the beginning of the course do not affect their cognitive learning outcomes. progressive inquiry clearly enhances the motivation levels of middle school students; however, the effect on the motivation level was more ambiguous among the upper secondary students. keywords: active learning; computer-supported collaborative learning; inquiry learning; learning outcomes; progressive inquiry info corresponding author email merja.kuisma@uta.fi doi: 10.14786/flr.v6i2.309 1. introduction inquiry-based learning, such as the progressive inquiry method, requires a great deal of effort from students and teachers, especially when information and communication technology (ict) is used as a learning tool (cerratto-pargman, järvelä, & milrad, 2012; winne, 1995). nevertheless, the benefits of inquiry-based learning surpass its drawbacks. one of these benefits is the enhancement of essential citizenship skills in the 21st century societies (banchi & bell, 2008; hakkarainen, palonen, paavola, & lehtinen, 2004). although the novelty of many of these skills is controversial, certain skills are commonly seen as necessary for coping with and succeeding in a knowledgeand information-based society (pauw, 2015). learners in 21st century societies need skills to acquire information, as well as inquire, analyze, transform, construct, compare, and experiment with knowledge (e.g., cerratto-pargman et al., 2012). despite studies that scrutinize the definitions of these skills, little research has focused on the content that should be used to teach these skills (pauw, 2015). this article seeks to contribute to this discussion by introducing geography as a context for higher-order thinking and progressive inquiry as a pedagogical model. research was conducted using a carefully planned quasi-experimental design that compared intervention and control groups, and investigated intact groups to reduce the reactive effects of the experimental procedure, thus increasing external validity (see dimitrov & rumrill, 2003). since the first author of this paper was in a position to act immediately when data was missing (e.g., a survey response), the data was completed right away by asking the student to fill in the missing form. the research design was made in such a way that it enabled the researchers to investigate two different developmental stages of learners in order to find differences in learners’ self-regulated learning skills, especially related to motivation, and their effects on learning outcomes. since self-regulated learning processes are developmental by nature (loyens, magda, & rikers, 2008), we compared two data sets: one from middle school and one from upper secondary school. according to wolters, pintrich and karabenic (2003, p. 5), “self-regulated learning is an active, constructive process whereby learners set goals for their learning and then attempt to monitor, regulate, and control their cognition, motivation, and behavior, guided and constrained by their goals and the contextual features in the environment.” based on an analysis of major research trends (wolters & taylor, 2012), it encompasses four phases (forethought/planning, monitoring, control/management, reaction/reflection) and four related areas of self-regulation (cognition, motivation, behavior, context) for each phase. in the current study, our focus is on the motivational aspects of self-regulated learning in control (or management) phase (pintrich, 2004; wolters, pintrich, & karabenic, 2003; wolters & taylor, 2012; zimmerman, 2000; zimmerman & schunk, 2008). one of the key features of geography as a school subject is that it provides content that can mediate the attainment of higher-order thinking skills, such as analyzing, synthesizing, and problem solving (leat, 1997; nagel, 2008; pauw, 2015). therefore, geography differs from other school subjects due to its analytic and synthetic nature. learners are challenged with questions about who, why, where, when and how this happen, as well as finding causal relationships within and between phenomena of human and physical geography (leat, 1997; 1998). as geography provides excellent possibilities for improving learners’ 21st century skills, courses of geography were chosen for this study. the following sections present the theories that underlie the research design of this study, followed by a description of the methods. 1.1 geographical thinking skills geography plays a central role in school curricula that educates adolescents for the future (pauw, 2015). geography deals with complex issues, such as globalization and global warming, which are included in the geography curricula of middle schools and upper secondary schools (finnish national board of education, 2015; pauw, 2015). moreover, thinking geographically is a skill that helps adolescents understand where something is, how the location affects its characters, and its relations to other phenomena (nagel, 2008). adolescents need knowledge about different cultures in order to understand global and local issues. in our study, pre-tests and post-tests were designed according to bloom’s taxonomy to determine students’ abilities to recollect factual knowledge, analyze information, and thinking synthetically (bloom, hastings, & madaus, 1981). moreover, multiple choice questions were designed to measure students’ abilities to recollect factual knowledge, and open-ended questions were created to measure their understanding and abilities to apply, analyze, synthesize, and assess knowledge. in order to succeed in their future careers and personal lives, adolescents require innovative and creative thinking skills, problem solving skills, and communication and collaboration skills that can be taught via geographic studies (leat, 1997; nagel, 2008). these skills were implemented into the design of the geography courses of this study. according nagel (2008, p. 354), “geography provides students an inexhaustible context for creativity in an interdependent world.” additionally, geography can teach metacognitive skills, thus making the learning processes visible to the learner, which can be transferred to other subject domains (leat, 1997; 1998). in this study, middle and upper secondary school students studied human and physical geography. the middle school students studied the geographical phenomena in the european context, while the upper secondary school students studied these phenomena from a global perspective. both age groups were expected to execute geographical thinking, and analyze and synthesize different phenomena collaboratively through a type of inquiry-based learning called progressive inquiry. 1.2 progressive inquiry as a teaching and learning model progressive inquiry is a teaching and learning model that falls under the umbrella of the active learning approach where students construct knowledge themselves (prince & felder, 2006). active learning activities involve learner-centered teaching strategies that seek to activate students, while progressive inquiry introduces students to a new way of creating knowledge that resembles the scientific inquiry process (paavola & hakkarainen, 2005). this is a cyclic process of creating a context for learning, determining the research questions, constructing working theories, seeking and deepening knowledge, conducting a critical assessment of knowledge advancement, and sharing expertise (muukkonen, hakkarainen, & lakkala, 1999). the theoretical framework of progressive inquiry consists of the following theories and models: (1) the knowledge-building theory of intentional learning and expertise (bereiter, 2002; scardamalia, 2002; scardamalia & bereiter, 1994); (2) the interrogative model of inquiry (hintikka, 1982, 1985), (3) the theory of expansive learning (engeström, 1999), and (4) the model of knowledge-creating companies (nonaka & takeuchi, 1995). the knowledge-building processes involved in the progressive inquiry model include monologic (“knowledge acquisition”), dialogic (“participation”) and trialogic (“knowledge creation”) levels of learning (paavola & hakkarainen, 2005). knowledge acquisition represents a subjective or mental view of human learning, participation represents an inter-subjective or social view of learning, and knowledge creation represents an interaction between individuals, communities, and shared knowledge-laden artifacts being developed (hakkarainen, 2009). the progressive inquiry approach underlines the key role of the active learner and collaboration when directing one’s behavior. it also presupposes that inquiry is a question-driven process of understanding what is needed for knowledge creation, that mediating artifacts can induce learning, and that learning is an expansive process where activities produce new activities. progressive inquiry provides learners with a high level of autonomy. to be successful in such a learning process, the learner requires various forms of support from the teacher, as well as self-regulated learning skills (winne, 1995). according to zimmerman (1998), the learner’s behavior should be intrinsic rather than extrinsic (e.g., regulated by the teacher, parents, or peers) because it enables self-regulated learning to take place. self-regulated learning skills become fortified as the learner gains experience with controlling his or her cognition, motivation, behavior (especially effort regulation), physical learning environment, and study schedule (winne, 1995; wolters, pintrich, & karabenick, 2003). self-regulatory skills are not fixed individual characteristics as they involve contextual, developmental, and individual constraints (pintrich, 1999; wolters, pintrich, & karabenick, 2003). when participating in collaborative learning, regulative processes can become even more complex through the interaction between self-regulative, co-regulative, and socially shared regulation processes (järvelä & hadwin, 2013). from the competence research perspective, all these three levels of learning contribute to the holistic development of the individual (cognitive, metacognitive) and subject-related (conceptual, functional) competences (le deist & winterton, 2005). self-regulation relates to metacognitive (“learning to learn”) competence development (nokelainen, kaisvuo, & pylväs, 2016) and is, therefore, a prerequisite of inquiry-based pedagogy because it is needed at the monologic level of knowledge-building processes in technology-mediated learning. co-regulation relates to the dialogic level and socially shared regulation is needed at the trialogic level. this study focuses on the motivational aspects of regulation. self-regulated learning processes are developmental by nature (loyens, magda, & rikers, 2008); therefore, we aim to compare the effects of progressive inquiry on the learning outcomes of two different age groups of adolescents. the progressive inquiry teaching and learning method—with its collaborative way of working—suits the pedagogical framework of geography education because geographical thinking involves question-making, multiple resources to seek answers, analytic and synthetic thinking, and creative problem solving. our first research question investigates whether this is true and whether it applies to different age groups of adolescents by scrutinizing the cognitive learning outcomes of the method. the second research question investigates the level of self-regulation skills and its relationship with cognitive learning outcomes. finally, we scrutinize the motivation level of the learners during the geography course. 1.2 research questions and hypotheses we wanted to know how progressive inquiry affected the attainment of cognitive learning goals and motivation levels of learners. motivation level is regarded as an affective learning ability. we also aimed to investigate the relationship between self-regulated learning skills and cognitive learning outcomes. another study used qualitative interviews to examine what kinds of learning narratives existed among learners (kuisma, 2018). in this study, we addressed the following research questions: rq1. does progressive inquiry affect cognitive learning results when studying geography at different levels of the education system? (1a) does the progressive inquiry method affect cognitive learning results when studying geography? (1b) do the cognitive learning results differ between middle school and upper secondary school students when they are taught using the progressive inquiry method? rq2. how do cognitive learning outcomes relate to self-regulated learning, specifically regulation of motivation? rq3. does the progressive inquiry method affect the motivation levels of students? first, we investigated if there was any statistically significant difference between the intervention group (progressive inquiry method and specific ict) and control group’s cognitive learning outcomes by testing the null hypothesis: h0: coloint = colocontr (colo= cognitive learning outcome; int= intervention group; contr = control group). we tested the data set of middle school students (sample 1) and upper secondary school students (sample 2) separately. second, we investigated the relationship between self-regulated learning skills, and teaching and learning methods with cognitive learning outcomes. we tested the null hypothesis, that there are no differences in variances between different groups of teaching and learning method, and levels of skills in self-regulated learning when investigated these groups’ relations to post-test results. additionally, we tested the null hypothesis h0: β1, β2, β3 = 0 for the regression equation colo = β0 + β1*tlm + β2*srl +β3*tlm*srl + ε (colo = cognitive learning outcome as post-test score, β0 = constant, tlm = teaching and learning method, srl = self-regulated learning skills, ε = random error). third, we investigated the effect of the progressive inquiry method on the level of motivation by null hypotheses that there are no differences in variances between different teaching and learning groups in motivation level. we investigated this matter at three different time points (t0, t1, and t2). 2. method 2.1 participants all of the participants (n = 253) participated voluntarily, and their schools were located in an urban area in central finland. we sent out invitations via email to participate in this study to all the middle school and upper secondary geography teachers and head masters in pirkanmaa region (area in central finland), and accepted all the volunteers. one hundred fifty-two informants (sample 1) studied at the middle school of the comprehensive school, and 101 informants (sample 2) studied at the upper secondary school. middle school participants of this study took part in a course on the human and physical geography of europe, the course topic for the upper secondary school students was geographical hazards. seventy-five students, among the younger participants of sample 1, were part of the intervention group that studied using the progressive inquiry method and specific information and communication technology. seventy-seven students were part of the control group that received teacher-centered training and a learning method that was different from ict used in the intervention group. students were taught by six different teachers in nine groups at two schools with an average group size of 18 students. most (86.8%) of the middle school students in sample 1 were 14 years of age, and there were slightly more girls (55.3%) than boys (44.7%). the schools were homogeneous in terms of the students’ school performance, and there were no statistically meaningful differences between the self-reported evaluations (f1,150 = .926, p = .873). furthermore, no classes required special admission criteria. forty-six students were part of the intervention group that studied with the progressive inquiry method and specific ict in the upper secondary school in sample 2, and 55 students were part of the control group that received teacher-centered teaching and a learning method that did not use specific ict tools. they were taught by three different teachers at three different schools in five groups, with an average group size of 22 students. sixty-two percent of the students were 17 years old, 16% were 18 years old, 13% were 16 years old, and 9% were 19 years old. there were more male students (72.3%) than females (27.7%) among the upper secondary school participants. this shows that the voluntary course of geographical hazards attracted more boys than girls. there were no statistically significant differences between the students in the intervention and control groups in terms of academic success (f1,99 = .028, p = .866). furthermore, there were no statistically significant differences between the three different schools in terms of academic success (f2,98 =1.446, p = .240). the participants from the upper secondary school differed from the participants from the middle school because they chose to carry out their studies in this particular educational setting and were selected to attend the school based on their previous school performance. therefore, the content and design of the geography courses differed between the middle and upper secondary schools. however, the intervention groups’ progressive inquiry teaching and learning model was similar, and the way the intervention group used the ict had similar objectives, even though developmental levels of the participants differed. 2.2 research instruments the research instruments used in this study included self-report questionnaires and tests. the timeline of these instruments is displayed in appendix 1. a pre-test and post-test were designed according to bloom’s taxonomy (bloom, hastings, & madaus, 1981) to measure the cognitive learning outcome. the tests had similar structures for both age groups, including 20 multiple choice questions that were designed to measure the students’ recollection of factual knowledge (maximum score of 20 points), and five open-ended questions designed to measure the participants’ understanding and abilities to apply, analyze, synthesize, and assess knowledge (maximum score of 10 points). the full upper secondary school pre-test is presented in appendix 2. the multiple-choice questions are formulated according to the following example: “which of the following concepts refer to an astronomical body that falls onto the surface of earth? a. asteroid, b. comet, c. meteorite, d. shooting star.” one example of the open-ended questions is as follows: “compare the threats of volcanic eruptions and climate change. what differences are there in the nature of these threats?” a university teacher of geography didactics and two middle school and upper secondary school teachers assessed the tests. finally, a pilot study was conducted with 21 middle and 18 upper secondary school students. the means (and standard deviations) for the pre-test and post-test completed by the middle school participants were 22.4 (sd = 4.0) and 20.9 (sd = 5.3), respectively. the means (and standard deviations) for the pre-test and post-test completed by the upper secondary school participants were 19.0 (sd = 3.4) and 20.4 (sd = 3.4), respectively. the maximum score for the tests was 30 points, which assured that the items were challenging without being excessively difficult. for the middle school participants in sample 1 (n=152), the cronbach's coefficients were α = .76 for the pre-test and α = .84 for the post-test. this ensured that the tests were reliable. for the upper secondary school participants in sample 2 (n=101), the cronbach’s coefficients were α = .56 for the pre-test and α = .63 for the post-test. regulation of motivation was measured with a 31-item survey developed by wolters and his colleagues (wolters, pintrich, & karabenick, 2003; wolters & benzon, 2013). it extends previous instruments (e.g., mslq, pintrich, smith, garcia, & mckeachie, 1993) by focusing on (wolters, 2003, p. 190) “the activities through which individuals purposefully act to initiate, maintain, or supplement their willingness to start, to provide work toward, or to complete a particular activity or goal (i.e., their level of motivation).” the instrument investigates motivation regulation with six dimensions: 1) regulation of value; 2) regulation of performance goals; 3) self-consequating; 4) environmental structuring; 5) regulation of situational interest; 6) regulation of mastery goals. students responded to the survey with a five-point likert scale (1=totally disagree, …, 5=totally agree). a sample item from the self-consequating factor is as follows: “i make a deal with myself that if i get a certain amount of the work done i can do something fun afterwards.” according to wolters and benzon (2013), bivariate correlations among six motivation regulation strategies have shown to be positive and fairly strong (.15-.75). in this study, a summative scale was calculated for the 31 items and discretized into five classes (weak, sufficient, satisfactory, good, excellent, see table 1 for details) in order to group students based on their motivational regulation. the reliability of the 31 variables was high in both samples (sample 1, α = .95; sample 2, α = .94). this indicates that all the variables measure the same construct, strategies for the regulation of motivation. the learners self-assessed their levels of motivation on a five-point likert scale (1=totally disagree, …, 5=totally agree) at three time points: at the beginning (t0), in the middle (t1), and at the end (t2) of the geography course. items were derived from pedagogically meaningful learning questionnaire (pmlq, nokelainen, 2006) that was originally designed to measure experiences on the ten categories of learning connected to the usability of learning material: 1) learner control; 2) learner activity; 3) collaborative learning; 4) goal orientation; 5) applicability in other contexts; 6) added value of the learning method; 7) motivation; 8) valuation of previous knowledge; 9) flexibility in certain time points during the investigation period; 10) feedback. in this study, only category for the motivation was applied. it is based on mslq (pintrich, smith, garcia, & mckeachie, 1993) and measures dimensions of intrinsic goal orientation, extrinsic goal orientations and meaningfulness of studies with four items. these items were summarized into one summative motivation variable. one example of these items is as follows: “the topics to be learnt at school are interesting to me.” the first measurement (t0) took place in the first lesson of the geography course. the second measurement (t1) was done after the first digital game took place in the intervention groups; and the third measurement (t2) took place after the second digital game (sample 1, middle school) or after the interactive leaflet was finished (sample 2, upper secondary school). based on the time point, the cronbach’s alpha varied from .55 to .68 (sample 1, middle school) and .63 to .72 (sample 2, upper secondary school). 2.2 procedures the rationale for this quasi-experimental study was sent to the ethics committee of the tampere region for revision. the committee approved this investigation as planned. all of the students answered the pre-tests at the beginning of the course and completed the post-test at the end of the course. they also filled in a self-report survey about their strategies for the regulation of motivation during their first lesson. furthermore, a survey measuring their motivational level was filled out at the beginning, in the middle, and at the end of the course. 2.2.1 instruction for middle school participants in the intervention group the researcher and teachers jointly planned the events in the geography course, and the subject was the human and physical geography of europe. the events of the intervention group were designed to take place as follows: the teacher presented the outline of the course, including its main contents, objectives, and assessments. then, each student was asked to choose a european country for his or her project work, and the students with the same country formed a pair. next, each pair wrote down what they already knew about the country and why they chose it, and were asked to develop a study plan including research questions to be answered. the digital learning platform, moodle, was used to write the study plans, comment on them, ask questions, and disseminate the best information sources with peers. the project work proceeded progressively by searching for information and seeking answers to the questions in the study plan and inventing new questions. the project also involved drawing maps for certain geographical topics, such as topography and livelihoods, and writing down how the map related to other maps and phenomena. in addition to their progressive investigation, the adolescents were asked to design simple digital games for their peers on two different topics. their peers then played each game by solving the geographical dilemmas. the teacher used an interactive whiteboard during the geography course, and the students used it when playing the interactive game. the task of the first game was to design a pair of vortexes for two different types of weathering (e.g., physical and chemical weathering), and in the second game to design a vortex pair for two types of climate (e.g., mediterranean and subarctic climate). at the end of the course, a tourism fair was organized in the classroom. half of the student pairs played the role of experts advertising their country to visitors, and then they switched roles. the maps and diagrams were presented at the fair along with drawings, pictures, and souvenirs that the students chose to display. the students were then asked to compare their original study plans with the outcomes of their work in order to make the learning visible. 2.2.2 instruction for upper secondary participants in the intervention group the researcher and teachers jointly planned the events in the geography course. the subject of the course was geographical hazards in the field of human and physical geography. the events for the intervention group were designed to take place as follows: the teacher presented the outline of the course, including the main contents, objectives, and assessment. then, each student was asked to choose one geographical hazard for his or her project, and the students with the same hazards formed a pair. next, each pair wrote down what they already knew about the hazard and why they chosen it, and developed a study plan with research questions. the digital learning platform, moodle, was used to write the study plans, comment on them, ask questions, and disseminate the best information sources to peers. the project would proceed progressively, starting with searching for information, followed by seeking answers to the questions in the study plan, and finally inventing new questions. the project work involved drawing at least one picture of the hazard, and adding augmented reality in the form of a video clip or audio file to the drawing. each pair wrote a report about their hazard and attached a drawing that including augmented reality. finally, an interactive leaflet was created with the various hazards, which supported the students when studying for the exam. in addition to their progressive investigation, the adolescents designed a simple digital game for their peers about two climatic hazards (climate change and ozone depletion). their peers played each game by solving the geographical dilemmas. the teacher used an interactive whiteboard during the geography course, and the students used it when playing the interactive game. at the end of the course, each pair gave a presentation of their work and displayed their videos or other supplementary material. the students used tablet computers; therefore, they used the technology that was attached to their devices to present their work. finally, the students were asked to compare their original study plans with the outcomes of their projects in order to make the learning visible. 2.2.3 instruction for the control groups the teacher presented the outline of the course, including its main contents, objectives, and assessments. then, the teacher introduced each topic, followed by exercises that were executed individually or with a peer. the students did not complete a project on a specific topic throughout the entire course. moreover, they did not use a digital learning platform, such as moodle, for planning or disseminating their ideas or information sources. they did not design a digital game in pairs, and the upper secondary school participants did not create a leaflet. in other words, the students in the control groups were not guided to work collaboratively to the same extent as the students in intervention group. moreover, they did not have the same autonomy in their learning exercises, which meant they received more teacher-centered instruction without as many technological tools. the most important difference when compared to the intervention groups is that these learners did not follow the question-driven procedure of the progressive inquiry model and they did not use educational technology tools to implement the collaborative studying approach into practice. 2.4 data analyses all of the data analyses were executed with spss, and the numeric data was systematically observed for outlier detection. in middle school data set, both the pre-test and post-test scores were skewed toward the high end of scale, whereas in the upper secondary school data set, the distribution of the pre-test was normal and the post-test score showed high kurtosis and was slightly skewed toward high scores. nevertheless, the skewed values were below one; therefore, the distributions were considered close enough to normal for parametric tests to be applied. 3. results 3.1 rq1: effects of the teaching and learning model on cognitive outcomes the following phrasing was used when presenting the results: the intervention groups comprised students who studied with progressive inquiry and specific ict tools, as described in chapters 2.3.1 and 2.3.2, whereas students in the control groups received teacher-centered instruction, as described in chapter 2.3.3. the middle school students in the intervention and control groups scored similar results on the pre-test (t = -.647, df = 150, p = .519) and levene’s test for homogeneity of variances indicated equal variances (f = .152, p = .697). the middle school students in the intervention group scored higher points (m = 22.0, sd = 4.4) on the post-test compared to the control group (m = 19.4, sd = 6.1). an independent samples t-test was executed to investigate the statistical significance between these mean values. the difference was statistically significant (t = 3.0, df = 139, p = .003). according to levene’s test, the difference between the variances in the intervention and control groups was statistically significant (f = 10.0, p = .002). the deviation in the post-test score was higher in the control group (sd = 6.1) than in the intervention group (sd = 4.4). similarly, there were no statistically significant differences in the pre-test scores between the upper secondary students in the intervention and control groups (t = -1.413, df = 99, p = .161), and levene’s test indicated equal variances (f = 1.207, p = .255). the students in the intervention group scored statistically higher points (m = 21.4, sd = 2.9) on the post-test compared to students in the control group (m = 19.6, sd = 3.6, t = 2.761, df = 99, p = .007). the equal variances were assumed based on levene’s test (f = 1.016, p = .316). the gain value was calculated based on the original pre-test and post-test scores because it normalized the progress in the learning outcome results. for example, when a student who scored high points on the pre-test, scored high points once again on the post-test, they did not receive the highest points of improvement. the absolute gain value (g) was calculated as g = posttest% pre-test% / 100%-pre-test%. in other words, the gain value measured students’ improvement between the pre-test and post-test. in this sample, the variance in the gain values was tested by applying the mann-whitney u-test to the nonparametric independent samples because of the skewed values and high kurtosis in the distribution of the variables. mann-whitney u-test was used to examine the difference between the gain values of the middle school and the upper secondary school intervention and control groups. results showed a statistically significant difference (p < .0001) in absolute gain values between intervention group (m = -0.15, sd = .65) and control group (m = -0.59, sd = 1.15) among middle school students. similar result was also found from the upper secondary school sample (intervention group: m = 0.22, sd = .31; control group: m = -0.03, sd = .34). this indicates that in both samples the intervention group students’ cognitive geographical skills improved more than of those in the control group. 3.2 rq2: relationships between regulation of motivation and cognitive learning outcomes a summative variable was constructed from the score that was derived from the 31 variables measuring strategies for the regulation of motivation. as the observations of the original variable were normally distributed (n = 253, m = 127.9, sd = 27.5, mdn = 129, mo = 129), it was decided to reconstruct five groups of equal relative frequencies. the group in the middle included the mean, median, and mode values (table 1). the variable of different teaching and learning groups in different educational levels was recoded into a new dummy variable, where a value of 1 equaled the progressive inquiry and a value 0 equaled the teacher-centered approach. table 1 relative frequencies of motivation regulation skills in five categories (n = 253) we tested the null hypothesis that there were no differences in the variances between the different groups of teaching and learning methods, and the levels of self-regulated learning skills when investigated these groups’ relations to post-test results in both samples. in sample 1 (middle school), levene’s test showed that the error variance of the dependent variable (post-test score) was equal across the investigated groups (f9, 141 = 1.813, p = .071), and the results from the regression analysis suggested that in this model (colo = β0 + β1*tlm + β2*srl +β3*tlm*srl + ε; colo = cognitive learning outcome as post-test score, tlm = teaching and learning method, srl = self-regulated learning skills), the teaching and learning methods did not explain the post-test score (f1, 141 = 3.866, p = .121) in a statistically significant way. neither the self-regulated learning skills alone (f4, 141 = .329, p = .847) nor scrutinized together with the teaching and learning method (f4, 141 = 2.148, p = .078) had any statistically significant relationship with the post-test score. the anova showed that the teaching and learning models explained only a part of the success on the post-test (η2 = .056, p = .005), while self-regulated learning skills explained even less (η2 = .020, p = .589). in sample 2 (upper secondary school), levene’s test showed that the error variance of the dependent variable (post-test score) was equal across the investigated groups (f9, 89 = 1.714, p = .097). the results of the regression analysis suggested that in this model (colo = β0 + β1*tlm + β2*srl +β3*tlm*srl + ε) the teaching and learning methods explained the post-test score (f1, 89 = 13.367, p = .017, η2 = .74) in a statistically significant way. neither the self-regulated learning skills alone (f4, 89 = 1.088, p = .468, η2 = .52) nor scrutinized together with the teaching and learning methods (f4, 89 = .641, p = .634, η2 = .03) had any statistically significant explanatory effect on the post-test score. the anova suggested that this model failed to explain much of the variation in the post-test scores because the teaching and learning methods explained only some of the success on the post-test (η2 = .090, p = .004), and self-regulated learning skills explained even less (η2 = .030, p = .596). even though the differences between the variances were not statistically significant, figure 1 shows how the post-test scores were higher among the middle school students with weak, sufficient, satisfactory, or good self-regulation skills, when the progressive inquiry method was used. students with high self-regulation skills scored slightly higher points on the post-test when they received teacher-centered instruction. additionally, figure 1 shows that among upper secondary school students, the post-test scores were higher for every level of self-regulation skills when the progressive inquiry method was used. figure 1. self-reported self-regulation skills in relation to post-test scores in the intervention and control groups at both education levels. 3.3 rq3: relationship between the level of motivation and the teaching and learning methods in sample 1 (middle school), levene’s test showed that the error variance of the dependent variable (level of motivation) was equal across the middle school students in the investigated groups (t0: f1,150 = 1.056, p = .306; t1: f1,150= .435, p = .511; t2: f1,150= 1.699, p = .194) and in sample 2 (upper secondary school) (t0: f1,99 = .36, p = .549; t1: f1,99= .003, p = .958; t2: f1,99= 4.838, p = .030). in the upper secondary school sample, the last measuring point (t2) showed that equal variances could not be assumed between the intervention and control groups. this should be recognized when interpreting the results. when analyzed using wilks’ lambda test, the results showed that the motivation level did not differ in a statistically significant way between the different time points (sample 1, p > .05; sample 2, p > .05); however, it did differ in a significant way between the teaching and learning models used with the middle school (f1,150 = 7038.8, p < .0001) and upper secondary school (f1,99 = 5859.8, p < .0001) students. in both samples, the teaching and learning models explained most of the variation in the students’ motivation levels (η2 = .98). figure 2 visualizes the motivation levels at three different time points. figure 2. self-reported motivation levels in the intervention groups that used the progressive inquiry learning model compared to the control groups, which received teacher-centered instruction in middle school and upper secondary schools (measured at time points t0, t1, and t2). 4. discussion and conclusions according to this study, the progressive inquiry learning model fits well with geographical studies both at the middle and upper secondary school levels. when investigating rq1, the students at both educational levels showed higher post-test scores when taught with the progressive inquiry method, and their cognitive learning outcomes improved significantly more than the control group. as the skills to analyze data and make syntheses were also being measured in preand post-tests, the results suggest that this type of progressive inquiry teaching and learning model, which enhances the higher-order thinking skills, fits well both education levels. this result was in line with previous research on computer-supported collaborative learning (cscl), which suggests that cscl technologies trigger positive changes in group dynamics by enhancing learning interactions and enabling sharing and constructing knowledge among team members (ludvigsen, lund, rasmussen, & säljö, 2011). it was also congruent with the finding that learning complex and challenging science topics improves when students are repeatedly trained to use their self-regulatory skills in hypermedia environments (greene & azevedo, 2007). since collaborative learning has become more common (o’donnell & hmelo-silver, 2013), research on engagement in learning has broadened from the context of individual learning to collaborative groups (järvelä, järvenoja, malmberg, isohätälä, & sobocinski, 2016). we require more information on the cognitive and socio-emotional interactions between team members in order to support different kind of learners, and the ways that self-, co-, and socially shared regulation affect the collaborative learning process. furthermore, this study provided evidence that supported the following claim: since students in upper secondary school are older and have undergone a selection procedure to attend the school, it is understandable that the progressive inquiry method would suit them better than younger students because their self-regulated learning skills would be more developed. this conclusion is derived from upper secondary school students’ higher absolute gain values, as well as the higher probability that the gain values would differ between the group that studied using the progressive inquiry method and the group that received teacher-centered instruction. the results from rq2 suggested that the self-regulated learning skills that the student possessed before the course did not affect their cognitive learning outcome in a statistically significant way in either the middle or upper secondary schools. this meant that the students were able to attain the necessary self-regulation skills during the course, and students with poorer self-regulated learning skills had equal possibilities to succeed when either the progressive inquiry method or teacher-centered method was applied. there was an interesting tendency suggesting that middle school students with the strongest self-regulatory learning skills scored the highest points on the post-test when they received teacher-centered instruction. the result was not statistically significant, but it would be useful to investigate, whether the students with a high level of self-regulation skills do not benefit from the progressive inquiry method as much as students with weaker self-regulation skills. it would be useful to measure students’ self-regulated learning skills in terms of regulation of motivation at the end of the course as well, in order to get more insight about whether the progressive inquiry method enhances these skills. however, since the progressive inquiry method enhances student autonomy when making choices about the order of tasks, proceeding with tasks, or scheduling their studies (kuisma, 2018), good results in cognitive learning outcomes point to potential improvement in these learning outcomes as well. this finding is in line with wolters and benzon (2013, p. 218) as they state that “efforts to improve students’ self-regulated learning may benefit from including motivational strategies into the instructional plan.” in future studies, we intend to use six factor structure of motivation regulation instead of a summative score, as wolters and benzon (2013) concluded after their empirical study that correlations of motivation regulation factors were not substantial enough to indicate a single underlying construct. the results from rq3 clearly show the positive effect of the progressive inquiry method on learners’ motivation levels in middle school. wolters and pintrich (1998) found in their study that students’ motivation level varied across several disciplines (mathematics, social studies, and english), but not their regulatory strategy use. although current study had no variance in discipline and did not measure change in regulatory processes, motivation levels were higher throughout the geography course compared to the control group. the collaborative activities in the form of designing and playing the digital games, along with the progressing project work, seemed to increase the motivation level of the middle school students. however, in the upper secondary school, the motivation levels of the intervention group were lower than those of the group that received teacher-centered training. the motivation levels of the intervention group continued to rise from one time point to another. finally, the motivation levels at the end of the course rose to the same levels as the control group. altogether, the upper secondary school students seemed to be less motivated by the progressive inquiry method than the younger students. inquiry learning can be used in classroom practices with or without ict, but it has been shown that learners’ abilities for scientific research and collaboration skills are fortified when ict is included (banchi & bell, 2008; hakkarainen et al., 2004). it has been suggested that the gap between the every-day use of ict and the way it is utilized at school can affect students’ motivation in higher education (hietajärvi et al., 2015). although socio-digital participation has proven to be a tangled web of relationships, it is suggested that complex technology-mediated knowledge practices can be useful for learning in educational settings when striving for deep learning. it would be interesting to investigate this web of causal relationships among middle and secondary school students. keypoints the progressive inquiry method and learning model suits geography education because it improved students’ cognitive learning outcomes and motivation levels. cognitive learning outcomes were improved at both education levels, i.e., middle school and upper secondary school. upper secondary school students profited more from the progressive inquiry method in terms of cognitive outcomes compared to middle school students. previous self-regulated learning skills had no effect on cognitive outcomes; therefore, the necessary regulation skills could be adopted during the course. the positive effect on motivation levels was evident in the middle school context. references banchi, h., & bell, r. (2008). the many levels of inquiry. science and children,46(2), 26–29. bereiter, c. (2002). education and mind in the knowledge age. hillsdale, nj: erlbaum. bloom, b. s., hastings, j. t., & madaus, g. f. (eds.) (1981). evaluation to improve learning. new york, ny: mcgraw-hill inc. cerratto-pargman, t., järvelä, s. m., & milrad, m. (2012). designing nordic technology-enhanced learning . internet & higher education, 15(4), 227–230. doi: 10.1016/j.iheduc.2012.05.001 dimitrov, d. m., & rumrill, p. d. (2003). pre-test-posttest designs and measurement of change. assessment & rehabilitation, 20(2), 159–165. greene, j. a., & azevedo, r. (2007). adolescents’ use of self-regulatory processes and their relation to qualitative mental model shifts while using hypermedia. journal of educational computing research, 36(2), 125–148. retrieved from https://doi-org.helios.uta.fi/10.2190/g7m1-2734-3jrr-8033 engeström, y. (1999). innovative learning in work teams: analyzing cycles of knowledge creation in practice. in y. engeström, r. miettinen, & r.-l. punamäki (eds.), perspectives on activity theory (pp. 377–404). cambridge, uk: cambridge university press. finnish national board of education. (2015). curriculum reform 2016. retrieved from http://www.oph.fi/english/current_issues/101/0/what_is_going_on_in_finland_curriculum_reform_2016 hakkarainen, k. (2009). a knowledge-practice perspective on technology-mediated learning. computer-supported collaborative learning, 4, 213–231. doi: 10.1007/s11412-009-9064-x hakkarainen, k., palonen, t., paavola, s., & lehtinen, e. (2004). communities of networked expertise: professional and educational perspectives . amsterdam, netherlands: elsevier. hietajärvi, l., tuominen-soini, h., hakkarainen, k., salmela-aro, k., & lonka, k. (2015). is student motivation related to socio-digital participation? a person-oriented approach. procedia – social and behavioral sciences, 171, 1156–1167. doi:10.1016/j.sbspro.2015.01.226 hintikka, j. (1982). a dialogical model of teaching. synthese, 51 (1), 39–59. hintikka, j. (1985). true and false logic of scientific discovery. in j. hintikka & f. vandamme (eds.), logic of discovery and logic of discourse (pp. 3–14). new york, ny: plenum press. järvelä, s., & hadwin, a. f. (2013). new frontiers: regulating learning in cscl. educational psychologist, 48(1), 25–39. doi:10.1080/00461520.2012.748006 järvelä, s., järvenoja, h., malmberg, j., isohätälä, j., & sobocinski, m. (2016). how do types of interaction and phases of self-regulated learning set a stage for collaborative engagement? learning and instruction, 43, 39–51. doi:10.1016/j.learninstruc.2016.01.005 kuisma, m. (2018). narratives of inquiry learning in middle school geographic inquiry class. international research in geographical and environmental education ,27(1), 85–98. doi: 10.1080/10382046.2017.1285137 leat, d. (1997). cognitive acceleration in geographical education. in m. williams & d. tilbury (eds.), teaching and learning geography (pp. 143–153). london, uk: routledge. leat, d. (1998). thinking through geography. cambridge, uk: chris kington publishing. le deist, f. d., & winterton, j. (2005). what is competence? human resource development international, 8(1), 27–46. doi: 10.1080/1367886042000338227 loyens, s. m. m., magda, j., & rikers, r. m. j. (2008). self-directed learning in problem-based learning and its relationships with self-regulated learning. educational psychology review, 20(4), 411–427. doi: 10.1007/s10648-008-9082-7 ludvigsen, s., lund, a., rasmussen, i., & säljö, r. (eds.) (2011). learning across sites. new tools, infrastructures and practices. oxford, uk: routledge. muukkonen, h., hakkarainen, k., & lakkala, m. (1999, december 12–15). collaborative technology for facilitating progressive inquiry: future learning environment tools. in c. hoadley & j. roschelle (eds.), proceedings of the computer support for collaborative learning (cscl) 1999 conference . stanford university, palo alto, california. mahwah, nj: lawrence erlbaum associates. nagel, p. (2008). geography: the essential skill for the 21st century. social education, 72(7), 354–358. retrieved from https://www.socialstudies.org/publications/socialeducation/november-december2008/geography-the-essential-skill-for-the-21st-century nokelainen, p. (2006). an empirical assessment of pedagogical usability criteria for digital learning material with elementary school students. educational technology & society, 9(2), 178-197. retrieved from https://pdfs.semanticscholar.org/ea96/b628f440642d72026c14710a67ccd06f41f1.pdf nokelainen, p., kaisvuo, h., & pylväs , l. (2016). self-regulation and competence in work-based learning. in m. mulder (ed.), competence-based vocational and professional education. bridging the worlds of work and education: bridging the worlds of work and education . (pp. 775-793). dortrecht: springer. doi: 10.1007/978-3-319-41713-4_36 nonaka, i., & takeuchi, h. (1995). the knowledge-creating company: how japanese companies create the dynamics of innovation . new york, ny: oxford university press. o’donnell, a. m., & hmelo-silver, h. e. (2013). what is collaborative learning: an overview. in c. e. hmelo-silver, c. a. chinn, c. k. k. chan, & a. o’donnell (eds.), the international handbook of collaborative learning(pp. 1–15). new york: routledge. paavola, s., & hakkarainen, k. (2005). the knowledge creation metaphor: an emergent epistemological approach to learning. science & education, 14(6), 535–557. doi: 10.1007/s11191-004-5157-0 pauw, i. (2015). educating for the future: the position of school geography. international research in geographical and environmental education, 24 (4), 307–324. doi: 10.1080/10382046.2015.1086103 pintrich, p. r. (1999). the role of motivation in promoting and sustaining self-regulated learning. international journal of educational research, 31(6), 459-470. doi: 10.1016/s0883-0355(99)00015-4 pintrich, p. r. (2004). a conceptual framework for assessing motivation and self-regulated learning in college students. educational psychology review,16(4), 385–407. http://doi.org/10.1007/s10648-004-0006-x pintrich, p. r., smith, d., garcia, t., & mckeachie, w. j. (1993). reliability and predictive validity of the motivated strategies for learning questionnaire (mslq). educational & psychological measurement, 53(3), 801–813. doi: 10.1177/0013164493053003024 prince, m. j., & felder, r. m. (2006). inductive teaching and learning methods: definitions, comparisons, and research bases. journal of engineering education, 95(2), 123–138. scardamalia, m. (2002). collective cognitive responsibility for the advancement of knowledge. in b. smith (ed.), liberal education in a knowledge society(pp. 67–98). chicago, il: open court. scardamalia, m., & bereiter, c. (1994). computer support for knowledge-building communities. the journal of the learning sciences,3, 265–283. retrieved from http://www.jstor.org/stable/1466822 winne, p. h. (1995). self-regulation is ubiquitous but its forms vary with knowledge. educational psychologist,30, 223–228. wolters, c. a. (2003). regulation of motivation: evaluating an underemphasized aspect of self–regulated learning. educational psychologist, 38, 189–205. wolters, c. a., & benzon, m. b. (2013). assessing and predicting college students’ use of strategies for the self-regulation of motivation. the journal of experimental education, 81(2), 199-221. doi: 10.1080/00220973.2012.699901 wolters, c. a., & pintrich, p. r. (1998). contextual differences in student motivation and self-regulated learning in mathematics, english, and social studies classrooms.instructional science, 26, 27–47. wolters, c. a., & taylor d. j. (2012). a self-regulated learning perspective on student engagement. in s. l. christenson, a. l. reschly, & c. wylie (eds.), handbook of research on student engagement (pp. 635-651). boston, ma: springer. wolters, c. a., pintrich, p. r., & karabenick, s. a. (2003, april). assessing academic self-regulated learning. paper presented at the conference on indicators of positive development: definitions, measures, and prospective validity, childtrends, bethesda, md. zimmerman, b. j. (1998). academic studying and the development of personal skill: a self-regulatory perspective. educational psychologist, 33(2/3), 73–86. zimmerman, b. j. (2000). self-efficacy: an essential motive to learn. contemporary educational psychology, 25(1), 82–91. doi:10.1006/ceps.1999.1016 zimmerman, b. j., & schunk, d. h. (2008). motivation: an essential dimension of self-regulated learning. in d. h. schunk, & b. j. zimmerman (eds.), motivation and self-regulated learning: theory, research, and applications (pp. 1–30). mahwah, nj: erlbaum. appendix 1 research design group 1: sample 1, middle school students group 2: sample 2, upper secondary school students appendix 2 pre-test in upper secondary school circle the option that you think is correct. (only one of the options is correct!) 1. which of the following concepts refer to an astronomical body that falls onto the surface of earth? asteroid comet meteorite shooting star 2. during the sunspot cycle when the sun is the most active, earth's magnetic field changes, so that a magnetic storm may destroy transformers, electronic devices and satellites. the movement of tectonic plates strengthens and volcanic activity causes threats to human activities. the slowing down of sea currents may cause global warming. the ozone layer in the upper atmosphere thins down significantly. 3. how large a portion of all earthquakes on earth take place in the pacific ocean and its surrounding areas? 10 % 25 % 50 % 80% 4. loss of human lives caused by earthquakes can be decreased by building a large undivided space under an apartment building, e.g. a parking hall. by making large buildings as complex in shape as possible. by attaching the floors of apartment buildings (ceilings) properly to the building's load-bearing walls by using steel reinforcement structures. by decreasing the use of adjusting structures. 5. a tsunami usually begins from an earthquake which is caused by an asteroid crash. is caused by human activities, e.g. building of dams. takes place in the inner parts of the continental lithosphere plate. raises the water mass above the whole earthquake area, causing a giant wave advancing in all directions. 6. the magnitude of an explosive volcanic eruption depends on the composition of magma. the steepness of the volcano's slope. population density on the slopes of the volcano. the location of the area with respect to the equator. 7. volcanic activity is useful to local people because geysers and craters attract tourists. a volcano that has erupted once will not erupt a second time, and it is possible to build a tourist resort in the crater. the volcanic mud slides from volcanos would make a luxurious spa treatment for tourists. it is easy to predict volcanic eruptions, and volcano eruption watching attracts tourists. 8. how does a cyclone form? it is formed when strong winds combine together. it is formed from warm seawater from a strong low pressure area. it is formed when a rain front approaching from the sea meets a continent. it is formed when a tsunami meets a continent. 9. which of the following countries have suffered the most from tropical cyclones? japan bangladesh united states australia 10. what are twisters? strong hurricanes which kill hundreds of people in the united states every year. tornado-like whirlwinds which occur in finland, for example. splinters of stones thrown into air during a volcanic eruption. earthquake waves. 11. floods cause less damage in ostrobothnia than in finnish lakeland. half of all deaths in the world caused by natural catastrophes. depletion of soil nutrients in the floodplains of rivers. more damage in natural rivers than in rivers that have been turned into straight canals. 12. hot weather can be a threat to human lives because air-conditioning devices are ineffective. hot weather is difficult to predict and there are no advance warnings. hot weather disrupts railway schedules. the temperature regulation system in a human body is working at its limits when the surrounding temperature rises to 40°c. 13. erosion, the wearing away of the soil, is accelerated if terrace plantations are built on steep slopes instead of fields that naturally follow to the shape of the slopes. if treeless areas are forested. due to logging of forests, agriculture and animal husbandry. due to thickening of the plant cover. 14. the most significant harmful effects from an unbalanced diet are more common in the population of working age than in children. are less significant than the harmful effects of excessive eating. are caused by lack of proteins, minerals, vitamins and fibre. are caused by insufficient caloric intake. 15. which disease spreads through dirty drinking water? typhoid fever malaria aids ebola 16. which of the following statements related to natural biodiversity is correct? biological diversity has increased over the last centuries. climate change decreases the natural biodiversity in arctic areas. genetic diversity is not part of biological diversity. out of all ecosystems, wetlands have suffered the least from decrease of biodiversity. 17. which one of the following statements related to climate change is incorrect? the average temperature on earth is estimated to rise 1–5 degrees celsius in a hundred years. the strengthening of the greenhouse effect is mainly due to carbon dioxide and other greenhouse gas emissions from usage of fossil fuels. climate change makes it easier to cultivate grain in regions around the equator. the surface level of oceans will rise due to climate change. 18. climate change can be slowed down by planning the infrastructure of an area so that motor vehicles are used as little as possible. by increasing one's own carbon footprint. by increasing methane emissions. by increasing consumption of fossil fuels in developing countries. 19. ozone depletion is less often noticeable above the antarctica than in the north pole area. a significant environmental problem in the lower levels of the atmosphere. a significant environmental problem in the upper atmosphere, the stratosphere. decreased due to freon and halon emissions. 20. the problem with urbanization is formation of smoke fog, or smog. decrease of floods as rain water is directed to storm water drains. merging of the residential areas of the well off and the disadvantaged people. increasing biodiversity. answer the following questions in writing: 21. compare the threats of volcanic eruptions and climate change. what differences are there in the nature of these threats? 22. how can slowing down of climate change be useful in sub-saharan africa? 23. there are regular floods in a certain area. what could cause the floods? how could the inhabitants decrease flood damages? which kind of damage caused by floods is the most harmful, in your opinion? explain briefly. filius et al publication2 frontline learning research vol.6 no. 2 (2018) 92 113 issn 2295-3159 promoting deep learning through online feedback in spocs renée m. filiusa, renske a.m. de kleijnc, sabine g. uijld, frans j. prinsc, harold v.m. van rijenb, diederick e. grobbeea auniversity medical center utrecht, julius center for health sciences and primary care, the netherlands buniversity medical center utrecht, biomedical sciences department, the netherlands cutrecht university, department of education, the netherlands dutrecht university, university college, the netherlands article received 8 march 2018/ revised 18 april/ accepted 6 september/ available online 12 november abstract higher education aims for deep learning and increasingly uses a specific form of online education: small private online courses (spocs). to overcome challenges that instructors face in order to promote deep learning through that format, the use of feedback may have significant potential. we interviewed eleven instructors and four students and organized a focus group to formulate scalable design propositions for instructors in spocs to promote deep learning. propositions have been formulated according to the cimo-logic. this study resulted in identification of four mechanisms by which the desired outcome (deep learning) can be achieved, which we describe here along with proposed interventions. results show that the “online learning interaction model” can be deepened with these mechanisms: 1) feeling personally committed, 2) asking and providing relevant feedback, 3) probing back and forth, and 4) understanding one’s own learning process. to activate these mechanisms, scalable feedback interventions are described in three categories. results at this relatively young field of spocs also show that feedback as a dialogical process may contribute to solving the current challenges of instructors in spocs to achieve deep learning with their students. keywords: online learning; deep learning; peer feedback; spocs; teaching/ learning strategies. info corresponding author mail: r.m.filius@uu.nl doi: 10.14786/flr.v6i2.350 1. introduction deep learning involves critical thinking, integrating what the student is learning with what he or she already knows and creating new connections (biggs, kember, & leung, 2001; entwistle, 1991; marton & saljo, 1997; hounsell, 1997) and is related to higher-order thinking skills. promoting deep learning is an important task for higher education (biggs & tang, 2011, nicolls, 2002, lynch, mc namara & seery, 2012, ramsden, 1992), which is increasingly conducted online (geitz, brinke, & kirschner, 2015). small private online courses (spocs) are a distinctive form of online education that is used in higher education, especially since the last decade (uijl, filius, & ten cate, 2017). recently, filius, de kleijn, uijl, prins, van rijen and grobbee (2018) found that instructors experience specific challenges when trying to promote deep learning in spocs. their study resulted in a description of five main challenges for instructors: alignment of learning activities, insights into student needs, adaptivity of teaching strategy, social cohesion, and creation of dialogue. these challenges are due to a lack of facial contact and visual cues, as online learning tends to involve mostly asynchronous written interaction, and to the fact that the course material is usually developed and set before the start of the course. to overcome such challenges to promoting deep learning in online education, the incorporation of feedback as a pedagogical strategy may have significant potential, which is currently not optimally exploited (lynch, mcnamara, & seery, 2012; rushton, 2005). following carless (2011), we take a broad definition of feedback as “all dialogue to support learning in both formal and informal situations” (askew & lodge, 2000, p. 1). this illustrates how we view feedback as a two-way form of interaction and not as a one-way comment from one to the other. it is generally agreed that feedback plays an important role in higher education (nicol & macfarlane-dick, 2006). feedback leads to the development of higher-order skills (davies & berrow, 1998), connecting new knowledge to what students already know and to knowledge construction (nicol, 2009). engaging students in peer feedback helps develop skills for reflection, self-regulation, and critical thinking (boud, 2001; dochy, segers, & sluijsmans, 1999; lin, liu, & yuan, 2001; p. m. sadler & good, 2006). feedback in spocs may be even more important than in face-to-face classes, because it increases the student-instructor interaction and student-student interaction, and thus compensates for the potential geographical disconnect in online courses that may affect students’ retention (dennen, aubteen darabi, & smith, 2007; richardson, koehler, besser, caskarlu, lim, and mueller, 2015). however, nowadays instructors are under pressure to provide high-quality feedback to students in a prompt manner, often to large and diverse cohorts (allan & bentley, 2012; nicol, 2009; planar & moya, 2016). and even though spocs involve small groups, the number of parallel courses that run at the same time and the diversity of students may be high, which makes the provision of feedback very time-consuming (crook, mauchline, maw, lawson, drinkwater, lundqvist, … and park, 2012). therefore, this study aims at exploring how challenges of instructors can be overcome when providing feedback by developing design propositions for instructors in spocs to promote deep learning. 2. design propositions to promote deep learning in spocs 2.1 deep learning the distinction between deep learning and surface learning as students’ approaches to studying has been supported by the results of previous research (biggs, kember, & leung, 2001; entwistle, 1991; marton & saljo, 1997). deep and surface learning are considered to be two extremes of a continuum. surface learning indicates that the learner simply memorizes new ideas. deep learning is defined as the process of actively integrating new ideas into the existing cognitive structure through critical thinking, integrating what is learned with what was already known, and creating new connections between concepts (aharony, 2006; biggs, 1999; hall, ramsay, & raven, 2004). according to garrison, anderson, and archer (2001), in order to promote deep learning, the whole person should be engaged–cognitively, socially and affectively–in the learning process. a deep learning approach is more likely to result in better retention and transfer of knowledge (ramsden & moses, 1992) and to lead to high-quality learning outcomes such as a good understanding of the discipline and critical thinking skills (athanassiou, mcnett, & harvey, 2003; athanassiou, 2003; biggs, 1999; booth, luckett, & mladenovic, 1999; lindblom-ylänne, 1999; ramsden & entwistle, 1983; trigwell, prosser, & waterhouse, 1999). students are unlikely to experience high-quality learning outcomes, or develop appropriate skills and competences through a surface approach to learning (hall et al., 2004). 2.2 using feedback to promote deep learning there are several instruments that instructors can use to promote deep learning, such as using concepts maps (hay, 2015), cross-cultural chat (osman & herring, 2007), podcasting (pegrum, bartle, longnecker, 2014) and online asynchronous discussions (du, harvard & li, 2005). one of the most powerful instruments for instructors to influence learning is feedback (hattie & timperley, 2007; kluger & denisi, 1996). we argue that feedback through dialogue between instructors and students, among peers, or perhaps even between student and computer may promote deep learning. the purpose of feedback is to reduce the discrepancies between the students’ current understanding or performance and the understanding or performance that is being aimed for (hattie & timperley, 2007). according to hattie and timperley, feedback is information provided by a source (e.g., teacher, peer, book, parent, self, experience) regarding aspects of one’s performance or understanding. however, once the feedback has been provided, the receiver needs to process and respond to the feedback. the way the student receives the feedback is just as important as how the provider intended the feedback (ilgen, fisher, & taylor, 1979). ilgen and colleagues composed a model in which they showed the student’s processing of feedback into different stages. emphasis was put on those aspects of feedback that influence: a) the way feedback is perceived, b) its acceptance by the recipient, and c) the willingness of the recipient to respond to the feedback (ilgen et al., 1979). in line with this model, according to nicol (2010), carless, salter, yang, and lam (2011), boud and molloy (2013), and planar and moya (2016), feedback can be viewed as two-directional and needs to constitute a dialogue between the person who facilitates it and the one who receives it. it must explicitly promote self-regulation and a proactive attitude on the part of the student towards it; at the same time, it needs to focus on the learning process and involve peers. according to geitz et al. (2015), feedback should be supported by dialogue and by activities that not only inform students about their current performance, but also teach them to seek and ask for feedback on future performances. this will put students more in control. it will also enable them to add meaning to the feedback and to discuss the feedback as equals with their peers. 2.3 the role of the instructor and the student this student-centered approach assumes that no longer the instructor, but the student, has become the center of the learning process. the instructor has become a facilitator who guides the learning process. garrison, anderson, and archer (2000) developed the community of inquiry framework that sheds more light on the role of the instructor to influence students’ deep learning approaches. in order to promote deep learning, the instructor should aim at three interdependent structural elements of the framework—social, cognitive, and teaching presence. social presence reflects the development of climate and interpersonal relationships in the community. cognitive presence provides a description of the progressive phases of practical inquiry leading to resolution of a problem or dilemma. teaching presence provides leadership throughout the course or study. these three elements that the instructor should focus on in online education show similarities with the “online learning interaction model” from ke and xie (2009). as garrison and colleagues (2000) focus on the teaching activities of the instructor, ke and xie focus on the learning activities of the students. both view interaction as a core indicator for deep learning. ke and xie (2009) distinguish three different types of interaction of students in an online course: 1) social interaction, 2) knowledge construction, and 3) regulation of learning. their model is based on concepts for deep learning in adult education and helps to examine the quality of online education. even though instructors may view interaction as essential to deep learning, given the high student-staff ratio it can be difficult for the instructor to engage in dialogue with students. thus, instructors look for alternative feedback strategies that are efficient and effective and less time-consuming (allan & bentley, 2012) and that can be implemented in spocs. for example, peer feedback strategies have shown to be beneficial to deep learning (anderson & rourke, 2002; boud, cohen, & sampson, 1999; moon, 2013). the combination of feedback strategies and the specific context of spocs will lead to a set of design propositions specifically useful for promoting deep learning in spocs. 2.4 design propositions and cimo-logic design propositions are heuristic statements about how and why a pedagogical intervention works in a certain context (plomp & nieveen, 2009). a design proposition is intended to be transparent, comprehensive, and described in such a way as to make clear under which conditions it lends itself to generalization for other contexts. in this study, the design propositions will be formulated according to the cimo-logic (van aken, 2007; van den akker, 1999) used in design literature (e.g. denyer, tranfield, & van aken, 2008) and several recent studies (bronkhorst, meijer, koster, & vermunt, 2011; brouwer, brekelmans, nieuwenhuis, & simons, 2012; dobber, akkerman, verloop, admiraal, & vermunt, 2012). a design proposition describes the specific context to which it applies, the intervention proposed, and the mechanism by which the desired outcome is achieved: cimo-logic. the causal relation between the intervention and outcome in the context is (potentially) more plausible when all cimo components are described (brouwer et al., 2012). this inclusion of context dependency and mechanisms triggered is why the cimo-logic is preferred over other ways of specifying design propositions that exist in the literature, which are often limited to specification of intervention and outcome. cimo-logic determines that a design principle has the following structure: “in this class of problematic contexts, use this intervention type to invoke these generative mechanism(s), to deliver these outcome(s)” (denyer et al., 2008, p. 395). for example, “if you have a spoc in which you want students to respond to each others’ contributions and try to look for common understanding (context), support group work (intervention type) to promote deep learning (intended outcome) through probing back and forth (mechanism).” figure 1 shows how the cimo-logic has been applied in this study. the context is defined by the specific challenges that instructors in spocs experience when aiming to promote deep learning. the contexts elucidate the context dependency of the intervention. interventions are purposeful measures (products, processes, or activities) that are formulated by the designer (or instructor) in order to solve a design problem or need (denyer et al., 2008; midgley, 2000), for example the need for deep learning. van aken (2004) indicates that the key question is not so much whether the intervention works, but what it is about the intervention that makes it work. why does an intervention lead to a certain outcome in a specific context? this has been described in the mechanisms. outcomes are the result of the interventions. figure 1. cimo logic (based on van den akker, 1999). 2.5 research question we believe taking a design proposition perspective in which interventions, outcomes, ánd mechanisms are investigated in relation to each other is rather unique, and will provide contextualized conclusions that have both practical and conceptual value. therefore, this paper aims to address the question: “how and why can deep learning in higher education spocs be promoted using scalable feedback interventions?” feedback interventions consist of information that is externally generated and includes tips for improvement (kluger & denisi, 1996). in this study, only feedback interventions that are scalable have been included, which refers to the requirement that it must be possible to increase the number of students involved without increasing the workload of the instructors. by investigating the mechanisms, we aim to answer why the intervention will (not) promote deep learning. 3. methods 3.1 design the study design was qualitative and exploratory and used individual interviews with instructors in spocs representing participants from different fields of study as well as students. since this study focuses on the design propositions for instructors, the interviews with the students have solely been used to substantiate the interviews with the instructors. this triangulation of the findings supported multiple perspectives rather than only the instructors’ perspective. moreover, a focus group representing experts from different disciplines was added. according to powell and single (1996), in cases where the existing knowledge of a subject is inadequate, as is the case here, the use of a focus group is especially useful and can be employed to gather diverse ideas about possible feedback interventions. the supportive, congenial, non-judgmental setting offered by the focus group enhanced the likelihood of collecting the diverse and spontaneous opinions that eluded the in-depth interviews (powell & single, 1996). the study was approved by the dutch ethical board for research in education (nvmo, the netherlands association for medical education, approval no. 210). the nvmo is an independent association that carries out activities for anyone involved in medical and health care education in the netherlands and flanders (belgium). 3.2 participants 3.2.1 individual interviews the data used in this study were taken from the same dataset as a previous study (filius et al., 2018) for the individual questions with the instructors. each study used different parts of this dataset. concerning the selection of participants, we aimed for maximal variation and theoretical sampling (guba, 1981). therefore, the first author asked the heads of the education and it departments at 4 institutions to recommend instructors and students from their institutions with experience in teaching or participating in spocs. from these recommendations we selected instructors and students in spocs with varying levels of experience in years of teaching in or following spocs. we expected age and experience to be relatively large influencers, more than for example backgrounds. in addition, we included instructors that we considered to be experts and who are known for being keynote speakers at relevant international conferences about online education. we expected them to have a broad view of developments among instructors and to increase the chance that we included as many experiences as possible. both the participating instructors and students represent different universities and virtual learning environments. a maximum of 2 of the same universities and a maximum of 2 of the same virtual learning environments were represented, which resulted in 10 different universities and 8 different virtual learning environments. the number of the purposive sample sizes of instructors has been determined by data saturation as the collection of more data appeared to have no additional interpretive worth (guest, bunce, & johnson, 2006). in the case of the students, we were looking for counter evidence for the findings of the interviews with the instructors. after four interviews we had not found any counter evidence and then we decided not conduct any additional interviews. all of the 11 invited instructors and 4 invited students agreed to be interviewed. all instructors (4 female and 7 male) were involved in teaching online courses in higher education. the average age of instructors was 51.8 years (sd=20.0), average teaching experience was 15 years (sd=19.6) compared with their average experience with spocs of 10.4 years (sd=6.8). six instructors had 2 years or less experience with spocs, the other 5 instructors had 10 years or more experience with spocs and online distance education. two of the instructors are also researchers in the field of online education. additionally, 4 students (3 female and 1 male) were involved, ranging in age from 28 to 52, with an average age of 39 years (sd=9). two of them had participated in just one spoc; the others had participated in more spocs, varying in duration and study load. 3.2.2 focus group session in total 10 professionals, other than the interviewed instructors, engaged in the focus group session. they were selected using specific-criterion sampling, which is a type of purposive sampling and selection method in which one concentrates on people with specific characteristics (palys, 2008). we interviewed 10 professionals from multiple disciplines and areas of expertise who are known for their open-mindedness, to fill in the gaps with more unconventional interventions. their ages ranged from 23 to 52 years of age. all of them work in art, technology, and/or education, a number of them being at the intersection of several disciplines. the three disciplines were evenly represented. their job positions are: journalist, artist, product manager of moocs, researcher, educational platform manager, educational technologist, and game designer. some of them were also students or instructors. all participants took part on a voluntary basis. 3.3 procedure 3.3.1 individual interviews participants were informed of the study’s purpose and approach both in the invitation e-mail and at the start of the interview. this included an explanation of the outcome ‘deep learning’ and ‘scalable interventions’. during the interviews, the interviewer asked each participant to name several examples of deep learning, to compare these with findings in literature to determine whether their understanding corresponded to our previous study. hardly any differences emerged in this respect. each participant signed a consent form. interviews were based on an open interview scheme following a qualitative approach (cohen, manion, & morrison, 2013; cresswell, 2007). this was done so as to do justice to the complexity of the topic as well as to the nature of encapsulated expert knowledge (berliner, 2001), since the in-depth nature of open interviewing allows the informants to answer from their own frame of reference (bogdan & biklen, 2003; cohen et al., 2013). the interviews lasted an average of one hour. the interview questions for instructors are shown in table 1. the same questions were asked of students, but from their perspective. questions were related to the cimo method by asking for the context, intervention used, mechanism activated, and outcome achieved. the deep learning process has been operationalised as the initiation of critical thinking, integrating what the student is learning with what he or she already knows and creating new connections. these three deep learning activities are mental processes which, when initiated, are considered as 'deep learning outcome'. specific attention was paid to what interventions have been used and which mechanisms triggered the deep learning activities. by subsequently asking for three statements or golden rules about providing feedback to promote deep learning, participants were encouraged to speak freely about their ideas on what might help to promote deep learning in spocs and why this might help. other questions asked to all participants were to prompt and/or probe for additional information. table 1 interview questions supplemented with probing questions 3.3.2 focus group session during the focus group session, a short introduction was provided to present the results of the interviews and to explain the definitions of spocs, feedback, and deep learning. participants were informed about the summarized results of the interviews in terms of contexts, mechanisms, and desirables outcomes. then they were asked to brainstorm in three rounds about the results of the interviews. every round involved different group compositions. following their suggestions in the small groups, they collaboratively discussed the interventions and mechanisms in more depth in order to conclude how feedback may promote deep learning in spocs. 3.4 analysis the analysis of the data proceeded in stages, using nvivo to code and retrieve the data. first, the interviews and focus group session were audio recorded and transcribed. to avoid misrepresentation and misinterpretation of interviewees’ statements, the transcript and a summary of the transcription were sent to each participant for member checking (poortman & schildkamp, 2012). the focus group participants received a report for verification, which was created based on the transcript and notes. all participants agreed with the transcribed content. second, the transcripts of the interviews and focus group session were inductively coded into meaningful categories by the first author, using open coding (cresswell, 2007). fragments of all texts of which the corresponding code was debatable according to the first researcher, which came to less than 2.5% of all texts, were discussed by the full research team. next, each meaningful category has been classified using the theme intervention and the theme mechanism, according to the cimo-logic. then the first author moved to more selective coding stages according to an iterative process. based on the previous round of analysis, the codes have been revised. on the basis of the data, some codes have been merged, deleted or reformulated. subsequently, all data were analyzed again, now with the new codes. considering the open and grounded nature of this analysis (bogdan & biklen, 2003) at every coding stage, all different categories were discussed by the research team until agreement on the categories’ content, as well as the codes, was reached. to enhance reliability in coding, an independent researcher also analyzed a random sample of approximately 10 percent of the data for calculating the inter-rater reliability. the percentage of agreement was 93%. internal validity was further enhanced due to the description of the results, which were context-rich, meaningful, and thick. external validity was promoted by including respondents’ quotes and by describing the coherence with the theoretical framework. reasoning from the perspective of the cimo-logic, the interventions that were derived from the data into meaningful categories are suggestions from respondents on how online feedback could overcome the problems mentioned in the specific context of a spoc. only interventions that are scalable, that is, without being very time-consuming, were selected as meaningful categories, in light of the constraints of shrinking staff budgets and expanding student numbers. for example, feedback interventions such as direct conversations with videoconferencing tools between instructor and student have been excluded for this reason, despite their potential in achieving deep learning. the mechanisms that derived from the data shed light on why interventions lead to the desired outcome, which is deep learning. each of the mechanisms was classified into one of the categories of ke and xie’s online learning interaction model (2009): 1) social interaction, 2) knowledge construction, and 3) regulation of learning. to ensure quality in all of the steps described, an audit was conducted by an independent researcher concerning all steps of data gathering and analysis (akkerman, admiraal, brekelmans, & oost, 2008). the audit had both a formative and a summative function. as a consequence, the auditor assessed the steps taken several times during the study and at the end of the study. this resulted in an audit report with questions and answers, mostly about the analysis of the data. for that reason, there have been some adjustments in the description of the analysis in this article. thereafter, the auditor reviewed the study again and affirmed it as being visible, comprehensible and transparent. according to the auditor, decisions are explicated and communicated, decisions have been substantiated and decisions are acceptable according to standard, values and norms. 4. results in the results of this study, we describe design propositions to overcome challenges in promoting deep learning in spocs according to the cimo-logic. the design propositions consist of the context (specific challenges in spocs) in which feedback interventions will trigger student mechanisms that will lead to the desired outcome (deep learning). we start with describing four main student mechanisms by which deep learning (the desired outcome) can be achieved in spocs (context). after that we address how these mechanisms can be triggered by feedback interventions. the letter after each quote refers to either an instructor (i) or student (s). where the suggestions of students were additional to those of instructors, it was explicitly mentioned that this originated from students. 4.1 student mechanisms the mechanisms in this study are the processes that are internal to the student. they describe how students engage in learning activities, which largely determines the quality of the learning outcomes they attain (vermunt & verloop, 1999). knowledge concerning the mechanisms sheds light on why interventions lead to the desired outcome, which is deep learning. mechanisms are: 1) feeling personally committed, 2) asking and providing relevant feedback, 3) probing back and forth, and 4) understanding one’s own learning process. each of the mechanisms has been categorized according to the online learning interaction model (ke & xie, 2009) as a) social, b) knowledge construction, or c) regulation. 4.1.1 mechanism 1: feeling personally committed (category: social) if students are personally addressed, they feel personally committed and accept the feedback more easily. according to the instructors, possibilities to do so in online education have not been optimally utilized yet. one of the instructors said: “one of the benefits of online learning, i think, is the transparency. because students write assignments, receive and give feedback, it is easy to get the picture: he is there, they are there, and those guys over there still don’t get it” (i8). another instructor explained, “here’s what i find is the benefit: in a classroom situation you rarely have the opportunity to ask, to focus on what every single student thinks or what every student is thinking about that question. in a classroom you only have a limited amount of time and you may have three, four, five students answer that question, but you don’t know what every student is thinking. in an online course you have the opportunity to get that student to respond to that–every single student to respond to that question and you have the opportunity to provide one-on-one feedback and ask those questions. in a face-to-face classroom i would never know those students who weren’t thinking… you only see the stars, basically” (i7). students will prefer to choose a deep learning approach once they feel personally committed, which can be achieved through tailored feedback: “the individualization, the differentiation that you can give to students in an online environment is so much greater than you can do in a face-to-face classroom.” (i4). 4.1.2 mechanism 2: asking and providing relevant feedback (category: knowledge construction) to learn how to focus on a deep learning approach, students indicate that it helps them to learn how to ask for feedback, but also how to provide peer feedback that promotes deep learning. instructors confirm this. one of them adds: “i think it is very instructional for students to provide feedback, for themselves. that they learn how to grade such a piece of work, what criteria are being used. and they will have to keep doing so, later in their life, when they are working at the university or elsewhere” (i6). students said that they haven’t been taught how to provide meaningful feedback and that it is hard to learn it oneself. instruction will thus be useful. compared to face-to-face education, students tended to ask for feedback more frequently, just because it is easier since there seems to be an opportunity 24 hours per day. both instructors and students also tend to provide feedback faster in online education because the virtual learning environment enables them to be very quick. both instructors and students think that this fast way of asking for and providing feedback may promote more of a surface approach to learning. and because the number of feedback requests is so high, it is difficult for students all to get involved in a dialogue with the instructor. an instructor explains how he deals with the large number: “we selected the most important issues and also some examples, and that was what we discussed” (i5). thus, according to the members of the focus group, it may help students to learn how to prioritize feedback requests. 4.1.3 mechanism 3: probing back and forth (category: knowledge construction) in order for deep learning to occur, students and instructors experienced a need for a back-and-forth probing to take place. by presenting ideas and getting feedback on these ideas by ping-ponging back and forth with peers and/or instructor, students thought deeply and got the opportunity to combine what they already knew with new knowledge. it required an environment in which students felt safe and interacted comfortably with each other and with the instructor. according to a student: “you need to build a relationship with each other in order to be motivated and to be able to accept the feedback, so someone must be open to receiving feedback” (s1). another instructor explained: “the feedback that works best is the feedback in which you can keep asking questions after each response from the student. as a dialogue. because that really forces the student to think deeply” (i1). back-and-forth probing can be either synchronous or asynchronous, but most respondents preferred it as synchronous: “it becomes snappier, it is easier to ask questions right away, to help the student to take the necessary steps and to think deeper” (i8). 4.1.4 mechanism 4: understanding one’s own learning process (category: learning regulation) both instructors and students expressed the view that deep learning can be promoted by letting students apply their knowledge, for example, in a scenario or case study. students will have to try to apply new information in other contexts, which enables them to create new knowledge and to make connections with prior knowledge and new concepts. they will have to go through various steps and receive feedback on each step. by doing so, they engage themselves in meaningful ways that enable them to reflect deeply on the learning activity and the feedback they have received. creating the right feedback for each step to be taken requires forward thinking. one of the instructors explains: “i found that very little deep learning occurs online anyway, unless there is some type of a scenario, or they have to apply it in a case study. in other words, it’s forward thinking. i would call it that the deep learning occurs when you have opportunities for forward thinking, forward looking. ‘what would you do if…? what would happen if…? what’s the projection if this?’ and it’s a little bit of what-if/then kind of thinking, that i think precedes all of the other learning. and without that, i don’t think that it really progresses further” (i4). this mechanism helps students to be prepared for opportunities to develop the capacity to regulate their own learning as they progress through higher education. 4.2 triggering mechanisms through feedback interventions the mechanisms described above can be triggered by several feedback interventions, which are described in three following categories: 1) feedback management 2) peer feedback, and 3) automatic feedback. the mechanisms and interventions are summarized in figure 2. figure 2. interventions and mechanisms according to the instructors and students. 4.2.1 feedback management interventions the feedback management interventions describe how to manage online the monitoring and provision of feedback to and among students in such a way that deep learning is promoted. for each intervention, the dominant mechanisms that were identified in this study are indicated in italics. intervention a: collect student information in advance in order to make students feel personally committed and to estimate what feedback is needed, it helps to collect student data before the start of the course. student data involved learning characteristics, such as the education level, results on a pre-test, information on expectations, personal learning objectives, and motivation. collecting this data benefited the feedback provided, because it enabled instructors to adjust their feedback to the needs of the students and thus make it more specific. the relatively convenient availability of student data in spocs compared to face-to-face education may compensate in part for the lack of facial contact. for instructors in spocs, knowing more about their students helped them to formulate the feedback to make it more tailored to the student’s needs. specific suggestions of how to implement this intervention mentioned in interviews and/or focus group have been added in table 2. table 2 specific suggestions of intervention a < intervention b: monitor progress using a dashboard instructors monitored students’ progress during their education using a dashboard. the dashboard provided the instructor with an analysis of student data such as their contributions to assignments and discussion forums, questions, completion rates, and grades. it enabled instructors to intervene during the course and provide specific personalized formative feedback, for example when students skipped certain necessary steps or when they tended to think in a wrong direction. according to instructors, receiving personalized feedback helps students to feel more personally committed and may help them to understand their own learning progress better–especially when the dashboard is visible to the students themselves, as members of the focus group suggested. specific suggestions of how to implement this intervention mentioned in interviews and/or focus group have been added in table 3. table 3 specific suggestions of intervention b intervention c: bring requests back to the essentials participants in the focus group suggested reducing the number of feedback requests and letting students prioritize the issues they want to receive feedback on. instructors suggested teaching students how to ask for the right feedback and guiding them during this learning process by reflecting on the type of feedback questions they ask. instructors in spocs expect this to be useful in aligning the learning activities with the learning goals and the assessment goals so that they all promote deep learning. moreover, it will help students to ask for (more) relevant feedback. specific suggestions of how to implement this intervention mentioned in interviews and/or focus group have been added in table 4. table 4 specific suggestions of intervention c intervention d: discuss and rate the quality of the feedback students suggest that teaching them how to provide relevant feedback may promote deep learning. to do so, participants in the focus group suggested letting students discuss and rate the quality of the feedback they provide and receive. by discussing and rewarding the quality of the feedback provided, students aim to learn how to focus on deep learning and how to increase the quality of their feedback. this might also help to give students recognition for the effort they make to provide good feedback. according to the interviewed respondents, feedback to promote deep learning should include many questions to elicit deep learning. a discussion could start with an instruction on, for example, what questions are helpful to promote deep learning, such as what-if/then questions; for example, “what would you do if…?” “what would happen if…?” “what’s the projection if this...?” specific suggestions of how to implement this intervention mentioned in interviews and/or focus group have been added in table 5. table 5 specific suggestions of intervention d 4.2.2 peer feedback types amongst the participants, peer feedback is considered as an appropriate and scalable intervention to activate the mechanism “asking and providing relevant feedback” and, once delivered in dialogue form, the mechanism “probing back and forth.” however, it may also trigger other useful mechanisms. dominant mechanisms have been indicated in italics below. different types of peer feedback to promote deep learning can be distinguished. intervention e: encourage asynchronous oral peer feedback (audio or video) instructors encouraged the involvement of peers in feedback processes and invited them to provide their feedback in spoken form. even though nearly all interviewed instructors and students used only written feedback in online education, several instructors and students mentioned the expected potential of oral peer feedback. it was quicker and more personal, and using the voice and inflection made it easier to be critical, to deliver bad and good news, and to add nuances. in contrast to the written feedback, it added the richness of tone of voice, and, by using video even of facial expressions, it made students feel personally committed and more connected to the course material. and according to one of the instructors, students listened to it more, because they typically accessed it on their smartphones and their tablets. specific suggestions of how to implement this intervention mentioned in interviews and/or focus group have been added in table 6. table 6 specific suggestions of intervention e intervention f: encourage written asynchronous peer feedback teaching students how to provide written peer feedback that is focused on deep learning and is provided as a dialogue was recommended by both instructors and students and confirmed by members of the focus group. this creates awareness about the type of feedback that can be given and stimulates critical thinking, questioning, and reflecting. when providing feedback in written form, there is more time to think about it thoroughly and to formulate it carefully. doing so in a dialogue form by probing back and forth, students can ask each other questions, reflect, and respond to each other, which encourages deep learning. compared to oral feedback, written peer feedback was found to promote deep learning even more effectively because of the more precise type of feedback students are able to provide. according to both the interviewed instructors and students, students often learn more from providing feedback than from receiving feedback. specific suggestions of how to implement this intervention mentioned in interviews and/or focus group have been added in table 7. table 7 specific suggestions of intervention f intervention g: support group work both instructors and students mentioned online group work as a learning method in which deep learning could be promoted through feedback. instructors steered the students towards different group assignments and stimulated personal commitment and interaction. by doing so, students felt motivated and encouraged to be engaged, to reflect and to explicate what they have learned. the instructor taught students to suspend their opinions to create a dialogue and to construct questions in such a way that higher-order thinking is necessary for the others to answer the questions. this not only stimulated back and forth probing, but also made students feel personally committed, which may have motivated students to work just a little harder. specific suggestions of how to implement this intervention mentioned in interviews and/or focus group have been added in table 8. table 8 specific suggestions of intervention g intervention h: provide organized synchronous feedback instructors organized sessions in which students discussed their work and their feedback. the synchronous character made students feel more personally committed than written feedback. the simultaneous communication enabled back and forth probing. the prompt feedback gave the students the opportunity to adjust their performance immediately. both students and instructors said that they appreciated the opportunity to ask for immediate clarification in such a way that the feedback process became a dialogue. specific suggestions of how to implement this intervention mentioned in interviews and/or focus group have been added in table 9. table 9 specific suggestions of intervention h 4.2.3 automatic feedback intervention i: add scenario-based multiple choice questions add scenario-based multiple choice questions, aimed at deep learning, to the course design. scenario-based multiple choice questions contain follow-up questions and may be represented by a tree structure. questions should be asked in such a way that students are encouraged to synthesize information, draw conclusions, and support findings, and reflect on them. online students appreciate multiple choice questions because of the active method and the immediate feedback that provides them with understanding of their own learning process. even though none of the respondents have experience with multiple choice questions specifically aimed at deep learning, most think it will be possible. it requires much precision and thinking very carefully about the questions and the responses, and will therefore be time-consuming during the development phase. however, if the number of students is large enough, the time investment will be worth it. a specific suggestion of how to implement this intervention mentioned in interviews and focus group have been added in table 10. table 10 specific suggestion of intervention i 5. discussion promoting deep learning is an important task for higher education, which is increasingly conducted online. spocs may be a form of online learning that has much potential for deep learning because of its small groups and relatively high number of interaction possibilities. in a previous study (filius et al., 2018), we showed that instructors experience specific challenges when trying to promote deep learning in spocs. this previous study resulted in a description of five main challenges: alignment of learning activities, insights into student needs, adaptivity in teaching strategy, social cohesion, and creating dialogue. to meet these challenges, the incorporation of feedback may have significant potential. therefore, the aim of this study was to provide scalable design propositions for instructors in spocs to promote deep learning through online feedback. design propositions have been formulated according to the cimo-logic. specific attention was paid to the mechanisms behind the interventions as they are central to the plausibility of a design principle (van aken, 2004). the results match with the categorization used by the online learning interaction model of ke and xie (2009), which also aims at deep learning. their three categories could be extended by the mechanisms found in this study. we suggest that interaction in the category “social interaction” may promote deep learning if it makes students feel personally committed. to make them feel that way, it helps students to receive adapted and individualized feedback. online learning interaction in the category “knowledge construction” should, in order to promote deep learning, be aimed at probing back and forth, as a dialogical process. this is in line with previous studies such as the work of rakoczy, harks, klieme, blum, and hochweber (2013), who indicate that receiving feedback is just as important as providing feedback. to fully exploit the feedback, students should be actively engaged in the feedback dialogue. in that same “knowledge construction” category we argue that it is important for students to learn how and when to ask for relevant feedback. this is supported by nicol (2010), who argues that getting students to request feedback, to respond to feedback, and to actively connect feedback to their assignments might result in students’ paying more attention to, and being able to use, instructor feedback. geitz et al. (2015) suggest that this may be explained by the fact that learning how and when to ask for exactly what type of feedback helps students to be more in control and to add personal meaning to the feedback. quality of feedback is important, but the quality of the interaction with the feedback may be even more important. moreover, it helps instructors to manage their time effectively. regarding the third category, “regulation of learning,” interaction to promote deep learning is especially useful when it provides students with more insight into their own learning process. this has been confirmed by other research: students must be equipped with the skills to think for themselves, to set their own goals, and to make improvements to their work while it is being produced (andrade, du, & mycek, 2010; molloy & boud, 2013; narciss, 2013; d. r. sadler, 2013). students need to develop awareness and responsiveness so they can detect anomalies or problems for themselves (d. r. sadler, 2013). according to topping (1998), these self-regulation skills provide students with skills that they will need not only during their higher education, but also during their future life. students who are more effective at selfregulation produce better feedback or are more able to use the feedback they generate to achieve their desired goals (butler & winne, 1995). interestingly, it is shown that peer feedback helps students to obtain these self-regulation skills even better than instructor feedback does (planar & moya, 2016). and peer feedback may be useful for instructors to manage their time effectively. with the current high student-staff ratio, it may be difficult for instructors to engage in dialogue with students. therefore we specifically aimed for scalable interventions. this possibly excluded several instructor-student interventions. results suggest that scalability occurs in three categories of interventions. the first category concerns feedback management, which seeks to reduce the range of tasks of the instructor or to better facilitate the instructor. as feedback should be adaptive in order to be effective (nicol, 2010) and adaptive feedback is considered to be a challenge in the specific context of spocs (filius et al., 2018), the interventions in this category offer possibilities to allow the provision of adaptive feedback to become more feasible. the second category concerns peer feedback, which has been shown to have much potential for promoting deep learning. boud et al. (1999) suggested that working with peers rather than with the instructor may promote higher-order thinking. anderson and rourke (2002) confirmed that discussions by peers were useful in achieving higher-order, but not lower-order, learning objectives because the controversial perspectives offered by other peers disturbed students’ initial understanding of the content and therefore prompted them to process it thoroughly. based on the results of this study, we suggest that the mechanisms found may play an important role in determining whether the peer feedback interventions will lead to deep learning. automatic feedback is the third category. although automatic feedback can be provided for most constructed response items (benson, 2010), the use to specifically promote deep learning has, to the best of our knowledge, not yet been fully explored. since both students and instructors have expectations that this may lead to deep learning, this paper may lead to further investigation of the use of automatic feedback to promote deep learning in spocs. how might instructors use the findings in this paper? one practical proposal is that instructors in spocs examine current feedback practices in relation to the interventions and mechanisms as described above. especially, we expect that combinations of several feedback interventions, triggering multiple mechanisms, may support deep learning in spocs. an examination of this kind might help identify where feedback practices might be strengthened. however, the design propositions presented here do not exhaust all interventions that instructors might perform to promote deep learning in spocs. they merely provide a starting point and emphasize the importance of framing feedback as a dialogical process with active engagement of students. the research challenge is to refine these design propositions, identify gaps, and gather further evidence about the potential of feedback to promote deep learning. learning in an online environment can constitute a positive springboard to the new role that instructors need to take on in an online education model where the student is at the center of the learning process (planar & moya, 2016). given this new role, it is crucial to develop and analyze learning methods that enable a greater amount of dialogue among the students in the learning process (planar & moya, 2016). this present study provides relevant insights into how and why deep learning can be promoted in spocs. since this study is exploratory in nature, we recommend to focus subsequent research on examining the findings on a larger scale. for a follow up study, we also recommend to aim at different methods for assessing deep learning, such as grades or academic performance in general. further on, we chose deliberately to focus primarily on the perspective of the instructors. therefore, the perspective of the students has been used only as a supplement and thus we have limited the number of students to four. a next study could include the perspective of the students and compare them with the findings in this study. future research should be aimed at how feedback interventions are better suitable for promoting deep learning while also taking into account the specific learning mechanisms that should be activated within the different contexts and the workload that instructors experience. future research could also include combinations with other instruments other than feedback, such as collaborative assignments and integrate earlier research into, for example, concept maps (hay, 2015), cross-cultural chat (osman & herring, 2007), podcasting (pegrum, bartle, longnecker, 2014) and online asynchronous discussions (du, harvard & li, 2005). as we explore the relatively young field of spocs, the results of this study show that feedback as a dialogical process may contribute to solving the current challenges of instructors in spocs to achieve deep learning with their students. specific attention has been paid to the mechanisms that are internal to the student and can be triggered by feedback interventions. findings concerning the mechanisms sheds light on why interventions lead to the desired outcome, which is deep learning. it is essential to continue this line of research and to explore systematically the implementation of the design principles, both on learning processes and on learning performance. acknowledgements the authors would like to thank rianne bouwmeester phd keypoints scalable design propositions for instructors in spocs to promote deep learning through online feedback four mechanisms by which deep learning can be achieved student mechanisms triggered by feedback interventions fostering of increased dialogue among the students quality of the interaction with feedback more important than quality of feedback itself references aharony, n. (2006). the use of deep and surface learning strategies among students learning english as a foreign language in an internet environment. british journal of educational psychology, 76(4), 851-866. http://dx.doi.org/10.1348/000709905x79158 akkerman, s., admiraal, w., brekelmans, m., & oost, h. (2008). auditing quality of research in social sciences. quality & quantity, 42 (2), 257-274. http://dx.doi.org/10.1007/s11135-006-9044-4 allan, r., & bentley, s. (2012, april). feedback mechanisms: efficient and effective use of technology or a waste of time and effort? paper presented at the stem annual conference,imperial college, london. anderson, t., & rourke, l. (2002). using peer teams to lead online discussions. journal of interactive media in education, 1, 1-21. andrade, h. l., du, y., & mycek, k. (2010). rubric-referenced self-assessment and middle school students’ writing. assessment in education: principles, policy & practice, 17(2), 199-214. http://dx.doi.org/10.1080/09695941003696172 askew, s., & lodge, c. (2000). gifts, ping-pong and loops–linking feedback and learning. in s. askew (ed.) (1st ed.)., feedback for learning(pp. 1-17). london: routledge, falmer. http://dx.doi.org/10.4324/9780203017678 athanassiou, n., mcnett, j. m., & harvey, c. (2003). critical thinking in the management classroom: bloom's taxonomy as a learning tool. journal of management education, 27(5), 533-555. http://dx.doi.org/10.1177/1052562903252515 benson, a. d. (2010). assessing participant learning in online environments. facilitating learning in online environments: new directions for adult and continuing education , number 100, 103, 69. http://dx.doi.org/10.1002/ace.120 berliner, d. c. (2001). learning about and learning from expert teachers. international journal of educational research, 35(5), 463-482. http://dx.doi.org/10.1016/s0883-0355(02)00004-6 biggs, j. (1999). what the student does: teaching for enhanced learning. higher education research & development, 18(1), 57-75. http://dx.doi.org/10.1080/07294360.2012.642839 biggs, j., kember, d., & leung, d. y. (2001). the revised two-factor study process questionnaire: r-spq-2f. british journal of educational psychology, 71(1), 133-149. http://dx.doi.org/10.1348/000709901158433 biggs, j., & tang, c. (2011). teaching for quality learning at university, berkshire: the society for research into higher education and open university press. bogdan, r., & biklen, s. k. (2003). qualitative research for education: an introduction to theories and methods. new york: pearson. booth, p., luckett, p., & mladenovic, r. (1999). the quality of learning in accounting education: the impact of approaches to learning on academic performance. accounting education, 8(4), 277-300. http://dx.doi.org/10.1080/096392899330801 boud, d. (2001). peer learning and assessment. in d. boud, r. cohen, & j. sampson (eds.), peer learning in higher education(1 sted., pp. 67-84). london: kogan page limited. boud, d., cohen, r., & sampson, j. (1999). peer learning and assessment. assessment & evaluation in higher education, 24 (4), 413-426. http://dx.doi.org/10.1080/0260293990240405 boud, d., & molloy, e. (2013). rethinking models of feedback for learning: the challenge of design. assessment & evaluation in higher education, 38(6), 698-712. http://dx.doi.org/10.1080/02602938.2012.691462 bronkhorst, l. h., meijer, p. c., koster, b., & vermunt, j. d. (2011). fostering meaning oriented learning and deliberate practice in teacher education. teaching and teacher education, 27(7), 1120-1130. http://dx.doi.org/10.1016/j.tate.2011.05.008 brouwer, p., brekelmans, m., nieuwenhuis, l., & simons, r. (2012). fostering teacher community development: a review of design principles and a case study of an innovative interdisciplinary team. learning environments research, 15(3), 319-344. http://dx.doi.org/10.1007/s10984-012-9119-1 butler, d. l., & winne, p. h. (1995). feedback and self-regulated learning: a theoretical synthesis. review of educational research, 65(3), 245-281. http://dx.doi.org/10.3102/00346543065003245 carless, d., salter, d., yang, m., & lam, j. (2011). developing sustainable feedback practices. studies in higher education, 36 (4), 395-407. http://dx.doi.org/10.1080/03075071003642449 cohen, l., manion, l., & morrison, k. (2013). research methods in education. london: routledge. http://dx.doi.org/10.4324/9781315456539 cresswell, j. (2007). qualitative inquiry and research design: choosing among five perspectives. london: sage publishing. crook, a., mauchline, a., maw, s., lawson, c., drinkwater, r., lundqvist, k., ... & park, j. (2012). the use of video technology for providing feedback to students: can it enhance the feedback experience for staff and students? computers & education, 58(1), 386-396. http://dx.doi.org/10.1016/j.compedu.2011.08.025 davies, r., & berrow, t. (1998). an evaluation of the use of computer supported peer review for developing higher-level skills. computers & education, 30(1), 111-115. http://dx.doi.org/10.1016/s0360-1315(97)00086-9 dennen, v. p., aubteen darabi, a., & smith, l. j. (2007). instructor-learner interaction in online courses: the relative perceived importance of particular instructor actions on performance and satisfaction. distance education, 28(1), 65-79. http://dx.doi.org/10.1080/01587910701305319 denyer, d., tranfield, d., & van aken, j. e. (2008). developing design propositions through research synthesis. organization studies, 29 (3), 393-413. http://dx.doi.org/10.1177/0170840607088020 dobber, m., akkerman, s. f., verloop, n., admiraal, w., & vermunt, j. d. (2012). developing designs for community development in four types of student teacher groups. learning environments research, 15(3), 279-297. http://dx.doi.org/10.1007/s10984-012-9116-4 dochy, f., segers, m., & sluijsmans, d. (1999). the use of self-, peer and co-assessment in higher education: a review. studies in higher education, 24(3), 331-350. http://dx.doi.org/10.1080/03075079912331379935 du, j., havard, b., & li, h. (2005). dynamic online discussion: task‐oriented interaction for deep learning. educational media international, 42(3), 207-218. http://dx.doi.org/10.1080/09523980500161221 entwistle, n. j. (1991). approaches to learning and perceptions of the learning environment. higher education, 22(3), 201-204. http://dx.doi.org/10.1007/bf00132287 fan, x., miller, b. c., park, k. e., winward, b. w., christensen, m., grotevant, h. d., & tai, r. h. (2006). an exploratory study about inaccuracy and invalidity in adolescent self-report surveys. field methods, 18(3), 223-244. http://dx.doi.org/10.1177/152822x06289161 filius, r.m., de kleijn, r.a.m., uijl, s.g., prins, f.j., van rijen, h.v.m. and grobbee, d.e. (2018). challenges concerning deep learning in spocs. international journal of technology enhanced learning, 10(1-2), 111-127. http://dx.doi.org/10.1504/ijtel.2018.088341 garrison, d. r., anderson, t., & archer, w. (2000). critical inquiry in a text-based environment: computer conferencing in higher education. internet and higher education, 2(2–3), 87−105. http://dx.doi.org/10.1016/s1096-7516(00)00016-6 garrison, d. r., anderson, t., & archer, w. (2001). critical thinking, cognitive presence, and computer conferencing in distance education. american journal of distance education, 15(1), 7-23. http://dx.doi.org/10.1080/08923640109527071 garrison, d. r., anderson, t., & archer, w. (1999). critical inquiry in a text-based environment: computer conferencing in higher education. the internet and higher education, 2(2), 87-105. http://dx.doi.org/10.1016/s1096-7516(00)00016-6 geitz, g., brinke, d. j., & kirschner, p. a. (2015). goal orientation, deep learning, and sustainable feedback in higher business education. journal of teaching in international business, 26(4), 273-292. http://dx.doi.org/10.1080/08975930.2015.1128375 guba, e. g. (1981). criteria for assessing the trustworthiness of naturalistic inquiries. ectj, 29(2), 75. guest, g., bunce, a., & johnson, l. (2006). how many interviews are enough? an experiment with data saturation and variability. field methods, 18(1), 59-82. hall, m., ramsay, a., & raven, j. (2004). changing the learning environment to promote deep learning approaches in first-year accounting students.accounting education, 13(4), 489-505. http://dx.doi.org/10.1080/0963928042000306837 hattie, j., & timperley, h. (2007). the power of feedback. review of educational research, 77(1), 81-112. http://dx.doi.org/10.3102/003465430298487 hay, d.b. (2007). using concept maps to measure deep, surface and non-learning outcomes. studies in higher education, 32(1), 39-57. http://dx.doi.org/10.1080/03075070601099432 hounsell, d. 1997. contrasting conceptions of essay-writing. in the experience of learning, edited by f. marton, d. hounsell, and n. entwistle, 106–125. edinburgh: scottish academic press. ilgen, d. r., fisher, c. d., & taylor, m. s. (1979). consequences of individual feedback on behavior in organizations. journal of applied psychology, 64(4), 349. http://dx.doi.org/10.1037/0021-9010.64.4.349 ke, f., & xie, k. (2009). toward deep learning for adult students in online courses. the internet and higher education, 12(3), 136-145. http://dx.doi.org/10.1016/j.iheduc.2009.08.001 kluger, a. n., & denisi, a. (1996). the effects of feedback interventions on performance: a historical review, a meta-analysis, and a preliminary feedback intervention theory. psychological bulletin, 119(2), 254-284. http://dx.doi.org/10.1037/0033-2909.119.2.254 lin, s. s., liu, e. z., & yuan, s. (2001). web-based peer assessment: feedback for students with various thinking-styles. journal of computer assisted learning, 17(4), 420-432. http://dx.doi.org/10.1046/j.0266-4909.2001.00198.x lindblom-ylänne, s. (1999). studying in a traditional medical curriculum-study success, orientations to studying and problems that arise. helsinki: printing house. lynch, r., mcnamara, p. m., & seery, n. (2012). promoting deep learning in a teacher education programme through self-and peer-assessment and feedback.european journal of teacher education, 35(2), 179-197. http://dx.doi.org/10.1080/02619768.2011.643396 marton, f., & saljo, r. (1997). approaches to learning. in f. marton, d. hounsell, & n. j. entwistle (eds.), the experience of learning: implications for teaching and studying in higher education (1sted., pp. 39-58). edinburgh: scottish academic press. midgley, g. (2000). systemic intervention. in midgley, g. (ed.),systemic intervention: philosophy, methodology, and practice(1 sted., pp. 113-133). new york: springer us. http://dx.doi.org/10.1007/978-1-4615-4201-8 molloy, e., & boud, d. (2013). changing conceptions of feedback. in e. molloy & d. boud (eds.), feedback in higher and professional education: understanding it and doing it well (1sted., pp. 11-33). london: routledge. moon, j. a. (2013). reflection in learning and professional development: theory and practice. london: routledge. http://dx.doi.org/10.4324/9780203822296 narciss, s. (2013). designing and evaluating tutoring feedback strategies for digital learning. digital education review,23, 7-26. nicol, d. (2009). assessment for learner self-regulation: enhancing achievement in the first year using learning technologies. assessment & evaluation in higher education, 34(3), 335-352. http://dx.doi.org/10.1080/02602930802255139 nicol, d. (2010). from monologue to dialogue: improving written feedback processes in mass higher education. assessment & evaluation in higher education, 35(5), 501-517. http://dx.doi.org/10.1080/02602931003786559 nicolls, g. (2002). developing teaching and learning in higher education.london: routeledge falmer. http://dx.doi.org/10.4324/9780203469231 osman, g., & herring, s. c. (2007). interaction, facilitation, and deep learning in cross-cultural chat: a case study. the internet and higher education, 10(2), 125-141. http://dx.doi.org/10.1016/j.iheduc.2007.03.004 palys, t. (2008). purposive sampling.the sage encyclopedia of qualitative research methods, 2, 697-698. pegrum, m., bartle, e., & longnecker, n. (2015). can creative podcasting promote deep learning? the use of podcasting for learning content in an undergraduate science unit. british journal of educational technology,46(1), 142-152. http://dx.doi.org/10.1111/bjet.12133 planar, d., & moya, s. (2016). the effectiveness of instructor personalized and formative feedback provided by instructor in an online setting: some unresolved issues. electronic journal of e-learning, 14(3), 196-203. plomp, t. & nieveen, n. (eds.) (2009). an introduction to educational design research: proceedings of the seminar conducted at the east china normal university, shanghai. enschede, the netherlands: slo – the netherlands institute for curriculum development. poortman, c., & schildkamp, k. (2012). alternative quality standards in qualitative research? quality & quantity, 46(6), 1727-1751. http://dx.doi.org/10.1007/s11135-011-9555-5 powell, r. a., & single, h. m. (1996). focus groups. international journal for quality in health care, 8(5), 499-504. http://dx.doi.org/10.1093/intqhc/8.5.499 rakoczy, k., harks, b., klieme, e., blum, w., & hochweber, j. (2013). written feedback in mathematics: mediated by students’ perception, moderated by goal orientation. learning and instruction, 27, 63-73. http://dx.doi.org/10.1016/j.learninstruc.2013.03.002 ramsden, p. (1992). learning to teach in higher education. london: routledge. http://dx.doi.org/10.4324/9780203413937 ramsden, p., & entwistle, n. (1983). understanding student learning.kent: croom helm. ramsden, p., & moses, i. (1992). associations between research and teaching in australian higher education. higher education, 23(3), 273-295. http://dx.doi.org/10.1007/bf00145017 richardson, j. c., koehler, a. a., besser, e. d., caskurlu, s., lim, j., & mueller, c. m. (2015). conceptualizing and investigating instructor presence in online learning environments. the international review of research in open and distributed learning, 16 (3), 256-297. http://dx.doi.org/10.19173/irrodl.v16i3.2123 rushton, a. (2005). formative assessment: a key to deep learning? medical teacher, 27(6), 509-513. http://dx.doi.org/10.1080/01421590500129159 sadler, d. r. (2013). opening up feedback. in s. merry, m. price, d. carless, & m. taras (eds.), reconceptualising feedback in higher education: developing dialogue with students (1sted., pp. 54-63). london: routledge. http://dx.doi.org/10.4324/9780203522813 sadler, p. m., & good, e. (2006). the impact of self-and peer-grading on student learning. educational assessment, 11(1), 1-31. http://dx.doi.org/10.1207/s15326977ea1101_1 topping, k. (1998). peer assessment between students in colleges and universities. review of educational research, 68(3), 249-276. http://dx.doi.org/10.3102/00346543068003249 trigwell, k., prosser, m., & waterhouse, f. (1999). relations between teachers’ approaches to teaching and students’ approaches to learning. higher education, 37(1), 57-70. uijl, s., filius, r., & ten cate, o. (2017). student interaction in small private online courses. medical science educator, 1-6. http://dx.doi.org/10.1007/s40670-017-0380-x van aken, j. e. (2004). management research based on the paradigm of the design sciences: the quest for field-tested and grounded technological rules. journal of management studies, 41(2), 219-246. http://dx.doi.org/10.1111/j.1467-6486.2004.00430.x van aken, j. e. (2007). developing organization studies as an applied science using a triple learning approach. paper presented at the third organization studies summer workshop, greece. van den akker, j. j. h. (1999). principles and methods of development research. in van den akker, j. j. h., branch, r. m. gustafson, k., nieveen, n. & plomp, t. (eds.),design approaches and tools in education and training(1 sted., pp. 1-14). dordrecht: springer netherlands. http://dx.doi.org/10.1007/978-94-011-4255-7 vermunt, j. d., & verloop, n. (1999). congruence and friction between learning and teaching. learning and instruction, 9(3), 257-280. http://dx.doi.org/10.1016/s0959-4752(98)00028-0 microsoft word rantavuori et al_publication.docx           frontline  learning  research  vol.4  no.  3  (2016)  1  -­‐  27   issn  2295-­‐3159       learning actions, objects and types of interaction: a methodological analysis of expansive learning among pre-service teachers juhana rantavuori1, yrjö engeström, lasse lipponen university of helsinki, finland article received 8 may / revised 24 march / accepted 6 april / available online 10 may   abstract the paper analyzes a collaborative learning process among finnish pre-service teachers planning their own learning in a self-regulated way. the study builds on culturalhistorical activity theory and the theory of expansive learning, integrating for the first time an analysis of learning actions and an analysis of types of interaction. we examine the theory of expansive learning as a possible conceptual and methodological framework for understanding this type of collaborative learning. the task of the paper is primarily methodological. we believe that cultural-historical activity theory needs to be turned into methods and procedures of systematic empirical analysis, and this article examines one such methodological solution. at the same time, we aim to uncover some substantive dynamics of expansive learning in collaborative teacher education oriented at openended problems and tasks. an almost complete expansive mini-cycle of learning actions appeared in the pre-service teachers’ meeting. however, an analysis of the steps of formation of the shared object revealed a more complex iterative process. as the expansive learning process moved epistemically from questioning to analysis, modeling and implementation, it also moved interactionally from coordination to cooperation and communication. yet there was no mechanical correspondence between specific learning actions and specific types of interaction. transitions and disturbances were crucial for the dynamics of expansive learning. a full assessment of a potentially expansive minicycle of learning calls for extending the time scale of the analysis. keywords: activity theory; expansive learning; learning actions; types of interaction; object; disturbances                                                                                                                           1  corresponding author: juhana rantavuori, center for research on activity, development, and learning, institute of behavioural sciences, faculty of behavioural sciences, po box 9 (siltavuorenpenger 1a), fi-00014 university of helsinki, finland. e-mail: juhana.rantavuori@helsinki.fi. doi: http://dx.doi.org/10.14786/flr.v4i3.174   rantavuori  et  al       | f l r     2   1. introduction open-ended and problem-based collaborative learning is becoming an increasingly important challenge for many contexts in which learners face complex problems for which pre-existing standard solutions are not sufficient (bereiter & scardamalia, 1993). we argue that it is not enough to promote collaborative and problem-oriented learning in general. theoretically ambitious models and empirically rigorous methods are needed for the design and assessment of such learning processes (see goldman, 2014). in this paper, we will analyze a collaborative learning process among finnish pre-service teachers. in this particular process, the pre-service teachers were responsible for planning their own learning actions and goals. we examine the theory of expansive learning (engeström, 2015) as a possible conceptual and methodological framework for understanding this type of learning. more generally, our study contributes to research on learning and interaction in activity systems, especially to how learning and interaction are connected in open-ended problem solving. activity systems are systems where people engage in solving problems or making or designing something (greeno, 2011; greeno & engeström, 2014). they are “dynamic, open, semiotic system(s) of meaningful actions and meaning-making processes” (lemke, 1990, p. 191). an activity system can be as small as an individual working with a computer, or as large as an organization having hundreds of employees. in our case, the activity system is a group of pre-service teachers, working on an open-ended problem solving task in a selfregulated way. the task of the paper is primarily methodological. we believe that cultural-historical activity theory needs to be turned into methods and procedures of systematic empirical analysis. therefore, the aim of the paper is to contribute to the construction of a methodology for analyzing dynamics of expansive learning. a new methodological framework created in this study is tested in the analysis of a planning meeting of a preservice teacher group. our study is focused on two important aspects of expansive learning, namely types and sequences of expansive learning actions (engeström & sannino, 2010) and types and sequences of object-oriented interaction (engeström, 2008; fichtner, 1984; raiethel, 1983). our aim is to understand what kinds of learning actions pre-service teachers conduct and in what types of interaction they engage in a collaborative learning process characterized by self-regulation and open-ended problem-solving. expansive learning actions have been studied in detail previously (e.g., engeström, rantavuori, & kerosuo, 2013; foot, 2001; nilsson, 2003; seppänen, 2004), and so have types of object-oriented interaction (e.g., de lange, 2011; saari; 1995). however, no studies have thus far combined learning actions and types of interaction into an integrated analysis. to fully understand the nature of open-ended and problem-based collaborative learning, and to develop the methodology of expansive learning, we need to combine these two analyses. no studies have done this up until the present. studies of expansive learning have often been based on interventions, such as change laboratories (virkkunen & newnham, 2013), deliberately designed to implement expansive learning (e.g., engeström et al., 2013). this was not the case in the process we analyze in this paper. in this sense, our case resembles an earlier study of innovative learning in industrial work teams (engeström, 2008, pp. 118–168). the assumption of these studies is that features of expansive learning may be found in processes in which the learners face a problem or task that needs to be defined by the learners themselves and has no predefined procedure to follow or correct solution to aim at. furthermore, these studies see an inherent tension and conflict of motives in these learning processes between the safe and easy but probably rather unproductive option of following the available routine script in dealing with the task on the one hand, and the risky and difficult but possibly very productive option of turning the task into a new, expanded object and way of working on the other hand. our study examines a single learning session. full-fledged cycles of expansive learning consist of mini-cycles which may be detected and fostered within single learning sessions or other rantavuori  et  al       | f l r     3   compact sequences of learning efforts. thus, from the point of view of the theory of expansive learning, our study addresses three interrelated methodological challenges: (a) combining and integrating for the first time an analysis of learning actions and an analysis of types of interaction, (b) examining possible features of expansive learning in a process which was not designed to accomplish expansive learning by deliberate intervention, and (c) examining possible evidence for a mini-cycle of expansive learning within a single learning session. in other words, the task of this article is to explore and elaborate on the explanatory potential of the theory of expansive learning in a context of learning to which it has not been usually applied, and to develop methodological tools for examining the potential of the theory in a systematic manner. added to this, the task of the article is also to show which role the mutual interaction between the participants plays in the expansive learning process. in what follows, we will first present the theoretical framework, the methodology used in the study, and the research questions. after that we describe the context of the study and the data collected. we then analyze our data in four sections, each devoted to one of our four research questions. finally, we discuss our findings and consider their methodological implications for the framework of expansive learning and for research on learning more generally. 2. theoretical framework 2.1 theory of expansive learning sfard (1998) suggested that there are two basic metaphors of learning competing for dominance: the acquisition metaphor and the participation metaphor. the key dimension underlying sfard’s dichotomy is derived from the question: is the learner to be understood primarily as an individual or as a community? this is an important dimension, largely inspired by the notion of community of practice put forward by lave and wenger (1991) and wenger (1998). however, an attempt to construct a one-dimensional conceptual space for the identification, analysis and comparison of theories is bound to eliminate too much of the complexity of the field of learning. the theory of expansive learning puts the primacy on communities as learners, on transformation and creation of culture, on horizontal movement and hybridization, and on the formation of theoretical concepts. in fact, from the point of view of expansive learning, both acquisition-based and participationbased approaches share much of the same conservative bias. both have little to say about transformation and creation of culture. both acquisition-based and participation-based approaches, depict learning primarily as one-way movement from incompetence to competence, with little serious analysis devoted to horizontal movement and hybridization. acquisition-based approaches may ostensibly value theoretical concepts, but their very theory of concepts is quite uniformly empiricist and formal (davydov, 1990). participation-based approaches are commonly suspicious if not hostile toward the formation of theoretical concepts, largely because these approaches, too, see theoretical concepts mainly as formal ‘bookish’ abstractions. so the theory of expansive learning must rely on its own metaphor: expansion. the core idea is qualitatively different from both acquisition and participation. in expansive learning, learners learn something that is not yet there. in other words, the learners construct a new object and concept for their collective activity, and implement this new object and concept in practice. traditional modes of learning deal with tasks in which the contents to be learned are well known ahead of time by those who design, manage, and implement various programs of learning. when whole collective activity systems, such as work processes and organizations, need to redefine themselves, traditional modes of learning are not enough. nobody knows exactly what needs to be learned. the design of rantavuori  et  al       | f l r     4   the new activity and the acquisition of the knowledge and skills it requires are increasingly intertwined. in expansive learning activity, they merge. relying on activity theory, the theory of expansive learning is foundationally an object-oriented theory. in other words, the object is both resistant raw material and the future-oriented purpose of an activity. the object is the true carrier of the motive of the activity. thus, in expansive learning activity, motives and motivation are not sought primarily inside individual subjects – they are in the object to be transformed and expanded. in educational settings, the students’ object is a contradictory unity of meaningful knowledge (use value) and grades (exchange value). a powerful object of learning has expansive potential to go beyond the exchange value, being typically an open-ended problem or challenge that has relevance for the learners not limited to reproducing predefined correct answers. such an object of learning typically also goes beyond verbal formulations, requiring transformative material actions of experimentation, modeling, and implementation in practice. the theory of expansive learning is based on the dialectics of ascending from the abstract to the concrete (engeström & sannino, 2010). this is a method of grasping the essence of an object by tracing and theoretically reproducing the logic of its development, that is, its historical formation through the emergence and resolution of its inner contradictions. a new theoretical idea or concept is initially produced in the form of an abstract, simple explanatory relationship, a germ cell. this initial abstraction is enriched and transformed step-by-step into a concrete system of multiple, constantly developing manifestations. in an expansive learning cycle, the initial simple idea is transformed into a complex object, a new form of practice. a successful expansive cycle produces a new theoretical concept – theoretically grasped practice – concrete in its systemic richness and multiplicity of manifestations. the expansive cycle begins with individual subjects questioning the accepted practice, and it gradually expands into a collective effort. in educational contexts, the most well-known example of ascending from the abstract to the concrete is davydov’s (1990) work on elementary school mathematics learning. for davydov, the germ cell of mathematics is real number, which is a particular case of a general relationship of quantities, where one of them is taken as a measure for computing the other. a number is obtained by the general formula a/c = n, in which n is any number, a is any object represented as a quantity, and c is any measure (davydov, 1990, pp. 361–362). from working out and operating with this foundational relationship, or abstract germ cell, davydov built a whole curriculum that resulted in a mastery of a rich and concrete diversity of mathematical phenomena and tasks (schmittau & morris, 2004). in subsequent studies of expansive learning, the learning challenge has often been more problematic, stemming from contradictions that need to be resolved. in these studies, the germ cell is initially not known by the instructor-interventionists themselves; it has to be discovered and modeled by the participants investigating and transforming their activity and knowledge domain (engeström & sannino, 2010). expansive learning may be described as a stepwise process that involves seven phases called learning actions. together these actions form an expansive cycle. this sequential model should be understood as an idealized tool for analyzing elements of expansive learning; real cycles of expansive learning do not neatly follow the order depicted in the theoretical model. process theories of learning are unavoidable to some extent prescriptive in that they advocate some optimal or desirable model of the learning process. this carries the risk of self-fulfilling prophecy, that is, as design-oriented researcher may impose his or her theoretical model on learners and instructors and seek confirmation for the model from evidence stemming from such pre-designed practice. there are good ways to keep this tendency in check (engeström & sannino, 2012). in the present study, the learning process was not designed to follow the theoretical model of expansive learning to begin with. an ideal-typical sequence of learning actions in an expansive cycle can be described as follows (engeström & sannino, 2010, p. 7). rantavuori  et  al       | f l r     5   the first action of an expansive cycle is that of questioning, criticizing, or rejecting some aspects of accepted practice and existing wisdom. the second action is that of analyzing the situation. analysis involves mental, discursive, or practical transformation of the situation in order to discover causes or explanatory mechanisms. analysis evokes “why” questions and explanatory principles. one type of analysis is historicalgenetic; it seeks to explain the situation by tracing its origination and evolution. another type of analysis is actual-empirical; it seeks to explain the situation by constructing a picture of its inner systemic relations. the third action is that of modeling the newly found explanatory relationship in some publicly observable and transmittable form. this means constructing an explicit, simplified model of the new idea that explains and offers a solution to the problematic situation. the fourth action is that of examining the model, running, operating, and experimenting on it in order to fully grasp its dynamics, potentials, and limitations. the fifth action is that of implementing the model, concretizing it by means of practical applications, enrichments, and conceptual extensions. the sixth and seventh actions are those of reflecting on and evaluating the process and consolidating its outcomes into a new, stable form of practice. the model of expansive learning is useful when we try to understand open-ended learning processes in which the problem and its solution are not predefined, and the participants must learn something that “is not yet there”, that is, to generate and appropriate culturally new practices and knowledge. expansive learning has mostly been studied in relatively long-term transformations and interventions. however, “largescale cycles involve numerous smaller cycles of learning actions” (engeström & sannino, 2010). such a mini-cycle may take place within a single intensive meeting of a group charged with a task of analyzing and solving a problem important for the development of its overall activity (e.g., engeström, 2008). although the theory of expansive learning proposes that full-fledged sequences of expansive learning actions typically take the shape of relatively predictable cycles, the cycle of expansive learning is not a universal formula of phases or stages. in fact, one probably never finds a concrete collective learning process which purely follows the ideal-typical model. the model is a heuristic conceptual tool derived from the logic of ascending from the abstract to the concrete. every time one examines or facilitates a potentially expansive learning process with the help of the model, one tests, criticizes and hopefully enriches the theoretical ideas of the model. learning processes are never purely expansive. they contain both expansive and non-expansive phases, steps forward and back, and digressions from expanding the object of activity (engeström et al., 2013). in the study of innovative learning in industrial work teams (engeström, 2008, pp. 118–168), two such non-expansive actions were identified, namely formulating/debating a problem and reinforcing existing practice. a change laboratory process in a finnish library (engeström et al., 2013) revealed three nonexpansive actions, namely informing, clarifying, and summarizing. in this study we followed the criteria of these previous studies for identifying the non-expansive learning actions. in expansive learning the emergence of a new expanding object is decisive. if such a new object was not found, the learning action was identified as non-expansive. these actions were then named descriptively, on the basis of their contents, without aiming at a theoretically systematic categorization. however, these non-expansive actions are not inimical or opposite to expansive learning, but unnecessary elements of the epistemic process of ascending from the abstract to the concrete. 2.2 object-oriented interaction the learning actions of the expansive cycle do not dictate what kinds of social interaction are involved in the learning process. to capture this aspect, we used the framework of three types of objectoriented interaction, namely coordination, cooperation, and communication. these three types of interaction rantavuori  et  al       | f l r     6   can be understood as qualitatively different types of epistemological subject–object–subject relations (raiethel, 1983; fichtner, 1984; engeström, 2008). one basic idea to define collaboration is to make a distinction between cooperation and collaboration. according to dillenbourg, baker, blaye, and o'malley (1996), cooperation is accomplished by the division of labor among the participants; each person is responsible for a portion of the problemsolving task. by contrast, collaboration is “a coordinated, synchronous activity that is the result of a continued attempt to construct and maintain a shared conception of a problem” (roschelle & teasley, 1995, p. 70). in this article cooperation and collaboration are used as specific concepts which are part of the analytical framework of three qualitative types of interaction. therefore, our intention is not to participate in the larger ongoing discussion concerning the use concepts of cooperation and collaboration in educational research. coordination is the “default” mode of interaction in groups, experienced as business-as-usual. in coordination, each participant focuses on and performs his or her own scripted role and tasks. the script, coded in written rules, plans, and agendas or engraved in tacitly assumed traditions, coordinates the participants’ actions as if from behind their backs, without being questioned or discussed. each participant has his or her own partial object or task; the possible shared object is not articulated and participants engage in dialogue mainly to maintain and adjust boundaries between their respective tasks and roles. cooperation is typically initiated when the participants face a discoordination, that is, a disturbance or problem that cannot be fixed simply by returning to the prescribed script. in cooperative interactions, participants focus on a shared problem, trying to find mutually acceptable ways to understand, conceptualize, and work on it. in this mode, the given script is temporarily suspended and actions are driven by the demands of the shared object. participants address each other dialogically and there is typically a marked increase in the intensity of the discourse, often manifested in overlapping talk and similar indications of increased engagement. cooperation may remain a mere attempt, typically when a participant initiates it but receives no or only minimal responses from the interlocutors. such an attempt often stands out as a disturbance in that it deviates from the standard script of the interaction. interaction may also take the shape of pseudo-cooperation. in this case, participants interact in a way that resembles cooperation; they address and respond to one another, often talking about something that is perceived as problematic. however, pseudo-cooperation focuses on a substitute object, often an “eternal issue” that can be discussed ad infinitum without ever approaching a resolution. pseudo-cooperation commonly resembles collective venting, sometimes also grumbling or complaining. communication is usually initiated when the participants experience recurring conflicts or breakdowns in their coordination and cooperation. in communication, the participants question and examine their own patterns of interaction in relation to their shared object. as a result, both the object and the script are reconceptualized. this type of self-reflective and transformative phases in interaction are rare and difficult to sustain without the mobilization of novel resources, such as shared documentation, plans, or outside help. overall, the framework of expansive learning calls attention to transitions between types of interaction. as the transitions are typically triggered by discoordinations, conflicts, ruptures and breakdowns, the analysis of types of interaction needs to pay special attention to these kinds of disturbances. often when coordination is interrupted or breaks down, it turns into a cooperation attempt or communication attempt which may or may not lead to a phase of full-fledged cooperation or communication. fluid, pulsating movement from coordination to cooperation and communication and back should be a hallmark of expansive learning characterized by a longitudinal effort to redefine the object of the collective activity. rantavuori  et  al       | f l r     7   2.3 object formation expansive learning is a process of identifying, articulating, reconceptualizing and expanding the object of the activity. in her activity-theoretical study of an elementary school teacher team planning and implementing an innovative curriculum unit, kärkkäinen (1999) identified three phases in the formation of the object of planning. shifts from one phase to the next one were described as turning points, characterized by clusters of disturbances and questioning. a simplified ideal-typical sequence of the formation of the object in expansive learning may be depicted with the help of figure 1. figure 1. ideal-typical phases of the formation of the object in expansive learning. in the first phase depicted in figure 1, the object of the activity may be in crisis due to fragmentation and routinization that prevent the practitioners from facing and embracing new challenges and opportunities in their activity. alternatively, the object may be in such an embryonic state of emergence that it is only vaguely and diffusely grasped and understood by the participants. in the second phase of figure 1, the participants articulate, conceptualize and model a new object for their activity. this new object is typically still a relatively abstract initial idea or principle, a “germ cell”, the expansive implications and potentials of which are not yet realized. in the third phase, the new object is expanded and made concrete, in other words, its manifold practical consequences, extensions, and applications are integrated into a complex totality. 3. research questions to analyze and understand the pre-service teachers’ collaborative learning process, we pose the questions enumerated in table 1. our research questions are driven by our methodological interest in examining the analytical potential of the framework of expansive learning with data from a learning context which was not deliberately designed to follow the guidelines of expansive learning. thus, the methodological questions in table 1 are of primary importance. the substantive questions may be read as tools with which the methodological questions are approached and made concrete. rantavuori  et  al       | f l r     8   table 1 research questions methodological research questions auxiliary substantive questions 1. how does the conceptual framework of expansive learning actions work in the analysis of data from a single session of collaborative learning not deliberately designed to follow the guidelines of expansive learning? 1. which expansive learning actions can be identified in the learning process of the pre-service teacher group? 2. how does the conceptual framework of the object formation work in the analysis of data from a single session of collaborative learning not deliberately designed to follow the guidelines of expansive learning? 2. how was the shared object formed in the learning process of the pre-service teacher group? 3. how does the conceptual framework of types of object-oriented interaction work in the analysis of data from a single session of collaborative learning not deliberately designed to follow the guidelines of expansive learning? 3. how were the types of interaction and transitions between them manifested during the collaborative learning process? 4. how does the integration of conceptual frameworks of expansive learning actions and types of interaction work in the analysis of data from a single session of collaborative learning not deliberately designed to follow the guidelines of expansive learning? 4. what was the relationship between expansive learning actions and types of interaction? 4. participants and context of the study the participants in the study were six pre-service teachers. they were enrolled in a class teacher education program (primary school level) with annual intake of ten students, with educational psychology as their major. the nearest equivalent to a term class teacher outside of finland is a primary school teacher (uk) or an elementary school teacher (usa). at the time of the data collection, the students were in their fourth year. in class teacher education at university of helsinki, students complete a master of arts (education), the completion of which takes approximately five years. the class teacher education at the university of helsinki consists of two different study programs. the major subject may be either education or educational psychology. the core contents of the major subject studies in educational psychology include working as a member of a group and interaction skills; learning, growth, and development; curriculum work and learning to deal with the reality of school life; as well as learning to conduct research. the students in this program study intensively as a small group approximately for three years, applying self-regulated, collaborative learning as one of their main approaches (see eteläpelto, littleton, lahti, & wirtanen, 2005; lipponen & kumpulainen, 2011). the pre-service teachers who participated in this study were thus already socialized into working and interacting within a pedagogical culture that built on collective discussion and collaboration on open-ended and largely self-designed tasks. their activity was that of a new type of university study characterized by self-directed collaborative planning and implementation. however, this new activity existed side by side with the traditional type of university study, characterized by individual rantavuori  et  al       | f l r     9   work on assignments given from above. a tension between these two scripts is an inherent feature of the activity analyzed here. in this article, we analyze a meeting of the pre-service teachers’ group at the beginning of a threemonth course. this course was part of the large study module called “multidisciplinary studies of school subjects taught in the comprehensive school”. during this study module students studied all 13 subjects which are taught in the primary school (grades 1–6). usually each subject is taught in its own separate course by the subject expert (teacher educator). in the teacher education program analyzed in this paper the entire study module was arranged in multidisciplinary way. in the beginning of the module the student group chose three multidisciplinary themes, that were “sustainable development”, “human being” and “time”. the selection of themes was a process were student group together created a joint conception of the important phenomena of the world. therefore this study module was also called “the deepening and widening of the world view”. under each theme one integrating course was created which consisted several school subjects and subject experts. the idea was that the students and subject expert would work together in collaborative way under the common integrating theme. the group was responsible for the planning and implementation of the contents and working procedures of each course. the first two courses of the study module (“sustainable development” and “human being”) were conducted during the second and third year. the last course (“time”) was conducted in the fourth year. the data of this study was collected from this last course. during the course the group investigated the concept of time from multiple disciplinary perspectives, integrating the subject disciplines of mother tongue, handicrafts, history, and multiculturalism into their design. as a final product of their course the members of the group agreed to produce a short theater performance. based on this initial plan they discussed the substantive idea of the theater play. they also discussed what kinds of expertise were needed in the course and invited appropriate experts (teacher educators) to join in the course. four teacher educators representing the subject disciplines listed above participated in the course in the role of experts and supervisors. the pre-service teachers and teacher educators all met as a group six times during the course. during the meetings general guidelines for the course were created, the students’ plans and ideas were discussed, and the final product was evaluated. during the course the pre-service teachers also met at least once a week without the subject experts to discuss their progress on the task and to prepare for the next meeting with subject experts. additionally, the students met some of the subject experts privately a few times during the course. 5. data collection and analysis our data corpus consists of six video-recorded meetings in which only the pre-service teachers were present, comprising a total of 12 hours of video. from this corpus, we selected the first officially scheduled two-hour meeting for detailed transcription and analysis. the selection was based on preliminary viewing of all the videos that resulted in content logs (jordan & henderson, 1995). we decided to focus on phases in which the pre-service teachers conducted planning and talked about planning. earlier studies of expansive learning (e.g., engeström, 2008, pp. 118–168) have demonstrated that features of expansive learning may be found when participants face an open-ended problem solving task, such as a need to plan something that is new for them. since an analysis combining the framework of expansive learning actions and the framework of types of interaction was new and needed to be carefully tested as a methodological solution, we decided to concentrate on a single meeting. focusing on a single meeting runs the risk that no meaningful mini-cycle of expansion is accomplished in such a limited time. our preliminary viewing of the video data convinced us that this meeting was rich in learning actions and types of interaction and would be worth a detailed analysis in spite of the risk. the procedure of our data analysis consisted of four steps, schematically depicted in figure 2. this figure 2 is a summary of the steps of our analysis, not a representation of the conceptual rantavuori  et  al       | f l r     10   structure of expansive learning. the four steps depicted in figure 2 stem from our specific research questions. they are not meant to represent a general procedure to be applied in all analyses of expansive learning. figure 2. steps in the analysis of the data. as a first step, we identified expansive and non-expansive learning actions in the meeting by (a) discerning the topical episodes based on their substantive contents, (b) analyzing the turns of talk within each topical episode in terms of actions and formulating a preliminary description of the actions, and (c) specifying the epistemic function of each action in the stream of learning actions. a learning action typically consisted of an interactive effort that contained more than one turn of talk but was usually shorter than a topical episode. learning actions which did not correspond to the characteristics of any of the expansive learning actions and did not contain an attempt at questioning or explicating the shared object were categorized as non-expansive. as a second step, we examined the succession of the learning actions in relation to the phases of the formation of the object. in other words, we checked which object the learning actions were directed at and what possible phases and turning points emerged in the formation of the object. as a third step, we identified types of interaction in the data. an interaction type for each topical episode was tentatively named by examining the nature of exchanges in the episode and by identifying possible shared or individual objects of the participants. next disturbances, that is, unintentional deviations from the script, were identified. finally, points of transition from one type of interaction to another were examined in greater detail. as a fourth step, to investigate the relationship between expansive learning actions and types of interaction, we brought the two analyses together. next we show briefly with help of transcript excerpts how the three analysis methods mentioned above were applied on the data. the students had agreed earlier that the main task for the course would a preparation of a short theater performance. thus the students needed to write together a script for the theater play. in the next excerpt (table 2) the students are discussing whether some common frames or guidelines are needed for the writing of the script. rantavuori  et  al       | f l r     11   table 2 an example the analyses of learning actions, types interaction and object formation turns transcription learning action type of interaction/ disturbance object formation 145 mark: shall we frame this in some way, i mean, if we go backwards in time [in the story], sort of... ae dist/coopa to 146 ann: well, somebody can go ten years forward [in his/her story], if tina goes 50 years forward [in her story]. ae dist/coopa to 147 john: i’m also getting curious whether we have some common guidelines or does everybody just choose “i will do this” or “i will do that.” is our plan again that i choose it [the story] to take place in ten years’ time, and you choose it [your story] to take place after 20 years. i don’t know if it makes any sense. ae dist/coopa to 148 mark: (inaudible) ae dist/coopa to 149 john: i tried to suggest this system with the panelists [teacher educators]. or are we going to go through any of those reference points with the panelists. that is a principled decision... ae dist/coopa to 150 ann: how about if one just begins working [his or hers own story] even if we others don’t know what the time or the place [where the story is situated]. if one could begin to create a personality or a role for the main character. then one does not necessarily need the time. it could the way to go forward with producing the fictional text. ae coord to legend: ae = analyzing: actual-empirical analysis; dist = disturbance; coopa = cooperation attempt; to = transitional object in this excerpt we identified one expansive learning action, namely actual-empirical analysis. during the learning action of analysis the painstaking process of problem finding and problem definition took place. mark (turn 145) highlighted the problem that common frames are needed for the joint writing process. john (turn 147) emphasized that it might be problematic if everybody could choose freely the topic for their writing. in the analysis of interaction this excerpt was seen as a disturbance. the group’s meeting started with coordination-type of interaction; each participant was concentrating on presenting their own idea and perspective. this coordination was disturbed as two participants, mark and john, made a cooperation attempt by challenging the group’s initial plan which they saw as too vague and non-specific. the cooperation attempt of mark and john did not get response from other participants and the interaction returned to the coordination mode (turn 150). in the analysis of object formation we concluded that in this excerpt the initial diffuse object, named “time”, was already transformed into the transitional object named “theater play”. in the next excerpt (table 3) the student group was talking about problems of collaboration and found a possible explanation from the group’s shared history. in our analysis of expansive learning this was identified to be a learning action of reflecting on the process. the student group was evaluating its own activity in a reflective way. in the analysis of interaction this sequence was identified to represent communication. john (turn 456) was tracing the problems in collaboration to the beginning of the group's life cycle, a phase in which the principle of individual freedom of choice dominated. john recognized that this habit of freedom of choice had now become a problem when the group needed to plan a collective project. rantavuori  et  al       | f l r     12   the initial way of working now became an obstacle to interaction and collaboration. in the analysis of object formation, the transitional object “theater play” was identified also in this excerpt. table 3 an example the analyses of learning actions, types interaction and object formation turns transcription learning action type of interaction/ disturbance object formation 456 john: just recently we were so excited and explaining to the lions [another student group] how we had such great freedom in the beginning. but however, that freedom is causing us problems now. although there is lots of freedom in this unit, it is unlikely that anyone in this unit will have as much freedom as we had. although, this [freedom] is a positive thing in many ways, one negative aspect [of it] is probably that we have sort of become the conquerors of the world, who can do whatever they feel like – and “that’s how i’m going to do it” r com to 457 tina: and whenever i’m up to it… r com to 458 john: and when i’m up to it. and if i’m not up to it, nobody can tell me that “you have to do it” r com to legend: r = reflecting on the process; com = communication; to = transitional object 6. expansive learning actions in our data, we could identify all the learning actions of the expansive cycle except consolidating the new practice. the absence of consolidation is an obvious consequence of focusing on a single meeting: the modeling of a new solution had just begun and the initial idea had not matured enough yet to be consolidated and generalized into a new and stable practice. the results of the analysis of learning actions are summarized in table 4. the meeting started with an episode that did not correspond to the characteristics of any of the expansive learning actions. in this episode, the pre-service teachers discussed practical preparations for the next meeting with subject experts without an attempt at questioning or explicating the object. we gave this non-expansive action the tentative name maintaining the existing practice to describe its character without making any particular theoretical assumptions. the notion of existing practice refers here to routine practices of planning and preparation within teacher education. rantavuori  et  al       | f l r     13   table 4 types and frequencies of expansive and non-expansive learning actions in the pre-service teachers’ meeting type of learning action number of learning actions number of turns of talk maintaining the existing practicea 1 65 questioning 1 8 analyzing: actual-empirical analysis 8 243 analyzing: historical analysis 1 2 modeling a new solution 2 6 examining the new model 5 45 implementing the new model 8 127 reflecting on the process 5 49 different topicb – 258 total 31 802 a non-expansive learning actions are indicated by italics. b conversation not related to the group’s assignment (planning of the course). as table 4 shows, the most common expansive learning actions in the meeting were analyzing, specifically actual-empirical analysis, and implementing the new model; both occurred 8 times. the large number of actions and speaking turns related to actual-empirical analysis indicates that problem finding and problem definition played a central role in the meeting – an emphasis to be expected at the beginning of the expansive learning process. interestingly enough, implementing the new model, reflecting on the process, and examining the new model formed the other dominant block of expansive learning actions. this indicates that instead of only focusing on the early learning actions of the expansive cycle, the group went indeed through an entire mini-cycle of expansive learning in the meeting. on the other hand, the low frequencies of questioning and modeling the new solution indicate that perhaps the shared object constructed in this first meeting was still only very preliminary and would invoke further questioning and re-modeling as the process went on. frequencies of expansive learning actions tell only a part of the story. the more important issue is the way in which the learning actions flow forward and form a meaningful order within a session. by meaningful order we refer to the general directionality of the theoretically formulated expansive cycle (see engeström et al., 2013). in table 5 we give a condensed overview of the progression of expansive learning actions and their contents in the meeting. rantavuori  et  al       | f l r     14   table 5 succession of expansive and non-expansive learning actions and their contents in the pre-service teacher group’s meeting turns of talk contents learning action 1–65 practicalities concerning the next meeting with subject experts are discussed mepa 66–73 tina: “have we completely forgotten the starting point?” q 74–80 planning of the theater play begins ae 81–98 division of instructional resources for the course ae 99–154 agreement on the joint writing task ae 155–173 setting the story in the future ae 174–175 the contents of the previous course considered as starting point ha 176–225 setting the story in the future (continued) ae 226–281 disagreement whether story should be situated in future or in history ae 282–312 negotiation on the starting point of the story ends up in deadlock ae 313–315 ann suggests that the theme “making a choice” should be in everyone’s story; she gets no response m 316–321 creation of a unified story seems impossible ae 322–324 ann demands again a response to her suggestion; this time other participants are responding m 325–343 ann’s idea is accepted and discussion begins on how to include “making a choice” in each participant’s story e 344–354 participants discuss the group’s way of working and state that collaboration is possible but it takes time r 355–358 need for the virtual learning environment (fle) to make things work is acknowledged i 359–362 sheila states that it is problematic if everyone can still write what one wants without any common frame r 363–372 realization that experts of different fields have different perspectives on important moments in history e 373–378 sheila emphasizes the need for a common starting point for the writing; the themes/topics are too general to guide the writing process r 379–384 realization that important turning points in history should be discussed with teacher educators e 385–402 decision that the shared plan should be moved into the virtual platform i 403–416 realization that preparing a theater play forces the participants to collaborate r 417–429 realization that jointly prepared questions for the expert interviews are needed i 430–434 decision to inform teacher educators about today’s decisions i 435–437 decision: we have to start using the fle [virtual learning environment] i 438–447 realization: what we teach today in school should be also relevant for the pupils in future e 448–461 john: we had such great freedom at the beginning and that freedom is causing us problems now r rantavuori  et  al       | f l r     15   461–467 decision: tina should send her text to everybody i 468–470 realization: we have to decide whom to interview i [471–713] [talk about subject matters unrelated to the planning of the next meeting and the course] [dt] 714–787 organizing the expert interviews and sending an email to subject experts i [788–802] [talk about practicalities unrelated to the planning of the next meeting and the course] [dt] legend: mep = maintaining existing practice; q = questioning; ae = analyzing: actual-empirical analysis; ha = analyzing: historical analysis; m = modeling a new solution; e = examining the new model; i = implementing the model; r = reflecting on the process; dt = different topic a non-expansive learning actions are indicated by italics. table 5 shows that the learning actions of the expansive cycle were taken by and large in the order predicted in the theory. to be sure, there were iterations, such as the sequence analyzing –> modeling –> analyzing –> modeling in turns 282–324. also, reflecting on the process was interspersed among actions of examining and implementing the new model in the latter part of the meeting. such iterations are not incompatible with the general model of the expansive cycle, but they represent an interesting challenge for further research. it seems that the expansive mini-cycle was in this case composed of two main parts. we might call these (1) working on the problem (turns 66 to 324) and (2) working on a new model (turns 325 to 470 and turns 714 to 787). during the first part, problem finding and problem definition and formulation of a tentative solution dominated the discussion. this included the learning actions of questioning, actualempirical and historical analysis, and modeling a new solution. during the second part, the solution idea was refined into practical applications and procedures. this included the learning actions of examining and implementing the new model and reflecting on the process. the learning action modeling a new solution formed a turning point and bridging phase between the two main parts. overall, the succession of learning actions in table 5 looks almost like a perfect expansive minicycle. however, closer scrutiny reveals that the cycle is not at all perfect. for this scrutiny, we need to trace the steps of the formation of the object. 7. phases of object formation the initial object of the work of the group was “time”. this was in general terms agreed upon in the group already in the spring. in the fall, before starting the officially scheduled meetings for planning and implementing the course, the pre-service teachers had an informal meeting in a café in which they came up with an idea of producing a small theater play as an outcome of the course. in their first officially scheduled meeting, the first non-expansive learning action, maintaining the existing practice (turns 63 to 65, excerpt 1), represents routine-like planning. it consisted of discussion about how to proceed, with no reference to the shared object. the pre-service teachers articulated their object first in terms of the theater play (turns 66 to 69). excerpt 1 63 sheila: have we planned at all the agenda for the next meetings? how about if everybody would prepare something for a certain meeting. how many are we... five? 64 john: six... tom [member of the group who is absent]. 65 sheila: tom, so we are six all together. how about if one or two people take charge of one meeting. or if it’s well structured, i don’t mind if everyone would prepare for a certain meeting a presentation. the we would use rantavuori  et  al       | f l r     16   three meeting for this and then we will have two presentations for each meeting. it [a meeting] is three hours, so it means one and half hours for each person. 66 tina: have we completely forgotten the starting point, or forgotten the idea that came up last time? well, you [addressing sheila] did not hear all of it. were you taking care of some other business at the time? 67 sheila: could you explain briefly your understanding of it? 68 tina: we were developing that idea of the theater play. 69 sheila: hm. in turn 66, tina challenged the group’s routine-like mode of working and reminded the participants of a shared starting point discussed in a preceding informal preliminary meeting: “have we completely forgotten the starting point, or forgotten the idea that came up last time?” this is the first articulation of the group’s emerging object: the theater play. however, the emerging object remained quite vague, as if a formal shell to be filled with contents. it was not yet a substantive principle or a “germ cell”. in this sense, we may characterize it as a transitory object. the second turning point in the formation of the object took place much later, starting from turn 313 (excerpt 2). the pre-service teachers had discussed the theater play idea for a lengthy period, circling around the idea that each participant would produce his or her own story and pondering on the difficulty of providing coherence and continuity to a text produced this way. ann then initiated actions of modeling in which the participants articulated the second version of their emerging consciously shared object. in this phase, the new object took the shape of the principle of “making a choice” – potentially a substantive germ cell for a new model. excerpt 2 313 ann: i might have a theme to suggest. 314 sheila: go ahead. 315 ann: what if there would be a shared theme of “making a choice” in all of these [individual stories]? that could be done in different ways. the consequences of the choice can be seen later in how the story develops. even though this can be difficult to execute. still, even if characters and situations [in the individual stories] were different, the “making of a choice” would be a connecting link [between the individual stories]. […] 322 ann: now i would like to hear comments about my recent idea. instead of just everybody being silent, i would like to hear some responses like: “i’m not sure...”, or “yes, sounds good...”, or “i would like to...”. 323 sheila: would you explain it briefly one more time? 324 ann: what we should decide now is the connecting element [between the stories]; if everyone starts to write on their own, the connecting element could be making a choice. in every story the theme would be making a choice. this would be visible always, as we move further in time… 325 tina:… it is choices that have impact… 326 john: …they are the ones that have impact. 327 mark: …how would we establish continuity between persons, or is it just any act of making a choice? 328 john: that’s just what we should create together. 329 ann: the continuity is in the fact that in what comes we will see the consequences of the previous choice. 330 john: and those of the previous, previous choices. 331 tina: like for example my choices. excerpt 2 is important in that the vague and diffuse initial object – the notion of time – and the formal transitional object of a theater play were now turned into a much more focused idea, that of making a choice. the notion of choice was connected to the original notion of time by realizing that choices have consequences that are revealed in time: “in what comes we will see the consequences of the previous choice.” from table 5 (section 6) one might infer that the new object, making a choice, was systematically examined and implemented from this point on. however, this was not the case. the phase that followed immediately after the examination of the newly articulated object of making a choice (turns 344 to 354) consisted of reflecting on the process, specifically on the possibility of genuine collaboration – but no reference was made to the idea of making a choice. the next phase (turns 355 to 358) focused on the rantavuori  et  al       | f l r     17   implementation of the plan by means of the virtual learning environment fle – again, with no reference to making a choice. in fact, until the very end of the meeting, the object of making a choice was not anymore mentioned by the participants. the actions of examining and implementing the model actually referred to the transitional object of the theater play, not to the principle of making a choice. the latter was as if forgotten, and the process circled back to the transitional object. in other words, the proposed germ cell was encapsulated, not elaborated on and expanded. the stepwise formation of the object in this meeting may be summarized with the help of figure 3. figure 3. actual steps in the formation of the object in the pre-service teachers’ meeting. the steps depicted in figure 3 testify to the iterative and non-linear character of expansive learning. in our previous study conducted in a library context (engeström et al., 2013), we identified such an iterative and non-linear loop of expansive learning cycle. in the first six sessions the occurrence of learning actions were in line with the general sequence of theoretical model of the expansive learning but in the last two sessions the expansive learning cycle started again from the beginning. in similar way, object formation does not follow the ideal-typical phases as formulated in figure 1 (section 2.3), and sometimes process can collapse and turn backwards. a single meeting is not likely to produce a neat full-fledged expansive cycle: “miniature cycles of innovative learning should be regarded as potentially expansive” (engeström, 2008). the potential is realized – or not realized – in the longer process. 8. types of interaction there are only few studies which have applied the framework of three types of object-oriented interaction. these studies have demonstrated (engeström, 2008, pp. 49–85; saari, 1995; de lange, 2011) that the most common type of interaction is coordination, the second most common is cooperation, and the rarest type is communication. further, these studies also revealed the important role of disturbances in the analysis of types of interaction. we identified all three main types of interaction – coordination, cooperation, and communication – in our data. we also found three phases of pseudo-cooperation. as shown in table 6, the most common type of interaction was coordination, comprising 227 turns of talk. 105 turns represented cooperation, and 47 turns pseudo-cooperation. communication occurred only in 10 turns. this low number of communication turns indicates that reconceptualizing the script and mode of interaction in relation to the shared object of activity was very challenging for the participants. rantavuori  et  al       | f l r     18   table 6 types of interaction in the meeting of the pre-service teachers’ group type of interaction phases turns coordination 5 227 cooperation 6 105 pseudo-cooperation 3 47 communication 1 10 total (types of interaction) 15 389 different topic 2 258 as pointed out above, a transition from one type of interaction to another often passes through a short phase of disturbances. disturbances may lead to disintegration, contraction, or expansion in the process. in our data, we identified a number of conflicts. in addition to those, we also examined cooperation attempts and communication attempts as disturbances. the frequencies of these disturbance types are presented in table 7. table 7 types and frequencies of disturbances in the student group’s meeting disturbance episodes turns conflict 3 10 cooperation attempt 7 40 communication attempt 5 32 total 15 82 table 8 presents the temporal succession of the types of interaction in the meeting. the idea of theater play as a transitional object was invoked in turns 66 to 73. the subsequent turns 74 to 144 represent a return to coordination. the participants brought up different resource issues (time, help from teacher educators, the virtual learning environment) that did not generate a common thread and problem to be jointly tackled. questions about the allocation of time for the preparation of the theater play were raised and ruminated about but not answered: “but how much time do we have to reserve for it [preparing of the theater play], extra days, for the work it out, because it takes...?” (turn 71) “how much time have we reserved? we have booked fridays from nine to three. after the panel meetings there is always time and...” (turn 81). this does not look very efficient; one might argue that it looks more like discoordination than coordination. however, the standard script of planning in meetings is often indeed inefficient, an example being prolonged episodes in which the participants try to agree on the date and time of the next meeting, each one bringing up disconnected concerns and constraints that make the decision-making look rather absurd. this way rantavuori  et  al       | f l r     19   coordination in meetings quite often comes close to its own limits; such episodes could easily collapse into discoordination or erupt into open conflict. as already noted in the discussion of table 5 (section 6), the group’s interaction seems to have consisted of two main parts. we might call the first part (turns 1 to 324) “coordinative interaction” and the second part (turns 325 to 470) “cooperative interaction”. characteristic to the first part was that the transitional object of theater play did not function as a truly shared object. the first part contained also a pseudo-cooperation phase and several cooperation attempts interpreted as disturbances. the second part is more problematic. there was a notable increase in cooperation and communication attempts. but as we know from the preceding section, after the brief phase of cooperation based on the object of making a choice, the remaining phases of cooperation and communication attempts were actually focused on the transitional object of theater play. in this light, the second part of the meeting was not simply continuation of the first part but rather circling back to the earlier object. table 8 types of interaction and disturbances in the student group’s meeting turns contents type of interaction / disturbance 1–65 practicalities concerning the next meeting with subject experts are discussed coordination 66–73 a disagreement between the participants of the common starting point cooperation attempta 74–144 a discussion of how to proceed with a joint preparation of a theater play coordination 145–149 a criticism that the guidelines for the joint writing task are missing cooperation attempt 150–154 a suggestion that the same protagonist in every story could be a link between different stories coordination 155–164 taking tina’s story as a common starting point pseudo-cooperation 165 is it possible to have something else than just science fiction in the story cooperation attempt 166–178 a development of the idea of the story that takes place in future pseudo-cooperation 179–189 a disagreement of how much one should put emphasis in future in his or her story cooperation attempt 190–213 a development of the story situated in future continues pseudo-cooperation 214–222 should we have a same central character in every story? cooperation attempt 223–225 john does not want to situate his story in the future conflict 226–311 unsuccessful attempts trying to find connecting theme for the shared story coordination 312 seems impossible to write a shared story conflict 313–315 ann is suggesting that in everyone’s story should be a one unified theme, which is “making a choice” but did not get response cooperation attempt rantavuori  et  al       | f l r     20   316–321 seems that participants only want to work individually without binding structure conflict 322–324 ann demands again a response for her suggestion more determined way and this time other participants are responding cooperation attempt 325–343 ann’s idea is accepted and discussion begins how to connect making a choice in each participant’s story cooperation 344–351 participants discuss the group’s way of working and state that collaboration is possible but it takes time communication attempt 352–358 john says that one needs to follow others work too if he/she wants that his/her story works cooperation 359–362 integrating theme is missing communication attempt 363–372 chosen perspectives for the story can sometimes be too narrow cooperation 373–378 sheila: our themes are topics are too general to guide our writing process communication attempt 379–406 different ideas concerning joint story writing are considered cooperation 407–416 previously we did not have a common goal which forces us now to collaborate communication attempt 417–447 preparing interviews; informing teachers; getting a virtual learning platform; visioning pupils’ needs in future cooperation 448–451 earlier we use to have only individual goals? communication attempt 452–461 john sees that in the early-stage of group’s work it was given so much freedom that collaboration is now difficult communication 461–470 tina’s story is chosen as one starting point and the interviews of the experts are organized cooperation [471–713] [talk about subject matters unrelated to the planning of the course] [different topic] 714–787 organizing the expert interviews and sending an email to subject experts coordination [788–802] [talk about practicalities unrelated to the planning of the course] [different topic] a disturbances are indicated by italics. 9. dynamics between learning and interaction as a result of the analysis of expansive learning actions, the meeting was tentatively divided into two main parts, “working on a problem,” and “working on a model”. in a similar way, in the analysis of types of interaction, the meeting was divided in two main parts, “coordinative interaction” and “cooperative interaction”. the transition from the first part to the second part took place at the same point in both analyses. rantavuori  et  al       | f l r     21   as the two analyses were merged in table 9 it was possible to divide the meeting into three parts. the first part of the meeting may be called “coordinated working on a problem”, the second part “transition from coordinated working to cooperative working”, and the third part “cooperative working on a model”. table 9 merged analyses of learning actions and types of interaction analysis of expansive learning analysis of interaction turns learning action turns type of interaction / disturbance 01–65 maintaining the existing practicea 01–65 coordination first part: “coordinated working on a problem” (turns 66–324) 66–73 questioning 66–73 cooperation attemptb 74–80 actual-empirical analysis 74–144 coordination 81–98 actual-empirical analysis | | 99–154 actual-empirical analysis | | | | 145–149 cooperation attempt | | 150–154 coordination 155–173 actual-empirical analysis 155–164 pseudo-cooperation | | 165 cooperation attempt 174–175 historical analysis 166–178 pseudo-cooperation 176–225 actual-empirical analysis 179–189 cooperation attempt | | 190–213 pseudo-cooperation | | 214–222 cooperation attempt | | 223–225 conflict 226–281 actual-empirical analysis 226–311 coordination 282–312 actual-empirical analysis 312 conflict second part: “transition from coordinated working to cooperative working” (turns 313–324) 313–315 modeling a new solution 313–315 cooperation attempt 316–321 actual-empirical analysis 316–321 conflict 322–324 modeling a new solution 322–324 cooperation attempt third part: “cooperative working on a model” (turns 325–470) 325–343 examining the new model 325–343 cooperation 344–351 reflecting on the process 344–351 communication attempt 352–354 examining the new model 352–358 cooperation 355–358 implementing the model | | 359–362 reflecting on the process 359–362 communication attempt 363–372 examining the new model 363–372 cooperation rantavuori  et  al       | f l r     22   373–378 reflecting on the process 373–378 communication attempt 379–384 examining the new model 379–406 cooperation 385–402 implementing the model | | 403–416 reflecting on the process 407–416 communication attempt 417–429 implementing the model 417–447 cooperation 430–434 implementing the model | | 435–437 implementing the model | | 438–447 examining the new model | | 448–461 reflecting on the process 448–451 communication attempt | | 452–461 communication 461–467 implementing the model 461–470 cooperation 468–470 implementing the model | | [471–713] [different topic] [471–713] [different topic] 714–787 implementing the model 714–787 coordination [788–802] [different topic] [788–802] [different topic] a non-expansive learning actions are indicated by italics. b disturbances are indicated by italics. table 9 would seem to indicate that as an expansive learning process moves epistemically from questioning to analysis, modeling, implementation and reflection on the process, it also moves interactionally from coordination to cooperation and at least attempted communication. on the other hand, there is no deterministic or mechanical correspondence between specific learning actions and specific types of interaction. epistemic actions that serve an expansive function from the point of view of the entire cycle may be performed in a coordinated manner that makes them look rather unproductive within their own limited confines. and superficially productive forms of interaction may in a closer analysis turn out to be phases of pseudo-cooperation that serve to avoid the core issues rather than tackle and solve them. in table 9, there is a long phase (74–312) containing several learning actions of actual-empirical analysis and one learning action of historical analysis. this phase consists of different types of interaction (coordination and pseudo-cooperation) and several disturbances (cooperation attempts and conflicts) which shows that mechanical correspondence between specific learning actions and specific types of interaction does not exist. in section 8, part of this phase (turns 74–144) is analyzed more detailed and it is possible to see how joint planning looks inefficient and unproductive as the participants brought up different resource issues that did not generate a common thread and problem to be jointly tackled. however, in our data the learning actions of questioning were typically interpreted as cooperation attempts. the learning actions of actual-empirical analysis were interactionally more heterogeneous, containing coordination and pseudo-cooperation types of interaction as well as several cooperation attempts and conflicts. the learning actions of modeling were typically interpreted as cooperation attempts, whereas the learning actions of examining and implementing the new model were typically interpreted as cooperation. the learning actions of reflecting on the process were typically categorized as communication attempts or, in one case, as communication. probably the most important lesson from the integrated analysis is the importance of transitions and disturbances. these may be small in terms of time and number of speaking turns, but they are crucial for the understanding of the dynamics of the learning process. the fact that the short sequence of turns 313 to 324 was identified as a turning point in both analyses testifies to this. rantavuori  et  al       | f l r     23   10. conclusions in this paper, we have examined the theory of expansive learning (engeström, 2015) as a conceptual and methodological framework for understanding open-ended and problem-based collaborative learning among finnish pre-service teachers. our study was focused especially on two aspects of expansive learning, namely types and sequences of expansive learning actions (e.g., engeström & sannino, 2010) and types and sequences of object-oriented interaction (engeström, 2008; fichtner, 1984; raiethel, 1983). the task of the paper is primarily methodological. we believe that cultural-historical activity theory needs to be turned into methods and procedures of systematic empirical analysis. therefore, the aim of the paper is to develop methodology for analyzing dynamics of expansive learning. a new methodological framework created in this study is tested in the analysis of planning meeting of a pre-service teacher group. our research questions are driven by our methodological interest in examining the analytical potential of the framework of expansive learning with data from a learning context which was not deliberately designed to follow the guidelines of expansive learning. for this reason, our methodological research questions are accompanied by auxiliary substantive questions. the substantive questions may be read as vehicles with which the methodological questions are approached and made concrete. the first methodological question of our study was: how does the conceptual framework of expansive learning actions work in the analysis of data from a single session of collaborative learning not deliberately designed to follow the guidelines of expansive learning? our analysis showed that it is possible to apply the framework of expansive learning actions on a learning process confined to a single meeting in which expansive learning was not deliberately induced. this indicates that expansive learning can take place as a naturally occurring process also without formative interventions such as the change laboratory method (engeström & sannino, 2010). expansive learning processes can be analyzed in minute detail at the level of conversational turns and episodes, and an almost complete mini-cycle of expansive learning can be fulfilled during a single meeting. our analysis of expansive learning actions showed that an almost complete expansive learning cycle appeared in pre-service teachers’ meeting. six out of the seven learning actions of the expansive learning cycle appeared in the meeting in a meaningful order. only the last expansive learning action, consolidating the new practice, was missing from the learning process – an outcome to be expected in light of the fact that our analysis was restricted to a mini-cycle accomplished within a single meeting. the low frequencies of the actions of questioning and modeling the new solution indicate that perhaps the shared object constructed in this first meeting was still only preliminary and would invoke further questioning and re-modeling as the process went on. it seems that the mini-cycle of expansive learning consisted of two main parts, namely “working on a problem” and “working on a new model”. the learning action of modeling a new solution functioned as a transition phase and bridge between the two main parts. an interesting methodological finding is that learning actions in this meeting roughly followed the same order as the theory proposed. there were some iterative back-and-forth movements between learning actions of analyzing and modeling and in similar fashion with learning actions of examining, implementing and reflecting (see engeström et al., 2013). this implies that learning actions may often take place in clusters in which the three first learning actions (questioning, analyzing and modeling) tend to occur together and, in a similar fashion, the next three learning actions (examining, implementing and reflecting) tend to occur together in iterative clusters. the two-part structure of the expansive learning cycle observed in this study also supports this finding. most theories of learning take the initial existence of a fairly clear problem, task or assignment as a given. it means that a phase of problem finding and definition of the object are not included in the focus of the analysis. in expansive learning this phase of “working on the problem” is essential. rantavuori  et  al       | f l r     24   on the other hand, many theories of learning also ignore or exclude the actions of implementation. jerome bruner (1974, p. 233) pointed out that if we really want to study the conditions of learning, we need to follow our subjects far longer than is usual in laboratory experiments or test-driven classrooms. we need to see what the learners will do with their new insights, how knowledge is turned into actions. in this respect, the actions of implementation in the second part of the process of expansive learning are of utmost importance. the second methodological question of our study was: how does the conceptual framework of the object formation work in the analysis of data from a single session of collaborative learning not deliberately designed to follow the guidelines of expansive learning? first, tracing the formation of the object is an indispensable methodological step in the analysis of expansive learning; and second, no matter how promising and powerful features of expansion we may find in a single learning session, a full assessment of such a mini-cycle requires an analysis of the entire multi-session learning process. an ideal-typical process of expansive object formation moves from a routinized, fragmented or diffuse initial object to a consciously articulated and shared “germ cell” object, to an expanded concrete object (figure 1, section 2.3). the steps of object formation were different in the meeting we analyzed. the initial diffuse object (“time”) was first transformed into a formal transitional object (“theater play”). subsequently a potential “germ cell” object (the principle of “making a choice”) was formulated by the participants – but it was abandoned and the participants returned to the transitional object (figure 3, section 7). in other words, what looked like a nearly perfect expansive mini-cycle turned out to be a more complex iterative process. the third methodological question of our study was: how does the conceptual framework of types of object-oriented interaction work in the analysis of data from a single session of collaborative learning not deliberately designed to follow the guidelines of expansive learning? the framework of types of objectoriented interaction seems a promising method for opening up the dynamics of collaborative learning. the analysis of interaction shows clearly that the participants were not just focused on completing the task but also on their mutual interaction. the diversity of types of interaction found in the data indicates that different types of interaction are needed especially when students are trying to complete a vague and open-ended task. all the three types of interaction – coordination, cooperation, and communication – occurred in the meeting of the pre-service teacher group. in addition, three phases of pseudo-cooperation were identified. this finding supports our conclusion that this form of learning was indeed rich and potentially powerful. we identified several transitions between types of interaction that were marked by disturbances, either in the form of conflicts or in the form of cooperation and communication attempts that were not picked up and sustained by other members of the group. also the analysis of types of interaction indicated that the meeting was divided in two parts, “coordinative interaction” and “cooperative interaction”. however, as pointed out above, the second part was not simply a continuation of the first part but more like circling back to the earlier transitional object. the fourth methodological question of our study was: how does the integration of conceptual frameworks of expansive learning actions and types of interaction work in the analysis of data from a single session of collaborative learning not deliberately designed to follow the guidelines of expansive learning? a key methodological finding of this study is that while necessary, both epistemic learning actions and types of interaction are in themselves insufficient windows into expansive learning. what is needed as a connecting link is tracing steps in the formation of the object. in our study the analysis of the formation of the object revealed that what looked like a neatly linear expansive mini-cycle was in fact a more complex and iterative movement between different versions of the object. the true expansive potential of a mini-cycle can only be discovered by extending the time scale and scope of the analysis. our analysis indicates that as an expansive learning process moves epistemically from questioning to analysis, modeling, implementation and reflection on the process, it also moves interactionally from coordination to cooperation and at least attempted communication. on the other hand, there is no rantavuori  et  al       | f l r     25   deterministic or mechanical correspondence between specific learning actions and specific types of interaction. the merging of the analysis of learning actions and the analysis of types of interaction led us to identify three parts in the meeting, namely “coordinated working on a problem”, “transition from coordinated working to cooperative working”, and “cooperative working on a model”. our integrated analysis highlights the importance of transitions and disturbances in expansive learning. these are often small in terms of time and number of speaking turns but crucial for the dynamics of the learning process. these kinds of findings concerning the complex character of expansive learning process have not been reported in previous studies and can thus serve as methodological supports for further research. in the theory of expansive learning, contradictions are seen as the driving force of transformation (engeström & sannino, 2010). although our analysis did not specifically focus on contradictions, we can see in the pre-service teachers’ meeting a pervasive tension between two scripts. the explicit script of the meeting was that of self-determined collaboration on a complex open-ended task, in which planning, design and implementation are unified. this script was challenged and interrupted by the traditional script of studying individually to complete given assignments, in which planning and design are reduced to technical and logistic arrangements. this tension seems to be behind the frequent disturbances observed in the meeting, and the group’s difficulties in constructing a new shared object may be understood against this background. the model of expansive learning is not a universal formula of phases or stages. one probably never finds a learning process that strictly follows the ideal-typical model of expansive learning. whenever one examines or facilitates a learning process with the help of the model, one tests, criticizes and hopefully enriches the theoretical ideas of expansive learning. the theory of expansive learning has mainly been applied to large-scale transformations in activity systems, often spanning a period of 2 or 3 years. in this study, however, an expansive learning cycle was applied to analyze a single learning session, lasting only two hours. this raises a critical question: can a mini-cycle of learning be characterized as expansive? our analysis demonstrates that a mini-cycle of innovative learning can be, to some extent, expansive. the emergence of such mini-cycles does not itself guarantee that a larger expansive cycle takes place. small cycles may remain isolated events, and the overall cycle of expansion may become stagnant or regressive or even fall apart. the occurrence of a full-fledged expansive learning cycle is challenging to achieve, and it typically requires concentrated effort of deliberate interventions. with these reservations in mind, the theory of expansive learning can be applied as a framework for analyzing small-scale innovative learning processes as well. moreover, our study contributes to research on learning and interaction in activity systems. most studies on activity systems focus either on learning or on interaction, keeping the two relatively separate. our aim was to produce a fine-grained analysis of the dynamics between expansive learning actions and types of interaction. the qualitative transition in the pre-service teachers’ learning took place at the same point in both learning actions and in types of interaction. this indicates that cycles of expansive learning actions and progressions of types interaction are closely intertwined. whether the dynamics and qualitative transitions are similar in other contexts as well, needs to be explored in future studies. keypoints a new methodological framework was created for analyzing dynamics of expansive learning. a new method is tested in the analysis of a planning meeting of a pre-service teacher group. tracing the formation of the object is an indispensable methodological step in the analysis of expansive learning. the analysis of group’s interaction highlights the importance of transitions and disturbances in expansive learning. rantavuori  et  al       | f l r     26   this study offers a methodological lens for examining innovative forms of learning in various contexts. references bereiter, c., & scardamalia, m. (1993). surpassing ourselves: an inquiry into the nature and implications of expertise. chicago: open court. bruner, j. s. (1974). beyond the information given. london: george allen & unwin ltd. davydov, v. v. (1990). types of generalization in instruction: logical and psychological problems in the structuring of school curricula. reston: national council of teachers of mathematics. dillenbourg, p., baker, m., blaye, a., & o’malley, c. (1996). the evolution of research on collaborative learning. in h. spada, & p. reimann (eds.), learning in humans and machines (pp. 189–211). oxford: elsevier science. engeström, y. (2008). from teams to knots: activity-theoretical studies of collaboration and learning at work. cambridge: cambridge university press. engeström, y. (2015). learning by expanding: an activity-theoretical approach to developmental research. cambridge: cambridge university press. engeström, y., rantavuori, j., & kerosuo, h. (2013). expansive learning in a library: actions, cycles and deviations from instructional intentions. vocations and learning, 6 (1), 81–106. doi:10.1007/s12186012-9089-6 engeström, y., & sannino, a. (2010). studies of expansive learning: foundations, findings and future challenges. educational research review, 5, 1–24. doi:10.1016/j.edurev.2009.12.002 engeström, y., & sannino, a. (2012). whatever happened to process theories of learning? learning, culture and social interaction, 1(1), 45–56. doi:10.1016/j.lcsi.2012.03.002 eteläpelto, a., littleton, k., lahti, j., & wirtanen, s. (2005). students’ accounts of their participation in an intensive long-term learning community. international journal of educational research, 43(3), 183– 207. doi:10.1016/j.ijer.2006.06.011 fichtner, b. (1984). co-ordination, co-operation and communication in the formation of theoretical concepts in instruction. in m. hedegaard, p. hakkarainen, & y. engeström (eds.), learning and teaching on a scientific basis: methodological and epistemological aspects of the activity theory of learning and teaching. aarhus: aarhus universitet, psykologisk institut. foot, k. (2001). cultural-historical activity theory as practical theory: illuminating the development of a conflict monitoring network. communication theory, 11(1), 56–83. doi:10.1111/j.14682885.2001.tb00233.x goldman, s. r. (2014). perspectives on learning: methodologies for exploring learning processes and outcomes. frontline learning research, 2(4), 46–55. doi:10.14786/flr.v2i4.117 greeno, j. g. (2011). a situative perspective on cognition and learning in interaction. in t. koschmann (ed.), theories of learning and studies of instructional practice (vol. 1, pp. 41–71). new york: springer. greeno, j. g., & engeström, y. (2014). learning in activity. in r. k. sawyer (ed.), the cambridge handbook of the learning sciences (2nd ed., pp. 128–147). cambridge: cambridge university press. jordan, b., & henderson, a. (1995). interaction analysis: foundations and practice. the journal of the learning sciences, 4, 39–103. kärkkäinen, m. (1999). teams as breakers of traditional work practices: a longitudinal study of planning and implementing curriculum units in elementary school teacher teams. helsinki: university of helsinki, department of education. de lange, t. (2011). formal and non-formal digital practices: institutionalizing transactional learning spaces in a media classroom. learning, media and technology, 36(3), 251–275. doi:10.1080/17439884.2011.549827 rantavuori  et  al       | f l r     27   lave, j., & wenger, e. (1991). situated learning: legitimate peripheral participation. cambridge: cambridge university press. lemke, j. l. (1990). talking science. norwood, nj: ablex. lipponen, l., & kumpulainen, k. (2011). acting as accountable authors: creating interactional spaces for agency work in teacher education. teaching and teacher education, 27(5), 812–819. doi:10.1016/j.tate.2011.01.001 nilsson, m. (2003). transformation through integration: an activity theoretical analysis of school development as integration of child care institutions and elementary school. karlskrona: blekinge institute of technology. raeithel, a. (1983). tätigkeit, arbeit und praxis. frankfurt am main: campus. roschelle, j., & teasley, s. (1995). the construction of shared knowledge in collaborative problem solving. in c. e. o’malley (ed.), computer-supported collaborative learning. heidelberg: springer-verlag. saari, e. (1995). voidaanko tutkimusryhmiä perustaa? tapaustutkimus valtion teknillisen tutkimuskeskuksen metallilaboratorion ryhmäkokeilusta vuosina 1989–1991. [could the research groups be founded? a case study of metals laboratory’s team experimentation in 1989–1991 at the technical research centre of finland] vtt tiedotteita 1627. espoo: vtt offsetpaino. schmittau, j., & morris, a. (2004). the development of algebra in the elementary mathematics curriculum of v.v. davydov. the mathematics educator, 8(1), 60–87. seppänen, l. (2004). learning challenges in organic vegetable farming: an activity-theoretical study of onfarm practices. helsinki: university of helsinki, institute for rural research and training. sfard, a. (1998). on two metaphors of learning and the dangers of choosing just one. educational researcher, 27(2), 4–13. doi:10.3102/0013189x027002004 virkkunen, j., & newnham, d. s. (2013). the change laboratory: a tool for collaborative development of work and education. rotterdam, the netherlands: sense publishers. wenger, e. (1998). communities of practice: learning, meaning and identity. cambridge: cambridge university press. microsoft word wegner&nückles_publication.docx             frontline  learning  research  vol.3  no.  4  (2015)  95  -­‐  109   issn  2295-­‐3159     training the brain or tending a garden? students’ metaphors of learning predict self-reported learning patterns elisabeth wegner, matthias nückles university of freiburg, germany article received 18 september / revised 30 november / accepted 6 december / available online 20 january abstract conceptions of learning are seen as an important factor in shaping students’ patterns of learning. however, conceptions are often implicit and difficult to assess. metaphors have been proposed as a method to assess conceptions, because metaphors are closely linked to the conceptual system. therefore, in our study we assessed which conceptions of learning are visible in students’ metaphors of learning and examined whether these metaphors predict differences in students’ learning patterns. altogether, n = 91 students of educational science from a german university filled in a questionnaire on their personal metaphors of learning, their learning strategy use, epistemological beliefs, and their motivation. four kinds of metaphors could be differentiated: regulation-related metaphors, learning as knowledge acquisition, learning as problem solving, or as personality development. a discriminant analysis revealed that students with personality development metaphors and with problem solving metaphors were more intrinsically motivated and more aware of the relativism of knowledge than students with regulationrelated or knowledge acquisition metaphors. students with personality development metaphors differed from students with problem solving metaphors in their stronger use of deep processing strategies, their lower extrinsic motivation and their stronger rejection of a dualism of knowledge. the study demonstrates that metaphors of learning are a suitable tool for assessing students’ conceptions of learning and gives new insights on using this innovative method as an assessment tool. keywords: conceptions of learning; metaphors; learning patterns; approaches to learning corresponding author: dr. elisabeth wegner, universität freiburg, institut für erziehungswissenschaft, rempartstr. 11, d-79085 freiburg, germany. phone: +49(0)761 / 203 97550, fax: : +49(0)761 / 203 2458, email: elisabeth.wegner@ezw.uni-freiburg.de doi: http://dx.doi.org/10.14786/flr.v3i4.212 wegner  &  nückles       | f l r     96   “learning is like rowing against the current. as soon as you stop, you drift back again.” benjamin britten (1913-76) “the roots of education are bitter, but the fruit is sweet” (aristotle, 384 -382 b.c.) 1. introduction a great number of proverbs tell us in in metaphors what learning is like, how learning occurs, and what the benefits of learning are. such metaphorical expressions have received a lot of attention from researchers from as diverse domains as philosophy (black, 1993), cognitive science (gick & holyoak, 1980) or cognitive linguistics (lakoff & johnson, 1980), because metaphors have been identified as being more than a deviation from the ‘normal use’ of language. instead, metaphors are closely linked to the way our conceptual system is structured, thus being one of the basic mechanisms in which we perceive the world (lakoff & johnson, 1980). in the context of cognitively oriented research, conceptual metaphors are usually defined as a situation or an object x that shares a similarity with a situation or an object y (“x is like y”). the situation or object x that is characterized by the metaphor is called the “target”, and the situation or object y that is the medium of comparison, the “source” of the metaphor. because conceptual metaphors are based on the detection of similarities of new experiences with familiar experiences, they help to understand novel information, concepts, or information (gentner & holyoak, 1997, p. 32). for example, britten’s metaphor of learning as rowing against the current helps to convey the importance of learning continuously. however, metaphors only partially structure an experience, because the target and source of a metaphor never match completely. obviously, the rowing metaphor leaves out important other aspects of learning, such as that learning produces positive outcomes, as in aristotle’s metaphor of education, or that learning requires the learner to link new information to existing knowledge, which becomes visible in a metaphor such as “learning is like weaving a net”. according to lakoff and johnson’s conceptual metaphor theory, the metaphors that are used also feed back into our conceptual systems. for example, the metaphors “time is a resource” and “work is a resource” bring us to the realisation that leisure time is also a resource, thus influencing our concepts of leisure to be perceived as a valuable good that must not be wasted (lakoff & johnson, 1980). thus metaphors act as a lens through which we perceive the world around us. landau, meier and keefer (2010) suggest that metaphors are so fundamental for human thinking, that in order to understand individuals’ actions with regard to abstract social concepts, such as justice, spirituality, or happiness, it is central to look at how individuals structure these concepts metaphorically: “…metaphor is a cognitive tool that people routinely use to interpret and evaluate information related to those abstract concepts. put simply, a metaphor-enriched perspective suggests that a complete account of the meanings people give to abstract, socially relevant concepts requires an understanding not only of their schematic knowledge about those concepts in isolation but also how they structure those concepts in terms of superficially dissimilar, relatively more concrete concepts.” (p.1047) therefore, we assume that metaphors could be an important tool to assess how students structure their concepts of learning. the aim of the current study was therefore to assess which kind of metaphors students use to describe learning and which impact the metaphors have on students’ learning. so far, there is only very little research assessing students’ metaphors of learning. however, we find ample research on students’ conceptions of learning. therefore, we will first outline findings on conceptions of learning and their role for how students learn. afterwards we will elaborate on how metaphors and conceptions might relate to each other and how metaphors have been used to assess conceptions. finally we will present evidence from our study indicating that indeed the metaphors that students use relate to their self-reported learning activities, their motivation and their epistemological beliefs, that is, their beliefs about knowledge and knowing. wegner  &  nückles       | f l r     97   1.1 conceptions of learning conceptions can be defined as an “individual’s personal and therefore variable response to a concept” (entwistle & peterson, 2004, p. 408). conceptions are usually understood as systems of beliefs (e.g. marton & säljö, 1976; richardson, 2007), which act as a filter for cognition (see pajares, 1992). the kind of conception of learning a student holds organizes the student's perception of learning environments, the interpretation of learning tasks, the expectations towards teaching staff and other students, motivation and also the choice of learning strategies (pajares, 1992). early studies on students’ conceptions of learning (säljö, 1979) differentiated between five different conceptions, ranging from reproductive conceptions such as understanding learning as the acquisition of factual information and as memorizing what has been learned, over learning as the application and use of knowledge, to meaning oriented conceptions such as understanding what has been learned and as seeing things in a different way. according to the phenomenographic perspective, conceptions are understood as qualitatively distinct categories, but “higher” conceptions such as “developing as a person” subsume lower conceptions, such as “acquisition of knowledge”. individuals develop towards more advanced conceptions (e.g. marton & säljö, 1976). other researchers do not assume a developmental order of distinct and developmental categories of conceptions (e,g. richardson, 2007). later research focused on how students’ understanding of learning is related to students’ use of learning strategies, their learning motivation and their epistemological beliefs. in this productive area of research, two merging research frameworks can be discerned (vanthournout, donche, gijbels & van petegem, 2014), namely the learning patterns framework (vermunt, 1996; vermunt & vermetten, 2004) and the approaches to learning framework (e.g. entwistle & peterson, 2004; entwistle & ramsden, 1983). both frameworks are based on the assumption that there are, on the one hand, different dimensions of learning on which students individually vary (such as their use of learning strategies, their learning motivation or their self-regulation strategies), but that on the other hand, these dimensions form systematic clusters, which are called learning patterns or approaches to learning. richardson (2011) assumes that students’ conceptions of learning are important for forming these systematic clusters. interestingly, in both research frameworks, we find a pattern that is characterized by an intrinsic interest in studying and in learning contents (deep approach / meaning-directed learning pattern). students with this pattern use deep processing strategies and have a high level of self-regulation, and according to vermunt (1996), this pattern is characterized by a mental model of learning as construction of knowledge. also, both frameworks describe an opposing pattern in which students have the major intention to cope with course requirements, are externally motivated, see contents as unrelated bits of knowledge and fail to see the meaning or value of the contents. this goes in hand with learning strategies that focus on rehearsal and involve little reflection, and also with a feeling of pressure and anxiety (surface approach / reproductiondirected learning pattern). this pattern is based on the mental model of learning as intake of knowledge (vermunt, 1996). both frameworks also describe, apart from these two more or less identical types of students, additional patterns or approaches. within the learning patterns framework, vermunt (1996) describes a type of students with an undirected learning pattern. this pattern is characterized by a lack of regulation, ambivalent motivation, and no identifiable mental model of learning. the other type of student described by vermunt (1996) are those with an application directed learning pattern, which is based on the mental model of learning as the use of knowledge, an intrinsic (vocational) orientation and concrete processing of information. the approaches to learning framework additionally includes a strategic approach (biggs, 1987; entwistle, tait, & mccune, 2000). this approach is characterized by a strong motivation to do well in the course and to complete the degree in order to accomplish personal goals. students with a strategic approach organize their studying well, manage their time effectively, and are alert to assessment requirements and criteria (virtanen & lindblom-ylänne, 2010). in a comprehensive review, vermunt and vermetten (2004) found that an undirected pattern/surface approach leads to the worst studying results; the best studying results are yielded by the meaning-directed pattern/deep approach. reproduction-directed pattern and application-directed pattern had no clear relation to studying success. wegner  &  nückles       | f l r     98   1.2 assessing conceptions of learning taken together, we can draw from research that conceptions play an important role in shaping students’ learning, and thus have an impact on their studying. however, assessment of conceptions is not as simple as it seems. as we have pointed out above, conceptions are partly implicit and therefore difficult to assess. interviews which could be used to assess also implicit aspects of conceptions are time consuming and are not suitable for large scale studies. often, questionnaires with dimensional assessment scales such as the inventory of learning styles (ils, vermunt, 1994) are used to assign students to distinct groups by using cluster analysis (e.g. parpala, lindblom-ylänne, komulainen, litmanen, & hirsto, 2010; entwistle & mccune, 2013; richardson, 2007). however, the technique of cluster analysis carries the risk of methodological artefacts, because general answer tendencies might account for correlations between two variables (richardson, 2011). for example, some persons tend to agree rather than to disagree on items (acquiescent response style), whereas others tend to choose extreme response categories on all scales. this can result in clusters not based on differences in the assessed dimensions, but on the general answer tendencies. another problem is that clusters can only be determined post hoc in large samples, but it is difficult to make an individual diagnosis of conceptions. consequently, assessment techniques are needed to determine conceptions of learning. given the important role of metaphors for our cognitive system, it comes as no surprise that recently in the area of teacher education and of higher education in general, metaphors have become increasingly popular for assessing implicit constructs such as conceptions (löfström, nevgi, wegner, & karm, 2015). 1.3 using metaphors for understanding conceptions to use metaphors to assess conceptions of learning, we need to take a closer look into how metaphors and conceptions are assumed to relate to each other, and how this has been exploited in research. unfortunately, educational researchers using metaphors often do not explicate which relation between metaphors and conceptions they assume. this is problematic because lakoff’s and johnson’s cognitive metaphor theory, which is still the most prominent metaphor theory, allows for different assumptions about the relation between metaphors and cognition. murphy (1996) describes a “strong version” of this theory, stating “that some concepts are not understood via their own representations but instead by (metaphoric) reference to a different domain” (p. 201). this would imply that a metaphor of learning is identical to a conception of learning. a person who describes learning as the construction of a skyscraper would then literally have the conception that learning is construction. in contrast, the “weak version” of cognitive metaphor theory assumes that both the source and the target concept of the metaphor are more or less developed separate cognitive structures. under this view, a certain conception is the reason why a person can identify features that are mappings between one's own conception and a certain metaphor (haser, 2005). thus, a person who has the conception of learning as a construction of knowledge would single out identical features between learning and building a skyscraper, but not between learning and eating, and thus prefer to use the metaphor of learning as building a skyscraper then as eating a cake as a descriptor. in research using metaphors for assessing conceptions we can find works based on the “strong” and the “weak" versions of cognitive metaphor theory. those researchers who are interested in examining the development and change of conceptions, for example in the context of educational development programs (e.g. bullough, 1991; clandinin, 1985), tend to argue on the base of a strong version of cognitive metaphor theory because they usually assume that changing the metaphor a person uses also leads to a change in the person’s conception. in contrast, researchers using metaphors mainly for assessment of conceptions (e.g. saban, kocbeker & saban, 2007; patchen & crawford, 2011) usually argue on the base of a weak version of the cognitive metaphor theory, assuming that metaphors help to express or to identify an underlying conception. based on the longstanding tradition on research on conceptions of teaching and learning (e.g. gow & kember, 1993; vermunt & vermetten, 2004), we assume that there are indeed underlying wegner  &  nückles       | f l r     99   conceptions that are separate from metaphors, and thus would adhere to a weak version of cognitive metaphor theory. we assume that the underlying conception enables or prompts a person to identify structural mappings between one's own conception and the metaphor. two principally different approaches can be discerned in assessing conceptions via metaphors (löfström et. al., 2015). on the one hand, researchers themselves generate metaphors and use them as a stimulus for assessing conceptions. for example, some researchers have developed questionnaires in which participants are asked to rate metaphors (e.g., lehmann, 2012). others have asked participants to reflect on preselected written metaphors (e.g., visser-wijnween, van driel, van der rijst, verloop & visser, 2009) or metaphorical pictures (ben-peretz, mendelson & kron, 2003), and analysed the participants’ responses with regard to the underlying conception. in both cases, the participants mapped preselected metaphors to their own conception. on the other hand, researchers also asked participants to produce metaphors on their own, and then analysed these metaphors according to their conceptual content. for example, saban et al. (2007) asked more than 1000 students to write down a metaphor on being a teacher and identified six dominant conceptual mappings for the metaphors: knowledge provider, craftsperson, facilitator, nurturer, counsellor and democratic leader. interestingly, these conceptual categories are similar to conceptions of teaching as described in the “teaching perspectives inventory” by collins and pratt (2011), namely, transmission, apprenticeship, developmental, nurturing and social reform. other studies (patchen & crawford, 2011; wegner & nückles, 2015a) classified metaphors based on the two scientific paradigms of learning as acquisition vs. as participation according to sfard (1998). only a few studies focus on students’ metaphors. inbar (1996) asked more than 400 students for metaphors on ‘being a student’ and on ‘teachers’. a great proportion of metaphors was related to feeling imprisoned in school, showing largely negative emotions towards school. marsch (2009) analysed high school students’ metaphors of biology learning. she found that most students conveyed an idea of learning as intake of knowledge. in a longitudinal study with students from educational science, wegner and nückles (2015b) found that students adapted their metaphors of learning to university learning culture in the course of their first year of studying. while in the first year, the most frequently used metaphor of learning was “collecting”, the most frequently used metaphor in the second year described learning as “discovering”. even though empirical studies do indicate that different views on teaching or learning are visible in metaphors, and there are theoretical arguments for a close relationship between metaphors and conceptions, there are few studies which really validate whether different metaphors also account for differences in underlying conceptions, and even less, whether they also account for differences in actual practice. moreover, all of the existing validation studies are case studies with very small samples, or just report data on selected cases illustrating their hypotheses (e.g. mahlios, massengill-­‐shaw, & barry, 2010; bullough, 1991; marsch, 2009; thomas & mcrobbie, 1999). some larger studies link metaphors of teaching to other self-reported data, but not to practice (wegner & nückles, 2015a; löfström & poom-valickis, 2013). thus, there is a need for empirical studies validating whether metaphors of learning can indeed be an indicator for conceptions of learning, and whether students’ metaphors of learning really relate to how students learn in terms of which learning strategies they use and what their motivation is. 1.4 summary and aims of the study in sum, we can conclude that students’ learning patterns are influenced by the individual understanding by students of what learning is, that is, their conceptions of learning. first evidence from studies on conceptions of teaching indicates that metaphors might also be an appropriate and helpful tool for assessing conceptions of learning, and that differences in metaphors of learning are also associated with differences in students’ learning practice. however, so far there are only few studies analysing the relation between metaphors and practice, and studies are only based on small sample sizes. in our study, we aimed at closing these gaps by (a) exploring whether the different conceptions of learning as they have been described in the literature are also visible in the metaphors that students use to describe learning, and (b) examining whether differences in the conceptual content of the metaphors account for differences in learning practice, such as the use of learning strategies, study motivation, and epistemological beliefs. wegner  &  nückles       | f l r     100   2. methods 2.1 participants and procedure ninety-one students of educational science from a german university took part in the study (78.1% female and 21.9% male, meanage =23.81 years, sdage =3.38). all students were first given a short example of what we meant by metaphor, and were then asked to write down their metaphors of learning. afterwards they filled-in questionnaires on learning-related measures. all measurements took place in university courses in the institute of educational science and were set at the beginning of a lesson. 2.2 questionnaires for assessing learning-related measures, we chose questionnaires on motivation, learning strategies and epistemological beliefs which are well established for german language speakers and which address central aspects of learning patterns and approaches to learning (for an overview of the scales and their reliabilities, see table 1). table 1 scales of the questionnaires, scale reliability (cronbach’s ɑ), mean values (m), standard deviation (sd) and number of items.   sample item ɑ m sd no. of items intrinsic motivation i don’t need a reward for completing the study tasks because they are fun. . 734 4.96 0.91 5 extrinsic motivation i will be quite proud when i have completed my degree. .726 5.82 0.96 3 organisation i draw tables and graphs in order to structure the contents of the subject. .791 3.68 0.61 8 elaboration i try to relate new concepts or theories to familiar concepts or theories. .742 3.67 0.55 8 critical thinking i examine whether theories, interpretations or conclusions are sufficiently grounded. .873 3.11 0.68 8 rehearsal i re-read my notes again and again. .827 3.14 0.75 7 metacognitive strategies before i start with learning, i try to plan which contents i do need to know and which i don’t. .745 3.59 0.47 11 time management i schedule time slots for studying. .897 2.99 0.95 5 learning with others i work on texts and tasks together with my colleagues. .849 3.41 0.77 7 relativism scientific research shows that there is one right answer to most problems. .635 1.69 .40 6 dualism if two scientists have a different opinion on a matter, one of them has to be wrong. .613 1.68 .43 4 wegner  &  nückles       | f l r     101   the use of learning strategies was assessed by seven scales of a german questionnaire (list; wild & schiefele, 1994) which is based on the motivated strategies for learning questionnaire (mslq; pintrich, smith, garcia, & mckeachie, 1993). participants rated on a five-point rating scale how often they engage in certain learning activities (“in the following, we would like to know about how you learn. you will find a list of learning activities. please indicate for each activity, how often it occurs when you are learning. you can rate the frequency between very seldom (1) and very often (5)”). the selected activities addressed four cognitive strategies (organization of contents, elaboration of contents, rehearsal, critical thinking), metacognitive strategies, use of time management strategies and the frequencies of learning with others. motivational orientation was assessed by two scales of the intrinsic motivation inventory (imi; deci & ryan, 2003), one on intrinsic motivation and one on the extrinsic value of studying in a version adapted to the context of higher education. students were instructed to rate the items on their seven-point rating scale ranging from completely disagree to completely agree (“please indicate for each statement how much you agree. […] these questionnaires are not evaluated! there are no “right” or “wrong” answers.”) epistemological beliefs in general were assessed by a german questionnaire on epistemological beliefs (köller, watermann, trautwein, & lüdtke, 2004). it comprises two dimensions, “dualism” (sample item: “if two scientists have a different opinion on a matter, one of them has to be wrong.”) and “relativism” (sample item: “scientific insights that seem true today can turn out to be wrong”). participants had to rate the statements on a four-point rating scale ranging from totally disagree (= 1) to totally agree (= 4). 2.3 assessment and analysis of metaphors following saban (saban et al., 2007), students had to answer the questions “learning is like… because…”. in order to enrich the answers, we added the question “the goal of learning is…”. metaphors were analysed following chi’s recommendations on coding verbal data (1997). two metaphors were excluded from the analysis because they were only fragments. one metaphor as a whole was defined as the unit of analysis, that is, the complete answer consisting of the source and explanation of the metaphor, because sometimes the same source was associated with different kinds of explanations (e.g. “learning is like food: you need it for survival” vs. "learning is like eating food: if you eat too much, you get sick”). we then inductively developed a system of categories within a team of two researchers. all decisions were also discussed within a larger research team, consisting of four researchers in total. as in other studies (e.g. inbar, 1996; leavy, mcsorley & boté, 2007) , we found a large amount of metaphors without conceptual content, but merely related to aspects of regulating one’s own learning and motivation, such as in "learning is like jumping into cold water. usually you don’t want to do it, but once you get started, it’s always good”. therefore, regulation-related metaphors were first separated from other metaphors. in the second step, the remaining metaphors were classified according to the conceptual content. we distinguished three different kinds of metaphors: learning as acquisition of knowledge vs. learning as problem solving vs. learning as development of personality (see table 2). for each category, a short description was written down with examples. then,half of the metaphors (n=43) were coded by a second independent person. interraterreliability as measured by cohen’s κ was very good (κ = .81). 3. results conceptions of learning as described in literature were visible in our metaphors. of the four categories of metaphors, knowledge acquisition was the most common (30.3%), followed closely by regulation-related metaphors (28.1%). personality development metaphors were described by 25.8% of the students, and only 15.7% of the students used metaphors which focused on learning as a prerequisite for solving problems (table 3). wegner  &  nückles       | f l r     102   table 2   categories of metaphors, description and anchoring examples for each category of metaphor in the next step, we determined whether students with different kinds of metaphors differed with regards to their epistemological beliefs, their study motivation, and their learning strategies. an overall manova with type of metaphor as independent measure, and epistemological beliefs, motivation and learning strategies as dependent measures showed a significant multivariate effect of metaphor type, f(33, 231) = 2.31, p < .001, η2=.25 (see table 3 for an overview of the descriptive data for the four kinds of metaphors). separate univariate anovas revealed significant differences for intrinsic motivation, f(3, 85) = 4.31, p < .01, η2=.13, for dualism f(3,85) = 2.78, p < .05, η2=.09, and the use of rehearsal strategies, f(3,85) = 4.31, p < .01, η2= .14. students with problem solving and development metaphors indicated a higher intrinsic motivation than students with regulation-related or knowledge acquisition metaphors (see table 3). students with personality development metaphors had the lowest scores on the dualism scale, while students with knowledge acquisition metaphors had the highest, indicating that students with knowledge acquisition metaphors believed much stronger that knowledge is either true or false than students with personality development metaphors. students with knowledge acquisition metaphors also had the strongest tendency to use rehearsal strategies, followed by students with regulation-related and personality development metaphors. students with problem-solving metaphors had the lowest scores on this scale. category description example n regulationrelated metaphors the metaphor and its explanation refer to self-regulation aspects and do not contain any information about cognitive processes or further goals in learning. “learning is like jumping into cold water. usually you don’t want to do it, but once you get started, it’s always good.” “learning is like climbing a mountain. some hills are steep, and others are easy to walk.” 25 (28.1%) acquisition of knowledge learning consists of the acquisition of something (=knowledge). there is no further indication that the acquired knowledge is used for something. “learning is like building a library with your own books. you start with one shelf and while you get more and more books you also need more shelves.” “learning is like solving a jigsaw puzzle … the goal is to solve the jigsaw puzzle and to get the complete picture.” 27 (30.3%) problem solving learning consists of the acquisition of something (= skills and knowledge) which are necessary to solve certain problems, to be prepared for future challenges or to be able to work in a certain job. „learning is like food – you need it for survival. without it you cannot deal with new problems.” “learning is like getting a closet with lots of clothes. at the beginning of your life you have only a few pieces of clothes, later you get more and more […]. the goal is to buy, to select, to sort, to categorize the clothes so you can use them and wear them when you need them.” 14 (15.7%) development of personality learning consists of developing something existing further, in order to develop one's own personality or new perspectives. “learning is like exploring other countries. you get to know new cultures and new perspectives, and you widen your horizon.” “learning is like a plant that is growing, because you thrive and prosper inside.” 23 (25.8%) total 89 (100%) wegner  &  nückles       | f l r     103   table 3 means and standard deviation for study motivation, epistemological beliefs and learning strategies for each group of metaphors regulationrelated acquisition of knowledge problem solving development of personality intrinsic motivation 4.70 (0.76) 4.66 (0.98) 5.36 (0.90) 5.33 (0.81) extrinsic motivation 5.70 (1.11) 5.91 (0.89) 6.11 (0.66) 5.70 (1.03) relativism 1.77 (0.42) 1.80 (0.44) 1.63 (0.42) 1.52 (0.29) dualism 1.66 (0.37) 1.82 (0.46) 1.75 (0.38) 1.49 (0.43) critical thinking 3.06 (0.75) 3.01 (0.58) 3.01 (0.71) 3.35 (0.71) learning with others 3.42 (0.68) 3.50 (0.86) 3.00 (0.62) 3.58 (0.80) elaboration 3.70 (0.53) 3.69 (0.54) 3.57 (0.48) 3.68 (0.65) organisation 3.53 (0.37) 3.80 (0.58) 3.68 (0.83) 3.71 (0.71) rehearsal 3.05 (0.75) 3.52 (0.57) 2.70 (0.82) 3.06 (0.75) metacognitive strategies 3.57 (0.43) 3.75 (0.55) 3.43 (0.46) 3.52 (0.42) time management 2.73 (0.78) 3.29 (0.95) 3.30 (0.92) 2.75 (1.06) to better understand the overall differences between the groups, and the patterns of motivation, epistemology and learning strategies for each group, we performed a discriminant analysis with epistemological beliefs, motivation, and learning strategies as predictors and the kind of metaphors as criterion. it resulted in three discriminant functions. the first discriminant function explained half of the variance, 56.8%, canonical r2 = .38; the second discriminant function explained one third of the variance, 33.2% canonical r2 = .26. the third discriminant function explained the remaining 9.9% of the variance, canonical r2 = .09. together, the three functions significantly differentiated between the metaphor types (wilk’s λ = .41, χ2(33) = 71.89, p = .000). after removing the first function, the remaining two functions still contributed significantly to the classification of the metaphors (wilk’s λ = .66, χ2(20) = 33.12, p = .03). however, the last function on its own could not differentiate between the metaphors. figure 1 shows the distribution of the four metaphors among the two separating functions. correlations of the predicting variables with each canonical discriminant function are given in table 4. a closer look at the discriminant functions revealed that the first function separated the students with regulation-related and with knowledge acquisition metaphors from the students with personality development and problem solving metaphors, whereas the second function mainly separated the students with problem solving metaphors from the students with personality development metaphors, see fig. 1. the first function was associated with high beliefs in the certainty of knowledge (i.e., low relativism), and with a low intrinsic motivation, see table 4 and fig. 2), thus indicating that students with knowledge acquisition and regulation-related metaphors were less intrinsically motivated and believed to a higher extent that knowledge is certain and unambiguous. the second function correlated positively with extrinsic motivation and the belief in the dualism of knowledge, and negatively with an extra preference for critical thinking, for learning with other students and for elaboration of contents (see table 4). this indicates that students with problem solving metaphors were more extrinsically motivated, believed more that knowledge was either wrong or right, and were less inclined to critically think about the contents or to discuss them with colleagues, than students with personality development metaphors. wegner  &  nückles       | f l r     104   figure 1. plot of the group centroids of the four metaphors with regard to the two discriminant functions. function 1 separates problem solving and personality development metaphors from the regulation-related and knowledge acquisition metaphors. function 2 separates personality development from the problem solving metaphors. table 4 correlations between discriminant functions and the predicting variables. bold print indicates the highest correlating function for each predictor variable function 1: instrinsic motivation (-) and variability of knowledge function 2: extrinsic motivation deep processing (-), dualism function 3: structured learning instrinsic motivation -.467 -.144 .191 relativism (general certainty beliefs) .299 .250 -.198 dualism (beliefs in simple knowledge) .179 .460 .140 learning with others .172 -.316 .263 critical thinking -.107 -.309 .140 external motivation -.068 .253 .149 elaboration .062 -.097 -.078 rehearsal .424 -.003 .702 organisation -.008 .051 .475 time management -.002 .421 .461 metacognitive strategies .269 .060 .357 knowledge acquisition regulation-related personality development problem solving wegner  &  nückles       | f l r     105   the last canonical discriminant function helped to differentiate the four groups only together with the second function. on this function, we found high loadings of measures indicating structured learning, such as the strategies of organization and rehearsal, metacognitive strategies and time management (see fig. 2). the function differentiated between regulation-related metaphors and knowledge acquisition metaphors, with students with knowledge acquisition metaphors showing more use of structured learning than students with regulation-related metaphors, that is, students with metaphors which just focus on aspects relating to the regulation of their learning or their motivation rather than on the results or the process of learning. however, as noted above, the third function could not discriminate between the groups on its own.   figure 2. z-standardized mean values for each of the metaphor categories. the variables of the first function are printed in black/bold (discriminating between the regulation-related and knowledge acquisition metaphors on the one hand, and the problem-solving and the personality development metaphors on the other hand). variables of the second function are given in hatched/italics (discriminating between problem solving metaphors and personality development metaphors). 4. discussion and conclusion in our study, we could distinguish four kinds of metaphors of learning, namely metaphors focusing on regulation aspects of learning, metaphors expressing the idea of learning as knowledge acquisition and the idea of learning as personality development, and metaphors focusing on learning as a prerequisite for problem solving. students’ metaphors of learning predicted different patterns of motivation, epistemology and use of learning strategies. students with problem solving and with personality development metaphors differed in their intrinsic motivation and their awareness for the tentativeness of knowledge from students wegner  &  nückles       | f l r     106   with knowledge acquisition and with regulation-related metaphors. students with personality metaphors could be separated from students with problem solving metaphors by their use of deep processing strategies, their belief in the dualism of knowledge and their extrinsic motivation. finally, students with knowledge acquisition metaphors had a tendency to engage more in structured learning activities than students with regulation-related metaphors, though not significantly so. metaphors of learning predicted study motivation, epistemological beliefs and learning strategies. this implies that metaphors can be used to detect differences in conceptions of learning. the predicted learning patterns mirror in some respects both the learning patterns and the approaches to learning model. personality development metaphors seem to predict a meaning-directed learning pattern or a deep approach, because students with personality development metaphors displayed a high intrinsic study motivation, a high awareness for the tentativeness and the complexity of knowledge and indicated to make much use of deep processing strategies. this finding confirms results from entwistle and mccune (2013), who found that there is a certain group of students that have a ‘disposition to understand for oneself’. this disposition seems to be based on the view of learning as development of personality. students with problem solving metaphors have similarities with students with an application-directed learning pattern as described by vermunt (1996), because the application directed mental model of learning is based on the use of knowledge as well. problem-solving metaphors were also associated with strong extrinsic motivation for studying, but were, other than students with the application-directed learning pattern, also more intrinsically motivated. on the other hand, students with problem-solving metaphors made only average use of concrete processing strategies such as elaboration of contents, which would have been expected in an application-directed learning pattern. knowledge acquisition metaphors seem to be similar to vermunts’ rehearsal-directed learning pattern, because they are also characterized by a mental model of intake of knowledge. as students with a rehearsal-directed learning approach, students with knowledge acquisition metaphors had an extrinsic study motivation and believed in the stability of knowledge. however, other than students with a rehearsal-directed learning pattern, students with knowledge acquisition metaphors in our sample also described the use of deep learning strategies and structured their learning activities strongly. in this respect, they seem more similar to the strategic approach described by biggs (1987), which is characterized by good organization, good time management, and alertness to the assessment requirements and criteria. this would indicate that acquisition and elaboration of knowledge are seen as the dominant requirement within the degree under consideration. finally, regulation-related metaphors have a great overlap with the undirected learning pattern. the metaphors did not convey a mental model of learning, students had little intrinsic motivation and they did not engage in deep or structured learning activities. this is interesting in several respects. on the one hand, these findings mirror those of other studies in which participants described metaphors with no apparent match to conceptions of teaching or learning. for example, in inbar's (1996) study with high school students, most metaphors were related to emotional aspects of learning and did not reveal anything about underlying conceptions of learning. similarly, in their study on teacher candidates’ metaphors of teaching, leavy et al. (2007) report a great number of metaphors that did “not refer to components central to the practice of teaching, but referred to what teaching meant to the individuals themselves (e.g. ‘teaching is like running a marathon; you train, sweat, and prepare for this great race but once you’re in it, you just keep going strong until the end’)” (p. 1226). in one group of the sample, 30% of the metaphors were ‘self-referential’. such self-referential metaphors can be found in many studies using metaphors (e.g. leavy et. al., 2007; zapata & lacorte, 2007; löfström & poom-valickis, 2013). findings from our study might be a first indicator that the participants who use such self-referential, emotional or motivational metaphors have not yet developed a differentiated explicable conception which can be communicated by a metaphor. considering the unorganized use of learning strategies in this group, the finding could be interpreted in the way that a lack of an elaborated conception of learning is a major problem for developing adequate learning strategies. consequently, to these students challenges of self-regulation are the most distinct experience of learning. however, further research is needed to confirm this hypothesis. wegner  &  nückles       | f l r     107   of course, some limitations have to be born in mind. again, our study only assessed self-report data on participants’ use of learning strategies in general. therefore, we do not know how students’ answers relate to their actual practice of learning or on what they think they should do. also, course requirements, which influence strongly how students actually learn, need to be considered (vermetten, lodewijks & vermunt, 1999). however, if students were biased in their answers on their learning strategy use, the differences in self-report data between the four kinds of metaphors indicate at least that students differ in what they think is a socially desirable answer. another limitation is that we assessed metaphors just in one context at one point of time. so we cannot draw conclusions about whether metaphors are stable across contexts or over time, as conceptions would be. nevertheless, metaphors seem to be a promising research tool which should receive further attention for research on conceptions of learning, because it seems indeed to matter whether students see learning as a matter of training their brains or tending their gardens. keypoints students’ “metaphors of learning” discriminated between different profiles of motivation, epistemological beliefs and use of learning strategies. different categories of metaphors could be linked to both learning patterns and approaches to learning. students describing learning in terms of personality development shared similarities with deep approach learners and meaning-directed learners. students focussing in their metaphors on only the regulation aspects of learning shared similarities with undirected learners. students describing learning in terms of knowledge acquisition shared similarities either with rehearsal-directed learners or with learners with a strategic approach. references biggs, j. b. (1987). student approaches to learning and studying. research monograph. australian council for educational research ltd., radford house, frederick st., hawthorn 3122, australia. ben-peretz, m., mendelson, n., & kron, f. w. (2003). how teachers in different educational contexts view their roles. teaching and teacher education, 19(2), 277–290. doi:10.1016/s0742-051x(02)00100-2 black, m. (1993). more about metaphor. in a. ortony (ed.), metaphor and thought (2nd. edition) (pp. 19– 41). bullough, r. v. (1991). exploring personal teaching metaphors in preservice teacher education. journal of teacher education, 42(1), 43–51. doi: 10.1177/002248719104200107 chi, m. t. h. (1997). quantifying qualitative analyses of verbal data: a practical guide. the journal of the learning sciences, 6(3), 271–315. doi: 10.1207/s15327809jls0603_1 clandinin, d. j. (1985). personal practical knowledge: a study of teachers' classroom images. curriculum inquiry, 15(4), 361-385. doi: 10.2307/1179683 collins, j. b., & pratt, d. d. (2011). the teaching perspectives inventory at 10 years and 100,000 respondents: reliability and validity of a teacher self-report inventory. adult education quarterly november, 61(4), p. 358-375. doi: 10.1177/0741713610392763. deci, e. l. & ryan, r. m. (2003). intrinsic motivation inventory (imi). retrieved in oct. 2012 from http://www.selfdeterminationtheory.org/questionnaires/ entwistle, n. j., & peterson, e. r. (2004). conceptions of learning and knowledge in higher education: relationships with study behaviour and influences of learning environments. international journal of educational research, 41(6), 407–428. doi:10.1016/j.ijer.2005.08.009 wegner  &  nückles       | f l r     108   entwistle, n. j., & ramsden, p. (1983). understanding student learning. london and canberra: croom helm. entwistle, n., & mccune, v. (2013). the disposition to understand for oneself at university: integrating learning processes with motivation and metacognition. british journal of educational psychology, 83(2), 267–279. doi:10.1111/bjep.12010 entwistle, n., tait, h., & mccune, v. (2000). patterns of response to an approaches to studying inventory across contrasting groups and contexts. european journal of psychology of education, 15(1), 33–48. doi:10.1007/bf03173165 gick, m. l., & holyoak, k. j. (1980). analogical problem solving. cognitive psychology, 12(3), 306–355. doi: 10.1016/0010-0285(80)90013-4 gentner, d., & holyoak, k. j. (1997). reasoning and learning by analogy: introduction. american psychologist, 52(1), 32–34. doi:10.1037/0003-066x.52.1.32 gow, l., & kember, d. (1993). conceptions of teaching and their relationship to student learning. british journal of educational psychology, 63(1), 20-23. doi: 10.1111/j.2044-8279.1993.tb01039.x haser, v. (2005). metaphor, metonymy, and experientialist philosophy: challenging cognitive semantics (vol. 49). walter de gruyter. inbar, d. e. (1996). the free educational prison: metaphors and images. educational research, 38(1), 77– 92. doi: 10.1080/0013188960380106 köller, o., watermann, r., trautwein, u. & lüdtke, o. (2004). wege zur hochschulreife in badenwürttemberg: tosca—eine untersuchung an allgemein bildenden und beruflichen gymnasien. opladen: leske + budrich. lakoff, g., & johnson, m. (1980). conceptual metaphor in everyday language. the journal of philosophy, 77(8), 453–486. doi: 10.2307/2025464 landau, m. j., meier, b. p., & keefer, l. a. (2010). a metaphor-enriched social cognition. psychological bulletin, 136(6), 1045-1067. 10.1037/a0020970 leavy, a. m., mcsorley, f. a., & boté, l. a. (2007). an examination of what metaphor construction reveals about the evolution of preservice teachers’ beliefs about teaching and learning. teaching and teacher education, 23(7), 1217–1233. doi: 10.1016/j.tate.2006.07.016 lehmann, b. (2012). entwicklung eines instruments zur erfassung unterrichtsbezogener metaphern [development of an instrument for assessing teaching-related metaphors]. in: faßhauer, u., fürstenau, b., wuttke, e. (eds.): berufsund wirtschaftspädagogische analysen – aktuelle forschungen zur beruflichen bildung. [analyses in pedagogy of economics and vocation – current research in vocational training] (p.127-139) opladen: leske + budrich. löfström, e., nevgi, a., wegner, e., & karm, m. (2015). images in research on teaching and learning in higher education. in: j. huisman & m. tight (eds.), theory and method in higher education research, volume 1. (p. 191-212). emerald group publishing limited. löfström, e., & poom-valickis, k. (2013). beliefs about teaching: persistent or malleable? a longitudinal study of prospective student teachers' beliefs. teaching and teacher education, 35, 104–113. doi:10.1016/j.tate.2013.06.004 mahlios, m., massengill-­‐shaw, d., & barry, a. (2010). making sense of teaching through metaphors: a review across three studies. teachers and teaching: theory and practice, 16(1), 49–71. doi:10.1080/13540600903475645 marsch, s. (2009). metaphern des lehrens und lernens: vom denken, reden und handeln bei biologielehrern. [metaphors of teaching and learning: thinking, talking and practice of biology teachers.] (dissertation) freie universität berlin, berlin. marton, f., & säljö, r. (1976). on qualitative differences in learning: i—outcome and process british journal of educational psychology, 46(1), 4–11. doi:10.1111/j.2044-8279.1976.tb02980.x murphy, g. l. (1996). on metaphoric representation. cognition, 60(2), 173-204.  10.1016/00100277(96)00711-1 pajares, f. m. (1992). teachers' beliefs and educational research: cleaning up a messy construct. review of educational research, 62(3), 307. doi: 10.3102/00346543062003307 wegner  &  nückles       | f l r     109   parpala, a., lindblom-ylänne, s., komulainen, e., litmanen, t., & hirsto, l. (2010). students' approaches to learning and their experiences of the teaching–learning environment in different disciplines. british journal of educational psychology, 80(2), 269–282. doi:10.1348/000709909x476946 patchen, t., & crawford, t. (2011). from gardeners to tour guides: the epistemological struggle revealed in teacher-generated metaphors of teaching. journal of teacher education, 62(3), 286–298. doi:10.1177/0022487110396716 pintrich, p. r., smith, d. a. f., garcia, t., & mckeachie, w. j. (1993). reliability and predictive validity of the motivated strategies for learning questionnaire (mslq). educational and psychological measurement, 53(3), 801–813. doi:10.1177/0013164493053003024 richardson, j. t. e. (2007). mental models of learning in distance education. british journal of educational psychology, 77(2), 253–270. doi:10.1348/000709906x110557 richardson, j. t. e. (2011). approaches to studying, conceptions of learning and learning styles in higher education. learning and individual differences, 21(3), 288–293. doi: 10.1016/j.lindif.2010.11.015 saban, a., kocbeker, b. n., & saban, a. (2007). prospective teachers' conceptions of teaching and learning revealed through metaphor analysis. learning and instruction, 17(2), 123–139. doi:10.1016/j.learninginstruc.2007.01.003 säljö, r. (1979). learning about learning. higher education, 8(4), 443-451. 10.1007/bf01680533 sfard, a. (1998). on two metaphors for learning and the dangers of choosing just one. educational researcher, 27(2), 4-13. 10.3102/0013189x027002004 thomas, g. p., & mcrobbie, c. j. (1999). using metaphor to probe students' conceptions of chemistry learning. international journal of science education, 21(6), 667–685. doi:10.1080/095006999290507 vanthournout, g., donche, v., gijbels, d., & van petegem, p. (2014). (dis)similarities in research on learning approaches and learning patterns. in d. gijbels, v. donche, j. t. e. richardson, & j. d. vermunt (eds.), learning patterns in higher education: dimensions and research perspectives (pp.1132). new york, oxford: routledge. vermunt, j.d. (1994). scoring key for the inventory of learning styles (ils) in higher education. tilburg: tilburg university. vermunt, j. d. (1996). metacognitive, cognitive and affective aspects of learning styles and strategies: a phenomenographic analysis. higher education, 31(1), 25–50. doi:10.2307/3447707 vermunt, j. d, & vermetten, y. (2004). patterns in student learning: relationships between learning strategies, conceptions of learning, and learning orientations. educational psychology review, 16(4), 359–384. doi:10.1007/s10648-004-0005-y vermetten, y. j., lodewijks, h. g., & vermunt, j. d. (1999). consistency and variability of learning strategies in different university courses. higher education, 37(1), 1-21.   doi:10.1023/a:1003573727713 virtanen, v., & lindblom-ylänne, s. (2010). university students’ and teachers’ conceptions of teaching and learning in the biosciences. instructional science, 38(4), 355–370. doi: 10.1007/s11251-008-9088-z visser-wijnveen, g., van driel, j., van der rijst, r., verloop, n., & visser, a. (2009). the relationship between academics' conceptions of knowledge, research and teaching–a metaphor study. teaching in higher education, 14(6), 673–686. doi: 10.1080/13562510903315340 wegner, e., & nückles, m. (2015a). knowledge acquisition or participation in communities of practice? academics’ metaphors of teaching and learning at the university. studies in higher education, 40(4), 624-643. doi: 10.1080/03075079.2013.842213 wegner, e. & nückles, m. (2015b). from eating to discovering: how metaphors of learning change during students’ enculturation. zeitschrift für hochschulentwicklung, 10 (4), 145-166. wild, k.-p., & schiefele, u. (1994). lernstrategien im studium: ergebnisse zur faktorenstruktur und reliabilität eines neuen fragebogens: [learning strategies of university students: factor structure and reliability of a new questionnaire.]. zeitschrift für differentielle und diagnostische psychologie, 15(4), 185–200. zapata, g. c., & lacorte, m. (2007). preservice and inservice instructors' metaphorical constructions of second language teachers. foreign language annals, 40(3), 521–534. doi:10.1111/j.19449720.2007.tb02873.x   schindler et al publication frontline learning research vol.7 no. 2 (2019) 23 39 issn 2295-3159 effectiveness of self-generation during learning is dependent on individual differences in need for cognition julia schindlera, simon schindlerb marc-andré reinhardb auniversity of würzburg, germany buniversity of kassel, germany article received 3 september 2018 / revised 18 february/ accepted 15 april / available online 7 may abstract self-generated information is better recognized and recalled than read information. this so-called generation effect has been replicated several times for different types of stimulus material, different generation tasks, and retention intervals. the present study investigated the impact of individual differences in learners’ disposition to engage in effortful cognitive activities (need for cognition, nfc) on the effectiveness of self-generation during learning. learners low in nfc usually avoid getting engaged in cognitively demanding activities. however, if these learners are explicitly instructed to use elaborate learning strategies such as self-generation, they should benefit more from such strategies than learners high in nfc, because self-generation stimulates cognitive processes that learners low in nfc usually tend not to engage in spontaneously. using a classical word-generation paradigm, we not only replicated the generation effect in free and cued recall but showed that the magnitude of the generation effect increased with decreasing nfc in cued recall. results are consistent with our assumption that learners higher in nfc engage in elaborate processing even without explicit instruction, whereas learners lower in nfc usually avoid cognitively demanding activities. these learners need cognitively demanding tasks that require them to switch from shallow to elaborate processing to improve learning. we conclude that self-generation is beneficial regardless of the nfc level, but our study extends the existing literature on the generation effect and on nfc by showing that self-generation can be particularly useful for balancing the learning disadvantage of students lower in nfc. keywords: desirable difficulties; generation effect; incidental learning; intentional learning; need for cognition info corresponding author: julia.schindler@uni-wuerzburg.de doi: 10.14786/flr.v7i2.407 1. introduction students often assume that learning strategies that are perceived as easy and effortless (e.g., rehearsal, rereading, or underlining) are highly effective. however, extant research suggests that under certain conditions learning is more effective when learners intentionally make the learning process more difficult (bjork and bjork 2011). specific difficulties, such as distributed learning sessions compared to massed learning (e.g., cepeda et al. 2006), interleaving topics (e.g., dunlosky et al. 2013), testing new knowledge (e.g., roediger and karpicke 2006), and self-generation of information (e.g., mcdaniel et al. 1988; slamecka and graf 1978), stimulate processes which are beneficial to the learning process. such difficulties often result in long-lasting memory for the learned material and make it easier to apply the acquired knowledge to new situations. thus, they are termed desirable difficulties (bjork 1994). the so-called generation effect (learners recall self-generated information better than read information) has been investigated extensively (for a meta-analysis, see bertsch et al. 2007). several extant studies used the classical word-generation paradigm by slamecka and graf (1978; see also e.g., mcdaniel et al. 1988). in these studies, learners were presented with word pairs consisting of a context word and a target word. in the read condition, learners read an associated word pair (saddle – horse). in the generate condition, they completed fragmented target words (saddle – h_ _ _ _) with the aid of the context word and a specific encoding rule (e.g., find the associated word). in subsequent learning tests, learners recalled and recognized generated target words better than read target words. the learning advantage of generated over read information has been replicated, for example, for different learning measures (recognition, cued recall, and free recall, e.g., mcdaniel et al. 1988; slamecka and graf 1978), for target and context words (mcdaniel and waddill 1990), for sentences (graf 1980, 1981), texts (e.g., doctorow et al. 1978; mcdaniel et al. 1986; mcdaniel et al. 2002), and numbers (e.g., gardiner and rowley 1984). it has been shown for immediate and delayed recall (e.g., schweickert et al. 1994; slamecka and frevreiski 1983), for withinvs. between-subject designs (fiedler et al. 1992), and for different generation rules (slamecka and graf 1978, exp.1 and 2). empirical studies like these suggest that self-generation might be a useful supplement to commonly used learning strategies in education. despite the extensive body of extant research on the generation effect, its general conditions of occurrence still need further clarification. mcdaniel and butler (2010) pointed out that self-generation is not necessarily beneficial for every learner and not appropriate for every type of learning material and criterial task. instead, they assume that complex interactions between the type of generation task, learner characteristics, learning material, and criterial task need to be considered when using self-generation to improve learning (see also mcdaniel and einstein 1989, 2005; einstein et al. 1990). the central claim of mcdaniel and butler’s contextual framework is that “desirable difficulties are those that stimulate processing that is not redundant with the processing spontaneously engaged by the learner (which […] will depend on learner characteristics, materials, or both) and that matches the demands of the criterial task” (p.179). in other words, self-generation can improve learning only when the generation task stimulates cognitive processes that go beyond the processes individual learners engage in spontaneously during learning or beyond processes encouraged by specific learning-material characteristics. one widely researched learner characteristic shown to differentially affect the degree that learners spontaneously engage in cognitive processing is a learner’s need for cognition. hence, need for cognition, in turn, is likely to affect the effectiveness of self-generation during learning. 2. need for cognition and the generation effect need for cognition (nfc) can be defined as a learner’s individual disposition to engage in effortful cognitive activities and to enjoy thinking and being cognitively challenged (cacioppo and petty 1982; cacioppo et al. 1984; cacioppo et al. 1986). high nfc is associated with thorough processing of arguments and argument quality (cacioppo et al. 1983; cacioppo et al. 1986), with thorough processing of task-relevant information (fleischhauer et al. 2014; reinhard 2010, exp. 1 & 2; reinhard and dickhäuser 2009; verplanklen et al. 1992), and thorough processing of learning materials (sadowski and gülgöz 1996). moreover, individuals high in nfc use more efficient learning strategies than learners low in nfc (cazan and indreica 2014), they tend to have better self-control during learning (bertrams and dickhäuser 2009; cazan and indreica 2014), and they are more willing to tackle difficult tasks (see et al. 2009; weißgerber et al. 2018). based on these findings, it is not surprising that individuals high in nfc recall learned information better than individuals low in nfc (cacioppo et al. 1983; kardash and noel 2000). they are more likely to solve complex problems or tasks (coutinho 2006; coutinho et al. 2005; nair and ramnarayan 2000) and they perform better on learning tests (heijne-penninga et al. 2010; sadowski and gülgöz 1996). consequently, individual differences in nfc are associated with academic achievement (luong et al. 2017). high nfc was found to be associated with course achievements (bertrams and dickhäuser 2009; sadowski and gülgöz 1996), university gpa (grade point average) (grass et al. 2017), course grades mediated by difficulty of learning material (leone and dalton 1988), and performance in exams mediated by self-regulated learning and deep information processing (cazan and indreica 2004). in sum, nfc seems to be directly or indirectly related to learner characteristics relevant for academic success and to different forms of academic performance and achievement measures (for a review see jebb et al. 2016; see also the meta-analyses by richardson et al. 2012 and von stumm and ackermann 2013). learners low in nfc are ‘cognitive misers’ (cacioppo et al. 1986; cacioppo et al. 1996) who usually avoid getting engaged in cognitively demanding activities. consistent with these findings, low nfc learners are found to be less willing to use elaborate learning strategies such as desirable difficulties in self-regulated learning than high nfc learners (weißgerber et al. 2018). in other words, learners low in nfc do not expend more cognitive resources on learning than necessary. however, if learners low in nfc are explicitly instructed to use elaborate learning strategies such as self-generation, they should benefit more from such strategies than learners high in nfc, because self-generation stimulates cognitive processes that learners low in nfc usually tend not to engage in spontaneously (mcdaniel and butler 2010). learners high in nfc, however, should readily engage in effortful cognitive processing of learning materials even when this is not explicitly required by the task (e.g., when just reading a text). for these learners, self-generation should contribute only weakly to their already elaborate processing. in sum, we assume that self-generation requires learners low in nfc to switch from less demanding shallow processing to elaborate processing, whereas learners high in nfc constantly use more elaborate processing strategies (see e.g., kardash and noel 2000). consequently, self-generation (compared to reading) should improve learning for learners low in nfc more strongly than for learners high in nfc. 2. the present study using a modified version of the classical word-generation paradigm by slamecka and graf (1978) and mcdaniel et al. (1988), the aim of the present study was to investigate the extent that effectiveness of self-generation in learning varies as a function of individual differences in nfc. learners were presented with word-pairs consisting of a context word and a target word. half of the presented word pairs consisted of incomplete target words which the learner needed to complete. (1) we aimed to replicate the generation effect, that is, we expected better recall for successfully generated target words than for read target words. (2) we expected to find a more pronounced generation effect for low nfc learners compared to high nfc learners. participants were also randomly assigned to one of two different learning settings. extant research found that the generation effect is more strongly pronounced in incidental than in intentional learning settings (bertsch et al. 2007). thus, to optimally investigate the expected interaction of learning condition (generation vs. reading) and individual differences in nfc, half of the participants were not informed about the learning test. however, learning in educational contexts is often intentional, for example when teachers and students purposefully use learning strategies to prepare for a learning test or to enhance the students’ learning outcome. thus, demonstrating that the effectiveness of self-generation differs as a function of individual differences in learners’ nfc not only in an incidental but also in an intentional learning setting would be highly relevant for adopting self-generation practices to applied educational contexts. hence, half of the participants were assigned to an intentional learning setting. 3. method 3.1 participants participants were 121 undergraduates, 19 grad students, and 3 non-students recruited at the campus of the university of kassel (germany). they came from varying disciplines with only 17 participants being undergraduate (n = 8) or graduate psychology students (n = 9). none of them surmised the exact purpose of our study. of the 143 participants in total (75 female, 68 male), 128 were native speakers of german. the age ranged from 17 to 55 with a mean age of 23.87 (sd = 4.73). all participants provided their written consent and were reimbursed with 5€ for their participation. 3.2 materials and procedure participants were tested individually or in groups of two to six in a laboratory. tasks and stimuli were presented on notebook computers. word-generation task. each participant was presented with 36 german word pairs in total consisting of a context word (e.g., kokon/ cocoon) and a semantically associated target word (e.g., raupe/ caterpillar). each target word belonged to one of six categories: fruit, body parts, clothing, animals, insects, and music instruments. six target words from each category were presented. half of the word pairs were complete (kokon – raupe), whereas the other half of the word pairs contained a fragmented target word (kokon – r_u_e) that participants were required to complete (varied within subjects). the number of slots indicated the number of missing letters in the generate condition. each word pair was presented in the middle of the notebook screen for a duration of 7 seconds with a 3 second interval between trials. participants were instructed to record the read and generated target words on a sheet of paper. half of them were instructed to memorize the target words for a later test (intentional learning setting, n = 71), and the other half was naive about the test (incidental learning setting, n = 72). each word pair occurred equally often in the generate and the read condition across participants. to ensure a balanced presentation of word pairs in both conditions, the 36 word pairs were divided into four blocks of nine word pairs. two blocks (18 word pairs) were presented in the generate and two blocks in the read condition for each participant. each block was paired equally often with each of the other three blocks in both conditions, which resulted in six stimulus lists. participants were randomly assigned to one of the six lists. presentation order of learning condition (generate-read vs. read-generate) was balanced across participants. word pairs were presented in randomized order within each learning condition. after the presentation of the final word pair, the experimenter collected the sheets of paper with the written target words. distractor task. after the word generation task, participants completed a computerized questionnaire on sleeping habits adopted from horne and ostberg (1975), which took participants about 5 minutes to complete. free and cued recall. following the distractor task, participants were asked to recall as many target words as possible within 5 minutes (free recall). after the free recall task, context words were presented for an additional 5 minutes in random order, and participants were asked to provide the target word to each context word (cued recall). need for cognition. participants completed the german 33-item need for cognition scale by bless et al. (1994). they read short statements (e.g., i really enjoy finding new solutions to problems; i prefer my life to be filled with puzzles that i must solve) and answered on a 7-point likert scale ranging from 1 (completely disagree) to 7 (completely agree). internal consistency of the nfc scale was high (cronbach’s  = .88). a mean nfc score was calculated for each participant (m = 4.84; sd = .68; min = 3.09; max = 6.27). additional measures. for the purpose unrelated to the present study, we administered the personal and global belief in a just world scale (dalbert 1999) and the academic self-concept scale (dickhäuser et al. 2002). control measures. participants in the incidental learning group were asked to indicate whether they had expected a test and prepared for it. in addition, all participants reported on a 7-point likert scale the extent that they found completing the fragmented target words difficult (ranging from 1 – not difficult at all to 7 – very difficult), the extent that they were motivated in identifying the fragmented target words in the learning phase and also in recalling the target words in the tests phase (ranging from 1 – not motivated at all to 7 – highly motivated). given that nfc is an indicator of an individual’s disposition to engage in effortful cognitive activities (e.g., cacioppo and petty 1982), nfc was expected to correlate with learners’ self-reported task-specific motivation but not with self-reported generation difficulty. finally, participants were asked to report any additional strategies they used (such as grouping target words into semantic categories or rehearsal) to memorize the target words during the learning phase. given extant findings that learners higher in nfc use more efficient learning strategies than learners lower in nfc (cazan and indreica 2014), we assumed that learners higher in nfc would use not only more strategies than learners lower in nfc but also more elaborate learning strategies. sociodemographic data were collected via an additional questionnaire. 4. results control measures. as expected, learners’ nfc correlated significantly with learners’ self-reported motivation for generating target words (r = .24, p = .004) and with their self-reported motivation to recall the target words in the test phase (r = .19, p = .02) but not with self-reported generation difficulty (r = -.14, p = .10). in the intentional learning group, 16 of the 71 learners reported the use of elaborate learning strategies such as grouping target words into categories (fruit, insects, body parts etc.). five learners reported the use of mnemonic strategies such as rehearsal or rereading of their recorded target words, and 50 learners reported not having used specific learning strategies. despite the unannounced test, 10 of the 72 learners in the incidental learning group reported to have noticed the semantic categories of the target words and that they tried to make use of the categories to generate target words, 2 rehearsed or reread the target words once, and 60 used no specific processing strategy. no differences in nfc were found among learners who reported the use of elaborate learning strategies, mnemonic strategies, and no strategies. forty-nine of the 72 incidental learners reported not having expected a test. there was no significant difference in the learners’ nfc between those who did and those who did not expect the unannounced learning test. although 23 incidental learners checked the box yes, i did expect a learning test in the final questionnaire, 19 of these participants reported in the open answer field on strategy use not to have used any learning strategy at all (i.e., they not even tried to memorize the word pairs) or they reported that they did not exactly prepare for a learning test, even if they surmised that the words would be important somehow later in the study. we will return to the four remaining incidental learners who have expected and prepared for the unannounced learning test when we report free and cued recall accuracy. accuracy of recording target words. in total, 90.1% of the generated and read target words have been recorded correctly in the learning phase. in the read condition, 96.7% have been recorded correctly. in the generate condition, 83.5% were generated successfully. no differences in generation accuracy were found between learning settings, and the accuracy did not decrease with decreasing nfc. although learners lower in nfc reported less motivation for generating target words than learners higher in nfc, they made no more errors generating target words than learners higher in nfc. 4.1 free and cued recall accuracy data analysis procedure. we estimated generalized linear mixed models (glmms) with a logit link function (dixon 2008) for free and cued recall accuracy as dependent variables. one word pair from list 3 (0.4% of the data) was excluded from the analysis, because of a technical error in displaying the word pair. the models were estimated and tested with the software packages lme4 (bates et al. 2014) and lmertest for r (kutznetsova et al. 2014). the number of possible iterations of the optimizer was increased to 100,000 to account for the models’ complexity. all significance tests were based on a type i error probability of .05. separate models were estimated for free recall accuracy and cued recall accuracy. to test for differences in learners’ free and cued recall performance as a function of learning condition, learning setting, and individual differences in nfc, learning condition and learning setting were included as contrast-coded predictor variables (learning condition: -1=read, 1 = generate; learning setting: -1 = incidental, 1 = intentional) and nfc as continuous grand-mean centered predictor variable in the glmms with free and cued recall accuracy as dependent variables. two-way and three-way interaction effects were estimated for all variables. in addition, the intercept and all main and interaction effects were estimated for target words that were recorded correctly in the learning phase (90.1% of the data). from a theoretical perspective, estimating learning outcomes for incorrectly recorded target words, which could not have been learned properly, would be pointless. moreover, from an applied educational perspective, teachers who want to use self-generation to improve student learning must ensure that their students are capable of generating the information from the planned lessons (mcdaniel and butler 2010) and that the critical information can be generated successfully. to this aim, accuracy of recording target words was included as another dummy-coded predictor variable with correctly recorded target words being the reference category (0 = correctly recorded target words, 1 = incorrectly recorded target words). finally, because participants and word pairs were sampled from a larger population, intercepts for persons and word pairs were allowed to vary randomly. descriptive statistics are provided in table 1. the parameter estimates for the fixed and random effects are provided in table 2. in the following sections, we focus on the main and interaction effects of learning condition, learning setting, and nfc for correctly recorded target words only. table 1 descriptive statistics for free recall and cued recall for generated and read target words and nfc during incidental and intentional learning (n=143) table 2 fixed effects and variance components in the glmm for free recall accuracy and cued recall accuracy free recall accuracy. the glmm analysis with free recall accuracy as dependent variable revealed a significant main effect of learning condition (β = 0.51, z = 14.34, p < .001) indicating that learners recalled significantly more generated than read target words. moreover, the analysis revealed a significant two-way-interaction of learning setting and nfc (β = 0.13, z = 2.20, p = .03). recall for target words increased significantly with increasing nfc when learning was intentional (β = 0.18, z = 2.45, p = .01) but not when learning was incidental (see figure 1). figure 1. estimated probability for accurately recalling target words (free recall) for incidental and intentional learning: simple slopes for need for cognition and differences between learning settings estimated at three different levels of need for cognition. in sum, we replicated the generation effect (hypothesis 1), although we found no significant differences in the magnitude of the generation effect as a function of individual differences in nfc (hypothesis 2). our findings additionally indicate that learners higher in nfc prepared better for the announced learning test than learners lower in nfc. cued recall accuracy. the glmm analysis with cued recall accuracy as dependent variable revealed significant main effects of learning condition (β = 0.85, z = 19.69, p <.001) and nfc (β = 0.22, z = 2.44, p = .01). both main effects were further qualified by a significant two-way interaction of learning condition and nfc (β = -0.09, z = -2.03, p = .04). although all learners recalled generated target words significantly better than read target words (generation effect), the simple effect of learning condition was more strongly pronounced for learners lower in nfc (nfc minus 1 sd: β = 0.93, z=15.11, p <.001) than for learners higher in nfc (nfc plus 1 sd: β = 0.76, z = 12.61, p <.001). in other words, learners lower in nfc benefited more from generating target words than learners higher in nfc (see figure 2). moreover, the simple slope for nfc was significant in the read condition (β = 0.31, z = 3.25, p = .001) but not in the generate condition. that is, improved test performance with increasing nfc was found only in the read condition, whereas individual differences in nfc did not affect cued recall accuracy in the generate condition (figure 2). figure 2. estimated probability for accurately recalling generated and read target words (cued recall): simple slopes for need for cognition and differences between learning conditions estimated at three different levels of need for cognition. in sum, these findings are consistent with hypotheses 1 and 2. we replicated the generation effect and learners lower in nfc benefited more from the generation effect than learners higher in nfc. moreover, these findings revealed no differences between learning settings. finally, to test if our results were driven or contorted by test expectancy in the incidental learning setting, the four learners who reported to have expected and prepared for a learning test by memorizing the presented word-pairs were treated as intentional learners in additional analyses of free and cued recall accuracy (intentional learning setting: n = 75, incidental learning: n = 68). this did not change the results in terms of levels of significance. 5. discussion the aim of the present study was to investigate the moderating effect of individual differences in learners’ nfc on the generation effect. we expected (1) to replicate the generation effect and (2) to find a more strongly pronounced generation effect for learners lower in nfc than for learners higher in nfc. the results of the present study supported hypothesis 1 and corroborated hypothesis 2 for cued recall accuracy as dependent variable. learners recalled more generated target words than read target words with both free and cued recall accuracy as dependent variables. these results are consistent with extant empirical research on the beneficial effects of generation (for an overview, see bertsch et al. 2007). however, our study extends the existing literature on the generation effect, because our findings showed that individual differences in learners’ nfc moderate the magnitude of the generation effect. as expected, the analysis of cued recall accuracy revealed that the generation effect was more strongly pronounced for learners lower in nfc than for learners higher in nfc. that is, learners lower in nfc benefited significantly more from generating target words than learners higher in nfc. this finding is consistent with the idea that desirable difficulties such as self-generation are beneficial when they stimulate cognitive processes that learners tend not to engage in spontaneously (e.g., see mcdaniel and butler 2010). our finding that learners lower in nfc recalled less read target words than learners higher in nfc indicates that learners lower in nfc are cognitive misers who processed the read word pairs shallower than learners higher in nfc. however, the same learners recalled generated target words as accurately as the learners higher in nfc. we assume that especially the learners lower in nfc benefited from the generation task (as indicated by a more strongly pronounced generation effect compared to learners higher in nfc), because self-generation stimulated more elaborate cognitive processing of the word pairs. in other words, self-generation required them to switch from shallow cognitive processing to more elaborate processing (kardash and noel 2000). however, with increasing nfc, learners increasingly engaged in elaborate cognitive processing of the learning material (even without explicit instruction) as indicated by increasingly improved recall of read target words. for these learners, self-generation becomes increasingly redundant to the extent that they already show elaborate processing independent of the specific task. consequently, learners higher in nfc benefit less from the generation task than learners lower in nfc. it is noteworthy that in the generation condition, there was no main effect of nfc. learners lower in nfc recalled as much target words as learners higher in nfc. in other words, self-generation helped the learners lower in nfc to close the learning gap on learners higher in nfc. the finding that cued recall accuracy for read target words improved with increasing nfc is consistent with extant empirical findings showing that learners high in nfc recall information better than learners low in nfc (e.g., cacioppo et al. 1983; heijne-penninga et al. 2010; sadowski and gülgöz 1996). however, our findings suggest that this disadvantage of learners lower in nfc can be balanced by using generative tasks that stimulate elaborate cognitive processing. in contrast to cued recall, hypothesis 2 was not supported by the free recall data. although we replicated the generation effect for free recall accuracy as dependent variable, individual differences in nfc were not found to moderate the magnitude of the generation effect. plausible explanations are the different task requirements of free and cued recall and how they match the kind of processing in the learning phase (e.g., see the contextual framework, mcdaniel and butler 2010 or transfer-appropriate processing, morris et al. 1977). to successfully generate a target word in the learning phase, learners were required to establish a mental link between the context word and the target word. this mental link could then be used as a scaffold to retrieve the generated target words from memory when context words were provided as cues in the cued recall task. in free recall, however, the mental links established during the generation task could not serve as scaffolds for target word retrieval without providing context words (note, in this context, the elaborate processing of the target word still improves learning in the generate compared to the read condition). this idea is supported by the finding that learners recalled about twice as much target words in cued recall compared to free recall (see descriptive statistics in table 1). note that learners at all nfc levels should have established elaborate mental links between a context word and the target word in the generate condition (because it was required by the task). in contrast, only learners high in nfc should have established such links between context and target words in the read condition. this interaction between learning condition and nfc, however, can only be seen when the criterial task draws upon these established mental links, that is, in cued but not in free recall. the idea of self-generation as scaffold to enhance memory for the target word by constructing a mental bridge between context and target word might suggest that self-generation is kind of an epistemic action („an external [i.e., not solely mental] action that an agent performs to change his or her own computational state” in contrast to pragmatic actions “whose primary function is to bring the agent closer to his or her physical goal”, kirsh and maglio 1994, pp. 514–515; see also kirsh 2006). from this perspective, one could argue that self-generation is an external action that alters the environment (here the learning material) and, thereby, adds to problem solving (here target-word memory). as was demonstrated for epistemic actions (kirsh and maglio; maglio and kirsh, 1996), it is only in hindsight, that the benefit of the additional and putatively unnecessary generation task becomes evident. in contrast to epistemic actions, however, desirable difficulties do not reduce working memory load, the number of cognitive steps involved in processing, or the probability of processing errors (see kirsh and maglio). instead, desirable difficulties are characterized by increasing cognitive effort in a way that is beneficial to learning. they are, by definition, no reduction of complexity. in this way, self-generation is clearly distinct from epistemic actions. the participants in our study were randomly assigned to one of two learning settings – an incidental and an intentional learning setting. learning in educational contexts is often intentional, for example when teachers and students purposefully use learning strategies to prepare for a test or to enhance the students’ learning outcome. hence, demonstrating that the two-way interaction of learning condition and nfc shows in an intentional learning setting would further corroborate the practical relevance of our findings. as expected, the finding that the generation effect was more strongly pronounced for learners lower in nfc than for learners higher in nfc (cued recall) did not differ between learning settings. neither the main effect of learning setting nor the interaction effects of learning setting with learning condition and nfc became significant. this result suggests that the compensatory effect of self-generation on target word memory of learners lower in nfc occurs independently of the learning setting in cued recall. in free recall, we found that learners higher in nfc recalled more target words than learners lower in nfc when learning was intentional. this finding suggests that learners higher in nfc voluntarily invested more cognitive resources on preparation for a test than learners lower in nfc (even when test performance had no actual consequences for their studies). this interpretation is consistent with extant studies demonstrating that learners higher in nfc are more willing to tackle difficult tasks than learners lower in nfc (see et al. 2009; weißgerber et al. 2018). we assume that (in addition to establishing mental links between context and target words) they might have tried to explicitly memorize the target words to be prepared for later recall. since free recall (in contrast to cued recall) assesses context-free retrieval of target words, deeper processing of target words in the learning phase led to increased recall accuracy for learners higher in nfc independent of learning condition. in sum, the different findings for free and cued recall obtained in our study can be explained by different task requirements of both criterial tasks and how each of them matched the kind of processing in the learning phase. finally, extant studies showed that learners higher in nfc use more efficient learning strategies than leaners lower in nfc (cazan and indreica 2014). hence, we assumed that learners higher in nfc would use more elaborate learning strategies than learners lower in nfc. however, participants self-reported use of elaborate, less elaborate, and no additional learning strategies did not vary as a function of individual differences in nfc. a likely explanation for this finding is that 7 seconds of word-pair presentation and 3 seconds of inter-stimulus interval are too short a time for most participants to properly administer additional learning strategies, let alone elaborate ones. this might be different for more complex learning material such as texts or algebraic word problems and remains to be investigated in future research. the findings reported in this study should be interpreted with possible limitations in mind. in everyday life, learners usually deal with learning material that is much more complex than isolated word pairs. moreover, most of the time, learners are unaware of the kind of criterial task for which to prepare, and when preparing for a test or exam, retention intervals are usually longer (several days or weeks) than just a few minutes as in most laboratory studies on the generation effect. despite these limitations, the findings of the present study have important theoretical and practical implications. the finding that learners recalled three times as much target words in the generate condition as they recalled in the read condition (see descriptive statistics, table 1) strongly suggests that self-generation might be a useful supplement to commonly used learning strategies in education. the findings of the present study, however, also indicate that educators should be prepared to find individual differences in the effectiveness of generative activities depending on learners’ characteristics. we demonstrated for the first time that the generation effect differs as a function of individual differences in nfc. for those high in nfc, self-generation contributes comparatively little to learning. it can, however, be highly beneficial for learners low in nfc (both in incidental and intentional learning settings). this suggests that self-generation can be used to systematically improve learning for those who are likely to fall behind their peers due to low engagement in effortful cognitive processing. since nfc is easily and quickly accessed in single learning settings as well as in classrooms, learners low in nfc and, thus, in special need for cognitively demanding learning instructions can (and should) be effortlessly identified. another important practical implication of our study is that the compensatory effect of self-generation becomes visible only when the generation task matches the requirements of the criterial task. when adopting self-generation as a learning strategy in educational contexts (e.g., school classes, educational books, or computerized learning environments), teachers, authors, and programmers should ensure that the generation task encourages cognitive processes relevant to the test, exam, or task for which the learners prepare (mcdaniel and butler 2010). the results of the present study raise some interesting future research questions. first, future research needs to replicate and extend the reported findings with more complex and naturalistic learning materials (e.g., math problems or expository texts), in more naturalistic settings (such as classrooms or in collaborative action learning), with different types of generation and criterial tasks and longer retention intervals. moreover, when using more complex learning material such as texts, other learner characteristics such as working memory, reading ability, creativity, learning goals, and openness to ideas should be considered alongside nfc to account for mutual variance that these components might share with nfc and to account for possible moderating effects to further optimize the use of self-generation in everyday learning settings. finally, the interaction of self-generation and nfc raises not only the question which further learner characteristics might affect the effectiveness of self-generation, but also which other desirable difficulties might be affected by individual differences in nfc. a possible candidate to look at might be the testing effect. access to and alteration of knowledge structures during learning, relearning, and retesting might be differential for learners high in nfc and those low in nfc. 6. conclusion the present study replicated the generation effect with a version of the classical word-generation paradigm by slamecka and graf (1978) and mcdaniel et al. (1988). learners recalled generated target words better than read target words. moreover, our study demonstrated for the first time that individual differences in learners’ nfc moderate the effectiveness of self-generation during learning. learners lower in nfc benefited significantly more from generating target words than learners higher in nfc when retrieval cues were provided in the test phase. this finding corroborates mcdaniel and butler’s (2010) explanation that desirable difficulties such as self-generation are only beneficial when they stimulate cognitive processes that learners tend not to engage in spontaneously. we assume that learners higher in nfc voluntarily engage in elaborate cognitive information processing even without explicit instruction, whereas learners lower in nfc need a cognitively demanding task that requires them to switch from shallow to elaborate cognitive processing to improve learning. the reported findings suggest that using self-generation in educational contexts is beneficial for learners at all levels of nfc, but it could be systematically used to improve learning for learners with a weak disposition to engage in cognitively demanding learning processes. keypoints desirable difficulty generation effect incidental learning intentional learning need for cognition acknowledgments the research presented in this article was supported by the federal state of hessen and its loewe research initiative desirable difficulties in learning (loewe: landes-offensive zur entwicklung wissenschaftlich-ökonomischer exzellenz [state offensive for the development of scientific and economic excellence]). we would like to thank our student assistants for assisting in data collection and coding. researchers who are interested in the stimulus material are invited to send an e-mail to the first or the second author. references bates, d., maechler, m., bolker, b., walker, s., christensen, r. h. b., & sigmann, h. (2014). lme4: linear mixed-effects models using eigen and s4 [software]. r-package version 1.1-6. retrieved may 1, 2014 from: http://cran.r-project.org/package=lme4 bertrams, a., & dickhäuser, o. (2009). high-school students' need for cognition, self-control capacity, and school achievement: testing a mediation hypothesis. learning and individual differences, 19(1), 135–138. doi:10.1016/j.lindif.2008.06.005 bertsch, s., pesta, b. j., wiscott, r., & mcdaniel, m. a. (2007). the generation effect: a meta-analytic review. memory & cognition, 35(2), 201–210. doi:10.3758/bf03193441 bjork, r. a. (1994). memory and metamemory considerations in the training of human beings. in j. metcalfe & a. p. shimamura (eds.), metacognition: knowing about knowing (pp.185–205). cambridge: mit press. bjork, e. l., & bjork, r. a. (2011). making things hard on yourself, but in a good way: creating desirable difficulties to enhance learning. in m. a. gernsbacher, r. w. pew, l. m. hough, & j. r. pomerantz (eds.), psychology and the real world: essays illustrating fundamental contributions to society (pp. 56–64). new york: worth publishers. bless, h., wänke, m., bohner, g., fellhauer, r. f., & schwarz, n. (1994). need for cognition: eine skala zur erfassung von engagement und freude bei denkaufgaben [need for cognition: a scale measuring engagement and happiness in cognitive tasks]. zeitschrift für sozialpsychologie, 25, 147–154. cacioppo, j. t., & petty, r. e. (1982). the need for cognition. journal of personality and social psychology, 42(1), 116–131. doi:10.1037/0022-3514.42.1.116 cacioppo, j. t., petty, r. e., & kao, c. f. (1984). the efficient assessment of need for cognition. journal of personality assessment, 48(3), 306–307. doi:10.1207/s15327752jpa4803_13 cacioppo, j t., petty, r. e., kao, c. f., & rodriguez, r. (1986). central and peripheral routes to persuasion: an individual difference perspective. journal of personality and social psychology, 51(5), 1032–1043. doi:10.1037/0022-3514.51.5.1032 cacioppo, j. t., petty, r. e., & morris, k. j. (1983). effects of need for cognition on message evaluation, recall, and persuasion. journal of personality and social psychology, 45(4), 805–818. doi:10.1037/0022-3514.45.4.805 cazan, a.-m., & indreica, s. e. (2014). need for cognition and approaches to learning among university students. procedia social and behavioral sciences, 127, 134–138. doi:10.1016/j.sbspro.2014.03.227 cepeda, n. j., pashler, h., vul, e., wixted, j. t., & rohrer, d. (2006). distributed practice in verbal recall tasks: a review and quantitative synthesis. psychological bulletin, 132(3), 354–380. doi:10.1037/0033-2909.132.3.354 coutinho, s. a. (2006). the relationship between the need for cognition, metacognition, and intellectual task performance. educational research and reviews, 1(5), 162–164. coutinho, s. a., wiemer-hastings, k., skowronski, j. j., & britt, m. a. (2005). metacognition, need for cognition and use of explanations during ongoing learning and problem solving. learning and individual differences, 15(4), 321–337. doi:10.1016/j.lindif.2005.06.001 dalbert, c. (1999). the world is more just for me than generally: about the personal belief in a just world scale’s validity. social justice research, 12(2), 79–98. doi:10.1023/a:1022091609047 dickhäuser, o., schöne, c., spinath, b., & stiensmeier-pelster, j. (2002). die skalen zum akademischen selbstkonzept: konstruktion und überprüfung eines neuen instrumentes [the academic self-concept scales: construction and evaluation of a new instrument].zeitschrift für differentielle und diagnostische psychologie, 23(4), 393–405. doi:10.1024//0170-1789.23.4.393 dixon, p. (2008). models of accuracy in repeated-measures designs. journal of memory and language, 59(4), 447–456. doi:10.1016/j.jml.2007.11.004 doctorow, m., wittrock, m. c., & marks, c. (1978). generative processes in reading comprehension. journal of educational psychology, 70(2), 109–118. doi:10.1037/0022-0663.70.2.109 dunlosky, j., rawson, k. a., marsh, e. j., nathan, m. j., & willingham, d. t. (2013). improving students’ learning with effective learning techniques: promising directions from cognitive and educational psychology. psychological science in the public interest, 14(1), 4–58. doi:10.1177/1529100612453266 einstein, g. o., mcdaniel, m. a., owen, p. d., & coté, n. c. (1990). encoding and recall of texts: the importance of material appropriate processing. journal of memory and language, 29(5), 566–581. doi: 10.1016/0749-596x(90)90052-2 fiedler, k., lachnit, h., fay, d., & krug, c. (1992). mobilization of cognitive resources and the generation effect.the quarterly journal of experimental psychology, section a, 45(1), 149–171. doi:10.1080/14640749208401320 fleischhauer, m., miller, r., enge, s., & albrecht, t. (2014). need for cognition relates to low-level visual performance in a metacontrast masking paradigm. journal of research in personality, 48, 45–50. doi:10.1016/j.jrp.2013.09.007 gardiner, j. m., & rowley, j. m. c. (1984). a generation effect with numbers rather than words. memory & cognition, 12(5), 443–445. doi:10.3758/bf03198305 graf, p. (1980). two consequences of generating: increased inter and intraword organization of sentences. journal of verbal learning and verbal behavior, 19(3), 316–327. doi:10.1016/s0022-5371(80)90248-0 graf, p. (1981). reading and generating normal and transformed sentences.canadian journal of psychology/revue canadienne de psychologie, 35(4), 293–308. doi:10.1037/h0081193 grass, j., strobel, a., & strobel, a. (2017). cognitive investments in academic success: the role of need for cognition at university. frontiers in psychology, 8, 790. doi:10.3389/fpsyg.2017.00790 heijne-penninga, m., kuks, j. b. m., hofman, w. h. a., & cohen-schotanus, j. (2010). influences of deep learning, need for cognition and preparation time on openand closed-book test performance. medical education, 44(9), 884–891. doi:10.1111/j.1365-2923.2010.03732.x horne, j. a., & ostberg, o. (1976). a self-assessment questionnaire to determine morningness-eveningness in human circadian rhythms. international journal of chronobiology, 4, 97–110. jebb, a. t., saef, r., parrigon, s., & woo, s. e. (2016). the need for cognition: key concepts, assessment, and role in educational outcomes. in a. lipnevich, f. preckel, & r. d. roberts (eds.), psychosocial skills and school systems in the twenty-first century: theory, research, and applications. the springer series on human exceptionality (pp.115¬¬–132). new york: springer. doi:10.1007/978-3-319-28606-8_5 kardash, c. a. m., & noel, l. k. (2000). how organizational signals, need for cognition, and verbal ability affect text recall and recognition. contemporary educational psychology, 25(3), 317–331. doi:10.1006/ceps.1999.1011 kirsh, d. (2006). distributed cognition. a methodological note. pragmatics and cognition, 14(2), 249–262). doi:10.1075/pc.14.2.06kir kirsh, d., & maglio, p. (1994). on distinguishing epistemic from pragmatic action. cognitive science, 18, 513–549. doi:10.1016/0364-0213(94)90007-8 kuznetsova, a., brockhoff, p. b., & christensen, r. h. b. (2014). lmertest: tests for random and fixed effects for linear mixed effect models (lmer objects of lme4 package) . r-package version 2.0-6. retrieved in june 2014 from: http://cran.r-project.org/web/packages/lmertest/index.html leone, c., & dalton, c. h. (1988). some effects of the need for cognition on course grades. perceptual and motor skills, 67(1), 175–178. doi:10.2466/pms.1988.67.1.175 luong, c., strobel, a., wollschläger, r., greiff, s., vainikainen, m.-p., & preckel, f. (2017). need for cognition in children and adolescents: behavioral correlates and relations to academic achievement and potential. learning and individual differences, 53, 103–113. doi:10.1016/j.lindif.2016.10.019 maglio, p. p. & kirsh, d. (1996). epistemic action increases with skill. in g. w. cottrell (ed.), proceedings of the eighteenth annual conference of the cognitive science society (pp. 391–396). mahwah, nj: erlbaum. mcdaniel, m. a., & butler, a. c. (2010). a contextual framework for understanding when difficulties are desirable. in a. s. benjamin (ed.), successful remembering and successful forgetting: a festschrift in honor of robert a. bjork (pp. 175-198). new york: taylor and francis. doi:10.4324/9780203842539 mcdaniel, m. a., & einstein, g. o. (1989). material-appropriate processing: a contextualist approach to reading and studying strategies. educational psychology review, 1(2), 113–145. doi:10.1007/bf01326639 mcdaniel, m. a., & einstein, g. o. (2005). material appropriate difficulty: a framework for determining when difficulty is desirable for improving learning. in a. f. healy (ed.), experimental cognitive psychology and its applications (pp. 73–85). washington, dc: american psychological association. doi:10.1037/10895-006 mcdaniel, m. a., einstein, g. o., dunay, p. k., & cobb, r. e. (1986). encoding difficulty and memory: toward a unifying theory. journal of memory and language, 25(6), 645–656. doi:10.1016/0749-596x(86)90041-0 mcdaniel, m. a., hines, r. j., & guynn, m. j. (2002). when text difficulty benefits less-skilled readers. journal of memory and language, 46(3), 544–561. doi:10.1006/jmla.2001.2819 mcdaniel, m. a., & waddill, p. j. (1990). generation effects for context words: implications for item-specific and multifactor theories. journal of memory and language, 29(2), 201–211. doi:10.1016/0749-596x(90)90072-8 mcdaniel, m. a., waddill, p. j., & einstein, g. o. (1988). a contextual account of the generation effect: a three-factor theory. journal of memory and language, 27(5), 521–536. doi:10.1016/0749-596x(88)90023-x morris, c. d., bransford, j. d., & franks, j. j. (1977). levels of processing versus transfer appropriate processing. journal of verbal learning and verbal behavior, 16(5), 519–533. doi:10.1016/s0022-5371(77)80016-9 nair, k. u., & ramnarayan, s. (2000). individual differences in need for cognition and complex problem solving. journal of research in personality, 34(3), 305–328. doi:10.1006/jrpe.1999.2274 reinhard, m.-a (2010). need for cognition and the process of lie detection. journal of experimental social psychology, 46(6), 961–971. doi:10.1016/j.jesp.2010.06.002 reinhard, m.-a., & dickhäuser, o. (2009). need for cognition, task difficulty, and the formation of performance expectancies. journal of personality and social psychology, 96(5), 1062–1076. doi:10.1037/a0014927 richardson, m., abraham, c., & bond, r. (2012). psychological correlates of university students’ academic performance: a systematic review and meta-analysis. psychological bulletin, 138(2), 353–387. doi:10.1037/a0026838 roediger, h. l., & karpicke, j. d. (2006). test-enhanced learning: taking memory tests improves long-term retention. psychological science, 17(3), 249–255. doi:10.1111/j.1467-9280.2006.01693.x sadowski, c. j., & gülgöz, s. (1996). elaborative processing mediates the relationship between need for cognition and academic performance. the journal of psychology, 130(3), 303–307. doi:10.1080/00223980.1996.9915011 schweickert, r., mcdaniel, m. a., & riegler, g. (1994). effects of generation on immediate memory span and delayed unexpected free recall.the quarterly journal of experimental psychology, section a, 47(3), 781–804. doi:10.1080/14640749408401137 see, y. h. m., petty, r. e., & evans, l. m. (2009). the impact of perceived message complexity and need for cognition on information processing and attitudes. journal of research in personality, 43(5), 880–889. doi:10.1016/j.jrp.2009.04.006 slamecka, n. j., & fevreiski, j. (1983). the generation effect when generation fails. journal of verbal learning and verbal behavior, 22(2), 153–163. doi:10.1016/s0022-5371(83)90112-3 slamecka, n. j., & graf, p. (1978). the generation effect: delineation of a phenomenon.journal of experimental psychology: human learning and memory, 4(6), 592–604. doi:10.1037/0278-7393.4.6.592 verplanken, b., hazenberg, p. t., & palenéwen, g. r. (1992). need for cognition and external information search effort. journal of research in personality, 26(2), 128–136. doi:10.1016/0092-6566(92)90049-a von stumm, s. & ackerman, p. l. (2013). investment and intellect: a review and meta-analysis. psychological bulletin, 139(4), 841–869. doi:10.1037/a0030746 weißgerber, c. s., reinhard, m.-a., & schindler, s. (2018). learning the hard way: need for cognition influences attitudes towards and self-reported use of desirable learning difficulties. educational psychology, 38(2), 176–202. doi:10.1080/01443410.2017.1387644 frontline learning research vol. 5 no. 3 special issue (2017) 66 80 issn 2295-3159 the flash-preview moving window paradigm: unpacking visual expertise one glimpse at a time damien litchfield1 & tim donovan2 1 edge hill university, uk 2 university of cumbria, uk article received 22 august / revised 15 december / accepted 23 march / available online 14 july abstract how we make sense of what we see and where best to look is shaped by our experience, our current task goals and how we first perceive our environment. an established way of demonstrating these factors work together is to study how eye movement patterns change as a function of expertise and to observe how experts can solve complex tasks after only very brief glances at a domain-specific image. the primary focus of this paper is to introduce an innovative gaze-contingent method called the ‘flash-preview moving window’ (fpmw) paradigm (castelhano & henderson, 2007), which was recently developed to understand our shared expertise in scene perception and how our first glimpse of a scene is used to guide our eye movement behaviour. in keeping with this special issue on visual expertise and medicine, this paper will highlight how the fpmw paradigm has the potential to resolve long-standing theoretical issues as to how, right from the very first glance, experts are able to process domain-specific images and guide their eye movements better than novices. since fpmw is a gaze-contingent eye-tracking method, the paper will first outline the current methodological and theoretical frontier, and how the fpmw paradigm bridges established methods used to investigate visual expertise. the paper will discuss a recent example in which the fpmw was employed to investigate medical image perception expertise for the first time (litchfield & donovan, 2016), and by discussing the insights and challenges this method offers, this should ultimately deepen our understanding of visual expertise. keywords: flash-preview moving window; eye movements; medical image perception; visual expertise; eye-tracking litchfield et donovan | f l r 67 1. introduction from the moment we open our eyes we see a rich visual world. we are unable to process all of this incoming information and so we must move our eyes several times a second to look at and process different aspects of our environment. how we make sense of what we see and where best to look is shaped by our experience, our current task goals and how we first perceive our environment (buswell, 1935; henderson, 2007; rayner, 2009; yarbus, 1967). the common purpose of most eye tracking studies is to explore how eye movement behaviour is regulated and to see how eye movement behaviour during specific tasks relates to underlying visual and cognitive processes (just & carpenter, 1984). when searching for an object in a newly presented image, the first fixation quickly (i.e., within 40-100ms) encapsulates the initial “gist” of the scene and receives some pre-attentive processing of basic features (rayner, smith, malcom, & henderson, 2009; võ & henderson, 2010). at this point visual properties such as colour, contour distribution and spatial frequency are likely to be processed (henderson & hollingworth, 1999; oliva, & schyns, 1997; schyns & oliva, 1994), with scene context and semantic information accessible depending on the duration of the first glimpse of the scene (fei-fei, iyer, koch, & perona, 2007; potter, 1976). it is within this initial glimpse of the scene that subsequent eye movements are guided (castelhano & henderson, 2007) with parafoveal and peripheral vision playing an important role in the early comprehension of the gist of a scene and in the detection of targets (henderson, pollatsek, & rayner, 1989). the visual system is able to integrate this visual input from sensory information with top-down information, which allows us to identify and locate any taskrelevant item(s) within the receptive field for saccadic targeting. however, the processes that underlie this integration of information are still heavily debated (see cohen, dennett, & kanwisher, 2016; tatler, 2009; torralba, oliva, castelhano, & henderson, 2006; wolfe evans, võ, & greene, 2011; zelinsky & schmidt, 2009). understanding how these visual processes become optimised with experience not only sheds light on these cutting edge issues, but also provides new ways of studying the development (and potential enhancement) of expertise (donovan & litchfield, 2013; litchfield & donovan, 2016). medical image perception is a domain of visual expertise that has long recognised the importance of processing the initial glimpse of an image (e.g., kundel & nodine, 1975), and given the rapid developments in scene perception research, it is fundamental that these respective research fields are reconciled as both inform each other (donovan & litchfield, 2013; drew, evans võ, jacobson, & wolfe, 2013a). a hallmark of expertise is that experts make better and faster decisions than novices (chase & simon, 1973; de groot, 1946/1965; gobet, 2015) and experts in medical image perception are no exception (for an excellent review of eye movements and visual expertise in medicine see reingold & sheridan, 2011; see also this special issue: fox & faulkner-jones, 2017; szulewski, kelton, & howes, 2017). moreover, examining the expertise-related differences in eye movement patterns reveals the types of visual processing and cognitive strategies that may underlie such expert performance. as mentioned above, our visual system integrates top-down information (e.g., knowledge, expectations) with bottom-up information (visual processing of the incoming image) to make sense of what we see and where to look. accordingly, experts in a specific domain can draw on their acquired knowledge to be more selective in the information they use to make decisions and are less likely to look at conspicuous but non-informative areas of a scene (for a recent meta-analysis, see gegenfurtner, lehtinan, & säljö, 2011). indeed, experts in medical image perception are much more selective than novices in deciding where to look (donovan & litchfield, 2013; krupinski, 1996; kundel & la follette, 1972; manning, ethell, & crawford, 2003; manning, ethell, & donovan, 2004; manning, ethell, donovan, & crawford, 2006), and exhibit efficient scanpaths that allow abnormalities to be detected quickly (krupinski, 1996; kundel, nodine, conant, & weinstein, 2007; kundel, nodine, krupinski, & mello-thoms, 2008), whilst minimizing unnecessary fixations (manning et al., 2006). consequently, this allows experts to inspect medical images faster than less experienced observers without compromising on decision accuracy (nodine, mello-thoms, kundel, & weinstein, 2002), and there is much to learn from viewing how experts search these complex images (litchfield, ball, donovan, manning, & crawford, 2010). these expert/novice differences can be interpreted by the prominent global-focal search model (nodine & kundel, 1987) or its recent formulation, the “holistic model” (kundel et al., 2007; see also the two-stage detection model by swensson, 1980). the holistic model proposes that within the first glimpse, litchfield et donovan | f l r 68 expert observers globally processes the medical image and subsequently make efficient search-related eye movements to potentially abnormal areas to support diagnostic decision-making. this means that prior to foveal search, experts process the low-level information relating to the present image and compare this with their extensive experience of previously viewed normal and abnormal medical images. by drawing on their knowledge of domain-specific visual representations (schema), experts rapidly recognise and coordinate their search for abnormalities and deploy search strategies based on the global information encapsulated within the initial “gist” of image viewing. if nothing is perturbed from the initial global impression, then subsequent search and discovery processing is engaged, which is reliant on individual feature search. there is substantial converging evidence to support the holistic model, and specifically that expert observers can rapidly process the initial glimpse of the image. the link between globally processing the initial glimpse and diagnostic performance was established by kundel and nodine’s (1975) tachistoscopic experiments, in which they found that even when images were presented for just 200ms, experts could still correctly detect 70% of abnormal images compared to 97% with no time constraints (see also carmody, nodine & kundel, 1981; evans, georgian-smith, tambouret, birdwell, & wolfe, 2013; mugglestone, gale, cowley & wilson, 1995; oestmann et al., 1988). since these ‘flash’ studies presented images so quickly that they prevented eye movements and yet performance was still above chance, this provided strong evidence that rapid processing of medical images must be contributing to diagnostic performance, aside from what can be gained by subsequent search and discovery processing. moreover this led kundel and nodine to argue that “visual search begins with a global response that establishes content, detects gross deviations from normal, and organizes subsequent foveal checking fixations” (kundel & nodine, 1975, p. 527), and this hypothesis has dominated medical image perception research in the decades since its inception. indeed, kundel et al. (2007) and nodine & mello-thoms (2010) proposed that the ability to optimise the processing of the initial glimpse is a hallmark of expertise as they suggested that visual search in medical image perception begins as: search & detect–recognize–decide but with experience, this develops into: recognize & detect–search–decide. however, it is important to note that in several of these ‘flash’ studies, such above-chance performance was only observed when they contained highly conspicuous nodules as performance was much worse for subtle nodules. for example, carmody et al. (1981) reported that detection of abnormal images did not improve beyond 180ms, but whereas abnormal images containing high and medium visibility nodules were detected often (100% and 83% respectively), abnormal images containing low visibility nodules were only correctly detected 53% of the time. similarly, oestmann et al. (1988) found that images containing subtle or obvious cancers were correctly detected 30% and 70% respectively when shown for just 250ms, whereas with unlimited viewing time detection increased to 74% and 98%, respectively. indeed, in free search conditions carmody et al., (1981) established that there was a relationship between nodule visibility and the frequency of comparative scans: the poorer the visibility of the nodule, the more likely that eye movements would be directed alternately to normal and abnormal regions to help distinguish pathology from normality. these findings are reflected in the holistic model, in that if the observer does not recognise perturbations in the image then search and discovery processing is undertaken to support diagnostic decisionmaking. but if the strength of the target signal influences detection so much, how do we know that global processing is involved and that it is not just rapid serial feature search instead? the answer lies with experiments that deliberately disrupted global or ‘holistic’ processing. for example, presenting segmented images (carmody, nodine, & kundel, 1980) or rotated images (oestmann, greene, bourgouin, linetsky, & llewellyn, 1993) impaired diagnostic performance compared to ‘global search’, where images were presented normally. in addition, these impairments in detection were independent of target conspicuity as they were found in images containing either obvious or subtle cancers (oestmann et al., 1993). what is problematic, however, is that since these disruption studies and the ‘flash’ studies had very limited sample size (3 to 4 observers), and predominately only tested experienced observers (common problems with many medical imaging studies), they actually provide very little direct evidence that experts are the only ones reliant on global processing, or whether in fact this disruption in global processing also impairs less experienced observers. this is an important distinction to make in understanding the litchfield et donovan | f l r 69 development of visual expertise. for example, it is one thing to infer that experts are able to globally process a domain-specific image and that this helps explain their expert performance, it is another to acknowledge that all observers potentially have access to global processing at some point (for example, in scene perception), and it is just that experts may have fine-tuned an existing system to a fit a particular domain, rather than created a new system entirely from scratch. this has both theoretical and practical implications as to how global processing is used, modified, and if vital to diagnostic performance, how it may be selectively trained to enhance performance. that said, a broader overview of the visual expertise literature shows there are multiple domains where increasing expertise enables observers to shift from local feature processing to holistic processing and thereby take into account the overall configuration of incoming information (gauthier, tarr, & bub, 2010), whether this is expertise in processing particular objects such as faces (gauthier & tarr, 1997) dogs (diamond & carey, 1986), cars (curby, glazek & gauthier, 2009), fingerprints (busey & vanderkolk, 2005), or potentially whole scenes (kelley, chun, & chua, 2003; werner & thies, 2000). aside from these disruption and flash studies, expert/novice eye movement studies in medical perception have been broadly supportive of the holistic model. a recurring finding of expert/novice studies is that not only can experts correctly identify more abnormalities than novices, the time taken to first fixate abnormalities (search latency) is faster for experts compared to novices (donovan & litchfield, 2013; krupinski, 1996; kundel, et al., 2007; kundel, nodine, krupinski, & mello-thoms 2008; nodine & mellothoms, 2010; reingold & sheridan, 2011). related to this, experts tend to make longer saccades than novices and these larger eye movements across the image help experts reach targets faster (krupinski, 1996; kundel, et al., 2007; kundel et al., 2008; manning et al., 2006). according to the holistic model the initial global analysis of the image is thought to initiate and guide search and so this efficiency in expert search behaviour is attributed to experts exploiting global processing that less experienced observers cannot (kundel et al., 2007). although low sample size is often a methodological constraint of these expertise studies, efforts have been made to combine time-to-first fixation data from a number of small sample mammography studies and this shows that faster search times are associated with expert performance (kundel et al., 2008). kundel et al. (2008) performed a mixture distribution analysis and found that over half of all cancers in mammography were fixated within 1 second, with the remaining cancers fixated in subsequent search. these two distributions of time-to-first fixations were taken as evidence to support the two information processing systems involved in the detection of abnormalities; 1) rapid initial holistic processing, and 2) the slower processing relating to search and discovery. however, as mentioned above, target conspicuity can drive early detection rates and therefore it does not appear to have been considered that instead of these distributions reflecting the two information processes in question, these distributions may simply be the time taken to find obvious and subtle cancers respectively. indeed, one of the actual studies used by kundel et al. (2008) found that time to first fixate cancers is dependent on the subtlety of the targets, with subtle cancers taking longer to be fixated (krupinski, 1995). in addition, we recently showed that unlike mammography, with chest x-rays only 33% of cancers were fixated within 1 second, whereas 56% of cancers were fixated within 2 seconds (donovan & litchfield, 2013). in both cases observers would have globally processed the image, but again, it is not clear whether the subsequent differences in search latencies arise because of the differences in how each image modality was globally processed, or simply because our chest x-ray images contained more subtle abnormalities which led to longer search times. as we have recently argued (donovan & litchfield, 2013; litchfield & donovan, 2016), the underlying methodological problem with using time-to-first fixation data is that it is obtained from eye tracking experiments under free viewing conditions, whereby the observer has constant access to the whole scene via peripheral vision. this makes it difficult to isolate the specific contribution of processing the initial glimpse of the scene on subsequent eye movement behaviour. scene perception research suggests that the initial representation evolves during scene viewing, and may provide a frame on which subsequent information can be added (friedman, 1979). visual and semantic information obtained from successive fixations on objects and other aspects of the scene can exert influence on future eye movements (henderson, weeks, & hollingworth, 1999). this poses a problem when it comes to examining eye movements during free viewing as it can be difficult to discriminate how eye guidance is affected by the initial scene litchfield et donovan | f l r 70 representation compared with this continuously updated representation. kundel et al. (2008) acknowledged that input is continuously gathered from the periphery during search, but are unable to specify precisely how the initial global processing interacts with the ongoing representation and how such information is integrated to guide search. as it stands, our understanding of visual expertise in medical image perception has been derived from different methodological techniques where eye movements were prevented in tachistoscopic ‘flash’ studies (kundel et al., 1975; carmody et al., 1981; oestmann et al., 1988), or by analysing time-tofirst fixation data in which the initial global impression is never dissociated from the ongoing scene representation (kundel et al., 2008). the recent meta-analysis by gegenfurtner et al. (2011) confirmed that across a range of visual domains, experts do indeed have shorter time-to-first fixations and make longer saccades compared to non-experts. but whilst this provides even more compelling evidence that experts have faster search latencies and can process domain-specific scenes better than novices, it is still unclear whether it is an expertise-specific advantage in global processing that specifically contributes to search guidance. moreover, recent research has shown that global processing may be exploited to rapidly categorize the image, but this does not mean such processing directly supports object recognition. evans et al. (2013) adopted the traditional ‘flash’ methodology and required experts and non-experts to rate whether flashed images were abnormal or not under different flash durations. in addition, a subset of observers also had to localize where they thought the abnormality was using a blank outline of the image. by dissociating detection from localization decisions, evans et al. found that experts could exploit this initial glimpse of the image better than non-experts regardless of the flash duration. critically however, all groups of observers were only at chance when it came to actually locating the abnormalities. although evans et al. did not measure eye movements, these findings are consistent with recent research on scene perception in that global processing appears to enable the gist of the scene to be extracted, but this is just part of a larger system of eye guidance (see tatler, 2009), and so it is not necessarily the case that global processing itself constrains subsequent search. given the aforementioned problems with flash and free viewing eye-tracking methodologies, we argued that a new methodology was needed that can dissociate initial scene processing from subsequent search and decision making (donovan & litchfield, 2013). we therefore proposed that the recently developed fpmw paradigm (castelhano & henderson, 2007) would be a suitable methodology for providing direct evidence of how the initial scene representation may guide search and decision-making and help isolate the contribution of domain-specific visual expertise, compared to our shared expertise at processing real world images in scene perception. 2. flash-preview moving window the fpmw paradigm (castelhano & henderson, 2007) draws on the ‘flash’ methodology that was previously discussed, but also the ‘moving window’ paradigm. the gaze-contingent ‘moving window’ paradigm was originally developed by mcconkie and rayner (1975) to investigate the visual span in reading. the ‘moving window’ allows the observer to continue making eye movements whilst information presented at fovea, parafovea and the periphery is systematically controlled by the researcher and is one of the most powerful eye-tracking methods we have available (rayner, 2009). this technique alters how much of the scene can be processed by overlaying a variable size mask that is tied to the central fixation position recorded from an eye-tracker and occludes the rest of the scene outside the gaze-contingent moving window. typically, the observer can examine a scene using their high-resolution fovea but are not able to use parafoveal and peripheral processing as all visual information outside the window is degraded, or simply turned into irrelevant material. much to the ire of participants, this can also be reversed so that the mask occludes the fovea, and so instead the scene can only be processed by parafoveal and peripheral vision. by changing the size of the window to the point that performance is significantly worse than no window conditions, researchers can systematically assess how much information can be detected and extensively processed outside the high-resolution fovea. using this paradigm it has been established that experience in reading a particular language changes the size and shape of the visual span (pollatsek, bolozky, well, & rayner, 1981), and that the visual span increases with expertise, whether this is in terms of reading skill litchfield et donovan | f l r 71 (rayner, slattery, & bélanger, 2010), or in other domains of visual expertise, such as chess (charness, reingold, pomplun, & stampe, 2001; reingold, charness, pomplun, & stampe, 2001; reingold & sheridan, 2011). in the fpmw paradigm (castelhano & henderson, 2007), observers are briefly shown a preview of the upcoming search scene and are subsequently asked to search for a particular target object whilst their peripheral vision is restricted to a gaze-contingent moving window. observers therefore have to rely on what they could process from the initial glimpse of the scene and use this initial representation to guide their search, alongside any other knowledge they have about the target and scene-type (e.g., what the target typically looks like and the likely location of the target given the scene-context). this method offers unique insights as to how the initial glimpse guides eye movement behaviour and draws on the strengths of the previously discussed methodologies. with fpmw, not only is it possible to systematically control what is presented in the scene preview and allow subsequent search-related eye movements to be made, the type of visual processing (e.g., foveal, parafoveal, peripheral) can also be systematically controlled during these eye movements. in the standard setup (see figure 1a for an example), observers begin by fixating a central fixation cross and then a preview of a scene or control image (for example a visual mask, or different scene type) is briefly flashed (typically for around 250ms), this is then replaced by a visual mask (for 50ms). a word then appears (for around 1000ms-2000ms) indicating which target object must be found. search then commences and the observer has 15 seconds to find the target whilst their peripheral vision is restricted to a small gaze-contingent moving window (between 2° and 5°). once the observer has found the target, they press a button whilst looking directly at the target (a gamepad is used to avoid latency delays from keyboards). subsequent analysis is used to identify whether the target object was correctly identified. this can be done by creating areas of interest (aoi) around the target object and establishing whether or not the observers gaze was within the specified aoi when the button was pressed, with analysis usually restricted to those trials where targets were correctly identified. as this is a gaze-contingent display method, it is important that the display screen is updated with as minimal delay as possible. as such, the fpmw should be presented on a monitor with a high refresh rate (>100hz, preferably crt to minimise delay) using a desktop eye-tracker with a high sampling rate (>500hz) and a chin rest should be used to maintain accuracy (calibrations should be accepted only if average visual angle error < 0.5°). figure 1. trial sequence of the ‘flash-preview moving window’ paradigm a) the traditional fpmw sequence where target identity is known only after the preview has been encoded b) the adapted trial sequence used by litchfield and donovan (2016), where target identity is known before the preview has been encoded. litchfield et donovan | f l r 72 there are a number of key parameters that can be manipulated in the fpmw depending on the research question (see section 3 below). specific eye movement metrics and performance measures can be obtained to establish the effects of the initial scene preview on subsequent search and decision-making (see table 1). several studies have now used fpmw to investigate scene perception and they all largely demonstrate scene preview benefits in search, whereby the time-to-first fixate targets (search latency) is faster with scene previews compared to control conditions, and other eye movement metrics reflect greater efficiencies in how windowed search is initiated and executed. note that time-to-first fixation is one of the primary metrics to establish how windowed search is guided by the initial glimpse of a scene using fpmw, however, this metric is distinct from time-to-first fixation used in free viewing conditions, as the latter do not control the level of parafoveal and peripheral processing during search. the fpmw paradigm (castelhano & henderson, 2007) has shown that scene preview benefits are obtained when the preview is identical to the subsequent search scene, and this benefit is maintained even if the preview is different in size. however, there is no benefit if the preview image is different to the search scene but belongs to the same scene category. moreover, the scene preview benefit exists even if the target object was not visible during the preview (i.e., digitally removed), but only found through windowed search, thereby confirming the benefit of scene-context processing, irrespective of any additional local target processing that could occur when targets are present in previews (castelhano & henderson, 2007; võ & henderson, 2010). the fpmw has also been used to demonstrate how semantically consistent and inconsistent objects are processed within scenes (castelhano & heaven 2011; võ & henderson, 2011), and how learned object function may guide attention aside from object features (castelhano & witherspoon, 2016). the ability to process the scene preview has been linked to individual differences in visual perceptual processing speed (võ & schneider, 2010), and the time-course of the initial representation derived from the scene preview has also been investigated. although 250ms is the standard preview duration, previews as short as 75ms can lead to search advantages compared to no preview conditions, and even as low as 50ms, provided the integration time between scene preview, target word and search was extended from 500ms to 3000ms (võ & henderson, 2010). moreover, preview benefits in search diminish with time, and in fact, the benefit of an initial glimpse of a scene facilitates search only up to the fourth fixation (hillstrom, schloley, liversedge, & benson, 2012). indeed, based on these latter fpmw findings we argued that the benefit of the initial glimpse in medical image perception would depend on the difficulty of finding abnormalities within this short time-frame (donovan & litchfield, 2013). taken together, this growing body of research using fpmw provides many useful insights into our shared expertise in scene perception and how we rapidly guide our eye movements based on just a single glimpse of the upcoming scene. to establish the contribution of domain-specific expertise in eye guidance, we recently applied the fpmw to medical image perception for the first time and examined whether experts are able to guide their eye movement behaviour more effectively than novices, solely upon seeing an initial glimpse (litchfield & donovan, 2016). in all experiments the scene preview was either an identical preview of the upcoming scene or a mask preview (i.e., random noise meaning no preview of the upcoming scene). we first compared performance and eye movement behaviour in a standard scene perception task that required objects to be found from real-world scenes. consistent with previous fpmw studies, both expertise groups showed identical scene preview benefits, in that their eye movements were more efficient at finding the targets in the scene preview condition compared to the mask preview condition. this is consistent with previous research showing that experts superior visual perceptual processing is domain-specific and does not transfer to domain general tasks that draw on shared expertise (nodine & krupinski, 1998; sowden, davies, & roling, 2000). the key comparison study, however, was how well these same expert and novice observers were able to find lung nodules (cancer) from chest x-rays in a subsequent fpmw experiment. we found the typical expertise effect in diagnostic performance, with experts being able to identify more cancers than novices. expert diagnostic performance was the same in scene preview as mask preview. in contrast, novice observers were actually worse at making diagnostic decisions when presented a scene preview prior to search. we interpreted these negative effects of preview on novice performance as being due to interference from distractors, in that by previewing the upcoming medical image novices were exposed to the potential ‘nodule-like’ distractors that are inherent in chest x-rays, but are in fact normal features. the fact that litchfield et donovan | f l r 73 expert’s decisions were not similarly impaired in scene preview suggested that their more elaborate knowledge of what nodules are ensured they were less biased by these ‘normal’ distractors. table 1. key dependent variables obtained using the ‘flash-preview moving window’ paradigm. dependent variable definition performance measures: accuracy (%) the % of targets correctly identified during search. reaction times (rt) the time taken from onset of windowed search until search is terminated via button press. search-related measures: time-to-first fixation (search latency) this is the total amount of time spent from the onset of the search scene up to (but not including) the first fixation on the target. this is considered one of the primary measures to establish that a scene preview benefit has been found in relation to search. number of fixations this is the total number of fixations made from the onset of the search scene up to (but not including) the first fixation on the target. although this measure is related to latency, it can help show whether observers are effectively selecting target candidates for fixation. scan path ratio. this is how much of the scene was explored. the scan path ratio is the length of the scan pattern through the scene until the first fixation on the target (total distance between all fixations from scene onset to the first fixation on target) divided by the most direct path to the object (distance from the central fixation point to the centre of the target object). as the ratio approaches 1 this indicates an optimal path to target location (henderson et al., 1999). first eye movement of search: initial saccadic latency this is the latency of the first eye movement of search. initial saccadic amplitude. this is the amplitude of the first eye movement of search. target processing measures: first fixation duration this is the duration of the initial fixation on the target and reflects the initial processing of the target. first gaze duration the total duration of all fixations on the target since it was first fixated until the gaze moves away from the target. total time on target this is the summation of all fixation durations on the target before a response button is pressed (including any refixations). litchfield et donovan | f l r 74 given the holistic model (kundel et al., 2007) suggests that experts should be better at searching for targets in these domain-specific scenes, we were surprised to find that there was only a weak scene preview effect in this medical perception task. more specifically, both novices and experts showed a slight search improvement with scene preview, with nodules fixated in fewer fixations and a borderline search latency improvement, however, with each search measure there was no expertise x preview interaction. when we unpacked the effect of scene preview we found that experts demonstrated more of a search benefit (-765 ms) than novices (-260 ms), but the scale of these scene preview search advantages were dwarfed in comparison to the scene preview benefits these same experts and novices demonstrated when searching for targets in scene perception (–1,282ms and –1,620ms respectively). medical abnormalities are more difficult to identify than targets used in scene perception research and so having such difficult targets to find could have contributed to this weaker scene preview search benefit. however, scene preview effects are still found even when the target is not even visible during scene previews (castelhano & henderson, 2007), and so it is likely that the learned spatial associations between target and scene are greatly contributing to the scene preview benefit and where best to look. if the targets in scene perception tasks have a more predictable location (a closer target-scene pairing) then previewing the scene-context would enable those learned target-scene spatial associations to be incorporated and guide eye movements. however, in our medical image perception task we highlighted that nodules may not have such a close target-scene spatial association (båth et al., 2005), and this may be why we observed such a weak scene preview effect in this medical image perception task. the broader implication to visual expertise is that experts would have learned these spatial associations whereas novices will have not – so whilst we only found weak expertise effects in our specific medical imaging study, it follows that much stronger scene preview effects should be found in tasks where there is a clearer target-scene spatial association for experts to learn, and subsequently exploit. whilst these issues are not well specified in the holistic model at all, they are clarified in greater detail in the increasing number of eye guidance models in scene perception (for an overview, see tatler, 2009). indeed, this once again highlights the growing need to reconcile visual expertise research in medical image perception with current work in scene perception. applying fpmw to medical image perception also raised a key methodological distinction between scene perception research and medical image perception tasks. in medical image perception observers are typically searching for the same type of target (e.g., nodule) within the same type of scene (e.g., chest x-ray), whereas in scene perception research the target and scene type are often randomised on each trial. as such, the latter situation may be maximising the advantage of seeing the scene preview relative to mask preview. to rule out the possibility that our weak scene preview effects were simply due to observers repeatedly searching for the same type of target from the same type of scene, a third fpmw experiment required novice and experienced observers to search through a random presentation of three different types of medical images (chest x-rays, brain images, skeletal images), each with their own specific target abnormalities (lung nodules, brain tumors, bone fractures). the rationale was that this should give experienced observers the opportunity to exploit their rapid understanding of these scenes compared to novices and demonstrate a stronger scene preview benefit with increasing expertise. we found that experienced observers were better overall than novices at identifying abnormalities, and their reaction times were faster than novices regardless of the preview condition. once again novices were worse at identifying targets if given a preview, but controversially, now even experienced observers were impaired at identifying target abnormalities if shown a scene preview before commencing search. furthermore, we observed no scene preview benefit in search whatsoever. although the impairments in accuracy for scene preview were relatively small (~3% drop in accuracy), the fact that we observed any such impairment in experienced observers seeing an initial glimpse of the scene goes against our conventional wisdom that processing the initial glimpse of the scene is beneficial to performance. to interpret these counter-intuitive findings and the fact that expert observers were not likewise impaired in experiment 2, we indicated that targets in medical image perception are difficult to identify even if directly fixating them (kundel et al., 1978; donovan & litchfield, 2013), and that by repeatedly searching the same type of scene expert observers may be able to attenuate the distractors that share similar features with pathology (cf. kompaniez-dunigan, abbey, boone, & webster, 2015). clearly, however, much more research is required to uncover the reasons behind these effects and to better litchfield et donovan | f l r 75 understand what processes govern whether an initial glimpse does, or does not, lead to more effective search and decision-making. the purpose of this paper was to highlight the current theoretical and methodological frontier and introduce the fpmw paradigm as a method for investigating visual expertise. however, this is not the first time that gaze-contingent eye-tracking methodologies have been used in medical image perception. not only were kundel and nodine early adopters of eye-tracking research (kundel & la follette, 1972; kundel, nodine, & carmody, 1978), they also undertook 'moving window' studies of their own in a bid to dissociate the role central and peripheral vision has on search and decision-making (kundel, nodine, & toto, 1984; kundel, nodine, & toto, 1991). although data was only based on 2 (kundel et al., 1984) or 4 expert observers (kundel et al., 1991), in their study a single chest x-ray was fully viewable at all times and lung nodules that were artificially added to different areas of the image were only visible when they fell within the specified moving window (either 1.5°, 3.5°, 5.5°, 8.5°, or no window). a single chest x-ray was selected for all trials so that the global features of the image were controlled and that the specific role of foveal, parafoveal and peripheral vision could be determined (kundel et al., 1984). even from these early studies, it became clear that reducing the size of the moving window to 1.5° dramatically impaired the detection of the nodules and that the time to first fixate nodules decreased with larger windows, consistent with what would later be established more clearly in chess expertise (reingold et al., 2001) and reading expertise (rayner et al., 2010). we therefore see the fpmw paradigm as a continuation of this long tradition to understand visual expertise using eye-tracking methodology alongside traditional ‘flash’ methodologies. although this fpmw methodology currently focuses on static image interpretation, and not on how dynamic medical images are interpreted (bertram, helle, kaakinen, & svedström, 2013; drew et al., 2013b; phillips et al., 2013) we are nevertheless excited to see how the fpmw will force us to re-evaluate what we think we know about visual expertise in medical image perception. unlike many of the other methodologies that can be used to investigate visual expertise and are discussed in this special issue, research has only just begun on using fpmw to understand visual expertise. as such, for instructive purposes the remainder of the paper provides an overview of the key parameters that can be manipulated in fpmw. 3. overview of key fpmw parameters 3.1 flash preview (content) a key strength of the fpmw paradigm is that what is presented in the preview does not have to correspond with what is subsequently searched for using the moving window. the baseline scene preview condition should be identical to the subsequent search scene, but otherwise, performance can be compared to a variety of control preview manipulations (including whether targets are even present or changed in previews). 3.2 flash preview duration this is how long the preview is presented for and the typical duration is 250ms. shorter durations can be used to obtain a scene preview benefit (see võ, & henderson, 2010), but if longer durations are used eye movements are likely to be made actually within the preview. litchfield et donovan | f l r 76 3.3 size of moving window this refers to the size of the moving window. the typical window size is between 2-5 degrees and reducing the size of window is likely to increase search times (cf. kundel et al., 1984). it is unknown at this time whether there is a relationship between flash preview duration and the size of the moving window. 3.4 sequence of preview in the majority of fpmw studies to date, the observer is first given a scene preview and then afterwards informed what target should be detected in search (figure 1a). this setup is useful to understand how newly activated target knowledge can be integrated with the currently held visual representation of the scene. however, there are many situations where observers already know the target item they are looking for before they actually see the critical search scene. indeed, in all of the medical image perception studies discussed, observers knew beforehand what was the target (e.g., cancer) before they actually saw the medical image. accordingly, litchfield and donovan (2016) created a modified fpmw sequence where target knowledge was activated before the preview was presented (figure 1b). note that robust scene preview benefits were found for scene perception using this modified sequence, even if the same was not the case when medical images were used. 3.5 integration time this refers to the time interval between the scene preview, the target knowledge and the onset of windowed search. integration can influence the extent to which the scene preview can be exploited to guide search (võ & henderson, 2010), and will also be affected by the sequence of preview. keypoints the ‘flash-preview moving window’ (fpmw) paradigm (castelhano & henderson, 2007) is a new gaze-contingent eye tracking technique that isolates the specific contribution a brief glimpse of a scene has on subsequent eye movement behaviour and decision-making. prevailing theories of medical image perception such as the holistic model (kundel et al., 2007) propose that experts exploit global processing of the initial glimpse of medical images to detect abnormalities and guide their eye movements better than novices. the fpmw has recently been used for the first time to directly assess the contribution of the initial glimpse on subsequent search behaviour as a function of expertise (litchfield & donovan, 2016), but found only weak expertise-specific advantages of processing scene previews. applying the fpmw to a medical image perception task demonstrated that novice observers and (in some situations) experienced observers, were worse at identifying targets when given a scene preview of the upcoming search scene. further research using fpmw in different visual expertise domains is required to build on these preliminary findings. litchfield et donovan | f l r 77 references båth, m., håkansson, m., börjesson, s., kheddache, s., grahn, a., ruschin, m., . . . månsson, l. g. (2005). nodule detection in digital chest radiography: introduction to the radius chest trial. radiation protection dosimetry, 114, 85–91. https://doi.org/10.1093/rpd/nch575 bertram, r., helle, l., kaakinen, j. k., & svedström, e. (2013). the effect of expertise on eye movement behaviour in medical image perception. plos one, 8, e66169. https://doi.org/10.1371/journal.pone.0066169 busey, t. a., & vanderkolk, j. r. (2005). behavioral and electrophysiological evidence for configural processing in fingerprint experts. vision research, 45, 431-448. https://doi.org/10.1016/j.visres.2004.08.021 buswell, g. (1935). how people look at pictures. university of chicago press. carmody, d. p., nodine, c. f., & kundel, h. l. (1980). global and segmented search for lung nodules of different edge gradients. investigative radiology, 15, 224–233. carmody, d. p., nodine, c. f., & kundel, h. l. (1981). finding lung nodules with and without comparative visual search. perception & psychophysics, 29, 594–598. https://doi.org/10.3758/bf03207377 castelhano, m. s., & heaven, c. (2011). scene context influences without scene gist: eye movements guided by spatial associations in visual search. psychonomic bulletin & review, 18, 890-896. https://doi.org/10.3758/s13423-011-0107-8 castelhano, m. s., & henderson, j. m. (2007). initial scene representations facilitate eye movement guidance in visual search. journal of experimental psychology: human perception and performance, 33, 753-763. https://doi.org/10.1037/0096-1523.33.4.753 castelhano, m. s., & witherspoon, r. l. (2016). how you use it matters: object function guides attention during visual search in scenes. psychological science, 27, 606-621. https://doi.org/10.1177/0956797616629130 charness, n., reingold, e. m., pomplun, m., & stampe, d. m. (2001). the perceptual aspect of skilled performance in chess: evidence from eye movements. memory & cognition, 29, 1146-1152. https://doi.org/10.3758/bf03206384 cohen, m. a., dennett, d. c., & kanwisher, n. (2016). what is the bandwidth of perceptual experience? trends in cognitive sciences, 20, 324-335. https://doi.org/10.1016/j.tics.2016.03.006 curby, k. m., glazek, k., & gauthier, i. (2009). a visual short-term memory advantage for objects of expertise. journal of experimental psychology: human perception and performance, 35, 94-107. https://doi.org/10.1037/0096-1523.35.1.94 de groot, a. d. (1946). het denken van den schaker. amsterdam: noord hollandsche. de groot, a. d. (1965). thought and choice in chess. the hague, netherlands: mouton. diamond, r., & carey, s. (1986). why faces are and are not special: an effect of expertise. journal of experimental psychology: general, 115, 107-117. https://doi.org/10.1037/0096-3445.115.2.107 donovan, t., & litchfield, d. (2013). looking for cancer: expertise related differences in searching and decision making. applied cognitive psychology, 27, 43-49. https://doi.org/10.1002/acp.2869 drew, t., evans, k. k., võ, m. l. h., jacobson, f. l., & wolfe, j. m. (2013a). informatics in radiology: what can you see in a single glance and how might this guide visual search in medical images? radiographics, 33, 263–274. https://doi.org/10.1148/rg.331125023 drew, t., vo, m. l. h., olwal, a., jacobson, f., seltzer, s. e., & wolfe, j. m. (2013b). scanners and drillers: characterizing expert visual search through volumetric images. journal of vision, 13, 3. https://doi.org/10.1167/13.10.3 evans, k. k., georgian-smith, d., tambouret, r., birdwell, r. l., & wolfe, j. m. (2013). the gist of the abnormal: above-chance medical decision making in the blink of an eye. psychonomic bulletin & review, 20, 1170-1175. https://doi.org/10.3758/s13423-013-0459-3 fei-fei, l., iyer, a., koch, c., & perona, p. (2007). what do we perceive in a glance of a real-world scene? journal of vision, 7, 1-29. https://doi.org/10.1167/7.1.10 fox, e. s., & faulkner-jones, b. e. (2017). eye-tracking in the study of visual expertise: methodology and approaches in medicine. frontline learning research, 5, 29-40. https://doi.org/10.14786/flr.v5i3.258. https://doi.org/10.1093/rpd/nch575 https://doi.org/10.1093/rpd/nch575 https://doi.org/10.1016/j.visres.2004.08.021 https://doi.org/10.3758/bf03207377 https://doi.org/10.3758/s13423-011-0107-8 https://doi.org/10.1037/0096-1523.33.4.753 https://doi.org/10.1177/0956797616629130 https://doi.org/10.3758/bf03206384 https://doi.org/10.1016/j.tics.2016.03.006 https://doi.org/10.1037/0096-1523.35.1.94 https://doi.org/10.1037/0096-3445.115.2.107 https://doi.org/10.1002/acp.2869 https://doi.org/10.1148/rg.331125023 https://doi.org/10.1167/13.10.3 https://doi.org/10.3758/s13423-013-0459-3 https://doi.org/10.1167/7.1.10 https://doi.org/10.14786/flr.v5i3.258 litchfield et donovan | f l r 78 friedman a. (1979). framing pictures: the role of knowledge in automatized encoding and memory for gist. journal of experimental psychology: general, 108, 316-355. https://doi.org/10.1037/00963445.108.3.316 gauthier, i., & tarr, m. j. (1997). becoming a “greeble” expert: exploring mechanisms for face recognition. vision research, 37, 1673-1682. https://doi.org/10.1016/s0042-6989(96)00286-6 gauthier, i., tarr, m., & bub, d. (2009). perceptual expertise: bridging brain and behavior. oxford university press. https://doi.org/10.1093/acprof:oso/9780195309607.001.0001 gegenfurtner, a., lehtinen, e., & säljö, r. (2011). expertise differences in the comprehension of visualizations: a meta-analysis of eye-tracking research in professional domains. educational psychology review, 23, 523-552. https://doi.org/10.1007/s10648-011-9174-7 gobet, f. (2015). understanding expertise: a multidisciplinary approach. london: palgrave. henderson, j. m. (2007). regarding scenes. current directions in psychological science, 16, 219-222. https://doi.org/10.1111/j.1467-8721.2007.00507.x henderson, j. m., & hollingworth, a. (1999). the role of fixation position in detecting scene changes across saccades. psychological science, 5, 438-443. https://doi.org/10.1111/1467-9280.00183 henderson, j. m., weeks, p. a., & hollingworth, a. (1999). the effects of semantic consistency on eye movements during complex scene viewing. journal of experimental psychology: human perception and performance, 25, 210-228. https://doi.org/10.1037/0096-1523.25.1.210 henderson, j. m., pollatsek, a., & rayner, k. (1989). covert visual attention and extrafoveal information use during object identification. perception & psychophysics, 45, 196-208. https://doi.org/10.3758/bf03210697 hillstrom, a. p., schloley, h., liversedge, s. p., benson, v. (2012). the effect of the first glimpse at a scene on eye movements during search. psychomic bulletin & review, 19, 204-210. https://doi.org/10.3758/s13423-011-0205-7 just, m. a., & carpenter, p. a. (1984). using eye fixations to study reading comprehension. in d. e. kieras & m. a. just (eds.), new methods in reading comprehension research (pp. 151–182). hillsdale: erlbaum. kelley, t. a., chun, m. m., & chua, k. p. (2003). effects of scene inversion on change detection of targets matched for visual salience. journal of vision, 3, 1-5. https://doi.org/10.1167/3.1.1 kompaniez-dunigan, e., abbey, c. k., boone, j. m., & webster, m. a. (2015). adaptation and visual search in mammographic images. attention, perception, & psychophysics, 77, 1081-1087. https://doi.org/10.3758/s13414-015-0841-5 krupinski, e. a. (1995). visual search of mammographic images: influence of lesion subtlety. academic radiology, 12, 965-969. https://doi.org/10.1016/j.acra.2005.03.071 krupinski, e. a. (1996). visual scanning patterns of radiologists searching mammograms. academic radiology, 3, 137-144. https://doi.org/10.1016/s1076-6332(05)80381-2 kundel, h. l., & la follette, p. s. (1972). visual search patterns and experience with radiological images. radiology, 103, 523–528. https://doi.org/10.1148/103.3.523 kundel, h. l., & nodine, c. f. (1975). interpreting chest radiographs without visual search. radiology, 116, 527–532. https://doi.org/10.1148/116.3.527 kundel, h.l., nodine, c.f., & carmody, d.p. (1978). visual scanning, pattern recognition, and decision making in pulmonary nodule detection. investigative radiology, 13, 175–181. https://doi.org/10.1097/00004424-197805000-00001 kundel, h. l., nodine, c. f., conant, e. f., & weinstein, s. p. (2007). holistic component of image perception in mammogram interpretation: gaze-tracking study 1. radiology, 242, 396-402. https://doi.org/10.1148/radiol.2422051997 kundel, h. l., nodine, c. f., krupinski, e. a., mello-thoms, c. (2008). using gaze-tracking data and mixture distribution analysis to support a holistic model. academic radiology, 15, 881–886. https://doi.org/10.1016/j.acra.2008.01.023 kundel, h. l., nodine, c. f., & toto, l. c. (1984). eye movements and the detection of lung tumors in chest images. in a. g. gale, & f. johnson (eds.). theoretical and applied aspects of eye movement research (pp. 297-304). north-holland: elsevier. https://doi.org/10.1037/0096-3445.108.3.316 https://doi.org/10.1037/0096-3445.108.3.316 https://doi.org/10.1016/s0042-6989(96)00286-6 https://doi.org/10.1093/acprof:oso/9780195309607.001.0001 https://doi.org/10.1007/s10648-011-9174-7 https://doi.org/10.1111/j.1467-8721.2007.00507.x https://doi.org/10.1111/1467-9280.00183 https://doi.org/10.1037/0096-1523.25.1.210 https://doi.org/10.3758/bf03210697 https://doi.org/10.3758/s13423-011-0205-7 https://doi.org/10.1167/3.1.1 https://doi.org/10.3758/s13414-015-0841-5 https://doi.org/10.1016/j.acra.2005.03.071 https://doi.org/10.1016/s1076-6332(05)80381-2 https://doi.org/10.1148/103.3.523 https://doi.org/10.1148/116.3.527 https://doi.org/10.1097/00004424-197805000-00001 https://doi.org/10.1148/radiol.2422051997 https://doi.org/10.1016/j.acra.2008.01.023 litchfield et donovan | f l r 79 kundel, h. l., nodine, c. f., & toto, l. c. (1991). searching for lung nodules. the guidance of visual scanning. investigative radiology, 26, 777-781. litchfield, d., ball, l. j., donovan, t., manning, d. j., & crawford, t. (2010). viewing another person’s eye movements improves identification of pulmonary nodules in chest x-ray inspection. journal of experimental psychology: applied, 16, 251-262. https://doi.org/10.1037/a0020082. litchfield, d., & donovan t. (2016). worth a quick look? initial scene previews can guide eye movements as a function of domain-specific expertise but can also have unforeseen costs. journal of experimental psychology: human perception and performance, 42, 982-994. https://doi.org/10.1037/xhp0000202 manning, d. j., ethell, s., & crawford, t. (2003). an eye-tracking afroc study of the influence of experience and training on chest x-ray interpretation. spie, 5034, 257–266. https://doi.org/10.1117/12.479985 manning, d. j., ethell, s. c., & donovan, t. (2004). detection or decision errors? missed lung cancer from the posteroanterior chest radiograph. british journal of radiology, 77, 231-235. https://doi.org/10.1259/bjr/28883951 manning, d. j., ethell, s. c., donovan, t., & crawford, t. j. (2006). how do radiologists do it? the influence of experience and training on searching for chest nodules. radiography, 12, 134-142. https://doi.org/10.1016/j.radi.2005.02.003 mcconkie, g. w., & rayner, k. (1975). the span of the effective stimulus during a fixation in reading. perception & psychophysics, 17, 578-86. https://doi.org/10.3758/bf03203972 mugglestone, m. d., gale, a. g., cowley, h. c., & wilson, a. r. m. (1995). diagnostic performance on briefly presented mammographic images. spie, 2436, 106–115. https://doi.org/10.1117/12.206840 nodine, c. f., & krupinski, e. a. (1998). perceptual skill, radiology expertise, and visual test performance with nina and waldo. academic radiology, 5, 603-612. https://doi.org/10.1016/s10766332(98)80295-x nodine, c. f., & kundel, h. l. (1987). the cognitive side of visual search in radiology. in j. k., o’regan, a. levy-schoen (eds). eye movements: from physiology to cognition (pp. 573–582). amsterdam: elsevier. https://doi.org/10.1016/b978-0-444-70113-8.50081-3 nodine, c. f., & mello-thoms, c. (2010). the role of expertise in radiologic image interpretation. in e. samei, e. krupinski (eds.), the handbook of medical image perception and techniques (pp. 139– 156), new york: cambridge university press. nodine, c. f., mello-thoms, c., kundel, h. l., & weinstein, s. p. (2002). time course of perception and decision making during mammographic interpretation. american journal of roentgenology, 179, 917923. https://doi.org/10.2214/ajr.179.4.1790917 oestmann, j. w., greene, r., bourgouin, p. m., linetsky, l., & llewellyn, h. j. (1993). chest “gestalt” and detectability of lung lesions. european journal of radiology, 16, 154-157. https://doi.org/10.1016/0720-048x(93)90015-f oestmann, j. w., greene, r., kushner, d. c., bourgouin, p. m., linetsky, l., & llewellyn, h. j. (1988). lung lesions: correlation between viewing time and detection. radiology, 166, 451-453. https://doi.org/10.1148/radiology.166.2.3336720 oliva, a., & schyns, p. g. (1997). coarse blobs or fine edges? evidence that information diagnosticity changes the perception of complex visual stimuli. cognitive psychology, 34, 72-107. https://doi.org/10.1006/cogp.1997.0667 phillips, p., boone, d., mallett, s., taylor, s. a., altman, d. g., manning, d., ... & halligan, s. (2013). method for tracking eye gaze during interpretation of endoluminal 3d ct colonography: technical description and proposed metrics for analysis. radiology, 267, 924-931. https://doi.org/10.1148/radiol.12120062 pollatsek, a., bolozky, s., well, a. d., & rayner k. (1981). asymmetries in the perceptual span for israeli readers. brain and language, 14, 174-180. https://doi.org/10.1016/0093-934x(81)90073-0 potter, m. c. (1976). short-term conceptual memory for pictures. journal of experimental psychology: human learning and memory, 2, 509-522. https://doi.org/10.1037//0278-7393.2.5.509 https://doi.org/10.1037/a0020082 https://doi.org/10.1037/xhp0000202 https://doi.org/10.1117/12.479985 https://doi.org/10.1259/bjr/28883951 https://doi.org/10.1016/j.radi.2005.02.003 https://doi.org/10.3758/bf03203972 https://doi.org/10.1117/12.206840 https://doi.org/10.1016/s1076-6332(98)80295-x https://doi.org/10.1016/s1076-6332(98)80295-x https://doi.org/10.1016/b978-0-444-70113-8.50081-3 https://doi.org/10.2214/ajr.179.4.1790917 https://doi.org/10.1016/0720-048x(93)90015-f https://doi.org/10.1148/radiology.166.2.3336720 https://doi.org/10.1006/cogp.1997.0667 https://doi.org/10.1148/radiol.12120062 https://doi.org/10.1016/0093-934x(81)90073-0 https://doi.org/10.1037/0278-7393.2.5.509 litchfield et donovan | f l r 80 reingold, e. m., & sheridan, h. (2011). eye movements and visual expertise in chess and medicine. in s. p. liversedge, i. d. gilchrist, s. everling (eds). the oxford handbook of eye movements (pp. 767–786). oxford: oxford university press. https://doi.org/10.1093/oxfordhb/9780199539789.013.0029 rayner, k. (2009). eye movements and attention in reading, scene perception, and visual search. the quarterly journal of experimental psychology, 62, 1457-1506. https://doi.org/10.1080/17470210902816461 rayner, k., & pollatsek, a. (1981). eye movement control during reading: evidence for direct control. the quarterly journal of experimental psychology, 33, 351-373. https://doi.org/10.1080/14640748108400798 rayner, k., slattery, t. j., & bélanger, n. n. (2010). eye movements, the perceptual span, and reading speed. psychonomic bulletin & review, 17, 834-839. https://doi.org/10.3758/pbr.17.6.834 rayner, k., smith, t. j., malcom, g. l., & henderson, j. m. (2009). eye movements and visual encoding during scene perception. psychological science, 20, 6-10. https://doi.org/10.1111/j.14679280.2008.02243.x schyns, p. g., & oliva, a. (1994). from blobs to boundary edges: evidence for time-and spatial-scaledependent scene recognition. psychological science, 5, 195-200. https://doi.org/10.1111/j.14679280.1994.tb00500.x sowden, p. t., davies, i. r., & roling, p. (2000). perceptual learning of the detection of features in x-ray images: a functional role for improvements in adults’ visual sensitivity? journal of experimental psychology: human perception and performance, 26, 379–390. https://doi.org/10.1037/00961523.26.1.379 swensson, r. g. (1980). a two-stage detection model applied to skilled visual search by radiologists. perception & psychophysics, 27, 11-16. https://doi.org/10.3758/bf03199899 szulewski, a., kelton, d., & howes, d. (2017). pupillometry as a tool to study expertise in medicine. frontline learning research, 5, 52-62. https://doi.org/10.14786/flr.v5i3.256 tatler, b. w. (2009). current understanding of eye guidance. visual cognition, 17, 777-789. https://doi.org/10.1080/13506280902869213 torralba, a., oliva, a., castelhano, m. s., & henderson, j. m. (2006). contextual guidance of eye movements and attention in real-world scenes: the role of global features in object search. psychological review, 113, 766-786. https://doi.org/10.1037/0033-295x.113.4.766 võ, m. l. h., & henderson, j. m. (2010). the time course of initial scene processing for eye movement guidance in natural scene search. journal of vision, 10, 1-13. https://doi.org/10.1167/10.3.14 võ, m. l. h., & henderson, j. m. (2011). object-scene inconsistencies do not capture gaze: evidence from the flash-preview moving window paradigm. attention, perception & psychophysics, 73, 1742–1753. https://doi.org/10.3758/s13414-011-0150-6 võ, m. l. h., & schneider, j. m. (2010). a glimpse is not a glimpse: differential processing of flashed scene previews leads to differential target search benefits. visual cognition, 18, 171-200. https://doi.org/10.1080/13506280802547901 werner, s., & thies, b. (2000). is" change blindness" attenuated by domain-specific expertise? an expertnovices comparison of change detection in football images. visual cognition, 7, 163-173. https://doi.org/10.1080/135062800394748 wolfe, j. m., võ, m. l.-h., evans, k. k., & greene, m. r. (2011). visual search in scenes involves selective and non-selective pathways. trends in cognitive sciences, 15, 77-84. https://doi.org/10.1016/j.tics.2010.12.001 yarbus, a. (1967). eye movements and vision. new york: plenum press. https://doi.org/10.1007/978-14899-5379-7 zelinsky, g. j., & schmidt, j. (2009). an effect of referential scene constraint on search implies scene segmentation. visual cognition, 17, 1004-1028. https://doi.org/10.1080/13506280902764315 https://doi.org/10.1093/oxfordhb/9780199539789.013.0029 https://doi.org/10.1080/17470210902816461 https://doi.org/10.1080/14640748108400798 https://doi.org/10.3758/pbr.17.6.834 https://doi.org/10.1111/j.1467-9280.2008.02243.x https://doi.org/10.1111/j.1467-9280.2008.02243.x https://doi.org/10.1111/j.1467-9280.1994.tb00500.x https://doi.org/10.1111/j.1467-9280.1994.tb00500.x https://doi.org/10.1037/0096-1523.26.1.379 https://doi.org/10.1037/0096-1523.26.1.379 https://doi.org/10.3758/bf03199899 https://doi.org/10.14786/flr.v5i3.256 https://doi.org/10.1080/13506280902869213 https://doi.org/10.1037/0033-295x.113.4.766 https://doi.org/10.1167/10.3.14 https://doi.org/10.3758/s13414-011-0150-6 https://doi.org/10.1080/13506280802547901 https://doi.org/10.1080/135062800394748 https://doi.org/10.1016/j.tics.2010.12.001 https://doi.org/10.1007/978-1-4899-5379-7 https://doi.org/10.1007/978-1-4899-5379-7 https://doi.org/10.1080/13506280902764315 microsoft word fryer_publication.docx frontline learning research vol.5 no. 4 (2017) 61-75 issn 2295-3159 corresponding author email: fryer@hku.hk doi: https://doi.org/10.14786/flr.v5i4.301 one more reason to learn a new language: testing academic self-efficacy transfer at junior high school luke k. fryera , w. l. quint oga-baldwinb athe university of hong kong, hong kong bwaseda university, japan article received 28 april 2017/ revised 10 october / accepted 28 october / available online 15 november abstract self-efficacy is an essential source of motivation for learning. while considerable research has theorised and examined the how and why of self-efficacy in a single domain of study, longitudinal research has not yet tested how self-efficacy might generalise or transfer between subjects such as mathematics, native and foreign language studies. the current study examined academic self-efficacy (two measurements 10 months apart) in three subjects (mathematics, native language, and foreign language) across students’ first year at junior high school. two studies were conducted, each including three schools (study-a: n=480; study-b: n=398) to support a test and retest of self-efficacy differences and interrelationships across the year of study. analyses of self-efficacy change presented a general pattern of significant, small declines in students’ selfefficacy for all three subjects. longitudinal latent analyses indicated a consistent moderate predictive effect from foreign language self-efficacy to native language self-efficacy. the pattern of declines, while consistent with research in western contexts is a source of concern. the transfer of selfefficacy from foreign to native language learning has potential educational and broader psychological implications. keywords: self-efficacy; mathematics; native language; foreign language; sem; longitudinal; junior high school fryer et oga-baldwin 62 | f l r 1. introduction self–efficacy is an essential individual difference for students at any level of education (honicke & broadbent, 2016). it mediates and is mediated by constructs that meta-analyses have established as powerful factors for achievement within formal education (e.g., teacher feedback and grade goals; hattie, 2009). ability beliefs have been a topic of attention within junior high school education (e.g., p. chen & zimmerman, 2007; pajares & graham, 1999) due in large part to the difficult transitionary (wigfield, eccles, maciver, reuman, & midgley, 1991) and developmental issues (e.g., eccles et al., 1993; roeser, eccles, & sameroff, 2000) that surround this period of formal education. the majority of past research in this area has been undertaken in western contexts (notable exceptions e.g., bong, 2001), raising questions about other educational systems. furthermore, while substantial research has examined (schunk & meece, 2006; usher & pajares, 2008) and sought to intervene (e.g., turner, & lapan, 2005; ramdass & zimmerman, 2008; barber, et al. 2015) in adolescent students’ domain specific self-efficacy, cross-domain self-efficacy research has yet to proceed further than cross-sectional examination (bong, 2001). while japanese elementary school education has been the subject of substantial theoretical and empirical research (cave, 2007; lewis, 1995; oga-baldwin, nakata, parker, & ryan, 2017; etc.), less is known about the student experience during and in the transition to japanese junior high school. this gap in our understanding is important at the moment given the on-going national curriculum reforms aimed at preparing japanese elementary school students for learning english as a foreign language during junior high school (mext, 2008, 2016). japan’s strong literacy and mathematics skills (oecd, 2013; 2016), are an area of national pride (aoki, 2016), with continued emphasis on reading and calculation ability. at the same time, early foreign language education is sometimes viewed with suspicion (otsu, 2004). the combination of these attitudes and current national reform efforts (mext, 2016), calls for an examination of the relationships between students’ beliefs about learning core subjects (mathematics, native and foreign languages) with an eye towards understanding how attitudes for one subject may affect another. until 2011, teaching english as a foreign language began in junior high school (mext, 2008); after this date it was taught in fifth and sixth year of elementary school, and treated as an additional field of study similar to art and music, but with no tests, grades, or assessments. from 2020, english will become a school subject in all elementary schools, with a standardised set of exams before students enter secondary school. it is thus particularly important that research in this area focus on critical periods of study like the transition to middle school prior to the implementation of the new curriculum. first year at secondary school in japan is therefore an ideal context for such a longitudinal investigation. results from such an examination might inform educators nationally and internationally regarding links between beliefs about learning core subjects. findings might also support national curriculum efforts in nearby countries seeking to strike a balance between these important subjects while maintaining their traditional strengths. the current research therefore undertook two studies to test changes and cross-domain relationships in students’ self-efficacy for learning mathematics, native and foreign languages during first-year junior high school. following this evaluation, the latent longitudinal interconnections between these domains of selfefficacy were tested. in these studies, we aimed to address both broad questions regarding potential declines in first-year junior high school students’ self-efficacy beliefs in the context of japan, as well as more generally assess the potential for transfer between self-efficacy for mathematics, native and foreign language studies during this critical transition to secondary school. 1.1 self-efficacy at school self-efficacy is one of the most researched non-cognitive factors within education for two important reasons. the first is its theoretical lineage stretching back to pivotal research regarding the role of expectancy within outcomes, and then being defined (bandura, 1977), elaborated (bandura, 1986), and fryer et oga-baldwin 63 | f l r rigorously tested across a range of domains (bandura, 1997). the second reason for its salient role within educational research is its powerful function within both achievement outcomes and its essential covariates (e.g., goal development and pursuit, persistence, effort and self-regulation; see bandura, 1993; zimmerman, 2000). bandura defined self-efficacy as a person’s judgement of their ability to organise and execute courses of action to attain a designated goal (bandura, 1997). perhaps one of the greatest challenges to a child’s academic self-efficacy is the transition to junior high school. it is essential that any discussion of this topic begin by acknowledging the less tangible, but still relevant, developmental issues children experience during adolescence both at home (steinberg, 1987) and at school (eccles, 1999; wigfield et al., 1991). in addition to these factors, junior high school learning environments present students with an array of social and academic challenges for students’ perceived selfefficacy. within their stage-environment fit programme of research, eccles and colleagues (e.g., 1993) noted a number of important differences between secondary and elementary school environments which might partially explain declines in students’ motivation to learn. for example, secondary school environments often exhibit more teacher control and less opportunity for students to make choices about their learning. furthermore, public evaluations and ability-groupings are also increasingly common. an additional issue for children transitioning to secondary school is that a major source of students’ academic self-efficacy comes from the grades they receive from their teachers (bandura, 1993). secondary school teachers use higher standards to assess students’ work than their elementary school counterparts (eccles & midgley, 1989). this issue combined with the lower self-efficacy that many teachers in junior high school report (important for supporting students’ adjustment to new learning environments; zee & koomen, 2016), relative to their elementary school counterparts across a broad spectrum of factors (eccles et al., 1993), means that students face numerous threats to their academic self-efficacy on entering junior high school. research with self-efficacy has generally focused on specific subjects such as mathematics or native language competencies (e.g., pajares & graham, 1999; pajares & miller, 1994). researchers have worked with specific domains of self-efficacy, even measuring self-efficacy for a broad range of junior high school subjects cross-sectionally (bong, 2001). while less often researched, the role of self-efficacy within foreign language learning is also well recognized. studies in language learning contexts have indicated that selfefficacy has important predictive relationships with achievement (mills, pajares, & herron, 2007), students’ attributions for success (hsieh & schallert, 2008), and interest in classroom tasks (fryer, ainley, & thompson, 2016). connecting this body of research across students’ studies, bong (2001) found support for strong relationships between self-efficacy for foreign language, math and students’ native language studies. consistent with empirical connections discussed to this point, bandura (1997) reviewed the conditions wherein competence judgements might generalise across activities. pajares (1997, p.3) added some detail to this principle of generalisation or transfer across domains: “i.e., the extent to which they relate to, or transfer across, different performance tasks or domains. for example, when differing tasks require similar subskills, judgments of capability to demonstrate the requisite subskills should predict the differing outcomes” and specifically in school contexts, “in school, students’ mathematics and verbal self-efficacy may generalize if the skills for each subject have been adequately taught and developed by a competent teacher.” similarly, schunk and dibenedetto (2016, p.43) have discussed how this transfer might take place within formal education, “educational conditions may foster general self-efficacy because school curricula are structured to promote positive transfer—new learning builds on prior learning.” they are quick to highlight, however, that “evidence of self-efficacy generality does not refute the domain-specific conceptualisation of self-efficacy.” both pajares (1997) and schunk and dibenedetto (2016) call for research into how and in what situations transfer of self-efficacy takes place. such research has the propensity to add to our understanding of both the development of self-efficacy beliefs and the competences/skills they are connected to. fryer et oga-baldwin 64 | f l r 1.2 japanese education and language learning empirical research and observational commentaries have noted the humanistic nature of japanese elementary schools (cave, 2007; oga-baldwin et al., 2017), while secondary schools are notably more rigid in their use of social control (cave, 2016; oga-baldwin & fryer, 2018). consistent with longstanding evidence from research within western education (e.g., eccles et al. 1993), students in this setting often suffer a general decrease in quality of motivation to learn beginning in junior high school (nishimura & sakurai, 2017), culminating for some in despondent attitudes toward learning at the end of high school (berwick & ross, 1989; sakai & kikuchi, 2009) and during university (fryer, et al., 2013). at the same time, while students’ motivation generally decreases throughout secondary school, japanese students’ performance on international tests of literacy and mathematics are consistently strong (oecd, 2016a). we may attribute some of this success to the comparatively high standards set in japanese schools (oga-baldwin & fryer, 2018), a feature common to most of the top performing asian countries (oecd, 2016b). while japan consistently ranks among the top performers for math, science, and reading (oecd, 2016a), students here also maintain a poor-to-middling rank as learners of foreign languages (education first, 2016). as mentioned previously, current policy is focused on improving the quality of foreign language education through earlier introduction (mext, 2008, 2016). at the same time, schools are now expected to promote high quality motivation to learn english (oga-baldwin & nakata, 2014). in practice, teachers are now expected to pay attention to both spoken aspects of the language as well as academic test preparation. interestingly, one of the explicit goals in the current curriculum is to further develop a greater understanding of japanese by learning new communication skills and strategies (mext, 2008). there is evidence to suggest that both positive and negative transfer between native and foreign language (both directions) takes place (cook, 2003), both contrary and consistent with anecdotal concerns regarding the negative implications of adding foreign languages to school curricula (otsu, 2004). when this transfer is negative, it is referred to as interference. one example of interference is when pronunciation or vocabulary knowledge from a previous language interferes with new pronounciation (see cheng & zhang, 2015, for a recent discussion and investigation). an example of positive transfer is when learning a foreign language’s vocabulary supports a deeper understanding of one’s native language vocabulary (e.g., cunningham & graham, 2000); this principle extends to the development of reading skills (gebauer, zaunbauer, & möller, 2013) in both directions, but strongest from foreign to native language reading fluency and comprehension. gebauer and colleagues (2013) concluded that the salient transfer from foreign to native language skills were a confirmation of cummins (1981, 1984) interdependence hypothesis. this theory suggests reciprocal connections between both languages, but stronger transfer from foreign to native, resulting partly from increased foreign language learning outside of formal school. given the interdependence of skills development and self-efficacy beliefs, it is reasonable to hypothesise similar relationships for perceived self-efficacy to learn foreign and native languages. transfer is not restricted to foreign language, native language, and mathematics; it is a complex area of research essential to understanding how learning across domains and contexts are connected (see belenky & schalk, 2014). better knowledge of transfer may indicate ways to support integrative knowledge across the formal education experience (mckeough, lupart, & marini, 2013). while support exists for hypothetical transfer between students’ self-efficacy for different school subjects, it has not yet been substantively addressed by longitudinal research to the authors’ knowledge. seminal cross-sectional research in korea (bong, 2001) observed strong correlations between mathematics, korean, and english, suggesting important connections. research is now needed to test whether transfer across academic domains takes place through students’ self-efficacy for learning mathematics, native and foreign language. fryer et oga-baldwin 65 | f l r 1.3 the current study the current research examined the development of first-year japanese junior high school students’ self-efficacy for learning mathematics, native and foreign language across a year of study. we were interested in both testing the quantitative change in students’ self-efficacy for these subjects and their potential cross-sectional and longitudinal relationships. these issues were addressed by conducting a pair of longitudinal studies at two sets of three japanese junior high schools across one academic year. based on the substantial existing research regarding the difficult transition many students experience adapting to junior high school (wigfield, et al. 1991), decreases in students’ self-efficacy were expected across their first year at junior high school (hypothesis-1). strong auto-lagged relationships were expected for self-efficacy within each subject, as prior self-efficacy plays an important role in future self-efficacy development in a domain of study (hypothesis-2). finally, anecdotal theorising in the japanese context suggests negative relationships between the increased foreign language studies (and therefore self-efficacy) and both students’ mathematics and native language academic self-efficacy (otsu, 2004). however, given the substantial covariance of the verbal subjects (japanese and foreign language studies) cross-sectionally (bong, 2001), evidence of potential positive skills transfer (foreign to native language as well as native to foreign; e.g., gebauer, zaunbauer, & möller, 2013), theory within the language education domain of crosslinguistic influence (cummins, 1981, 1984), and socio-cognitive theorising regarding the potential generalising (bandura, 1993) or transfer (pajares, 1997) of self-efficacy across domains, we expect positive cross-sectional and longitudinal connections between the self-efficacies for learning the two languages (hypothesis-3). figure 1 presents the fully-forward longitudinal latent model we propose to test. figure 1. longitudinal math, japanese and english self-efficacy fully-forward model to be tested note: t1 = time-1 (may, 2016); t2 = time-2 (march, 2017) fryer et oga-baldwin 66 | f l r 2. methods 2.1 sample and context for the current research, two studies each collected data from first-year students at three junior high schools in western japan: study a n=480 (female=236) and study b n=398 (female=186). all students were 12-13 years old. research participation was coordinated through meetings with the board of education, school principals, and teachers. student participation was voluntary, and parents were notified regarding student involvement in the study. six schools located in two rural-suburban districts agreed to participate in the study. the municipalities were largely representative of japan as a whole (japan statistics bureau, 2017). surveys for both study a and b were administered during the 2016-2017 school year, two times and 10 months apart, at the beginning and end of the academic year. ethical oversight was included in the review process for the jsps grant-in-aid for scientific research. 2.2 instrumentation in the current study, self-efficacy for each subject of mathematics, japanese and english as a foreign language was measured by five items from the patterns of adaptive learning inventory (midgley et al., 2000). students self-reported their agreement with scale items (e.g., i am certain i can master the skills taught in class this year; even if the work is hard, i can learn it) across the likert formatted scale, from “i don’t think so at all” (1) to “i think so” (6). the scale for each subject had a stem to differentiate the items (e.g., in my english class this year...; in my japanese class this year...; in my mathematics class this year...). 2.3 analysis missing data were assessed prior to analyses (<1%). missing data were coded and then accounted for by full information maximum likelihood estimation within mplus 7.2 (muthén & muthén, 1998-2015). for both studies in the current research, analyses began with confirmatory factor analysis of each of the three domains of self-efficacy. fit for the structural equation models were assessed with incremental (comparative fit index; cfi) and absolute (root mean square error of approximation; rmsea) measures model fit. acceptable/good fit were based on cfi values above .90/.95 (mcdonald & marsh, 1990) and rmsea values below .05/.08 (browne & cudeck, 1992). analyses then proceeded with reliability (raykov's rho; raykov, 2009) and then invariance estimations. based on chen (2007), invariance between time-1 and time-2 measures of self-efficacy in the same domain was perceived to be tenable if rmsea increased by less than .015 and cfi did not change more than .01. latent inter-correlations between all variables and mean differences between time-1 and time-2 were calculated for both samples from the two studies. at the final stage, fully-forward structural equation modelling was conducted on both samples (figure 1). beta (β) coefficients were relied on to interpret structural equation modelling findings. we followed keith’s (2015) suggested guidelines for the interpretation of beta coefficients in research on influences on learning, β below 0.05 are interpreted as “too small to be considered meaningful”; those above 0.05 are considered “small but meaningful”; those above 0.10 are considered “moderate”; and those above 0.25 are considered “large”. fryer et oga-baldwin 67 | f l r 3. results 3.1 descriptive findings and analysis of changes in self-efficacy perceptions table 1 presents the latent inter-correlations between all variables in each study, a and b. the correlations were predictably strong, but presented no threat of multi-collinearity (i.e., no r > .90; tabachnick & fidell, 2007). t-tests were conducted (time1-time2) for both studies, for each of mathematics, foreign and native language (bonferroni adjustments undertaken). both studies presented significant (p < .01), small (.01 > d < .2; cohen, 1988; sawilowsky & shlomo, 2003; effect-size calculated using lakens, 2013) decreases in self-efficacy for each mathematics, japanese and english across the yearlong studies. all scales demonstrated excellent reliability in both studies (> .90; devellis, 2012). table 1. latent correlations, means, sds and raykov’s rho reliability coefficient for study a and b t1 english se t1 math se t1 japanese se t2 english se t2 math se t2 japanese se t1 english se t1 math se .71/.72 t1 japanese se .73/.73 .72/.68 t2 english se .62/.59 .48/.44 .50/.47 t2 math se .46/.47 .63/.58 .44/.47 .64/71 t2 japanese se .52/.50 .48/40 .58/.55 .69/.74 .73/72 means 3.83/3.75 3.85/3.88 3.84/3.80 3.78/3.53 3.64/3.49 3.70/3.57 sds 1.24/1.18 1.28/1.31 1.20/1.29 1.23/1.18 1.28/1.26 1.21/1.23 raykov’s rho .91/.91 .93/.94 .93/.94 .90/.92 .93/.94 .94/.94 note: study a / study b; se = self-efficacy; t1 = time-1 (may, 2016); t2 = time-2 (march, 2017) 3.2 construct validity and invariance testing for all latent modelling, error covariances were included for items with identical wording crosssectionally. fit (table 2) for configural and longitudinal models was acceptable (rmsea<.08, cfi>.90). invariance tests for each domain of self-efficacy, for both studies, suggested that the assumption of invariance was tenable (table 2). fryer et oga-baldwin 68 | f l r table 2. fit for configural, invariance and fully-forward models configural model test japanese invariance math invariance english invariance fully-forward sem study a cfi .96 .96 .96 .96 .96 rmsea .060 .061 .061 .061 .060 chi-square 984.224 (352) 989.197 (356) 985.838 (356) 989.198 (356) 984.22 (352) study b cfi .97 .97 .97 .97 .97 rmsea .050 .050 .055 .055 .060 chi-square 775.263 (352) 776.848 (356) 779.56 (356) 778.27 (356) 775.26 (352) 3.3 longitudinal internal and external tests the longitudinal fully-forward test of the model (figure 1) with samples from study a and study b presented consistent results. strong auto-lagged βs for all three domains of self-efficacy were observed. a single significant cross-lagged β was found in both studies, from time-1 english to time-2 japanese. figure 2 presents the finalised model with all significant and non-significant βs. figure 2. model test (a) and replication (b) of mathematics, japanese, english self-efficacy cross-lagged test note: no small βs were observed. fryer et oga-baldwin 69 | f l r 4. discussion the current studies examined the longitudinal differences and relationships between first-year junior high school students’ self-efficacy for learning mathematics, japanese and a foreign language across one academic year. analysis of self-efficacy change for each subject in both studies confirmed significant small decreases for all three subjects (hypothesis-1). large auto-regressive βs were observed from all time-1 to time-2 variables (hypothesis-2). strong correlations were present between variables cross-sectionally and longitudinally. the only significant cross-lagged βs presented by the latent longitudinal model were from foreign language self-efficacy to native language self-efficacy (hypothesis-3). this final result builds on past cross-sectional findings (bong, 2001), which, like the present study, indicated substantial cross-sectional relationships between mathematics, native and foreign language academic self-efficacy. these crosssectional findings are narrowed to a singular cross-domain relationship, which suggests transfer from foreign language academic self-efficacy to native language academic self-efficacy. all of the findings (correlational and changes in self-efficacy beliefs for each subject over time) were consistent across study a and study b, supporting their reliability in the context of japanese junior high school. 4.1 theoretical implications while no comparable longitudinal study of academic self-efficacy has been undertaken for the current study to build on, bong’s (2001) study examined similar sources of self-efficacy cross-sectionally. in her study, however, she organised them into verbal and quantitative constructs for an internal/external frame of reference test. while such an organisation is consistent with general theory regarding ability beliefs, the current results suggest the connection between related (verbal) areas of self-efficacy might have a more nuanced relationship than simply converging together on a higher-order factor. the current studies’ results suggest that self-efficacy for a new language supports native language self-efficacy. essentially, these findings indicate that improving self-efficacy in one domain can support selfefficacy for learning in another close domain (e.g., languages). in the current study, consistent with language learning theory (cummins, 1981; 1984) and recent reading skills research (gebauer, et al., 2013) the crossdomain transfer was significant for perceived foreign language learning self-efficacy to perceived native language learning self-efficacy. it is possible, consistent with cummins theory of language skills transfer, that the dominant connections from foreign to native language self-efficacy might be due to supplementary foreign language studies undertaken outside of school. it might also be the partial result of recent reforms to elementary school english education aimed at stimulating students’ interest in foreign language studies, which can in turn support self-efficacy (see ainley, buckley, & chan, 2009). replication of the current studies’ findings in another context might provide an avenue for overcoming the diminishing returns of prior achievement for future self-efficacy (i.e., wood & bandura, 1989). the decline observed in students’ self-efficacy across three core subjects, in two studies across six junior high schools, is consistent with u.s. findings highlighting the complex transition to junior high school. the current results suggest that, consistent with western findings (eccles, 1999; eccles & midgley, 1993), students experienced a decline in important motivations to learn during the transition to secondary school. further research is necessary to determine what aspects of the learning environment influence selfefficacy and whether this decline persists or is only present during students’ first year. 4.2 practical implications first, it is reasonable to assume that many japanese junior high school students experience widespread, small declines in self-efficacy for learning during their first year. several instructional (zimmerman, 2000) and parenting (bandura, 1993) strategies are available to support the mitigation of this fryer et oga-baldwin 70 | f l r decline. in the classroom, modelling cognitive strategies, helping students set proximal goals and then verbally supporting said goals, and providing frequent feedback on learning activities can play a role in supporting students’ self-efficacy. second, rather than playing a role in the decline of core subjects like native language and mathematics, the current study’s latent lagged results suggest that studying a foreign language can have either a non-significant predictive effect (i.e., mathematics) or a substantive predictive positive effect (i.e., native languages) on students’ academic self-efficacy. as reviewed, perceptions of selfefficacy are essential predictors of achievement across a broad array of domains, suggesting that enhancing foreign language self-efficacy is one additional means of supporting native language self-efficacy and thereby native language skills. our findings suggest that perceived self-efficacy transfer between subjects that rely on similar skills, much like skills and competences themselves, is possible. specifically, the longitudinal findings from both studies suggest that the transfer is likely to be one way: from foreign to native language self-efficacy. it might therefore be worthwhile for educators’ and, more importantly, national curriculum developers’ to consider whether students are coming to secondary school with robust self-efficacy for their foreign language studies. questions regarding the state of students’ self-efficacy for foreign language studies when entering secondary school may also have implications for other countries as well. in an interesting and potentially relevant side-note, many countries with strong foreign language skills (notably countries in northern europe; education first, 2016) also score highly on the pisa native language reading test (oecd, 2016a). 5. future directions and limitations further testing at other levels of education in japan and internationally are called for. it is important for future studies to control for prior achievement as well as self-efficacy. within the current research programme, latent curve analysis of the full dataset for the current studies in two years (five points) will clarify whether the present results are a part of longor short-term trends. longitudinal person-centred analyses are necessary to move beyond mean-based estimates and examine the broader self-efficacy developmental pattern. finally, quasi-experimental interventions are necessary to assess any attempt at mitigating the declining self-efficacy observed in the current study. 6. conclusion substantial research has examined and sought to intervene in students’ self-efficacy during junior high school. however, research examining self-efficacy across domains to this point has been constrained to cross-sectional work (i.e., bong, 2001). in the present longitudinal research, we therefore undertook to examine the interplay of academic self-efficacy across three domains of study. the current study presents two preliminary conclusions for future research to test and build on. firstly, students experienced broad declines in academic self-efficacy across their first year of junior high school. second, enhanced selfefficacy for learning a new language predicted the development of greater native language learning selfefficacy. this finding has potentially significant implications for early foreign language learning and selfefficacy development in schools more broadly. in practice, given the widely acknowledged relationship between self-efficacy and both persistence and achievement (bandura, 1997), anything schools can do to enhance students’ self-efficacy for learning a foreign language will not only support the development of a population competent in a second language, but also support students’ confidence in their native language learning. fryer et oga-baldwin 71 | f l r we encourage examination and replication of our study. if our findings have clear external validity, then japan and countries in similar situations should consider embarking on truly ambitious foreign language education paths, ensuring students arrive at junior high school with confidence in their second/foreign language. the current pair of studies suggests that we can do this knowing that the experience might, through self-efficacy transfer, also support robust native language skills. the current studies’ findings are in alignment with theory (cummins, 1981, 1984) and empirical research (gebauer et al., 2013) in the context of second/foreign language learning. they are also consistent with theorising about how self-efficacy might generalise or transfer between sub-skills and domains of study (bandura, 1997; pajares, 1997). the present findings therefore have potential implications for how we support student learning in the specific context of native and foreign language studies. furthermore, we hope these initial findings are enough to get the field thinking about the possibilities such connections create and encourage tests in both a broader range of domains and contexts. keypoints two longitudinal studies examined academic self-efficacy for three subjects during first year at japanese junior high school students experienced widespread declines in academic self-efficacy during their first year foreign language academic self-efficacy supported native language academic self-efficacy across the first year of junior high school analysis of change and latent modelling results were consistent across the two studies conducted acknowledgments this research was supported by jsps grant-in-aid for scientific research (c) 16k02924. fryer et oga-baldwin 72 | f l r 7. references ainley, m., buckley, s., & chan, j. (2009). interest and efficacy beliefs in self-regulated learning. in m. wosnitza, s. karabenick, a. efklides, & p. nenniger (eds.), contemporary motivation research: from global to local perspectives (pp. 207-228). aoki, m. (2016, december 6). japan’s 15-year-olds perform well in pisa global academic survey. the japan times. http://www.japantimes.co.jp/news/2016/12/06/national/japans-15-year-olds-performwell-pisa-global-academic-survey/ bandura, a. (1977). self-efficacy toward a unifying theory of behavioral change. psychological review, 84, 191-215. doi:10.1016/0146-6402(78)90002-4 bandura, a. (1986). social foundations of thought and action: a social cognitive theory. new york: pearson. bandura, a. (1993). perceived self-efficacy in cognitive development and functioning. educational psychologist, 28, 117–148. doi:10.1207/s15326985ep2802_3 bandura, a. (1997). self-efficacy: the exercise of control. new york: freeman. berwick, r., & ross, s. (1989). motivation after matriculation: are japanese learners of english still alive after exam hell? jalt journal, 11, 193–210. belenky, d. m., & schalk, l. (2014). the effects of idealized and grounded materials on learning, transfer, and interest: an organizing framework for categorizing external knowledge representations. educational psychology review, 26, 27-50. doi: 10.1007/s10648-014-9251-9 bong, m. (2001). between-and within-domain relations of academic motivation among middle and high school students: self-efficacy, task value, and achievement goals. journal of educational psychology, 93, 23-34. doi: 10.1037/0022-0663.93.1.23 browne, m., & cudeck, r. (1992). alternative ways of assessing model fit. sociological methods & research, 21, 230-258. doi:10.1177/0049124192021002005 cave, p. (2007). primary school in japan: self, individuality and learning in elementary education. new york: routledge. cave, p. (2016). schooling selves: autonomy, interdependence, and reform in japanese junior high education. chicago: university of chicago press. chen, f. f. (2007). sensitivity of goodness of fit indexes to lack of measurement invariance. structural equation modeling, 14, 464-504. doi: 10.1080/10705510701301834 chen, p., & zimmerman, b. (2007). a cross-national comparison study on the accuracy of self-efficacy beliefs of middle-school mathematics students. journal of experimental education, 75(3), 221-244. doi:10.3200/jexe.75.3.221-244 cheng, b., & zhang, y. (2015). syllable structure universals and native language interference in second language perception and production: positional asymmetry and perceptual links to accentedness. frontiers in psychology, 6. doi:10.3389/fpsyg.2015.01801. cohen, j. (1988). statistical power analysis for the behavioral sciences. (2nd ed.) hillsdale, nj: erlbaum. cook, v. (2003). effects of the second language on the first (vol. 3): multilingual matters. cummins, j. p. (1983). heritage language education: a literature review. toronto: ministry of education.
 cummins, j. p. (1984). bilingualism and special education: issues in assessment and pedagogy. clevedon, uk: multilingual matters.
 cunningham, t. h., & graham, c. r. (2000). increasing native english vocabulary recognition through spanish immersion: cognate transfer from foreign to first language. journal of educational psychology, 92, 37-49. doi:10.1037//0022-0663.92.1.37 devellis, r. f. (2012). scale development: theory and application (3rd ed.). thousand oaks, ca: sage. eccles, j. s. (1999). the development of children ages 6 to 14. future of children, 9(2), 30-44. doi:10.2307/1602703 eccles, j. s., & midgley, c. (1989). stage/environment fit: developmentally appropriate classrooms for early adolescents. in r. e. ames & c. ames (eds.), research on motivation in education (vol. 3, pp. 139-186). san diego, ca: academic press. fryer et oga-baldwin 73 | f l r eccles, j. s., midgley, c., wigfield, a., buchanan, c. m., reuman, d., flanagan, c., & mac iver, d. (1993). development during adolescence: the impact of stage-environment fit on young adolescents' experiences in schools and in families. american psychologist, 48, 90-101. doi:10.1037/0003066x.48.2.90 education first. (2016). ef english proficiency index 2016: japan. available from: http://www.ef.co.uk/epi/regions/asia/japan/
 fryer, l. k., ainley, m., & thompson, a. (2016). modelling the links between students' interest in a domain, the tasks they experience and their interest in a course: isn't interest what university is all about? learning and individual differences, 50, 157-165. doi:10.1016/j.lindif.2016.08.011 fryer. l. k., carter, p., ozono, s., nakao, k., & anderson, c. j. (2013) instrumental reasons for studying in compulsory english courses: i didn’t come to university to study english, so why should i? innovation in language learning and teaching. 8, 239-256 doi: 10.1080/17501229.2013.835314 gebauer, s. k., zaunbauer, a. c. m., & möller, j. (2013). cross-language transfer in english immersion programs in germany: reading comprehension and reading fluency. contemporary educational psychology, 38(1), 64-74. doi:10.1016/j.cedpsych.2012.09.002 hattie, j. c. (2009). visible learning: a synthesis of over 800 meta-analyses relating to achievement. london & new york: routledge, taylor & francis. hsieh, p.-h. p., & schallert, d. l. (2008). implications from self-efficacy and attribution theories for an understanding of undergraduates’ motivation in a foreign language course. contemporary educational psychology, 33, 513-532. doi:10.1016/j.cedpsych.2008.01.003 honicke, t., & broadbent, j. (2016). the influence of academic self-efficacy on academic performance: a systematic review. educational research review, 17, 63–84. http://doi.org/10.1016/j.edurev.2015.11.002 japan statistics bureau (2017). japan statistical yearbook 2017. available from http://www.stat.go.jp/english/data/nenkan/index.htm. keith, t. z. (2015). multiple regression and beyond: an introduction to multiple regression and structural equation modelling (2nd ed.). new york: routledge. lakens, d. (2013). calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and anovas. frontiers in psychology, 4:863. doi:10.3389/fpsyg.2013.00863 lewis, c. c. (1995). educating hearts and minds: reflections on japanese preschool and elementary education. new york: cambridge university press. mcdonald, r. p., & marsh, h. w. (1990). choosing a multivariate model noncentrality and goodness of fit. psychological bulletin, 107, 247-255. doi:10.1037/0033-2909.107.2.247 mckeough, a., lupart, j. l., & marini, a. (2013). teaching for transfer: fostering generalization in learning: routledge. mext. (2008). explanatory commentary for the elementary school curriculum guidelines: foreign language activities. tokyo: kyoiku shuppan. mext. (2016). gaikokugo waakingu guruupu-niokeru torimatome (an) betten shiryo. retrieved from http://www.mext.go.jp/b_menu/shingi/chukyo/chukyo3/058/siryo/__icsfiles/afieldfile/2016/09/14/137 3448_1.pdf midgley, c., maehr, m. l., hruda, l. z., anderman, e., anderman, l., freeman, k. e., & urdan, t. (2000). manual for the patterns of adaptive learning scales. ann arbor, 1001, 48109-41259. mills, n., pajares, f., & herron, c. (2007). self-efficacy of college intermediate french students: relation to achievement and motivation. language learning, 417-442. doi: 10.1111/j.14679922.2007.00421.x/full nishimura, t., & sakurai, s. (2017). journal of applied developmental psychology. journal of applied developmental psychology, 48, 42–48. doi:10.1016/j.appdev.2016.11.004 oecd. (2016a). pisa 2015 results (volume i): excellence and equity in education, oecd publishing, paris. doi:10.1787/9789264266490-en oecd. (2016b). pisa 2015 results (volume ii): policies and practices for successful schools, oecd publishing, paris. doi:10.1787/9789264267510-en fryer et oga-baldwin 74 | f l r oga-baldwin, w. l. q., & fryer, l. k. (2018). growing up in the walled garden: motivation, engagment, and the japanese educational experience. in g. a. d. liem & t. s. hong (eds.), student motivation, engagement, and growth: asian insights. routledge. oga-baldwin, w. l. q., & nakata, y. (2017). engagement, gender, and motivation: a predictive model for japanese young language learners. system, 65, 151–163. doi:10.1016/j.system.2017.01.011 oga-baldwin, w. l. q., nakata, y., parker, p., & ryan, r. m. (2017). contemporary educational psychology. contemporary educational psychology, 49, 140–150. doi: 10.1016/j.cedpsych.2017.01.010 otsu, y. (2004). is elementary school english necessary? tokyo: keio gijuku daigaku shuppan kai. pajares, f., & graham, l. (1999). self-efficacy, motivation constructs, and mathematics performance of entering middle school students. contemporary educational psychology, 24, 124-139. doi: 10.1006/ceps.1998.0991 pajares, f., & miller, m. d. (1994). role of self-efficacy and self-concept beliefs in mathematical problem solving: a path analysis. journal of educational psychology, 86, 193-203. doi: 10.1037/00220663.86.2.193 ramdass, d., & zimmerman, b. j. (2008). effects of self-correction strategy training on middle school students' self-efficacy, self-evaluation, and mathematics division learning. journal of advanced academics, 20(1), 18-41. doi: 10.4219/jaa-2008-869 raykov, t. (2009). evaluation of scale reliability for unidimensional measures using latent variable modeling. measurement and evaluation in counceling and devlelopment, 42(3), 223-232. doi:10.1177/0748175609344096 roeser, r. w., eccles, j. s., & sameroff, a. j. (2000). school as a context of early adolescents' academic and social-emotional development: a summary of research findings. elementary school journal, 100, 443-471. doi: 10.1086/499650 sakai, h., & kikuchi, k. (2009). an analysis of demotivators in the efl classroom. system, 37(1), 57–69. doi: 10.1016/j.system.2008.09.005 steinberg, l. (1987). the impact of puberty on family relations: effects of pubertal status and pubertal timing. developmental psychology, 23, 451-460. doi: 10.1037/0012-1649.23.3.451 sawilowsky, shlomo s. (2003). a different future for social and behavioral science research, journal of modern applied statistical methods, vol 2, 128-132. doi: 10.22237/jmasm/1051747860 schunk, d. h., & meece, j. l. (2006). self-efficacy development in adolescence. in f. pajares and t. urdan. self-efficacy beliefs of adolescents (pp. 71-96). usa: iap schunk, d. h., & dibenedetto, m. k. (2016). self-efficacy theory in education. in k. r. wentzel & d. b. miele (eds.), handbook of motivation at school (pp. 34-54). new york, ny: taylor & francis. tabachnick, b. g., & fidell, l. s. (2007). using multivariate statistics (5th ed.). boston: pearson education. taboada barber, a., buehl, m. m., kidd, j. k., sturtevant, e. g., richey nuland, l., & beck, j. (2015). reading engagement in social studies: exploring the role of a social studies literacy intervention on reading comprehension, reading self-efficacy, and engagement in middle school students with different language backgrounds. reading psychology, 36, 31-85. doi: 10.1080/02702711.2013.815140 turner, s. l., & lapan, r. t. (2005). evaluation of an intervention to increase non-traditional career interests and career-related self-efficacy among middle-school adolescents. journal of vocational behavior, 66, 516-531. doi: 10.1016:j.jvb.2004.02.005 usher, e. l., & pajares, f. (2008). sources of self-efficacy in school: critical review of the literature and future directions. review of educational research, 78, 751-796. doi: 10.3102/0034654308321456 wigfield, a., eccles, j. s., maciver, d., reuman, d. a., & midgley, c. (1991). transitions during early adolescence changes in childrens domain-specific self-perceptions and general self-esteem across the transition to junior-high-school. developmental psychology, 27, 552-565. doi: 10.1037/00121649.27.4.552 wood, r., & bandura, a. (1989). social cognitive theory of organizational management. academy of management review, 14, 361-384. doi: 10.5465/amr.1989.4279067 fryer et oga-baldwin 75 | f l r zee, m., & koomen, h. m. y. (2016). teacher self-efficacy and its effects on classroom processes, student academic adjustment, and teacher well-being. review of educational research, 86, 981-1015. doi:doi:10.3102/0034654315626801 zimmerman, b. j. (2000). self-efficacy: an essential motive to learn. contemporary educational psychology, 25, 82-91. doi:10.1006/ceps.1999.1016 codepen broda publication frontline learning research vol.8 no. 4 (2020) 52 73 issn 2295-3159 assessing the predictive nature of teacher and student writing self-regulation discrepancy michael brodaa, eric ekholma & sharon zumbrunna adepartment of foundations of education, virginia commonwealth university, usa article received 12 june 2019 / revised 9 december / accepted 4 march 2020/ available online 17 july abstract in this study, we examine the extent to which the discrepancy between teacher-reported and student-reported self-regulatory behaviours during writing were associated with students’ end-of-year writing grades after controlling for student writing ability and other demographic characteristics. results of our study, conducted with a sample of 201 middle grades students enrolled in a large, comprehensive suburban school district in the mid-atlantic u.s., suggest a significant and positive relationship between teacher discrepancy and grades, after controlling for writing ability, student self-regulation, gender, race/ethnicity, and ses. this has clear implications for the classroom, as it suggests that even after accounting for student difference in terms of ability background, and demographics, the effort that teachers perceive their students making in the fall are still associated with students’ year-end performance in their class. this represents some of the first frontline evidence of the predictive relationship between self-regulation discrepancy and student achievement in writing. keywords: literacy; self-regulation; achievement; teaching; hlm info corresponding author email: xmdbroda@vcu.edu doi: https://doi.org/10.14786/flr.v8i4.505 1. introduction at least since rosenthal and jacobson’s (1968) publication of pygmalion in the classroom, researchers, educators, and the popular media have been interested in the role that teacher expectations play in relation to student academic performance. depending upon the research examined or the person describing the research, teacher expectations may either have a substantial impact on later student achievement (often referred to as the “self-fulfilling prophecy” effect), or they may have a trivial impact (jussim & harber, 2005). as jussim and harber (2005) point out, the reality is somewhere between these two extremes; teacher expectations do seem related to student achievement in certain circumstances. in this study, we examine the relationship between teacher expectations and student grades in one circumstance the extent to which teachers’ discrepant expectations of students’ self-regulatory behaviors during writing were associated with students’ writing grades, after controlling for writing ability and student background characteristics. as the use of observational rating scales that measure students’ self-regulatory behaviors continues to increase in k-12 classrooms (duckworth & yeager, 2015), the study of how student and teacher perceptions of these behaviors relate to one another, and how they might systematically differ, is a crucial emerging area of research. 1.1 review of relevant literature 1.1.1 teacher expectations as a matter of course, teachers form expectations of their students. these expectations are not a priori good or bad, nor does it seem likely that they are eradicable. issues arise, however, when teachers hold biased or otherwise discrepant expectations about students and these discrepant expectations contribute to educational inequities (e.g. auwarter & aruguete, 2008; mckown & weinstein, 2002; rist, 1970). discrepant expectations refer to teacher overor underestimates of a student’s ability on a given attribute when compared to another indicator of that same attribute (e.g. harvey, suizzo, & jackson, 2016; hinnant et al., 2009; jussim et al., 1996). for example, if a teacher believes a student has relatively little aptitude for algebra but the student performs well on a standardized algebra assessment, this would indicate a discrepant expectation in the form of an underestimation of algebra ability. discrepant expectations can be considered quantitative variables that have both a direction and a magnitude; that is, they can be either overor underestimates and can represent different degrees of overor underestimation (madon et al., 1997). below, we briefly review the research literature on discrepant teacher expectations related to student ability and motivation. 1.1.2 expectations related to ability much of the research examining teacher expectations has focused on the relationship between teacher expectations of student ability and subsequent student academic performance across several content areas (e.g. brophy, 1983; de boer et al., 2010; hinnant et al., 2009; jussim et al. , 1996; madon et al., 1997; rosenthal & jacobson, 1968). such research often examines the “self-fulfilling prophecy” phenomenon in which teachers’ initial expectations of student academic ability influence later student academic performance by causing students to live up (or down) to the teacher’s expectations of them (jussim & harber, 2005). generally, results from this line of research suggest that teacher expectations do predict later student performance, particularly for stigmatized or vulnerable students, although the magnitude and durability of these expectation effects over time remains unclear (de boer et al., 2010; hinnant et al., 2009; jussim & harber, 2005). additionally, in an early meta-analysis of self-fulfilling prophecy studies, brophy (1983) found that stronger expectation effects emerged when teacher expectations were manipulated early in a school year compared to when they were manipulated later. that is, when teachers’ expectations of students were influenced by unreliable or irrelevant information (e.g. race/ethnicity or invalid assumptions made by others) before these teachers had an opportunity to form realistic expectations of students based off more relevant information (e.g. classroom achievement), these discrepant expectations tended to have a larger effect on subsequent student performance. teachers seem to form expectations about students based on information beyond students’ previous academic performance. more specifically, previous research has found that student characteristics such as gender, ethnicity, and socioeconomic status (ses) all contribute to teacher expectations (e.g. auwarter & aruguete, 2008; brophy, 1983; jussim & harber, 2005; mckown & weinstein, 2002; rist, 1970; tenenbaum & ruck, 2007). research examining teacher expectations differing by student gender mostly suggests that teachers tend to hold higher expectations of female students in general (e.g. de boer et al., 2010; harvey et al., 2016; hinnant et al., 2009); however, other studies have found that gender-based expectations may be contingent upon the content area in question. for example, mckown and weinstein (2002) found negative expectation effects for female students in math but not in reading. findings from studies examining differences in teacher expectations by ethnicity predominantly indicate that teachers have higher expectations for white students than for stigmatized minority students (jussim & harber, 2005; jussim et al., 1996; tenenbaum & ruck, 2007). in a series of meta-analyses on ethnicity and teacher expectations, tenenbaum and ruck (2007) found that teachers had higher academic expectations of white students than they did of african american students (d = .25) or latinx students (d = .46), although they held slightly lower academic expectations of white students than they did of asian american students (d = -.17). additionally, findings suggested that teachers tended to speak more positively about (d = .31) and to (d = .21) white students than to african american and/or latinx students. in contrast, a large-scale longitudinal study conducted by de boer and colleagues (2010) found no expectation effects based on student ethnicity. however, this study was conducted in the netherlands where the cultural climate may meaningfully differ from that of the united states. research also suggests that teachers tend to expect less of students from lower socio-economic status (ses) backgrounds than of students from more economically-advantaged backgrounds (auwarter & aruguete, 2008; jussim, et al., 1996; rist, 1970). in a landmark early study examining the relationship between teacher expectations and student ses, rist (1970) observed a group of african american children throughout their kindergarten year and during portions of their firstand second-grade years. in short, rist found that students from higher-ses backgrounds tended to align with the teacher’s idealized version of a successful student. the teacher, in turn, seemed to have higher expectations of students who matched her stereotype of how a successful student looked and behaved, and these differential expectations manifested themselves in how the teacher treated students as well as the opportunities the students were afforded in the classroom. more recent work further supports this notion that teachers often have higher expectations of students from high-ses backgrounds than they do of students from low-ses backgrounds (jussim & harber, 2005; jussim et al., 1996), with some evidence suggesting that teachers may hold disproportionately low expectations for boys from more economically-disadvantaged backgrounds (auwarter & aruguete, 2008). it is important to note that differences in teacher expectations by student gender, race, ethnicity, socioeconomic status, or other student characteristics are not necessarily indicative of bias, nor do they always produce self-fulfilling prophecies (jussim & harber, 2005). for teacher expectations to be biased, they must be systematically different about students based on certain student demographic characteristics (e.g. ethnicity) and they must be inaccurate (madon et al., 1997). previous research suggests that, even when teacher expectations for students differ according to student demographic characteristics, these expectations may accurately reflect real differences in student ability (madon et al., 1998). in such circumstances, teacher expectations may be systematically different but unbiased. 1.1.3 expectations related to other factors often, research investigating teacher expectations has focused on discrepancies between teacher perceptions of student ability and more objective measures of students’ capabilities (e.g. brophy, 1983; hinnant et al., 2009; mckown & weinstein, 2002). for example, hinnant and colleagues (2009) examined discrepancies in teachers’ perceptions of student academic ability and children’s scores on two subtests of the woodcock johnson psycho-educational battery (woodcock & johnson, 1977). some researchers have included measures of student motivation along with a standardized achievement score when estimating discrepancy in teacher expectations to control for additional student-level factors that might influence performance outcomes (e.g. de boer et al., 2010; madon et al., 1997). further, other lines of research have focused on the discrepancy between teacher-perceived motivation and student-reported motivation (e.g. harvey et al., 2016). for example, harvey and colleagues (2016) used the residuals from a model in which teacher reported student self-efficacy was regressed on student reported self-efficacy to calculate a discrepancy score. they found that this discrepancy variable predicted students’ year-end grade in math and reading even after controlling for student ability using a standardized test score. this result suggests that teacher perceptions of student competencies other than academic ability may meaningfully relate to student outcomes. given these findings, it seems possible that teachers’ discrepant expectations of students in several areas may related to indicators of student academic performance. in the current study, we investigate the extent to which discrepancies in teachers’ perceptions of student writing self-regulatory behaviors relate to student writing/english language arts (ela) grades. although research resoundingly affirms that students’ self-regulation predicts numerous measures of academic success (e.g. zimmerman, 2013), we are aware of no work that examines how discrepancies between teacher perceptions of student self-regulation and students’ perceptions of their own self-regulation might relate to student academic success. 1.1.4 grades in this study, we focus on the extent to which teachers’ perceptions of student writing self-regulatory behaviors predict students’ later writing/ela grades. therefore, we do not focus on whether students fulfill teachers’ expectations, but rather on how discrepant expectations of student self-regulatory behaviors might persist throughout the year and manifest themselves in students’ grades. according to brookhart and colleagues (2016), grades refer to “the symbols assigned to individual pieces of student work or composite measures of student performance on student report cards” (p. 804). however, the meaning of “performance” seems to vary quite a bit between teachers, and performance often represents more than standardized academic achievement or ability (bowers, 2011; brookhart et al., 2016; casillas et al., 2012; mcmillan, 2001; willingham et al., 2002). previous research on teachers’ grading practices has found that, although indicators of standardized achievement and prior grades tend to account for the largest amount of variance in teacher-assigned grades, factors such as student effort, motivation, improvement, and even behavior can also influence grading practices (e.g. bowers, 2011; casillas et al., 2012; mcmillan, 2001). in a study of more than 4,000 students, casillas and colleagues (2012) found that psychosocial and behavioral measures (e.g. motivation, self-regulation) were as useful for predicting high school gpa as were prior grades, although standardized achievement measures were the strongest predictors of gpa. this suggests that grades might provide a summary of cognitive and conative student characteristics (brookhart et al., 2016). given that grades inform many high-stakes decisions made about students, such as decisions related to grade promotion, high school graduation, and college admission, it is critical that we understand which performances and competencies grades represent. further, if grades are to serve as fair and valid indicators of student performance, we must not only understand the factors that contribute to grades but also how potential misperceptions of these factors relate to grades. failure to address this second point could lead not only to inaccurate inferences about individual students based on their grades, but perhaps also systemic educational inequities depending upon possible patterns in these misperceptions. 1.1.5 writing self-regulation in academic contexts, self-regulation refers to a proactive process or set of processes that students employ either to learn or to produce an artifact demonstrating knowledge (pintrich & de groot, 1990; winne & hadwin, 1998; zimmerman, 2008). these processes may include setting learning goals, using appropriate strategies, monitoring learning, and maintaining motivation throughout a learning task (winne & hadwin, 1998; zimmerman, 2008). although research suggests that promoting student self-regulation leads to increased achievement across several academic domains (zimmerman, 2013), self-regulation may be particularly important in writing (graham & harris, 2000; graham, harris, & mason, 2005, graham & perin, 2007; santangelo, harris, & graham, 2016). writing is often a complex, prolonged process involving multiple recursive components, including planning, drafting, and revising. given this, along with the difficulty of writing, it is no surprise that writing proficiently requires high levels of self-regulation (e.g. graham, 2018; graham & harris, 2000; hayes & flower, 1980; hayes & flower, 1986; zimmerman & riesemberg, 1997). indeed, several prominent models of writing emphasize many of the metacognitive processes implicated in effective self-regulation (hacker, 2018; hayes, 2012; hayes & flower, 1986). engaging in self-regulatory behaviors, such as goal-setting and self-monitoring, may allow students to better navigate the complexities of a given writing task and may, in turn, positively influence writing-related beliefs (graham & harris, 2000). for example, multiple studies have demonstrated a positive relationship between student self-reported writing self-regulation and self-efficacy (collie et al., 2016; ekholm et al., 2015; zimmerman & bandura, 1994; zumbrunn et al., 2016). further, a large body of research overwhelmingly shows that teaching students writing self-regulatory strategies leads to considerable improvements in writing performance (graham et al., 2005; graham et al., 2015; graham et al., 2012; graham & perin, 2007). thus, self-regulation seems to play a critical role in the writing classroom. of particular relevance for this study is how teachers perceive students’ writing self-regulatory behaviors. as mentioned previously, research on grading practices indicates that teachers often take students’ self-regulatory behaviors into account when assigning grades (brookhart et al., 2016). however, as is the case with perceptions of academic ability and self-efficacy, teachers may err in how accurately they perceive students’ self-regulatory behaviors. although some writing self-regulatory behaviors are easy to observe, others may be more covert. for example, planning may be easily observed via a graphic organizer; however, strategies such as engaging in positive self-talk or monitoring progress toward goals may be harder for teachers to accurately infer. given the difficulties inherent in observing some of these self-regulatory behaviors, it is possible that teachers unconsciously rely on (possibly irrelevant) student characteristics when making inferences about students’ self-regulation. thus, in this study, we are specifically interested in the extent to which teachers’ assessments of students’ self-regulation align with students’ own assessments, and further, the extent to which the discrepancy in those assessments might also be predictive of writing outcomes. if discrepancy measures, such as the one we explore here, are predictive of writing outcomes above and beyond typical measures of self-regulation, we believe that this measure can have real value for researchers and practitioners. 1.1.6 the present study in the present study, we investigated discrepancies between teacher and student ratings of student writing self-regulation. to extend the literature on teacher expectations, we were interested in the relations between these discrepant expectations and student demographic characteristics, end-of-year writing/ela grades, and student writing achievement. the following research questions guided the study: 1.does the average discrepancy between teacher and student perceptions of writing self-regulation differ according to student demographics, including gender, race/ethnicity, and socioeconomic status? 2. to what extent do discrepant teacher expectations of student writing self-regulation uniquely predict student grades after accounting for prior writing achievement and student demographic variables? 3. does the relationship between discrepant expectations and writing grades differ across different student demographic groups? 2. methods 2.1 participants all participants in the study were part of a three-year longitudinal study in a large, suburban school district in virginia. to qualify for inclusion in the current study, participants from the longitudinal study must have had a score on a standardized statewide writing test, which at time this study was conducted were administered to students in 5th, 8th, and 10th grade. this resulted in the inclusion of 201 students for whom we had a test score from either the 5th, 8th, or 10th grade test, as well as both teacher and student ratings of writing self-regulation and an end-of-year writing grade. both student and teacher self-regulation ratings were required to calculate our measure of discrepancy described below. the sample consisted of 91 females (45%) and 110 males (45%), with 83 (41%) students identifying as african american, 73 (36%) identifying as white, 35 (17%) identifying as latinx, and 10 (5%) identifying as another ethnicity or multiracial. additionally, 8 (4%) of these students received special education services, 1 (.50%) was classified as an english language learner (ell), 15 (7%) were classified as gifted, and 59 (29%) were classified as economically disadvantaged. 2.2 data sources 2.2.1 state standardized writing test the virginia standards of learning (sol) standardized writing test was administered each spring to students in 5th, 8th, and 10th grade.1 each test is intended to assess the state’s writing standards not only for the grade level in which the test is given but also for all grades between the current test year and the previous writing test. for example, the 8th grade test assesses writing standards for 6th, 7th, and 8th grades. although the exact standards assessed on the tests differ according to the grade in which they are given, all writing standards are subsumed under two broad categories: 1) “research, plan, compose, and revise for a variety of purposes” and 2) “edit for correct use of language, capitalization, punctuation, and spelling.” each test consists of two subtests: a multiple-choice/technology enhanced item (mc/tei) subtest and an open-ended constructed response (e.g. personal narrative, persuasive essay) subtest. scores from each subtest contribute equally to the student’s overall score, and these total scores can range from 0 to 600. the lower threshold for a “proficient” score (i.e. a passing score) is 400, and the lower threshold for an “advanced” score is 500. according to the 2014-2015 virginia standards of learning technical report (virginia department of education, 2015), scores on the writing tests demonstrated good reliability for all combinations of mc/tei tests and writing prompts (stratified alpha range .84 .88). there was one circumstance in which participants might have multiple test scores. because the longitudinal study from which these participants were recruited took place over 3 years, there is a cohort of students for whom we had both an 8th and a 10th grade test score (i.e. students who were in 8th grade during the first year of data collection and 10th grade during the third year of data collection). we chose to use the 8th grade test scores for this group because there was less missing data on other measures at this measurement point than at the 10th grade measurement point. 2.2.2 student writing/english language arts grades in this study, grades represent a student’s end-of-year grade in writing/english language arts. for students in elementary school, writing grades comprised the several standards-based criteria: writes for a variety of purposes; edits writing for correct grammar, capitalization, punctuation, and spelling; and demonstrates growth in word study knowledge. for students in high school, ela grades included teachers’ judgments of students’ progress on state standards related to both reading and writing. across all grade levels, grades were reported by the school division in a letter-grade format (e.g. a, b, c), including both pluses and minuses. after examining the distribution of all grade categories, and recognizing that including all letter grades with pluses and minuses would result in a highly uneven and unbalanced grade distribution, we instead collapsed grades into three distinct categories: 1 = c, d, or f, 2 = b, and 3 = a. this categorization roughly divides the sample into thirds as demonstrated by the grade category barplot found in figure 1. figure 1. distribution of writing grades. balancing model complexity was also a consideration, since any schema beyond two categories would require an ordinal or multinomial logistic regression and adding many more categories would further complicate the interpretation of our models. after examining the observed distribution of grades and weighing the additional complexity of modeling additional categories, we arrived at a three-category construct. 2.2.3 student-reported writing self-regulation measure to measure students’ assessment of their self-regulation, we used eight items from the larger writing self-regulation aptitude scale (wsras; ekholm et al., 2015), which was originally intended for use with college students. the scale asks students to rate their perceived writing self-regulative behaviors on a scale from 1 (almost never) to 4 (almost always), and it assesses the self-regulated learning processes of goal setting, planning, self-monitoring, attention control, emotion regulation, self-instruction, and help-seeking for writing. for the current study, the items, “i make my writing better by changing parts of it” and “i tell myself i did a good job when i write my best,” were added to include the self-regulation processes of self-evaluation and self-imposed contingencies. additionally, slight changes in language were made to the original items to ensure the developmental appropriateness of the scale. all items for this scale are available in appendix a. mcdonald’s omega (mcdonald, 1970; mcneish, 2017) for scores on this measure was .78. 2.2.4 teacher-reported student writing self-regulation measure the teacher-reported student writing self-regulation measure (trswsr; zumbrunn, 2014) is a three-item scale that asks teachers to assess the frequency with which their students plan their writing, revise their writing, and persist through difficulties during writing. all items are measured on a scale of 1 (never) to 10 (always). mcdonald’s omega (mcdonald, 1970; mcneish, 2017) for scores on this measure was .88. since teachers were asked to assess the self-regulation of several students in their classes, we opted to make this scale considerably shorter than the student-report self-regulation measure to avoid overburdening teachers. both the studentand teacher-report writing self-regulation measures were collected in the fall of the school year. 2.2.5 student demographics several demographic predictors were included as covariates in this study, including gender, socioeconomic status, and race/ethnicity. all data for these measures was provided by the participating school division, and the operationalization of each is consistent with divisionand state-level practices. gender was operationalized here as male vs. female. socioeconomic status was defined as whether a student was classified by the virginia department of education as economically disadvantaged.2 students’ race/ethnicity was operationalized into five groups: white, african american, asian, latinx, or multiracial. any student who identified as having hispanic or latino ethnicity (regardless of racial identification) was classified as latinx. due to small sample limitations, demographics such as ell status, gifted status, and special education identification were not included in these analyses. 2.3 data analysis 2.3.1 estimating discrepancy scores we modify a procedure common in the teacher expectation literature to estimate discrepancies in teacher expectations (e.g. harvey et al., 2016; hinnant et al., 2009). first, we estimated confirmatory factor analysis (cfa) models for teacherand student-reported self-regulation and then obtained predicted values of the standardized latent teacherand student-report variables for each student. then, we subtracted the value of the student-report latent score from the teacher-report latent score to obtain our estimate of teacher discrepancy. our approach differs slightly from harvey et al. (2016), who estimated discrepancy scores using the residuals of a linear model with student self-regulation predicting teacher self-regulation. by using latent variables to represent our discrepancy scores, we obtained scores that had less measurement error compared to using observed measures alone (kline, 2016). the discrepancy variable represents the difference between a student’s self-rating of self-regulation and the teacher’s evaluation of the student’s self-regulation after accounting for measurement error. negative values of the discrepancy variable represent cases where teacher ratings are lower than student ratings (i.e. underestimations), whereas positive discrepancy values represent cases where teacher ratings are higher than student ratings (i.e. overestimations). 2.3.2 group comparisons to determine the extent to which teacher discrepancy scores differ by group, we use independent samples t-tests for comparisons by gender and ses, and a one-way anova for comparisons by race/ethnicity. 2.3.3 multilevel modeling because the students in this study were nested within classrooms, we performed multilevel analyses using a two-level generalized linear mixed model with a multinomial distribution and a generalized logistic link function (agresti, 2012). this is the multilevel complement to the single-level multinomial logistic regression (cohen et al., 2003; long & freese, 2014). given the ordinal nature of our outcome measure (course grades), we arrived at the generalized multinomial logit after testing the adequacy of an ordinal logistic regression. our initial model failed the brant wald test for proportional odds (p = .004) (brant, 1990), which suggested that a multinomial distribution better characterized the pattern of responses than an ordinal logistic regression. to obtain model estimates, one outcome category (grade of a) was used as the reference category, and simultaneous models were fit comparing log odds of a student earning a (grade of b) or (grade of c and below) relative to a (grade of a). the overall modeling approach included six steps. in step 1, writing sol scores, a measure of prior performance, was included. our predictor of interest, student-teacher discrepancy score, was included in step 2. in step 3, students’ self-regulation score was introduced. in step 4, a vector of student demographic covariates was introduced, including district-reported measures of race, ethnicity, gender, and whether a student qualifies as economically disadvantaged. step 5 tested race/ethnicity, gender, and economic disadvantaged status as moderators of student/teacher discrepancy using interaction effects. finally, in step 6, a between-classroom (teacher level) covariate, average discrepancy, was introduced in addition to the variables included in steps 1-5. this model building approach, sequentially adding level one predictors, followed by level 2 predictors, is recommended by hox (2010) and allows for a more thorough examination of how studentand teacher-level predictors function in relationship to the outcome. it also facilitates a richer understanding of the extent to which blocks of covariates explain additional residual variance at the withinor between-person level. by simultaneously fitting studentand teacher-level equations, multilevel modeling can more accurately partition betweenand within-classroom variance. given the relatively small number of clusters (teachers) in this study (n = 18), we chose to limit our model to include only a random intercept; all slopes were treated as fixed. this choice is justified given the significant power demands and increasing complexity of estimation that occur with each additional random effect. all multilevel analyses were conducted using mplus 8.0 (muthén & muthén, 1998 2017). 2.3.4 testing model robustness using fixed effects to test the robustness of our model to unexplained heterogeneity at the teacher level, we also used a fixed effects regression model with a multinomial logistic link (allison, 2009; wooldridge, 2016), clustering by teacher, which applies a fixed effects transformation to all within-teacher predictors. this approach effectively removes all between-teacher heterogeneity, which is important given that unobserved or unexplained teacher characteristics may confound our interpretation of the relationship between teacher discrepancy score and grades. thus, if results are consistent between our preferred mlm model and the more restrictive fixed effects model, we may assume that there are no significant teacher-level confounding variables that we have missed. we also used a regular multinomial logit model (without random effects) along with cluster-robust standard errors. this is an additional approach for handling clustered data, and while not as preferred3 as our approach and fixed effects, it provides an additional useful comparison point for the robustness of our model to our specifications. 2.3.5 missing data all cases included in this study were complete, therefore, this study did not have any missing data. our design, which relied on the calculation of a teacher discrepancy score that required an observed student and teacher score for self-regulation, necessitated the exclusion of non-complete cases. 3. results 3.1 estimating discrepancy scores we began our analysis by conducting a confirmatory factor analysis with our latent measures of studentand teacher-reported self-regulation. this model demonstrated acceptable to good model fit according to the criteria recommended by hu and bentler (1999) (cfi = .95, rmsea = .069, 90% ci [.047, .091]). having confirmed that our model accurately represents the underlying covariance in our data, we then subtracted the predictions (tsr – ssr) from our cfa model to create our teacher discrepancy score. a histogram of this new variable is found in figure 1. given the formula used to calculate it, students with a negative discrepancy score are those whose self-rating was higher than that of their teacher (resulting in a negative score). similarly, students with a positive discrepancy are those whose self-rating was lower than that of their teacher (resulting in a positive score). students with discrepancy values at or near zero are those whose self-rating was equivalent to that of their teacher. figure 2. distribution of student-teacher discrepancy. 3.2 group differences in discrepancy scores using our model-predicted estimate of teacher discrepancy, we first conducted a series of mean comparisons to investigate whether this construct differs according to common student demographics. using independent-samples t-tests, we compared students according to gender (males vs. females), and socioeconomic status (students classified as economically disadvantaged vs. those not classified). results indicated that females had a significantly more positive teacher discrepancy than males (t = 3.67, df = 199, p <.001). this suggests that compared to students’ self-evaluations, teachers tend to overestimate females’ self-regulation and under-estimate males’ self-regulation in their writing. no significant differences were found by economically disadvantaged status (p = .54). we also used a one-way analysis of variance (anova) to examine whether, on average, teacher discrepancy ratings differed according to students’ identified race/ethnicity. the overall model was not significant (f(4, 196) = 1.28, p = .28), which suggests that average discrepancy scores do not differ significantly by student race/ethnic group. 3.3 zero-order correlations next, we estimated a set of zero-order correlations to assess the interrelationships between our measured scales, grades, and standardized writing score. results can be found in table 1. several significant correlations emerged, including a positive and correlation between sol score and final grades (r = .58, p < .001). teacher discrepancy was positively correlated with sol score (r = .39, p < .001) and end-of-year grade (r = .50, p < .001). student-reported self-regulation was not significantly correlated with either sol score (r = .04, p = .69) or end of year grade (r = .11, p = .33). table 1 pairwise correlations among variables of interest notes. sol = virginia standards of learning exam. eoy = end of year. ssr = student-reported self-regulation. tsr = teacher-reported self-regulation. disc = discrepancy between tsr and ssr. a all correlations with eoy grade, an ordinal variable, are spearman rank-order correlations. * p < .05. *** p < .001. 3.4 multilevel modeling next, we estimated a series of multilevel models with students modeled at level 1, nested within teachers at level 2, and final grades as the outcome. covariates in our final model include sol scores, discrepancy scores, student self-regulation, and the student demographics described above. descriptive statistics for this model can be found in table 2. table 2 descriptive statistics for variables of interest notes. n = 201. data are presented as mean (sd) for continuous measures, and n (%) for categorical measures. we used a hierarchical model building approach and arrived out our final model in six steps. for the sake of parsimony, we only interpret the final main effect model results here, however, full estimates of all models are available in table a1 in the appendix. results are presented in two blocks – first, we interpret the coefficients that compare the odds of earning an a vs. a b. second, we interpret the coefficients that compare the odds of earning an a vs. a c or below. a visual summary of all coefficients can be found in figure 3. figure 3. visual summary of multilevel model estimates. 3.4.1 comparing the odds of earning a vs. b higher sol scores were significantly associated with lower odds of earning a b vs. earning an a (b = -0.04, p < .001). in standardized units, for each one-sd increase in sol score, the odds of earning a b (vs. an a) decreased by about 45 percent. teacher discrepancy score was significant and associated with lower odds of earning a b (b = -.53, p < .001). for each one-sd increase in teacher discrepancy (suggesting that teachers overrate students’ self-regulation relative to their self-assessment), the odds of earning a b (vs. an a) decrease by about 21 percent. student-reported self-regulation was not found to be a significant predictor of grades when controlling for sol score and teacher discrepancy score (b = -.11, p =.91). along with the three predictors of interest interpreted above, a vector of student demographics, including gender, race/ethnicity, and economically disadvantaged status were included in the final model. females were predicted to have 221 percent higher odds of earning a b vs. an a compared to males (b = 1.17, p < .001), and economically disadvantaged students were predicted to have 458 percent higher odds of earning a b vs. an a compared to non-economically disadvantaged students (b = 1.72, p < .018) after controlling for sol score, teacher discrepancy, and student self-report . in step 5, we introduced a series of discrepancy-by-demographic interaction effects to test for statistical moderation. we found that none of our demographic variables moderated the relationship between discrepancy score and grades. finally, we also introduced a teacher-level covariate, average discrepancy score, to examine whether a given teacher’s tendency to overor under-rate self-regulation relative to their students was associated with differential odds of earning a b vs. an a. we found that it was (b = -1.38, p =.028), which suggests that for each one-unit increase in a teacher’s average discrepancy, we would expect the odds of earning a b vs. an a to decrease by about 75 percent. 3.4.2 comparing the odds of earning a vs. c or below as might be expected, nearly all predictors had coefficient estimates of similar magnitudes and directions when comparing the odds of earning a c or below vs. an a. increases in sol scores (b = -0.04, p <.001) and teacher discrepancy scores (b = -0.94, p <.001) were both significant and associated with decreased odds of earning a c or below vs. an a. thus, increases in these measures would predict higher odds of receiving an a. the estimates for a teachers’ average discrepancy at the classroom level functioned similarly and were significant (b = -1.67, p <.001), which suggested again that as a teachers’ average tendency to overrate their students’ self-regulatory behaviors increased, they were also more likely to assign a given student a high grade. finally, gender appeared to function similarly for this outcome category, as females had significantly higher odds of earning a c or below vs. an a compared to males with similar sol scores and teacher discrepancy scores (b = 1.53, p <.001). several predictors did not function equivalently when comparing the odds of earning an a vs. a c or below, and this is likely why the model overall did not satisfy the proportional odds assumption required to use ordinal (vs. multinomial) logistic regression. economic disadvantage was not found to be a statistically significant predictor, although the magnitude and direction were similar to the b vs. a estimates. additionally, students who identified as multiracial were found to have significantly higher odds of earning a c or below vs. an a compared to white students (b = 2.11, p = .004), although given the especially small subsample of multiracial students in this study, we would advise caution in over-interpreting the significance of this result. 3.5 robustness checks comparing alternative model specifications to test the robustness of our model to unexplained or confounding variables at the teacher level, we also employed a fixed-effects panel regression model, which uses a fixed-effects transformation to eliminate all teacher-level heterogeneity, and a regular multinomial regression model with cluster-robust standard errors. these represent alternative approaches to operationalizing the model we tested, and as such, we would expect the coefficient estimates derived from all three to be quite similar. to test this assumption, we repeated our analysis two more times and compared the results. the covariates used (except for average discrepancy, which could not be included in the fixed effects model) and the modeling process proceeded as with the multilevel formulation. results were generally very consistent, with discrepancy remaining a significant and positive predictor of student grades, although the magnitude of the discrepancy coefficient was slightly larger in our preferred mlm approach. a comparison of the estimates for the discrepancy measure can be found in table 3. table 3 sensitivity analysis of discrepancy estimates comparing mlm, re, and olr-cr models notes. mlm = multilevel model. re = random effects ordinal logit model. olr-cr = ordinary logistic regression model with cluster (teacher)-robust standard errors. all estimates are presented as log odds coefficients. *** p< .001. 4. discussion the purpose of this study was to examine the relations between students’ reports of their own writing self-regulation, teacher reports of student writing self-regulation, standardized writing achievement, student writing/english language arts (ela) grades, and student demographics. more specifically, we were interested in better understanding discrepancies between student and teacher reports of self-regulation, including how these discrepancies might relate to writing outcomes as well as how they might manifest differently among student demographic groups. when comparing discrepancy scores across student demographic characteristics, we found that females had significantly higher discrepancy scores than males, indicating that, relative to students’ self-ratings, teachers tend to overestimate females’ writing self-regulation and underestimate males’ writing self-regulation. using multilevel modeling, nesting students within classrooms, we found that the discrepancy between teachers’ beginning of the year evaluation of students’ self-regulation and students’ own evaluation of their self-regulation is a significant and positive predictor of students’ final grades in writing/ela. further, we found that this discrepancy remains predictive of grades even after accounting for a range of covariates, including students’ demographic background and prior achievement in writing, suggesting that this discrepancy is a durable construct that is not subsumed by other common predictors of academic performance. as prior research has shown (brookhart et al., 2016; casillas et al., 2012; willingham et al., 2002), grades represent an amalgamation of several factors, including student ability, knowledge, motivation, persistence, and behavior. across this body of research, student ability regularly emerges as the strongest single predictor of grades. unsurprisingly, this was the case in this study as well. students’ end-of-year writing/ela grades were more highly correlated with standardized writing achievement scores than with any other predictors examined in the current study. similarly, in our final regression model, standardized writing achievement made the strongest unique contribution to the prediction of students’ grades. nevertheless, our discrepancy score variable was correlated with students’ final grades (r = .50) and uniquely predicted students’ grades after accounting for the influence of other predictors, including writing achievement. this finding echoes that of previous research (e.g. harvey et al., 2016) indicating that teachers’ perceptions of students’ behavioral and psychological characteristics meaningfully relate to grading decisions, even after these perceptions have been adjusted to account for similar information from other sources. we also found that female students tend to have larger positive discrepancy scores than male students. there are two possible explanations for this. first, female students may be more critical of their self-regulation in writing. if female students tend to underestimate their own self-regulation, then even accurate teacher estimations would yield positive discrepancy scores because of how these scores were calculated. inversely, teachers may overestimate female students’ self-regulation in relation to their male peers. this explanation is consistent with prior research investigating gender-based motivational differences across academic domains (e.g. meece et al., 2006; pajares et al., 1999). according to this body of research, females are, on average, more motivated in ela disciplines than are their male counterparts. if teachers are familiar with this trend or otherwise hold high a priori expectations of female writers, they may overestimate female students’ writing self-regulation. however, it should be noted that these findings conflict somewhat with results reported by mckown & weinstein (2002), who found no teacher expectation effects for females in the domain of reading. contrary to the results of several previous studies suggesting teachers hold differential expectations of students based on their ethnicity and/or socioeconomic status (auwarter & aruguete, 2008; rist, 1970; tenenbaum & ruck, 2007), we did not find any significant differences in self-regulation discrepancy scores between these groups. further, these demographic variables did not significantly predict student grades. one possible explanation for this has to do with when we administered our fall survey. the school district in which this study was conducted begins its academic year immediately following labor day, and teachers rated students’ writing self-regulation in october. according to brophy (1983), teachers may rely more heavily on student demographic characteristics to inform expectations very early in the school year, when they do not yet have more relevant information about these students. given that the teachers in this study had several weeks to form expectations of their students from these students’ writing behaviors, it seems possible that they based these expectations more on their observations of students’ writing proficiencies than on student ethnicity. this study has a number of important limitations worth highlighting. first, all participating students and teachers were recruited from a single school division, albeit a demographically diverse one. therefore, it is possible that factors unique to this district, such as district-level policies relating to grading practices, could influence the results we report here. second, although our modeling strategy accounts for some between-teacher differences in discrepancy scores and grading practices, previous research suggests that teachers vary in the factors they consider when assigning student grades (e.g. mcmillan, 2001). we did not have enough teachers to investigate many teacher-level variables relating to student grades, so we could not investigate typical grading practices as a predictor. this is a potential area for future research. third, the grades variable we examine here refers to the students’ grade in writing/english language arts. therefore, this grade likely reflects factors such as reading ability, behaviors, and motivation in addition to writing ability, behaviors, and motivation. although reading and writing scores are typically highly correlated (stotsky, 1983), they are nevertheless different domains. most importantly, our work here operationalizes writing self-regulation in a very specific (and perhaps narrow) way, focusing on the extent to which students plan, revise, and persist in writing despite challenges. we recognize that these three behaviors represent only a thin slice of the larger set of behaviors that encompass writing self-regulation. future research should seek to explore associations involving other key components of self-regulation, and should explore (perhaps qualitatively) how teachers themselves conceptualize writing self-regulation to ensure that important behaviors are not being overlooked. in addition, future work should consider the relationship between self-regulation discrepancy and other measures of writing performance and achievement, especially measures such as standardized writing prompts that can be independently scored and evaluated and do not involve as much subjective input from teachers. this would help to further establish the validity of self-regulation discrepancy as a concept distinct from teacheror studentevaluated self-regulation in isolation, especially given the strong pairwise correlations that we observed in this study between discrepancy, teacher-reported self-regulation, and writing grades. 4.1 conclusion the results of this study raise several important questions. most notably, this work demonstrates that teachers’ initial perceptions of student writing self-regulation are predictive of students’ eventual performance in language arts class. this finding alone may not raise eyebrowsafter all, students who demonstrate higher levels of self-regulation early on are also more likely to end up with higher grades in class. this study, however, goes a step further. we find that the discrepancy between teacher and student perceptions remain predictive of performance, even after controlling for students’ prior performance and their own evaluation of their self-regulation. in other words, when teachers tend to overestimate a student’s self-regulation, that student is more likely to end up with higher grades, even after controlling for the student’s demonstrated ability or own perceptions of self-regulation. this finding is particularly interesting in the context of writing, where many (but certainly not all) measures of writing self-regulation (e.g. planning, revising, avoiding distractions) are usually behaviors that can be easily observed by teachers in situ. the writing self-regulation scale used here focuses largely on the use of writing strategies and processes, such as planning and revising, and on writing persistence, such as writing for longer periods of time while remaining focused. however, the observability of these behaviors does not seem to mitigate the discrepancies that can emerge between student and teacher reports, and further, these discrepancies remain a positive and significant predictor of writing achievement after accounting for other measures of writing ability. given these findings, what new questions arise for writing researchers? a key question seems to be if writing self-regulation discrepancy exists, and is not explained by typical background characteristics, what is driving it? how do teachers come to over-estimate (or students under-estimate) writing self-regulation? and how does this systematic discrepancy in turn lead to differential performance grades in writing, even after accounting for prior student academic performance? one possible avenue of exploration might be mixed-methods research that systematically identifies students with low and high discrepancy scores, and then uses interviews, focus groups, or field observation to obtain more nuanced understandings of the classroom and individual processes at play. this type of work would further unpack the complex associations revealed in this study. so where does this leave writing teachers? the task of accurately assessing developing writers is a challenging one, and surely it involves careful observation of how students are implementing new writing strategies. however, the results may suggest some caution to those who might view these behaviors as a relatively “clean” measure of writing achievement. even when observing these processes, subtle biases can emerge that can shape teachers’ future assessment and evaluation of their students. keypoints we find a significant and positive relationship between student-teacher self-regulation discrepancy and ela grades in a sample of 201 middle school students in mid-atlantic region the us. this relationship holds even after controlling for prior student achievement and background/ demographic characteristics. these findings provide preliminary evidence of the durable nature of student-teacher discrepancy as a predictor of more subjective academic outcomes, such as grades. acknowledgments this research was supported by grants from the virginia commonwealth university presidential research incentive program and the virginia commonwealth university foundation langschultz fund. the views expressed in this paper are those of the authors and do not necessarily reflect the views of the granting agencies or organizations. footnotes 1 since then, virginia has eliminated the 5th grade writing test. the 8th and 10th grade tests remain in use. 2 the virginia department of education assigns “economically disadvantaged” status each year for students who: 1) are eligible for free/reduced price meals, 2) receive temporary aid for needy families (tanf), or 3) are eligible for medicaid. 3 we found that adjusting standard errors for clustering was less preferred than either a fixed effects or multilevel approach because it only impacts the size of the estimated standard errors, and not the magnitude of the coefficients. the other two methods we tested can adjust both standard errors and coefficient estimates for clustering. references agresti, a. (2012). categorical data analysis (3rd ed.). new york, ny: john wiley & sons. allison, p. d. (2009). fixed effects regression models. thousand oaks, ca: sage. auwarter, a. e., & aruguete, m. s. (2008). effects of student gender and socioeconomic status on teacher perceptions. the journal of educational research, 101(4), 242-246. https://doi.org/ 10.3200/joer.101.4.243-246 brant, r. (1990). assessing proportionality in the proportional odds model for ordinal logistic regression. biometrics, 46, 1171–1178. https://doi.org/10.2307/2532457 brookhart, s. m., guskey, t. r., bowers, a. j., mcmillan, j. h. smith, j.k., smith, l. f., … & welsh, m. e. (2016). a century of grading research: meaning and value in the most common educational measure. review of educational research, 86(4), 803-848. https://doi.org/10.3102%2f0034654316672069 brophy, j. e. (1983). research on the self-fulfilling prophecy and teacher expectations. journal of educational psychology, 75(5), 631-661. https://doi.org/10.1037/0022-0663.75.5.631 bowers, a. j. (2011). what’s in a grade? the multidimensional nature of what teacher-assigned grades assess in high school. educational research and evaluation, 17(3), 141-159. https://doi.org/10.1080/13803611.2011.597112 casillas, a., robbins, s., allen, j., kuo, y. l., hanson, m. a., & schmeiser, c. (2012). predicting early academic failure in high school from prior academic achievement, psychosocial characteristics, and behavior. journal of educational psychology, 104(2), 407-420. https://doi.org/10.1037/a0027180 cohen, j., cohen, p., west, s. g., & aiken, l. s. (2003). applied multiple regression/correlation analysis for the behavioral sciences (3rd ed.). lawrence erlbaum associates. collie, r. j., martin, a. j., & curwood, j. s. (2016). multidimensional motivation and engagement for writing: construct validation with a sample of boys. educational psychology, 36(4), 771-791. https://doi.org/10.1080/01443410.2015.1093607 de boer, h., bosker, r. j., & van der werf, m. p. (2010). sustainability of teacher expectation bias effects on long-term student performance. journal of educational psychology, 102(1), 168-179. https://doi.org/10.1037/a0017289 duckworth, a. l., & yeager, d. s. (2015). measurement matters: assessing personal qualities other than cognitive ability for educational purposes. educational researcher, 44(4), 237–251. https://doi.org/10.3102/0013189x15584327 ekholm, e., zumbrunn, s., & conklin, s. (2015). the relation of college student self-efficacy toward writing and writing self-regulation aptitude: writing feedback perceptions as a mediating variable. teaching in higher education, 20(2), 197-207. https://doi.org/10.1080/13562517.2014.974026 graham, s. (2018). a writer(s) within community model of writing. in c. bazerman, v. berninger, d. brandt, s. graham, j. langer, s. murphy, p. matsuda, d. rowe, & m. schleppegrell, (eds.), the lifespan development of writing (pp. 271-325). national council of english. graham, s., & r. harris, k. (2000). the role of self-regulation and transcription skills in writing and writing development . educational psychologist, 35(1), 3-12. https://doi.org/10.1207/s15326985ep3501_2 graham, s., harris, k. r., & mason, l. (2005). improving the writing performance, knowledge, and self-efficacy of struggling young writers: the effects of self-regulated strategy development. contemporary educational psychology, 30(2), 207-241. https://doi.org/10.1016/j.cedpsych.2004.08.001 graham, s., harris, k. r., & santangelo, t. (2015). research-based writing practices and the common core: meta-analysis and meta-synthesis. the elementary school journal, 115(4), 498-522. https://doi.org/10.1086/681964 graham, s., mckeown, d., kiuhara, s., & harris, k. r. (2012). a meta-analysis of writing instruction for students in the elementary grades. journal of educational psychology, 104(4), 879-896. https://doi.org/10.1037/a0029185 graham, s., & perin, d. (2007). a meta-analysis of writing instruction for adolescent students. journal of educational psychology, 99(3), 445-476. https://doi.org/10.1037/0022-0663.99.3.445 hacker, d. (2018). a metacognitive model of writing: an update from a developmental perspective. educational psychologist, 53 (4), 220-237. https://doi.org/10.1080/00461520.2018.1480373 harvey, k. e., suizzo, m. a., & jackson, k. m. (2016). predicting the grades of low-income–ethnic-minority students from teacher-student discrepancies in reported motivation. the journal of experimental education, 84(3), 510-528. https://doi.org/10.1080/00220973.2015.1054332 hayes, j. r. (2012). modeling and remodeling writing. written communication, 29(3), 369-388. https://doi.org/10.1177/0741088312451260 hayes, j. r., & flower, l. s. (1980). identifying the organization of writing processes. in l. gregg & e. steinberg (eds.), cognitive processes in writing (pp. 3 –30). lawrence erlbaum associates, inc. hayes, j. r., & flower, l. s. (1986). writing research and the writer. american psychologist, 41(10), 1106-1113. https://doi.org/10.1037/0003-066x.41.10.1106 hinnant, j. b., o’brien, m., & ghazarian, s. r. (2009). the longitudinal relations of teacher expectations to achievement in the early school years. journal of educational psychology, 101(3), 662–670. https://doi.org/10.1037/a0014306 hox, j. j. (2010). multilevel analysis: techniques and applications (2nd ed.). routledge. hu, l. t., & bentler, p. m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives. structural equation modeling: a multidisciplinary journal, 6(1), 1-55. https://doi.org/10.1080/10705519909540118 jussim, l., eccles, j., & madon, s. (1996). social perception, social stereotypes, and teacher expectations: accuracy and the quest for the powerful self-fulfilling prophecy. in advances in experimental social psychology (vol. 28, pp. 281-388). academic press. jussim, l., & harber, k. d. (2005). teacher expectations and self-fulfilling prophecies: knowns and unknowns, resolved and unresolved controversies. personality and social psychology review, 9(2), 131-155. https://doi.org/10.1207/s15327957pspr0902_3 kline, r. (2016). principles and practice of structural equation modeling (4th ed.) guilford press. long, j. & freese, j. (2014). regression models for categorical dependent variables using stata (3rd ed.). stata press. madon, s., jussim, l., & eccles, j. (1997). in search of the powerful self-fulfilling prophecy. journal of personality and social psychology, 72(4), 791-809. https://doi.org/10.1037/0022-3514.72.4.791 madon, s., jussim, l., keiper, s., eccles, j., smith, a., & palumbo, p. (1998). the accuracy and power of sex, social class, and ethnic stereotypes: a naturalistic study in person perception. personality and social psychology bulletin, 24(12), 1304-1318. https://doi.org/10.1177/01461672982412005 mcdonald, r. p. (1970). the theoretical foundations of principal factor analysis, canonical factor analysis, and alpha factor analysis. british journal of mathematical and statistical psychology, 23(1), 1-21. https://doi.org/10.1111/j.2044-8317.1970.tb00432.x mckown, c., & weinstein, r. s. (2002). modeling the role of child ethnicity and gender in children's differential response to teacher expectations. journal of applied social psychology, 32 (1), 159-184. https://doi.org/10.1111/j.1559-1816.2002.tb01425.x mcmillan, j. h. (2001). secondary teachers' classroom assessment and grading practices. educational measurement: issues and practice, 20(1), 20-32. https://doi.org/10.1111/j.1745-3992.2001.tb00055.x mcneish, d. (2018). thanks coefficient alpha, we’ll take it from here. psychological methods, 23(3), 412-433. https://doi.org/10.1037/met0000144 meece, j. l., glienke, b. b., & burg, s. (2006). gender and motivation . journal of school psychology, 44(5), 351-373. https://doi.org/10.1016/j.jsp.2006.04.004 muthén, l. k., & muthén, b. o. (1998-2017). mplus user’s guide (8th ed.). los angeles, ca: muthén & muthén. pajares, f., miller, m. d., & johnson, m. j. (1999). gender differences in writing self-beliefs of elementary school students. journal of educational psychology, 91(1), 50-61. https://doi.org/10.1037/0022-0663.91.1.50 pintrich, p. r., & de groot, e. v. (1990). motivational and self-regulated learning components of classroom academic performance. journal of educational psychology, 82(1), 33-40. https://doi.org/10.1037/0022-0663.82.1.33 rist, r. (1970). student social class and teacher expectations: the self-fulfilling prophecy in ghetto education. harvard educational review, 40(3), 411-451. https://doi.org/10.17763/haer.40.3.h0m026p670k618q3 rosenthal, r., & jacobson, l. (1968). pygmalion in the classroom. the urban review, 3(1), 16-20. https://doi.org/10.1007/bf02322211 salahu-din, d., persky, h., & miller, j. (2008). the nation’s report card: writing 2007 (nces report no. 2008-468). national center for education statistics. santangelo, t., harris, k. r., & graham, s. (2016). self-regulation and writing: an overview and meta-analysis. in c. macarthur, s. graham, & j. fitzgerald (eds.), handbook of writing research (vol. 2, pp. 174– 193). guilford. stotsky, s. (1983). research on reading/writing relationships: a synthesis and suggested directions. language arts, 60(5), 627-642. url: https:// www.jstor.org/stable/41961512 tenenbaum, h. r., & ruck, m. d. (2007). are teachers' expectations different for racial minority than for european american students? a meta-analysis. journal of educational psychology, 99(2), 253-273. https://doi.org/10.1037/0022-0663.99.2.253 virginia department of education (2015). virginia standards of learning assessments technical report: 2014-2015 administration cycle . retrieved from http://www.doe.virginia.gov/testing/test_administration/technical_reports/sol_technical_report_2014-15_administration_cycle.pdf willingham, w. w., pollack, j. m., & lewis, c. (2002). grades and test scores: accounting for observed differences . journal of educational measurement, 39(1), 1-37. https://doi.org/10.1111/j.1745-3984.2002.tb01133.x winne, p. h., & hadwin, a. f. (1998). studying as self-regulated learning. in hacker, d. j, dunlosky, j., & graesser, a.c. (eds.), metacognition in educational theory and practice (pp. 277-304). routledge. woodcock, r., & johnson, m.b. (1977). the woodcock-johnson psycho-educational battery. dlm teaching resources. wooldridge, j. m. 2016. introductory econometrics: a modern approach (6th ed.). cengage. zimmerman, b. j. (2008). investigating self-regulation and motivation: historical background, methodological developments, and future prospects. american educational research journal, 45(1), 166-183. https://doi.org/10.3102/0002831207312909 zimmerman, b. j. (2013). theories of self-regulated learning and academic achievement: an overview and analysis. in b. j. zimmerman & d. h. schunk (eds.), self-regulated learning and academic achievement: theory, research, and practice (pp. 10-45). routledge. zimmerman, b.j., & bandura, a. (1994). impact of self-regulatory processes on writing course attainment. american educational research journal, 31(4), 845 – 862. https://doi.org/10.3102/00028312031004845 zimmerman, b., & riesemberg, r. (1997). becoming a self-regulated writer: a social cognitive perspective. contemporary educational psychology, 22, 73–101. https://doi.org/10.1006/ceps.1997.0919 zumbrunn, s. (2014, february). perceived writing climate as a predictor of student writing self-efficacy and self-regulation. paper presented at the writing research across borders conference, paris, france. zumbrunn, s., marrs, s., & mewborn, c. (2016). toward a better understanding of student perceptions of writing feedback: a mixed methods study. reading and writing, 29 (2), 349-370. https://doi.org/10.1007/s11145-015-9599-3 appendix – supplementary tables table a1 results from multilevel multinomial logistic regression predicting writing grades notes. n = 201. sol = standards of learning test score. disc = discrepancy score. econ_dis = student classified as economically disadvantaged. int1 = sol x discrepancy interaction effect. int2 = srwsr x discrepancy interaction effect. int3 = female x discrepancy interaction effect. int4 = econ_dis x discrepancy interaction effect. int5 = asian x discrepancy interaction effect. int6 = african american x discrepancy interaction effect. int7 = latinx x discrepancy interaction effect. int8 = multiracial x discrepancy interaction effect. coefficients reported here are log odds coefficients (logits). microsoft word pesu et al_publication.docx ! ! ! ! ! frontline)learning)research)vol.4)no.)3)(2016))92)=)109) issn)2295=3159)) ! *corresponding author at: department of psychology, p.o. box 35, university of jyväskylä, fin-40014, finland, email address: laura.a.pesu@jyu.fi doi: http://dx.doi.org/10.14786/flr.v4i3.249! the development of adolescents’ self-concept of ability through grades 7-9 and the role of parental beliefs laura pesu*, kaisa aunola, jaana viljaranta, & jari-erik nurmi university of jyväskylä, finland article received 7 april / revised 4 july / accepted 8 july / available online 20 july! abstract this study examined the development of adolescents’ self-concept of ability in mathematics and literacy during secondary school, and the role that mothers’ and fathers’ beliefs concerning their child’s abilities play in this development. also examined was whether the role of mothers’ and fathers’ beliefs about their adolescent child’s ability in mathematics and literacy differs according to the adolescent’s gender and level of performance. a total of 231 adolescents and their mothers and fathers were followed up across secondary school. the results showed, first, that adolescents’ self-concept of ability declined slightly from grade 7 to grade 9 in both mathematics and literacy. second, mothers’ and fathers’ beliefs about their adolescent child’s abilities in grade 7 predicted the child’s subsequent self-concept in grade 9, but only in mathematics. third, the role of mothers’ beliefs in their child’s self-concept of mathematics ability was found to be stronger among high-performing than low-performing adolescents. keywords: self-concept of ability; secondary school; mother’s beliefs; father’s beliefs pesu%et %al % % % | f l r ! ! 93! 1. introduction students’ self-concept of ability in different academic domains, that is, the knowledge and perceptions individuals have of themselves in a particular subject area (bong & skaalvik, 2003; brunner, keller, hornung, reichert, & martin, 2009) influences their academic performance and the academic careerrelated choices they make (eccles et al. 1983; marsh, trautwein, lüdtke, köller, & baumert, 2005; valentine, dubois, & cooper, 2004; wigfield, eccles, schiefele, roeser, & davis-kean, 2006). since these self-conceptions guide students’ actual performance at school and hence their future education and related decisions, it is important to identify the factors that support the development of self-concept, particularly during the critical period of adolescence when self-concept of ability typically declines (nagy et al., 2010; wigfield et al., 1997). because the development of self-concept of ability has been suggested to be linked to interaction with other people (dermitzaki & efklides, 2000), such as parents, the present study examined the development of self-concept of ability in literacy and mathematics among 231 finnish adolescents from grade 7 to grade 9, and the role that mothers’ and fathers’ beliefs about their children’s abilities play in this development. also investigated was whether children’s gender and level of performance influence the possible associations between parental beliefs and their child’s self-concept of ability. 1.1 self-concept of ability recent research has led to an understanding that self-concept is multidimensional and hierarchical in nature and is formed in social comparison and in communication with significant others (bong & skaalvik, 2003). thus, academic self-concept may be different for the domains of mathematics and verbal skills, for example (arens, yeung, craven, & hasselhorn, 2011). previous research has shown that mathematics and verbal self-concepts are almost uncorrelated although achievement in mathematics and verbal subjects substantially correlate (marsh, 1990; marsh, byrne, & shavelson, 1988). the internal/external frame of reference (i/e) model focuses on explaining why this is. according to the i/e model academic self-concept in a specific school subject is formed in relation to two comparison processes that are called “frames of reference” (marsh & yeung, 2001). in the external (normative/social comparison) frame of reference a student compares his/her own performance in a particular domain (e.g. mathematics) with her/his perception of other students’ performance in this domain. in the internal (ipsative-like) reference a student compares his/her own performance in a particular domain (e.g. mathematics) with his/her performance in other school subjects (e.g. literacy). the actual self-concept in a particular school domain is formed in these simultaneous comparison processes. thus, if a student is poor in mathematics compared to other students in his/her class (external comparison), but in comparison to his/her performance in other school subjects is doing better in mathematics than in other subjects, his/her mathematics self-concept can be good. based on the internal/external frame of reference (i/e) model, as well as previous empirical studies showing that mathematics and verbal self-concept domains are distinct (arens et al., 2011), in the present study selfconcept is approached subject-specifically. the expectancy-value theory by eccles et al. (1983) provides a theoretical framework for selfconcept in the academic setting. according to the expectancy-value theory (eccles et al., 1983; eccles & wigfield, 1995; wigfield & eccles, 2000) individuals’ performance in school and their academic choices are explained not only by the extent to which they value the activity in question, but also by the expectancies they have for success in that activity (wigfield & eccles, 2000). according to the theory, students’ selfconcept of ability, that is, the individual’s perception of his or her competence in a certain academic domain, influences the expectancies students have and, through these expectancies, different academic outcomes, such as performance (wigfield & eccles, 2000). theoretically, self-concept of ability is distinct from expectancy of success: self-concept of ability focuses on present ability while expectancies focus on the future. however, empirically these two concepts have not been found to be separate (eccles et al., 1983; wigfield & eccles, 2000). pesu%et %al % % % | f l r ! ! 94! previous research has shown that students’ self-concept of ability plays an important role in academic environments by directing behavior and effort in learning situations (e.g., atkinson, 1964; bandura, 1986; eccles et al., 1983; wigfield et al., 2006). students who believe in their abilities and expect that they can and will do well in a task are much more likely to perform better and to engage in an adaptive manner in such academic tasks than students who do not believe in their abilities and expect to fail in a certain task (chapman, tunmer, & prochnow, 2000; eccles et al., 1983; pintrich & schunk, 2008). similar results have been found among both younger school-aged children (chapman et al., 2000) and adolescents (caprara, vecchione, alessandri, gerbino, & barbaranelli, 2011; eccles et al., 1983), and in different academic domains, such as math (chiu & klassen, 2010; eccles et al., 1983) and literacy (chapman et al., 2000; chiu & klassen, 2009). among adolescents, self-concept of ability has further been found to predict career choices. it has been shown, for example, that students who have greater confidence in their math abilities are more likely to aspire to math-related careers than students whose confidence in their math abilities is lower (eccles, 2007). several studies have shown that the development of self-concept of abilities is a continuous process that starts at the very beginning of the school career. young students typically have very positive, and even unrealistic, perceptions of their abilities during the first years of primary school (aunola, leskinen, onatsuarvilommi, & nurmi, 2002), but as they grow older, their perceptions of their abilities become more realistic and more negative (jacobs, lanza, osgood, eccles, & wigfield, 2002). one important phase for the development of self-concept of ability is early adolescence (preckel, niepel, schneider, & brunner, 2013). during this time, many physical changes and changes in a person’s environment and social context take place. at the same time an educational transition usually takes place the transition to secondary school. this transition means changes in adolescents’ everyday social contexts, in the ways adolescents get feedback in school and in their frames of reference (see wigfield et al., 2006). the rates of self-concept of ability in mathematics and literacy have been shown to decline during elementary and secondary school (e.g. wigfield, eccles, maciver, reuman, & midgley, 1991). because the earlier studies on the topic have mainly been carried out in the us (eccles et al., 1983; nagy et al., 2010), australia (nagy et al., 2010; watt, 2004), or germany (nagy et al., 2010; preckel et al., 2013), it is not known, however, whether the results on the tendency of self-concept of ability to decline during the transition to secondary school apply to other cultural and educational settings. consequently, the first aim of the present study was to examine the developmental changes in self-concept of mathematics and literacy abilities during secondary school in finland. the characteristics of the finnish school system differ from school systems in some other countries. in finland, children start their education by attending pre-school in the year they turn 6. in the year of their 7th birthday children start compulsory comprehensive school which is divided into a lower level (i.e., elementary school; grades 1-6) and an upper level (i.e., secondary school; grades 7-9). in finnish secondary schools all students are taught at the same academic level and students do not need to make decisions whether to take higher or lower level courses. this characteristic of finnish school system is different from, for example, the system in germany where students need to decide which achievement-based secondary school track they take (gniewosz & noack, 2012). because in finland the compulsory courses are at the same level for everyone! both highand low-performing students are studying in the same classrooms. moreover, in finnish comprehensive school education extra attention is paid to support particularly those students who have difficulties in their learning. the fact that finnish school system includes well-developed support services for students suffering, for example, from learning difficulties has been suggested to partly explain finnish students’ academic success in worldwide pisa-studies (välijärvi et al., 2007). overall, the fact that in finland all students are taught at the same academic level independent of their level of performance or motivation and that extra attention is paid to support students with learning difficulties may positively impact the students’ self-concept development, particularly among students showing lower performance. pesu%et %al % % % | f l r ! ! 95! 1.2 the role of parents previous studies have shown that ability-related self-concepts develop in interaction with one’s environment, and are affected by evaluations of and feedback from parents (bong & skaalvik, 2003; eccles et al. 1983; gniewosz, eccles, & noack, 2014; shavelson, hubner, & stanton, 1976). according to the expectancy-value model proposed by eccles and colleagues (1983), parental beliefs about their children’s abilities may affect children’s self-concept of ability development through at least two mechanisms (see e.g. eccles, 1993). first, parents may directly tell their children what they think the child is good at (jacobs & eccles, 2000). second, parents can also provide different learning opportunities for their children based on their beliefs about their children’s abilities (jacobs & eccles, 2000). children then interpret this information from their parents and incorporate it into their self-concept of ability (jacobs & eccles, 2000). there is also strong empirical evidence for the assumption that parents’ beliefs about their children’s academic performance affect children’s subject-specific self-concept of ability (eccles parsons, adler, & kaczala, 1982; frome & eccles, 1998; gniewosz, eccles, & noack, 2012; jacobs, 1991; mcgrath & repetti, 2000; phillips, 1987). for example, parents’ beliefs in their child’s success in the literacy domain have been found to be positively related to sixth-grade children’s self-concept of their literacy ability (frome & eccles, 1998). similar results have been found in the domain of mathematics (eccles parsons et al., 1982; gniewosz et al., 2012). although the importance of parental beliefs in the formation of children’s self-concept of mathematics and literacy ability is widely acknowledged, there is some evidence that the role of parental beliefs in the development of students’ self-concept may vary with age (e.g., gniewosz et al., 2012). for example, pesu, viljaranta and aunola (2016) found that teachers’ beliefs played a bigger role than parents’ beliefs in first-grade students’ self-concept of mathematics and literacy ability development. gniewosz et al. (2012), in turn, found that the effects of maternal child-related competence beliefs on students’ mathematics self-concept increased during the secondary school transition, whereas the effect of grades decreased. interestingly, after the school transition the impact of maternal competence beliefs decreased and the impact of grades increased. when interpreting the previous results on the topic it should be noted that although longitudinal procedures were applied when predicting children’s self-concept of ability by parental beliefs, children’s self-concept of ability may also play a role in parental beliefs. the studies focusing on the role of parental beliefs in students’ self-concept of abilities have also found some gender differences. for example, it has been shown that parents typically think that boys are better at mathematics than girls (eccles parsons et al., 1982; eccles & jacobs, 1987; gunderson, ramirez, levine, & beilock, 2012), independently of children’s actual performance in mathematics (eccles, 1993; eccles parsons et al., 1982). this has been shown to impact girls’ self-perceptions in mathematics (jacobs, 1991). conversely, parents tend to think that girls do better in literacy (gniewosz et al., 2014). although there are studies focusing on these mean-level-differences in parental beliefs concerning boys and girls, less is known, however, whether there is variability in the relations among parental beliefs and their children’s self-concept of ability between boys and girls. according to simpkins, fredricks, and eccles (2012) it is important to study whether the associations among the indicators vary as a function of gender because socialization and cognitive theories suggest that the associations are not similar for boys and girls. according to these theories adolescents most likely act in a similar way as people who are most similar to themselves (maccoby, 1998). this suggests that mothers may have a stronger impact on their daughters than to their sons and fathers on their sons than their daughters (maccoby, 1998). furthermore, testing the moderating effect of gender is important because it may have important implications for interventions (simpkins et al., 2012): if parental beliefs have different impact on boys and girls self-concept of ability, the interventions should take this into consideration when thinking the best ways to support girls and boys. whether the effect of parental beliefs about their children’s abilities on children’s self-concept development is affected by the child’s gender is thus far, however, underexplored. alongside gender it has been recently suggested that the child’s level of performance may also impact the association between adults’ beliefs and students’ self-concept of ability (pesu et al., 2016). in the study by pesu et al. (2016), the impact of teachers’ beliefs on first-grade students’ self-concept of pesu%et %al % % % | f l r ! ! 96! mathematics and reading ability was different depending on the level of student’s performance: among highperforming students, teachers’ beliefs had a positive impact on students’ self-concept of mathematics and reading ability, whereas among low-performing students, teachers’ beliefs did not have this positive impact. pesu et al. (2016) suggested that one explanation for this differential impact of teacher beliefs is that highperforming children are more prone to be affected by adults’ beliefs than low-performing children as (owing to their cognitive abilities) they are able to make more accurate interpretations of adults’ feedback and their own performance. also, bohlmann and weinstein (2013) argued that children’s cognitive reasoning skills affect the way they perceive, interpret, and attribute meaning to teachers’ actions. thus, it can be that students who have better cognitive skills are better able to interpret adults’ feedback overall. however, the differential role that parental beliefs have on student self-concept, depending on the student’s level of performance, has not to our knowledge been investigated among older children like secondary school students. finding out differences in the associations between parental beliefs and students’ self-concept of ability depending on students’ level of performance might have important implications for interventions. one further limitation of earlier research is that the majority of studies on the role of parental beliefs have focused on the role of mothers (for exceptions, see frome & eccles, 1998; gniewosz & noack, 2012; pesu et al., 2016), to the relative neglect of the role of fathers’ beliefs. however, it might be that mothers and fathers play a different role in their children’s self-concept development (frome & eccles, 1998; macgrath & repetti, 2000; maccoby, 1998). consequently, the second aim of the present study was to investigate the role of mothers’ and fathers’ beliefs about their adolescent children’s abilities in mathematics and literacy in the development of adolescents’ self-concept of ability during secondary school. further, possible differences in these associations depending on the adolescent’s gender, on the one hand, and level of performance, on the other, were investigated. the research questions were: a) to what extent finnish adolescents’ self-concept of mathematics and literacy ability change during secondary school? based on earlier literature, we hypothesized that self-concept of mathematics and literacy ability decline during grades 7-9 (nagy et al., 2010; wigfield et al., 1991). b) do parental beliefs concerning adolescents’ abilities predict the development of adolescents’ self-concept of literacy and mathematics ability during grades 7-9? we hypothesized that mothers’ and fathers’ beliefs positively predict adolescents’ subsequent self-concept of literacy and mathematics ability (e.g. frome & eccles, 1998;!gniewosz et al., 2012). c) are there differences in the associations between parental beliefs and adolescents’ selfconcept of abilities depending on adolescents’ a) gender, b) level of performance? we set two alternative hypotheses concerning the gender differences in the associations. as the first hypothesis, we hypothesized that the associations of mothers’ beliefs with adolescents’ selfconcept of ability are stronger among girls than among boys whereas the associations of fathers’ beliefs with adolescents’ self-concept of ability are stronger among boys than among girls, as suggested by the socialization model (maccoby, 1998). as the second hypothesis, we hypothesized that gender does not play a role in the connections between mothers’/fathers’ beliefs and self-concept of ability because previous studies have not found these kinds of gender differences (pesu at al., 2016; simpkins et al., 2012).! based one previous results by pesu et al. (2016), we also hypothesized that the role of mothers’/fathers’ beliefs in self-concept of ability is stronger among highthan low-performing students. pesu%et %al % % % | f l r ! ! 97! 2. method 2.1 participants the present study is a part of a longitudinal study (the jyväskylä entrance into primary school (jeps) study (nurmi & aunola, 1999–2009)) focusing on students’ academic and motivational development from the beginning of the school career until the end of comprehensive school. the sample comprised students from two medium-sized districts (urban or semi-urban areas) in central finland. the present study focuses on the data obtained from the adolescents and their parents when the former were in the 7th and 9th grades. the participants were 231 students in grade 7 and 221 in grade 9 (in grade 7: 114 girls and 117 boys, in grade 9:107 girls and 114 boys) and their mothers (n = 221) and fathers (n = 191). the adolescents filled in questionnaires on their self-concept of ability in the spring of the 7th grade and again in the spring of the 9th grade. performance in mathematics and literacy was assessed by tests in the spring term of the 7th grade. all questionnaires and tests were performed during regular school hours in classroom group situations by trained investigators. mothers and fathers were asked to fill in mailed questionnaires concerning their beliefs about their child’s performance in mathematics and literacy in the spring of the grade 7. the response rate was 96 % for mothers and 83% for fathers. the families participating in the study were to some extent more educated than the finnish population overall (statistics finland, 2010): 11.5% of mothers and 12.1% of fathers had no vocational education, 26.6% of mothers and 38.4% of fathers had a vocational education, and 61.9% of mothers and 49.6% of fathers had a degree from an institution of higher learning (e.g., polytechnic) or university. at the beginning of the 7th grade, 68,3% of the children were living in a nuclear family, 13,5% were living in a blended family, and 9,1% were living in a single parent household. 2.2 measures 2.2.1 self-concept of ability in literacy and mathematics students’ self-concept of ability in mathematics and literacy was measured with a questionnaire based on the ideas presented by eccles and wigfield (1995). students were asked to answer three questions, separately for mathematics and literacy (how good are you at mathematics / literacy? how good do you think you are at mathematics / literacy compared to the other students in your class? how hard are assignments related to mathematics / literacy for you (revised)) on a 5-point likert-scale. self-concept of ability in mathematics and literacy were scored separately by calculating the mean of the three items in each case. the cronbach’s alpha reliabilities for self-concept in mathematics in grade 7 and grade 9 were .87 and .89, respectively, and for self-concept in literacy .81 and .81, respectively. 2.2.2 adolescents’ performance in mathematics adolescents’ performance in mathematics was assessed with the group-administered ktlt test (räsänen & leino, 2005), which is a standardized math test for grades 7-9 (13-16 years). the test consists of 40 mathematical tasks (basic calculation and equation tasks, word problems, geometry tasks, measurement tasks), to be done individually. one point was given for each correct answer. the test was administered with a 45-minute time limit. the internal reliability of the test in the present data was .86. the internal reliability of the test in the normative data (n = 1,157) has been shown to be 0.88 (räsänen & leino, 2005). the test has also been shown to correlate with other measures of mathematical skills (r = 0.61–0.78, p < 0.001; räsänen & leino, 2005). pesu%et %al % % % | f l r ! ! 98! 2.2.3 adolescents’ performance in literacy adolescents’ performance in literacy was measured by three subtests taken from the test of word reading, spelling and reading comprehension (holopainen, kairaluoma, nevala, ahonen, & aro, 2004): a) in the first spelling error task, participants were asked to mark with a vertical line on 100 words typed on a sheet of paper as many spelling errors (an extra, missing, or wrong letter in a word) as they could identify in 3.5 minutes. the score was the number of correctly detected errors. the test-retest reliability for the subtest has been shown to be 0.83 (holopainen et al., 2004). b) in the second word chain test, the participants were asked to separate understandable words in a word chain by drawing a line between the words. a total of 100 words were presented in chains of four words with no spaces between them. the adolescents were allowed 3.5 minutes to find the end of one word and the beginning of a new word in each chain and to mark it with a vertical line. the test was scored as the number of correctly found words. the test-retest reliability of the subtest has been shown to be 0.84 (holopainen et al., 2004). c) in the reading comprehension test, the participants were asked to read a four-page long story (the hounds of the village, written by finnish author veikko huovinen), in which 52 words had been changed so that they did not fit in with the story (i.e., they were in contradiction with the meaning of the sentence, paragraph or larger text context). the participants were asked to underline all the inappropriate words they could find. a point was given for each correctly underlined word. the time limit for the subtest was 45 minutes. the sum score of the standardized three subtest scores was taken as the measure of literacy performance. the cronbach’s alpha reliability of the sum score was .81. 2.2.4 mothers’ and fathers’ beliefs about their child’s performance in literacy/mathematics mothers’ and fathers’ beliefs were measured at the end of the 7th grade with 2 items (e.g. how well do you think your child is doing in literacy/mathematics at the moment? how well do you think your child will do in literacy/mathematics in the future?) using a 4-point likert-scale. the cronbach alpha reliabilities of the scale were .92 (literacy) and .93 (mathematics) among mothers and .92 (literacy) and .93 (mathematics) among fathers. 2.2.5. analyses strategy the analyses were carried out along the following steps. first, the developmental changes in adolescents’ self-concepts of mathematics and literacy abilities from grade 7 to grade 9, and possible gender differences in these changes, was investigated by repeated measures anova. second, hierarchical regression analyses were carried out to examine whether parents’ beliefs about their adolescent children’s abilities in mathematics and literacy play a role in the development of adolescents’ self-concept of mathematics and literacy ability during secondary school and whether the role of parental beliefs differs according to the adolescent’s gender or level of performance. in these analyses, adolescents’ self-concept of ability in a specific school subject in the spring of the ninth grade (time 2) was predicted by their selfconcept of ability in that subject in the spring of the seventh grade (time 1), academic performance in that subject in the seventh grade (time 1), gender, and mothers’ or fathers’ beliefs about their child’s abilities in the spring of the seventh grade (time 1). each variable was entered stepwise in the analysis. the effects of mothers’ and fathers’ beliefs were tested in separate analyses. in order to determine whether any connection existed between mothers’/fathers’ beliefs and the adolescents’ subsequent level of self-concept of ability was influenced by the adolescents’ gender or by the adolescents’ level of performance, the related interaction terms (gender x belief or academic performance x belief) were added to the analysis in the last step. each interaction term was tested in a separate analysis. the analysis was carried out separately for self-concept of mathematics ability and self-concept of literacy ability. in order to be able to examine the effects of the pesu%et %al % % % | f l r ! ! 99! interaction terms, all the predictor variables were standardized before being added to the regression models and before calculating any interaction terms. the missing data was handled pairwise. 3. results the means (m), standard deviations (sd), and pearson product-moment-correlations of the study variables are shown in table 1. the results of repeated measures anova showed that adolescents’ self-concept of mathematics ability slightly declined from grade 7 (m = 3.41, sd = 0.82) to grade 9 (m = 3.30, sd = 0.94; f (1, 202) = 5.87, p < .05). their self-concept of literacy also slightly declined during this period (time 1: m = 3.54, sd = 0.70; time 2: m = 3.44, sd = 0.72; f (1, 203) = 3.86, p = .05). self-concept of literacy ability was higher among girls than boys across the measurement points (f (1, 202) = 21.14, p < .001), whereas self-concept of math ability was higher among boys than girls (f (1, 202) = 6.23, p < .05). no gender differences in the change in self-concepts from grade 7 to grade 9 were, however, evident. pesu%et %al % % % | f l r ! ! 100! table 1 intercorrelations, means, and standard deviations for the study variables variables 1. 2. 3. 4. 5. 6. 7. 8. 9. 10. 11. 1. self-concept literacy t1 2. self-concept literacy t2 .48c 3. self-concept math t1 .12 .20b 4. self-concept math t2 .08 .29c .71c 5. performance literacy t1 .38c .45c .23b .23b 6. performance math t1 .26c .27c .59c .51c .60c 7. gender -.21b -.32c .17b .14a -.36c .00 8. mother belief literacy t1 .46c .40c .29c .23b .57c .46c -.30c 9. mother belief math t1 .09 .16a .64c .62c .39c .64c -.01 .53c 10. father belief literacy t1 .35c .35c .25b .24b .51c .48c -.28c .56c .38c 11. father belief math t1 .01 .06 .54c .55c .32c .56c .02 .40c .67c .58c m 3.56 3.47 3.36 3.30 46.34 21.52 2.90 2.76 2.87 2.84 sd .69 .74 .84 .94 14.43 6.33 .79 .80 .69 .76 note. a = p < .05. b = p < .01. c = p < .001. t1 = time 1, t2= time 2 pesu%et %al % % % | f l r ! ! 101! 3.1 math-related self-concept the results of the hierarchical regression analyses for mathematics-related self-concept (see table 2) showed, first, that individual differences in self-concept of mathematics ability were relatively stable from grade 7 to grade 9. second, mothers’ (β = .28, p < .001) and fathers’ (β = .22, p < .001) beliefs about their child’s abilities predicted adolescents’ subsequent self-concept of mathematics ability at the end of grade 9, after controlling for the previous levels of self-concept of mathematics ability and mathematics performance: the higher the beliefs parents had about their child’s mathematics ability in grade 7, the better the adolescents’ self-concept of mathematics ability was in grade 9. finally, the connections between adolescents’ self-concept of mathematics ability and mothers’ belief in mathematics was found to be different depending on the adolescent’s level of mathematics performance (β = .13, p < .01). to examine this interaction effect further, aiken and west’s (1991) procedure was used. in this procedure, simple slopes for the mothers’ belief variable in the prediction of adolescents’ mathematics self-concept were calculated and presented using standardized scores separately for adolescents who showed either low (–1 sd) or high (+1 sd) levels of mathematics performance. the results are shown in figure 1. the results showed that among high-performing adolescents, mothers’ beliefs positively predicted subsequent self-concept of mathematics ability, whereas among low-performing adolescents this positive effect of mothers’ beliefs was weaker. the impact of parental beliefs was similar for boys and girls. table 2 the results of hierarchical regression analyses for mathematics related self-concept at time 2 (standardized betas) predictor mathematics related self-concept at time 2 step1 β step2 β step3 β step4 β step5 β a. self-concept (time 1) .71*** .71*** .62*** .50*** .47*** b. gender .02 .04 .06 .07 c. performance (time 1) .14* .03 .06 d. beliefs d1. beliefs mother (time 1) .28*** .30*** d2. beliefs father (time 1) .22** .23** e. interaction terms b x d1 -.16 b x d2 .05 c x d1 .13** c x d2 .04 r2 = .51 r2 = .51 r2 = .52 r2 = .55-.561 r2 = .55-.571 note 1. *** p < .001, ** p < .01, * p < .05 the effects of mothers’ and fathers’ beliefs were each tested in separate analyses. similarly all interaction terms were tested in separate analyses. 1 r2 varies depending on which variables are included into the model as predictor variables. pesu%et %al % % % | f l r ! ! 102! figure 1. the impact of mothers’ beliefs on students’ mathematics related self-concept among low, medium and high-performing students 3.2 literacy-related self-concept the results of hierarchical regression analyses (see table 3) showed, first, that individual differences in self-concept of literacy ability were relatively stable through grades 7-9. the results showed further that, after controlling for the previous level of self-concept and literacy performance, mothers’ or fathers’ beliefs did not predict adolescents’ self-concept of literacy ability. no parental belief x gender or parental belief x performance interaction effects were found either. pesu%et %al % % % | f l r ! ! 103! table 3 the results of hierarchical regression analyses for literacy related self-concept at time 2 (standardized betas) predictor literacy related self-concept at time 2 step1 β step2 β step3 β step4 β step5 β a. self-concept (time 1) .48*** .43*** .34*** .32*** .32*** b. gender -.23*** -.15* -.15* -.15* c. performance (time 1) .27*** .24** .24** d. beliefs d1. beliefs mother (time 1) .08 .08 d2. beliefs father (time 1) .08 .09 e. interaction terms b x d1 .21 b x d2 .03 c x d1 .02 c x d2 .08 r2 = .23 r2 = .28 r2 = .33 r2 = .34. r2 = .34 note 1. *** p < .001, ** p < .01, * p < .05 the effects of mothers’ and fathers’ beliefs were each tested in separate analyses. similarly all interaction terms were tested in separate analyses. 4. discussion the present study aimed to contribute to the literature on students’ self-concept of ability by examining, first, to what extent developmental changes in self-concept of mathematics and literacy abilities occur among finnish students across secondary school and, second, what role mothers’ and fathers’ beliefs play in the development of adolescents’ self-concept of mathematics and literacy ability during this period. furthermore, whether the possible associations of parental beliefs with adolescents’ self-concepts of abilities are influenced by adolescents’ gender or level of performance was investigated. the results showed first that both self-concept of mathematics and literacy ability slightly declined during secondary school. second, mothers’ and fathers’ beliefs about their child’s abilities predicted changes in the adolescents’ self-concept of ability, but only in mathematics: the higher the beliefs parents had about their child’s mathematics ability in grade 7, the better the adolescents’ subsequent self-concept of mathematics ability was in grade 9. furthermore, the role of mothers’ beliefs in adolescents’ self-concept of mathematics ability was found to be particularly strong among those adolescents who showed a high level of mathematics performance. finally, gender did not have an effect on the connections between parental beliefs and adolescents’ self-concept of ability development in mathematics or literacy. 4.1. the development of self-concept of ability the results of this study showed first that adolescents’ self-concept slightly declined during secondary school among both girls and boys. this result is consistent with previous results reported among us (nagy et al., 2010; wigfield et al., 1991), german (nagy et al., 2010) and australian students (nagy et pesu%et %al % % % | f l r ! ! 104! al., 2010) and suggest that also in the finnish context the secondary school years are an important time for the development of self-concept. the period of the transition to secondary school brings many changes in adolescents’ lives. their everyday social contexts change, the ways they get feedback at school change, and their frames of reference change (see wigfield et al., 2006). it is noteworthy, however, that in the present study the decline in self-concept was only minor. one explanation for there being only a slight decline in self-concept can be found in the finnish national curriculum guidelines, according to which teachers should focus on motivating both boys and girls equally to learn and to help them build a positive self-concept. thus, it is possible that since finnish teachers are aware of the importance of supporting self-concept construction, students receive much support from their school in this area, and thus show less of a decline in self-concept during adolescence. 4.2. the role of mothers’ and fathers’ beliefs in self-concept of ability development the results of the present study showed further that mothers’ and fathers’ beliefs predicted students’ self-concept of mathematics ability development across secondary school: the higher parental beliefs at the beginning of secondary school, the higher the adolescent’s self-concept in mathematics at the end of secondary school. the results are in line with eccles et al.’s expectancy-value theory which suggests that parental beliefs affect their children’s self-concept of ability (eccles parsons et al., 1982; frome & eccles, 1998; lau & pun, 1999; mcgrath & repetti, 2000). previous empirical research on the role of parents in students’ self-concept of ability development, however, has mainly focused on the role of mothers’ beliefs whereas that of fathers’ has received less attention. however, there is some evidence that mothers and fathers both play a role in their children’s self-concept of ability development in both mathematics and literacy (frome & eccles, 1998; gniewosz et al., 2014) at least among sixth-grade students (frome & eccles, 1998) and fifthto seventh-graders (gniewosz et al., 2014). the results of the present study also indicate that in secondary school both mothers’ and fathers’ beliefs have an impact on adolescents’ self-concept of ability in the domain of mathematics. this result adds to the literature since previous studies on the role of parents have not focused on this particular age group. however, the present results are inconsistent with those of previous research insofar as mothers’ and fathers’ beliefs did not play a role in their adolescents’ self-concept development in the domain of literacy. there are several possible explanations for this result. first, it is possible that achievement feedback is less clear in literacy than in mathematics, which would help explain why parents had a more evident role in adolescents’ self-concept development in mathematics. another possibility is that because mathematics is typically considered a more difficult school subject than literacy, and because there is a clearer declining trend in the self-concept of mathematics ability, the self-concept of mathematics ability is particularly prone to external feedback. third, previous studies showing connections between parental beliefs and their children’s self-concept of ability development have been conducted in cultural settings other than finland. research has revealed that finnish children attain fluency in native language reading and writing earlier, by the end of the first school year (seymour, aro, & erskine, 2003) than for example english-speaking children, whose rate of literacy skills development is more than twice as slow (seymour et al., 2003). slower literacy skills development has been attributed to fundamental linguistic differences in syllabic complexity and orthographic depth (seymour et al., 2003). for this reason, finnish parents might involve themselves less in their children’s literacy-related studies than mathematics studies, also later on. thus, parents might have less information about their children’s success in literacy than in mathematics and thus less influence on their children’s self-concept in literacy than in mathematics. the results of the present study showed, finally, that the role of mothers’ beliefs about their adolescent child’s mathematics ability was dependent on the level of the adolescent’s performance: mothers beliefs were positively related to their children’s self-concept of mathematics ability among high-performing adolescents but less so among low-performing adolescents. these results are in line with the results of pesu et al. (2016), who found that the role of teachers’ beliefs on first graders’ self-concept of mathematics and reading ability differed depending on the level of the student’s performance: teachers’ beliefs had a positive pesu%et %al % % % | f l r ! ! 105! impact on students’ self-concept of mathematics and reading ability only among high-performing students, not among low-performing students. there are several possible explanations for this result that mothers’ beliefs play a role, particularly among high-performing children. first, it might be that mothers communicate their beliefs, even where they are equally positive, differently to children whose levels of performance are different. thus, the effect of mothers’ beliefs would be different for children who perform differently at school. second, it could be that students interpret mothers’ cues about their beliefs differently depending on their level of performance. bohlmann and weinstein (2013) argued that children’s cognitive abilities influence their perceptions and interpretations of teachers’ actions. it is possible that children’s cognitive abilities influence their perceptions of external feedback overall. thus, it could be that high-performing adolescents are cognitively better able to accurately perceive and interpret mothers’ beliefs (see also, pesu et al., 2016). the present study showed that gender had no effect on the development of self-concept of ability in either mathematics or literacy. this result is consistent with previous studies showing similar patterns in the development of self-concept in boys and girls (e.g. nagy et al., 2010). the results of the present study showed further that gender did not influence the relationship between mothers’ and fathers’ beliefs and adolescents’ self-concept of ability development. since finland can be considered an egalitarian culture (chiu & klassen, 2009; chiu & klassen, 2010), there might be fewer gender differences overall. in an egalitarian culture, individuals are taught to view, value, and act towards one another as equals based on their common humanity (chiu & klassen, 2009; chiu & klassen, 2010). people learn these practices and values through formal and informal socialization, including through schooling (chiu & klassen, 2009; chiu & klassen, 2010). finnish culture is also considered as having little characteristics of a masculine culture (chiu & klassen, 2009; chiu & klassen, 2010). in masculine cultures males are typically favored in higher status roles, and women have lower income (cheung & chan, 2007). because gender roles are rigid in masculine cultures, this may lead, for example, girls to value mathematics learning less, devote less time to studying mathematics and have lower mathematics self-concept than boys (hofstede, 2003; wigfield, tonks, & eccles, 2004). as finland is considered an egalitarian and less masculine culture than many other cultures (chiu & klassen, 2009; chiu & klassen, 2010), finnish children grow up in a society where boys and girls are treated more equally than in cultures that are less egalitarian. this may explain why the present study did not show any gender differences in girls’ and boys’ self-concept of abilities and why the impact of parental beliefs was similar for boys and girls. 4.3. limitations this study has its limitations. first, the study was carried out in just one educational setting, finland. as it is possible that parental beliefs play a different role in students’ self-concept of abilities in different educational settings and cultures, further cross-cultural research on the topic is needed. second, even though a longitudinal procedure was used in the present study, it might be that some third factor not controlled for explains the predictions found. one should, therefore, be cautious before making any judgements about the possible causality of the results. third, the measure for mothers’ and fathers’ beliefs included two questions only. in future research measurements including more items to measure parental beliefs should be used to replicate the results found here. overall, the results of this study suggest that during secondary school finnish adolescents’ selfconcepts of mathematics and literacy ability undergo a slight decline, and that in the domain of mathematics both mothers’ and fathers’ beliefs about their children’s abilities play a role in the development of adolescents’ self-concept of ability. it is important that both mothers and fathers know what role they play in the formation of their children’s self-concepts of ability. because parents receive information about their children’s success at school indirectly, i.e. via grades and feedback from teachers, they might tend to think they do not have much of a role in their children’s academic-related life. it is important that schools and teachers in particular inform parents about the crucial role they can have on their children’s self-concept development in different academic domains. teachers and school personnel should inform parents about the pesu%et %al % % % | f l r ! ! 106! ways in which they, both mothers and fathers, could support their children and their children’s developing self-concepts. keypoints mothers’ and fathers’ child-specific ability beliefs predicted adolescents’ self-concept of mathematics ability. mothers’ and fathers’ beliefs did not predict adolescents’ self-concept of literacy ability. the relations between mothers’ beliefs and adolescents’ self-concept of mathematics ability varied according to adolescents’ performance: mothers beliefs were positively related to their children’s self-concept of mathematics ability among high-performing adolescents but less so among low-performing adolescents. references aiken, l. s., & west, s. g. (1991). multiple regression: testing and interpreting interactions. newbury park, ca: sage. arens, a. k., yeung, a. s., craven, r. g., & hasselhorn, m. (2011). the twofold multidimensionality of academic self-concept: domain specificity and separation between competence and affect components. journal of educational psychology, 103, 970-981. doi: 10.1037/a0025047 atkinson, j. w. (1964). an introduction to motivation. princeton, nj: van nostrand. aunola, k., leskinen, e., onatsu-arvilommi, t., & nurmi, j-e. (2002). three methods for studying developmental change: a case of reading skills and self-concept. british journal of educational psychology, 72, 343-364. doi: 10.1348/000709902320634447 bandura, a. (1986). social foundations of thought and action: a social cognitive theory. englewood cliffs, nj: prentice-hall. bohlmann, n., & weinstein, r. (2013). classroom context, teacher expectations, and cognitive level: predicting children’s math ability judgments. journal of applied developmental psychology, 34, 288298. doi: 10.1016/j.appdev.2013.06.003 bong, m., & skaalvik, e. m. (2003). academic self-concept and self-efficacy: how different are they really? educational psychology review, 15, 1–40. doi: 10.1023/a:1021302408382 brunner, m., keller, u., hornung, c., reichert, m., & martin, r. (2009). the cross-cultural generalizability of a new structural model of academic self-concepts. learning and individual differences, 19, 387-403. doi:10.1016/j.lindif.2008.11.008 caprara, g. v., vecchione, m., alessandri, g., gerbino, m., & barbaranelli, c. (2011). the contribution of personality traits and self-efficacy beliefs to academic achievement: a longitudinal study. british journal of educational psychology, 81, 78-96. doi: 10.1348/2044-8279.002004 chapman, j. w., tunmer, e. t., & prochnow, j. e. (2000). early reading-related skills and performance, reading self-concept, and the development of academic self-concept: a longitudinal study. journal of educational psychology, 92, 703-708. doi: 10.1037/0022-0663.92.4.703 cheung, h. y., & chan, a. w. h. (2007). how culture affects female inequality across countries. journal of studies in international education, 11, 157−179. doi: 10.1177/1028315306291538 chiu, m. m., & klassen, r. m. (2009). calibration of reading self-concept and reading achievement among 15-year-olds: cultural differences in 34 countries. learning and individual differences, 19, 372-386. doi:10.1016/j.lindif.2008.10.004 chiu, m. m., & klassen, r. m. (2010). relations of mathematics self-concept and its calibration with mathematics achievement: cultural differences among fifteen-year-olds in 34 countries. learning and instruction, 20, 2-17. doi:10.1016/j.learninstruc.2008.11.002 pesu%et %al % % % | f l r ! ! 107! dermitzaki, i., & efklides, a. (2000). aspects of self-concept and their relationship to language performance and verbal reasoning ability. the american journal of psychology, 113, 621–637. doi: 10.2307/1423475 eccles, j. s. (1993). school and family effects on the ontogeny of children’s interests, self-perceptions, and activity choices. in j. e. jacobs & r. dienstbier (eds.), developmental perspectives on motivation (pp. 145-208). university of nebraska press. eccles, j. s. (2007). where are all the women? gender differences in participation in physical science and engineering. in s. j. ceci & w. m. williams (eds.), why aren’t more women in science? top researchers debate the evidence (pp. 199–210). washington, dc: american psychological association. doi:10.1037/11546-016 eccles, j. s., adler, t. f., futterman, r., goff, s. b., kaczala, c. m., meece, j. l., & midgley, c. (1983). expectancies, values, and academic behaviors. in j. t. spence (ed.), achievement and achievement motivation (pp. 75–146). san francisco: w. h. freeman. eccles parsons, j., adler, t. f., & kaczala, c. m. (1982). socialization of achievement attitudes and beliefs: parental influences. child development, 53, 310–321. doi: 10.2307/1128973 eccles, j. s., & jacobs, j. e. (1987). social forces shape math attitudes and performance. in m. r. walsh (ed.), the psychology of women: ongoing debates (pp. 341-354). new haven, us: yale university press. eccles, j. s., & wigfield, a. (1995). in the mind of the actor: the structure of adolescents’ achievement task values and expectancy-related beliefs. personality and social psychology bulletin, 3, 215-225. doi: 10.1177/0146167295213003 frome, p. m., & eccles, j. s. (1998). parents’ influence on children’s achievement-related perceptions. journal of personality and social psychology, 74, 435–452. doi: 10.1037/0022-3514.74.2.435 gniewosz, b., eccles, j. s., & noack, p. (2012). secondary school transition and the use of different sources of information for the construction of the academic self-concept. social development, 21, 537-557. doi: 10.1111/j.1467-9507.2011.00635.x gniewosz, b., eccles, j. s., & noack, p. (2014). early adolescents’ development of academic self-concept and intrinsic task value: the role of contextual feedback. journal of research on adolescence, 25, 1-15. doi: 10.1111/jora.12140. gniewosz, b., & noack, p. (2012). mamakind or papakind? [mom’s child or dad’s child]: early adolescents’ parental preferences in intergenerational academic value transmission. learning and individual differences, 22, 544-548. doi:10.1016/j.lindif.2012.03.003 gunderson, e. a., ramirez, g., levine, s. c., & beilock, s. i. (2012). the role of parents and teachers in the development of gender-related math attitudes. sex roles, 66, 156-166. doi: 10.1007/s11199-011-99962 hofstede, g. (2003). culture’s consequences. thousand oaks, ca: sage. holopainen, l., kairaluoma, l., nevala, j., ahonen, t., & aro, m. (2004). lukivaikeuksien seulontatesti nuorille ja aikuisille. [dyslexia screening test for youth and adults]. jyväskylä: jyväskylän yliopistopaino. jacobs, j. e. (1991). influence of gender stereotypes on parent and child mathematics attitudes, journal of educational psychology, 83, 518–527. doi: 10.1037/0022-0663.83.4.518 jacobs, j. e., & eccles, j. s. (2000). parents, task values, and real-life achievement-related choices. in c. sansone & j. m. harackiewicz (eds.), intrinsic and extrinsic motivation: the search for optimal motivation and performance (pp. 405–439). san diego, ca: academic press, inc. jacobs, j. e., lanza, s., osgood, d. w., eccles, j. s., & wigfield, a. (2002). changes in children’s selfcompetence and values: gender and domain differences across grades one through twelve. child development, 73, 509-527. doi: 10.1111/1467-8624.00421 lau, s., & pun, k-t. (1999). parental evaluations and their agreement: relationship with children’s selfconcepts. social behavior and personality: an international journal, 27, 639–650. doi: 10.2224/sbp.1999.27.6.639 maccoby, e. e. (1998). the two sexes: growing up apart, coming together. cambridge, ma: belknap press. pesu%et %al % % % | f l r ! ! 108! marsh, h. w. (1990). a multidimensional, hierarchical self-concept: theoretical and empirical justification. educational psychology review, 2, 77-172. doi: 10.1007/bf01322177 marsh, h. w., byrne, b. m., & shavelson, r. (1988). a multifaceted academic self-concept: its hierarchical structure and its relation to academic achievement. journal of educational psychology, 80, 366-380. doi: 10.1037/0022-0663.80.3.366 marsh, h. w., trautwein, u., lüdtke, o., köller, o. & baumert, j. (2005). academic self-concept, interest, grades, and standardized test scores: reciprocal effects models of causal ordering. child development, 76, 397-416. doi: 10.1111/j.1467-8624.2005.00853.x marsh, h. w., & yeung, a. s. (2001). an extension of the internal/external frame of reference model: a response to bong (1998). multivariate behavioral research, 36, 389-420. doi: 10.1207/s15327906389420 mcgrath, e. p., & repetti, r. l. (2000). mothers’ and fathers’ attitudes toward their children’s academic performance and children’s perceptions of their academic competence. journal of youth and adolescence, 29, 713–723. doi: 10.1023/a:1026460007421 nagy, g., watt, h. m. g., eccles, j., trautwein, u., lüdtke, o., & baumert, j. (2010). the development of students’ mathematics self-concept in relation to gender: different countries, different trajectories? journal of research on adolescence, 20, 482-506. doi: 10.1111/j.1532-7795.2010.00644.x pesu, l., viljaranta, j., & aunola, k. (2016). the role of parents’ and teachers’ beliefs in children’s selfconcept development. journal of applied developmental psychology, 44, 63-71. doi: 10.1016/j.appdev.2016.03.001 phillips, d. a. (1987). socialization of perceived academic competence among highly competent children. child development, 58, 1308–1320. doi: 10.2307/1130623 pintrich, p. r. & schunk, d. h. (2008). motivation in education. theory, research and applications (3rd ed.). new jersey: pearson education. preckel, f., niepel, c., schneider, m., & brunner, m. (2013). self-concept in adolescence: a longitudinal study on reciprocal effects of self-perceptions in academic and social domains. journal of adolescence, 36, 1165-1175. doi: 10.1016/j.adolescence.2013.09.001 räsänen, p., & leino, l. (2005). ktlt. laskutaidon testi. opas yksilö-tai ryhmämuotoista arviointia varten. seymour, p. h., aro, m., & erskine, j. m. (2003). foundation literacy acquisition in european orthographies. british journal of psychology, 94, 143-174. doi: 10.1348/000712603321661859 shavelson. r. j., hubner, j. j., & stanton, g. c. (1976). self-concept: validation of construct interpretations. review of educational research, 46, 407-441. doi: 10.2307/1170010 simpkins, s. d., fredricks, j. a., & eccles, j. s. (2012). charting the eccles’ expectancy-value model from mothers’ beliefs in childhood to youths’ activities in adolescence. developmental psychology, 48, 1019-1032. doi: 10.1037/a0027468 statistics finland (2010). differences between municipalities in educational level of population were still considerable in 2009. helsinki: statistics finland. retrieved march 20, 2016, from http:// www.stat.fi/til/vkour/2009/vkour_2009_2010-12-03_tie_001_en.html valentine, j. c., dubois, d. l., & cooper, h. (2004). the relation between self-beliefs and academic achievement: a meta-analytic review. educational psychologist, 39, 111-133. doi: 10.1207/s15326985ep3902_3 välijärvi, j., kupari, p., linnakylä, p., reinikainen, p., sulkunen, s., törnroos, j., & arffman, i. (2007). the finnish success in pisa-and some reasons behind it: pisa 2003. 2. jyväskylän yliopisto, koulutuksen tutkimuslaitos. watt, h. m. g. (2004). development of adolescents’ self-perceptions, values, and task-perceptions according to gender and domain in 7ththrough 11th-grade australian students. child development, 75, 15561574. doi: 10.1111/j.1467-8624.2004.00757.x wigfield, a., & eccles, j. s. (2000). expectancy-value theory of achievement motivation. contemporary educational psychology, 25, 68–81. doi: 10.1006/ceps.1999.1015 wigfield, a., eccles, j. s., maciver, d., reuman, d. a., & midgley, c. (1991). transitions during early adolescence: changes in children’s domain-specific self-perceptions and general self-esteem across the pesu%et %al % % % | f l r ! ! 109! transition to junior high school. developmental psychology, 27, 552-565. doi: 10.1037/00121649.27.4.552 wigfield, a., eccles, j. s., schiefele, u, roeser, r. w., & davis-kean, p. (2006). development of achievement motivation. in n. eisenberg, w. damon & r. m. lerner (eds.), handbook of child psychology: vol. 3, social, emotional, and personality development (6th ed.) (pp. 933-1002). hoboken, nj, us: john wiley & sons inc. wigfield, a., eccles, j. s., yoon, k. s., harold, r. d., arbreton, a. j. a., freeman-doan, c., & blumenfeld, p. c. (1997). change in children’s competence beliefs and subjective task values across the elementary school years: a 3-year study. journal of educational psychology, 89, 451-469. doi: 10.1037/0022-0663.89.3.451 wigfield, a., tonks, s., & eccles, j. s. (2004). expectancy value theory in cross-cultural perspective. in d. m. mcinerney & s. van etten (eds.), big theories revisited (pp. 165-198). charlotte, nc:iap. microsoft word boonen et al_publication.docx ! ! ! ! ! frontline)learning)research)vol.4)no.)5)(2016))55)<)82) issn)2295<3159)) ! it’s not a math lesson we’re learning to draw! teachers’ use of visual representations in instructing word problem solving in sixth grade of elementary school anton j. h. boonen1, helen c. reed, judith schoonenboom & jelle jolles vrije universiteit amsterdam, the netherlands article received 3 march / revised 30 may / accepted 13 october / available online ??? december abstract non-routine word problem solving is an essential feature of the mathematical development of elementary school students worldwide. many students experience difficulties in solving these problems due to erroneous problem comprehension. these difficulties could be alleviated by instructing students how to use visual representations that clarify the problem structure and the relations between solution-relevant elements (so-called visual-schematic representations). research shows that instructional effectiveness depends largely on teachers’ mathematical knowledge for teaching. teachers’ knowledge of visual representations is therefore essential to instructing word problem comprehension in this way. as there is little to no literature investigating teachers’ practices in this area, the goal of the present study is to examine teachers’ use of visual representations to support non-routine word problem solving. eight mainstream elementary school teachers implemented an innovative approach focused on the use of visual-schematic representations. after a short training, teachers were able to produce these representations during instruction. however, some teachers seemed unclear about what these representations comprise and what function they serve within the word problem solving context. teachers seemed to base their use of representations on personal preferences rather than on an optimal fit with the word problem characteristics. these aspects need to be addressed in teacher training and professional development programs.! this study makes an unique contribution to research in the important and problematic area of word problem solving in regular classrooms. the results of this study are relevant for educational researcher, teachers, and teacher educators who deal with difficulties in instructing mathematical word problems. keywords: word problem solving instruction; visual representations; teacher modelling; mathematical knowledge for teaching !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 1 anton j. h. boonen and helen c. reed (both authors contributed equally), dep. of pedagogical & educational sciences, section educational neuroscience, faculty of behavioural and movement sciences & learn! research institute for learning and education, vrije univ. amsterdam, van der boechorststraat 1, 1081 bt amsterdam, the netherlands. email: a.j.h.boonen@vu.nl doi: http://dx.doi.org/10.14786/flr.v4i5.245 boonen%et %al % % % ! ! ! 56! ! 1. introduction on one side of a scale there are three pots of jam and a 100g weight. on the other side there are a 200g and a 500g weight. the scale is balanced. what is the weight of a pot of jam? in contemporary math education, word problems like this are frequently offered to students. learning to solve these so-called non-routine word problems is an essential feature of mathematical development (depaepe, de corte, & verschaffel, 2010; jiminéz & verschaffel, 2014; swanson, lussier, & orosco, 2013). in this study, non-routine word problems are defined as challenging problems set in realistic contexts, that require understanding, analysis and interpretation. they are not simple computational tasks embedded in words; they require an appropriate selection of strategies and decisions that lead to a logical solution (van garderen & montague, 2003). students’ difficulties in solving non-routine word problems are common in contemporary math classrooms (boonen, van der schoot, van wesel, de vries, & jolles, 2013; hegarty & kozhevnikov, 1999; krawec, 2010; van garderen & montague, 2003) and are widely recognized by both teachers and mathematics education researchers (boonen, van wesel, jolles, & van der schoot, 2014; prenger, 2005; van garderen, 2006; van garderen & montague, 2003). several instructional programs, for example cognitive strategy instruction, have been developed that provide support in word problem solving for low-performing students (see jitendra, 2002; jitendra et al., 2013; montague, 2003; montague, warger, & morgan, 2000). these programs have often been used in research settings involving researchers working with individuals or small groups. however, also instructional programs that are implemented in mainstream classrooms are developed and tested (e.g., boonen & jolles, 2015; csíkos, szitányi, & kelemen, 2012; verschaffel, de corte, van vaerenbergh, lasure, bogaerts & ratinckx, 1999; willis & fuson, 1988). in the study of verschaffel et al., (1999), for example, a learning environment for teaching and learning how to model and solve mathematical word problems was developed and tested in four mainstream classrooms of 5th grade students. this is important, given that mainstream schools in many countries are becoming more inclusive, with greater numbers of students with mild to severe learning difficulties in classrooms (jitendra & star, 2012; sharma, loreman, & forlin, 2012). dutch mainstream elementary school teachers, for example, report on average a quarter of their students as having special educational needs (van der veen, smeets, & derriks 2010). it would therefore be of benefit to teachers if they had instructional approaches at their disposal designed to support learning in areas that many students find difficult in this case, nonroutine word problem solving. so, contrary to previous studies that are implemented in mainstream classrooms, the present study is not focused on student modelling and solving mathematical word problems (i.e., csíkos et all., 2012; verschaffel et al., 1999) but examines how teachers implement an innovative approach to instructing word problem solving based on the use of visual representations during whole-class teaching in mainstream sixth grade classrooms. in sixth grade, students are expected to be able to solve a wide variety of non-routine problems of increasing difficulty. the challenge in enabling students to tackle such problems is considerable, so instructional support at this grade level is particularly appropriate. 1.1. what is difficult about solving non-routine word problems? non-routine word problems cannot be solved using any fixed algorithm or set of prescribed procedures (elia, van den heuvel-panhuizen, & kovolou, 2009; pantziara, gagatsis, & elia, 2009). in other words, in non-routine word problem solving there is no convenient model or solution path that is readily available to apply to solving a problem. according to verschaffel, greer and de corte (2000), solving word problems in general involves several phases, starting with understanding the situation that is described in the word problem boonen%et %al % % % ! ! ! 57! ! and forming a situation model. the first phase concerns (1) reading the word problem; (2) understanding the problem text; and (3) make a complete and coherent representation of the situation described in the problem. the next phase involves mathematizing, i.e., translating the situation model into mathematical form by (4) identifying the relevant numerical and linguistic components; (5) identifying the relationships between these components; and (6) expressing these by mathematical equations. the last three phases of the word problem solving process are (7) executing the mathematical equations; (8) interpret the outcome and formulate an answer; and (9) evaluate the solution (krawec, 2010; lewis & mayer, 1987; verschaffel et al., 2000). although this study is particularly interested in making a complete and coherent representation of the situation described in the problem of a word problem, which is the first phase in the word problem solving process, all the other phases described in verschaffel et al. (2000) are taken into account in the teaching intervention that is implemented. altogether, solving non-routine word problems thus involves carrying out and integrating several cognitive activities that involve a non-trivial amount of related information. research shows that the difficulties experienced by many students in solving word problems arise not from their inability to execute computations, but from difficulties in understanding the problem text, identifying solution-relevant components and the relations between them, and making a complete and coherent representation of the situation described in the problem (boonen et al., 2013; carpenter, corbitt, kepner, lindquist, & reys, 1981; cummins, kintsch, reusser & weimer, 1988; krawec, 2010; lewis & mayer, 1987). hence, erroneous word problem solutions are frequently a consequence of errors in the problem comprehension phase rather than in the problem solution phase. 1.2. previous research into supporting word problem solving from the perspective presented above, providing support for the problem comprehension phase should be particularly beneficial to students’ problem solving performance. yet, based on our observations in educational practice, word problem solving in the classroom as well as teacher training often focuses on the solution phase. existing research-based programs to support word problem solving of low-performing students incorporate both problem solving phases in the form of prescribed cognitive strategies (i.e., reading and understanding the problem, analyzing the information presented, developing logical solution plans, evaluating solutions) presented in a sequence of steps (jitendra & star, 2012; jitendra et al., 2009; krawec, 2012; montague et al., 2000). in these programs, support for the problem comprehension phase typically includes strategies for understanding the problem text (e.g., by paraphrasing the text and underlying relevant information) and strategies for identifying and representing the underlying problem structure by means of a visual representation. the assumption is that a visual representation should clarify the problem structure by making the numerical, linguistic and spatial relations between solution-relevant elements visible, which consequently facilitates understanding of the problem and identification of the computations to be performed (boonen et al., 2014; krawec, 2010, 2012). thus, using visual representations during problem comprehension could be an effective way to support word problem solving (van garderen & montague, 2003). two current research-based approaches to using visual representations in the problem comprehension phase can be distinguished. the first is to provide students with specific visual representations for specific types of problem, namely routine word problems (i.e., word problems in which a convenient model or solution path is readily available to apply while solving a problem). in these word problems students are then encouraged to reflect on the similarities and differences between problem types and the corresponding visual representations (jitendra, dipipi, & perron-jones, 2002; jitendra & star, 2012; jitendra et al., 2009). although this approach can be successful for teaching students how to tackle routine problems with an identical structure (e.g., ‘compare’ problems as: mary has 5 marbles. john has 8 marbles. how many more marbles does john have than mary?), it boonen%et %al % % % ! ! ! 58! ! is usually not possible to match non-routine word problems to a fixed representation type. teaching students to use only one specific visual representation for each type of problem is, moreover, risky, as it may lead to an inflexible and rigid use of representation strategies (jitendra, griffin, haria, leh, adams, & kaduvettoor, 2007; van dijk, van oers, & terwel, 2003; van dijk, van oers, terwel, & van den eeden, 2003a). the second approach – typically for non-routine word problems – defines visual representation in general heuristic terms: this should be done either mentally or by making a drawing. indeed, it is characteristic for non-routine problems that they cannot be represented in a prescribed way; in such conditions, the use of a heuristic approach may seem appropriate. it is a false assumption, however, that students know how to translate such a heuristic into a useful visual representation. students generally do not know what to draw, when, under which circumstances and for which types of problems (jitendra et al., 2007; jitendra et al., 2009). in summary, there are important shortcomings in current approaches to using visual representations to support comprehension of non-routine word problems. on the one hand, approaches that teach fixed visual representations for problem solving are insufficient to deal with the situationspecific structure of non-routine problems. on the other hand, heuristic approaches do not provide sufficient guidance for students to produce useful visual representations. 1.3. visual representations for comprehension of non-routine word problems these shortcomings could be addressed by teaching students to produce visual representations that accurately depict the situation-specific structure of non-routine word problems. such representations should present a complete and coherent model of the relations between all solutionrelevant problem components. we refer to these as accurate visual-schematic representations. it is important to note that such representations can incorporate standard mathematical models, such as a bar model, pie chart or number line, but that they often comprise freely constructed drawings. examples of an accurate bar model and ‘own’ construction are shown in table 1(a) and (b) respectively for the weighing scale problem presented earlier. research shows that accurate visual-schematic representations facilitate problem comprehension, help to identify the required calculation processes and thereby contribute to successful problem solving (see e.g., boonen et al., 2013, 2014; hegarty & kozhevnikov, 1999; van garderen, 2006; van garderen & montague, 2003). for example, the visual-schematic bar model in table 1(a) demonstrates that an accurate depiction of problem structure can greatly reduce calculation demands: simply by depicting the given quantities in the correct relation to each other, it can instantly be seen that each pot of jam must weigh 100g + 100g. this stands in contrast to more commonly known visual representations, namely pictorial and arithmetical representations, both of which frequently accompany word problem solving in mathematical text books. pictorial representations contain a detailed image of some element of the problem text (e.g., an object or a person) without identifying relations between problem elements or the required calculations (table 1[e]). these representations have been found to negatively influence the word problem solving process (boonen et al., 2014; berends & van lieshout, 2009; hegarty & kozhevnikov, 1999; krawec, 2010, 2012; van garderen, 2006; van garderen & montague, 2003). arithmetical representations (e.g., proportion tables, table 1[f]) are intended to support the calculation processes required to compute answers. this type of representation is generally introduced in the problem solution phase but does not contribute to problem comprehension; thus, when the problem is not well understood, arithmetical representations frequently contain erroneous information and lead to incorrect answers (boonen et al., 2013; cummins et al., 1988; krawec, 2010). boonen%et %al % % % ! ! ! 59! ! table 1 examples of types of visual representation for the weighing scale problem representation type example (a) accurate visual-schematic: bar model (b) accurate visual-schematic: own construction (c) inaccurate visual-schematic: bar model (d) inaccurate visual-schematic: own construction (e) pictorial (f) arithmetical: proportion table ! 100 200 500 3 x boonen%et %al % % % ! ! ! 60! ! 1.4. teachers' knowledge of visual representations theoretical perspectives on teaching students mathematics increasingly emphasize that teachers must possess adequate ‘mathematical knowledge for teaching’ (mkt), namely the collective mathematical knowledge, skills and attitudes needed to support student learning (ball, thames, & phelps, 2008). this includes both pedagogical content knowledge (knowing a variety of effective ways to present and represent mathematical content, taking account of learner characteristics and common misconceptions and difficulties in learning the subject matter; shulman, 1987) as well as specialized mathematical subject matter knowledge a teacher needs to teach particular content. with respect to visual representations, teachers need to recognize what is involved in using particular representations and when they are appropriate to use (ball et al., 2008). in the context of the present study, teachers need to be aware that visual-schematic representations should be used to support the first phase of the word problem solving process (i.e., problem comprehension) and that arithmetical representations are only appropriate in the problem solution phase. moreover, teachers should be able to use multiple representations and link different representations to each other and to underlying ideas (ball et al., 2008; dreher & kuntze, 2015). research shows that the effectiveness of teachers’ mathematical instructional practices depends largely on the quality of teachers’ mkt (ball et al., 2008; hill et al., 2008). unfortunately, there is relatively little literature about teachers’ understanding of and competence with visual representations, and we have found no literature that specifically addresses how teachers teach students to construct visual representations to support word problem comprehension. it cannot be assumed that teachers are able to do this; indeed, visual representations are reported to be problematic for teachers. for example, orrill, sexton, lee and gerde (2008) reported that middle grade mathematics teachers in the us are uncomfortable with visual representations, and that this relates to their incomplete knowledge about using and interpreting such representations. turner (2008) found that beginning elementary school teachers in the uk frequently have difficulty in choosing and using visual representations (e.g., number lines and hundred squares), and that their choices are based on superficial attractiveness rather than the suitability of the representations for the mathematics they want children to learn. in germany, dreher and kuntze (2015) found that even secondary school mathematics teachers do not fully understand the role and use of different forms of visual representations for learning about and teaching fractions. in the present case, in which visual representations should be used to support comprehension of non-routine word problems, teachers may not know what kind of representations should be made or in which phase of the problem solving process to use them. teachers may also have difficulty in constructing visual representations accurately (i.e., correctly and completely). incorrect and/or incomplete visual-schematic representations are referred to as inaccurate visual-schematic representations (e.g., table 1[c] and [d]). furthermore, research shows that it is more effective to teach students to construct their own visual representations than to provide them ready-made, as this contributes to skill adaptivity (van dijk et al., 2003; van dijk et al., 2003a). thus, teaching needs to focus on the construction process (i.e., how to make the representation), rather offering a representation as a given entity. finally, teachers should encourage students to use visual representations in a diverse, adaptive/flexible and functional way. this refers to being able to use different kinds of visual representations and to switch between them such that the representation fits the structural characteristics of the problem and is useful for helping to solve it. indeed, the ability to deal flexibly with multiple representations and move adaptively between them is seen as being essential for successful mathematical problem solving (e.g., acevedo nistal, van dooren, clarebout, elen & verschaffel, 2009; dreher & kuntze, 2015). however, as noted above, these aspects are frequently problematic for teachers (dreher & kuntze, 2015; orrill et al., 2008; turner, 2008). in short, although using accurate visual-schematic representations to support the problem boonen%et %al % % % ! ! ! 61! ! comprehension phase of word problem solving has considerable potential for improving problem solving performance, research is needed that examines how teachers implement an approach centred on the use of these visual representations. it is important to establish this point, as it is critical to the viability of this approach for supporting word problem solving in schools, as well as providing important indications for teacher professionalization programs. 1.5. the present study the goal of the present study is to examine teachers’ use of visual representations when implementing a teaching intervention for supporting non-routine word problem solving that focuses on constructing accurate visual-schematic representations. the teaching intervention is embedded within a sequence of problem solving steps (comparable to the programs mentioned above) that reflect the problem solving phases described earlier. it is important to note that, just as both problem solving phases are essential for effective problem solving, so all steps are intended to be carried out fully in the prescribed order for each problem treated. the general goal of the teaching intervention is to teach students cognitive strategies for solving non-routine word problems. the specific goal is to teach students to construct accurate visualschematic representations of problem structure and to encourage them to select and use visual representations in a functional (i.e., helpful for solving the problem) and diverse and adaptive/flexible way. diversity refers to the varied use of visual representations. adaptivity/flexibility, on the other hand, refers to choosing flexibly between available representations (i.e., teachers’ capability to decide on the representation that is the most appropriate, see heinze, star & verschaffel, 2009). an overview of the steps and the relation of each step to problem solving phase is provided in box 1 (based on montague et al., 2000). the key interest of the present study, that is the construction of accurate visual-schematic representations (implemented in the third, i.e., visualization step), is indicated in bold print. the teaching intervention comprises eight lessons that make use of teacher modelling (i.e., the teacher who acts as the instructor of the intervention thinks aloud while demonstrating a cognitive activity), student modelling (i.e. sixth grade pupils think aloud while demonstrating a cognitive activity under the guidance of their teacher) and independent student practice (see methods section). we focus on teacher modelling, as teachers’ use of visual representations is expected to be most visible here. an important aspect of the teaching intervention is that teachers are encouraged to implement it in a way that is compatible with their own manner of teaching (rogers, 2003). this makes it possible to study natural diversity in teachers’ behaviours. boonen%et %al % % % ! ! ! 62! ! box 1: word problem solving steps problem solving phase step content step 1. problem comprehension: understanding text read the problem carefully all the way through each sentence of the text is studied for comprehension and not just to search for numbers and key words (such as more than, times, as much as, etc.). step 2. problem comprehension: understanding text meaning, identifying problem components, identifying relations understand the text: " put it in your own words " imagine the situation " underline important information " what is being asked? deep understanding of the text is stimulated by a sequence of four substeps. the text is paraphrased (i.e., put into own words), the described situation is imagined mentally, solution-relevant information needed for solving the problem is underlined, and the problem solver asks him/herself what question is to be answered. step 3. problem comprehension: representing problem structure visualize the problem structure: make a drawing of the problem situation an accurate visual-schematic representation of the text is made. this contains correct and complete depictions of spatial, linguistic and numeric relations between all solution-relevant elements of the text. step 4. problem solution: determining operations hypothesise a plan to solve the problem: " what kind of problem is it? (+, -, x, : ) " what do you need to calculate? the number of solution steps and type(s) of operation (addition, subtraction, multiplication and/or division) are derived from the schematic representation. the required calculation is written down in standard symbolic notation. step 5. problem solution: executing computations compute the required operation the required calculations are performed. step 6. check your answer the computation is checked and it is considered whether the result is a plausible answer to the question asked. boonen%et %al % % % ! ! ! 63! ! 1.6. research questions to meet the goals of the present study, the following research questions are posed: a) what attention do teachers give to visualization when implementing the teaching intervention, when do they use visual representations in the word problem solving process and to what purpose? given the focus of the teaching intervention on the use of visual representations, teachers are expected to pay most attention to the visualization step of the problem solving process, to use visual representations to structure problem elements and the relations between them, and to embed this within the full sequence of prescribed steps. however, it is possible that teachers focus on other parts of the problem solving process, that they use visual representations at other points of the process and/or for other purposes (e.g., illustrating unfamiliar words in the text, calculating answers), or that they do not use visual representations at all. it is also possible that steps are omitted, combined or performed out of order, which may result in the visualization step not being embedded in the prescribed sequence as intended. b) what kinds of visual representations do teachers use and how diverse and adaptive/flexible is this use of visual representations? teachers are expected to demonstrate a varied use of visual representations (i.e., diversity) and to offer different kinds of representations for solving a problem (i.e., adaptivity/flexibility). however, it is possible that teachers use representations in a limited and fixed way and/or that they do not consider different ways of representing problems. c) what is the quality of the representation process and of the visual representations used? teachers are expected to model the representation process transparently, correctly and completely. however, it is possible that teachers do not make their reasoning transparent (e.g., offering a visual representation without explaining which elements of the problem should be represented or without explaining how the representation can be used to solve the problem), and/or that their reasoning is incorrect (e.g., naming and/or using visual representations wrongly) and/or incomplete (e.g., naming a visual representation that can be used but not indicating why). teachers are also expected to construct visual representations that correctly and completely depict the relations between all solution-relevant problem components. however, the visual representations made could be incorrect (i.e., containing erroneous problem components or relations) and/or incomplete (i.e., missing components and/or relations). the visual representations used are also expected to be suitable and useful (i.e., functional) for solving the problem. however, it is possible that visual representations do not fit problem characteristics well (e.g., a number line for solving a problem about percentages) and/or that they contain excess information that is not relevant for and could interfere with solving the problem. 1.7. study approach and relevance the present research is performed within the context of a teaching intervention for supporting nonroutine word problem solving that involved eight mainstream teachers recruited through purposive sampling. these teachers had a positive attitude towards mathematics, were confident about teaching mathematics and using visual representations, and were motivated to participate in and contribute to research in this area. it is well established that a lack of motivation and interest can severely impact the way in which teachers implement educational innovations in regular classrooms. the personal willingness of teachers to adopt and integrate innovations into their classroom practice is of crucial importance (evers, brouwers, & tomic, 2002; ghaith & yaghi, 1997; hermans, tondeur, van braak, & valcke, 2008; rogers, 2003). thus, by limiting participation to individuals with these qualities, results are obtained under favourable conditions in boonen%et %al % % % ! ! ! 64! ! which teacher behaviour is not negatively influenced by motivational factors. this allows behaviours to be analysed on the basis of the specified criteria, with known affective confounders excluded. as far as we know, this is the first study to focus on mainstream teachers’ use of visual-schematic representations in whole-class word problem solving instruction. the study thus makes a unique contribution to research in the area of word problem solving, as well as to literature on teachers’ understanding of and competence with visual representations. insights in these areas are important for improving instructional support and teacher training, given the relation between teachers’ mkt and instructional effectiveness (ball et al., 2008) together with the emphasis on word problem solving and the use of visual representations stated in the mathematics curricula of many countries and international assessments (e.g., department for education, 2013; mullis & martin, 2013; nctm, 2000; noteboom, 2009; oecd, 2013). 2. methods 2.1. participants directors of elementary schools in the central provinces of the netherlands were approached with information about the teaching intervention and a request to participate in this research. eight mainstream sixth grade teachers from four schools subsequently volunteered to participate. these teachers indicated that they were motivated to implement the teaching intervention and contribute to this research, that they had a positive attitude towards mathematics, that they were confident about teaching mathematics and that they believed themselves to be competent in using visual representations in the math lesson. table 2 presents the background characteristics of the participating teachers. table 2 background characteristics of participating teachers sex highest qualification levela number of years teaching days per week teaching school typeb teacher 1 f 1 20 5 2 teacher 2 f 1 6 5 2 teacher 3 m 1 13 4 2 teacher 4 m 1 27 4 2 teacher 5 f 2 14 5 1 teacher 6 m 1 13 4 2 teacher 7 m 1 39 5 2 teacher 8 f 1 6 3 2 note. a 1 = bachelor of applied sciences (teacher education for primary schools) 2 = master of applied sciences (teacher education for primary schools); b 1 = urban 2 = provincial parents of students in the classes of the participating teachers were informed that their child would participate in the study and that they could withhold permission for their child to participate. no parents took boonen%et %al % % % ! ! ! 65! ! up this option; thus, all participating classes were intact. 2.2. context of the research: the teaching intervention three weeks prior to commencement of the teaching intervention, teachers attended one afternoon session of 60 minutes that presented the aims and rationale of the study and introduced the intervention. the purpose of a second afternoon session of 120 minutes two weeks later was to reveal how to execute the problem solving steps. both training sessions were facilitated by one of the researchers who executed the research. professional development materials that were presented during the second afternoon session contained: (a) an overview of the most important national and international literature regarding word problem solving and word problem solving instruction; (b) a presentation and explanation of each of the problem solving steps; (c) a description of how to perform each step on the basis of several examples. furthermore, teachers received an elaborated lesson protocol that contained a fully scripted model of how to apply the steps for each word problem to be treated. the problems were obtained from regular math textbooks and were fully authentic. rather than use the scripts verbatim, however, teachers were encouraged to use their own explanations and elaborations during the intervention and to implement it in a way that was compatible with their own teaching approach. the eight lessons of the teaching intervention were subsequently delivered by all participating teachers over the course of four weeks, with two lessons of 40 minutes duration per week. these lessons replaced regular lessons within the standard math curriculum. on-going assistance from the research team was available throughout the study duration. each student in the class of a participating teacher received a textbook with the problems to be treated, an exercise book and a prompt card that depicted the problem solving steps in a visually attractive way. the first four lessons made use of teacher modelling. in these lessons, three to six word problems were intended to be modelled by the teacher, with 22 problems to be modelled in total. twenty of the teacher-modelled problems were selected for analysis. the selection of these problems was based on the following criteria: the 20 problems were non-routine word problems which required two or more solution steps and could not be solved using a fixed algorithmic method (see appendix i for some examples of the word problems modelled by the teachers)2. it should be noted that for all word problems more than one form visual representation could be made. the two problems that were excluded were not representative of the types of non-routine problems with which students have difficulties as they required only one solution step and could be solved using a fixed algorithmic method3. the last four lessons made use of student modelling followed by independent student practice. instructional support was gradually faded out within and across lessons so that students could ultimately take control of their own problem solving work. 2.3. study measures and analyses the complete teaching intervention (i.e., 8 lessons) was recorded on video for each of the participating teachers. all recordings of the first four lessons (i.e., in which teacher modelling took place) were subsequently viewed by two members of the research team, and the way in which the teacher modelled the problem solving steps (see box 1) for each of the selected problems was recorded, described and coded according to the measurement and analysis scheme presented in the following paragraphs. inter-rater !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 2!note that not all teachers modelled all of these problems due to time constraints.! 3 the two word problems that were excluded from the analysis: 1. there is a skate competition in the park. one lap is 0.8 kilometers. bob skated 30 laps. how many kilometers did bob skate? 2. the cyclists cycle laps of 1.3 kilometers in the city center. how many kilometers did they cycle after 55 laps? ! boonen%et %al % % % ! ! ! 66! ! reliability of the coding was computed as krippendorff’s alpha coefficient for nominal (here, dichotomous) and ordinal data. the obtained coefficients for nominal variables ranged from .86 to 1.00 and for ordinal variables from .91 to .98, indicating high reliability (krippendorff, 2004). 2.3.1. rq1 (attention to visual representations, usage and purpose) for each teacher-modelled problem and for each problem solving step, it was recorded: (1) whether and when the teacher performed the step; (2) step duration in minutes; (3) whether the teacher used a visual representation during the step; (4) for what purpose the visual representation was used; (5) any other factors (e.g., lesson or class management) impacting teachers’ attention to or use of visualization. where steps were combined (i.e., performed simultaneously as opposed to sequentially), the time spent on the combination was allocated in equal proportions to each of the constituent steps. the total time spent on each step by each teacher across all teacher-modelled problems was then calculated, giving an indication of the attention given to each step and consequently the relative attention given to visualization compared to the other steps. next, the data from (1), (3), (4) and (5) above were reviewed and discussed by two members of the research team. this resulted in the identification of patterns specifying when visual representations were used during the problem solving process and to what purpose. 2.3.2. rq2 (kinds of visual representations used, diversity and adaptivity/flexibility) teachers’ behaviours relevant to their use of different types and forms of representations were recorded. for each visual representation used, the type of representation (i.e., pictorial, arithmetical, visualschematic) and form of representation (i.e., bar model, pie chart, number line, proportion table, own construction (i.e., a freely constructed drawing, see table 1 [b] and [d]) was recorded; more than one representation could be recorded per problem, if applicable. for each teacher, the following measures were then derived from these data: • diversity: to examine the extent to which teachers demonstrated a varied use of visual representations simpson’s d was calculated (simpson, 1949). simpson’s d is an estimation of the effective number of species, here, the effective number of representation types used by a teacher. in this case, d ranges from 1 (only one representational form used) to 5 (all representation forms used evenly). note that pictorial representations were not included in calculating d. • adaptivity/flexibility: the extent to which teachers demonstrated adaptive/flexible use of visual representations by offering different representations to solve one problem was calculated as the number and percentage of problems modelled for which the teacher used more than one form of arithmetical and/or visual-schematic representation (e.g., bar model plus proportion table). 2.3.3. rq3 (quality of representation processes and representations) teachers’ behaviours relevant to the quality of their representation processes and representations were recorded. for each arithmetical and visual-schematic representation used, the quality of the representation process was rated in terms of: • transparency, i.e., the extent to which the teacher explicitly explained (or asked students to explain) step-by-step reasoning about what information is essential to represent for solving the given problem and how to represent this (1 = reasoning implicit and unexplained, 2 = reasoning partially explained, 3 = reasoning fully explained); • correctness, i.e., whether or not the reasoning underlying the representation process was correct; • completeness, i.e., whether or not the reasoning underlying the representation process was complete. for each arithmetical and visual-schematic representation used, the quality of the representation itself was rated in terms of: • functionality, i.e., the extent to which the representation was suitable and useful for solving the given problem (1 = not suitable/useful, 2 = suitable/useful but includes solution-irrelevant information, boonen%et %al % % % ! ! ! 67! ! 3 = suitable/useful and includes only solution-relevant information); • correctness, i.e., whether or not all solution-relevant elements were correctly represented; • completeness, i.e., whether or not all solution-relevant elements were included and not missing. on the basis of these data, the average transparency rating, percent correct and complete representation processes, average functionality rating and percent correct and complete representations were then calculated for each teacher across all arithmetical and visual-schematic representations he/she used. note that, although pictorial representations were recorded, measures of diversity, flexibility, functionality and quality of the representation process and representation were not reported for this representation type. the production of pictorial representations was, after all, not the main focus of the teaching intervention that had to implemented by the teachers. 3. results 3.1. rq1. what attention do teachers give to visualization, when do they use visual representations in the word problem solving process and to what purpose? although teachers were trained to perform each of the problem solving steps in the prescribed order for each problem, none of them consistently followed the full sequence of steps. all teachers occasionally omitted steps completely (often the hypothesise and check steps) or partially (particularly paraphrasing the text and imagining the specific situation described therein). furthermore, teachers frequently concatenated steps, often combining visualize with the understand and/or compute step. the consequences of this for teachers’ use of visual representations are examined later. table 3 presents the time allocation for each teacher and each step across all teacher-modelled problems. note that teachers did not spend equivalent amounts of time on modelling (range = 41111 minutes) and did not model the same number of problems (range 4-15). all but two teachers frequently solicited student input during the modelling process and engaged in extensive interaction with students, particularly during the understand, visualize and compute steps. consequently, most time was spent on these steps. it can be seen that four of the eight teachers spent the most time on understanding the problem text (understand), two on both visualising the problem structure (visualize) and calculating answers (compute), and two on calculating answers (compute). thus, although all teachers were aware of the key role of visualization, most spent more time on other parts of the problem solving process. for example, two teachers extensively discussed the general context of each problem (e.g., what kinds of things one can buy in a supermarket) and six teachers placed considerable emphasis on calculating answers. boonen%et %al % % % ! ! ! 68! ! table 3 teachers’ time allocation (minutes) per step over all teacher-modelled problems step teacher teacher 2 teacher 3 teacher 4 teacher 5 teacher 6 teacher 7 teacher 8 total read 5.5 8.1 6.1 6.7 5.8 1.4 2.9 6.3 42.8 understand 20.6 48.0 24.8 8.4 24.5 6.0 17.8 13.1 163.2 visualize 8.6 15.7 15.8 16.4 21.0 11.3 13.9 20.0 122.7 hypothesise 6.7 11.0 7.2 1.9 12.7 5.2 3.2 1.7 49.6 compute 10.6 23.4 27.3 16.5 22.0 14.7 4.7 20.0 139.2 check 1.7 4.8 4.7 3.9 5.9 2.5 5.0 28.5 total 53.7 111.0 85.9 53.8 91.9 41.1 42.5 66.1 546 number of problems modelled 11 15 14 10 9 7 4 9 note. time-overlap between solution steps (e.g., read & understand) might be present but is, for reasons of clarity, not reported. regarding when visual representations are used in the problem solving process and to what purpose, four patterns of usage were identified: (1) understand & visualize; (2) visualize; (3) visualize & compute; (4) compute. the upper part of table 4 shows the amount of time spent (absolute and relative) on visualization following each of these patterns. the lower part of the table shows the corresponding types of visual representations (i.e., pictorial, arithmetical, visual-schematic) used. individual differences between teachers are apparent. the first pattern was observed when the problem text contained a lot of information or when students did not fully comprehend what was being asked. visual representations (specifically, visual-schematic) were then occasionally used to help clarify the situation described in the text; the steps understand and visualize were then combined. half of the teachers showed this pattern of behaviour. the second pattern observed (i.e., visualize) occurred once the problem text was fully comprehended. visual representations (pictorial, arithmetical and/or visual-schematic) were then sometimes used to separately depict the problem structure; these representations were not further used in calculating answers. all but one teacher demonstrated this pattern of usage. with the third pattern, visual representations were used to both structure the problem and help calculate answers; the steps visualize and compute were then combined. all but one teacher used visual representations in this way. for four teachers, this was the main way in which they used visual representations and one teacher used them only in this way. two distinct approaches were identified. with the first approach, a visual-schematic representation was used for structuring the problem and calculating answers. an example is shown in figure 1. all teachers who showed the visualize & compute usage pattern worked in this way on at least one problem. the second approach used by two teachers used arithmetical representations (specifically, proportion tables) for this purpose. boonen%et %al % % % ! ! ! 69! ! finally, with the fourth pattern (i.e., compute), visual representations (specifically arithmetical) were used only to calculate answers. this often followed the situation in which a pictorial or visualschematic representation was used in the visualize step, or followed directly from the text comprehension step (understand) in cases where the visualize step was omitted. six teachers used visual representations in this way. table 4 patterns of visual representation usage over all teacher-modelled problems step rep type teacher 1 teacher 2 teacher 3 teacher 4 teacher 5 teacher 6 teacher 7 teacher 8 time allocation per usage pattern (minutes) understand & visualize 3.4 (8%) 4.8 (14%) 5.9 (13%) 1.7 (6%) visualize 6.1 (32%) 7.3 (18%) 4.4 (10%) 3.1 (9%) 12.6 (27%) 6.6 (22%) 13.9 (75%) visualize & compute 5.1 (26%) 13.4 (33%) 22.7 (53%) 21.7 (61%) 11.0 (24%) 13.0 (44%) 40.0 (100%) compute 8.1 42%) 16.7 (41%) 16.0 (37%) 5.7 (16%) 16.5 (36%) 8.2 (28%) 4.7 (25%) number of visual representations per usage pattern understand & visualize visualschematic 1 1 1 1 visualize pictorial 5 1 1 arithmetical 1 visualschematic 4 8 5 4 8 2 3 visualize & compute arithmetical 2 7 visualschematic 3 3 8 13 2 4 3 compute arithmetical 2 2 5 1 1 2 note. visual representations were never used when reading the problem text (read), hypothesising the required calculations (hypothesise) or checking the answer (check). boonen%et %al % % % ! ! ! 70! ! this week one kilogram of cheese costs €8.90. how much does a piece weighing 300 grams cost? (teacher 4) figure 1. example visual-schematic representation used to calculate answers. 3.2. rq2. what kinds of visual representations do teachers use and how diverse and flexible is this use of visual representations? there were considerable differences between teachers with respect to the type and form of visual representations used. regarding types, (i.e., pictorial, arithmetical, visual-schematic) all but one teacher used mainly visual-schematic representations. though teachers were aware of the importance of these visual representations, pictorial and arithmetical representations were also used, however (see table 5). table 5 teachers’ use of representation types over all teacher-modelled problems rep type teacher 1 teacher 2 teacher 3 teacher 4 teacher 5 teacher 6 teacher 7 teacher 8 number of problems 11 15 14 10 9 7 4 9 number of reps 9 19 20 19 14 9 4 10 pictoriala 0 (0%) 5 (26%) 0 (0%) 0 (0%) 1 (7%) 0 (0%) 1 (25%) 0 (0%) arithmeticala 2 (22%) 2 (11%) 7 (35%) 1 (5%) 2 (14%) 2 (22%) 0 (0%) 7 (70%) visualschematica 7 (78%) 12 (63%) 13 (65%) 18 (95%) 11 (79%) 7 (78%) 3 (75%) 3 (30%) diversity 1.59 2.80 3.17 1.24 4.12 3.52 1.80 1.85 adaptivity/ flexibilityb 1 (9%) 1 (7%) 2 (14%) 1 (10%) 3 (33%) 1 (14%) 1 (25%) 1 (11%) note. apercentages are of the number of visual representations used per teacher; bpercentages are of the number of problems modelled by the teacher. figure 2 shows teachers’ relative use of different representational forms (i.e., bar model, pie chart, number line, proportion table, own construction, pictorial). note that though pictorial representations are considered a type, they are included here to give a complete picture of the visual representations used. the most frequent form of visual-schematic representation was the bar model, which was used by all but one teacher. number lines were used by four teachers, while pie charts were very infrequently used, by four boonen%et %al % % % ! ! ! 71! ! teachers. the only form of arithmetical representation used was proportion tables. this was used by all but one teacher. figure 2. teachers’ use of representational forms over all teacher-modelled problems. the extent to which teachers exhibited a varied use of representational forms (i.e., diversity) differed between teachers, varying from 1.24 to 4.12 on a scale of 1 (only one representational form used) to 5 (all representational forms used equally) (see table 5). excluding pictorial representations, one teacher used five representational forms including own constructions, two teachers used four forms, four teachers used three forms and one teacher used two forms. some teachers verbally expressed and demonstrated a preference for a certain representational form. two teachers showed a strong preference for bar models (although one of these systematically used them incorrectly, as will be discussed later) and two teachers had a strong preference for proportion tables. quotations that reflect teachers’ representation preferences are reported in box 2. box 2: teachers’ representation preferences (translated form dutch) teacher 3: “you all know that i am a big fan of proportion tables”; “i always use a proportion table. you start somewhere and end up with a proportion table” teacher 4: “i can draw another bar model here”; “i feel another bar model coming on!” teacher 8: “a proportion table is great to use for nearly all sorts of problems” teachers demonstrated a low to medium degree of adaptivity/flexibility in representation use. all offered multiple representational forms (arithmetical and visual-schematic) on the same problem at least once (see table 5). five teachers used different forms of visual-schematic representation in this way. four combined a bar model with their own schematic drawing of the problem structure; two offered different visual-schematic mathematical models (number line, bar model and/or pie chart). five teachers used both visual-schematic and arithmetical representations on the same problem at least once. remarkably, three teachers modelled at least one problem without using any visual representation at all. teachers rarely compared the use of different representational forms or reflected critically on what different kinds of representation contribute to word problem solving. this was particularly the case for teachers who had an expressed preference for using a specific form, who showed a tendency to use their preferred form irrespective of problem characteristics. 0" 5" 10" 15" 20" 25" 1" 2" 3" 4" 5" 6" 7" 8" nu m be r"o f"r ep re se nt a7 on s" teacher" boonen%et %al % % % ! ! ! 72! ! 3.3. rq3. quality of representation processes and visual representations teachers were expected to model the representation process transparently, correctly and completely, whereby they should emphasize reasoning processes and understanding the actions undertaken. thus, teachers were expected to explicitly explain which problem elements should be modelled and how, which representational forms are suitable to do this and why, and which actions they take in constructing the visual representations step-by-step. all eight teachers showed these kinds of behaviours to a considerable degree, with average transparency ratings of 2.11-3.00 on a three-point scale (see table 6). the extent to which teachers provided this reasoning themselves or stimulated students to provide it varied between teachers, however. two teachers also encouraged students to suggest which visual representations to use. table 6 quality of representation processes and visual representations teacher 1 teacher 2 teacher 3 teacher 4 teacher 5 teacher 6 teacher 7 teacher 8 number of reps 9 19 20 19 14 9 4 10 representation processes (reasoning) transparencya 2.11 2.43 2.70 2.84 2.62 2.89 3.00 2.90 correctnessb 33% 100% 100% 95% 100% 100% 100% 100% completenessb 100% 86% 90% 95% 85% 100% 100% 100% visual representations functionalitya 1.56 2.64 2.85 2.63 2.69 3.00 3.00 3.00 correctnessb 33% 100% 100% 89% 92% 89% 100% 100% completenessb 100% 79% 85% 89% 77% 100% 100% 100% accuracy of visual representation types arithmeticalc 0% 0% 100% 100% 50% 100% 100% visual-schematicc 29% 92% 77% 83% 73% 86% 100% 100% notes. aaveraged over all arithmetical and visual-schematic representations; bpercentage of all arithmetical and visual-schematic representations; cpercentage correct and complete visual representations of this type. correctness of reasoning in the representation process was high (95% and above) for seven of the eight teachers. remarkably, one teacher (teacher 1) reasoned correctly on only 33% of visual representations. the main source of error appeared to be an incorrect understanding of what a bar model is: she drew boxes that she called bar models within which she wrote the sum in symbolic notation. this teacher along with three others also used a bar model form as a proportion table, with the numbers written in the bars bearing no relation to the spatial-numerical relations. another teacher (teacher 7) expressed the belief that it is not possible to make a visualization for every type of problem, though his reasoning was correct for the representations that he did make. completeness of reasoning was also high (95% and above) for five boonen%et %al % % % ! ! ! 73! ! teachers, while the remaining three teachers demonstrated complete reasoning on 85-90% of the visual representations used. there were differences between teachers concerning the quality of the visual representations used. all teachers produced at least one accurate (i.e., correct and complete) visual-schematic representation, though accuracy varied from 29% to 100% for this representation type (see last row of table 6). only two teachers were fully accurate on all visual-schematic representations, while the teacher who demonstrated low correctness of reasoning produced accurate visual-schematic representations only 29% of the time. four teachers constructed visual-schematic representations that were incorrect (i.e., depicting erroneous problem elements or relations) and four teachers constructed visual-schematic representation that were incomplete (i.e., missing solution-relevant elements and/or relations). examples of accurate and inaccurate visualschematic representations are given in figure 3. the arithmetical representations of four teachers were all accurate (i.e., correct and complete), but three teachers produced incorrect or incomplete arithmetical representations (one teacher never used these visual representations, see table 5). yesterday, fenna had €337.65 in her bank account. her new statement shows that her grandmother deposited €45 for her school report and she earned €11.75 for babysitting. how much money does she have in her account now? (teacher 5) (teacher 1) (a) accurate (correct and complete) (b) inaccurate (incorrect) figure 3. examples of (a) accurate and (b) inaccurate visual-schematic representations. the extent to which teachers invested in understanding what was being asked in the problem text particularly when the text contained much information directly impacted the functionality (i.e., suitability and usefulness) of the visual representations used. with one exception, teachers’ visual representations generally had good functionality, with average ratings of 2.63-3.00 on a three-point scale (see table 6). however, the visual representations of the teacher who showed low correctness of reasoning had lower functionality (average rating 1.56), which could be expected. functionality was also lower when teachers did not correctly identify the solution-relevant elements; in these cases, they tended to produce visual representations that contained excess, irrelevant information. figure 4(a) shows an example of such a visualschematic representation, with unnecessary calculation of intermediate arrival and departure times. furthermore, teachers used bar models more frequently than other visual-schematic representations, even on problems involving percentages that were suitable for a pie chart, or addition and subtraction problems that were suitable for using a number line. figure 4(b) shows an example of a bar model used (inaccurately) when a number line would have been more suitable. boonen%et %al % % % ! ! ! 74! ! antoine and bertrand travel with the train from utrecht to paris via den bosch, roosendaal, and brussels. they depart at 13.38 from utrecht central station. after 28 minutes they arrive in den bosch. they have to wait 13 minutes before the train to roosendaal leaves. the journey to roosendaal takes 52 minutes. in roosendaal they have 23 minutes to change to the train to brussels. the journey to brussels takes 1 hour and 10 minutes. after 55 minutes, the train departs from brussels to paris: a journey of 1 hour and 22 minutes. at what time do antoine and bertrand arrive in paris? (teacher 8) (a) (teacher 4) (b) figure 4. examples of (a) visual-schematic representation containing excess information and (b) bar model used to solve addition problem. 4. discussion the goal of this study was to examine teachers’ use of visual representations when implementing a teaching intervention for supporting non-routine word problem solving. the teaching intervention focused on the construction of accurate visual-schematic representations, embedded within a sequence of six problem solving steps. it differs from previous research into supporting word problem solving (e.g., jitendra, 2002; montague, 2003; montague et al., 2000) in several respects relating to both the educational setting and the way in which it is implemented. first, the teaching intervention was not implemented with low-performing students with special educational needs, but in regular, mainstream classrooms. second, the intervention was not designed for individual or small group instruction, but for use in whole-class teaching. third, the intervention was not carried out by researchers but by mainstream teachers. furthermore, while the intervention was based on existing stepwise strategy instruction programs to support non-routine word problem solving, it incorporated an important innovation. where existing programs define visual representation in heuristic terms, the present teaching intervention specifically defines the criteria that should be satisfied: a visual representation should clarify the problem structure by making the numerical, linguistic and spatial relations between solution-relevant elements visible. the study boonen%et %al % % % ! ! ! 75! ! therefore makes an unique contribution to research in the important and problematic area of word problem solving in regular classrooms. to address the study objective, three research questions were posed. we first examined teachers’ attention to visualization when implementing the teaching intervention and determined when they use visual representations in the word problem solving process and to what purpose. answering this question provides insights into teachers’ use of visual representations. we also investigated what kinds of visual representation teachers use, the extent to which teachers show a diverse and adaptive/flexible use of representations, and the quality of the representation process and of the visual representations used. this information could inform research, teacher training and professionalization with respect to improving teachers’ expertise with visual representations. 4.1. teachers’ attention to visualization and its role in word problem solving teachers were trained in the use of the teaching intervention and were aware of its focus on visualization (i.e., the construction of a visual representation of problem structure). it could therefore be expected that teachers would pay the most attention to the visualize step, in comparison to the other problem solving steps. this was not the case, however: most teachers spent most time on other parts of the problem solving process, namely understanding problem text and computing answers. extensive time spent on the understand step could be explained by the fact that teachers regularly did not apply the intended strategies for clarifying what was asked (e.g., paraphrasing the text and imagining the described situation) but instead spent a lot of time discussing irrelevant contextual information. this indicates that teachers are unfamiliar with and therefore need to master strategies for supporting text comprehension if they are to effectively support word problem solving of students. teachers’ focus on correctly performing the required arithmetical computations is not surprising, given that instructional methods commonly used in mainstream classrooms and teacher training award much more attention to calculating correct answers than to understanding the problem text. consequently, teachers are generally used to spending most time on the solution phase of the word problem solving process, and this practice is likely to have been perpetuated in the current setting. teachers routinely combined the visualize step with the understand or compute step. this stands in sharp contrast to existing research-based programs for supporting word problem solving of low-performing students, which assume that word problem solving is a sequential (i.e., step-by-step) process (krawec, 2010; montague et al., 2000). the results of the present study show that, in authentic classroom settings, word problem solving need not always occur sequentially steps can be combined such that problem comprehension and solution interact and emerge together. theories of word problem solving therefore need to recognize the potentially iterative nature of the process, particularly for non-routine problems. 4.2. teachers’ representational use and diversity and adaptivity/flexibility therein most teachers made ample use of visual-schematic representations (i.e., visual representations that represent the problem structure, the solution-relevant elements and the relations between them), though some seemed unclear about what these representations comprise and what function they serve within the word problem solving context. this is in line with research findings from other countries (e.g., dreher & kuntze, 2015; orrill et al., 2008; turner, 2008). moreover, some teachers also made frequent use of arithmetical representations specifically, proportion tables. in contrast to visual-schematic representations, arithmetical representations support only calculation processes rather than problem comprehension. thus, when the problem text is not well understood, using only this kind of representation bears the risk that it does boonen%et %al % % % ! ! ! 76! ! not contain the solution-relevant elements of the problem to be solved. even after training, teachers appeared not always to be aware of this difference in the use of these types of representations. teachers appeared to have strong preferences for using one particular form of visual representation, namely a bar model or proportion table. given these preferences, it is not surprising that teachers showed limited diversity and adaptivity/flexibility in representation use. this could be explained by the fact that these representational forms are frequently offered in math textbooks in elementary schools and teacher education. teachers consequently may feel more comfortable using them, as they encounter them more often and have more knowledge about them. this may also explain the finding that teachers rarely considered or compared the suitability of the representations used. nonetheless, an imbalanced and/or inflexible use of representations can be problematic when teachers are unable to respond appropriately to students’ needs (jitendra et al., 2007) (e.g., students may find a number line more helpful than a bar model for certain sorts of problems) or to problem characteristics (e.g., a pie chart is more suitable than a bar model for solving a problem involving percentages). thus, the diverse and adaptive/flexible use of visual representations is clearly an issue to be addressed in teacher training and professionalization. 4.3. the quality of the representation process and the visual representations produced the finding that teachers demonstrated medium to high transparency in the way in which they provided explicit, step-by-step reasoning about what information should be represented and how to represent it is compatible with the nature of contemporary math education, which emphasizes reasoning processes and understanding (barnes, 2005; van den heuvel-panhuizen, 2003; webb, van der kooij, & geist, 2011). several teachers extensively involved students in this process; the explicit interaction between teacher and students is also one of the underlying principles of contemporary math education (van den heuvelpanhuizen, 2003; webb et al., 2011). it is worrying, though, that while the majority of teachers generally demonstrated high correctness of reasoning, one teacher appeared to hold a fundamental misconception of what a bar model is. furthermore, the reasoning of half of the teachers was not always complete: visual representations were introduced but not further or fully explicated. such incomplete reasoning is risky if students do not know how to use the representation in question: misconceptions can then arise that can be difficult to correct (hill et al., 2008). a correct and complete representation process is therefore essential but is not always exhibited in teachers’ natural behaviours. with respect to the quality of the representations themselves, only two teachers were consistently able to construct representations that both correctly and completely contained all solution-relevant elements, in spite of the training received. it is highly likely that other mainstream teachers (who have not been explicitly trained in this area) would also have difficulty in producing correct and complete visual representations. furthermore, some teachers were unable to come up with a suitable visual representation for some problems and seemed to think that they were making a visual representation when in fact they were sometimes merely structuring the words in the text. clearly, teachers’ expertise in this area needs to be improved. also, some teachers produced representations containing excess, irrelevant information, largely as a consequence of insufficient understanding of what was being asked. such representations make calculating the answer more difficult and error-sensitive than necessary, and make it unclear what information is solution-relevant and what is not. this suggests that teachers, as well as students, need to develop effective strategies for understanding what is being asked in non-routine word problems. boonen%et %al % % % ! ! ! 77! ! 4.4. implications for teacher professionalization in using visual representations to support word problem solving based on these findings, it can be concluded that the use of visual representations to support word problem solving should be given much more attention in teacher education and teacher professionalization programs. while the quality of the representation processes and the representations produced by most of the teachers in this study was reasonable, some misconceptions and inappropriate use of certain representational forms were observed. it could be argued that anything less than teachers’ full mastery in this area is undesirable (cf. ball et al., 2008). furthermore, the limited diversity and adaptivity/flexibility are matters of concern. while routine, algorithmic problems can be solved by applying pre-existing templates, non-routine problems require a problem-specific approach. in these cases, a limited and inflexible use of representations can result in an ineffective and inefficient problem solving process. teachers therefore need to possess a broad repertoire of visual representations and understand their conditions of use, so that they are able to offer representations that both match problem characteristics and support students’ needs, as posited in the concept of mkt (ball et al., 2008). it is important, therefore, to develop teachers’ knowledge of the characteristics and purpose of different types and forms of visual representation, as well as understanding when and how to use them to support word problem solving. this notion should be seen in conjunction with the work of disessa (2002) on the existence of metarepresentational competence of students in mathematics and science. metarepresentational competence (mrc) refers to the full complex of abilities dealing with representational issues. it concerns the ability to design new representations, including both creating representations and judging the adequacy for particular purposes. it also includes understanding how representations work, how to work out representations for different purposes, and what the purposes of representations are (disessa, 2002; verschaffel, reybrouck, jans & van dooren, 2010). in addition to the studies of disessa (2002) and verschaffel et al. (2010) which are focused on students’ level of mrc, the findings of this study show that it is also important to look at teachers’ development of a true representational literacy (see disessa, 2002, p. 105). metarepresentational competence should, among others, have a prominent place in teacher education and teacher professionalization programs (e.g., dreher &kuntze, 2015; orill et al., 2008; verschaffel et al., 2010). in addition, attention needs to be paid to mastering strategies for supporting text comprehension. teachers need to learn how to identify the solution-relevant information in the word problem and, on the basis of that information, how to derive the specific questions that have to be answered. this should prevent teachers from spending too much time on irrelevant details which do not facilitate and may even hinder problem comprehension and solution. 4.5. directions for future research the participants in this study were recruited through purposive sampling of teachers who were motivated to participate in and considered themselves competent to contribute to research in this area. this ensured that results were obtained under favourable conditions in which teacher behaviour is not negatively influenced by motivational factors that are known to undermine the way in which teachers implement educational innovations in regular classrooms (e.g., evers et al., 2002; ghaith & yaghi, 1997; hermans et al., 2008; rogers, 2003). consequently, the findings cannot be directly generalized to the many mainstream teachers who have a negative attitude towards mathematics and are not confident about teaching mathematics (e.g., bursal & paznokas, 2006; isiksal, curran, koc, & askun, 2009; swars, daane, & giesen, 2006). nonetheless, we expect that since even motivated and confident teachers experience difficulties in using visual representations (such as limited diversity, adaptivity/flexibility and representational quality), difficulties experienced by other mainstream teachers might be more prominent. thus, it is important to address the competences and needs of these teachers in future research. teachers were encouraged to use their own explanations and elaborations, rather than a fully scripted boonen%et %al % % % ! ! ! 78! ! lesson protocol. in this way, teachers could implement the teaching intervention in a way that is compatible with their own teaching approach and beliefs about teaching. this is important for the successful implementation and feeling of ownership of educational innovations (ketelaar, beijaard, boshuizen, & den brok, 2012). nonetheless, it makes instruction vulnerable to potential shortcomings in teachers’ skills. thus, even when teachers believe they master the skills necessary to implement an instruction correctly, they should be provided with explicit training in its key ingredients and be given the opportunity to adopt and consolidate new skills before it can be implemented in the classroom (bitan-friedlander, dreyfus, & milgrom, 2004). to this end, training should be lengthy enough to provide teachers with enough time to internalize change, that is, accept the innovation, acquire the necessary skills and be prepared to implement it (bitan-friedlander et al., 2004). possibly, the present training was not of sufficient duration so that the desired level of competence was not attained. future research should investigate the duration and intensity of training required to achieve this level of competence. finally, some teachers believed themselves to be competent while in fact they were not, with one teacher even holding a fundamental misconception. as realistic beliefs about one’s personal competence can positively influence individuals’ willingness to invest in training and education (beets, flay, vuchinich, acock, li, & allred, 2008; han & weiss 2005), it would be of interest to investigate how teachers can be helped to develop realistic beliefs about their pedagogical and didactical proficiency in the use of visual representations. keypoints visual-schematic representations can be used to support word problem solving however, teachers were unclear what function visual-schematic representations serve some teachers were unclear about what visual-schematic representations comprise. teachers had preferences for using particular forms of visual representation. diverse and adaptive/flexible visual representation use is an important issue for teacher training. references acevedo nistal, a., van dooren, w., clarebout, g., elen, j., & verschaffel, l. (2009). conceptualising, investigating and stimulating representational flexibility in mathematical problem solving and learning: a critical review. zdm the international journal on mathematics education, 41, 627-636. doi: 10.1007/s11858-009-0189-1 ball, d. l., thames, m. h., & phelps, g. (2008). content knowledge for teaching: what makes it special? journal of teacher education, 59, 389-407. doi: 10.1177/0022487108324554 barnes, h. (2005). the theory of realistic mathematics education as a theoretical framework for teaching low attainers in mathematics. pythagoras, 61, 42-57. doi: 10.4102/pythagoras.v0i61.120 beets, m. w., flay, b. r., vuchinich, s., acock, a. c., li, k.-k., & allred, c. (2008). school climate and teachers’ beliefs and attitudes associated with implementation of the positive action program: a diffusion of innovations model. prevention science, 9, 264-275. doi: 10.1007/s11121-008-0100-2 berends, i. e., & van lieshout, e. (2009). the effect of illustrations in arithmetic problem-solving: effects of increasing cognitive load. learning and instruction, 19, 345–353. doi: http://dx.doi.org/10.1016/j.learninstruc.2008.06.012 bitan-friedlander, n., dreyfus, a., & milgrom, z. (2004). types of “teachers in training”: the reactions of primary school science teachers when confronted with the task of implementing an innovation. teaching and teacher education, 20, 607-619. doi: 10.1016/j.tate.2004.06.007 boonen%et %al % % % ! ! ! 79! ! boonen, a. j. h., van der schoot, m., van wesel, f., de vries, m. h., & jolles j. (2013). what underlies successful word problem solving? a path analysis in sixth grade students. contemporary educational psychology, 38, 271-279. doi: http://dx.doi.org/10.1016/j.cedpsych.2013.05.001 boonen, a. j. h., van wesel, f., jolles, j., & van der schoot, m. (2014). the role of visual representation type, spatial ability, and reading comprehension in word problem solving: an item-level analysis in elementary school children. international journal of educational research, 68, 15-26. doi: http://dx.doi.org/10.1016/j.ijer.2014.08.001 boonen, a. j. h., & jolles, j. (2015). comprehension & visualization: teaching students to solve word problems. research & reviews: journal of educational studies, x, 1-4. bursal, m., & paznokas, l. (2006). mathematics anxiety and preservice elementary teachers’ confidence to teach mathematics and science. school science & mathematics, 106, 173-180. doi: 10.1111/j.19498594.2006.tb18073.x carpenter, t. p., corbitt, m. k., kepner, h. s., lindquist, m. m., & reys, r. e. (1981). national assessment. in e. fennema (ed.), mathematics education research; implications for the 80's (pp. 2238). reston, va: national council of teachers of mathematics. csíkos, c., szitányi, j., & kelemen, r. (2012). the effects of using drawings in developing young children’s mathematical word problem solving: a design experiment with third-grade hungarian students. educational studies in mathematics, 81, 47-65. doi:10.1007/s10649-011-9360-z cummins, d. d., kintsch, w., reusser, k., & weimer, r. (1988). the role of understanding in solving word problems. cognitive psychology, 20, 405-438. doi: http://dx.doi.org/10.1016/0010-0285(88)90011-4 depaepe, f., de corte, e., & verschaffel, l. (2010). teachers' approaches towards word problem solving: elaborating or restricting the problem context. teaching and teacher education, 26, 152-160. doi: 10.1016/j.tate.2009.03.016 department for education. (2013). national curriculum in england; mathematics programmes of study: key stages 1 and 2. retrieved from www.gov.uk/government/publications/national-curriculum-in-englandmathematics-programmes-of-study. disessa, a. (2002). students’ criteria for representational adequacy. in k. gravemeijer, r. lehrer, b. van oers, & l. verschaffel (eds). symbolizing, modeling and tool use in mathematics education (pp. 105129). dordrecht, bosten, londen: kluwer academic publishers. dreher, a., & kuntze s. (2015). teachers' professional knowledge and noticing: the case of multiple representations in the mathematics classroom. educational studies in mathematics, 88, 89-114. doi: 10.1007/s10649-014-9577-8 elia, i., van den heuvel-panhuizen, m., & kolovou, a. (2009). exploring strategy use and strategy flexibility in non-routine problem solving by primary school high achievers in mathematics. zdm, 41, 605-618. doi: 10.1007/s11858-009-0184-6 evers, w. j. g., brouwers, a., & tomic, w. (2002). burnout and self-efficacy: a study on teachers’ beliefs when implementing an innovative educational system in the netherlands. british journal of educational psychology, 72, 227-243. doi: 0.1348/000709902158865 ghaith, g., & yaghi, h. (1997). relationships among experience, teacher efficacy, and attitudes toward the implementation of instructional innovation. teaching and teacher education, 13, 451-458. doi: 10.1016/s0742-051x(96)00045-5 han, s. s., & weiss, b. (2005). sustainability of teacher implementation of school-based mental health programs. journal of abnormal child psychology, 33, 665–679. doi: 10.1007/s10802-005-7646-2. hegarty, m., & kozhevnikov, m. (1999). types of visual–spatial representations and mathematical problem solving. journal of educational psychology, 91, 684–689. doi: http://dx.doi.org/10.1037/00220663.91.4.68 heinze, a., star, j. r., & verschaffel, l. (2009). flexible and adaptive use of strategies and representations in mathematics education. zdm, 41, 535-540. doi:10.1007/s11858-009-0214-4 hermans, r., tondeur, j., van braak, j., & valcke, m. (2008). the impact of primary school teachers’ educational beliefs on the classroom use of computers. computers & education, 51, 1499-1509. doi: 10.1016/j.compedu.2008.02.001 boonen%et %al % % % ! ! ! 80! ! hill, h. c., blunk, m. l., charalambous, c. y., lewis, j. m., phelps, g. c., sleep, l., & ball, d. l. (2008). mathematical knowledge for teaching and the mathematical quality of instruction: an exploratory study. cognition and instruction, 26, 430-511. doi; 10.1080/07370000802177235 isiksal, m., curran, j. m., koc, y., & askun, c. s. (2009). mathematics anxiety and mathematical selfconcept: considerations in preparing elementary-school teachers. social behavior and personality, 37, 631-644. doi: 10.2224/sbp.2009.37.5.631 jiminez, l., & verschaffel, l. (2014). development of children’s solutions of non-standard arithmetic word problem solving. revista de psicodidáctica, 2014, 19, 93-123. doi: 10.1387/revpsicodidact.7865 jitendra, a.k. (2002). teaching students math problem-solving through graphic representations. teaching exceptional children, 34, 34-38. doi: 10.1177/004005990203400405 jitendra, a.k., dipipi, c. m., & perron-jones, n. (2002). an exploratory study of schema based word problem solving instruction for middle school students with learning disabilities: an emphasis on conceptual and procedural understanding. the journal of special education, 36, 23–38. doi: http://dx.doi.org/10.1177/00224669020360010301 jitendra, a. k., griffin, c. c., haria, p., leh, j., adams, a., & kaduvettoor, a. (2007). a comparison of single and multiple strategy instruction on third-grade students' mathematical problem solving. journal of educational psychology, 99, 115-127. doi: 10.1037/0022-0663.99.1.115 jitendra, a. k., petersen-brown, s., lein, a. e., zaslofsky, a. f., kunkel, a. k., jung, p.-g., & egan, a. m. (2013). teaching mathematical word problem solving: the quality of evidence for strategy instruction priming the problem structure. journal of learning disabilities, xx, 1-22. doi: 10.1177/0022219413487408 jitendra, a. k., & star, j. r. (2012). an exploratory study contrasting highand low achieving students' percent word problem solving. learning and individual differences, 22, 151-158. doi: 10.1016/j.lindif.2011.11.003 jitendra, a. k., star, j. r., starosta, k., leh, j. m., sood, s., caskie, g., … mack, t. r. (2009). improving seventh grade students’ learning of ratio and proportion: the role of schema-based instruction. contemporary educational psychology, 34, 250-264. doi: 10.1016/j.cedpsych.2009.06.001 ketelaar, e., beijaard, d., boshuizen, h., & den brok, p. j. (2012). teachers’ positioning towards an educational innovation in the light of ownership, sense-making and agency. teaching and teacher education, 28, 273-282. doi: 10.1016/j.tate.2011.10.004 krawec, j. l. (2010). problem representation and mathematical problem solving of students with varying abilities (doctoral dissertation, university of miami). miami. krawec, j. l. (2012). problem representation and mathematical problem solving of students of varying math ability. journal of learning disabilities, xx, 1-13. doi: 10.1177/0022219412436976 krippendorff, k. (2004). reliability in content analysis: some common misconceptions and recommendations. human communication research, 30, 411-433. doi: http://dx.doi.org/10.1111/j.1468-2958.2004.tb00738.x lewis, a. b., & mayer, r. e. (1987). students’ miscomprehension of relational statements in arithmetic word problems. journal of educational psychology, 79, 363-371. doi: 10.1037/0022-0663.79.4.363 montague, m. (2003). solve it! a practical approach to teaching mathematical problem solving skills. va: exceptional innovations, inc. montague, m., warger, c., & morgan, t. h. (2000). solve it! strategy instruction to improve mathematical problem solving. learning disabilities research & practice, 15, 110-116. doi: 10.1207/sldrp1502_7 mullis, i.v.s. & martin, m.o. (eds.). (2013). chestnut hill, ma: timss & pirls international study center, boston college. national council of teachers of mathematics (nctm). (2000). principles and standards for school mathematics. reston, va: national council of teachers of mathematics. noteboom, a. (2009). fundamentele doelen rekenen-wiskunde [key objectives for arithmetic and math]. enschede, the netherlands: slo, national expertise center for curriculum development. oecd (2013). pisa 2012 assessment and analytical framework: mathematics, reading, science, boonen%et %al % % % ! ! ! 81! ! problem solving and financial literacy. pisa, oecd publishing. orrill, c. h., sexton, s., lee, s.-j., & gerde, c. (2008). mathematics teachers’ abilities to use and make sense of drawn representations. in the international conference of the learning sciences 2008: proceedings of icls 2008. mahwah, nj: international society of the learning sciences. prenger, j. (2005). taal telt! een onderzoek naar de rol van taalvaardigheid en tekstbegrip in het realistische rekenonderwijs. [language counts! a study into the role of linguistic skill and text comprehension in realistic mathematics education]. doctoral dissertation, university of groningen, the netherlands. pantziara, m., gagatsis, a. & elia, i. (2009). using diagrams as tools for the solution of non-routine mathematical problems. educational studies in mathematics, 72, 39-60. doi: 10.1007/s10649-0099181-5 rogers, e. m. (2003). diffusion of innovations (5th ed.). new york: free press. schoppek, w., & tulis, m. (2010). enhancing arithmetic and word-problem solving skills efficiently by individualized computer-assisted practice. the journal of educational research, 103, 239-252. doi: 10.1080/00220670903382962 sharma, u., loreman, t., & forlin, c. (2012). measuring teacher efficacy to implement inclusive practices. journal of research in special educational needs, 12, 12-21. doi: 10.1111/j.1471-3802.2011.01200.x shulman, l. s. (1987). knowledge and teaching: foundations of the new reform. harvard educational review, 57, 1-22. doi: http://dx.doi.org/10.17763/haer.57.1.j463w79r56455411 simpson, e. h. (1949). measurement of diversity. nature, 163, 688. doi: http://dx.doi.org/10.1038/163688a0 swanson, h. l., lussier, c. m., & orosco, m. j. (2013). cognitive strategies, working memory, and growth in word problem solving in children with math difficulties. journal of learning disabilities, xx, 1-20. doi: 10.1177/0022219413498771 swars, s. l., daane, c. j., & giessen, j. (2006). mathematics anxiety and mathematics teacher efficacy: what is the relationship in elementary preservice teachers? school science and mathematics, 106, 306-315. doi: 10.1111/j.1949-8594.2006.tb17921.x turner, f. (2008). beginning elementary teachers’ use of representations in mathematics teaching, research in mathematics education, 10, 209-210. doi: 10.1080/14794800802233795 van den heuvel-panhuizen, m. (2003). the didactical use of models in realistic mathematics education: an example from a longitudinal trajectory on percentage. educational studies in mathematics, 54, 9-35. doi: 10.1023/b:educ.0000005212.03219.dc van der veen, i., smeets, e., & derriks, m. (2010). children with special educational needs in the netherlands: number, characteristics and school career. educational research, 52, 15-43. doi: 10.1080/00131881003588147 van dijk, i. m. a. w., van oers, h. j. m., & terwel, j. (2003). providing or designing? constructing models in primary math education. learning and instruction, 13, 53-72. doi: http://dx.doi.org/10.1016/s0959-4752(01)00037-8 van dijk, i. m. a. w., van oers, b., terwel, j., & van den eeden, p. (2003a). strategic learning in primary mathematics education: effects of an experimental program in modeling. educational research and evaluation, 9, 161-187. doi: http://dx.doi.org/10.1076/edre.9.2.161.14213 van garderen, d. (2006). spatial visualization, visual imagery, and mathematical problem solving of students with varying abilities. journal of learning disabilities, 39, 496–506. doi: http://dx.doi.org/10.1177/00222194060390060201 van garderen, d., & montague, m. (2003). visual–spatial representation, mathematical problem solving, and students of varying abilities. learning disabilities research & practice, 18, 246–254. doi: http://dx.doi.org/10.1111/1540-5826.00079 verschaffel, l., de corte, e., lasure, s., van vaerenbergh, g., bogaerts, h., & ratinckx, e. (1999). learning to solve mathematical application problems: a design experiment with fifth graders. mathematical thinking and learning, 1, 195-229. doi: �http://dx.doi.org/10.1207/s15327833mtl0103_2 boonen%et %al % % % ! ! ! 82! ! verschaffel, l., greer, b., & de corte, e. (2000). making sense of word problems. swets and zeitlinger: lisse. verschaffel, l., reybrouck, m., jans, c., & van dooren, w. (2010). children's criteria for representational adequacy in the perception of simple sonic stimuli.cognition and instruction, 28, 475-502. webb, d. c., van der kooij, h., & geist, m. r. (2011). design research in the netherlands: introducing logarithms using realistic mathematics education. journal of mathematics education at teachers college, 2, 47-52. willis, g. b., & fuson, k. c. (1988). teaching children to use schematic drawings to solve addition and subtraction word problems. journal of educational psychology, 80, 192-201. doi: http://dx.doi.org/10.1037/0022-0663.80.2.192! appendix i: examples of word problems modelled by the teachers a) this week one kilogram of cheese costs €8.90. how much does a piece weighing 300 grams cost? b) yesterday, fenna had €337.65 in her bank account. her new statement shows that her grandmother deposited €45 for her school report and she earned €11.75 for babysitting. how much money does she have in her account now? c) antoine and bertrand travel with the train from utrecht to paris via den bosch, roosendaal, and brussels. they depart at 13.38 from utrecht central station. after 28 minutes they arrive in den bosch. they have to wait 13 minutes before the train to roosendaal leaves. the journey to roosendaal takes 52 minutes. in roosendaal they have 23 minutes to change to the train to brussels. the journey to brussels takes 1 hour and 10 minutes. after 55 minutes, the train departs from brussels to paris: a journey of 1 hour and 22 minutes. at what time do antoine and bertrand arrive in paris? d) a truck has a loading box of 4.5 metres long, 2.5 metres wide and 0.5 metres high. the truck drives back and forth to a construction site with sand and gravel. because the truck does not want to lose too much sand and gravel, the truck is filled to the brim and not higher. there lies 144 m3 sand at the construction site. how often should the truck drive back and forth with the sand? e) the students from the upper grades of elementary school de zonnewijzer have a sports day. all children of grade 3, 4, 5 and 6 are participating. grade 3 has 20 students, grade 4 has 25 students, grade 5a has 21 students, grade 5b has 23 students, grade 6a has 19 students and grade 6b has 26 pupils. at the sports day in each grade a teacher is present. how many litre packs of milk must be purchased if every child and every teacher drinks one glass of 0.2 litres of milk? microsoft word iordanou_publication.docx           frontline  learning  research  vol.4  no.  5  (2016)  106  -­‐  119   issn  2295-­‐3159       corresponding author: kalypso iordanou, university of central lancashire, 12 -14 university avenue, pyla, 7080 larnaka, cyprus. email address: kiordanou@uclan.ac.uk doi: http://dx.doi.org/10.14786/flr.v4i5.252   from theory of mind to epistemic cognition. a lifespan perspective kalypso iordanou university of central lancashire article received 25 april / revised 17 august / accepted 5 september / available online 19 january   abstract although a sizeable body of research now exists in epistemic cognition, it tends to stand apart from other aspects of cognition and cognitive development. here it is proposed to situate epistemic cognition in a context of its roots and development as a dimension of cognitive development more generally. the present paper draws a strong continuous link between the earliest understanding of other minds, examined under the theory of mind, and the tasks that confront adults throughout the lifespan – that of interpreting evidence and coordinating it with what they already take to be true. the primary focus is the how question of knowledge change. to gain insight into this question, it is proposed to focus on epistemic activity in action. it is suggested here that the standards for knowledge formation and revision, which are closely connected with epistemic understanding of theory-evidence coordination, change developmentally. another major change proposed is that the process increasingly comes under conscious control. keywords: epistemic cognition, cognitive development, argumentation, theory of mind iordanou           | f l r     107   1. problem how do people know? how do people form beliefs? how do people revise beliefs? are there developmental differences in this regard? these questions have long been a concern of psychologists, philosophers and educators. their answers can be found in writing on the topic of epistemic cognition (bendixen & rule, 2004; greene, muis, & pieschl, 2010; greene, sandoval, & bråten, 2016; muis, bendixen, & haerle, 2006; perry, 1970). following chinn, buckland, and samarapungaran (2011) and greene et al. (2016), epistemic cognition is defined here as “cognition of or relating to knowledge” (greene et al., 2016, p. 3). although a sizeable body of research now exists under this heading, it tends to stand apart, with few connections to other aspects of cognition and cognitive development. here i propose a broader view, situating epistemic cognition in a context of its roots and development as a dimension of cognition and cognitive development more generally. my primary focus is on the mechanism question, the ‘how’ question of knowledge change. i propose that this change only comes about through application of one's epistemic cognition in practice, which consists of forming and revising claims. this is a continuous process through life and there is reason to think that the nature of the process changes, with mechanisms and standards for knowledge formation and revision changing developmentally. a major change, i propose, is that the process increasingly comes under conscious control. to gain insight into the how question of knowledge change, i propose focusing on epistemic activity in action – the application of epistemic cognition. available models of the developmental progression of epistemic cognition offer a very general stage-like description of this progression with little attention to mechanism. they began with perry’s (1970) study of harvard undergraduate students, hardly a broad sample of the population. for many years, research following perry’s work continued the study of adolescent and adult samples, with no reference to the developmental origins of their thought. an assumption that epistemic beliefs emerge abruptly in adolescence and remain unchanged thereafter – a non-developmental account – seems unwarranted, considering the cognitive development that occurs along so many other dimensions during the years between early childhood and adolescence. researchers are thus left with few answers to the key questions of how epistemic conceptions emerge and how they continue to develop. although researchers have gone on to address many other important questions, such as how epistemic cognition is related to academic performance (muis, kendeou, franko, 2011; stømsø, bråten, & britt, 2011), little work has been done to further our understanding of its development. largely standing today is the 2004 conclusion drawn by bendixen and rule, “currently there is neither a unified model of epistemological understanding to guide research, nor a single model that clearly articulates the relationship between personal epistemology and how epistemological beliefs change and develop” (p. 69). a fuller developmental account is essential not only for expanding understanding at a theoretical level but also for its educational implications, by identifying means to support development of sophisticated epistemic cognition. it is not the case that little is known about cognitive development in the first decade of life, and more specifically development potentially relevant to epistemic cognition. in particular, there is now an extensive literature, particularly in the field of developmental psychology, on children’s theory of mind (tom), addressed to how young children understand their own and others’ minds. however, like epistemic cognition research, tom research has been confined to a particular age range – in the case of tom the first years of life − with very little work addressed to older children. research on older children’s higher-order tom tends to have an atheoretical quality, placing more emphasis on the application of second-order understanding and how it affects other aspects of development rather than on the development of a comprehensive theory addressing issues such as how change occurs (miller, 2012). in understanding developing knowledge about knowing, then, there exists a conspicuous gap consisting of the decade between early childhood and adolescence. my aim is to fill this gap by identifying a continuous development and in doing so to examine the nature of the process of change. iordanou           | f l r     108   2. a model of development of epistemic cognition: what develops? drawing on chinn’s et al. (2011) model of epistemic cognition, i propose that the epistemic standards that individuals employ change developmentally. epistemic processes refer to strategies and other activities by which one can achieve knowledge. epistemic standards refer to the standards used to evaluate knowledge claims (chinn et al., 2011). the literature on epistemic cognition has focused on examining people's beliefs about knowledge and knowing – known as their epistemic beliefs (greene et al., 2016; kitchener, 2002). the model proposed here extends the literature significantly by proposing the examination of application of people’s epistemic beliefs (epistemic activity in action), rather than focusing on epistemic beliefs themselves. agreeing with sandoval (2005) that students’ beliefs about their own knowing may differ from their beliefs about scientists’ knowing, i extend this idea by proposing that the application of students’ epistemic beliefs in practice may differ from their epistemic beliefs. i further advocate that the distance between epistemic activity in action and epistemic cognition decreases as one acquires increasing awareness and control of each. given the focus of the present work on epistemic activity in action, the development of epistemic standards is a focus, although it is acknowledged that other components of epistemic cognition (i.e., epistemic processes, values) also develop and interact with epistemic standards (clement et al., 2015). insights from research in tom, testimony and argumentation contribute to addressing the question of what develops in the epistemic realm and supports knowledge change. 2.1. epistemic standards change developmentally 2.1.1. epistemic standards in early childhood the origins of epistemic cognition are identifiable in the early childhood achievements examined under the theory-of-mind literature (kuhn, cheney & weinstock, 2000). even young children form and revise beliefs. they are just not aware of doing so. for example, preschoolers have the tendency to report they have always known information they have just learned (taylor, esbensen & bennett, 1994). what influences young children to adopt and revise beliefs?  a growing literature on testimony provides insights regarding young children’s standards in judging the credibility of the source of new information. standards that young children employ include an informant’s expertise, age, power, group membership and relationship with the child, with children showing preference for informants who are experts in the domain that the information is related to, to older informants, to authority figures and to those who have close familiarity with or belong to the same group (harris & carriveau, 2011; mills, 2013). for example, children prefer to seek and endorse information from native-accented speakers (kinzler, corriveau & harris, 2011). other epistemic standards employed by young children include an informant’s record of accuracy (harris & carriveau, 2011) and an informant’s confidence about their knowledge (jaswal & malone, 2007). research examining young children’s reasoning with peers also offers insights regarding standards that young children employ in modifying their beliefs. this line of research shows that even three-year-olds use evidence (e.g., this is ice) to justify their claims (kӧymen, rosenbaum & tomasello, 2014). notably, research shows that young children’s standards change with age. remarkable differences have been reported between the age of three and four. while three-year-olds show preference for egocentric standards, such as familiarity with the informant (corriveau, harris, et al., 2009), four-year-olds prefer more objective and more germane standards, such as the informant’s history of reliability. for example, four-year-olds show preference toward informers who have been reliable in their past performance, even when they have to reject familiar individuals who weren’t reliable in recent judgments in favour of reliable strangers (corriveau & harris, 2009). children by the age of four show greater sensitivity to the number and kind of errors made by an informant (mills, 2013) and are able to distinguish experts based on their domain of expertise, showing preference for one expert over another depending on the issue they are dealing with and experts’ domain of expertise, compared to three-year-olds (koenig & jaswal, 2011; sobel & corriveau, 2010). for example, when children were presented with a new dog, they preferred to ask the dog expert rather than a novice about the name of the dog (koenig & jaswal, 2011). furthermore, five-year-olds show better understanding of how iordanou           | f l r     109   relevant facts can be used to affect knowledge change in others, as evidenced by the production of more justifications in their dialogues. importantly, they are also more open to changing their knowledge than three-year-olds (kӧymen, rosenbaum & tomasello, 2014). this change during the third to fourth year of life takes place at the same time as major developmental milestones are observed in children’s cognitive development, as manifested in their achievements in the false belief task. this co-incidence supports the more general position proposed here that development of epistemic cognition should be situated in cognitive development more generally. 2.1.2. epistemic standards in middle childhood the epistemic standards employed by elementary-school children remain predominantly egocentric. barzilai and zohar (2012), examining via think-alouds how elementary school students judged the trustworthiness of websites, found that the predominant epistemic standard employed was personal authority – asking, for example, their mom. more objective and rigorous epistemic standards, such as website author’s expertise, scientific evidence or author biases were very rarely employed. yet, during elementary school years, there is a developing appreciation of the epistemic standard of judging epistemic products (e.g., arguments, models) on the basis of their fit to evidence. pluta, chinn, and duncan (2011) asked elementary school students to generate a list of criteria to evaluate scientific models and found that a quarter reported criteria relating to model fit with evidence (although other criteria were more commonly reported). although elementary-school students show a developing appreciation for data to support their claims, they show preference for data from their own knowledge or experience rather than more objective scientific evidence (amsel & brock, 1996; anderson, chinn, change, waggoner, & yi, 1997; kuhn & moore, 2015). kuhn and moore (2015) examined how elementary-school students used evidence in their dialogues to convince their peers to change beliefs about a social science and a physical science topic. they found that, even though a list of relevant shared evidence was available, about 90% of the evidence that students employed came from their personal knowledge and experience. similar results were observed in research examining how elementary-school students deal with evidence that disconfirms their prior beliefs. amsel and brock (1996) examined children’s behaviour when, in their experimentation, they encountered findings that contradicted their prior beliefs; they found that children in elementary childhood failed to use the new evidence to change their beliefs. students typically are biased in evaluating evidence; they tend to ignore evidence that contradicts their knowledge or distort evidence to fit their existing theories (chinn & brewer, 1993). much of the research on scientific reasoning reports similar results (lehrer & schauble, 2015; sandoval, sodian, koerber, & wong, 2014). 2.1.3. epistemic standards in adolescence in adolescence, attention to objective data increases, although data based on personal knowledge remains a more predominant epistemic standard. subjective epistemic standards, such as agreement with one’s own knowledge, are predominant in adolescents’ judgments about the trustworthiness of sources (mason, boldrin, & ariasi, 2010) and about the veracity of knowledge claims (mason, ariasi, & boldrin, 2011). iordanou and constantinou (2015) examined how 15and 16-year-olds argue with peers who hold opposing views on a socio-scientific issue, in a knowledge-rich learning environment. participants’ dialogue transcripts were analyzed in terms of the overall use of evidence, the amount of evidence per argument and per counterargument, the function of evidence use and the accuracy of the evidence employed. only a quarter of adolescents’ dialogue units contained evidence. adolescents employed evidence most of the time to support their own position rather than to weaken the opposing position. in terms of the epistemic standards employed, these older adolescents, like middle-school students (kuhn & moore, 2015), used undocumented evidence claims from personal knowledge to support their claims. eighty percent of adolescents’ dialogue units made claims based on personal knowledge. besides limited use of evidence in argument production, limited employment of rigorous epistemic standards regarding evidence-claim coordination has been documented during argument evaluation. iordanou, muis and kendeou (2014) examined, using the think-aloud methodology, the processes that iordanou           | f l r     110   adolescents engage in when reading a text, focusing particularly on on-line processing of evidence. adolescents rarely judged the credibility of evidence using epistemic standards such as the number of empirical studies which support a particular finding, the methodology used to produce a finding (e.g. whether the scientific method was used) or the fit of a claim to evidence. 2.1.4. epistemic standards in adulthood the growing literature examining college students’ judgments of trustworthiness of different information sources provides some insight regarding adults’ epistemic standards. adults show the ability to evaluate experts from different disciplines (e.g. biologist, chemist, earth scientist), by estimating the extent to which they might possess relevant knowledge about a specific science topic (bromme & thomm, 2015) – an ability not shown by young children. undergraduate students consider official documents as more credible sources of information than newspapers (bråten, strømsø & salmerón, 2011). however, they place unwarranted faith in textbooks (wineburg, 1991). examining undergraduate students’ dialogues with peers to gain insight into their epistemic activity in action, iordanou and constantinou (2014) have observed that even adults do not employ evidence consistently to support their claims. in iordanou and constantinou (2014) study, adults’ percentage of usage of evidence which functioned to support their claims was only 25% and the percentage of usage of evidence which functioned to weaken other’s claims was even less − 18%. kuhn (2016) in an effort to gain a better understanding of the factors underlying the limited use of evidence in argumentation, examined whether individuals’ limitations in conceptions of both evidence and causality may constrain their potential to employ evidence in argumentation. in that study, adults were presented with a scenario and were asked to choose among three options the one that could serve as the strongest evidence against an opponent’s claim. findings show that half of the adult participants chose the option which included no evidence and simply made a contrasting causal assertion, showing limitations in appreciation and application of epistemic standards pertaining to evidence-claim coordination, even in adulthood. similarly, kuhn et al. (2000) found that only half of a group of adults consisting of undergraduate students, college students and professionals reached an evaluativist way of thinking, that is an understanding that knowledge evolves through coordination of theory with data. the only exception was the group of experts − all of whom exhibited an evaluativist mode of thinking. limitations in reasoning about evidence have been observed not only in laypersons’ reasoning but also in scientists’ reasoning, such as confirmation bias in interpreting evidence in order to provide support to favourite theories. yet, despite these limitations, experts’ epistemic standards are more in line with the rigorous standards employed in formal science. scientists employ rigorous epistemic standards and practices (e.g., peer review, statistics), while they also reflect and revise those standards that make the distinction of the strongest theories at a particular time and the growth of knowledge possible (chinn & buckland, 2012). self-reflection on the way of knowing and the standards that one employs is according to habermas the most comprehensive way of knowing, and one of the highest criteria employed by doctoral examiners to judge thesis quality, compared to either the empirical-analytical way of knowing which places emphasis on facts, the objective elements of knowing, or the historical-hermeneutic way of knowing which stresses interpretation, the more subjective elements of knowing (clement et al., 2015). 2.2. epistemic understanding of evidence and theory-evidence coordination both develop underlying age changes observed in epistemic standards is changing epistemic understanding regarding evidence and its coordination with theory, which also undergoes development. three-year-olds make highly subjective judgments (wildenger, hofer & burr, 2010) and attribute thinking as reflection of external reality. one of the landmarks in this developmental progression of epistemic understanding of evidence and evidence-theory coordination is the understanding that evidence is different from a claim, which is reflected in pre-schoolers’ success in the false belief task by the age of four (perner & davies, 1991). this success reflects understanding that different information leads to different beliefs. this iordanou           | f l r     111   understanding also entails the understanding that evidence is different from information; information only becomes evidence in relation to a claim. during middle childhood, the understanding that one piece of evidence is amenable to different interpretations is achieved (lalonde & chandler, 2002). the understanding that different individuals can assign different meanings to the same stimulus (carpendale & chandler, 1996) reflects achievement of more mature understanding than earlier success in tom tasks, since it involves an understanding that different beliefs could result from the same input, not different inputs. in other words, middle-school children realize not only that people can form different beliefs when they have access to different information, as was the case with tom tasks, but also when they have access to the same information. middle school students showed also a better understanding of evidence, reflected in their ability to distinguish between causes and reasons, than pre-schoolers. astington, pelletier and homer (2002) found that seven-year-olds exhibited better ability in distinguishing between the cause of a situation and a person’s reason for believing it than pre-schoolers; this ability was related with second-order false-belief understanding, that is, their awareness that people have beliefs about the content of others’ minds. yet, understanding of human knowing is not yet fully developed, as it is not applied consistently nor with appropriate justification (eisbach, 2004). also, even though middle-school children are able to understand multiple interpretations of simple stimuli, which offer clear-cut dual interpretations, such as ambiguous pictures (lalonde & chandler, 2002), nonetheless, when explicitly asked to respond to stimuli that do not offer any facilitative, perceptual cues of the existence of alternative interpretation, as is the case in most real-life situations, they are not able to do so. sandoval and millwood (2005), examining high school students’ written explanations for problems on natural selection, found that adolescents made noninterpretive references to data (e.g., the graph shows x). this finding suggests that even high school students believe that “claims are not distinct from data but are somehow embodied in them, that a particular graph or table or other inscription directly represents some aspect of the natural world and consequently has but one meaning” (p. 49, sandoval & millwood, 2005). children’s limited understanding of the fact that physical or other phenomena are not self-explanatory, but rather are amenable to different interpretations, is also reflected in their preference for direct observation as a means for knowing. when elementary school students were asked to explain how they could become more certain about what happened in a historical event or about the cause of frog deformities, for which there are contradictory accounts, most students reported that eyewitness accounts (e.g. talk to anyone who was around at that time) would be sufficient to provide an explanation. only a few elementary school students reported that investigation and interpretation of evidence can provide insights to what happened or what is the cause of the problem, respectively (iordanou, 2016; kuhn, iordanou, pease & wirkala, 2008). the understanding that evidence supports claims has its roots in early childhood – even young children draw on evidence from their personal experience to support or contradict claims (kӧymen, rosenbaum & tomasello, 2014; wildenger, hofer & burr, 2010). yet, this understanding is not fully developed even by adulthood. research in the area of argumentation shows that the epistemic understanding that evidence can be employed to offer support to theories precedes the development of the understanding that evidence also plays the important role of weakening claims (iordanou & constantinou, 2014, 2015; kuhn, zillmer, crowell, & zavala, 2013). the understanding that evidence can be used to weaken claims is related to understanding that evidence can have different interpretations and that evidence can have different functions in relation to different claims. 2.3. epistemic beliefs about standards vs. application of epistemic standards examining the development of epistemic cognition reveals two paradoxes. the first is the commonly encountered one of a discrepancy between beliefs and their expression in action. in other words, there appears to be a discrepancy between individuals’ beliefs regarding epistemic standards and the application of epistemic standards in practice. for example, although children might show to endorse the epistemic belief that experts in a domain are more reliable than non-experts, as seen in their preference between experts when iordanou           | f l r     112   there are clear differences between them concerning the degree of their prior knowledge, children generally adopt in non-problematic fashion, during epistemic action, information from experts in different knowledge domains (harris & koening, 2006). elementary school students appear to adopt the epistemic belief, when asked, that the epistemic standard of model fit with evidence is useful to evaluate scientific models (pluta, chinn, & duncan, 2011), nonetheless, there is evidence that they do not employ this epistemic standard in action (iordanou, muis, & kendeou, 2014; sandoval et al., 2014). in tom tasks, when judging others’ mental states, adults underestimated the probability that a more ignorant other would search incorrectly as a result of holding a false belief, even though they were aware of the difference between their own and the other’s perspective (zhang et al., 2010; birch & bloom, 2007). in addition, in tasks entailing evaluation of texts, stømsø, bråten and britt (2011) found that, although some undergraduate students reported that they endorse the epistemic belief of justification of knowledge based on evidence, when asked to indicate the criteria on which they have based their judgments of trustworthiness in epistemic action, they reported both advanced criteria – content – but also less advanced ones – their own opinion. similarly, iordanou, muis, and kendeou (2014) reported a discrepancy between adolescents’ and adults’ epistemic knowledge and their epistemic activity in action. in that study, although some individuals acknowledged that they endorse the epistemic belief that evaluation and interpretation of evidence is central for knowing, when directly asked how they could become more certain about their knowledge, they did not engage spontaneously in evaluation of evidence during epistemic action, when they were reading a text. focusing on a particular age, we also observe lack of consistency in the application of epistemic standards. for example, pre-schoolers do not show consistency in using the epistemic standard of an informant’s history of errors, including the number and kind of errors made when choosing informants (mills, 2013); neither do they show consistency in assigning test questions correctly to different experts (aguiar, stoess & talyor, 2012). in the examples presented above an inconsistency between individual’s epistemic beliefs about standards and the application of those epistemic standards has been observed, as well as an inconsistency in the application of epistemic standards. individuals’ epistemic action is not always consistent with their epistemic beliefs. the second paradox appears in examining epistemic activity across the lifespan. although very young children show competence with respect to a particular epistemic criterion, older individuals exhibit limitations in the application of the same criterion. for example, even though some research findings show that pre-schoolers are able to judge an informant’s credibility based on the quality of the informant’s argument rather than on his or her power (castelain, bernard, van der henst, & mercier, 2015), other findings show that most undergraduate students do not engage in evaluation of arguments, examining, for example, whether scientific evidence supports a knowledge claim while researching information on the web (mason, boldrin, & ariasi, 2010) or when reading a text (iordanou, muis, & kendeou, 2014). it is proposed here that the mechanism behind development of epistemic cognition, which explains the two paradoxes described above, is the development of individuals’ epistemic awareness of their epistemic beliefs and conscious control of application of epistemic standards, an issue that we discuss below. 2.4. understanding of epistemic standards and control of their application develop and support epistemic cognition in action studying students engaging in dialogic argumentation over time offers insights regarding how both knowledge and epistemic cognition change. iordanou and constantinou (2015), employing the micro-genetic method, a powerful method for understanding epistemic cognitive development (sandoval, 2014), examined how students use evidence to influence the beliefs of their peers. eleventh graders, working with a partner, engaged in electronic argumentative dialogues with classmates who held an opposing view on the topic and in some evidence-focused reflective activities, based on transcriptions of their dialogues. another sixteen 11th graders, who studied the data base in the learning environment for the same amount of time as experimental-condition students but did not engage in an argumentative discourse activity, served as a comparison condition. the findings of this study were consistent with findings of other studies (iordanou & constantinou, 2014; kuhn & moore, 2015) in showing that after extensive engagement in argumentative iordanou           | f l r     113   activities, students exhibited a shift from presenting their “right”, self-evident theories of how things are, without providing any data to support their argument beyond presenting their personal opinions, to employing data to support their positions and offering alternative interpretations for a particular piece of evidence. in addition, students developed an appreciation of the epistemic understanding that evidence can be used to weaken others’ claims, which appears to be a more challenging developmental achievement than understanding that evidence can be used to support one’s own claims. finally, students made more specific reference to evidence and its source after sustained engagement in argumentative activities, a finding which is also consistent with other studies (iordanou & constantinou, 2014), suggesting that the process of coordinating evidence with claims, and the awareness of the need to do so, came under increasing conscious control over time. the analysis of participants’ dialogues over the course of the intervention provided further support to this suggestion. in particular, the micro-genetic analysis showed that, in addition to the increase observed in the use of evidence and the function of evidence employed, an increase was observed in students’ meta-level statements regarding evidence (e.g., ‘‘give us some evidence’’, ‘‘you have not provided evidence’’) over the course of the intervention, revealing a developing epistemological understanding of the epistemic standard of evaluating a theory based on its fit to evidence. similar results were observed in chinn, duschl, duncan, buckland, and pluta’s (2008) study, where middle school students engaged in argumentation and reflective activities aimed at constructing, revising, and evaluating scientific models on the basis of evidence, over the course of an academic year. by the end of the intervention, students in the experimental condition exhibited greater advances not only in their ability to effectively coordinate models and evidence, but also in their understanding of epistemic criteria. a shift was observed from non-evidential criteria (e.g., have words and pictures) to evidential criteria, linking models to evidence. the findings of iordanou and constantinou (2015) and chinn et al.’s (2008) studies have two important implications. the first implication is that dialogic argumentation can offer a suitable setting for studying students’ epistemic activity in action and gaining a better understanding of how epistemic cognition changes. the second implication is that argumentation appears to be a promising pathway to support the development of epistemic cognition (iordanou, 2016; iordanou, kendeou, & beker, 2016; sandoval, 2005). engagement in argumentation is a fruitful way for making tacit epistemic beliefs, reflected first in epistemic action, explicit, as well as for changing epistemic beliefs (iordanou, 2016). the work of iordanou (2016) showed that engagement in dialogic argumentative activities supported the development of more evaluativist epistemic beliefs, that is an understanding that knowledge evolves through coordination of theory with data and through evaluation, the position found to be best supported by argument and evidence would be determined to have more merit compared to alternative positions (kuhn et al., 2000). the increasing acquisition of awareness and conscious control of application of epistemic standards proposed here, and reflected in the iordanou and constantinou’s (2015) findings as well as in findings from studies on testimony and tom (corriveau & harris, 2009; koenig & jaswal, 2011; mills, 2013; sobel & corriveau, 2010), are in line with other findings in cognitive development showing a developing metacognitive monitoring from childhood to adolescence (kitsantas & zimmerman, 2002; roderer & roebers, 2014; van der stel & veenman, 2010). tom research examining adults’ eye movements, while they were following a director’s instructions for moving items, shows that adults initially interpreted the director’s instructions egocentrically, just like children, but were faster and more effective in correcting a wrong interpretation (epley, morewedge & keysar, 2004). findings like this one suggest that adults’ better metacognitive control is what enables them to “correct” their egocentric errors and exhibit superior behaviour than do children (apperly, warren, andrews, grant, & todd, 2011). 2.5. specificity of epistemic standards there is ample evidence pointing to the domain-specificity of epistemic cognition (muis, bendixen, & haerle, 2006). individuals’ epistemic cognition differs across domains (kuhn et al., 2000) and advancement in epistemic cognition in one domain does not necessarily transfer in other domains (iordanou, iordanou           | f l r     114   2010; 2016; hofer, 2004). in iordanou’s (2016) study, notable differences were observed in the epistemic standards employed between a social science topic and a physical science topic. when elementary school students were asked to justify their knowledge of a physical science topic – dinosaurs’ extinction – and a social science topic – home-schooling −, the majority of the students reported scientific evidence to justify their knowledge in the physical science topic, while they employed claims from general knowledge or personal experience (e.g. “you don’t have friends at home”) to justify their knowledge in the social science topic. domain differences were also observed in both participants’ epistemic beliefs and epistemic activity in action between different knowledge domains in iordanou et al.’s (2014) study. young adolescents and adults in that study engaged in more epistemic processing of evidence in the history domain than in the science domain. in particular, they engaged more in judging an evidence’s credibility while reading a text in the history domain, than in the science domain. behind this domain-specificity of epistemic cognition, reside domain-specific challenges regarding the development of epistemic cognition. kuhn et al. (2008) have suggested that in the social domain the major challenge in achieving sophisticated epistemic cognition is different from the challenge in the science domain. in a word, in the social domain, the challenge is to come to terms with the concern that human interpretation plays an unmanageable, overpowering role, while in the science domain, the major challenge is to recognize that human interpretation plays any role at all. in the science domain, the entry of human interpretation into what was previously regarded as direct perception of a single reality must be recognized and come to be understood in positive terms. human construction of alternative possibilities (multiple representations of truth, or theories) needs to be coordinated with empirical evidence, in an ongoing process that constitutes scientific work. in the social domain, in contrast, human interpretation is more readily recognized and the danger is one of a permanent stall in a radical relativism, with the evil of subjectivity seen as overpowering the quest for any knowledge beyond subjective opinion. epistemic standards also must be examined as a function of context. students’ epistemic standards differ when reflected in essays versus dialogues. in the kuhn and moore (2015) study, middle-school students used more evidence from their own personal knowledge and experience in their dialogues, about 90%, than in their essays, 40%, suggesting the dialogue was a more authentic experience for them. finally, specific content also introduces variation in standards. here, more research is needed. bråten, strømsø, and salmerón, (2011) found that readers with low topic knowledge failed to employ the most appropriate epistemic standards, whereas bromme and thomm (2015) found that adults’ judgments regarding reliable informants were not related to participants’ prior knowledge, general science knowledge or their study subject. similarly, mason, boldrin and ariasi (2010) found that prior knowledge was not related to epistemic activity in action, whereas iordanou, muis, and kendeou (2014) found that individuals’ prior knowledge predicted their epistemic cognition in action. 3. conclusions and future research the question of how knowledge changes as individuals progress through the lifespan requires better answers. research findings point to differences between children and adults in the way they make judgments. for example, tenney, small, kondrad, jaswal, and spellman (2011) found that adults take into consideration information regarding informants’ calibration, that is how well one’s confidence matches one’s likelihood of being correct, whereas children ignore this information and tend to rely more on an informant’s confidence. the review presented here proposes that epistemic understanding of theory-evidence coordination develops gradually and different forms of understanding develop at different ages. also, the present paper presents evidence showing that there is a discrepancy between epistemic beliefs and their expression in action. with age and expertise understanding of epistemic standards and control of their application develop and support epistemic cognition in action. there is a need for more developmental research, especially longitudinal studies, to enhance our understanding of the forms that epistemic cognitive development take and to explain why epistemic development occur or fail to occur. iordanou           | f l r     115   to satisfy the quest for a better understanding of epistemic cognitive development, there is a need for new measures that would allow us to examine more deeply and thoroughly what develops (chinn et al., 2011). dynamic instruments that examine individuals’ epistemic cognition as a dynamic, complex construct need to be employed. some promising measures are think-aloud protocols (hofer, 2004), eye-tracking techniques, collaborative discussions and computer-based learning environments (greene, muis, & pieschl, 2010), all of them employed in micro-genetic investigations. there is also need for a better understanding of the specificity of epistemic cognition. research suggests that epistemic cognition has both general and context-specific elements (muis, bendixen, & haerle, 2006; sinatra, kienhues, & hofer, 2014). some aspects of epistemic cognition, such as the appreciation of evidence, transfer across contexts (iordanou & constantinou, 2014; 2015), whereas other aspects, such as the epistemic criteria that individuals employ for adopting and revising claims appear to be domain specific (iordanou, 2016; kuhn et al., 2008). taking into account the complex and multifaceted nature of epistemic cognition, future research needs to examine the specificity question of epistemic cognition at a more finegrained level, addressing questions such as how epistemic standards vary across conditions and why this is the case. finally, there is a need for future research to examine how the development of epistemic cognition in action can be supported. engagement in dialogic argumentation appears a promising pathway of supporting understanding that there is no single self-evident truth and that multiple interpretations may exist of the same phenomenon as the human mind plays an active role in ascribing meaning to the world (carpendale & lewis, 2006; iordanou & constantinou, 2015; moshman, 2004; walker, wartenberg, & winner, 2012). engagement also in explicit reflection about the role of evidence in reasoning and about epistemic standards are promising means for supporting an appreciation of the role of evidence in forming and revising knowledge (chinn & buckland, 2012; iordanou & constantinou 2014; 2015). future research should examine such methods further. in summary, the purpose of the present paper has been to draw a strong continuous link between the earliest understanding of other minds and the tasks that confront adults throughout the life span – that of interpreting evidence and coordinating it with what they already take to be true, in a manner over which they exercise conscious control. adults continue to do so imperfectly to be sure (kuhn, 2016) but their skill has developed from earlier levels and has the potential to continue to develop. i propose that epistemic cognition builds on increasing awareness and epistemic understanding of theory-evidence coordination and of the role of the human mind in interpreting reality. the standards for knowledge formation and revision are closely connected with epistemic understanding of theory-evidence coordination and change across the lifespan, as well as control of their application. competence in understanding theory-evidence coordination and the role of the human mind in knowing has its roots in early tom achievements and proceeds gradually from there towards more and more mature and complete understanding, in a process that ideally never ends. future research should go beyond a focus on what people believe and increase attention not only on how people choose what to believe, but on how these standards for choice themselves evolve and are applied within reallife contexts. lastly, addressing the question of how researchers and educators can best support individuals’ development in these respects promises to have profound consequences for people’s lives. keypoints the how question of knowledge change is examined evidence from tom and evidence theory coordination literature are examined a focus on epistemic activity in action is proposed the standards for knowledge formation and revision change developmentally the process increasingly comes under conscious control iordanou           | f l r     116   references agruiar, n. r., stoess, c. j., & taylor, m. (2012). the development of children’s ability to fill the gaps in their knowledge by consulting experts. child development, 83(4), 1368-81. amsel, e., & brock, s. (1996). the development of evidence evaluation skills. cognitive development, 11, 523-550. doi: http://dx.doi.org/10.1111/j.1467-8624.2012.01782.x. anderson, r. c., chinn, c., chang, j., waggoner, m., & yi, h. (1997). on the logical integrity of children’s arguments. cognition and instruction, 15(2), 135-167.  doi: http://dx.doi.org/10.1207/s1532690xci1502_1 apperly, i. a., warren, f., andrews, b. j., grant, j., & todd, s. (2011). developmental continuity in theory of mind: speed and accuracy of belief-desire reasoning in children and adults. child development, 82(5), 1691-1703. doi: http://dx.doi.org/10.1111/j.1467-8624.2011.01635.x astington, j. w., pelletier, j., & homer, b. (2002). theory of mind and epistemological development: the relation between children's second-order false-belief understanding and their ability to reason about evidence. new ideas in psychology, 20(2), 131-144.  doi: http://dx.doi.org/10.1016/s0732118x(02)00005-3 barzilai, s., & zohar, a. (2012). epistemic thinking in action: evaluating and integrating online sources. cognition and instruction, 30(1), 39-85. doi: http://dx.doi.org/10.1080/07370008.2011.636495 bendixen, l., & rule, d. (2004). an integrative approach to personal epistemology: a guiding model. educational psychologist, 39(1), 69-80. doi: http://dx.doi.org/10.1207/s15326985ep3901_7 birch, s. a. j., & bloom, p. (2007). the curse of knowledge in reasoning about false beliefs. psychological science, 18(5), 382-386. doi: http://dx.doi.org/10.1111/j.1467-9280.2007.01909.x bråten, i., britt, m. a., strømsø, h. i., & rouet, j. (2011). the role of epistemic beliefs in the comprehension of multiple expository texts: toward an integrated model. educational psychologist, 46(1), 48-70. doi: http://dx.doi.org/10.1080/00461520.2011.538647 bråten, i., strømsø, h. i., & salmerón, l. (2011). trust and mistrust when students read multiple information sources about climate change. learning and instruction, 21(2), 180-192.  doi: http://dx.doi.org/10.1016/j.learninstruc.2010.02.002 bromme, r., & thomm, e. (2015). knowing who knows: laypersons’ capabilities to judge experts’ pertinence for science topics. cognitive science, 38(8) 1-12.  doi: http://dx.doi.org/10.1111/cogs.12252 carpendale, j. i., & chandler, m. j. (1996). on the distinction between false belief understanding and subscribing to an interpretive theory of mind. child development, 67(4), 1686-1706. doi: http://dx.doi.org/10.1111/j.1467-8624.1996.tb01821.x carpendale, j., & lewis, c. (2006). how children develop social understanding. oxford: blackwell. castelain, t., bernard, s., van der henst, j.-b., & mercier, h. (2015). the influence of power and reason on young maya children's endorsement of testimony. developmental science, 18(1), 1-10. doi: http://dx.doi.org/10.1111/desc.12336 chinn, c. a., buckland, l. a., & samarapungavan, a. (2011). expanding the dimensions of epistemic cognition: arguments from philosophy and psychology. educational psychologist, 46(3), 141-167.  doi: http://dx.doi.org/10.1080/00461520.2011.587722 chinn, c. a., & brewer, w. f. (1993). the role of anomalous data in knowledge acquisition: a theoretical framework and implications for science instruction. review of educational research, 63, 1-49. doi: http://dx.doi.org/10.2307/1170558 chinn, c. a., & buckland, l. a. (2012). model-based instruction: fostering change in evolutionary conceptions and in epistemic practices. in k. s. rosengren, s. k. brem, e. m. evans, & g. m. sinatra (eds.), evolution challenges: integrating research and practice in teaching and learning about evolution (pp. 211-232). oxford: oxford university press.  doi: http://dx.doi.org/10.1093/acprof:oso/9780199730421.003.0010 chinn, c. a., duschl, r. a., duncan, r. g., buckland, l. a., & pluta, w. j. (2008). a microgenetic classroom study of learning to reason scientifically through modeling and argumentation. in g. iordanou           | f l r     117   kanselaar, j. van merriënboer, p. kircshner, &t. de jong (eds.), international perspectives in the learning sciences: creating a learning world. proceedings of the 8th international conference for the learning sciences-(vol. 3, pp. 14-15). utrecht, the netherlands: international society of the learning sciences. clement, n., lovat, t., holbrook, a., kiley, m., bourke, s., paltridge, b., ... & mcinerney, d. m. (2015). exploring doctoral examiner judgements through the lenses of habermas and epistemic cognition. in theory and method in higher education research (pp. 213-233). emerald group publishing limited. doi: http://dx.doi.org/10.1108/s2056-375220150000001010 corriveau, k. h., & harris, p. l. (2009). preschoolers continue to trust a more accurate informant 1 week after exposure to accuracy information. developmental science, 12, 188–193.  doi: http://dx.doi.org/10.1111/j.1467-7687.2008.00763.x corriveau, k.h., harris, p.l., meins, e., ferneyhough, c., arnott, b., elliott, l., liddle, b., hearn, a., vittorini, l. & de rosnay, m. (2009). young children’s trust in their mother’s claims: longitudinal links with attachment security in infancy. child development, 80(3), 750-761. doi: http://dx.doi.org/10.1111/j.1467-8624.2009.01295.x eisbach, a. o. (2004). children’s developing awareness of diversity in people’s trains of thought. child development, 75(6), 1694-1707.  doi: http://dx.doi.org/10.1111/j.1467-8624.2004.00810.x epley, n., morewedge, c. k., & keysar, b. (2004). perspective taking in children and adults: equivalent egocentrism but differential correction. journal of experimental social psychology, 40, 760-768. doi: http://dx.doi.org/10.1016/j.jesp.2004.02.002 greene, j. a., muis, k. r., & pieschl, s. (2010). the role of epistemic beliefs in students’ self-regulated learning with computer-based learning environments: conceptual and methodological issues. educational psychologist, 45(4), 245-257. doi: http://dx.doi.org/10.1080/00461520.2010.515932 greene, j. a., sandoval, w. a., bråten, i. (eds.). handbook of epistemic cognition. new york, ny: routledge. harris, p. l., & corriveau, k. h. (2011). young children's selective trust in informants. philosophical transactions of the royal society b: biological sciences, 366(1567), 1179-1187. doi: http://dx.doi.org/10.1098/rstb.2010.0321 harris, p. l., & koening, m. a. (2006). trust in testimony: how children learn about science and religion. child development, 77(3), 505-534. doi: http://dx.doi.org/10.1111/j.1467-8624.2006.00886.x hofer, b. k. (2004). epistemological understanding as a metacognitive process: thinking aloud during online searching. educational psychologist, 39(1), 43-55. doi: http://dx.doi.org/10.1207/s15326985ep3901_5 iordanou, k. (2010). developing argument skills across scientific and social domains. journal of cognition and development, 11(3), 293-327. doi: http://dx.doi.org/10.1080/15248372.2010.485335 iordanou, k. (2016). developing epistemological understanding through argumentation in scientific and social domains. zeitschrift für pädagogische psychologie.  30(2-3), 109-119.  doi: http://dx.doi.org/10.1024/1010-0652/a000172 iordanou, k., & constantinou. c. p. (2014). developing pre-service teachers’ evidence-based argumentation skills on socio-scientific issues. learning & instruction, 34, 42-57.  doi: http://dx.doi.org/10.1016/j.learninstruc.2014.07.004 iordanou, k., & constantinou. c. p. (2015). supporting use of evidence in argumentation through practice in argumentation and reflection in the context of socrates learning environment. science education, 99, 282–311.  doi: http://dx.doi.org/10.1002/sce.21152 iordanou, k., kendeou., p., & beker, k. (2016). argumentative reasoning. in w. sandoval, j. greene, & i., bråten. (eds). handbook of epistemic cognition, (39-53). new york, ny: routledge.   doi: http://dx.doi.org/10.4324/9781315795225 iordanou, k., muis, k., & kendeou, p. (2014). epistemic understanding and meta-level processing of evidence when reading a text. paper presented at the earli sig2 conference. amsterdam, the netherlands. iordanou           | f l r     118   jaswal, v. k., & malone, l. s. (2007). turning believers into skeptics: 3-year-olds' sensitivity to cues to speaker credibility. journal of cognition and development, 8(3), 263-283.  doi:   http://dx.doi.org/10.1080/15248370701446392 kinzler, k. d., corriveau, k. h., & harris, p. l. (2011). children’s selective trust in native-­‐accented speakers. developmental science, 14(1), 106-111.   doi: http://dx.doi.org/10.1111/j.14677687.2010.00965.x kitsantas, a., & zimmerman, b. j. (2002). comparing self-regulatory processes among novice, non-expert, and expert volleyball players: a microanalytic study. journal of applied sport psychology, 14, 91–105. doi: http://dx.doi.org/10.1080/10413200252907761 koenig, m. a., & jaswal, v. k. (2011). characterizing children’s expectations about expertise and incompetence: halo or pitchfork effects? child development, 82(5), 1634-1647. doi: http://dx.doi.org/10.1111/j.1467-8624.2011.01618.x köymen, b., rosenbaum, l., & tomasello, m. (2014). reasoning during joint decision-making by preschool peers. cognitive development, 32, 74-85.  doi: http://dx.doi.org/10.1016/j.cogdev.2014.09.001. kuhn, d. (2016). a role for reasoning in a dialogic approach to critical thinking. topoi, 1-8. doi: http://dx.doi.org/10.1007/s11245-016-9373-4 kuhn, d., cheney, r., & weinstock, m. (2000). the development of epistemological understanding. cognitive development, 15, 309–328. doi: http://dx.doi.org/10.1016/s0885-2014(00)00030-7 kuhn, d., iordanou, k., pease, m., & wirkala, c. (2008). beyond control of variables: what needs to develop to achieve skilled scientific thinking? cognitive development, 23, 435–451.   doi: http://dx.doi.org/10.1016/j.cogdev.2008.09.006 kuhn, d., & moore, w. (2015). argumentation as core curriculum. learning: research and practice, 1(1), 66-78. doi: http://dx.doi.org/10.1080/23735082.2015.994254 kuhn, d., zillmer, n., crowell, a., & zavala, j. (2013). developing norms of argumentation: metacognitive, epistemological, and social dimensions of developing argumentive competence. cognition and instruction, 31(4), 456-496. doi: http://dx.doi.org/10.1080/07370008.2013.830618 lalonde, c. e., & chandler, m. j. (2002). children’s understanding of interpretation. new ideas in psychology, 20(2-3), 163-198.  doi:  http://dx.doi.org/10.1016/s0732-118x(02)00007-7 lehrer, r., & schauble, l. (2015). the development of scientific thinking. handbook of child psychology and developmental science. 2(16),1-44. (edited work) doi: http://dx.doi.org/10.1002/9781118963418.childpsy216 mason, l., ariasi, n., & boldrin, a. (2011). epistemic beliefs in action: spontaneous reflections about knowledge and knowing during online information searching and their influence on learning. learning and instruction, 21, 137-151. doi: http://dx.doi.org/10.1016/j.learninstruc.2010.01.001 mason, l., boldrin, a., & ariasi, n. (2010). searching the web to learn about a controversial topic: are students epistemically active? instructional science, 38, 607-633.   doi: http://dx.doi.org/10.1007/s11251-008-9089-y miller, s. a. (2012). theory of mind: beyond the preschool years. new york, ny: psychology press. mills, c. m. (2013). knowing when to doubt: developing a critical stance when learning from others. developmental psychology, 49(3), 404-418. doi: http://dx.doi.org/10.1037/a0029500 moshman, d. (2004). from inference to reasoning: the construction of rationality. thinking and reasoning, 10(2), 221 – 239. doi: http://dx.doi.org/10.1080/13546780442000024 muis, k. r., bendixen, l. d., & haerle, f. c. (2006). domain-generality and domain-specificity in personal epistemology research: philosophical and empirical reflections in the development of a theoretical framework. educational psychology review, 18(1), 3-54. doi: http://dx.doi.org/10.1007/s10648-006-9003-6 muis, k. r., kendeou, p., & franco, g. m. (2011). consistent results with the consistency hypothesis? the effects of epistemic beliefs on metacognitive processing. metacognition and learning, 6, 45-63.  doi: http://dx.doi.org/10.1007/s11409-010-9066-0 perner, j., & davies, g. (1991). understanding the mind as an active information processor: do young children have a “copy theory of mind”? cognition, 39, 51-69.   doi: http://dx.doi.org/10.1016/00100277(91)90059-d iordanou           | f l r     119   perry, w. g. (1970). forms of intellectual and ethical development in the college years: a scheme. new york: holt, rinehart and winston. pluta, w. j., chinn, c. a., & duncan, r. g. (2011). learners’ epistemic criteria for good scientific models. journal of research in science teaching, 48(5), 486-511. doi: http://dx.doi.org/10.1002/tea.20415 roderer, t., & roebers, c. m. (2014). can you see me thinking (about my answers)? using eye-tracking to illuminate developmental differences in monitoring and control skills and their relation to performance. metacognition and learning, 9(1), 1-23. doi: http://dx.doi.org/10.1007/s11409-013-9109-4 sandoval, w. a. (2005). understanding students’ practical epistemologies and their influence on learning through inquiry. science education, 89, 634-656.  doi: http://dx.doi.org/10.1002/sce.20065 sandoval, w. a. (2014). science education's need for a theory of epistemological development. science education, 98(3), 383-387.  doi: http://dx.doi.org/10.1002/sce.21107 sandoval, w. a., & millwood, k. a. (2005). the quality of students' use of evidence in written scientific explanations. cognition and instruction, 23(1), 23-55. doi: http://dx.doi.org/10.1207/s1532690xci2301_2 sandoval, w. a., sodian, b., koerber, s., & wong, j. (2014). developing children's early competencies to engage with science. educational psychologist, 49(2), 139-152.   doi:   http://dx.doi.org/10.1080/00461520.2014.917589 sinatra, g. m., kienhues, d., & hofer, b. k. (2014). addressing challenges to public understanding of science: epistemic cognition, motivated reasoning, and conceptual change. educational psychologist, 49(2), 123-138.  doi:  http://dx.doi.org/10.1080/00461520.2014.916216 sobel, d. m., & corriveau, k. h. (2010). children monitor individuals’ expertise for word learning. child development, 81(2), 669-679. doi: http://dx.doi.org/10.1111/j.1467-8624.2009.01422.x stømsø, h. i., bråten, i., & britt, m. a. (2011). do students’ beliefs about knowledge and knowing predict their judgement of texts’ trustworthiness? educational psychology, 31 (2), 177-206. doi: http://dx.doi.org/10.1080/01443410.2010.538039 taylor, m., esbensen, b. m., & bennett, r. t. (1994). children's understanding of knowledge acquisition: the tendency for children to report that they have always known what they have just learned. child development, 65, 1581-1604.   doi: http://dx.doi.org/10.1111/j.14678624.1994.tb00837.x tenney, e. r., small, j. e., kondrad, r. l., jaswal, v. k., & spellman, b. a. (2011). accuracy, confidence, and calibration: how young children and adults assess credibility. developmental psychology, 47(4), 1065-1077. doi: http://dx.doi.org/10.1037/a0023273 van der stel, m., & veenman, m. v. (2010). development of metacognitive skillfulness: a longitudinal study. learning and individual differences, 20(3), 220-224.   doi: http://dx.doi.org/10.1016/j.lindif.2009.11.005 walker, c. m., wartenberg t. e., & winner e. (2012). engagement in philosophical dialogue facilitates children's reasoning about subjectivity. developmental psychology, 2(1), 1-10. doi: http://dx.doi.org/10.1037/a0029870 wildenger, l. k., hofer, b. k., & burr, j. e. (2010). epistemological development in very young knowers. in l. d. bendixen, & f. c. fleucht (eds.), personal epistemology in the classroom: theory, research and implications for practice (pp. 220-257). cambridge: cambridge university press.   doi: http://dx.doi.org/10.1017/cbo9780511691904.008 wineburg, s. s. (1991). on the reading of historical texts: notes on the breach between school and academy. american educational research journal, 28(3), 495-519.  doi: http://dx.doi.org/10.3102/00028312028003495 zhang, t., zheng, x., zhang, l., sha, w., deák, g., & li, h. (2010). older children’s misunderstanding of uncertain belief after passing the false belief test. cognitive development, 25, 158-165. doi: http://dx.doi.org/10.1016/j.cogdev.2009.12.001 morena -esteva et al publication frontline learning research vol.6 no. 3 (2018) 72 84 issn 2295-3159 application of mathematical and machine learning techniques to analyse eye tracking data enabling better understanding of children’s visual cognitive behaviours enrique garcia moreno-estevaa, sonia l. j. whiteb, joanne m. woodc, alex a. black c adepartment of education, university of helsinki, finland b faculty of education, queensland university of technology, australia c faculty of health, queensland university of technology, australia article received 9 may 2018/ revised 16 september/ accepted 17 october/ available online 7 december abstract in this research, we aimed to investigate the visual-cognitive behaviours of a sample of 106 children in year 3 (8.8 ± 0.3 years) while completing a mathematics bar-graph task. eye movements were recorded while children completed the task and the patterns of eye movements were explored using machine learning approaches. two different techniques of machine-learning were used (bayesian and k-means) to obtain separate model sequences or average scanpaths for those children who responded either correctly or incorrectly to the graph task. application of these machine-learning approaches indicated distinct differences in the resulting scanpaths for children who completed the graph task correctly or incorrectly: children who responded correctly accessed information that was mostly categorised as critical, whereas children responding incorrectly did not. there was also evidence that the children who were correct accessed the graph information in a different, more logical order, compared to the children who were incorrect. the visual behaviours aligned with different aspects of graph comprehension, such as initial understanding and orienting to the graph, and later interpretation and use of relevant information on the graph. the findings are discussed in terms of the implications for early mathematics teaching and learning, particularly in the development of graph comprehension, as well as the application of machine learning techniques to investigations of other visual-cognitive behaviours. graph interpretation; eye tracking; machine learning; mathematics education; gaze metrics info corresponding author: enrique.garciamoreno-esteva@helsinki.fi doi: https://doi.org/10.14786/flr.v6i3.365 1. introduction eye tracking is rapidly becoming an established technique for investigating the cognitive processes involved in learning mathematics and other subjects. the use of eye tracking makes it possible to make implicit visual-cognitive behaviours explicit, in order to better understand learning processes and subsequently inform educational practice (lai, et al., 2013). previous methods, such as interviewing and think aloud protocols, have been adopted to understand the different approaches and strategies used in a range of problem-solving tasks, however, there are limitations to these approaches. first, interviews undertaken following task completion are reliant on a participant accurately remembering and recalling specific steps. second, think aloud protocols assume the participant has the cognitive flexibility to think out loud while also engaging in the problem-solving task (rosenzweig, krawec, & montague, 2011). for example, to understand how young children engage in mathematics problem-solving tasks, the additional cognitive demand of think aloud protocols could potentially have an adverse impact on task performance. indeed, kotsopoulos and lee (2012) used modified think aloud protocols and real time naturalistic analysis of students completing mathematical problem-solving tasks, and reported that problem-solving often broke down in the early stages of understanding the task. the limitations with many of these approaches led van gog, paas, van merriënboer and witte (2005) to highlight the need for research methods to understand different problem-solving processes and supported the use of eye tracking methods for more complex tasks that involve a sequence of cognitive steps. an extensive body of eye tracking research has focused on the link between visual gaze and information processing. for example, how a student looks at a diagram is influenced by their preliminary intuitions and the conceptions activated by the task context (knoblich, ohlsson & raney, 2001). eye tracking research has revealed that more experienced problem-solvers (experts) can identify task relevant visual information more rapidly than less experienced individuals (novices), and their visual attention (eye fixation scanpaths) tend to be more focused on relevant than irrelevant regions of the visual stimulus (gegenfurtner, lehtinen & säljö, 2011; tsai, hou, lai, liu, & yang, 2012); these objective findings were corroborated by self-reported accounts of participants completing the task (tsai, et al., 2012). another study reported that novices display significantly more shifts in visual attention than experts and have longer gaze sequences (or scanpaths) for a given problem-solving task (kim, aleven & dey, 2014). furthermore, psychology research indicates that the order of fixations affects cognitive functions, such as memory (e.g. bochynska & laeng, 2015; rinaldi, brugger, bockisch, bertolini, girelli, 2015). bochynska and laeng (2015) used eye tracking in a visuospatial memory recognition task and found that both the spatial information and the order of fixations were important for visuospatial memory formation, with increased accuracy in trials where the elements were presented serially, in the same order as in the participant’s original fixation scanpath. the current research, utilised machine learning techniques to make sense of the order structure of the multiple scanpaths of primary school children while they completed a graph problem-solving task, in order to better understand the visual cognitive behaviours involved. the rationale for using machine learning analysis is that it allows us to examine the sequential (temporal) structure of the data. this is unlike the approach adopted in most eye tracking studies and allows novel insight into the underlying behaviours of children while completing these graph tasks. while there are other methods to examine the temporal structure of data, we felt our approach was the most suitable for our research purpose. our method is also faster and more automated than many other traditional methods. in order to interpret the scanpaths, our approach used the sequential framework of children's data comprehension, described by curcio (2010) which comprises; understanding, interpretation, and prediction with data. the first level of comprehension, ‘understanding’, requires reading of the information explicitly stated in the presented data (e.g. graph). the second level of comprehension, ‘interpretation’, requires reading between the data and integration of the presented information, this level of comprehension requires specific skills, such as comparison or computation (e.g. addition, subtraction, multiplication, division). the third level of comprehension, ‘prediction’, requires reading beyond the data and the use of existing knowledge to make inferences/predictions from the data. when making predictions, the information is neither explicitly nor implicitly presented in the graph. the bar graph task used in the present study required each child to engage in the first two levels of data comprehension: reading the question and basic details of the graph (understanding), and then reading between the different elements of information (interpretation) in order to complete the computation and arrive at the correct solution. most importantly, with graph interpretation, both visual and cognitive integration is required (ratwani, trafton & boehm-davis, 2008), where integration of relevant visual attributes on the graph, such as labels, pattern recognition, or other spatial features contribute to higher order visual clusters of information, that are compared to create a coherent representation and response. the aim of this research was thus to use mathematical and machine learning based analysis of eye tracking data to better understand the visual and cognitive behaviours associated with the completion of a graph task in a sample of children in year 3. using a desktop eye tracking system, children completed a mathematics task that involved comprehension of a bar graph. the scanpaths and the accuracy of individual responses to the task were analysed to identify the visual cognitive behaviours when completing the graph task, for example, what happens when children are confronted with such a task, which features children look at, and the order in which relevant information is accessed. the research also aimed to determine whether the gaze patterns for children who correctly or incorrectly completed the task were different, and if so, to identify the characteristics of the different visual cognitive behaviours. such information will enable more reliable inferences about the implicit visual cognitive behaviours of children when engaging in a graph problem-solving task and other cognitive tasks. 2. method 2.1 participants participants included 106 year 3 children (58 females, 48 males; mean age 8.8 ± 0.3 years) who completed a graphical mathematics task. children were from three primary schools in south east queensland, australia. data collection occurred in the last half of the school year, and the graph task formed part of a larger study that involved eye tracking while children completed a series of mathematics and reading tasks. all data collection occurred in a quiet room near the respective classrooms. the study was approved by university human research ethics committee that operates within the australian national statement on ethical conduct in human research. approval to conduct research in queensland state schools was also granted by the queensland government, department of education and training. 2.2 task the design of the graph task was based on the year 3 australian mathematics curriculum where children are interpreting and comparing data displays (acara, 2016). as presented in figure 1, the graph task included: a) a bar graph, where the height of each bar indicated the number of hours worked by sarah during a given week; b) a labelled coordinate system, where the x-axis had the week number labels, and the y-axis had numbers corresponding to hours; c) a sentence indicating sarah’s hourly wage; d) another sentence indicating the question related to sarah’s wages in week 3. 2.3 apparatus a screen-based tobii eye tracker (tx300) operating at 300 hz recorded the eye movements of the children as they completed the graph task. the task was presented on a 23 inch screen and participants sat comfortably (without restraint) at a working distance of approximately 60cm. the calibration used the tx300 nine point calibration procedure, with re-calibration conducted for any points where calibration was recorded as poor (denoted in red). only when the calibration procedure was completed for all nine points (denoted in green) was the eye tracking task started. events were detected using a dispersion based algorithm for detecting fixations. a fixation was defined as static eye movements with gaze positions remaining within a visual angle of 1.6° for at least 100 milliseconds (tobii technology, 2014). for the graph task, the 106 participants had a mean tracking percentage of 87.8 ± 9.8%. the output of the eye tracking device is typically a sequence of coordinate pairs relative to the scene that the participant is viewing, together with a time stamp when the locus of gaze is positioned at a given coordinate pair. 2.4 data extraction the visual stimulus (the graph on the screen; figure 1) was subdivided into regions known as areas of interest (aoi). initially, aoi sequences were generated from the sequence of fixations, as determined by the dispersion-based algorithm. the first step involved determining which were the most critical aois and labelling them as a1 (wage information (part of sentence below the graph)), a2 (week number (part of sentence below the graph)), a3 (week 3 bar), and a4 (number region containing the number of hours corresponding to week 3). these aois were considered as the minimum number of areas that needed to be fixated to complete the graph task successfully. the other aois were labelled as somewhat critical (b) and less critical (c). the categorisation of b areas as somewhat critical was derived as part of an iterative process, incorporating typical classroom practice and qualitative inspection of the scanpaths. typical classroom practice, or the procedural frameworks used to support children engaging with graphs, usually includes reading the title and axis labels as an initial orientation to the graph; these areas were therefore labelled as somewhat critical b areas. furthermore, qualitative inspection of scanpaths revealed that when reading the sentences below the graph, some children failed to read the full sentence and did not access the most critical information (a1 and a2), but did read the first part of each of the two sentences. to enable identification of this type of incomplete reading behaviour, the sentences were separated into two different areas, indicating somewhat critical (b) and most critical (a) information. finally, the less critical c areas were classified because they represented procedural checking that may be promoted as part of classroom practice, for example, systematically checking all bars on the bar graph before providing a response. figure 1. the graph task visual stimulus partitioned into areas of interest (aois). in this diagram, dashed lines indicate the original stimulus size and solid lines indicate the alignment of displacement zones for each aoi. the size of each aoi regardless of whether they were categorised as a, b or c, was calculated using two criteria; first, the aoi had to allow for vertical and horizontal displacement errors in the visual scanpath, and second, the aois could not overlap (holmqvist, et al., 2011). while some studies involving children (e.g. sasson & elison, 2012) have implemented a two degree vertical and horizontal displacement zone centred over relevant visual stimuli (one degree above and one degree below the stimulus), this was not possible with the current graph task as it would have resulted in substantial overlap of aois. thus wherever possible, a 1.9 degree vertical and 0.5 degree horizontal displacement zone was centred over the stimulus and in the event of overlap, the horizontal displacement zone was reduced to 0.15 degrees, which occurred for the following aois: a1 and b1, a2 and b2. to avoid overlap between c2 and b5 aois, b5 had a 1.5 degree centred vertical displacement, whereas the c2 vertical displacement zone could not be centred over the stimulus, so the c2 aoi was defined as 0.5 degree above and 0.95 degree below the stimulus. to distinguish fixations on the y-axis, a4 and b4 had a 1.9 degree vertical displacement, and a larger 1 degree horizontal displacement. the data sequences that were obtained initially were based on fixations and then transformed into dwell sequences by coalescing contiguous identical elements in the sequence. thus, for example, if in a fixation sequence we have b2, a1, a1, c3, this would give rise to the dwell sequence b2, a1, c3. in our analysis, sequences of dwells were used. the data included 106 data sequences or scanpaths (known as the inputs both terminologies are used to reflect what is typically used in the literature), one for each child, and the corresponding 106 answers to the graph question (the outputs), were categorised as correct (1) or incorrect (0). the data sequences (sequences of aoi names) could thus be grouped into two classes, one being the sequences of children who responded correctly to the task, and the other being the sequences of children who responded incorrectly. 2.5 machine learning techniques machine learning approaches were used to find a representative sequence of aois that would provide qualitative information about the visual cognitive characteristics associated with task performance. two methods are described that provide means by which an average scanpath can be extracted from a set of multiple sequences or scanpaths obtained from children performing the same task. method one applies a naïve bayesian approach to calculate the most probable vector for each class (using two possible alternatives to managing variable feature vector length), and method two applies a single iteration of a k-means approach to calculate the central vector for each class using an edit distance metric. the use of these techniques will form the basis for future research that characterises the relevant scanpaths of children to potentially guide identification of their ability to effectively perform graph tasks. 2.5.1 method one – most probable vector for each class the naïve bayes model or classifier “learns” from the data sequences, which are known as feature vectors, which is a commonly used term in the machine learning literature (murphy, 2012). a data sequence can be defined as a finite list of representational feature objects or values, which can be names, numbers or symbols, as long as the possible set of representational names or values is finite. in the case of scanpaths, the feature values are the aoi names. the classes in the model are the children who either correctly or incorrectly completed the graph task. learning a naïve bayes model consists of calculating the class conditional probabilities of the feature values. for each class and each feature, the number of occurrences of the feature value is divided by the number of feature vectors in the class (whether the child completed the task correctly or incorrectly) in question. this yields the class conditional probability for the possible values of a given feature in the given class. the “classifier” obtained from the model consists of probability distributions of the feature values determined by the class conditional probabilities for the values of each feature in each class. for example, suppose that we have feature vectors (each one in brackets) in a class: (a, b, c, a, b), (b, b, c, c, b), and (c, a, b, c, a). then, the class conditional probabilities for the second feature value (in bold) of this class are, for feature value a, 1/3; for feature value b, 2/3; and for feature value c, 0 because it does not appear in the second feature position. for a given sequence or feature vector in a given class, all the class conditional probabilities are multiplied and the product is multiplied by the class prior probability. in our analysis we assumed what is called a uniform prior, that is, a prior probability value of ½ for each class. the final products, one for each class, are the joint probabilities. to obtain the most probable feature vector for a class (i.e., the most probable or average scanpath or data sequence for the given class) the feature value that has the highest probability of occurring in a feature and class, as given by the class conditional probabilities, is selected for each feature and class. in the above example, the resulting vector would be (a, b, c, c, b). in the first position, all values have probability 1/3 so we pick one at random, say a. in the second through fifth positions, b, c, c, and b have probability 2/3 whereas the other symbols in each position have probability 1/3 or 0. the resulting list, or vector, of most probable feature values is the most probable feature vector (or average scanpath or data sequence) for that class. the vectors obtained for each class provide qualitative information regarding the children’s task performance according to the previously established criterion, which indicates whether the graph task was completed correctly or incorrectly. the feature vectors can have different lengths, which can lead to biases; there are two alternatives for managing variable feature vector lengths. the first way to deal with this length variability is to consider those feature vectors in a class which are shorter than the longest one in that class, as feature vectors containing missing data. the class conditional probabilities are calculated as outlined in the example. if a feature is considered as a coordinate or location in a vector, for a given feature and class, the probability of a feature value is the number of times that value occurs in any of the feature vectors, at the given location, divided by the total number of feature vectors in the given class. if a feature vector is too short to contain any feature values, then it will not contribute to the numerator of the preceding quantity, as in the example above. the alternative approach to handling the length variability is to turn all the feature vectors in a given class into vectors of the same length by “padding” the shorter vectors in the class with a dummy value to the right, up to the length of the longest feature vector in the class. padding a short vector to the right means that, after the last value of the short feature vector, a new padded value is added to the end of the list of feature values until it is the same length as the longest feature vector in the class. the “padding feature value” instances are different to all original feature values occurring in the data feature vectors. as an example, if the longest feature vector has a length of 5 (a, b, a, c, a), and we have a feature vector of a length of 3 (a, b, c), then, in positions 4 and 5 of the short feature vector we introduce the padding feature value, say x, so that instead of (a, b, c) we end up with (a, b, c, x, x). the class conditional probabilities are calculated as before, noting that in some cases, the most probable feature value might turn out to be the padding feature value. in the case of our analysis we used the first method to deal with length variability because, among other more technical reasons, it allowed us to have strings that contain only aoi names. 2.5.2. method two – the central vector for each class the second overall method of analysis involves taking a measure of the “distance” between one feature vector and another one with the levenshtein edit distance metric. the edit distance between two feature vectors is given as the number of feature deletions, insertions, and replacements that are required to transform one of the vectors into the other one. the method then works by finding a feature vector, which is called a central feature vector, that minimises the total edit distance between the central feature vector and each of the other feature vectors. methods to compute edit distance and thus a central feature vector are well understood in the computer science or mathematics literature (see skiena, 2010 for a review). these central feature vectors also reflect, on average, what is happening with all the sequences, and should be generally similar to those obtained with method one the most probable vector. 3. results the results of method one show that classifying multiple scanpaths, particularly, dwell area of interest (daoi) sequences with naïve bayes works extremely well (where a dwell is a continuous series of fixations within the same aoi). this in itself is informative, as it means that the elements of a daoi sequence are not necessarily correlated, that is, after looking at any one particular aoi, the observer may look at any other aoi with equal likelihood. both method one and method two allowed us to produce “virtual scanpaths” (the most probable vector and the central vector scanpaths) that demonstrate what children are doing when they are solving the problem correctly or incorrectly. 3.1 using a naïve bayes classifier to obtain most probable vectors using a naïve bayes classifier we were able to get a perfect classification, where our training error rate was zero. since predictive analyses was not our goal, we did not write our own cross validation routine. however, with a commercial package that handles missing data very differently (mathematica, wolfram research), we obtained a training error rate of .15 and a “leave one out” cross validation error rate of .27. this suggested that the aois in the gaze sequences are uncorrelated along the sequences and obey the naïve bayes assumption. what makes the use of a naïve bayes classifier attractive to us is the possibility of generating “virtual scanpaths” that are “most probable”. since for each feature (each coordinate of the scanpath sequences) and for each class we obtain a probability distribution of the possible aois, the most probable one can be selected. this provides us with an exemplary sequence that provides good qualitative data regarding what the children are doing while they solve the task, depending on whether they do so correctly or incorrectly. on average, the lengths of the scanpaths of children in either group are largely similar (69 ± 31 aois for children who responded correctly, vs 77 ± 39 for children who responded incorrectly), although they are slightly shorter for children who completed the task correctly. the qualitative information provided by the average scanpaths becomes even more evident with further manipulation and post-processing of the average scanpaths. if we merge the aois which are contiguous and identical in the average scanpath into a single instance, and remove dwells on blank spaces, the average scanpaths become much shorter, and this varies between the two classes. the correct children’s post processed (merged) average scanpath is as follows: {c1,b1,a1,b2,a2,b2,a2,a3,b2,a3,b2,a3,a1,a3,a1,a3,a1,a2,a1,a3,a1,a3,a1} and the incorrect children’s post processed (merged) average scanpath is as follows: {c2,c1,b1,a1,b1,a1,b1,b2,b1,b2,a1,b2,c2,b2,a1,c2,b2,a3,b2,c2,a3,a1,a3,a1,c1,c2,b2,a2,a3,b2,a1, b2,a1,a3,c2,b3,c2,b2,a3,a2} as can be observed, the incorrect children’s merged average scanpath is almost twice as long (40 aois) as that of correct children (23 aois). it is notable that a4 does not feature in the merged average scanpath for either the correct or incorrect children. this is likely to indicate that while children may have accessed a4 in their scanpath, there is not one clear position in the participants’ sequences where a4 features, therefore it does not appear in the average scanpath. the percentage of overall aoi accessed in the merged average scanpath are presented in figure 2. this shows that the gaze of the incorrect children wandered more than the gaze of correct children. it is also interesting to note that the incorrect children spread their attention relatively evenly across the a, b, and c areas, whereas the correct children exhibited greater visual attention (i.e. highest percentage of dwells) on the more critical a areas. figure 2. percentage of aoi access in the merged average scanpath, comparing correct and incorrect children. a aois: most critical, b aois: somewhat critical, c aois: less critical. examining the merged average scanpath as a sequence of dwells in areas a, b and c, demonstrates a more detailed characterisation of the visual cognitive behaviours (figure 3). correct children more rapidly accessed and maintained attention on the a areas. whereas children who were incorrect did not exhibit the same rate of attention on a areas, but rather appeared to be uncertain as to where to focus their visual attention, with many shifts between a, b and c areas with a larger number of dwells. figure 3. merged average scanpath, showing which aoi were accessed by correct and incorrect children over time. 3.2 finding the central data item in each class and what it reveals the average scanpath derived by method two (central data item) determined that correct children had the following sequence: {a1,b5,b1,b1,b1,a1,a1,a1,c2,a1,b2,b2,b2,b2,a2,b2,a2,a2,a1,b3,c2,c1,a3,a3,a1,a3,b3,a4,a3,a4,a3,a3,b2,a2,a1,a1} while the sequence for incorrect children was as follows: {c1,b1,b1,a1,a1,b2,b2,b2,b2,b2,b2,b2,a3,c1,c2,b3,a3,a4,c1,b5,b2,a2,b2,b2,b2,b2,a2,c3,c1,a4,c2,a3,b1, a1,a2,a1} the average scanpath sequences for correct and incorrect children have the same sequence length (36 values). in terms of percentage of dwells in the most critical areas, incorrect children, in contrast to the correct children, spent a lot of time looking at the less critical areas (b and c) (figure 4), which is similar to the results of method one (the most probable vector). figure 4. percentage of aoi access using the central feature vector, comparing correct and incorrect children. a aois: most critical, b aois: somewhat critical, c aois: less critical. the central feature vectors allowed us to learn how the critical areas were accessed. if we remove all of the non-critical (b and c) areas, leaving only the a areas, a distinction between the two vectors which is crucial to our understanding of the children’s task behavior becomes evident. following this step, we obtained the following correct sequence: or more simply, without the non-critical areas: correct response sequence: {a1,a1,a1,a1,a1,a2,a2,a2,a1,a3,a3,a1,a3,a4,a3,a4,a3,a3,a2,a1,a1} and, the incorrect response sequence: {a1,a1,a3,a3,a4,a2,a2,a4,a3,a1,a2,a1}. figure 5 summarises the sequence of dwells on most critical areas (a1-a4) and other non-critical areas (b-c) for the modified sequences. looking at the modified sequences, it is evident that area a4, which includes the values along the y-axis of the task graph, is infrequently visited by incorrect children, and, as anybody who learns math as a student or as a professional is aware, knowing what is being counted is critical to understanding and solving graph problems correctly. moreover, the correct children’s inspection of the most critical areas is in a predictable order (a1, a2, a3, a4). conversely, a less predictable order was evident in the modified sequence of the incorrect children. figure 5. modified average scanpath, showing the sequence in which the most critical aois (a1-a4) and other non-critical areas (b-c) were accessed by correct and incorrect children over time. discussion in this paper we discuss two novel methods of analysing eye tracking data to better understand the visual and cognitive behaviours associated with completion of a graph task in a sample of children in year 3. each method produced two strings (one for the correct children and one for the incorrect) which were compared. the initial findings are significant because they demonstrate how the methods developed by our research group can assist in characterising the visual behaviour of children who correctly or incorrectly answered the graph task. the results of both methods of analysis reveal differences in gaze patterns between the correct and incorrect groups. this discussion will summarise the key findings for each method, compare the different analytic approaches and characterise the visual cognitive behaviours identified. method one used a naïve bayes classifier to obtain most probable vectors. the perfect classification training rate with our software of 100% suggests that the eye tracking data for this task obeyed the naïve bayes assumption, and that the aois in the scanpath were uncorrelated, and an average scanpath can plausibly be determined. we believe that this is likely to be task dependent, however, further research is required in order to fully understand this issue. we did not do cross-validation with our software, but used a commercial package (mathematica, wolfram research) which showed a training error rate of .15 and a leave one out cross validation of .27, which given the small amount data and relatively large number of features, is still very good. we are optimistic that with more data, this rate would drop. with method one, comparisons of the average number of dwells in the sequences of each class identified a slightly higher number of dwells for the incorrect children, compared to the correct children (77 vs 69 aois, respectively). this higher number of dwells for the incorrect children is in accord with the findings of kim et al. (2014) who found that novices display significantly more shifts in visual attention than experts and have longer gaze sequences (or scanpaths) for a given problem-solving task, than their more experienced counterparts. in our study, the correct children exhibited fewer dwells and fewer shifts in visual attention than their counterparts who were incorrect. this is further evidenced in the merged average scanpath, where correct children had a higher percentage of dwells in the most critical a areas, whereas the incorrect children had dwells distributed across the most critical and less critical (a-c) areas. interestingly, these patterns reflect those of gegenfurtner, et al., (2011) and tsai, et al., (2012) regarding the behaviours of more experienced problem solvers who were shown to identify task relevant features of visual information more rapidly than less experienced individuals, and their visual attention focused more on relevant than irrelevant regions of the visual stimulus. our research was also able to characterise the sequence of rapid dwells on relevant/critical areas in the children who completed the graph task correctly, as shown in figure 3. method two used the central data item to determine the average scanpath, as well as a procedure to isolate the sequence of dwells on the most critical a areas. this isolation of the sequence of dwells on a areas was revealing in terms of understanding whether there was a logical sequence in accessing the most critical information. the children who responded correctly followed a logical sequence of dwells (based on our initial categorisation of which areas were most or least critical), that progresses from a1 to a2 to a3, then a4 (figure 5). conversely, the children who responded incorrectly, accessed a2 relatively late in the sequence. a2 represented the area surrounding the main question ‘how much did she earn in week 3?’ with a2 specifically representing ‘week 3’. delays in accessing this information may result in poorly focused attention to the relevant aspects of the graph, resulting in the child not knowing which bar they should be using in their calculation. relating this to curcio’s (2010) framework it is the first level in the data comprehension that is delayed, or out of sequence for the children who responded incorrectly. the initial understanding and reading of information in a logical sequence is the essential foundation to guide subsequent visual attention for interpreting and integrating the relevant graphical information and successfully completing the task. this finding is consistent with kotsopoulos and lee (2012), who reported that student problem-solving most often broke down in the early stages of understanding a mathematics problem-solving task. it may be that if initial understanding or orientation is not achieved early, the visual dwells become more frequent and variable in location, where children may be looking for information to help them understand the requirements of the task. it may also be that the positioning of the question, either above or below the graph, would impact on the visual behaviours and task response. in comparing the results of method one and method two, there are similarities and differences. both methods produce an average scanpath where the correct children have a larger percentage access of a areas, whereas incorrect children distribute their dwells across a, b and c areas, demonstrating that they may not be identifying the most critical areas. the average scanpath derived using the most probable vector indicated that children who provided correct responses had a slightly shorter average scanpath (69 aois) compared to the children who were incorrect (77 aois). this distinction was not evident in the average scanpath derived using the central data item, both groups had an average scanpath of 36 aois. a further difference between the two methods was associated with dwells on a1-a4 aois. the average scanpath (central vector) included dwells on all of the most critical areas (a1-a4) and implied a logical dwell sequence, whereas the average scanpath (most probable vector) did not include an a4 dwell. as noted in the results for method one (most probable vector), the omission of a4 from the average scanpath could be an indicator of variation in how children accessed this scale information, so it did not appear using the most probable vector method. for example, children may have counted up the gridlines within the body of the graph to determine the two hours worked in week 3. or alternatively, children may have looked more generally at the scale of the y-axis, but not necessarily specifically at a4. the source of this variation would need to be investigated with eye tracking data from another task. there are limitations to both methods and in the characterisation of the average scanpath for children who responded correctly and incorrectly. as with all averaging procedures, there is some loss of individual detail, however, the aim of this study was to characterise the visual cognitive behaviours of those children who completed the task correctly versus those children who were incorrect. method one, using a naïve bayes classifier to determine the most probable vector, might be limited by the data used in the analysis unless, as was the case in our study, the eye tracking data obeys the naïve bayes assumption of an uncorrelated aoi sequence. future research should replicate this task, or a close variant, in order to obtain test data to validate the models produced here. overall, this research has applied novel machine learning and mathematical analyses to characterise the visual cognitive behaviours of year 3 children engaged in a mathematics graph task. the resulting characterisations support the importance of initial understanding of the presented task and identification of the most critical information. while identification of the most critical information may be part of a typical classroom practice, this reinforcement, and the logical sequence of visual access of the most critical information, may be beneficial for children who feel less confident. keypoints machine learning techniques take the analysis of eye tracking data to a new level. in particular, they allow for understanding of phenomena in the dimension of time in relation to the order structure of data. spending time looking at critical areas and accessing critical areas in a logical order is important for completing the task successfully. spending too much time looking at less critical areas, and more dwells, is indicative of probable failure to successfully complete the task. for teaching purposes, it is important to identify critical areas and to break down the process of problem-solving into recognisable steps which can be carried out procedurally. acknowledgments this research was financially supported by the ian potter foundation (ref: 20140415). thank you to the schools, teachers, parents and children for their interest and involvement in this research. egme wishes to acknowledge the on-going discussions with nora mcintyre, at psychology in education research centre in the university of york, with whom valuable discussions ensued about the possible measures discussed here, and about classifying gaze patterns according to the subjects state of mind; and the ongoing support of prof. markku s. hannula at the faculty of educational sciences of the university of helsinki during this research. sw was supported by an australian research council decra (de160100830) during the preparation of this manuscript. references australiancurriculum, assessment and reporting authority (2016). the australian curriculum:mathematics,v8.2.retrievedfrom http://www.australiancurriculum.edu.au/ bochynska,a., &laeng,b. (2015).trackingdown thepathofmemory:eyescanpathsfacilitatetheretrievalofvisuospatialinformation. cognitiveprocessing,16(suppl. 1)159–163.doi:10.1007/s10339-015-0690-0 curcio,f.r.(2010).developingdata-graphcomprehensioningradesk-8 (3rdedition.).reston,va:thenationalcouncilofteachersofmathematics. gegenfurtner,a.,lehtinen,e.,&säljö,r.(2011). expertisedifferencesinthecomprehensionofvisualizations:ameta-analysisofeye-trackingresearchinprofessionaldomains. educationalpsychologyreview,23 (4),523-552.doi:10.1007/s10648-011-9174-7 holmqvist,k.,nystrom,m.,andersson,r.,dewhurst,r.,jarodzka,h.,&vandeweijer,j.(2011). eyetracking:acomprehensiveguidetomethodsandmeasures .newyork,ny:oxforduniversitypress. kim,s.,aleven,v.,&dey,a.k.(2014,april). understandingexpert-novicedifferencesingeometryproblemsolvingtasks:asensor-basedapproach .paperpresentedatthechi'14extendedabstractsonhumanfactorsincomputingsystems,toronto, ontario,canada.doi:10.1145/2559206.2581248 knoblich,g.,ohlsson,s.,&raney,g.e.(2001).aneyemovementstudyofinsightproblemsolving. memory&cognition,29(7),1000-1009.doi:10.3758/bf03195762 kotsopoulos,d.,&lee,j.(2012).anaturalisticstudyofexecutivefunctionandmathematicalproblem-solving. thejournalofmathematicalbehavior,31 ,196-208doi:10.1016/j.jmathb.2011.12.005 lai,m.l.,tsaim.-j.,yang,f.-y.,hsu,c.-y.,liut.-z.,lees.w.-y.,leem.-h.,chiou,g.-l.,liang,j.c.andtsaic.-c.(2013).areviewofusingeye-trackingtechnology in exploringlearningfrom2000 to2012.educationalresearchreview,10, 90-115.doi:10.1016/j.edurev.2013.10.001 murphy,k.p.(2012). machinelearning:aprobabilisticperspective,cambridge:mitpress. ratwani,r,m.,trafton,j.g.,&boehm-davis,d.a.(2008).thinkinggraphically:connectingvisionandcognitionduringgraphcomprehension. journalofexperimentalpsychology:applied,14 ,36-49.doi:10.1037/1076-898x.14.1.36 rinaldi,l.,brugger,p.,bockisch,c.j.,bertolini,g.,&girelli,l.(2015).keepinganeyeonserialorder:ocularmovementbindspaceandtime .cognition,142,291-298.doi:10.1016/j.cognition.2015.05.022 rosenzweig,c.,krawec,j.,&montague,m.(2011).metacognitivestrategyuseofeighth-gradestudentswithandwithoutlearningdisabilitiesduringmathematicalproblemsolving: athink-aloudanalysis.journaloflearningdisabilities,44 ,508-520.doi:10.1177/0022219410378445 sasson,n.j.,&elison,j.t.(2012). eyetrackingyoungchildrenwithautism. journalofvisualizedexperiments,61,e3675.doi:10.3791/3675 skiena,s.(2010).thealgorithmdesignmanual (2ndedition),newyork:springer sciencebusinessmedia tobii technology.(2014).usermanual:tobiitx300eyetracker ,revision2.sweden:tobiitechnology. tsai,m.-j.,hou,h.-t.,lai,m.-l.,liu,w.-y.,&yang,f.-y.(2012).visualattentionforsolvingmultiple-choicescienceproblem:aneye-trackinganalysis. computers&education,58 (1),375-385.doi:10.1016/j.compedu.2011.07.012 vangog,t.,paas,f.,van merriënboer,j.j.,&witte,p.(2005).uncoveringtheproblem-solvingprocess:cuedretrospectivereportingversusconcurrentandretrospectivereporting. journalofexperimentalpsychology:applied,11(4),237-244.doi:10.1037/ harteis et al publication frontline learning research vol.6 no. 3 (2018) 37 56 issn 2295-3159 do we betray errors beforehand? the use of eye tracking, automated face recognition and computer algorithms to analyse learning from errors christian harteisa, christoph fischera, torben töniges b, britta wrede b apaderborn university, germany btechnical faculty, bielefeld university, germany article received 14 may/ revised 25 september/ accepted 26 september/ available online 7 december abstract preventing humans from committing errors is a crucial aspect of man-machine interaction and systems of computer assistance. it is a basic implication that those systems need to recognise errors before they occur. this paper reports an exploratory study that utilises eye-tracking technology and automated face recognition in order to analyse test persons’ emotional reactions and cognitive load during a computer game and learning through trial and error. computer algorithms based on machine learning and big data were tested that identify particular patterns of test persons’ gaze behaviour and facial expressions that antecede errors in a computer game. the results show that emotions and learning from errors are positively correlated and that gaze behaviour and facial expressions inform about the errors that follow. however, the algorithms still need to be improved through further studies to be suitable for daily use. this research is innovative in its use of mathematical formulae to operationalise learning through errors and the use of computer algorithms to predict errors in human behaviour in trial-and-error situations. keywords: face recognition; eye tracking; emotions; learning from errors info: mail corresponding author christian.harteis@upb.de doi: https://doi.org/10.14786/flr.v6i3.370 1. introduction: research problem working life becomes increasingly complex and challenging, particularly through technological development and digitalisation that aim to enable flexible work processes. as a result, the organisation of work is changing, as well as – in consequence – working tasks and working tools. these become more difficult and the need for efficacy may generate time pressures. under these conditions, the risk of errors arises. estimations differ of the amount of worktime spent on errors in enterprises but are as high as half of the entire worktime (hofman & frese, 2011). on the one hand, systems engineering can strive to develop intelligent systems that prevent humans from committing errors. on the other hand, research into complex systems has revealed that it is not possible for human activity to avoid errors completely. hence, learning from errors becomes a relevant issue because it is at least possible to avoid the repetition of errors (harteis & bauer, 2014). there is no contradiction in simultaneously trying to develop systems that prevent errors (as far as possible) and postulating learning from errors, because both issues are interrelated. rather, understanding how to learn from errors is a precondition for developing man-machine interaction that assists in error avoidance. hence, the main research problem addressed here is how to understand learning from errors in order to provide safety through error prevention in man-machine interaction. theoretically, learning from errors has the following preconditions, none of which is trivial in the context of work (bauer & mulder, 2013; oser & spychiger, 2005): (a) the error has to be identified, (b) feedback to the acting person has to occur and (c) reflection and cause analyses have to result in the creation of negative knowledge – that is, knowledge about how things are not shaped and how processes do not work (gartmeier, bauer, gruber, & heid, 2008; oser, naepflin, hofer, & aerni, 2012). in addition, during these processes, the individual concernment of the failing person has to occur. “concernment refers to the emotional reaction, in which the error embarrasses the actor in a certain way. such an emotional reaction adds value to the experience of the error situation” (harteis & bauer, 2014, p. 710). this added value attaches sufficient importance to the error to initiate the cause analysis – subjectively unimportant events can easily be neglected – and adds authority to the knowledge resulting from reflecting on the cause analyses, which ultimately supports appropriate storage in the memory by amplifying the episodic memory and, thus, learning from error that prevents from its repetition (oser et al., 2012). however, while knowing that concernment resulting in emotional engagement is a crucial precondition for learning from errors, there is so far no evidence about the kind of valence of emotions (e.g. positive or negative) that best supports learning from errors. to sum up: any kind of emotional reaction in an error situation can be considered a basic precondition for learning from errors. it was oser who identified situations that almost – but not finally – ended up with errors as incidents of interest for their suggestion of a sense of failure (oser, müller, obex, volery, & shavelson, 2018; oser & obex, 2015). of course, in order to prevent errors, it is important not only to learn from errors but also from incidents in which errors nearly occur. in general, the crucial moment of learning from near misses is the emotional reaction that arises when somebody realises that an error is about to occur or that an error almost happened. cause analysis and reflection upon the incident play a similar role here as they do for learning from errors. however, investigations into this sense of failure revealed that emotional reactions that accompany (almost) error situations do not necessarily occur as a reaction to the incident itself but may arise shortly before the error occurs (oser & volery, 2012). whereas emotional reactions after an error tend toward embarrassment, cognitive load is considered to be the reason for an emotional reaction before an (almost) error situation (de jong, 2010). cognitive load refers to the working load within the limited capacity of the short-term memory (sweller, 1994). when cognitive load becomes too big, the actor’s capacity for information processing and problem perception decreases so that errors become probable. to investigate the research problem stated above, several issues have to be considered: emotional reactions and cognitive load are important phenomena in relation to the occurrence of errors or near misses and are important indicators that can help to prevent upcoming errors. the challenge, of course, lies in how best to operationalise and measure these phenomena. empirical research on learning from errors and near misses has to date applied self-reporting methods, that is, interviews and questionnaires. investigating error situations in work contexts is particularly challenging because companies tend to avoid publishing business processes. there are studies investigating employees’ attitudes towards errors at work (e.g. hetzner, gartmeier, heid, & gruber, 2011) which make use of standard self-report questionnaires (e.g. the error orientation questionnaire – rybowiak, garst, frese, & batinic, 1999) and there are studies investigating ways of dealing with error situations in daily working life which apply self-report questionnaires or interviews (e.g. harteis, bauer, & gruber, 2008). the current state of research thus has to acknowledge the following problems: studies operating with questionnaires focus either on general attitudes towards errors or they introduce a constructed error situation (e.g. through case stories or vignettes) and ask for potential reactions. neither option provides any insight into how a person actually behaves and reacts in a real error situation. in addition, whether the constructed situation causes a similar emotional engagement to real situations remains a matter for speculation. studies focusing on incidents that participants actually experienced usually ask test persons for episodes in which an error occurred and ask them to describe how the people concerned dealt with this situation. however, it is difficult to relate different cases described by different test persons to each other because the cases themselves represent error situations of different dimensions, because it is unclear how representative the described cases are for the test persons’ (work) environment and because the descriptions are probably subjectively biased. by their nature, self-reports feature only those mental and emotional processes that test persons are aware of and can remember. those studies therefore neglect whatever may remain unconscious or cannot be recalled. obviously, there is a research gap in the studies that investigate learning from errors, requiring a study that (a) on the one hand reliably provides stable and repeatable error situations for an entire sample, (b) does not depend on subjective biases and memory performances, and (c) is able to grasp unconscious emotional reactions. this study aims to test particular online measures of emotional reactions and cognitive load. 2. research questions and study context in order to reach the research aims, research questions were formulated that address theoretical issues and issues of online measurement. since field studies have to accept the problems described above, a laboratory setting appears appropriate to establish conditions that are identical for all test persons and exactly repeatable. admittedly, a laboratory environment lacks authenticity. however, it should be acceptable as long as test persons develop concernment when failing during the experiment. hence, the present study used a regular jump-and-run computer game (give up 2 by armor games) and controlled for test persons’ involvement. 2.1 research questions the research questions for this study can be separated into thematic (rq 1 and rq 2) and methodological (rq 3 and rq 4) questions. rq 1. are there emotional reactions and indicators for cognitive load to be found that precede errors? the hypothesis to be tested is that emotional reactions and/or cognitive load precede errors. an answer to this question is relevant for the intention to anticipate errors before their occurrence. rq 2. is the quality of learning related to emotional reactions? the hypothesis to be tested is that better learners show stronger emotional reactions than worse learners do. this question tests oser’s theory about the importance of concernment for learning from errors. rq 3. are there appropriate online measures that indicate emotional reactions? this study also aims to test particular online measures for their suitability for educational research questions. rq 4. is there a specific measure for the quality of learning from errors? a computer game with exactly repeatable conditions allows the combination of indicators for learning from errors and degree of difficulty. 2.2 description of the computer game give up 2 is a computer game in which the player operates a figure that needs to overcome obstacles and dangers on various levels of increasing difficulty. two kinds of error can occur in this game: (a) the player fails to overcome the obstacle with the figure or (b) the figure gets attacked by weapons and dies. in that case, the player has to start from the beginning of the respective level again. the course of action always remains the same, i.e. each player faces the same conditions. video 1. demonstration of the computer game this game permits the introduction of different test persons into comparable problem settings that provoke errors. it provides a competitive scenario that should motivate the test persons to perform as well as possible. 3. methods of data collection, analyses and challenges this main section of this paper describes the procedures of data collection, data preparation and data analysis. first, a flowchart illustrates the sequence of data processing. then, the system configuration and the sample description provide insight into the way the data collection was realised (3.1). the raw data were then prepared (3.2) for further analyses, exploring the predictors of errors (3.3) and learning from errors (3.4). the description of the analyses applied here also comprises a discussion of the challenges, because novel approaches were tested. figure 1 shows a flowchart of the procedures that were applied for this investigation. they comprise regular and well-established approaches to eye and face analysis utilising particular parameters and procedures that will be explained within this section. the flowchart presents the sequence showing how the data were prepared and analysed. figure 1. flow chart of data acquisition and processing. for eye analyses, pupil diameters were used to grasp cognitive load, and face analyses used standard tools (i.e. the openface toolkit and affectiva framework) that apply particular parameters to identify emotional reactions. these data were aggregated to temporal bindings per item, which were then used for machine learning. the result of the data-mining process is a trained classifier for the prediction of an error. 3.1 test procedure data collection was realised with the remote eye-tracking system smi red250, using the software versions experiment center 3.7.60 and begaze 3.7.42 and a video camera (logitech c922 pro stream) installed at the top of the stimuli screen within a laboratory, with a headrest and steady artificial lightning, i.e. robust laboratory conditions with no brightness differences between test persons (holmqvist et al., 2011). additionally, the stimulus itself did not vary a lot in luminesce. thirty-eight test persons with varying experience in computer gaming voluntarily took part in the experiment. table 1 sample description. before starting the experiment, the test persons filled in the consent form, received a video introduction to the game and were asked to confirm that they understood the game and the operation of the system. they also completed a questionnaire before and after the experiment describing their gaming experience and their engagement with the game measured by the personal involvement inventory (zaichkowsky, 1994). table 1 indicates that the test persons were sufficiently engaged. the test persons played the game for five minutes on the stimuli presentation screen using three keys of a keyboard, and they were observed by a video camera and remote eye-tracking system. the following in situ-data were generated and utilised: game video. via screen recording, a video of the game was generated, including all inputs of the test persons during the game. facial video. the video camera recorded the test person’s face during the game. pupil diameter. the remote eye-tracking system constantly recorded the test person’s right pupil diameter during the game. changes in pupil diameter apply as indicators of cognitive load (szulewski, kelton, & howes, 2017). 3.2 data preparation the synchronisation of these data was realised by time stamps implemented through the eye-tracking software. the challenge for subsequent analyses was to derive meaningful information from these data. therefore, they were further edited. as a first step, comprehensive annotations were added to the game video: errors. every event in which a test person failed (i.e. being hit by an object or failing to overcome an obstacle) was marked as an ‘error’. successes. every event in which a test person succeeded in avoiding an object or overcoming an obstacle was marked as a ‘success’. 3.3 analyses exploring predictors of errors (emotional reactions and cognitive load) the face videos were analysed by applying the affectiva framework (mcduff, el kaliouby, & picard, 2015) and the openface toolkit (baltrušaitis, robinson, & morency, 2016). based on millions of facial recording and facial images, these frameworks are able to extract crucial facial landmarks, such as eyebrows or mouth contours (see left side of video 2). these extracted points on each image are used to extract the head pose, gaze and facial action units (aus). the facial action coding system (facs, ekman & friesen, 1978) is a taxonomy for classifying various facial behaviours (e.g. au1: inner brow raise), and the frameworks used here were able to classify up to 17 of these aus (wolf, 2015). the affectiva framework also combines different aus to build up expression classes that are more abstract, such as emotional valence (i.e. the positive or negative nature of an emotion) and engagement (the expressiveness of the emotions). these frameworks, particularly in combination, provide derived data on crucial landmarks of the face representing emotional reactions as well as their valence and engagement. the next step of analysis aimed at identifying the precursors of errors. data-mining procedures were applied that utilised the derived data from the emotional reactions and a binary classification was modelled. all recorded videos were divided into snippets of 300 milliseconds in length with an overlap of 150 milliseconds. the positive class within the classification was modelled as ‘error predication’ and all snippets occurring before errors were marked. the negative class was modelled as ‘all the rest’ and all other data were assigned accordingly. to avoid noisy data, the negative class was filtered by removing those video snippets that followed an error event. the data were split into training (75%) and testing (25%) sets, each set containing the same percentage of positive and negative data. for each frame of the 300-millisecond snippets, the following 28 items were extracted by openface and affectiva: 17 aus i.e. facial expressions. valence. engagement. gaze angle x + gaze angle y. extracted in radians and averaged for both eyes. a person looking from the left to the right results in a change of gaze of angle x, while a person looking from up to down results in a change of gaze of angle y. if a person is looking straight ahead, both angles will be close to 0 (see figure 2). normalised pupil diameter. if the surrounding lighting conditions are steady, the pupil diameter (see figure 3) can be used as an indicator of cognitive workload: the wider the diameter, the higher the workload of the person (beatty, 1982; krejtz, duchowski, niedzielska, biele, & krejtz, 2018; laeng, sirois, & gredebäck, 2012; szulewski, kelton, & howe, 2017). head pose tx, ty, tz (location of the head). head rotation rx (pitch), ry (yaw), rz (roll) – see figure 4. figure 2. gaze angles. figure 3. pupil diameter. figure 4. head rotation. figure 5. example of a linear regression for items within 300ms frame. to take temporal dynamics into account, these items were further processed. for each of the 28 items, the temporal and stochastic variations were extracted. figure 5 shows an exemplary schematic representation: the red crosses represent items occurring within the respective timeframe. for each item, the following features were extracted: linear regression parameters (l0, l1) maximum value (max) minimum value (min) mean standard deviation (std) this results in 6 features per item, totalling 168 features, which were used as inputs for the machine training procedure. genetic programming (olson, urbanowicz, andrews, lavender, kidd & moore, 2016) was used to train an optimised machine-learning pipeline on the training set that would subsequently be used to evaluate the testing set. the pipeline consisted of multiple steps, such as feature selection, preprocessing, model selection and parameter optimisation. ultimately, the learned classifier could be used to classify unknown video snippets with the aim of identifying the precursors of errors. while this description of procedures describes the general ratio of analyses, combinations and sets of features were also varied in order to explore answers for the research questions raised above. video 2. demonstration of aggregated data. 3.4 analyses exploring learning from errors at this point, two kinds of derived data had been generated: errors and successes on the one hand, and emotional reactions on the other. as a third kind of derived data, the quality of learning from errors had to be identified. disciplines that investigate learning through a formalised lens have developed the idea of calculating a learning curve to indicate quality of learning (jaber, 2016; yelle, 1979). applied to the computer game, the learning curve (orange) can be defined by indicating the errors and successes (y-axis) of a specific task for all attempts (x-axis) at solving the task on the different game levels (separated by vertical lines; see figure 6). figure 6. example of a learning curve.   as a computer game requires a quite specific kind of learning of one particular skill (i.e. mastering the task), a formalised perspective for differentiating qualities of learning appears appropriate. an overall comparison of the ratio of successes and errors would provide a measurement of the overall performance of the player but would grant no insight into skill development. in order to grasp this, it is necessary to evaluate the performance progression over time. however, changes in performance indicated by changes in the slope of the learning curve can be illustrated by an angle (in figure 6: α and β). these angles indicate improvement if they have a positive value and a decline if they have a negative value. the design of the computer game provides a remarkable increase in playing difficulty at levels 2 and 6. each time the difficulty increases, the player has to readjust their behaviour and thereby improve their skill. hence, at levels 2 and 6, we can expect relatively more errors, compared to successes, than at levels 3, 4 and 5. the introduction of a new task problem provides the player with an opportunity to learn or improve their skill. the learning curve is thus steeper at levels 3, 4 and 5 than at levels 2 and 6. a possible approach to calculate learning quality contrasts the angles of the learning curve between levels 2 and 5 and between levels 5 and 6. level 5 represents the last easy playing level; the learning curve is considered to be steepest here. levels 2 and 6 represent difficult playing levels. looking at the absolute slope of the learning curve would be biased in favour of test persons who started with an already high skill level; however, it is not the intention to measure a test person’s absolute skill level but their skill development. hence, the slope at level 2 defines the baseline skill the test person shows after the first increase in difficulty. the following three levels of a similar difficulty provide the test persons with opportunities to improve on the lower level of difficulty. the difference between the slopes at levels 2 and 5 (angle α) represents the skill increase (or decrease) during this phase of the game. however, focusing on this difference alone would be biased in favour of test persons who performed very weakly at level 2 but who improved at level 5. as the change in skill is also relevant for difficult tasks, it is necessary to consider the development during level 6 – represented in the difference between the slope at levels 5 and 6 (angle β). a consideration of both angles corrects the bias of angle α towards a weak performance at level 2 combined with a good performance at level 5. hence, the derived data on learning quality considers the relative increase in successes between levels 2 and 5 in contrast to the relative increase in errors between levels 5 and 6. to make this measurable, the two angles α and β are calculated, where α is the angle between the slopes of levels 2 and 5 and β is the angle between the slopes of levels 5 and 6. the quality of learning can thus be defined by the following formula: learning = α + β hence, at this step of data preparation, the following derived data was available in order to answer the research questions: hperformance: errors and successes hemotional reactions: valence and engagement hcognitive load: changes in pupil diameter hquality of learning: learning = α + β the methodological challenges were to overcome the weaknesses of previous research on learning from errors as discussed in the section above. one of the major concerns raised there was that previous research does not inform about factual learning from errors. of course, a computer game provides quite a specific scenario for learning from errors. however, the major advantage is that it provides stable and repeatable conditions for all test persons. the gaming situation provides an opportunity to establish experimental conditions that appropriately motivate test persons to perform as well as possible while allowing them to commit errors without serious consequences. hence, the computer game provides an experimental scenario in which actual reactions to error can be observed. there is a further concern about the reliability and validity of data on learning from errors. in this experiment, particular online measurements were implemented to indicate learning from errors. it is a challenge, of course, to derive meaningful data from the data which are themselves extracts from raw data. the quality of these derived data will be discussed in later sections. 4. results the presentation of the results follows the sequence of research questions listed above. besides a t-test for group comparisons, the quality of the trained classifier resulting from data-mining and machine-learning procedures will be illustrated by using standard big data indicators, namely, receiver operating characteristic (roc) curves (fawcett, 2006) and confusion matrices (congalton, 1991). rq 1: are there emotional reactions and indicators for cognitive load to be found that precede errors? to answer this question, the previously described method was used to train a classifier on the whole training set features. the genetic programming reveals that the best result can be achieved with a gradient boosting classifier (friedman, 2001). the corresponding roc curve (see figure 7) shows the performance of the trained classifier. the roc curve visualises the diagnostic capability of a classifier: therefore, the true positive rate (sensitivity; i.e. an error is correctly predicted) is plotted against the false positive rate (probability of false error predictions) while varying the different thresholds of the classifier. the finally received operating characteristic of the trained classifier can be seen in figure 7. figure 7. roc curve of the optimal classifier. the dashed line represents the baseline of random guessing (i.e. the pure chance of a wrong or correct error prediction is 0.5). an optimal classifier (i.e. each prediction is correct) would receive an roc curve with an area of 1.0. the trained classifier here received a roc curve with area of 0.85. hence, with the current training set, it was possible to train a classifier able to precede errors sufficiently. changes in pupil diameters were considered as an indicator of cognitive load. the 300ms timeframes before errors and those before successes were examined and linear regressions for the changes in the test persons’ pupil diameters before errors and before successes were calculated for each test person individually. consequently, all test persons’ beta-coefficients can be put into a t-test for independent samples distinguishing errors and successes. table 2 presents the results of this t-test. table 2 two-sided t-test comparing changes in pupil diameters before errors and successes. in mean, changes in pupil diameter in error situations were significantly larger than changes in success situations. rq 2: is the quality of learning related to emotional reactions? since the quality of learning was operationalised through the formula developed above, a median split was applied to distinguish two groups of test persons: better learners and worse learners. for this calculation, only those n = 19 test persons could be considered who finished level 6 in the game and whose face recognition was successful. on this basis, a t-test reveals differences in emotional reactions (i.e. engagement and valence). table 3 two-sided t-test between better and worse learners. table 3 shows the results of the t-test that confirm the theoretical expectations: better learners show significantly higher emotional reactions in terms of engagement and valence. this means that better learners show emotional reactions of a stronger amount or intensity than worse learners do, and they also tend to a higher extent towards positive emotions than worse learners do. rq 3: are there appropriate online measures that indicate emotional reactions? to answer this research question, the features that were used for training the error classifier can be analysed in more detail. an importance ranking of all features was calculated, based on chi-square statistical analyses between each feature and class. a chi-square analysis test is able to measure dependence between stochastic variables. for classification, this test can be used to obtain a measure of dependence between the features and the two classes of the classifier. the features that are most likely to be independent of class receive a low score and the features that are most likely to be dependent of class receive a high score. the more dependent a feature of the class is, the better this particular feature is for use in classification – in our case, for predicting an error. figure 8 shows the 20 most important features for the prediction of errors. figure 8. ranking of most important features for the prediction of errors. it is remarkable here that only eye blink (au45), gaze (gaze angle x), cognitive load (pupil diameter), nose wrinkle (au09) and lip stretcher (au20) are represented in this top ranking. this means that those five items bear the majority of information required for the classification of errors. this calculation of the importance of singular features for the prediction of errors reveals that not only can emotional reactions be considered as relevant precursors of errors but also that pupil diameters can be interpreted as indicators of cognitive load (i.e. significant changes within 300 milliseconds before an error occurs). hence, a combination of pupil diameter features and facial video features contributes to the improvement of the classification of errors. in order to assess the increase in quality through this combination, the described machine pipeline was trained in two different ways: option 1 considers all 168 features and option 2 considers all features except the 6 pupil diameter features. figure 9 shows confusion matrices for both options (option 1, left side; option 2, right side). figure 9. confusion matrices. these confusion matrices cover a four-field table. the x-axis marks the prediction of an event (0 = no error prediction; 1 = error prediction) and the y-axis marks the actual outcome of an event (0 = no error; 1 = error). hence, the first and fourth quadrants indicate correct predictions, while the second and third indicate incorrect predictions. hence: the upper left quadrant refers to true negative predictions, the upper right quadrant to false positive ones, the lower left quadrant to false negative ones, and the lower right quadrant to true positive predictions. the figure also comprises the absolute number of cases (#) and the probability of correct/incorrect predictions. the comparison of both options reveals that the prediction of the first quadrant (i.e. the correct prediction of no error occurring) and the fourth quadrant (i.e. the correct prediction of an error) are slightly better if the pupil diameter features are also considered. in addition, the probability of detecting an error (prediction of no error occurring but error occurs – third quadrant) can be improved (from 38% down to 32%) if pupil diameters are used as well. hence, considering pupil diameter features in addition to facial expression features slightly improves the overall score and the practical usage of the entire classification. the better performance can also be seen in the roc analysis. the roc curve of the classifier using all features can be seen above (figure 7). the roc curve of the classifier excluding the pupil diameter features is shown in figure 10. figure 10. roc curve for classifier excluding pupil diameter features. the roc curve area for the classifier that excludes pupil diameter features is 0.80, whereas the roc curve area for the classifier considering all features is 0.85 (see figure 7). for the purpose of this study, this combination of indicators resulting from online measurement can be considered an appropriate measure of emotional reactions. it is important to emphasise that we did not strive to distinguish between different emotions but were simply interested in any kind of emotional reaction. rq 4: is there a specific measure for the quality of learning from errors? as the computer game provides similar tasks of increasing difficulty, it is plausible to assume that test persons improve by time, trials and errors. as a crucial measure of the change in performance (i.e. learning), the addition of angles α and β of the learning curve was chosen (see figure 6). table 3 reveals the descriptive statistics of the test persons’ performance. table 4 descriptive statistics for learning. the test persons varied substantially in their performances. in total, four test persons scored positively, ten test persons scored negatively and five test persons scored zero. for this particular setting, the learning curve indicates if and how an individual improves or develops during the course of a computer game of increasing difficulty. it indicates successful and unsuccessful attempts and thus can be considered a measure of the quality of learning from errors during the game. 5. critical reflection on data quality and validity the study comprises two variables: learning from errors and emotional reactions. the discussion and critical reflection on data quality and validity will focus on each variable separately. finally, a reflection follows on the relevance of data generated by a computer game for gaining insight into learning from errors. 5.1 data on learning from errors the construal of learning from errors follows quite a specific approach: the quality of learning – indicated through a learning score – was constructed as individual development along several trial-and-error attempts within a regular jump-and-run computer game. the scores on learning appear to tend towards the negative side (see negative mean within table 3), which requires a careful interpretation and must not be confused with a decrease in knowledge or negative learning. on the one hand, and as described above, the angles that result in the learning score were deliberately chosen because they refer to moments when the game’s difficulty increased substantially (at levels 2 and 6). it is part of a regular performance to fail initially with an increase of difficulty and then to adapt to the new learning. such regular performance results in a negative turn in the learning curve. on the other hand, only those test persons could be considered for the analyses who were able finally to complete this level of increased difficulty. this implies that even the test person with the lowest learning score in the sample was able to master this task of increased difficulty – albeit while making the highest number of errors within the sample. this measurement of learning faces two major limitations. first, test persons failing to master level 6 within the limited time could not be included in the measurement of the learning curve because the relation of successes and failures within level 6 could be determined only if a test person completed this level successfully within the given 5 minutes. such a time-based cut-off results in a flawed learning curve for the last level. future research attempts may allow as much time as a test person needs to cope with the increased difficulty. second, very good test persons who master all challenges without any failures show a learning score of 0, because the learning score reflects individual development. test persons who perform consistently well and thus do not show any difference in their error and success rates between the different levels do not develop in the sense of the measurement applied here. a learning score of 0 can be considered a ceiling effect because the task was not challenging enough for these test persons. hence, the data on learning from errors here indicate individual development during the run of a computer game which presumed that test persons fail occasionally during the run of the game. individuals who performed constantly at the same (high or low) level received a learning score of 0. the learning score does not therefore provide information about the quality of performance but about the quality of individual development; that is the crucial aspect of learning from errors in the context of this setting and the important focus of the online data used here. of course, the data would provide additional potential for focusing on learning from errors by directly connecting incidents when a test person initially fails and later succeeds during the game. however, since the focus of the study was on exploring opportunities to predict errors before they occur, the decision was made to include as many data as possible for the machine training. considering combinations of initial failures and subsequent successes would have decreased the number of observed cases dramatically. in addition, there would also be alternative data available that reflect a test person’s quality of performance (e.g. score, time, number of successes or failures), but such information does not tell us anything about learning from errors. 5.2 data on emotional reactions without doubt, the face is an important means of expressing emotional reaction. the recorded video data made it possible to identify a variety of aus that indicate emotional reactions. these indicators are based on analyses of fixpoints based on computer algorithms. in real face-to-face communications, the human mind is capable of processing the fine nuances of facial expressions unconsciously in order to interpret reactions appropriately. however, observational studies with human observers would probably not be able provide those kinds of data reliably. hence, the kind of facial analyses provided in this study can be considered to be an advance. this study – as already mentioned – did not aim at differentiating between different kinds of emotional reaction, however. this can be seen as a limitation, but for the context of the research questions raised here, this limitation does not have an impact on the data quality either for analysing learning from errors or for predicting errors because negative emotional reactions (e.g. fear) can limit human behaviour and situational perception in a similar way to positive reactions (e.g. euphoria). as the results reveal, the features considered here for developing a classifier for errors are sufficient to predict an error before it occurs – in the context of the video game that was part of the investigation. the choice of features resulted from an exploratory procedure of data mining that searched for relevant patterns in the training set and then confirmed the choice in the test set of the sample. it is difficult to judge if the resulting values of a 70% correct error prediction, 80% correct prediction of non-occurring errors and an roc curve of 0.85 are sufficient to justify applying these instruments in the technical context of man-machine-interaction. the acceptability of values probably depends on the range of application (e.g. high security areas or back-up systems). however, given that the classifier quite often predicted an error even though no error occurred (20%, in total > 4,500 cases, see figure 8) it would still seem inappropriate for application in real contexts. one explanation might be that test persons showed emotional reactions but still managed to avoid making an error in the video game. further machine-learning procedures may help to reduce this kind of wrong prediction. 5.3 relevance of data generated through a computer game learning from errors during a computer game may be seen as very different from learning from errors in real life, particular workplace settings. indeed, learning from errors at work occurs within an organisational error culture (putz, schilling, & kluge, 2012) that cannot be transferred to a laboratory setting. however, learning during work and learning during a computer game share important similarities: in both situations, learning occurs as a by-product of the intention to reach a goal. in both cases, there is no curriculum and no instruction that guides the acting but simply the intention to reach the goal successfully. learning from errors in real-life contexts varies widely across occasions and individuals. hence, it is difficult to identify the general characteristics of learning from errors empirically because situations are not comparable enough. a computer game, by contrast, provides stable conditions across all test persons and makes it possible to observe learning processes by looking at the tasks at hand. it should therefore provide enough insights into general processes related to learning from errors that can claim relevance to learning from errors in real-life contexts. 6. conclusions first, the findings reveal on the one hand that non-specified emotional reactions antecede test persons’ failures to overcome an obstacle or avoid an attack. second, the findings reveal that test persons who show a beneficial pattern of emotional reactions – that is, a higher extent of engagement and a positive valence – achieve higher learning scores than do test persons with an awkward pattern of emotional reactions. hence, they confirm oser’s theory about the importance of emotions for learning from errors (oser & spychiger, 2005; oser & volery, 2012). for the field of learning from errors in working life, this finding reveals the importance of the organisational culture (schein, 2004) and team climate in workplaces (edmondson, 1999). these require the appropriate social conditions that accept emotional reactions without generating disadvantages for the failing person. conditions that fail to provide such an environment tend to provoke the concealment, disregard and thus repetition of errors (marsick & watkins, 2003). on the other hand, given the extent to which the results fit into theoretical patterns, the findings also indicate that the tested way of gathering online data is a promising one. certainly, the procedures applied in this study have potential for improvement, as discussed above. as long as there are no repeat studies applying the same or similar measures, we do not know much about the validity of such measures. the experiences from this study suggest the need to repeat it under improved circumstances in two respects. first, a time limit is to be avoided; all test persons should receive as much time as they require to master level 6 of this game – as long as it appears reasonable to expect that each test person is able to cope with the difficulties of level 6 within a reasonable time. second, a repeat of this study should use a larger sample that would make it possible to connect initial failure with subsequent success directly in order to permit a focus on and analysis of concrete cases of learning from errors. in addition, a repeat of this study with a larger sample would provide an appropriate set of data to test the quality of the classifier found herein. keypoints emotional reactions and cognitive load precede errors. measures of automated face recognition generate data coherent with literature on emotions. the combination of face recognition and eye-tracking data can be used to predict errors before they occur. online measurements confirm keypoints of the theory of learning from errors. references baltrusaitis, t., robinson, p., & morency, l. p. (2016). openface. an open source facial behavior analysis toolkit. in wacv (ed.), 2016 ieee winter conference on applications of computer vision (pp. 1-10). lake placid: ieee. doi: 10.1109/wacv.2016.7477553 bauer, j., & mulder, r. h. (2013). engagement in learning after errors at work: enabling conditions and types of engagement. journal of education and work, 26(1), 99–119. beatty, j. (1982). task-evoked pupillary responses, processing load, and the structure of processing resources. psychological bulletin, 91(2), 276-292. congalton, r. g. (1991). a review of assessing the accuracy of classifications of remotely sensed data. remote sensing of environment,37(1), 35-46. de jong, t. (2010). cognitive load theory, educational research, and instructional design: some food for thought. instructional science , 38(2), 105-134. edmondson, a. (1999). psychological safety and learning behavior in work teams. administrative science quarterly,44(2), 350-383. ekman, p., & friesen, w. (1978). facial action coding system: a technique for the measurement of facial movements . sunnyvale: consulting psychologist press. fawcett, t. (2006). an introduction to roc analysis. pattern recognition letters,27(8), 861-874. friedman, j. h. (2001). greedy function approximation: a gradient boosting machine. the annals of statistics, 29(5), 1189-1232. gartmeier, m., bauer, j., gruber, h., & heid, h. (2008). negative knowledge: understanding professional learning and expertise. vocations and learning: studies in vocational and professional education, 1 (2), 87–103. harteis, c., & bauer, j. (2014). learning from errors at work. in s. billett, c. harteis & h. gruber (eds.), international handbook of research in professional and practice-based learning (pp. 699-732). dordrecht: springer academics. harteis, c., bauer, j., & gruber, h. (2008). the culture of learning from mistakes: how employees handle mistakes in everyday work. international journal of educational research, 47(4), 223–231. hetzner, s., gartmeier, m., heid, h., & gruber, h. (2011). error orientation and reflection at work. vocations and learning: studies in vocational and professional education , 4(1), 25-39. hofman, d. a., & frese, m. (2011). errors, error taxonomies, error prevention, and error management: laying the groundwork for discussing errors in organisation. in d. a. hofmann & m. frese (eds.), errors in organisations(pp. 1–43). london: routledge. holmqvist, k., nyström, m., andersson, r., dewhurst, r., jarodzka, h., & van de weijer, j. (2011). eye tracking. a comprehensive guide to methods and measures. oxford: oxford university press. jaber, m. y. (ed.). (2016). learning curves: theory, models, and applications. boca raton: crc press. krejtz, k., duchowski, a. t., niedzielska, a., biele, c., & krejtz, i. (2018). eye tracking cognitive load using pupil diameter and microsaccades with fixed gaze. plos one,13(9), e0203629. laeng, b., sirois, s., & gredebäck, g. (2012). pupillometry. a window to the preconscious? perspectives on psychological science, 7(1), 18-27. marsick, v. j., & watkins, k. e. (2003). demonstrating the value of an organization’s learning culture: the dimension of the learning organization questionnaire. advances in developing human resources, 5 (2), 132-151. mcduff, d., el kaliouby, r., & picard, r. w. (2015, september). crowdsourcing facial responses to online videos. in ieee (ed.), affective computing and intelligent interaction (acii), 2015 (pp. 512-518). piscataway township: ieee. olson r.s., urbanowicz r.j., andrews p.c., lavender n.a., kidd l.c., & moore j.h. (2016). automating biomedical data science through tree-based pipeline optimization. in g. squillero & p. burelli (eds.), applications of evolutionary computation. evoapplications 2016. lecture notes in computer science (pp. 123-137). cham: springer. oser, f., müller, s., obex, t., volery, t., & shavelson, r. j. (2018). rescue an enterprise from failure: an innovative assessment tool for simulated performance. in o. zlatkin-troitschanskaia, m. toepper, h. a., c. lautenbach & c. kuhn (eds.), assessment of learning outcomes in higher education(pp. 123-144). springer, cham. oser, f., näpflin, c., hofer, c., & aerni, p. (2012). towards a theory of negative knowledge (nk): almost-mistakes as drivers of episodic memory amplification. in j. bauer & c. harteis (eds.), human fallibility. the ambiguity of errors for work and learning (pp. 53–70). dordrecht: springer. oser, f., & obex, t. (2015). gains and losses of control: the construct “sense of failure” and the competence to “rescue an enterprise from failure”. empirical research in vocational education and training, 7(1), 3. oser, f., & spychiger, m. (2005). lernen ist schmerzhaft. beltz: weinheim. oser, f., & volery, t. (2012). "sense of failure" and "sense of success" among entrepreneurs: the identification and promotion of neglected twin entrepreneurial competencies . bern: skbf. putz, d.,schilling, j., & kluge, a. (2012). measuring organizational climate for learing from errors at work. in j. bauer & c. harteis (eds.), human fallibility(pp. 107-123). dordrecht: springer. rybowiak, v., garst, h., frese, m., & batinic, b. (1999). error orientation questionnaire (eoq): reliability, validity, and different language equivalence. journal of organizational behavior, 20, 527-547. schein, e. h. (2004). organizational culture and leadership. san francisco: jossey-bass. sweller, j. (1994). cognitive load theory, learning difficulty, and instructional design. learning and instruction, 4(4), 295-312. szulewski, a., kelton, d., & howes, d. (2017). pupillometry as tool to study expertise in medicine. frontline learning research, 5(3), 55-65. wolf, k. (2015). measuring facial expression of emotion. dialogues in clinical neuroscience,17(4), 457-462. yelle, l. e. (1979). the learning curve: historical review and comprehensive survey. decision sciences,10(2), 302-328. zaichkowsky, j. l. (1994). the personal involvement inventory: reduction, revision, and application to advertising. journal of advertising, 23(4), 59-70. baltrusaitis, t., robinson, p., & morency, l. p. (2016). openface. an open source facial behavior analysis toolkit. in wacv (ed.), 2016 ieee winter conference on applications of computer vision (pp. 1-10). lake placid: ieee. doi: 10.1109/wacv.2016.7477553 bauer, j., & mulder, r. h. (2013). engagement in learning after errors at work: enabling conditions and types of engagement. journal of education and work, 26(1), 99–119. beatty, j. (1982). task-evoked pupillary responses, processing load, and the structure of processing resources. psychological bulletin, 91(2), 276-292. congalton, r. g. (1991). a review of assessing the accuracy of classifications of remotely sensed data. remote sensing of environment,37(1), 35-46. de jong, t. (2010). cognitive load theory, educational research, and instructional design: some food for thought. instructional science , 38(2), 105-134. edmondson, a. (1999). psychological safety and learning behavior in work teams. administrative science quarterly,44(2), 350-383. ekman, p., & friesen, w. (1978). facial action coding system: a technique for the measurement of facial movements . sunnyvale: consulting psychologist press. fawcett, t. (2006). an introduction to roc analysis. pattern recognition letters,27(8), 861-874. friedman, j. h. (2001). greedy function approximation: a gradient boosting machine. the annals of statistics, 29(5), 1189-1232. gartmeier, m., bauer, j., gruber, h., & heid, h. (2008). negative knowledge: understanding professional learning and expertise. vocations and learning: studies in vocational and professional education, 1 (2), 87–103. harteis, c., & bauer, j. (2014). learning from errors at work. in s. billett, c. harteis & h. gruber (eds.), international handbook of research in professional and practice-based learning (pp. 699-732). dordrecht: springer academics. harteis, c., bauer, j., & gruber, h. (2008). the culture of learning from mistakes: how employees handle mistakes in everyday work. international journal of educational research, 47(4), 223–231. hetzner, s., gartmeier, m., heid, h., & gruber, h. (2011). error orientation and reflection at work. vocations and learning: studies in vocational and professional education , 4(1), 25-39. hofman, d. a., & frese, m. (2011). errors, error taxonomies, error prevention, and error management: laying the groundwork for discussing errors in organisation. in d. a. hofmann & m. frese (eds.), errors in organisations(pp. 1–43). london: routledge. holmqvist, k., nyström, m., andersson, r., dewhurst, r., jarodzka, h., & van de weijer, j. (2011). eye tracking. a comprehensive guide to methods and measures. oxford: oxford university press. jaber, m. y. (ed.). (2016). learning curves: theory, models, and applications. boca raton: crc press. krejtz, k., duchowski, a. t., niedzielska, a., biele, c., & krejtz, i. (2018). eye tracking cognitive load using pupil diameter and microsaccades with fixed gaze. plos one,13(9), e0203629. laeng, b., sirois, s., & gredebäck, g. (2012). pupillometry. a window to the preconscious? perspectives on psychological science, 7(1), 18-27. marsick, v. j., & watkins, k. e. (2003). demonstrating the value of an organization’s learning culture: the dimension of the learning organization questionnaire. advances in developing human resources, 5 (2), 132-151. mcduff, d., el kaliouby, r., & picard, r. w. (2015, september). crowdsourcing facial responses to online videos. in ieee (ed.), affective computing and intelligent interaction (acii), 2015 (pp. 512-518). piscataway township: ieee. olson r.s., urbanowicz r.j., andrews p.c., lavender n.a., kidd l.c., & moore j.h. (2016). automating biomedical data science through tree-based pipeline optimization. in g. squillero & p. burelli (eds.), applications of evolutionary computation. evoapplications 2016. lecture notes in computer science (pp. 123-137). cham: springer. oser, f., müller, s., obex, t., volery, t., & shavelson, r. j. (2018). rescue an enterprise from failure: an innovative assessment tool for simulated performance. in o. zlatkin-troitschanskaia, m. toepper, h. a., c. lautenbach & c. kuhn (eds.), assessment of learning outcomes in higher education(pp. 123-144). springer, cham. oser, f., näpflin, c., hofer, c., & aerni, p. (2012). towards a theory of negative knowledge (nk): almost-mistakes as drivers of episodic memory amplification. in j. bauer & c. harteis (eds.), human fallibility. the ambiguity of errors for work and learning (pp. 53–70). dordrecht: springer. oser, f., & obex, t. (2015). gains and losses of control: the construct “sense of failure” and the competence to “rescue an enterprise from failure”. empirical research in vocational education and training, 7(1), 3. oser, f., & spychiger, m. (2005). lernen ist schmerzhaft. beltz: weinheim. oser, f., & volery, t. (2012). "sense of failure" and "sense of success" among entrepreneurs: the identification and promotion of neglected twin entrepreneurial competencies . bern: skbf. putz, d.,schilling, j., & kluge, a. (2012). measuring organizational climate for learing from errors at work. in j. bauer & c. harteis (eds.), human fallibility(pp. 107-123). dordrecht: springer. rybowiak, v., garst, h., frese, m., & batinic, b. (1999). error orientation questionnaire (eoq): reliability, validity, and different language equivalence. journal of organizational behavior, 20, 527-547. schein, e. h. (2004). organizational culture and leadership. san francisco: jossey-bass. sweller, j. (1994). cognitive load theory, learning difficulty, and instructional design. learning and instruction, 4(4), 295-312. szulewski, a., kelton, d., & howes, d. (2017). pupillometry as tool to study expertise in medicine. frontline learning research, 5(3), 55-65. wolf, k. (2015). measuring facial expression of emotion. dialogues in clinical neuroscience,17(4), 457-462. yelle, l. e. (1979). the learning curve: historical review and comprehensive survey. decision sciences,10(2), 302-328. zaichkowsky, j. l. (1994). the personal involvement inventory: reduction, revision, and application to advertising. journal of advertising, 23(4), 59-70. heitzman et al frontline learning research vol. 7 no 4 (2019) 1 24 issn 2295-3159 facilitating diagnostic competences in simulations: a conceptual framework and a research agenda for medical and teacher education nicole heitzmann1, tina seidel2, ansgar opitz1, andreas hetmanek2, christof wecker3, martin fischer4, stefan ufer1, ralf schmidmaier4, birgit neuhaus1, matthias siebeck4, kathleen stürmer5, andreas obersteiner6, kristina reiss2, raimund girwidz1 & frank fischer1 1ludwig-maximilians-universität münchen, germany 2technische universität münchen, germany 3universität hildesheim, germany 4 university hospital, ludwig-maximilians-universität münchen, germany 5 universität tübingen, germany 6 pädagogische hochschule freiburg, germany article received 10 june 2018 / revised 7 may / accepted 30 august / available online 22 october abstract we propose a conceptual framework which may guide research on fostering diagnostic competences in simulations in higher education. we first review and link research perspectives on the components and the development of diagnostic competences, taken from medical and teacher education. applying conceptual knowledge in diagnostic activities is considered necessary for developing diagnostic competences in both fields. simulations are considered promising in providing opportunities for knowledge application when real experience is overwhelming or not feasible for ethical, organizational or economic reasons. to help learners benefit from simulations, we then propose a systematic investigation of different types of instructional support in such simulations. we particularly focus on different forms of scaffolding during problem-solving and on the possibly complementary roles of the direct presentation of information in these kinds of environments. two sets of possibly moderating factors, individual learning prerequisites (such as executive functions) or epistemic emotions and contextual factors (such as the nature of the diagnostic situation or the domain) are viewed as groups of potential moderators of the instructional effects. finally, we outline an interdisciplinary research agenda concerning the instructional design of simulations for advancing diagnostic competences in medical and teacher education keywords: diagnostic competence; simulation; medical education; instructional support; learning through problem-solving; scaffolding; teacher education info corresponding author: nicole.heitzmann@psy.lmu.de doi: 10.14786/flr.v7i4.384 1. introduction diagnostic competences are important goals of many academic programmes, including medicine and teacher education programmes. recently, the scientific understanding of the structure of diagnostic competences as well as their measurement have improved substantially (e.g., abs, 2007; heitzmann, 2014; herppich et al., 2018; tariq & ali, 2013). however, with a better understanding of diagnosing and diagnostic competences, it became apparent how even advanced students and young professionals struggle to apply their knowledge to individual cases. this issue led to an emphasis on the question of how we can design and provide learning opportunities where higher education students can learn to diagnose in their fields. there is a need for more or less authentic situations in which conceptual knowledge can be acquired and applied to problems. however, providing real practice opportunities may not always be the most effective approach, because highly complex and dynamically changing real-life situations can be overwhelming for novice learners (grossman et al., 2009). therefore, approximations to real-life practice using simulations may sometimes be more effective in a higher education programme than just increasing the opportunity for real-life practice (e.g., stegmann, pilz, siebeck, & fischer, 2012). in this manuscript, we address how simulations can be designed to allow learners to make the most of a simulation. design features vary broadly across simulations in different studies. for instance, simulations differ in whether reflection phases are included or whether additional information can be accessed during the simulation. so far, we do not know enough about the instructional design of effective simulations, that is, which forms of instructional support (e.g., forms of scaffolding, explicit presentation of information) are linked to better learning processes or learning outcomes. there is also a broad scope of diagnostic situations which can be simulated. the simulation can be targeted to support learning to diagnose in the context of a patient interview or in direct interaction with one or several students in class. simulations sometimes afford interaction with documents and materials rather than with people, as for example in diagnosing tentative causes of a patient’s fever from an x ray picture or diagnosing the causes of a failure to solve specific tasks in mathematics by analysing the method of calculation a student has used. the cognitive requirement in diagnosing might be quite different depending on these as well as other specific features of a situation. thus, the question of the extent to which the effects of instructional designs generalize across different types of simulated diagnostic situations is relevant. the effectiveness of instructional support is also known to depend on interindividual differences. most obviously, this is the case with different levels of prior knowledge, where maladjusted instruction has been shown to be not only ineffective but sometimes even detrimental to learning (kalyuga, 2007). other individual learning prerequisites include motivation, emotion or general cognitive abilities. what do we know about how interindividual differences in these variables influence the advancement of diagnostic competences in simulations? and what should we know in order to be able to design simulation-based learning environments which cater to the needs and preferences of individual learners? however, before we can begin to address these design features and questions, we need a clear conceptualization of the competences that we want to foster with simulations. this will form the basis for further design steps. therefore, the first goal of this article is to identify and discuss important components and relationships between components regarding diagnostic activities in the fields of both teacher and medical education. based on a discussion of commonalities and differences between the fields, we propose a conceptual framework and sketch a research agenda addressing the facilitation of diagnostic competences in simulation-based learning environments. we first propose a conceptual framework, as displayed in fig. 1, outlining an approach to fostering diagnostic competences with simulations which will guide our work on the aforementioned design features and questions. the main body of the article comprises an elaboration of the framework’s components. first, we address the two components of diagnostic competences in the framework: the knowledge which learners need to possess and the processes they need to master. second, we elaborate on learning through simulations and on ways to support this process in order to maximize learning gains. third, we introduce a range of relevant learning prerequisites and contextual factors. in all three parts, we selectively review the research in medical and teacher education and explain how this research ties into the construction of the framework. in the second part of the article, we briefly propose sets of hypotheses and a possible methodological approach for a research programme based on the conceptual framework. the ultimate goal of the research programme is to advance the conceptual framework into a theory of learning to diagnose in simulations that can inform the design of simulation-based learning environments in medical and teacher education and possibly beyond. 2. arguments for analysing teachers’ diagnostic competences and relating them to medical doctors’ diagnostic competences in this section, we will first provide a rationale for why diagnostic competences are relevant not only in the medical context but also for teachers. we propose that despite differences in the professional practice in both fields, there are also substantial commonalities. we then explicate the premises for a conceptual framework. diagnosing means “recognizing exactly” or “differentiating” and is associated with the activities and processes of classifying causes and forms of phenomena (“diagnosing”, n.d.). these causes and forms are often not directly observable; they are latent or hidden and need to be identified through observable cues and by drawing inferences. typically, diagnoses are needed as decision points for action and for interventions targeted at improving a problem. professions develop, maintain and teach diagnostic activities and procedures which are considered reliable and trustworthy, and they also develop quality standards for diagnosing and for diagnoses. in this article, we consider two professions. one is a well-known example with respect to diagnosing, namely medicine. the other, teacher education, is less well-known for diagnosing, but research has increasingly referred to diagnosing as a relevant professional activity. in both medicine and education, diagnostic activities are considered components of professional problem-solving. in education, diagnosing is often regarded as a crucial precondition and an element of adaptive teaching, for example when a teacher assesses the performance of a student (schrader, 2009, 2011). some researchers in the field of teacher education have been critical about the use of the term “teacher diagnostics”, particularly regarding the possibly negative effects of wrongly labelling students in the school context (borko, roberts, & shavelson, 2008). also, in the medical field, researchers are considering the possible negative side effects of diagnoses, which can sometimes lead patients onto destructive paths (e.g. croft et al., 2015). at times, terms such as “teacher assessment”, “teacher judgment” and “teacher decision-making” have been used instead in order to find more nuanced ways to describe the tasks and processes involved in continuously gaining and evaluating knowledge about students. however, these alternative conceptualizations continue to include teacher diagnostic activities (herppich et al., 2018), and many current research approaches now use the term “teachers’ diagnostic competences” (glogger-frey, herppich, & seidel, 2018). of course, one may ask whether these professions and research strands share commonalities which go beyond terminological similarities. the link between medicine and teacher education seems risky at first glance. there are many salient differences between the professional practice of teachers and that of medical doctors, for example in terms of group size and the duration of diagnostic situations, relevant knowledge bases and the standardization of processes. we propose in this paper that despite these differences, there are also substantial commonalities between the two professions with respect to diagnosing, which warrant a joint conceptual framework as well as joint research programmes building on the framework. this proposition builds on the following premises. first, an underlying argument is based on the idea of knowledge communities. humans try to improve ways of generating knowledge about their world to achieve, among other things, a higher accuracy in diagnoses of problems and, ultimately, more adequate follow-up actions. in this process, humans are capable of overcoming their intuitive epistemic mechanisms by developing more systematic knowledge creation activities (e.g., stanovich, 2011). knowledge communities maintain, share, and teach the competences needed to engage in these more systematic knowledge creation activities. they also contribute to developing quality criteria like standards for diagnosis procedures, reporting diagnoses and discussing outcomes of the diagnostic process with other people. professions can be considered such knowledge communities. at this point, one might object to our comparison of professions by arguing that diagnostic processes in the medical profession show a higher degree of standardization compared to those in the educational domain. however, we do not think this difference makes our comparison invalid. we assume that the field of education would benefit from a higher degree of standardization in some diagnostic situations, because standardization could lead to more reliable and valid diagnostic results than intuitive approaches. however, medicine entails common diagnostic situations where highly standardized and algorithmic diagnostic activities are just not feasible. an example would be primary care, when doctors need to diagnose and treat many patients in short time frames. we thus assume that standardization is a relevant dimension for both professions, doctors and teachers, but a higher degree of standardization cannot generally be considered a sign of more advanced professionalization. second, it seems plausible that effects of instructional interventions may rather generalize to cognitively similar diagnostic situations across the domains than to cognitively dissimilar situations within one domain (see kirschner, verschaffel, star, & van dooren, 2017). consider, for instance, instructional guidance for learning to interpret evidence from several different sources when a teacher tries to determine a student’s level of mathematical understanding or misconceptions leading to his mistakes, compared with a general medical practitioner trying to determine the causes of a patient’s symptoms by reading the patient’s file, conducting a physical examination and using newly generated laboratory parameters. it seems plausible that both medical students and pre-service teachers can benefit from guidance for generating a set of early hypotheses which help them prioritize evidence and search for and select necessary evidence in a goal-oriented manner. however, this type of support might be less helpful in the contexts of collaborative diagnosing, for instance, between an internist and a radiologist. in the latter context, medical students’ and young doctors’ critical difficulties in sharing relevant information based on their meta-knowledge of the other specialists’ knowledge, tasks and instruments were identified (tschan et al., 2009). an effective instructional intervention for information sharing might be quite different from instruction in support of early hypothesis generation. third, there are also some striking commonalities between teachers’ and physicians’ professional practice. they both engage in decision-making related to valued characteristics in other people (that is, health and education), they both must integrate various different knowledge bases and they both must master diagnostic activities like generating hypotheses or drawing conclusions (fischer et al., 2014). and in both domains, higher education programmes increasingly aim at advancing their students’ competences through hands-on experiences and reflections on practical experience (e.g., grossman & mcdonald, 2008). fourth, the research strands from these two domains also have interesting commonalities. most importantly, they consider diagnosing as part of professional action related to cases. the diagnostician uses observations and data collected with a specific goal in mind. a diagnosis, then, is based on data and is the result of a systematic and reflective process (helmke, 2010). understood this way, diagnosing is not restricted to medical decision-making in the context of patients’ diseases; instead, diagnosing can be considered more generally as the goal-oriented collection and interpretation of case-specific or problem-specific information to reduce uncertainty in order to make medical or educational decisions. after explicating why we think it makes sense to analyse the diagnostic competences of medical practitioners and teachers with a joint conceptual framework, we now want to introduce this conceptual framework in more detail. 3. conceptual framework in our conceptual framework (fig. 1), we consider diagnosing as a process of goal-oriented collection and integration of case-specific information to reduce uncertainty in order to make medical or educational decisions. figure 1. fostering diagnostic competences with simulations: a conceptual framework of factors. different knowledge bases involved in diagnosing exist. the conceptual framework builds on a recent attempt to integrate the different types of classification into a two-dimensional classification. this classification distinguishes knowledge facets (content knowledge, pedagogical content knowledge and pedagogical knowledge) and knowledge types (knowing that, knowing how and knowing why and when) (förtsch et al., 2018). the conceptual framework further entails diagnostic activities like generating hypotheses or evaluating evidence. diagnostic competences are defined as individual dispositions enabling people to apply their knowledge in diagnostic activities according to professional standards to collect and interpret data in order to make high-quality decisions. domains differ in what is considered an appropriate engagement in these diagnostic activities and to what degree these activities are formalized and standardized. in the following sections, each component is explained in more detail. first, we will have a closer look at the professional knowledge bases and at diagnostic activities. 3.1. diagnostic competences: professional knowledge and diagnostic activities – two strands of research in medical and teacher education in both domains, medicine and teaching in primary and secondary schools, there is (1) research focusing predominantly on the types and combinations of knowledge needed for successful diagnoses, for example conceptual knowledge and strategic knowledge. in both domains, there is also (2) research focusing on the process of diagnosing and the different activities diagnosticians engage in to derive accurate and well-justified diagnoses, for example generating hypotheses or evaluating evidence. these two strands of research will be reviewed in this section. 3.1.1. research focusing on the professional knowledge base for diagnosing a common differentiation of the professional knowledge base in medicine suggests two knowledge types: (a) biomedical knowledge and (b) clinical knowledge (patel, evans, & groen, 1989). (a) biomedical knowledge includes knowledge about the normal functioning of the human body as well as pathological mechanisms or processes in the causation of diseases (boshuizen & schmidt, 1992; kaufman, yoskowitz, & patel, 2008). (b) clinical knowledge is knowledge about symptoms or symptom patterns of diseases, their typical courses and factors indicating a high likelihood of a particular disease, such as patient characteristics or environmental factors. it also comprises knowledge about appropriate therapeutic treatments (van de wiel, boshuizen, & schmidt, 2000). research in medical expertise and in medical education has focused on how these knowledge types are used when diagnosing patient cases. one main interest has been how novices, intermediates and experts differ in their application of these types of knowledge. the main goal of this comparison was to understand the processes and changes involved in expertise development. cognitive psychologists working in the context of medical expertise and medical education have suggested models which advance the mere classification of knowledge according to content. for example, stark, kopp and fischer (2011) distinguished conceptual knowledge from practical knowledge. conceptual knowledge concerns concepts and their interrelations, whereas practical knowledge describes how conceptual knowledge is used in problem-solving situations. practical knowledge is further divided into the knowledge of steps which can be taken in problem-solving (strategic knowledge) and knowledge about the conditions of the successful application of these steps (conditional knowledge). the distinction between conceptual and practical knowledge has received validating support from empirical studies inside and outside the medical domain (see heitzmann, 2014, for an overview). research in medical education also offers specific accounts of the changes that the knowledge bases undergo in the course of developing expertise. the most prominent change is that of encapsulation into illness scripts, as developed by boshuizen and colleagues (1992). this account proposes the qualitative entanglement of biomedical and clinical knowledge as a key to better understanding the development of diagnostic expertise (e.g., boshuizen & schmidt, 2008; schmidt & boshuizen, 1993; woods, 2007). through repeated confrontation with clinical cases, biomedical knowledge gets “encapsulated” (mamede et al., 2012). that these two types of knowledge are encapsulated means that biomedical knowledge gets interconnected and integrated with clinical features. learners connect the knowledge of underlying biomedical mechanisms with symptoms of a disease, with patient characteristics and with the conditions under which a certain disease emerges. this encapsulation yields so-called “illness scripts” (charlin, boshuizen, custers, & feltovich, 2007; schmidt & rikers, 2007). biomedical knowledge in an encapsulated form is still important for a coherent understanding of a disease, even though it might not reach conscious attention (woods, 2007). an illness script also contains knowledge of the relations between different diseases as well as of cases of a disease the physician has previously encountered (schmidt & rikers, 2007). shulman (1987) has proposed a differentiation of three components of the professional knowledge base: (a) content knowledge about the subject matter being taught, such as knowledge of mathematics or biology, (b) pedagogical content knowledge about the teaching and learning in a subject area, for example typical misconceptions in physics or how to explain certain grammatical structures, and (c) pedagogical knowledge about teaching and learning independent of the subject area, including for example knowledge about classroom management, student motivation or learning strategies. current research addresses the roles of these knowledge types (tröbst et al., 2018) in professional action, including diagnosis. important open questions include how the different components of professional knowledge interact (e.g., tröbst et al., 2018) and possibly change through experience and how different knowledge components are restructured and possibly integrated (“amalgamation”, tröbst et al., 2018) through knowledge application in decision-making and problem-solving contexts during the development of teaching expertise. stürmer, seidel, & holzberger (2016) have recently hypothesized that knowledge transformation mechanisms in teacher expertise development are similar to those hypothesized for medical expertise development. in particular, these mechanisms can be characterized as a systematic restructuring rather than a mere accumulation of knowledge. as a consequence of this restructuring, performance increases while the consciously available knowledge decreases, similar to the process of encapsulation hypothesized in research on clinical reasoning. a cross-domain foundation for empirical research on how to support the development of diagnostic competences requires a cross-domain conception of professional knowledge. while the distinction of conceptual and practical (strategic and conditional) knowledge appears applicable to the professional knowledge of teaching as well, the question whether the distinction of content knowledge, pedagogical content knowledge, and pedagogical knowledge has a counterpart in medicine is currently under discussion. recently, there have been attempts to integrate the knowledge classifications across domains (förtsch et al., 2018; hargreaves, 2000). förtsch et al. (2018) proposed a two-dimensional model integrating what they call facets of knowledge (content knowledge, pedagogical content knowledge and pedagogical knowledge) and types of knowledge (knowing that, knowing how and knowing why and when). the integration of pedagogical knowledge and pedagogical content knowledge in models of medical diagnosing particularly challenges established viewpoints. for instance, förtsch et al. (2018) emphasize the importance of scrutinizing and communicating diagnostic processes and outcomes with colleagues as well as with patients and their families as an example of the need for this integration of pedagogical knowledge and pedagogical content knowledge into an overall model. 3.1.2. research with a focus on diagnostic activities knowledge is applied in diagnostic activities in order to produce a diagnosis. diagnostic activities can be characterized in different ways (e.g., barrows & pickell, 1992; gräsel, 1997; moskowitz, kuipers, & kassirer, 1988). there are repertories of professional activities in which the different bodies of knowledge can be applied. teacher education has emphasized teachers’ assessments of their students’ knowledge and skills. this research has focused on the accuracy of teachers’ judgments of students’ learning prerequisites and outcomes in primary and secondary school (klug, bruder, kelava, spiel, & schmitz, 2013; spinath, 2005; südkamp, kaiser, & möller, 2012) as well as on the role of teachers’ diagnoses in formative assessment (e.g., glogger-frey et al., 2018; bennett, 2011; hattie, 2003). however, diagnostic activities in teaching have not yet been systematically investigated (e.g., herppich et al., 2018; shulman, 2015). we introduce a taxonomy of epistemic activities which was developed in a multidisciplinary group of scientists from different fields (biology, mathematics, medicine, psychology and computer sciences, fischer et al., 2014). the scientists agreed that the activities in the taxonomy were all relevant for generating knowledge in their domains. the taxonomy differentiates eight activities which we illustrate with examples from medicine and teaching: (1) problem identification (e.g. a patient reports non-specific symptoms such as shortness of breath; a student wrongly answers a question in class); (2) questioning (e.g. a doctor asks what could be the reason for the symptoms; a teacher asks what could be the reason for a student’s error); (3) hypothesis generation (e.g. a doctor suspects a specific disease, such as a pulmonary embolism; the teacher suspects a specific misconception); (4) the construction and redesign of artefacts (e.g. a medical report which prescribes further examination; the development of a task which provides insight into the presence of a misconception); (5) evidence generation (e.g. conducting required further examination, for example through computed tomography; observation of the student’s solution of the task); (6) evidence evaluation (e.g. evaluation of the computed tomography with radiological signs of a pulmonary embolism; evaluation of the solution of the task with some but not all of the signs for the hypothesized misconception); (7) drawing conclusions (e.g. deciding that the most likely cause of the patient’s symptoms is a pulmonary embolism; deciding that the most likely reason for the student’s error is the assumed misconception, which impedes further learning); and (8) communication and scrutinization (e.g. a medical report with the diagnosis of a pulmonary embolism for another doctor; informing another teacher about the discovered misconception held by a certain student so that teacher can adapt the teaching). as one can see, the eight epistemic activities can be translated into eight diagnostic activities, because diagnosing is a form of knowledge generation. diagnosing may not always require all eight epistemic-diagnostic activities, and no generally valid order is assumed for these eight activities. rather, the number of steps and their order depend on the specificities of the case, the situation and the expertise of the individual diagnostician, including information available from previous diagnostic processes. doctors or teachers may skip activities (probably because of preceding knowledge encapsulation; boshuizen, schmidt, custers, & van de wiel, 1995) or may repeat activities in a circular fashion. the number of activities or their order may, therefore, not be reliable indicators of the quality of the diagnostic process. rather, the quality of the single activities and their combinations in sequences may be better indicators. the quality criteria applied to an activity and to sequences of activities are not general but are typically implied by professional standards in a specific domain. it became clear early in research on diagnostic processes and strategies in medicine that diagnostic activities, such as the eight activities elaborated above, do not sufficiently characterize the requirements for appropriate cognitive performance. research on the development of expertise in medicine indicated that inexperienced doctors seem to engage in most or all of these diagnostic activities, whereas experts seem to skip many steps and arrive at their conclusions more quickly and intuitively, a finding which can be explained by the encapsulation process introduced above. it should be noted that in diagnostic situations with an emphasis on generating hypotheses, other diagnostic activities than the eight outlined above, which focus on confirming hypotheses, can come into play. in these hypothesis-generating situations, the focus is on filtering ongoing information regarding its relevance for the diagnostic task. then a chain occurs of inferences between observations and knowledge schemata, which provide explanations of observational patterns and their relationships. these processes have been identified in teacher education using the concept of professional vision (gaudin & chaliès, 2015; sherin, jacobs, & philipp, 2011). in professional vision research, these diagnostic processes are referred to as noticing relevant information and as knowledge-based reasoning, including both aspects of explanations as well as predictions. even if not every single diagnostic situation is covered by them, classifications of diagnostic activities, such as the epistemic activities described above (fischer et al., 2014), can be used as a descriptive framework with a high degree of validity for a broad variety of diagnostic situations within and across domains. however, the set of activities should not be understood as a revival of the scientific method, a style of scientific thinking focusing solely on hypothetico-deductive arguments which has dominated many science education curricula (kind & osborne, 2017). different professions develop different standards for what it means to engage effectively and efficiently in these epistemic-diagnostic activities, how explicitly they should be conducted and how their results should be documented and communicated. for example, in the medical domain, there are explicit expectations regarding which diagnostic activities are considered necessary and in which order they should ideally be conducted (e.g. generating hypotheses based on some early information, then systematically collecting or generating evidence to test and exclude alternative hypotheses; kassirer, 2010). medicine has also developed practices and associated standards on how the process and the result of the diagnosing should be documented in written reports to colleagues. in the domain of teaching in schools, the diagnostic process of teachers in their classrooms is far less formalized, and with few exceptions, reports to colleagues are somewhat informal and are oral as opposed to following a schema or being written. the eight diagnostic activities, based on the eight epistemic activities presented above, can provide a common language with which to compare what people do to generate knowledge in differently structured diagnostic situations and in different domains; what their emphases are; what parts of the process are more formalized; which parts are communicated and how, etc. based on the suggested (joint) notion of diagnostic competences for medical and teacher education, we will now shift the focus to possibilities for facilitating the development of diagnostic competences in higher education. 3.2. advancing diagnostic competences with simulations the application of knowledge in diagnostic activities is considered crucial to the development of diagnostic competences (e.g., kolodner, 1992). possibilities for the application of knowledge are typically part of real-life professional practice and not, as such, part of higher education programmes. it was shown in several domains that professionals engage with problem-solving in a specific way to further develop their competences (ericsson, 2004). to develop professional expertise, successful learners show similar ways of practising, while they differ especially in their intentional, focused and repeated practising of the challenging aspects of a task. this idea has been used in medical as well as in teacher education (berliner, 2001). in both fields, real-life practice might not be the most promising opportunity for developing diagnostic competences for students who are still in higher education programmes. in fields in which real-life situations are highly complex and dynamic, novices such as students can easily be overwhelmed. under these circumstances, real-life approximation-of-practice (grossman et al., 2009) can help to avoid the overwhelming of learners by reducing the complexity of the practical situation and by providing additional guidance (gartmeier et al., 2015). going beyond representations which illustrate practice for students, approximations enable engaging the learners with important aspects of practice (grossman et al., 2009). several instructional approaches have been developed which can guide the design of learning environments allowing the practising of professional tasks (e.g., van merriënboer & kirschner, 2018). in this article, we focus on educational simulations as one particularly promising approach which can serve as approximations-of-practice. there are economical, practical and ethical reasons for using simulations in educational settings (ziv, wolpe, small, & glick, 2003). in contrast to engaging in real practice, approximations-of-practice using simulations allows for engagement in critical but rare situations which do not often happen during internships (or, if they do happen, the intern is not the one asked to professionally address the issue). approximations-of-practice through simulations also allow learners to engage in very difficult situations requiring repeated practice. trying a difficult technique or intervention repeatedly is typically impossible or is at least undesirable in real encounters with patients or students. thus, the main challenge is to discover how simulations can help overcome the shortcomings of real practices, such as low base rates in certain critical situations or the lack of opportunities for repeated practice. however, another main challenge in designing simulation-based learning environments is learning how simulations can offer opportunities for knowledge application and practice without overwhelming the learner. in the remainder of this section, we briefly introduce the aim and different types of simulations before reviewing research on using simulations to measure and facilitate diagnostic competences. a simulation is a model of a natural or artificial system with certain features which can be manipulated. the aim of a simulation is to arrive at a better understanding of the interconnections of the variables in the system or to put different strategies for controlling the system to the test (wissenschaftsrat, 2014, building on frasson & blanchard, 2012; shannon, 1975). in research on learning and instruction and medical education, two different types of simulations can be distinguished by their functions: 1) in discovery or inquiry learning, simulations represent a segment of reality. through manipulating variables, learners can build knowledge about those variables and their interconnections, for example chemical reactions (linn, lee, tinker, husic, & chiu, 2006), in simulated systems (de jong, 2006). this type of simulation is suitable for acquiring conceptual knowledge of the interconnection of entities in complex systems (de jong, hendrikse, & meij, 2010). 2) in simulation training, learners get the opportunity to act in a simulation and thereby develop complex skills. examples may include pilots’ training in landing an airplane (landriscina, 2011) or the enactment of a surgery in medical education (al-kadi & donnon, 2013). this type of simulation has been used successfully in different areas, for example in economics (ahn, 2008), in management (stewart, williams, smith-gratto, black, & kane, 2011), in nursing (smith & barry, 2013) and in medical education (siebeck, schwald, frey, röding, stegmann, & fischer, 2011; stegmann et al., 2012; cook et al., 2012; okuda et al., 2009). there are no strict boundaries between these two types of simulations. for example, through explorations of a simulation in an inquiry environment, learners can build conceptual knowledge about the interconnections between variables and later apply this knowledge in training to further develop their diagnostic or intervention-related competences. the second type of simulation seems more appropriate for fostering diagnostic competences, because in higher education, the opportunities for knowledge application are rather scarce. with this focus, a simulation is a learning environment in which (1) a segment of reality (e.g. a professional situation) is presented in a way which enables engagement in diagnostic activities (e.g. through presenting the behaviour of a student or the results of a diagnostic test of a patient). in a simulation of this type, (2) learners’ actions influence the further development of the system. thus, a central goal of learning with simulations is to provide training opportunities in which learners can take diagnostic actions in cases with a certain similarity to cases in their professional practice (gartmeier et al., 2015; shavelson, 2013). in this sense, both digital simulations and role-play (e.g. taking on the role of a doctor or a patient in an anamnestic role-play or of a teacher or student in a diagnostic interview) have been used as simulation-based learning environments. in medical education, learning via simulations in skill labs is already common practice (peeraer et al., 2007). there are empirical studies with simulations, such as role-play with standardized patients (bokken, linssen, scherpbier, van der vleuten, & rethans, 2009), and the findings show that students benefit not only from directly interacting with a standardized patient but also from observing the interaction of a peer with a standardized patient (stegmann et al., 2012). a wealth of empirical studies have recently been summarized in meta-analyses, establishing a medium positive effect for simulation-based learning in contrast to other forms of training for medical areas as different as clinical reasoning or resuscitation (cook, 2014). in teacher education, there have been studies which assessed diagnostic competences with a computer-based simulation of a virtual classroom (kaiser, helm, retelsdorf, südkamp, & möller, 2012; südkamp, möller, & pohlmann, 2008) or teachers’ recommendations for the type of school a student should choose after primary school (gräsel & böhmer, 2013). however, the purpose of those simulations was not to foster the development of diagnostic competences but rather to measure the learners’ level of diagnostic competence. there has also been emphasis in teacher education on so-called clinical experience (grossman, 2010) and clinical simulations (dotger, 2013), which aim at advancing pre-service teachers’ knowledge and skills. de coninck, valcke, ophalvens and vanderlinde (2019) recently reviewed research on simulations in teacher education. one important conclusion was that simulations need to be embedded in instructionally well-designed learning environments to be effective. building on earlier work (e.g., gartmeier, bauer, fischer, karsten, & prenzel, 2011), de coninck et al. (2019) further developed an approach to using simulations in learning environments. their goal was to facilitate teachers’ competence in engaging in effective communication with parents. they suggested four instructional principles for simulation-based learning environments (e.g. invoke cyclical process, including simulation-based experience, feedback and reflection) and found preliminary evidence for the perceived usefulness of online and face-to-face simulations which were designed using these principles. although this line of research is highly promising, it has not yet yielded systematic experimental evidence with objective measures of competence which would help establish a model of instructional design for simulations that is based on a theory. to sum up, both medical and teacher education have developed conceptual accounts of the pedagogical potentials of simulations and have brought forward empirical studies indicating that simulations can provide opportunities for engaging students in realistic activities and that they can help both in assessing and in advancing competences, including diagnostic competences. research on using simulations to advance competences has had a longer tradition and has yielded more empirical studies in medical education than in teacher education. therefore, there is more systematic knowledge available in medical education on the question of the effective design of simulation-based learning environments. while it is obvious from this available knowledge that there are clear advantages to the use of simulations, there is also theory and evidence that learners can be overwhelmed by being involved in solving complex problems. this is especially true for learners with little or no prior knowledge necessary to tackle the task (e.g., renkl, 2014). in the next section, we will therefore analyse how learners can be effectively guided while learning with simulations. 3.3. additional instructional support in simulation-based learning environments there is ample evidence in research on simulations that learners benefit from additional instructional support (cook et al., 2013; lazonder & harmsen, 2016, wouters & van oostendorp, 2013). cook et al. (2013) analysed experimental studies in medical education regarding the effectiveness of 13 different design features. for instance, they studied whether the simulated scenarios were highly interactive, whether they differed widely in difficulty or offered a broad clinical variation and what the effects of repetitive, distributed and collaborative practice were. effects were greater for more interactive, repetitive and distributed experiences. a broader range of difficulty and clinical variation as well as the availability of feedback and collaborative practice were only effective on some of the outcome measures considered but did not show consistent positive effects. though interesting, the design features which cook et al. (2013) considered were focused more on the training regimen applied in the simulation, that is, how long and in which sequence learners should engage in which kinds of scenarios. however, the meta-analysis did not address how learners were instructionally supported while dealing with the simulated scenarios, that is how their problem-solving within the scenarios was supported. we therefore conclude that there is urgently missing evidence from primary studies with respect to more theoretically grounded mechanisms of instructional support. but a promising instructional design concept in this respect is scaffolding, which is adaptive support during the solution of a task. during scaffolding, a teacher, peer or computer-based system takes over elements of the task, so the learner must carry out only a part of the task within their reach (wood, bruner, & ross, 1976). scaffolding thus enables a learner to solve problems they would not be able to solve without that support (quintana et al., 2004; wood et al., 1976). the main functions of scaffolding include directing the attention to relevant but not obvious aspects of a task, reducing the degrees of freedom during the solution of a task, supporting learners in focusing on the goal, highlighting critical features of the task, supporting learners in dealing with failure and frustration and demonstrating important steps (wood et al., 1976; see also pea, 2004). a central aspect of scaffolding is that it is flexibly adapted to a learner’s progress. the goal of scaffolding is to support the learner in a way which increasingly makes self-regulated problem-solving possible (cf. wecker & fischer, 2011). in the context of simulations, we consider three forms of scaffolding as particularly relevant. 1) prompting during the solution of a task. to facilitate learners’ task solutions, they can be supported with prompts (or hints), a basic form of scaffolding. numerous empirical studies provide evidence for the effectiveness of support during the solution of tasks with respect to the application of knowledge. for instance, by means of socio-cognitive scaffolds (fischer, kollar, stegmann, & wecker, 2013), it is possible to support learners who must solve problems in collaboration with others. scaffolds which guide learning partners to reciprocally refer to each other’s contributions were found to be particularly effective (vogel, wecker, kollar, & fischer, 2017). prompts supporting the learners’ metacognition can also help them monitor their own learning processes (bisra, liu, nesbit, salimi, & winne, 2018). in some cases, metacognitive prompts were only successful in combination with other kinds of instructional support (berthold, nückles, & renkl, 2007; roll, holmes, day, & bonn, 2012; stark, tyroller, krause, & mandl, 2008). 2) role-taking. in simulated diagnostic situations, learners can take on different roles: (a) the role of the diagnosing person (a physician or a teacher), (b) the person with the characteristic to be diagnosed (a patient or a student) or (c) an observer. it is often regarded as a necessary prerequisite for competence development that learners actively engage with learning tasks (schank, berman, & macpherson, 1999). therefore, it seems plausible that the role of the diagnosing person can particularly foster the development of diagnostic competences. but a particular challenge for diagnostic situations is that it is often helpful if the diagnosing individual takes the perspective of the patient or the student. in teaching, the diagnosis of a misconception is facilitated if the diagnostician manages to relate a student’s observable solution strategies to potential misconceptions. hence, to take the perspective of the student may be an important part of the task of diagnosing misconceptions. therefore, role-taking in a simulation can be regarded as a form of part-task practice (i.e. engaging just in the emulation of a student’s cognitive processing that constitutes part of the diagnostic task; see also 4c/id, e.g. van merriënboer & kirschner, 2018, with respect to part-task practice), which reduces the complexity of the task. hence, in line with the definition of scaffolding mentioned above, role-taking can be regarded as a form of scaffolding. empirical findings show that acting in the role of the person with the characteristic to be diagnosed can have a positive effect on the development of diagnostic competences. the quality of perspective-taking is fostered by directly telling the learners to take over a role (goeze, zottmann, vogel, fischer, & schrader, 2014). research on vicarious learning (cf. van gog & rummel, 2010) has also provided some evidence suggesting that the role of the observer can positively affect diagnostic competences. in a study in which doctor–patient communication was simulated, medical students learned just as much whether they participated actively or just observed the simulation (stegmann et al., 2012). further systematic research is needed to answer the question of how role-taking affects the development of diagnostic competences. 3) reflection phases. another promising form of scaffolding in simulations is adding reflection phases with the goal of reducing the time pressure which otherwise often exists in diagnostic situations in which a doctor or a teacher needs to act immediately. thus, a more detailed planning of further steps is possible. reflection can help learners adequately generate courses of action. to monitor learners’ own learning and performance is another important goal of competence development. besides feedback from other individuals (hattie & timperley, 2007), another important source of information for monitoring their own performance is feedback which learners generate themselves in reflection phases (nicol & macfarlane-dick, 2006). in reflection phases, instructional material or an instructor can encourage learners to think about the goals of a procedure, to analyse their own performance and to plan further steps. adding reflection phases to foster medical diagnostic competences has turned out to be a successful form of scaffolding (ibiapina, mamede, moura, eloi-santos, & van gog, 2014; mamede et al., 2012; mamede et al., 2014). however, these initial findings from medical education are far from robust and conclusive. and reflection phases can be of various lengths, can be more or less structured by guiding questions, can take place in individual or collaborative settings and can also take place before, during or after working with the simulation. the extent to which the effects of reflection phases generalize to diagnostic competences in teaching also remains an open question. explicit presentation of information. scaffolded problem-solving is often contrasted with learning through lectures and book-reading. this explicit presentation of information can provide learners with additional conceptual or strategic knowledge, for example through additional lecture videos, which learners can access before, during or after work with the simulation. an adequate professional knowledge base which learners can apply to diagnostic tasks plays a major role in an early phase of competence development (anderson & lebiere, 1998; schank et al., 1999; vanlehn, 1996). there is also evidence from learning with worked examples, in which learners with unfavourable learning prerequisites (e.g. low prior knowledge or poor self-regulatory skills, van gog & rummel, 2010) can benefit from the explicit presentation of information (e.g., moreno, 2004). there is also research on the question of how phases of learner-directed inquiry and discovery activities can be productively combined with the explicit presentation of information. research seems to suggest that there is a “time for telling” (schwartz & bransford, 1998; wecker, rachel, heran-dörr, waltner, wiesner, & fischer, 2013); however, systematic research on combining phases of explicit presentation of information, scaffolding and unguided problem-solving in simulation-based learning is scarce. to conclude, existing research on complex learning environments, including simulations, suggests that learning is more effective when the environments include instructional support (lazonder & harmsen, 2016; wouters & van oostendorp, 2013). we hence suggest investigating the effects of different approaches to instructional support in simulations, with a focus on concepts with a strong theory-based foundation. a promising example of such an instructional concept is scaffolding. accordingly, we integrate into our conceptual framework different forms of scaffolding which we consider particularly promising in the context of learning to diagnose in simulations; these include support during problem-solving activities via prompts as well as role-taking and reflection phases (see the box on the left in fig. 1). we also integrate the explicit presentation of information; there is “a time for telling” about domain concepts and strategies, and the timing does not seem to be arbitrary. integrating both into the framework, the different forms of scaffolding and the explicit presentation of information, will enable research on their combination or succession in simulation-based learning, for example prompts which remind learners to implement a strategy they developed in the reflection phase or reflection phases which follow a phase of self-directed problem-solving with the explicit presentation of information introducing the relevant concepts for reflection. 3.4. learning prerequisites and contextual factors as moderators of instructional effects the effects of simulations and instructional guidance on diagnostic competences will likely depend on differences between learners and on the characteristics of the simulated situation. the conceptual model entails important individual prerequisites and contextual factors as possible moderators of the effects of simulations and guidance (fig. 1). 3.4.1. learning prerequisites it is unlikely that a simulation-based learning environment or any kind or combination of instructional support could have equally positive effects for all learners regardless of their learning prerequisites. in research on learning and instruction, there is evidence of the moderating effects of individual learner characteristics. such effects have been called aptitude-treatment interactions (snow, 1991). not much is known so far about aptitude-treatment interaction effects in learning with simulations. a comprehensive research programme on instructional support for the development of diagnostic competences in simulation-based learning environments must address the question of how important these prerequisites are for novel, dynamic contexts. in what follows, we suggest some promising cognitive, affective and personality-related moderators. these may serve as starting points for more systematic research on how learning prerequisites moderate the effects of simulation-based learning with and without additional instructional support. possibly the most important single cognitive prerequisite is prior knowledge (gerard, matuk, mcelhaney, & linn, 2015). instructional support approaches which are highly effective for learners with low prior knowledge can be much less effective for or even detrimental to learners with more advanced prior knowledge (expertise reversal effect, kalyuga, 2007). in novel tasks (where specific prior knowledge and experience are absent), general cognitive abilities have been shown to be highly predictive for learning. in particular, learners with high levels of fluid intelligence seem to have an advantage when engaging in new and complex tasks (cf. schwaighofer, 2015). so-called executive functions have also come into focus in research on learning and instruction. executive functions are responsible for self-regulation and include the ability to update and monitor representations in working memory, the ability to shift between parts of the task or between different mental representations and the ability to inhibit automatic or dominant responses (miyake & friedman, 2012). these functions seem to be particularly important in learning contexts (schwaighofer, fischer, & bühner, 2015). while instructional research has focused on working memory functions for two decades now (e.g. koopmann-holm & o’connor, 2017), shifting and inhibition have received much less attention. recently, shifting has been shown to moderate the effects of instructional design features in complex learning with multiple sources of information (schwaighofer, bühner, & fischer, 2016). learners with lower shifting abilities benefitted more from scaffolding with worked examples than learners with higher shifting abilities. schwaighofer et al. (2015) put forward the hypothesis that in complex learning environments, it is the coordinated interplay of the executive functions rather than one of these abilities in isolation which is responsible for successful learning. also, motivational and affective factors may be key, for example interest (rotgans & schmidt, 2014), self-efficacy (zimmerman, 2000) or epistemic emotions (pekrun & linnenbrink-garcia, 2012). these variables have been included in studies both as moderating factors of the effect of complex learning on the development of competences as well as outcomes in their own right. however, they have hardly been deployed systematically in research on learning with simulations. simulations for learning as well as serious games have nourished hopes for increased motivation and the growing interest of learners (sailer, hense, mandl, & klevers, 2013; wouters, van nimwegen, van oostendorp, & van der spek, 2013). learning through interacting with a complex and dynamic simulation differs substantially from following a well-structured lecture or reading a chapter in a textbook. it seems likely that personality traits (e.g. openness to experience, costa & mccrae, 1985) and individual preferences, such as the tolerance of ambiguity (e.g., hancock, roberts, monrouxe, & mattick, 2015), do play a role in how individuals engage in and learn from simulations and in how they experience and benefit from additional instructional support (e.g. prompts or role-plays). recently, relationships have been found between some personality traits, like diligence, and the development of competence (cf. pellegrino & hilton, 2013). but there is hardly any systematic research which has investigated the potentially moderating effects of relatively stable personality traits on the development of competences in simulation-based learning. consequently, our conceptual framework for guiding research on learning to diagnose with simulations systematically addresses individual learning prerequisites (fig. 1). these entail cognitive factors (e.g. prior professional knowledge, executive functions) as well as motivational and affective factors (e.g. interest, self-efficacy, epistemic emotions). the framework also includes personality traits (e.g. openness to experience). the benefit of identifying the dependencies of instructional effects on pre-existing interindividual differences between learners is obvious; the systematic inclusion of these factors will likely yield scientific knowledge which helps answer the question of for whom specific instructional designs are likely to be effective. scientific knowledge on the moderating effects of individual learning prerequisites can therefore contribute to making simulation-based learning environments more adaptive and personalized. individual differences are not the only group of moderating factors we assume in our conceptual model. the other group of potential moderators are contextual (fig. 1) and will be introduced in the next section. 3.4.2. contextual factors typically, instructional researchers are alert to differences between ill-structured and well-structured domains with respect to effective instruction (spiro, feltovich, jacobson, & coulson, 1992), but they often assume that similar instructional effects can be obtained across different domains as long as the task structures are comparable. however, it is not at all straightforward that analogical instructional effects are obtained in different diagnostic situations within and across domains. the nature of the diagnostic situation is constituted by the features apparent in the situation where the diagnosis takes place. no established classification for the nature of diagnostic situations is available to date, not even in medicine, where research on diagnosing has a longer tradition than in teaching. heuristically diagnostic situations can differ along two dimensions, (1) interaction-based vs. document-based diagnosis and (2) individual vs. collaborative diagnosis. 1) interaction vs. document-based. the differentiation of interaction-based and document-based diagnosis is based on the main source of information necessary for the diagnosis. in interaction-based diagnosis (e.g. in anamnestic situations with a patient or during a diagnostic interview with a student concerning arithmetic skills), this information is gathered from the interaction with the patient or student. there is a high likelihood of time pressure and little opportunity for reflection. in document-based diagnosis (e.g. diagnosis based on a patient’s laboratory findings or diagnosis of a student’s competence level based on the analysis of written assignments), the necessary information is available in written, pictorial or video format. in these situations, there is typically no time pressure, and reflection is possible. the distinction between interaction-based and document-based diagnosis is a relevant distinction from a practical point of view. both types of situations might entail different demands, and hence, different instructional support might be necessary. 2) individual vs. collaborative. the necessity for collaboration and communication with other professionals during the diagnostic process characterizes a second dimension of diagnostic situations. in many such situations, an individual diagnostician arrives at a diagnosis. however, there are also situations in which collaboration is necessary and routine, for example consultations of medical experts from different specialties on complex patient cases. empirical studies provide initial evidence that the social regulation of diagnostic activities causes difficulties even for advanced practitioners (christensen et al., 2000; tschan et al., 2009). so far, however, there is little empirical research on the competencies necessary for such collaborations or on ways to foster them (cf. kiesewetter, fischer, & fischer, 2017). in schools, collaborative diagnoses do not (yet) seem to be part of teachers’ daily routines. however, there are instances, for example with students having multiple learning difficulties in several subjects, in which teachers must collaboratively diagnose a student’s situation. another still unexplored case of collaborative diagnosis in schools is the collaboration between teachers of different subjects in diagnosing cross-domain competencies, for example a biology and a physics teacher diagnosing students’ scientific reasoning skills based on their encounters with these students in their own lessons. our conceptual framework includes the diagnostic situation and the domain as contextual moderators (fig. 1). the tasks which are considered diagnostic tasks vary greatly, even within a single domain. regarding diagnostic situations, the conceptual framework will initially entail two heuristic dimensions along which diagnostic situations can be distinguished: interaction-based vs. document-based and individual vs. collaborative diagnosing. these two dimensions are not regarded as the only important aspects of diagnostic situations but are meant to serve as a starting point. including them in a research framework allows the question of generalizability across types of simulated diagnostic situations (such as individual vs. collaborative diagnosis) and domains (such as medicine and teaching) to be addressed more systematically. 4. from a conceptual framework to a theory: a research agenda in the preceding sections of this article, we laid out a conceptual framework (fig. 1) which may possibly guide joint research programmes of medical and teacher education in analysing and facilitating diagnostic competences. one primary goal of this research programme would be to advance the conceptual framework from being just a joint language to becoming a joint theory with specified relationships between its concepts. such a research programme would first need to identify or develop and validate simulations as approximations of medical and teaching practices. validation studies need to generate evidence that the simulation corresponds to the simulated situation, at least to a certain degree. such a programme would also have to agree on ways to measure the different professional knowledge bases and their coordinated application as well as their engagement in diagnostic activities (e.g. evaluating evidence) and their diagnostic quality, for instance, the simulations’ accuracy (e.g., hege, kononowicz, kiesewetter, & foster-johnson, 2018; südkamp et al., 2012) and efficiency (e.g., braun et al., 2017). these measurements will be functional in improving our understanding of the causal relationships of instructional support features (e.g. prompts, role-play, reflection phases) and their combination in simulation-based learning environments with the development of diagnostic competences. causal hypotheses are that instructional support leads to qualitatively improved engagement in diagnostic activities in simulation-based learning environments and that improved engagement in diagnostic activities, in turn, leads to an advancement of diagnostic competences. to test these types of hypotheses, experimental mediation designs with measures of the diagnostic activities or sequences thereof during the learning phase seem promising. combinations of scaffolding, with explicit presentation of information at varying points in time, enable the testing of hypotheses regarding the optimal timing for such explicit presentation of information in the context of scaffolded diagnostic activities. including cognitive, motivational and affective learning prerequisites in the research programme allows the assessment of how these prerequisites moderate the size or even the direction of instructional effects. for instance, the expertise reversal effect (kalyuga, 2007), well established for example-based learning, may be tested for other forms of scaffolding. generating more systematic scientific knowledge on the role of these prerequisites will eventually enable more adaptive designs of simulations and the instructional guidance they entail. regarding the contextual factors of the diagnostic situations, it seems promising to apply the same instructional approach to two or more diagnostic situations which differ markedly, to learn about how specific or general an instructional effect is. for example, the effects of reflection phases can be investigated for highly dynamic interaction-based situations and for less dynamic document-based diagnosing. with respect to domain, it would seem an appropriate first step to use the same instructional approach in environments simulating structurally similar diagnostic situations (e.g. document-based and collaborative diagnosing) in at least two different domains. a hypotheses which can be tested is the claim that the effects of simulations and instructional guidance differ between different diagnostic situations within domains, probably even more than between the same types of diagnostic situations in different domains. to seriously address the challenges of involving different fields such as the medical field and teacher education, a research programme also implies interdisciplinary research collaboration. in order to be able to compare the findings from different studies, the data structure needs to be consistent across corresponding research projects, and some of the variables introduced earlier need to be measured in identical or at least comparable ways (e.g. motivation, cognitive load). by using learning prerequisites, diagnostic situations and domains as potential moderators of the overall effects, meta-analytic techniques could be used to synthesize the results from several experimental studies on the effects of various approaches of instructional support on the development of diagnostic competences. research programmes, such as the one outlined here, will be instrumental in further advancing the conceptual framework into a theoretical account of a way of instructionally supporting the development of diagnostic competences in simulation-based learning environments. such research programmes can also contribute to persisting issues in research on learning and instruction, such as the interaction of knowledge and reasoning; the differential effects of different types of scaffolding for differently advanced learners; the domain specificity of instructional guidance; and the role of the authenticity of learning environments in higher education for the development of competence. keypoints diagnostic competences are considered crucial in many professions, including the medical and teaching professions. the article proposes a new way to conceptualize diagnostic competences across professional domains, including both outcome and process-related indicators. based on research from medical education, this article proposes that simulations may be effective in fostering diagnostic competences in medical and teacher education. instructional guidance for learning in simulations requires further research. scaffolding with prompts, roles and reflection phases seems to be promising. the effects of simulations and additional instruction are likely to depend on learning prerequisites as well as on the type of situation the simulation represents. this article proposes a conceptual model and an interdisciplinary research agenda involving scholars from teaching and medical education. acknowledgement the research for this article was funded by the german research association (deutsche forschungsgemeinschaft, dfg) (for2385) references abs, h. j. (2007). überlegungen zur modellierung diagnostischer kompetenz bei lehrerinnen und lehrern [considerations of modelling teachers’diagnostic competence]. in m. lüders & j. wissinger (eds.), forschung zur lehrerbildung: kompetenzentwicklung und programmevaluation [research on teacher education. competence development and program evaluation] (pp. 63–84). waxmann verlag. ahn, j.-h. (2008). application of the experiential learning cycle in learning from a business simulation game. e-learning, 5 (2), 146–156. https://doi-org.emedien.ub.uni-muenchen.de/10.2304/elea.2008.5.2.146 al-kadi, a. s., & donnon, t. (2013). using simulation to improve the cognitive and psychomotor skills of novice students in advanced laparoscopic surgery: a meta-analysis. medical teacher, 35(sup1), 47-55. doi: 10.3109/0142159x.2013.765549 anderson, j. r., & lebiere, c. j. (1998). the atomic components of thought. mahwah, nj: lawrence erlbaum associates. doi: 10.4324/9781315805696 barrows, h. s., & pickell, g. c. (1992). developing clinical problem solving skills: a guide to more effective diagnosis and treatment . new york, ny: ww norton & co. bennett, r. e. (2011). formative assessment: a critical review.assessment in education: principles, policy & practice, 18(1), 5–25. https://doi.org/10.1080/0969594x.2010.513678 berliner, d. c. (2001). learning about and learning from expert teachers. international journal of educational research, 35, 463–482. doi: 10.1016/s0883-0355(02)00004-6 berthold, k., nückles, m., & renkl, a. (2007). do learning protocols support learning strategies and outcomes? the role of cognitive and metacognitive prompts. learning and instruction, 17(5), 564–577. https://doi.org/10.1016/j.learninstruc.2007.09.007 bisra, k., liu, q., nesbit, j. c., salimi, f., & winne, p. h. (2018). inducing self-explanation: a meta-analysis. educational psychology review, 1-23. doi:10.1007/s10648-018-9434-x bokken, l., linssen, t., scherpbier, a., van der vleuten, c., & rethans, j.-j. (2009). feedback by simulated patients in undergraduate medical education: a systematic review of the literature. medical education, 43(3), 202–210. https://doi.org/10.1111/j.1365-2923.2008.03268.x borko, h., roberts, s. a., & shavelson, r. (2008). teachers’ decision making: from alan j. bishop to today. in p. clarkson & n. presmeg (eds.), critical issues in mathematics education (1. aufl. ed., pp. 37–67). s.l.: springer-verlag. boshuizen, h. p. a., & schmidt, h. (1992). on the role of biomedical knowledge in clinical reasoning by experts, intermediates and novices. cognitive science, 16(2), 153–184. https://doi.org/10.1016/0364-0213(92)90022-m boshuizen, h. p. a., & schmidt, h. (2008). the development of clinical reasoning expertise. in j. higgs, m. a. jones, s. loftus, & n. christensen (eds.), clinical reasoning in the health professions (3rd ed., pp. 113–122). amsterdam: elsevier (butterworth heinemann). boshuizen, h. p. a., schmidt, h., custers, e., & van de wiel, m. (1995). knowledge development and restructuring in the domain of medicine: the role of theory and practice. learning and instruction, 5(4), 269–289. https://doi.org/10.1016/0959-4752(95)00019-4 braun, l. t., zottmann, j. m., adolf, c., lottspeich, c., then, c., wirth, s., ... & schmidmaier, r. (2017). representation scaffolds improve diagnostic efficiency in medical students. medical education, 51(11), 1118-1126. doi: 10.1111/medu.13355 charlin, b., boshuizen, h. p. a., custers, e. j., & feltovich, p. j. (2007). scripts and clinical reasoning. medical education, 41(12), 1178–1184. doi: 10.1111/j.1365-2923.2007.02924.x christensen, c., larson, j. r., jr, abbott, a., ardolino, a., franz, t., & pfeiffer, c. (2000). decision making of clinical teams: communication patterns and diagnostic error. medical decision making: an international journal of the society for medical decision making , 20(1), 45–50. doi: 10.1177/0272989x0002000106 cook, d. a. (2014). how much evidence does it take? a cumulative meta-analysis of outcomes of simulation-based education. medical education, 48(8), 750–760. https://doi.org/10.1111/medu.12473 cook, d. a., brydges, r., hamstra, s. j., zendejas, b., szostek, j. h., wang, a. t., … hatala, r. (2012). comparative effectiveness of technology-enhanced simulation versus other instructional methods: a systematic review and meta-analysis. simulation in healthcare, 7(5), 308–320. doi: 10.1097/sih.0b013e3182614f95 cook, d. a., hamstra, s. j., brydges, r., zendejas, b., szostek, j. h., wang, a. t., … hatala, r. (2013). comparative effectiveness of instructional design features in simulation-based education: systematic review and meta-analysis. medical teacher, 35(1), 867-898. doi: 10.3109/0142159x.2012.714886 costa, p. t., & mccrae, r. r. (1985). the neo personality inventory. journal of career assessment, 3(2), 123–129. croft, p., altman, d. g., deeks, j. j., dunn, k. m., hay, a. d., hemingway, h., ... & riley, r. d. (2015). the science of clinical practice: disease diagnosis or patient prognosis? evidence about “what is likely to happen” should shape clinical practice. bmc medicine, 13, 1-8. doi: 10.1186/s12916-014-0265-4 de coninck, k., valcke, m., ophalvens, i., & vanderlinde, r. (2019). bridging the theory-practice gap in teacher education: the design and construction of simulation-based learning environments. in k. hellmann, j. kreutz, m. schwichow, & k. zaki (eds.), kohärenz in der lehrerbildung (pp. 263-280). wiesbaden: springer vs. diagnosing. (n.d.) in cambridge dictionary online. retrieved from https://dictionary.cambridge.org/dictionary/english/diagnosing . retrieved 11:24, may 25, 2019 de jong, t. (2006). technological advances in inquiry learning. science, 312(5773), 532–533. https://doi.org/10.1126/science.1127750 de jong, t., hendrikse, p., & meij, h. van der. (2010). learning mathematics through inquiry: a large-scale evaluation. in m. j. jacobson & p. reimann (eds.), designs for learning environments of the future (pp. 189–203). springer us. retrieved from http://link.springer.com/chapter/10.1007/978-0-387-88279-6_7 dotger, b. h. (2013). i had no idea: clinical simulations for teacher development. information age publishing inc. ericsson, k. a. (2004). deliberate practice and the acquisition and maintenance of expert performance in medicine and related domains. academic medicine, 79(10), 70–81. doi: 10.1097/00001888-200410001-00022 fischer, f., kollar, i., stegmann, k., & wecker, c. (2013). toward a script theory of guidance in computer-supported collaborative learning. educational psychologist, 48(1), 56–66. doi: 10.1080/00461520.2012.748005 fischer, f., kollar, i., ufer, s., sodian, b., hussmann, h., pekrun, r., … eberle, j. (2014). scientific reasoning and argumentation: advancing an interdisciplinary research agenda in education. frontline learning research, 2(2), 28–45. https://doi.org/10.14786/flr.v2i2.96 förtsch, c., sommerhoff, d., fischer, f., fischer, m., girwidz, r., obersteiner, a., ... & seidel, t. (2018). systematizing professional knowledge of medical doctors and teachers: development of an interdisciplinary framework in the context of diagnostic competences. education sciences, 8(4), 207. doi: 10.3390/educsci8040207 frasson, c., & blanchard, e. (2012). simulation-based learning. in n. seel, encyclopedia of the sciences of learning (pp. 3076–3080). boston: springer. gartmeier, m., bauer, j., fischer, m. r., karsten, g., & prenzel, m. (2011). modellierung und assessment professioneller gesprächsführungskompetenz von lehrpersonen im lehrer-elterngespräch [modeling and assessment of teachers' professional competence for parent-teacher conversations]. in o. zlatkin-troitschanskaja (ed.). stationen empirischer bildungsforschung (pp. 412-424). vs verlag für sozialwissenschaften. doi: 10.1007/978-3-531-94025-0_29 gartmeier, m., bauer, j., fischer, m. r., hoppe-seyler, t., karsten, g., kiessling, c., … prenzel, m. (2015). fostering professional communication skills of future physicians and teachers: effects of e-learning with video cases and role-play. instructional science, 43(4), 443–462. https://doi.org/10.1007/s11251-014-9341-6 gaudin, c., & chaliès, s. (2015). video viewing in teacher education and professional development: a literature review. educational research review, 16, 41-67. doi:10.1016/j.edurev.2015.06.001 gerard, l., matuk, c., mcelhaney, k., & linn, m. c. (2015). automated, adaptive guidance for k-12 education. [review]. educational research review, 15, 41-58. doi: 10.1016/j.edurev.2015.04.001 glogger-frey, i., herppich, s., & seidel, t. (2018). linking teachers’ professional knowledge and teachers’ actions: judgment processes, judgments and training. teaching and teacher education, 76(1), 176-180. doi:10.1016/j.tate.2018.08.00 goeze, a., zottmann, j. m., vogel, f., fischer, f., & schrader, j. (2014). getting immersed in teacher and student perspectives? facilitating analytical competence using video cases in teacher education. instructional science, 42(1), 91–114. doi: 10.1007/s11251-013-9304-3 gräsel, c. (1997). wir können auch anders: problemorientiertes lernen an der hochschule [we can also do things differently: problem-oriented learning in higher education]. in h. gruber & a. renkl (eds.), wege zum können. determinanten des kompetenzerwerbs. bern: huber. gräsel, c., & böhmer, i. (2013). die übergangsempfehlung nach der grundschule. welche informationen nutzen lehrerinnen und lehrer für die entscheidung? [school tracking decisions after primary school: which information do teachers use for their decision?]. in n. mcelvany & h. g. holtappels (eds.), empirische bildungsforschung. theorien, methoden, befunde und perspektiven. festschrift für wilfried bos (pp. 235–248). münster u.a.: waxmann. grossman, p. (2010). learning to practice: the design of clinical experience in teacher preparation (policy brief) . washington, dc: american association of colleges for teacher education and national education association. grossman, p., & mcdonald, m. (2008). back to the future: directions for research in teaching and teacher education. american educational research journal, 45(1), 184-205. doi:10.3102/0002831207312906 grossman, p., compton, c., igra, d., ronfeldt, m., shahan, e., & williamson, p. (2009). teaching practice: a cross-professional perspective. teachers college record, 111(9), 2055–2100. hancock, j., roberts, m., monrouxe, l., & mattick, k. (2015). medical student and junior doctors’ tolerance of ambiguity: development of a new scale. advances in health sciences education, 20(1), 113–130. https://doi.org/10.1007/s10459-014-9510-z hargreaves, d. h. (2000). the production, mediation and use of professional knowledge among teachers and doctors: a comparative analysis. in oecd (ed.), knowledge management in the learning society (pp. 219-238). paris: oecd publishing. hattie, j. (2003). formative and summative interpretations of assessment information. retrieved from http://www.education.auckland.ac.nz/webdav/site/education/shared/hattie/docs/formative-and-summative-assessment-(2003).pdf. hattie, j., & timperley, h. (2007). the power of feedback. review of educational research, 77(1), 81–112. https://doi-org.emedien.ub.uni-muenchen.de/10.3102/003465430298487 hege, i., kononowicz, a. a., kiesewetter, j., & foster-johnson, l. (2018). uncovering the relation between clinical reasoning and diagnostic accuracy–an analysis of learner's clinical reasoning processes in virtual patients. plos one, 13(10), e0204900. doi: 10.1371/journal.pone.0204900. heitzmann, n. (2014, january 22). fostering diagnostic competence in different domains (text.phdthesis). ludwig-maximilians-universität münchen. retrieved from http://edoc.ub.uni-muenchen.de/16862/ helmke, a. (2010). unterrichtsqualität und lehrerprofessionalität diagnose, evaluation und verbesserung des unterrichts [instructional quality and teacher professionalism. diagnosis, evaluation and improvement of instruction] . seelze-velber: klett/kallmeyer. herppich, s., praetorius, a.-k., förster, n., glogger-frey, i., karst, k., leutner, d., … südkamp, a. (2018). teachers’ assessment competence: integrating knowledge-, process-, and product-oriented approaches into a competence-oriented conceptual model. teaching and teacher education. https://doi.org/10.1016/j.tate.2017.12.001 ibiapina, c., mamede, s., moura, a. s., eloi-santos, s., & van gog, t. (2014). effects of free, cued, and modelled reflection on medical students’ diagnostic competence. medical education, 48, 796-805. doi: 10.1111/medu.12435. kaiser, j., helm, f., retelsdorf, j., südkamp, a., & möller, j. (2012). zum zusammenhang von intelligenz und urteilsgenauigkeit bei der beurteilung von schülerleistungen im simulierten klassenraum [on the relation of intelligence and judgment accuracy in the process of assessing student achievement in thesimulated classroom]. zeitschrift für pädagogische psychologie / german journal of educational psychology , 26(4), 251–261. https://doi.org/10.1024/1010-0652/a000076 kalyuga, s. (2007). expertise reversal effect and its implications for learner-tailored instruction. educational psychology review, 19(4), 509–539. doi: 10.1007/978-1-4419-8126-4_12 kassirer, j. p. (2010). teaching clinical reasoning: case-based and coached. academic medicine: journal of the association of american medical colleges , 85(7), 1118–1124. doi: 10.1097/acm.0b013e3181d5dd0d kaufman, d., yoskowitz, a., & patel, v. l. (2008). clinical reasoning and biomedical knowledge: implications for teaching. in j. higgs, m. a. jones, s. loftus, & n. christensen (eds.), clinical reasoning in the health professions (pp. 137–150). amsterdam: elsevier health sciences. kiesewetter, j., fischer, f., & fischer, m. r. (2017). collaborative clinical reasoning—a systematic review of empirical studies.journal of continuing education in the health professions, 37(2), 123-128. doi: 10.1097/ceh.0000000000000158 kind, p. m., & osborne, j. (2017). styles of scientific reasoning a cultural rationale for science education? science education, 101 (1), 8–31. https://doi.org/doi:10.1002/sce.21251 kirschner, p. a., verschaffel, l., star, j., & van dooren, w. (2017). there is more variation within than across domains: an interview with paul a. kirschner about applying cognitive psychology-based instructional design principles in mathematics teaching and learning. zdm, 49 (4), 637-643. doi: 10.1007/s11858-017-0875-3 klug, j., bruder, s., kelava, a., spiel, c., & schmitz, b. (2013). diagnostic competence of teachers: a process model that accounts for diagnosing learning behavior tested by means of a case scenario. teaching and teacher education: an international journal of research and studies , 30, 38–46. https://doi.org/10.1016/j.tate.2012.10.004 kolodner, j. l. (1992). an introduction to case-based reasoning. artificial intelligence review, 6(1), 3–34. https://doi.org/10.1007/bf00155578 koopmann-holm, b., & o’connor, a. (2017). working memory. crc press. landriscina, f. (2011). simulation and learning: the role of mental models. in n. seel, encyclopedia of the sciences of learning (pp. 3072–3075). springer. lazonder, a. w., & harmsen, r. (2016). meta-analysis of inquiry-based learning: effects of guidance. review of educational research, 86(3), 681–718. https://doi.org/10.3102/0034654315627366 linn, m. c., lee, h.-s., tinker, r., husic, f., & chiu, j. l. (2006). teaching and assessing knowledge integration in science. science, 313(5790), 1049–1050. https://doi.org/10.1126/science.1131408 mamede, s., van gog, t., moura, a. s., de faria, r. m. d., peixoto, j. m., rikers, r. m. j. p., & schmidt, h. g. (2012). reflection as a strategy to foster medical students’ acquisition of diagnostic competence. medical education, 46(5), 464–472. https://doi.org/10.1111/j.1365-2923.2012.04217.x mamede, s., van gog, t., sampaio, a. m., de faria, r. m. d., maria, j. p., & schmidt, h. g. (2014). how can students’ diagnostic competence benefit most from practice with clinical cases? the effects of structured reflection on future diagnosis of the same and novel diseases. academic medicine january 2014, 89(1), 121–127. https://doi.org/10.1097/acm.0000000000000076 miyake, a., & friedman, n. p. (2012). the nature and organization of individual differences in executive functions: four general conclusions. current directions in psychological science, 21(1), 8–14. https://doi.org/10.1177/0963721411429458 moreno, r. (2004). decreasing cognitive load for novice students: effects of explanatory versus corrective feedback in discovery-based multimedia. instructional science, 32(1), 99–113. doi: 10.1023/b:truc.0000021811.66966.1d moskowitz, a. j., kuipers, b. j., & kassirer, j. p. (1988). dealing with uncertainty, risks, and tradeoffs in clinical decisions: a cognitive science approach. annals of internal medicine, 108(3), 435. doi: 10.7326/0003-4819-108-3-435 nicol, d. j., & macfarlane-dick, d. (2006). formative assessment and self-regulated learning: a model and seven principles of good feedback practice. studies in higher education, 31(2), 199–218. doi: 10.1080/03075070600572090 okuda, y., bryson, e. o., demaria, s., jacobson, l., quinones, j., shen, b., & levine, a. i. (2009). the utility of simulation in medical education: what is the evidence? mount sinai journal of medicine: a journal of translational and personalized medicine , 76(4), 330–343. https://doi.org/10.1002/msj.20127 patel, v. l., evans, d. a., & groen, g. j. (1989). biomedical knowledge and clinical reasoning. in d. a. evans & v. l. patel (eds.), cognitive science in medicine: biomedical modeling. (pp. 53–112). cambridge, ma us: the mit press. pea, r. d. (2004). the social and technological dimensions of scaffolding and related theoretical concepts for learning, education, and human activity. the journal of the learning sciences, 13(3), 423–451. doi: 10.1207/s15327809jls1303_6 peeraer, g., scherpbier, a., remmen, r., hendrickx, k., van petegem, p., weyler, j., & bossaert, l. (2007). clinical skills training in a skills lab compared with skills training in internships: comparison of skills development curricula. education for health, 20(3), 125. pekrun, r., & linnenbrink-garcia, l. (2012). academic emotions and student engagement. in handbook of research on student engagement (pp. 259–282). springer, boston, ma. https://doi.org/10.1007/978-1-4614-2018-7_12 pellegrino, j. w., & hilton, m. l. (2013). education for life and work: developing transferable knowledge and skills in the 21st century . washington, d.c: national academies press. quintana, c., reiser, b. j., davis, e. a., krajcik, j., fretz, e., duncan, r. g., … soloway, e. (2004). a scaffolding design framework for software to support science inquiry. journal of the learning sciences, 13(3), 337–386. https://doi.org/10.1207/s15327809jls1303_4 renkl, a. (2014). toward an instructionally oriented theory of example-based learning. cognitive science, 38(1), 1–37. https://doi.org/10.1111/cogs.12086 roll, i., holmes, n. g., day, j., & bonn, d. (2012). evaluating metacognitive scaffolding in guided invention activities. instructional science, 40(4), 691–710. https://doi.org/10.1007/s11251-012-9208-7 rotgans, j. i., & schmidt, h. g. (2014). situational interest and learning: thirst for knowledge. learning and instruction, 32, 37–50. https://doi.org/10.1016/j.learninstruc.2014.01.002 sailer, m., hense, j., mandl, h., & klevers, m. (2013). psychological perspectives on motivation through gamification. ixd&a, 19, 28–37. schank, r. c., berman, t. r., & macpherson, k. a. (1999). learning by doing. instructional-design theories and models: a new paradigm of instructional theory , 2, 161–181. schmidt, h. g., & boshuizen, h. p. a. (1993). on acquiring expertise in medicine. educational psychology review, 5(3), 205–221. https://doi.org/10.1007/bf01323044 schmidt, h. g., & rikers, r. m. j. p. (2007). how expertise develops in medicine: knowledge encapsulation and illness script formation. medical education, 41(12), 1133–1139. doi: 10.1111/j.1365-2923.2007.02915.x schrader, f.-w. (2009). anmerkungen zum themenschwerpunkt diagnostische kompetenz von lehrkräften [the diagnostic competency of teachers]. zeitschrift für pädagogische psychologie, 23(3), 237–245. https://doi.org/10.1024/1010-0652.23.34.237 schrader, f.-w. (2011). lehrer als diagnostiker [teachers as diagnosticians]. in e. terhart, h. bennewitz, & m. rothland (eds.), handbuch der forschung zum lehrerberuf (pp. 683–698). münster: waxmann. schwaighofer, m. (2015). kognitive basisfunktionen: rolle für den wissenserwerb und trainierbarkeit [basic cognitive functions: their role for knowledge acquisition and transer]. (text.phdthesis). ludwig-maximilians-universität münchen. retrieved from https://edoc.ub.uni-muenchen.de/17906/ schwaighofer, m., fischer, f., & bühner, m. (2015). does working memory training transfer? a meta-analysis including training conditions as moderators. educational psychologist, 50(2), 138–166. https://doi.org/10.1080/00461520.2015.1036274 schwaighofer, m., bühner, m., & fischer, f. (2016). executive functions as moderators of the worked example effect: when shifting is more important than working memory capacity. journal of educational psychology, 108(7), 982–1000. https://doi.org/10.1037/edu0000115 schwartz, d. l., & bransford, j. d. (1998). a time for telling. cognition and instruction, 16(4), 475–5223. https://doi.org/10.1207/s1532690xci1604_4 shannon, r. e. (1975). systems simulation: the art and science. prentice-hall. shavelson, r. j. (2013). on an approach to testing and modeling competence. educational psychologist, 48(2), 73–86. https://doi.org/10.1080/00461520.2013.779483 sherin, m. g., jacobs, v. r., & philipp, r. a. (eds.). (2011). mathematics teacher noticing: seeing through teachers' eyes. new york: routledge. shulman, l. s. (1987). knowledge and teaching: foundations of the new reform. harvard educational review, 57(1), 1–23. https://doi.org/10.17763/haer.57.1.j463w79r56455411 shulman, l. s. (2015). pck: its genesis and exodus. in re-examining pedagogical content knowledge in science education (pp. 3–13). new york: routledge. siebeck, m., schwald, b., frey, c., röding, s., stegmann, k., & fischer, f. (2011). teaching the rectal examination with simulations: effects on knowledge acquisition and inhibition. medical education , 45(10), 1025–1031. doi: 10.1111/j.1365-2923.2011.04005.x smith, s. j., & barry, d. g. (2013). the use of high-fidelity simulation to teach home care nursing. western journal of nursing research, 35(3), 297–312. https://doi.org/10.1177/0193945911417635 snow, r. e. (1991). aptitude-treatment interaction as a framework for research on individual differences in psychotherapy. journal of consulting and clinical psychology, 59(2), 205. http://dx.doi.org/10.1037/0022-006x.59.2.205 spinath, b. (2005). akkuratheit der einschätzung von schülermerkmalen durch lehrer und das konstrukt der diagnostischen kompetenz [accuracy of teacher judgments on student characteristics and the construct of diagnostic competence]. zeitschrift für pädagogische psychologie, 19 (1), 85–95. https://doi.org/10.1024/1010-0652.19.12.85 spiro, r. j., feltovich, p. j., jacobson, m. j., & coulson, r. l. (1992). cognitive flexibility, constructivism, and hypertext: random access instruction for advanced knowledge acquisition in ill-structured domains. in t. m. duffy & d. h. jonassen (eds.), constructivism and the technology of instruction: a conversation. (pp. 57–75). hillsdale, nj england: lawrence erlbaum associates, inc. stark, r., tyroller, m., krause, u.-m., & mandl, h. (2008). effekte einer metakognitiven promptingmaßnahme beim situierten, beispielbasierten lernen im bereich korrelationsrechnung [effects of a prompting intervention in situated, example-based learning in the domain of correlation]. zeitschrift für pädagogische psychologie, 22(1), 59–71. https://doi.org/10.1024/1010-0652.22.1.59 stanovich, k. (2011). rationality and the reflective mind. oxford university press. stark, r., kopp, v., & fischer, m. r. (2011). case-based learning with worked examples in complex domains: two experimental studies in undergraduate medical education. learning and instruction, 21(1), 22–33. https://doi.org/10.1016/j.learninstruc.2009.10.001 stegmann, k., pilz, f., siebeck, m., & fischer, f. (2012). vicarious learning during simulations: is it more effective than hands-on training? medical education, 46(10), 1001–1008. https://doi.org/10.1111/j.1365-2923.2012.04344.x stewart, a. c., williams, j., smith-gratto, k., black, s. s., & kane, b. t. (2011). examining the impact of pedagogy on student application of learning: acquiring, sharing, and using knowledge for organizational decision making. decision sciences journal of innovative education , 9(1), 3–26. https://doi.org/10.1111/j.1540-4609.2010.00288.x stürmer, k., seidel, t., & holzberger, d. (2016). intra-individual differences in developing professional vision: preservice teachers’ changes in the course of an innovative teacher education program. instructional science, 44(3), 293–309. https://doi.org/10.1007/s11251-016-9373-1 südkamp, a., kaiser, j., & möller, j. (2012). accuracy of teachers’ judgments of students’ academic achievement: a meta-analysis. journal of educational psychology, 104(3), 743–762. https://doi.org/10.1037/a0027627 südkamp, a., möller, j., & pohlmann, b. (2008). der simulierte klassenraum. eine experimentelle untersuchung zur diagnostischen kompetenz [the simulated classroom: an experimental study on diagnostic competence]. zeitschrift für pädagogische psychologie, 22(3–4), 261–276. tariq, m., & ali, s. a. (2013). clinical reasoning and dual mental processing in diagnostic competence. journal of the college of physicians and surgeons--pakistan: jcpsp , 23(10), 689–690. https://doi.org/10.2013/jcpsp.689690 tröbst, s., kleickmann, t., heinze, a., bernholt, a., rink, r., & kunter, m. (2018). teacher knowledge experiment: testing mechanisms underlying the formation of preservice elementary school teachers’ pedagogical content knowledge concerning fractions and fractional arithmetic. journal of educational psychology, 110(8), 1049-1065. doi: 10.1037/edu0000260 tschan, f., semmer, n. k., gurtner, a., bizzari, l., spychiger, m., breuer, m., & marsch, s. u. (2009). explicit reasoning, confirmation bias, and illusory transactive memory: a simulation study of group medical decision making. small group research. https://doi.org/10.1177/1046496409332928 van de wiel, m., boshuizen, h. p. a., & schmidt, h. g. (2000). knowledge restructuring in expertise development: evidence from pathophysiological representations of clinical cases by students and physicians. european journal of cognitive psychology, 12 (3), 323–355. doi: 10.1080/09541440050114543 van gog, t., & rummel, n. (2010). example-based learning: integrating cognitive and social-cognitive research perspectives. educational psychology review, 22(2), 155–174. https://doi.org/10.1007/s10648-010-9134-7 van merriënboer, j. j. g., & kirschner, p. a. (2018b). 4c/id in the context of instructional design and the learning sciences. in f. fisher, c. e. hmelo-silver, s. r. goldman, & p. reimann (eds.), international handbook of the learning sciences (pp. 169–179). new york: routledge. vanlehn, k. (1996). cognitive skill acquisition. annual review of psychology, 47(1), 513–539. https://doi.org/10.1146/annurev.psych.47.1.513 vogel, f., wecker, c., kollar, i., & fischer, f. (2017). socio-cognitive scaffolding with computer-supported collaboration scripts: a meta-analysis. educational psychology review, 29(3), 477–511. https://doi.org/10.1007/s10648-016-9361-7 wecker, c., & fischer, f. (2011). from guided to self-regulated performance of domain-general skills: the role of peer monitoring during the fading of instructional scripts. learning and instruction, 21(6), 746–756. https://doi.org/10.1016/j.learninstruc.2011.05.001 wecker, c., rachel, a., heran-dörr, e., waltner, c., wiesner, h., & fischer, f. (2013). presenting theoretical ideas prior to inquiry activities fosters theory-level knowledge. journal of research in science teaching, 50(10), 1180–1206. https://doi.org/10.1002/tea.21106 wissenschaftsrat. (2014). bedeutung und weiterentwicklung von simulation in der wissenschaft [importance of the development of simulations in science] . dresden. wood, d., bruner, j. s., & ross, g. (1976). the role of tutoring in problem solving. journal of child psychology and psychiatry, 17(2), 89–100. https://doi.org/10.1111/j.1469-7610.1976.tb00381.x woods, n. n. (2007). science is fundamental: the role of biomedical knowledge in clinical reasoning. medical education, 41 (12), 1173–1177. https://doi.org/10.1111/j.1365-2923.2007.02911.x wouters, p., van nimwegen, c., van oostendorp, h., & van der spek, e. d. (2013). a meta-analysis of the cognitive and motivational effects of serious games. journal of educational psychology, 105(2), 249. doi: 10.1037/a0031311 wouters, p., & van oostendorp, h. (2013). a meta-analytic review of the role of instructional support in game-based learning. computers & education, 60(1), 412-425. doi: 10.1016/j.compedu.2012.07.018 zimmerman, b. j. (2000). self-efficacy: an essential motive to learn. contemporary educational psychology, 25(1), 82–91. https://doi.org/10.1006/ceps.1999.1016 ziv, a., wolpe, p. r., small, s. d., & glick, s. (2003). simulation-based medical education: an ethical imperative. academic medicine: journal of the association of american medical colleges , 78 (8), 783–788. doi: 10.1097/00001888-200308000-00006 frontline learning research vol.5 no. 2 (2017) 36-59 issn 2295-3159 corresponding author: julia morinaj, institute of educational science, university of bern, fabrikstrasse 8, 3012 bern, switzerland, iuliia.morinaj@edu.unibe.ch doi: http://dx.doi.org/10.14786/flr.v5i2.298 school alienation: a construct validation study julia morinaja, jan scharfb, alyssa grecub andreas hadjarb, tina haschera, kaja marcina auniversity of bern, switzerland buniversity of luxembourg, luxembourg article received 27 march / revised 8 june / accepted 16 july / available online 25 july abstract early identification of school alienation is of great importance for students’ educational outcomes and successful participation in society. this study examined the psychometric characteristics of a newly developed assessment instrument, the school alienation scale (sals), to measure school alienation among primary and secondary school students. the sals consists of three school-related domains, namely, classmates, teachers, and learning. based on the responses of swiss (1) and luxembourgish (2) students from two schoolspecific cohorts — primary (grade 4; n1=486, n2=503) and secondary schools (grade 7; n1=550, n2=534), we assessed instrument reliability, validity, and cross-cultural equivalence. the scale showed evidence of reliability and internal validity across two samples, confirming that the hypothesized first-order three-factor model fits the data better than several alternative models. the results of measurement invariance tests revealed that the measurement model operated equally well for primary and secondary school students in both countries. the construct validity of the sals was additionally supported by demonstrated criterion-related validity. specifically, school alienation domains were negatively associated with positive attitudes to and enjoyment in school; social problems in school were positively related to alienation from classmates and teachers. our key contributions to the measurement of school alienation are the disclosure of the core domains of school alienation, development of a reliable and valid instrument, and justification for its use. therefore, the results of this study have important implications for further theoretical work in alienation research and contribute to comparative research by examining the construct of school alienation in different educational settings. keywords: school alienation; construct validity; criterion validity; measurement invariance mailto:iuliia.morinaj@edu.unibe.ch http://dx.doi.org/10.14786/flr.v5i2.298 morinaj et al. 37 | f l r 1. introduction education crucially supports successful participation in society and develops the existing body of knowledge. a knowledge base and skills acquired through schooling provide individuals opportunities to act effectively in a rapidly changing world. additionally, only a well-educated population can contribute significantly to the community and the economy (becker, 1994; seetanah, 2009; vila, 2000; zhang & zhuang, 2011). for these reasons, societies have a genuine interest in providing young people various educational opportunities. however, children’s learning begins when they are born (krumboltz, 2009) when they start observing their environment, and it continues throughout their lifetime. indeed, the majority of young children come to school filled with curiosity, creativity, and a strong desire to learn (lumsden, 1994). however, there is substantial evidence that students’ intrinsic academic motivation and interest in learning at school significantly decline over time (eccles & midgley, 1990; gottfried, fleming, & gottfried, 2001; oecd, 2004). not only do these processes co-occur with considerable social, physical, cognitive, emotional, and behavioral changes during adolescence (eccles, brown, & templeton, 2008; schunk & meece, 2005), but they are also associated with significant changes in family relations, peer affiliation, school and home environments (schunk & meece, 2005). such shifts can be accompanied by school alienation, delinquency, and dropping out of high school (eccles & gootman, 2002), pointing to a mismatch between adolescents’ needs and their environments (archambault, janosz, morizot, & pagani, 2009b; eccles & midgley, 1989). in other words, adolescents whose environments do not fulfill their needs are more likely to become psychologically and physically disengaged, and eventually alienated from school (eccles & roeser, 2009; gutman & eccles, 2007). the rising interest in school alienation has led to a better understanding of the phenomenon per se (hadjar & lupatsch, 2010; hascher & hadjar, 2017; hascher & hagenauer, 2010). much less attention, however, has been paid to the development and validation of instruments that measure school alienation. to ensure meaningful inferences from a theoretical construct and to justify instrument use for further research and praxis, it is of utmost importance to carefully design and validate the construct of school alienation (arnold, arad, rhoades, & drasgow, 2000; clark & watson, 1995). moreover, measurement validation is inevitably bound to theory development and theory testing (zumbo, 2009). thus, construct validity, based on theoretical and statistical evidence, lies at the heart of the present study. in this study, we introduce a new multidimensional perspective on school alienation. in particular, we developed the school alienation scale (sals) as an assessment instrument used in the binational research project, school alienation in switzerland and luxembourg (sasal), to identify students’ negative attitudes toward classmates, teachers, and learning during primary and secondary education. such measurement enables gathering important data about the issues students face in daily school life. unlike previous studies on school alienation (e.g., martin, 2008; newmann, 1992; rovai & wighting, 2005), which consider a single academic setting (e.g., american secondary schools, virtual classrooms), we examined the development of school alienation in different educational settings — switzerland and luxembourg. the selection of these countries, which have some differences and share some similarities, is meaningful from the perspective of comparative educational research. in addition, comparative research enables testing the instrument in various settings and provides additional evidence of construct validity. the main objectives of this research are to gain a thorough understanding of the phenomenon of school alienation and to develop and validate a new theoretically-based measure of school alienation. here, we first describe the major concepts of school alienation, followed by the guiding theoretical issues regarding instrument construction. second, we provide data on instrument reliability, validity, and crosscultural equivalence, based on research conducted in luxembourg and the swiss canton of bern. parallel analyses were conducted across primary and secondary school students and across countries to demonstrate the replicability of findings. finally, we discuss the study’s results and implications for future research. morinaj et al. 38 | f l r 1.1 conceptualization of school alienation the term alienation comes from the latin alienatus, meaning “estranged,” which in turn originated from alius, meaning “other” or “another” (watt, 2000). the concept of alienation has been introduced from sociological, psychological, philosophical, theological, and historical perspectives and applied to various contexts. for example, marx primarily focused on the concept of alienation in the economic system. moreover, subsequent alienation research has shown that the problem of alienation also arises in highly centralized and formalized organizations (blauner, 1964; aiken & hage, 1966), family life (kelly & johnston, 2001), work settings (hirschfeld & feild, 2000; shantz, alfes, & truss, 2014) as well as religious (exline, yali, sanderson, 2000), political (finifter, 1970; seeman, 1975; pantoja & segura, 2003), and medical contexts (young, 1984). although alienation appears to be a strictly contextual construct (safipour, schopflocher, higginbottom, & emami, 2011), at a general level it may be characterized as a kind of estrangement, distancing, or separateness from a former or a normal state, leading to some sort of loss (railton, 2013). so far, relatively little is known about alienation in educational settings. however, recent studies essentially contributed to the development of school alienation research, specifying the characteristics of educational and social learning environments necessary for the prevention of school alienation (dickey, 2004; hadjar, backes, & gysin, 2015; hadjar & lupatsch, 2010; hascher & hadjar, 2017; hascher & hagenauer, 2010; osin, 2009). oftentimes, youngsters who like school at the beginning may later become estranged from learning and develop negative attitudes toward school, hurtling into school alienation, and over time, even into dropping out (archambault et al., 2009b; eccles & alfeld, 2007). different students, whether they study in elite private or inner-city schools, undoubtedly experience the same problem (sidorkin, 2004). moreover, it is especially in adolescence that youth often deal with numerous stressful situations at school (safipour et al., 2011). for example, students who see little practical value in learning and its relevance outside of school, who experience poor relationships with teachers, or who suffer because of classmates may feel like outsiders. in the framework of the sasal project, we conceptualize this feeling of estrangement from the social (i.e., classmates and teachers) and academic aspects of schooling (i.e., learning), including cognitive and emotional components, as school alienation (hascher & hadjar, 2017). in a school mostly free of alienation, we would expect students to maintain courteous relationships with classmates and school staff members and to engage actively and meaningfully in classroom and school activities. suffice it to say that educational alienation cannot be dismissed as a temporary aberration, as it is rather a part of education (sidorkin, 2004). there is a strong likelihood that adolescents who are alienated from school will not accomplish their basic educational goals (archambault, janosz, fallu, & pagani, 2009a). as a result, alienated students leave school with numerous negative experiences, including deviant behaviors, difficulties in fitting in, low participation in school activities, depression, running away, early sexual activity, conflicts with families, school withdrawal, limited education, failed attachment to school, or a lack of interest in further academic qualification (brown, higgins, & paulsen, 2003a; farrow, 1991; frey, ruchkin, martin, & schwab-stone, 2009; hascher & hagenauer, 2010; tarquin & cook-cottone 2008). school alienation can therefore lead to a process of exclusion from a society that is increasingly based on learning. hence, the early diagnosis and understanding of school alienation are of great importance for the community, including educators, school staff members, and parents (brown et al., 2003a; stamm, kost, suter, holzinger, & stroezel, 2011). currently, most measurements of school alienation draw upon some of the alienation categories addressed by seeman (1959) and dean (1961), including powerlessness, meaninglessness, normlessness, isolation, and self-estrangement (e.g., brown, higgins, pierce, hong, & thoma, 2003b; mau, 1992). for example, recent studies explicitly focused on the four alienation dimensions most closely associated with the school context, namely, powerlessness, meaninglessness, normlessness, and social estrangement (brown et al., 2003b; çağlar, 2013; kocayörük & simsek, 2015; mau, 1992). similarly, another study developed an alienation-based framework for student experience in higher education, concentrating on the three categories, such as powerlessness, meaninglessness, and self-estrangement (barnhardt & ginns, 2014). despite the relative popularity of recognizing the above-mentioned categories as an alienation construct, only a very morinaj et al. 39 | f l r limited number of studies have measured it (brown et al., 2003b). some researchers have addressed the phenomenon of school alienation by investigating its correlates. correlations are demonstrated repeatedly, for example, between school alienation and student participation (altenbaugh, engel, & martin, 1995; carlson, 1995; newmann, 1992), between school alienation and academic achievement (johnson, 2005; reinke & herman, 2002), between school alienation and teacher and peer support (altenbaugh et al., 1995; ghaith, shaaban, & harkous, 2007; hascher & hagenauer, 2010), and between school alienation and school withdrawal (liu, 2010). murdock (1999) has focused on the behavioral aspect of school alienation, based on engagement in school tasks and self-reported disciplinary problems. further, the phenomenon of school alienation was also studied in the context of motivation research, especially in regard to student (dis)engagement. some researchers used the term disengagement as a synonym for school alienation (altenbaugh et al., 1995), although later they described a vicious cycle in which alienation encourages further disengagement leading to more alienation. alienation was also described as a subdimension of academic amotivation (legault, green-demers, & pelletier, 2006). another empirical study viewed engagement as a positive counterpart to alienation (case, 2008). however, the shortcomings of those approaches lie in the finding that the two constructs are not symmetrically opposite (schabracq & cooper, 2003). a variety of viewpoints in alienation research indicates a lack of consistency among the operationalizations of school alienation. in accordance, it may not be surprising that there is little consistency in its measurement (barnhardt & ginns, 2014; brown et al., 2003b; çağlar, 2012; carlson, 1995; hyman, cohen, & mahon, 2003; johnson, 2005; mau, 1992; murdock, 1999). although most instruments recognized the multidimensional nature of school alienation, it is still obscure whether school alienation is a general or rather a domain-specific construct and what are its key elements. moreover, there is a lack of information regarding the validity and reliability of existing instruments that measure school alienation. the identification of factors underlying student feelings of school alienation and alleviation of unfavorable consequences is important for understanding the true impact alienation may have on school success. in the alienation literature, several reasons of school alienation are discussed. oftentimes, student alienation is conceptualized as the result of three experiences. first, school subjects appear irrelevant to some students and they rarely see the connection between learning at school and their social realities (çağlar, 2013). as a result, they estrange themselves from the learning process and develop negative attitudes toward the school (altenbaugh et al., 1995). second, teachers behave differently toward students in the same classroom (babad, 1992; hughes, gleason, & zhang, 2005; jussim & eccles, 1992; jussim, eccles, & madon, 1996; mckown & weinstein, 2008). for example, teachers tend to provide more emotional support, more favorable feedback, and more challenging learning opportunities to high achievers. at the same time, they expect troublesome behavior from difficult students and impose more restrictive actions upon those students (baker, 1998). accordingly, as a reaction, students may respond to these disciplinary actions with more disobedience which in turn may lead to further alienation and distancing from teachers (kagan, 1990). third, students feel not accepted by their classmates, experiencing non-fulfilment of their psychological and social needs (baker, 1998). peer rejection can result in student academic and socioemotional misfortune in a school environment as well as a lack of identification with the school system (buhs & ladd, 2001; ladd, 1999; ladd, birch, & buhs, 1999). in other words, alienated students may withdraw from peers, teachers, learning, and eventually school itself (altenbaugh et al., 1995). research on student withdrawal has also distinguished between the academic (e.g., learning, academic performance) and social (e.g., interaction with classmates and teachers) domains of the educational institution, suggesting that despite close interrelation between the two domains, a student may be integrated into one domain without being sufficiently integrated into the other (tinto, 1975, 1993). these theoretical propositions suggest that in a school setting students can be alienated in academic and/or social domains. considering the multidimensional nature of the school alienation phenomenon and that attitudes are usually directed toward specific situations, objects, or behavior (honkanen, verplanken, & olsen, 2006), this study suggests a more holistic approach. specifically, we propose that school alienation consists of three school-related domains — classmates, teachers, and learning — considering them as interrelated but relatively independent dimensions. in other words, we address school alienation as a construct of a domain-specific morinaj et al. 40 | f l r nature. accordingly, we developed the school alienation scale (sals) as a carefully constructed, comprehensive, psychometrically sound, and brief standardized assessment instrument that can be used to diagnose students’ negative attitudes toward classmates, teachers, and learning during primary and secondary education. more precisely, it measures whether students are alienated from classmates, and/or teachers, and/or learning at school. students may become alienated in one, two, or all three domains. in the latter case, students experience negative attitudes toward both social actors and academic aspects of schooling. 1.2 rationales for questionnaire construction prior to constructing the sals to assess student alienation from school, we articulated several basic principles. first, we conceptualized the target construct to assure the coherence of the initial item pool. to be specific, the three school alienation dimensions of classmates, teachers, and learning comprised both emotional (i.e., a student’ feelings toward school) and cognitive (i.e., a student’s beliefs, perceptions, knowledge, assumptions, and judgments of school) aspects. second, we specified the association between these aspects, taking into account that historically, cognitive and emotional processes were studied separately (hilgard, 1980; snow, corno, & jackson, 1996). for years, psychologists have been trying to solve the “chicken and egg” causality dilemma. however, recent research is based on the idea that cognition cannot be separated from individual’s emotions and feelings; it tends to examine them as interwoven psychological processes, one leading to another, in an ongoing cycle (dai & sternberg, 2004; dweck, mangels, & good, 2004; eisenberg, 2014; linnenbrink & pintrich, 2004; pessoa, 2008). there is scientific evidence that positive emotions stimulate rigorous thinking, facilitating application of existing knowledge (isen, 2004). moreover, emotions are essential not only in the relationship between individuals and their environment, but also in group processes, such as group members (for example, a school class) sharing similar emotional experiences (aritzeta et al., 2016). when members of a class feel they are a part of the class, they are most likely to pursue common goals (e.g., learning, getting good marks, or graduating). kreitler (2013) likewise perceived emotions as motivational forces that influence activation and functioning of cognition. in other words, emotion and cognition are different sides of the same coin, working jointly in the process of intellectual functioning (dweck et al., 2004; wimmer, 2013; zihl, szesny, & nickel, 2013), and do not act independently of each other (kreitler, 2013). therefore, in this study we scrutinized the complex interplay between emotion and cognition in alliance, and emphasized the difficulty of setting boundaries between them. furthermore, we also decided whether school alienation should be measured with general or subjectspecific items. recent studies addressed school alienation not as subject-specific, but rather as a general negative orientation toward social actors and/or learning in school (hadjar & lupatsch, 2010; hascher & hagenauer, 2010; hadjar, backes, & gysin, 2015). consequently, in this study we decided to measure school alienation at a more general level. nevertheless, we acknowledge that students’ feelings and beliefs may vary depending on school subject (goetz, frenzel, pekrun, & hall, 2006). 1.3 aim and hypotheses the current study was designed to illustrate psychometric characteristics of a newly developed selfreport instrument to assess school alienation among primary and secondary students, based on certain theoretical and conceptual considerations. validity theorists emphasize the importance of constructing a cohesive validity argument that integrates several sources of evidence to support the construct validity and use of the instrument (aera, apa, & ncme, 2014). the unitary concept of validity still remains a widely used approach (chan, 2014; sireci, 2009). therefore, we sought to provide a consistent approach for examining the construct validity of the sals. we first assessed the reliability of the scale in terms of internal consistency of the items. we then explored the factorial structure of the sals and tested several competing models of school alienation by means of confirmatory factor analysis. next, we estimated measurement invariance across two cohorts, two countries, and gender to ensure that the instrument possessed the same psychometric morinaj et al. 41 | f l r characteristics across various groups. finally, we assessed criterion-related validity, as an additional evidence of construct validity, to investigate the relationship between the sals and external criteria. with respect to a suggested construct validity approach, we tested the following hypotheses. recognizing the multidimensional nature of school alienation, we assumed the sals developed for diagnosing student alienation from classmates, teachers, and learning to be reliable (hypothesis 1). regarding the factorial structure of the sals, our proposition that students can be alienated from classmates, teachers, or learning at school implied that the instrument should depict a three-factor structure. hence, the measurement model tested here hypothesized that responses to the sals could be explained by three first-order factors — classmates, teachers, and learning (see figure 1, model 3). it was expected that the specified model could be verified and fit the data better than several alternative models (hypothesis 2). furthermore, with the goal of following individuals over time and comparing groups, an instrument should measure the same construct having identical structure across various groups. therefore, we assessed measurement invariance across different contexts (i.e., switzerland and luxembourg), grades (i.e., grade 4 and grade 7), and gender. by establishing measurement invariance, we provide additional evidence of construct validity (van de schoot, lugtig, & hox, 2012). we assumed the questionnaire was equally suitable for both primary and secondary school students (hypothesis 3) in both switzerland and luxembourg (hypothesis 4) and for both boys and girls (hypothesis 5). however, it is not sufficient to provide only measurement invariance for evaluating validity of the measure. criterion-related validity was examined as important evidence of construct validity (ariño, 2003; messick, 1995; vandenberg & lance, 2000). the validity of the measurement is also indicated by the correspondence of the relationships between the new measure and other variables to theoretical expectations. together with the sals, we administered other scales, including the well-being in school scale (hascher, 2007). in particular, we examined associations between the domains of school alienation and demonstrably valid dimensions of well-being (hascher, 2008; hascher & hagenauer, 2011). in keeping with previous research on the association of alienation and well-being (crinson & yuill, 2008; dekel & tuval-mashiach, 2012; hall-lande, eisenberg, christenson, & neumark-sztainer, 2007; ifeagwazi, chukwuorji, & zacchaeus, 2015; moreno, & de roda, 2003; osin, 2009; vahedi & nazari, 2011), we expected positive attitudes to school and enjoyment in school to be significantly negatively related to the domains of school alienation (hypothesis 6). prior research has shown that when students experience positive emotions during school-related activities and interaction with people involved in those activities their alienation decreases (ifeagwazi et al., 2015). wellbeing is, thus, viewed as a resource for coping with negative impacts on learning and individual development (hascher, 2011, 2012). this implies that social problems in school should be positively related to alienation from classmates and teachers (hypothesis 7). in accordance with hascher (2010), social discrimination in school was found to diminish student well-being and escalate social problems in school. accordingly, the present study was conducted to accomplish three major goals. first, given little agreement on the conceptualization of school alienation, we sought to identify the key elements of school alienation. second, considering generally acknowledged multidimensional nature of school alienation, we wanted to verify whether school alienation is a general or a domain-specific construct. in addition, we analyzed the interplay between emotion and cognition. third, we also sought to empirically validate the new theoretically-based instrument to measure school alienation. a carefully designed and psychometrically sound research instrument, based on theory and statistical analysis, would make a valuable contribution to alienation research from theoretical, practical as well as methodological standpoints. morinaj et al. 42 | f l r 2. method 2.1 participants and procedures data were collected from two cohorts of primary and secondary school students, attending grade 4 and grade 7, respectively. the sample in this study (n = 2,073) consisted of 486 students from primary school (47.3% male; mage= 10.3 years [sd = .98]) and 550 students from secondary school (45.2% male; mage= 13.0 years [sd = .55]) from the swiss canton of bern; and 503 primary school students (54.9% male; mage= 9.7 years [sd = .75]) and 534 secondary school students (57.8% male; mage= 12.7 years [sd = .65]) from luxembourg. fifty percent of the students in the swiss canton of bern and about 75% of students in luxembourg had a migration background (including firstand second-generation immigrants). both countries have stratified education systems, offering track selection at the secondary level (edk, 2015; menje & university of luxembourg, 2015). the majority of secondary school students in switzerland studied in schools with various school tracks and ability groups. just over half of secondary school students (55%) studied in the middle track (sek), 36% in the lower track (real), and 8% in the upper track (spezsek). on average, primary school students perceived themselves as having good academic performance (mperf = 4.32 [sd = .46] on a scale from 1 = poor to 5 = very good). secondary school students defined themselves as having above-average academic performance (mperf = 3.69 [sd = .57] on a scale from 1 = poor to 5 = very good). for the luxembourgish sample, the largest percentage of students (34%) were from the general secondary track (enseignement secondaire), 25.3% from the technical secondary track (enseignement secondaire technique), 23.2% from the lowest technical secondary track (régime préparatoire/modulaire), and 17% were from the project track proci (projet pilote “cycle inferieur” de l’enseignement secondaire technique), a comprehensive track within technical secondary education with heterogeneous ability levels. primary school students, on average, perceived themselves as having good academic performance (mperf = 4.05 [sd = .53]), whereas secondary school students described themselves as having above-average academic performance (mperf = 3.72 [sd = .58]). contacts to schools were established through school principals and teachers, who were provided general information about the project, its goals, and benefits of participating. in the former case, principals decided together with teachers about their participation in the study. in the latter case, we contacted teachers directly. once agreement was reached, we informed school principals about teachers’ consent to engage in a project. random sampling of schools or classes was not possible, because of a high risk of dropout. a main sampling criterion was heterogeneity in regard to institutional school characteristics, school composition (migrant population, ability levels), and industrial versus rural areas. in collaboration with teachers, students received informed consent forms addressed to their parents, indicating voluntary participation, assurance of anonymity, and confidentiality. small incentives were given to participants during school break time and after the survey, because incentives were found to motivate respondents and thus increase the response rate (singer & couper, 2008). 2.2 measures the new paper-and-pencil school alienation questionnaire was developed by applying a wellestablished systematic framework designed to produce reliable and valid scales (hinkin, tracey, & enz, 1997). the first crucial step in scale construction involved the creation of the initial pool of items meant to cover each content area relevant to the target construct and the sample. the items were developed by experienced theorists and researchers in the field based on the existing instruments (hadjar & lupatsch 2010; hascher & hagenauer 2010). the wording of each item was thoroughly discussed, to eliminate potential sources of constructirrelevant variance. a typical item represented a statement in regard to general positive or negative student feelings and thoughts toward classmates, teachers, and learning (e.g., “what we learn in school is boring”). morinaj et al. 43 | f l r prior to application, a pilot study was conducted to test the developed instrument in luxembourg and the swiss canton of bern. the pilot study examined how well the new measure reflected our expectation regarding its psychometric properties, and evaluated whether the construct of interest operated equivalently across two groups. after an initial assessment of the scale’s reliability and validity, irrelevant, similar, or ambiguous items were eliminated from the study (e.g., “when i get bad grades, i do not feel good”, “i wonder why working with a partner and group work are meaningful”). accordingly, the final version of the sals (consisting of 39 items, with 12-14 items per school alienation domain) was administered to assess school alienation among primary and secondary school students. a typical item represented a statement in regard to student feelings and thoughts toward classmates, teachers, and learning. an example of an item of alienation from the classmates scale is: “in my class i feel like someone who doesn’t fit in”; alienation from the teachers scale: “i do not feel taken seriously by my teachers”; alienation from the learning scale: “i don’t find pleasure in learning at school”. for each item, students responded on a 4-point likert scale with the endpoints 1 = disagree to 4 = agree. we chose this scale format to reach a specific respondent opinion; therefore, we eliminated the mid-point on the scale. in addition, shorter scales are relatively quick to use (preston & colman, 2000). however, prior research has shown that there is often a give-and-take between scale reliability and ease of administration (as cited in østerås et al., 2008). several items in the survey were phrased in the reverse to force participants to carefully read the questions and to prevent the tendency to agree with a statement or respond in the same pattern throughout the questionnaire. 3. results 3.1 descriptive statistics and reliability analysis prior to conducting statistical analyses, we calculated the amount of missing data across the variables analyzed in this study to ensure robust data analysis. the largest amount of missing data was less than 3%, which meant that accurate statistical estimates could be obtained (acuña & rodriguez, 2004). the reliability of the sals was measured in terms of the internal consistency of the items and analyzed by calculating cronbach’s alpha (hypothesis 1), using the statistical package for the social sciences (spss) software version 23. using subsamples from primary and secondary school students in the swiss canton of bern, we conducted a reliability analysis by means of stepwise exclusion of the items with the lowest corrected item-total correlation. we then applied a series of factor analyses for all four subsamples to verify the factorial structure of the scale. based on factor analysis and internal consistency of the items, 24 items, with 8 items per each school alienation domain, were selected for the final scale. the descriptive statistics for these 24 items and the school alienation scales are shown in table 1. the cronbach’s alpha coefficients ranged from α = .74 to α = .88 for all study participants. the results of the reliability analyses confirmed hypothesis 1. table 1 descriptive statistics of school alienation scales canton of bern primary school (n = 486) secondary school (n = 550) m sd α s(se) k(se) m sd α s(se) k(se) classmates 1.48 .43 .79 1.27(.11) 1.45(.22) 1.52 .46 .83 1.45(.11) 2.50(.21) teachers 1.42 .45 .77 1.58(.11) 3.29(.22) 1.57 .48 .79 1.11(.11) 1.39(.21) learning 1.54 .52 .87 1.34(.11) 2.14(.22) 1.83 .56 .88 1.13(.11) 1.63(.21) luxembourg primary school (n = 503) secondary school (n = 534) m sd α s(se) k(se) m sd α s(se) k(se) classmates 1.56 .50 .74 1.29(.11) 1.58(.22) 1.60 .52 .84 1.25(.11) 1.30(.21) teachers 1.61 .59 .79 1.10(.11) 0.89(.22) 1.77 .59 .83 0.90(.11) 0.66(.21) learning 1.54 .56 .82 1.49(.11) 2.23(.22) 1.87 .60 .86 0.65(.11) 0.17(.21) note. m = mean; sd = standard deviation; α = cronbach’s alpha; s = skewness; k = kurtosis; se = standard error. morinaj et al. 44 | f l r intercorrelations of the school alienation factors for primary and secondary school students in both countries are reported in table 2. correlation coefficients displayed a positive association among the three domains of school alienation and represented small to high effects, indicating that the factors were conceptually and statistically distinguishable. table 2 correlations among the domains of school alienation canton of bern luxembourg 1 2 3 1 2 3 1. classmates .37**(.45**) .23**(.23**) .35** (.39**) .24** (.27**) 2. teachers .44**(.55**) .50**(.56**) .38** (.48**) .49** (.56**) 3. learning .25**(.29**) . 40** (.44**) .27** (.33**) .51** (.57**) note. values below the diagonal represent intercorrelations for primary school students (n = 486 students, canton of bern; n = 503 students, luxembourg) and values above the diagonal represent intercorrelations for secondary school students (n = 550 students, canton of bern; n = 534 students, luxembourg). latent correlations revealed by cfas are presented in parentheses. **p < .01. 3.2 factorial structure of the sals prior to analyzing the factorial structure of the sals, the data were screened to assure their appropriateness for confirmatory factor analysis (cfa) and the statistical estimator. the results from a shapiro-wilk test of normality, which considers the values of both skewness and kurtosis simultaneously, revealed that the data significantly deviate from a normal distribution (p < .05). some indicators demonstrated substantial non-normality, thereby weakening the assumption of multivariate normality. having non-normal continuous indicators, we applied an estimator with robust standard errors and a chi-square test statistic, i.e., robust maximum likelihood estimation method (mlr), to achieve reliable statistical results (brown, 2015). this allows evaluating models with missing data and provides parameter estimates, which are robust to nonnormality (brown, 2015). all analyses were based on covariance structures and applied robust maximum likelihood estimation. by applying cfa using software package mplus version 7.31, we assessed the internal validity of the sals. in other words, we explored whether the sals combined the variables that were collinear and belonged to the same underlying variable due to their mutual association. we estimated a range of the most significant fit indices as measures of model fit (schermelleh-engel, moosbrugger, & müller, 2003), including the chisquare test statistic, the root mean square error of approximation (rmsea), the comparative fit index (cfi), and the standardized root mean square residual (srmr). the following typical criteria were applied to assess the adequacy of model fit: cfi of close to .95, rmsea and srmr of close to .05 (hu & bentler, 1999; schermelleh-engel et al., 2003), and df < 2 (hair, anderson, tatham, & black, 1995; schermelleh-engel et al., 2003). in addition, we performed the chi-square difference tests to compare several competing models against each other. in the first confirmatory factor analysis, we assessed the hypothesized measurement model (see figure 1, model 3). it incorporated the three factors of school alienation, i.e., classmates, teachers, and learning, each consisting of 8 items. within each domain, several items were used to indicate an emotional component and other items were used to indicate a cognitive component. the three factors were free to correlate. as can be seen in table 3, the hypothesized model indicated a very good fit for primary as well as secondary school students in both luxembourg and the swiss canton of bern. completely standardized item loadings on the three factors of school alienation were all statistically significant and above .50 (primary school in both morinaj et al. 45 | f l r countries) and .57 (secondary school in both countries). latent correlations among the three school alienation factors were moderate, reaching on average .50 for all subsamples. figure 1. first-order one-factor (model 1), second-order with three facets included (model 2), hypothesized first-order three-factor (model 3), and second-order with six facets included (model 4) models potentially underlying the sals. 3.3 competing models of school alienation facilitating the assessment of the factorial validity of the sals, we evaluated three alternative models against the hypothesized model 3 (see figure 1, models 1, 2, 4). the models were subsequently compared in regard to their model fit by applying chi-square difference testing using the satorra-bentler scaled chi-square, in which the usual chi-square statistic is divided by a scaling correction to better approximate chi-square under non-normality (brown, 2015). in model 1, we assessed whether the data could be better described by an overall construct of school alienation in which no factors of school alienation were assumed. in model 2, we specified that three school alienation domains load on one higher-order school alienation factor. model 4 postulated that emotional and cognitive components comprised each of the three school alienation domains. as can be seen in table 3, none of the competing models provided a better fit to the data than the hypothesized first-order three-factor model for both primary and secondary school students in either luxembourg or the swiss canton of bern. in both groups, chi-square difference tests revealed that the comparison between model 1 and model 3 was highly significant (p < .01), indicating that model 3 fits the data better than model 1. comparing model 2 and model 3, in all subgroups except for primary school in the canton of bern, the chi-square difference test was statistically non-significant; this indicated that both models fit statistically equally well. the corrected difference test between model 3 and model 4 generated statistically model 1 model 2 model 3 model 4 morinaj et al. 46 | f l r insignificant values (p > .01), indicating that the fit of the restricted model (model 3) was not significantly worse than the fit of the unrestricted model (model 4; see table 3). in other words, both models function equally across the swiss and the luxembourgish primary and secondary school samples. thus, the null hypothesis of equal fit for both models cannot be rejected. in addition, a confirmatory factor analysis of model 4 revealed latent correlations between emotional and cognitive components in each of the three domains of school alienation. these correlations almost reached a perfect positive correlation (above .90), which indicates that both components undergo the same directional change. on the basis of these findings and theoretical considerations, we believe that emotional and cognitive components could be viewed as a single factor, as specified in model 3. in keeping with hypothesis 2, the hypothesized model 3 is verified and fits the data better than alternative models. table 3 goodness of fit indices for the four models of the sals note. see models in figure 1. cfi = comparative fit index; rmsea = root mean square error of approximation; srmr = standardized root mean squared residual; δ and δdf = changes in chi-square and degrees of freedom between the hypothesized and competing models. canton of bern model  df df cfi rmsea srmr model comparison δ δdf p primary school (n = 486) model 1 996.28 246 4.05 .70 .08 .10 model 2 417.11 244 1.71 .93 .04 .06 2 vs. 3 25.91 1 <.01 model 3 401.29 243 1.65 .94 .04 .05 3 vs. 1 319.84 3 <.01 model 4 401.44 242 1.66 .94 .04 .05 4 vs. 3 0.46 1 ns secondary school (n = 550) model 1 1377.88 246 5.60 .69 .09 .11 model 2 471.49 244 1.93 .94 .04 .05 2 vs. 3 0.27 1 ns model 3 471.56 243 1.94 .94 .04 .05 3 vs. 1 597.97 3 <.01 model 4 468.28 242 1.94 .94 .04 .05 4 vs. 3 2.95 1 ns luxembourg model  df df cfi rmsea srmr model comparison δ δdf p primary school (n = 503) model 1 825.85 246 3.36 .73 .07 .08 model 2 360.34 244 1.48 .95 .03 .05 2 vs. 3 0.30 1 ns model 3 360.56 243 1.48 .95 .03 .05 3 vs. 1 197.72 3 <.01 model 4 356.80 242 1.47 .95 .03 .05 4 vs. 3 2.92 1 ns secondary school (n = 534) model 1 1465.46 246 5.96 .65 .10 .11 model 2 451.97 244 1.85 .94 .04 .05 2 vs. 3 0.14 1 ns model 3 452.10 243 1.86 .94 .04 .05 3 vs. 1 789.81 3 <.01 model 4 447.44 242 1.85 .94 .04 .05 4 vs. 3 4.05 1 ns morinaj et al. 47 | f l r 3.4 measurement invariance pursuing the goals of evaluating the recently constructed research instrument and facilitating crosscultural research, we tested measurement invariance across primary and secondary school students in luxembourg and the swiss canton of bern (hypotheses 3 and 4) using multigroup confirmatory factor analyses. it is broadly recognized as one of the most successful and multilateral approaches to testing measurement invariance across specific groups of interest (steenkamp & baumgartner, 1998). examining measurement invariance through a series of models, we investigated whether the instrument possesses the same psychometric characteristics across various groups (milfont & fischer, 2010; schmitt & kuljanin, 2008). this procedure included testing hierarchically arranged models: (a) a configural invariance model, (b) a metric invariance model, (c) a scalar invariance model, and (d) a residual invariance model. first, configural invariance indicated whether respondents from luxembourg and the swiss canton of bern considered the construct of interest in a similar fashion, assuming the same item to be associated with the same latent factor in each group (milfont & fischer, 2010). in this model, none of the parameters were constrained to be equal across the groups. all subsequent models in testing for invariance were compared against the fit of the configural model, serving as a baseline value (byrne, 2012). at each step of the measurement invariance test procedure, earlier specified constraints remained in force. in the next step, metric invariance was required to test the equivalence of item factor loadings across the two groups, namely, whether participants in different groups assigned the same meaning to the items of a latent construct. previous research has shown that at least partial metric invariance must be confirmed prior to conducting the subsequent tests (vandenberg & lance, 2000). furthermore, scalar invariance, constraining item intercepts to be the same across groups, was a prerequisite for the comparison of latent means between groups. finally, having all previous constraints in place, we applied residual invariance to test the equivalence of measurement error for each item between groups (byrne, 2012; chen, sousa, & west, 2005; milfont & fischer, 2010). following the recommendations for testing measurement invariance with an adequate sample size that is equal between the groups (n > 300; chen, 2007), we compared the three fit indices (i.e., cfi, rmsea, and srmr) across all invariance models. in compliance with the suggested guidelines, for testing metric invariance, a change of ≥ -.010 in cfi, complemented by a change of ≥ .015 in rmsea, or a change of ≥.030 in srmr indicates noninvariance; for testing scalar or residual invariance, a change of ≥ -.010 in cfi, in addition to a change of ≥ .015 in rmsea, or a change of ≥.010 in srmr indicates noninvariance. in both primary and secondary school samples, δcfi, δrmsea, and δsrmr were evidently below the recommended cutoff values, and each of the four invariance models produced an acceptable fit (see table 4). the chi-square difference test between model 3 and model 2 in both samples was significant (δ δdf p δ δdf prespectively). given that the test was based on large samples (nprim = 989; nsec = 1,084) and the decrease in cfi was below .01, we concluded that there was no substantial difference between luxembourgish and swiss students on the intercepts of the measured variables (see ong & van dulmen, 2006). based on the three indicators, metric, scalar, and residual invariance between the luxembourgish and swiss primary and secondary school students was confirmed, revealing the equivalence of loadings, intercepts, and residuals across the two countries. consistent with hypotheses 3 and 4, results suggest that the questionnaire works equally well for primary and secondary school students in both luxembourg and switzerland. finally, after confirming that the instrument possesses the same psychometric characteristics across two countries, we tested measurement invariance across male and female students in primary and secondary schools (hypotheses 5). based on the goodness-of-fit statistics, all four models provided an acceptable fit (all rmseas were below .05 and all cfis were above .90; schermelleh-engel et al., 2003). thus, the assumption of measurement invariance cannot be rejected, indicating that the sals is measurement invariant for gender across primary and secondary schools. morinaj et al. 48 | f l r table 4 measurement invariance tests for the three-factor model of school alienation across swiss and luxembourgish primary and secondary school students model overall fit indices model comparison comparative fit indices  df cfi rmsea srmr δ  δdf δcfi swiss (n = 486) and luxembourgish (n = 503) primary school sample 1. configural invariance 761.35 486 .94 .04 .05 2. metric invariance 803.92 507 .94 .04 .05 2 vs. 1 40.66 21 .00 3. scalar invariance 915.07 531 .93 .04 .06 3 vs. 2 130.43* 24 .01 4. residual invariance 929.35 537 .93 .04 .06 4 vs. 3 13.31 6 .00 swiss (n = 550) and luxembourgish (n = 534) secondary school sample 1. configural invariance 923.31 486 .94 .04 .05 2. metric invariance 949.64 507 .94 .04 .05 2 vs. 1 28.31 21 .00 3. scalar invariance 1099.31 531 .93 .05 .06 3 vs. 2 170.53* 24 .01 4. residual invariance 1103.59 537 .93 .04 .06 4 vs. 3 5.79 6 .00 note. cfi = comparative fit index; rmsea = root mean squared error of approximation; srmr = standardized root mean square residual; δ difference between the comparison and nested model. δrmsea and δsrmr were explicitly below the recommended cutoff values. *p < .001. 3.5 external validity of the sals to provide additional evidence of construct validity, we estimated the relationship between the sals and external criteria by means of correlational evidence. criterion-related evidence is viewed as a gold standard in validation research (chan, 2014; kane, 2001; sireci, 2009). well-being dimensions were chosen as criterion variables from hascher’s (2004) questionnaire on subjective well-being in school (donat, peter, dalbert, & kamble, 2016; hascher, 2007, 2008, 2010; hascher & hagenauer, 2011). the scale was validated in four countries, namely, germany, switzerland, the czech republic, and the netherlands (hascher, 2007). specifically, we examined the relationship between the school alienation domains and (a) enjoyment and positive attitudes to school (hypothesis 6) and (b) social problems in school (hypothesis 7). these welldefined measures demonstrated high quality and reliability in prior research (hascher, 2008; hascher & hagenauer, 2011). the strength of the relationships between the sals and the external criteria based on a spearman’s test is presented in table 5. strength of association indices was interpreted according to cohen (1992). the pattern of correlations between the school alienation domains and well-being dimensions was highly alike between the luxembourgish and swiss samples. in keeping with our expectations, there were weak to strong significant negative associations between the domains of school alienation and the three positive well-being dimensions (positive attitudes toward school, enjoyment in school, and academic selfconcept). specifically, correlation coefficients between each of the domains of school alienation and positive attitudes toward school ranged from r = -.24 (p < .01) to r = -.59 (p < .01) in the swiss canton of bern and from r = -.26 (p < .01) to r = -.55 (p < .01) in luxembourg. furthermore, there were weak to medium positive associations between alienation from classmates/teachers and social problems in school in both countries (r = .17 – .45, p < .01). in addition, a small positive correlation was discovered between alienation from classmates/teachers and worries in school, whereas physical complaints had a weak positive association with all three domains of school alienation. morinaj et al. 49 | f l r table 5 correlations between the factors of school alienation and dimensions of school-specific well-being canton of bern luxembourg classmates teachers learning classmates teachers learning positive attitudes -.24**(-.28**) -.30**(-.33**) -.59**(-.54**) -.26**(-.29**) -.32**(-.41**) -.51**(-.55**) enjoyment -.23**(-.17**) -.34**(-.30**) -.43**(-.42**) -23**(-.28**) -.30**(-.44**) -.37**(-.48**) self-concept -.19**(-.23**) -.22**(-.25**) -.26**(-.28**) -.11* (-.18**) -.27**(-.30**) -.24**(-.28**) worries .09*(.19**) .14**(.16**) -.02 (.03) .05 (.13**) .16**(.14**) .05 (.02) physical complaints .11*(.17**) .16**(.21**) .05 (.11**) .08 (.24**) .21**(.34**) .18**(.13**) social problems .35**(.39**) .22**(.17**) .10*(.08) .36**(.45**) .30**(.26**) .20**(.09*) note. values without brackets represent correlations for primary school students (n = 486 students, canton of bern; n = 503 students, luxembourg) and values within brackets represent correlations for secondary school students (n = 550 students, canton of bern; n = 534 students, luxembourg). *p < .05. **p < .01. 4. discussion and conclusion this study investigated the reliability and validity of a newly developed scale intended to measure school alienation (school alienation scale, sals). school alienation was operationalized as student alienation from the most relevant domains in the context of schooling, namely, classmates, teachers, and learning. to our knowledge, to date there has been a dearth of well-crafted evaluations of student alienation from school (brown et al., 2003b). thus, this research enriches previous studies on school alienation and presents a comprehensive and psychometrically sound standardized measurement instrument to assess school alienation. the instrument consistently demonstrated strong psychometric properties, external validity, and measurement invariance across two countries — luxembourg and switzerland. we applied a binational approach to validate the construct of school alienation in two countries with different educational systems, thus contributing to comparative research and the generalizability of the study results. our study furthermore considers both social actors and academic aspects of schooling as well as the state-of-the-art research on the interplay between emotional and cognitive processes (dai & sternberg, 2004; linnenbrink & pintrich, 2004; pessoa, 2008). school alienation emerged in three distinct but closely related domains. in accordance with the theory, the results of this study supported prior research that school alienation stems from three experiences. first, alienation from teachers reflects the idea that poor interactions between students and teachers, whom students may perceive as being unmotivated, nonsupportive, and disrespectful, facilitate development of alienation (hascher, 2010; kagan, 1990; legault et al., 2006). second, students alienated from classmates are exposed more than most adolescents to social isolation and have fewer important relationships with other people; they experience academic misfortune and a lack of identification with a school environment (buhs & ladd, 2001; ladd, birch, & buhs, 1999). third, students alienated from learning are less likely to participate in school lessons and learning activities; they experience a negative association with the educational environment and perceive school life as worthless (çağlar, 2013; carlson, 1995). however, social relations might mediate the relationship between negative school experience and alienation (hascher, 2010). in each of the three school alienation domains, we have found an almost perfect positive correlation between emotional and cognitive components, indicating that cognition and emotion are interwoven psychological processes and undergo the same directional change. the results of the reliability analysis supported hypothesis 1, indicating the sals to be a reliable measure of school alienation. the three subscales of the sals, namely, alienation from classmates, teachers, and learning, displayed acceptable to very good reliability for both primary and secondary school students in both countries. furthermore, a series of confirmatory factor analyses supported the postulated structure of the sals, affirming its theoretical and scale validity. the hypothesized first-order three-factor model was morinaj et al. 50 | f l r particularly superior to three alternative models (hypothesis 2 confirmed). this finding contributes to alienation research, indicating that students can be alienated from classmates, teachers, and/or learning at school. although the three domains of school alienation were found to be correlated, there is no solid ground for a one-factor solution, implying that the factors of school alienation are conceptually and statistically distinguishable. this argues that no strong hierarchical structure underlies the sals factors. we, therefore, provided additional evidence that school alienation can be addressed as a multidimensional domain-specific construct. furthermore, even though the proposed competing model (in which emotional and cognitive components were considered as separate facets of school alienation domains; model 4) fitted the data reasonably well, this study revealed that emotion and cognition are interwoven psychological processes. nearly perfect positive correlations between emotional and cognitive components provided evidence that emotional and cognitive processes are closely related and cannot work independently of each other (dai & sternberg, 2004; eisenberg, 2014; kreitler, 2013). integrated understanding of intellectual functioning might be achieved through applying neurobiological approaches, which are generally disregarded in psychological research (dai & sternberg, 2004). applying the longitudinal cohort-sequential design, we will significantly improve the validity of conclusions suggested in this study. with regard to hypotheses 3, 4, and 5, this study revealed the measurement invariance of the threefactor model across luxembourgish and swiss primary and secondary school students and across gender (hypotheses 3, 4, and 5 confirmed). these findings suggested that respondents from luxembourg and the swiss canton of bern considered the construct of school alienation similarly and assigned the same meaning to the items. support for scalar invariance indicated that the latent means can be significantly compared across groups (milfont & fischer, 2010). finally, the same level of measurement error was found for each item between groups, confirming residual invariance. thus, the hypothesized three-factor model was found to be appropriate for both primary and secondary school students and may be used with both populations, regardless of student’s gender. in keeping with the idea that alienated youth are exposed to both physical and emotional disease (farrow, 1991), the domains of school alienation could be associated with the dimensions of well-being. consistent with hypotheses 5 and 6, the sals had a significant negative correlation to enjoyment and positive attitudes to school and a positive correlation with social problems in school. these significant moderate correlations were similar across the swiss and luxembourgish samples, providing additional evidence of the new instrument’s validity. thus, these findings indicated that the sals had robust criterion-related validity; that is, the more individuals feel good and evaluate their current situation positively, the lower the level of their alienation (vahedi & nazari, 2011). for example, when students appreciate activities and settings as well as value people in these settings, their alienation decreases and their psychological well-being increases (ifeagwazi et al., 2015). in contrast, social discrimination in the classroom was found to lessen student wellbeing and intensify social problems in school (hascher, 2010), because it is especially in adolescence that students experience a keen desire for peer interactions and want to be accepted by peers (rubin, bukowski, & parker, 1998; safipour et al., 2011). social alienation, in turn, was found to be one of the possible risk factors for health (safipour et al., 2011). research on well-being suggested that social support, positive interactions, and integration into the classroom maintained by teachers and classmates might stimulate student well-being (hascher, 2010), thus diminishing their alienation from school and increasing their probability of becoming productive citizens. applying a longitudinal design, further research could explore the transition from childhood to adulthood, when young people replace childish behavior with more mature behavior and which in turn may affect adolescent’s psychological well-being (as cited in safipour et al., 2011). although the hypothesized model and research hypotheses were supported, we identified several limitations of the present study. first, participation in this study was voluntary, assuming that the study sample might be biased. given that prior to filling out a questionnaire we obtained written parent permission, it is possible that alienated students were underrepresented in the current study sample. we also recognize that there are limitations in regard to the instrument development process. some items in the initial scale administered to students during the pilot study behaved differently for the swiss and luxembourgish samples, which resulted in substantial scale reduction. although each item represented a particular situation, participants morinaj et al. 51 | f l r in luxembourg and the swiss canton of bern might interpret items differently according to their feelings and thoughts toward classmates, teachers, and learning. since we were able to show that the final sals operated in the same way across primary and secondary school students in luxembourg and the swiss canton of bern, it is reasonable to expect cultural differences between the initial and new populations. we suggest that future studies include cognitive interviews with the participants during the pilot testing phase to obtain students’ understanding of items and indicate problematic items for the new sample, or ascribe any difficulties with the test instructions to cultural differences. further, cognitive interviews provide researchers with additional evidence regarding differences in item responses among participants from different samples. however, we assume that cognitive interviews could have influenced students’ daily school life, because think-aloud conditions trigger participants to reflect upon their relationships with classmates and teachers as well as their attitudes toward learning. in addition to the questionnaire, it would be useful to include a student standardized diary, in which students could specify their feelings and perceptions regarding classmates, teachers, and learning. this research method, which is strongly recommended to be applied more frequently in educational settings (schmitz & wiese, 2006), would allow us to observe students’ attitudes over time between different measurement time-points. asking students to report their experiences in diaries would enable us to gain insight into participants’ stances and give the students an opportunity for self-reflection. we believe that after rigorous preparation, both a questionnaire and a student diary might be administered using a computer. moreover, despite the confirmation of the proposed model, further domains of alienation, such as parental alienation (kelly & johnston, 2001; kocayörük & şimşek, 2015), may be useful for the analysis of school alienation. separation or divorce may lead to parental alienation syndrome expressed in children’s estrangement from their parents. in this sense, alienated children experience unreasonable negative feelings and beliefs toward one of their parents as a reaction to parents’ divorce (kelly & johnston, 2001). this syndrome might undermine not only parent-child and other interpersonal relationships, but also students’ school performance and engagement with the school environment (kocayörük & şimşek, 2015). in addition, parental attitudes toward school may also affect students’ feelings about school (ames, de stefano, watkins, & sheldon, 1995). children tend to imitate parents’ behavior and later adopt it as their own (parsons, adler, & kaczala, 1982). for example, parents’ active participation and interest in what their children are learning have been found to enhance students’ value of education and their academic motivation (gonzalez-dehass, willems, & holbein, 2005). longitudinal studies can reveal the development of school alienation, its causes, and consequences for students and education in general. to conclude, the present study indicates that the sals can be used as a valid and reliable research instrument in a variety of educational settings. the main finding of this study is that the hypothesized firstorder three-factor model, incorporating three school-related domains, namely, alienation from classmates, from teachers, and from learning, was confirmed. this finding is especially noteworthy because the disclosure of the core elements of school alienation, conceptually and statistically distinguishable, may contribute to further research and praxis. we also found that emotional and cognitive components of school alienation domains cannot be considered separately. the results of this study further show a negative association between school alienation domains and positive attitudes toward school; and a positive association between school alienation and social problems in school. this is the first study to clarify the defining elements of school alienation supporting the theoretically derived assumptions about the phenomenon of school alienation (see hascher & hadjar, 2017). the results demonstrate the importance of detecting alienation symptoms among students in each of the three school alienation domains, as well as their causes and consequences. understanding the school alienation phenomenon at different stages of development represents an important step in designing school-based prevention strategies to increase student well-being and academic success. while we know that the construct of school alienation operates equivalently in different educational contexts, it is necessary to observe the development of school alienation across time. in subsequent research, we need to apply a longitudinal cohort-sequential design and broaden our understanding of the school alienation phenomenon. morinaj et al. 52 | f l r keypoints we developed and validated a new 24-item questionnaire, the school alienation scale (sals), to measure school alienation among primary and secondary school students. we disclosed the core elements of the school alienation construct: alienation from classmates, teachers, and learning. we contributed to comparative research by validating the construct of school alienation in different educational settings. carefully constructed, psychometrically sound, and brief multidimensional assessment instrument is practical for use in future studies. acknowledgments we would like to thank the students and teachers who participated in the sasal project. we also express our deepest gratitude to the research assistants who supported us during data collection and data entry processes. our sincere appreciation is extended to pd dr. gerda hagenauer, for her invaluable advice and expertise in data analysis. this study is supported by the swiss national science foundation under grant 100019l_159979 in switzerland and the luxembourg national research fund under grant inter/snf/14/9857103 in luxembourg. references acuña, e., & rodriguez, c. (2004). the treatment of missing values and its effect on classifier accuracy. in d. banks, f. r. mcmorris, p. arabie, & w. gaul (eds.), classification, clustering, and data mining applications (pp. 639-647). springer berlin heidelberg. doi:10.1007/978-3-642-17103-1_60 aiken, m., & hage, j. (1966). organizational alienation: a comparative analysis. american sociological review, 31(4), 497-507. altenbaugh, r. j., engel, d. e., & martin, d. t. (1995). caring for kids: a critical study of urban school leavers. london: the falmer press. american educational research association, american psychological association, and national council on measurement in education. (2014). standards for educational and psychological testing. washington, dc: american psychological association. ames, c., de stefano, l., watkins, t., and sheldon, s. (1995). teachers’ school-to-home communications and parent involvement: the role of parent perceptions and beliefs. retrieved from: http://eric.ed.gov/?id=ed383451 archambault, i., janosz, m., fallu, j. s., & pagani, l. s. (2009a). student engagement and its relationship with early high school dropout. journal of adolescence, 32(3), 651-670. doi:10.1016/j.adolescence.2008.06.007 archambault, i., janosz, m., morizot, j., & pagani, l. (2009b). adolescent behavioral, affective, and cognitive engagement in school: relationship to dropout. journal of school health, 79(9), 408-415. doi:10.1111/j.1746-1561.2009.00428.x ariño, a. (2003). measures of strategic alliance performance: an analysis of construct validity. journal of international business studies, 34(1), 66-79. doi:10.1057/palgrave.jibs.8400005 aritzeta, a., balluerka, n., gorostiaga, a., alonso-arbiol, i., haranburu, m., & gartzia, l. (2016). classroom emotional intelligence and its relationship with school performance. european journal of education and psychology, 9(1), 1-8. doi:10.1016/j.ejeps.2015.11.001 http://eric.ed.gov/?id=ed383451 http://eric.ed.gov/?id=ed383451 morinaj et al. 53 | f l r arnold, j. a., arad, s., rhoades, j. a., & drasgow, f. (2000). the empowering leadership questionnaire: the construction and validation of a new scale for measuring leader behaviors. journal of organizational behavior, 21(3), 249-269. babad, e. (1992). teacher expectancies and nonverbal behavior. in r. s. feldman (ed.), applications of nonverbal behavioral theories and research (pp. 167-190). hillsdale, nj: erlbaum. baker, j. a. (1998). are we missing the forest for the trees? considering the social context of school violence. journal of school psychology, 36(1), 29-44. barnhardt, b., & ginns, p. (2014). an alienation-based framework for student experience in higher education: new interpretations of past observations in student learning theory. higher education, 68(6), 789-805. doi: 10.1007/s10734-014-9744-y becker, g. s. (1994). human capital revisited. in g. s. becker (ed.), human capital: a theoretical and empirical analysis with special reference to education (pp. 15-28). chicago: the university of chicago press. blauner, r. (1964). alienation and freedom: the factory worker and his industry. the sociological quarterly, 6(1), 83-85. doi:10.2307/2574777 brown, t. a. (2015). confirmatory factor analysis for applied research. new york: guilford publications. brown, m. r., higgins, k., & paulsen, k. (2003a). adolescent alienation what is it and what can educators do about it? intervention in school and clinic, 39(1), 3-9. doi:10.1177/10534512030390010101 brown, m. r., higgins, k., pierce, t., hong, e., & thoma, c. (2003b). secondary students’ perceptions of school life with regard to alienation: the effects of disability, gender and race. learning disability quarterly, 26(4), 227-238. buhs, e. s., & ladd, g. w. (2001). peer rejection as antecedent of young children’s school adjustment: an examination of mediating processes. developmental psychology, 37(4), 550-560. doi:10.2307/1593636 byrne, b. m. (2012). structural equation modeling with mplus: basics, concepts, applications, and programming. new york: routledge. çağlar, ç. (2012). development of the student alienation scale (sas). education and science, 37(166), 195205. çağlar, ç. (2013). the relationship between the perceptions of the fairness of the learning environment and the level of alienation. eurasian journal of educational research, 50, 185-206. carlson, t. b. (1995). we hate gym: student alienation from physical education. journal of teaching in physical education, 14(4), 467-477. doi:10.1123/jtpe.14.4.467 case, j. m. (2008). alienation and engagement: development of an alternative theoretical framework for understanding student learning. higher education, 55(3), 321-332. doi:10.1007/s10734-007-9057-5 chan, e.k.h. (2014). standards and guidelines for validation practices: development and evaluation of measurement instruments. in b. d. zumbo & e. k. h. chan (eds.), validity and validation in social, behavioral, and health sciences (pp. 9-24). switzerland: springer international publishing. doi:10.1007/978-3-319-07794-9_2 chen, f. f. (2007). sensitivity of goodness of fit indexes to lack of measurement invariance. structural equation modeling, 14(3), 464-504. doi:10.1080/10705510701301834 chen, f. f., sousa, k. h., & west, s. g. (2005). teacher's corner: testing measurement invariance of secondorder factor models. structural equation modeling, 12(3), 471-492. doi:10.1207/s15328007sem1203_7 clark, l. a., & watson, d. (1995). constructing validity: basic issues in objective scale development. psychological assessment, 7(3), 309-319. doi:10.1037/1040-3590.7.3.309 cohen, j. (1992). a power primer. psychological bulletin, 112(1), 155–159. doi:10.1037/00332909.112.1.155 crinson, i., & yuill, c. (2008). what can alienation theory contribute to an understanding of social inequalities in health? international journal of health services, 38(3), 455–469. doi:10.2190/hs.38.3.e dai, d. y., & sternberg, r. j. (2004). beyond cognitivism: toward an integrated understanding of intellectual functioning and development. in d. y. dai & r. j. sternberg (eds.), motivation, emotion, and cognition: integrative perspectives on intellectual functioning and development (pp. 3-38). mahwah, nj: lawrence erlbaum. https://doi.org/10.2307/2574777 https://doi.org/10.1177/10534512030390010101 https://doi.org/10.2307/1593636 https://doi.org/10.1123/jtpe.14.4.467 http://dx.doi.org/10.1080/10705510701301834 http://dx.doi.org/10.1207/s15328007sem1203_7 http://psycnet.apa.org/doi/10.1037/0033-2909.112.1.155 http://psycnet.apa.org/doi/10.1037/0033-2909.112.1.155 https://doi.org/10.2190/hs.38.3.e morinaj et al. 54 | f l r dean, d. g. (1961). alienation: its meaning and measurement. american sociological review, 5, 753-758. doi:10.2307/2090204 dekel, r., & tuval-mashiach, r. (2012). multiple losses of social resources following collective trauma: the case of the forced relocation from gush katif. psychological trauma: theory, research, practice and policy, 4(1), 56–65. doi:10.1037/a0019912 dickey, m. (2004). the impact of web-logs (blogs) on student perceptions of isolation and alienation in a webbased distance-learning environment. open learning, 19(3), 279-291. doi:10.1080/0268051042000280138 donat, m., peter, f., dalbert, c., & kamble, s. v. (2016). the meaning of students’ personal belief in a just world for positive and negative aspects of school-specific well-being. social justice research, 29(1), 73102. doi:10.1007/s11211-015-0247-5 dweck, c. s., mangels, j. a., good, c. (2004). motivational effects on attention, cognition, and performance. in d. y. dai & r. j. sternberg (eds.), motivation, emotion, and cognition: integrative perspectives on intellectual functioning and development (pp. 41-55). mahwah, nj: lawrence erlbaum. fredricks, j. a., blumenfeld, p. c., & paris, a. h. (2004). school engagement: potential of the concept, state of the evidence. review of educational research, 74(1), 59-109. doi:10.3102/00346543074001059 eccles, j. s., & alfeld, c. a. (2007). not you! not here! not now. in r. k. silbereisen & r. m. lerner (eds.), approaches to positive youth development (pp. 133–156). thousand oaks, ca: sage eccles, j. s., brown, b. v., & templeton, j. l. (2008). a developmental framework for selecting indicators of well-being during the adolescent and young adult years. in b. v. brown (ed.), key indicators of child and youth well-being: completing the picture (pp. 197–236). mahwah, nj: erlbaum. eccles, j., & gootman, j. a. (2002). community programs to promote youth development. washington, dc: national academy press. doi:10.17226/10022 eccles, j. s., & midgley, c. (1989). stage-environment fit: developmentally appropriate classrooms for young adolescents. in c. ames & r. ames (eds.), research on motivation in education (pp. 139-186). san diego: academic press. eccles, j. s., & midgley, c. (1990). changes in academic motivation and self-perception during early adolescence. in r. montemayor, g. adams, & t. p. gullotta (eds.). from childhood to adolescence: a transitional period? (pp. 134-155). newbury park: sage publications. eccles, j. s., & roeser, r. w. (2009). schools, academic motivation, and stage‐environment fit. in r.m. lerner & l. steinber (eds.), handbook of adolescent psychology (pp. 404-434). hoboken, nj: john wiley & sons. doi:10.1002/9780470479193.adlpsy001013 eisenberg, n. (2014). altruistic emotion, cognition, and behavior. new york, ny: psychology press. exline, j. j., yali, a. m., & sanderson, w. c. (2000). guilt, discord, and alienation: the role of religious strain in depression and suicidality. journal of clinical psychology, 56(12), 1481-1496. doi:10.1002/10974679(200012)56:12<1481::aid-1>3.0.co;2-a farrow, j. a. (1991). youth alienation as an emerging pediatric health care issue. american journal of diseases of children, 145(5), 491-492. doi:10.1001/archpedi.1991.02160050015002 finifter, a. w. (1970). dimensions of political alienation. american political science review, 64(02), 389410. doi:10.2307/1953840 frey, a., ruchkin, v., martin, a., & schwab-stone, m. (2009). adolescents in transition: school and family characteristics in the development of violent behaviors entering high school. child psychiatry and human development, 40(1), 1-13. doi:10.1007/s10578-008-0105-x ghaith, g. m., shaaban, k. a., & harkous, s. a. (2007). an investigation of the relationship between forms of positive interdependence, social support, and selected aspects of classroom climate. system, 35(2), 229240. doi:10.1016/j.system.2006.11.003 goetz, t., frenzel, a. c., pekrun, r., & hall, n. c. (2006). the domain specificity of academic emotional experiences. the journal of experimental education, 75(1), 5-29. gonzalez-dehass, a. r., willems, p. p., & holbein, m. f. d. (2005). examining the relationship between parental involvement and student motivation. educational psychology review, 17(2), 99-123. doi:10.1007/s10648-005-3949-7 https://doi.org/10.3102/00346543074001059 https://doi.org/10.1002/1097-4679(200012)56:12%3c1481::aid-1%3e3.0.co;2-a https://doi.org/10.1002/1097-4679(200012)56:12%3c1481::aid-1%3e3.0.co;2-a morinaj et al. 55 | f l r gottfried, a. e., fleming, j. s., & gottfried, a. w. (2001). continuity of academic intrinsic motivation from childhood through late adolescence: a longitudinal study. journal of educational psychology, 93(1), 313. doi:10.1037/0022-0663.93.1.3 gutman, l. m., & eccles, j. s. (2007). stage-environment fit during adolescence: trajectories of family relations and adolescent outcomes. developmental psychology, 43(2), 522-537. doi:10.1037/00121649.43.2.522 hadjar, a., backes s., & gysin., s. (2015). school alienation, patriarchal gender-role orientations and the lower educational success of boys. a mixed-method study. masculinities and social change, 4(1), 85116. doi:10.17583/msc.2015.1319 hadjar, a., & lupatsch, j. (2010). der schul(miss)erfolg der jungen. kölner zeitschrift für soziologie und sozialpsychologie, 62(4), 599-622. doi:10.1007/s11577-010-0116-z hair, j. f., anderson, r. e., tatham, r. l., & black, w. c. (1995). multivariate data analysis. nj: prenticehall. hall-lande, j.a., eisenberg, m.e., christenson, s.l., & neumark-sztainer, d. (2007). social isolation, psychological health, and protective factors in adolescence. adolescence, 42(166), 265–286. hascher, t. (2004). wohlbefinden in der schule. münster: waxmann. hascher, t. (2007). exploring students’ well-being by taking a variety of looks into the classroom. hellenic journal of psychology, 4(3), 331-349. hascher, t. (2008). quantitative and qualitative research approaches to assess student well-being. international journal of educational research, 47(2), 84-96. doi:10.1016/j.ijer.2007.11.016 hascher, t. (2010). wellbeing. in p. peterson, e. baker & b. mcgaw (eds.), international encyclopedia of education (pp. 732-738). oxford: elsevier. doi:10.1016/b978-0-08-044894-7.00633-3 hascher, t. (2011). wellbeing. in s. järvelä (ed.), social and emotional aspects of learning (pp. 99-106). oxford: elsevier. hascher, t. (2012). well-being and learning in school. in n. m. seel (ed.), encyclopedia of the sciences of learning (pp. 3453-3456). heidelberg: springer. hascher, t., & hadjar, a. (2017). school alienation – a review of theoretical approaches and educational research. sasal-manuscript, university of bern/university of luxembourg. hascher, t., & hagenauer, g. (2010). alienation from school. international journal of educational research, 49(6), 220-232. doi:10.1016/j.ijer.2011.03.002 hascher, t., & hagenauer, g. (2011). schulisches wohlbefinden im jugendalter–verläufe und einflussfaktoren. in a. ittel, h. merkens, & l. stecher (eds.), jahrbuch jugendforschung (pp. 15-45). wiesbaden: vs. doi:10.1007/978-3-531-93116-6_1 hilgard, e. r. (1980). the trilogy of mind: cognition, affection, and conation. journal of the history of the behavioral sciences, 16(2), 107-117. doi:10.1002/1520-6696(198004)16:2<107::aidjhbs2300160202>3.0.co;2-y hinkin, t. r., tracey, j. b., & enz, c. a. (1997). scale construction: developing reliable and valid measurement instruments. journal of hospitality & tourism research, 21(1), 100-120. doi:10.1177/109634809702100108 hirschfeld, r. r., & feild, h. s. (2000). work centrality and work alienation: distinct aspects of a general commitment to work. journal of organizational behavior, 21, 789-800. doi: 10.1002/10991379(200011)21:7<789::aid-job59>3.0.co;2-w honkanen, p., verplanken, b., & olsen, s. o. (2006). ethical values and motives driving organic food choice. journal of consumer behaviour, 5(5), 420-430. doi: 10.1002/cb.190 hu, l. t., & bentler, p. m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives. structural equation modeling: a multidisciplinary journal, 6(1), 1-55. doi:10.1080/10705519909540118 hughes, j. n., gleason, k. a., & zhang, d. (2005). relationship influences on teachers' perceptions of academic competence in academically at-risk minority and majority first grade students. journal of school psychology, 43(4), 303-320. doi:10.1016/j.jsp.2005.07.001 http://dx.doi.org/10.1080/10705519909540118 https://dx.doi.org/10.1016%2fj.jsp.2005.07.001 morinaj et al. 56 | f l r hyman, i., cohen, i., & mahon, m. (2003). student alienation syndrome: a paradigm for understanding the relation between school trauma and school violence. the california school psychologist, 8(1), 73-86. doi:10.1007/bf03340897 ifeagwazi, c. m., chukwuorji, j. c., & zacchaeus, e. a. (2015). alienation and psychological wellbeing: moderation by resilience. social indicators research, 120(2), 525-544. doi:10.1007/s11205-014-06021 isen, a. m. (2004). some perspectives on positive feelings and emotions: positive affect facilitates thinking and problem solving. in a. s. r. manstead, n. frijda, & a. fischer (eds.), feelings and emotions: the amsterdam symposium (pp. 263-281). cambridge, ny: cambridge university press. doi.10.1017/cbo9780511806582.016 jussim, l., & eccles, j. s. (1992). teacher expectations ii: construction and reflection of student achievement. journal of personality and social psychology, 63(6), 947-961. doi:10.1037/00223514.63.6.947 jussim, l., eccles, j., & madon, s. (1996). social perception, social stereotypes, and teacher expectations: accuracy and the quest for the powerful self-fulfilling prophecy. in m. p. zanna (ed.). advances in experimental social psychology (pp. 281–388). san diego, ca: academic press. doi:10.1016/s00652601(08)60240-3 kagan, d. m. (1990). how schools alienate students at risk: a model for examining proximal classroom variables. educational psychologist, 25, 105–125. doi:10.1207/s15326985ep2502_1 kane, m. t. (2001). current concerns in validity theory. journal of educational measurement, 38(4), 319-342. kelly, j. b., & johnston, j. r. (2001). the alienated child: a reformulation of parental alienation syndrome. family court review, 39(3), 249-266. doi: 10.1111/j.1745-3984.2001.tb01130.x kocayörük, e., & şimşek, ö. f. (2015). parental attachment and adolescents' perception of school alienation: the mediation role of self-esteem and adjustment. the journal of psychology, 150(4), 405421. doi:10.1080/00223980.2015.1060185 kreitler, s. (2013). cognition and motivation: forging an interdisciplinary perspective. cambridge university press. krumboltz, j. d. (2009). the happenstance learning theory. journal of career assessment, 17(2), 135-154. doi:10.1177/1069072708328861 ladd, g. w. (1999). peer relationships and social competence during early and middle childhood. annual review of psychology, 50(1), 333-359. doi:10.1146/annurev.psych.50.1.333 ladd, g. w., birch, s. h., & buhs, e. s. (1999). children's social and scholastic lives in kindergarten: related spheres of influence? child development, 70(6), 1373-1400. doi:10.1111/1467-8624.00101 legault, l., green-demers, i., & pelletier, l. (2006). why do high school students lack motivation in the classroom? toward an understanding of academic amotivation and the role of social support. journal of educational psychology, 98(3), 567. doi:10.1037/0022-0663.98.3.567 linnenbrink, e. a., & pintrich, p. r. (2004). role of affect in cognitive processing in academic contexts. in d. y. dai & r. j. sternberg (eds.), motivation, emotion, and cognition: integrative perspectives on intellectual functioning and development (pp. 57-87). mahwah, nj: lawrence erlbaum. doi:10.4324/9781410610515 liu, r. (2010). alienation and first-year student retention. professional file, 116, 1-18. lumsden, l. s. (1994). student motivation to learn. research roundup, 10(3),1-5. martin, j. (2008). pedagogy of the alienated: can freirian teaching reach working-class students? equity & excellence in education, 41(1), 31–44. doi:10.1080/10665680701773776 mau, r. y. (1992). the validity and devolution of a concept: student alienation. adolescence, 27(107), 731741. mckown, c., & weinstein, r. s. (2008). teacher expectations, classroom context, and the achievement gap. journal of school psychology, 46(3), 235-261. doi:10.1016/j.jsp.2007.05.001 menje & university of luxembourg (2015). bildungsbericht luxemburg 2015. band 2: analysen und befunde. luxembourg: menje. https://doi.org/10.1017/cbo9780511806582.016 http://dx.doi.org/10.1207/s15326985ep2502_1 https://doi.org/10.1177/1069072708328861 https://doi.org/10.1146/annurev.psych.50.1.333 http://dx.doi.org/10.4324/9781410610515 http://dx.doi.org/10.1080/10665680701773776 morinaj et al. 57 | f l r messick, s. (1995). validity of psychological assessment: validation of inferences from persons' responses and performances as scientific inquiry into score meaning. american psychologist, 50(9), 741-749. doi:10.1037/0003-066x.50.9.741 milfont, t. l., & fischer, r., (2010). testing measurement invariance across groups: applications in crosscultural research. international journal of psychological research, 3(1), 111-121. moreno, e. s., & de roda, a. b. l. (2003). social psychology of mental health: the social structure and personality perspective. the spanish journal of psychology, 6(1), 3–11.
 murdock, t. b. (1999). the social context of risk: status and motivational predictors of alienation in middle school. journal of educational psychology, 91(1), 62-75. doi:10.1037//0022-0663.91.1.62 newmann, f. m. (1992). student engagement and achievement in american secondary schools. new york: teachers college press. ong, a. d., & van dulmen, m. h. (2006). oxford handbook of methods in positive psychology. new york, ny: oxford university press. pantoja, a. d., & segura, g. m. (2003). does ethnicity matter? descriptive representation in legislatures and political alienation among latinos. social science quarterly, 84(2), 441-460. doi:10.1111/15406237.840201 parsons, j. e., adler, t. f., & kaczala, c. m. (1982). socialization of achievement attitudes and beliefs: parental influences. child development, 53(2), 310-321. doi:10.2307/1128973 pessoa, l. (2008). on the relationship between emotion and cognition. nature reviews neuroscience, 9(2), 148-158. doi:10.1038/nrn2317 preston, c. c., & colman, a. m. (2000). optimal number of response categories in rating scales: reliability, validity, discriminating power, and respondent preferences. acta psychologica, 104(1), 1-15. doi:10.1016/s0001-6918(99)00050-5 railton, p. (2013). alienation, consequentialism, and the demands of morality. in r. shafer-landau (ed.), ethical theory: an anthology (pp. 441-457). john wiley & sons, inc. reinke, w. m., & herman, k. c. (2002). creating school environments that deter antisocial behaviors in youth. psychology in the schools, 39(5), 549-559. rovai, a. p., & wighting, m. j. (2005). feelings of alienation and community among higher education students in a virtual classroom. internet and higher education, 8(2), 97–110. doi:10.1016/j.iheduc.2005.03.001 rubin, k. h., bukowski, w., & parker, j. g. (1998). peer interactions, relationships, and groups. handbook of child psychology, 3(5), 619-700. safipour, j., schopflocher, d., higginbottom, g., & emami, a. (2011). the mediating role of alienation in self-reported health among swedish adolescents. retrieved from: http://www.tandfonline.com/doi/pdf/10.3402/vgi.v2i0.5805?needaccess=true schabracq, m., & cooper, c. (2003). to be me or not to be me: about alienation. counselling psychology quarterly, 16(2), 53-79. schunk, d. h., & meece, j. (2005). self-efficacy development in adolescence. in f. pajares & t. urdan (eds.), self-efficacy beliefs during adolescence (pp. 71-96). greenwich, ct: information age publishing. schermelleh-engel, k., moosbrugger, h., & müller, h. (2003). evaluating the fit of structural equation models: tests of significance and descriptive goodness-of-fit measures. methods of psychological research online, 8(2), 23-74. schmitt, n., & kuljanin, g. (2008). measurement invariance: review of practice and implications. human resource management review, 18(4), 210-222. doi:10.1016/j.hrmr.2008.03.003 schmitz, b., & wiese, b. s. (2006). new perspectives for the evaluation of training sessions in self-regulated learning: time-series analyses of diary data. contemporary educational psychology, 31(1), 64-96. doi:10.1016/j.cedpsych.2005.02.002 seeman, m. (1959). on the meaning of alienation. american sociological review, 24, 783-791. doi:10.2307/2088565 seeman, m. (1975). alienation studies. annual review of sociology, 1, 91-123. doi:10.1146/annurev.so.01.080175.000515 http://www.tandfonline.com/doi/pdf/10.3402/vgi.v2i0.5805?needaccess=true http://www.tandfonline.com/doi/pdf/10.3402/vgi.v2i0.5805?needaccess=true http://dx.doi.org/10.1016/j.hrmr.2008.03.003 morinaj et al. 58 | f l r seetanah, b. (2009). the economic importance of education: evidence from africa using dynamic panel data analysis. journal of applied economics, 12(1), 137-157. shakespeare, w. (2000). romeo and juliet. oxford, uk: oxford university press (originally published in 1599). shantz, a., alfes, k., & truss, c. (2014). alienation from work: marxist ideologies and twenty-first-century practice. the international journal of human resource management, 25(18), 2529-2550.doi: 10.1080/09585192.2012.667431 sidorkin, a. m. (2004). in the event of learning: alienation and participative thinking in education. educational theory, 54(3), 251-262.doi:10.1111/j.0013-2004.2004.00018.x singer, e., & couper, m. p. (2008). do incentives exert undue influence on survey participation? experimental evidence. journal of empirical research on human research ethics, 3(3), 49-56. doi:10.1525/jer.2008.3.3.49 sireci, s. g. (2009). packing and unpacking sources of validity evidence: history repeats itself again. in r. w. lissitz (ed.), the concept of validity: revisions, new directions and applications (pp. 19–38). charlotte, nc: information age publishers.
 snow, r. e., corno, l., & jackson d. (1996). individual differences in affective and conative functions. in d. c. berliner & r. c. calfee (eds.), handbook of educational psychology (pp. 243-310). new york: macmillan. steenkamp, j. b. e., & baumgartner, h. (1998). assessing measurement invariance in cross-national consumer research. journal of consumer research, 25(1), 78-90. doi:10.1086/209528 tarquin, k., & cook-cottone, c. (2008). relationships among aspects of student alienation and self concept. school psychology quarterly, 23(1), 16-25. tinto, v. (1975). dropout from higher education: a theoretical synthesis of recent research. review of educational research, 45(1), 89-125. doi:10.3102/00346543045001089 tinto, v. (1993). leaving college: rethinking the causes of student attrition. chicago: university of chicago press. oecd. (2004). learning for tomorrow's world: first results from pisa 2003. paris: oecd publishing. doi:10.1787/9789264006416-en osin, e. (2009). subjective experience of alienation: measurement and correlates. gesellschaft fur logotherapie und existenzanalyse, 1(26), 16-23. østerås, n., gulbrandsen, p., garratt, a., benth, j. š., dahl, f. a., natvig, b., & brage, s. (2008). a randomised comparison of a four-and a five-point scale version of the norwegian function assessment scale. health and quality of life outcomes, 6(14), 1-9. doi:10.1186/1477-7525-6-14 stamm, m., kost, j., suter, p., holzinger, m., & stroezel, h. (2011). dropout ch: schulabbruch und absentismus in der schweiz. zeitschrift für pädagogik, 57(2), 187-202. vahedi, s., & nazari, m. a. (2011). the relationship between self-alienation, spiritual well-being, economic situation and satisfaction of life: a structural equation modeling approach. iranian journal of psychiatry and behavioral sciences, 5(1), 64-73. van de schoot, r., lugtig, p., & hox, j. (2012). a checklist for testing measurement invariance. european journal of developmental psychology, 9(4), 486-492. doi:10.1080/17405629.2012.686740 vandenberg, r. j., & lance, c. e. (2000). a review and synthesis of the measurement invariance literature: suggestions, practices, and recommendations for organizational research. organizational research methods, 3(1), 4-70. doi:10.1177/109442810031002 vila, l. e. (2000). the non-monetary benefits of education. european journal of education, 35(1), 21-32. doi:10.1111/1467-3435.00003 watt, i. (2000). essays on conrad. cambridge university press. doi:10.1017/cbo9780511485343 wimmer, m. (2013). motivation, cognition, and emotion: a phylogenetic-interdisciplinary approach. in s. kreitler (ed.), cognition and motivation: forging an interdisciplinary perspective (pp. 137-157). cambridge university press. young, i. m. (1984). pregnant embodiment: subjectivity and alienation. journal of medicine and philosophy, 9(1), 45-62. doi:10.1093/jmp/9.1.45 http://dx.doi.org/10.1080/09585192.2012.667431 https://dx.doi.org/10.1525%2fjer.2008.3.3.49 https://doi.org/10.3102/00346543045001089 https://doi.org/10.1017/cbo9780511485343 https://doi.org/10.1093/jmp/9.1.45 morinaj et al. 59 | f l r zhang, c., & zhuang, l. (2011). the composition of human capital and economic growth: evidence from china using dynamic panel data analysis. china economic review, 22(1), 165-171. doi:10.1016/j.chieco.2010.11.001 zihl, i., szesny, n., & nickel, t. (2013). cognition in the context of psychopathology: a selective review. in s. kreitler (ed.), cognition and motivation: forging an interdisciplinary perspective (pp. 76-95). cambridge university press. doi:10.1017/cbo9781139021463.007 zumbo, b. d. (2009). validity as contextualized and pragmatic explanation, and its implications for validation practice. in r. w. lissitz (ed.), the concept of validity: revisions, new directions and applications (pp. 65–82). charlotte, nc: information age publishers.
 https://doi.org/10.1016/j.chieco.2010.11.001 van laer et elen publication frontline learning research vol.6 no. 3 (2018) 228 issn 2295-3159 towards a methodological framework for sequence analysis in the field of self-regulated learning stijn van laer a& jan eelena aku leuven, belgium article received 13 may 2018 / revised 30 august / accepted 26 september / available online 19 december abstract in recent decades, conceptualizations and operationalizations of self-regulated learning (srl) have shifted from srl as an aptitude to srl as an event. alongside this shift, increased technological capability has introduced computer log files to the investigation of srl, uncovering new research avenues. one such avenue investigates the time-related characteristics of srl through learners’ behavioural sequences. although sequence analysis is still relatively new in srl research, other fields have fruitful traditions in its application and may serve as a basis for applications in the field of srl. ten years of investigating srl through sequence analysis have produced a wide range of methodological approaches. while this variety of methods illustrates the diversity of opportunities, it also indicates the lack of consensus regarding the most appropriate approaches often resulting in difficult to understand methods and non-transparent ways of reporting. since the introduction of sequence analysis in the field of srl, researchers have been emphasizing the need for a methodological framework to guide its application. yet, to date, no such framework has been proposed, hindering our progress through (1) transparent methods and (2) comparative studies to (3) empirical and ecological applications. to help overcome this issue, this manuscript discusses the basis of a methodological framework for the use of sequence analysis in srl research. we first make a case for why such a framework is necessary; secondly, we propose a set of guidelines which could serve as a starting point for the construction of a framework. keywords: computer log files; sequence analysis; self-regulated learning; methodological framework info corresponding author mail stijn.vanlaer@kuleuven.be doi: https://doi.org/10.14786/flr.v6i3.367 acknowledgments we would like to acknowledge the support of the project “adult learners online” funded by the agency for science and technology (project number: sbo 140029), who made this research possible. 1. introduction over the last five decades, multiple theoretical conceptualizations and practical operationalizations have been proposed for self-regulated learning (srl), shifting the focus from srl as an aptitude to srl as an event (e.g., endedijk, brekelmans, sleegers, & vermunt, 2016; panadero, klug, & järvelä, 2016; winne, 2016). besides this shift, technological developments have meant that computer log files now have a role to play in investigations of learners’ srl. from both theoretical and practical perspectives, computer log files are an interesting avenue for investigating learners’ srl (e.g., azevedo & hadwin, 2005; winne, 2005; zimmerman & schunk, 2001). on the one hand, their sequenced structure means that computer log files possess time-related characteristics relevant to the current conceptualization of srl as an event (e.g., azevedo, 2014; ben-eliyahu & bernacki, 2015; molenaar & järvelä, 2014). on the other hand, their unobtrusive nature enables us to observe traces of srl in learners’ behaviour in ecologically valid contexts (e.g., bourbonnais et al., 2006; hine, 2011). while sequence-based analysis has only become popular as a means of investigating the time-related characteristics of srl within the last ten years, other fields of research (e.g., bioinformatics, chemistry, marketing, and sociology) have longstanding traditions in the use of such analyses. insights gained from these fields may serve as a basis for applying sequence analysis in investigations of srl (e.g., perer & wang, 2014; winne & baker, 2013). a decade of log file sequence analysis in srl research has produced a large amount of relevant work (e.g., azevedo, taub, & mudrick, 2015; bannert, molenaar, azevedo, järvelä, & gašević, 2017; molenaar & järvelä, 2014; roll & winne, 2015) and a variety of methodological approaches. while theory-driven approaches for example prefer to recode log files to theoretically meaningful events, predefine the length of an ideal sequence, or set the threshold for significance (e.g., roll & winne, 2015; winne, 2010), data-driven approaches often prefer to extract the most common sequences from the data, regardless of their content and length (e.g., bannert et al., 2017; beheshitha, gašević, & hatala, 2015). differences with regard to the statistical analyses used can also be found. some researchers investigate the occurrence of for example particular sub-sequences as varying from learner to learner and apply multi-level analysis (e.g., taub, azevedo, bouchet, & khosravifar, 2014; taub, azevedo, bradbury, millar, & lester, 2017),while others focus on clusters of learners and instead apply chi-square analysis (e.g., van laer & elen, 2016; van laer, jiang, & elen, 2018) or variance analysis. still others argue that statistical analysis based on sub-sequences is insufficient to establish a full picture of learners’ learning patterns and prefer to use stochastic models based on the entire sequences to operationalize the investigation of learners’ behaviour (e.g., bannert, sonnenberg, mengelkamp, & pieger, 2015; sonnenberg & bannert, 2015). this multitude of approaches demonstrates not only the diversity of opportunities, but also the lack of consensus regarding the most appropriate methods. this lack of consensus often results in fragmentation, leading to non-transparent research practices and research reports, hampering the validation and testing of methods and thus the advancement of the investigation of srl through sequence analysis. in line with this observation, researchers have been emphasizing the need for a methodological framework to guide the application of log file sequence analysis in srl research since 2014 (e.g., azevedo, 2014; bannert, reimann, & sonnenberg, 2014; molenaar & järvelä, 2014). such a methodological framework could, on the one hand, provide a decision-tree-like approach to choosing which analysis to perform when (schnaubert, heimbuch, & bodemer, 2016) and, on the other hand, offer guidelines for reporting on each of the steps taken and considerations made. yet, to date, no methodological frameworks have been proposed (e.g., segedy & biswas, 2015; winne, 2014), hindering our ability to validate, duplicate, and so to demonstrate progress in the use of sequence analysis in srl research and our search for the most appropriate methods (kuhn, 2012). therefore, in this manuscript we discuss a methodological framework for the application of sequence analysis in the field of srl. to do so, we first make a case for why such a methodological framework is necessary. secondly, we propose a set of guidelines which may serve as a starting point for the construction of a framework. with a methodological framework in place, the investigation of time-related characteristics in srl using sequence analysis could evolve towards (1) the use of transparent methods, (2) comparative studies, and (3) empirical and ecological applications, supporting both research and practice. in what follows, we first define sequence analysis, elaborate on its link to srl and introduce the most common phases in its operationalization, providing an illustrative example from one of our own studies. the illustrative example used in this manuscript is not intended as a good practice, but a demonstration of the complexity of sequence analyses and the decisions to be made. at the end of this introductory section, we outline the operational efforts made in the search for tangible proof of progress in sequences analysis for the investigation of learners’ srl as a method. based on insights gathered from the introductory section, the second section proposes a set of guidelines upon which framework construction can be based. in the third and final section, we elaborate on the implications for research and practice and suggest further directions in the construction of a methodological framework for sequence analysis in the field of srl. 1.1 sequence analysis a sequence (β) is an ordered list of elements (β = < a, c, b, d, e, g, c, e, d, b, g >) (zhou, xu, nesbit, & winne, 2010). such elements can be physical, behavioural, or conceptual in nature. the analysis of a sequence makes it possible to discover hidden time-related relations between different sequences, parts of these sequences, and the individual elements within these sequences (antunes & oliveira, 2001). sequence analysis therefore is indispensable in many application domains (e.g., bioinformatics, chemistry, marketing, sociology, and education) (liu, dev, dontcheva, & hoffman, 2016) and approaches are plentiful. for example in bioinformatics sequence analysis is the process of investigating a deoxyribonucleic acid (dna) sequence to understand its features, function, structure, or evolution (e.g., lubahn et al., 1988; stackebrandt & goebel, 1994). in chemistry, sequence analysis comprises the determination of the sequence of a polymer formed of several monomers (e.g., martin, shabanowitz, hunt, & marto, 2000; van krevelen & te nijenhuis, 2009). in marketing, sequence analysis on its turn is often used in analytical customer relationship management applications, such as next product to buy models (e.g., kumar, venkatesan, & reinartz, 2004; prinzie & van den poel, 2007). in sociology, sequence methods are increasingly used to study life-course and career trajectories, patterns of organizational and national development, conversation and interaction structure, and the problem of work and family synchrony (e.g., bonin, vogel, & campbell, 2014; stark & vedres, 2012). finally, in recent years the field of education also gained interest in the investigation of sequence data. methods have been increasingly used in the context of data analysis to investigate learning processes (reimann, markauskaite, & bannert, 2014). one distinct area of learning research in which sequence-analysis methods have been used is srl-research, in particular for studying regulation and metacognition in students' learning through computer log files (e.g., azevedo, moos, johnson, & chauncey, 2010; zhou et al., 2010). computer log files gathered from learners’ interaction with online learning environments are the most know and potentially the least obtrusive way of gathering data with regard to learners’ srl behaviour. such log files are gathered through clickstreams. clickstreams are also known as click paths, or the route that learners choose when clicking or navigating through an online learning environment. a clickstream is a list of pages visited by a learner, presented in the order the pages are visited (also defined as the 'succession of mouse clicks' that each learners makes). based on the sequence of the visited pages, researchers attempt to map learners’ srl processes. 1.2 self-regulated learning and its measurement as learning in general is seen as an activity performed by learners rather than something happening to them as result of instruction (e.g., bandura, 1989; oliver & trigwell, 2005) it entails a self-regulated process through means of which learners’ regulate their behaviour according to the instructional demands (zimmerman & schunk, 2001). to be successful learners, learners need to self-regulate their learning. this assumption is evidenced by a substantial body of literature showing scores on self-regulated-learning-related variables to be strongly positive correlated and to have causal relations with scores on performance-related variables (e.g., daniela, 2015; lin, coburn, & eisenberg, 2016). the theoretical conceptualization of srl evolved from srl as an aptitude to srl as an event. the aptitude approach on the one hand sees srl as in-person, across situations (aggregated over or abstracted from behaviour), and stable from a certain age onwards (e.g., veenman, 2007; winne & perry, 2000). the event approach, on the other hand, conceptualizes srl as a cyclical process unfolding in roughly three phases (forethought, performance and evaluation) (e.g., boekaerts, 1992), influenced by variables internal and external to the learner (e.g., winne & hadwin, 1998). additionally, the event approach also sees srl as covert in nature and so requires inferencing through learners’ behaviours and behavioural consequences (e.g., veenman, prins, & verheij, 2003). although both approaches are still used in research, over the past three decades the event approach gained considerable interest and dominated the investigation of learners’ srl (see: puustinen & pulkkinen, 2001). in line with the shift of conceptualization of srl, also the conceptualization of measurement approaches to capture it evolved. measurements shifted from single measurements administered before or after the execution of a task to continuous measurements administered during the execution of the task (winne & perry, 2000). the latter type of measurement is referred to as an on-line measurement whereas the former is referred to as the off-line measurement of srl (pintrich, 2004). following the shift in conceptualization of the srl concept, the use of offline measurements based on learners’ perceptions (e.g., self-reports) came under stress (endedijk et al., 2016). this is mainly because these types of measurements assume learners are capable to predict, reflect, or estimate in general terms (prior or after a task) how they will act in a certain context and subsequently rely on learners’ perceptions about their own srl rather than on the actual account of srl they exhibit (veenman, bavelaar, de wolf, & van haaren, 2014). the interest in sequence data and sequence analysis in srl taps in into the cyclical nature of srl and has been particularly fuelled by improvements in technical capabilities. the recording of learning-related behavioural data that are suitable for quantitative analysis has become almost effortless and unobtrusive for the learners in computer-based learning environments (winne, nesbit, & popowich, 2017), making this type of data particularly interesting for both practice and research. examples of this usefulness are, the mining of theory-based patterns from big data to identifying srl strategies in massive open online courses (maldonado-mahauad, pérez-sanagustín, kizilcec, morales, & munoz-gama, 2018), the finding of traces of srl in activity streams (cicchinelli et al., 2018), or the assessment of online learning material and its relation to learners’ quantitative behaviour patterns and their effects on motivation and learning performance (yang, li, & xing, 2018). applications are plentiful. 1.3 sequence analysis in the field of self-regulated learning in the field of srl in general sequence analysis refers to a sequence as an ordering of observable behavioural events preceded and followed by an unknown behavioural state (e.g., du, plaisant, spring, & shneiderman, 2016; köck & paramythis, 2011). simply put, each change of state is an event, and each event implies a change of state (müller, studer, gabadinho, & ritschard, 2010). for example, an assumed behavioural state could be reading a content page in an online learning environment, while clicking the calendar tool would be an observable behavioural event that changes the behavioural state of a learner to viewing the calendar page. through the investigation of ordered observable behavioural events (sequence), researchers try to gain insights in the unknown behavioural states learners are in (molenaar & järvelä, 2014). this investigation leads to three types of research questions: (1) questions about the nature of the sequences of the observed events, (2) questions about variables that affect those sequences, and (3) questions about the variables affected by the sequences (abbott & tsay, 2000). to gain a brief insight into the variety of approaches that can be used to handle each of these questions, below we illustrate how the investigation of them can be operationalized. this will be done based on three common phases in the investigation of sequential data in the field of srl (e.g., liu et al., 2016; zhou, 2016). these phases are: (1) the pre-processing phase, (2) the mining and characterization phase, and (3) the analysis phase. secondly, we provide an illustrative example highlighting (1) the complexity of sequence analysis, (2) the theoretical and methodological choices and considerations to be made, and (3) the reporting of the methods used, illustrating the need for agreed upon frameworks to be able to conduct and report sequence analyses transparently. 1.3.1 data structure as the raw computer log file data gathered functions as the input and absolute basis for sequence analysis (coronel & morris, 2016), we will start with the description of the data structure of sequence data before elaborating on the different phases of sequence analysis itself. the most common data format of raw computer-log-file data is the time stamped event (tse) format (gabadinho, ritschard, mueller, & studer, 2011). a tse-dataset contains at least three columns: (1) the timestamp of the observable behavioural event, (2) a personal identifier of the learner, and (3) an event name (of the observable behavioural event). examples of event names are the names of each element in the online learning environment (e.g., discussion form, content page, exercise, etc.), areas on the screen learners clicked, or clicking actions (e.g., caprotti, 2017; cicchinelli et al., 2018; maldonado-mahauad et al., 2018). 1.3.2 pre-processing in a first phase of the sequence-analysis method the raw data is pre-processed (e.g., zhou, 2016). this pre-processing phase generally consists of two steps. the first step relates to the question: “is there a need for recoding the raw data?” if there is a need (i.e., deductive approaches) to recode the raw data, this happens using an action library. such a library specifies the links between the observed events and the recoded, conceptualized concept. examples of these practices include action libraries based on the strict recoding of events or clusters of events using srl theories (e.g., winne et al., 2017). this depends on whether or not the action library is based on think-aloud coding schemes or on pure theoretical conceptualizations (e.g., bannert et al., 2015; taub et al., 2017). another illustration is the coding of computer log files using a tool-related coding scheme (e.g., lust, 2012; lust, vandewaetere, ceulemans, elen, & clarebout, 2011; siadaty, gašević, & hatala, 2016). if there is no need (i.e., inductive approaches) to recode the data, the raw data (as is) can be used (e.g., kurki, järvenoja, järvelä, & mykkänen, 2017; van laer & elen, 2016). the second step in the pre-processing phase involves assigning an ordered list of events to each learner, resulting in a single sequence per user (gabadinho et al., 2011). while the chronological ordering of the observed events suffices for the investigation of the sequential nature of such a sequence, the investigation of the temporal characteristics of a sequence will require the calculation and addition of the distance (time) between consecutive events to the sequence. based on the compilation of a single sequence per learner, sub-sequences and models can be mined and characterized. 1.3.3 mining and characterization after the raw data is pre-processed to a single sequence per learner, research questions related to the characteristics of sequences can be investigated. research questions are plentiful and pertain to the investigation of either whole sequence (β) or sub-sequences (α). a sub-sequence (α) is part of a sequence (β) if the sub-sequence (α) can either directly (α = < b, d, e, g >) or indirectly (α = < c, e, g, e >) be formed from the sequence (β = < a, c, b, d, e, g, c, e, d, b, g >) (zhou et al., 2010). the most common approach to mining and characterizing sequences and sub-sequences is called the algorithmic approach (e.g., kinnebrew, loretz, & biswas, 2013; perez et al., 2017; poole, lambert, murase, asencio, & mcdonald, 2016). this approach assumes the relation between the different events is unknown and therefore attempts to create meaning from the events that have already occurred by investigating the statistical relationships among them (breiman, 2001). efficient algorithms for discovering these characteristics have been proposed in statistical literature. the prominent algorithms are those of bettini, wang, and jajodia (1996), srikant and agrawal (1996), mannila, toivonen, and verkamo (1997), zaki (2001) and masseglia, teisseire, and poncelet (2002). all the algorithms require parameter settings. examples of these parameter settings are (1) time constraints of the occurrence of an event or sub-sequence, (2) a method for counting the occurrences of events and sub-sequences, and (3) a threshold for the identification of frequently occurring events and sub-sequences. once the parameters are defined, of-the-shelve software tools makes it possible to apply algorithmic approaches and to identify typical sequences (models), frequent events, and frequent sub-sequences. common platforms for performing this identification include prom (process mining workbench), developed by van der aalst (2016), spam (sequential pattern mining) by ayres, flannick, gehrke, and yiu (2002), or the traminer (trace mining in r) package developed for r-statistics by gabadinho et al. (2011). an extensive overview of algorithmic tools can be found in slater, joksimović, kovanovic, baker, and gasevic (2017). besides the algorithmic approach described above, there are also other approaches (for an extensive overview see: poole et al. (2016)). examples are theory-driven approach (e.g., cleary, 2011) which hypothesize the characteristics of sequences and sub-sequences and stochastic approach focussing on whole-sequence modelling (e.g., biswas, jeong, kinnebrew, sulcer, & roscoe, 2010; jeong, biswas, johnson, & howard, 2010). once sequences and sub-sequences are mined for and characterized they can be used as variables in statistical trials. 1.3.4 analysis when investigating sequences in the light of srl, we may be interested to know how sequences or sub-sequences are impacted by variables internal or external to the learners (e.g., winne & baker, 2013). for example when providing an instructional intervention to learners, we might want not only to see the change in learners’ learning outcomes but also the change in the occurrence of particular sequences or sub-sequences. another example might be that we want to compare sequences or sub-sequence of learners with low or high motivation (e.g., duffy & azevedo, 2015; jovanović, gašević, dawson, pardo, & mirriahi, 2017). in other words we may want to explore which sequences or sub-sequences discriminate most when different groups’ sub-sequences or averaged sequence are compared. to answer such questions, various approaches have been proposed for the incorporation of sub-sequences and sequences as dependent variables. the approach of studer, mueller, ritschard, and gabadinho (2010) consists of measuring the strength of association of each sequence or sub-sequence with the considered covariate and selects the sequence or sub-sequences with the strongest association. the association is measured with the pearson independence chi-square. the most discriminant one is the one with the highest chi-square. another approach proposed by kinnebrew et al. (2013) relies on multiple comparisons by t-test statistics between groups based on the considered covariate. the t-test is not used to prove that the groups of sequences differ. instead, it is employed as a heuristic for identifying more interesting sub-sequences in an exploratory analysis. this is done for example by determining with 95% confidence that a frequent sub-sequence is shown to be different between the groups. besides these common methods multi-level modelling (e.g., taub et al., 2016), regression analyses (e.g., segedy, kinnebrew, & biswas, 2015), or spearman correlation analysis (e.g., kizilcec, pérez-sanagustín, & maldonado, 2017) are also used. besides the investigation of variables influencing sequences, we can also investigate the influence of sequences on another variable. in the field of srl an example could be the impact of the occurrence of a specific sequence on group performance (e.g., molenaar & chiu, 2015). such research questions investigate the dissimilarities between different sequences (e.g., abbott & tsay, 2000; aisenbrey & fasang, 2010). these dissimilarities are commonly measured using the optimal matching edit distance. the optimal matching edit distance is defined as the minimal cost of transforming one sequence into the other (e.g., biemann & wolf, 2009; mazon, rossi, & toledo, 2014). the transformation operations considered by the optimal matching edit distance are (1) the insertion / deletion cost and (2) a change in the temporal distance resulting in the transformation from one sequence or sub-sequence to another. event dependent costs can be specified both for the insertion/deletion of an event as well as for a one-unit change in the temporal distance of given events. both the insertion / deletion and temporal distance cost result in a distance matrix between sequences themselves. this matrix can then be used in classification methods as well as in scaling methods to investigate the relation between various sequences (e.g., maldonado-mahauad et al., 2018; segedy et al., 2015). 1.4 an illustrative example of sequence analysis earlier we provided a condensed overview of different choices to be made at each phase of the sequence analysis process. to further illustrate the complexity of sequence analysis, the choices to be made, and the reporting of the methods used we provide an example of a study applying sequence analysis. in the study presented, we investigated the impact of reflection cues on learners’ srl. an event approach to srl was chosen focussing on srls’ cyclical, influenceable, and covert nature. srl was operationalized through learners’ learning behaviour and learners’ learning outcomes. two research questions were addressed: the first one investigated the impact of reflection cues on (a) learners’ learning behaviour and (b) on learners’ learning outcomes. the second research question investigated how learners’ learning outcomes related to by learners’ learning behaviour. to answer these questions, a 2x2 mixed factorial design was applied and data was gathered from 60 learners in second chance adult education. half of the group was exposed to additional cues for reflection; the learners in the control group were not. learners’ behavioural data existed of computer-log-file data gathered through an online learning environment in an ecologically valid setting. learners’ learning outcomes were assessed through cognitive (domain knowledge), motivational (goal orientation), and metacognitive (learning effort and learning confidence) tests and questionnaires. the computer-log-file data gathered had the tse-format. the event names were actions learners’ could perform in the online learning environment (i.e., post in the discussion forum; submit assignment; etc.). as a unit of analysis we used the entire eight week course and instructional stability throughout the eight weeks was described using the instrument of van laer and elen (2018). no validated operationalizations of sequence analysis based on the conceptualizations of the cyclical, influenceable, or covert nature of srl could be retrieved to direct the operationalization of our investigation. to deal with the issue of the lack of operationalizations, we decided to follow an approach staying as close to the observed data as possible. an inductive rather than a deductive approach was followed to avoid non-transparent alignment between conceptualization and operationalization. in line with this approach, we limited the assumptions made by (1) taking only into account observed overt events, (2) focussing only on the sequential aspects of computer-log-file data, and by (3) not conceptualizing the evolution of srl through the course of srl resulting on directly observable patterns via frequent sub-sequences rather than the extraction of behavioural models. the pre-processing of the data resulted in one (eight week long, +/10000 events) sequence of ordered raw behavioural events per learner. no recoding was applied, nor was time between events calculated. in the mining and characterization the traminer algorithm (gabadinho et al., 2011) was used in r-statistics to investigate learners’ sequences through the investigation of directly observable patterns via frequent sub-sequences. the identification of frequent sub-sequences was based on (1) the time constraints of the occurrence of events in the observed sub-sequences, (2) a counting method for counting the occurrences of sub-sequences, and on (3) a threshold for the identification of frequently occurring sub-sequences. as only directly observable sub-sequences were targeted, the parameter for the distance between events was set to one, representing that only events directly observed before or after a certain event could be seen as part of a sub-sequence. the counting method chosen was selected arbitrary, based on the occurrence of sub-sequences over the different learners. the frequency threshold was set to 25% meaning that at least 25% of the learners should exhibit the sub-sequence to be counted as frequently occurring. 688 frequent sub sequences were observed. next, we investigated frequent sub-sequences’ relationship with (1) the condition learners were in (impact of cues on behaviour) and (2) learners’ learning outcomes (relations between outcomes and behaviour). in the analysis phase, the frequent sub-sequences were used as dependent variables. for this analysis, chi-square tests were used containing the frequent sub-sequences for discriminating the groups and the variables that defines the groups (condition and learners’ learning outcomes). based on these tests, the effect sizes were calculated using cramer’s v. the cramer’s v expresses the relation between a certain discriminating frequent sub-sequence and the learners’ characteristics and is reported in a value between zero and one. the closer to one the higher the relation is. cohen (1988) refers to small (≤.30), medium (≥.30 and ≤50), and large (≥.50) effect sizes. with regard to the first research question dealing with the investigation of the impact of reflection cues on (a) learners’ learning behaviour and (b) on learners’ learning outcomes, learners in the experimental condition were shown to make significantly more use of sub-sequences consisting of events related to assignments and tasks, communication, and assessment. furthermore, both conditions showed a significant increase in domain knowledge and learning confidence and a decrease in performance goal approach. learners in the experimental condition who received cues for reflection scored significantly higher on performance goal approach compared to the learners in the control condition. as for the interaction effect between time and condition, learners in the experimental condition scored significantly higher for performance avoidance approach compared to their counterparts. this result was unexpected in the light of the aim of the study (van laer et al., 2018). finally, with regard to the second research question dealing with the investigation of how learners’ learning outcomes related to by learners’ learning behaviour, it became clear that changes in learning behaviour seemed to be linked to learning outcomes (performance avoidance approach). results showed that differences in learners’ learning behaviour were observed when learners had different performance avoidance approach scores. 1.5 towards tangible proof of progress as research aims at either at building or testing theory, the research cycle moves from description, to explanation, to testing with repeated iterations through this cycle (van der merwe, 2013). throughout this iterative process, descriptive models are expanded into explanatory frameworks that are tested against reality until they are eventually developed into theories as research study builds upon research study. the result is to validate and add confidence to previous findings, or else invalidate them and force researchers to develop more valid or more complete theories (meredith, 1993). in this way both (1) theoretical conceptualizations of the theory under investigation and their (2) operationalization through measurements are continuously updated and refined. as illustrated throughout the different paragraphs presented above, different operationalizations of sequence analysis can be made. the illustrative example has shown one of these operationalizations. to be able to monitor methodologies’ evolution towards tangible proof of progress and so secure the iterative research cycle, the literature on advances in research methodology (e.g., beach & pedersen, 2013; lupia & alter, 2014) proposes three indications of such an evolution. the first one is the transparency of the method applied (moravcsik, 2014). transparently reported methods permit scholars to assess research and to communicate with one another. unless other scholars can examine evidence, parse the analysis, and understand the processes by which evidence and theories were chosen, why should they trust and thus expend the time and effort to scrutinize, critique, debate, or extend existing research? as demonstrated earlier, a lot of explorative work on the use of sequence analysis has been done, yet most of the studies do not seem to report in detail on the different phases of sequence analysis or on the parameter settings involved in each of them, hampering a thorough study of the method applied. when literature on the investigation of srl through sequence analysis reports on the log file data structure (e.g., biswas, roscoe, jeong, & sulcer, 2009; duffy & azevedo, 2015; lazakidou & retalis, 2010), it often does this through elaborating on the events traced: clicks, pages, specific cognitive or metacognitive activities, and so on. despite the information on the events traced, additional information on the structure of the data, such as the timestamp interval or type of timestamps, session identifiers, etc., is hardly provided. this information is important to distinguish which pre-processing steps are possible or desirable (e.g., calculation of time between events, grouping of learners or individuals, etc.). with regard to this pre-processing phase, in the best cases researchers acknowledge they developed a set of filters or recoding algorithms to remove irrelevant information from the raw log files, with the aim of presenting the relevant information in a compact format that is suitable for further analysis (e.g., jeske, backhaus, & stamov roßnagel, 2014; paans, molenaar, segers, & verhoeven, 2018). nonetheless they hardly ever elaborate on which information was discarded and what made the researcher assume this information could be classified as irrelevant. when for example action libraries are used (e.g., bannert et al., 2014; goldberg et al., 2014) researchers elaborate on the coding scheme, but lack to state how the ‘raw’ events are recoded and what the reliability of this recoding was like. without this information it is impossible to distinguish which coding scheme is most reliable and works best for what data. in line with this, no studies seem to be available which argue for the selection of a certain coding scheme or elaborate on why a certain coding scheme is preferable over another. with regard to the mining and characterization of sequences, most of the current literature seems to indicate which algorithms are used to identify or mine (frequent) sub-sequences. nonetheless, the authors rarely seem to address the assumptions underlying the algorithm used (e.g., balderas, dodero, palomo-duarte, & ruiz-rube, 2015; lan & lu, 2017) or the procedure followed to select the appropriate algorithm (e.g., kizilcec et al., 2017; maldonado-mahauad et al., 2018), let alone the parameter settings applied when mining for sub-sequences. the same is the case when the (frequent) sub-sequences are used as dependent or independent variables. the analysis methods used to answer similar research questions vary from researcher to researcher (ahmadpour & khaasteh, 2017; cerezo, esteban, sánchez-santillán, & núñez, 2017; chen, breslow, & deboer, 2018). traditional cluster analysis and predictive apriori algorithms are used to identify sets of successful learner and environmental characteristics impacting performance, without any explanation of why a certain approach might be considered superior to another. the multitude of methods and the observation of the lack of transparency make the studies unreplicable. the second characteristic relates to the availability of comparative research designs (e.g., bureau & salomonsen, 2012; peterson, 2005). comparison is one of the most powerful tools used in intellectual inquiry, since an observation made repeatedly is given more credence than a single observation. put simply, as argued by mills, van de bunt, and de bruijn (2006) the main goal of comparative research is to search for or identify variance or similarity. although there is quite some comparative research on the measurement of srl, the majority of it focusses at best on the comparison of online behaviour event measurements (i.e., sequence analysis) with offline perception event measurements (i.e., self-reports) (e.g., cho & yoo, 2017; hadwin, nesbit, jamieson-noel, code, & winne, 2007). even at the most basic level of comparison, namely the use of different coding schemes to recode the ‘raw’ event captured in log files, there seems to be hardly any evidence on which coding scheme results in the most accurate results under which conditions (azevedo, 2014). although there are useful summaries of approaches and tools (e.g., slater et al., 2017) as well as ample ideas on how to apply sequence analysis (e.g., azevedo et al., 2010; winne, 2018; winne et al., 2017), no discussion of different sequence analysis methods could be found in the field of srl. based on this observation, research comparing different sequence analysis approaches for srl seems to be missing. this would lead to the identification of commonalities and differences between methods adding to the validation of the method. the third and final characteristic is the application of a method in empirical ecologically valid settings (e.g., chambless & ollendick, 2001; rotter, 1954). only when methods can be applied in different contexts and situations, they can propel and more importantly validate investigations. although there have been attempts in the field of self-regulated learning to operationalize insights drawn from experimental settings in ecologically valid empirical contexts, these attempts are mainly based on a mixture of insights obtained from the experimental setting, accompanied by a data-driven approach to overcome the gaps left by the experimental approach (e.g., parameter settings, coding events, identification of sub-sequences) (e.g., hsu, 2018; ifenthaler, gibson, & dobozy, 2018; taub, azevedo, bradbury, millar, & lester, 2018). no clear attempts to apply and transfer insights between settings seem to be made so far. in summary, it becomes clear that a lot of work still needs to be done. it seems that none of the three indications for tangible proof of progress already has been achieved for the use of sequence analysis in the field of srl. in line with this finding already in 2014, roger azevedo (2014) pointed out that researchers investigating sequence data recoded the data different, made diverse statistical and theoretical assumptions regarding the data collected, and that too easily inferences were drawn from the sequential and temporal unfolding data. his call for action was in vane and repeated multiple times (e.g., molenaar & järvelä, 2014; winne et al., 2017), supported by the expression of the need for standards and frameworks to align the investigation of sequence data in the field of srl (e.g., bannert et al., 2015). partly, the aim of this manuscript is to add to the body of literature calling for standards, protocols, and frameworks that can be tested and validated. by providing general guidelines contributing to a framework for the use and reporting of sequence analysis for srl this manuscript aims to propel the establishment of sequence analysis as a research method. 1.6 problem statement as illustrated in the current section, sequence analysis in the field of srl is an umbrella term covering a large variety of approaches. with regard to the log file data format and the pre-procession phase it often seems unclear how researchers devise and deploy the data scrubbing, cleansing, recoding, or the cleaning processes (e.g., clarke, 2016; müller, naumann, & freytag, 2003; rahm & do, 2000). with regard to the mining and characterization phase and to the analysis phase hardly any explicit references seem to be made to the basis of specific decisions. current research makes it hard to distinguish which parameter settings were derived from literature or which ones were set arbitrary, nor why this is the case. no arguments are given on why a certain approach is considered above another one (e.g., poole et al., 2016; stark & vedres, 2012). the multitude of approaches and considerations does not yet seem to be condensed into a transparent methodological framework for sequence analysis, nor do these approaches seem to contribute yet to tangible proof of progress in the investigation of learners’ srl using sequence analysis based on computer log files. without such a framework it is impossible to test, falsify or modify approaches to sequence analysis and so spark the investigation of learning through sequence analysis. in the next section we propose a set of guidelines functioning as a potential starting point for the construction of a transparent methodological framework to communicate approaches of sequence analysis in the field of srl. 2. guidelines on the use of sequence analysis in the field of self-regulated learning as illustrated throughout the introductory section of this manuscript, the operationalization of sequence analysis to investigate learners’ srl using computer log files is shaped through many practical choices and theoretical assumptions. whereas in the previous sections we highlighted the need for a methodological framework, in what follows we propose a set of guidelines functioning as a potential starting point for the construction of such a framework. we offer guidelines in two main areas. the first one relates to the alignment of the conceptualization and the operationalization of the different components of srl. the second one relates to the enactment of the operationalization of the selected sequence analysis approach. 2.1 alignment of conceptualization and operationalization singleton, straits, and straits (1999) see alignment of conceptualization and operationalization as one of the key features to scientific and methodological success. they refer to the process of conceptualization as the act of defining the different components of a phenomenon under investigation and to operationalization as the practical result of the conceptualization act. in line with this definition, we explore the operational impact of the current conceptualization of srl. as presented in the introduction, current conceptualizations of srl focus on its cyclical, influenceable, and covert nature (e.g., winne & hadwin, 1998). in what follows we relate these three general conceptions to practical consequences in the operationalization of the selected sequence analysis approach chosen. even when these three conceptions are not at the basis of the conceptualization of srl, the conceptions below might shed light on the relation between on the one hand the conceptualization of the phenomenon under investigation and the selected operationalization of sequence analysis. 2.1.1 the cyclical nature of self-regulated learning the idea of srl phases unfolding in different cyclical phases raises questions concerning (1) the dynamics of these cyclical srl process, (2) the sequential patterns within it, as well as (3) the development of the cycle over time. each of these questions bring the notion of sequentiality and temporality to the discourse on srl (e.g., molenaar & järvelä, 2014). while in the literature on srl the ‘temporal’ and ‘sequential’ notion is often used intertwined (knight, wise, & chen, 2017), literature on sequence analysis from other fields of research makes a clear distinction. temporality refers to the passage of elapsed time and comes with a collection of related concepts such as duration, rate, and acceleration (blikstein, 2011; haythornthwaite & gruzd, 2012). sequentiality is used to refer to the order of events and transitions between different events, without explicit reference needed to duration or passages of time (biswas et al., 2010; halatchliyski, hecking, goehnert, & hoppe, 2014). as a first consideration, when for example incorporating only sequential characteristics of srl, the construction of a single sequence per learner in the pre-processing phase will consist of the chronological ordering of events, assuming time between each events’ timestamp of secondary interest. when in contrast also temporal characteristics of srl are taken into account, the construction of learners’ single sequences will include the calculation of the time between the consecutive events’ timestamps and the inclusion of these calculations in further analyses. the latter poses additional conceptual questions to the status of this calculation. does it for example represents a single hidden unknown state or is it instead an indication of involvement with the environment? a second consideration relates to the developmental characteristic of the behaviour observed. when srl-development over time is assumed (e;g., andrade & evans, 2015; huang, klein, & beck, 2017) a whole-sequence approach might be preferred over a sub-sequence approach that does not consider such a development (in the time frame of investigation) (winne & hadwin, 1998). both will affect further analyses. as demonstrated above while briefly reviewing the conceptualization of the cyclical nature of srl, a clear link between (1) the conceptualization of the sequential and temporal characteristics of srl and its practical operationalization and (2) the developmental characteristics of srl over time and its operationalization in practice seem necessary to be able to study the sequence analyses applied. 2.1.2 the influenceable nature of self-regulated learning with regard to how srl comes to be, recent event theories regard srl as influenced by variables internal and external to the learner (veenman, van hout-wolters, & afflerbach, 2006). in general research identifies three major sets of internal variables influencing srl: cognitive (e.g., zimmerman, 1986, 1990, 1998; zimmerman & pons, 1986), metacognitive (borkowski, carr, rellinger, & pressley, 1990; pressley, levin, & mcdaniel, 1987) and motivational variables (e.g., butler & winne, 1995; schraw, crippen, & hartley, 2006; schraw & moshman, 1995; zimmerman, 2000). a substantial body of literature identifies external variables at different grainsize levels influencing learners’ srl. dignath and büttner (2008) for example point in their meta-analysis that (1) instruction of cognitive strategies (i.e., rehearsal, elaboration, and organizational strategies) affected learners’ srl significantly. the same was observed for (2) instruction of metacognitive strategies (i.e., planning, monitoring, and evaluation), (3) promoting metacognitive reflection, and (4) instruction of motivation strategies. another example is the literature review of van laer and elen (2016) identifying seven attributes of learning environments that support learners’ srl. the combinations of the abovementioned internal and external variables make up the timeframe in which srl needs to be investigated. under this conceptualization, each change in either variables internal and / or external to the learner will influence learners’ srl (e.g., greene & azevedo, 2007; winne & hadwin, 1998). thus without the appropriate timeframe in which to investigate learners’ srl insights might be hard to gather. so, the main consideration with regard to the influenceable nature of srl relates to the unit of analysis. the size of the allowed timeframe affects the operationalization of sequence analysis for example through the parameter settings while mining and characterizing sequence and sub-sequence. when for example both internal and external variables are regarded as stable throughout the investigative trial the timeframe might stretch over the entire trial. if in contrast the variables are assumed to be variable at a certain rate, the timeframe might want to match this rate as much as possible. as illustrated before the characterization of sequences and sub-sequences relies amongst others on specifications with regard to the time constraints of the occurrence of an event or sub-sequence. the time dimension raises questions concerning the maximal distance between events (molenaar & järvelä, 2014). moreover, we may consider two or more events form a relevant sequence or sub-sequence only if they occur within a given distance of each other (maximal timespan). for example when they occur in a time window in which both variables within and external to the learner are assumed to be constant. for example completing an exercise right after viewing a content related page might not mean the same as completing that same exercise twenty events after viewing the content related page (e.g., du et al., 2016; jovanovic, pardo, mirriahi, dawson, & gašević, 2017). to conclude, it seems that to be able to study the sequence analyses applied, the operationalization of the unit of analysis (e.g., through parameter settings) as conceptualized through the influenceable nature of srl needs to be elaborated. elaborating on the relation between unit of analysis and parameter settings allows us to assess the suitability of the decisions made. 2.1.3 covert nature of self-regulated learning current conceptualizations assume srl operates at different levels of the cognitive system and so regulates lower order cognitive processes that, in turn, shape learners’ overt cognitive behaviour (roth, ogrin, & schmitz, 2016). this conceptualization results in the assumption that srl occurring in each of the srl phases cannot be directly observed as it manifests in overt cognitive behaviours (williamson, 2015) and through behavioural consequences like learners’ learning outcomes (veenman & alexander, 2011). for instance, when a learner recalculates the outcome of a mathematical equation, it is assumed that a srl monitoring or evaluation process must have preceded this overt cognitive activity of recalculation. as illustrated in the introductory section, conceptualization of the covert nature of srl can be based on the relation between the overt behavioural events with srl-influencing constructs (e.g., tool-use, engagement, etc.) or directly with srl-related activities (e.g., goal-setting and planning, monitoring, etc.) (e.g., azevedo et al., 2015; bannert et al., 2015; lust, 2012). depending on the conceptualization of (a) the covert nature of srl and (b) how this nature can be uncovered through the overt behavioural events observed through computer log files, a link will be constituted with the operationalization of this covert nature through the establishment of action libraries (zhou, 2016). when an action library is used, overt behavioural events (or sets of events) are recoded into more meaningful learner behaviour or srl behaviour. although this practice is common when following a deductive research approach, action libraries often have a different level of granularity as they are developed through other online event measurements (i.e., think aloud trials) (azevedo, 2014). comparing the micro-level approach followed through using computer log files with the codes abstracted from think aloud trails might be problematic given the different grain-size (e.g., al mamun, lawrie, & wright, 2017). in summary, when operationalizing the covert nature of srl via recoding data using different grainsizes (i.e., computer log files and think aloud data), clearly the relationship between observed behaviour and assigned codes needs to be made explicit and communicated transparently. 2.2 enactment of the operationalization of sequence analysis proctor, powell, and mcmillen (2013) see enactment as the step following the conceptualization and operationalization of a research method. they identify enactment as the systematic application of the operationalized concepts. as illustrated before, the quest for tangible proof of progress lies in transparently reported methods and procedures permitting scholars to assess research and to communicate with one another (e.g., beach & pedersen, 2013; lupia & alter, 2014). with regard to the enactment of the operationalization of the selected sequence approach, we focus on two components: systematic account of the operationalization and transparent parameter settings. 2.2.1systematic account of the operationalization the operationalization of sequence analysis starts from gathered data with a specific structure and unfolds roughly in three phases. in the first phase, the pre-processing phase, a single sequence per learner is constructed. in the second phase, sequences and frequent sub-sequences are mined and characterized. identified sequences and sub-sequences can function as either dependent or independent variables. although general approaches such as the one described above have been usefully proposed by for example zhou (2016) and liu et al. (2016), current research fails to go much beyond the vagueness level of this general approach. as demonstrated a multitude of decisions need to be made in the chain of sequence analysis (roll & winne, 2015). keeping systematicity and transparency in mind, detailed accounts of each of these phases would increase replicability (e.g., moravcsik, 2014). systematic accounts of the enactment of the operationalization of sequence analysis might start with a description of the raw data gathered. this not only means reporting on the environment in which the data was gathered, but also the actual structure of the dataset extracted, including for example the database structure. in the pre-processing phase systematicity and transparency might be accomplished through elaborating (among others) on the data cleaning process, the recoding procedure (if applicable), and the transformations applied (if applicable). with regard to the mining and characterization of the sequences and sub-sequences this might be done through a clear account of the different steps taken in the mining and characterization process, including for example a detailed explanation of the algorithm used and the parameters set. finally in the analysis phase transparency and systematicity might be accomplished by the presentation of the output format of the previous phase and the analysis approach chosen (with its key figures). 2.2.2 transparent parameter settings it is clear that the conceptualization of a theory cannot account for each variation in operationalization nor for the justification of each parameter setting (bannert et al., 2017). nonetheless, decisions need to be made to derive useful approaches. regardless of the inductive or deductive approach to the conceptualization of srl, transparency and courage on the part of the researchers is essential to report in detail which parameter settings are derived from theory and which are arbitrary (e.g., tsai, shen, & tsai, 2011). the degree to which this is possible on either side is irrelevant to the argument, as long it is clear which parameter settings are used for what reason or considering what assumption. an example of such a practice was presented in the illustrative case. no theoretical evidence could be found to determine the frequency threshold for identifying frequent sub-sequences. in this case, it was reported that an arbitrary cut-off was set at 25%. 3. implications and conclusions although the use of sequence analysis in the field of self-regulation is one of the last decade, a lot of valuable work has been done to propel the investigation of learners’ srl through sequence analysis. despite these efforts, no methodological framework seems to be available for the systematic application and reporting of sequence analysis in the field of srl. because of the lack of such a framework, tangible proof of progress is difficult to achieve and so the evolution of sequence analysis to investigate srl seems to be hampered. to illustrate the need for such a methodological framework we provided in the introduction of this manuscript a brief overview of the variability of current operationalizations and illustrated one such approach. from the pre-processing phase onwards, over the mining and characterization phase, up to the use of the identified sequences and sub-sequences as dependent and independent variables, a multitude of conceptual and operational choices need to be made. therefore, in this manuscript we aimed to foster discussion on a methodological framework for the application of sequence analysis in the field of srl which would make replication, falsification, and validation possible. to do so, in addition to the case built in the introduction section, we proposed in the previous section a set of guidelines functioning as a potential starting point for the construction of a framework. these guidelines were centred on two key areas. the first area focussed on the alignment of the conceptualization of the different components of srl and the operationalization of the selected sequence analysis approach (e.g., singleton et al., 1999). the other area focussed on the enactment of the operationalization of the selected sequence analysis approach (e.g., proctor et al., 2013). with regard to the former four guidelines were proposed relating to the current conceptualization of srl that relate to: (1) the sequential and temporal characteristics of srl; (2 the development through time of srl; (3) the unit of analysis imposed by the factors influencing srl; (4) the matching-granularity as linked to the covert nature of srl. with regard to the enactment of the operationalization of the sequence analysis approach two guidelines are proposed related to: (1) the systematic account of the operationalization; and (2) the transparent communication of parameter settings. although this manuscript does not pretend to provide solutions nor to be exhaustive with regard to possible approaches to sequences analysis in the field for srl, on the one hand it highlights the need for a transparent and systematic methodological approach. on the other hand, it also identifies guidelines that might function as a basis for further construction of such a methodological framework for the use of sequence analysis in the field of srl. keeping in mind the nature of the guidelines provided, we might wonder whether the guidelines are simply ‘common sense’ and applicable to many other research methods. the latter is certainly the case, yet guidelines have not been constructed before for the investigation of sequence analysis for srl. with regard to the ‘ordinariness’ of the guidelines provided in this manuscript, there is an abundance of systematic methodological literature reviews (e.g., kallio, pietilä, johnson, & kangasniemi, 2016; kelly, lesh, & baek, 2014; mertens, 2014) illustrating that guidelines similar to the ones suggested in this manuscript might be beneficial for a broad range of research methods and also that without such guidelines, the quest for tangible proof of progress is more than likely to be unsuccessful. 3.1 implications when a methodological framework is in place the investigation of time-related characteristics in srl, using sequence analysis might evolve to (1) transparent methods, (2) comparative studies, and (3) empirical and ecological applications, supporting both research and practice. such a framework enables researchers to use the framework to describe and compare current approaches to sequence analysis. a solid description of these approaches, their reproduction and validation first in similar, later in different contexts. the former will allow us to apply the insights gathered not only under very strict conditions in one particular situation but is likely to foster empirical and ecologically valid trails. the latter would be useful for researchers and practitioners using sequence analysis for example to inform the design of their course. by having a methodological framework at their disposal, selecting the most appropriate sequence analysis approach for their needs is facilitated. the sooner we are able to compare, validate, and establish sequence analysis methods, the more quickly we can make progress in the investigation of srl through learners’ learning behaviour. 3.2 further directions as it was not our aim to present a fully developed methodological framework for the use of sequence analysis in the field of srl, in future investigations it might be interesting to further detail the conceptual assumptions related to the investigation of time-related characteristics of srl and their relation to methodological operationalizations. this could be done by incorporating more theoretical research on the investigation of self-regulated learning and extract the possible methodological consequences for the operationalization of the theoretical conceptualizations proposed. another avenue might be the integration of non-content-related research on sequence analysis to further investigate the operational possibilities of the method to further investigate the conceptual assumptions made by choosing a particular approach over another. by doing so a protocol, standard, or framework can be established as a method for the use of sequence analysis which can then be tested, validated and modified to further optimize the use of sequence analysis for the investigation of srl. 3.3 conclusions first, the manuscript built a case for the need of a methodological framework for the application of sequence analysis in the field of self-regulated learning. secondly, it provided a ground for further discussion on the construction of such methodological framework, raising questions about both the conceptualization and the operationalization of sequence analysis in the field of srl. additionally it provides guidelines and possible directions supporting sequence analysis as a method in the field of srl. applying sequence analysis as a method in the field of srl in a more systematic and transparent way might support the development of the method towards more transparency, comparative studies, and empirical and ecological applications, supporting both research and practice. as demonstrated throughout the manuscript it seems that it is not the amount of data gathered that will help us gain insights, but rather the way we analyse them and the thoroughness of that analysis. this manuscript by no means implies that the overview of operationalizations is exhaustive or complete, nor does is pretend to provide a best practice or example through the illustrative example. instead, it aimed to provide a transparent and verifiable framework to discuss the method we see as potentially powerful for investigating learners’ srl. such a framework challenges the assumptions made, approaches taken, and thus propels the investigation of learners’ srl through computer log files to new heights. in conclusion, the guidelines proposed and their underlying call for transparency and systematicity through the construction of a methodological framework for the use of sequence analysis in the field of srl can potentially transcend the use of sequence analysis for computer log files and go as far as to other log file methods investigating the temporal and sequential nature of phenomena. as the literature on the investigation of for example eye movement (e.g., kiefer, giannopoulos, raubal, & duchowski, 2017; lorigo et al., 2008) or skin conduction (e.g., el‐sheikh, 2007; haufler et al., 2017) log files seems to experience similar issues, also these fields of research might benefit from the general guidelines formulated in the manuscript presented (e.g., kuhn, 2012). although the focus in this manuscript is only on one rather specific approach to sequence analysis, the quest for transparency and systematicity is one that relates to all investigations, especially when it comes to new complex methods that require abundant data processing in order to make them meaningful. keypoints substantial increase in the use sequence-analysis methods for investigating self-regulated learning. there is hardly any tangible proof of progress of sequences analysis as a method. alignment of conceptualization and operationalization and transparent enactment is desirable. when such a framework is in place, sequence analysis is more likely to grow as a method. acknowledgments we would like to acknowledge the support of the project “adult learners online” funded by the agency for science and technology (project number: sbo 140029), which made this research possible. references abbott, a., & tsay, a. (2000). sequence analysis and optimal matching methods in sociology: review and prospect. sociological methods & research, 29(1), 3-33. doi: 10.1177/0049124100029001001 ahmadpour, z., & khaasteh, r. (2017). writing behaviors and critical thinking styles: the case of blended learning. khazar journal of humanities & social sciences, 20(1). aisenbrey, s., & fasang, a. e. (2010). new life for old ideas: the" second wave" of sequence analysis bringing the" course" back into the life course. sociological methods & research, 38(3), 420-462. al mamun, m. a., lawrie, g., & wright, t. (2017). factors affecting student engagement in self-directed online learning module. paper presented at the proceedings of the australian conference on science and mathematics education andrade, m. s., & evans, n. w. (2015). developing self-regulated learners. esl readers and writers in higher education: understanding challenges, providing support , 113. antunes, c. m., & oliveira, a. l. (2001). temporal data mining: an overview.paper presented at the kdd workshop on temporal data mining. ayres, j., flannick, j., gehrke, j., & yiu, t. (2002). sequential pattern mining using a bitmap representation.paper presented at the proceedings of the eighth acm sigkdd international conference on knowledge discovery and data mining. azevedo, r. (2014). issues in dealing with sequential and temporal characteristics of self-and socially-regulated learning. metacognition and learning, 9(2), 217-228. azevedo, r., & hadwin, a. f. (2005). scaffolding self-regulated learning and metacognition–implications for the design of computer-based scaffolds: springer. azevedo, r., moos, d. c., johnson, a. m., & chauncey, a. d. (2010). measuring cognitive and metacognitive regulatory processes during hypermedia learning: issues and challenges. educational psychologist, 45(4), 210-223. doi: 10.1080/00461520.2010.515934 azevedo, r., taub, m., & mudrick, n. (2015). technologies supporting self-regulated learning. the sage encyclopedia of educational technology, 731-734. balderas, a., dodero, j. m., palomo-duarte, m., & ruiz-rube, i. (2015). a domain specific language for online learning competence assessments. international journal of engineering education, 31(3), 851-862. bandura, a. (1989). human agency in social cognitive theory. american psychologist, 44(9), 1175. bannert, m., molenaar, i., azevedo, r., järvelä, s., & gašević, d. (2017). relevance of learning analytics to measure and support students' learning in adaptive educational technologies. paper presented at the proceedings of the seventh international learning analytics & knowledge conference. bannert, m., reimann, p., & sonnenberg, c. (2014). process mining techniques for analysing patterns and strategies in students’ self-regulated learning. metacognition and learning, 9(2), 161-185. bannert, m., sonnenberg, c., mengelkamp, c., & pieger, e. (2015). short-and long-term effects of students’ self-directed metacognitive prompts on navigation behavior and learning performance. computers in human behavior, 52, 293-306. beach, d., & pedersen, r. b. (2013). process-tracing methods: foundations and guidelines: university of michigan press. beheshitha, s. s., gašević, d., & hatala, m. (2015). a process mining approach to linking the study of aptitude and event facets of self-regulated learning. paper presented at the proceedings of the fifth international conference on learning analytics and knowledge. ben-eliyahu, a., & bernacki, m. l. (2015). addressing complexities in self-regulated learning: a focus on contextual factors, contingencies, and dynamic relations. metacognition and learning, 10(1), 1-13. doi: 10.1007/s11409-015-9134-6 bettini, c., wang, x. s., & jajodia, s. (1996). testing complex temporal relationships involving multiple granularities and its application to data mining. paper presented at the proceedings of the fifteenth acm sigact-sigmod-sigart symposium on principles of database systems. biemann, t., & wolf, j. (2009). career patterns of top management team members in five countries: an optimal matching analysis. the international journal of human resource management, 20(5), 975-991. doi: 10.1080/09585190902850190 biswas, g., jeong, h., kinnebrew, j. s., sulcer, b., & roscoe, r. (2010). measuring self-regulated learning skills through social interactions in a teachable agent environment. research and practice in technology enhanced learning, 5(02), 123-152. biswas, g., roscoe, r., jeong, h., & sulcer, b. (2009). promoting self-regulated learning skills in agent-based learning environments. paper presented at the proceedings of the 17th international conference on computers in education. blikstein, p. (2011). using learning analytics to assess students' behavior in open-ended programming tasks. paper presented at the proceedings of the 1st international conference on learning analytics and knowledge. boekaerts, m. (1992). the adaptable learning process: initiating and maintaining behavioural change. applied psychology, 41(4), 377–397. doi: 10.1111/j.1464-0597.1992.tb00713.x bonin, f., vogel, c., & campbell, n. (2014). social sequence analysis: temporal sequences in interactional conversations. paper presented at the cognitive infocommunications (coginfocom), 2014 5th ieee conference on. borkowski, j. g., carr, m., rellinger, e., & pressley, m. (1990). self-regulated cognition: interdependence of metacognition, attributions, and self-esteem. dimensions of thinking and cognitive instruction, 1, 53-92. bourbonnais, s., hamel, e. b., lindsay, b. g., liu, c., stankiewitz, j., & truong, t. c. (2006). method, system, and program for merging log entries from multiple recovery log files: google patents. breiman, l. (2001). statistical modeling: the two cultures (with comments and a rejoinder by the author).statistical science, 16(3), 199-231. bureau, v., & salomonsen, h. h. (2012). comparing comparative research designs. butler, d. l., & winne, p. h. (1995). feedback and self-regulated learning: a theoretical synthesis. review of educational research, 65(3), 245-281. doi: 10.2307/1170684 caprotti, o. (2017). shapes of educational data in an online calculus course. journal of learning analytics, 4(2), 76-90. cerezo, r., esteban, m., sánchez-santillán, m., & núñez, j. c. (2017). procrastinating behavior in computer-based learning environments to predict performance: a case study in moodle. frontiers in psychology, 8, 1403. chambless, d. l., & ollendick, t. h. (2001). empirically supported psychological interventions: controversies and evidence. annual review of psychology, 52(1), 685-716. doi: 10.1146/annurev.psych.52.1.685 chen, x., breslow, l., & deboer, j. (2018). analyzing productive learning behaviors for students using immediate corrective feedback in a blended learning environment. computers & education, 117, 59-74. cho, m.-h., & yoo, j. s. (2017). exploring online students’ self-regulated learning with self-reported surveys and log files: a data mining approach. interactive learning environments, 25(8), 970-982. cicchinelli, a., veas, e., pardo, a., pammer-schindler, v., fessl, a., barreiros, c., & lindstädt, s. (2018). finding traces of self-regulated learning in activity streams. clarke, r. (2016). big data, big risks. information systems journal, 26(1), 77-90. doi: 10.1111/isj.12088 cleary, t. j. (2011). emergence of self-regulated learning microanalysis. handbook of self-regulation of learning and performance, 329-345. coronel, c., & morris, s. (2016). database systems: design, implementation, & management: cengage learning. daniela, p. (2015). the relationship between self-regulation, motivation and performance at secondary school students. procedia-social and behavioral sciences, 191, 2549-2553. doi: 10.1016/j.sbspro.2015.04.410 dignath, c., & büttner, g. (2008). components of fostering self-regulated learning among students. a meta-analysis on intervention studies at primary and secondary school level. metacognition and learning, 3(3), 231-264. doi: 10.1007/s11409-008-9029-x du, f., plaisant, c., spring, n., & shneiderman, b. (2016). eventaction: visual analytics for temporal event sequence recommendation. paper presented at the visual analytics science and technology (vast), 2016 ieee conference on. duffy, m. c., & azevedo, r. (2015). motivation matters: interactions between achievement goals and agent scaffolding for self-regulated learning within an intelligent tutoring system. computers in human behavior, 52, 338-348. el‐sheikh, m. (2007). children's skin conductance level and reactivity: are these measures stable over time and across tasks? developmental psychobiology, 49(2), 180-186. endedijk, m. d., brekelmans, m., sleegers, p., & vermunt, j. d. (2016). measuring students’ self-regulated learning in professional education: bridging the gap between event and aptitude measurements. quality & quantity, 50(5), 2141-2164. gabadinho, a., ritschard, g., mueller, n. s., & studer, m. (2011). analyzing and visualizing state sequences in r with traminer. journal of statistical software, 40(4), 1-37. goldberg, b., sottilare, r., roll, i., lajoie, s., poitras, e., biswas, g., . . . long, y. (2014). enhancing self-regulated learning through metacognitively-aware intelligent tutoring systems: boulder, co: international society of the learning sciences. greene, j. a., & azevedo, r. (2007). a theoretical review of winne and hadwin's model of self-regulated learning: new perspectives and directions. review of educational research, 77(3), 334-372. doi: 10.3102/003465430303953 hadwin, a. f., nesbit, j. c., jamieson-noel, d., code, j., & winne, p. h. (2007). examining trace data to explore self-regulated learning. metacognition and learning, 2(2-3), 107-124. halatchliyski, i., hecking, t., goehnert, t., & hoppe, h. u. (2014). analyzing the main paths of knowledge evolution and contributor roles in an open learning community. journal of learning analytics, 1(2), 72-93. haufler, a. j., lewis, g. f., davila, m. i., westhelle, f., gavrilis, j., bryce, c. i., . . . mcdaniel, w. (2017). biobehavioral insights into adaptive behavior in complex and dynamic operational settings: lessons learned from the soldier performance and effective, adaptable response (spear) task. frontiers in medicine, 4, 217. haythornthwaite, c., & gruzd, a. (2012). exploring patterns and configurations in networked learning texts. paper presented at the system science (hicss), 2012 45th hawaii international conference on. hine, c. (2011). internet research and unobtrusive methods. social research update(61), 1. hsu, t.-c. (2018). behavioural sequential analysis of using an instant response application to enhance peer interactions in a flipped classroom. interactive learning environments, 26(1), 91-105. huang, n., klein, m., & beck, a. (2017). measuring student teachers development of metacognition and self-regulated learning in professional dialogue. ecer 2017. ifenthaler, d., gibson, d., & dobozy, e. (2018). informing learning design through analytics: applying network graph analysis. australasian journal of educational technology, 34(2). jeong, h., biswas, g., johnson, j., & howard, l. (2010). analysis of productive learning behaviors in a structured inquiry cycle using hidden markov models. paper presented at the educational data mining 2010. jeske, d., backhaus, j., & stamov roßnagel, c. (2014). self‐regulation during e‐learning: using behavioural evidence from navigation log files. journal of computer assisted learning, 30(3), 272-284. jovanović, j., gašević, d., dawson, s., pardo, a., & mirriahi, n. (2017). learning analytics to unveil learning strategies in a flipped classroom. the internet and higher education, 33, 74-85. doi: 10.1016/j.iheduc.2017.02.001 jovanovic, j., pardo, a., mirriahi, n., dawson, s., & gašević, d. (2017). an analytics-based framework to support teaching and learning in a flipped classroom. learning analytics in the classroom: translating learning analytics research for teachers. oxon: routledge . kallio, h., pietilä, a. m., johnson, m., & kangasniemi, m. (2016). systematic methodological review: developing a framework for a qualitative semi‐structured interview guide. journal of advanced nursing, 72 (12), 2954-2965. kelly, a. e., lesh, r. a., & baek, j. y. (2014). handbook of design research methods in education: innovations in science, technology, engineering, and mathematics learning and teaching : routledge. kiefer, p., giannopoulos, i., raubal, m., & duchowski, a. (2017). eye tracking for spatial research: cognition, computation, challenges. spatial cognition & computation, 17(1-2), 1-19. doi: 10.1080/13875868.2016.1254634 kinnebrew, j. s., loretz, k. m., & biswas, g. (2013). a contextualized, differential sequence mining method to derive students' learning behavior patterns. jedm| journal of educational data mining, 5(1), 190-219. kizilcec, r. f., pérez-sanagustín, m., & maldonado, j. j. (2017). self-regulated learning strategies predict learner behavior and goal attainment in massive open online courses. computers & education, 104, 18-33. doi: 10.1016/j.compedu.2016.10.001 knight, s., wise, a. f., & chen, b. (2017). time for change: why learning analytics needs temporal analysis. journal of learning analytics, 4(3), 7-17. köck, m., & paramythis, a. (2011). activity sequence modelling and dynamic clustering for personalized e-learning. user modeling and user-adapted interaction, 21(1-2), 51-97. doi: 10.1007/s11257-010-9087-z kuhn, t. s. (2012). the structure of scientific revolutions: university of chicago press. kumar, v., venkatesan, r., & reinartz, w. (2004). a purchase sequence analysis framework for targeting products, customers and time period. forthcoming in journal of marketing. kurki, k., järvenoja, h., järvelä, s., & mykkänen, a. (2017). young children’s use of emotion and behaviour regulation strategies in socio-emotionally challenging day-care situations. early childhood research quarterly, 41, 50-62. lan, m., & lu, j. (2017). assessing the effectiveness of self-regulated learning in moocs using macro-level behavioural sequence data. paper presented at the emoocs-wip. lazakidou, g., & retalis, s. (2010). using computer supported collaborative learning strategies for helping students acquire self-regulated problem-solving skills in mathematics. computers & education, 54(1), 3-13. lin, b., coburn, s. s., & eisenberg, n. (2016). self-regulation and reading achievement. the cognitive development of reading and reading comprehension, 67-86. liu, z., dev, h., dontcheva, m., & hoffman, m. (2016). mining, pruning and visualizing frequent patterns for temporal event sequence analysis. paper presented at the proceedings of the ieee vis 2016 workshop on temporal & sequential event analysis. lorigo, l., haridasan, m., brynjarsdóttir, h., xia, l., joachims, t., gay, g., . . . pan, b. (2008). eye tracking and online search: lessons learned and challenges ahead. journal of the association for information science and technology, 59 (7), 1041-1052. doi: 10.1002/asi.20794 lubahn, d. b., joseph, d. r., sar, m., tan, j.-a., higgs, h. n., larson, r. e., . . . wilson, e. m. (1988). the human androgen receptor: complementary deoxyribonucleic acid cloning, sequence analysis and gene expression in prostate. molecular endocrinology, 2(12), 1265-1275. doi: 10.1210/mend-2-12-1265 lupia, a., & alter, g. (2014). data access and research transparency in the quantitative tradition. ps: political science & politics, 47(1), 54-59. doi: 10.1017/s1049096513001728 lust, g. (2012). opening the black box. students' tool-use within a technology-enhanced learning environment: an ecological-valid approach. lust, g., vandewaetere, m., ceulemans, e., elen, j., & clarebout, g. (2011). tool-use in a blended undergraduate course: in search of user profiles. computers & education, 57(3), 2135-2144. doi: 10.1016/j.compedu.2011.05.010 maldonado-mahauad, j., pérez-sanagustín, m., kizilcec, r. f., morales, n., & munoz-gama, j. (2018). mining theory-based patterns from big data: identifying self-regulated learning strategies in massive open online courses. computers in human behavior, 80, 179-196. mannila, h., toivonen, h., & verkamo, a. i. (1997). discovery of frequent episodes in event sequences. data mining and knowledge discovery, 1(3), 259-289. doi: 10.1023/a:1009748302351 martin, s. e., shabanowitz, j., hunt, d. f., & marto, j. a. (2000). subfemtomole ms and ms/ms peptide sequence analysis using nano-hplc micro-esi fourier transform ion cyclotron resonance mass spectrometry. analytical chemistry, 72(18), 4266-4274. doi: 10.1021/ac000497v masseglia, f., teisseire, m., & poncelet, p. (2002). real time web usage mining with a distributed navigation analysis. paper presented at the research issues in data engineering: engineering e-commerce/e-business systems, 2002. ride-2ec 2002. proceedings. twelfth international workshop on. mazon, j. m., rossi, j. d., & toledo, j. (2014). an optimal matching problem for the euclidean distance. siam journal on mathematical analysis, 46(1), 233-255. doi: 10.1137/120901465 meredith, j. (1993). theory building through conceptual methods.international journal of operations & production management, 13(5), 3-11. mertens, d. m. (2014). research and evaluation in education and psychology: integrating diversity with quantitative, qualitative, and mixed methods : sage publications. mills, m., van de bunt, g. g., & de bruijn, j. (2006). comparative research: persistent problems and promising solutions. international sociology, 21(5), 619-631. doi: 10.1177/0268580906067833 molenaar, i., & chiu, m. m. (2015). effects of sequences of socially regulated learning on group performance. paper presented at the proceedings of the fifth international conference on learning analytics and knowledge. molenaar, i., & järvelä, s. (2014). sequential and temporal characteristics of self and socially regulated learning. metacognition and learning, 9(2), 75-85. doi: 10.1007/s11409-014-9114-2 moravcsik, a. (2014). transparency: the revolution in qualitative research. ps: political science & politics, 47(1), 48-53. doi: 10.1017/s1049096513001789 müller, h., naumann, f., & freytag, j.-c. (2003). data quality in genome databases. müller, n. s., studer, m., gabadinho, a., & ritschard, g. (2010). analyse de séquences d'événements avec traminer.paper presented at the egc. oliver, m., & trigwell, k. (2005). can ‘blended learning’be redeemed. e-learning, 2(1), 17–26. doi: 10.2304/elea.2005.2.1.17 paans, c., molenaar, i., segers, e., & verhoeven, l. (2018). temporal variation in children's self-regulated hypermedia learning. computers in human behavior. panadero, e., klug, j., & järvelä, s. (2016). third wave of measurement in the self-regulated learning field: when measurement and intervention come hand in hand. scandinavian journal of educational research, 60(6), 723-735. doi: 10.1080/00313831.2015.1066436 perer, a., & wang, f. (2014). frequence: interactive mining and visualization of temporal frequent event sequences. paper presented at the proceedings of the 19th international conference on intelligent user interfaces. perez, s., massey-allard, j., butler, d., ives, j., bonn, d., yee, n., & roll, i. (2017). identifying productive inquiry in virtual labs using sequence mining. paper presented at the international conference on artificial intelligence in education. peterson, r. a. (2005). problems in comparative research: the example of omnivorousness. poetics, 33(5), 257-282. doi: 10.1016/j.poetic.2005.10.002 pintrich, p. r. (2004). a conceptual framework for assessing motivation and self-regulated learning in college students. educational psychology review, 16(4), 385-407. doi: 10.1007/s10648-004-0006-x poole, m. s., lambert, n., murase, t., asencio, r., & mcdonald, j. (2016). sequential analysis of processes. the sage handbook of process organization studies, 254. pressley, m., levin, j. r., & mcdaniel, m. a. (1987). remembering versus inferring what a word means: mnemonic and contextual approaches. prinzie, a., & van den poel, d. (2007). predicting home-appliance acquisition sequences: markov/markov for discrimination and survival analysis for modeling sequential information in nptb models. decision support systems, 44(1), 28-45. doi: 10.1016/j.dss.2007.02.008 proctor, e. k., powell, b. j., & mcmillen, j. c. (2013). implementation strategies: recommendations for specifying and reporting. implementation science, 8(1), 139. doi: 10.1186/1748-5908-8-139 puustinen, m., & pulkkinen, l. (2001). models of self-regulated learning: a review. scandinavian journal of educational research, 45(3), 269-286. doi: 10.1080/00313830120074206 rahm, e., & do, h. h. (2000). data cleaning: problems and current approaches. ieee data eng. bull., 23(4), 3-13. reimann, p., markauskaite, l., & bannert, m. (2014). e‐research and learning theory: what do sequence and process mining methods contribute? british journal of educational technology, 45(3), 528-540. roll, i., & winne, p. h. (2015). understanding, evaluating, and supporting self-regulated learning using learning analytics. journal of learning analytics, 2(1), 7-12. roth, a., ogrin, s., & schmitz, b. (2016). assessing self-regulated learning in higher education: a systematic literature review of self-report instruments. educational assessment, evaluation and accountability, 28(3), 225-250. doi: 10.1007/s11092-015-9229-2 rotter, j. b. (1954). social learning and clinical psychology. schnaubert, l., heimbuch, s., & bodemer, d. (2016). extracting selection strategies: comapring measures to analyze sequential data. paper presented at the earli sig27 measuring learning online, oulu, finland. schraw, g., crippen, k. j., & hartley, k. (2006). promoting self-regulation in science education: metacognition as part of a broader perspective on learning. research in science education, 36(1-2), 111-139. doi: 10.1007/s11165-005-3917-8 schraw, g., & moshman, d. (1995). metacognitive theories. educational psychology review, 7(4), 351-371. doi: 10.1007/bf02212307 segedy, j. r., & biswas, g. (2015). towards using coherence analysis to scaffold students in open-ended learning environments. paper presented at the aied workshops. segedy, j. r., kinnebrew, j. s., & biswas, g. (2015). using coherence analysis to characterize self-regulated learning behaviours in open-ended learning environments.journal of learning analytics, 2(1), 13-48. siadaty, m., gašević, d., & hatala, m. (2016). associations between technological scaffolding and micro-level processes of self-regulated learning: a workplace study. computers in human behavior, 55, 1007-1019. doi: 10.1016/j.chb.2015.10.035 singleton, r., straits, b. c., & straits, m. (1999). approaches to social research oxford university press. new york and oxford. slater, s., joksimović, s., kovanovic, v., baker, r. s., & gasevic, d. (2017). tools for educational data mining: a review. journal of educational and behavioral statistics, 42(1), 85-106. doi: 10.3102/1076998616666808 sonnenberg, c., & bannert, m. (2015). discovering the effects of metacognitive prompts on the sequential structure of srl-processes using process mining techniques. journal of learning analytics, 2(1), 72-100. srikant, r., & agrawal, r. (1996). mining sequential patterns: generalizations and performance improvements. paper presented at the international conference on extending database technology. stackebrandt, e., & goebel, b. (1994). taxonomic note: a place for dna-dna reassociation and 16s rrna sequence analysis in the present species definition in bacteriology. international journal of systematic and evolutionary microbiology, 44 (4), 846-849. stark, d., & vedres, b. (2012). social sequence analysis. the emergence of organizations and markets, 347-364. studer, m., mueller, n. s., ritschard, g., & gabadinho, a. (2010). classer, discriminer et visualiser des séquences d'événements. paper presented at the egc. taub, m., azevedo, r., bouchet, f., & khosravifar, b. (2014). can the use of cognitive and metacognitive self-regulated learning strategies be predicted by learners’ levels of prior knowledge in hypermedia-learning environments? computers in human behavior, 39, 356-367. taub, m., azevedo, r., bradbury, a. e., millar, g. c., & lester, j. (2017). using sequence mining to reveal the efficiency in scientific reasoning during stem learning with a game-based learning environment. learning and instruction. taub, m., azevedo, r., bradbury, a. e., millar, g. c., & lester, j. (2018). using sequence mining to reveal the efficiency in scientific reasoning during stem learning with a game-based learning environment. learning and instruction, 54, 93-103. taub, m., mudrick, n. v., azevedo, r., millar, g. c., rowe, j., & lester, j. (2016). using multi-level modeling with eye-tracking data to predict metacognitive monitoring and self-regulated learning with c rystal i sland. paper presented at the international conference on intelligent tutoring systems. tsai, c.-w., shen, p.-d., & tsai, m.-c. (2011). developing an appropriate design of blended learning with web-enabled self-regulated learning to enhance students' learning and thoughts regarding online learning. behaviour & information technology, 30(2), 261-271. doi: 10.1080/0144929x.2010.514359 van der aalst, w. m. (2016). process mining: data science in action: springer. van der merwe, w. a. j. (2013). towards a conceptual model of the relationship between corporate trust and corporate reputation. university of pretoria. van krevelen, d. w., & te nijenhuis, k. (2009). properties of polymers: their correlation with chemical structure; their numerical estimation and prediction from additive group contributions : elsevier. van laer, s., & elen, j. (2016). adults’ self-regulatory behaviour profiles in blended learning environments and their implications for design. technology, knowledge and learning, 1-31. van laer, s., & elen, j. (2018). an instrumentalized framework for supporting learners’ self-regulation in blended learning environments. in m. j. spector, b. b. lockee, & m. d. childress (eds.), learning, design, and technology: an international compendium of theory, research, practice, and policy : springer, cham. van laer, s., jiang, l., & elen, j. (2018). the effect of cues for reflection on learners’ self-regulated learning through changes in learners’ learning behaviour and outcomes. computers & education, under review. veenman, m. v. (2007). the assessment and instruction of self-regulation in computer-based environments: a discussion. metacognition and learning, 2(2), 177-183. veenman, m. v., & alexander, p. (2011). learning to self-monitor and self-regulate. handbook of research on learning and instruction, 197-218. veenman, m. v., bavelaar, l., de wolf, l., & van haaren, m. g. (2014). the on-line assessment of metacognitive skills in a computerized learning environment. learning and individual differences, 29, 123-130. doi: 10.1016/j.lindif.2013.01.003 veenman, m. v., prins, f. j., & verheij, j. (2003). learning styles: self‐reports versus thinking‐aloud measures. british journal of educational psychology, 73(3), 357-372. veenman, m. v., van hout-wolters, b. h., & afflerbach, p. (2006). metacognition and learning: conceptual and methodological considerations. metacognition and learning, 1(1), 3-14. williamson, g. (2015). self-regulated learning: an overview of metacognition, motivation and behaviour. winne, p. (2016). self-regulated learning. sfu educational review, 1(1). winne, p. h. (2005). key issues in modeling and applying research on self‐regulated learning. applied psychology, 54(2), 232-238. winne, p. h. (2010). improving measurements of self-regulated learning. educational psychologist, 45(4), 267-276. winne, p. h. (2014). issues in researching self-regulated learning as patterns of events. metacognition and learning, 9(2), 229-237. doi: 10.1007/s11409-014-9113-3 winne, p. h. (2018). theorizing and researching levels of processing in self‐regulated learning. british journal of educational psychology, 88(1), 9-20. winne, p. h., & baker, r. s. (2013). the potentials of educational data mining for researching metacognition, motivation and self-regulated learning. jedm| journal of educational data mining, 5(1), 1-8. winne, p. h., & hadwin, a. f. (1998). studying as self-regulated learning. metacognition in educational theory and practice, 93, 27–30. winne, p. h., nesbit, j. c., & popowich, f. (2017). nstudy: a system for researching information problem solving. technology, knowledge and learning, 22(3), 369-376. doi: 10.1007/s10758-017-9327-y winne, p. h., & perry, n. e. (2000). measuring self-regulated learning. yang, x., li, j., & xing, b. (2018). behavioral patterns of knowledge construction in online cooperative translation activities. the internet and higher education, 36, 13-21. doi: 10.1016/j.iheduc.2017.08.003 zaki, m. j. (2001). spade: an efficient algorithm for mining frequent sequences. machine learning, 42(1-2), 31-60. doi: 10.1023/a:1007652502315 zhou, m. (2016). data pre-processing of student e-learning logs information science and applications (icisa) 2016(pp. 1007-1012): springer. zhou, m., xu, y., nesbit, j. c., & winne, p. h. (2010). sequential pattern analysis of learning logs: methodology and applications. handbook of educational data mining, 107, 107-121. zimmerman, b. j. (1986). becoming a self-regulated learner: which are the key subprocesses? contemporary educational psychology, 11, 307–313. zimmerman, b. j. (1990). self-regulated learning and academic achievement: an overview. educational psychologist, 25(1), 3–17. doi: 10.1207/s15326985ep2501_2 zimmerman, b. j. (1998). academic studing and the development of personal skill: a self-regulatory perspective. educational psychologist, 33 , 73–86. zimmerman, b. j. (2000). self-efficacy: an essential motive to learn. contemporary educational psychololgy, 25(1), 82-91. doi: 10.1006/ceps.1999.1016 zimmerman, b. j., & pons, m. m. (1986). development of a structured interview for assessing student use of self-regulated learning strategies. american educational research journal, 23(4), 614-628. doi: 10.3102/00028312023004614 zimmerman, b. j., & schunk, d. h. (2001). self-regulated learning and academic achievement: theoretical perspectives : routledge. eckstein frontline learning research vol.7 no. 2 (2019) 1 22 issn 2295-3159 production and perception of classroom disturbances – a new approach to investigating the perspectives of teachers and students boris ecksteina auniversity of teacher education st. gallen, switzerland article received 9 september 2018/ revised 2 november / accepted 5 march/ available online 25 march abstract classroom disturbances impair the quality of teaching and learning, and they can be a source of strain for both teachers and students. some studies indicate, however, that not everyone involved gets equally disturbed by the same occurrences. altogether, there is still little solid knowledge about the teachers’ and the students’ subjective perception of disturbance. moreover, rater effects may have confounded the findings available. addressing these desiderata, the sugus study investigates two elements of classroom disturbances within an interactionist framework: the incidence of deviant behaviour shown by particular target students, and the intensity of disturbance as subjectively perceived by teachers, by classmates, and by the targets themselves. for this purpose, we conducted a questionnaire survey among 85 primary-school class teachers and 1412 students. the data were analysed by means of a two-level correlated trait – correlated method minus one [ct-c(m-1)] model. this relatively novel statistical procedure has only rarely been applied in educational research so far. it made it possible to determine the respondents’ common view on classroom disturbances as well as the rater-specific perspectives. the results indicate that increasing deviance coincides with increasing distraction and annoyance – but mainly in a relatively small intersection of the different perspectives. beyond that, the analysis revealed substantial rater effects which explain 30 to 61% of variance in teacher ratings, for instance. the author discusses likely reasons why disturbances are perceived so divergently. keywords: classroom disturbances; deviant behaviour; subjective perception of disturbance; rater effects; ct-c(m-1) modelling info corresponding author email: boris.eckstein@phsg.ch doi: 10.14786/flr.v7i2.411 1 two main elements of classroom disturbances classroom disturbances arise from inappropriate student or teacher behaviour (montuoro & lewis, 2015), albeit not everyone in class gets equally distracted or annoyed (eckstein, grob, & reusser, 2016). this implies that classroom disturbances consist of an objective core that the persons involved may perceive differently (eckstein, 2018). this argument is theoretically well-founded, but there is only little empirical evidence regarding commonalities and differences between the teachers’ and the students’ distinct perceptions. this paper presents a new methodological approach to investigating three perspectives on the objective core of classroom disturbances: the self-perception of students, the perception of teachers, and the perception of classmates. building on two main lines of theory and research, the focus lies on deviant student behaviour and on the persons’ subjective perception of disturbance. 1.1 deviant student behaviour a first line of theory and research on classroom disturbances has focused on student behaviour that deviates from specific classroom rules (e.g. chattering) or from common socio-moral conventions (e.g. insolence). studies on incidence rates from several countries largely agree that the most frequent forms of deviant student behaviour are relatively minor discipline problems whereas aggressive and dissocial behaviours are considerably rarer (beaman, wheldall, & kemp, 2007; crawshaw, 2015). many studies investigated ontogenetically determined risk factors of deviant student behaviour, e.g. impulsivity (carroll, houghton, taylor, west, & list-kerz, 2006). other studies examined proximal causes and preventions like teaching styles (godwin et al., 2016; sherman, rasmussen, & baydala, 2008), or classroom management (emmer & sabornie, 2015). this body of research has contributed important knowledge of one aspect of classroom disturbances. however, not all studies considered that the same behaviours can be differently perceived, interpreted, and judged from distinct perspectives (crawshaw, 2015). 1.2 teachers’ and students’ subjective perception of disturbance a second line of theory and research has revealed differential perceptions of classroom disturbances: most teachers feel stressed when their students frequently show deviant behaviours (e. little, 2005), which may even result in burnout (kokkinos, 2007). many students perceive deviant behaviour of their classmates as disturbing too (infantino & little, 2005), but they usually deem it less troubling than teachers (montuoro & lewis, 2015). while most students admit that they get distracted by their classmates’ deviance, not all of them claim to feel annoyed (schönbächler et al., 2009). students who behave deviantly themselves often realise that this may disturb the others, but they worry much more about their own problems (preuss-lausitz, 2005). altogether, this means that teachers and students perceive classroom disturbances differently according to distinct frames of perception (wettstein, ramseier, scherzinger, & gasser, 2016). role-specific traits influence the teachers’ and the students’ perception: teachers are accountable for all classroom interactions whereas the students are mainly concerned with their own learning, motivation and emotional wellbeing. this implies distinct valences and normative expectations that affect the way in which teachers and students perceive disturbances (wettstein, scherzinger, & ramseier, 2018). moreover, individual traits of students (wettstein, ramseier, & scherzinger, 2018) and teachers (hamre, pianta, downer, & mashburn, 2008) influence their perception: teachers with decreasing self-efficacy beliefs, for instance, feel increasingly strained by deviant student behaviour (arbuckle & little, 2004; dicke et al., 2014). finally, contextual traits affect the perception: in instructional settings with rigid rules (zevenbergen, 2001) or in classes with a low collective level of disturbances (makarova, herzog, & schönbächler, 2014), occasional occurrences of deviance are perceived as strongly disturbing. 2 holistic conceptualisation of classroom disturbances in the sugus study the sugus study integrates the above introduced two lines of theory and research. we conceptualise classroom disturbances as a co-constructed, interactionist phenomenon (eckstein, luger, grob, & reusser, 2016). figure 1 illustrates this holistic understanding in a theoretical model (eckstein, grob, et al., 2016; after wettstein, 2012). the arrows symbolise causal effects. the key argument is depicted in the upper part of the figure: classroom disturbances are constituted by two main elements – deviant student behaviour, and subjective perception of disturbance. we assume that deviance commonly disturbs teachers and students, yet the context as well as role-specific and individual traits of the “disturbed” affect the intensity in which they get distracted and/or annoyed. the present article focuses on this presumed core mechanism. beyond that, the model outlines how the assumed interaction could continue: teachers or students might react to behaviours that they have perceived as disturbing, e.g. with rebukes. the “disturber”, in turn, deems such behavioural reactions more or less fair or humiliating – depending on the context and his/her personal traits. the model’s circular structure visualises that student behaviour and the way in which it is perceived, interpreted, and judged are contextualised in a history of preceding interactions (doyle, 2006). figure 1. interactionist model of the production and perception of classroom disturbances. 3. rater effects as a methodological consequence of differential perceptions rater effects stem from characteristics of the raters, the rating instrument, and the rating environment (wolfe, 2004). in research on classroom disturbances, the above mentioned distinct frames of perception almost inevitably entail such rater effects: teachers, students, or external observers perceive and thus assess disturbances according to role-specific, individual and/or contextual conditions. not all studies that focused on deviant student behaviour accounted for these issues sufficiently (crawshaw, 2015). as a consequence, some findings available cannot be considered as purely objective information. extreme cases are rater biases, caused, for example, by halo effects (hoyt, 2000), prejudiced, selective attention (hofer, 1986), expectations triggered by labels like “adhd” (ohan, visser, strain, & allen, 2011), or self-serving strategies (fishbein & ajzen, 2010). further rater effects may be caused by differential opportunities in terms of the perceptibility of certain disturbances: if students behave deviantly during group work outside the classroom, for instance, only the present classmates will notice but not the teacher. moreover, high-inference rating instruments that require subjective interpretations from the raters amplify rater effects (hoyt & kerns, 1999; südkamp, kaiser, & möller, 2012). examples from research on classroom disturbances are vague formulations like “troublesome behaviours”, likert-type frequency scales without clearly defined options (e.g. “seldom – sometimes – often”), or ratings of the whole class instead of individual students. 4. aims 4.1 general objectives of the research design we designed the sugus study in order to investigate classroom disturbances according to our holistic conceptualisation; and we aimed to control for rater effects (eckstein, grob, et al., 2016). for these purposes, we conducted a multi-perspective survey in which teachers and students reported on individual students of their class (target students): the teachers described all students in their class, the students described themselves plus four randomly assigned classmates (eckstein, luger, grob, & reusser, 2018). a first objective of the sugus study was to measure the incidence of deviant student behaviour as unbiased as possible: the respondents assessed the targets’ behaviours on a low-inference rating scale that we had developed in order to reduce the amount of potential rater effects. a second objective was to measure the respondents’ subjective perception of disturbance: the respondents assessed the intensity of distraction and annoyance the targets had caused according to their perspectives (eckstein, grob, et al., 2016). a third objective has been to analyse the relationship between deviance and perception of disturbance in each perspective and in the raters’ common view respectively. as the sugus study simultaneously investigates the behaviour of individual target students as well as three different perspectives on these targets, it breaks new ground in researching the (objective) production and the (subjective) perception of classroom disturbances. 4.2 analysis strategy and research questions the survey yielded a complex multitrait-multimethod data set which has been analysed by means of the “correlated trait – correlated method minus one” [ct-c(m-1)] approach (eid, lischetzke, nussbeck, & trierweiler, 2003). this is a special variant of structural equation modelling which analyses the influence of latent traits on manifest indicators as well as the impact of the methods applied, e.g. different types of rating. a ct-c(m-1) model estimates trait factors and method factors. a trait factor comprises the amount of the measurements’ variance that is consistent across all methods applied. a method factor, by contrast, comprises the amount of variance which is specific to this method. one method is selected as comparison standard (reference method) for which no method factor is modelled. as a consequence, there is one method factor less than methods applied, hence “m-1” (eid et al., 2008). psychologists developed this modelling technique originally and applied it, for instance, in personality research with multirater designs in order to analyse the consistency and the specificity of different perspectives on the mood of target persons (carretero-dios, eid, & ruch, 2011). educational researchers applied this approach only recently in a few studies to analyse commonalities and differences between teacher and student ratings, e.g. regarding inclusion (venetz, zurbriggen, & schwab, 2017). the ct-c(m-1) model in this paper analyses teacher, peer and self-ratings of classroom disturbances. the model’s trait factors are the intersection of the three perspectives and represent, thus, the raters’ common view. in order to quantify the extent of this common view, consistency coefficients are calculated. as opposed to this, the model’s method factors comprise method-specific divergences from the common view. these divergences can be interpreted as rater effects which are quantified by specificity coefficients. furthermore, the model analyses correlations. the correlations between the trait factors estimate the strength of the constructs’ relation in the common view. they indicate how strongly deviant behaviour correlates with the perceived intensity of disturbance – in the shared perspective of all raters. the correlations between the method factors, by contrast, estimate the similarity of rater-specificities across different constructs. they represent the generalisability of rater effects. the analyses provide answers to the following research questions: (q-1) consistency of deviance ratings: to what extent are teacher, peer and self-ratings of deviant student behaviour consistent, and to what extent are they rater-specific? because the ratings rest on a low-inference instrument, a larger extent of consistency than specificity is expected. (q-2) specificity of the subjective perception of disturbance: to what extent are teacher, peer and self-ratings of the intensity of disturbance consistent, and to what extent are they rater-specific? because perception primarily concerns the raters, a larger extent of specificity than consistency is expected. (q-3) relation of deviance and perception of disturbance in the raters’ common view: how strongly do the trait factors of deviance and perception of disturbance correlate? medium to strong correlations are expected: increasing deviance coincides with increasing distraction and annoyance. (q-4) generalisability of rater-effects: how strongly do the method factors correlate? medium correlations are expected: a rater-specific divergence in one construct coincides with analogous divergences in the other constructs. 5. method 5.1 sample and conduct of the survey in summer 2016, we conducted a written survey in 85 primary school classes (90.2% fifth grade, 9.8% mixed grades; students’ mean age: 11.73 years, sd: 0.52 years). all 85 class teachers and 1412 students out of a total of 1687 participated in the study. 275 students did not participate because they did not want to or because their parents did not allow it. the survey took place twice with one week in between and lasted a whole lesson each time. this research design had been pretested in a pilot study in 2014 (eckstein, reusser, grob, & hofstetter, 2015). the surveys’ main focus was on individual target students: the teachers reported on all students in their class (teacher ratings), the students described themselves (self-ratings) plus four randomly assigned classmates (peer ratings). the participating students acted both as raters and targets, the non-participating students were rated only by their teachers. on both occasions, the same raters and targets were paired together with the aid of personalised questionnaires: the raters’ and the targets’ names were printed on a tear-off strip. several measures were taken to guarantee the respondents anonymity. supervising members of the project team reported that the students had been in a good mood after the survey (eckstein et al., 2018). 5.2 instruments the questionnaire consisted of a general part (e.g. regarding teaching styles) and of a specific part that focused on the target students. this paper addresses the target-specific instruments exclusively; they were identical for teacher, peer and self-ratings except for the item wording (see appendix a). 5.2.1 deviant student behaviour in the first part of the survey, we asked the respondents how frequently the target students had behaved deviantly in the preceding two weeks. the instrument comprised 18 items. the answering format consisted of six categories (“never” to “5 times”) plus an option for free answers (“more frequently, namely: …”). the factorial structure covered two dimensions (eckstein et al., 2018): undisciplined behaviour: peer ratings: <α> = .79, teacher ratings: α = .85, self-ratings: α = 64; 8 items, e.g. “[name of the target] talked to another child during the lesson although the students were supposed to be quiet.” dissocial behaviour: peer ratings: <α> = .87, teacher ratings: α = .86, self-ratings: α = .81; 10 items, e.g. “[name of the target] insulted another child in class.” 5.2.2 subjective perception of disturbance one week later, the respondents assessed the intensity of disturbance the targets had recently caused according to their subjective perception. they rated nine statements on a four-point rating scale (“strongly disagree” to “strongly agree”). eight of these nine items were included in the factor analyses, which led to two dimensions (eckstein et al., 2018): affective perception of disturbance: peer ratings: <α> = .87, teacher ratings: <α> = .79, self-ratings: <α> = .70; 4 items, e.g. “[name of the target] … annoyed me.” cognitive perception of disturbance: peer ratings: α = .92, teacher ratings: α = .93, self-ratings: α = .79; 4 items, e.g. “[name of the target] … distracted me from the lesson.” 5.3 the ct-c(m-1) modelling technique following eid et al. (2008) and carretero-dios et al. (2011), a two-level ct-c(m-1) model was set up and calculated in mplus 8.0 (muthén & muthén, 2017). the model comprises four target-specific traits (undisciplined and dissocial behaviour, affective and cognitive perception of disturbance) which have been measured by three different methods (teacher, peer and self-ratings). the most important principles of the modelling are explained in this section. further explanations follow in the results section and in appendix b. level 2 (l2) is the level of the target students (nl2 = 1677) where the teacher ratings, the self-ratings, and the aggregated peer ratings have been modelled. level 1 (l1), by contrast, is the level of unique peer ratings (nl1 = 5811). each target student received 3.47 peer ratings on average. these unique peer ratings are nested within targets – they have been aggregated at l2 into error-free random intercepts (true scores). figure 2 provides a sketch of the model: ovals symbolise latent factors, boxes represent manifest indicators (item parcels) with tiny grey arrows illustrating the measurement errors. the symbols are labelled with acronyms; in each case, the first letter fits the constructs’ denotation (e.g. “u”: undisciplined behaviour). blue refers to teacher ratings, labelled with the letter “t”. yellow refers to self-ratings, labelled with “s”, and red refers to peer ratings, labelled with “p”. the trait factors are white as they represent the shared perspective of all raters. the labels’ numbers refer to the item parcels which are numbered consecutively (see appendix a). unidirectional arrows illustrate factor loadings (λ), double arrows represent covariances (ψ). figure 2. sketch of the two-level ct-c(m-1) model. i will explain the modelling in detail by an exemplary look at the construct “undisciplined behaviour”. the white oval “u_trait” at l2 symbolises the latent trait factor of the targets’ undisciplined behaviour. u_trait unites the proportion of variance that is shared across all types of rating and thus represents the raters’ common view on the targets’ behaviour. the blue oval “u_mt” stands for the method factor of the teacher ratings; the yellow oval “u_ms” represents the method factor of the self-ratings. these method factors comprise the proportion of variance that is specific to those two types of rating (not shared with other types of rating). the blue boxes “ut1–ut3” symbolise the indicators of the teacher ratings, the yellow boxes “us1–us3” represent the indicators of the self-ratings. these indicators are influenced by the trait factor and by each according method factor. the influences are estimated by the factor loadings “λ”. the small red ovals “up1–up3” at l2 illustrate the aggregated peer ratings. unlike the teacher and self-ratings, these are not indicators at l2 but error-free random intercepts. no method factor was modelled for these aggregated peer ratings because they have been selected as the reference method with which the other types of rating are compared (comparison standard). it implies further that the peer ratings’ l2-scores are determined by the trait exclusively (no rater effect influences the aggregated peer ratings). the red boxes “up1–up3” at l1 represent the manifest indicators of the unique peer ratings. the black dots on the edge of these boxes indicate that they have been aggregated into random intercepts at l2. each target student was rated by 3.47 peers on average. these unique peer ratings per target may differ from one another. the red oval “u_pl1” at l1 illustrates a unique method factor of the unique peer ratings. this factor comprises the specific proportion of variance that is due to interindividual differences among the multiple peer raters per target (neither shared among peers, nor shared with other types of rating). it estimates the rater effect which influences the unique peer ratings. altogether, u_trait corresponds to the true scores per target, measured by the error-free aggregated peer ratings. in addition, u_trait encompasses the proportion of the teacher ratings and self-ratings that is consistent with the aggregated peer ratings. that is to say, u_trait is the intersection of the different perspectives. this intersubjectively consistent measurement can be considered as an approximation to an objective information about the targets undisciplined behaviour. the method factors comprise the disagreement among the raters. the other three constructs have been modelled in the same way. the white ovals illustrate the trait factors of dissocial behaviour (d_trait), affective perception of disturbance (a_trait), and cognitive perception of disturbance (c_trait). the blue ovals represent the teachers’ method factors (_mt); the yellow ovals symbolise the self-ratings’ method factors (_ms); the red ovals at l1 represent the peer ratings’ unique method factors (_pl1). the red boxes at l1 stand for the unique peer ratings (dp1–dp3, ap1–ap2, cp1–cp2); the small red ovals at l2 illustrate the aggregated scores (labelled identically as the unique ratings). the blue boxes represent the teacher ratings (dt1–dt2, at1–at2, ct1–ct2); the yellow boxes represent the self-ratings (ds1–ds2, as1–as2, cs1–cs2). further measures have been adopted: the correlations between trait and method factors of the same trait-method unit are explicitly set to zero, because these factors are uncorrelated by definition (eid et al., 2008). the factor loadings of the peer ratings are set equal at both levels to avoid cluster biases (jak, oort, & dolan, 2013) which does not worsen the model fit (chen, 2007). the 30 manifest indicators are item parcels. parcelling was necessary for reducing the complexity of the model. this is justified because the analysis primarily aimed to estimate the relations between the constructs (t. d. little, cunningham, shahar, & widaman, 2002). because the data are not normally distributed (eckstein, grob, & reusser, 2017), the distinctive mplus estimator with robust standard errors “mlr” (finney & distefano, 2013) was applied. the mlr estimator uses all information available to model missing data (muthén & muthén, 2017). 6. results the two-level ct-c(m-1) model fits the data well (𝜒2(mlr) = 1360.18, df = 372, p < .001; rmsea = .021; cfi = 0.93; srmrl1 = 0.04; srmrl2 = 0.04). figure 3 shows the estimated standardised factor loadings and correlations. because all included parameters are significant (p < .05), the usual marking with asterisks (*) has been omitted for the sake of clarity. furthermore, only correlations with |r| > 0.20 are displayed (a complete list of all correlations is provided in table 2). all factor loadings are significant and for the most part within an acceptable range. the standardised trait-factor loadings of the aggregated peer ratings amount to 1.00 because they are perfectly explained by the trait factor (measurement errors and rater effects are completely at l1). compared to this, the freely estimated loadings of the teacher indicators on the trait factors are lower (.42 ≤ λ ≤ .70). the loadings of the self-ratings are even lower (.18 ≤ λ ≤ .40). this is a first indication that the aggregated peer ratings (reference method) converge more strongly with the teacher ratings than with the self-ratings. table 1 presents further results. as for the means (m), the target students on average only rarely showed undisciplined behaviour (.68 ≤ m ≤ 1.68) and even more rarely dissocial behaviour (.22 ≤ m ≤ .44). that the means are very low becomes evident if one keeps in mind that the rating scale was not limited to a fixed maximal number of incidents but included an option for free answers. given these low mean incidences, it is not surprising that the target students on average were considered to be only little annoying (.27 ≤ m ≤ .70) and distracting (.39 ≤ m ≤ .53). the theoretical means of these two scales amount to 1.5. the unstandardized trait factor loadings indicate level differences between the types of rating (geiser, eid, west, lischetzke, & nussbeck, 2012): loadings greater than 1.00 indicate that the pertaining self-ratings or teacher ratings are higher than the average peer rating (comparison standard); loadings lesser than 1.00 indicate lower selfor teacher ratings compared to the peer ratings. in sum, the teachers on average reported more incidents of undisciplined behaviours but fewer incidents of dissocial behaviours than the peers. furthermore, the teachers on average reported a greater intensity of distraction but a lesser intensity of annoyance compared to the peers. the self-ratings on average are lower than the peer ratings as regards all four constructs. figure 3. standardised parameters according to the two-level ct-c(m-1) model estimation. using the formulas proposed by eid et al. (2008), indicator-specific variance components were manually calculated. consistency (c) is a measure of interrater agreement that quantifies the intersection between the raters’ distinct perspectives. it expresses the extent to which the aggregated peer ratings (reference method) explain the variance in the teacher ratings, in the self-ratings, and in the unique peer ratings. specificity (s), by contrast, measures the rater effects: the coefficient is the proportion of variance in the ratings that is not consistent with the aggregated peer ratings. technically speaking, it quantifies the overestimation or underestimation of other ratings compared to the average peer ratings. the reliability coefficient ( ω ) is the ratio of explained variance to total variance. this implies that unreliability (unr = 1ω ) can be accounted for by the measurement error. consistency, specificity, and unreliability add up to 1.00 per indicator, and their proportions are to be interpreted in percentages when multiplied by 100. the components’ meaning can be illustrated with an example from table 1: the variance in the teacher ratings in ut1 is explained by the aggregated peer ratings to an extent of 45% (consistency). 38% of variance are method-specific, that is to say due to rater effects. the remaining 17% of variance can be accounted for by measurement errors (unreliability). some of the indicator-specific reliability coefficients are rather low. but as regards the whole constructs, reliability is considerably higher. figures 4 to 7 illustrate the variance components for all constructs per method. the bottom-most, grey part of the pillars represents unreliability of the constructs (.06 ≤ unr ≤ .29). the green part above stands for consistency (.07 ≤ c ≤ .54). the uppermost, red part displays specificity (.30 ≤ s ≤ .70). table 1 factor variance, indicator-specific means, unstandardized factor loadings, and variance components note. “factor variance” indicated for teacher ratings, self-ratings and unique peer ratings relates to the (unique) method factors. as for the aggregated peer ratings (reference method), “factor variance” relates to the trait factor. the factor loadings are non-standardised values. the loadings of the unique peer ratings on the unique method factors are equivalent to the loadings of the aggregated peer ratings on the trait factors (no cluster bias). reliability of the error-free aggregated peer ratings amounts to 1.00 because they are completely explained by the trait factor (errors and method effects are completely at l1). the self-ratings are marked by low consistency but high specificity over all constructs. this becomes most obvious in the construct “dissocial behaviour”: only 7% of variance in the self-ratings can be explained by the aggregated peer ratings (reference method) whereas 70% of variance are rater-specific and can thus be explained by rater effects. the variance in teacher ratings of the construct “undisciplined behaviour” can be explained to an extent of 54% by the aggregated peer ratings (consistency). with respect to the construct “dissocial behaviour”, by contrast, consistency only amounts to 26%. this discrepancy is unexpected because both constructs were measured with analogous low-inference scales that were assumed to lead to similar results. moreover, we expected larger consistency coefficients for all deviance ratings. also contrary to our expectations, the teacher ratings of the construct “cognitive perception of disturbance” are to an extent of 47% consistent with the aggregated peer ratings. since this is a measure of subjective perceptions, we expected a lower consistency coefficient. the unique peer ratings are to an extent of 18 to 31% consistent with the aggregated scores. in itself, the values in this range are rather low. what is even more surprising, however, is that all consistency coefficients of the unique peer ratings are lower than those of the teacher ratings. this means that the teacher ratings converge more strongly with the peers’ aggregated scores than the unique peer ratings – out of which the aggregated scores had originally been calculated. table 2 shows a complete list of the correlations between the latent factors. significant results (p < .05) are marked by an asterisk (*). the table is made up of sectors (i to vii), which are described in the following. table 2 interfactor correlations sector i shows the correlations of the unique method factors. all coefficients are positive and significant (.33* ≤ r ≤ .68*). thus, the rater-specificity of the unique peer ratings is largely generalizable: if a peer rater overestimates a target in one construct (compared to the aggregated score), this peer tends to overestimate the same target in other constructs too. or vice versa: underestimation in one construct coincides with underestimation in the other constructs. the correlations between the method factors of the teacher ratings (sector iv) and between the method factors of the self-ratings (sector vii) can be interpreted analogously. sector ii includes the correlations between the trait factors which estimate the strength of the constructs’ relation in the raters’ common view (measured by the relatively small intersection of the different types of rating). all these correlations are positive and significant: the more frequently targets show undisciplined behaviour, the more frequently they behave in dissocial ways too (r = .75*). besides, the more frequently the targets show undisciplined behaviour, the more annoyed (r = .77*) and distracted (r = .91*) the respondents get. the same holds true with respect to dissocial behaviour (.75* ≤ r ≤ .82*). although such findings were expected, the actual correlations are so high that they raise questions concerning discriminant validity. sector iii displays the correlations between the trait factors and the method factors of the teacher ratings. the highest positive correlation indicates a tendency that targets who show increasingly frequent indiscipline according to the common view (u_trait) are perceived as increasingly (r = .24*) annoying in their teacher’s specific perspective (a_mt). the correlations between the trait factors and the method factors of the self-ratings in sector v can be interpreted in analogous ways, but they are rather low or not significant. sector vi lists the correlations between the method factors of the teacher ratings and the self-ratings. none of the coefficients is significant. this means that the teacher ratings and the self-ratings do not share any variance that they do not share with the aggregated peer ratings. 7. discussion 7.1 summary and conclusion according to our theoretical framework, classroom disturbances consists of an objective core that not all persons involved perceive as equally disturbing (eckstein, grob, et al., 2016). we assume that deviant student behaviour commonly distracts and annoys teachers and students, yet the intensity of a perceived disturbance is affected by role-specific, individual, and contextual conditions. the sugus study investigates this interactionist phenomenon by means of a multi-perspective survey. based on teacher, peer and self-ratings we measured the incidence of deviant student behaviour and the respondents’ subjective perception of disturbance. the aim of this paper has been to analyse commonalities and differences between the three perspectives with a two-level ct-c(m-1) model. four research questions have been pursued and can be answered as follows: (q-1) consistency of the deviance ratings. because of our low-inference instrument, we assumed to obtain rather unbiased ratings of deviant student behaviours. therefore, we supposed that the teacher, peer and self-ratings would be largely consistent. in other words, we expected only minor rater effects. this hypothesis needs to be rejected. the extent to which the ratings are consistent merely amounts to 7 to 54%. as against this, rater effects explain between 30 and 70% of the measurements’ variance. only the teacher ratings of the construct “undisciplined behaviour” are more consistent (54%) than specific (30%). in all other cases, the rater-specific divergences make up the larger proportion. these unexpectedly large rater effects may in part be accounted for by the respondents’ role-specific frames of perception (scherzinger, wettstein, & wyler, 2017). as far as the self-ratings are concerned, the students may have underestimated the frequency of their own dissocial behaviour in terms of a self-serving strategy (70% specificity). this interpretation is supported by the low means of the self-ratings and by the unstandardized trait factor loadings of the self-ratings which are lower than those of the peer ratings. the rater effects that influence the teacher ratings of dissocial behaviour (61% specificity) may be explained by differential opportunities in terms of its perceptibility: students probably hide dissocial behaviours due to fear of sanctions so that their teacher does not perceive every incident. this would explain why the teachers on average reported fewer incidents of dissocial behaviour compared to the average peer ratings according to the unstandardized trait factor loadings. undisciplined behaviour, by contrast, is socially less disapproved and therefore less hidden and better perceptible for the teachers (30% specificity). according to the unstandardized trait factor loadings, the teachers on average reported more incidents of indiscipline than the peers. this may be due to the teachers’ sensitivity to disciplinary issues in class. maybe this implies further that the peers underestimated the indiscipline of their classmates. (q-2) specificity of the subjective perception of disturbance. the assumption was that specificity is larger than consistency. this hypothesis is supported by the results: up to 70% of variance can be accounted for by rater effects. solely regarding the teachers’ cognitive perception of disturbance, the ratio is less pronounced than expected (47% consistency vs. 48% specificity). this implies that teachers and peer raters on average agreed to 47% on the intensity of distraction the targets had caused. this finding is surprising because it rests on highly subjective measurements so that the proportion of specificity was expected to be larger. in all other cases, specificity clearly outweighs consistency as expected. this becomes particularly evident in the self-ratings (63 to 65% specificity). the low means and the unstandardized trait factor loadings indicate that the students probably underestimated themselves in this respect compared to the average peer ratings, which might be once more a consequence of self-serving biases. the rater effects that influence the teachers’ affective perception of disturbance (53% specificity) can be outlined by the level difference revealed by the unstandardized trait factor loadings: the teachers on average described the targets as less annoying than the peers. the low means imply that the teachers described most students as not annoying. maybe these ratings are a consequence of the teachers’ professional ethos. (q-3) relation of deviance and perception of disturbance in the raters’ common view. we expected medium to strong correlations between the four trait factors. this hypothesis is supported by the results: the more frequently the targets show deviant behaviour (according to the common view), the more annoyed and distracted the respondents get. some of these correlations are so high, however, that they raise the question of whether the constructs discriminate sufficiently. it has to be considered, for instance, whether the trait factors “undisciplined behaviour” and “cognitive perception of disturbance” should be merged into one super factor. this merging would mean to equate the incidence of indiscipline and the intensity of distraction as two undistinguishable aspects of the same phenomenon. we refrained from doing so because the merging of factors should result in theoretically explicable constructs (kleinke, schlüter, & christ, 2017). in our view, the potential super factor would represent an obscure mixture of two distinct constructs. furthermore, the constructs discriminate well in the raters’ specific perspectives (the method factors correlate much less strongly). (q-4) generalisability of rater-effects. the assumption was that the rater effects tend to be similar across the different constructs. the results support this hypothesis. the analyses showed medium to high correlations between the rater-specific method factors (.19* ≤ r ≤ .76*). raters who overestimated a target in one construct tended to overestimate this target regarding the other constructs as well (compared to the aggregated peer ratings). the same holds true vice versa: underestimation in one construct coincides with underestimation in the other constructs. this indicates that the rater effects are largely generalisable. 7.2 limitations the aggregated peer ratings have been selected as reference method because multiple ratings per target were available so that occasional exaggerations and trivialisations were assumed to even out. therefore, the peer ratings were supposed to be more precise than the other types of ratings. however, more information would be needed to verify this assumption, e.g. ratings from external observers (wettstein, scherzinger, et al., 2018). as an alternative, the generalisability of ratings might be increased in longitudinal study designs, e.g. by means of the experience sampling method (zurbriggen, venetz, & hinni, 2018). 7.3 implications the analyses have revealed that rater effects clearly dominated the teachers’ and the students’ reports of classroom disturbances. this finding supports the key argument of our theoretical model (eckstein, grob, et al., 2016): classroom disturbances consist of an objective core (deviant student behaviour) that the persons involved may perceive, interpret and judge differently (subjective perception of disturbance). the detected differences between the teachers’ and the students’ perspectives, for example, indicate role-specific frames of perception which most likely can be explained by their distinct tasks, aims and normative expectations. the teachers, for instance, need to notice and to moderate even minor forms of indiscipline in order to prevent more serious disturbances because of their pedagogical mandate. this is probably one reason why they reported more incidents of undisciplined behaviours than the students on average. furthermore, the differences between the multiple peer ratings per target student indicate that individual traits affect the raters’ perception, interpretation and judgement of disturbances. these interindividual differences may be explained partly by the raters’ general sensitivity to disturbance, as some students (and teachers) get more easily distracted or annoyed by the same occurrences than others (eckstein, 2018). another part of the ratings’ specificity may be explained by the quality of the raters’ relationship to the targets which evolved from their preceding interactions (doyle, 2006). it can be assumed, for example, that students observe close friends in class differently than classmates who they dislike. in the case of the teachers, hofer (1986) revealed that they monitor their students according to preconceived categories like “the disturber” or “the top pupil” which can lead to biased perceptions and thus to inadequate reactions. referring to the labelling approach (becker, 1963), such mechanisms can be interpreted as a reciprocal dependency of the production and the perception of classroom disturbances. the presumed long-term process of this interactionist phenomenon is visualised by our model’s circular structure. beyond that, the large extent of the rater effects might be associated with a general epistemological problem: considering the subjectivity of human perception, it seems questionable whether classroom disturbances can be investigated in the sense of objective matters of fact. a methodological attempt to deal with this problem in future studies could consist in complementing surveys with video studies (janík & seidel, 2009) using direct behaviour ratings (christ, riley-tillman, & chafouleas, 2009). this much information might enable researchers at least to further approximate the objective core of classroom disturbances. researchers without the resources to apply such mixed method designs, by contrast, need to select the most appropriate source of information with respect to the aims of their study (kunter & baumert, 2006). in any case, the application of low-inference instruments seems recommendable. the consistency coefficients of our deviance scale amount up to 54% which is larger than usual interrater agreement in classroom research (wagner et al., 2016). as a final point, the findings have practical implications: the large extent of the rater effects implies that the labelling of single “problem students” (hunt et al., 1989) is questionable from an ethical point of view. therefore, it is crucial for teachers to be aware of possible biases of their own perception so that they are able to reflect self-critically on their judgements on their students. furthermore, teachers who are aware that classroom disturbances are an interactionist problem may reflect consciously on various intervention strategies (thommen & wettstein, 2007). 7.4 research perspectives the two-level ct-c(m-1) model that has been presented in this paper forms the basis for further analyses in the context of the sugus study. as a next step, we plan to investigate causes and preventions of classroom disturbances according to our theoretical model. however, these analyses need to overcome further methodological challenges (koch, holtmann, bohn, & eid, 2017). the ct-c(m-1) modelling technique could henceforth serve as an advantageous tool in educational research, because it estimates the common view of teachers and students as well as rater-specific divergences. it could be applied, for instance, in studies on instructional quality (pham et al., 2012) in order to investigate causes of interrater (dis-)agreement. keypoints the common view of teachers and students on classroom disturbances is only marginal. teacher, peer and self-ratings of deviant student behaviour are consistent to an extent of 7 to 54%. rater effects explain up to 70% of variance in the teachers’ and the students’ perception of disturbance. the rater effects are generalisable: teachers and students who overestimate a student in one aspect overestimate this student in other aspects too. acknowledgments i would like to thank the swiss national science foundation for supporting the sugus study (project no.: 100019_152722) and the project leader prof. em. dr. kurt reusser for his substantial support and his invaluable encouragement.furthermore, i would like to thank two anonymous reviewers for providing judicious comments that helped to improve this articles’ quality. references arbuckle, c., & little, e. (2004). teachers’ perceptions and management of disruptive classroom behaviour during the middle years. australian journal of educational and developmental psychology, 4, 59–70. beaman, r., wheldall, k., & kemp, c. (2007). recent research on troublesome classroom behaviour: a review. australasian journal of special education, 31(1), 45–60. doi:10.1080/10300110701189014 becker, h. s. (1963). outsiders. studies in the sociology of deviance. new york, ny: the free press. carretero-dios, h., eid, m., & ruch, w. (2011). analyzing multitrait-mulitmethod data with multilevel confirmatory factor analysis: an application to the validation of the state-trait cheerfulness inventory. journal of research in personality, 45(2), 153–164. doi:10.1016/j.jrp.2010.12.007 carroll, a., houghton, s., taylor, m., west, j., & list-kerz, m. (2006). responses to interpersonal and physically provoking situations. educational psychology, 26(4), 483–498. doi:10.1080/14616710500342424 chen, f. f. (2007). sensitivity of goodness of fit indexes to lack of measurement invariance. structural equation modeling, 14(3), 464–504. doi:10.1080/10705510701301834 christ, t. j., riley-tillman, t. c., & chafouleas, s. m. (2009). foundation for the development and use of direct behavior rating (dbr) to assess and evaluate student behavior. assessment for effective intervention, 34(4), 201–213. doi:10.1177/1534508409340390 crawshaw, m. (2015). secondary school teachers’ perceptions of student misbehaviour: a review of international research, 1983 to 2013. australian journal of education, 59(3), 293–311. doi:10.1177/0004944115607539 dicke, t., parker, p. d., marsh, h. w., kunter, m., schmeck, a., & leutner, d. (2014). self-efficacy in classroom management, classroom disturbances, and emotional exhaustion: a moderated mediation analysis of teacher tandidates. journal of educational psychology, 106(2), 569–583. doi:10.1037/a0035504 doyle, w. (2006). ecological approaches to classroom management. in c. evertson & c. weinstein (eds.), handbook of classroom management: research, practice and contemporary issues (pp. 97–125). mahwah, nj: erlbaum. eckstein, b. (2018). unterrichtsstörungen: eine frage der perspektive? [classroom disturbances: a question of perspective?]. in s. schwab, g. tafner, s. luttenberger, h. knauder, & m. reisinger (eds.), von der wissenschaft in die praxis? zum verhältnis von forschung und praxis in der bildungsforschung (pp. 78–92). münster: waxmann. eckstein, b., grob, u., & reusser, k. (2016). unterrichtliche devianz und subjektives störungsempfinden. entwicklung eines instrumentariums zur erfassung von unterrichtsstörungen [deviant classroom behavior and subjective perception of disturbance. development of an instrument to assess classroom disturbances]. empirische pädagogik, 30(1), 113–129. eckstein, b., grob, u., & reusser, k. (2017). production and perception of classroom disturbances. personal and contextual facets . paper presented at the earli biennial main conference. education in the crossroads of economy and politics, tampere (fin). eckstein, b., luger, s., grob, u., & reusser, k. (2016). sugus – studie zur untersuchung gestörten unterrichts. ergebnisbericht der hauptstudie – anonymisierte fassung [sugus study to investigate classroom disturbances. result report of the main study anonymised version] . zürich: universität zürich. eckstein, b., luger, s., grob, u., & reusser, k. (2018). sugus: technischer bericht der quantiativen teilstudie. studiendesign, stichprobe und skalendokumentation [sugus: technical report of the quantitative study. study design, sample and documentation of the scales] . zürich: universität zürich. eckstein, b., reusser, k., grob, u., & hofstetter, a. (2015). sugus – studie zur untersuchung gestörten unterrichts. kurzer ergebnisbericht der vorstudie – anonymisierte fassung [sugus study to investigate classroom disturbances. short result report of the pilot study anonymised version] . zürich: universität zürich. eid, m., lischetzke, t., nussbeck, f., & trierweiler, l. (2003). separating trait effects from trait-specific method effects in multitrait-multimethod models: a multiple-indicator ct-c (m-1) model. psychological methods, 8(1), 38–60. doi:10.1037/1082-989x.8.1.38 eid, m., nussbeck, f., geiser, c., cole, d. a., gollwitzer, m., & lischetzke, t. (2008). structural equation modeling of multitrait-multimethod data: different models for different types of methods. psychological methods, 13(3), 230–253. doi:10.1037/a0013219 emmer, e. t., & sabornie, e. j. (2015). handbook of classroom management(2nd ed.). new york, ny: routledge. finney, s. j., & distefano, c. (2013). nonnormal and categorical data in structural equation models. in g. r. hancock & r. o. mueller (eds.), structural equation modeling: a second course(2nd ed., pp. 439–492). charlotte, nc: information age publishing. fishbein, m., & ajzen, i. (2010). predicting and changing behavior. the reasoned action approach. new york, ny: psychology press. geiser, c., eid, m., west, s. g., lischetzke, t., & nussbeck, f. w. (2012). a comparison of method effects in two confirmatory factor models for structurally different methods. structural equation modeling: a multidisciplinary journal, 19(3), 409-436. godwin, k. e., almeda, m. v., seltman, h., kai, s., skerbetz, m. d., baker, r. s., & fisher, a. v. (2016). off-task behavior in elementary school children. learning and instruction, 44, 128–143. doi:10.1016/j.learninstruc.2016.04.003 hamre, b. k., pianta, r. c., downer, j. t., & mashburn, a. j. (2008). teachers’ perceptions of conflict with young students: looking beyond problem behaviors. social development, 17(1), 115–136. doi:10.1111/j.1467-9507.2007.00418.x hempel-jorgensen, a. (2009). the construction of the “ideal pupil” and pupils’ perceptions of “misbehaviour” and discipline: contrasting experiences from a low-socio-economic and a high-socio-economic primary school. british journal of sociology of education, 30(4), 435–448. doi:10.1080/01425690902954612 hofer, m. (1986). sozialpsychologie erzieherischen handelns [social psychology of education] . goettingen: hogrefe. hoyt, w. t. (2000). rater bias in psychological research: when is it a problem and what can we do about it? psychological methods, 5(1), 64-86. doi:10.1037/1082-989x.5.1.64 hoyt, w. t., & kerns, m.-d. (1999). magnitude and moderators of bias in observer ratings: a meta-analysis. psychological methods, 4(4), 403–424. hunt, d., carline, j., tonesk, x., yergan, j., siever, m., & loebel, j. (1989). types of problem students encountered by clinical teachers on clerkships. medical education, 23(1), 14–18. infantino, j., & little, e. (2005). students’ perceptions of classroom behaviour problems and the effectiveness of different disciplinary methods. educational psychology, 25(5), 491–508. doi:10.1080/0144341050004654 jak, s., oort, f. j., & dolan, c. v. (2013). a test for cluster bias: detecting violations of measurement invariance across clusters in multilevel data. structural equation modeling, 20(2), 265–282. doi:10.1080/10705511.2013.769392 janík, t., & seidel, t. (eds.). (2009). the power of video studies in investigating teaching and learning in the classroom . münster: waxmann. kleinke, k., schlüter, e., & christ, o. (2017). strukturgleichungsmodelle mit mplus. eine praktische einführung [structural equation modelling with mplus. a practical introduction] (2nd ed.). berlin: de gruyter. koch, t., holtmann, j., bohn, j., & eid, m. (2017). explaining general and specific factors in longitudinal, multimethod, and bifactor models: some caveats and recommendations. psychological methods. doi:10.1037/met0000146 kokkinos, c. m. (2007). job stressors, personality and burnout in primary school teachers. british journal of educational psychology, 77, 229–243. doi:10.1348/000709905x90344 kunter, m., & baumert, j. (2006). who is the expert? construct and criteria validity of student and teacher ratings of instruction. learning environments research, 9(3), 231-251. doi:10.1007/s10984-006-9015-7 little, e. (2005). secondary school teachers’ perceptions of students’ problem behaviours. educational psychology, 25(4), 369–377. doi:10.1080/01443410500041516 little, t. d., cunningham, w. a., shahar, g., & widaman, k. f. (2002). to parcel or not to parcel: exploring the question, weighing the merits. structural equation modeling, 9(2), 151–173. doi:10.1207/s15328007sem0902_1 makarova, e., herzog, w., & schönbächler, m.-t. (2014). wahrnehmung und interpretation von unterrichtsstörungen aus schülerperspektive sowie aus sicht der lehrpersonen [perception and interpretation of classroom disturbances in the perspectives of students and teachers]. psychologie in erziehung und unterricht, 61(2), 127–140. montuoro, p., & lewis, r. (2015). student perceptions of misbehavior and classroom management. in e. t. emmer & e. j. sabornie (eds.), handbook of classroom management(2nd ed., pp. 344–362). new york, ny: routledge. müller, c. m., & hofmann, v. (2016). does being assigned to a low school track negatively affect psychological adjustment? a longitudinal study in the first year of secondary school. school effectiveness and school improvement: an international journal of research, policy and practice, 27 (2), 95–115. doi:10.1080/09243453.2014.980277 muthén, b. o., & muthén, l. k. (2017). mplus user’s guide(8th ed.). los angeles, ca: muthén & muthén. ohan, j., visser, t. a. w., strain, m. c., & allen, l. (2011). teachers’ and education students’ perceptions of and reactions to children with and without the diagnostic label “adhd”. journal of school psychology, 49, 81–105. doi:10.1016/j.jsp.2010.10.001 pham, g., koch, t., helmke, a., schrader, f.-w., helmke, t., & eid, m. (2012). do teachers know how their teaching is perceived by their pupils? procedia-social and behavioral sciences, 46, 3368-3374. doi:10.1016/j.sbspro.2012.06.068 preuss-lausitz, u. (2005). verhaltensauffällige kinder integrieren. zur förderung der emotionalen und sozialen entwicklung [integrating children with behavioural problems. fostering the emotional and social development] . weinheim: beltz. scherzinger, m., wettstein, a., & wyler, s. (2017). unterrichtsstörungen aus der sicht von schülerinnen und schülern und ihren lehrpersonen. ergebnisse einer interviewstudie zum subjektiven erleben von störungen [classroom disturbances in the perspectives of students and teachers. findings of an interview study on the subjective experience of disturbances] vierteljahresschrift für heilpädagogik und ihre nachbargebiete, 86 (1), 70–83. doi:10.1026/0049-8637/a000159. schönbächler, m.-t., makarova, e., herzog, w., altin, ö., känel, s., lehmann, v., & milojevic, s. (2009). klassenmanagement und kulturelle heterogenität: ergebnisse 2. forschungsbericht nr. 37. [classroom management and cultural heterogeneity: results 2. research report no. 37] . bern: universität bern. sherman, j., rasmussen, c., & baydala, l. (2008). the impact of teacher factors on achievement and behavioural outcomes of children with attention deficit/hyperactivity disorder (adhd): a review of the literature. educational research, 50(4), 347–360. doi:10.1080/00131880802499803 südkamp, a., kaiser, j., & möller, j. (2012). accuracy of teachers’ judgments of students’ academic achievement: a meta-analysis. the journal of educational psychology, 104(3), 743–762. doi:10.1037/a0027627 thommen, b., & wettstein, a. (2007). toward a multi-level-analysis of classroom disturbances. european journal of school psychology, 5 (1), 65-82. venetz, m., zurbriggen, c., & schwab, s. (2017). what do teachers think about their students’ inclusion? consistency of selfand teacher reports . paper presented at the earli biennial main conference. education in the crossroads of economy and politics, tampere (fin). wagner, w., gollner, r., werth, s., voss, t., schmitz, b., & trautwein, u. (2016). student and teacher ratings of instructional quality: consistency of ratings over time, agreement, and predictive power. journal of educational psychology, 108(5), 705–721. wettstein, a. (2012). a conceptual frame model for the analysis of aggression in social interactions. journal of social, evolutionary, and cultural psychology, 6(2), 141–157. doi:10.1037/h0099218 wettstein, a., ramseier, e., & scherzinger, m. (2018). eine mehrebenenanalyse zur schülerwahrnehmung von störungen im unterricht der klassenund einer fachlehrperson [a multilevel analysis of the students’ perception of classroom disturbances during the instruction of class teachers and subject teachers]. psychologie in erziehung und unterricht, 65(1), 1–16. doi:10.2378/peu2018.art01d wettstein, a., ramseier, e., scherzinger, m., & gasser, l. (2016). unterrichtsstörungen aus lehrerund schülersicht. aggressive und nicht aggressive störungen im unterricht aus der sicht der klassen-, einer fachlehrperson und der schülerinnen und schüler [classroom disturbances in the perspectives of teachers and students. aggressive and non-aggressive disturbances during instruction in the view of class teachers, subject teachers, and students]. zeitschrift für entwicklungspsychologie und pädagogische psychologie, 48 (4), 171–183. doi:10.1026/0049-8637/a000159 wettstein, a., scherzinger, m., & ramseier, e. (2018). unterrichtsstörungen, beziehung und klassenführung aus lehrer-, schülerund beobachterperspektive [classroom disturbances, relationship, and classroom management in the perspectives of teachers, students and external observers]. psychologie in erziehung und unterricht, 65(1), 58-74. doi:10.2378/peu2018.art04d wolfe, e. w. (2004). identifying rater effects using latent trait models. psychology science, 46, 35-51. zevenbergen, r. (2001). mathematics, social class, and linguistic capital. in b. atweh, h. forgasz, & b. nebres (eds.), sociocultural research on mathematics education(pp. 201–215). mahwah, nj: erlbaum. zurbriggen, c., venetz, m., & hinni, c. (2018). the quality of experience of students with and without special educational needs in everyday life and when relating to peers. european journal of special needs education, 33(2), 205–220. doi:10.1080/08856257.2018.1424777 appendix a1. undisciplined behaviour: item wording and parcelling appendix a2. dissocial behaviour: item wording and parcelling note: items d16, d17, and d18 included an extra prompt: “how was it shortly before or after the lesson, for example on your way to school or during the break?” the intention behind this addition was to address dissocial forms of behaviour that had occurred outside the classroom but might have had consequences during the lesson, for example if the students involved had still been emotionally upset. appendix a3. affective perception of disturbance: item wording and parcelling note. items a01r and a09r had been positively worded so that the raters could describe the targets in a favourable way in order to prevent a negative stigmatisation of the target student due to the survey. the original answers were afterwards recoded into inverted values (r). appendix a4. cognitive perception of disturbance: item wording and parcelling appendix b. mplus input file of the two-level ct-c(m-1) model microsoft word garrote_publication.docx           frontline  learning  research  vol.5  no.  1  (2017)  1  -­‐  15   issn  2295-­‐3159         the relationship between social participation and social skills of pupils with an intellectual disability: a study in inclusive classrooms ariana garrote1 university of zurich, switzerland article received 5 august / revised 31 october / accepted 1 november / available online 26 january abstract researchers claim that a lack of social skills might be the main reason why pupils with special educational needs (sen) in inclusive classrooms often experience difficulties in social participation. however, studies that support this assumption are scarce, and none include pupils with an intellectual disability (id). this article seeks to make an important contribution to this discussion. the social skills and social participation of pupils with id and their typically developing (td) peers in 38 general education classrooms were assessed with multidimensional instruments. the analyses indicate that the majority of pupils with id were not popular but were socially accepted and had friends. additionally, no significant relationship was found between social skills and the social participation of pupils with id, although such pupils had lower levels of social skills compared with their td peers. thus, it appears that pupils with id do not require high levels of social skills to be befriended or accepted by classmates. in contrast, social skills were associated with popularity and social acceptance within the group of td pupils. in fact, popular td pupils had the highest level of social skills. these findings support the assumption that in addition to low levels of social skills, there must be other mechanisms that influence the social participation of pupils with id in inclusive classrooms. keywords: social skills; social participation; special educational needs; sociometric status; friendship                                                                                                                           1 contact information: ariana garrote, department of educational sciences, university of zurich, freiestrasse 36, 8032 zürich, switzerland. e-mail: agarrote@ife.uzh.ch doi: http://dx.doi.org/10.14786/flr.v5i1.266 garrote     | f l r     2   1. introduction children learn a wide range of social behaviours and skills in their social interactions with peers. their experiences within this social context influence their socio-emotional development and their later adjustment during adulthood (rubin, bukowski, & parker, 2006). this phenomenon is one reason why the un convention recommends the social participation of pupils with disabilities in the community and the classroom (united nations, 2006). consequently, an increasing tendency to include pupils with special educational needs (sen) in general education schools is observed internationally (koster, nakken, pijl, & van houten, 2009; pijl, frostad, & flem, 2008; ruijs & peetsma, 2009). however, the mere presence of these pupils in general education classrooms does not automatically result in successful social participation. pupils with sen are less involved in social interactions with peers, less accepted and more frequently rejected than their typically developing (td) peers (avramidis, 2013; estell et al., 2008; feldman, carter, asmus, & brock, 2015; garrote, 2016; grütter, meyer, & glenz, 2015; huber, 2006; koster, pijl, nakken, & van houten, 2010; nepi, fioravanti, nannini, & peru, 2015; pijl & frostad, 2010; rotheram-fuller, kasari, chamberlain, & locke, 2010). on one hand, studies demonstrate that td pupils do not like to work with their low-achieving peers, including peers with sen (huber & wilbert, 2012; krull, wilbert, & henneman, 2014; monchy, pijl, & zandberg, 2004). thus, teachers tend to avoid mixing pupils with sen and their td peers for certain activities, such as group work, and therefore, pupils with sen lack shared learning experiences. in fact, feldman et al. (2015) found that due to their teacher’s planning students with severe disabilities, included in general education classrooms, were not present in most of the classes and were not physically close enough to their peers if present. therefore, pupils with sen not only miss opportunities to interact and to become involved in social relationships but also do not receive the chance to learn and practice social skills with their td peers. on the other hand, researchers hypothesize that the difficulties pupils with sen experience with social participation may be due to their lack of social skills (e.g., avramidis, 2013; huber & wilbert, 2012; pijl et al., 2008; schwab, gebhardt, krammer, & gasteiger-klicpera, 2015). however, the relationship between the social participation of pupils with sen and their social skills has not been thoroughly studied, and the concepts of social participation and sen vary from study to study. typically, the samples of pupils with sen are heterogeneous and include pupils with different needs and problems (e.g., children with behavioural problems, with learning disabilities, or with physical disabilities). therefore, valid knowledge on the topic is lacking. this study contributes to closing this research gap. it investigates to what extent the social participation of pupils with sen in general education classrooms is associated with their social skills. social participation and social skills were measured in a large sample of pupils with a diagnosed intellectual disability (id) and their td peers. this study contributes the following novelties to the literature. first, although there are a number of studies investigating the social participation of pupils with sen enrolled in inclusive classrooms, there is little knowledge about influencing factors. thus, this study focusses on social skills, which are assumed to influence the social participation of pupils with sen. second, the heterogeneity of the group of pupils with sen was decreased by limiting the sample to pupils with id. third, social participation was measured using a multidimensional approach that included aspects of social relationships (i.e., friendships) and social acceptance (i.e., popularity and rejection), which are two important dimensions of social participation (bossaert, colpin, pijl, & petry, 2013; koster et al., 2009). fourth, to measure social skills, an empirically well-studied scale with a theoretical foundation was used. fifth, until now, no findings have been reported regarding the relationship between the social participation and the social skills of pupils with id in inclusive classrooms. finally, this study shall give new insights about processes influencing social participation in inclusive classrooms, to derive important implications for further research and for the development of suitable interventions. garrote     | f l r     3   1.1. social skills there is no commonly accepted concept of social skills (or competences). however, there is consensus regarding the connection between social skills and successful social interactions as well as the ability to establish and maintain positive social relationships. socially competent individuals are described as being able to use social interactions to satisfy their goals and needs while considering the needs and goals of others (groeben, perren, stadelmann, & klitzing, 2011; perren, forrester-knauss, & alsaker, 2012; rosekrasnor, 1997). this definition differentiates between social skills that are important for the self and social skills that are oriented towards the others. malti and perren (2016) term these two dimensions of social skills selfand other-oriented. initiating and maintaining social interactions, leadership skills and the ability to set limits with peers are considered to be self-oriented skills because they aim at satisfying one’s own needs. other-oriented social skills, such as helping, caring, and cooperating, are based on considering the interests and benefits of others in social interactions. whereas deficits in selfand other-oriented social skills have been associated with negative peer relations, peer rejection, and victimization (bellini, peters, benner, & hopf, 2007; perren et al., 2012; malti & perren, 2016; henricsson & rydell, 2006), engaging in prosocial behaviour and being able to initiate social interactions can help children become positively involved with peers (fabes, martin, & hanish, 2009; henricsson & rydell, 2006; perren, argention-groeben, stadelmann, & klitzing, 2016; rubin et al., 2006). 1.2. social skills of pupils with sen there is evidence that certain groups of pupils with sen are less socially skilled than td children (gresham & macmillan, 1997). for example, pupils with autism spectrum disorders (asd) have difficulties with self-oriented social skills, such as initiating social interactions, which places them more at risk of being socially isolated (bellini et al., 2007). however, intervention studies have demonstrated that these children can benefit from the supportive behaviour of their td peers with respect to the development of social interaction skills (camargo et al., 2014; whalon, conroy, martinez, & werch, 2015). in addition, engaging in other-oriented social skills, such as cooperative and prosocial behaviour, can also have a positive impact on the social participation of pupils with asd (kamps et al., 2002). similar results have been found for pupils with intellectual disabilities (goldstein, english, shafer, & kaczmarek, 1997), learning disabilities (kavale & forness, 1996), and behavioural problems (frederickson & turner, 2003). however, few studies have investigated the relationship between social skills and the social participation of pupils with sen, and the findings are ambiguous. for example, frostad and pijl (2007) found a weak relationship between the social skills (cooperative behaviour and empathy) and the social position (acceptance, friendships and membership in a subgroup) of pupils with sen. however, a separate examination of the group of pupils with learning difficulties and the group of pupils with behaviour problems changed the results. while there was no relation between the social position and the social skills of pupils with learning difficulties, there was a significant relationship for pupils with behavioural problems. the authors concluded that a low level of empathy (i.e., exhibiting concern and respect for the feelings and viewpoint of others) might be an explanation of social difficulties only for pupils with behavioural problems. further, schwab et al. (2015) described a link between self-rated social participation and prosocial behaviour. students with sen (not specified) in secondary schools felt less socially included and reported lower levels of prosocial behaviour than their td peers. thus, the authors interpreted that the poor social participation of pupils with sen might be associated with their reported low levels of prosocial behaviour. this study attempts to contribute to the clarification of the relationship between the social skills and the social participation of pupils with sen in general education classrooms. to obtain more focused findings, the sample of pupils with sen consisted only of children with a diagnosed id. these pupils were compared with selected td peers with respect to social acceptance, social relationships, and selfand otheroriented social skills. the following research questions were addressed: garrote     | f l r     4   a) how does the social participation of pupils with id in general education classrooms appear in terms of social relationships and social acceptance? b) how do peers and teachers rate the social skills of pupils with id compared with popular, rejected, and isolated td pupils? c) is there a relationship between the social participation of pupils with id and their social skills? 2. method 2.1. sample the sample consisted of 38 inclusive primary classrooms of the germanand french-speaking parts of switzerland. a total of 692 firstto fourth-graders (aged m = 97.92; sd = 10.51) participated in the study. only general education classrooms that included pupils with id were allowed in the study.2 therefore, in each classroom at least one pupil was officially diagnosed by a school psychologist as having id. based on this diagnosis special education resources were allocated to support these pupils3. in a first step, all 692 participants were individually interviewed regarding their social involvement in the classroom and the social skills of their peers. in a second step, popular, rejected, and isolated pupils were identified based on the collected data (for a detailed description, see the analyses). finally, social skills and peer relationships were estimated by teachers for the selected td pupils (n = 89; 51.6% female) and for the sample of pupils with id (n = 43; 39.5% female). 2.2. measures social acceptance, social relationships, and social skills were assessed with teacher questionnaires and by individual pupil interviews. the instruments used for the interviews were developed and piloted to be suitable for pupils with id. 2.2.1 social acceptance and social relationships to assess social acceptance, the rating method was applied. all pupils rated how much they liked to play with each classmate on a five-point-scale using smileys (1 = l = “i do not like to play with this classmate at all” to 5 = j = “i like to play with this classmate a lot”). each classmate’s name was read to the participants and presented on an individual card. the social relationships of pupils within the classroom were assessed using the nomination procedure. all participants were asked to nominate their regular playmates in the classroom (“with whom in your classroom do you play the most?”). the number of sameand cross-gender nominations was unlimited. teachers also estimated the social acceptance and relationships of the pupils. the teacher version of the selfand other-oriented social competences questionnaire (socomp; see perren et al., 2012) includes the subscale positive peer relationships (5 items; α = 0.89). the items in this subscale primarily address                                                                                                                           2 pupils with id were full-time members of the general education classrooms. in each classroom, a special education teacher was present for four to 14 hours per week (m = 7.22, sd = 3.36). 3 the school psychologists applied the criterion iq < 75 to diagnose id at the time of the data collection. here, it is important to mention that iq tests have their limitations when it comes to assessments of children with first language different from the test language or who are in a difficult affective state. in addition, a dichotomous categorisation of td pupils and pupils with id does not reflect disability as a social construct. nevertheless, this categorisation is used in the present study because it is common in practice and affects processes of social participation.   garrote     | f l r     5   friendships (e.g., “has at least one good friend”) and social acceptance (e.g., “is generally popular among peers”). 2.2.2 social skills the social skills were rated by teachers and peers. the teachers were requested to estimate the social skills of their pupils with id and selected td pupils (popular, rejected, and isolated; for criteria, see 2.3.1) using the socomp questionnaire. the teachers were not informed regarding the selection criteria of the td pupils they rated, which means they were unaware of their sociometric status (popular, rejected, or isolated). all of the socomp items were estimated on a three-point scale (0 = “not true at all” to 2 = “definitely true”). two dimensions of social skills were assessed: self-oriented social skills (α = 0.88) with the three subscales leadership (3 items; e.g., “organizes, suggests play activities to peers”), setting limits (3 items; e.g., “refuses unreasonable requests from others”), and social participation (4 items; e.g., “converses with peers easily”), and other-oriented social skills (α = 0.88), including the two subscales prosocial (5 items; e.g., “frequently helps other children”) and cooperative behaviour (5 items; e.g., “compromises in conflicts with peers”). the other-oriented social skills of the pupils were also estimated by peers with two questions regarding cooperative and prosocial behaviour (α = 0.83). all of the participants rated on a five-point-scale with smileys (1 = l = “i do not agree at all” to 5 = j = “i totally agree”) four randomly selected classmates with respect to how well they could work with them and how helpful they were. 2.3. analyses 2.3.1 social status popular, rejected, and isolated pupils were selected for the study based on individual interviews with all of the participants. to identify popular and rejected pupils, the rating data were analysed. pupils who received the highest score on the scale (5) from most of their peers (at least one standard deviation above average) were categorized as popular. those pupils who received the lowest score on the scale (1) from most of their peers (at least one standard deviation above average) were categorized as rejected. to identify isolated pupils, indegree and outdegree scores were calculated based on the nomination data with ucinet (borgatti, everett, & freeman, 2002). children who were not nominated by any classmate and did not nominate anyone as a playmate (indegree = 0 = outdegree) were categorized as isolated. however, pupils were only identified as isolated if at least 80% of the pupils of the classroom had participated in the study. this cut-off criterion was established to respect the interdependency of network data (huisman & steglich, 2008; robins, pattison, & woolcock, 2004). table 1 shows the distribution of sociometric status among td pupils and pupils with id. in 38 classrooms, 89 td pupils were identified as popular, rejected, or isolated. in the case of 14 pupils, there was an overlap of the sociometric statuses “rejected” and “isolated”. for 46.5% of the pupils with id, a classification into status groups was possible. while most of these pupils were in the category rejected, five were identified as isolated and only one as popular. the rest of the pupils with id (53.5%) were “average”, which means they were not rejected, isolated, or popular. garrote     | f l r     6   table 1 sociometric status of pupils with id and their td peers sample n popular rejected isolated isolated and rejected average pupils with id 43 1 14 4 1 23 td pupils 89 38 19 18 14 2.3.2 social acceptance and social relationships several scores for social acceptance (popularity and social rejection) and social relationships (number of friendships and at least one friend) were computed and standardized for each participant. therefore, the rating data, the nomination data, and the subscale on positive peer relationships of the teacher questionnaire were analysed. the rating data were used to calculate the scores for social acceptance. the popularity score consists of the sum of the highest rating on the scale (5) received by all classmates, whereas the social rejection score corresponds to the sum of the lowest ratings on the scale (1) received by all classmates. in addition, a mean score of social acceptance was calculated for each pupil with all ratings from 1 to 5 received by peers. for each pupil, two social relationship scores were calculated. one was the number of reciprocated friendships. the other was a dichotomized score, which represents whether the pupil had at least one friend or none. according to the common practice, reciprocal nominations were defined as friendships (hymel, vaillancourt, mcdougall, & renshaw, 2004). because social network data are strongly influenced by the participant ratio, the values were divided by the maximum number of possible friendships in the classroom ((n * (n – 1)) / 2). the teacher’s view of the social relationships and social acceptance of the pupils was represented with standardized sums of the socomp dimension of positive peer relationships. 2.3.3 social skills the social skills were estimated by teachers and peers. for the teacher’s perspective, the standardized sums of the two dimensions selfand other-oriented social skills of the socomp questionnaire were calculated for each pupil. based on the peer ratings regarding the cooperative and prosocial behaviour of classmates, a standardized mean score was calculated for each pupil. this score represents the otheroriented social skills from the peer’s perspective. 3. results 3.1. social acceptance and social relationships practically all popular td children had at least one friend and on average the highest number of friends compared with the other pupils (table 2). in addition, more than half of the rejected td pupils had friends. a majority of the pupils with id (63%) had at least one reciprocal friend. interestingly, this was the case for most of the rejected and for most of the average pupils with id but not for the only popular pupil in this sample. garrote     | f l r     7   table 2 friendships of pupils with id and their td peers friendships group n m (sd) at least one friend (%) pupils with id 43 .98 (1.08) 27 (63) popular 1 rejected 14 .71 (.47) 10 (71) average 23 1.39 (1.27) 17 (74) td pupils 89 1.18 (1.48) 48 (54) popular 38 2.32 (1.53) 35 (92) rejected 19 .89 (.83) 13 (68) isolated pupils are not represented in this table. 3.2. social skills the relationship between the social skills of pupils with id, their sociometric status, and their friendships were analysed with nonparametric tests and correlations. in a first step, comparisons within the status groups of pupils with id and of td pupils were performed. in a second step, status groups of pupils with id were compared with status groups of td pupils. 3.2.1 pupils with id according to teacher reports, popular and accepted pupils with id (n = 24) did not differ significantly from rejected and isolated pupils with id (n = 19) in their social skills (table 3). in addition, only weak correlations were found between the mean acceptance score of pupils with id and their selforiented (r = 0.119) and other-oriented (r = 0.197) social skills. however, differences were found with respect to their positive peer relationships (u = 103.50, z = -3.07, p = 0.002). consistent with their sociometric status, rejected and isolated pupils with id were estimated by teachers to be less accepted and less popular among their peers. additionally, these pupils were perceived by their peers as having lower levels of other-oriented social skills (u = 140.00, z = -2.15, p = 0.031) than the popular and accepted pupils with id. in line with this result, the correlation between the other-oriented social skills rated by peers and the mean acceptance score was high (r = 0.662). however, rejected and isolated pupils with id did not have significantly fewer friends than popular and accepted pupils with id. for further analyses, the sample of pupils with id was split into two groups: pupils with at least one friend (n = 27; 62% accepted and 37% rejected) and pupils without friends (n = 16; 62% accepted and/or isolated, 31% rejected, and 6% popular). regarding social acceptance and rejection, the two groups did not differ significantly. in addition, based on peer and teacher estimations, no significant differences were found regarding self-oriented or other-oriented skills. surprisingly, the groups were also similar in their positive peer relationships reported by teachers. there was only a significant difference regarding the item “has at least one good friend” (u = 127.00, z = -2.52, p = 0.012), which indicates that teachers notice when pupils with id do not have friends. garrote     | f l r     8   table 3 social skills mean values (sd) of pupils with id teacher-reported peer-reported group n self-oriented other-oriented other-oriented pupils with id 43 9.35 (3.92) 11.53 (4.39) 3.08 (.77) popular & average 24 10.29 (4.20) 12.42 (4.37) 3.35 (.62) rejected & isolated 19 8.58 (3.42) 10.42 (4.27) 2.74 (.82)* with friends 27 9.52 (3.98) 11.63 (4.85) 3.09 (.71) without friends 16 9.56 (3.97) 11.38 (3.63) 3.06 (.89) the values displayed are not standardised. the teacher-reported values reported range from 0 to 20, and the peerreported values range from 1 to 5. the indicated significant differences refer to a comparison with the value in the row above (* p < 0.05). 3.2.2 td pupils according to the teacher reports, popular td pupils (n = 38) had significantly higher levels of selforiented social skills (u = 588.50, z = -3.16, p = 0.002), other-oriented social skills (u = 305.50, z = -5.53, p = 0.000), and positive peer relationships (u = 142.50, z = -6.98, p = 0.000) than rejected and isolated td pupils (n = 51). additionally, rejected and isolated td pupils had fewer friends (u = 312.50, z = -6.60, p = 0.000) and were estimated by peers as exhibiting a lower level of other-oriented social skills (u = 107.00, z = -7.15, p = 0.000) than popular td pupils (table 4). in addition, high correlations were found between the mean acceptance score and the other-oriented social skills reported by teachers (r = 0.616) and by peers (0.727). however, the correlation between the mean acceptance score and the self-oriented social skills was low (r = 0.254). there were significant differences between td pupils with at least one friend (n = 48; 73% popular and 27% rejected) and without friends (n = 41; 48% rejected, 43% isolated, and 7% popular). td pupils with friends were estimated by teachers as having higher self-oriented (u = 1406.50, z = 3.49, p = 0.000) and other-oriented social skills (u = 1449.00, z = 3.85, p = 0.000). similar results emerged regarding peerreported other-oriented social skills (u = 1504.50, z = 4.28, p = 0.000). thus, according to teacher and peer reports, td pupils without friends tend to be less socially competent than td pupils with friends. in addition, the teacher reports displayed differences regarding positive peer relationships (u = 1578.50, z = 4.98, p = 0.000). this outcome indicates that teachers notice when pupils do or do not have friends. additionally, td pupils with and without friends differed significantly with respect to social acceptance (u = 1611.00, z = 5.16, p = 0.000) and rejection (u = 626.00, z = -2.82, p = 0.005). thus, popular td pupils tend to have friends, whereas rejected td pupils tend to be friendless. these accentuated results might have appeared because the subsample consists of pupils of extreme status groups: the most popular and the most rejected pupils in the classroom. this fact must be considered when interpreting the results.   garrote     | f l r     9   table 4 social skills mean values (sd) of td pupils teacher-reported peer-reported group n self-oriented other-oriented other-oriented td pupils 89 13.02 (5.42) 14.03 (5.15) 3.60 (.97) popular 38 15.08 (4.56) 17.50 (3.11) 4.35 (.56) rejected & isolated 51 11.49 (5.54)** 11.45 (4.86)*** 3.04 (.82)*** with friends 48 14.90 (4.57) 15.92 (4.58) 3.97 (.83) without friends 41 10.83 (5.55)*** 11.78 (4.91)*** 3.17 (.94)*** the values displayed are not standardised. the teacher-reported values range from 0 to 20, and the peer-reported values range from 1 to 5. the indicated significant differences refer to a comparison with the value in the row above (*** p < 0.001. ** p < 0.01). 3.2.3 comparison of pupils with id with td pupils compared with rejected td pupils (n = 19), rejected pupils with id (n = 14) had a significantly lower level of self-oriented social skills (u = 54.50, z = -2.87, p = 0.003; table 5). however, they did not differ with respect to their other-oriented social skills and positive peer relationships reported by their teachers. in addition, the result of a comparison of all pupils with id (n = 43) with the rejected and/or isolated td pupils (n = 51) revealed no significant differences regarding social skills, whereas regarding social relationships, a difference was found: an advantage for pupils with id. the teachers estimated the positive peer relationships of pupils with id on a higher level than of rejected and/or isolated td pupils (u = 777.00, z = -2.44, p = 0.015). in addition, pupils with id had significantly more friends than rejected and/or isolated td pupils (u = 660.50, z = -3.68, p = 0.000). table 5 social skills mean values (sd) of pupils with id and td pupils teacher-reported peer-reported group n self-oriented other-oriented other-oriented rejected pupils with id vs. 14 7.93 (3.45) 10.93 (4.51) 2.71 (.78) rejected td pupils 19 13.26 (5.34)** 10.79 (4.76) 2.95 (.77) pupils with id vs. 43 9.35 (3.92) 11.53 (4.39) 3.08 (.77) popular td pupils 38 15.08 (4.56)*** 17.50 (3.11)*** 4.35 (.56)*** rejected & isolated td pupils 51 11.49 (5.54) 11.45 (4.86) 3.04 (.82) the values displayed are not standardised. the teacher-reported values range from 0 to 20, and the peer-reported values range from 1 to 5. *** p < 0.001. ** p < 0.01. a different set of results emerged from the comparison of pupils with id (n = 43) with popular td pupils (n = 38). significant differences in all aspects were found. the pupils with id were estimated to have lower levels of self-oriented social skills (u = 295.00, z = -4.95, p = 0.000), other-oriented social skills (u = garrote     | f l r     10   212.00, z = -5.75, p = 0.000), positive peer relationships (u = 169.00, z = -6.29, p = 0.000), and peerreported other-oriented social skills (u = 99.00, z = -6.79, p = 0.000) than popular td pupils. in addition, pupils with id had significantly fewer friends than popular td pupils (u = 162.50, z = -6.19, p = 0.000). finally, a fishers’ z-test was performed to compare pupils with id (n = 43) with their td peers (n = 89) regarding the correlations between social acceptance and social skills. only the correlation between social acceptance and the other-oriented social skills rated by the teachers was significantly higher (z = 2.711, p = 0.003) in the group of td pupils (r = 0.616) than in the group of pupils with id (r = 0.197). this result indicates that pupils with id are not less accepted if they have lower levels of other-oriented social skills. in contrast, socially accepted td pupils have higher levels of other-oriented social skills. in addition, the correlations between other-oriented social skills rated by peers and social acceptance were high for both groups and in turn did not differ significantly. this outcome means that for pupils with id as well as for td pupils being socially accepted was related to being perceived by peers as being helpful and cooperative. in contrast, for both groups, the self-oriented social skills were weakly correlated with social acceptance. 4. discussion until now, little was known regarding the social participation of pupils with id enrolled in general education classrooms. in addition, no studies have investigated the role that social skills may play in social participation. this study contributes to clarifying the relationship between the social skills and the social participation of pupils with id in inclusive classrooms. therefore, the social relationships and social acceptance of these pupils were analysed in relation to their selfand other-oriented social skills. in addition, a comparison of social skills was performed between pupils with id and td pupils experiencing a more or less positive social participation. this study reveals that most pupils with id enrolled in inclusive classrooms were not popular but were accepted by their peers. in addition, a majority of these pupils, including those who were rejected, had reciprocal friends. these findings agree with studies which report that not all pupils with sen are at risk of being isolated or rejected in general education classrooms (e.g., avramidis, 2013; frostad, mjaavatn, & pijl, 2011; frostad & pijl, 2007; grütter et al., 2015; koster et al., 2010; schwab, 2015). regarding social skills, in this study, pupils with id were generally less socially competent than their td peers. however, no significant association was found with having friends. pupils with id who had friends exhibited a similar level of social skills as pupils with id who were isolated or did not have reciprocal friendships. thus, it seems that their ability to form and maintain friendships was not influenced by a lack of social skills, and they even appear to possess skills that benefit interaction with peers. according to the literature, children require a number of basic social skills to form friendships (gest, graham-bermann, & hartup, 2001; sebanc, 2003). thus, it seems that the majority of the pupils with id in this sample must possess these basic social skills. this finding is promising because peer relationships can positively contribute to the socio-emotional adjustment of children (gifford-smith & brownell, 2003; murray & greenberg, 2006). in addition, the social skills of pupils with id were not always related to their social acceptance. from the perspective of peers, rejected pupils with id were estimated as being less cooperative and prosocial than accepted pupils with id. however, from the perspective of teachers, the social skills of rejected pupils with id did not differ from those of accepted pupils with id. similar results were presented in a study by frostad and pijl (2007), who found no significant relationship between the social skills of pupils with learning disabilities and their social acceptance or their social relationships. in addition, the two variables were only weakly related when the entire sample of pupils with and without sen was analysed. however, pupils with sen had lower levels of social skills than their td peers. garrote     | f l r     11   in this study, pupils with id were also compared regarding their social skills with td pupils experiencing more or less difficulties in their social participation in classroom. while the social participation of popular td pupils seemed to be satisfactory, rejected and isolated td pupils experienced difficulties in their social participation. in a first step, the rejected and isolated td pupils were compared with the entire sample of pupils with id. as expected, there were no differences between the two groups with respect to social skills. however, pupils with id had fewer difficulties in building and maintaining social relationships than rejected and isolated td pupils. this outcome means that although pupils with id had levels of social skills similarly low to those of their rejected and isolated td peers, they tended to have more friends. however, the latter finding could be biased because isolated pupils are friendless by definition. thus, only rejected pupils with id and rejected td pupils were compared. these two groups displayed similarities in their peer relationships and their other-oriented social skills. however, rejected pupils with id had lower levels of self-oriented social skills than their rejected td peers. based on these results, the obvious conclusion is that pupils with id can experience more positive social relationships than certain of their td peers but tend to have more difficulties in setting limits, initiating social interactions and leading than their rejected td peers. otherwise, the pupils with id have levels of social skills similarly low to those of their rejected and isolated td peers. this result means that in general education classrooms not only pupils with sen but also certain of their td peers lack social skills and can therefore also be at risk of social exclusion. consequently, interventions to foster social participation should not only focus on pupils with sen but also should involve the entire class. in fact, peer-mediated learning activities that involve all pupils can positively influence the social interactions in inclusive classrooms (fuchs, fuchs, mathes, & martinez, 2002; jacques, wilton, & townsend, 1998). in a second step, a comparison between pupils with id and popular td pupils was performed. significant differences between the two groups were found. pupils with id had fewer friends and exhibited lower levels of selfand other-oriented social skills than popular td pupils. this contrast between the two groups was perhaps because practically none of the pupils with id were popular. consequently, it could be argued that pupils with id are not as popular as certain of their td peers because of their lack of social skills. on one hand, a strong positive relationship was found between the level of other-oriented social skills and the social acceptance of td pupils. this outcome could indicate that for td pupils providing particular consideration to the needs and goals of others can positively influence their social acceptance or popularity in a group. a similar finding was reported by gest et al. (2001). in their study, children who were perceived by teachers and peers as socially skilled were more popular among peers. on the other hand, among pupils with id, there was a weak relationship between social skills and social acceptance. more specifically, only peer rated other-oriented social skills were related to social acceptance, but not selfand other-oriented social skills reported by teachers. that is, if pupils with id were rejected by their peers, it was probably not only because of their low level of social skills. thus, there must be other mechanisms influencing the social participation of pupils with id in inclusive classrooms, for example, the achievement level of these pupils. in fact, krull et al. (2014) and nepi et al. (2015) found a relationship between low academic achievement levels and low social acceptance by peers. further, on a group level, classroom composition and group norms can also play a crucial role regarding the social participation of individuals in inclusive classrooms (garrote, 2016; grütter et al., 2015). but, additional studies are required to support these findings. the findings of this study are a contribution to current knowledge on the social participation and the social skills of pupils with id in inclusive classrooms. nevertheless, several limitations of this study should be mentioned. first, a specific concept of social skills was chosen. while such a choice makes the findings more conclusive within a study, comparisons with other studies using different concepts are challenging. second, when interpreting the results, it must be considered that correlations between aspects of social participation and social skills might appear because of overlapping concepts. for example, the aspect of social interactions can be found (in a different function) in relation to the assessment of social skills and of social participation. third, the high correlations between social acceptance and peer-rated other-oriented social skills might result from the assessment method. participants were requested to rate how much they liked to play with their classmates. subsequently, they rated how helpful their classmates were and how well garrote     | f l r     12   they could work with them. if we consider the cognitive process of dissonance (festinger, 1957), it is expected that pupils will rate their classmates consistently or that the two variables will correlate highly. fourth, in this study, groups were compared regarding social skills. the question of which basic social skills pupils with sen require to have friends or be accepted by their peers remains unanswered. fifth, pupils with id were compared to td pupils of extreme status groups: the most popular and the most rejected pupils in the classroom. this was due to the study design, in which detailed data collection was restricted to a number of pupils in the sample. sixth, a dichotomous categorisation of pupils with id and td pupils was used to investigate possible hindering factors in the social participation of pupils with id. although, this categorisation does not reflect disability as a social construct, it is commonly used in practice to allocate special education resources to support individual pupils. in addition, how special educational support is implemented (e.g., in a resource-room) can enhance the perception of pupils with id as an out-group and influence processes of social participation in inclusive classrooms. to shed light into social processes influenced by this dichotomy used in practice further studies are needed. in conclusion, the findings support the assumption that social skills are only one possible explanation why certain pupils with id are more at risk of being less socially involved in their classrooms than their td peers. having friends or being rejected did not seem to depend on the low level of social skills of pupils with id. thus, there must be other factors that have a stronger influence on the social participation of these pupils. possible factors may be identified on the individual level (e.g., having the label “id”). however, group processes should be considered as well. focusing on group processes rather than on individual characteristics has been a promising approach regarding the development and evaluation of interventions to foster social interactions among pupils with and without sen (whalon et al., 2015). indeed, facilitating social participation requires the effort and engagement of all, including peers and teachers (farmer, mcauliffe lines, & hamm, 2011; garrote, sermier dessemontet, & moser opitz, 2017; gest & rodkin, 2011). this change of perspective could also be beneficial for research on social participation. this approach demands from researchers focusing more on variables on classroom or group level rather than solely on individual characteristics. in fact, social participation of individuals in inclusive classrooms is very likely to vary as a function of group and individual factors. keypoints most pupils with id have friends and are accepted by classmates in inclusive classrooms. generally, social skills might play a role in the social participation of pupils but are not the only influencing factor. for pupils with id in primary classrooms, high levels of self-oriented and other-oriented social skills are not necessary to be socially accepted and have friends. acknowledgements this work was supported by the swiss national science foundation [grant number 146086]. references avramidis, e. (2013). self-concept, social position and social participation of pupils with sen in mainstream primary schools. research papers in education, 28(4), 421–442. doi:10.1080/02671522.2012.673006 garrote     | f l r     13   bellini, s., peters, j. k., benner, l., & hopf, a. (2007). a meta-analysis of school-based social skills interventions for children with autism spectrum disorders. remedial and special education, 28(3), 153– 162. doi:10.1177/07419325070280030401 borgatti, s. p., everett, m. g., & freeman, s. f. (2002). ucinet 6 for windows: software for social network analysis. natick: analytic technologies, inc. bossaert, g., colpin, h., pijl, s. j., & petry, k. (2013). truly included? a literature study focusing on the social dimension of inclusion in education. international journal of inclusive education, 17(1), 60–79. doi:10.1080/13603116.2011.580464 camargo, s., rispoli, m., ganz, j., hong, e., davis, h., & mason, r. (2014). a review of the quality of behaviorally-based intervention research to improve social interaction skills of children with asd in inclusive settings. journal of autism and developmental disorders, 44(9), 2096-2116. doi:10.1007/s10803-014-2060-7 estell, d. b., jones, m. h., pearl, r., van acker, r., farmer, t. w., & rodkin, p. c. (2008). peer groups, popularity, and social preference: trajectories of social functioning among students with and without learning disabilities. journal of learning disabilities, 41(1), 5–14. doi:10.1177/0022219407310993 fabes, r. a., martin, c. l., & hanish, l. d. (2009). children's behaviors and interactions with peers. in k. h. rubin, w. m. bukowski, & b. p. laursen (eds.), social, emotional, and personality development in context. handbook of peer interactions, relationships, and groups (pp. 45–62). new york: guilford press. farmer, t. w., mcauliffe lines, m., & hamm, j. v. (2011). revealing the invisible hand: the role of teachers in children's peer experiences. journal of applied developmental psychology, 32(5), 247–256. doi:10.1016/j.appdev.2011.04.006 feldman, r., carter, e. w., asmus, j., & brock, m. e. (2015). presence, proximity, and peer interactions of adolescents with severe disabilities in general education classrooms. exceptional children, 82(2), 192– 208. doi:10.1177/0014402915585481 festinger, l. (1957). a theory of cognitive dissonance. stanford, california: stanford university press. frederickson, n., & turner, j. (2003). utilizing the classroom peer group to adress children's social needs: an evaluation of the circle of friends intervention approach. the journal of special education, 36(4), 234–245. frostad, p., mjaavatn, p. e., & pijl, s. j. (2011). the stability of social relations among adolescents with special educational needs (sen) in regular schools in norway. london review of education, 9(1), 83– 94. doi:10.1080/14748460.2011.550438 frostad, p., & pijl, s. j. (2007). does being friendly help in making friends? the relation between the social position and social skills of pupils with special needs in mainstream education. european journal of special needs education, 22(1), 15–30. doi:10.1080/08856250601082224 fuchs, d., fuchs, l. s., mathes, p. g., & martinez, e. a. (2002). preliminary evidence on the social standing of students with learning disabilities in pals and no–pals classrooms. learning disabilities research & practice, 17(4), 205–215. doi:10.1111/1540-5826.00046 garrote, a. (2016). soziale teilhabe von kindern in inklusiven klassen. empirische pädagogik, 30(1), 67– 80. garrote, a., sermier dessemontet, r., & moser opitz, e. (2017). facilitating the social participation of pupils with special educational needs in mainstream schools: a review of school-based interventions. educational research review, 20, 12–23. doi:10.1016/j.edurev.2016.11.001 gest, s. d., graham-bermann, s. a., & hartup, w. w. (2001). peer experience: common and unique features of number of friendships, social network centrality, and sociometric status. social development, 10(1), 23–40. doi:10.1111/1467-9507.00146 gest, s. d., & rodkin, p. c. (2011). teaching practices and elementary classroom peer ecologies. journal of applied developmental psychology, 32(5), 288–296. doi:10.1016/j.appdev.2011.02.004 gifford-smith, m. e., & brownell, c. a. (2003). childhood peer relationships: social acceptance, friendships, and peer networks. journal of school psychology, 41(4), 235–284. doi:10.1016/s00224405(03)00048-7 garrote     | f l r     14   goldstein, h., english, k., shafer, k., & kaczmarek, l. (1997). interaction among preschoolers with and without disabilities: effects of across-the-day peer intervention. journal of speech, language & hearing research, 40(1), 33–48. retrieved from http://search.ebscohost.com/login.aspx?direct=true&db=c8h&an=1998076502&site=ehost-live gresham, f. m., & macmillan, d. l. (1997). social competence and affective characteristics of students with mild disabilities. review of educational research, 67(4), 377–415. retrieved from http://www.jstor.org/stable/1170514 groeben, m., perren, s., stadelmann, s., & klitzing, k. (2011). emotional symptoms from kindergarten to middle childhood: associations with selfand other-oriented social skills. european child & adolescent psychiatry, 20(1), 3-15. doi:10.1007/s00787-010-0139-z grütter, j., meyer, b., & glenz, a. (2015). sozialer ausschluss in integrationsklassen: ansichtssache? psychologie in erziehung und unterricht, 62(1), 65. doi:10.2378/peu2015.art05d henricsson, l., & rydell, a.-m. (2006). children with behaviour problems: the influence of social competence and social relations on problem stability, school achievement and peer acceptance across the first six years of school. infant and child development, 15(4), 347–366. doi:10.1002/icd.448 huber, c. (2006). soziale integration in der schule?!: eine empirische untersuchung zur sozialen integration von schülern mit sonderpädagogischem förderbedarf im gemeinsamen unterricht. marburg: tectum-verlag. huber, c., & wilbert, j. (2012). soziale ausgrenzung von schülern mit sonderpädagogischem förderbedarf und niedrigen schulleistungen im gemeinsamen unterricht. empirische sonderpädagogik, (2), 147– 165. huisman, m., & steglich, c. (2008). treatment of non-response in longitudinal network studies. social networks, 30(4), 297–308. doi:10.1016/j.socnet.2008.04.004 hymel, s., vaillancourt, t., mcdougall, p., & renshaw, p. (2004). peer acceptance and rejection in childhood. in p. k. smith & c. h. hart (eds.), blackwell handbook of childhood social development. oxford, uk: blackwell publishing ltd. jacques, n., wilton, k., & townsend, m. (1998). cooperative learning and social acceptance of children with mild intellectual disability. journal of intellectual disability research, 42(1), 29–36. doi:10.1046/j.1365-2788.1998.00098.x kamps, d., royer, j., dugan, e., kravits, t., gonzalez-lopez, a., garcia, j., . . . garrison kane, l. (2002). peer training to facilitate social interaction for elementary students with autism and theirs peers. council for exceptional children, 68(2), 173–187. kavale, k. a., & forness, s. r. (1996). social skill deficits and learning disabilities: a meta-analysis. journal of learning disabilities, 29(3), 226–237. doi:10.1177/002221949602900301 koster, m., nakken, h., pijl, s. j., & van houten, e. (2009). being part of the peer group: a literature study focusing on the social dimension of inclusion in education. international journal of inclusive education, 13(2), 117–140. doi:10.1080/13603110701284680 koster, m., pijl, s. j., nakken, h., & van houten, e. (2010). social participation of students with special needs in regular primary education in the netherlands. international journal of disability, development and education, 57(1), 59–75. doi:10.1080/10349120903537905 krull, j., wilbert, j., & hennemann, t. (2014). soziale ausgrenzung von erstklässlerinnen und erstklässlern mit sonderpädagogischem förderbedarf im gemeinsamen unterricht. empirische sonderpädagogik, 6(1), 59–75. malti, t., & perren, s. (eds.). (2016). soziale kompetenz bei kindern und jugendlichen: entwicklungsprozesse und förderungsmöglichkeiten (2nd ed.). stuttgart: kohlhammer. monchy, m. de, pijl, s. j., & zandberg, t. (2004). discrepancies in judging social inclusion and bullying of pupils with behaviour problems. european journal of special needs education, 19(3), 317–330. doi:10.1080/0885625042000262488 murray, c., & greenberg, m. t. (2006). examining the importance of social relationships and social contexts in the lives of children with high-incidence disabilities. the journal of special education, 39(4), 220–233. doi:10.1177/00224669060390040301 garrote     | f l r     15   nepi, l. d., fioravanti, j., nannini, p., & peru, a. (2015). social acceptance and the choosing of favourite classmates: a comparison between students with special educational needs and typically developing students in a context of full inclusion. british journal of special education, n/a. doi:10.1111/14678578.12096 perren, s., argention-groeben, m., stadelmann, s., & klitzing, k. von. (2016). selbstund fremdbezogene soziale kompetenzen: auswirkungen auf das emotionale befinden. in t. malti & s. perren (eds.), soziale kompetenz bei kindern und jugendlichen. entwicklungsprozesse und förderungsmöglichkeiten (2nd ed., pp. 91–110). stuttgart: kohlhammer. perren, s., forrester-knauss, c., & alsaker, f. d. (2012). selfand other-oriented social skills: differential associations with children’s mental health and bullying roles. journal for educational research online / journal für bildungsforschung online; vol 4, no 1 (2012): assessment and development of social competence, (1), 99–123. retrieved from http://www.j-e-r-o.com/index.php/jero/article/view/306 pijl, s. j., & frostad, p. (2010). peer acceptance and self-­‐‑concept of students with disabilities in regular education. european journal of special needs education, 25(1), 93–105. doi:10.1080/08856250903450947 pijl, s. j., frostad, p., & flem, a. (2008). the social position of pupils with special needs in regular schools. scandinavian journal of educational research, 52(4), 387–405. doi:10.1080/00313830802184558 robins, g., pattison, p., & woolcock, j. (2004). missing data in networks: exponential random graph (p∗) models for networks with non-respondents. social networks, 26(3), 257–283. doi:10.1016/j.socnet.2004.05.001 rose-krasnor, l. (1997). the nature of social competence: a theoretical review. social development, 6(1), 111–135. doi:10.1111/j.1467-9507.1997.tb00097.x rotheram-fuller, e., kasari, c., chamberlain, b., & locke, j. (2010). social involvement of children with autism spectrum disorders in elementary school classrooms. journal of child psychology and psychiatry, 51(11), 1227–1234. doi:10.1111/j.1469-7610.2010.02289.x rubin, k. h., bukowski, w. m., & parker, j. g. (2006). peer interactions, relationships, and groups. in n. eisenberg (ed.), handbook of child psychology (6th ed., pp. 571–645). hoboken, n.j: john wiley & sons. ruijs, n. m., & peetsma, t. t. d. (2009). effects of inclusion on students with and without special educational needs reviewed. educational research review, 4(2), 67–79. doi:10.1016/j.edurev.2009.02.002 schwab, s. (2015). social dimensions of inclusion in education of 4th and 7th grade pupils in inclusive and regular classes: outcomes from austria. research in developmental disabilities, 43-44, 72–79. doi:10.1016/j.ridd.2015.06.005 schwab, s., gebhardt, m., krammer, m., & gasteiger-klicpera, b. (2015). linking self-rated social inclusion to social behaviour. an empirical study of students with and without special education needs in secondary schools. european journal of special needs education, 30(1), 1–14. doi:10.1080/08856257.2014.933550 sebanc, a. m. (2003). the friendship features of preschool children: links with prosocial behavior and aggression. social development, 12(2), 249–268. doi:10.1111/1467-9507.00232 united nations (2006). convention on the rights of persons with disabilities and optional protocol. new york: united nations. whalon, k. j., conroy, m. a., martinez, j. r., & werch, b. l. (2015). school-based peer-related social competence interventions for children with autism spectrum disorder: a meta-analysis and descriptive review of single case research design studies. journal of autism and developmental disorders, 45(6), 1513–1531. doi:10.1007/s10803-015-2373-1 codepen vriesema frontline learning research vol.8 no. 3 special issue (2020) 126 139 issn 2295-3159 experience and meaning in small-group contexts: fusing observational and self-report data to capture self and other dynamics christine calderon vriesemaa, & mary mccaslinb auniversity of california, santa barbara, usa buniversity of arizona, usa article received 17 may 2019 / revised 15 november/ accepted 1 january / available online 30 march abstract self-report data have contributed to a rich understanding of learning and motivation; yet, self-report measures present challenges to researchers studying students’ experiences in small-group contexts. rather than using self-report data alone, we argue that fusing self-report and observational data can yield a broader understanding of students’ small-group dynamics. we provide evidence for this assertion by presenting mixed-methods findings in three sections: (a) self-report data alone, (b) observational data alone, and (c) the fusion of both data sources. we rely on 101 students’ self-reported experiences as well as observational (i.e., audio) data of students working in their group (n = 24 groups). in section order, we found that (1) students’ self-reported small-group behavior predicted their end-of-study reported anxiety and emotion; (2) coded observational data captured five types of group dynamics that students can engage in; and (3) students’ initial group-level characteristics predicted their real-time group dynamics, and observed group regulation activity predicted students’ self-reported anxiety, emotion, and regulation moving forward. thus, while self-report and observational data alone can each increase our understanding of student motivation and learning processes, pursuing both in tandem more effectively captures the give-and-take among students, how these experiences evolve over time, and the personal meanings they can afford. keywords: self-report; observation; small-group dynamics; motivation; co-regulation info corresponding author email: vriesecn@uwec.edu doi: https://doi.org/10.14786/flr.v8i3.493 1. introduction instruments measuring students’ motivation and learning processes have contributed to a rich understanding of students’ experiences in school. yet, the extent to which self-report measures adequately capture the learning process for all students across varying contexts remains an important concern (e.g., urdan & bruchmann, 2018). for researchers studying specific instructional contexts, self-report data can pose challenges to investigating motivation and strategy use in small groups. namely, self-report measures make it difficult to investigate how students’ reported behavior and emotion occur in real-time and in relation to other people in their immediate environments. when students work together, each person brings unique experiences and characteristics into their small groups. how identity, disposition, motivation, and readiness to learn impact group functioning—and how individuals are impacted by their interactions with others over time—reflects a dynamic process that self-report data alone cannot capture. to better understand this complex learning environment, we pursued a longitudinal, mixed-methods study of 101 students’ small-group experiences during six math lessons (n = 24 groups from two third-grade and two fifth-grade classrooms). students completed self-report measures at pretest and at posttest. throughout the study, students also completed an instrument describing their individual small-group behaviors after each lesson. finally, after completing the study, students responded to items asking them how they would feel if their teacher asked them to get into small groups again. in addition to these self-report data, our project included real-time audio data of students working in their small groups. selected results of this study were briefly discussed as part of a larger chapter focusing on the guiding theoretical perspective (mccaslin & vriesema, 2018); we present the full study here for the first time. the present special issue aims to better understand the impact of self-report data on theory and practice (dinsmore & fryer, 2020). we contribute to this goal by specifically addressing two of the three guiding questions: how does the use of self-report constrain the analytical choices made with self-report data, and how do the interpretations of self-report data influence interpretations of findings? we situate both questions within the context of small-group research. we begin by briefly introducing the guiding theory. we then present our study’s findings in three sections depicting what we learn from self-report data alone, observational data alone, and integrating both data sources. 1.1 theory the co-regulation model (mccaslin, 2009) that guides this research is a motivation perspective positing that learners are social, have a basic need for participation and validation (mccaslin & burross, 2008), and differ in how and what they participate (mccaslin et al., 2016). influenced by vygotskian tenets, this theory describes how three sources of influence function together to inform emergent identity. these sources are cultural (e.g., norms, challenges), social (e.g., relationships, opportunities), and personal (e.g., readiness to learn, disposition). students bring their personal backgrounds and characteristic adaptations to the classroom; yet, the opportunities presented to students and the relationships formed throughout their schooling experiences can shape who students become. given the dynamic processes described within this theoretical perspective, small-learning groups present an opportune setting to study emergent identity. students in small groups each bring varying achievement levels, dispositions, and motivation to the task. however, the nature of the small-group instructional setting requires that students work together toward a common goal and negotiate challenges when necessary. when students work with each other across multiple occasions, small groups provide an opportunity to understand how student identity informs their work with other classmates and how these shared classroom experiences can shape student identity moving forward. some scholars have hailed small-group learning formats as the success story of educational psychology (johnson & johnson, 2009). small-group activities can enhance student thinking and learning of both formal (e.g., math) and informal (e.g., appropriate social skills, motivated student engagement) content and skills (e.g., elias & schwab, 2006; hadwin et al., 2018; webb, 2008). however, while small-group learning has demonstrated benefits, there also are concerns that not all small-group activities are beneficial nor do all group members experience them similarly (rogat et al., 2013; webb, 2013). naturalistic observational studies examining the processes that actually occur within small groups and what students make of them are relatively scarce. extant research, however, suggests their importance (e.g., hadwin & järvelä, 2011; tan et al., 2005; webb, 2013). therefore, the dynamic processes occurring within small-group settings necessitate dynamic methodologies to study them. asking students about their experiences in small groups can yield important information regarding students’ interpretations of events; and, researchers can investigate how students’ personal characteristics associate with these self-report data. however, self-report data alone cannot capture the give-and-take of small-group interactions. yet, observational data alone also can fail to capture the full student experience. in the case of observation-only data, researchers rely on their own interpretations of events and fail to capture students’ own self-reported experiences of the events. thus, combining self-report and observational data provides a foundation for more fully understanding how individual characteristics inform small-group dynamics, and how these dynamics inform student identity moving forward. to illustrate these points in finer detail, we present three sections that discuss (a) self-report data, (b) observational data, and (c) the fusion of both data sources in our research. 2. section 1: self-report data this section relies on self-report data to illustrate how students’ reported small-group behavior associated with their characteristics at pretest and posttest. first, we describe how students’ pretest characteristics—their teacher-ranked math readiness, self-reported anxiety and emotional adaptation (i.e., context-dependent emotion and coping strategies; mccaslin et al., 2016)—associated with their self-reported small-group behavior. second, we show how self-reported group behavior predicted students’ posttest anxiety, emotional adaptation, and reports of how they would feel if their teachers asked them to get into small groups again. to contextualize these results, we first describe the relevant method information. 2.1 procedure students (n = 101) completed the pretest (october) and posttest (january) surveys that measured their anxiety and emotional adaptation. teachers also ranked each of their students on mathematics achievement at pretest. at the end of the study, students completed an instrument asking them how they would feel if their teacher asked them to get into groups again. throughout the study, students also completed short instruments immediately after each small group lesson to indicate their behavior during the lesson. we present students’ average reported behavior (i.e., the average across the six lessons) below in order to enhance clarity of the results. 2.2 data sources 2.2.1 what school is like for me (wslm) wslm is the test anxiety scale, a well-known, well-researched, and well-critiqued instrument (pekrun, 2006; zeidner & matthews, 2005) adapted from sarason et al., 1958). wslm asks students to agree or disagree with 18 sentences describing anxious thoughts and feelings. cronbach’s alpha for wslm was α = .76 at pretest and .71 at posttest. 2.2.2 school situations (ss) school situations (ss; burggraf, 1993) is an adaptation of the test for self-conscious affect (tosca), a dispositional measure originally designed for adults and subsequently revised by tangney and colleagues to include children (e.g., tangney et al., 1995). the ss inventory asks students to use a five-point scale to endorse sentences in response to 12 written vignettes that portray routine school challenges within three contexts: whole class, small group, or private/individual. sentences are behavioral representations of emotions (guilt, shame, or pride) and coping strategies (externalize, normalize). rather than consider the five ss scales (pride, guilt, shame, normalize, externalize) independently, as originally designed, we used five unique emotional adaptation profiles identified in previous research (mccaslin et al., 2016) for our analyses. the five profiles were: (1) distance and displace: the student attempts to withdraw from a difficult situation to care for the self and/or attempts to blame other people or things to find relief from feelings of shame; (2) regret and repair: the student attempts to repair or fix the situation and to care for the self through normalizing the event in order to find relief from feelings of guilt; (3) inadequate and exposed: the student assumes responsibility and blame for mistakes or difficulties without engaging in self-care or displacement strategies in response to negative emotion; (4) proud and modest: the student acknowledges success, but tempers feelings of pride with humility; and (5) minimize and move on: the student adopts a ‘just keep going, do not dwell, look beyond it’ escape response to mistakes and difficult situations. at pretest, cronbach’s alpha was .75, .79, .70, .72, and .64 for distance and displace, regret and repair, inadequate and exposed, proud and modest, and minimize and move on, respectively. in the same order, internal consistency reliability at posttest was .75, .87, .70, .75, and .66, respectively. 2.2.3 how i was in group today (how i was) how i was presented 20 sentences to students and asked them to underline any that described their behavior in their group that day. sentences comprised three scales (mccaslin et al., 1994): (1) enhancing: sentences that represent engagement from which other group members may benefit; (2) neutral: sentences that represent participation that is neither active nor withdrawn; and (3) interfering: sentences that describe preoccupation with concerns of the self. the interfering scale consisted of items suggesting that students withdrew from or were unable to participate in small group activity (e.g., “my stomach felt funny”; “my head hurt”) rather than engaging in behaviors that actively distracted or interfered with others in small group. therefore, we subsequently refer to this scale as “withdrawn” to clarify this distinction. 2.2.4 how i felt how i felt was designed to capture students’ thoughts and feelings when the teacher said it was time to get into their small group. it consisted of six items that described positive and negative emotional experiences in three relative domains: cognitive, affective, and physiological. interested (cognitive), happy (affective mood), and relaxed (physiological) comprised the “positive” scale (α = .85). confused (cognitive), sad (affective mood), and nervous (physiological) comprised the “negative” scale (α = .82). students used a 3-point scale (not at all, a little bit, a lot) to respond to each item. 2.3 results 2.3.1 pretest student characteristics and self-reported small-group behavior students’ pretest anxiety and emotional adaptation did not associate with students’ self-reported small-group behavior. however, students with higher initial math readiness reported greater use of neutral regulation strategies, such as listening, during their small groups (r = -.27, p = .007; higher numbers indicate lower rank in math readiness). 2.3.2 self-reported small-group behavior and posttest student characteristics we pursued a series of multiple regression analyses that controlled for students’ reported pretest anxiety and pretest emotional adaptation. we did not control for group membership (e.g., using fixed effects models) because we believed that this might yield decontextualized results. in this paper, we focused on exploring how group processes shaped individual processes and vice versa; thus, we did not account for group membership in order to work toward this goal. however, we did attempt to cluster errors at the group level in our regression analyses in order to account for the shared variance within groups. unfortunately, we did not have a sufficient number of participants for the number of groups in our study to run this analysis effectively. as a result, we proceeded to use traditional multiple regression analyses here and subsequently in the paper. results indicated that students’ self-reported behavior in their small-groups predicted students’ posttest anxiety (f(9, 72) = 4.16, p < .001; r2 = .34, adjusted r2 = .26), as well as two emotional adaptation profiles: regret and repair (f(9, 71) = 5.22, p < .001; r2 = .40, adjusted r2 = .32) and inadequate and exposed (f(9, 71) = 3.00, p = .004, r2 = .28, adjusted r2 = .18). specifically, reported use of enhancing regulation during small group predicted less anxiety at posttest (β = -2.32, p = .035). use of withdrawn regulation also predicted lower endorsement of regret and repair and inadequate and exposed emotional adaptation at posttest (β = -0.25, p = .01; β = -0.22, p = .051, respectively). 2.3.3 self-reported small-group behavior and posttest anticipated affect we pursued a series of multiple regression analyses that controlled for students’ pretest anxiety and emotional adaptation to determine how students’ self-reported behavior during small group predicted their anticipated affect at posttest (i.e., when they imagined the teacher asking them to get into small groups again). students’ self-reported behavior in small groups predicted their endorsement of both positive and negative affect (f(9,75) = 3.00, p < .001, r2 = .32, adjusted r2 = .24; f(9,75) = 3.45, p = .001, r2 = .29, adjusted r2 = .21, respectively). reported enhancing behavior during small group predicted greater anticipated positive affect (β = 0.50, p < .001); in contrast, reported withdrawn behavior predicted greater anticipated negative affect (β = 0.42, p < .001). 2.4 constrained analytical choices and interpretations overall, the self-report data indicated how students’ reported small-group behavior associated with their personal characteristics and attitudes at pretest and posttest. specifically, average student-perceived enhancing behavior across the six lessons predicted lower anxiety at posttest and greater positive affect at the end of the study when students imagined getting into small groups again. in contrast, students who described themselves as withdrawn during their small groups felt more negative emotion when they imagined getting into small groups again. student-perceived withdrawn behavior also predicted less endorsement of inadequate and exposed and regret and repair emotional adaptation; thus, while withdrawing from participation might mitigate the potential for experiencing shame in small-group settings, it also prevents students from potentially developing strategies for overcoming interpersonal challenges with peers. although the interpretations of self-report data provided insight into how students’ perceived small-group behavior associated with their personal characteristics and expectations (e.g., affect), there are several important limitations. first, our analyses were constrained by individual-level data. the data allowed us to examine how students’ self-reported behavior associated with their pretest and posttest outcomes; yet, students do not participate in their small groups alone. the constrained data sources prevented a more complete understanding of the give-and-take among students in these settings. second, our interpretations of the data relied purely on student reports. students’ individual interpretations of their classroom activities are vital to understanding their emergent identity; however, finding ways to corroborate self-report data with real-time data can enhance understanding of selfand other-awareness in small-group dynamics. 3. section 2: observational data while section 1 illustrated associations with students’ self-reported small-group behavior, section 2 depicts students’ actual behavior during their small groups. in section 2, we describe the types of co-regulation dynamics that emerged during students’ small-groups lesson and how the dynamics associated with the groups’ average achievement on the small-group tasks; the group is the unit of analysis. we present the observational results after describing the relevant procedures and coding systems. 3.1 procedure three researchers independently analyzed, transcribed, and verified audio data of small-group interactions for three lessons (representing the beginning, middle, and end of the six lessons) for each group (n = 24 groups). the three researchers remained unaware of the larger study. two complementary observation systems were developed for analyzing the audio data. we describe the coding systems below. 3.2 data sources 3.2.1 group behavior checklist (gbc) the first system, the gbc, is a lower-inference observation instrument that captured the range of onand off-task behaviors that students displayed when working with others in small groups. this study used four gbc variable domains: (a) planning, (b) problem solving, (c) help-seeking, and (d) feedback. coding was completed in 30-second intervals. in total, 2,180 intervals were coded with the gbc. the percentage of exact agreement (91%) among coders was calculated on three coding pairs over three lessons. 3.2.2 group environment summary (ges) the second system, the ges, is a higher-inference system that captured students’ interpersonal and affective dynamics and expressed intrapersonal coping strategies. variable domains included group affective climate; giggle/laugh bursts; and types of aggressive, protective, regressive/escape, and somatic expressed coping behaviors. coding was completed in two halves: at the mid-point and end of each lesson. the percentage of exact agreement was 73% among three coding pairs over three lessons. see mccaslin and colleagues (2011) for more complete documentation of audio enhancement and transcription procedures; mccaslin and vega (2013) for coding system design, procedures, and application in the pilot study; and vega (2014) for implementation decisions for the revised system. 3.2.3 group achievement student activity worksheets completed “by the group” as part of each lesson were scored and verified by two math educators for correctness. percentage correct represented students’ small-group achievement for the lesson material. 3.3 results 3.3.1 group dynamics we represented the gbc and ges data as the percentage of intervals in which a behavior occurred. we then subjected the data from both observation systems to a principal components analysis using varimax rotation in order to develop an understanding of overall group dynamics from our discrete coding categories. results yielded five independent factors that collectively accounted for 61.92% of the variance in student small-group interactive behavior. factors, in order of magnitude, were: conflict and control, working together, resource drain, edgy compliance, and scuffle and confusion (see table 1 for example behaviors and percentage variance explained by factor). we consider these five distinct co-regulation dynamics that students can engage in while in small groups. table 2 real-time co-regulation dynamics note. only the top three positively loaded items for each factor are listed in the table. total number of items varied by factor: n = 18 for factor 1; n = 11 for factor 2; n = 7 for factor 3; n = 9 for factor 4; n = 5 for factor 5. the exploratory factor analysis yielded 8 cross-loaded items: 5 items loaded in the opposite direction, and 3 items loaded in the same direction. the five small-group dynamics factors can be organized into relatively task-focused, other-focused, or the fusion of the two perspectives. in task-focused contexts, student dynamics primarily centered on the academic activity at hand, whereas dynamics in other-focused contexts reflected an emphasis on one’s group members. in addition, we can consider how types of coping behaviors typically associated with individual student behavior—aggressive, protective, and regressive—emerged as characteristics of group co-regulation dynamics. please see figure 1 for a visual representation of how the joint activity varied across the five small-group dynamics. figure 1. small-group regulation foci note. this figure was adapted from mccaslin and vriesema (2018). the working together dynamic fused the demands of task and peers in small group learning within a protective press. group members could ask for assistance, disagree with each other, and offer suggestions and solutions without concern for personal safety. in comparison, an aggressive press encompassed both the relatively task-involved edgy compliance dynamic (in which provocative and aggressive behaviors were related to attempts to meet task demands) and the other-involved conflict and control dynamic (in which aggression and protection behaviors consumed group attention). finally, the scuffle and confusion dynamic in the disorganized pursuit of task demands and the resource drain of needy peers were each marked by regressive, or relatively immature, co-regulation dynamics. taken together, these profiles did not represent particular groups per se; rather, they represented the types of co-regulation dynamics—that is, the types of observed behavior (e.g., communication patterns, coping strategies)—that emerged during the students’ time in small groups. please see table 2 for the means and standard deviations for the co-regulation dynamics. table 3 descriptive statistics for students’ co-regulation dynamics note. means and standard deviations reflect the percentage of time students spent engaging in each of the co-regulation dynamics. 3.3.2 group achievement scuffle and confusion negatively associated with the average percentage correct on group task activities (r = -.46, p = .024). group achievement did not associate with conflict and control (r = -.02, p = .929), working together (r = .09, p = .670), resource drain (r = .08, p = .718), or edgy compliance (r = .21, p = .321). 3.4 constrained analytical choices and interpretations we presented this section depicting observational data alone for two reasons. first, the extensive coding systems provided a framework for understanding the types of dynamics that can emerge in small groups. the researchers observed student behavior that informed behavioral co-regulation ranging from the ‘ideal’ working together dynamics to the aggressive give-and-take between students (e.g., conflict and control) to the disorganized task pursuits that embodied scuffle and confusion. these interpretations, therefore, yielded a broader understanding of students’ systematically observed behavior during their small-group lessons. second, we presented the observational data alone to illustrate that even with real-time data of students working in their small groups, we fail to understand what these dynamics can mean for students’ identity moving forward. relying on observational data alone constrained our analyses to correlations between small-group dynamics and group achievement on the small-group tasks. while this has the important benefit of aligning with the types of data available to teachers when they use small groups in their instruction, we argue that fusing observational and self-report data can provide a more nuanced understanding of students’ experiences in small groups as well as insight into what these classroom experiences can mean for students’ emergent identity. we provide evidence for this argument in the next section. 4. section 3: fusing self-report and observational data rather than constraining self-report data to individual-level analyses or relying solely on group-level observational data, section 3 first illustrates how group-level characteristics associated with real-time group dynamics. specifically, we took the average of group members’ individual characteristics, such as emotional adaptation, to determine how the group’s overall approach to learning tasks associated with their group functioning during the small-group lessons. second, we describe how the give-and-take of small-group dynamics predicted students’ individual characteristics at posttest (i.e., their anxiety, emotional adaptation, and anticipated affect). third, while we have confidence in the reliability of our systematic coding procedures, researcher perceptions of small group dynamics may not coincide with student perceptions of their small-group experiences. therefore, we also present results that depict the alignment between studentand researcher-perceived group behaviors. 4.1 results 4.1.1 pretest group characteristics and real-time group co-regulation dynamics we created group-averaged scores for anxiety, emotional adaptation, and math readiness in order to determine how group composition associated with co-regulation dynamics. anxiety. group-averaged anxiety did not associate with students’ group regulation dynamics. emotional adaptation. group-averaged endorsement of inadequate and exposed positively associated with working together dynamics (r = .47, p = .021). group-averaged endorsement of distance and displace positively associated with engagement in resource drain dynamics (r = .41, p = .048). finally, group-averaged endorsement of proud and modest negatively associated with edgy compliance group dynamics (r = -.42, p = .039). math readiness. groups with a greater concentration of higher-ranked math students were more likely to display working together and edgy compliance dynamics (r = -.21, p = .049; r = -.23, p = .028 respectively). in contrast, groups with a greater concentration of lower-ranked math students were more likely to engage in scuffle and confusion dynamics (r = .37, p = .001). 4.1.2 real-time group co-regulation dynamics and posttest student outcomes we pursued a series of multiple regression analyses controlling for students’ pretest characteristics to determine how real-time group dynamics predicted students’ self-reported anxiety, emotional adaptation, and anticipated affect at posttest. anxiety. small-group dynamics predicted students’ self-reported anxiety, f(11, 66) = 3.97, p < .001; r2 = .40, adjusted r2 = .30. specifically, participating in groups that displayed working together and resource drain co-regulation dynamics predicted lower posttest student anxiety (β = -.27, p =.011; β = -.27, p = .012, respectively). although students may have used different strategies in the two co-regulation dynamics, receiving help from peers in both contexts may have associated with lower anxiety at posttest. emotional adaptation. group dynamics predicted students’ endorsement of regret and repair at posttest, f(11, 65) = 4.09, p < .001; r2 = .41, adjusted r2 = .31. in particular, edgy compliance dynamics predicted greater regret and repair at posttest (β = .21, p = .038). group dynamics did not predict the other four emotional adaptation profiles. anticipated affect. we explored whether small-group dynamics predicted how students would feel if their teachers asked them to get into their small groups again. students’ observed small-group dynamics predicted their self-reported anticipated positive affect (f(11, 69) = 2.08, p < .001; r2 = .25, adjusted r2 = .13). specifically, participating in groups that displayed working together and edgy compliance dynamics predicted greater anticipated positive affect reported by students at end of the study (β = .43, p = .004, and β = .30, p = .041, respectively). small-group dynamics did not predict anticipated negative affect. 4.1.3 alignment between self-report and observational group data to examine the alignment between student and researcher perceptions, we (a) created group-averaged how i was scores for the same lessons for which we had coded data and then (b) examined the associations between self-reported group behavior and observed group behavior. group is the unit of analysis (n = 24). group-averaged reported enhancing behavior positively associated with working together and resource drain co-regulation dynamics (r = .27, p = .010; r = .31, p = .003, respectively). group-averaged reported withdrawn behavior positively associated with both resource drain (r = .24, p = .022) and conflict and control (r = .23, p = .027) dynamics. finally, group-averaged reported withdrawn behavior also negatively associated with the working together dynamic (r = -.29, p = .004). 4.2 interpretations fusing the self-report and real-time observation data provided two main insights into students’ experiences in small groups. first, results indicated that students were aware of themselves within their groups. the self-reported small group behavior aligned with the systematic observation data. for example, self-reported enhancing behavior positively associated with working together co-regulation, while self-reported withdrawn behavior associated with greater conflict and control and less working together co-regulation. further underscoring the alignment between the two data sources, resource drain co-regulation—dynamics in which group members expressed needs (e.g., by asking for materials, attention, etc.)—associated with both self-reported enhancing and withdrawn behavior. in this instance, groups consisted of members who provided help (enhancing) to those who needed and/or wanted it (withdrawn). thus, students as young as grade three appear to accurately self-monitor, and students as old as grade five appear willing to accurately report their small-group behavior. second, results indicated that students’ emotional adaptation and academic readiness were important features of small-group dynamics and personal learning. for example, participation in edgy compliance co-regulation dynamics, unpleasant as it may have been, predicted an increase in students’ posttest endorsement of the regret and repair emotional adaptation profile. this suggested that students were not only aware of the behavior exhibited in their small group but that they learned about themselves and others from that experience. in this instance, interpersonal dynamics informed intrapersonal endorsements that appeared to move the student away from their prior experience toward the person they wanted to become—the person who feels badly when failing to support another and works to make amends. students’ academic readiness also provided evidence for the press between intraand interpersonal dynamics. groups with higher-ranked students, for example, were more likely to display working together and edgy compliance co-regulation dynamics. this suggests that students with higher math readiness have the potential to direct their resources in more productive (e.g., offering suggestions, asking questions) and less productive (e.g., bragging, refusing others’ participation) ways. yet, in spite of the different real-time interaction patterns, the self-report data indicated that students learned from these experiences and that the small groups shaped students’ posttest characteristics. as noted with edgy compliance co-regulation, these dynamics predicted greater endorsement of positive regulation strategies moving forward (regret and repair); and, for the groups already working together, these dynamics predicted lower anxiety at posttest. 5. discussion in this paper, we addressed two of the three guiding questions in this special issue: how does the use of self-report constrain the analytical choices made with self-report data, and how do the interpretations of self-report data influence interpretations of findings? our goal was to illustrate the benefits and challenges of using self-report data to understand students’ experiences in and attitudes towards small groups. overall, using self-report data alone provided insight into students’ self-awareness and perceptions of their own behavior during small group. while useful for capturing the perceived student experience, relying purely on self-report data constrained our inquiry to individual-level analyses in ways that ignored the mutual give-and-take between the individual and their group members. thus, the primary challenge of using self-report data to study small groups is that we fail to capture the dynamic processes that are inherent in these social learning tasks. in other words, self-report data can illustrate how personal sources of influence (e.g., math readiness, emotional adaptation) shape students’ experiences and emergent identities; yet, we fail to also learn how students shape—and are shaped by—the mutual press between personal and social sources of influence in real time. although we emphasized the role of self-report data for understanding students’ experiences in small groups, this paper also identified strengths and limitations of observation-only data. for example, our real-time data corroborated research by ladd and colleagues (2014) in which students reported the (lack of) positive small-group behaviors displayed by their peers. the researchers noted that “substantial proportions of participants received average ratings that were so low…as to imply that they “seldom” or “never” exhibited such skills during collaborative classroom activities” (p. 169). our data provided insight into which behaviors and skills these peers might engage in instead. working together is great when it happens, but the pursuit of joint activity, in which disagreements are respected, questions appreciated, peer elaborations valued, and understandings deepened, does not represent the reality of the array of small group dynamics. instead, groups also display aggressive and regressive coping behaviors that result in more or less effective group functioning. yet, while observation-only data yielded a broader understanding of real-time behavior in small groups, we nevertheless failed to capture students’ perceived experiences within these dynamics. rather, by fusing self-report and observational data, we learned that small-group co-regulation dynamics were saturated with social and self-conscious emotions, and the uneven regulation of those emotions often did not proceed smoothly or turn out well. overall, students differed in their typical need to cope with learning difficulty, but coping with lack of control and uncertainty are part of what it means to be in a small group for most. to be in a small group with peers—classmates who vary in their own learning and social skills (ladd et al., 2014; rogat et al., 2013)—can exacerbate or attenuate that reality. thus, like others in this special issue (e.g., rogiers et al., 2020; van halem et al., 2020), we argue that using multiple data sources can yield broader understandings of student behavior in classroom settings. 5.1 considerations the focus on self-report data in this paper and special issue warrants further discussion of survey data in particular. first, some critiques of self-report data question whether participants can provide accurate responses to researchers’ survey items. these concerns broadly reflect literature showing that participants sometimes fail to understand their own motives (nisbett & wilson, 1977) or provide opinions about events that did not happen (bishop et al., 1980). however, the present study provides evidence that students in grades three and five can (and are willing to) provide accurate reports of their time in small groups. students’ self-reported individual group behavior aligned with researcher-observed behaviors taking place during the small-group activities. for researchers, this suggests the utility of using self-report data in research on small-group processes, particularly when these survey measures can be corroborated with other data sources. furthermore, in line with prior recommendations (corno, 2011), teachers may also want to consider using brief surveys to better understand how small-group activities unfolded in their classrooms. of course, while this strategy may help teachers to refine these activities in their classes, future research will also need to determine whether students’ responses vary depending on whether students are reporting co-regulation dynamics to their teachers or to researchers. second, in addition to considering participants’ understanding of their own attitudes, some researchers express concerns about using self-report data due to the surveys themselves. these criticisms acknowledge that there may be aspects of any given survey that can prevent participants from responding as accurately as possible (duckworth & yeager, 2015). in the current paper, the internal consistency reliability coefficients generally provided one source of evidence for using these survey measures in our analyses. however, it is important to acknowledge that one of our five school situations factors, minimize and move on, fell below the recommended .70 for cronbach’s alpha (nunnally, 1978). thus, even though this factor—capturing student escape strategies—was identified in previous research using this same instrument (mccaslin et al., 2016), we encourage researchers to replicate this work to determine whether the five school situations factors emerge in their own samples, or whether some strategies do not translate across all contexts. in spite of the relatively low internal consistency for the minimize and move on factor, we are confident in our self-report measures, again due to the alignment between the survey and observational data in this study. 5.2 conclusion in sum, conceptions of small-group members in terms of their cooperative or regulatory skill set is a start that is likely to sputter without recognition of the fullness of individuals who have personal histories and concerns that make them more and less vulnerable to threat (frijda, 2008) and making threats. this calls for expanding conceptions of small-group cooperative skills and dispositions of individuals to include, for example, considerations of power and influence among group members. our observations of demanding behavior and provocative exchange suggest that students can consider power from a perspective of coercion and control rather than one of support and positive influence (keltner, 2016). students’ personal concerns and heightened perceptions of threat are part of power and influence dynamics. both are better understood within deliberate consideration of conflicts that may underlie and result from them. we do students a disservice when we fail to acknowledge the fullness of the task of working and learning with others. we also miss an opportunity to fully learn from the potential of small-group learning for students’ personal growth and well-being when we fail to use multiple research methods fluidly. thus, while self-report and observational data alone can each increase our understanding of student motivation and learning processes, pursuing both in tandem can yield richer understandings of students’ classroom activity, how these experiences evolve over time, and how that matters in the dynamics of being and becoming a student. keypoints students’ self-reported behavior during small group predicted their reported end-of-study anxiety and anticipated emotion. real-time audio data indicated five distinct types of co-regulation dynamics that students can engage in within small groups (e.g., communication patterns, coping, etc.). students’ initial group-level characteristics predicted their real-time co-regulation dynamics. co-regulation dynamics during small group predicted individual students’ self-reported end-of-study anxiety, anticipated emotion, and emotional adaptation. the real-time audio data corroborated students’ self-reported behavior during small group. references bishop, g. f., oldendick, r. w., tuchfarber, a. j., & bennett, s. e. (1980). pseudo-opinions on public affairs. the public opinion quarterly, 44(2), 198-209. burggraf, s. a. (1993). school situations. unpublished manuscript. bryn mawr, pa: bryn mawr college. corno, l. (2011). studying self-regulation habits. in h. d. schunk, & b. zimmerman (eds.), handbook of self-regulation of learning and performance (pp. 361-375). new york: routledge. duckworth, a. l., & yeager, d. s. (2015). measurement matters: assessing personal qualities other than cognitive ability for education purposes. educational researcher, 44(4), 237-251. https://doi.org/10.3102/0013189x15584327 elias, m. j., & schwab, y. (2006). from compliance to responsibility: social and emotional learning and classroom management. in c. m. evertson & c. s. weinstein (eds.), handbook of classroom management: research, practice, and contemporary issues (pp. 309-341). mahwah, nj: lawrence erlbaum associates. frijda, n. h. (2008). the psychologists’ point of view. in m. lewis, j. m., haviland-jones, & l. f. barrett (eds.), handbook of emotions, 3rd ed. (pp. 68-87). new york: guilford press. fryer, l. k., & dinsmore, d. l. (2020). the promise and pitfalls of self-report: development, research design and analysis issues, and multiple methods. frontline learning research, 8(3), 1–9. http://doi.org/10.14786/flr.v8i3.623 hadwin, a. f., & järvelä, s. (2011). introduction to a special issue on social aspects of self-regulated learning: where social and self meet in the strategic regulation of learning. teachers college record, 113(2), 235-239. hadwin, a. f., järvelä, s., & miller, m. (2018). self-regulation, co-regulation, and shared regulation in collaborative learning environments. in d. h. schunk & j. a. greene (eds.), handbook of self-regulation of learning and performance (pp. 83-06). new york, ny: routledge. johnson, d. w., & johnson, r. t. (2009). an educational psychology success story: social interdependence theory and cooperative learning. educational researcher, 38, 365-379. https://doi.org/10.3102/0013189x09339057 keltner, d. (2016). the power of paradox: how we gain and lose influence. new york, ny: penguin press. ladd, g. w., kochenderfer-ladd, b., visconti, k. j., ettekal, i. sechler, c. m., & cortes, k. i. (2014). grade-school children’s social collaborative skills: links with partner preference and achievement. american educational research journal, 51(1), 152-183. https://doi.org/10.3102/0002831213507327 mccaslin, m. (2009). co-regulation of student motivation and emergent identity. educational psychologist, 44(2), 137-146. https:// doi.org/10.1080/00461520902832384 mccaslin, m., & burross, h. (2008). student motivational dynamics. teachers college record, 110(11), 2319-2340. mccaslin, m., tuck, d., waird, a., brown, b., lapage, j., & pyle, j. (1994). gender composition and small-group learning in fourth-grade mathematics. elementary school journal, 94, 467-482. mccaslin, m., & vega, r. i. (2013). peer co-regulated learning, emotion, and coping in small-group learning. in s. phillipson, k. y. l. ku, s. n. phillipson (eds.), constructing educational achievement: a sociocultural perspective (pp. 118-135). new york, ny: routledge. mccaslin, m., & vriesema, c. c. (2018). co-regulation: a model for classroom research in a vygotskian perspective. in d. m. mcinerney & g. a. d. liem (eds.), big theories revisited 2: research on sociocultural influences on motivation and learning. charlotte, nc: information age publishing. mccaslin, m., vega, r. i., anderson, e. e., calderon, c. n., labistre, a. m. (2011). tabletalk: navigating and negotiating in small-group learning. in d. mcinerney, r. walker, g. liem (eds.), sociocultural theories of learning and motivation: looking back, looking forward (pp. 191-222). charlotte, nc: information age publishing. mccaslin, m., vriesema, c. c., & burggraf, s. (2016). making mistakes: emotional adaptation and classroom learning. teachers college record, 118(2). nisbett, r. e., & wilson, t. d. (1977). telling more than we can know: verbal reports on mental processes. psychological review, 84(3), 231-259. https://doi.org/10.1037/0033-295x.84.3.231 nunnally, j. c. (1968). psychometric theory (2nd edition). new york: mcgraw-hill. pekrun, r. (2006). the control-value theory of achievement emotions: assumptions, corollaries, and implications for educational research and practice. educational psychology review, 18, 315-341. https://doi.org/10.1007/s10648-006-9029-9 rogat, t. k., linnenbrink-garcia, l., & didonato, n. (2013). motivation in collaborative groups. in c. e. hmelo-silver, c. a. chinn, c. k. k. chan, & a. m. o’donnell (eds.), the international handbook of collaborative learning (pp. 250-267). new york: taylor & francis. rogiers, a.; merchie, e., & van keer, h. (2020). opening the black box of students’ text-learning processes: a process mining perspective. frontline learning research, 8(3), 40–62. http://doi.org/10.14786/flr.v8i3.527 sarason, s. b., davidson, k. s., lighthall, f. f., & waite, r. r. (1958). a test anxiety scale for children. child development, 29(1), 105-113. tan, i. g. c., sharan, s., & lee, c. k. e. (2007). group investigation effects on achievement, motivation, and perceptions of students in singapore. the journal of educational research, 100(3), 142-154. https://doi.org/10.3200/joer.100.3.142-154 tangney, j. p., burggraf, s. a., & wagner, p. a. (1995). shame-proneness, guilt-proneness, and psychological symptoms. in j. p. tangney & k. w. fischer (eds.) self-conscious emotions: the psychology of shame, guilt, embarrassment, and pride (pp. 343-367). ny: guilford press. urdan, t., & bruchmann, k. (2018). examining the academic motivation of a diverse student population: a consideration of methodology. educational psychologist, 53(2), 114-130. https://doi.org/10.1080/00461520.2018.1440234 van halem, n., van klaveren, c. p. b. j., drachsler, h., schmitz, m., & cornelisz, i. (2020). tracking patterns in self-regulated learning using students’ self-reports and online trace data. frontline learning research, 8(3), 142-164. http://doi.org/10.14786/flr.v8i3.497 vega, r. i. (2014). the role of student coping in the socially shared regulation of learning in small groups. unpublished doctoral dissertation. tucson, az: university of arizona. webb, n. m. (2008). learning in small groups. in t. l. good (ed.), 21st century education: a reference handbook (vol. 2., pp. 203–211). thousand oaks, ca: macmillan. webb, n. m. (2013). information processing approaches to collaborative learning. in c. e. hmelo-silver, c. a. chinn, c. k. k. chan, & a. m. o’donnell. (eds.), the international handbook of collaborative learning (pp. 19-40). new york: taylor & francis. vekkaila et al publication frontline learning research vol.7 no. 1 (2019) 51 64 issn 2295-3159 how do doctoral students in stem fields engage in scientific knowledge practices? jenna vekkailaa, viivi virtanena,jani kukkolab, liezel frick c, kirsi pyhältöa a centre for university teaching and learning (hype), university of helsinki, helsinki, finland b department of philosophy, history and arts, university of helsinki, finland c department of curriculum studies, stellenbosch university, stellenbosch, south africa. article received 1 august 2018/ revised 25 november / accepted 31 january / available online 22 february abstract knowledge creation is at the core of scientific endeavour. as early career researchers, doctoral students take part in knowledge creation through engaging in various knowledge practices and make their original contribution to knowledge, and become experts in their particular domain. however, our understanding of what doctoral knowledge practices entails is still insufficient. for this study, a total of 34 doctoral students from stem fields, including natural sciences, bioand environmental sciences and medicine were interviewed to gain a better understanding of the kinds of knowledge practices in which doctoral students in the sciences engage. the data were collected with semi-structured interviews, which were qualitatively content analysed. the results showed that the participants mostly described activities that were established everyday knowledge practices of the researcher community (75 %), whereas practices that were innovative (25 %), entailing transformation of the current practices and developing new ones, were less often reported. moreover, the practices were typically collective, involving the students, their supervisors or other members of their research groups (67 %). further investigation showed that the participants were typically actively engaged in knowledge practices (79 %) rather than just adapting existing ones (13 %). perceiving oneself as a bystander was even less typical (8 %). the significance of this study lies in exploring doctoral students’ self-reported knowledge practices in stem fields, and demonstrates that they perceive themselves as actively and collaboratively engaged in creating knowledge. keywords: doctoral training; doctoral student; qualitative study; knowledge practice; stem fields info corresponding author email: jenna.vekkaila@helsinki.fi doi 10.14786/flr.v7i1.393 1. introduction knowledge creation is at the core of scientific endeavour. doctoral students are key players in knowledge creation within any university or discipline since they contribute to the endeavour by producing an original contribution in the form of doctoral dissertation, and by extending the knowledge boundaries of a particular discipline (see e.g., the united kingdom quality assurance agency for higher education, 2008). therefore, they should also be a key interest to both universities and disciplinary communities that stand to benefit from such advances in knowledge. knowledge creation takes place through knowledge practices, entailing various disciplinary research activities such as data collection, analysis, article writing, elaboration of concepts and theories, planning a research project, and presenting research. in stem fields (the abbreviation stem referring to science, engineering, technology and mathematics will be used in this article) such practices are suggested to be typically collective (hakkarainen et al., 2014): doctoral research in stem fields is typically focused on solving shared research problems related to a supervisor’s research projects, pursuing article-based dissertations that consist of co-authored internationally refereed journal articles, and working intensively in relatively strong researcher communities, including several doctoral students, postdocs, and academic staff. yet, not all the researcher communities in the stem fields embrace collective knowledge practices, nor do all doctoral students have similar access to such practices even if they may exist in their communities. accordingly, in order to create an optimal learning environment for knowledge creation for doctoral students in stem fields, a better understanding of the knowledge practices, and ways in which the students engage in these practices during their studies, is needed. the study aims to contribute to bridging the gap in the literature in the field by exploring the kinds of knowledge practices in which doctoral students in stem fields engage during their studies. the knowledge practices are explored in the framework of socio-constructivist views of learning (see e.g. sfard, 1998; paavola, lipponen, & hakkarainen, 2004) by drawing on the seminal work on” knowledge building” by nonaka and taceukhi (1995), engeström (1999), and bereiter (2002). 1.1 knowledge practices as key for knowledge creation knowledge creation is a socially embedded endeavour (john-steiner, 2000), rooted in the researcher community typically comprising of supervisors, other senior scholars, post-doctoral researchers, doctoral students, and both national and international researcher networks (mcalpine & norton, 2006; pyhältö & keskinen, 2012) sharing the same object of activity and knowledge artefacts such as research interest, frameworks, and/or methods. this has several consequences. the knowledge creation takes place in the researcher communities via knowledge practices, which are the socially created ways in which scientists think, interact, and engage in their day-to-day work (brew et al., 2011; mcalpine & åkerlind, 2010) while carrying out research enquiries. such practices entail, for instance, various methods employed in research, frameworks utilised, research designs carried out, and scientific writing genres applied. as a result, doctoral student learning is highly embedded in the knowledge practices, not only in terms of knowledge acquisition (a mental process of individual learning) and knowledge participation (a process of being socialised into an epistemic community), but also in terms of the deliberate process of creating new knowledge that has the potential to transform the student’s ways thinking and behaving (hakkarainen et al., 2004, 2013). prior empirical research on doctoral research knowledge practices is very limited. few prior studies indicate that doctoral students do engage to different extents in different kinds of knowledge practices, and that differences between the researcher communities in this regard occur. hakkarainen and his colleagues (2013), for instance, showed that doctoral students in cutting edge research groups in medicine and in natural sciences were most typically engaged in collective inquiry efforts. in a more recent study (2014) on leaders of national centers of excellence in the sciences, it was shown that professors aimed at cultivating the pursuit of collectively shared research objects, the pursuit of externally reviewed co-authored journal articles, and were focused on collective supervision (hakkarainen et al., 2014). the findings imply that in the best, cutting-edge research communities, the aim of such communities is often to deliberately involve doctoral students in their collective knowledge practices – including the co-construction of goals, reciprocal monitoring and planning of research, and the shared regulation of joint cognitive processes in complex problem-solving (e.g., hadwin & oshige, 2011; volet, vauras, & salonen, 2009), co-authoring, hard work and intensive training – to become members of the research communities (florence & yore, 2004; kamler 2008; hakkarainen et al. 2013). through sustained engagement, new doctoral students are gradually socialized into the knowledge practices that at its best allow them to work at the frontiers of knowledge and transform their ways of thinking and behaving (holmes, 2004). a great deal of this kind of learning takes place through horizontal (between-peer) (see fenge, 2012) and vertical (between newcomers and senior researchers) knowledge sharing. engaging in the cutting-edge knowledge practices allows doctoral students’ co-evolvement and co-development along with their research problems and co-authored articles, and eventually ‘authoring themselves’ as full-members of top researcher communities (holland et al., 1998). however, knowledge practices should not be seen as a singular construct. accordingly, at least distinction between the established knowledge practices (commonly known in the community that everyone needs to master) and innovative practices (that are typically novel or recently transformed), can be made (hakkarainen et al., 2004). moreover, the practices may vary from individual to collective, and from routine practices related to supporting knowledge building to more fluid and innovative practices, which foster the solving of emergent and novel problems (e.g., hakkarainen et al., 2013) that mediate progress towards new scientific discoveries. the practices can also be more or less explicit and intentional. established knowledge practices are more often tacit, since they are well mastered by the established members of the researcher community than in the case of newly developed innovative practices that often still require extra effort to maintain. in addition, the practices are not static in nature; instead, they constantly and intentionally evolve and change in the interplay between individuals and their communities (lave & wenger, 1991). learning about these practices and how to participate in them is essential for becoming a scientist (e.g., becher & trowler, 2001). the knowledge practices are to a certain extent context dependent. in stem fields, solving complex problems through laboratory or fieldwork often requires intensive group-based collaborative knowledge practices (cumming, 2009; delamont & atkinson, 2001) and expertise is distributed among the various researchers at different career phases. this is especially typical in large-scale research projects with many staff members and where a variety of research instruments are utilised (e.g., furner, 2003). moreover, researcher groups often develop their own set of distinctive knowledge practices that evolve over the time. the knowledge practices of the researcher community determine to a great extent not only the quality of their research outputs, but also what the doctoral students learn during their studies, and the overall quality of the doctoral experience. in addition, individual variation between the students in how they engage in these practices is likely to occur. accordingly, in order to understand the influence such practices may have on the students, we need to explore what kinds of knowledge practices doctoral students engage in, and how they engage in them. 1.2 doctoral student engagement in knowledge practices doctoral students themselves can engage differently within the knowledge practices provided by the researcher communities (hopwood, 2010; mathieson, 2011). they can, for instance, adopt, adapt, or withdraw from the practices, and their involvement or lack thereof may eventually modify the practices i.e. display agentic behaviours (hopwood, 2010; pyhältö & keskinen, 2012). this includes working with others to expand the “object of activity”, by recognizing the motives and resources of others, interpreting them, and aligning one’s own responses to these interpretations with the responses of others involved while expanding knowledge in terms of the doctoral project (pyhältö & keskinen, 2012). because sense of agency, while internal, is always constructed in a physical, social, and cultural context, the researcher community can either promote or hinder doctoral students’ sense of agency (o’ meara, terosky, & neumann, 2008). variation across the doctoral students in their experienced ability to exercise their agency is based on a variety of personal, social, and organizational resources and demands at hand (o’ meara & campbell, 2011). therefore, an important aspect of developing relational agency is having an opportunity to participate and contribute (greeno, 2006; lipponen & kumpulainen, 2011; pyhältö, pietarinen & soini, 2012) to the knowledge practices of the researcher community (hancock, hughes, & walsh, 2017). this requires creating the kinds of practices in which doctoral students are seen and treated as accountable researchers. however, the students’ active and responsive collaboration with the researcher community that makes it possible to expand understanding and create new knowledge cannot be taken for granted (pyhältö & keskinen, 2012). hence students can display various degrees of agency in the knowledge practices provided by the researcher community ranging from active and interactional agent, to obedient employer, whose task is to learn “the rules of the game” and carry tasks given by the senior members of the researcher community. the ability of doctoral students to participate in knowledge creation is shown to be determined both by individual attributes, such as their motivation, skills, and ability to carry out agentic behaviour (jazvac-martek, chen, & mcalpine, 2011; mcalpine & amundsen, 2009; see also bandura, 2001; hadwin & oshige, 2011), as well as researcher community attributes, such as the way in which doctoral students are introduced to the community, the quality and quantity of supervisory and researcher community support provided, and the nature of practices in the given community (e.g., delamont & atkinson, 2001; gardner, 2007; golde, 2010; jazvac-martek et al., 2011). at its best, from the beginning of their doctoral processes students are involved in the knowledge practices, which are focused on the knowledge objects that enhance both knowledge and associated practices (hakkarainen, et al., 2004; walker et al., 2008). it has been suggested that in order for doctoral students to create new ideas, they first need a foundation for their creative actions – that is, they must master the existing frameworks, their rules and limits (frick & brodin, 2014). thus, engaging in shared and innovative knowledge practices enables doctoral students to surpass their individual limitations and create new ideas (e.g., walker et al., 2008). this further results in changes both in the relationship between the researchers and their working environment, and in shared knowledge objects (e.g., hakkarainen et al., 2004; lave & wenger, 1991). yet, pyhältö & keskinen (2012) found that doctoral students in behavioural sciences, humanities and medicine rarely displayed agentic behaviours within their researcher communities. prior studies imply that the knowledge practices of scientific communities play a central role in the process of learning to become a scientist, yet our understanding of the nature and function of knowledge practices, especially among stem field doctoral students, is insufficient. 2. aim of the study since no research (empirical or non-empirical) exists on doctoral students’ knowledge practices in stem areas, the aim of this study was to gain a better understanding of the kinds of knowledge practices in which doctoral students in stem fields engage during their doctoral process. in order to reach the aim, the study addressed two complementary research questions; firstly, the kinds of knowledge practices reported by the doctoral students were identified, and secondly, the ways in which students participated into these practices were explored. 3. methods 3.1 finnish doctoral education in stem fields finnish doctoral education in the sciences (pyhältö, stubb, & tuomainen, 2011) is based on the european model. conducting doctoral thesis research is embedded in the activities of the research community. the doctorate involves a dissertation and its public defence. it is complemented with coursework (total 60-80 ects) that is based on personal study plans, typically including international conferences and some methodological studies. doctoral education in finland is outlined in more detail by pyhältö, nummenmaa, soini, stubb, and lonka (2012). our study includes science, medicine, and bioand environmental sciences. considering academic research they all can be viewed as natural sciences in which research is based on empirical evidence from observation and experimentation with mathematics as crucial partner. in finland physics, chemistry, and biology are the contents of entrance examination into studying medicine, and in the research-intensive university the master students from biology often do their doctorates in medicine. hence, we use in this paper the abbreviation stem with medicine included. in finnish science, medicine, and bioand environmental sciences, the most common type of doctoral thesis is a summary of articles. each doctoral student is required to publish from three to five articles in peer-reviewed international journals. the articles are often co-authored with the supervisors. doctoral students in these fields usually work on their phds full time, and the typical completion time varies from four to six years. the key distributives of doctoral education in the faculties of bioand environmental sciences, medicine and science are reported in table 1. there are some differences between the faculties. most doctoral students conduct their work alone in science whereas in medicine and bioand environmental sciences the work is usually conducted in the research group. further, the science students are less often engaged in the doctoral programs compared to their colleagues in the other two faculties. yet the graduation time is shorter in science and medicine than in bioand environmental sciences. the original survey data for the analysis was collected in 2011 with a broad range of disciplines included (pyhältö, stubb, & tuomainen, 2011), and in light of that, the three faculties share a quite similar system of doctoral education. however, particularly the differences in research group status may have effect on the knowledge practices identified from the doctoral students’ interviews. table 1 the structure of doctoral education in the faculties under study: doctoral students’ (n) membership of doctoral program and research group, typical form of conducting thesis, and typical graduation time (see original data; pyhältö, stubb, & tuomainen, 2011). 3.2 participants a total of 34 doctoral students from stem fields (7 participants from the natural sciences, 7 participants from medicine, and 20 participants from the bioand environmental sciences) participated in the study. they were all conducting their research and theses at a large research-intensive finnish university. all the participants had a master’s degree; most of the participants (n=29) were full-time doctoral students and five were part-time. all the participants were pursuing a summary of articles, but they were in different phases of their doctoral process: five were in the beginning of the doctoral process, meaning that they were typically launching their research projects, collecting or analysing data, or writing their first or second article. nine of the participants were in the middle part of the process, which typically included data analysis and writing a third or fourth article. most of the participants (n=16) were in the last part of the process, which typically meant finalizing the last articles and the summary of the articles. four participants had already defended their doctoral theses. all the participants were interviewed on a voluntary basis. 3.3 data collection data were collected by employing semi-structured interviews (e.g., kvale, 2007). the interview protocol was designed to investigate the doctoral students’ experiences of their thesis processes and their views of themselves within these processes (stubb, pyhältö, & lonka, 2014). all interviews were conducted by members of the authors’ research group. the interviews lasted from 22 minutes to almost three hours. the interviews were recorded and transcribed verbatim. 3.4 analysis the interview data were qualitatively content analysed (e.g., creswell, 2012) by relying on an abductive strategy (e.g., morgan, 2007). hence, when categorising the data, observations and prior understanding based on theories were repeatedly assessed in relation to each other by combining data-grounded (harry, sturges, & klingner, 2005; mills, bonner, & francis, 2006) and theory-guided analysis strategies (creswell, 2012) in order to acquire the most accurate possible understanding of doctoral students’ experiences of knowledge practices. the analysis included four complementary phases. at first, all text segments related to knowledge practices were identified. these included all doctoral students’ expressions of conducting research work alone or together with other researchers. the criteria for the text segments, which where coded as experiences of knowledge practices, were that they involved a description of research activities and the object of activity (e.g., data collection, analysis, article writing, elaboration of concepts and theories, planning a research project, presenting research). the analysis resulted in 192 text segments from 34 interviews that were included in the further analysis. the units ranged from a couple of sentences to a dozen sentences. secondly, the knowledge practices identified in the first phase were coded according to the quality of the practices into two exclusive categories by applying a model proposed by hakkarainen et al. (2004): 1): established practices: including text segments in which the practices are reported to be commonly known in the community, or practices that everyone needs to adapt; and 2) innovative practices: including text segments in which the practices are reported to be modified from the existing or new practices. thirdly, the knowledge practices were further categorised into two categories based on whether they were described as individual or collective. the analysis yielded two categories: 1) individual knowledge practices, consisting of descriptions of working alone with the research; and 2) collective knowledge practices, consisting of reports of community-based activities in which two or more researchers are involved. at the fourth analytical phase, all the knowledge practices were coded further into three categories according to how the doctoral students described their roles in the practices: 1) active, containing expressions of being an intentional participant who can affect the activities; 2) adaptation, containing descriptions of being a passive participant in the practice or doing activities that someone else has ordered them to do; and 3) bystander, containing reports of not being involved in or having an unorganized perception of one’s role in the practice. the analysis process was conducted by the first, the second and the third authors. the categories derived from the analysis were critically assessed by the research group at the end of each analysis phase in order to enhance the trustworthiness and credibility of the analysis and results (e.g., miles & huberman, 1994). in the few cases of disagreement, a consensus of final categorization was reached through discussion amongst the researchers. to increase the reliability of our analysis parallel coding was carried out with 67% of the data (total of 129 text segments) independently by two co-authors. the inter-rater reliability for each of the analysis phase were: the agreement range was 100 % (first phase); 81 % (second phase), 95 % (third phase) to 74 % (fourth phase). the few cases of disagreement (particularly phase 2 and 4) we relied on coding of the co-author who had background in the stem-field research, since we presumed that she was more familiar with the knowledge practices of the stem fields. in the findings section, we provide direct quotations from the participants’ descriptions, translated from finnish to english. the quotations were selected to illustrate the particular category as well as to highlight the differences between the categories. for each category, there were several potential illustrative quotations from each discipline available. the most comprehensive quotations were chosen from each category while keeping at the same time track that all disciplines were equally represented. 4. results the doctoral students described a variety of knowledge practices (f=192). the practices ranged from individual work with research instruments to dialogues about theories and observation, as well as shared problem solving and making new discoveries. the reported practices also differed in terms of how established or innovative they were, as described by the participants. furthermore, the participants described their role in the reported practices in varying ways. 4.1 established and innovative knowledge practices the majority of all the knowledge practices reported were established everyday practices cultivated by the researcher community (75 %). such practices typically involved mastering and using research instruments and methods, defining and planning the research topics and processes. the students also described practices related to scientific writing and publishing. one of the participants described such a practice in the following way: in the beginning, my time was spent grasping the laboratory practices and that sort of stuff. and i do use all of them quite diversely, the different laboratory techniques, i mean. and they are demanding—i didn’t even learn them at first. so, they do require a slow and steady pace to get the hang of them. (medicine 2) occasionally, the students described established practices resulting in a discovery. the established practices resulting the discoveries were often cultivated and sustained by the researcher communities for long time. in these cases, the method itself was established and well-known, but it was used in a way that resulted in originality, as the following excerpt shows: in a way, the techniques i use in that [study] are the ones that have been the practice in many laboratories for a long time, but in many places these are no longer used. still, our group has trusted that this is the way to go… in the end we took a real risk and it turned out to be fruitful, and of course that was really motivating. (bio 2) sometimes the participants reported practices that were innovative, entailing transformation of the current practices and developing new ones. these practices were less typical (25 %) compared to established practices. innovative practices are the ones of utmost interest with knowledge creation at issue. innovative practices typically emerged in situations where established ones did not work or did not provide solutions for the problems faced. accordingly, they were characterised by learning from errors. the innovative practices were either reformed or modified from established ones, or new practices that were just created for solving novel problems. these practices were typically related to developing research ideas and theoretical observations, solving empirical problems and mastering research techniques, as well as getting results and making discoveries. the new ways of doing were often associated with aiming at, or actually constructing, new knowledge: a new research idea, theoretical observation, or scientific discovery. occasionally, the students faced research related problems and found solutions on their own: so i have been testing different techniques as a kind of pioneer work, as there hasn’t been anyone in the research group who’s used these methods… i have, kind of, made these tools up for myself and that is the reason why it has taken such a long time… and i often find myself at a dead end. (bio 18) 4.2 collective and individual knowledge practices the participants described that the knowledge practices involved not only the students themselves, but also others, typically their supervisors, peers or other researchers from their researcher groups. hence, the reported practices were mostly collective (67 %). resulting from the fact that research was often carried out in research groups (table 1). the students reported that engaging in conceptual discussion and working with theoretical ideas, defining and planning their research work, as well as mastering research techniques and writing a publication were the kinds of practices that involved their supervisors, other senior researchers and peers. for instance, the supervisors and colleagues were often active in providing suggestions and guidelines for their students in choosing their research topics, as well as planning their research processes: depending on the article i’ve been working on, many of my colleagues have collaborated with me…discussing how to do this and this. depending on the research questions, the procedure has been different with different people. so maybe this says something about the multidisciplinarity we’ve had, having all these different people with their different viewpoints involved in a single project. (natural science 4) further, while typically the collective practices were also described as established activities, interestingly, there were more descriptions of innovative activities among collective practices than among individual practices. in stem fields, solving complex problems through laboratory or field research often requires intensive group-based collaborative research practices resulting that not only knowledge creation but also researcher development is highly embedded in intensive group-based collaborative research practices. one participant expressed how he had started to develop his own research ideas and increasingly became involved in dialogues throughout the doctoral process: in the beginning, it was mostly the supervisors who had the ideas—that we could do things this or that way. but the longer it took, the more i got into the practicalities and learned to deal with them. after that, i’ve been able to think more about what i want to research next, and to bring more and more of my own ideas to the brainstorming. (medicine 3) participants described a third of all the reported practices as individual (33 %), such as working on their own and how they learned to use research methods, instruments and devices through individual study or experimentation. the students also described their individual responsibilities and the challenges they faced, such as experiences of being without support, as one participant describes: so if the group does something for the first time. i feel that it’s sort of my responsibility. and it slightly burdens me, because i haven’t got any training for that. and the supervisors, they are clearly not able to help. then you feel quite alone. (medicine 2) 4.3 the engagement of doctoral students in knowledge practices further investigation showed that the students typically experienced being actively engaged in knowledge practices (79 %). hence, they perceived themselves as active actors and intentional participants who were able to affect activities and make decisions in the practice at hand. this is a key for cultivating relational agency both in terms of the engaging in knowledge creation in order to deliver original research output as well as in terms of becoming full member of the researcher community. active engagement was described in established and innovative, collective and individual practices related to using and mastering research methods, instruments, or techniques, as well as working with conceptual and theoretical problems, and developing and sharing ideas: i just went through and compared the comments, and there was this sort of eureka moment i had, that maybe all i need to do is to decide for myself. i realized that i have my own opinion about where this should go, and it was very close to what one of my supervisors thought as well. but, it was also against the view of my other supervisor. but then again, in the end, i just made my own decision. and i came to the conclusion that even in the so-called ‘hard sciences’ there really isn’t always exactly one truth to follow. (natural science 3) in the following, another participant expresses how he had an active role and control over his research work: i’ve been given a lot of room for my own self-guidance, and my own thoughts and implementations. and i’ve never really had any difficulties in getting my own thoughts about what i wanted to do heard. so in that regard it’s been quite rewarding, and i’ve been given the opportunity to do plenty of different kinds of things. (bio 18) doctoral students less frequently described adapting existing activities and ideas. such experiences were only occasionally reported among all the instances of knowledge practices (13 %). a characteristic of these experiences was that the students considered themselves to be passive participants who were doing activities and work that someone else wanted or had ordered them to do. accordingly, developing relational agency is not easy or self-evidently resulted from carrying out doctoral research. if doctoral students are not given opportunity to participate and contribute actively to the knowledge practices of their researcher community, including opportunity for experimenting and even to fail, also the opportunities for learning to become a researcher are limited. in some cases, the students believed that the way others carried out the activities was not meaningful for them. such roles were often described in association with established and collective practices. furthermore, these descriptions were typically related to planning the research topic and process, developing research ideas, as well as writing an article. one of the participants expressed his adaptive role in choosing and conducting his doctoral research in the following way: i was told that i should be working on this doctoral dissertation topic. they needed a candidate for it. and i just went and started in that project where they had the opening for one more student and was stuck there. i could not choose the topic myself, but partly there was kind of pressure to have this topic that was worth four academic articles. and i need those four. but in this situation, it is not that i can just creatively come up with something to research. something i’d find interesting to look further into by myself. but it’s just not possible. if the thing you’re working on doesn’t sound [to the supervisors] like it’s going to be good enough, then it’s not worth spending your time on, and you would be told to work on other stuff. kind of from the top down. (medicine 4) doctoral students rarely considered themselves bystanders (8 %), in other words, observers who were left outside the practice. yet experiencing oneself as bystander can be considered highly problematic since it limits doctoral student’s learning both in terms of becoming independent researcher, and in delivering original research output. in the cases where this did occur, they described seemingly unclear and unorganised perceptions of their role in the practice. the bystander role was typically associated with established and collective practices related to, for instance, defining and discussing the research topic and plan, designing the research questions or writing a publication. one participant described his role of a bystander in the following way: then, our clinician actually wrote the paper because i did not have the right clinical background for that. (bio 6). 5. discussion our results show that established knowledge practices played an important role in cultivating doctoral students’ insights into their research, developing creative thoughts and behaviours, enabling them to define a problem space and to solve them. such practices are typically well tested and cultivated over long period of time by the researcher community, and hence provide a grounding for its knowledge creation. engaging doctoral students in these baseline knowledge practices is key both for becoming full member of the community and teaching them about the research and disciplinary practices, as mastery of existing ideas and tools are often a precondition for creativity. the established knowledge practices served as the basis for making discoveries and, hence, the creation of new knowledge (see also sternberg & lubart, 1999). accordingly, the established knowledge practices provided a vehicle for introducing and engaging doctoral students into the researcher communities, i.e. socialising the student as a novice into the academic community (becher & trowler, 2001). they also provided a starting point for researcher development. in addition, the existence and extent of the reported innovative practices evident from our dataset is encouraging, since it implies that doctoral students are contributing to novel ways of working, and the transformation of their respective fields of study (trafford & leshem, 2009; wellington, 2012). engaging in such practices provides also opportunity to learn from the mistakes, and use them as opportunity to further cultivate the established practices. more importantly, it allows doctoral students learn how to develop new collective knowledge practices in order to create knowledge in their field. doctoral students are frequently found to face academic isolation and a lack of academic connections (see for example ali & kohun, 2006; austin, 2009) that hinder their progress. our results on knowledge practices suggest that the doctoral students in the stem fields engaged primarily in collective practices. this is partly explained by the fact that majority of the participants engaged in the intensive research group collaborations and conducted article-based dissertations including typically co-authored articles with senior members of the group. however, the result cannot be reduced into the research group status, since the majority of students in sciences reported that they did not carry out their dissertation work in the research group. accordingly, rather than being matter of the structure, it seems to be a matter of the quality of knowledge practices developed in which doctoral students engaged in that matters. the argument follows that students can be actively engaged in collective practices even though they are not formally carrying out their dissertation work in the group. this finding is in accordance with our prior results on medical, humanities, and behavioural science doctoral students, which showed that more than half of the students perceived themselves as members of a scholarly community and its practices. no statistically significant differences were detected in the previous study even though in medicine the majority of the students worked in a research group and carried out article-based dissertations, while in the humanities the students were more likely to follow a monograph dissertation format and were not formally engaged in research groups (pyhältö, stubb, & lonka, 2009). even tough working collectively does not guarantee that doctoral students won’t experience feelings of isolation, the data presented in this article can be considered encouraging since the advantages of collaboration to productivity, and thus also to knowledge creation, have been emphasised (see for example becher & trowler, 2001). the results also suggest that the doctoral students typically considered themselves active participants in knowledge practices instead of mere adapters or bystanders. this provides a good grounding for cultivating the doctoral students’ relational agency within their researcher communities in terms of knowledge practices that may contribute to eventual knowledge creation. given that research groups and collective work are typical in stem areas, this finding is not surprising. yet, it partly contradicts some of our earlier findings suggesting that a minority of doctoral students in humanities, behavioural sciences, and medicine perceived themselves as active relational agents in their own researcher communities (pyhältö & keskinen, 2012). however, our results also imply that not all the students enjoy equal opportunities to exercise relational agency. moreover, it is important to note that active engagement in knowledge practices in their researcher groups does not guarantee an active role in other activities of the group or in other researcher communities. as the stakes are high for doctoral students to complete their studies in a timely manner, our findings about doctoral students’ active role imply that doctoral education provided engaging learning environment for knowledge creation for the majority of our participants (see also frick, 2010). however, taking on the role of adapter or bystander should not be viewed with outright suspicion, as mastery that supports knowledge creation requires an understanding of existing knowledge and an immersion in the field before such a field can be extended or transformed through new and original work (dewett et al., 2005; sternberg & lubart, 1999). yet, if students spend the majority of their time as either adapters or bystanders, in which case they may become stuck in these roles, or resort to mimicking others’ knowledge work rather than creating their own contribution and transforming the field in so doing it can be considered highly problematic (kiley, 2009). the significance of this study lies in exploring doctoral students’ self-reported knowledge practices in stem fields, and shows that they typically perceive themselves as actively and collaboratively engaged in the practices through transforming their respective fields of study. moreover, the study indicates that doctoral knowledge creation embedded in knowledge practices in the studied stem areas is not only an individual cognitive endeavour. instead, it is also a collective process, which takes place in a broader scientific community, not exclusively limited in conducting doctoral dissertation in the research group 5.1 methodological considerations the strength of the chosen qualitative design was that it enabled a multifaceted and deep investigation of doctoral students’ experiences of knowledge practices. in addition, the multiphase analysis enabled investigation of the knowledge practices from various perspectives. however, one problem with the used retrospective approach is that it exposes the memory effect (cox & hassard, 2007), potentially resulting in difficulties for participants in recalling their experiences (kvale, 2007). at the same time, use of the retrospective approach ensured that the participants had a chance to deeply reflect on their experiences and recall the most significant past events (kvale, 2007). the majority of the participants were in the middle or in the last part of the doctoral process and, because of their experience, they have had more opportunities to be involved in and gain experience with various kinds of knowledge practices. the interview data were collected from 34 doctoral students in the stem fields from a large research-intensive university in finland. because of the distinctive features of the disciplines included (e.g., lindblom-ylänne et al., 2006) and the limited sample size, the results should be generalised to other fields and other countries with caution. knowledge creation practices evolve over time and, hence, further research is needed to explore the knowledge practices among researcher communities from different domains and from a longitudinal perspective. 5.2 implications for doctoral education our results indicate that doctoral students can have active roles and be intentional participants in various scientific knowledge practices. further, the findings suggest that active engagement in knowledge practices can be enabled by supporting doctoral students to influence or direct the surrounding. this requires further developing strategies that promote the intentional participation of students in scientific activities and practices (pyhältö & keskinen, 2012; zhao & kuh, 2004). active engagement can be supported through environments that enable doctoral students to share their knowledge and expertise with others, take more responsibility for and ownership of their research activities, and perceive themselves as contributing members of their community (e.g., dunlap, 2006; mcalpine & amundsen, 2009). for instance, the more experienced members of the researcher communities, such as supervisors, senior researchers and post-doctoral fellows, could support and encourage doctoral students to take increasingly more ownership and responsibility for planning, monitoring and evaluating the everyday practices of knowledge creation. such practices, according to our results, could be planning and conducting actual research work, theoretical problem solving, and dialogues on research ideas. supporting the active role of doctoral students in knowledge practices is likely to be an investment in the quality of future academic work. at best, active doctoral students will become autonomous scientists who create new, high-quality knowledge. keypoints the study aims to contribute to the doctoral education literature by exploring the kinds of knowledge practices in which doctoral students in the stem fields engage during their studies. the significance of this study lies in exploring doctoral students’ self-reported knowledge practices. this study demonstrates that the doctoral students perceive themselves as actively and collaboratively engaged in the knowledge practices. the study concludes that active engagement in knowledge practices can be enabled by supporting doctoral students to influence or direct the surrounding knowledge creation activities. references ali, a., & kohun, f. (2006). dealing with isolation feelings in is doctoral programs. international journal of doctoral studies, 1 (1), 21–33. austin, a.e. (2009). cognitive apprenticeship theory and its implications for doctoral education: a case example from a doctoral program in higher and adult education. international journal for academic development, 14(3), 173–183. https://doi.org/10.1080/13601440903106494 bandura, a. (2001). social cognitive theory: an agentic perspective. annual review of psychology, 52, 1–26. https://doi.org/10.1146/annurev.psych.52.1.1 becher, t. & trowler, p. r. (2001). academic tribes and territories. intellectual enquiries and the culture of disciplines (2nd ed.). open university press, buckingham. bereiter, c. (2002). education and mind in the knowledge age. mahwah, nj: lawrence erlbaum associates. brew, a., boud, d., & namgung, s. u. (2011). influences on the formation of academics: the role of the doctorate and structured development opportunities. studies in continuing education, 33(1), 51–66. https://doi.org/10.1080/0158037x.2010.515575 cox, j. w., & hassard, j. (2007). ties to the past in organization research: a comparative analysis of retrospective methods. organization, 14(4), 475–497. https://doi.org/10.1177/1350508407078049 creswell, j. (2012). qualitative inquiry & research design. choosing among five approaches (3rd ed.). london: sage publishers. cumming, j. (2009). the doctoral experience in science: challenging the current orthodoxy. british educational research journal, 35(6), 877–890. https://doi.org/10.1080/01411920902834191 delamont, s., & atkinson, p. (2001). doctoring uncertainty: mastering craft knowledge. social studies of science, 31(1), 87–107. https://doi.org/10.1177/030631201031001005 dewett, t., shin, s. j., toh, s. m., & semadeni, m. (2005). doctoral student research as a creative endeavour. college quarterly, 8(1), 1–20. dunlap, j. c. (2006). the effect of a problem-centered, enculturating experience on doctoral students’ self-efficacy. interdisciplinary journal of problem-based learning, 1(2), 19–48. https://doi.org/10.7771/1541-5015.1025 engeström, y. (1999). activity theory and individual and social transformation. in y. engeström, r. miettinen, & r. punamäki (eds.), perspectives on activity theory. learning in doing: social, cognitive and computational perspectives (pp. 19–38). cambridge: cambridge university press. fenge, l. a. (2012). enhancing the doctoral journey: the role of group supervision in supporting collaborative learning and creativity. studies in higher education, 37(4), 401–414. https://doi.org/10.1080/03075079.2010.520697 florence, m. k., & yore, l. (2004). learning to write like a scientist: co-authoring as an enculturation task. journal of research in science teaching, 41(3), 637–668. https://doi.org/10.1002/tea.20015 frick, b. l. (2010). creativity in doctoral education: conceptualising the original contribution. in c. nygaard, n. courtney & c.w. holtham (eds.), teaching creativity – creativity in teaching. oxfordshire: libri publishing. frick, b. l. & brodin, e. m. (2014). developing expert scholars: the role of reflection in creative learning. in e. shiu (ed.), creativity research: an interdisciplinary and multidisciplinary research handbook (pp. 312–333). london: routledge. furner, j. (2003). little book, big book: before and after little science, big science: a review article, part i. journal of librarianship and information science, 35(2), 115–125. https://doi.org/10.1177/0961000603352006 gardner, s. k. (2007). “i heard it through the grapevine”: doctoral student socialization in chemistry and history. higher education, 54(5), 723–740. https://doi.org/10.1007/s10734-006-9020-x golde, c. m. (2010). entering different worlds. socialization into disciplinary communities. in s. k. gardner & p. mendoza (eds.), on becoming a scholar. socialization and development in doctoral education (pp. 79–95). virginia, usa: stylus publishing, llc. greeno, j. g. (2006). authoritative, accountable positioning and connected, general knowing: progressive themes in understanding transfer. the journal of learning sciences, 15, 537–547. http://dx.doi.org/10.1207/s15327809jls1504_4 hadwin, a., & oshige, m. (2011). self-regulation, coregulation and socially shared regulation: exploring perspectives of social in self-regulated learning theory. teachers college record, 113(2), 240–264. hakkarainen, k., hytönen, k., makkonen, j., seitamaa-hakkarainen, p., & white, h. (2013). interagency, collective creativity, and academic knowledge practices. in a. sannino & v. ellis (eds.), learning and collective creativity: activity-theoretical and sociocultural studies (pp. 77–95). london: routledge. hakkarainen, k., palonen, t., paavola, s., & lehtinen, e. (2004), communities of networked expertise. professional and educational perspectives . amsterdam: elsevier. hakkarainen, k. p., wires, s., keskinen, j., paavola, s., pohjola, p., lonka, k., & pyhältö, k. (2014). on personal and collective dimensions of agency in doctoral training: medicine and natural science programs. studies in continuing education, 36(1), 83–100. https://doi.org/10.1080/0158037x.2013.787982 hancock, s., hughes, g., & walsh, e. (2017). purist or pragmatist? uk doctoral scientists’ moral positions on the knowledge economy. studies in higher education, 42(7), 1244–1258. https://doi.org/10.1080/03075079.2015.1087994 harry, b., sturges, k. m., & klingner, j. k. (2005). mapping the process: an exemplar of process and challenge in grounded theory analysis. educational researcher, 34(2), 3–13. https://doi.org/10.3102/0013189x034002003 holland, d., lachiocotte, w., skinner, d., & cain, c. (1998). identity and agency in cultural worlds. cambridge, ma: harvard university press holmes, l. (2004). challenging the learning turn in education and training. journal of european industrial training, 28(8/9), 625–638. https://doi.org/10.1108/03090590410566552 hopwood, n. (2010). a sociocultural view of doctoral students’ relationships and agency. studies in continuing education, 32(2), 103–117. https://doi.org/10.1080/0158037x.2010.487482 jazvac-martek, m., chen, s., & mcalpine, l. (2011). tracking the doctoral student experience over time: cultivating agency in diverse spaces. in l. mcalpine & c. amundsen (eds.), doctoral education: research-based strategies for doctoral students, supervisors and administrators (pp. 17–36). netherlands: springer. john-steiner, v. (2000). creative collaboration. oxford university press. kamler, b. (2008). rethinking doctoral publication practices: writing from and beyond the thesis. studies in higher education, 33(3), 283–294. https://doi.org/10.1080/03075070802049236 kiley, m. (2009). identifying threshold concepts and proposing strategies to support doctoral candidates. innovations in education and teaching international, 46(3), 293–304. https://doi.org/10.1080/14703290903069001 kvale, s. (2007). doing interviews. london: sage publications. lave, j., & wenger, e. (1991). situated learning: legitimate peripheral participation. cambridge: university press. lindblom-ylänne, s., trigwell, k., nevgi, a., & ashwin, p. (2006). how approaches to teaching are affected by discipline and teaching context. studies in higher education, 31(3), 285–298. https://doi.org/10.1080/03075070600680539 lipponen, l., & kumpulainen, k. (2011). acting as accountable authors: creating interactional spaces for agency work in teacher education. teaching and teacher education, 27, 812–819. http://dx.doi.org/10.1016/j.tate.2011.01.001 mathieson, s. (2011). developing academic agency through critical reflection: a sociocultural approach to academic induction programmes. international journal for academic development, 16(3), 243–256. https://doi.org/10.1080/1360144x.2011.596730 mcalpine, l., & amundsen, c. (2009). identity and agency: pleasures and collegiality among the challenges of the doctoral journey. studies in continuing education, 31(2), 109–125. https://doi.org/10.1080/01580370902927378 mcalpine, l., & norton, j. (2006). reframing our approach to doctoral programs: an integrative framework for action and research. higher education research & development, 25(1), 3–17. https://doi.org/10.1080/07294360500453012 mcalpine, l., & åkerlind, g. (2010). academic practice in a changing international landscape. in l. mcalpine & g. åkerlind (eds.), becoming an academic. international perspectives (pp. 1–15). united kingdom: palgrave macmillan. miles, m. b., & huberman, a. m. (1994). qualitative data analysis (2nd ed.). thousand oaks, ca: sage publications. mills, j., bonner, a., & francis, k. (2006). the development of constructivist grounded theory. international journal of qualitative methods, 5(1), 25–35. https://doi.org/10.1177/160940690600500103 morgan, d. l. (2007). paradigms lost and pragmatism regained: methodological implications of combining qualitative and quantitative methods. journal of mixed methods research, 1(1), 48–76. https://doi.org/10.1177/2345678906292462 nonaka i, takeuchi h. (1995). the knowledge creating company. oxford university press: new york. o’ meara, k., terosky, a.l. & neumann, a. (2008). faculty careers and work lives: a professional growth perspective. ashe higher education report, 34 (3). san fransisco, ca: jossey-bass. o’ meara, k, & campbell, c. m. (2011). faculty sense of agency in decisions about work and family. the review of higher education, 34 (3), 447–476. http://dx.doi.org/10.1353/rhe.2011.0000 paavola, s., lipponen, l., & hakkarainen, k. (2004). models of innovative knowledge communities and three metaphors of learning. review of educational research, 74(4), 557–576. https://doi.org/10.3102/00346543074004557 pyhältö, k., & keskinen, j. (2012). doctoral students’ sense of relational agency in their scholarly communities. international journal of higher education, 1(2), 136–149. pyhältö, k., nummenmaa, a. r, soini, t., stubb, j., & lonka, k. (2012). research on scholarly communities and development of scholarly identity in finnish doctoral education. in s. ahola & d. m. hoffman (eds.), higher education research in finland. emerging structures and contemporary issues (pp. 337–357). jyväskylä: jyväskylä university press. pyhältö, k., pietarinen, j., & soini, t. (2012). do comprehensive school teachers perceive themselves as active professional agents in school reforms?. journal of educational change, 13(1), 95–116. pyhältö, k., stubb, j., & lonka, k. (2009). developing scholarly communities as learning environments for doctoral students. international journal for academic development, 14(3), 221–232. https://doi.org/10.1080/13601440903106551 pyhältö, k., stubb. j., & tuomainen, j. (2011). international evaluation of research and doctoral education at the university of helsinki to the top and out to society. summary report on doctoral students’ and principal investigators’ doctoral training experiences. retrieved from http://wiki.helsinki.fi/display/evaluation2011/survey+on+doctoral+training sfard, a. (1998). on two metaphors for learning and the dangers of choosing just one. educational researcher, 27(2), 4–13. https://doi.org/10.1080/13601440903106551 sternberg, r. j., & lubart, t. i. (1999). the concept of creativity: prospects and paradigms. in r.j. sternberg (ed.), handbook of creativity (pp. 3–15). cambridge: cambridge university press. stubb, j., pyhältö, k., & lonka, k. (2014). conceptions of research: the doctoral student experience in three domains. studies in higher education, 39(2), 251–264. https://doi.org/10.1080/03075079.2011.651449 trafford, v. & leshem, s. (2009). doctorateness as a threshold concept. innovations in education and teaching international, 46(3), 305–316. https://doi.org/10.1080/14703290903069027 united kingdom quality assurance agency for higher education (2008). the framework for higher education qualifications in england, wales and northern ireland . mansfield: linney direct. volet, s., vauras, m., & salonen, p. (2009). selfand social regulation in learning contexts: an integrative perspective. educational psychologist, 44(4), 215–226. https://doi.org/10.1080/00461520903213584 walker, g. e., golde, c. m., jones, l., conklin bueschel, a., & hutchings, p. (2008). the formation of scholars. rethinking doctoral education for the twenty-first century . san francisco, usa: jossey-bass. wellington, j. (2012). searching for doctorateness. studies in higher education, 38(10), 1490–1503. https://doi.org/10.1080/03075079.2011.634901 zhao, c., & kuh, g. d. (2004). adding value: learning communities and student engagement. research in higher education, 45(2), 115–138. codepen 1.zanden frontline learning research special issue vol.9 no.2 (2021) 9 27 issn 2295-3159 relationships between teacher practices in secondary education and first-year students’ adjustment and academic achievement j. a. c. van der zanden1,2, e. denessen1a. h. n. cillessen 1 & p. c. meijer2 abehavioural science institute, radboud university, the netherlands bradboud teachers academy, radboud university, the netherlands article received 24 april 2020/ revised 4 september / accepted 16 december / available online 12 march abstract to ease the transition to university, preparation in secondary school is often seen as a first step. this study investigated longitudinal relationships between teacher practices in secondary education (i.e., emotional support, autonomy support, and student-centred teacher practices) and first-year students’ academic achievement and social and emotional adjustment at university. we focused on students’ perceptions of their teachers’ practices to, on the one hand, take individual differences into account and, on the other hand, to investigate differences in teacher practices between schools. in a three-wave longitudinal study, 235 students were followed from their final year of secondary school to the end of the first year at university. the results indicated that teacher practices related to students’ social and emotional adjustment across the transition to university, but not to their academic achievement. specifically, we found that perceived teachers’ emotional support was related to students’ social adjustment at university whereas autonomy support was associated with emotional adjustment. differences in teacher practices between schools were quite small. this study indicated that teachers in secondary education might play a pivotal role in preparing students for university. this role goes beyond preparing students for academic achievement, as teachers may have a long-term impact on first-year students’ social and emotional adjustment. keywords: student success; transition from secondary to university education; university preparation; longitudinal study info corresponding author email: p.vanderzanden@docentenacademie.ru.nl doi: https://doi.org/10.14786/flr.v9i2.665 1. introduction by starting a university study, students leave their familiar secondary school to enter a new life-sphere in which they are faced with several academic, social, and intrapersonal changes inherent to university life. most students are well able to deal with those changes, but for a substantial number of students this is not the case. a longitudinal study with first-year students in the united kingdom showed that 60% found it hard to get used to their new university life when they just made the transition (nightingale et al., 2013). most of these students felt better adjusted after three months in the first year, but a group of 31% remained poorly adjusted even six months after starting their first year. those students also demonstrated low academic achievement in the first year. postareff, mattsson, lindblom-ylanne, and hailikari (2016) interviewed finnish students about their experiences in the first year and found that almost 50% of them experienced negative emotions such as feeling tired, stressed out, and overwhelmed by their studies. a subset of these students (17% of the total sample) also demonstrated poor academic achievement. these adjustment and achievement difficulties are worrying because multiple studies have shown that students’ first-year experiences are associated with their future psychological well-being and academic pathway (e.g., allen & robbins, 2008; crede & niehorster, 2012). to ease the transition to university, preparation in secondary education is often seen as an important first step (conley, 2010; noyens, donche, coertjens, van petegem, 2017). thus far, extensive research has shown that it is important ‘how’ first-year students enter university, for example with which secondary school grades and learning skills (see for an overview, e.g., richardson, abraham, & bond, 2012). however, it is largely unknown how students are prepared for university in secondary schools and how that relates to their achievement and adjustment at university. our study contributes to this literature by exploring how teacher practices in secondary education are associated with students’ achievement and adjustment at university. in the current research, we adopted the self-determination theory (sdt; deci & ryan, 1985, 2000) to explore whether teacher practices in secondary education that provide relevant conditions for students’ basic psychological need-satisfaction can lead to better academic achievement, and social and emotional adjustment in the first year at university. 1.1 outcomes in university: first-year student academic achievement and adjustment in the current policy and research discourse on the effects of higher education, it is common to equate student success primarily with academic achievement. this tradition is driven by a neo-liberal view of education accompanied by a strong emphasis on accountability, efficiency, and performance (glastra & van middelkoop, 2018; zajda & rust, 2016). in this study, we focused on students’ grade point average (gpa) as a measure of their academic achievement. traditionally, this is the most commonly used indicator of academic achievement (e.g., caskie, sutton, & eckhardt, 2014; robbins et al., 2004), which is found to predict students’ educational attainment and persistence throughout university (e.g., allen & robbins, 2008; garcía-ros, pérez-gonzález, cavas-martínez, & tomás, 2019). first-year student success might not only be defined in terms of performance. as described by evans, forney, guido, patton, and renn (2010) and mayhew et al. (2016) in their books on college student development, psychosocial theories can help to understand how students’ feelings, behaviours, and relationships change at university. the transition to university not only entails academic challenges (e.g., working more independently and regulating their own learning), but also social (e.g., meeting new people) and intrapersonal challenges (e.g., increasing responsibility and developing a student identity; gale & parker, 2014). students should be able to manage these challenges in order to successfully adjust to the university environment. most research on university adjustment is based on the work of baker and siryk (1989) who defined emotional adjustment as the level of psychological (e.g., worries) and physical distress (e.g., headaches), and social adjustment as how well students are able to deal with the social demands of university life (credé & niehorster, 2012). as suggested by credé and niehorster (2012) and robbins et al. (2004), some researchers see students’ social and emotional adjustment as important university outcomes on their own while others see them as key determinants of students’ academic achievement. a previous meta-analysis points to small relationships between students’ social and emotional adjustment and academic achievement (credé & niehorster, 2012). additionally, in an earlier review study, we found that high social and emotional adjustment does not necessarily relate to high academic achievement in the first year (van der zanden, denessen, cillessen, & meijer, 2018). as such, we see both outcomes as desirable and valued outcomes in their own right. 1.2 preparing students for university in secondary education preparation is argued to be very important in order to support students to successfully make the transition to university (conley, 2010; kyndt, donche, trigwell, & lindblom-ylänne, 2017). in the netherlands, the pre-university track of secondary education plays a major role in this preparation. previous studies have already focused on what teachers in secondary education do to prepare students for university. based on interviews with dutch secondary school teachers, van rooij and jansen (2018) found that teachers aimed to prepare students for university by talking with them about specific degree programmes and studying in general, and promoting their study skills, research skills, attitude of inquiry, and independence. further, for the united states conley (2010) showed that teachers from schools that are successful in preparing students for university contribute to students’ cognitive strategies (e.g., problem solving), self-management skills, and academic behaviours (e.g., task prioritising) in their lessons. although the abovementioned studies showed ways in which teachers can contribute to student university preparation, they provide limited insight into long-term relationships between teachers’ practices in secondary education and students’ academic achievement and adjustment at university. in the current research, we aimed to expand the scope of previous research by adopting the self-determination theory (sdt; deci & ryan, 1985, 2000) to explore whether teachers’ practices in secondary education that provide relevant conditions for students’ basic psychological need-satisfaction can lead to better adjustment and achievement at university. in this regard, the transition is seen as an ongoing process (i.e., instead of a single event) which unfolds over time and across school contexts (ellerbrock & kiefer, 2013). according to the sdt, fulfilment of students’ basic psychological needs of autonomy (i.e., perceiving one’s behaviours as originating from the self), competence (i.e., being able to produce important outcomes), and relatedness (i.e., developing meaningful relationships with others) contributes to higher intrinsic motivation and better learning (deci & ryan, 1985; niemiec & ryan, 2009). teachers can play a key role in supporting students’ needs, and thereby contributing to students’ learning and development (niemiec & ryan, 2009). in line with previous studies that adopt a sdt approach in relation to educational transitions (e.g., ellerbrock & kiefer, 2013; madjar & cohen-malayev, 2016), we assumed that school contexts that meet students’ needs promote a smooth transition from one school context to the next. applied to the transition to university, this means that when students’ needs for autonomy, competence, and relatedness are met in secondary education, students can already develop dispositions (i.e., more positive learner identities, self-confidence, and autonomous motivation) that are relevant for university education. in this way, it might be easier for students to adjust to the university environment as they can rely on them at university. teachers in secondary education can foster students’ needs in various ways. based on the sdt, we focused in the current study on the following three teacher practices: emotional support, autonomy support, and student-centred teaching. in the next paragraphs, we describe these teacher practices and how they relate to dispositions that might be useful at university more elaborately. 1.2.1. teachers’ emotional support sdt posits that to support students’ basic psychological needs, and the need for relatedness in particular, it is important that teachers provide emotional support to students (niemic & ryan, 2009). this refers to the extent to which teachers show affection, express their understanding, dedicate resources (e.g., aid, time), and are available in case of need (stroet, opdenakker, & minnaert, 2013). previous studies have shown that student perceptions of teachers’ emotional support not only contributed to their actual engagement in learning, motivation, and achievement (see the reviews of roorda, koomen, spilt, & oort, 2011; stroet et al., 2013), but might also have a lasting effect across a school transition (e.g., langenkamp, 2010). focused on the transition from middle to high school, langenkamp (2010) found that positive relationships with middle school teachers prevented students from failing courses in the first year at high school. she explained that teachers’ emotional support before the transition might foster students’ love of learning and skills to find similar teacher relationships in the new school environment which might influence students’ academic achievement afterwards. 1.2.2. teachers’ autonomy support teachers can also support students’ needs according to sdt by providing autonomy support (niemiec & ryan, 2009; stroet et al., 2013). this entails encouraging students to pursue their own goals and providing students with choice and freedom in study activities or giving a rationale when choice is constrained. additionally, autonomy-supportive teachers use inviting language (e.g., ‘you can’) instead of controlling language (e.g., ‘you should’) and try to avoid explicit rewards and punishments. several studies have shown that autonomy support positively related to students’ academic achievement and well-being, for example, among norwegian secondary school students (diseth & samdal, 2014) and university students (ljubin-golub, rijavec, & olčar, 2020). further, it contributes to the development of dispositions among secondary school students that are highly valued in university environments such as academic efficacy and mastery goal orientation (kenny, walsh-blair, blustein, bempechat, & seltzer, 2010) and self-regulated learning (sierens, vansteenkiste, goossens, soenens, & dochy, 2009). additionally, in a recent study with university students, hernández, moreno-murcia, cid, monteiro, & rodrigues (2020) found that perseverance mediated the relationship between perceived autonomy supportive behaviours by teachers and academic achievement. it seems that if students perceived their teachers as autonomy-supportive they demonstrate a stronger desire to endorse in self-regulated behaviours and to complete quality work, which results in higher academic achievement. 1.2.3. student-centred teaching student-centred teaching refers to practices in which the teacher is the facilitator of the learning process, characterized by the stimulation of knowledge construction, cooperative work, authentic assignments, and opportunities for self-regulated learning (prosser & trigwell, 1999). need-supportive teaching behaviours share theoretical notions with a constructivist, student-centred approach to teaching (e.g., stroet, opdenakker, & minnaert, 2015; cents-boonstra, lichtwarck-aschoff, denessen, aelterman, & haerens, 2020). for example, stroet et al. (2015) found that in constructivist student-centred classrooms teachers more often provided individual guidance to students, helped students in directing their learning processes, fostered the relevance of learning activities, and did not express their dissatisfaction. a basic assumption in the approaches to teaching literature (trigwell, prosser, and taylor, 1994) is that approaches to teaching influence students’ learning approaches and their subsequent learning outcomes. previous research (e.g., beausaert, segers, & wiltink, 2013) has shown that student-centred teaching stimulated students to adopt a deep approach to learning, which is characterized by the search for meaning in a task and an intrinsic motive to attain understanding. this learning approach is often advocated as most appropriate for learning in higher education (e.g., vanthournout, gijbels, coertjens, donche, & van petegem, 2012). research on the relationship between this type of learning and academic achievement have shown contradictory results. nonetheless, studies have generally found a positive and weak relationship between a deep approach to learning and students’ academic achievement (see for literature overviews, e.g., richardson et al., 2012; watkins, 2001). further, torenbeek, jansen, and hofman (2011a) studied how the teaching approach at secondary school corresponded to the teaching approach at university, and how that related to the academic achievement of first-year students. they only found positive results when teaching at university was a bit more student-centred (i.e., teacher exerts less control) than in secondary school. when students joined a university program with less student-centred teaching than in secondary education, their academic achievement was lower. additionally, related to students’ emotional adjustment, applying a deep approach to learning related to less study-related stress (asikainen, salmela-aro, parpala, & katajavuori, 2019). 1.3 present study the aim of our study was to explore longitudinally how specific teacher practices in secondary education related to first-year university students’ adjustment and achievement. the following research question was addressed: how do students’ perceptions of teachers’ emotional support, autonomy support, and student-centred teaching in secondary education relate to students’ academic achievement, and social and emotional adjustment in the first year at university? hereby we focused on students’ perceptions of their secondary education teachers’ practices because previous research suggests that how students perceive their teachers’ practices, rather than the practices themselves, affect their learning and development (prosser & trigwell, 1999). this also enabled us to investigate individual differences in students’ perceptions, while, at the same time, allowed an examination of the degree of consensus among students from the same school about their teachers’ preparation practices. herewith, the study combined two levels of diversity needs as it is located at the boundary of the micro-level of student variability and the meso-level of the learning environment in secondary education. the innovative character of this study consists of moving beyond the individual by considering the role of secondary education teachers in the transition to university. further, it used a longitudinal design with three measurement moments across the transition from secondary school to university. 2. method 2.1 transition to university in the netherlands in the netherlands, secondary education is divided into different levels, namely pre-vocational secondary education, senior general secondary education, or pre-university education. around the age of 12, students with the highest achievement levels are tracked in the pre-university level of secondary education. normally, this track takes six years to complete (grade 7 to 12), but the actual duration may differ if students resit or skip classes or enrol after they have already completed senior general secondary education. graduating from the pre-university track—consisting of both school and mandatory national examinations—grants direct access to degree programmes at research universities 1 . additionally, most universities are public institutions with standard tuition fees determined by the ministry of education, culture and science (nuffic, 2015). in 2016, 75% of the pre-university students directly made the transition to university after graduating (vsnu, 2017). the other students enrolled at professional higher education or did not enrol in post-secondary school education. 2.2 sample eight schools in the south-east of the netherlands participated in our study. data were collected at the end of students’ final year at secondary school (time1, april 2016), two months after their transition to university (time2, november 2016), and after finishing the first year at university (time3, october 2017). at time1, 496 grade 12 students (51% female, ± 18 years of age) participated. of these 496 students, 54 students indicated that they were not enrolled in university education after obtaining their secondary school degree. of the remaining 442 students, 314 students (57% female) completed the questionnaire at time2. these students started their studies in 19 different cities at 21 different educational institutions: 93% of the participants entered a research university and 7% enrolled at professional higher education. four students went to study abroad and were not included in the analyses. we did not find differences between students who did and did not participate at time2 on the secondary school they attended, educational profile in secondary school, and social and emotional adjustment in secondary school. at time3, 235 students participated (61% female). of these, 201 students had continued their study after the first year, 32 had switched to another study or university, and 2 had stopped studying. anovas did not reveal differences between students who did and did not participate at time3 on their academic achievement, and social and emotional adjustment in secondary school. next to participant dropout over time, we noticed missing rates per variable ranging from 0 to 13% (n = 31). 2.3 measures data were collected by means of questionnaires. figure 1 provides an overview of the research design and the times at which the variables were measured. figure 1. overview of research design and times at which the variables were measured 1 as the results of students’ final exams in secondary school are announced in june, we asked students to report these results at time2. as they concern students’ academic achievement in secondary school, we refer to them in this manuscript as ‘academic achievement time1’. academic achievement. at time2, students reported their grade point average (gpa) in subjects in which they did final exams in secondary school. these subject gpas were averaged to form one composite score for students’ secondary school gpa. at time3, students were asked to report their average grade attained at courses in the first year at university. in the netherlands, grades are measured on a scale of 1 to 10, with 10 being excellent and 5.5 the minimal score to pass exams. social and emotional adjustment. students’ social and emotional adjustment was measured with two subscales of the student adjustment to college questionnaire (baker & siryk, 1989), which was translated and validated in the dutch language by beyers and goossens (2002). the 6-item ‘emotional adjustment’ scale measures the level of psychological and physical stress students experienced at secondary school (time1) and university (time2 and time3). an example item was ‘i have been feeling in good health lately’. the 4-item ‘social adjustment’ scale measures how well students were able to deal with the interpersonal demands of secondary school (time1) and university (time2 and time3) (e.g., ‘i had several close ties at school/university’). answers were given on a 5-point scale (1 = strongly disagree, 5 = strongly agree). cronbach’s α for the emotional adjustment scale was .76 at time1, .76 at time2, and .78 at time3. for social adjustment, cronbach’s α was .80, .77, and .82 for time1, time2, and time3, respectively. teachers’ emotional support. the ‘relatedness’ scale from the teacher as social context questionnaire (tasc; belmont, skinner, wellborn, & connell, 1988; sierens et al., 2009) was used to measure whether students perceived their teachers as emotionally supportive. the scale comprises eight items (e.g., ‘my teachers know me well’) and students were instructed to respond to the items with respect to the teachers who teach them in their final year of secondary education. for each item students indicated on a 5-point scale how many of their teachers showed the behaviours (1 = none of my teachers, 5 = all of my teachers). we opted for this scale because if more teachers show the behaviour, students have more opportunities to get used to these practices and to develop associated skills and strategies which might make the transition to university easier. cronbach’s α was .80. teachers’ autonomy support. the 8-item ‘autonomy’ scale from the tasc (belmont et al., 1988) was used to measure the extent to which students perceived that their teachers support their autonomy. an example item was ‘my teachers give me a lot of choices about how to do my schoolwork’. answers were given on a 5-point scale (1 = none of my teachers, 5 = all of my teachers) and cronbach’s α was .72. student-centred teaching. the 8-item ‘student-centred approach’ scale from the approaches to teaching inventory (trigwell & prosser, 2004)—translated to dutch by beausaert et al. (2013)—was used to measure whether students perceived their teachers’ practices as learning-centred. an example item was ‘my teachers create opportunities for us to discuss our changing understanding of the subject’. again, a 5-point scale was used and cronbach’s α was .68. 2.4 procedure approval for this study was granted by the ethics committee of the faculty of social sciences of our university. after obtaining permission from the principals of the eight secondary schools, informed consent was obtained from the pre-university students and their parents. at time1, students filled out the questionnaire (± 30 minutes) during regular classes in secondary education. students could report their e-mail address so that we could contact them for data collection at time2 and time3. at these time points, students received an online questionnaire by e-mail which took them 10 minutes to complete. to increase our response rate at time2 and time3, students were sent two online reminders to complete the survey. further, a small incentive for participation was offered. five gift cards (10 euro) were raffled among the participants at time2 and time3. 2.5 statistical analyses following an examination of descriptive statistics and correlations, we investigated how teachers’ practices in secondary education were associated with first-year students’ academic achievement and social and emotional adjustment. for students’ academic achievement—measured at two time points—a mediation analysis was conducted. the analysis was performed with the process macro version 3.0 for spss 23.0 developed by hayes (2018). in this approach, effects are evaluated with bias corrected bootstrap 95% confidence intervals. these intervals are significant when the upper and lower bounds of the interval do not contain zero. further, mediation effects can be significant regardless of the significance of the total effect (hayes, 2018). because our independent variables were moderately correlated, we reported structure coefficients in addition to beta coefficients, as suggested by kraha, turner, nimon, reichwein zientek, and henson (2012). structure coefficients provide insight in the bivariate relationship between an independent variable and the observed effect without influence of other variables. social and emotional adjustment were measured at all three time points. this allowed the measurements to be correlated within individuals. to account for these correlated observations, we performed a multilevel model for change—as suggested by singer and willett (2003, chapter 7, page 243) —in which we modelled within subject effects (time), between subject effects (emotional support, autonomy support, student-centred teaching), and interactions between them (e.g., time*emotional support). including interaction effects in the analyses enabled us to model the influence of the teaching practices on students’ adjustment on each of the three measurement moments. in a multilevel model for change, the model’s random effects are embodied in its error covariance structure. we used the ‘unstructured’ function to account for unequal correlations and variances between time points. the ‘unstructured’ function allows each element of the error covariance structure to take on the value that the data demand (singer & willett, 2003). thus, it is an appropriate method when you have just a few waves of data collection. to enhance interpretation of our parameters, we recentered the three teachers’ practices by subtracting 1 from students’ scores, to create a scale from 0 to 4 instead of 1 to 5. the analyses were performed with the ‘nlme’ package in r (pinheiro, bates, debroy, & sarkar, 2018). 3. results 3.1 preliminary analysis the means and standard deviations for all measures are listed in table 1. significant gender differences were found. male students rated themselves higher than female students did on emotional adjustment at time1 (d = .61), time2 (d = .49), and time3 (d = .39). further, male students perceived their teachers as more student-centred (d = .27) than female students. hence, gender was included as a control variable in all analyses. further, schools differed in the extent to which students perceived their teachers’ practices as emotionally-supportive, autonomy-supportive, and student-centred (table 1). however, the intraclass correlations (iccs) indicated much more variation within schools than between schools, as less than 9% of the variance in students’ perceptions of teacher practices could be explained by differences between schools. as shown in table 2, it was found that students with a higher gpa in secondary school obtained a better gpa in the first year at university (r = .53, p < .01). for students’ adjustment, we saw different patterns of correlations for emotional and social adjustment. the correlations between emotional adjustment at time1, time2, and time3 ranged between .56 and .60, indicating that students with higher emotional adjustment in secondary school obtained better emotional adjustment at university. correlations for social adjustment at the three time points varied more; the correlation between time2 and time3 was stronger (r = .43, p < .01) than the correlations of time1 with time2 (r = .15, p < .01) and time3 (r = .23, p < .01). further, teacher practices in secondary education showed no significant correlations with students’ academic achievement, but small correlations with social and emotional adjustment.   table 1 descriptive statistics with gender and school differences for all study variables measured on time1 (n = 496), time2 (n = 314), and time 3 (n = 235) note: aa = academic achievement, sa = social adjustment, ea = emotional adjustment, es = emotional support, as = autonomy support, sc = student-centred; * p < .01. table 2 correlations among the study variables measured on time1 (n = 496), time2 (n = 314), and time 3 (n = 235) note: aa = academic achievement, sa = social adjustment, ea = emotional adjustment, es = emotional support, as = autonomy support, sc = student-centred; * p < .01. 3.2 relations between teacher practices in secondary school and first-year academic achievement with process, we examined whether teachers’ emotional support, autonomy support, and student-centred teaching related to students’ first-year gpa (measured at time3), and whether these effects were mediated by students’ gpa in secondary education. as shown in table 3, none of the teacher practices was significantly related to students’ gpa in secondary education (r² = .03), nor to their first-year university gpa (r² = .01). there was a direct effect of students’ gpa in secondary education on their first-year university gpa, b = .60, p < .01, r² = .28. structure coefficients were reported in table 3, but because the teacher practices were not significantly correlated with students’ gpa at secondary school or at university, the regression coefficients were not suppressed by correlations between independent variables. table 3 unstandardized regression coefficients (b) with standard errors (se) and structure coefficients (rs) estimating students’ academic achievement in secondary education (mediator) and in the first year at university (n = 191) note: es = emotional support, as = autonomy support, sc = student-centred, aa = academic achievement. a model 1 shows the effects of the three teacher practices in secondary education on first-year students’ academic achievement (controlled by gender); model 2 shows the effects of the three teacher practices and students’ academic achievement in secondary education on first-year students’ academic achievement (controlled by gender); * p < .05. 3.3 relations between teacher practices in secondary school and social adjustment in table 4, the reference category was time1. on time1, students’ initial level of social adjustment was estimated to be 4.35. the within-subject effects of time showed that on average students’ social adjustment significantly declined from time1 to time2 (b = -.57, p < .01), and from time1 to time 3 (b = -.45, p < .01). we also performed an analysis with time2 as reference category, revealing that students’ social adjustment increased from time2 to time3 (b = .12, p = .01). next, the between-subject effects of the three teacher practices on students’ social adjustment were investigated. as can be seen in table 4, students who perceived their secondary school teachers as emotionally supportive were more likely to experience social adjustment in general, b = .32, p < .01. as we were especially interested in whether teacher practices in secondary school related to students’ social adjustment at university, we included interactions between emotional support and time in our model to see whether the effects differed across the three measurement moments. emotional support was significantly related to students’ social adjustment at time1, b = .37, p < .01. the interaction effects were not significant, indicating that the effect of emotional support on time2 and time3 is not significantly different from the effect on time1. further investigation in which we changed the reference category (from time1 to time2 and time3, respectively) showed that emotional support was significantly associated with emotional adjustment at time 2, b = .23, p < .01, but not at time3, b = .18, p = .08. this indicates that the effect of teachers’ emotional support on social adjustment extended to university, but faded away over time.   table 4 results of fitting a multilevel model for change on students’ social adjustment data note: es = emotional support, as = autonomy support, sc = student-centred; * p < .05. 3.4 relations between teacher practices in secondary school and emotional adjustment the first model, including within-subject effects of time, indicated that on time1 students’ initial level of emotional adjustment was 3.59 (table 5). the within-subject effects of time2 and time3 were non-significant, indicating that on average students’ emotional adjustment did not differ across the measurement moments. next, results of the second model with both within-subject and between-subject effects showed that teachers’ emotional support, b = .16, p < .01, and autonomy support, b = .20, p < .01, were related to students’ emotional adjustment in general (table 5). this means that students who perceived their teachers’ practices as emotionally supportive and autonomy-supportive were more likely to report higher emotional adjustment scores in secondary school and in the first year at university. as we were especially interested in whether teacher practices in secondary school related to students’ emotional adjustment at university, we included interactions between these variables (emotional support and autonomy support) and time in our model to see whether the effects differed across the three measurement moments (table 5, right column). the results showed that teachers’ emotional support contributed to students’ emotional adjustment at secondary school (time1), b = .25, p < .01. the interaction effects between emotional support and time2, and emotional support and time3 were both significant (b = -.23, p <.01, and b = -.19, p = .04, respectively) indicating that the effect of emotional support on emotional adjustment at time2 and time3 is significantly lower than at time1. further investigation in which we changed the reference category (from time1 to time2 and time3, respectively) showed that emotional support indeed did not have a significant effect on students’ emotional adjustment at time2 (b = .02, p = .82) and time3 (b = .06, p = .52). regarding the interactions between autonomy support and time, the results showed that autonomy support was associated with emotional adjustment at time1, b = .19, p < .01. the interaction effects were not significant, indicating that the effect of autonomy support on time2 and time3 is not significantly different from the effect on time1. additional analyses in which we changed the reference category showed that autonomy support also was associated with emotional adjustment at time2, b = .18, p = .02, and time3, b = .33, p < .01. this indicates that autonomy support had both an effect on students’ emotional adjustment in secondary school and in the first year at university. table 5 results of fitting a multilevel model for change on students’ emotional adjustment data note: es = emotional support, as = autonomy support, sc = student-centred; * p < .05 4. discussion 4.1 discussion of main findings previous research underlined the pivotal role of teachers in secondary education in preparing students for university (e.g., conley, 2010; van rooij & jansen, 2018). our study aimed to expand the scope of prior research by adopting a sdt approach to investigate how students’ perceptions of specific teacher practices in secondary education related to first-year students’ adjustment and achievement at university. with a longitudinal study including three measurement moments, we focused on student development over time and across school contexts. the results indicated that perceived teacher practices in secondary education related to students’ social and emotional adjustment across the transition to university, but not to their academic achievement. specifically, we found that perceptions of teachers’ emotional support contributed to first-year students’ social adjustment, whereas autonomy support related to students’ emotional adjustment at university. this study indicated that the secondary education learning environment is not school-specific. instead, students perceive variability in features of the learning environment. these individual differences might arise because, for example, students from the same school have different teachers depending on the subjects they chose for their final exams, students perceive the same teacher behaviours differently (e.g., because of different needs) or because teachers behave differently towards different students (e.g., providing differentiated instruction). investigating students’ perceptions allowed us to take this individual diversity into account. the finding that students from the same school can have rather different perceptions of their teachers’ practices emphasizes the need of taking student variability into account in the transition to university, instead of focusing on which secondary school students attend. our study adds to the growing body of literature concerning the sdt in relation to educational transitions (e.g., ellerbrock & kiefer, 2013; madjar & cohen-malayev, 2016). it seems that when students perceive autonomy and emotional support in secondary education, they might already develop dispositions and skills that are relevant for university education, through which it might be easier for them to adjust to the university environment. this corresponds to conley’s (2010) view who argued that creating a learning environment in secondary education that enables students to already internalize behaviours, such as working autonomously and taking responsibility for their learning, eases the transition to university, and to research of langenkamp (2010) who described that positive relationships with teachers before a transition provide students with skills to find similar supporting relationships after a transition. however, with the present study we did not gain insight in the underlying mechanisms through which teacher practices in secondary education might enhance adjustment in the first year at university. this might be an interesting avenue for future research, for example by means of a qualitative retrospective study in which students reflect on their secondary school experiences and how they help them adjust to the university environment, or by investigating in a quantitative longitudinal study whether variables as autonomous motivation and academic self-confidence mediate the relationship between teacher practices and university outcomes. for academic achievement we found in line with previous research (richardson et al., 2012) that students who obtained higher grades in secondary education also obtained a higher gpa in the first year at university. neither of the three secondary school teacher practices included in this study related to university gpa, explaining only 1% of the variance. although this might suggest that teacher practices in secondary education do not matter for academic achievement at university, there may be other explanations for these findings. for example, we investigated the effects of teacher practices in secondary education on the academic achievement of first-year students from different study programs together, because previous research pointed to general differences between secondary school and university education. for example, oolbekkink-marchand, van driel, and verloop (2014) found that secondary education teachers attached more value to shared regulation of learning activities with their students, whereas university teachers expected their students to work independently and to regulate their own learning. nevertheless, the students in our sample enrolled in more than 100 different study programs that are likely to differ from each other (see, e.g., lindblom-ylänne, trigwell, nevgi, & ashwin, 2006). because of such differences, it is possible that teacher practices in secondary education have differential effects among study programs. torenbeek, jansen, and hofman (2011b) found, for example, that first-year students from secondary schools with strong teacher control obtained more credit points in so-called ‘soft’ sciences (e.g., psychology), whereas they did not find such effect for students in ‘hard’ sciences (e.g., life science and technology). because of the huge number of programs that the students in our sample were enrolled in, we were not able to analyse the effects of these programs. future research could focus on the variation between study programs to gain deeper insight into the effects of teacher practices in secondary education on students’ academic achievement in different university contexts. further, our results indicated that student-centred teacher practices did not relate to students’ adjustment or achievement at university. a possible explanation for this is that student-centred teacher practices may take many forms and different factors in the student-centred environment might encourage or discourage students’ learning, for example the assessment mode, students’ workload, or type of feedback (baeten, kyndt, struyven, & dochy, 2010; postareff, mattsson, & parpala, 2018). this indicates that it is not a student-centred environment in itself but specific aspects in the environment that might contribute to student outcomes, which were not specifically included in our assessment of student-centred teacher practices. 4.2 limitations and recommendations for future research in interpreting our results, some limitations should be considered. first, we focused on students’ perceptions of teacher practices in secondary education, rather than teachers’ self-assessments or classroom observations. this is because multiple studies have shown that how students perceive their teachers’ practices affect their learning and development, instead of the practices in itself (e.g., prosser & trigwell, 1999; stroet et al., 2013). further, we asked students how they perceived the practices of their teachers in general, rather than using domainor subject-specific (e.g., math or english) measures. nevertheless, what one teacher does can have a large impact on a student irrespective of other teachers’ behaviours (e.g., a physics teacher’s practices for a student who pursues a physics degree). this is not reflected in the current scale. for future research, it is interesting to explore whether the perceived quantity or quality of teacher practices in secondary education matter for the transition to and success at university. further, we recommend to add other measures of teacher practices next to student perceptions (e.g., teacher interviews or classroom observations) to gain richer insights in how the three included teacher practices manifest themselves in secondary education. this might also be useful for professional development programmes for (student) teachers on how to prepare students for university success. a second limitation of this study is that we did not explicitly focus on students’ needs in relation to teacher practices. however, by focusing on students’ perceptions of teachers’ behaviours, we implicitly take diversity in student needs into account. students’ perceptions of their teachers’ behaviours might be different because, for example, students perceive the same teacher behaviours differently because of different needs, or because teachers adapt their behaviour to individual student needs. in future research, it might be interesting to investigate whether students’ needs in general and the fit between students’ needs and teachers’ practices (in secondary education) in specific influence the transition to university. it might be, for example, that students whose secondary school and university fit the individual’s developmental needs experience a more positive transition than students who experience a misfit between their learning environments and needs. third, because we focused in this study on student success – defined as students’ academic achievement and social and emotional adjustment – we focused on students who continued their studies in the second year. students who dropped out or switched studies were not taken into account in our analyses, as these are complex phenomena in itself. focusing on student dropout and switching behaviour would be an interesting avenue for future research, involving a different way of data collection, which might shed an additional light on how teacher behaviours might contribute to how students experience the first year. additionally, it is possible that students who obtained lower grades in secondary school did not respond to our requests to fill out the questionnaires at time2 and time3, and therefore, we should interpret our findings with caution. fourth, the present study was limited by the methods we used. the internal reliability of our scales ranged from moderate to good. especially the scales to measure teachers’ practices in secondary education (e.g., student-centred teaching practices) may need further refinement. furthermore, although variables were measured over time, it is still difficult to make causal claims because we have neither been able to control (many) possible confounding variables, nor to include a control group. our study supports causal theories (e.g., self-determination theory which assumes that perceived emotional support contributes to students’ well-being; niemiec & ryan, 2009), but strong inferences about causal direction cannot be made. this is especially a concern at time1 because there perceived teachers’ practices and students’ adjustment are measured simultaneously. an interesting avenue of further investigation is to examine the nature and complexity of these associations allowing for more conclusive statements about the associations between (perceived) teacher practices in secondary education and students’ adjustment and achievement. further, while we performed a multilevel model for change other analyses can be used as well, for example a latent-growth model. the statistical models of these analyses are identical (willett, 2004) but latent-growth models might be useful because they permit the modelling of change in several domains simultaneously. fifth, this study took place in the netherlands which might have consequences for the generalisability of the findings. secondary education in the netherlands is a differentiated system in which students are tracked by ability into specific educational programmes. at first sight, our findings are therefore specifically transferable to countries that also have a differentiated secondary education system (i.e., academic and vocational streaming) such as austria, belgium, and japan (chmielewski, 2014). however, they might also be applicable to other educational systems in which some form of differentiation by student achievement levels is used in the curricula. for example, in the united states and australia, courses in one subject are often offered at varying levels of difficulty within a school (chmielewski, 2014). the high-ability level courses might be comparable to the courses in the dutch pre-university track. in future research it would be useful to perform our research in other educational systems and/or countries to gain more insight into the generalisability of the findings. 4.3 conclusion and practical implications this study indicated that teachers in secondary education can play a role in preparing students for university success. this role goes beyond preparing students for academic achievement at university, as teachers also have a long-term impact on first-year students’ social and emotional adjustment. very general teacher behaviours as emotional support and autonomy support related to students’ social and emotional adjustment in the first year at university. the results provide useful insights for educational policy in secondary education. one way in which secondary education usually prepares students for university is by enabling them to pass the standardised secondary school examinations and to obtain high grades. the average grade of students in secondary education indeed is strongly correlated with their academic achievement in the first year at university. in this way, secondary education already plays a large role in preparing students for first-year academic success. however, there is more to student success than academic performance in which secondary education plays a role. student preparation for university also has a social-emotional dimension which means paying attention to how students cope with problems and emotions. our findings argue for an increasing awareness of the role of secondary education teachers play in student preparation for university that goes beyond mere academic preparation, and the increasing embeddedness of this multi-domain perspective in educational policy. teachers in secondary schools can employ a range of different practices to prepare students for university success. most obviously, they can stimulate the academic achievement of future first-year students by ensuring that they pass their final secondary school examinations with good grades. a previous meta-analysis by richardson et al. (2012) showed that students’ secondary school gpa correlated .40 with university gpa. in addition, it is important that students perceive their learning environment as autonomy-supportive and emotionally-supportive. according to stroet et al. (2015), teachers can support students’ autonomy, for example, by creating opportunities for students to work in their own way and by incorporating students’ interests into their lessons. teachers may provide emotional support by showing interest to what students are saying and to what is of importance for them. by incorporating those general teacher behaviours in secondary education, teachers can help students to make a smooth transition to university. however, we should bear in mind that these effects might be subtle as the correlations between perceived emotional and autonomy support and adjustment at university ranged from .16 to .27. next to (teachers in) secondary education, university faculty also play a major role in students’ transition to university. more contact and cooperation between secondary education and university is needed to equip university faculty with greater insight into what they can expect from incoming students, on which they can build during their courses in the first year at university. they can, for example, provide autonomy support in the first year – thereby creating a continuous development process with secondary education – to allow students to internalise behaviours as autonomous and independent learning which make it easier for them to adjust to the university environment. in this way, university education can be better attuned to students’ previous academic experiences and developed skills, to facilitate first-year students’ success keypoints this study investigated how teachers’ practices in secondary education related to first-year students’ academic achievement and social and emotional adjustment. in a three-wave longitudinal study, 235 students were followed from their final year of secondary school to the end of the first year at university. results indicated that emotional support related to students’ social adjustment at university whereas autonomy support was associated with emotional adjustment. no relationships between teachers’ practices in secondary education and students’ first-year academic achievement were found. this study indicated that secondary education teachers’ practices might have a long-term impact on first-year students’ adjustment at university. acknowledgments we thank ben pelzer for his helpful statistical advice. footnotes 1 some degree programmes have additional requirements such as compulsory courses in pre-university education or additional admittance policies (e.g., medicine). references allen, j., & robbins, s. b. (2008). prediction of college major persistence based on vocational interests, academic preparation, and first-year academic performance. research in higher education, 49, 62-79. doi:10.1007/s11162-007-9064-5 asikainen, h., salmela-aro, k., parpala, a., & katajavuori, n. (2020). learning profiles and their relation to study-related burnout and academic achievement among university students. learning and individual differences, 78. doi:10.1016/j.lindif.2019.101781 baeten, m., kyndt, e., struyven, k., & dochy, f. (2010). using student-centred learning environments to stimulate deep approaches to learning: factors encouraging or discouraging their effectiveness. educational research review, 5, 243-260. doi:10.1016/j.edurev.2010.06.001 baker, r. w., & siryk, b. (1989). student adaptation to college questionnaire (sacq): manual. los angeles, california: western psychological services. beausaert, s. a. j., segers, m. s. r., & wiltink, d. p. a. (2013). the influence of teachers’ teaching approaches on students’ learning approaches: the student perspective. educational research, 55, 1-15. doi:10.1080/00131881.2013.767022 belmont, m., skinner, e., wellborn, j., & connell, j. (1988). teacher as social context: a measure of student perceptions of teacher provision of involvement, structure, and autonomy support . rochester, new york: university of rochester. beyers, w., & goossens, l. (2002). concurrent and predictive validity of the student adaptation to college questionnaire in a sample of european freshman students. educational and psychological measurement, 62, 527-538. doi:10.1177/00164402062003009 caskie, g. i. l., sutton, m. c., & eckhardt, a. g. (2014). accuracy of self-reported college gpa: gender-moderated differences by achievement level and academic self-efficacy. journal of college student development, 55, 385-390. doi:10.1353/csd.2014.0038 cents-boonstra, m., lichtwarck-aschoff, a., denessen, e., aelterman, n., & haerens, l. (2020). fostering student engagement with motivating teaching: an observation study of teacher and student behaviours. research papers in education, doi:10.1080/02671522.2020.1767184 chmielewski, a. k. (2014). an international comparison of achievement inequality in withinand between-school tracking systems. american journal of education, 120, 293-324. doi:10.1086/675529 conley, d. t. (2010). key principles of college and career readiness. in d. t. conley (ed.), college and career ready. helping all students succeed beyond high school (pp. 104-132). san francisco, california: jossey-bass. credé, m., & niehorster, s. (2012). adjustment to college as measured by the student adaptation to college questionnaire: a quantitative review of its structure and relationships with correlates and consequences. educational psychology review, 24, 133-165. doi:10.1007/s10648-011-9184-5 deci, e. l., & ryan, r. m. (1985). intrinsic motivation and self-determination in human behavior. new york, ny: plenum press. doi:10.1007/978-1-4899-2271-7 deci, e. l., & ryan, r. m. (2000). the “what” and “why” of goal pursuits: human needs and the self-determination of behavior. psychological inquiry, 11, 227-268. doi: 10.1207/s15327965pli1104_01 diseth, a., & samdal, o. (2014). autonomy support and achievement goals as predictors of perceived school performance and life satisfaction in the transition between lower and upper secondary school. social psychology of education, 17, 269-291. doi:10.1007/s11218-013-9244-4 ellerbrock, c. r., & kiefer, s. m. (2013). the interplay between adolescent needs and secondary school structures: fostering developmentally responsive middle and high school environments across the transition. high school journal, 96, 170-194. evans, n. j., forney, d. s., guido, f. m., patton, l. d., & renn, k. a. (2010). student development in college: theory, research, and practice. san francisco, california: jossey-bass. garcía-ros, r., pérez-gonzález, f., cavas-martínez, f. & tomás, j. m. (2019). effects of pre college variables and first year engineering students’ experiences on academic achievement and retention: a structural model. international journal of technology and design education, 29, 915-928. doi: 10.1007/s10798-018-9466-z gale, t., & parker, s. (2014). navigating student transition in higher education: induction, development, becoming. in h. brook, d. fergie, m. maeorg & d. michell (eds.), universities in transition: foregrounding social contexts of knowledge in the first year experience (pp. 13-40). adelaide, south australia: university of adelaide press. glastra, f., & van middelkoop, d. (2018). studiesucces in het hoger onderwijs: van rendement naar maatschappelijke relevantie. delft, the netherlands: eburon. hayes, a. f. (2018). introduction to mediation, moderation, and conditional process analysis: a regression-based approach (2nd ed.) . new york, ny: guilford press. hernández, e. h., moreno-murcia, j. a., cid, l., monteiro, d., & rodrigues, f. (2020). passion or perseverance? the effect of perceived autonomy support and grit on academic performance in college students. international journal of environmental research and public health, 17. doi:10.3390/ijerph17062143 kenny, m. e., walsh-blair, l. y., blustein, d. l., bempechat, j., & seltzer, j. (2010). achievement motivation among urban adolescents: work hope, autonomy support, and achievement-related beliefs. journal of vocational behavior, 77, 205-212. doi:10.1016/j.jvb.2010.02.005 kraha, a., turner, h., nimon, k., reichwein zientek, l., & henson, r. k. (2012). tools to support interpreting multiple regression in the face of multicollinearity. frontiers in psychology, 3, 1-16. doi:10.3389/fpsyg.2012.00044 kyndt, e., donche, v., trigwell, k., & lindblom-ylänne, s. (2017). new perspectives on learning and instruction. higher education transitions theoryand research. new york, ny: routledge. langenkamp, a. g. (2010). academic vulnerability and resilience during the transition to high school: the role of social relationships and district context. sociology of education, 83, 1-19. doi:10.1177/0038040709356563 lindblom‐ylänne, s., trigwell, k., nevgi, a., & ashwin, p. (2006). how approaches to teaching are affected by discipline and teaching context. studies in higher education, 31, 285-298. doi:10.1080/03075070600680539 ljubin-golub, t., rijavec, m., & olčar, d. (2020). student flow and burnout: the role of teacher autonomy support and student autonomous motivation. psychological studies, 65, 145-156. doi:10.1007/s12646-019-00539-6 madjar, n., & cohen-malayev, m. (2016). perceived school climate across the transition from elementary to middle school. school psychology quarterly, 31, 270-288. doi:10.1037/spq0000129 mayhew, m. j., rockenbach, a. b., bowman, n. a., seifert, t. a., wolniak, g. c., pascarella, e. t., & terenzini, p. t. (2016). how college affects students: 21st century evidence that higher education works. san francisco, california: jossey-bass. niemiec, c. p., & ryan, r. m. (2009). autonomy, competence, and relatedness in the classroom: applying self-determination theory to educational practice. theory and research in education, 7, 133-144. doi:10.1177/1477878509104318 nightingale, s. m., roberts, s., tariq, v., appleby, y., barnes, l., harris, r. a., dacre-pool, l., & qualter, p. (2013). trajectories of university adjustment in the united kingdom: emotion management and emotional self-efficacy protect against initial poor adjustment. learning and individual differences, 27, 174-181. doi:10.1016/j.lindif.2013.08.004 noyens, d., donche, v., coertjens, l., & van petegem, p. (2017). transitions to higher education: moving beyond quantity. in e. kyndt, v. donche, k. trigwell, & s. lindblom-ylänne (eds.), higher education transitions. theory and research. new york, ny: routledge. nuffic. (2015). education system the netherlands. the dutch education system described. retrieved 13 march, 2018 from https://www.nuffic.nl/documents/459/education-system-the-netherlands.pdf. oolbekkink-marchand, h. w., van driel, j. h., & verloop, n. (2014). perspectives on teaching and regulation of learning: a comparison of secondary and university teachers. teaching in higher education, 19, 799-811. doi:10.1080/13562517.2014.934342 pinheiro, j., bates, d., debroy, s., & sarkar, d. (2018). nlme: linear and nonlinear mixed effects models. r package version 3.1-131.1. retrieved 2 september, 2017 from https://cran.r-project.org/package=nlme. postareff, l., mattsson, m., lindblom-ylänne, s., & hailikari, t. (2016). the complex relationship between emotions, approaches to learning, study success and study progress during the transition to university. higher education, 73, 1-17. doi:10.1007/s10734-016-0096-7 postareff, l., mattsson, m., & parpala, a. (2018). the effects of perceptions of the teaching-learning environment on the variation in approaches to learning – between-student differences and within-student variation. learning and individual differences, 68, 96-107. doi:10.1016/j.lindif.2018.10.006 prosser, m., & trigwell, k. (1999). understanding learning and teaching: the experience in higher education . buckingham, united kingdom: the society for research into higher education. richardson, m., abraham, c., & bond, r. (2012). psychological correlates of university students' academic performance: a systematic review and meta-analysis. psychological bulletin, 138, 353-387. doi:10.1037/a0026838 robbins, s. b., lauver, k., le, h., davis, d., langley, r., & carlstrom, a. (2004). do psychosocial and study skill factors predict college outcomes? a meta-analysis. psychological bulletin, 130, 261-288. doi:10.1037/0033-2909.130.2.261 roorda, d. l., koomen, h. m. y., spilt, j. l., & oort, f. j. (2011). the influence of affective teacher–student relationships on students’ school engagement and achievement: a meta-analytic approach. review of educational research, 81, 493-529. doi:10.3102/0034654311421793 sierens, e., vansteenkiste, m., goossens, l., soenens, b., & dochy, f. (2009). the synergistic relationship of perceived autonomy support and structure in the prediction of self-regulated learning. british journal of educational psychology, 79, 57-68. doi:10.1348/000709908x30439 singer, j. d., & willet, j. b. (2003). applied longitudinal data analysis: modeling change and event occurrence. new york, ny: oxford university press. stroet, k., opdenakker, m.-c., & minnaert, a. (2013). effects of need supportive teaching on early adolescents' motivation and engagement: a review of the literature. educational research review, 9, 65-87. doi:10.1016/j.edurev.2012.11.003 stroet, k., opdenakker, m.-c., & minnaert, a. (2015). need supportive teaching in practice: a narrative analysis in schools with contrasting educational approaches. social psychology of education, 18, 585-613. doi:10.1007/s11218-015-9290-1 torenbeek, m., jansen, e. p. w. a., & hofman, w. h. a. (2011a). the relationship between first‐year achievement and the pedagogical‐didactical fit between secondary school and university. educational studies, 37, 557-568. doi:10.1080/03055698.2010.539780 torenbeek, m., jansen, e. p. w. a., & hofman, w. h. a. (2011b). how is the approach to teaching at secondary school related to first-year university achievement? school effectiveness and school improvement, 22, 351-370. doi:10.1080/09243453.2011.577788 trigwell, k., & prosser, m. (2004). development and use of the approaches to teaching inventory. educational psychology review, 16, 409-424. doi:10.1007/s10648-004-0007-9 trigwell, k., prosser, m., & taylor, p. (1994). qualitative differences in approaches to teaching first year university science. higher education, 27, 75-84. doi:10.1007/bf01383761 van der zanden, p. j. a. c., denessen, e., cillessen, a. h. n., & meijer, p. (2018). domains and predictors of first-year student success: a systematic review. educational research review, 23, 57-77. doi: 10.1016/j.edurev.2018.01.001 van rooij, e. c. m., & jansen, e. p. w. a. (2018). “our job is to deliver a good secondary school student, not a good university student.” secondary school teachers’ beliefs and practices regarding university preparation. international journal of educational research, 88, 9-19. doi:10.1016/j.ijer.2018.01.005 vanthournout, g., gijbels, d., coertjens, l., donche, v., & van petegem, p. (2012). students' persistence and academic success in a first-year professional bachelor program: the influence of students' learning strategies and academic motivation. education research international, 1-10. doi:10.1155/2012/152747 vsnu. (2017). meer vwo'ers naar de universiteit. retrieved 2 august, 2018 from https://www.vsnu.nl/nl_nl/nieuwsbericht/nieuwsbericht/277-meer-vwo-ers-naar-de-universiteit.html. watkins, d. (2001). correlates of approaches to learning: a cross-cultural meta-analysis. in r. j. sternberg & l. f. zang (eds.), perspective on thinking, learning, and cognitive (pp. 165-195). mahwah: lawrence erlbaum. willett, j. b. (2004). investigating individual change and development: the multilevel model for change and the method of latent growth modeling. research in human development, 1, 31-57. doi: 10.1080/15427609.2004.9683329 zajda, j., & rust, v. (2016). current research trends in globalisation and neo-liberalism in higher education. in j. zajda & v. rust (eds.), globalisation and higher education reforms (pp. 1-23). switzerland: springer. codepen 2heinrichspublication frontline learning research special issue vol.8 no.5 (2020) 5 23 issn 2295-3159 happy-victimizing in adolescence and adulthood – empirical findings and furtherperspectives karin heinrichsa, eveline gutzwiller-helfenfingerb,brigitte latzkoc,gerhard minnameierd & bettina döringe auniversity of education upper austria, austrian buniversity of duisburg-essen, germany cuniversity of leipzig, germany dgoethe university frankfurt am main, germany ecity of hannover, bereich migration und integration, hannover, germany article received 11 june 2018 / revised 7 october/ accepted 12 december 2019 / available online 1 july 2020 abstract research on the happy victimizer phenomenon has mainly focused on preschool and schoolchildren, with a few studies also including adolescents and young adults. the main finding is that young children, despite knowing that harming somone is wrong, ascribe positive feelings to perpetrators and offer hednonistic justifications, interpreted as a lack of moral motivation. only at age 9 or 10 do almost all children ascribe negative feelings to perpetrators. according to the developmental transition hypothesis, the phenomenon should disappear in late childhood. however, reasoning patterns resembling that of the happy victimizer have been found in studies with adolescents and young adults, challenging that hypothesis. we present findings from four studies involving adolescents and young adults to give an overview of the patterns found and the measurement approaches used. finally, we critically discuss the limitations of those studies and raise some core theoretical and methodological issues that remain to be resolved, some of them being addressed in the remaining papers of this special issue. the four studies and the paper are innovative in that (a) situational factors are included in the measurement of the moral reasoning patterns; (b) new reasoning patterns are identified in the context of an extended measurement approach; and (c) the moral reasoning patterns are investigated in their own right and not used as potential explanatory variables for behaviour, as has been the main focus of research on the happy victimizer phenomenon so far. keywords: happy victimizer pattern, adolescence, adulthood, situation specifity info corresponding author email: karin.heinrichs@ph-ooe.at doi: https://doi.org/10.14786/flr.v8i5.385 1. introduction the present paper focuses on the development of morality in the context of evaluating moral rule transgressions. our conceptualisation of morality refers to the prescriptive (or normative) aspect of morality (as distinct from the descriptive aspect) and relates to a code of behaviour which – if specific requirements are met – might be endorsed by all rational individuals (gert, 2012). moreover, the moral domain can be distinguished from the domain of social conventions and personal issues (smetana, 2006). moral issues refer to behavioural choices affecting the rights and welfare of others, that is, the “right and good”, requiring humans to show benevolence and kindness towards others (gibbs, 2003), with the aim of not harming, protecting, or restoring others’ welfare (gutzwiller-helfenfinger, 2015). to realise this, we need to overcome our own, egocentric, self-interested viewpoint and take a more “objective”, moral point of view lying outside ourselves (baier, 1965). often, people experience moral conflicts in situations where following their own needs and desires would entail violating the rights and welfare of others. transgressing a moral rule therefore means that others’ rights or welfare are harmed, for example by hitting or stealing from another person. research within the happy victimizer phenomenon more closely investigates the way children, adolescents, and adults make sense of situations where a protagonist breaks a moral rule. the first empirical study investigating the happy victimizer phenomenon as such was presented by nunner-winkler and sodian (1988), although a few earlier studies had already addressed children’s emotion expectancies regarding a variety of social events (e.g., barden, zelko, duncan, & master, 1980). nunner-winkler and sodian (1988) found that four-year-old children stated that an individual has positive emotions after a rule transgression, even though they knew that the rule transgression was wrong. by contrast, eight-year-old children, after having judged the transgression as wrong, attributed negative emotions to the transgressor. according to functionalist theories of emotions, emotions are understood as internal control and evaluation systems motivating human behaviour (bretherton, fritz, zahn-waxler, & ridgeway, 1986, p. 530). therefore, nunner-winkler and sodian (1988) interpreted the pattern of judging the transgression as wrong while attributing positive emotions to the transgressor as an indicator of lower moral motivation in early compared to later childhood. while this pattern has been confirmed in numerous studies, research involving late childhood and adolescence is rare and has yielded incongruent results regarding the development of moral motivation. for instance, while nunner-winkler (2008) showed an increase of moral motivation between the ages of four and 23 years, krettenauer (2011) as well as malti and buchmann (2010) found no age differences in the strength of moral motivation during adolescence. however, none of the studies presented focused on the development of moral motivation from late childhood to adolescence and adulthood. yet there seems to be evidence that at least some adults show psychological patterns similar to the happy victimizer phenomenon (hvp) (nunner-winkler, 2007), that is, deciding in favor of a moral transgression (“defection”) in order to fulfil one’s own needs while experiencing positive emotions (and no remorse). additionally, findings from other disciplines report patterns of moral decision-making quite similar to the basic structure of the hvp. this includes research from behavioural economics (e.g., andreoni & bernheim, 2009; dana, weber & kuang, 2007; list, 2007), research on moral harzard and free riding in teams (anesi, 2009) or research on academic cheating (klein schiphorst, 2013). thus, we might hypothesise that the hvp has probably not been overcome during childhood. in this case, the associated developmental transition hypothesis (nunner-winkler, 1993; krettenauer, malti, & sokol, 2008) cannot be confirmed. moreover, the question arises to what extent the hvp emerges in later phases of moral development; and if so, what it might mean if even adults show patterns like hv, that is, feeling happy when breaking rules. oser and reichenbach (2005) pointed out that there might be other, complementary patterns of moral decision-making and emotion attributions. the authors emphasised that there are people who act in line with their deontic moral judgments and obey moral rules, but at the same time feel unhappy, for example, because they expect to suffer disadvantages compared to others who prefer to stick to their own needs in similar situations. oser and reichenbach call this pattern the “unhappy moralist” (um). other possible patterns are those of the happy moralist (hm) and the unhappy victimizer (uv) (oser & reichenbach, 2005). finally, former research indicated that the percentage of people applying patterns of moral decision-making like the hv varies depending on the quality of the conflict presented as well as on how the judgments were measured (nunner-winkler, 2013). to gain deeper insights into the conditions under which the different patterns emerge, this paper goes beyond a mere description of the patterns themselves and addresses whether and to what extent patterns of moral decision-making, in particular the hv, vary depending on situational determinants and procedures of measurement. according to nunner-winkler (2013), situational determinants can refer to (a) the type of (moral) judgment elicited; (b) the type of story (and conflict) presented; and (c) a combination of both. in this paper, results from four empirical studies are presented. some of the findings in studies 1, 3, and 4 have already been published in other contexts, but are discussed here from new angles. first, we introduce results on patterns of moral decision making and emotion attributions in adolescence (chapter 2) and adulthood (chapter 3). the aim is (a) to find out whether the hv and complementing patterns of moral decision-making and emotion attributions (re-)emerge after childhood; and (b) to ascertain that these patterns do not represent the original happy victimizer phenomenon as identified in childhood. second, we offer insights regarding situational determinants of patterns of moral-decision making and emotion attributions. third, results and methodological limitations of the studies presented are critically discussed to stimulate further research and theory building (chapter 4). finally, we introduce potential explanations of these patterns and provide first suggestions regarding the potential impact of the patterns on action and behaviour (chapter 5). 2. the happy victimizer pattern in adolescence in this section, we present two recent studies investigating patterns of moral judgments, action decisions, emotion attributions and the respective justifications in adolescents following the happy victimizer tradition. a special focus lies on the role of situational cues and the way they (may) result in situation specific patterns of moral reasoning and decision-making. the study by döring (2013) investigated the hv pattern within a standardised, representative survey with children and adolescents in fourth, seventh and ninth grade. in a questionnaire, students were presented two situations designed as moral conflicts and asked to write down their moral decisions, reasons and emotions for each situation. in the study by gutzwiller & perren (2015, 2016) a qualitative, scenario-based measure was used requiring participants to give written answers to a situation involving a passive moral temptation. 2.1 study 1: the happy victimizer between childhood and adolescence the study by döring (2013) was intended to explore to what extend the different patterns of moral decision-making, in particular the hv pattern, emerged; how these patterns could be explained in adolescence; and whether an ongoing linear growth of moral motivation could be observed. in an earlier study, krettenauer (2011) had hypothesised an increase of moral motivation with age, because, firstly, children perceive values as external and, secondly, “individual conscience becomes more salient” with age (krettenauer, 2011, p. 311). however, contrary to predictions, no increase of moral motivation had been found for students of seventh, ninth and eleventh grade (krettenauer, 2011). another study by nunner-winkler (2008) had revealed an age-related increase between childhood and early adulthood. however, that study had not included adolescents. therefore, in the study by döring (2013), a representative student sample was used to investigate the developmental pathway of moral motivation (as indicated by the proportion of happy victimizers) between late childhood and adolescence. we give a brief description of the methodology used and discuss some central results. 2.1.1 method a representative, standardised student survey was used to investigate children in the transition between childhood and adolescence. students of fourth, n=1.221, m (sd) = 10.05 (0.45), seventh, n=815, m (sd) = 13.18 (0.54), and ninth grade, n=2.891, m (sd) = 15.18 (0.58) were asked to fill in a questionnaire on two moral conflicts. the first was a vignette with the following story: “imagine you offered your bike for sale. you want to sell it for 400 euros. a young man is interested. he bargains with you and you agree on 320 euros. then he says: `sorry, i don’t have the money on me; i’ll quickly run home to get it. i’ll be back in half an hour.’ you say: ‘agreed, i’ll wait for you.’ shortly after he is gone, another customer shows up who is willing to pay the full price”. the second vignette was the following: “imagine that you have found a purse with 100 euros in it and an identity card of the owner” (malti & buchmann, 2010, p. 142). in this assessment, situational determinants referred to the use of two different stories and associated conflicts (breaking an oral contract vs. finding money that belongs to someone else). subsequent to reading the two moral conflicts, participants were asked what they themselves would do in the described situation (action decision); why they would do it (reason); and how they would feel (emotion). if a participant decided to wait for the first customer because of moral reasons (e.g., “because i promised”) and reported positive emotions, s/he was categorised as happy moralist (hm). the happy victimizer (hv) category was used if a person decided to sell the bike to the second customer because of the money and felt good about that decision. 2.1.2 results chi² tests were used to compare the percentage of hvs between age groups. for both moral conflicts the percentage of hvs was higher in ninth (27 % vs. 8.6 %) than in fourth grade (11.9 % vs. 8.7 %). seventh grade results were not significantly different from ninth but from fourth grade (22.6 % vs. 8.7 %). compared to the results for hv, the percentage of hm was higher in fourth than in ninth grade (cramer’s v [bike-conflict] = .19, p<.001; cramer’s v [money conflict] = .21, p<.001). therefore, the results showed a decrease of moral motivation, that is, more participants being categorised as hvs in the adolescent than in the childhood subsample. 2.1.3 discussion the results are limited as the study was only cross-sectional and calls for replication in a longitudinal design. however, the results indicating a decrease of moral motivation are in line with research on identity development and the age crime curve (see döring, 2013). equally, some critical considerations regarding the assessment of the hv pattern are necessary. first, döring (2013) used two moral conflicts from earlier studies (malti & buchmann, 2010; nunner-winkler, meyer-nikele & wohlrab, 2006) to assess the hv. however, the two moral conflicts were correlated at a low level only. one of the reasons for the low correlation might be that one conflict was not only between a moral rule and a personal need; the moral rule was also congruent with a judicial issue (money conflict), suggesting that situational cues played a role in the perception of the conflicts. this might have enhanced the perceived quality of the moral conflict, resulting in fewer adolescents showing the hv pattern in that conflict. also, the low correlation between the conflicts can be seen as an indication of situation specifity, that is, the specifics of a given moral conflict influencing its interpretation and the subsequent moral judgment. a second issue is the importance of others for decision making. if it is possible that an immoral decision becomes public, fewer adolescents have positive emotions while making an immoral decision in an anonymous setting. another possible explanation would be in line with motivational psychology: motivation varies and is rather a state than a trait (cf. vallerand, 2000). this explanation is underlined by the fact that moral motivation is only moderately (but significantly) correlated with moral identity and empathy (döring, 2013). regarding the relationship between moral motivation and violent behaviour, the analysis by döring (2013) revealed a significant relationship even after controlling for other important predictors of violent deviant behaviour in a logistic regression model (e.g. sex, self-control, deviant peers, intrafamiliar violence), thus confirming earlier research. 2.2 study 2: adolescents’ moral reasoning patterns in the context of a passive moral temptation research within the happy victimizer paradigm (hvp) has consistently made use of scenarios (vignettes) of moral or morally relevant situations to assess children’s, and more recently, adolescents’ and adults’, moral reasoning (e.g., nunner-winkler, 2007). scenarios have mostly involved proactive, obvious transgressions of moral rules (e.g., stealing, hitting, excluding someone, etc.) where the protagonist intentionally displays behaviours harming others (e.g., nunner-winkler & sodian, 1988). until very recently, passive moral temptations where a protagonist does not intend to transgress and only realises that s/he might do so as a result of specific circumstances (gutzwiller-helfenfinger & perren, 2015; 2016; heinrichs et al., 2015) do not seem to have been investigated. thus, getting too much change money, the situation used in our study, differs dramatically from proactively stealing money from someone. the former situation is more ambiguous and open, especially as it is not directly related to a negative duty, like, for example “thou shalt not steal”, which clearly applies to the latter. at best, it may be related to positive duties like “you should help someone in need”; and even then, the realisation that a shop assistant ending up with a negative balance can be seen as a person in need does not come easily. as negative duties carry a stronger moral obligation than positive duties (e.g., belliotti, 1981), this adds to the ambiguity of the situation. scenarios measures used within the hvp often describe a transgression as “fait accompli” and do not require participants to make deontic judgments, that is, they do not leave the situation open and ask what the protagonist should do. however, including a deontic judgment makes it possible to assess participants’ initial constructions of a given moral or morally relevant situation and to gain insight into the features of the situation that are both salient and relevant to them (cf. wainryb, brehl, & matwin, 2005). still, asking for a deontic judgment (“what should x [the protagonist] do”) does not necessarily represent the course of action chosen for oneself. therefore, it is necessary to include both a deontic judgment and own action decision in addition to a judgment of the hypothetical transgression itself (“x kept the money. is it okay or not to keep the money”) to have a fuller representation of participants’ construction of the situation. the different judgment conditions can thus be conceptualised as situational variations (cf. nunner-winkler, 2013). accordingly, the reasoning patterns resulting from the different moral judgments need to be considered in order to have a fuller understanding of adolescents’ moral meaning-making. in this study, adolescents’ moral reasoning was investigated as part of a prospective longitudinal study on teens’ use of electronic devices and social media and its relation to social behaviour (netteen; e.g., sticca, ruggieri, alsaker, & perren, 2013). in this chapter the moral reasoning patterns of adolescents responding to a passive moral temptation scenario are explored. one question of particular interest is whether we will find participants actually saying that one should / they would keep the money (self-perspective) or that keeping the money would be okay (other perspective), respectively. based on earlier research within the happy victimizer paradigm, even very young children display moral rule knowledge by saying that stealing etc. is not okay. however, the ambiguity of the situation in the present scenario might make it less clear for participants to perceive the moral obligation included. both the theoretical rationale and the findings presented here have not been published so far. 2.2.1 method 331 14-year-old swiss secondary i students (48% male) from 25 classrooms participated in the study, which included four assessment points within 2 years. data from t4 are presented. data collection took part in a classroom setting. students gave written answers to open-ended questions in a scenario involving a passive moral temptation (“change money”) with a protagonist buying a bicycle light in a shop and receiving too much change money (10 swiss francs). participants had to make three different moral judgments: (a) a deontic judgment on what the protagonist should do (keep the money, give back the money or both), justify their judgment, attribute (an) emotion(s) to the protagonist and justify their attribution(s) (deontic); (b) indicate what they themselves would do in the given situation (keep the money, give back the money or both; self-judgment), justify their decision, attribute (an) emotion(s) to themselves, and justify their attribution(s); and (c) after being told that the protagonist had transgressed the moral rule (i.e., kept the excess change) to judge the transgression (okay or not okay or both), justify their judgment, attribute (an) emotion(s) to the protagonist, and justify their attribution(s) ( “classical” happy victimizer condition). emotion attributions consisted of a response scale (happy, proud, indifferent, sad, angry, anxious, ashamed), with students marking the correct response (cf. malti, gasser, & buchmann, 2009). justifications of judgments and emotion attributions were content analysed using categories from research within the happy victimizer paradigm and research on moral disengagement. the coding process included deductive coding based on existing categorisations (e.g., bandura, barbaranelli, caprara, & pastorelli, 1996; gutzwiller-helfenfinger, 2009; menesini, sanchez, fonzi, ortega, costabile, & lo feudo, 2003) and inductive and abductive coding leading to the development of new or extensions of existing categories (e.g., perren & gutzwiller-helfenfinger, 2012). justifications were coded by two well-trained coders who were blind to the other data of the study. inter-rater reliability (18% of scenarios) was high (percentage of perfect agreement = .88). 2.2.2 results for all moral judgments (deontic, self, transgression), almost no one chose the “both” categories representing a state of indecision. therefore, the “both” categories were excluded from further analyses. in a first step, frequencies regarding judgments and emotion attributions were analysed. regarding morally appropriate judgments (in terms of following a moral rule and thereby not harming another’s rights or welfare), roughly three quarters of participants said that jan/a should give the money back (76.1 %), that they themselves would give it back (76.1%), and that it was not okay that jan/a kept the money (71.9%). of those who said that jan/a should give the money back, 79.1% attributed positive and 11.6% negative emotions, while 9.2% attributed indifference to jan/a. of those who said that they themselves would give the money back, 72.5% attributed positive and 14.2% attributed negative emotions, whereas 13.3% attributed indifference to themselves. of those who said that it was not okay that jan/a kept the money, 15.3% attributed positive and 68.2% attributed negative emotions, while 16.5% attributed indifference to jan/a. overall, 36 (10.9%) participants could be identified as displaying the “classical” happy victimizer pattern, that is, judging the hypothetical protagonist’s transgression as wrong while attributing positive emotions to the transgressor. the predominant justification of the positive emotion attributed was hedonism, mostly relating to jan/a’s having more money now, being used by 19 (52.8%) participants (out of a total of 36 participants). regarding the counterpart, that is, morally inappropriate judgments (in favour of breaking the moral rule and thereby harming another’s rights or welfare), roughly a quarter of participants said that jan/a should keep the money (23.9%), that they themselves would keep the money (23.9%), and that it was okay that jan/a kept the money (24.9%). of those who said that jan/a should keep the money 70.5% attributed positive and 5.1% attributed negative emotions, whereas 24.4 attributed indifference to jan/a. of those who said that they themselves would keep the money, 54.5% attributed positive and 9.1% attributed negative emotions, whereas 36.4% attributed indifference to themselves. of those who said that it was okay that jan/a kept the money, 64.1% attributed positive and 6.4% attributed negative emotions, while 29.5% attributed indifference to jan/a. participants displayed also a “new” victimizer reasoning pattern, that we call “happy transgressor” (okay to keep the money and positive emotions; cf. latzko & gutzwiller-helfenfinger, 2014) to distinguish it from the classical happy victimizer pattern (see table 1). overall, 44 (13.3%) participants displayed this reasoning pattern in the deontic judgment (deontic happy transgressor: “s/he should keep the money and would feel good about it”); 42 (12.7%) did so when indicating what they themselves would do in the given situation (happy transgressor self: “i would keep the money and feel good about it”); and 50 (15.1%) when judging the rule transgression (happy transgressor misdeed: “it is okay for him/her to keep the money, and s/he feels good about it”). the predominant justification of the positive emotions attributed was again hedonism, relating to jan/a or oneself having more money now: for the deontic happy transgressor, 28 (63.4%) participants (out of a total of 44), for the happy transgressor self, 25 (59.5%) participants (out of a total of 42), and for the happy transgressor misdeed, 30 (60%) participants (out of a total of 50) gave hedonistic justifications. table 1 the happy transgressor patterns and the classical happy victimizer pattern in the “change” vignette second, to investigate the relationship between the use of the classical happy victimizer pattern and the new transgressor reasoning patterns, crosstabulations were calculated. a significant relationship between reasoning patterns was found for: deontic happy transgressor with happy transgressor self, χ2(1, 331) = 177.84, p = .000, with 33 participants using both patterns; deontic happy transgressor with happy transgressor misdeed, χ2 (1, 331) = 152.93, p = .000, with 34 participants using both patterns; and happy transgressor self with happy transgressor misdeed, χ2 (1, 331) = 162.64, p = .000, with 34 participants using both patterns. no relationship was found between the happy transgressor patterns (deontic and self) and the classical happy victimizer pattern. the happy transgressor misdeed and the classical happy victimizer pattern are mutually exclusive, making analyses regarding their relationship superfluous. when the total use of immoral reasoning patterns was considered, we found that 234 (70.7%) of participants did not use any of the patterns, while 53 (16%) used one, 13 (3.9%) used two, and 31 (9.4%) used three patterns. 2.2.3 discussion using a passive moral temptation scenario, the moral reasoning patterns of adolescents judging a vignette where a protagonist (jan/a) gets too much money back were explored. about three quarters of participants made morally appropriate judgments and mostly attributed compatible emotions, saying that jan/a should give the money back and would feel good doing so, that they themselves would give the money back and feel good doing so, and that it was not okay if jan/a kept the money and that she would feel bad. still, about a quarter of participants made morally inappropriate judgments, actually saying that jan/a should keep the money, that they themselves would keep the money, and that it was okay that jan/a kept the money. emotions attributed were mostly positive, indifference was attributed in roughly a quarter or a third of the cases, while negative emotions were least frequently attributed. this is a surprisingly high proportion of morally inappropriate judgments made in favour of transgressing in combination with predominantly positive emotions (cf. heinrichs et al., 2015). a possible explanation lies in the nature of the scenario chosen: the passive nature of the moral temptation (with the protagonist not intending to transgress and being thrown into the situation) in combination with the inclusion of a positive duty (helping someone in need or giving back what is not one’s property) might have led these participants to construct the basic situation (receiving too much change) as being not or only weakly related to moral rules. the picture becomes more distinct when we consider the happy victimizer and the happy transgressor patterns. between one tenth and one sixth of participants displayed one of these reasoning patterns: the classical happy victimizer (it is not okay that jan/a kept the money but s/he felt good doing so), deontic happy transgressor (one should keep the money and would feel good doing so), happy transgressor self (participants would keep the money and feel good doing so), and happy transgressor misdeed (it is okay that jan/a kept the money and s/he felt good doing so). the predominant justification given for the attribution of positive emotions within all patterns was hedonism, mostly referring to the fact that oneself or jan/a now had more money. thus, it seems that having more money as a result of the shop assistant’s error when one did not intend to transgress in the first place was the salient and relevant construction for these participants. we argue that judging what jan/a should do versus what oneself would do versus judging whether it is okay that jan/a kept the money represent different situations or at least situational variations. actually, the differentiation between other-as-perpetrator and self-as-perpetrator (e.g., keller, lourenço, malti, & saalbach, 2003) was an important milestone in the research on moral motivation (e.g., gasser et al., 2013). moral reasoning in the context of self-as-perpetrator is seen as representing higher self-relevance and therefore a more valid indicator of moral motivation in terms of giving priority to moral over non-moral values or solutions. accordingly, it is important to investigate the relationship between the patterns. our results indicate that there were meaningful bivariate relationships between the three happy transgressor patterns, suggesting a relative stability in the use of at least two of them. conversely, no relationship was found between the deontic or self happy transgressor patterns and the classical happy victimizer pattern. it seems that making meaning of the scenario when everything is still open (what should jan/a do? what would you do?) differs from judging the transgression by jan/a (a hypothetical perpetrator) as a given fact. that self-as-perpetrator (happy transgressor self) is associated with other-as-perpetrator for the deontic happy transgressor but not for the classical happy victimizer confirms the potential crucial role of the openness of the situation, that is, before the transgression is introduced. judgments of the transgression alone offer only limited insights into individuals’ construction and interpretation of a situation. this underlines the importance of studying a wider array of moral reasoning patterns in order to gain a deeper insight into adolescents’ moral meaning-making. the results presented suggest that it is important to move beyond the happy victimizer pattern to understand adolescents’ (and potentially adults’) moral reasoning and to include not only judgments and emotion attributions (and justifications) relating to a rule transgression but also deontic and self judgments in situations including a moral temptation. 3. the happy victimizer pattern in adulthood in the previous section results on moral decision-making in adolescence confirmed that, contrary to previous assumptions based on conceptions of the hvp as a developmental transition in early childhood, not all teenagers leave happy victimizing behind. however, earlier studies indicated that the hv or similar structures emerge in adulthood, too, but mainly among people showing deviant behaviour (krettenauer, asendorpf & nunner-winkler, 2013; nunner-winkler, 2013). moreover, experimental studies where participants play dictator games reveal that a relevant percentage of adults take choices that benefit themselves to the detriment of their fellow players. only about 20 percent share in a fair manner (50/50 split) in the standard condition (fehr & schmidt, 2006; forsythe, horowitz, savin, & sefton, 1994). research about team work also confirmed that free riders (people who accept or even intend to benefit from the work of a group while contributing less themselves) are quite common (dingel, wei, & huq, 2013). whether unfair intentions or behaviour are punished or whether selfish behaviour is accepted even in the case of inequality seems to depend on situational cues (fehr & schmidt, 1999). thus, it is an open question to what extent patterns resembling the hv may appear in adulthood and how they are connected to deviant behaviour. furthermore, the studies on free riding or sharing mentioned above may point to the fact that “deviant” behaviour is not limited to situations of high moral intensity like moral dilemmas, but rather begins in situations of low moral intensity. some kinds of “deviant” behaviour or moral transgressions seem to be widely spread and accepted as “normal”, like for example violating speed limits, tax fraud, white lies or – particularly in some cultural contexts – corruption. maybe we have to acknowledge that some kinds of “immoral behaviour” are to some extent omnipresent in adults’ everday life. therefore, it seems to be fruitful to investigate the hv in adulthood to disentangle the various facets of moral functioning. in particular, we want to study how situational cues relating to different contexts cause intrapersonal variations of moral reasoning as well as emotion attributions and to address the impact of measurement methods on the identification of the hv in adulthood. 3.1 study 3 3.1.1 method in this study 271 students of economics and business education filled in paper-and-pencil questionnaires. they were confronted with four different situations, all assumed to provoke patterns of moral decision-making and embedded in an economically relevant context: buying a new tv set as a private consumer (tv sale), making a difficult decision as an entrepreneur of a newly founded start-up (start-up), being tempted while getting too much change when buying something in a tourist shop (change), or asking for a refund of (unwarranted) traveling expenses from one’s company (travel costs). participants were asked to decide between two alternative courses of action and to reason why they would choose this alternative if they had to act themselves (self judgment; condition [b] in study 2, see section 2.1.1) as well as to state how they would feel (from “good” to “bad” on a four-point likert scale). the participants were coded as displaying the hv if they decided in favour of the transgression and attributed good or rather good emotions; as unhappy victimizers (uv) if they opted for the transgression and attributed rather bad or bad emotions; as happy moralists (hm) if they opted for obeying the rule and attributed good or rather good emotions; and as unhappy moralists (um) if they decided for obeying the rule and attributed rather bad or bad emotions. 3.1.2 results among the 271 participants we found varying proportions of hv patterns across the four situations: from 31.1% percent (“start-up”), 36.6 % (“tv sale”) and 41.3 % (“travel costs”) up to 49.2% (“change”). also, the amount of hm patterns differed across situations: from 35.4% (“change”) to 47.2 (“tv sale” and “travel costs”) up to 61.8% (“start-up”). the “change” situation elicited the highest proportion of decisions towards victimizing (uv and hv: 63%), the start-up situation elicited the lowest proportion of victimizing stragegies (uv and hv: 34.7%) (for more details see heinrichs et al., 2015). moreover, we did not observe strong correlations between the situational ratings, which points to substantial intra-individual variation across situations (see also study 1 above). because of the nominal level of measurement we calculated the pearson’s contingency coefficient c as well as cramer’s v on a subsample of those cases where the agency type could be identified for all four situations (n = 254). the monte carlo method was used because some cells had fewer than five cases. table 2 cramer’s v and contingency coefficient c for all four agency types (hv, uv, hm, um) across the four situations i) these results have already been published in german (see heinrichs et al., 2015), but in this paper, they are discussed compared to the results towards the hv in adolescence (study 1 and 2). ii) *** refers to p < 0.001, ** to p < 0.01 and * to p < 0.05. in three out of six possible comparisons we found weak, but significant indicators of consistency (see table 2). 3.1.3 discussion all in all, the results confirm that up to 63% of the participants intended to trangsgress moral rules at least in certain situations and up to 49.2% would feel good or rather good to do so. thus, these findings clearly point out that happy victimizing emerged in adulthood, too, and seemed to be quite common. however, we did not observe strong correlations between the situational ratings, which points to substantial intra-individual variation across situations (see also study 1 above). it seems that the emergence of these patterns depended on the stimulating situation. thus, patterns of moral decisions making did not show interpersonal consitency across situations. however, these results are limited to situations embedded in economic contexts and therefore mainly point to negative financial consequences for other persons or a company. similar to study 2, the situations include some forms of temptation rather than systematically planned deviant behaviour. it is possible that the patterns of moral reasoning might change if other situational contexts like for example contexts of prosocialness are used or if bodily or psychological harm to the victims is included. 3.2 study 4 3.2.1 method study 4 explores how the patterns of moral reasoning and emotion attributions differ depending on the kind of moral judgment elicited: deontic judgment (“what should the protagonist do?”), self judgment (“what would you do?”), classical hv condition (rating whether a transgression already committed by the protagonist is okay or not okay). study 4 included a paper-and-pencil survey requiring students of economics and business education (n=233) to rate two different situations, both potentially provoking moral transgressions. one situation referred to the context of a start-up enterprise (see study 3, in section 3.1), the other was a slighty adapted version of the “change money” vignette used in study 2 (the classmate was replaced by a colleague, see section 2.2). as in study 2 participants had to make different moral judgments and attribute emotions for both situations. they were asked to make their decisions from three different perspectives: deontic judgment, self judgment, and classical happy victimizer condition. accordingly, study 4 explored situational variations in moral reasoning patterns relating to decision perspective and situational variations (“start-up” and “change money”). in addition to their decisions, the participants had to indicate how the protagonist (i.e., in the deontic judgment and the classical hv condition) or they themselves (self judgment condition) would feel if they acted as they decided. they had to rate emotions on a 4-point likert scale ranging from good (1) to bad (4). ratings of 1 and 2 were coded as happy, 3 and 4 as unhappy. additionally, we differed between degrees of hv reasoning patterns: if positive emotions were attributed, the «pure» hv was assigned; if rather positive emotions were attributed, the «moderate» hv was assigned. 3.2.2 results for our analyses, we were interested in the cases where participants basically argued in favour of the transgression, that is, said that the protagonist should transgress (deontic judgment), that they themselves would transgress (self judgment), and that it was okay that the protagonist had transgressed (classical hv condition). accordingly, the classical hv pattern where participants said that it was not okay that the protagonist had transgressed (e.g., kept the change money) but felt good about it was not included in our analyses. when using the classical hv condition, we found that 35.4 % of the participants said that the transgression was “o.k.” in the start-up-story, and 77.3 % did so in the change-money-story. however, among those only 9 out of 233 (4 %) attributed positive or rather positive emotions in the start-up-story thus displaying a a pattern corresponding to the happy transgressor misdeed in study 2 (see table 1). in the “change money” vignette 31 out of 220 (14.1%) participants displayed this pattern. in contrast, when asking what the participants would do (self judgment) the proportion of participants preferring the transgression was higher in the start-up-story (n=107; 46.9 %) than in the change money-story (n=30; 13.6%). when including emotion attributions, the number of transgression-friendly patterns (corresponding to the happy transgressor self in study 2) was as follows: in the start-up context 67 out of 228 students (29.4%) displayed the pattern, in the change money-story 22 out of 220 (10.0%) did so. the picture regarding transgression-friendly patterns in the deontic judgment condition was similar to that displayed in the self judgment condition: among the adult participants of study 4, the number of participants deciding in favour of the transgression in the deontic condition was 126 (54.3%) in the start-up-story and 29 (12.8%) in the change money-story. when including emotion attributions, 67 out of 126 (28.9%) participants in the start-up-story and 20 out of 29 (8.9%) in the change money-story attributed positive or rather positive emotions and thus displayed a pattern corresponding to the deontic happy transgressor pattern in study 2. table 3 mean differences for deontic and self judgment across situations (t-tests; p>0.005) to gain deeper insights into potential differences in emotions attributions between the various reasoning patterns in the three conditions (deontic judgment, self judgment, classical hv condition), paired samples t-tests were performed. this enabled us to study whether the means of the emotions attributed (1: good to 4: bad) of those persons who preferred to transgress the rule (i.e., victimise) differed significantly from the means of those who preferred not to transgress the rule (i.e., moralise) (bonferroni correction included). results showed no significant differences between deontic judgment and the self judgment, or the deontic judgment and the classical hv condition. however, contrasting the victimisers’ emotion attributions between the deontic judgment (“one should break the rule”) and the self judgment (“i would break the rule”) middle to high effect sizes were found (start-up: d = 0.68; change money: d = 1.33). conducting a priori power analyses revealed that the mean differences would have been significant in the start-up situation if the sample had been larger than 46 and in the change money if the sample had been larger than 16, respectively (see table 3). moreover, t-tests for comparing victimising in the classical hv condition (happy transgressor misdeed) with victimising in the self judgment (happy transgressor self) confirmed that significantly more negative (i.e., morally appropriate) emotions were attributed to the perpetrator in the classical hv condition (happy transgressor misdeed) than in the self judgment condition (happy transgressor self; start-up: t = 11.90, df = 64; p < 0.0001; change money: t = 22.30, df = 154; p < 0.0001). 3.2.3 discussion first of all, in study 4, the number of participants displaying the hv pattern differs between the start-up and the change money story. in this regard, these results confirm previous findings of situation specificity of moral reasoning. considering the results in more detail, we see that in the classical condition there are less hv patterns in the start-up situation than in the change money story. in contrast, in regard to deontic and self judgments we found more people preferring victimising in the start-up situation than in the change money story. a similar picture occurs among all participants preferring victimizing as well as in the subgroup of those who preferred victimizing and feel happy. thus, both situational cues of the stories as well as the method of measuring the decision towards obeying or breaking the moral rule seem to impact moral judging and emotion attributions. in the classical condition high rates of victimizing patterns emerged, but the judgments of most of the participants who preferred transgressing attributed bad or rather bad emotions and insofar the patterns have to be coded as uv. maybe attributing negative emotions to another’s transgression indicates that, if the participants have to act by themselves, they would switch from breaking to preferring to obey the moral rule and – hopefully – would feel good about it. this needs to be addressed in future research. however, maybe, these results are more a consequence of the measurement approach used than an indication of specific moral functioning. participants had to retrospectively rate another person’s behaviour that had already been enacted. the question arises whether this method of projection works better for children than adults who normally are able to take another person’s perspective. altogether, the results suggest interaction effects between stories and methods of measurement. and they encourage us to differ between types of patterns – as related to a given story and method of measurement (associated prompts)– as was done in study 2 (deontic happy transgressor, happy transgressor self, happy transgeressor misdeed, classical happy victimizer) and to develop hypotheses about the relevance and meaning of these patterns of moral decision-making and emotion attributions among adults. one direction to continue is to think about the relevance of the happy victimizer and happy transgressor patterns for acting. in study 4, we found that emotion attributions of those participants who favoured the transgression differed according to situational stimuli (start-up story vs. change money story). the rate of happy victimizer and transgressor patterns obviously varied depending on the the way the decision was measured. regarding the deontic and the self judgment conditions – both referring to decisions about fictive actions or action in the future – more participants showed transgression-friendly patterns than in the classical condition. moreover, the findings indicate – in line with results for children (keller et al., 2003; nunner-winkler, 2013) that self judgments go together with more positive emotions than deontic judgments. this seems to be plausible if we assume that self judgments represent intentions to act that include a self commitment towards one way of acting, maybe after having struggled with ambivalence or even inner conflicts (see heinrichs, kärner & reinke, this issue). 4. limitations of research within the happy victimizer paradigm four limitations are discussed relating to both the studies presented and the state of the art of research on the hv in adolescence and adulthood. 4.1 using projection to assess moral judgments and emotions former research has already pointed out that the percentage of individuals identified as displaying the hv varies depending on the methods used to assess patterns of moral decision-making as well as of emotion attributions. in particular, different results emerged if participants were asked to decide what the protagonist in an open-ended story should do (other-perspective) or what he or she himself would do (self-perspective) in that situation (nunner-winkler, 2013). nunner-winkler and colleagues originally used stories already including a transgression by a protagonist and asked children to judge that action. this method of projection seems to work well with children at the age of 4. children at that age are assumed not to be able to differ between their own and others’ perspectives (barden et al., 1980). however, there are serious doubts whether this method can be applied in adolescence (see 2.2) or adulthood (see 3.2). older participants in our own pilot interview studies came up with comments clearly indicating that they could not simply impose their own decisions and emotions upon the protagonist. moreover, they pointed out that they could not imagine how the protagonist felt because they were not that person and therefore did not know the protagonist’s thoughts or feelings. especially the results of study 2 showed that different moral judgments involving different perspectives yielded differential reasoning patterns in adolescents, some of them not described previously (the happy transgressor patterns, see chapter 2.2.2; see also gutzwiller-helfenfinger & latzko, this issue). therefore, our results confirm that self and other perspectives need to be included when studying the moral reasoning patterns of adolescents and adults. moreover, in particular the categories suggested in study 2 represent four types of patterns of moral decision-making and related emotion attributios in morally relevant situations. the three new happy transgressor patterns need to be studied more deeply. 4.2 using questionnaires to assess the hv in all our studies we used questionnaires with standardised and half-standardised questions. standardised questions were used for basic judgments (deontic, action decision, transgression, emotion attribution), while open-ended questions were used to gain insights into participants’ reasoning about these judgments and attributions. however, in studies 3 and 4 participants often gave quite short answers, making content analyses difficult. thus, it might be fruitful to conduct mixed-methods and experimental studies which allow for a more sensitive and multi-variant assessment of moral reasoning patterns while at the same time making the inclusion of larger samples possible. as we started studying the hv in adulthood, we conducted interviews and later moved on to questionnaires to investigate the context sensitivity of hv. we confirmed that hv varies intra-individually across situations. therefore, it might be interesting to again conduct interviews as part of mixed-methods studies to gain deeper insights into the way participants reason regarding judgments, courses of action, and emotion attributions, in particular, to explore the meaning participants make for the different categories as deontic happy transgressor, happy transgressor self, happy transgressor misdeed and classical happy victimizer. 4.3 relevance of patterns of moral decision-making for acting as discussed above, some of our results indicate that there are people who agree on the morally inadequate solution independent of the procedure of measurement: they claim that the protagonist should break the rule; they confirm that they themselves would also transgress; and they assess the transgression as being okay. they argue consistently in favour of the transgression. maybe they would also act accordingly in real life. however, asking for judgments, intentions or action decisions via questionnaires and interviews using hypothetical scenarios has rather low external validity when it comes to predicting moral behaviour. moreover, within hv research, issues of social desirability have been frequently raised and critically discussed. in the context of questionnaires and interviews using participants’ self-reported judgments, we only assess what participants tell the researchers what they would do and how they judge certain ways of acting, but we do not know what they really would do if they experienced similar situations in real life. moreover, in our studies motivation, volition, and emotions were not assessed directly, but inferred from self-reports. it is important to find creative ways of investigating the impact of the identified patterns of moral decision-making on actual behaviour. as already mentioned, an experimental approach including the systematic variation of situational cues and conditions may prove promising. such an approach has been used in past studies on the hvp in children (e.g., nunner-winkler & sodian, 1988). 5. research on the hv – future perspectives the findings from our studies clearly show that patterns of moral decision-making and emotion attributions do vary among adolescents and adults due to situational determinants. thus, the hvp can no longer be interpreted as a developmental stage that will be overcome during childhood. this is an important step for this field of research; and there are many options how to move further ahead and make theoretical as well as empirical progress (lakatos, 1978). first, there is a lack of research how to explain the finding that differential types of moral decision-making and emotion attributions emerge across situations, indicating intra-individual variation. the remaining papers in this special issue present theoretical approaches that represent differential theoretical perspectives reflecting whether patterns of moral deicsion-making reflect emotional (moral) development, specific cognitive structures, or are dependent on volitional processes or intentions. second, our findings confirm that patterns of moral decision-making depend on situational cues. however, the data presented call for deeper investigations of the kind of situational factors that contribute to individuals’ changes regarding judgments, decisions, and emotion attributions across situations. earlier studies for example pointed out that being treated unfairly or having a position differing from that of a relevant other may cause changes in moral decision-making (heinrichs et al., 2015). additional factors influencing the interpretation of morally relevant situations, like for example personal determinants, need to be considered in order to gain a fuller picture of the stability versus change in adolescents’ and adults moral reasoning patterns. third, there is a lack of longitudinal research to explore the relative contribution of developmental changes, personal determinants, and situational and contextual factors. accordingly, variables like metacognitive capacities, attribution styles (weiner, 2013), the moral self (krettenauer, 2011), or moral climate need to be included in future prospective longitudinal research. fourth, if we assume that people prefer to have positive feelings, and that negative feelings might indicate unfulfilled needs or inner conflicts, it will also be interesting to study changes from patterns including the attribution of negative emotions towards patterns associated with positive emotions. possibly, the attribution of negative emotions points to a motivation to change personal or situational determinants, maybe for individual development. finally, research on the hv and other patterns of moral decision-making faces the challenge of confirming whether these patterns of moral decision-making identified in the context of interviews and paper-and-pencil questionnaires are of high external validity and relevant for acting in everyday life (heinrichs et al., this issue; minnameier, this issue). therefore, an interdisciplinary perspective on research on moral decision-making and the hv should be fostered, requiring a cooperation for example between psychologists, educationalists, economists, sociologists, and philosophers. moreover, the hv could be investigated in a variety of contexts and domains of everyday (working) life to further explore similarities and differences in moral reasoning patterns across contexts and domains (see e.g., heinrichs & wuttke, 2016; minnameier, heinrichs, & kirschbaum, 2016). the question raised at the beginning, namely whether the hvp, originally identified in early childhood, at least diminishes during adolescence, focused on a developmental perspective. this starting point is relevant particularly in and for education. to discuss the implications of these results with respect to education requires us to raise both normative and empirical issues. regarding normative issues, the question arises what patterns of moral decision-making should be aimed at in educational contexts. parche-kawik (2003) analysed theoretical approaches in organisational theory, business ethics, and education, and found that they include the assumption that individuals are or ought to be able to argue on kohlberg’s stages five or six in morally relevant conflicts. however, other authors (e.g. beck 2016, minnameier, 2018; this issue) argue that acting to one’s own advantage is not only quite common, but also an appropriate way of acting if the established rules work as “moral institutions”. moreover, the hv is claimed to represent one possible way of cooperation and of implementing moral ways of acting (minnameier et al., 2016; pies, 2009). thus, apart from the empirical question how people reason in real-life moral conflicts, there is also the need for an intensive discussion on the normative question regarding the kind of argumentation that is adequate with respect to moral functioning and education. keypoints contrary to earlier theorising and research, happy victimizer reasoning patterns can consistently be found in adolescents and adults. situation specificity of patterns of moral decision-making and emotion attributions is confirmed. content analysis leads to a new set of categories regarding patterns of moral decision-making, such as the deontic happy transgressor, the happy transgressor self, the happy trangressor misdeed, and the classical happy victimizer. the findings encourage further research to explore the meaning-making underlying the patterns as well as their relevance for acting. acknowledgments we are grateful that matthias guerts provided substantial contributions to the analyses of study 4 within his bachelor’s thesis at the goethe-university frankfurt/main, germany. we are also grateful for julia vogel’s and carmen amrein’s substantial contribution to analyses in study 2 within a bachelor’s (julia vogel) and a master’s thesis (carmen amrein) at the university of teacher education of lucerne, switzerland. qualitative analyses in study 2 were co-funded by the university of teacher education of lucerne, switzerland. references andreoni, j., & bernheim, b. d. (2009). social image and the 50–50 norm: a theoretical and experimental analysis of audience effects. econometrica, 77, 1607-1636. https://doi.org/10.3982/ecta7384 anesi, v. (2009). moral hazard and free riding in collective action. social choice and welfare, 32(2), 197-219. doi: 10.1007/s00355-008-0318-8 baier, k. (1965). the moral point of view. new york: random house. bandura, a., barbaranelli, c., caprara, g. v., & pastorelli, c. (1996). mechanisms of moral disengagement in the exercise of moral agency. journal of personality and social psychology, 71, 364-374. https://doi.org/10.1037/0022-3514.71.2.364 barden, r. c., zelko, f. a., duncan, s. w. & master, j. c. (1980). children’s consensual knowledge about the experimental determinants of emotion. journal of personality and social psychology, 39, 968-976. https://doi.org/10.1037/0022-3514.39.5.968 beck, k. (2016). individuelle moral und beruf: eine integrationsaufgabe für die ordnungsethik? [individudal morality and profession: an integrative task for ordnungsethik?] in g. minnameier (hrsg.), ethik und beruf: interdisziplinäre zugänge (s. 41-54). bielefeld: w. bertelsmann verlag. belliotti, r. a. (1981). positive and negative duties. theoria, 47 (2), 82-92. bretherton, i., fritz, j., zahn-waxler, c., & ridgeway, d. (1986). learning to talk about emotions: a functionalist perspective. child development, 57, 529-548. doi: 10.2307/1130334 dana, j., weber, r. a., & kuang, j. x. (2007). exploiting moral wiggle room: experiments demonstrating an illusory preference for fairness. economic theory, 33(1), 67-80. doi: 10.1007/s00199-006-0153-z dingel, m.j., wei, w. & huq, a. (2013). cooperative learning and peer evaluation: the effect of free riders on team performance and the relationship between course performance and peer evaluation. journal of the scholarship of teaching and learning, 13(1), 45-56. döring, b. (2013). the development of moral identity and moral motivation in childhood and adolescence. in k. heinrichs, f. oser & t. lovat (eds.), handbook of moral motivation. theories, models, applications (pp. 289-307). rotterdam: sense publishers. fehr, e., & schmidt, k. m. (1999). a theory of fairness, competition, and cooperation. the quarterly journal of economics, 114, 817-868. https://doi.org/10.1162/003355399556151 fehr, e. & schmidt, k. (2006). the economics of fairness, reciprocity and altruism – experimental evidence and new theories. in s.-c. kolm & j. m. ythier (eds.), handbook of the economics of giving, altruism and reciprocity (pp. 615-691). amsterdam: elsevier. https://doi.org/10.1016/s1574-0714(06)01008-6 gasser, l., gutzwiller-helfenfinger, e., latzko, b., & malti, t. (2013). do moral emotion attributions motivate moral action? a selective review of the literature. in k. heinrichs, t. lovat, & f. oser (eds.),handbook of moral motivation. theories, models, applications (pp. 307-322). rotterdam: sense publishers. gert, (2012). the definition of morality. in e. n. zalta (ed.), the stanford encyclopedia of philosophy (fall edition). [http://plato.stanford.edu/archives/fall2012/entries/morality-definition/] gibbs, j. c. (2003). moral development and reality. thousand oaks etc.: sage. gutzwiller-helfenfinger, e. (2009). moral disengagement bei jugendlichen. kodiermanual zur auswertung des fragebogens zum moralischen verständnis . [moral disengagement in adolescents. coding manual for a moral understanding questionnaire]. unpublished manual. teacher training university of lucerne, switzerland. gutzwiller-helfenfinger, e. (2015). not unlearning to care – healthy moral development as a precondition for nonkilling. in r. bahtijaragic bach & j. e. pim (eds.), nonkilling balkans (pp. 139-169). sarajevo: university of sarajevo, faculty of philosophy & honolulu: center for global nonkilling. gutzwiller-helfenfinger, e., & latzko, b. (2020). happy victimizing in emerging adulthood: reconstruction of a developmental phenomenon? frontline learning research, 8(5), 47-69. https://doi.org/10.14786/flr.v8i5.382 gutzwiller-helfenfinger, e., & perren, s. (2016). the relationship between adolescents’ bully-victim problems and their use of mechanisms of moral disengagement in the context of passive moral temptations. paper presented in the symposium moral disengagement in the production of progression (chairs: k. runions & e. gutzwiller-helfenfinger). 22nd world meeting of the international society for research on aggression (isra), sydney (australia), july 19-23, 2016. gutzwiller-helfenfinger, e., & perren, s. (2015). adolescents’ evaluations of passive moral temptations – relations to bully-victim problems. paper presented in the symposium morality and bully-victim problems (chairs: l. kollérova & d. strohmeier). 17th european conference on developmental psychology (ecdp), braga (portugal), september 8-12, 2015. heinrichs, k. & wuttke, e. (2016). mangelnde financial literacy der kunden als moralische herausforderung beim verkauf von finanzprodukten? – eine kontextspezifische analyse im licht der happy-victimizer-forschung [lack of consumers’ financial literacy as a challenge for sales in financial insustry – a context specific analysis in light of happy victimizer research]. in g. minnameier (hrsg.). ethik und beruf – interdisziplinäre zugänge (s. 199-214). bielefeld: bertelsmann. heinrichs, k., kärner, t. & reinke, h. (2020). an action-theoretical approach to the ‘happy victimizer’ pattern – exploring the role of moral disengagement strategies on the way to action , frontline learning research, 8(5), 24-46. doi: https://doi.org/10.14786/flr.v8i5.386 heinrichs, k., minnameier, g., gutzwiller-helfenfinger, e. & latzko, b. (2015). „don’t worry, be happy“? – das happy-victimizer-phänomen im berufsund wirtschaftspädagogischen kontext [the happy victimizer phenomenon in a vocational and business educational context]. zeitschrift für berufsund wirtschaftspädagogik, 111(1), 31-55. forsythe, r., horowitz, j., savin, n., & sefton, m. (1994). fairness in simple bargaining experiments. games and economic behavior, 6(3), 347-369. keller, m., lourenço, o., malti, t., & saalbach, h. (2003). the multifaceted phenomenon of ‘happy victimizers’: a cross-cultural comparison of moral emotions. british journal of developmental psychology, 21 , 1–18. doi: 10.1348/026151003321164582 klein schiphorst, a. t. (2013). students’ justifications for academic cheating and empirical explanations of such behavior. social cosmos, 4(1), 57-63. krettenauer, t. (2011). the dual moral self: moral centrality and internal moral motivation. the journal of genetic psychology, 172(4), 309-328. https://doi.org/10.1080/00221325.2010.538451 krettenauer, t., asendorpf, j. b., & nunner-winkler, g. (2013). moral emotion attributions and personality traits as long-term predictors of antisocial conduct in early adulthood findings from a 20-year longitudinal study. international journal of behavioral development, 37(3), 192-201. doi: 10.1177/0165025412472409 krettenauer, t., malti, t., & sokol, b. (2008). the development of moral emotions and the happy victimizer phenomenon: a critical review of theory and applications. european journal of developmental science, 2, 221-235. doi: 10.3233/dev-2008-2303 lakatos, i. (1978). the methodology of scientific research programmes: volume 1: philosophical papers . cambridge university press. latzko, b., & gutzwiller-helfenfinger, e. (2014). happy victimizer im erwachsenenalter: rekonstruktion eines phänomens? [the happy victimizer in adulthood: reconstruction of a phenomenon?] vortrag an der 23. tagung des arbeitskreises moral (deutschsprachige moralforscherinnen und moralforscher), hannover, 9.-11. januar 2014. list, j. a. (2007). on the interpretation of giving in dictator games. journal of political economy, 115(3), 482-493. doi: https://doi.org/10.1086/519249 malti, t., & buchmann, m. (2010). socialization and individual antecedents of adolescents‘ and young adults‘ moral motivation. journal of youth and adolescence, 39, 138-149. doi: 10.1007/s10964-009-9400 malti, t., gasser, l., & buchmann, m. (2009). aggressive and prosocial children’s emotion attributions and moral reasoning. aggressive behavior, 35(1), 90–102. doi: 10.1002/ab.20289 menesini, e., sanchez, v., fonzi, a., ortega, r., costabile, a., & lo feudo, g. (2003). moral emotions and bullying: a cross-national comparison of differences between bullies, victims and outsiders. aggressive behavior, 29(6), 515–530. https://doi.org/10.1002/ab.10060 minnameier, g. (2020). explaining happy victimizing in adulthood – a cognitive and economic approach, frontline learning research, 8(5), 70-91. https://doi.org/10.14786/flr.v8i5.381 minnameier, g. (2018). reconciling morality and rationality: positive learning in the moral domain. in o. zlatkin-troitschanskaia, g. wittum & a. dengel (eds.). positive learning in the age of information (plato) a blessing or a curse? (pp. 347-361). wiesbaden: springer. doi: 10.1007/978-3-658-19567-0 minnameier, g., heinrichs, k., kirschbaum, f. (2016). sozialkompetenz als moralkompetenz – wirklichkeit und anspruch? [social competence as moral competence – reality or expectation?] zeitschrift für berufsund wirtschaftspädagogik, 112(4), 636-666. nunner-winkler, g. (2013). moral motivation and the happy victimizer phenomenon. in k. heinrichs, f. oser & t. lovat (eds.), handbook of moral motivation. theories, models, applications (pp. 267-288). rotterdam: sense publishers. nunner-winkler, g. (2008). die entwicklung des moralischen und rechtlichen bewusstseins von kindern und jugendlichen [the development of moral and legal awareness in childhood and adolescence]. forensische psychiatrie, psychologie, kriminologie, 2, 146-154. https://doi.org/10.1007/s11757-008-0080-x nunner-winkler, g. (2007). development of moral motivation from childhood to early adulthood. journal of moral education, 36(4), 399-414. https://doi.org/10.1080/03057240701687970 nunner-winkler, g. (1993). die entwicklung moralischer motivation [the development of moral motivation]. in w. edelstein, g. nunner-winkler & g. noam (hrsg.), moral und person (s. 278-303). frankfurt a.m.: suhrkamp. nunner-winkler, g. & sodian, b. (1988). children’s understanding of moral emotions. child development, 59, 1323-1338. doi: 10.2307/1130495 nunner-winkler, g., meyer-nikele, m., & wohlrab, d. (2006). integration durch moral: moralische motivation und ziviltugenden jugendlicher [integration through morality: moral motivation and civic virtues in adolescence]. wiesbaden: vs-verlag. oser, f., & reichenbach, r. (2005). moral resilience: what makes a moral person so unhappy. in w. edelstein & g. nunner-winkler (2005). morality in context (pp. 203-224). north holland publishing: elsevier. parche-kawik, k. (2003). den homo oeconomicus bändigen. zum streit um den moralisierungsbedarf marktwirtschaftlichen handelns. [to handle the homo oeconomicus. contribution to the discussion about the need for morality in economic markets] frankfurt/m.: peter lang-verlag. perren, s., & gutzwiller-helfenfinger, e. (2012). cyberbullying and traditional bullying in adolescence: differential roles of moral disengagement, moral emotions, and moral values. european journal of developmental psychology, 9(2), 195-209. doi: 10.1080/17405629.2011.643168 pies, i. (2009). moral als heuristik: ordonomische schriften zur wirtschaftsethik [morality as heuristics: ordnonomic papers on business ethics]. berlin: wvb. smetana, j. g. (2006). social domain theory: consistencies and variations in children’s moral and social judgments. in m. killen & j. smetana (eds.), handbook of moral development (pp. 119–154). mahwah, nj: lawrence erlbaum. sticca, f., ruggieri, s., alsaker, f., & perren, s. (2013). longitudinal risk factors for cyberbullying in adolescence. journal of community & applied social psychology, 23(1), 52-67. https://doi.org/10.1002/casp.2136 vallerand, r. j. (2000). deci and ryan's self-determination theory: a view from the hierarchical model of intrinsic and extrinsic motivation. psychological inquiry, 11(4), 312-318. wainryb, c., brehl, b. a., & matwin, s. (2005). being hurt and hurting others: children's narrative accounts and moral judgments of their own interpersonal conflicts. monographs of the society for research in child development, 70, 1–114. https://doi.org/10.1111/j.1540-5834.2005.00354.x weiner, b. (2013). ultimate and proximal (attribution-related) motivational determinants of moral mehaviour. in k. heinrichs, f. oser & t. lovat (eds.), handbook of moral motivation. theories, models, applications (pp. 99-112). rotterdam: sense publishers. knöchelmann et al frontline learning research vol.7 no. 4 (2019) 58 65 issn 2295-3159 adults’ ability to interpret covariation data presented in bar graphs depends on the context of the problem nina knöchelmanna, sabine kruegera, anita flacka, christopher osterhaus a aludwig-maximilians-universität münchen, germany article received 28 march 2019/ revised 15 may / accepted 10 october/ available online 3 december abstract the ability to correctly interpret data is an important skill in modern knowledge societies. the present study investigates adults’ ability to interpret covariation data presented in bar graphs. drawing on previous findings that show that the problem context influences the interpretation of contingency tables (grounded and concrete problems are easier than abstract ones) and based on findings from the literature on motivated reasoning (confirming problems are easier than disconfirming ones), we present n = 111 undergraduates with bar graphs in either grounded (confirming or disconfirming) or abstract contexts. our results show that only grounded problems in confirming contexts are easier than abstract ones; grounded problems in disconfirming contexts are more challenging than abstract ones. overall, the interpretation of bar graphs is difficult: even in our sample of educated college students, correct performance did not exceed 50%. our results support earlier findings regarding the context dependency of data-interpretation skills, and they suggest that relatively minor task variations have an impact on reasoners’ interpretations of bar graphs. keywords: data interpretation; bar graphs; problem context; confirming; disconfirming info corresponding author: c.osterhaus@psy.lmu.de. doi: 10.14786/flr.v7i4.471 1. introduction data interpretation is a key feature of scientific thinking, and it is an important skill—not only in schools, but also in everyday life, where people need to consider complex data when making far-reaching decisions (e.g., when making decisions in the context of elections, investments, or about medical treatments). although basic abilities in data interpretation are already present in elementary school children (koerber, mayer, osterhaus, schwippert, & sodian, 2015), even adults have difficulties to correctly interpret complex data about covariation (e.g., saffran, barchfeld, sodian, & alibali, 2016). for instance, when asked to interpret data presented in 2x2 contingency tables, reasoners frequently fail to use the correct strategy, which involves a comparison of the conditional probabilities across the rows. often, reasoners use simpler strategies: they compare the absolute frequencies across two cells (compare-two strategy; shaklee & tucker, 1980) or they try to find an anchor in the data (a simple ratio between two cells, such as 1:1 or 2:1) to which they compare the ratio between the other two cells (anchor-and-compare strategy; osterhaus, magee, saffran, & alibali, 2019). previous work has shown that reasoners’ successful interpretations of contingency tables depend on two characteristics of the task, which are the symmetry of the problem (symmetric vs. asymmetric; saffran et al., 2016) and the problem context (grounded vs. abstract; osterhaus et al., 2019). symmetric problems involve a comparison between two candidate causes (i.e., x1 leads to y, x2 leads to y), whereas in asymmetric problems, a candidate cause is compared to a control group, to which no intervention is applied (x leads to y, not-x leads to y). symmetric problems are easier, because they seem to equally draw reasoners’ attention to all four cells. comparing grounded to abstract contingency problems, research has shown that grounded problems (i.e., problems involving a concrete context and cover story) are easier, because they seem to afford a quicker access to long-term memory and the easier access of pragmatic cognitive schemas that might support reasoning (osterhaus et al., 2019). contingency tables are an effective way to present data regarding the relation between two dichotomous variables. they are, however, not the most common form of data presentation that people encounter in their daily lives. more frequent are visualizations, such as bar graphs, which people are presented with far more often, for instance in the media. bar graphs facilitate information processing by presenting covariation data in a spatial organization that allows reasoners to quickly grasp the relation between two variables, and also, they allow to outsource cognitive processes to an additional perceptual route (hegarty, 2011). research shows that already young kindergarten and elementary school children can read off information about relations from graphs, resulting in an update of existing beliefs in response to the data (koerber, osterhaus, & sodian, 2017). although bar graphs are a common and effective way of presenting data, adults’ ability to interpret this form of data presentation is not well understood, and it is unclear if task variations (like the ones observed in the study of contingency tables) have a similar impact on people’s interpretation and strategy choice on these problems. the present study, therefore, investigates adults’ ability to interpret data that is presented in bar graphs. following earlier findings regarding the influence of context on the interpretation of contingency tables (osterhaus et al., 2018), we present participants with bar graphs that are embedded in either a grounded or abstract context, and that can only be solved correctly by using the conditional-probabilities strategy (comparing conditional probabilities across conditions). based on prior findings from the study of contingency tables (osterhaus et al., 2019), we hypothesized that grounded problems afford correct interpretations relative to abstract ones. drawing on the literature on reasoning biases in data interpretation (chinn & brewer, 1993) and on motivated reasoning (klaczynski, 2001), our study also explores if the beneficial effect of grounded problems is stable across confirming and disconfirming contexts. it is reasonable to assume that grounded problems are only easier when they are confirming, that is, when they lead to an activation of prior knowledge that is in line with the data presented. research has shown that people are less likely to take seriously an implausible covariation between two factors when there is no plausible causal explanation (koslowski, 1996), and so we expect disconfirming contexts to be, in turn, more difficult. disconfirming contexts present reasoners with causal relations that seem implausible given their prior knowledge, which may result in a distortion, rather than an affordance, of their correct interpretation. 2. methods 2.1 participants the sample comprised 111 university students who were (in majority) recruited from two large german research universities (n = 111; 91 females, 19 males; 1 participant did not disclose their gender). informed consent was obtained from all participants. participants were recruited via social media and through advertising in the universities. participants received either course credit for their participation or they entered a lottery to win a voucher for a bookstore. 2.2 design we used a within-subjects design with three groups (confirming vs. disconfirming vs. abstract context) in which participants interpreted a set of 3 x 3 (9) bar graphs. a post-hoc power analysis conducted with g*power (faul, erdfelder, lang, & buchner, 2007) showed a power of 1-β = 1.00 (α = 0.05, effect size f = 0.71). 2.3 materials we used three different bar graphs for each of the three conditions (confirming vs. disconfirming vs. abstract contexts). all problems were presented in the asymmetric form to provide the highest difficulty level. the frequencies displayed in the bar graphs were taken from a prior study with contingency tables (saffran et al., 2016; see table 1). we only included cell frequencies that resulted in problems that could exclusively be solved by using the conditional-probabilities strategy (but not any other less sophisticated strategy). this way, participants’ correct solutions are indicative of their use of the conditional-probabilities strategy, which is the only strategy that guarantees the correct interpretation independently of the exact cell frequencies. the problems should, therefore, be of comparable difficulty—especially because we chose cell frequencies < 1,000 in order to keep computing demands at an acceptable level. the bar graphs were designed in microsoft® excel (for an example, see figure 1) and they were presented in an online questionnaire. for every bar graph, participants were asked to decide (based on the data presented) if a given intervention (present, absent) has a positive, negative, or no effect at all on a given dichotomous outcome variable (effect present, absent). for the confirming problems, participants interpreted data that showed a positive relation between tutoring (in maths, physics, or chemistry) and grade improvement (i.e., the number of students whose grades did or did not improve); for the disconfirming problems, participants interpreted data that showed a negative or no relation between doing sports (jogging, cycling, or swimming) and the improvement of people’s fitness level (i.e., the number of people whose fitness level did or did not improve); and for the abstract problems, participants interpreted data that showed a positive, negative or no effect between an abstract candidate cause (x, y, or z) and an abstract outcome variable x (i.e., the number of instances where x is present or absent). the specific problem contexts in the grounded condition were chosen because we expected them to elicit equally strong expectations regarding the direction of the causal relation (e.g., tutoring improves grades, doing sports improves fitness level). table 1 overview of frequencies used in the problems and correct solutions per problem note. the contingency tables under ‘problem’ display the treatment condition in the columns (treatment +/ no treatment -) and the outcome in the rows (improvement +/ no improvement -). 2.4 procedure participants responded to the nine items in an online questionnaire (social science survey). confirming, disconfirming, and abstract problems were presented in a randomized order; this order was the same for all participants. there were no minimum or maximum time boundaries for inclusion in the study (only three participants took longer than one hour to click through the problems). for the remaining participants, the completion of the survey took on average 6.5 minutes (sd = 7.1). 2.5 coding participants were given 1 point per problem when they offered the correct interpretation of the data, and 0 points when they did not. a sum score was calculated, so that participants could obtain between 0 and 3 points for each of the three conditions. the correct response for each item was determined by comparing the conditional probabilities regardless of the size of the difference. if the two conditional probabilities for ‘improved’ were equal and p (improved | treatment) = p (improved | no treatment), there was no effect. if p (improved | treatment) > p (improved | no treatment), there was a positive effect; if p (improved | treatment) < p (improved | no treatment), there was a ‘negative effect’. figure 1. example item (grounded, confirming context). 3. results 3.1 correct solutions on average and across all conditions, participants provided correct solutions on 4.23 of the 9 problems (sd = 2.48, min = 0, max = 9). performance differed substantially across items, ranging from 11% to 68% correct (see table 1). 3.2 influence of context: confirming vs. abstract vs. disconfirming a repeated-measures analysis of variance (anova) was conducted to test for the influence of context. the assumption of sphericity was met (mauchly’s w = 0.951, p = 0.06). the data were not distributed normally, but the repeated-measures anova tends to be robust against violations of the assumption of normality. the analysis revealed significant differences between the three conditions, f(2, 222) = 73.66, p < .001, η2 = 0.40. a planned contrast (problem context) indicated no significant difference between the grounded (confirming and disconfirming) and abstract problems, t(110) = -1.718, p = .96. a set of post-hoc tests, however, revealed significant differences between all three conditions: the confirming condition (m = 1.9, sd = 1.1) was significantly easier than the abstract condition (m = 1.5, sd = 0.9), t(110) = 5.11, p < .001, cohen’s d = .48, which in turn was significantly easier than the disconfirming condition (m = 0.8, sd = 1.0), t(110) = 7.59, p < .001, cohen’s d = .72 (see figure 2). figure 2. the average amount of correct solutions (out of a maximum of 3) per condition (confirming vs. abstract vs. disconfirming contexts). error bars display standard errors (se). 4. discussion bar graphs are a commonly-used visualization to present data about covariation. despite their common use, the findings of the present study show that they are difficult to be correctly interpreted, even by educated undergraduate students. this is especially true when the data are presented in abstract or disconfirming contexts. the influences of problem context (grounded vs. abstract) have not previously been investigated in reasoners’ interpretation of bar graphs. findings from the study of people’s interpretation of contingency tables suggest that grounded problems are easier than abstract ones because they may support reasoning by enabling a quicker access to long-term memory and the easier access of pragmatic cognitive schemas. the present study shows that indeed grounded problems lead to a higher number of correct responses. extending on previous work, our results, however, show that this finding only holds when a confirming context is used (i.e., a context that presents participants with a causal relation that is plausible given their prior knowledge). when, in turn, a disconfirming context is used (i.e., a context that presents participants with a causal relation that is implausible given their prior knowledge), correct performance declines. the decline in performance for disconfirming contexts is in line with the literature on reasoning biases in scientific reasoning (chinn & brewer, 1993) and motivated reasoning (klaczynski, 2000), and it suggests that participants are guided by their prior knowledge when solving problems like the ones presented in the current study. on all disconfirming problems, the use of a simple strategy (e.g., compare two) resulted in an interpretation that seemed plausible given participants’ prior knowledge. on the confirming problems, the use of the same simple strategy resulted in an interpretation that seemed implausible. it seems likely that participants generally used less sophisticated strategies first and did not invest further cognitive resources if their judgement was in line with their expectations based on prior knowledge. when, however, an initially not sophisticated interpretation resulted in an implausible, conflicting outcome, participants were likely inclined to pay more attention to the data and to explore various interpretations. compared to abstract contexts, the use of grounded problems thus has a general impact on performance. whether or not this influence is beneficial or detrimental, however, depends on the specific context that is used and its fit with participants’ prior knowledge. correct performance was, in the present study, overall low: the undergraduate students in our sample solved only an average of 47% of the problems correctly. this percentage is higher than the chance of correct guessing and it is similar to undergraduate students’ interpretation of contingency tables, for which comparable numbers of correct solutions were found (e.g., 54% in osterhaus et al., 2019). although correct performance is difficult to compare across studies (in our study, only the conditional probabilities strategy led to correct responses; in other studies, some problems can be solved with simpler strategies), the present findings suggest that the interpretation of bar graphs is, in contrast to popular belief, not substantially easier than the interpretation of data that is presented in contingency tables. a limitation of the present study is that the conditional probabilities were not equal across conditions. that is, in order to confirm or disconfirm participants’ prior beliefs, we used different causal directions and strengths across conditions. future work should keep these probabilities constant to assure that they are not confounding factors that may potentially drive the effect. all of the items that we used in the present study, however, could only be solved with the conditional-probabilities strategy. research has shown that reasoners’ use of this strategy is quite consistent across contingency tables with different conditional probabilities (osterhaus et al., 2019), and so we do not expect the discrepancies in conditional probabilities to have caused the substantial differences in correct solutions (and strategy use) between the conditions. the difficulties in interpreting bar graphs that we documented in the present study are a finding that needs to be stressed. people are often presented with bar graphs (e.g., in the media, in patient brochures, in investment information, etc.), and overestimating reasoners’ ability to correctly interpret this information is likely to result in poor decision making. future research should, therefore, address the question of how to foster this important ability, thereby increasing scientific literacy and helping people to draw correct inferences from diverse forms of data, including bar graphs. keypoints the correct interpretation of bar graphs is difficult, even for educated college students. relatively minor tasks variations have an impact on reasoners’ interpretation of the data. grounded problems in confirming contexts afford correct interpretations relative to abstract contexts. grounded problems in disconfirming contexts distort reasoners’ interpretations. acknowledgments the research was conducted by nina knöchelmann, sabine krueger, and anita flack in the context of a bsc psychology seminar at the ludwig-maximilians-universität münchen (lmu), under the direction of christopher osterhaus. the research was supported by a “teaching through research” grant provided by lmu (forschungsförderung “lehre@lmu”), awarded to matthias stadler and christopher osterhaus. we are grateful to all participants for their friendly collaboration and support of this research. references chinn, c. a., & brewer, w. f. (1993). the role of anomalous data in knowledge acquisition: a theoretical framework and implications for science instruction. review of educational research, 63, 1-49. doi:10.3102/00346543063001001 faul, f., erdfelder, e., lang, a.-g., & buchner, a. (2007). g*power 3: a flexible statistical power analysis program for the social, behavioral, and biomedical sciences. behavior research methods, 39, 175–191. doi:10.3758/brm.41.4.1149 hegarty, m. (2011). the cognitive science of visual-spatial displays: implications for design. topics in cognitive science, 3(3), 446–474. doi:10.1111/j.1756-8765.2011.01150.x klaczynski, p. a. (2000). motivated scientific reasoning biases, epistemological beliefs, and theory polarization: a two-process approach to adolescent cognition. child development, 71, 1347–1366. doi:10.1111/1467-8624.00232 koerber, s., osterhaus, c., & sodian, b. (2017). diagrams support belief revision. frontline learning research, 5(1), 76–84. doi:10.14786/flr.v5i1.265 koerber, s., mayer, d., osterhaus, c., schwippert, k., & sodian, b. (2015). the development of scientific thinking in elementary school: a comprehensive inventory. child development, 86, 327–336. doi:10.1111/cdev.12298 koslowski, b. (1996). theory and evidence: the development of scientific reasoning. cambridge, ma: mit press. osterhaus, c., magee, j., saffran, a., & alibali, m. w. (2019). supporting successful interpretations of covariation data: beneficial effects of variable symmetry and problem context. quarterly journal of experimental psychology, 72, 994–1004. doi:10.1177/1747021818775909 saffran, a., barchfeld, p., sodian, b., & alibali, m. w. (2016). children’s and adults’ interpretation of covariation data: does symmetry of variables matter? developmental psychology, 52, 1530–1544. doi:10.1037/dev0000203 shaklee, h., & tucker, d. (1980). a rule analysis of judgments of covariation between events. memory & bars display standard errors (se). hanin et van nieuwenhoven publication frontline learning research vol.6 no. 2 (2018) 39 65 issn 2295-3159 teaching the problem-solving process in a progressive or a simultaneous way: a question of making sense? vanessa hanina, catherine van nieuwenhoven a a université catholique de louvain, belgium article received 27 september 2017/ revised 12 december / accepted 10 august / available online 30 august abstract over the past two decades, the perennial low success rates of elementary students in math problem solving and the difficulties experienced by teachers in helping their students with this type of task has become quite a hot topic. in response, several instructional interventions aiming to develop an expert and reflexive approach to problem solving have been designed. however, these interventions are based on two contrasting teaching approaches, either teaching the components of the problem-solving process at the same time or teaching them one at the time. a meticulous analysis of the literature indicates that studies that have compared these two teaching approaches have focused primarily on undergraduate students. moreover, they have mainly been assessed in terms of cognitive outcomes. yet, recent studies stress the importance of analyzing the cognitive, motivational and emotional processes involved in problem-solving learning together in order to gain a full understanding of the process. addressing these limitations is essential to enhance our understanding of problem-solving learning and to design more effective interventions. this paper focuses on this issue by investigating whether teaching the problem-solving process in all its complexity or one component at a time is preferable in terms of cognitive, motivational and emotional outcomes. this issue is handled for both novice and expert solvers. data were gathered among 267 upper elementary students. findings showed that both teaching approaches support the shortand long-term acquisition of cognitive problem-solving strategies, regardless of the student’s profile. however, beneficial emotional and motivational outcomes occur only when the problem-solving process is taught in all its complexity, that is, makes sense for the learner. novice solvers made less use of the help-seeking strategy and persisted more. keywords: mathematics problem-solving; emotion regulation strategy; heuristic strategy; novice and expert; teaching practices info corresponding author mail: vanessa.hanin@uclouvain.be doi: 10.14786/flr.v6i2.333 1. introduction in a societal context that lays increasing emphasis on the need for analytical and complex task-resolution skills (depaepe, de corte, & verschaffel, 2010; ntcm, 2010), it is important to question their development and acquisition in the academic context. in terms of mathematics education, curriculum designers have stressed that students develop meaningful mathematical skills, motivate themselves and learn how to deal appropriately with situations encountered in their everyday life through problem-solving tasks (ntcm, 2010). however, both research studies (demonty, blondin, matoul, baye, & lafontaine, 2013; demonty & fagnant, 2014) and international tests (oecd, 2016) have brought to light students’ perennial low success rates in math problem solving. the alarming international test scores stimulated researchers to design educational interventions that aim to develop an expert and reflexive approach to problem solving. most educational research (blum, 2011; de corte, 2012; fagnant & jaegers, 2018; tzohar-rozen & kramarski, 2014) agrees that the development of an expert and reflexive approach to problem solving occurs through the mastery of specific heuristic strategies embedded in an overall metacognitive approach. yet, different assumptions concerning the importance of facilitating realistic meaningful experiences, as opposed to facilitating cognitive processing, lead to different pedagogical approaches. on the one hand, research on the development of an expert and reflexive approach to mathematical problem solving conducted in the field of mathematics instruction stresses the importance of giving all students, whether they are experts or novices, realistic, meaningful, challenging and complex problem-solving tasks, supporting a “simultaneous” teaching approach (blum, 2011; depaepe et al., 2010; van dooren, verschaffel, greer, de bock, & crahay, 2010). on the other hand, studies conducted in the field of cognitive psychology and, more specifically, anchored within the framework of cognitive load theory, have shown that, for novice learners, it is preferable to teach one element at a time to avoid overwhelming their working memory (blayney, kalyuga, & sweller, 2010; pollock, chandler, & sweller, 2002). thus, if research in mathematics instruction supports approaching the problem-solving process in its full complexity, the work carried out in the field of cognitive psychology supports the opposite approach. in addition, there is a lack of understanding of the role of affective components and motivational aspects in the problem-solving process, as both of the traditions mentioned above focus (more or less widely) on cognitive aspects. such understanding therefore appears crucial for the integration of the two approaches. and this is especially the case given that recent studies have pointed out the necessity of considering cognitive, emotional and motivational dimensions when dealing with teaching and learning issues (ahmed, van der werf, kuyper, & minnaert, 2013; pekrun, 2014; tzohar-rozen & kramarski, 2014). the present study aims to extend previous research by overcoming these limitations and enhancing our understanding of the effects on cognitive, motivational and emotional dimensions of teaching the problem-solving process to upper elementary students. more precisely, this paper aims to test the effects of these two teaching approaches (“simultaneous” versus “gradual-all together”) on the frequency of use of heuristic strategies, the valence of the emotions felt, the kind of emotion regulation strategies used and the level of persistence. in addition, given that novice and expert solvers may differ regarding their learning process, a comparison between the two teaching approaches was made for these two learner profiles (muir, beswick, & williamson, 2008; zimmerman & campillo, 2003). 2. theoretical foundations and relevant empirical work first, we examine the problem-solving process. then, the relation of cognitive load theory to the two teaching approaches studied in the present paper is described. subsequently, the role played by emotions in problem-solving tasks and the necessity of developing functional emotion regulation strategies are examined. this section closes with discussion of one important motivational dimension for problem-solving learning, that is, persistence. 2.1 the problem-solving process nowadays, scholars acknowledge that the development of expertise in mathematical problem solving requires the reconceptualization of mathematical problems as exercises in mathematical modeling; that is, considering the statement of a problem as the description of a situation in everyday life that can be modeled mathematically (blum & niss, 1991; de corte, verschaffel, & masui, 2004; fagnant, demonty, & lejonc, 2003). mathematical modeling is viewed as a complex and cyclic process involving a number of phases. in this regard, based on a thorough analysis of the existing literature in the problem-solving field (blum, 2011; fagnant & demonty, 2005; fuchs, fuchs, prentice, burch, hamlett, et al., 2003 ; lucangeli, tressoldi, & cendron, 1998), hanin and van nieuwenhoven (2016) identified, eight heuristic strategies of particular importance in solving non-routine problems that delineate the problem-solving process (figure 1): (1) building a representation of the problem, that is, specifying the relevant contextual and numerical information, the unknown (what we are looking for), and the relationships between these elements via a drawing, a table or a reformulation (situation model); (2) estimating the answer a priori, that is, approximating the solution of the problem by identifying the kind of response that is needed (a measure, a price, etc.), by rounding the numbers, by imagining what the solution is not or by giving an order of magnitude; (3) using one’s knowledge, namely, identifying the mathematical structure of the problem, asking whether one has already solved a similar problem and how one did it (identification of the mathematical model); (4) planning, that is, breaking down the problem into steps; (5) executing the necessary calculations, that is, translating each step of the solution plan by an appropriate mathematical operation and executing it to arrive at a final mathematical result; (6) verifying the relevance of the operations chosen, ensuring compliance with instructions and checking the accuracy of the calculations; (7) interpreting the outcome, that is, making sure that the solution makes sense with regard to the problem statement (plausibility of the solution) and that the solution is congruent with the a prior estimate; (8) for a satisfactory interpretation, the solution is communicated. the two teaching approaches compared in the present study are based on this problem-solving process. it should also be noted that the linear arrangement gives a timeline to be used in the teaching of the solving process. the problem-solving process must be seen as cyclical and constituted of a back-and-forth between the different heuristic strategies. figure 1 diagram of the problem-solving process 2.2 the contribution of cognitive load theory cognitive load theory views the resolution of complex, non-routine mathematical problems as high in element interactivity, where the elements are the heuristic strategies (pollock et al., 2002; van merriënboer & sweller, 2005). in such tasks students must process the different heuristic strategies in working memory simultaneously for learning to occur, since these heuristic strategies are all tightly linked. as many elements must be processed in working memory simultaneously, there is high intrinsic cognitive load . yet we know that working memory is limited. usually the individual’s cognitive architecture constructs schemas in order to handle the problem of working memory overload. schemas organize a large number of elements and take their interactivity into account, while acting as a single element, that is, without overloading working memory. however, these schemas result from a preliminary construction. and therein lies the problem. the construction of such schemas is the result of the simultaneous processing of all of the elements and therefore includes the elements’ interactivity. however, due to working memory limitations, novice learners can only process a few of the elements in working memory simultaneously, which prevents the construction of schemas, that is learning, from taking place. thus, as processing all of the elements simultaneously in working memory is not effective (in the case of elements that are high in interactivity), there remains the option of processing element by element in working memory and thereby eliminating the interactions among them, but at the cost of a reduced understanding. on this point, empirical evidence has shown that it is more beneficial for non-expert learners to begin with a “one-by-one element” approach and, once each element has been examined in isolation, to deal with the full set of elements in an interactive way (what we call the “gradual-all together” approach) rather than to present, from the outset, the material in its full complexity (what we call the “simultaneous” approach) (pollock et al., 2002; van merriënboer & sweller, 2005). regarding expert learners, empirical studies have highlighted that they perform equally well regardless of the teaching approach used, because they already possess relevant problem-solving schemas (pollock et al., 2002). 2.3 the contribution of emotion theory complex math tasks are known to generate negative emotions (hanin, & van nieuwenhoven, 2018; op’t eynde, de corte, & mercken, 2004). however, tasks lacking sufficient challenge also generate negative emotions, mostly boredom (pekrun, goetz, daniels, stupnisky, & perry, 2010). thus, depending on the teaching approach implemented, both “novice” and “expert” problem-solvers might be affected by negative emotions. negative emotions are known to be detrimental to learning. more precisely, studies have reported that negative emotions foster the use of rigid, detail-oriented, and analytical approaches, divert a part of the available cognitive resources from the task, and promote external regulation (pekrun, 2014). not only do students feel negative emotions when dealing with complex math problems but, in addition, they do not regulate these emotions (de corte, depaepe, op’t eynde & verschaffel, 2011). on this point, hanin et al. (2017) highlighted six strategies used by upper elementary school children to regulate their negative emotions when solving math problems. “negative self-talk” involves focusing on the negative aspects of the situation, by dramatizing them, by constantly thinking them over or by convincing oneself that they are beyond one’s control. “dysfunctional avoidance” involves task avoidance, despite the fact that task completion is beneficial in the long run. “emotion expression” refers to the social sharing of one’s emotions. “task utility self-persuasion” involves convincing oneself of the personal utility of the task despite the fact that the task generates unpleasant emotions. “help seeking” concerns seeking peer and teacher assistance. finally, “brief attentional relaxation” involves releasing attention for a few seconds through distraction or relaxation. this strategy covers two sub-strategies, namely, a positive form of distraction and physical relaxation. of these six, students considered the first three to be maladaptive and the other three were judged to be adaptive (hanin, gregoire, mikolajczak, fantini-hauwel, & van nieuwenhoven, 2017). however, although scholars agree on the crucial role played by emotions in problem-solving learning and performance, to our knowledge, no study so far has examined the effect of cognitive training programs on emotions and emotion regulation strategies. 2.4 the contribution of motivation theory alongside emotions, a motivational variable of particular interest when solving mathematical problems is persistence (montague & applegate, 2000), that is, “the behavioral strength that fosters, despite the impediments encountered, the continuation of the actions required by the engagement” (brault-labbé & dubé, 2008, p. 731, our translation). studies conducted by montague and applegate (2000) suggested that task difficulty has a direct influence on middle school students’ persistence. they postulated that “some students "shut down" cognitively when they perceive problems as difficult or when information-processing demands seem excessive” (p. 225). although little research has examined this concept in the context of compulsory schooling, on the basis of the above findings, we assume that learning all the heuristic strategies within the same problem-solving task, that is, as interactive, will raise more difficulties than learning them one at a time. consequently, the “simultaneous” approach might undermine students’ persistence, where the “gradual-all together” approach may have no effect or even support students’ persistence. nevertheless, no study has examined the effect on students’ persistence of specific cognitive teaching approaches to problem solving. 3. aims and hypotheses this study investigates whether taking into account motivational and emotional dimensions leads to the same results as those observed in studies comparing the same two teaching methods but from a purely cognitive perspective. in other words, this study seeks to clarify the following question: “are the principles of cognitive load theory still valid when motivational and emotional dimensions come into play?”. this general issue is dealt with through two research aims. the first aim seeks to clarify whether it is better for learning to handle the problem-solving process in its full complexity – the “simultaneous” approach – or to teach one heuristic strategy at the time the “gradual-all together” approach. studies undertaken within the framework of cognitive load theory have stressed that the “simultaneous” approach consumes much more of the student’s cognitive resources than the “gradual-all together” approach. moreover, studies have also shown that complex tasks give rise to negative emotions. therefore, we expected that the “gradual-all together” group would display a better understanding of the heuristic strategies taught, persist longer in the face of difficulties, feel fewer negative emotions, use appropriate emotion regulation strategies and perform better in problem solving than the “simultaneous” group and the control group (traditional approach) . the second aim examines whether it is appropriate to adopt a different teaching approach according to the learner’s level of expertise in problem solving. consistent with cognitive load theory, we hypothesized that the “gradual-all together” approach would lead to better learning outcomes for “novice” problem-solvers as compared to the “simultaneous” approach and the traditional approach. more precisely, we expected “novices” in the “gradual-all together” group to benefit more from the training program than their peers in the “simultaneous” and control groups. with regard to “expert” problem-solvers, as tasks lacking sufficient challenge are a source of negative emotions (pekrun, 2014; pekrun et al., 2010), we supposed that the “simultaneous” approach would suit them better. these hypotheses were tested through the implementation of a cognitive training program in mathematical problem-solving (section 4.2). as mentioned in the introduction, this comparison fits into a larger project. by examining the role of the motivational and emotional aspects of the problem-solving process, this paper also seeks to integrate otherwise disconnected lines of inquiry and, thereby, to contribute to advancing existing knowledge in the learning research field. it seems crucial to better understand the processes at work during mathematical problem-solving tasks in order to provide practitioners with the most adaptive solutions. 4. method 4.1 participants a sample of 267 upper elementary students took part in the first part of the present study. they came from seven french-speaking belgian schools located in different cities and had different socio-economic backgrounds and levels of performance. the “gradual-all together” group was made up of 86 students (m age = 10.8; sd = 0.89), 48.8% of which were girls; the “simultaneous” group consisted of 79 students (m age = 10.5; sd = 0.66), of which 53.2% were girls. with respect to the control group, it was composed of 102 students (m age = 10.7; sd = 0.87), of which 45.1% were girls. for the second part of the present study, only students with a “novice” or “expert” problem-solver profile, defined on the basis of their problem-solving performance at time 1 (pretest), were sampled. students with an average score below 0.5 out of 1 were called “novice” problem-solvers whereas those with an average score above or equal to 0.8 out of 1 were labelled “expert” problem-solvers. this selection was consistent with the teachers’ assessments. in each group, 43 novice problem-solvers and 12 expert problem-solvers participated in this second study. 4.2 training program the two teaching approaches examined in the present study are based on a similar training program aiming at developing among students an expert and reflexive approach to problem solving (description available in appendix a). with respect to the two teaching approaches tested, as described above, they differed in how the heuristic strategies were taught. the “gradual-all together” approach consisted in teaching one heuristic at a time. in each problem, between one and two heuristics were addressed. it was only when all the heuristic strategies had been examined individually that the learner had to implement, within the same problem, the full set of heuristics, that is, to address them as interactive. conversely, in the “simultaneous” approach, the heuristic strategies were all taught in the first problem while solving it. then, for each subsequent problem, the student had to apply the full set of strategies in a flexible way. during the same period, with a similar frequency, the students in the control group worked on the statement of the same problems. however, unlike the teachers of the “gradual-all together” and of the “simultaneous” groups, the teachers of the control group received no methodological instructions. the problems were therefore handled according to the teacher’s usual instructional practices for problem solving. the intervention extended over five weeks at the rate of two non-routine problems a week plus three weeks of pretest, post-test, and follow-up assessment (see table 1). in addition, note that, for the sake of ecological validity, the training programs were implemented by the regular classroom teachers. table 1 illustration of the research design. note. w1 = week 1; w2 = week 2, etc. as outlined by durlak and dupre (2008), one cannot interpret the results of the implementation of a training program without first ensuring that it has been delivered as planned. therefore, the fidelity of implementation of the training program was evaluated. first, a check-list that contained step-by-step instructions was provided for each lesson. as the teacher completed a step, he or she had to check it off. examination of the checklists showed that teachers completed 95% of the steps as prescribed. second, teachers were asked to keep a diary with their feelings, students’ responsiveness, potential amendments, and any other piece of information considered to be relevant. on this point, while the teachers reported being convinced of the relevance of the training program, they also pointed out that it was too intense. this observation was seen in particular from the teachers of the “simultaneous” group, who reported: “i think that what we did has had an impact, i mean i saw a change, changes in behavior, such as greater ease getting to work, greater autonomy, etc. but, we were very restricted in terms of time and thus i think that it did not leave them the time to digest the information because it was too fast, i think that it was the main problem. this must be done in the longer term”. third, the teachers received a half-day’s training during which the components of the training program were outlined. a manual containing the description of the training program components and the lesson plans as well as detailed examples of anticipated correct representations, procedures and solutions was given to each teacher. finally, meetings were organized to exchange positive experiences and difficulties encountered as well as to take stock of the past lessons and to plan those to follow. 4.3 measures first of all, it is important to point out that the three groups completed all of the measures presented below on three occasions: prior to the intervention, immediately after it, and then six weeks later. problem-solving performance was assessed by means of a performance test made up of three non-routine problems. this test was designed on the basis of our expertise in mathematics teaching and on fagnant and demonty’s (2005) textbook. the students’ performance was appraised by a global score obtained by averaging their scores for the three items on a binary scale (0 = wrong answer, 1 = right answer). the test took on average 30 minutes to complete. it is noteworthy that for each measurement point, the same three mathematical structures were used to design the problem statements, and only the presentation of the problem was modified. in addition, students were asked to indicate all their reasoning and calculations on their sheet, not to erase anything, but to cross out if necessary. heuristic strategies were measured on the basis of students’ written products. more precisely, we scrutinized their pre-test, post-test, and follow-up tests for traces of the application of five of the eight heuristic strategies taught, namely, building a representation, using one’s knowledge, planning, checking the outcome and the procedure, and interpreting the outcome. a score for each heuristic strategy was computed according to a binary scale (0 = missing; 1 = present). note that the presence of each heuristic strategy was measured and not the accuracy. persistence was appraised by means of an adapted and translated version of the “attitude toward mathematics survey” scale (fredricks & mccolskey, 2012), one of the few existing instruments offering a persistence scale that is isolated from behavioral engagement. the adapted version measures persistence in mathematical problem solving through 8 items (e.g. “when i have trouble understanding a math problem, i re-examine it until i understand it”) rated on a 4-point likert scale (1 = (almost) never to 4 = (almost) always). the internal consistency of the global score was satisfactory to good (pretest: α = .65, posttest: α = .87, follow-up test: α = .81). the emotions experienced by the students while solving a mathematical problem were evaluated through a questionnaire presenting facial expressions. these latter included positive emotions (enjoyment, pride, relief), and negative emotions (boredom, fear, anger, hopelessness, shame, worry, frustration, and nervousness) most frequently experienced by elementary and secondary students when dealing with problem-solving tasks (op’t eynde et al., 2004; pekrun, goetz, & frenzel, 2005). students were asked to indicate to what extent they felt each emotion when solving a math problem using a 5-point likert scale (1 = never to 5 = always). the internal consistency of the global score for positive emotions (pretest: α = .78, posttest: α = .74, follow-up test: α = .74) as well as that of the global score for negative emotions (pretest: α = .82, posttest: α = .82, follow-up test: α = .87) was good. emotion regulation strategies were appraised using the children’s emotion regulation scale in mathematics (cers-m) designed and validated by hanin et al. (2017). this questionnaire has shown good psychometric properties with belgian upper elementary students. it consists of 14 items, rated on a 4-point likert scale (ranging from 1= (almost)never to 4 = (almost) always) and targets six strategies used by 5th and 6th graders to regulate their emotions when solving math problems, namely, task utility self-persuasion (e.g. “even if i do not like solving math problems, i tell myself that it is important to do so in order to be able to understand them and thereby to succeed”), help-seeking (e.g. “i ask the teacher to help me to solve the problem”), brief attentional relaxation (e.g. “i put down my pencil for a few seconds and stretch my arms”), emotion expression (e.g. “i tell my neighbor that the problem makes me angry, sad, hopeless, or bored”), negative self-talk (e.g. “i tell myself that it is terrible not being able to solve the problem and that i am sure that it only happens to me”), and dysfunctional avoidance (e.g. “in order not to experience an unpleasant moment, i tell myself that i will solve the problem later”). the cers-m subscales showed acceptable internal consistency for the three measurement times in the present sample with scale reliabilities ranging from .64 to .84. 4.4 analysis to achieve our two research goals, three procedures were carried out. first, mixed-model group (“gradual-all together” vs. “simultaneous” vs. control) x time (time 1 vs. time 2 vs. time 3) repeated measures analyses of variance (anovas) were performed on each measure, with time as the within-subject factor and group as the between-subjects factor. this gave us a general overview of the variables for which the three groups were distinguishable. second, these first results were refined by examining, for all of the variables under consideration, short-term changes (between time 1 and time 2), long-term changes (between time 1 and time 3) and post-intervention changes (between time 2 and time 3). repeated measures anovas were therefore performed. third, for each of the variables presenting a significant short-term, long-term or post-intervention development, paired t-tests were performed in order to characterize the development of each group. we used the bonferroni correction for the t-tests in order to avoid a potential alpha error inflation (field, 2013). we report here the corrected p-values. in addition, in order to have a more accurate understanding of the effects, we reported the effect sizes. the latter were computed on the basis of cumming’s (2012) recommendations, that is, the calculation of cohen’s d to which we applied a correction to remove the effect size’s bias. cohen’s d was calculated using the standard formula: 5. results 5.1 question 1: in terms of learning, would it be better to teach non-routine problem solving according to a “simultaneous” approach or a “gradual-all together” approach? in order to check for any baseline differences at pretest between the three groups, a univariate anova was performed for the variables under consideration. the three groups differed regarding only the emotion expression strategy (see appendix b). 5.1.1 overall effect of the problem-solving intervention significant group x time interactions were found for persistence (f(2,264)=4.244,p=.002,partial η^2=.06). with respect to heuristic strategies, findings revealed significant group x time interactions for three out of the five strategies measured, namely, building a representation of the problem (f(2,264)=25.778,p <.001,partial η^2=.19); using one’s knowledge (f(2,264)=12.971,p <.001,partial η^2=.10) and planning (f(2,264)=10.588,p <.001,partial η^2=.09). regarding the emotion regulation strategies, two of them presented a significant group x time interaction, namely, help-seeking (f(2,264)=3.643,p=.006,partial η^2=.04) and negative self-talk (f(2,264)=2.747,p=.028,partial η^2=.03). however, it is noteworthy that while the effect sizes are moderate to large for the heuristic strategies, they are quite small for the emotion regulation strategies. 5.1.2 shortand long-term effects and change dynamics with respect to short-term changes (between time 1 and time 2), the three groups differed for persistence; for three heuristic strategies, namely, building a representation of the problem, using one’s knowledge and planning; and for two emotion regulation strategies, help-seeking (significant trend) and negative self-talk (see appendix e). a finer analysis showed that, contrary to our expectations, only the “simultaneous” group stood out regarding the persistence variable (t(78)=-4.816,p<.001,d=.78) by displaying a significant increase. however, this finding must be put into perspective, given the relatively small effect size. with respect to heuristic strategies, the three groups presented a significant increase in the use of the strategy of building a representation of the problem. on this point, the “gradual-all together” group (t(85)=-10.184,p<.001,d=.1.57) and the “simultaneous” group (t(78)=-13.711,p<.001,d=2.09) displayed larger effect sizes as compared to the control group (t(101)=-3.494,p=.003,d=.42). with regard to the planning strategy, both the “gradual-all together” group (t(85)=-5.809,p<.001,d=.92) and the “simultaneous” group (t(78)=-6.538,p<.001,d=1.02) displayed a significant improvement. the same results were observed regarding the strategy of “using one’s knowledge”, that is, there was a significant increase for both the “gradual-all together” group (t(85)=-5.706,p<.001,d=.99) and the “simultaneous” group (t(78)=-7.633,p<.001,d=1.21). as far as emotion regulation strategies are concerned, it turned out that the “simultaneous” group resorted less frequently to the help-seeking strategy at time 2 than at time 1 (t(78)=3.155,p=.006,d=.40). furthermore, while a significant group x time interaction was found regarding the negative self-talk strategy, fine-grained t-test analyses revealed that none of the three groups showed significant development. with regard to long-term changes (between time 1 and time 3), four out of the six significant short-term differences found at time 2 compared with time 1 were also significant over the longer term, at time 3 (see appendix e). specifically, the same three heuristic strategies that showed short-term changes also presented significant increases at time 3. the same behavior was observed for the three strategies, that is, a significant increase for both the “gradual-all together” group (building a representation of the problem: t(85)=-7.671,p<.001,d=1.08; using one^' s knowledge: t(85)=-5.023,p<.001,d=.81; planning: t(85)=-4.815,p<.001,d=.77) and the “simultaneous” group (building a representation of the problem: t(78)=-12.972,p<.001,d=1.88; using one^' s knowledge: t(85)=-6.189,p<.001,d=.98; planning: t(85)=-5.251,p<.001,d=.79). finally, with regard to emotion regulation strategies, results indicated a significant decrease in the use of the help-seeking strategy by both the “gradual-all together” group (t(85)=3.341,p=.003,d=.40) and the “simultaneous” group (t(78)=5.440,p<.001,d=.69). however, it is noteworthy that the effect sizes are rather small. in order to capture solely what happened after the intervention and in doing so to distinguish clearly between short and long-term changes, the difference between time 2 and time 3 was investigated (see appendix e). on this point, it appeared that the “gradual-all together” group made substantially less use of the heuristic strategy of building a representation (t(85)=3.428,p=.003,d=.38). nevertheless, the effect sizes are very tenuous. finally, a significant decrease in the use of the emotion expression strategy by the “gradual-all together” group must also be noted (t(85)=2.621,p=.033,d=.25). 5.1.3 discussion our first research question investigated whether it is better for learning to teach the problem-solving process in its full complexity, that is, to teach all of the problem-solving heuristic strategies within the same problem-solving task (the “simultaneous” approach), or to teach one heuristic strategy at a time and only after that to implement them simultaneously within the same task (the “gradual-all together” approach). findings suggest that solving two non-routine mathematical problems weekly without specific methodological instructions is already effective, as reflected in the control group’s increased use between time 1 and time 2 of the ‘building a representation of the problem’ heuristic strategy. however, this change did not persist after the end of the training program (no significant difference between time 1 and time 3). in this respect, hanin and van nieuwenhoven (2016) have shown that this heuristic is one of the few to be traditionally taught in math classrooms, which is confirmed by the descriptive statistics at time 1 (see appendix b). consequently, this strategy is not new for the students; they have already practiced it. on this point, scholars have shown that repeated practice and confrontation with the learning material promote knowledge and skills assimilation and internalization (anderson, 1981; piaget, 1978). thus, when students become familiar with the representation heuristic strategy, which occurs faster thanks to previous practice, they may no longer feel the need to write their representation down and may prefer to do it mentally. teaching problem solving according to a “gradual-all together” approach appeared to be more effective at the cognitive, metacognitive and emotional levels than the traditional approach. at the cognitive level, results highlighted a short-term significant increase in the use of three heuristic strategies, namely, building a representation of the problem, using one’s knowledge, and planning. while the improvement in the two last strategies persisted over time, a decrease in the use of the ‘building a representation of the problem’ heuristic strategy was recorded between time 2 and time 3. this observation supports the familiarization-internalization assumption put forward earlier. in effect, through the training program proposed, not only did students have the opportunity to implement this strategy many times, they also received information about how to implement it, when it is the most convenient strategy, and what it is used for (the www & h rule; veenman, van hout-wolters, & afflerbach, 2006). this method of instruction is known to promote students’ understanding and familiarization with the learning material (tzohar-rozen & kramarski, 2017; veenman et al., 2006). as a result, we may wonder why the other two strategies, namely, using one’s knowledge and planning, were not also internalized by the students. in this regard, let us mention that the ‘building a representation of the problem’ strategy is relevant for each problem. therefore, students had implemented it on many occasions. moreover, descriptive statistics at time 1 (see appendix b) showed that of the three heuristic strategies, ‘building a representation of the problem’ was the only one that had really been exploited in the traditional classrooms before the beginning of the training program. consequently, students had had more opportunities to practice it. conversely, the planning strategy will not be invoked if the problem requires only one or two stage(s) to be solved. with respect to the ‘using one’s knowledge’ strategy, as claimed by fuchs et al. (2003), this will be fully used only at the point when the learner has encountered problems presenting different mathematical structures, has had the opportunity to identify these structures and has attempted these types of problems several times in order to build a typology of the problems’ mathematical structures. at this stage, it is thus more the premises of the strategy that are observed. at the emotional level, the help-seeking strategy appeared to be used to a lesser extent in the long run by the students in the “gradual-all together” group. such a finding would suggest that the “gradual-all together” approach, through appropriate tooling and scaffolding, fostered among students a more responsible and autonomous approach to problem solving, in other words, a self-regulated approach. in effect, in such an approach, students rely less on their teacher and peers (help-seeking strategy) (allal, 2007). nonetheless, contrary to our supposition, the “gradual-all together” approach did not enhance students’ level of persistence. the “simultaneous” approach to problem solving stood out as beneficial not only in terms of the cognitive and metacognitive aspects but also at the emotional and motivational levels. in effect, a short-term increase in the use of three heuristic strategies, namely, building a representation of the problem, using one’s knowledge, and planning, as well as maintenance over time of the level reached were observed among the students of the “simultaneous” group. thus, unlike their counterparts in both the control and the “gradual-all together” groups, the students in the “simultaneous” group seemed to still need to write down their representation of the problem at time 3. this observation, although surprising at first sight, does not invalidate the familiarization-internalization hypothesis. as reported by the teachers, the intensity of the training program, especially in its “simultaneous” version, may not have given students the opportunity to assimilate and internalize entirely the heuristic strategies taught. in this respect, several scholars have pointed out that the duration, the intensity and the frequency of the intervention’s activities impact the findings (durlak, weissberg, dymnicki, taylor, & schellinger, 2011; greenberg, domitrovich, graczyk, & zins, 2005). in addition, the teaching approach itself may, in a complementary way, explain this observation. in the “simultaneous” approach, students’ awareness was raised to deal with the various heuristic strategies in a coordinated way. more precisely, in this approach the representation strategy is viewed as a step determining all the others, that is, the step on which the subsequent heuristic strategies are built. consequently, writing down one’s representation may facilitate the continuation of the problem-solving process. in addition, a short-term decrease in the use of the help-seeking strategy as well as a maintenance over time of this lower level were observed among the “simultaneous” group. contrary to the “gradual-all together” group, for which the decrease appeared only at time 3, in the “simultaneous” group, the same decrease was already observed at time 2. so, contrary to our supposition, approaching the problem-solving process in its full complexity is more beneficial than a step-by-step approach. in other words, it is not enough to equip the learner adequately; the problem-solving approach must make sense for him/her. in this respect, not only did the “simultaneous” approach lead to faster changes in the development of self-regulated behaviors as compared to the “gradual-all together” approach, it was also associated with a short-term increase in persistence. this finding suggests that students are able to handle and to overcome the difficulties encountered, which in turn supports the hypothesis that the “simultaneous” approach contributes to the development of self-regulated behaviors. it follows from this first investigation that the “simultaneous” approach stands out as the most promising one in that it impacts the cognitive, metacognitive, emotional and motivational dimensions of problem-solving learning. however, is the “simultaneous” approach to problem solving appropriate for all profiles of students? 5.2 question 2: is it appropriate to adopt a different teaching approach according to the learner’s level of expertise in problem solving? findings regarding novice problem-solvers are presented first, followed by those for expert learners. let us mention in advance that there were no baseline differences between the three groups of novice problem-solvers (see appendix c). 5.2.1 overall effect of the problem-solving intervention on novices significant group x time interactions were found for three heuristic strategies, namely, building a representation of the problem (f(2,126)=12.830,p <.001,partial η^2=.18), using one’s knowledge (f(2,126)=4.651,p=.001,partial η^2=.07), and planning (f(2,126)=5.054,p=.001,partial η^2=.08). additionally, significant interactions were also found regarding two emotion regulation strategies, namely, help-seeking (f(2,126)=4.949,p=.001,partial η^2=.09) and negative self-talk (f(2,126)=3.484,p=.009,partial η^2=.06). finally, the persistence dimension also showed a significant interaction (f( 2,126)=3.294,p=.013,partial η^2=.09). 5.2.2 shortand long-term effects and change dynamics regarding novices regarding short-term changes, the three groups of novice problem-solvers differed in terms of three heuristic strategies (building a representation of the problem, using one’s knowledge, and planning). in addition, the three groups were also distinguishable on three emotion regulation strategies (that is, task utility self-persuasion, help-seeking, negative self-talk) as well as on persistence (see appendix f). in-depth analyses highlighted a substantial improvement for the three heuristic strategies presenting a significant increase, for both the “gradual-all together” group (building a representation of the problem: t(42)=-6.525,p<.001,d=1.36; using one’s knowledge: t(42)=-3.925,p<.001,d=.98; planning: t(42)=-4.702,p<.001,d=1.17) and the “simultaneous” group (t(42)=-10.898,p<.001,d=2.31; t(42)=-4.611,p<.001,d=.99; t(42)=-4.370,p<.001,d=.94; respectively). as regards emotion regulation, a significant decrease in the use of the help-seeking strategy was recorded among the “simultaneous” group (t(42)=3.204,p=.009,d=.59). unexpectedly, the control group displayed a significant decrease in the use of the negative self-talk strategy (t(42)= 2.727,p=.030,d=.44). however, the magnitudes of the effects regarding emotion regulation strategies are rather weak. in addition, note that although repeated measures anovas indicated a significant interaction regarding the task utility self-persuasion strategy, a finer-grained examination by means of t-tests showed that none of the three groups presented a significant development on this dimension. finally, and contrary to our expectations, only the “simultaneous” group displayed a significant rise in persistence (t(42)=-4.026,p=.003,d=.98). with regard to long-term changes, four out of the six significant differences found at time 2 compared with time 1 were also significant at time 3 (see appendix f). first, the two experimental groups displayed a significant improvement regarding the use of heuristic strategies. more precisely, both the “gradual-all together” group (t(42)=-6.561,p<.001,d=1.14 ; t(42)=-3.415,p=.003,d=.74 ; t(42)=-4.760,p<.001,d=1.10;respectively) and the “simultaneous” group (t(42)=-8.545,p<.001,d=1.70 ;t(42)=-4.128,p<.001,d=.86 ; t(42)=-3.553,p=.003,d=.71 ; respectively) displayed a significant improvement in the use of three heuristic strategies, namely, building a representation of the problem, using one’s knowledge, and planning. in addition, regarding emotion regulation strategies, the “simultaneous” group used the help-seeking strategy much less (t(42)=5.509,p<.001,d=.88). with respect to what happened after the intervention, although repeated measures anovas indicated significant group x time interactions regarding persistence and the negative self-talk strategy, a finer-grained examination via t-tests revealed that there were no significant differences between time 2 and time 3, for any of the three groups. regarding expert problem-solvers, there were no baseline differences between the three groups except for the emotion expression strategy (see appendix d). 5.2.3 overall effect of the problem-solving intervention regarding experts significant group x time interactions were found for two heuristic strategies, namely, building a representation of the problem (f(33,2 )=5.072,p=.002,partial η^2=.28), and planning (f(33,2 )=3.026,p=.026,partial η^2=.19). note that interactions with a significant trend were observed for both the strategy of using one’s knowledge (f(33,2 )=2.533,p=〖.051〗^t,partial η^2=.16) and persistence (f(33,2)=2.841,p=.07,partial η^2=.49). 5.2.4 shortand long-term effects and change dynamics regarding experts as regards short-term changes, the three groups of expert problem-solvers differed on the same cognitive and motivational variables as the novice problem-solvers, that is, persistence, building a representation of the problem, using one’s knowledge and planning (see appendix g). however, unlike novices, the three groups of expert learners were not distinguishable on the emotional dimension. both the “gradual-all together” group (t(11)"=-10.757," p<.001"," d=3.96) and the “simultaneous” group (t(11)=-6.280,p<.001,d=2,47) displayed a significant improvement in regard to the heuristic strategy of building a representation of the problem. the same findings were found regarding the strategy of using one’s knowledge ("gradual-all together\" " group: t(11)=-3.985,p=.009,d=1.66; “simultaneous” group: t(11)=-5.063,p<.001,d=1.92). with respect to the planning strategy, a substantial improvement was observed within the “simultaneous” group (t(11)=-4.710,p=.001,d=2.03). in addition, a significant improvement in persistence was recorded among the “simultaneous” group (t(11)=-4.977,p=.048,d=1.84). with respect to long-term changes, the four significant differences found at time 2 compared with time 1 were still significant at time 3 (see appendix g). as regards the heuristic strategies, both the “gradual-all together” group and the “simultaneous” group used the same three strategies significantly more: building a representation of the problem ("“gradual-all together” group: t" (11)=-4.183,p=.006,d=1.83;”"simultaneous” " group ": t" (11)=-8.848,p<.001,d=2.99), using one’s knowledge ("“gradual-all together” group: t" (11)=-3.40,p=.021,d=1.35;”"simultaneous”" group ": t" (11)=-3.644,p=.012,d=1.38) and planning ("gradual-all together” group: t" (11)=-3.191,p=.027,d=1.12;”"simultaneous” " group ": t" (11)=-2.809,p=.048,d=1.38). finally, a significant increase in persistence was observed among the “simultaneous” group (t(11)=-4.314,p=.006,d=.74). 5.2.5 discussion our second research question examined whether it is productive for the teacher to adopt a different teaching approach (“gradual-all together” vs “simultaneous”) according to the learner’s level of expertise in non-routine problem solving. contrary to what we expected, our findings revealed that while the “gradual-all together” approach and the “simultaneous” approach are equally beneficial regarding the cognitive and metacognitive dimensions, the second approach is more effective with regard to the emotional and motivational aspects of problem-solving learning, regardless of the student’s profile. first, with respect to the cognitive dimension, novice and expert problem-solvers in both approaches displayed a significant increase and maintenance over time of the level reached for three heuristic strategies (building a representation of the problem, using one’s knowledge, and planning). however, only these three heuristic strategies showed this positive development. several assumptions may be put forward to explain why the heuristic strategies of “checking the outcome(s) and the procedure” and “interpreting the outcome” did not experience greater impact by the intervention. with respect to the former, checking is both time-consuming and cognitive resource-consuming in that the student must repeat calculations (outcome checking) and read over his work (procedure checking). the “interpreting the outcome” heuristic strategy, according to the descriptive statistics, presented a steady but not significant increase between the three times of measurement. this slower growth may be partially explained by teachers’ beliefs. nowadays it is widely acknowledged that teachers’ beliefs influence their teaching (beswick, 2006; van der sandt, 2007; wilkins, 2008). in this respect, several scholars have shown that a substantial proportion of teachers, when confronted with non-routine problems for which a realistic answer is expected, display a strong tendency to exclude realistic considerations; in other words, they believe that realistic considerations have no place in math classrooms (depaepe et al., 2010). on this basis, we hypothesize that if, before presentation of the problem-solving process to the teachers, we had carried out a critical analysis of their misconceptions, the heuristic strategy of “interpreting the outcome” would have been used significantly more by their students. however, the positive growth in the use of this heuristic strategy within the three groups suggests that the nature of the problems proposed (requiring students to make sense of the outcome), had already brought about a change, although a slower one than if the problems had been preceded by an analysis. in a complementary way, these two strategies of checking and interpreting are located at the end of the problem-solving process. in this respect, teachers of both the “simultaneous” and the “gradual-all together” groups reported that the training program schedule was very intense, leaving little time for students to master each of the heuristic strategies. consequently, it is possible that the last heuristic strategies were less developed. second, as for emotion regulation strategies, only novice learners experienced significant changes. on this point, it appeared that novices in the “simultaneous” group made less use of the help-seeking strategy. this finding suggests that, as previously mentioned, the “simultaneous” approach supports the development of more autonomous management of the problem-solving process. moreover, a short-term significant decrease in the use of the “negative self-talk” strategy was observed among the novices in the control group. however, this decrease did not last over time. if this observation may sound a little surprising at first sight, as previously mentioned, scholars have shown that the repeated practice of an activity or of a set of knowledge improves the learner’ skills and performance (anderson, 1981; fayol, 2006) and, as a result, diminishes anxiety and emotional internalization behaviors (hampel, meier, & kümmel, 2008; kanfer, ackerman, & heggestad, 1996). thus, the simple weekly processing of two mathematical problems seems enough to reduce significantly the internalization of negative emotions. in our view, the intensity of both the “simultaneous” and the “gradual-all together” approaches may explain why such a decrease was not significant for these two approaches. further, we hypothesize that the short duration of the intervention accounts for the non-maintenance over time of the decrease in “negative self-talk” observed within the control group. third, concerning the motivational dimension, a short-term significant increase in persistence was observed among both novice and expert “simultaneous” learners. although this improvement faded after the end of the training program for the former, it persisted over time for the latter. this finding may be explained by both the intensity and the short duration of the training program. more vulnerable students, such as novice problem-solvers, may not have had sufficient time to deeply assimilate the heuristic strategies taught. consequently, at the end of the training program, they may not have felt better equipped to solve mathematical problems and their persistence fell back to its baseline level. additionally, the lack of challenge or, at least, the low level of challenge involved in the “gradual-all together” approach may account for the stability of the persistence level observed among both novice and expert problem-solvers experiencing this approach. in this connection, a study conducted among seventh grade students underscored a positive and quite strong relationship between persistence in problem-solving tasks and a taste for challenging tasks (malmivuori, 2006). similarly, wolters and rosenthal (2000) showed that enhancing eighth grade students’ interest in working on a task by making it more challenging or meaningful increased their persistence for the task. in short, consistent with the work done on mathematics instruction, it is clear that a training program that approaches problem solving in its full complexity, and thereby puts the focus on an understanding of the process, is cognitively, metacognitively, motivationally and emotionally more fruitful for both the novice and the expert problem-solver than the “gradual-all together” approach. 6. general discussion and conclusion for decades, problem solving has constituted a real stumbling block both for students and for teachers who report having difficulty helping their students with this type of task (fagnant, dupont, & demonty, 2016). while many scholars have developed training programs that aim to develop an expert and reflexive approach to problem solving among elementary and secondary students, they have recommended different teaching approaches without identifying the most effective one (blum, 2011; de corte et al., 2004; fagnant & demonty, 2005; tzohar-rozen & kramarski, 2017). yet, in order to improve the teaching of non-routine problem solving and so improve students’ problem-solving learning and performance, it is important to compare the effectiveness of these various approaches. the present paper contributes to advancing the existing knowledge about problem-solving instruction by examining whether it is more fruitful to teach the problem-solving process in all its complexity (the “simultaneous” approach) or one heuristic strategy at a time (“the gradual-all together” approach). first, our findings highlight that, while both learning approaches support the acquisition of cognitive strategies over the shortand the long-term, if student motivation and a positive emotional rapport with problem-solving tasks are taken into account, the “simultaneous” approach, and thus, the maintenance of complexity, is more beneficial, for both novice and expert students. in this sense, our results support the work done in mathematics instruction that stresses the importance of proposing realistic, meaningful, challenging and complex problem-solving tasks. as these findings contradict current classroom practices, they are of particular importance. on this point, research has shown that in order to facilitate task completion for their students, teachers reduce complex tasks to micro-tasks, especially for students with a “novice” profile (demonty & fagnant, 2014; depaepe et al., 2010). these micro-tasks require students to apply procedures that are meaningless for the task requested. for example, the task of making a representation of a problem does not make sense per se, unless the student then goes on to solve the problem afterward. so, this study draws attention to the fact that the difficulties experienced by students in problem solving have more to do with the inadequacy of the tasks proposed to prepare them for managing complexity than with the management of the complexity itself. concretely, it emphasized that, for a novice to become proficient at solving complex tasks, such as mathematical problems, requires repeated confrontation with such tasks. in this respect, ericsson’s theory of deliberate practice (ericsson, 2008; ericsson, krampe, & tesch-römer, 1993) has shown, through several empirical studies, that considerable practice, in the sense of the number of hours devoted to the practice of the competence that one wishes to acquire, is a prerequisite for the automation of such competence. a practical spin-off would be to increase the practice of solving complex tasks in classes. a second contribution of our study to advancing existing knowledge regards heuristic strategies. it follows from our results that the strategies taught are not acquired at the same pace by the learner. some strategies, less familiar, based on inadequate belief or requiring more cognitive resources, take longer to be integrated. this information is critical for designing more effective training programs, as each strategy occupies a specific and central place in the problem-solving process. a study conducted with students of the same age emphasized that to increase more rapidly students’ use of the “checking” and the “interpreting” strategies, it is necessary to add emotional and motivational support to the cognitive and metacognitive intervention (hanin, & van nieuwenhoven, 2018b). this observation is in line with self-regulated learning theories, which postulate that motivational beliefs and emotions play a key role not only in initiating the learning process but also in sustaining the learner’s efforts throughout the process (boekaerts, 2011; zimmerman, 2011). a third contribution pertains to the emotional aspects of problem-solving tasks. our findings reveal that differences regarding emotional variables between the three groups are quite limited. this study highlights that nurturing one aspect of self-regulation (in the present case, the cognitive one) has little effect on the other ones (here the emotional aspect). so, it would seem that to induce emotional regulation among learners, it is necessary to explicitly teach emotional knowledge and skills. this echoes and supports a recent exploratory study conducted among fifth graders by tzohar-rozen and kramarski (2017). in this way, the present study adds nuance to the empirical studies conducted so far, which have shown that learners who receive metacognitive support for problem-solving tasks display greater general motivation as well (hoffman, 2010; kramarski & gutman, 2006). this means, from a conceptual point of view, that if the processes involved in the regulation of cognition share common points with those involved in the regulation of motivation, they are distinct from those entangled in the regulation of emotions. this observation makes it possible to refine and add nuance to current literature on the subject. fourth and last, we used an original method that consists in establishing a dialogue between two research fields that deal with similar issues but that are not usually involved in an interdisciplinary context with the purpose of advancing existing knowledge about problem solving and allowing new empirical insights. priolet (2014) talked about an “integrative theoretical framework” to designate “the mobilization of works related to mathematics instruction and to both the psychology of learning and of development” (p. 60, our translation). more precisely, in addition to being based on an instructional analysis, our findings underline the necessity of presenting to students teaching-learning situations that take into account their cognitive, emotional and motivational functioning and the contextual features in which the learning takes place in order to be truly functional (maury, 2001; priolet, 2014). the present study confirms that we cannot offer students a training program that has “just” been thought of “mathematically”. so, not only does adopting such an integrative perspective allow for better understanding of the processes at work during mathematical problem-solving tasks and thereby for drawing firmer conclusions for practical guidance, but it is also essential for understanding problem-solving competence in all its complexity. in addition, this study draws attention not only to the importance of providing learning activities that make sense to the learner but also to the difference between “the learner’s academic fulfillment” and “the learner’s academic performance”. in other words, our findings suggest that developing students’ heuristic strategies, emotion regulation strategies and motivation, on the one hand, and increasing his/her performance on the other, are not always compatible processes. the weight placed by our politicians on both national and international tests is a good reflection of current concerns. however, the perennial low success rates, mentioned in the introduction, suggest that this may not be a very productive avenue. maybe it is time to ask whether schools should be more concerned about educating “achievers” or long-term, engaged learners. while the present study yields promising results for both the educational and research perspectives, several limitations that call for further investigations must be noted. first, several changes did not persist, took place quite slowly or were fairly weak. this suggests the need for a training program that is less intense (where the lessons are more spaced out in time), of longer duration (where the learner has more opportunities to implement what he/she has learned in new problem situations) and that directly addresses the emotional dimension (by including lessons on emotions and emotion regulation strategies). such a training program might lead to better uptake of all the heuristic strategies taught, to better emotional regulation, and consequently to better performance. second, as this is, to our knowledge, the first study to investigate these questions at the level of compulsory education, it would be interesting to replicate it with other samples in order to strengthen the stability of the present findings, especially as it partially questions the assumptions of cognitive load theory. third, as previously mentioned, the study would benefit from a more fine-grained measure of the heuristic strategies. on this point, in order to have a better understanding of the link between the heuristic strategies and the performance score, adding a measure of the accuracy of the implementation of each heuristic strategy would be enlightening. fourth, in the present study expert and novice solvers were assigned in terms of high and low performance, as is the case in many studies of novice and expert learners (bassok 2003; brand et al., 2003; muir et al., 2008). however, defining expertise in problem solving exclusively in terms of performance is restrictive. consequently, it would be interesting in future studies to more thoroughly conceptualize these two learner profiles. fifth and last, it is noteworthy that the results regarding the expert problem-solvers depend on the sample size. this was rather small, which makes it necessary to replicate the present study with a bigger sample of expert problem-solvers. disclosure of interest the authors declare that they have no conflicts of interest concerning this article and approve the final article. keypoints examines two contrasting approaches to teaching the problem-solving process; shows that both approaches support the acquisition of cognitive strategies; identifies meaningfulness of approach as key for emotional and motivational benefits to occur; highlights the fact that the emotional dimension concerns only novice problem-solvers; makes it possible to design more targeted and more efficient pedagogical interventions. references ahmed, w., van der werf, g., kuyper, h., & minnaert, a. (2013). emotions, self-regulated learning, and achievement in mathematics: a growth curve analysis. journal of educational psychology, 105 (1), 150-161. doi :10.1037/a0030160 . allal, l. (2007). régulations des apprentissages: orientations conceptuelles pour la recherche et la pratique en éducation [regulation of learning: conceptual guidelines for research and practice in education]. in l. allal & l. mottier lopez (eds.), régulation des apprentissages en situation scolaire et en formation [regulation of learning in educational and instructional situations] (pp. 7-24). brussels: de boeck. anderson, j. r. (1981). cognitive skills and their acquisition. hillsdale, nj: lawrence erlbaum associates. bassok, m. (2003). analogical transfer in problem solving. in j. e. davidson & r. j. sternberg (eds.), the psychology of problem solving (pp. 343-369). new york, ny: cambridge university press. beswick, k. (2006). the importance of mathematics teachers' beliefs. the australian mathematics teacher, 62(4), 17-22. blayney, p., kalyuga, s., & sweller, j. (2010). interactions between the isolated-interactive elements effect and levels of learner expertise: experimental evidence from an accountancy class. instructional science, 38(3), 277-287. doi: 10.1007/s11251-009-9105-x. blum, w. (2011). can modelling be taught and learnt? some answers from empirical research. in g. kaiser, w. blum, r. borromeo ferri, & g. stillman (eds.), trends in teaching and learning mathematical modelling(pp. 15-30). new york, ny: springer. blum, w., & niss, m. (1991). applied mathematical problem solving, modelling, applications, and links to other subjects. state, trends and issues in mathematics instruction. educational studies in mathematics, 22(1), 37-68. doi: 10.1007/bf00302716. boekaerts, m. (2011). emotions, emotion regulation, and self-regulation of learning. in b. j. zimmerman & d. h. schunk (eds.), handbook of self-regulation of learning and performance (pp. 408–425). new york, ny: routledge. brand, s., reimer, t., & opwis, k. (2003). effects of metacognitive thinking and knowledge acquisition in dyads on individual problem solving and transfer performance. swiss journal of psychology, 62 (4), 251-261. doi: 10.1024/1421-0185.62.4.251. brault-labbé, a., & dubé, l. (2008). engagement, surengagement et sous-engagement académiques au collégial: pour mieux comprendre le bien-être des étudiants [academic commitment, over-commitment and under-commitment at secondary school: understanding students’ wellbeing]. revue des sciences de l’éducation, 34(3), 729-751. doi: 10.7202/029516ar. cumming, g. (2012). understanding the new statistics. effect sizes, confidence, intervals, and meta-analysis. new york, ny: routledge. de corte, e. (2012). constructive, self-regulated, situated and collaborative (cssc) learning: an approach for the acquisition of adaptive competence. journal of education, 192(2/3), 33-47. de corte, e., depaepe, f., op’t eynde, p., & verschaffel, l. (2011). students’ self-regulation of emotions in mathematics: an analysis of meta-emotional knowledge and skills. zdm, 43(4), 483-495. de corte, e., verschaffel, l., & masui, c. (2004). the clia-model: a framework for designing powerful learning environments for thinking and problem solving. european journal of psychology of education, 19 (4), 365-384. demonty, i., blondin, c., matoul, a., baye, a., & lafontaine, d. (2013). la culture mathématique à 15 ans. premiers résultats de pisa 2012 en fédération wallonie-bruxelles [mathematical knowledge at 15 years old. first results of pisa 2012 in the wallonia-brussels federation]. les cahiers des sciences de l’education, 34, 1-26. demonty, i., & fagnant, a. (2014). tâches complexes en mathématiques : difficultés des élèves et exploitations collectives en classe [complex mathematical tasks: students’ difficulties and whole-class practice].education et francophonie, 42(2), 173-189. doi : 10.7202/1027912ar. depaepe, f., de corte, e., & verschaffel, l. (2010). teachers' approaches towards word problem solving: elaborating or restricting the problem context. teaching and teacher education, 26(2), 152-160. doi: 10.1016/j.tate.2009.03.016 . durlak, j. a., & dupre, e. p. (2008). implementation matters: a review of research on the influence of implementation on program outcomes and the factors affecting implementation. american journal of community psychology, 41(3-4), 327-350. doi : 10.1007/s10464-008-9165-0. durlak, j. a., weissberg, r. p., dymnicki, a. b., taylor, r. d., & schellinger, k. b. (2011). the impact of enhancing students' social and emotional learning: a meta-analysis of school-based universal interventions. child development, 82(1), 405-432. doi: 10.1111/j.1467-8624.2010.01564.x. elia, i., van den heuvel-panhuizen, m., & kolovou, a. (2009). exploring strategy use and strategy flexibility in non-routine problem solving by primary school high achievers in mathematics, zdm mathematics education, 41(5), 605-618. ericsson, k. a. (2008). deliberate practice and acquisition of expert performance: a general overview. academic emergency medicine, 15, 988-994. ericsson, k. a., krampe, r. t., & tesch-römer, c. (1993). the role of deliberate practice in the acquisition of expert performance. psychological review, 100, 363-406. fagnant, a., & demonty, i. (2005). résoudre des problèmes : pas de problème! guide méthodologique et documents reproductibles. 10/12 ans [solving problems: no problem! methodological guide and reproducible documents. 10/12 years]. brussels : de boeck. fagnant, a., demonty, i., & lejonc, m. (2003). la résolution de problèmes: un processus complexe de « modélisation mathématique » [problem-solving: a complex process of “mathematical modeling”]. bulletin d’informations pédagogiques, 54, 29-39. fagnant, a., dupont, v., & demonty, i. (2016). régulation interactive et résolution de tâches complexes en mathématiques [interactive regulation and complex tasks in mathematics]. in l. mottier lopez & w. tessaro (eds.), le jugement professionnel au cœur de l’évaluation et de la régulation des apprentissages [professional judgment at the core of evaluation and regulation of learning] (pp. 229-251 ). berne: peter lang. fagnant, a., & jaegers, d. (2018). soutenir l’autorégulation cognitive et développer les compétences en résolution de problèmes: une étude exploratoire en fin d’enseignement primaire [supporting cognitive self-regulation and developing problem-solving skills: an exploratory study at the end of primary education]. in s. cartier & l. mottier lopez (eds.), soutien à l'apprentissage autorégulé en contexte scolaire: perspectives francophones [support for self-regulated learning in the school context : francophone perspectives] (pp. 161-181). quebec: presses de l’université du québec. fayol, m. (2006). un esprit pour apprendre [a mind to learn]. in e. bourgeois & g. chapelle (eds.), apprendre et faire apprendre [to learn and to teach] (pp.53-67). paris: puf. field, a. (2013). discovering statistics using ibm spss statistics (4th ed.). london: sage publications. fredricks, j. a., & mccolskey, w. (2012). the measurement of student engagement: a comparative analysis of various methods and student self-report instruments. in s. l. christenson, a. l. reschly, & c. wylie (eds.), handbook of research on student engagement(pp. 763-782). new york, ny: springer us. fuchs, l. s., fuchs, d., prentice, k., burch, m., hamlett, c. l., et al. (2003). explicitly teaching for transfer: effects on third-grade students’ mathematical problem solving. journal of educational psychology, 95(2), 293-305. greenberg, m. t., domitrovich, c. e., graczyk, p. a., & zins, j. e. (2005). the study of implementation in school-based preventive interventions: theory, research, and practice . washington, dc: center for mental health services, substance abuse and mental health administration, u.s. department of health and human services. hampel, p., meier, m., & kümmel, u. (2008). school-based stress management training for adolescents: longitudinal results from an experimental study. journal of youth and adolescence, 37 (8), 1009-1024. doi: 10.1007/s10964-007-9204-4. hanin, v., & van nieuwenhoven, c. (2016). evaluation d’un dispositif pédagogique visant le développement de strategies cognitives et métacognitives en résolution de problème en première secondaire [evaluation of a training-program aiming at the development of cognitive and metacognitive strategies in problem-solving among grade one students]. e-jiref, 2(1), 53-88. hanin, v., grégoire, j., mikolajczak, m., fantini-hauwel & van nieuwenhoven, c. (2017). children’s emotion regulation scale in mathematics (cers-m): development and validation of a self-reported instrument. psychology, 8(13), 2240-2275. doi : 10.4236/psych.2017.813143. hanin, v., & van nieuwenhoven, c. (2018a). evaluation d’un dispositif d’enseignement apprentissage en resolution de problèmes mathématiques: evolution des comportements cognitifs, métacognitifs, motivationnels et émotionnels d’un résolveur novice et expert [evaluation of a training-program in mathematical problem-solving : evolution of cognitive, metacognitive, motivational and emotional behaviors of an expert and a novice solver]. e-jiref, 4(1), 37-66. hanin, v., & van nieuwenhoven, c. (2018b). developing an expert and reflexive approach to problem-solving: the place of emotional knowledge and skills. psychology, 9(2), 280-309. doi: 10.4236/psych.2018.92018. hoffman, b. (2010). i think i can, but i’m afraid to try: the influence of self-efficacy and anxiety on problem-solving efficiency. learning & individual differences, 20, 276-283. kanfer, r., ackerman, p. l., & heggestad, e. d. (1996). motivational skills & self-regulation for learning: a trait perspective. learning and individual differences, 8(3), 185-209. doi: 10.1016/s1041-6080(96)90014-x . kramarski, b., & gutman, m. (2006). how can self-regulated learning be supported in mathematical e-learning environments? journal of computer assisted learning, 22(1), 24-33. lucangeli, d., tressoldi, p. e., & cendron, m. (1998). cognitive and metacognitive abilities involved in the solution of mathematical word problems: validation of a comprehensive model. contemporary educational psychology, 23(3), 257-275. malmivuori, m. l. (2006). affect and self-regulation. educational studies in mathematics, 63(2), 149-164. doi: 10.1007/s10649-006-9022-8. maury, s. (2001). didactique des mathématiques et psychologie cognitive : un regard comparatif sur trois approches psychologiques [didactics of mathematics and cognitive psychology: a comparative look at three psychological approaches]. revue française de pédagogie, 137, 84-89. montague, m., & applegate. (2000). middle school students’ perceptions, persistence, and performance in mathematical problem solving.learning disability quarterly, 23(3), 215-227. doi: 10.2307/1511165. muir, t., beswick, k., & williamson, j. (2008). “i’m not very good at solving problems”: an exploration of students’ problem solving behaviours. the journal of mathematical behavior, 27(3), 228-241. doi: 10.1016/j.jmathb.2008.04.003 . national council of teachers of mathematics (nctm). (2010). why is teaching with problem solving important to student learning? reston: national council of teachers of mathematics. novick, l. r. (1988). analogical transfer, problem similarity, and expertise. journal of experimental psychology: learning, memory, and cognition, 14 (3),510–520. doi: 10.1037/0278-7393.14.3.510 . oecd. (2016). pisa 2015 results. excellence and equity in education (volume i). paris: oecd publishing. op’t eynde, p., de corte, e., & mercken, i. (2004). pupils’ (meta)emotional knowledge and skills in the mathematics classroom. paper presented at the annual meeting of the american educational research association (aera) , san diego. pekrun, r. (2014). emotions and learning. educational practices series. belley, france: international academy of education. pekrun, r., goetz, t., daniels, l. m., stupnisky, r. h., & perry, r. p. (2010). boredom in achievement settings: control-value antecedents and performance outcomes of a neglected emotion. journal of educational psychology, 102(3), 531-549. doi: 10.1037/a0019243. pekrun, r., goetz, t., & frenzel, a.c. (2005). achievement emotions questionnaire-mathematics (aeq-m). user’s manual. unpublished document. university de munich, munich. piaget, j. (1978). success and understanding. cambridge, ma: harvard university press. pollock, e., chandler, p., & sweller, j. (2002). assimilating complex information. learning and instruction, 12(1), 61-86. doi: 10.1016/s0959-4752(01)00016-0 . priolet, m. (2014). enseignement-apprentissage de la résolution de problèmes numériques à l’école élémentaire: un cadre didactique basé sur une approche systémique [teaching-learning of digital problem-solving in elementary schools: a didactic framework based on a systemic approach]. education & didactique, 8(2), 59-86. tzohar-rozen, m., & kramarski, b. (2014). metacognition, motivation and emotions: contribution of self-regulated learning to solving mathematical problems. global education review, 1(4), 76-95. tzohar-rozen, m., & kramarski, b. (2017). meta-cognition and meta-affect in young students: does it make a difference on mathematical problem solving? teachers college record, 119(13). van der sandt, s. (2007). research framework on mathematics teacher behaviour: koehler and grouws' framework revisited.eurasia journal of mathematics, science & technology, 3(4), 343-350. van dooren, w., verschaffel, l., greer, b., de bock, d., & crahay, m. (2010). la modélisation des problèmes mathématiques [modeling mathematical problems]. in m. crahay & m. dutrévis (eds.), psychologie des apprentissages scolaires [psychology of school learning](2nd ed., pp. 199-220). bruxelles: de boeck. van merriënboer, j. j., & sweller, j. (2005). cognitive load theory and complex learning: recent developments and future directions. educational psychology review, 17(2), 147-177. doi: 10.1007/s10648-005-3951-0. veenman, m. v. j., van hout-wolters, b., & afflerbach, p. (2006). metacognition and learning: conceptual and methodological considerations. metacognition and learning, 1(1), 3-14. doi: 10.1007/s11409-006-6893-0. wilkins, j. l. (2008). the relationship among elementary teachers’ content knowledge, attitudes, beliefs, and practices. journal of mathematics teacher education, 11(2), 139-164. doi: 10.1007/s10857-007-9068-2. wolters, c. a., & rosenthal, h. (2000). the relation between students’ motivational beliefs and their use of motivational regulation strategies. international journal of educational research, 33(7-8), 801-820. doi: 10.1016/s0883-0355(00)00051-3 . zimmerman, b. j. (2011). motivational sources and outcomes of self-regulated learning and performance. in b. zimmerman & d. schunk (eds.), handbook of self-regulation of learning and performance (pp. 49-64). new york, ny: routledge. zimmerman, b. j., & campillo, m. (2003). motivating self-regulated problem solvers. in j. e. davidson & r. j. sternberg (eds.), the psychology of problem solving(pp. 233-262). new york, ny: cambridge university press. appendix a. training program’s features. the training program implemented in the present study aimed to develop an expert and reflexive approach to problem solving among students by means of the following components: (1) each regular teacher familiarized his/her students with the eight heuristic strategies/stages depicted in figure 1. this familiarization was performed according to the www & h rule (veenman et al., 2006), which consists in teaching each heuristic strategy by specifying the what (what it consists of), the why (its usefulness), the when (the most relevant point in the problem-solving process at which to implement it), and the how (the way to implement it correctly); (2) the teachers used an open-ended methodology to foster diversity of problem representations, modelings, strategies and procedures (fagnant & demonty, 2005); (3) students were trained and scaffolded to regulate their own problem-solving process in an increasingly autonomous way; (4) the non-routine problems chosen were realistic (i.e., problems were anchored in fifth-grade students’ experiential worlds), complex (i.e., problems made it necessary to implement a mathematical modeling process), and open-ended (i.e., problems could be correctly represented, modeled, and solved by taking different paths), as suggested by de corte et al. (2004). furthermore, as the objective of this training program was the development of a process to solve non-routine problems, only application problems were selected. moreover, except for the first problem, which was solved in groups of five students, problems were solved individually and followed by a whole-class discussion. codepen 5. clercq frontline learning research special issue vol.9 no.2 (2021) 96 120 issn 2295-3159 bridging contextual and individual factors of academic achievement: a multi-level analysis of diversity in the transition to higher education mikaël de clercq, benoit galand, virginie hospel & mariane frenay1 auniversité catholique de louvain, belgium article received 15 may 2020/ revised 14 october / accepted 31 october/ available online 12 march abstract the transition to higher education has been extensively documented in the literature. in this line, many individual variables were identified as strong predictors of academic achievement. yet, this literature suffers from one main limitation; contextual factors have often been left out of the investigation. the majority of studies have tested the impact of individual characteristics assuming that the effects are the same in different programs. however, differences between institutions or programs could result in specific learning contexts leading to different adjustment processes. as an attempt to overcome this limitation, the current study has investigated the impact of both individual and contextual factors on academic achievement through a multifactorial multilevel analysis. the analyses were carried out on 1,173 freshmen from 21 study programs. results highlighted that 15% of variation in students’ achievement was found between programs. aspects of curriculum organization that contributed to academic achievement were gender ratio, opportunities given for practice and class size. besides, seven individual factors were also predictive of academic achievement in the multifactorial approach: past performance, socioeconomic status, self-efficacy beliefs, value, mastery goal structure, study time and paid job. finally, significant random effects were identified for peer support, course value, attendance and external engagement (i.e. commitment in extra-academic activities). the implications and limitations of this study are discussed. by connecting individual and contextual predictors of academic achievement this study intends to endorse a frontline approach regarding the transition to higher education. keywords: academic achievement, learning environment, teaching practices, contextual factors info corresponding author email: mikael.declercq@uclouvain.be doi: https://doi.org/10.14786/flr.v9i2.671 1. introduction in the light of political imperatives for study success and higher education expansion, the transition to higher education (he) and more precisely student’s first-year experience has received increased attention (gale & parker, 2014). the challenge is to understand the process which can lead students to the academic achievement, and factors embedded in it. the effects of many variables have been investigated on academic achievement: socioeconomic status (sirin, 2005), past performance (elias & macdonald, 2007), attendance (credé, roch & kieszczynka, 2010), self-efficacy (bandura, 1997), peer support (dennis, phinney & chuateco, 2005), informed choice (husman and lens 1999) and perception of teaching practices (lizzio, wilson & hadaway, 2007). despite a vast body of research, most studies addressed individual predictors of academic achievement without considering organizational diversity among institutional contexts. as achievement process is mainly considered as universal, environmental characteristics are overlooked. yet, as suggested by some authors (pascarella and terenzini, 2005; kuh, kinzie, schuh & whitt, 2011), distinctiveness of countries, institutions or programs in admission procedures, student body, requirements and assessment could result in specific learning environments leading to specific achievement process. for example, we can postulate that the learning context of a student registered in medicine in the united kingdom will largely differ from those of a student enrolled in a sociology program in france. moreover, achievement among the first year at the university is a complex process involving a series of interrelated factors. for instance, de clercq, galand and frenay (2017) found that the impact of self-efficacy beliefs on achievement depends on students’ past performance and their socioeconomic status. the negative effect of low confidence in his or her ability to succeed could be offset by high past performance and privileged social background. therefore, it is important to understand the relative importance of these factors when they are considered together. current research widely acknowledges the need to endorse multifactorial approaches of academic achievement (coertjens, brahm, trautwein, & lindblom-ylänne, 2017; de clercq, galand, dupont, & frenay, 2013a; richardson, abraham, & bond, 2012; van petegem, coertjens, donche, & noyens, 2017). this paper can be understood as a frontline step forward in the consideration of the transition to he through a multifactorial and multi-level analysis of students’ achievement process. as achievement can be conceived as an index of a successful transition, the investigation of the process leading to achievement during the first year at the university can be considered as an interesting perspective on student’s transition. more precisely, performing a multi-level analysis, the present study aimed at investigating the main effects of the factors embedded in students’ achievement processes, considering both individual and program characteristics. this study investigated the predictors of a successful transition by connecting individual and contextual levels of investigation of students’ achievement processes. 1.1 toward a conceptual framing of the achievement process based on literature about first-year experience at the university, three categories of achievement predictors can be distinguished: background factors, experience of the learning environment and psychosocial factors (de clercq et al., 2013a, de clercq, galand & frenay, 2020; richardson et al., 2012; schneider, & preckel, 2017). more precisely, a conceptual framing of the achievement process can be proposed from the integration of previous works and theories on the first-year experience. recent work embedded in an adapted approach of tinto’s model of student departure also suggested that both individual backgrounds, psychological attributes, and perceived institutional factors should be considered together when addressing the first year in he (schaeper, 2019). another recent research grounded in the self-system model of motivational development (skinner, wellborn, & connell, 1990) and social cognitive theory of career and educational choice (lent & brown, 2019) posited that student achievement will be directly influenced by psychological factors (engagement and motivation) which are, in turn, impacted by the experience of the learning environment (de clercq et al., 2020). this study also considered backgrounds factors as important distal predictors of academic achievement. this conceptual framing is also consistent with price’s 4p model of student’s learning outcomes in he (price, 2014). this model, focused on approach to learning, arguing that a student’s characteristics (presage) will determine his or her perceptions of the learning context (perceptions) which will in turn influence the student’s approach of the learning task (process) which will finally predict his or her outcomes (product). based on these studies and the above-mentioned theoretical frameworks, we suggested that achievement would be a three-stage process including backgrounds factors, experience of the learning environment and psychological factors. more precisely, psychological factors are expected to have the most proximal impact on academic achievement. the experience of the learning environment is supposed to have a distal impact, and the background factors are expected to be the most distal category of variables. the categories of predictors are described in figure 1 hereunder. figure 1. conceptual modelling of the achievement process moreover, beyond these three categories of variables, the objective characteristics of the context were added to the model. based on the i-e-o model (astin, 2012), a recent study highlighted the necessity to consider both student perceptions and institutional characteristics when investigating college experience (georgianna et al., 2020). this argument can be related to broader theories such as the ecological systems theory (bronfenbrenner, 1992; 2005). this theory postulates that (1) individual development arises through transactions with the environment and that (2) different environmental nested levels can be considered. ecological systems theory can also be related to studies using a three-level framework of he (enders, 2004; munge, et al., 2018; nuñez, & kim, 2012; perna, & thomas, 2008; taylor & ali, 2017) which argued that the microsystem (students’ individual characteristics) will be affected by the meso(institutional context) and the macrosystem (educational policy). among the first year at the university, the institutional context (mesosystem) can be influential in several ways. schaeper (2019) proposed that it could impact student experience as an objective reality and as a subjective shared perception of the learning environment. the proposed framework therefore encompassed both objective and subjective dimensions of the context: the above-mentioned experience of the learning environment and the organizational characteristics of the institutional context such as course timetable, class size or composition of student body. this additional category considers for organizational aspects of the context (georgianna et al., 2020; pascarella and terenzini, 2005; kuh et al., 2011). 1.2 the organizational aspects of the institutional context van der hulst and jansen (2002) were the first to investigate the question of academic achievement using multi-level analysis. focusing on 1,578 engineering students from twelve different cohorts of students in three engineering disciplines, they showed that three percent of the variation in student’s achievement was due to the curriculum. they also found a significant impact of the spread of study activities over the year and instruction characteristics. van den berg and hofman (2005) also used multi-level analysis to estimate the variation in study progress from one course to another among 1,800 freshmen from 60 different courses. they showed that five percent of variation was due to factors related to courses, such as course timetable or course evaluation. this study also highlighted the importance of study load and opportunity for practice (problem-based learning) for achievement. yet, the results did not highlight any effect of examination characteristics. jansen (2004) also questioned the influence of curriculum characteristics on academic success in six different departments. she demonstrated the significant effect of curriculum organization characteristics such as study load and tutorial hours on academic success. these studies provided interesting first findings about the effects of organizational aspects on academic achievement. yet, they are still too scarce to provide a clear picture of this impact. beyond organizational aspects of institutional context, some authors also found that individual predictors of achievement varied from one study program to another (de clercq et al., 2013a; lizzio, wilson, & simons 2002, millet, 2003; 2012). for instance, lizzio and colleagues (2002) found that the impact of prior achievement was only significant in one of three faculties (science) investigated. de clercq and colleagues (2013a) found that peer support, intention to persist and time spent to study were only significant predictors for science program whereas, extrinsic motivation was only a negative predictor of achievement in a physical education program. these results suggest furthering investigations on educational environment. 1.3 the role of background factors student background factors (such as gender, age, parental educational level, socioeconomic status, and past academic performance or standardized achievement test score) have been extensively studied in the literature about the first year at the university (van rooij, brouwer, fokkens-bruinsma, jansen, donche, & noyens, 2018). from a theoretical point of view, several models supported the importance of background characteristics in the first year at the university. for example, they are conceived as pre-entry attributes of tinto’s theory of departure (tinto, 1997) and as pre-college traits of pascarella and terenzinni’s model of change (2005). many studies have indicated that women performed better than men in he (lekholm & cliffordson, 2008), and that age is negatively related to academic performance (farsides & woodfield, 2007). yet, the meta-analysis of richardson and colleagues (2012) on first-year students estimated very low corrected correlations of these factors with academic achievement (between .03 and .04). another line of research tackled students’ educational or socioeconomic background. students from lower socioeconomic backgrounds were found to have lower achievement (rodríguez-hernández, cascallar, e., & kyndt, 2020; sirin, 2005). several studies based on social capital theory (bourdieu, 1986) identified social capital as an explanation of how socioeconomic status is related to academic achievement. yet, this idea was recently questioned and criticized (rodríguez-hernández et al., 2020). the meta-analysis of richardson and colleagues (2012) estimated a corrected correlation of .15 with academic achievement. however, some studies failed to replicate this result, claiming that socioeconomic status has no direct effect on achievement at university when the impact of past performance is controlled (sackett, kuncel, arneson, cooper, & waters, 2009). in this idea, a mediating mechanism can be postulated: socioeconomic status would determine prior school background (high school grade, school specialization, educational track…) which would subsequently be predictive of academic achievement. past academic performance (i.e. secondary-school diploma or standardized achievement test scores) has also been an extensive topic of investigation. four meta-analyses (poropat, 2009; richardson et al., 2012; robbins et al., 2004; sackett et al., 2009) corroborated the stable and important link between past performance and academic achievement (corrected correlation between .20 and .35). this factor is mainly identified as the most powerful predictor – sometimes the only significant one – of achievement in college or at university (díaz, glass, arnkoff, & tanofsky-kraff, 2001; perry, hladkyj, pekrun, & pelletier, 2001; vandamme, superby, & meskens, 2005). beyond traditional background factors, another line of research recently emerged from the literature, focusing on student’s study choice process. several studies and theories, such as future time perspective theory (husman & lens, 1999) and educational choice implementation (germeis & verschueren, 2007; germeijs, luyckx, notelaers, goossens & verschueren, 2012), have highlighted the importance of the information process. first findings concluded that students who made an informed and thoughtful study choice attain higher academic achievement, were more satisfied with their courses, and applied more adaptive study strategies (biémar, philippe, & romainville, 2003; lens, simons, & dewitte, 2002). however, this field of research remains underdeveloped and needs more investigation to support these first findings. despite a vast body of research on background factors, this category of predictors is far from explaining all the variance in achievement, suggesting that other variables could also play a role. indeed, robbins et al. (2004) showed in a meta-analysis that background factors account approximately for 22 percent of the variance in academic achievement. these authors also confirmed that other kinds of factors, namely psycho-social factors, make a significant incremental contribution in predicting college achievement. such results were corroborated by dollinger, matyja and huber (2008) who insisted on the added value of considering background factors together with psychosocial ones. allen and colleagues (2010) found that some psychosocial factors could have some effect sizes comparable to those of background factors. moreover, recent research insisted on the added value to move beyond backgrounds variables and to consider the experience of the learning environment in the investigation of the first year in he (schaeper, 2019). this experience could have a strong impact on academic achievement (schneider, & preckel, 2017). 1.4 the role of the experience of the learning environment grounded in a first contextual consideration of academic achievement, some studies put forward the necessity to consider students’ experience of the university (lizzio et al., 2007). the experience of university, defined as the way the student connects to the specific university context, could have an impact on engagement and achievement (rayle, kurpius, & arredondo, 2006). several studies found a relationship between perceptions of the learning environment and academic achievement (lee & burkam, 2003; lizzio, et al., 2002; lizzio, et al., 2007). in he literature, a large amount of energy was devoted to the understanding of perceived teaching practices considered as a paramount component of students’ experience of the university (schneider, & preckel, 2017). the australian authors lizzio and colleagues (2002, 2007) specifically focused on the importance of the perceptions and the relationships with the teachers. their studies indicated that the perception of good teaching was positively associated with academic grades whereas the perception of a heavy workload and inappropriate assessment were negatively associated. such results were questioned by patrick, ryan and kaplan (2007) who found an indirect link between student/teacher relationship and achievement. teacher support would rather have a direct impact on students’ self-efficacy beliefs and motivation. some authors supported that, beyond the quality of the relationship with the teacher, what was important for student success is the teaching climate (anderman & patrick, 2012). according to the goal theory (ames, 1992), students’ subjective perceptions of the values and messages conveyed by the learning environment would play a significant role in their engagement and achievement (anderman & patrick, 2012). more precisely, two constructs can be distinguished: mastery goal structure and performance goal structure. according to anderman & patrick (2012), mastery goal structure encompasses with students’ perception that “learning and understanding are valued and that success is indicated by personal improvement” (p.181). conversely, performance goal structure taps the student’s perception that “achievement and success entail outperforming others or surpassing normative standards” (anderman & patrick, 2012; p.181). several authors found a positive association of mastery goal structure and academic achievement in high school context (bong, 2005; greene, miller, crowson, duke & akey, 2004; meece, anderman & anderman, 2006; roeser, midgley & urdan, 1996). studies also investigated these variables in he and showed a positive link between mastery goal, task value and achievement (church, elliot & gable, 2001; de clercq, et al., 2020; karabenick, 2004, zusho, karabenick, bonney & sims, 2007). however, empirical findings about the relation between these constructs and academic achievement in he are still scarce and deserve further consideration (de clercq et al., 2020). although this literature attempted to address the contextual nature of the first-year experience, it only focuses on individual perceptions of the learning environment. the studies mentioned above did not use multi-level analyses which also impede the interpretation of their findings and conclusions regarding the experience of the program. according to marsh and colleagues (2012), the aggregation of the individual perceptions of the students is needed to overcome idiosyncratic bias and obtain an accurate estimation of shared learning environment. such methodology was employed by some researchers in order to investigate the impact of the actual curriculum characteristics on academic success (jansen, 2004; van den berg & hofman, 2005; van der hulst & jansen, 2002). our work aimed to apply this approach to the investigation of the impact of learning environment experience on academic achievement. 1.5 the role of psychosocial factors based on contemporary educational (e.g. tinto, 1997) and motivational theories (e.g; eccles and wigfield, 2002), robbins and colleagues (2004) delimited psychosocial factors as the motivational, behavioural and social variables at work during the academic year. several factors were mentioned such as academic self-efficacy beliefs, task value, intention to persist, social support and commitment. social cognitive theory (bandura, 1997) claimed that confidence in one’s ability and chance of success are important in the prediction of achievement. this assumption has been substantiated by related works such as expectancy-value theory (eccles and wigfield, 2002). from an empirical lens, numerous studies supported the claim that academic self-efficacy is a key construct in student achievement in he (bruinsma, 2004; chemers, hu, & garcia, 2001; elias & macdonald, 2007). beyond its link with academic achievement, several authors (bong and skaalvik, 2003; torres and solberg, 2001) showed that self-efficacy beliefs foster social integration, intention to persist and engagement, while reducing the perceived stress of college students. the meta-analysis of richardson and colleagues (2012) identified self-efficacy beliefs as one of the most important predictors of academic achievement (corrected correlation between .28 and .67 depending to the measure used). grounded in expectancy-value theory (eccles & wigfield, 2002), task value is also identified as an important motivational construct in the achievement process. task value can be defined as students’ perception of the importance, the utility, the cost and the interest of a task (pintrich & de groot, 1990; eccles, 2005). empirical investigations in he mainly highlighted that task value had positive effects on study time, effort, and academic achievement (bong, 2005; bruinsma, 2004, neuville, frenay & bourgeois, 2007, pintrich & de groot, 1990, pintrich 1999, pintrich 2003). the meta-analysis of robbins and colleagues (2004) found a moderate link between task-value and academic achievement (corrected correlation of .25). another line of research focused on the intention to persist (cabrera, nora, & castaneda, 1993; hausmann, ye, schofield, & woods, 2009). this variable is at the core of tinto’s attrition theory (1997) as one of the most direct predictors of students’ persistence. moreover, intention to persist has been empirically found to be related to academic persistence (hausmann et al., 2009; vallerand, fortier & guay, 1997;). several authors also pointed to a positive association between the intention to persist and academic achievement (dadeppo, 2009; neuville et al., 2007; van rooij, jansen, & van de grift, 2018). more precisely, de clercq and colleagues (2013a) highlighted in hierarchical regressions that intention to persist significantly determined academic achievement while controlling for the impact of backgrounds factors, experience of the learning environment and other psychological factors. neuville and colleagues (2007) also demonstrated through structural equation modeling that intention to persist was a significant predictor of academic achievement. yet, the causal relationship between these two constructs is still unclear and it could also be assumed that intention to persist is a consequence of achievement. another body of research focused on the social layer of the transition to he. freshmen are separated from their previous social networks and need to create new ones. this is one of the first difficulties that these students have to cope with (schmitz et al., 2010). self-determination theory (ryan & deci, 2017) supported that an individual needs to feel socially accepted and recognized in order to actively engage in a specific task. the same goes for the student entering at the university who needs to feel personally supported by his peers in order to achieve academically. in this line, peer support also emerged as an individual predictor of academic achievement in some studies (dennis et al., 2005; hackett, betz, casas, & rocha-singh, 1992; larose, robertson, roy, & legault, 1998). two meta-analyses (richardson et al., 2012; robbins et al., 2004) assessed a weak effect size of peer support on academic achievement (corrected correlation between .09 and .11). for some authors, the impact of peer support on academic achievement may be indirect. for example, torres and solberg (2001) showed that peer support had a direct impact on self-efficacy beliefs, study time and anxiety but no direct effect on achievement. finally, some scholars stated that behavioural commitment is the most proximal predictors of academic performance (credé et al., 2010). in this line, dollinger and colleagues (2008) concluded that class attendance accounts for 6 to 10 percent of achievement variance among undergraduate students. similar conclusions were drawn by pirot and de ketele (2000) among first-year university students. moreover, a meta-analysis from credé and colleagues (2010) revealed a strong relationship between class attendance and academic marks (r = .41), leading the authors to claim that class attendance is the single best predictor of achievement. study time, another indicator of commitment, was found to be moderately correlated with final grades in the first year at university (vandamme et al., 2005; van den berg and hofman, 2005). the meta-analysis of robbins et al. (2004) found a direct weak link between study time and academic marks (corrected correlation of .12). finally, some studies (torenbeek, jansen & hofman, 2010; van den berg & hofman, 2005) investigated the external engagement (commitment in extra-academic activities such as paid work). these studies found a negative association with academic achievement. 1.6 aims of the study as an attempt to overcome limitations of the literature on academic achievement at the first year at the university, this study endorsed a multifactorial and contextual approach regarding the achievement process. students’ achievement process is considered as a proxy to assess a successful transition to he. this procedure can therefore be conceived as a frontline approach of transition to he by bridging subjective and objective aspects of the institutional context to background and psychosocial individual factors. moreover, the study also aimed at analyzing the contextual differences among several study program. this consideration made an innovative investigation of the meso-level of diversity during the first year in he possible. more precisely, the aims were fourfold. first, the study aimed at identifying the main predictors of academic achievement when their impacts are considered together. regarding the above-mentioned literature, past performance (díaz et al., 2001), self-efficacy beliefs (richardson et al., 2012) and behavioural engagement (credé et al., 2010) are supposed to remain important predictors of academic achievement in multifactorial analyses. from a theoretical point of view, behavioral engagement is supposed to be the most proximal predictor of achievement in our theoretical framework. conversely, the effect of gender, age (richardson et al., 2012), socio-economic status (sackett et al., 2009), peer support (torres and solberg, 2001) and performance goal structure (anderman & patrick, 2012) on academic achievement are supposed to become non-significant predictors in such a multifactorial approach. these factors are also theoretically supposed to be distal and non-direct predictors of achievement in our theoretical framework. second, in order to highlight the importance of considering the contextual diversity in the transition to he, the study aimed to estimate the magnitude of the variations of academic achievement across different study programs. van der hulst & jansen (2002) namely found three percent of variation across study program in engineering. van den berg & hoffman (2005) found five percent of variation across the course regarding study progress. we expected higher variation across study programs because the contextual differences in the programs investigated in the current study (different disciplines, number of students, teaching practices, academic demands…) are much bigger than in research cited above. third, the aim of the study was also to understand which characteristics of the programs could explain variation in achievement. to do so, organizational characteristics of the study programs (class size, composition of the student body, course timetable) and aggregated measures of perceived learning environment (teacher-student relationship, goal structure) were considered at the program level. third, we assumed that the impact of several achievement predictors on academic achievement would significantly vary from one program to another. considering the work of lizzio and colleagues (2002) and de clercq and colleagues (2013a), different effects of past performance, intention to persist, motivation, peer support and behavioural engagement on achievement are expected.   2. methodology 2.1 sample 1,173 freshmen participated in the study. they came from 21 different study programs at the université catholique de louvain in belgium, namely medicine, dentistry, veterinary, psychology, philosophy, economics, management, law, engineering, biology, geology, physics, chemistry, agronomy, political science, communication, computer science, architecture, physical education and history. the sample was composed of 53% of female. the median age was 18 years old and scattered as follows: 16% of 17 years old freshmen, 52% of 18 year old, 20% of 19 year old, 7 % of 20 year old and 5% older than 20 year old. 2.2 administration and measures first several study program characteristics were retrieved from institutional records. the following information was obtained: gender ratio, proportion of tutorials1 in the program (hours of lecture/tutorials ratio), average class size and mean age. then, two self-report questionnaires were administered during regular lectures at the beginning and during the academic year. students were informed that they were free to participate in the study and that gathered information will be kept confidential. moreover, the study was approved by the ethics commission of the university. the first questionnaire administered in september assessed variables that characterized students at the entrance of the university. students reported their gender (0= male; 1= female), their age (1= 17; 2=18; 3=19; 4=20; 5= older than 20) and their high school grade (1= 60-70%; 2= 70-80%; 3= 80-90%; 4= more than 90%). their socioeconomic status was also measured, through cultural and material resources as used in pisa studies (oecd, 2012). exploratory factor analysis demonstrated one single factor (eigenvalue higher than 1) which explains more than forty percent of the variance. according to tabachnick and fidell (2007), factor loadings ranged from good to excellent (.57 to .72). factor score was extracted to create an overall ses indicator. finally, informed choice was estimated by the number of sources consulted by the students to choose their study (6 different sources; e.g. teachers, former students, counsellor). the second questionnaire administered in november encompassed variables that characterized students during the academic year (self-efficacy beliefs, task value, attendance, peer support, intention to persist) and perceived teaching practices. all scales were adapted from galand & frenay (2005). students answered on a five-point scale from 1 = totally disagree to 5 = totally agree. self-efficacy beliefs were measured through seven items (e.g., ‘as long as i do my work, i’m sure i can succeed this year’; α = .72). task value was assessed through 12 items (e.g., ‘i am really interested in the courses’ α =.83). intention to persist was assessed using/by means of 5 items (e.g., ‘i will continue my study no matter if i pass or not the year”; α =.87). peer support was measured through 7 items (e.g., ‘i know i can rely on some other students to help me’; α =.80). two items tapped student’s time spent studying (e.g., ‘how many hours per week do you usually spent on academic work?’; α =.66), two items assessed attendance (e.g., ‘how often do you attend the lecture?’; α =.74) and two dichotomous items tapped external engagement (e.g., ‘do you need to have a paid job to finance your study”). teaching practices were measured through three different scales adapted from galand & frenay (2005). students answered on a five-point scale from 1 = totally disagree to 5 = totally agree. the quality of teacher/student relationship was estimated by four items (e.g., ’teachers listen to our needs‘; α =.71). students’ perceived mastery structure was assessed through seven items (e.g., ‘teachers first emphasized and focused on student’s understanding‘; α =.67). finally, students’ perceived performance structure was measured through six items (e.g., ‘teachers mainly help high achieving students’; α =.69). in order to obtain accurate indices of shared learning environment, students’ perceptions of teaching practices were aggregated to the study program level. therefore, we also assessed the icc(1) (i.e., proportion of variance occurring at the classroom level) and the icc(2) (i.e., the level of inter-rater agreement between students’ rating within classrooms) (lüdtke et al., 2009). whereas icc(1) values greater than .10 (10% of variance occurring at the program level) suggest the presence of enough program-level variability to support multilevel analyses, icc(2) values can be interpreted as traditional estimate of reliability (marsh et al., 2012). the results revealed sufficient variability at the program level, and adequate levels of reliability (quality of teacher/students relationship: icc(1) = .161; icc(2) = .914; perceived mastery structure: icc(1) = .105; icc(2) = .867; perceived performance structure: icc(1) = .182; icc(2) = .925). finally, academic achievement was measured at the end of the year (august) using students’ gpa (final overall percentage). in the french speaking belgian tertiary educational system, achievement is measured through the average percentage for all courses at the end of the academic year. this final score was collected from department records and used as an overall indicator of achievement. an overall final percentage of sixty is required to succeed. 2.2 analytical procedure multilevel analyses were performed using the hlm7 software using a step-by-step procedure. the significance test of deviance reduction was conducted at the 5% level. apart from gender, all variables introduced in the analyses were grand mean-centred. first, an empty model was run to estimate the proportion of between programs variance. next, in the model 1, study program characteristics were entered at level 2. in model 2, background characteristics were added at the individual level. in order to estimate the impact of the learning environments, students’ perception of teaching practices were introduced at level 2 (model 3) but not at the individual level. such procedure is supported by marsh and colleagues (2012) who highlighted that the most appropriate measure of learning climate consists in the aggregated students’ perceptions introduce at program level. in model 4, psychosocial factors were added to the model at the individual level. finally, in order to estimate the variation of the effects of level 1 predictors on achievement across the programs, slope variations were introduced in model 5. reduction of the residual variance between programs and within programs is presented in the tables of the results section. effect sizes for level 1 factors were calculated using reyes, brackett, rivers, white and salovey (2012, p. 706) formula (δ= γ/(√(r00+σ^2)); “where γ is the association between the predictor and outcome variables, and the denominator is the standard deviation of the outcome variable, where τ00 and σ2 are the betweenand within-groups variances”) and directly integrated in the results section.  3. results table 1 presented descriptive statistics for the level 1 variables and their correlations with academic achievement. age, high school grade and socioeconomic status were the most related to academic achievement. table 1 descriptive statistics for variables included in the analyses and correlation with academic achievement. note. *p<.05;**p<.01;***p<.001. most relationships were weak. the correlation matrixes among the independent variables at the individual and the program level were respectively provided in table 2 and 3. table 2 correlations among independent variables at the individual level. the table 3 highlighted that perceived performance goal structure was positively related program with age and proportion of girls among the program. it was also negatively related to students’ perceived quality of the relationship with the teachers. table 3 model 1 enclosed study program characteristics and reported 9.3% of between-program variance explanation. more precisely, study program’s average “class” size was a significant negative predictor of academic achievement. model 2 included individual variables. introduction of these variables explained 18% of within-classroom variance and 25% of between-classroom variance. past performance was a significant predictor of academic achievement (δ=0.37). students with more privileged socioeconomic background (δ=0.06) and more thoughtful study choice process (δ=0.02) were also found to achieve slightly better. detailed information about model 2 can be found in table 2. another intuitive way of interpreting the results could be adopted by referring directly to the coefficient. for instance, the coefficient of 7.64 means that an increase of one point out of the past-performance likert scale will induce an increase of 7.64% on final average percentage. table 4 results of multilevel analyses for models 1, 2 and 3. note. nstudents=1173; nprograms=21 *p<.05;**p<.01;***p<.001. the perceptions of teaching practices aggregated at the program level were added in the third model (model 3) and their main effects on achievement were tested. the results can be found in table 3. this model improved the between-program variance explanation by 15.6%. the results showed better academic achievement in study programs in which teaching practices were perceived as putting an emphasis on the mastery of the content and self-enhancement (δ=0.20). in model 4, psychosocial factors were added to the analysis at level 1. such variables provided incremental within-classroom variance explanation of 7.4%. four variables were found to have a significant impact on academic achievement. self-efficacy beliefs were related to better achievement at the end of the year (δ=0.15). the more students attributed value to the courses, the better they performed at the end of the year (δ=0.10). students who spent more time studying also reached higher performance (δ=0.08). finally, students involved in a paid job experienced more difficulties to perform (δ=0.22). detailed information about the model can be found in table 5. in order to estimate the variation of the impact of past performance, intention to persist, value, peer support, attendance, study time and external engagement, slope variations were released in model 5. this last model provided 4.6% of additional within-program variance explanation. the final model explained 30% of the within-program variance and 49.9% of the between-program variance. significant slope differences were found regarding intention to persist, value, peer support and attendance, disclosing that the relationship between these variables and achievement fluctuates between programs. it is also worth noting that lecture/tutorial hours ratio appeared as a significant negative predictor of achievement in this final model. 4. discussion the transition to he is a complex experience partially determined by the characteristics of the new environment. however, multifactorial and contextualized approaches have received relatively little empirical consideration to date impeding our understanding of the contextual diversity into he. this study went a step further by looking at this phenomenon through the specific lens of students’ achievement process. by using multilevel analysis, the current study estimated the magnitude of achievement variation across study programs and depicted the impact of background, psychosocial and environmental factors while considered together. 4.1 a multifactorial consideration of achievement process the results first corroborate that both psychosocial and background variables need to be considered (dollinger et al., 2008; robbins et al., 2004). psychosocial factors provide around nine percent of additional explained variance regarding achievement which is in line with the meta-analysis of robbins and colleagues (2004). this finding supports the necessity to consider these categories of variables together. backgrounds factors emerge as the most predictive category of predictors which highlighted the crucial importance of students’ preparation for he as assumed by nicholson’s (1990) model of transition cycles. a large part of achievement can therefore already be explained before entering the university. out of the results, the most important predictors are past performance, self-efficacy beliefs and behavioural engagement. the more the student is confident in his abilities to succeed and spends time studying, the better he or she will achieve academically. this is in line with the literature which highlighted the unequivocal links of these variables with academic achievement (credé et al., 2010, díaz et al., 2001;richardson et al., 2012). course value also proves to remain a significant predictor of achievement in this multifactorial analysis. this result is in line with the meta-analysis of robbins and colleagues (2004) that found a moderate effect on achievement. more surprisingly, socioeconomic status also remains a significant predictor of achievement whereas empirical and theoretical literature mainly identified this factor as having a distal and non-direct impact on achievement. this remaining effect could be explained by the specific nature of the belgian educational system. the belgian educational system is characterized as a highly unequal educational system where socioeconomic status has a strong impact on educational pathway since the very beginning of primary education (cattonar & mangez, 2014; maroy & dupriez, 2000). students from privileged background go to the best schools and develop the most advanced scholastic competences. at the same time, access to tertiary education (including he) is widely unconstrained and costs very low tuition fees. this specific configuration increases the importance of first-year students’ socioeconomic status. students from underprivileged background are free to enter at the university. however, they are very poorly prepared to face the requirement of the program due to their disadvantageous educational pathway. 4.2 combining micro and meso-levels of investigation as little is known about the contextual grounding of academic achievement, a first step was to assess the magnitude of achievement’s variation from one program to another. as expected, the variation of 15% is quite large and higher than previous estimates of van der hulst and jansen (2002) regarding engineering students or van den berg and hoffman (2005) on study progress variation across courses. consistently with previous studies (chickering & reisser, 1993), we can suggest with some confidence that study programs largely differ in learning environment which could explain achievement variations. therefore, the diversity of the environment would deserve more attention in future research. looking at the explanation of between program variance, three categories of variable are spotted as sources of explanation: curriculum characteristics, background variables and teaching environment. first, class size and lecture/tutorial hours’ ratio are both negative predictors of academic achievement in the final model. if a student is in a program with a large number of students and a lot of tutorials, he or she is going to achieve less compared to students in smaller programs and fewer tutorials. the result concerning the tutorials contradicts the results of van den berg and hofman (2005) who found a positive association between progress and opportunity for practice. this unexpected result can be explained by the level of investigation of the study. van den berg and hofman (2005) worked at the course level whereas this study analyzed the program level. thus, we assume that study programs that are very challenging and selective during the first year (i.e. medicine, engineering, physics) also offer more tutorials, and that thus, the composition of the program is relevant. second, background factors are identified as significant explanation of between-program variations in achievement. such finding could be interpreted as follows: the different programs do not deal with the same patterns of students. from one program to another, student background characteristics significantly vary which could explain between-program variations and highlight that contextual diversity is accompanied by individual diversity. according to our results, significant background variation can be found in past performance, socioeconomic status and informed choice. this assumption is in line with studies emphasizing the heterogeneity of the student body among first-year students (fenollar et al., 2007; heikkilä, niemivirta, nieminen & lonka, 2011). recent studies of de clercq and colleagues (2017; 2020) identified different student’s entrance profiles endorsing specific achievement processes. these results pointed to the importance of considering diversity in the student body when addressing academic achievement issue. they also lend credence to this individual heterogeneity and show its connection with contextual diversity. third, consistently with previous research (anderman & patrick, 2012; lizzio et al., 2007), the overall perception that teachers valued learning and understanding in a program (i.e. mastery goal structure) explains a significant part of achievement variation between programs. the more teachers are perceived as supporting in-depth learning and self-improvement, the higher the students will achieve. this result highlights that students perceived academic environment has a real impact on the outcomes they are able to achieve. it also suggests that the consideration of teaching practices is paramount to provide a comprehensive depiction of contextual diversity among first-year students in universities. yet, it is worth noting that only mastery goal structure appears to be significant in contrast to performance goal structure and the quality of the relationship with the teacher. these variables are also moderately correlated. a teaching environment characterized by high mastery goal structure also tends to be defined by rather low performance goal structure and good teacher relationship. so, we can postulate that these other aspects of teaching practices are indirectly related on achievement. while mastery goal structure is the most important when investigating academic achievement, it couldn’t be the same with other outcomes variables such as retention or wellbeing. this underdeveloped field of investigation would deserve more attention in the literature. by emphasizing the specificity of the environment, the findings question the generalizability of the results found in one context to other ones. as suggested by millet (2012), several achievement predictors would be context-specific and their effects cannot be directly transferred to another environment. this assumption is supported by significant slope variations found in our analyses, notably about attendance, intention to persist, value and peer support. the differential effect of attendance could be explained by the nature of the assessment procedures. a student registered in a study program which mainly assessed students through evaluations mapping on knowledge (e.g. law program) would not require much behavioural and cognitive participation to perform. it is not the case for a student registered in another study program which employs more complex assessment tasks directly depending on students’ active participation during the lesson (problem-based evaluation task, for example; engineering program). such reasoning was partially corroborated by de clercq, galand and frenay (2013b) who proved that the impact of deep processing strategies differed according to the way the student was assessed. the nature of the courses in the program could explain the differential effect of the value of the course and the students’ intention to persist. students enrolled in programs mainly composed of very general courses (e.g. psychology program) might require a more thorough motivation to succeed than students in very applied and specific courses (e.g. dentistry, veterinary). finally, the differential effect of peer support could be partially explained by the difficulty of the interactions with the students in the program. a study program mainly characterized by lectures held in large auditoriums (e.g. economics) could result in more difficulty for the students to develop social connections than programmes with smaller class sizes (e.g. architecture, history programmes). perceived peer support could, therefore, become crucial for students in these big, rather anonymous learning conditions. 4.3 a widened theoretical perspective on the diversity in the transition to he our findings do support the theoretical assumption of the three-level framework of he (enders, 2004; munge, thomas, & heck, 2018; taylor & ali, 2017) that micro-level characteristics of the transition cannot be fully understood without considering meso-level features of he. this assumption also meets bronfenbrenner’s (1992, 2005) ecological systems model which describes individual development as shaped by multiple, nested layers of context. transferred to the he system, the results emphasize the importance of considering both immediate and more distal environmental settings and how they interact with individual characteristics in order to better understand student transition to he (taylor & ali, 2017; mclinden 2017). moreover, the relevance of the four categories of variables investigated was highlighted, each of them being composed of significant predictors of achievement. this finding lends credence to the proposed conceptual framework. it also supports several other theoretical models such as price’s 4p model of student’s learning (2014) or de clercq and colleagues’ (2020) model of the achievement process which both highlighted the necessity to consider background factors, psychological variables and perception of the context together. more precisely, the combination of individual and contextual variables allows for an explanation of 30 percent of individual achievement variance and almost 50 percent of program variance. the selection of variables can therefore be considered as relevant. in sum, our approach can be considered as a valid modelling of achievement process. yet, our analytical procedure does not allow to fully validate our conceptual framework. the multi-level analysis only focuses on the direct predictors of achievement and did not consider for indirect effects. the temporal unfolding of the achievement process was therefore left out of the scope of the study. there is, thus, room for improvement in the validation and the modelling of the proposed theoretical framework. following this idea, the investigation of diversity should be expanded beyond the micro and meso-level. future research should integrate differences at (1) the micro-level of the individual student experience, (2) the meso-level of the institutional context and (3) the macro-level of the wider education system. 4.4 limitations and further considerations this study’s limitations are threefold. first, a limitation of this study lies in the choice of the variables investigated and the measures employed. the aim was to focus on the contextual nature of academic achievement. to do so, the multilevel analyses allowed us to estimate the magnitude of achievement variation across programs and the explanative power of the perceived learning environment. however, this approach was constrained and did not provide much information about the structural and administrative characteristics of the context. only four curriculum characteristics were investigated whereas a lot of other indicators (e.g., difference in admission procedures, requirements, assessment, workload, administrative functioning,…) could have improved our understanding of the contexts. according to kuh and colleagues (2011), the detailed investigation of these objective characteristics would go a step further in this consideration. to do so, such an approach could be inspired by the above-mentioned dutch studies which specifically addressed this question (jansen, 2004; van der hulst & jansen, 2002). moreover, the subjective experience of the environment was only measured by perceived teaching practices. this could be further investigated by adding the social layer of the environment (schaeper, 2019) and the nature of the learning tasks composing the program (assessment, perceived workload). finally, despite the multifactorial approach of academic achievement, complementary variables could have been considered at the individual level. for instance, learning strategies would have been an interesting variable to include in the scope of the analysis. according to the meta-analysis of richardson and colleagues (2012), adding cognitive factors would be incrementally predictive of achievement. moreover, millet (2003, 2012) found significant variations of the effect of cognitive variables across the program. following this idea, the results of this study should be replicated including cognitive factors in the modelling. second, a question remains concerning the level of investigation adopted. the analyses were carried out at the study program level. conversely, van den berg & hoffman (2005) realized analyses at the course level. such scopes of investigation refer to different questions and realities. further studies could broaden the scope to the university level and provide complementary results on the variations observed. another possibility would be to include different levels of investigation directly in the same analysis. a multilevel analysis endorsing student, course, program and university levels could be of major interest regarding the contextual question of academic achievement. the macro-level of diversity could also have been considered in order to provide a comprehensive investigation. however, such analysis would require a massive international data collection and important collaborative work that lacks in our approach. third, the specific perspective of the paper can be questioned. we postulate that achievement can be conceived as a proxy of a successful transition. therefore, the investigation of the process leading to achievement among the first year at the university was considered as an interesting perspective on student’s transition. yet, this perspective is very specific and leaves several aspects of the transition out of the picture. several authors argued that the transition is a dynamic process involving several outcomes such as achievement, adjustment, retention, learning and personal fulfilment (crider, calder, bunting, & forwell, 2015; dangoisse, de clercq, van meenen, chartier, & nils, 2019; kovač, 2015). therefore, the transition to he cannot be fully understood through the achievement process and through cross-sectional designs (coertjens et al., 2017). an interesting future perspective would be to complement the current approach with a multilevel design addressing other layers of the transition (retention, learning) and considering the longitudinal nature of the first-year experience. in this idea, a study of de clercq and colleagues (2018) proposed several important moments to consider in order to capture the dynamic nature of the first year at the university, such as the very first week of the transition that constitutes a crucial moment of adjustment for the students. in conclusion, the paper highlighted the strong potential of an approach that bridges the micro and meso-level of the diversity to the transition to he. more investigation is needed to unveil the potential of such an approach and move toward an even more fine-tuned understanding of the first-year experience. from a practical point of view, the study highlighted the necessity to consider both individual and contextual heterogeneity when supporting the students in their transition to he. the findings sustained a more program-specific promotion of academic achievement accounting for different patterns of students. keypoints contextual diversity explains 15 percent of academic achievement among the first year at the university. individual predictors of academic achievement significantly varied across study programs. both curriculum characteristics and perceived teaching practices are significant predictors of academic achievement. study programs varied in student body which highlights that micro-level diversity is partially determined by meso-level diversity. footnotes 1 in our educational system, tutorials encompass small group learning activities such as practical exercises and workshops. references allen, j., robbins, s. b., & sawyer, r. (2010). can measuring psychosocial factors promote college success? applied measurement in education, 23(1), 1-22. doi:10.1080/08957340903423503 ames, c. (1992). classrooms: goals, structures, and student motivation. journal of educational psychology, 84(3), 261-271. doi:10.1037/0022-0663.84.3.261. doi:10.1037/0022-0663.84.3.261 anderman, e. m., & patrick, h. (2012). achievement goal theory, conceptualization of ability/intelligence, and classroom climate. in s. l. christenson, a. l. reschly, & c. wylie (eds.), handbook of research on student engagement (pp. 173-191): springer us.doi:10.1080/08957340903423503 astin, a. w. (2012). assessment for excellence: the philosophy and practice of assessment and evaluation in higher education . rowman & littlefield publishers. bandura, a. (1997). self-efficacy: the exercise of control. new-york: freeman. biémar, s., philippe, m.-c., & romainville, m. (2003). l'injonction au projet: paradoxale et infondée ? approche longitudinale du choix d'études supérieures. l'orientation scolaire et professionnelle, 32, 31-51. https://doi.org/10.4000/osp.3167 bong, m. (2005). within-grade changes in korean girls' motivation and perceptions of the learning environment across domains and achievement levels. journal of educational psychology, 97(4), 656-672. https://doi.org/10.1037/0022-0663.97.4.656 bong, m., & skaalvik, e. m. (2003). academic self-concept and self-efficacy: how different are they really? educational psychology review, 15(1), 1-40. https://doi.org/10.1023/a:1021302408382 bourdieu, p. (1986). the forms of capital. in j. richardson (ed.). handbook of theory and research for the sociology of education (pp. 241–258). new york, ny: greenwood. bronfenbrenner, u. (1992). ecological systems theory. london: jessica kingsley publishers. bronfenbrenner, u. (2005) (ed.) making human beings human: bioecological perspectives on human development. thousand oaks, ca: sage publications. bruinsma, m. (2004). motivation, cognitive processing and achievement in higher education. learning and instruction, 14(6), 549-568. retrieved from http://www.sciencedirect.com/science/article/b6vfw-4dtm0pn-1/2/4b3c023a60881ffb6536046597aff795 cabrera, a. f., nora, a., & castaneda, m. b. (1993). college persistence: structural equations modeling test of an integrated model of student retention. the journal of higher education, 64(2), 123-139. retrieved from http://www.jstor.org/stable/2960026 cattonar, b., & mangez, e. (2014). codages et recodages de la réalité scolaire. pisa dans la presse écrite de belgique francophone. revue internationale d’éducation de sèvres, (66), 61-70. https://doi.org/10.4000/ries.3999 chemers, m. m., hu, l.-t., & garcia, b. (2001). academic self-efficacy and first-year college student performance and adjustment. journal of educational psychology, 93(1), 55-64. doi:10.1037//0022-0663.93.1.55 chickering, a. w., & reisser, l. (1993). education and identity. the jossey-bass higher and adult education series . san francisco: jossey-bass inc. church, m. a., elliot, a. j., & gable, s. l. (2001). perceptions of classroom environment, achievement goals, and achievement outcomes. journal of educational psychology, 93(1), 43. doi:10.1037//0022-0663.93.1.43 coertjens, l., brahm, t., trautwein, c., & lindblom-ylänne, s. (2017). students’ transition into higher education from an international perspective. higher education, 73(3), 357-369. doi:10.1007/s10734-016-0092-y credé, m., roch, s. g., & kieszczynka, u. m. (2010). class attendance in college: a meta analytic review of the relationship of class attendance with grades and student characteristics. review of educational research, 80(2), 272-295. doi: 10.3102/0034654310362998 crider, c., calder, c. r., bunting, k. l., & forwell, s. (2015). an integrative review of occupational science and theoretical literature exploring transition. journal of occupational science, 22(3), 304-319. doi:10.1080/14427591.2014.922913 dadeppo, l. m. w. (2009). integration factors related to the academic success and intent to persist of college students with learning disabilities. learning disabilities research & practice, 24 (3), 122-131. doi:10.1111/j.1540-5826.2009.00286.x dangoisse, f., de clercq, m., van meenen, f. v., chartier, l., & nils, f. (2019). when disability becomes ability to navigate the transition to higher education: a comparison of students with and without disabilities. european journal of special needs education, 1-16. doi:10.1080/08856257.2019.1708642 de clercq, m., galand, b., dupont, s., & frenay, m. (2013a). achievement among first-year university students: an integrated and contextualised approach. european journal of psychology of education, 28(3), 641-662. doi:10.1007/s10212-012-0133-6 de clercq, m., galand, b., & frenay, m. (2013b). learning processes in higher education: providing new insights into the effects of motivation and cognition on specific and global measures of achievement. in d. gijbels, v. donche, j. t. e. richardson, & j. d. vermunt (eds.), learning patterns in higher education: dimensions and research perspectives : taylor & francis. de clercq, m., galand, b., & frenay, m. (2017). transition from high school to university: a person-centered approach to academic achievement. european journal of psychology of education, 32(1), 39-59. doi:10.1007/s10212-016-0298-5 de clercq, m., galand, b., & frenay, m. (2020). one goal, different pathways: capturing diversity in processes leading to first-year students' achievement. learning and individual differences. 1-21. doi: https://doi.org/10.1016/j.lindif.2020.101908 de clercq, m., roland, n., brunelle, m., galand, b., & frenay, m. (2018). the delicate balance to adjustment: a qualitative approach of student’s transition to the first year at university.psychologica belgica, 58(1), 67-90. doi: http://doi.org/10.5334/pb.409 dennis, j. m., phinney, j. s., & chuateco, l. i. (2005). the role of motivation, parental support, and peer support in the academic success of ethnic minority first-generation college students. journal of college student development, 46(3), 223-236. doi: 10.1037/0033-295x.98.2.224 díaz, r. j., glass, c. r., arnkoff, d. b., & tanofsky-kraff, m. (2001). cognition, anxiety, and prediction of performance in 1st-year law students. journal of educational psychology, 93(2), 420-429. doi:10.1037/0022-0663.93.2.420 dollinger, s. j., matyja, a. m., & huber, j. l. (2008). which factors best account for academic success: those which college students can control or those they cannot? journal of research in personality, 42(4), 872-885. doi: 10.1016/j.jrp.2007.11.007 eccles, j. s. (2005). subjective task value and the eccles et al. model of achievement-related choices. in c. s. dweck & a. j. elliot (eds.), handbook of competence and motivation (pp. 105-121). new york, ny, us: guilford publications, inc. eccles, j. s., & wigfield, a. (2002). motivational beliefs, values, and goals. annual review of psychology, 53, 109-132. doi:https://doi.org/10.1146/annurev.psych.53.100901.135153 elias, s. m., & macdonald, s. (2007). using past performance, proxy efficacy, and academic self-efficacy to predict college performance. journal of applied social psychology, 37(11), 2518-2531. doi: 10.1111/j.1559-1816.2007.00268.x enders, j. (2004). higher education, internationalisation, and the nation-state: recent developments and challenges to governance theory. higher education, 47(3), 361-382. doi:https://doi.org/10.1023/b:high.0000016461.98676.30 fenollar, p., romajn, s., & cuestas, p. j. (2007). university students' academic performance: an integrative conceptual framework and empirical analysis. british journal of educational psychology, 77(4), 873-891. retrieved from 10.1348/000709907x189118 galand, b., & frenay, m. (2005). l'approche par problèmes et par projets dans l'enseignement supérieur: impact, enjeux et défis. louvain-la-neuve : presses universitaires de louvain. gale, t., & parker, s. (2014). navigating change: a typology of student transition in higher education. studies in higher education, 39 (5), 734-753. doi:10.1080/03075079.2012.721351 georgianna l. martin, matthew j. smith, william c. takewell & allison miller (2020) revisiting our contribution: how interactions with student affairs professionals shape cognitive outcomes during college, journal of student affairs research and practice, 57:2, 148-162, doi: 10.1080/19496591.2019.1631834 germeijs, v., luyckx, k., notelaers, g., goossens, l., & verschueren, k. (2012). choosing a major in higher education: profiles of students’ decision-making process. contemporary educational psychology, 37 (3), 229-239. doi:10.1016/j.cedpsych.2011.12.002 germeijs, v., & verschueren, k. (2007). high school students' career decision-making process: consequences for choice implementation in higher education. journal of vocational behavior, 70(2), 223-241. doi:doi: 10.1016/j.jvb.2006.10.004 greene, b. a., miller, r. b., crowson, h., duke, b. l., & akey, k. l. (2004). predicting high school students' cognitive engagement and achievement: contributions of classroom perceptions and motivation. contemporary educational psychology, 29(4), 462-482. doi:10.1016/j.cedpsych.2004.01.006 hackett, g., betz, n. e., casas, j. m., & rocha-singh, i. a. (1992). gender, ethnicity, and social cognitive factors predicting the academic achievement of students in engineering. journal of counseling psychology, 39(4), 527-538. doi:10.1037/0022-0167.39.4.527 hausmann, l., ye, f., schofield, j., & woods, r. (2009). sense of belonging and persistence in white and african american first-year students. research in higher education, 50(7), 649-669. doi:10.1007/s11162-009-9137-8 heikkilä, a., niemivirta, m., nieminen, j., & lonka, k. (2011). interrelations among university students’ approaches to learning, regulation of learning, and cognitive and attributional strategies: a person oriented approach. higher education, 61(5), 513-529. doi:10.1007/s10734-010-9346-2 husman, j., & lens, w. (1999). the role of the future in student motivation. educational psychologist, 34(2), 113-125. doi:https://doi.org/10.1207/s15326985ep3402_4 jansen, e. p. w. a. (2004). the influence of the curriculum organization on study progress in higher education. higher education, 47(4), 411-435. doi:10.1023/b:high.0000020868.39084.21 karabenick, s. a. (2004). perceived achievement goal structure and college student help seeking. journal of educational psychology, 96(3), 569. doi:10.1037/0022-0663.96.3.569 kovač, v. b. (2015). “transition: a conceptual analysis and integrative model.” in transitions in the field of special education: theoretical perspectives and implications for practice , edited by d. l. cameron and r. thygesen, 19–33. new york, ny: waxmann. kuh, g. d., kinzie, j., schuh, j. h., & whitt, e. j. (2011). student success in college: creating conditions that matter: john wiley & sons. larose, s., robertson, d. u., roy, r., & legault, f. (1998). nonintellectual learning factors as determinants for success in college. research in higher education, 39(3), 275-297. retrieved from http://search.ebscohost.com/login.aspx?direct=true&db=aph&an=771178&site=ehost-live lee, v. e., & burkam, d. t. (2003). dropping out of high school: the role of school organization and structure. american educational research journal, 40(2), 353-393. retrieved from 10.3102/00028312040002353 lent, r. w., & brown, s. d. (2019). social cognitive career theory at 25: empirical status of the interest, choice, and performance models. journal of vocational behavior, 115, 103316. doi: https://doi.org/10.1016/j.jvb.2019.06.004 lens, w., simons, j., & dewitte, s. (2002). from duty to desire. in pajares, f., and urdan, t. (eds.), academic motivation of adolescents, information age. publishing, greenwich, ct, pp. 221–245. lizzio, a., wilson, k., & simons, r. (2002). university students’ perceptions of the learning environment and academic outcomes: implications for theory and practice. studies in higher education, 27, 27-52. doi:10.1080/0307507012009935 9 lizzio, a., wilson, k., & hadaway, v. (2007). university students' perceptions of a fair learning environment: a social justice perspective. assessment & evaluation in higher education, 32(2), 195-213. doi: 10.1080/02602930600801969 maroy, c., & dupriez, v. (2000). la régulation dans les systèmes scolaires : proposition théorique et analyse du cadre structurel en belgique francophone. revue française de pédagogie, (130), 73-87. retrieved august 11, 2020, from www.jstor.org/stable/41201545 marsh, h. w., lüdtke, o., nagengast, b., trautwein, u., morin, a. j., abduljabbar, a. s., & köller, o. (2012). classroom climate and contextual effects: conceptual and methodological issues in the evaluation of group-level effects. educational psychologist, 47(2), 106-124.doi: https://doi.org/10.1080/00461520.2012.670488 mclinden, m. (2017). examining proximal and distal influences on the part-time student experience through an ecological systems theory. teaching in higher education, 22(3), 373-388. doi:10.1080/13562517.2016.1248391 meece, j. l., anderman, e. m., & anderman, l. h. (2006). classroom goal structure, student motivation, and academic achievement. annual reviews psychology, 57, 487-503. doi:10.1146/annurev.psych.56.091103.070258 millet, m. (2003). les étudiants et le travail universitaire [students and academic work ]. lyon : presses universitaires de lyon. millet, m. (2012 ). l’échec des étudiants de premiers cycles dans l’enseignement supérieur enfrance. retours sur une notion ambiguë et descriptions empiriques. in m. romainville & c. michaut (dir.). réussite, échec et abandon dans l’enseignement supérieur (p. 69-88). bruxelles: de boeck. munge, b., thomas, g., & heck, d. (2018). outdoor fieldwork in higher education: learning from multidisciplinary experience. journal of experiential education, 41(1), 39-53. doi:10.1177/1053825917742165 neuville, s., frenay, m., schmitz, j., boudrenghien, g., noël, b., & wertz, v. (2007). tinto's theoretical perspective and expectancy-value paradigm: a confrontation to explain freshmen's academic achievement. psychologica belgica, 47(1), 31-50. retrieved from http://search.ebscohost.com/login.aspx?direct=true&db=psyh&an=2007-17171-002&site=ehost-live nicholson, n. (1990). the transition cycle: causes, outcomes, processes and forms. fisher, in s. & cooper c. l. (eds.) on the move: the psychology of change and transition, new york: john wiley. nuñez, a. m., & kim, d. (2012 ). building a multicontextual model of latino college enrollment: student, school, and state-level effects. the review of higher education, 35(2), 237-263. doi: 10.1353/rhe.2012.0004 oecd. (2012). pisa 2012 results: what students know and can do?. pascarella, e., & terenzini, p. (2005). how college affects students. vol. 2, a third decade of research. san francisco: jossey-bass. patrick, h., ryan, a. m., & kaplan, a. (2007). early adolescents' perceptions of the classroom social environment, motivational beliefs, and engagement. journal of educational psychology, 99(1), 83. retrieved from 10.1037/0022-0663.99.1.83 perna, l. w., & thomas, s. l. (2008). theoretical perspectives on student success: understanding the contributions of the disciplines. ashe-higher education reader report. san francisco, ca: jossey-bass. perry, r. p., hladkyj, s., pekrun, r. h., & pelletier, s. t. (2001). academic control and action control in the achievement of college students: a longitudinal field study. journal of educational psychology, 93 (4), 776-789. doi:10.1037/0022-0663.93.4.776 pintrich, p. r. (1999). the role of motivation in promoting and sustaining self-regulated learning. international journal of educational research, 31, 459-470. doi:10.1016/s1041-6080(99)80007-7 pintrich, p. r. (2003). a motivational science perspective on the role of student motivation in learning and teaching contexts. journal of educational psychology, 95(4), 667-686. doi:10.1037/0022-0663.95.4.667 pintrich, p. r., & de groot, e. v. (1990). motivational and self-regulated learning components of classroom academic performance. journal of educational psychology, 82(1), 33-40. doi:10.1037/0022-0663.82.1.33 pirot, l., & de ketele, j. m. (2000). l'engagement académique de l'étudiant comme facteur de réussite à l'université. étude exploratoire menée dans deux facultés contrastées. revue des sciences de l'éducation, 26(2), 367-394. doi: 10.7202/000127 poropat, a. e. (2009). a meta-analysis of the five-factor model of personality and academic performance. psychological bulletin, 135 (2), 322-338. doi:10.1037/a0014996 price, l. (2014). modelling factors for predicting student learning outcomes in higher education. in d. gijbels, v. donche, j. t. e. richardson, & j. d. vermunt (eds.), learning patterns in higher education: dimensions and research perspectives : taylor & francis. rayle, a. d., kurpius, s. e. r., & arredondo, p. (2006). relationship of self-beliefs, social support, and university comfort with the academic success of freshmen college women. journal of college student retention: research, theory and practice, 8 (3), 325-343. retrieved from rayle@coe.ufl.edu 10.2190/r237-6634-4082-8q18 reyes, m. r., brackett, m. a., rivers, s. e., white, m., & salovey, p. (2012). classroom emotional climate, student engagement, and academic achievement. journal of educational psychology, 104(3), 700. doi:10.1037/a0027268 richardson, m., abraham, c., & bond, r. (2012). psychological correlates of university students' academic performance: a systematic review and meta-analysis. psychological bulletin, 138(2), 353-387. doi:10.1177/0016440206200300910.1177/001644020620030092002-01523-009 rodríguez-hernández, c. f., cascallar, e., & kyndt, e. (2020). socio-economic status and academic performance in higher education: a systematic review. educational research review, 29, 100305. doi: https://doi.org/10.1016/j.edurev.2019.100305 roeser, r. w., midgley, c., & urdan, t. c. (1996). perceptions of the school psychological environment and early adolescents' psychological and behavioral functioning in school: the mediating role of goals and belonging. journal of educational psychology, 88(3), 408-422. doi:10.1037/0022-0663.88.3.408 ryan, r. m., & deci, e. l. (2017). self-determination theory: basic psychological needs in motivation, development, and wellness : guilford publications. schaeper, h. (2019). the first year in higher education: the role of individual factors and the learning environment for academic integration. higher education. doi:10.1007/s10734-019-00398-0 sackett, p. r., kuncel, n. r., arneson, j. j., cooper, s. r., & waters, s. d. (2009). does socioeconomic status explain the relationship between admissions tests and post-secondary academic performance? psychological bulletin, 135(1), 1-22. doi:10.1037/a0013978 schneider, m., & preckel, f. (2017). variables associated with achievement in higher education: a systematic review of meta-analyses. psychological bulletin, 143(6), 565. doi:https://doi.org/10.1037/bul0000098 sirin, s. r. (2005). socioeconomic status and academic achievement: a meta-analytic review of research. review of educational research, 75(3), 417-453. doi: 10.3102/00346543075003417 skinner, e. a., wellborn, j. g., & connell, j. p. (1990). what it takes to do well in school and whether i've got it: a process model of perceived control and children's engagement and achievement in school. journal of educational psychology, 82(1), 22-32. doi: http://dx.doi.org/10.1037/0022-0663.82.1.22 tabachnick bg and fidell ls (2007) using multivariate statistics. fifth edition. boston: pearson education inc. taylor, g., & ali, n. (2017). learning and living overseas: exploring factors that influence meaningful learning and assimilation: how international students adjust to studying in the uk from a socio-cultural perspective. education sciences, 7(1), 35. https://doi.org/10.3390/educsci7010035 tinto, v. (1997). classrooms as communities. exploring the educational character of student persistence. journal of higher education, 68 (6), 599-623. doi:10.1080/00221546.1997.11779003 torenbeek, m., jansen, e., & hofman, a. (2010). the effect of the fit between secondary and university education on first‐year student achievement. studies in higher education, 35(6), 659-675. doi:10.1080/03075070903222625 torres, j. b., & solberg, v. s. (2001). role of self-efficacy, stress, social integration, and family support in latino college student persistence and health. journal of vocational behavior, 59(1), 53-63. retrieved from 10.1006/jvbe.2000.1785 vallerand, r.-j., fortier, m.-s., & guay, f. (1997). self-determination and persistence in a real-life setting: toward a motivational model of high school dropout. journal of personality and social psychology, 72 (5), 1161-1176. doi:10.1037/0022-3514.72.5.1161 vandamme, j.-p., superby, j. f., & meskens, n. (2005). freshers’achievement : prediction methods and influent factors. paper presented at the higher education, multijuridictionality and globalisation conference, mons, belgique. van den berg, m. n., & hofman, w. h. a. (2005). student success in university education: a multi-measurement study of the impact of student and faculty factors on study progress. higher education, 50(3), 413-446. retrieved from http://www.jstor.org/stable/25068105 van der hulst, m., & jansen, e. (2002). effects of curriculum organisation on study progress in engineering studies. higher education, 43(4), 489-506. doi:10.1023/a:1015207706917 van petegem, p., coertjens, l., donche, v., & noyens, d. (2017). transitions to higher education: moving beyond quantity. in e. kyndt, et al (eds), higher education transitions: theory and research (pp. 3-12). new york: routledge. van rooij, e., brouwer, j., fokkens-bruinsma, m., jansen, e., donche, v., & noyens, d. (2018). a systematic review of factors related to first-year students success in dutch and flemish higher education. pedagogische studiën, 94(5), 360-405. retrieved from https://repository.uantwerpen.be/docman/irua/cebc4c/149722.pdf. van rooij, e. c. m., jansen, e. p. w. a., & van de grift, w. j. c. m. (2018). first-year university students’ academic success: the importance of academic adjustment. european journal of psychology of education, 33(4), 749-767. doi:10.1007/s10212-017-0347-8 vermunt, j. (2005). relations between student learning patterns and personal and contextual factors and academic performance. higher education, 49(3), 205-234. doi:10.1007/s10734-004-6664-2 zusho, a., karabenick, s. a., bonney, c. r., & sims, b. c. (2007). contextual determinants of motivation and help seeking in the college classroom . in r. p. perry & j. c. smart (eds.), the scholarship of teaching and learning in higher education: an evidence-based perspective (pp. 611-659). dordrecht: springer netherlands. codepen jansen frontline learning research vol. 9 no. 1 (2021) 44 65 issn 2295-3159 don't just judge the spelling! the influence of spelling on assessing second-language student essays thorben jansen1, cristina vögelin2,nils machts1, stefan keller2 & jens möller1 1institute of psychology of learning and instruction, kiel university, germany 2institute for educational sciences, university of basel, switzerland article received 27 september 2019 / revised 17 december 2020/ accepted 21 december / available online 11 february 2021 abstract when judging subject-specific aspects of students’ texts, teachers should assess various characteristics, e.g., spelling and content, independently of one another since these characteristics are indicators of different skills. independent judgments enable teachers to adapt their classroom instruction according to students’ skills. it is still unclear how well teachers meet this challenge and which intervention could be helpful to them. in study 1, n = 51 pre-service teachers assessed four authentic english as a second language (esl) essays with different overall text qualities and different qualities of spelling using holistic and analytic rating scales. results showed a negative influence of the experimentally manipulated spelling errors on the judgment of almost all textual characteristics. in study 2, an experimental prompt was used to reduce this judgment error. participants who were made aware of the judgment error caused by spelling errors formed their judgments in a less biased way, indicating a reduction of bias. the determinants of the observed effects and their practical implications are discussed. keywords: teacher judgments; second language writing; halo effect; spelling, writing assessment info corresponding author: email: tjansen@ipl.uni-kiel.de doi: https://doi.org/10.14786/flr.v9i1.541 1. introduction for the last two decades, second language education has received increasing attention in europe. european citizens are expected to be proficient in two second-languages in addition to their first language when they complete secondary education (european commission, 2008). thus, there is a need for high-quality second language instruction in schools and close monitoring of its effectiveness. an important part of second language education is writing instruction, as writing is a key competence for higher education in the globalized world in general and tertiary education in particular (keller, 2013). the importance of writing, in turn, requires competent teachers who have a good grasp of text quality, know how to give helpful feedback, and can assign fair and objective judgments based on transparent criteria. accurate teacher judgments of students’ achievement are, therefore, an important aspect of effective instruction (elliott, lee, & tollefson, 2001), especially when teachers align their lessons with the perceived competencies of their students (brookhart, 2011, 2013; herppich et al., 2017; ruiz‐primo & furtak, 2007). also, accurate teacher judgments serve as the primary source of information for students to evaluate their performances and affect their self-perception (zimmermann, möller, & köller, 2018). in the case of assessing students’ second-language essays, there is scant research on teacher judgments. previous research on teacher judgments neglected subjectand genre-specific achievements and focussed on teachers’ ability to judge students’ competencies in general (südkamp, kaiser, & möller, 2012). there are aspects specific to second language writing assessment which might distort teacher judgments (vögelin, jansen, keller, machts, & möller, 2019). for example, teachers need to distinguish between several independent characteristics, such as spelling, content, or organization, when assessing students’ essays (bae & bachman, 2010; cooksey, freebody, & wyatt-smith, 2007; flower & hayes, 1981; hyland, 2008). empirical studies have found that textual characteristics, such as the number of spelling errors, influence teacher judgments of other independent characteristics, which is defined as a halo effect (murphy & reynolds, 1988; saal, downey, & lahey, 1980). the term halo effect describes the impact of one characteristic on the judgment of an independent characteristic. a halo effect occurs, for example, when the number of spelling errors affects content-related judgments of a text. halo effects are particularly likely in the assessment of english as a second language (esl) student texts since these texts typically contain more spelling errors than first-language student texts (flor, futagi, lopez, & mulholland, 2015), and because teachers focus more on those errors (cumming, kantor, & powers, 2002). when writing in a second language, students face larger challenges at the formal level (spelling, grammar, etc.) in addition to the content-related level (content, organization, etc.). as a result, the difference between content-related and formal criteria of second-language texts is typically greater than in first-language texts (hamp-lyons, 1991; weigle, 2002), and the correlation between judgments of these levels is typically lower in second-language than in first-language texts (bae & bachman, 2010; wind, stager, & patil, 2017). however, no studies investigated those halo effects on the judgment of esl student texts resulting from spelling. such halo effects could have serious consequences for learners as they could lead to misjudgments of other writing aspects, such as content, organization, argumentation, or structure of an essay, and thus lead to students losing motivation (urhahne, 2015; zhou & urhahne, 2013) and confidence (artelt, 2016). additionally, they could be a major problem in language instruction in school as students with a deficit or strength in formal language were in danger of being underor overestimated by their teachers additionally, they could be a major problem in language instruction in school as students with a weakness or strength at the formal level would be in danger of being underor overestimated by their teachers. students might receive wrong grades or teachers might not plan subsequent lessons adequately when aligning their teaching to students’ perceived competencies. in the worst case, students could be assigned to unsuitable educational tracks based on biased scoring (schrader, 2013). furthermore, concerning many written high-stakes exams in foreign education in secondary school and standardized second-language tests, halo effects from spelling could be crucial for certain students erroneously not being admitted to tertiary education. recent reviews and meta-analyses on teachers’ judgment accuracy have stressed the urgent need for intervention studies on improving the quality of teachers’ judgment (kaufmann, 2020; urhahne & wijnia, 2021). facilitating diagnostic competencies is a relatively new field with only nine empirical studies, none of which contain concrete interventions or suggesting methods to foster the quality of teachers’ judgments on students’ written performances (chernikova et al., 2019). to strengthen the literature, we used a two-step approach to detect a judgment bias in the first study and reduce it in the second. the two studies presented in this paper represent each of the steps. in study 1, we examined whether the number of spelling errors in esl argumentative essays influences pre-service teachers’ assessments of other textual characteristics using genre-specific analytic rating scales. in study 2, we adapted an intervention from the research of professional rater training to the context of teacher education to investigate whether distortions in judgment can be reduced using a prompt alerting teachers’ to possible halo effects. 2. the influence of spelling errors on the assessment of students’ texts teacher assessment of students’ achievement is moderated by (i.e., varies as a function of) a number of variables. the heuristic model of judgment accuracy by südkamp et al. (2012) systematizes these moderators of teacher judgment accuracy, which is defined as the correspondence between teacher judgments of students’ achievement and students’ actual achievement. the model systematizes the moderators into four factors: it differentiates teacher characteristics (e.g., teaching experience, specialist knowledge, and pedagogical knowledge), test characteristics (e.g., relevant characteristics of a test, genre of a test, reliability of a test), student characteristics (e.g., performance, gender, motivation, age), and characteristics of the judgment to be made (e.g., the specificity of the domain to be judged, number of levels on a rating scale). when it comes to assessing writing, a special judgment characteristic is that teachers should use multiple pre-set scoring criteria and make the criteria available to the students (huot, 1996). hence, it is necessary to define which text characteristics should and should not influence raters’ judgments a rating scale. raters considering text characteristics in the way that is defined on the rating scale contributes to a correct judgment, whereas considering characteristics contrary to a rating scale – or confusing different aspects – contributes to a distorted judgment. for example, the number of spelling errors should influence the judgment of general text quality and spelling of a text; however, it should not influence the judgments of the content or organization (parr & timperley, 2010). the heuristic model of judgment accuracy by südkamp et al. (2012) describes which factors moderate teacher judgment accuracy. however, because the described model is a heuristic model, it does not describe the judgments’ estimation process. it is useful to investigate the estimation process of judgments to investigate which factors influence judgment accuracy. dual-process models of social information processing describe the estimation and the information processing that underlies it (herppich et al., 2017; karst, dotzel, & dickhäuser, 2018), such as the continuum model of impression formation processes (fiske & neuberg, 1990). this model describes individuals’ information processing as a continuum between a heuristic strategy and a controlled strategy to process information. the heuristic or shortcut strategy is assumed to be relatively automatic and to require little cognitive effort. for example, fiske and neuberg (1990) showed that individuals tend to use more heuristic information processing when they interpret target attributes to fit well into a category. using this strategy leads to a simplification of the information in order to handle the complexity of the judgment more easily, thus increasing the possibility of halo effects. using the controlled strategy, the individual collects more information and uses algebraic rules to integrate the information into a judgment. this strategy should minimize halo effects and will be used if individuals attend closely to target attributes. empirical studies have shown that teachers collect information about text characteristics in student texts and integrate them correctly into their judgment. student texts with few spelling errors were assessed more positively (birkel & birkel, 2002) compared to texts with numerous spelling errors. teachers also assess texts’ general quality more positively if they contained fewer errors, regardless of whether participants read only one (rafoth & rubin, 1984; rezaei & lovorn, 2010) or 24 texts (barkaoui, 2010). further, empirical studies have shown that textual characteristics influencing the judgment on scales that should not be influenced by the respective text characteristics indicate heuristic information processing. for example, teachers who were supposed to judge the content only assessed texts with several spelling errors as more negative concerning content than texts with fewer spelling errors (marshall, 1967; rezaei & lovorn, 2010; scannell & marshall, 1966). scannell and marshall (1966) and marshall (1967) asked students and teachers to judge the overall quality of a history essay solely based on its contents. the authors prepared several versions of the essay, which only differed in spelling, punctuation, and grammatical errors. texts with more than three spelling errors per 100 words were assessed more negatively with regard to content than texts with three or fewer spelling errors. similarly, rezaei and lovorn (2010) found that the number of spelling, grammatical, and punctuation errors influenced the judgment of the content, although teachers had been explicitly instructed to assess content only. the influence of spelling errors was also seen in the judgment of esl student writing. raforth and rubin (1984) and sweedler-brown (1993) prepared different versions of one or six second-language essays by students that differed only in the number of spelling errors. the participants were supposed to assess content, organization, structure, and grammar in addition to the overall quality. in both studies, the number of spelling errors negatively affected the judgment of all assessment dimensions. thus, several studies have shown that the number of spelling errors influences the assessment of textual characteristics, even if they should not. however, no study exists on the assessment of esl argumentative essays in an upper-secondary school context. the results of these studies are only applicable to a limited extent since the studies differ from an esl school context concerning teachers’ characteristics, students’ characteristics, and the judgment to be made, which could influence the assessment (südkamp et al., 2012). in many studies, the texts did not originate from pupils but from university students (barkaoui, 2010; freedman, 1979; rafoth & rubin, 1984; rezaei & lovorn, 2010; sweedler-brown, 1993) and professional raters but teachers assessed them (freedman, 1979; rafoth & rubin, 1984; sweedler-brown, 1993; wolfe, song, & jiao, 2016). moreover, unlike in a school context, participants lacked the opportunity to compare texts because they only read a single text (marshall, 1967; rafoth & rubin, 1984; rezaei & lovorn, 2010; scannell & marshall, 1966). this paper presents two studies focusing on fostering teachers’ diagnostic competencies by reducing halo-effects caused by spelling errors on other aspects in the assessment of advanced argumentative l2 essays from upper-secondary school. we examined two sequential research questions with two sequential studies using the same material. firstly, do halo effects distort pre-service teachers’ ratings of students’ texts? secondly, can a prompt reduce this halo-effect? study 1 addressed the phenomenon in an experimentally controlled within-subject design, examining how different numbers of spelling errors affected the judgment on rating scales that should and should not be influenced by spelling. study 2 examined whether a prompt could reduce the erroneous influence of spelling errors. as we show in detail below, we divided the participants in study 2 into a control and an intervention group. participants in both groups were presented the same material as in study 1. participants in the intervention group additionally received a prompt that alerted them to possible halo-effects arising from the number of spelling errors. 3. study 1 this study aimed to investigate whether and to what extent spelling influences the assessment of different textual characteristics. for this, participants assessed four esl argumentative essays of higher or lower overall quality in an experimental study. the research team prepared two versions of each text that differed only in the number of spelling errors. pre-service teachers were then asked to evaluate these texts with a holistic scale for the texts’ overall assessment and seven analytic scales with detailed descriptors of the different levels. we tested the following hypotheses: hypothesis 1: texts of low overall quality are assessed more negatively on the holistic and all analytic scales than texts of higher quality. hypothesis 2: texts with more spelling errors are assessed more negatively than texts with few spelling errors on the scale that should be influenced by spelling. hypothesis 3: texts with few spelling errors are assessed more positively than texts with many spelling errors on the scales that should not be influenced by spelling (halo effect). 3.1 method participants assessed student texts in a digital instrument called the student inventory asset (see figure 1; vögelin, jansen, keller, & möller, 2018). this instrument was developed on the basis of earlier work by kaiser, möller, helm, and kunter (2015). in this computer-based tool, participants read student texts on the left-hand side and see the rating scales displayed on the screen’s right-hand side. each participant assessed four english student texts (two with high and two with low overall text quality as well as two with few spelling errors and two with many spelling errors). texts were presented in a randomized sequence. participants first assessed the texts’ general quality on a holistic scale before they assessed seven individual characteristics on analytic scales. figure 1. assessment in the student inventory asset. 3.2 sample n = 51 pre-service teachers participated in this study. the required sample size for analyzing the effects within the subjects was calculated with g*power (faul, erdfelder, lang & buchner, 2007). based on other findings with the student inventory asset (jansen, vögelin, machts, keller, & möller, 2019; kaiser et al., 2015; vögelin et al., 2019), a moderate effect size of d = 0.60 was expected so that, at a power of b = .90, the required sample size was n = 33. the samples consisted of pre-service english teachers who were recruited from master seminars at universities kiel and basel. the average age of the participants was m = 29.41 (sd = 7.25) years. 62.7% of the subjects were female. 3.3 variables we used a 2x2 experimental design with two independent variables that were varied within the subjects: overall quality (low vs. high) and a number of spelling errors (“few errors” vs. “many errors”). the dependent variables were the holistic and analytic assessments. 3.3.1 text quality text quality varied in two levels: “high” and “low”. we used expert ratings to operationalize text quality, which is a common approach in writing research (meadows & billington, 2010; royal-dawson & baird, 2009; scanell & marshall, 1966). experts from the school of teacher education (location anonymized) evaluated 15 student texts together with the asset research team using the naep rating scale (driscoll, avallone, orr, & crovo, 2010). as a result, we chose two “stronger” and two “weaker” texts from the sample, each of roughly equivalent overall quality and adjusted them to the same text length. the weaker texts exhibited levels of work that received a failing grade, i.e., considered insufficient by all experts regarding the unit’s learning goals. the stronger texts that were chosen showed levels of work that were considered to have surpassed the learning goals and received good to excellent ratings by experts. the texts came from students who had been learning english as a second language for five years and had been instructed for four weeks on the form, structure, and content of argumentative essays. at the end of the teaching unit, the students wrote an argumentative essay in 90 minutes answering the following writing prompt: “do you agree or disagree with the following statement? as humans are becoming more dependent on technology, they are gradually losing their independence.” all texts were of similar length (between 450 and 465 words) since text length significantly influences text assessment (wolfe et al., 2016). 3.3.2 number of spelling errors this study included both spelling and punctuation errors in our manipulation since relevant rubrics often list spelling and punctuation issues in the same category (usually “language mechanics”). we varied the variable within the texts similar to raforth and rubin (1984) and sweedler-brown (1993) in two steps (“few errors” vs. “many errors”). in the text version with few errors, spelling errors were reduced to one error per 100 words, whereas in the text version with many errors, spelling errors were raised to seven errors per 100 words. the numbers (frequency) of spelling errors were derived from the mean number of spelling errors in the high and low-quality texts of corpus overall. thus, we aimed to ensure that the manipulation occurred within the range of the students’ language competencies. table 1 displays the types of spelling and punctuation errors we integrated into text variations with a low quality of spelling. table 1 types of spelling and punctuation errors 3.3.3 text assessment participants assessed each text on the six-level holistic scale of the national assessment of educational progress (driscoll et al., 2010). this study further employed genre-specific analytic rating scales (for an overview, see jansen, vögelin, machts, keller, köller, & möller, 2021). these seven scales were based on the 6 + 1 trait model (culham, 2003) as well as the test in english for educational purposes (teep) (weir, 1988), and we adjusted them to address genre-specific characteristics of argumentative essays (hyland, 1990; zemach & stafford-yilmaz, 2008). the seven scales were frame of essay: introduction and conclusion, body of essay: internal organization of paragraphs, support of arguments, spelling and punctuation, grammar, vocabulary, and overall task completion. each dimension contained four levels with detailed descriptors, with higher levels indicating a positive assessment (see appendix). we used two two-factor multivariate variance analyses with repeated measurements (manova) to analyze the data with subsequent post-hoc tests. in doing so, we conducted two separate manova, one for the scales that should be influenced by spelling (holistic scale, spelling and punctuation) and one for the scales that should not be influenced by spelling (frame of essay, body of essay, support of arguments, grammar, vocabulary, and overall task completion). when an effect of spelling occurred on scales that should be influenced by the spelling, it showed that teachers were able to recognize the spelling errors. by contrast, the analyses were considered as halo-effects when an effect of spelling occurred on scales that should not be influenced by the spelling. we also analyzed the data using nonparametric tests, which confirmed our results. therefore, we only present the manova results. 3.4 results table 2 shows descriptive results of the assessments. the analysis of the two scales that should be influenced by spelling (holistic scale and spelling and punctuation) showed significant multivariate main effects for spelling errors (f(2, 49) = 66.03, p < .001) and text quality (f(2, 49) = 65.56, p < .001), and no interaction between text quality and spelling errors (f(2, 49) = 1.93, ns) for judgments. results of univariate post-hoc tests (see table 2) showed effects for spelling and text quality on both scales. low-quality texts were judged more negatively than high-quality texts, and texts with many spelling errors were judged more negatively than texts with few spelling errors. our findings thus supported hypotheses 1 and 2. the analyses of the scales that should not be influenced by spelling (frame of essay, body of essay, support of arguments, grammar, vocabulary, and overall task completion) also showed significant multivariate main effects of text quality (f(6, 45) = 32.30, p < .001), spelling errors (f(6, 45) = 7.41, p < .001), and no interaction between text quality and spelling errors (f(6, 45) = 1.45, ns). the results showed that texts of low quality were judged significantly more negatively on all scales than texts of high quality (see table 2). this result further supported hypothesis 1. the post-hoc analysis for the effect of spelling errors showed that texts with many spelling errors were judged more negatively than texts with few spelling errors on five out of six scales (the exception was frame of essay) that should not be influenced by spelling. these findings thus supported hypothesis 3. 3.3.3 text assessment participants assessed each text on the six-level holistic scale of the national assessment of educational progress (driscoll et al., 2010). this study further employed genre-specific analytic rating scales (for an overview, see jansen, vögelin, machts, keller, köller, & möller, 2021). these seven scales were based on the 6 + 1 trait model (culham, 2003) as well as the test in english for educational purposes (teep) (weir, 1988), and we adjusted them to address genre-specific characteristics of argumentative essays (hyland, 1990; zemach & stafford-yilmaz, 2008). the seven scales were frame of essay: introduction and conclusion, body of essay: internal organization of paragraphs, support of arguments, spelling and punctuation, grammar, vocabulary, and overall task completion. each dimension contained four levels with detailed descriptors, with higher levels indicating a positive assessment (see appendix). we used two two-factor multivariate variance analyses with repeated measurements (manova) to analyze the data with subsequent post-hoc tests. in doing so, we conducted two separate manova, one for the scales that should be influenced by spelling (holistic scale, spelling and punctuation) and one for the scales that should not be influenced by spelling (frame of essay, body of essay, support of arguments, grammar, vocabulary, and overall task completion). when an effect of spelling occurred on scales that should be influenced by the spelling, it showed that teachers were able to recognize the spelling errors. by contrast, the analyses were considered as halo-effects when an effect of spelling occurred on scales that should not be influenced by the spelling. we also analyzed the data using nonparametric tests, which confirmed our results. therefore, we only present the manova results. 3.4 results table 2 shows descriptive results of the assessments. the analysis of the two scales that should be influenced by spelling (holistic scale and spelling and punctuation) showed significant multivariate main effects for spelling errors (f(2, 49) = 66.03, p < .001) and text quality (f(2, 49) = 65.56, p < .001), and no interaction between text quality and spelling errors (f(2, 49) = 1.93, ns) for judgments. results of univariate post-hoc tests (see table 2) showed effects for spelling and text quality on both scales. low-quality texts were judged more negatively than high-quality texts, and texts with many spelling errors were judged more negatively than texts with few spelling errors. our findings thus supported hypotheses 1 and 2. the analyses of the scales that should not be influenced by spelling (frame of essay, body of essay, support of arguments, grammar, vocabulary, and overall task completion) also showed significant multivariate main effects of text quality (f(6, 45) = 32.30, p < .001), spelling errors (f(6, 45) = 7.41, p < .001), and no interaction between text quality and spelling errors (f(6, 45) = 1.45, ns). the results showed that texts of low quality were judged significantly more negatively on all scales than texts of high quality (see table 2). this result further supported hypothesis 1. the post-hoc analysis for the effect of spelling errors showed that texts with many spelling errors were judged more negatively than texts with few spelling errors on five out of six scales (the exception was frame of essay) that should not be influenced by spelling. these findings thus supported hypothesis 3. table 2 means, standard deviations, and univariate analyses of variance in study 1 for the variables spelling errors (“few errors” vs. “many errors”) and text quality (“high” vs. “low”) 3.5 discussion study 1 showed that pre-service teachers correctly differentiated between two quality levels of spelling and overall text quality when evaluating learners’ writing competencies. this result is in line with previous studies showing similar results: teachers collect information about text characteristics in student texts and integrate them correctly into their judgment (birkel & birkel, 2002; barkaoui, 2010; rafoth & rubin, 1984; rezaei & lovorn, 2010). more importantly, it demonstrated that spelling errors had a considerable influence on teacher judgments when they assessed text characteristics, which should be evaluated separately from spelling errors. all text characteristics, besides spelling and punctuation, were identical in both versions, yet they were assessed more negatively when texts included more spelling errors. pre-service teachers’ judgments regarding the characteristics body of essay, support of arguments, grammar, vocabulary, and overall task completion were all subject to halo effects. according to the continuum model (fiske & neuberg, 1990), these halo effects could indicate the use of a heuristic information processing strategy. we assume that to assess these characteristics, the pre-service teachers needed to focus on all areas of the text simultaneously, which made it hard to pay attention to the individual aspects of assessment and thus more difficult to distinguish between the analytic criteria. only the characteristic frame of essay (i.e., whether a text had a clear introduction and conclusion) was not influenced by the number of spelling errors, possibly indicating a stronger use of the controlled information processing strategy. the category frame of essay may be less prone to halo effects since its assessment is limited to the introduction and conclusion of the text, which made it easier to attend focus on the specific descriptors while excluding distractions of spelling. this might explain why pre-service teachers employed more controlled information processing when assessing this particular text characteristic. one could debate whether it is theoretically possible for teachers—or raters—to disentangle different rating criteria completely (sadler, 2009). while rating criteria should be non-overlapping, there is ample research showing that empirical scores for rating criteria intercorrelate (huot, 1996). in terms of formative feedback and judgment validity, this overlap has limits, however. when the quality of spelling negatively influences the assessment of structure or argumentation in an essay, this indicates a halo-effect and a distortion of judgment, which is in line with previous research (marshall, 1967; rezaei & lovorn, 2010; scannell & marshall, 1966). telling a student to improve the overall quality of a text by correcting the spelling is sound advice. telling a student that her essay uses poor argumentation when the problem is poor spelling, by contrast, is unsuitable or misleading feedback. therefore, it is a key task for teacher education to alert teachers to such distortion effects and show them ways of avoiding or reducing them. according to the continuum model (fiske & neuberg, 1990), the use of the heuristic strategy can be reduced if evaluators spend more attention on target attributes of the assessment, possibly preventing halo effects. this could be encouraged by a prompt instructing evaluators to assess every analytic scale as a different, independent category. previous studies show that additional attention to the assessment’s target attributes increases the reliability (lovorn & rezaei, 2011) and accuracy (chamberlain & taylor, 2011; dempsey, pytlikzillig, & bruning, 2009; meadows & billington, 2010) of the assessment based on analytic scales. a recent meta-analysis on fostering diagnostic competence of pre-service teachers (chernikova et al., 2019) identified large positive effects for the prompts on pre-service teachers’ judgments. however, the meta-analysis contained only nine studies, including no study fostering the assessment of students’ written performances. to fill this research gap, study 2 examines whether a prompt could reduce the halo-effect of spelling. 4. study 2 study 2 tested an intervention to reduce the halo effect found in study 1. the participants in study 2 were randomly split into a control group, which replicated study 1, and an experimental group. before the assessment, the participants in the experimental group were given a verbal prompt that instructed them to pay attention to the possible influence of spelling errors on their assessment of other text characteristics. this study tested the following hypotheses: hypothesis 1: texts of low overall quality are assessed more negatively on the holistic and all analytic scales than texts of higher quality. hypothesis 2: texts with more spelling errors are assessed more negatively than texts with few spelling errors on the scale that should be influenced by spelling. hypothesis 3: texts with few spelling errors are assessed more positively than texts with many spelling errors on the scales that should not be influenced by spelling (halo effect). hypothesis 4: the prompt reduces the halo effect of spelling. 4.1 method 4.2 sample n = 66 pre-service teachers participated in study 2. the required sample size for analyzing the interaction between effects within and between the subjects was calculated with g*power (faul, erdfelder, lang, & buchner, 2007). the halo effect size (d = 0.63) averaged across all scales found in study 1 was expected so that, at a power of b = .90, a sample size of n = 54 was required. the samples consisted of pre-service teachers of english, who were recruited from seminars at universities (locations anonymized). the average age of the participants was m = 24.31 (sd = 4.90) years, and 66% were female. the sample group was randomly split into two groups: the “prompt group” (n = 30) and the “control group” (n = 36). 4.3 variables 4.3.1 text quality we employed the same four texts of two quality levels that were used in study 1. 4.3.2 number of spelling errors as in study 1, we varied the number of spelling errors at two levels: “few errors” (1 error per 100 words) and “many errors” (7 errors per 100 words). 4.3.3 prompt in the student inventory asset, participants of the treatment group were shown the following prompt: “empirical studies have shown that teachers are influenced to a large degree by spelling errors while assessing text quality” before being asked to assess the student texts. the control group saw the prompt “please judge the texts in a balanced and fair way”. the first prompt was thus specifically aimed at alerting participants to possible halo effects emanating from spelling, while the second prompt asked for fair and unbiased assessment only in a very general fashion. 4.3.4 text assessment as in study 1, participants assessed each of the four texts on the six-level holistic scale of the national assessment of educational progress (driscoll et al., 2010) and on seven four-level genre-specific analytic scales (culham, 2003; weir, 1988). 4.4 results we used two three-factor multivariate variance analyses with repeated measurements to analyze the data with subsequent contrast tests. in doing so, we reported the results separately for the scales that should or should not be influenced by spelling. table 3 shows the descriptive parameters of study 2. the analyses of the scales that should be influenced by spelling (holistic scale and spelling and punctuation) showed significant multivariate main effects for the number of spelling errors (f(2, 63) = 146.40, p < .001) and the overall quality of the texts (f(2, 63) = 97.60, p < .001), and no main effect of the prompt (f(2, 63) = 1.13, ns). the multivariate analysis further showed no interaction effects for the text quality and the number of spelling errors (f(2, 63) = 0.24, ns), the text quality and the prompt (f(2, 63) = 0.15, ns), the number of spelling errors and the prompt (f(2, 63) = 2.46, ns), and no three-way interaction between the factors of spelling errors, quality, and prompt (f(2, 63) = 1.58, ns). the results of the post-hoc analyses showed univariate effects for spelling and for text quality on both scales in the assumed direction (see table 4). as in study 1, the results supported hypothesis 1 and 2. additionally, the analyses of the scales that should not be influenced by spelling (frame of essay, body of essay, support of arguments, grammar, vocabulary, and overall task completion) showed significant multivariate main effects for the number of spelling errors (f(6, 59) = 28.43, p < .001), the quality of the texts (f(6, 59) = 48.43, p < .001), and the prompt (f(6, 59) = 1.32, ns). most importantly, as expected, there was also a significant interaction between the number of spelling errors and the prompt (f(6, 59) = 2.34, p < .05). the multivariate analysis showed no interaction between the spelling errors and the quality (f(6, 59) = 0.66, ns), the quality and the prompt (f(6, 59) = 0.64, ns), or the number of spelling errors, the quality, and the prompt (f(6, 59) = 1.42, ns). univariate post-hoc tests were calculated for significant multivariate effects on the scales that should not be influenced by spelling (see table 4). for text quality, results indicated that texts of low overall quality were assessed more negatively on all scales than texts of high quality. these findings supported hypothesis 1. the effect sizes were, again, consistently high. similarly, texts with many spelling errors were assessed more negatively on all scales than were texts with few spelling errors. the data in this study also supported the appearance of halo effects formulated in hypothesis 3: pre-service teachers’ judgments on all scales were influenced by the number of spelling errors. the test for the effect of the prompt (hypothesis 4) was a one-sided, post-hoc analysis of the interaction between spelling and prompt. it showed significant results on the scales support of arguments, vocabulary, and overall task completion. the prompt reduced the halo effect of the number of spelling errors on these scales (see table 3). table 3 means and standard deviation in study 2 for the variables spelling errors (“few errors” vs. “many errors”) and text quality (“high” vs. “low”); split for the prompt group and the control group table 4 post-hoc analyses of the multivariate spelling, text quality, and prompt*spelling effects. 4.5 discussion study 2 aimed to test whether a prompt was effective at reducing halo-effects of spelling on other independent aspects of student writing, thus improving the quality of their assessments. participants saw a prompt that alerted them to possible distortion effects of spelling before the assessment. we first examined whether the effects of text quality and spelling errors, including halo-effect, occurred in the same way as in study 1. results showed that pre-service teachers considered text quality and spelling errors when assessing esl argumentative essays, replicating results from study 1: texts of low overall quality, and many spelling errors, were assessed more negatively with regard to these two characteristics than texts of high quality and with few spelling errors. moreover, the results showed that the number of spelling errors influenced teachers’ judgments on all scales that should not be influenced by spelling errors. hence, a halo effect was seen on all assessment scales and across all participants. these findings are in line with study 1. this replication is a particular strength of the study and an important part of good scientific practice, especially in the so-called “replication crisis” (bakker, van dijk, & wicherts, 2012; open science collaboration, 2015). study 2 also expanded study 1 by implementing a prompt to reduce halo effects. the prompt had an effect on participants’ assessment of the scales support for arguments, vocabulary, and overall task completion. however, it did not have a significant effect on the assessment of frame of essay, body of essay, and grammar. interpreting this fact is one of the more challenging aspects of the study. we used the continuum model of impression formation processes (fiske & neuberg, 1990) as a theoretical model, in which information processing is described as a continuum between a heuristic and a controlled strategy to process information. the model predicts that individuals will use more heuristic information processing when target attributes are interpreted to fit well into a category. one could argue that participants found it easier to judge support for arguments, vocabulary, and overall task completion because they had clear notions of the amount of support, number of topic-specific words, or conventions that were required for this type of text. this would mean that it was possible for them to apply the controlled strategy of information processing and still handle the complexity of the judgment on these scales. by contrast, one could argue frame of essay, body of essay, and grammar were more difficult to judge because they were less clearly defined or less familiar to this type of participant (pre-service teachers). thus, these scales would have been too complex for using the controlled strategy, rendering the prompt less effective. this interpretation is in line with the finding that error-explaining prompts are not effective when teachers’ cognitive load is too high (heitzmann, fischer, & fischer, 2018) or for teachers with low professional knowledge (chernikova et al., 2019). in both cases, teachers need more instructional guidance, which would be especially relevant for pre-service teachers such as the ones who participated in our study. however, more studies are needed to investigate what makes scales too complex to use controlled information processing, and what knowledge teachers might require to reduce the complexity of judgments. 5. general discussion the two studies aimed to examine two sequential research questions: do halo-effects of spelling errors distort pre-service teachers’ ratings of students’ texts? can a prompt reduce this halo-effect? regarding the first research question, the results in study 1 and 2 showed halo-effects of spelling errors distorting pre-service teachers’ ratings of students’ texts. this finding is in line with our hypothesis and the research on the influence of spelling errors on the assessment of students’ texts (marshall, 1967; rezaei & lovorn, 2010; scannell & marshall, 1966). further, our studies presented three novelties that complement existing research: first, the sample consisted of pre-service teachers rather than of professional raters. halo-effects in professional rating settings are a well-known phenomenon, and various interventions are used to safeguard against their distorting consequences (see myford & wolfe, 2009). in contrast, no safeguards exist for pre-service teachers, and hence, this judgment error could persist until teachers start working in a school. based on our findings, the strong need to examine possibilities of reducing pre-service teachers’ halo-effects becomes evident. second, in comparison to previous research (marshall, 1967; rafoth & rubin, 1984; rezaei & lovorn, 2010; scannell & marshall, 1966), teachers assessed texts with the same overall quality but different amounts of spelling errors in our studies. during the text assessment, the participants saw examples of high-quality texts with few and many spelling errors and thus were able to experience these as two related but ultimately distinct categories of quality. nevertheless, the halo-effect occurred and extended to textual aspects clearly unrelated to spelling, such as the framing of an essay or the internal structure of paragraphs. seeing that even high-quality texts can contain spelling errors does not seem to protect pre-service teachers against halo-effects. the third novelty is that the texts analysed in our study were written by upper-secondary school students and not by university students, which indicates that the halo-effects of spelling errors are a more widespread problem than previous research has suggested (barkaoui, 2010; freedman, 1979; rafoth & rubin, 1984; rezaei & lovorn, 2010; sweedler-brown, 1993). following these three novelties, the conclusion for our first research question is the following: halo-effects are a problem in upper-secondary esl writing education and there is a need for an intervention to reduce it. we addressed this need in our second study. the results of the second study showed that a prompt could reduce the halo-effect of spelling. this result is in line with our hypotheses, the assumptions of the continuum model of impression formation processes (fiske & neuberg, 1990) and the research of fostering diagnostic competencies of pre-service teachers. we have chosen the prompt because of the assumption, derived from the continuum model, that drawing attention to target attributes of the assessment can reduce halo-effects. we interpret the effect of the prompt as support for this assumption and the usefulness of the continuum model to describe judgment processes. the prompt’s effect size was similar to the mean prompt’s effect size in a meta-analysis on fostering teachers’ diagnostic competencies (chernikova et al., 2019). however, in studies summarized in meta-analyses, the prompts were usually part of larger interventions. the results of our study contradict those of the only other study investigating prompting effects solely on facilitating teachers’ diagnostic competencies (heitzmann et al., 2018), which showed that prompts alone tend to hamper the quality of teachers’ diagnoses and are only effective in combination with adaptive feedback. heitzmann et al. (2018) explained the hampering effects with the high cognitive load teachers experienced when processing the prompt. hence, it could be more efficient and economical to use prompts that elicit little cognitive load, like in the actual study. the research in the relatively new field of fostering pre-service teachers’ diagnostic competencies opens a promising avenue for further research. it contains a simple but innovative suggestion for fostering the quality of teacher assessments of written students’ performances, combining interdisciplinary research from second language learning research and psychology, and showing beneficial prompting effects without additional adaptive feedback (heitzmann et al., 2018). ’the particular strength of our study is that it showed a way of fostering the assessment quality that could be integrated into teacher education economically. further, our results can help teachers in a way that encourages them to ask themselves: did i misjudge my students’ who have spelling difficulties, and did i overlook their strengths because of their spelling errors? we concluded for both research questions that domain-specific judgment characteristics, like analytic writing assessments, could trigger judgment errors, which can and should be detected and reduced to raise second language instruction quality. a limitation of our study is that we did not randomize the order of holistic and analytic scoring in the studies. we used the scoring order recommended by singer and lemahieu (2011) and did not want to include a scoring order leading to more distracting interferences among the different scales. we included both a holistic rating scale and an analytic rating scale since both scales are commonly employed in practice (weigle, 2002), and our study aimed to investigate halo effects in relation to achievement-relevant and achievement-irrelevant characteristics. the experimental setting used in the student inventory asset assures internal validity due to its strict variable control and randomized allocation of text qualities and spelling errors. however, one should keep in mind that real-life assessment situations differ considerably from this simplified experimental research design (keller, 2013). hence, caution should be exercised when applying these insights into real-life situations. another limitation is the text selection: four texts on two quality levels and with two independent degrees of spelling errors were assessed in the studies. while this selection made it possible to reach a high internal validity of our study, it makes it difficult to generalize the results, and it remains unclear whether prompts are also effective in school settings. in school, teachers have a larger amount of texts available that cover a wide continuum of qualities and in which the quality levels and the number of spelling errors correlate with one another (bae, bentler, & lee, 2016; lai, wolfe, & vickers, 2015). furthermore, the sample in our studies consisted of pre-service esl teachers only. the applicability of the results to experienced teachers is limited since experienced teachers may discriminate between assessment scales more accurately due to their greater knowledge and experience, and hence may be less influenced by halo effects. however, to this date, no study demonstrated the difference between pre-service and experienced teachers’ judgments in assessing student texts (meadows & billington, 2010; royal‐dawson & baird, 2009). we decided to conduct our study with pre-service teachers to maximize the possible outreach of the prompt: if the prompt reduced halo-effects in teacher’ education, the effects could be beneficial throughout teachers’ professional career. to generalize our findings, we encourage other researchers to conduct similar experiments with experienced teachers. for the scientific discussion of performance assessment, previous research and the outlined studies indicate that the influence of spelling errors should also be investigated in other subjects, such as science education. further studies could systematically investigate other forms of support that help teachers reduce halo effects in the assessment. in practice, results indicate that teachers judge the quality of student texts and thus the competencies of students in certain areas more erroneously when texts contain many spelling errors. these judgment errors can both systematically disadvantage students and interfere with competence-adjusted lesson planning. thus, it is important to inform pre-service and experienced teachers of typical difficulties in writing assessment and support them in assessing students’ written performances objectively. hence, assessment scales and training programs such as those offered by the student inventory asset should be further improved and used in empirical studies of the complex psychological processes underlying text assessment. further, they should be made available to both pre-service and experienced teachers as training instruments. keypoints examining teachers’ holistic and analytic writing assessment english as a second-language (esl) essays from upper-secondary school spelling errors triggered halo effects on analytic rating scales a prompt could reduce these halo effects the first study shows the judgment error; the second reduces it funding this work was supported by the swiss national science foundation [165483; 100019l162675] and the german research foundation (dfg) [mo648/25-1] references bae, j., & bachman, l. f. (2010). an investigation of four writing traits and two tasks across two languages. language testing, 27(2), 213–234. https://doi.org/10.1177/0265532209349470 bae, j., bentler, p. m., & lee, y. s. (2016). on the role of content in writing assessment. language assessment quarterly , 13(4), 302–328. https://doi.org/10.1080/15434303.2016.1246552 bakker, m., van dijk, a., & wicherts, j. m. (2012). the rules of the game called psychological science. perspectives on psychological science , 7(6), 543–554. https://doi.org/10.1177/1745691612459060 barkaoui, k. (2010). explaining esl essay holistic scores: a multilevel modeling approach. language testing , 27(4), 515–535. https://doi.org/10.1177/0265532210368717 birkel, p., & birkel, c. (2002). wie einig sind sich lehrer bei der aufsatzbeurteilung? eine replikationsstudie zur untersuchung von rudolf weiss [how concordant are teachers’ essay scorings? a replication of rudolf weiss’ sudies]. psychologie in erziehung und unterricht, 49(3), 219–224. brookhart, s. m. (2011). educational assessment knowledge and skills for teachers. educational measurement: issues and practice , 30(1), 3–12. https://doi.org/10.1111/j.1745-3992.2010.00195.x brookhart, s. m. (2013). the use of teacher judgement for summative assessment in the usa. assessment in education: principles, policy & practice , 20(1), 69–90. https://doi.org/10.1080/0969594x.2012.703170 chamberlain, s., & taylor, r. (2011). online or face‐to‐face? an experimental study of examiner training. british journal of educational technology , 42(4), 665–675. https://doi.org/10.1111/j.1467-8535.2010.01062.x chernikova, o., heitzmann, n., fink, m.c. et al. (2019). facilitating diagnostic competencies in higher education—a meta-analysis in medical and teacher education. educ psychol rev, 32, 157–196. https://doi.org/10.1007/s10648-019-09492-2 cooksey, r. w., freebody, p., & wyatt-smith, c. (2007). assessment as judgment-in-context: analysing how teachers evaluate students’ writing 1. educational research and evaluation, 13(5), 401–434. https://doi.org/10.1080/13803610701728311 culham, r. (2003). 6+ 1 traits of writing: the complete guide . new york: scholastic inc. cumming, a., kantor, r., & powers, d. e. (2002). decision making while rating esl/efl writing tasks: a descriptive framework. the modern language journal , 86(1), 67–96. https://doi.org/10.1111/1540-4781.00137 dempsey, m. s., pytlikzillig, l. m., & bruning, r. h. (2009). helping preservice teachers learn to assess writing: practice and feedback in a web-based environment. assessing writing , 14(1), 38–61. https://doi.org/10.1016/j.asw.2008.12.003 driscoll, d. p., avallone, a. p., orr, c. s., & crovo, m. (2010). writing framework for the 2011 national assessment of educational progress . washington, dc: national assessment governing board, us dept. of education. elliott, j., lee, s. w., & tollefson, n. (2001). a reliability and validity study of the dynamic indicators of basic early literacy skills-modified. school psychology review , 30(1), 33–49. european commission (2008). multilingualism an asset for europe and a shared commitment. retrieved from http://eur-lex.europa.eu/legal-content/en/txt/?uri=uriserv:ef0003 fiske, s. t., & neuberg, s. l. (1990). a continuum of impression formation, from category-based to individuating processes: influences of information and motivation on attention and interpretation. in m. p. zanna (ed.), advances in experimental social psychology (vol. 23, pp. 1–74). new york: ny: academic press. flor, m., futagi, y., lopez, m., & mulholland, m. (2015). patterns of misspellings in l2 and l1 english: a view from the ets spelling corpus. bergen language and linguistics studies , 6, 107–132. https://doi.org/10.15845/bells.v6i0.811 flower, l., & hayes, j. r. (1981). a cognitive process theory of writing. college composition and communication , 32(4), 365–387. https://doi.org/10.2307/356600 freedman, s. w. (1979). how characteristics of student essays influence teachers’ evaluations. journal of educational psychology , 71(3), 328–338. https://doi.org/10.1037/0022-0663.71.3.328 hamp-lyons, l. (1991). assessing second language writing in academic contexts . chestnut st., norwood: ablex publishing corporation. heitzmann, n., fischer, f., & fischer, m. r. (2018). worked examples with errors: when self-explanation prompts hinder learning of teachers diagnostic competences on problem-based learning. instructional science, 46(2), 245–271. https://doi.org/10.1007/s11251-017-9432-2. herppich, s., praetorius, a. k., förster, n., karst, k., leutner, d., behrmann, l., . . . südkamp, a. (2017). teachers’ assessment competence: integrating knowledge-, process-, and product-oriented approaches into a competence-oriented conceptual model. teaching and teacher education. (76), 1–13. https://doi.org/10.1016/j.tate.2017.12.001 huot, b. (1996). toward a new theory of writing assessment. college composition and communication, 47(4), 549-566. https://doi.org/10.2307/358601 hyland, k. (2008). second language writing. new york: cambridge university press. https://doi.org/10.1017/s0261444808005235 jansen, t., vögelin, c., machts, n., keller, s., & möller, j. (2019). das schülerinventar asset zur beurteilung von schülerarbeiten im fach englisch: drei experimentelle studien zu effekten der textqualität und der schülernamen [the student inventory asset for judging students performances in the subject english: three experimental studies on effect of text quality and student names]. psychologie in erziehung und unterricht , 66(4), 303–315. https://doi.org/10.2378/peu2019.art21d jansen, t., vögelin, c., machts, n., keller, s., köller, o., & möller, j. (2021). judgment accuracy in experienced versus student teachers: assessing essays in english as a foreign language. teaching and teacher education, 97, 103216. https://doi.org/10.1016/j.tate.2020.103216 kaiser, j., möller, j., helm, f., & kunter, m. (2015). das schülerinventar: welche schülermerkmale die leistungsurteile von lehrkräften beeinflussen [the student inventory: how student characteristics bias teacher judgments]. zeitschrift für erziehungswissenschaft, 18(2), 279–302. https://doi.org/10.1007/s11618-015-0619-5 kaufmann, e. (2020). how accurately do teachers judge students? re-analysis of hoge and coladarci (1989) meta-analysis. contemporary educational psychology, 63, 101902. https://doi.org/10.1016/j.cedpsych.2020.101902 keller, s. (2013). integrative schreibdidaktik englisch für die sekundarstufe: theorie, prozessgestaltung, empirie . tübingen: gunter narr verlag. lai, e. r., wolfe, e. w., & vickers, d. (2015). differentiation of illusory and true halo in writing scores. educational and psychological measurement, 75(1), 102–125. https://doi.org/10.1177/0013164414530990 lovorn, m. g., & rezaei, a. r. (2011). assessing the assessment: rubrics training for pre-service and new in-service teachers. practical assessment, research & evaluation , 16(16), 1–18. marshall, j. c. (1967). composition errors and essay examination grades re-examined. american educational research journal , 4(4), 375–385. meadows, m., & billington, l. (2010). the effect of marker background and training on the quality of marking in gcse english . manchester: aqa centre for education research and policy. murphy, k. r., & reynolds, d. h. (1988). does true halo affect observed halo? journal of applied psychology , 73(2), 235–238. https://doi.org/10.1037/0021-9010.73.2.235 open science collaboration (2015). estimating the reproducibility of psychological science. science , 349(6251). https://doi.org/10.1126/science.aac4716 parr, j. m., & timperley, h. s. (2010). feedback to writing, assessment for teaching and learning and student progress. assessing writing , 15(2), 68–85. https://doi.org/10.1016/j.asw.2010.05.004 rafoth, b. a., & rubin, d. l. (1984). the impact of content and mechanics on judgments of writing quality. written communication , 1(4), 446–458. https://doi.org/10.1177/0741088384001004004 rezaei, a. r., & lovorn, m. (2010). reliability and validity of rubrics for assessment through writing. assessing writing , 15(1), 18–39. https://doi.org/10.1016/j.asw.2010.01.003 royal‐dawson, l., & baird, j. a. (2009). is teaching experience necessary for reliable scoring of extended english questions? educational measurement: issues and practice , 28(2), 2–8. https://doi.org/10.1111/j.1745-3992.2009.00142.x ruiz‐primo, m. a., & furtak, e. m. (2007). exploring teachers’ informal formative assessment practices and students’ understanding in the context of scientific inquiry. journal of research in science teaching , 44(1), 57–84. https://doi.org/10.1002/tea.20163 saal, f. e., downey, r. g., & lahey, m. a. (1980). rating the ratings: assessing the psychometric quality of rating data. psychological bulletin , 88(2), 413–428. https://doi.org/10.1037/0033-2909.88.2.413 sadler, d. r. (2009). indeterminacy in the use of preset criteria for assessment and grading. assessment & evaluation in higher education, 34 (2), 159-179. https://doi.org/10.1080/02602930801956059 scannell, d. p., & marshall, j. c. (1966). the effect of selected composition errors on grades assigned to essay examinations. american educational research journal, 3(2), 125–130. südkamp, a., kaiser, j., & möller, j. (2012). accuracy of teachers’ judgments of students’ academic achievement: a meta-analysis. journal of educational psychology , 104(3), 743–762. https://doi.org/10.1037/a0027627 sweedler-brown, c. o. (1993). esl essay evaluation: the influence of sentence-level and rhetorical features. journal of second language writing , 2(1), 3–17. https://doi.org/10.1016/1060-3743(93)90003-l urhahne, d., & wijnia, l. (2021). a review on the accuracy of teacher judgments. educational research review , 32, 100374. https://doi.org/10.1016/j.edurev.2020.100374 vögelin, c., jansen, t., keller, s., machts, n., & möller, j. (2019). the influence of lexical features on teacher judgments of esl argumentative essays. assessing writing , 39, 50–63. https://doi.org/10.1016/j.asw.2018.12.003 vögelin, c., jansen, t., keller, s., & möller, j. (2018). the impact of vocabulary and spelling on judgments of esl essays: an analysis of teacher comments. the language learning journal. advance online publication. https://doi.org/10.1080/09571736.2018.1522662 weigle, s. c. (2002). assessing writing.: cambridge language assessment series . cambridge: cup. weir, c. (1988). the specification, realization and validation of an english language proficiency test. in hughes a. (ed.), testing english for university study. elt documents 127 (pp. 45–110). london: modern english publications in association with the british council. wind, s. a., stager, c., & patil, y. j. (2017). exploring the relationship between textual characteristics and rating quality in rater-mediated writing assessments: an illustration with l1 and l2 writing assessments. assessing writing, 34, 1–15. https://doi.org/10.1016/j.asw.2017.08.003 wolfe, e. w., song, t., & jiao, h. (2016). features of difficult-to-score essays. assessing writing , 27, 1–10. https://doi.org/10.1016/j.asw.2015.06.002 zimmermann, f., möller, j., & köller, o. (2018). when students doubt their teachers’ diagnostic competence: moderation in the internal/external frame of reference model. journal of educational psychology , 110(1), 46–57. https://doi.org/10.1037/edu0000196 . appendix: analytic scales for argumentative essays frame of essay: introduction and conclusion 4 effective introduction with “hook” and “thesis statement”; effective conclusion summarizing main arguments 3 mostly effective introduction with either “hook” or “thesis statement”; mostly effective conclusion summarizing main arguments 2 introduction and/or conclusion identifiable but only partly effective 1 both introduction and conclusion not clearly identifiable or mostly ineffective body of essay: internal organization of paragraphs 4 paragraphs are well-organized and coherent throughout 3 paragraphs are mostly well-organized and coherent 2 paragraphs are partly well-organized and coherent 1 paragraphs are not well-organized and incoherent support of arguments 4 author uses a variety of different examples to support her/his argument and fully explains their relevance to the topic 3 author uses different examples to support her/his argument and mostly explains their relevance to the topic 2 author uses a few examples to support her/his argument and partly explains their relevance to the topic 1 author uses repetitive examples to support her/his argument and their relevance to the topic is mostly unclear spelling and punctuation 4 author uses mostly correct spelling and punctuation 3 author uses mostly correct spelling and punctuation, with few distracting errors 2 author uses partly correct spelling and punctuation, with some distracting errors 1 author uses partly correct spelling and punctuation, with many distracting errors grammar 4 author uses a variety of complex grammatical structures, few grammar mistakes 3 author uses some complex grammatical structures, grammar mostly correct 2 author uses few complex grammatical structures, grammar partly correct 1 author uses few or no complex grammatical structures, grammar mostly incorrect vocabulary 4 author uses sophisticated, varied vocabulary throughout 3 author mostly uses sophisticated, varied vocabulary 2 author partly uses sophisticated, varied vocabulary, sometimes repetitive 1 author uses little sophisticated, varied vocabulary, often repetitive overall task completion 4 text fully conforms to the conventions of an argumentative essay, thus fully completing the task 3 text mostly conforms to the conventions of an argumentative essay, thus mostly completing the task 2 text partly conforms to the conventions of an argumentative essay, thus partly completing the task 1 text does not conform to the conventions of an argumentative essay, thus not completing the task codepen heinimaki et al publication frontline learning research vol.8 no. 2 (2020) 65 89 issn 2295-3159 core and activity-specific functional participatory roles in collaborative science learning olli-pekka heinimäkia, simone voletb, marja vaurasa adepartment of teacher education, university of turku, finland bschool of education, murdoch university, australia article received 26 march 2019/ revised 6 september / accepted 9 march 2020/ available online 21 april abstract prior research on the significance of roles in collaborative learning has explored their impact when they are pre-assigned to group members. in this article, it is argued that focusing on assigned roles downplays the spontaneous, emergent, and interactional nature of roles in small task groups and that this focus has limited the development of generalizable frameworks aimed at understanding the impact of roles in and across collaborative learning settings. a case is built for the importance of focusing on the functional participatory roles enacted during collaborative learning and for conceptualising these roles as emergent, dynamic, and evolving in situ (first claim). further, a flexible conceptual framework for the analysis and understanding of such roles across diverse collaborative science-learning activities is proposed, based on the assumption that during collaborative learning, both core and activity-specific roles are enacted (second claim). the core roles resemble each other across activities as they associate closely with the nature of the science discipline itself, whereas the activity-specific roles vary across activities as their emergence is dependent on the affordances, demands, and characteristics of the particular activity and environment. data from three diverse science-learning environments, including four totally or partly student-led collaborative science activities, were scrutinized to establish the degree of empirical support for this assumption and, thereby, the conceptual usefulness of the proposed framework. the contributions of the framework for future research of collaborative science learning are discussed. keywords: roles; collaborative learning; process-data analysis; situative approach; science learning info corresponding author email: olpehe@utu.fi doi: 10.14786/flr.v8i2.469 1. introduction totally or partly student-led (referred to simply as student-led hereafter) collaborative learning activities are increasingly in use across educational levels and disciplines. in science classrooms, such activities aim to stimulate students’ deep learning and engagement through co-construction of science knowledge with peers (ucan & webb, 2015; webb, 2008). student-led collaborative learning activities are often characterized by their informal and open-ended nature (volet, summers, & thurman, 2009), and recent research has emphasized the emergent, interactive and dynamic nature of group processes as they unfold during such activities (e.g., hadwin, järvelä, & miller, 2018; hilpert & marchand, 2018). the inherently dynamic nature of collaborative learning forms an integral part of the situative perspective (greeno, 1998, 2006), which is concerned by the complex activity systems constituted by individual cognitive agents interacting with each other and with their environment (including tools, technological artefacts, tasks, communities, etc.). studies grounded in this perspective have explored how individuals and context operate jointly to produce outcomes in situ (turner & nolen, 2015). as students interact with each other, the task and the learning context, they constantly enact roles that are expected to have a major impact on the success of the collaboration. the importance of roles in collaborative groups has long been recognised, as educational research examining the benefits of pre-assigning desirable roles to group members goes back several decades (e.g., johnson & johnson, 1989; slavin, 1996). this line of research has re-gained momentum in recent computer-supported collaborative learning (cscl) research (e.g., strijbos, martens, jochems, & broers, 2004; gu, shao, guo, & lim, 2015). however, although roles are recognised as fundamental in group dynamics and for effective group work (forsyth, 2014), prior research in educational settings has tended to concentrate on the roles that are pre-assigned and scripted for students. as a consequence, an understanding of the roles that emerge spontaneously and are enacted dynamically in situ during student-led collaborative activities is still lacking. in light of the increased importance of collaborative groups in educational settings and teams in the workplace, this gap needs to be addressed. such research requires methodologies that capture and analyse “learners-in-the-context” (nolen, horn, & ward, 2015, p. 237). the insight gained from the fragmented line of research on roles, conceptualised as emergent and dynamic (e.g., volet, vauras, salo, & khosa, 2017; lehmann-willenbrock, beck, & kauffeld, 2016; sarmiento & shumar, 2010), has enhanced our understanding of real-life collaborative learning but this is not without methodological challenges. one challenge originates from the fact that empirical studies on the dynamics of collaborative learning groups (not unlike other research fields) have involved the constant development of new coding systems or the revision of existing ones, aimed at capturing the specific behaviours and interactions assumed to be influenced by the characteristics of the particular situation being studied (volet & summers, 2013). even though a data-driven approach is much needed to complement a theory-driven one, this raises challenges for comparing findings obtained across multiple settings, activities and groups. as a consequence, the development of more generalizable coding systems has been complicated (nolen et al., 2015; turner & nolen, 2015; volet & summers, 2013). as argued by volet and summers, there is a pressing need to develop coding systems that are both sensitive enough to capture the specific characteristics of the data in question and general enough to address scaling issues. to address the conceptual gaps in prior research on roles (see strijbos & laat, 2010), a flexible conceptual framework for understanding and analysing emergent task-related functional participatory roles in and across collaborative science-learning activities is tentatively proposed. this framework acknowledges that functional participatory roles are emergent, spontaneously enacted during collaborative learning activities as well as dynamic and evolving in situ (first claim). it also assumes that both core and activity-specific roles are enacted by students during collaborative science learning (second claim). to provide some validation for the proposed framework and address the methodological challenge related to generalisation of the findings, data from three distinct science-learning environments are presented and scrutinized to explore the extent of empirical support for the assumption of core and activity-specific functional participatory roles enacted during collaborative science learning. to build the conceptual case for the two claims, the next section reviews in turn the variety of role typologies and frameworks in the literature, and the range of empirical studies on roles in collaborative learning research. 1.1 variety of role typologies and frameworks it is easy to agree with moxnes (1999) that “[i]n research of small groups … an extraordinary range of roles have been suggested” (p. 110) as one can find a large repertoire of role typologies and frameworks in the extant small-group literature. however, no generally endorsed typology or framework exists (stewart, fulmer, & barrick, 2005). therefore, there is still a call for the establishment of role frameworks, typologies, and coding systems to enhance our understanding and facilitate the analysis of roles in different small-group contexts, such as collaborative learning groups. the majority of prior work in the field has been conducted in work-related contexts and with different type of small task groups. pioneering work, from which many studies have drawn since across disciplines, consists of benne and sheats’ (1948, reprinted 2007) typology of functional roles of group members, which identifies as many as 27 distinct roles (e.g., information giver, procedural technician, follower), and the work of bales (1951; see also bales & slater, 1955), who classified group members roles into two distinct foci (i.e., task-specialist and socio-emotional specialist). one subsequent and well-known typology is the classification of team roles (e.g., team worker, coordinator, implementer) by belbin (e.g., 1993), who argued that high-performing teams need a balanced mix of different roles to be represented in the team. later, with a typology of productive roles (e.g., supporter, proposer, recorder), chiu (2000) linked roles with three strategic behavioural dimensions that an individual can apply in social interaction: i) evaluation of previous action (supportive, critical, unresponsive), ii) knowledge content (contribution, repetition, null), and iii) invitational form (command, question, statement). chiu’s framework has provided helpful grounding for empirical investigations into how individual actions, and the quality of such actions, interact dynamically with the actions of other group members when the group is working toward a joint goal (volet et al., 2017). to sum up, research to date has identified two broader types of roles in task groups: task and socio-emotional roles (driskell, driskell, burke, & salas, 2017; forsyth, 2014). the framework proposed in this article focuses on task-related roles (i.e., roles relating to the work of the group toward task achievement), notwithstanding that socio-emotional roles are expected to influence the enactment of task-related roles and, in turn, the quality of task achievement. because of this extant multiplicity of typologies and frameworks, driskell et al. (2017) aimed recently to trace core team roles by integrating existing role typologies of work teams. this endeavour was based on the findings of another study (see gregory, shimone, burke, & salas, 2015) in which, according to driskell et al. (2017), 23 unique team-role typologies were discovered in a pool of 139 research papers via a comprehensive literature review, yielding the identification of 164 distinct roles. even though some of the variations turned out to be terminology-based (similar roles had been labelled differently), roles unique to certain team-role contexts were also identified. as the bulk of these reviewed typologies covered roles in work-related contexts (driskell et al., 2017), no straightforward generalizations can be drawn to collaborative learning contexts. however, these findings offer hypothetical grounding to assume that there may be core roles underpinning collaborative science learning, but also roles that are more specific for certain science learning activities. 1.2 roles in collaborative learning research while the main aim of research on roles in work-related contexts has been to identify roles that are crucial for organizations and teams, and to explore individual preferences and abilities to take over those roles (e.g., belbin, 1993), collaborative learning research has taken a different approach on the study of roles by focusing mainly on pre-assigned and scripted roles. pre-determined roles for learners have long been suggested as a valuable way to improve the quality of collaboration and group performance because they foster productive individual inputs and more equal participation in collaborative task completion (e.g., cohen, 1994; johnson & johnson, 1989; slavin, 1996). more recently, the value of pre-determining roles for learners has gained significant momentum in the field of cscl, with evidence that it can promote successful collaborative learning in challenging environments, such as when group members interact via technology (e.g., cesareni, cacciamani & fujita, 2016; cheng, wang, & mercer, 2014; de wever, van keer, schellens, & valcke, 2009, 2010; gu et al., , 2015; morris et al., 2010; pozzi, 2011; schellens, van keer, de wever, & valcke, 2007; strijbos et al., 2004). for instance, gu et al. (2015) designed a role structure for university students including six assigned roles (starter, supporter, arguer, questioner, challenger, timer), accompanied with prompts how to play each role, to promote meaningful engagement as well as active participation during problem solving in a cscl-environment. similarly, morris et al. (2010) have used scripted roles such as predictor, summarizer, questioner, and clarifier to scaffold students collaboration in the cscl-environment gstudy. other researchers have investigated the effects of role assigning to collaborative learning outcomes. for example, strijbos et al. (2004) investigated the effects of pre-scribed roles on collaboration and group performance, compared to groups for which roles were not pre-assigned (i.e., groups had to self-organise and coordinate their collaborative activities). in that study, groups of university students had to undertake a group project, where group communication was carried out via e-mail. for the groups in the assigned roles condition, four roles with specific tasks were developed and assigned: project planner, communicator, editor, and data collector. the findings indicated that assigned roles had no effect on group performance in terms of group-level grade, but elicited more task-related statements and increased students’ awareness about collaboration. the positive effect of assigning roles has been also documented in research on knowledge construction; for example, pre-scribed roles of starter, summarizer, moderator, theoretician and source searcher in asynchronous discussion boards were found positively linked to levels of social knowledge construction in a study of de wever et al. (2009; see also 2010; schellens et al., 2007). while the importance of studying roles in authentic settings has been highlighted (morris et al., 2010), pre-determining roles have posed some thorny issues in research on both cscl and face-to-face learning, and this regardless of their benefits and positive effects. first, as assigned roles commonly involve only a given set of roles that have already been considered meaningful for productive collaboration prior to the activity, roles that may emerge spontaneously and naturally during the activity remain unacknowledged. these may include productive roles but also a whole range of roles stretching all the way to the possible impact of more counterproductive roles (hogan, 1999; lehmann-willenbrock et al., 2016). second, it can be argued that allowing a student to play one single role at a time during an activity is too static and not adequate to characterize individuals’ behaviour in dynamic contexts where roles are naturally in a state of constant flux among group members (salazar, 1996; sarmiento & shumar, 2010). third, the fact that assigned roles are typically designed and operationalized with a particular activity and group of students in mind has restrained the establishment of more scalable analytical frameworks in the field (strijbos & laat, 2010). such arguments stress the importance of shifting the focus of research on roles in collaborative learning from fixed roles to roles that spontaneously emerge and evolve, and of developing frameworks that are suitable to understand and analyse such roles across collaborative learning settings. although there has been overall much less research on type of roles that emerge spontaneously and are self-adopted by the students (i.e., emergent role approach, see strijbos & weinberger, 2010), this research has yielded some empirical support for the arguments and shortcomings detailed above. a few of these studies has been conducted in collaborative science learning context. for instance, hogan’s study (1999) investigated group discussions of secondary school students, whose task was to engage to co-construction of knowledge to make sense and explain science phenomena they had observed in a laboratory setting. through video-observations, she discerned eight roles enacted spontaneously by the students in this group activity: four of the roles were found to promote co-construction (promoter of reflection, contributor of content knowledge, modeller, mediator), and four, also spontaneously enacted were found to impede co-construction (promoter of acrimony, distractor, promoter of simple task completion or unreflective acceptance of ideas, reticent). in that study, however, roles were conceptualized as consistent patterns of behaviours over time, which did not allow a fine-grained investigation of how roles fluctuate dynamically in situ among the group members (see also maloney, 2007). two recent studies (volet et al., 2017; volet, jones, & vauras, 2019) demonstrated the benefits of adopting such a fine-grained analytical approach. for example, the study of volet et al. (2017) revealed that students were able to flexibly enact wide range of roles during the process of collaborative concept mapping. moreover, in the groups that produced higher-quality concept maps, students displayed greater flexibility in the enactment of roles featuring deeper cognitive processing of the science content (e.g., knowledge provider, challenger) compared to the groups of students that produced concept maps of a lower quality. one major limitation of empirical studies on emergent roles, however, is that they rarely extend beyond a single learning setting, which limits the generalisability of the conceptual framework underpinning the study and its findings. a few attempts, however, have been to develop more cohesive and scalable frameworks of emergent roles in certain specific collaborative learning contexts. one of the most comprehensive efforts to date is by strijbos and laat (2010), who, after reviewing prior cscl role literature, proposed a conceptual framework of roles at three levels: i) micro (role as a task), ii) meso (role as a pattern), and iii) macro (role as a stance). furthermore, they used two asynchronous cscl datasets to illustrate the viability of the conceptual framework on the macro level, that is, “[a]n individual’s participative pattern based on their attitude towards the task and collaborative learning” (p. 497). the participative stances consist of four roles at the small-group level (e.g., ghost, over-rider) and the correspondence of those roles to larger groups (cf. e.g., lurker, generator). as the framework of strijbos and laat (2010) was developed in the context of cscl, its adoptability in research into face-to-face learning would need to be established. in addition, the approach of participative stances may not be sufficient enough if the aim is to focus on roles that are expected to fluctuate and be enacted dynamically in situ during the activity (cf. micro-level). notwithstanding, scrutinization of several datasets (strijbos & laat, 2010) seems logical in order to identify possible commonalities and discrepancies also between such roles and in different types of learning context. thereby, a cross-dataset approach was adopted in the present study, to explore the emerging nature of core and activity-specific functional participatory roles in collaborative science learning. 2. understanding emerging functional participatory roles in collaborative science learning following benne and sheats (1948, 2007), forsyth (2014), and oliveira, boz, broadwell, and sadler (2014), it is posited that task-roles naturally emerging in situ are largely functional, in the sense that they emerge due to group members’ attempts to address and fulfil task-related demands of the ongoing activity. moreover, and consistent with marcos-garcía, monés, and dimitriadis (2015) and volet et al., (2017), when roles emerge, they manifest themselves through individual participation in group interaction, thereby become observable. thus, the aim in the present study was to contribute to research on the emergence of roles by proposing a conceptualization of functional participatory roles, which can be defined as the specific strategies and behaviours used by an individual in a particular situation (cf. volet et al., 2017). this stresses the spontaneous, dynamic, and interactive nature of roles in small task groups (lehmann-willenbrock et al., 2016; salazar, 1996), and is consistent with a categorization of roles as understood at the micro-level (see strijbos & laat, 2010). in sum, the main aim of this article is to empirically scrutinize the usefulness of the proposed conceptual framework of spontaneously enacted core and activity-specific functional participatory roles in collaborative science learning. a key assumption underlying the framework is that some task-related roles resemble each other across collaborative learning activities. these are thus called core roles, whereas other roles that display more specific characteristics and vary across science activities are called activity-specific roles. consequently, consistent with their common function across collaborative science-learning tasks (i.e., learn and understand the science content), the core roles are assumed to be more intimately linked to the nature of the science discipline itself, while activity-specific roles are expected to be dependent on affordances and characteristics of the learning activity, environment, and task demand. the proposed framework is therefore in line with some researchers’ claims that there are both commonalities and discrepancies in roles emerging in diverse small-group settings (e.g., driskell et al., 2017; strijbos & laat, 2010). data from three diverse learning environments (two involving one and the third involving two learning activities) were scrutinized to examine the extent of empirical support for the assumption of emerging, functional, core and activity-specific participatory roles in student-led collaborative science learning. 2.1 a search for empirical support the datasets scrutinized for empirical support for this assumption come from three recent studies. each of these studies aimed to explore the conceptual usefulness of the construct of roles in order to understand individual contributions in productive, student-led collaborative science learning. the first study addressed the scarce evidence of emergent roles in collaborative science learning (volet et al, 2017). an original coding system was developed to analyse naturally emerging roles in collaborative concept mapping. a second study (heinimäki, salo, & vauras, 2019) extended the conceptual grounding by adapting the original coding system to the study of a virtual collaborative science-learning environment. finally, a third study adapted the original coding system to analyse roles in hands-on collaborative science-learning activities (volet et al., 2019). in order to capture the expected constant fluctuation of roles from ongoing group interaction, all three studies employed video-based role analyses conducted at the turn level, meaning that the identification of roles was based on discrete verbal and non-verbal individual contributions (i.e., utterance, nodding one’s head, etc.). inter-rater reliability between two trained independent coders (targeting to at least 20% of all the analysed turns of each activity) yielded a “substantial” to an “almost perfect” result (see landis & koch, 1977, p. 165) across datasets. in the next section, each dataset is presented in turn, introduced by a brief summary of the aim, design, methodology and key findings of the study that generated the data and a description of how the coding system used to analyse that dataset was developed or adapted. in addition, a chosen data excerpt from each dataset is provided to illustrate sequences of group interactions, in which emerging functional participatory roles were enacted. this is followed by a synthesis of the outcomes from the three datasets and a comprehensive overview of the indicators of core and activity-specific roles identified in each dataset, with illustrative examples of each role. all the names mentioned in the excerpts are pseudonyms. 2.1.1 dataset 1: veterinary science students co-constructing a concept map of a real-life clinical case the first dataset is derived from a study in which the aim was to use two analytical approaches to understand individual contributions to productive collaborative learning from clinical case-based assignments (volet et al., 2017). how patterns of self-adopted roles could explain qualitative differences between groups in task performance was explored. the research site was a mandatory physiology unit, specifically a case-based group assignment designed to provide second-year veterinary science students with early exposure to and opportunities to learn from a randomly assigned real-life clinical case. small peer groups of five to six students were required to investigate and learn from their case in their own time over a six to seven week period (see khosa, volet, & bolton, 2010, 2014; khosa & volet, 2013, 2014; vauras, volet, & nolen, 2019; volet et al., 2017). towards the end of that period, the groups were invited to construct together and without time limit, a meaningful conceptual map of their respective clinical case. a set of cards featuring their case was provided as well as a pen to draw either unidirectional arrows (representing cause and effect relationships) or bidirectional arrows (representing inter-related relationships) between concepts. this totally student-led activity was video-taped, and the video-footage used to identify and analyse the enacted roles emerging naturally during the activity. the quality of the groups’ concept maps was assessed through comparison with the maps of clinical experts. based on this evaluation, two groups with high and two groups with low quality conceptual maps were chosen for the role analysis. this sampling made it possible to explore how the different self-adopted roles may have contributed to explain the differences in quality of group performance, and it provided some empirical support for the functionality of the role-coding system across differently performing groups. the main findings were that roles reflecting deep cognitive processing of the science content, in contrast to just sharing opinions without reference to science, and flexible enactment of roles with such a cognitive focus (see descriptions below), were more common in the groups that produced concept maps of a higher quality. development of the role-coding system for dataset 1. the original coding system was developed primarily based on the typologies of functional roles by benne and sheats (1948, 2007) and of productive roles by chiu (2000), and were further adapted to frame the role data in the particular science-learning context. after several rounds of data scrutinization (see volet et al., 2017 for full description), altogether, ten discrete task-roles were identified and allocated to three broader foci representing respectively, roles that focus on science content (knowledge seeker/provider, information seeker/giver), evaluations on science content and previous actions (challenger, supporter, follower), and personal opinions and viewpoints about the activity (opinion seeker/giver). these roles are next briefly described: the information seeker and information giver feature sharing of task-related facts and information or seeking this kind of information; the knowledge seeker and knowledge provider delve deeper into the content, for example, through providing more elaborated scientific knowledge such as explanations, effects or relations between concepts. ‘providing’ (not ‘giving’) knowledge was used to stress that knowledge cannot be ‘handed over’ to a contrastingly ‘lower’ level of facts and information as it is a product of deeper cognitive content processing (see alexander, 2018); the challenger takes a critical standpoint towards suggestions, statements, and actions. contrary to negative connotations the term criticize may carry, the challenger aims to contribute mainly positively to the group’s performance by, for example, seeking further justifications or suggesting alternative solutions, which could lead to a better quality of task performance as a group; the supporter endorses statements, ideas, and actions. the supporter may, for example, state previous comments in other words to bring further clarity and, simultaneously, show support for and agreement with the contributions of others; the follower also expresses agreement through, for example, utterances such as “yeah”, “ok”, or non-verbal behaviour (e.g., a head nod). however, the follower fails to offer more constructive contributions in relation to task completion (cf. null action; see chiu, 2000), which, thereby, differentiates it distinctly from the supporter role; and the opinion seeker and opinion giver produce statements and comments related mostly to procedural matters, such as how to proceed and what to do next. the term ‘opinion’ stresses that the statements in question can be counted as a personal viewpoint regarding the matter at hand as they are not justified by any scientific content (such as facts or knowledge) and, thus, not conceived as (observationally apparent) attempts to contribute to deeper science-based discourse. data excerpt from dataset 1. the first excerpt is taken from a situation during concept mapping of the clinical case, where the group is trying to understand the relationship between two concepts. this particular group of five students were successful in their efforts to co-construct a meaningful concept map, since their group product was the one evaluated as of the highest quality of all the four groups investigated in the study (see volet et al., 2017). this is in line with the following excerpt, which demonstrates how the students enacted flexibly functional participatory roles that focused on making sense of the science content needed to solve the problem at hand. data excerpt 1. veterinary science students co-constructing a concept map. for starters, matt initiated the discussion based on renee’s question about the relationship of weight loss and azotaemia, but he did not provide scientific backup for his suggestions, thus stayed in opinion sharing realm. however, this triggered a sequence of self-adopted roles among group members that is manifested in group interaction through intensive exchange of ideas, arguments, challenging of previous statements and co-construction of knowledge. after having a while listened this exchange, thea contributed to the discussion by providing a solution to the question. after thea justified her claim with even more convincing evidence in the capacity of knowledge provider role, the group quickly found a common consent on the matter, and started applauding thea for her solution. 2.1.2 dataset 2: senior high school general science students experimenting in a virtual environment the second dataset comes from a study in which roles were used to evaluate the collaborative science learning of senior high school students in a virtual learning environment (heinimäki et al., 2019). the study aimed to establish a coding system of functional participatory roles for this learning context, and use data excerpts to demonstrate the viability of that coding system for fine-grained analyse of roles from group interactions. further, analytical descriptions of these data excerpts from a standpoint of roles were demonstrated to provide meaningful insights into understanding the process and quality of collaborative interactions in different task situations. the participants were senior high school general science students enrolled in advanced-level courses in chemistry and biology. during three lessons (75–90 minutes each), small student groups participated in a virtual expedition on a research vessel in a virtual web-based science-learning environment, virtual baltic sea explorer (vibse), designed to provide students with a realistic research context and an inspiring set of tools to co-construct integrated knowledge with their peers in two disciplines: chemistry and biology (see pietarinen, vauras, laakkonen, kinnunen, & volet, 2019; vauras, telenius, yli-panula, iiskala, pietarinen, & kinnunen, 2017; vauras et al., 2019). each group worked at their own table in a face-to-face setting with a laptop, which was used to operate the virtual learning environment. the scrutinization of roles in the study of heinimäki et al. (2019) included video observations of six groups of triads undertaking a study on the effects of ph changes on a certain species of copepods. in doing this, the groups had to design their study, set hypotheses, carry out experiments by using authentic data from marine biologists, analyse the results, and draw conclusions based on their outcomes. the groups conducted this activity largely autonomously, but teacher support was available if needed. finally, each group produced a presentation describing their study and conclusions (i.e., group outcome) (see vauras et al., 2019). the quality of the group outcome was assessed (on the 6-level scale) by science experts in biology and chemistry; the criteria comprised the structure of the presentation, understanding of the task, hypotheses, research plan, conclusions and the quality of scientific language used in the presentation. the groups selected for the present article represented six diverse groups in terms of the level of the group outcome. the same groups were used in heinimäki et al. (2019), the article that reported the validation of the coding system of the roles enacted by senior high school students. development of the role-coding system for dataset 2. the development of the coding system for the data started with an evaluation of the functionality of the original role-coding system (volet al., 2017) to analyse roles emerging in the different type of science learning context. to accommodate these discrepancies, roles were investigated using data-driven methods, too. the original coding system proved usable as each of its roles were identified from the data at hand as well. however, new types of roles also emerged, calling for the inclusion of additional roles in the coding system under construction. in the end, five activity-specific roles were identified and incorporated into the coding system. in addition, some of the original sub-categories of task-roles (see section 2.1.1) were slightly modified, and a category of experiment and process-focused roles was formed to better fit the context of the experimental activity. accordingly, some fine tuning was made to the conceptualization of some distinct roles to accommodate them better to the data at hand. five newly emerged activity-specific roles were closely linked both to the functional demands of conducting the task in the virtual environment and the digital tools involved in the activity: the navigator is responsible for moving the group around (e.g., between research phases) in vibse by using the mouse and the keyboard of the laptop provided for the group; the attention focuser attempts to focus the attention of the group members on something related to the task content; the recorder performs activities related to keeping records of decisions made and manual controlling of variables in the virtual laboratory; the dictator is closely linked with recording activities; for example, when the outcomes of a joint group discussion are dictated to be further recorded (e.g., what to write in group’s presentation); and the technological contributor performs actions, gives instructions and asks questions related to use of technology in relation to task performance. data excerpt from dataset 2: the second excerpt is drawn from a phase of the activity where the group engages to interpret the results of their experiments. this group was categorized as a high-outcome group based on the quality of the presentation they later gave to the rest of the class (see heinimäki et al., 2019). the following excerpt not just illustrates the interactive nature of functional participatory roles self-adopted in the situation, but also the role the affordances and constrains of the virtual learning environment played in shaping group interaction. data excerpt 2. high school general science students interpreting the results of their virtual experiment. the excerpt emerges at a point when the group is about getting stuck discussing irrelevant issues. at this point, ellen encourages the group to move forward with the task. the opinion giver role adopted by ellen demonstrates the impact of the virtual environment and the technology involved in this specific activity. the activity-specific roles of navigator, recorder and attention focused enacted after this, further illustrate the impact of the activity on the enacted roles. the emergence of this pattern of roles was triggered by the fact that the groups had to operate with only one mouse, keyboard and screen, which means that only one student at a time could use the technology (in this excerpt sofia). this environmental constraint had implications on the roles sofia enacted (e.g., navigator, recorder) in situ, but also impacted on the other students who had to describe very clearly the actions they wanted the person operating the equipment to perform during the evolving activity (e.g., opinion giver role adopted by ellen). in addition, the excerpt illustrates the emergent nature and affordances of the activity as the group conducts their experiments in the virtual laboratory. for instance, the group received some results immediately, and in this case the results were unexpected. the first reaction of ellen was to understand the scientific reasons for it (knowledge seeker), whereas paula just settled by stating that they performed poorly as a group with no scientific interest (opinion giver). 2.1.3 dataset 3: preservice primary school teachers’ conducting hands-on science experiments the third dataset is derived from a study in which preservice primary school student teachers’ attitudes and productive engagement in collaborative science learning within a mandatory introductory science course was investigated (volet et al., 2019). productive engagement was investigated through the task-focused roles spontaneously enacted by the students during two hands-on science laboratory activities grounded in a scientific inquiry approach and undertaken in small groups (see pino-pasternack & volet, 2018; volet et al., 2019). the first activity focused on learning about chemical reactions, which was investigated through fair tests with small ‘rockets’. the groups had to design their test, develop research questions and hypotheses, choose research variables, conduct their test, and interpret the results; this activity was also aimed at promoting students’ understanding about the fundamentals of carrying out a scientific investigation. the second activity was more exploratory; the groups investigated how to make electric circuits with play dough and various materials that were provided. this activity aimed mainly at generating meaningful exchange of questions and ideas among the students as they tried to interpret their observations. the laboratory sessions lasted approximately two hours. the groups were video-taped, and emergent task-related roles were analysed from four groups (four students each) that participated in both activities. the groups typically comprised a somewhat heterogeneous mix of students regarding prior science skills and attitudes towards learning science (see volet et al., 2019). the role analysis revealed that, overall, roles focusing on science content and experimenting-related activities were more commonly enacted in the first, more structured activity (‘rocket’), whereas opinion sharing focused roles were in comparison more prevalent in the second, more exploratory activity (‘circuit’). the analysis of roles also revealed finer distinctions in relation to attitude-, and group-related differences in the quality of individual and group engagement in the two activities (see volet et al., 2019). development of the role-coding system for dataset 3. a procedure similar to that described for the previously described dataset (see section 2.1.2) was followed: this included establishing the relevance of the original coding system (volet et al., 2017) for the data at hand and data-driven investigation of other activity-specific roles which may have emerged during the two science activities. again, all the roles described by volet et al. (2017) were identified, but some minor modifications that were quite similar to those described in the case of dataset 2 were implemented to adapt the coding system for this specific study. altogether, three activity-specific roles were identified. again, these roles largely interweaved with the procedural task performance, and materials/tools involved in the hands-on experiments. the reader reads instructions or other information aloud for the group. this was typically from the lab manual that was provided to each student; the procedural contributor focuses on processes and procedures such as by providing general comments about the materials, filling in answers in the lab manual, and taking notes, which do not include contributions involving (observably apparent) deep cognitive processing of the science content (e.g., scientific explanations); and the observation maker conducts rather straightforward observations towards, for example, what is occurring during the hands-on experiment. data excerpt from dataset 3: the third excerpt is from the activity which focused on chemical reactions (‘rocket’), and is taken from a specific phase in which the group was planning their experiment. a few utterances showing agreement and other low-level contributions (i.e., follower role) were removed to shorten the excerpt. the number of self-adopted roles focusing on science content within that group was the lowest across all four groups investigated in that study (see volet et al., 2019), which is also consistent with the dialogue in the following excerpt. data excerpt 3. preservice primary school teachers’ planning a hands-on science experiment. the quality of this group interaction remained at low level throughout the excerpt in reference to science, as students enacted primarily roles that were in the opinion and procedural realms. for instance, as they went on with planning and decision making, the students relied largely on their personal opinions about the matter. a few times an attempt was made to encourage a science-based discourse through the role of information giver, but these events did not trigger any sustained and deep science-based argumentation or knowledge co-construction. in this respect, the observed pattern of enacted roles within this group was very different compared to the exchange illustrated earlier with veterinary students in data excerpt 1. 2.2 summary the investigation of task-related roles in each of the three datasets provided support for the assumption of spontaneously enacted core and activity-specific roles during collaborative science learning. nine roles were commonly found across all learning environments and could therefore be conceptualized as core roles. table 1 provides a definition and indicators for each of these core roles, with examples from the different datasets as illustrations. these definitions and indicators are, however, not assumed to be totally rigid and allow for some flexibility. for example, some slight adaptations were made to the conceptualisation of the roles of knowledge provider/seeker and opinion giver/seeker in both datasets 2 and 3 to reach a better correspondence with the functions that these roles actually played in those activities (e.g., the nature of acquired knowledge slightly varied depending on the characteristics of an activity). furthermore, and consistent with the assumption, there was evidence that some roles were only enacted in certain activities and could therefore be conceptualized as activity-specific roles. table 2 provides a definition and indicators of these roles, with examples from the different datasets. in the tables, core and activity-specific roles were grouped under a broad classification of task-related roles, although more detailed classifications into sub-categories would be plausible. for example, the task-related roles identified in the initial study (volet et al., 2017) were allocated to three distinct foci, each capturing a different form of engagement in the collaborative science-learning activity (see section 2.1.1). these sub-categories were, however, re-arranged in the analysis of the two subsequent datasets in order to better fit with the context of experimental activities. importantly, rigid sub-categories were avoided in order to keep the framework as flexible as possible. table 1 overview of core roles in collaborative science learning table 2 overview of activity-specific roles in collaborative science learning 3. discussion student-led collaborative learning in science has become increasingly common at different levels of education. yet, to date, the roles spontaneously enacted by students during these types of collaborative activities have received limited empirical attention – prior research focusing mainly on pre-assigned roles in small task groups (e.g., gu et al., 2015; morris et al., 2010; pozzi, 2011; strijbos et al., 2004; for a review, see cohen, 1994). by focusing on a single learning setting and typically downplaying the spontaneous, emergent and interactional nature of roles (see oliveira et al., 2014; lehmann-willenbrock et al., 2016; salazar, 1996), prior research has limited the development of generalizable conceptual frameworks aimed at understanding the impact of roles in and across collaborative learning activities as they unfold in real-time. in this article, empirical support was provided for the importance of paying attention to the emergent and interactional nature of roles in student-led collaborative learning in science and the nature of these roles. two claims are made. the first claim is that like other key processes in small-task-group learning, such as socially shared metacognitive regulatory processes (e.g., iiskala, volet, lehtinen, & vauras, 2015), and emotion and motivation regulation (e.g., järvenoja, järvelä, & malmberg, in press), there is a need to understand how naturally emerging functional participatory roles meaningfully contribute to learning and task completion. shifting the emphasis from pre-assigned fixed roles to roles as naturally emerging, evolving, and dynamic in situ is consistent with greeno’s (2006) conceptualization of learning in activity systems “in which learners interact with each other and with material, informational, and conceptual resources in their environment” (p. 92). the empirical support gathered in the present study through the identification of functional participatory roles enacted in the three datasets supports the claim that roles emerge during a group activity through ongoing and intertwined interaction between the group and its environment. the second claim is grounded in the assumption, supported by empirical evidence from the three datasets, that during collaborative science-learning activities, group members spontaneously enact two types of functional, participatory task-related roles: a. the core roles reflect inherently the nature of the science discipline (see anderson, 2007; duschl & hamilton, 2011 for elaboration on the nature of science) and, therefore, the uptake of these roles is at the heart of productive collaborative science learning; and b. the activity-specific roles depend on the characteristics of the activity, for example, the task demands, context and settings, and, therefore, these roles play an important part in successful collaborative science learning. data from the three diverse science-learning environments, including four collaborative science activities, provided empirical support for the assumption that in collaborative science learning, some task-related roles are commonly found across all science-learning activities (i.e., core roles) while others are only found in certain science-learning activities (i.e., activity-specific roles). in spite of notable contextual differences in some aspects (i.e., educational level, student characteristics, learning environment, learning task), similar core roles were observed in the processing of the science content by student groups across environments, domains, and activities. in contrast, the majority of activity-specific roles were linked to the more technical or practical aspects of performing and accomplishing the learning task (i.e., recording progress, navigating in the virtual environment, reading aloud from the lab materials, etc.). for example, the range of activity-specific roles that were identified, such as navigator, technological contributor, and observation maker, emerged and were enacted largely due to specifically situated affordances and constraints related to carrying out the particular activity. these two empirically supported claims led to the proposal of a flexible conceptual framework for the analysis and understanding of the core and activity-specific functional participatory roles emerging during collaborative science learning (see figure 1 figure 1. a conceptual framework of the core and activity-specific functional participatory roles emerging during collaborative science learning. this figure incorporates all the elements of the framework and its underlying claims. reading from the bottom-up, the figure shows how during science-learning activities carried out in small groups, students enact functional participatory roles which are emergent, task-related, dynamic, and evolve in situ. the middle part of the figure distinguishes the two types of roles that can emerge during an activity, core and activity-specific, and their characteristics and functionality regarding learning and task completion are specified. the dotted lines with bi-directional arrows between each type of role and the science-learning activity stress the interactive, dynamic, and intertwined nature of roles during an activity. the arrow with a question mark above the box of roles indicates the expected link to the quality of collaborative science learning. although this aspect was not examined in the present study, other research has provided some empirical support for this relationship (e.g., volet et al., 2017). the study of volet et al. (2017), however, did not distinguish between core and activity-specific roles. on the grounds that the emergence of core roles is conceptualized as reflecting inherently the nature of the science discipline, whereas the emergence of activity-specific roles is assumed to depend on the characteristics of the particular science activity, it can be argued that both core and activity-specific roles play an important part in successful collaborative science learning, but how these roles are taken up by group members during an ongoing activity has an impact on the quality of collaborative science learning. to sum up and conclude, the contributions of the proposed framework for research on collaborative science learning can be understood as threefold: a. conceptually, shifting the focus from fixed roles to roles that are naturally emerging and evolving in situ is consistent with the literature characterising collaborative learning as highly dynamic (hadwin et al., 2018; hilpert & marchand, 2018). the adoption of a situative approach to better understand the task-related, functional participatory roles that naturally emerge during collaborative activities addresses an earlier, yet still valid, call that researchers need to acknowledge the significance of roles in a “comprehensive social-psychology theory” (hare, 1994, p. 434). this approach also addresses hoadley’s (2010) call for the need to understand “where roles come from, and how they might emerge” (hoadley, 2010, p. 545), which he claimed was not always clearly stated in prior research; b. methodologically, identifying two types of functional participatory roles, that is, those commonly enacted during collaboration across science activities and those that are activity-specific, the framework adds value to existing frameworks, which have examined empirically the impact of roles mainly in single learning settings. thus, the proposed framework addresses the calls for more scalable analytical frameworks that are flexible and sensitive to data specificity as well (cf. volet & summers, 2013). in the present case, a degree of cross-dataset generalizability is reached through the identification of core roles commonly found in diverse collaborative science-learning activities, and sensitivity to data through activity-specific roles that are dependable on the characteristics of the activity; and c. empirically, as indicated in figure 1, the framework is expected to provide potential for further analysis of the significance of roles for productive collaborative science learning through investigations of how core and activity-specific roles that are relevant for successful task completion are enacted during collaborative science-learning activities. thus far, there is some evidence that certain core roles, such as knowledge provider, knowledge seeker, and challenger, can meaningfully explain, for example, attitudes towards learning, engagement in deep learning, and the quality of collaborative science outcomes (volet et al., 2017, 2019). although these findings offer reasonable validation for future use of the framework beyond single datasets when the aim is to obtain insights into high-quality collaborative science learning, more research is needed to further establish these relations and gain deeper insight into how activity-specific roles operate in conjunction with core roles in order to strengthen productive collaborative science learning. in addition, future research aimed at understanding better the importance of core and activity-specific roles in collaborative science-learning bears also practical implications for science education. teachers may not be used to guiding students during collaborative inquiry learning, since this form of teaching does not necessarily fit the typical structures and norms of classrooms and schools that they are familiar with. many are often challenged as to how to provide adequate support for student groups, in particular if students are novices without a repertoire of disciplinary practices (vauras et al., 2019; see also kirschner, sweller, & clark, 2006). in this respect, a better understanding of how enacted, intertwined functional participatory roles impact on the quality of learning and task completion may provide some guidance to teachers using collaborative learning activities in classrooms. as spontaneously enacted roles serve potentially more readily observable indicators of the quality of groups’ productive engagement than, for instance, regulatory processes, such understanding can help teachers to meaningfully calibrate their support and scaffolding to student groups during ongoing collaborative activities. for instance, if non-beneficial or even detrimental role patterns start emerging within a group, an informed teacher could try to intervene in the ongoing interaction early on, before the harmful patterns escalate any further. teachers can thus also utilize such understanding to inform, model and encourage flexible enactment of the most productive functional participatory roles during collaborative activities. an appreciation of the significance of self-adopted roles could also help teachers and developers design new, facilitative features in collaborative science-learning activities and environments. furthermore, such an appreciation may contribute to address the possible mismatch between design and reality, meaning that “roles-as-intended” can differ from “roles-as-enacted” (hoadley, 2010, p. 553). thus, the observation of roles emerging spontaneously during an activity may be worthwhile exploring further with a view to establish if a particular learning environment actually promotes desired types of behaviours and learning (e.g., high-level co-construction of knowledge; see e.g., volet et al., 2009; webb, 2008), i.e. it was designed for or if it actually feeds the uptake of “not-intended” roles. finally, it is not surprising that as learning activities are increasingly taking place across diverse settings and contexts, there have been calls for research that unravels both the unique and parallel features across such activities (ludvigsen, lund, rasmussen, & säljö, 2011). the proposed framework addresses this call by highlighting the emerging functional participatory roles, both common and unique, found in collaborative science-learning, and at the same time, by making a contribution to the broader ongoing debate over the interplay of domain-generic and domain-specific learning in science (see duschl & hamilton, 2011). 3.1 limitations and other considerations it is acknowledged that the empirical support for the assumption of core and activity-specific roles in collaborative science learning is based only on three small datasets. more and larger datasets will be needed for further validation of the proposed conceptual framework and this will be especially important to validate the claim that core roles can, in fact, be found across a wide range of collaborative science-learning activities. the outcome of future investigations may perhaps also lead to a reconsideration of the overall composition of the core roles presented in the present study. for example, given that collaborative learning activities typically involve some kind of recording (i.e., note taking, writing from a dictation, generating a research report, etc.), it could be that bundling up all these behaviours together for analytical purposes would be consistent with the proposal to consider the recorder as a “classic functional role” (morris et al., 2010, p. 816; see e.g., benne & sheats, 1948, 2007). however, a reductionist approach could end up being problematic as the nature of recording-related activities and behaviours tend to vary extensively between learning environments (cf. e.g., the recorder in dataset 2 and the procedural contributor in dataset 3), which could then lead to the risk of oversimplification and the inhibiting of the process of fully capturing the rich activity-specific data in question. one limitation to the generalizability of the proposed framework is the fact that the data provided as empirical support involved only certain types of collaborative science activities, and was restricted to face-to-face interactions. nevertheless, although straightforward generalizations cannot be drawn, it is reasonable to expect that this framework could be applicable to some other domains, or at least to collaborative science learning in other stem (science, technology, engineering, and mathematics) disciplines. when it comes to asynchronous and distance-learning environments, where group functioning and individual behaviour can somewhat vary compared to face-to-face activities, the proposed framework and its derived coding systems could possibly fall short of capturing certain roles that are characteristic of these environments. thus, this could lead to a call for a different kind of approach, such as additionally focusing on the macro-level of roles, as suggested by strijbos and laat (2010). it would be intriguing to put to the test how the framework would operate in cscl environments where students interact synchronously, such as in chats (sarmiento & shumar, 2010). it would be especially important to explore what kind of activity-specific roles emerge in these types of environments as the datasets scrutinized in the present article suggest that the number of roles needed to execute task-related functions can be higher in virtual environments compared to more “conventional” learning environments (see table 2). this is noteworthy from the standpoint of learning and group performance because the task-based demands of mastering multiple different roles can be challenging for the group members (cf. role flexibility; see benne & sheats, 1948, 2007; forsyth, 2014). one more limitation is linked to the possibility that individuals’ behaviour and roles can differ in larger groups compared to smaller sized groups (forsyth, 2014), which may possibly influence the enactment of roles (hare, 1994; strijbos & laat, 2010). this suggests that the applicability of the framework should be explored outside of the small-group context. furthermore, as the samples of students involved in the present study were senior high school and university students, the applicability of the framework and its derived set of core and activity-specific roles should be extended to research with groups of younger students and possibly to workplace environments. finally, as the proposed framework focuses exclusively on functional participatory roles that are task-related, future research may also explore the relevance of socio-emotional roles. it could be argued that socio-emotional roles are neither core nor activity-specific but rather emerge at the interface of learner characteristics and activity specificity or that the enactment of these roles is influenced by the atmosphere and orientation of the group towards the collaborative task. for example, if a learning activity is perceived by some students as not taking into consideration their personal goals, interest, motivation, prior skills, and knowledge, this could lead to feelings of frustration and lower engagement, even the development of negativity towards the activity or collaboration with others, and vice versa in terms of more positive socio-emotional roles (cf. strijbos & laat, 2010). this may also influence how task-related roles are enacted and, in turn, the quality of the group’s task performance. there is no doubt that the potential of the proposed framework to represent the significance of roles in authentic, real-life collaborative learning situations will need further validation and exploration. keypoints there is a scarcity of scalable frameworks aimed at understanding the significance of roles in and across collaborative science-learning activities. the functional participatory roles enacted by group members in collaborative learning are conceptualized as emergent, dynamic, and evolving in situ. a flexible conceptual framework for the analysis of emergent functional participatory roles across collaborative learning activities is proposed. the framework builds on the assumption that during collaborative learning, both core roles and activity-specific roles are enacted. three collaborative science-learning datasets were scrutinized to establish the degree of empirical support for the proposed framework. acknowledgments this research was supported by grant no. 274117 from the academy of finland, awarded to the third author, and by the australian research council under the discovery award (dp150101142), awarded to the second author. references alexander, p. a. (2018). information management versus knowledge building: implications for learning and assessment in higher education. in o. zlatkin-troitschanskaia, m. toepper, h. a. pant, c. lautenbach, & c. kuhn (eds.), assessment of learning outcomes in higher education: cross-national comparisons and perspectives , 43–56. cham: springer international publishing. https://doi.org/10.1007/978-3-319-74338-7_3 anderson, c. h. (2007). perspectives on science learning. in s. k. abell & n. g. lederman (eds.), handbook of research on science education. new jersey: lawrence erlbaum associates. bales, r. f. (1950). interaction process analysis: a method for the study of small groups . cambridge: allison-wesley. bales, r. f., & slater, p. e. (1955). role differentiation in small decision-making groups. in t. parsons & r. f. bales (eds.), family, socialization and interaction process (pp. 259–306). illinois: the free press. belbin, m. (1993). team roles at work. oxford: butterworth-heinemann. benne, k. d., & sheats, p. (2007, reprinted). functional roles of group members. group facilitation, 8, 30–35. https://doi.org/10.1111/j.1540-4560.1948.tb01783.x benne. k. d., & sheats, p. (1948). functional roles of group members. journal of social issues, 4, 41–49. https://doi.org/10.1111/j.1540-4560.1948.tb01783.x cesareni, d., cacciamani, s., & fujita, n. (2016). role taking and knowledge building in a blended university course. international journal of computer-supported collaborative learning, 11 (1), 9–39. https://doi.org/10.1007/s11412-015-9224-0 cheng, b., wang, m., & mercer, n. (2014). effects of role assignment in concept mapping mediated small group learning. internet and higher education, 23, 27–38. https://doi.org/10.1016/j.iheduc.2014.06.001 chiu, m. m. (2000). group problem-solving processes: social interactions and individual actions. journal for the theory of social behavior, 30(1), 27–49. https://doi.org/10.1111/1468-5914.00118 cohen, e. g. (1994). restructuring the classroom: conditions for productive small groups. review of educational research, 64(1), 1–35. https://doi.org/10.3102/00346543064001001 de wever, b., keer, h. van, schellens, t., & valcke, m. (2010). roles as a structuring tool in online discussion groups: the differential impact of different roles on social knowledge construction. computers in human behavior, 26(4), 516–523. https://doi.org/10.1016/j.chb.2009.08.008 de wever, b., van keer, h., schellens, t., & valcke, m. (2009). structuring asynchronous discussion groups: the impact of role assignment and self-assessment on students’ levels of knowledge construction through social negotiation. journal of computer assisted learning, 25(2), 177–188. https://doi.org/10.1111/j.1365-2729.2008.00292.x. driskell, t., driskell, j. e., burke, c. s., & salas, e. (2017). team roles: a review and integration. small group research, 48 (4), 482–511. https://doi.org/10.1177/1046496417711529 duschl, r., & hamilton, r. (2011). learning science. in r. e. mayer & p. a. alexander (eds.), handbook of research on learning and instruction (pp. 78–107). new york & london: routledge. forsyth, d. r. (2014). group dynamics (6th ed.). belmont: wadsworth cengage learning. greeno, j. g. (1998). the situativity of knowing, learning, and research. american psychologist, 53, 5–26. https://doi.org/10.1037/0003-066x.53.1.5 greeno, j. g. (2006). learning in activity. in r. k. sawyer (ed.), the cambridge handbook of the learning sciences (pp. 79–96). new york, ny, us: cambridge university press https://doi.org/10.1017/cbo9781139519526.009 gu, x., shao, y., guo, x., & lim, c. p. (2015). designing a role structure to engage students in computer-supported collaborative learning. internet and higher education, 24, 13–20. 10.1016/j.iheduc.2014.09.002 hadwin, a. f., järvelä, s., & miller, m. (2018). self-regulation, co-regulation and shared regulation in collaborative learning environments. in d. h. schunk, & j. a. greene (eds.).handbook of self-regulation of learning and performance (2 nd ed.) (pp. 83–106). new york, ny. hare, a. p. (1994). types of roles in small groups. small group research, 25(3), 433–448. https://doi.org/10.1177/1046496494253005 heinimäki, o-p., salo, a-e., & vauras, m. (2019). luonnontieteiden yhteisöllisessä tietokoneavusteisessa oppimisessa omaksuttujen funktionaalisten osallistumisen roolien luokittelun kehittely [development of a classification for functional participatory roles enacted during computer-supported collaborative science learning]. psykologia, 54(04), 236–254. hilpert, j. c., & marchand, g. c. (2018). complex systems research in educational psychology: aligning theory and method. educational psychologist, 53(3), 185–202. https://doi.org/10.1080/00461520.2018.1469411 hoadley, c. (2010). roles, design, and the nature of cscl. computers in human behavior, 26(4), 551–555. https://doi.org/10.1016/j.chb.2009.08.012 hogan, k. (1999). sociocognitive roles in science group discourse. international journal of science education, 21(8), 855–882. https://doi.org/10.1080/095006999290336 iiskala, t., volet, s., lehtinen, e., & vauras, m. (2015). socially shared metacognitive regulation in asynchronous cscl in science: functions, evolution and participation. frontline learning research, 3(1), 78–111. https://doi.org/10.14786/flr.v3i1.159 järvenoja, h., järvelä, s., & malmberg, j. (in press). supporting groups’ emotion and motivation regulation during collaborative learning. learning and instruction. https://doi.org/10.1016/j.learninstruc.2017.11.004 johnson, d. w., & johnson, r. t. (1989). cooperation and competition: theory and research. edina: interaction book company. khosa, d. k., & volet, s. e. (2013). promoting effective collaborative case-based learning at university: a metacognitive intervention. studies in higher education, 38, 870–889. http://dx.doi.org/10.1080/03075079.2011.604409 khosa, d. k., & volet, s. e. (2014). productive group engagement in cognitive activity and metacognitive regulation during collaborative learning: can it explain differences in students’ conceptual understanding? metacognition and learning, 9, 287–307. https://doi.org/10.1007/s11409-014-9117-z khosa, d. k., volet, s. e., & bolton, j. r. (2010). an instructional intervention to encourage effective deep collaborative learning in undergraduate veterinary students. journal of veterinary medical education, 37, 368–375. https://doi.org/10.3138/jvme.37.4.369 khosa, d., volet, s. e., & bolton, j. (2014). clinical case-based learning in health sciences: analysis of collaborative concept mapping processes and reflections. journal of veterinary medical education , 41, 406–417. https://doi.org/10.3138/jvme.0314-035r1 kirschner, p., sweller, j., & clark, r. (2006). why minimal guidance during instruction does not work: an analysis of the failure of constructivist, discovery, problem-based, experiential and inquiry-based teaching. educational psychologist, 41, 75–86. https://doi.org/10.1207/s15326985ep4102_1 landis, j. r., & koch, g. g. (1977). the measurement of observer agreement for categorical data. biometrics, 33, 159–174. doi:10.2307/2529310 lehmann-willenbrock, n., beck, s. j., & kauffeld, s. (2016). emergent team roles in organizational meetings: identifying communication patterns via cluster analysis. communication studies, 67(1), 37–57. https://doi.org/10.1080/10510974.2015.1074087 ludvigsen, s., lund, a., rasmussen, i., & säljö, r. (2011). introduction. in s. ludvigsen, a. lund., i. rasmussen, & r. säljö (eds.), learning across sites: new tools, infrastructures and practices (pp. 1–13). new york: routledge. maloney, j. (2007). children’s roles and use of evidence in science: an analysis of decision-making in small groups. british educational research journal, 33(3), 371–401. https://doi.org/10.1080/01411920701243636 marcos-garcía, j. a., martínez-monés, a., & dimitriadis, y. (2015). despro: a method based on roles to provide collaboration analysis support adapted to the participants in cscl situations. computers and education, 82, 335–353. https://doi.org/10.1016/j.compedu.2014.10.027 morris, r., hadwin, a. f., gress, c. l. z., miller, m., fior, m., church, h., & winne, p. h. (2010). designing roles, scripts, and prompts to support cscl in gstudy. computers in human behavior, 26 (5), 815–824. https://doi.org/10.1016/j.chb.2008.12.001 moxnes, p. (1999). understanding roles: a psychodynamic model for role differentiation in groups. group dynamics: theory, research, and practice, 3(2), 99–113. https://doi.org/10.1037/1089-2699.3.2.99 nolen, s. b., horn, i. s., & ward, c. j. (2015). situating motivation. educational psychologist, 50(3), 234–247. https://doi.org/10.1080/00461520.2015.1075399. oliveira, a. w., boz, u., broadwell, g. a., & sadler, t. d. (2014). student leadership in small group science inquiry. research in science and technological education, 32(3), 281–297. https://doi.org/10.1080/02635143.2014.942621 pietarinen, t., vauras, m., laakkonen, e., kinnunen, r., & volet, s. (2019). high school students’ perceptions of affect and collaboration during virtual science inquiry learning. journal of computer assisted learning, 35, 334–348. https://doi.org/10.1111/jcal.12334 pino-pasternak, d., & volet, s. (2018). evolution of pre-service teachers’ attitudes towards learning science during an introductory science unit. international journal of science education, 40(12), 1520–1541. https://doi.org/10.1080/09500693.2018.1486521 pozzi, f. (2011). the impact of scripted roles on online collaborative learning processes. international journal of computer-supported collaborative learning , 6(3), 471–484. doi: 10.1007/s11412-011-9108-x salazar, a. j. (1996). an analysis of the development and evolution of roles in the small group. small group research, 27(4), 475–503. https://doi.org/10.1177/1046496496274001 sarmiento, j. w., & shumar, w. (2010). boundaries and roles: positioning and social location in the virtual math teams (vmt) online community. computers in human behavior, 26(4), 524–532. https://doi.org/10.1016/j.chb.2009.08.009 schellens, t., van keer, h., de wever, b., & valcke, m. (2007). scripting by assigning roles: does it improve knowledge construction in asynchronous discussion groups ? international journal of computer-supported collaborative learning, 2 , 225–246. https://doi.org/10.1007/s11412-007-9016-2 slavin, r. (1996). research on cooperative learning and achievement: what we know, what we need to know. contemporary educational psychology , 69(1), 43–69. https://doi.org/http://dx.doi.org/10.1006/ceps.1996.0004 stewart, g. l., fulmer, i. s., & barrick, m. r. (2005). an exploration of member roles as a multilevel linking mechanism for individual traits and team outcomes. personnel psychology, 58, 343–365. https://doi.org/10.1111/j.1744-6570.2005.00480.x strijbos, j. w., & de laat, m. f. (2010). developing the role concept for computer-supported collaborative learning: an explorative synthesis. computers in human behavior, 26(4), 495–505. https://doi.org/10.1016/j.chb.2009.08.014 strijbos, j. w., & weinberger, a. (2010). emerging and scripted roles in computer-supported collaborative learning. computers in human behavior, 26(4), 491–494. https://doi.org/10.1016/j.chb.2009.08.006 strijbos, j.-w., martens, r. l., jochems, w. m. g., & broers, n. j. (2004). the effect of functional roles on group efficiency: using multilevel modeling and content analysis to investigate computer-supported collaboration in small groups. small group research, 35(2), 195–229. https://doi.org/10.1177/1046496403260843 turner, j. c., & nolen, s. b. (2015). introduction: the relevance of the situative perspective in educational psychology. educational psychologist, 50(3), 167–172. https://doi.org/10.1080/00461520.2015.1075404 ucan, s., & webb, m. (2015). social regulation of learning during collaborative inquiry learning in science: how does it emerge and what are its functions? international journal of science education, 37(15), 2503–2532. https://doi.org/10.1080/09500693.2015.1083634 vauras, m., telenius, m., yli‐panula, e., iiskala, t., pietarinen, t., & kinnunen, r. (2017). virtuaalinen tutkimusmatka luonnontieteelliseen osaamiseen [virtual exploration into learning science]. in h. savolainen, r. vilkko, & l. vähäkylä (eds.), oppimisen tulevaisuus [ future of learning], 24–35. porvoo: gaudeamus. vauras, m., volet, s., & nolen, s. (2019). supporting motivation in collaborative learning: challenges in the face of an uncertain future. in e. gonida & m. lemos (eds.), motivation in education at a time of global change: theory, research, and implications for practice (pp. 187–203). new york: emerald. volet, s., & summers, m. (2013). interpersonal regulation in collaborative learning activities: reflections on emerging research methodologies. in s. volet & m. vauras (eds.), interpersonal regulation of learning and motivation: methodological advances (pp. 204–220). london & new york: routledge. volet, s., jones, c., & vauras, m. (2019). preservice primary teachers’ science learning: effects of within-group diversity of attitudes on productive engagement. learning and individual differences, 73, 79–91. https://doi.org/10.1016/j.lindif.2019.05.002 volet, s., summers, m., & thurman, j. (2009). high-level co-regulation in collaborative learning: how does it emerge and how is it sustained? learning and instruction, 19(2), 128–143. https://doi.org/10.1016/j.learninstruc.2008.03.001 volet, s., vauras, m., salo, a. e., & khosa, d. (2017). individual contributions in student-led collaborative learning: insights from two analytical approaches to explain the quality of group outcome. learning and individual differences, 53, 79–92. https://doi.org/10.1016/j.lindif.2016.11.006 webb, m. (2008). learning in small groups. in t. l. good (ed.), 21st century education: a reference handbook (pp. 203–211). thousand oaks: sage publications. codepen heemskerken publication frontline learning research vol.8 no. 6 (2020) 38 58 issn 2295-3159 students’ observed engagement in lessons, instructional activities, and learning experiences christina hubertina helena maria heemskerk1,2, lars-erik malmberg1 1department of education, oxford university, united kingdom 2institut für psychologie, universität bern, switzerland article received 3o january 2020/ revised 24 july 2020/ accepted 26 august/ available online 16 september abstract in order to expand previous intraindividual studies of student engagement we investigated students' observed engagement (i.e., onand off-task behaviour), instructional activities (i.e., teacher-led whole class, individual work, pair-work, student-teacher interaction, assessment, and ”other”), and self-reported learning experiences (cognitive engagement, difficulty, competence, emotional engagement, positive and negative emotions), within lessons during one calendar week. eighteen fourth and fifth grade target students (mage=10.1, sd=0.44) were observed every 30 sec during two to four lessons each day for five school days (engagement and instructional activities), on average 66.05 times per lesson (sd=19.16, range=15-80, nobs=14,994) between 9-18 lessons during a week. simultaneously, students provided 1-3 electronic questionnaire self-reports per lesson (mself_report=35.1, sd=12.6, range=19-52, nself_report=631). we regressed observed engagement (0 = off-task, 1 = on-task) on self-reported learning experiences using 3-level (time-points nested in lessons, nested in students) bayesian logistic regression models in brms. observed engagement diminished during lessons, and was predicted by higher cognitive engagement, and instructional activities. as compared to teacher-led instruction, engagement was higher during individual tasks, teacher-supported tasks, and assessments. overall self-reported and observed engagement within lessons converged, supporting their use in intraindividual research keywords: intraindividual; engagement; observation; ecological momentary assessment; bayesian info corresponding author email: christina.heemskerk@psy.unibe.ch doi: https://doi.org/10.14786/flr.v8i6.613 1. introduction students' engagement is essential for learning. “academic motivation refers to an individual’s inclination, energy, direction, and drive with respect to learning and achievement (martin, ginns and papworth, 2017). academic engagement refers to the thoughts, behaviours, and emotions that reflect this inclination, energy, and drive (fredricks, blumenfeld and paris, 2004; martin et al., 2017). thus, engagement may be considered the manifestation of motivation—the thoughts, actions, and emotions that an individual undertakes or experiences as a result of his or her motivation (martin, ginns, et al., 2017).” (collie and martin, 2019: 2). in cross-sectional and longer-term longitudinal studies, engagement has been moderately to strongly associated with academic performance. roorda et al. (2011) report a correlation of r=0.29 between engagement and academic achievement, whilst godwin and fisher (2011) report that correlations between time-on-task and learning outcomes range from 0.13 to 0.71, making student engagement an important focus of research. to expand previous studies, we had three objectives. we first investigated how observed engagement varied within and between lessons, and between students. we expected engagement to vary mostly between individual students, with some situation-specific variation. in doing so, we go beyond studies in which engagement is observed during one lesson or subject (e.g. grieco, jowers and bartholomew, 2009; jarrett et al., 1998), at the class level (e.g. pöysä et al., 2017; barros, silver and stein, 2009), or reported on by teachers (e.g. carlson et al., 2015; barros et al., 2009) or parents (e.g. watson et al., 2019). second, we investigated how observed engagement was related to observed instructional activities. consistent with findings from previous studies we expected on-task behaviour to be more prevalent during one-to-one student-teacher interactions, and pair-work than during teacher-led instruction (e.g. godwin et al., 2013). third, we investigated how observed engagement was related to self-reported learning experiences students’ thoughts, actions, and emotions reported at the beginning, middle, and end of each lesson. this goes beyond studies in which self-reports are either collected randomly throughout the school day, or systematically once a lesson. we expected within-lesson variability in on-task behaviour, and higher levels thereof, to be associated with more cognitive and emotional engagement, higher competence, less difficulty, more positive, and fewer negative emotions (e.g. malmberg, woolgar and martin, 2013b). we used the bayesian estimation technique in r, implemented in brms (bürkner, 2017; 2018), as this technique is well suited for rich multilevel data (14,994 time-points nested in 56 lessons nested in 18 students) with sparse observations at the highest level (hox, van de schoot and matthijsse, 2012). 1.1 variation in observed engagement a growing body of intraindividual studies focuses on situation-specific engagement (e.g. patall et al., 2016), showing large variation in engagement within students, between one situation and another. although some studies include student-reported perceptions of the teacher (e.g., autonomy support; (tsai et al., 2008)), or global observations of classroom interaction quality (pöysä et al., 2019), there is to the best of our knowledge no previous study of students’ observed intraindividual engagement during real-time instructional activities over an extended time-period. observation studies differ in focus: either on the quality of interaction or on observable behaviours. while observation of interaction quality typically involves some degree of inference about the underlying qualities over a relatively longer sequence of time, observations of behaviours typically involve less inference and shorter but repeated observation timeframes. momentary time sampling of behavioural engagement in young children has been found to be most accurate at short intervals (5-15 seconds) when compared to continuous recording (zakszeski, hojnoski and wood, 2017). observations of quality are typically domain-specific, focusing on agency and communion (e.g. wubbels et al., 2015), classroom organisation (e.g., providing time to learn, minimizing disruptions), instructional support (e.g., providing explanations that support cognitive development) and emotional support (e.g., warmth, closeness) (e.g. praetorius, lenske and helmke, 2012). while instructional support is associated with academic performance, and emotional support with fewer behavioural problems (la paro and pianta, 2000), all these dimensions are positively associated with student engagement (e.g. malmberg et al., 2010). in the present study we focused on the teachers’ instructional activities, which can be considered an aspect of classroom organisation. 1.2 observed engagement and instruction format students' engagement can be subdivided into behavioural, emotional, and cognitive components (fredricks et al., 2004). the key variable in the current study is behavioural engagement, as behavioural engagement can be assessed through observations in a naturalistic environment. in studies of classroom behaviour, the terms 'on-task', 'engagement', and 'attention' have been used quite interchangeably (fredricks et al., 2004). common measures are time spent on-task or off-task (godwin et al., 2013; pellegrini, huberty and jones, 1995; pellegrini and davis, 1993), with off-task often further categorised by type of behaviour displayed (rock, 2005; pellegrini and davis, 1993), or the source of distraction (godwin and fisher, 2011). following fredricks et al. (2004), off-task can be grouped into active/disruptive (fidgeting, unnecessary or excessive movement) and passive/withdrawn (staring, lack of participation, dozing off) behaviours (e.g. rock, 2005). what constitutes on-task behaviour is inherently linked to the instructional strategy, which defines how students can carry out activities. in our study, on-task behaviour was operationalised as 'displaying goal-directed and task-appropriate behaviours', showing a clear link to the type of task set by the teacher. whilst working in pairs, conversation with a partner is part of the activity, whereas in independent work it is not. particularly for students in the primary school age certain task types are more conducive to high levels of on-task behaviour, such as allowing for talking and inviting collaboration. it has also been suggested that certain instruction formats are easier for teachers to supervise. godwin et al. (2013) found that individual work and whole-group instruction with children seated at their desks were negatively correlated with on-task behaviour (r=-0.018 and r=-0.113 respectively), and these formats accounted for almost 37% of instruction time. in contrast, paired work was positively correlated (r=0.032), yet students spent only approximately 18% of their time working in pairs. their findings are consistent with studies showing positive effects of group-work on primary school students’ self-regulation (dignath, buettner and langfeldt, 2008). hence, to investigate the association between observed and experienced engagement of students, it is necessary to take account of the teacher’s instructional strategy. recently, this was confirmed by pöysä et al. (2019), who found that the quality of classroom organisation was associated with students' self-reported behavioural and cognitive engagement. however, the observations were carried out in 20-minute segments at the classroom level, rather than the situational level. there is a need for studies comparing observed situation-specific behavioural engagement and self-reported engagement with learning in ordinary lessons in school. 1.3 self-reported learning experiences following schmitz (schmitz, 2006; schmitz and skinner, 1993) there is a current surge in intraindividual (process) educational research due to user-friendly self-report instruments in handheld computers. data are collected in experience sampling and ecological momentary assessment studies. there is a growing body of intraindividual studies focusing on situation-specific engagement (martin et al., 2015; malmberg and martin, 2019; patall et al., 2017; 2016; pöysä et al., 2019; shernoff et al., 2016). in the present study, participants completed situation-specific ratings on their cognitive and emotional engagement, as well as competence belief, task difficulty, and emotional state. current process approaches (schmitz, 2006; hamaker, 2012) emphasize the advantages of investigating processes in real-time, through multiple self-reports as these are less prone to retrospection bias and are more contextually relevant than cross sectional surveys (wilhelm, perrez and pawlik, 2012). in our study students reported at the beginning, middle, and end of each lesson during one week’s time, which enabled us to get snapshots of their perceptions of the task at hand (task difficulty), their current knowledge or understanding of the topic being studied (competence belief), how much thought they put into the task (cognitive engagement), whether they liked the subject or not (emotional engagement), and their emotions (positive and negative affect). 1.3.1 self-reported engagement school engagement is understood as a tripartite construct, its sub-components (cognitive, behavioural and emotional engagement) associated with each other (fredricks et al., 2004; wang, willett and eccles, 2011). low levels in any one of the three engagement domains has been shown to relate to unsuccessful outcomes in school, making all three – and their interaction – of interest to researchers and educators. the interplay between behavioural and cognitive engagement with learning (e.g., effort exertion, task-focus), competence beliefs (e.g., how good a student thinks he or she is at school, in a particular subject or at a task), and emotional engagement (e.g., subject liking), is posed in several theoretical models, for example, self-determination (ryan and deci, 2000), engagement (fredricks et al., 2004; roorda et al., 2011), and self-regulation frameworks (boekaerts and corno, 2005). cognition and cognitive activity are widely assumed to influence behaviour (martin, 2007). it is therefore not surprising that cognitive and behavioural engagement are related constructs. martin (2007) found that in 12,237 secondary school students (aged 12-18) adaptive cognitions and behaviours correlated (r=0.78), as did maladaptive cognitions and behaviours (r=0.68). taking this into account, wang et al. (2011) propose a model of school engagement including three second-order constructs: behavioural, emotional, and cognitive engagement. using data from 1,103 ethnically diverse american middle-school students (8th grade), they found a correlation of r=0.70 between cognitive and behavioural engagement in their sample. 1.3.2 competence beliefs and task difficulty although competence beliefs and task difficulty are related to each other, they are distinct. in the literature, competence belief is referred to as self-concept (marsh, 1990; shavelson, hubner, & stanton, 1976), self-efficacy (bandura, 1997), agency beliefs (little, 1998), and control beliefs (skinner, zimmer-gembeck and connell, 1998); converging around an individual’s sense of agency (bandura, 2008), their self-perceived capacity to fulfil a goal. for example, self-concept is positively and strongly related to academic performance (self-concept: d=0.43; hattie, 2009), the association being stronger when matched within a domain (valentine, dubois and cooper, 2004). the situational equivalents of competence beliefs have been termed mastery experiences (bandura, 1997), and competence beliefs (malmberg et al., 2013a; tsai et al., 2008). students gauge their competence beliefs based on their history of successes and failures, providing a basis for evaluating whether subsequent tasks are deemed difficult and whether they have what it takes to succeed. situation-specific competence belief and task difficulty have previously been found to be negatively and differentially associated (malmberg et al., 2013a). moreover, situation-specific competence belief was positively and differentially associated with situation-specific effort exertion (malmberg et al., 2013a). on this basis, we expected higher levels of on-task behaviour to be associated with higher levels of self-reported competence, and lower levels of self-reported difficulty. 1.3.3 subject enjoyment and learning behaviour the enjoyment students experience in schoolwork is related to their behaviour in school. martin (2007) found that school enjoyment correlated with adaptive academic cognitions as well as adaptive academic behaviours (r=0.74 and r=0.64 respectively). and pietarinen, soini and pyhältö (2014) describe the way higher levels of emotional engagement in primary school students contribute to greater cognitive and behavioural engagement, and subsequently to higher achievement. likewise, hospel, galand and janosz (2016) identified moderate correlations between behavioural engagement a multi-faceted construct comprising both on-task and off-task behaviours – and enjoyment, and specifically between enjoyment and classroom participation; arguably the strongest indicator of on-task behaviour (r=0.39 and r=0.43 respectively). in an earlier study, den brok et al. (2005) reported positive associations between the interpersonal closeness teachers fostered in class and their students’ enjoyment of the subject taught, as well as students’ self-reported effort during lessons. finally, wang et al. (2011) report a correlation of 0.89 between emotional and behavioural engagement. these results indicate that enjoyment/emotional engagement is an important factor to include when investigating behavioural engagement. some studies have included enjoyment as a positive emotion in momentary sampling, along with other achievement emotions (e.g. hospel et al., 2016). in this study, as observations took place during a range of subjects, enjoyment was expected to vary mostly between lessons, and have limited variability within lessons. thus, as we postulated that enjoyment is related to – and varies between – the subjects students are studying and not so much within lessons, we separated enjoyment from the other emotions in the questionnaire, since we expected those to vary to a greater extent within lessons. 1.3.4 emotions and learning emotion states in the classroom affect not only pupils’ psychological well-being, but also their cognitive, motivational, and regulatory processes involved in learning and achievement (pekrun, 2006; goetz et al., 2006). a subset of emotions, called achievement emotions, can be defined as emotions related to achievement activities or achievement outcomes (pekrun, 2006). the circumplex model of emotions (watson and tellegen, 1985) categorises emotions along the dimensions of valence (pleasant vs. unpleasant) and activation (activating vs. deactivating), making it possible to distinguish four broad groups of emotions: positive activating, (e.g., enthusiasm, pride), positive deactivating (e.g., relief, relaxation, nostalgia), negative activating (e.g., anxiety, anger, and shame), and negative deactivating (e.g., boredom, hopelessness). emotions are relevant to learning and behaviour, as they can be task-promoting or task-inhibiting (newton, 2013; inkinen et al., 2013). hospel et al. (2016) found that different emotions relate to different types of classroom behaviour; behavioural engagement as a global measure (encompassing participation, following instructions, absenteeism, withdrawal, and disruptive behaviours) was positively related to positive emotions, and negatively to negative emotions. each of the emotions examined in their study related to each of the five sub-domains of behavioural engagement in different ways. for example, anger related positively to all off-task behaviours (absenteeism, withdrawal, and disruptive behaviours), whereas sadness was only significantly related to absenteeism and withdrawal, but not disruptive behaviours (hospel et al., 2016). in our study we focus on on-task behaviours, rather than off-task behaviours, which can be seen as a combination of participation and following instructions. hospel et al. (2016) found significant and moderate positive correlations with participation as well as following instructions for positive activating emotions (interest, hope), and weak to moderate negative correlations for negative activating emotions (anger, anxiety) and participation. negative deactivating emotions (boredom, sadness) were weakly to moderately negatively correlated with participation. no positive deactivating emotions were included in their measures. unlike moods and personality traits, which are relatively much longer lasting, emotion states fluctuate throughout the day; they can last from a few seconds to a few hours (newton, 2013). besides being comparatively stable, traits are assumed to be mostly person-specific, whereas states are eventor situation-specific (hamaker, nesselroade and molenaar, 2007). thus, pupils’ emotions at the start of a lesson may not be the same as at the end of the lesson, depending on the events and people encountered in the lesson. these fluctuations in emotion states in turn contribute to variations in (task) behaviour. moreover, in the present study we focus on students’ achievement emotions and on-task behaviour across a range of different academic domains. goetz et al. (2016) found that emotions are related differently to different school subjects; for example, they found moderate positive correlations between enjoyment and english, but no significant correlation for french. thus, achievement emotions are related to on-task behaviour, and they vary between subjects. this highlights the importance of accounting for emotional state during learning, when examining on-task behaviour in the classroom. 2. method 2.1 sample and procedure data were collected in three classrooms from three different primary schools in southeast england (uk). a total of 62 students completed a modified version of the learning experience questionnaire (malmberg et al., 2013a) 1-3 times per lesson (at the beginning, after 20 minutes, and at the end, unless the lesson ended within five minutes of the previous questionnaire). this continued for two-to-four lessons per day (between 9:50 and 16:00), for five school days (one calendar week). a sub-sample of six students per class were observed. target students were selected by the class teacher, to provide an even split of boys and girls where possible, and include a range of attainment levels as assessed by the class teacher (under-achieving, average achievement, and high-achieving compared to age-related expectation). mean age of the target students was 10.1 years (sd=0.44, range 9.2-10.6), further participant characteristics are provided in table 1. the 18 target students did not differ from the 46 non-observed students with regard to gender (2[1]=0.03; p=0.87), age (t[58]= -0.93; p=0.36), or within-classroom standardized academic performance (t[60]=0.51; p=0.61). these target students were observed in the same order, with five seconds in between. thus, each of the target students was observed every 30 seconds, up to a maximum of 80 observations per 40-minute lesson. the 18 students were observed 14,994 times in total, and provided 631 self-reports, in a total of 56 lessons. observations averaged 66.05 events per lesson (sd=19.16, range=15-80, nobs=14,994), nested in an average of 12.61 lessons per student (sd=2.75, range=9-18). students provided on average 2.48 reports per lesson (sd=0.72, range=1-3), on average 6.74 reports per day (sd=2.43, range=2-12), totalling 35.1 reports per student (sd=12.6, range=19-52). teachers reported on students' academic performance, relative to age-related expectation, in each subject covered. missing data occurred only for subject-specific attainment (3.1%), and in these cases generic attainment was used instead. ethical approval was provided by the departmental research ethics committee at the oxford university department for education. parents provided informed written consent. table 1 target participant characteristics 2.2 measures 2.2.1 observations the observer (the first author) looked up each 5 seconds to observe the engagement of each target student in turn, giving a cycle of 30 sec. after a brief observation ‘snapshot’, the type of engagement was recorded into four categories: (1) focussed on-task behaviour, (2) non-disruptive off-task behaviour, (3) off-task disruptive behaviour, and (4) absent. for the current analyses we collapsed the two off-task categories into one, giving a binary on-task variable (0=off-task, 1=on-task). ‘absent’ ratings – when a student was not in the classroom at the time their observation turn was due – were discarded for the current analyses. students were rated on-task if they displayed goal-directed behaviour to carrying out the task as instructed by the teacher. if their behaviour was not task-appropriate or goal-directed they were rated 'off-task'. if they displayed goal-directed and task-appropriate actions at the same time as inappropriate actions (e.g. yawning whilst writing) and the inappropriate actions did not interfere with task completion, they were rated 'on-task'. where the behaviour could not definitively be appointed as on-task or off-task it was rated as 'other'. this category also included actions such as getting a drink or going to the toilet, since these are biological needs, even though they can be used as task-avoiding behaviours. at each observation point the classroom organisation was recorded, coded into six categories (1) teacher-led whole class instruction, (2) individual student work, (3) student pair or small-group work, (4) student-teacher one-to-one interaction, (5) assessment, and (6) ”other” (see table 2). table 2. instructional activities and engagement note: six target students per class were observed every 30 sec, between two to four lessons per day, for five days. tasks coded as ‘teacher-led whole class instruction’ included time when pupils sat on the carpet or at desks, focused on the white board, television, or teacher at the front. typically, they were expected to listen, process the information presented to them, or answer questions verbally/in writing. individual tasks were those where the teacher had expressed the expectation that pupils work independently, without assistance from or discussion with other pupils. pupils would generally be expected to focus on the materials on their own desk, along with supporting resources on the white board or wall displays. any interaction with peers was coded as 'off-task'. if teachers allowed collaboration, and pupils were expected to work as a pair or group to produce one collective piece of work, this was coded as pair or small group work. if pupils were expected to produce one piece of work, but copied into each individual book, this was also considered pair or small group work. if pupils were allowed to consult each other, but were expected to produce their own, unique final piece of work, this was coded as individual work, but task-related discussion with peers was coded as 'on-task'. student-teacher one-to-one interaction was coded separately, due to the increased likelihood of being on-task during this interaction if it was a case of receiving additional instruction or help in task-completion. however, if a pupil was reprimanded for being off-task, or encouraged to get back on-task, and would not comply with the request, this was still coded as 'off-task' as instructions were not followed. assessment was coded separate from individual work to reflect the increased pressure on pupils to work in absolute silence, and to take into account the difference in motivation between day-to-day learning tasks and test-taking. finally, the 'other' category was utilised for situations where the whole class or an individual student was engaged in activities/scenarios without a specific learning outcome, yet with behavioural expectations. examples of this are waiting for resources to be handed out, transitions between tasks, tidying the classroom, or waiting for the teacher’s assistance with a task. as shown in table 1, students carried out individual work most of the time (38%) and interacted individually with the teacher the least (4%). students were relatively more engaged during interaction with the teacher (96%), and relatively less during whole-class teacher led instruction (76%). we carried out interrater agreement of the observations in another sample ( ϰ=0.80) (heemskerk et al., 2019). 2.2.2 self-reported learning experiences at the beginning, middle, and end of each 40-minute lesson, students completed the brief self-report questionnaire. the questionnaire included 14 items using five-point scales. cognitive engagement was measured with three items (”at the moment ... how much effort are you putting into completing your task?”, ”how focused are you on your task?” and ”how much effort are you putting into keeping focused on your task?”), using the following scale: 5 = very much, 4 = quite a lot, 3 = a bit, 2 = not really, 1 = not at all (internal consistency by lesson: mα = 0.93, sdα = 0.03; and by time-point: mcdonald's ωtime-point= 0.87) (mcdonald, 1999). the following single-item measures were included, all using the same five-point scale as above: task difficulty ”the task you are doing at the moment... how difficult is your task?”; competence belief ”how good are you at this task?”; and subject liking “how much do you like this subject?”. following (watson and tellegen, 1985) and pekrun (2006), emotions were measured with eight items (”how are you feeling at the moment?”). four items focused on positive emotions, two activating (i.e., alert, enthusiastic) and two deactivating (relaxed, and calm), and four on negative emotions, two activating (i.e., frustrated, angry) and two deactivating (i.e., bored, tired). the items were answered on five-point scales: 5 = very much, 4 = quite a lot, 3 = a bit, 2 = not really, 1 = not at all. factor analysis did not support a four-factor solution, but a two-factor solution fitted reasonably. one factor represented positive emotions (mα = 0.58, sdα = 0.09; ωtime-point= 0.61), and one negative emotions (mα=0.80, sdα=0.02; ωtime-point =0.75). higher values for all our measures indicated more engagement, more difficulty, higher competence, and more positive and more negative emotions respectively. teachers reported on students' performance in relation to age-related expectations in all subjects on five-point scales (1 = well below expectation, 2 = below expectation, 3 = at age-related expectation, 4 = above expectation, 5 = well above expectation). the teacher reports were standardized within each class (m=0, sd=1). as subject-specific reports were available, we were able to link students' academic performance with each subject they did during the observed lessons, with the exception of 3.1% of observations, where a generic attainment score was used. students’ teacher-reported school-subject-specific performance is based on students’ prior performance (i.e., earlier in the school year) and hence would be treated as a covariate rather than learning outcome. academic performance was also found to be associated with students’ weekly-average situation-specific ratings of competence (r=0.46), effort exertion (r=0.15) and task difficulty (r= -0.20) in malmberg et al. (2013a). 2.2.3 analytical procedures we carried out all analyses using multilevel logistic regression models (merlo et al., 2006; moineddin, matheson and glazier, 2007; mood, 2010; rozi et al., 2017) in the brms r-package (bürkner, 2017; 2018). we specified a series of three-level logistic regression models in which time-points (t) were nested within lessons (l), nested in students (s). as our focus is on associations (see table 3) between engagement in lessons (30 sec intervals), we centred all self-reported predictors within lessons (brincks et al., 2017). r code is provided in the supplementary materials (s1). table 3. correlations between predictor variables. note: *** significant at p<0.001; ** significant at p<0.01; * significant at p<0.05. we first specified a variance component model (model 1) in which we partitioned the variance into between-lesson variance and between-student variance (equation 1). multilevel logistic regression models have no variance at the lowest level (merlo et al., 2006); there is no random effect for the time-points (tls). the value of the lowest level is fixed at π^2/3=3.29. in model 2 we added the within-lesson centred predictor (-1 = first 10 min, 0 = mid 20 min, 1 = last up to 10 min) of time (timetls (timels) (equation 2). in model 3 we added task-type, using teacher-led group-instruction as baseline and the other types as dummy-coded predictors: 'ind' = individual student work, 'pair' = students working in pairs or small groups, 'tea' = student-teacher one-to-one interaction, 'ass' = assessment, and 'oth' = other (equation 3). in model 4 we added self-reported learning experiences, centred within lessons: 'cogn' = cognitive engagement, 'diff' = task difficulty, 'comp' = competence belief, 'like' = emotional engagement, 'pos' = positive affect, 'neg' = negative affect (equation 4). in model 5 we added teacher-reported academic performance, 'zperf', in each school-subject (equation 5). the bayesian technique estimates the probability of a parameter given the data (p(θ | data)), rather than the probability of the data given the model (p(data | model)), as in null-hypothesis significance testing. it is suitable for situations when maximum likelihood might be underpowered, as it does not rely on large-sample theory (zitzmann et al., 2016). we specified uninformative priors for all parameters to mimic maximum likelihood estimates (zitzmann et al., 2016). all models were specified running four chains of 2000 iterations (of which 1000 were warm-up). models converged well, indicated by r ̂ values <1.01 and efficient sample sizes were more than 100 times the number of chains (https://mc-stan.org/misc/warnings.html). estimated posterior distributions and convergence of chains are presented in the supplementary materials (s2). 3. results 3.1 research aim 1: the variation of engagement within and between lessons, and between students. in the baseline model, as shown in table 4, there was more variance between students than between lessons, and the most variance within lessons. in the second model we can see that observed engagement decreased during lessons (b= -0.21, credibility interval [-0.30, -0.17], odds ratio (or)=0.79). 3.2 research aim 2: the relationship between observed engagement and instructional activities. the variables entered into models 3-5 were first checked for multicollinearity. no variables caused concern as none correlated at >0.6 (field, 2013) (see table 3). compared to direct instruction (teacher-led instruction), students were more engaged during individual work (b=0.37, [0.26, 0.48], or=1.45), pair-work (b=0.29, [0.11, 0.48], or=1.34), when interacting one-to-one with their teacher (b=2.55, [2.10, 3.04], or=12.81), during assessments (b=0.36, [0.10, 0.61], or=1.45), and other activities (b=1.17, [0.97, 1.37], or=3.22). 3.3 research aim 3: the relationship between observed engagement and self-reported learning experiences. in the fourth model, students were observed to be more behaviourally engaged when they experienced that they were more cognitively engaged than on average during the lesson (b=0.18, [0.03, 0.33], or=1.20). they were also more behaviourally engaged when they felt less competent than on average during the lesson (b= -0.20, [-0.31, -0.08], or=0.83). task difficulty, emotional engagement, affect, and teacher-reported prior academic performance did not predict observed engagement (model 5). for interpretation, estimated marginal means are presented in figure 1. the distribution of onand off-task observations across the predictor variable is visible in the black dots along the ‘0’ and ‘1’ values of the y-axis, as on-task is a binary variable and only can be rated 0 or 1. each predictor has a relevant distribution of values along the x-axis, and the regression line indicates the probability (value on the y-axis) of the observation being ‘on-task’, given the value of the predictor on the x-axis. the shaded area around the regression line shows the credibility interval. a narrower shaded area indicates greater certainty of the estimate. posterior distribution plots and chain plots for model 5 can be found in the supplementary materials (s2).  table 4. effects of time, learning activities, learning experiences, and performance on behavioural engagement note: posterior estimates not including zero in the credibility intervals indicated in bold. all estimates are from brms 2.10.0 (2019-08-27) figure 1. estimated marginal means for fixed effects. all estimates are from brms 2.10.0 (2019-08-27) 4. discussion using bayesian estimation, we investigated how observed engagement (1) varied within and between lessons, and between students, (2) was associated with instructional activities and (3) was associated with self-reported learning experiences (controlling for academic performance). most variance was found within lessons, between students, and least between lessons. observed student engagement decreased from the beginning to the end of lessons, and associated with the type of learning activity in the class (one-to-one student-teacher interaction > “other” activities > individual work > assessment > pair/group work > teacher-led whole class instruction). observed engagement was higher when students reported higher cognitive engagement and lower competence belief. 4.1 variation in observed engagement within and between lessons, and between students investigating situation-specific motivation in schools, martin et al. (2015) identified substantial variation in student engagement within days and between students. in line with their findings, we found considerable variation at the situation-level (within lessons in our study), and between students. direct comparison of the level of variation is not meaningful, as our data structure diverged by further differentiating within-lesson and between-lesson observations and self-reports, whilst not including day-to-day variation. however, whereas martin et al. found the greatest degree of variation at the between-student level, we found greater situation-specific variation. our study might uniquely suggest that the more fine-tuned the manner in which we observe behaviour and ask for self-reports, the more variance we find at this level (30 second intervals for observations, 15-20 minutes for self-reports). the greater situational variance in on-task behaviour in our sample may be related to the age of the participants; our sample consisted of primary rather than secondary school students, and it is suggested that with increased age and metacognitive strategy use, the dependence of engagement on contextual factors decreases (fredricks et al., 2004). 4.2 instruction format and observed engagement in their study of instruction format in relation to on-task behaviour in fourto eight-year-old children, godwin et al. (2013) found that students were least on-task during individual work and whole-class instruction, and most on-task during small-group work and testing. their categories did not include ‘one-to-one support’ or ‘other’. our results also indicated higher levels of on-task behaviour during testing, and lower levels during whole-class instruction. however, we found that students were more on-task during individual work, and less during pair/group work. this may relate to the age of participants, as our sample was older than in the study by godwin et al., and they might be better able to cope with working independently. godwin et al. (2013) suggest that certain instructional formats present greater possibility for students go off-task, as they are more difficult for teachers to supervise. this would be a logical explanation for the relatively high level of on-task behaviour in one-to-one instructional settings in the present study. in the present study, 33% of observations involved teacher-led, whole-class instruction, and 38% involved independent work. only 8% of the observed time was spent working in pairs or small groups. due to the small proportion of time spent working in small groups, our results for this category of task should be interpreted with caution. similar to the findings of godwin et al. (2013), overall associations between observed engagement (on-task) and instructional activities were small. 4.3 self-reported learning experiences and observed engagement within-lesson increases in cognitive engagement were expected to be positively related to observed behavioural engagement, and this was confirmed by the data. competence belief unexpectedly related negatively to behavioural engagement. this may be an indication that pupils felt a greater need to pay attention and participate in classroom activities at times when they perceived their competence at the task in hand to be lower than average. we found no significant relationship between on-task behaviour and within-lesson changes in task difficulty, emotional engagement, or emotion states. as expected, emotional engagement may vary more strongly between lessons than within lesson. emotion states may have been too difficult for children of this age to accurately identify and rate, despite efforts to only include emotions in the questionnaire that children in this age category would be familiar with. another possibility is that the differential effects of activating and deactivating emotions of the same valence may be masking one another by including only two emotion state factors (positive and negative affect) in our analyses. however, factor analysis did not support a four-factor structure as suggested in the circumplex model (positive activating, positive deactivating, negative activating, negative deactivating (watson and tellegen, 1985)) in our data, but rather a two-factor structure (positive and negative affect). due to the repeated measurements taken, it was decided to keep the questionnaire as short as possible. this, in combination with the limited vocabulary of the age group for naming emotions, led to only two emotions being included for each quadrant of the circumplex model. it may be that, with three or four emotions per quadrant, a four-factor model would have been supported by the data, and these in turn might have been related to behavioural engagement in this age group. overall associations between observed engagement (on-task) and self-reported learning experiences were small. 4.4 limitations the present study had some limitations. first, while we had rich data at the situational level (n = 14,994) the sample size at the student-level was small (n = 18). although three diverse schools were included in the current sample (state-funded, voluntary-aided, and private), further sociodemographic data for participants was not collected. further replications of our findings in larger samples are needed, including the collection of individual sociodemographic data. second, while we had rich data on instructional activity, we did not investigate the quality of the student-teacher interaction during the instructional activity. such future studies could show whether quality of student-teacher interaction varies across instructional activities. for example, emotional support could be higher in one-to-one student-teacher interaction (la paro and pianta, 2000). third, while we did not carry out interrater agreement assessments with the current version of the observation instrument, we did so in heemskerk et al. (2019). as cognitive engagement cannot be measured by observations, we relied on self-reports for this. finally, while we focussed on exploration of variance within lessons, further analyses could focus on differences between lessons. 4.5 conclusion observed engagement varied more greatly within lessons than between lessons, and diminished over time during lessons. on-task behaviour was predicted by higher cognitive engagement and instructional activities. as compared to teacher-led instruction, engagement was higher during individual tasks, teacher-supported tasks, and assessments. overall self-reported (cognitive) and observed (behavioural) engagement within lessons converged, supporting their use in intraindividual research. keypoints student engagement varies substantially within lessons and between students. more variance in behaviour is found at the situational level with more fine-tuned observation and self-report methods. student engagement is related to instructional activities, with the highest engagement during one-to-one instruction, assessments, and independent work. educators must carefully consider their use of whole-group instructions, as this accounted for 33% of observed time and was associated with the lowest level of engagement. observed engagement is positively related to self-reported engagement and negatively related to competence belief. references bandura, a. (1997) self-efficacy: the exercise of control, new york: w.h. freeman. bandura, a. (2008) toward an agentic theory of the self. in: marsh, h.w., craven, r.g. and mcinerney, d.m. (eds) self-processes, learning, and enabling human potential: dynamic new approaches. charlotte, nc: information age publishing, pp.15-49. barros, r.m., silver, e.j. and stein, r.e. (2009) 'school recess and group classroom behavior'. pediatrics 123(2) pp.431-436. doi: 10.1542/peds.2007-2825. boekaerts, m. and corno, l. (2005) 'self‐regulation in the classroom: a perspective on assessment and intervention'. applied psychology 54(2) pp.199-231. doi: 10.1111/j.1464-0597.2005.00205.x. brincks, a.m., enders, c.k., llabre, m.m., bulotsky-shearer, r.j., prado, g. and feaster, d.j. (2017) 'centering predictor variables in three-level contextual models'. multivariate behavioral research 52(2) pp.149-163. doi: 10.1080/00273171.2016.1256753. bürkner, p.-c. (2017) 'brms: an r package for bayesian multilevel models using stan'. journal of statistical software 80(1) pp.1-28. doi: 10.18637/jss.v080.i01. bürkner, p.-c. (2018) 'advanced bayesian multilevel modeling with the r package brms'. the r journal 10(1). doi: 10.32614/rj-2018-017. carlson, j.a., engelberg, j.k., cain, k.l., conway, t.l., mignano, a.m., bonilla, e.a., geremia, c. and sallis, j.f. (2015) 'implementing classroom physical activity breaks: associations with student physical activity and classroom behavior'. preventive medicine 81 pp.67-72. doi: 10.1016/j.ypmed.2015.08.006. collie, r.j. and martin, a.j. (2019) motivation and engagement in learning. oxford research encyclopedia of education. new york: oxford university press. den brok, p., levy, j., brekelmans, m. and wubbels, t. (2005) 'the effect of teacher interpersonal behaviour on students' subject-specific motivation'. the journal of classroom interaction 40. dignath, c., buettner, g. and langfeldt, h.-p. (2008) 'how can primary school students learn self-regulated learning strategies most effectively?'. educational research review 3(2) pp.101-129. doi: 10.1016/j.edurev.2008.02.003. field, a. (2013) discovering statistics using ibm spss statistics, london: sage. fredricks, j.a., blumenfeld, p.c. and paris, a.h. (2004) 'school engagement: potential of the concept, state of the evidence'. review of educational research 74(1) pp.59-109. doi: 10.3102/00346543074001059. godwin, k.e., almeda, m., petroccia, m., baker, r.s. and fisher, a.v. (2013) 'classroom activities and off-task behavior in elementary school children', 35th annual conference of the cognitive science society. austin, tx, cognitive science society. pp.2428-2433. godwin, k.e. and fisher, a.v. (2011) 'allocation of attention in classroom environments: consequences for learning', 33rd annual conference of the cognitive science society. austin, tx, cognitive science society. pp.2806-2811. goetz, t., pekrun, r., hall, n. and haag, l. (2006) 'academic emotions from a social-cognitive perspective: antecedents and domain specificity of students' affect in the context of latin instruction'. british journal of educational psychology 76 pp.289-308. doi: 10.1348/000709905x42860. goetz, t., sticca, f., pekrun, r., murayama, k. and elliot, a.j. (2016) 'intraindividual relations between achievement goals and discrete achievement emotions: an experience sampling approach'. learning and instruction 41(c) pp.115-125. doi: 10.1016/j.learninstruc.2015.10.007. grieco, l.a., jowers, e.m. and bartholomew, j.b. (2009) 'physically active academic lessons and time on task: the moderating effect of body mass index'. medicine and science in sports and exercise 41(10) pp.1921-1926. doi: 10.1249/mss.0b013e3181a61495. hamaker, e.l. (2012) why researchers should think "within-person": a pragmatic rationale. in: mehl, m.r. and conner, t.s. (eds) handbook of research methods for studying daily life. london: guilford press, pp.43-61. hamaker, e.l., nesselroade, j.r. and molenaar, p.c.m. (2007) 'the integrated trait–state model'. journal of research in personality 41(2) pp.295-315. doi: 10.1016/j.jrp.2006.04.003. hattie, j. (2009) visible learning, abingdon: routledge. hospel, v., galand, b. and janosz, m. (2016) 'multidimensionality of behavioural engagement: empirical support and implications'. international journal of educational research 77 pp.37-49. doi: https://doi.org/10.1016/j.ijer.2016.02.007 . hox, j.j., van de schoot, r. and matthijsse, s. (2012) 'how few countries will do? comparative survey analysis from a bayesian perspective'. survey research methods 6(2) pp.87. doi: 10.18148/srm/2012.v6i2.5033. inkinen, m., lonka, k., hakkarainen, k., muukkonen, h., litmanen, t. and salmela-aro, k. (2013) 'the interface between core affects and the challenge– skill relationship'. journal of happiness studies 15 pp.891-913. doi: 10.1007/s10902-013-9455-6. jarrett, o.s., maxwell, d.m., dickerson, c., hoge, p., davies, g. and yetley, a. (1998) 'impact of recess on classroom behavior: group effects and individual differences'. the journal of educational research 92(2) pp.121-126. doi: 10.1080/00220679809597584. la paro, k.m. and pianta, r.c. (2000) 'predicting children's competence in the early school years: a meta-analytic review'. review of educational research 70(4) pp.443-484. doi: 10.3102/00346543070004443. little, t.d. (1998) sociocultural influences on the development of children’s action-control beliefs. in: heckhausen, j. and dweck, c.s. (eds) motivation and self-regulation across the life span. cambridge, usa: cambridge university press, pp.281-315. malmberg, l.-e., hagger, h., burn, k., mutton, t. and colls, h. (2010) 'observed classroom quality during teacher education and two years of professional practice'. journal of educational psychology 102(4) pp.916-932. doi: 10.1037/a0020920. malmberg, l.-e. and martin, a.j. (2019) 'processes of students’ effort exertion, competence beliefs and motivation: cyclic and dynamic effects of learning experiences within school days and school subjects'. contemporary educational psychology 58 pp.299-309. doi: 10.1016/j.cedpsych.2019.03.013. malmberg, l.-e., walls, t.a., martin, a.j., little, t.d. and lim, w.h.t. (2013a) 'primary school students' learning experiences of, and self-beliefs about competence, effort, and difficulty: random effects models'. learning and individual differences 28 pp.54-65. doi: https://doi.org/10.1016/j.lindif.2013.09.007 . malmberg, l.-e., woolgar, c. and martin, a.j. (2013b) 'quality of measurement of the learning experience questionnaire for personal digital assistants'. international journal of quantitative research in education 1(3) pp.275-296. martin, a.j. (2007) 'examining a multidimensional model of student motivation and engagement using a construct validation approach'. the british journal of educational psychology 77(2) pp.413-440. doi: 10.1348/000709906x118036. martin, a.j., ginns, p. and papworth, b. (2017) 'motivation and engagement: same or different? does it matter?'. learning and individual differences 55 pp.150-162. doi: 10.1016/j.lindif.2017.03.013. martin, a.j., papworth, b., ginns, p., malmberg, l.-e., collie, r.j. and calvo, r.a. (2015) 'real-time motivation and engagement during a month at school: every moment of every day for every student matters'. learning and individual differences 38(supplement c) pp.26-35. doi: https://doi.org/10.1016/j.lindif.2015.01.014 . mcdonald, r.p. (1999) test theory: a unified treatment, hillsdale: erlbaum. merlo, j., chaix, b., ohlsson, h., beckman, a., johnell, k., hjerpe, p., råstam, l. and larsen, k. (2006) 'a brief conceptual tutorial of multilevel analysis in social epidemiology: using measures of clustering in multilevel logistic regression to investigate contextual phenomena'. journal of epidemiology and community health 60(4) pp.290. doi: 10.1136/jech.2004.029454. moineddin, r., matheson, f. and glazier, r.h. (2007) 'a simulation study of sample size for multilevel logistic regression models'. bmc med. res. methodol. 7(1). doi: 10.1186/1471-2288-7-34. mood, c. (2010) 'logistic regression: why we cannot do what we think we can do, and what we can do about it'. european sociological review 26(1) pp.67-82. doi: 10.1093/esr/jcp006. newton, d.p. (2013) 'moods, emotions and creative thinking: a framework for teaching'. thinking skills and creativity 8(supplement c) pp.34-44. doi: https://doi.org/10.1016/j.tsc.2012.05.006 . patall, e.a., vasquez, a.c., steingut, r.r., trimble, s.s. and pituch, k.a. (2016) 'daily interest, engagement, and autonomy support in the high school science classroom'. contemporary educational psychology 46 pp.180-194. doi: 10.1016/j.cedpsych.2016.06.002. patall, e.a., vasquez, a.c., steingut, r.r., trimble, s.s. and pituch, k.a. (2017) 'supporting and thwarting autonomy in the high school science classroom'. cognition and instruction 35(4) pp.337-362. doi: 10.1080/07370008.2017.1358722. pekrun, r. (2006) 'the control-value theory of achievement emotions: assumptions, corollaries, and implications for educational research and practice'. educational psychology review 18(4) pp.315-341. doi: 10.1007/s10648-006-9029-9. pellegrini, a.d. and davis, p. (1993) 'relations between children's playground and classroom behaviour'. british journal of educational psychology 63(1) pp.88-95. doi: 10.1111/j.2044-8279.1993.tb01043.x. pellegrini, a.d., huberty, p.d. and jones, i. (1995) 'the effects of recess timing on children's playground and classroom behaviors'. american educational research journal 32(4) pp.845-864. doi: 10.3102/00028312032004845. pietarinen, j., soini, t. and pyhältö, k. (2014) 'students’ emotional and cognitive engagement as the determinants of well-being and achievement in school'. international journal of educational research 67 pp.40-51. doi: https://doi.org/10.1016/j.ijer.2014.05.001 . pöysä, s., vasalampi, k., muotka, j., lerkkanen, m.-k., poikkeus, a.-m. and nurmi, j.-e. (2017) 'variation in situation-specific engagement among lower secondary school students'. learning and instruction 53(supplement c) pp.64-73. doi: https://doi.org/10.1016/j.learninstruc.2017.07.007 . pöysä, s., vasalampi, k., muotka, j., lerkkanen, m.k., poikkeus, a.m. and nurmi, j.e. (2019) 'teacher–student interaction and lower secondary school students’ situational engagement'. british journal of educational psychology 89(2) pp.374-392. doi: 10.1111/bjep.12244. praetorius, a.-k., lenske, g. and helmke, a. (2012) 'observer ratings of instructional quality: do they fulfill what they promise?'. learning and instruction 22(6) pp.387-400. doi: 10.1016/j.learninstruc.2012.03.002. rock, m.l. (2005) 'use of strategic self-monitoring to enhance academic engagement, productivity, and accuracy of students with and without exceptionalities'. journal of positive behavior interventions 7(1) pp.3-17. doi: 10.1177/10983007050070010201. roorda, d.l., koomen, h.m.y., spilt, j.l. and oort, f.j. (2011) 'the influence of affective teacher–student relationships on students’ school engagement and achievement'. review of educational research 81(4) pp.493-529. doi: 10.3102/0034654311421793. rozi, s., mahmud, s., lancaster, g., hadden, w. and pappas, g. (2017) 'multilevel modeling of binary outcomes with three-level complex health survey data'. open journal of epidemiology 7(1) pp.17. doi: 10.4236/ojepi.2017.71004. ryan, r. and deci, e. (2000) 'self-determination theory and the facilitation of intrinsic motivation, social development, and well-being'. american psychologist 55(1) pp.68-78. doi: 10.1037/0003-066x.55.1.68. schmitz, b. (2006) 'advantages of studying processes in educational research'. learning and instruction 16(5) pp.433-449. doi: 10.1016/j.learninstruc.2006.09.004. schmitz, b. and skinner, e. (1993) 'perceived control, effort, and academic performance: interindividual, intraindividual, and multivariate time-series analyses'. journal of personality and social psychology 64(6) pp.1010. doi: 10.1037/0022-3514.64.6.1010. shernoff, d.j., kelly, s., tonks, s.m., anderson, b., cavanagh, r.f., sinha, s. and abdi, b. (2016) 'student engagement as a function of environmental complexity in high school classrooms'. learning and instruction 43(c) pp.52-60. doi: 10.1016/j.learninstruc.2015.12.003. skinner, e.a., zimmer-gembeck, m.j. and connell, j.p. (1998) 'individual differences and the development of perceived control introduction and overview'. monographs of the society for research in child development 63(254). doi: 10.2307/1166220. tsai, y.-m., kunter, m., lüdtke, o. and trautwein, u. (2008) day-to-day variation in competence beliefs: how autonomy support predicts young adolescents’ felt competence. in: marsh, h.w., craven, r.g. and mcinerney, d.m. (eds) self-processes, learning, and enabling human potential: dynamic new approaches. charlotte, nc: information age pub., pp.119-120. valentine, j.c., dubois, d.l. and cooper, h. (2004) 'the relation between self-beliefs and academic achievement: a meta-analytic review'. educational psychologist 39(2) pp.111-133. doi: 10.1207/s15326985ep3902_3. wang, m.-t., willett, j.b. and eccles, j.s. (2011) 'the assessment of school engagement: examining dimensionality and measurement invariance by gender and race/ethnicity'. journal of school psychology 49(4) pp.465-480. doi: 10.1016/j.jsp.2011.04.001. watson, a., timperio, a., brown, h., hinkley, t. and hesketh, k.d. (2019) 'associations between organised sport participation and classroom behaviour outcomes among primary school-aged children'. plos one 14(1). doi: 10.1371/journal.pone.0209354. watson, d. and tellegen, a. (1985) 'toward a consensual structure of mood'. psychological bulletin 98(2) pp.219. doi: 10.1037/0033-2909.98.2.219. wilhelm, p., perrez, m. and pawlik, k. (2012) conducting research in daily life. in: mehl, m.r. and conner, t.s. (eds) handbook of research methods for studying daily life. london: guilford press, pp.62-86. wubbels, t., brekelmans, m., den brok, p., wijsman, l., mainhard, t. and van tartwijk, j. (2015) teacher-student relationships and classroom management. in: emmer, e.t. and sabornie, e.j. (eds) handbook of classroom management. second edition ed. new york: routledge, pp.363-386. zakszeski, b.n., hojnoski, r.l. and wood, b.k. (2017) 'considerations for time sampling interval durations in the measurement of young children's classroom engagement'. topics in early childhood special education 37(1) pp.42-53. doi: 10.1177/0271121416659054. zitzmann, s., lüdtke, o., robitzsch, a. and marsh, h.w. (2016) 'a bayesian approach for estimating multilevel latent contextual models'. structural equation modeling: a multidisciplinary journal 23(5) pp.661-679. doi: 10.1080/10705511.2016.1207179. microsoft word vangrieken et al_publication.docx frontline learning research vol.5 no. 4 (2017) 1 41 issn 2295-3159 corresponding author: katrien vangrieken, centre for research on professional learning & development, corporate training and lifelong learning, university of leuven, dekenstraat 2 (box 3772), 3000 leuven, belgium. email: katrien.vangrieken@kuleuven.be. doi: http://dx.doi.org/10.14786/flr.v5i4.297 group, team, or something in between? conceptualising and measuring team entitativity katrien vangrieken, anne boon, filip dochy, & eva kyndt university of leuven, belgium article received 9 march / revised 21 july / accepted 27 august / available online 16 october abstract the current gap between traditional team research and research focusing on non-strict teams or groups such as teacher teams hampers boundary-crossing investigations of and theorising on teamwork and collaboration. the main aim of this study includes bridging this gap by proposing a continuum-based team concept, describing the distinction between strict teams and mere collections of individuals as the degree of team entitativity. the concept entitativity is derived from social psychology and further developed and integrated in team research. based upon this concept and core team definitions, the defining features shaping teams’ degree of entitativity are determined: shared goals and responsibilities; cohesion (task cohesion and identification); and interdependence (task and outcome). furthermore, a questionnaire is developed to empirically grasp these features. the questionnaire is tested in two waves of data collection (n1=1,320; n2=731). based upon a combination of classical test theory analyses (exploratory and confirmatory factor analyses) and item response theory analyses, the questionnaire is developed. the final questionnaire consists of three factors: shared goals and cohesion, task interdependence, and outcome interdependence. further psychometric analyses include the investigation of validity, longitudinal measurement invariance, and test-retest reliability. this manuscript describes frontline research by: (1) developing a new conceptualisation transcending the variety in terminology and definitions used in team and group research by creating a shared language and opening up team research to non-strict teams and (2) combining two methodological traditions regarding questionnaire development and validation (classical test theory and item response theory). keywords: team; team entitativity; questionnaire development; classical test theory; item response theory vangrieken et al 2 | f l r 1. introduction as collaboration and teamwork appear to be indispensable in the current society and organisational life, ample research has focused on the investigation of teams or groups and their processes in order to gain insight into how employee professional development and performance can be improved (mathieu, maynard, rapp, & gilson, 2008; sundström, mcintyre, halfhill, & richards, 2000). research on collaborative processes shows a large diversity in types and structures of collaboration and the terminology that is used to describe them. these include for example group, team, network, community of practice, and professional (learning) community. the latter are especially prevalent in the educational context as the educational counterpart of the learning organisation concept proposed by senge (1990). while the traditional and archetypically defined team concept is more common in research in industry and organisations, educational research tends to focus on more flexible concepts such as communities. both seem to be mostly separated streams of research: research in schools rarely includes frameworks that were developed and elaborately investigated in traditional team research and the latter tends to neglect teacher teams and collaboration. team research tends to focus on a specific type of team, meeting strict criteria included in team definitions such as the one suggested by cohen and bailey (1997) (e.g., decuyper, dochy, & van den bossche, 2010; edmondson, 1999; mathieu, heffner, goodwin, salas, & cannon-bowers, 2000; salas, burke, & cannon-bowers, 2000). they define a team as an interdependent collection of individuals, sharing responsibility for the outcomes of the team, who both see themselves and are seen by others as a social entity (i.e., a team) that is part of a larger social system (such as a business unit or organisation). however, in practice not all teams meet these definitions and they are not always teams in the strict sense of the word (such as teacher teams). despite the increasing importance of team and collaborative work in schools, team researchers tend to keep away from this context because teacher teams often do not fit traditional team criteria, which hampers a straightforward application of team theories and frameworks. at the same time, research on the topic of teacher collaboration tends to investigate this phenomenon more or less in general rather than to focus on more deep-level quality features and dynamics of such collaborative endeavours. moreover, it is prone to lacking both conceptual and methodological rigour: the terminology and conceptualisations of different types of teacher collaboration are often ill-defined (vangrieken, dochy, raes, & kyndt, 2015) and literature on teacher collaboration tends to be criticised for its lack of methodological rigour (crow & pounder, 2000). in line with the recommendation of wageman, gardner, and mortensen (2012), these gaps indicate the need for a flexible definition of a real team. wageman et al. (2012) stress the value of being able to transfer insights and approaches developed in team research to new kinds of collaboration. in order to make this possible, there is a need for a new and continuum-based team concept. this study aims to tackle these gaps by (1) introducing a continuum-based conception of a team and (2) making this continuum measurable. addressing the first aim, this study introduces the concept team entitativity to bridge the conceptual gap between real teams and other types of collaboration that are not strict teams. team entitativity refers to the teamness of a team or the degree to which a group of individuals meets the criteria of a team. the conception and definition of team entitativity is based upon a combination of two mostly separated streams of research (research on group perception and research on team dynamics). crossing the boundaries between team research and research on teacher collaboration by introducing a more flexible team concept could be very beneficial for both fields, expanding the scope of traditional team research and strengthening the rigour of research on teacher collaboration. in order to reach the second aim, a questionnaire was developed. this makes it possible to measure the conceptual gap and to empirically bridge research on teams and other groupings which was not done before. in a first step towards bridging these fields empirically, the psychometric quality of the team entitativity instrument was assessed in the context of teacher teams in secondary education in flanders (belgium). vangrieken et al 3 | f l r 2. the grey area of teacher teams as suggested earlier, teacher teams are often excluded from traditional team research because they mostly do not meet the strict theoretical team definitions (smith, 2009; vangrieken, dochy, raes, & kyndt, 2013). elements such as interdependence appear to be difficult to achieve in teacher collaboration and teacher teams tend to operate more as aggregates of individuals bonded by social ties, focusing on psychological safety and social cohesion (moolenaar, 2010; ohlsson, 2013; smith, 2009; vangrieken, dochy, & raes, 2016). in these cases, teacher teams operate as co-acting groups in which members believe to work in a team but actually work in ways that do not match the basic notion of teamwork (lyubovnikova et al., 2015). the core of their job, teaching, is mostly performed individually in their classroom. ‘egg-crate schools’ characterised by teacher isolation in classrooms and the accompanying culture of individual work, together with teachers’ preference for individually exercised autonomy may withhold a collaborative, teaming culture to take shape in education (gajda & koliba, 2008; lortie, 1975, in westheimer, 2008; somech, 2008). related to the complexity of realising teacher teamwork in its fullest form, there is a grey area surrounding teacher teams. the construct does not appear as uniform but it takes different forms and shapes. vangrieken et al. (2013, 2015) present a comprehensive teacher team typology attempting to provide an overview of these different possible forms. as such, teacher teams can have different tasks that range from a focus on policy-making and school-level decisions (e.g., governance/management, innovation and school reform) to teamwork focused on classroom education, teacher professional development, and problem-solving (e.g., instruction, pedagogy, teacher learning). these different tasks entail a different level of depth of collaboration: while a focus on a mere material or practical task involves only superficial discussions, teams working on improving instruction or teacher learning entail more deep-level discussions of teachers’ practice and pedagogical beliefs or motives. furthermore, teacher teams can be organised in different ways: disciplinary or interdisciplinary, within or cross grade levels, and long term or temporary. in practice, this often creates a complicated structure of various teams and sub-teams in schools. for example, in (flemish) secondary schools teachers are mostly organised in disciplinary subject departments or teams (cross grades). especially in bigger schools, these subject departments sometimes consist of different grade-level subject teams (teachers who teach the same subject in the same grade). however, in smaller schools – and especially in primary education – the whole teaching staff may be organised in one overarching team. this focus on whole-school collaboration can for example also be found in literature on professional (learning) communities (e.g., birenbaum, kimron, & shilton, 2001; leonard & leonard, 2001). a final dimension derives from the fact that the construct teacher team presents itself as a continuum, ranging from completely individualised teacher work to teamwork in its fullest form (smith, 2009; vangrieken et al., 2015). hence the team entitativity construct is especially valuable in this context to capture this collaborative continuum. 3. conceptualising team entitativity the first part of this study focuses on defining team entitativity. starting from the origins of the entitativity construct in social psychology research (campbell, 1958), the concept is further developed and integrated in team research. 3.1 origins of the construct ‘entitativity’ ‘entitativity’ resides from gestalt theory and social psychology and was originally introduced by campbell (1958) who defined it as “the degree of being entitative. the degree of having the nature of an entity, of having real existence” (campbell, 1958, p. 17). he focused on the analysis of social aggregates as entities vangrieken et al 4 | f l r and on explaining why some groups are considered real groups and others mere aggregates of individuals. the construct received resonance in different studies on group perception yet still remained unclear in definition, leading to a variety in the way entitativity is interpreted, defined, manipulated, and measured (hamilton, sherman, & castelli, 2002). in order to categorise this variety in the use and interpretation of entitativity, a distinction along two axes can be made: on the one hand a distinction between an outsider and an insider perspective and on the other hand a focus on essence or agency. the first axis defines a distinction proposed by yzerbyt, corneille, and estrada (2001). an outsider perspective focuses on non-group members’ (outsiders) perceptions of the degree of entitativity that a collection of individuals has. originally, entitativity focused on outsider perceptions as it was coined by campbell (1958) to investigate the way people use available information cues to assess the extent to which an aggregate of individuals has the quality of being an entity or a group. these information cues were derived from the principles of gestalt theory and include proximity, similarity, common fate, and pregnance (good continuation or good figure). hence the focus was on visual cues that can be observed by outsiders who are not part of the entity. as opposed to the outsider perspective, studies including an insider perspective focus on properties of the group, as they are experienced and perceived by the members themselves (carpenter et al., 2008; jans, postmes, & van der zee, 2011). these are not necessarily observable by non-group members. the extent to which these properties are perceived to be present is presumed to indicate the group’s degree of entitativity. for example, carpenter et al. (2008) ascribed the fact that some groups are more group-like than others to the perceived degree of cohesion, unity, interdependence, conformity, and fit between group members. from an insiders’ perspective, jans et al. (2011) criticised the proposition that similarity a core aspect of entitativity in an outsider perspective is a key predictor of perceived entitativity and argue that within-group differences do not necessarily hinder a sense of cohesion and unity. the issue of whether or not similarity is a key predictor of entitativity brings us to the next axis defining a distinction between essence and agency (brewer, hong, & li, 2004; rutchik, hamilton, & sack, 2008). essence theory perceives entitative groups as being characterised by a fixed and inherent shared essence (brewer et al., 2004). the criterion of similarity is the source of groupness, as group perception is based on members sharing certain properties (jans et al., 2011). the focus is on innate fixed personality traits of group members and this perception can mostly be found in social cognitive research on stereotypes (e.g., crawford, sherman, & hamilton, 2002; yzerbyt et al., 2001). opposed to this static meaning of entitativity, agency theory proposes a dynamic view in which entitativity is based on perceived common goals and intentions of groups (brewer et al., 2004). the nature of the group and its boundaries are no longer perceived as static but malleable and changeable over time. dynamic properties (e.g., shared goals and coordination) are to a lesser extent formal organisational structures; instead they refer to psychological states and relationships among group members (brewer et al., 2004). in this dynamic perspective, similarity of group members is not needed to perceive a group as an entity (jans et al., 2011). 3.2 team entitativity research on entitativity rarely focuses on work groups, but often investigates large social entities, such as for example race or nationalities (e.g., castano, yzerbyt, paladino, & sacchi, 2002; dasgupti, banaji, & abelson, 1999; yzerbyt et al., 2001). some studies investigate a broad array of entities, ranging from a family, a committee, to people at a bus stop or even as broad as women as a category (e.g., crawford & salaman, 2012; hamilton et al., 2002; svirydzenka, sani, & bennett, 2010). this study further develops the construct for application in research on work groups and teams and as such introduces team entitativity. 3.2.1 teams-in-theory: defining a team in order to clearly conceptualise team entitativity as the degree to which an aggregate of individuals is a team, it needs to be clear what it means to be a team. however, in team literature (a lot of) different definitions are used. katzenbach and smith (1993) define a team as “a small number of people with complementary skills who are committed to a common purpose, set of performance goals, and approach for vangrieken et al 5 | f l r which they hold themselves mutually accountable” (p. 45). a similar and integrative definition was proposed by cohen and bailey (1997): a team is a collection of individuals who are interdependent in their tasks, who share responsibility for outcomes, who see themselves and who are seen by others as an intact social entity embedded in one or more larger social systems (for example, business unit or the corporation), and who manage their relationships across organizational boundaries (p. 241). similarly, hackman (2012) defined teams as: intact social systems whose members work together to achieve a common purpose. they have clear boundaries that distinguish members from non-members. they work interdependently to generate a product for which members have collective, rather than individual, accountability. and they at least have moderate stability, which gives members time to learn how to work well together (p. 437). integrating these definitions, a team can be defined as “a distinguishable collection of individuals, who identify themselves as a team and interact as a team to reach certain shared goals for which they share responsibility and hold themselves mutually accountable. members are jointly committed to the common purpose and task and are interdependent in their tasks and outcomes” (vangrieken et al., 2015). often the terms team and group are used interchangeably, both in research and practice (lyubovnikova, west, dawson, & carter, 2015; west & lyubovnikova, 2012). however, in this study they are considered to be separate constructs. while they do partially share the same characteristics – they are delineated social units working in larger organisations – they are not the same concepts (main, 2007). the main difference between both includes the fact that the term group is more broadly used and defined: all teams are groups but not all groups can be called teams, as they do not always meet all criteria included in team definitions (van den bossche, gijselaers, segers, & kirschner, 2006; vangrieken et al., 2015). while a team needs to meet certain criteria included in the definition provided above in order to actually be a real team, groups do not. the concept group in itself can have a variety of meanings and is used to refer to different entities that strongly vary in their properties (hamilton et al., 2002). groups can be broadly defined as collections of individuals that share a common social categorisation and identity (raes, kyndt, decuyper, van den bossche, & dochy, 2015). the term is often used in definitions of various social categorisations, such as communities of practice (cop) or professional learning communities (plc). in these cases, the definitions add criteria that need to be met in order for a group to be called a cop or plc. in conclusion, the term group is mostly used as an umbrella term capturing different possible social categorisations. in this study, a team is defined as one specific subtype of a group meeting certain team-criteria. 3.2.2 towards a continuum-based team concept: team entitativity previous team research already started to challenge the universal use of the team concept, distinguishing real teams from for example co-acting groups (lyuobovnikova, west, dawson, & carter, 2015) or pseudo-teams (west & lyuobovnikova, 2012). however, these conceptions still refer to a black-and-white categorical interpretation: teams are real teams or pseudo-teams. however, as suggested by wageman et al. (2012), changes in the environment of teams (e.g., increasing virtual and globally dispersed work and collaboration) challenge the boundaries of the traditionally defined team. rather than using defining features as delineators to distinguish real teams from nonor pseudo-teams, it is valuable to investigate these features as dynamic constructs in their own right (wageman et al., 2012). hence, team entitativity is proposed here as a concept that is worth investigating as a separate variable. it is introduced as a dynamic characteristic of teams and how team members themselves perceive their team. in line with entitativity literature, teams are presented as a continuum. collections of individuals or groups possess a degree of team entitativity, which is dynamic in nature. a team’s perceived degree of team entitativity is dynamic and malleable, depending on for example time, team development, and contextual influences. this is in line with the conception of groups as social systems that are dynamically engaged with their contexts (hackman, 2012; wageman et al., 2012). hence, vangrieken et al 6 | f l r when positioning our conceptualisation of team entitativity on the essence – agency axis described earlier, it matches the latter perspective. when discussing the features that determine whether a group of people is a real team, different aspects are mentioned. these include (structural) interdependence, shared objectives, team reflexivity, boundedness (i.e., it is clear who is part of the team and who is not), and stability of membership (buljac, van woerkom, & van wijngaarden, 2013; lyubovnikova et al., 2015; wageman et al., 2012; west & lyubovnikova, 2012). in a changing society and environment for teams, the meaning and importance of these features may have changed. for example, in globally dispersed teams the boundaries of team membership may not be that clear and the fact that people tend to be part of multiple teams and the existence of complicated multiteam systems consisting of different teams and subteams may challenge the criterion of stable team membership. moreover, supposedly structural team features such as structural interdependence (i.e., residing from the design of work itself) are actually dynamic phenomena that are increasingly in the hands of the team members themselves (wageman et al., 2012). hence, an important characteristic of the teamness of groups or teams includes that it strongly resides from choices and perceptions of the members themselves. west and lyubovnikova (2012) also suggested asking the individual team members to what extent the core characteristics of a team are present. thus, regarding the insider – outsider axis of entitativity, the focus here is on an insider perspective starting from how members themselves perceive their team. furthermore, because of our focus on work teams rather than social entities (such as families) our definition of team entitativity is task-focused. the latter can be seen as a core characteristic of work teams (salas, bowers, & cannon-bowers, 1995). work teams can be categorised as common identity groups that focus on attachment to the group as a whole, based upon identification with the group, its goal, and its purpose (sassenberg, 2002). the task focus and the fact that members are brought together to complete shared goals distinguishes them from for example groups of friends. the latter are described as common bond groups that are focused on interpersonal attachment between group members based upon positive attitudes towards other members (prentice, miller, & lightdale, 1994; sassenberg, 2002). hence, the focus here is not on these groups that focus on (social) interpersonal attraction but on task-focused work teams in which attachment is based on a common goal and purpose. previous research found mixed results concerning the effects of certain social aspects (e.g., social cohesion) on team functioning, not always finding beneficial outcomes (e.g., van den bossche et al., 2006; lima, 2001). this supports our focus on task-related rather than social aspects of teams. 3.2.3 defining features the degree of team entitativity a collection of individuals possesses, is determined by the degree to which different defining features of a team are present. the defining features included in the conceptualisation put forward here are derived from the aforementioned team definitions and the most prominent criteria rising from entitativity literature (campbell, 1958; carpenter et al., 2008; hamilton et al., 2002). three main features of team entitativity can be distinguished, acting as forces binding team members together in a task-focused team: (a) having shared goals and responsibilities, (b) cohesion, and (c) interdependence among team members. shared goals and responsibilities. entitativity literature uses common goals and outcomes, shared purpose, and common fate as criteria or cues for the degree of entitativity (campbell, 1958; carpenter et al., 2008; crawford & salaman, 2012; hamilton & sherman, 1996; lickel et al., 2000). furthermore, to be defined as a team, groups of individuals have to share common goals (katzenbach & smith, 1993; kozlowski & bell, 2003). moreover, highly entitative teams are characterised by a shared responsibility of the members for the goals and outcomes to be reached and for the required team process (cohen & bailey, 1997; katzenbach & smith, 1993). members of these teams hold themselves mutually accountable for the common goals (katzenbach & smith, 1993). thus in case of high levels of team entitativity, members have to reach a shared outcome in the team and have predefined team goals. the team is held mutually accountable for the attainment of these goals or outcomes and the team process by which they are reached. cohesion. the degree of cohesion is a second defining feature of entitativity (e.g., carpenter et al., vangrieken et al 7 | f l r 2008; cavazza, pagliaro, & guidetti, 2014; spencer-rodgers, hamilton, & sherman, 2007). sometimes entitativity is even reduced to cohesion (hamilton, 2007). group cohesion is defined as the result of all the forces that bind members to each other and to their team and act upon members to stay in the group (festinger, 1950; guzzo & shea, 1992). festinger (1950) suggests three components of cohesion: (a) attraction to the group (i.e., interpersonal attraction or social cohesion), (b) task commitment (i.e., task cohesion, extent to which individual goals are shared with or enabled by the group), and (c) group pride (i.e., members experiencing positive affect from being associated with the group) (mcleod & von treuer, 2013). the first component relates to the interpersonal conception of group cohesion – based on mutual positive attitudes among members that is the origin of attachment in common bond groups such as groups of friends (sassenberg, 2002). as the focus of this conceptualisation is on work teams and task-related aspects, this interpersonal or social cohesion was not included. what binds members of a common identity group – such as a work group is not this interpersonal attraction but a focus on identification with the goal and purpose of the group (sassenberg, 2002). hence, the focus here is on the two other components suggested by festinger (1950): task cohesion (cf. task commitment) and identification with the team (cf. group pride). task cohesion includes the degree of shared commitment among members to achieve a goal that can be reached solely by the collective efforts of the group (van den bossche et al., 2006). carron (1982) defined cohesion as “a dynamic process that is reflected in the tendency for a group to stick together and remain united in the pursuit of its goals and objectives” (p. 124). hence, in the case of a high degree of team entitativity, teams are characterised by a shared commitment to the tasks of the team, the latter acting as the force that holds team members together. the second aspect of cohesion, identification with the team, includes on the one hand the degree of identification of the members with the team and on the other hand the degree to which team members feel that (membership of) the team is important for their job. both aspects are seen as important features determining the degree of entitativity (crawford & salaman, 2012; lickel et al., 2000; meneses, ortega, navarro, & de quijano, 2008). similarly, high levels of team identification and perceptions of being part of the team are important features of a team (smith, 2009). team identification can be defined as “awareness and attractions towards an interacting group of interdependent members, by self-identified members of that group” (bouas & arrow, 1996, p. 155-156). as the focus is on the task, the importance of the team to its members is defined here as the perceived degree of importance of team membership and team tasks or goals to members with regard to their job. thus, in the case of a high degree of team entitativity, team members experience a sense of affinity with their team and perceive membership as important for their job functioning. interdependence. a third defining feature includes the degree to which team members are interdependent. different authors investigating entitativity describe interdependence as an important criterion (carpenter et al., 2008; mcgrath, 1984). moreover, working interdependently is a key characteristic of teams (cohen & bailey, 1997; kozlowski & bell, 2003). in short, interdependence means that team members do not only depend on their own actions but also on the actions of other members in order to perform their tasks or reach certain outcomes (weick, 1976). wageman (1995) argues that interdependence can originate from different sources: task inputs (e.g., skill distribution, resources), work processes, the way goals are defined, and the way performances are rewarded. van der vegt, emans, and van de vliert (1998) distinguish between task and outcome interdependence in an attempt to capture these different sources of interdependence. the common task or goal of teams can vary and can be of a different nature. while the focus here is on work teams (i.e., teams that are brought together to perform certain work-related tasks), people can also be collaborating with a focus on a specific aim for learning. the latter are mostly referred to as communities of practice (cop) or professional learning communities (plc) and are more often found in the educational context compared to business (dufour & eaker, 1998; lave & wenger, 1991; vangrieken, meredith, packer, & kyndt, 2017). the nature and meaning of interdependence may be different in work teams compared to learning teams because of the different focus of these teams. the main aim of work teams includes collaborating on and performing one or more tasks. while learning behaviours that may occur in these teams are essential for successful teamwork and task performance, learning is not their sole aim. conversely, learning teams such as plcs present a framework for organising professional development initiatives in collaboration. the vangrieken et al 8 | f l r conceptualisation of task and outcome interdependence described below mainly applies to work teams. task interdependence can be defined as “the interconnections between tasks such that the performance of one definite piece of work depends on the completion of other definite pieces of work” (van der vegt et al., 1998; p. 127). it emanates from the tasks in the team referring to the extent to which interaction, coordination, and collective action of team members are needed to complete tasks and members must rely on their fellow team members to perform their tasks effectively (guzzo and shea, 1992; saavedra, earley, & van dyne, 1993; van der vegt, emans, & van de vliert, 1999; wageman, 1995). thus in highly entitative teams, members strongly feel that they need each other to perform their own tasks and that communication and coordination are needed to realise the team task. van der vegt et al. (1998) describe outcome interdependence as “the extent to which team members believe that their personal benefits and costs depend on successful goal attainment by other team members” (p. 130). it derives from “the degree to which significant consequences of the work – such as goal attainments and tangible rewards – are contingent on collective performance” (wageman, 1995, p.146). hence, outcome interdependence is high when all members profit from the outstanding performance of fellow team members (van der vegt et al., 1998). in essence, it refers to the degree to which the team is characterised by team or individual reward structures (whether individual outcomes depend on joint or personal performance) (beersma, homan, van kleef, & de dreu, 2013). in short, outcome interdependence refers to the degree to which team members perceive that how well other team members perform has a positive influence on them and the degree to which the team is characterised by team reward structures. 3.2.4 individual perceptions versus convergence as suggested above, our conceptualisation of team entitativity is focused on individual team members’ perceptions of dynamic team properties that are mostly not objectively observable but rise from subjective experiences and perceptions. this entails that individual members’ perceptions of their teams’ degree of entitativity may differ. as the focus here is on perceptions of a team level construct, one can wonder what it means for the team when members have a different perception of their team’s degree of team entitativity. the question is whether a collection of individuals can be described as highly entitative when members’ perceptions vary. when individuals’ perceptions vary, the average perception of all members may not be a good representation of a group’s degree of team entitativity. averaging out individual members’ perceptions ignores the existence of individual variability. hence, we propose the degree of convergence or agreement on team entitativity within the group, the extent to which there is a shared perception, to also indicate a team’s degree of team entitativity. this is in line with the fact that the development of a shared vision was put forward as an important characteristic of teams (salas et al., 2000). thus, it can be assumed that members of a highly entitative team, strongly meeting such team characteristics, develop a shared vision on their team’s degree of entitativity. kozlowski and chao (2012) and kozlowski (2015) discuss how such sharedness comes to exist and how teamlevel phenomena and concepts (such as team entitativity) emerge. they argue that different forms of emergence can take place. the composition forms, based on convergent dynamics, are especially applicable here. translated to the concept of team entitativity, it suggests that as individuals interact in their team, their perception of team entitativity becomes more and more homogeneous and converges. hence the extent to which individual team members’ perceptions of entitativity (of the defining features) are shared, may also indicate the degree of teamness the team possesses. vangrieken et al 9 | f l r 4. towards a measurement of team entitativity 4.1 research aim a clear operationalisation of entitativity lacks in literature and there is no standard set of items that is used for measurement (lickel et al., 2000). moreover, as suggested above, our conceptualisation of team entitativity differs from the traditional entitativity construct. hence, besides defining team entitativity as a theoretical construct, the present study aimed at developing a measurement instrument based upon the conceptualisation presented above. as mentioned earlier, the development of this instrument was applied to a specific type of teams-inpractice: teacher teams in secondary education in flanders (belgium). more specifically, the focus was on subject groups. these are structural units within the school that gather teachers who teach the same or closely related subjects (i.e., disciplinary teams). they collaborate concerning subject-related matters, such as the curriculum and student evaluation. it was opted to focus on subject groups because they are meaningful collaborative units within each school that are organised to collaborate on core teaching issues (i.e., curricular issues as well as didactical-pedagogical matters). however, the content of their collaboration can vary strongly across groups, some focusing on making superficial practical arrangements, others discussing more profound pedagogical and didactical matters related to the subject. 4.2 procedure this study consists of three steps. first, in the questionnaire development phase, an item pool was generated based upon the conceptualisation of team entitativity. building from this item pool, the measurement instrument was designed and presented to teachers for feedback. second, in the test phase, this instrument was tested in a large-scale quantitative study. the questionnaire was distributed among teachers in secondary schools in flanders (in dutch). in the analyses, two methodological approaches were combined: classical test theory (ctt) to assess the underlying factor structure of the questionnaire and item response theory (irt) for item-level analyses. adding irt analyses to the traditionally performed ctt analyses has several benefits: (a) besides providing information on how the scale functions, irt analyses provide detailed information on how the items function; (b) they provide a subject-independent assessment of the items and the scale, resulting in comparable results across groups from the same population; (c) they provide insight in the range of measurement precision of the items and the scale, demonstrating the range of the underlying construct in which the scale is most reliable and thus best at discriminating among individuals. as the irt techniques used here require unidimensionality of the questionnaire under investigation, the factor structure first needs to be assessed. in a next step, item-level analyses were performed using irt, providing more detailed information on the item-level and range of effectiveness of the scales. further analyses were performed to investigate the psychometric quality of the instrument. finally, the degree to which team members have a shared perception of their team’s degree of team entitativity was investigated. in the third phase, the resulting 15-item instrument was retested to assess longitudinal measurement invariance, test-retest reliability, and predictive validity. in order to assess the discriminant and predictive validity of the team entitativity instrument, its relationship to two other well-established constructs was investigated. in order to test the discriminant validity of the measure, the distinction between team entitativity and psychological safety was assessed. the latter is often investigated in team research and includes feeling safe to express oneself without fear of damaging one’s self-image, status, or career (kahn, 1990). in environments that are perceived as psychologically safe, team members are not punished for asking for help or admitting mistakes and are prepared to present new ideas, ask questions, and express concerns because – due to the safe environment – they are not afraid to make a mistake in doing so (edmondson, 1999; 2008). similar to team entitativity, it can be described as a belief of the team members about the interpersonal context of the team (vangrieken et al., 2016). psychological safety has been demonstrated to be related to positive team processes and outcomes, such as team learning behaviours (e.g., edmondson, 1999; van den bossche et al., 2006; vangrieken et al., 2016). vangrieken et al 10 | f l r furthermore, in order to test the predictive validity of the measure, the relationship between team entitativity and team work engagement was assessed. team work engagement (twe) is “a positive, fulfilling, work-related and shared psychological state characterised by team work vigor, dedication and absorption which emerges from the interaction and shared experiences of the members of a work team” (torrente, salanova, llorens, & schaufeli, 2012, p. 107). it contains three aspects: vigour, dedication, and absorption. team vigour is defined as “high levels of energy and an expression of willingness to invest effort in work and persistence in the face of difficulties” (costa, passos, & bakker, 2012, p. 5). team dedication refers to “a shared strong involvement in work and an expression of a sense of significance, enthusiasm, inspiration, pride and challenge while doing so” (costa et al., 2012, p. 6). finally, team absorption includes a “shared focused attention on work, whereby team members experience and express difficulties detaching themselves from work” (costa et al., 2012, p. 6). it is hypothesised that team entitativity may function as a team resource that, based upon the rationale of the job demands-resources (jd-r) model (demerouti, bakker, nachreiner, & schaufeli, 2001; schaufeli & bakker, 2004) fosters team work engagement. the motivational process included in the jd-r model describes the rise of (team) work engagement as triggered by job resources. these include physical, psychological, social, or organisational job aspects that can help to achieve work goals, reduce job demands, and encourage personal growth and development (demerouti, et al., 2001; schaufeli & bakker, 2004). team entitativity can be described as such a (team) resource that fosters twe. this is in line with the model of albrecht (2012) and torrente, salanova, llorens, and schaufeli (2012), suggesting that team resources such as team climate and teamwork are positively related to (team) engagement. 5. questionnaire development for each of the defining features of team entitativity, literature was searched for existing questionnaires. in this first step, items were collected from a variety of existing (sub)scales (table 1). from this item pool those items that matched the conceptualisation of one of the defining features were selected and adjusted to fit our conceptualisation. in a next step, additional items were composed to cover those aspects of team entitativity that were not represented in the item pool. this led to the first version of the team entitativity questionnaire including 27 items. given our research design combining a ctt and irt approach, two considerations were important in this questionnaire development phase. first of all, in order to be able to select the best functioning items, it is advised to start off with a sufficiently large selection of items. hence, we started with a broad selection of 27 items, aiming to end up with a more parsimonious instrument with approximately half of the original items. secondly, based upon the irt framework it is important to include a broad spectrum of items with varying levels of the construct (i.e., location of the item), referring to the degree of team entitativity you need to have in order to endorse the item. this indicates that items should be situated on different locations of the underlying team entitativity continuum. while some items should only be endorsed by teachers part of highly entitative teams, other items should be easier to endorse in order to distinguish between different teams situated at the lower end of the continuum. when measuring psychological constructs, it is more complicated to judge the levels or location of items compared to when assessing for example math skills (in the latter case the level or location refers to the difficulty of items). attempting to meet this requirement, (a) a variety of items assessing the same feature were included and (b) where possible some items were formulated more strongly – including more entitativity requirements – than others (e.g., “to me, the objective of this team is unclear” vs. “i think that, as a team, we have a clear and shared objective”; “the team members collaborate to achieve a shared mission” vs. “in my opinion, within this team, we have a clear and shared mission, which we all try to achieve as a team”; “for my job as a teacher, i consider it important to be part of this team” vs. “in order to be a good teacher, i have to collaborate with the other members of my team”). this version was presented to 42 teachers to evaluate the items. their remarks, feedback, and suggestions for improvement were used in order to assess whether the items were meaningful for the context of teachers and to further refine the formulation. overall, all items were perceived as relevant. the teachers made suggestions to reduce the complexity of some items by shortening them and adjusting the terminology. vangrieken et al 11 | f l r as these were all minor adjustments, the questionnaire remained rather context-independent. the resulting instrument includes five items assessing shared goals and responsibilities, 10 items measuring cohesion (four assessing task cohesion, and six assessing identification), and 12 items aimed to capture interdependence (of which seven measure task interdependence and five outcome interdependence) (see appendix for a list of the items). table 1 original questionnaires authors measured construct original number of items number of questions adapted aubé and rousseau (2005) team goal commitment 3 1 campion, medsker, and higgs (1993) task interdependence interdependent feedback and rewards 3 3 2 1 carless and de paola (2000) task cohesion 4 3 henry, arrow, and carini (1999) behavioral group identification 4 2 kiggundu (1983) initiated and received task interdependence 15 and 13 see van der vegt, emans, and van de vliert (1999) pearce and gregersen (1991) experienced reciprocal interdependence 5 1 van der vegt and bunderson (2005) collective team identification 4 2 van der vegt, emans, and van de vliert (1999) task and outcome interdependence 5 and 8 5 van der vegt and janssen (2003) perceived task interdependence 5 3 van der vegt, emans, and van de vliert (1998) initiated and received task interdependence outcome interdependence 4 and 4 6 2 and 2 1 wageman, hackman, and lehman (2005) compelling direction 2 2 6. testing the questionnaire the instrument was tested in a quantitative study consisting of two waves of data collection among 37 secondary schools in flanders (a region of belgium with approximately 6,500,000 inhabitants). in the first wave, 1,677 teachers completed the questionnaire. as one of our aims included the investigation of team members’ shared perception of team entitativity, it is important that a sufficient number of members per team completed the questionnaire. however, as suggested in section 2, teachers are often part of multiple subject teams. in order not to overload these teachers with multiple questionnaires, they were asked to fill out the questionnaire only once (for the team they are most strongly involved with). this choice was made to make sure that the research was practically feasible. the downside of this decision includes that this reduces the potential pool of participants when considering the team level, reducing response rates of some of the teams. furthermore, being very restrictive regarding the required response rate would probably magnify bias in the sample as it could be hypothesised that more cohesive teams with a higher degree of entitativity have a larger probability of reaching a high response rate. hence, in order to find a balance between practical feasibility and vangrieken et al 12 | f l r the need for sufficient variety in team entitativity on the one hand, and the need for a sufficiently high participation rate per team on the other hand, only data of teachers part of the subject teams of which at least 60 per cent (teams with less than 10 members) or 50 per cent (teams with 10 or more members) of the members completed the questionnaire, were retained for analyses. hence, the data of 1,320 teachers part of 227 subject groups in 35 different schools were used. in the second wave of data collection, 1,178 of the teachers participating in the first wave (70.24%) filled out the questionnaire. the same criteria regarding participation rate per team were used for our second wave, resulting in a sample of 731 teachers (nested in 139 subject groups from 31 different schools) suitable for analyses. 6.1 method 6.1.1 instrument the distributed team entitativity measure contained 27 dutch items. furthermore, a measure for psychological safety was included to assess discriminant validity. the dutch version of the seven-item psychological safety instrument of edmondson (1999) (e.g., van den bossche et al., 2006) was used and slightly adapted so the items were formulated from the individual’s point of view (e.g., “i feel that is difficult to ask other team members for help”, instead of “it is difficult to ask other team members for help”). finally, a measure for team work engagement (twe) was included in order to assess predictive validity of the team entitativity measure. twe was assessed with the short version of the utrecht work engagement scale (uwes) for teams (torrente et al., 2012). items of the team entitativity and psychological safety questionnaire were answered on a 6-point likert scale ranging from 1=‘completely disagree’ to 6=‘completely agree’. twe was measured using a 7-point likert scale ranging from 1=‘never’ to 7=‘always’. background information was asked about the individual teachers and their team. this included: gender, age, years of teaching experience, information on their appointment, the size of their team, and the number of years they have been part of the team. regarding teachers’ type of appointment, they were asked to indicate whether or not they are appointed ad interim (i.e., a substitute teacher with a temporary contract) and whether or not they have a permanent appointment in their school. 6.1.2 sample characteristics table 2 provides an overview of the sample characteristics. the majority of the respondents indicated to have structurally planned meetings with their team once or a few times per trimester (first wave=83.35%; second wave=84.54%). however, when asked about informal meetings and consultation with colleagues of their team, teachers reported higher frequency levels. most teachers reported to consult their team colleagues several times a month to several times a week (first wave=65.91%; second wave=67.44%), a substantial part of the teachers even reported daily consultations (first wave=13.33%; second wave=13.41%). vangrieken et al 13 | f l r table 2 sample characteristics first wave (n = 1,320) second wave (n = 731) gender female (%) 808 (61.21%) 442 (60.47%) male (%) 512 (38.79%) 289 (39.53%) age mean 41.13 40.47 min – max 21 – 62 21 – 62 sd 10.52 10.58 teaching experience mean 16.33 16.14 min – max 0 – 46 0 – 40 sd 10.56 10.60 type of appointment ad interim (%) 115 (8.71%) 69 (9.44%) permanent appointment (%) 1056 (80%) 574 (78.52%) team size mean 7.26 7.08 min – max 3 32 3 32 sd 4.37 4.44 years of membership mean 12.33 12.36 min – max 0 – 40 0 – 39 sd 9.48 9.88 6.1.3 analyses because the instrument assesses teachers’ perceptions of their team, a first step was investigating the degree of team level variance and associated appropriateness and feasibility of multilevel analyses. this was investigated by means of the intraclass correlation (icc(1) and icc(2)) coefficients and design effects of the items. the icc coefficients assess group reliability. icc(1) refers to the proportion of the total variance that can be explained by team membership (newman & sin, 2009). as stated by dyer, hanges, and hall (2005) multilevel analyses may provide minimal practical benefits and may be difficult or impossible to estimate when icc(1)s are smaller than .05. icc(1) values of the items ranged from .010 to .132 (m=.059), 12 items having an icc(1) value below .05, and design effects were all below 2 (m=1.286, min = 1.049, max = 1.636). the icc(2) coefficients indicate the reliability of the group means (bartko, 1976; newman & sin, 2009). icc(2) values of the items range from .13 to .53 (m=.35), indicating that aggregate group means are not reliable because of high levels of within-group variance. given the low icc(1) and icc(2) values and design effects, it was opted to perform analyses at the individual level. in order to take the clustering in our data into account, the “type is complex” command was used in the mplus analyses. while this does not specify a model on two levels (individual and team), it takes the clustering of the data into account when computing standard errors and chi-square tests of model fit. dimensionality. in order to assess the structure of the questionnaire, the sample of the first wave was split in two stratified subsamples, a development sample (n1= 662) and a validation sample (n2=658). stratification was based upon school on the one hand and team size on the other hand: all schools and team vangrieken et al 14 | f l r sizes are equally represented in both samples. the development sample was used to investigate the structure of the questionnaire by means of exploratory factor analyses (efa, maximum likelihood estimation with robust standard errors). oblique rotation (direct oblimin) was used because it was assumed that the different factors are related as they are derived from the same underlying construct. moreover, as suggested in the conceptualisation of team entitativity, the different features are – based on theory – assumed to be related. next, the identified structure was assessed in the validation sample by means of confirmatory factor analysis (cfa) and hierarchical cfa (maximum likelihood estimation with robust standard errors). internal consistency was examined using the cronbach’s alpha coefficient. in the analyses, the complex data structure (i.e., clustered) was taken into account, computing robust standard errors and chi-square statistics. item-level analyses. next, the full sample of the first wave (n=1,320) was used in order to assess the performance of the individual items within the derived structure using irt analyses. a graded response model was used, which estimates two parameters for each item: a discrimination parameter (slope) and a location parameter. these analyses were performed for each factor of the structure derived from efa and cfa analyses separately. final questionnaire. based upon the combination of the structural analyses and item-level analyses, the final questionnaire was composed and again assessed in the validation subsample of the first wave (n2=658) using cfa and internal consistency analyses. moreover, shared variance of the items within each factor and discriminant validity of the factors were assessed in line with the guidelines of fornell and larcker (1981). furthermore, in order to investigate the degree to which team members have a shared perception of their team’s degree of team entitativity, within-group agreement (rwg) scores were investigated (james, demaree, & wolf, 1984). finally, based upon the retest of the questionnaire, longitudinal measurement invariance and predictive validity of the questionnaire were assessed. the calculation of the inter-item correlations, kaiser-meyer-olkin (kmo) measure of sampling adequacy, and bartlett’s test of sphericity to check the suitability of our data for efa, was done in spss version 23. factor analyses and analyses investigating longitudinal measurement invariance and validity (discriminant and predictive) were performed in mplus version 7.4 using “type is complex” to take nesting of our data into account. the satorra-bentler correction was used when comparing the different models assessing longitudinal measurement invariance (satorra & bentler, 2010). irt analyses were performed using parscale. all other analyses were performed in r version 3.2.1 (r core team, 2015) using the package psych (revelle, 2012). 7. results 7.1 classical test theory: analysing the structure 7.1.1 exploring the structure first, inter-item correlations were assessed in the development sample (first wave) to detect redundant items that do not increase construct coverage of the scale. of each item pair demonstrating a correlation above .75, one item was omitted. both items of each of these sets of items were situated in the same defining feature of team entitativity. hence, all features remained sufficiently represented. the selection of these items was based upon conceptual reasons as the retained items more clearly covered the intended content. next, it was assessed whether our data were suitable for efa. sample size was sufficiently large (n= 662), the kaisermeyer-olkin (kmo) measure of sampling adequacy equalled .920 and the bartlett’s test of sphericity was significant (χ2=7823.62, df=253, p<.001). this indicates that the inter-correlation matrix contains enough common variance to make efa useful. efa were performed on the development sample (n=662) in order to investigate the structure of the questionnaire. based upon the eigenvalues (i.e., larger than one), the scree plot, and conceptual reasons, a vangrieken et al 15 | f l r three-factor solution proved to be most valid. moreover, all items that were included had to have a loading larger than .40 on one of the factors and were not allowed to have high cross-loadings. the latter means that items had to have relatively low loadings on other factors (below .30) and that the difference between an item’s loading on one factor and its loading on another factor was not allowed to be below .20. eight items not meeting these criteria were omitted. the resulting factor solution is presented in table 3. table 3 results exploratory factor analysis (oblique rotation) items factor 1 factor 2 factor 3 shared1 .625 .084 -.147 shared2 .815 -.026 -.022 shared3 .882 -.037 -.002 shared4 .721 .045 .075 taskcoh2 .809 .032 .024 taskcoh4 .866 -.081 .053 ident2 .716 .149 -.023 taskint2 .049 .573 .073 taskint3 -.036 .526 .197 taskint4 -.036 .736 -.002 taskint6 .008 .842 -.016 outint1 .045 .725 -.028 outint3 -.080 .024 .734 outint4 .049 -.011 .767 outint5 .102 .054 .631 estimation method: maximum likelihood with robust standard errors the first factor includes seven items referring to shared goals and responsibilities, and cohesion (task cohesion and identification) and is labelled shared goals and cohesion. the items intended to measure the importance of the team to its members (theoretically assumed to be part of the cohesion feature), are not represented in this solution. the second factor includes five items primarily referring to task interdependence. although this factor includes one item (outint1 in the appendix) originally intended to assess another defining feature outcome interdependence evaluation of its content demonstrated it to be closely related to task interdependence. the final factor includes three items referring to outcome interdependence. vangrieken et al 16 | f l r 7.1.2 confirming the structure the structure derived from efa was assessed in the validation sample (n=658) (first wave) by means of cfa. the ratio of sample size to number of items exceeded the ratio of 10:1, indicating that our data were suitable for cfa (hair, black, babin, anderson, and tatham, 2006). furthermore, the ratio of sample size to number of free parameters to be estimated in the model exceeded the required minimum of 10:1 (bentler & chou, 1987). cfa resulted in a good fit of the assumed model (χ2/df =2.85; df=87; cfi=.952; tli=.942; rmsea=.053 [90 % ci [.045; .061]]; srmr=.044) since these fit indices met the generally accepted norms for cfa (brown & cudeck, 1993; hu & bentler, 1999). table 4 presents an overview of the resulting cfa solution. table 4 results (hierarchical) confirmatory factor analysis item regression weight standard error standardised regression weight critical ratioa cronbach’s alpha shared goals and cohesion .91 shared1 1.000 b .591 b shared2 1.307 .095 .755 13.792 shared3 1.457 .105 .848 13.883 shared4 1.367 .114 .796 11.965 taskcoh2 1.483 .133 .832 11.156 taskcoh4 1.579 .149 .836 10.609 ident2 1.170 .091 .700 12.877 task interdependence .82 taskint2 1.000 b .739 b taskint3 .933 .065 .643 14.443 taskint4 .754 .083 .621 9.073 taskint6 .972 .066 .822 14.696 outint1 .789 .069 .662 11.488 outcome interdependence .81 outint3 1.000 b .673 b outint4 1.281 .079 .856 16.301 outint5 1.246 .080 .787 15.543 hierarchical cfa team entitativity .89 shared goals and cohesion 1.000 b .605 b task interdependence 1.828 .273 .831 6.694 outcome interdependence 1.284 .222 .577 5.787 estimation method: maximum likelihood with robust standard errors aall critical ratios: p<.001 bvalue fixed at 1.00 for model identification purpose, hence no standard error was computed to test the assumption that all three factors are indicators of the same underlying construct, a hierarchical cfa was performed. in this model, the same factor model as included in the regular cfa was vangrieken et al 17 | f l r tested with an additional latent factor encompassing the three other factors. again, fit indices confirmed an appropriate fit of the presumed factor model (χ2/df=2.85, df=87; cfi=.952; tli=.942; rmsea=.053 [90 % ci [.045; .061]]; srmr=.044). results can be found in table 4. 7.1.3 internal consistency all cronbach’s α values were sufficient and did not significantly increase if one of the items would be dropped (table 4). moreover, item-whole correlations were assessed to investigate whether the included items correlated sufficiently with the scale as a whole. it was opted to assess corrected item-whole correlations as these correct for item overlap and scale reliability (revelle, 2012). for shared goals and cohesion corrected item-whole correlations ranged from .61 to .84. for task interdependence these scores ranged from .62 to .81 and for outcome interdependence from .66 to .81. moreover, for the hierarchical team entitativity factor corrected item-whole correlations ranged from .42 to .75. 7.2 item response theory: item-level analyses the next step of our analyses included the application of irt analyses to assess the performance of the items included in the questionnaire. a two-parameter logistic model was used. the location parameter indicates the level of the construct: the larger the location of an item, the more of the measured construct a respondent must have in order to endorse the item (edelen & reeve, 2007). items with a different location provide information at different points of the underlying construct or trait (i.e., shared goals and cohesion; task interdependence; outcome interdependence). while items with a low location parameter are best at discriminating among people at lower trait levels, items with a high location parameter can best discriminate at higher trait levels (reise, ainsworth, & haviland, 2005). the mean of the location parameter is set to zero. the discrimination parameter refers to the extent to which the item is related to the underlying construct and how good the item is in discriminating among different individuals (at the level of the location parameter) (edelen & reeve, 2007; reise et al., 2005). as unidimensionality is a prerequisite to perform these analyses, irt analyses were performed for each factor separately. for each factor, following results will be discussed: (a) the test information curve (tic), demonstrating the degree of reliability or measurement precision of the scale at different levels of the underlying measured construct; (b) estimates of the parameters; and (c) the itemcharacteristic curves, demonstrating the relationship between someone’s response to an item and his or her level of the underlying measured construct. the tic (figures 1, 3, and 5) demonstrates the psychometric information of the test (i.e., its ability to differentiate among individuals) at each level of the underlying construct. this indicates the instrument’s measurement precision or how well it functions at the different trait levels. accordingly, the standard error of measurement (red curve in the graph) is inversely related to the tic: the higher the information level of the tic, the lower the standard error. the psychometric information provided by the test is thus analogous to test reliability in ctt. however, while in the latter case reliability is the same for all individuals, in irt analysis this can vary depending on the position of the individual on the trait continuum (reise et al., 2005). in the results presented below, the tic for each factor is shown together with the distribution of the underlying trait in the sample. the latter (a histogram) indicates how many people in each area of the construct are present in the sample. two aspects analogous to the two parameters (location and discrimination) – are important when interpreting the tics. first of all, the location of the curve along the continuum of the underlying construct indicates in what areas of the scale the instrument provides most information. secondly, the height of the curve demonstrates how much information the test provides for the different trait levels. an information level of five (y-axis in the graph presenting the tic) corresponds to a reliability of .80 (blue line in the figures); a level of 10 indicates a reliability of .90 (orange line). the area in which these reliability levels are reached are indicated with dotted vertical lines (respectively blue or orange) in the distribution histograms. this indicates the range of people for whom the scale is reliable at a level of .80 or .90. the item-characteristic curve describes the relationship between an individual’s position on the continuum of the measured latent construct (i.e., shared goals and cohesion; task interdependence; outcome vangrieken et al 18 | f l r interdependence) and the probability that he or she will indicate a particular response for an item that is aimed to measure that construct (reise et al., 2005). in polytomous response models (used here), the graphs for each item include a curve for each response category. the matrix plots (figure 2, 4, and 6) show the itemcharacteristic curves of the items included in each respective factor. as indicated, each curve represents one of the response categories (1 to 6). the difference in discrimination is demonstrated in the steepness of the curves for each response category and the degree to which these curves are distinguished rather than overlapping. the location of the items is demonstrated in the position of the curves on the horizontal axis (continuum of the underlying construct), items with curves mostly situated on the left side are easier to endorse compared to items on the right side. 7.2.1 shared goals and cohesion one non-functioning item had to be omitted from the factor shared goals and cohesion in order for the scale to function appropriately (ident2). the response categories of this item violated the assumption of ordinality. the tic of this scale (figure 1) demonstrates that the scale is most reliable at the lower and middle range of the underlying trait. figure 1. tic and distribution underlying trait shared goals and cohesion. vangrieken et al 19 | f l r a reliability of .90 (i.e., 10 on the information axis) is reached between trait levels -2 and approximately 0.20 and a reliability of .80 (i.e., 5 on the information axis) is reached between approximately -2.30 and 0.60. when interpreting this range it is important to compare this to the distribution of the underlying trait (figure 1), indicating that a large part of the sample can be assessed reliably using this scale. this is reflected in the slope (discrimination) and location parameters (table 5). table 5 irt parameter values item discrimination sd location sd shared goals and cohesion shared1 0.935 0.027 -1.422 0.033 shared2 1.753 0.052 -0.821 0.019 shared3 2.375 0.076 -0.741 0.015 shared4 1.649 0.048 -0.682 0.020 taskcoh2 1.939 0.058 -0.692 0.017 taskcoh4 1.826 0.053 -0.599 0.018 task interdependence ident6 0.569 0.015 -0.746 0.051 taskint2 1.187 0.032 -0.690 0.026 taskint3 0.923 0.024 -0.118 0.032 taskint4 1.286 0.036 -1.109 0.025 taskint6 2.126 0.067 -0.841 0.017 outint1 1.247 0.035 -1.079 0.026 outcome interdependence outint3 1.136 0.031 0.718 0.027 outint4 2.043 0.059 0.517 0.016 outint5 1.350 0.036 0.286 0.023 the range of the location parameters demonstrates that most of the item response categories are endorsed by respondents with lower than average shared goals and cohesion. this confirms that the scale is most useful in discriminating among individuals at the lower end of the trait continuum. the item parameters described above are visually presented in the item characteristic curves in figure 2. vangrieken et al 20 | f l r figure 2. matrix plot item characteristic curves shared goals and cohesion. 7.2.2 task interdependence compared to the structural analyses, irt analyses of this factor indicated that one item should be retained in order for this scale to function well (ident6). without this item, irt analyses did not run correctly and the other items did not function appropriately. this item was selected to be included because of conceptual reasons: the content matches with the other items in the factor and is an added value to the overall content of the factor. the tic of task interdependence (figure 3) demonstrates that a reliability of .80 is reached between a trait level of approximately -2.40 and 0.40. the bottom half of figure 3 indicates which part of the sample can be assessed with this level of reliability. vangrieken et al 21 | f l r figure 3. tic and distribution underlying trait task interdependence. the parameter values (table 5) again confirm that this scale provides most information and thus is most reliable in the lower and middle range of the trait. the item-characteristic curves for the task interdependence items are presented in figure 4. vangrieken et al 22 | f l r figure 4. matrix plot item characteristic curves task interdependence. 7.2.3 outcome interdependence the tic of outcome interdependence presented in figure 5 demonstrates that a reliability of .80 is reached between trait levels of approximately of -0.50 and 1.30. the bottom half of figure 5 again indicates which part of the sample can be assessed with this level of reliability. vangrieken et al 23 | f l r figure 5. tic and distribution underlying trait outcome interdependence. hence, compared to the tics of the previous factors, this scale provides most information at higher ends of the trait continuum. this is confirmed by the parameter values (table 5). figure 6 shows the itemcharacteristic curves for the outcome interdependence items. vangrieken et al 24 | f l r figure 6. matrix plot item characteristic curves outcome interdependence. 7.3 combining ctt and irt the solution found appropriate in the irt analyses was retested using the principles of ctt. this means that item ident2 was omitted from the first factor and ident6 was included in factor two. for these analyses the validation sample of the first wave (n2 = 658) was used. 7.3.1 cfa and internal consistency the results in table 6 demonstrate that the combined solution resulted in a good fit of the assumed model with our data. results of the hierarchical cfa indicated an appropriate fit (χ2/df=3.07; df=87; cfi=.945; tli= .934; rmsea=.056 [90 % ci [.048; .064]]; srmr=.050). all cronbach’s α indicators (table 6) were sufficient. for shared goals and cohesion, corrected item-whole correlations ranged from .58 to .85. for task interdependence these scores ranged from .46 to .80 and for the outcome interdependence factor from .66 to .81. moreover, for the hierarchical team entitativity factor corrected item-whole correlations ranged from .43 to .73. vangrieken et al 25 | f l r table 6 results (hierarchical) confirmatory factor analysis combined solution item regression weight standard error standardised regression weight critical ratioa cronbach’s alpha shared goals and cohesion .90 shared1 1.000 b .575 b shared2 1.342 .099 .754 13.609 shared3 1.505 .112 .853 13.428 shared4 1.420 .120 .805 11.830 taskcoh2 1.526 .139 .833 10.991 taskcoh4 1.609 .154 .829 10.426 task interdependence .82 ident6 1.000 b .482 b taskint2 1.354 .131 .747 10.334 taskint3 1.249 .136 .642 9.168 taskint4 .987 .132 .606 7.472 taskint6 1.280 .131 .808 9.809 outint1 1.070 .118 .670 9.079 outcome interdependence .81 outint3 1.000 b .672 b outint4 1.282 .079 .856 16.267 outint5 1.248 .081 .788 15.478 hierarchical cfa: team entitativity .89 shared goals and cohesion 1.000 b .610 b task interdependence 1.422 .233 .850 6.115 outcome interdependence 1.306 .217 .577 6.024 estimation method: maximum likelihood with robust standard errors aall critical ratios: p<.001 bvalue fixed at 1.00 for model identification purpose, hence no standard error was computed 7.3.2 shared variance and discriminant validity in order to investigate the shared variance of the items within the latent variables of the model, the squared multiple correlations (r2) of the items were assessed. for shared goals and cohesion, these ranged from .331 to .728, with a mean of .609, indicating that the six items in this scale account for 60.9% of the variance in these items. for task interdependence, 44.53% of the variance in the items is accounted for in this factor (range of r2 from .232 to .653). finally, for outcome interdependence the three items in this scale account for 60.18% of the variance in these items (range of r2 from .452 to .733). to assess discriminant validity, the guidelines of fornell and larcker (1981) were followed. they state that the average variance extracted by the latent factor should be higher than the variance explained by the correlation with another factor. discriminant validity is proven when the root of the average variance vangrieken et al 26 | f l r extracted (ave) exceeds the inter-factor correlations. the square root of the ave exceeded the correlation between each respective factor and other latent factors (table 7). next, discriminant validity was assessed in relation to psychological safety. cfa (again controlling for the clustering in the data) indicated an acceptable fit of the psychological safety instrument with our data after allowing covariance between one pair of items (χ2/df=4.31, cfi=.944; tli=.909; rmsea =.071 [90 % ci [.053; .091]]; srmr=.045). the standardised cronbach’s α coefficient equalled .84. the square root ave for each team entitativity factor exceeded the correlation between psychological safety and both the subscales as well as hierarchical scale of team entitativity (table 7), supporting discriminant validity of the instrument. however, the high correlation between shared goals and cohesion, and psychological safety indicates that both constructs are strongly related. experiencing your team as cohesive around shared goals and tasks appears to be strongly related to experiencing psychological safety in your team. this indicates the need for caution when investigating the role of these concepts in predicting other team processes or outcomes. table 7 correlation between team entitativity and psychological safety 1. 2. 3. 4. 1.shared goals and cohesion .78 2.task interdependence .48*** .67 3. outcome interdependence .31*** .43*** .78 4. psychological safety .61*** .26*** .05 .65 note: ***p<.001 diagonal: square root average variance extracted 7.3.3 convergence of team members’ perceptions as suggested earlier, our conceptualisation and operationalisation of team entitativity focuses on the individual team members’ perception of their team. this assumes that individuals’ responses within a team may vary, which was confirmed by the low icc(1) and icc(2) values and design effects. the extent to which these perceptions vary can also be seen as an indicator of the team’s degree of team entitativity. teams are supposed to develop a shared vision, inter alia on their own functioning (salas et al., 2000). the degree to which such a shared vision is present was investigated using the indicator for within-group agreement rwg proposed by james, demaree, and wolf (1984). rwg ranges from zero (complete absence of agreement) to 1 (full agreement). given the existing variety in individual team members’ perceptions of team entitativity, it is valuable to assess within-group agreement rather than calculating the average degree of entitativity of a team. aggregating individuals’ perceptions of their team’s degree of team entitativity, as is often done in research on team constructs, would not do justice to the complexity of team phenomena such as team entitativity. this neglects the individual-level variance, which was demonstrated to be substantial in this sample (see below). within-group agreement of each team was calculated for each factor, descriptive statistics are shown in table 8. in this case agreement on outcome interdependence appears to be the lowest and has most variation. vangrieken et al 27 | f l r table 8 descriptive statistics and correlations within-group agreement m sd min. max. 1. 2. 3. 1. rwg shared goals and cohesion .85 .18 0 1 2. rwg task interdependence .86 .15 0 1 .36*** 3. rwg outcome interdependence .72 .29 0 1 -.06 .23* note: **p<.01, ***p<.001 moreover, the degree of variance on both the within and the between level was investigated. results demonstrated that most of the variation appeared to be situated on the individual level. the within-level variances equated to .79, .63, and 1.24 while the between-level variances are .087, .051, and .055 for respectively shared goals and cohesion, task interdependence, and outcome interdependence. hence, most of the variance in the sample can be attributed to individual-level differences within teams while relatively little variance can be explained by the clustering in different teams. combining the insights derived from the rwg scores and distribution of variance across the within and between level, a more nuanced perspective on convergence of perceptions can be put forward. while moderate to relatively high rwg scores indicate that team members’ perceptions of team entitativity are quite similar within teams (members tend to converge to some extent), analysing the variance components demonstrates that team-level variance is low. hence, most of the differences in scores can still be attributed to differences between individuals rather than differences between teams. this is also demonstrated in the broad range of rwg scores (ranging from complete disagreement to complete agreement). this range indicates that in a substantial amount of teams, members’ perceptions of team entitativity tend to diverge and are very different. this indicates the value of also investigating team entitativity as a construct varying at the individual level – as most of the variance was situated on this level – with within-group agreement as a team-level indicator of convergence of entitativity rather than using it to assess an arbitrary cut-off argument for aggregation. 7.4 retest of the questionnaire 7.4.1 dropout analyses anova analyses were used to assess whether the mean team entitativity scores of teachers who were included in the second wave (n=731) differed significantly from those not included in the retest (n=589). teachers who did not participate in the second wave, scored slightly lower on shared goals and cohesion (f=14.05, df=1318, p>.01, η2=.01) and task interdependence (f=4.06, df=1318, p<.05, η2<.01) in the first wave. however, the low effect sizes indicate an overall limited effect of attrition on the results. no significant differences were found for outcome interdependence (f=.66; df=1318; p=.42, η2<.001). 7.4.2 internal consistency retest all cronbach’s α values were sufficient and did not significantly increase if one of the items would be dropped. for shared goals and cohesion (α=.90), corrected item-whole correlations ranged from .56 to .88. for task interdependence (α=.82) these scores ranged from .40 to .80 and for outcome interdependence (α=.83) from .76 to .78. moreover, for the hierarchical factor (α=.89) values ranged from .37 to .73. 7.4.3 longitudinal measurement invariance next, it was investigated whether the measurement of team entitativity is equivalent over time, vangrieken et al 28 | f l r meaning that the same construct with the same structure is measured. for each factor it was tested whether factor loadings and intercepts are equal over time (coertjens, donche, de maeyer, vanthournout, & van petegem, 2012). this was done by means of testing and comparing three models: (1) a baseline model (i.e., the basic model structure is invariant, factor loadings can differ over time), (2) a model including invariant loadings over time (i.e., metric invariance, items are interpreted in a similar way), and (3) a model including invariant intercepts over time (i.e., scalar invariance). the results (table 9) indicate that for shared goals and cohesion, factor loadings and intercepts were invariant. for task interdependence, loadings can be assumed invariant over time (δχ2=4.41, δdf=5, p=.492). the chi-square difference test for investigating invariance of intercepts was significant (δχ2=19.49, δdf=5, p=.002). because this test is influenced by the large sample size, the difference in cfi was checked, confirming intercept invariance (δcfi=.006). for outcome interdependence, the assumption of invariant loadings was again confirmed (δχ2=1.10, δdf=2, p=.576). however, only partial intercept invariance was reached. assessment of the items demonstrated that mostly outint3 violated the assumptions. when estimating a partial intercept invariance model (freeing the constraint on the intercept of outint3) an improved model fit was found. the difference in cfi (δcfi=.003) confirmed partial intercept invariance. this indicates caution when using sum scores of outcome interdependence when making comparisons over time. table 9 longitudinal measurement invariance model description χ2(df) cfi rmsea δχ2(δdf) p-value δcfi shared goals and cohesion baseline 150.55 (47) .967 .055 invariant loadings 161.45 (52) .965 .054 9.85 (5) .079 .002 invariant intercepts 171.53 (57) .964 .052 6.77 (5) .238 .001 task interdependence baseline 171.31 (47) .948 .060 invariant loadings 177.38 (52) .948 .057 4.41 (5) .492 .000 invariant intercepts 196.25 (57) .942 .058 19.49 (5) .002 .006 outcome interdependence baseline 5.85(5) .999 .015 invariant loadings 7.11 (7) 1.000 .005 1.10 (2) .576 .001 invariant intercepts 32.48 (9) .979 .060 26.94 (2) <.001 .021 partial intercept invariance 11.22 (8) .997 .023 4.27 (1) .039 .003 note: the satorra-bentler correction was used to calculate δχ2 and the p-value. 7.4.4 test-retest reliability correlations between the first and second wave equalled .69 (p<.001) for shared goals and cohesion, .62 (p<.001) for task interdependence, and .53 (p<.001) for outcome interdependence. as stated by kyndt et al. (2014), these moderate values can indicate that the underlying constructs in themselves are not stable over time. this is in line with the conceptualisation of team entitativity as a dynamic construct. moreover, testretest reliability assumes that the whole group of participants changes in the same way over time and does not take individual differences in these evolutions into account. however, the focus of our conceptualisation on an insider perspective looking at individuals’ perceptions and high levels of individual-level variance (demonstrated in low icc(1), icc(2), and design effects), indicate extensive individual differences. 7.4.5 predictive validity in order to check predictive validity, it was assessed whether the different aspects of team entitativity (first wave) are related to twe (second wave). cfa (taking the clustering in our data into account) indicated vangrieken et al 29 | f l r an appropriate fit of the twe instrument with our data (χ2/df=3.00; df=23; cfi=.987; tli= .981; rmsea=.055 [90 % ci [.041; .070]]; srmr=.017). all team entitativity features were significantly correlated to the twe components (table 10). with regard to the within group agreement variables, a small positive significant correlation was found between agreement concerning shared goals and cohesion and task interdependence on the one hand and dedication on the other hand and a small negative correlation between agreement concerning outcome interdependence and absorption. hence, in line with the theoretical assumptions of the job-demands resources model, the three defining features of team entitativity were positively related to twe measured at a later point in time. the relationship was strongest for shared goals and cohesion. this confirms predictive validity of the scales. however, less compelling results were found regarding within-group agreement; the latter does not demonstrate strong predictive validity in relation to twe in this sample. table 10 correlation between team entitativity and team work engagement shared goals and cohesion task interdependence outcome interdependence rwg shared goals and cohesion rwg task interdependence rwg outcome interdependence vigour .51*** .32*** .24*** .06 .04 -.08 dedication .55*** .29*** .14*** .10* .06* -.08 absorption .50*** .31*** .25*** .06 .03 -.10* note: *p<.05, **p<.01, ***p<.001 8. conclusion and discussion team entitativity was introduced here in order to tackle the conceptual gap between real teams and other types of groupings or pseudo-teams. the latter tend to be prominent in the educational sector as teacher teams are often teams-in-name only, leaving them neglected in traditional team research. the team entitativity concept tackles this gap by proposing a continuum-based conception of teams, referring to the degree to which a collection of individuals possesses the quality of a team and thus meets the criteria included in team definitions. in this way, it transcends the black-and-white conception in team research, opening up team research to various team types that were previously excluded. to determine the degree of team entitativity, three defining features were described based upon a combination of traditional team definitions and research on entitativity. these features include: shared goals and responsibilities, cohesion (task cohesion and identification with the team), and interdependence (task and outcome interdependence). previous research on team learning and team dynamics theorised and empirically confirmed the relationship between these defining features – especially cohesion and interdependence and the occurrence of successful team processes and learning behaviours (e.g., cohen & bailey, 1997; edmondson, 2002; mullen & copper, 1994; runhaar, ten brinke, kuijpers, wesselink, & mulder, 2014; van den bossche et al., 2006; wageman, 1995). hence, it can be argued that groups or teams with a higher degree of team entitativity are more proficient regarding team learning and successful team processes. this indicates that team entitativity is vangrieken et al 30 | f l r not only a descriptive term – describing what defines a team but also includes a normative value referring to a quality mark of groups and teams. however, the positive effects and added value of these entitativity features depends on certain conditions in a group’s context. first of all, there needs to be a reason to collaborate and work in teams as opposed to loosely coupled individuals. when there is no collective task or goal to start with and individuals can perform their tasks individually, there is no need for a highly entitative team to tackle these individualised tasks. this demonstrates the value of the task dimension in the teacher team typology proposed by vangrieken et al. (2013). the content of the task to be performed matters – whether it is limited to making mere practical arrangements or deep-level discussion of teaching practice; and not all tasks require the same level of team entitativity (meneses et al., 2008). this also demonstrates the importance of the first defining feature – especially having shared goals – as a baseline indication of a team’s degree of entitativity. theoretically, it can be assumed to be a condition sine qua non for high levels of team entitativity as being interdependent requires teams to have a shared task and goal. this appears to be especially important in the case of teams in education. the fact that collaboration in teacher teams is often restricted to a focus on practical arrangements (plauborg, 2009) and that teachers’ core task is mostly performed individually makes realising real teacher teams especially challenging. secondly, structural features such as shared goals and interdependence are subject to human interpretation and choices made by the team members. for example, a shared goal can be handled by collectively tackling the task at hand or by dividing the task in different parts that can be solved individually by the different team members. while structural features could be assessed more or less objectively, this is not the case for the subjective reality of individual team members – how they perceive and cope with these features of teamwork. however, this subjective reality does have an influence on teamwork that goes beyond objective structural features. hence, structural features are often dynamic phenomena that are shaped by team members’ choices (wageman et al., 2012). this confirms the conception of team entitativity from an insiders’ perspective as a dynamic construct. in line with the conception of team entitativity, a questionnaire was developed to empirically grasp the collaborative continuum among teacher teams. the psychometric quality was assessed by a combination of two research traditions: classical test theory (structural analyses) and item response theory (item-level analyses). structurally, the questionnaire consists of three factors, partially confirming the theoretically assumed defining features but also demonstrating their interrelatedness. the first factor, shared goals and cohesion, includes the first and part of the second theoretical feature. it refers to having a shared goal and task for which members share responsibility (shared goals and responsibility) and are collectively committed to realise it (task cohesion) in a unified team. hence shared goals and responsibility on the one hand and cohesion on the other hand empirically mostly cluster together. the aspect of identification is no longer directly represented in the final questionnaire, which may be related to the task focus of the conceptualisation and operationalisation of team entitativity. identification, feeling part of the team, may be interpreted as a social component and in this way does not really match with the task-focus. the second factor assesses task interdependence, referring to whether team members feel they need each other to perform their tasks, all members need to contribute, help each other, and communicate in order to realise the task of the team (task interdependence), and effective team collaboration matters to the members in the sense of their own job functioning (importance). as can be seen in this definition, the factor includes an aspect of importance, which was theoretically assumed to be part of the feature of cohesion. this is not surprising given the task-focus: the importance of a team with regard to members’ task performance is very closely related to task interdependence, also indicating a sense that the team matters for one’s job functioning. the final factor refers to outcome interdependence: the degree to which the team is evaluated and rewarded as a whole and whether team performance and goal attainment matter to the members as they influence their outcomes or evaluation. hierarchical cfa confirmed that while three aspects can be distinguished, they all relate to the same underlying team entitativity construct. this confirms the interrelatedness of the defining features in practice. irt analyses demonstrated that the range of strongest reliability of the scales is situated at the lower end, meaning that they mostly discriminate between people in the lower range of team entitativity. this was especially the case for shared goals and cohesion and task interdependence. when looking at the distributions vangrieken et al 31 | f l r of the underlying traits for the factors of team entitativity (figures 1, 3, and 5), it is demonstrated that a relatively large part of the sample can be captured reliably. furthermore, the convergence or the degree to which a shared perception of team entitativity was present was assessed by means of within-group agreement on each of the factors. in the teacher context the degree of outcome interdependence appears to be the least shared team construct: individual members’ perceptions within teams vary strongest. moreover, the variance components were analysed and demonstrated low team level variance. this indicates the value of combining analyses of rwg scores and variance components to get a full picture of convergence of team members’ perceptions. this low degree of team level variance combined with a large array of rwg scores (ranging from 0 to 1) indicates that while members’ perceptions overall tend to converge, most of the variance in team entitativity can be attributed to the individual level and teams vary with regard to their degree of convergence. this demonstrates the need for caution when using mean rwg scores of variables as an indicator for the appropriateness of aggregating the data to the team level. rather, it seems to be more valuable to use within-group agreement as a variable indicating convergence in its own right. moreover, the results suggest taking both the variance in within-group agreement and the degree of team level variance into account to assess at what level the variance in team entitativity is situated. despite the relevance of the conceptualisation and assessment of team entitativity as a team-level construct, these analyses of the variance components (and related icc(1) and icc(2) coefficients) indicated the lack of a strong team-concept. on the team level, team entitativity refers to the degree to which the team is perceived to be entitative, having the quality of a real team. however, these perceptions appear to be mostly explained by individual level differences. this confirms that team members’ perceptions of their team’s degree of entitativity can vary strongly and indicates the value of team entitativity as an individual-level concept. in the latter case, it also reflects individuals’ team identification or their willingness to contribute to the team’s task and goals. this relates to the insider perspective on team entitativity and the fact that defining team features are strongly influenced by choices and perceptions of individual members. moreover, the lack of membership stability and clear boundedness suggested by wageman et al. (2012) together with complicated organisation structures (such as in schools, described in section 2) with multiteam systems consisting of various teams and subteams, or members being part of multiple teams at the same time confirms the value of also investigating team entitativity as an individual level concept. hence the conception of team entitativity inherently reflects a multilevel construct and thus applies to the individual as well as the team level. while team concepts are traditionally conceptualised as having most meaning on the higher – team – level this indicates the value of team entitativity as an individual-level construct. the lack of a strong team concept in the analyses may partially be explained by the specific sample under investigation: teacher teams. as indicated in section 2, the educational context is mostly characterised by a focus on individual and independent work, with a tendency to emphasise teachers’ individual autonomy rather than teamwork. this results in a strong diversity in teacher collaboration and teachers’ attitudes towards teamwork. the applicability of the concept to the individual level makes it possible to also map this individual variability among teachers. 8.1 limitations this study has some limitations of which one needs to be aware when interpreting the results. first of all, the focus of this study was on teacher teams and the questionnaire was administrated in dutch. while team entitativity presents a valuable framework to conceptualise and investigate the diversity in teacher collaboration and teamwork, it is also more broadly applicable to teams in other (professional) contexts. hence, an important issue for future research includes validating the questionnaire in other contexts and assessing whether the same structure can be found. to foster translation to other contexts, all items were formulated more or less context-independent and not specifically focused on teacher teams. this makes it possible to validate the instrument in different vocational contexts. this would enhance the strength and applicability of the instrument and make it possible to make comparisons between different professions. moreover, an english version of the instrument should be validated in order to enhance more widespread use. in order to foster this, the dutch items were translated to english by a professional academic translation agency. translating the vangrieken et al 32 | f l r questionnaire, both to other professional contexts and different languages would make it possible to internationally compare the degree of team entitativity, both in education as well as in services and industry. however, given that the concept and questionnaire presented here may translate differently to different cultural and linguistic contexts, caution is warranted when comparing results across contexts. hence, when using the instrument in different cultural or linguistic contexts and different sectors, it is important to assess measurement invariance of the instrument across these different contexts in order to assess whether comparison of results is warranted. furthermore, given the low icc(1) and icc(2) coefficients in this sample, the questionnaire was analysed solely at the individual level. although the clustering in the data was taken into account when analysing the structure of the questionnaire (providing robust standard errors and chi-square values), no teamlevel models were investigated. hence, the conceptually multilevel nature of the construct could not yet be investigated as such in the analyses of the questionnaire. future research using the questionnaire should investigate the iccs in the sample under investigation in order to check whether multilevel analyses are deemed appropriate, investigating the structure of the questionnaire at the individual and team level. this could contribute to the understanding of the multilevel nature of team entitativity. if sufficient team-level variance is present for performing multilevel analyses, these could capture team entitativity both as an individual and a team level construct. furthermore, analyses of discriminant validity indicated a high correlation between the shared goals and cohesion scale and psychological safety. hence, there is a need for caution when investigating both constructs as predictors of other processes and outcomes, and it would be advised to assess this correlation and – if required take it into account in the analyses to avoid multicollinearity. as suggested above, irt analyses demonstrated rather low location parameters, indicating that the range of strongest reliability is mostly focused at the lower end of the underlying trait. this must be taken into account when using the questionnaire and analysing the results: in this case, most reliable statements can be made with regard to people in the lower end of team entitativity and the scale is less reliable to distinguish between different people at the higher end of the continuum. this can partially be a consequence of the fact that teacher teams are often not highly entitative teams but often tend to be teams in name only. the lack of a larger portion of cases and variety in the higher end of the team entitativity scale makes it difficult to validate the instrument for this higher end of the continuum. when using the instrument in other professional contexts, this should be reassessed. the range of reliability of each scale should be compared to the distribution of the underlying trait in the sample under investigation to assess which range of the sample is reliably measured. an important area for future research regarding this instrument thus includes applying it to samples with a larger diversity in the higher range of the scale (i.e., more entitative teams). this would make it possible to further develop and validate the instrument as more reliable at the higher end of the team entitativity continuum. finally, analyses regarding longitudinal measurement invariance demonstrate only partial intercept invariance for outcome interdependence. this must be taken into account when analysing sum scores for this scale. when using analyses starting from the items rather than sum scores (structural equation modelling and latent growth curve analyses), this can be modelled in the analysis. thus when using this scale in longitudinal designs, change over time can be investigated using latent growth curve models. 8.2 contributions and future research the introduction of team entitativity, conceptually as well as empirically, contributes to research in different ways. this study helps in transcending three types of boundaries or gaps: (1) conceptual, (2) methodological, and (3) context-related. first of all, the concept team entitativity makes it possible to break conceptual boundaries between different research traditions. it creates a shared language when talking about groups and teams and thus builds a conceptual bridge between different research traditions using a variety in terminology (i.e., those focusing on strict teams versus those including a more broad perspective on groups and collaboration). vangrieken et al 33 | f l r secondly, this study transcends methodological boundaries by combining ctt and irt analyses. although irt analyses are rarely applied in the development and validation of questionnaires in this research area, it provides additional information that gets lost when relying solely on ctt. different aspects of this study demonstrate the added value of this combination. first, while cronbach’s α (ctt) provides information on the reliability of the scale overall, it does not indicate for who (i.e., in which range of the trait continuum) the scale is most reliable. irt analyses provide more nuanced insights as they give additional information on the range of the trait continuum along which the scale is most reliable. second, while ctt is focused on scalebased analyses, providing insights in the structure of the questionnaire, irt-analyses are item-based and provide information on the functioning of each item. hence, combining ctt and irt combines the best of both worlds: structural information on the scale and item-level analyses. a third line of demarcation is related to the research context. there appears to be a clear distinction and lack of boundary crossing between traditional team research and research on teacher collaboration and teamwork. teacher teams are mostly excluded from team research and studies investigating teacher collaboration rarely include theoretical and methodological approaches developed and well-validated in team research. the latter makes research on teacher collaboration vulnerable to criticism for lacking both conceptual and methodological rigour. introducing team entitativity makes it possible to transcend these boundaries and to foster cross-pollination between traditional team research and research in education. this fosters the inclusion of teacher teams (and other types of non-strict teams) in team research. the team entitativity concept makes it possible to conceptualise and measure the variety in teamness, investigate whether differences can be found, and to assess whether the theories, models, and research methods used in this tradition also apply to teams with varying degrees of entitativity. it thus becomes possible to further replicate earlier research in a broader context. this contribution is in line with one of the recommendations made by wageman et al. (2012). they argued that defining team features in research should be used to investigate various collaborations rather than rule out certain types of collaboration because they do not fit the traditional team definition. moreover, the dynamic nature of the team entitativity conceptualisation presented here suggests the need for longitudinal investigations. both the defining features and the degree of convergence of team members’ perceptions may vary through time. furthermore, breaking these boundaries makes it possible to strengthen the theoretical underpinnings and methodological approaches of research on teacher collaboration. while teacher collaboration has proven to be valuable (vangrieken et al., 2015), deep-level research in underlying dynamics and characteristics of this collaboration lacks. the team entitativity framework and integration of insights from team research makes it possible for future research to tackle this lack of theorising and empirical deep-level investigation of collaborative endeavours among teachers. moreover, future research could answer questions such as: does it matter to be more team-like? which defining feature matters most and is this different in different contexts? as hypothesised above, in the context of teachers for whom individual work still mostly consumes the major part of their job, baseline entitativity features such as having shared goals may be most important. moreover, a specific conceptual application could be the inclusion of the concept in meta-analyses of team and group research. if sufficient information is present in the primary studies, the degree of team entitativity of the groupings could be estimated based upon the description of the sample and other background information on the groups. starting from the identified features included in the instrument, assessment of this background information by multiple judges could lead to an estimation of the mean degree of team entitativity in the study. this could then be included as a moderator in the meta-analyses, making it possible to compare studies on different team and group types and to assess the influence of team entitativity. vangrieken et al 34 | f l r keypoints we conceptualise the gap between team and group by introducing team entitativity or the degree to which defining team features are present in a group we combine two theoretical frameworks by introducing and integrating entitativity that originates from social psychology research into team research team entitativity is measured by the degree of shared goals and cohesion, task interdependence, outcome interdependence, and convergence of perceptions analyses combine traditional classical test theory and item response theory analyses, which provides a clear added-value and complementary insights the concept and instrument make it possible to map the variety of groups and teams and assess the applicability of traditional frameworks herein references albrecht, s. l. (2012). the influence of job, team and organizational level resources on employee well-being, engagement, commitment and extra-role performance: test of a model. international journal of manpower, 33, 840-853. doi: 10.1108/01437721211268357 aubé, c., & rousseau, v. (2005). team goal commitment and team effectiveness: the role of task interdependence and supportive behaviors. group dynamics: theory, research and, practice, 9, 189204. doi:10.1037/1089-2699.9.3.189 bartko, j. j. (1976). on various intraclass correlation reliability coefficients. psychological bulletin, 83, 762765. doi:10.1037/0033-2909.83.5.762 beersma, b., homan, a. c., van kleef, g. a., de dreu, c. k. w. (2013). outcome interdependence shapes the effects of prevention focus on team processes and performance. organizational behaviour and human decision processes, 121, 194-203. doi:10.1016/j.obhdp.2013.02.003 bentler, p. m., & chou, c. (1987). practical issues in structural modelling. sociological methods & research, 16, 78–117. doi:10.1177/0049124187016001004 birenbaum, m., kimron, h., & shilton, h. (2011). nested contexts that shape assessment for learning: schoolbased professional learning community and classroom culture. studies in educational evaluation, 37, 35-48. doi:10.1016/j.stueduc.2011.04.001 bouas, k. s., & arrow, h. (1996). the development of group identity in face-to-face and computer-mediated groups with membership change. computer supported cooperative work, 4, 153–178. doi:10.1007/bf00749745 brewer, m., hong, y., & li, q. (2004). dynamic entitativity: perceiving groups as actors. in v. yzerbyt, c. m. judd, & o. corneille (eds.), the psychology of group perception: perceived variability, entitativity, and essentialism (pp. 25-38). new york, ny: psychology press. brouwer, p. (2011). collaboration in teacher teams (doctoral dissertation, utrecht university, the netherlands). retrieved from http://igitur-archive.library.uu.nl/dissertations/2011-1110200504/brouwer.pdf brown, m. w., & cudeck, r. (1993). alternative ways of assessing model fit. in a. bollen & j.s. long (eds.), testing structural equation models (pp. 136-162). california: sage publications inc. buljac, m., van woerkom, m., & van wijngaarden, j. d. h. (2013). are real teams healthy teams? journal of healthcare management, 58, 92-109. retrieved from hdl.handle.net/1765/50224 campbell, d. t. (1958). common fate, similarity, and other indices of the status of aggregates of persons as social entities. behavioral science, 3, 14-25. doi:10.1002/bs.3830030103 campion, m. a., medsker, a., & higgs, a. c. (1993). relations between work group characteristics and effectiveness: implications for designing effective work groups. personnel psychology, 46, 823-850. vangrieken et al 35 | f l r doi:10.1111/j.1744-6570.1993.tb01571.x carless, s. a., & de paola, c. (2000). the measurement of cohesion in work teams. small group research, 31, 107-118. doi:10.1177/104649640003100104 carpenter,s., fortune, j.l., delugach, h.s., etzkom, l.h., utley, d.r., farrington, p.a., & virani, s. (2008). studying team shared mental models. in p.j. ågerfalk, h. delugach & m. lind (eds.), proceedings of the 3rd international conference on the pragmatic web: innovating the interactive society (pp. 41-48). new york, ny: acm. doi:10.1145/1479190.1479197 carron, a. v. (1982). cohesiveness in sport groups: implications and considerations. journal of sport psychology, 4, 123-138. doi:10.1123/jsp.4.2.123 castano, e., yzerbyt, v., paladino, m., & sacchi, s. (2002). i belong, therefore, i exist: ingroup identification, ingroup entitativity, and ingroup bias. personality and social psychology bulletin, 28, 135-143. doi:10.1177/0146167202282001 cavazza, n., pagliaro, s., & guidetti, m. (2014). antecedents of concern for personal reputation: the role of group entitativity and fear of social exclusion. basic and applied psychology, 36, 365-376. doi:10.1080/01973533.2014.925453 coertjens, l., donche, v., de maeyer, s., vanthournout, g., & van petegem, p. (2012). longitudinal measurement invariance of likert-type learning strategy scales: are we using the same ruler at each wave? journal of psychoeducational assessment, 30, 577-587. doi:10.1177/0734282912438844 cohen, s. g., & bailey, d. e. (1997). what makes teams work: group effectiveness research from the shop floor to the executive suite. journal of management, 23, 239-290. doi:10.1177/014920639702300303 costa, p. l., passos, a. m., & bakker, a. b. (2012). team work engagement: considering team dynamics for engagement (working paper no. 12-06). retrieved from business research unit website http://bruunide.iscte.pt/repec/pdfs/12-06.pdf crawford, m. t., & salaman, l. (2012). entitativity, identity, and the fulfillment of psychological needs. journal of experimental social psychology, 48, 726-730. doi:10.1016/j.jesp.2011.12.015 crawford, m. t., sherman, s. j., & hamilton, d. l. (2002). perceived entitativity, stereotype formation, and the interchangeability of group members. journal of personality and social psychology, 83, 1076-1094. doi:10.1037//0022-3514.83.5.107 crow, m. g., & pounder, d. g. (2000). interdisciplinary teacher teams: context, design, and process. educational administration quarterly, 36, 216-254. doi:10.1177/0013161x00362004 dasgupti, n., banaji, m. r., & abelson, r. p. (1999). group entitativity and group perception: associations between physical features and psychological judgments. journal of personality and social psychology, 77, 991-1003. doi:0022-3514/99/$3.00 de witte, h., verhofstadt, e., & omey, e. (2007). testing karasek’s learning and strain hypotheses on young workers in their first job. work & stress, 21, 131-141. doi:10.1080/02678370701405866 decuyper, s., dochy, f., & van den bossche, p. (2010). grasping the dynamic complexity of team learning: an integrative model for effective team learning in organisations. educational research review, 5, 111133. doi:10.1016/j.edurev.2010.02.002 demerouti, e., bakker, a. b., nachreiner, f., & schaufeli, w. (2001). the job demands-resources model of burnout. journal of applied psychology, 80, 499-512. doi:10.1037/0021-9010.86.3.499 dyer, n. g., hanges, p. j., & hall, r. j. (2005). applying multilevel confirmatory factor analysis techniques to the study of leadership. the leadership quarterly, 16, 149-167. doi:10.1016/j.leaqua.2004.09.009 edelen, m. o., & reeve, b. b. (2007). applying item response theory (irt) modeling to questionnaire development, evaluation, and refinement. quality of life research, 16, 5-18. doi:10.1007/s11136-0079198-0 edmondson, a. c. (1999). psychological safety and learning behavior in work teams. administrative science quarterly, 44, 350-383. doi:10.2307/2666999 edmondson, a. c. (2002). the local and variegated nature of learning in organisations: a group-level perspective. organisation science, 13, 128-146. doi:10.1287/orsc.13.2.128.530 edmondson, a. c. (2008). the competitive imperative of learning. harvard business review, 86(7-8), 60-67. doi:10.1109/emr.2014.6966928 vangrieken et al 36 | f l r festinger, l. (1950). informal social communication. psychological review, 57, 271-282. doi:10.1037/h0056932 fornell, c., & larcker, d.f. (1981). evaluating structural equation models with unobservable variables and measurement error. journal of marketing research, 18, 39–50.
 doi:10.2307/3151312 gajda, r., & koliba, c. j. (2008). evaluating and improving the quality of teacher collaboration: a field-tested framework for secondary school leaders. nassp bulletin, 92, 133–153. doi:10.1177/0192636508320990 guzzo, r. a., & shea, g. p. (1992). group performance and intergroup relations in organizations. in m. d. dunnette, l. h. hough (eds.). handbook of industrial and organizational psychology, vol. 3 (pp. 269313). consulting psychologists press: palo alto. hackman, j. r. (2012). from causes to conditions in group research. journal of organizational behavior, 33, 428-444. doi:10.1002/job.1774 hair, j., black, w., babin, b., anderson, r., & tatham, r. (2006). multivariate data analysis (6th ed.). upper saddle river, pearson prentice hall.
 hamilton, d. l. (2007). understanding the complexities of group perception: broadening the domain, 37, 1077-1101. doi:10.1002/ejsp.436 hamilton, d. l., & sherman, s. (1996). perceiving persons and groups. psychological review, 103, 336-355. doi:10.1037//0033-295x.103.2.336 hamilton, d. l., sherman, s.j., & castelli, l. (2002). a group by any other name – the role of entitativity in group perception. european review of social psychology, 12, 139-166. doi:10.1080/14792772143000049 henry, k. b., arrow, h., & carini, b. (1999). a tripartite model of group identification: theory and measurement. small group research, 30, 558-581. doi:10.1177/104649649903000504 hu, l., & bentler, p.m. (1999). cut-off criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives. structural equation modeling, 6, 1-55. doi:10.1080/10705519909540118 james, l. r., demaree, r. g., & wolf, g. (1984). estimating within-group interrater reliability with and without response bias. journal of applied psychology, 69, 85-98. doi:10.1037/0021-9010.69.1.85 jans, l., postmes, t., & van der zee, k. i. (2011). the induction of shared identity: the positive role of individual distinctiveness in groups. personality and social psychology bulletin, 37, 1130-1141. doi:10.1177/0146167211407342 kahn, w. a. (1990). psychological conditions of personal engagement and disengagement at work. the academy of management journal, 33, 692-724. doi:10.2307/256287 katzenbach, j. r. & smith, d. k. (1993). the discipline of teams. harvard business review, 71, 111-120. kiggundu, m. n. (1983). task interdependence and job design: test of a theory. organizational behavior and human performance, 31, 145-172. doi:0030-5073/83/020145-28$03.00/0 kozlowski, s. w. j. (2015). advancing research on team process dynamics: theoretical, methodological, and measurement considerations. organizational psychology review, 5, 270-299. doi:10.1177/2041386614533586 kozlowski, s. w. j., & bell, b. s. (2003). work groups and teams in organizations. in w. c. borman, d. r. ilgen, & r. j. klimoski (eds.). handbook of psychology: vol. 12. industrial and organizational psychology (pp. 333-375). london: wiley. kozlowski, s. w. j., & chao, g. t. (2012). the dynamics of emergence: cognition and cohesion in work teams. managerial and decision economics, 33, 335-354. doi:10.1002/mde.2552 kyndt, e., janssens, i., coertjens, l., gijbels, d., donche, v., & van petegem, g. (2014).vocational education students’ generic working life competencies: developing a self-assessment instrument. vocations and learning, 7, 365-392. doi:10.1007/s12186-014-9119-7 leonard, p., & leonard, l. (2001b). the collaborative prescription: remedy or reverie? international journal of leadership education, 4, 383-399. doi:10.1080/13603120110078016 lickel, b., hamilton, d. l., wieczorkowska, g., lewis, a., sherman, s. j., & uhles, a. n. (2000). varieties of groups and the perception of group entitativity. journal of personality and social psychology, 78, 223246. doi:10.1037//0022-3514..78.2.223 vangrieken et al 37 | f l r lima, j.a. (2001). forgetting about friendship: using conflict in teacher communities as a catalyst for school change. journal of educational change, 2, 97-122. doi:10.1023/a:1017509325276 lortie, d. c. (1975). schoolteacher: a sociological study. london: university of chicago press. lyubovnikova, j., west, m. a., dawson, j. f., & carter, m. r. (2015). 24-karat or fool’s gold? consequences of real team and co-acting group membership in healthcare organizations. european journal of work and organizational psychology, 24, 929-950. doi:10.1080/1359432x.2014.992421 main, k. (2007). a year-long study of the formation and development of middle school teaching teams (doctoral dissertation, griffith university, meadowbrook, australia). https://www120.secure.griffith.edu.au/rch/file/64a6473e-3a2b-f149-bd30-6e2033bbef0f/1/02whole.pdf mathieu, j. e., heffner, t. s., goodwin, g. f., salas, e., & cannon-bowers, j. a. (2000). the influence of shared mental models on team process and performance. journal of applied psychology, 85, 273-283. doi:10.1037//0021-9010.85.2.273 mathieu, j., maynard, m.t., rapp, t., & gilson, l. (2008). a review of recent advancements and a glimpse into the future. journal of management, 34. doi:10.1177/0149206308316061 mcgrath, j. e. (1984). groups: interaction and performance. englewood cliffs, nj: prentice-hall. mcleod, j., & von treuer, k. (2013). towards a cohesive theory of cohesion. international journal of business and social research, 3(12), 1-11. retrieved from http://www.thejournalofbusiness.org/ meneses, r., ortega, r., navarro, j., & de quijano, s. d. (2008). criteria for assessing the level of group development (lgd) of work groups: groupness, entitativity, and groupality as theoretical perspectives. small group research, 39, 492-514. doi:10.1177/1046496408319787 moolenaar, n. m. (2010). ties with potential: nature, antecedents, and consequences of social networks in school teams (doctoral dissertation, university of amsterdam, the netherlands). retrieved from http://dare.uva.nl mullen, b., & copper, c. (1994). the relation between group cohesiveness and performance: an integration. psychological bulletin, 115, 210-227. doi:10.1037//0033-2909.115.2.210 newman, d. a., & sin, h. (2009). how do missing data bias estimates of within-group agreement? sensitivity of sdwg, cvwg, rwg(j), rwg(j)*, and icc to systematic nonresponse. organizational research methods, 12, 113-147. doi:10.1177/1094428106298969 ohlsson, j. (2013). team learning: collective reflection processes in teacher teams. journal of workplace learning, 25, 296-309. doi:10.1108/jwl-feb-2012-0011 pearce, j. l., & gregersen, h. b. (1991). task interdependence and extrarole behaviour: a test of the mediating effects of felt responsibility. journal of applied psychology, 76, 838-844. doi:10.1037/00219010.76.6.838 prentice, d. a., miller, d. t., & lightdale, j. r. (1994). asymmetries in attachments to groups and to their members: distinguishing between common-identity and common-bond groups. personality and social psychology bulletin, 20, 484-493. doi:10.1177/0146167294205005 raes, e., kyndt, e., decuyper, s., van den bossche, p., & dochy, f. (2015). an exploratory study of group development and team learning. human resource development quarterly, 26, 5-30. doi:10.1002/hrdq.21201 reise, s. p., ainsworth, a. t., & haviland, m. g. (2005). item response theory: fundamentals, applications, and promise in psychological research. current directions in psychological science, 14, 95-101. doi:10.1111/j.0963-7214.2005.00342 r core team (2015). r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. retrieved from http://www.rproject.org revelle, w. d. (2012). psych: procedures for personality and psychological research. evanston: northwestern university. retrieved from http://personality-project.org/r/psych.manual.pdf. runhaar, p., ten brinke, d., kuijpers, m., wesselink, r., & mulder, m. (2014). exploring the links between interdependence, team learning and a shared understanding among team members: the case of teachers facing an educational innovation. human resource development international, 17, 67-87. doi:10.1080/13678868.2013.856207 rutchik, a. m., hamilton, d. l., & sack, j. d. (2008). antecedents of entitativity in categorically and vangrieken et al 38 | f l r dynamically construed groups. european journal of social psychology, 38, 905-921. doi:10.1002/ejsp.555 saavedra, r., earley, p. c., & van dyne, l. (1993). complex interdependence in task-performing groups. journal of applied psychology, 78, 61-72. doi:0021-9010/93/$3.00 salas, e., bowers, c. a., & cannon-bowers, j. a. (1995). military team research: 10 years of progress. military psychology, 7, 55–75. doi:10.1207/s15327876mp0702_2 salas, e., burke, c. s., & cannon-bowers, j. a. (2000). teamwork: emerging principles. international journal of management reviews, 2, 339-356. doi:10.1111/1468-2370.00046 sassenberg, k. (2002). common bond and common identity groups on the internet: attachment and normative behavior in on-topic and off-topic chats. group dynamics: theory, research, and practice, 6, 27-37. doi:10.1037//1089-2699.6.1.27 satorra, a., & bentler, p. m. (2010). ensuring positiveness of the scales difference chi-square test statistic. psychometrika, 75, 243-248. doi:10.1007/s11336-009-9135-y schaufeli, w. b., & bakker, a. (2004). job demands, job resources and their relationship with burnout and engagement: a multi-sample study. journal of organizational behavior, 25, 293-315. doi:10.1002/job.248 smith, g. (2009). if teams are so good… science teachers’ conceptions of teams and teamwork (doctoral dissertation, queensland university of technology, australia). retrieved from http://eprints.qut.edu.au/31734/1/gregory_smith_thesis.pdf somech, a. (2008). managing conflict in school teams: the impact of task and goal interdependence on conflict management and team effectiveness. educational administration quarterly, 44, 359-390. doi:10.1177/0013161x08318957 spencer-rodgers, j., hamilton, d. l., & sherman, s. j. (2007). the central role of entitativity in stereotypes of social categories and task groups. journal of personality and social psychology, 92, 369-388. doi:10.1037/0022-3514.92.3.369 sundström, e., mcintyre, m., halfhill, t., & richards, h. (2000). work groups: from the hawthorne studies to work teams of the 1990s and beyond. group dynamics: theory, research, and practice, 4, 44-67. doi:10.1037//1089-2699.4 svirydzenka, n., sani, f., & bennet, m. (2010). group entitativity and its perceptual antecedents in varieties of groups: a developmental perspective. european journal of social psychology, 40, 611-624. doi:10.1002/ejsp.761 torrente, p., salanova, m., llorens, s., & schaufeli, w. b. (2012). from “i” to “we”: the factorial validity of a team work engagement scale. in j., neves, & s.p. gonçalves (eds.). occupational health psychology: from burnout to well-being (pp. 333-355). lisboa: alth psychology. van den bossche, p., gijselaers, w.h., segers, m., kirschner, p.a. (2006). social and cognitive factors driving teamwork in collaborative learning environments: team learning beliefs and behaviors. small group research, 37, 490-521. doi:10.1177/1046496406292938 van der vegt, g. s., & bunderson, j. s. (2005). learning and performance in multidisciplinary teams: the importance of collective team identification. the academy of management journal, 48, 532-547. doi:10.2307/20159674 van der vegt, g. s., emans, b., & van de vliert, e. (1998). motivating effects of task and outcome interdependence in work teams. group & organization management, 23, 124-143. doi:10.1177/1059601198232003 van der vegt, g. s., emans, b., & van de vliert, e. (1999). effects of interdependencies in project teams. the journal of social psychology, 139, 202-214. doi:10.1080/00224549909598374 van der vegt, g. s., & janssen, o. (2003). joint impact of interdependence and group diversity on innovation. journal of management, 29, 729-751. doi:10.1016/s0149-2063(03)00033-3 vangrieken, k., dochy, f., & raes, e. (2016). team learning in teacher teams: team entitativity as a bridge between teams-in-theory and teams-in-practice. european journal of psychology of education, 31, 275298. doi:10.1007/s10212-015-0279-0 vangrieken, k., dochy, f., raes, e., & kyndt, e. (2013). team entitativity and teacher teams in schools: vangrieken et al 39 | f l r towards a typology. frontline learning research, 1, 86-98. doi:10.14786/flr.v1i2.23 vangrieken, k., dochy, f., raes, e., & kyndt, e. (2015). teacher collaboration: a systematic review. educational research review, 15, 17-40. doi:10.1016/j.edurev.2015.04.002 vangrieken, k., meredith, c., packer, t., & kyndt, e. (2017). teacher communities as a context for professional development: a systematic review. teaching and teacher education, 61, 47-59. doi:10.1016/j.tate.2016.10.001 wageman, r. (1995). interdependence and group effectiveness. adminsitrative science quarterly, 40, 145180. doi:0001-8392/95/4001-0145/$1 .00. wageman, r., gardner, h., & mortensen, m. (2012). the changing ecology of teams: new directions for teams research. journal of organizational behavior, 33, 301-315. doi:10.1002/job.1775 wageman, r., hackman, j. r., & lehman, e. (2005). team diagnostic survey: development of an instrument. the journal of applied behavioral sciance, 41, 373-398. doi:10.1177/0021886305281984 weick, k. (1976). educational organizations as loosely coupled systems. administrative science quarterly, 21, 1-19. west, m. a., & lyubovnikova, j. (2012). real teams or pseudo-teams? the changing landscape needs a better map. industrial and organizational psychology, 5, 25-55. doi:10.1111/j.1754-9434.2011.01397.x westheimer, j. (2008). learning among colleagues: teacher community and the shared enterprise of education. in m. cochran-smith, s. feiman-nemser, & j. mcintyre (eds.). handbook of research on teacher education (pp. 756-782). reston, va and lanham, md: association of teacher educators and rowman. yzerbyt, v. corneille, o., & estrada, c. (2001). the interplay of subjective essentialism and entitativity in the formation of stereotypes. personality and social psychology review, 5, 141-155. doi:10.1207/s15327957pspr0502_5 zaccaro, s. j. (1991). nonequivalent associations between forms of cohesiveness and group-related outcomes: evidence for multidimensionality. the journal of social psychology, 131, 387-399. doi:10.1080/00224545.1991.9713865 vangrieken et al 40 | f l r appendix overview of items team entitativity item shared1 to me, the objective of this team is unclear a shared2 i think that, as a team, we have a clear and shared objective shared3 within this team, i have the feeling that we are all responsible for the achievement of our mission, as well as for the way in which we do that shared4 within this team, i think that we rely on each other’s contributions for the achievement of our mission shared5 in order to achieve our objective, i think that all team members should contribute taskcoh1 the team members collaborate to achieve a shared mission b taskcoh2 in my opinion, within this team, we have a clear and shared mission, which we all try to achieve as a team taskcoh3 i think that the members of this team are committed to our mission b taskcoh4 within this team, i feel a strong and shared involvement in our mission ident1 i feel connected with this team and its members b ident2 i feel part of this team ident3 the mission and objectives of this team are also relevant to and important for my job as a teacher ident4 i do not consider the work i do for his team as an intrinsic part of my job a ident5 for my job as a teacher, i consider it important to be part of this team ident6 whether or not the team achieves its objectives does not really influence me as a teacher a taskint1 i think that teachers mainly work on their own and that they do not really need to collaborate with team members a taskint2 in order to be a good teacher, i have to collaborate with the other members of my team taskint3 my team members’ accomplishments influence the way i function and what i achieve at work. similarly, what i accomplish at work also has an impact on my team members taskint4 i think that a lot of communication and coordination between the team members is needed in order to achieve the desired results in this team taskint5 within this team, i think that we need each other’s knowledge and advice in order to complete our tasks successfully b taskint6 within this team, i think that we need each other’s help and support in order to complete our tasks successfully vangrieken et al 41 | f l r taskint7 within this team, i notice that we all depend on each other to perform our tasks well outint1 i benefit from my team members doing their job well outint2 as a teacher, when the colleagues in my team work really well, i am put in a disadvantage a outint3 the feedback i receive on how well i work is mainly based on information about how well the entire team functions outint4 my team’s achievements are a considerable part of my job appraisals and assessments outint5 i have the feeling that, in the end, we and our achievements are assessed as a team a these items need to be reverse coded prior to analyses. b these items were omitted from the questionnaire prior to the efa based upon inter-item correlations. note: the items were translated for publication purposes. the original and validated questionnaire is in dutch. items that were retained in the final version of the instrument are indicated in bold and italics. codepen hietajarvi et al publication frontline learning research vol.8 no. 1 (2020) 33 55 issn 2295-3159 are schools alienating digitally engaged students? longitudinal relations between digital engagement and school engagement lauri hietajärvia, kirsti lonkaa, kai hakkarainena, kimmo alhob & katariina salmela-aroa a faculty of educational sciences, university of helsinki b department of psychology and logopedics, faculty of medicine, university of helsinki article received 29 november 2018 / revised 15 december 2019/ accepted 2 february 2020/ available online 20 february abstract this article examined digital learning engagement as the out-of-school learning component that reflects informally emerging socio-digital participation. the gap hypothesis proposes that students who prefer learning with digital technologies outside of school are less engaged in traditional school. this hypothesis was approached from the framework of connected learning, referring to the process of connecting self-regulated and interest-driven learning across formal and informal contexts. we tested this hypothesis with longitudinal data. it was of interest how digital engagement, operationalized as a general digital learning preference, wish for digital schoolwork, and their interaction, is related to traditional school engagement. this was examined both cross-sectionally in three time points and longitudinally across three years. the participants were 1,705 (43.7% female) 7th–9th graders (13-15 years old) from 27 schools in helsinki, finland. we explored the structure of correlations between latent constructs at each time point separately, and finally, to evaluate longitudinal relations between digital engagement and school engagement we specified latent cross-lagged panel models. the results indicate that students holding a stronger general digital learning preference experienced higher schoolwork engagement, both contemporaneously and over time, indicating successful connected learning. however, the results also showed support for the gap hypothesis: students who preferred digital learning but did not have the chance to digitally engage at school, experienced a decrease in school engagement over time. the article shows that there is a need to examine the reciprocal interactive processes between the learners and their social ecologies inside and outside school more closely. keywords: connected learning; digital engagement; schoolwork engagement; gap hypothesis; longitudinal analysis info corresponding author: lauri.hietajarvi@helsinki.fi doi: 10.14786/flr.v8i1.437. 1. introduction connecting learning across in-school and out-of-school contexts has been a continuous challenge in education (e.g., malcolm, hodkinson, & colley, 2003) and the novel informal learning opportunities provided by digital media have highlighted tensions between informal and formal practices of learning (ito et al., 2013). in general, previous research shows that the more students spend time engaging with digital media the more skills they are able to acquire (eu kids online, 2014). when these informally cultivated digital practices are successfully connected with academic learning practices, such students are likely to flourish and extend their potentials (ito et al., 2013). yet, finnish young people are not provided sufficient structured support in school for cultivating advanced digital competences, as digital technologies are used at finnish school seldom and mostly for shallow training of basic digital skills (european parliament, 2015; european commission, 2017; hakkarainen, hietajärvi, alho, lonka, & salmela-aro, 2015). in this condition, an increased gap or misfit between a digitally engaged learner and the learning environment may occur. the gap hypothesis proposes that students who prefer learning with digital technologies outside of school are less engaged in traditional school. this is problematic because school engagement is crucial for students’ learning, academic development, and well-being (salmela-aro & upadyaya, 2012; upadyaya & salmela-aro, 2013). thus, promoting practices that connect informal and formal learning as well as support school engagement should be the main goals of modern pedagogical practices. however, in comparison to other pisa countries, technology-enhanced pedagogies appear to be not so widely adopted in finnish schools (oecd, 2015), and utilizing digital technologies successfully in education calls for transformations in the social practices of schooling, which appear to be happening very slowly (hakkarainen, 2009). such transformations are nevertheless needed, as the schooling system should prepare students for the current technology-rich innovation-driven society that calls for collaborative solving of complex non-routine problems and cultivating associated personal and social competences. toward that end, it is important to learn creative and academic practices of using socio-digital technologies (hakkarainen et al., 2015). the conventional individualist, acquisition-oriented, and teacher-centered educational practices prevailing at school are considered to be a major hindrance to creating such a workforce (robinson, 2011). the present study addresses the conditions of continuity and discontinuity between these informal and formal contexts of learning, and how these can be seen as either indicators of connected learning or the ‘gap’ and how such interconnections are reflected in learners’ school engagement. the assumption of the gap between adolescents’ digital and academic engagement is not a new, it originates from prensky’s (2001) introduction of the controversial concept of digital natives (see also bennett & maton, 2010). the argument is that due to extended socialization in using socio-digital technologies, adolescents are often very comfortable with various socio-digital tools and applications and are able to fluently learn novel applications (hakkarainen et al., 2015). it is developmentally significant that young generations have cognitively socialized to a radically different social and technological environment than the older generations (wexler, 2006). the earlier and the more intensively young people adapt to the transforming cognitive, social, and cultural environment, the stronger the impact of this environment on their intellectual, emotional, behavioural, and social engagement is likely to be (moisala et al, 2016a; 2016b; ritella & hakkarainen, 2012). the gap hypothesis stems from the idea that in schools, digital immigrants, who are not sharing the same ‘language’, are teaching digital natives (hakkarainen et al, 2015; prensky, 2001). although prensky’s digital natives are teachers of today, it is suggested that the gap between some students’ progressive use of digital media outside of the classroom and the traditional pedagogies of most schools is still growing (ito et al., 2013). however, the gap can emerge for various reasons and can reflect various psychological processes. for instance, the students’ out-of-school interests and competencies may not be socially recognized leading to experiences of withdrawal and disengagement (rajala, kumpulainen, hilppö, paananen, & lipponen, 2015). it is also possible that the students’ out-of-school practices of working with learning and knowledge are critically different from the traditional practices of school (kumpulainen & sefton-green, 2012, mcfarlane, 2015). such situation may cause discontinuities, for instance, between individual versus social learning, externally regulated teaching versus self-initiated inquiry learning, and working with pre-digested textbook context versus navigating through open knowledge and media spaces. despite the controversy over the original concepts of digital natives and digital immigrants, it seems that there indeed are gaps between connecting (digital) learning across in-school and out-of-school contexts. 2. digital engagement and school engagement the early empirical findings tapping into the concept of digital natives revealed that students do not share the same experiences and competencies with digital media, ranging from students that engage in a wide range of digital activities do not participate in similar activities at all, with a spread of moderate participators in between (bennett & maton, 2010). a year-long ethnographic investigation of ito and colleagues (2010; see also barron, 2006) revealed diverging but partially overlapping genres of socio-digital participation. most adolescents use digital technologies for shallow friendship oriented hanging out with an extended network of friends. a much smaller proportion of young people use the emerging socio-digital technologies for pursuing their interests messing around with like-minded peers at social and digital networks. a significant but relatively small group of young people are geeking out by developing their technological and creative socio-digital competences (li, hietajärvi, palonen, salmela-aro, & hakkarainen, 2017). presumably interest-driven socio-digital participation fosters learning and development of young people and may assist in cultivating considerable student expertise (olson & bruner, 1996) concerning digital learning and activity. by relying on bereiter and scardamalia’s (1993) notion of progressive problem solving and hatano and inagaki’s (1992) adaptive expertise, hakkarainen and colleagues (2000) constructed measures for assessing to what extent young people have developed such crucial aspects of student expertise as putting effort to using digital technologies at the edge of competences and enjoying working with challenging problems with digital technologies. hence, a significant proportion of young people pursue their interest online, are supported by their more competent peers and have learned considerable digital competences through intensive socio-digital participation. it is also typical for young people to learn through active, although not always very deep, personal and social exploration rather than learn by passively consuming pre-determined information (hietajärvi, seppä, & hakkarainen, 2016; li et al, 2017). from the viewpoint of gap hypothesis, it is claimed that active socio-digital participators and especially those who have developed sophisticated peer learning and digital competences in informal contexts, may not get sufficient social recognition of their capabilities and, therefore, might feel alienated and experience mismatch between their personal and school practices, indicating inadequate person-environment fit. consequently, this has been suggested to be among the factors contributing to lower school engagement of those actively digitally engaging students, pointing to the gap between adolescents’ digital and school-related engagement (see, e.g., prensky 2001; halonen, hietajärvi, lonka ,& salmela-aro, 2017; kumpulainen & sefton-green, 2012; salmela-aro, muotka, alho, hakkarainen, & lonka, 2016a; selwyn, 2006). empirically, there appears, however, to be both continuities and discontinuities between young peoples’ digital learning engagement and school engagement. case studies have described students' informal interest-driven learning practices that can both facilitate and obstruct academic engagement (deng, connelly & lau, 2016; gurung & rutledge, 2014), and have also suggested that integrating practices of informal digital learning engagement in schoolwork can enhance student engagement (clements, 2015; esteves, 2012). larger scale quantitative studies, although scarcer, point to a similar direction: digital participation is related to both self-directed learning and student engagement (laird & kuh, 2005; rashid & ashgar, 2016). previous studies supporting the gap hypothesis, in turn, have suggested that that students’ reporting more cynicism towards school also reported that they would be more engaged in their schoolwork if they were able to use more digital technologies (halonen et al., 2017; salmela-aro et al., 2016a). hakkarainen and colleagues (2000) already reported similar finding among a large sample of finnish primary and secondary students. yet, other studies offer both positive and negative relations between out-of-school digital learning engagement and student engagement depending on the actual activities (hietajärvi, salmela-aro, tuominen, hakkarainen, & lonka, 2019; junco, 2012a, 2012b). in terms of adopting digital pedagogies in school, previous studies indicate that integrating digital technologies and media in education in general appears to offer mainly positive results regarding engagement and performance (chen, lambert, & guidry, 2010; junco, 2011; sung, chan, & liu, 2016; tamim, bernard, borokhovski, abrami, & schmid, 2011). 3. the conceptual framework in the present study, the gap hypothesis is approached from the framework of connected learning (ito et al, 2013; kumpulainen & sefton-green, 2012). connected learning refers to the process of connecting adolescents’ self-regulated and interest-driven learning (barron, 2006) across formal and informal contexts, in the reciprocal interactive processes between the learners and their social ecologies (nardi & o’day, 2000). connected learning is anchored to interest rather than mere friendship-driven socio-digital participation (ito et al., 2010). connected learning emerges when young people find contexts for pursuing their interests, network with like-minded peers, and when academic institutions recognize the value of informally developed knowledge and competences and allow interest-driven learning to be relevant in school (ito et al., 2013). further, it involves extensive peer-to-peer learning and supports learning processes relevant for academic achievements, civil activity, and, perhaps, also for professional career. connected learning takes into account a wide range of learning contexts, whereas in the present study, we focused on digital learning engagement as the out-of-school learning component. more precisely, we conceptualize digital learning engagement as reflecting informally emerging socio-digital participation (hakkarainen, hietajärvi, alho, lonka, & salmela-aro, 2015) including a considerable degree of self-regulated (boekarts & minnaert, 1999, see also panadero & järvelä, 2015) learning embedded on the contexts of their peer supported interest-driven learning ecologies (barron, 2006). it needs to be noted that although self-regulated learning is an internal process, it is embedded in digital and social context and environment and it is assisted and influenced by social interaction (panadero & järvelä, 2015). as such, we consider digital learning engagement as situated within the ecologies of connected learning. specifically, we operationalized digital learning engagement with two constructs. more precisely, we operationalized digital learning engagement as being expressed through having a digital learning preference, a preference for cultivation of adaptive student expertise concerning digital learning and problem-solving (hakkarainen et al, 2000) as well as showing different degrees of wish for digital schoolwork, that is, wish for connecting this digital learning to the context of school. in addition, in this study we extend the gap hypothesis to contribute to the wider research on school engagement (fredricks, blumenfeld & paris, 2004), operationalized as schoolwork engagement (upadyaya & salmela-aro, 2013), a generally positive disposition towards school and schoolwork that is considered a key outcome and indicator of connected learning (ito et al., 2013). specifically, we utilized the framework of connected learning to provide novel empirical information of the antecedents of school engagement combining it with the framework of demands-resources model extended to academic well-being (salmela-aro & upadyaya, 2014). in doing so, the present study attempts to provide insights into the psychological processes underlying successful experiences of connected learning, or, conversely, an experience of disengagement. according to this model, possible relations between digital learning engagement and school engagement can be viewed as resulting from the balance between the psychological demands of the situation and the resources available to overcome these demands, conceptualized over two processes, the energy-depleting process and the motivational process (demerouti, bakker, nachreiner & schaufeli, 2001; salmela-aro & upadyaya, 2014). digital learning engagement can function over both these processes by increasing the demands and depleting energy (e.g., multitasking, interruptions, cognitive load) or providing extended resources cultivating engagement (e.g. knowledge building and utilization, peer support) (barron, 2006; ito et al., 2013; salmela-aro & upadyaya, 2014). congruence between digital learning engagement and school engagement can be conceptualized as a condition of successful connected learning, in which the resources gained in out-of-school digital activities are successfully connected to in-school learning. in turn, the gap can be used to conceptualize a condition of discontinuity where these interest-driven digital learning practices collide with, for instance, strong external regulation and teacher-centered practices, and informally developed competencies and practices of learning are not utilized in school. similar discontinuity may follow if students, in their informal practices, have learned to rely on intensive peer-to-peer learning but are expected to work mostly alone at school. these discontinuities and contradictions in the possibilities to utilize the resources gained in out-of-school learning may consequently decrease students’ engagement with schoolwork. 4. research aim and hypotheses previous research indicates that active digital learning engagement is related to both positive and negative school-related outcomes, but there is little knowledge concerning the conditions by which these positive or negative outcomes come to be. moreover, most previous studies were conducted in higher education context, and no longitudinal designs have come to our knowledge. therefore, the present study was conducted in upper comprehensive school and with a longitudinal design. more precisely, the present study empirically focuses on the question of how digital learning engagement is related to school engagement both cross-sectionally and over time, while taking into account how wishing to use more, or less, digital technologies in schoolwork moderates the longitudinal relationship. the present research aimed to examine processes of continuity and discontinuity between out-of-school digital learning engagement and school engagement, utilizing both the broader framework of connected learning combined with the demands-resources model (salmela-aro & upadyaya, 2014) as tools to interpret the processes underlying the gap. towards that end, by combining the framework of connected learning with the demands-resources model (salmela-aro & upadyaya, 2014) we expected that (hypothesis 1) having a digital learning preference would be reflected as also having a positive disposition towards school indicated by a positive relation to schoolwork engagement (both cross-sectionally and longitudinally). the positive relation was expected due to the supposedly increased psychological resources (salmela-aro & upadyaya, 2014) resulting from the informal self-regulated and connected learning happening in the process (barron, 2006, see also hietajärvi et al., 2019). in contrast, we expected that (hypothesis 2) reporting a higher wish for digital schoolwork, that is, experiencing a discontinuity between the practices of one’s informal interest-driven learning and academic learning, would be negatively related to schoolwork engagement both cross-sectionally and longitudinally (kumpulainen & sefton-green, 2012). in particular, we hypothesized that (hypothesis 3) the negative relation from a wish for digital schoolwork and school engagement would be explained especially by interaction between digital learning preference and the wish for digital schoolwork, resulting in a condition of discontinuity. in other words, we expected that the combination of having a digital learning preference without the possibility to connect it to academic learning would lead to lower school engagement. 5. method 5.1 participants there was a total of 1,705 (43.7% male) participants from 27 schools in helsinki. the data were collected annually in spring following the same students across grades 7 (age ~14 yrs., n = 1272), 8 (age ~15 yrs., n = 1150) and 9 (age ~16 yrs., n = 903) over the upper comprehensive school. of all participants 1,090 (64%) participated in the study at least at two time points over the data collection period and 530 (31%) participated in all the waves. the participants filled in a self-report questionnaire on digital engagement and academic well-being. participation in the study was voluntary and informed consent forms were collected from the students and from their parents. data collection was organized as a convenience sample, that is, all teachers in the schools that were able to organize data collection administered the questionnaires during school hours and all students that were attending at the time of data collection and were willing to take the questionnaire were included as participants. because of the data collection procedure the reasons for attrition may be due to either the schools or the teachers’ inability to incorporate the data collection into their timeframe, the students being absent during the data collection or unwilling to respond. despite this, the number of students that participated in at least two waves was satisfactory. the study protocol was approved by the university of helsinki ethical review board in the humanities and social and behavioural sciences. 5.2 measures the means, standard deviations and measures of internal consistency for all constructs used in this study are shown in table 1. 5.2.1 digital learning engagement digital learning engagement was conceptualized as, on one hand, showing orientation towards learning with technologies in general, and on the other hand, expressing enthusiasm towards using more socio-digital technologies in formal schoolwork. these were measured with two constructs both measured on a scale from 1 (= not at all true) to 5 (= very true). digital learning preference (dlp; hakkarainen et al., 2000; see also halonen et al., 2017) was measured with four items that assessed having a preference towards learning and solving problems with digital technologies (e.g. “it’s fun to learn to use digital technologies, because it offers continuously new challenges”). rather than merely assessing interest in using digital technology, the items have been designed so that they trace students’ orientation toward learning with and about digital technology, and, thereby related to cultivation of adaptive student expertise concerning digital learning and problem-solving. wish for digital schoolwork (wds; hakkarainen et al., 2000; see also halonen et al., 2017; salmela-aro et al, 2016a) was measured with three items that directly tapped into the gap hypothesis by assessing motivation and possibilities towards using more digital technologies in school and its perceived effect on school engagement with three items (e.g., “i’m more engaged in my schoolwork when i’m able to use digital technologies”). in other words, the scale assessed whether or not the student favoured the use of more digital technologies in schoolwork. higher scores indicated a stronger wish towards using more technologies in schoolwork. table 1 raw means, standard deviations and measures of internal consistencies of the constructs used in the models 5.2.2 school engagement school engagement was assessed using the schoolwork engagement inventory (i.e., eda abbreviated from energy, dedication, and absorption; salmela-aro & upadyaya, 2012) measuring a trait-like long-term study-related positive state of mind. the inventory consists of three subscales, each including three items, measuring energy (e.g., “when i study, i feel i’m bursting with energy”), dedication (e.g., “i am enthusiastic about my studies”), and absorption (e.g., “time flies when i’m studying”). however, schoolwork engagement is often specified as a unidimensional measurement model representing a generally positive study-related frame of mind (salmela-aro & upadyaya, 2012). the items were rated on a scale ranging from 1 (‘never’) to 7 (‘every day’). 5.3 analysis strategy we followed an analysis strategy in which we first ran preliminary analyses for screening the data and ensuring that the latent constructs we used carried the same meaning across gender and time. second, we analysed gender differences in the means of latent constructs. third, to test hypotheses 1 and 2, we examined the partial correlations between our latent constructs separately in each point of measurement and specified the longitudinal model. finally, to test hypothesis 3 we added the latent interaction to the longitudinal model. the more detailed steps in the analysis procedure are described as follows. first, as preliminary analysis, the data were screened for the number and patterns of missing values using the ibm statistical package for social sciences, version 25 (spss). the missing values were assessed longitudinally and at each time point separately. second, we specified and tested the measurement model using a confirmatory factor analysis approach (cfa). residuals of the same items were allowed to be correlated over time. the analyses were conducted using mplus 8.0 (muthén & muthén, 2018) in conjunction with r and rstudio (r core team, 2018) with the package mplusautomation (hallquist & wiley, 2018). maximum likelihood with standard errors robust for non-normality (mlr) was used as the estimator and missing data was handled with full information maximum likelihood estimation (fiml). the complex survey data option (muthén & muthén, 2018; see also asparouhov & muthén, 2006; muthén & satorra, 1995) was used in all analyses to correct for non-independence at the class level. invariance of the measurement model across the factor structure (configural), factor loadings (metric) and item intercepts (scalar) was tested to ensure that the measures held the same meaning across gender and over time. the model fits (see e.g. hu & bentler, 1998) were evaluated based on the chi-square value as well as the root mean square error of approximation (rmsea) with an approximate acceptable cutoff value of less than .08, standardized root mean residual (srmr) with an approximate cutoff or less than .08, and, incremental indexes such as the comparative fit index (cfi) and the tucker-lewis index (tli) with approximate acceptable cutoff values of greater than .9. in evaluating measurement invariance, we relied on the conventional criteria of evaluating change in rmsea and cfi (chen, 2007; for more discussion on measurement invariance testing, see putnick & bornstein, 2016). finally, after confirming that the measurement model represented sufficiently the same constructs across both gender and time, we explored mean differences across gender by regressing each latent factor on gender. to test the cross-sectional parts of hypotheses 1 and 2, that is, to evaluate how digital learning engagement and school engagement are related within a time point, we explored the gender-controlled partial correlations between the latent variables. this was done by visualizing the latent variable partial correlations by plotting the variables as nodes in a ebicglasso-regularized network (epskamp & fried, 2018) using the r-package qgraph (epskamp, cramer, waldorp, schmittmann, & borsboom, 2012). partial correlations are presented, so the edges in the latent partial correlation network can be interpreted similarly as regression path coefficients, as they are controlled for gender as well as each other, but without assuming any direction of effects. this type of modelling allows for a powerful measurement error corrected modelling and visualization of contemporaneous relations between latent variables when the direction of effects cannot be inferred from the data (guyon, falissard, & kop, 2017). then, to test the longitudinal parts of hypotheses 1 and 2, that is, to evaluate how digital learning engagement and school engagement are related over time, we specified latent longitudinal panel models (l-clpm; little, preacher, selig & card, 2007). the clpm is especially useful for identifying the relations between variables across time and can be applied to identify a possible causal relationship between variables measured at different time points. the clpm accounts for stability over time through the inclusion of autoregressive parameters. more precisely, the autoregressive effects describe the stability of individual differences from one measurement point to the next, whereas the cross-lagged effects describe the effect of a variable on another measured at a later occasion. taking into account autoregressive effects, cross-lagged effects in the present study can be interpreted as predicting change over time (selig & little, 2012). moreover, our models were specified using a latent measurement model, so that the variables were free of measurement error (little et al., 2007). further, in our models mean differences across gender were controlled for by regressing each latent variable on gender. the model was specified with both autoregressive and crossed paths specified between successive time points and the paths from time 1 to time 2 were constrained equal with paths from time 2 to time 3 to achieve a simpler model. the cost of acquiring a simpler model in comparison to an unconstrained model was evaluated by examining the change in fit indices rmsea and cfi as well as the bayesian information criterion (bic), which penalizes complexity (raftery, 1995). finally, to test hypothesis 3, that is, to examine the presence of a condition of discontinuity, we included latent interactions between digital learning preference and a wish for digital schoolwork as predictors of schoolwork engagement. the latent interactions were estimated using the latent moderated structural equations (lms) approach (klein & moosbrugger, 2000) implemented in mplus (muthén & muthén, 2018) as a maximum likelihood-based approach, which, in general can be viewed as recommendable (see e.g. marsh, wen & hau, 2004). 6. results 6.1 preliminary results there were less than 10% missing values overall at each time point. based on little’s mcar test, the data were not missing completely at random longitudinally (χ2 (7346) = 7702.84, p = .003). looking at each time point separately the data were missing completely at random at time 1 and time 2 (χ2 (805) = 823.04, p = .322; t2: χ2 (488) = 527.56, p = .105), whereas at time 3 the mcar assumption did not hold (χ2 (388) = 509.86, p < .001). 6.1.1 measurement model first, we specified a baseline measurement model in the first time point and continued to test for measurement invariance across gender. the baseline model (see figure 1) fitted the data acceptably (χ2(101)=611.21, χ2 scaling correction factor (cf) = 1.20, p < .001, rmsea = .063, cfi = .956, tli = .948, srmr = .032). figure 1. the baseline measurement model for time 1 with unstandardized factor loadings. we then proceeded to evaluate the measurement invariance across gender and time (model fit indices as well as the factor loadings and r2 for the scalar longitudinal measurement model are presented in the appendices). the results indicated that there were no considerable differences in the measurement model between boys and girls and that the constructs we measured did not change in their meaning over time (for details see appendix a). then, to control for mean differences across gender, we regressed each latent factor on gender. the model fit did not decline and the modification indices did not suggest any direct effects between the factor indicators and gender. regarding mean differences across gender (see table 2) the model indicated that male participants scored higher in digital learning preference than female participants and wished for digital schoolwork at each time point, whereas in schoolwork engagement there were no gender differences. table 2 mean differences across gender in the latent variables note: gender code: 0, ‘female’; 1, ‘male’. the estimate to be interpreted as how the ‘male’ group differs from the ‘female’ group. 6.2 cross-sectional relations between digital learning engagement and school engagement to answer how digital learning preference and a wish for digital schoolwork are related to schoolwork engagement cross-sectionally, we examined the contemporaneous partial correlations between the latent variables. the correlation coefficients were extracted from the latent measurement model with scalar invariance constraints and gender as a covariate. the latent variables showed a similar pattern of relations with each other at all time points as can be inferred from figure 2 (for zero-order correlations see appendix c). the results indicated that when controlled for each other, digital learning preference was positively related to both a wish for digital schoolwork and schoolwork engagement, supporting hypothesis 1. however, when controlled for digital learning preference, a wish for digital schoolwork was negatively related to schoolwork engagement, supporting hypothesis 2. figure 2. cross-sectional gender-controlled latent variable partial correlation networks. note: dlp, digital learning preference; wds, wish for digital schoolwork; eda, schoolwork engagement. network estimated with ebicglasso regularization (see epskamp & fried, 2018). nodes placed by fruchterman-reingold algorithm (fruchterman & reingold, 1991). blue indicates positive correlations, red negative, and the width and colour of the edges correspond to the absolute value of the correlations: the stronger the correlation, the thicker and more saturated the edge (see espkamp et al., 2012). this pattern of partial correlations gives a reason to suspect that the effect of digital learning preference on schoolwork engagement might be moderated by a wish for digital schoolwork, in which case there would be discontinuity between out-of-school digital learning and the possibilities to connect this to school. in other words, how the digital learning preference is related to school engagement would depend on the participants’ motivation and possibilities in using technologies in school, or the students’ personal digital learning practices’ fit with the pedagogical practices at school. 6.3 longitudinal relations between socio-digital participation and academic well-being to answer how digital learning preference and a wish for digital schoolwork are related to schoolwork engagement longitudinally, we specified the latent longitudinal panel model. the model fitted the data well (χ2 (1096) = 3151.59, cf = 1.11, p < .001, rmsea = .034, cfi = .944, tli = .940, srmr = .038) and the model fit did not differ considerably from the unconstrained measurement model. in addition, bayesian information criterion (bic) favoured the model (bic = 142676.60) over the more complex unconstrained model (bic = 142703.41). we then included the hypothesized latent interaction. the interaction term was statistically significant and resulted in a slightly better fit as indicated by the log-likelihood chi-square difference (∆χ2 = 7.23, df = 1, p = .007). thus, the results presented are from the model with stationarity assumed and the latent interaction included. all unstandardized structural model parameters are presented in table 3 and statistically significant structural parameters are illustrated in figure 3. table 3 autoregressive and cross-lagged parameters and latent interactions of the longitudinal panel model the model indicated that, first, the effects of the same construct on itself over time were moderately strong and somewhat carried over for two years, indicating that the constructs are quite stable over time. beyond these autoregressive effects the model also revealed that digital learning preference predicted higher schoolwork engagement, supporting hypothesis 1. a wish for digital schoolwork had only a weak and statistically insignificant negative effect on schoolwork engagement. however, their interaction predicted schoolwork engagement negatively in line with hypothesis 3. figure 3. structural parameters of the longitudinal panel model. note: cross-lagged paths with positive coefficients highlighted in blue, negative in red. only paths significant at p < .05 illustrated for clarity. gender effects omitted for clarity. a closer inspection of the interaction (see figure 4) indicated that a wish for digital schoolwork predicted lower future schoolwork engagement only for those students that had reported a higher digital learning preference. for those students reporting lower digital learning preference there appeared to be no effect between a wish for digital schoolwork and schoolwork engagement. figure 4. schoolwork engagement regressed on wish for digital schoolwork with differing levels (±1 sd) of digital learning preference. note: the blue line represents high digital learning preference (+1sd), the red line low digital learning preference (-1sd); dashed lines represent 95% confidence intervals. 7. discussion the present study aimed to investigate connected learning in terms of examining the conditions of continuity and discontinuity emerging between students’ self-regulated and interest-driven informal learning in the digital contexts and their school engagement. previous research has shown that contextual factors such as parental affect, teacher support and mastery-supportive classroom atmosphere promote higher school engagement (upadyaya & salmela-aro, 2013), and the present study sheds light on how orientation toward learning through digital tools and orientation toward developing socio-digital competencies might be related to the equation. the present results support the gap hypothesis, but also reveal signs of connected learning, connecting digital learning engagement with schoolwork engagement. more precisely, we expected that having an digital learning preference would be reflected in having a positive study-related state of mind (hypothesis 1), that wishing to use more digital tools in schoolwork would be negatively related to school engagement (hypothesis 2), and that especially the combination of an individual disposition towards digital learning and a lack of contextual support would lead to lower school engagement (hypothesis 3) and thus represented a condition of discontinuity. the present cross-sectional results give support for all our hypotheses, whereas the present longitudinal data only supports hypotheses 1 and 3. cross-sectionally, we observed that digital learning preference is positively correlated to both a wish for digital schoolwork and schoolwork engagement. wishing for more digital schoolwork, in turn, was negatively correlated to schoolwork engagement. longitudinally, digital learning preference predicted higher schoolwork engagement across time and the interaction between digital learning preference and a wish for digital schoolwork predicted later schoolwork engagement negatively. the directions of the effects appeared very clear: schoolwork engagement did not predict increases or decreases in either digital learning preference or wish for digital schoolwork, indicating that these indeed work as antecedents for school engagement. the present results might be interpreted in various ways. from the viewpoint of connected learning framework, the positive association between digital learning engagement and school engagement might indicate successes in connecting learning across informal and formal contexts. this is understandable because the items in the digital learning preference scale trace the participants’ orientation toward cultivating their adaptive expertise of learning through digital technologies; such efforts may generalize from digital to other spheres of learning. this is also in line with previous studies showing a positive relation between digital and academic engagement (laird & kuh, 2005; rashid & ashgar, 2016), and gives support to the previous findings with a longitudinal component included. thus, the present results also indicate that adopting more sophisticated digital practices and competencies might help students build novel resources for schoolwork, thus contributing to higher school engagement over time in line with the demands-resources model (hietajärvi et al., 2019; salmela-aro & upadyaya, 2014). such resources could involve students’ spontaneous use of digital technologies for academic purposes, such as seeking school-relevant knowledge from the internet or reciprocally helping one another in schoolwork (hietajärvi, seppä, & hakkarainen, 2016, li et al., 2017). further, given that especially students with high digital learning preference and low wish for digital schoolwork experienced higher later school engagement, we can argue that informal digital learning engagement, if recognized and taken into account in schools, can foster connected learning (ito et al., 2013). data from our related studies indicate, however, that digital technologies were used rather infrequently at school and mostly for pretty basic purposes at the time of collecting the present data (halonen et al., 2017). in previous research, especially those digitally oriented students who appeared to feel inadequate and alienated at school wished for an opportunity to use more digital technologies at their schoolwork (salmela-aro et al., 2016a). the results of the present study also revealed evidence of the gap, indicating that connecting learning across contexts is challenging. the results of this study indicate that not being able to incorporate the prior experiences and practices of students into the formal learning environment creates, for some students, an experience of discontinuity contributing to feelings of disengagement (rajala et al, 2015). these were students who reported to have cultivated high levels of digital expertise, and could even be described as ‘geeking out’ (ito et al., 2010). for other, digitally less engaged students, wish for digital schoolwork appeared irrelevant regarding changes in school engagement over time. taken together, both the cross-sectional and longitudinal results showed support for a more nuanced understanding of the gap hypothesis: students who, on one hand, hold a disposition for learning with and about digital tools, but on the other hand, experience being not capable of deploying this competence in school, experience decline in their school engagement. conversely, it is noteworthy that holding a stronger orientation towards learning with digital tools contributed to a higher school engagement possibly due to the increased psychological resources gained in the process, given that the students’ needs in terms of digital schoolwork are in congruence. 7.1 methodological reflections and limitations the present study used a large longitudinal sample of adolescents, which strengthens the inferences that can be drawn from the analyses. however, the sample was not representative; it was a convenience sample collected from helsinki and, therefore, cannot be generalizable across finland or beyond. a replication study with a representative sample, possibly with students of different ages, various parts of finland and from different academic contexts (high school, vocational school), would be needed. the data were based on self-reports, so the actual digital participation practices of the students were not assessed. further, the present investigation addressed mostly young people’s informal digital activities because out-of-school socio-digital participation was far more intensive that within school one. thus, we were not able to actually trace the actual behaviour of participants either with digital technologies or in school; nor were we able to actually examine the pedagogical practices or technology use in school. we cannot, for instance, say anything about how students with different levels of digital learning preference approach and behave when working on learning tasks or what kind of learning tasks they are presented with, including digital technologies or not. we acknowledge that this can make all the difference and thus future studies should include multi-level data of students’ informal and formal learning activity. moreover, future studies should better approach the qualitative differences in adolescents’ digital learning engagement, for instance, by mixed methods data on their interest-driven pursuits, and why and how they engage with digital technologies to support them (for this kind of initial efforts, see hietajärvi et al., 2016; li et al., 2017). further, school engagement is but a one indicator of academic functioning, and we should look into different ways of conceptualizing the emergence of connected learning or the gap, future studies should also take into account academic achievement, as well as indicators other than school engagement regarding motivation and well-being. the cross-lagged panel model we used to analyse the longitudinal relations allowed us to examine how digital learning engagement predicted change in school engagement over time, and as such gives us grounds to present inferences regarding the temporal ordering on these effects with a large number of participants. the methodological choice of using latent variables allowed us to model the relations without measurement error (e.g., little et al., 2007) but we did not separate between-participant and within-participant variances, which affects the inferences that we can draw, especially limiting stronger causal inferences (hamaker, kuiper, & grasman, 2015). the inclusion of the latent interaction allowed us to model the gap conditional between individual and perceived contextual factors, but without a multi-level setting and actually tracing the school-level practices we cannot truly establish detailed aspects of the postulated gap between the students’ informal digital practices and the schools’ digital pedagogical practices. there can be considerable differences between teachers, schools, districts and countries in how digital tools are implemented in teaching and learning. consequently, the gap is likely to be more observable in some schools than others and with some students rather than others. moreover, it appears that the interplay between students’ digital and academic engagement is complex rather than straightforward. national efforts of digitalizing practices of learning and teaching in finland are likely to have significant impact on future manifestations of the gap to be revealed by collecting empirical data. after we carried out the present study, finnish matriculation examination (the only high-stake test in finland) has been digitalized and major efforts of digitalization of school are underway so as to meet societal challenges and overcome the gap. our research network has developed novel self-report instruments for tracing in details the extent and focus of within school use of digital technologies and associated pedagogic approaches together with young people’s informal socio-digital practices. further, looking for the gap hypothesis only through students who use more, or less, digital technologies or how much they would like to use digital technologies does not appear to be fruitful. it is crucial to collect detailed longitudinal data of students’ transforming informal and formal ecologies of socio-digital participation that are likely to change from one cohort to the next. moreover, investigating relations between digital and school engagement calls for investigating qualitatively different levels (or genres) of socio-digital participation because interest-driven creative and academic practices are more likely to foster young people’s learning and development. moreover, students come from different backgrounds and contexts, which is reflected in how they experience digital learning engagement in terms of learning and how these experiences collide or connect with the educational practices of school (howard, ma, & yang, 2016; ito et al., 2013). this can position students unequally and contribute to a digital competence gap. for instance, investigations of barron and colleagues (2009) indicate that students who have cultivated advanced levels of digital competence often come from advantaged homes and have parents who provide structured support for school as well as digital learning, and foster the development interest-driven technical and creative capabilities. disadvantaged students, in turn, may not only have limited parental guidance of the school activity but also restricted access to tools, practices, or social support relevant for building advanced digital competences increasing dangers of educational exclusion or exposure to maladaptive patterns of digital engagement. to counter these risks, all students should be provided with tools to cultivate their digital practices and capitalize on the connected learning possibilities. from the schools’ viewpoint, the problem to be solved is how the pedagogical solutions around digital tools are implemented so that the students’ out-of-school practices are acknowledged to best support students’ personal and collaborative learning and development within a network of connected learning. toward that end, it appears critical to engage students learning by using digital technologies for sustained collaborative effort of building and creating knowledge and media (paavola & hakkarainen, 2014). 8. conclusion our results are among the first to directly assess the gap hypothesis with larger scale quantitative and longitudinal data. the results indicate that the gap hypothesis was supported under the following condition: students who express a disposition to solve problems and learn with digital technologies out of school and, would prefer to use more technologies for learning in school, experience a discontinuity between out-of-school and in-school learning and report lower later school engagement. however, the results also indicate that students who hold a disposition towards digital learning but do not experience a discontinuity in connecting it to schoolwork, experience higher later school engagement. the finding that digital learning preference predicted higher school engagement provides indirect evidence for some aspects of connected learning in schools instead of the gap. based on the results we emphasize that the manifestation of the gap vs. connected learning is dependent on multiple factors, both individual (the level of digital and school engagement) and contextual (the prevailing digital-pedagogic practices of schools). these need to be taken into account in future research. multiple methodologies and richer data sets are needed to reveal more about the adolescents’ truly connected learning experiences – or the lack of them. this study was one step forward in demonstrating and understanding the complexity of this issue. there is an obvious need to examine the reciprocal interactive processes between the learners and their social ecologies inside and outside school more closely in order to support the intellectual development and school engagement of our youth. towards that end it seems essential to also enhance the educational practices in schools so that the informal learning gained in out-of-school digital engagement can be recognized and supported across all students. keypoints cross-sectional and longitudinal results showed both continuity and discontinuity in connecting out-of-school digital learning and school. digital learning preference was related to higher schoolwork engagement. wish for digital schoolwork was related to lower schoolwork engagement. students who preferred digital learning experienced increased schoolwork engagement over time, especially when connected to in-school digital schoolwork. students who preferred digital learning but did not have sufficient possibilities to connect it to schoolwork, experienced decline in schoolwork engagement. acknowledgements this research was funded by: the academy of finland project “mind the gap—between digital natives and educational practices,” pi kirsti lonka (grant #265528); the academy of finland project “bridging the gaps—affective, cognitive, and social consequences of digital revolution for youth development and education,” pi katariina salmela-aro (grant #308351) and pi kirsti lonka (grant #308352); the strategic research council project “growing mind: educational transformations for facilitating sustainable personal, social, and institutional renewal in the digital age,” pi kai hakkarainen (grant #312527) and team leader kimmo alho (grant #312529); and, the academy of finland project “#agents – young people’s agency in social media,” pi katariina salmela-aro (grant # 320371). references asparouhov, t., & muthen, b. (2006). comparison of estimation methods for complex survey data analysis. mplus web notes. url: https://www.statmodel.com/download/surveycomp21.pdf barron, b. (2006). interest and self-sustained learning as catalysts of development: a learning ecology perspective. human development, 49, 193–224. doi: 10.1159/000094368 barron, b., martin, c. k., takeuchi, l., & fithian, r. (2009). parents as learning partners in the development of technological fluency. international journal of learning and media, 1, 55–77. doi: 10.1162/ijlm.2009.0021 bennett, d. a. (2001). how can i deal with missing data in my study? australian and new zealand journal of public health, 25, 464–469. doi: 10.1111/j.1467-842x.2001.tb00294.x bennett, s. & maton, k. (2010). beyond the “digital native” debate: towards a more nuanced understanding of students’ technology experiences. journal of computer assisted learning, 26, 321–331. doi: 10.1111/j.1365-2729.2010.00360.x bereiter, c. & scardamalia, m. (1993). surpassing ourselves: an inquiry into the nature and implications of expertise. chicago, il: open court. boekaerts, m., & minnaert, a. (1999). self-regulation with respect to informal learning. international journal of educational research, 31, 533–544. doi: 10.1016/s0883-0355(99)00020-8 chen, f. f. (2007). sensitivity of goodness of fit indices to lack of measurement invariance. structural equation modeling, 14, 464–504. doi: 10.1080/10705510701301834 chen, p. s. d., lambert, a. d., & guidry, k. r. (2010). engaging online learners: the impact of web-based learning technology on college student engagement. computers & education, 54, 1222–1232. doi: 10.1016/j.compedu.2009.11.008 clements, j. c. (2015). using facebook to enhance independent student engagement: a case study of first-year undergraduates. higher education studies, 5, 131–146. doi: 10.5539/hes.v5n4p131 demerouti, e., bakker, a. b., nachreiner, f., & schaufeli, w. b. (2001). the job demands-resources model of burnout. journal of applied psychology, 86, 499–512. doi: 10.1037/0021-9010.86.3.499 deng, l., connelly, j., & lau, m. (2016). interest-driven digital practices of secondary students: cases of connected learning. learning, culture and social interaction, 9, 45–54. doi: 10.1016/j.lcsi.2016.01.004 epskamp, s., cramer, a. o., waldorp, l. j., schmittmann, v. d., & borsboom, d. (2012). qgraph: network visualizations of relationships in psychometric data. journal of statistical software, 48, 1–18. doi: http://dx.doi.org/10.18637/jss.v048.i04 epskamp, s., & fried, e. i. (2018). a tutorial on regularized partial correlation networks. psychological methods, 23(4), 617-634.. doi: 10.1037/met0000167 epskamp, s., rhemtulla, m., & borsboom, d. (2017). generalized network psychometrics: combining network and latent variable models. psychometrika, 82, 904–927. doi: 10.1007/s11336-017-9557-x esteves, k. k. (2012). exploring facebook to enhance learning and student engagement: a case from the university of philippines (up) open university. malaysian journal of distance education, 14. eu kids online (2014) eu kids online: findings, methods, recommendations (deliverable d1.6). eu kids online, london, united kingdom: london school of economics. european commission. (2017). the digital competence framework 2.0. retrieved from https://ec.europa.eu/jrc/en/digcomp/digital-competence-framework/ european parliament. (2015). innovative schools: teaching and learning in the digital era workshop documentation. brussels, belgium: european parliament. fredricks, j.a., blumenfeld, p.c., & paris, a.h. (2004). school engagement: potential of the concept, state of the evidence. review of educational research, 74, 59–109. doi: 10.3102/00346543074001059 fruchterman. t.m.j. & reingold, e.m. (1991). graph drawing by force-directed placement. software: practice and experience, 21, 1129–1164. doi: 10.1002/spe.4380211102 gurung, b., & rutledge, d. (2014). digital learners and the overlapping of their personal and educational digital engagement. computers & education, 77, 91–100. doi: 10.1016/j.compedu.2014.04.012 guyon, h., falissard, b., & kop, j.-l. (2017). modeling psychological attributes in discussion: network analysis vs. latent variables. frontiers in psychology, 8, 798. doi: 10.3389/fpsyg.2017.00798 hakkarainen, k. (2009). a knowledge-practice perspective on technology-mediated learning. international journal of computer-supported collaborative learning, 4, 213–231. doi: 10.1007/s11412-009-9064-x hakkarainen, k., ilomäki, l., lipponen, l., muukkonen, h., rahikainen, m., tuominen, t., et al. (2000). students’ skills and practices of using ict: results of a national assessment in finland. computers & education, 34, 103–117. doi:10.1016/s0360-1315(00)00007-5 hallquist, m. n. & wiley, j. f. (2018). mplusautomation: an r package for facilitating large-scale latent variable analyses in mplus. structural equation modeling, 1–18. doi: 10.1080/10705511.2017.1402334. halonen, n., hietajärvi, l., lonka, k., & salmela-aro, k. (2016). sixth graders’ use of technologies in learning, technology attitudes and school well-being. the european journal of social & behavioural sciences. 18, 2307–2324. doi: 10.15405/ejsbs.205 hamaker, e. l., kuiper, r. m., & grasman, r. p. (2015). a critique of the cross-lagged panel model. psychological methods, 20, 102. doi: 10.1037/a0038889 hatano, g. & inagaki, k. (1992). desituating cognition through the construction of conceptual knowledge. in p. light & g. butterworth (eds.), context and cognition. ways of knowing and learning (pp. 115–133). new york, new york: harvester. hietajärvi, l., salmela-aro, k., tuominen, h., hakkarainen, k., & lonka, k. (2019). beyond screen time: multidimensionality of socio-digital participation and relations to academic well-being in three educational phases. computers in human behavior, 93, 13–24. doi: 10.1016/j.chb.2018.11.049 howard, s. k., ma, j., & yang, j. (2016). student rules: exploring patterns of students’ computer-efficacy and engagement with digital technologies in learning. computers & education, 101, 29–42. doi: 10.1016/j.compedu.2016.05.008 hu, l. t., & bentler, p. m. (1998). fit indices in covariance structure modeling: sensitivity to underparameterized model misspecification. psychological methods, 3, 424. ito, m., baumer, s., bittanti, m., cody, r., herr-stephenson, b., horst, h. a., et al. (2010). hanging out, messing around, and geeking out. cambridge, massachusetts: the mit press. ito, m., gutiérrez, k., livingstone, s., penuel, b., rhodes, j., salen, k., ... & watkins, s. c. (2013). connected learning: an agenda for research and design. irvine, california: digital media and learning research hub. jenkins, h. (2009). confronting the challenges of participatory culture: media education for the 21st century. cambridge, massachusetts: mit press. junco, r. (2012a). the relationship between frequency of facebook use, participation in facebook activities, and student engagement. computers & education, 58, 162–171. doi: 10.1016/j.compedu.2011.08.004 junco, r. (2012b). too much face and not enough books: the relationship between multiple indices of facebook use and academic performance. computers in human behavior, 28, 187–198. doi: 10.1016/j.chb.2011.08.026 junco, r., heiberger, g., & loken, e. (2011). the effect of twitter on college student engagement and grades. journal of computer assisted learning, 27, 119–132. doi: 10.1111/j.1365-2729.2010.00387.x klein, a., & moosbrugger, h. (2000). maximum likelihood estimation of latent interaction effects with the lms method. psychometrika, 65, 457–474. doi: 10.1007/bf02296338 kumpulainen, k., & sefton-green, j. (2012). what is connected learning and how to research it? international journal of learning, 4, 7–18. doi: 10.1162/ijlm_a_00091 laird, t. f. n., & kuh, g. d. (2005). student experiences with information technology and their relationship to other aspects of student engagement. research in higher education, 46, 211–233. doi: 10.1007/s11162-004-1600-y li, s., hietajärvi, l., palonen, t., salmela-aro, k., & hakkarainen, k. (2017). adolescents’ social networks: exploring different patterns of socio-digital participation. scandinavian journal of educational research, 61, 255–274. doi: 10.1080/00313831.2015.1120236 little, t. d., preacher, k. j., selig, j. p., & card, n. a. (2007). new developments in latent variable panel analyses of longitudinal data. international journal of behavioral development, 31, 357–365. doi: 10.1177/0165025407077757 maccallum, r.c., browne, m.w., & cai, l. (2005). testing differences between nested covariance structure models: power analysis and null hypotheses. psychological methods, 11, 19–35. doi: 10.1037/1082-989x.11.1.19 malcolm, j., hodkinson, p., & colley, h. (2003). the interrelationships between informal and formal learning. journal of workplace learning, 15, 313–318. doi: 10.1108/13665620310504783 marsh, h. w., wen, z., & hau, k. t. (2004). structural equation models of latent interactions: evaluation of alternative estimation strategies and indicator construction. psychological methods, 9, 275. doi: 10.1037/1082-989x.9.3.275 mcfarlane, a. (2015). authentic learning for the digital generation: realising the potential of technology in the classroom. london: routledge. moisala, m., salmela, v., hietajärvi, l., carlson, s., vuontela, v., lonka, k., ... & alho, k. (2016a). gaming is related to enhanced working memory performance and task-related cortical activity. brain research , 1655, 204–215. doi: 10.1016/j.brainres.2016.10.027 moisala, m., salmela, v., hietajärvi, l., salo, e., carlson, s., salonen, o., ... & alho, k. (2016b). media multitasking is associated with distractibility and increased prefrontal activity in adolescents and young adults. neuroimage, 134, 113–121. doi: 10.1016/j.neuroimage.2016.04.011 muthén, l. k., & muthén, b. o. (2018). mplus: statistical analysis with latent variables: user's guide [version 8]. los angeles, california: muthén & muthén. muthén, b., & satorra, a. (1995). complex sample data in structural equation modeling. sociological methodology, 25, 267–316. doi: 10.2307/271070 nardi, b., & o’day, v. (2000). information ecologies: using technology with heart. cambridge, massachussets: mit. oecd. (2015). students, computers and learning: making the connection. paris, france: pisa, oecd publishing. doi: 10.1787/9789264239555-en olson, d.r. & bruner, j. s. (1996) folk psychology and folk pedagogy. in d. r. olson, d.r. & n. torrance (eds.) the handbook of education and human development. new models of learning, teaching and schooling (pp. 9–27). malden, massachusetts: blackwell publisher. orvis, k. l. (ed.). (2008). computer-supported collaborative learning: best practices and principles for instructors: best practices and principles for instructors. hershey, new york, new york: igi global. paavola s. & hakkarainen, k. (2014). trialogical approach for knowledge creation. in tan s-c., jo, h.-j., & yoe, j. (eds.), knowledge creation in education (pp. 53–72). singapore: springer. panadero, e. & järvelä, s. (2015). socially shared regulation of learning: a review. european psychologist, 20, 190–203. doi: 10.1027/1016-9040/a000226 prensky, m. (2001). digital natives, digital immigrants part 1. on the horizon, 9, 1–6. doi: 10.1108/10748120110424816 putnick, d. l., & bornstein, m. h. (2016). measurement invariance conventions and reporting: the state of the art and future directions for psychological research. developmental review, 41, 71–9. doi: 10.1016/j.dr.2016.06.004 r core team (2018). r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. url https://www.r-project.org/. raftery, a. e. (1995). bayesian model selection in social research. sociological methodology, 111-163. rajala, a., kumpulainen, k., hilppö, j., paananen, m., & lipponen, l. (2015). connecting learning across school and out-of-school contexts: a review of pedagogical approaches. in o. erstad, k. kumpulainen, å. mäkitalo, k. p. pruulmann-vengerfeldt, & t. jóhannsdóttir (eds.), learning across contexts in the knowledge society. (pp. 15-35) rotterdam, the netherlands: sense publishers. rashid, t., & asghar, h. m. (2016). technology use, self-directed learning, student engagement and academic performance: examining the interrelations. computers in human behavior, 63, 604–612. doi: 10.1016/j.chb.2016.05.084 ritella, g. & hakkarainen, k (2012). instrument genesis in technology mediated learning: from double stimulation to expansive knowledge practices. international journal of computer-supported collaborative learning, 7, 239–258 doi: 10.1007/s11412-012-9144-1. robinson, k. (2011). out of our minds. learning to be creative. westford, massachusetts: capstone publishing inc. salmela-aro, k. (2017). dark and bright sides of thriving–school burnout and engagement in the finnish context. european journal of developmental psychology, 14, 337–349. doi: 10.1080/17405629.2016.1207517 salmela-aro, k., muotka, j., alho, k., hakkarainen, k., & lonka, k. (2016a). school burnout and engagement profiles among digital natives in finland: a person-oriented approach. european journal of developmental psychology, 13, 704–718. doi: 10.1080/17405629.2015.1107542 salmela-aro, k., & upadaya, k. (2012). the schoolwork engagement inventory. european journal of psychological assessment, 28, 60–67. doi: 10.1027/1015-5759/a000091 salmela-aro, k., & upadyaya, k. (2014). school burnout and engagement in the context of demands–resources model. british journal of educational psychology, 84, 137–151. doi: 10.1111/bjep.12018 salmela-aro, k., upadyaya, k., hakkarainen, k., lonka, k., & alho, k. (2016b). the dark side of internet use: two longitudinal studies of excessive internet use, depressive symptoms, school burnout and engagement among finnish early and late adolescents. journal of youth and adolescence, 1–15. doi: 10.1007/s10964-016-0494-2 sawyer, k. (2014). the cambridge handbook of the learning sciences. cambridge, massachusetts: cambridge university press. selig, j. p., & little, t. d. (2012). autoregressive and cross-lagged panel analysis for longitudinal data. in b. laursen, t. d. little, & n. card (eds.). handbook of developmental research methods (pp. 265–278). new york, new york: guilford press. selwyn, n. (2006). exploring the ‘digital disconnect’ between net savvy students and their schools. learning, media and technology, 31, 5–17. doi: 10.1080/17439880500515416 sung, y. t., chang, k. e., & liu, t. c. (2016). the effects of integrating mobile devices with teaching and learning on students’ learning performance: a meta-analysis and research synthesis. computers & education, 94, 252–275. doi: 10.1016/j.compedu.2015.11.008 tamim, r. m., bernard, r. m., borokhovski, e., abrami, p. c., & schmid, r. f. (2011). what forty years of research says about the impact of technology on learning: a second-order meta-analysis and validation study. review of educational research, 81, 4–28. doi: 10.3102/0034654310393361 upadyaya, k., & salmela-aro, k. (2013). development of school engagement in association with academic success and well-being in varying social contexts: a review of empirical research. european psychologist, 18, 136–147. doi: 10.1027/1016-9040/a000143 wexler, b. e. (2006). brain and culture. neurobiology, ideology, and social change. cambridge, massachusetts: the mit press . appendices appendix a model fit indices in the measurement invariance testing appendix b unstandardized factor loadings and r2 of the longitudinal measurement model with scalar invariance constraints appendix c latent variable correlations of the longitudinal measurement model with scalar invariance constraints appendix d materials to reproduce the results can be found here: https://osf.io/2hk3y/ codepen davis frontline learning research vol.9 no. 1 (2021) 30 43 issn 2295-3159 exploring differences in psychological well-being and self-regulated learning in university student success sarah k. davis1 & allyson f. hadwin1 1university of victoria, victoria, bc, canada article received 3o october 2019/ revised 4 december 2020 / accepted 5 december / available online 25 january 2021 abstract worldwide, there are increasing concerns about postsecondary students’ mental health and how student success is implicated. previous research has established psychological well-being and self-regulated learning are important components of student success, however, there is a paucity of research examining the interplay between these factors during a semester-long course. in this study, 118 students in a learning-to-learn elective university course completed nine weekly online planning and reflection tools. students planned for a study session, completed an academic engagement and a psychological well-being measure, then reflected on a challenge faced and described the strategy chosen to overcome that challenge. findings revealed (a) students who reported always attaining their goals also reported higher overall psychological well-being, and (b) within-person patterns of psychological well-being and academic engagement over time may affect regulatory responses to challenge or vice versa. implications for theory, research, and practice are discussed. keywords: goal attainment; process mining; psychological well-being; self-regulated learning; student success. info corresponding author email: skdavis@uvic.ca doi: https://doi.org/10.14786/flr.v9i1.581 1. introduction university students’ mental health is a growing concern globally. north american postsecondary students report: (a) feeling exhausted by academic work, and (b) experiencing levels of stress and anxiety compromising mental health, academic learning, and personal success (acha, 2018). one out of four australian university students experiences high levels of distress (larcombe et al., 2015), and in the uk, 78% of postsecondary students reported experiencing problems with their mental health in the past year (national union of students, 2015). across europe, findings are mixed: university students’ mental health tends to be better than the rest of the general population, however more students are reporting struggling with mental illness in the past 15 years (rückert, 2015). these high levels of distress could be due to any number of challenges at university. however, the consequences of poor mental health on postsecondary students is clear: mental health concerns are a common reason given by university students who take a temporary leave of absence or drop out altogether (yorke & longden, 2008). preventing this attrition is daunting because few students experiencing mental health challenges seek help (acha, 2018). in addition, the problems students experience at university may be compounded by the challenges they encounter while attempting to engage with and master coursework. specifically, while completing coursework, students report encountering problems with motivation and beliefs, planning and goal setting, well-being, emotion, and cognition (hadwin et al., 2019). these challenges interfere with student success at university. however, there is a paucity of research examining how academic challenges encountered during learning affect mental health at university. 1.1 mental health mental health, distinct from mental illness, refers to a state of well-being in which individuals cope with stressors, work productively, and contribute to society (who, 2016). in keyes’ dual-continua model, mental illness and mental health do not exist as opposite ends of a single continuum, but rather as distinct, correlated axes suggesting mental health is a separate state (see figure 1; keyes, 2005, 2013). including both hedonic (i.e., positive feeling defined as emotional well-being) and eudaimonic (i.e., positive functioning defined as psychological and social well-being) perspectives in defining mental health is vital for understanding overall human well-being (deci & ryan, 2008). there are three factors in keyes’ mental health model: psychological, social, and emotional well-being (keyes, 2002). in sum, mental health is how individuals perceive and evaluate their own affective states, and psychological and social functioning. figure 1. keyes’ dual-continua model of mental health. figure from “promoting and protecting positive mental health: early and often throughout the lifespan,” by c. l. m. keyes in c. l. m. keyes (ed.), mental well-being: international contributions to the study of positive mental health (p. 17), 2013, springer netherlands. copyright 2013 by c. l. m. keyes. reprinted with permission. 1.2 psychological well-being for this exploratory study, we focused on psychological well-being (pwb) as the mental health factor of interest because pwb may be particularly important to student success and learning at university (howell, 2009). this is because pwb characterizes the process of living and functioning well and actualizing human potential (i.e., eudaimonia; ryan et al., 2008) which are particularly relevant to university functioning. pwb captures myriad concepts related to eudaimonia, including self-acceptance, positive relations with others, personal growth, life purpose, autonomy, environmental mastery (keyes, 2013; ryff & keyes, 1995; ryff & singer, 1998), and relatedness, competence, engagement, and meaning (diener et al., 2010). the specific concepts captured by pwb may differ depending on what conceptualization and/or measure is used. for example, in self-determination theory, ryan and deci (2001) explain eudaimonic living is fostered by pursuing intrinsic goals, satisfying basic psychological needs for competence and relatedness, being mindful and acting with awareness, and behaving autonomously. this is in comparison to other conceptualizations, for example, ryff & singer (1998) whose widely used six dimensions indicate the presence of pwb. in addition, others define psychological well-being as being synonymous with happiness (e.g., hills & argyle, 2001), or hedonic well-being indicating positive or negative affect or satisfaction with life. as this study uses keyes’ theoretical framework, psychological well-being in this study is defined as how individuals perceive the quality of their functioning in life, or eudaimonia (keyes, 2013). the role of pwb is of particular interest in student success research due to the recent shift in the field from only focusing on symptoms of mental disorders at university (e.g., acha, 2018), to understanding the factors contributing to students’ pwb at university. previous research indicates students’ high well-being in high school predicts high well-being in the first weeks of university, and well-being decreases during a university semester (de coninck et al., 2019). in addition, students’ optimism is the best predictor of high pwb and lower levels of psychological distress (burris et al., 2009), and student involvement in campus organizations and sports has a positive effect on fourth year pwb (kilgo et al., 2016). current approaches to research on pwb contribute greatly to the understanding of pwb at university. however, gaps in the field include considering how pwb fluctuates over time at university and the interplay of pwb, learning, and student success. 1.3 student success and self-regulated learning at university historically definitions of student success focus on attaining a degree at the institution of attendance (kuh et al., 2007). contemporary definitions of student success are moving from defining success only at the institutional level to defining success at the student level. a meta-analysis of student success research defined success in the first year of university through the three domains of critical thinking, academic achievement, and socio-emotional well-being (van der zanden et al., 2018). this multidimensional view recognizes students may define success for themselves in different ways. thus, this current study operationalizes student success at an even finer-grained level: student success is when students attain self-set goals (e.g., academic, social, etc.) to self-determined standards of excellence by exercising strategic metacognitive monitoring and control of behaviors, emotions, motivation, and cognition within and across study sessions. self-regulated learning (srl) is vital for student success because self-regulation is ubiquitous. at university, self-regulating learners take control of their own learning, motivation, affect, and behaviors while striving to attain their own academic and personal goals (schunk & greene, 2018; zimmerman, 1989; zimmerman & schunk, 2001). the vast amount of information and choices in university can easily become overwhelming and students need to be active participants in their learning by engaging in regulating their learning, rather than by being passive recipients of information (pintrich, 2004). in srl research, challenges provide opportunities for both researchers and students to examine regulated learning as students are trying to attain goals (hadwin & winne, 2012). to become strategic learners, students (a) proactively take control of their learning by setting goals, (b) progressively develop metacognitive awareness, (c) monitor and evaluate their learning conditions, and (d) adapt their approaches when needed (winne, 2001; zimmerman, 1989). increasing metacognitive knowledge and self-monitoring skills through srl can help students overcome academic challenges and effectively develop coping strategies to deal with them (zimmerman & martinez-pons, 1990). challenges are central to university and may hinder or constrain pwb and/or learning, however limited research examines the interplay between srl and pwb around academic challenges. from previous research on srl and psychopathology at university, (a) students who experienced high levels of psychological distress may be unable to persist when they experience failure or challenges to complete academic tasks (brackney & karabenick, 1995); and (b) medical students who report using more srl strategies also reported lower rates of depression (van nguyen et al., 2015). in a study on mental health and srl using keyes’ (2002, 2005) conceptualization found, students with flourishing mental health also had the highest levels of overall adaptive academic functioning, defined as having a growth mindset, setting mastery goals, not procrastinating, and having high self-control (howell, 2009). finally, students’ use of effective motivation regulation strategies indirectly affected academic performance and emotional well-being (grunschel et al., 2016). salient components of pwb include a sense of autonomy and life purpose, and as such, goal setting and attainment are critical. for this reason, examining pwb and srl may provide further insight into the specific role of pwb in student success. 2. purpose and research questions this study aimed to examine the interplay between pwb and srl as students plan for and reflect on their approaches to attaining self-set academic goals over nine consecutive weeks. we had two research questions: (a) does pwb differ between groups of students with varying goal attainment?, and (b) how do patterns of regulation over the semester differ between a student who consistently attains weekly study goals (i.e., high goal attainment) and a student who does not (i.e., low-moderate goal attainment)?based on the findings from howell (2009), we hypothesized students with higher goal attainment will also have higher pwb, and they will regulate their learning around challenges differently than students with lower pwb. 3. methods 3.1 participants students from across the university were enrolled in an undergraduate elective course on learning strategies for university success in the fall semester of 2017. this educational psychology course taught the theory, research and practice of strategic learning, motivation, and behaviour with a self-regulated learning lens framed around winne & hadwin’s (1998) srl model. students attended one 90-minute lecture and one 90-minute lab section each week and were enrolled in at least one other course concurrently. consenting participants in this study were 140 students. we had two criteria for inclusion. first, students who missed ⅓ or more of the weekly srl diary tool (n = 22) were excluded from analysis because weekly concurrent data from these students was too sparse to examine patterns in their psychological well-being, academic engagement, or goal attainment. second, students were excluded from analysis if 50% or more of their weekly srl diary tools were completed within 1 hour or less since this activity required them to plan for, conduct and then reflect upon a 1-2 hour academic study session. the remaining 118 students fit these criteria. participants had a mean age of 19.12 years, 58% of students were female, 70% were first year students, and 90% of students reported english was their first language. 3.2 data 3.2.1. srl diary tool the purpose of the weekly srl diary tool was to encourage students to commit to one study session per week and practice engaging in a selfregulatory cycle to plan for, reflect on, and learn from each study session. diary tools are a useful instrument for measuring srl over time because they can help students raise their metacognitive awareness of their studying (schmitz et al., 2011). a narrative response constructor in the weekly diary tool prompted students to identify and reflect on a main challenge encountered that week (see figure 2). figure 2. items assessing students’ metacognitive awareness of their weekly main challenge and their regulatory response to that challenge. students completed the srl diary tool in two parts, planning and reflecting. in the planning session, the academic engagement and pwb measures create an opportunity for students to do an overall check-in on themselves for the previous week. to compute the group mean for academic engagement and pwb scores, student’s weekly scores were averaged to compute one within-person average for each student, and these scores were averaged to create the group grand mean. for the academic engagement measure, students answered six questions either yes or no about their engagement in all their academic courses for the past week (see appendix a). items 1-4 captured four aspects of behavioural engagement and items 5 and 6 captured cognitive engagement. cronbach’s alpha was .64 for the academic engagement scale (see fredericks et al., 2004). the psychological well-being measure (see appendix a) was adapted from rush & grouzet (2012) and has 10 items where students rated each item on a 6-point likert scale from 1 not at all to 7 very much. cronbach’s alpha was .85 for the pwb scale in this study. 3.2.2. indicators of srl this study uses students’ goal attainment and the challenge and strategy reflection as indicators of srl. for goal attainment, each week after completing the 1-2 hour study session, students reflected on their self-set goal and indicated if they (a) did attain, or (b) did not attain their goal. for analysis, a score of 1 was used to indicate the goal had been attained and 0 was used to indicate it had not. taken from previous research on goal attainment in the online srl diary tool (hadwin et al., 2019), we divided students up into three groups based on their goal attainment score which was calculated by the proportion of weeks the goal was reported to have been attained. natural breaks in the distribution of goal attainment proportions (see figure 3) resulted in three groups: (a) low/moderate attainers reporting attaining goals 33-78% of the time (n = 49), (b) high attainers reported attaining goals 86-89% of the time (n = 32), and (c) always attainers reported attaining their goals 100% of the time (n = 37). descriptives for the three groups are reported in table 1. figure 3. histogram of within-person mean proportion of self-set goals reported attained over nine weeks used to create three goal attainment groups. table 1 descriptives for the three goal attainment groups for challenges and strategies, these lists were generated after reviewing and categorizing open ended text-based challenge statements and strategies identified by students in earlier iterations of the course (hadwin et al., 2019). for analysis purposes, challenges were grouped into 10 distinct categories including: (a) motivation, (b) planning, (c) strategy, (d) cognition, (e) environment, (f) vocabulary and expression, (g) culture, (h) emotion, (i) mental health and well-being, (j) health and wellness, and (k) other challenge not in the list. rather than asking about specific techniques or tactics (e.g., highlighting, elaborative interrogation, etc), strategy choices focused on types of regulatory actions. for analysis purposes, strategies were grouped into 9 categories including: (a) persisting, (b) goal management, (c) strategy adjustment, (d) help seeking, (e) emotion regulation, (f) changing effort, (g) task understanding, (h) passive strategies, and (i) other strategy. 3.2.3 academic performance variables two academic performance variables were computed for the three groups. students’ final course grade reflects comprises coursework completed during the semester of the learning-to-learn course, including a final exam. the final exam tested students on their knowledge of course concepts through multiple choice questions and was worth 25% of their final course grade. grades on both items could range from 0% to 100%. 3.3 procedures all procedures were approved by the institution’s human research ethics board and all students used in data consented to participate through implied consent by enrolling in the course and not withdrawing from the research study. there was no incentive for consenting to participate in the research. data were collected as part of regular course activities graded for participation but not for content. participants completed part of the weekly srl diary tool in their lab section and finished them independently for homework before the next lab meeting. 4. results rq1: does pwb differ between groups of students with varying goal attainment? in examining the groups for differences in pwb, the low/moderate goal attainment group had the lowest pwb score of the three groups and was significantly different only from the always goal attainment group (see table 4). pwb was positively correlated to academic engagement (r = .605, p < .001) and goal attainment (r = .414, p < .001), meaning that higher levels of pwb were associated with higher levels of academic engagement and goal attainment. academic engagement was positively correlated to goal attainment (r = .538, p < 0.001), meaning that higher levels of academic engagement were associated with higher levels of goal attainment. a one-way analysis of variance (anova) determined there were significant differences between the three goal attainment groups for pwb (f (2,115) 6.497, p = .002). this corresponded to an effect size of η 2 = .10 indicating 10% of the variance in pwb scores was predictable from goal attainment group membership. a tukey post hoc test (α = .05) revealed the pwb score was significantly lower for the low/moderate goal attainment group (m = 45.27) than the always goal attainment group (m = 5.15). the high goal attainment group (m = 47.73) did not differ significantly from either group. table 2 mean pwb scores for each goal attainment group rq2: how do patterns of regulation over the semester differ between a student who consistently attains weekly study goals (i.e., high goal attainment) and a student who does not (i.e., low-moderate goal attainment)? the anova established differences between the pwb of the low/moderate group and the always goal attainment group. next, we examined within-person patterns of pwb, academic engagement, and srl for two sample students. process mining is a new method used to gain insight into students’ regulatory patterns and processes (see bannert et al., 2013). previous research has used process mining maps to aggregate student data by groups, but they can also be used to map individual students’ data over time to uncover patterns representative of dominant student profiles (e.g., rogiers et al, 2020). due to the highly individualized nature of the online srl diary tool, we did not have expectations of “correct” sequences of student responses. for example, if a student reported a motivation challenge, there are several strategies they may have chosen rather than only one correct strategy to choose. thus, we did not aggregate students’ process mining maps but rather we chose one student from each group whose individual mean of pwb was the closest to the group mean and created a process mining map for each of these students. we hypothesized these students would show different patterns of regulating their learning over time. videos 1 and 2 show the process mining maps for student lm (mean pwb = 45.22) from the low/moderate group and student al (mean pwb = 51.89). the process mining maps show each student’s self-reported academic engagement, pwb, challenge (in capital letters), and weekly strategy over nine weeks. pwb mean scores were divided into three categories: low pwb = 10-30, moderate pwb = 31-50, and high pwb = 51-70. academic engagement was also divided into three categories: low engagement = 1-2, moderate engagement = 3-4, and high engagement = 5-6. challenges are reported in capital letters. both challenges and strategies were grouped according categories outlined in 3.2.2 in this paper. as seen in the dynamic version of the map (see video 1), each moving dot indicates the sequence of the four categories as selected each week by the student and the 9 weeks appear simultaneously. the green dot indicates the goal was attained and the red dot indicates the goal was not attained. hovering over the dot will reveal the week of each individual dot. the darker blue colour of the box indicates the higher frequency of selection by the student and the numbers by the lines between boxes indicate the number of times a path occurred. for example, for student lm, at the start of the video, the 3 weeks where the student had high engagement, the student attained all their goals. student lm’s most common challenge was motivation (n = 4) and the most common path was between moderate engagement and moderate pwb (n = 6). in student al’s video (see video 2), the video shows frequent high engagement with all goals attained. student al’s most common challenge (n = 4) and the most common path was between high engagement and high pwb (n = 5). video 1: process mining map for student lm from the low/moderate goal attainment group student lm from the low/moderate group began the course with high engagement and pwb, but strategy choices of passive and help-seeking led to frequent motivation challenges, leading to moderate engagement and pwb toward the end of the course (see video 1). student al from the always group also began the course with high engagement and pwb, but the strategy choice of persist led to moderate pwb (see video 2). when strategy, motivation, or cognition were the dominant challenges, strategies chosen by student al led to high engagement, followed by high pwb toward the end of the course. video 2: process mining map for student al from the always goal attainment group 5. discussion this study aimed to examine the interplay between pwb and srl as students plan for and reflect on their approaches to attaining self-set academic goals over nine consecutive weeks. we used both between-person and within-person approaches. two main findings from this study are highlighted: (a) students’ between-person pwb differs according to self-reported goal attainment, and (b) students’ within-person patterns of regulatory responses provide insight into the interplay between pwb and srl. this study found students who reported always attaining their goals had higher pwb than students who reported low/moderate levels of goal attainment. this is consistent with boudreaux & ozer’s (2012) finding that students who have high goal attainment both report success in pursuing multiple goals and have high life satisfaction and positive affect (i.e., emotional well-being). the finding from this study adds to the student success literature that students who always attain their self-set study goals also have high pwb. feeling purpose in life and attaining goals is a part of the six dimensions of pwb (ryff & singer, 1998). similarly, from an srl perspective, attaining self-set goals indicates students are regulating their learning effectively. we did not specifically examine students’ task perceptions or quality of the goals students’ attained, two essential aspects of the phases of regulating learning (hadwin & winne, 2012), therefore future research could examine how task perceptions and subsequent goal attainment are implicated in both academic engagement and pwb. our second main finding was within-person patterns of pwb and academic engagement over time may affect regulatory responses to challenge or vice versa. regulatory responses in this study comprised the challenges and strategies students reported while attempting to attain self-set goals. previous research examining srl and mental health at one time point also found students with better mental health had a mastery-approach goal orientation (howell, 2009). examining the process maps of student lm and student al indicates pwb and engagement may be interacting with challenges and strategies. for example, student lm seemed to try a variety of strategies when they had a motivation challenge but continued to have motivation challenges during the semester, ending with moderate engagement and pwb compared to the start of the semester. student al also experienced motivation challenges but tried several other strategies and in turn experienced high engagement and high pwb consistently throughout the semester. drawing on these findings, does high engagement or pwb fuel regulatory responses or do regulatory responses fuel high engagement or pwb? as srl processes and strategies change over time and within-person, future research should continue to examine srl and pwb or mental health as within-person processes. this will also allow for robust interventions based on students’ individual patterns, rather than comparing students to each other. 5.1 implications in srl, metacognitive awareness is central to learning and student success. at university, students must contend with multiple goals (e.g., academic, social, financial) to attain success. this study contributes to theory by indicating high pwb may be advantageous for students regulating their learning, or vice versa. personal growth, life purpose, and autonomy are some of the dimensions of pwb (ryff & singer, 1998), therefore these larger-grained states may be important for students to be aware of as they attempt to regulate their learning, specifically around challenges. due to the paucity of research on mental health and srl, there are many opportunities for future research to examine the interplay of srl and pwb or mental health at university for student success. importantly, as students may benefit from extending their metacognitive awareness to their engagement and pwb while learning, interventions should be designed with both students and researchers in mind. by examining the process mining maps of two students, we were able to see patterns in students’ pwb, engagement, challenges, and strategies over nine weeks. molenaar et al. (2019) found providing elementary students with personalized visualizations based on their learning improved regulating practice behaviour, learning transfer, and relative monitoring accuracy. thus, even though we did not show students their maps in this study, future research could show students their process mining maps of their learning. importantly, this process needs to be supported by srl as performance feedback alone is not always sufficient for students to translate their data into increased engagement in srl (butler & winne, 1995). for example, students may be able to realize they are not engaged in their courses or attaining their goals, or they may recognize their pwb is also lower, and select a course of action. alternatively, if students realize their pwb is low, this could help them to see the connection between their pwb, engagement, and/or goal attainment. this metacognitive awareness could also help students employ different strategies, such as revisiting their task perceptions, or revising their goals. when students engage in weekly regulatory planning and reflection, assessing their own pwb and engagement may provide easy access for students to take this data, examine it, and make changes as necessary. specifically, visualizations of adaptive and maladaptive regulation patterns, such as through process mining maps, may help students identify for themselves where and when to make changes. 5.2 limitations this study examined data from one semester of an undergraduate elective learning-to-learn course and may not be generalizable to other courses. students in this course could have been actively making changes to their learning approaches, potentially affecting the findings. future research could examine the interplay between pwb and srl in other undergraduate courses to see how student’ regulatory responses vary when they are not taking a course on how to effectively manage their srl processes and strategies. also, this study only examined one aspect of mental health, pwb, rather than all three aspects of psychological, social, and emotional well-being. understanding students’ pwb at university is a salient issue. however, as pwb is only one part of mental health, considering all aspects of mental health may further help students increase their self-awareness and success simultaneously. future research could use a comprehensive mental health scale, such as the mental health continuum-short form (mhc-sf; keyes, 2009) to investigate further the interplay between the three factors of mental health and srl. next, only two students’ process maps were examined for within-person differences. this limits the generalizability of the findings but does create opportunities for designing visualizations to assist students who are trying to improve either their learning approaches, their pwb and/or mental health, or both. this study relied on self-reported data only. this is important as this data can be useful for students to reflect on their own learning, but future research should incorporate objective data (e.g., trace data) to triangulate results. in particular, we relied on students’ self-reported data of their goal attainment and even students in the low group were reported attaining goals most of the time (see figure 3). while students’ perceptions of their own learning are salient in srl and student success, triangulating other objective goal attainment data would be prudent. finally, the very moderate reliability for the academic engagement measure , indicating there may be a large error variance. the purpose of this measure was a weekly checklist where students could indicate whether or not they engaged in the specific behaviour (e.g., attended classes), however caution should be exercised in interpreting the findings of this measure due to the measurement error present. 6. conclusion in sum, to tackle the bigger issue of mental health at university, this study indicates students who report attaining their goals more often also have higher pwb. leveraging srl processes and strategies around academic challenges may also help students’ pwb and engagement or vice versa. as engagement and pwb fluctuate over time, being aware of regulatory patterns may help students engage in more metacognitive control and strategic action. this is because, for students to be active in their srl, they need ways to record and track data about their learning in order to identify patterns, monitor their approaches, and make adaptations as necessary (winne, 2005). self-regulating learners already regulate their behaviour, cognition, motivation, and emotion to reach goals (winne & hadwin, 1998), therefore extending their metacognitive awareness and control to their pwb may also be advantageous for student success. finally, analyzing within-person patterns of pwb and srl processes may offer the most opportunity for interventions with high utility and applicability by students to be successful at university. keypoints students’ mental health at university is a growing global concern. this study examines psychological well-being, one of the three factors of mental health, and srl at university. process mining maps provide valuable insight to psychological well-being and students’ regulatory responses to academic challenges. acknowledgments this research was supported by a social sciences and humanities research council (sshrc) of canada insight research grant 435-2012-0529 (pi: hadwin) and 435-2018-0440 (pi: hadwin); and a sshrc doctoral fellowship (s. k. davis). we would also like to thank ramin rostampour for his creation of the process mining maps. references american college health association (2018). american college health association-national college health assessment ii: undergraduate student reference group data report fall 2018. silver spring, md: american college health association. retrieved from: https://acha.org/documents/ncha/ncha-ii_fall_2018_undergraduate_reference_group_data_report.pd bannert, m., reimann, p., & sonnenberg, c. (2013). process mining techniques for analysing patterns and strategies in students’ self-regulated learning. metacognition and learning, 9, 161-185. https://doi.org/10.1007/s11409-013-9107-6 boudreaux, m. j., & ozer, d. j. (2013). goal conflict, goal striving, and psychological well-being. motivation and emotion, 37, 433-443. https://doi.org/10.1007/s11031-012-9333-2 brackney, b. e. & karabenick, s. a. (1995). psychopathology and academic performance: the role of motivation and learning strategies. journal of counseling psychology, 42(4), 456-465. https://doi.org/10.1037/0022-0167.42.4.456 burris, j. l., brechting, e. h., salsman, j., & carlson, c. r. (2009). factors associated with the psychological well-being and distress of university students. journal of american college health, 57(5), 536-544. https://doi.org/10.3200/jach.57.5.536-544 butler, d.l. and winne, p.h. (1995). feedback and self-regulated learning: a theoretical synthesis. review of educational research, 65(3), 245-281. https://doi.org/10.3102/00346543065003245 conley, c. s., durlak, j. a., & kirsch, a. c. (2015). a meta-analysis of universal mental health prevention programs for higher education students. prevention science, 16, 487-507. https://doi.org/10.1007/s11121-015-0543-1 deci, e. l., & ryan, r. m. (2008). facilitating optimal motivation and psychological well-being across life's domains. canadian psychology, 49(1), 14. https://doi.org/10.1037/0708-5591.49.1.14 de coninck, d., matthijs, k., & luyten, p. (2019). subjective well-being among first-year university students: a two-wave prospective study in flanders, belgium. student success, 10(1), 33-45. https://doi.org/10.5204/ssj.v10i1.642 diener, e., wirtz, d., tov, w., kim-prieto, c., choi, d., oishi, s., & biswas-diener, r. (2010). new well-being measures: short scales to assess flourishing and positive and negative feelings. social indicators research, 97(2), 143-156. https://doi.org/10.1007/s11205-009-9493 -y grunschel, c., schwinger, m., steinmayr, r., & fries, s. (2016). effects of using motivational regulation strategies on students' academic procrastination, academic performance, and well-being. learning and individual differences, 49, 162-170. https://doi.org/10.1016/j.lindif.2016.06.008 fredricks, j. a., blumenfeld, p. c., & paris, a. (2004). school engagement: potential of the concept: state of the evidence. review of educational research, 74(1), 59–119. https://doi.org/10.3102%2f00346543074001059 hadwin, a. f., & winne, p. h. (2012). promoting learning skills in undergraduate students. in m. j. lawson & j. r. kirby (eds.), the quality of learning: dispositions, instruction, and mental structures (pp. 201-227). cambridge university press. https://doi.org/10.1017/cbo9781139048224.013 hadwin, a. f., davis, s. k., bakhtiar, a., & winne, p.h. (2019). academic challenges as opportunities to learn to self-regulate learning. in h. askell-williams & j. orrell (eds). problem solving for teaching and learning book: a festschrift for emeritus professor mike lawson (pp. 34-48). routledge. https://doi.org/10.4324/9780429400902 hills, p. & argyle, m. (2001). emotional stability as a major dimension of happiness. personality and individual differences, 31(8), 1357-1364. https://doi.org/10.1016/s0191-8869(00)00229-4 howell, a. j. (2009). flourishing: achievement-related correlates of students’ well-being. the journal of positive psychology, 4(1), 1-13. https://doi.org/10.1080/17439760802043459 keyes, c. l. m. (2002). the mental health continuum: from languishing to flourishing in life. journal of health and social behavior, 43(2), 207-222. https://doi.org/10.2307/3090197 keyes, c. l. m. (2005). mental illness and/or mental health? investigating axioms of the complete state model of health. journal of consulting and clinical psychology, 73(3), 539-548. https://doi.org/10.1037/0022-006x.73.3.539 keyes, c. l. m. (2009). atlanta: brief description of the mental health continuum short form (mhc-sf). retrieved from: https://www.aacu.org/sites/default/files/mhc-sfenglish.pdf keyes, c. l. m. (2013). promoting and protecting positive mental health: early and often throughout the lifespan. in c. l. m. keyes (ed.) mental well-being: international contributions to the study of positive mental health (pp. 3-28). springer netherlands. https://doi.org/10.1007/978-94-007-5195-8 kilgo, c. a., mollet, a. l., & pascarella, e. t. (2016). the estimated effects of college student involvement on psychological well-being. journal of college student development, 57(8), 1043-1049. https://doi.org/10.1353/csd.2016.0098 kuh, g. d., kinzie, j., buckley, j. a., bridges, b. k., & hayek, j. c. (2007). piecing together the student success puzzle: research, propositions, and recommendations. ashe higher education report (vol. 32). john wiley & sons. https://doi.org/10.1002/aehe.3205 larcombe, w., finch, s., sore, r., murray, c. m., kentish, s., mulder, r. a., lee-stecum, p., baik, c., tokatlidis, o. & williams, d. a. (2016) prevalence and socio-demographic correlates of psychological distress among students at an australian university, studies in higher education, 41(6), 1074-1091. https://doi.org/10.1080/03075079.2014.966072 molenaar, i., horvers, a., dijkstra, r., & baker, r.s. (2020). personalized visualizations to promote young learners' srl: the learning path app. in proceedings of the tenth international conference on learning analytics & knowledge (lak '20). association for computing machinery, new york, ny, usa, 330–339. https://doi.org/10.1145/3375462.3375465 national union of students (2015). mental health poll. retrieved from: https://www.nusconnect.org.uk/resources/mental-health-poll-2015 pintrich, p. r. (2004). a conceptual framework for assessing motivation and self-regulated learning in college students. educational psychology review, 16, 385-407. https://doi.org/10.1007/s10648-004-0006-x rogiers, a., merchie, e., & van keer, h. opening the black box of students’ text-learning processes: a process mining perspective. frontline learning research, 8(3), 40-62. https://doi.org/10.14786/flr.v8i3.527 rückert, h. (2015). students׳ mental health and psychological counselling in europe. mental health & prevention, 3(1-2), 34-40. https://doi.org/10.1016/j.mhp.2015.04.006 rush, j. & grouzet, f. m. e. (2012) it is about time: daily relationships between temporal perspective and well-being, the journal of positive psychology, 7(5), 427-442. https://doi.org/10.1080/17439760.2012.713504 ryan, r. m., & deci, e. l. (2001). on happiness and human potentials: a review of research on hedonic and eudaimonic well-being. annual review of psychology, 52, 141-166. https://doi.org/10.1146/annurev.psych.52.1.141 ryan, r. m., huta, v., & deci, e. l. (2008). living well: a self-determination theory perspective on eudaimonia. journal of happiness studies, 9, 139-170. https://doi.org/10.1007/s10902-006-9023-4 ryff, c. d., & keyes, c. l. m. (1995). the structure of psychological well-being revisited. journal of personality and social psychology, 69(4), 719-727. https://doi.org/10.1037/0022-3514.69.4.719 ryff, c. d., & singer, b. (1998). the contours of positive human health. psychological inquiry, 9, 1-28. https://doi.org/10.1207/s15327965pli0901_1 schmitz, b., klug, j., & schmidt, m. (2011). assessing self-regulated learning using diary measures with university students. in zimmerman, b. j., schunk, d. h. (eds.), handbook of self-regulation of learning and performance (pp. 251-266). routledge. https://doi.org/10.4324/9780203839010.ch16 schunk, d. h. & greene, j. a. (2018). historical, contemporary, and future perspectives on self-regulated learning and performance. in d. h. schunk & j. a. greene (eds.),handbook of self-regulation of learning and performance (2 nd ed.) (pp. 1-15). routledge. https://doi.org/10.4324/9781315697048.ch1 van der zanden, p. j. a. c, denessen, e., cillessen, a. h. n., & meijer, p. c. (2018). domains and predictors of first-year student success: a systematic review. educational research review, 23, 57-77. https://doi.org/10.1016/j.edurev.2018.01.001 van nguyen, h., laohasiriwong, w., saengsuwan, j., thinkhamrop, b., & wright, p. (2015). the relationships between the use of self-regulated learning strategies and depression among medical students: an accelerated prospective cohort study. psychology, health & medicine, 20(1), 59–70. https://doi.org/10.1080/13548506.2014.894640 winne, p. h., & hadwin, a. f. (1998). studying as self-regulated learning. in d. j. hacker, j. dunlosky, & a. c. graesser (eds.),  metacognition in educational theory and practice. (pp. 277-304). lawrence erlbaum.    https://doi.org/10.4324/9781410602350-19 winne, p. h. (2001). self-regulated learning viewed from models of information processing (p. 164). in b. j. zimmerman & d. h. schunk (eds.), self-regulated learning and academic achievement: theoretical perspectives (2nd ed., pp.153–189). lawrence erlbaum. https://doi.org/10.1007/978-1-4612-3618-4 winne, p. h. (2005). a perspective on state-of-the-art research on self-regulated learning. instructional science, 33, 559-565. https://doi.org/10.1007/s11251-005-1280-9 world health organization. (2016). mental health: strengthening our response. [fact sheet].   retrieved from: http://www.who.int/mediacentre/factsheets/fs220/en/   yorke, m., & longden, b. (2008). the first-year experience of higher education in the uk. york: higher education academy. retrieved from: https://www.heacademy.ac.uk/system/files/fyefinalreport_0.pdf zimmerman, b. j. (1989). a social cognitive view of self-regulated academic learning. journal    of educational psychology, 81(3), 329-339. https://doi.org/10.1037/0022-0663.81.3.329 zimmerman, b. j., & martinez-pons, m. (1990). student differences in self-regulated learning: relating grade, sex, and giftedness to self-efficacy and strategy use. journal of educational psychology, 82(1), 51-59. https://doi.org/10.1037//0022-0663.82.1.51 zimmerman, b. j. & schunk, d. h. (eds.) (2001). self-regulated learning and academic achievement: theoretical perspectives (2nd ed.). lawrence erlbaum. https://doi.org/10.1007/978-1-4612-3618-4 appendix a: measures academic engagement items frontline learning research vol. 5 no. 3 special issue (2017) 123 138 issn 2295-3159 visual expertise as embodied practice jonas ivarsson1 university of gothenburg, sweden article received 28 april / revised 21 november / accepted 23 march / available online 14 july abstract this study looks at the practice of thoracic radiology and follows a group of radiologists and radiophysicists in their efforts to find, discuss, and formulate issues or troubles ensuing the implementation of a new radiographic imaging technology. based in the theoretical tradition of ethnomethodology it examines the local endogenous practices pertaining to the radiologists’ expertise in the interpretation of visual representations and tries to explicate the ways in which they draw upon various resources in order to accomplish their professional tasks. as the study is addressing the topic of visual expertise it also aims to do so in terms that acknowledge that all expertise is rooted in embodied practices. the analysis follows a case of what is called the enacted production of radiological reasoning. one of the central features of the described work is the manner in which it is carried out by way of the living present body of an expert. the experienced radiologist interweaves anatomical and technological terminology with visual representations and gestures in such a way that none of these components can be said to be superfluous to the argumentation. as a consequence, we should appreciate gestures and embodied actions as important means through which expertise become organised. these are parts of a repertoire of methods through which the experts learn their profession. in addition, gestures can also become enrolled in the re-negotiation of expertise in the face of new challenges. keywords: visual expertise; radiology; ethnomethodology; body; gesture 1 contact information: jonas ivarsson, department of education, communication and learning, university of gothenburg, sweden. e-mail: jonas.ivarsson@gu.se doi: http://dx.doi.org/10.14786/flr.v5i3.253 http://dx.doi.org/10.14786/flr.v5i3.253 ivarsson | f l r 124 1. introduction this study looks at practitioners in thoracic radiology. it follows a group of radiologists and radiophysicists in their efforts to find, discuss, and formulate issues or troubles ensuing from the implementation of a new radiographic imaging technology. by foregrounding this case, where professionals grapple to overcome some difficulties in interpreting new forms of radiographs, it becomes possible to examine the matter of visual expertise through a dual lens; it simultaneously presents us with a specialised area of expertise as something performed and as something talked about by the very same practitioners. the profession of radiology has typically been described as the technical art of visually perceiving structures and pathologies by way of radiographs. it has been portrayed as a solitary practice where for instance the position of the radiologist’s eye in relation to the image can be examined for the ways that it will impact on the detection of pathologies (kundel, nodine, & toto, 1991). in this traditional view of radiological practice, the body plays an intriguingly subordinate role, and expertise in diagnosing x-rays is described as grounded in deep forms of cognitive processing (lesgold et al. 1988). when radiology in this way becomes prefaced by its function as visual assessments of representational objects, both the practitioners and their patients seem to figure merely as dis-embodied phantoms. the objective of this study then is to revisit this isolated focus on eyes and perceiving retinas and to bring the body back into the study of visual expertise. the general model of perception as a process where sensation and movement are seen as intrinsically tied to visual understandings of form is itself not new (cf., myers, 2008). ideas of this kind have been advanced in the theoretical works of such scholars as maurice merleau-ponty (1962) and james gibson (1968, 1979). this particular study though, will draw on insights generated within the tradition of ethnomethodology and studies of talk-in-interaction. as will be argued, this means that the analysis seeks to describe the work of the practitioners in its discipline specific details. it looks at the local endogenous practices pertaining to the radiologists’ expertise in the interpretation of visual representations and the ways in which they draw upon various resources, including the body, in order to accomplish their professional tasks. 2. ethnomethodology ethnomethodology is a form of social inquiry, “dedicated to explicating the ways in which collectivity members create and maintain a sense of order and intelligibility in social life” (have, 2004, p. 14). the tradition was founded by harold garfinkel in the 50s and 60s and one of his key publications is “studies in ethnomethodology” which was published 1967. this book began with a densely phrased description of the enterprise that nevertheless captures much of what then came to be expounded: ethnomethodological studies analyze everyday activities as members’ methods for making those same activities visibly-rational-and-reportable-for-all-practical-purposes, i.e., ‘accountable’, as organizations of commonplace everyday activities. (garfinkel, 1967, p. vii) rather than offering a method of study, ethnomethodology turns an eye towards the methods used by members and makes those into its object of study. as a consequence, the analyses are not primarily aimed at generating new knowledge. rather such studies seek to explicate what is already known and shared within a targeted group. a possible objection to this restriction in scope could be that analyses of this kind would not amount to much. however, through the close descriptions and detailed accounts provided by the analyses, also non-members can be granted partial access to the inner workings of a practice. in this way the ethnomethodological analysis can also become a form of pedagogy (garfinkel, 2002). ivarsson | f l r 125 one feature, central to the focus on members’ methods in garfinkel’s (1967) writing, is also the idea of accountability. the notion of accountability practices was originally borrowed from the legalities surrounding businesses but in the hands of ethnomethodology it came to be applied to the entirety of social life. the idea here is that members of a practice design their actions, for others, so as to make those actions visible as what they are. for instance, a pedestrian aiming to cross a busy street usually makes sure that this “crossing-the-street” becomes a witnessable thing in the world for others (especially drivers) to see and relate to. it is in this sense that ethnomethodology has come to speak of social life as constituting a “witnessable order”. this view embodies a radical methodological departure from other social inquiries that work from a belief in an underlying or hidden order, that can only be uncovered with the application of specific sociological methods or theoretical concepts (livingston, 2008). to ethnomethodology, social order is available for all to see and analyse, it is as available to laymen as it is to professionals. this general approach is immensely useful when studying a vast array of social situations and actions. nevertheless, when we move into domains of specialised professional practice some methodological complications arise. first of all, professional actions such as operations on and in a physical or symbolic environment are social in a special sense. when for instance a dentist is clearing out a root-canal with a file, the physical actions are done as parts of a medical procedure and are primarily done to do that job. simultaneously, those same actions are also accountable actions within the medical practice of root canal treatments. as the manual domain-specific operations are executed they become witnessable by other practitioners. this means that if those actions are carried out incorrectly (according to the standards upheld by the profession) they can be called out, reprimanded or made into a case of medical malpractice. the upshot here is that some forms of professional conduct may chiefly be designed for, and thus accessible to, other professionals. this condition can make the study of expert performance and reasoning more difficult (for a further discussion see lynch, 1993). there are different ways that this methodological difficulty has been managed. in their studies of archaeological excavations (goodwin, 1994), architectural reasoning (lymer, 2010) or gallbladder surgery (koschmann, lebaron, goodwin, & feltovich, 2011) the authors all chose to focus on educational settings, arrangements where skilled practitioners were instructing students in the professional forms of seeing and reasoning. in these settings the important things to be seen and known as professional objects became articulated and thereby also rendered available for an overhearing analyst. another possibility is instead that the analyst becomes a skilled member of the practice. in this tradition we find the highly detailed and most insightful analyses of such practices as improvised jazz (sudnow, 1978), mathematical proving (livingston, 1999) and the organisation of turn-taking in surfing (liberman, 2015). in these latter studies, the authors can be seen adhering to what garfinkel (2002) termed the “unique adequacy requirement of method”, the policy dictating that an analyst must also be competent in the very methods that he or she is studying. while the present study is adopting neither of these two approaches in any traditional sense the case to be analysed has nevertheless been selected and worked on with the above-mentioned complications in mind. any analyst interested in the expertise pertaining to radiology is presented with a daunting challenge due to the obscurity of the work—as part of an ordinary day’s work the practice of assessing radiographs is typically carried out in solitude. for this reason, the analysis presented here will focus on a specially organised meeting where a group of radiologists met with a group of radiophysicists to talk about their current skills in detecting pulmonary nodules2 and some possible limitations of those skills. thus, the visual expertise within the group of radiologists was itself a topic for the discussion. whereas this situation presented more talk than would a solitary reading of radiographs, its analysis would still require some competency in the radiological matters discussed. in order to make any such analysis possible the author has collaborated with the group of radiologists and radiophysicists to the extent of doing joint analyses and co-authoring papers over several years. this form of research can be seen as an example of what garfinkel (2002) calls “hybrid studies of work”, a form of study that focuses on member’s 2 a pulmonary nodule is a small round or oval-shaped growth in the lung which is possibly malignant. ivarsson | f l r 126 methods in its discipline-specific details. grounded in the experience of this and other inter-disciplinary collaborations it is also suggested here that the requirement of unique adequacy could be somewhat reconsidered. rather than standing as a requirement pertaining to each and every individual it is perhaps better to consider the competency of the analysing collaborative or research team as a whole. the circumstances that brought about this specific collaboration and the studied situation will be addressed after some notes on the study of embodied interaction. 2.1 studying knowledge and the body there is a growing body of ethnomethodological studies that address interaction in everyday and workplace settings that also attend to embodied and material aspects of the communicative situations. as argued by charles goodwin (2000) traditional analytic and disciplinary boundaries have tended to isolate language from its environment. in order to avoid such separations he has aimed to “provide a systematic framework for investigating the public visibility of the body as a dynamically unfolding, interactively organized locus for the production and display of meaning and action” (2000, p. 1490). not only are bodies seen as central to the investigation of human action, they are implicated at a fundamental level in the very skills that people come to possess: “as the active body acquires skills, those skills are stored, not as representations in the mind, but as dispositions to respond to the solicitations of the situation” (streeck, 2015, p. 422). in this way knowledge has a tacit, gestural and even muscular side (griesemer, 2004), but these aspects are easily neglected once ideas becomes established and mastered on a personal level. in relation to medical practice, “bodies” become implicated in a multitude of ways. stefan hirschauer (1991) has for instance addressed the link between physicians, patients’ bodies, and anatomical representations by looking at the work done during surgery. he argues that surgeons have to acquire two bodies in their education: their own trained body and the abstract body as learnt from textbooks and other representations. when learning about anatomy griesemer sees the origin of ideas to lie much “in the coordination of the senses, particularly sight and touch” (2004, p. 440). this interest for the relationship between hands and object in surgery has also shifted the focus of the observation “away from visual and cognitive models toward a focus on what happens at the interface of hands and instruments” (prentice, 2005, p. 840). at this interface, gestures can play different important roles (kendon, 1997). streeck (2009) has distinguished what he calls six “ecologies of gesture” and two of these are most relevant in this context. first that gestures can select and elaborate features and significances of the world within sight and thereby orient participants to the visible environment beyond the reach of the hands. second, that gestures can evoke phenomena that are not present and depict imaginary and abstract worlds. in a study of how brain neuroscientists work with digital fmri scans alac (2008) discusses how the ‘seeing’ of images is an embodied process that is achieved through a coordination of the visual information generated by the technical instruments with the world of meaningful actions and practical problem solving. in this work “gesture, talk and the manipulation of the digital screen function together as techniques for managing perception” (2008, p. 493). furthermore, as “the gestures participate in the interpretive act as an embodied enactment of the process of change” (p. 494) alac argues that the neuroscientists display a way of seeing images that involves the hands as well as the eyes. a much similar argument is presented by slack, hartswood, procter, and rouncefield (2007) in their study of the diagnostic work of mammography. the authors characterize the reading of mammograms as “lived work” which is encompassed by “the arrangement of mammograms, gesturing and pointing to features on mammograms, manipulating mammograms, and annotations” (2007, p. 176). they stress the importance of appreciating the social nature of the work and the embodied nature of reading and annotation. when the studied experts, in their practices of seeing/reading mammograms, incorporate such things as hands, pencils and gestures one should note that “these techniques are not ad hoc workarounds but repertoires of manipulations that are an integral part of the embodied practice of realizing phenomena as what they accountably are” (2007, p. 182). ivarsson | f l r 127 in relation to how professionals acquire skills during training, koschmann and lebaron (2002) investigated learners in various medical professions and their use of gestures in articulating their knowledge. the authors distinguish two ways that the notion of “articulating knowledge” can be conceived. first, gestures can be seen to reveal something about the learner’s current understanding. on the other hand, gesture can be treated “not only as an external manifestation of understanding but also as reflecting a constructive process of connection making” (2002, p. 252). it is with this latter view in mind that we turn to the present study and its interest in how gestures can become active means through which visual expertise is built and enacted—how gestures are performed in the service of sense making (cf., koschmann, lebaron, goodwin, zemel, & dunnington, 2007), both for self and others. 3. background to the studied setting the analysed material stems from a collaborative research project carried out in radiology and radiation physics. on a general level, the project was addressing how advancements in imaging technologies were challenging existing forms of expertise and thereby imposing the development of new criteria and methods of interpretation. more specifically, the empirical material concerns the work following the introduction of a new radiographic technology, called tomosynthesis. at the time of its implementation, tomosynthesis was recognized to have considerable advantages over ordinary chest radiography. in a first study, it was shown that the detection of pulmonary nodules was significantly higher for tomosynthesis than for chest radiography when used by experienced thoracic radiologists (vikgren et al., 2008). on the other hand, compared to the technology of computed tomography (ct), chest tomosynthesis has a limited depth resolution, which was considered a disadvantage for interpreting pathologies (johnsson et al., 2010). furthermore, since tomosynthesis at that time was a new technology, the knowledge of how to best analyse the resulting images was limited. as a response to this new situation, a subsequent study was arranged. in this study, six observers analysed the same group of tomosynthesis cases (n. 89) for presence of pulmonary nodules in two reading sessions, with the purpose of measuring the difference in performance due to learning with feedback between the two sessions. the reading sessions were separated by a collective review session, at which the observers were given feedback on their analyses on an additional set of tomosynthesis cases (n. 25). the collective session also served the purpose of identifying pitfalls and formulating suggestions on how to avoid them (asplund et al., 2011; rystedt, ivarsson, asplund, johnsson, & båth, 2011). the present investigation will focus on the interaction between the participants during the collective review session. 3.1 recording and data processing the review session lasted for almost six hours and was recorded with two cameras. a primary high definition camera was aimed at two projector screens set side-by-side displaying tomosynthesis and ct images. in order to help discriminate between the voices of the active participants at the session, a secondary standard definition camera was also installed and aimed at the group. originally, the view provided by this camera was not intended to be included in the analysis. nevertheless, after the fact, this recording was found to display a number of interesting features and was subjected to further analysis. in order for the radiologists to properly carry out their work, of making discernments on the screens, the ambient light in the room had to be kept at a minimum. for the purposes of the recording, this resulted in dark and grainy video images that are difficult to print as stills on a page. when viewed this way, the embodied behaviours of the participants are easily lost. thus, the very phenomenon that this study seeks to explore evades a simple re-presentation. in order to overcome this analytic problem the events have been represented in a sequence of drawings. this is something that has now become common practice in many ivarsson | f l r 128 video studies (e.g., goodwin, 2007b; lindwall & ekström, 2012; melander & sahlström, 2009). rather than simply tracing the images provided by the video stills, the processing here has proceeded in a more roundabout way. after transcription (elan), episodes selected for further analysis have been digitally reenacted using 3d modelling software (poser). poser is a virtual film studio that centres on depicting the human figure in three-dimensional form. it allows for the recreation of some features of the setting, but primarily lets the user control the bodies of digital mannequins by way of their orientations, gestures and gazes. the outputted renderings have later been retouched (photoshop) and compiled together with the transcribed talk (indesign). aside from proffering visually clear output, this procedure has had two main analytic advantages. first, there is an analogy to how the transcription of spoken interaction compels the analyst to focus on details of the talk that would not be attended to under normal conditions. by engaging in the exact replication of body postures, the flexing of joints, the placement of limbs and the like, the analyst can get a handle on the situation that the simple tracing of outlines cannot offer. the three dimensions of the recorded bodies are momentarily recovered in the process. the second advantage works mostly in benefit of the reader. the images selected for presentation are not necessarily tethered to the camera view. if a different angle provides a better view of an unfolding action, it can be selected effortlessly, and, thereby help the reader to get a better understanding of the events as they took place. whereas the original video is still understood as the primary data (the transcripts and) the images should be seen in their capacity as analytic representations—on a par with the textual analysis itself. 4. analysis in order to begin with the analysis, the relation between the studied session and ordinary practice have to be clarified. during an ordinary day’s work, the major task facing the radiologists typically concerns diagnostic work in clinical practice and their formulation of recommendations tailored to referring physicians. in line with this characterisation, the efforts undertaken during the review session could be understood as diagnostic work of a second order. the things to be diagnosed not only had to do with suspected nodules, but also most centrally concerned errors in the work of finding nodules. out of this exploration, recommendations informing further first order diagnostic work had to be formulated. in the materials, both of these aspects can be discerned. in some instances, the participants work towards the formulation of what things accountably are (as anatomical structures). in yet other instances there are attempts at formulating difficulties pertaining to the very process of diagnosis—difficulties thus instigated by the new technology. as will become evident, these two processes of diagnostic reasoning are closely connected, and, to a varying degree, involve interesting forms of embodied conduct. the episode that we will examine in detail mainly follows an extended argument made by one of the senior radiologists. in order to enhance the readability of the 42 seconds long sequence the transcripts have been segmented into a number of figures (1–7). the labelling of these is merely meant to provide some clues as regards the evolving topic of the talk-in-interaction. it should also be noted that the separate images represent one continuous sequence with no omissions at the joints. 4.1 the tricky thing just prior to the sequence, the group has been discussing some general features of the technology of tomosynthesis and the potential benefits of discovering centrally placed tumours with patients that also have ivarsson | f l r 129 pleural plaques3. this part of the discussion is concluded with the notion that such a prospect will very much depend on the location of the pathology. at this point the radiologist anna opens up a somewhat different, but still related, point in relation to the specific materials that are currently displayed on the screens. as will be clear, on the topical level this stretch of talk is replete with the reported troubles of perception and understanding. in the short sequence examined here, the expression “to perceive” occurs no less than three times, and “to understand” is set up as a contrast to “believe”. the example is thus an endogenous formulation that speaks about some perceptual difficulties generated by the introduction of the new technology. the visual expertise of radiological diagnosis is thus both being demonstrated here and made into the very topic for the discussion. however, rather than raising this as a general type of problem (which the cognitively associated terms could suggest) it is cast as a “setting’s trouble”, a form of problem that builds on, and refers to, the knowledge and practices shared by parties to that setting. still, as we will see, the articulation of this trouble is not straightforward, nor is it done by mere talk. anna commences by, what the sociologist doug maynard calls, an “embodied telling of a seeing” (2006, p. 107). figure 1. establishing the referential grounds. anna starts her contribution with the word “but”, a disjunction marker which could be heard as making a slight shift in topic in relation to previous talk. what follows is the formulation of this topic, i.e. “the tricky thing with tomosynthesis”. this initial characterization of a trouble, functions as a preface to a telling (sacks, 1974) and thus as a “prospective indexical” (goodwin, 1996), an indexical expression whose referent is to be specified in ensuing talk. 3 pleural plaques are characterised by areas of fibrous thickenings on the lining of the lungs. although benign (not cancerous) they are the most common indication of significant exposure to asbestos. ivarsson | f l r 130 the group sits across from two separate projector screens set side-by-side, one showing the ct and the other the tomosynthesis data. in [1-3] anna’s attention is directed at the rightmost screen showing the tomosynthesis image [4] and she also makes a brief, but fully extended, pointing gesture towards this screen [3]. next, [5] there is a shift in direction: topically, bodily and referentially. she moves her already extended arm to the left so to point at the adjacent projector screen [6]. the words “a thing like that” also begins to specify the indexical referent more precisely. from the recording it is evident that the two objects pointed to here are both regarded as visible and publically available for everyone present at the session. the very finding of those objects is not the primary concern in this instance. however, it should be acknowledged that this was not always the case. the issue of discovering possible pathologies is a prerequisite for any subsequent diagnostic work and the increase in detection rates was also one of the critical features of tomosynthesis (vikgren et al., 2008). in this short passage, the two referential objects pointed to in succession (here marked by the added arrows in the tomosynthesis section image [4] as well as the ct counterpart [6]) become unified in that they are treated as denoting a single physical structure situated elsewhere. this “thing” is a previously discussed pathology: a plaque with a pleural basis4. figure 2. representing digital manipulations. with the arm extended towards the ct screen and her hand held flat, anna makes two cutting or slicing movements. although done at a distance, this gesture builds on and gives meaning to a specific structure of the environment, namely the ct image [6]. it is an environmentally coupled gesture (goodwin, 2007a) which is done to indicate a cutting across the visible plaque [7]–[8]. furthermore, this gesture follows the anatomical plane known as the sagittal plane which roughly divides the patient’s left and right sides. her comment speaks about what-we-all-see in this image, as the object being in the centre of the lung [8]. through the response token from maria (“m”), there is confirmation that the argument is being followed thus far. 4 although pathological, a plaque is not a pulmonary nodule and mistaking it for one, as some of the radiologists had done, constitutes a case of a “false positive.” the ensuing reasoning exhibited here is aimed at minimizing such mistakes in the future. ivarsson | f l r 131 having both secured the attention of the group and established the referential grounds for her further work, anna returns to the main topic of “the tricky thing”. it is now developed into “the particular thing about depth” [9]. while keeping her forearm in place, she turns her wrist so that the back of the hand faces the screens (in terms of the anatomical planes this would be analogous to shifting from the sagittal to the coronal plane), and, as the word “depth” is produced with emphasis, she simultaneously moves the hand away from her. the significance of this gesture should be understood in relation to its material and social environment. as argued by lebaron and streeck: it is our contention that gesture — certainly descriptive or ‘iconic’ gesture — necessarily involves indexical links to the material world, even though these links are rarely established or explicated in the communicative situation itself. rather, in conversational contexts that are detached from the talked-about world, participants must fill in encyclopaedic knowledge (ranging from universal bodily experiences to highly specific cultural practices) to see and recognize gestures. (lebaron & streeck, 2000, p. 131) the particular movement made by anna [9] thus represents a case of such a highly specific cultural practice, an action that is most central to professional radiologists. in addition to indicating a movement in the ventral direction (toward the front of the patient) it also resembles one way of navigating in a set of section images. this is one of the central methods through which the sense of volume and location is built. in this way, the gestural action is not simply an embellishment to the talk (kendon, 1972), rather, it works to indexically tie the meaning of the word “depth” to a material and everyday radiological practice, known and shared amongst the participating radiologists. the gesture becomes part of what koschmann and colleagues (2007) characterise as a gestural formulation, which in its design displays the speaker’s analysis of whom she is addressing; the gesture is selected and shaped because of its presumed recognisability to the members of the setting (schegloff, 1972). figure 3. perceiving pleura. in [10] anna starts a new construction with “when”. this clause begins an attempt to establish the setting for the exposition to come. she keeps her hands flat before her as if reading a plane x-ray or tomosynthesis image. after a couple of cut-of’s and pauses she changes the sentence frame and in [11-12] again refers to the tomosynthesis image. she makes a deictic gesture toward the left-hand screen [11], then returns to regarding the flat image [12]. now the “tricky thing” is related to the act of perceiving the location of the object. the use of the adverb “pleurally” also performs classificatory work: since their search for nodules delimits their interest to objects that are located inside the lung, anything found in the pleura, the layer covering the lung, is to be disregarded in this particular task. ivarsson | f l r 132 after the clause “to perceive that that one is situated pleurally” there is another try at starting a new clause with “when”. also this time the attempt is abandoned and we get a further qualification of the anatomical basis for the trouble; “since the chest vaults” [13]. again we find gesturing that is closely coupled with the practice of thoracic radiology. anna is continuously amalgamating the materials of anatomical and technological concepts, visual representations and embodied gestures alongside her otherwise vernacular talk. the gesture accompanying the entire stretch of talk [in 13] is repeated twice, and, in effect, maps this vaulting feature on to a generalized body. the frame of reference taken is one of an external observer where the object is created before her eyes. but in the visual contrast between the flattened image and the portrayed 3-d object we get an early hint of a complication which becomes articulated next. figure 4. subjective involvement. next, the consequences of this rounded shape of the lung is commented on. the act of perception is again introduced into the talk, and this time it is formatted as reported speech “where in the lung am i”. simultaneous to the verbal comment about location anna taps her chest twice [14]. at this point, her own body is enrolled as a new referential ground. she has thereby established a transition in the frame of reference, from that of an external reader of images to that of an idealized patient. in addition to the two screens, and the gesture space in front of her, she is now also designing her comments so that they make sense in relation to her upper body. in [15], the third attempt at starting with a “when”-construction is brought to its completion. together with her pointing gestures, it does the job of building a contrast between two locations [15-17]. in the two demonstrations, she is also timing her pointings with the deictic terms “he:re” [16] and “the:re” [17] (hindmarsh & heath, 2000). by shifting her gaze to where she is pointing she highlights the gesture for her ivarsson | f l r 133 interlocutors, but at the same time makes the gesture a tool for her own understanding (lebaron & koschmann, 2003). so, what is the role of her body, and the possible reasons for bringing it into the interaction here? for one thing, it should be seen as doing communicative work in relation to the other participants. it has on the one hand, a rhetorical side, as a developing embodied argument (mirivel, 2011) which is clearly recipient designed (schegloff, 1972) for present parties. however, what anna has to do is not merely to “read” the images at hand and present that reading to her colleagues. in the context of the self-reflective situation set up by the team, she’s struggling to express her current understanding of the relation between the unique (forthis-person) three-dimensional space constituted by the patient’s body, and how this is first mediated via the imaging technology and later represented in the radiographs. some of the work of formulating this relationship has a speculative or exploratory quality to it. and in the combination of these two aspects of communication-cum-speculation we find the specific features of the setting. the extensive gesturing carried out by anna becomes a way of organizing her reasoning so that it is made available to her peers. it thereby takes the form of a provisional radiological reasoning done of and for the setting. furthermore, it is done not so much to diagnose this patient, as to provide materials for generic formulations that speak to the renegotiation of this group’s expertise (for an analysis of this practice see lymer et al., 2014). this form of publically oriented professional reasoning is partly done by way of a subjective involvement with the graphical materials (cf., ochs, gonzales, & jacoby, 1996). the subjective involvement is accomplished by linguistic as well as gestural means: by grammatically placing herself in the patient’s body (e.g., “where in the lung am i” [14]), and, by gesturally positioning the observed structure as if this was located in her own body [16] & [17]. the interchangeability of these frames of reference further suggests close links between talk, gesture and the material environment. it is not because the professionals routinely handle the bodies of patients that a physical body is being involved in the argument here—radiologists predominantly work on representational objects. however, here and now anna’s body provides a threedimensional structure aiding the installation of a specific contrast. in other words, her body is made to double as a scaffold in the developing formulation of how tomosynthesis depicts volumetric information. figure 5. perceiving depth. immediately following the establishment of the two separate locations there is a third iteration speaking about the activity of perceiving [18] and [19]. this time, and in comparison to [9], [11] and [14], the formulation has become more succinct: “to perceive the depth in tomosynthesis”. the previous “vaulting” gesture is also reused and laminated with the described problem of perception [18], as is the “depth/ventral” gesture again overlaid with the word “depth” [19]. according to mcneil and levy (1993), when gestural forms are reused in this way they often mark the reappearance of a particular plot element, ivarsson | f l r 134 character, or narrative value. such gestural cohesion, as they call it, is thus an aspect of the process of creating and maintaining topical cohesion across turns at talk (mcneill & levy, 1993). here we find the two themes of, first, the shape of the thoracic cavity, and, second, the work of navigating in a stack of images to be gesturally reintroduced. in [20] anna further qualifies the categorization, or the distinction, that is at stake. next, she returns to, and elaborates on, the contrast. figure 6. creating contrast. with her right hand held vertically at the front of her chest, the first location is specified as “in the front” [21], whereupon she provides the alternate location. introduced as something surprising she points further back and to the side of the chest. this time she is using her left index finger while the right hand remains flat on the top of her chest [22]. through this particular configuration of talk, body and embodied action, the contrasting locations have now become publically visible. the contrast being created here revisits the contrast described earlier [15-17] and we see again some recycled gestures. where the problem before was characterized as one of accurately assessing depth within an imagined image, the problem is here made more vivid by illustrating how the vagaries of the tomosynthetic image could radically skew the evaluation of location in a human body. figure 7. wrapping up. ivarsson | f l r 135 finally, in [23-25] anna summarizes the argument and spells out the trouble in a non-technical terminology. this concludes her extended contribution. another participant, lena, adds a comment about using the ribs and when they come into focus as a useful method for determining location. the comment is briefly elaborated and four of the participating radiologists then close this particular discussion on the note that “it is still difficult.” 5. discussion the interest of this article has been to address the topic of visual expertise and to do so in terms that acknowledge that all expertise is rooted in embodied practices. to this end the episode discussed above has served as an example of what we can call the enacted production of radiological reasoning. albeit a special case, the expertise of radiological diagnosis that comes into play here is first demonstrated through the actions performed, but it is also talked about in the studied discussion. in this talk, terms such as ‘seeing’, ‘understanding’ and ‘perceiving’ figure as members’ matters. to quote slack and his colleagues on this phenomenon: to be sure, members speak of seeing, noticing and other topics that are grist to the mentalists’ mill, but we have shown that they do so not in an isolated context (neither behind the skull nor as atomic ‘cognisers’) but in a manner where terms such as ‘seeing’, ‘noticing’ and so on are practical members’ achievements, achieved in and through natural language and embodied conduct. (slack et al., 2007, p. 192) these ‘practical achievements’, or actions, show us a corporeal side of the expertise in interpreting visual representations. not only are visual phenomena—in this case the existence of a pleurally based plaque—made into something observable and reportable through the deployment of a professional language. one of the central features of the above illustration is the manner in which this work is also carried out by way of the living present body of the expert. the experienced radiologist interweaves anatomical and technological terminology with visual representations and gestures in such a way that none of these components can be said to be superfluous to the argumentation. furthermore, the sequence encompasses, not gestures as a general phenomenon, but as specialized embodied conduct indexical (i.e. uniquely fitted) to projected images, practical actions, or specific locations in patient-bodies. so, by building on, and referring to, the matters and routines known and shared by the parties to the setting, these gestural actions also work to anchor the meaning of the talk in material and everyday radiological practice. in this capacity, the gestures act as aids in the bridging of interpretative gaps between the radiographic renderings and what those come to mean as professionally relevant objects. in the studied case these interpretative difficulties have been aggravated, due to the extraordinary situation of the newly introduced radiographic technology of tomosynthesis. however, this situation also allows for several fruitful observations. what is pulled into view are some transmutations, or movements from one medium to another, from the technologically mediated body of the patient into formulations of members’ understandings. and at this intersection we come very close to what is ordinarily regarded as expertise. we find exhibited production procedures through which the body of the patient is coordinated with the skilled body of the practitioner. without downplaying the relevance of functioning eyes and brains, the approach exemplified here can help us to appreciate gestures and embodied actions as critical means through which visual expertise becomes organised. these are parts of a repertoire of methods through which the radiologists learn their profession, and, as is evident here, they can also become enrolled in the renegotiation of expertise in the face of new challenges. ivarsson | f l r 136 transcription legend (0.5) numbers in parentheses indicate silence, represented in tenths of a second. (.) a dot in parentheses indicates a “micropause.” a hyphen after a word or part of a word indicates a cut-off or self-interruption. :: colons are used to indicate the prolongation or stretching of the sounds just proceeding them. word underlining is used to indicate some form of stress or emphasis. [ separate left square brackets on two successive lines indicate the onset of overlapping [ or simultaneous talk. > < the combination of “more than” and “less than” symbols indicate that the talk between them is compressed or rushed. acknowledgements the study has been funded by the swedish research council (2015-03621) and has been a part of the letstudio—a strategic initiative for promoting interdisciplinary research within the learning sciences at the university of gothenburg. i would like to express my gratitude to all partners involved in this inspiring collaborative undertaking. references alac, morana. (2008). working with brain scans: digital images and gestural interaction in fmri laboratory. social studies of science, 38(4), 483-508. doi: 10.1177/0306312708089715 asplund, sara, johnsson, åse a, vikgren, jenny, svalkvist, angelica, boijsen, marianne, fisichella, valeria, flinck, agneta, wiksell, åsa, ivarsson, jonas, rystedt, hans, månsson, lars gunnar , kheddache, susanne, & båth, magnus. (2011). learning aspects and guidelines regarding detection of pulmonary nodules and developing quality criteria for chest tomosynthesis. acta radiologica, 52, 503-512. doi: 10.1258/ar.2011.100378 garfinkel, harold. (1967). studies in ethnomethodology. englewood cliffs, nj: prentice-hall. garfinkel, harold. (2002). ethnomethodology's program: working out durkheim's aphorism. lanham: rowman & littlefield publishers. gibson, james j. (1968). the senses considered as perceptual systems. london. gibson, james j. (1979). the ecological approach to visual perception. boston, ma: houghton mifflin. goodwin, charles. (1994). professional vision. american anthropologist, 96(3), 606-633. goodwin, charles. (1996). transparent vision. in e. ochs, e. a. schegloff & s. thompson (eds.), grammar and interaction. cambridge, ma: cambridge university press. goodwin, charles. (2000). action and embodiment within situated human interaction. journal of pragmatics, 32, 1489-1522. goodwin, charles. (2007a). environmentally coupled gestures. in s. duncan, j. cassel & e. t. levy (eds.), gesture and the dynamic dimensions of language (pp. 195-212). philadelphia: john bejamins. goodwin, charles. (2007b). participation, stance, and affect in the organization of activities. discourse and society, 18(1), 53-73. griesemer, james. (2004). three-dimensional models in philosophical perspective. in s. de chadarevian & n. hopwood (eds.), models: the third dimension of science (pp. 433-442). stanford, ca: stanford university press. ivarsson | f l r 137 have, paul ten. (2004). understanding qualitative research and ethnomethodology. london: sage. hindmarsh, jon, & heath, christian. (2000). embodied reference: a study of deixis in workplace interaction. journal of pragmatics, 32, 1855-1878. hirschauer, stefan. (1991). the manufacture of bodies in surgery. social studies of science, 21, 279-319. johnsson, åse a, vikgren, j, svalkvist, a, zachrisson, sara, flinck, a, boijsen, m, kheddache, s, månsson, l g, & båth, magnus. (2010). overview of two years of clinical experience of chest tomosynthesis at sahlgrenska university hospital. radiation protection dosimetry, 1-6. kendon, adam. (1972). some relationships between body motion and speech. in a. siegman & b. pope (eds.), studies in dyadic communication (pp. 177-210). new york: pergamon press. kendon, adam. (1997). gesture. annual review of anthropology, 26, 109-128. koschmann, timothy, & lebaron, curtis. (2002). learner articulation as interactional achievement: studying the conversation of gesture. cognition and instruction, 20(2), 249-282. koschmann, timothy, lebaron, curtis, goodwin, charles, & feltovich, p. (2011). 'can you see the cystic artery yet?' a simple matter of trust. journal of pragmatics, 43(2), 521-541. koschmann, timothy, lebaron, curtis, goodwin, charles, zemel, alan, & dunnington, gary. (2007). formulating the triangle of doom. gesture, 7(1), 97-118. kundel, harold l, nodine, calvin f, & toto, lawrence. (1991). searching for lung nodules. the guidance of visual scanning. investigative radiology, 26, 777-781. lebaron, curtis, & koschmann, timothy. (2003). gesture and the transparency of understanding. in p. glenn, c. lebaron & j. mandelbaum (eds.), studies in language and social interaction (pp. 102-112). mahwah, nj: lawrence erlbaum. lebaron, curtis, & streeck, jürgen. (2000). gestures, knowledge, and the world. in d. mcneill (ed.), language and gesture (pp. 118-138). cambridge: cambridge university press. lesgold, a., rubinson, h., feltovich, p., glaser, r., klopfer, d., & wang, y. (1988). expertise in a complex skill: diagnosing x-ray pictures. in m. chi, r. glaser, & m. farr (eds.), the nature of expertise (pp. 310-342). london: routledge. liberman, kenneth. (2015). turn-taking in the surfer's lineup. surfertoday.com. http://www.surfertoday.com/surfing/12275-turn-taking-in-the-surfers-lineup-an-academic-analysisby-kenneth-liberman lindwall, oskar, & ekström, anna. (2012). instruction-in-interaction: the teaching and learning of a manual skill. human studies, 35, 27-49. livingston, eric. (1999). cultures of proving. social studies of science, 29(6), 867-888. livingston, eric. (2008). ethnographies of reason. hampshire: ashgate. lynch, m. (1993). scientific practice and ordinary action: ethnomethodology and social studies of science. cambridge: cambridge university press. lymer, gustav. (2010). the work of critique in architectural education. göteborg: acta universitatis gothoburgensis. lymer, gustav, ivarsson, jonas, rystedt, hans, johnsson, åse allansdotter, asplund, sara, & båth, magnus. (2014). situated abstraction. from the particular to the general in second order diagnostic work. discourse studies, 16(2), 182-212. doi: 10.1177/1461445613514674 maynard, douglas w. (2006). cognition on the ground. discourse studies, 8(1), 105-115. mcneill, david, & levy, elena t. (1993). cohesion and gesture. discourse processes, 16(4), 363-386. doi: 10.1080/01638539309544845 melander, helen, & sahlström, fritjof. (2009). in tow of the blue whale: learning as interactional changes in topical orientation. journal of pragmatics, 41(8), 1519-1537. doi: http://dx.doi.org/10.1016/j.pragma.2007.05.013 merleau-ponty, maurice. (1962). phenomenology of perception. london: routledge & kegan paul. mirivel, julien. (2011). embodied arguments: verbal claims and bodily evidence. in j. streeck, c. goodwin & c. lebaron (eds.), embodied interaction: language, and body in the material world (pp. 254-263). cambridge: cambridge university press. http://www.surfertoday.com/surfing/12275-turn-taking-in-the-surfers-lineup-an-academic-analysis-by-kenneth-liberman http://www.surfertoday.com/surfing/12275-turn-taking-in-the-surfers-lineup-an-academic-analysis-by-kenneth-liberman http://dx.doi.org/10.1016/j.pragma.2007.05.013 ivarsson | f l r 138 myers, natasha. (2008). molecular embodiments and the body-work of modeling in protein crystallography. social studies of science, 38(2), 163-199. doi: 10.1177/0306312707082969 ochs, elinor, gonzales, patrick, & jacoby, sally. (1996). "when i come down i'm in the domain state": grammar and graphic representation in the interpretive activity of physicists. in e. ochs, e. a. schegloff & s. a. thompson (eds.), interaction and grammar (pp. 328-369). cambridge, ma, england: cambridge university press. prentice, rachel. (2005). the anatomy of a surgical simulation: the mutual articulation of bodies in and through the machine. social studies of science, 35(6), 837-866. doi: 10.1177/0306312705053351 rystedt, hans, ivarsson, jonas, asplund, sara, johnsson, åse allansdotter, & båth, magnus. (2011). rediscovering radiology. new technologies and remedial action at the worksite. social studies of science, 41(6), 867-891. doi: 10.1177/0306312711423433 sacks, harvey. (1974). an analysis of the course of a joke's telling in conversation. in r. bauman & j. sherzer (eds.), explorations in the ethnography of speaking (pp. 337-353). cambridge: cambridge university press. schegloff, emanuel a. (1972). notes on a conversational practice: formulating place. in d. sudnow (ed.), studies in social interaction (pp. 75-119). new york: macmillan. slack, roger, hartswood, mark, procter, rob, & rouncefield, mark. (2007). cultures of reading: on professional vision and the lived work of mammography. in s. hester & d. francis (eds.), orders of ordinary action (pp. 175-193). aldershot: ashgate. streeck, jürgen. (2009). gesturecraft. the manufacture of meaning. amsterdam: john benjamins. streeck, jürgen. (2015). embodiment in human communication. annual review of anthropology, 44(1), 419-438. doi: doi:10.1146/annurev-anthro-102214-014045 sudnow, david. (1978). ways of the hand: the organization of improvised conduct. cambridge, ma: harvard university press. vikgren, jenny, zachrisson, sara, svalkvist, angelica, johnsson, åse a , boijsen, marianne, flinck, agneta, kheddache, susanne, & båth, magnus. (2008). comparison of chest tomosynthesis and chest radiography for detection of pulmonary nodules: human observer study of clinical cases. radiology, 249, 1034-1041. frontline learning research vol. 5 no. 3 special issue (2017) 1-13 issn 2295-3159 methodologies for studying visual expertise andreas gegenfurtner1 & jeroen j. g. van merriënboer2 1 technische hochschule deggendorf, germany 2 maastricht university, the netherlands article received 14 july / revised 23 july / accepted 23 july / available online 31 july abstract visual expertise can be defined as maximal adaptation to the requirements of a visionintensive task. the process of developing a “good eye” in vision-intensive tasks is proposed, indicated, and elaborated by various measures contingent on diverse methodological arenas, all of which attempt to advance our understanding of what constitutes visual expertise. the aim of this special issue is to provide a reflection on this methodological pluralism and to offer a discussion of the affordances and constraints of some of these methodological approaches. specifically, grounded on the medical domain, this special issue brings together a selection of nine articles that discuss cognitiveneurosciences, receiver operating characteristics (roc) analysis, eye tracking, pupillometry, the flash-preview moving window paradigm, the combination of eye tracking data and verbal report data, the use of interviews and verbal protocols, ethnomethodology, and the expert performance approach. two commentaries conclude the special issue. as an introduction, this article presents a comparative metaphorical mapping of visual expertise research. metaphors are a useful tool for mirroring in simple terms the often complex paradigms underlying theory and applied research practice. we first identify four metaphors used in the analysis of visual cognition: activation, detection, inference, and practice. these metaphors are described with an empirical example and discussed to elicit (partly tacit) assumptions associated with prototypical method decisions. we then link the proposed metaphorical mapping to the contributions in this special issue. keywords: visual expertise; expert performance; visual cognition; professional vision; methods in learning research. 1 corresponding author: andreas gegenfurtner, technische hochschule deggendorf, institut für qualität und weiterbildung, dietergörlitz-platz 1, 94469 deggendorf, germany, email: andreas.gegenfurtner@th-deg.de doi: http://dx.doi.org/10.14786/flr.v5i3.316 mailto:andreas.gegenfurtner@th-deg.de http://dx.doi.org/10.14786/flr.v5i3.361 gegenfurtner et van merriënboer 2 | f l r 1. introduction this special issue is devoted to research on visual expertise. visual expertise can be defined as maximal adaptations to the requirements of a vision-intensive task. examples of vision-intensive tasks include the identification of different types of fish (boucheix & lowe, in press; jarodzka, scheiter, gerjets, & van gog, 2010) or the detection of abnormalities in microscopic specimen (helle, nivala, kronqvist, gegenfurtner, björk, & säljö, 2011; krupinski, graham, & weinstein, 2013). in many professions, visual material constitutes an important part of the epistemic resources used for conducting professional work (gegenfurtner, nivala, säljö, & lehtinen, 2009; goodwin, 1994; gruber & degner, 2016; palonen, boshuizen, & lehtinen, 2014; säljö, 2012). consequently, newcomers need to learn, appropriate, and master the skills associated with domain-specific visual tasks (gegenfurtner, lehtinen, jarodzka, & säljö, 2017a; kok, van geel, robben, & van merriënboer, 2017; seppänen & gegenfurtner, 2012; szulewski, gegenfurtner, sivilotti, howes, & van merriënboer, in press). past research has employed very different strategies to examine these learning processes of novices as well as the processes and practices underlying the superior performance of domain experts. different research methodologies are associated with different underlying epistemologies (damsa et al., in press). therefore, the phenomenon of visual expertise is approached from different perspectives that correspond with different assumptions about what constitutes the allegorical “good eye” of an expert. because expertise and expert performance are largely domain-specific (gegenfurtner & seppänen, 2013; gruber & degner, 2016), this special issue focuses on one example domain: medicine. medicine has been chosen here because many medical specialties rely on the analysis of visual material, such as the human skin, x-rays, pathological slides, or electrocardiograms. although the focus is on medicine, the lessons learned from a reflection on methodologies for studying visual expertise can inform other vision-intensive domains to advance our understanding of an expert’s professional vision (goodwin, 1994). some investigators in medicine, particularly kundel, nodine, and carmody (1978), in a now classic study, suggest that the highly perceptual nature of image comprehension requires intensive processing of visual data through oculomotor activity that guides signal detection and decision-making. others, particularly lesgold and colleagues (1988), in what is now also a classic study, put less emphasis on the perceptual aspect; rather, they suggest that visual expertise is mainly the function of cognitive inference that aligns schemata from episodic memory consistent with the perceptual features detected. much medical research done in both traditions has been reviewed elsewhere (boshuizen & schmidt, 2008; ericsson, 2004; gegenfurtner, kok, van geel, de bruin, jarodzka, szulewski, & van merriënboer, 2017; gegenfurtner, siewiorek, lehtinen, & säljö, 2013; patel, arocha, & zhang, 2005). recently, alternative approaches have been formulated; these suggest that a good eye is indicated by neurophysiologic events in certain brain areas (bilalić, 2017; gegenfurtner, kok, van geel, de bruin, & sorger, 2017b; haller & radue, 2005), and is accomplished through situated social discourse (ivarsson, 2017; johansson, lindwall, & rystedt, 2017; koschmann, lebaron, goodwin, zemel, & dunnington, 2007). in short, the allegory of having a good eye in medicine is proposed, indicated, and elaborated by various measures contingent on diverse methodological arenas that all attempt to advance our understanding of what constitutes visual expertise. the main aim of this special issue is to provide a reflection on this methodological pluralism. for this purpose, we have invited scholars from various professional backgrounds to contribute an article that introduces a particular methodological approach used to study visual expertise. each article is devoted to one approach (or to a combination of approaches) and discusses its affordances and constraints for empirically analyzing visual expertise. these approaches are: cognitive-neurosciences (gegenfurtner et al., 2017b), gegenfurtner et van merriënboer 3 | f l r receiver operating characteristics analysis (krupsinki, 2017), eye tracking (fox & faulkner-jones, 2017), pupillometry (szulewski, kelton, & howes, 2017), the flash-preview moving window paradigm (litchfield & donovan, 2017), the combination of eye tracking data and verbal report data (helle, 2017), the use of interviews and verbal protocols (van de wiel, 2017), ethnomethodology (ivarsson, 2017), and the expert performance approach (williams, fawyer, & hodges, 2017). the special issue closes with two commentaries (boucheix, 2017; jarodzka & boshuizen, 2017). before we introduce each contribution, we would like to present a framework that aims to structure the pluralism of methodologies in visual expertise research. 2. four metaphors in visual expertise research our goal is to present a framework in which the different metaphors of having a “good eye” can be considered as mutually constituting the richness we have on ideas, concepts, and theories relating to visual expertise in medicine and beyond. we have no intent to judge some methodological traditions as being more valuable than others, nor do we intend to unify them in some abstract way. although the idea of what was termed interactive complexification (alexander, schallert, & reynolds, 2009) is considered meaningful for highlighting the confluence of factors that determines “any given aspect” of the product and the process of learning to diagnose medical images, we believe there is, at times, also a place for simple answers. we hope that the framework put forward in this article will provide such a simple answer that can be used as a glass through which we look at the phenomenon of visual expertise. the framework addresses the questions of how to analyze visual expertise and how to elicit tacit assumptions underlying common research practice. by discussing four examples drawn from radiology, we identify four metaphors that constitute our framework on visual expertise research. despite its groundwork in medical literature, the framework also has implications beyond medicine, due to its common interest in learning and comprehension in technology-rich visionintensive contexts. the framework is built around four metaphors. metaphors are a useful tool for mirroring in simple terms the often complex paradigms underlying theory and applied research practice. an example of how metaphors are able to elicit complex, different, yet partly tacit, assumptions on a commonly studied phenomenon can be found in “learning as acquisition”, “learning as participation”, and “learning as knowledge creation” (paavola, lipponen, & hakkarainen, 2002; sfard, 1998; see also hakkarainen, palonen, paavola, & lehtinen, 2004). as sfard (1998, p. 4) notes, “metaphors are the most primitive, most elusive, and yet amazingly informative objects of analysis”. we believe that their value and power stems from the fact that metaphors converge and portray, in a snapshot format, what took years of scientific discourse to develop; this allows frank presentation of positions and their entailments in a condensed way and invites a critical (re)consideration of accepted and perhaps unreflected practice. of course, metaphors are simple and simplistic; there is no claim that they attempt to depict all of the breadth and depth of what often is a complex epistemology. gegenfurtner et van merriënboer 4 | f l r table 1 a comparative metaphorical mapping of visual expertise as activation, detection, inference, and practice activation detection inference practice indicators of visual expertise neurophysiologic activity eye movements verbal reports representational practices unit of analysis individual individual (social) individual and social sociotechnical place of visual cognition neural network system optic system (distributed) memory system activity system analytic time span milliseconds to seconds seconds minutes to few hours minutes to decades associated methodology cognitive neuroscience roc analysis; eye tracking methodology protocol analysis; interviews ethnomethodology, ethnography below, we identify, exemplify, and discuss four metaphors often used when analyzing visual expertise; these metaphors are seen as four of the many dimensions possessed by one with “a good eye”. we reflect on these metaphors in terms of how they contribute to research devoted to examining visual expertise. it is hoped that the value and significance of this methodological reflection will be that it helps to map a scattered and fragmented field of inquiry. table 1 serves as the guiding framework for the comparison of the four metaphors, including their methodological entailments: activation, detection, inference, and practice. 2.1. visual expertise as activation metaphor research adhering to the activation metaphor uses measures of neurophysiologic activity as an indication of visual expertise. in the activation metaphor, there is a strong emphasis on the neurological and biological basis of our humanness (alexander et al., 2009; see also meltzoff, kuhl, movellan, & sejnowski, 2009). this emphasis might originate from the widely held belief that “information is stored in neural networks in the brain, and that human behavior arises from extremely complex communication between neurons in these networks and also between separate networks or assemblies” (sauseng & klimesch, 2008, p. 1003). this neural network system is seen as the place where visual cognition and expertise “occurs”. hence, visual expertise is indicated by neural activity. specifically, this activity can be measured by an electroencephalograph (eeg) as the electric current in axons; by a magnetoencephalograph (meg) as the magnetic field induced by those electric currents; by a positron emission tomography (pet) as the blood flow distribution in the cells; or by functional magnetic resonance imaging (fmri) scanning as differences in cellular oxygen consumption. whenever information stored in neural networks is used for cognitive processes, neural activation can be measured with one of those tools. for example, if a radiologist formulates a diagnosis based on a patient’s medical image, an fmri scanner could be used to indicate the processes of this radiologist’s visual cognition. these processes are extremely fast; the best conventional apparatus currently available—the eeg scanner—is able to trace this activity with a temporal resolution in the range of milliseconds (sauseng & klimesch, 2008). an empirical example prototypical for the activation metaphor can further illustrate epistemological and methodological premises. gegenfurtner et van merriënboer 5 | f l r haller and radue (2005) investigated differences in neuronal activations of radiologists and laypersons in reading radiologic and non-radiologic images. using functional magnetic resonance (fmri) imaging, the brain scans showed that radiologic images evoked stronger activations in the brains of radiologists than in those of laypersons, with the bilateral middle and inferior temporal gyrus, bilateral medial and middle frontal gyrus, and left superior and inferior frontal gyrus being particularly affected. these regions are generally assumed to be linked to the encoding and storing of memory of visual objects and events. hence, this finding seems to imply that what is seen on the presented image is automatically referenced to memorized images, indicating an unconscious, stimulus-driven indexical relation between the pictorial representation and the corresponding mental representation. being prototypical for research in neuroscience, haller and radue (2005) used technological images as stimuli in a series: stimuli were shown for 2.5 seconds followed by a fixation cross for 8.5 seconds to compensate for blood oxygenation level-dependent signal delay. subjects gazed at the stimuli series with an immobilized head in darkened and (electrical and auditory) noise-protected rooms. settings of this kind are highly controlled. these types of controls are necessary because neural measures are highly sensitive: activation should be shown in response to the stimulus only. strong controls therefore aim to guarantee bias-free recordings. at this point in time, the activation metaphor for visual expertise has rarely been used in medical diagnosis studies (gegenfurtner et al., 2017b). however, the coming together of learning research and neuroscience is beginning to form an exciting new field (ansari & coch, 2006; de jong, van gog, jenks, manlove, van hell, jolles, et al., 2009). neuroscience has the potential to trace implicit and experiential learning before it can be observed in behavior. this can help us understand when, how, and why learning occurs. in particular, the how and why can be tackled within these highly controlled settings. certainly, it is not a novel statement that behaviors, such as diagnosing a medical image, that appear similar on the surface may involve very different cognitive/perceptual mechanisms underlying this behavior. neuroscience, in combination with the learning sciences, now provides a new avenue for tackling these issues, to further understand visual expertise (bilalić, 2017; gegenfurtner et al., 2017b). 2.2. visual expertise as detection metaphor detection can be defined as “determining whether a simple, featurally defined stimulus is present in, or absent from, the visual field” (smith & ratcliff, 2009, p. 283). a central premise of research using the detection metaphor is uncertainty; that is, the degree to which a subject is able to discriminate between signal (the stimulus of interest) and noise (background stimuli distracting visual attention, thus causing decisionmaking under conditions of uncertainty). in medical image diagnosis, the signal would be a tumor, while noise would be (healthy) organic material surrounding the tumor. clearly, in pictures as complex as radiographs, with an abundance of structures, forms, and elements displayed in manifold shadings of grey, and with the typical presence of technical artefacts, detection of a tumor is a challenging task. tasks of this kind are frequently used to quantify the ability of discerning between signal and noise. two approaches are prevalent: eye-tracking methodology and receiver operating characteristic (roc) analysis. below, an empirical example by kundel, nodine, conant, and weinstein (2007), which combined eye-tracking with roc analysis, can illuminate prototypical premises of each approach. kundel et al. (2007) investigated rapid initial fixations (detections) on abnormalities on mammograms. briefly, they found that more experienced radiologists showed global perceptual processes that helped them detect the abnormality (malignant breast lesions) in less than a second. in contrast, less experienced radiologists showed search-to-find strategies that took considerably longer to first fixate the abnormality. gegenfurtner et van merriënboer 6 | f l r expertise differences in these two groups were indicated by eye-tracking and by roc analysis. with respect to eye-tracking, the recording of eye movements is usually used to visualize the scan paths of observers. in kundel’s study, radiologists with more experience had longer saccades and fewer fixations than did less experienced radiologists. with respect to roc analysis, detectability was significantly higher for observers with more experience than for those with less experience. detectability is a measure that quantifies the sum of true positives and true negatives, divided by the sum of all positives and negatives in a detectability value, da. a conclusion can be made that the detection metaphor indicates that novice diagnosticians develop a good eye through medical education in terms of their ability to discriminate a potential signal from background noise. this ability can be quantified and expressed mathematically in a formula that allows comparison of observers at the individual or the group level. studies in the detection metaphor, which usually employ eyetracking methodology and/or roc analysis, indicate that superior visual cognition can be characterized as a high decision-speed accuracy relation: visual perception changes with a rise in experience, from a relatively slow search-to-find mode to a global holistic mode. this change then increases sensitivity (i.e., proportion of correctly identified abnormalities), specificity (i.e., proportion of correctly identified healthy tissue), and thus accuracy of the detection performance. usually, the analytic time span is somewhat longer than the time span of cognitive neuroscience studies. the work of kundel et al. (2007), which can be seen as a prototypical example, reported an average search time of 26.90 seconds, and a median time to first fixate the abnormality (the signal) of 1.13 seconds. an exciting direction for further research adopting the detection metaphor is a focus on the transfer of expertise. transfer is a concept to describe how skills in one field are applicable in a second field (gegenfurtner et al., 2010; quesada-pallarès & gegenfurtner, 201). in the context of expertise research, eye tracking (and other methods as well) can be used to study how visual expertise transfers from a domain-specific, typical, or routine task to a domain-general, atypical, or novel task (gegenfurtner & seppänen, 2013; gegenfurtner et al., 2017c). another exciting research direction is collaborative gaze. traditionally, eye-tracking studies have focused on individual observers as the unit of analysis (seppänen & gegenfurtner, 2012); however, developments of eye-tracking technology and analytic algorithms now allow collaborative gaze studies (e.g., sangin, molinari, nüssli, & dillenbourg, 2008). it will be fascinating, from an epistemological point of view, to follow the coping of tension between attentional detection as an individual quantifiable performance, notable in mathematical functions (i.e., smith & ratcliff, 2009), on one hand, and detection as collaborative achievement and intersubjective meaning-making (much in line with koschmann & zemel’s, 2009, notion of discovery as occasioned production), on the other hand. 2.3. visual expertise as inference metaphor lesgold and colleagues (1988, p. 336) speculated that radiological diagnosis “is largely a matter of cognitive inference. that is, given a set of findings (perceptual features), one has to determine which diseases are consistent with those findings. if more than one disease is consistent, then one either looks further, (…) or suggests additional medical tests to discriminate among the possibilities”. two issues are striking in this initial quote. first, lesgold emphasizes cognition and memory processes in diagnosing medical images. back in the 1980s, this was not customary in the medical literature. although there have been pioneering studies on cognitive processes (e.g., patel & groen, 1986; boshuizen, 1989), most focused on perceptual processes (based on arnheim, 1969, see also the section on the detection metaphor). what lesgold indicated and empirically tested was thinking as an essential function in medical diagnosis. second, in this quote and elsewhere in his chapter, lesgold emphasized how vision and cognition correlated in the formation and evaluation of diagnostic decisions: experienced radiologists build mental representations that guide perception. the literature now gegenfurtner et van merriënboer 7 | f l r knows a variety of rhetorical functions built from verbal protocols to describe those mental representations, among them lesgold’s schemata, encapsulated scripts (boshuizen & schmidt, 2008), e-mops (kolodner, 1983), or sustains (love, medin, & gureckis, 2004). in the next paragraph, we describe one example on the correlation of vision and cognition in medical image diagnosis that is still informed by lesgold’s (one might tend to say: seminal) speculation of cognitive inference. morita and colleagues (2008) investigated how perceptual and conceptual processing interrelates in the diagnosis of computer tomograms (ct). shortly, they found that expert compared to novice ct readers verbalized more findings, more hypotheses, and more perceptual activities. importantly, experts verbalized many perceptual features during conceptual activities, and verbalized conceptual words during perceptual processing. put differently, this indicates that experts retrieved and used knowledge from memory based on information that they saw on the ct image, which iteratively stimulated looking at the image based on knowledge coded in memory (be it in the form of encapsulated scripts, e-mops, or schemata). from a methodological point of view, it would be tempting to criticize the neglect to use eye movement recordings; this would have allowed highly specific, quantifiable measures of perceptual activity. yet, verbal protocols can also be used as indicators for visual expertise (helle, 2017; van de wiel, 2017). usually, as prototypically shown in this example, protocols are collected for a duration of up to few hours and then are analyzed with a focus on cognitive mechanisms. from this perspective, the inference metaphor on visual expertise clearly emphasizes the cognitive parts of the interrelated process. morita and colleagues (2008) decided on individual ct readers as a unit of analysis. however, protocols can also be used at a group level (greeno, 2006; see simpson & gilhooly, 1997, for an example in cardiology) to indicate collaborative negotiations and intersubjective meaning-making. the inference metaphor in visual expertise research can answer, in two respects, the question regarding what develops in novice diagnosticians that moves them toward higher accuracy. first, knowledge develops; an extensive knowledge base is the foundation for expert performance and for rapid inference of coded memory to detected visible features. second, the perceptual-conceptual processing linkage develops. morita and colleagues have demonstrated that protocol measures are a valid tool for eliciting cognitive mechanisms underlying ct diagnoses that guide, and are guided by, perceptual detection. epistemologically, the inference metaphor appears to account for both signal detection and the alignment of knowledge from memory (inference) that is consistent to what is detected. nevertheless, methodologically, it lacks the precise and timesensitive measures such as eye-tracking gaze recordings or cortical oscillation eegs. this is because researchers rely on explicit, conscious think-aloud utterances from participants and these cannot account for their underlying implicit, non-conscious processing. hence, the sole use of protocols—be it at an individual or at an agglomerated group level—risks the resemblance of linguistic descriptions that play a rhetorical function in describing and illustrating phenomena; examples of these fancy rhetorical functions, which simply cannot be fully validated by protocol analysis alone, are provided above (sustains and the like). 2.4. visual expertise as practice metaphor finally, the last metaphor we identify as being frequently used in visual expertise research is the practice metaphor. sociocultural practice generates semantic structures of information that shape and are shaped by sequentially unfolding activity in relevant manners for a domain of scrutiny, such as laparoscopy or sports (gegenfurtner & szulewski, 2016; koschmann et al., 2007). as a starting point, we present the following quote from carsetti (2004, p. 307) that we found interesting enough to use to introduce our reflection on the practice metaphor: “a percept is something that lives and becomes, it possesses a biological complexity gegenfurtner et van merriënboer 8 | f l r which is not be explained simply in terms of the computations by a neural network classifying on the basis of very simple mechanisms”. this quote has two interesting elements. first, it emphasizes the lived nature of visual cognition, or what livingston (1986) referred to as the lived work of reading. we will present an empirical example in the next paragraph that elaborates on this notion. livingston, in a series of ethnographic field descriptions, highlighted the sociability of practices that constitute intersubjective thinking and acting. as such, the author provided a look that differed from looks “behind the skull” (garfinkel, 1967) or from “computations by a neural network” (carsetti, ibid). the second interesting element in this quote is that it seems to align neuroscientific work with labels such as “simply” and “very simple”. we lack authority and motivation to judge such a judgment about the simpleness of neuroscience as being itself simplistic. nevertheless, it illustrates the position of this author that something that focuses only on neural activation is unable to account for the full complexity of visual expertise (interestingly, compare the quote of sauseng & klimesch, 2008, starting the activation metaphor section, where the complexity of neural network communication is emphasized). certainly, it is a matter of definition what “complex” is or what shall be allowed—based on which methodological and epistemological considerations—to have “complexity”. making such (maybe tacit, maybe deliberate) assumptions explicit is one of the purposes of the framework in table 1. to further illustrate the methodological entailments of the practice metaphor, we describe an example elaborating on the lived work of mammography. slack, hartwood, procter, and rouncefield (2007) highlight how diagnosing a mammogram is reflexively contingent on artful practices, in which multiple readers interact and intersubjectively constitute breast biographies. central in their analyses are practices. goodwin (1994, 2000) indicated that seeing and interpreting what is seen are not exclusively cognitive processes located in the individual brain (cf. activation and detection metaphors); rather, seeing is a socially situated activity accomplished through the deployment of a range of historically matured discursive practices. these practices constitute visual expertise in goodwin’s terms, and they are negotiated around a common object of disciplined perception (ivarsson, 2017; lindwall & lymer, 2008; stevens & hall, 1998), in the study of slack and colleagues (2007): pictorial representations of the breast produced by an x-ray apparatus. slack et al. identified practices such as arranging mammograms, manipulating, annotating, gesturing, and pointing that contribute to the lived work of doing a radiologic diagnosis. these representative practices (greeno, 2006) unfold within an activity system, in many cases temporally over the course of minutes, but their sociogenesis stretches over the course of decades (such as the material resources used; i.e., pictures produced by x-ray technology). hence, analysis of visual expertise using the practice metaphor adopts a different analytic time span than does, for example, analysis using the activation metaphor; and it adopts a sociotechnical system as the unit of analysis that explicitly accounts for the mediating role of technology (burri & dumit, 2008; gegenfurtner, 2013; säljö, 2012; siewiorek & gegenfurtner, 2010). it is essentially the focus on embodied talk-in-interaction—talk between people (knogler et al., 2013) and between humans and non-human objects (gibson, 1979; suchman, 2007)—that makes the practice metaphor a useful tool to analyze visual expertise and to generate a practicebased theorizing about the work of experts. 3. structure of the special issue devoted to research on visual expertise, this special issue consists of 12 contributions in three sections: (a) this introduction sets the stage for the unfolding reflection, (b) 9 articles reflect on diverse methodological approaches, and (c) 2 commentaries by boucheix (2017) as well as jarodzka and boshuizen (2017) close the special issue. these two concluding commentaries offer a detailed discussion of each of the nine articles and gegenfurtner et van merriënboer 9 | f l r synthesize their lessons in novel ways. hence, an in-depth summary of the nine articles is redundant here. however, we would like to link the methodological approaches outlined and discussed in the nine articles with the comparative metaphorical mapping presented in table 1. first, the activation metaphor corresponds with cognitive-neurosciences. prominent methods within cognitive-neurosciences are functional magnetic resonance imaging and electroencephalography. gegenfurtner et al. (2017b) reflect on these methods, their potential advantages, and their risks for studying visual perceptual expertise in medicine. second, the detection metaphor corresponds with roc analysis and eye tracking methodology. krupinski (2017) discusses roc analysis. several articles cover eye tracking. fox and faulkner-jones (2017) offer a general discussion. szulewski et al. (2017) focus on pupillometry. litchfield and donovan (2017) reflect on the benefits of the flash-preview moving window paradigm. and helle (2017) explores how eye tracking and verbal report data can be combined. this is already a bridge toward the third metaphor: inference. most prominently, the inference metaphor corresponds with verbal report data and interviews. the latter is introduced and discussed by van de wiel (2017). finally, the practice metaphor corresponds with ethnomethodology. ivarsson (2017) offers an ethnomethodological reflection of visual expertise as embodied practice. as a contribution that overarches the four metaphors, williams and colleagues (2017) describe how the expert performance approach can inform research designs intending to study visual expertise. different methods lead to different answers based on different indicators (damşa et al., in press). it is a reflection on these indicators, and more generally on the methodological entailments behind seemingly different metaphors, that can help raise awareness of each metaphor’s (epistemological and pragmatic) potentiality and contingency, and that can thus advance our research practice in medical education and the learning sciences. sfard (1998) noted that a combination of learning metaphors yields to more robust findings than does a non-combination. the detection metaphor seems to be currently dominant in the field of visual expertise research (which is also reflected in the number of articles in this special issue). further developments in the field could profit from combining this particular metaphor with alternative metaphors. the special issue is thus a first step towards reaching this ambitious goal (see also the commentaries by boucheix, 2017, and jarodzka and boshuizen, 2017, as well as gegenfurtner et al., 2013; lehtinen, 2012; säljö, 2009). probably, research aimed at method triangulation will advance the field more than discussions about the superiority of a particular method (carsetti, 2004; sauseng & klimesch, 2008). this is in line with research in times that have been labeled the decade of synergy (bransford et al., 2006). yet, combining methods is neither trivial nor simple. it is an essential task for future research on visual expertise to explore the synergies between metaphors to avoid the dangers associated with choosing just one metaphor. keypoints the special issue discusses the methodological pluralism of research on visual expertise. this introduction offers a comparative metaphorical mapping of visual expertise research. the mapping includes four metaphors: activation, detection, inference, and practice. combining metaphors will help the further development of visual expertise research. gegenfurtner et van merriënboer 10 | f l r references alexander, p. a., schallert, d. l., & reynolds, r. e. (2009). what is learning anyway? a topographical perspective reconsidered. educational psychologist, 44, 176-192. doi:10.1080/00461520903029006 ansari, d., & coch, d. (2006) bridge over troubled waters: education and cognitive neuroscience. trends in cognitive sciences, 10, 146-151. doi:10.1016/j.tics.2006.02.007 arnheim, r. (1969). visual thinking. berkeley: university of california press. bilalić, m. (2017). the neuroscience of expertise. cambridge: cambridge university press. boshuizen, h. p. a. (1989). de ontwikkeling van medische expertise: een cognitief-psychologische benadering. maastricht: maastricht university. boshuizen, h. p. a., & schmidt, h. g. (2008). the development of clinical reasoning expertise; implications for teaching. in j. higgs, m. jones, s. loftus, & n. christensen (eds.), clinical reasoning in the health professions (pp. 113-121). oxford: butterworth-heinemann. boucheix, j.-m. (2017). the interplay between methodologies, tasks and visualisation formats in the study of visual expertise. frontline learning research. boucheix, j.-m., bonnetain, e., avena, c., & freysz, m. (2010). benefits of learning technologies in medical training, from full-scale simulators to virtual reality and multimedia presentations. in j. p. didier, e. bigand, & a. vinter (eds.), rethinking physical and rehabilitation medicine (pp. 171-191). new york: springer. boucheix, j.-m., & lowe, r. (in press). generative processing of animated partial depictions fosters fish identification skills: eye tracking evidence. le travail humain. bransford, j., stevens, r., schwartz, d., meltzoff, a., pea, r., roschelle, j., et al. (2006). learning theories and education. toward a decade of synergy. in p. a. alexander & p. h. winne (eds.), handbook of educational psychology (pp. 209-244). mahwah, nj: erlbaum. burri, r. v., & dumit, j. (2008). social studies of scientific imaging and visualization. in e. j. hackett, o. amsterdamska, m. lynch, & j. wajcman (eds.), handbook of science and technology studies (pp. 297317). cambridge, ma: mit press. carsetti, a. (2004). the embodied meaning. self-organization and symbolic dynamics in visual cognition. in a. carsetti (ed.), seeing, thinking, and knowing. meaning and self-organization in visual cognition and thought (pp. 307-327). dordrecht: kluwer. damsa, c. i., froehlich, d. e., & gegenfurtner, a. (in press). reflections on empirical and methodological accounts of agency at work. in m. goller & s. paloniemi (eds.), agency at work: an agentic perspective on professional learning and development. new york: springer. de jong, t., van gog, t., jenks, k., manlove, s., van hell, j., jolles, j., et al. (2009). explorations in learning and the brain. on the potential of cognitive neuroscience for educational science. berlin: springer. ericsson, k. a. (2004). deliberate practice and the acquisition and maintenance of expert performance in medicine and related domains. academic medicine, 79, s70-s81. doi:10.1097/00001888-20041000100022 fox, s. e., & faulkner-jones, b. e. (2017). eye-tracking in the study of visual expertise: methodology and approaches in medicine. frontline learning research. doi: doi:10.14786/flr.v5i3.258 garfinkel, h. (1967). studies in ethnomethodology. englewood cliffs, nj: prentice hall. gegenfurtner, a. (2013). transitions of expertise. in j. seifried & e. wuttke (eds.), transitions in vocational education (pp. 305-319). opladen: budrich. gegenfurtner, a., kok, e., van geel, k., de bruin, a., jarodzka, h., szulewski, a., & van merriënboer, j. j. g. (2017a). the challenges of studying visual expertise in medical image diagnosis. medical education, 51, 97-104. doi:10.1111/medu.13205 gegenfurtner, a., kok, e. m., van geel, k., de bruin, a. b. h., & sorger, b. (2017b). neural correlates of visual perceptual expertise: evidence from cognitive neuroscience using functional neuroimaging. frontline learning research. doi:10.14786/flr.v5i3.259 gegenfurtner et van merriënboer 11 | f l r gegenfurtner, a., lehtinen, e., jarodzka, h., & säljö, r. (2017c). effects of eye movement modeling examples on adaptive expertise in medical image diagnosis. computers & education, 113, 212-225. doi:10.1016/j.compedu.2017.06.001 gegenfurtner, a., nivala, m., säljö, r., & lehtinen, e. (2009). capturing individual and institutional change: exploring horizontal versus vertical transitions in technology-rich environments. in u. cress, v. dimitrova, & m. specht (eds.), learning in the synergy of multiple disciplines. lecture notes in computer science (pp. 676-681). berlin: springer. doi:10.1007/978-3-642-04636-0_67 gegenfurtner, a., & seppänen m. (2013). transfer of expertise: an eye-tracking and think-aloud study using dynamic medical visualizations. computers & education, 63, 393-403. doi:10.1016/j.compedu.2012.12.021 gegenfurtner, a., siewiorek, a., lehtinen, e., & säljö, r. (2013). assessing the quality of expertise differences in the comprehension of medical visualizations. vocations and learning, 6, 37-54. doi:10.1007/s12186-012-9088-7 gegenfurtner, a., & szulewski, a. (2016). visual expertise and the quiet eye in sports comment on vickers. current issues in sport science, 1, 108. doi:10.15203/ciss_2016.108 gegenfurtner, a., vauras, m., gruber, h., & festner, d. (2010). motivation to transfer revisited. in k. gomez, l. lyons, & j. radinsky (eds.), learning in the disciplines: icls2010 proceedings (vol. 1, pp. 452459). chicago, il: international society of the learning sciences. gibson, j. j. (1979). the ecological approach to visual perception. boston, ma: houghton. goodwin, c. (1994). professional vision. american anthropologist, 96, 606-633. doi:10.1525/aa.1994.96.3.02a00100 goodwin, c. (2000). practices of seeing: visual analysis: an ethnomethodological approach. in t. van leeuwen & c. jewitt (eds.), handbook of visual analysis (pp. 157-182). london: sage. greeno, j. a. (2006). learning in activity. in r. sawyer (ed.), the cambridge handbook of the learning sciences (pp. 79-96). cambridge, ma: cambridge university press. gruber, h., & degner, s. (2016). expertise und kompetenz [expertise and competence]. in m. dick, w. marotzki, & h. mieg (eds.), handbuch professionsentwicklung (pp. 173-180). bad heilbrunn: klinkhardt. hakkarainen, k., palonen, t., paavola, s., & lehtinen, e. (2004). communities of networked expertise: professional and educational perspectives. amsterdam: elsevier. haller, s., & radue, e. w. (2005). what is different about a radiologist’s brain? radiology, 236, 983-989. doi:10.1148/radiol.2363041370 helle, l. (2017). prospects and pitfalls in combining eye-tracking data and verbal reports. frontline learning research. doi:10.14786/flr.v5i3.254 helle, l., nivala, m., kronqvist, p., gegenfurtner, a., björk, p., & säljö, r. (2011). traditional microscopy instruction versus process-oriented virtual microscopy instruction: a naturalistic experiment with control group. diagnostic pathology, 6, s81-s89. doi:10.1186/1746-1596-6-s1-s8 ivarsson, j. (2017). visual expertise as embodied practice. frontline learning research. doi:10.14786/flr.v5i3.253 jarodzka, h., & boshuizen, h. p. a. (2017). unboxing the black box of visual expertise in medicine. frontline learning research. jarodzka, h., scheiter, k., gerjets, p., & van gog, t. (2010). in the eyes of the beholder: how experts and novices interpret dynamic stimuli. learning and instruction, 20, 146-154. doi:10.1016/j.learninstruc.2009.02.019 johansson, e., lindwall, o., & rystedt, h. (2017). experiences, appearances, and interprofessional training: the instructional use of video in post-simulation defbriefings. international journal of computersupported collaborative learning, 12, 91-112. doi:10.1007/s11412-017-9252-z knogler, m., gegenfurtner, a., & quesada pallarès, c. (2013). social design in digital simulations: effects of single versus multi-player simulations on efficacy beliefs and transfer. in n. rummel, m. kapur, m. nathan, & s. puntambekar (eds.), to see the world and a grain of sand: learning across levels of space, time, and scale (vol. 2, pp. 293-294). madison, wi: international society of the learning sciences. gegenfurtner et van merriënboer 12 | f l r kok, e. m., van geel, k., van merriënboer, j. j. g., & robben, s. g. f. (2017). what we do and do not know about teaching medical image interpretation. frontiers in psychology, 8, 309. doi: 10.3389/fpsyg.2017.00309 kolodner, j. l. (1983). towards an understanding of the role of experience in the evolution from novice to expert. international journal of man-machine systems, 19, 497-518. doi:s0020-7373(83)80068-6 koschmann, t., lebaron, c., goodwin, c., zemel, a., & dunnington, g. (2007). formulating the triangle of doom. gesture, 7, 97-118. doi:10.1075/gest.7.1.06kos koschmann, t., & zemel, a. (2009). optical pulsars and black arrows: discoveries as occasioned productions. journal of the learning sciences, 18, 200-246. doi:10.1080/10508400902797966 krupinski, e. a. (2017). receiver operating characteristics. frontline learning research. doi: doi:10.14786/flr.v5i3.250 krupinski, e. a., graham, a., r., & weinstein, r. s. (2013). characterizing the development of visual search expertise in pathology residents viewing whole slide images. human pathology, 44, 357-364. doi:10.1016/j.humpath.2012.05.024 kundel, h. l., nodine, c. f., & carmody, d. (1978). visual scanning, pattern recognition, and decisionmaking in pulmonary nodule detection. investigative radiology, 13, 175-181. kundel, h. l., nodine, c. f., conant, e. f., weinstein, s. p. (2007). holistic component of image perception in mammogram interpretation: gaze-tracking study. radiology, 242, 396-402. doi:10.1148/radiol.2422051997 lehtinen, e. (2012). learning of complex competences: on the need to coordinate multiple theoretical perspectives. in a. koskensalo, j. smeds, r. de cillia, & á. huguet (eds.), language: competencies change contact (pp. 13-27). berlin: lit. lesgold, a., rubinson, h., feltovich, p., glaser, r., klopfer, d., & wang, y. (1988). expertise in a complex skill: diagnosing x-ray pictures. in m. t. h. chi, r. glaser, & m. j. farr (eds.), the nature of expertise (pp. 311-342). hillsdale, nj: erlbaum. lindwall, o., & lymer, g. (2008). the dark matter of lab work: illuminating the negotiation of disciplined perception in mechanics. journal of the learning sciences, 17, 180-224. doi:10.1080/10508400801986082 litchfield, d., & donovan, t. (2017). the flash-preview moving window paradigm: unpacking visual expertise one glimpse at a time. frontline learning research. doi:10.14786/flr.v5i3.269 livingston, e. (1986). the ethnomethodological foundations of mathematics. london: kegan paul. love, b. c., medin, d. l., & gureckis, t. m. (2004). sustain: a network model of category learning. psychological review, 111, 309-332. doi:10.1037/0033-295x.111.2.309 meltzoff, a. n., kuhl, p. k., movellan, j., & sejnowski, t. j. (2009). foundations for a new science of learning. science, 325, 284-288. doi:10.1126/science.1175626 morita, j., miwa, k., kitasaka, t., mori, k., suenaga, y., iwano, s., et al. (2008). interactions of perceptual and conceptual processing: expertise in medical image diagnosing. international journal of humancomputer studies, 66, 370-390. doi:10.1016/j.ijhcs.2007.11.004 paavola, s., lipponen, l., & hakkarainen, k. (2002). epistemological foundations for cscl: a comparison of three models of innovative knowledge communities. in g. stahl (ed.), computer support for collaborative learning: foundations for a cscl community. cscl 2002 proceedings (pp. 24-32). hillsdale, nj: erlbaum. palonen, t., boshuizen, h. p. a, & lehtinen, e. (2014). how expertise is created in emerging professional fields. in s. billett, t. halttunen, & m. koivisto (eds.), promoting, assessing, recognizing and certifying lifelong learning: international perspectives and practices (pp. 131-150). new york: springer. patel, v. l., arocha, j. f., & zhang, j. (2005). thinking and reasoning in medicine. in k. j. holyoak (ed.), the cambridge handbook of thinking and reasoning (pp. 727-750). cambridge, ma: cambridge university press. patel, v. l., & groen, g. j. (1986). knowledge-based solution strategies in medical reasoning. cognitive science, 10, 91-116. doi:10.1207/s15516709cog1001_4 gegenfurtner et van merriënboer 13 | f l r quesada-pallarès, c., & gegenfurtner, a. (2015). toward a unified model of motivation for training transfer: a phase perspective. zeitschrift für erziehungswissenschaft, 18 (suppl. 1), 107-121. doi:10.1007/s11618014-0604-4 säljö, r. (2009). learning, theories of learning, and units of analysis in research. educational psychologist, 44, 202-208. doi:10.1080/00461520903029030 säljö, r. (2012). literacy, digital literacy and epistemic practices: the co-evolution of hybrid minds and external memory systems. nordic journal of digital literacy, 7, 5-19. sangin, m., molinari, g., nüssli, m., & dillenbourg, p. (2008). how learners use awareness cues about their peer knowledge? insights from synchronized eye-tracking data. international conference of the learning sciences proceedings. sauseng, p., & klimesch, w. (2008). what does phase information of oscillatory brain activity tell us about cognitive processes? neuroscience and biobehavioral reviews, 32, 1001-1013. doi:10.1016/j.neubiorev.2008.03.014 seppänen, m., & gegenfurtner, a. (2012). seeing through a teacher’s eyes improves students’ imaging interpretation. medical education, 46, 1113-1114. doi:10.1111/medu.12041 sfard, a. (1998). on two metaphors for learning and the dangers of choosing just one. educational researcher, 27, 4-13. doi:10.3102/0013189x027002004 siewiorek, a., & gegenfurtner, a. (2010). leading to win: the influence of leadership style on team performance during a computer game training. in k. gomez, l. lyons, & j. radinsky (eds.), learning in the disciplines: icls2010 proceedings (vol. 1, pp. 524-531). chicago, il: international society of the learning sciences. simpson, s. a., & gilhooly, k. j. (1997). diagnostic thinking processes: evidence from a constructive interaction study of electrocardiogram (ecg) interpretation. applied cognitive psychology, 11, 543-554. doi:10.1002/(sici)1099-0720(199712)11:6<543::aid-acp486>3.0.co;2-c slack, r., hartswood, m., procter, r., & rouncefield, m. (2007). cultures of reading: on professional vision and the lived work of mammography. in s. hester & d. francis (eds.), orders of ordinary action. respecifying sociological knowledge (pp. 175-193). aldershot: ashgate. smith, p. l., & ratcliff, r. (2009). an integrative theory of attention and decision-making in visual signal detection. psychological review, 116, 283-317. doi:10.1037/a0015156 stevens, r., & hall, r. (1998). disciplined perception: learning to see in technoscience. in m. lampert & m. l. bunk (eds.), talking mathematics in school: studies of teaching and learning (pp. 107-149). new york: cambridge university press. suchman, l. (2007). human-machine reconfigurations. plans and situated actions (2nd ed.) cambridge, ma: cambridge university press. szulewski, a., gegenfurtner, a., howes, d., sivilotti, m., & van merriënboer, j. j. g. (in press). measuring physician cognitive load: validity evidence for a physiologic and a psychometric tool. advances in health sciences education. doi:10.1007/s10459-016-9725-2 szulewski, a., kelton, d., & howes, d. (2017). pupillometry as a tool to study expertise in medicine. frontline learning research. doi:10.14786/flr.v5i3.256 van de wiel, m. (2017). examining expertise using interviews and verbal protocols. frontline learning research. doi:10.14786/flr.v5i3.257 williams, m. a., fawver, b., & hodges, n. j. (2017). using the ‘expert performance approach’ as a framework for examining and enhancing skill learning: improving understanding of expert learning. frontline learning research. doi:10.14786/flr.v5i3.267 codepen 3heinrichs publication frontline learning research special issue vol.8 no.5 (2020) 24 46 issn 2295-3159 an action-theoretical approach to the happy victimizer pattern – exploring the role of moral disengagement strategies on the way to action karin heinrichsa, tobias kärnerb & hannes reinkec auniversity of education upper austria, austria buniversity of hohenheim, stuttgart, germany cuniversity of bamberg, germany article received 11 june 2018 / revised 3 may 2020 / accepted 7 may / available online 1 july abstract research in moral education demonstrates the pattern referred to as happy victimising (hv) does not emerge only among children. adults also transgress moral rules and might feel good doing so; however, research reveals the hv pattern emergence is context specific. in contrast to findings among young children in whom the hv pattern was interpreted as a lack of motivation and thus a developmental stage, it is an open question as to what happy victimising in adulthood means and how such patterns affect intentions as an important step towards action. this paper offers an action-theoretical approach, allowing for reconstruction of the process of intention formation, as well as a systematic discussion of results from two separate lines of research: (1) research on patterns of moral decision-making, such as the hv, and (2) research on moral disengagement. additionally, a survey study provides insights into what intentions, emotion attributions, and moral disengagement strategies adults display in situations of low moral intensity, and whether they indicate consistent or contradictory patterns across situations. results indicate intra-personal consistency regarding patterns of moral decision-making, but also show there are participants who vary these patterns across situations. moral disengagement strategies were shown to have context-specific use, at least in regard to their subcategories. regarding education, this study encourages not only a focus on strengthening the moral self or autonomous moral judgement but also on paying attention to actions and person-situation interactions. this might be useful to implement environments that support reduced application of moral disengagement strategies. keywords: moral transgression, happy victimizer, moral acting, moral disengagement info corresponding author email: karin.heinrichs@ph-ooe.at doi: https://doi.org/10.14786/flr.v8i5.386 1. introduction the frequency of economic scandals, such as diesel-gate, tax evasion, corruption, fraud, or fake-shops for breathing masks during the covid-19 pandemic indicates that even people who seem to be friendly and empathic at first glance may transgress moral rules almost as frequently as others, depending on the situation, context, or one’s role. at least occasionally, people do not follow conventional rules, moral standards, or principles and, simultaneously, ignore others’ perspectives in favour of fulfilling their own or their companies’ needs. while this may lead to the assumption that moral transgression and self-centeredness are a basic human phenomenona that emerge in adulthood, research also shows that moral education can be successful in developing socio-moral competencies (e.g., lind, 2019; weinberger & frewein, 2019) and creating a moral atmosphere. further, fostering empathy and social perspective-taking helps to increase prosocial behaviour (bandura, 2016; eisenberg, fabes, & spinrad, 2007; malti et al., 2016). from an educational perspective, it seems important to keep searching for effective ways to develop the competencies adults need to effectively deal with morally relevant situations in their everyday (working) lives. in routine as well as odd situations, people have to balance their own interests with those of others. thus, to discuss aims of moral education across the lifespan, results of empirical research should be considered that contribute to explaining how people act, what personal and situational factors determine whether someone acts in line with or contradictory to moral standards, and how such competencies can be developed up to adulthood and beyond. research on moral psychology shows that agents simultaneously know about moral rules and attribute positive emotions in cases of transgression. this pattern of ethical decision-making, here called ‘happy victimising’, was initially detected among children approximately four years old; however, recent research has shown that this pattern also emerges among adolescents and adults (e.g., heinrichs, gutzwiller-helfenfinger, latzko, minnameier, & döring, this issue; heinrichs, minnameier, latzko & gutzwiller-helfenfinger, 2015; nunner-winkler, 2007; 2013; minnameier, heinrichs, & kirschbaum, 2016; minnameier & schmidt, 2013). furthermore, there is evidence of different manifestations of ethical decision-making and emotion attributions, such as the patterns of ‘happy victimizing’ (hv; transgressing moral rules, attributing positive emotions), ‘unhappy victimizing’ (uv; transgressing moral rules, attributing negative emotions), ‘happy moralizing’ (hm; obeying moral rules, attributing positive emotions), and ‘unhappy moralizing’ (um; obeying moral rules, attributing negative emotions) these patterns are at least applied in economically relevant situations, and emerge to varying extents, depending on the measurement methods used (gutzwiller-helfenfinger & latzko, this issue). they also vary intra-personally across situations, and can influence individual actions (döring, 2013; gasser, gutzwiller-helfenfinger, latzko, & malti, 2013). thus, these patterns of ethical decision-making might contribute to explaining adults’ deviant behaviours within different social, private, and work-related contexts. however, there is a lack of empirical evidence and theoretical foundations on how internal patterns of decision-making, moral judgments, and emotion attributions can be modelled to determine the processes of action in morally (and economically) relevant situations. this is important to study across different developmental stages; however, this article focuses on adults’ patterns of moral decision-making and emotion attributions as indicators of the valence of intentions, and thus as predictors for actions. to bridge the gap between judgment and action, we applied the process model of judging and acting (heinrichs, 2005). this model provides a theoretical framework to gain deeper insight into relevant situational and individual determinants of action processes. this model further offers a detailed reconstruction of the action formation process, from interpreting a perceived situation to implementing a behaviour. particularly, it provides ideas on relevant steps in the first phase of acting, which concludes with intention formation. referring to this process model, one obstacle to forming a high-valence intention is a lack of self-commitment to one’s preferred way of behaving. this lack may often become apparent in situations with conflicting values or goals, or when an individual perceives ambivalence in cognitive or emotional states, and often appears in uv and um patterns. to overcome this lack of commitment and form a high-valence intention, the process model posits that people use cognitive control strategies (heinrichs, 2005). to specify what cognitive control strategies might be helpful in morally relevant situations where the agent has to choose between transgressing against or obeying a moral rule, it is necessary to refer to further research. therefore, we refer to bandura`s concept of ‘moral disengagement’ strategies (mds; bandura, 1990, 2016). bandura and colleagues suggest (volitional) control strategies, and indicate that these strategies deactivate self-sanctions that would normally support moral action. individuals using mds can thus make a choice other than the ‘moral’ course of action, as these self-regulation strategies enable agents to reach a state of emotional well-being while causing negative (severe) consequences for others (osofsky, bandura, & zimbardo, 2005). in specifying the role of mds in the process of forming an intention, however, it remains questionable whether and to what extent people really apply mds in certain situations; when they decide for or against a moral transgression; and whether they feel committed to choosing ‘victimising’ or ‘moral behaviours’. focusing on mds in particular situations seems important, as moral reasoning, moral emotions, and patterns of moral decision-making vary intra-personally across situations, and can therefore be considered results of person-situation interactions. therefore, it might be fruitful to study how people use mds within the context of moral transgressions. bandura and colleagues mainly studied mds as individual tendencies across situations; thus, they looked for intra-personal consistency. contrastingly, we focused on the application of mds in particular situations during the action process, particularly during the first step towards acting: the sub-process of forming an intention. we aimed to provide deeper insights into whether individuals who intend to make moral transgressions (uv, hv) also apply mds. if they use mds in terms of (moral) reasoning for their preferences, in line with the process model, it could be assumed that mds might have functioned during the process of forming intentions towards a preference of moral transgression, increasing the level of commitment, and the probability of acting in line with one’s intention. thus, this paper provides theoretical ideas and the first empirical data on the hv pattern in adulthood, exploring mds use in morally relevant situations. theoretically, the presented study refers to an action-based approach (heinrichs, 2005) that allows for the presentation of theoretical ideas on how patterns of moral decision-making and emotion attributions (hv, uv, hm, um), as well as mds, may affect intentions in morally relevant situations, and specifically in situations that provoke decisions to follow or break a moral rule. the results of this questionnaire study among students can provide insight in differences in intentions, attributed emotions, and frequency, as well as qualities and intra-personal differences, of mds use across morally relevant situations in a work-related context. the paper is structured as follows. section 2.1 provides basic information on the state of research in terms of empirical findings on patterns of moral decision-making among adults. in this context, a rationale is provided for why this study focuses on situations of low moral intensity that are assumed to provoke victimisation at a higher rate than dilemmas and that are omnipresent in everyday (working) life. section 2.2 explicates basic assumptions of an action-theoretical approach to reconstruct patterns of moral decision-making as potential intentions to act in morally relevant situations. we therefore chose situations in consumer and business contexts, as adolescents and adults are familiar with these situations in everyday or working life. furthermore, such situations might have the potential to elucidate inner conflicts related to transgressing against moral rules in favour of economic or personal interests. section 2.3 explains theoretical foundations and empirical findings related to mds and posits that mds may represent cognitive strategies that can be linked to intentions to obey or transgress against a moral rule. based on this theoretical foundation, a survey study was conducted to explore whether and to what extent students apply patterns of moral decision-making and emotion attributions across situations, as well as whether and to what extent students use mds when choosing whether to victimise others (sections 3 and 4). finally, implications for further research on decision-making, behaviour, and disengagement in morally relevant situations and implications for moral education are discussed (section 5). 2. theory 2.1 happy victimising in adulthood the developmental psychologists nunner-winkler and sodian (1988) coined the term ‘happy victimizer phenomenon’, and explained it as a lack of moral motivation at an early stage of moral development. as the rate of people displaying this ‘happy victimizer phenomenon’ decreases in later age groups (from eight years onwards), it was assumed that this pattern is caused by the absence of a link between cognition and emotion, and therefore might be overcome during later stages of moral development (krettenauer, malti, & sokol, 2008; nunner-winkler, 1993). however, further studies revealed that such patterns also emerged to a considerable extent in adolescence (döring, 2013; heinrichs et al., this issue) and even in adulthood (heinrichs et al., 2015; nunner-winkler, 2007; 2013; minnameier, heinrichs, & kirschbaum, 2016; minnameier & schmidt, 2013). however, empirical studies have revealed that these patterns of moral decision-making in adulthood do not characterise a person-specific method of moral judgment, applied consistently across situations, with hardly any exceptions, as was assumed in the ‘happy victimizer phenomenon’ of early childhood. moreover, in adulthood, these patterns are described as varying intra-personally across situations, to an extent that had not been expected based on previous theoretical assumptions. simultaneously, recent findings indicate a significant small or medium effect that still points to personal preferences towards patterns of moral decision-making (heinrichs, et al., this issue; malti & krettenauer, 2013). thus, patterns of moral decision-making seem to result from person-situation interaction, but are more affected by situational determinants than developmental psychology had assumed. moreover, empirical research in moral and developmental psychology has previously confirmed that the proportion of moral decisions reflecting the hv pattern varies depending on situational cues, particularly on the degree of the moral conflict, or whether people are encouraged to make judgments using a selfor others’ perspective (keller, lourenço, malti, & saalbach, 2003; nunner-winkler, 2013; malti & krettenauer, 2013). however, these situational variations were mostly discussed as being dependent on measurement methods (heinrichs et al., 2015; this issue; gutzwiller-helfenfinger & latzko, this issue; nunner-winkler, 2013). following empirical results on intra-personal variations in moral decision-making, in this paper, we do not use the term ‘happy victimizer phenomenon’, as previously described as emerging in earlier stages of childhood development. moreover, we differentiate ‘patterns’ of moral decision-making and emotion attributions (see above; heinrichs et al., this issue). it has been assumed that hv—as well as uv, hm, and um—displays intra-personal variations in moral decision-making and managing moral emotions across situations. however, until now, there has been no satisfying empirical evidence supporting this, but rather a need for research on the personal and situational determinants that trigger the use or intra-personal change in these patterns. furthermore, there is also a lack of theoretical approaches to explain hv in adulthood. however, there is evidence that these patterns are important insofar as they are linked to individual actions. empirical findings confirm that hv is a relevant pattern in the context of bullying (gasser et al., 2013), and characterises bullies or bully victims. moreover, the hv pattern is related to deviant adolescent behaviour (döring, 2013). research on counterproductive behaviour in organisational contexts indicates the relevance of individual moral values, judgments, and moral sensibility (moore, detert, treviño, baker, & mayer, 2012). thus, patterns of moral decision-making, such as hv, uv, hm, and um, might further contribute to explaining deviant behaviours among adults, within different social, private, and work-related contexts. therefore, they may also be relevant from the perspectives of moral education, vocational education, human resource development, and organisational behaviour. 2.2 an action-based approach to moral decision-making this paper primarily contributes to theoretical progress in explaining determinants of the action process in morally relevant situations and, in particular, determinants of moral intentions. therefore, two lines of research are linked to each other: research on patterns of moral decision-making, like the hv pattern, and research on mds. the process model of acting (heinrichs, 2005) functions as a theoretical framework which can be used to reconstruct and specify the action processes, from constituting a situation to forming an intention, implementation, conduct, and evaluation. this model was developed as a theoretical framework integrating esser’s (1996) social psychological model of ‘definition of the situation’ and heckhausen’s rubikon model (gollwitzer, 1996; heckhausen, gollwitzer, & weinert, 1987), and is based on a set of assumptions. it allows for analysing the interaction between personal and situational determinants on the way from perceiving selected situational cues to forming an intention and behaviour (see figure 1). thus, in line with approaches to moral judgment and action in the post-kohlbergian tradition, it does not focus on the development of personal determinants, but points to applying patterns of decision-making, reasoning, or acting to morally relevant situations and contexts (krebs & denton, 2005; lapsley & narvaez, 2005; for a summary, see also heinrichs, 2010). figure 1. the process model of acting: from constituting a situation to forming an intention (heinrichs, 2005) the model’s basic assumptions are as follows (heinrichs, 2005): an action is determined by an initial situation constituted by the individual. if he or she experiences a difference between is and ought in respect to moral issues, he or she perceives a morally relevant ‘problem’1 . experiencing such a problem motivates further judgments and actions, and marks the important starting point of the action process. insofar as actions are determined by a situation subjectively constituted at the beginning of the process, the constituted situation is determined by personal and situational conditions. the action process is reconstructed and theoretically divided into four phases: (1) forming an intention, (2) planning, (3) implementation/conduct, and (4) evaluation. the central output of the first phase is an intention. the agent is willing and feels committed to realising and achieving an aim. this means he or she has made a decision towards aims or actions. sometimes, the individual already has the aim linked to a concrete action plan; otherwise, the concrete action will have to be specified during the implementation phase. this intention might be formed if the agent has perceived a problem, if he or she feels confident he or she can find a solution (at least in the future), if he or she is motivated to contribute to solving the problem, and, moreover, if a status of self-commitment was developed and expresses the volitional power to overcome barriers during the implementation phase (see figure 1). this means that the intention, as a measurable product of inner processes, has to be connected to the status of commitment and to be of notable valence. regarding patterns of moral decision-making (hv, uv, um, hm), forming an intention is the first of four phases in the action process. if a person has experienced a morally relevant problem (defined as a subjectively perceived gap between is and ought), then this individual might perceive a tension between obeying a moral rule and fulfilling his or her personal needs, and between different aims or actions. he or she might experience an ambivalence between cognitive, emotional, or motivational states, particularly if he or she attributes negative emotions towards his or her preferred way of behaving (uv or um). if the person perceives an inner conflict, he or she might not yet feel committed to one aim or action. to form an intention and progress into action, he or she must then decide among alternatives. referring to action theory, cognitive control strategies, in the sense of volitional strategies, play a major role in increasing commitment (heckhausen, 1987; heinrichs, 2005; 2013; sokolowski, 1993; 1996). volitional strategies support dealing with inner conflict or ambivalence in such a way that, finally, a state of commitment to one out of several possible actions could be achieved. in relation to patterns of moral decision-making, this means that the person feels committed to obeying or transgressing against a moral rule. it is assumed that such a state of self-commitment can only be reached if the individual anticipates being able to cope with upcoming negative consequences or conflicts. as research shows that patterns such as hv emerge depending on the quality of a moral conflict, it can be assumed that there is a need to apply cognitive control strategies across situations that might vary according to the extent of moral intensity. a situation’s moral intensity depends on how an individual perceives elements of reality (see figure 1; for more details on the concept of moral intensity, see jones, 1991). facing some ‘situational’ prompts, individuals may experience (intense) internal conflicts and high moral intensity. this could be expected, for example, in moral dilemmas that studies in moral psychology—especially in the kohlbergian tradition—focused on, and that are supposed to seldomly emerge in everyday life. contrastingly, situations of low moral intensity are assumed to be more frequent. individuals perceive low moral intensity if they do not recognise an internal conflict, or if they quite easily make a decision towards one preferred action. many people may perceive low moral intensity, for example, when transgressing against a moral rule only has (mild) negative consequences such as increased economic costs or treating others slightly unfairly, rather than causing physical or psychological harm. in line with the process model of acting, we may assume that if people experience an inner conflict, forming an intention might be a matter of reflective (vs. intuitive) data processing. in moral conflicts and situations of high moral intensity, an individual might perceive a need for self-commitment and apply cognitive control strategies. conversely, in situations of low to moderate moral intensity, an individual might experience a smaller difference between is and ought. people may display tendencies towards one action or another more easily and intuitively, based on automatic (cognitive and affective) processes of moral motivation (haidt & craig, 2008; heinrichs, 2005; 2013; rothmund & baumert, 2014). however, even in cases of less conscious modes of data processing in which an individual has a clear preference for an action, for example in situations of low moral intensity or situations the individual has faced before, cognitive control strategies are assumed to play an important role in building commitment. taking situations of low intensity into account in research on moral actions may be important from an educational perspective. one possible scenario is that people who choose victimising in situations of low moral intensity may become used to it or even continue to transgress morally, even in situations of moderate or high moral intensity. contrastingly, people who accept victimising in situations of low moral intensity might switch to obeying moral rules in situations of higher moral intensity. however, studying the development of hv and its determinants is an important question for further research and not the aim of this paper. therefore, these considerations encouraged us to study patterns of moral decision-making in situations of low moral intensity as a first step and a matter of moral sensibility (thoma & bebeau, 2013; tirri, 1999). thus, the process model of acting theoretically allows one to specify the role of cognitive control strategies as part of forming an intention. it is also assumed that intentions may differ in valence, that is, in strength of commitment or in their volitional power. further, the volitional power of one’s intentions impacts how barriers need to be overcome during the implementation phase (heckhausen et al., 1987; heinrichs, 2005). however, the process model of acting is first limited to theoretically reconstructing a sequence of input and output of inner sub-processes. admittedly, an individual cannot be conscious of these psychological processes; instead, it is assumed that the individual is at least potentially aware of the content or results of subprocesses, such as intentions, emotion attributions, or cognitive control strategies, applied in a particular situation (nisbett-wilson-thesis; see neuweg, 1999). thus, identifying relevant content or output of subprocesses, as mentioned above, could serve to systematically develop hypotheses concerning the links between them as results of subprocesses of actions in certain situations, such as the thesis that people who decide to victimise others and feel happy may have applied control strategies and deactivated self-sanctions, particularly regarding a specific situation. moreover, the process model of acting does not provide concepts specifying different kinds or qualities of cognitive control strategies; however, self-regulation theory does. the concept of mds (bandura, 1990; 2002; 2016) focuses on mechanisms people apply when choosing non-moral actions, such as transgressing against moral rules or victimising others. thus, in this paper, mds are chosen to specify mechanisms that supposedly support people in decision-making when experiencing internal conflict or ambivalence. applying mds may support the formation of intentions, even to engage in ‘immoral’ behaviours, such as moral transgressions or victimising others. 2.3 moral disengagement strategies during the last two decades, mds have been studied in various contexts (bandura, 2016), including those related to situations of high moral intensity (e.g., mcalister, bandura, & owen, 2006) and lower moral intensity, such as in work-related contexts (moore et al., 2012), sports (boardley & kavussanu, 2007), or connected to leadership issues (detert, treviño, & sweitzer, 2008). empirical findings reveal that mds affect prosocial behaviours and transgressions in childhood, adolescence (bandura, caprara, barbaranelli, pastorelli, & regalia, 2001), and adulthood (detert et al., 2008; fida, paciello, tramontano, fontaine, barbaranelli, & farnese, 2015; moore et al., 2012; osofsky, bandura, & zimbardo, 2005). research reveals that mds as a personal trait impacts ethical and unethical behaviour in a wide range of morally relevant situations—not only in situations of high moral intensity, such as moral dilemmas, but also in situations of lower moral intensity (moore et al., 2012). bandura assumed that “self-sanctions keep conduct in line with internal standards” (bandura, 1990, p. 28). otherwise, “disengagement of moral self-sanctions enables people to compromise their moral standards and still retain their senses of moral integrity” (bandura, 2016, p. 2). he differentiated four loci of mds: locus of the behaviour, agent of action, outcomes of action, and recipients affected by action. each of these loci indicates strategies allowing an individual to ignore moral standards and follow non-moral values (bandura, 2016; osofsky et al., 2005). the findings clearly indicate that mds have the power to specifically explain unethical behaviour. however, mds are mostly measured by using a scale to understand the propensity as a trait or tendency to apply mds in adolescence (bandura et al., 2001), and an adapted version for adults (moore et al., 2012). findings have confirmed individuals’ propensity to use mds as a predictor of unethical behaviour. furthermore, bandura reported that mds was triggered by specific contextual factors (bandura, 2002; moore et al., 2012). nevertheless, there is still a lack of insight into what type of mds are applied in certain situations and whether mds preference varies across situations. according to the process model of acting as explained above, mds might be particularly important in cases of ambivalence or conflicting aims or intentions. additionally, mds may affect forming an intention not only during reflective data processing but is assumed also to function as a filter in information processing when an individual intuitively commits to a non-moral action. an individual might make a decision based on heuristics, habits, or routines in everyday life, especially in cases of lower ambivalence and lower moral intensity, or when he or she does not have the opportunity to reflect, or accepts a suboptimal solution (esser, 1996; heinrichs, 2005). additionally, it can be assumed, in line with the mds approach, that people with a high personal tendency towards mds might develop a set of justifications consistent with their moral self, to manage internal ambivalence and conflicts in these morally relevant situations. mds may function as cognitive control (or volitional) strategies and support forming an intention towards victimising or obeying moral rules. thus, the application of mds used in a particular situation may (sometimes) become visible if people are asked for the reasons behind their preferred actions. 2.4 research questions to summarise, this paper offers theoretical approaches intended to contribute to explaining patterns of moral decision-making, such as the hv pattern, along with the process model of acting (heinrichs, 2005) and the concept of mds (bandura, 2016). the theoretical considerations focus on a procedural perspective of acting, particularly on reconstructing how happy or unhappy people are who intend to break or obey moral rules and, thus, show patterns of intentions with varying valence. it is assumed that patterns of moral decision-making and emotion attributions represent results of processes determined by personal and situational conditions, and may vary interpersonally and situationally. in addition to this theoretical approach to reconstructing the hv pattern, this paper is intended to empirically explore whether situational stimuli of low moral intensity may provoke intrapersonal variation of hv or uv patterns across situations. moreover, it is intended to gain insights and explore whether and to what extent adults apply mds in given situations. the following research questions are addressed: (1) to what extent do adults apply victimising strategies in situations of low moral intensity? (2) do patterns of moral decision-making and emotion attributions among adults vary intra-personally across situations of low intensity? (3) to what extent do adults apply mds to justify victimisation in (different) situations of low moral intensity? 3. method 3.1 data collection and sample characteristics in total, 587 university students from goethe university, frankfurt (germany; n = 201) and the university of bamberg (germany; n = 344) were surveyed using self-report questionnaires. thirty-four students were guest students from other universities and eight students did not report their university affiliation. on average, students had studied in total for 3.3 semesters (sd = 1.8, min. = 1, max. = 12). our sample comprised 213 male and 364 female students (10 students did not provide information regarding gender), with a mean age of 22.3 (sd = 2.9) years. thus, participants were emerging adults. of the observed students, 28.3% were studying to become teachers, 47.4% were studying economics, and 11.6 % were studying business education and educational management. participant filled in a self-report, paper-and-pencil questionnaire. questionnaires were provided in different university courses (e.g., educational psychology in the subject of teacher education studies, basics of scientific work, business ethics); thus, we used convenience sampling. in the questionnaire, students were confronted with descriptions of morally relevant situations. the stimuli used in this study did not focus on extreme moral conflicts or dilemmas, such as the death penalty, but focused on situations of lower moral intensity that emerge in everyday life. negative consequences of victimising others were limited to economic effects, such as high costs or losing money, or neglecting values relevant to social interactions, like honesty, trust, or legality. in the given cases, bodily harm or even death were not focused on as relevant consequences if the rule was disobeyed. moreover, we were interested in whether adults deal with such situations of lower moral intensity using sophisticated heuristics, in line with mds. in response to open-ended questions, the students were asked to make decisions, anticipate their own emotions, and provide reasons for their decisions and emotions; more precisely, we asked for their intentions. this way of capturing the hv pattern has been described from a self-perpetrator perspective as “self-judgments” (keller et al., 2003; yuill, pearson, pearbhoy, & van den ende, 1996; see also heinrichs et al., this issue). patterns of moral decision-making (here representing patterns of intentions) were coded (hv, uv, um, hm). to explore whether cognitive control strategies, for example mds, play a role in the action process, a content analysis of participants’ answers to open-ended questions regarding morally relevant decisions was conducted. the results are based on qualitative data (not the scale of mds). consistent with the concept of moral disengagement, data analysis was limited to participants who decided to transgress. applying mds is assumed to indicate perceived ambivalence, an output of moral decision-making in respect to selected situations. moreover, this operationalisation of ‘applied mds’ indicates that mds play a role in the action process and intention formation, either before committing to a particular action, or afterwards to justify a previously made decision. 3.2 operationalisation of constructs 3.2.1 patterns of moral decision-making to identify the situation-specific patterns of moral decision-making in terms of hm, um, hv, and uv, we used descriptions of two hypothetical situations of low moral intensity. in this study, moral intensity varied only regarding one criterion reported by jones (1991): the quality of the relationship between perpetrator and victim. other criteria that could potentially cause situations to be perceived as differing in moral intensity remained consistent across situations included in the survey. in situation 1 (‘travel costs’: an employee has to decide how to act confronted with the temptation to claim travel costs without having expenses), we chose a legal person (an organisation) as the victim, and in situation 2 (‘change’: a person receives to much change and has to decide to give the money back or not), we chose a natural person as the victim (for further description of the two situations, see the appendix). hm, um, hv, and uv were coded based on participants’ decisions as a relevant part of intentions. participants answered the question ‘what would you do’ (give the money back as the moral strategy vs. keep the money as the victimising strategy; self-judgment perspective). additionally, they rated their corresponding emotional state (‘how would you feel?’) by choosing one of the following options: very good, rather good, rather bad, or very bad. the ratings were re-coded as ‘happy’ (very good and rather good) and ‘unhappy’ (very bad and rather bad). therefore, the hm pattern was operationalised by keeping a moral rule and feeling (very or rather) happy, and the um pattern was defined by keeping a moral rule and feeling (very or rather) bad. furthermore, the hv pattern was defined by violating a moral rule and feeling (very or rather) happy, and the uv pattern was defined by violating a moral rule and feeling (very or rather) bad. 3.2.2 mechanisms of moral disengagement in the context of moral decision-making, participants were asked to provide reasons for their decisions. on that basis, answers to open-ended questions were coded to organise the given reasons via a coding scheme for mechanisms of mds, as adopted from the work of bandura, barbaranelli, caprara, and pastorelli (1996). in line with the idea of mds, only the answers of participants who decided to violate the moral rule (hv or uv) were considered. in total, the coding scheme consisted of the eight mechanisms of mds and a category ‘others’, as described in table 1. the reported coding scheme was the basis for coding participants’ reasons for their decisions. the coding categories were a priori defined theoretically and coding rules were determined, thus constituting the theoretical basis of our analysis (creswell, 2014; schreier, 2012). two independent researchers performed the coding. both were well-trained with sample codes. within a training round, the coders coded the participants' responses. when codes did not match, the respective sense units were discussed and assigned to a category by reaching a consensus. the category system was then further differentiated and validated. thus, the coding procedure was guided by the standard procedure of qualitative content analysis as described mayring (2015). to assess the inter-rater reliability of the coding, 64 reasons given by the participants (33 cases of situation 1 and 31 cases of situation 2), corresponding to almost 23 % of the overall applicable 280 cases, were coded by the two independent coders. cohen's kappa score was 0.613 for the cases using situation 1, and 0.713 for those using situation 2. for all the 64 double-coded cases, cohen’s kappa reached 0.7. therefore, the situation-specific coding, as well the overall coding, showed satisfactory inter-rater reliability that could be classified as “substantial” (range from 0.61 to 0.8), according to the corresponding ranges of kappa with respect to landis and koch (1977). cases coded in category ‘9 others’ were discussed individually by two coders within consensus validation and, if possible, were assigned to one of the categories 1 to 8. table 1 coding scheme for mechanisms of mds 4. results 4.1 variation of patterns of moral decision-making across situations first, analyses were conducted to answer research questions 1 and 2. the results (see table 2) indicated that all patterns of moral decision-making can be found in the sample. most participants chose hm (situation 1: 66.4 %; situation 2: 73.6 %). however, 10.7 % in situation 2 (n = 61) and up to 19.4 % (n = 110) in situation 1 showed the hv pattern. thus, victimising patterns can be applied in this sample of adults. table 2 intra-personal connections and variations of patterns of moral decision-making note. pearson χ2 = 152.092, df = 9, p < 0.001; 19 participants had a missing value for at least one of the two situations. regarding research question 2, cramer’s v (0.299; p < 0.001; pearson χ2 = 152.092, df = 9, p < 0.001) indicated a moderate-sized intrapersonal consistency of patterns of moral decision-making. however, there were participants who changed their pattern of moral decision-making across situations. for example, more than 25% out of students who rated hm in situation 2 rated victimising in situation 1. 4.2 mechanisms of mds as reasons for violating a moral rule along with the theoretical assumption of mds, only those participants who chose victimising strategies were integrated into the content analyses to answer research question 3. the findings indicated that mds were applied in the situations presented by those participants who chose victimising (situation 1: n = 159; situation 2: n = 111). descriptive frequency analyses showed that all categories of mds were used across both situations, though some mechanisms were not used in situations 2. however, frequency of the different mds varied across situational stimuli (see figures 2 and 3). figure 2. mechanisms of md as reasons for violating a moral rule: situation 1—travel costs figure 3. mechanisms of mds as reasons for violating a moral rule: situation 2—change 5. discussion and conclusions 5.1 main findings and limitations the process model of acting offers a framework to elucidate the role of cognitive control strategies, particularly mds, to forming intentions towards following or breaking a moral rule in situations of low moral intensity. further, it can provide theoretical progress and allows for a better understanding of patterns of moral decision-making and emotion attributions, such as hv, uv, um, and hm (aims and valence), in morally relevant situations. thus, this approach allows for the integration of perspectives of other theoretical approaches to hv in adulthood, as presented by minnameier (this issue) and gutzwiller-helfenfinger and latzko (this issue). the action-based perspective presented in this paper consider cognitive and emotional as well as volitional processes. additionally, the present study offers empirical results underlying the theoretical assumptions of the action-based approach to the hv pattern in situations of low moral intensity. the results related to research questions 1 and 2 merely support the findings of former studies regarding adults’ patterns of moral decision-making, as adults decided to victimise, and victimising emerged as a result of person-situation interactions. the patterns showed significant intrapersonal consistency; however, at the same time some participants showed variations in patterns across situations (heinrichs et al., 2015). moreover, the results of the present study enrich former empirical research on patterns of moral decision-making in adulthood, particularly by focusing on patterns of intentions in situations of low moral intensity. additionally, mds were assessed in the context of moral transgressions, not as a personal tendency. qualitative content analysis provided codes of mds within answers to open-ended questions, and showed that students who chose victimising applied mds in the two given situations of low moral intensity, to a relevant extent. frequency analyses of the different categories of mds indicated that mds use differed in quality and quantity across situations and between participants who attribute positive or negative emotions (between participants who show patterns of uv and hv; figures 2 and 3). in future research, it could be interesting to collect data from a bigger sample to go beyond descriptive methods of analyses and test whether these differences in mds use can be confirmed as significant effects triggered by situational conditions or as a personal tendency of mds use. however, the present study does not provide valid empirical evidence, but rather empirical insights regarding patterns of moral decision-making (um, uv, hv, hm) and mds use. thus, our empirical approach obviously has various limitations that are discussed comprehensively, to point out the potential of the presented approach to gain theoretical and empirical progress in future research (lakatos, 1978). 5.1.1 limitations and further perspectives regarding methods and data collection the present data were collected to examine to what extent moral decision-making, emotion attribution, and mds use emerged as intra-personally consistent or varying across situations. assuming that patterns of moral decision-making, as well as mds use, are the results of person-situation interactions, only two situational stimuli were used for comparisons between different situational conditions. however, no personal determinants or traits were included to control for personal conditions. thus, the results only allow for developing a hypothesis that patterns of moral decision-making might show up with intra-personal consistency, to a particular extent, while also indicating there is intra-personal variation across situations that has not yet been explained. regarding mds use, the results were mostly limited to a descriptive level. mds use was coded based on the participants’ answers to open-ended questions, in line with the theoretical assumption that they would only be used if a person had also chosen ‘victimising’. that led to a reduced sample of the coded mds: 159 participants for situation 1, 111 participants for situation 2, and only 32 participants who expressed mds in both situations. thus, this study does not provide reliable data on intrapersonal stability or variation of mds use. intra-personal consistent use vs. variation of mds use across situations of low moral intensity should be studied in future investigations. additionally, the present study is limited to situations of low moral intensity, predominantly characterised as situations of temptation. thus, the results do not provide information on how people react in situations of high moral intensity. furthermore, decisions were measured using the self-judgment perspective (‘what would you do?’ and ‘how would you feel?’; heinrichs et al., this issue) as indicators of participants’ intentions (aims, valence). however, in future research, it could be fruitful to use other methods to detect whether and to what extent participants experience ambivalence or internal conflict, and to what extent they feel committed to one method of action. moreover, participants were asked how they would act and feel in hypothetical situations presented as text-based stimuli. thus, the decisions they made in this study do not necessarily correspond to the behaviour they would display in real life. furthermore, there is a need to develop more realistic settings of data collection, allowing for a better understanding of decision-making, emotions, and actions. this might be possible, for example, in the field or with experimental studies. the results of the present investigation might also differ from those studies, if the participants were encouraged to reflect on their intentions as related to stimuli presented in an interview and embedded in interpersonal conversation, or to come up with their own narratives telling their experiences or motives in obeying or breaking moral rules in their everyday lives. moreover, from the very beginning of hv research, the measurement procedure has been criticised for not really identifying emotions, but emotion attributions or emotion justifications. it would be important to develop sophisticated and valid measures of moral emotions in the context of patterns of moral decision-making (see heinrichs et al., 2015; gutzwiller-helfenfinger & latzko, this issue). 5.1.2 limitations and further perspectives in regard to displaying the process of acting therefore, to reflect on the present study, the results call for further research to provide deeper insights into the process of intention formation and further sub-processes of moral actions. it would be interesting to know whether and to what extent individuals differ in subjectively constituted problems or in assessing the moral intensity of situations, depending on their individual moral principles or values, on selfand social sanctions, and on non-moral values and needs. mds use and hv or uv patterns may emerge, at least in different forms, depending on whether ambivalence or internal conflict was perceived, or whether the individual managed to cope with his or her negative emotions (bandura, 2016). to build commitment to one preferred way of acting (victimising or moral acting), the agent has to activate mechanisms of self-regulation based on social or self-sanctions (bandura, 2016). nevertheless, we must admit that the way of measuring intentions, emotions, or applied mds regarding the given situations is far from validly displaying the sequence of inner sub-processes. in future research, specific experimental studies may offer, for example, data on forming an intention, attributing emotions, or justifying decisions under contrasting conditions. 5.1.3 limitations and further perspectives from a developmental perspective furthermore, this paper mainly focused on the process of acting rather than on individual development. the present study only captures one measurement point and includes students representing emerging adults. to develop implications for moral education, it would be relevant to at least study the development of relevant personal determinants, such as mds. research on mds has provided interesting results from a developmental perspective. osofsky, bandura, and zimbardo (2005) indicated a low, but significant correlation between mds and age. older participants reported higher levels of mds. however, a correlation between age and mds only provides superficial indications and calls for a deeper understanding of underlying processes. in respect to this, bandura offered a more sophisticated assumption, that the development of mds is in line with a change of preference from social sanctions in childhood and earlier years to self-sanctions among adults (bandura, 2016). this idea is quite consistent with the basic assumption in the kohlbergian tradition of studying moral development. kohlberg claims social sanctions to be important at the preconventional level. at the conventional level, not one person but a social group or system, may sanction deviant or immoral behaviour. at the postconventional level the agent is assumed to have developed into a person with autonomous judgement and a strong ‘moral self’, committed to a hierarchy of values, and able to reflect on moral problems from a perspective of legitimacy. thus, the relevance of social sanctions seems to decrease; however, self-sanctions might increase on the way to higher moral stages. however, there is empirical evidence that such an individual-constructivist (vs. social-constructivist) idea of moral development towards an autonomous moral self ignores phenomena such as situational impact on moral judgements and patterns of hv (as well as hm, um, hm) during adolescence and adulthood (beck & parche-kawik, 2004; heinrichs et al., 2015; krebs & denton, 2005). 5.2 conclusions the action-theoretical approach presented in this paper provides a perspective to explain hv as a pattern emerging within the action process, particularly regarding immoral behaviour among adults. people choose to jeopardise their own standards and seem to feel happy about a decision in favour of a moral transgression if they manage to find ways to cope with inner conflicts and get strongly committed towards their intention. this idea is in line with bandura, who presented the concept of moral disengagement as an indicator of lacking self-regulation concerning cases of immoral behaviour. to conclude, the three approaches, the action-theoretical model, the hv pattern in adulthood as a pattern emerging during the process of acting, as well as the concept of moral disengagement, address the same basic question: how can people act in ways that contradict their moral principles, at least in some situations, without feeling distress or having a bad conscience? this study indicates that we have to acknowledge that breaking moral rules or standards is quite common, even among adults, at least as studied here in situations of lower moral intensity. otherwise, findings have revealed that people do commit to fairness, sharing, and others’ well-being (fehr & fischbacher, 2004; fehr & schmidt, 1999, 2006). however, the results presented are far from being valid for drawing evidence-based conclusions concerning moral education. nevertheless, the approach reconstructs (happy or unhappy) victimising in an action-based perspective as a result of intention formation, and provides long-term perspectives for education that also differ from those discussed in this special issue on the hv pattern in adulthood and adolescence (gutzwiller-helfenfinger & latzko, this issue; minnameier, this issue). to reduce mds use or foster reflective use of strategies like mds offers an approach that differs from common methods of moral education, such as supporting moral judgment competence with respect to cognitive development (minnameier, 2012), fostering moral expertise (narvaez & lapsley, 2005), or developing moral emotions (gutzwiller-helfenfinger & latzko, this issue). additionally, from an educational perspective, the results of the present study only pointed to the empirical ‘is’ in decision-making, intentions, emotion attributions, and moral disengagement in morally relevant situations. there is an additional need to discuss aims considering also norms and values and legitimate curricula. it seems important to reflect on whether students should be encouraged to follow their moral ideals in certain situations, even if they must accept great (personal) disadvantages. perhaps it would be preferable to enable them to balance their own needs and those of others. minnameier would argue that sometimes, it is morally adequate to behave as a strategic moralist (minnameier, this issue; minnameier, heinrichs & kirschbaum, 2016). he states, for example, that to implement trustful cooperation in the long term, depending on the situational conditions, an individual may have to adopt his or her way of acting towards moral standards of partners or one’s environment. sometimes it may be recommended to look for strategies (e.g., in cooperative games) that may lead to moral behaviours in a sense of an overarching moral aim, such as developing trustful cooperation and preventing others from pursuing only their own needs and getting rich at the expense of the individual in the long-term. overall, there are considerably contrasting positions concerning the aims of moral education. it is discussed controversially whether moral education should focus on struggling to promote autonomously judging moral agents and accept that some of them may end up as unhappy moralists (oser & reichenbach, 2005), or whether it would be better to foster individuals who are able to choose ways of acting that may lead to showing hv patterns when considering the situational conditions for implementing morality and acting as strategic moralists. however, there seems to be a common position underlying this discourse of ‘successfully engaging in meaningful, positive and caring relationships is both a prerequisite for and consequence of successful teaching and learning processes (malti, häcker, & nakamura, 2009). moreover, meaningful relationships are especially important in a globalised society […]‘ (gutzwiller-helfenfinger & heinrichs, this issue) to create a ‘moral atmosphere’ (kohlberg, 1984), supporting moral actions and the development of socio-moral competencies. keypoints adults tending towards moral transgression used various moral disengagement strategies (mds) across situations. fostering a reflective use of mds might support formation of intentions with high valence at least in morally relevant situations of low moral intensity. abbreviations hm = happy moralizing pattern; hv = happy victimizing pattern; mds = moral disengagement strategies; um = unhappy moralizing pattern; uv = unhappy victimizing pattern footnotes 1 the definition of problem used here mainly points to the individually constituted discrepancy between is and ought. in terms of problem-solving approaches (e.g. following dörner, 1979), it includes tasks and problems (for a more sophisticated discussion, see heinrichs, 2005). references bandura, a. (1990). mechanisms of moral disengagement in terrorism. in reich, w. (ed.), origins of terrorism: psychologies, ideologies, theologies, states of mind (pp. 161–191). cambridge: woodrow wilson center press. bandura, a. (2002). selective moral disengagement in the exercise of moral agency. journal of moral education, 31(2), 101–119. https://doi.org/10.1080/0305724022014322 bandura, a. (2016). moral disengagement: how people do harm and live with themselves. new york, ny: worth publishers. bandura, a., barbaranelli, c., caprara, g. v., & pastorelli, c. (1996). mechanisms of moral disengagement in the exercise of moral agency. journal of personality and social psychology, 71(2), 364–374. https://doi.org/10.1037/0022-3514.71.2.364 bandura, a., caprara, g. v., barbaranelli, c., pastorelli, c., & regalia, c. (2001). sociocognitive self-regulatory mechanisms governing transgressive behaviour. journal of personality and social psychology, 80(1), 125–135. https://doi.org/10.1037/0022-3514.80.1.125 beck, k., & parche-kawik, k. (2004). das mäntelchen im wind? zur domänenspezifität moralischen urteilens [the wave in the wind? towards domain specificity of moral judging]. zeitschrift für pädagogik, 50(2), 244–265. boardley, i. d., & kavussanu, m. (2007). development and validation of the moral disengagement in sport scale. journal of sport and exercise psychology, 29(5), 608–628. https://doi.org/10.1123/jsep.29.5.608 creswell, j. w. (2014). research design: qualitative, quantitative, and mixed methods approaches (4th ed.). los angeles: sage. detert, j. r., treviño, l. k., & sweitzer, v. l. (2008). moral disengagement in ethical decision-making: a study of antecedents and outcomes. journal of applied psychology, 93(2), 374–391. https://doi.org/10.1037/0021-9010.93.2.374 döring, b. (2013). the development of moral identity and moral motivation in childhood and adolescence. in k. heinrichs, f., oser, & t. lovat (eds.), handbook of moral motivation. theories, models, applications (pp. 289–305). rotterdam: sense publishers. dörner, d. (1979). problemlösen als informationsverarbeitungsprozess [problem solving as information processing]. stuttgart: kohlhammer. eisenberg, n., fabes, r. a., & spinrad, t. l. (2007). prosocial development. in n. eisenberg, w. damon, & r. m. lerner (eds.), handbook of child psychology: social, emotional, and personality development (p. 646–718). john wiley & sons inc. https://doi.org/10.1002/9780470147658.chpsy0311 esser, h. (1996). die definition der situation [definition of the situation].kölner zeitschrift für soziologie und sozialpsychologie, 48(1), 1–34. fehr, e., & schmidt, k. m. (1999). a theory of fairness, competition, and cooperation. quarterly journal of economics, 114(3), 817–868. https://doi.org/10.1162/003355399556151 fehr, e., & schmidt, k. (2006). the economics of fairness, reciprocity and altruism: experimental evidence and new theories. in: s.-c. kolm, & j. m. ythier (eds.), handbook of the economics of giving, altruism and reciprocity (pp. 615–691). amsterdam: elsevier. fehr, e., & fischbacher, u. (2004). third-party punishment and social norms. evolution and human behaviour, 25(2), 63–87. https://doi.org/10.1016/s1090-5138(04)00005-4 fida, r., paciello, m., tramontano, c., fontaine, r. g., barbaranelli, c., & farnese, m. l. (2015). an integrative approach to understanding counterproductive work behaviour: the roles of stressors, negative emotions, and moral disengagement. journal of business ethics, 130(1), 131–144. https://doi.org/10.1007/s10551-014-2209-5 gasser, l., gutzwiller-helfenfinger, e., latzko, b., & malti, t. (2013). moral emotion attributions and moral motivation. in k. heinrichs, f. oser, & t. lovat (eds.), handbook of moral motivation. theories, models and applications (pp. 304–320). rotterdam: sense publishers. gollwitzer, p.m. (1996). rubikonmodell der handlungsphasen. [rubikon model of action phases] in j. kuhl, & h. heckhausen (eds.), motivation, volition und handlung (pp. 531–582). göttingen: hogrefe. gutzwiller-helfenfinger, e., & heinrichs, k. (2020). the happy victimizer pattern in adulthood – state of the art and contrasting approaches: introduction to the special issue. frontline learning research, 8(5), 1-4. https://doi.org/10.14786/flr.v8i5.681 gutzwiller-helfenfinger, e., & latzko, b. (2020). happy victimizing in emerging adulthood: reconstruction of a developmental phenomenon? frontline learning research, 8(5), 47-69. https://doi.org/10.14786/flr.v8i5.382 haidt, j., & craig j. (2008). the moral mind: how five sets of innate intuitions guide the development of many culture-specific virtues, and perhaps even modules. in p. carruthers, s. laurence, & s.p. stich (eds.), the innate mind: foundations and the future, evolution and cognition (vol. 3; pp. 367–391). oxford: oxford university press. heckhausen, h., gollwitzer, p.m., & weinert, f. e. (1987). jenseits des rubikon [beyond the rubicon]. berlin: springer. heinrichs, k. (2005). urteilen und handeln – ein prozessmodell und seine moralpsychologische spezifizierung . [a process model of judging and acting and its moral psychological specification] (vol. 12). frankfurt a. m.: peter-lang-verlag. heinrichs, k. (2010). urteilen und handeln in der moralischen entwicklung [judging and action in a developmental perspective]. in b. latzko, & t. malti (eds.), moralentwicklung und moralerziehung in kindheit und adoleszenz (s. 69–86). göttingen: hogrefe. heinrichs, k. (2013). moral motivation in the light of action theory. in k. heinrichs, f. oser, & t. lovat (eds.),handbook of moral motivation. theories, models and applications (pp. 623–657). rotterdam: sense publishers. heinrichs, k., gutzwiller-helfenfinger, e., latzko, b., minnameier, g., & döring, b. (2020). happy victimizing in adolescence and adulthood – empirical findings and further perspectives, frontline learning research, 8(5), 5-23. https://doi.org/10.14786/flr.v8i5.385 heinrichs, k., minnameier, g., latzko, b., & gutzwiller-helfenfinger, e. (2015). „don’t worry, be happy“? – das happy-victimizer-phänomen im berufsund wirtschaftspädagogischen kontext [“don’t worry, be happy”? – the happy victimizer phenomenon in the context of vocational and business education]. zeitschrift für berufsund wirtschaftspädagogik, 111(1), 32–55. jones, t. m. (1991). ethical decision-making by individuals in organisations: an issue contingent model. academy of management review, 16(2), 366–395. https://doi.org/10.5465/amr.1991.4278958 keller, m., lourenço, o., malti, t., & saalbach, h. (2003). the multifaceted phenomenon of ‘happy victimizers’: a cross‐cultural comparison of moral emotions. british journal of developmental psychology, 21(1), 1–18. https://doi.org/10.1348/026151003321164582 kohlberg, l. (1984). the psychology on moral development: the nature and validity of moral stages (vol. ii). new york: harper & row. krebs, d. l., & denton, k. (2005). toward a more pragmatic approach to morality: a critical evaluation of kohlberg's model. psychological review, 112(3), 629–649. https://doi.org/10.1037/0033-295x.112.3.629 krettenauer, t., malti, t., & sokol, b. (2008). the development of moral emotions and the happy victimizer phenomenon: a critical review of theory and applications. european journal of developmental science , 2, 221–235. https://doi.org/ 10.3233/dev-2008-2303 lakatos, i. (1978). the methodology of scientific research programmes. cambridge, ma: cambridge university press. landis, j. r., & koch, g. g. (1977). the measurement of observer agreement for categorical data. biometrics, 33(1), 159–174. https://doi.org/ 10.2307/2529310 lapsley, d. k., & narvaez, d. (2005). moral psychology at the crossroads. character psychology and character education, in d. k. lapsley, & c. power (eds.), character psychology and character education (pp. 18–35). notre dame: university of notre dame press. latzko, b., & malti, t. (2010). children's moral emotions and moral cognition: towards an integrative perspective.new directions for child and adolescent development, 129, 1-10. https://doi.org/10.1002/cd.272 lind, g. (2019). moral ist lehrbar!: wie man moralisch-demokratische fähigkeiten fördern und damit gewalt, betrug und macht mindern kann [morality can be taught! how we can foster moral-democratic abilities and reduce violence, fraud and power]. berlin: logos. malti, t., & krettenauer, t. (2013). the relation of moral emotion attributions to prosocial and antisocial behaviour: a meta‐analysis. child development, 84, 397–412. https://doi.org/10.1111/j.1467-8624.2012.01851.x malti, t., häcker, t., & nakamura, y. (2009). sozial-emotionales lernen in der schule [socioemotional learning in schools]. zurich: pestalozzianum. malti, t., averdijk, m., zuffianò, a., ribeaud, d., betts, l. r., rotenberg, k. j., & eisner, m. p. (2016). children’s trust and the development of prosocial behavior. international journal of behavioral development, 40(3), 262–270. https://doi.org/10.1177/0165025415584628 mayring, p. (2015). qualitative inhaltsanalyse: grundlagen und techniken [qualitative content analysis: basics and techniques]. weinheim: beltz. mcalister, a. l., bandura, a., & owen, s. v. (2006). mechanisms of moral disengagement in support of military force: the impact of sept. 11. journal of social and clinical psychology, 25(2), 141–165. h ttps://doi.org/10.1521/jscp.2006.25.2.141 minnameier, g. (2020). explaining happy victimizing in adulthood – a cognitive and economic approach. frontline learning research, 8(5),70 91. https://doi.org/10.14786/flr.v8i5.381 minnameier, g. (2012). a cognitive approach to the ‘happy victimizer’. journal of moral education, 41(4), 491–508. https://doi.org/10.1080/03057240.2012.700893 minnameier, g., heinrichs, k., & kirschbaum, f. (2016). sozialkompetenz als moralkompetenz – wirklichkeit und anspruch? [social competence as moral competence – reality and demand?]. zeitschrift für berufsund wirtschaftspädagogik, 112(4), 636–666. minnameier, g., & schmidt, s. (2013). situational moral adjustment and the happy victimizer. european journal of developmental psychology , 10(2), 253–268. https://doi.org/10.1080/17405629.2013.765797 moore, c., detert, j. r., treviño, l., baker, v. l., & mayer, d. m. (2012). why employees do bad things: moral disengagement and unethical organisational behaviour. personnel psychology, 65(1), 1–48. https://doi.org/10.1111/j.1744-6570.2011.01237.x narvaez, d., & lapsley, d. k. (2005). the psychological foundations of everyday morality and moral expertise. in d. k. lapsley, & c. power (eds.), character psychology and character education (pp. 140–165). notre dame: university of notre dame press. neuweg, g. h. (1999). könnerschaft und implizites wissen zur lehr-lerntheoretischen bedeutung der erkenntnisund wissenstheorie michael polanyis [expertise and tacit knowledge – the relevance of michael polanyi’s theory of knowledge for teaching and learning]. münster: waxmann. nunner-winkler, g. (2013). moral motivation and the happy victimizer phenomenon. in: k. heinrichs, f. oser, & t. lovat (eds.), handbook of moral motivation. theories, models, applications (pp. 267–287). rotterdam: sense publishers. nunner-winkler, g. (2007). development of moral motivation from childhood to early adulthood. journal of moral education, 36(4), 399–414. https://doi.org/10.1080/03057240701687970 nunner-winkler, g. (1993). die entwicklung moralischer motivation [the development of moral motivation]. in w. edelstein, g. nunner-winkler, & g. noam, g. (eds.), moral und person (pp. 278–303). frankfurt am main: suhrkamp. nunner-winkler, g., & sodian, b. (1988). children's understanding of moral emotions. child development, 59, 1323–1338. https://doi.org/10.2307/1130495 oser, f. k., & reichenbach, r. (2005). moral resilience – the unhappy moralist. in w. edelstein, & g. nunner-winkler (eds.), morality in context, advances in psychology (pp. 203–224). north holland: elsevier. osofsky, m. j., bandura, a., & zimbardo, p. g. (2005). the role of moral disengagement in the execution process. law and human behaviour, 29(4), 371–393. https://doi.org/10.2307/1130495 rothmund, t., & baumert, a. (2014). shame on me implicit assessment of negative moral self-evaluation in shame-proneness. social psychological and personality science, 5(2), 195–202. https://doi.org/10.1177/1948550613488950 schreier, m. (2012). qualitative content analysis in practice. los angeles: sage. sokolowski, k. (1993). emotion und volition [ emotion and volition]. göttingen: hogrefe. sokolowski k. (1996). wille und bewusstheit [will and consciousness]. in: j. kuhl, & h. heckhausen (eds.), enzyklopädie der psychologie: themenbereich c theorie und forschung, serie iv motivation und emotion, band 4 motivation, volition und handlung (pp. 485–530). göttingen: hogrefe. thoma, s. j., & bebeau, m. j. (2013). moral motivation and the four component model. in k. heinrichs, f. oser, & t. lovat (eds.), handbook of moral motivation. theories, models and applications (pp. 49–67). rotterdam: sense publishers. tirri, k. (1999). teachers' perceptions of moral dilemmas at school. journal of moral education, 28(1), 31–47. https://doi.org/10.1080/030572499103296 weinberger, a. & frewein, k. (2019). vake (values and knowledge education) als methode zur integration von werterziehung im fachunterricht in heterogenen klassen beruflicher schulen: förderung von kognitiven und affektiven zielen [vake (values and knowledge eduation): a measure for integrating values education in domain specific lessons in heterogeneous classes in vocational schools: fostering cognitive and affective goals]. in k. heinrichs & h. reinke (eds), heterogenität in der beruflichen bildung. im spannungsfeld von erziehung, förderung und fachausbildung (s. 181–194). reihe wirtschaft – beruf – ethik. bielefeld: wbv. appendix hypothetical situations for the assessment of moral decision-making [introduction for the participants] the following section shows various descriptions of situations. we would like to ask you to read through these situations carefully and to imagine yourself in the described situations. there are always exactly two alternative actions. please decide in favour of one of them by marking the appropriate alternative with a cross. afterwards, you will be asked to give reasons for your decision. please indicate the main reasons in any case. in the end, you will be asked how you would have felt by acting as stated in each case. situation 1—travel costs please imagine the following situation: you are working for a large international company and attend a national meeting within the framework of your employment company. accidentally, a friend hast to go to the same city as you do. he offers to give you a ride. you accept his offer willingly. during the return journey, you have a nice chat. nobody from your company knows or has noticed that you didn’t take your own car, and therefore have no expenses of your own. the travel expenses for this journey would be 50 euros. the formula for the travel expenses report says that only real travel costs are refundable. the friend can settle up his travel expenses himself, so that you don’t have to refund him anything. please put yourself into this situation and decide what you would do: ( ) you claim the travel expenses of 50 euros. ( ) you don’t claim the travel expenses of 50 euros. please give reasons for your decision… [these statements from participants are the basis for coding the mechanisms of mds] how would you feel if you would had really acted like that? please mark only one possible answer. ( ) very good ( ) rather good ( ) rather bad ( ) very bad situation 2—change please imagine the following situation: due to the fact that you speak spanish very well, you go on holiday in spain. you have earned the money for your holiday by working in a factory. during a day trip by bus to a distant city, you buy a handmade wooden sculpture for your parents in a craft shop. it is 50 euros. you pay with a 200 euro bill, and leave the shop. as you count your change outside the shop, you notice that the seller has given you back four 50 euro bills instead of 150 euros, inadvertently. please put yourself in this situation and decide what you would do: ( ) you return 50 euros to the seller. ( ) you keep the 50 euros. please give reasons for your decision… [these statements from participants are the basis for coding the mechanisms of mds] how would you feel if you would had really acted like that? please mark only one possible answer. ( ) very good ( ) rather good ( ) rather bad ( ) very bad codepen jorion et al frontline learning research vol.8 no. 6 (2020) 77 87 issn 2295-3159 uncovering patterns in collaborative interactions via cluster analysis of museum exhibit logfiles natalie jorion1, jessica roberts2, alex bowers3, mike tissenbaum4, leilah lyons5, vishesh kumar4 & matthew berland4 1psi services, llc, usa 2carnegie mellon, usa 3teachers college columbia university, usa 4university of wisconsin—madison, usa 5university of illinois at chicago, usa article received 30 december 2019 / revised 10 may/ accepted 12 october/ available online 4 november abstract a driving factor in designing interactive museum exhibits to support simultaneous users is that visitors learn from one another via observation and conversation. researchers typically analyze such collaborative interactions among museumgoers through manual coding of liveor video-recorded exhibit use. we sought to determine how log data from an interactive multi-user exhibit could indicate patterns in visitor interactions that could shed light on informal collaborative learning. we characterized patterns from log data generated by an interactive tangible tabletop exhibit using factors like “pace of activity” and the timing of “success events.” here we describe processes for parsing and visualizing log data and explore what these processes revealed about individual and group interactions with interactive museum exhibits. using clustering techniques to categorize museumgoer behavior and heat maps to visualize patterns in the log data, we found distinct trends in how users solved the exhibit. some players seemed more reflective, while others seemed more achievement oriented. we also found that the most productive sessions occurred when players occupied all four areas of the table, suggesting that the activity design had the desired outcome of promoting collaborative activity. keywords: collaborative interactions; game-based learning; heat maps; cluster analysis; log files; informal learning environments info corresponding author email: talie.jorion@gmail.com doi:https://doi.org/10.14786/flr.v8i6.597 1. objectives while the design community harbors enthusiasm for log data analysis, many open questions exist about what this decontextualized data can tell us about collaborative interactions in open-ended tinkering activities. multi-user touch-tables are gaining popularity in informal spaces like museums for their ability to allow groups of users to interact with the exhibit in a social setting simultaneously (block et al., 2015; davis et al., 2013). such exhibit usage affords interactions conducive to learning, such as observing and exploratory tinkering (roberts & lyons, 2017; yoon et al., 2012). these productive behaviors are observable in live interactions, but observations that span multiple weeks produce too much data to analyze manually (snibbe, 2006). we hypothesize that log data from such a tabletop interactive exhibit could help identify usage patterns that reveal insights into the collaborative interactions. to that aim, we analyzed logfile data from a collaborative tabletop exhibit, oztoc (tissenbaum et al., 2017, figure 1), through hierarchical cluster analysis (hca) and heatmap visualizations. the game’s objective is to catch fish by creating circuits with different colors and numbers of leds. the game space has four separate play stations by design to facilitate collaborative inquiry. the players can create circuits on their own or with the help of others. figure 1. the oztoc exhibit. the scoreboard is in the background. in this study, we define collaboration as the process by which two or more agents share pooled understanding to work on a problem. problems that are complex or ill-defined are well-suited to collaborative problem solving and can lead to individual learning. we conceptualize learning through a constructivist lens as actively building on prior knowledge through experimentation and reflection. interactive exhibits allow players to make inferences about a game’s underlying principles to refine their game tactics. much research examining collaborative interaction in museums focuses on visitor dialogue, analyzing how learners construct and negotiate meaning through talk as they explore an exhibit (e.g., martin et al., 2018; tissenbaum et al., 2017). analyses frequently contextualize this dialogue in relation to visitors’ physical interactions with each other and the exhibit (long et al., 2019; roberts et al., 2018). these studies rely on dataand time-intensive video recordings, limiting the scope of most studies to small sample populations. the increase of technology-based exhibits in museums in recent years can enable an up-close analysis of exhibit interactions via touch events captured in log files. while some work has shown promise in revealing interaction through touch events alone (evans et al., 2016), researchers have not established practical methods for measuring collaborative interactions through such files. this study used cluster analysis heatmaps to group patterns of interactions and visualize log data. cluster analysis is a descriptive data mining technique in which vectors of data for individuals within a dataset are patterned based on similarity (alfredo et al., 2010; bowers, 2010; romesburg, 1984). in this study, we used hierarchical cluster analysis (hca), which researchers have shown work well with education data (bowers, 2010; lee et al., 2016). here for the hca, we used k-means clustering with average linkage as the agglomeration method, which builds up a hierarchical “tree” of log entries by recursively merging those entries that have higher-than-average pairwise similarities. the resulting branches indicate the dominant clusters of patterns present in the logs. additionally, we visualized the data with hca heatmaps (bowers, 2010; lee et al., 2016). cluster analysis heatmaps are a well-established means to visualize high-dimensionality patterns in data while retaining and displaying individual data points in their context and examining fine-grained patterns in the data and identifying overall informative clusters (wilkinson & friendly, 2009). cluster analysis is incredibly helpful for exploring under-studied phenomena, like student programming styles (berland et al., 2013), expert and novice practices in maker activities (blikstein, 2013), and how students use exploratory learning environments (amershi & conati, 2009). researchers have also used heatmaps to highlight patterns in teamwork and participation (upton & kay, 2009). 2. methods 2.1 logfile data using the adage system (lyons, 2014), the oztoc table tracks the movement, addition, and removal of blocks (or circuit “components”) from the table. it also records actions performed with the blocks while on the table (e.g., connections between components). the program logs each event with a timestamp and x-y coordinates for four types of circuit component blocks (resistors, batteries, leds, and timers). we computed the seconds from the timestamps since the last event to determine the movement rate of pieces, which we use as a proxy for the players’ level of activity. play stations, playerids, and sessionids were assigned as described next. 2.2 identifying unique users understanding user behavior through logfiles requires first determining how to demarcate unique users in both space and time (block et al., 2015; tissenbaum et al., 2016). oztoc does not constrain users to delimited stations around the table, but the visual design divides the table into four labeled play stations associated with player feedback displayed on a large scoreboard screen. informal observations indicated that users tended to respond to these design cues and occupy one “play station,” limiting block manipulation and circuit-construction activities to the labeled space. we confirmed this by charting the point coordinates from all the events to the table (figure 2). figure 2. block placements during play, color-coded by component type, fall into four quadrants on the touch-table. this map confirms that players rarely worked in the middle of the table, which parallels findings from similar studies of shared tabletop use (martínez maldonado et al., 2010). this was important for us to confirm because the features we cluster require that the block manipulations and circuit completion events can be attributed to a specific user. we also had to identify how to demarcate players by time since oztoc runs continually and does not have a dedicated start and stop event. to determine when visitors likely entered or exited the exhibit, the designers analyzed board actions on a single day using a script with inactivity intervals ranging from 10 seconds to 120 seconds. they found that the inactivity metric did not change significantly, ranging from 45 to 120 seconds. to validate the 45-second cutoff time, they hand-labeled a 2-hour sample of video data and found that it was 100% accurate (tissenbaum et al., 2016). therefore, lags over 45 seconds for all four player spaces on the table triggered the creation of a new group of users. 3. data sources the data analyzed here are from an open-ended engineering exhibit situated at the new york hall of science (lyons et al., 2015). in the exhibit, visitors design and build glowing fishing lures to attract simulated bioluminescent fish (figure 3). to catch all the different fish, players must experiment with creating circuits with different colors and numbers of leds. wooden blocks represent resistors (1), batteries (2), timers (3), and different colored leds (4). participants make circuit connections by bringing the blocks’ positive and negative terminals in contact with one another. creating a successful circuit causes the leds to glow and lures the fish attracted to that light out for cataloging. figure 3. depiction of the oztoc exhibit. this study tests the validity and utility of cluster heatmapping for revealing group interaction patterns by analyzing records from a single afternoon of exhibit usage. in the 133 minutes of recorded interactions during that period, 6,077 unique events were logged as described above. 4. results with players and play stations defined, we restructured the dataset into 10-second intervals of “playtime.” figure 4 displays players as rows and each of the fifty 10-second timeslices as columns, with a brighter red indicating 20 or more actions within those 10 seconds, and white indicates no data. “actions” are defined as any logged event in that player’s playstation during that 10 seconds. the clusters were formed by cross comparing the activity levels in these player “trajectories” to investigate whether players engage with the table with the same patterns of intensity over time. the annotations to the heatmap’s right are labels and were not used to group the player trajectories. “#circuits” represent the total number of circuits created by that player, with the brighter red indicating more circuits created. “#circuits per sec” shows the total number of circuits created per second. the last four annotation columns on the far right indicate how many other simultaneous players were present with that player. figure 4 shows that when individual player trajectories are clustered by similarity, there are at least three different patterns with different “outcomes” (measured by circuits created). the longest and most active clusters only occur when there are four players. the players depicted in cluster 3 (figure 4, bottom) use oztoc for 10 seconds to one minute. these players usually play with few other simultaneous players and make few (if any) circuits. cluster 1 (figure 4, top) players have longer play sessions in which there is a consistent level of activity, and players make many circuits. cluster 2 (figure 4, middle) players make the most circuits and have the highest activity per 10-second intervals. these players are typified by somewhat shorter play sessions than cluster 1 and end with rapid tile movement. figure 4. a heatmap revealing clusters based on raw counts of actions over time. we then standardized the activity counts using the grand mean across the dataset to reveal differences in activity better. whereas figure 4 uses raw numbers, figure 5 uses z-scores of the number of movements. each row displays a player, and each of the fifty 10-second timeslices as columns, with annotations to the right. brighter blue indicates movement increasingly below the mean, with the darkest blue indicating no movement. darker red indicates a higher average movement of tiles above the mean. in addition, this heatmap was annotated with yellow dots to indicate when the player completed a circuit (although we did not use this information to cluster the player trajectories). these dots reveal the extent to which players are planning before completing a circuit. there seems to be a relationship between the number of moves and completed circuits, and particularly whether players tend to complete more circuits after periods of high or low frequencies of movements. figure 5. a heatmap revealing clusters of player trajectories based on z-standardized levels of activity over time. figure 5 shows that at least three different types of players (though standardizing the scores cause some individual players to shift between clusters in the two different analyses, so cluster 1 and cluster 4 do not comprise the same players). while clusters 4 and 5 show players who had three companions, and cluster 6 represents the cluster of players typically in small groups who only interact minimally with oztoc, cluster 5 players alternate between blue and red, indicating that they pause often. these players stay at the table longer than cluster 4 and make fewer circuits overall. moreover, many of the yellow dots (completed circuits) are in or followed by blue intervals, indicating a pause likely due to the player watching the visual effect triggered by a completed circuit. cluster 4 has the most red with the lowest blue, indicating continual active interaction with the tiles (indicating more “achievement-oriented” behavior). these players also have rising tile movement across their session, often peaking with yellow dot “streaks” of circuit completion. this group completes more circuits than cluster 5, in which there are fewer and more randomly distributed yellow dots. we expected to find variations in duration and circuit creation rates among users of this exhibit and larger groups to have richer interactions. however, the apparent clustering of these interaction patterns and their strong relationship to the number of concurrent players at the table were surprising. according to this data, the more players involved in the game, the more time they spend tinkering, and the more circuits they create, suggesting the game fosters collaborative behavior. however, the data did not necessarily show that the game fostered constructivist learning, as some players seemed to be more reflective (cluster 5), while others seemed to be achievement-oriented (cluster 4). a closer investigation of which players belonged to the same play sessions revealed that cluster 4 and cluster 5 players are present in sessions together, but the cluster 6 players are mostly in sessions with other cluster 6 players. four of the six cluster 6 players in groups of four belonged to the same session, suggesting as a group, they may not have understood the game or that their interaction ended prematurely. the other two were members of 5-person sessions, suggesting that the player may have begun at the end of the group’s session and left with the group shortly after starting. 5. limitations given that we inferred player actions from log file data in this study and did not validate each action for the entirety of the data, some of the inferences may be inaccurate. for example, a new player could have made a move within the 45-second cutoff interval, and then two players would have erroneously been labeled as one player. it is also possible that one player could move to the play station of another user. however, given that we validated some of these inferences by hand-coding a sample of video data with a high degree of accuracy, we assume that these instances are few and make a negligible impact on the overall inferences we make in this study. moreover, we do not want to overstate claims about the degree to which the activity traces provide evidence for players demonstrating constructivist learning interactions. we would need supplementary video or interview data to corroborate that the most productive sessions were ones in which there was constructive dialogue indicating that the players were developing more accurate conceptual models of circuit building. 6. discussion making sense of collaborative learning in unstructured activities is a complicated task. however, the difficulty of this problem should not overshadow the importance of deriving methodologies to find evidence that a system is working as intended. this study demonstrates how we derived data for collaborative learning in an interactive museum exhibit using unstructured data. first, we defined what data would constitute evidence for learning. our metric was not just a simple measure of “success” (i.e., total circuits created) but of “productive behavior” (the relationship of the player moves to the game outcomes). to do so, we investigated player behavior over time to evaluate the extent to which the game fostered “productive” sessions. we assume that when participants make fewer moves to create a unique working circuit, they have a better conceptual understanding of the circuits’ mechanics. making fewer moves would contrast to “unproductive” behavior in which players either do not realize the game’s goal or consistently use a brute-force approach with no changes in behavior. we found distinct interaction patterns related to the movement rate of the tangible blocks and the number of completed tasks (here, the construction of working circuits). second, we accounted for different amounts of time the users interacted with the game and standardized the time blocks, so they were comparable across users. third, we found patterns in behavior using cluster analysis and heatmaps, which helps to visualize large and complex datasets. the visualizations revealed that the most productive interactions occurred when the table was fully occupied with four simultaneous users, suggesting that collaboration led to more successful goal-oriented actions. this provides evidence that the exhibit fostered collaborative interactions as the designers intended (lyons et al., 2015). although this method is not as reliable as a traditional individual assessment, the information derived is richer and can provide more useful information on learning strategies and how systems affect user behavior patterns on a large scale. 7. implications for future work this analysis presents several avenues for future investigations. we have hypothesized that the pauses seen with cluster 5 players represent planning or strategizing before completing a circuit. further investigation can determine how players’ completion events and pauses relate to companions’ activities. for example, the visualization could be reorganized to display players together by a session with each player offset according to their start times to indicate productive group collaboration. this could indicate whether a user’s completion of a circuit prompts echoing of the same circuit by others. furthermore, these visualizations help flag events for finer-grained qualitative analysis to understand the interactions happening “above the table” among participants (tissenbaum et al., 2017). if further work shows that players in a cluster demonstrate similar challenges, this could help us provide real-time participant support. such information could alert museum facilitators when to intervene when visitors are frustrated, consequently increasing visitor dwell time and overall domain learning. this method of using log data to identify user strategies and states to provide just-in-time scaffolding can be a useful method for facilitators looking to improve user engagement in other less-structured museum spaces and online learning environments. keypoints heatmaps can be a helpful way of visualizing logfile data of user interactions and finding clustering trends in collative user behavior for this museum exhibit, the more players participated in the game, the longer the session lasted, suggesting that the activity promoted collaborative activity in terms of problem-solving, two styles of playing emerged: reflective versus more achievement-oriented acknowledgments this research was made possible by the national science foundation data consortium fellowship and the new york hall of science. references alfredo, v., félix, c., & àngela, n. (2010). clustering educational data. in c. romero, s. ventura, m. pechenizkiy & r. s. j. d. baker (eds.), handbook of educational data mining (pp. 75-92). boca raton, fl: crc press. https://doi.org/10.1201%2fb10274-8 amershi, s., & conati, c. (2009). combining unsupervised and supervised classification to build user models for exploratory learning environments. jedm| journal of educational data mining, 1(1), 18-71. berland, m., martin, t., benton, t., petrick smith, c., & davis, d. (2013). using learning analytics to understand the learning pathways of novice programmers. journal of the learning sciences, 22(4), 564-599. https://doi.org/10.1080%2f10508406.2013.836655 blikstein, p. (2013, april). multimodal learning analytics. in proceedings of the third international conference on learning analytics and knowledge (pp. 102-106). acm. https://doi.org/10.1145%2f2460296.2460316 block, f., hammerman, j., horn, m., spiegel, a., christiansen, j., phillips, b., ... & shen, c. (2015, april). fluid grouping: quantifying group engagement around interactive tabletop exhibits in the wild. in proceedings of the 33rd annual acm conference on human factors in computing systems (pp. 867-876). acm. https://doi.org/10.1145%2f2702123.2702231 bowers, a. j. (2010). analyzing the longitudinal k-12 grading histories of entire cohorts of students: grades, data-driven decision making, dropping out and hierarchical cluster analysis. practical assessment, research, and evaluation, 15(1), 7. davis, p., horn, m. s., schrementi, l., block, f., phillips, b., evans, e. m., ... & shen, c. (2013). going deep: supporting collaborative exploration of evolution in natural history museums. in proceedings of 10th international conference on computer supported collaborative learning . evans, a. c., wobbrock, j. o., & davis, k. (2016, february). modeling collaboration patterns on an interactive tabletop in a classroom setting. in proceedings of the 19th acm conference on computer-supported cooperative work & social computing (pp. 860-871). lee, j. e., recker, m., bowers, a., & yuan, m. (2016, june). hierarchical cluster analysis heatmaps and pattern analysis: an approach for visualizing learning management system interaction data. in edm (pp. 603-604). long, d., mcklin, t., weisling, a., martin, w., guthrie, h., & magerko, b. (2019). trajectories of physical engagement and expression in a co-creative museum installation. in proceedings of the 2019 on creativity and cognition (pp. 246-257). https://doi.org/10.1145%2f3325480.3325505 lyons, l. (2014). exhibiting data: using body-as-interface designs to engage visitors with data visualizations. in learning technologies and the body (pp. 197-212). routledge. lyons, l., tissenbaum, m., berland, m., eydt, r., wielgus, l., & mechtley, a. (2015, june). designing visible engineering: supporting tinkering performances in museums. in proceedings of the 14th international conference on interaction design and children (pp. 49-58). https://doi.org/10.1145%2f2771839.2771845 martin, k., horn, m., & wilenksy, u. (2018). ant adaptation: a complex interactive multitouch game about ants designed for museums. in constructionism conference. martínez maldonado, r., kay, j., & yacef, k. (2010, november). collaborative concept mapping at the tabletop. in acm international conference on interactive tabletops and surfaces (pp. 207-210). acm. https://doi.org/10.1145%2f1936652.1936690 roberts, j., banerjee, a., hong, a., mcgee, s., horn, m., & matcuk, m. (2018, april). digital exhibit labels in museums: promoting visitor engagement with cultural artifacts. in proceedings of the 2018 chi conference on human factors in computing systems (pp. 1-12). https://doi.org/10.1145%2f3173574.3174197 roberts, j., & lyons, l. (2017). the value of learning talk: applying a novel dialogue scoring method to inform interaction design in an open-ended, embodied museum exhibit. international journal of computer-supported collaborative learning , 12(4), 343-376. https://doi.org/10.1007%2fs11412-017-9262-x romesburg, h. c. (1984). cluster analysis for researchers. lifetime learning publications. snibbe, a. c. (2006). drowning in data.stanford social innovation review, 4(3), 39-45. tissenbaum, m., berland, m., & lyons, l. (2017). dclm framework: understanding collaboration in open-ended tabletop learning environments. international journal of computer-supported collaborative learning, 12 (1), 35-64. https://doi.org/10.1007%2fs11412-017-9249-7 tissenbaum, m., kumar, v., & berland, m. (2016). modeling visitor behavior in a game-based engineering museum exhibit with hidden markov models. international educational data mining society. upton, k., & kay, j. (2009, june). narcissus: group and individual models to support small group work. in international conference on user modeling, adaptation, and personalization (pp. 54-65). springer, berlin, heidelberg. https://doi.org/10.1007%2f978-3-642-02247-0_8 wilkinson, l., & friendly, m. (2009). the history of the cluster heat map. the american statistician, 63(2), 179-184. https://doi.org/10.1198%2ftas.2009.0033 yoon, s. a., elinich, k., wang, j., steinmeier, c., & tucker, s. (2012). using augmented reality and knowledge-building scaffolds to improve learning in a science museum. international journal of computer-supported collaborative learning, 7 (4), 519–541. https://doi.org/10.1007%2fs11412-012-9156-x andressen et a l publication frontline learning research vol. 7 no 3 (2019) 1 – 26 issn 2295-3159 processing and learning from multiple sources: a comparative case study of students with dyslexia working in a multiple source multimedia context anette andresena, øistein anmarkrud a ladislao salmerónb, ivar bråtena auniversity of oslo, norway buniversity of valencia, spain article received 22 january / revised 2 may/ accepted 2 june / available online 16 july abstract this study investigated how four 10th-grade students with dyslexia processed and integrated information across web pages and representations when learning in a multiple source multimedia context. eye movement data showed that participants’ processing of the materials varied with respect to their initial exploration of the web pages, their overall processing time, and the linearity of their processing patterns, with post-learning interviews indicating the deliberate, strategic considerations underlying each participant’s processing pattern. eye movement data in terms of fixation duration and percentage of regressions also corroborated the findings of formal, diagnostic assessments. finally, it was found that participants differed with respect to how much factual information they learned from working with the materials and how well they were able to integrate information across the web pages and representations, with results suggesting particular problems with learning factual information and, at the same time, constructing a coherent mental representation of the issue, as well as with drawing on textual information in the integration process. this study brings together two research areas that essentially have been kept apart in theory and research, that is, dyslexia and multimedia learning, and it provides unique information about the role of individual differences in multiple source multimedia contexts. keywords: multiple source use; dyslexia; eye-tracking; strategic processing; multimedia learning corresponding author: oistein.anmarkrud@isp.uio.no doi: 10.14786/flr.v7i3.451 1. introduction due to technological developments, human learning is becoming increasingly multi-representational (ainsworth, 2018). accordingly, learning in school has become much more than reading and understanding textbooks. one reading context where students encounter multimedia information on a regular basis is the internet, making it possible for students to benefit from texts, pictures, animations, films, and interactive graphs. hence, the internet has become an invaluable learning tool for students, providing them with a vast amount of multimedia information that they can use for academic purposes (e.g., kammerer, meier, & stahl, 2016; kingsley & tancock, 2013; mason, junyent, & tornatora, 2014; van strien, brand-gruwel, & boshuizen, 2014). however, although the abundant information available just a finger swipe or mouse click away has brought new affordances for learning, it also comes with some caveats. successful learning from the internet requires learners to integrate task-relevant and reliable information across different representations (e.g., pictures, videos, and texts), web pages, and perspectives, as well as with their own prior knowledge (e.g., bråten, braasch, & salmerón, in press; cho, woodward, & li, 2017; deschryver, 2015; rouet & britt, 2014). presumably, learning from multimedia materials, that is, the construction of a coherent mental representation based on information from different types of media, requires considerable working memory resources (e.g., irrazabal, saux, & burin, 2016; schüler, scheiter, & van genuchten, 2011; sweller, ayres, & kalyuga, 2011). in explaining the relationship between multimedia learning and working memory, mayer’s (2003, 2014a) influential cognitive theory of multimedia learning draws on limited-capacity (baddeley, 1995, 2000; just & carpenter, 1992) and dual-channel theories (baddeley, 1995; clark & paivio, 1991; paivio, 1971, 1986). the limited-capacity theory assumes that human working memory is limited in capacity and, thus, can process only a certain amount of information at a time. if processing demands exceed this limited capacity, the likely result is cognitive overload, which reduces learning. the dual-channel theory assumes that visual and auditory information is processed in separate channels in working memory and that these channels operate independently from each other, each with its own capacity. hence, when these two channels are combined, such as when a multimedia learning context involves both a text (visual channel) and a narration (auditive channel), the learner will be able to process more information simultaneously than when the learning context involves a combination of text and pictures, since both text and pictures have to be processed in the visual channel. however, current web pages, often containing text, pictures, animations with narration, audio files, and so forth, have the potential to overload both the visual and the auditive channel in working memory (knoop-van campen, segers, & verhoven, 2018; schüler et al., 2011). when students use the internet for educational purposes, they often visit several web pages that may present overlapping, complementary, and conflicting information (cho, afflerbach, & han, 2018; cho et al., 2017; salmerón, strømsø, kammerer, stadtler, & van den broek, 2018), with successful learning demanding integration across different representations (e.g., text, video, and picture) and web pages. although this can represent a great challenge for students regardless of their reading skills, poor readers may be particularly vulnerable in such learning contexts. still, surprisingly little is known about how students with dyslexia handle such multi-representational information (anmarkrud, brante, & andresen, 2018; knoop-van campen et al., 2018; mccarthy & swierenga, 2010), compared to the knowledge that exists about typically developing readers’ multimedia learning (mayer, 2014c). this study is frontline because it uniquely contributes to both multimedia learning and dyslexia, integrating and broadening the research agenda in both areas and providing new insights into what it means to be a struggling reader in the 21st century. one pertinent question is how the combined processing demands imposed by the multimedia context along with the processing demands of reading could affect comprehension and learning from multimodal materials for readers with dyslexia. hence, the present study aimed to provide a detailed description of some of the potential challenges that students with dyslexia may experience when trying to integrate information across representations (i.e., texts, pictures, and videos) and web pages by means of a comparative case study of four adolescents with dyslexia working on a socio-scientific issue in a digital environment. as such, this study draws on theoretical and empirical work on information processing in multimedia learning and attempts to extend that work to the challenges faced by readers with dyslexia in a multimedia learning context. 1.1 developmental dyslexia dyslexia is a specific learning disability characterized by difficulties with accurate and/or fluent word recognition, poor decoding skills, and spelling difficulties (lyon, shaywitz, & shaywitz, 2003). the associated reading and writing difficulties are typically a result of a deficit in the phonological component of language (e.g., harm & seidenberg, 1999; ramus et al., 2003). dyslexia is found to affect between 3 and 7% of the population (hulme & snowling, 2009) and is assumed to be of neurobiological origin (e.g., shaywitz & shaywitz, 2008). although there is no “cure” for dyslexia, several studies have shown that many children with dyslexia can develop reading skills comparable to typically developing readers when provided necessary support and high-quality remedial reading instruction (e.g., hulme & snowling, 2009; torgersen, 2001; torgersen et al., 2001). however, there are substantial differences in reading skills among students with dyslexia, and those towards the severe end of the spectrum may struggle with reading long into adolescence and adulthood. further, individuals with dyslexia have often been found to have working memory problems (e.g., avons & hanna, 1995; barbosa, miranda, santos, & bueno, 2009; melby-lervåg, lyster, & hulme, 2012), not only with the processing of information in a phonological code but also with central executive domains and the processing of visual information (e.g., fischbach, könen, rietz, & hasselhorn, 2014; menghini, finzi, carlesimo, & vicari, 2011; smith-spark & fisk, 2007). it is assumed that the working memory deficits often seen among individuals with dyslexia may contribute to difficulties constructing a coherent representation of a text during reading, independent of their difficulties with phonological coding (e.g., berninger, raskind, richards, abbot, & stock, 2008; borella, carretti, & pelegrina, 2010; follmer, 2018; smith-spark & fisk, 2007). however, several studies have found substantial within-group-heterogeneity in working memory capacity among students with dyslexia (e.g., gathercole, alloway, willis, & adams, 2006; jeffries & everatt, 2004; smith-spark & fisk, 2007). 1.2 learning from multiple representations in a digital context up to the early 2000s, research examining reading comprehension and text-based learning typically involved a reader encountering a single text, usually on paper. with the development of new and user-friendly information technologies, such as the internet, the conception of what constitutes a typical reading situation has changed considerably (bråten et al., in press; cho et al., 2018; fox & alexander, 2017; leu, kiili, & forzani, 2016; salmerón et al., in press). whereas conventional printed textbooks typically include text and various forms of illustrations and as such can be labelled multimedia materials, the digital learning contexts of today provides a variety of representations (e.g., animations, videos, audio files, simulations) in addition to text and pictures. research on multimedia learning originally focused on learning from a combination of text and pictures in an offline context (butcher, 2014; mayer, 2014b). however, given the technological developments in recent decades, multimedia learning has come to refer to the combination of any type of words and visual displays, regardless of whether the learning occurs in a non-digital or digital context. several cognitive models have been developed to explain multimedia learning, with mayer’s (2001, 2014b) cognitive theory of multimedia learning being the most influential. the basic assumption of this model is that multimedia learning rests on a cognitive system with multiple memory stores, with a working memory system of limited capacity considered an essential processing component. also, the model posits that good multimedia learning requires the integration of information from various representations and that comprehension and learning can be hampered by the constraints of the human cognitive system, particularly working memory (e.g., chan & unsworth, 2011; mayer & moreno, 2010; schüler et al., 2011). a significant body of research indicates that using multiple representations in academic learning contexts may be beneficial (e.g., butcher, 2014; cuevas, fiore, & oser, 2002; rieber, tzeng, & tribble, 2004). however, poorly designed multimedia learning environments can increase working memory load, leading to reduced learning. one example is when an additional representation (e.g., a picture added to a text) does not contain new information. in such a case, learners will have to waste additional processing capacity on the redundant information from the added representation without gaining any new knowledge. this is often referred to as the redundancy effect and has been found to interfere with learning (e.g., gerjets, scheiter, opfermann, hesse, & eysink, 2009; pociask & morrison, 2008; torcasio & sweller, 2010). another example is that presenting information by means of more than two representations in and of itself can increase processing demands, particularly when the representations are physically or temporally disparate (ayres & sweller, 2014). in such a situation, learners would have to split their attention between the different representations and fill the “gaps” between representations by drawing inferences before integrating information across the representations. this is referred to as the split-attention effect, and it can potentially reduce learning, especially among learners with reduced working memory capacity (fenesi, kramer, & kim, 2016). in brief, the use of multimedia has the potential to increase learning when multimedia environments are designed according to the limitations of the human information processing system. however, ill-structured multimedia environments have the potential to reduce learning, compared to single-media environments, due to increased load on working memory. 1.3 dyslexia and learning from multiple representations thus far, few studies have been conducted on multimedia learning among students with dyslexia. in a recent study, knoop-van campen et al. (2018) examined how multiple representations affected learning and study time among 11-year-old students with dyslexia compared to typically developing readers. the participants worked with three user-paced multimedia lessons on the topics of balance in nature, motion, and global warming, and participants were divided into three conditions: 1) pictures + text, 2) pictures + audio, and 3) pictures + text + audio. the participants studied every lesson once, and learning was assessed at two time points; at one immediate posttest and a delayed posttest one week later. the results showed that in the picture + text condition, the dyslexic participants spent statistically significantly longer time working with the materials than did the non-dyslexic participants. there were no statistically significant differences in study time between the two groups in the two other conditions, and there was no statistically significant correlation between study time and learning in any of the conditions. further, there were no main effects of condition, group, or working memory capacity on learning measured at the immediate and delayed posttests, although participants with dyslexia had statistically significantly lower scores on the working memory measure compared to the participants without dyslexia. a study by maccullagh, bosanquet, and badcocks (2017) highlights the challenges students with dyslexia may have integrating information across representations. interviews were conducted with 13 university students with dyslexia to investigate how they used an available multimedia tool consisting of video-recorded lectures in combination with other representations such as animations and text boxes to compensate for their reading difficulties. several participants reported that recorded lectures that they viewed with this tool were challenging to follow when all the different representations were combined. the participants reported that they often had to go through the online lectures several times to be able to benefit from them (e.g., just listen to the lecturer the first time, pay attention to the animation and text boxes the second time, and take notes the third time). in two studies, alty and colleagues (alty, al-sharrah, & beacham, 2006; beacham & alty, 2006) examined the effects that different combinations of representations, such as textual and visual materials (e.g., diagrams) as well as audio files (voice over), had on the learning of statistics among students with dyslexia. in the first study, alty et al. (2006) compared students with and without dyslexia across three conditions: one group received the learning materials as text only, one group received the materials as diagrams + voice over, and one group received the materials as text + diagrams. the results showed, contrary to expectations, that the students with dyslexia in the text-only condition significantly outperformed the students with dyslexia in the two other conditions with regard to learning, whereas the students without dyslexia in the diagrams + voice-over condition performed better than the students without dyslexia in the two other conditions. given the somewhat surprising finding concerning the participants with dyslexia, beacham and alty (2006) conducted the same experiment with a larger sample of students with dyslexia, this time without a non-dyslexic control group. the results in this second study corroborated the findings from the original study; again, the participants in the text-only condition outperformed the students in the two other conditions regarding learning, despite reporting that this was the least preferred version of the learning materials. 1.4 integration as strategic activity according to cho and afflerbach (2017), integration of information across representations and web pages requires strategic activity in the service of creating meaning. more generally, comprehension strategies may be defined as intentional attempts to control and modify meaning construction during learning (cf., afflerbach, pearson, & paris, 2008). if the semantic overlap between representations or sources is high, integration may rely on automatic processing (myers & o'brien, 1998). if not, integration will have to rely on deliberate, strategic activity (kurby, britt, & magliano, 2005), with the execution and monitoring of strategies drawing on working memory resources. presumably, multimedia learning that requires strategic processing will be particularly challenging for learners with dyslexia, who also must spend considerable working memory resources on more basic reading processes. of note is also that multimedia materials, such as online sources, are not necessarily designed according to multimedia principles (e.g., mayer, 2014b), that is, to reduce the load on working memory. several studies clearly have indicated that strategic behavior, such as coordinating representations and actively searching for meaning in different sources, can facilitate the integration of information into a coherent and rich mental model when working with multimedia materials (e.g., azevedo & cromley, 2004; greene, moos, azevedo, & winters, 2008; moreno & mayer, 2000). 1.5 the present study although it has been argued that learning in multimedia contexts, such as the internet, could be beneficial for struggling readers because text is supplemented with other representations (castek et al., 2011; henry et al., 2012), learning in such contexts also may represent particular challenges for struggling readers. this is because a combination of processing demands associated with word reading and processing demands associated with integrating information across web pages and representations may lead to cognitive overload (chan & unsworth, 2011). in this study, we extended previous research by exploring variations in the processing patterns of adolescent readers with dyslexia who worked with conflicting web pages containing multiple representations. additionally, we set out to explore how these processing patterns were related to cognitive differences among participants and to their performance on post-reading learning and integration tasks. to be able to address these issues in depth, we opted for a comparative case study design (yin, 2009) combining quantitative and qualitative data. specifically, following the logic of a comparative case approach (campbell, 2012; yin, 2009), we selected four cases that were analysed independently before they were compared and contrasted. the study of these cases was guided by two research questions. first, to what extent do different processing patterns displayed by students with dyslexia when reading multimodal information represent deliberate, strategic activity? we expected that processing time would be related to the severity of the participants’ reading difficulties, with more severe reading difficulties associated with longer processing time due to the time spent on reading the texts on the web pages. moreover, prior research with typically developing readers in multimedia learning contexts has shown that the order in which representations are processed, and the transitions between representations, can influence learning outcomes (e.g., mason, pluchino, & tornatora, 2016; mason, scheiter, & tornatora, 2017). two different approaches to the integration of textual and pictorial information in multimedia contexts have been described in the literature. the first approach, picture-to-text processing, involves a brief inspection of pictures before processing textual material, with a quick examination of a picture providing a global spatial representation of the topic that, in turn, can scaffold comprehension of the text material. recent research has found a positive effect of the picture-to-text approach (eitel, scheitel, & schüler, 2013; eitel, scheiter, schüler, nyström, & holmqvist, 2013; mason et al., 2017). the second approach, text-to-picture processing, involves processing textual material first, which may help readers focus on the essential elements of a picture subsequently. this approach also has been found to facilitate learning in multimedia contexts (hegarty & just, 1993). which of these approaches is the most efficient is probably a matter of the complexity of the information conveyed by the different representations, with the representation containing the least complex information preferably processed first (eitel & scheiter, 2015; mason et al., 2017). considering that our participants were students with dyslexia, we expected that participants with a deliberate, strategic processing pattern would use the picture-to-text approach to try to compensate for their word reading problems. second, how are processing patterns related to individual differences among participants and to their learning from and integration of multimodal information presented on different web pages? several studies have indicated that students with dyslexia who have developed sufficient word decoding skills through adequate remedial reading instruction may display reading comprehension almost on a par with students without dyslexia (e.g., bishop & snowling, 2004; de olivera, da silva, dias, sebra, & macedo, 2014; torgersen, 2001; torgersen et al., 2001). although there are very few studies examining comprehension of graphics (both static and motion) among individuals with dyslexia, available studies do not indicate that they have particular problems extracting information from pictorial representations (abtahi, 2012; roca, tejero, & insa, 2018; taylor, duffy, & hughes, 2007). hence, we expected that the participants would be able to gain factual knowledge from all representations and web pages used in this study. however, because integration of information from multiple representations imposes considerable processing demands on working memory, and because students with dyslexia have been found to display working memory deficits, we expected that integrating information across representations and web pages would be a profound challenge for our participants. further, we expected that the participants would be inclined to draw more on information from pictures and videos than on information conveyed by textual material when trying to construct a coherent mental representation of the learning materials. this research is based on a sample of 22 tenth-graders with dyslexia who participated in a study investigating differences in multiple source use between students with and without dyslexia (andresen, anmarkrud, & bråten, 2019). that study indicated that the group of students with dyslexia was heterogeneous with respect to working memory capacity and reading skill. hence, to examine in depth how differences in these two key competencies might influence the processing of and learning from multiple multimedia sources in a digital environment, we selected participants with dyslexia who varied with respect to working memory capacity and reading skill for this comparative case study. 2. method 2.1 participants the participants were four norwegian adolescents, ranging in age from 15 years 9 months to 15 years 11 months. in norway, if there is concern about a student’s reading proficiency, the student will be assessed by an educational-psychological service (eps). if appropriate, the students will be diagnosed with dyslexia based on test results, classroom observations, and interviews with parents and teachers. students who receive remedial reading instruction according to the special needs education act are usually reassessed every other year. thus, the four participants in this study were diagnosed by experts at the eps within the last two years. their diagnoses were based on criteria included in the definition of dyslexia proposed by lyon et al. (2003). this means that the four participants displayed difficulties in word recognition, phonological processing, and spelling (lyon et al., 2003). specifically, all participants were assessed with standardized diagnostic test batteries called logos (høien, 2014) or stas (klinkenberg & skaar, 2003), which are frequently used in norway and other scandinavian countries to diagnose dyslexia. on these test batteries, all participants with dyslexia scored below the 15th percentile on subtests measuring reading fluency, word identification, phonological processing, and spelling and were simultaneously within the normal range on subtests measuring listening comprehension. none of the participants had comorbid conditions such as attention deficit disorder, language impairment, or more general learning disabilities. all participants had normal or corrected to normal vision. see table 1 for relevant background information about each of the four participants. table 1 background information about the four participants note measured with logos (høien, 2014), 2measured with stas (klinkenberg & skaar, 2003), 3measured with a norwegian adaption of swanson and trahan’s (1992) working memory span task, 4compared to the mean scores of a sample of 528 norwegian 10th and 11th graders (anmarkrud & ferguson, 2011). participant 1 was diagnosed with dyslexia in 4th grade. the latest eps assessment showed that this participant still had substantial reading difficulties, with very low scores on subtests measuring phonological word reading (nonword reading; 1.2 percentile), orthographic reading (0.1 percentile), and reading fluency (0.2 percentile). based on the thorough reading assessment at the eps, participant 1 was by far the weakest reader among the four participants. as seen in table 1, participant 1 also displayed a very limited working memory capacity. participant 2 was diagnosed with dyslexia in 6th grade. on the latest eps assessment, participant 2 scored approximately two grades below the current grade level on subtests measuring phonological reading (nonword reading), orthographic reading, and reading fluency. however, presumably due to very good language comprehension skills, participant 2 seemed to be able to compensate for the word reading problems and had grade-appropriate reading comprehension scores. hence, participant 2 could be characterized as a poor word decoder with relatively good comprehension skills. participant 3 was not diagnosed with dyslexia until 8th grade. in the report from the eps, this participant was described as a very motivated and academically sound student with excellent learning abilities (e.g., a relatively high working memory capacity). due to these strengths, participant 3 had been able to conceal and compensate for the word reading problems for many years, and it was not until entering 8th grade, where the reading materials in school became increasingly more complex, that the reading difficulties became very visible. the eps assessment showed that participant 3´s scores on tests measuring phonological reading (nonword reading), orthographic reading, and reading fluency were equivalent to what is typically found among students two years younger. participant 4 was diagnosed with dyslexia in 5th grade. the results on the latest eps assessment showed that this participant mastered a phonological word reading task (i.e., could read the nonwords correctly) but was very slow on this task compared to typically developing peers. participant 4 demonstrated substantial difficulties on subtests measuring phonological awareness, orthographic reading, and reading fluency, with scores well below the 15th percentile on these subtests. table 2 content of the three web pages 2.2 learning materials participants were given access to a researcher-generated internet site presented in an offline mode that was titled “sunbathing and health”. the site contained three different web pages about the controversial issue of sun exposure and health. these web pages presented two main perspectives: sun exposure is beneficial, and sun exposure is harmful. each web page contained a title and a lead paragraph explaining the overall content of that page and then presented a video, a short text, and a picture, in that order. the first page contained information about the nature of ultraviolet radiation, different wavelength bands, how ultraviolet radiation is measured, and how different types of ultraviolet radiation affect the skin. the second page presented research arguing that sun exposure is healthy because it increases the production of vitamin d, which can protect against cancer, particularly in inner organs. the third page focused on the harmful effects of sun exposure due to increased risk of skin cancer, particularly basal cell carcinoma and melanoma, and explained that sun exposure cannot be considered a safe source of vitamin d. the main idea units of the different representations (i.e., the texts, the videos, and the pictures) were unique, which made it possible to trace each idea unit in participants’ post-reading answers back to a particular representation on a particular web page. of note is that the learning materials were designed in accordance with design principles for multimedia learning (mayer, 2014a). thus, we took the spatial contiguity principle (e.g., austin, 2009; johnson & mayer, 2012) into consideration by presenting text, videos, and pictures near each other on the web pages, and we omitted redundant information across the different representations in accordance with the redundancy principle (e.g., mayer, heiser, & lonn, 2001; moreno & mayer, 2002). table 2 provides an overview of the content of the three web pages. the texts that were included on the web pages (one on each page) contained 83, 92, and 90 words, and ranged in readability from 37 to 40 (see table 3). these readability scores were based on björnsson’s (1968) formula, taking word length and sentence length into consideration. this formula yields readability scores ranging from approximately 20 (very easy text) to approximately 60 (very difficult text). vinje (1982) reported that textbooks used in norwegian upper-secondary school had a readability score of approximately 42 and that public information texts from the norwegian government had a readability of 45. table 3 descriptive information about the text on each of the three web pages 2.3 measures 2.3.1 topic knowledge measure. to assess students’ knowledge about the topic of sun exposure and health, both before and after working on the learning materials described above, we used a 12-item multiple-choice test. the items referred to concepts and information central to the issue of sun exposure and health that were discussed on the three web pages. because the same measure was administered both before and after participants worked on the learning materials, learning gain could be calculated by subtracting the number of correct responses out of 12 on the first occasion from the number of correct responses out of 12 on the second occasion. a preliminary version of the topic knowledge measure was reviewed by a professor of medical biochemistry at the university of oslo who was not part of the project, which resulted in only minor modifications to the response alternatives of a few items. sample items from the topic knowledge measure are displayed in appendix a. in the larger sample of students with dyslexia from which the four participants were selected, the internal consistency reliability (kuder richardson 20) for scores on this measure was .62. 2.3.2. working memory measure. working memory was measured using a norwegian adaptation of swanson and trahan’s (1992) working memory span task (braasch, bråten, strømsø, & anmarkrud, 2014). this measure is derived from daneman and carpenter’s (1980) original reading span test. twelve sets of unrelated sentences were read aloud with a 2-second interval between each sentence. the sets gradually increased from two to five sentences. participants were tasked to simultaneously a) answer a comprehension question about an unknown sentence after the final sentence was read, and b) remember the final words from each of the sentences. for each of the 12 trials, participants were awarded 1 point if they correctly answered the comprehension question and one additional point for each of the final words they recalled. if participants failed to answer the comprehension question correctly, they did not receive any points for that set regardless of how many final words they recalled. internal consistency reliability (cronbach’s α) for scores on this measure in the larger sample of students with dyslexia from which the four participants were selected was .73. 2.3.3 multiple source integration task. multiple source integration was assessed by asking the four participants to respond orally to two open-ended questions modelled on the integrative short essay tasks used by rukavina and daneman (1996) to measure students’ understanding of a controversial scientific issue. of note is that this approach also has been used effectively in several previous studies of multiple source integration (e.g., barzilai & ka’adan, 2017; bråten, anmarkrud, brandmo, & strømsø, 2014; ferguson & bråten, 2013).the first question was, “could you explain the relationship between sun exposure, health, and illness?” the second question was, “could more than one view on the relationship between sun exposure, health, and illness be correct? yes or no? if yes, why? if no, why not?” following rukavina and daneman (1996), we considered our first question to indirectly require participants to integrate different perspectives across web pages and representations, or, at least, to consider each perspective’s claims and explanations. our second question was considered to directly require participants to pit perspectives against each other, measuring how well they could reason about the issue in terms of the claims and explanations presented across web pages and representations. the oral responses were audio-taped and transcribed before they were scored. following andresen et al. (2019), the responses were scored in three steps. in the first step, we coded responses to both questions based on the extent to which participants integrated the two main perspectives represented in the materials (i.e., sun exposure is healthy vs. sun exposure is harmful), regardless of the web pages and representations they drew upon in their responses. on the indirect integrative question, participants could obtain scores between 0 and 5. a score of 0 was given for no response or for irrelevant information. a score of 5 was given for mentioning the two main perspectives and providing elaborate explanations or reasons for both perspectives as well as relating the two perspectives to each other by comparing and/or contrasting them and trying to reconcile them. inter-rater reliability was established in the larger sample from which the four participants were selected. in this process, the first and second authors independently scored a random selection of 50% of participant responses to the first question, initially agreeing on 80% and resolving all disagreements through discussion. on the direct integrative question, we first coded whether participants recognized that the main perspectives were not mutually exclusive and might be reconciled (i.e., whether participants answered “yes” or “no” to the question). second, we coded to what extent participants could explain and reconcile the two perspectives (i.e., when they answered “yes”) and to what extent they could select one of the perspectives and provide explanation or reason for that perspective (i.e., when they answered “no”). again, scores could range from 0 to 5. a score of 0 was given when participants answered “no” to the question without providing any further justification for their answer. a score of 5 was given when participants answered “yes” to the question, mentioned the two perspectives, provided elaborate explanations or reasons for both, and related the two perspectives to each other by explaining how they may be reconciled. again, the reliability of the coding was established in the larger sample from which the participants were selected. the first and second authors independently scored a random selection of 50% of the answers to the second question, initially agreeing on 83% and resolving all disagreements through discussion. participants’ scores on the two integrative questions were collapsed, which means that their scores after this step could range from 0 to 10. table 4 presents the entire coding system used for scoring the oral responses in the first step. table 4 coding system used in the first step of the scoring of the oral responses in the second step, we assessed the extent to which participants drew on information from the three different web pages and the different types of representations on each web page (i.e., text, video, and picture) when constructing their oral responses to the two questions. because the main idea units of each representation were unique, we could trace an idea unit included in an oral response back to a particular representation on a particular web page. inter-rater reliability of this coding also was established in the larger sample by the first and second authors who independently coded a random selection of 50% of the responses to both questions and initially agreed on the origin of 85% of the idea units. all disagreements were resolved through discussion. in the second step, participants could obtain scores between 1 and 2.98. in addition to a constant of 1, participants were awarded a score of 0.33 for each web page and a score of 0.11 for each representation (i.e., text, video, or picture) that they used in their responses. for example, a participant who included idea units from the text and the video on the first web page would obtain a score of 0.33 for the web page and a score of 0.11 each for the text and the video, so this participant’s score would be 1.55 (including the constant). if this participant additionally drew on the video on the second web page, he or she would obtain a score of 0.33 for that web page and 0.11 for that video, resulting in a score of 1.99 (including the constant). we awarded a score of 0.33 for each web page and a score of 0.11 for each representation because we considered the entire web site to consist of three parts (i.e., web pages), which were again divided into three representations. in the third step, we computed each participant’s total multiple source integration score by multiplying the participant’s score from the first and second steps. in this way, we considered both the integration of the main perspectives (step one) and the coverage of the learning materials (step two) when assessing multiple source integration. thus, on the multiple source integration task, participants could obtain a maximum score of 29.8 (i.e., 10 x 2.98). the reason we added a constant of 1 to each participant’s score in the second step was to avoid any participant obtaining a total multiple source integration score that was lower than the score obtained in step one. 2.3.4 apparatus and analysis of eye-tracking data. when working with the learning materials, gaze data were collected by means of a tobii x2-60 eye-tracking device. the tobii x2-60 is a screen-based eye tracker that records gaze data at a sampling rate of 60 hz. data collection took place in a quiet room, and direct sunlight to the screen and the student’s head was restricted to avoid distorting reflections. students sat approximately 60 cm from the screen, which was a t540p lenovo laptop with a 15.6” monitor and 1920 x 1080 resolution. a nine-point calibration task was used to ensure reliable eye-tracking. the process was repeated until the average deviation dropped below 0.5º. eye-tracking data were analyzed with two different approaches. first, participants’ patterns of processing were examined with regard to sequence of processing (linear vs. nonlinear processing patterns) and time spent on the various web pages and representations. second, we performed finer grained analyzes of eye movements during the reading of the texts on the three web pages. eye-tracking data were analyzed by means of tobii studio software. specifically, we established text paragraphs as the main areas of interest (aoi). within these, we examined the number of fixations, average duration of fixations (measured in milliseconds), and percentage of regressions (saccadic movements to the left, excluding carriage returns). the first time participants read a paragraph was defined as first-pass reading, while any subsequent reread was considered second-pass reading. in the reading literature, higher average fixation duration and a high percentage of regressions are considered indicators of comprehension difficulties (rayner, chace, slattery, & ashby, 2006; schotter, tran, & rayner, 2014). 2.3.5 follow-up interview. finally, when participants had finished working with the learning materials and responded to the post-reading measures, we replayed the recordings of their eye movements and used these recordings as stimuli to gain insight into whether there were any deliberate reasons for their processing patterns when working with the learning materials. the main topic of these follow-up interviews was the order in which the different web pages and representations where processed and whether the observed processing pattern was representative of how they would typically approach a web page in an academic setting. the follow-up interviews were transcribed, and the various utterances were given a time stamp making it possible to connect an utterance to an incident in the recordings of the eye movements. 2.4 procedure data collection took place in two different sessions, approximately two weeks apart. the first author collected all data in participants’ home schools either during the school day or directly afterwards. all instructions, questionnaire items, and questions were read aloud to the participants while they had their own printed copies in front of them. this procedure was followed to reduce the effects of participants’ reading difficulties on the various measures. in the first session, the topic knowledge and working memory measures were individually administered in that order. this session lasted approximately 40 minutes. in the second session, participants individually studied the three web pages on a laptop (see above) with the following instruction read aloud: “sun exposure and health is a topic of current interest. imagine that you are supposed to hold an oral presentation on this topic for your fellow students. here are three web pages you can use to prepare the presentation. you may have 30 minutes studying these three pages; please do not take any notes. you can move between the pages as much as you would like and read them in the order that you choose”. when finished working on the three web pages, the participants responded to the topic knowledge measure for a second time before answering the oral multiple source integration task. finally, the follow-up interview was conducted. participants were given a gift certificate of 300 nok (approximately 35 us $) and were offered an individual course in strategic internet reading as a reward for their participation. this study was carried out in accordance with the recommendations of the norwegian national research ethics committees. the protocol was approved by the norwegian centre for research data. in norway, the norwegian centre for research data functions as a national ethics committee approving all studies within the social sciences. all subjects as well as their parents gave written informed consent for participation in the study. 3. results to address the first research question regarding possible variations in the processing patterns, we used eye tracking to examine their processing of the web pages and representations. we also performed a more detailed analysis of eye movements during text reading, focusing on differences among participants regarding number of fixations, fixation duration, and regressions. figure 1 displays the processing pattern of the four participants when working with the learning materials in the second session. one difference that could be observed concerned how they initially explored the web pages. participant 3 started the session by using 20 seconds to quickly click through all three pages, before returning to the first page and going through the three web pages more thoroughly. participant 4, on the other hand, went straight to page one, going through this page before moving to page two and page three. on each web page, both participant 1 and participant 2 started with the text before moving on to the video and the picture, in that order. a second difference was overall processing time, with participant 1 spending a total of 23 minutes and the other three spending between 6 minutes and 30 seconds and 7 minutes and 45 seconds on the three web pages. interestingly, the time participant 1 spent on reading the texts on the three web pages (15 minutes) constituted the difference in total processing time between participant 1 and the other participants. a third observable difference between the participants was the sequence of processing the web pages, specifically the degree of linearity. the processing patterns of participant 3 and participant 4 can be categorized as linear in the sense that they both processed the various representations in the order in which they appeared on the web pages (i.e., video, text, and picture). although the processing patterns of these two students may look identical on a surface level, there were important differences on a more detailed level. thus, participant 3 always started out by reading the lead paragraph explaining the content of the page before going through each web page in a linear pattern. participant 4 skipped the lead paragraph on all pages and went directly to the video before examining the text and picture in a linear pattern on every page. in the follow-up interview, participant 3 said that this quick examination of the three web pages was a deliberate strategy used to get an overview of the content of the web pages. when asked about the processing order of the representations in the follow-up interview, both participant 3 and participant 4 explained that they believed that those who make web pages probably have a reason for the design of a page; therefore, they just processed the representations in the order in which they appeared on each web page. figure 1. processing patterns of the four participants on the three web pages the processing patterns of participant 1 and participant 2 can be described as nonlinear because they did not process the representations in the order in which they appeared on the web pages. a noteworthy difference between participant 1 and participant 2 was that the former systematically went back and reread the text and re-examined the pictures on each web page after the initial processing of the site. although, as seen from the processing patterns, they both prioritized the text (i.e., read the text first), their reasons for this were different. participant 1 explained in the follow-up interview that reading the text first was done deliberate to get some knowledge of the content before watching the video; a strategic activity used to maximize the comprehension of the video. participant 1 applied this strategy on all three web pages. participant 2, on the other hand, expressed that the text was usually the representation that contained the key information on a web page, and, therefore, participant 2 would always start with the text when entering a web page. this participant also gave another strategic reason for such a processing pattern during the follow-up interview. due to the reading difficulties, participant 2 sometimes experienced that “things were unclear” after the reading of a text on a web page, hoping to clarify misunderstandings by watching a video or looking at pictures if such representations were available on a web page. regarding the second research question, we examined whether processing patterns were related to differences among participants with respect to working memory capacity, reading skills, and topic knowledge, as well as to their post-reading performances on the topic knowledge measure and the multiple source integration task. table 5 number of fixations, percentage of regressions, and fixation duration during reading note. 1fixation durations are measured in milliseconds. outlier fixations (defined as participants’ mean fixation + 2 sd) are replaced by participants’ median fixation. the results of the eye-tracking analyses are displayed in table 5, with results corroborating the findings of the eps assessments regarding the reading levels of the participants. each participant’s average fixation duration (see “average forward fixation duration” in table 5) and percentage of regressions across the three web pages were compared with the reading of students with normative development (rayner, ardoin, & binder, 2013) and with a sample of students with dyslexia (prado, dubois, & valdois, 2007). participant 1 read texts at a substantially slower rate (average fixation duration 476 msc) but did not display more regressions (31.59% regressions) than what has been reported in children with dyslexia (325 msc, 31% regressions). participant 1 was also the only participant who reread all three texts. as can be seen in table 5, both the 1st and the 2nd pass were characterized by long fixations and many regressions. participant 4’s reading behavior (average fixation duration 310 msc, 39.11% regressions) was similar to what has been reported for children with dyslexia. finally, participant 2 and participant 3 read the texts substantially faster (with average fixation duration 200 and 190 msc, and with 26.06% and 16.20% regressions, respectively), which closely resembles the behaviors typically reported for children without dyslexia of a similar age (average fixation duration 230-250 msc, 22% regressions) (rayner et al., 2013). however, previous research has found that the transparency of the orthography can influence fixation time during reading; the deeper the orthography, the longer the fixations (bahnmueller, huber, nuerk, göbel, & moeller, 2016; rau, moll, snowling, & landerl, 2015; van roy & pretorius, 2013). norwegian is a transparent orthography compared to english or french, and this should be considered when interpreting the reading rates of participant 2 and participant 3 in relation to the norms, which are established with englishor french-speaking children. 3.1 learning gain and multiple source integration as displayed in table 6, the participants started out with a similar amount of topic knowledge, with participant 1 and participant 3 receiving a score of 4, which is equivalent to a z-score of -1.32 compared to a norm sample of 528 norwegian 10th and 11th graders without dyslexia (anmarkrud & ferguson, 2011). please note that all participants in the norm sample responded to the same topic knowledge measure both before and after studying the same information about sun exposure and health as did the four participants in the current study, and that the pre-reading topic knowledge, post-reading topic knowledge, and learning gain of the participants in the current study were compared to the pre-reading topic knowledge, post-reading topic knowledge, and learning gain obtained by the students in the norm sample. participant 2 and participant 4 scored 5, which is equivalent to a z-score of -.93 when compared to the norm sample. however, there were differences between the participants with respect to learning gain, indicating the amount of factual knowledge they were able to gain from working with the learning materials. participant 1 read at a much slower rate than what is often seen among students with dyslexia but reread all texts thoroughly. this participant had a learning gain of 6, ending up with a post-reading topic knowledge score of 10 (z-score -.14). compared to the norm sample, this is a substantial learning gain, equivalent to a z-score of 1.45. participant 3 and participant 4 both ended up with a post-reading topic knowledge score of 9 (z-score -.72), with their learning gains of 5 and 4 equalling z-scores of .99 and .53, respectively, when compared to the norm sample. hence, the results indicated that these three participants were all able to extract factual knowledge from the representations and web pages included in the learning materials. however, participant 2, who read relatively fast and also displayed relatively low working memory capacity, did not seem to gain much factual knowledge from the learning materials, increasing the topic knowledge score by only one point and ending up with a post-reading topic knowledge score of 6 (z-score -2.48), indicating a relatively small learning gain compared to the norm sample (z-score -.86). on the multiple source integration task, higher scores required that readers integrated information across web pages and representations (i.e., texts, videos, and pictures) into a coherent mental representation of the issue of sun exposure and health. given a potential maximum score of 29.8, participant 1 clearly struggled with this task and obtained a score of only 3.76, which is equivalent to a z-score of -1.76 when compared to a sample of norwegian 10th graders previously responding to this task (authors 1). participant 3 and participant 4 obtained scores of 5.64 and 9.28, equalling z-scores of -1.46 and -.88, respectively. in contrast, participant 2, who gained little factual knowledge, received a multiple source integration score of 16.80 (z-score .35), thus performing above the average of the norm sample. hence, the results suggested that none of the four participants was able to learn factual knowledge from the three web pages (indicated by the post-reading topic knowledge scores) and at the same time construct a coherent mental representation of the issue in question (indicated by the scores on the multiple source integration task). thus, spending cognitive processing capacity on integrating information across different web pages and representations seemed to have left little capacity for learning factual information from the same web pages and representations, and vice versa. table 6 participant scores on the learning and integration measures note. 1the z-scores are based on comparison with the pre-reading topic knowledge, post reading topic knowledge, and learning gain scores of a sample of 528 norwegian 10th and 11th graders (anmarkrud & ferguson, 2011). 2the z-scores are based on comparison with a sample of 44 norwegian 10th graders (andresen et al., 2019). table 7 representations drawn on in the integration task finally, even though the eye-tracking data indicated that the four participants processed all the representations on all the web pages, there was a clear pattern regarding which representations the participants drew on in the multiple source integration task. as displayed in table 7, none of the participants drew on the text material when answering the multiple source integration task; they all based their oral responses on videos and, to some degree, information from the pictures. hence, since all the participants processed the texts, the inability to draw on information from the texts in the integrations task could not be ascribed to a lack of processing of the texts, but rather the inability to integrate the information extracted from the texts with information from other representations across the three web pages. 4. discussion the present study examined how four adolescents with dyslexia processed information from multiple representations in a digital context and how they learned from and integrated information across different representations and web pages. our first research question concerned differences in processing patterns and whether differences in processing patterns represented differences with respect to deliberate strategic activity. while all the participants provided reasons for the strategic approaches they chose when working with the learning materials, there seemed to be differences regarding the sophistication of these strategic approaches. two of the participants, participant 3 and participant 4, processed the various representations in the same order as they appeared on the web pages, with the rationale for this linear processing approach being that those who made web pages probably had good reasons for the order in which the representations appeared. both participant 1 and participant 2, on the other hand, started with the text before processing the graphic representations. previous research has revealed that students with dyslexia may be aware of their reading problems but seem to have a limited strategic repertoire to compensate for these difficulties (furnes & norman, 2015). moreover, studies conducted with traditional paper-based reading have consistently found differences between students with and without reading difficulties regarding knowledge about which strategies to use and when to use them (e.g., baker & beall, 2009; furnes & norman, 2015; roeschl-heils, schneider, & van kraayenoord, 2003). in their review of the literature, anderson and ambruster (1984) reported that, compared to students without reading difficulties, students with reading difficulties to a lesser degree planned their reading ahead, integrated information, and reread text when they noticed comprehension problems. hence, it is interesting that only one of the participants (participant 2) explained a strategic approach by referring to the reading difficulties. this participant thus read the text first, and then examined pictures and videos to clarify misunderstandings that might have arisen while reading due to the reading difficulties. although participant 1 did not verbalize any particular strategic reason for the meticulous reading, and rereading, of the texts on the three web pages, it is conceivable that this approach reflected the experience as a struggling reader and reasoning about what could do to compensate for the difficulties. hence, inconsistent with what we expected, a strategic approach to the learning materials was taken by the two participants who started with text before moving towards the graphic representations, in accordance with a text-to-picture processing approach (hegarty & just, 1993; mason et al., 2017). the prioritization of text, regarding both the time spent with text and the decision to process text first, can also reflect the text superiority effect (e.g., corriveau, einav, robinson, & harris, 2014; einav, robinson, & fox, 2012; eyden, robinson, einav, & jaswal, 2013), which implies that children tend to put more trust and emphasis on information from written information compared to other types of information, especially in an academic context. the approaches of participant 3 and participant 4 can be described as less sophisticated and more passive. simply following the order of the web pages in a linear approach can reflect a type of “outsourcing” of the decisions regarding how to work with the learning materials to those who made the web pages. participants 2 and 3 read through the texts on the web pages quickly, as compared to what has been established as standards for students with dyslexia (prado et al., 2007; rayner et al., 2013). there are several reasons why one should be careful in interpreting this relatively fast reading pace as reflective of good word reading skills. first, the oft-cited standards are based on readers with dyslexia who are younger than those who participated in our study. second, previous research has found that the transparency of the orthography can influence fixation duration during reading. the fact that norwegian is a relatively transparent orthography compared to the orthographies within which the standards have been established could be a reason for the slight mismatch between our reading data and the established standards. there are currently no standards for average fixation duration, number of fixations, or amount of regressions based on reading in norwegian. third, previous research indicates that reading on a screen in and of itself can lead to a faster reading pace than reading on paper (e.g., trakhman, alexander, & berkowitz, in press; van de vijver & harsveld, 1994). our second research question concerned the relationship between processing patterns, individual differences in reading abilities and working memory, and learning and integration in a digital multimedia context. the results showed that three of the participants (1, 3, and 4) had substantial learning gains, also when compared to the learning gains of a norm sample of students without dyslexia. hence, these three participants were able to learn factual knowledge from representations and web pages, and this knowledge was sufficient to answer post-reading questions in a multiple-choice format. however, these participants were to a very limited degree able to integrate information across representations and web pages into a coherent mental representation that reconciled the opposing perspectives covered in the learning materials. participant 2, on the other hand, received a relatively good score on the integration task, also when compared to a norm sample previously responding to this task; however, this participant ended up with a small learning gain, both compared to the other participants in this case study and to the norm sample. as previously described, learning multimedia materials can place a high demand on the capacity-limited working memory system, and multimedia learning therefore is expected to be a challenge for students with dyslexia because they often have working memory problems. one plausible interpretation of our results is that none of the participants were able to learn information on a propositional level, necessary to answer the multiple-choice questions, and at the same time aggregate the main ideas from the representations and web pages into an integrated understanding of the issue. thus, it seemed that for these participants, the interaction with the multimedia materials resulted in an “either – or” learning outcome; either the details or the bigger picture, but not both. presumably, this result is due to the combination of reduced working memory capacity and word reading problems, with reading putting too heavy a load on an already limited working memory and leaving too little working memory capacity to both remember detailed information on a propositional level and integrate information (anmarkrud et al., 2018; hulme & snowling, 2009; melby-lervåg et al., 2012). accordingly, in previous research on paper-based reading, it has been found that both word recognition and comprehension processes compete for working memory capacity (e.g., stanovich, 1986). however, this study does not allow us to draw conclusions about whether it is dyslexia or limited working memory capacity that is causing the integration difficulties. in contrast to our results, as well as what has been found in previous research (e.g., alty et al., 2006; beacham & alty, 2006; maccullagh et al., 2017), knoop-van campen and colleagues (2018) did not find that readers with dyslexia had difficulties integrating information across different representations compared to peers without dyslexia. one likely explanation for inconsistent results in research on multimedia learning among students with dyslexia is the age of participants. the participants in knoop-van campen and colleagues’ (2018) study were 11-year-old children, while the participants in this and other studies (alty et al., 2006; beacham & alty, 2006; maccullagh et al., 2017) have been university students or adolescents in upper-secondary school. among 11-year-old children, including children with dyslexia and their typically developing peers, working memory is not yet fully developed (schneider, 2011). hence, similarity in working memory capacity in the two groups (i.e., students with and without dyslexia) could have resulted in a lack of difference in multimedia learning. the large learning gain of participant 1 also requires some explanation. despite a very limited working memory capacity and severe reading difficulties, participant 1 ended up with the largest learning gain among the four participants, a learning gain that was also well above the mean learning gain of the norm sample consisting of students without dyslexia. knoop-van campen and colleagues (2018) found that self-paced work with learning materials allowed the participants with dyslexia in their study to spend the time necessary to learn the materials properly. thus, those authors found that the participants with dyslexia used significantly more time with the learning materials in the text condition compared to the control group, but this extra time washed out the expected differences in learning. in line with this, participant 1’s slow and meticulous processing and reprocessing of the materials seem to have compensated for the shortcomings regarding working memory capacity and reading abilities compared to the other three participants. finally, although eye-tracking data revealed that all the participants processed all representations on all pages, it is intriguing that none of the participants drew on the text material in their answers on the integration task. they were, however, able to draw on information from the text when answering the multiple-choice questions. this implies that keeping information from the texts in working memory, and at the same time integrating this information with information from the other representations, seems to be a major challenge for readers with dyslexia. it should be noted that we used a portable eye-tracker without a chin rest with a sampling rate of 60 hz, which is a somewhat lower sampling rate than the current standard in reading research. albeit being a limitation of the present study, the relatively large areas of interest (i.e., paragraphs) used in the analyses of eye movements make it unlikely that the sampling rate had a large impact on the quality of the data. 5. conclusion our comparative case study indicates that students with dyslexia may approach conflicting web pages containing various representations in different ways. additionally, while they may process information strategically to improve their learning, they may fail to integrate information across different web pages and representations. to the best of our knowledge, this is the first study of dyslexic readers’ integration of information across representations and web pages. the few studies that exist on the online reading of students with dyslexia have focused more on the learning outcomes of internet reading (castek et al., 2011), internet reading as a community of practice (henry et al., 2012), dyslexia-friendly interfaces (mccarthy & swierenga, 2010), and the difference between reading printed and digital text (schneps, thomson, chen, sonnert, & pomplum, 2013). hence, knowledge about how readers with dyslexia process information in a digital multimedia context is essentially lacking. it should be noted that this lack of research makes it difficult to draw firm conclusions regarding readers with dyslexia based on previous work, since this group of readers is almost invisible in research on information processing in multimedia contexts. however, this study brings together two research fields that have traditionally been kept separate in research and theory, that is, the study of dyslexia and the study of learning in a multimedia environment. in the present information society, where the vast majority of adolescents, including those with dyslexia, use multimedia materials such as the internet as an important information source in their school work, this study represents a timely integration of important areas of research that, hopefully, will inspire much further work. keypoints students with dyslexia worked with multiple web pages and representations eye-tracking and interviews showed differences in processing and strategy use processing was related to individual differences and learning and integration learning and integrating information at the same time were problematic acknowledgments thanks are due to shane colvin and arild moland for help in creating the learning materials. references abtahi, m.s. (2012). interactive multimedia learning object (imlo) for dyslexic children. procedia – social and behavioral sciences, 47, 1206-1210. doi:10.1016/j.sbspro.2012.06.801 afflerbach, p., pearson, p.d., & paris, s.g. (2008). clarifying differences between reading skills and reading strategies. the reading teacher, 61, 364-373. doi:10.1598/rt.61.5.1 ainsworth, s. (2018). multiple representations and multimedia learning. in f. fischer, c.e. hmelo-silver, s.r. goldman, & p. reiman (eds.), international handbook of the learning sciences (pp. 96-105). new york: routledge. alty, j. l., al‐sharrah, a., & beacham, n. (2006). when humans form media and media form humans: an experimental study examining the effects different digital media have on the learning outcomes of students who have different learning styles. interacting with computers, 18 , 891–909. doi:10.1016/j.intcom.2006.04.002 anderson, t.h., & ambruster, b.b. (1984). studying. in p.d. pearson, m. kamil, r. barr, & p. rosenthal (eds.), handbook of reading research (1st ed., pp. 657–679). white plains, ny: longman. andresen, a., anmarkrud, ø., & bråten, i. (2019). investigating multiple source use among students with and without dyslexia. reading and writing, 32, 1149-1174. https://doi.org/10.1007/s11145-018-9904-z anmarkrud, ø., brante, e.w., & andresen, a. (2018). potential processing challenges of internet use among readers with dyslexia. in j.l.g. braasch, i. bråten, & m.t. mccrudden (eds.), handbook of multiple source use (pp. 117-132). new york: routledge. anmarkrud, ø., & ferguson, l.e. (2011). working memory and topic knowledge of norwegian 10th and 11th graders . unpublished data set. oslo: faculty of educational sciences, university of oslo. austin, k.a. ( 2009). multimedia learning: cognitive individual differences and display design techniques predict transfer learning with multimedia learning modules. computers & education, 53, 1339-1354. doi:10.1016/j.compedu.2009.06.017 avons, s.e., & hanna, c. (1995). the memory-span deficit in children with specific reading-disability – is speech rate responsible? british journal of developmental psychology, 13, 303-311. doi: 10.1111/j.2044-835x.1995.tb00681.x ayres, p., & sweller, j. (2014). the split-attention principle in multimedia learning. in r.e. mayer (ed.), the cambridge handbook of multimedia learning (pp. 206-226). new york: cambridge university press. azevedo, r., & cromley, j.g. (2004). does training on self-regulated learning facilitate students’ learning with hypermedia? journal of educational psychology, 96, 523-535. doi: 10.1037/0022-0663.96.3.523 baddeley, a. (1995). working memory. oxford: clarendon press. baddeley, a.d. (2000). the episodic buffer: a new component of working memory? trends in cognitive science, 4, 417-423. doi: 10.1016/s1364-6613(00)01538-2 bahnmueller, j., huber, s., nuerk, h.c., göbel, s.m., & moeller, k. (2016). processing multi-digit numbers: a translingual eye-tracking study. psychological research, 80, 422-433. doi: 10.1007/s00426-015-0729-y baker, l., & beall, l.c. (2009). metacognitive processes and reading comprehension. in s.e. israel & g.g. duffy (eds.), handbook of research on reading comprehension (pp. 373–388). new york: routledge. barbosa, t., miranda, m.c., santos, r.f., & bueno, o.f.a. (2009). phonological working memory, phonological awareness, and language in literacy difficulties in brazilian children. reading and writing, 22, 201-218. doi: 10.1007/s11145-007-9109-3 barzilai, s., & ka’adan, i. (2017). learning to integrate divergent information sources: the interplay of epistemic cognition and epistemic metacognition. metacognition and learning, 12, 193-232. doi: 10.1007/s11409-016-9165-7 beacham, n.a., & alty, j.l. (2006). an investigation into the effects that digital media can have on the learning outcomes of individuals who have dyslexia. computers & education, 47, 74-93. doi: 10.1016/j.compedu.2004.10.006 berninger, v. w., raskind, w., richards, t., abbott, r., & stock, p. (2008). a multidisciplinary approach to understanding developmental dyslexia within working-memory architecture: genotypes, phenotypes, brain, and instruction. developmental neuropsychology, 33, 707-744. doi: 10.1080/87565640802418662 bishop, d.v.m., & snowling, m.j. (2004). developmental dyslexia and specific language impairment: same or different? psychological bulletin, 130, 858-886. doi: 10.1037/0033-2909.130.6.858 björnsson, c. h. (1968). läsbarhet [readability]. stockholm: liber. borella, e., carretti, b., & pelegrina, s. (2010). the specific role of inhibition in reading comprehension in good and poor comprehenders. journal of learning disabilities, 43, 541-552. doi: 10.1177/0022219410371676 braasch, j.l.g., bråten, i., strømsø, h.i., & anmarkrud, ø. (2014). incremental theories of intelligence predict multiple document comprehension. learning and individual differences, 31, 11-20. doi: 10.1016/j.lindif.2013.12.012 bråten, i., anmarkrud, ø., brandmo, c., & strømsø h.i. (2014). developing and testing a model of direct and indirect relationships between individual differences, processing, and multiple-text comprehension . learning and instruction, 30, 9-24. doi: 10.1016/j.learninstruc.2013.11.002 bråten, i., braasch, j.l.g., & salmerón, l. (in press). reading multiple and non-traditional texts: new opportunities and new challenges. in e.b. moje, p. afflerbach, p. enciso, & n.k. lesaux (eds.), handbook of reading research (vol. v). new york: routledge. butcher, k.r. (2014). the multimedia principle. in r.e. mayer (ed.), the cambridge handbook of multimedia leraring (2nd ed., pp. 174-205). new york: cambridge university press. campbell, s. (2012). comparative case study. in a.j. mills, g. durepos, & e. wiebe (eds.), encyclopedia of case study research (pp. 175-176). thousand oaks, ca: sage. castek, j., zawilinski, l., mcverry, j.g., o'byrne, w.i., & leu, d.j. (2011). the new literacies of online reading comprehension: new opportunities and challenges for students with learning difficulties. in c. wyatt-smith, j. elkins, & s. gunn (eds.), multiple perspectives on difficulties in learning literacy and numeracy (pp. 91-110). new york: springer. chan, e., & unsworth, l. (2011). image-language interaction in online reading environments: challenges for students' reading comprehension. australian educational researcher, 38, 181-202. doi: 10.1007/s13384-011-0023-y cho, b.-y., & afflerbach, p. (2017). an evolving perspective of constructively responsive reading comprehension strategies in multilayered digital text environments. in s.e. israel (ed.), handbook of research on reading comprehension (2nd ed., pp. 109-134). new york: guilford. cho, b.-y., afflerbach, p., & han, h. (2018). strategic processing in accessing, comprehending, and using multiple sources online. in in j.l.g. braasch, i. bråten, & m.t. mccrudden (eds.), handbook of multiple source use (pp. 133-150). new york: routledge. cho, b.-y., woodward, l., & li, d. (2017). examining adolescents' strategic processing during online reading with a question generating task. american educational research journal, 54, 691-724. doi: 10.3102/0002831217701694 clark, j.m., & paivio, a. (1991). dual coding theory and education. educational psychology review, 3, 149-210. doi: 10.1007/bf01320076 corriveau, k.h., einav, s., robinson, e.j., & harris, p.l. (2014). to the letter: early readers trust print-based over oral instructions to guide their actions. british journal of developmental psychology, 32, 345-358. doi: 10.1111/bjdp.12046 cuevas , h.m., fiore , s.m., & oser, r.l. (2002). scaffolding cognitive and metacognitive processes in low verbal ability learners: use of diagrams in computer based training environments. instructional science, 30, 433–464. doi: 10.1023/a:1020516301541 daneman, m., & carpenter, p.a. (1980). individual differences in working memory and reading. journal of verbal learning and verbal behavior, 19, 450-466. doi: 10.1016/s0022-5371(80)90312-6 de olivera, d.g., da silva, p.b., dias, n.m., sebra, a.g., & macedo, e.c. (2014). reading component skills in dyslexia: word recognition, comprehension, and processing speed. frontiers in psychology, 5: 1339. doi: 10.3389/fpsyg.2014.01339 deschryver, m. (2015). higher order thinking in an online world: toward a theory of web-mediated knowledge synthesis.teachers college record, 116, 1-44. http://www.tcrecord.org id number: 17692 einav, s., robinson, e.j., & fox, a. (2012). take it as read: origins of trust in knowledge gained from print. journal of experimental psychology, 114, 262-274. doi: 10.1016/j.jecp.2012.09.016 eitel, a., & scheiter, k. (2015). picture or text first? explaining sequence effects when learning with pictures and text. educational psychology review, 27,153–180. doi: 10.1007/s10648-014-9264-4 eitel, a., scheiter, k., & schüler, a. (2013). how inspecting a picture affects processing of text in multimedia learning. applied cognitive psychology, 27, 451–461. doi: 10.1002/acp.2922 eitel, a., scheiter, k., schüler, a., nyström, m., & holmqvist, k. (2013). how a picture facilitates the process of learning from text: evidence for scaffolding. learning and instruction, 28, 48–63. doi: 10.1016/j.learninstruc.2013.05.002 eyden, j., robinson, e.j., einav, s., & jaswal, v.k. (2013). the power of print: children’s trust in unexpected printed suggestions. journal of experimental child psychology, 116, 593-608. doi: 10.1016/j.jecp.2013.06.012 fenesi, b., kramer, e., & kim, j.a. (2016). split-attention and coherence principles in multimedia instruction can rescue performance for learners with lower working memory capacity. applied cognitive psychology, 30, 691-699. doi: 10.1002/acp.3244 ferguson, l.e., & bråten, i. (2013). student profiles of knowledge and epistemic beliefs: changes and relations to multiple-text comprehension. learning and instruction, 25, 49-61. doi: 10.1016/j.learninstruc.2012.11.003 fischbach, a., könen, t., rietz, c. s., & hasselhorn, m. (2014). what is not working in working memory of children with literacy disorders? evidence from a three-year-longitudinal study. reading and writing, 27, 267-286. doi: 10.1007/s11145-013-9444-5 follmer, d. j. (2018). executive function and reading comprehension: a meta-analytic review. educational psychologist, 53, 42-60. doi: 10.1080/00461520.2017.1309295 fox, e., & alexander, p.a. (2017). text and comprehension. in s.e. israel (ed.), handbook of research on reading comprehension (2nd ed., pp. 335-352). new york: guilford. furnes, b., & norman, e. (2015). metacognition and reading: comparing three forms of metacognition in normally developing readers and readers with dyslexia. dyslexia, 21, 273-284. doi: 10.1002/dys.1501 gathercole, s.e., alloway, t.p., willis, c., & adams, a.-m. (2006). working memory in children with reading disabilities. journal of experimental child psychology, 93, 265-281. doi: 10.1016/j.jecp.2005.08.003 gerjets, p., scheiter, k., opfermann, m., hesse, f.w., & eysink, t.h. (2009). learning with hypermedia: the influence of representational formats and different levels of learner control on performance and learning behavior. computers in human behavior, 360–370. doi: 10.1016/j.chb.2008.12.015 greene, j.a., moos, d.c., azevedo, r., & winters, f.i. (2008). exploring differences between gifted and grade-level students’ use of self-regulatory learning processes with hypermedia. computers & education, 50, 1069-1083. doi:10.1016/j.compedu.2006.10.004 harm, m.v., & seidenberg, m.s. (1999). phonology, reading acquisition, and dyslexia: insights from connectionist models. psychological review, 106, 491-528. doi: 10.1037/0033-295x.106.3.491 hegarty, m., & just, m.a. (1993). constructing mental models of machines from text and diagrams. journal of memory and language, 32,717–742. doi: 10.1006/jmla.1993.1036 henry, l.a., castek, j., o'byrne, w.i., & zawilinski, l. (2012). using peer collaboration to support online reading, writing, and communication: an empowerment model for struggling readers. reading & writing quarterly, 28, 279-306. doi: 10.1080/10573569.2012.676431 høien, t. (2014). logos teoribasert diagnostisering av lesevansker [logos theory based assessment of reading difficulties] . bryne, norway: logometrica. hulme, c., & snowling, m. j. (2009). developmental disorders of language learning and cognition. chichester: wiley-blackwell. irrazabal, n., saux, g., & burin, d. (2016). procedural multimedia presentations: the effects of working memory and task complexity on instruction time and assembly accuracy. applied cognitive psychology, 30, 1052-1060. doi: 10.1002/acp.3299 jeffries, s., & everatt, j. (2004). working memory: its role in dyslexia and other specific learning disabilities. dyslexia, 10, 196-214. doi: 10.1002/dys.278 johnson, c.i., & mayer, r.e. (2012). an eye movement analysis of the spatial contiguity effect in multimedia learning. journal of experimental psychology: applied, 18, 178-179. doi: 10.1037/a0026923 just, m.a., & carpenter, p.a. (1992). a capacity theory of comprehension: individual differences in working memory. psychological review, 99, 122-149. doi: 10.1037/0033-295x.99.1.122 kammerer, y., meier, n., & stahl, e. (2016). fostering secondary-school students' intertext model formation when reading a set of websites: the effectiveness of source prompts. computers & education, 102, 52-64. doi: 10.1016/j.compedu.2016.07.001 kingsley, t., & tancock, s. (2013). internet inquiry: fundamental competencies for online comprehension. the reading teacher, 67, 389-399. doi: 10.1002/trtr.1223 klinkenberg, j.e., & skaar, e. (2003). stas: standardisert test i avkoding og staving [stas: standarized test of decoding and spelling] . hønefoss, norway: ringerike ppt. knoop-van campen, c.a.n., segers, e., & verhoeven, l. (2018). the modality and redundancy effects in multimedia learning in children with dyslexia. dyslexia, 24, 140-155. doi: 10.1002/dys.1585 kurby, c.a., britt, m.a., & magliano, j.p. (2005). the role of top-down and bottom-up processes in between-text integration. reading psychology, 26, 335–362. doi: 10.1080/02702710500285870 leu, d.j., kiili, c., & forzani, e. (2016). infividual differences in the new literacies of online research and comprehension. in p. afflerbach (ed.), handbook of individual differences in reading (pp. 259-272). new york: routledge. lyon, g.r., shaywitz, s.e., & shaywitz, b.a. (2003). a definition of dyslexia. annals of dyslexia, 53, 1-14. doi: 10.1007/s11881-003-0001-9 maccullagh, l., bosanquet, a., & badcock, n. (2017). university students with dyslexia: a qualitative exploratory study of learning practices, challenges, and strategies. dyslexia, 23, 3-23. doi: 10.1002/dys.1544 mason, l., junyent, a.a., & tornatora, m.c. (2014). epistemic evaluation and comprehension of web-source information on controversial science-related topics: effects of a short-term instructional intervention. computers & education, 76, 143-157. doi: 10.1016/j.compedu.2014.03.016 mason, l., pluchino, p., & tornatora, m.c. (2016). using eye-tracking technology as an indirect instruction tool to improve text and picture processing and learning. british journal of educational technology, 47, 1083-1095. doi: 10.1111/bjet.12271 mason, l., scheiter, k., & tornatora, m.c. (2017). using eye-movements to model the sequence of text-picture processing for multimedia comprehension. journal of computer assisted learning, 33, 443-460. doi: 10.1111/jcal.12191 mayer, r.e. (2001). multimedia learning. new york: cambridge university press. mayer, r.e. (2003). the promise of multimedia learning: using the same instructional design methods across different media. learning and instruction, 13, 125-139. doi: 10.1016/s0959-4752(02)00016-6 mayer, r.e. (2014a). introduction to multimedia learning. in r.e. mayer (ed.), the cambridge handbook of multimedia learning (pp. 1-24). new york: cambridge university press. mayer, r.e. (2014b). cognitive theory of multimedia learning. in r.e. mayer (ed.), the cambridge handbook of multimedia learning (pp. 43-71). cambridge: cambridge university press. mayer, r.e. (ed.) (2014c), the cambridge handbook of multimedia learning. new york: cambridge university press. mayer, r.e., heiser, h., & lonn, s. (2001). cognitive constraints on multimedia learning: when presenting more material results in less understanding. journal of educational psychology, 93, 187-198. doi: 10 1037i/0022-0663 93.1 187 mayer, r.e., & moreno, r. (2010). techniques that reduce extraneous cognitive load and manage intrinsic cognitive load during multimedia learning. in j.l. plass, r. moreno, & r. brünken (eds.), cognitive load theory (131-152). new york: cambridge university press. mccarthy, j.e., & swierenga, s.j. (2010). what we know about dyslexia and web accessibility: a research review. universal access in the information society, 9, 147-152. doi: 10.1007/s10209-009-0160-5 melby-lervåg, m., lyster, s.a.h., & hulme, c. (2012). phonologival skills and their role in learning to read: a meta-analytic review. psychological bulletin, 138, 322-352. doi: 10.1037/a0026744 menghini, d., finzi, a., carlesimo, g. a., & vicari, s. (2011). working memory impairment in children with developmental dyslexia: is it just a phonological deficity? developmental neuropsychology, 36, 199-213. doi: 10.1080/87565641.2010.549868 moreno, r., & mayer, r.e. (2000). engaging students in active learning: the case for personalized multimedia messages. journal of educational psychology, 92, 724-733. doi: 10.1037//0022-06m.92.4.724 moreno, r., & mayer, r.e. (2002). verbal redundancy in multimedia learning: when reading helps listening. journal of educational psychology, 94, 156-163. doi: 10.1037//0022-0663.94.1.156 myers, j.l., & o'brien, e.j. (1998). accessing the discourse during reading. discourse processes, 26, 131-157. doi: 10.1080/01638539809545042 paivio, a. (1971). imagery and verbal processes. new york: oxford university press. paivio, a. (1986). mental representations: a dual-coding approach. new york: oxford university press. pociask, f.d., & morrison, g.r. (2008). controlling split attention and redundancy in physical therapy instruction. educational technology research and development, 56, 379–399. doi: 10.1007/s11423-007-9062-5 prado, c., dubois, m., & valdois, s. (2007). the eye movements of dyslexic children during reading and visual search: impact of the visual attention span. vision research, 47, 2521-2530. doi: 10.1016/j.visres.2007.06.001 ramus, f., rosen, s., dakin, s.c., day, b.l., castellote, j.m., white, s., & frith, u. (2003). theories of developmental dyslexia: insights from a multiple case study of dyslexic adults. brain, 126, 841-865. doi: 10.1093/brain/awg076 rau, a.k., moll, k., snowling, m.j., & landerl, k. (2015). effects of orthographic consistency on eye movement behavior: german and english children and adults process the same words differently. journal of experimental child psychology, 130, 92-105. doi: 10.1016/j.jecp.2014.09.012 rayner, k., ardoin, s.p., & binder, k.s. (2013). children's eye movements in reading: a commentary. school psychology review, 42, 223-233. rayner, k., chace, k.h., slattery, t.j., & ashby, j. (2006). eye movements as reflections of comprehension processes in reading. scientific studies of reading, 10, 241-255. doi: 10.1207/s1532799xssr1003_3 rieber , l.p., tzeng, s.-c., & tribble, k. (2004). discovery learning, representation, and explanation within a computer-based simulation: finding the right mix. learning and instruction, 14, 307–323. doi: 10.1016/j.learninstruc.2004.06.008 roca, j., tejero, p., & insa, b. (2018). accident ahead? difficulties of drivers with and without reading impairment recognizing words and pictograms in variable message signs. applied ergonomics, 67, 83-90. doi: 10.1016/j.apergo.2017.09.013 roeschl-heils, a., schneider, w., & van kraayenoord, c.e. (2003). reading, metacognition, and motivation: a follow-up study of german students in grades 7 and 8. european journal of psychology of education, 18, 75–86. doi: 10.1007/bf03173605 rouet, j.-f., & britt, m.a. (2014). multimedia learning from multiple documents. in r.e. mayer (ed.), the cambridge handbook of multimedia learning (2nd ed., pp. 813-841). new york: cambridge university press. rukavina, i., & daneman, m. (1996). integration and its effect on acquiring knowledge about competing scientific theories from text. journal of educational psychology, 88, 272-287. doi: 10.1037/0022-0663.88.2.272 salmerón, l., strømsø, h.i., kammerer, y., stadtler, m., & van den broek, p. (2018). comprehension processes in digital reading. in m. barzillai, j. thomson, s. schroeder, & p. van den broek (eds.), learning to read in a digital world (pp. 91-120). amsterdam: john benjamins. schneider, w. (2011). memory development in childhood. in u. goswami (ed.), the wiley-blackwell handbook of childhood cognitive development (2nd ed., pp. 347-376). malden, ma: wiley-blackwell. schneps, m.h., thomson, j.m., chen, c., sonnert, g., & pomplum, m. (2013). e-readers are more effective than paper for some with dyslexia. plos one, 8, e75634. doi: 10.1371/journal.pone.0075634 schotter, e.r., tran, r., & rayner, k. (2014). don’t believe what you read (only once): comprehension is supported by regressions during reading. psychological science, 25, 1218-1226. doi: 10.1177/0956797614531148 schüler, a., scheiter, k., & van genuchten, e. (2011). the role of working memory in multimedia instruction: is working memory working during learning from text and pictures? educational psychology review, 23 , 389-411. doi: 10.1007/s10648-011-9168-5 shaywitz, s.e., & shaywitz, b.a. (2008). paying attention to reading: the neurobiology of reading and dyslexia. development and psychopathology, 20, 1329-1349. doi: 10.1017/s0954579408000631 smith-spark, j.h., & fisk, j.e. (2007). working memory functioning in developmental dyslexia. memory, 15, 34-56. doi: 10.1080/09658210601043384 stanovich, k.e. (1986). matthew effects in reading: some consequences of individual differences in the acuistion of literacy. reading research quarterly, 21, 360-407. doi: 10.1598/rrq.21.4.1 swanson, h.l., & trahan, m.f. (1992). learning disabled readers' comprehension of computer mediated text: the influence of working memory, metacognition, and attribution. learning disabilities research and practice, 7, 74-86. sweller, j., ayres, p., & kalyuga, s. (2011). cognitive load theory. new york: springer. taylor, m., duffy, s., & hughes, g. (2007). the use of animation in higher education teaching to support students with dyslexia. education + training, 49, 25-35. doi: 10.1108/00400910710729857 torcasio, s., & sweller, j. (2010). the use of illustrations when learning to read: a cognitive load theory approach. applied cognitive psychology, 24, 659–672. doi: 10.1002/acp.1577 torgersen, j.k. (2001). the theory and practice of intervention: comparing outcomes from prevention and remediation studies. in a. fawcett & r. nicolson (eds.), dyslexia: theory and good practice (pp. 185-201). london: fulton. torgersen, j.k., alexander, a.w., wagner, r.k., rashotte, c.a., voeller, k., conway, t., et al. (2001). intensive remedial instruction with severe reading disabilities: immediate and long-term outcomes from two instructional approaches. journal of learning disabilities, 34, 33-58. doi: 10.1177/002221940103400104 trakhman, l.m.s., alexander, p.a., & berkowitz, l.e. (in press). effects of processing time on comprehension and calibration in print and digital mediums. the journal of experimental education. advance online publication. doi: 10.1080/00220973.2017.1411877 van de vijver, f.j., & harsveld, m. (1994). the incomplete equivalence of the paper-and-pencil and computerized versions of the general aptitude test battery. journal of applied psychology, 79, 852–859. doi: 10.1037/0021-9010.79.6.852 van roy, b. & pretorius, e.j (2013). is reading in an agglutinating language different from an analytic language? an analysis of isizulu and english reading based on eye movements. southern african linguistics and applied language studies, 31, 281-297. doi: 10.2989/16073614.2013.837603 van strien, j.l.h., brand-gruwel, s., & boshuizen, h.p.a. (2014). dealing with conflicting information from multiple nonlinear texts: effects of prior attitudes. computers in human behavior, 32, 101-111. doi: 10.1016/j.chb.2013.11.021 vinje, f.e. (1982). journalistspråket [ the journalist language]. fredrikstad, norway: institute for journalism. yin, r.k. (2009). case study research: design and methods (4th ed.). thousand oaks, ca: sage. codepen 2. willems frontline learning research special issue vol.9 no.2 (2021) 28 49 issn 2295-3159 predicting freshmen’s academic adjustment and subsequent achievement: differences between academic and professional higher education contexts jonas willems1, tine van daal1, peter van petegem1, liesje coertjens2 & vincent donche1 1university of antwerp, belgium 2uclouvain, belgium article received 18 march 2020 / revised 28 august/ accepted 4 december/ available online 12 march 2021 abstract this study tests an integrative model, which delineates how students’ academic motivation, academic self-efficacy and learning strategies (processing strategies and regulation strategies) at the end of secondary education impact academic adjustment in the first semester of the first year of higher education (fyhe) and subsequent academic achievement at the end of the fyhe, in two types of he programmes. more precisely, the present study explores the extent to which the explanatory values of aforementioned determinants of academic adjustment and academic achievement differ across academic (providing more theoretical and scientific education) and professional (offering more vocational education that prepares students for a particular occupation, such as nursing) programmes. hereto, multiple-group sem analyses were carried out on a longitudinal dataset containing 1987 respondents (academic programmes: n=1080, 54.4%; professional programmes: n=907, 45.6%), using mplus 8.3. results indicate differences in the predictive power of determinants under scrutiny between professional and academic contexts. firstly, learning strategies and motivational variables at the end of secondary education have more predictive power in the prediction of fyhe academic adjustment in the academic programmes than in professional programmes. secondly, our results indicate that academic adjustment in the first semester of the fyhe influences academic achievement to a bigger extent in professional programmes than in academic programmes. moreover, these differences across he contexts were found after controlling for prior education. implications of the findings are discussed. keywords: learning strategies; motivation; academic adjustment; first-year academic achievement; programme diversity info corresponding author email: jonas.willems@uantwerpen.be doi: https://doi.org/10.14786/flr.v9i2.647 1. introduction over the years, democratisation of higher education (he) around the world has led to a substantial increase and diversification of the student population enrolling in he (schuetze & slowey, 2002). this seems to be accompanied by low study success rates, early drop-out and study delay of students in the first year of higher education (fyhe). for example, in flanders (dutch speaking part of belgium), only 48.6% of freshmen successfully complete their required coursework in the fyhe (declercq & verboven, 2014). this has extensive psychological and financial costs for the individual student, the family and society (oecd, 2013). as such, more insight into factors that facilitate freshmen’s transition process to he can give rise to an increase in the academic achievement of these students (briggs, clark, & hall, 2012). in recent decades, several lines of research have argued that non-cognitive factors such as learning strategies and motivational and adjustment variables are important determinants of students’ academic achievement in the fyhe (e.g. bailey & phillips, 2016; credé & kuncel, 2008; richardson, abraham, & bond, 2012; robbins et al., 2004). this body of research, however, has been mostly carried out in academically oriented he programmes (offering more theoretical and scientific education), leaving professionally oriented he contexts (offering more vocationally oriented education that prepares students for a specific occupation) rather underexplored (for an exception, see vanthournout, gijbels, coertjens, donche, & van petegem, 2012). this paucity of research in professional he contexts is certainly remarkable, given that a significant part of adolescents worldwide enrols in professional he programmes (oecd, 2009), for example, in flanders, 54.4% of the he student population participate in professional he (flemish government, 2019). moreover, previous research points out that institutional and disciplinary differences might influence the interrelationships between variables in a predictive model of academic achievement (e.g. de clercq et al., 2013). simply assuming that the aforementioned determinants of academic achievement have the same predictive value in professional and academic fyhe contexts, thus, seems to neglect this important source of meso-level diversity. therefore, this study sets out to investigate to what extent the predictive power of learning strategies (processing strategies and regulation strategies), motivational variables and academic adjustment in predicting fyhe students’ academic achievement differs across academic and professional he contexts, using an integrative, longitudinal research design. in what follows, we firstly describe how he in flanders is organised, after which we briefly describe the main constructs and expected relationships under study – albeit as we will point out have been investigated in predominantly academic he contexts, with little attention to academic adjustment as an intervening variable for academic achievement in the fyhe. 2. research context: flemish he system as in many other european he systems (such as germany, the netherlands, finland, denmark and portugal), flemish he is provided by two types of institutions: universities and university colleges. universities offer academically oriented he programmes, which provide mainly theoretical and scientific education. they typically prepare students for a succeeding master programme and correspond to the bologna two-cycle programmes (bachelor and master, encompassing a total of 4 or 5 years; the bologna declaration, 1999). university colleges, on the other hand, are specialised institutions that organise so called ‘professional bachelor programmes’, which are mainly designed for learners to acquire the knowledge, skills and competencies specific to a particular occupation, such as nursing or social work (camilleri, delplace, frankowicz, hudak, & tannhäuser, 2014). these vocational programmes offer a direct access to the labour market and are in line with the bologna first cycle programmes (one cycle of 3 years). academic and professional bachelor programmes have different aims and expectations of students, and typically differ from each other with regard to their curricular organisation. in professional programmes, theory and practice are combined through the use of student-centred learning methodologies such as: simulations, working with real-life materials and workplace learning settings (e.g. long-term internships, machinery to repair, assignments for translators, see also camilleri et al. 2014). in academic programmes, then, subject matter is more abstract and often less practical. also, the teaching speed is higher, research activities and large-scale lectures are more common, and more independent learning and scientific research attitudes are expected from students in these academic programmes (van rooij et al., 2017). 3. the pivotal role of academic adjustment in predicting academic achievement academic adjustment is generally described as the extent to which a student successfully copes with the various educational demands and characteristics of the new he environment, and comprises components such as motivation to learn, taking action to meet academic demands, having a clear sense of purpose, management of expectations, and general satisfaction with the academic environment (baker, mcneil, & siryk, 1985; baker & siryk, 1999; gerdes & mallinckrodt, 1994). today, it is well established from a variety of studies that academic adjustment is imperative in the prediction of students’ academic achievement in the fyhe: students who are more academically adjusted drop out less often (bean, 1980; kuh, kinzie, buckley, bridge, & hayek, 2006) and achieve better grades (bailey & phillips, 2016; petersen, louw, & dumont, 2009; prospero & vohra-gupta, 2007; severiens & wolff, 2008; wintre et al., 2011). considering this importance of academic adjustment in the prediction of freshmen’s academic achievement, it is not surprising that a considerable number of studies on the first-year transition experience treat this construct as an important outcome in its own right (e.g. garriott, love, & tyler, 2008; rice, vergara, & aldea, 2006). this latter body of research has unveiled that students’ learning strategies and motivational variables, on their turn, have considerable impact on first-year academic adjustment (e.g. baker, 2004; cazan, 2012). moreover, previous research in academic he contexts has suggested that academic adjustment is a mediator of the effects of several learning strategies and motivational variables on academic achievement (petersen, louw, & dumont, 2009; van rooij, jansen, & van de grift, 2018), which further highlights the pivotal role of the academic adjustment construct in first-year students’ transition process. for instance, van rooij and colleagues (2018) found that intrinsic (autonomous) motivation and self-regulated study behaviour did not influence academic achievement directly, but through academic adjustment. however, the work of peterson et al. (2009), who also investigated the mediating role of first-year students’ academic adjustment, suggests that this construct is not a “pure” mediator on academic achievement. indeed, these authors found that the effects of students’ intrinsic motivation and identified regulation (together autonomous motivation) and self-esteem were mediated by adjustment, while extrinsic regulation (controlled motivation) and academic overload (being unable to cope with the academic workload) had a direct impact on academic achievement. this rationale leads us to the integrative conceptual model adopted in the present study, which delineates that students’ learning strategies (deep processing, surface processing, self-regulation, lack of regulation), academic motivation, and academic self-efficacy have an impact on academic adjustment in the first semester of the fyhe and subsequent academic achievement (fig. 1). acknowledging that academic adjustment might not be a pure mediator on academic achievement (peterson et al., 2009), the model also includes direct paths between the exogenous variables and academic achievement. furthermore, as it is clear from previous studies that students’ prior secondary education tracks might influence academic adjustment and achievement as well (e.g. de clercq et al., 2013; vermunt, 2005), in the present study, we have included this factor as a control variable. finally, for the design of the present study, we adhere to the suggestion of van rooij et al. (2018) that research on this matter should be conducted longitudinally and should “start measuring motivational and behavioural variables in secondary school and investigate how they relate to adjustment and student success outcomes later in university” (p. 763). figure 1. conceptual model of learning strategies and motivational variables affecting academic adjustment and subsequent academic achievement. from the above, it is clear that one might expect a positive relationship between academic adjustment and academic achievement. the subsequent paragraphs further detail on the hypothesised interrelations between motivational and learning related variables on the one hand, and academic adjustment and achievement on the other hand. the hypothesised directions of the associations in the conceptual model (fig.1) are summarised in table 1. 4. motivational and learning related determinants of academic adjustment and academic achievement 4.1 academic motivation it has previously been observed that fyhe students’ academic motivation (deci & ryan, 2000) is linked to academic adjustment. for instance, clark, middleton, nguyen, & zwick (2014), petersen et al. (2009) and van rooij et al. (2018) reported a positive relation between types of autonomous motivation and academic adjustment. amotivation, on the other hand, has been found to be associated with lower academic adjustment (baker, 2004). research further shows that students who are more autonomously motivated have higher academic achievement than students who are more controlled motivated or more amotivated (e.g. bailey & phillips 2016; guay, ratelle, roy, & litalien, 2010). finally, a number of studies revealed a negative association between amotivation and students’ academic achievement (e.g. prospero & vohra-gupta, 2007; vanthournout et al., 2012). 4.2 academic self-efficacy academic self-efficacy (further shortened to self-efficacy) is defined as individuals’ beliefs that they can successfully perform given academic tasks at designated levels (schunk, 1991). this construct has repeatedly been identified as one of the strongest determinants of academic achievement in the fyhe (e.g. richardson et al. 2012; robbins et al. 2004). furthermore, cazan (2012) and chemers, hu, & garcia (2001) demonstrated that self-efficacy was strongly and positively related with academic adjustment. van rooij et al. (2018), however, did not find a significant relationship between self-efficacy and academic adjustment, after controlling for intrinsic motivation, self-regulation and degree programme satisfaction. thus, the specific role of self-efficacy, especially after controlling for additional concepts, remains inconclusive. 4.3 learning strategies in the learning pattern model, developed by vermunt (1998), learning strategies are described to encompass both cognitive processing strategies and regulation strategies. firstly, processing strategies refer to those thinking activities and study skills students apply whilst studying (vermunt, 1998). generally, two qualitatively different types of cognitive processing are discerned in educational literature, namely deep and surface processing (e.g. vermunt & donche, 2017; vermunt & vermetten, 2004). deep processing refers to the use of learning activities that lead to meaningful learning and in-depth understanding of the learning content, such as relating and structuring. surface processing refers to the use of learning activities like memorizing that lead to the learning of superficial features of a study task, also described as root learning (vermunt & vermetten, 2004). traditionally, it has been argued that the use of deep processing strategies leads to high academic achievement, while surface processing strategies entail lower academic achievement (vermunt & donche, 2017). two important meta-analyses in the field corroborate this idea, as they found those relationships to be significant, albeit small (dent & koenka, 2016; richardson et al., 2012). however, the findings of studies on the direction of the relationship between cognitive processing and academic achievement in studies can also be inconclusive. for instance, it has been argued that surface learning, in some situations, might be advantageous to the learner (see dinsmore & alexander (2012) for an elaborate exposition). further, associations between cognitive processing strategies and academic adjustment have been less investigated. nevertheless, previous research has demonstrated that students who lack appropriate study skills in he are at risk of having problems with their academic adjustment (abbott-chapman, hughes, & wyld, 1992). therefore, we hypothesise that deep and surface processing strategies will be related to academic adjustment. secondly, regulation strategies are defined as those activities that students use to harness their cognitive processing strategies (schunk & zimmerman, 2012). students who are more self-regulated are able to actively steer their own learning processes through activities such as planning tasks, monitoring progress, and diagnosing problems. lack of regulation, on the other hand, refers to an absence of clarity on how to steer the learning process (vermunt & donche, 2017). several studies have found self-regulation to be positively related to both academic adjustment (cazan, 2012; hurtado et al., 2007; van rooij et al., 2018) and academic achievement (e.g. dent & koenka, 2016; richardson et al., 2012). although lack of regulation has repeatedly been found to affect academic achievement in a negative fashion (e.g. donche & van petegem, 2010; vermunt, 2005), to our knowledge, there has been no investigation of the relationship between lack of regulation and academic adjustment in professional and academic programmes. considering the ‘deficit’-character of the construct, we theoretically expect that lack of regulation is negatively associated to academic adjustment. table 1 hypothesised directions of the relationships under study 5. exploring programme diversity: a meso-level study when reviewing the literature on determinants of academic adjustment and academic achievement, it becomes apparent that relatively few studies have tackled these relationships in the specific setting of professional he. indeed, the vast majority of studies on the relationships under scrutiny (fig.1) have been carried out in academically oriented programmes (e.g. petersen et al., 2009; van rooij et al., 2018). several scholars, however, have established that disciplinary differences influence the learning environments wherein students reside in terms of, for instance; requirements of students, assessment systems, study goals and teaching methods (becher, 1994; braxton & hargens, 1996; young, 2010). moreover, previous research has suggested that such variations in environments might influence interrelationships between non-cognitive variables and academic achievement. an example of this disciplinary diversity is provided by de clercq et al. (2013) who investigated whether freshmen’s background, study choice process, experience of the university, motivational beliefs, learning strategies, and behavioural engagement had similar predictive power in two university disciplines: science and physical education. the authors found several differences in the effects of those determinants; in physical education courses, for example, self-efficacy was the most powerful predictor of academic achievement, whereas intention to persist was the most powerful determinant in the science discipline. another study by lizzio, wilson, & simons (2002) showed that relationships between university students’ prior achievement, learning strategies and academic achievement varied between faculties of humanities, science, and commerce. this also concurs with the study by fonteyne, duyck & de fruyt (2017), who found that the predictive power of background, cognitive, personality, metacognitive, self-efficacy and motivational factors on academic achievement varied considerably across various academic study disciplines, such as psychology, criminology, history, and pharmaceutical sciences. however, students’ academic adjustment, as a possible mediator in further understanding the effects of entry characteristics on academic achievement, was not taken into account in these studies. moreover, professionally oriented programmes have distinctive aims and expectations of students and typically adopt different didactical approaches than academically oriented programmes (camilleri et al., 2014; see also “research context: flemish he system”). we therefore expect that such programme diversity could influence relationships in a predictive model of academic achievement as well. determinants of academic adjustment and academic achievement might, thus, have a dissimilar predictive value in both contexts. therefore, the aim of this study is to explore this programme diversity, by comparing the impact of the different determinants depicted in the conceptual model (fig.1), in professionally and academically oriented programmes. the following two research questions are central to this study: rq1: to what extent is the explanatory value of secondary students’ academic motivation, academic self-efficacy and learning strategies (processing and regulation strategies) in the prediction of first-year he academic adjustment, different between professional and academic fyhe programmes? rq2: to what extent is the explanatory value of secondary students’ academic motivation, academic self-efficacy, learning strategies and subsequent first-year academic adjustment, in the prediction of first-year he academic achievement, different between professional and academic fyhe programmes? 6. method 6.1 respondents & procedure the data stem from a longitudinal research project on students’ transition from secondary to he in flanders. in this project, students from 32 randomly selected secondary schools (offering a mixture of secondary education (se) tracks; general, arts, technical and vocational) participated and were followed up until the second year of he. at the end of their last year of se, students completed questionnaires (both online and paper and pencil) measuring their academic motivation, self-efficacy and learning strategies (wave 1: may/june 2011, n=2839). informed consent and contact information for future research was obtained from 84.1% of these students (n=2387). data obtained from the flemish government show that 1987 (83.3%) of students who had given their informed consent transitioned to he, constituting the final sample for this study. a small majority of these students (n=1080, 54.4%) started an academic bachelor programme, whereas 907 (45.6%) opted for a professional bachelor programme. at a second wave, during the first semester of the fyhe, students’ academic adjustment was mapped out in an online questionnaire. several communication channels were used to reach respondents (letter, e-mail, sms, phone call after repeated non-response), making this an intensive data collection that extended over three months (october december 2011). in the second wave of the study, 604 (30.4%; academic programmes: n=331; professional programmes: n=273) of the 1987 students who transitioned to he completed the academic adjustment scale. table 2 shows the proportions of respondents’ prior education tracks in both academic and professional he programmes, and compares these with the corresponding student proportions in the actual flemish population in academic year 2010-2011, as reported by glorieux, laurijssen, & sobczyk (2014). as no students from the vocational se track completed the academic adjustment scale in he, this group of students is not represented in the present study. table 2 proportion of prior education study tracks, in the sample and the flemish population, across academic and professional he programmes 6.2 measures students’ motivational characteristics and learning strategies at the end of se (wave1) were mapped out using scales of the self-report questionnaire ‘learning strategy and motivation questionnaire’, which was previously validated in flanders (lemo; donche & van petegem, 2008). academic motivation. controlled motivation was operationalised by six items, of which ‘i am motivated to study, because i am supposed to do this’ is an example (α=.73). autonomous motivation was also measured by six items, for instance, ‘i am motivated to study, because i want to learn new things’ (α=.83). finally, amotivation was measured using three items, such as ‘i am motivated to study…honestly, i don’t know; i feel like i’m wasting my time in school’ (α=.78). five answering categories were given, ranging from ‘not important at all’ to ‘very important’. self-efficacy is defined more specifically as a student’s perception of having the necessary knowledge and skills to carry out learning tasks. the self-efficacy scale exists of four items, for instance ‘i have confidence in the way in which i study’ (α=.84). items were scored on a five-point likert scale ranging from ‘completely disagree’ to ‘completely agree’. learning strategies. we opted to measure qualitatively different cognitive processing and regulation strategies. more concretely, surface processing was measured by the ‘memorizing’-scale (e.g. ‘i memorise lists of characteristics of a certain phenomenon’; 4 items; α=.67), and deep processing was measured by the ‘critical processing’-scale (e.g. ‘i try to understand the interpretations of experts in a critical way’; 4 items; α=.73). on the level of regulation strategies, we measured self-regulation (e.g. ‘in addition to the compulsory subject matter, i read other books or texts that have to do with the subject matter’; 4 items; α=.64), and lack of regulation (e.g. ‘i notice that it is difficult for me to determine whether i have sufficiently mastered the subject matter; 4 items; α=.70). all items are scored ranging from 1 (i never or hardly ever do this) to 5 (i almost always do this). academic adjustment (wave2) was measured using the ‘adjustment’-scale (6 items, α=.76), validated by torenbeek and colleagues (2010, 2011). an item example is: ‘i have experienced some difficulties in adjusting to the teaching approach of my current study programme (reverse coded)’. items were scored on a five-point likert scale, ranging from ‘completely disagree’ to ‘completely agree’. academic achievement data from students at the end of the fyhe (wave3) was obtained from the flemish government. it was conceptualised as study progress, which is the ratio of credits (ects study points) earned by a student versus the credits attempted by that student. credits are earned when a course is passed, which is when a student scored a minimum of 10 out of 20 on the evaluation for that course. prior education. data on students’ prior education tracks was obtained from the flemish government. flemish se is provided for young people aged 12 to 18 in four tracks: general se, technical se, artistic se, and vocational se. as mentioned above, the present study was not able to include students from the vocational se track. although differences in the educational tracks in general se exist, students from the general se track tend to be more prepared to enrol in an academic he study programme. all students in the flemish educational system are free to access either professional or academic he programs after se. for use in further analysis, this variable was dummy coded (0=general track, 1=arts/technical track). 6.3 analysis in the present study, multiple-group structural equation modelling (sem; byrne, 2016) is used to examine the fit of the conceptual model illustrated in figure 1, and to conduct cross-group comparisons between university college students and university students (rq1 and rq2). all analyses were carried out in mplus (version 8.3). in all models, the maximum likelihood estimator with robust standard errors (mlr) was used, which is robust to non-normality of observations (muthén & muthén, 2010). this method also allows for missing data handling; using the complete sample by incorporating data from respondents that did not participate in every wave, which has been found to provide better results in terms of unbiased estimates and statistical power (enders, 2011). global fit of the models is assessed using the ‘comparative fit index’ (cfi) and ‘root mean square error of approximation’ (rmsea). a model has excellent fit when cfi has a value above .95, and rmsea has a value less than .05. a model has acceptable fit, with a cfi-value above .90, and rmsea value is less than .08 (hu and bentler, 1999). a prerequisite to conducting substantive comparisons between groups is the establishment of measurement invariance across those groups (vandenberg & lance, 2000). hence, in a first step, we carried out multiple-group measurement invariance testing (meredith, 1993) to seek evidence that our measurement instrument operates equivalently across the two groups under scrutiny (i.e. do university college students and university students understand all the scales and items in a similar way?). hereto, four steps were undertaken, in each of which more restricted confirmatory factor analysis (cfa) models are estimated (byrne, 2016; vandenberg & lance, 2000): (1) a configural invariance model, wherein only the number of factors and the factor-loading pattern are equivalent across groups. in this stadium, there are no equality constraints imposed on the parameter estimates of the model; (2) a metric invariance model requires that only factor loadings are equal across groups; (3) a scalar invariance model, wherein intercepts are constrained as well; and (4) a strict invariance model, finally, imposes equality constrains on the error variances across groups (brown, 2014; gregorich, 2006). when metric invariance is established, this means that the different constructs in the measurement model have the same meaning in the two groups. scalar invariance, then, implies that the means of the scales across both groups can be compared. researchers generally agree that assessing scalar invariance is sufficient for establishing measurement invariance (milfont & fischer, 2010). in every proceeding step, the invariance of the factor structure was evaluated by comparing the fit of the more restricted model with the fit of the less restricted model (byrne, 2016). to this end, we examined changes in the following fit indices: cfi and rmsea. a decrease in cfi of .01 or more (cheung & rensvold, 2002) and an increase in rmsea of .015 or more (chen, 2007) was considered as evidence that the invariance hypothesis should be rejected1 . if scalar measurement invariance was not attained, the model was tested for partial scalar invariance (byrne, shavelson, & muthén, 1989). hereto, modification indices were examined to identify possible item(s) that induced the lack of equivalence, after which the particular intercepts of these items were allowed to differ between groups. byrne et al. (1989) and steinmetz, schmidt, tina-booh, wieczorek, & schwartz (2009) suggest that a minimum of two intercepts have to be equal across groups to establish partial scalar invariance of a scale. after this initial testing of the measurement models and their equivalence over students from different types of bachelor programmes, the structural model (fig. 1) was as stated above tested using a multiple-group sem approach. this allowed us to compare (a) the standardised regression parameter estimates and (b) the explained variances in the endogenous variables in both groups of students. in this study p < .05 is used as a criterion of statistical significance. further, in order to more accurately compare the predictive power in terms of explained variance of the factors under study on their respective outcomes across academic and professional programmes, we controlled for prior education, and, subsequently, scrutinised the incremental values of these factors over prior education. during analyses, we encountered a multicollinearity problem that was induced by high latent correlations between the ‘lack of regulation’ and ‘self-efficacy’ variables in the academic (r=-.815) as well as in the professional (r=-.651) bachelor programme group of respondents. this multicollinearity issue clearly influenced the analyses as theoretically inconceivable parameter estimates emerged when both constructs were included in one predictive model. since ‘lack of regulation’ and ‘self-efficacy’ are theoretically clearly distinctive and both concepts have considerable impact on first-year academic adjustment and academic achievement, we opted to retain both variables in the analysis. hence, we preferred to break down the model under scrutiny (fig.1) in two components containing either learning strategies or motivational factors as determinants of academic adjustment and subsequent academic achievement (fig.2). figure 2. split of the conceptual models of learning strategies and motivational variables affecting academic adjustment and subsequent academic achievement. 7. results 7.1 measurement invariance multiple-group measurement invariance analyses were conducted on the two conceptual models (see figure 2 in the method section). in the next paragraphs, we first provide an overview of the results for the learning strategies variables and subsequently for the motivational variables. 7.1.1 learning strategies after adding one error covariance term in the lack of regulation scale, and one in the academic adjustment scale, the configural model for learning strategies showed adequate fit (see table 3). next, inspection of the metric model shows that the hypothesis of invariant factor loadings was not rejected (∆cfi=.003, ∆rmsea=0). after constraining the intercepts (scalar model), however, model fit decreased significantly (∆cfi=.016, ∆rmsea=-.003). we needed to relax constraints on one intercept of the lack of regulation scale, and one of the surface processing scale to improve model fit sufficiently. subsequently, imposing equality constrains on the error variances across groups did not significantly reduce fit (see table 3). thus, results for this model suggest (1) metric invariance for all scales in the model, (2) scalar invariance for deep processing, self-regulation and academic adjustment, and (3) partial scalar invariance for lack of regulation and surface processing. table 3 results from measurement invariance tests for learning strategies and academic adjustment *** p<.001; a lr=one item of lack of regulation scale freed; b sp=one item of surface processing scale freed; c the reference point for the calculation of these values is the metric model. 7.1.2 motivational variables in order to achieve adequate fit for the configural model of the motivational variables (see table 4), we added three error covariances in the autonomous as well as in the controlled motivation scale. furthermore, one error covariance was added in the self-efficacy scale and one in the academic adjustment scale (the same as in the learning strategies model). as can be seen in table 4, the results from the measurement invariance tests, thus, provide evidence that strict invariance is established for autonomous and controlled motivation, amotivation, self-efficacy, and academic adjustment. table 4 results from measurement invariance tests for motivational variables and academic adjustment 7.2 multiple-group sem 7.2.1 learning strategies in a next step, a multiple group sem analysis containing the learning strategies component (model 1 in figure 2) provided satisfactory fit (cfi=.911, rmsea=.031). parameter estimates of this model (table 5) demonstrate that students’ first-year academic achievement is related with their prior education (dummy coded; 0=general se) in both the academic (β=-.149, p<.001) and professional (β=-.215, p<.001) contexts, indicating that students from more academically preparing study tracks in secondary education (general education), achieve better in first-year he. further, it appears that academic adjustment is a direct significant (positive) determinant of academic achievement in academic (β=.290, p<.001) programmes, as well as in professional (β=.334, p<.001) programmes. academic adjustment is, in its turn, significantly and negatively associated with lack of regulation in both types of programmes (acad.: β=-.430, p<.001; prof.: β=-.253, p=.004). finally, prior education predicts academic adjustment in professional programmes (β=-.248, p<.001), but not in academic programmes (β=-.096, p=.135). table 5 results of the multiple-group sem analysis for learning strategies; standardised parameter estimates and explained variances (r²) in academic and professional programmes a dummy coded: 0=general se in order to accurately compare the predictive power, in terms of explained variance, of learning strategies, motivational variables and academic adjustment on their respective outcomes across academic and professional programmes, we contrasted the incremental values of these factors over prior education, which are calculated in table 6 (i.e. δr² learning strategies in the prediction of academic adjustment, and δr² learning strategies + adjustment in the prediction of academic achievement). results show that the larger regression coefficients of learning strategies in the prediction of academic adjustment in the academic group, are also reflected in the explained variances (δr² learning strategies). learning strategies at the end of se predict over double the amount of variance in first-semester academic adjustment in academic programmes (19.4%) in comparison to professional programmes (7.2%). table 6 further shows that in professional he, 3.1% more variance in academic achievement is explained by learning strategies and academic adjustment (11.1%), compared to the academic context (8%). further analyses (see 7.2.3 incremental value of academic adjustment) will show that this latter difference in incremental value can be completely attributed to the increase in explained variance of academic adjustment, not learning strategies. table 6 calculation of incremental value (δr²) of learning strategies and academic adjustment over prior education a as also reported in table 5. 7.2.2. motivational variables the motivational variables model (model 2 in figure 2) had satisfactory fit as well (cfi=.924, rmsea=.039). parameter estimates of this model (table 7) demonstrate that prior education relates to adjustment in academic (β=-.128, p=.043) as well as in professional (β=-.276, p<.001) programmes. next, self-efficacy had a significant positive effect on academic adjustment in the academic programmes (β=.266, p<.001), but not in the professional programme group (β=.088, p=.285). controlled motivation was significantly and negatively related with adjustment in both types of programmes (acad.: β=-.188, p=.024; prof.: β=-.189, p=.028). academic adjustment, subsequently, positively predicted academic achievement in the academic programmes (β=.215, p=.001) as well as in the professional programmes (β=.333, p<.001). further, next to the indirect effect through academic adjustment in the academic group, self-efficacy also has a direct positive impact on academic achievement in the academic (β=.134, p=.002) and in the professional group (β=.117, p=.014). finally, prior education was also related to academic achievement in both types of programmes (acad.: β=-.160, p<.001; prof.: β=-.213, p<.001). table 7 results of the multiple-group sem analysis for motivational variables: parameter estimates and explained variances in academic and professional programmes similar to the learning strategies model, we observed that the incremental value of motivational variables in the prediction of academic adjustment, over prior education (table 8), is larger in academic programmes (12.9%) relatively to professional programmes (9.3%). however, this difference across programmes (3.6%) is smaller than that in the learning strategies model (12.2%). furthermore, motivational variables and academic adjustment together, in the professional programmes were able to predict 13% of variance in academic achievement, over and above prior education. this is 4.1% more variance than was predicted in the academic contexts, where the increase in explained variance of these variables over prior education was 8.9%. again, this latter difference in incremental value is attributable to the increase in explained variance of academic adjustment (see 7.2.3 incremental value of academic adjustment). table 8 calculation of incremental value (δr²) of motivational variables and academic adjustment over prior education a as also reported in table 7. 7.2.3 incremental value of academic adjustment in predicting academic achievement, over prior education, learning strategies and motivation the motivational variables model (table 7) showed that, in addition to academic adjustment, self-efficacy was a significant direct predictor of academic achievement. moreover, we know that the non-significant effects of the determinants of academic achievement presented in table 5 and 7 have a small impact on the reported explained variances as well in addition to the significant effects. hence, these considerations bring about the question of what the ‘net’ impact of academic adjustment on academic achievement is in terms of explained variance, and to what extent this is different between academic and professional programmes. therefore, we also examined a model containing only the direct impact of the learning strategies and motivational variables on academic achievement (models b in table 9 and 10), wherein academic adjustment was not included as predictor of achievement. this allowed us to estimate the explained variances (r²) of learning strategies and motivational variables in academic achievement, and consequently the incremental value (δr²) of academic adjustment in the predictive model. our calculations show that the difference in explained variance in academic achievement in the learning strategies model, between academic and professional contexts (δr²=3.1%, see table 6), can be completely ascribed to the incremental value of academic adjustment, as can be seen in model c of table 9 (9.8% 6.7%). indeed, learning strategies have identical incremental value over prior education in both academic and professional programmes (model b, table 9: δr²=1.3). table 9 calculation of incremental value (δr²) of academic adjustment in predicting academic achievement, over prior education and learning strategies closer inspection of table 10, then, shows that the difference of 4.1% explained variance in academic achievement in the motivational variables model (see table 7), can also be attributed to the fact that academic adjustment predicts more variance in professional programmes (δr²=9.4%), relative to academic programmes (δr²=3.9%); a difference of 5.5%. however, given that learning strategies have a larger incremental value over prior education in academic programmes (δr²=5%), in comparison to professional programmes (δr²=3.6%) a difference of 1.4% this larger predictive value of academic adjustment in professional programmes is slightly compensated. table 10 calculation of incremental value (δr²) of academic adjustment in predicting academic achievement, over prior education and motivation 8. discussion and conclusion this study set out to explore whether the determinants in our conceptual model (fig. 1) have dissimilar predictive power in professional he programmes, in comparison with more academically oriented programmes. more specifically, we examined (1) to what extent the explanatory value of secondary students’ academic motivation, academic self-efficacy and learning strategies in the prediction of first-year he academic adjustment, was different between these two types of programmes, and (2) whether secondary students’ academic motivation, academic self-efficacy, learning strategies and subsequent first-year academic adjustment, are similarly or differently predictive for academic achievement within the two different he programmes. we examined these relationships and differences in explanatory value of determinants, controlling for students’ prior education track. in what follows, we discuss the resulting parameter estimates of the multi-group sem-models in relation to previous research, after which we further focus on the differences in predictive power between academic and professional he contexts. 8.1 identified relationships in academic and professional he contexts our results indicate that academic adjustment in the first semester of the fyhe exerted the largest influence on academic achievement in both academic and professional programmes; students who feel more academically adjusted to their new learning environment in the first semester of he will obtain a higher percentage of their credits at the end of their first year. this result corroborates the findings of several previous studies in academic he contexts (bailey & phillips, 2016; petersen et al., 2009; prospero & vohra-gupta, 2007; severiens & wolff, 2008; wintre et al., 2011). the only other direct and positive significant predictor of academic achievement in both he contexts was self-efficacy which is also in line with former studies (richardson et al., 2012; robbins et al., 2004). however, as our study went beyond the consideration of separate direct effects by modelling the relationship between variables within integrated models, we could also identify some motivational and learning strategy variables that influenced academic achievement indirectly, through their impact on academic adjustment. firstly, students in academic programmes who had more confidence in their way of studying (self-efficacy) at the end of se felt more academically adapted in the first semester of the fyhe, which contradicts the findings by van rooij et al. (2018), but are in line with findings from other research (cazan, 2012; chemers et al., 2001). this latter relationship was non-significant in professional programmes. further, for both academic and professional he programmes, results confirm the hypotheses that students who have difficulties in steering their own learning process (lack of regulation) and those for whom the drivers for studying were more determined by external sources (controlled motivation) at the end of se, felt less adapted to their new learning environment. a strength of the present study is that all the above relationships were present after controlling for the expected effect of the prior se education track which students followed. in contrast to previous research, several variables were not significantly associated with academic adjustment in either he programmes: autonomous motivation (clark et al., 2014; petersen et al., 2009; van rooij et al., 2018), amotivation (baker, 2004) and self-regulation (cazan, 2012; hurtado et al., 2007; van rooij et al., 2018). the hypothesis that deep and surface processing at the end of se influences fyhe academic adjustment (abbott-chapman et al., 1992) was not supported by our data either. thus, similar to the study by van rooij et al. (2018), which highlighted the pivotal role of academic adjustment in predicting achievement in university, we found academic adjustment to be an important mediator of the effects of several motivational variables and learning strategies on academic achievement. interestingly, however, while van rooij and colleagues did not find evidence of self-efficacy being related with academic adjustment nor with academic achievement, the present study found that students’ self-efficacy of studying, especially in he academic programmes, to be positively associated with academic achievement both directly and indirectly, through academic adjustment. 8.2 differences between academic and professional programmes in line with previous work on disciplinary diversity of he programmes influencing interrelationships between non-cognitive variables and academic achievement (de clercq et al., 2013; fonteyne et al., 2017; lizzio et al., 2002), the present study provides empirical evidence that also he programme diversity (academic vs. professional) influences relationships in predictive models of academic achievement. the fact that the size of regression coefficients varies over academic and professional programmes, is an important first indication that learning strategies, motivational variables, and academic adjustment affect their respective outcomes differently in both he contexts. furthermore, one relationship between self-efficacy and academic adjustment was significant in academic programmes, but not in professional programmes. the value of investigating the diversity between he programmes on the meso-level, is also further evidenced by the differences in explained variance across both groups. firstly, learning strategies and motivational variables at the end of se seem to have more predictive power in the prediction of fyhe academic adjustment in the academic context (motivational variables: δr²=12.9%; learning strategies: δr²=19.4%) than in the professional context (motivational variables: δr²=9.3%; learning strategies: δr²=7.2%). in this light, especially lack of regulation and self-efficacy proved to have important differential effects in both contexts. secondly, our results suggest that academic adjustment in the first semester of he influences academic achievement to a bigger extent in professional programmes than in academic programmes. indeed, the incremental value of academic adjustment on academic achievement in terms of explained variance, was relatively larger in the professional programmes (motivational variables model: δr²= 9.4%; learning strategies model: δr²=9.8%), compared to the academic programmes (motivational variables model: δr²= 3.9%; learning strategies model: δr²= 6.7%). finally, after controlling for prior education, students’ learning strategies seem to have an identical predictive power regarding academic achievement in both he programmes (δr²=1.3%), while motivational variables and more specifically self-efficacy, predicts academic achievement to a slightly larger extent in academic versus professional programmes (acad.: δr²=5%; prof.: δr²=3.6%). 8.3 implications, limitations and perspectives the results of this study indicate that scholars investigating students’ transition into he should be attentive for the possible influences of the specific he programme or educational context wherein research is carried out. indeed, our study strengthens the idea that predictive models of academic achievement developed in academic programmes should not be imprudently applied in professional programmes considering that the predictive power of the included variables in this study varies across the two contexts. further research is needed to accurately establish (1) why learning strategies and motivational variables at the end of se seem to affect fyhe academic adjustment to a bigger extent in academic programmes than in professional programmes and (2) why academic adjustment seems to have more predictive power in the prediction of fyhe academic achievement in the professional context relative to the academic context. several theoretical models point at the importance of both the learning environment and student characteristics in the development of students’ learning processes (e.g. biggs, 1987), motivation (e.g. deci & ryan, 2000) and fyhe adjustment (e.g. tinto, 1975). in this regard, it should be noted that, although we did control for students’ prior education which is an important student characteristic, it remains unclear whether the unveiled differences in predictive power of the abovementioned determinants emerge from differences between the he programmes (contextual), or differences between students within the two systems (individual). with reference to the contextual characteristics, it remains unclear whether and how specific aspects of the learning environments under scrutiny (e.g. aims, expectations of students, assessment methods or didactical approaches) might have been moderating the relationships in the predictive model. possibly, compared to students in academic contexts, students in professional contexts might get a clear-cut indication of how well they are functioning, much earlier on in their programme (due to, for instance, different evaluation periods/feedback loops). this, then, would entail professional bachelor students to be able to make a more accurate assessment of their own academic adjustment, which could explain the larger effect of academic adjustment on achievement in professional contexts. another hypothesis we introduce here, is that there might be more demanding requirements related to regulation in academic he programmes, so that a higher level of lack of regulation at the end of se impacts academic adjustment to a larger extent in academic he programmes, relative to professional contexts. some limitations of the present study need to be highlighted. firstly, although this study adopts a longitudinal research design and, thus, enables further understanding of the directionality of effects, it does not allow for causal interpretation (cohen, manion, & morrison, 2011). the results and conclusions should therefore be interpreted cautiously. a second limitation concerns the adopted academic adjustment scale. this measure was developed in an academic context, and therefore, we cannot guarantee that this measure is a valid representation of academic adjustment in professional contexts. indeed, it is conceivable that the academic adjustment construct in professional programmes comprises sub-facets specific to that context, which are not accounted for in the present study. nonetheless, the items of the adopted scale are drawn up rather generally (another item example is: “i have the impression that i experience too many difficulties for my studies in higher education.”), and measurement invariance analyses have shown that the meaning of items of the adopted academic adjustment scale is equivalent for students from the two types of programmes (full scalar invariance was established). finally, although this study demonstrated the significance of diversity on the level of study programmes (academic vs. professional) in predicting academic adjustment and achievement, additional studies need to explore the effect of discipline diversity within the professional he context (becher, 1994). this would help us to establish a greater degree of accuracy and understanding of the matter at hand. based upon the results of this study, we suggest that (especially professional) he administrators should be attentive to the pivotal role of freshmen’s academic adjustment in the first semester of the fyhe. also, next to academic adjustment, coaching and guidance initiatives aimed at facilitating the academic transition in the fyhe should particularly target students’ self-efficacy, lack of regulation, and controlled motivation, which especially in academic he programmes – have considerable impact on academic adjustment. this is promising, as previous longitudinal research has shown that these latter variables are malleable, and not fixed pre-entry characteristics (e.g. vermunt & donche, 2017). keypoints the present study provides empirical evidence that he programme diversity (academic vs. professional he programmes) influences the relationships in predictive models of academic achievement. learning strategies and motivational variables at the end of se have more predictive power in the prediction of fyhe academic adjustment in academic he programmes, relative to professional he programmes. academic adjustment in the first semester of the fyhe influences academic achievement to a bigger extent in professional programmes than in academic programmes. these differences across he contexts were found after controlling for prior education, and adopting a longitudinal, integrative study design. references abbott-chapman, j. a., hughes, p. w., & wyld, c. (1992). monitoring student progress: a framework for improving student performance and reducing attrition in higher education. hobart: national clearinghouse for youth studies. bailey, t. h., & phillips, l. j. (2016). the influence of motivation and adaptation on students’ subjective wellbeing, meaning in life and academic performance . higher education research & development, 35(2), 201–216. https://doi.org/10.1080/07294360.2015.1087474 baker, s. r. (2004). intrinsic, extrinsic, and amotivational orientations: their role in university adjustment, stress, well-being, and subsequent academic performance. current psychology, 23(3), 189-202. https://doi.org/10.1007/s12144-004-1019-9 baker, r. w., mcneil, o. v., & siryk, b. (1985). expectation and reality in freshman adjustment to college. journal of counseling psychology, 32(1), 94. https://doi.org/10.1037/0022-0167.32.1.94 baker. r.w.. & siryk. b. (1999). student adaptation to college questionnaire (sacq): manual. los angeles: western psychological services. bean, j. p. (1980). dropouts and turnover: the synthesis and test of a causal model of student attrition. research in higher education, 12(2), 155-187. https://doi.org/10.1007/bf00976194 becher, t. (1994). the significance of disciplinary differences. studies in higher education, 19(2), 151-161. https://doi.org/10.1080/03075079412331382007 biggs, j. b. (1987). student approaches to learning and studying. research monograph. hawthorn: australian council for educational research ltd. braxton, j. m., & hargens, l. l. (1996). variation among disciplines: analytical frameworks and research. in j. c. smart (ed.),higher education: handbook of theory and research (vol. 11) (pp. 1-46). new york: agathon press. briggs, a. r., clark, j., & hall, i. (2012). building bridges: understanding student transition to university. quality in higher education, 18(1), 3-21. https://doi.org/10.1080/13538322.2011.614468 brown. t. a. (2014). confirmatory factor analysis for applied research. new york: guilford publications. byrne. b. m. (2016). structural equation modelling with amos: basic concepts. applications. and programmeming . new york: routledge. byrne. b. m.. shavelson. r. j.. & muthén. b. (1989). testing for the equivalence of factor covariance and mean structures: the issue of partial measurement invariance. psychological bulletin. 105(3). 456–466. https://doi.org/10.1037/0033-2909.105.3.456 camilleri, a. f., delplace, s., frankowicz, m., hudak, r., and tannhäuser, a. c. (2014). professional higher education in europe. characteristics, practice examples and national differences . malta: knowledge innovation centre. cazan. a. m. (2012). self-regulated learning strategies predictors of academic adjustment. procedia-social and behavioral sciences. 33(1). 104-108. https://doi.org/10.1016/j.sbspro.2012.01.092 chemers, m. m., hu, l. t., & garcia, b. f. (2001). academic self-efficacy and first year college student performance and adjustment. journal of educational psychology, 93(1), 55-64. https://doi.org/10.1037/0022-0663.93.1.55 chen. f. f. (2007). sensitivity of goodness of fit indexes to lack of measurement invariance. structural equation modeling. 14 (3). 464-504. https://doi.org/10.1080/10705510701301834 cheung. g. w. & rensvold. r. b. (2002). evaluating goodness-of-fit indexes for testing measurement invariance. structural equation modeling. 9(2). 233-255. https://doi.org/10.1207/s15328007sem0902_5 clark, m. h., middleton, s. c., nguyen, d., & zwick, l. k. (2014). mediating relationships between academic motivation, academic integration and academic performance. learning and individual differences, 33(1), 30-38. https://doi.org/10.1016/j.lindif.2014.04.007 cohen, l., manion, l., & morrison, k. (2011). research methods in education (6th edition). london: routledge. credé, m., & kuncel, n. r. (2008). study habits, skills, and attitudes: the third pillar supporting collegiate academic performance. perspectives on psychological science, 3(6), 425-453. https://doi.org/10.1111/j.1745-6924.2008.00089.x de clercq, m., galand, b., dupont, s., & frenay , m. (2013). achievement among first-year university students: an integrated and contextualised approach. european journal of psychology of education, 28(3), 641–662. https://doi.org/10.1007/s10212-012-0133-6 deci, e. l., & ryan, r. m. (2000). the" what" and" why" of goal pursuits: human needs and the self-determination of behavior. psychological inquiry, 11(4), 227-268. https://doi.org/10.1207/s15327965pli1104_01 declercq, k., & verboven, f. (2014). enrollment and degree completion in higher education without ex ante admission standards . leuven: faculty of economics and business. dent, a. l., & koenka, a. c. (2016). the relation between self-regulated learning and academic achievement across childhood and adolescence: a meta-analysis. educational psychology review, 28 (3), 425–474. https://doi.org/10.1007/s10648-015-9320-8 dinsmore, d. l., & alexander, p. a. (2012) . a critical discussion of deep and surface processing: what it means, how it is measured, the role of context, and model specification. educational psychology review, 24(4), 499-567. https://doi.org/10.1007/s10648-012-9198-7 donche, v., & van petegem, p. (2008). the validity and reliability of the short inventory of learning patterns. in e. cools, h. van den broeck, c. evans, & t. redmond (eds.), style and cultural differences: how can organisations, regions and countries take advantage of style differences (pp. 49-59). gent, belgium: vlerick leuven gent management school. donche, v., & van petegem, p. (2010). the relationship between entry characteristics, learning style and academic achievement of college freshmen. in m. poulson (ed.), higher education: teaching, internationalisation and student issues (pp. 277–288). new york: nova science publishers. enders, c. k. (2010). applied missing data analysis. london: guilford press. flemish government (2019). hoger onderwijs in cijfers [higher education in numbers]. retrieved 24 march, 2020, from https://onderwijs.vlaanderen.be/nl/hoger-onderwijs-in-cijfers . fonteyne, l., duyck, w., & de fruyt, f. (2017). program-specific prediction of academic achievement on the basis of cognitive and non-cognitive factors. learning and individual differences, 56(1), 34-48. https://doi.org/10.1016/j.lindif.2017.05.003 garriott, p. o., love, k. m., & tyler, k. m. (2008). anti-black racism, self-esteem, and the adjustment of white students in higher education. journal of diversity in higher education, 1(1), 45-58. https://doi.org/10.1037/1938-8926.1.1.45 gerdes, h., & mallinckrodt, b. (1994). emotional, social, and academic adjustment of college students: a longitudinal study of retention. journal of counseling & development, 72(3), 281-288. https://doi.org/10.1002/j.1556-6676.1994.tb00935.x glorieux, i., laurijssen, i., & sobczyk, o. (2014). de instroom in het hoger onderwijs van vlaanderen: een beschrijving van de huidige instroompopulatie en een analyse van de overgang van secundair onderwijs naar hoger onderwijs. [the inflow into higher education in flanders. a description of the current enrolment population and an analysis of the transition from secondary education to higher education]. leuven: steunpunt studieen schoolloopbanen. gregorich. s. e. (2006). do self-report instruments allow meaningful comparisons across diverse population groups? testing measurement invariance using the confirmatory factor analysis framework. medical care. 44(11). doi: 78-94. 10.1097/01.mlr.0000245454.12228.8f guay, f., ratelle, c., roy, a., & litalien, d. (2010). academic self-concept, autonomous academic motivation, and academic achievement: mediating and additive effects. learning and individual differences, 20(6), 644–653. https://doi.org/10.1016/j.lindif.2010.08.001 hu, l. t., & bentler, p. m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives.structural equation modeling: a multidisciplinary journal, 6(1), 1-55. https://doi.org/10.1080/10705519909540118 hurtado, s., han, j. c., sáenz, v. b., espinosa, l. l., cabrera, n. l., & cerna, o. s. (2007). predicting transition and adjustment to college: biomedical and behavioral science aspirants’ and minority students’ first year of college. research in higher education, 48(7), 841-887. https://doi.org/10.1007/s11162-007-9051-x iacobucci, d. (2010). structural equations modeling: fit indices, sample size, and advanced topics. journal of consumer psychology, 20(1), 90-98. https://doi.org/10.1016/j.jcps.2009.09.003 kuh, g. d., kinzie, j., buckley, j. a., bridges, b. k., & hayek, j. c. (2006). what matters to student success: a review of the literature. commissioned report for the national symposium on postsecondary student success. washington, dc: national postsecondary education cooperative. lizzio, a., wilson, k., & simons, r. (2002). university students' perceptions of the learning environment and academic outcomes: implications for theory and practice. studies in higher education, 27 (1), 27-52. https://doi.org/10.1080/03075070120099359 meredith. w. (1993). measurement invariance. factor analysis and factorial invariance. psychometrika. 58(4). 525-543. https://doi.org/10.1007/bf02294825 milfont, t. l., & fischer, r. (2010). testing measurement invariance across groups: applications in cross-cultural research. international journal of psychological research, 3(1), 111-130. https://doi.org/10.21500/20112084.857 muthén. l. k.. & muthén. b. o. (2010). mplus. statistical analysis with latent variables. user’s guide (6th ed.) . los angeles: muthén & muthén. oecd (2009), education at a glance 2009: oecd indicators, paris: oecd publishing. oecd (2013), education at a glance 2013: oecd indicators, paris: oecd publishing. petersen, i. h., louw, j., & dumont, k. (2009). adjustment to university and academic performance among disadvantaged students in south africa. educational psychology, 29(1), 99-115. https://doi.org/10.1080/01443410802521066 prospero,m., & vohra-gupta, s. (2007). first generation college students: motivation, integration, and academic achievement. community college journal of research and practice, 31 (12), 963–975. https://doi.org/10.1080/10668920600902051 rice, k. g., vergara, d. t., & aldea, m. a. (2006). cognitive-affective mediators of perfectionism and college student adjustment. personality and individual differences, 40(3), 463-473. https://doi.org/10.1016/j.paid.2005.05.011 richardson, m., abraham, c., & bond, r. (2012). psychological correlates of university students’ academic performance: a systematic review and meta-analysis. psychological bulletin, 138(2), 353–387. https://doi.org/10.1037/a0026838 robbins, s. b., lauver, k., le, h., davis, d., langley, r., & carlstrom, a. (2004). do psychosocial and study skill factors predict college outcomes? a meta-analysis. psychological bulletin, 130(2), 261–288. https://doi.org/10.1037/0033-2909.130.2.261 schuetze, h. g., & slowey, m. (2002). participation and exclusion: a comparative analysis of non-traditional students and lifelong learners in higher education. higher education, 44(3-4), 309-327. https://doi.org/10.1023/a:1019898114335 schunk, d. h. (1991). self-efficacy and academic motivation. educational psychologist, 26(3-4), 207-231. https://doi.org/10.1080/00461520.1991.9653133 schunk, d. h., & zimmerman, b. j. (2012). motivation and self-regulated learning: theory, research, and applications . new york: routledge. severiens, s., & wolff, r. (2008). a comparison of ethnic minority and majority students: social and academic integration, and quality of learning. studies in higher education, 33(3), 253-266. https://doi.org/10.1080/03075070802049194 steinmetz, h., schmidt, p., tina-booh, a., wieczorek, s., & schwartz, s. h. (2009). testing measurement invariance using multigroup cfa: differences between educational groups in human values measurement. quality & quantity, 43(4), 599-616. https://doi.org/10.1007/s11135-007-9143-x the bologna declaration (1999). the european higher education area: joint declaration of the european ministers of education, 19 june 1999, bologna. tinto, v., (1975). dropout from higher education: a theoretical synthesis of recent research. review of educational research, 45(1), 89-125. https://doi.org/10.3102/00346543045001089 torenbeek. m. (2011). hop. skip and jump? the fit between secondary school and university (phd diss.). university of groningen. groningen. torenbeek. m.. jansen. e.. & hofman. a. (2010). the effect of the fit between secondary and university education on first‐year student achievement. studies in higher education. 35(6). 659-675 . https://doi.org/10.1080/03075070903222625 van rooij, e., brouwer, j., fokkens-bruinsma, m., jansen, e., donche, v., & noyens, d. (2017). a systematic review of factors related to first-year students' success in dutch and flemish higher education. pedagogische studiën, 94(5), 360-405. van rooij, e. c., jansen, e. p., & van de grift, w. j. (2018). first-year university students’ academic success: the importance of academic adjustment. european journal of psychology of education, 33(4), 749-767. https://doi.org/10.1007/s10212-017-0347-8 vandenberg. r. j.. & lance. c. e. (2000). a review and synthesis of the measurement invariance literature: suggestions. practices. and recommendations for organizational research. organizational research methods. 3(1). 4-70. https://doi.org/10.1177/109442810031002 vanthournout, g., gijbels, d., coertjens, l., donche, v., & van petegem, p. (2012) . students’ persistence and academic success in a first-year professional bachelor programme: the influence of students’ learning strategies and academic motivation. education research international, 1–10. https://doi.org/10.1155/2012/152747 vermunt, j. d. (1998). the regulation of constructive learning processes. british journal of educational psychology, 68(2), 149–171. https://doi.org/10.1111/j.2044-8279.1998.tb01281.x vermunt, j. d. (2005). relations between student learning patterns and personal and contextual factors and academic performance. higher education, 49(3), 205-234. https://doi.org/10.1007/s10734-004-6664-2 vermunt, j. d., & donche, v. (2017). a learning patterns perspective on student learning in higher education: state of the art and moving forward. educational psychology review, 29(2), 269–299. https://doi.org/10.1007/s10648-017-9414-6 vermunt, j. d., & vermetten, y. j. (2004). patterns in student learning: relationships between learning strategies, conceptions of learning, and learning orientations. educational psychology review , 16(4), 359-384. https://doi.org/10.1007/s10648-004-0005-y wintre. m. g.. dilouya. b.. pancer. s. m.. pratt. m. w.. birnie-lefcovitch. s.. polivy. j.. & adams. g. (2011). academic achievement in first-year university: who maintains their high school average? higher education. 62(4). 467-481. https://doi.org/10.1007/s10734-010-9399-2 young, p. (2010). generic or discipline‐specific? an exploration of the significance of discipline‐specific issues in researching and developing teaching and learning in higher education. innovations in education and teaching international, 47(1), 115-124. https://doi.org/10.1080/14703290903525887 codepen introduction frontline learning research special issue vol.9 no.2 (2021) 18 issn 2295-3159 from micro to macro: widening the investigation of diversity in the transition to higher education mikaël de clercq1ellen jansen2taiga brahm3& elke bosse4 1université catholique de louvain, belgium 2university of groningen, the netherlands 3university of tuebingen, germany 4his institute for higher education development, germany info corresponding author doi: https://doi.org/10.14786/flr.v9i2.783 introduction transition into higher education (he) remains at the forefront of policy and practice in education worldwide (gale & parker, 2014). transition as a process (nicholson, 1990) in which individuals move from one stage to another may cause stress and discomfort that possibly lead to negative outcomes. transition into he is a particularly challenging process for the student due to a large variety of difficulties and requirements which could impede study success (trautwein & bosse, 2017). moreover, increasing student numbers and diversity in european he have reinforced concerns about study success in general and the successful transition to university in particular (abbott-chapmann, 2006, 2011; vossensteyn et al., 2015; wolter, 2013). consequently, it is important to further develop our understanding of factors that can contribute to a successful and less stressful transitions into higher education for a diverse student body. in this special issue, we go beyond considering individual factors, such as student characteristics (micro level). we go beyond student diversity, to investigate also the impact of the learning environment/ institution (meso level) and national educational policies (macro level). each study contributes to this endeavour by connecting two of the three levels of higher education. defining and framing transition to he the notion of transition lacks a clear-cut definition in the literature. for example, zittoun (2008, 2009) endorsed a developmental approach of transition and considered it as a brutal rupture implying major change in attitudes and behavior in order to adjust to a new environment. hutchinson’s (2005) life course perspective rather conceived transition as a continuous progressive change in individual status. gale and parker (2014) defined the transition as the « ability to navigate change » (p.4). finally, briggs, clark, and hall (2012) depicted the transition as an adjustment process to a major change in life. further extending these conceptions, we define transition into he as a process of instability and, thus, a rupture period, which leads to a qualitative evolution regarding students’ academic and social integration. transitions into he can have different transition situations: from secondary to higher education, from home country to abroad, from vocational/professional to university higher education. furthermore, the transitions can result in different realities, for example, students choose study programmes with varying aims; they study diverse disciplines, and are confronted with other students who each bring their own realities to campus. defining and framing diversity in he recent attempts to clarify the notion of diversity in he (balloo, 2018; bosse, 2015; de clercq, galand, & frenay, 2020; winstone & hulme, 2019) highlight that in the current literature diversity mostly encompassed students’ variability in background characteristics. in this special issue, we propose that diversity in he goes beyond students’ heterogeneity and can be extended to the variability in the features of the learning environment, as well as higher education institutions (tremblay, lalancette, & roseveare, 2012). in fact, the distinctive characteristics of the student in interaction with his or her learning context both determine if the transition will be unsettling and unproductive or a transformative experience leading to achievement and self-fulfilment (ecclestone, biesta, & hughes, 2010). in he research, diversity can therefore be defined as the sum of the interrelated differences in students’ backgrounds and in context characteristics that influence the process of transition into he and, consequently, study success. diversity can be viewed through different lenses, each focusing on a certain range of related factors that influence the process of transition into he. many studies use a three-level framework of he, distinguishing micro, meso and macro level approaches of the transition into he (enders, 2004; munge, thomas, & heck, 2018; taylor & ali, 2017; tremblay et al., 2012; vavoula & sharples, 2009). these three main levels of analyses will be considered as a guideline to extend our conception of diversity. to situate the different research perspectives, the notion of diversity will include differences at (1) the micro-level of the individual student experience, (2) the meso-level of the learning environment or institutional context and (3) the macro-level of the wider education system and global context. this calls for innovative methods that consider the different levels of investigation and possibly capture the relation between the levels. different approaches to diversity in the transition into he while the body of research on the transition to he is steadily growing (see for example the special issues or books by coertjens, brahm, trautwein, & lindblom-ylänne, 2017; jenert, postareff, brahm, & lindblom-ylänne, 2015; kyndt, donche, trigwell, & lindblom-ylänne, 2017), the focus on student diversity in the first-year of he can provide new perspectives, from a methodological point of view as well as content wise. most studies investigating the transition into he focus on one specific individual lens of diversity and do not address its interaction with the learning environment. considering different layers of differences together with advanced techniques (e.g. mixed method, social network analysis, multi-level analysis) promises perspectives on research of diversity. this widened analytical lens will constitute an important theoretical contribution for a better understanding of transition to he. the research up-to-now offers valuable insights into student diversity at the micro-level analysis of the transition into he. with respect to the students’ individual development, it examines student success correlates (e.g., motivation, learning strategies; for a review: richardson, abraham, & bond, 2012) and identifies relevant student profiles with distinct experiences of the he environment (de clercq, galand, & frenay, 2017; hailikari, tuononen, & parpala, 2016; van rooij, jansen, & van de grift, 2017; vanthournout, coertjens, gijbels, donche, & van petegem, 2013). he research also explores how non-traditional and international students experience their encounter with he (bathmaker & thomas, 2009; christie, tett, cree, hounsell, & mccune, 2008; holmegaard, madsen, & ulriksen, 2017; hope, 2017; jansen, suhre, & andré, 2017). by analysing the institutional contexts (e.g. the learning environment, study programmes), the research takes the meso-level of diversity in the transition into he into account (de clercq, galand, dupont, & frenay, 2013; jansen, 2004; schaeper, 2020). in fact, the respective studies widen the perspective on diversity by considering context diversity including organisational factors in order to investigate their impact on the transition to he and study success. these studies offer new insights into the role of diverse study contexts on the transition to he. to complement these perspectives on diversity, studies often consider organisational diversity in terms of types of he institutions (powell & solga, 2011; van de werfhorst, & mijs, 2010) or institutional characteristics (e.g., completion/entry rates, institutional structure) (chen, 2012; hauschildt, vögtle, & gwosć, 2018). such studies allow shedding light on inequalities and the structural barriers on the macro-level investigation of diversity in the transition into he (schuetze & slowey, 2002). the existing literature provides valuable insights, yet there are three main limitations in the field. • first, the insights into the role of diversity in the transition process remain scarce and underdeveloped in the literature (de clercq et al., 2020). for instance, it is not yet clear how certain diversity characteristics are related to particular transition experiences (winstone & hulme, 2019). further investigation of diversity as a student characteristic as well as a characteristic of the institutional context and the organisation is therefore needed. • second, the research approaches are fragmented due to the distinct perspectives on diversity at the different levels of the he system (bosse, 2015), resulting in difficulties to integrate findings from these different levels and to discuss practical implications (noyens, donche, coertjens, & van petegem, 2017). • third, it is not yet clear how new/advanced methodologies can address the analyses of the different levels of the he system not only one layer at a time but instead to integrate these multiple, nested layers. complex methods, such as multi-level analysis, mixed methods or social network analysis may therefore provide new insights. furthermore, transition is not a state but a process (tett, cree, & christie, 2017) and thus requires longitudinal approaches alongside analysis of personal and situational states to properly investigate the role of diversity (brouwer, jansen, krijnen, & warrens, 2021; noyens et al., 2017). connecting different levels of analysis of diversity in he to overcome the limitations of the existing literature in the field of he, this special issue aims at entering a novel pathway to investigate diversity to he by a) further investigating the role of diversity at different levels of he, b) connecting previously unconnected investigations of diversity in the transition into he at micro, meso and macro level and c) developing methodological designs that can tackle this broadened perspective of diversity in the transition into he. thus, the special issue does not only aim at widening the perspective on research concerning the role of diversity in transition processes, it also aims to contribute to further methodological developments in the field. as current research widely acknowledges the need to endorse a multidimensional approach in he research (coertjens et al., 2017; richardson et al., 2012), the special issue provides research that employs a methodology that goes beyond a single-factor analysis of the transition issue and uses cutting-edge data analysis to investigate different levels of diversity in the transition to he. the common aim of the studies composing the special issue is to connect either two or more levels (micro/meso/macro) in transition research or different types of diversity by using a range of methodologies (quantitative, qualitative, mixed-methods, and beyond). these studies therefore contribute collectively and individually to a widened understanding of diversity. each study included in this special issue is intended to provide a specific approach of the role of diversity in the transition into he by connecting different levels of diversity through a complex methodology. five studies of the special issue propose different ways of connecting micro and meso levels of analyses of diversity, with two studies looking into the characteristics of secondary schools in relation to student characteristics and students’ adjustment in the first year and three studies investigating factors at micro and meso level within the first year of higher education. together, these studies provide new insights into the role of diversity in the transition process and new perspectives to integrate these levels in a same approach. the study of van der zanden, denessen, cillessen, and meijer investigates longitudinally the relationships between teacher practices in secondary education (meso level of analysis) and first-year student academic achievement and social and emotional adjustment at university (micro level of investigation). results show that teachers in secondary education might play a pivotal role in preparing students for university. this role goes beyond preparing students for academic achievement, as teachers may have a long-term impact on first-year student social and emotional adjustment. willems, van daal, van petegem, coertjens and donche compare two distinct he contexts: professional and academic programs (meso level of analysis). they investigate how student’s psychosocial variables at the end of secondary education impact academic adjustment (micro level of analysis). the multigroup structural equation modelling demonstrates that learning strategies and motivational variables at the end of secondary education have higher predictive power in the academic context than in the professional context and that academic adjustment in the first semester influences academic achievement to a bigger extent in professional than in academic programs. these findings highlight that the role of individual characteristics for academic success differs from one context to another. an integration of micro and meso levels of analysis is proposed in the study of jenert and brahm by combining a person-centred quantitative and a longitudinal qualitative approach. while the quantitative study identifies three profiles of first-year students that demonstrate the individual diversity of the student body, the qualitative study finds very different reactions to the characteristics and events of the first-year environment among students of the three profiles. from a practical point of view, the findings show how different students perceive similar situations in distinctive and even contradictory ways. it emphasizes the need for more customized support structures during the first year of higher education that go beyond the usual distinction of traditional and non-traditional students. also, the study of bohndick, jänsch, bosse, and barnat endorsed a person-centred approach (micro level of analysis) but combined with structural equation modelling on the perception of institutional requirements (meso level of analysis). while latent profile analysis failed to demonstrate relevant diversity regarding the perception of institutional requirements, the structural equation modelling showed that these perceptions largely depend on the self-efficacy and volition. from a practical point of view, this study not only suggests to use the differences in the perception of requirements as a guideline for the design of support activities, but also to support the students’ self-efficacy and volition. the study of de clercq, hospel, galand and frenay compare the impact of student’s psychological variables (micro level of analysis) on academic adjustment among 21 different study programs (meso level of analysis). the multilevel analyses show significant variations of success rate between the programs. it highlights the diversity of the student body in the program. this means that the programs are not composed of the same student body. instead, how students perceive the characteristics of the context explains success variation. finally, engagement, motivation and social support do not have the same impact on academic adjustment from one program to another. from a practical point of view, the findings suggest that students’ experience of the transition to he widely differs from one program to another because these programs do not attract the same types of students and they do not provide the same learning context. students’ support therefore needs to be specific with regard to the characteristics of the program and to the type of students that are part of the program. the paper of balloo and winstone provides a methodological demonstration about how institutions can carry out nuanced analyses of their institutional data by combining a micro and a meso level of analysis on diversity. in four illustrative examples this primer provides tools which can empower university staff to perform such an investigation without requiring specialist software or expertise. the findings can inform the design of context-specific interventions that focus on reducing achievement gaps. in doing so, institutions can enhance the evidence-based understanding of potential reasons for differential study success during transition to he. dalhberg, vigmo and surian extend the scope of investigation by connecting the micro and macro level of analysis of diversity in the transition to he. they first compared two higher education institutions in sweden and italy (macro level of investigation) regarding their institutional policies to widened participation of migrant students. second, they investigate ethnographically generated student narratives regarding their individual transition to higher education from a migrant perspective (micro level of analysis). by combining these two perspectives, the authors show how policy ideas about widening participation and transition become visible in the students’ narratives and how they shape students’ experiences of participation, normalization, and marginalization in their own hei. this study highlights the need to follow-up on the impact of educational policies as it discovers the mismatch between the support provided to students and their needs and challenges when transitioning to he. as the whole is greater than the sum of the parts, every study carried out specific investigations that are constituting the complementary pieces of the final discussion of the special issue. the discussion article of this special issue connects the findings of the papers and relates them to the new zealand educational he context. the discussion paper of van der meer reopens the question of diversity toward the concept of “the whole student”. this concept can serve to take an “helicopter view” on each study contribution and to address the practical question of the “right” support to provide to a heterogeneous student body. the discussion also focuses on the methodological developments in the field of transitioning to he as well as on further developing the theoretical model of addressing diversity in he. keypoints contributes to developing an understanding of diversity as a multidimensional construct encompassing characteristics of students, learning environments, he institutions and he systems. provides a comprehensive theoretical framework as well as empirical findings that connect individual, context and organisational diversity factors of the transition into he. exploits and reflects on advanced methodologies (e.g. social network analysis, multi-level analysis, mixed methods etc) to address the different levels of diversity, always bearing in mind that transition processes into he are complex in nature, and therefore need an appropriate (multidimensional) methodology. responds to a new challenge in the field of he based on the current policy push to expand he participation and graduation by reflecting diversity in a comprehensive way. references abbott-chapmann, j. (2006). moving from technical and further education to university: an australian study of mature students. journal of vocational education & training, 58(1), 1–17. abbott-chapmann, j. (2011). making the most of the mosaic. facilitating post-school transitions to higher education of disadvantaged students. the australian educational researcher , 38(1), 57–71. balloo, k. (2018). in-depth profiles of the expectations of undergraduate students commencing university: a methodological analysis. studies in higher education , 43(12), 2251–2262. bathmaker, a. m., & thomas, w. (2009). positioning themselves: an exploration of the nature and meaning of transitions in the context of dual sector fe/he institutions in england. journal of further and higher education , 33(2), 119–130. bosse, e. (2015). exploring the role of student diversity for the first-year experience. zeitschrift für hochschulentwicklung , 10(4), 45–66. briggs, a. r. j., clark, j., & hall, i. (2012). building bridges. understanding student transition to university. quality in higher education, 18(1), 3–21. brouwer, j., jansen, e., krijnen, w., & warrens, m. (2021). longitudinal data analysis of emotions of first-year students in an international degree programme. in e. braun, r. esterhazy, & r. kordts-freudinger (eds.), research on teaching and learning in higher education (pp. 129–142). münster: waxmann. chen, r. (2012). institutional characteristics and college student dropout risks: a multilevel event history analysis. research in higher education , 53(5), 487–505. christie, h., tett, l., cree, v. e., hounsell, j., & mccune, v. (2008). ‘a real rollercoaster of confidence and emotions’: learning to be a university student. studies in higher education , 33(5), 567–581. coertjens, l., brahm, t., trautwein, c., & lindblom-ylänne, s. (2017). students’ transition into higher education from an international perspective. higher education , 73(3), 357–369. de clercq, m. d., galand, b., dupont, s., & frenay, m. (2013). achievement among first-year university students: an integrated and contextualised approach. european journal of psychology of education, 28(3), 641–662. de clercq, m., galand, b., & frenay, m. (2017). transition from high school to university: a person-centered approach to academic achievement. european journal of psychology of education, 32(1), 39–59. de clercq, m., galand, b., & frenay, m. (2020). one goal, different pathways: capturing diversity in processes leading to first-year students’ achievement. learning and individual differences, 81, 101908. ecclestone, k., biesta, g., & hughes, m. (2010). transitions and learning through the lifecourse . new york: routledge. enders, j. (2004). higher education, internationalisation, and the nation-state: recent developments and challenges to governance theory. higher education , 47(3), 361–382. gale, t., & parker, s. (2014). navigating change: a typology of student transition in higher education. studies in higher education , 39(5), 734–753. hailikari, t., tuononen, t., & parpala, a. (2016). students’ experiences of the factors affecting their study progress: differences in study profiles. journal of further and higher education , 42(1), 1–12. hauschildt, k., vögtle, e. m., & gwosć, c. (2018). social and economic conditions of student life in europe: eurostudent vi 2016-2018: synopsis of indicator . bielefeld: w. bertelsmann verlag. holmegaard, h. t., madsen, l. m., & ulriksen, l. (2017). why should european higher education care about the retention of non-traditional students? european educational research journal, 16(1), 3–11. hope, j. (2017). cutting rough diamonds’: the transition experiences first generation students in higher education. in e. kyndt, v. donche, k. trigwell, & s. lindblom-ylänne (eds.), higher education transitions: theory and research (pp. 85–99). london, new york: routledge. hutchison, e. (2005). the life course perspective: a promising approach for bridging the micro and macro worlds for social work. families in society: the journal of contemporary social services , 86(1), 143–152. jansen, e., suhre, c., & andré, s. (2017). transition to an international degree programme: preparedness, first-year experiences and study success of students from different nationalities. in e. kyndt, v. donche, k. trigwell, & s. lindblom-ylänne (eds.), higher education transitions: theory and research (pp. 47–65). london, new york: routledge. jansen, e. p. w. a. (2004). the influence of the curriculum organization on study progress in higher education. higher education , 47(4), 411–435. jenert, t., postareff, l., brahm, t., & lindblom-ylänne, s. (2015). editorial: enculturation and development of beginning students. zeitschrift für hochschulentwicklung, 10(4), 9–21. kovač, v. b. (2015). transition: a conceptual analysis and integrative model. in d. lansing cameron & r. thygesen (eds.), transitions in the field of special education: theoretical perspectives and implications for practice (pp. 19–34). münster: waxmann. kyndt, e., donche, v., trigwell, k., & lindblom-ylänne, s. (eds.) (2017). higher education transitions: theory and research. london, new york: routledge. munge, b., thomas, g., & heck, d. (2018). outdoor fieldwork in higher education: learning from multidisciplinary experience. journal of experiential education , 41(1), 39–53. nicholson, n. (1990). the transition cycle: causes, outcomes, processes and forms. in fischer, s., cooper, c. l. (ed.), on the move: the psychology of change and transition (pp. 83–108). hoboken, nj: wiley. noyens, d., donche, v., coertjens, l., & van petegem, p. (2017). transitions to higher education: moving beyond quantity. in e. kyndt, v. donche, k. trigwell, & s. lindblom-ylänne (eds.), higher education transitions: theory and research (pp. 3–13). london, new york: routledge. powell, j. j., & solga, h. (2011). why are higher education participation rates in germany so low? institutional barriers to higher education expansion. journal of education and work , 24(1-2), 49–68. richardson, m., abraham, c., & bond, r. (2012). psychological correlates of university students’ academic performance: a systematic review and meta-analysis. psychological bulletin , 138, 353–387. schaeper, h. (2020). the first year in higher education: the role of individual factors and the learning environment for academic integration. higher education , 79, 95–110. schuetze, h. g., & slowey, m. (2002). participation and exclusion: a comparative analysis of non-traditional students and lifelong learners in higher education. higher education , 44(3/4), 309–327. taylor, g., & ali, n. (2017). learning and living overseas: exploring factors that influence meaningful learning and assimilation: how international students adjust to studying in the uk from a socio-cultural perspective. education sciences , 7(1), 35. tett, l., cree, v. e., & christie, h. (2017). from further to higher education: transition as an on-going process. higher education , 73(3), 371–406. trautwein, c., & bosse, e. (2017). the first year in higher education critical requirements from the student perspective. higher education , 73, 371–387. tremblay, k., lalancette, d., & roseveare, d. (2012). assessment of higher education learning outcomes feasibility study report: volume 1 design and implementation. paris: oecd. retrieved from http://hdl.voced.edu.au/10707/241317 van de werfhorst, h. g., & mijs, j. j. (2010). achievement inequality and the institutional structure of educational systems: a comparative perspective. annual review of sociology , 36, 407–428. van rooij, e. c. m., jansen, e. p. w. a., & van de grift, w. j. c. m. (2017). secondary school students’ engagement profiles and their relationship with academic adjustment and achievement in university. learning and individual differences , 54, 9–19. vanthournout, g., coertjens, l., gijbels, d., donche, v., & van petegem, p. (2013). assessing students’ development in learning approaches according to initial learning profiles: a person-oriented perspective. studies in educational evaluation , 39, 33–40. vavoula, g., & sharples, m. (2009). meeting the challenges in evaluating mobile learning: a 3-level evaluation framework. international journal of mobile and blended learning (ijmbl) , 1(2), 54–75. vossensteyn, h., kottmann, a., jongbloed, b., kaiser, f., cremonini, l., stensaker, b., hovdhaugen, e., & wollscheid, s. (2015). dropout and completion in higher education in europe: main report. retrieved from http://doc.utwente.nl/98513/1/dropout-completion-he_en.pdf winstone, n. e., & hulme, j. a. (2019). ‘duck to water’ or ‘fish out of water’? diversity in the experience of negotiating the transition to university. in s. lygo-baker, i. m. kinchin, & n. e. winstone (eds.), engaging student voices in higher education: diverse perspectives and expectations in partnership (pp. 159–174). cham: springer international publishing. wolter, a. (2013). massification and diversity: has the expansion of higher education led to a changing composition of the student body?: european and german experiences. in i. repac & p. zaga (eds.), higher education reforms: looking back looking forward (pp. 202–220). berlin: peter lang. zittoun, t. (2008). learning through transitions: the role of institutions. european journal of psychology of educa , 23(2), 165–181. zittoun, t. (2009). dynamics of life-course transitions: a methodological reflection. in j. valsiner, m. p. c. molenaar, lyra, c. d. p. m., & n. chaudhary (eds.), dynamic process methodology in the social and developmental sciences (pp. 405–430). new york: springer us. codepen winne frontline learning research vol.8 no. 3 (2020) 164 173 issn 2295-3159 commentary: a proposed remedy for grievances about self-report methodologies philip h. winnea asimon fraser university, canada abstract this special issue’s editors invited discussion of three broad questions. slightly rephrased, they are: how well do self-report data represent theoretical constructs? how should analyses of data be conditioned by properties of self report data? in what ways do interpretations of self-report data shape interpretations of a study’s findings? to approach these issues, i first recap the kinds of self-report data gathered by researchers reporting in this special issue. with that background, i take up a fundamental question. what are self-report data? i foreshadow later critical analysis by listing facets i observe in operational definitions of self-report data: nature of the datum, topic, property, setting or context, response scale, and assumptions setting a stage for analyzing data. discussion of these issues leads to a proposal that ameliorates some of them: help respondents become better at self reporting. keywords: self-report data; likert scale; think aloud protocol info corresponding author email: winne@sfu.ca doi https://doi.org/10.14786/flr.v8i3.625 1. the landscape of self-reports represented in this special issue the most common forms of self-report data are surveys and think-aloud protocols. the former were popular among articles in this special issue. only one study used think aloud procedures. chauliac et al. (2020) asked participants to rate how frequently they applied various cognitive processes while studying textual information. durik and jenkins (2020) administered survey items calling for respondents to the degree to which they agreed with statements describing interest in several areas of study: astronomy, biology, math and psychology; and a single item inviting respondents to classify how certain they were about the set of ratings of interest. fryer and nakao (2020) investigated several judgments students made about a course. topics ranged over interest, meaningfulness and other dimensions. their focus was on possible differences arising due to different quantitatively described response formats – a conventional labeled categorical (likert) scale, slider, swipe and visual analog. iaconelli and wolters (2020) administered three surveys inquiring about respondents’ ratings of agreement and confidence about motivational constructs and self-regulated learning. moeller et al. (2020) gathered participants’ ratings of the degree to which statements about interest in, emotional investment in, and value of a course applied to them. rogiers et al. (2020) administered a survey instrument on which participants rated how much they agreed with statements describing their use of several cognitive and metacognitive processes. these researchers also gathered participant’s think aloud accounts about how participants studied an informative text. van halem, et al. (2020) administered the well-known motivated strategies for learning questionnaire which includes various subscales of likert items describing constructs within the arenas of motivation, cognition and metacognition. participants in vriesema and mccaslin’s (2020) study responded to a survey including likert response items about anxiety and selected from among a set of 20 sentences ones that described perceptions about participation in a small group activity. various features can more thoroughly discriminate the nature of self-reporting as a process and data these studies represent. in the next section, i propose a typology for these and other self-report methodologies. 2. facets of a self-report datum among facets of self-report data, primus inter pares (first among equals) is reliance on language. a self-report datum is a participant’s verbal utterance (think aloud) or recorded response to a spoken (interview) or written (survey, diary, experience sampling) invitation to describe a state or an event. the invitation to self report and the report itself are couched in language. no psychometric computation can remedy or precisely quantify indefiniteness arising from the dependence of self-report data on inherent elasticity and nuance of natural language. consequently, validity of interpretations grounded in self-report data is lessened in proportion to this source of unreliability. a variety of measures reported in this special issue and elsewhere in the literature that researchers intend to parallel or contrast to self-report data are not self-report measures according to this conceptualization. examples include: electrodermal signals, heart rate, various electromagnetic records of brain activity and data derived from tracking eye gear. topics described by self-report data range very widely. two main divisions are apparent: states and events. states may be internal, for example, a mood or physical condition. vriesema and mccaslin’s (2020) participants made a dichotomous decision whether their “stomach felt funny” or “head hurt” during a group activity. states also can be external, such as a characteristic of a learning situation or the availability of needed information. iaconelli and wolter’s (2020) participants rated desire for more time to complete schoolwork and finishing assignments right before deadlines. events are marked by changes internal to the respondent (increased anxiety, decreased certainty the effectiveness of a studying tactic) and in the environment (reviewing previously studied content, searching for information using an online search engine). students in rogiers et al. (2020) study described marking important words and rehearsing information. to the degree that language is indefinite, so, too, are self-report data. included within articles published in this special issue are: affective states and emotional reactions (e.g., interest, enjoyment, worry, pride), physiological states (stomach upset), behavioral events (summarizing content, copying notes, engagement with group members), and cognitive and metacognitive events (rehearsing content, judging extent and qualities of learning, planning). other studies range further afield. variance in the meaning of these features almost surely differs among participants within a study and across contexts. what kinds of tasks are considered schoolwork and which are not? what time interval between finishing an assignment and a deadline is “right before”? what makes words important? how do individuals' constructions of these concepts differ in ways that matter to the research questions each study investigated? properties of topics respondents self report about also vary. included in the research reported in this special issue are: frequency of an event, intensity of states, certainty of knowledge or about a future state, fit to one’s view of self, typicality of the event or state, and appropriateness of some behavior or feeling relative to a setting. for example, moeller et al (2020) experience sampling probes asked respondents to rate their understanding of, liking for, effort put toward learning, annoyance at learning, emotional cost to learn and future value of learning particular subject matter content. the wider literature adds to liberally to these topics. invitations to self report refer to a setting or context relative to which the report is forged. settings or contexts range over two dimensions. one dimension is whether a specific setting frames the self-report. examples are an experience in the immediate past, such as a just-completed session of collaborative work or a class period that has just finished. moeller and colleagues (2020) minimized the time interval with their experience sampling method, as did chauliac et al. (2020) by administering their survey after participants had just completed a studying task. the other end of this spectrum is a generalized setting, such as studying or life in school. some items in the motivated strategies for learning questionniare used in van halem et al (2020) study ask respondents to consider “a typical course.” durik and jenkins (2020) asked students to rate how strongly they had “always been fascinated with mathematics” and how much they were “really looking forward to learning more about mathematics.” the second dimension of the setting or context is whether that setting or context is one the respondent has personally experienced as opposed to one the respondent is asked to imagine. a prime example of a personally experienced context is the popular think aloud protocol like that used in the study by rogiers et al. (2020). participants work on a task and talk about what they think or do while they engage with it or, sometimes, retrospectively, quite soon after the participant disengages from the task. two variants of this latter case vary the delay between task completion and self report. under the methodology of experience sampling, like that in moeller et al.’s (2020) study, learners are notified at random points in time, commonly by a “beeper” or mobile phone notification, to write out or make an audio recording about an experience just completed. under a diary methodology, learners self report at periodic intervals, such as on the weekend. when respondents are asked to imagine a setting or context, they predict what would be the case if the experience actually happened. various response scales are used as metrics for self-report data. fryer and nakao’s (2020) study illustrates a direct investigation of this. in cases where respondents;’ utterances are analyzed by researchers who do not adopt an a priori classification, the researcher invents a categorical metric based on responses generated by one or multiple respondents. categorical bins of data are sometimes described as themes. in this case, respondents do not know when they respond how their self reports will be binned, so responses can not be biased by a response format. in other cases, respondents are fully aware how their self reports are “scored” because the response scale is used to provide a response. some scales call for selecting one option from a set of categories; sex, academic major and race are examples. other scales are ordinal, such as the extensively used likert scale. here, the response is expressed as a relative quantity, e.g., strongly agree/disagree, rarely/almost always, or very unlike/like me. when response scales are ordinal, an important decision the researcher makes is whether to permit the respondent to declare neutrality or indifference by offering an odd number of ratings versus an even number of ratings which typically precludes that option. some researchers ask for actual counts of events. i believe respondents cannot be accurate in this case unless their memory is perfect, a rare if not implausible quality. 3. assumptions, issues and complaints regarding self-report data i stipulate for sake of argument that respondents respond to invitations to self report to the best of their abilities. to borrow a common phrase from television shows about witnesses in american court cases, when respondents self report, they intend to “tell the truth, the whole truth and nothing but the truth.” likert scaled data almost always are analyzed using conventional arithmetic operations, such as forming a subscale by summing or averaging responses to several items. this was true of studies reported in this special issue. arithmetic operations require data have properties of a continuous interval scale. that is, units are differentiable (e.g., on a 100-point scale, 67 is different from 68) and differences between any adjacent pair of scale points is assumed to measure the same amount (e.g., the same difference is spanned between 67 and 68 and between 97 and 98). researchers usually sidestep this issue by reasoning this way – adding a large number of ordinal responses across separate items, say ratings of 1 to 5 over 10 items, elongates the scale so that it approximates a continuous scale with intervals of equal size. in my example, summing over 10 items increases the maximum scale length to 50 which affords a sufficient approximation to properties needed to use conventional arithmetic. this sleight of hand requires all the items represent one underlying dimension. to avoid the “apples and oranges” problem, the scale must be unidimensional to warrant adding responses. researchers attempt to test this requirement by computing a measure of internal consistency reliability or dimensionality. common approaches include computing cronbach’s alpha coefficient or doing a principal components or factor analysis. however, such computations put the cart before the horse. the assumption to be tested must hold to do these calculations because they require applying arithmetic operations to response values. the ploy is therefore tautological. as well, it often suffers important flaws (see liddell & kruschke, 2018). when respondents are invited to think aloud, write diary entries or otherwise generate free responses, all but a tiny fraction of studies bin instances by aggregating across respondents. rogiers et al. created 11 bins to differentiate text-learning strategies. implicit in this practice is a critical assumption, namely, variance across respondents and points-in-time/within-task are unimportant. this is analogous to the issue just discussed about likert data. as well, a tacit assumption rarely explicitly addressed by the researcher is there are no interactions among facets. that is, self reports in one bin do not correlate with others in a different bin and the kind of self report placed in one bin is not conditioned on another bin. for example, solicitations are independent of replies, recognition of an obstacle has no relation to solutions that overcome obstacles or admissions that obstacles cannot be overcome. my sample is small but i have never observed researchers test both assumptions. other assumptions, usually unstated and commonly untested, underlie researchers’ interpretations about bins of utterances or factors/components generated by a quantitative method. an utterance or survey item is assigned to one and only one bin, subscale or factor/component when the researcher’s analysis signals it can belong to multiple bins (fryer & nakao, 2020; rogiers et al., 2020). this practice simplifies analyses but may well oversimplify what respondents mean. when self reports that could be classified into multiple bins/components/factors are excluded from some of those containers and assigned to a single bin, respondents’ data are biased by that operational definition. all self report data become available because a researcher invites respondents to report. before issuing the invitation, the researcher cannot know whether information encapsulated in the requested report would have existed absent the researcher’s invitation. this raises a conundrum. some self reports may be outright fabrications. on being asked to report about something that was not in a respondent’s working memory or awareness, respondents may reply only to fulfill a social demand created by the researcher’s request. if this is the case, instrumentation creates a state or event that is reported rather than externalizing a report about a state or event that existed before the respondent was invited to describe it. consider this item from rogiers et al.’s (2020) scale about metacognitive monitoring: “i managed to learn the text in a good way.” the time point for a respondent to make this judgment is unspecified so a learner may honestly respond about monitoring during study or make the judgment when the survey is administered. valid interpretations about how a learner processes information will be challenged. this possibility raises a perplexing issue particularly in the context of agentic behavior like self-regulated learning (srl; winne, 2018). an agent’s behavior or experience can be altered as the agent becomes aware of characteristics of that behavior and experience. so, would a learner have considered the topic and properties of an experience or event if s/he had not been asked? in what cases and to what degrees are self reports possibly epiphenomenal? regarding think aloud reports, ericsson and simon (1993; see also fox et al., 2011) were acutely aware of the foregoing possibility. after scrutinizing research available at the time, they concluded: “with great consistency, this evidence demonstrates that verbal data are not in the least epiphenomenal but instead are highly pertinent to and informative about subjects” cognitive processes and memory structures” (p. 220). but they also note an important caveat. when subjects verbalize directly only the thoughts entering their attention as part of performing the task, the sequence of thoughts is not changed by the added instruction to think aloud. however, if subjects are also instructed to describe or explain their thoughts, additional thoughts and information have to be accessed to produce these auxiliary descriptions and explanations. as a result, the sequence of thoughts is changed, because the subjects must attend to information not normally needed to perform the task (p. xiii). rogiers et al. (2020) wisely engaged participants in a practice session to clarify the process of thinking aloud. they also adopted the common practice of prompting participants to “‘verbalize everything that you are doing or thinking’ or ‘keep thinking aloud’” when the researcher perceived there were “(a) meaningful silences or (b) certain nonverbal behaviours took place (i.e., frowning, repeatedly turning the text page, staring).” it might be argued this challenges ericsson and simon’s caveat. translating a state or event that has non-linguistic form into language involves considering what words are best to use. it is unclear whether words validly represent the learner’s experience and whether monitoring the qualities of that translation adds information into the cognitive arena not present before the learner was prompted to think aloud. every paper-and-pencil or computer-delivered questionnaire item asks respondents to describe properties of a thought, such as its generality, frequency or intensity. sometimes, respondents unintentionally make up answers. a relevant case arose in a study comparing self reports to logs of online behavior (winne & jamieson-noel, 1982). in this study, learners using software to study could view objectives for learning by clicking a button. the software logged this action if the learner clicked the button. after studying and taking an achievement test, we asked learners how helpful they found the objectives as guides to learning. several participants responded the objectives were helpful but the log of their data showed they never accessed the objectives. a more recent study reported less blatant but still important differences between online (logged) behaviors and self reports about those behaviors: “self-reports on prospective questionnaires show poor across method convergence with on-line thinking-aloud and observational data, obtained from students solving mathematics problems. these results are in line with earlier multi-method studies for reading and mathematics (see above). likewise, self-reports on the retrospective questionnaire do not converge with observational data” (veenman & van cleef, 2019, p. 698). a further complication arises when researchers gather self-report data to characterize srl. i model an elemental srl event as an if-then production. conditions (ifs) are the context in which a learner applies a particular cognitive operation to particular information (then). what was just addressed relates to thens, an action learners perform. on the other side of this model, every self report methodology specifies conditions, ifs, the context within which the learner is to reply. as noted earlier, these may be set out for the learner in general terms, e.g., “this course” as in the motivated strategies for learning questionnaire (see van halem et al., 2020) or “the lecture contents of the past couple of minutes” (moeller et al., 2020, p. 6). when a learner responds, it is reasonable to infer some particular features of a setting, ifs, influence the learner’s response. what are those features? is it reasonable to aggregate data across learners if there is variance in those features across learners? when the context is described for respondents as “this course,” to all learners characterize “the” course in sufficiently similar ways to warrant treating responses as if the conditions are the same? researchers lack data about what specifically each learner construes as ifs when responding to most survey items. assuming those conditions are the same when respondents have the same response, identical thens, is an instance of a logical fallacy, post hoc ergo prompter hoc. the same fallacy applies when researchers compute stability coefficients (test-retest) to characterize reliability of self-report data. a great number of studies using self-report data correlate those data with other, usually, outcome variables such as achievement or satisfaction with a long-term prior experience. in the studies published in this special issue and throughout the literature, language commonly used to express such results casts self report data as “accounting for” or “explaining” variance in another variable. these and analogous phrases implicitly but invalidly refer to the construct represented by self-report data as a cause. misleading phrasing about correlational findings that invites understanding them as causal is a common gaffe (robinson et al., 2007). this matter becomes even more muddled when statistical features of self-report data are not recognized. first, like all other data, self-report inherently have noise. some variance arises due to randomness and this attenuates the magnitude of relations. it might be suggested this challenge could be met by correcting for attenuation, but that operation actually muddles interpretation (see winne & belfry, 1982). second, it is often the case that self-report data are among other predictors used to predict an outcome. multiple regression and similar models are common statistical methods applied in this case, as was true for studies in this special issue (chauliac et al., 2020; durik & jenkins, 2020; van halem et al., 2020; moeller et al., 2020; vriesema & mccaslin, 2020). in these types of analysis, each predictor is residualized for every other predictor. interpretations of results of this analysis almost always fail to acknowledge the statistical output describes a residualized variable, not the original (winne, 1983). accurate phrasing relating to a beta coefficient such as, “self-reported elaboration residualized for the self-reported importance of other studying methods, sex of respondent and an indication of academic ability” appears in print as “elaboration … .” validity is even more strained by this oversight because it is not explicit to readers that the relation concerns a mathematically residualized construct that is not the same as the original construct. 4. conclusions the editors asked: how well do self-report data represent theoretical constructs? how should analyses of data be conditioned by properties of self-report data? in what ways do interpretations of self-report data shape interpretations of a study’s findings? “construct” is the pivotal word in the opening question of this trio. data, whether self reported and otherwise, are realized through operations a researcher designs and the correspondence between that operational definition and its implementation that realizes data. as implied by the title of gitelman’s (2013) edited volume, “raw data” is an oxymoron, data are not raw in the sense of lacking bias. theory is the muse that inspires gathering particular data in particular ways. from this perspective, all data inherently have some bias because they originate in a particular theory that conceptualizes constructs and characterizes forms for representing them as data. features of the physical, mathematical and psychological realms shape how researchers can obtain data sought to investigate a theory in ways shaped by multiple theories. the veracity with which realized self report data represent a particular construct depends on instrumentation writ large – instructions about how and when to respond, the setting in which a respondent reports, response format, and other values for facets described previously. this concern with generalizability was a core message of cronbach et al. (1972) views of generalizability, their extension to classical test theory’s notion of reliability. self-report data (as well as other forms of data) should be addressed not from a perspective of “the” reliability but a recognition of need to investigate generalizability. the fundamental question is: which facets and over what range of a facet’s values is unwanted (or uninterpretable) variance introduced into data (see winne, 2018)? in the context of self-report data, questions about facets and their relations to generalizability arise quickly. what is judged an “authentic” setting and what is not? who decides? do instructions cause a respondent to report memories as best they can be recalled or invent a false but plausible “memory” that fits the invitation to self report? was the protocol designed to invite a self report realized as designed (sometimes labelled fidelity of treatment implementation)? what facets of an operational design generate or suppress variance in the self-report data? replies to questions about how well data represent constructs will not be simple. while it is taxing, researchers should test facets of self-report data within their research. triangulation with non-self-report data is one approach, as illustrated in several articles appearing in this special issue. but triangulation is a hedge against issues of generalizability rather than an escape from them. identifying relevant facets is a precursor to improving opportunities to validly interpret data and analyses of data. analyses of self-report data can be improved. a key is to consider which laws of mathematics or, for categorical data, logic apply to self-report data as they are operationalized. the time to raise this question when designing research rather than after data have been gathered. researchers should more explicitly realize and describe in their publications the calculus applied to self-report data. mathematical manipulations of data are part of the data’s operational definition. this aspect of operational definitions have bearing on opportunities to validly interpret results of calculations. i illustrated this for cases where self-report data join other predictors in a multiple regression model. data residualized in such models are transformations of “data-at-the-first-instance” (a cumbersome label i hope may spark a good replacement for the more common label “raw data”). it will be helpful to readers if researchers remind them all components of operational definitions. an addendum to this recommendation is to entertain, when operational definitions of self-report data are being engineered, alternative analytical methods that might be applied to self-report data. for example, with likert-scaled data from surveys, what are trade-offs if items are examined using a principal components analysis versus a cluster analysis? two further considerations can be raised when analyzing self-report data in verbal forms, such as transcripts of think aloud data and replies to interview questions that are freely structured by the respondent. commonly, researchers labor to identify theoretically sensible bins or themes, then investigate interrater agreement before dividing up work on the corpus. at both stages, some, often large portions of data are discarded because those reports don’t fit bins. respondents, however, believed their accounts were relevant to the researcher’s invitation to report. disregarding these data represents another form of bias and this practice may mask or misconstrue what respondents deemed representative of their theory. second, bins of self-report data sometimes disregard temporal development and contingencies among bins. this may be particularly important in think-aloud data where unfolding episodic content is central to the respondent’s experience. self-report data should be investigated for contingencies and trajectory beyond statically binning segmented self-reports. the final question the editors posed asked about ways interpretations of self-report data shape interpretations of a study’s findings. at first blush, this might seem trivial if the word “findings” is taken to mean what appears in the discussion section of an article or chapter. those “findings” are interpretations of analyses of self-report (and other) data, so interpretations of self-report data are directly related to the study’s findings. a different perspective reflects what i observe more frequently in the research literature. a great deal of research that analyzes and interprets self-report data is carried out to investigate a research hypothesis. hypotheses are shaped in the first place by a theory the researcher chooses and uses to guide the investigation. as previously discussed, theoretical lenses shape decisions about what data merit collecting in the first place, instrumentation used to gather those data and warrants for features of analyses of data. it is important to keep in mind that theory sharpens some phenomena, blurs others and renders the rest invisible by classifying them as unimportant. my answer to the editors’ final question is that a study’s findings inevitably and substantially are shaped by a researcher’s interpretations about what self-report data are worth collecting. in other words, findings in the past shape theories that shape interpretations in the future. 5. coda science may someday develop instruments that accurately “read” human brain activity in a way that can reveal exactly what a person is thinking. current instruments – face readers, gaze trackers, and other physiological sensors – in my judgment, are not capable of that task. for the time being, learning science must rely partly on what people tell about their thoughts and feelings. when researchers collect self-report data, they depend on the respondent to “know thyself.” in less quaint terms, the respondent is a critical cog in a system that generates self-report data. i forecast learning science can better understand self report data and more prudently use them by seeking a fuller account about why knowing one’s self is difficult, and how people can more fully and more accurately come to know themselves. it follows that one approach to remedying some grievances i presented is to investigate how to help respondents – the key component within a system of instrumentation that develops self-report data –improve self reporting. acknowledgments foundations for this article were developed with financial support provided over many years by the social sciences and humanities council of canada and simon fraser university. references berger, j.-l., & karabenick, s. a. (2016). construct validity of self-reported metacognitive learning strategies. educational assessment, 21(1), 19-33. https://doi.org/10.1080/10627197.2015.1127751 cronbach, l. j., gleser, g. c., nanda, h., & rajaratnam, n. (1972). the dependability of behavioral measurements. new york: wiley. chauliac, m., catrysse, l., gijbels, d., & donche v. (2020). it is all in the surv-eye: can eye tracking data shed light on the internal consistency in self-report questionnaires on cognitive processing strategies? frontline learning research, 8(3), 26 – 39. https://doi.org/10.14786/flr.v8i3.489 durik, a. m., & jenkins j. s. (2020). variability in certainty of self-reported interest: implications for theory and research. frontline learning research, 8(3) 85-103. https://doi.org/10.14786/flr.v8i3.491 ericsson, k. a., & simon, h. a. (1993). protocol analysis: verbal reports as data (revised edition). mit press. fox, m. c., ericsson, k. a., & best, r. (2011). do procedures for verbal reporting of thinking have to be reactive? a meta-analysis and recommendations for best reporting methods. psychological bulletin , 137, 316-44. https://doi.org/10.1037/a0021663 fryer, l. k., & nakao k. (2020). the future of survey self-report: an experiment contrasting likert, vas, slide, and swipe touch interfaces. frontline learning research, 8(3),10-25. https://doi.org/10.14786/flr.v8i3.501 gitelman, l. (ed.). (2013). “raw data” is an oxymoron. the mit press: cambridge, ma. https://doi.org/10.7551/mitpress/9302.001.0001 van halem, n., van klaveren, c., drachsler h., schmitz, m., & cornelisz, i. (2020). tracking patterns in self-regulated learning using students’ self-reports and online trace data. frontline learning research, 8(3), 140-163. https://doi.org/10.14786/flr.v8i3.497 iaconelli, r., & wolters c.a. (2020). insufficient effort responding in surveys assessing self-regulated learning: nuisance or fatal flaw? frontline learning research, 8(3), 104 – 125. https://doi.org/10.14786/flr.v8i3.521 karabenick, s. a., woolley, m. e., friedel, j. m., ammon, b. v., blazevski, j., bonney, c. r., , de groot, e., gilbert, m. c., musu, l., kempler, t. m., & kelly, k. l. (2007). cognitive processing of self-report items in educational research: do they think what we mean? educational psychologist, 42, 139–151. https://doi.org/10.1080/00461520701416231 liddell, t. m., & kruschke, j. k. (2018). analyzing ordinal data with metric models: what could possibly go wrong? journal of experimental social psychology, 79, 328-348. https://doi.org/10.1016/j.jesp.2018.08.009 moeller, j., viljaranta, j., kracke, b., & dietrich, j. (2020). disentangling objective characteristics of learning situations from subjective perceptions thereof, using an experience sampling method design. frontline learning research, 8(3), 63-84. https://doi.org/10.14786/flr.v8i3.529 robinson, d. h., levin, j. r., thomas, g. d., pituch, k. a., & vaughn, s. (2007). the incidence of “causal” statements in teaching-and-learning research journals. american educational research journal, 44(2), 400-413 https://doi.org/10.3102/0002831207302174 rogiers, a., merchie, e., & van keer h. (2020). opening the black box of students’ text-learning processes: a process mining perspective. frontline learning research, 8(3), 40 – 62. https://doi.org/10.14786/flr.v8i3.527 veenman, m. v. j., & van cleef, d. (2018). measuring metacognitive skills for mathematics: students’ self-reports versus on-line assessment methods. zdm, 51(4), 691-701. https://doi.org/10.1007/s11858-018-1006-5 vriesema, c.c., & mccaslin, m. (2020) experience and meaning in small-group contexts: fusing observational and self-report data to capture self and other dynamics. frontline learning research, 8 (3), 126-139. https://doi.org/10.14786/flr.v8i3.493 winne, p. h. (1983). distortions of construct validity in multiple regression analysis. canadian journal of behavioural science, 15, 187-202. https://doi.org/10.1037/h0080736 winne, p. h., & belfry, m. j. (1982). interpretive problems when correcting for attenuation. journal of educational measurement , 19, 125-134. https://www.jstor.org/stable/1434905?seq=1#metadata_info_tab_contents winne, p. h., & jamieson-noel, d. l. (2002). exploring students’ calibration of self-reports about study tactics and achievement. contemporary educational psychology, 27, 551-572. https://doi.org/10.1016/s0361-476x(02)00006-1 winne, p. h. (2018). paradigmatic issues in state-of-the-art research using process data.frontline learning research 6, 250-258. https://doi.org/10.14786/flr.v6i3.551 codepen strohmaier et al frontline learning research vol.8 no. 1 (2020) 16 32 issn 2295-3159 a comparison of self-reports and electrodermal activity as indicators of mathematics state anxiety. an application of the control-value theory anselm r. strohmaiera, anja schiepe-tiskab & kristina m. reiss a b aheinz nixdorf chair of mathematics education, tum school of education, technical university of munich, munich, germany bcentre for international student assessment, technical university of munich, munich, germany article received 11 november 2018 / revised 13 december 2019/ accepted 24 january 2020 / available online 19 february abstract abstract: in the present study with 86 undergraduate students, we related trait mathematics anxiety (ma) with two indicators of state anxiety: self-reported state anxiety and electrodermal activity (eda). extending existing research, we included appraisals of control and perceived value in hierarchical multiple regression analyses in accordance with the control-value theory of achievement emotions (pekrun, 2006). results showed that trait ma predicted self-reported state anxiety, while no additional variance was explained by including control and value. in contrast, we found no significant relation between trait ma and physiological state anxiety, but a significant, negative three-way interaction effect with control and value. regression coefficients indicated that trait ma predicted physiological state anxiety, but only in the presence of negative perceived control and positive perceived value. thus, our results support the control-value theory for physiological state anxiety, but not for self-reports. they emphasize the need to distinguish between trait and state ma, the advantages of adopting the control-value theory, and the benefits of using eda recording as a supplemental assessment method for state anxiety. keywords: mathematics anxiety, electrodermal activity, galvanic skin response, control-value theory, state anxiety. info corresponding author anselm.strohmaier@tum.de doi 10.14786/flr.v8i1.427 1. introduction mathematics anxiety (ma) has a substantial impact on many students’ academic and personal lives. it influences achievement in mathematics tests and classes (hembree, 1990; ma, 1999; namkung, peng, & lin, 2019). moreover, students with high ma avoid mathematics in everyday life as well as in career and academic choices (dowker, sarkar, & looi, 2016; ma, 1999). ma is common across countries, cultures, and ages (dowker et al., 2016; lee, 2009). in the 2012 study of the programme for international student assessment (pisa), 30% of students reported that they felt helpless when doing a mathematic problem (oecd, 2013b). at the same time, ma is a problem of increasing relevance. on average across oecd countries, ma increased significantly from pisa 2003 to pisa 2012 (oecd, 2013b). thus, for educational research, it is important to understand how ma affects students when doing mathematics. research has elaborated the distinction between (momentary) state anxiety (mastate) and (habitual) trait mathematics anxiety (matrait), assessed through separate self-reports, but the findings left their relationship ambiguous (goetz, bieg, lüdtke, pekrun, & hall, 2013). hence, merely assessing matrait cannot exhaustively explain how ma affects mathematical activities momentarily. then again, directly assessing mastate provides a challenge, because self-reports of state emotions might be unreliable (pekrun & bühner, 2014). among other physiological measures, electrodermal activity (eda; also referred to as galvanic skin response; gsr) had sporadically been used as an indicator for mastate in the 1980s, but its relationship with self-reports of mastate or to matrait remained unclear. in this paper, we addressed this research gap by combining two novel approaches. first, we included and compared both self-reports and eda as measures of mastate. second, we used the control-value theory of achievement emotions (pekrun, 2006) as a framework to test their relation to matrait. accordingly, we included appraisals of control and perceived value as moderators of the relation between matrait and mastate. 1.2 mathematics anxiety ma “involves feelings of tension and anxiety that interfere with the manipulation of numbers and the solving of mathematical problems in a wide variety of ordinary life and academic situations” (richardson & suinn, 1972, p. 551). ma has an adverse effect on cognitive resources, independent of actual abilities (ashcraft, 2007; ashcraft & kirk, 2001; maloney et al., 2013). ashcraft and kirk (2001) found that in a mental addition task, undergraduates with high ma showed a smaller working memory capacity that led to an increase in reaction time and errors. this first finding started an intensive line of research, largely confirming direct effects of ma on performance (for overviews, see dowker et al., 2016; suárez-pellicioni, núñez-peña, & colomé, 2016). this influence is not limited to working memory capacity. for example, maloney, ansari, and fugelsang (2011) found that high ma students suffer from low-level numerical deficits, like a less precise representation of numerical magnitude. although most studies refer to ma as a unidimensional construct, a number of studies reported evidence that it consists of more than one factor, most prominently a cognitive component (“worry”) and an affective component (“emotionality”; e.g., ho et al., 2000; lukowski et al., 2016; wigfield & meece, 1988). these studies typically analyzed the factorial structure of questionnaires and related the dimensions to cognitive outcomes like mathematical achievement (e.g., ashcraft & ridley, 2005; lukowski et al., 2016). 1.3 trait and state mathematics anxiety while there are a large number of studies on ma, very few of them differentiate between mastate and matrait (goetz, bieg, lüdtke, pekrun, & hall, 2013; goldin, 2014). however, this distinction arguably is important when focusing on the effects of ma during mathematical activities. self-reports of matrait refer to multiple, generalized mathematical situations (bieg, goetz, wolter, & hall, 2015). in contrast, mastate refers to the specific, current situation. therefore, reports of matrait might be a good predictor for long-term effects of ma on learning or career and course choices (dowker et al., 2016) but do not necessarily accurately predict mastate during specific mathematical activities like tests or classes. when investigating the effects of ma during such activities, directly addressing mastate seems to be more appropriate. studies investigating the role of emotions in mathematics and of ma in particular predominantly focus on trait emotions rather than state emotions (goetz et al, 2013; goldin, 2014). accordingly, an extensive number of findings have been gathered on effects, individual differences, and precursors of matrait (dowker et al., 2016). in contrast, there are fewer studies on mastate, often using qualitative analyses (goldin, 2014). yet, specific mechanisms explaining the impact of mastate have rarely been reported for mathematics (dowker, 2016). the relationship between mastate and matrait is ambiguous. on the one hand, a number of studies indicate a strong positive relation. for high matrait students, mastate is considered a key explanation for a lower working memory capacity (ashcraft & moore, 2009; beilock, 2008). in his meta-analysis, hembree (1990) reports a mean correlation of r = .42 between matrait and state anxiety. however, state anxiety was not necessarily assessed during specific mathematical activities in the four reported studies (e.g. plake & parker, 1982). on the other hand, some studies indicate that there is a notable discrepancy between matrait and mastate. goetz et al. (2013) found that girls systematically report higher levels of matrait, but that this difference is not present in reports of mastate during mathematics tests or classes. this difference between reports of matrait and mastate is largely explained by individual beliefs and perceptions of competence (bieg et al., 2015; goetz et al., 2013). another reason for differences between matrait and mastate might be that ma negatively affects achievement in mathematics through long-term avoidance behavior, but not during mathematical activities per se (dowker et al., 2016): to avoid aversive consequences, mastate can even enhance motivation momentarily and lead to an increase in effort and strategy use during mathematics tests (eysenck & calvo, 1992; eysenck, derakshan, santos, & calvo, 2007). this indicates that matrait does not necessarily induce mastate. in general, state anxiety can have various cognitive and motivational-affective effects on learning and performance. zeidner (2014) lists 15 specific deficits in information processing during learning caused by anxiety, which are likely to be transferable to mastate. this includes cognitive deficits in areas like information encoding, information storage and processing, and information retrieval and production. moreover, state anxiety is associated with physiological reactions. however, this has not been described for mastate in particular, but only for state anxiety in general. per definition, state anxiety is a “transitory emotional state consisting of feelings of apprehension, nervousness, and physiological sequelae such as an increased heart rate or respiration” (wiedemann, 2015, p. 808). among other aspects, state anxiety is thus characterized by increased arousal and activation of the autonomic nervous system (steimer, 2002; wiedemann, 2015). accordingly, state anxiety does not only cause cognitive deficits, but also a physiological reaction. in sum, existing studies mostly focus on matrait, while its relation to mastate is left ambiguous. thus, to better understand how ma affects learning not only over a longer period of time but also momentarily, additional research is needed. this refers both to the question of the relation between matrait and mastate, as well as to the mechanisms and precursors of mastate in particular. in the following, we propose a theoretical framework for investigating these questions. 1.4 the control-value theory the control-value theory of achievement emotions (pekrun, 2006) characterizes predictors of achievement emotions, including state anxiety. it states that appraisals of control and the perceived subjective value of an achievement situation are the most proximal predictors of achievement emotions. a low appraisal of control and a simultaneous high perceived value of the task are key determinants of state anxiety. in contrast, trait emotions, environmental factors, or former achievement are considered distal factors and are assumed to have a mostly indirect effect on state emotions. according to the control-value theory, matrait should therefore predict mastate mostly indirectly, in association with low appraisals of control and a high subjective value. several empirical studies support aspects of the control-value theory in mathematics (e.g., niculescu, tempelaar, dailey-hebert, segers, & gijselaers, 2015). frenzel, pekrun, and goetz (2007) found that matrait is associated with a pattern of low competence beliefs paired with high achievement values in mathematics. extending the scope, research about attitudes and beliefs about competence in mathematics offers plenty of evidence supporting the control-value theory for other mathematical achievement emotions (for an overview, see goldin et al., 2016). however, to our knowledge, no study implemented both matrait and mastate as well as appraisals of control and perceived value in one model. 1.5 assessing mathematics state anxiety to assess mastate, research has mostly focused on qualitative research (see goldin, 2014, for an overview). these approaches included retrospective interviews and videotaping, but the reliability of these methods has been questioned (goldin, 2014). using a more quantitative approach, goetz et al. (2013) proposed short self-reports that could be used both for measuring anxiety during tests as well as during classes. the advantage of self-reports is that they can be used conveniently for experience-sampling and might be more reliable than observations. however, self-reports about achievement emotions might disrupt the current activity (goldin, 2014). moreover, it is questionable if self-reports can reflect an accurate evaluation of current emotions. in general, self-reports can only cover aspects of emotions that a person is aware of, depend on the use of language, and are subject to systematic biases, e.g. social desirability (pekrun & bühner, 2014). consequently, other researchers have attempted to use physiological measures to directly investigate ma in performance situations (dowker et al., 2016; hannula, 2016), predominantly using neuropsychological methods (e.g., lyons & beilock, 2012; pletzer, kronbichler, nuerk, & kerschbaum, 2015). these studies revealed that ma activates brain areas linked to fear processing, disgust and pain processing, but they did not distinguish between matrait and mastate (artemenko, daroczy, nuerk, 2015; suárez-pellicioni et al., 2016). state anxiety in general is associated with arousal and stress and with physiological reactions due to the activation of the autonomic nervous system. this leads to an increased heart rate and respiration, among other physiological reactions (steimer, 2002; wiedemann, 2015). this also holds for state anxiety in the context of education (zeidner, 2014). therefore, these specific physiological reactions can be assumed to be an indicator for mastate. some studies have assessed heart rate or cortisol secretion to monitor stress levels during mathematical tests (dew, galassi, & galassi, 1984; faust 1992, as cited in ashcraft, 2002; mattarella-micke, mateo, kozak, foster, & beilock, 2011; pletzer, wood, moeller, nuerk, & kerschbaum, 2010; sarkar, dowker, & cohen kadosh, 2014). these studies produced mixed results. mattarella-micke et al. (2011) showed that cortisol secretion can be associated with high performance (for low ma students) or with low performance (for high ma students), probably associated with a working memory overload. in contrast, pletzer et al. (2010) did not find a correlation between cortisol secretion and reports of ma, but used self-reports of matrait, not mastate. a relation between heart rate and state anxiety has been shown in various fields (e.g. kantor, endler, heslegrave, & kocovski, 2001), but has rarely been used in mathematics. faust (as cited in ashcraft, 2002) reported changes in heart rate when a highly math-anxious group performed mathematics tests of increasing difficulty. in contrast, dew, galassi, and galassi (1984) found no substantial relation between heart rate and matrait or mastate. in addition to heart rate, dew, galassi, and galassi (1984) observed physiological arousal during a timed mathematics test assessing participants’ eda. eda are fluctuations in skin conductance due to an increase in sweat gland activity. since sweat gland activity is associated with the autonomic nervous system activity, eda is an established method to assess physiological reactions to arousal, concerns, or stress (boucsein, 2012; naveteur & freixa i baqué, 1987; nikula, 1991). in his overview of the method, boucsein (2012) extensively reviewed applications and correlates of various measures of eda. he concludes that eda “can be regarded as a valid indicator for the strength of – mostly negative – emotions, for observing the course of psychological stress, and for objectively determining coping efficacy” (p. 521). therefore, eda can indicate state anxiety by detecting associated physiological reactions (boucsein, 2012). eda has recently been used to observe emotions during educational processes like self-regulated and multimedia learning (dindar et al., 2019; mudrick, taub, azevedo, price, & lester, 2017) and reading (meer, breznitz, & katzir, 2016). dew et al. (1984) used various measures of eda and different scales to assess matrait and mastate, but found no relation between eda and mastate, and only a small relation between eda and mastate for one of their measures of eda. as a possible explanation, they acknowledge that the challenge of comparing cognitively experienced anxiety and physiologically experienced anxiety might need a larger sample than their 31 students. moreover, their study design did not include a baseline measure, which is generally advisable for data quality (boucsein, 2012) and could indicate if eda is indeed influenced by a mathematical test context. thus, while their theoretical assumptions seem well-founded, the authors argue that their data was not sufficient for a meaningful interpretation (dew et al., 1984). in conclusion, mastate has been assessed through qualitative methods, self-reports, and physiological measures. physiological reactions are a vital aspect of anxiety in general and arguably of mastate in particular, but previous research has not provided clear results concerning the relation between self-reports and physiological measures of mastate, or the relation between mastate and matrait in general. 1.6 the present research so far, we have discussed that the relation between mastate and matrait is not yet fully understood. in performance situations, mastate might be stronger related to processes influencing mathematical thinking, like a reduction of working memory capacity. therefore, taking into account mastate seems important when analyzing effects of ma, but it can be assessed in different ways. while self-reports of mastate are easy to obtain, they might suffer from systematic biases. as an alternative, some studies used physiological measures of stress and arousal instead of self-reports to assess mastate in mathematical performance situations. yet, these studies did either not address both mastate and matrait or, in the case of dew et al. (1984), did not show clear results. moreover, no study did yet include appraisals of control or value to describe the relation between mastate and matrait in accordance with the control-value theory. we consider this a considerable gap in research on ma. we assume that the approach by dew and colleagues (1984) to use eda as an indicator for mastate is more promising today, because the possibilities to record and analyze eda have greatly improved. particularly, the innovations in eda recording offer better possibilities in observing the association between eda and matrait, since they allow to assess mastate more reliable and in an authentic environment. at the same time, using the control-value theory offers a better theoretical framework for the correlation between mastate and individual antecedents. it has been supported by a number of studies using self-reports and other methods to assess state anxiety, but to our knowledge, the control-value theory has not yet been utilized to analyze precursors of eda. 1.7 hypotheses in the present study, we investigated the relation of matrait with two indicators for mastate, the physiological measure eda and self-reported state anxiety. we assessed mastate both in a baseline context (a relaxation exercise) and a mathematics test. first, we assumed that the mathematics test would lead to an increase in both measures (hypothesis 1) and thus indicate that anxiety is successfully induced by the mathematics test. second, we assumed that there is a relation between self-reported mastate and eda (hypothesis 2). moreover, we anticipated that our findings would replicate the direct association between self-reported mastate and matrait (goetz et al., 2013; hypothesis 3a). we expected to find a similar relation between eda and matrait, since eda should reveal physiological arousal, which in turn is an indicator of mastate (hypothesis 3b). according to the control-value theory, appraisals of control and subjective value were included as predictors. we expected that this would confirm the relation between these appraisals and both measures of mastate (hypotheses 4a and 4b). finally, the relation between mastate and matrait should be higher when students report low control and high perceived value. thus, we expected a negative three-way interaction between matrait, appraisals of control, and perceived value, on both measures of mastate, respectively (hypotheses 5a and b). 2. method 2.1 sample and procedure 95 undergraduate students participated in the study. they gave written informed consent before participation. the study was conducted according to the ethical principles of psychologists and code of conduct of the american psychological association from 2017. an ethics approval was not required by institutional guidelines or national regulations, in line with the guidelines of the german research foundation. due to technical difficulties, 5 participants had to be excluded from the sample. additionally, we excluded 4 students because of deviations of more than 3 sd in one of the assessed measures. the remaining participants were 86 undergraduate students (53 female) from programs other than mathematics, ranging from engineering to nutritional science. mathematics students were not recruited as participants to avoid a bias in their beliefs and attitudes towards mathematics, as well as in their mathematical skills. the mean age was 23.2 years (sd = 4.07). participants were recruited on campus and were paid 15 eur for participation. during recruitment and before the experiment any indication of a mathematical content of the study was avoided. the study was described as a study investigating eda during various tasks. the individual sessions of the experiment took place in an office at the university containing only two tables, two chairs, and a closed closet. at the beginning of the experiment, the experimenter made participants familiar with the wristband assessing eda. she then put the device on the wrist of the participant’s non-dominant hand and fitted it comfortably. after recording had started, participants were presented a 5-minute relaxation exercise via headphones. the exercise facilitated relaxation through breathing exercises, accompanied by an audio track that included sounds from nature to help promote a relaxing environment for the participant. when the participant removed the headphones after the exercise, the experimenter immediately presented the first questionnaire assessing state-anxiety. after the participant finished the questionnaire, a first mathematical test was presented. the participant was asked to read the instruction carefully and then wait for the signal to start. all participants had 10 minutes to solve the test and received a short notice after 8 minutes. after the test, the participant answered the second state-questionnaire. the procedure was repeated for a second mathematics test. at the end of the experiment, trait and demographic data were assessed. 2.2 mathematics tests both mathematics tests consisted of six items. eleven items were taken from a pool of released items from the pisa-study (oecd, 2013a); one item was adopted from the trends in international mathematics and science study (timss, international association for the evaluation of educational achievement [iea], 2013). since research suggests that anxiety might have a larger influence for cognitively demanding tasks (ching, 2017; faust, ashcraft, & fleck, 1996), we composed both tests to be fairly difficult. the overall solution rate of 42% (sd = 21%) suggests that the tests were appropriately demanding. the items covered a broad range of mathematical problems, ranging from geometry to statistics. they were based on the concept of mathematical literacy and therefore covered mathematical competencies beyond mere factual knowledge. the tasks required knowledge that all students should have achieved by the end of their compulsory education. for an overall achievement score, we coded each item according to the coding instructions from pisa and timss (0 = incorrect, 0.5 = partially correct, 1 = correct; oecd, 2013a; iea, 2013) and calculated a sum score for all 12 items. 2.3 study measures we assessed matrait using the anxmat-scale developed for the pisa-studies (five items, e.g. “i feel helpless when doing a mathematics problem”, α = .87; oecd, 2005). participants answered on a 4-point likert scale from 1, strongly disagree to 4, strongly agree. we assessed self-reported mastate twice during the experiment according to goetz et al., 2013, asking if participants felt anxious in the previous situation (1, definitely not to 4, definitely). appraisals of control and perceived values were assessed after both tests and were task-specific. for appraisals of control, we used two items accounting for the controllability and probability of outcomes (e.g. “i think my competence in this area is …”, α = .78) on a 7-point likert-scale (1, low to 9, high; engeser & rheinberg, 2008; pekrun & perry, 2014). appraisals of perceived value were assessed with the four-item cognitive preferences-scale by kehr, von rosenstiel, and bles (1997) on a 7-point likert-scale (e.g. “it is important to me to solve the exercises”; 1, not at all to 9, very much; α = .85). for eda data collection during the relaxation exercise and the tests, we used an empatica e4 wristband. the wristband is worn like a watch and measures skin conductance with two stainless steel electrodes at the inner wrist. the exosomatic non-invasive sensor applies a very small, non-perceptible alternating current with a peak value of 100 μ a at 1v with an 8hz frequency. the 4 hz signal is recorded on an integrated flash memory. 2.4 eda data analyses eda signals consist of two components. the tonic signal is influenced by medium-term factors like room temperature or physiological characteristics of the individual. it provides a level of skin conductance that is rather stable within some seconds. even though the tonic signal can be an indicator for stress or anxiety, the phasic component of the signal is suited better to compare eda between individuals and is commonly used as an indicator for state anxiety (boucsein, 2012). phasic components of the eda signal are usually called responses, since they reflect a short peak in the signal. responses can be specific responses to a stimulation, for example a bursting balloon. however, there are phasic responses that are not associated to any specific external stimulation, hence nonspecific. the frequency of these nonspecific responses in skin conductance is associated with stress and anxiety and is one of the most common measures for eda (boucsein, 2012). the phasic and the tonic components of an eda signal overlap and need to be decomposed for analyses. data processing was carried out using matlab (v9.2.0) and the matlab-based software ledalab (v3.4.9). the software applies continuous decomposition analysis to extract the phasic signal (benedek & kaernbach, 2010). after the extraction, any peak in the phasic signal bigger than .01 μ s is counted as a response (boucsein, 2012). for both phases of the experiment (relaxation and test), the number of events is then summed up and divided by the duration of the phase in minutes. the result is the frequency of nonspecific skin conductance responses per minute (scr.freq). scr.freq served as the measure for physiological mastate. 2.5 analyses for hypothesis 1, we conducted a repeated measures anova to test for differences in state anxiety during the relaxation and the test. to assess the relation between the two measures of mastate and their relation to matrait (hypothesis 2 and 3), we calculated the correlations controlling for gender, achievement, and the respective baseline measures (see sect. 3.1). for hypotheses 4 and 5, we adopted a 5-step hierarchical multiple regression model for both measures of mastate as outcome variables (self-reported and physiological mastate). all predictors except gender were z-standardized before the analyses. in step 1, we included the control variables as predictors. in step 2, we additionally included matrait. in accordance with the control-value theory, step 3 included appraisals of control and subjective value. in step 4, we included the interaction term between control and subjective value. finally, step 5 included the interaction terms between matrait and appraisals of control and subjective value, respectively. additionally, we included the three-way interaction between matrait, control, and subjective value. 3. results 3.1 control variables gender differences exist between self-reports of ma (dowker et al., 2016). moreover, because of physiological differences in the sweat gland density and activity, women tend to display a weaker eda reactivity than men (boucsein, 2012). accordingly, our results revealed significant gender differences, with females showing weaker eda, t(84) = 2.93, p = .004, reporting higher matrait, t(84) = -2.38, p = .020, and lower control, t(84) = 2.37, p = .020. no significant gender differences were found regarding self-reports of mastate, t(84) = -0.49, p = .626, and perceived value t(84) = 0.02, p = .984. because of this general influence of gender, we included gender as a control variable in all following analyses. in addition, achievement is associated both with trait anxiety (ma, 1999) and with physiological reaction (mattarella-micke et al., 2011). in our data, we similarly found a significant relation between the test score and reports of matrait, r(86) = -.28, p = .008, self-reports of mastate, r(86) = -.31, p = .004, and control, r(86) = .51, p = .000, respectively, but no significant relation between the test score and eda, r(86) = .15, p = .177, and perceived value, r(86) = .07, p = .518, respectively. since our analyses focused on the interplay of matrait and mastate, irrespective of achievement, we also controlled for the test score in the following analyses. for both measures of mastate (self-reports and eda), we used the data from the relaxation exercise as respective baseline measures. 3.2 main analyses 3.2.1 descriptive results table 1 provides the means and standard deviations for matrait and appraisals of control and perceived value. additionally, mean scores and standard deviations for both measures of mastate during the relaxation exercise and the test are included. for both measures, mastate was significantly higher during the test compared to the relaxation exercise, confirming hypothesis 1. while physiological mastate increased from 15.43 events per minute to 20.04 events per minute (f(85) = 10.53, p = .002, 2 = .10), self-reported anxiety increased from 1.37 to 1.62 (f(85) = 17.23, p < .001, 2 = .17). table 1 descriptive statistics and differences between mastate in relaxation exercise and tests note. the unit for physiological state mathematics anxiety is scr.freq [1/min]. **p < .01 ***p < .001. 3.2.2 correlations table 2 provides correlations between all measures. all correlations were controlled for gender and test score and for the respective mastate baseline during the relaxation exercise. contrary to hypothesis 2, no significant correlation was observed between the two measures of mastate (r = .06, p = .63). matrait showed a moderate and significant correlation with self-reported mastate (r = .34, p = .002), but not with physiological mastate (r = .08, p = .48), which supports hypothesis 3a, but not 3b. including appraisals of control and perceived values, matrait correlated moderately and significantly with control and value (r = -.38, p < .001; r = .24, p = .029). a significant, moderate correlation emerged between appraisals of control and self-reported mastate (r = -.29, p = .008), but not physiological mastate (r = .05, p = .65). in contrast, appraisals of the perceived value were significantly related to physiological mastate (r = .29, p = .007), but not to self-reported mastate (r = .16, p = .15). appraisals of control and perceived value showed no significant relation (r = -.01, p = .95). table 2 correlations between measures of anxiety and appraisals of control and perceived value note. correlations of the two measures of state mathematics anxiety are controlled for their respective baseline. all correlations are controlled for gender and test score. n = 86. *p < .05 **p < .01 ***p < .001. 3.2.3 hierarchical multiple regression results of the hierarchical multiple regressions are reported in table 3. it displays only the predictors added in each step. for the full hierarchical models, see appendix a.1. for the two regressions, we used the two measures of mastate as outcome measures respectively. inclusion of the control variables explained 58% of the variance in physiological mastate during the test (p < .001), and 24% of the variance in self-reported mastate (p < .001). for self-reported mastate, step 2 revealed a significant relation between matrait and self-reported mastate ( β = 0.32, p = .002) that explained additional 9% of the variance in self-reported mastate (p = .002). step 3 did not confirm a relation between appraisals of control or perceived value and self-reported mastate ( β = -0.20, p = .090; β = 0.09, p = .37). step 4 did not reveal an interaction effect of control x value ( β = -0.00, p = .98), and step 5 revealed no three-way interaction effect of matrait x control x value ( β = -0.15, p = .20). similarly, the interaction effects matrait x control and ma x value were not significant ( β = -0.06, p = .59; β = -0.23, p = .053). these findings do not support hypothesis 4a or 5a. overall, the predictors explained 39% of the variance in self-reported mastate (p < .001). we conducted the same hierarchical multiple regression for physiological mastate. contrary to self-reported mastate, step 2 did not reveal a significant relation with matrait ( β = 0.06, p = .48). in step 3, adding appraisals of control and perceived value increased the r2 significantly by 4% (p = .033). in this step, perceived value had a significant positive relation with physiological mastate ( β = 0.19, p = .009), while no relation was found for control ( β = -0.04, p = .65). again, step 4 did not reveal an interaction of the control and value on mastate ( β = 0.04, p = .65). contrary to self-reported mastate, step 5 revealed a negative three-way interaction effect of matrait x control x value ( β = -0.23, p = .008), while the interaction effects matrait x control and matrait x value were not significant ( β = -0.06, p = .44; β = -0.01, p = .96). these effects explained an additional 4% of the variance in physiological mastate (p = .041). the three-way interaction effect is displayed in figure 1 (right). for comparison, figure 1 (left) displays the non-significant interaction for self-reported mastate. because there was no significant direct relation between matrait and physiological mastate, the slopes are less steep in figure 1 (right) than for self-reported mastate. however, it illustrates that the slope of matrait on physiological mastate increases for students appraising low control and high perceived value at the same time. these results support hypothesis 5b, but not hypothesis 4b. overall, the predictors explained 66% of the variance in physiological mastate (p < .001). table 3 hierarchical multiple regression analyses for physiological mastate and self-reported mastate figure 1. relation between matrait and mastate in dependence of control and perceived value. 4. discussion 4.1 measures of state anxiety in mathematics tests in line with hypothesis 1, we found significant differences between the relaxation exercise and the test for both measures of mastate. this implies that the mathematics test induced anxiety compared to the relaxation exercise. however, descriptive analyses showed that self-reports of mastate were relatively low in our study. this might be due to the fact that the experiment was a low-stakes test for the participants. we would assume that our result might emerge even stronger in a high-stakes test situation. in contrast to our hypothesis, students’ self-reports about mastate and their physiological mastate were not significantly related. our assumption had been that even though self-reports and physiological measures might differ to some extent, they should still refer to a similar mastate and hence be related. judging from our results, the two measures might refer to conceptually different aspects of mastate. some researchers suggest that matrait is a multidimensional construct, usually differentiating between a cognitive and an affective dimension (lukowski et al., 2016; wigfield & meece, 1988). similarly, physiological mastate and self-reported mastate as assessed in this study might refer to different facets of mastate. consequently, they might not necessarily be related. for example, eda might be more associated with arousal and an affective, emotional dimension of mastate. in contrast, self-reports might be more related to a cognitive dimension of mastate that is associated with worries and cognitive resources (ashcraft & moore, 2009; beilock, 2008; liebert & morris, 1967). future research could include a multi-dimensional assessment of ma to address this possibility. additionally, the measures might differ because of their differing mode of assessment (pekrun & bühner, 2014). self-reports might not be able to paint an adequate picture of achievement emotions, especially for a highly physiological emotion like anxiety (pekrun & bühner, 2014). furthermore, self-reports of mastate might be subject to biases like social desirability (pekrun & bühner, 2014) or stereotypes (goetz et al., 2013). 4.2 the relation between mastate and matrait in line with goetz et al. (2013), we found a significant relation between matrait and self-reported mastate which was within the range of previous findings reported by hembree (1990). students with higher matrait also reported higher mastate during a mathematical test. however, we did not find a relation between matrait and physiological mastate. this finding is contrary to hypothesis 3b but is in line with previous findings by dew et al. (1984). dew et al. (1984) proposed two explanations. first, the results might be viewed as questioning the construct validity of matrait scales. since these scales have been further validated since then and worked as expected with regard to self-reported mastate, this explanation seems unlikely. alternatively, since students reporting mastate need to evaluate their perceived anxiety cognitively, it is assumed that they might in part refer to generalized beliefs about mathematics. this might include the same resources as their evaluation of matrait (bieg, goetz, & lipnevich, 2014; goetz et al., 2013), or students might even refer directly to their matrait when trying to evaluate mastate. this would increase the relation between self-reported mastate and matrait, but not between physiological mastate and matrait. 4.3 the control-value theory according to the control-value theory (pekrun, 2006), mastate should primarily be determined by appraisals of control and perceived value. these appraisals should also moderate the relation between matrait and mastate. for the two measures of mastate, the application of the control-value theory in the present study produced diverging results. both matrait and self-reported mastate were related to appraisals of control. nevertheless, the hierarchical multiple regression did not produce signs that appraisals of control or value play an important role for the relation between matrait and self-reported mastate. rather, this relation seemed to be direct. hence, we did not find support for the control-value theory for self-reported mastate. as was proposed above, the relation between matrait and self-reported mastate might be increased by the similar mode of assessment. the resulting direct relation could overweight a possible indirect effect of appraisals of control and perceived value. in contrast, physiological mastate showed a different pattern. in opposition to self-reported mastate, we did not find a direct correlation with matrait. however, we found strong support for the control-value theory in this second multiple regression analysis. first, perceived value was related to mastate, independent of matrait. second, including appraisals of control and perceived value explained additional 8% of variance of physiological mastate, which indicates a substantial contribution to its emergence. third, the interplay between matrait, control, and value also was observed as expected. as illustrated in figure 1 (right), high matrait was related to high mastate, but only when students appraised their control low and their perceived value high. this effect is in line with the control-value theory, since matrait is considered a distal antecedent, whereas appraisals of control and perceived value are considered proximal causes of mastate. these results further support the notion that the causal relation between matrait and the two measures of mastate might be conceptually different. 4.4 limitations using eda comes with some immanent limitations, and only some of them can be overcome. for example, eda is subject to physiological gender differences. this inhibits its practicality for inquiring the gender gap in ma. even when controlling for a baseline value, differences in reactivity exist. in general, a large variance between students’ eda makes comparisons more difficult. in our study, we assessed the baseline value during a relatively short period of time. a more reliable value could be obtained through several hours or days of baseline recordings (boucsein, 2012). of course, such a study requires much more time. lastly, even though ma is common in students of all ages, our specific sample cannot be overgeneralized. it needs to be verified if eda recording can be useful in schools and for specific groups of students, for example high-anxiety students or younger students. more generally, our test did not seem to induce a very strong emotional reaction. in order to generalize our findings to high-stakes testing which might cause more mastate, additional studies are needed. moreover, we followed goetz at al. (2013) in using a single-item scale to assess self-reported mastate. this keeps the disruption of the participants at a minimum but might result in some inaccuracies. our results indicate that the scale was working properly, but future studies might try to assess mastate at more occasions or check if the one-item scale is appropriately precise. similarly, a number of different questionnaires exist to assess ma and general test anxiety. comparing these questionnaires regarding their relation to eda, particularly regarding cognitive and affective dimensions of these scales, could help to explain the absent relation between self-reported and physiological mastate. the relation between eda and physiological arousal has been well established by previous research (boucsein, 2012). however, other factors than mastate might additionally influence eda during mathematics tests. future research could incorporate additional state measures that assess cognitive load or situational motivation to further narrow down the processes associated with eda reactivity, and might support these findings through qualitative data like interviews or think-aloud-protocols. until the validity of eda as a measure of physiological mastate is fully understood, results will always require a cautious discussion of limitations and different explanations. the control-value theory is generalizable to various achievement emotions, including both trait and state emotions (pekrun, 2006). in the current cross-sectional study, we focused on mastate as an outcome, and task-specific appraisals of control and perceived value as moderators. consequently, we considered matrait as a distal predictor. however, future studies could also consider matrait as an outcome itself. for analyzing effects of general appraisals of control and perceived value towards mathematics as predictors of matrait, longitudinal designs would be more advantageous. 4.5 conclusion our study combined several innovative approaches that have emerged in research on ma within the last years. with the distinction between matrait and mastate, we differentiated between two different facets of ma. further, through the adoption of the control-value theory, we compared eda recordings and common self-reports as a tool for observing mastate and investigated their unique relations to matrait. overall, we found that eda was related to matrait, but that this relation only got visible when taking appraisals of control and perceived value into account. students reporting high matrait were not necessarily more physiologically anxious during mathematical activities. rather, a pattern of appraisals of low control and high perceived value accompanied that relation. hence, with regard to eda, our results were in line with the control-value theory, which on the other hand was not supported by self-reported measures of mastate. in sum, our findings match the plea by goetz at al. (2013) to consequently distinguish between matrait and mastate in research on ma, as well as to additionally include physiological data in assessing emotions in learning. furthermore, our results indicate that self-reports and physiological measures might refer to different aspects of mastate. thus, our results support theoretical considerations and empirical findings that self-reports of mastate should be interpreted cautiously. ultimately, we cannot decide if self-reports or eda captured actual mastate. rather, the two measures both seem to be related to matrait, but in different ways. therefore, we cannot conclude that eda can make self-reports obsolete, but we propose that the assessment of eda can provide additional information about underlying affective aspects of mastate. because of recent technical advances in recording and analyses of eda, the method seems to offer a convenient addition to the common practice of self-reports. furthermore, the advances in eda-recording offer the possibility to conduct studies in the classroom during regular classes with hardly any disruption. we believe that our study can be a first step into this promising direction of in vivo research on ma. as a next step, the relation to mathematical achievement should be investigated. in the recent study, we used a mathematics test, the goal of which was to trigger ma, but that was not designed to diagnose mathematical achievement in detail. our preliminary results indicate that achievement might be differently associated with self-reported and physiological mastate, but a study that assesses mathematical performance in more detail is needed to shed light on this question. additionally, achievement under conditions of mastate and no mastate should be compared in a within-subject design, since eda shows a notable variance between subjects. similarly, using tests that are not mathematical could help to distinguish how specific mastate is linked to mathematics. at the same time, using eda for other domains or test anxiety in general, possibly using the control-value theory, might be an interesting and fruitful perspective for future research. lastly, the relation to working memory capacity, which has proven to be a key factor in the effects of ma, should be taken into account. ultimately, this knowledge could be used to design longitudinal and intervention studies that use eda to observe the role of mastate for learning processes or create ways to decrease mastate in mathematics tests, possibly without necessarily tackling matrait. with a number of questions remaining unanswered, our study is merely a first step in including eda as an indicator for mastate. nevertheless, the results illustrate that self-reports only comprise one perspective on the multi-faceted phenomenon of mathematics anxiety, and that including eda can be uniquely insightful. keypoints we did not find a correlation between eda and measures of state anxiety or trait mathematics anxiety, respectively. self-reported state anxiety correlated significantly with trait anxiety independent of appraisals of control and perceived value, which is in contrast to the control-value theory. in line with the control-value theory, trait mathematics anxiety predicted physiological state anxiety when high perceived value and low control of the achievement situation were reported. acknowledgments we would like to thank ashley l. johnson and kathrin ebenhöh for their contributions during data collection. this research was funded by the federal ministry of education and research (bmbf) and the standing conference of the ministers of education and cultural affairs of the länder in the federal republic of germany (kmk) [grant number zib2016]. references artemenko, c., daroczy, g., nuerk, h.-c. (2015). neural correlates of math anxiety an overview and implications. frontiers in psychology 6, 1333. doi:10.3389/fpsyg.2015.01333 ashcraft, m. h. (2002). math anxiety: personal, educational, and cognitive consequences. current directions in psychological science 11(5), 181-185. doi:10.111/1467-8721.00196 ashcraft, m. h. (2007). is math anxiety a mathematical learning disability? in d. b. berch & m. m. m. mazzocco (eds.), why is math so hard for some children? the nature and origins of mathematical learning difficulties and disabilities (pp. 329-348). baltimore: brookes. ashcraft, m. h., & kirk, e. p. (2001). the relationships among working memory, math anxiety, and performance. journal of experimental psychology: general, 130(2), 224–237. doi:10.1037//0096-3445.130.2.224 ashcraft, m. h., & moore, a. m. (2009). mathematics anxiety and the affective drop in performance. journal of psychoeducational assessment, 27(3), 197-205. doi:10.1177/0734282908330580 ashcraft, m. h., & ridley, k. s. (2005). math anxiety and its cognitive consequences: a tutorial review in j. i. d. campbell (ed.), handbook of mathematical cognition (pp. 315-327). new york, ny: psychology press. beilock, s. l. (2008). math performance in stressfull situations. current directions in psychological science, 17(5), 339–343. doi:10.1111/j.1467-8721.2008.00602.x benedek, m., & kaernbach, c. (2010). a continuous measure of phasic electrodermal activity. journal of neuroscience methods, 190(1), 80-91. doi:10.1016/j.jneumeth.2010.04.028 bieg, m., goetz, t., & lipnevich, a. a. (2014). what students think they feel differs from what they really feel academic self-concept moderates the discrepancy between students’ trait and state emotional self-reports. plos one, 9(3). doi:10.1371/journal.pone.0092563 bieg, m., goetz, t., wolter, i., & hall, n. c. (2015). gender stereotype endorsment differentially predicts girls’ and boys’ trait-state discrepancy in math anxiety. froniers in psychology, 6(1404). doi:10.3389/fpsyg.2015.01404 boucsein, w. (2012). electrodermal activity (2 ed.). new york: springer. ching, b. h.-h. (2017). mathematics anxiety and working memory: longitudinal associations with mathematical performance in chinese children. contemporary educational psychology, 51, 99-113. doi:10.1016/j.cedpsych.2017.06.006 dew, k. m. h., galassi, j. p., & galassi, m. d. (1984). math anxiety: relation with situational test anxiety, performance, physiological arousal, and math avoidance behaviour. journal of counseling psychology, 31 (4), 580-583. doi:10.1037/0022-0167.31.4.580 dowker, a., sarkar, a., & looi, c. y. (2016). mathematics anxiety: what have we learned in 60 years? frontiers in psychology, 7, 508. doi:10.3389/fpsyg.2016.00508 dindar, m., malmberg, j., järvelä, s., haataja, e., & kirschner, p. a. (2019). matching self-reports with electrodermal activity data: investigating temporal changes in self-regulated learning. education and information technologies. doi:10.1007/s10639-019-10059-5 engeser, s., & rheinberg, f. (2008). flow, performance and moderators of challenge-skill balance. motivation and emotion, 32(3), 158-172. doi:10.1007/s11031-008-9102-4 eysenck, m. w., & calvo, m. g. (1992). anxiety and performance: the processing efficiency theory. cognition and emotion, 6(6), 409. doi:10.1080/02699939208409696 eysenck, m. w., derakshan, n., santos, r., & calvo, m. g. (2007). anxiety and cognitive performance: attentional control theory. emotion, 7(2), 336-353. doi:10.1037/1528-3542.7.2.336 faust, m. w., ashcraft, m. h., & fleck, d. e. (1996). mathematics anxiety effects in simple and complex addition. mathematical cognition, 2, 25-62. doi:10.1080/135467996387534 frenzel, a. c., pekrun, r., & goetz, t. (2007). girls and mathematics a „hopeless“ issue? a control-value approach to gender differences in emotions towards mathematics. european journal of psychology of education, 22(4), 497-514. doi:10.1007/bf03173468 goetz, t., bieg, m., lüdtke, o., pekrun, r., & hall, n. c. (2013). do girls really experience more anxiety in mathematics? psychological science, 24(10), 2079–2087. doi:10.1177/0956797613486989 goldin, g. a. (2014). perspectives on emotion in mathematical engagement, learning, and problem solving. in r. pekrun & l. linnenbrink-garcia (eds.), international handbook of emotions in education (pp. 391-414). new york: routledge. goldin, g. a., hannula, m. s., heyd-metzuyanim, e., jansen, a., kaasila, r., lutovac, s., . . . zhang, q. (eds.). (2016). attitudes, beliefs, motivation and identity in mathematics education. an overview of the field and future directions . hamburg: springer open. hannula, m. s. (2016). introduction. in g. a. goldin, m. s. hannula, e. heyd-metzuyanim, a. jansen, r. kaasila, s. lutovac, p. di martino, f. morselli, j. a. middleton, m. pantziara, & q. zhang (eds.), attitudes, beliefs, motivation and identity in mathematics education. an overview of the field and future directions (pp. 1-2). hamburg: springer open. hembree, r. (1990). the nature, effects and relief of mathematics anxiety. journal for research in mathematics education, 21. doi:10.2307/749455 ho, h.-z., senturk, d., lam, a. g., zimmer, j. m., hong, s., okamoto, y., . . . wang, c.-p. (2000). the affective and cognitive dimensions of math anxiety: a cross-national study. journal for research in mathematics education, 31(3), 362-379. doi:10.2307/749811 international association for the evaluation of educational achievement [iea] (2013). timss 2011 assessment. chestnut hill, ma: timss & pirls international study center. kantor, l., endler, n. s., heslegrave, r. j., kocovski, n. l. (2001). validating self-report measures of state and trait anxiety against a physiological measure. current psychology 20(3), 207-215. doi: 10.1007/s12144-001-1007-2 kehr, h. m., von rosenstiel, l., & bles, p. (1997). zielbindung, subjektive fähigkeiten und intrinsische motivation [goal commitment, subjective abilities, and intrinsic motivation]. paper presented at the 16th colloquium of motivational psychology, potsdam, germany. lee, j. (2009). universals and specifics of math self-concept, math self-efficacy, and math anxiety across 41 pisa 2003 participating countries. learning and individual differences, 19(3), 355–365. doi:10.1016/j.lindif.2008.10.009 liebert, r. m., & morris, l. w. (1967). cognitive and emotional components of test anxiety: a distinction and some initial data. psychological reports, 20(3), 975-978. doi:10.2466/pr0.1967.20.3.975 lukowski, s. l., ditrapani, j., jeon, m., wang, z., schenker, v. j., doran, m. m., . . . petrill, s. a. (2016). multidimensionality in the measurement of math-specific anxiety and its relationship with mathematical performance. learning and individual differences. doi:10.1016/j.lindif.2016.07.007 lyons, i. m., & beilock, s. l. (2012). mathematics anxiety: separating the math from the anxiety. cerebral cortex, 22, 2102-2110. doi:10.1093/cercor/bhr289 ma, x. (1999). a meta-analysis of the relationship between anxiety toward mathematics and achievement in mathematics. journal for research in mathematics education, 30(5), 520-540. doi:10.2307/749772 maloney, e. a., ansari, d., & fugelsang, j. a. (2011). the effect of mathematics anxiety on the processing of numerical magnitude. quarterly journal of experimental psychology, 64(1), 10-16. doi:10.1080/17470218.2010.533278 maloney, e. a., schaeffer, m. w., & beilock, s. l. (2013). mathematics anxiety and stereotype threat: shared mechanisms, negative consequences and promising interventions. research in mathematics education, 15(2), 115–128. doi:10.1080/14794802.2013.797744 mattarella-micke, a., mateo, j., kozak, m. n., foster, k., & beilock, s. l. (2011). choke or thrive? the relation between salvary cortisol and math performance depends on individual differences in working memory and math-anxiety. emotion, 11(4), 1000-1005. doi:10.1037/a0023224 meer, y., breznitz, z., & katzir, t. (2016). calibration of self-reports of anxiety and physiological measures of anxiety while reading in adults with and without readig disability. dyslexia, 22, 267-284. doi:10.1002/dys.1532 mudrick, n. v., taub, m., azevedo, r., price, m. j., & lester, j. (2017). can physiology indicate cognitive, affective, metacognitive, and motivational self-regulated learning processes during multimedia learning? paper presented at the annual meeting of the american educational research association (aera), san antonio, tx. namkung, j. m., peng, p., & lin, x. (2019). the relation between mathematics anxiety and mathematics performance among school-aged students: a meta-analysis. review of educational research, 89(3), 459–496. doi:10.3102/0034654319843494 naveteur, j., & freixa i baqué, e. (1987). individual differences in electrodermal activity as a function of subjects’ anxiety. personality and individual differences, 8(5), 615-626. doi:10.1016/0191-8869(87)90059-6 niculescu, a. c., tempelaar, d., dailey-hebert, a., segers, m., gijselaers, w. (2015). exploring the antecedents of learning-related emotions and their relations with achievement outcomes. frontline learning research 3 (1), 1-17. doi:10.14786/flr.v3i1.136 nikula, r. (1991). psychological correlates of nonspecific skin conductance responses. psychophysiology, 28(1), 86-90. doi:10.1111/j.1469-8986.1991.tb03392.x ng, e. l., & lee, k. (2015). effects of trait test anxiety and state anxiety on children's working memory task performance. learning and individual differences, 40, 141-148. doi:10.1016/j.lindif.2015.04.007 oecd. (2005). pisa 2003 technical report. retrieved from http://www.oecd.org/edu/school/programmeforinternationalstudentassessmentpisa/35188570.pdf oecd. (2013a). pisa 2012 released mathematics items. retrieved from https://www.oecd.org/pisa/pisaproducts/pisa2012-2006-rel-items-maths-eng.pdf oecd. (2013b). pisa 2012 results: ready to learn: students ’engagement, drive and self-beliefs (volume iii): pisa, oecd publishing. pekrun, r. (2006). the control-value theory of achievement emotions: assumptions, corollaries, and implications for educational research and practice. educational psychology review, 18(4), 315–341. doi:10.1007/s10648-006-9029-9 pekrun, r., & bühner, m. (2014). self-report masures of academic emotions. in r. pekrun & l. linnenbrink-garcia (eds.), international handbook of emotions in education (pp. 561-579). new york: routledge. pekrun, r., lichtenfeld, s., marsh, h. w., murayama, k., & götz, t. (2017). achievement emotions and academic performance: longitudinal models of reciprocal effects. child development, 88(5), 1653-1670. doi:10.1111/cdev.12704 pekrun, r., & perry, r. p. (2014). control-value theory of achievement emotions. in r. pekrun & l. linnenbrink-garcia (eds.), international handbook of emotions in education (pp. 120-141). new york: routledge. plake, b. s., & parker, c. s. (1982). the development and validation of a revised version of the mathematics anxiety rating scale. educational and psychological measurement, 42(2), 551-557. doi:10.1177/001316448204200218 pletzer, b., wood, g., moeller, k., nuerk, h. c., & kerschbaum, h. h. (2010). predictors of performance in a real-life statistics examination depend on the individual cortisol profile. biological psychology, 85, 410-416. doi:10.1016/j.biopsycho.2010.08.015 pletzer, b., kronbichler, m., nuerk, h.-c., & kerschbaum, h. h. (2015). mathematics anxiety reduces default mode network deactivation in response to numerical tasks. frontiers in human neuroscience, 9, 202. doi:10.3389/fnhum.2015.00202 richardson, f. c., & suinn, r. m. (1972). the mathematics anxiety rating scale: psychometric data. journal of counseling psychology, 19(6), 551–554. doi:10.1037/h0033456 sarkar, a., dowker, a., & cohen kadosh, r. (2014). cognitive enhancement or cognitive cost: trait-specific outcomes of brain stimulation in the case of mathematics anxiety. the journal of neuroscience, 34, 16605–16610. doi:10.1523/jneurosci.3129-14.2014 steimer, t. (2002). the biology of fearand anxiety-related behaviors. dialogues in clinical neuroscience, 4(3), 231-249. suárez-pellicioni, m., núñez-peña, m. i., & colomé, à. (2016). math anxiety: a review of its cognitive consequences, psychophysiological correlates, and brain bases. cognitive, affective, & behavioral neuroscience, 16(1), 3-22. doi:10.3758/s13415-015-0370-7 wiedemann, k. (2015). anxiety and anxiety disorders. in j. d. wright (ed.), international encyclopedia of the social & behavioral sciences (second edition) (pp. 804-810). amsterdam: elsevier. wigfield, a., & meece, j. l. (1988). math anxiety in elementary and secondary school students. journal of educational psychology, 80 (2), 210-216. doi:10.1037/0022-0663.80.2.210 zeidner, m. (2014). anxiety in education. in r. pekrun & l. linnenbrink-garcia (eds.), international handbook of emotions in education (pp. 120-141). new york: routledge. table a.1 full hierarchical multiple regression analyses for physiological mastate and self-reported mastate codepen chauliac et al frontline learning research vol.8 no. 3 special issue (2020) 26 39 issn 2295-3159 it is all in the surv-eye: can eye tracking data shed light on the internal consistency in self-report questionnaires on cognitive processing strategies? margot chauliaca, leen catryssea,david gijbelsa,vincent donchea auniversity of antwerp, belgium article received 13 may 2019 / revised 18 february / accepted 23 february / available online 30 march abstract although self-report questionnaires are widely used, researchers debate whether responses to these types of questionnaires are valid representations of the respondent’s actual thoughts and beliefs. in order to provide more insight into the quality of questionnaire data, we aimed to gain an understanding of the processes that impact the completion of self-report questionnaires. to this end, we explored the process of completing a questionnaire by monitoring the eye tracking data of 70 students in higher education. specifically, we examined the relation between eye movement measurements and the level of internal consistency demonstrated in the responses to the questionnaire. the results indicated that respondents who look longer at an item do not necessarily have more consistent answering behaviour than respondents with shorter processing times. our findings indicate that eye tracking serves as a promising tool to gain more insight into the process of completing self-report questionnaires. keywords: eye tracking; cognitive processes; survey research; self-report questionnaires; working memory capacity info corresponding author email: margot.chauliac@uantwerpen.be doi: https://doi.org/10.14786/flr.v8i3.489 1. introduction in self-report questionnaires respondents are asked to answer questions about themselves and as an instrument they are widely used to measure beliefs, attitudes, feelings and opinions in diverse fields of research (singleton & straits, 2009). this also holds for the domain of research on learning and instruction where self-report questionnaires are often used to map student learning. important assets of these questionnaires are that they are easy to administer in both small and large groups and that their use is time and cost-effective. however, despite the reliability, validity and advantages self-report questionnaires might offer, a critical stance towards their use is required to gain more insight into students' cognitive processing strategies (dinsmore & alexander, 2012). many researchers argue that respondents are, consciously or unconsciously, not always able to respond accurately to these questions (schellings, 2011; schellings & van hout-wolters, 2011). this inability may influence the consistency by which a respondent scores the different items of a questionnaire, and thus the reliability of its outcomes (richardson, 2004, 2013; veenman, 2011; veenman & van hout-wolters, 2005). in order to assess the quality of the retrieved data, it is critical to examine whether the responses to questionnaires are valid representations of respondents’ actual thoughts and beliefs (schwarz, 2007; tourangeau et al 2000). thus, when a specific set of items focuses on the same topic (i.e. a scale mapping a specific belief), the responses to these items need to be representative of the respondent’s beliefs. an individual answering pattern on a set of items that is consistent with the underlying scale leads to reliable survey data. when this is not the case, one can start questioning how the respondent completed the survey and to what extent this is related to the consistency of their answers. generally, there is a black box concerning the processes in participants’ completion of self-report questionnaires. gaining an understanding of these processes could help to provide additional insight into the quality of survey data. however, this area of interest, and in particular, the process of completing the questionnaires, has been under-examined in the literature so far. in this study, we use eye tracking to examine the processes at play when completing self-report questionnaires that aim to map students’ cognitive processing strategies. in particular, we focus on the specific strategies that students use while processing items. 2. theoretical framework 2.1 cognitive processing when completing self-report questionnaires surveys have a long history in educational research (marsden & wright, 2010; rossi et al 1983). despite its long history, it is only from 1980 onwards that cognitive psychology started to enter the field of survey research. the focus shifted from the outcomes of the questionnaire to the cognitive processing activities that were at play when completing questionnaires (fowler, 2014; willis & miller, 2011). however, despite the development of the cognitive aspects of survey methodology (casm), the focus was still on examining how cognitive processes could influence the outcomes of the questionnaires, instead of investigating how the underlying processes while completing self-report questionnaires could be related to the reliability of their outcomes. following the casm-movement, multiple theoretical models were developed to grasp the processes at work in the reading of questions and providing answers to these questions (jobe & herrmann, 1996): these included the four-stage model by tourangeau (1984), the autobiographical question-answering model from schwarz (1990), the flexible processing model by willis, et al (1991) and the information processing model of self-report item response (karabenick et al., 2007). all these models share the common feature of attempting to grasp the complexity of completing survey questionnaires by distinguishing important stages that respondents go through in order to generate an answer. most researchers agree that respondents must give meaning to a question and be able to retrieve necessary information from their memory. only then they will be able to make an informed decision and choose a congruent response option (karabenick et al., 2007; tourangeau, 1984). all these models are characterised by the fact that the described stages do not have to follow each other in a linear sequence. one can move back and forth so that there may be iterations and overlap between the steps. it is even possible that one or some of the steps are weakly conducted or completely missing. nevertheless, one can only expect substantive answers when respondents thoroughly conduct all cognitive processes when answering a question (krosnick, 1991; krosnick & alwin, 1987). in the field of survey research, the comprehension of the question is an important prerequisite for achieving meaningful results. a crucial step is, therefore, to design a questionnaire such that all respondents understand the items in the same way as the researcher intended (neuert, 2016). previous research already demonstrated that comprehension problems could arise or that respondents may satisfice while completing the survey (krosnick & alwin, 1987). another important factor that plays a role is the capacity of a respondent’s working memory. working memory concerns the limited amount of information that can be processed and temporarily stored in the memory while performing complex cognitive tasks (baddeley & hitch, 1974). krosnick (1991) argues that working memory is limited and that respondents are unable to give the latter options as much attention as the ones they consider initially. moreover, respondents may differ in cognitive ability to complete survey questions. this can influence the eventual results (gathercole & alloway 2013; krosnick, 1991). until now, a lot of research has been done to gain more insight into the problems that might arise while completing self-report questionnaires (galesic et al 2008; graesser et al 2006; lenzner et al 2011). in the past, researchers made use of cognitive interviewing techniques such as think-aloud protocols and verbal probing to get a grip on the difficulties that might arise (collins, 2003). the think-aloud protocol is a data-gathering method in which respondents are asked to verbalise their thought processes during or after doing a specific task. verbal probing is a cognitive interviewing technique where questions are designed to elicit specific information that is usually not provided by respondents. these cognitive interviews provide a suitable methodology for examining the extent to which tools of inquiry capture the experiences of students in a valid and reliable manner (beatty & willis, 2007; desimone & le floch, 2004; presser et al., 2004). however, despite the benefits cognitive interviews have to offer, they do not allow researchers to look directly into processing behaviour while respondents complete in the questionnaire. 2.2 eye tracking as an eye-opener in survey research previous research has shown that eye tracking can help gain more insight into the black box of the processes of completing self-report questionnaires (galesic et al., 2008; redline & lankford, 2001). via this relatively unobtrusive instrument, one can track the implicit processes at play while completing questionnaires. eye tracking research has a long tradition in studying cognitive processing during reading and other information processing tasks (duchowski, 2007; neuert, 2016; rayner, 1998). more recently, the technique has also been introduced into the field of survey methodological research to study cognitive processes while answering survey questions (lenzner et al 2010; neuert, 2016). in previous research, eye tracking has been used to study, among other topics, the visual designs of branching instructions (redline & lankford, 2001), different response formats (lenzner et al., 2014), response order effects (galesic et al., 2008), the effects of question wording (graesser et al., 2006; lenzner et al., 2011) and the cognitive processes associated with answering rating scale questions (menold et al., 2014). however, these were mainly experimental studies that focused on the aspects of the questionnaire that could lead to difficulties in processing. by investigating the potential burden the questions might bring, one focuses on the possible constraints of the survey. however, the effects these difficulties have on the actual process of completing the questionnaire have not yet been addressed. the relationship between eye movements and cognitive processing is based on two assumptions: the immediacy assumption and the eye-mind assumption. the immediacy assumption states that a visual stimulus on which the eyes fixate is processed immediately. the eye-mind assumption postulates that as long as the stimulus is fixated, it is mentally processed. thus, both assumptions suggest that eye movements provide direct information about what is processed and the amount of cognitive effort that is involved (just & carpenter, 1980). although eye tracking cannot help in making a concrete distinction between the different stages respondents might go through when completing self-report questionnaires, it does provide insight into the entire process that evolves in the time period between the reading of the stimulus — in this case, the self-report question — and the giving of an answer. the duration a respondent spends processing gives an indication of the cognitive effort the respondent put into the processing. longer fixation times could, for example, be associated with a deeper and more effortful cognitive processing or may be an indicator of comprehension problems (holmqvist et al., 2011). 2.3 present research in this study, we will include eye tracking as an online measure in order to gain an understanding of the cognitive processes that are active while completing self-report questionnaires. by taking a closer look at eye tracking data, we strive to examine whether the underlying processes are possibly explanatory indicators of the internal consistency by which the respondent completed the questionnaire. after all, internal consistency is one of the most critical prerequisites in obtaining meaningful results from survey data. consistency is determined by how similar a respondent answers questions that belong to the same scale. our study aims to answer the following two research questions: 1. to what extent is there a relation between the consistency in answering behaviour and eye movement measures when completing a self-report questionnaire? 2. to what extent is there a relationship between the consistency in answering behaviour, eye movement measurers and a respondent's working memory capacity when completing a self-report questionnaire? in the assumption that the time a respondent spends fixating on an area of the survey item more or less corresponds to the time this area is processed (staub & rayner, 2007), the time taken to choose an answering option can be an indicator of the cognitive effort that was invested in arriving at this answer or judgment (fazio, 1990). therefore, we hypothesise that there could be a link between the cognitive processing taking place when scoring the items of a self-report questionnaire and the internal consistency of the scales. based on the previous findings on working memory capacity (krosnick, 1991), we expect an interplay between students’ working memory capacity and the cognitive process taking place when completing a questionnaire. 3. methodology 3.1 participants the sample consisted of 92 bachelor students from a social science faculty. students were recruited during regular lectures and all participated on a voluntary basis. before the start of the experiment, we received their consent, which was approved by the ethics committee for social sciences and humanities of the participating university. all participants had a normal or corrected-to-normal vision and had dutch as their native language. due to issues that are common in eye tracking research (e.g. problems with the calibration of the eye tracker, and a lack of responses to the survey questions [see e.g. holmqvist et al., 2011]) we lost data from 22 respondents. data from 10 respondents were excluded due to technical issues; data from 12 respondents were left out because of poor quality of the eye tracking data. after this data cleaning, the data from 70 participants were included in the statistical analyses. to thank the students for their participation, they received two cinema tickets. 3.2 materials and procedure the self-report questionnaire data were collected as part of a larger project about learning from texts and the completing of questionnaires where we recorded eye movements to gain insight into the processing behaviour of participants. after the reading of each text, a validated self-report questionnaire was completed to measure students’ task-specific processing strategies. a task-specific version of the ils-sv questionnaire was developed based on the original version (donche & van petegem, 2008; vermunt & donche, 2017). this version contained four scales about cognitive processing strategies and consisted of sixteen items that mapped how participants process information when reading a particular text. students had to read the question, select the answering category of their choice, and state their answer out loud. answering options ranged from 1 = 'i rarely or never do this' to 5 = 'i almost always do this'. all survey items were answered consecutively without the possibility of changing the given answer. apart from completing the self-report questionnaire, the students’ working memory capacity was measured by means of the automated operation span task (aospan). according to unsworth et al. (2005), the aospan is a reliable and valid test for measuring the working memory capacity that can be used in various research domains. participants were required to solve a series of mathematical operations while trying to retain a set of unrelated letters. the aospan is mouse-driven, calculates scores automatically and requires little to no intervention from the experimenter (unsworth et al., 2005). in order to be sure that participants were not only focusing on remembering the letters, a 85% accuracy criterion was imposed for solving the mathematical problems (unsworth et al., 2005). the aospan provides two scorings, an absolute credit scoring and a partial credit scoring. since partial credit scoring is preferred over the absolute all-or-nothing scoring, we made use of the latter (conway et al., 2005). the mean score for all respondents was 60.11 (sd = 9.85). the score for this working memory capacity test was normally distributed, and for further analysis, we made use of standardised scores. 3.3 eye tracking equipment to measure students’ eye movements, we made use of the tobii pro x3-120 eye tracker, which alternates between bright and dark pupil eye tracking in a predefined, systematic way. this eye tracker had a sampling frequency of 120 hz (binocularly), which made it possible to take a closer look at the fixation durations. the eye tracker was secured to a 17.3-inch monitor with a resolution of 1.920 x 1.080 pixels. every participant sat at about 60 cm from the screen and the eye tracker. to minimise the influence of student movement, we employed a chinrest. tobii technology (stockholm, sweden) reported a gaze accuracy of 0.4°, gaze precision of 0.24° and a total system latency of fewer than 11 milliseconds for this eye tracker. the eye movements were recorded with tobii-studio (3.4.8) software. 3.4 consistency in response behavior the first indication of consistency in response behaviour is the cronbach alpha coefficient, which was calculated for each of the four scales. the consistency levels for the four scales were .67 .68, .69 and .65, respectively (table 1). these results show an acceptable internal consistency level for four-item scales. since only a small number of items are used per scale, and given the sensitivity of the cronbach's alpha for the number of items, a cut-off value of .60 is considered sufficient (cortina, 1993; pallant, 2007). as respondents can differ in the way they score the separate items of a specific scale, thus showing diversity in scoring behaviour across items, we categorised their rating behaviour for each scale. for all respondents, four consistency indicators were created (one for each scale), making distinctions between raters using the same answering category for all items on a scale or raters showing more diversity in the use of answering categories, by making, for instance, use of at least two different answering categories. the consistency indicator ranged from 1 (consistent answering pattern) to 4 (very diverse answering pattern). the questionnaire did not include reversed items, so this could not serve as an explanation for diversity in answering categories. table 1 ils-sv scales, number of items, item examples and reliability (internal consistency) 3.5 analysis we used the tobii fixation filter for fixation identification, which is an implementation of a classification algorithm proposed by olsson (2007). it uses a velocity threshold (35 pixels per window) and a distance threshold (35 pixels) (olsen, 2012). eye movement data were analysed at the item level. the question and response field for each item in the survey were considered as a combined area of interest (aoi). for each aoi or item in the survey, the total fixation duration and the total fixation count were calculated separately. to control for the length of aoi's, the total fixation duration measure was normalised by calculating a milliseconds-per-character measure (ariasi et al., 2017; catrysse et al., 2016; yeari et al., 2016). the total fixation count measure was normalised by calculating a count-per-character measure. in addition, we logarithmically transformed these measures because they are heavily skewed (catrysse et al., 2018; holmqvist et al., 2011; lo & andrews, 2015). to check the distribution of the dependent measures, the fitdistrplus package was used (delignette-muller & dutang, 2015). the eye movement data were analysed with linear mixed effects models (lmm) with the lme4 package (bates et al 2015) in r (r core team, 2014) and with the rstudio interface. mixed-effects models are statistical models that incorporate random and fixed effects (baayen, 2008; baayen et al., 2008). subjects, items and scales were considered as crossed random effects (baayen, 2008; baayen et al., 2008). the analysis was conducted at the item level and was based on 1,120 data points (70 students each processing 16 items). separate models were fitted for the total fixation duration and the total fixation count. two models per measure were fitted: (1) an lmm with subjects, subscales and items as random effects and consistency in answering behaviour as a fixed effect and (2) an lmm with subjects, subscales and items as random effects and consistency in answering behaviour and working memory capacity as fixed effects. the interactions between the fixed effects were also incorporated into the second model. 4. results 4.1 the relation between consistency in answering behaviour and eye movement measures in order to answer the first research question, we report the means and standard deviations for the eye movement measures in table 2 in relation to the consistency in answering behaviour. for example, students who were very consistent in their answering behaviour on a certain scale, that is, choosing the same answering option for each item, looked on average 8.74 seconds at an item and the corresponding answering options, and made on average 31.99 fixations on an item and answering options. table 2 descriptive statistics for the number of different answering options per scale in relation to the eye movement measures note: untransformed eye movement measures reported. in the next step, we examined the relation between the consistency in answering behaviour and eye movement measures. we analysed the data with linear mixed effect models. for the total fixation duration, the parameter estimates indicated that there was a significant effect of consistency in answering behaviour on the total fixation duration for an item (table 3). more specifically, the results showed that a student who chose two or three different answering categories looked longer at an item than a student who only opted for one answering category. a student who chose four categories did not look longer at the items than a student who chose only one answering category. overall, the parameter estimates showed that students with less consistent scoring behaviour spend more time on processing the items and answering options. however, this was not the case for students who picked four different answering options on a scale. this implies that there seems to be a turning point in the effect of correlation between consistency in answering behaviour and students’ eye movement measures. table 3 parameter estimates of the random and fixed effects for the random intercept model for total fixation duration and total fixation count note: significant values are in bold. for the total fixation count, the estimate of the intercept had a negative value of -0.90. this is due to the log transformation of the count-per-character measure, which causes small values (<1) to turn into negative values. moreover, we were mainly interested in the potential change in the fixation count, rather than in its absolute value. therefore, this negative value was not problematic for the interpretation of our results. the results for the fixation count are similar as for the total fixation duration. a student who chose two or three answering categories made more fixations on an item than a student opting for only one answering category. a student who chose four categories did not make more fixations on the items than a student who picked only one answering category. 4.2 the relation between consistency in answering behaviour, working memory capacity and eye movement measures to answer the second research question on the relationship between consistency in answering behaviour, working memory capacity and eye movement measures, we updated the mixed effects model of table 5 and added working memory capacity as a fixed effect in a new model (table 4). both for the total fixation duration and total fixation count, we did not find any significant effect of working memory capacity. we can thus conclude that working memory capacity in this study has no interference with students’ eye movement measures when completing this self-report questionnaire. table 4 parameter estimates of the random and fixed effects for the random intercept model for total fixation duration and total fixation count including working memory capacity note: ao: answering option(s) — wmc: working memory capacity — significant values are in bold. 5. discussion although self-report questionnaires are widely used to map students’ processing strategies, there still is a lacuna in the knowledge about the processes at play while respondents complete these questionnaires. by gaining insight into these processes, we want to provide evidence for the debate about the often-reported reliability issues of self-report questionnaires (richardson, 2004, 2013; veenman, 2011; veenman & van hout-wolters, 2005). in this exploratory study, we used eye tracking in order to unobtrusively track the processes that are at play while completing a self-report questionnaire on cognitive processing strategies. previous research mainly focused on cognitive difficulties that might arise when processing the questionnaire, and therefore the questionnaire’s potential limitations to accurately grasp respondents' opinions and beliefs (galesic et al., 2008; graesser et al., 2006; lenzner et al., 2014; menold et al., 2014; redline & lankford, 2001). how these difficulties affect the process of completing the questionnaire, and how they influence the questionnaire’s reliability, are two questions that have not been addressed before. based on previous research stating that the time a respondent spends fixating on a specific area is more or less equal to the time this area is being processed, the processing time is assumed to be a good indicator of the invested cognitive effort (fazio, 1990; staub & rayner, 2007). therefore, we believe that there could be a link between the cognitive processing taking place when scoring the items of a survey and the internal consistency of the scored scales from a questionnaire. this concept of consistency is important, given the fact that when a respondent's answering behaviour is not consistent, one can thus start questioning the reliability of the survey data. we first examined the relationship between the consistency in answering behaviour and eye movement measures. our results demonstrate that the consistency in answering behaviour is significantly related to the total fixation duration for an item. the more a respondent’s answers differ in one scale, the longer the respondent looks at the items compared to those who only opt for one answering option. however, no significant difference was found between the respondents choosing one response option and the ones opting for four different answers for items belonging to the same scale. given these results, there seems to be a turning point in the effect of consistency in answering behaviour. results suggest that too much pondering over a question does not lead directly to a more consistent answering behaviour. on the contrary, when the respondents spend more time processing a question, they might be trying to process the question more thoroughly to come to an appropriate and thus consistent response, but they just do not succeed in doing so. secondly, we aimed to gain more insight into the relation between eye movement measures, answering behaviour and the working memory capacity of the respondent. previous research on the working memory demonstrated that its capacity is limited and that respondents may therefore not give each answering option as much attention as the one they considered initially (gathercole & alloway 2013; krosnick, 1991). therefore, we hypothesised that we would find less consistent answering behaviour for the students with a lower working memory capacity. however, both for the total fixation duration as well as for the total fixation count, we did not find any significant effect of working memory capacity. this could be because memory distortions do not play a significant role when this self-report questionnaire is being completed immediately after completing the task that the questionnaire referred to. 6. limitations and directions for future research although our findings show that eye tracking is a promising technique to gain more insight into the process of completing self-report questionnaires, we want to emphasise the exploratory nature of this study and point at some limitations and directions for future research. the process of completing questionnaires is an extremely complex process. different theoretical models try to distinguish different stages that possibly play a role when a respondent is cognitively processing a question (see e.g. karabenick et al., 2007; tourangeau, 1984). in our study, we considered the question as well as the answering options as one area of interest. this choice allows for an indication of the total time taken until one decides and thus completes the process of filling in the item. more specifically, by focussing on the survey item in its entirety, we took all stages of the different theoretical models into account. in future research, it would be interesting to separate this area into two distinct areas of interest — the question and the answering options — in order to investigate which possible influence each of these areas has on the internal consistency. this would also allow us to further separate the different stages of the theoretical models. however, as these stages do not follow a linear path, separating into different areas of interest will lead to a loss of information. when analysing the question, we could, for example, consider whether different reading processes lead to different outcomes in internal consistency. hereto, it would also be important to take other eye tracking measures into account. in our study, we made use of the total fixation duration and the fixation count to map the whole process. however, analysing merely the question would allow us to use other measures such as first pass fixations and second pass fixations (hyönä et al.,2003; jarodzka & brand-gruwel, 2017) which could possibly shed some light on further difficulties the respondents encountered. next to looking at the question itself, it could also be clarifying to look at how the respondent processes the different the answering options. the way in which a respondent ponders over a question — merely focusing on one answering option or considering each of the five possibilities — could potentially elucidate their answering behaviour. another constraint of this study is that no use was made of complementary data. the use of merely eye tracking data might not provide us with the necessary insight into the reasons why students who respond in a less consistent way take more time to respond to the items, whereas using a multi-method approach to look at the data could help us put different pieces of the puzzle together (catrysse et al., 2018). as we already know from previous research, a longer reading time can be an indication of several different cognitive processes such as (1) high-level or deeper cognitive processing (ariasi & mason, 2011; holmqvist et al., 2011; penttinen et al., 2013), (2) strategic attempts to resolve comprehension problems or further text comprehension (ariasi et al., 2017; hyönä & lorch, 2004; hyönä, lorch, & kaakinen, 2002; hyönä et al., 2003; kinnunen & vauras, 1995), (3) comprehension monitoring (van gog & jarodzka, 2013), (4) difficulty with text passages (rayner, et al., 2006) and (5) attempts to reinstate information into working memory in order to elaborate or rehearse that information (hyönä & lorch, 2004). however, further research is needed to investigate whether the current insight in the field of text reading also hold for the process of completing survey questionnaires. a last observation is that when completing the questionnaire, respondents were asked to state their given answer out loud after every question. knowing that researchers were monitoring their answers could possibly have had an influence on the natural process of completing the questionnaire. for future research, it would therefore be necessary to look at the processes that are at play without verifying for the responses given by respondents. moreover, it was impossible for the respondents to change their answer on certain questions. once they provided an answer, the next question was immediately projected without an opportunity for the respondent to change their mind. considering this in future research, one will be able to search for doubts and changes in the answering process. 7. conclusions notwithstanding certain limitations, our exploratory study was able to show that eye tracking offers important research perspectives that helped us gain more insight into the cognitive processes at play in the process of completing a self-report questionnaire. it also gave us insight into how these processes are related to the consistency by which the survey has been completed. by lifting a corner of the veil that lies over survey research, we now not only know that a longer processing time is not necessarily linked to more consistent answering behaviour, but also that there is a turning point in which longer processing does not lead to more consistency in answering behaviour. keypoints the use of eye tracking to record the process of completing self-report questionnaires appears to be a promising tool to gain more insight herein. respondents who look longer at the item in question do not necessarily have more consistent answering behaviour than respondents who spend less time answering the questions. when answering self-report questionnaires, there seems to be a turning point in which a longer focus on the item does not lead to a more consistent answering pattern. references ariasi, n., hyönä, j., kaakinen, j., & mason, l. (2017). an eye-movement analysis of the refutation effect in reading science text. journal of computer assisted learning, 33(3), 202-221. https://doi.org/10.1111/jcal.12151 ariasi, n., & mason, l. (2011). uncovering the effect of text structure in learning from a science text: an eye-tracking study. instructional science, 39(5), 581-601. https://doi.org/10.1007/s11251-010-9142-5 baayen, r. (2008). analyzing linguistic data: a practical introduction to statistics using r . cambridge: cambridge university press. baayen, r., davidson, d., & bates, d. (2008). mixed-effects modeling with crossed random effects for subjects and items. journal of memory and language, 59(4), 390-412. https://doi.org/10.1016/j.jml.2007.12.005 baddeley, a., & hitch, g. (1974). working memory. in g. h. bower (ed.), psychology of learning and motivation (vol. 8, pp. 47-89): academic press. bates, d., mächler, m., bolker, b., & walker, s. (2015). fitting linear mixed-effects models using lme4. journal of statistical software, 67(1), 1-48. https://doi.org/10.18637/jss.v067.i01 beatty, p., & willis, g. (2007). research synthesis: the practice of cognitive interviewing. public opinion quarterly, 71(2), 287-311. https://doi.org/10.1093/poq/nfm006 catrysse, l., gijbels, d., & donche, v. (2018). it is not only about the depth of processing: what if eye am not interested in the text? learning and instruction, 58, 284-294. https://doi.org/10.1016/j.learninstruc.2018.07.009 catrysse, l., gijbels, d., donche, v., de maeyer, s., van den bossche, p., & gommers, l. (2016). mapping processing strategies in learning from expository text: an exploratory eye tracking study followed by a cued recall. frontline learning research, 4(1), 1-16. https://doi.org/10.14786/flr.v4i1.192 collins, d. (2003). pretesting survey instruments: an overview of cognitive methods. quality of life research, 12(3), 229-238. https://doi.org/10.1023/a:1023254226592 conway, a., kane, m., bunting, m., hambrick, d., wilhelm, o., & engle, r. (2005). working memory span tasks: a methodological review and user’s guide. psychonomic bulletin & review, 12(5), 769-786. https://doi.org/10.3758/bf03196772 cortina, j. m. (1993). what is coefficient alpha? an examination of theory and applications. journal of applied psychology, 78(1), 98-104. https://doi.org/10.1037/0021-9010.78.1.98 delignette-muller, m. l., & dutang, c. (2015). fitdistrplus: an r package for fitting distributions. journal of statistical software, 64(4), 1-34. https://doi.org/10.18637/jss.v064.i04 desimone, l., & le floch, k. (2004). are we asking the right questions? using cognitive interviews to improve surveys in education research. educational evaluation policy analysis, 26(1), 1-22. https://doi.org/10.3102/01623737026001001 dinsmore, d., & alexander, p. (2012). a critical discussion of deep and surface processing: what it means, how it is measured, the role of context, and model specification. educational psychology review, 24(4), 499-567. https://doi.org/10.1007/s10648-012-9198-7 donche, v., & van petegem, p. (2008). the validity and reliability of the short inventory of learning patterns. in e. cools, h. van den broeck, & t. redmond (eds.), style and cultural differences: how can organisations, regions and countries take advantage of style differences (pp. 49-59). ghent: vlerick leuven ghent management school. duchowski, a. (2007). eye tracking methodology: theory and practice. london: springer. fazio, r. (1990). multiple processes by which attitudes guide behavior: the mode model as an integrative framework. in m. p. zanna (ed.), advances in experimental social psychology (vol. 23, pp. 75-109). new york: academic press. fowler, f. (2014). survey research methods 5th edition. thousand oaks: sage publications. galesic, m., tourangeau, r., couper, m., & conrad, f. (2008). eye-tracking data: new insights on response order effects and other cognitive shortcuts in survey responding. public opinion quarterly, 72(5), 892-913. https://doi.org/10.1093/poq/nfn059 gathercole, s., & alloway , t. (2013). de invloed van het werkgeheugen op het leren: handelingsgerichte adviezen voor het basisonderwijs . amsterdam: swp, amsterdam. graesser, a., cai, z., louwerse, m., & daniel, f. (2006). question understanding aid (quaid) a web facility that tests question comprehensibility. public opinion quarterly, 70(1), 3-22. https://doi.org/10.1093/poq/nfj012 holmqvist, k., nyström, m., andersson, r., dewhurst, r., jarodzka, h., & van de weijer, j. (2011). eye tracking: a comprehensive guide to methods and measures. oxford: oxford university press. hyönä, j., & lorch, r. (2004). effects of topic headings on text processing: evidence from adult readers' eye fixation patterns. learning and instruction, 14(2), 131-152. https://doi.org/10.1016/j.learninstruc.2004.01.001 hyönä, j., lorch, r., & kaakinen, j. (2002). individual differences in reading to summarize expository text: evidence from eye fixation patterns. journal of educational psychology, 94(1), 44-55. https://doi.org/10.1037//0022-0663.94.1.44 hyönä, j., lorch, r., & rinck, m. (2003). eye movement measures to study global text processing. in r. hyönä (ed.), the mind's eye: cognitive and applied aspects of eye movement research (pp. 313-334). amsterdam: elsevier science. jarodzka, h., & brand-gruwel, s. (2017). tracking the reading eye: towards a model of real-world reading. journal of computer assisted learning, 33(3), 193-201. https://doi.org/10.1111/jcal.12189 jobe, j., & herrmann, d. (1996). implications of models of survey cognition for memory theory. basic applied memory research, 2, 193-205. https://doi.org/10.1023/a:1023279029852 just, m., & carpenter, p. (1980). a theory of reading: from eye fixations to comprehension. psychological review, 87(4), 329. https://doi.org/10.1037/0033-295x.87.4.329 karabenick, s., woolley, m., friedel, j., ammon, b., blazevski, j., bonney, c., . . . kempler, t. (2007). cognitive processing of self-report items in educational research: do they think what we mean? educational psychologist, 42(3), 139-151. https://doi.org/10.1080/00461520701416231 kinnunen, r., & vauras, m. (1995). comprehension monitoring and the level of comprehension in highand low-achieving primary school children's reading. learning and instruction, 5(2), 143-165. https://doi.org/10.1016/0959-4752(95)00009-r krosnick, j. (1991). response strategies for coping with the cognitive demands of attitude measures in surveys. applied cognitive psychology, 5(3), 213-236. https://doi.org/10.1002/acp.2350050305 krosnick, j., & alwin, d. (1987). an evaluation of a cognitive theory of response-order effects in survey measurement. public opinion quarterly, 51(2), 201-219. https://doi.org/10.1086/269029 lenzner, t., kaczmirek, l., & galesic, m. (2011). seeing through the eyes of the respondent: an eye-tracking study on survey question comprehension. international journal of public opinion research, 23(3), 361-373. https://doi.org/10.1093/ijpor/edq053 lenzner, t., kaczmirek, l., & galesic, m. (2014). left feels right: a usability study on the position of answer boxes in web surveys. 32 (6), 743-764. https://doi.org/10.1177/0894439313517532 lenzner, t., kaczmirek, l., & lenzner, a. (2010). cognitive burden of survey questions and response times: a psycholinguistic experiment. applied cognitive psychology, 24(7), 1003-1020. https://doi.org/10.1002/acp.1602 lo, s., & andrews, s. (2015). to transform or not to transform: using generalized linear mixed models to analyse reaction time data. frontiers in psychology, 6, 1171. https://doi.org/10.3389/fpsyg.2015.01171 marsden, p., & wright, j. (2010). handbook of survey research 2nd edition. bingley: emerald group publishing. menold, n., kaczmirek, l., lenzner, t., & neusar, a. (2014). how do respondents attend to verbal labels in rating scales? field methods, 26(1), 21-39. https://doi.org/10.1177/1525822x13508270 neuert, c. (2016). eye tracking in questionnaire pretesting. olsen, a. (2012). the tobii i-vt fixation filter. olsson, p. (2007). real-time and offline filters for eye tracking. pallant, j. (2007). spss survival manual: a step by step guide to data analysis using spss for windows 3th edition : maidenhead: open university press. penttinen, m., anto, e., & mikkilä-erdmann, m. (2013). conceptual change, text comprehension and eye movements during reading. research in science education, 43(4), 1407-1434. https://doi.org/10.1007/s11165-012-9313-2 presser, s., couper, m., lessler, j., martin, e., martin, j., rothgeb, j., & singer, e. (2004). methods for testing and evaluating survey questions. public opinion quarterly, 68(1), 109-130. https://doi.org/10.1093/poq/nfh008 rayner, k. (1998). eye movements in reading and information processing: 20 years of research. psychological bulletin, 124(3), 372-422. https://doi.org/10.1037/0033-2909.124.3.372 rayner, k., chace, k., slattery, t., & ashby, j. (2006). eye movements as reflections of comprehension processes in reading. scientific studies of reading, 10(3), 241-255. https://doi.org/10.1207/s1532799xssr1003_3 redline, c. d., & lankford, c. (2001). eye-movement analysis: a new tool for evaluating the design of visually administered instruments (paper and web). proceedings of the survey research methods section of the american statistical association . richardson, j. (2004). methodological issues in questionnaire-based research on student learning in higher education. educational psychology review, 16(4), 347-358. https://doi.org/10.1007/s10648-004-0004-z richardson, j. (2013). research issues in evaluating learning pattern development in higher education. studies in educational evaluation, 39(1), 66-70. https://doi.org/10.1016/j.stueduc.2012.11.003 rossi, p., wright, j., & anderson, a. (1983). handbook of survey research. sample surveys: history, current practice, and future prospects . san diego: academic press. schellings, g. (2011). applying learning strategy questionnaires: problems and possibilities. metacognition and learning, 6(2), 91-109. https://doi.org/10.1007/s11409-011-9069-5 schellings, g., & van hout-wolters, b. (2011). measuring strategy use with self-report instruments: theoretical and empirical considerations. metacognition and learning, 6(2), 83-90. https://doi.org/10.1007/s11409-011-9081-9 schwarz, n. (1990). assessing frequency reports of mundane behaviors: contributions of cognitive psychology to questionnaire construction. in research methods in personality and social psychology. (pp. 98-119). thousand oaks, ca, us: sage publications, inc. schwarz, n. (2007). cognitive aspects of survey methodology. applied cognitive psychology, 21(2), 277-287. https://doi.org/10.1002/acp.1340 singleton, r., & straits, b. (2009). approaches to social research 5th edition. oxford: oxford university press. staub, a., & rayner, k. (2007). eye movements and on-line comprehension processes. in g. gaskell (ed.), the oxford handbook of psycholinguistics: oxford university press. tourangeau, r. (1984). cognitive sciences and survey methods. in t. jabine, m. straf, j. tanur, & r. tourangeau (eds.), cognitive aspects of survey methodology: building a bridge between disciplines (pp. 73-100). washington, dc: national academy press. tourangeau, r., rips, l., & rasinski, k. (2000). the psychology of survey response. cambridge: cambridge university press. unsworth, n., heitz, r., schrock, j., & engle, r. (2005). an automated version of the operation span task. behavior research methods, 37 (3), 498-505. https://doi.org/10.3758/bf03192720 van gog, t., & jarodzka, h. (2013). eye tracking as a tool to study and enhance cognitive and metacognitive processes in computer-based learning environments. in r. azevedo & v. aleven (eds.), international handbook of metacognition and learning technologies (pp. 143-156). new york: springer. veenman, m. (2011). alternative assessment of strategy use with self-report instruments: a discussion. metacognition and learning, 6(2), 205-211. https://doi.org/10.1007/s11409-011-9080-x veenman, m., & van hout-wolters, b. (2005). the assessment of metacognitive skills: what can be learned from multi-method designs? in c. artelt & b. moschner (eds.), lernstrategien und metakognition: implikationen für forschung und praxis (pp. 77-99): münster: waxmann. vermunt, j., & donche, v. (2017). a learning patterns perspective on student learning in higher education: state of the art and moving forward. educational psychology review, 29(2), 269-299. https://doi.org/10.1007/s10648-017-9414-6 willis, g., & miller, k. (2011). cross-cultural cognitive interviewing: seeking comparability and enhancing understanding. field methods, 23(4), 331-341. https://doi.org/10.1177/1525822x11416092 willis, g., royston, p., & bercini, d. (1991). the use of verbal report methods in the development and testing of survey questionnaires. applied cognitive psychology, 5(3), 251-267. https://doi.org/10.1002/acp.2350050307 yeari, m., oudega, m., & van den broek, p. (2016). the effect of highlighting on processing and memory of central and peripheral text information: evidence from eye movements. journal of research in reading, 40(4), 365-383. https://doi.org/10.1111/1467-9817.12072 codepen schick frontline learning research vol.9 no. 1 (2021) 1 29 issn 2295-3159 senior medical student attitudes towards patient communication and their development across the clinical elective year – a q-methodology study kristina schick1, martin gartmeier1& pascal o. berberat1 atechnical university munich, tum school of medicine, tum medical education center, germany 4 november 2019 / article revised 2 december 2020 / accepted 5 december / available online 13 january 2021 abstract to be proficient in communicating with patients, physicians need specified knowledge, skills and attitudes. until now, medical educators have mostly focused on undergraduate students’ communication knowledge and skills in training and assessment. attitudes towards communication with patients have been researched less frequently, but it is plausible that they also influence physicians’ behaviours in many ways. the present study investigates the communication-focused attitudes of senior medical students and their development through the clinical elective year using an innovative approach based on q-methodology. we conducted a q-methodology study using statements from the kalamazoo communication skills assessment form. a total of 47 final-year medical students documented their attitudes towards communication by sorting these statements in regard to their importance into a normal distribution grid in medical interviews. our innovative approach included three time points during the elective year at which these statements were sorted, with only a slight decrease of participants. we applied a q-factor analysis and found three attitude profiles that were structurally stable over time. attitude profile #1 focused on providing information and fostering shared decision making; profile #2 focused on the patients’ concerns and emotions while meeting the patients’ demands for sufficient information. finally, the focus of attitude profile #3 was on using appropriate conversation techniques to structure communication and gather sufficient information. overall, the respondents assigned increasing importance to building good relationships and making shared decisions with patients over time. statements about structuring conversations and communication techniques were evaluated as less important by the end of the clinical elective year. keywords: communication competence; attitudes; q-methodology; longitudinal study. info corresponding author: email: kristina.schick@tum.de doi: https://doi.org/10.14786/flr.v9i1.583 1. introduction for most physicians, communication with patients is a frequent and important aspect of their daily clinical work. it has been estimated that physicians engage in around 150,000 to 200,000 patient conversations during their careers (fallowfield et al., 2002). moreover, the importance of good communication has been empirically demonstrated, for example, through its potential to improve doctor-patient relationships, facilitate patient compliance and positively influence physicians’ well-being (epstein & street, 2011; ha et al., 2010; levinson et al., 2010). consequently, medical education researchers have frequently investigated the question of how communication skills can be improved through didactic interventions (kurtz et al., 2016; rider & nawotniak, 2010). we adopt a complementary perspective in the present study, with our research question seeking to address how young physicians’ attitudes towards communication develop over a one-year period of practical clinical work early in their career (i.e., during their clinical electives). we advance three strategies to substantiate the innovative character of this research question and our approach to answering it. attitudes are an important yet not well researched aspect of a physician’s ability to communicate with patients. in drawing upon q-methodology, we also employ an innovative methodological approach to measure attitudes towards communicative behaviour in doctor-patient dialogues. moreover, we focus on informal learning about communication in workplace settings as an often-neglected learning strategy. in approaching the subject in this way, we valuably amend the existing literature, which puts a strong emphasis on the efficacy of formal learning in acquiring communication skills. we adopt a longitudinal perspective comprising three measurements, which is an innovative approach in research using q-methodology. in doing so, we are able to provide a detailed, empirically based account of the dynamics involved in young physicians’ attitudes towards communication across a period characterised by informal learning in a novel work environment. in elaborating on these points in the following, we will further illustrate the innovative character of the present study. 1.1 attitudes towards physician-patient communication currently, much research on physician-patient communication focuses on the concept of communication skills (e.g., kurtz et al., 2016). communication skills are conceptualised as concrete behaviours (e.g., listening attentively to a patient’s statements or negotiating an agenda with a patient), which—if appropriately applied in communication with patients—are elements of a successful physician-patient encounter. in one particular model, rider (2010) describes a number of essential skills, including opening the conversation, gathering information and fostering shared decision making. opening a conversation involves a warm welcome, after which an agenda upon which the patient and the doctor could mutually agree is set, which allows the patient to present their concerns without interruption (rider, 2010). gathering information entails questioning techniques, history taking and consideration of the possible physical and psychical aspects of the complaints and illness. gathering information is followed by finding a comprehensive solution due to the concept of shared decision making (loh & simon, 2009). shared decision making consists of a mutual agreement on the diagnoses and treatment plans to develop a decision tailored to the patient (loh & simon, 2009). an alternative conceptual approach is to think of communication as a professional competence (epstein & hundert, 2002; mehay & burns, 2009). the idea here is that physicians manage to integrate different personal resources in order to perform competently in professional situations. with regard to communication with patients, such resources are skills (as described above), knowledge (of appropriate communication techniques as well as domain-specific, biomedical knowledge) and attitudes towards communication (e.g., seeing it as an essential part of any medical treatment; blömeke et al., 2015; hartig, 2008; schick et al, 2019). because skills are regarded as one aspect of professional competence, these perspectives complement each other. research on physician-patient communication, however, has focused mainly on the skills aspect, for instance, through behaviour-oriented methods of assessment, such as objective structured clinical examinations (epstein, 2007; kurtz et al., 2016). in order to advance alternative perspectives that contribute to a more holistic understanding of the physician’s role as a communicator, shedding light on the aspect of attitudes is valuable. flocke et al. (2002), for instance, distinguished four communication styles of physicians: biopsychosocial, biomedical, person-focused and high physician control (flocke et al., 2002). these styles differ in the degrees of focus placed on psychosocial dimensions and the patient’s disease and are related to the attitude patterns described by roter et al. (1997). in both studies, inferences about physicians’ attitudes were drawn from external observations. we argue that a more direct research approach involving physicians as respondents is just as plausible. in the following, we will elaborate upon the attitude construct and the relationship between attitudes and behaviour. research has drawn upon various definitions of attitudes. first, two types of attitudes can be distinguished: (1) general attitudes towards specific targets like physical objects, groups, policies and events, and (2) “attitudes toward performing specific behaviours with respect to an object or target” (ajzen et al., 2019). second, some researchers argue that attitudes are stable and difficult to change, while others assume that attitudes can change over time and are influenced by information (bohner & dickel, 2011). these properties apply to both types of attitudes, which include general attitudes and attitudes towards a particular behaviour. in both cases, attitudes are defined as evaluations “of an object of thought” (bohner & dickel, 2011). such evaluations can be differentiated into three aspects: cognitive, affective and behavioural. the cognitive aspect refers to persuasions about an object. a person can be convinced about positive or negative attributes of the object of thought through the elaboration of information (haddock & maio, 2014). the affective aspect describes emotions and feelings towards the object, and the behavioural aspect refers to behavioural patterns associated with the object (haddock & maio, 2014). an assumption connected to the behavioural aspect is that concrete experiences with the attitude’s object can influence attitudes and lead to attitude change (bohner & dickel, 2011). our study focuses on attitudes towards behaviour as described in the theory of planned behaviour by ajzen (1991). this theory poses that behaviour is influenced by intention, which is itself influenced by attitudes towards behaviour, subjective norms and perceived behavioural control. finally, these three aspects are influenced by beliefs. beliefs are an individual’s convictions regarding the positive or negative consequences of certain behaviours. the conglomerate of positive and negative contributions results in either positive (if positive convictions outweigh negative ones) or negative attitudes (if vice versa) towards the behaviour (ajzen et al., 2019). empirical studies have shown that there is a medium relationship between behaviour-related attitudes and actual behaviour (ajzen et al., 2019). these studies assume that the more precise the measurement of the behaviour-related attitude, the more the attitude can be said to predict the behaviour. other different models are helpful in explaining the relationship between attitudes and behaviour, the most relevant one being the mode model. in this model, two attitude-to-behaviour processes are described: the spontaneous and the deliberative processes (ajzen et al., 2019). the decision of which process to activate depends on the motivation and the opportunity to elaborate on information. if the motivation is high and the opportunity to elaborate on information is afforded, the deliberative process will be activated. the deliberative process activates a particular attitude based upon an individual act in the given situation. according to this idea, a specific behaviour is consistent with an attitude. however, if there is no opportunity to elaborate on information, a spontaneous process will be activated. whether the behaviour is in accordance with the attitude or not depends on the strength of the attitude. a strong attitude will be activated automatically, and the presented behaviour will be related to the attitude. however, if it is a weak attitude, there is no automatic activation, and the presented behaviour might not be in line with the attitude (ajzen et al., 2019; fazio & olsen, 2014). a further approach to describe the relationship between attitudes and behaviour is the reflective-impulsive model (strack & deutsch, 2004). this two-process model, comparable to the mode model, describes a reflective approach and an impulsive approach. the reflective approach contains an elaborate processing of information and a conscious activation of a behaviour. however, the impulsive approach activates behaviour automatically without considering related information (strack & deutsch, 2004). to investigate attitudes towards patient-centred communication, rating scales have frequently been used. one example is a study on whether cancer patients should be told the truth about their disease (grassi et al., 2000; locatelli et al., 2013). such studies are valuable in facilitating understanding of how effective interventions can be designed to improve communication. established rating scales for assessing physicians’ communication-related attitudes are the communication skills attitude scale and the patient-provider orientation scale (krupat et al., 2000; rees et al., 2002). however, certain biases described in the methods-focused literature apply specifically to attitude-related studies (e.g., social desirability; cross, 2005; klooster et al., 2008). for this reason, different research methodologies should be used to mutually compensate for their disadvantages and to achieve a more balanced picture over time. in the present study, we draw upon the q-methodology (brown, 1996; stephenson, 1936; watts & stenner, 2012), which is tailored to study attitudes while being less prone to effects like social desirability and the tendency to the middle or extreme values (cross, 2005). in a q-study, participants are asked to order a set of statements about an attitude’s object into a normal distribution grid. therefore, the statements mutually depend upon each other. q-methodology is an individual-focused method dedicated to identifying attitude patterns shared by groups of respondents. in contrast, factor analysis (also called the r-method) is variables-focused; its purpose is to identify the underlying factors in sets of variables (cross, 2005; watts & stenner, 2005). the q-methodology has already been applied in health care research, for example, by muddiman et al., 2019, who investigated the views of medical trainees about being a good doctor. across all the attitude profiles described in this study, the participants judged aspects of good conversational and interpersonal skills as an important part of being a good doctor. 1.2 informal learning of communication skills in the workplace a further innovative aspect of the present study is its focus on how young physicians’ attitudes towards communication develop through the processes of workplace learning. such learning is embedded in daily clinical practice, is not guided by a teacher or instructor (but may be supported by a colleague or supervisor) and often occurs implicitly (tynjälä, 2008). workplace learning is one of the most important ways through which young physicians develop their communication styles with patients (archer et al., 2008; bombeke et al., 2012; brown, 2010). in the research on workplace learning, it is common to distinguish between formal, informal and non-formal learning (eraut, 2000; tynjälä, 2008). formal learning takes place “off the job” in a structured course system with a teacher or trainer purposefully striving to improve learners’ job-specific skills and/or knowledge (eraut, 2000; manuti et al., 2015). in contrast, informal learning often takes place spontaneously in implicit and unplanned ways. the goals and outcomes of informal learning are not officially documented (kyndt et al., 2009). non-formal learning combines aspects of formal and informal learning: like informal learning, non-formal learning activities take place in a work environment and are associated with work activities. however, non-formal learning is more organized than informal learning (e.g., through a mentor or coach supporting the learner; eraut, 2000). thus, non-formal learning is an individual process and is tailored to the needs of the trainee. defined learning goals of non-formal learning often focus on practical learning experiences (eraut, 2000). drawing upon these definitions, clinical electives (which are the final step of undergraduate medical education (ume)) are best described as a non-formal learning setting featuring some formal and informal elements. the participants in clinical electives are still regarded as senior medical students (rather than as young physicians); as such, they are still learners. each student in the clinical electives is assigned a senior physician as a mentor who provides advice, feedback and, if necessary, instruction on the hows and whys of clinical practice. therefore, the learning is tailored to the needs of the trainee. giroldi et al. (2017) investigated the work-related conditions that help trainees to improve their clinical communication skills. beneficial conditions comprise a psychologically safe environment allowing for different approaches to be tried out, opportunities to learn communication from proficient role models (like a mentor) and time to reflect about one’s communication practices (edmondson, 1999; giroldi et al., 2017). bombeke et al. (2012) investigated the transfer of skills acquired in communication seminars with simulated patients to a workplace setting with real patient contact. the authors indicated that there was an increase in positive attitudes towards develop patient-centred communication skills when combining simulated training with workplace learning. these findings are in line with suggestions by archer et al. (2008), who emphasised the importance of considering attitudes and their development during medical education. during their clinical electives, many medical students are confronted with real (and not simulated) patients on a regular basis. it is plausible that while conducting medical interviews for the first time independently, students may experiment with different techniques and strategies and thereby develop their own routines and preferences regarding communication, yet with no relationship to any formal curriculum (berings et al., 2005). in addition, young medical students often have many opportunities to observe and learn from more experienced physicians as they communicate with patients in the clinical setting. through observation and imitation, students might adopt specific behaviours. in a qualitative study, giroldi et al. (2017) have shown that the interaction between the trainee and trainer caused a change of the frame of reference towards communication with patients. if the trainees have positive experiences with these, routines along with attitudes matching the respective behaviours (bohner & dickel, 2011; haddock & maio, 2014) might be established. we therefore assume that such learning has a great impact on the development of communication competence and also influences the respective attitudes of young physicians. since little evidence exists in this respect, the present study fills a gap in the literature by tracing the development of communication-related attitudes during clinical electives. 1.3 logitudinal perspective on the development and change of attitudes finally, we argue that the present study is innovative because it adopts a longitudinal perspective in examining how attitudes towards communication develop. few empirical studies have pursued such a research focus (bombeke et al., 2011; woloschuk et al., 2004), which is disappointing because much medical education research on communication draws upon the implicit assumption that developmental processes occur in this area and that they can be influenced, for instance, through training (kurtz et al., 2016) or non-formal learning. insights into the naturally occurring processes of change over time regarding attitudes towards communication would therefore greatly amend this strand of research. different models exist that describe the development and change of attitudes. one model is the elaboration likelihood model (elm), which describes the process of elaborating arguments towards an object of thought and (re-)evaluating this object through the consideration of (new) information (bohner & dickel, 2011; stroebe, 2014). besides elaborating on (new) information, another two-path model referring to a heuristic approach is the heuristic-systematic model (hsm). the systemic path is comparable with the elm. information is elaborated in a complex procedure. the heuristic path refers to simple strategies like trusting experts, empirical statistics or even sympathetic people. people use the holistic approach if they lack the time to evaluate an object in an elaborative way (stroebe, 2014). a more recent model is the past model (past attitudes are still there), developed by petty et al. (2006). in this case, it is assumed that old attitudes are re-evaluated and, if necessary, corrected. the old attitude will be marked as invalid. furthermore, the theory of cognitive dissonance is one of the most popular with respect to attitude changes. the theory describes the phenomenon in which a person has to be in cognitive consonance between their attitudes and their behaviour. even if the attitudes are disaccorded with their behaviour or even their expected behaviour, cognitive dissonance emerges. according to the informal learning character of the clinical electives, we assume that the theory of cognitive dissonance is applicable insofar as the students have initial attitudes towards behavioural aspects of doctor-patient communication. at the wards, they will experience different situations and reactions to their behaviour. if this behaviour is contrary to their own attitudes, cognitive dissonance might occur. they have to adjust their attitude or their behaviour to avoid cognitive dissonance (haddock & maio, 2014). in the present study, we compare senior medical students’ attitudes towards communication with patients measured at three points in time: at the beginning, in the middle and at the end of their clinical elective year to investigate the following two research questions: 1) which attitude profiles of senior medical students regarding physician-patient communication can be differentiated at the beginning of the final year of medical school? 2) how do these attitude profiles develop during the students’ clinical electives in the final year of medical school? following the suggestions of watts and stenner (2005), we will explain the methodological approach used in the present study by (1) starting with a general overview of q-methodology, (2) describing the q-set design and its content, and (3) reviewing the participants of our study. we subsequently explain (4) how we administered the q-sort materials and (5) analysed the collected data. following the methods section, we present our results. first, the profiles of the first measurement point are presented, followed by the presentation of the profiles’ development across the clinical electives. the article concludes by discussing the results based on the research questions and the current state of research in the discussion. also, the limitations and further research perspectives are addressed in the discussion section. 2. methods 2.1 general overview of q-methodology william stephenson introduced q-methodology as a way to measure attitudes (stephenson, 1936; watts & stenner, 2005, 2012). this method combines aspects of quantitative and qualitative research and focuses on identifying groups of individuals with specific attitude profiles who are prevalent in larger samples (brown, 1996). q-method studies investigate such attitude profiles by means of card-sorting procedures (mchugh et al., 2018; mchugh et al., 2019; muddiman et al., 2019). the cards show different statements (the q-set), which respondents sort into a normal distribution grid (cf. figure 1). then, the participants judge each statement regarding its importance with respect to a specific criterion. as was mentioned above, this forced-choice procedure minimises certain response biases, such as tendencies towards the middle or one end of a scale (müller & kals, 2004). the different responses were analysed using q-factor analysis. its purpose is to differentiate the attitude profiles based on similarities between the collected q-sorts. the different profiles are narratively described according to how the collected statements are ordered by the respondents (watts & stenner, 2012). figure 1. q-sort grid 2.2 q-set design and content the q-set of our study consisted of the 34 anchor statements (as) of the german version of the kalamazoo communication skills assessment form for students’ self-appraisal (kcsafd-self, cf. table 2) (rider, 2010; schick et al., 2019). in its original use, the global items were assessed by raters, and the corresponding anchor statements were used as reference points for the assessment of the global items. the scale showed good internal consistency (cronbach’s α = .79; schick et al., 2019); it comprehensively describes key communication skills in doctor-patient dialogues and thus meets the requirements for a substantial q-set. in our study, we used the as instead of the nine global items because the as represent concrete behaviours relevant in doctor-patient dialogues. the nine global items were “building interpersonal relationships” (4 as), “taking the patients’ perspectives” (2 as), “showing empathy” (4 as), “opening the conversation” (3 as), “providing closure” (5 as), “gathering information” (4 as), “communicating information” (5 as), “sharing information” (3 as) and “reaching agreement” (4 as). the distribution grid (figure 1) consisted of seven columns ranging from “less important” (-3 to -1) to “important” (0) and “very important” (+1 to +3). the students had to select two items for the categories “+/3”, four items for the categories “+/2”, six items for the categories “+/1” and ten items for the category “0”. 2.3 participants and study design our target group consisted of senior medical students at the technical university of munich in the clinical elective period of their final year in 2017/2018. in germany, the clinical elective period lasts 48 weeks and is divided into three 16-week terms. each medical student has to conduct one term each in surgery and internal medicine and can choose where to spend the third term. we introduced our study at the welcome session for the local clinical elective students and invited them to participate. students were asked to take part in three measurements: in the first week (t1), after 24 weeks (t2) and close to the end of their clinical elective period (t3, after 46 weeks). the sample was a convenience one. participation was voluntary and rewarded with a book voucher of increasing value after completing the three q-sorts (t1 = 15 eur; t2 = 30 eur; t3 = 60 eur). a total of 47 final-year undergraduate medical students participated. this number decreased slightly over time (cf. table 1). ten students participated only at t1, another student participated at t1 and t2, and 36 students participated at all three measurement points. during the year, no further students were included in the study. we conducted a chi2 test to analyse the group differences between students who only participated in t1 with students who participated also at all three measurement points on the basis of their q-sort profile in t1. the chi2 test showed no significant difference between the students who participated at all three measurement points and those who only participated in t1, (ꭓ2 (2) = 2.137, p = .343). the ethics committee of the klinikum rechts der isar of the technical university of munich approved the study (project number: 482/17 s). table 1 descriptive statistics of sample size note. n = sample size; m = mean; sd = standard deviation; t1 = at the beginning of the electives; t2 = in the middle of the electives; t3 = at the end of the electives; 2.4 administering the q-sort procedure we implemented the q-study using the flashq template (hackert & braehler, 2007), which was adapted to our study purposes. the respective hyperlink was distributed to the participants via email at all three measurement times. each participant received an individual code, which remained stable across the different measurements and allowed matching of the corresponding q-sorts. the participants were asked the following guiding question when sorting their q-set: which aspects do you consider very important / important / less important in conducting a good medical interview with a patient? first, the participants reviewed all statements in random order and pre-sorted them into three columns (from “very important” to “less important”). then, students were asked to further sort the statements into the normal distribution grid (figure 1). finally, we asked the students to justify their choice of the so-called characterising statements (+/3) in a short open response format. 2.5 statistical analyses we ran a q-factor analysis using the r package “qmethod” (zabala, 2018). in order to determine the attitude profiles for each time point based on the collected q-sorts (t1-3), we conducted the following steps in the analysis: 1. we considered: a. the explained variance; b. eigenvalues (sum squared loading of all q-sorts on one factor, with ev ≤ 1 as acceptable); c. humphrey’s rule, which “states that a factor is significant if the cross-product of its two significant loadings (ignoring the sign) exceeds twice the standard error” (watts & stenner, 2012, p. 107); d. the amount of included individual q-sorts; e.the correlation between the attitude profiles at t1-3; and f. the interpretative content of the profiles to find the most suitable number of profiles for each time point. 2. the attitude profiles were interpreted qualitatively by identifying so-called characterising, distinguishing and consensus statements. the characterising statements were those at the end-points of the normal distribution grid labelled with the values “+/-3”. the distinguishing statements significantly differentiated the profiles from each other, whereas the consensus statements represented common statements between profiles. furthermore, the following categorical scheme was used: a. we ordered the statements of the q-set (see 2.2) by the corresponding nine global items of the kalamazoo communication skills assessment form (kcsaf) (table 2). b. these nine global items were clustered into four categories based on their content: interpersonal relationships and empathy, opening and closure of the conversation, gathering and communicating information, and shared decision making (table 2). 3. on the basis of the nine global items, we calculated the means of the q-values of the respective statements. the mean values of the categories were combined into one curve for each profile (figure 2). 4. the individual-to-individual correlations were calculated by using the factor loadings of each respondent’s individual q-sort at the different time points of the particular factor. the higher the correlation, the more similar the individual q-sorts were at the different time points. the development of the profiles was analysed using a combined quantitative and qualitative approach: 1. to estimate the degree of similarity between profiles, we used the number of distinguishing items. in the traditional q-methodology approach, the decision on how many q-sorts to extract is based on the number of distinguishing items. a smaller number of distinguishing items indicates greater similarity between different q-sorts. we used this approach to match similar profiles of the three measurement points. the profiles with the lowest number of distinguishing items were matched. for this calculation, we used an adopted version of the “qmethod” r-packages (zabala, 2018), where we entered the profiles of each measurement point and compared the position of the statements in each profile per measurement point. 2. we calculated correlations using the factor loadings of the participants in the corresponding profiles at each time point. 3. we interpreted the similarities and differences between the profiles at the three time points on the basis of the categorical scheme in a qualitative way, as described above. table 2 q-values of the three profiles at three time points (t1-t3) statements note. the wording of the original items of the gap kalamazoo communication skills assessment form: clinician (rider, 2010, pp. 74–76) has been slightly adapted for the q-set statements. 3. results the first research question concerns medical students’ attitude profiles regarding physician-patient communication at the beginning of the final year of medical school. the q-factor analysis determined a three-factor solution as most suitable for our data. table 3 presents the quantitative criteria of the two-, three-, and four-factor solutions (for further information, see also section 2.5). all three-factor solutions provided acceptable quantitative criteria, but the two factor solution included too many q-sorts per factor. this led to a lower degree of differentiation between the participants. the four-factor solution provided two factors with less than six participants; this outcome did not comply with the recommendations of watts and stenner (2012). therefore, the three-factor solution provided good quantitative and qualitative criteria to interpret the three factors. for the three-factor solution, the correlations between the profiles were r = .339 between profiles 1 and 2, r = .560 between profiles 1 and 3, and r = .306 between profiles 2 and 3 at t1. the three profiles explained 44.46% of the total variance. in the following sections, we narratively describe the profiles at t1 according to the four categories of interpersonal relationships and empathy, opening and closure of conversation, gathering and communicating information and shared decision making. the three profiles and their development are presented in figure 2. table 3 overview of psychometric properties of different factor solutions. figure 2. profile curves of all three time points based on means and standard deviations of q-values 3.1 profiles at the beginning of the final year of medical school (t1) table 4 distribution of students based on attitude profiles at all three measurement points note. n = number of participants; m = mean; sd = standard deviation 3.1.1 attitude profile #1: providing information and fostering shared decision making 3.1.1.1 interpersonal relationships and empathy respondents in profile #1 emphasised that the external appearance of the students should be appropriate, but regarded addressing the patient’s concerns and needs (7: +2; 11: +1; 26: +1; 29: -1; 27: -2)1 as less important for demonstrating empathy and building a relationship with the patient. regarding perspective taking, exploring the patient’s concerns and expectations (21: 0) was considered more important than addressing circumstances of the patient’s life (15: -2). 3.1.1.2 opening and closure of conversation individuals in profile #1 found opening and closing the conversation (28: -2; 13: -3) to be a less important aspect of a good doctor-patient dialogue. 3.1.1.3 gathering and communicating information one aspect individuals in this profile found very important was communicating accurate information. they also emphasised the importance of providing sufficient information to the patient and clearly explaining opportunities for diagnostic and therapeutic plans (18: +3; 2 +2). however, gathering information was rated as less important by the students in this profile. they stressed the importance of clarifying details and summarising during the dialogue (6: +1; 17: +1), but regarded appropriate questioning techniques as less important (10: -2; 20: -3). 3.1.1.4 shared decision making. shared decision making emerged as another important aspect in profile #1. sharing information in an understandable way (12: +3) and ensuring that the patients understood the problems (16: +1) were of high importance in this profile. the students considered the patient’s agreement with therapeutic plans (8: +2; 31:+2) as very important for a successful consultation. however, further concerns and needs of the patients were regarded as less important; also, to consider or even include the families in reaching agreements played a subordinate role (3: -1; 25: -1).  3.1.2 attitude profile #2: focusing on the patients’ concerns and emotions while meeting their desire for sufficient information 3.1.2.1 interpersonal relationships and empathy according to the respondents in profile #2, the development of a trustful interpersonal relationship is a very important aspect of physician-patient conversations. the students emphasised the importance of responding to the patients’ concerns (1: +2; 11: +3) and emotions (34: +1). empathy and compassion also played a very important role in this profile (30: +2). moreover, the students considered acknowledging the concerns, doubts and sorrows of patients regarding their complaints to be very important (21: +2). 3.1.2.2 opening and closure of conversation opening the conversation was also considered essential by respondents in profile #2: they rated the importance of “the patient should have the time to tell his or her story without interruption” as high (28: +1). the students also considered the behaviour of closing the conversation by answering open questions as important (14: +1). however, arranging additional appointments seemed less important in profile #2 (22: -1; 24: -1). 3.1.2.3 gathering and communicating information participants in profile #2 considered gathering information in a structured way using appropriate conversation techniques to be less important (6: -1; 10: -3; 17:0; 20: -3). moreover, the students evaluated communicating concrete information about the diseases and their possible progress (2: -1; 19: -2; 23: -2) as less important. 3.1.2.4 shared decision making sharing information was rated as very important for profile #2 students. they considered it important to explain information in an understandable way (12: +3) and consider the patient’s understanding of the problem and the desire for information (5: +2; 16: +1). in comparison to sharing information, reaching agreements played a secondary but still important role. the patient’s agreement with the diagnostic and treatment plans was considered to be very important (8: +1). 3.1.3 attitude profile #3: using appropriate conversation techniques to structure the communication and gather sufficient information 3.1.3.1 interpersonal relationships and empathy the respondents in this profile emphasised greeting the patients and showing interest in their concerns (1: +3). however, further aspects of building a good interpersonal relationship seemed to be less important for the profile (27: -3; 26: -1). the students also ranked empathy and compassion as being less important for a successful conversation with patients (29: -2). only behaving appropriately was emphasised as very important (7: +1). the students also considered patients’ perspectives on the disease to be less important (15: -1; 21: 0). 3.1.3.2 opening and closure of conversation the behavioural aspects of opening the conversation were assessed to be less important for the students in profile #3 (9: -2; 28: 0). in contrast, behaviours related to closing the conversation were seen to be more important, especially clarifying open questions and summarising at the end of the visit (14: +3; 33: +1). 3.1.3.3 gathering and communicating information one priority of profile #3 was gathering information and using appropriate techniques to structure the conversations (6: +2; 10: +2). the students’ preferences showed that they rated the importance of communicating information about the diagnostics and treatment plans (18: +1) higher than providing information about the severity and progress of the diseases (19: -1; 23: -2). 3.1.3.4 shared decision making for sharing information, the students considered the patient’s understandings of the offered information to be very important (5: +2; 12: +2; 16: +1). the students found that obtaining the patient’s agreement with the treatment had high importance, which matched that of offering enough information for reaching agreements (8: +1). in contrast, examining the patient’s comprehension of the treatment plans (31: -3) was considered less important. 3.2 development of attitude profiles across the clinical electives (t2–t3) our second research question was designed to address how the attitude profiles of senior medical students develop during their clinical electives in the final year of medical school. at t2 and t3, the q-factor analysis indicated that a three-factor solution was most suitable in this case as well. at t2, the three factors explained 46.87% of the total variance. at the end of the clinical electives (t3), the explained variance of the three factors was 48.99%. in the following sections, the development of the profiles is described according to the four categories of interpersonal relationships and empathy, opening and closure of conversation, gathering and communicating information and shared decision making. figure 3. migration of participants between profiles across the clinical electives 3.2.1 development of attitude profile #1 ten medical students remained in profile #1 at t2 (see figure 3). the average individual-to-individual correlation was r = .44 at t2. the correlation of the profile between t1 and t2 was r = .83. six medical students who represented attitude profile #1 at t1 and t2 were also clustered in profile #1 at t3. four students returned to profile #1 at t3 (see figure 3). the individual-to-individual correlation was r = .60 at t3. the correlation between the profiles of t2 and t3 was r = .85. 3.2.1.1 interpersonal relationships and empathy at t1, profile #1 members agreed with the importance of establishing and maintaining good interpersonal relationships in conversation with patients to a moderate to pronounced degree. as is apparent from figure 2, the profile #1 curve shows an overall increase of agreement with the items concerning interpersonal relationships during the electives. in detail, a slight decrease of agreement around the middle and an increase at the end of the clinical electives regarding the demonstration of empathy and compassion (7: t1 = +2; t2 = +1; t3 = +2)2 and consideration of the patients’ emotions and feelings (29: t1 = -1; t2 = -2; t3 = -1) is evident. in particular, the respondents’ attitudes towards considering patients’ perspectives changed over time. at the end of the clinical electives, agreement with consideration of the living circumstances of patients was higher (15: t1 = -2; t2 = 0; t3 = 0), as was agreement with listening to the patients’ concerns and sorrows (21: t1 =0; t2 = 0; t3 = +1). 3.2.1.2 opening and closure of conversation in the eyes of the profile #1 members, opening and closing the conversation still played a subordinate role in interaction with patients, especially in comparison with communicating information and shared decision making. at t2, the profile members showed higher preferences regarding aspects such as clarifying open questions (14: t1 = 0; t2 = +1; t3 = +1). however, closing the conversation and saying thanks to the patient were seen as less important throughout the clinical elective year (13: t1 = -3; t2 = -1; t3 = -2). 3.2.1.3 gathering and communicating information profile #1 was relatively stable over time regarding agreement with gathering and communicating information. the art of asking received less attention in the middle of the electives, but it again received more attention at the end of the electives (6: t1= +1; t2 = -2; t3 = +1). during their electives, the students considered communicating information about the severity and progress of the disease to be less important (19: t1 = 0; t2 = -1, t3 = 0; 23: t1 = 0; t2 = -1; t3 = -1). the aspects of clarifying future options of care and communicating enough information to patients remained very important for respondents in this profile (2: t1 = +2, t2 = +2, t3 = +2; 18: t1 = +3, t2 = +2, t3 = +3). 3.2.1.4 shared decision making the students still considered sharing information to be important for assessing the patient’s understanding and need for more information (16: t1 = +1; t2 = +1; t3 = +2). this profile also focused on reaching agreements with patients about treatment plans at t2 and t3 (8: t1 = +2, t2 = +3, t3 = +2). however, considering the patient’s needs increased in importance at the end of the clinical electives (3: t1 = -1; t2 = -2; t3 = 0). 3.2.2 development of attitude profile #2 six students were already in profile #2 at t1 (see figure 3). the individual-to-individual correlation of the attitude profile was r = .49 at t2. the correlation between the profile at t1 and t2 was r = .86. four students remained in attitude profile #2 during the clinical electives, and two medical students returned to profile #2 at t3 (see figure 3). the individual-to-individual correlation of the attitude profile was r = .60 at t3. the correlation between t2 and t3 was r = .82. 3.2.2.1 interpersonal relationships and empathy respondents in this profile put great emphasis on building a trustworthy interpersonal relationship at t1; this preference remained stable over time with only slight adjustments. students considered it important to respond explicitly to the feelings and emotions of the patients (27: t1 = -1; t2 =0; t3 =+1; 29: t1 = 0; t2 = +2; t3 = 0). overall, respondents regarded considering patients’ concerns and sorrows as important, yet seemed to struggle with this evaluation throughout the duration of the electives (21: t1 = +2; t2 = 0; t3 = +1). a comparable struggle was observable in the attitudes towards an appropriate behaviour regarding the nature of conversations (7: t1 = 0; t2 = +3; t3 = +1). 3.2.2.2 opening and closure of conversation the statements on opening and closing the conversation were estimated to be important throughout all three time points. the respondents regarded summarising the conversation at the end of the visit as increasingly important towards the end of the clinical elective (33: t1 = 0: t2 = 0: t3 = +2). however, closing the conversation and expressing gratitude for the patient’s confidence were estimated as less important during the clinical electives (13: t1 = 0; t2 = -2; t3 = -1). 3.2.2.3 gathering and communicating information the category gathering and communicating information was judged to be less important at t1. this judgement remained stable over time. some adjustments could be observed, while clarifying details with specific questions decreased in importance (6: t1 = -1; t2 = -1; t3 = -2). the students were reluctant to provide specific information about the severity of the disease as early as at t1. this assessment became even more pronounced over the course of the electives (19: t1 = -2; t2 = -1; t3 = -3). 3.2.2.4 shared decision making the students estimated consideration of the patient’s understandings to be a less important part of sharing information in t2 and t3 compared to its high importance at t1 (5: t1 = 2, t2 = 1, t3 = 0; 16: t1 = 1; t2 = 0; t3 = 0). the aspect of reaching agreement remained stable over time. only the aspect of considering the family’s perspective in the agreement process increased in importance during the clinical electives (25: t1 = 0, t2 = 0, t3 = +1). 3.2.3 development of attitude profile #3 at t2, three students were also in profile #3 at the beginning of the clinical electives (see figure 3). the individual-to-individual correlation was r = .53 at t2. the correlation between the profile in t1 and t2 was r = .49. at t3, three medical students who were in profile #2 at t1 had moved to profile #3 in t3. no student stayed in profile #3 during the clinical electives (see figure 3). the individual-to-individual correlation was r = .38 at t3. the correlation of profile #3 at t2 and t3 was r = .46. 3.2.3.1 interpersonal relationships and empathy as was already observed at t1, the students in profile #3 regarded building a trustworthy relationship as quite unimportant across the three measurements (27: t1 = -3, t2 = -1; t3 = -3; 1: 26: t1 = -1, t2 = 0, t3 = 0; 1: t1 =+3, t2 = +1, t3 = +1). they also did not agree on the relevance of considering the patients’ concerns and sorrows, their living situations, or even their living circumstances at all three time points. the aspect of empathy, however, especially behaving appropriately in a given situation, obtained more concern (29: t1 =-2, t2 = -1, t3 = 0; 7: t1 = +1, t2 = +2, t3 = +2). 3.2.3.2 opening and closure of conversation profile #3 showed an emphasis on closing the discussion, while opening the discussion played a secondary role at t1. setting an agenda at the beginning of the visit became more important at t3 (9: t1 = -2, t2 = -3, t3 = 0). summarising and closing at the end of the conversation seemed to be especially important to the students (33: t1 = +1, t2 = -2, t3 = +1; 13: t1 = 0, t2 = -2, t3 = +1). however, clarifying open questions decreased in importance over time (14: t1 =3, t2 = 0, t3 = -1). overall, the importance of behavioural aspects related to opening and closing the conversation played a secondary role in this profile. 3.2.3.3 gathering and communicating information the profile initially placed great emphasis on gathering information through appropriate conversation techniques. these attitudes were less preferred at t2 and t3 (10: t1 = +2, t2 = -2, t3 = -1; 20: t1 = 0, t2 = -3, t3 = -1). providing appropriate information was considered to be more important during the electives, but showed a decrease at the end of the electives (18: t1 = +1; t2 = +3, t3 = -1; 23: t1 = -2, t2 = +1, t3 = 0, 19: t1 = -1, t2 = +1, t3 = +1). overall, agreement with communicating information to patients increased during the clinical electives. 3.2.3.4 shared decision making the aspects of sharing information with patients and ensuring patient comprehension were considered to be more important over time compared to the beginning of the electives (12: t1 =+2; t2 = +3, t3 = +3; 16: t1 = +1; t2 = -1, t3 = +2). the students consistently emphasised the importance of the patient’s comprehension of the treatment plans (31: t1 = -3, t2 = 0, t3 = +3). in contrast, the families’ perspectives were considered to be less important in the decision-making process (25: t1 = -1, t2 = 0, t3 = -2). 4. discussion our study investigates attitudes towards different physician behaviours in communications with the patients of senior medical students during their final year of ume. as our first research question, we examined the attitude profiles of senior medical students at the beginning of their clinical electives in the final year of ume. we identified three different attitude profiles that focused either on communicating information and shared decision making (profile #1), fostering the doctor-patient relationship (profile #2) or structuring the conversation and gathering sufficient information (profile #3). these attitude profiles were developed in accordance with the research findings of roter et al. (1997), who differentiate five types of physicians using a cluster analysis of questionnaire data into “narrowly biomedical”, “expanded biomedical”, “biopsychosocial”, “psychosocial” and “consumerist” types. comparable profiles could be identified by flocke et al. (2002). instead of five profiles, flocke et al. (2002) described four different profiles: “biopsychosocial”, “biomedical”, “person-focused” and “high physician control”. attitude profile #1 considered informed shared decision making as the most important goal of the doctor-patient visit. this profile is comparable to “biopsychosocial” physicians who mainly elaborate on the concerns and medical conditions of the patients and ensure the patients’ understanding their illnesses and possible treatment plans (flocke et al., 2002; roter et al., 1997). the second profile focused on establishing a trustworthy interpersonal relationship between the patient and the physician. flocke et al. (2002) describe a group of physicians as “person-focused” if the physician is focusing more on the patients’ personal and emotional conditions than on the diseases themselves. the same profile was identified by roter et al. (1997), who referred to it as “psychosocial”. the physicians in attitude profile #3 considered behaviours of using appropriate conversation techniques to gather information and communicate accurate information as most important at t1. these physicians can be compared to those in the “biomedical” group (flocke et al., 2002) or the “narrowly biomedical” profile (roter et al., 1997). this profile includes physicians who concentrate on gathering and communicating information and are less focused on interpersonal issues and fostering an appropriate relationship (flocke et al., 2002; roter et al., 1997). the profiles “high-physician control” (flocke et al., 2002) and “consumerist” (roter et al., 1997) could not be found in our study. the profile “high-physician control” focuses on a doctor-centred communication style. in this view, the doctor leads the conversation and does not respond to the specific needs of patients (flocke et al., 2002). this approach to communicating with patients is not taught in ume. the predominant focus in ume communication seminars is on a patient-centred approach to communication (langewitz, 2012). a similar issue could be the reason why we could not find a profile similar to the “consumerist” one proposed by roter et al. (1997). this profile describes physicians who see themselves as consultants. they provide information to the patient while neglecting (psycho-)social exchange. a main aspect of today’s communication training is gathering information to achieve a holistic view of the patient’s concerns and needs and providing information to foster shared decision making (langewitz, 2012). as our second research question, we investigated the development of medical students’ attitude profiles during their clinical electives. first, we interpreted profile changes on the content level, and second, we interpreted migration between profiles. the most stable attitude profile on the content and migration levels was #2. the medical students in this profile assigned high relevance to building interpersonal relationships and considering patients’ concerns regarding their illnesses. also, an emphasis on statements about shared decision making, gathering and communicating information, and the opening and closure of conversation remained stable over time. profile #2 showed the least amount of migration across the clinical electives, which could account for the stability of the profile on the content level. this attitude pattern seemed to represent a successful approach to doctor-patient communication, which remained relatively stable during clinical electives. for the students with attitude profile #1, empathetic behaviour played an increasingly important role over time. at t1, the students judged general statements about empathetic behaviour to be more important, but statements about the individual needs of the patients were considered to be less important. over time, however, the students judged statements referring to the patients’ perspectives about their illnesses and concerns as more important for a successful doctor-patient dialogue. the development of profile #1 during the clinical electives was characterised by a high degree of migration from profile #1 to profile #3 at t2 and a constant shift of profile #1 participants in profile #3 at t3. this could be attributed to the stronger correlation of profiles #1 and #3 than the correlation between these two profiles and profile #2. in other words, the higher similarity of profiles #1 and #3 could substantiate the migration between the profiles. it is plausible that the two profiles overlapped at a certain point and that the factor loadings of the students on the two profiles were more similar than the loadings on profile #2. during the clinical electives, their q-sorts changed slightly, so that they moved from one profile to the other at t2 and t3, respectively. attitude profile #3 showed the highest degree of change in our study. participants in this profile rated statements about structuring conversations with the use of appropriate information techniques to be more important at t1. over time, their preferences changed towards considering the provision of information about the illnesses and their treatment options to be more important. finally, at the end of the electives (t3), shared decision making was prioritised, as were understanding the patients’ perspectives and establishing interpersonal relationships. nevertheless, the attitudes in this profile were still more physicianthan patient-centred. as mentioned in the introductory part of this text, we conceptualise the clinical electives as a non-formal learning setting. the purpose of the clinical electives is to allow young physicians to transfer the theoretical knowledge acquired during ume into practice. the medical students regularly and responsibly engage in doctor-patient dialogues and (should regularly) receive feedback from their supervisors. this feedback could support students in learning from their experiences. aside from conducting conversations themselves, the student can just as well observe dialogues between other physicians and their patients. this kind of observation can foster attitude changes as the students learn from role models, which has been described as the heuristic path in the hsm (stroebe, 2014). through reflection (i.e., after a working day), attitudes could change with the elaboration of new information and experiences associated with the observed conversation approach (bohner & dickel, 2011; stroebe, 2014). another explanation for the changes we observed in the profiles could be that the students were urged to behave in particular ways by the work demands imposed upon them. if the required behaviours contradicted their attitudes, they may have been urged to modify their attitudes to avoid cognitive dissonance (rather than to try to change their clinical practice work; haddock & maio, 2014). these possible attitude changes could be reflected in the changes within the profiles at the three time points as well as in the migration between the profiles at the different time points. in comparing the three profiles and their development, it became apparent that the aspects of building a trustworthy interpersonal relationship and showing empathy increased in importance during the clinical electives. in particular, understanding the patients’ perspectives became increasingly more important over time. the q-methodology defines highly correlated profiles as comparable profiles (watts & stenner, 2012; weber et al., 2009). all profiles showed medium to high correlations at the different time points. profiles #1 and #2 were highly correlated at the different time points, whereas profile #3 showed a medium to high correlation with the other profiles. according to watts and stenner (2012), highly correlated profiles can be assumed to have high similarity. also, the individual-to-individual correlations (which can be indicative of how the q-sorts of the individual participants were comparable over time) showed medium to high degrees of correlation at the time points. however, these findings were not significant due to the small number of participants per profile. however, in q-methodology, both qualitative and quantitative data are considered in the analyses. the quantitative data are just one part of the method to consider for interpretation, but the qualitative interpretation of the profiles also plays a role in the analyses. the participants as well as the factor structure were very stable over time, and this was especially evident in profile #2. students with this profile were already focusing on empathy and interpersonal relationships already at the beginning of their clinical electives. these findings are in line with those of smith et al. (2017), who measured the development of empathy in the first three years of medical school in a longitudinal study by means of the jefferson scale of physician empathy (jspe) and the questionnaire of cognitive and affective empathy (qcae). the results of the qcae showed an increase in taking the patients’ perspectives and emotional contagion. however, empathy measured by jspe showed a decrease during medical school. another meta-analysis revealed that there was a measurable decline of empathy as assessed with the jspe (spatoula et al., 2019). while these contradictory issues have been discussed in medicine and medical education, there is currently no consensus on whether, how and why empathy changes over time (spatoula et al., 2019). in our study, we could only make assumptions about the increased focus on empathy items in the development of all the profiles. the theory of attitude change posits that an attitude can change if one has had positive experiences with the attitude object. students may have had good experiences in their clinical electives associated with interacting empathetically with patients and therefore may have been able to strengthen interpersonal relationships. finally, increases in the importance of empathy and shared decision making were apparent in all profiles. involving the patient in the decision-making process is regarded as important for the establishment of a successful doctor-patient relationship (charles et al., 1997). all profiles that emerged from our study exhibited good psychometric properties. the explained variances of the factor solutions for each time point were above 40% and therefore in a good range (watts & stenner, 2012). in addition, the correlations between the profiles at the different time points were medium to high (field, 2013). on this basis, we felt it was safe to interpret and elaborate on the development of the factors. nonetheless, some limitations regarding the interpretation of the results need to be addressed. although our sample size was adequate for the q-methodology approach (watts & stenner, 2012), we could only study a convenience sample. the medical students participated voluntarily and were rewarded with book vouchers totalling 105 eur if they participated at all three time points. the migration of students between the profiles was another limitation. as is apparent from our description of the results, there was a certain degree of change regarding the individual students that formed the attitude profiles at different points in time. we therefore argue that we could only trace the more general dynamics and trends occurring in our sample when describing how the profiles transformed over time. we also asked the students which statements they considered very important / important / less important in order to conduct a good medical interview with a patient. the aim of this question was to focus on the students’ attitudes and not on their behaviour in doctor-patient dialogues. however, in accordance with the theory of attitude development, experiences of past behaviour also influence attitudes. therefore, it is impossible to say exactly whether the students in our study predominantly referred to their attitudes, their previous behaviour or their perceived communication skills. beyond these aspects, numerous individual-level dynamics emerged, which we cannot fully describe here. further issues, such as demographic properties and process variables, may emerge during the clinical electives as relevant predictors of individual transitions between profiles. in addition, our sample size was not large enough to find significant results or to interpret them accordingly. moreover, it is evident that the q-sort technique, just like any empirical method, is also affected by measurement error. one point worth noting here is that the structure of the q-sort grid (cf. figure 1) limits the degrees of freedom regarding the sorting of statements. this means that respondents could, for instance, not assign the highest level of importance to several items, even if this would perfectly reflect their personal viewpoints. moreover, our study only provides evidence on how the medical students’ attitude profiles developed over time. future studies should focus on the question of how these profiles relate to students’ actual behaviours in communication with patients. to conclude, this study investigated the attitude profiles of medical students at the beginning of their electives and their development over time using q-methodology. the profiles differed in regard to their emphasis on supporting shared decision making, establishing trustworthy relationships and structuring conversations. the importance of the empathetic aspect and shared decision making increased over time in all the profiles. future research should investigate the interdependence between the attitudes and observed behaviours of the students as physicians in order to gain more information about the relationship between these variables in medical interviews. keypoints we investigate how senior medical students’ attitudes towards doctor-patient communication develop through informal, workplace-based learning. the use of q-methodology to assess these attitudes in a longitudinal perspective is an innovative methodological approach. our study reveals three different attitude profiles that show different developmental patterns over time. acknowledgements the study was part of the project äkhom, reference number: 01pk1501c and the project volea, reference number: 16dhb2133, both funded by the german ministery of education and research. footnote 1(statement number: q-value) 2(statement number: t1 = q-value; t2 = q-value; t3 = q-value) references ajzen, i. (1991). the theory of planned behavior. organizational behavior and human decision processes , 50(2), 179–211. https://doi.org/10.1016/0749-5978(91)90020-t ajzen, i., fishbein, m., lohmann, s., & albarracin, d. (2019). the influence of attitudes on behavior. in d. albarracín & b. t. johnson (eds.), the handbook of attitudes (pp. 197–255). routledge. archer, r., elder, w., hustedde, c., milam, a., & joyce, j. (2008). the theory of planned behaviour in medical education: a model for integrating professionalism training. medical education , 42(8), 771–777. https://doi.org/10.1111/j.1365-2923.2008.03130.x berings, m. g. m. c., poell, r. f., & simons, p. r.-j. (2005). conceptualizing on-the-job learning styles. human resource development review , 4(4), 373–400. https://doi.org/10.1177/1534484305281769 blömeke, s., gustafsson, j.-e., & shavelson, r. j. (2015). beyond dichotomies: competence viewed as a continuum. zeitschrift für psychologie / journal of psychology , 223(1), 3–13. https://doi.org/10.1027/2151-2604/a000194 bohner, g., & dickel, n. (2011). attitudes and attitude change. annual review of psychology, 62, 391–417. https://doi.org/10.1146/annurev.psych.121208.131609 bombeke, k., symons, l., vermeire, e., debaene, l., schol, s., winter, b. de, & van royen, p. (2012). patient-centredness from education to practice: the ‘lived’ impact of communication skills training. medical teacher, 34(5), e338-48. https://doi.org/10.3109/0142159x.2012.670320 bombeke, k., van roosbroeck, s., winter, b. de, debaene, l., schol, s., van hal, g., & van royen, p. (2011). medical students trained in communication skills show a decline in patient-centred attitudes: an observational study comparing two cohorts during clinical clerkships. patient education and counseling, 84(3), 310–318. https://doi.org/10.1016/j.pec.2011.03.007 brown, j. (2010). transferring clinical communication skills from the classroom to the clinical environment: perception of a group of medical students in the united kingdom. academic medicine , 85, 1052–1059. https://doi.org/10.1097/acm.0b013e3181dbf76f brown, s. r. (1996). q methodology and qualitative research // q methodology and qualitative research. qualitative health research , 6(4), 561–567. https://doi.org/10.1177/104973239600600408 charles, c., gafni, a., & whelan, t. (1997). shared decision-making in the medical encounter: what does it mean? (or it takes at least two to tango). social science & medicine , 44(5), 681–692. https://doi.org/10.1016/s0277-9536(96)00221-3 cross, r. m. (2005). exploring attitudes: the case for q methodology. health education research , 20(2), 206–213. https://doi.org/10.1093/her/cyg121 edmondson, a. (1999). psychological safety and learning behavior in work teams. administrative science quarterly , 44(2), 350. https://doi.org/10.2307/2666999 epstein, r. m. (2007). assessment in medical education. the new england journal of medicine , 356, 387–396. https://doi.org/10.1056/nejmra054784 epstein, r. m., & hundert, e. m. (2002). defining and assessing professional competence. jama, 287 (2), 226. https://doi.org/10.1001/jama.287.2.226 epstein, r. m., & street, r. l. (2011). the values and value of patient-centered care. annals of family medicine , 9(2), 100–103. https://doi.org/10.1370/afm.1239 eraut, m. (2000). non-formal learning and tacit knowledge in professional work. british journal of educational psychology , 70(1), 113–136. https://doi.org/10.1348/000709900158001 fallowfield, l., jenkins, v., farewell, v., saul, j., duffy, a., & eves, r. (2002). efficacy of a cancer research uk communication skills training model for oncologists: a randomised controlled trial. the lancet , 359(9307), 650–656. https://doi.org/10.1016/s0140-6736(02)07810-8 fazio, r. h., & olsen, m. a. (2014). the mode model: attitude-behavior processes as a function of motivation and opportunity. in j. w. sherman, b. gawronski, & y. trope (eds.), dual-process theories of the social mind (pp. 155–171). guilford publications. field, a. (2013). discovering statistics using ibm spss statistics: and sex and drugs and rock ‘n’ roll (4th edition). mobilestudy. sage. flocke, s. a., miller, w. l., & crabtree, b. f. (2002). relationship between physician practice style, patient satisfaction, and attributes of primary care. the journal of family practice , 51(10). giroldi, e., veldhuijzen, w., geelen, k., muris, j., bareman, f., bueving, h., van der weijden, t., & van der vleuten, c. (2017). developing skilled doctor-patient communication in the workplace: a qualitative study of the experiences of trainees and clinical supervisors. advances in health sciences education: theory and practice. advance online publication. https://doi.org/10.1007/s10459-017-9765-2 grassi, l., giraldi, z., messina, e. g., valle, e., & cartei, g. (2000). physicians’ attitudes to and problems with truth-telling to cancer patients. support care cancer , 8, 40–45. https://doi.org/10.1007/s005209900067 ha, j. f., surg anat, d., & longnecker, n. (2010). doctor-patient communciation: a review. the ochsner journal , 10, 38–43. hackert, c., & braehler, g. (2007). flashq [computer software]. http://www.hackert.biz/flashq/home/ haddock, g., & maio, g. r. (2014). einstellungen [attitudes]. in k. jonas, w. stroebe, m. hewstone, & m. reiss (eds.), springer-lehrbuch. sozialpsychologie (6th ed., pp. 197–229). springer. https://doi.org/10.1007/978-3-642-41091-8_6 hartig, j. (2008). psychometric models for the assessment of competencies. in j. hartig, e. klieme, & d. leutner (eds.), assessment of competencies in educational contexts (pp. 70–90). hogrefe & huber. klooster, p. m. ten, visser, m., & jong, m. d.t. de (2008). comparing two image research instruments: the q-sort method versus the likert attitude questionnaire. food quality and preference, 19(5), 511–518. https://doi.org/10.1016/j.foodqual.2008.02.007 krupat, e., rosenkranz, s. l., yeager, c. m., barnard, k., putnam, s. m., & inui, t. s. (2000). the practice orientations of physicians and patients: the effect of doctor-patient congruence on satisfaction. patient education and counseling , 39, 49–59. https://doi.org/10.1016/s0738-3991(99)00090-7 kurtz, s., silverman, j., & draper, j. (2016). teaching and learning communication skills in medicine (2nd ed.). crc press. http://gbv.eblib.com/patron/fullrecord.aspx?p=4711378 kyndt, e., dochy, f., & nijs, h. (2009). learning conditions for non‐formal and informal workplace learning. journal of workplace learning, 21(5), 369–383. https://doi.org/10.1108/13665620910966785 langewitz, w. (2012). zur erlernbarkeit der arzt-patienten-kommunikation in der medizinischen ausbildung [physician-patient communication in medical education: can it be learned?]. bundesgesundheitsblatt, gesundheitsforschung, gesundheitsschutz , 55(9), 1176–1182. https://doi.org/10.1007/s00103-012-1533-0 levinson, w., lesser, c. s., & epstein, r. m. (2010). developing physician communication skills for patient-centered care. health affairs (project hope), 29(7), 1310–1318. https://doi.org/10.1377/hlthaff.2009.0450 locatelli, c., piselli, p., cicerchia, m., & repetto, l. (2013). physicians’ age and sex influence breaking bad news to elderly cancer patients. beliefs and practices of 50 italian oncologists: the g.i.o.ger study. psycho-oncology , 22(5), 1112–1119. https://doi.org/10.1002/pon.3110 loh, a., & simon, d. (2009). gemeinsame entscheidungsfindung von arzt und patient: das konzept des „shared decision making“ [shared decision making between doctor and patient: the concept of “shared decision making”]. in t. langer & m. w. schnell (eds.), das arzt-patient patient-arzt gespräch (pp. 143–152). hans marseille verlag gmbh. manuti, a., pastore, s., scardigno, a. f., giancaspro, m. l., & morciano, d. (2015). formal and informal learning in the workplace: a research review. international journal of training and development , 19(1), 1–17. https://doi.org/10.1111/ijtd.12044 mchugh, n., baker, r., biosca, o., ibrahim, f., & donaldson, c. (2019). who knows best? a q methodology study to explore perspectives of professional stakeholders and community participants on health in low-income communities. bmc health services research , 19(1), 35. https://doi.org/10.1186/s12913-019-3884-9 mchugh, n., van exel, j., mason, h., godwin, j., collins, m., donaldson, c., & baker, r. (2018). are life-extending treatments for terminal illnesses a special case? exploring choices and societal viewpoints. social science & medicine (1982) , 198, 61–69. https://doi.org/10.1016/j.socscimed.2017.12.019 mehay, r., & burns, r. (2009). miller’s pyramid of clinical competence (1990). https://www.essentialgptrainingbook.com/ch29/ muddiman, e., bullock, a. d., hampton, j. m., allery, l., macdonald, j., webb, k. l., & pugsley, l. (2019). disciplinary boundaries and integrating care: using q-methodology to understand trainee views on being a good doctor. bmc medical education , 19(1), 59. https://doi.org/10.1186/s12909-019-1493-2 müller, f. h., & kals, e. (2004). q-sort technique and q-methodology: innovative methods for examining attitudes and opinions. forum qualitative sozialforschung / forum: qualitative social research , 5(2). http://www.qualitative-research.net/index.php/fqs/article/download/600/1302 petty, r. e., tormala, z. l., briñol, p., & jarvis, w. b. g. (2006). implicit ambivalence from attitude change: an exploration of the past model. journal of personality and social psychology , 90(1), 21–41. https://doi.org/10.1037/0022-3514.90.1.21 rees, c., sheard, c., & davies, s. (2002). the development of a scale to measure medical students’ attitudes towards communication skills learning: the communication skills attitude scale (csas). medical education , 36(141-147). https://doi.org/10.1046/j.1365-2923.2002.01072.x rider, e. a. (2010). interpersonal and communication skills. in e. a. rider & r. h. nawotniak (eds.), a practical guide to teaching and assessing the acgme core competencies (2nd ed., pp. 1–137). hcpro, inc. rider, e. a., & nawotniak, r. h. (eds.). (2010). a practical guide to teaching and assessing the acgme core competencies (2nd edition). hcpro, inc. roter, d., stewart, m., putnam, s. m., lipkin, m., stiles, w., & inui, t. s. (1997). communication patterns of primary care physicians. jama , 277, 350–356. https://doi.org/10.1001/jama.1997.03540280088045 schick, k., berberat, p. o., kadmon, m., harendza, s., & gartmeier, m. (2019). german language adaptation of the kalamazoo communication skills assessment form (kcsaf): a multi-method study of two cohorts of medical students. zeitschrift für pädagogische psychologie, 33(2), 135–147. https://doi.org/10.1024/1010-0652/a000241 smith, k. e., norman, g. j., & decety, j. (2017). the complexity of empathy during medical school training: evidence for positive changes. medical education, 51(11), 1146–1159. https://doi.org/10.1111/medu.13398 spatoula, v., panagopoulou, e., & montgomery, a. (2019). does empathy change during undergraduate medical education? a meta-analysis. medical teacher , 1–10. https://doi.org/10.1080/0142159x.2019.1584275 stephenson, w. (1936). the inverted factor technique. british journal of psychology , 26, 344–361. https://doi.org/10.1111/j.2044-8295.1936.tb00803.x strack, f., & deutsch, r. (2004). reflective and impulsive determinants of social behavior. personality and social psychology review , 8(3), 220–247. https://doi.org/10.1207/s15327957pspr0803_1 stroebe, w. (2014). strategien zur einstellungsund verhaltensänderung [attitude and behavior change strategies]. in k. jonas, w. stroebe, m. hewstone, & m. reiss (eds.), springer-lehrbuch. sozialpsychologie (6th ed., pp. 231–268). springer. tynjälä, p. (2008). perspectives into learning at the workplace. educational research review , 3(2), 130–154. https://doi.org/10.1016/j.edurev.2007.12.001 watts, s., & stenner, p. (2005). doing q methodology: theory, method and interpretation. qualitative research in psychology , 2(1), 67–91. https://doi.org/10.1191/1478088705qp022oa watts, s., & stenner, p. (2012). doing q methodological research: theory, method and interpretation . sage. weber, t., danielson, s., & tuler, s. (2009). using q method to reveal social perspectives in environmental research . greenfield, ma. social and enviromental research institute. www.serius.org/pubs/qprimer.pdf woloschuk, w., harasym, p. h., & temple, w. (2004). attitude change during medical school: a cohort study. medical education , 38(5), 522–534. https://doi.org/10.1046/j.1365-2929.2004.01820.x zabala, a. (2018). package “qmethod” (version 1.5.4) [computer software]. https://cran.r-project.org/web/packages/qmethod/qmethod.pdf codepen laine et al publication frontline learning research vol.8 no. 2 (2020) 90 108 issn 2295-3159 individual interest and learning in secondary school stem education erkka laine b, marjaana veermansa, andreas gegenfurtner b& koen veermansa auniversity of turku, finland bdeggendorf institute of technology, germany article received 27 february 2019/ revised 23 december/ accepted 26 march 2020 / available online 29 april abstract interest research offers different hypotheses about the association between interest and learning outcomes. the standard hypothesis proposes that interest predicts learning outcomes: people acquire new knowledge about a topic they find interesting. the affective by-product hypothesis assumes that learning predicts interest: by learning something, people develop an interest in this topic. finally, the reciprocal hypothesis states that interest and learning covary. this longitudinal study aimed to test the predictive validity of these three hypotheses in the context of secondary school stem education. the participants were 104 finnish 7th grade students aged 12-14. data were collected at three times during the school year through questionnaires and grade evaluations in mathematics and biology. a partial least squares (pls) path modeling approach was used to determine the relationships between interest and course grades across the three measurement points: at the beginning of the autumn semester, at the beginning of the spring semester, and after the spring semester at the end of the school year. the results differed between the autumn and spring semesters: during the autumn semester, students’ interest predicted their grades, whereas during the spring semester, grades predicted their interest. these findings indicate that the relationships between students’ individual interest towards science and mathematics with learning vary. as a practical implication, more focus should be put on when and what type of performance feedback is given to students with differing interest profiles. keywords: interest; learning; stem; partial least squares (pls) path modeling info corresponding author: email: etlain@utu.fi doi: https://doi.org/10.14786/flr.v8i2.461 1. introduction previous research literature has shown that students’ interest in learning science, technology, engineering, and mathematics (stem) varies at different ages (osborne, simon & collins, 2003; hofer, 2010). the most notable change usually takes place during the transition from elementary to secondary school (krapp & prenzel, 2011). during that period, some students seem to start losing interest in investing effort into stem learning while the interest of others takes a deeper and more long-lasting form. this decline in interest is particularly alarming from the perspective of how modern societies will be able to respond to a multitude of science related challenges in the future. for example, the european commission’s report “does the eu need more stem graduates” (european commission, 2015) estimates that the need for high qualification stem jobs will generally increase throughout europe by 2025. this is expected to coincide with a diminishing number of low qualification level jobs due to increased digitalization and robotics. although long-term predictions of this kind are highly uncertain, it is likely that the labour market will continue shifting towards more knowledge-intensive employment and the demand for highly-skilled jobs will increase (ec, 2015). this will pose challenges for educational systems in terms of how to ensure that students will not only learn relevant skills and knowledge in schools but can also be fostered to develop a lasting interest towards science. the latter is particularly important because early interest in stem careers has been found to predict student persistence in science and the choice of a science-related major in college (tai, liu, maltese, & fan, 2006). aside from encouraging adolescents to pursue careers in stem professions, stem education is also important from the perspective of giving students the qualifications to become scientifically literate citizens in their adult lives. this includes not only providing adequate knowledge content for different stem topics, but also inspiring a long-lasting interest in it, as well as a basic understanding of the scientific method. to be able to do this means that the teaching provided in schools should be both cognitively satisfying, and at the same time encourage students to adopt a positive and curious mindset towards stem. there are several possible reasons behind the declining trends in students’ interest to learn stem subjects. it may be that the way school education is organized, the curriculum or the quality and type of instruction do not provide enough support for students’ interest to develop further. another explanation relates to some of the psychological demands that adolescents confront in their lives, which may cause them to view academic learning as less important compared to other aspects of life. thirdly, it may be that students’ view of themselves as learners, their ideal self-concept, may become separated from stem domains and hence cause them to become disinterested in investing effort in those topics (krapp & prenzel, 2011). the relationship between interest and learning has been studied widely in the past research literature (krapp & prenzel, 2011) and recently some researchers have turned their focus on examining the directionality of the two concepts. while the prevalent view has been that interest is an antecedent for learning, some researchers have taken a different approach and aimed to determine how the acquisition of knowledge on a certain subject influences learners’ interest towards it (rotgans & schmidt, 2017b). if it would be so that acquiring knowledge about a school subject in fact generates and predicts students’ interest later on, then this would have consequences on how schools should design their pedagogical approaches. this would emphasize the importance of cognitive support in education and perhaps give less weight on constantly trying to come up with new and more exciting ways to get students to engage with the study topic. the aim of this study was to contribute to this discussion by examining how the relationships between students’ interest and learning in stem subjects develop and vary during the time period of one school year. although the relationship between students’ interest in stem and their academic achievement has already been studied longitudinally (i.e. köller et al., 2001), the time points at which interest and achievement outcomes were measured have often spanned over many years. in addition, the focus of these studies has mostly been on how interest predicts students’ course choices or college majors in the stem domain, and a more detailed view on what takes place during a single school year is needed. for this reason, the current study concentrated on revealing predictive patterns within a time frame that would be long enough to see changes but at the same time short enough to see interaction patterns. 1.1 the phases of interest development interest can be conceptualized as a phenomenon that arises from the interaction between a person and his or her environment (hidi & renninger, 2006), and which produces experimental modes that have both positive cognitive and affective qualities (krapp & prenzel, 2011). the cognitive qualities might include for example personally meaningful goals or viewing the activity or topic as valuable for the person’s future. affective qualities, then, might for example relate to feeling enjoyment when interacting with the activity, or engaging deeply with the topic at hand. theoretical literature usually acknowledges two different types of interest, namely situational and individual interest. these two types differ from each other in how much they are based on affect, knowledge and value, as well as the temporal duration of interest. situational interest is viewed to be more affect based and temporarily fleeting, whereas individual interest is considered to connect more with the individual’s values and acquired knowledge on the subject, and be more stable over time (krapp, hidi & renninger, 1992; hidi & renninger, 2006; renninger & hidi, 2011; rotgans & schmidt, 2017b). both of these forms of interest can be viewed to represent different analytical levels, where situational interest refers to the actual and on-going process of engaging an activity, and individual interest the relatively stable tendency to invest time and effort in the topic of interest (krapp & prenzel, 2011). current interest theory often further divides these two different types of interest into several sub-categories which are thought to reflect their developmental phases. krapp (2002) uses a three-category model, which divides interest into three phases of development, namely emerging situational, stabilised situational, and individual interest. hidi and renninger (2006) further add to this model a fourth phase by dividing individual interest into emerging and well-developed individual interest. a person’s interest is theorized to develop through each of these phases consecutively, starting from the situational interest being triggered and leading to well-developed individual interest through continued engagement with the topic, as well as increased value and knowledge acquisition (hidi & renninger, 2006; renninger & hidi, 2011). this study concentrates on individual interest; how its emerging or well-developed phases might manifest in students’ self-reported levels of interest to study mathematics and biology during the duration of one school year. in particular, it aims to see how their interest levels reflected their grades in these subjects, and whether the grades they received predict their interest levels later on. the period of one school year was chosen as the duration of the study because it represents the natural annual cycle of students’ schoolwork, and at the same time is long enough for changes to take place in their reported interest levels or received grades, which could in turn affect their grades or reported interest levels. the context of this study was closely tied to a real-life setting in the school’s everyday life, and the aim was to acquire a general understanding of what happens to students’ interest during this one year of formal education. following the four-phase model by hidi and renninger (2011), individual interest was conceptualized as consisting of two phases, namely emerging individual interest, and well-developed individual interest. these two phases represent the last stages of interest development, when the activity, topic or domain is viewed to be personally valued, relates to the person’s existing knowledge structures, and is intrinsically motivated (hidi & renninger, 2006; krapp & prenzel, 2011). an emerging individual interest is characterized not only by positive feelings, but also by stored value and knowledge. the activity itself is of value to the person and he or she would engage it in any case if given the option to choose. this phase is usually viewed to be self-generated, although it may at times require external support from peers or experts and can, in the context of education, be affected by instructional conditions or the learning environment. this can then lead to the last phase of the model, namely well-developed individual interest which is viewed as a psychological state of interest as well as a more or less enduring tendency to engage the topic or object of interest. it has very much the same attributes as the emerging individual interest, but with more stored value and knowledge on the topic or activity. it is also viewed to be mostly self-regulated, but can benefit from instructional designs and learning environments that offer opportunities to gain knowledge through challenging tasks and interaction. although, the development of interest should be viewed as a continuum, separating different phases in it can have theoretical and practical value. an actually operating interest can become generated either through an already existing disposition, i.e. individual interest or through the special conditions that take place in a teaching or learning situation, i.e. the interestingness of the situation (krapp & prenzel, 2011). individual interest is affected by the environmental and situational factors that take place in different learning situations. school education, with its separate lessons in different school subjects, can be viewed as a continuum in which a student’s situational interest towards the study topic may change from one time to another. a student’s experiences of focused attention and positive emotions in one learning situation may affect his or her interest in another situation later on, and gradually start to develop towards a more sustained form of interest. meaningfulness of the task and personal involvement are seen as the pre-requisite for a person to acquire an individual interest towards a domain (hidi & renninger, 2006). anchoring the phases into different intra-individual processes that take place throughout one’s interest development can for example, help teachers to adjust their teaching according to the different needs of students (krapp & prenzel, 2011). in this study the focus was on examining the relationships between students’ individual interest in the subjects of mathematics and biology, and their knowledge acquisition on those subjects. the aim was to examine how the predictive relationships between interest and knowledge acquisition might vary during the course of one school year. we chose to concentrate on measuring individual interest on the subject level, since this is the dominant level on which students’ learning outcomes are measured in schools. there exist also more finely grained methods to measure individual interest, for example on the sub-domain level (e.g. geometry in mathematics, or photosynthesis in biology). however, by concentrating our analyses on the level of school subjects, we aimed to provide information relevant to the work of educators and teachers, who could also benefit from the information that the results provide. 1.2 interest and learning interest has been found to have facilitating and mediating effects on learning outcomes. this has been observed in different contexts and settings, such as writing (albin, benton, & khramtsova, 1996), studying psychology (harackiewicz, durik, barron, linnenbrink & tauer, 2008), learning statistics (hay, callingham & carmichael, 2015), and reading science texts (ainley, hidi & berndorff, 2002). in addition, interest has been found to have a positive connection with other motivational factors such as mastery goals (harackiewicz et al., 2008), utility value (gegenfurtner, knogler & schwab, 2020), and self-efficacy (ainley, 2012). the traditional view in the research literature has seen interest as the prerequisite, or at least the facilitator of learning. in this study, we will call this the standard hypothesis as it is the most common way to define the relationship between interest and learning (rotgans & schmidt, 2017b; 2017c). the main idea behind this is that in order for learning to take place the individual has to become interested in the topic either through the support of one’s individual interest or through arousing situational factors. individual interest can even in its less-developed form help to generate situational interest through its relation to students’ prior knowledge and value for the subject. in school education context students encounter study domains and topics that do not necessarily rank high on their list of individual interests, but they may view the information provided at the lesson as important or realize its value for completing their studies successfully. a student who is able to connect with the study content and develop strategies for working with it is more likely to start developing curiosity questions towards the topic. these curiosity questions in turn increase the student’s sense of possibilities for learning and increase the perceived value of the studied topic, which may in time become realized as more well-developed individual interest (renninger, 2000). many pedagogical approaches, such as inquiry learning (renninger et al., 2014), problem-based learning (rotgans & schmidt, 2011), multi-user virtual environments (chen et al., 2015), and game-based learning (knogler, harackiewicz, gegenfurtner, & lewalter, 2015; rodríguez‐aflecht et al., 2018) have been developed with the aim of increasing learners’ interest towards the studied topic. all these approaches rely, at least in part, on the idea that by making the learning situation more engaging and enjoyable to the student it increases their interest, and ultimately leads to better learning outcomes. however, recent research literature has highlighted the fact that there exists incongruence between the direct effects of interest in learning, and that the empirical findings are not as self-evident as previously suggested (köller, baumert, & schnabel, 2001; nieswandt, 2007; tapola, veermans, & niemivirta, 2013). despite a significant amount of empirical research existing on interest-enhancing practices in educational settings, a considerable portion of the studies report only partial improvements or small effect sizes. this has been especially evident among sub-groups of students with differing levels of pre-existing interest (hulleman & harackiewicz, 2009; renninger et al., 2014; rodríguez‐aflecht et al., 2018). this has caused some researchers to question whether the relationship has been viewed from the correct perspective. in their recent work, rotgans and schmidt (2017a; 2017b) raise the question of what would happen if the relationship between interest and learning, or knowledge acquisition, is reversed so that learning precedes interest. this proposition gains some theoretical support from the notion that individual interest may help students to remain situationally interested during learning situations, and according to the four-phase model, individual interest is more knowledge-based and less reliant on affective fluctuations (krapp & prenzel, 2011; renninger, 2000). hence it would be plausible to expect that some amount of learning has to take place before the student can start developing a longer lasting individual interest towards the study topic. in this study we call this the affective by-product hypothesis. the third hypothesis stemming from the work of rotgans and schmidt (2017b) posits that interest and learning both affect each other’s development reciprocally. the lack of research on the directionality of the interest-learning relationship had already surfaced some years earlier. for example, köller et al. (2001) already raised the question whether academic achievement necessarily follows interest, or that it could be the other way around so that those students who feel more competent in the study subject could also generate interest towards it more easily. in the reciprocal hypothesis, interest and learning are seen to interact with each other in varying degrees over time, so that their relative emphasis in the learning process differs from time to time. in their analyses, rotgans and schmidt (2017b) did not find support for the so-called standard hypothesis that individual interest would precede learning. instead they found a significant path coefficient (standardized β=0.20, p < .05) between the knowledge measured at first time point, and the interest measured at second time point; thus, supporting the affective by-product hypothesis. this means that the amount of knowledge students had at the beginning of the experiment seemed to influence their interest towards the topic at the end. lastly, they did not find support for the reciprocal hypothesis that interest and knowledge acquisition affect each other reciprocally. the purpose of this current study was to test these hypotheses in a classroom environment over the duration of one school year (9.5 months). 2. research question and hypotheses this study aimed to explore longitudinally how students’ individual interest relates to their learning outcomes in mathematics and biology during a school year. for this purpose, and based on the theoretical discussion that was presented earlier, this study addressed the following research question: what is the relationship between interest and learning in mathematics and biology education? to examine the relationship between students’ interest and learning we formulated three theoretically derived hypotheses. these hypotheses and their relation to the theoretical model can be seen in table 1. the partial least squares (pls) structural equation model that was constructed to examine the relationship between the variables can be seen in figure 1, along with the three groups of hypotheses. table 1 the three models and the hypotheses figure 1. hypothesized pls model for individual interest and grades in mathematics and biology. in the first group of hypotheses based on previous research literature (hidi & renninger, 2006; ainley, hidi, & berndorf, 2002), interest was expected to predict learning. for this, we formulated the so-called standard hypotheses. this theoretical approach suggests that interest development precedes learning and because of this, students’ interest at time points 1 and 2 should predict their grades at time points 2 and 3 respectively. h1.1: interest at time 1 predicts learning outcomes at time 2. h1.2: interest at time 2 predicts learning outcomes at time 3. second, based on the findings by rotgans and schmidt (2017b), learning was expected to predict interest. for this we formulated the so-called affective by-product hypothesis. if individual interest would require some initial knowledge acquisition to become generated, then the students’ grades at time point 2 should predict their interest at the end of the school year at time 3. h2: learning outcomes at time 2 predict interest at time 3. the third possible line of reasoning was based on the idea that knowledge and interest may influence each other reciprocally. in other words, interest would facilitate students’ learning, and therefore increased knowledge would in turn cause interest to increase (rotgans & schmidt, 2017b). to test this, we formulated the reciprocal hypothesis in which interest was expected to be reciprocally related to learning outcomes at the same time points. since pls analysis does not allow testing bidirectional relationships at the same time, we formulated two separate models which we called primary and alternative models. these two models differed only in terms of the direction of the simultaneous relationship between interest and grades at time points 2 and 3. the idea was that if interest and learning had reciprocal relationships during the autumn semester and spring semesters, then these relationships would manifest in the pls analyses as predictive connections between the two variables at simultaneous time points 2 and 3, or be indicated by changes in these predictive relationships during the school year. h3: interest and learning outcomes predict each other at one time point or over time 3. method 3.1 participants and study design the participants were 104 (53 girls, 51 boys) 7th grade students aged 12-14 from six different classes in a lower secondary school in southern finland. the study setting was longitudinal, following the same students through their first year of lower secondary school. formal consent was obtained from parents for the students’ participation, and students without a letter of consent were excluded from the study as participants. since the collection of those letters was organized by the different class teachers there is no exact number on how many students were excluded due to no consent. it can however be estimated that the number was not high, most probably less than 10%. data collection occurred at three time points. time 1 was at the beginning of the autumn semester (0.5 months after starting the school year), time 2 was at the beginning of the spring semester (4.5 months), and time 3 was at the end of the spring semester (9.5 months). students’ interest in mathematics and biology were measured on all the time points, and their learning on time points 2 and 3. in the school, teaching was divided into 5 periods each lasting about 8 weeks. each class of students received the same amount of teaching during the school year, but the courses may have taken place in different periods. this also explains why mathematics and biology were chosen as the subjects, since they were the only two stem subjects that the students received teaching during both autumn and spring semester. although there was no additional demographic data collected from the participants, it can be said that schools in finland are generally very homogenous in terms of student population. majority of them receive their funding through public sources, which could give support for the sample being representative. it is also not customary to collect such demographic data in finnish schools, and in terms of the focus point of this study, these types of data were not relevant. 3.2 measures 3.2.1 individual interest in mathematics and biology an instrument of tapola et al. (2013) was used to assess students’ individual interest in mathematics and biology. to measure interest in mathematics, a single item was used (“how interested are you in mathematics”) with a five-point scale, ranging from 1 (not at all interested) to 5 (very interested). similarly, a five-point scale single item, ranging from 1 (not at all interested) to 5 (very interested) was used to measure students’ individual interest in biology (“how interested are you in biology”) with a five-point scale, ranging from 1 (not at all interested) to 5 (very interested). single-item scales have previously been used to measure interest (ainley, 2006; palmer, 2009; tulis & ainley, 2011; tapola, jaakkola, & niemivirta, 2014) and since we were concentrating on students’ interest in these study subjects on a generalised level, we adopted this view of measuring. the situation in which the students were asked about their interest took place outside learning situations during the school day. this arrangement reduced the possibility of students connecting the question to any particular learning situation in their everyday studies. 3.2.2 learning outcomes students’ learning outcomes were measured by grade level evaluation after each semester on time points 2 and 3. grade evaluation was done by the subject teachers and was based on students’ learning and performance throughout the whole duration of each semester. in the finnish system students’ evaluation is based on both formative and summative assessment. the teachers observe the students’ development throughout the duration of the course and most often also use final tests to evaluate students’ learning outcomes at the end of each course. these both forms of evaluation are used to determine the course grades for each student. in this study students’ grades were used to indicate their skill and knowledge levels, and these were indicated by grades from 4 (failed) to 10 (excellent). 3.3 data analysis correlation analysis and pls modelling were used in the data analysis. the correlation analysis provides a more global view on relations between individual variables while the pls modelling provides an integrated model combining all variables in one model. descriptive statistics and correlation analyses were carried out using spss version 21 (ibm 2012). the structural equation models were modelled using the warppls software. partial least squares (pls) is a structural equation modeling (sem) technique which can simultaneously test the measurement model through relationships between indicators and their corresponding constructs, and the structural model through relationships between constructs (gil-garcia, 2008). pls is efficient when working with small sample sizes and complex models, and it does not assume the data to be normally distributed (hair, hult, ringle & sarstedt, 2017). instead of assessing overall model fit, pls is an approach for predicting relationships in a model which were the focus in this study. 4. results 4.1 correlational analyses the correlation analyses presented in table 2, show that students’ individual interest at time 1 was positively correlated with their grades at time 2 and time 3 in both mathematics and biology. in mathematics time 1 interest and time 2 grade had a moderate positive correlation (r(101) = .41, p < .01). in biology there was a moderate positive correlation between time 1 interest and time 2 grade (r(102) = .28, p < .01). these are in line with the standard hypothesis h1.1. table 2 correlation of individual interest and subject grades in mathematics and biology however, interest at time 2 did not correlate with grade at time 3 in either subjects, which is not in line with the standard hypothesis 1.2. mathematics grade at time 2 had a moderate correlation with mathematics interest at time 3 (r(91) = .44, p < .01). similar correlation was also found in biology (r(92) = .30, p < .01). these results support the affective by-product hypothesis h2. regarding the reciprocity of interest and learning outcomes, interest in mathematics at time 2 was significantly correlated with mathematics grade at time 2 (r(79) = .25, p < .05) albeit the correlation being rather small. interest at time 3 also significantly but also more sizeably correlated with mathematics grade at time 3 (r(91) = .47, p < .01. in biology correlation between interest time 2 and biology grade at time 2 was nonsignificant, but interest at time 3 and grade at time 3 had a positive correlation (r(92) = .39, p < .01). these findings only very weakly support the reciprocal hypothesis h3. students’ individual interest exhibited signs of stability across measurement times. in mathematics time 1 interest correlated strongly with time 2 interest (r(80) = .53, p < .01) and time 3 interest (r(92) = .66, p < .01). the relationship remained at a similar level from time 2 to time 3 (r(73) = .61, p < .01) in biology the results were similar but the effect sizes were smaller. interest in biology at time 1 had a moderate correlations with time 2 interest (r(80) = .43, p < .01) and time 3 interest (r(92) = .45, p < .01). from time 2 to time 3 there was a strong correlation between the interest measures (r(73) = .52, p < .01). overall these results indicate some level of stability, but also of changes. in both mathematics and biology the grades across time were also highly correlated. in mathematics the effect size between time 2 and time 3 grades was (r(100) = .83, p < .01), and in biology (r(102) = .74, p < .01). 4.1.1 mean level differences in students’ interests and grades between classes because the participants were spread to six different classes there was a possibility for the data to be nested differently in these classes. for this, intraclass correlation coefficients (icc) were estimated for each of the measured variables for each of the six classes in order to control that their variances did not differ significantly from each other. the icc values were estimated through variance components estimation using the type iii sum of squares. specific cut-off values for icc that would require the use of multilevel methods usually range from 0.10 (e.g., lee, 2000; koo & li, 2016) to 0.25 (e.g., bowen & guo, 2011). the analysis revealed excellent reliability (icc < .10) for interest in mathematics, interest in biology, and mathematics grade at all three measurement points and for biology grade at time 2. for biology grade at time 3 the icc estimate was 0.11. the low icc values, and based on previous literature about the homogeneity of finnish school classes in terms of mathematics learning (brezovszky et al., 2019), indicated that the classes were homogenous enough, and that there was no need for multi-level methods to be used in the analyses. 4.2 partial least squares (pls) modeling the main differences between pls modelling and more traditional methods of structural equation modelling, such as cbm-sem or regressions based on sum-scores, are how they treat the latent variables included in the model. in cb-sem the constructs are considered as common factors that explain the covariation between the indicators that are associated to the constructs. when estimating the model parameters in cb-sem, the scores of these common factors are not known or needed. in pls-sem the constructs are represented through proxies; the weighted composites of indicator variables to that particular construct. this relaxes the assumption that all the covariation between the sets of indicators are caused by a common factor, and also facilitates accounting for measurement error, which gives it an advantage when compared with multiple regression using sum scores. another advantage that pls-sem has, is its ability to produce a single specific score for each composite of each observation by establishing weights for each proxy. pls-sem estimates coefficients that aim to maximize the r2 values of each target variable, giving it the ability to estimate predictive patterns between the model constructs. hence, pls-sem is a suitable method when the aim of the research is to develop theory and explain variance between the constructs (hair et al., 2017) for the pls analysis, two sets of hypothesized path models, the primary and the alternative models were constructed for both mathematics, and biology. the models consisted of the individual interest variable (mathematics or biology) at time points 1, 2, and 3, and learning outcome variables (grades) at time points 2 and 3. the primary path models, one in each subject domain, were aimed at clarifying the standard hypotheses h1.1 and h1.2 and the affective by-product hypothesis h3. in addition, alternative models were constructed in order to complement the analyses on the behalf of the reciprocal hypothesis h3. 4.2.1 collinearity assessment to check for collinearity in the structural model, the average collinearity variance inflation factor (avif) values were obtained from the model analyses. this was done by looking at each predictor construct separately and estimating how much their variance is artificially increased by other predictor constructs in the model. an avif value higher than 5 exhibits a critical value (hair et al., 2017). the avif values were 1.24 for the primary mathematics model, and 1.53 for the alternative mathematics model. in biology the avif values were 1.07 for the primary model, and 1.31 for the alternative model. these results showed that collinearity was not a critical issue in any of the structural models. 4.2.2 coefficient of determination to evaluate the predictive power of a structural model, a commonly used measure is the coefficient of determination value r2. the coefficient represents the amount of variance that is explained by all of the exogenous constructs in the model that are linked to a certain endogenous construct. (hair et al., 2017). in this study all of the exogenous constructs consisted of one-item measures, some of which also functioned as endogenous constructs. therefore, the r2 values obtained in the analyses represent the amount of variance in a given construct explained by all of the constructs linked to it in the model. the links between constructs and the hypothesized directions of the predictive effects are indicated by arrows, and can overall be seen in figure 1, and specified by the primary and alternative models, as well as by the mathematics and biology domains in figures 2, 3, 4, and 5. the average r2 value for the primary mathematics model was .37, and .37 for the alternative model. in biology these values were .28 for the primary model, and .29 for the alternative model. among the primary models for mathematics and biology the lowest r2 value was obtained in the biology grade at time 2 construct (r2 = .09) meaning that only 9% of its variance was explained by the two constructs linked to it, namely interest in biology at time 1 and time 2. the highest value was obtained in the mathematics grade at time 3 construct where the r2 value was .69. while this construct had three explaining constructs linked to it, namely interest in mathematics at time 2 and time 3, and mathematics grade at time 2 making a higher r2 value understandable the difference is still considerable. in the alternative models for mathematics and biology the results were very similar, with the lowest r2 value again being with biology grade at time 2 (r2 = .08), and the highest with mathematics grade at time 3 (r2 = .66). 4.2.3 predictive relevancy of the models following the recommendation of hair et al., (2017) stone-geisser’s q2 values were obtained through a blindfolding procedure in order to examine the models’ predictive relevancies. in pls path modeling, predictive relevancy means that the model also accurately predicts data that has not been used in the model estimation. in the blindfolding procedure data is re-used by deleting data points systematically and providing them a prediction of their original values, by treating them as missing values in the model. these values are then compared to the original data in order to determine the prediction error between the predicted data points and the true omitted data points. the sum of squared prediction errors is used to calculate the q² value. q2 values higher than 0 suggest that the endogenous construct is relevantly predicted by the model (hair et al., 2017). the q2 values for mathematics variables ranged between .19 and .68 in both the primary and the alternative models, and for biology between .08 and .58 in the both models which means that all models had predictive relevancy for the constructs. 4.3 results from the structural equation models in this section the results of the partial least squares analyses are presented. the results of the primary pls models can be seen in figure 2 for mathematics, and figure 3 for biology. in addition, the results for the alternative models that aimed to complement the analyses for the reciprocal hypotheses h3, by converting the direction of the relationships between interest measures and learning outcomes at simultaneous time points, are presented in figure 4 for mathematics, and figure 5 for biology. 4.3.1 standard hypotheses examining the first research hypotheses of whether interest predicts learning outcomes, the focus was on the primary model, which can be seen in figures 2 and 3. students’ interest in mathematics at time 1 predicted their mathematics grade at time 2 (β = 0.42, p < .01). in biology a similar pattern was found with time 1 interest also predicting biology grade in time 2 (β = 0.26, p < .01). however, the predictive relationship between interest at time 2 and grades at time 3 were found to be nonsignificant in both subjects. these findings supported the standard model hypothesis h1.1 but not h1.2. figure 2. primary partial least squares model of individual interest in mathematics and mathematics grades. figure 3. primary partial least squares model of individual interest in biology and biology grades. 4.3.2 affective by-product hypotheses examining the second research question of whether learning outcomes predict interest the results were again obtained from the primary model. this choice was made because the direction of the predictive relationships concerning the affective by-product hypothesis h2 did not differ between the primary and the alternative models. results of the analyses showed that mathematics grade at time 2 predicted interest in mathematics at the end of the school year at time 3 (β = 0.32, p < .01). similarly, biology grade at time 2 predicted interest in biology at time 3 (β = 0.25, p < .01). these findings supported affective by-product hypothesis h2. 4.3.3 reciprocal hypotheses results from the primary model analyses revealed that in both mathematics and biology, students’ interest at time 2 did not predict their grades at time 2 as the results were nonsignificant. however, there was a significant predictive effect from interest to grades at time 3 in both mathematics (β = 0.15, p < .05) and biology (β = 0.16, p < .05). based on these results the reciprocal hypothesis h3 was supported only at time 3. as it was already mentioned earlier, pls analysis does not allow testing bidirectional relationships at the same time. this shortcoming was compensated by formulating an alternative model in which the direction of the relationships between interest and grades were reversed at time 2 and time 3. in the alternative model the results were similar to the standard model. in both mathematics and biology, students’ grades at time 2 did not predict their interest at time 2 as the results were nonsignificant. this meant that no support was found for hypothesis h3.1. at time point 3 students’ grades did, however, predict their interest in both mathematics (β = 0.32, p < .01) and biology (β = 0.36, p < .01) again providing partial support for the reciprocal hypothesis h3. figure 4. alternative partial least squares model of individual interest in mathematics and mathematics grades figure 5. alternative partial least squares model of individual interest in biology and biology grades. 5. discussion the current study examined how students’ individual interest in mathematics and biology related to their learning outcomes in these subjects. the study objective was to test the predictive validity of three different sets of hypotheses that were introduced by rotgans and schmidt (2017b) in the context of mathematics and biology education in secondary schools: the standard hypothesis, the affective by-product hypothesis, and the reciprocal hypothesis. when looking at the predictive validity of the standard hypothesis, namely whether or not interest was a predictor of learning outcomes, the results differed in the autumn and spring semester. in the autumn semester students’ interest had a predictive effect on learning outcomes in both mathematics and biology. this finding was also supported by the results of the correlational analyses. when looking at the spring semester the results were quite different. in this semester, the predictive effect of interest towards learning outcomes was not found in either of the two subjects, and the correlations also became non-significant. this indicated at best partial support for the standard hypotheses that interest is an antecedent for knowledge acquisition. in the second hypothesis, the framing was reversed, and the focus was on whether students’ learning outcomes predicted their individual interest. in this alternative model both mathematics and biology students’ grades at time 2 predicted their interest at the end of the school year. in addition, the correlation analyses revealed moderate positive correlations on both subjects across the two measurement times. in other words, students who received higher grades at the mid-semester evaluation were more likely to express higher levels of interest in the subject at the end of the school year. these results offer support for the affective by-product hypothesis. similar to the standard hypotheses, the results for reciprocal hypotheses of interest and learning differed between semesters. at time 2, interest did not predict learning outcomes at the same time point in either of the two subjects. correlational analyses showed only a weak positive relationship between interest and grade in mathematics, while in biology this was not observed. however, at the end of the school year at time 3 interest did have a predictive effect on grades in both subjects and correlations were moderate. similar pattern was also found in the alternative model where students’ grades at time 2 did not predict their interest at the same time point, but at time 3 they did. when looking over the course of the whole year however both models show a path from interest to interest through grades. this seems to indicate that the relationship between students’ interest and learning outcomes did not stay the same throughout the school year, which is in line with the reciprocity view. to summarize, the standard hypothesis was supported only during either the autumn semester or the spring semester, but not throughout the school year. this finding was consistent across the two subjects. thus, the affective by-product hypothesis was supported, but because it could only be tested across two measurement times between time 2 and time 3, namely spring semester, drawing too strict conclusions of these results would be premature, especially considering the fact that there were indications for reciprocity both on one time point and across time points one explanation for these findings could relate to the differences between the measurement situations. at time 1 measurement point the students had just started their journey through secondary education and had changed to a new school with new teachers and new classmates. in this kind of environment, new external factors may have affected their interest. this aligns with the idea that in the interplay between interest and learning interest goes through an evolutional process. at first, when a student has insufficient knowledge about the topic, situational interest needs to be triggered and re-triggered before individual interest develops. over time the effect of situational interest decreases, and more value and knowledge based individual interest becomes the decisive factor in learning (rotgans & schmidt, 2017c). changes during this process (e.g. grades) may reinforce or disrupt this development. 5.1. practical implications interest and its relation to other motivational factors as well as learning outcomes have been studied extensively in the past (e.g. ainley et al., 2002; hidi & renninger, 2006; harackiewicz et al., 2008; krapp & prenzel, 2011). previous literature has often concluded that helping students become interested in the topic at hand would also help them to achieve better learning outcomes. however, evidence for interest directly predicting learning outcomes has largely been missing and this has given rise to alternative interpretations of the role interest has in the learning process. following the ideas of rotgans and schmidt (2017b; 2017c) this study tested three sets of hypotheses as regards to the possible relationship that interest and learning might have during secondary school students’ school year. based on the results we present two conclusions that should be taken into consideration in future research. firstly, the relationship between individual interest and learning found in this study underline that this relationship is not stable throughout time, but can exhibit changes, both positive and negative, during the time period of one school year. over longer periods of time the relationship may vary; a student might, for instance, express individual interest towards a topic or a subject but may not be able to achieve the learning outcomes he or she wished for, which may affect interest. from the perspective of educators these findings point out the importance of designing curriculums and learning environments so that they offer experiences of success for each student. aiming for high academic standards in schools is of course a desirable goal for education but it can also lead to some adverse effects if it is seen as the sole purpose of teaching. receiving low grades can have a negative impact on students’ self-perceived ability and may quell their interest towards the study subject (baumert, schnabel & lehrke, 1998). supporting students’ interest in studies should in itself be viewed as an important goal since it has been found to relate to their career choices later on in life (maltese & harsh, 2015). supporting students’ interest towards mathematics and science is also relevant from the perspective of 21st century skills, since societies are becoming increasingly technology-driven, and navigating in them in the future requires skills that are many times taught in stem subjects in schools. technology may also offer solutions for the issue of students becoming disengaged in stem learning because of negative performance assessment. digital learning environments can offer more individualized feedback to each learner and at the same time offer more accurate support for learning as well as learning tasks that are more finely balanced in terms of difficulty. 5.2. theoretical implications one theoretical contribution that this study has to offer is the longitudinal setting spanning over a whole school year that revealed patterns that are not easy to explain within the existing frameworks. previous research literature has usually focused only on narrow time frames where hard to expect any interest development or very long time frames where development may occur, but fluctuations may also easily stay out of sight. there is no clear consensus in the research literature what would constitute an appropriate timeline for interest development. in their study, knogler et al. (2015) carried out a science study intervention over the course of three weeks. their findings suggest that situational interest is as its name indicates, situational, and does not transfer to other learning situations to great extent. the case with individual interest is less clear and especially the process of how and when situational interest starts to evolve into individual interest is subject to debate. rotgans and schmidt (2017c) criticize the four-phase model of interest development (hidi & renninger, 2006; renninger & hidi, 2011) for being too simplified and vague. their suggestion is that situational and individual interests differ, especially in the way they are connected to knowledge. situational interest arises from a knowledge gap that the person wants to fill, whereas individual interest can only start to develop once the person has acquired some knowledge of the object of interest. within this view situational interest is not something that only precedes individual interest, but always exists, depends on past experiences of interest and knowledge development, and may influence further interest and knowledge development in the situation. the findings of this study seem to indicate that mathematics and biology differ somewhat in terms of stability of interest as well as in relation to learning outcomes. in biology (m = 3.33) students rated their interest, on average, higher than in mathematics (m = 2.88) throughout the school year. this is in line with findings from previous research that biology is the more popular science subject (baram-tsabari, 2015) among school students. however, when looking at the correlations, the students’ interest in biology did not correlate as strongly across time points as it did in mathematics, thus indicating lower level of stability. what is perhaps even more surprising is that students’ interest in biology at the beginning of school year did not correlate with their grade either at time 2 or time 3. one explanation for this could be that biology as a subject is less clear than mathematics, causing interest in the subject to be also less stable. 5.3. future directions and limitations the measurement used to measure individual interest in this study consisted of only one item, and although similar one-item instruments have been used in previous studies (tapola et al., 2013; tapola et al., 2014), it still warrants the question of how would a more fine-grained instrument have affected the results. as data for this study was collected as a part of a larger research project the choice to limit the questionnaire items was practical; keeping the questionnaire compact enough not to risk overburdening the participants. in the future, it would be recommended to widen the instrument so that it could take into consideration the value component of interest. combining this with qualitative methods, such as interviews, would provide a better understanding of the processes that affect students’ interest development over time. one limitation could also be that information about students’ learning outcomes were obtained only twice during the school year. although students’ grades are the normal way of evaluating school performance, it may be that the pressure for students to receive a better final grade during the spring semester is greater than during the first half of the school year. because their initial knowledge levels on stem subjects were not controlled at the beginning of the school year, there was no exact way of knowing how much learning had taken place between time points 1 and 2. however, the participants were 7th graders who were on their first year of secondary education that normally lasts for three years. it can be argued that the pressure to perform well increases towards the end of 9th grade, when they need to start making choices about their future education, and the possible negative consequences of that also probably have a bigger effect. another limitation relates to the results of the pls analyses. some of the r-squared values in the model are quite low which means that the model was able to explain only a relatively small amount of the variance in some of the variables. for example, variances in students’ time 2 grades were only explained by 19% through their individual interest in both mathematics and biology. this leaves the question of what other, perhaps latent variables, would account for the rest of the variance. in addition, and perhaps more surprisingly, a large portion of variance in students’ interest at the end of the school year was left unexplained by this model. this calls for further research to investigate what other factors, external and internal, affect students’ interest development during the school year. keypoints the predictive validity of three hypotheses on the relationship between interest and learning during a school year were tested interest and learning outcomes had a reciprocal relationship that alternated during the school year. future studies would benefit from combining a longitudinal setting with more detailed student profiles. as a practical implication, instead of just grading, offering students' more detailed feedback on their performance across learning situations may foster their interest towards stem. new learning technologies could provide support for teachers to receive information on their students' interest development and offer possibilities for more individualized learning paths. acknowledgements this research was partially supported by the finnish cultural foundation. references ainley, m. (2012) students’ interest and engagement in classroom activities. in s. christenson, a. reschly, & c. wylie (eds.), handbook of research on student engagement (pp. 283–302). springer. https://doi.org/10.1007/978-1-4614-2018-7_13 ainley, m. d., hidi, s., & berndorff, d. (2002). interest, learning, and the psychological processes that mediate their relationship . journal of educational psychology, 94(3), 1–17. https://doi.org/10.1037/0022-0663.94.3.545 albin m. l., benton s. l., khramtsova i. (1996). individual differences in interest and narrative writing. contemporary educational psychology, 21(4), 305–324. https://doi.org/10.1006/ceps.1996.0024. baram-tsabari, a. (2015). promoting information seeking and questioning in science. in k. a. renninger, m. nieswandt, & s. hidi (eds.), interest in mathematics and science learning (pp. 135–152). american educational research association. https://doi.org/10.3102/978-0-935302-42-4 baram-tsabari, a., & yarden, a. (2010) quantifying the gender gap in science interest.international journal of science and mathematics education, 9(3), 523–550. https://doi.org/10.1007/s10763-010-9194-7 baumert, j., schnabel, k., & lehrke, m. (1998). learning math in school: does interest really matter? in l. hoffmann, a. krapp, k. a. renninger, & j. baumert (eds.), interest and learning (pp. 327–336). kiel: ipn. bong, m., lee, s. k., & woo, y.-k. (2015). the roles of interest and self-efficacy in the decision to pursue mathematics and science. in k. a. renninger, m. nieswandt, & s. hidi (eds.), interest in mathematics and science learning (pp. 33–48). american educational research association. https://doi.org/10.3102/978-0-935302-42-4 bowen, n. k., & guo, s. (2011). structural equation modeling. oxford university press. https://doi.org/10.1093/acprof:oso/9780195367621.001.0001 brezovszky, b., mcmullen, j., veermans, k., hannula-sormunen, m. m., rodríguez-aflecht, g., pongsakdi, n., laakkonen e., & lehtinen, e. (2019). effects of a mathematics game-based learning environment on primary school students’ adaptive number knowledge. computers & education, 128, 63–74. https://doi.org/10.1016/j.compedu.2018.09.011 chen, j. a., tutwiler, m. s., metcalf, s. j., kamarainen, a., grotzer, t., & dede, c. (2016). a multi-user virtual environment to support students’ self-efficacy and interest in science: a latent growth model analysis. learning and instruction, 41, 11–22. https://doi.org/10.1016/j.learninstruc.2015.09.007 european commission (ec)(2015). does the eu need more stem graduates? publications office of the european union. https://doi.org/10.2766/000444 gegenfurtner, a., knogler, m., & schwab, s. (2020). transfer interest: measuring interest in training content and interest in training transfer. human resource development international, 23(2), 146–167. https://doi.org/10.1080/13678868.2019.1644002 gil-garcia, j. r. (2008). using partial least squares in digital government research. in g. d. garson & m. khosrow-pour (eds.), handbook of research on public information technology (vol. 1, pp. 239–253). information science reference. https://doi.org/10.4018/978-1-59904-857-4.ch023 hair, j. f., hult, g. t. m., ringle, c. m., & sarstedt, m. (2017). a primer on partial least squares structural equation modeling (2nd edition.). los angeles: sage. harackiewicz, j. m., durik, a. m., barron, k. e., linnenbrink-garcia, l., & tauer, j. m. (2008). the role of achievement goals in the development of interest: reciprocal relations between achievement goals, interest, and performance. journal of educational psychology, 100(1), 105–122. https://doi.org/10.1037/0022-0663.100.1.105 hay, i., callingham, r., & carmichael, c. (2015). interest, self-efficacy, and academic achievement in a statistics lesson. in k. a. renninger, m. nieswandt, & s. hidi (eds.), interest in mathematics and science learning (pp. 203–224). american educational research association. https://doi.org/10.3102/978-0-935302-42-4 hidi, s., & renninger, a. (2006). the four-phase model of interest development. educational psychologist, 41(2), 111–127. https://doi.org/10.1207/s15326985ep4102_4 hofer, m. (2010). adolescents’ development of individual interests: a product of multiple goal regulation? educational psychologist, 45 (3), 149–166. https://doi.org/10.1080/00461520.2010.493469 hulleman, c. s., & harackiewicz, j. m. (2009). promoting interest and performance in high school science classes. science, 326 (5958), 1410–1412. https://doi.org/10.1126/science.1177067 knogler, m., harackiewicz, j. m., gegenfurtner, a., & lewalter, d. (2015). how situational is situational interest? investigating the longitudinal structure of situational interest. contemporary educational psychology, 43, 39–50. https://doi.org/10.1016/j.cedpsych.2015.08.004 koo, t., & li, m. (2016). a guideline of selecting and reporting intraclass correlation coefficients for reliability research. journal of chiropractic medicine, 15(2), https://doi.org/10.1016/j.jcm.2016.02.012. krapp, a., hidi, s., & renninger, k. a. (1992). interest, learning and development. in k. a. renninger, s. hidi, & a. krapp (eds.), the role of interest in learning and development (pp. 3–25). erlbaum. krapp, a., & prenzel, m. (2011). research on interest in science: theories, methods, and findings. international journal of science education, 33(1), 27–50. https://doi.org/10.1080/09500693.2010.518645 köller, o., baumert, j., & schnabel, k. (2001). does interest matter? the relationship between academic interest and achievement in mathematics. journal for research in mathematics education, 32(5), 448–470. https://doi.org/10.2307/749801 lee, v. e. (2000). using hierarchical linear modeling to study social contexts: the case of school effects, educational psychologist, 35(2), 125–141, https://doi.org/10.1207/s15326985ep3502_6 maltese, a. v., & harsh, j. a. (2015). students’ pathways of entry into stem. in k. a. renninger, m. nieswandt, & s. hidi (eds.), interest in mathematics and science learning (pp. 203–224). https://doi.org/10.3102/978-0-935302-42-4 nieswandt, m. (2007). student affect and conceptual understanding in learning chemistry. journal of research in science teaching, 44(7), 908–937. https://doi.org/10.1002/tea.20169. osborne, j., simon, s., & collins, s. (2003). attitudes towards science: a review of the literature and its implications. international journal of science education, 25, 1049–1079. https://doi.org/10.1080/0950069032000032199 renninger, k. a. (2000). individual interest and its implications for understanding intrinsic motivation. in c. sansone & j. m. harackiewicz (eds.), intrinsic and extrinsic motivation: the search for optimal motivation and performance (pp. 373–404). academic press. https://doi.org/10.1016/b978-012619070-0/50035-0 renninger, k. a., austin, l., bachrach, j. e., chau, a., emmerson, m. s., king, b. r., riley, k. r., & stevens, s. j. (2014). going beyond whoa! that’s cool! achieving science interest and learning with the ican intervention. in s. karabenick & t. urdan (eds.), motivation-based learning interventions: advances in motivation and achievement series (vol. 18, pp. 107–140). https://doi.org/10.1108/s0749-742320140000018003 renninger, k. a., & hidi, s. (2011). revisiting the conceptualization, measurement, and generation of interest. educational psychologist, 46(3), 168–184. https://doi.org/10.1080/00461520.2011.587723 rodríguez‐aflecht, g., jaakkola, t., pongsakdi, n., hannula-sormunen, m., brezovszky, b., & lehtinen, e. (2018). the development of situational interest during a digital mathematics game. journal of computer assisted learning, 34(3), 259–268. https://doi.org/10.1111/jcal.12239 rotgans, j. i., & schmidt, h. g. (2011). situational interest and academic achievement in the active-learning classroom. learning and instruction, 21, 58–67. https://doi.org/10.1016/j.learninstruc.2009.11.001 rotgans, j. i., & schmidt, h. g. (2017a). how individual interest influences situational interest and how both are related to knowledge acquisition: a microanalytical investigation. the journal of educational research,. 111(5), 530–540. https://doi.org/10.1080/00220671.2017.1310710 rotgans, j. i. & schmidt, h. g. (2017b). the relation between individual interest and knowledge acquisition. british educational research journal, 43(2), 350–371. https://doi.org/10.1002/berj.3268 rotgans, j. i. & schmidt, h. g. (2017c). the role of interest in learning: knowledge acquisition at the intersection of situational and individual interest. in p. a. o'keefe & j. m. harackiewicz (eds.) the science of interest (pp. 69–93). cham: springer. tai, r. h., liu, c. q., maltese, a. v., & fan, x. (2006). planning early for careers in science. science, 312(5777), 1143–1144. https://doi.org/10.1126/science.1128690 tapola, a., jaakkola, t., & niemivirta, m. (2014). the influence of achievement goal orientations and task concreteness on situational interest. journal of experimental education, 82(4), 455–479. https://doi.org/10.1080/00220973.2013.813370 tapola, a., veermans, m., & niemivirta, m. (2013). predictors and outcomes of situational interest during a science learning task. instructional science, 41(6), 1047–1064. https://doi.org/10.1007/s11251-013-9273-6 uitto, a., juuti, k., lavonen, j., & meisalo, v. (2006). students’ interest in biology and their out-of-school experiences. journal of biological education, 40(3), 124–129, https://doi.org/10.1080/00219266.2006.9656029 frontline learning research vol. 6 no. 1 (2018) 1-18 issn 2295-3159 corresponding author: lisette hornstra, department of education, utrecht university, po box 80140, 3508 tc, utrecht, the netherlands, t.e.hornstra@uu.nl doi:10.14786/flr.v6i1.305 a dual pathway of student motivation: combining an implicit and explicit measure of student motivation lisette hornstraa, antoinette kamsteega, sara pota, & lydia verheija autrecht university, the netherlands article received 22 october 2017 / revised 18 january / accepted 8 february / available online 2 march abstract abundant research in social psychology shows human behaviour is guided by beliefs through two pathways, a deliberate and automatic pathway. research on student motivation has thus far focused mostly on the deliberate pathway and consequently almost exclusively relied on explicit measures (i.e. self-reports of motivation) to assess student motivation and subsequently predict student behaviour and achievement. the purpose of this study was to examine whether student motivation is associated with students’ behavioural engagement and school grades through dual pathways by assessing motivation with a newly developed implicit measure and an explicit measure. participants were 139 students in year 3 of secondary education (58% female, m = 14.8 years). motivation was assessed with an implicit association test (iat) as well as an explicit measure (self-report). behavioural engagement was assessed by teacher ratings, and school grades were reported by students. the explicit and implicit measures of student motivation were not significantly correlated, suggesting that both measures tap into different aspects of student motivation. furthermore, structural equation analyses revealed that students’ explicit and implicit motivation were positively associated with school grades. neither motivation measure was associated with teacher ratings of behavioural engagement. this study contributes to existing research by showing that an implicit measure of student motivation can predict unique variation in school grades in addition to an explicit measure. as such, the current study provides initial support for a dual pathway model of student motivation. keywords: motivation; engagement; implicit measure; dual pathway model mailto:t.e.hornstra@uu.nl hornstra et al 2 | f l r 1. introduction there is widespread consensus among researchers and educators that motivation to learn is a powerful factor contributing to engagement and achievement. self-determination theory (sdt) (deci & ryan, 1985; ryan & deci, 2000a) states that intrinsic motivation facilitates higher levels of engagement and achievement, whereas extrinsic types of motivation can lead to maladaptive learning behaviours and outcomes. previous research has empirically supported these claims and found positive reciprocal associations between intrinsic motivation and engagement and achievement (see for example green et al., 2012; guay, ratelle, liem, & litalien, 2010; korpershoek, 2016; michou,vansteenkiste, mouratidis, & lens, 2014; taylor et al., 2014; walker, greene, & mansell, 2006). even though motivation has been found to facilitate achievement and vice versa, most studies only found weak or modest associations between motivation-related constructs on the hand, and behavioural engagement and achievement on the other hand (for reviews, see for example cerasoli, nicklin, & ford, 2014; richardson, abraham, & bond, 2012) instead of the powerful relationships that are often assumed. these modest findings may be due to the specific focus of previous studies on deliberate motivational processes. yet, abundant research in social psychology shows that most human behaviour is not only predicted by conscious, deliberate processes. that is, in situations where people do not have an opportunity to deliberate, their beliefs operate in a more reactive or automatic way. when the automatic pathway is followed, beliefs are automatically activated and guide behaviour without conscious awareness or deliberation. hence, it is believed that human behaviour can be predicted by two pathways, a deliberate and an automatic pathway (e.g., chaiken & trope, 1999; gawronski & bodenhausen, 2006; 2011). this may also apply to students’ motivation for school. in everyday classroom situations, students may not always deliberately choose their actions or plan their strategies. students’ motivation will probably often impact their behaviours in more automatic or reactive ways. by assessing motivation with an implicit as well as an explicit measure, this study aims to examine whether student motivation predicts behavioural engagement and school grades through dual pathways. 1.1 student motivation sdt (deci & ryan, 1985; ryan & deci, 2000b) provides an integrative framework of student motivation. students who are intrinsically motivated engage in an activity because the activity in itself evokes interest or pleasure or because they identify with reasons for performing an activity. extrinsic motivation comes from external pushes, reinforcement, or internal pressures that cause feelings of obligation or guilt. extrinsic motivation comes in various forms that vary in their degree of relative autonomy. extrinsic motivation is fully external when students feel controlled by others or by contextual pressures to engage in an activity they would not otherwise want to engage in (i.e., external regulation). alternatively, students can also pressure themselves to engage in an activity out of guilt, shame, or concerns about what others might think of them (i.e., introjected regulation). this type of regulation involves a higher level of autonomy compared to external regulation. in case of identified regulation, autonomy is even higher. for example when students are motivated for an activity because they consider it useful for their future careers or they recognize the importance of the skills they might develop through that activity. external and introjected regulation are associated with undesirable behaviours such as unwillingness or passive compliance. these two types of regulation have also been referred to as controlled motivation, whereas identified regulation and intrinsic motivation are referred to as autonomous motivation and are associated with more beneficial outcomes, such as greater behavioural engagement (e.g., ryan & deci, 2000a, 2000b; vansteenkiste, lens, & deci, 2006, vansteenkiste et al., 2009). in this study, behavioural engagement is defined as students’ effort, attention, and persistence with regard to their schoolwork (e.g., skinner, furrer, marchand, & kindermann., 2008). previous research supports the assumption that autonomous motivation is positively associated with behavioural engagement and achievement, although these associations are in general not very strong. recent meta-analyses (cerasoli et al., 2014; richardson et al., 2012) found small to medium correlations of r = .17 hornstra et al 3 | f l r and r = .21 between intrinsic motivation and academic achievement in school settings. a positive relationship between autonomous motivation and behavioural engagement has also been found in previous research. autonomous or intrinsic motivation has for example been associated with higher-quality learning (grolnick & ryan, 1987; vansteenkiste, simons, et al., 2004), the use of effective learning strategies (michou et al., 2014), and class participation (green et al., 2012). however, most studies on the relationship between motivation and behavioural engagement have used self-reports to assess both constructs, which may lead to common-method bias (e.g., podsakoff, mackenzie, lee, & podsakoff 2003). that is, correlations based on similar methods may overestimate the actual strength of the association. previous research has indeed found much stronger associations of self-reported motivation with self-reported behavioural engagement compared to studies that included different measures, such as students’ self-reported motivation and teacher ratings of behavioural engagement (e.g., skinner & belmont, 1993; skinner, chi, et al., 2012). this shows that previous studies may have overestimated the associations between self-reported motivation and behavioural engagement. hence, it can be concluded from prior research that self-reports of motivation can only explain students’ behavioural engagement or achievement to a limited extent. 1.2 a dual pathway model the associative-propositional evaluation (ape) model (gawronski, 2006; 2007; 2011) is a dual pathway model that describes how beliefs guide human behaviours. according to the ape-model, beliefs can be activated upon encountering a relevant stimulus, leading to an automatic reaction. alternatively, in other instances, a deliberate propositional process may follow the activation of the belief in which a person consciously reflects on the validity of a belief before acting upon it. especially within the domain of prejudice, the ape-model has been studied and supported by empirical evidence (gawronski, peters, brochu, & strack, 2008, or for an overview see gawronski & bodenhausen, 2011). this dual pathway may also apply to student motivation. that is, students’ motivational beliefs may also guide student behaviour either automatically or in more deliberate ways. students may hold different types of motivational beliefs, ranging from controlled to more autonomous. according to the dual pathway model, these beliefs can be activated automatically or they can be deliberated upon. in school, students encounter numerous moments on a daily basis during which motivational beliefs (e.g., ‘i do not enjoy this type of task’) may be activated and they can either engage in their schoolwork or not. oftentimes they will not have the opportunity nor the willingness to deliberate and think about what behaviours they will express. consequently, automatic processes will often guide students’ behavioural engagement or performance. take for example a student who mostly endorses controlled motivational beliefs (“i only do my schoolwork, because i have to.”). many situations in school may automatically activate this belief, and as a result, he/she often may not fully engage in adaptive learning behaviours or perform to the best of his/her abilities. however, when this student has the opportunity to deliberate, he/she may also endorse more autonomous reasons for engaging in schoolwork. hence, when there is opportunity to deliberate, a second pathway may be followed and this student may realize the importance of doing her schoolwork and put in more effort after all. 1.3 measurement of student motivation in educational science, explicit self-reports are typically used to measure student motivation (e.g., zimmerman, 2006). self-reported motivation or ‘explicit motivation’ only captures those aspect of students’ beliefs about their motivation that they are willing to report, that they can reflect upon, and that they are able to describe accurately. hence, the value of such introspectively derived explicit measurements may be limited due to a variety of factors, including social desirability (lepper, corpus, & iyengar, 2005), or limited awareness, opportunity, or ability to translate mental beliefs into a self-report (nosek, hawkins, & frazier, 2011; rudman, 2011). to circumvent such problems, reaction-time measures, commonly referred to as ‘implicit measures’, have become increasingly popular in other domains of social psychology over the last decades (fazio & olson, 2003). implicit measures aim to capture beliefs that automatically guide behaviour. hornstra et al 4 | f l r many implicit measures are reaction time measures which are administered on a computer and assess associations stored in memory (greenwald mcghee, & schwartz, 1998). by unobtrusively assessing the strength of these associations (for example the extent to which a person associates a certain attitude object with ‘positive’ or ‘negative’), more or less automatically activated beliefs are assessed (nosek et al., 2011). thereby these measures aim to capture associations in memory that, when activated, can automatically cause affective or behavioural responses (greenwald et al.,1998; nosek et al., 2011; rudman, 2011; sherman, gawronski & trope, 2014). implicit measures, have indeed been shown to be much less susceptible to social desirability and self-presentation bias than explicit measures (e.g., steffens, 2004). in a variety of domains, implicit measures are found to add to the prediction of variations in human behaviour that are not accounted for by self-report measures (for an overview, see for example greenwald, poehlman, uhlmann, & banaji, 2009). implicit measures may also be a suitable instrument to assess motivation. because these measures assess associations stored in memory, they could also assess, for example, the extent to which students associate themselves with enjoyment of schoolwork. as such, they may provide an alternative to self-reports for assessing students’ motivational beliefs. one of the most widely used implicit measures is the implicit association test (iat) of greenwald et al. (1998). the iat assesses the association between various concepts by asking participants to repeatedly pair two concepts. the more strongly the participant associates two concepts, the faster the participant will respond when this particular pair of concepts is presented. the iat is often used to assess people’s attitudes, for example participants’ positive versus negative attitudes toward ethnic minority versus majority people (e.g., mcconnell & leibold, 2001; van den bergh, denessen, hornstra, voeten, & holland, 2010) or political preferences (e.g., galdi, arcuri, & gawronski, 2008). in addition to these attitude-iat’s, identityiat’s have been developed to assess how participants associate themselves with a certain target concept. gray et al. (2011) for example found that participants who associated themselves more strongly with alcohol, engaged in more drinking behaviours. the predictive validity of the iat has been demonstrated for various domains (greenwald, et al., 2009), including ethnic prejudice (connell & leibold, 2001), voting behaviour (galdi, et al., 2008), substance use (rooke, hine, & thorsteinsson, 2008), and consumer behaviours (friese, wänke, & plessner, 2006). in addition, there have been a few studies that examined motivation in an implicit manner. for example, several studies demonstrated that activating motivational beliefs or goals can affect subsequent behaviour or performance. that is, bargh, gollwitzer, lee-chai, barndollar, and trötschel (2001) have shown that when a goal to perform well (versus a goal to cooperate with others) was activated, respondents’ performance on an intellectual task increased. furthermore, in a study by levesque and pelletier (2003), respondents were primed with words representing intrinsic motivation, extrinsic motivation, or neither. respondents who were primed with intrinsic motivation enjoyed a puzzle task more and performed better than respondents in the control condition. respondents primed with extrinsic motivation enjoyed the task less and performed less well. in addition, burton, lydon, d’alessandro, & koestner (2006) found that priming students with intrinsic motivational words, led to greater well-being. moreover, an (implicit) lexical decision test predicted subsequent course performance six weeks later. finally, in a set of two laboratory experiments, keatley, clarke, and hagger (2013) used an identity-iat to assess undergraduate students’ motivation and examined whether their implicit autonomous motivation predicted the duration respondents worked on an unsolvable task. students’ implicit motivation predicted task persistence beyond the prediction by explicit measures. although the studies described here only included undergraduate students – thereby limiting the scope of these findings – these studies show the potential role of implicit motivation in predicting behavioural or performance outcomes. 1.4 relations between implicitly and explicitly measured beliefs previous research in other domains found weak correlations between explicit and implicit measures (e.g., nosek et al., 2011; rudman, 2011). these weak correlations could indicate low concurrent validity of both measures, but could also suggest that both type of measures tap into different aspects of one’s beliefs. hornstra et al 5 | f l r that is, implicit measures aim to assess beliefs that are activated without deliberation. these may be automatically activated beliefs which the respondent may not even consider to be valid (e.g., an association between ‘black man’ and ‘criminal). yet, this association can still affect one’s behaviour, for example stepping back when one encounters a black person (gawronski & bodenhausen, 2011). explicit measures on the other hand assess beliefs that one has deliberated upon. after deliberation, one may express a belief (e.g., ‘negative evaluations of black people are wrong’) that may be inconsistent with the automatically activated belief. this explicitly assessed belief may also be predictive of one’s behaviour (talking in a friendly manner to a black person). hence, implicitly and explicitly assessed beliefs are not necessarily consistent with one another. moreover, implicitly and explicitly assessed beliefs are found to be predictive of different types of behaviour. namely, self-report measures are typically more predictive of planned and strategic behaviours (i.e. deliberate pathway), whereas reaction-time measures are more predictive of non-verbal and immediate behaviours (i.e., automatic pathway) (sherman et al., 2014). 1.5 this study the present study aims to examine whether a dual pathway model can also be applied to motivation of students in secondary education to predict students’ behavioural engagement and school grades. at this age, motivation of many students starts to develop unfavourable (e.g. opdenakker, maulana, & den brok, 2012). previous studies on implicit motivation (bargh et al., 2001; burton et al., 2006; keatley et al., 2013; levesque & pelletier, 2003) have exclusively focused on undergraduate students. the results of these studies may not be generalizable to high school students, for whom – contrary to undergraduate students who chose their course of study – school is compulsory. as such, it is important to examine whether the findings of these previous studies can be extended to other educational contexts. moreover, to our knowledge, prior studies have only assessed how implicit affects participants’ behaviour or performance during tasks performed in a laboratory setting. the present study includes more ecologically valid outcome measures, and is among the first to examine whether an implicit measure of student motivation is associated with students’ actual behaviours in class and their school grades. this study adds to research on student motivation by assessing whether an innovative and new instrument to assess student motivation can be used as an alternative to explicit motivation measures –which have been shown to have limited predictive validity. moreover, if this implicit measure of motivation is indeed associated with students’ behavioural engagement and school grades, targeting maladaptive motivational associations may be a fruitful approach for motivational interventions, for example by priming associations that are considered adaptive for learning. in the present study, students’ motivational beliefs will be assessed implicitly, by means of a reaction-time measure, and explicitly, by self-report. for the sake of readability, we use the terms ‘implicit motivation’ and ‘explicit motivation’, although students’ motivational beliefs are not necessarily implicit or explicit in nature. to be precise, the term ‘implicit’ and ‘explicit’ refer to the way the beliefs were measured and the pathways through which they are assumed to affect behaviour. the present study examined to what extent students’ motivation predicts behavioural engagement and school grades, thereby aligning with the ape model that assumes that beliefs guide subsequent behaviours (gawronski, 2006) and with prior studies which found causal effects of motivation on subsequent achievement (e.g., green et al., 2012; guay et al., 2010). however, we do not assume that these relationships are unidirectional. previous research has shown reciprocal relationships between motivation and achievement (taylor et al., 2014). as such, we assume that associations between motivation and engagement/school grades are indicative of reciprocal relationships between these constructs. the following hypotheses were addressed in this study: hypothesis 1: implicit and explicit motivation are positively, but weakly correlated. as previous research mostly found weak correlations between explicit and implicit measures (e.g., nosek et al., 2011; rudman, 2011), a positive, but weak association is expected between both measures of motivation. hornstra et al 6 | f l r hypothesis 2. implicit motivation uniquely predicts teacher ratings of students’ behavioural engagement and school grades in addition to explicit motivation. previous research (e.g. sherman et al., 2014) indicated that explicit beliefs tend to be more predictive of planned and strategic behaviours and implicit beliefs more predictive of non-verbal and immediate behaviours. as behavioural engagement and school grades both comprise and result from a complex variety of planned and immediate behaviours, it is expected that implicit and explicit motivation are both predictive of behavioural engagement and school grades. therefore we expect that explicit and implicit motivation both explain unique variations in behavioural engagement and school grades. 2. methods 2.1 respondents a sample of 139 students (59 male, 80 female) from six classes from two different schools participated. they attended year 3 (grade 9) of general secondary education. this track is attended by approximately 23% of secondary school students in the netherlands. it can be positioned between prevocational education and pre-university education (attended by approximately 56% and 20% of the secondary school population, respectively) (ministry of education, culture, and science, 2014). the mean age of the students was 14.8 years (sd = 0.67). the majority of students had a dutch or western background (95.7%). 2.2 instruments 2.2.1 explicit motivation the self-regulation questionnaire academic (srq-a) (ryan & connell, 1989) was administered to assess students’ explicit motivation for school. this measure is rooted in sdt. it assesses the extent to which students’ school-related behaviours are autonomously regulated. it consists of four subscales with 32 items in total that are answered on a four-point scale ranging from not true at all (1) to very true (4). the items were preceded by a question, for example ‘why do i work on my schoolwork?’. the four subscales are intrinsic regulation (e.g., ‘because i enjoy doing my schoolwork.’), identified regulation (‘e.g., ‘because it’s important to me to work on my classwork.’) , introjected regulation (‘because i’ll be ashamed of myself if it didn’t get done.’), and external regulation (‘because i want the teacher to think i’m a good student.’). a confirmatory factor analyses revealed that a model with two factors, that is autonomous motivation (consisting of the items of the subscales intrinsic regulation and identified regulation) and controlled motivation (consisting of the items of the subscales introjected regulation and external regulation) fitted the data reasonably well (cfi= .89; rmsea= .070; srmr= .088) and outperformed alternative models. internal consistencies, as indicated by cronbach’s alpha, were good: autonomous motivation, α= .84 and controlled motivation, α= .86. 2.2.2 implicit motivation an implicit association test (iat) (greenwald et al., 1998) was administered to assess the extent to which students associate autonomous versus controlled reasons for making schoolwork with their perception of themselves. the iat measures the strength of associations by comparing reaction times to different pairings of concepts. specifically, the iat works as follows: the strength of automatic associations between a target category (e.g., autonomous or controlled reasons for schoolwork) and an identity category (e.g., ‘me’ or ‘not me’) is inferred from the relative speed with which one sorts stimulus words into these categories hornstra et al 7 | f l r correctly. to represent the target categories, we used the category labels ‘making schoolwork because i want to’ and ‘making schoolwork because i have to’ to represent autonomous and controlled motivation respectively. the corresponding words of both categories were terms that could be associated with autonomous motivation (“wanting”, “fun”, “voluntary”, “interesting”, “important”) and words that could be associated with controlled motivation (“obligation”, “boring”, “control”, “required”, “pressure”). identityrelated words were used (“i”, “myself”, “self” or “they”, “them”, “their”) that either belonged to the category ‘me’ or to ‘not me’. it was expected that students with higher autonomous implicit motivation would associate themselves more strongly with ‘making schoolwork because i want to’ and find it easier to classify the stimulus words into the correct categories – hence, respond more quickly – when ‘making schoolwork because i want to’ and ‘i’ were presented on the same side. prior to this study, a small pilot (n= 17 students) was conducted in which a large set of words were presented that respondents could classify as belonging to the two categories. only words that were exclusively listed to belong to one of these categories and not to both were used in this iat. the iat consisted of seven blocks (see fig. 1). in the first practice block, respondents were shown a series of words that appeared in the middle of the screen that either represented autonomous or controlled motivation. participants had to correctly classify these words in the categories ‘making schoolwork because i want to’ on the left side of the screen by pressing the ‘e’ key on the laptop or in the category ‘making schoolwork because i have to’ on the right side of the screen by pressing the ‘i’ key. in the second block, identity-related words were shown that needed to be classified as ‘me’ or ‘not me’ with the same keys. the third and fourth block were congruent test blocks and the aforementioned categories were combined. during these blocks, both motivation-related words and identity-related words were presented on the screen and needed to be classified in the correct categories with the ‘e’ and ‘i’ keys. these test blocks were followed by a practice block and two incongruent test blocks in which the motivation categories were switched to the other sides of the screen. reaction times were measured for each response. the response latencies for the first two congruent test blocks were compared to the response latencies for the latter two incongruent test blocks. it was expected that a stronger association between two concepts paired together (e.g., ‘autonomous motivation’ and ‘me’) would result in shorter reaction times when compared to other pairs. given the relatively small sample size, the order of blocks was not counterbalanced. the meta-analysis by greenwald et al (2009) indicates that using a fixed order of blocks does not affect predictive validity. in general, iat’s are found to have good test-retest reliability, good convergent validity with other implicit measures (cunningham, preacher, & banaji, 2001), and good predictive validity (e.g. greenwald et al., 2009). the internal consistency of this iat, calculated by the method described by bosson, swann, and pennebaker (2000), was α= .74 which indicates good reliability of the measure. the scoring procedures recommended by greenwald, nosek and banaji (2003) were used to calculate the standardized d-score. trials greater than 10,000 milliseconds indicate that respondents may have been distracted and were deleted. subjects who responded extremely fast (<300 milliseconds) on more than 10% of the trials were not included in the analyses (i.e., those who were simply hitting keys as fast as possible). in addition, the iat scores of students with less than 60% correct responses were not included in subsequent analyses . consequently, the scores of 19 students were excluded. positive iat scores indicated a higher level of autonomous motivation for schoolwork and negative iat scores indicated a higher level of controlled motivation for schoolwork. hornstra et al 8 | f l r figure 1. screen shots of the iat during (a) practice block 1, (b) practice block 2, (c) test block 3-4, (d) practice block 5, and (e) test block 6-7. (b) (a) (c) making schoolwork because i have to making schoolwork because i want to interesting making schoolwork because i want to me making schoolwork because i have to not me myself me not me them making schoolwork because i want to making schoolwork because i have to making schoolwork because i want to me making schoolwork because i have to not me fun boring (e) (d) hornstra et al 9 | f l r 2.2.3 behavioural engagement. teachers (n= 6) rated the behavioural engagement, i.e. their effort, attention, and persistence, of each student. we used this method instead of students’ self-reports, because explicit motivation was also measured by self-reports. assessing both constructs by similar measures would likely result in an overestimation of the correlation between explicit motivation and behavioural engagement (i.e. ‘common method bias’, podsakoff et al., 2003). we used the behavioural engagement scale of skinner, kindermann, and furrer (2008) who based their measure on wellborn (1991). a back-translation procedure was used to translate the items to dutch. the scale consisted of five items per student to be answered on a four-point likert scale ranging from totally not applicable to this student (1) to totally applicable to this student (4). an example item is “when this student doesn’t do well, he/she works harder”. the reliability of this scale was α= .92. 2.2.4 school grades students reported their average course grade for three core subjects in the curriculum, dutch, english, and mathematics. self-reported grades are considered to reflect actual grades with reasonable accuracy, especially in academic domains (kuncel, credé, & thomas, 2005). the average grade across these three core subjects was calculated for each student. the grades can range from 1 to 10, with 10 representing the highest grade. 2.3 procedure passive parental consent was obtained prior to data collection. no parents objected to participation. during data collection, students were visited by a researcher in a computer room. they received a brief explanation by the researcher and a brief instruction on paper, after which they could turn on the computer. all instruments were administered online. they were first presented with the iat, followed by the explicit motivation questionnaire, questions on demographics and their school grades. simultaneously, their teachers filled out the ratings of behavioural engagement for each student. 2.4 data-analyses for explicit motivation and school grades, missing data were limited (2.2%-4.3% missing data). for implicit motivation and behavioural engagement, more data were missing (13.7% and -18.7% respectively). missingness was due to students who were excluded because of high error rates on the iat and because one teacher did not fill out the ratings of behavioural engagement. because missingness could not be considered completely at random (mcar), missing data were handled by using the full information maximum likelihood procedure (fiml) (schafer & graham, 2002). to test the first hypothesis on the association between the implicit and explicit motivation measure, the correlation between both measures was calculated. furthermore, to test the second hypothesis, which states that both implicit and explicit motivation predicted behavioural engagement and school grades, structural equation analyses were performed in mplus 7.4 (muthén & muthén, 1998-2017). the dependent variables (behavioural engagement and school grades) were estimated as latent factors based on the observed scores. in the analysis, we controlled for several factors associated with the outcome variables, i.e., gender, grade repetition, and minority background. these were entered as dummy variables in the analysis. explicit (autonomous and controlled) and implicit motivation were added as observed predictors to the model. given that we corrected for error in the implicit motivation, when calculating the iat scores, measurement error was also taken into account for the explicit measure, by correcting for attenuation. the significance of the coefficients for the different predictor variables was tested using wald tests (z tests). the set level of significance was 5%. model fit was evaluated with chi-square difference tests, the rmsea, the cfi, and by hornstra et al 10 | f l r the standardized root mean square residual (srmr). a significant chi-square difference indicates whether or not model fit significantly improved or worsened. an rmsea below 0.05 indicates good fit of a model and scores between .05 and .08 indicate reasonable fit. scores above .10 indicate poor fit. a cfi above 0.90 indicates good fit of a model. a srmr value below .08 is generally considered a good fit (hu & bentler, 1999). 3. results 3.1 descriptive statistics and correlations in table 1, descriptive statistics are reported. the implicit measure has a positive mean, suggesting that on average students implicitly endorse autonomous motivation over controlled motivation. table 1 descriptive statistics n m sd min max explicit motivation – autonomous 136 2.28 0.39 1.14 3.43 explicit motivation – controlled 136 2.24 0.39 1.06 3.11 implicit motivation 120 0.26 0.48 -1.56 1.36 behavioural engagement 113 2.72 0.76 1.00 4.00 school grades 133 6.42 0.80 4.30 10.00 total 139 with regard to the first hypothesis, weak correlations were expected between the implicit and explicit measures of motivation. the correlations reported in table 2 show that implicit motivation was not significantly correlated with either autonomous or controlled explicit motivation (r = .12; p = .191; r = .13; p = .172, respectively). in addition, table 2 shows that behavioural engagement and school grades were significantly positively correlated (r = .20; p = .042). none of the motivation measures were correlated with behavioural engagement (explicit – autonomous: r = 0.17; p = .069; explicit – controlled: r = .03; p = .723; implicit: r = .28; p = .788). with respect to school grades, it was found that both explicit measures were not significantly correlated with school grades (explicit – autonomous: r = .10; p = .246; explicit – controlled: r = -.02; p = .831). implicit motivation was, however, positively associated with school grades (r = .19; p = .037). table 2 correlations 1. 2. 3. 4. 5. 1. explicit motivation autonomous 1.00 2. explicit motivation controlled .63** 1.00 3. implicit motivation .12 .13 1.00 4. behavioural engagement .17 .03 .03 1.00 5. school grades .10 -.02 .19* .20* 1.00 * p<0.05; ** p<0.01; *** p<0.001 hornstra et al 11 | f l r 3.2 prediction of behavioural engagement and school grades the second hypothesis stated that both implicit and explicit motivation would predict behavioural engagement and school grades. table 3 presents a summary of the structural equation model with the different measures of motivation (explicit motivation – autonomous, explicit motivation – controlled, and implicit motivation) as predictors of both behavioural engagement and school grades. gender, grade repetition, and minority background were entered as control variables. to test whether the predictive value differed of implicit and explicit motivation differed, we constrained the associations between both types of motivation and the outcome measures to be equal for implicit and explicit motivation. neither for behavioural engagement, nor for school grades, this worsened model fit (δχ2 (1) = .611, p = .805 and δχ2 (1) = .031, p = .860, respectively). therefore, the final model indicated that the relations between motivation on the one hand and school grades and behavioural engagement on the other hand were equal for implicit motivation and explicit (autonomous) motivation. the final model with equality constrains fitted the data well: χ2(68) = 102.927, p = .004; cfi = .93; rmsea = .065; srmr = .093. table 3 reports the final model. note that even though the unstandardized coefficients were constrained to be equal, the standardized coefficients slightly differ for implicit and explicit motivation. the results of the final model indicated that, after controlling for gender, grade repetition, and minority background, neither explicit, nor implicit motivation were significantly associated with behavioural engagement (explicit motivation – autonomous: β = .20, p = .227; explicit motivation – controlled: β = -.29, p = .307; implicit motivation: β = -.20, p = .227). contrarily, school grades were positively predicted by explicit autonomous motivation as well as implicit motivation (both β = .26, p = .025). explicit controlled motivation did not significantly predict students’ school grades (β= -.22, p = .305). hence, the hypothesized relation between both types of motivation and behavioural engagement could not be confirmed, but hypothesis 2 was confirmed for the association between implicit and explicit motivation and school grades. that is, both implicit and autonomous motivation were found to be positive predictors of school grades. the corresponding effect sizes (es = .17 and es = .24), which were based on the standardized coefficients, indicate small to medium effect sizes. when implicit motivation was not included as a predictor in the final model, the background characteristics and both aspects of explicit motivation explained 4.0% of the variance in school grades. this was raised to 9.7% after adding implicit motivation to the model, indicating that implicit motivation explained an additional 5.7% of the variance in school grades. table 3 summary of structural equation model for variables predicting behavioural engagement and school grades ( n= 139) behavioural engagement school grades unstandardized (se) standardized (se) unstandardized (se) standardized (se) gender (girl) .09 (.17) .05 (.10) .02 (.14) .01 (.12) grade repetition -.93*** (.24) -.42*** (.09) -.19 (.16) -.13 (.11) minority background -.17 (.58) -.03 (.11) .41 (.38) .12 (.12) explicit motivation autonomous .20 (.17) .09 (.07) .26* (.12) .17* (.08) explicit motivation controlled -.29 (.29) -.13 (.12) -.22 (.21) -.15 (.14) implicit motivation .20 (.17) .11 (.09) .26* (.21) .23* (.11) r2 .19 .10 * p<0.05; ** p<0.01; *** p<0.001 hornstra et al 12 | f l r 4. discussion the aim of this study was to examine whether motivation is associated with students’ behavioural engagement and school grades through a dual pathway model. previous research focused almost exclusively on the deliberate pathway by measuring motivation with self-reports that require respondents to be consciously aware of their motivational beliefs and to be able to accurately reflect on those beliefs. by developing and employing an implicit measure of student motivation, we were able to assess students’ implicit motivational beliefs. in line with our expectations we found that the explicit and implicit measure of student motivation both added to the prediction of students’ school grades. this suggests that automatic, non-deliberate motivational processes and deliberate motivational processes play a role in predicting secondary students’ school grades. neither motivation measure was associated with teacher ratings of behavioural engagement. by showing that an implicit measure of student motivation adds to the prediction of students’ school grades, this study contributes to existing research. to our knowledge, this study was the first study examining how implicit motivation is related to actual student outcomes in the classroom, and the first to examine this with secondary school students. as such, the present study has shown that in actual school settings, motivation is associated with students’ school grades through different pathways. in addition, this study also provides further support for basic assumptions of sdt (e.g., deci & ryan, 2000a), as we found that (implicit) autonomous motivation is more beneficial for school grades as compared to controlled motivation. this shows that these basic assumptions of sdt can be extended to implicit processes as well. together, the results are a first step toward supporting the existence of a dual pathway model of student motivation. the outcomes show that automatically activated motivational beliefs may impact students’ school grades through an implicit pathway, although effect sizes can be considered to be modest. these results are in line with previous studies in many other domains of human functioning in which it was shown that peoples’ implicit beliefs can explain unique variations in behaviour, beyond what is explained by selfreports (greenwald et al. 2009). the results also indicated that the implicit and explicit measure of motivation were not significantly correlated even though we aimed for a high degree of conceptual correspondence. this aligns with prior research on relations between implicit and explicit measures (e.g., hoffmann, gawronski, gschwendner, le, schmitt, 2005; nosek et al., 2011). low correspondence between both types of measures can indicate independence of the constructs that are measured by both types of measures or can be caused by several other factors, including, method-related characteristics, bias in explicit self-reports, or limited awareness, opportunity, or ability to translate mental beliefs into a self-report (hoffmann et al., 2005). further research is needed to examine whether both measures assess independent belief systems regarding students’ motivation for school which could be differentially predictive of different types of student outcomes or, alternatively, whether explicit measures are more strongly affected by the aforementioned methodological limitations. neither explicit nor implicit motivation was significantly associated with behavioural engagement, contrary to previous studies (green et al., 2012; guay et al., 2010; korpershoek, 2016; michou et al., 2014; taylor et al., 2014; walker, et al., 2006). previous studies typically used self-reports to assess both explicit motivation and behavioural engagement and found more substantial correlations between explicit motivation and behavioural engagement. studies that did not use similar measures, but included teacher ratings of behavioural engagement instead, as we did in our study, found weaker, but nevertheless significant relations between explicit motivation and behavioural engagement (e.g., skinner & belmont, 1993; skinner, chi, et al., 2012). the absence of a significant relation between motivation and behavioural engagement in our study may be accounted for by a lack of power. that is, there may be weak relationships which could only be detected with a larger sample size. in addition, it is important to note that the effect sizes for the association between both types of motivation and school grades were only small to medium. one factor that could account for the modest effect sizes may be the level of specificity at which students’ motivation was assessed in this study. that is, the focus of the present study was on students’ general motivational dispositions regarding their school work. hornstra et al 13 | f l r as such, this study was able to demonstrate that students’ implicit motivational dispositions are associated with their school grades. yet in future studies, it may also be of interest to study students’ implicit and explicit motivational processes at a more specific level, focusing on domain-specific, task-specific, or situation-specific motivation. given that implicit beliefs can be activated and effect behaviour within a specific situation (e.g. bargh et al, 2001), a more specific approach may potentially yield more substantial effect sizes. in addition, given the modest effect sizes, it is important to take into consideration that there was still a substantial degree of variance in school grades that could not be accounted for by either their implicit or explicit motivation, and is caused by other factors beyond the scope of the present study. 4.1 limitations some limitations of this study need to be acknowledged. first of all, because of the cross-sectional nature of this study, we cannot draw any conclusions on causal directions. based on previous research (e.g., taylor et al., 2014), we assume that students’ motivation, both implicit and explicit, are reciprocally associated with school grades. that is to say, higher school grades will likely also increase students’ implicit and explicit motivation for their schoolwork. second, our sample was relatively small given that small effects were to be expected. as such, our study may not have sufficient power to reveal weak relationships. this study can be considered to be a first step in examining whether a dual process model applies student motivation. however, we recommend follow-up studies with larger samples, as well as longitudinal designs in order to find further support for the proposed model. third, iat measures have some limitations and/or disadvantages. although iat’s can measure implicit associations, these associations do not necessarily have to be implicit because the participant can be aware of their associations (fazio & olson, 2003). in addition to that, some participants might be able to discover what associations are being measured during the test. even though they might be aware of this, chances of manipulation of the test are small (de houwer, 2002). another possible limitation of the iat is that it only measures relative preference for the two concepts (brunel, greenwald & washington, 2004). even though someone can have a stronger association with autonomous motivation compared to controlled motivation, this does not say anything about the absolute strength of their motivation. fourth, our sample was restricted to students in general secondary education. as such, the variation in student motivation and both outcome measures may have been larger if students from other tracks would have been involved, which could also have increased the strength of the relationships between motivation and the outcomes variables. finally, the implicit and explicit motivation measures were administered in a computer room during regular classroom hours. if feasible, individual administration would be preferable to ensure that students can fill out the iat and questionnaires quietly without any disruptions. 4.2 conclusions and future research this study is among the first to show that implicit motivation can predict students’ school grades. to better understand how implicit motivation affects school grades, more in-depth research is needed on the psychological processes and behaviours that are evoked by implicit motivation. in addition, research on explicit motivation suggested that a wide range of individual, background, and contextual characteristics affects explicit student motivation (e.g., goodenow, 1993, hornstra, van der veen, &peetsma, 2016; hornstra, van der veen, peetsma, & volman, 2015a; 2015b; shernoff & schmidt, 2007; vansteenkiste et al., 2012; stroet et al., 2013). more research is also needed on individual and contextual antecedents of students’ implicit motivation, to gain a better understanding of how educators can facilitate optimal student functioning. to summarize, the current study provided initial support for a dual pathway model of student motivation and showed that an implicit measure of student motivation can predict unique variation in school grades in addition to an explicit measure. consequently, future research on student motivation would benefit from incorporating both explicit and implicit measures of student motivation. hornstra et al 14 | f l r keypoints we hypothesized a dual-pathway model of student motivation and examined if motivation guides behaviour through a deliberate and an automatic pathway. motivation of 139 high school students was assessed with an explicit measure (self-reports) and a newly developed implicit measure (iat). implicit motivation and explicit motivation both predicted students’ school grades. neither explicit nor implicit motivation predicted teacher reports of students’ behavioural engagement. motivation can affect student achievement through an implicit automatic pathway and a deliberate pathway. references bargh, j. a., gollwitzer, p. m., lee-chai, a., barndollar, k., & trötschel, r. (2001). the automated will: nonconscious activation and pursuit of behavioral goals. journal of personality and social psychology, 81, 1014–1027. doi:10.1037/0022-3514.81.6.1014 bosson, j. k., swann, w. b., & pennebaker, j. w. (2000). stalking the perfect measure of implicit selfesteem: the blind men and the elephant revisited? journal of personality and social psychology, 79, 631–643. doi:10.1037/0022-3514.79.4.631 brunel, f. f., tietje, b. c., & greenwald, a. g. (2004) is the implicit association test a valid and valuable measue of implicit consumer social cognition? journal of consumer psychology, 4, 385-404. doi:10.1207/s15327663jcp1404_8 burton, k. d., lydon, j. e., d’alessandro, d. u., & koestner, r. (2006). the differential effects of intrinsic and identified motivation on well-being and performance: prospective, experimental, and implicit approaches to self-determination theory. journal of personality and social psychology, 91, 750–762. doi:10.1037/0022-3514.91.4.750 cerasoli, c. p., nicklin, j. m., & ford, m. t. (2014). intrinsic motivation and extrinsic incentives jointly predict performance: a 40-year meta-analysis. psychological bulletin, 140, 980–1008. doi:10.1037/a0035661 chaiken, s., & trope, y. (eds.). (1999). dual-process theories in social psychology. new york: guilford press. cunningham, w. a., preacher, k. j., & banaji, m. r. (2001). implicit attitude measures: consistency, stability, and convergent validity. psychological science, 12, 163–170. doi:10.1111/14679280.00328 deci, e. l., & ryan, r. m. (1985). intrinsic motivation and self-determination in human behaviour. new york: plenum de houwer, j. (2002) the implicit association test as a tool for studying dysfunctional associations in psychopathology: strengths and limitations. journal of behavior therapy and experimental psyciatry, 33, 115-133. doi:10.1016/s0005-7916(02)00024-1 fazio, r. h., & olson, m. a. (2003). implicit measures in social cognition research: their meaning and use. annual review of psychology, 54, 297–327. doi:10.1146/annurev.psych.54.101601.145225 friese, m., wänke, m., & plessner, h. (2006). implicit consumer preferences and their influence on product choice. psychology and marketing, 23, 727–740. doi:10.1002/mar.20126 froiland, j. m., & worrell, f. c. (2016). intrinsic motivation, learning goals, engagement, and achievement hornstra et al 15 | f l r in a diverse high school. psychology in the schools, 53, 321–336. doi:10.1002/pits.21901 galdi, s., arcuri, l., & gawronski, b. (2008). automatic mental associations predict future choices of undecided decision-makers. science, 321, 1100–1102. doi:10.1126/science.1160769 gawronski, b., & bodenhausen, g. v. (2006). associative and propositional processes in evaluation: an integrative review of implicit and explicit attitude change. psychological bulletin, 132, 692–731. doi:10.1037/0033-2909.132.5.692 gawronski, b., & bodenhausen, g. v. (2007). unraveling the processes underlying evaluation: attitudes from the perspective of the ape model. social cognition, 25, 687–717. doi:10.1521/soco.2007.25.5.687 gawronski, b., & bodenhausen, g. v. (2011). the associative–propositional evaluation model. advances in experimental social psychology, 44, 59–127. doi:10.1016/b978-0-12-385522-0.00002-0 gawronski, b., peters, k. r., brochu, p. m., & strack, f. (2008). understanding the relations between different forms of racialp: a cognitive consistency perspective. personality and social psychology bulletin, 34, 648–665. doi:10.1177/0146167207313729 goodenow, c. (1993). classroom belonging among early adolescent students: relationships to motivation and achievement. the journal of early adolescence, 13, 21–43. doi:10.1177/0272431693013001002 green, j., liem, g. a. d., martin, a. j., colmar, s., marsh, h. w., & mcinerney, d. (2012). academic motivation, self-concept, engagement, and performance in high school: key processes from a longitudinal perspective. journal of adolescence, 35, 1111–1122. doi:10.1016/j.adolescence.2012.02.016 greenwald, a. g., mcghee, d. e., & schwartz, j. l. k. (1998). measuring individual differences in implicit cognition: the implicit association test. journal of personality and social psychology, 74, 1464– 1480. doi:10.1037/0022-3514.74.6.1464 greenwald, a. g., nosek, b. a., & banaji, m. r. (2003). understanding and using the implicit association test: i. an improved scoring algorithm. journal of personality and social psychology, 85, 197–216. doi:10.1037/0022-3514.85.2.197 greenwald, a. g., poehlman, t. a., uhlmann, e. l., & banaji, m. r. (2009). understanding and using the implicit association test: iii. meta-analysis of predictive validity. journal of personality and social psychology, 97, 17–41. doi:10.1037/a0015575 grolnick, w. s., & ryan, r. m. (1987). autonomy in children's learning: an experimental and individual difference investigation. journal of personality and social psychology, 52, 890-898. doi:10.1037/0022-3514.52.5.890 guay, f., chanal, j., ratelle, c. f., marsh, h. w., larose, s., & boivin, m. (2010). intrinsic, identified, and controlled types of motivation for school subjects in young elementary school children. british journal of educational psychology, 80, 711–735. doi:10.1348/000709910x499084 guay, f., ratelle, c. f., roy, a., & litalien, d. (2010). academic self-concept, autonomous academic motivation, and academic achievement: mediating and additive effects. learning and individual differences, 20, 644–653. doi:10.1016/j.lindif.2010.08.001 guthrie, j. t., klauda, s. l., & ho, a. n. (2013). modeling the relationships among reading instruction, motivation, engagement, and achievement for adolescents. reading research quarterly, 48, 9–26. doi:10.1002/rrq.035 hofmann, w., gawronski, b., gschwendner, t., le, h., & schmitt, m. (2005). a meta-analysis on the correlation between the implicit association test and explicit self-report measures. personality and social psychology bulletin, 31, 1369–1385. doi:10.1177/0146167205275613 hornstra, l., van der veen, i., & peetsma, t. (2016). domain-specificity of motivation: a longitudinal study hornstra et al 16 | f l r in upper primary school. learning and individual differences, 51, 167–178. doi:10.1016/j.lindif.2016.08.012 hornstra, l., van der veen, i., peetsma, t., & volman, m. (2015a). innovative learning and developments in motivation and achievement in upper primary school. educational psychology, 35(5), 598–633. doi:10.1080/01443410.2014.922164 hornstra, l., van der veen, i., peetsma, t., & volman, m. (2015b). does classroom composition make a difference: effects on developments in motivation, sense of classroom belonging, and achievement in upper primary school. school effectiveness and school improvement, 26), 125–152. doi:10.1080/09243453.2014.887024 keatley, d., clarke, d. d., & hagger, m. s. (2012). investigating the predictive validity of implicit and explicit measures of motivation in problem-solving behavioural tasks. british journal of social psychology, 52, 510–524. doi:10.1111/j.2044-8309.2012.02107.x korpershoek, h. (2016). relationships among motivation, commitment, cognitive capacities, and achievement in secondary education. frontline learning research, 4, 28-43. doi :10.14786/flr.v4i3.182 lepper, m. r., corpus, j. h., & iyengar, s. s. (2005). intrinsic and extrinsic motivational orientations in the classroom: age differences and academic correlates. journal of educational psychology, 97, 184–196. doi:10.1037/0022-0663.97.2.184 levesque, c., & pelletier, l. g. (2003). on the investigation of primed and chronic autonomous and heteronomous motivational orientations. personality and social psychology bulletin, 29, 1570–1584. doi:10.1177/0146167203256877 mcconnell, a. r., & leibold, j. m. (2001). relations among the implicit association test, discriminatory behaviour, and explicit measures of racial attitudes. journal of experimental social psychology, 37, 435–442. doi:10.1006/jesp.2000.1470 michou, a., vansteenkiste, m., mouratidis, a., & lens, w. (2014). enriching the hierarchical model of achievement motivation: autonomous and controlling reasons underlying achievement goals. british journal of educational psychology, 84, 650–666. doi:10.1111/bjep.12055 ministry of education culture, and science (2014). kerncijfers 2009-2013: onderwijs, cultuur, en wetenschap [core figures 2009-2013: education, culture and science]. the hague: ministry of education, culture and science muthén, l. k., muthén, b. (1998-2017). mplus user’s guide (7th ed.). los angeles, ca: author. kuncel, n. r., crede, m., & thomas, l. l. (2005). the validity of self-reported grade point averages, class ranks, and test scores: a meta-analysis and review of the literature. review of educational research, 75, 63–82. doi:10.3102/00346543075001063 nosek, b. a., hawkins, c. b., & frazier, r. s. (2011). implicit social cognition: from measures to mechanisms. trends in cognitive sciences, 15, 152–159. doi:10.1016/j.tics.2011.01.005 opdenakker, m.-c., maulana, r., & den brok, p. (2012). teacher–student interpersonal relationships and academic motivation within one school year: developmental changes and linkage. school effectiveness and school improvement, 23, 95–119. doi:10.1080/09243453.2011.619198 podsakoff, p. m., mackenzie, s. b., lee, j.-y., & podsakoff, n. p. (2003). common method biases in behavioural research: a critical review of the literature and recommended remedies. journal of applied psychology, 88, 879–903. doi:10.1037/0021-9010.88.5.879 reeve, j., & tseng, c.-m. (2011). agency as a fourth aspect of students’ engagement during learning activities. contemporary educational psychology, 36, 257–267. doi:10.1016/j.cedpsych.2011.05.002 richardson, m., abraham, c., & bond, r. (2012). psychological correlates of university students' academic hornstra et al 17 | f l r performance: a systematic review and meta-analysis. psychological bulletin, 138, 353-387. doi: 10.1037/a0026838 rooke, s. e., hine, d. w., & thorsteinsson, e. b. (2008). implicit cognition and substance use: a metaanalysis. addictive behaviours, 33, 1314–1328. doi:10.1016/j.addbeh.2008.06.009 rudman, l. a. (2011). implicit measures for social and personality psychology. london: sage. doi:10.4135/9781473914797 ryan, r. m., & connell, j. p. (1989). perceived locus of causality and internalization: examining reasons for acting in two domains. journal of personality and social psychology, 57, 749–761. doi:10.1037/00223514.57.5.749 ryan, r. m., & deci, e. l. (2000a). intrinsic and extrinsic motivations: classic definitions and new directions. contemporary educational psychology, 25, 54–67. doi:10.1006/ceps.1999.1020 ryan, r. m., & deci, e. l. (2000b). self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. american psychologist, 55, 68–78. doi:10.1037/0003066x.55.1.68 schafer, j. l., & graham, j. w. (2002). missing data: our view of the state of the art. psychological methods, 7, 147–177. doi:10.1037/1082-989x.7.2.147 sherman, j. w., gawronski, b., & trope, y. (eds.). (2014). dual-process theories of the social mind. new york: guilford publications. shernoff, d. j., & schmidt, j. a. (2007). further evidence of an engagement–achievement paradox among u.s. high school students. journal of youth and adolescence, 37, 564–580. doi:10.1007/s10964-0079241-z skinner, e. a., & belmont, m. j. (1993). motivation in the classroom: reciprocal effects of teacher behaviour and student engagement across the school year. journal of educational psychology, 85, 571–581. doi:10.1037/0022-0663.85.4.571 skinner, e. a., chi, u., & the learning-gardens educational as. (2012). intrinsic motivation and engagement as “active ingredients” in garden-based education: examining models and measures derived from self-determination theory. the journal of environmental education, 43, 16–36. doi:10.1080/00958964.2011.596856 skinner, e. a., kindermann, t. a., & furrer, c. j. (2008). a motivational perspective on engagement and disaffection: conceptualization and assessment of children’s behavioural and emotional participation in academic activities in the classroom. educational and psychological measurement, 69, 493-525. doi: 10.1177/0013164408323233 steffens, m. c. (2004). is the implicit association test immune to faking? experimental psychology, 51, 165–179. doi:10.1027/1618-3169.51.3.165 stroet, k., opdenakker, m.-c., & minnaert, a. (2013). effects of need supportive teaching on early adolescents’ motivation and engagement: a review of the literature. educational research review, 9, 65–87. doi:10.1016/j.edurev.2012.11.003 taylor, g., jungert, t., mageau, g. a., schattke, k., dedic, h., rosenfield, s., & koestner, r. (2014). a self-determination theory approach to predicting school achievement over time: the unique role of intrinsic motivation. contemporary educational psychology, 39, 342–358. doi:10.1016/j.cedpsych.2014.08.002 van den bergh, l., denessen, e., hornstra, l., voeten, m., & holland, r. w. (2010). the implicit prejudiced attitudes of teachers: relations to teacher expectations and the ethnic achievement gap. american educational research journal, 47, 497–527. doi:10.3102/0002831209353594 vansteenkiste, m., lens, w., & deci, e. l. (2006). intrinsic versus extrinsic goal contents in self hornstra et al 18 | f l r determination theory: another look at the quality of academic motivation. educational psychologist, 41, 19–31. doi:10.1207/s15326985ep4101_4 vansteenkiste, m., sierens, e., soenens, b., luyckx, k., & lens, w. (2009). motivational profiles from a self-determination perspective: the quality of motivation matters. journal of educational psychology, 101, 671–688. doi:10.1037/a0015083 vansteenkiste, m., sierens, e., goossens, l., soenens, b., dochy, f., mouratidis, a., aelterman, n., haerens, l., & beyers, w. (2012). identifying configurations of perceived teacher autonomy support and structure: associations with self-regulated learning, motivation and problem behaviour. learning and instruction, 22(6), 431–439. doi:10.1016/j.learninstruc.2012.04.002 vansteenkiste, m., simons, j., lens, w., sheldon, k. m., & deci, e. l. (2004). motivating learning, performance, and persistence: the synergistic effects of intrinsic goal contents and autonomysupportive contexts. journal of personality and social psychology, 87, 246–260. doi:10.1037/00223514.87.2.246 walker, c. o., greene, b. a., & mansell, r. a. (2006). identification with academics, intrinsic/extrinsic motivation, and self-efficacy as predictors of cognitive engagement. learning and individual differences, 16, 1–12. doi:10.1016/j.lindif.2005.06.004 wellborn, j.g. (1991). engaged and disaffected action: the conceptualization and measurement of motivation in the academic domain. unpublished doctoral dissertation, university of rochester, new york. zimmerman, b. j. (2008). investigating self-regulation and motivation: historical background, methodological developments, and future prospects. american educational research journal, 45, 166–183. doi:10.3102/0002831207312909 lisette hornstraa, antoinette kamsteega, sara pota, & lydia verheija autrecht university, the netherlands abstract keywords: motivation; engagement; implicit measure; dual pathway model 1. introduction 1.1 student motivation 1.2 a dual pathway model 1.3 measurement of student motivation 1.4 relations between implicitly and explicitly measured beliefs 1.5 this study 2. methods 2.1 respondents 2.2 instruments 2.2.1 explicit motivation 2.2.2 implicit motivation 2.2.3 behavioural engagement. 2.2.4 school grades 2.3 procedure 2.4 data-analyses 3. results 3.1 descriptive statistics and correlations table 1 table 2 3.2 prediction of behavioural engagement and school grades table 3 4. discussion 4.1 limitations 4.2 conclusions and future research keypoints references publication larmuseau et al frontline learning research vol.7 no. 2 (2019) 57 74 issn 2295-3159 combining physiological data and subjective measurements to investigate cognitive load during complex learning charlotte larmuseauaa, jan cornelisb, piet desmet a & fien depaepe a a itec, imec research group at ku leuven, etienne sabbelaan 51, kortrijk, belgium b imec, kapeldreef 75, leuven, belgium article received 27 august 2018 / revised 2 december / accepted 2 may / available online 10 may abstract cognitive load theory is one of the most influential theoretical explanations of cognitive processing during learning. despite its success, attempts to assess cognitive load during learning have proven difficult. therefore, in the current study, students’ self-reported cognitive load after the problemsolving process has been combined with measures of physiological data, namely, electrodermal activity (eda) and skin temperature (st) during the problem-solving process. data was collected from 15 students during a high and low complex task about learning and teaching geometry. this study first investigated the differences between subjective and physiological data during the problemsolving process of a high and low complex task. additionally, correlations between subjective and physiological data were examined. finally, learning behavior that is retrieved from log-data, was related with eda. results reveal that the manipulation of task complexity was not reflected by physiological data. nevertheless, when investigating individual differences, eda seems to be related to mental effort. keywords: cognitive load; physiological data; electrodermal activity; skin temperature; complex learning info corresponding author: rcharlotte.larmuseau@kuleuven.be doi: 10.14786/flr.v7i2.403 1. introduction as society and work environments become more complex it is increasingly relevant that learning environments mirror this complexity of the real world (jonassen, 2000; kirschner, ayres & chandler, 2011; merrill, 2009; van merriënboer, kirschner & kester, 2003). nevertheless, a risk of complex learning environments is that the cognitive load imposed by the complex learning tasks is often excessive (larmuseau, elen & depaepe, 2018; van merriënboer & sluijsmans, 2009). this phenomenon can be explained by cognitive load theory (clt) introduced by sweller (1994). clt uses current knowledge about the human cognitive architecture as a baseline to develop the instructional design for complex learning environments (martin, 2014). clt distinguishes three types of cognitive load, intrinsic, extraneous and germane load (brunken, plass & leutner, 2003; paas, tuovinen, tabbers & van gerven, 2010; sweller, 2010). the level of intrinsic load is assumed to be determined by the complexity of the task or learning material and cannot be directly altered by the instructional designer. extraneous load is mainly imposed by instructional procedures that are suboptimal, whereas germane load refers to the learners’ working memory resources available to deal with the complexity of the task or learning material (sweller, 2010). both extraneous and germane load can by facilitated by the instructional designer. an instructional designer should find a balance between keeping the matter sufficiently challenging but still within the cognitive capacities of the learner. exceeding learners’ cognitive capacities can induce cognitive overload which could hamper learning. specifically, this means that when the content is very complex due to high element interactivity (i.e., the amount of interrelations between knowledge, procedures, formulas etc.) which affects intrinsic load, instructional designers should keep extraneous load to a minimum (e.g., by providing clear instructions, provide embedded support) and subsequently foster germane load (kirschner, kester & corbalan, 2011; sweller, 2010). in order to align the instructional design with students’ cognitive abilities, we should be able to measure cognitive load during complex learning. former studies investigated cognitive load by using subjective measurements such as self-reported questionnaires (boekaerts, 2017; zheng & cook, 2012). those self-reported questionnaires have some important disadvantages (e.g., subjective measures, assumption of constant workload capacity, see section 2.2 ; deleeuw & mayer, 2008; raaijmakers, baars, schaap, paas & van gog, 2017; spanjers, van gog & van merriënboer, 2012). as a result, more researchers show interest in using objective, real-time measures. physiological measures provide objective data and can be unobtrusively collected while dealing with a task or learning material. moreover, physiological data might provide an indication of changes in cognitive functioning throughout the process of solving a task (boekaerts, 2017). former studies already indicated that electrodermal activity (eda) and skin temperature (st) can be linked to different levels of task complexity (haapalainen, kim, forlizzi & dey, 2010; nourbakhs, wang, chen & calvo, 2012; shi, ruiz, taib, choi & chen, 2007). nevertheless, it is unclear whether these physiological measures are related to self-reported intrinsic load, extraneous load, germane load and the overall mental effort during complex problem solving (leppink, paas, van der vleuten, van gog & van merriënboer, 2013). therefore, in the current study, a high and low complex task was developed relating to the learning and teaching of geometry. the complexity of the task was manipulated by increasing the element interactivity for the high complex task (sweller, 2010). in both tasks the same amount of support was provided. data was retrieved using self-reported questionnaires to measure students’ experienced intrinsic load, extraneous load, germane load and mental effort. this distinction between the different types and mental effort was made because the different types of cognitive load concerns mental load induced by task complexity and instructional design, whereas mental effort invested covers the overall amount of cognitive processing for a particular task (paas et al., 2003). the subjective measures were combined with physiological data through wrist-worn wearables containing both eda and st. the purpose of this study was threefold. first, we investigated differences in the experienced cognitive load and the physiological data while solving a high and low complex task. secondly, we examined whether individual differences of subjective measurements are related to individual differences of physiological data for the high and low complex task. finally, we described whether peaks (i.e., eda) and/or drops (i.e., st) of physiological data are related to specific events (e.g., consultation of support) that took place during the problem solving process. 2. theoretical framework 2.1 cognitive load theory clt is concerned with the instructional implication of the interaction between the complexity and instructional design of the learning material and human cognitive architecture (sweller, 2010). basically, the human cognitive architecture consists of an effectively unlimited long-term memory, which interacts with a working memory that has limited processing capacity (kirschner et al., 2011; sweller, 1994). long-term memory contains cognitive schemata that are used to store and organize knowledge. learning occurs when information is successfully processed in working memory and when new schemas are created or incorporated into consisting schemas in the long-term memory. as the processing capacity of the working memory is so limited, overcoming individual working memory limitations by instructional manipulations has been the main focus of clt (sweller, van merriënboer & paas, 1998). cognitive load can be defined as a multidimensional construct representing the load that performing a particular task, imposes on the learners’ cognitive system (paas et al., 2010). clt claims that the cognitive load that learners experience can be intrinsic, extraneous or germane (sweller, 2010). the level of intrinsic load for a particular task is assumed to be determined by the inherent difficulty of a certain topic and the level of element interactivity of the learning material in relation a student’s prior knowledge. the more elements that interact, the more intrinsic processing is required for coordinating and integrating the material and the higher the working memory load (de leeuw & mayer, 2008; sweller, 2010). working memory load is not only imposed by the intrinsic complexity of the material that needs to be learned, it can also be imposed by the instructional design. for instance, unclear instructional procedures can impose extraneous load. extraneous processing means that the learner engages in cognitive processing that does not support the learning objective (de leeuw & mayer, 2008; glogger-frey, gaus & renkl, 2017; van merriënboer & sluijsmans, 2008; sweller, 2010). instructional design techniques that reduce extraneous load (e.g., fading support) should ensure that students devote less attention to irrelevant aspects of the task. subsequently, more cognitive capacity can be allocated to the actual learning objective (ciernak, scheiter & gerjets, 2009; mayer & moreno, 2010; sweller, ayres & kalyugo, 2011). meanwhile, intrinsic and extraneous load depend on the characteristics of the learning tasks or the instructional design, germane load is more concerned with the cognitive characteristics of the learner. more specifically, it refers to the working memory resources that are available to engage in knowledge elaboration processes and argumentation (sweller, 2010). accordingly, in order to optimize learning, learning tasks should be aligned with the learner’s cognitive capabilities (schmeck, opfermann, van gog, paas & leutner, 2015; sweller, 2010). measuring cognitive load during complex learning should provide more insight into how to align instructional design with students’ cognitive capabilities. 2.2 subjective measurements of cognitive load self-reports for measuring cognitive load are subjective measurements consisting of unidimensional and multidimensional scales. unidimensional subjective rating scales have been used intensively in research and have been identified as reliable and valid estimators of cognitive load (boekaerts, 2017; chang & yang, 2010; leppink et al., 2013; paas, 2003). the paas’s nine-point mental effort rating scale has been most frequently used in cognitive load research (chen et al., 2016; paas, 1992). paas’s nine-point mental effort rating scale requires learners to rate their mental effort immediately after completing a task (paas, 1992). mental effort refers to the cognitive capacity that is allocated to accommodate the demands imposed by a task (paas et al., 2003). according to paas, learners can introspect the amount of mental effort invested during a learning task. subsequently, paas claims that the learner’s assessment can be used as an index of overall cognitive load (chen et al., 2016). nevertheless, this unidimensional scale gives little insight into the influence of the complexity of the task and the influence of the instructional design on cognitive load (boekaerts, 2017; de bruin & van merriënboer, 2017; klepsch, schmitz & seufert, 2017; leppink et al., 2013). accordingly, leppink et al. (2013) and klepsch et al. (2017), developed a subjective cognitive load scale in which they used multiple items for each type of cognitive load in order to get more specific information about intrinsic load, extraneous load and germane load. despite the frequent use of self-reported scales to assess cognitive load, some critiques have been raised. firstly, subjective measurements are based on the assumption that students are able to introspect on their cognitive processes and accordingly are able to self-report on their experienced cognitive load (boekaerts, 2017; schmeck et al., 2015). secondly, as subjective scales are often administered after the learning task, subjective scales do not capture variations in load over time. taking into account these limitations, it might be more interesting to combine subjective measurements with real-time objective cognitive load information (boekaerts, 2017; chen et al., 2016; zheng & cook, 2012). 2.3 physiological measures of cognitive load the physiological approach for cognitive load measurement is based on the assumption that any change in the human cognitive functioning is reflected in the human physiology. subsequently, in contrast to subjective measurements, physiological measures are continuous and measured at a high frequency (e.g., every second) and with a high precision (chen et al., 2016). given the close relationship between cognitive load and neural systems, human neurophysiological signals are seen as promising avenues to measure cognitive load (boekaerts, 2017; chen et al., 2016). former research has investigated the relationship between learners’ cognitive load and their physiological behaviour. the physiological measures that have been used to investigate cognitive load are among others heart rate by electrocardiography (ecg), brain activity by electroencephalography (eeg), eye activity (e.g., blink rate, pupillary dilation), eda, heat flux and st (antonenko, paas, grabner & van gog, 2010; haapalainen et al. 2010; scharinger, soutschek, schubert & gerjets, 2015; smets et al., 2018; zagermann, pfeil & reiterer, 2016). although a lot of physiological data, such as brain and eye activity, has been proven to be highly effective for measuring cognitive load, these types of physiological data often requires expensive sophisticated equipment that is highly obtrusive in measuring cognitive activities, especially in ecological valid contexts (chen et al., 2016; scharinger et al., 2015). possible solutions to collect physiological data in an unobtrusive way is by means of wrist-worn wearables. these wearables can easily capture different physiological data such as eda and st and are less expensive compared to more sophisticated measures of physiological data (chen et al., 2016). eda involves measuring the electrical conductance of the skin through sensors attached to the wrist. skin conductivity varies with changes in skin moisture level (i.e., sweating) and can reveal changes in the sympathetic nervous system (sns). the slowly changing part of the eda signal is called the skin conductance level (scl) and is a measure of psychophysiological activation. scl can vary substantially between and within individuals. a fast change in the eda signal (i.e., a peak) occurs in reaction to a single stimulus and is called galvanic skin response (gsr; braithwaite, watson, jones & rowe, 2013). research has linked gsr variation to stress and sns arousal. as a person becomes more or less stressed, the gsr increases or decreases respectively (hoogerheide, renkl, logan, paas & van gog, 2019; liapis, katsanos, sotiropoulos, xenos & karousos, 2015, smets et al., 2018). additionally, research has also linked gsr readings to cognitive activity, claiming gsr responses increase when more cognitive load is experienced (ikehara & crosby, 2005; nourbakhs et al, 2012; setz et al., 2010; shi et al., 2007, yousoof & sapiyan, 2013). the study of nourbakhs, wang, chen and calvo (2015) captured gsr data of 13 and 16 participants from different reading and arithmetic tasks. the arithmetic tasks contained four difficulty levels, whereas the reading task contained three difficulty levels. results of anova indicated that both mean gsr and accumulated gsr yielded significantly different results throughout different task difficulty levels. shi et al. (2007) investigated 11 subjects when dealing with four tasks divided in four distinct levels of cognitive load. results revealed insignificant differences across the interactive models for mean gsr, but significant differences when using accumulated gsr. yousoof and sapiyan (2013) investigated whether cognitive load could be detected by mean eda. in this experiment 7 subjects had to solve three different programming tasks that were different in terms of complexity. yousoof and sapiyan found no conclusive results for mean gsr, indicating that the variation among the subjects was very different during one task. in addition to eda, st can also reflect changes in sns. research claims that acute stress triggers peripheral vasoconstriction, causing a rapid, short-term drop in skin temperature. moreover, stress can also cause a more delayed skin warming, providing two opportunities to quantify stress (herborn et al., 2015; karthikeyan, murugappan & yaacob, 2012; shusterman, anderson & barnea, 1997; smets et al., 2018; vinkers, et al., 2013). little research has used st to assess cognitive load. nevertheless, the study of haapalainen et al. (2010) investigated the cognitive load of 20 subjects through gsr and heat flux data (i.e., rate of heat transfer). the subjects had to solve six elementary cognitive tasks that differed in difficulty. afterwards, haapalainen et al. (2010) evaluated the performance of each of the features in assessing cognitive load using personalised machine learning techniques (i.e., naïve bayes classifier). results indicated that they did not obtain satisfactory results for gsr. by contrast, they did find that across all participants heat flux was shown to be an indicator of differences in cognitive load. the findings of former studies indicate that eda and st can indicate differences in cognitive load, but none of these studies related physiological data with self-reported cognitive load. 2.4 research aims to conclude, physiological measures have some important advantages when compared to subjective measurements. these measures are more objective (i.e., not dependent on students’ perceptions), multidimensional (i.e., different physiological measures are sensitive to different cognitive processes), unobtrusive (i.e. no additional requirements), implicit (i.e., collect data while students are working on their tasks) and continuous (i.e. provide information of cognitive processes during learning). nevertheless, it can be difficult to interpret physiological data. therefore, it would be interesting to investigate whether there is a relationship between subjective measurements of cognitive load and physiological data. the following research questions are formulated: • rq1: does the manipulation of the level of complexity of a task, based on element interactivity, result in differences in perceived cognitive load and mental effort when controlled for prior knowledge? • rq2: does the manipulation of the level of complexity of a task, based on element interactivity, result in differences in physiological data, when controlled for prior knowledge? • rq3: is there a relationship between individual differences in self-reported data and individual differences of physiological data for a high and low complex task? rq4: is there a relationship between the physiological data of one learner and his/her interactive behaviour during the problem solving process? 3. methodology 3.1 participants and study design participants were 15 future primary school teachers of which ten were female and five male (age between 18-24). all participants were first year bachelor students (i.e., second semester). the study was highly ecologically valid as the study was orchestrated by the students’ lecturer of the teaching mathematics course unit. moreover, the intervention was integrated into the students’ study program (i.e., primary school teacher training). the intervention consisted of a within-subject design and was conducted online in the moodle learning management system (lms). the intervention took place in the auditorium of their faculty where students could solve the tasks individually on their own computer among their fellow students. this session was supervised by their lecturer and a researcher. students first received an online questionnaire of which the timeframe (+/five min.) to complete the first questionnaire was used as an adaption period in order to stabilize the wearable signals (i.e., baseline measurement). next, all students had to solve a high complex and a low complex task on preparing a lesson in geometry as shown in figure 1. in order to control for order effects, (a) half of the subjects were exposed to the high complex task during the first session and the low complex task during the second session, whereas for (b) the other half, the sequence was vice versa. more specifically, eight students started with the high complex task and seven students started with the low complex task. 3.2 high and low complex tasks the high and low complex tasks were developed in moodle lms. the scope of both tasks was designing a lesson preparation on the circumference of a circle for primary school children. this subject matter was not yet covered in previous lessons. both tasks contained six elements where both aspects of pedagogical content knowledge; pck (i.e., inductive teaching strategy, choose teaching materials to support your lesson, aligning the topic of the lesson with the flemish curriculum and integration of differentiation in your lesson in the classroom) and content knowledge; ck (i.e., formula of the circumference of the circle) were addressed. the complexity of the high complex task was manipulated based on element interactivity (sweller, 2010). in the high complex task students had to coordinate and integrate six elements consisting of ck and pck in order to write a course preparation about the circumference of the circle, whereas the low complex task consisted of six questions where each element was addressed separately (see figure 1). during both problems, the same support consisting of procedural and supportive information was provided. an example of procedural information can also be found in figure 1 in the second box. procedural information is provided just-in-time and concise. supportive information is much more comprehensive and is comparable to the background theory. both procedural and supportive information can be consulted by clicking on the words in italics. figure 1. high complex task, question of the low complex task and an example of the procedural information 3.3 students’ prior knowledge information about students’ prior knowledge was gathered in the first semester during their examination. students were tested on their knowledge of pck (mean = 63.5%, sd = 19.7) and ck (mean = 72.2%, sd = 27.8). content was (teaching) mathematics in general and geometry in particular. examples of test-items can be found in figure 2. all tests were corrected by the instructor of the course unit. we have no insight into the prior knowledge of one student who participated in the study, which means that we can include an indicator of prior knowledge for 14 students in the analysis. figure 2: example questions of the prior knowledge test 3.4 subjective measurements for the measurement of cognitive load a validated instrument developed by leppink et al. (2013) was used for the measurement of intrinsic, extraneous and germane load. the questionnaire was translated into the specific context of the present study as shown in table 1. the questionnaire consisted of a 7-point likert scale (i.e., ranging from “totally disagree” to “totally agree”). reliability was determined through cronbach’s α in order to investigate the overall consistency of the constructs (schreiber, nora, stage, barlow & king, 2006). confirmatory factor analysis (cfa) was not conducted due to the small sample size, but former research has validated the questionnaire and has proven that the questionnaire is reliable (leppink et al., 2013). additionally, the paas’s nine-point mental effort rating scale was added to the questionnaire (paas, 1992). table 1 survey items and reliability of the constructs 3.5 physiological data to measure physiological data including eda and st, 15 students were monitored with wrist-worn wearables as shown in figure 2. these wearables were able to sense gsr with a high dynamic range (.05-20µs) at the lower side of the wrist and the output was accurate within a frame of approximately 1 second. st was acquired at the upper side of the wrist at a frequency of 32 hz and the output was accurate within a frame of approximately 1 second at 0.1 °c. before analysing the physiological data, a number of procedures were carried out. firstly, a confidence indicator (ci), with values ranging from 0 to 1, monitors whether the sensor is correctly attached to the body. values of ci lower than .80 were ignored as this indicates low quality of the data due to incorrect sensor attachment (+/.01% per individual). secondly, visual analysis of the signal was conducted for both eda and st. artefacts were removed 20s before and after the artefact and an interpolation over the gap was performed. thirdly, large differences in skin conductance among individuals can occur (yousoof & sapiyan, 2013). therefore, to counteract the variation between subjects, the eda and st data of each individual participant were standardized, bringing the mean of each signal to 0 and its variance to 1. fourthly, time domain features were analysed and mean eda and st were calculated as shown in figure 3. figure 3. standardized mean eda and s 3.6 log-data log-data was retrieved from the moodle learning management system (lms). the lms-system automatically keeps tracks of user activity (i.e., every min) and session. log-data was divided into several events, namely: (1) start the task; reading instructions, (2) writing an answer, (3) consultation of support and (4) submission; reviewing the answer. 3.7 analysis this study first investigated the differences between a high and low complex task for both the subjective measurements and physiological data (i.e., rq1, rq2). therefore, both subjective measurements and physiological data were tested on the normality assumption. results of the shapiro-wilk tests reveal that both subjective measurements and physiological measurements were normally distributed. as we were interested in the mean differences between the high and low complex task of both the self-reported and physiological data, controlled for prior knowledge (i.e., both pck and ck), order effect (see section 3.1), we conducted a linear mixed model (lmm) incorporating pck, ck and order as fixed factors and measurement time as a repeated measure (two-level for rq1 and three-level for rq3). when conducting lmm, the restricted maximum likelihood method (reml) was applied (baayen, davidson & bates, 2008). based on findings of rq1 and rq2, this study investigated the individual differences in the self-reported data of cognitive load for a high and low complex task, and how this relates to individual differences in physiological data (rq3). cohen’s d was calculated when differences were significant in order to have insight into the effect sizes (lecroy & krysik, 2007). a bivariate correlation analysis was conducted in order to find relationships between physiological data and subjective measurements of cognitive load. fourthly, as the advantage of physiological data is that it is measured continuously, this study investigated whether there are relationships between specific events (i.e., consultation of support) based on log-data and peaks (i.e., spontaneous fluctuations per s) of eda and drops of st (i.e., rq4). given the small sample size, the analysis more descriptive. 4. results 4.1 research question 1 descriptive statistics of the subjective measurements as shown in table 2 reveal that students reported on average higher intrinsic load, extraneous load and mental effort during the high complex task in comparison with the low complex task. results furthermore indicate that students reported higher germane load during the low complex task which was expected. table 2 descriptive statistics of the subjective measurements of the high and low complex task in order to investigate differences in the perceived cognitive load and mental effort (i.e., rq1), lmm was conducted incorporating pck, ck and ‘order effect’ as fixed factors and time as a two-level repeated measurement. pairwise comparison of the different measurements of intrinsic load, extraneous load, germane load and mental effort are indicated in table 3. results reveal that intrinsic load differed significantly across phases. f(1,13) = 6.43, p = .03. pairwise comparison reveals that intrinsic load was significantly higher (m = .86, p = .03) during the high complex task with cohen’s d = .88. when investigating the fixed factors, there was no significant effect of both pck, f(1,10) = .05, p = .82 and ck, f(1,10) = .43, p = .53. moreover, no significant order effect was found f(1,10) = 12, p = 74. as expected, results reveal no significant difference for extraneous load across phases f(1,13) = 17, p = .69. pairwise comparison reveals no significant mean difference (m = -.05, p = .90) between the high and low complex task for extraneous load. results of the fixed effects reveal no significant effect of pck f(1,10) = .04, p = .84, ck f(1,10) = .17, p = .69, and order f(1,10) = 1.58, p = .24. results for germane load indicate no significant differences across phases f(1,13) = 1.21, p = .29. pairwise comparison reveals no significant mean difference for germane load (m = -.18, p = .29) between the high and low complex task. results of the fixed effects indicate no significant effects for pck, f(1,10) = .00, p = .96 and ck, f(1,11) = .01, p = .93. moreover, no order effect was found, f(1,10) = 1.39, p = .2. finally, results revealed that mental effort was different across phases. mean difference of mental effort between the high and low complex task was significant (m = 1.43, p = 00) in the predicted direction with cohen’s d = 1.52. no significant effects of pck, f(1,11) = 2.39, p = .15 and ck, f(1,11) = 2.84, p = .12. additionally, no order effect, f(1,10) = .27, p = 62 was found. table 3 pairwise comparison of subjective measurements controlled for prior knowledge (i.e., pck, ck) and order effect 4.2 research question 2 descriptive statistics of the physiological data can be found in table 4. mean eda is lower during the high complex task compared to the low complex task. mean st is lower during the high complex task. table 4 descriptive statistics of the standardized physiological data in order to investigate the differences of physiological data between the baseline measurement, high and low complex task (i.e., rq2), lmm was conducted incorporating pck, ck, order effect as fixed factors and time as a three-level repeated measurement. results indicate that differences were found for mean eda across the different phases f(2,26) = 6.56, p = .01. pairwise comparison of the different measurements of mean eda are indicated in table 5. results of pairwise comparison reveals that the mean difference between the baseline measurement and high complex task phase is significant in the predicted direction (m = -.60, p = .05) with cohen’s d = .19. moreover, the mean difference is significant between the baseline measurement and the low complex task (m = -1.05, p = .00) with cohen’s d = .14. results reveal that no significant mean difference was found between the high and low complex task (m = -.45, p = .14). moreover, the mean difference was in the unexpected direction. when investigating the fixed factors, there was a non-significant main effect of both pck f(1,10) = .18, p = .68 and ck f(1,10) = .81, p = .36. additionally, there was a significant effect of order f(1,10) = 7.62, p = .02, which indicates an order effect. no significant differences were found for mean st across the different measurements, f(2,26) =.16, p = .85. pairwise comparison reveals no significant mean differences between baseline measurement and the high complex task (m = 1.02, p = .61), baseline measurement and the low complex task (m = .87, p = .66), and between the high and low complex task (m = -.15, p = .94). nonetheless, all mean differences were in the expected direction. when investigating the fixed effects, there was a non-significant main effect of both pck f(1,10) = .00, p = .97 and ck f(1,10) = .12, p = .74. additionally, there was no significant order effect, f(1,10) = .45, p = 52. table 5 pairwise comparison of physiological data controlled for prior knowledge and order 4.3 research question 3 results of rq1 reveal significant differences for perceived intrinsic load and mental effort. rq3 investigates the relationship between the individual differences of intrinsic load, mental effort and physiological data. results are displayed in table 6 and reveal that mental effort is significantly positive correlated with mean eda (r = .58, p = .03) for the high complex task. nevertheless, no significant positive correlation was found between mean eda and mental effort for the low complex task. no significant results were found for st. table 6 correlations between standardized physiological data and subjective measurements for the high complex task and low complex task. 4.4 research question 4 in the final rq4, this study investigates the relationship between physiological data and specific events retrieved from log-data and eda peaks. an example of such relationships is shown in figure 4. table 7 gives an overview of the amount of relationships between specific events and eda peaks. in contrast to eda, no conclusive relationships were found between st (i.e., drops) and specific events. st for most participants increased throughout the intervention as illustrated in figure 5. figure 4. electrodermal activity related to log-data of participant 15 figure 5: skin temperature related to log-data of participant 15 table 7 the relationship between specific events and eda peaks 5. discussion 5.1 research question 1 this study attempted to firstly investigate the difference of subjective measurements of cognitive load between a high and low complex task (i.e., rq1). results reveal that the students indicate higher perceived intrinsic load for the high complex task when compared with the low complex task. this indicates that the manipulation of complexity based on element interactivity was successful. additionally, students indicated that the perceived mental effort was higher during the high complex task. effect sizes of both intrinsic load and mental effort were high (>.80) indicating that the manipulation of complexity had an impact (lecroy & krysik, 2007). this reveals that students invested more mental effort into solving the high complex task in order to maintain performance at a constant level (paas et al., 2003). this is also in line with clt, since the high complex task was high in element interactivity and possibly required a lot of cognitive processing (van merriënboer & sweller, 2005). no significant differences were found for extraneous load between both tasks. this finding was expected as the instructions for both tasks were of the same level of difficulty. additionally, no significant differences were found for germane load, indicating that both tasks enhanced students’ understanding of the content at a similar level. this was in line with our expectations as the content and available support of both tasks was the same. 5.2 research question 2 secondly, this study aimed at investigating whether we can use physiological data to distinguish between the two complexity levels of the task. when investigating mean eda, results reveal that significant differences were found between both tasks and the baseline measurement. these findings indicate that both tasks result in a higher mean eda. nevertheless, effect sizes were very small (< .20), indicating that task complexity only had a minimal impact on mean eda (lecroy & krysik, 2007). moreover, no significant differences were found for mean eda between the high and low complex. these results are in line with the findings of the study of haapalainen et al. (2010), which also revealed no significant differences for eda between six tasks of different levels of difficulty. moreover, against expectations, descriptive statistics reveal that mean eda was higher during the low complex task, when compared with the high complex task. these unexpected findings may be induced by the order effect. this order effect may reduce a clear difference between the eda during the high and low complex task. moreover, visual analysis reveals that for the majority of all participants, skin conductance rises throughout the intervention (i.e., drift). since, more participants had the low complex at the end, this might indicate that results are biased by drift. this indicates the need for the current study to also examine eda peaks as these peaks are not affected by drift (rq4). when investigating mean st no significant mean differences were found for mean st across all different phases. nevertheless, descriptive statistics reveal that st was higher during the baseline measurement period. moreover, st was higher during the low complex task compared with the high complex task. this could indicate that st is related to task complexity as research indicated that st declines relative to a trigger event (ikehara & crosby, 2005). current findings indicate that mean eda and mean st might be indicators of changes of cognitive load, but cannot be used to detect differences in task complexity. nevertheless, there is no clear link between st and cognitive load. accordingly, correlations between individual differences in the perceived intrinsic load, mental effort and physiological data for a high and low complex task are investigated (rq3). 5.3 research question 3 a third aim of this study was to investigate whether we can relate subjective measures of the perceived intrinsic load and mental effort (i.e., based on findings of rq1) with physiological data (i.e., mean eda and st) during a high and low complex task. findings reveal that mental effort positively correlates with mean eda for the high complex intervention. nevertheless, we did not find a significant correlation between mean eda and mental effort during the problem-solving process of the low complex task. results might also be influenced by the fact that skin conductance was rising throughout the intervention. in addition, most students first solved the high complex task. no significant correlations between mean st and self-reported data were found. this finding could be due to the fact that st shows a very slow rise and decline in temperature change relative to the trigger event. therefore, it might be difficult to relate st to self-reports (ikehara & crosby, 2005). since, there seems to be a relationship between eda and mental effort and since st drops can be related to specific events, we investigated the relationship between physiological data and learning behaviour retrieved from log-data. 5.4 research question 4 in order to investigate the relationship between physiological data and learning behaviour. log-data was investigated and divided into four main events, namely, reading instructions, writing an answer, consulting support and reviewing the answer. results reveal that there seems to be a relationship between specific learning behaviour and eda peaks. moreover, results reveal that more peaks were registered during the high complex task, when compared with the low complex task, which indicates a different result compared to rq2. when investigating the intensity of the peaks, findings reveal that the peaks that are related to the events ‘submission’ are more intense. this might explain, besides the occurrence of drift, why mean eda was higher during the low complex task. possibly, results may have been influenced by the fact that the low complex task was presented as a test-format, which might induced more intensive peaks when students submitted their task. when investigating relations between peaks and events it seems that during the high complex task, peaks are more frequently related to cognitive processes (e.g., reading instructions, consulting support and writing) when compared with the low complex task (e.g., submission). for instance, when investigating the event ‘consultation of support’ more in detail, peaks were related to students (n = 4) watching a video that explains the circumference of a circle. this is line with previous research indicating that gsr responses are associated with effortful cognitive processing during multimedia learning (antonietti, colombo & di nuzzo, 2015). additionally, hardly any peaks were found for the low complex task during writing, which is in line with the study of mudrick, taub, azevedo, price & lester (2017). mudrick et al. (2017) investigated multimedia learning and indicated that the lowest amount of gsr responses were retrieved when answering multiple choice questions, suggesting that this might require less cognitive processing. this finding is also in line with the study of hoogerheide et al. (2018) indicating that mean eda was significantly lower during the problem-solving process of a practice problem, when compared with teaching a practice problem in an authentic learning situation. these exploratory findings indicate that the intensity of eda signals might be more related to the type of learning activities. in line with previous findings of rq2 and rq3, no conclusive results were found for st. nevertheless, on the basis of data visualisation of all students we could see that for the largest number of participants (i.e., 8 students), st is lower during the high complex task, which is in line with findings of rq2. 5.5 limitations and further research despite the merits of the study in terms of indicating that individual differences in experienced mental effort can indicate individual differences in eda, there are some important limitations that should be mentioned. firstly, results must be approached carefully as multiple analyses on the same dependent variable were conducted which can increase the chance of committing a type 1 error (roth, 1999). secondly, as we were investigating physiological data, we were obliged to implement a within-subject design. this is required when investigating skin conductance, as skin conductance can vary markedly between individuals (braithwaite et al., 2013). nevertheless, the within subject design had some important disadvantages. since the same learning materials were taught within both the high and low complex task, students might have learned from the previous task and therefore perceived the high complex task as less difficult. this is turn might have influenced skin conductance and skin temperature, and may be a reason why there was no clear difference between the high and low complex task. this problem can be addressed in future studies by addressing different topics. moreover, future studies should offer more different tasks of different levels of complexity, and also create more conditions in order to increase the amount of measurements. this could provide a better understanding of possible correlations between mental effort and mean eda. a third important limitation, when investigating skin conductance is drift, a continuous increase of the intensity of the signal. it is important to distinguish drift from important shifts in real tonic processes (braithwaite et al., 2013). nevertheless, this distinction between drift and real tonic processes is not always entirely clear. this emphasizes the need of an accurate baseline measurement. the baseline measurement in the current study could be optimized by giving the participants a moment of relaxation. given the small sample size we decided not to remove data of participants. instead, in this study we have additionally investigated the peaks of skin conductance (as these are no subject of drift) and related them to specific events in the learning environment. nevertheless, it can be advisable to remove data of participants on the basis of drift in larger datasets. moreover, a larger sample size would also allow us to investigate patterns between eda peaks and specific events in the learning environment (e.g., reading instructions) while using quantitative methods. finally, as the study did not take place in a lab setting but in the classroom of the students, a lot of confounding factors unrelated to cognitive load may cause clouds in the data such as a lecturer entering the classroom and students leaving the classroom when finished. these events are likely to degrade the accuracy of cognitive load measurement by gsr (i.e., eda). nevertheless, the ecological valid setting also has advantages such as authenticity of the results (schmuckler, 2001). moreover, as the content was part of students’ training program, students were encouraged to thoroughly solve the tasks, which is reflected in the task performance. 6. conclusion this study attempted to firstly investigate the difference of subjective measures of cognitive load and physiological data (i.e., mean eda and st) between a high and low complex task in an ecologically valid setting. students indicated that they perceived higher intrinsic load during the high complex task and that the high complex task required more mental effort. this indicates that task complexity can be manipulated based on element interactivity. nevertheless, complexity was not reflected by differences in physiological data (i.e., mean eda and st). accordingly, in a next phase this study investigated correlations between perceived intrinsic load, mental effort and physiological data. results revealed a positive correlation between mean eda and mental effort during the high complex task. nevertheless, no significant correlations were found for the low complex task. preliminary results of a more descriptive analysis showed that peaks of eda during the high complex task were more frequently related to cognitive processes when compared with the low complex task (i.e., submitting the task). the latter finding might explain the significant relationship between mental effort and mean eda. future research should replicate similar studies while using larger sample sizes to verify these findings. additionally, the relationship between eda and the type of learning behaviour (i.e., retrieved from log-data) should not be overlooked. keypoints preliminary results indicate that mean eda is correlated with self-reported mental effort. results indicate that perceived intrinsic load can be manipulated based on element interactivity, which is in line with the cognitive load theory. it is important for future research to investigate correlations between subjective measurements and physiological data while using large sample sizes. when investigating eda, it is important to investigate peaks of skin conductance in combination with specific events retrieved from log-data. this might reveal patterns and provide more insight into the influence of the learning behaviour on skin conductance. references antonenko, p., paas, f., grabner, r., & van gog, t. (2010). using electroencephalography to measure cognitive load. educational psychology review, 22. 425-438. doi:10.1007/s10648-010-9130-y antonietti, a., colombo, b., & di nuzzo, c. (2015). metacognition in self-regulated multimedia learning: integrating behavioural, psychophysiological and introspective measures. learning, media and technology, 40. 187-209. doi:10.1080/17439884.2014.933112 baayen, r. h., davidson, d. j., & bates, d. m. (2008). mixed-effects modeling with crossed random effects for subjects and items. journal of memory and language, 59. 390-412. doi:10.1016/j.jml.2007.12.005 boekaerts, m. (2017). cognitive load and self-regulation: attempts to build a bridge. learning and instruction, 51, 90–97. doi:10.1016/j.learninstruc.2017.07.001 braithwaite, j., watson, d., jones, r., & row, m. (2013). a guide for analysing electrodermal activity (eda) & skin conductance responses (scrs) for psychological experiments. psychophysiology, 49, 1017–1034. doi:10.1017.s0142716405050034 brunken, r., plass, j. l., & leutner, d. (2003). direct measurement of cognitive load in multimedia learning. educational psychologist, 38. 53-61. doi : 10.1207/s15326985ep3801_7 chang, c. c., & yang, f. y. (2010). exploring the cognitive loads of high-school students as they learn concepts in web-based environments. computers and education, 55. 673-680. doi: 10.1016/j.compedu.2010.03.001 chen, f., zhou, j., wang, y., yu, k., arshad, s. z., khawaji, a., & conway, d. (2016). robust multimodal cognitive load measurement. human-computer interaction series. doi: 10.1007%2f978-3-319-31700-7 cierniak, g., scheiter, k., & gerjets, p. (2009). explaining the split-attention effect: is the reduction of extraneous cognitive load accompanied by an increase in germane cognitive load? computers in human behavior, 25. 315-324. doi: 10.1016/j.chb.2008.12.020 de bruin, a. b. h., & van merriënboer, j. j. g. (2017). bridging cognitive load and self-regulated learning research: a complementary approach to contemporary issues in educational research. learning and instruction, 51. 1-9. doi: 10.1016/j.learninstruc.2017.06.001 deleeuw, k. e., & mayer, r. e. (2008). a comparison of three measures of cognitive load: evidence for separable measures of intrinsic, extraneous, and germane load. journal of educational psychology, 100, 223-234. doi: 10.1037/0022-0663.100.1.223 glogger-frey, i., gaus, k., & renkl, a. (2017). learning from direct instruction: best prepared by several self-regulated or guided invention activities? learning and instruction, 51. 25-35. doi:10.1016/j.learninstruc.2016.11.002 hoogerheide, v., renkl, a., fiorella, l., paas, f., & van gog, t. (2018). enhancing example-based learning: teaching on video increases arousal and improves problem-solving performance.journal of educational psychology, 211. 45-56. doi: 10.1037/edu0000272 haapalainen, e., kim, s., forlizzi, j. f., & dey, a. k. (2010). psycho-physiological measures for assessing cognitive load. proceedings of the 12th acm international conference on ubiquitous computing. doi: 10.1145/1864349.1864395 herborn, k. a., graves, j. l., jerem, p., evans, n. p., nager, r., mccafferty, d. j., & mckeegan, d. e. f. (2015). skin temperature reveals the intensity of acute stress. physiology and behavior, 1. 225-230. doi: 10.1016/j.physbeh.2015.09.032 ikehara, c., & crosby, m. (2005). assessing cognitive load with physiological sensors. proceedings of the 38th hawaii international conference on system sciences. doi: 10.1109/hicss.2005.103 jonassen, d. h. (2000). toward a design theory of problem solving.educational technology research and development, 48. 63-85. doi: 10.1007/bf02300500 karthikeyan, p., murugappan, m., & yaacob, s. (2012). descriptive analysis of skin temperature variability of sympathetic nervous system activity in stress. journal of physical therapy science, 24. 1341-1344. doi: 10.1589/jpts.24.1341 kirschner, p. a., ayres, p., & chandler, p. (2011). contemporary cognitive load theory research: the good, the bad and the ugly. computers in human behavior, 27. 99-105. doi: 10.1016/j.chb.2010.06.025 kirschner, f., kester, l., & corbalan, g. (2011). cognitive load theory and multimedia learning, task characteristics and learning engagement: the current state of the art. computers in human behavior, 27 , 1-4. doi: 10.1016/j.chb.2010.05.003 klepsch, m., schmitz, f., & seufert, t. (2017). development and validation of two instruments measuring intrinsic, extraneous, and germane cognitive load. frontiers in psychology, 8. doi: 10.3389/fpsyg.2017.01997 larmuseau, c., elen, j., & depaepe, f. (2018). the influence of students’ cognitive and motivational characteristics on students’ use of a 4c/id-based online learning environment and their learning gain. in lak'18:international conference on learning analytics and knowledge, march 7–9, 2018, sydney, nsw, australia. acm, new york, ny, usa , 10 pages. doi: 10.1145/3170358.3170363 lecroy, c. w., & krysik, j. (2007). understanding and interpreting effect size measures. social work research, 31. 243-248. doi: 10.1093/swr/31.4.243 leppink, j., paas, f., van der vleuten, c. p. m., van gog, t., & van merriënboer, j. j. g. (2013). development of an instrument for measuring different types of cognitive load. behavior research methods, 45. 1085-1072. doi: 10.3758/s13428-013-0334-1 liapis, a., katsanos, c., sotiropoulos, d., xenos, m., & karousos, n. (2015). recognizing emotions in human computer interaction: studying stress using skin conductance. in lecture notes in computer science (including subseries lecture notes in artificial intelligence and lecture notes in bioinformatics). 255-262. doi: 10.1007/978-3-319-22701-6_18 martin, s. (2014). measuring cognitive load and cognition: metrics for technology-enhanced learning. educational research and evaluation, 20. 592-621. doi: 10.1080/13803611.2014.997140 mayer, r. e. (2014). incorporating motivation into multimedia learning. learning and instruction, 29. 171–173. doi: 10.1016/j.learninstruc.2013.04.003 mayer, r. e., & moreno, r. (2003). nine ways to reduce cognitive load in multimedia learning. educational psychologist, 38. 43-52. doi: 10.1207/s15326985ep3801_6 merrill, d. (2009). first principles of instruction. in instructional-design theories and models, 50. 43-59. doi: 10.4324/9780203872130 mudrick, n. v., taub, m., azevedo, r., price, m. j., & lester, j. (2017). can physiology indicate cognitive, affective, metacognitive, and motivational self-regulated learning processes during multimedia learning? paper presented at the annual meeting of the american educational research association (aera), san antonio, tx. nourbakhsh, n., wang, y., chen, f., & calvo, r. a. (2012). using galvanic skin response for cognitive load measurement in arithmetic and reading tasks. proceedings of the 24th conference on australian computer-human interaction ozchi ’12 . doi: 10.1145/2414536.2414602 paas, f. (1992). training strategies for attaining transfer of problem solving skills in statistics: a cognitive load approach. journal of educational psychology, 84, 429–434. doi: 10.1037/0022-0663.84.4.429 paas, f., van gog, t., & sweller, j. (2010). cognitive load theory: new conceptualizations, specifications, and integrated research perspectives. educational psychology review, 2. 115-121. doi: 10.1007/s10648-010-9133-8 paas, f., tuovinen, j., tabbers, h., & van gerven, p. w. m. (2010). cognitive load measurement as a means to advance cognitive load theory. educational psychologist, 38. 63-71. https://doi.org/10.1207/s15326985ep3801 raaijmakers, s. f., baars, m., schaap, l., paas, f., & van gog, t. (2017). effects of performance feedback valence on perceptions of invested mental effort. learning and instruction, 51. 35-46. doi: 10.1016/j.learninstruc.2016.12.002 roth, a. j. (1999). multiple comparison procedures for discrete test statistics. journal of statistical planning and inference, 82. 101-117. doi: 10.1016/s0378-3758(99)00034-8 scharinger, c., soutschek, a., schubert, t., & gerjets, p. (2015). when flanker meets the n-back: what eeg and pupil dilation data reveal about the interplay between the two central-executive working memory functions inhibition and updating. psychophysiology. 1293-1304. doi: 10.1111/psyp.12500 schmeck, a., opfermann, m., van gog, t., paas, f., & leutner, d. (2015). measuring cognitive load with subjective rating scales during problem solving: differences between immediate and delayed ratings. instructional science, 43. 93-114. doi: 10.1007/s11251-014-9328-3 schmuckler, m. a. (2001). what is ecological validity? a dimensional analysis. infancy, 2, 419–436. doi: 10.1207/s15327078in0204_02 schreiber, j. b., nora, a., stage, f. k., barlow, e. a., & king, j. (2006). reporting structural equation modeling and confirmatory factor analysis results: a review. the journal of educational research, 99, 323-338. doi: 10.3200/joer.99.6.323-338 setz, c., arnrich, b., schumm, j., la marca, r., tröster, g., & ehlert, u. (2010). discriminating stress from cognitive load using a wearable eda device. ieee transactions on information technology in biomedicine, 14. 410-417. doi: 10.1109/titb.2009.2036164 shi, y., ruiz, n., taib, r., choi, e., & chen, f. (2007). galvanic skin response (gsr) as an index of cognitive load. in chi ’07 extended abstracts on human factors in computing systems chi ’07 . doi: 10.1145/1240866.1241057 shusterman, v., anderson, k. p., & barnea, o. (1997). spontaneous skin temperature oscillations in normal human subjects. the american journal of physiology, 273. doi: 10.1152/ajpregu.1997.273.3.r1173 smets, e., velazquez, e. r., shiavone, g., chakroun, i., d’hondt, e., de raedt, … van hoof, c. (2018). large-scale wearable data reveal digital phenotypes for daily-life stress detection. digital medicine, 67. doi: 10.1038/s41746-018-0074-9 spanjers, i. a. e., van gog, t., & van merriënboer, j. j. g. (2012). segmentation of worked examples: effects on cognitive load and learning.applied cognitive psychology, 26. 353-358. doi: 10.1002/acp.1832 sweller, j. (1994). cognitive load theory, learning difficulty, and instructional design. learning and instruction, 4. 295-312. doi: 10.1016/0959-4752(94)90003-5 sweller, j. (2010). element interactivity and intrinsic, extraneous, and germane cognitive load. educational psychology review, 22. 123-138. doi: 10.1007/s10648-010-9128-5 sweller, j., ayres, p., & kalyuga, s. (2011). cognitive load theory. explorations in the learning sciences, instructional systems and performance technologies. doi: 10.1007/978-1-4419-8126-4 sweller, j., van merrienboer, j., & paas, f. (1998). cognitive architecture and instructional design. educational psychology review, 3. 251-196. doi: 10.1023/a:1022193728205 van merrienboer, j. j. g., kirschner, p. a., & kester, l. (2003). taking the load off a learner’s mind: instructional design for complex learning. educational psychologist, 38, 5–13. doi: 10.1207/s15326985ep3801_2 van merriënboer, j. j. g., & sluijsmans, d. m. a. (2009). toward a synthesis of cognitive load theory, four-component instructional design, and self-directed learning. educational psychology review, 21. 55–66. doi: 10.1007/s10648-008-9092-5 vinkers, c. h., penning, r., hellhammer, j., verster, j. c., klaessens, j. h. g. m., olivier, b., & kalkman, c. j. (2013). the effect of stress on core and peripheral body temperature in humans. stress, 16. 520-520. doi: 10.3109/10253890.2013.807243 yousoof, m., & sapiyan, m. (2013). measuring cognitive load for visualizations in learning computer programming-physiological measures. ubiquitous and communication journal, 8. 1410-1426. retrieved from : https://pdfs.semanticscholar.org/bdbd/8af1870e956e4a727e2449897077266fa8e5.pdf zagermann, j., pfeil, u., & reiterer, h. (2016). measuring cognitive load using eye tracking technology in visual computing. proceedings of the sixth workshop on beyond time and errors on novel evaluation methods for visualization. doi: 10.1145/2993901.2993908 zheng, r., & cook, a. (2012). solving complex problems: a convergent approach to cognitive load measurement. british journal of educational technology, 43. 233-246. doi: 10.1111/j.1467-8535.2010.01169.x codepen bakhtiar publication frontline learning research vol.8 no. 2 (2020) 1 34 issn 2295-3159 dynamic interplay between modes of regulation during motivationally challenging episodes in collaboration aishah bakhtiara, allyson f. hadwinb auniversity of victoria, canada article received 19 agustus 2019/ revised 12 december / accepted 21 february/ available online 24 march abstract the cognitive and social demands of collaboration can raise significant motivation challenges. task progression relies on team members strategically taking control of the problems and adapting accordingly. theory indicates that productive collaboration involves groups using three modes of regulation: self-regulation, co-regulation, and socially shared regulation. despite research demonstrating the occurrence of all three modes in collaboration, it is unclear how these modes interact and how co-regulation supports the emergence of self and shared-regulation of motivation. the study aimed to examine the role co-regulation played in dynamically stimulating the emergence of selfand shared-regulation of motivation. a cross-case comparison was conducted between two groups who experienced high levels of motivation challenges but achieved contrasting perceptions of the overall team learning productivity. during analysis, groups’ dynamic regulatory processes within the online environment were visually represented using a tool called the chronologically-ordered representation for tool-related activity (cordtra). findings demonstrate that co-regulation of motivation may afford and thwart the emergence of selfand shared-regulation, and these processes interacted with the group’s situational challenges and the regulatory skills group members possessed. comparisons between the two groups indicated that groups' motivation regulation should (a) match the demands of the challenges at hand, (b) be positively supported by group members through co-regulation, and (b) involve a more varied strategic responses so that the group may continue to learn and co-construct knowledge effectively as a team. keywords: motivation; self-regulated learning; co-regulation; socially shared regulation; cordtra diagram info corresponding author: email: aishah@uvic.ca doi: 10.14786/flr.v8i2.561 1. introduction in recent years, there has been a surge of research examining students’ collaboration in physical classrooms and online learning environments (puntambekar, erkens, & hmelo-silver, 2011; järvelä & hadwin, 2013). research on collaborative learning (a) investigates how group members co-construct and build upon each other’s knowledge and ideas and (b) examines groups' and individual members' regulation of cognition, behaviour, motivation, and emotions in the learning task (järvelä & hadwin, 2013). the focal point of the second line of research is the dynamic interaction between individual and group regulation. hadwin, järvelä, and miller (2011, 2018) posited that regulation in collaboration involves three modes: self-regulation, socially shared regulation, and co-regulation. self-regulation refers to individual learners’ processes of regulating their own cognition, behaviour, motivation, and emotions for personal goals or in the service of the group goal. socially shared regulation (referred to as shared regulation hereafter) involves the group jointly regulating their cognitive, behavioural, motivational, and emotional states toward a shared outcome. in line with the sociocultural lens of hadwin, järvelä, et al.’s (2018) work, the emergence of selfand shared-regulation are temporarily supported or thwarted through co-regulation (hadwin, järvelä, et al., 2018). this process involves the reciprocal interaction between the personal, social, and cultural elements, which may facilitate or frustrate the internalization of regulatory skills or information at the individual and group level (mccaslin & vriesema, 2018). thus, co-regulation is more than just capable individuals directing or supporting the learning of less capable individuals. it can be initiated by a single individual, multiple individuals, tools, or the physical task, and it can be directed toward either individuals or the group. however, in many studies, co-regulation is often limited to prompts and guides directed toward individual learners (e.g., didonato, 2013; järvelä, malmberg, & koiveneimi, 2016; panadero & järvelä, 2015). limited research investigates co-regulation as a process that benefits or undermines either selfor shared-regulation. in response to this shortcoming, hadwin, järvelä, et al. (2018) revised their definition and conceptualization of co-regulation form the original (hadwin et al., 2011) depiction to emphasize the importance of those prompts being taken up by individuals or groups in co-regulatory episodes. addressing this scarcity requires an in-depth and in-context investigation of the intertwined individual and group processes. most research on collaboration concerns cognitive regulation and cognitive outcomes. rogat, linnenbrink, & didonato (2013) conclude that motivation regulation and motivational outcomes have been largely ignored in the collaborative learning literature. even when designing computer support for collaborative learning, the motivational aspect of collaboration is often neglected (belland, kim, & hannafin, 2013). research, however, shows that motivation is instrumental in directing and stimulating cognitive processes which are needed when learners work on joint tasks (rogat et al., 2013). research demonstrates that shared regulation of motivation affords opportunities for the groups to engage in other forms of regulation, particularly those that involve regulating differences in perspectives and understanding related to the cognitive aspect of the task (järvelä, malmberg, & koivuneimi, 2016; malmberg, järvelä, järvenoja, & panadero, 2015). nonetheless, working in small groups can raise significant motivation challenges. inadequate responses to these challenges can lower learners’ motivation to engage in future collaboration (ruiz ulloa & adams, 2004). this negative consequence counters the very aim of collaborative learning which is to promote high engagement with the learning materials (blumenfeld, kempler, & krajcik, 2006). motivation challenges have been frequently reported to be a barrier to collaborating with others (järvenoja, volet, & järvelä, 2013). these challenges may originate from multiple sources including variations in individual members’ motivation and attitude toward the collaborative task (e.g., wosnitza & volet, 2014), interpersonal dynamics related to work ethics and personalities (e.g., tucker & abbasi, 2016), and difficulties in negotiating multiple perspectives, ideas, and goals related to the task (e.g., barron, 2003). although motivation challenges are common, it is unclear how multiple individuals with diverse motivational beliefs and goals negotiate their motivation for the group task. it has been proposed that motivation is fostered, influenced, and maintained through an ongoing process of co-regulation (mccaslin & hickey, 2001). hence, this study aimed to investigate the role co-regulation plays in stimulating selfand shared regulation of motivation in collaboration. 1.1 co-occurrence of self, co, and shared-regulation research indicates that self-, co-, and shared-regulation co-emerge and mutually reinforce one another. panadero, kirschner, järvelä, malmberg, and järvenoja (2015) demonstrated the link between individual members' self-regulation skills and the groups' ability to socially share their regulation of learning, motivation, and emotions. based on self-report data about individual and group regulation, panadero et al. found that groups composed of individuals with greater self-regulation skills in managing their cognition, motivation, and emotion were more likely to demonstrate better shared-regulation during collaboration. similarly, bakhtiar, webster, & hadwin (2018) found that a group who was more successful in adapting as a team had individuals who (a) were more punctual with their individual assignments and (b) had a better understanding of the task instructions. in pino-pasternak, whitebread, and neale’s (2018) study with pre-schoolers, group members possessing higher global self-regulation skills were found to actively co-regulate their peers and consequently increase team productivity. overall, research suggests that for shared processes to occur efficiently, individual members must actively self-regulate and support the regulation of others in the service of the group goal. group dynamics and processes have also been shown to play a role in promoting and shaping groups' shared regulation of motivation. by analyzing groups' interaction, dialogue, and behaviour, research demonstrates that positive socioemotional interaction is associated with instances of high-quality and more adaptive shared regulation (isohätälä, jävenoja, & järvelä, 2017; rogat & linnenbrink-garcia, 2011). positive socioemotional interaction has been operationalized as involving group members being attentive to one another, acknowledging others’ contributions, and respectfully soliciting opinions. research also demonstrates that co-regulation when performed in a directive, rather than facilitative, manner can negatively influence the group dynamics and shared processes (rogat & adam-wiggins, 2015). rogat and adam-wiggins characterized directive co-regulation as involving the co-regulator asking individuals to comply with specific ideas and approaches without much consideration of the other person’s point of view. in contrast, facilitative co-regulation considers the well-being and motivation of team members and is often pro-actively activated to avoid experiencing challenges that might disrupt progress. these studies highlight the importance of positive reciprocal contributions and participations among group members when regulating as a team. a review on shared regulation indicates that the phenomenon has often been characterized through qualitative data on groups’ conversations during collaboration (panadero & järvelä, 2015). the authors point out a trend suggesting that collaboration with higher instances of co-regulation characterized by one individual taking the lead has been less optimal for performance than collaboration with higher instances of shared regulation. note, however, only a few of the studies reviewed by panadero and järvelä (2015) provided evidence for the link between performance and shared regulation. moreover, shared regulation in those studies tended to focus on regulation that is more cognitive-focused rather than motivational or emotional-focused. as observed by the same authors, the theoretical distinctions between co-regulation and shared regulation are rarely made in the studies they reviewed (see also hadwin et al., 2011). since the review, research evidences mixed findings about the frequency effect of coand shared-regulation on group performance. for example, schoor and bannert (2012) reported no difference in the amount of shared regulation between highand low-performing dyads. regardless of these findings, we argue that the frequency of “sharedness” is less important than activating appropriate regulation modes to fit the situational demands and challenges. this argument is particularly relevant in motivation regulation because research indicates that self-regulating one's motivation and co-regulating group members' motivation alone can be sufficient without needing shared-regulation to frequently occur (järvenoja, järvelä, & malmberg, 2017). in järvenoja et al.’s study, co-regulation of motivation involved one or several group members trying to “increase motivation within the group or offered support, advice or encouragement when group members expressed a lack of motivation or negative feelings affecting the group work” (p. 6). for instance, when an individual experienced a lack of interest, having another group member suggest making the task personally relevant was enough to increase the individual’s motivation. fundamentally, self-, co-, and shared regulation are activated based on the needs and the types of challenges the situation presents (hadwin, järvelä, et al., 2018). 1.2 regulating during motivationally challenging situations hadwin, järvelä, et al.’s (2011, 2018) conceptualized self-, co-, and shared-regulation as being driven by the same cognitive architecture described in winne and hadwin (1998). this cognitive architecture is represented by the acronym copes (conditions, operations, products, evaluations, and standards) which describes several interacting elements during a process of regulation. the copes-typology highlights the need for viewing regulation as triggered by specific contextual demands or challenges that stall goal progress. specifically, a person’s conditions create contexts for regulation. in collaboration, conditions can exist in three areas: (a) self, (b) group, and (c) task and context (hadwin, järvelä, et al., 2018). self-conditions include personal characteristics, beliefs, and histories individuals bring to the task. in contrast, group conditions include individual perceptions of their group’s shared characteristics, knowledge, norms, and histories. task and context conditions include individual perceptions about the situation, including affordances and constraints in the task and environmental context. all three conditions dynamically interact and produce complex contexts for regulation. conditions are realized through metacognitive monitoring either self-activated or nudged by others or tools. researchers may introduce metacognitive prompts to encourage students to stop and reflect on their current conditions or challenges and promote students to proactively take control of their situations (e.g., järvelä et al., 2015; järvenoja et al., 2017). evidence suggests that prompts and awareness tools targeting motivation and emotions introduced early in the collaboration can be effective in supporting groups’ socioemotional experience which is essential in productive collaborative learning (järvenoja et al., 2017; näykki, isohätälä, järvelä, pöysä-tarhonen, & häkkinen, 2017). in theory, when students reflect on their conditions and evaluate them as warranting further interventions, learners may exert control by engaging cognitive operations to process and manipulate relevant information. in terms of motivation regulation, the information is motivation-related (winne & hadwin, 2008; winne & marx, 1989). for example, an individual may search for information about their efficacy belief, boost that belief by assuring they can do the task, and focus on that information to persist in the task (srl). likewise, individuals can help others recognize information about their motivation and use it to support others’ regulation (corl). when regulating as a team, group members may articulate personal concerns regarding the group’s overall engagement and negotiate to work toward one common goal. the results of operations create products that can manifest cognitively (e.g., increased understanding of task goal and value), behaviourally (e.g., persistence in the task), and emotionally (e.g., in the state of flow) during the task. learners then construct judgments or evaluations of the products by comparing them to specified or perceived standards or goals. research demonstrates that groups’ regulatory processes are different under varying levels of challenges (sobocinski, malmberg, & järvelä, 2017). sobocinski et al. examined the temporality of the regulatory phases (forethought, performance, and reflection) in high and low challenge events across eleven collaborative groups. high and low levels of challenges were determined based on groups' aggregated responses related to their cognitive, motivational, and emotional states collected at the beginning of each collaborative session. the observed moment-by-moment regulatory actions in each challenge level were then fed into a process mining software to detect the most common sequence of actions. findings showed that the frequency of regulatory processes was similar in high and low challenge level events. however, the sequential pathways of the processes slightly differed: during a high level of challenge, groups switched between forethought and performance more often. although there was little information on the types of difficulties the groups experienced and the intended purpose of engaging in the frequent switch, sobocinski et al.’s findings suggest that more intense challenges require learners to adapt and refine their approaches more frequently. experiencing more challenges, however, does not necessarily equate to performing poorly in the task. groups who experienced similarly high levels of challenges may end up with different outcomes as each group’s situational affordances and constraints may vary (järvenoja, järvelä, & malmberg, 2015). researchers argue that successful regulated learners are identified by their ability to overcome situation-specific challenges; these individuals often exhibit more accurate metacognitive awareness of their situations and possess a more varied strategy repertoire (hadwin & winne, 2012). responses that are not fitting with the task demands or the challenges present in the task would theoretically be less effective than responses that directly address the situational needs and challenges (hadwin & winne, 2012). examining what many learners do in response to highly challenging situations may bring to light critical features of strategic actions that may be more adaptive for learning. motivation challenges in collaboration emerge from a variety of individual and situational sources (rogat et al., 2013; järvelä, järvenoja, & veermans, 2008). bakhtiar, hadwin, and järvenoja (2018) argued that there are several types of motivation challenges. the challenges can be categorized into (a) behaviour-based—difficulties related to effort initiation and maintenance; (b) cognitive-based—difficulties related to cognitive beliefs about one’s competence, task value, and goals; (c) affective-based—difficulties related to task enjoyment; and (d) externally-based—difficulties related to environmental distractors pulling away one’s attention from the task. these challenges can lower individuals’ and the group’s motivation to engage in a collaborative task. bakhtiar, hadwin, et al. (2018) found that group members tended to resort to behaviour control strategies to address difficulties related to group members’ effort initiation and participation in the task. these strategies were in the form of providing social support, focusing on task completion, and checking each other’s progress. behaviour control strategies were more dominant regardless of the data indicating that cognitive control strategies, such as planning and setting or revising goals, to be better at reducing the likelihood of encountering the same type of motivation challenge again. the authors’ analyses were based on data about individual members’ perceptions of their strategy use. it is still unclear how motivation regulation strategies are performed under coor shared-regulation modes. this current study addressed this gap by exploring motivation regulation strategies learners deployed individually (srl), for others (corl), and collectively as a team (ssrl). 1.3 how can the dynamic interplay between modes of regulation be captured? investigating how individual team members dynamically negotiate their motivation for the group task involves examining the specific actions that unfold during groups’ episodes of motivation regulation (järvenoja et al., 2018). such an undertaking often requires researchers to intricately code group members’ utterances across multiple dimensions of regulation to make sense of the multi-layered processes (see volet & vauras, 2013). these dimensions can include micro-processes (planning, monitoring, adapting), targets (behaviour, cognition, motivation), and modes of regulation (see bakhtiar, webster, et al., 2018; rogat & linnenbrink-garcia, 2011). however, code frequencies are typically examined according to one dimension at a time, which undermines the dynamic interactions between the coded processes. our aim to examine the dynamic interplay between modes of regulation during motivationally challenging episodes implicates finding a method that can holistically capture (visually) the complex processes. the use of each mode of regulation must also be contextualized by including information about (a) the individuals involved in the regulation, and (b) the specific motivation regulation strategy enacted within the episode. thus far, previous studies have explored three approaches for capturing groups’ regulatory processes using (a) process mining tools (e.g., malmberg et al., 2015), (b) social network analysis (e.g., wijga, endedijk & veldkamp, 2019), and (c) probabilistic decision pathways (hadwin, bakhtiar, & miller, 2018). in these approaches, groups’ regulatory activities are visually represented to aid the analysis of related events: process miners generate common sequences of regulatory events, social network analyzers produce a network of relationships between group members’ regulatory activities, and probabilistic pathways are represented using directed graphs to map learners’ transition from one event to another. one limitation of using these approaches lies in their inability to holistically capture the activities involved in multiple episodes of regulation. at best, the visualization focuses on either the dynamic person-to-person relationships (person-focused) or the frequently traversed sequence of actions (activity-focused) aggregated across episodes of regulation. in this study, we explored a method for visually representing the dynamic interplay between individual and group motivation regulation called the chronologically-ordered representation for tool-related activity (cordtra) diagram (see hmelo-silver, chernobilsky, & jordan, 2008). the cordtra was initially designed to capture group members’ interactions with support tools in a computer-supported collaborative learning environment. for instance, a diagram may show that, within a defined time block, specific individuals are viewing a discussion board while also generating ideas aloud. the cordtra diagram is adaptable to data about group regulation because the diagram can visually represent fine-grained and multi-dimensional codes associated with regulation. specifically, the diagram is a scatterplot time-series of coded data that can represent individual members and specific regulatory actions. arranging each group member’s activities chronologically allows researchers to unpack how individual and group regulation evolve during the collaboration. cordtra diagrams provide opportunities to examine the larger patterns that emerge between codes and foster holistic visualization of the dynamic processes. during pattern identification, researchers may cycle back and forth between the diagram and the actual conversation data for further context on the specific utterances involved during those activities. 1.4 purpose the purpose of this study was to examine the role co-regulation of motivation plays in dynamically stimulating the emergence of selfand shared-regulation of motivation in two groups with contrasting perceptions of their overall team learning productivity. a cross-case analysis was conducted between two groups who experienced similarly high-level motivation challenges but demonstrated contrasting outcomes in terms of their team learning behaviour or effectiveness in co-constructing ideas and knowledge during collaboration. groups were compared based on three guiding research questions: i. what motivation challenges trigged regulation? ii. how were selfand shared-regulation used in relation to co-regulation? iii. what strategies were used during episodes of motivation regulation? 2. methods 2.1 case study a case study method was chosen for conducting the cross-case comparisons for three reasons. first, the dynamic interplay between groups’ self-, co-, and shared regulation in response to motivation challenges is a complex social phenomenon, and it is not clear how the dynamic process occurs during learning. case studies have been argued to be appropriate for answering the how questions (yin, 2003). second, case studies involve in-depth and in-context empirical investigations in which there is no manipulation of behaviour. instead of examining the effect of a specific manipulation, case studies involve using multiple data sources to understand the phenomenon in question. the collaborative interaction examined in this study was not manipulated as we were interested in investigating learners’ responses to naturally occurring motivationally challenging events. third, while case studies are not meant for generalization to populations, findings are generalizable to theoretical propositions (yin, 2003). in this study, we tested (and consequently informed) the theory relating to self-, co-, and shared regulation as proposed in hadwin, järvelä, et al. (2018). 2.2 the collaborative project participants were enrolled in an undergraduate elective educational psychology course on learning strategies and academic success. students in the course came from a wide range of disciplines and levels of academic achievements. in the course, students were required to complete a graded semester-long collaborative project in groups of three to five. the project required groups to produce and present a strategy for tackling one aspect of university learning in the form of an infographic. students approached the project in eight work sessions distributed across the semester: individual planning task, group planning task, three online discussion sessions, building the infographic, presenting the final infographic, and individual reflection task. within each work session, cognitive and metacognitive prompts were introduced to promote productive collaborative learning. these prompts were also used as a data collection tool focusing on attaining students’ metacognitive reflections and decisions regarding their motivation in the collaboration. cumulatively, the project was worth 35% of the course mark. details of the work sessions are described below. 2.2.1 individual planning task this activity was designed to guide students through a process of systematically unpacking information about the project and using that understanding to strategically plan for the task. specifically, individual students (a) summarized their understanding of the collaborative project, (b) evaluated examples of past students’ infographics, (c) reflected on their goals and motivation for the task, and (d) constructed a plan for tackling the main motivation challenge they anticipated experiencing in the collaboration. students were given 15 minutes to complete this task during class. group assignment using individual planning scores. the individual planning task was worth 5% of the course mark. an excellent planning showed that the student leveraged the task to get ready for the project and provided thoughtful answers. short cryptic answers indicated poor planning. planners were graded by two course instructors who cross-checked their marking for inconsistencies. students within course sections were assigned to groups based on their planning scores. specifically, students were distributed to include group members with high (5%), moderate (3-4%), and low (1-2%) planning scores. this distribution was so that all groups had similar variability in terms of how well prepared its members were. 2.2.2 group planning task group members met in person during group planning. guided by eight open-ended prompts, groups were tasked to plan their approaches and interactions for the project. an example of a prompt included: what is our strategy for preparing for each part of the project? groups were given 20 minutes of class time to plan and were encouraged to record their meeting minutes. group planning was not graded. consequently, the meeting minutes were shallow, making them unusable for data analysis. 2.2.3 online discussions (3 sessions) in the next three sessions, groups collaboratively discussed and documented their ideas related to their group’s infographic using shared online documents, google docs. each session had a different topic focus assigned by the instructors, but the topic builds upon one another to help students gradually build their group’s final infographic. each session was worth 5% of the project mark. in all three sessions, groups were required to use an online text-based chat tool (google hangouts) to interact and discuss their ideas. essentially, during each session, each student had two online applications opened on their individual computer: (a) google hangouts for communicating with group members and (b) google docs for recording ideas and resources collaboratively (figure 1). groups were given 60 minutes of class time to work on each topic. incomplete work was assigned to be completed outside of class time using the same google tools. groups were given a week to submit work from each session. figure 1. example of a group’s online work session. google hangouts environment (left) and google docs environment (right). 2.2.4 building the infographic groups built an infographic face-to-face during class time. in this phase, groups turned their resources and ideas (collected in the previous online discussion sessions) into an infographic format. the infographic was produced within the online environment using available templates in google slides. group members were able to concurrently access, edit, and track changes to their infographic. 2.2.5 presenting the infographic in the final week, groups presented their finished infographic to other students in their course section. the final infographic and the accompanying group presentation were worth 10% of the project mark. 2.2.6 individual reflection task at the end of the project, students reflected on the overall collaborative experience which included (a) assessment of their team learning behaviour and the final product, (b) reflection on experienced challenges, and (d) a plan for improving their future collaborations. this portion of the assignment was worth 5% of the project mark. students were graded on the overall thoughtfulness of their reflection. 2.3 data sources figure 2 provides an overview of the collaborative project and the data sources gathered for this study. descriptions of each data source follow. figure 2. overview of the work sessions in the collaboration project (above) and the data sources gathered for this study (below). bolded data sources were used for the case sampling strategy. 2.3.1 individual pre-task motivation and planned response during individual planning, students completed a narrative constructor tool to reflect on their motivation for the task (figure 3). for each prompt, students select a response option from a drop-down menu. the tool asked individuals to (a) rate their task motivation, (b) provide a reason for that rating, (c) evaluate whether that motivation was a problem, (d) provide a goal for regulation, and (e) describe a strategy response. if the strategy was not one of the options, students had to specify their planned strategy in a text field. figure 3. a narrative constructor tool for collecting data about individual pre-task motivation and planned strategic response. 2.3.2 chat records groups’ text-based conversations during the three online discussion sessions served as the primary data source because they provided evidence of groups’ self-, co-, and shared-regulation. for each text-based utterance, google hangouts provided a timestamp, date, and the username belonging to the utterance. 2.3.3 scores for active collaboration in each online discussion for each online discussion session, the course instructors rated each group’s quality of collaborative conversations. groups received a score representing their teamwork quality including team members’ attempts at actively co-constructing ideas and knowledge as a group. groups received a full mark (2 out of 2) if every member in the team actively contributed in meaningful ways. a half mark (1 out of 2) was given if there was evidence that group members were trying to contribute in some ways but lacked active co-construction of ideas and knowledge. groups received a score of zero if they did not meet online to complete the assigned work. 2.3.4 activity logs associated with the shared online documents students' activity logs within each google document were automatically collected in the online environment. each log contained information about (a) who viewed or edited a document, (b) the document name, and (c) when the activity was performed. for each group member, the percentage of editing work the person completed relative to their group in each shared document was calculated. for example, when a person edited the group’s document 5 times out of the group’s total of 50 times, the person’s percent editing work for that work session would be 10%. this calculation was used to gauge each member’s contribution to each discussion session 2.3.5 ratings of experienced motivation challenges at the end of the project, each member rated the extent to which they experienced several types of motivation challenges during collaboration. three items were used to collect ratings about motivation challenges including (a) difficulties in getting individuals to participate, getting the work done, and staying on task (behaviour-based); (b) difficulties with cognitive beliefs about one’s competence, task value, and goals (cognitive-based) and (c) difficulties in maintaining positive attitudes and emotions about the task, the group, and the situation (affective-based). ratings were collected on a five-point likert scale from 1 = not a problem to 5 = major problem. for each student, ratings on all three items were summed to produce an individual score of experienced motivation challenges. at the group level, group members’ scores were averaged to produce the group’s mean ratings of experienced motivation challenges. within-group agreement regarding the challenges was denoted using the standard deviation of the group's mean score. 2.3.6 main motivation challenge and its regulation students were subsequently asked to reflect on one main motivation challenge their group encountered during the collaboration (figure 4). this reflection required students to construct a narrative about (a) the main motivation challenge the group experienced, (b) the strategy used to address the challenge, (c) why the strategy was selected, (d) the strategy effectiveness, and (e) who was involved in enacting the strategy. figure 4. a narrative constructor tool for collecting data about the main motivation challenge and strategy response. 2.3.7 ratings of the team learning behaviour in the reflection phase, group members also evaluated their team learning behaviour using a nine-item team learning behaviour questionnaire adopted from van den bossche, gijselaers, segers, and kirschner (2006). scores on the questionnaire describe the degree to which group members construct, co-construct, and build upon each other’s ideas and contributions. higher ratings indicate a more positive team learning behaviour demonstrated during collaboration. an example of an item on the questionnaire included, “my team members elaborated on each other's information and ideas.” all responses were collected on a seven-point likert scale ranging from 1 = strongly disagree to 7 = strongly agree. 2.4 case sampling strategy two groups were sampled from a pool of 20 groups using two criteria. (1) informed by research that indicates challenging events as opportunities to exercise regulation, both groups must have experienced similarly high levels of motivation challenges thereby having had similar opportunities to regulate motivation. (2) the two groups must have maximally different perceptions of the overall team learning productivity (i.e., team learning behaviour ratings). this latter criterion was chosen to identify the types of regulatory actions that were more successful and less successful in influencing team motivation and productivity. accordingly, groups were selected based on data collected at the end of the collaborative project related to group members’ post-collaboration judgment of (a) the level of motivation challenge experienced during collaboration and (b) the overall team learning behaviour their group demonstrated. the sample was identified in stages. first, groups with missing data from at least one member were removed, leaving 12 cases remaining. next, groups’ ratings of motivation challenges were categorized as: (a) high, including scores falling above 0.5 sd from the mean (n = 5), (b) medium, including scores between -0.5 to 0.5 sd around the mean (n = 1), and (c) low, including scores falling below 0.5 sd from the mean (n = 6). demarcations of the standard deviations were based on the average group ratings of motivation challenges across all groups in the sample (m = 6.15, sd = 1.71). next, to obtain the richest regulation data, we retained groups in the high challenge category (n = 5) assuming these groups had the most opportunities to regulate motivation challenges. finally, from the 5 cases, the group with the highest (group f) and the group with the lowest (group k) team learning behaviour score were selected because they represented maximally different group perception of the overall team learning productivity. the difference in team learning behaviour scores between the two selected groups was more than one standard deviation away based on the overall mean ratings of team learning behaviour across all groups (m = 56.21; sd = 6.27). hereafter, the group with a more positive perception of their overall team learning productivity was referred to as the high productivity group, and the group with a less positive perception of their overall team learning productivity was referred to as the low productivity group. table 1 summarizes the characteristics of the case groups. table 1 characteristics of each case group note: excluding group performance score, parentheses refer to the within-group standard deviation associated with the mean. productivity was defined as group members’ aggregated perceptions of their team’s overall ability to build upon on each other’s ideas and contributions and work efficiently as a team. 3. analysis data analysis was conducted in four steps. this systematic approach was essential for maintaining the chain of evidence and establishing rigour (seale & silverman, 1997). 3.1 data gathering the first step involved becoming familiarized with the data sources. elaborative running records of the groups’ conversations and activity logs were created by matching the timestamp of the activity logs and the chat conversations, providing a summary of all the events that occurred during each group’s online collaboration. individual members’ self-reports were also reviewed and cross-checked with the objective data. 3.2 coding of chat data. groups’ chat records from their online discussion sessions were coded in four waves (see appendix a and b for coding schemes). the unit of coding was at the episode level. given the nature of the chat data, a code could be assigned to a single utterance, multiple utterances that may have been separated in between another code, or several consecutive utterances as long as the utterances were still related to one purpose of actions. codes were mutually exclusive. in each wave, 30% of the data were given to a second-rater (a senior research assistant) for a reliability check. inter-rater reliability was determined by cohen’s kappa statistics. indices between .80 to .90 were considered strong reliability (mchugh, 2012). wave 1 : task-focused versus socioemotional-focused episodes. first, group conversations were segmented into task-focused and socioemotional-focused episodes following järvenoja et al.’s (2017) coding scheme. task-focused episodes referred to utterances related to task details, including domain related knowledge or items to include in the group’s shared product. socioemotional-focused episodes referred to utterances about motivation and emotions related to individuals, a group of individuals, task features, progress, or product. in the high productivity group, 39.59% of the utterances were socioemotional-focused. comparably, in the low productivity group, 36.38% of the utterances were socioemotional-focused. inter-rater reliability index for coding in this wave was cohen’s k = .89. wave 2: types of motivation challenges. the socioemotional conversations were further scrutinized for utterances depicting motivation challenges. 2.15% of socioemotional utterances in high productivity group and 7.46% of socioemotional utterances in low productivity group were not considered related to motivational challenges. the eliminated utterances pertain to off-task conversations, where group members shared about personal events (e.g., leaving town to see a family in the weekend) or general remarks about their university courses. groups’ remaining socioemotional utterances were broken down into several episodes of motivation challenges. in other words, the beginning of a challenge episode is marked by utterances that led to the emergence of a motivation challenges or related challenges (e.g., unclear task goals with low engagement), and the end of a challenge episode is marked by utterances that led to a dissolution of a challenge or related challenges. there were instances where a dissolution was not reached; the end of those episodes was marked by group members’ shifts in the topic of conversations. utterances that depict instances of motivation challenges were identified by demonstrated difficulties initiating or engaging in the task, or explicit statements about encountering a motivation hurdle. motivation challenge utterances were deductively categorized into four types: (a) behavioural—effort initiation and task persistence, (b) cognitive—competency beliefs, task value and utility, and goals, (c) affective—interest, enjoyment, positive and negative emotions, and (d) external—environmental distractors (bakhtiar, hadwin, et al., 2018). inter-rater reliability index of the motivation challenge coding was cohen’s k = .83. wave 3 : modes of regulation. instances of self-, co-, and shared regulation enacted as a response to the motivation challenges were then identified. coding in this wave was guided by the coding scheme in bakhtiar, webster, et al. (2018). next, instances of self-, co-, and shared-regulation were grouped into a larger motivation regulation episode to identify when the regulatory actions for a challenge began and ended. it was possible to have more than one regulation mode within one motivation regulation episode because a challenge can be addressed using multiple modes at the same time. inter-rater reliability index was cohen’s k = .81. wave 4: motivation regulation strategies. in the last wave, the types of motivation regulation strategies enacted within the regulation episodes were identified. coding in this wave was guided by bakhtiar (2019) and bakhtiar, hadwin, et al. (2018) coding scheme outlining types of motivation regulation strategies in online collaborative contexts. inter-rater reliability index was cohen’s k = .85. 3.3 code revision and visualization. next, codes were reviewed, revised as necessary, and visually represented in chronological order on a cordtra diagram (e.g., figure 5). the preparation of the diagram was guided by hmelo-silver et al. (2008). however, three modifications were made to the diagram. first, instead of listing code names in the legend, each row represents a code. second, episode breaks were added to show that the episodes were not necessarily on a continuous timeline. third, horizontal lines between categories of codes were added for clarity. each group’s processes were visually represented on three levels: (a) person—showing each group member's initiation of and participation in the regulation, (b) mode—showing whether self-, co-, or shared was involved, and (c) strategy—showing the types of motivation regulation strategies applied in the episode. 3.4 summarization of data across sources. for each group, the frequencies of self-, co-, and shared-regulation instances were summed across episodes. then, the group’s percentage of co-regulation that promoted selfand shared-regulation were summed by inspecting the episodes represented on the group’s cordtra diagram. group’s conversation data were cross-checked for further context about the regulatory actions demonstrated within each episode. for example, in episode 4 in figure 5, person b can be seen co-regulating person m by using a social support strategy. the co-regulation promoted person m’s self-regulation who later regulated by expressing his emotions to person b. note that person b’s co-regulatory utterances were interleaved by person m’s self-regulatory utterances; however, this is considered one instance of co-regulation and one instance of self-regulation. the dashed lines indicate the beginning of episodes and episode numbers are marked on the top edge of the lines. codes’ sequences on the diagrams were also examined in terms of their chronological information (how early or late a code occurs), in terms of their overlap with other codes, and if a specific sequence of codes frequently occurs in multiple episodes. for instance, in figure 5, expressing emotions strategy (green diamond) precedes social support strategy (yellow diamonds). there is also a pattern of social support strategy (yellow diamonds) being used during co-regulatory utterances (pink triangles) and commonly activated by person b (blue circles). at the end of step 4 of data analysis, each group's data across all sources were synthesized and a cross-data summary for each group was prepared. figure 5. example of a cordtra diagram adapted for this study. the x-axis represents the timestamp for each conversation turn. one conversation can have multiple codes. the y-axis represents three categories of codes separated by horizontal lines: types of motivation regulation strategies (top), regulation modes (middle), and group members involved (bottom). each dashed line represents the beginning of an episode, with the corresponding episode number marked above the line. 3.5 case comparisons the last step involved a cross-case analysis of the key similarities and differences between the two groups as guided by the research questions. this analysis was recursive involving identifying emerging themes, looking for data reference that either support or contradict the themes, and revising the overall themes. 4. findings 4.1 what motivation challenges triggered regulation? the regulation of learning model used as a framework for this study (hadwin, järvelä, et al., 2018) highlighted the importance of contextualizing students’ regulation by examining the conditions that triggered regulatory actions. for this reason, groups’ motivational challenges were first examined. analysis of data revealed two areas of differences: (a) task participation level and (b) motivation challenges triggered by selfversus group conditions. 4.1.1 task participation level in both groups, there were imbalances in terms of each member’s percentage of editing work. some group members contributed to the product notably more than others (table 2). however, these imbalances were less evident in the high productivity group compared to the low productivity group. data indicated that all members in the high productivity group were involved in all collaborative sessions. in the second and third sessions, the group actively discussed their ideas and knowledge as evident by the teacher-rated collaboration score. group members' reflections indicated overall satisfaction with group members’ contribution frequency. in contrast, the low productivity group started strong; they actively collaborated in the first session and were given a full mark for active collaboration. unfortunately, some group members' participation dwindled thereafter. editing work became significantly uneven coupled with one missing member in session 2 and two missing members in session 3 (0% contribution). in their post-collaboration reflection, only person h, k, and p were acknowledged as the main contributors of the group product. table 2 summary of task participation for each group 4.1.2 motivation challenges triggered by selfversus group conditions as shown in table 3, the most prevalent motivation challenge in the high productivity group was cognitive-based (f = 9), but the most prevalent challenge in the low productivity group was behaviour-based (f = 18). specifically, the high productivity group experienced challenges mostly originating from self conditions (personal motivational beliefs and interest in the task) whereas the low productivity group experienced challenges mostly arising from group conditions (inter-individual interactions or lack thereof). the high productivity group’s motivation challenges included lacking confidence or task purpose and feeling less motivated to engage in the task because the task was viewed as cognitively challenging. members of this group also expressed negative feelings in the form of defeat and lacking task enjoyment. occasionally, the group was challenged with low participation from a couple of individuals due to the individuals’ work demands in other courses. one group member summarized in her reflection: sometimes you get a group (like this one) that doesn't seem to have a lot of passion and motivation to learn about the project or to present it in a really meaningful way. this makes it hard to make improvements and enjoy yourself. we had one of these groups, it could be because the class wasn't the most important on our schedules (person b, individual reflection data). when asked to identify the group’s salient motivation challenge, members in the high productivity group stated challenges relating to procrastinating on task completion (1 member), finding time to work together (2 members), and experiencing technological glitches (1 member). in addition, the group’s average pre-task motivation showed a small score variation between members, suggesting individuals in the group were similar in terms of their willingness to engage in the task (mean = 3.75, sd = 0.50). this similarity may be associated with the group’s lower instances of encountering conflicts related to differences in inter-individual motivational goals. in contrast, the low productivity group’s motivation challenges included issues with conflicting opinions, priorities, and goals. at times, group members appeared to be working at cross-purposes, particularly when the group failed to reach an agreement with regards to what they should be doing. one group member described in her reflection: it was very hard to contact members due to some just not replying/engaging in the group chat. our infographics template had been changed three times due to conflict over which one best suit our topic (person h, individual reflection data). also prevalent in the low productivity group were behaviour-based motivation challenges related to some individuals being off-task, not showing up to group meetings, and avoiding the task altogether. when individual judgments about the group’s salient motivation challenge were examined, at least one member mentioned difficulties dealing with differences in opinions. others noticed challenges related to getting members to complete the task and communicating with them in the online platform. moreover, upon examining the group’s pre-task motivation, individuals in this group were more variable in their responses. the group’s mean motivation level was 3.20 (sd = 1.30). these differences may have increased the likelihood of experiencing conflicts in task goals and priorities. despite differences in the types of motivation challenges, both groups had a similar number of motivation challenges (f = 23 in the high productivity group and f = 25 in the low productivity group). table 3 types and frequencies of motivation challenges in each group 4.2 how were selfand shared-regulation used in relation to co-regulation? 4.2.1 frequencies of self-, co-, and shared-regulation the raw frequencies of self-, co-, and shared-regulation in each group were first calculated (table 4). based on the proportional frequencies, the high productivity group used self-regulation the most (44.1%) followed by co-regulation (38.2%) and shared-regulation (17.6%). the higher instances of self-regulation may be due to the nature of the motivation challenges reported which often originated from self conditions. in contrast, considering the low productivity group often experienced challenges mostly arising from group conditions, one would assume the group activated shared regulation most often; however, this was not the case. the low productivity group activated co-regulation the most (51.3%) while their within-group proportions of selfand shared regulation were relatively equal (23.1% and 25.6% respectively). this group’s co-regulation can be seen as an attempt to promote regulation at both the individual and the group level, which may be necessary when individual group members were not proactively regulating on their own. 4.2.2 percent co-regulation that stimulated selfand shared-regulation the percentage of co-regulation that (a) promoted self-regulation, (b) promoted shared-regulation, and (c) did not lead to any regulation was examined (italicized in table 4). these percentages were calculated based on within-episode patterns shown on each group's cordtra diagram (figure 6 and 7). on each diagram, circles refer to co-regulation that promoted self-regulation and rectangles refer to co-regulation that promoted shared-regulation. the findings indicate that more than half of the high productivity group’s co-regulation stimulated the emergence of self-regulation (61.5%, circles in figure 6). in the same group, the percentage of co-regulation that was ignored and that promoted the emergence of shared-regulation were relatively similar (23.1% and 15.4% respectively). in contrast, co-regulation in the low productivity group was followed equally by shared regulation (40%), self-regulation (30%), no regulation (30%). however, between the two groups, the proportion of co-regulation leading to no responses was fewer in the high productivity group than the low productivity group (23% vs 30%). these non-responses were found in episode 7, 17, and 20 in the high productivity group diagram (figure 6), and in episode 2, 4, 7, 10, 21, and towards the end of 22 in the low productivity group diagram (figure 7). table 4 frequency (percentage) of regulation modes in each group note: italicized numbers refer to the breakdown of co-regulation that followed self-regulation (srl), socially shared-regulation (ssrl), or no regulation. figure 6. the high productivity group cordtra diagram. coded utterances are chronologically ordered based on their conversation turns. dotted lines represent the beginning of a motivation regulation episode and numbers on the x-axis represent the episode number. circles refer to co-regulation that promoted self-regulation, and rectangles refer to co-regulation that promoted shared-regulation. figure 7. the low productivity group cordtra diagram. coded utterances are chronologically ordered based on their conversation turns. dotted lines represent the beginning of a motivation regulation episode and numbers on the x-axis represent the episode number. circles refer to co-regulation that promoted self-regulation, and rectangles refer to co-regulation that promoted shared-regulation 4.2.3 qualitative differences in co-regulation between groups groups’ cordtra diagram demonstrates the dynamic interactions between individuals involved in groups’ episodes of co-regulation. when the overlap between co-regulation and each individual member was examined, one similarity between the groups was that they both had two dominant co-regulators taking turns in co-regulating their peers. these individuals were person b and person j in the high productivity group (figure 6), and person k and person h in the low productivity group (figure 7). however, when conversation data were cross-examined, the groups differed in the socioemotional tone accompanying the co-regulation. groups’ conversation data showed that the high productivity group’s co-regulation was often proactive and seemed to facilitate group members to work efficiently on the task. for example, when the group was tasked to read a long paper in preparation for the third collaborative session, person a proactively checked on whether her group members have completed the task (see excerpt below). she began the conversations with an expression that indicates cohesion before prompting group members to check on their progress. in response, group members openly expressed their struggles with the task which was followed by some members attempting to help. person a: hey gang! has everyone looked at the reading? person b: hey! yes i did, i jotted down some notes on the first 4 techniques too. it’s a long paper ugh! person j: very long lol person a: i was having difficulty finding it on the page. i forgot to look it over [emoticon] person j: it's ok. did u manage to find it now? person b: i think it is posted to part 3 description in contrast, the low productivity group’s co-regulation was more directive and reactive to specific actions or inactions. the directive tone of their co-regulation might have thwarted appropriation of selfand shared regulation. for instance, in the except below, person a co-regulated the group to take on a specific approach to doing the task, which involved limiting one person to edit the group’s google docs. this action may have been interpreted as person a trying to avoid responsibility, as evident in the log data showing person a’s low to non-existent involvement in the project. person a’s co-regulation may have thwarted the emergence of shared regulation in his group; when some members attempted to edit the work as a group, person a quickly brought them back to his original plan. person a: it’s easier if we let one person does it (writing out answers). k did it last time, so he can do it again this time. person k: ok. [person h attempted to help person k expand his ideas in the google docs] person a: hey. can we please stick to one person editing. it makes the computer slower for every extra person in there. overall, between the two cases, findings suggest that the number of co-regulators may not matter as much as the affective tone accompanying the co-regulation, because tone can either promote or thwart selfand/or shared-regulation. 4.3 what strategies were used during episodes of motivation regulation? the overlap between regulation modes and strategy types in groups’ cordtra diagrams revealed some differences. groups seemed to differ in terms of the variation in their strategy response and in terms of the types of strategies they enacted in a socially shared way. table 5 summarizes each groups’ motivation regulation strategies according to their associated modes of regulation. table 5 frequency of motivation regulation strategies in the form of self-, co-, and shared regulation between the high productivity and low productivity group note: one regulation episode mode may involve more than one type of strategy. frequency calculation was performed by examining the overlap between each regulation mode and each strategy type in groups’ cordtra diagrams. the most common strategy in the high productivity group was behaviour control (f = 18), followed closely by the frequency of cognitive (f = 13) and emotion control (f = 7) strategies, 2(2) = 4.84, p = .09. the group performed very few environment control strategies (f = 2) involving manipulation of physical task features such as the timing of the online session or the technology. when the micro-level strategies were examined, the high productivity group frequently engaged in social support behaviour (f = 10) encouraging or facilitating the participation of others by seeking and providing help, promoting openness, accommodating needs, and modelling specific tactics for completing the task. the most frequently observed shared regulation in this group was in the form of social support. other micro-level strategies in the high productivity group included fundamental regulatory processes such as planning and goal setting (f = 7) and checking in or monitoring (f = 6). as it was defined in the coding scheme, the planning strategy involved constructing, negotiating, or aligning task perceptions and goals; it did not include planning logistics such as figuring out when to meet. planning strategies in the high productivity group were mostly in the form of co-regulation where individuals were temporarily guided to think about task goals and purposes. there was only one shared planning strategy enacted in response to an expression of low task confidence and understanding. similarly, the high productivity group’s monitoring strategies mostly involved individuals temporarily guiding others to monitor task features or progress. one stark difference between the groups’ strategies related to the frequency of emotion control strategies: this type of strategy was quite frequent in the high productivity group but not evident at all in the low productivity group. specifically, individual members in the high productivity group self-regulated their low motivation and lack of task enjoyment by openly expressing their emotions (f = 6). during one episode involving emotion expression, another member in the group attempted to co-regulate that peer's feeling by reassuring her the group had it under control. interestingly, two members in the high productivity group had planned to use emotion control strategies for motivation hurdles they anticipated at the beginning of the collaboration: person r planned to "keep a positive attitude," and person j planned to “make the task interesting.” in comparison, the low productivity group activated behaviour control strategies (f = 24) most frequently compared to cognitive (f = 13) and environment control (f = 6) strategies, 2(2) = 11.61, p = .01. the low productivity group’s behaviour control strategies mostly involved checking-in or behaviour-monitoring strategies. most of the check-ins were co-regulatory involving directing specific individuals to the task because of a lack of contribution from those individuals. under the cognitive control category, the low productivity group’s most common strategy was information processing that was activated using the co-regulation mode. the information processing strategy involved guiding others to engage with the learning materials to gain a better understanding of the domain knowledge necessary for completing the task efficiently. person k who was dominant in performing the co-regulation stated in his planner that he planned to “demonstrate to others how to work on the task and would like to work on his leadership skills.” on other occasions when group members' engagement and attention derailed, the low productivity group focused on managing the environmental or logistical aspects of the task as a team (shared regulation) rather than addressing the task goals together. planning and setting goals were not as frequent in the low productivity group (f = 3) compared to the high productivity group (f = 6), and none of the planning was performed in a shared way. given the lack of joint involvement in constructing task goals and understanding, it was not surprising to find the low productivity group revising their goals and plans more frequently (f = 3) compared to the high productivity group (f =7). overall, except for managing the external environment, the high productivity group tended to use more varied types of motivation regulation strategies focusing on managing behaviour, thoughts, and emotions. the high productivity group demonstrated social support in response to motivation challenges but also frequently engaged in fundamental regulatory processes such as planning and monitoring. in contrast, the low productivity group mostly focused on managing behaviour and thoughts. the behaviour control strategy the low productivity group used was highly related to conducting check-ins on others’ behaviour. the same group engaged in less frequent strategic planning or goal setting but had to revise their previously unshared goals a couple of times along the way. 5. discussion this study aimed to examine the role co-regulation played in dynamically stimulating the emergence of selfand shared-regulation motivation in two groups with contrasting perceptions of their overall team learning productivity. this study demonstrated that the cordtra diagram was suitable for representing collaborative groups’ regulatory actions because the intertwined processes between individuals and the group can be simultaneously considered. findings also demonstrated that co-regulation of motivation may afford and thwart the emergence of group members’ selfand shared-regulation of motivation, and these processes interacted with the group’s situational challenges and the regulatory skills or strategies group members possessed. specifically, case comparisons indicated that groups’ motivation regulation should (a) match the demands of the challenges at hand, (b) be positively supported by group members through co-regulation, and (b) involve a more varied strategic responses so that the group may continue to learn and co-construct knowledge effectively as a team. 5.1 match between situational demands (conditions) and modes of regulation examining the context that triggered regulation matters. rather than focusing on the frequency of self-, co-, and shared regulation, we examined the modes of regulation in tandem with the motivation challenges (conditions) that stimulated responses. if the conditions were ignored when groups’ proportions of selfand shared-regulation were compared, we would have concluded that the low productivity group enacted more shared regulation than self-regulation and the high productivity group enacted more self-regulation than shared regulation. this conclusion would contradict previous research demonstrating an association between higher instances of shared regulation and better collaboration outcomes (see panadero & järvelä, 2015). however, when the types of motivation challenges that triggered regulation were examined, we found the high productivity group experienced more challenges originating from self conditions, such as personally lacking task enjoyment. individuals within this group did not necessarily share the same motivation challenges. hence, a high proportion of regulation in the high productivity group involved individuals self-regulating their own motivation and supporting or co-regulating the motivation of the struggling group members. in contrast, the low productivity group experienced more motivation challenges originating from group conditions relating to inter-individual conflicts and a lack of group interactions. the group, however, failed to activate higher instances of shared regulation despite some members’ frequent attempts to promote regulation through co-regulation. some instances of co-regulation thwarted the formation of shared-regulation, particularly when the co-regulator did not consider the other group members' perspectives. also, rather than taking personal responsibility to proactively self-regulate in the face of motivation challenges, a few group members in the low productivity group seemed to rely on others to regulate for them. this finding supports the theory that self-regulation is a necessary ingredient in collaboration; when individuals are not proactively self-regulating, coand shared-regulation would be less relevant or effective. the focus of shared motivation regulation in both groups also seemed to differ. shared motivation regulation in the high productivity group targeted the wellbeing of the group, and group members focused on creating a supportive environment by exchanging feedback, negotiating needs, and jointly finding solutions to manage their challenges. in contrast, the low productivity group's shared motivation regulation tended to focus on environmental-based motivation challenges; the group focused on changing the physical conditions by finding a different meeting time where everyone could be more engaged. this finding is in line with previous research that found low-performing groups often focused on controlling the external challenges such as difficulties navigating the online collaborative environment and using technologies (e.g., malmberg et al., 2015). in contrast, the same study found that high-performing groups were more active in managing the cognitive, motivational, as well as social aspects of their collaboration—similar to findings related to the high productivity group in this study. the differences in motivation challenges experienced in both groups call for a more refined conceptualization of motivation, especially as it is experienced in collaborative contexts. in this study, motivation challenges were conceptualized as circumstances in which individuals’ or the group’s level of willingness to engage in the task was compromised. factors influencing motivation are diverse and not limited to traditionally discussed motivational beliefs such as self-efficacy. motivation challenges may originate from behavioural (e.g., effort initiation and maintenance), cognitive (e.g., efficacy or task purpose), emotional (e.g., boredom), and environmental (e.g., technological support) difficulties experienced before and during the task (bakhtiar, hadwin, et al., 2018). accordingly, the way learners respond to different challenges may also vary as not all motivation challenges require the same type of strategic response. we argue that understanding the nuances around groups’ reactions to different challenges can help researchers and educators provide a more targeted support for students' motivation when learning in a team. more research needs to consider the contexts that triggered learners’ regulation, rather than focusing on examining modes of regulation (frequency or sequence) in isolation from the situational demands that triggered it. 5.2 supporting group members’ regulation through co-regulation co-regulation, as conceptualized by hadwin, järvelä, et al. (2018), refers to affordances and constraints stimulating appropriation of strategic planning, enactment, reflection, and adaptation by individuals or the group. therefore, selfand shared-regulation should be examined with the co-regulatory affordances and constraints within the learning environment. findings in this study demonstrate that co-regulation of motivation may mediate the productivity of selfand shared-regulation. between the two groups, the low productivity group exhibited more co-regulation that was directive than facilitative (see also bakhtiar, webster, et al., 2018; rogat & adam-wiggins, 2015). likely because of the directive regulation, group members were more reluctant to respond and productive selfor shared-regulation were less likely to transpire. this study’s findings also suggest that whether the co-regulatory role is distributed amongst all members or dominant in specific team members had less of an effect on team learning. instead, team learning seemed to be negatively influenced when the co-regulation was communicated in a directive and an undesirable way. future research with a larger pool of participants is needed to examine which of these factors (distribution of co-regulatory roles or socioemotional tone) have more influence on group learning and regulation. currently, there is a great interest in examining the extent to which groups' regulation is shared because “sharedness” is indicative of group cohesion and so may influence task performance (iiskala, vauras, lehtinen, & salonen, 2011). panadero and järvelä (2015) in their review also recommended future research to examine the best conditions that promote socially shared regulation in collaborative learning. before pursuing such investigation, one missing link must be addressed: similar to the concept of group cohesion, limited research has attempted to articulate how a group transitions from a collection of individuals (srl) to acting as a collective entity (ssrl). one suggested possibility is that active self-regulation across group members must simultaneously be observed. when all group members are metacognitively aware of the group’s needs and challenges and are invested in taking control of the situations, shared-regulation is more likely to emerge (järvelä et al., 2015). otherwise, shared-regulation is possible when some individuals activate co-regulation by bringing the group’s attention to the needs and challenges that needed to be regulated. co-regulation may be a necessary metacognitive process for temporarily supporting (sometimes constraining) groups’ shared regulation. hence, future research should investigate the types of co-regulation that afford or constrain the activation of shared regulation. similarly, research is needed to examine the types of co-regulation that facilitate the development of self-regulatory competence for collaborative learning (e.g., didonato, 2013). 5.3 variation in motivation regulation strategies findings also indicate that the high productivity group exhibited more varied responses to motivation challenges, including behaviour-, cognitive-, and emotion-control strategies. the variation in strategic responses observed in this group may point to the importance of groups (a) being more flexible to readily adapt when one strategy does not work, and (b) regulating across cognitive, behavioural, and affective dimensions of regulation rather than focusing on one aspect alone (see rogat & linnenbrink-garcia, 2011). in contrast, the low productivity group demonstrated a limited strategy repertoire, which is a common finding in novice groups (see bakhtiar, hadwin, et al., 2018). the low productivity group tended to use strategies focused on correcting individual members’ behaviour in the task; none of their strategies was related to improving the group’s emotional experiences. the group tended to focus on increasing members’ contributions and getting the task done without considering individual members’ thoughts, beliefs, and feelings. the lack of emotion control strategies may have contributed to experiencing less productive team learning. previous research demonstrated that emotion control strategies in the form of openly expressing negative and positive emotions could help build psychological trust amongst team members and that trust is a strong predictor of group members’ ability to co-construct knowledge and ideas in the task (van den bossche et al., 2006). moreover, research argues that planning is critical in setting the stage for more effective learning (hadwin, bakhtiar, et al., 2018). during planning, learners generate goals and standards, making it easier to monitor task progress. in response to motivation challenges, the groups in this study engaged in planning and goal setting at different frequencies. planning (either done alone, for others or together as a team) was one of the most common strategies in the high productivity group. the low productivity group engaged in planning at half the frequency of the other group and none of the planning related to developing shared task goals. thus, it was not surprising to find the low productivity group working at cross-purposes and having to revise their goals and plans a couple of times during their collaboration. the link between group motivation and planning in the form of figuring out task goals should be explicitly examined in future studies involving a larger sample size of collaborating groups. 5.4 limitations one limitation of this study concerns the possible underestimation of self-regulation of motivation. some strategies may be internal to an individual and so may not have been observed during the group’s interactions. xu and du (2013) found that a high proportion of the variation in students' motivation regulation strategies was at the individual level, suggesting minimal group influences. strategies such as telling myself that i can successfully attain a set goal (i.e., promotion of goal striving or goal self-talk) are difficult if not impossible to observe unless the individuals themselves explicitly mention using the strategy. unlike the strategies found in previous research (bakhtiar, hadwin, et al., 2018; xu & du, 2013), several strategies were not observed in these two cases. the strategies include enhancing task interest, changing one's thought about the value or utility of the task, and administering rewards for accomplishing a goal. such strategies may have been covertly enacted by individuals and so were not evident during groups' conversations. data in this study might underestimate individual motivation challenges, particularly when these were not shared in conversations. the collaborative project was scripted to involve eight work sessions with accompanying cognitive and metacognitive prompts that guided students’ collaboration. the instructional design of the collaboration may have alleviated some group coordination challenges and other more intense motivation challenges (see hadwin, bakhtiar, et al., 2018). as a result, the level of motivation challenges of the groups in this sample experienced may be lower than groups with no such guided supports. the effect of using the cognitive and metacognitive prompts on groups’ overall team learning productivity could not be examined in this study. however, this study demonstrates the types of prompts and guidance tools that can be designed and adapted to support students’ motivation and learning during collaboration. adopting a case study design meant the findings in terms of the differences between the group who were effective and not effective in managing the motivation challenges they experienced during collaboration may not be evident in larger sample of groups. as described in the group assignment approach, the groups in this study were made relatively equal in terms of their levels of task preparation and understanding of the task demands. while the differences between the two cases may not be due to their levels of task preparation, it may be possible that the differences were due to other factors beyond our control such as specific group members’ previous experience collaborating or overall self-regulation skills acquired prior to the study (see panadero et al., 2015). in addition, new to research in motivation regulation, groups’ dynamic regulatory responses to motivation challenges were visually represented using cordtra diagrams. the diagrams allowed us to holistically examine how the strategic actions among group members go on and off, and how selfand shared-regulation of motivation were used in relation to co-regulation. as seen in the diagrams, we selected groups' motivation regulation episodes and ordered them chronologically while ignoring activities that were not considered an episode of motivation regulation. this approach may suggest that one episode of motivation regulation informs the next one, but it may be possible that the activities in the motivation regulation episodes were influenced by the activities that were not considered motivational and so were not represented on the diagram (see järvenoja, näykki & törmänen, 2019; rogat & linnenbrink-garcia, 2011; ucan & webb, 2015). future research may attempt to construct cordtra diagrams with a more continuous timeline, showing groups' regulation across different areas and with non-regulatory activities. 5.5 implications by analyzing the dynamic interplay between selfand shared-regulation as supported or thwarted by co-regulation, the current study contributes to uncovering how individual and group motivation regulation processes evolve in collaboration. findings indicated that shared-regulation was limited and superficial when active self-regulation across group members was infrequently observed. even when individuals attempted to co-regulate towards shared-regulation, such efforts were often unhelpful if the individuals themselves were not ready to play an active role in regulating and if the co-regulator ignored other group members' ideas and contributions in opposition to their own. hence, supports for motivation in collaboration need to consider two elements. first, individual group members may need to be supported to engage their metacognition, beginning from the fundamental processes of constructing task perceptions and specific goals which are critical for directing and motivating individuals toward the task. when regulation at the individual level is productive, it will likely transfer to productive regulation at the group level. second, group members may not necessarily know how to effectively co-regulate others without projecting directive statements, particularly when conversations occur in an online environment. supports in the form of metacognitive sentence starters (e.g., morris et al., 2010) geared toward supporting motivation in online collaborations should be explored to improve groups' motivation regulation processes. finally, this study is one of limited studies to identify motivation regulation strategies students conduct in the form of coand socially shared regulation, particularly as they occur during online collaboration. findings indicated that frequently adopted motivation co-regulation strategies included guiding individuals or the groups to check on their product and progress, responding to members’ concerns and needs, and reminding others about a plan or a goal. strategies that were performed in a shared way included collectively supporting each other’s task engagement and modifying the environmental features of the task together. this set of findings provides a basis for re-operationalizing strategies relevant in collaborative contexts, rather than relying on motivation research that tended to focus on solo learning activities. as it was demonstrated, individual and social regulatory processes dynamically interact as groups move through a task, consequently, influence the expressions of motivation regulation strategies in such contexts. keypoints co-regulation of motivation may afford and thwart the emergence of group members’ selfand shared-regulation of motivation, and these processes interacted with groups’ situational challenges and regulatory skills. the cordtra diagram as a visualization tool was suitable for representing collaborative groups’ regulatory actions because the intertwined processes between individuals and the group can be simultaneously considered. compared to the low productivity group, individuals in the high productivity group (a) were actively self-regulating their motivation, (b) were positively co-regulating their group members’ motivation, and (a) demonstrated more varied types of strategies in response to motivation challenges. acknowledgements support for this research was provided by an insight grant for research to hadwin, a.f., & winne, p.h. from the social sciences and humanities research council of canada (435-2012-0529). we thank hakase hayashida for writing the phyton code to generate the cordtra diagrams in this paper. references bakhtiar, a. (2019). regulating self, other’s, and group motivation in online collaboration (doctoral thesis, university of victoria, canada). retrieved from https://dspace.library.uvic.ca/handle/1828/11354. bakhtiar, a., hadwin, a.f., & järvenoja, h. (2018). contextual differences of students’ motivation regulation strategies in a collaborative project . paper presented at the bi-annual international conference on motivation, aarhus, denmark. bakhtiar, a., webster, e. a., & hadwin, a. f. (2018). regulation and socio-emotional interactions in a positive and a negative group climate. metacognition and learning, 1–34. https://doi.org/10.1007/s11409-017-9178-x barron, b. (2003). when smart groups fail. journal of the learning sciences, 12(3), 307–359. https://doi.org/10.1207/s15327809jls1203_1 belland, b. r., kim, c., & hannafin, m. j. (2013). a framework for designing scaffolds that improve motivation and cognition. educational psychologist, 48(4), 243-270. https://doi.org/10.1080/00461520.2013.838920 blumenfeld, p. c., kempler, t. m., & krajcik, j. s. (2006). motivation and cognitive engagement in learning environments. in r. k. sawyer (ed.), the cambridge handbook of: the learning sciences (pp. 475-488). new york, us: cambridge university press. didonato, n. c. (2013). effective self-and co-regulation in collaborative learning groups: an analysis of how students regulate problem solving of authentic interdisciplinary tasks. instructional science, 41(1), 25-47. https://doi.org/10.1007/s11251-012-9206-9 hadwin, a. f., bakhtiar, a., & miller, m. (2018). challenges in online collaboration: effects of scripting shared task perceptions. international journal of computer-supported collaborative learning , 13(3), 301-329. https://doi.org/10.1007/s11412-018-9279-9 hadwin, a. f., järvelä, s., & miller, m. (2011). self-regulated, co-regulated, and socially shared regulation of learning. in b. j. zimmerman, & d. h. schunk (eds.), handbook of self-regulation of learning and performance (pp. 65-84). new york, ny: routledge. https://doi.org/10.4324/9780203839010 hadwin, a. f., järvelä, s., & miller, m. (2018). self-regulation, co-regulation and shared regulation in collaborative learning environments. in d. schunk & j. greene (eds.). handbook of self-regulation of learning and performance 2nd edition (pp. 83-106). new york, ny: routledge. https://doi.org/10.4324/9781315697048 hadwin, a. f., & winne, p. h. (2012). promoting learning skills in undergraduate students. enhancing the quality of learning: dispositions, instruction, and mental structures , 201-227. https://doi.org/10.1017/cbo9781139048224.013 hmelo-silver, c. e., chernobilsky, e., & jordan, r. (2008). understanding collaborative learning processes in new learning environments. instructional science, 36(5-6), 409-430. https://doi.org/10.1007/s11251-008-9063-8 iiskala, t., vauras, m., lehtinen, e., & salonen, p. (2011). socially shared metacognition of dyads of pupils in collaborative mathematical problem-solving processes. learning and instruction, 21 (3), 379-393. https://doi.org/10.1016/j.learninstruc.2010.05.002 isohätälä, j., järvenoja, h., & järvelä, s. (2017). socially shared regulation of learning and participation in social interaction in collaborative learning. international journal of educational research, 81, 11–24. https://doi.org/10.1016/j.ijer.2016.10.006 järvelä, s., & hadwin, a. f. (2013). educational psychologist new frontiers: regulating learning in cscl. educational psychologist, 48(1), 25–39. https://doi.org/10.1080/00461520.2012.748006 järvelä, s., järvenoja, h., & veermans, m. (2008). understanding the dynamics of motivation in socially shared learning. international journal of educational research, 47(2), 122–135. https://doi.org/10.1016/j.ijer.2007.11.012 järvelä, s., kirschner, p. a., panadero, e., malmberg, j., phielix, c., jaspers, j., ... & järvenoja, h. (2015). enhancing socially shared regulation in collaborative learning groups: designing for cscl regulation tools. educational technology research and development, 63(1), 125-142. https://doi.org/10.1007/s11423-014-9358-1 järvelä, s., malmberg, j., & koivuniemi, m. (2016). recognizing socially shared regulation by using the temporal sequences of online chat and logs in cscl. learning and instruction, 42, 1–11. https://doi.org/10.1016/j.learninstruc.2015.10.006 järvelä, s., volet, s., & järvenoja, h. (2010). research on motivation in collaborative learning: moving beyond the cognitive-situative divide and combining individual and social processes. educational psychologist, 45(1), 15–27. https://doi.org/10.1080/00461520903433539 järvenoja, h., järvelä, s., & malmberg, j. (2015). understanding regulated learning in situative and contextual frameworks. educational psychologist, 50(3), 204–219. https://doi.org/10.1080/00461520.2015.1075400 järvenoja, h., järvelä, s., & malmberg, j. (2017). supporting groups’ emotion and motivation regulation during collaborative learning. learning and instruction (online). https://doi.org/10.1016/j.learninstruc.2017.11.004 järvenoja, h., järvelä, s., törmänen, t., näykki, p., malmberg, j., kurki, k., ... & isohätälä, j. (2018). capturing motivation and emotion regulation during a learning process. frontline learning research, 6(3), 85-104. https://doi.org/10.14786/flr.v6i3.369 järvenoja, h., näykki, p., & törmänen, t. (2019). emotional regulation in collaborative learning: when do higher education students activate group level regulation in the face of challenges?. studies in higher education, 44(10), 1747-1757. https://doi.org/10.1080/03075079.2019.1665318 järvenoja, h., volet, s., & järvelä, s. (2013). regulation of emotions in socially challenging learning situations: an instrument to measure the adaptive and social nature of the regulation process. educational psychology, 33(1). https://doi.org/10.1080/01443410.2012.742334 malmberg, j., järvelä, s., järvenoja, h., & panadero, e. (2015). promoting socially shared regulation of learning in cscl: progress of socially shared regulation among highand low-performing groups. computers in human behaviour, 52, 562–572. https://doi.org/10.1016/j.chb.2015.03.082 mccaslin, m., & hickey, d. t. (2001). educational psychology, social constructivism, and educational practice: a case of emergent identity. educational psychologist, 36(2), 133-140. https://doi.org/10.1207/s15326985ep3602_8 mccaslin, m., & vriesema, c. c. (2018). co-regulation: a model for classroom research invygotskianian perspective. in g. a. d. liem, & d. m. mcinerney, (eds.). big theories revisited 2 (pp. 154-180). charlotte, nc: iap. mchugh m. l. (2012). interrater reliability: the kappa statistic. biochemia medica, 22(3), 276-82. https://doi.org/10.11613/bm.2012.031 näykki, p., isohätälä, j., järvelä, s., pöysä-tarhonen, j., & häkkinen, p. (2017). facilitating socio-cognitive and socio-emotional monitoring in collaborative learning with a regulation macro script–an exploratory study. international journal of computer-supported collaborative learning , 12(3), 251-279. https://doi.org/10.1007/s11412-017-9259-5 panadero, e., & järvelä, s. (2015). socially shared regulation of learning: a review. european psychologist. https://doi.org/10.1027/1016-9040/a000226 panadero, e., kirschner, p. a., järvelä, s., malmberg, j., & jarvenoja, h. (2015). how individual self-regulation affects group regulation and performance: a shared regulation intervention. small group research, 46(4), 431–454. https://doi.org/10.1177/1046496415591219 pino-pasternak, d., whitebread, d., & neale, d. (2018). the role of regulatory, social, and dialogic dynamics on young children’s productive collaboration in group problem solving. new directions for child and adolescent development, 162, 1–26. https://doi.org/10.1002/cad.20262 puntambekar, s., erkens, g., & hmelo-silver, c. (eds.). (2011). analyzing interactions in cscl: methods, approaches and issues. new york, ny: springer. https://doi.org/10.1007/978-1-4419-7710-6 rogat, t. k., & adams-wiggins, k. r. (2015). interrelation between regulatory and socioemotional processes within collaborative groups characterized by facilitative and directive other-regulation. computers in human behaviour, 52, 589–600. https://doi.org/10.1016/j.chb.2015.01.026 rogat, t. k., & linnenbrink-garcia, l. (2011). socially shared regulation in collaborative groups: an analysis of the interplay between quality of social regulation and group processes. cognition and instruction, 29(4), 375–415. https://doi.org/10.1080/07370008.2011.607930 rogat, t. k., linnenbrink-garcia, l., & didonato, n. (2013). motivation in collaborative groups. in c. hmelo-silver, c. chinn, c. chan, & a. o’donnell (eds.), international handbook of collaborative learning (pp. 250–267). new york, ny: routledge. https://doi.org/10.4324/9780203837290 ruiz ulloa, b. c., & adams, s. g. (2004). attitude toward teamwork and effective teaming. team performance management, 10(7/8), 145-151. https://doi.org/10.1108/13527590410569869 schoor, c., & bannert, m. (2012). exploring regulatory processes during a computer-supported collaborative learning task using process mining. computers in human behaviour, 28(4), 1321–1331. https://doi.org/10.1016/j.chb.2012.02.016 seale, c., & silverman, d. (1997). ensuring rigour in qualitative research. the european journal of public health, 7(4), 379-384. https://doi.org/10.1093/eurpub/7.4.379 sobocinski, m., malmberg, j., & järvelä, s. (2017). exploring temporal sequences of regulatory phases and associated interactions in lowand high-challenge collaborative learning sessions. metacognition and learning, 12(2), 275–294. https://doi.org/10.1007/s11409-016-9167-5 tucker, r., & abbasi, n. (2016). bad attitudes: why design students dislike teamwork. journal of learning design, 9(1), 1–20. https://doi.org/10.5204/jld.v9i1.227 ucan, s., & webb, m. (2015). social regulation of learning during collaborative inquiry learning in science: how does it emerge and what are its functions?. international journal of science education, 37(15), 2503-2532. https://doi.org/10.1080/09500693.2015.1083634 van den bossche, p., gijselaers, w. h., segers, m., & kirschner, p. a. (2006). social and cognitive factors driving teamwork in collaborative learning environments team learning beliefs and behaviours. small group research, 37(5), 490–521. https://doi.org/10.1177/1046496406292938 volet, s., & vauras, m. (2013). interpersonal regulation of learning and motivation: methodological advances . new york, ny: routledge. https://doi.org/10.4324/9780203117736 wijga, m., endedijk, m. d., & veldkamp, b. p. (2019). social regulation at the workplace: different modes of regulation and variation in quality. in a. bakhtiar and m. wijga (organisers), self-, co-, and shared regulation: what do they look like in different contexts? why do they matter? symposium conducted at the bi-annual meeting of the european association for research on learning and instruction, aachen, germany. winne, p. h., & hadwin, a. f. (1998). studying as self-regulated engagement in learning. in d. hacker, j. dunlosky, & a. graesser (eds.), metacognition in educational theory and practice (pp. 277-304). hillsdale, nj: lawrence erlbaum. https://doi.org/10.4324/9781410602350 winne, p. h., & hadwin, a. f. (2008). the weave of motivation and self-regulated learning. in d. h. schunk & b. j. zimmerman (eds.), motivation and self-regulated learning: theory, research and applications (pp. 298-314). new york, ny: lawrence erlbaum. https://doi.org/10.4324/9780203831076 winne, p. h., & marx, r. w. (1989). a cognitive-processing analysis of motivation within classroom tasks. research on motivation in education, 3, 223-257. wosnitza, m., & volet, s. (2014). trajectories of change in university students’ general views of group work following one single group assignment: significance of instructional context and multidimensional aspects of experience. european journal of educational psychology, 29, 101–115. https://doi.org/10.1007/s10212-013-0189-y xu, j., & du, j. (2013). regulation of motivation: students’ motivation management in online collaborative groupwork. teachers college record, 115(10), 1–16. yin, r. k. (2003). case study research: design and methods (3rd edition). california, la: sage publication. appendix a coding scheme used in wave 1 to wave 3 1 järvenoja, järvelä, and malmberg (2017) 2 bakhtiar, hadwin, and järvenoja (2018) 3 bakhtiar, webster, and hadwin (2018) appendix b coding scheme used in wave 4: motivation regulation strategy codepen draijer frontline learning research vol.8 no. 4 (2020) 18 36 issn 2295-3159 the multidimensional structure of interest jael draijera, arthur bakkera, esther slota & sanne akkermana autrecht university, the netherlands article received 21 october 2019/ revised 7 may 2020/ accepted 17 may / available online 15 june abstract there is increasing attention for interest as a powerful, complex, and integrative construct, ranging in appearance from entirely momentary states of interest to longer-term interest pursuits. developmental models have shown how these situational interests can develop into individual interests over time. as such, these models have helped to integrate more or less separate research traditions and focus the attention of the field more on the developmental dynamics. this, however, also raises subsequent questions, one being how development can be understood in terms of interest structure. the developmental models seem to suggest that development occurs roughly along the line of six dimensions, which we summarize as the dimensions of historicity, value, agency, frequency, intensity, and mastery. using an experience sampling method that was implemented in a smartphone application, we prompted 94 adolescents aged 13 to 16 (60% female) to rate each interest they experienced during two weeks on these six dimensions. a latent profile analysis on 1247 interests showed six distinct multidimensional patterns, indicating both a homogeneous and heterogeneous structure of interest. four homogeneous patterns were indicated by more or less equal levels on all six dimensions in varying degrees, and contained 86% of the interests. two heterogeneous patterns were found, describing variations of interest that are interpreted and discussed. these results endorse the complexity of the construct of interest and provide suggestions for identifying different manifestations of interest. keywords: interest; situational interest; individual interest; structure of interest; latent profile analysis info corresponding author email: j.m.draijer@uu.nl doi: https://doi.org/10.14786/flr.v8i4.577 1. introduction interest has for long been seen as a powerful basis for learning and good education. more than a century ago, interest was already seen as a guarantee for effortless attention (arnold, 1906; dewey, 1913) and research over the years has shown that interest goes together with motivation to further engage in certain topics or activities, relates strongly to academic achievement across subject areas, school types, and age groups (ainley, hidi, & berndorff, 2002; harackiewicz, durik, barron, linnenbrink-garcia, & tauer, 2008; hidi & renninger, 2016; schiefele, krapp, & winteler, 1992) and to more sustainable curriculum choices and career decisions (harackiewicz, barron, tauer, & elliot, 2002; harackiewicz et al., 2008; köller, baumert, & schnabel, 2001). currently, interest is appreciated particularly because it generates a learning process that appears natural and intrinsic to the person, and therefore fits with new models of education that depart from person-centered, life-wide and connected learning approaches (e.g., barron, 2006; ito et al., 2018; walkington, 2013). this new appreciation has sparked research focusing on how interests may be evoked (e.g., renninger, bachrach, & hidi, 2019) and sustained in education (harackiewicz et al., 2008; hedges, 2019). additionally, recent research has shown that also interests outside the school context can be catalysts for meaningful learning processes, with the potential to develop considerable knowledge and skills (ito et al., 2018; krapp, 2002; renninger & hidi, 2016). moreover, out-of-school interests can be of equal value to students in making study and career choices, an interest in gaming for example leading a student to consider choosing computer science (erstad & silseth, 2019; holmegaard, 2015, vulperhorst, van der rijst, & akkerman, 2020). recognition of the powerful role of out-of-school interests has also generated educational research on how to involve students’ existing interests into education (e.g., hinton & kern, 1999; reber, canning, & harackiewicz, 2018). besides being powerful, interest is also a complex construct; it integrates cognition, motivation and affect (hidi & renninger, 2006, renninger & leckrone, 1991; sachisthal et al., 2018; schiefele & rheinberg, 1997) in a person’s perception of and preference to engage with a specific object of interest (arnold, 1906; krapp, 2002). moreover, key scholars have proposed that interests may manifest in different ways, ranging from a relatively fleeting person-object relation, referred to as situational interest, to a long-lasting predisposition towards specific person-object relations, referred to as individual interest (krapp, hidi, & renninger, 1992). during the last two decades scholars have attempted to capture the spectrum between these two extreme manifestations of interest, proposing developmental models to define the shades in between (e.g., hidi & renninger, 2006; krapp, 2002). the current study aims to contribute to the understanding of this spectrum by conducting a more detailed investigation of the structure of interest in various manifestations, so as to advance our understanding of interest as well as provide a more nuanced basis for how to evoke and sustain interests in educational settings and materials. in the following we first describe how key scholars have theorized different manifestations of interests and what appears as underlying dimensions of this development. 1.1 manifestations of interest the conceptualization of interest shows a long tradition with increasing nuance in how interest can be manifested (krapp, 2002). the two aforementioned manifestations of interest have long been regarded and studied as independent constructs, with one tradition focusing on individual interest (e.g., dewey, 1913; strong, 1927) and one focusing on situational interest (e.g., hidi & baird, 1986). however, krapp (2002) proposed that situational interests could under the right conditions develop into individual interests, as such unifying them as two kinds of the same phenomenon. he proposed a three-phase developmental model using two phases of situational interest (catch and hold) earlier defined by mitchell (1993). hidi and renninger (2006) extended this to a four-phase model of interest development, splitting up individual interest into two phases referred to as emerging and well-developed individual interest. most distinctive in these four phases is the increasing duration and strength of interest as it progresses. the threeand four-phase models led the field to recognize the wide variation in manifestations of interest, but more detailed descriptions of these manifestations are warranted. descriptions of the four phases in the literature suggest multiple indicators of interest (e.g., value, knowledge) underlying all phases. each subsequent phase is described to be mainly characterized by higher levels of these indicators (e.g., value builds up with each subsequent phase), yet this does not seem to apply to every indicator. for instance, when describing the role of external support in interest development, hidi and renninger (2006, 2011) emphasize that both emerging and well-developed individual interests are “typically but not exclusively self-generated” (2006, p. 115) and might still benefit from external support. these various descriptions are relevant as they suggest there might be more complexity in the dimensional structure of interest: high levels on one dimension may not necessarily indicate high levels on other dimensions. in the current study we aim to clarify this multidimensionality of the construct of interest. first, we look at the literature of abovementioned key scholars in developmental interest theories (i.e. suzanne hidi, ann renninger, and andreas krapp) to identify the dimensions that seem key in describing different manifestations of interests. these dimensions are then used to measure a diverse set of interests and the relations between the dimensions are examined. 1.2 dimensions of interest to identify key dimensions of interests we searched for indicators that are suggested to vary across and distinguish between interests. some indicators of interest, like focused attention, are described to be always present to a similar degree (dewey, 1913; krapp et al., 1992; schiefele, 2009) and are therefore not key to distinguish manifestations of different interests. using the work of aforementioned key scholars, we arrived at six dimensions which can be considered continuous dimensions by means of which to differentiate diverse manifestations of interests. 1.2.1 historicity the first of these dimensions is the time a person has engaged in an interest. triggered situational interests are short-term (hidi & renninger, 2006), whereas individual interests are assumed to be relatively enduring (krapp et al., 1992) and can persist over the course of years. the term historicity is used as a counterpart of novelty and signifies the quality of being historical or long-term and thus already meaningful to a person in some way (bruner, 1990). the historicity of an interest can therefore be seen as a dimension underlying interest, which can be used to distinguish between interests in different manifestations. 1.2.2 agency the degree to which an interest is internally or externally triggered and maintained is a second characteristic that can be used to distinguish between different manifestations of interest. whereas a situational interest is considered to be largely externally triggered and maintained, an individual interest is regarded as mainly internally managed and pursued by the person (krapp et al., 1992). several terms have been employed to refer to this aspect of internal management, for example external support needed as opposed to self-generated (hidi & renninger, 2006), voluntary and independent engagement (renninger & hidi, 2016). in the current study we employ the term agency to denote the extent to which a person is the agent of their own interest, both in triggering and pursuing the interest. 1.2.3 value high value attributed to the object of interest is often taken as an indicator of an individual interest (hidi & renninger, 2006; krapp, 1999). value refers to the personal significance of the object of interest because of various possible reasons, for example because it is considered relevant to a person’s enjoyment, development or future goals (krapp, 2002). hidi and renninger (2006) have used the concept of stored value which builds up over time as an interest moves from situational to individual phases. 1.2.4 frequency an additional indicator of the distinction between different interests is the frequency of engagement. it is assumed that a well-developed individual interest is more frequently engaged with than a situational interest (renninger & hidi, 2016), as persons with an individual interest actively seek repeated reengagement with this interest (hidi & renninger, 2006). as engagement in a situational interest is externally triggered, it is assumed that the frequency of engagement for these interests is lower. 1.2.5 intensity even though every interest is characterized by focused attention, engagement in some interests is more intense than engagement in others (renninger & hidi, 2016). especially with well-developed interests, engagement is characterized as more intense concentration on what one is doing, a loss of self-awareness and distortion of the perception of time, which is also called flow (nakamura & csikszentmihalyi, 2014; renninger & hidi, 2016). as flow is often regarded as a specific state that requires high perception of skills and challenge only, we refer here to intensity of engagement as a continuous dimension of interest. 1.2.6 mastery finally, interest is closely associated to gaining knowledge or expertise about the interest related contents and activities, in other words gaining mastery of the object. key scholars have taken high levels or depth of knowledge as a characteristic of individual interest (hidi & renninger, 2006; renninger, 2000; renninger & su, 2019), whereas other scholars have stressed the possibility of combinations of high interest (as indicated by high attributed value) and low knowledge (alexander et al., 1994; tobias, 1994). in the current study we chose to include mastery as a potential dimension of interest, allowing to explore how it relates to other dimensions in describing the structure of interest. 1.3 current study following from the above, historicity, value, agency, frequency of engagement, intensity, and mastery can all be seen as distinctive, yet continuous dimensions by means of which to differentiate interests with diverse dimensional structures. every interest can be considered as positioned somewhere on these dimensions at a particular moment in time. although consistent with descriptions and developmental models of situational and individual interests (hidi & renninger, 2006; krapp, 2002), previous literature has not given insight into the way these dimensions are associated, thereby implying that these dimensions manifest themselves homogeneously. this would mean that when measuring the six dimensions for various interests, the dimensions are always related in roughly the same way (high value also means high mastery and high flow, etc.). a more heterogeneous structure of interest would be indicated by distinct patterns of association between the dimensions for different interests, for example by a cluster of interests that show low mastery but high value, whereas another cluster of interests perhaps shows reversed relations between these dimensions. in the current study we aim to clarify the structure of interest by examining the relations between these multiple dimensions using a bottom-up approach. accordingly, the research question is: what dimensional structure underlies the construct of interest in terms of the dimensions historicity, value, agency, frequency, intensity, and mastery? insight into the structure of interest can aid methodological, theoretical and practical understanding of this complex construct. if the structure of interest points towards homogeneity, it can be measured efficiently by focusing on a limited set of dimensions; if there is evidence for a multidimensional structure this will need to be reflected in interest measurements. a more detailed characterization of different manifestations of interests can aid researchers and educators in searching for ways to trigger situational interest and stimulate individual interests in school subjects, disciplines and future professions. for example, if this research provides evidence for different manifestations of novel (“situational”) interests, further research could study the different ways in which these manifestations would develop and can be aided to sustain over time. 2. methods as the aim of the current study is to explore the construct of interest, we measured all possible daily-life interests without limiting to certain content domains. to this end the current study employs an experience sampling method (esm): a research methodology in which participants receive signals to answer questions about their situated experiences at set or randomized points in time (csikszentmihalyi, larson, & prescott, 1977). the study was approved by the ethics committee of the faculty of social and behavioural sciences of utrecht university. 2.1 participants the participants of this study were 94 high school students from four different dutch schools, aged 13 to 16 years old (m = 14.53, sd = 0.45), of which 60% were girls. the schools were recruited through advertisements in educational journals for practitioners. within each school all ninth graders were invited to participate, and 145 students volunteered (20% of all ninth graders). a stratified sampling strategy based on class and gender was used to select the final sample. the students were offered financial compensation (€25) for taking part in the study and their parents were asked to sign permission forms. 2.2 procedure esm is used to study participants in their natural environments through self-report and generally provides a more accurate representation of reality (csikszentmihalyi & larson, 2014). in the current study esm was implemented in a smartphone application, named intin, to study all interests that high school students experience during their daily lives (akkerman & bakker, 2012–2014; akkerman & bakker, 2019). the participants used intin for two consecutive weeks in november 2015. they received a 1.5-hour instructional briefing prior to the data collection, during which they discussed what interests are and how to use the application. additionally, they practiced with filling in the application and were stimulated to ask questions. table 1 the six semantic scales as used in intin, in dutch (italicized) and translated to english 2.3 instrument the application intin was used to record participants’ interests. at the start of the data collection period, users were asked to enter all their existing interests into the application. the participants were free to give any name to their interests, without pre-determined categories. this procedure resulted in a list of the topics and activities that the adolescents considered their interests. the dimensions underlying the construct of interest were presented to participants when they added an interest to the application. the six dimensions were formulated as semantic scales (table 1) and participants had to rate their interest on sliders with scores from 0 to 100. the starting point of each slider was at 50, participants had to touch and move each slider to be able to continue. during the two weeks of data collection, intin prompted users every two waking hours (including school hours) to report whether they had engaged with an interest in the past hours and, if so, users were asked several questions on how they engaged with this interest (this moment-to-moment data is not used in the current study). users could either select the interests in their existing list, or add a new interest to the list. if a new interest was added, users were asked to rate this interest on the six scales as well, thus resulting in a final list of their existing interests and all interests they encountered during the two-week period. to gain some insight into the interests that were reported in this study, we coded the content domains of the interests using four categories: school & work, leisure & social, maintenance, and self & world. the first category contained all interests that participants reported in the school domain, like school courses and topics, within-school projects (e.g., school musical) and jobs. leisure & social contains interests in the leisure domain, such as hobbies, sports, media, and social activities. maintenance contains interests that are considered to be part of daily routines, such as eating, cooking, and cleaning. the self & world category consists of interests in public issues, religion, and future plans (e.g., career choices). cohen’s kappa was calculated for 10% of the data (125 interests) to determine interrater reliability. results indicated strong agreement between the two coders, κ = .87, p < .001. 2.4 data analysis in total 1247 interests were added to the application by the participants, with an average of 13 interests per student. on average, each interest was engaged with three times within the two weeks of data collection. during data collection some technical difficulties occurred, which resulted in missing interests for 26 students, ranging from 1 to 11 missing interests. no systematic mechanism could be found to explain why these interests were missing and it is unlikely that the missingness was related to the structure of the interests on the dimensions. as entire interests were missing, these were regarded as a form of unit nonresponse and data imputation was not applicable. to evaluate the impact of the participants with missing interests on the results, analyses described below were run with and without these participants. the results were robust in the sense that the difference between including or excluding these participants was very small: the same profile solution was chosen as best representing the data, and the resulting profiles were almost identical in shape (the dimension-means in each profile differing by 0.14 on average) and size (maximum 2 percent difference). therefore, we decided not to exclude any participants and included everyone in the final analyses. all analyses were conducted with mplus 7.2 (muthén & muthén, 2015). since the research question of this study was concerned with estimating parameters at the interest-level and not the individual level, we chose not to conduct multi-level analyses. however, as interests were nested within individuals, all analyses used in this study took nesting and non-independence of observations into account by adjusting standard errors and chi-square tests of model fit (type=complex in mplus; muthén & muthén, 2015). an explorative latent profile analysis (lpa; lazarsfeld & henry, 1968) was conducted on the 1247 interests to assign interests to categories based on empirically distinctive patterns of scores on the six semantic scales. these categories of interests with a similar pattern are called profiles. we ran one to eight-profile models and compared these to identify which of them described the data best. if a one-profile model would suit the data best, that would indicate a homogeneous structure of interest; if one of the multi-profile models would describe the data best, this would indicate a more heterogeneous structure in which dimensions are in interplay to form distinct multidimensional patterns. the model is displayed in figure 1. the variances were restricted to be equal across profiles conform the mplus default, since our hypothesis is primarily focused on differences between different profile solutions and the shape of the final solution. in addition, assessing variance differences between groups would have made the interpretation of the results unnecessarily convoluted and testing any hypotheses regarding variances may require getting a larger sample. to determine the model that would best represent the data, we used the criteria advocated by meeus, van de schoot, keijsers, schwartz, and branje (2010), with the exception of using blrt (bootstrap lo–mendel–rubin likelihood ratio test) as this is unsuitable to models with nested data. first, a solution with k + 1 profiles should demonstrate better model fit than the previous solution of k profiles, indicated by a lower bayesian information criterion (bic) value. in the current study we also include aic and adjusted bic. second, in the chosen solution the profile separation should be reasonable and interests should be able to be assigned to a profile accurately, indicated by entropy level of minimally 0.70. third, the additional profile of a k +1 model should be a meaningful addition, and not a slight variation on one of the previously identified profiles. as recommended by marsh, lüdtke, trautwein, and morin (2009), the relation to theory, the nature of the groups and interpretation of the results are also taken into account when choosing the solution that makes most sense. using wald tests for equality of parameters on the lpa models (as recommended by asparouhov & muthén, 2007), we evaluated whether each scale contributed to the distinction between profiles. figure 1. lpa model with six continuous indicators 3. results to study the multidimensional structure of interest, we first compared results of all lpa-models. if applicable, we discuss the interest profiles of the chosen model regarding the profile patterns and the content domains of the interests in each profile. 3.1 latent profile analysis first we investigated the results of the lpa. the aic, bic and adjusted bic decrease as the number of profiles increases (see table 2). this indicates that every model with an additional profile is an improvement on the previous model. entropy values are above .70 for all solutions, which indicates good profile separation and means that interests can be assigned to profiles accurately in all models. as model fit indices were not conclusive on the number of profiles in the data, we also looked at the additional value of the new profiles and profile interpretability (marsh et al., 2009; meeus et al., 2010). in figure 2 it can be noted that from the seventh profile onwards no new theoretically meaningful profiles emerged. the seventh profile is a variation of the fourth profile: the pattern is similar with slightly higher values on the scales. therefore, it does not represent a conceptually new category of interests. the eighth profile that emerged was hard to interpret and contained only 25 interests, and was therefore not deemed a relevant pattern. because of these considerations we did not judge the seventh and eighth profile as extremely meaningful in this data and chose the six-profile model as most useful and valid in describing variations in the data at this time. the six-profile model met our criteria and showed distinctive profiles that we deemed interpretable, meaningful and interesting to further explore. we do however acknowledge that the fit indices are not conclusive and stress the importance of future research in exploring other possible variations of interest. looking at the shapes of the resulting profiles, it can be noted that the largest contrast is between the historicity dimension and the other five dimensions. as historicity cannot decrease over time for a certain interest, contrary to the other dimensions, it can be seen as the odd one out conceptually also. we therefore assessed whether the same results would be found if historicity was excluded from the analyses. the results of this analysis are displayed in the appendix a and they also suggest a profile solution with both homogeneous and heterogeneous structures. 3.2 six-profile model the six-profile model contains four flat profiles, which we named according to their position (figure 3a), and two irregular profiles which we named according to their shape (figure 3b). all semantic scales contribute to the distinction between the profiles as indicated by the significant wald statistics (table 3). as a further assessment of validity, we checked the distribution of profiles across participants. a person’s interests fall on average into 3.5 different profiles, which provides support that the multiple profiles are not subject to major individual differences. for the sake of interpretation, profile membership is fixed in the remainder of the article (meaning that each interest is assigned to a profile by the highest posterior class-membership probability). figure 2. distribution of the profiles across the six scales, for 1-profile to 8-profile solutions. table 2 log-likelihood, model fit indices and entropy measures for the 1to 8-profile models (n = 1247). figure 3. flat (a) and irregular (b) profiles in the six-profile solution. table 3 means, variances and wald test statistic of the semantic scales in the 6-profile model note. * = significant at p < .001, df = 5. 3.2.1 flat profiles the four profiles with a horizontal pattern contain interests with a homogeneous structure: these interests were given approximately equal scores on all six dimensions. the top flat profile contains approximately 17% of the reported interests in the data and is indicated by a pattern of high scores on all scales. these interests were familiar to the participants and internally driven; they were valued highly and engaged in frequently. participants reported that their engagement in these interests was intense and that they experienced mastery. the high flat profile has the same shape with slightly lower scores on all scales. this profile is the largest in the data and contains almost 40% of the interests. the middle flat profile is indicated by scores around 60 on all dimensions: these interests were rated lower than the first two profiles on the dimensions but have the same shape. the low flat profile contains interests that the participants rated very low on the six dimensions. these interests were reported to be new, not valued very highly, mostly controlled by others, not experienced very frequently, not intense and not mastered. 3.2.2 irregular profiles the two irregular shaped profiles contain interests with a heterogeneous structure. the irregular m-shaped profile (further abbreviated as m-profile) has a pattern of lower scores on self-reported historicity, frequency, and mastery, and higher scores on value, agency, and intensity. this profile thus indicates a group of interests that were quite new to participants, were not engaged in very often and were not mastered, but were regarded as quite important, intense, and internally controlled. the other distinctive pattern is the irregular w-shaped profile (further abbreviated as w-profile), which is indicated by very high scores on self-reported historicity and lower scores on the other items, especially on intensity of engagement. this profile contains interests that the participants reported they had had for a while, but were not valued very highly, not internally controlled, were not engaged in very frequently nor intensely, and were not mastered. 3.2.3 content domains to aid further interpretation of the multidimensional patterns, we categorized the interests according to content domain. this categorization serves to further describe the profiles and gain more understanding of their composition. overall, the largest proportion of interests concerned leisure & social interests (74%), followed by school & work (16%), maintenance (8%) and self & world (2%). figure 4 displays the proportions of interests within each category for every profile. every profile contains interests of every content category. the top flat profile contains mainly leisure & social interests (88%), as do the high flat profile (82%) and the mid flat profile (65%). the low flat profile mainly consists of both leisure & social (46%) and school & work (42%) interests. a small amount of the interests in the flat profiles are coded as maintenance (5–11%) and self & world (1–6%). the irregular m-profile contains a large amount of leisure & social (53%) and school & work (34%) interests. of the interests in the m-profile, 9% was coded as self & world interests. notable amongst these are interests related to the terrorist attacks in paris that took place right before the data collection in november 2015. the irregular w-profile consists mostly of leisure & social interests (75%). notable amongst these are a large proportion of interests in watching tv (30% of the interests in the w-profile). figure 4. the proportion of school & work, leisure & social, maintenance, and self & world interests within one profile. figure 5. interests of tyler in each profile. 3.3 interest portrait to further illustrate the six profiles, we discuss the interests of one of the boys in the sample. tyler is a 14-year old boy who reported interests that were representative for those in the total sample (figure 5). tyler reported no interests in the top flat profile. his interests in hanging out with friends and field hockey were assigned to the high flat profile based on the semantic scales: their levels of value, agency, frequency, intensity and mastery are approximately equal. tyler’s interest in gaming was assigned to the middle flat profile: it shows the same pattern on the dimensions as the first two interests, but with lower scores. compared to field hockey, tyler experienced gaming as less internally driven, engaged in it less often, had a less intense engagement and had low mastery-feelings about it. tyler’s interest in music was classified as a low flat interest: this interest was quite new to him and he gave it low scores on all scales. tyler reported listening to music on several occasions and playing the drums and the piano once during a music lesson. tyler’s interests in snowboarding and curriculum choices were both assigned to the irregular m-profile. both interests were quite new to him, but he valued them highly, regarded them as quite internalized, engaged in them quite frequently and intensely, and had some mastery of these subjects. his interest in curriculum choices was very relevant to him: he had to make some curriculum choices at school soon and clarified in intin that he thought these choices might determine the jobs he could apply for later. his new interest in snowboarding shows the same pattern on the dimensions. it seems that tyler had an interest in sports in general, and despite being very novel the interest in snowboarding was regarded as internally driven and valuable. tyler’s reported interests in watching television and reading a book were interests he had experienced for a while. these interests were assigned to the irregular w-profile: they were not valuable to him, he regarded them as quite externally supported, did not engage in them frequently or intensely, nor did he feel like he mastered these interests. tyler watched television several times a week but did not find particular elements interesting about the programs he watched. he read a book on several occasions, but despite finding the topic interesting he did not value reading very highly. 4. discussion the current study aimed to clarify the structure of interest by investigating the relations between six dimensions as derived from the literature. we measured interests that high school students experienced in their daily lives to answer the following research question: what dimensional structure underlies the construct of interest in terms of the dimensions historicity, value, agency, frequency, intensity, and mastery? using a latent profile analysis we found evidence for a multidimensional structure of interests, as the latent profile analysis with four flat and two irregular patterns seemed to describe the data best. the flat profiles of interests indicate a homogeneous multidimensional structure: the scores on the six dimensions were approximately equal for interests with this pattern. however, the irregular patterns suggest a heterogeneous structure of interest where high scores on some dimensions go together with low scores on other dimensions. these patterns stress the importance of a multidimensional approach when investigating experiences of interest, both in scientific endeavors and educational contexts. the homogeneous structure was most prevalent in the data. as the flat profiles in the six-profile model together contain 86% of the interests, it appears they demonstrate a typical structure of interest in which all the dimensions are rated equally high or low. for interests with this homogeneous structure, scores on the historicity-scale give an indication that high levels on the other dimensions have been established over time: a persistent interest generally has high levels on the dimensions and a new interest generally has low levels. this would also support the potential of situational interests in developing into more enduring interests, as suggested by krapp (2002) and hidi and renninger (2006). however, as we did not investigate changes in the patterns of association over time, we can only speculate about this. even though many of the interests in the data conform to a homogeneous structure on the dimensions, the irregular profiles in lpa-models show more distinctive patterns of association between the dimensions. in the six-profile model two irregular patterns were identified. first, the m-profile contained new interests with a pattern of high attributed value and experienced intensity, a relatively internal agency, and medium levels on frequency and mastery. these interests are equally new as the interests in the low flat profile, but were rated much higher on all dimensions. the irregular m-profile interests might have developed very rapidly into interests that are important to the student, or the interest was considered valuable and internally controlled when the student first engaged in it. the notion of attributing high value to a relatively new interest has been discussed by renninger (2000), though very briefly. she stated that a subset of interests consists of low knowledge and high potential value, which she calls attraction. interests with such attraction over time may develop into an individual interest. however, renninger does not elaborate on why the (potential) value of these interests is so high. an explanation for the heterogeneous structure of the interests in the m-profile may lie in the concept of identification as used by krapp (2002). even if the psychological state of interest is initially triggered by environmental elements, the interests in the m-profile might show high compatibility with the identity, values, goals, or other interests of the person, which might cause the value, agency, and intensity of this interest to be very high. for example, the interests in the terrorist attacks in paris that were assigned to the irregular m-profile might have been very relevant to the values and attitudes of these adolescents: they felt it was important to read and talk about this recent event. in the case of tyler, curriculum choices are very relevant to the goal of finding a good job in due time, which makes it plausible that he values the interest in curriculum choices very highly. tyler’s new interests are related to his existing goals, which may cause him to identify with these interests rapidly (hofer, 2010). alternatively, the new interests are connected to existing individual interests, such as tyler’s interest in snowboarding and sports in general. this explanation is in line with azevedo (2018) who states that situational interests might show continuities with previous experiences, which plays a role in their further development. the interests with a w-pattern are interests that participants have had for a long time and show a pattern of low value, agency, frequency, mastery, and especially low levels of intensity: the experience of the person is not intense and they do not get absorbed when engaging with the interest. within flow theory, the inverse of flow (intense engagement) is described as apathy or boredom, the result of low to moderate skills and low challenge. nakamura and csikszentmihalyi (2014) found that passive leisure activities and chores are commonly associated with apathy or boredom, which could provide an explanation for the large proportion of tv interest assigned to the irregular w-profile. the large portion of school interests in this profile might also be explained by flow theory: students might have low skills and experience little challenge with regard to these interests, and feel boredom or apathy in relation to them. nevertheless, although these interests might be regarded as boring, students still reported these activities or topics as interests. a possible explanation for the pattern of w-shaped interests comes from sociocultural theory, implying that interests are not always centered on the content but can also be a means to participate in a community one belongs to (azevedo, 2006, 2018; digiacomo et al., 2018; greeno, collins, & resnick, 1996). perhaps it is not the content of a w-shaped interest, but the participation in a community through this interest that makes the activity or topic interesting. for example, watching television at night with family might be interesting not necessarily because of the content of the program, but rather because it is a means to be part of the family. alternatively, in the case of passive leisure interests, it could be that a person is not interested in an activity because of the content, but because of the effect that engaging with the interest has on one’s mood and energy. kleiber, larson, and csikszentmihalyi (2014) categorize watching television as a relaxed leisure activity, “a type of experience that may restore one’s energy and spirit, but does not require exertion of effort” (p. 472). this restoration of energy might be exactly what makes the adolescents interested in this activity. the findings of this study shed more light on how the dimensions underlying the four-phase model of interest development (hidi & renninger, 2006) relate to one another in different interests. the flat patterns found are clearly represented in the four-phase model and seem very consistent with the descriptions of the four phases. the current study adds the irregular patterns to this model, which demonstrate more distinct ways in which the underlying dimensions can be interrelated. where hidi & renninger (2006) already expected some irregularities (e.g., by including ‘typically though not exclusively’ for some indicators), our results confirm and elaborate on these. it is recommended that future research investigates the complex interplay of dimensions in longitudinal studies in order to further address the relation of these patterns of interest to the developmental model. it is important to bear in mind that the results of the current study relied on self-reported measurement, which required respondents to be aware of interests in order to report them. renninger and hidi (2016) state that “respondents in an early phase of interest development may not be in a position to respond to questions about the level of their interest; they may not be conscious that their interest has been triggered” (p. 62). however, the low flat profile provides evidence that participants were in fact able to report on their recently externally triggered interests. providing the students with adequate instruction and keeping small, two-hour intervals between measurements has aided detection of externally triggered interests and put participants in a position to report on these. another note of caution regarding the interpretation of the results could be the limited sample in terms of number and age. even though the number of participants was relatively small (n = 94), data analysis was performed on the level of interests. as 1247 interests were reported, the sample size was deemed large enough to provide sufficient power for latent profile analysis (tein, coxe, & cham, 2013). regarding the limited age range of the participants, one might argue that interest domains differ across age groups (e.g., tracey, robbins, & hofsess, 2005) and this might limit generalizability of the results. even though we acknowledge that adolescents’ interest domains change over time, we have no indication that the structure of interest is different across age groups (as also noted by hidi & renninger, 2006). moreover, the current study captured a wide variety of interests reflecting different domains of daily life, which makes it probable that we captured the comprehensive construct of interest. the holistic and integrative approach of this study has allowed us to draw conclusions that can inform educational practice and interest research. with regard to school interests, we see multiple academic interests in all six profiles (though in varying quantities), and therefore the general conclusions of the study also yield for academic interests. the school interests categorized into the top and high flat profiles are especially relevant, as they challenge the generally held belief that students commonly lack individual interest in school subjects (as also found by slot, akkerman, & wubbels, 2019). as a consequence, educators might do well to determine the baseline subject-related interests of their students, in order to judge the necessity of triggering new, situational interest, or the possibility of appealing to existing, individual interests. the same goes for out-of-school interests, which can be important for study choices and hold learning potential (holmegaard, 2015; ito et al., 2018; vulperhorst, van der rijst, & akkerman, 2020). for example, research has shown that involving students’ personal interests in school can improve homework completion rates, task interest and engagement (e.g., hinton & kern, 1999; reber, canning, & harackiewicz, 2018). when assessing the interests of students, educators are advised to get a detailed understanding of the interest in terms of the dimensions, as for example the fact that an interest is very novel does not mean it is not important to the student (has high value). when educators gain a more detailed understanding of students’ interests this can provide them with useful information about the students’ out-of-school learning experiences, and avoid incorrect judgments about (the lack of) student interest. in addition, this study may help educators to identify students’ potential and latent interests and to observe how these may vary before deciding upon a way to evoke or sustain students’ interests in school. we propose more research into differences between (academic) interests assigned to the different profiles, for example in the object or educational context of these interests, to arrive at a better understanding of how to aid development of interests in classrooms. with regard to interest research, the novel bottom-up approach of this study contributes to the field in several ways. firstly, this study has demonstrated the necessity of measuring interests on multiple dimensions: single indicators are not sufficient to measure the experience of interest due to its heterogeneity. additionally, our findings suggest some additions to the developmental models of interest development so far (e.g., hidi & renninger, 2006; krapp, 2002), by demonstrating heterogeneous manifestations of interest. we do however stress that the fit indices in the current study were not conclusive on the exact number of interest variations in the data and the statistical support for our final model is not unambiguous. hence we do not claim that these six patterns are the only potential profiles or manifestations of interest. we do however maintain that all found profiles are strikingly distinct in both shape and content, and that the irregular profiles are well interpretable and theoretically meaningful. taken together this supports the notion of heterogeneity in manifestations of interest. we strongly recommend follow-up research to perform similar analyses on another dataset to confirm or refine these variations. to inform developmental models even more, we also recommend that researchers investigate the interplay of dimensions in a longitudinal manner to explicate multidimensional developmental relations, for example by using latent transitional modeling (lta). this may reveal how interests within both homogeneous and heterogeneous profiles develop over time. the findings of this study may help to understand and acknowledge the multidimensional experience of students’ interests in educational settings. this multidimensionality calls for an idiosyncratic perspective on interest, focusing on the specific structure of every particular person-object relation. only through this multidimensional structure we can fully grasp the powerful construct of interest and refine ways to evoke and sustain them in educational contexts. keypoints based on key scholars we identified six interest dimensions: historicity, value, agency, frequency, intensity, mastery using these dimensions in lpa, a variety of interest manifestations is captured by latent profiles the profiles demonstrate both a homogeneous and heterogeneous structure of interest interest measures should account for multidimensionality references ainley, m., hidi, s., & berndorff, d. (2002). interest, learning, and the psychological processes that mediate their relationship. journal of educational psychology, 94(3), 545–561. https://doi.org/10.1037//0022-0663.94.3.545 akkerman, s. f., & bakker, a. (2012–2014). interest in science: development across sites of learning. els starting grant, utrecht university, the netherlands. akkerman, s. f., & bakker, a. (2019). persons pursuing multiple objects of interest in multiple contexts. european journal of psychology of education, 34(1), 1–24. https://doi.org/10.1007/s10212-018-0400-2 alexander, p. a., kulikowich, j. m., & jetton, t. l. (1994). the role of subject-matter knowledge and interest in the processing of linear and nonlinear texts. review of educational research, 64(2), 201–252. https://doi.org/10.2307/1170694 arnold, f. (1906). the psychology of interest (i). psychological review, 13(4), 221–238. asparouhouv, t., & muthén, b.o. (2007). wald test of mean equality for potential latent class predictors in mixture modeling. retrieved from http://www.statmodel.com/download/ meantest1.pdf on april 21, 2016. azevedo, f. s. (2006). personal excursions: investigating the dynamics of student engagement.international journal of computers for mathematical learning, 11, 57–98. https://doi.org/10.1007/s10758-006-0007-6 azevedo, f. s. (2018). an inquiry into the structure of situational interests. science education, 102(1), 108–127. https://doi.org/10.1002/sce.21319 barron, b. (2006). interest and self-sustained learning as catalysts of development: a learning ecology perspective. human development, 49(4), 193–224. https://doi.org/10.1159/000094368 bruner, j. s. (1990). acts of meaning. cambridge, ma: harvard university press. csikszentmihalyi, m., & larson, r. (2014). validity and reliability of the experience-sampling method. in m. csikszentmihalyi (ed.), flow and the foundations of positive psychology (pp. 35–54). dordrecht, the netherlands: springer. csikszentmihalyi, m., larson, r., & prescott, s. (1977). the ecology of adolescent activity and experience. journal of youth and adolescence, 6(3), 281–294. https://doi.org/10.1007/978-94-017-9094-9_12 dewey, j. (1913). interest and effort in education. boston, ma: houghton mifflin. digiacomo, d. k., van horne, k., van steenis, e., & penuel, w. r. (2018). the material and social constitution of interest. learning, culture and social interaction, 19, 51–60. https://doi.org/10.1016/j.lcsi.2018.04.010 erstad, o., & silseth, k. (2019). futuremaking and digital engagement: from everyday interests to educational trajectories. mind, culture, and activity, 26(4), 309–322. https://doi.org/10.1080/10749039.2019.1646290 greeno, j. g., collins, a. m., & resnick, l. b. (1996). cognition and learning. in d. berliner & r. calfee (eds.), handbook of educational psychology (pp. 15–41). new york: macmillian. harackiewicz, j. m., barron, k. e., tauer, j. m., & elliot, a. j. (2002). predicting success in college: a longitudinal study of achievement goals and ability measures as predictors of interest and performance from freshman year through graduation. journal of educational psychology, 94(3), 562–575. https://doi.org/10.1037/0022-0663.94.3.562 harackiewicz, j. m., durik, a. m., barron, k. e., linnenbrink-garcia, l., & tauer, j. m. (2008). the role of achievement goals in the development of interest: reciprocal relations between achievement goals, interest, and performance. journal of educational psychology, 100(1), 105–122. https://doi.org/10.1037/0022-0663.100.1.105 hedges, h. (2019). the “fullness of life”: learner interests and educational experiences. learning, culture and social interaction, 23. https://doi.org/10.1016/j.lcsi.2018.11.005 hidi, s., & baird, w. (1986). interestingness—a neglected variable in discourse processing. cognitive science, 10(2), 179–194. https://doi.org/10.1207/s15516709cog1002_3 hidi, s., & renninger, k. a. (2006). the four-phase model of interest development. educational psychologist, 41(2), 111–127. https://doi.org/10.1207/s15326985ep4102_4 hinton, l. m., & kern, l. (1999). increasing homework completion by incorporating student interests. journal of positive behavior interventions, 1(4), 231–241. https://doi.org/10.1177/109830079900100405 holmegaard, h. t. (2015). performing a choice-narrative: a qualitative study of the patterns in stem students’ higher education choices. international journal of science education, 37(9), 1454–1477. https://doi.org/10.1080/09500693.2015.1042940 ito, m., martin, c., pfister, r. c., rafalow, m. h., salen, k., & wortman, a. (2018). affinity online: how connection and shared interest fuel learning. nyu press. kleiber, d., larson, r., & csikszentmihalyi, m. (2014). the experience of leisure in adolescence. in m. csikszentmihalyi (ed.), flow and the foundations of positive psychology (pp. 467–474). dordrecht, the netherlands: springer. köller, o., baumert, j., & schnabel, k. (2001). does interest matter? the relationship between academic interest and achievement in mathematics. journal for research in mathematics education, 32(5), 448–470. https://doi.org/10.2307/749801 krapp, a. (2002). structural and dynamic aspects of interest development: theoretical considerations from an ontogenetic perspective. learning and instruction, 12(4), 383–409. https://doi.org/10.1016/s0959-4752(01)00011-1 krapp, a., hidi, s., & renninger, k. a. (1992). interest, learning, and development. in k. a. renninger, s. hidi, & a. krapp (eds.), the role of interest in learning and development (pp. 3–25). hillsdale, nj: lawrence erlbaum associates, inc. lazarsfeld, p.f., & henry, n.w. (1968). latent structure analysis. boston: houghton mill. marsh, h. w., lüdtke, o., trautwein, u., & morin, a. j. (2009). classical latent profile analysis of academic self-concept dimensions: synergy of person-and variable-centered approaches to theoretical models of self-concept.structural equation modeling: a multidisciplinary journal, 16(2), 191–225. https://doi.org/10.2307/30047170 meeus, w., van de schoot, r., keijsers, l., schwartz, s. j., & branje, s. (2010). on the progression and stability of adolescent identity formation: a five‐wave longitudinal study in early‐to‐middle and middle‐to‐late adolescence. child development, 81(5), 1565–1581. https://doi.org/10.1111/j.1467-8624.2010.01492.x mitchell, m. (1993). situational interest: its multifaceted structure in the secondary school mathematics classroom. journal of educational psychology, 85(3), 424–436. https://doi.org/10.1037/0022-0663.85.3.424 muthén, l.k., & muthén, b.o. (2015). mplus user’s guide. 7th ed. los angeles, ca: muthén & muthén. nakamura, j., & csikszentmihalyi, m. (2014). the concept of flow. in m. csikszentmihalyi (ed.), flow and the foundations of positive psychology (pp. 239–263). dordrecht, the netherlands: springer. reber, r., canning, e. a., & harackiewicz, j. m. (2018). personalized education to increase interest. current directions in psychological science, 27(6), 449–454. https://doi.org/10.1177/0963721418793140 renninger, k. a. (2000). individual interest and its implications for understanding intrinsic motivation. in c. sansone & j. m. harackiewicz (eds.), intrinsic and extrinsic motivation: the search for optimal motivation and performance (pp. 373–404). new york, ny: academic press. renninger, k. a., bachrach, j. e., & hidi, s. e. (2019). triggering and maintaining interest in early phases of interest development. learning, culture and social interaction, 23, 100260. https://doi-org /10.1016/j.lcsi.2018.11.007 renninger, k. a. & hidi, s. (2016). the power of interest for motivation and engagement. new york, ny: routledge. renninger, k. a., & leckrone, t. (1991). continuity in young children’s actions: a consideration of interest and temperament. in l. oppenheimer & j. valsiner (eds.), the origins of action: interdisciplinary and international perspectives (pp. 205–238). new york, ny: springer-verlag. renninger, k. a., & su, s. (2019). interest and its development, revisited. in r. m. ryan (ed.), the oxford handbook of human motivation (pp. 205–225). oxford, england: oxford university press. sachisthal, m. s. m., jansen, b. r. j., peetsma, t. t. d., dalege, j., van der maas, h. l. j., & raijmakers, m. e. j. (2019). introducing a science interest network model to reveal country differences. journal of educational psychology, 111(6), 1063–1080. https://doi.org/10.1037/edu0000327 schiefele, u. (1992). topic interest and levels of text comprehension. in k. a. renninger, s. hidi, & a. krapp (eds.), the role of interest in learning and development (pp. 151–182). hillsdale, nj: lawrence erlbaum. schiefele, u. (2009). situational and individual interest. in k. wentzel, a. wigfield, & d. miele (eds.), handbook of motivation at school (pp. 197–222). new york, ny: routledge. schiefele, u., krapp, a., & winteler, a. (1992). interest as a predictor of academic achievement: a meta-analysis of research. in k. a. renninger, s. hidi, & a. krapp (eds.), the role of interest in learning and development (pp. 183–212). hillsdale, nj: lawrence erlbaum. schiefele, u. & rheinberg, f. (1997). motivation and knowledge acquisition: searching for mediating processes. in m. l. maehr & p. r. pintrich (eds.), advances in motivation and achievement, volume 10 (pp. 251–301). greenwich, uk: jai press. slot, e., akkerman, s., & wubbels, t. (2019). adolescents’ interest experience in daily life in and across family and peer contexts. european journal of psychology of education, 34(1), 25–43. https://doi.org/10.1007/s10212-018-0372-2 strong, e. k., (1927). vocational interest test. educational record, 8, 107–121. tein, j. y., coxe, s., & cham, h. (2013). statistical power to detect the correct number of classes in latent profile analysis.structural equation modeling: a multidisciplinary journal, 20(4), 640–657. https://doi.org/10.1080/10705511.2013.824781 tobias, s. (1994). interest, prior knowledge, and learning. review of educational research, 64(1), 37–54. https://doi.org/10.3102/00346543064001037 tracey, t. j., robbins, s. b., & hofsess, c. d. (2005). stability and change in interests: a longitudinal study of adolescents from grades 8 through 12. journal of vocational behavior, 66(1), 1–25. https://doi.org/10.1016/j.jvb.2003.11.002 vulperhorst, j. p., van der rijst, r. m., & akkerman, s. f. (2020). dynamics in higher education choice: weighing one’s multiple interests in light of available programmes. higher education, 79, 1001–1021. https://doi.org/10.1007/s10734-019-00452-x walkington, c. a. (2013). using adaptive learning technologies to personalize instruction to student interests: the impact of relevant contexts on performance and learning outcomes. journal of educational psychology, 105(4), 932–945. https://doi.org/10.1037/a0031882 appendix 5-dimension model excluding historicity table a.1 log-likelihood, model fit indices and entropy measures for the 1to 7-profile models excluding historicity (n = 1247). figure a.1. profiles in the six-profile solution. winne publication frontline learning research vol 6 no. 3 special issue (2018) 250-258 issn 2295-3159 discussion paradigmatic issues in state-of-the-art research using process data philip h. winnea afaculty of education, simon fraser university, canada abstract learning science is enthusiastically adopting new instruments to gather physiological and other forms of event data to represent mental states and series of them that reflect processes. in an attempt to provoke more thought about this kind of research, i suggest paradigmatic issues relating to data, analyses of them and interpretations of results. i advocate we not label these data as “objective.” instead, we share a subjective interpretation of them. i argue propositions about validity need more nuance. bounds on generalization related to so-called ecological validity are rarely empirically justified. when researchers transform raw data before analysis and when analytic methods partition variance, interpretations of results omit key qualifications. i posit emotion and motivation be positioned in theory as moderators rather than mediators because agentic, self-regulating learners make and revise knowledge by choosing forms of cognitive engagement in a context where they interpret arousal. i note that researchers’ anchor interpretations of process data in learners’ accounts. this creates a tautology that troubles usual notions of reliability. finally, i recommend research involving process data turn more toward helping learners identify conditions of learning that spark arousal so learners can regulate motivation and emotion. this leads to a surprise: treating learners as individuals and helping them identify triggers of arousal may recommend learning science cast emotions and motivation as epiphenomena. keywords: validity; trace data; agency; motivation; emotion; self-regulated learning, learning science paradigm info. corresponding author: : philip h. winne, faculty of education, simon fraser university, burnaby, british columbia v3h 4r2, canada doi: https://doi.org/10.14786/flr.v6i3.551 paradigmatic issues in state-of-the-art research using process data a great deal of recent research has investigated relations between learners’ affective states, brain states and other physiologically-related variables to traditional measures of achievement, motivation and upcoming indicators of cognitive processing called traces. articles in this special issue represent a broad and high-quality sample of these efforts. they boldly explore newer approaches to gathering data, tackle challenges in analyzing data with unconventional properties, and suggest new views of frameworks to account for learning processes that create outcomes. it is almost certain that no research study is conceptually faultless or methodologically perfect. interpretations and implications researchers draw about their results arise in that context and, thus, are debatable. in this article, i do make conventional critiques about whether this or that instrument or analytic approach is faulty or whether another is likely more appropriate. instead, i describe from my perspective paradigmatic issues about these kinds of research. my aim is to provoke thinking not about any particular study but about fundamental characteristics of this up-and-coming line of research. 1. process and trace data are a step forward but not the truth one description often applied to process data is they are objective. depending on what one means by “objective” this is a valid interpretation or it is wrong. it is fine to label process data as objective in the sense that “reasonable” observers can agree whether an event occurred. i prefer to think of this as shared subjectivity rather than objectivity. a second sense of the concept of objectivity is wrong. from this perspective, data are conceptualized as incapable of having any other value. i elaborate. data such as a gaze duration or a button click are verifiable – a learner’s gaze settled on a particular area of interest for k or more milliseconds. a button was clicked. the metric is nominal and boundaries are definite. for data like these the question is: what are properties the counting metric? if each gaze period or each click is identical, each event may count as “1” and a total period of gazing at something in particular or the sum of clicks can be identified by adding each instance. in this context where data are labeled “objective” theoretical constructs must be considered. first, i take as axiomatic – believed without proof – learners are agents. they are in control of what they think about. i fend off an immediate counterclaim about mental activities that might be considered a “cognitive reflex.” some mental activities are genuine reflexes. the startle response is an example. genuine reflexes like these are very infrequent in everyday learning situations, so i treat them as rare and genuine anomalies. other apparent cognitive reflexes are learned cognitive routines that have been automated through extensive experience, particularly practice with feedback. learners are supposed to develop and automate such routines and use them as they study and collaborate. understanding others’ speech at everyday rates of utterance is an example. other examples include number facts such as 4 × 8 = 32, raising a hand to be recognized in discussion, and subvocalizing roy g biv to name colors in the visible spectrum in order of their wavelengths from shorter to longer. as well, education encourages students to disassemble other cognitive routines that are disciplinary or social misconceptions. examples are: denominators must be identical to multiply fractions, females are not adept at math and maintenance rehearsal is the best tactic to promote recall . because learners are agents, observers need more information than gaze duration or clickstream events to validly interpret an observation. suppose a learner gazing at a particular region in a diagram of an electric circuit could use software to identify that region (e.g., enclose it in an ellipse) and tag it “confusing.” these extra data – the region enclosed plus the tag – signal what the learner was thinking that caused gaze to linger. the learner judged (metacognitively monitored) she was challenged to understand something about that part of the circuit. she was motivated to make a permanent record of that state of mind, so she drew an ellipse. quite likely, she plans to search later – ask the teacher or a peer, comb the internet – to locate information about this and other content tagged “confusing” to resolve these confusions. suppose a button in a software tool was labeled “see more ….” a click of that button signals the learner is seeking additional or elaborative information beyond what is presented in the current display. clicking the button represents interest, an expectation useful material will be found at the resource linked to the button, and an efficacy expectation understanding can be enhanced by accessing that new information. these examples illustrate trace data (winne, 1982). the best instances of trace data copule an a verifiable event – gaze lingering for a measured time on a particular bit of information, a button clicked, a tag applied – with a convincing theoretical claim about what that event “means.” when learners generate trace data without having to do much more than they normally would do as they study – when the data are ambient in the sense means used to generate data are integrally involved in ways learners normally engage with information (see winne, teng, chang, lin, marzouk, nesbit, patzak, raković, samadi, & vytasek, 2019) – observers have the additional information needed to construct well founded inferences. note, however, such inferences are nonetheless grounded in a theoretical framework. an event datum or string of event data create an opportunity to ask, why did that event occur? why is the string of events shaped as it is? learning science is keen to notice events and characterize strings of them. but event data are not objective. people notice anomalous events because they are unexpected and interesting with respect to subjective schemas that describe the world. when researchers use instrumentation to notice an event, theories underlie the mechanisms that allow the instrument to record the event. in short, process data and especially trace data are inherently and inescapably subjective. what we mean by the label “objective” is really that we share subjectivitity. 2. claims about validity often overreach validity is a concept often misinterpreted. as messick (1989) argued, validity is not a property of an instrument, a protocol or a setting. settings, instruments and protocols do not have validity. validity is more nuanced. it is a property of an inference or an interpretation. people construct inferences and interpretations. i consider three important cases. 2.1 to what does validity apply? statements of relationship or causation – the findings reported about research in learning science – have a degree validity. that degree is proportional to the extent constructs named in stating a finding correspond to operational definitions that describe how data were generated. one form of this relationship was just discussed in regard to trace data. other less obvious cases need address. transformations of data also transform constructs suppose a researcher transforms raw data. two common kinds of transformations populate our research literature. one is transformations of scale, such as a log transformation of continuous scores or an arcsine transformation of proportions. these are often used to reshape a distribution of data so it more closely matches a gaussian (normal) distribution that is better suited to many inferential statistical methods. a second widely used transformation of data is statistically partitioning and removing variance from a variable. examples are standardized partial weighting (e.g., regression) coefficients in a linear modeling analysis. both scale and variance partitioning transformations of data change raw scores into new scores. these new transformed scores are then analyzed by a statistical or machine learning method. results of these analytical methods are then framed using words identical to the untransformed data. this is careless phrasing. the fault is not with the numerical work. it has been rendered appropriate by the transformation. the fault lies in failing to recognize transformed scores introduce an additional operational definition, the transformation. unlike careful attention to reporting the span of a scale for reporting responses to questionnaire items, the number of options offered in multiple-choice items or sampling rates for physiological processes that vary continuously across time, changes to data wrought by numerical transformations are almost never taken into account. why is this important? numerically transformed data submitted to analytical methods represent a different construct than the construct represented by raw data. but, this difference is not represented in descriptions and interpretations of analytical results. researchers don’t write, “the base 10 log version of our control variable … .” the nomological network changes when data are transformed (see winne, 1983). correlations of raw scores with transformed scores always are less, sometimes much less, than 1.00. correlations of raw scores with other anchor variables differ, sometimes considerably, from correlations of transformed scores with those anchor variables. learning science is misled whenever the full operational definition of data is not recognized. 2.2 concerns for ecological validity are rarely empirically justified a third concern about validity relates to the concept of ecological validity. variations of this claim are common in the literature of learning science: “because learners carry out tasks in settings where they learn every day, a study’s findings have ecological validity.” it follows: “when conditions drift away from those characterizing an everyday setting, e.g., from a regular classroom to a lab, findings lose ecological validity in proportion to the drift.” the error of this claim about lessened generalizability is twofold. first, worry that findings disintegrate in a new setting very rarely have grounding in any or sufficient research demonstrating that factors differentiating the settings actually affect findings. it is too casually presumed without empirical backing that such and such a factor differentiating settings is a cause or moderator of an effect. two examples are our literature’s practice of reporting where a sample originates (e.g., a western canadian university) and the proportions of females and males in a sample. to my knowledge, no study has demonstrated geographic locations of canadian universities moderate variance in findings other than the residential addresses of participants in studies. or, if one study’s participants are 60% female and another study’s participants are 48% female, where are studies demonstrating sex of participant is a proven moderator of cognitive processing or emotion? in general, our research traditions have a very sparse catalog of factors that influence the values of population parameters as a function of factors such as location, proportion of females and males, language first spoken, ethnic heritage and the like (see winne, 2017). in fact, claims about “ecological validity” are more guesswork. there are straightforward repairs for this fault. researchers can analyze their data to investigate the extent to which ecological factors moderate findings in their study. to do this places two burdens on research. first, such factors need to be considered at the outset of a study and measured. second, sample sizes need to be larger. another solution is to scour the literature for findings that demonstrate an ecological factor does moderate what was investigated in a current study. lastly, researchers could refrain from speculating about the extent to which the generalizability of findings may be limited by ecological factors lacking empirical backing. 3. arousal and emotion moderate learning emotion has become a “hot” (pun intended) topic on learning science. how might emotions relate to learning? i argue they are moderators not mediators. knowledge is fashioned by cognitively operating on information, e.g., generating an explanation for causal influences improves comprehension and memory (bisra, liu, nesbit, salimi, & winne, 2018). affect or emotion can moderate a process like self explanation and other cognitive operations applied to information in two ways. first, an affective state or emotional thought may be one factor a learner judges when deciding whether to apply a cognitive operation to particular information. this influences what a learner learns not because the learner is in a particular emotional state but because a cognitive process is or is not applied to particular information when the learner experiences a particular emotion. second, the contents of long-term memory are multidimensional. learners can judge the accuracy, thoroughness, and reliability of their knowledge (although they sometimes err in the accuracy of their judgments). knowledge in long-term memory is associated with a variety of other contents, for example, contextual descriptions about when a proposition was added to memory, what other knowledge it relates to and an affective stance regarding that proposition. elaborations like these form a network, and characteristics of the cognitive network correlate to what learners perceive, learn and can recall. an affective experience or emotional thought is elaborative content. these experiences augment results of “cold” cognitive operations and are added into the network of long-term memory. cognitive operations are the causes of learning. affect and emotional experiences are content that may moderate cognitive operations. several articles in this special issue set a stage for an intriguing deduction. a learner may not be aware a particular emotion is aroused while learning but, nonetheless, that state of arousal may influence cognition or it’s manifestation in interpersonal interactions. thus, what learners learn may vary without the learner’s awareness. what research has yet to explore is whether it is helpful to alert a learner that “now” is the time to regulate arousal. a complication to such a study is learners’ capacities to change one affective state into another, that is, to regulate affect. alerts about tacit emotions will be effective in proportion to this aptitude. future research might probe how well learners can regulate affect and whether learners can learn to regulate affect to benefit learning. 4. puzzles about proxy measures of emotion physiological signals and facial displays are proxies for learners’ experiences of emotions and other motivational constructs. measures generated by instruments that represent these experiences are fundamentally different. measures developed using physiological sensors used to detect arousal, for example, blood flow or eda, rely on comparisons to a baseline. researchers declare deviations from a subjectively set threshold to mark arousal or recovery. deviations span time. in contrast, facial displays are measured by configurations of points on a face at a point in time. these measures are absolute; a configuration is matched or it is not. time is irrelevant to the measurement. this contrast invites a question: what roles does time or rate of change play in a learner’s perception of affect? do learners perceive affect as variance in arousal over time or do they perceive affect as step function, on or off? what implications, if any, might this have for theorizing about affect and for learners approach to regulating affect? 4.1 a vexing tautology both kinds of proxies for affect are validated by asking learners to label their arousal. this creates a tautology. a researcher observes a measurement and asks the learner, “what is your emotion now?” thereafter, that measurement is taken as a signal of the named emotion. how can the researcher be sure? the learner said so. why is this tautology an issue? a benefit of instrumentation used to gather proxy data is its unobtrusiveness. the learner does not have to be interrupted to report states of arousal or variation in affect. however, beyond instrumenting arousal, researchers are interested to identify valence. today’s state-of-the-art instruments can’t do that without confronting the tautology just described. 5. is reliability relevant? when a learner reports and proxy measurements do not correlate highly, two cases need sorting. first, one or the other – the learner or the instrument – may be biased. learners may be reluctant to report some affects. social demand does not disappear in the lab, work groups or classrooms. also, instruments may be miscalibrated. detecting and correcting bias in learners’ reports requires ground truth. the previously mentioned tautology undermines faith there is ground truth. a second case to consider when learners’ reports do not highly correlate with proxy measurements is when the learner, the instrument or both are unreliable. data generated by traditional instruments – responses to survey questions, answers to achievement test items, and the like – almost always prompt researchers to investigate the reliability of those data. following cronbach, gleser, randa and rajaratnam (1972), reliability is a matter of identifying sources of variance in scores, identifying which of those factors (or facets) cannot be explained or controlled, and quantifying the contribution those “nuisance” sources make to variance in scores. that is, unreliability arises because nuisance factors introduce erratic variance in learners’ reports or an instrument’s measurements. some instruments generating proxy data about arousal, for example, sensors registering eda, can be examined for reliability using other mechanical systems with known reliability. this helps to disassociate erratic variability in measurements researchers use to gauge learners’ arousal. any residual variance attributed to extraneous factors is then ascribed to the learner. but, if the learner is not biased (i.e., reliably misreports about arousal), can the learner be wrong in declaring their experience of affect? again, the tautology occludes interpretation. which source of data is the reliable source? 6. final thoughts: applying findings and the evolution of learning science learners’ interpretations of states of arousal create their emotions and motivations. they will report emotions and motivations when asked. researchers gather and interpret proxy data about these states. both learners’ reports and researchers’ proxies correlate with what learners learn. learners’ emotions and motivations do not create knowledge or change knowledge. knowledge is created and amended when learners apply cognitive operations to information. emotions and motivations are factors, mixed with other information, learners weigh in choosing which information to operate on and which cognitive operations to apply to selected information. according to this logic, emotions and motivations are moderators but not causes of learning. if this logic is correct, what are implications for helping learners become better learners? i encourage research aim to identify conditions learners perceive that arouse them. then, identify which of those conditions correlate positively and which correlated negatively with what learners learn. with a more precise map of relations between conditions learners perceive about their instructional context and learning results, instructional designs can more dependably be engineered to offer learners experiences in which learners can more productively self-regulate learning. this account of learning leads to a potential surprise. suppose conditions of instruction that arouse a particular learner are known. suppose further conditions positively related to this learner’s achievements can be incorporated into the instructional context and those that negatively correlate with learning can be removed. can learning science now ignore a learner’s motivations and emotions. is the ultimate goal of learning science finding dependable principles for engineering features of instruction that are devoid of constructs about inner states like emotions and motivations? should learning science strive to locate theoretical constructs such as emotions and motivations to the category of epiphenomena? this line of thinking explicitly and emphatically acknowledges students are individuals. averaging out individual differences to describe effects at the level of a randomly formed group pervades research in learning science. i argue this is a mistake in attempts to model learners as self-regulating agents (winne, 2017). learners are agents (winne, 2018). they choose how to learn. they are in control. what matters is their perception of instructional conditions. each learner’s perception emerges from a history of individual experiences about which conditions merit attention, what those conditions represent and predict about the present context, and what is a path to goals that context. the surprise i reveal about learning science and guidance it can generate for designing instruction rests on the axiom that learners are agents. instructional design therefore must be sensitive to, responsive to and supportive of learners as individuals. in this quest, instrumentation and methods like those described in this special issue are a boon. with energetic attention to methodological and interpretive issues affecting how data are interpreted, learning science carried out using methods and with instrumentation described in this special issue offers great opportunity to elevate understandings about each learner as an individual, about relationships between each learner’s perceptions of their learning environment, and what learners learn in those contexts. key points process data are better described in terms of shared subjectivity rather than as objective. transformations of data change constructs and too often are not recognized in reporting effects. concerns about ecological validity often lack empirical grounding. emotion and motivations are moderators, not causes of learning. grounding interpretations of process data on learners’ reports is tautological, undermining traditional concerns about reliability. surprisingly, progressive research relating process data to achievement may render emotions and motivations to the category of epiphenomena. acknowledgements foundations for this article were developed with financial support provided over many years by the social sciences and humanities council of canada and simon fraser university. references bisra, k., liu, q., nesbit, j. c., salimi, f., & winne, p. h. (2018). inducing self-explanation: a meta-analysis . educational psychology review, 30, 703-725. cronbach, l. j., gleser, g. c., nanda, h., & rajaratnam, n. (1972). the dependability of behavioral measurements. new york: wiley. messick, s. (1989). validity. in r. l. linn (ed.), educational measurement (3rd ed., pp. 13-103). new york: macmillan winne, p. h. (1982). minimizing the black box problem to enhance the validity of theories about instructional effects. instructional science, 11, 13-28 winne, p. h. (1983). distortions of construct validity in multiple regression analysis. canadian journal of behavioural science, 15, 187-202 . winne, p. h. (2017). leveraging big data to help each learner upgrade learning and accelerate learning science. teachers college record, 119(3), 1-24. winne, p. h. (2018). cognition and metacognition within self-regulated learning. in d. schunk & j. greene (eds.),handbook of self-regulation of learning and performance. (2 nd ed., pp. 36-48). new york, ny: routledge. winne, p. h., teng, k., chang, d., lin, m. p-c., marzouk, z., nesbit, j. c., patzak, a., raković, m., samadi, d., & vytasek , j. (2019). nstudy: software for learning analytics about processes for self-regulated learning. journal of learning analytics, 6 , 95-106. codepen 7. dahlberg frontline learning research special issue vol.9 no.2 (2021) 145 169 issn 2295-3159 widening participation? (re)searching institutional pathways in higher education for migrant students the cases of sweden and italy giulia messina dahlberg1, sylvi vigmo1 & alessio surian2 1university of gothenburg, sweden 2university of padova, italy article received 17 may 2020/ revised 30 november / accepted 4 december / available online 11 march 2021 abstract the aim of this study is to shed light on the ways in which transitions and support are framed in policy contexts in relation to widening participation in higher education (he) in sweden and italy. more specifically, this study investigates the ways in which the discourse about the inclusion of migrant students in he is framed in relation to the kinds of support for this group offered in two higher educational institutions, in sweden and italy. furthermore, the study sheds light on the ways in which policy ideas about transition and widening participation are enmeshed in the students’ narratives and how they affect their experiences of participation, normalization and marginalization in he. the analysis includes two datasets: i) national policy, laws and regulations and webpages of a selection of national universities and university colleges; and ii) ethnographically generated data that builds upon a case-study design and consists of audio recordings of informal discussions and interviews with students. we are, in this study, interested in framing diversity in terms of a move beyond the naturalization of hegemonic stances where labelled “others” (e.g. based on cultural/ethnic background, functionality, socio-economic status) are treated as essentialized or mutually exclusive categories. one of the central, frontline contributions of this study, lies in its attempts to analytically scrutinise processes of inclusion and marginalisation that include a broad analytical gaze. this allowed us to analyse the mismatch between the range of support provided, and the actual needs and challenges that migrant students meet in their transition and participation to higher education in two european countries. keywords: transitions; widening participation; migrant students; higher education; vertical case studies. info corresponding author email: giulia.messina.dahlberg@gu.se doi: https://doi.org/10.14786/flr.v9i2.655 1. introduction 1.1 disentangling diversity: the case of the intersection of higher education with migration due to the migration waves in the past three decades in the geopolitical spaces of europe, student heterogeneity has gained new dimensions in addition to social and economic mobility. european countries have faced and are facing similar challenges with regard to the inclusion and integration of migrants in all sectors of society, not least in higher education (he). mobility is here understood as a human condition wherein contemporary migration is no longer framed as a “unidirectional migrant passage” (guo, 2015, p.9) but rather as “the multiple and circular migration across transnational spaces” (p. 7). such a precarious condition, we argue, may afford and/or prevent individuals’ opportunities for full participation and socialization in society in a range of different ways. over the past century, higher educational institutions have started to engage with a compelling agenda for the inclusion and integration of an increasingly “diverse” student population. in the geopolitical space of sweden, issues of openness and inclusion have marked educational policy since world-war ii, not least in regard to higher education. a government proposal “the open university” (prop. 2001/02:15) discusses widening recruitment and participation where fundamental issues include the need for the student population to reflect the makeup of society generally, and a belief that all levels of diversity, including the heterogeneity of ideas, beliefs and approaches, is key for achieving academic excellence. in italy, higher educational institutions’ policies concerning the inclusion and promotion of diversity, are informed by regulations and resources by the ministry of interior and by the ministry of research and university. for instance, in 2020, migrants who were granted international protection by the ministry of interior were offered 100 study grants to access italian universities. furthermore, unhcr produced a specific manifesto for the inclusive university that addresses the inclusion and integration of refugee students and academics (unhcr, 2019), signed by 43 italian universities so far. there are, however, several tensions between how the discourse about diversity, integration and widening participation gets framed in policy documents like the ones briefly presented above, and the ways in which practices of inclusion and support play out in the everyday life of students and academics. in the present study, the university is understood as a complex site where political, economic and social interests intersect (bacevic, 2019; tummons & beach, 2019). this involves the idea that the university is playing a major role in the global corporatisation, marketisation and massification of (higher) education on the one hand (giroux, 2010; tight, 2019), and the changing and growing societal expectations on he to support democracy and equity in the 21st century on the other (see also fox, baker, charitonos, jack, & moser-mercer, 2020). furthermore, the fields of intercultural studies and global education have yet to come to terms with diverging understandings of cultural differences according to reference frameworks that prioritize value rubrics, cultural mobility and “effective” cross-cultural communication in contrast with social mobility and critical decolonial perspectives (walsh, 2012). the research presented in this study is an attempt to shed light on and unpack some of these tensions and contradictions. taking the above as points of departure, we are, in this study, interested in framing diversity in terms of a move beyond the naturalization of hegemonic stances where labelled “others” (e.g. based on cultural/ethnic background, functionality, socio-economic status) are treated as essentialized and, even worse, mutually exclusive categories. we argue for a conceptualisation of identity, and therefore also diversity, in terms of an intricate process in which the agency of human beings does not exclusively reside in the single individual, but is also always situated, i.e. related to contexts, from specific micro-moments, to macro-structures, of which policies are an important part (see also ecclestone, biesta & hughes, 2010). diversity, from such a line of thinking, is something that gets done in a practice that is entangled with other practices, both at the micro and macro analytical level. finding (new) ways to explore and understand practices from such analytical positions is, we argue, of crucial importance when dealing with the study of the many dimensions of life for marginalized groups in society. in such a line of thinking, practices of being and becoming a student are linked to institutional contexts, social regulations and many other practices in which the social and the material are entangled (gherardi, 2017). such a “relational epistemology” (law, 1994), we argue, highlights the performative dimension of sociomateriality and provides us with the theoretical tools for the study of identity as being in a constant state of becoming, i.e. fluid, ever changing and situated, rather than essentialistic, fixed or static (hancock, 2016). and yet, “fixed” binary categories (e.g. either you are a migrant or not) are sine qua non conditions and play a crucial role in the ways in which support services, special programmes and educational arrangements are planned and operationalised in policy as well as in practice. furthermore, diversity is always entrenched in a fabric of ideologies and hegemonies (see e.g. messina dahlberg & bagga-gupta, 2019). who is “diverse” is often framed as in need of some sort of institutionalised support in policy (bagga-gupta, messina dahlberg & winther, 2016; bagga-gupta, messina dahlberg & vigmo, 2020), and this is thought of and done differently in different communities. the aim of this study is to shed light on the ways in which policy ideas about widening participation, diversity and inclusion in he are related to practices in the study of students’ institutional pathways in two european countries by answering the following questions: what are the ways in which the discourse about the inclusion of students in he is framed in relation to the kinds of support for migrant or refugee students offered in two higher educational institutions, in sweden and italy? in what ways are policy ideas about transition and widening participation enmeshed in the students’ narratives and how do they affect their experiences of participation, normalization and marginalization in he? given the aim and the research questions outlined above, we use a comparative approach to illustrate the ways in which widening participation, diversity and mobility are both context-sensitive, local phenomena, but also respond to a global discourse of growing societal expectations on he to include a diverse student population. the analysis is based on two selected european countries, sweden and italy, located at the antipodes of the european map. italy’s geographical position is strategically central for migration waves to europe that originate from the global south. swedish immigration has historically been characterised by labour migration and an integration policy largely based on notions of equality and diversity (in terms of multiculturality and integration, rather than assimilation, see also kupský, 2017). thus, the two countries’ ideological and political agendas build upon different cultural and historical frames of reference, that, in turn, affect the ways in which inclusion and integration are framed in policy and implemented in practice. such perspectives, we argue, deserve further scrutiny from a comparative and multi-scale perspective. the rather ambitious endeavour to investigate policies and their mutual relations to practice in two different nation-states across time and space has been undertaken in the study of two datasets. these have allowed us to partially connect different levels of analysis from micro to macro levels and to provide further insights on the ways in which students’ lived experiences are mutually entangled with current national and supranational policies about transitions and widening participation in he. 1.2. universities as boundary spaces there is an extensive body of research concerning the intersection of higher education and migration, where the focus lies on migrant students’ lived experience. students are often framed in the literature in terms of being in a “betwixt space”, a space between home and he, which risks leading to a sense of not belonging anywhere (dunwoodie, kaukko, wilkinson, reimer & webb, 2020; hope, 2017) or as undergoing a “transitional experience” in the physical context of the university (macfarlane, 2016). similarly, the work by gennep and maskalaki (middleton, 2018), focuses on “liminality” as the observation and analysis of factors clustering around the “rites of passage” that mark an individual’s transition from one status to another; and on “emplacement”, i.e. a sense of common identity and place in relation to boundary spaces (such as institutional educational spaces). political and policy initiatives that aim to widen participation in he have been investigated in relation to how potential students experience and navigate ways into academic studies, and what barriers can negatively affect pathways into he (farenga, 2018; lopez gavira & moriña, 2015; macfarlane, 2016). furthermore, challenges for first-generation as well as low-income students are linked to conditions for guidance and support during the attempt by higher educational institutions to reach these student groups (brown, wohn & ellison, 2016). however, access to information needs contextualization and “translation” from a more knowledgeable person to situate and link that information to students’ everyday practice and social network practices, that, thus, become important access points for the student (2016). a qualitative interview study in germany (schneider, 2018), focuses on syrian asylum seekers and refugees (asrs) applying to german higher education. the general entry requirements were linked to constraining financial costs given in the asylum seekers’ and refugees’ accounts of their socioeconomic situation, and were closely connected to an identity marker of being a refugee. schneider (2018) draws the conclusion that the formal university admission requirements would “morph into practically insurmountable barriers” (p. 468). other barriers were the language requirements, and the total lack of recognizing previously achieved skills and qualifications, unless academic documents were translated into german for verification. further insights provided through research reinforce the necessity to widen the perspective to integrate the multifaceted social practices present in students’ diverse life trajectories (hope, 2017; kyndt, donche, trigwell & lindblom-ylänne, 2017; trigwell, 2017). taylor and harris-evans (2018) highlight that the “entangled” and “irregular” transition processes must be revisited and reconceptualized to involve students’ lived realities. similarly, gale and parker (2014) point to a lack of research on the “transition as becoming” metaphor, and argue that students’ lives and their realities should become involved. this also means ensuring that other actors, who influence the transition processes, are included to address the complexities found in the non-linear processes of transition across time and space (kyndt, et al., 2017). revisiting mainstream norms, and how these impact conditions for transitions processes, calls for more diverse approaches and flexibility concerning the investigation of how students receive support (taylor & harris-evans, 2018), and what it means for “becoming, being and achieving” as a student in he, as hope (2017) succinctly puts it. adopting a linear perspective for how to adapt and conform to being a student in he, departs from the notion that the student her/himself is involved in the attempt to fit the university student norm (taylor & harris-evans, 2018). to address these challenges, a more fine-grained theoretical as well as holistic approach can contribute to balance perspectives and support the development of transition practices (kyndt et al., 2017). to conclude, the research presented in this section highlights a central challenge for migrant students, i.e. the realization and recognition of the value of their competences and abilities which are now not incorporated in the set of skills that seem to be required to participate in and complete mainstream higher educational programmes. there is, however, a glaring lack of research that focuses on how such a deficit perspective varies and impacts on students’ educational pathways in terms of transitions and participation in different disciplines and professional/vocational programmes (mangan & winter, 2017; harvey & mallman, 2019). comparative studies that focus on educational landscapes in countries in europe that have received large amounts of migrants over the past decades, are also scarce. furthermore, a general concern raised in the literature is the epistemological and methodological issue of integrating different aspects that have an impact on students’ educational pathways and transitions to he across analytical levels (gravett, kinchin, & winstone, 2020). this article aims at contributing and filling this gap in the analysis of two datasets, where the intersections between policy and practice are scrutinized in a comparative study between sweden and italy. in the remainder of the paper, section 2 discusses the vertical case study as the methodological approach. the results are presented in section 3. the study ends with a discussion of the results and their implications for the inclusion and integration of a diverse student population in higher education 2. method: vertical case study the overarching issue that this study aims at taking onboard is to investigate the ways in which policy ideas are recontextualized in local sites of practice. to address this aim and the research questions we use a multi-sited (marcus 1995), comparative approach to ethnographic studies of educational practices that is no longer bound to a specific “group” or a community, but, and in line with our take on diversity and identity outlined above, “reframes the basic unit of cultural analysis as processual and iterative” (eisenhart, 2017, p.142), thus expanding the contextualization of ethnography to shed analytical focus on how policies, practices and institutionalized arrangements in one site, move and are entangled to other contexts and timescales. bartlett and vavrus (2014) called such an approach “vertical case study”, but they emphasise that the term “vertical” is a reminiscence of how they initially conceptualised the approach that now incorporates, besides the vertical, also horizontal and transversal elements of comparison. figure 1: illustration of the vertical case study approach (adapted from bartlett & vavrus, 2014). given the aim of this study, such a comparative approach deals, on the horizontal axis (figure 1), with a variety of contexts that includes case studies with students from the spaces of sweden and italy at selected universities. university a (unia) and university c (unic) are located in sweden, while university b (unib) is located in italy. the transversal axis represents policy implementation across time and that, we argue, represents the dimension that may assist us in the endeavour to come away from research methods that, more or less exclusively, focus upon a particular singular group or population, to reframe interpretive logics and representational techniques as travelling “across time, space and level, rather than as the characteristics or lifeways of bounded groups” (eisenhart, 2017, p.135). the data for the vertical case study outlined here was created through two complementary methods: analysis of policy of support (dataset 1) and case studies of three emblematic students from a broader dataset of 13 participants (dataset 2). the first method was employed to analyse policy content and its connection across levels and sites. the second method was used to analyse the students’ narratives of their lived experiences of transitions to he. these methods used together allowed us to deal with micro and macro levels of analysis and to shed light on the relations across levels but also on their potential to engage with multiple spatial fields or sites, represented by the blue, waved line in figure 1. thus, the multi-sited ethnographies reported in this study offer means of examining the ways in which the participants in the case studies make sense of their transitions towards becoming university students when the conditions to achieve such transformations are constituted and constrained by their connection with policy, educational activities, support services, and technologies. the policy of support (dataset 1) corresponds to the vertical axis of comparison (macro-micro). the analysis of dataset 2 connects the horizontal axes (policy implementation across sites) in terms of the analysis of students’ narratives of transitions and participation in higher education at their local practices (micro-level). the nature and process of data creation of the two datasets are detailed further in the next two sub-sections. 2.1 dataset 1 the nature of dataset 1, created using an ethnographic approach by means of connection across websites and documents, lends itself to envisage the analytical scaling model from macro to micro in terms of entanglements, rather than levels or layers. this logic also gives further fuel to the debate about the notion of analytical levels and on the kinds of (dis)advantages that it may bring, in relation to an alternative perspective according to which there are always macro aspects in the micro and vice versa (see e.g. faubion & marcus, 2009; tsing, 2005). dataset 1 includes a selection of the provision of support services in he as it is framed and described in webpages and policy documents in a selection of national universities in sweden (unia) and italy (unib). figure 2. vertical entanglement of policy. overview of the principal nodes and connections. figure 2 illustrates the entanglements from the first searches using a selection of keywords in the different webpages (e.g. the “entry points” box in light grey). the illustration develops in two opposite vertical directions, starting from the entry points box in the middle. on the top half, the path in the italian data is represented, while the bottom half shows the principal nodes in the swedish data. figure 2 could be envisaged as an hourglass, of which the entry points box constitutes the narrowest part. the metaphor of the hour glass also allows us to illustrate the vertical movement “across scales” as well as search words that have been a central aspect in a distillation process in terms of “an operation of extraction” (marres & weltevrede, 2013, p. 2019) that, we argue operationalises the process of creating representative and relevant data. figure 2 represents different aspects of this process. firstly, it shows the paths undertaken, from the first search words, to the documents found in the searches. secondly, it is an illustration of the tensions that exist in the attempt to present different parts of the data according to a logic of scales, from micro to macro, or local to global, and vice versa. figure 2 also visualises the results of the searches that will be presented in further details in section 3.1. 2.2 dataset 2 in dataset 2, we specifically draw upon data from two projects, led by the two co-authors in this study. both projects focus upon students’ transitions to he and consist of ethnographically generated data that build upon a case-study design (middleton, 2018; yin, 2018). the data includes audio recordings of informal discussions and interviews with 13 participants as well as document data in settings in the nation-states of sweden and italy. the data focused upon in this dataset includes three semi-structured interviews in particular, that were carried out with three participants for dataset 2, two in sweden and one in italy. the participants represent the breadth among the student group focused upon in this study (migrant and refugees) and two of them are students at the universities that were selected for the creation of dataset 1. dataset 2 corresponds to the horizontal axis of comparison in the vertical case study approach, i.e. the comparison of policy implementation at the local level, across sites. the participants’ backgrounds range from recent asylum seekers, to students migrating in early childhood and those migrating as teenagers. the set of procedures that guided the analysis of dataset 2 were inductively oriented, and aimed at identifying “patterns across the dataset” (braun & clarke, 2006; braun, clarke & weate, 2016). firstly, we familiarized with the data by revisiting the transcriptions in several collaborative steps, as a reflective process, while ensuring to remain open to alternative patterns and in search of overlapping themes in the recorded interviews. this was followed by further refinements by critically approaching our initially suggested themes, and as a result of continued analysis (braun & clarke, 2006; braun, clarke & weate, 2016), we reached the results presented in section 3. the portraits based on the demographic parameters of the cases used in this study are spelt out below. m, born in turkey, is now in her mid 20s. m came to sweden as a toddler. though she has gone through the swedish educational system, her parents were working, not aware of the expectations of parents to engage in supporting and encouraging children’s learning also from home. the parents’ level of education is very low, and according to herself, she comes from a so-called socio-economic exposed area. her education path before entering university was marked by challenges, starting with the fact that she learned to read very late. she has been awarded a bachelor degree in pedagogy, which was the programme that opened the door to university studies, but not her first choice. after m gained her bachelor degree, she did not apply for a job, and, now more confident as a university student, she decided to continue her studies. at present she is finalizing her second bachelor in the field of organization and staff development, focusing on human resources, with a strong foundation in sociology and human work science, at uni c, sweden. f, who is now in her late 30s, was born in the kurdish part of iran. she came to sweden at thirteen, after being smuggled across the border between iran and turkey, and was sent to turkish prison for a period, as she was deemed to be an adult. f was brought up by her grandparents, while her father had immigrated to sweden for political reasons. due to her grandparents preventing her from attending the village school, f only had sporadic periods at school. f studied on her own at home, and took tests up to year six without her grandparents’ knowledge. in sweden, f was placed in a preparation class for a short time, and then placed in a class according to her age and not based on skills. as f did not understand what teachers were saying, the years in secondary school were lost. after attending adult classes (upper secondary level) to get grades required for university studies, f applied for freestanding courses that were later compiled into a bachelor degree in psychology, at unia. after her bachelor exam, f has worked at the swedish migration agency for a couple of years, handling new migrants’ and immigrants’ applications, in particular due to f’s familiarity with kurdish and persian as her two first languages. at present she is finalising her master programme in public administration, management and guidance, at unia, sweden. a is a 34 years old syrian man, with refugee status who at the time was enrolled in the second year of studies at political science bachelor degree at the uni b, italy. he had arrived in italy after deserting the syrian army and a transition period in lebanon where he taught himself english and had found a way to relate to the wider world and to earn some money by teaching arabic to english speaking people. in syria, after secondary school, he had been forced to study computer science although it was not what he wanted to do. his secondary school marks did not allow him to study political science (his first choice) at university. computer science was the only possible choice. he is a gifted storyteller, interested in writing with a focus on reporting about syria. he supports himself by translating, teaching arabic, and when necessary by working as a waiter. 2.3 analysis across axes of comparison a vertical case study approach presents a set of challenges in terms of producing (re)presentation of the results of the analysis that include important details that, put together, make a coherent whole. another challenge has been to re-think the “global/local antinomy” (bartlett & vavrus, 2014, p.134) in order to become sensitive in our analysis about the ways in which policy gets its materiality not only through inscriptions, but also through encounters and practices. finally, even our focus on a particular group and selected nation states has had methodological implications, in terms of moving away from clearly bounded research sites on the one hand, and using the same bounded notions of nation-states or specific named-groups, on the other. in order to attend to the complexity of the different dimensions in a vertical case study, not least in the use of two datasets, we used several techniques. we engaged with the analyses of datasets recursively, in that they shed mutual light to one another as the themes and interesting patterns were identified in the data. for instance, the analyses of selected chunks in the data were discussed in data sessions in relation to the ways in which they contributed to shed light on the different dimensions of a widening participation agenda in both countries, i.e. how they differed or overlapped. thus, another important gap that this study aims at filling by this frontline mix-method approach is the one between the “official and local enactment [of policy] and the proliferation of unintended consequences” (fenwick, richard & sawchuk, 2011, p.115). in addition, a key particular element of such a generative and collaborative process at a vertical level, across micro and macro, and at a horizontal level, across sites, was to inductively identify key overarching meaning patterns that were further screened through thematic analysis (braun & clarke, 2006, 2019; braun., clarke, & weate, 2016). the results of the analyses are presented in section 3, in two separated sub-sections, starting with the vertical axis of comparison and a focus on dataset 1 in section 3.1, to move to the horizontal axis and the analysis of policy implementation across sites in the narratives of our selected case studies in dataset 2, in section 3.2. the final section, 4, presents an overarching discussion and some implications that go beyond the nation-state focus of sweden and italy. 3. results 3.1 tracing vertical policy entanglements (macro-micro): unia and unib in this section, we present the results from the analysis of dataset 1, i.e. provision of support services for migrants and refugees. focus lies on the vertical axis of comparison (macro-micro) starting from the “entry points” (see figure 1) in the search conducted in each university. we start by presenting the result from the entry points in first unia and then unib across the different documents to trace their entanglements across scales and sites on the vertical and horizontal axes of comparison (see figure 1 & 2). 3.1.1 unia the analysis of unia webpages shows that issues of access and inclusion overlap across different groups where the categories of “migrants”, “refugees”, “newly arrived” “asylum seekers” blend together and are not mutually exclusive in the swedish data. figure 3 is an illustration of one of the results from the entry points in the unia searches. figure 3. university a. capture of the full-size screenshot of the website (left). zooming in the central area of the page (right). translation from swedish in the text box. the page that contains information about a mentoring programme for newly arrived migrants (see figure 3) is addressed to university students that have the opportunity, by participating in this programme, to support and guide the newly arrived students in their transition to he. the students are matched with mentors who can avail of an introduction course during which they learn about leadership, communication and intercultural competence. the text also aims at making this event (that is offered regularly every term at unia) and the role as mentor into something that many students aspire to participate in. mutual learning is raised as part of the process for both mentors and “adepts”, but “på ett roligt och nyttigt sätt” (sw:1 “in a fun and useful way”). another webpage at unia provides information about initiatives for newly arrived and integration that imply a number of educational paths especially tailored for “utländska” (sw: foreigners) and “nyanlända” (sw: newly arrived). the common denominator of the ways in which such paths are presented in the documents is the use of formulations that put the spotlight on the orientation of such programmes towards specific groups and with a specific aim. for instance, the national programme for school and pre-school teachers especially targeting migrants and newly arrived are called “snabbspår” (sw: fast track). similar programmes for other professions imply that the applicant has a competence and/or a formal education from the country of origin that can be validated as the corresponding qualifications required to enter the programme according to mainstream swedish requirements. this process is called “validering av reell kompetens” (sw: validation of actual competence) and is currently a much-debated issue in the higher educational landscape in sweden because it puts the spotlight on the mainly pragmatic and instrumental view of competence, especially when it is related to a particular professional occupation. this is discussed in depth in the swedish report by the swedish council for higher education (uhr): “kan excellens uppnås i homogena studentgrupper”? (sw: can excellence be achieved in homogeneous student groups?) from 2016 (uhr, rapport 2016). in the uhr document (2016), universities are reported to have raised issues concerning the need for the establishment of a network with other authorities to allow “nyanlända” to rapidly enter he. thus, the issue of “speed” when it comes to migrant students transitioning in and out of he is prominent in the swedish data. one year after the publication of the uhr report (2016), the issue of widening participation (rather than only recruitment) was also included as part of national policy (promemoria u2017/03082/uh), as a necessity for successful transitions across educational pathways. this implies a use of resources that take cognizance of the different functions in he, from study counselling and the routines for evaluation of “actual competence” (see above) to the kinds of pedagogical efforts required in this endeavour. the 2017 promemoria caused a vivid debate among university faculty in sweden at the end of the same year and resulted in a cancellation of the proposition to change the formulation in the swedish higher education act from widening recruitment to widening participation. 3.1.2 unib in the data from unib (italy) the page about inclusion for refugees can be accessed through a search in the university website using the term “refugees” (see figure 2). the page about inclusion and reception of refugees is rather representative of unib webpages: it provides a short framing text with a section called “resources and opportunities” that includes a series of internal links to other pages (see figure 4). figure 4. university b. capture of the full-size screenshot of the website (left). translation from swedish in the text boxes (right). these links lead to pages with information about, for instance, “scholars at risk”, student scholarships, and the manifesto for the inclusive university2 (see also figure 3). when entering the page about the manifesto, words like “accoglienza” “diritti” “parità” and “inclusione” (it:3 welcoming, rights, equity and inclusion) are used to frame an agenda of inclusion, openness and widening participation in which “scuola e università offrono un'importante opportunità per i giovani rifugiati, rappresentando un passaggio fondamentale nel loro percorso di inclusione sociale”4 . according to this logic, the university (along with school) becomes the fundamental site where a “rite of passage” is made possible, from a status of “refugee” to “university student”. once in this page, a link, embedded in the body of the text, leads the reader to a pdf document, where the official manifesto is presented. the text is a programmatic document and presents general principles and outlines that the universities that participate in it share a commitment to follow. also in this document, that comprehends students, researchers and teachers with a refugee status, the terminology about hospitality, inclusion and integration dominates the text. diversity is framed as the appreciation of cultural difference and as a source of academic enrichment for the university. the manifesto ends with a quote to the canto xvii in dante’s paradiso (heaven) in the divine comedy. the chosen verses allude to dante’s exile and the challenges of being forced to leave one’s home. tu lascerai ogne cosa diletta più caramente; e questo è quello strale che l'arco de lo essilio pria saetta. tu proverai sì come sa di sale lo pane altrui, e come è duro calle lo scendere e 'l salir per l'altrui scale.5 (dante, paradiso, canto xvii) the manifesto contains a number of external links to other unhcr documents, and the inhere initiative (he supporting refugees in europe) (www.inhereproject.eu) that includes several longer documents, e.g. the guidelines for university staff members 6 and the good practice catalogue 7. the inhere recommendations represent an international document that aims at providing general guidelines to the 29 participating universities in europe. the “good practice catalogue” focuses on a number of issues and challenges for an individual with a refugee status that aim at participating in the academic life as a student, teacher or researcher. such issues include, for instance, support for recognition of previous degrees, “equitable and wide” access, financial support, language and bridging courses, integration, employment and employability, online learning, humanitarian work and collaboration among institutions and across sectors. the section about access to he links to article 27 in a eu directive (2011/95/eu) on standards for the qualification of third-country nationals or stateless persons as beneficiaries of international protection, according to which all eu member states shall provide the same conditions for access as third country nationals who are legally residents. furthermore, the section explores further the issue of access and its boundaries, more specifically: granting equitable and wide access to higher education involves more than providing tuition-free degrees, and a multitude of diverse policy tools and institutional measures exist for non-traditional or disadvantaged learners. in addition to scholarships and financial support, measures may target refugees via outreach activities, and also beyond recruitment to provide general information to the potential refugee students about the higher education system and its opportunities, consulting them through mentoring programmes and helping them navigate through the application procedures. (https://www.inhereproject.eu/wp-content/uploads/2017/08/inhere-gpc_en.pdf.pdf p.6) support to promote and grant access goes, according to this logic, beyond policy and financial support. also, in the extract above, the refugee status is framed as belonging to another dimension that should be kept separated from the “ordinary” support for the inclusion of “non-traditional” or “disadvantaged” learners. who these latter groups include more specifically is not clear in the document, but what makes the refugee status out of the ordinary is students’ need to be informed, guided and mentored to find their way in the intricate path of higher education, not least when it comes to the application procedures. issues of inclusion and support are seemingly “differently special” for refugees as compared to other groups at risk of being marginalized. such specificity for the status of a person who has been granted asylum, that was forced, against her/his will, to suddenly leave all that she/he cares for “to taste the bitterness of others’ bread” as the quote to dante’s canto succinctly illustrates, is a rather prominent feature in the unib data and the further documents at national and international level that have been accessed in the search process. this trend is also reinforced by the terminology used in the documents, irrespectively of the language variety (italian or english), that effectively draws from semantic traits that allude to hospitality, generosity, openness, as gifts from the virtuous and prosperous donor to a person facing integration difficulties at different levels. the refugee is a person whose cultural and educational background should be welcomed, respected and elevated as a source of stimulating exchange and mutual learning. to conclude, based on our comparison, the results of the analysis of dataset 1 sheds light on the different ways in which migrant students are framed as in need of support, mentoring, welcome, or other special arrangements specifically tailored for students who differ from the mainstream in the form of individuals who are purported to belong to a homogeneous, monolingual norm related to a nation-state. the virtuous character of the host country in unib (italy) in terms of hospitality and generosity is less visible in the swedish data, and is instead replaced by a logic of efficiency and speed in the process, from reception to integration. within the local higher education institution, as well as across universities and nation-states, the results show that different services and targeted programmes (addressing students’ mobility on one side, and students with a migrant background on the other) run the risk of creating parallel welcome and support approaches concerning student conditions that share the same “cultural diversity” core issue. the results also contribute to the discussion on the relevance of conceptualizing the study of diversity and of transitions to he in scalar terms (global/local) (see also enders, 2004). the emergence of global or supra-national policies like arqus and inhere, relates to a discourse of globalization and internationalization of he (along with the idea of the migrant student as in need of special support), and offers insights on the patterns created by national governmental policies in the everyday lives of students. it is to shed light on this latter dimension that we move to the next result section of this study. 3.2 tracing students’ patterns of transitions to higher education (policy implementation across sites) in this sub-section we present the results from the analysis of dataset 2, i.e. the interviews and discussions together with three students, two in sweden and one in italy. this dataset sheds light on the second question posed in this study, i.e. the investigation of the ways in which policy ideas about transition and widening participation are enmeshed in the students’ narratives and how they affect the students’ experiences of participation, normalization and marginalization in he. 3.2.1 “breaking the bubble”: transitions as conquering seemingly unreachable spaces the transitional paths to higher education involved leaving previous familiar and local spaces, out of the comfort zone, to encounter and develop an understanding of mainstream norms that underpin studies at university level. to conquer these unknown spaces, a shift from the local and micro perspective, to a new context of individual growth, i.e. learning what this implied in practice, became a prerequisite for transitions. it is evident how previous patterns of participation in education affect and become entangled with how experiences of transitioning to higher education evolve over time, as illustrated in the narratives. when m accounted for what it is like being a student at the university, she referred to her experiences when she first entered the bachelor programme in pedagogy at unia. at the time of applying, she was less sure about what to study, and just accepted the programme that she had put first on her ranking list. as she described herself as coming from a socio-economic exposed area, she attended secondary and upper secondary level in the same suburb, with the same people. the transition to university was a challenge in many ways. her account illustrated the social and contextual boundaries, and how previous life trajectories impact transition to university. m and her friends decided to apply elsewhere to break the bubble. m mentioned her transition to university studies as something very big in her world, and that now she was a grown-up. this change also implied thinking about what she was saying. now it’s the university world, now think differently, behave maturely, maybe not say everything you are thinking. it was like entering a new identity, a new role. now you need to present a façade that you don’t really want to from within. here it becomes relevant to frame m’s narrative in terms of “transition as becoming” (gale & parker, 2014, see also biesta, 2016), i.e. as a transforming process of identity requiring the learning of “the rites of passage”, to adapt to and participate according to norms for participation as a student. these changes also impacted on m’s ways of being with others, and, in order to challenge herself, she made other decisions, to learn and develop as a person. in the case of m, the “breaking the bubble” experience, has given her self-confidence. it is evident in her account that such a transition for her, has meant to learn how to be as well as behave in accordance to the norms set for university studies, norms that have a high currency in terms of societal expectations and the ways in which she could handle different situations in her everyday life. at the beginning of her university studies, she went for people like her, to stay in a comfort zone with people sharing similar life trajectories as herself. university studies entailed placing herself in uncomfortable situations, and everything felt strange. i felt as if i had come to an empty, it was like creating, finding a new empty sheet that i had to fill in, and start writing and shape this, be courageous and dare to be brave when being in uncomfortable situations that you will meet. the notion of the “empty sheet” as a metaphor for emptiness that needed to be “filled” by education, is similar to the ways in which f described her own stance to transitioning, and being responsible for creating a space for her student identity. i have learned what academic is expected to look like, you do not have to feel inferior in certain contexts, like i did before i started my studies at the university, when i was hired as a receptionist, where i noticed that people were thinking about me, you have no education, i was nothing to them. the metaphor of “breaking the bubble” into a world that is partly concealed and surrounded by an aura of enchanting mystery is relevant to understand the kinds of shortcomings, mismatches and at times, absurd situations that our cases faced when dealing with transition to he. here, gatekeeping is a key aspect of this process as it means, for the students, to be able to see opportunities as well as who, what, where and when these can become accessible realities. 3.2.2 gatekeeping on possibilities, expectations, and ambitions to overcome “insurmountable barriers” the processes of transitioning from previous education to enter an individual pathway through the university system, and with expectations of enrolment, were enmeshed with several unforeseeable challenges at a macro policy level. with less or little insight into university policy regulations, and gatekeeping enabled by existing structures, overcoming barriers, and crossing boundaries, were interdependent with encountering key persons who could support becoming a legitimate participant in higher education. starting with low grades from the upper secondary level, f met with difficulties when applying to the programme she was interested in, a bachelor programme in psychology to which she was not admitted. somebody informed her about the possibility of applying to freestanding courses that could be compiled into an exam later. f referred to this as her opportunity, an opportunity that also added extra stress during the whole study period, as she was not guaranteed a place in these courses. each term, each course was framed by this uncertainty. it would have been easier for me if i would have been admitted to the programme from the beginning. that would have reduced my stress during all terms, on my own trying to find out what courses i could apply to and keep my fingers crossed hoping to be admitted. this was the case for each course. it’s not that easy to be admitted to a programme when you don’t have a pass with distinction on everything . below, f referred back to her experiences working at the swedish migration agency, to make sense of her own challenges to transition to higher education. f recalled the many accounts given by refugees and immigrants entering sweden, and going through the required processes. f pointed to swedish bureaucracy as seriously hampering conditions and as time-consuming. the swedes don’t like when they hear an immigrant say i want to study, i want to study, i want to have a good job and things like that. i feel as if they become envious at once, you have come here and you are not expected to, you are to work with some sh-t job, you are not expected to have a better job than us, or something in that direction. f herself had experiences from this kind of attitude as well as hearing from immigrants she has met. i have experienced this kind of feeling everywhere, not only linked to myself, when an immigrant has said i want to do this, and this, many swedes often smile scornfully. maybe that is also one reason why you don’t get help, it is not, people come from other countries, do not think about educating yourself, it is all about finding a job. that is what they want, not educate them. policies of inclusion and integration were here instantiated in f’s accounts of the ways in which she and other “immigrants” were faced with the normalized expectations wherein quick access to paid labour was the primary scope of any educational path, rather than individual growth. while this may be the aim and focus for other student groups, when pondering the possibility to start an academic programme of study, in the case of marginalized groups, the choice of the academic path seemed to obey a norm where practical gains, rather than preferences, dominated. a similar issue, but with a different outcome, was present in the data from italy and in the case of a, whose tertiary education records from syria show that he attended a computer science programme. such records are misleading and far less telling and relevant when compared with the informal education and professional competences that he recently acquired. computer science was not what a wanted to study at the university. according to the syrian formal education system, a’s secondary school marks did not allow him to enrol into the political science degree at the university. therefore, in order to pursue his studies further, computer science was the only possible choice. to a, political science and media/journalism were his passion and main drive, in relation to his studies. while escaping syria and living in lebanon, a taught himself english from scratch, mainly by watching youtube videos. a developed his competences as teacher by using english to offer face-to-face, as well as online arabic lessons. at the same time, he started to develop his skills as a journalist and translator. thus, a’s set of competences that were developed through informal and nonformal learning, i.e. not to be found in a’s formal records, included his fluency in a range of languages as well as his experience as a teacher and journalist that were not recognized in the formal process of enrolment in a higher education programme, although they may be relevant features in a’s prospective study path and future career profile. however, although very long and difficult, the “processo di riconoscimento dei crediti formativi” (it: educational credits recognition process) eventually resulted in a’s enrolment in the political science programme in unib. transitioning from a life pre-he to in-he implies several challenges and constraints, that include both existential and ontological dimensions as well as pragmatic and with immediate consequences for one’s quality of life. for a this concerned practical/logistical aspects of living and studying in unib as well as aspects of academic and content organisation. the latter implied an understanding of the rationale behind the disciplinary composition of the political science bachelor degree course. unib offered a to be tutored by another student during this phase. according to a, this peer support proved to be helpful. when i applied, the tutor (from morocco) helped me a lot, we became friends. the fact that he knows my language and he knows the rules is very useful. it helped me to understand what the study plan is. it is even more useful when he studies the same course. for instance, for a it was hard to see the relevance of the enhanced role of the history area, especially history beyond contemporary and modern events. as a student with a syrian schooling background this was an unexpected feature of the course. it was relevant to be able to talk about it and to be introduced to the rationale behind this curriculum choice by somebody who could speak “his language” and would understand a’s specific schooling background. the relationship between students’ multilingual competence and transitioning towards legitimate patterns of participation in he was a recurrent theme in the accounts. 3.2.3 overlooking multilingualism as a bridge for transitioning to he in spite of the european commission’s multilingualism policy8 , now extended beyond the european languages and spatial borders to include multilingualism mirroring a society increasingly impacted by migration, there was little evidence in the narratives that indicated available local policy support practices for language support. on the contrary, the participants’ lived experience illustrated a lack of understanding at university level of the present systemic barriers and how these affected the conditions for studying negatively. f, who said she lost several years of schooling, made frequent references to how language use in university studies was challenging. the references made to lack of linguistic capital necessary for being able to succeed, were distinguished by lack of support to develop these skills, as a constant feature of her lived reality as a student. to address this barrier, f developed strategies for how to cope with the advanced scientific language in the course literature by adopting online resources. often and most of us who come from these countries are not that good in english, we don’t have it with us. so, when you finally learn swedish, you have problems with english as well. what has helped me was that now there is google translate, so today there are more things around. i don’t think i would have managed if it was like when i came to sweden, when these resources didn’t exist. i have hardly studied any english. when f compared herself with fellow students, who had attended the swedish school system, she reasoned about writing, and the higher demands she had to struggle with, and the strategies she adopted. f has been forced to put at least the double efforts into her writing, and she constantly has to remind herself about the different characters of assignments, and what is required from the different writing genres in the various assignment formats. the concrete strategy for handling the reading and understanding texts in english, became copy-and-paste in google translate. f is aware of its shortcomings but she said that without this resource, she would not have succeeded. f’s experiences highlight a consistent barrier and constraint characterized in particular by their linguistic competence, which the migrants were unable to capitalize upon due to “insufficient pedagogical and relational approaches” (harvey & mallman, 2019). furthermore, in their study, harvey and mallman (2019) highlight that the linguistic capital of new migrants is the “most difficult for [them] to realize the potential of, due to insufficient pedagogical and relational approaches within the institution” (p.10). due to the different epistemological and ontological dimensions that are at the core of different disciplines and their traditions, our analysis shows that linguistic competence (along with other, so-called general competences), is related to the voice and tradition (in terms of expectations) of a discipline. in other words, different educational programmes and the variety of course subjects and tasks therein, put different expectations on students’ capability to use an academic language e.g. crafting an argument, applying specific, technical terminology and abstract concepts. while such expectations may seem to lie beyond other, so-called, general competences, such as transcultural sensitivity or critical and analytical thinking and proficiency in language production, is central to provide the students with the “right” tools for legitimate participation in terms of “learning to talk the talk” as seedhouse succinctly puts it (seedhouse, 2008; see also lave & wenger, 1991). in the field of language studies, neologisms like superdiversity and translanguaging (vertovec, 2017, garcia, 2009) have been introduced to acknowledge the totality of the linguistic repertoire of individuals, especially when dealing with southern multilingualism, i.e. the kind of linguistic competences that are usually not recognized in institutions in the global north. however, translanguaging has lately become a synonym of a pedagogical approach, more than an analytical construct (pavlenko, 2018; bagga-gupta & messina dahlberg, 2018). this is in line with the invitation formulated by olson (2003) to formal education institutions to re-consider the ways they mediate “between the formal institutions of society law, government, economy, science and the interests, beliefs, and intentions of persons” (p.285). in unib (italy), a suggested that being able to study and perform the exam in english would help migrant students to better understand the textbooks and to stay on track with the exam schedule. luckily some professors have accepted to allow students to have their exam in english. examples include professors teaching english, history of international relations, history of political doctrines, culture and religion, sociology. but most of the teaching materials are in italian. it would be important to have handbooks in english as well. the issue of the language of study is closely related to the issue of the pace and the grading of the study career. in another case, a professor did not accept to have the exam in english although he would accept to do an oral exam. this is a step forward because even if i am able to study in italian, i am very slow when i have to write in italian so i am wasting time even if i know the answer (and therefore getting lower marks). holding an oral exam would be better than a written exam, according to a. part of the rationale for opting for an oral rather than a written exam lies in the fact that writing the answers to the exam’s questions in italian could be a slower process, even when the content of the answer is known by the student. to be “on track” and follow the programme of study in terms of, for instance, being successful to get the right amount of ects each academic semester is a constantly present dimension in the life of all students, but is especially the case when students have been granted support on account of their socio-economic status as newly arrived migrants. 3.2.4 what and who is framing the support? academic support contradictions while reasoning about ways into higher education and the support therein, m (unic, sweden) argued that more attention should be given to the individual, her or his context, to find options available or alternatives forward. i value higher education a lot, but it doesn’t mean that it shapes your identity. it should not feel impossible to study, to me it was all about finding my way into studies. it is so multidimensional, it is not possible just to connect with having low grades, and this person cannot continue studies. grades are not enough, you need a personal meeting. m’s account pointed to potential students’ ways of navigating their way to university studies (see e.g. farenga, 2018; lopez gavira & moriña, 2015; macfarlane, 2016) and the challenges on a formal, policy level when life trajectories and complexities found in transitioning to university studies do not take into account dimensions from lived realities, that diverge from the norm expressed in structures and regulations for access, and application processes as presented in the analysis of dataset 1. one example of such a mismatching was provided by a (unib, italy) who recalled that during his first academic year he was entitled to have his meals for free in the university canteens. to do so he had to certify his low income. at the end of the academic year, he was asked to pay the money back because during that year he was not able to pass enough exams, i.e. the equivalent of 24 ects. as a result, in his own words: i am paying the money back through several instalments. i paid for one part of it. every 2-3 months i am paying part of it. i probably still owe 4-600 euros. i respect the rule. this kind of support was made available when it was actually relevant to the student’s needs, but at a time (the very first year) when he was likely to be unable to comply with the conditions attached to the support that was being offered. these types of mismatches between the institutional attempt to provide support and the actual conditions of those who might benefit from such support, provide evidence of the need for he institutions to acquire a more comprehensive and complex view of the condition of migrant students and to involve them in negotiating and defining more appropriate and viable ways to address students’ equality. in the remainder of the paper, we will further discuss the issues that arose in the analyses of both datasets, as well as the pedagogical implications that may derive from the study of the entanglements of complex phenomena across analytical scales. 4. discussion the aim of this study was to investigate the intersections between policy and practice in relation to widening participation and transitions to higher education in two european countries that have faced and are still facing challenges related to the integration and inclusion of students whose identity positionings are not in conformity with the mainstream. furthermore, this study is an attempt to take a step further in finding relevant analytical paths as well as methodologies that may shed light on these phenomena by bringing together and juxtaposing macro and micro analytical levels. focus has lied on the textures of practices (gherardi, 2006) in which migrant students are entangled with in their everyday life as participants in he. we have focused upon the analyses of two datasets, one including national and international policy in selected online webpages, and the other case studies that focus on three students, enrolled in three higher educational institutions in sweden and italy. a comparison of the ways in which the selected universities in the respective countries deal with the challenges connected to the recent wave of migration related to global crises, has been incorporated in a research design in which focus lies at the intersection of a number of dimensions, one stretching across nation-states (nationally and locally) and the other across analytical scales (macro and micro). there are a number of reflections that we can make from the analyses carried out in this study. firstly, higher education pedagogy in times of transition requires to meet the kinds of societal challenges related to an (higher) education for the masses (see e.g. tight, 2019), as well as uncertain career paths and relevance especially related to diverse student cohorts. secondly, what we have framed here under the common term of migrant students is not a homogenous group in either country. and, yet, services of “special” support for students are only available for those who can produce evidence of their status of belonging to named social groups, like refugees, newly arrived migrants, but also, more in general, students whose abilities do not conform with the current formal understanding of what a “normal student” may be. thus, if we accept the theoretical argument that identity and diversity are dimensions of human life that are constantly fluctuating and related to historical and cultural understandings of what different labels are and may mean, what does that leave us with, in terms of the study of a complex assemblage like the university (bacevic, 2019)? the problem of “one higher education for all” automatically implies the creation of parallel systems within the systems, or, also, alternative solutions for the inclusion of a wide and diverse student population. we argue that one important pedagogical implication for universities globally, is to reflect on their role in the purported “widening participation” agenda that takes into account its political dimensions at the macro level (in terms of policy and infrastructures) as well as the micro level (in terms of what is done in the classroom). this study offers a substantial contribution to the research that attempts to analytically scrutinise processes of inclusion and marginalisation with a broad analytical gaze, that allowed us to analyse, among other things, a “catch 22 situation”, i.e. of mismatch between the range of support, including policies of inclusion from educational institutions, and the actual needs and challenges that the individuals belonging to the group focused upon here, meet in their transition and participation in higher education (in two european countries) (see also dangoisse, de clercq, van meenen, chartier & nils, 2019; dunwoodie et al., 2020). furthermore, sociomaterial analyses like the one conducted in this study question the assumption that policy and standards developed at a national level and implemented in the local level have separated logics or that there exists “an ontological distinction between the scalar level of the local, regional, national and global” (fenwick, edwards & sawchuck, 2011). the resulting image, thus, is far from sharp. real problems, as ingold (2018) succinctly puts it, seldom arise with a self-contained solution already inside them. real problems have no solution. for instance, in our analysis of the italian data, being granted free-meal vouchers during the first year of study, provided that the student completes a certain amount of ects, implies, rather than the solution to a problem of accessibility and equity, the creation of further issues, when the student, who was not in the position to complete the credits during the first year, must reimburse the costs of all vouchers. the problem of access to academic language is present in the analysis of both datasets in terms of a lack of formal support services. as an alternative to such a formal support, the students reported to use digital technology and internet access to find ways to eventually enter and participate in he in legitimate ways. in the case of f, in sweden, google translate was a crucial component for her to be able to compensate for a lack of support in the development of an academic language where a familiarity with english, as well as swedish, constitutes an important dimension. similarly, in the italian data, a was able to learn english by himself, watching youtube videos. furthermore, the analysis of policy ideas has also shown that the promotion of mobility (including social mobility) in he, as it is framed in the two countries, has important consequences for the inclusion of marginalized groups and their transitions to he. one important ideology that has clearly emerged in the analysis of the swedish and italian data is the notion of “speed”. speed is here understood as the sine qua non condition for transitions for all students. it is especially present in the discourse about the inclusion of marginalized groups and, more specifically, for students with a diverse ethnic background, like newly arrived migrants and refugees. while the rhetoric around speed differs in the analyses of policy in the two countries, (being rather prominent and more straightforwardly presented in the swedish data) it permeates the expectation of a widening participation agenda in both countries. speed means to get swiftly enrolled in the “right” higher educational programme, in getting university ects, a degree and eventually a relevant job position. getting everything “right” seems to be a rather prominent issue in the ways in which our cases have handled their transitions to he, in terms of a “breaking the bubble” experience. according to this logic, the university becomes an enchanted place that has the potential to open up new possibilities that arise at the students’ horizon. here, connection to key persons (often friends and/or peers at the university) that possess a practical knowledge about the most successful path towards inclusion and swift transitions to he, was of paramount importance for all our cases in both contexts. practices of transitions from a life pre-he to in-he are related to the ways in which competences (both formal and nonformal) are acknowledged in policy documents and this is framed differently in the two countries. while in sweden practices of recognition of informal competences, so-called “reell kompetens”, have been present for a long time, at least on the outset, within the scope to widen the recruitment to higher education, this feature was not present in the italian dataset 1. the notions of “reell kompetens” (sw: informal competence) and “crediti formativi” (it: formative credits) differ in that the former recognizes the informal dimension of the competence acquired, whether the latter does not. a declared policy of recognition of informal competence in sweden, however, does not mean that higher education is automatically more open or that the process of the inclusion and recognition of these competences in the student formal records is without issues. as we have seen in f’s report, this is far from the actual situation for all students that lack formal credits or marks in their lives pre-he. in fact, a eventually did manage to enrol in political science at unib (italy), albeit after a long process. in other words, standardized practices of recognition of previous competences (be them gained through formal education or not) are, in our analysis of both datasets and across countries, rather than the solution to the issue of widening recruitment and participation, the result of efforts to accommodate an ideology in which mobility and fluidity of individuals and ideas has high currency, especially in europe after the bologna agreement at the end of the 1990s. this is also related to issues of integration, wherein the understanding of diversity varies in the two contexts of this study: from openness as a result of generosity and humanitarian efforts, where diversity is welcome as a source of mutual enrichment (italy); to openness as the necessary condition to create diverse groups in academia (sweden) that will, in turn, reflect the makeup of society at large, where, according to this logic, diversity is the norm. to conclude, the analytical argument that constitutes one of the central, frontline contribution of this study, is that, in order to shed light on the kinds of opportunities that arise at the students’ horizon, it is relevant to compare and understand the ways in which policy implementation and appropriation takes place in practice and across sites. however, one important result that emerged in such an endeavour is that the boundaries between these dimensions are not to be conceived as dividing membranes that clearly mark what is global and what is local, what is macro and micro, or what is labelled as belonging to one nation-state or another. rather, the vertical case study approach used here has shown the complex entanglements of, for instance, policies of standardization, and the ways in which students make sense of situations in which such policies become practices of standardization in terms of, for instance i) standards in relation to what the students should deliver to pass a course and ii) standards of the support delivered to help the students to do so. transitions and widening participation, and the challenges that these have entailed for the students, are, in other words, not only affected by policy, they are, in fact, produced by a policy of inclusion. our analyses have confirmed and carefully illustrated the general discourse in policy where migrant students are individuals whose specific experiences make access to and participation in he distinct for them (perry & mallozzi, 2011). this deficit perspective still frames the issue as a one-way “aid” relation between the higher education organization and the students in both countries. our analysis suggests that alternative frames could become part of a more adaptable and sensitive approach to issues related to how to accommodate a diverse student population in he. such alternative frames include a shift in perspective that takes into account the students’ competences and (formal and nonformal) educational backgrounds in relation to their study and career plans. this could be achieved by including and analysing students’ autobiographical data to allow a better understanding of their communities of interest, transnational networks, and skills (see also ünlüsoy & de haan, 2020). this informal and nonformal curriculum indicates the importance for he institutions to make available appropriate autobiographical and competence recognition and validation tools. furthermore, being aware of such skills would offer university staff and faculty opportunities to find ways forward in the curricula’s international dimensions as well as to explore and to acknowledge potential contributions by migrant students to develop courses and assessment approaches across global and local dimensions. keypoints processes of inclusion in he at the macro level, inevitably incorporate processes of marginalization at the micro level. boundaries between analytical scales (macro-micro, global-local) are not hermetical. a rhetoric of speed in transitioning to and across higher education permeates the expectations of what successful transition to he may entail. proficiency in language production is crucial to provide migrant students with the “right” tools for legitimate participation. higher education institutions need to elaborate appropriate competence and autobiographical recognition and validation tools for migrant students. footnotes 1 sw: = original in swedish. 2https://www.unhcr.it/wp-content/uploads/2019/11/manifesto-delluniversita-inclusiva_unhcr.pdf 3 it: = original in italian. 4it: school and university offer an important opportunity for young refugees, thus representing a crucial transition in their path of social inclusion. (our translation) 5 it: you shall leave everything you love most dearly: /this is the arrow that the bow of exile shoots first. /you shall know the bitter (salty) taste/ of others’ bread, and know/ how hard a path it is to continuously/ descend and ascend others’ stairs. (our translation) 6 https://www.inhereproject.eu/wp-content/uploads/2018/09/inhere_guidelines_en.pdf 7https://www.inhereproject.eu/wp-content/uploads/2017/08/inhere-gpc_en.pdf.pdf 8https://ec.europa.eu/education/policies/multilingualism/about-multilingualism-policy_en list of abbreviations asrs asylum seekers and refugees he higher education inhere higher education supporting refugees in europe initiative uhr universitetsoch högskolerådet (sw: swedish council for higher education) unia university a (sweden) unib university b (italy) unic university c (sweden) unhcr united nations high commissioner for refugees unicore university corridors for refugees project references bagga-gupta, s., messina dahlberg, g. & winther, y. (2016) disabling and enabling technologies for learning in higher education for all: issues and challenges for whom? informatics, 3(21). doi: 10.3390/informatics3040021 bagga-gupta, s., messina dahlberg, g., vigmo, s. (2020). equity and social justice for whom and by whom in contemporary higher education? situated-distributed policies of inclusion/integration in sweden. learning and teaching. the international journal of higher education in the social sciences, 13 (3), 82-110. doi: 10.3167/latiss.2020.130306 bagga-gupta, s. & messina dahlberg, g. (2018). meaning-making or heterogeneity in the areas of language and identity? the case of translanguaging and nyanlända (newly-arrived) across time and space. international journal of multilingualism, 15(4), 383-411. doi: 10.1080/14790718.2018.1468446 bacevic, j. (2019). with or without u? assemblage theory and (de)territorialising the university. globalisation, societies and education, 17(1), 78-91. doi: 10.1080/14767724.2018.1498323 bartlett, l., & vavrus, f. (2014). transversing the vertical case study: a methodological approach to studies of educational policy as practice. anthropology and education quarterly, 45(2), 131-147. doi: 10.1111/aeq.12055 biesta, g. j. j. (2016). good education in an age of measurement: ethics, politics, democracy. london: routledge. braun, v., & clarke, v. (2006). using thematic analysis in psychology, qualitative research in psychology, 3(2), 77-101. doi: 10.1191/1478088706qp063oa braun, v., & clarke, v. (2019). reflecting on reflexive thematic analysis. qualitative research in sport, exercise and health, 11(4), 589-597. doi: 10.1080/2159676x.2019.1628806 braun, v., clarke, v. & weate, p. (2016). using thematic analysis in sport and exercise research. in b. smith & a. c. sparkes (eds.), routledge handbook of qualitative research in sport and exercise (pp. 191-205). london: routledge. brown, m., g., wohn, d. y., & ellison, n. (2016). without a map: college access and the online practices of youth from low-income communities. computers & education, 92-93, 104-116. doi: 10.1016/j.compedu.2015.10.001 dangoisse, f., clercq, m. d., meenen, f. v., chartier, l., & nils, f. (2020). when disability becomes ability to navigate the transition to higher education: a comparison of students with and without disabilities. european journal of special needs education, 35(4), 513-528. doi: 10.1080/08856257.2019.1708642 dunwoodie, k., kaukko, m., wilkinson, j., reimer, k., & webb, s. (2020). widening university access for students of asylum-seeking backgrounds:(mis) recognition in an australian context. higher education policy, 1-22. ecclestone, k., biesta, g. & hughes. m. (2010) (eds.) transitions and learning through the lifecourse. london: routledge eisenhart, m. (2017). a matter of scale: multi-scale ethnographic research on education in the united states. ethnography and education, 12 (2), 134-147. doi: 10.1080/17457823.2016.1257947 enders, j. (2004). higher education, internationalisation, and the nation-state: recent developments and challenges to governance theory. higher education, 47(3), 361-382. doi: 10.1023/b:high.0000016461.98676.30 eu directive 2011/95/eu. directive 2011/95/eu of the european parliament and the council. official journal of the european union l337/10 en . https://eur-lex.europa.eu/lexuriserv/lexuriserv.do?uri=oj:l:2011:337:0009:0026:en:pdf farenga, s. a. (2018). early struggles, peer groups and eventual success: an artful inquiry into unpacking transitions into university of widening participation students. widening participation and lifelong learning, 20(1), 60-78. doi: 10.5456/wpll.20.1.60 faubion, j. d., & marcus, g. e. (2009) (eds.). fieldwork is not what it used to be. learning anthropology’s methods in a time of transition. ithaca: cornell university press. fenwick, t., edwards, r., & sawchuck, p. (2011). emerging approaches to educational research. tracing the sociomaterial. london: routledge. fox, a., baker, s., charitonos, k., jack, v., & moser-mercer, b. (2020). ethics-in-practice in fragile contexts: research in education for displaced persons, refugees and asylum seekers. british educational research journal. doi: 10.1002/berj.3618. gale, t., & parker, s. (2014) navigating change: a typology of student transition in higher education. studies in higher education, 39 (5), 734-753, doi: 10.1080/03075079.2012.721351 garcía, o. (2009). bilingual education in the 21st century: a global perspective. oxford: blackwell. gherardi, s. (2017). sociomateriality in posthuman practice theory. in a. hui, t. schatzki, & e. shove (eds.), the nexus of practices. connections, constellations, practitioners (pp. 38-51). london: routledge. giroux, h. a. (2010). bare pedagogy and the scourge of neoliberalism: rethinking higher education as a democratic public sphere, educational forum, 74(3), 184–196. doi: 10.1080/00131725.2010.483897 gravett, k., kinchin, i. m., & winstone, n. e. (2020). frailty in transition? troubling the norms, boundaries and limitations of transition theory and practice. higher education research & development, 1-17. doi: 10.1080/07294360.2020.1721442 guo, s. (2015). the changing nature of adult education in the age of transnational migration: toward a model of recognitive adult education. in s. guo & e. lange (eds.), transnational migration, social inclusion and adult education (pp. 7–17). new directions for adult and continuing education, no. 146. san francisco, ca: jossey-bass. hancock, a-m. (2016). intersectionality. an intellectual history. oxford: oxford university press. harvey, a., & mallman, m. (2019). beyond cultural capital: understanding the strengths of new migrants within higher education. policy futures in education, 17(5), 657-673. doi: 10.1177/1478210318822180 hope, j. (2017). cutting rough diamonds: the transition experiences first generation students in higher education. in e. kyndt, v. donche, k. trigwell, and s. lindblom-ylänne (eds.),higher education transitions – theory and research (pp. 85-100) . new york: routledge. ingold, t. (2017). anthropology and/as education. new york: routledge. kupský, a. (2017). history and changes of swedish migration policy. journal of geography, politics and society, 7(3), 50–56. doi: 10.4467/24512249jg.17.027.7183 kyndt, e., donche, v., trigwell, k., & lindblom-ylänne, s. (2017). understanding higher education transitions: why theory, research and practice matter. in e. kyndt, v. donche, k. trigwell, and s. lindblom-ylänne (eds.),higher education transitions – theory and research (pp. 306-319) . new york: routledge. lave, j., & wenger, e. (1991). situated learning: legitimate peripheral participation. cambridge, ma: cambridge university press. law, j. (2004). after method: mess in social research. london and new york: routledge. lopez gavira, r. & moriña, a. (2015). hidden voices in higher education: inclusive policies and practices in social science and law classrooms. international journal of inclusive education, 19(4), 365-378. doi: 10.1080/13603116.2014.935812 macfarlane, k. (2016). transition through immersion in he: an evaluation of how a transition and immersion programme for school pupils embeds a culture of the university experience for key stakeholders. widening participation and lifelong learning, 18(3), 63-73. doi: 10.5456/wpll.18.3.63 mangan, d. & winter, l. a. (2017). (in)validation and (mis)recognition in higher education: the experiences of students from refugee backgrounds. international journal of lifelong education, 36(4), 486-502. doi: 10.1080/02601370.2017.1287131 marcus, g. e. (1995). ethnography in/of the world system: the emergence of multi-sited ethnography. annual review of anthropology, 24, 95-117. doi: 10.1146/annurev.an.24.100195.000523 marres, n., & weltevrede, e. (2013). scraping the social? issues in live social research. journal of cultural economy, 6(3), 313-335. doi: 10.1080/17530350.2013.772070 messina dahlberg, g. & bagga-gupta, s. (2019). on the quest to “go beyond” a bounded view of language. research in the intersections of the educational sciences, language studies and deaf studies domains 1997-2018. deafness and education international, 21(2-3), 74-98. doi: 10.1080/14643154.2018.1561782 middleton, a. (2018). reimagining spaces for learning in higher education. london: palgrave. olson, d.r. (2003). psychological theory and educational reform. how school remakes mind and society. cambridge: cambridge university press. pavlenko, a. (2018). superdiversity and why it isn’t. reflections on terminological innovations and academic branding. in s. breidbach, l. küster & b. schmenk (eds.).sloganizations in language education discourse (pp. 142-168) . bristol: multilingual matters. perry, k. h., mallozzi, c. a. (2011). ‘are you able … to learn?’: power and access to higher education for african refugees in the usa. power and education, 3(3), 249–262. doi: 10.2304/power.2011.3.3.249 promemoria u2017/03082/uh. brett deltagande i högskoleutbildning. available at: http://www.regeringen.se/rattsdokument/departementsserien-och-promemorior/2017/07/brett deltagande-i-hogskoleutbildning/ prop. 2001/02:15. den öppna högskolan. available at: https://www.riksdagen.se/sv/dokument-lagar/dokument/proposition/den-oppna-hogskolan_g p0315d2 ramsay, g., baker, s. (2019). higher education and students from a refugee background: a meta-scoping study. refugee survey quarterly. doi: 10.1093/rsq/hdy018 schneider, l. (2018). access and aspirations: syrian refugees’ experiences of entering higher education in germany. research in comparative & international education, 13(3, 457-478. doi: 1 0.1177/1745499918784764 seedhouse, p. (2008). learning to talk the talk: conversation analysis as a tool for induction of trainee teachers. in s. garton & k. richards (eds.) professional encounters in tesol (pp. 42-57). basingstoke: palgrave macmillan. taylor, c. a., & harris-evans, j. (2018). reconceptualising transition to higher education with deleuze and guattari. studies in higher education, 43, 1254–1267. doi: 10.1080/03075079. 2016.1242567 tight, m. (2019). mass higher education and massification. higher education policy, 2019(32), 93-108. doi: 10.1057/s41307-017-0075-3 trigwell, k. (2017). transitions within university: concepts and cases. in e. kyndt, v. donche, k. trigwell, and s. lindblom-ylänne (eds). higher education transitions – theory and research (pp. 121-130). new york: routledge. tsing, a. (2005). friction: an ethnography of global connection. princeton: princeton university press. tummons, j. & beach, d. (2019). ethnography, materiality, and the principle of symmetry: problematising anthropocentrism and interactionism in the ethnography of education, ethnography and education, 15(3), 286–299. doi: 10.1080/17457823.2019.1683756 uhr rapport. (2016). kan excellens uppnås i homogena grupper? en redovisning av regeringsuppdraget att kartlägga och analysera lärosätenas arbete med breddad rekrytering och breddat deltagande. available at: https://www.uhr.se/globalassets/_uhr.se/publikationer/2016/uhr-kan-excellens-uppnas-i-homogena-studentgrupper.pdf unchr (2019). manifesto dell’università inclusiva. available at: https://www.unhcr.it/wp-content/uploads/2019/11/manifesto-delluniversita-inclusiva_unhcr.pdf ünlüsoy, a., & de haan, m. (2020). turkish-dutch teens’ networked configurations for learning. frontline learning research, 8(2), 109-130. doi: 10.14786/flr.v8i2.423 vertovec, s. (2017). super-diversity. london: routledge. walsh, c. (2012). interculturalidad crítica/pedagogía decolonial. revista de educação técnica e tecnológica em ciências agrícolas, 3 (6), 25-42. yin, r. k. (2018). case study research and applications . design & methods. 6th ed. thousand oaks: sage. microsoft word li_proofs.docx frontline learning research vol. 9 no. 4 (2021) 76 91 issn 2295-3159 the frequency of emotions and emotion variability in selfregulated learning: what matters to task performance? shan li, juan zheng & susanne p. lajoie1 department of educational and counselling psychology, mcgill university, montreal, qc, canada article received 30 june 2021 / article revised 4 september 2021 / accepted 17 october / available online 5 november abstract emotion variability and its relationship to performance is an underexplored area of research both inside and outside the realm of medical education. we address this gap by examining the relative importance of the frequency of emotions and emotion variability that occurred in specific phases of self-regulated learning (srl) in predicting students’ performance. specifically, 23 medical students were recruited to complete the task of diagnosing a virtual patient in a hospitalsimulated environment. students’ facial expressions were video-recorded and were classified into basic emotions. we calculated the frequency of emotions and emotion variability at each srl phase: forethought, performance, and selfreflection. findings revealed that both the frequency of emotions and emotion variability influenced clinical reasoning performance, but they functioned differently in different srl phases. moreover, emotion variability negatively predicted performance regardless of which srl phases it was tied to. this study helps shift the focus of research from the effect of emotions on performance to the joint effect of emotion and emotion variability, which has the potential to address the inconsistency in emotion-related research findings. although we situate the study in the context of clinical reasoning, findings from this research inform the research of emotion in learning and instruction for other domains. furthermore, this study lays the foundation for future advances in emotion-related study designs since the introduction of emotion variability leaves many questions unanswered and shows promise for new research directions. keywords: emotion; emotion variability; self-regulated learning; srl phases 1 corresponding author: susanne p. lajoie, department of educational and counselling psychology, mcgill university, canada, email address: susanne.lajoie@mcgill.ca ; doi:https://doi.org/10.14786/flr.v9i4.901 li, zheng, et lajoie 77 | f l r 1. introduction medical students experience a range of emotions in clinical settings. for instance, they may experience surprise during the diagnoses of some patients with unexpected symptoms; they may experience joy and even relief when a correct diagnosis is reached, or; they may worry about their patient’s emotions and expectations when given bad news. these emotions, and others, undoubtedly affect a medical student’s thoughts, behaviours, and performance. given that emotions change over time, an inherent attribute of emotion is emotion variability (barrett, 2009), which is defined as the fluctuations in emotional states (oliver & simons, 2004). for instance, two medical students may have the same level of positive emotions but differ from one another in their emotion variability, with one student changing his or her emotions frequently and the other person changing such emotions rarely. research on the role of emotions in clinical reasoning is still nascent, but it is beginning to draw more attention from medical education researchers (artino, holmboe, & durning, 2012; lajoie, zheng, & li, 2018; mcconnell et al., 2016). nevertheless, in a review of the medical education literature, no research has been conducted on the relationship between emotion variability and student ability to diagnose patients. therefore, it is necessary to develop an understanding of how emotion and emotion variability jointly affect students’ performance. the purpose of this study is to examine the roles of emotion and emotion variability in clinical reasoning and to compare their relative importance in predicting diagnostic performance. although we situated this study in the context of clinical reasoning, this study informs future research on the role of emotion variability in learning and instruction across different disciplines. emotion variability and its relationship to performance is an underexplored area of research both inside and outside the realm of medical education (gruber, kogan, quoidbach, & mauss, 2013). moreover, this study takes the dynamic aspect of emotions (i.e., emotion variability) into account, which has the potential to address the inconsistencies in emotion-related research findings. lastly, this study presents new research directions by assessing the relative importance of the many aspects of emotions. 2. theoretical rationale 2.1 emotion-related research research on emotions has flourished in various disciplines and contexts, as reflected in the abundant classification of emotions distinguishing between basic emotions, achievement emotions, epistemic emotions, and social emotions, as well as the emergence of numerous advanced emotion detection techniques (izard, 2007; pekrun & stephens, 2010). according to ekman (1992), basic emotions (anger, disgust, fear, happiness, sadness, and surprise) are innate, universal, automatic, and fast neurological and behavioural responses to environmental challenges. for this paper, we limit our exploration to basic emotions since they are universal and may function as building blocks for more complex emotions. in addition, basic emotions are likely to be activated in medical settings that are by nature high-stakes (lajoie, zheng, li, jarrell, & gube, 2019; tracy & randles, 2011). furthermore, basic emotions can be captured in situ as behavioural measures (e.g., facial expression detection) provide an unobtrusive real-time emotional expression trace (lajoie et al., 2019). we operationalized basic emotions by their frequency as they occurred in situ since the frequency of emotions has been extensively investigated in various contexts (barrett, 2006). specifically, the frequency of emotions refers to how often a person experiences a dominant affect. the literature regarding the relationship between the frequency of emotions and performance is mixed (mcconnell et al., 2016). it is our contention that the frequency of emotions alone does not provide a complete picture of how emotions affect learning and performance. in this study, we make an effort to examine the joint effect of emotion (i.e., the frequency of emotions) and emotion variability (i.e., how much it varies) on problem-solving li, zheng, et lajoie 78 | f l r performance to advance a complete understanding of how different aspects of emotion influence diagnostic performance. 2.2 emotion variability emotion variability represents the dynamic nature of emotion above and beyond what researchers can gain from exploring the overall frequency of emotion. appraisal approaches to emotion, which proposes that emotions arise from a meaningful analysis of a situation, account for variability in emotional responses (scherer, 2005). moreover, some researchers argued that emotion variability is an inherent attribute of emotion since each emotion remains more or less constant over time (barrett, 2009). as of yet there are no empirical studies of the influence of emotion variability on clinical reasoning; however, there are studies from other fields, such as psychological health, that present two competing perspectives on whether emotion variability impedes or promotes performance (gruber et al., 2013). one perspective argues that greater variability in emotions is related to poorer outcomes (thompson, boden, & gotlib, 2017) since substantial changes in emotions deplete an individual's mental resources and passions (xu, martinez, van hoof, eljuri, & arciniegas, 2016). the other perspective contends that emotion variability has a positive influence on outcomes, especially students’ wellbeing (kashdan & rottenberg, 2010). researchers who hold this latter perspective consider emotion variability as adaptive in that individuals demonstrate that they can respond adaptively to changing environments or task conditions. our study can add new empirical evidence that would shed light on these competing perspectives. 2.3 emotion and self-regulated learning while there is a pressing need to examine the effect of emotion and emotion variability simultaneously in one study, it is noteworthy that such examinations should be situated in the context of human thinking and learning, which by its nature, is guided by contemporary learning theories (artino et al., 2012). self-regulated learning (srl) is a well-established learning theory that provides researchers with a theoretical framework to study how students regulate their behaviours, cognition, motivation, and emotion to achieve personal goals (li & lajoie, 2021; pintrich, 2000; zimmerman, 2000). according to zimmerman (2000), srl consists of three phases: forethought, performance, and self-reflection. the three phases are structurally interrelated and cyclically sustained in students’ learning process. in the forethought phase, students analyse the task, set personal goals, and plan appropriate strategies for the attainment of their predetermined goals. while the forethought phase is a preparation step for srl, the performance phase is an actual executing process in which students control and monitor their learning activities. the self-reflection phase involves students' self-judgement and reaction to the performance. as learners progress through these three phases, their performance yields specific emotions. emotions can serve to sustain and moderate the srl processes, or they may hinder such processes, which could significantly affect learning outcomes (lajoie et al., 2018; pekrun & stephens, 2010). instead of viewing emotion as a by-product of srl, ben-eliyahu (2019) considered emotions as an important target of learning, i.e., achieving emotional wellbeing in learning or problemsolving. on that account, ben-eliyahu (2019) proposed an academic emotional learning (ael) cycle to understand how learners actively monitor and modify their emotions in srl using various self-regulated emotion (sre) strategies. in line with predominant srl models, sre occurs in “a weakly sequenced recursive cycle” (ben-eliyahu, 2019, p 92), in which earlier emotional experiences update conditions on which a student experiences a subsequent emotion. emotions can and should be examined in the phases of srl in a similar way to that of behaviours, cognitive and metacognitive activities. as illustrated in the modal model of emotion (gross, 2013), a psychologically relevant situation triggers focused attention from individuals to assess the meaning of that situation in light of relevant goals and consequently generates emotional responses. as person-situation transactions unfold over time, emotion fluctuates in learning and problem-solving, depending on how individuals perceive and li, zheng, et lajoie 79 | f l r assess the evolving situation (john & gross, 2006). according to schutz and davis (2000), the forethought phase is a preparatory phase whereby students determine how to best prepare for a task and how to actually prepare for the task. therefore, the most prevalent emotions during this phase are anticipatory emotions (e.g., anxiety, distress, hope, or fear) that are triggered by a perceived threat or challenge. the forethought phase is “a period of mixed emotions” (schutz & davis, 2000, p. 249). in the performance phase, students’ emotions are affected by their subjective judgement of the task, the students’ level of preparation, and perceived ability to solve any potential problems that might occur during the task. there has been a tremendous amount of research examining the emotion of anxiety in testing situations in the phase of performance; however, few studies have documented students’ frequent emotions while solving real-world problems (pekrun et al., 2002). the emotions that are most prevalent in the self-reflection phase tend to be either harmful (e.g., disappointed) or beneficial (e.g., happy), depending on how well students performed the task and the attributions they made about the performance (schutz & davis, 2000). the interplays between emotions and srl strategies (e.g., evaluation and monitoring) are beginning to receive attention from researchers (ahmed, van der werf, kuyper, & minnaert, 2013; lajoie et al., 2019). for instance, lajoie et al. (2019) examined the differences in the co-occurrence of emotions and srl strategies between high and low performing medical students as they diagnosed a virtual patient. ben-eliyahu and linnenbrink-garcia (2013, 2015) investigated emotion regulation from within the srl perspective and its relationship with srl strategies and learning achievement. nevertheless, the relationships between the construct of emotion variability, srl phases, and learning performance have been under-researched. 2.4 the current study the relationship between the valence of emotions (i.e., positive versus negative) and performance is starting to be addressed in the context of clinical reasoning (artino, hemmer, & durning, 2011; artino et al., 2012; harley et al., 2015; lajoie et al., 2018, 2019; mcconnell & eva, 2012). it is noteworthy, however, that positive emotions are not always beneficial and negative emotions are not always detrimental (pekrun & stephens, 2010). for instance, anxiety initially increases one’s vigilance of a task situation whereby they allocate more resources on the task to perform well. therefore, it may generate more consistent results when examining the relationships between students’ performance and the frequency of emotions rather than different categories of discrete emotions. there are also several other considerations that have led us to focus on the frequency of emotions instead of discrete emotions. first, students’ emotional experiences could be better described by the frequency of emotions, given that students may experience a mixed emotional state at a certain time. moreover, there is extensive literature on the relationships between discrete emotions and performance. we contend that the frequency of emotions may provide additional insights into how the features of emotions predict performance. furthermore, examining the frequency of emotions minimizes the influences of individual differences since students respond to even the same situation with different emotions (siemer, mauss, & gross, 2007). in addition, we argue that emotion variability plays a role in determining students’ performance. to the best of our knowledge, few attempts have been made in the literature to investigate emotion variability in the context of clinical reasoning. in particular, this study is one of the first to examine how the frequency of emotions and emotion variability in different srl phases contribute to students’ diagnostic performance simultaneously. this study is exploratory in nature; as such, we cannot propose any specific hypotheses. in sum, this study aims to address this gap by answering the following research questions: do the frequency of emotions and emotion variability affect students' performance in clinical reasoning? if so, what is the relative importance of the frequency of emotions and emotion variability that occurred in specific srl phases in predicting performance? li, zheng, et lajoie 80 | f l r 3. methods 3.1 participants considering that medical students are unique in that medical schools are highly selective, one consideration for determining the sample size is the eligible participants we are able to recruit. in addition, the number of samples should not be smaller than the number of variables since we used the assessment method of averaging over orderings to compare the relative importance of predictor variables (bi, 2012). we described the method of averaging over orderings below. as a result, we recruited 23 medical students (10 males and 13 females) in a large public research university in canada. the students were in their second year of medical study, with an average age of 24.8 (sd = 3.97). they had completed a 7-week module on endocrinology, metabolism, and nutrition. therefore, the participants shared a similar level of knowledge on the problem-solving scenarios that were designed specifically for this study. it is important to note that we had six predictive variables in this study, i.e., the frequency of emotions and emotion variability in each of the three srl phases. our sample size was three times larger than the number of the predictive variables, indicating the sample size of this study was adequate. the study was approved by the local research ethics board (reb) office and consent forms made students aware of the general purpose, procedures, and data collection processes (e.g., their facial expressions would be recorded) of this study. students stated that they felt comfortable diagnosing virtual patient cases in technology-rich environments. participants had the option to withdraw their consent and discontinue their involvement at any time. all participants completed the study; however, the emotional data of two participants were not used due to technical problems. 3.2 task and learning context students were tasked with diagnosing one virtual patient case (i.e., amy case) in bioworld (lajoie, 2009). bioworld is a simulation platform that helps medical students practice clinical reasoning skills. as shown in figure 1, students were first presented with a description of the patient’s profile and symptoms, whereby they collected useful evidence for their diagnosis. they were required to formulate diagnostic hypotheses and to confirm or disconfirm their hypotheses by re-reading the case description, ordering lab tests, or searching the library within the system. in particular, clinical laboratory test results are important parameters in the decision-making process of diagnosis, and bioworld system provides a full list of laboratory tests such as urine and blood tests. the online library that is embedded in the platform provides students information on diseases and corresponding diagnostic procedures; hence, students with varying levels of declarative knowledge get support based on their needs. it is worth mentioning that students can store the results of laboratory tests or searching results in the evidence palette for future reference. in essence, the evidence palette provides a mechanism for students to monitor what and how much information they have gathered for diagnosis. after submitting a final hypothesis, students engage in the self-reflection phase of problem-solving. they check the relevance of collected evidence items and lab tests to their final hypothesis, indicating which evidence items and lab tests are neutral to, support or against their diagnosis. in addition, students are asked to rank evidence items and lab tests based on their importance to the final hypothesis. they end the task by writing a case summary of their clinical reasoning processes. li, zheng, et lajoie 81 | f l r figure 1. the interface of bioworld in this study, we chose the amy case based mainly upon students’ prior knowledge. the amy case was created by a content expert and validated by two other experts who had expertise in medicine and learning sciences. the correct diagnosis for the virtual patient case was diabetes mellitus (type i). 3.3 procedure a training session was provided to the students to familiarize themselves with the bioworld system prior to the experiment. immediately afterward, students were asked to diagnose the virtual patient independently. specifically, students performed the task individually in a lab class where they could ask research assistants for help if they encountered technical issues. students’ facial expressions were video recorded while they solved the case. participants’ problem-solving traces were automatically captured by the log files, which included the timestamp and duration of each activity. it took approximately 40 minutes for participants to complete the study. 3.4 performance, srl, and emotion measures diagnostic performance referred to the extent of evidence match between the participant’s and an expert’s solutions. in particular, an expert’s solution was pre-configured in the bioworld platform. therefore, the participant’s solution was automatically assessed in terms of evidence match with the expert’s solution as soon as he/she submitted a final diagnosis. as such, performance is a continuous variable. log files were used to categorize clinical reasoning behaviours (e.g., order lab tests, prioritize evidence items) into three srl phases based on the coding scheme developed by li et al. (2018) (see table 1). evidence table participants store their selected symptoms from the case description lab tests participants order clinical lab tests to confirm or disconfirm their hypothesis online library participants search online library within bioworld to obtain more information manage hypothesis participants propose one or more hypotheses belief meter confidence level of a correct diagnosis patient case description li, zheng, et lajoie 82 | f l r table 1 the coding scheme for analyzing srl behaviors of clinical reasoning srl phases clinical behaviours code description forethought collecting evidence items co collecting evidence items from the patient description by recalling one’s prior knowledge pertaining to the symptoms performance raising/managing hypotheses ra outlining a single or multiple diagnostic hypothesis based on the collected evidence items adding tests ad conducting medical lab tests searching library se searching for particular information in the library for additional explanations self-reflection categorizing evidence/results ca checking the relevance of evidence items and lab test results towards a specific hypothesis (i.e., whether the evidence/tests in support, against, or neutral of one hypothesis) linking evidence/results li justifying the probability of a hypothesis being correct to the disease prioritizing evidence/results pr ranking evidence items and lab test results according to their importance to a hypothesis summarization for final diagnosis su making the final diagnosis by writing a summarization students’ emotions expressed in each srl phase were analysed using facereader 6 software (den uyl, van kuilenburg, & lebert, 2005). facereader is a facial expression recognition software that can classify students' facial expressions into one of the six basic emotions: happy, sad, angry, surprised, scared, and disgusted. in facereader, each emotion is expressed as a value between 0 and 1 for a target face. an expression is scored “dominant” when its intensity is significantly higher than that of all others. the output of facereader consists of two columns of information, i.e., timestamps and corresponding emotional states. as such, researchers can identify the duration of an emotional state and when it changes to a different emotion over the course of problem-solving. given that facereader detects emotions at a fine-grained temporal resolution in an automated fashion, it provides researchers a unique tool to examine the frequency of emotions and emotion variability that other emotion measures cannot afford. the software analyses facial expressions with an accuracy of 95% (noldus information technology, 2015). the accuracy of this software has also been verified by several researchers in empirical psychological studies (chentsova-dutton & tsai, 2010; harley et al., 2015). the frequency of emotions and emotion variability were calculated for each srl phase. specifically, we calculated the frequency of emotions by dividing the number of emotions an individual expressed in a specific srl phase by the length of that phase. moreover, we calculated the probability distribution of the six types of basic emotions for each of the three srl phases. for instance, the probability of happiness is .25 if an individual experiences 100 emotions in total in the performance phase, among which 25 are happy. afterward, emotion variability was calculated using the shannon entropy formula (jack, garrod, & schyns, 2014): h (p1, …, pa) = ∑ 𝑝i log (𝑝i)! "#$ li, zheng, et lajoie 83 | f l r where 𝑝i is the probability of emotional state i appearing in a certain stream of emotions. shannon entropy provides a mathematical way to quantify the randomness of emotional states. the minimum value of emotion variability is zero, indicating that students’ emotion never varies. the maximum value of emotional variability is 2.58 (the base-2 logarithm of the six possible emotions) indicating the highest variability. 3.5 data analysis to address our research questions, we applied the assessment method of averaging over orderings proposed by lindeman, merenda, and gold (lmg) (lideman, merenda, & gold, 1980) to compare the relative importance of the six predictive variables (i.e., emotion frequency – forethought, emotion frequency – performance, emotion frequency – reflection, emotion variability – forethought, emotion variability – performance, emotion variability – reflection) in predicting diagnostic performance. in particular, the lmg approach examines the proportion of variance explained by each variable of interest, considering both its direct effect and its effect when combined with the other variables. the lmg method also takes the dependence on orderings into consideration instead of adding the predictors in the model randomly (lideman et al., 1980). the formulae denoted as lmg can be written as: 𝐿𝑀𝐺(𝑥%) = 1 𝑝! 1 𝑠𝑒𝑞𝑅&6({𝑥%})9𝑟; ' )*'+,-!-"./ (1) 𝑠𝑒𝑞𝑅&6({𝑥%})9𝑆%(𝑟); = 𝑅&6{𝑥%} ∪ 𝑆%(𝑟); − 𝑅&(𝑆%(𝑟)) (2) 𝑆%(𝑟) represents the set of predictors entered the linear model before the predictor 𝑥% in the order 𝑟, while 𝑠𝑒𝑞𝑅&6({𝑥%})9𝑆%(𝑟); represents the portion of 𝑅& (sum of squares) allocated to predictor 𝑥% in the order 𝑟. for example, for the six explanatory variables (p = 6), there are 720 different orderings (6! = 6 x 5 x 4 x 3 x 2 x 1 = 720) and 720 different estimations, i.e., sequential sum of squares. the relative importance of the explanatory variable is the mean of the 720 estimations. a review of different methods for assessing relative importance suggests that lmg is one of the most computer-intensive methods and is widely adopted by researchers (grömping, 2006). to assess if a variable is clearly different from the others in terms of relative importance, we used bootstrap percentile cis (confidence intervals) to assess the variability of the estimates. the confidence intervals for differences show if differences in contributions can be considered statistically significant. specifically, 1000 bootstrap samples were requested in analyses to estimate the cis for differences between relative contributions. 4. results 4.1 descriptive analysis the mean values of the frequency of emotions, emotion variability, and performance were shown in table 2. we also examined the bivariate correlations between those variables. the results of the descriptive analysis showed that students were generally in neutral affective states in the forethought phase, while they expressed their emotions strongly in the performance and self-reflection phases. the frequency of emotions in the performance phase highly correlated with the frequency of emotions in the self-reflection phase. moreover, the emotion variability in the performance phase highly correlated with that in the self-reflection phase. regarding the correlations between emotional variables and performance, the frequency of emotions in the self-reflection phase was only the factor that significantly associated with performance. li, zheng, et lajoie 84 | f l r table 2 descriptive analysis of the frequency of emotions, emotion variability, and performance variable mean sd 1 2 3 4 5 6 7 1.emotion-forethought 2.10 1.95 2.emotion-performance 11.64 7.60 .12 3.emotion-self-reflection 6.38 6.57 .39 .46* 4.emotion variability-forethought 1.23 .50 .25 .19 .17 5.emotion variability-performance 1.56 .41 .24 .13 .24 .24 6.emotion variability-self-reflection 1.41 .48 .02 .17 .32 .21 .68** 7.performance 53.00 13.82 .07 .04 .49* -.36 -.11 -.08 note: the values in the last seven columns are pearson correlations. the symbols * and ** indicates that correlation is significant at the .05 and .01 levels (2-tailed), respectively. in addition, we calculated the average durations of the three srl phases and the mean values of the frequency of the six basic emotions, for the sake of helping readers gain a comprehensive understanding of students’ emotional experience (see table 3). in general, the frequencies of all the six types of emotions increased from the forethought phase to the performance phase, and then decreased as students moved from the performance phase to the self-reflection phase. table 3 the average duration of the three srl phases and the mean values of the frequency of the six basic emotions duration sd angry disguste d happ y surprise d sad scared forethought 3.15 1.99 0.84 0.31 0.14 0.62 0.07 0.10 performance 17.06 7.83 5.63 2.18 0.99 2.20 0.31 0.33 self-reflection 8.76 5.53 2.68 0.94 0.62 1.87 0.10 0.16 note: the first two columns show the average duration of each srl phase and corresponding standard deviation of the durations. the values in the remaining columns are the mean frequencies of a specific emotion, i.e., the number of a specific emotion per minute. 4.2 relative importance analysis the total proportion of variance in students’ diagnostic performance explained by the six emotion-related predictors was 50.67%. according to cohen (1988), the effect size is large (> .26). considering that clinical reasoning is an active and constructive process whereby students need to manipulate their cognitive, metacognitive, affective, and motivational aspects of problem-solving, this result has demonstrated the crucial roles of the frequency of emotions and emotion variability in the context of diagnosing patients. in particular, the results in table 4 showed that the frequency of emotions in the forethought and performance phases of self-regulated learning negatively predicted performance (r = -.66 and r = -.32, respectively), while the frequency of emotions in the self-reflection phase positively predicted performance (r = 1.55). in general, emotion variability negatively predicted performance regardless of which srl phases it was tied to. li, zheng, et lajoie 85 | f l r table 4 relative importance of emotion and emotion variability in predicting performance variable r2 (lmg) 95% ci average coefficient (standardized coefficient) emotion-forethought .0114 [.0024, .1322] -.66 (-.09) emotion-performance .0172 [.0052, .1076] -.32 (-.17) emotion-self-reflection .3071 [.0646, .6325] 1.55 (.74) emotion variability-forethought .1458 [.0258, .3371] -10.70 (-.39) emotion variability-performance .0102 [.0061, .1664] -.70 (-.02) emotion variability-self-reflection .0150 [.0106, .2738] -5.33 (-.19) note: emotion-forethought, emotion-performance, and emotion-self-reflection refer to the frequency of emotions in the forethought, performance, and self-reflection phase, respectively; lmg, also known as r2, is an index of the relative importance of predictive variables in predicting the outcome variable; 95% ci = 95% confidence interval. similar to conventional regression coefficients, average coefficients indicate the relationship between each of the predictors and a dependent variable. however, the predictors must be uncorrelated in regression models, while the predictors do not have to meet the multicollinearity assumption when estimating their coefficients with the lmg method. the lmg method takes the dependence of the predictors into account and produces what are called average coefficients. moreover, the results in table 4 and figure 2 indicated that the frequency of emotions in the self-reflection phase (r2 = .3071) and emotion variability in the forethought phase (r2 = .1458) were the two most important factors in predicting performance. the confidence intervals for differences between relative contributions, as shown in table 5, revealed that the difference in the frequency of emotion between the self-reflection phase and the forethought phase, as well as the difference between the selfreflection phase and the performance phase, were statistically significant. specifically, the frequency of emotions in the self-reflection phase was significantly more important than that of the other two srl phases in predicting clinical reasoning performance. furthermore, there were no significant differences between the three srl phases in terms of students’ emotion variability. there were also no significant differences between the frequency of emotions and emotion variability regardless of the srl phases. figure 2. visualization of the relative importance of the factors in predicting performance. note: a1e, a2e, and a3e refer to the frequency of emotions in the forethought, performance, and self-reflection phase, respectively; a1ev, a2ev, and a3ev refer to emotion variability in the forethought, li, zheng, et lajoie 86 | f l r performance, and self-reflection phase, respectively; diagnostic performance was indicated by the evidence match between students’ and experts’ solutions. the total proportion of variance explained by the model was 50.67%. table 5 bootstrap confidence intervals for differences between relative contributions difference 95% ci a1e-a2e -.0058 [-.0860, .1014] a1e-a3e -.2958 * [-.6101, -.0245] a1e-a1ev -.1344 [-.3207, .0548] a1e-a2ev .0012 [-.1366, .0939] a1e-a3ev -.0036 [-.2606, .0988] a2e-a3e -.2900 * [-.6071, -.0250] a2e-a1ev -.1286 [-.3051, .0266] a2e-a2ev .007 [-.1302, .0661] a2e-a3ev .0022 [-.2483, .0554] a3e-a1ev .1614 [-.1887, .5695] a3e-a2ev .2970 [-.0410, .5905] a3e-a3ev .2921 [-.1788, .5862] a1ev-a2ev .1356 [-.0802, .2977] a1ev-a3ev .1307 [-.1643, .2821] a2ev-a3ev -.0049 [-.2028, .0997] note: * indicates that ci for difference does not include 0. a1e, a2e, and a3e refer to the frequency of emotions in the forethought, performance, and self-reflection phase, respectively; a1ev, a2ev, and a3ev refer to emotion variability in the forethought, performance, and self-reflection phase, respectively 5. discussion this research revealed that the frequency of emotions in both the forethought and performance phases of srl negatively predicted performance in clinical reasoning, while the frequency of emotions in the self-reflection phase positively predicted performance. a potential explanation is that a high frequency of emotions in the forethought and performance phases may interfere with students’ cognition and motivation. it is highly possible that students who failed to regulate their emotions in these two phases could not concentrate on their decision-making and executing behaviours. an exception is the frequency of emotions in the self-reflection phase, a period when students receive and make sense of the collected evidence and test results. according to schutz and davis (2000), emotions in the selfreflection phase can be due to cognitive appraisals about the progress of their diagnosis, which influence how students adjust their problem-solving strategies and final performance. moreover, the work of beneliyahu (2019) offered another explanation for these findings. specifically, students’ affective inclination moderates how activity emotions are linked to performance. however, there is a possibility that the moderation effect declines as learning or problem-solving process unfolds. therefore, the effect of affective inclination on the emotions in the self-reflection phase is different from the forethought and performance phases. li, zheng, et lajoie 87 | f l r in addition, this study found that emotion variability had a negative influence on diagnostic performance, regardless of srl phases it was tied to. this finding highlights the importance of keeping a stable emotional state to guarantee high performance, which has been corroborated by previous research (thompson et al., 2017). this finding also has implications for teaching and learning regarding emotion regulation. as argued by ben-eliyahu (2019), “teaching learners to regulate emotions to facilitate potential benefits during intellectual work can and should be an explicit target during learning” (p. 86). it must be noted that emotions occur in specific situations as responses to the features of external environment and internal appraisals (gross, 2013). the finding of this study is influenced by the features of clinical reasoning and the unique characteristics of medical students. therefore, this finding needs to be interpreted appropriately in other contexts. moreover, the frequency of emotions and emotion variability were found to jointly affect diagnostic performance. this result is not surprising since prior research has found both variables are important when considered independently (barrett, 2006; gruber et al., 2013; kashdan & rottenberg, 2010; mcconnell et al., 2016; thompson et al., 2017). a unique contribution of this study is that we examined the two variables simultaneously in one study, which could help gain a better understanding of the role of emotions in learning. in addition to the novelty of the research context, this study also provided methodological insights regarding the measurement of emotions at a fine-grained temporal resolution and the analysis of emotion variability. further insights are that the two most important factors in predicting performance were the frequency of emotions in the self-reflection phase and emotion variability in the forethought phase. the frequency of emotions in the self-reflection phase may be due to students' awareness that their final diagnosis is correct. variability in the forethought phase may be due to the open consideration of all possibilities for a diagnosis, making students less stable or secure in their ability and perhaps more emotionally variable. however, future research is needed to shed light on how and why these two factors affect performance. in sum, findings from this research alert medical researchers to the important role that emotion plays in clinical reasoning. medical teachers and students alike should be aware of the close ties between the frequency of emotions, emotion variability, and diagnostic performance. moreover, findings from this research resonate with the call for more studies on emotion-related regulation in the context of clinical reasoning. medical students would benefit from the training programs regarding how to redirect, control, and modify emotional arousals to perform adaptively in emotionally arousing situations. for instance, this study suggested that restraining substantial changes in emotions across the whole diagnostic process and reappraising emotions in the early stages of clinical reasoning (i.e., forethought and performance phase) may lead to better diagnostic reasoning outcomes. this finding partially aligns with prior work on emotion regulation that suppression was found positively related to positive deactivated emotions in certain contexts, and positive deactivated emotions (e.g., relief) were indicators of high performance (ben-eliyahu & linnenbrink-garcia, 2013). it is noteworthy that researchers’ interest in emotion-related regulation has grown enormously in recent years, and there is substantial research on emotion regulation strategies (eisenberg et al., 2018; leblanc, essau, & ollendick, 2017). such strategies include, but are not limited to, changing how one appraises a situation, regulating the demands of familiar settings, selecting adaptive response alternatives, encoding of internal emotion cues, and accessing to coping resources (siemer et al., 2007). it would be fruitful for future studies to examine which types of emotion regulation strategies are effective in clinical reasoning, taking the particular context and the characteristics of individuals into account. 6. conclusion this study examined the relative importance of the frequency of emotions and emotion variability in the three srl phases (i.e., forethought, performance, and self-reflection) in predicting diagnostic performance. we found that both the frequency of emotions and emotion variability matter to students' performance, and they functioned differently in the clinical reasoning process. this study helps shift the focus of research from the role of emotions in learning alone, towards the joint effects of li, zheng, et lajoie 88 | f l r emotions and emotion variability on performance. this study has also methodological insights with regards to the analysis of emotion variability and the relative importance of different emotion features. it is important to note that this study is not without limitations. we examined the influence of the frequency of emotions in each srl phase on diagnostic performance rather than the frequency of each specific emotion, for example, happy and angry. it is possible that one type of emotion has greater predictive power than another emotion. in addition, students experience a wide range of emotions in academic settings that may go beyond the scope of basic emotions. therefore, it is necessary to advance methodological innovations in capturing different categories of emotions in situ for future studies. furthermore, students may have a mixed emotional state, whereas we recognized one dominant emotion at a given time due to the constraints of emotion measures. lastly, a larger cohort of medical students is needed to verify the generalizability of our findings as they diagnose patients with different levels of complexities. as a closing remark, this study lays the foundation for future advances in emotion-related study designs since the introduction of emotion variability leaves many questions unanswered and shows promise for new research directions. for example, does affective inclination affect emotion variability (ben-eliyahu, 2019)? what are the differences between emotion regulation and the regulation of emotion variability? are appraisals sufficient causes of emotion variability? how does emotion variability affect learning in the long run? it would also be interesting to take the valence of emotions into account when investigating the frequencies and variabilities of emotions. for instance, the frequency of positive emotions or the variability of negative emotions may tell an in-depth story about students’ performance differences. another related direction for future research is to assess the relative importance of the valence, frequency, and variability of emotions to learning performance. in addition, it is promising to examine the variability in emotional intensity and its relationships with cognition, metacognition, and learning performance. li, zheng, et lajoie 89 | f l r key points this study is one of the first to explore the variability aspect of emotions in srl (selfregulated learning), which has the potential to open new research directions. we identified students’ real-time emotions from their facial expressions as they diagnosed virtual patients in a computer-simulated environment. we calculated the frequency of emotions and emotion variability at each srl phase: forethought, performance, and self-reflection. we used the assessment method of averaging over orderings to examine the joint effects of emotion and emotion variability on clinical reasoning performance. we found that both emotion and emotion variability affected students’ performance and they functioned differently in the srl process. emotion variability negatively predicted performance regardless of which srl phases it was tied to. acknowledgments this research was funded by the fonds de recherche du québec société et culture (frqsc) and the social sciences and humanities research council of canada (sshrc) references ahmed, w., van der werf, g., kuyper, h., & minnaert, a. (2013). emotions, self-regulated learning, and achievement in mathematics: a growth curve analysis. journal of educational psychology, 105(1), 150–161. http://dx.doi.org/10.1037/a0030160 artino, a. r., hemmer, p. a., & durning, s. j. (2011). using self-regulated learning theory to understand the beliefs, emotions, and behaviors of struggling medical students. academic medicine, 86(10), s35–s38. https://doi.org/10.1097/acm.0b013e31822a603d artino, a. r., holmboe, e. s., & durning, s. j. (2012). can achievement emotions be used to better understand motivation, learning, and performance in medical education? medical teacher, 34(3), 240–244. https://doi.org/10.3109/0142159x.2012.643265 barrett, l. f. (2006). solving the emotion paradox: categorization and the experience of emotion. personality and social psychology review, 10(1), 20–46. https://doi.org/10.1207/s15327957pspr1001_2 barrett, l. f. (2009). variety is the spice of life: a psychological construction approach to understanding variability in emotion. cognition and emotion, 23(7), 1284–1306. https://doi.org/10.1080/02699930902985894 ben-eliyahu, a. (2019). academic emotional learning: a critical component of self-regulated learning in the emotional learning cycle. educational psychologist, 54(2), 84–105. https://doi.org/10.1080/00461520.2019.1582345 ben-eliyahu, a., & linnenbrink-garcia, l. (2013). extending self-regulated learning to include selfregulated emotion strategies. motivation and emotion, 37(3), 558–573. https://doi.org/10.1007/s11031-012-9332-3 ben-eliyahu, a., & linnenbrink-garcia, l. (2015). integrating the regulation of affect, behavior, and cognition into self-regulated learning paradigms among secondary and post-secondary students. metacognition and learning, 10(1), 15–42. https://doi.org/10.1007/s11409-014-9129-8 bi, j. (2012). a review of statistical methods for determination of relative importance of correlated predictors and identification of drivers of consumer liking. journal of sensory studies, 27(2), 87– 101. https://doi.org/10.1111/j.1745-459x.2012.00370.x chentsova-dutton, y. e., & tsai, j. l. (2010). self-focused attention and emotional reactivity: the role of culture. journal of personality and social psychology, 98(3), 507–519. https://doi.org/10.1037/a0018534 li, zheng, et lajoie 90 | f l r cohen, j. (1988). statistical power analysis for the behavioral sciences (2nd ed.). lawrence erlbaum associates. den uyl, m., van kuilenburg, h., & lebert, e. (2005). facereader: an online facial expression recognition system. in proceedings of the 5th international conference on methods and techniques in behavioral research (vol. 2005, pp. 589–590). eisenberg, n., spinrad, t. l., & valiente, c. (2018). emotion‐related self‐regulation and children’s social, psychological, and academic functioning. in diversity in harmony-insights from psychology: proceedings of the 31st international congress of psychology (pp. 268–295). wiley online library. ekman, p. (1992). an argument for basic emotions. cognition and emotion, 6(3–4), 169–200. https://doi.org/10.1080/02699939208411068 grömping, u. (2006). relative importance for linear regression in r: the package relaimpo. journal of statistical software, 17(1), 1–27. https://doi.org/10.18637/jss.v017.i01 gross, j. j. (1998). the emerging field of emotion regulation: an integrative review. review of general psychology, 2(3), 271–299. https://doi.org/10.1037/1089-2680.2.3.271 gross, j. j. (2013). emotion regulation: conceptual and empirical foundations. in j. j. gross (ed.), handbook of emotion regulation (2nd ed., pp. 3–20). guilford publications. gruber, j., kogan, a., quoidbach, j., & mauss, i. b. (2013). happiness is best kept stable: positive emotion variability is associated with poorer psychological health. emotion, 13, 1–6. http://dx.doi.org/10.1037/a0030262 harley, j. m., bouchet, f., hussain, m. s., azevedo, r., & calvo, r. (2015). a multi-componential analysis of emotions during complex learning with an intelligent multi-agent system. computers in human behavior, (48), 615–625. https://doi.org/10.1016/j.chb.2015.02.013 izard, c. e. (2007). basic emotions, natural kinds, emotion schemas, and a new paradigm. perspectives on psychological science, 2(3), 260–280. https://doi.org/10.1111/j.1745-6916.2007.00044.x jack, r. e., garrod, o. g. b., & schyns, p. g. (2014). dynamic facial expressions of emotion transmit an evolving hierarchy of signals over time. current biology, 24(2), 187–192. https://doi.org/10.1016/j.cub.2013.11.064 john, o. p., gross, j. j. (2006). individual differences in emotion regulation. in j. j. gross (ed.), handbook of emotion regulation (1st ed., pp. 351–372). guilford publications. kashdan, t. b., & rottenberg, j. (2010). psychological flexibility as a fundamental aspect of health. clinical psychological review, 30(7), 865–878. https://doi.org/10.1016/j.cpr.2010.03.001 lajoie, s. p. (2009). developing professional expertise with a cognitive apprenticeship model: examples from avionics and medicine. in k. a. ericsson (ed.), development of professional expertise: toward measurement of expert performance and design of optimal learning environments (pp. 61–83). new york: cambridge university press. lajoie, s. p., zheng, j., & li, s. (2018). examining the role of self-regulation and emotion in clinical reasoning: implications for developing expertise. medical teacher, 40(8), 842–844. https://doi.org/10.1080/0142159x.2018.1484084 lajoie, s. p., zheng, j., li, s., jarrell, a., & gube, m. (2019). examining the interplay of affect and self regulation in the context of clinical reasoning. learning and instruction, 101219. https://doi.org/10.1016/j.learninstruc.2019.101219 leblanc, s., essau, c. a., & ollendick, t. h. (2017). emotion regulation: an introduction. in c. a. essau, s. leblanc, & t. h. ollendick (eds.), emotion regulation and psychopathology in children and adolescents. (1st ed., pp. 3–17). oxford: oxford university press. https://doi.org/10.1093/med:psych/9780198765844.003.0001 li, s., & lajoie, s. p. (2021). cognitive engagement in self-regulated learning: an integrative model. european journal of psychology of education, 1–20. https://doi.org/10.1007/s10212-021-00565x li, s., zheng, j., poitras, e., & lajoie, s. (2018). the allocation of time matters to students’ performance in clinical reasoning. in r. nkambou, r. azevedo, & j. vassileva (eds.), lecture notes in computer sciences (pp. 110–119). springer international publishing ag, part of springer nature. https://doi.org/10.1007/978-3-319-91464-0_11 li, zheng, et lajoie 91 | f l r lideman, r., merenda, p., & gold, r. (1980). introduction to bivariate and multivariate analysis scott. scott foresman: glenview, il, usa. mcconnell, m. m., & eva, k. w. (2012). the role of emotion in the learning and transfer of clinical skills and knowledge. academic medicine, 87(10), 1316–1322. https://doi.org/10.1097/acm.0b013e3182675af2 mcconnell, m. m., monteiro, s., pottruff, m. m., neville, a., norman, g. r., eva, k. w., & kulasegaram, k. (2016). the impact of emotion on learners application of basic science principles to novel problems. academic medicine, 91(11), 58–63. https://doi.org/10.1097/acm.0000000000001360 noldus information technology. (2015). reference manual: facereader version 6.1. wageningen, the netherlands: noldus information technology international headquarters. oliver, m. n. i., & simons, j. s. (2004). the affective lability scales: development of a short-form measure. personality and individual differences, 37, 1279–1288. https://doi.org/10.1016/j.paid.2003.12.013 pekrun, r., goetz, t., titz, w., & perry, r. p. (2002). academic emotions in students’ self-regulated learning and achievement: a program of qualitative and quantitative research. educational psychologist, 37(2), 91–105. https://doi.org/10.1207/s15326985ep3702_4 pekrun, r., & stephens, e. j. (2010). achievement emotions: a control-value approach. social and personality psychology compass, 4(4), 238–255. https://doi.org/10.1111/j.17519004.2010.00259.x pintrich, p. r. (2000). the role of goal orientation in self-regulated learning. in m. boekaerts, p. r. pintrich, & m. zeidner (eds.), handbook of self-regulation (1st ed., pp. 451–502). san diego, ca: us: academic press. scherer, k. r. (2005). what are emotions? and how can they be measured? social science information, 44(4), 695–729. https://doi.org/10.1177/0539018405058216 schutz, p. a., & davis, h. a. (2000). emotions and self-regulation during test taking. educational psychologist, 35(4), 243–256. https://doi.org/10.1207/s15326985ep3504_03 siemer, m., mauss, i., & gross, j. j. (2007). same situation-different emotions: how appraisals shape our emotions. emotion, 7(3), 592–600. https://doi.org/10.1037/1528-3542.7.3.592 thompson, r. j., boden, m. t., & gotlib, i. h. (2017). emotional variability and clarity in depression and social anxiety. cognition and emotion, 31(1), 98–108. https://doi.org/10.1080/02699931.2015.1084908 tracy, j. l., & randles, d. (2011). four models of basic emotions: a review of ekman and cordaro, izard, levenson, and panksepp and watt. emotion review, 3(4), 397–405. https://doi.org/10.1177/1754073911410747 xu, s., martinez, l. r., van hoof, h., eljuri, m. i., & arciniegas, l. (2016). fluctuating emotions: relating emotional variability and job satisfaction. journal of applied social psychology, 46, 617– 626. https://doi.org/10.1111/jasp.12390 zimmerman, b. j. (2000). attaining self-regulation: a social cognitive perspective. in m. boekaerts, p. r. pintrich, & m. zeidner (eds.), handbook of self-regulation (1st ed., pp. 13–39). san diego (ca): academic press. https://doi.org/10.1016/b978-012109890-2/50031-7 codepen beck publication frontline learning research vol.8 no. 6 (2020) 1 37 issn 2295-3159 ensuring content validity of psychological and educational tests – the role of experts klaus becka ajohannes gutenberg-university mainz, germany article received 18 june 2019/ revised 29 july 2020/ accepted 14 august/ available online 4 september abstract many test developers try to ensure the content validity of their tests by having external experts review the items, e.g. in terms of relevance, difficulty, or clarity. although this approach is widely accepted, a closer look reveals several pitfalls need to be avoided if experts’ advice is to be truly helpful. the purpose of this paper is to exemplarily describe and critically analyse widespread practices of involving experts to ensure the content validity of tests. first, i offer a classification of tasks that experts are given by test developers, as reported on in the respective literature. second, taking an exploratory approach by means of a qualitative meta-analysis, i review a sample of reports on test development (n = 72) to identify the common current procedures for selecting and consulting experts. results indicate that often the choice of experts seems to be somewhat arbitrary, the questions posed to experts lack precision, and the methods used to evaluate experts’ feedback are questionable. third, given these findings i explore in more depth what prerequisites are necessary for their contributions to be useful in ensuring the content validity of tests. main results are (i) that test developers, contrary to some practice, should not ask for information that they can reliably ascertain themselves (“truth-apt statements”). (ii) average values from the answers of the experts which are often calculated rarely provide reliable information about the quality of test items or a test. (iii) making judgements about some aspects of the quality of test items (e.g. comprehensibility, plausibility of distractors) could lead experts to unreliable speculations. (iv)when several experts respond by giving incompatible or even contradictory answers, there is almost always no criterion enabling a decision between them. in conclusion, explicit guidelines on this matter need to be elaborated and standardised (above all, by the aera, apa, and ncme “standards”). keywords: test development; content validity; expert; qualitative meta-analysis info corresponding author email: beck@uni-mainz.de doi: https://doi.org/10.14786/flr.v8i6.517 “if all experts are united, caution should be exercised.” albert einstein 1. introduction in the latest edition of the standards for educational and psychological testing (aera, apa & ncme, 2014) the multi-dimensionality of test validity is stressed once again. validity is considered not merely an attribute of a single test or a given assessment procedure, but rather of the interpretation of test scores 1, be it in the context of diagnosing individuals, identifying group properties, determining causal relationships between latent traits, or developing practical educational measures. when taking this approach to validity several facets emerge which might affect the accuracy of conclusions drawn from a measurement result. the issue of validity has been explored extensively in the literature. in this paper, one particular facet of validity essential to adequate test score interpretations is explored: content validity.2 a basic tenet of test development is that a test instrument should measure what it claims to measure. therefore, the content of such instruments has to reflect or correspond as accurately as possible to the real-world issues it is intended to assess. if this condition is not satisfied, interpretation of the outcomes of an assessment will be inappropriate. the difficulty begins with defining what is to be measured, especially in regard to latent attributes, which is usual in the fields of psychology and education. following boring (1923) some psychologists have taken the view that the latent object measured is defined by what the assessment instrument measures—a variant of early behaviourism also held in a certain sense by psychological constructionists in making use of operational definitions (e.g., van der maas, kan, & borsboom, 2014). with good reason it was argued early on in seminal debates that this interpretation is not acceptable (e.g., miles, 1957). of course, test developers have at least a rough idea of the object they want to measure when creating a new test. that idea represents or designates the content in relation to which validity only can be examined. the clearer this idea, the more precisely content validity can be investigated.3 and the greater the validity of test content being assessed, the lower the risk of inadequate interpretations of test scores (given the satisfactory quality of all other properties that a test should exhibit). thus, content validity is an essential attribute of all psychological measurement instruments (lynn, 1986), and understanding how it can be achieved is of paramount importance. for this reason, test developers are interested in enhancing the content validity of their instruments. it is widely accepted as evidence of having reached this goal if experts agree that this claim has been met. in section 2, i give a brief overview of the current practices of involving experts in the development of tests. here i explore the rationale for involving experts in the process of ensuring content validity of tests or test items. additionally, i describe the cognitive processes that experts have to go through when they do their job. in section 3, i offer a categorisation schema for the numerous tasks test developers ask experts to perform. i conduct a qualitative meta-analysis of a sample of 72 reports on test development and present a summary of the procedures test developers took when consulting experts to ensure the validity of the content of test items or tests. in section 4, i describe the wide variety of practices currently employed and explain why they are often unreliable. in section 5, i discuss some of the main problems encountered when consulting experts on test development. in section 6, i draw conclusions from the main findings of the data analyses and specify the need for improvement of current practices in consulting experts on test development. 2. experts’ advice on how to ensure content validity: theoretical background when developing psychological or educational tests the typical way to ensure content validity is to make use of the expertise 4 of professionals in the field (allen & yen, 2002; reynolds, livingston, & willson, 2009). by a priori assumption, experts can determine whether or not the content of a test is “sensitive” to variation in the formation of a latent attribute that the test or one of its items is directed toward. due to the innumerable situations a person may find him or herself in, this latent attribute could invoke an infinite number of perceptions, thoughts, reactions, or actions. in this sense, a group of test items is understood as a sample of situational constellations to which a person might react (or on which a person might act) relatively adequately (kerlinger, 1986; anastasia & urbina, 1997). in terms of ensuring the content validity of a test, an expert is asked to judge, assess, or rate the extent to which the items of a test adequately represent the infinite number of situational constellations by which the latent attribute to be measured stimulates a certain behaviour in test takers. this implies that the expert needs to do the following: identify an understanding of the latent attribute to be measured; determine the universe of possible situational constellations by which the latent attribute is activated; consider all the behavioural patterns being caused by the latent attribute in dealing with these situational constellations; relate these constellations to the test and its items and to the respective latent attribute to be stimulated, that is, the construct, and examine whether these relationships are adequate in terms of the aim(s) of the test; determine whether the test items will function as adequate stimulators for different groups of test takers (e.g., in terms of language, level of education, social background, etc.); provide feedback to the test developer(s) in a way that allows them to make decisions concerning the validity of the test items (grant & davis, 1997, p. 272, col. 1-2). experts can give feedback informally, for example during structured interviews, or formally, for example in written reports or surveys with “yes” or “no” ratings of acceptability or ratings on likert scales. anderson et al. (2015, p. 22), for example, recommend following an online likert-based procedure when working with large data banks and large pools of experts, and using latent trait models “to control for between-rater severity, evaluate interrater consistency, and provide item-level diagnostic statistics.” usually, several experts are consulted to evaluate content validity. if experts are scattered around a country (e.g. across the us) they might have to convene in person to discuss their ratings, if different, until consensus on the (good or bad) quality of an item or a test is reached or a majority vote is taken. the greater the number of experts consulted, the more likely it is that they do not meet in person, and that quantifying procedures are used to obtain information from them (lawshe, 1975; thorn & deitz, 1989). as a qualitative alternative, delphi studies may be conducted in which experts state their opinions in written form, then the test developer bundles those opinions and distributes them repeatedly until consensus, or at least a majority consensus, is reached (see e.g., lohse-bossenz, kunina-habenicht, & kunter, 2013; messmer & brea, 2015). test developers also often assign different tasks to experts. for example, they pose different questions to the experts, sometimes asking one group to judge all items as a whole and the other group to judge each item separately, or sometimes by recruiting multiple groups of experts and asking each group to evaluate one validity-related aspect of the items only (see e.g., aydın & uzuntiryaki, 2009, p. 871; mesmer-magnus et al., 2010, pp. 514-515; jenßen, dunekacke, & blömeke, 2015, pp. 21-22; cf. appendix 2, no. 6). such aspects include the representativeness, clarity, relevance, conciseness, or technical adequacy of single items or the test as a whole (see table 1). when consulting more than one expert, the question arises as to how to deal with different or even incommensurable judgements rendered by them. if the experts’ feedback is qualitative, the test developer may have to grapple with diverging arguments, weigh them, and finally decide which are the most plausible. here, the question arises as to what quality the arguments should have in order to override the opinion of at least some experts, and why the test developer should ask experts at all if in the end he or she will be making the final decision himor herself. if the experts’ feedback is quantitative (e.g., ratings on likert scales), usually an arithmetic mean of these ratings along with its standard deviation is computed and then checked to determine whether certain meaningful thresholds in these values have been exceeded or underrun. again, the test developer has to make a decision, namely, to set the boundaries of the range within which, in the end, an arithmetic mean or a standard deviation shall be acceptable (see e.g., jenßen, dunekacke, & blömeke, 2015, pp. 22-23; cf. appendix 2, no. 6). again, the question arises as to which arguments justify the determination of those boundaries. the aforementioned procedures are followed in an attempt to reach at least broad consensus among external experts, and also between those experts and the test developer(s) on the adequacy of test items. such consensus tends to be used as a justification for the inclusion of the items and as an argument for the overall quality of the test (even though it is based on the implicit fallible, if not false, assumption that this is a valid and authoritative indicator of the appropriateness of the content of a test). these approaches also seem to be fully in line with the rationale for ensuring validity by “evidence based on test content” declared by the standards management committee of aera, apa, and ncme (2014, pp. 14-15), which exemplify the role of experts by four functions: • assigning test items to content-specific categories or facets of an occupation (p. 14, col. 2); • judging “the representativeness of the chosen set of items” (p. 14, col. 2); • rating the “relative importance, criticality and/or frequency” (p. 14, col. 2) of job observation-based item pools; and • avoiding construct-irrelevant sources of variance caused by inadequate wording of items (p. 15, col. 1; see also standards 1.9, p. 25 and 4.8, p. 88.). it is unclear whether these four functions are meant to be exclusive. presumably, this is not the case because there is no mention of why experts should not be allowed to perform more than these four kinds of tasks related to ensuring content validity. messick stated that experts’ “judgements should focus on such features as readability level, freedom from ambiguity and irrelevancy, appropriateness of keyed answers and distractors, relevance and demand characteristics of the task format, and clarity of instructions” (messick, 1987, p. 55). similarly, kane (2013a, p. 5), citing angoff, pointed out (and did not criticise) that “(m)ost of the early tests of mental ability and many current standardised tests of various kinds … have been justified primarily in terms of “a review of the test content by subject matter experts” (angoff, 1988, p. 22). elsewhere, kane stressed that he did not claim “that the blessing of an expert committee is, in itself, adequate for the validation of an achievement-based [test; k.b.] interpretation, but [he; k.b.] would expect an achievement test to pass this kind of challenge” (kane 2013b, p. 121). though the meaning of validity has changed over time (baker, 2013; shepard, 2013) and now is derived from an argument-based approach, the facets of validity as previously discussed have not disappeared. their status has only been altered by now being subordinated to the modern understanding of validity as a problem of quality of “interpretation and use of test scores” (kane, 2013a, p. 2).5 it is worth pointing out that test developers draw on experts to ensure not only content validity but also construct validity (factorial validity, convergent validity, and discriminant validity), concurrent validity, cognitive validity 6 or other variations of validity. 7 in addition, experts are consulted to judge or to give advice with regard to statistical aspects of test development. moreover, they are needed to assess the adequacy of test takers’ answers to open-ended questions on a test. 8 experts’ additional opinions are in demand as assessments in education, at least in the united states, seem to shift from measurement tools to policy levers establishing test-based accountability policies, a development which launches new test functions and is considered to be a matter of consequential validity (henig, 2013; shepard, 2013; welner, 2013; aera, apa, & ncme, 2014, p. 14). in this paper only content validity in the sense described above is explored. 3. types of information needed from experts test developers do not always consult experts to ensure content validity. in some cases, they feel confident that they are able to fulfil this quality criterion without help because they consider themselves to be experts. further, they might not know or have contact with the “very best” experts who are willing and able to get involved in the development of their test(s). however, if they decide to consult experts, they may have several questions they want to pose to them. in table 1 is a list of 31 types of tasks that experts might be asked to perform. the list of types of tasks for experts was developed in three successive stages. first, ten studies were evaluated in terms of the tasks experts were asked to perform. these tasks were grouped under general headings and according to the stages of the test development process in which they were performed. then, as each subsequent study was analysed, any new task for experts found was placed under one of the general headings if possible or was put into a new category with a new heading. in the end, the categories were reviewed and any possible task for experts i felt, from my own experience with test development, was still missing from the list was added. overall, 31 types of tasks for experts were identified. thus, this inductive and deductive approach to identifying tasks for experts is not intended to be understood as an exhaustive list. table 1 experts’ contributions to ensure content validity note. * tu: truth-apt, but unknown: no objective knowledge available yet. to analyse procedural practices, differentiation is made provisionally among three types of information9 test developers need. the first type is truth-apt10 in the analytical sense of propositional logic. this means that a statement is capable of being true or false and that it is objectively (i.e., intersubjectively) testable whether this statement is in fact true or false (indicated in table 1 as to). for example, experts may be asked whether an attractor on a multiple choice test is correct, the answer being either true or not true (e.g., 7f). the second type is truth-apt as well, but it states information which is not yet known (indicated as tu). for example, if asked how many traffic accidents occurred on a particular holiday throughout the world, or how often a certain cognitive performance has to be shown by an apprentice per work day, there will be a true answer but it is not available (is unknown) as long as this matter has not been researched (e.g. 4f). the third type asks for an expert’s opinion (i.e., a valuation, estimation, or perception) regarding, for example, the relevance or understandability or feasibility of a test item or an entire test (see e.g., 4a, 7b, 8a). such statements appear as individual judgements and therefore are inherently subjective (indicated with s). they may differ from person to person without possible rebuttal, meaning one can only agree or disagree with them. statements of this type cannot be “true” or “false”, and it would not make sense to attribute the property “truth-apt” to them. likewise, it would not make sense to take a position on a “true” statement by saying that one does not like it or that one is ready to agree with it. as shown below, these distinctions are of fundamental importance when dealing with questions to, and answers from, experts. 4. procedures for ensuring content validity to gain insight into the procedures adopted by test developers to ensure content validity by a qualitative meta-analysis i reviewed a sample of 72 published reports. in germany between 2009 and 2015 two large research programmes funded by the ministry of education and research were launched to promote the development of modelling and measuring competences and skills in the fields of academic education (kokohs) and vocational education (ascot). in kokohs and its thematic context 24 project groups have been researching the measurement of general and domain-specific academic competences (pant et al., 2016). the kokohs programme received roughly €13 million in funding.11 in ascot and its thematic context six project groups have been developing mainly it-based test instruments in the fields of commerce, mechanics, and health care for use in vocational education (beck, landenberger, & oser, 2016). this initiative received approximately €7 million in funding.12 initial findings of kokohs and ascot were published in 2014 and 2015. both programmes were treated as the first sections of a far-reaching research project, and were exclusively devoted to test development.13 within that period findings from five other projects on test development conducted in germany (funded by other entities) were published and have therefore been included in the sample. i also accessed two relevant american journals: ‘educational and psychological measurement’ and ‘measurement and evaluation in counselling and development’. unlike many other journals in the area of diagnostics, they offer a special section for reports on test development in various subject areas, thus allowing easy access to the material of interest. i explored volumes from 2009 to 2015 of both journals, the same period during which the projects in germany were being conducted. thus, sampling was convenient and random: “convenient” with respect to the sources from which the test development reports were drawn and “random” for the single reports which were funded independently of each other within large research programmes (germany) or published independently of each other in the two selected journals (us). as the aim of this study is not to present representative data on its subject, but rather to give an exploratory overview of the range of procedures dealing with the inclusion of experts in ensuring content validity of tests or test items, this approach to sampling does not result in a relevant bias. the overall sample included 72 reports on test development, 14 which i analysed as to how content validity is treated, whether experts were involved, and if so, how the experts were selected, how many of them were consulted, and what they were asked to contribute to ensure content validity. from this perspective the test content (whether psychological or educational, whether focusing on teachers, students, or apprentices), national context, and publication medium (journal or book chapter) were irrelevant, as the procedures to ensure content validity can be described and compared independently of these conditions. 15 of the 72 reports, 16 did not rely on experts to ensure content validity (see appendix 1, col. 3(a): “n”): in two of them, experts’ advice was purposefully avoided because the researchers felt capable of assessing the quality of their test (appendix 1, no. 53, 72), and in the others 16 no mention was made of the use of external experts. in the remaining 56 reports (78% of the sample) some information was provided concerning the use of experts and their functions in the process of ensuring content validity. in 10 of them17 experts were consulted but no clear description was given about the task(s) they performed. overall, none of the reports in the sample provided an adequate description of the criteria used to determine the exact role and contribution of experts in ensuring content validity (cf. appendix 1 col. 3 to 6). only three of the reports 18 came close to doing so. in 18 reports only one type of information was requested from experts (cf. appendix 1, col. 7) whereas in others several types of information were sought (e.g., appendix 1, no. 25 and 35): between seven and eight questions were posed to each expert.19 overall, experts were asked 113 questions related to all types of information relevant to content validity (cf. appendix 3) over the course of all test development projects in the sample. their support was requested mostly for items concerning domain-specific characteristics (cf. appendix 3, col. 4, type no. 4: 41 of all 113 requests) and quality of wording (type no. 7: 33 requests). some information types listed in table 1 did not seem to be needed very often from external experts: type no. 1: ‘definition of a content area’ was requested twice only (cf. appendix 3, col. 4), type no. 6: ‘representativeness of an item sample’ was requested five times. other types of relevant information were not requested at all.20 across all projects in which the type of questions posed to experts was reported, the majority of the experts’ responses fell into the category “subjective” that is requiring approval or disapproval (table 2).21 table 2 distribution of experts’ responses to all types of questions posed of the 56 reports that included the involvement of experts, 40 provided information on the number of experts they contacted. in four other reports experts were said to have been consulted but the number was not stated. in two other reports only the number of one subgroup of experts was stated. overall, at least 1,261 experts plus an undisclosed additional number are claimed to have been involved in the 46 projects in which a number of experts consulted was stated. this means that test developers asked an average of at least 27 experts for help to ensure the content validity of their instrument.22 it is not surprising that the majority of these experts (approximately 70%) had a university background. table 3 offers an overview of the experts who were consulted about test item content. table 3 distribution of experts according to fields* note. *assignment to the categories was based on information given by test developers, which sometimes was inconclusive and therefore made assignments somewhat arbitrary. nevertheless, this table gives a general impression of the number and background of those experts. here, an impression is given of the cost arising from consulting experts regarding content validity. it can be assumed that it takes an expert approximately one hour to answer one question concerning one test instrument.23 from the data (appendix 1, a combination of col. 4 and 7) the number of questions posed to experts across the 48 projects in which they were involved is known, and the number of hours of work required by them totals 4,439.24 usually experts do not receive any monetary compensation for their work. in none of the reports in our sample was a fee for the work of experts mentioned. this might be plausible with respect to university students because one might surmise optimistically that by analysing test items they achieve a certain learning gain as remuneration for their work. however, even if the amount of their contribution is subtracted, 2,177 working hours of non-student experts are left (corresponding to 272 eight-hour days of work or approximately 54 weeks of work, i.e., more than one year of full-time work). this amount of work was done almost completely by academics either at or outside the university. all of the latter were probably used to being paid relatively well. therefore, in monetary terms their contribution was considerable. in pilot studies and empirical investigations, it is common to describe in detail the procedures followed to draw samples. one might expect the same applies to selecting experts for their input on test content. this expectation seems to be justified, especially since their work is essential to the quality of a newly developed measurement instrument which may be administered to thousands of test takers, and test results may be used as an important basis for decisions affecting numerous stakeholders. therefore, i investigated whether the test developers were aware of the area(s) of expertise of their potential experts and what method they used to recruit them. the findings are listed in appendix 1, col. 3(b, c) and are summarised in table 4. table 4 descriptions given of procedures for selecting experts* note. * of the 56 projects involving experts; ** see headings in appendix 1, col. 3. in most of the reports, test developers reported at least some idea about the expertise they needed. this was obvious because they described the tasks they gave to the experts and/or the questions they asked them. in 14 reports the test developers mentioned something vaguely about how they selected their experts; in 13 of those, this was done after mentioning the expertise needed. in 12 projects involving experts nothing was mentioned about this issue.25 in addition, in empirical studies it is standard procedure to report on the results yielded from the pilot sample. likewise, test developers could provide a qualitative report on the feedback they received from their experts and also a summative, quantitative report, especially when multiple experts were involved. in table 5 is an overview of the reports in which experts’ feedback was described (cf. appendix 1, col. 6). table 5 reports on experts’ feedback* note. * of the 56 reports including experts; ** see heading in appendix 1, col. 6. it should be kept in mind that test developers often are not inclined to give a complete account of their endeavours to ensure content validity with the aid of experts (berk, 1990). only one report in the sample (appendix 1, no. 6) provided full details about how experts were involved according to the quality criteria explained above (cf. table 4). three other reports (appendix 1, no. 17, 50, 55) come reasonably close to this standard, which is by no means excessively high, but rather customary in all other empirical issues. however, the question arises as to whether and when experts should be consulted at all when developing tests. 5. discussion: who is an expert and what can he or she contribute to ensure content validity? the analysis of the outcomes of typical empirical studies of test development revealed that no recognised standard has been established on what and how to report on the involvement of experts in the development of measurement instruments. rather, authors tend to deal unsystematically with such matters, and if they include details, they tend to be incomplete. one might assume that in some cases such details are omitted from reports due to limited space; however, a few brief details on this procedure could be expected. in our sample, such details were not provided in any report. however, this problem may also be judged differently in the editorial boards of the relevant journals. the obligation to report on experts’ contributions is one formal aspect of treating matters concerning the involvement of experts in test development. more important and particularly fundamental are the aims of involving them. from the overview given in the preceding section, it is clear that the inclusion of experts to ensure content validity of test items or tests as a whole cannot guarantee that this objective is reached.26 the main issues are as follows: (1) it is not always clear whether the advice of an expert or even multiple experts is necessary and/or could be helpful at all. (2) “experts” usually are not individually selected by controlling for their expertise/qualification. (3) experts also are not assessed on the effort they make to answer the questions posed to them. (4) variations or even contradictions in the answers of experts often are not carefully analysed, but rather are levelled out by procedures of averaging or excluding minority votes, regardless of the type of answers given (cf. table 1). (5) revised test items are often not returned to experts for review. (6) it remains unclear whether and to what extent the expertise of experts may be overruled by the expertise of the test developer(s). in the following, these issues are discussed in more detail. 5.1 asking experts questions requiring truth-apt answers in general, questions asking for truth-apt answers could be dealt with by one competent expert if not by the test developer, who usually is an expert him or herself (cf. app. 1, no. 53, 72). as mentioned above (section 3), it is necessary to differentiate between information which is already objectively known (denoted by to in table 1) and information which is still unknown but in principle could be known after having been investigated, usually by researchers (denoted by tu in table 1). the following must be considered: (a) if, for example, the mathematical correctness of an attractor in an mc item is to be examined or if it is necessary to know whether certain content is part of a relevant regular curriculum, the answer can be true or false only (to; cf. no. 4e, 4g 5b, 7f, 7g, 8c, 8d in table 1). in clear cases like these it does not make sense to ask multiple experts. if there is one and only one true answer, and if only highly competent experts are involved, the test developer will get only one answer from all experts consulted. otherwise, not all experts consulted are real experts, and their dissent might be used to exclude those who provide incorrect answers from further questioning (grant & davis, 1997, p. 273, col. 2). obviously, it would be a mistake to understand the dissent of experts on questions such as these as a matter of differing opinions or positions, and then to dissolve it by having a majority vote or something similar. (b) the situation is different if, for example, the frequency or likeliness of the appearance of a particular vocational operation in a specific domain is at issue. usually, the determination of quantities of this type (no. 4f, 6, 8g in table 1), though truth-apt, might be unknown and will be a matter of further research (tu). experts normally do not initiate research projects to find answers to test developers’ questions; rather, they will give an estimation of how the correct answer might read. in this case, it might be advisable (but not imperative) to ask several experts to give their estimation of the correct answer, and then to take the mean, mode, or another reasonable average value from their specifications. however, if the test developers have any valid information or a justified opinion of the range within which the experts’ answers should be, they must exclude some answers as outliers, unless these alone would justify a study of this particular question to be launched. caution should be exercised if a test item asks the test taker for his or her assumption about what others (including experts) believe to be the frequency or likeliness of this type of issue. again, this could be known if investigated; however, this finding normally is not available. therefore, cases like these also fall into the category tu (see 4f in table 1). (c) another type of information requested from experts requires separate consideration. imagine, as is often the case, a question posed to experts about whether a certain distractor in an mc item will be considered plausible or implausible by test takers. with respect to one particular test taker, the expert’s response (yes or no) is truth-apt but the expert usually does not know if the answer to the question actually is true or false (therefore: tu). however, for numerous anonymous test takers an answer is even more tricky because some of the test takers may perceive that distractor as plausible and some of them may not. the answer also is unknown (tu) as in (b) but, even worse, it will arise from a process of comparatively higher inference than in (b). the expert has to speculate on the type of test taker who considers the distractor to be plausible and to estimate the percentage of test takers of this type in the unknown multitude of possible test takers. in this case, an expert’s reaction to the test developer’s question is much more questionable than when speculating about only one single test taker (and all the more precarious than his or her answer to a question concerning the frequency or likeliness of observable occurrences of vocational operations as in (b)). in any case, experts’ answers to questions like this will be highly uncertain and therefore not really useful in terms of ensuring content validity. therefore, this type of question should not be posed to experts at all. (d) criteria of type no. 4a through d (table 1) are expressions of an inevitable subjective valuation (s). this follows from the fact that two experts in the same domain or occupational activity may judge the relevance, importance, representativeness, or dangerousness of a given task differently, but their answers cannot be proven to be right or wrong. even if the experts’ answers to this type of question are the same, they only demonstrate consensus in their personal feelings (as viewers of a piece of artwork might do). however, unity in feelings does not change the logical character of their statements as valuations. (e) if experts are asked to predict future developments or events (e.g., by evaluating the correctness/falsity of an answer concerning an mc item; no. 7d, table 1), later on their answers might turn out to be true or false (as in (a) above). in this case the correct answer is not yet known, not even by researchers, and therefore cannot be given at the time the question is posed. thus, it would be senseless to ask experts questions concerning the quality of such an answer option. mc item options of this type are flawed and should never be included. however, the situation is different if test items request forecasts based on theoretical models. for example, nobody knew in 2007 what the consequences would be if a very large bank were to go bankrupt. nevertheless, it is imaginable that developers of tests of economic understanding might have posed a question like this to experts at that time. if experts were asked what they thought the most likely course of economic development would be, in the respective item stem there should have been the constraints ceteris paribus (“all else being equal”) or rebus sic stantibus (“things thus standing”), which is common in economics and law. further, if relying on a particular theory of international trade (e.g., in the tradition of the famous model of ricardo) the correct answer could have been known (to).27 for some test developers, this is reason enough to consult several or even numerous experts. again, however, one competent expert (i.e., a person with knowledge in the respective domain) should be able to tell whether an option offered in an mc item like this is correct or incorrect. this case is either to or tu depending on whether completely all (to) or only some (tu) conditions/variables of the further development of circumstances are given. for a competent test developer, experts’ assistance is not necessary with to and not helpful with tu , because in the latter case in principle no future development can be excluded. initially, one could conclude that to items generally do not need to be judged by experts unless the test developer feels the need to verify that his or her solution is correct. in this case input from one expert should be enough. conversely, tu items should be excluded from all tests because their content deals with uncertain circumstances. test developers and test users must know that, as a matter of principle, in these cases the content of the answers given by test takers (as well as by experts) cannot be judged as right or wrong. it is quite another issue if not the knowledge of test takers, but rather their ability to handle uncertain situations, is to be assessed. their responses to tu items then might be judged as being acceptably precise, or more or less circumspect, or as being within an acceptable amount of time or indexed by an acceptable caveat of accuracy, and so on. this leads to another type of information to be obtained from experts: their opinion about the appropriateness of an item to stimulate test takers to show their competence in dealing with the various dimensions of uncertainty (no. 8a, 8b in table 1). 5.2 asking experts for subjective reactions the alternative to asking experts for truth-apt statements is to ask them for statements indicating their approval or disapproval (acceptance or rejection) of an item, also referred to as consent statements.28 the basis of such statements is not objectivity in the sense usually meant when referring to interpersonal verifiability. rather, these statements are based on subjective mental states, subjective preferences, subjective appraisals or something similar. they can be identified easily by their indispensable recourse to the speaker himor herself, even if this recourse is not made explicit. for example, if an expert says: “with respect to criterion x item no. i is fine,”29 he or she means: “this is item no. i and looking at it with respect to criterion x my impression is positive.” we also refer to such judgements as valuations. although valuations are subjective in character, they are different from norms (i.e., demands or claims30). declarations of this type are usually followed by an exclamation mark. they consist of two elements: a description of a state of affairs as the object of an order and the order itself, such as “this is a certain state of affairs and i demand that it be maintained (established or eliminated)”31. in the context of our literature review we may call these statements “decisions” because a decision is an explicit or implicit statement about a personal will, wish, or at least hope.32 by contrast, valuations do not necessarily imply a normative claim. 33 (f) in the course of test development experts often are asked for statements indicating their approval (grant & davis, 1997, p. 272, col. 1; table 2, col. 3). for example, they are asked to decide whether a particular item belongs to the domain a test is aimed at. without an exact definition of the notion domain and a clear criterion for the delimitation of domains, this statement cannot be true or false (shavelson, gao, & baxter, 1995). when answering such questions experts might be led by their intuitions,34 which undoubtedly differ among individuals. in table 1 is a list of various situations in which experts are asked or forced to make a decision. 35 the vexing problem for the test developer then is to decide how varying decisions of a group of experts should be treated. should the test developer follow the majority? how many experts are needed to get a substantial contribution to the decision on the content validity of an item? why should the majority decision or any decision be considered the best? should the test developer ask the experts not only to make a decision but also to disclose the criterion on which their decision is based? what if the experts cannot name that criterion? what if the experts can, and the developer gets different decisions based on different criteria? would it be a good idea to keep only those items on which experts are in complete agreement? as an example, huang and lin (2015; cf. appendix 1, no. 71) developed their inventory for measuring student attitudes toward calculus for college students. in working on content validity, they asked experts to judge the clarity of the 24 items on the test. one of the items read: “i think it is difficult to learn calculus.” is the criterion of clarity fulfilled? how would you decide? yes or no? the experts decided “no”. therefore, the item was reworded to: “calculus is not difficult for me.” do you agree that clarity is now given? the group of experts consisted of “three educators with expertise in instrument development, one psychological counsellor, and two professors who have each taught calculus for more than 10 years” (2015, p. 113, col. 1-2). the authors did not report on the number of experts’ decisions on clarity after rewording the item. however, it is plausible that not all of them responded in the same way. the question posed to the experts by huang and lin does not solicit a truth-apt answer. judging clarity in the context of item construction obviously is a matter of subjectivity because clarity is a two-place predicate meaning “clarity for someone”. the example above deals with clarity for college students, but why did huang and lin not ask college students, who are the genuine experts, to answer this question? indeed, the two authors conducted a pilot study with a sample of students, but they did not ask them to judge the items on the criterion of clarity. rather, they administered their test to them. however, to get the necessary information on clarity they should have asked the students which of the two alternative wordings to them was clearer. what would the consequences have been if they had asked students (as experts) in this way, and if they had not obtained a clear majority for one of the two alternatives? if one of the two alternatives had been favoured, would the answers given by the minority have been considered as valid as those of the majority? would it have been necessary to investigate relevant psychological differences between members of the minority group and those of the majority group? how could this have been done? would validated tests have been needed to detect and identify these differences? hence, is the result of validating test items an infinite regress? finally, what does it mean for the validity of test interpretations if experts find a particular item only partly clear? would it be fair to force all future test takers to respond to such an item? how many items would remain in an item pool if all those items failing to gain consensus among experts with regard to clarity (or any other predicate of the same kind) were removed? pragmatic issues may come into play such as the costs and benefits of developing a test in terms of money, energy, and time. how much effort has to be invested to ensure the content validity of items with respect to clarity (as in the example above) and all other “s aspects” (in accordance with table 1 and appendix 3)? do financial constraints diminish the quality of test content or corrupt it totally? in other words, is content validity a quality which may vary gradually (as has been the prevailing view for a long time; gay 1980)?36 can it be assumed that if this is so, the content validity correlates (strongly) with the risk of inadequate, or rather, invalid, test interpretations? thus, the value of experts’ decisions on item validity is debatable. from what has been mentioned so far, the experts base their decisions on definitions (von savigny, 1971), expected effects of items on test takers, or other causal relations. in the latter case, it does not make sense to ask experts for judgements, because they are overtaxed by the expectation to give an appropriate answer to a question requiring empirical knowledge (to). as mentioned in the example above, often the test takers are the true experts (see table 1, no. 5a); however, in test development they usually are treated as test subjects within a pilot study, which is different from using them as experts for their decision on the quality of test items (vogt, king, & king, 2004; grant & davis, 1997, p. 273, col. 2). (f) finally, asking experts to give valuations seems to be the most common way to involve experts in ensuring content validity. experts are usually presented with scales (often likert type) on which they state their personal valuations of a certain criterion to be adopted for a test item. for example, experts are asked to rate the relevance of an item (no. 4a in table 1), its level of difficulty (no. 5aa), its understandability (no. 7b), its appropriateness (no. 5ad), and/or the extent to which they fully agree/agree/partly agree/disagree with a statement characterising the respective item under a certain condition (e.g., to be gender neutral (no. 5ab) or culturally unbiased (no. 5ac)). afterwards, the experts’ answers are usually mapped numerically, presuming that they vary continuously and therefore can be averaged arithmetically. then, the resultant value is re-translated empirically and interpreted according to the description of the grades of the given scale. if this value falls into the range of a rather positive valuation (probably higher than the arithmetic mean) the respective item, in the absence of any other obstacle, will be accepted by the test developers. otherwise, this item will be modified or even omitted. although this procedure generally is considered appropriate, on closer examination it is anything but meaningful for various reasons. (aa) often it is far from clear what the differences between full grades on a given scale mean in terms of item quality. this is all the more true for values lying between full grades, as is usually the case for the arithmetic mean. for example, imagine that (as in jenßen, dunekacke, and blömeke 2015, pp. 22-23; appendix 1, no. 6) 24 experts are asked to rate whether a given item is a good representation of all theoretically possible items. what does it mean if on a four-point scale of 1 not at all, 2 rather no, 3 rather yes and 4 totally one item gets a mean of 2.4 and a second item a mean of 2.8? in this case, the authors decided to omit the first one because its mean lies beyond the middle of the scale, that is, 2.5, whereas the second item has been kept but had to undergo extensive revision. thus, the difference of four-tenths between 2.4 and 2.8 is significant but the same difference between 3.5 and 3.9 does not lead to any consequences. obviously, in this case, the requirement of equidistance of data is not considered adequately. (ab) as an alternative one could consider whether the mode or the median would offer better interpretations of the experts’ answers. both have a relatively clear meaning as they mirror a verbalised subjective feeling in standardised (i.e., not a real individual’s) wording. this is an advantage over the computation of an arithmetic mean usually resulting in a decimal which does not represent any of the answers given by experts. however, if the modal value exceeds the value next to it by one or two only,37 it might be difficult to decide that this small deviation justifies the decision to prefer the mode as an adequate interpretation of their opinion. moreover, if the median value resulting from the experts’ answers lies between the next lower value on the left and the next higher value on the right, one could hardly reason that this value represents the average opinion of the experts. choosing one of either measures would exclude all divergent interesting or relevant reasons the experts might have had in mind when answering test developers’ question(s). (ac) the question remains as to whether it makes any sense at all to reduce the valuations of experts to an average. first, test developers may deal with arithmetic means quite arbitrarily. in addition, differing expert valuations may be of different quality due to judgement tendencies known from research on rating biases (e.g., leniency, central tendency, halo effect, etc.; cf. beck, 1987, pp. 184-186), and controlling for this is very time-consuming and expensive. furthermore, one expert might have better reasons for his or her valuation than another, an issue that refers to the fact that even valuations may include a more or less distinct portion of cognition emerging, for example, from relevant experience 38 or knowledge. moreover, the answers experts give might be influenced by their motivation to cooperate (which might be influenced by monetary incentives), their commitment to supporting the test developers, or the time they are prepared to invest in the processing time of the tasks they were given from possibly completely unknown people (i.e., the test developers) whom they are not obliged to help. 6. conclusion to summarise, the role of experts in ensuring the validity of the content of tests and the methods used by test developers to obtain experts’ feedback is often dubious. experts’ involvement in the process of test development can vary significantly. usually, more than one expert is consulted and often experts do not agree in their valuations, decisions, or even truth claims. thus, the test developers have to reclaim responsibility for ensuring content validity, which they initially intended to transfer to their experts. perhaps in an attempt to avoid giving responsibility to experts who differ in their judgements, often without being aware of it, test developers try to “average out” the procedures discussed above, which they believe permit interpretation of experts’ incommensurable feedback. moreover, if test developers are lucky enough to have consensus among their experts, they still cannot be sure the content of their test is valid because consensus is no guarantee for correctness in any sense (popper, 1972). finally, test developers often do not seem to have a clear understanding of what they are asking experts to judge (e.g., the colloquial, and therefore ambiguous, terms such as relevance, difficulty, and clarity of test items). in addition, they usually do not have clear criteria against which they can assess the knowledge of external experts. test developers might feel tempted to make use of experts’ competences to legitimise their claim that their test is of good quality. however, in many cases it is the test developers themselves who are the best experts, and as such they do not really need to get advice from others on the many questions which are routinely posed to so-called experts in the procedures of test development and, in particular, of ensuring content validity. if there is a real need for advice from experts, selecting them is not easy and needs much more diligence than is routinely applied (grant & davis, 1997, p. 270, col. 1-2). moreover, formulating adequate questions for experts is even more challenging. it would be a grave misunderstanding to believe that getting experts involved in an attempt to ensure content validity is a matter of representativeness in any sense. rather, it is a matter of careful selection of people (or a single person) having the competence to contribute to the quality of a test. test developers should consider themselves lucky if they have tests at hand that allow them to identify competent experts. as mentioned above, this idea would lead immediately to an infinite regress. if test developers feel it necessary to consult experts, there should be no reason not to publish the names of those experts along with the names of the test developers or at least, with their consent, to communicate their identity and lift the verbal veil of anonymity associated with the sweeping talk of “experts”.39 thus, they not only would share responsibility but also receive appropriate recognition for their contribution to the resultant product. the standards for educational and psychological testing (aera, apa & ncme, 2014) highlight the importance of reporting on the process of developing tests and are guidelines with which test developers are expected to comply. it seems that editors and reviewers of journals pay little attention to the details about how experts contribute to test development, even though standard 1.9 gives quite helpful hints on how to report on the inclusion of experts in the different stages of test development (not only to ensure content validity; p. 25 col. 2-26, col.1) and standard 7.5 states “test documents should record … the nature of judgements made by subject matter experts (e.g., content validation linkages)” (p. 126, col. 2). however, the standards do not go into the details concerning the role experts play in providing evidence of content validity and do not elaborate on the procedures, problems, or fallacies surrounding the selection of experts and interpreting and employing their advice (pp. 14-15). with that in mind, one might conclude that further discussion is needed on the guidelines and principles to be observed when making use of the advice of experts to ensure the validity of the content of tests. keypoints the paper offers a comprehensive overview of experts’ involvement in ensuring content validity of tests. a review of reports on test development (n = 72) reveals a lack of information regarding the process of selecting experts and the methodological treatment of their qualitative and quantitative input. a discussion on whether and when to consult or not to consult experts to ensure the content validity of tests. the current standards for educational and psychological testing (aera, apa & nmce, 2014) should be enhanced by offering detailed guidelines regarding the involvement of experts in test development to ensure content validity. footnotes 1 in the 1985 edition of the “standards” it says: “the inferences regarding specific uses of a test are validated, not the atest itself” (p. 9). 2 though being judged as technically incorrect and therefore theoretically not helpful in the early years after inception, the concept of content validity has survived and still enjoys attention in the relevant literature (guion, 1977; sireci, 1998). 3 the first author to develop a quantification of content validity is presumably lawshe (1975). several modifications of his measure have been developed (wilson, pan, & schumsky, 2012). 4 n this paper the meaning of the terms expert and expertise are not as specialised as in ericsson and smith (1991). rather, the terms are used more generally to convey the meaning of comprehensive and authoritative knowledge of or skill in a particular area. 5 to be more exact, kane emphasises that this terminology gives equal weight to interpretation and use, as in his formulation “interpretation/use argument” (iua) (kane, 2013a, p. 2). 6 though related to content validity, the construct cognitive validity points in particular to mental processes of test takers, while content validity is devoted mainly to domain-specific aspects of wording and the purport of test items. as a matter of fact, “cognitive validity” in particular designates the problem of whether an item elicits thought processes of adequate complexity (field, 2013; smith, 2017). nevertheless, the relationship between the two constructs needs to be clarified. 7 with regard to types of validity there are many more circulating in the literature. see e.g., messick (1990). newton and shaw (2014, pp. 7-8) list 151 “(k)inds of validity that have been proposed over the decades” (p. 8, table 1.3). 8 this task often is done by raters. in a sense they also are experts. however, their function is clearly distinct from that of experts involved in test development. nevertheless, white’s (2018) sophisticated discussion on performance standards for raters could be transferred and applied in part to the issue of experts involved in ensuring the content validity of tests. 9the reason for this simplification is that, as a rule, reports on test development do not explore this issue in depth or even touch upon it. more specific differentiations are discussed below (section 5). 10 there is a longstanding and endless philosophical discussion on the notion of and theories surrounding truth. in the present context truth-aptness is understood as the meaning argued by correspondence theorists in the framework of critical rationalism in the sense of popper, or analytic philosophy in the sense of quine; see e.g., jackson, oppy, & smith, 1994; dodd, 2002. 11 around 220 researchers involved in roughly 70 single projects conducted in 12 federal states of germany. detailed information on measuring instruments developed in kokohs is given by zlatkin-troitschanskaia et al. 2020. see also: https://www.kompetenzen-im-hochschulsektor.de/kokohs-2011-2015/. 12 including more than 12,000 test takers at approximately 300 schools in 13 federal states of germany. for more information on ascot, see: https://www.bmbf.de/pub/berufsbildungsbericht_2016_eng.pdf. 13 in the meanwhile, the follow-up phases have started being dedicated to questions of application and implementa¬tion (kokohs ii [https://www.blogs.uni-mainz.de/fb03-kokohs-eng/] and ascot+ [https://www.ascot-vet.net/de/forschungs-und-transferinitiative-ascot.html]). 14 kokohs: 18 reports; ascot: 4 reports; other projects conducted in germany: 5 reports; educational & psychological measurement: 22 reports; and measurement & evaluation in counselling & development: 23 reports. 15 this holds also for the selection of experts in the german programmes kokohs and ascot. though different in regard to their research topics, they both had to deal with the question of content validity. 16 appendix 1, no. 3, 16, 23, 24, 26, 30, 38, 46, 48, 49, 51, 56, 59, 60. 17 see appendix 1, no. 8, 12, 15, 41, 42, 47, 65 – 68. 18 see appendix 1, col. 6, no. 17 (missing only a qualitative report on the results of consultation with experts), no. 50 (missing a quantitative report) and col. 3(c), no. 55 (missing information on criteria and method(s) for selecting experts). 19 in the 46 projects reporting on the questions posed to experts, on average two or three questions were asked. 20 types no. 4 d, f; 5 ab, 5b, 7h, 8c, d, e, g. 21 the figures in table 2 are calculated according to the frequencies given in table 1, col. 4. 22 the number of experts consulted ranged from one to 306 (appendix 1, no. 67 and 17 respectively): 1-9 experts included in 21 projects, 10-49 experts in 13 projects, and more than 50 experts in 6 projects. 23 this depends on the number and type of items of the respective instrument, which can be within a broad range. 24 in seven of these projects only the number of experts consulted was mentioned, but not the number of questions posed to them. in these cases, i surmised that only one question was posed. in nine projects the number of questions posed to experts was mentioned but not the number of experts consulted, and therefore these were omitted from this calculation. 25 i interpreted the texts to the best of my ability. 26 this question clearly differs from the problems that arise in connection with the statistically analyzable selection of items or the assessment of their reliability. experts can also be consulted for this. however, such aspects are not the focus of this paper. 27 see for example the test of economic literacy (third edition) by walstad and rebeck (2001) including several items on further economic development subject to the conditions of the theory of competitive markets. 28 languages comprise several types of statements (e.g., interrogative, imperative, exclamative). as differentiated in linguistics truth-apt statements and statements requiring consent are subgroups of declarative statements. 29 unfortunately, statements like this may be interpreted as relaying a truth-apt message such as “item no. i is correct” (in the sense of no. 7f in table 1). this is due to the polysemy of most linguistic signs. additionally, the grammar of indo-european languages allows a speaker to use exactly the same grammatical structure to utter a truth-apt and a consent statement (e.g., “this is green.” and “this is wonderful.”). usually, the situational context eliminates this semantic deficit, but in written scientific texts it is necessary to make use of distinct and unequivocal formulations. 30 in linguistics, referred to as commands, i.e., imperative sentences. 31 e.g.: “this is item no. j and looking at it in respect to criterion y i strongly recommend it be reformulated!” 32 thus, definitions are nothing other than conditional norms originating from a personal decision about the meaning of a certain term supplemented by the condition if other members of the given language community are ready to join this proposal. 33 this becomes clear if one thinks of a person who valuates the length of days as being too short or the utility of a certain test as being unnecessary. 34 intuitions, in this case, may be based on and chosen from a large number of possible criteria. even if the experts agree on those criteria, verbalised intuitions are nothing other than the expression of individual preferences. 35 no. 1, 2, 4c, 7d. 36 in the 1980s a method called the index of item-objective congruence was developed for dealing with multiple experts’ judgements (thurn & dietz 1989, referring to rovinelli & hambleton, 1977). this index is calculated on the basis of qualitative expert judgments (an item is “definitely” a measure of a domain, “definitely not” a measure of a domain, or no decision on this question is possible: +1, 0, -1) resulting in a measure ranging from +1.0 to -1.0. the authors suggest a minimum index of +0.7 for good items (1989, p. 342), but they do not discuss whether or to what extent index values < 1.0 diminish the quality of decisions based on items with an index value of 0.7. 37 because the number of experts involved often is below 20 (see appendix 1, col. 4), it is not unlikely that a result like this can be reached. 38 it is often assumed that the longer the duration of service teachers have behind them the higher is their expertise. but, as siedentop and eldar (1989) have convincingly argued, this is by no means certain. duration of experience is not a reliable indicator for expertise. 39 for example, walstad and rebeck (2001) give the names and affiliations of all experts involved in the development of their “test of economic literacy” (pp. 3-4, 68). acknowledgements i would like to thank two anonymous reviewers and the editor who provided me with valuable advice on the presentation of my findings. references aera, apa, & ncme (1985). standards for educational and psychological testing. washington: apa. aera, apa, & ncme (2014). standards for educational and psychological testing. washington: aera. allen, m. j., yen, w. m. (2002). introduction to measurement theory (2nd ed.). prospect heights, il: waveland press. anastasi a., urbina s. (1997). psychological testing (7th ed.). new york, ny: prentice hall. anderson, d., irvin, s., alonzo, j., & tindal, g. a. (2015). gauging item alignment through online systems while controlling for rater effects. educational measurement: issues and practice, 34(1), 22–33. angoff, w. h. (1988). validity: an evolving concept. in h. wainer & h. braun (eds.), test validity (pp. 9–13). hillsdale, nj: lawrence erlbaum. baker, e. (2013). the chimera of validity. teachers’ college record, 115(9). beck, k. (1987). die empirischen grundlagen der unterrichtsforschung [the empirical foundations of research on classroom teaching. a critical analysis of the descriptive power of observation methods]. goettingen: hogrefe. beck, k., landenberger, m. & oser, f. (eds.) (2016). technoligiebasierte kompetenzmessung in der beruflichen bildung [technology-based measurement in vocational education and training]. bielefeld: bertelsmann. berk, r. (1990). importance of expert judgment in content-related validity evidence. western journal of nursing research, 12(5), 659–671. doi: org/10.1177/019394599001200507 brennan, r. l. (2013). commentary on “validating the interpretations and uses of test scores”. journal of educational measurement, 50(1), 74-83. dodd, j. (2002). truth. analytic philosophy, 43(4), 279-291. doi: 10.1111/1468-0149.00270 [https://onlinelibrary.wiley.com/doi/abs/10.1111/1468-0149.00270] ericsson, k. a. & smith, j. (1991). prospects and limits of the empirical study of expertise: an introduction. in k. a. ericsson & j. smith (eds.), toward a general theory of expertise (pp. 1-39). new york: cambridge univ. press. field, j. (2013). cognitive validity. in a. geranpayeh & l. taylor, examining listening. research and practice in assessing second language listening (pp. 77-151). cambridge, uk: cambridge univ. press gay, l. r. (1980). educational evaluation and measurement: competencies for analysis and application . columbus, oh: charles e. merrill. jackson, f., oppy, g. & smith, m. (1994). minimalism and truth aptness. mind, 103(411), 287-302. [https://www.jstor.org/stable/2253741?seq=1#page_scan_tab_contents] grant, j. s. & davis, l. l. (1997). selection and use of content experts for instrument development. research in nursing & health, 20, 269–274. guion, r. m. (1977). content validity: the source of my discontent. applied psychological measurement, 1(1), 1-10. doi.org/10.1177/014662167700100103 henig, j. r. (2013). the politics of testing when measures “go public”. teachers’ college record, 115(9), 1-11. kane, m. t. (2013a). validating the interpretations and uses of test scores. journal of educational measurement, 50(1), 1–73. kane, m. t. (2013b). validation as a pragmatic, scientific activity. journal of educational measurement, 50(1), 115–122. kerlinger f. n. (1986). foundations of behavioral research (3rd ed.). new york, ny: holt, rinehart, & winston lawshe, c. h. (1975). a quantitative approach to content validity. personnel psychology, 28, 563–575. lynn, m. (1986). determination and quantification of content validity. nursing research, 35, 382–385. maas, van der, h. l. j., kan, k.-j. & borsboom, d. (2014). intelligence is what the intelligence test measures. seriously. journal of intelligence, 2(1), 12-15. doi: https://doi.org/10.3390/jintelligence2010012 messick, s. (1987). validity. ets research report series. vol. 1987, issue 2, 1-108. http://onlinelibrary.wiley.com/doi/10.1002/j.2330-8516.1987.tb00244.x/abstract; date accessed 2018/05/06; doi: 10.1002/j.2330-8516.1987.tb00244.x) messick, s. (1990). validity of test interpretation and use. ets research report series. https://eric.ed.gov/?id=ed395031; date accessed 2019/03/30. newton, p. e. & shaw, s. d. (2014). validity in educational and psychological assessment. los angeles: sage. pant, h. a., zlatkin-troitschanskaia, o., lautenbach, c., toepper, m. & molerov, d. (eds.) (2016). modelling and measuring competencies in higher education – validation and methodological innovations (kokohs) – overview of the research projects (kokohs working papers, 11). berlin & mainz: humboldt university & johannes gutenberg university. http://www.kompetenzen-im-hochschulsektor.de/617_deu_html.php; date accessed 2018/03/20. popper, k.r. (1972). objective knowledge. oxford: clarendon. reynolds c. r., livingston r. b., willson v. (2009). measurement and assessment in education (2nd ed.). upper saddle river, nj: pearson. rovinelli, r. j. & hambleton, r. k. (1977). on the use of content specialists in the assessment of criterion-referenced test item validity. dutch journal for educational research, 2, 49-60. savigny, von, e. (19712). grundkurs im wissenschaftlichen definieren [basic course on scientific defining] . münchen: dtv. shavelson, r. j., gao, x. & baxter, g. p. (1995). on the content validity of performance assessments: centrality of domain specification. in m. birenbaum & f. douchy (eds.), alternatives in assessment of achievements, learning process, and prior knowledge (pp. 131–141). boston: kluwer academic. shepard, l. a. (2013). validity for what purpose? teachers’ college record, vol. 115(9), p. 1-12. http://www.tcrecord.org id number: 17116, date accessed: 2019/01/23. siedentop, d. & eldar, e. (1989). expertise, experience, and effectiveness. journal of teaching physical education, 8, 254-260. sireci, s. g. (1998). the construct of content validity. social indicators research, 45(1), 83-117. doi:org/10.1023/a:100698552 smith, m. d. (2017). cognitive validity: can multiple-choice items tap historical thinking processes? american educational research journal, 54(6), 1256-1287. doi: 10.3102/0002831217717949 thorn, d. w. & deitz, j. c. (1989). examining content validity through the use of content experts. the occupational therapy journal of research, 9, 334-346. vogt, d. s., king, d. w. & king, l. a. (2004). focus groups in psychological assessment: enhancing content validity by consulting members of the target population. psychological assessment, 16(3), 231-243. walstad, w. b. & rebeck, k. (2001). test of economic literacy. third edition. new york: national council on economics education. welner, k. g. (2013). consequential validity and the transformation of tests from measurement tools to policy tools. teachers’ college record, 115(9), p. 1-6. http://www.tcrecord.org id number: 17115, date accessed: 06.05.2015. white, m. c. (2018). rater performance standards for classroom observation instruments. educational researcher, 47(8), 492-501. wilson, f. r., pan, w. & schumsky, d. a. (2012). recalculation of the critical values for lawshe’s content validity ratio. measurement and evaluation in counseling and development, 45(3) 197–210. zlatkin-troitschanskaia, o., pant, h. a., nagel, th.-m., molerov, d., lautenbach, c. & toepper, m. (eds.) (2020). portfolio of kokohs assessemnts. test instruments for modelling and measuring domain-specific and generic competencies of higher education students and graduates. mainz & berlin. https://www.wihoforschung.de/_medien/downloads/kokohs_kompetenztest-verfahren_englisch.pdf appendix 1. reports on test development with regard to including experts* * chosen from two programs funded by the german ministry of education and research devoted to the development of instruments for measuring competences (academic students: kokohs: pant et al. 2016; apprentices: ascot: beck et al. 2016) plus individual projects conducted in germany on test development (time span 2014-2015) and two journals with a focus on the development of measurement instruments in psychology/education (time span 2009-2015): educational and psychological measurement (sage, los angeles et al.) and measurement and evaluation in counselling and development, (sage, los angeles et al.) ° not specified in further detail appendix 2. sources a projects conducted as part of the kokohs program on modeling and measuring competences/skills in higher education (2014 – 2015) [in alphabetical order according to acronym] (1) siebert-ott, g., decker, l., kaplan, i., & macha, k. (2015). akademische textkompetenzen bei studienanfängern und fortgeschrittenen studierenden des lehramtes (akatex) – kompetenzmodellierung und erste ergebnisse der kompetenzerfassung. in u. riegel, s. schubert, g. siebert-ott & k. macha (hrsg.), kompetenzmodellierung und kompetenzmessung in den fachdidaktiken (s. 257-273). münster: waxmann. akatex (2) lohse-bossenz, h., kunina-habenicht, o., & kunter, m. (2013). the role of educational psychology in teacher education: expert opinions on what teachers should know about learning, development, and assessment. european journal of psychology of education, 28, 1543-1565. doi: 10.1007/s10212-013-0181-6. bilwiss (3) hammer, s., carlson, s.a., ehmke, t., koch-priewe, b., koeker, a., ohm, u., rosenbrock, s., & schulze, n. (2015). kompetenz von lehramtsstudierenden in deutsch als zweitsprache. in s. blömeke & o. zlatkin-troitschanskaia, kompetenzen von studierenden (s. 33-54). zeitschrift für pädagogik. beiheft 61. weinheim: beltz juventa. dazkom (4) eggert, s. & boegeholz, s. (2009). students' use of decision‐making strategies with regard to socioscientific issues: an application of the rasch partial credit model. science education, 94(2), 230-258. exmo (5) fritsch, s., berger, s., seifried, j., bouley, f., wuttke, e., schnick-vollmer, k., & schmitz, b. (2015). the impact of university teacher training on prospective teachers’ ck and pck – a comparison between austria and germany. empirical research in vocational education and training, 7(4), 1-20. http://www.ervet-journal.com/content/7/1/4; doi: 10.1186/s40461-015-0014-8. komewp (6) jenßen, l., dunekacke, s., & blömeke, s. (2015). qualitätssicherung in der kompetenzforschung. in s. blömeke & o. zlatkin-troitschanskaia, kompetenzen von studierenden (s. 11-31). zeitschrift für pädagogik. beiheft 61. weinheim: beltz juventa. komma (7) schwippert, k., braun, e., prinz, d., schaeper, h., fickermann, d., pfeiffer, j., & brachem, j.-c. (2014). kompaed tätigkeitsbezogene kompetenzen in pädagogischen handlungsfeldern. die deutsche schule, 106(1), 72-84. kompaed (8) trempler, k., hetmanek, a., wecker, c., kiesewetter, j., wermelt, m., fischer, f., fischer, m., & graesel, c. (2015). nutzung von evidenz im bildungsbereich. validierung eines instruments zur erfassung von kompetenzen der informationsauswahl und bewertung von studien. in s. blömeke & o. zlatkin-troitschanskaia, kompetenzen von studierenden (s. 144-166). zeitschrift für paedagogik. beiheft 61. weinheim: beltz juventa. kompare (9) neumann, i., rösken-winter, b., lehmann, m., duchhardt, c., heinze, a., & nickolaus, r. (2015). measuring mathematical competences of engineering students at the beginning of their studies. peabody journal of education, 90(4), 465-476. doi: 10.1080/0161956x.2015.1068054. kom@ing (10) schroeder, s., richter, t. & hoever, i. (2008). getting a picture that is both accurate and stable: situation models and epistemic validation. journal of memory and language, 59, 237-255. doi: 10.1016/j.jml.2008.05.001. koswo (11) hartmann, s., upmeier zu belzen, a., krüger, d., & pant, h. a. (2015). scientific reasoning in higher education. constructing and evaluating the criterion-related validity of an assessment of preservice science teachers’ competencies. zeitschrift für psychologie, 223(1), 47-53. doii: 10.1027/2151-2604/a000199. ko-wadis (12) bender, e., hubwieser, p., schaper, n., margaritis, m., berges, m., ohrndorf, l., magenheim, j., & schubert, s. (2015). towards a competency model for teaching computer science. peabody journal of education, 90(4), 519-532. doi: 10.1080/0161956x.2015.1068082. kui (13) winter-hölzl, a., waeschle, k., wittwer, j., watermann, r., & nueckles, m. (2015). entwicklung und validierung eines tests zur erfassung des genrewissens studierender und promovierender der bildungswissenschaften. in s. blömeke & o. zlatkin-troitschanskaia, kompetenzen von studierenden (s. 185-202). zeitschrift für pädagogik. beiheft 61. weinheim: beltz juventa. lssced (14) tiede, j., grafe, s., & hobbs, r. (2015). pedagogical media competencies of preservice teachers in germany and the united states: a comparative analysis of theory and practice, peabody journal of education, 90(4), 533-545. doi: 10.1080/0161956x.2015.1068083. m3k (15) taskinen, p. h., steimel, j., gräfe, l., engell, s., & frey, a. (2015). a competency model for process dynamics and control and its use for test construction at university level, peabody journal of education, 90(4), 477-490. doi: 10.1080/0161956x.2015.1068074. mokomasch (16) schladitz s., gross, j., & wirtz, m. (2015). konstruktvalidierung eines tests zur messung bildungswissenschaftlicher forschungskompetenz. in s. blömeke & o. zlatkin-troitschanskaia, kompetenzen von studierenden (s. 167-184). zeitschrift für pädagogik. beiheft 61. weinheim: beltz juventa. profile-p (17) steuer, g., engelschalk, t., joestl, g., roth, a., wimmer, b., schmitz, b., schober, b., spiel, c., ziegler, a., & dresel, m. (2015). kompetenzen zum selbstregulierten studium. in s. blömeke & o. zlatkin-troitschanskaia, kompetenzen von studierenden (s. 203-225). zeitschrift für pädagogik. beiheft 61. weinheim: beltz juventa. pro-srl (18) beck, k., landenberger, m. & oser, f. (hrsg.) (2016). technologiebasierte kompetenzmessung in der beruflichen bildung. bielefeld: bertelsmann.. wiwikom b projects of the german ascot program on measuring and modelling vocational competences in different occupations (2014 – 2015) [in alphabetical order according to acronym] (19) wuttke, e., seifried, j., brandt, s., rausch, a., sembill, d., martens, t., & wolf, k. (2015). modellierung und messung domänenspezifischer problemlösekompetenz bei angehenden industriekaufleuten – entwicklung eines testinstruments und erste befunde zu kognitiven kompetenzfacetten. zeitschrift für berufs und wirtschaftspaedagogik, 111(2), 189-207. dompl-ik (20) walker, f., link, n., & nickolaus, r. (2015). berufsfachliche kompetenzstrukturen bei elektronikern für automatisierungstechnik am ende der berufsausbildung. zeitschrift für berufsund wirtschaftspaedagogik, 111(2), 222-241. koko-ea (21) schmidt, t., nickolaus, r., & weber, w. (2014). modellierung und entwicklung des fachsystematischen und handlungsbezogenen fachwissens von kfz-mechatronikern. zeitschrift für berufsund wirtschafts¬paedagogik, 110, 549-574. koko-kfz (22) wittmann, e., weyland, u., nauerth, a., döring, o., rechenbach, s., simon, j., & worofka, i. (2014). kompetenzerfassung in der pflege älterer menschen – theoretische und domänenspezifische anforderungen der aufgabenmodellierung. in j. seifried, u. fasshauer & s. seeber (hrsg.), jahrbuch der berufsund wirtschaftspädagogischen forschung 2014 (s. 53-66). opladen: barbara budrich. tema c other research projects conducted in germany on test development in the field of (higher) education (2014 – 2015) [in alphabetical order according to first author] (23) esslinger, g. (2015). kommas beim lesen verarbeiten können – eine vernachlässigte teilkompetenz allgemeiner lesefähigkeit. in u. riegel, s. schubert, g. siebert-ott, & k. macha (hrsg.), kompetenzmodellierung und kompetenzmessung in den fachdidaktiken (s. 274-292). münster: waxmann. (24) filipiak, a. & reis, o. (2015). was lernen studierende in der systematischen theologie? kompetenzdiagnostik in der religionslehrerbildung. in u. riegel, s. schubert, g. siebert-ott, & k. macha (hrsg.), kompetenzmodellierung und kompetenzmessung in den fachdidaktiken (s. 227-241). münster: waxmann. (25) kuhn, c. (2014). fachdidaktisches wissen von lehrkräften im kaufmännisch-verwaltenden bereich. modellbasierte testentwicklung und validierung. landau: vep (26) lindl, a. & kloiber, h. (2015). erste schritte zur kompetenzmessung von lateinlehrkräften. in u. riegel, s. schubert, g. siebert-ott, & k. macha (hrsg.), kompetenzmodellierung und kompetenzmessung in den fachdidaktiken (s. 293-305). münster: waxmann. (27) vogler, j., messmer, r., & allemann, d. (2017). das fachdidaktische wissen und können von sportlehrpersonen (pck-sport). german journal of exercise and sport research, 47(4), 335-347.   d journal educational and psychological measurement, sec. “validity studies”, 2009(1) – 2015(6) [in chronological order of release date] (28) ordoñez, x. g., ponsoda, v., abad, f. j., & romero, s. j. (2009). measurement of epistemological beliefs: psychometric properties of the eqebi test scores. educational and psychological measurement, 69(2), 287-302. doi: 10.1177/0013164408323226 (29) liu, o. l., minsky, j., ling, g., & kyllonen, p. (2009). using the standardized letters of recommendation in selection: results from a multidimensional rasch model. educational and psychological measurement, 69(3), 475-492. doi: 10.1177/0013164408322031 (30) lei, p.-w., wu, q., diperna, j., & morgan, p. l. (2009). developing short forms of the earli numeracy measures. comparison of item selection methods. educational and psychological measurement, 69(5), 825-842. doi: 10.1177/0013164409332215 (31) aydın, y. ç. & uzuntiryaki, e. (2009). development and psychometric evaluation of the high school chemistry self-efficacy scale. educational and psychological measurement, 69(5), 868-880. doi: 10.1177/0013164409332213 (32) hulpia, h., devos, g., & rosseel, y. (2009). development and validation of scores on the distributed leadership inventory. educational and psychological measurement, 69(6), 1013-1034. doi: 10.1177/0013164409344490 (33) cadiz, d., sawyer, j. e., & griffith, t. l. (2009). developing and validating field measurement scales for absorptive capacity and experienced community of practice. educational and psychological measurement, 69(6), 1035-1058. doi: 10.1177/0013164409344494 (34) fletcher, th. d. & nusbaum, d. n. (2010). development of the competitive work environment scale: a multidimensional climate construct. educational and psychological measurement, 70(1), 105-124. doi: 10.1177/0013164409344492 (35) luttrell, v. r., callen, b. w., allen, c. s., wood, m. d., deeds, d. g., & richard, d. c. s. (2010). the mathematics value inventory for general education students: development and initial validation. educational and psychological measurement, 70(1), 142-160. doi: 10.1177/0013164409344526 (36) myers, n. d., chase, m. a., beauchamp, m. r., & jackson, b. (2010). athletes’ perceptions of coaching competency scale ii-high school teams. educational and psychological measurement, 70(3), 477-494. doi: 10.1177/0013164409344520 (37) mesmer-magnus, j., murase, t., dechurch, l. a., & jiménez, m. (2010). coworker informal work accommodations to family: scale development and validation. educational and psychological measurement, 70(3), 511-531. doi: 10.1177/0013164409355687 (38) linnenbrink-garcia, l., durik, a. m., conley, a. m., barron, k. e., tauer, j. m., karabenick, s. a., & harackiewicz, j. m. (2010). measuring situational interest in academic domains. educational and psychological measurement, 70(4), 647-671. doi: 10.1177/0013164409355699 (39) teo, t. (2010). the development, validation, and analysis of measurement invariance of the technology acceptance measure for preservice teachers (tampst). educational and psychological measurement, 70(6), 990-1006. doi: 10.1177/0013164410378087 (40) mcdermott, p. a., fantuzzo, j. w., warley, h. p., waterman, c., angelo, l. e., gadsden, v. l., & sekino, y. (2001). multidimensionality of teachers’ graded responses for preschoolers’ stylistic learning behavior: the learning-to-learn scales. educational and psychological measurement, 71(1), 144-169. doi: 10.1177/0013164410387351 (41) cheng, y.-y., chen, l.-m. liu, k.-s., & chen, y.-l. (2011). development and psychometric evaluation of the school bullying scales: a rasch measurement approach. educational and psychological measurement, 71(1), 200-216. doi: 10.1177/0013164410387387 (42) gable, r. k., ludlow, l. h., mccoach, d. b., & kite, s. l. (2011). development and validation of the survey of knowledge of internet risk and internet behavior. educational and psychological measurement, 71(1), 217-230. doi: 10.1177/0013164410387389 (43) nilsson, j. e., marszalek, j. m., linnemeyer, r. m., bahner, a. d., & misialek, l. h. (2011). development and assessment of the social issues advocacy scale. educational and psychological measurement, 71(1), 258-275. doi: 10.1177/0013164410391581 (44) curşeu, p. l. & schruijer, s. g. (2012). decision styles and rationality: an analysis of the predictive validity of the general decision-making style inventory. educational and psychological measurement, 72(6), 1053-1062. doi: 10.1177/0013164412448066 (45) warner, j. a., koufteros, x., & verghese, a. (2014). learning computerese: the role of second language learning aptitude in technology acceptance. educational and psychological measurement, 74(6), 991-1017. doi: 10.1177/0013164414520629 (46) paulhus, d. l. & dubois, p. j. (2014). application of the overclaiming technique to scholastic assessment. educational and psychological measurement, 74(6), 975-990. doi: 10.1177/0013164414536184 (47) kersting, n. b., sherin, b. l., & stigler, j. w. (2014). automated scoring of teachers’ open-ended responses to video prompts: bringing the classroom-video-analysis assessment to scale. educational and psychological measurement, 74(6), 950-974. doi: 10.1177/0013164414521634 (48) nezhnov, p., kardanova, e., vasilyeva, m., & ludlow, l. (2015). operationalizing levels of academic mastery based on vygotsky’s theory: the study of mathematical knowledge. educational and psychological measurement, 75(2), 235-259. doi: 10.1177/0013164414534068 (49) dimitrov, d. m., raykov, t., & al-qataee, a. a. (2015). developing a measure of general academic ability. educational and psychological measurement, 75(3), 475-490. e journal measurement and evaluation in counselling and development, sec. “assessment, development, and validation”, 2009(1) – 2015(4) [in chronological order by release date] (50) kim, b. s., soliz, a., orellana, b., & alamilla, s. g. (2009). latino/a values scale. development, reliability, and validity. measurement and evaluation in counseling and development, 42(2), 71-91. doi: 10.1177/0748175609336861 (51) tovar, e., simon, m. a., & lee, h. b. (2009). development and validation of the college mattering inventory with diverse urban college students. measurement and evaluation in counseling and development, 42(3), 154-178. doi: 10.1177/0748175609344091 (52) pistole, m. c. & roberts, a. (2011). measuring long-distance romantic relationships: a validity study. measure. measurement and evaluation in counseling and development, 44(2), 63-76. doi: 10.1177/0748175611400288 (53) kopp, j. p., zinn, t. e., finney, s. j., & jurich, d. p. (2011). the development and evaluation of the academic entitlement questionnaire. measurement and evaluation in counseling and development, 44(2), 105-129. doi: 10.1177/0748175611400292 (54) kim, s.-h., sherry, a. r., lee, y.-s., & kim, c.-d. (2011). psychometric properties of a translated korean adult attachment measure. measurement and evaluation in counseling and development, 44(3), 135-150. doi: 10.1177/0748175611409842 (55) adelson, j. l. & mccoach, d. b. (2011). development and psychometric properties of the math and me survey: measuring third through sixth graders’ attitudes toward mathematics. measurement and evaluation in counseling and development, 44(4), 225-247. doi: 10.1177/0748175611418522 (56) neto, f. (2012). the satisfaction with sex life scale. measurement and evaluation in counseling and development, 45(1), 18-31. doi: 10.1177/0748175611422898 (57) sun, q., ng, k.-m., & wang, c. (2012). a validation study on a new chinese version of the dispositional hope scale. measurement and evaluation in counseling and development, 45(2), 133-148. doi: 10.1177/0748175611429011 (58) locke, b. d., mcaleavey, a. a., zhao, y., lei, p.-w., hayes, a., castonguay, l. g., li, h., tate, r., & lin, y.-c. (2012). development and initial validation of the counseling center assessment of psychological symptoms–34. measurement and evaluation in counseling and development, 45(3), 151-169. doi: 10.1177/0748175611432642 (59) erford, b. t. & alsamadi, s. c. (2012). the screening test for emotional problems–parent report (step-p). studies of reliability and validity. measurement and evaluation in counseling and development, 45(3), 170-180. doi: 10.1177/0748175611432643 (60) hardesty, p. h. & richardsin, g. b. (2012). the structure and validity of the multidimensional social support questionnaire. measurement and evaluation in counseling and development, 45(3), 181-196. doi: 10.1177/0748175612441214 (61) lim, s. y. & chapman, e. (2013). an investigation of the fennema-sherman mathematics anxiety subscale. measurement and evaluation in counseling and development, 46(1), 26–37. doi: 10.1177/0748175612459198 (62) ng, p., su, x. s., chan, v., leung, h., cheung, w., & tsun, a. (2013). the reliability and validity of a campus caring instrument developed for undergraduate students in hong kong. measurement and evaluation in counseling and development, 46(2), 88–100. doi: 10.1177/0748175612467463 (63) jacobs, k. & struyf, e. (2013). measuring integrated socioemotional guidance at school: factor structure and reliability of the socioemotional guidance questionnaire (seg-q). measurement and evaluation in counseling and development, 46(3), 159–177. doi: 10.1177/0748175613481978 (64) rantanen, a. p. & soini, h. s. (2013). development of the response observation system. measurement and evaluation in counseling and development, 46(4), 247–260. doi: 10.1177/0748175613484041 (65) balkin, r. s., harris, n. a., freeman, s. j., & huntington, s. (2014). the forgiveness reconciliation inventory: an instrument to process through issues of forgiveness and conflict. measurement and evaluation in counseling and development, 47(1), 3–13. doi: 10.1177/0748175613497037 (66) boudreaux, d. j., dahlen, e. r., madson, m. b., & yowell, e. b. (2014). attitudes toward anger management scale: development and initial validation. measurement and evaluation in counseling and development, 47(1), 14–26. doi: 10.1177/0748175613497039 (67) montes, s. a., ledesma, r. d., garcia, n. m., & poó, f. m. (2014). the mindful attention awareness scale (maas) in an argentine population. measurement and evaluation in counseling and development, 47(1), 43–51. doi: 10.1177/0748175613513806 (68) carey, j., brigman, g., webb, l., villares, e., & harrington, k. (2014). development of an instrument to measure student use of academic success skills: an exploratory factor analysis. measurement and evaluation incounseling and development, 47(3), 171–180. doi: 10.1177/0748175613505622 (69) dominguez espinosa, a. & van de vijver, f. j. r. (2014). an indigenous social desirability scale. measurement and evaluation in counseling and development, 47(3), 199–214. doi: 10.1177/0748175614522267 (70) ahn, c. m., ebesutani, c., & kamphausr. w. (2014). a psychometric analysis and standardization of the behavior assessment system for children-2, self-report of personality, college version, among a korean sample. measurement and evaluation in counseling and development, 47(3), 226–244. doi: 10.1177/0748175614531797 (71) huang, y.-c. & lin, s.-h. (2015). development and validation of an inventory for measuring student attitudes toward calculus. measurement and evaluation in counseling and development, 48(2), 109–123. doi: 10.1177/0748175614563314 (72) jenkins-guarnieri, m. a., vaughan, a. l., & wright, s. l. (2015). development of a self-determination measure for college students: validity evidence for the basic needs satisfaction at college scale. measurement and evaluation in counseling and development, 48(4), 266-284. doi: 10.1177/0748175615578737 appendix 3. frequency of requests addressed to experts note. * cf. table 1: to: truth-apt; tu: truth-apt, but unknown; s: subjective microsoft word wegerif_publication.docx             frontline  learning  research  vol.4  no.  4  special  issue  (2016)  59  -­‐  61   issn  2295-­‐3159       doi: http://dx.doi.org/10.14786/flr.v4i4.319   commentary: expanding conceptualizations for the study of learning rupert wegerif university of cambridge, uk there are always theoretical assumptions involved in research. theoretical assumptions determine which phenomena are visible and which are invisible and they make different educational goals and pedagogical strategies either thinkable or unthinkable. this collection of articles expands the dialogue about educational research by exploring a range of different ways of conceptually framing education. each different conceptualisation reveals a different set of objects and relationships and so opens up different ways of thinking about educational goals. for example, tsafrir goldberg and baruch schwarz bring emotion into the frame in a study which suggests that conventional approaches to education that ignore emotion do so at a cost. goldberg and schwarz suggest, with evidence, that by engaging with emotion explicitly we can improve the quality of argumentation and so have a cognitive learning gain. this example illustrates the value of seeing education through a new conceptual frame. implicit in their account is the potential to understand educational goals in a new way, switching from thinking of education in terms of cognitive development to thinking of it in terms of emotional development. a similar account can be given of each of the articles in this special issue. each claims to enrich our understanding by foregrounding a different aspect of the complex whole of education thereby enabling us to think in new ways about the nature of education and the goals of education. implicit in this first paragraph, and in the introduction to this special issue, is the assumption that expanding the dialogue about education with a range of new conceptualisations is a good thing. but there is an obvious possible challenge to that assumption. surely the goal of scientific research should not be to proliferate a range of perspectives but to tell us which one is true, or at least more true than the others? this special issue offers us the choice to look at education through the framework of patterns of space and time (giuseppe ritella, beatrice ligorio and kai hakkarainen), social practices (william penuel, daniela digiacomo, katie van horne and ben kirshner), ecology (crina damsa and alfredo jornet) imagination (jaakko hilppö, antti rajala, tania zittoun, kristiina kumpulainen and lasse lipponen) as well as the already mentioned focus on emotion (tsafrir goldberg and baruch schwarz). but in this proliferation of possible perspectives the question must arise which one or which ones should we choose to focus upon and on what basis can we make that choice? in order to distinguish science from what he called ‘pseudo science’, karl popper used the example of astrology (popper, 1963). astrology offers a complex language for describing reality in terms of forces that are said to underlie and control both emotions and events. i do not think that astrology could have survived so long or be so popular if it did not serve some kind of function and that function seems to be helping people feel as if they understand their lives. the example of astrology teaches us that we need to distinguish between the comforting illusion of mastery provided by any analytic framework whatsoever and a more scientific understanding that actually works because it is not simply a social construction but based on correspondence with some sort of underlying reality or transcendental context. commentary  wegerif       60 | f l r     the social anthropologist marshall sahlins pointed out the danger of ontology recapitulating methodology in social science research (1976, p. 89). he made this quip as a criticism of ecological analysis. his point was that first we see everything in terms of interacting organisms and then, after a long complex study, we conclude that everything seems to be interacting organisms. the same could apply to chronotopic analysis. first we assume that everything is patterns of space and time and then we conclude our study by pointing out ‘so you see everything turns out to be patterns of space and time’. the point is not, can we look at things this way, but what do we gain if we look at things this way? given that there are always an infinite number of possible ways of looking at things why should we invest in this one? part of the claim made as to the value of chronotopic analysis is as a contrast to assuming the hegemony of a single ‘objective’ or ‘physical’ frame of space and time. this is perhaps just a way of acknowledging that different perspectives can be different worlds, not simply points of view within an already given world. this is clear in the example that ritella et al give of an education project with aboriginal canadians where local space-time conflicted creatively with scientific space-time. for bakhtin, the originator of the concept of chronotope, dialogues were not so much dialogues between people as dialogues between chronotopes. if dialogues are dialogues between chronotopes then there must be a chronotope of chronotopes. bakthin called this ‘great time’: basically the idea of the dialogic space within which different cultures can learn from each other across apparent external distances in space and time. great time is precisely not the idea of an encompassing context that is monologic, like the idea of an objective overarching context of space and time. but how can we think this dialogic meta-context? perhaps we cannot think it but we can feel it. in a typically gnomic sentence in his late notes bakhtin refers indirectly to his concept of great time when he writes: ‘the unspoken truth in dostoevsky (christ’s kiss). the problem of silence.’ (bakhtin, 1986, p. 148). this refers us to a story within a story. in doestoevsky’s the brothers karamazov ivan tells the tale of how christ, having returned to earth in seville at the time of the counter-reformation, has been arrested by the spanish inquisition. the grand inquisitor visits his cell and explains his crime. apparently christ’s mission of freeing humanity was cruel since most humans could not cope with such freedom. instead they needed the meta-narrative provided by the church. christ listened in silence to his condemnation and responded only with a single kiss. in this image of christ’s silent kiss we have a radically different way of thinking about education. it suggests a goal beyond the merely cognitive. the ability to understand others’ perspectives without becoming lost in them. an understanding that is not verbally articulated but takes the form of love. i suspect that this is a goal for education that follows logically from adopting a chronotopic frame of analysis but i wonder what the authors of the paper think about that. perhaps we need a proliferation of different conceptualisations because we find ourselves within education and we cannot have an overview. there is no single master voice or ruling chronotope that we can rely upon. but this does not mean that anything goes and we should treat astrology as if it was science. goldberg and schwarz again show a way forward. not content simply with conceptualising education differently they turn their conceptualisation into an experimental study that tests its value. their comparison of three methods of education suggest, as a working hypothesis at least, that some ways of educating lead to a more fruitful combination of emotion and cognition than other ways of educating. in my view this is the sort of study that all these different conceptualisation could do and ought to do in order to demonstrate their fruitfulness. it is not enough to say, look this is interesting and seems to make sense – astrologers could say the same. implicit in putting forward a way of conceptualising education is the claim that this is fruitful for education and if so i think this needs to be explored and demonstrated further with empirical research. popper used the example of astrology as a pseudo-science in order to make a contrast with real science. pseudo-sciences like astrology, psychoanalysis, marxism and so on were so vague in their claims that they could not be falsified by any evidence. it is hard to see how conceptualising education in terms of space and time or ecology or imagination or social practice could ever be directly falsified. however this does not mean that such conceptualisations are not part of the larger dialogue of educational science. lakatos argued, against popper, that the different ways of conceptualising things represented by different scientific ‘programmes’ did in fact compete and evolve even though they could not be falsified directly. the authors of commentary  wegerif       61 | f l r     the different ways of conceptualising education in these studies are all claiming that their way of thinking is fruitful for education. once we acknowledge in all humility that we do not know the truth and can never know it completely then it becomes plausible that a range of different ways of approaching the true might be useful and might in fact, in their own different ways, all be partly true or partly good even while remaining, from our limited point of view, apparently incompatible with each other. in other words the proliferation of many conceptualisations of learning reflects what bakhtin might call a polyphonic truth where the truth is not to be looked for in any one voice but in the ongoing dialogue. normally design-based research seems to assume a given world and just manipulate a few variables within that world. but there is no reason why we cannot extend the design-based research approach to include dialogue between conceptualisations of education. ultimately we will only know if these different ways of conceptualising education are useful or not if they lead to fruitful consequences. and how do we judge if a consequence is fruitful you ask? i can see why you want clear criteria but that is not how it works. the truth of education is not ultimately a question of theory but a question to be answered by life and by how we live together and construct our future lives together. references bakhtin, m.m. (1986) speech genres and other late essays. trans. vern w. mcgee. austin, tx: university of texas press. popper, k. (1963) conjectures and refutations. abingdon: routledge. sahlins, m. (1976). culture and practical reason. university of chicago press. microsoft word proofsstenalt.docx frontline learning research vol. 9 no. 3 (2021) 52 68 issn 2295-3159 info corresponding author: maria hvid stenalt, department of science education, university of copenhagen, denmark email: mhs@ind.ku.dk doi: https://doi.org/10.14786/flr.v9i3.697 digital student agency: approaching agency in digital contexts from a critical perspective maria hvid stenalt department of science education, university of copenhagen, denmark article received 3 august 2020 / article revised 12 april 2021/ accepted 1 july / available online 30 july abstract developing student agency is a critical aspect of higher education and, in particular, digital education. in this sense, the capacity to understand what constitutes agency in digital contexts of education and evaluate students’ digital agency is now crucial. in contrast to traditional approaches to student agency in digital contexts that subsume technologies to educational intentions, media research has illustrated a more complex interplay between humans and technology. drawing on this insight, the paper argues for a more critical disposition to digital student agency, wherein relational, cultural, and technological dynamics are central to agency. specifically, the article proposes a framework for digital student agency that distinguishes five critical domains to student agency in digital contexts: (1) agentic possibility, (2) digital self-representation, (3) data uses, (4) digital sociality, and (5) digital temporality. the article concludes by outlining the implications of the framework for educational practice and academic research around student agency and student learning. specifically, adopting the framework implies changes in how we investigate student agency in digital contexts and enables critical investigations of student-centred teaching practices. keywords: student agency, networked publics, learning, educational technologies, digital education stenalt 53 | f l r 1. introduction cultivating individuals’ capacity to intervene in and transform given frames of action is key in higher education strategies and teaching that seeks to develop human beings’ capacity to act agentically (damşa et al., 2010; klemenčič, 2017; oecd, 2018). as research exploring student agency in higher education increases, recognition is growing that integration of agency-supportive practices and learning environments is essential for the cultivation of agentic students in higher education (jääskelä, poikkeus, et al., 2020; marín et al., 2020; toom et al., 2017). in many ways, the prevalence of digital technologies in contemporary higher education offers new grounds for supporting student agency. digitalisation plays a vital role in expanding the range of course delivery formats, from campus-based delivery to fully online, hybrid, or blended courses, facilitating flexibility on the learner’s part (kirkwood & price, 2011). moreover, online courses can present students with consequential choices and adapt ‘both the course curriculum and assessment to accommodate those choices’ (lindgren & mcdaniel, 2012, p. 346). this expansion of student choice and redistribution of initiative from educational institutions to the learner is seen to support student agency (bandura, 2002; irvine et al., 2013). another example is learning analytics. as jääskelä, heilala, et al. (2020) explained, collecting and analysing educational data can provide feedback on student progress, promote students’ agentic awareness, and tailor education to students’ needs. moreover, using technologies to facilitate student-centred learning and teaching links to agency. as defined by klemenčič et al. (2020, p. 33), student-centred education concerns ‘the capability of students to participate in, influence, and take responsibility for their learning pathways and environments, in order to achieve the expected learning outcomes’. providing students with opportunities for active participation and involvement in their own learning is central for much use of digital technologies such as student response systems that enable immediate student feedback and digital platforms that support collaborative processes. crucially, such teaching practices are discursively linked to the empowerment of students (starkey, 2019). while several studies investigating digital technologies include student agency, only few pay attention to underpinning their utilisation of the concept theoretically (marín et al., 2020). some studies have provided a clear theoretical account of student agency, but these studies tend to subsume the digital to educational intentions and settings. for example, irvine, code, and richards (2013) framed multiaccess learning as an opportunity for student choice, enabling face-to-face students and distancelearning students to access course materials and other participants and personalise learning. in this study, the context of the course and social processes were presented in terms of delivery formats, overall student characteristics, and concerns of the teacher. lindgren and mcdaniel (2012) explored whether online learning can be improved by employing narratives and student agency as prominent design features. similar to irvine et al. (2013), the specific interplay between agency and the context of action was subsumed in representations of learning activities, the steps involved in completing assignments, and the specific technologies used. in a study by luo et al. (2019), the role of student agency in a course utilising a flipped learning format was presented and analytically considered in relation to the educational instructions. hamilton and friesen (2013) describe this approach to technology as technological instrumentalism, where the technology is ‘interpreted in light of this or that pedagogical framework or principle and measured against how well they correspond in practice to that framework or principle, and technologies are neutral means employed for ends determined independently by their users’ (hamilton & friesen, 2013, p. 3). as bayne (2015) has reasoned, the disassembly of technology from social activity overlooks the epistemological consequences of technology use. crucially, it isolates the social and meaning-making aspects of digital education from the material (fenwick & landri, 2012). studies that focus on digital agency (passey et al., 2018; shonfeld et al., 2017) emphasise agency as a requirement for and through education. more specifically, digital agency refers to having the necessary digital competencies, digital confidence, and digital accountability to control and adapt to the digital world as an individual. however, this framing of digital agency pays little attention to agency in education and how the digital affects humans. stenalt 54 | f l r 2. aim the relationship between agency and digital contexts of learning has been given little attention in research addressing student agency and digital technology. a tendency has been to subsume the digital to educational intentions, frame digital agency as competencies required to control the digital world, or skip definitions of agency. these perspectives neither consider how the digital affects agency in contexts of learning nor offer support for developing agency-supportive teaching practices. to overcome this gap, this papers sets out an alternative conceptualisation of the digital in relation to student agency. it builds upon insights from media research and frames student actions in digital contexts as socially, technologically, and contextually configured and shaped through student negotiations of agentic power and will. thus, this theoretical paper combines understandings of the digital from media research with versions of agency from student agency research to cultivate empirical investigations that account for the agency-structure interplay in digital contexts of education and consider the possible implications of technology use. also, by drawing more explicitly on media research, it becomes possible to add a relational and participatory dimension to agency that allows us to identify and analyse the conflicts that might lead students to maladaptive practices to protect themselves from forms of digital participation. examples from research and discussions of implications are included to illustrate the theoretical approach suggested. the starting point of the paper is not to develop a new definition of agency as such but to broaden understandings of agency as a primarily individual phenomenon with a more relational understanding of agency (burkitt, 2016; stenalt, 2021), building on the premises of the digital. moreover, while emerging into an arena with blurry boundaries between human and nonhuman agency (fenwick & landri, 2012; leonardi, 2010), the paper focuses on students as the primary research object and uses conceptualisations of human agency as a stepping stone. lastly, the paper should not be seen as attempting to present a complete framework for understanding and supporting student learning in digital contexts. instead, it is intended to serve as a supplement to research and designs for learning. with this in mind, the remainder of the article proceeds in four iterative stages. first, it maps distinct approaches to and dimensions of agency. second, key dynamics of digital engagement from media research are presented. third, the key ideas are used to develop a theoretical framework that includes five domains and suggests a nuanced way of approaching the topic of digital student agency. fourth, having worked through the different domains, the paper discusses possible implications for practice and research. 3. what is student agency? defining student agency is by no means straightforward because student agency has been conceptualised in various ways within higher education literature (nieminen et al., 2021). in constructing the outset of this research, this paper draws on the frameworks typically adopted in higher education research (jääskelä et al., 2016; klemenčič, 2015). in the context of higher education, a sociological approach typically pays attention to the dualistic interplay between humans and structure and the ways power structures and structural factors impact human agency (hitlin & long, 2009). agency is understood as something that actors always have; however, the possibility of fulfilling personal goals is conditional on external structures. higher education studies conceptualising agency in this way tend to foreground individual students’ pursuit of personal, educational success within a context of macrosocial structures. for instance, calitz et al. (2016) examined patterns of unequal participation for working-class first-generation students at a south african university, applying sen’s (2005) capability approach to make sense of the relationship between a person and the social forces that can hinder or enable them to convert resources into capabilities. such forces include physical or mental stenalt 55 | f l r disabilities, variations in available nonpersonal resources, environmental variations, and differences in relative social positions. arkoudis and tran (2007) applied positioning theory and notions of moral agency (harré & van langenhove, 1999) to move away from static and stereotyped descriptions of chinese international students studying in australian higher education and toward understandings that depict these students as actively involved in meaning-making. analysing how students intentionally position themselves in relation to their lecturers and university expectations, these authors highlighted issues of mismatching between institutional expectations and students’ struggles. sociocognitive studies rest on a nondualistic, agentic perspective. in it, human development involves an agent intentionally influencing their life circumstances through individual psychological processes of development or knowledge acquisition (bandura, 2006). bandura (2006) argued, ‘through cognitive self-regulation, humans can create visualised futures that act on the present; construct, evaluate, and modify alternative courses of action to secure valued outcomes; and override environmental influences’ (p. 164). studies proposing this conceptualisation of agency foreground the psychological exercise of self-management through self-reflection, self-regulation, and self-efficacy (bandura, 2001, 2006; zimmerman, 1995). self-management is understood to depend on students’ ability to generate goals for their engagement through cognitive representations of desired future states that match their individual strengths and preferences (zimmerman, 1995). self-efficacy comprises students’ beliefs about their capability to succeed in what is required of them (bandura, 2006). bandura (2006) described the sources of self-efficacy to be (a) enactive mastery experiences (actual performances); (b) social comparison and modelling based on observation of others and their successes and failures (vicarious experiences); (c) forms of persuasion, both verbal and otherwise if within realistic bounds; and (d) interpretation of one’s own physiological and affective states. for instance, malmberg and hagger (2009) investigated supportive and instructional agency beliefs, defined as the perceived ability to facilitate others’ learning. nye et al. (2011) framed agency from a sociocognitive position as a willingness to engage with the curriculum. by contrast, studies that use a sociocultural framework focus on how individuals use the resources they have available in their sociocultural context (eteläpelto, 2017). jääskelä, poikkeus, et al. (2020) defined a subject-centred sociocultural understanding of student agency as ‘a student’s experience of having access to or being empowered to act through personal, relational, and participatory resources, which allow him/her to engage in purposeful, intentional, and meaningful action and learning in study contexts’ (p. 2). sociocultural frameworks have been used to explore course-related experiences of agency (jääskelä et al., 2016) and to map students’ agency profiles and their connection to students’ perceptions of teaching practices (jääskelä, poikkeus, et al., 2020). some scholars study agency from a life-course perspective, directed towards students’ professional futures (soini et al., 2015; toom et al., 2017), and pay attention to students’ past, present, and future. this research refers to emirbayer & mische’s (1998) understanding of agency as ‘the temporally constructed engagement by actors of different structural environments—the temporal relational contexts of action—which, through the interplay of habit, imagination, and judgment, both reproduces and transforms those structures in interactive response to the problems posed by changing historical situations’ (p. 970). the interplay in a specific context is teased out by exploring how students enter into relationships with surrounding people, places, meanings, and events; and students’ actual interactions with their context. the temporal perspective, in particular, has been adopted in recent research into the development of professional agency (soini et al., 2015; toom et al., 2017). in addition, it underpins merrill’s (2014) study in which agency is understood within contexts and time; individuals engage in changing their future. harris et al. (2018) discussed student agency’s temporal nature in the context of assessment, where student resistance to assessment or assessment decisions might draw on past experiences. stenalt 56 | f l r 3.1. limitations in current approaches to student agency even if existing studies address significant aspects of student agency, some argue that research remains limited by typically investigating student agency through only some of the aspects related to agency (jääskelä, heilala, et al., 2020; jääskelä et al., 2016). instead, jääskelä et al. (2016) argued that a holistic operationalisation of agency includes personal, relational, and participatory domains. the personal domain of agency connects to individuals’ resources and disposition towards being agentic. in particular, it involves self-efficacy and competence beliefs. the relational domain of agency refers to relational recourses for learning such as the learning climate, peer support, and power relations. participatory resources involve the contextual dimension of agency, comprising the extent to which students perceive themselves to have the opportunity to participate, influence, and make choices in learning and within the learning context. while the holistic operationalisation of agency might be thought of as supporting the complexity of the construct, the framework remains limited in explaining how, for example, opportunities for active participation emerge and how peers become resources for learning. more specifically, in terms of the purpose of this article, the framework is less helpful in understanding how the digital becomes a resource for students, how the digital affects students and vice versa. further, others argue that there is a need to be more sensitive towards ‘the intentional projects of individual and collective agents, and how these projects are enabled and constrained’ (ashwin, 2012, p. 21). as klemenčič (2015) has described, ‘student agency is conceptualised as a process of student actions and interactions during studentship, which encompasses variable notions of agentic orientation (“will”), the way students relate to past, present and future in making choices of action and interaction, and of agentic possibility (“power”), that is their perceived power to achieve intended outcomes in a particular context of action and interaction’ (p. 16). against this concern, student agency research needs to analytically differentiate agentic possibilities from agentic orientation and account for how agentic possibilities emerge to students and how student orientation influences student actions. this paper contends that this issue requires critical thought – particularly in light of the optimistic accounts of the ways digital technologies benefit student learning (garrison & kanuka, 2004) and the tendency to frame digital learning activities from an educational perspective. 4. central dynamics of digital engagement there is clearly a need to cultivate a better sense of agency within digital contexts and the ways the distinct digital features influence students’ agency. this paper turns to media research to theorise how the digital configures the environment in a way that has the potential to shape students’ engagement. the research of interest examines social network sites, and while networking socially or for professional purposes may not dominate digital technologies use in formal educational contexts, they serve many of the same functions. they allow people to connect and share content with other people than close friends and family, and they help people gather for a purpose. 4.1 networked publics within social network sites, people are expected to act as networked individuals, maintaining various networks of people and resources that can be navigated as required to meet specific needs (boyd, 2010; wellman et al., 2003; wellman & rainie, 2013). as baym and boyd (2012) explained, people who use social media ‘juggle multiple layers and kinds of audiences, bringing into being multiple and diverse kinds of publics, counterpublics, and other emergent social arrangements’ (pp. 321–322). boyd (2010) makes the important observation that it is useful to think of such sites and the practices unfolding here as networked publics. as stated: stenalt 57 | f l r ‘networked publics are publics that are restructured by networked technologies. as such they are simultaneously (1) the space constructed through networked technologies and (2) the imagined collective that emerges as a result of the intersection of people, technology, and practice’ (boyd, 2010, p. 39). according to livingstone (2005), networked publics are constructed by a space and collection of people or bounded by a shared performance or object. while the terminology emphasises networks, it is not the primary purpose of many social sites (boyd & ellison, 2007), and people might not use digital technologies to connect with others (gourlay, rodríguez-illera, barberà, et al., 2021). 4.2 sharing of profiles: locus of interaction and self-representation most social network sites allow participants to generate personal profiles, and boyd (2010) argues that profiles act as the locus of interaction and represent the individual. following this line of thinking, mccosker (2017) has described social sites as performative spaces where people actively construct their identities and profiles for an audience. because individuals’ profiles are objects of selfrepresentation attention, comparison, negotiation, and remix (papacharissi, 2011), constantly editing or remixing oneself is an essential online practice (papacharissi, 2012). in addition to profiles being a site of self-management, profiles are also a site of external control. indeed, self-representation depends on the content and activity selected by the specific media (mccosker, 2017). as illustrated by mccosker (2017), one’s name, username, profile image, ‘about’ information, and relationships (follows, followers) appear to be a standardised basis for personal identifiers. in contrast, data on education, birthday, age, and other websites are less common features in profiles. with external structures restricting selfrepresentation, users often engage in actions of resistance or playfulness to overcome reductive versions of themselves (mccokser, 2017). hence, research has found that performing identity and sociality online involves a constant balancing of social benefits with privacy costs (papacharissi, 2012). thus, while networking sites enable individuals to enact social roles, play and resistance strategies are central to staging a digital profile that conceals the aspects of oneself that the individual would like not to be shared. as digital profiles and data become contested, it is also becoming critical to consider personal privacy. we can think of privacy or expressive privacy as the protection of acts of speech or activity that express self-identity or personhood (ess, 2015). according to ess (2015), a space of expressive privacy is required if individuals are to reflect and critique alternatives. mechanisms implemented in social media allow users to control access to the data they generate to some extent. profiles and usergenerated data may be open source, available for many to access and potentially use, or closed, accessible only to the student or the teacher. yet, privacy settings are enacted in several ways in digital contexts, beyond settings of open/closed or individual/all, such as the ‘following mechanism’ of reciprocal or nonreciprocal kind: either mutual acceptance of each other must be in place to read each other’s contributions, or one-way following is accepted. 4.3 data-based content: persistent, replicable, scalable, and searchable according to boyd (2010), networked individuals are also challenged to manage the persistence, replicability, scalability, and searchability of their profiles and data in digital environments. the first property, persistence, suggests that individuals’ contributions, such as text or expressions, are easily captured and stored. in fact, many systems operate on persistence by default, making previous unmediated moments of communication persistent. second, boyd (2010) argued that technology has increased the replicability of content. because of this and the ease with which one can modify original content, what is original and replicated is hard to discern. third, boyd (2010) mentioned scalability as an affordance of networking sites. scalability is the potential to enhance the distribution of content or who has access to it. however, what is amplified through broad distribution is not always what the content owner would have chosen. lastly, boyd emphasises searchability as an affordance, based on stenalt 58 | f l r the premise that technology use leaves digital traces that can be located to a person or an object. these characteristics imply that profiles and individuals’ activities are difficult to erase and easy to share with known and unknown others. participants’ digital contributions to interactions, then, persist and can easily be retrieved, duplicated, and redistributed across various contexts. when considering the ways the audiences of personal data emerge to humans, this becomes increasingly important. as boyd (2010) stated, ‘in unmediated spaces, it is common to have a sense for who is present and can witness a particular performance. the affordances of networked publics change this’ (p. 49). not knowing one’s audience makes it difficult to make a contribution that considers others’ reactions. 4.4 affect: blurring of relational investment and tools for engagement a crucial underpinning of social media is affective investments on the users’ part. as people engage with digital technologies, they are frequently offered choices for expressing emotions and social bonds. indeed, what makes social networking sites unique is that they allow individuals to connect with others and articulate and enact sociality through affect and affective technical features (boyd & ellison, 2007). technical features such as comments, views, or likes are referenced as social buttons (gerlitz & helmond, 2013). social buttons comprise affective statements such as ‘great’ or affective states such as ‘feeling amused’ as reactions to others’ comments, which are shared with a particular group of users. hence, they allow users not only to share or recommend content but also to share social connections and affect. affective features such as those mentioned transform user affect and spontaneous responses into comparable acts of engagement and signs of social connections that structure digital performance and are critical to social networking sites (boyd, 2010; gerlitz & helmond, 2013). this includes how personal data might act as objects of memories or mementoes on social sites (lupton, 2020). as noted by lupton (2020), data which are shared are archived, and can be revived to connect people with their past connections and activities. yet, while affective interactional acts are critical in terms of maintaining or increasing users’ digital engagement, the so-called like economy ‘is facilitating a web of positive sentiment in which users are constantly prompted to like, enjoy, recommend and buy as opposed to discuss or critique’ (gerlitz & helmond, 2013, p. 1362). thus, a characteristic of networked sites is how networked communication entails a ‘panoply of affective attachments: articulations of desire, seduction, trust and memory; sharp jolts of anger and interest; political passions; investments of time, labor, and financial capital; and the frictions and pleasures of archival practices’ (paasonen et al., 2015, p. 1). as paasonen (2018) reasoned, ‘if time, attention and data are the price that people pay, or that which they hand over in order to access social media, then affective ripples are part of that which is gained in return’ (p. 9-10). 4.5 time: when does something happen online? networked interactions or exchanges can also be seen as underpinned by temporalities that structure activities and influence peoples’ abilities to access and share data. while measures of time such as clocks and calendars dominate, in many instances, time in digital contexts is difficult to pinpoint. ‘network time’ has been proposed by hassan (2007) as a way of opening up the idea of many temporal possibilities. hassan (2007) saw the internet as an inherently asynchronous space where nothing occurs simultaneously. instead, the internet offers multiple spectra of temporalities. operating systems might respond to input at high speed and immediately, but there will always be some temporal lag. the internet connection might be slow or fail, and we can spend seconds, minutes, or hours waiting for content to appear. moreover, the digital space includes different time zones, which challenges the idea of time consistency. network time, then, is referred to as ‘a digitally compressed clock-time’ (hassan 2003, p. 233), which is time that has exploded into a million different time fractions. stenalt 59 | f l r due to this tension, it makes sense to interrogate time less regarding its logical meaning and more in its existential meaning (bennett & burke, 2018; lash, 2001). here, hassan (2007) used adam’s (2008) notion of timescapes to illustrate the range of experiences of time possible. the concept of timescapes captures ‘that we cannot embrace time without simultaneously encompassing space and matter, that is, without embodiment in a specific and unique context’ (adam, 2008, p. 1). following adam’s work, time is constitutive of seven elements: (a) time frame, referring to a bounded unit with a beginning and an end; (b) temporality, referring to the unfolding of time and the direction it takes; (c) timing, referring to when or something happening at a specific time; (d) tempo, referring to the speed at which something happens; (e) duration, referring to considerations of how long something takes; (f) sequence, referring to the order of things such as actions; and (g) temporal modalities, referring to when something happens (in the past, present, or future). crucially, these forms of time are linked to digital behaviour and sense-making. 4.7 implications while student agency research has focused on the individual aspects of human actions and subsumed the digital context to educational intentions, media studies offer significant insight into digital behaviour as socially, culturally, and technologically dependent. this is not to say that technologies used in education have the same precise features as social network technologies or that the purpose of technology use is the same. instead, this paper suggests taking the elements into account as central dynamics forming part of engagement in the digital world to shed light on why students engage the way they do in digital interactions. thus, the value of constructing digital technologies in education as social network sites is analytical. it directs our attention to practices as being informed by the dynamics of networked publics. as such, it involves paying attention to the interweaving of technology and humanity and how these are interconnected with other practices and relations (markham, 2018), rather than focusing on the digital as a tool (focusing on cultural practices in or of digital contexts; markham, 2018) or a medium (viewing the digital context as a cultural space in which one can be present and feel absorbed; markham, 2018). against this background, approaches to digital student agency need to include critical understandings of the way agency is constructed and constrained. this involves the observable settings and features and how the interplay between social, cultural and technological aspects emerge to students. 5. digital student agency a critical framework the paper continues to sketch out a digital student agency framework, which considers the complex construct of student agency in relation to the distinct digital features. the proposed framework involves five domains: agentic possibility, digital self-representation, data uses, digital sociality, and digital temporalities (see table 1). it is important to stress that each domain is critical in orientation, developed to identify and understand the ways student agency is constrained. this coincides with selwyn’s (2010) argument that we need nuanced and thick descriptions of technology use. given the framework’s critical nature, it is not intended to present a normative ideal for how a digital context for learning should be to facilitate student agency. for example, it does not state that online peer feedback should be conducted in a certain way, using a specific technology with specific features. being critical means exploring and understanding the implications that settings and the cultural and social context might have on student agency. some practical examples of how each domain might be approached are described as actions. stenalt 60 | f l r table 1 ‘digital student agency’ framework (disa) domain key questions actions agentic possibility what power do students have to achieve the intended outcomes in the particular context of action and interaction? identifying sources and resources of agency in the interaction analysing the ways sources emerge to students during the interaction determining how the object of engagement links to student trajectories identifying student possibility to influence the object of engagement and the interaction required evaluate the level of access that students have to influence the object of engagement digital selfrepresentation how can students manage and adapt their selfrepresentation in the digital context of learning? identifying how and where options for managing students’ profiles are constructed and processed analysing how students’ profiles are visible to others (peers, teachers, managers) exploring the implications data uses how are student data circulated or recirculated? identifying how and where student contributions are generated and circulated identifying how students can manage their data identifying who has access to student data and when identifying how students’ data can be used, including purposes extending the original intent evaluating the implications of the uses digital sociality how is sociality constructed, and how can students manage sociality? identifying the means that are available for communication and cultivating a sense of sociality identifying how and where sociality is constructed determining how students can manage whom they interact with, how they interact, and the purpose of the interaction analysing the role of socialisation in terms of how it affects the interaction or the outcome of the interaction digital temporalities how are student actions constructed in terms of time? identifying the digital temporalities and analysing the underpinning of time: is it structured by individual students, groups of students, teaching staff, or a digital system? stenalt 61 | f l r analysing how time materialises to students digitally explore the implications of time from a student perspective the first domain mentioned in the framework – agentic possibility – is critical to identifying the possibilities for cultivating agency in the specific context of action. it pays attention to students’ effective opportunities and freedom to do what they reason to make sense within education. it also includes assessing the sources of agency available in the particular context. sources refer here to the broad array of internal dispositions, relational, and contextual sources or resources known from mainstream agency frameworks. while the sources cultivate student agency in higher education contexts (jääskelä, poikkeus, et al., 2020; toom et al., 2017), research has also made clear that technology and educational instructions can hinder access to sources of agency (stenalt, 2021). moreover, ashwin and mcvitty (2015) have emphasised that educational planning of student engagement includes stratified and directive access to knowledge and influence. thus, the first domain is the impetus for discussions of how the educational purpose and guidelines configure student agency. the second domain – digital self-representation – involves identifying the particular ways students are placed in digital interactions and the type of self-representation possible in the specific exchange. while digital system profiles are relatively easy to locate, students might perceive other data types to represent them. for example, students might be identified through the following: • contributions with their signature • an alias • the order of appearance • pictures • oral expressions of authorship in-class supplementing the digital contribution • familiarity with students’ style of expression once the profiles have been identified, we can begin to think about how students can control their self-representation and explore how it might influence student behaviour. while social media often enable users to adjust to their privacy settings (waterloo et al., 2018), these choices are likely to be distributed to platform managers or teachers within educational contexts. students, then, are managed as a collective group and assumed to have the same relationship with each of their peers. following bayne et al. (2019), students might turn to acts of resistance if they cannot control their selfrepresentation. similarly, research has found limited options for self-representation to decrease student engagement (stenalt, 2021). the third domain – data uses – examines the life of student data. for example, does data stick around in the sense that others can access the information? does it have a value for students in extending the particular interaction and interaction outcomes? because learning activities aligned with the formal assessment are of high importance to students (biggs & tang, 2011), critical academic data, which disappears, might lead to student frustration. automated personalisation of educational offerings using students’ learning data as the stepping stone can be seen by students to reduce them to numbers and restrict their access to resources (tsai et al., 2020). it is crucially important to address the conflicting beliefs and opportunities available for social learning in digital contexts. understanding the type of sociality that an interaction fosters or requires is the fourth domain’s concern. the domain also supports explorations of expected affective investment in relation to successful participation and students’ actual affective investment – allowing insight into the way sociality is constituted. for example, successful collaboration is seen to require coregulation (volet et al., 2009), which involves ‘individuals’ various attempts to affect each other’s motivation, stenalt 62 | f l r emotional state, cognitive actions, etc. for their own purpose or others’ benefits, or alternatively to coordinate their actions for a shared purpose’ (järvenoja et al., 2013, p. 35). while social and communicative activities are essential for maintaining a positive group climate (janssen et al., 2012), such opportunities and features might not be sufficiently devised digitally. at the same time, research has found that students prefer technologies that ease the logistics of their life, rather than technologies that support collaborative work (henderson et al., 2017). this challenges ideals of digital collaboration. relatedly, educational framings of interactions may stress social learning but provide limited opportunities for students to affect the object of engagement. here, we can draw on ashwin and mcvitty (2015) model to explore the potential impact. in light of this, decoding the imagined collective, the actual collective, and the role of the collective should be an essential element of understanding digital educational spaces. the last domain – digital temporalities – enables consideration of the temporalities involved in the context and how they affect students. following from hassan and others, students’ relation to others and content in digital contexts is not a universal thing but instead processes that develop from being engaged and shaped by structures that include the educational instructions and technology used. as a simple example, a blog post is, by default, typically invisible for others during the writing of the post. the invisibility ceases after data has been published, where it is structured in relation to the logic of a calendar. it remains visible to students until the month or week changes or until a sufficient number of other data entries are made. while the data visually disappears as time moves on, it can be retrieved later on through search mechanisms. in contrast, entries by students through kahoot, which can be used to facilitate quizzes with students in the same room, fall into sequences, are visible for all students for a limited time (the duration of a sequence), and cannot be retrieved later on by the individual student. taking a broader perspective, technologies and educational choices produce different relationships to time and content. time settings also risk labelling certain groups of students rather than helping them achieve their potential. students, for example, who struggle to manage their time and meet deadlines in an online course with various activities might see themselves as lacking self-discipline and less capable of studying (bennett & burke, 2018). for students with small children at home, an online course might be more challenging to attend than a course at the campus, reducing the proposed benefit of providing student flexibility (kirkwood & price, 2014). additionally, the temporality of digital interactions can affect students’ approaches to learning. lash (2001), for example, has described how technological forms of life that are sped up can result in content being devaluated within hours or days (lash, 2001). what lash (2001) points to is how experiences of limited time or moving fast forward can lead to a higher degree of student insecurity because their attention may become directed to the consequences of the present for the future rather than looking to the past to explain the present. reeve and jang (2006) found that limited time to engage and a teacher monopolising time correlated negatively with student experience of autonomy. this, at least, invites us to consider how the speed of digital interactions and the distribution of time relate to student self-regulation and self-efficacy. 6. future approaches to student agency recent events such as covid-19 and the turn towards distance education have increased the awareness of digital technologies for education (williamson et al., 2020). due to the challenges of making online teaching feel meaningful and relevant to students, approaches that help develop digital student agency are needed. in light of this, the paper offers three recommendations. first, the paper suggests that the digital should not be assumed to support agency by default. rather than subsuming agency to technical affordances or pedagogical guidelines, the digital should be confronted as something constituted by several interrelated digital domains nested in social relations. stenalt 63 | f l r key to this, the framework offers a disposition that encourages critical investigations of students’ agency possibilities in digital contexts, how they may emerge to students, and affect their actions. following this, the framework moves beyond simply stating online tools and educational intentions when exploring digital contexts of education to in-depth investigations of the dynamics configuring digital environments. in specific, the framework enables explorations of the invitational quality of the digital (adams & thompson, 2016) and decentring the object of inquiry (pink et al., 2017), allowing researchers to give an account of the mode of being constructed through digital technologies and the way agency is constrained. taking the domains into account, it becomes clear how digital technologies can mediate our being in the world and direct our attention in a certain way, affecting how we come to know the content (rosenberger, 2017a, 2017b). thus, technology use fundamentally involves substantive interventions for the context into which it is embedded. in this light, the digital is neither a neutral mediator of information nor independent of the system into which it is adopted. the use of the domains allows us to raise a range of questions to explore the interplay between human-technology in educational contexts, including the following: • how do educational practices constrain student agency? • what forms of settings do students see as appropriate in cultivating their agency and making connections with others to learn with digital technology? • how do students make sense of data produced and shared online? • what forms experiences of time, and how do temporalities affect students’ engagement? second, the logic of digital engagement suggests that agency in digital contexts needs to be understood less as an individual phenomenon and more as a relational phenomenon to balance digital ways of being. the relational approach to agency points to the notion of ‘relational agency’ (burkitt, 2016), whereby agency describes ‘people producing particular effects in the world and on each other through their relational connections and joint actions, whether or not those effects are reflexively produced’ (burkitt, 2016, p. 323). the key point of relational agency is not to understand the digital from a single student perspective but to understand that sense-making of the digital is associated with the different forms of relationships being confronted and constituted in the digital context. as burkitt (2016) stated, ‘it is not simply relations, and the objects and entities produced in relations, that is the focus of study; more specifically, it is social relations and the mode of life humans produce through them, including material culture and technology, that relational sociologists need to bring into the analysis’ (p. 331). in light of this, the framework can build an understanding of the forms of power affecting student agency and how to organise learning that considers the challenges to student agency in digital contexts. third, the complexity of agency in digital contexts highlights the need to consider how technology-supported student-centred environments emerge to students. while jääskelä, poikkeus, et al. (2020) have identified a relationship between high levels of student agency and perceptions of courses as student-centred, more knowledge is needed of the features constituting a student-centred environment to develop models that can guide the design of such environments. for instance, to what degree does student-centredness depend on high levels of privacy and how can privacy be materialised in designs? here, the framework can be used to help identify the underlying features of a digital studentcentred environment from a student perspective. 7. conclusion because digital education connects to student agency, it is important that we understand the various qualities and capabilities of digital contexts of learning and how they affect student agency. this article has outlined how approaches to understanding the digital in existing student agency research are stenalt 64 | f l r too narrow and therefore miss out on important insight derived from the field of media research. instead, the paper proposes a digital student agency framework for developing a better understanding of the human-technology interplay that pays attention to the ways relational, cultural, and technological dynamics constitute agency. the proposed framework suggests understanding digital student agency through five domains formed by media research and research into student agency in higher education. the framework is an initial attempt at addressing the interplay between student agency and digital contexts of learning. therefore, it invites testing and critique of the framework and its domains. developing knowledge of digital ways of being in educational contexts is complex. in doing so, it makes sense to look at student agency as a basis for exploring the student perspective and working out realistic accounts of digital encounters and ways to support students in digital contexts. thus, the task proposed by this paper is not to better digital education and practices by promoting certain ideals of higher education teaching and learning but by considering the ways students’ trajectories and social connections might challenge students in mundane digital contexts. as digital ways of engaging and managing students become ubiquitous, the distinction between learning as private and learning as public will become blurry. the considerations presented by the framework will not only be constrained to specific digital teaching-learning interactions but will be part of students’ everyday life. keypoints the article: provides a critical approach to student agency in digital contexts of higher education expands current student agency research by moving beyond an agency-context dichotomy towards understandings of agency-context as interrelated offers a ‘digital student agency’ framework that distinguishes five significant domains: (1) agentic possibility, (2) digital self-representation, (3) data uses, (4) digital sociality, and (5) digital temporalities references adam, b. (2008). of timescapes, futurescapes and timeprints. in l. university (ed.), lüneburg talk web 070708. adams, c., & thompson, t. l. (2016). attending to objects, attuning to things. in t. l. t. cathrine adams (ed.), researching a posthuman world (pp. 23-56). palgrave macmillan. https://doi.org/10.1057/9781-137-57162-5_2. arkoudis, s., & tran, l. t. (2007). international students in australia: read ten thousand volumes of books and walk ten thousand miles. asia pacific journal of education, 27(2), 157-169. https://doi.org/10.1080/02188790701378792. ashwin, p. (2012). analysing teaching-learning interactions in higher education. accounting for structure and agency. london, new york, continuum. ashwin, p., & mcvitty, d. (2015). the meanings of student engagement: implications for policies and practices. in a. m. curaj, liviu; pricopie, remus; salmi, jamil; scott, peter (ed.), the european higher education area between critical reflections and future policies (pp. 343-359). springer open. https://doi.org/10.1007/978-3-319-20877-0_23. bandura, a. (2002). growing primacy of human agency in adaptation and change in the electronic era. european psychologist, 7(1), 2. https://doi.org/10.1027/1016-9040.7.1.2. bandura, a. (2006). toward a psychology of human agency. perspectives on psychological science, 1(2), 164-180. https://doi.org/10.1111/j.1745-6916.2006.00011.x. stenalt 65 | f l r baym, n. k., & boyd, d. (2012). socially mediated publicness: an introduction. journal of broadcasting & electronic media, 56(3), 320-329. https://doi.org/10.1080/08838151.2012.705200. bayne, s. (2015). what's the matter with ‘technology-enhanced learning’? learning, media and technology, 40(1), 5-20. https://doi.org/10.1080/17439884.2014.915851. bayne, s., connelly, l., grover, c., osborne, n., tobin, r., beswick, e., & rouhani, l. (2019). the social value of anonymity on campus: a study of the decline of yik yak. learning, media and technology, 44(2), 92-107. https://doi.org/10.1080/17439884.2019.1583672. bennett, a., & burke, p. j. (2018). re/conceptualising time and temporality: an exploration of time in higher education. discourse: studies in the cultural politics of education, 39(6), 913-925. https://doi.org/10.1080/01596306.2017.1312285. biggs, j., & tang, c. (2011). teaching for quality learning at university (4 ed.). open university press. boyd, d. (2010). social network sites as networked publics: affordances, dynamics, and implications. in z. papacharissi (ed.), a networked self (pp. 47-66). routledge. boyd, d. m., & ellison, n. b. (2007). social network sites: definition, history, and scholarship. journal of computer‐mediated communication, 13(1), 210-230. https://doi.org/10.1111/j.10836101.2007.00393.x. burkitt, i. (2016). relational agency: relational sociology, agency and interaction. european journal of social theory, 19(3), 322-339. https://doi.org/10.1177/1368431015591426. calitz, t. m. l., walker, m., & wilson-strydom, m. (2016). theorising a capability approach to equal participation for undergraduate students at a south african university. perspectives in education, 34(2), 57-69. https://doi.org/10.18820/2519593x/pie.v34i2.5. damşa, c. i., kirschner, p. a., andriessen, j. e., erkens, g., & sins, p. h. (2010). shared epistemic agency: an empirical study of an emergent construct. the journal of the learning sciences, 19(2), 143-186. https://doi.org/10.1080/10508401003708381. emirbayer, m., & mische, a. (1998). what is agency? american journal of sociology, 103(4), 962-1023. https://doi.org/10.1086/231294. ess, c. (2015). new selves, new research ethics. in h. fossheim & h. ingierd (eds.), internet research ethics (pp. 48-76). cappelen damm akademisk. eteläpelto, a. (2017). emerging conceptualisations on professional agency and learning. in m. p. goller, susanna (ed.), agency at work an agentic perspective on professional learning and development (1 ed., pp. 183-201). springer. doi:10.1007/978-3-319-60943-0_10. fenwick, t., & landri, p. (2012). materialities, textures and pedagogies: socio-material assemblages in education. pedagogy, culture & society, 20(1), 1-7. https://doi.org/10.1080/14681366.2012.649421. garrison, d. r., & kanuka, h. (2004). blended learning: uncovering its transformative potential in higher education. the internet and higher education, 7(2), 95-105. https://doi.org/10.1016/j.iheduc.2004.02.001. gerlitz, c., & helmond, a. (2013). the like economy: social buttons and the data-intensive web. new media & society, 15(8), 1348-1365. https://doi.org/10.1177/1461444812472322. gourlay, l., rodríguez-illera, j. l., barberà, e., bali, m., gachago, d., pallitt, n., jones, c., bayne, s., hansen, s. b., hrastinski, s., jaldemark, j., themelis, c., pischetola, m., dirckinck-holmfeld, l., matthews, a., gulson, k. n., lee, k., bligh, b., thibaut, p., vermeulen, m., nijland, f., vrielingteunter, e., scott, h., thestrup, k., gislev, t., koole, m., cutajar, m., tickner, s., rothmüller, n., bozkurt, a., fawns, t., ross, j., schnaider, k., carvalho, l., green, j. k., hadžijusufović, m., hayes, s., czerniewicz, l., knox, j., & networked learning editorial, c. (2021). networked learning in 2021: a community definition. postdigital science and education. https://doi.org/10.1007/s42438-02100222-y hamilton, e., & friesen, n. (2013). online education: a science and technology studies perspective/éducation en ligne: perspective des études en science et technologie. canadian journal of learning and technology/la revue canadienne de l’apprentissage et de la technologie, 39(2). harré, r., & van langenhove, l. (1999). positioning theory: moral contexts of intentional action. blackwell oxford. stenalt 66 | f l r harris, l. r., brown, g. t., & dargusch, j. (2018). not playing the game: student assessment resistance as a form of agency. the australian educational researcher, 45(1), 125-140. https://doi.org/10.1007/s13384-018-0264-0. hassan, r (2003). network time and the new knowledge epoch. time & society, 12(2-3), 226-241. https://doi.org/10.1177/0961463x030122004. hassan, r. (2007). 24/7: time and temporality in the network society, stanford university press. henderson, m., selwyn, n., & aston, r. (2017). what works and why? student perceptions of ‘useful’digital technology in university teaching and learning. studies in higher education, 42(8), 1567-1579. https://doi.org/10.1080/03075079.2015.1007946. hitlin, s., & long, c. (2009). agency as a sociological variable: a preliminary model of individuals, situations, and the life course. sociology compass, 3(1), 137-160. https://doi.org/10.1111/j.17519020.2008.00189.x. irvine, v., code, j., & richards, l. (2013). realigning higher education for the 21st century learner through multi-access learning. journal of online learning and teaching, 9(2), 172. janssen, j., erkens, g., kirschner, p. a., & kanselaar, g. (2012). task-related and social regulation during online collaborative learning. metacognition and learning, 7(1), 25-43. https://doi.org/10.1007/s11409010-9061-5 järvenoja, h., volet, s., & järvelä, s. (2013). regulation of emotions in socially challenging learning situations: an instrument to measure the adaptive and social nature of the regulation process. educational psychology, 33(1), 31-58. https://doi.org/10.1080/01443410.2012.742334. jääskelä, p., heilala, v., kärkkäinen, t., & häkkinen, p. (2020). student agency analytics: learning analytics as a tool for analysing student agency in higher education. behaviour & information technology, 1-19. https://doi.org/10.1080/0144929x.2020.1725130. jääskelä, p., poikkeus, a.-m., häkkinen, p., vasalampi, k., rasku-puttonen, h., & tolvanen, a. (2020). students’ agency profiles in relation to student-perceived teaching practices in university courses. international journal of educational research, 103. https://doi.org/10.1016/j.ijer.2020.101604. jääskelä, p., poikkeus, a. m., vasalampi, k., valleala, u. m., & rasku-puttonen, h. (2016). assessing agency of university students: validation of the aus scale. studies in higher education, 1-19. https://doi.org/10.1080/03075079.2015.1130693 kirkwood, a., & price, l. (2011). enhancing learning and teaching through technology: a guide to evidencebased practice for academic developers. h. e. academy. http://oro.open.ac.uk/32489/ kirkwood, a., & price, l. (2014). technology-enhanced learning and teaching in higher education: what is ‘enhanced’and how do we know? a critical literature review. learning, media and technology, 39(1), 6-36. https://doi.org/10.1080/17439884.2013.770404. klemenčič, m. (2015). what is student agency? an ontological exploration in the context of research on student engagement. in m. klemenčič, s. bergan, & r. primožič (eds.), student engagement in europe: society, higher education and student governance. (pp. 11-29). council of europe higher education series no. 20. strasbourg: council of europe publishing. klemenčič, m. (2017). from student engagement to student agency: conceptual considerations of european policies on student-centered learning in higher education. higher education policy, 30(1), 69-85. https://doi.org/10.1057/s41307-016-0034-4. klemenčič, m., pupinis, m., & kirdulytė, g. (2020). mapping and analysis of student-centred learning and teaching practices: usable knowledge to support more inclusive, high-quality higher education (neset analytical report). publications office of the european union. http://dx.doi.org/10.2766/67668. lash, s. (2001). technological forms of life. theory, culture & society, 18(1), 105-120. https://doi.org/10.1177/02632760122051661. leonardi, p. m. (2010). digital materiality? how artifacts without matter, matter. first monday, 15(6). https://doi.org/10.5210/fm.v15i6.3036. lindgren, r., & mcdaniel, r. (2012). transforming online learning through narrative and student agency. educational technology & society, 15(4), 344-355. livingstone, s. (2005). on the relation between audiences and publics. in s. livingstone (ed.), audiences and publics: when cultural engagement matters for the public sphere. (2 ed., pp. 17-41). intellect books. stenalt 67 | f l r luo, h., yang, t., xue, j., & zuo, m. (2019). impact of student agency on learning performance and learning experience in a flipped classroom. british journal of educational technology, 50(2), 819-831. https://doi.org/10.1111/bjet.12604. lupton, d. (2020). data selves: more-than-human perspectives. polity press. malmberg, l. e., & hagger, h. (2009). changes in student teachers' agency beliefs during a teacher education year, and relationships with observed classroom quality, and day-to-day experiences. british journal of educational psychology, 79(4), 677-694. https://doi.org/10.1348/000709909x454814. marín, v. i., de benito, b., & darder, a. (2020). technology-enhanced learning for student agency in higher education: a systematic literature review. interaction design and architecture(s) journal ixd&a, 45, 15-49. markham, a. n. (2018). ethnography in the digital internet era. in n.k. denzin & y.s. lincoln (eds.), sage handbook of qualitative research (5 ed., pp. 650-668). sage. mccosker, a. (2017). data literacies for the postdemographic social media self. first monday, 22(10). https://doi.org/10.5210/fm.v22i10.7307. merrill, b. (2014). determined to stay or determined to leave? a tale of learner identities, biographies and adult students in higher education. studies in higher education, 40(10), 1859-1871. https://doi.org/10.1080/03075079.2014.914918 nieminen, j. h., tai, j., boud, d., & henderson, m. (2021). student agency in feedback: beyond the individual. assessment & evaluation in higher education, 1-14. https://doi.org/10.1080/02602938.2021.1887080. nye, a., hughes-warrington, m., roe, j., russell, p., deacon, d., & kiem, p. (2011). exploring historical thinking and agency with undergraduate history students. studies in higher education, 36(7), 763-780. https://doi.org/10.1080/03075071003759045. oecd. (2018). the future of education and skills: education 2030. oecd education working papers. oecd paris, france. papacharissi, z. (2011). conclusion. a networed self. in z. papacharissi (ed.), a networked self (pp. 304318). routledge. papacharissi, z. (2012). without you, i'm nothing: performances of the self on twitter. international journal of communication, 6, 18. passey, d., shonfeld, m., appleby, l., judge, m., saito, t., & smits, a. (2018). digital agency: empowering equity in and through education. technology, knowledge and learning, 23(3), 425-439. https://doi.org/10.1007/s10758-018-9384-x. paasonen, s. (2018). affect, data, manipulation and price in social media. distinktion: journal of social theory, 19(2), 214-229. https://doi.org/10.1080/1600910x.2018.1475289. paasonen, s., hillis, k., & petit, m. (2015). networks of transmission: intensity, sensation, value. in k. hillis, s. paasonen, & m. petit (eds.). networked affect. (pp. 1-24) cambridge: mit press. pink, s., sumartojo, s., lupton, d., & heyes labond, c. (2017). empathic technologies: digital materiality and video etnography. visual studies, 32(4), 371-381. https://doi.org/10.1080/1472586x.2017.1396192. reeve, j., & jang, h. (2006). what teachers say and do to support students' autonomy during a learning activity. journal of educational psychology, 98(1), 209-218. https://doi.org/10.1037/00220663.98.1.209. rosenberger, r. (2017a). the ict educator’s fallacy. foundations of science, 22(2), 395-399. https://doi.org/10.1007/s10699-015-9457-4. rosenberger, r. (2017b). notes on a nonfoundational phenomenology of technology. foundations of science, 22(3), 471-494. https://doi.org/10.1007/s10699-015-9480-5. selwyn, n. (2010). looking beyond learning: notes towards the critical study of educational technology. journal of computer assisted learning, 26(1), 65-73. https://doi.org/10.1111/j.13652729.2009.00338.x. sen, a. (2005). human rights and capabilities. journal of human development, 6(2), 151-166. https://doi.org/10.1080/14649880500120491 stenalt 68 | f l r shonfeld, m., passey, d., appleby, l., judge, m., saito, t., smits, a., khablan, s., & starkey, l. (2017). digital agency to empower equity in education: summary report. in: rethinking learning in a digital age. edusummit 2017. (pp. 39-45). soini, t., pietarinen, j., toom, a., & pyhältö, k. (2015). what contributes to first-year student teachers’ sense of professional agency in the classroom? teachers and teaching, 21(6), 641-659. https://doi.org/10.1080/13540602.2015.1044326. starkey, l. (2019). three dimensions of student-centred education: a framework for policy and practice. critical studies in education, 60(3), 375-390. https://doi.org/10.1080/17508487.2017.1281829. stenalt, m. h. (2021). researching student agency in digital education as if the social aspects matter: students’ experience of participatory dimensions of online peer assessment. assessment & evaluation in higher education, 46(4), 644-658. https://doi.org/10.1080/02602938.2020.1798355. toom, a., pietarinen, j., soini, t., & pyhältö, k. (2017). how does the learning environment in teacher education cultivate first year student teachers' sense of professional agency in the professional community? teaching and teacher education, 63, 126-136. https://doi.org/10.1016/j.tate.2016.12.013. tsai, y.-s., perrotta, c., & gašević, d. (2020). empowering learners with personalised learning approaches? agency, equity and transparency in the context of learning analytics. assessment & evaluation in higher education, 1-14. https://doi.org/10.1080/02602938.2019.1676396. volet, s., vauras, m., & salonen, p. (2009). self-and social regulation in learning contexts: an integrative perspective. educational psychologist, 44(4), 215-226. https://doi.org/10.1080/00461520903213584. waterloo, s. f., baumgartner, s. e., peter, j., & valkenburg, p. m. (2018). norms of online expressions of emotion: comparing facebook, twitter, instagram, and whatsapp. new media & society, 20(5), 18131831. https://doi.org/10.1177/1461444817707349. wellman, b., quan-haase, a., boase, j., chen, w., hampton, k., díaz, i., & miyata, k. (2003). the social affordances of the internet for networked individualism. journal of computer-mediated communication, 8(3). https://doi.org/10.1111/j.1083-6101.2003.tb00216.x wellman, b., & rainie, l. (2013). if romeo and juliet had mobile phones. mobile media & communication, 1(1), 166-171. https://doi.org/10.1177/2050157912459505 williamson, b., eynon, r., & potter, j. (2020). pandemic politics, pedagogies and practices: digital technologies and distance education during the coronavirus emergency. learning, media and technology, 45(2), 107-114. https://doi.org/10.1080/17439884.2020.1761641. zimmerman, b. j. (1995). self-regulation involves more than metacognition: a social cognitive perspective. educational psychologist, 30(4), 217-221. https://doi.org/10.1207/s15326985ep3004_8. codepen hirt et al frontline learning research vol.8 no. 4 (2020) 74 111 issn 2295-3159 types of social help-seeking strategies in different and across specific task stages of a real, challenging long-term task and their role in academic achievement carmen nadja hirta, yves karlena,francesca suterb & katharina maag merkib auniversity of applied sciences and arts northwestern switzerland, school of education, switzerland buniversity of zurich, institute of education, switzerland article received 24 february 2020 / revised 9 june / accepted 17 july / available online 5 august abstract social help seeking (shs) is an important strategy for successful self-regulated learning at all school levels. the aim of this longitudinal study is threefold: to ascertain the existence of different types of shs strategies in various task stages of creating an individual academic paper, examine the extent to which these types of shs strategies change in the course of that challenging long-term task and analyse the extent to which these types are relevant to academic achievement. this examination extends previous studies by adopting a task-specific, person-centred development perspective on shs outside regular classroom instruction. in particular, we explore shs types in the context of a real, long-term task, whereby aspects neglected in previous studies (need for help, help sources based on specific issue areas) are used for type creation and test for differences in academic achievement. three online questionnaires were completed by 603 upper secondary school-level students (62.9% female) with a mean age of 17.3 (sd = .71) within one school year. latent class analyses, latent transition analyses (lta) and non-parametric procedures (kruskal-wallis h test, post hoc dunn-bonferroni test) were performed. different shs types were identified (independents, fac-tual supervisor-focused, factual supervisor-focused and motivational family-focused, motivational family-focused, and factual and motivational family-focused) and found to vary over different task stages. moreover, lta indi-cated a considerable change between the shs types over time. nevertheless, no significant differences in achievement emerged between the types per task stage, thus reflecting the adage, 'there is more than one way of doing it'. keywords: social help seeking, longitudinal study, latent class analyses, latent transi-tion analyses, academic achievement info corresponding author email: carmen.hirt@fhnw.ch doi: https://doi.org/10.14786/flr.v8i4.627 1. introduction learners usually encounter difficulties while working on a challenging task. one strategy for overcoming a difficulty that hinders further operations involves seeking help (nelson-le gall, 1985; newman, 2000). help seeking (hs) is an important self-regulatory strategy (e.g. newman, 2000; pintrich & zusho, 2002) at all school levels (järvelä, 2011), and it represents an important external resource management strategy for achieving goals (schenke, lam, conley & karabenick, 2015). some authors include seeking help from both personal and non-personal sources in their definition of hs (e.g. aleven, stahl, schworm, fischer & wallace, 2003); other authors such as zimmerman and moylan (2009) consider hs as "[...] a social form of information seeking" (p. 303). the social aspect of seeking help is emphasized in the latter, which distinguishes this strategy from information seeking (e.g. books, internet and computer-supported, interactive learning environments). the hs strategy is used when a task cannot be completed on its own, which is also expressed by nelson-le gall (1985) in her explanation of hs as an "adaptive alternative to individual problem solving" (p. 66). as the focus in the present study is on seeking help from other persons, information seeking from non-real personal supporters is excluded. the term social help seeking (shs) is therefore used to underline this conceptual distinction. shs differs from other self-regulation strategies (schworm & fischer, 2006) because it requests interaction with other people such as teachers, peers and parents (ellis, 1997; karabenick & newman, 2010; newman, 2000), thereby confirming its uniqueness, whereas technologically mediated hs can also be social if the presence of the other individual is real (e.g. phone call, e-mail) (karabenick & newman, 2010). one consequence of the social-interactive character of the described form of help seeking is that it makes the shs process vulnerable to a variety of influences (karabenick & berger, 2013). most of the studies available on shs aimed to determine whether help is sought, for what reasons and from whom and concluded that not every shs is equally beneficial to learning (wolters, pintrich & karabenick, 2003). making a distinction between the more and less productive forms of shs is therefore important. thus far, these forms of shs have been primarily studied in the context of regular instruction, strongly based on the goals, reasons or orientations of the learners. this effort has resulted in the forms or types of instrumental shs, executive shs and shs avoidance, with different effects on achievement, both in variableand person-centred approaches. however, shs does not only occur in the setting of regular instruction. school tasks for completion outside of regular instruction can also lead to difficulties and thus to shs processes, where peers and teachers do not necessarily serve as source, but other people such as parents are possibly more likely to be available. although the contextual resources can determine the rate and effectiveness of shs (karabenick & gonida, 2018), they have been hardly examined for educational tasks in the context beyond regular instruction. the present study therefore seeks to examine shs in an educational task outside the classroom environment, which consists in writing a compulsory school leaving certificate paper at the upper secondary school level. this paper is written individually or in groups outside of class during extracurricular time over approximately one year, and it entails three task stages (development, implementation and final stages). in this context, we aim to examine whether we can group shs strategies into certain types at different stages during the entire process of developing the individual academic paper, through the adoption of a person-centred approach to focus on students as a complex system of interacting components. for the creation of these types, we particularly focus on the perceived need for help, the different issue areas that arise, and various contact persons, as these aspects, aside from the goal of shs, have been demonstrated to have a crucial influence on shs (karabenick & gonida, 2018; karabenick & knapp, 1988, 1991; makara & karabenick, 2013). currently, studies linking shs strategies to specific stages of a complete task process are non-existent, and research with a developmental perspective is sparse, even though a development perspective is especially relevant when considering issues related to the need for help and the selection of shs sources (karabenick & gonida, 2018). therefore, we examine the issue of whether and in what way shs strategies change over a longer period and identify the specific shs strategies that are the most successful in view of the need for help, different issue areas, and various contact persons considering different types. 1.1 social help seeking as a processual learning strategy depending on the author, shs is described as a process that includes five to seven steps. the following seven steps can be defined according to different authors: at the beginning of the process, the student must realise that he or she has a problem (1) and will not make any progress without help (2). ascertaining that help is needed does not automatically induce shs behaviour, as motivational, cognitive and social factors mediate students’ shs behaviour in a specific context (schworm & fischer, 2006). if the decision to seek help is made, the student sets an shs goal (3) that can be executive (i.e. asking for a direct answer or solution) or instrumental (asking for hints) in nature (karabenick & knapp, 1991; karabenick & newman, 2010; schworm, 2018). the next step is to search for potential helpers (4). this decision is influenced by various characteristics of the possible helper on the one hand, and of the person seeking help, on the other (ryan, pintrich & midgley, 2001; ryan & shin, 2011; schworm & fischer, 2006). based on the type of assistance requested (5) and received (6), the learner must ultimately decide whether such assistance has served its purpose or whether he or she should seek further help and in doing so, start a new shs cycle (7) (nelson-le gall, 1985). these steps may be conducted in a different order, and individual steps can also be undertaken in parallel (nelson-le gall, 1985). similar to zimmerman's (2002) self-regulation model, in which reflections serve to decide whether students can continue learning in a certain manner or whether adaptations are necessary, a parallel assumption for the shs process is that the final evaluation may engender an adaptation of future shs (karabenick & berger, 2013). for example, if a student could not find the desired help from a particular person, he or she may ask another individual for help in the future (nelson-le gall, 1981). therefore, shs places high demands on the reflection of previous processes and consequently on the decision making of the person seeking help. 1.2 previous types of social help seeking and their role in academic achievement regarding previous research on shs forms and types, a distinction can be made between variable-centred and person-centred approaches to highlight both the theoretical and the empirical backgrounds. variable-centred approaches focus on the relationship of the variables themselves, whereas person-centred approaches underscore the subpopulations of people identified by similar value patterns on a set of variables (finney, barry, horst & johnston, 2018). by examining the pattern of values across variables, a person is typified in a holistic meaning (magnusson, 1998). therefore, the goal of using a person-centred approach is to represent the types (nagin, 2005) and thereby to provide a "view of the person as a system of interacting components" (robins, john & caspi, 1998, p. 135). in contrast, a variable-centred approach does not provide information on the combination of dimensions within the person, but it establishes relationships between aggregated dimensions (finney et al., 2018). up to now, shs research within variable-centred approaches mainly distinguishes three forms (reasons or orientations, sometimes also referred to as types) of shs in the classroom context that chiefly emphasise the goals of learners (e.g. butler, 1998, 2006; nadler, 1998; nelson-le gall, 1981; ryan, patrick & shim, 2005). first, instrumental (autonomous, adaptive, appropriate) shs highlights the improvement of individual knowledge and competence, and here in-depth learning transpires (karabenick & newman, 2006). second, executive (expedient, dependent) shs strives to finalise solutions to avoid exertion. the third form is to avoid shs, even if help is needed. the last two forms are less conducive to sustainable learning (karabenick & newman, 2010). thus, shs not only implies an act of dependence but also represents an adaptive and strategically beneficial process (butler, 1998; karabenick, 1998; nelson-le gall, 1981). given the potential benefit of shs, various studies in this context have focused on the determinants of persons and situations (wolters et al., 2003) that affect whether help is sought, for what reasons and by whom (butler, 1998, 2006; butler & neuman, 1995; elliot & church, 1997; elliot & mcgregor, 2001; karabenick, 1998; nelson-le gall, 1981; ryan & pintrich, 1997; ryan et al., 2001). the few studies with person-centred approaches (finney et al., 2018; karabenick, 2003), whereby shs types were analysed, are of particular interest here and are therefore described in more detail. karabenick (2003) examined students using five hs scales, namely, instrumental and executive help seeking, help-seeking avoidance, help-seeking threat, and formal versus informal help seeking, whereby the latter points to the aspect of shs. all items were formulated hypothetically in relation to a possible difficulty that might occur (e.g. "getting help would be one of the first things i would do if i were having trouble in this class" (karabenick, 2003, p. 55)). on this basis, a hierarchical cluster analysis was conducted, which yielded four types of student shs: strategic/adaptive/formal (17%), strategic/adaptive/informal (25%), non-strategic (36%) and avoidant (23%). although this outcome broadly confirmed the division into the three forms of shs for the classroom context (e.g. butler, 2006; nadler, 1998; nelson-le gall, 1981), it resulted in an expansion by two specific source levels (i.e. informal versus formal). the types were subsequently examined with regard to their differences in achievement. the most clearly identified cluster was the instrumental shs type (strategic/adaptive/formal) with a preference for help from formal sources (teachers), which had the highest performance level, whereas the avoidant type exhibited the lowest performance level (karabenick, 2003). finney et al. (2018) investigated whether they could replicate karabenick’s (2003) four-type solution if they employed a more general measurement of shs. they initially examined incoming first-year college students using modified items already utilised by karabenick (2003) and related them to the general school context (i.e. across all classes during one semester). similar to karabenick’s (2003) approach, the items were formulated in such a way that the need for help should be controlled by the formulation of the question about what the students would do if they needed help (finney et al., 2018). in addition, the items were verbalised in relation to a hypothetical problem. in contrast to karabenick (2003), finney et al. (2018) produced three shs types and identified only minor differences between the individual types. hence, the four types from the first study could not be replicated, although the authors recognised a possible reason for this case in the more general measurement of shs (i.e. across several classes). these reports argue in favour of examining shs types in a much stronger task-related manner. in a second study with upper class students and the same instruments, finney et al. (2018) indicated that the types had similar profiles to the types of first-year college students. additionally, with regard to dissimilarities in achievement, finney et al. (2018) established that the instrumental and formal source-related shs dimension was positively related to optimal outcomes and negatively related to non-optimal outcomes; the avoidance, threat, and executive shs dimensions had inverse relations. altogether, the strong orientation toward the goals of shs in previous research on shs types in the classroom context becomes clear, which can demonstrate that different forms or types of shs based on the goals make dissimilar contributions to learning and thus to the achievement of the learners, both with variableand person-centred approaches (e.g. butler, 2006; karabenick, 2003; ryan et al., 2005). thus far, the focus has been primarily on the purpose of shs: to explore the shs forms and types of learners. we intend to broaden this focus through an in-depth analysis of the need for help, different perceived issue areas and various contact persons, as these factors have also emerged as central influencing factors in shs. 1.3 central influencing factors in social help seeking the need for help can be considered as a key element to understand the part played by shs in the learning process (karabenick & knapp, 1991). as awareness of the need for help constitutes the starting point of shs (ryan et al., 2001) shs "should be directly related to the learners' perceived need for help" (karabenick & gonida, 2018, p. 422). when learners are aware of what they need to learn/work effectively, they are able to take action to meet the demands of the task (nelson-le gall, 1981). karabenick and knapp (1988) demonstrated the high importance of the learner's need for help in their early analyses of students in various courses. thus, the need for help was positively related to the frequency of asking for help. moreover, this relationship was curvilinear (karabenick & knapp, 1988) so that students with a high need and students with a very low need for help were the least likely to seek help. karabenick and knapp (1988) attributed the fact that students in great need seek less help to cognitive and emotional hindrances, especially helplessness. this curvilinear relationship was also established for the link between the stated need and the grades expected by learners (karabenick & knapp, 1988). in a later study, karabenick and knapp (1991) again examined the need for help and the shs of students. they concluded that the need for help was strongly associated with shs. if grades were included, the researchers demonstrated that the need for help was inversely related to grades, and that grades were inversely related to shs. in contrast to their previous study, no significant quadratic but a linear trend could be found for the relationship between help seeking and grades. hence, karabenick and knapp (1991) revealed that learners with a low need for help sought less help and overall performed better, which they attributed to these learners' increased use of learning strategies. overall, the results indicate that when learners need help, they are more likely to seek it. however, the authors note that the relationships presented merely demonstrated what learners would do if they were confronted with a problem (karabenick & knapp, 1991). the results of studies with variable-centred approaches illustrate the relationship between the perceived need for help and shs. the inclusion of the perceived need for help in studies of shs also with person-centred approaches is therefore essential. if the help is needed and recognised as such, shs also depends on different contact persons (e.g. makara & karabenick, 2013) who are selected on the basis of the perceived issue area (e.g. boldero & fallon, 1995). contextual resources can determine the rate and effectiveness of shs (karabenick & gonida, 2018). within different courses or tasks, challenges in different sub-topics can appear for which help is needed and sought. various early studies indicated that different people asked for help depending on the issue area for which help is required. for example, family and peers are primarily asked for help with psychological, personal challenges (boldero & fallon, 1995; rickwood, 1995; tinsley, de st. aubin & brown, 1982). however, academic advisors or teachers are more often asked for help with career-related aspects (tinsley et al., 1982). therefore, the characteristics of the person providing help (e.g. confidence [newman, 2000], knowledge [stroebe, hewstone, codol & stephenson, 2013] and care [ryan & shim, 2012]) play a central role. within the challenging task of developing an individual academic paper over a longer period, huber, lehmann and husfeldt (2011) identified different thematic areas that are relevant for writing a school leaving certificate paper and can therefore cause difficulties as well. most of the pupils turned to the supervisor for challenges regarding the content and structure of the paper. a few students asked the supervisor for help with formulating the research question, formal principles (e.g. footnotes, bibliography, citations) and information sources. in terms of working methods, for questions about the timetable and organisation of work and about writing the paper (writing process), even fewer learners turned to the supervisor. students were the least likely to ask their supervisors for help with their choice of topic or with overcoming a crisis. the supervisor is evidently not selected as the first contact person for challenges in all of the issue areas. the authors revealed that other contact persons were also used when difficulties arose, including parents and peers. however, the specific issue to which the shs referred regarding these sources was unclear. from these results with variable-centred approaches, shs must be viewed in the light of the challenges in different areas, and the choice of the persons providing help is based on what kind of issue is causing difficulty. 1.4 research deficits in summary, shs has a strong social-interactive character, in which the need for help of the person seeking assistance as well as the context (task, issue area) and the contextual resources (contact persons) based on this context can have a crucial impact. previous research on shs types reported similar results across different school levels, with instrumental shs types with a preference for help from formal sources performing better than executive shs types or avoiders. there has been a strong focus to date on the goals of shs and potential helpers in the classroom context, both in person-centred (finney et al., 2018; karabenick, 2003) and variable-centred approaches (butler, 1998, 2006; butler & neuman, 1995; elliot & church, 1997; elliot & mcgregor, 2001; nelson-le gall, 1981; ryan & pintrich, 1997; ryan et al., 2001). however, several research desiderata can be identified. the few studies that dealt with shs types in a person-centred approach usually included only one measurement point in their analyses or merely depicted a cross-section and recorded shs in relation to different university courses across several different tasks. today, few studies with a developmental perspective are available (karabenick & gonida, 2018), and reports argue in favour of examining shs types in a much stronger task-related manner. furthermore, the need for help has a connection with shs. this factor must be considered in the formation of shs types. thus far, this case has hardly occurred in explicit terms. the items in the investigations to date were formulated hypothetically in relation to a possible undefined difficulty that might occur. the perceived level of need was not integrated. additionally, shs can vary based on challenges in diverse task and issue areas. therefore, shs must be considered as a task-situated process in relation to the difficulties actually experienced in different issue areas. moreover, most of the analyses especially focus on shs behaviour in the classroom context. as learning transpires not only in the classroom, shs processes should also be examined for educational tasks beyond the context of regular instruction that allow for a high degree of self-regulation; one example is writing an academic paper during extracurricular time, in which teachers or peers are not always present but other contacts may play a role. 1.5 present study: questions and hypotheses the aim of this study is to extend previous investigations by adopting a task-specific, longitudinal perspective and focusing on shs in the context of a real, challenging, long-term, academic task outside of regular classroom instruction. the study extends previous research by examining shs strategies along different types (person-centred approach), focusing on the perceived need for help, various sources and specific areas in which challenges can arise. three questions are investigated. first, do different types of shs strategies exist in each task stage of developing an academic paper outside the classroom context (q1)? due to the theory on shs with dissimilar decision possibilities (nelson-le gall, 1981), different characteristics of influence on the part of the person providing help (e.g. ryan & shim, 2012), and as students are confronted with various challenges in each creation stage (backhaus & tuor, 2008), we expect to find diverse shs types per task stage (hypothesis 1). second, what is the extent to which the types of shs strategies change during the course of the creation process (q2)? following zimmerman (2002), we assume that experiences in one stage can result in a change in shs in a subsequent stage. thus, despite the assumptions of finney et al. (2018), our presupposition is the changeability of shs behaviour over time based on different sub-tasks per stage and previous experiences. consequently, we hypothesise that students may transfer between different shs types (hypothesis 2). furthermore, assumptions in an explorative manner can be made regarding possible changes, as previous theoretical and empirical foundations are insufficient: students who received help from their paper supervisor that they found useful in the first stage need less help with the academic paper in the following stages (hypothesis 3); however, if they do not discuss the concept broadly with the supervisor, these students will need help at a later stage (hypothesis 4). we also expect that some students will need help throughout all stages, relating to different issues, and will ask various contact persons for help (hypothesis 5). third, what is the degree of importance of these shs types per task stage in academic achievement (q3)? based on previous findings of both variable-centred (butler, 1998, 2006; karabenick & knapp, 1988, 1991; nadler, 1998; nelson-le gall, 1981; ryan et al., 2005) and person-centred approaches (finney et al., 2018; karabenick, 2003), we assume that the shs types are ultimately related to different academic achievement (hypothesis 6).   2. context and methods 2.1 context in switzerland, where this study was conducted, the upper secondary school level has different tiers that are oriented toward various professional tracks. grammar schools (isced levels 3–4) represent the track with a strong emphasis on academic learning that prepares students for university. toward the end of this track, students must write a compulsory school leaving certificate paper, which is referred to as the matura thesis. this academic paper significantly contributes to the grade on the final exit examination, which in turn, if passed, guarantees unrestricted admission (with the exception of medical studies in switzerland) to a university of applied sciences or university (swiss federal council & edk, 1995). after completion of the upper secondary school leaving certificate, students have the ability to access new knowledge, develop their curiosity, imagination and communication skills, and work alone and in groups (swiss federal council & edk, 1995, art. 5). as the matura thesis is a demanding task, it requires self-regulated learning. however, it also has the potential to promote self-regulated learning competencies (huber, husfeldt, lehmann & quesel, 2008). students are given approximately one year's time to complete their papers. the papers are written individually or in groups, outside of class during extracurricular time. the developing process of the matura thesis consists of three task stages, each containing different key activities. the first stage is concept development (t1), in which a research question is identified; the learners have ample freedom to select the topic and plan implementation to answer the research question. the second phase refers to the implementation stage (t2), in which students start creating their paper. in the final stage (t3), students finish their matura thesis, evaluate and revise it, and finally submit it. 2.2 participants all study procedures complied with the human subjects' guidelines of the swiss national science foundation. students had the opportunity to withdraw their participation at any time. this longitudinal study began with 1,250 students (55.9% female) with an average age of 17.5 (sd = .81) at 12 urban and rural upper secondary schools. the number of participants was later reduced based on the criteria below. to determine possible changes between shs types and task stages, we had to reduce the sample based on two arguments. first, we included in the analyses only those students who filled out all three online questionnaires (t1, t2, t3) needed for this study. this criterion provided a longitudinal sample of n = 713 (63.0% female) students with an average age of m = 17.4 (sd = .76) over three measurement points. second, the academic paper can be written alone or in a group. to be able to compare the hs of individuals, we were interested in those students who wrote their paper on their own. we consequently excluded those students who wrote their thesis in pairs or groups. this criterion resulted in the reduction of 110 students. the final sample of this study consisted of n = 603 students (62.9% female) with an average age of 17.3 (sd = .71). up to 91.7% of these learners stated that they were born in switzerland, and 86.4% specified that they most often spoke german/swiss german at home. we compared the sample of this study with the characteristics of the population of students in grammar schools in the german-speaking part of switzerland. the sample reflected the population of this type of school with respect to gender, age, nationality and native language. 2.3 measurements the participants answered all the items in the online self-report questionnaires. for these analyses, the perceived frequency of asking the supervisor, family and peers for help, as well as the issue areas and the related contact persons were quantified at three measurement times, retrospectively for one task stage each. the perceived need for help was calculated at one measurement time, retrospectively for the three task stages of creating the academic paper. this approach resulted in 72 collected shs variables (24 variables per stage). all of the aspects included are explained in more detail in the following sections. 2.3.1. perceived need for help after the students had submitted their matura thesis (t3), we asked them about their retrospective perceived need for help in the different stages of the creation process. we asked, "to what extent were you reliant on help to continue working through the whole process?" (3 items, 1 item for each stage: "during the [fill in the stage] stage, i was...."). the responses ranged from 1 (not at all reliant on help) to 6 (entirely reliant on help). 2.3.2. perceived frequency of asking for help the perceived frequency of asking for help was measured thrice with the following question during the entire process, retrospectively for each stage (t1, t2, t3): "how often did you request help from the following persons?" participants responded on a 6-point likert scale from 1 (never) to 6 (very often) for supervisor, family and peers, and it was respectively adjusted for each creation stage (3 × 3 items). 2.3.3. issue areas and contact persons to identify the individuals from whom the students asked for help relating to specific issue areas, we asked them the following question retrospectively for each stage (t1, t2, t3) at three different points in time: "who did you ask for help concerning the following issue areas?" we gave the students different possible contact persons (nobody, supervisor (formal), family (informal), peers (informal), and others) whom they could select relating to different given issue areas (5 × 4 × 3 items). most of the issue areas were obtained from huber et al.'s (2011) study, as they proved to be relevant for writing a school leaving certificate paper. however, to ensure the comparability of the types across the phases, those issue areas that turned out to be relevant and thus possibly leading to difficulties for all three phases were integrated into the present analyses, namely, information source, working methods, timetable and organisation of the paper (factual issue areas), and motivation and resolving crises (motivational issue areas). for this question, multiple answers were possible (1 = selected, 2 = not selected). 2.3.4. academic achievement academic achievement was also recorded as the students’ grades on their academic papers. the paper supervisors and the second assessors evaluated the papers. we were given access to the students’ official grades, which in switzerland range from 1 to 6, with 6 being the highest grade. 2.4 statistical analyses as a person-centred approach, we conducted latent class analyses (lca) to ascertain the existence of different shs types at each task stage of generating an academic school leaving certificate paper. latent class analysis is a statistical method for classifying individuals into homogeneous subgroups (latent classes) (finch & bronk, 2011). to calculate the results of the lca, we used mplus version 8.1 (muthén & muthén, 1998-2017). to consider missing values, the full information maximum likelihood method (fiml) was used. the proportion of missing values of the variables was on average 1.99% (min = 0.50%, max = 5.14%). the maximum likelihood robust estimator (mlr) was used to account for a possible deviation from the multivariate normal distribution. with the categorical command in mplus, we specified that the issue area and contact person items were ordered categorical (dichotomous in this case) variables, whereas all the other variables (perceived need for help, perceived frequency of asking for help) were treated as metric variables. we were able to conduct a mixed distribution analysis (lca in this case) in mplus by using the command analysis of type = mixture. with mixed distribution analysis, we assumed that the population of students was composed of different subgroups (latent classes, hs types). the model quality statistics and parameter estimates of the lca were determined in mplus using the maximum likelihood estimation method. the log-likelihood value is a measure of the probability of the data provided by the model, and it serves as a basis for the calculation of further model fit tests or indices (finch & bronk, 2011). as the final number of classes is unknown beforehand, different models with a changing number of classes have to be tested and compared in terms of various statistical and non-statistical criteria (sample size, interpretability of the classes, average latent class probabilities). the entropy (e) assesses the quality of the measurement as a whole (asparouhov & muthén, 2014), and it should have a value greater than .80 (rost, 2006). furthermore, the values of the akaike information criterion (aic) and of the bayesian information criterion (bic) were included in the analyses, whereby the lower values correspond to the better fitting model (geiser, 2011). the lo-mendell-rubin (lmr) and the bootstrap likelihood ratio test (blr) were also utilised to identify the suitable number of classes. if the lmr and the blr tests are significant, the model with k classes represents the data better than the model with k-1 classes (finch & bronk, 2011). the differences between the classes per stage regarding perceived need for help, perceived frequency of asking for help, issue areas and contact persons, and achievement were analysed using nonparametric methods, namely, the kruskal-wallis test and post hoc dunn-bonferroni test for categorical and metric variables (martens, 2003), as the data contained outlier values that were retained due to inconstancy (sheskin, 2011). the kruskal-wallis h test is based on rank numbers assigned to individual characteristic values. finally, the sums of the rank numbers were calculated for each group (shs type per stage) and were analysed for significant differences (martens, 2003). we conducted a latent transition analysis (lta) to examine the changes in the course of the learning process. latent transition analysis can be considered as a longitudinal extension of the lca. it uses an autoregressive relationship to link the latent class variables of several measurement points (muthén & asparouhov, 2011). the latent transition probabilities, which are of particular interest in the analyses of the lta results, indicate the probability of being assigned to a specific class at time t based on the assignment to a class at time t-1 (muthén & asparouhov, 2011). further details can be found in the next section. 3. results 3.1 descriptive statistics table 1 presents the descriptive statistics for all categorical variables. the values represent the percentages of selected contact persons per issue area and the measurement points. table 2 presents the descriptive statistics for all metric variables. table 1 descriptive statistics for the categorical variables note. t1 = concept development stage; t2 = implementation stage; t3 = final stage; n = number of cases; 100% minus the percentage of the selected response option corresponds to the percentage of the not selected response option. table 2 descriptive statistics for the metric variables note. t1 = concept development stage; t2 = implementation stage; t3 = final stage; n = number of cases; m = mean; sd = standard deviation; perceived need for help: 1 = not at all reliant on help, 6 = entirely reliant on help; perceived frequency of asking for help: 1 = never, 6 = very often; academic achievement: 1 = lowest grade, 6 = highest grade. 3.2 identification and description of the help-seeking types per stage (lca) table 3 lists the models that were best suited to the data and their criteria to statistically evaluate the model fit for the different class solutions generated by the lca. the analyses were based on the aforementioned 72 collected shs variables (i.e. 24 variables per stage), referring to perceived need for help, perceived frequency of asking for help, and the issue areas and contact persons (see section 2.3). table 3 statistical fit indices for the most appropriate class solutions at different measurement times note. t1 = concept development stage; t2 = implementation stage; t3 = final stage; bic = bayesian; aic = akaike; e = entropy; be = boundary estimates (logit thresholds that were set at the extreme values -15.000 and 15.000); plmr = significance of the lo-mendell-rubin; pblr = significance of the bootstrap likelihood ratio test. initially, the 4-class solution for measurement point t1, the 5-class solution for t2 and the 6-class solution for t3 seem to be ideal (see table 3). however, closer inspection apparently indicated that these solutions contained many boundary estimates, thus making interpretation difficult. even though the bic is the most frequently used decision criterion, it sometimes refers to theoretically implausible solutions, especially for large samples, and results in the overestimation of class numbers (specht, luhmann & geiser, 2014). hence, the criteria were only evaluated in combination. the 3-class solution is preferred at all measurement times due to the interpretability and uniqueness of the classes, statistical fit indices, average latent class probabilities and the smallest number of boundary estimates. the following sections describe the shs types in more detail for each stage of the creation process of the matura thesis. first, the differences between the classes per measurement point are listed. as both metric and binary individual items were included in the lca, the results are then listed separately for each task stage. 3.2.1. social help-seeking types in the concept development stage (t1) the first classes refer to the concept development stage. three classes could be identified, which mainly differ in the categorical and metric variables examined (see table 4). table 4 differences between classes in the concept development stage (t1) note. χ2 = chi-square (df = 2); kruskal-wallis test and post hoc dunn-bonferroni test for categorical and metric variables. figure 1 shows the mean values of the metric class variables for the three latent classes, and table 5 presents the respective probabilities for the categorical class variables. the first class containing 20.6% (n = 124) of the students surveyed (see figure 1) mostly had values that were in-between the ranges of the other classes. with an average value of m = 3.53 (sd = 1.16), they had a self-reported dependence on help that ranked above the second class and below the third class. similarly, the self-reported perceived frequency of asking for help was in the middle range for the supervisor (m = 3.89, sd = 1.08) and for the family (m = 3.01, sd = 1.49) compared to the other classes. they exhibited a slightly higher value (m = 2.78, sd = 1.33) only regarding perceived frequency of asking peers for help. however, they had the highest value on seeking help from the supervisor. this result already indicated that students in this first class especially asked their supervisor for help, if at all. the results of the categorical variables (see table 5) confirmed this initial assumption and revealed that for the factual issue areas (information source and working methods), the primary person asked for help was the supervisor. these results explicated the class that we named factual supervisor-focused. figure 1. mean values of the respective metric class variables in the concept development stage (t1, n = 603). perceived need for help: 1 = not at all reliant on help, 6 = entirely reliant on help. perceived frequency of asking: 1 = never, 6 = very often. class 1 = factual supervisor-focused, class 2 = independents, class 3 = factual supervisor-focused and motivational family-focused. the second class containing 31.3% (n = 189) of the students (see figure 1) had the overall lowest values compared to the other classes in the concept development stage. these students gained an average value of m = 3.22 (sd = 1.16). according to self-reports, the students were minimally dependent on help. the perceived frequency of asking for help from the supervisor revealed the lowest value for all the classes with an average value of m = 3.59 (sd = 1.08). the values for the perceived frequency of asking the family m = 2.98 (sd = 1.49) and peers m = 2.33 (sd = 1.33) for help were the lowest in comparison to the other classes as well. for motivational aspects (motivation and resolving crises), the students did not ask anyone for help (see table 5). these findings explained the class that we named independents. the third class (see figure 1), which contains 48.1% (n = 290) of the students, was the largest of the three classes in the concept development stage. compared to the other two classes, the students in the third class stated on average that they were most dependent on help (m = 3.65, sd = 1.16) and that they sought help most from the supervisor (m = 4.00, sd = 1.08) and the family (m = 4.00, sd = 1.49). this group was found in a middle range (m = 2.77, sd = 1.33) only when seeking help from peers. for factual issue areas (information source, working methods, and timetable and organisation of work, see table 5), the students in this group stated that they turned to supervisors. in contrast, for motivational aspects (motivation and resolving crises), the family was the primary contact person. these outcomes clarified the class name factual supervisor-focused and motivational family-focused table 5 probabilities for the respective categorical class variables in the concept development stage (t1) note. class 1 = factual supervisor-focused; class 2 = independents; class 3 = factual supervisor-focused and motivational family-focused; in bold type = probabilities with a value > 0.500. 3.2.2. social help-seeking types in the implementation stage (t2) three classes could also be found for the implementation stage. table 6 shows the differences between classes in the implementation stage. as in the concept development stage, the metric class variables were examined in more detail before exploring the categorical variables for each of the classes at this stage. figure 2 shows the mean values of the metric class variables for the three latent classes, and table 7 presents the respective probabilities for the categorical class variables. the first class containing 31.9% (n = 192) of the students (see figure 2) had the lowest overall values compared to the other classes. with an average value of m = 2.98 (sd = 1.02), students in this class had a low self-reported perceived need for help. the mean values for perceived frequency of seeking help for the two contact groups family (m = 2.78, sd = 1.44) and peers (m = 1.92, sd = 1.27) were the lowest overall that could be found at this stage. the mean value for self-reported perceived frequency of asking for help from the supervisor was at a medium level (m = 3.64, sd = 1.02) compared to the others. once again, this finding could be confirmed by the probability table (see table 7), which showed that the students in the first class of this stage had individually indicated that they had not asked anyone for help. this case had already emerged in the concept development stage, and no value above 0.50 could be found for the issue area 'information source'. for this reason, this class was called independents. table 6 differences between classes in the implementation stage (t2) note . χ2 = chi-square (df = 2). the second class (see figure 2) formed the largest group in the implementation stage with 42.9% of the students (n = 259). it seemed to be highly similar to the third class from the first stage. figure 2 shows that the second class had by far the highest levels of self-reported perceived need for help (m = 3.66, sd = 1.02) and perceived frequency of asking for help from the supervisor (m = 4.04, sd = 1.02), family (m = 3.94, sd = 1.44) and peers (m = 2.55, sd = 1.27) compared to the other two classes at this stage. a closer examination of the issue areas in combination with the contact persons (see table 7) apparently indicated that the students in the second class primarily turned to the supervisor for fact-related issue areas (information source, working methods, timetable and organisation of work). however, they seemed to prefer to ask the family for help in overcoming crises (motivation and resolving crises). we named this class factual supervisor-focused and motivational family-focused. the third class (25.2%, n = 152) constituted the smallest class (see figure 2). here, the students showed an average value of m = 3.25 (sd = 1.02) for self-reported need for help and therefore were in the medium range. the perceived frequency of help seeking was also in the medium range (family, m = 3.44, sd = 1.43; peers, m = 2.18, sd = 1.27), except for perceived frequency of asking the supervisor, which was the lowest in comparison (m = 3.56, sd = 1.02). this class also stated that they had not turned to anyone except for help on motivational issue areas (motivation and resolving crises, see table 7). according to the students, the family was consulted in such cases. this rationale underlies the naming of this class as motivational family-focused. figure 2. mean values of the respective metric class variables of the implementation stage (t2, n = 603). perceived need for help: 1 = not at all reliant on help, 6 = entirely reliant on help. perceived frequency of asking: 1 = never, 6 = very often. class 1 = independents, class 2 = factual supervisor-focused and motivational family-focused, class 3 = motivational family-focused. table 7 probabilities for the respective categorical class variables in the implementation stage (t2) note. class 1 = independents; class 2 = motivational family-focused; class 3 = factual supervisor-focused and motivational family-focused; in bold = probabilities with a value > 0.500. 3.2.3 social help-seeking types in the final stage (t3) we also found three classes in the last stage of the creation process (about one month before the submission of the paper). the differences between classes in the final stage are presented in table 8. figure 3 includes the mean values of the metric class variables for the three latent classes and table 9 presents the respective probabilities for the categorical class variables. table 8 differences between classes in the final stage (t3) note: χ2 = chi-square (df = 2). an examination of the metric class variables (see figure 3) revealed that the first class (28.4%, n = 171) had the lowest mean values compared to the other two classes. the relatively low dependence on help (m = 3.63, sd = 1.05) indicated by the students also corresponded to perceived frequency of seeking help from other persons (supervisor, m = 3.03, sd = 1.27; family, m = 3.54, sd = 1.31; peers, m = 1.99, sd = 1.39). the probabilities regarding the categorical class variables (see table 9) similarly confirmed this finding, as the students in this class were apparently the most likely not to have asked anyone for help in the issue areas surveyed. hence, we named the class independents. the second class (39.4%, n = 238, see figure 3) had a self-reported need for help of m = 3.98 (sd = 1.05). the perceived frequency of seeking help from the supervisor was m = 3.20 (sd = 1.27), m = 4.27 (sd = 1.31) for the family, and m = 2.69 (sd = 1.39) for the peers (see figure 3). based on the metric class variables, these values were middle ranged compared to the other two classes. however, if one considers the probabilities of the categorical variables (see table 9), the family was clearly asked for help in particular with the issue area of timetable and organisation of work and also when motivational aspects were involved (motivation and resolving crises). the family thus played a central role in this class, and we therefore called it motivational family-focused. the family was likewise central to the third class (32.2%, n = 194, see figure 3), especially in matters of solving problems relating to issues in timetable and organisation of work and overcoming crises (see table 9). the mean values for the metric class variables (see figure 3) were the highest compared to the other two classes (perceived need for help, m = 4.30, sd = 1.05; perceived frequency of asking supervisor, m = 3.62, sd = 1.27; perceived frequency of asking family, m = 4.79, sd = 1.31; perceived frequency of asking peers, m = 2.75, sd = 1.39). given this family-based orientation, we named this class factual and motivational family-focused. figure 3. mean values of the respective metric class variables in the final stage (t3, n = 603). perceived need for help: 1 = not at all reliant on help, 6 = entirely reliant on help. perceived frequency of asking: 1 = never, 6 = very often. class 1 = independents, class 2 = motivational family-focused, class 3 = factual and motivational family-focused. table 9 probabilities for the respective categorical class variables of the final stage (t3) note: class 1 = independents; class 2 = motivational family-focused; class 3 = factual and motivational family-focused; in bold = probabilities with a value > 0.500. 3.3 changes in social help-seeking types over time (lta) before proceeding to the results of the lta, several methodological notes need to be provided. the first is that "[if] the same measurement model is used across all time points (e.g., lca) and the same number and type of classes are used, it is reasonable to explore measurement invariance" (nylund, 2007, p. 44). further, nylund (2007) notes that a full measurement invariance is not always plausible, depending on which classes have emerged from the lca. in this investigation, the lca produced the same number but not the same types of classes over time. one of the three classes seemed to occur at all three measurement points (independents), and two classes transpired in each case at two measurement points (factual supervisor-focused and motivational family-focused, motivational family-focused). therefore, model 1 was constructed first, in which the thresholds and intercepts became restricted over time for the same occurring classes. model 1 was finally compared with a second model (model 2), in which the thresholds and intercepts were freely estimated over time. due to the large number of cells in the frequency table, the chi-square test could not be computed. a model comparison using the aic and bic provides information on the particular model that is more suitable for the available data, whereby the lower values correspond to the better fitting model (geiser, 2011). as an additional dimension, the entropy, which should have a value greater than .80, was used (rost, 2006). finally, the more suitable model was used for interpreting the latent transition probabilities. table 10 summarizes the values of the model comparison: model 1 had no restrictions toward model 2, which had restrictions for constant classes over time. the model comparison indicated that the two models only differed marginally from each other. as the entropy value for model 1 was more favourable, this model was used for further analyses. this decision can also be justified theoretically, because although some classes remained the same over time, this does not signify that all persons remained in the same classes. table 10 statistical fit indices for the most appropriate model: comparison of the freely estimated model (model 1) with the restricted model (model 2) note. bic = bayesian; aic = akaike; a = independents, t1/2, t2/1, t3/1; factual supervisor-focused and motivational family-focused, t1/3, t2/2; motivational family-focused, t2/3, t3/2. figure 4 shows the latent transition probabilities for the total sample from t1 to t2 and from t2 to t3. the probabilities of 0.11 and 0.27 (t1 to t2), and 0.41 (t2 to t3) in bold type indicated a relatively low stability of the intraindividual behavioural style of students in the classes of the same name over time. first, the probabilities for a change in the shs type from the concept development stage (t1) to the implementation stage (t2) were considered. persons who belonged to the factual supervisor-focused class (class 1) in the concept development stage (t1) switched to independents (class 1) for the implementation stage (t2) with a 62% probability and to the motivational family-focused class (class 3) with a probability of 16%. as the factual supervisor-focused only existed for t1, we cannot speak of permanent members of this group. all these students changed to another shs group for the implementation stage (t2): 62% changed to independents, 23% to factual supervisor-focused and motivational family-focused and 16% to motivational family-focused. students who belonged to the independents (class 2) in the concept development stage (t1) were likely to change to the factual supervisor-focused and motivational family-focused (class 2) for the implementation stage (t2) with a 66% probability and to the motivational family-focused (class 3) with a probability of 13%. however, approximately 11% remained in the class of independents. persons who belonged to the class factual supervisor-focused and motivational family-focused (class 3) switched to independents (class 1) for the implementation stage (t2) with a probability of 28% and to the class motivational family-focused (class 3) with a probability of 45%. about 27% remained in the class factual supervisor-focused and motivational family-focused (class 2). second, we examined the probabilities of changing the shs type from the implementation stage (t2) to the final stage (t3). students in the class independents (class 1) had a probability of 41% of remaining in this class for the final stage, and they did not change to motivational family-focused (class 2). however, they were 59% likely to switch to factual and motivational family-focused (class 3). persons who belonged to the factual supervisor-focused and motivational family-focused (class 2) in the implementation stage (t2) did not switch to independents (class 1), but they were 46% likely to join the motivational family-focused (class 2) in the final stage. nevertheless, 54% moved to the factual and motivational family-focused (class 3). persons belonging to t2 in the class of motivational family-focused (class 3) neither switched to independents nor stayed in the same class in the final stage. a full 100% of this class 3 changed to factual and motivational family-focused for the final stage (t3). figure 4. transition probabilities based on the different shs types and measurement points (n = 568). only changes are shown. t1 = concept development stage, t2 = implementation stage, t3 = final stage. (1), (2), (3) = class number per measurement point. in bold type = remained in the same class. 3.4 disparities between the social help-seeking types per stage regarding academic achievement to answer the third question, a kruskal-wallis h test was run to determine any differences in grades between the three shs types per stage and therefore ascertain whether the shs type to which the students belonged per stage played a role in academic performance. the distributions of grades were similar for all shs types in all three stages, as assessed by the visual inspection of the box plots. the results indicated that the median values did not differ statistically significantly between the shs types in the concept development stage (t1), χ2(2) = 2.20, p = .333, or in the implementation stage (t2), χ2(2) = 1.36, p = .506, or in the final stage (t3), χ2(2) = 0.54, p = .763 (see table 11). table 11 disparities between shs types per stage regarding academic achievement note. t = measurement point; n = number of cases; mdn = median (1 to 6); χ2 = chi-square; df = degrees of freedom; p = level of significance (asymptotic, two-sided). no multiple comparisons due to non-significant differences between samples in the overall test. 4. discussion 4.1 overall results the main objectives of this study were to adopt a person-centred approach to gain in-depth insights into how young adults seek help in dealing with a real, challenging long-term task; to attempt to break it down into specific types of shs strategies in view of the need for help, different issue areas and various contact persons; and to integrate the role of academic achievement. swiss students at different grammar schools at the upper secondary level participated in the study and thus provided us with insights into their help-seeking behaviour. the aim of this study was threefold. the first aim was to ascertain the existence of different shs types in each task stage of creating an academic paper outside the classroom context (q1). the shs process is characterized by many different decision-making options (nelson-le gall, 1981). in addition, shs is dependent on various characteristics of the person giving help (e.g. ryan & shim, 2012). based on these points and the diverse challenges that the learners encounter in each creation phase, we expected to find different shs types in each stage of the long-term task under investigation (hypothesis 1). this hypothesis can be confirmed: not every stage includes all the same shs types. the lcas per stage indicate that one of the three classes occurs in all stages of the paper creation process, namely, independents. this group is referred to as avoiders in karabenick (2003) because of their low values in terms of the quantity of shs. however, we chose to call them independents instead, because in addition to the low perceived frequency of asking for help, they also report less perceived need for help. the members of this class could be students who already have some previous knowledge and therefore encounter fewer challenges. this low perceived need for help and the associated low perceived frequency of asking for help confirm the results of karabenick and knapp (1988, 1991), in which students with a very low need are the least likely to seek help. the extent to which this case is due to competencies in self-regulated learning (karabenick & knapp, 1991) would require further examination. students who are factual supervisor-focused and motivational family-focused only appear in the first two stages. these group members seem to be highly uncertain about what to exactly do. furthermore, they do not seem to know for sure how to motivate themselves in the concept development and the implementation stage, based on their comparatively highest values for both stages regarding perceived need and perceived frequency of seeking help from the supervisor concerning factual issues and from the family regarding motivational issues. however, the underlying reasons for these generally increased values would have to be verified in further investigations. in contrast, the motivational family-focused group only appears in the last two stages, which makes sense, as maintaining motivation and persisting is especially relevant for long-term tasks, and challenges in this regard might arise more toward the middle and the end of the process (ulmi, bürki, verhein & marti, 2017). research has also revealed that motivational aspects are preferably discussed with trusted persons (boldero & fallon, 1995; fallon & bowles, 1999; nelson le gall, gumerman & scott-jones, 1983). factual supervisor-focused students only appeared in the concept development stage. motivational challenges do not seem to be an issue for this group. in addition, the factual and motivational family-focused appear in only one of the stages, namely, the final stage. these students mainly report asking the family for help with motivational issues and issues in work organisation. however, a striking aspect of the final stage is that none of the class members has a strong orientation toward supervisors. work on the paper should have further progressed by this time; hence, the inhibition threshold for asking the supervisor might naturally rise, as the supervisor is also the assessor (bonati & hadorn, 2009). another possibility is that the ensuing questions no longer concern the area of competence or responsibility of the supervisor, such as proofreading the work. overall, the supervisors do not seem to be perceived as the contact persons with whom crises or motivational challenges can be discussed. this inference particularly confirms the findings of boldero and fallon (1995), who highlighted that personal issues are most likely to be discussed with family or peers. thus, compared to karabenick’s types (2003), we have found a hybrid type in which informal and formal sources are asked for help, namely, the factual supervisor-focused and motivational family-focused. this type appears because we focused on the specific issue areas for certain contact persons, which in turn underlines the relevance of this differentiated perspective. the second aim was to determine the extent to which these shs types change during the course of the creation process of a long-term task (q2). in accordance with zimmerman (2002), we postulated that experiences in one stage can induce a change in self-reported behaviour in the subsequent stage. additionally, different problems can arise in long-term tasks, which in turn can have an influence on shs. we consequently assumed that the students can switch between the shs types and that these types do not have to be stable over time (hypothesis 2), as was expected by finney et al. (2018). this hypothesis can be confirmed, as the lta indicates a diligent switching between shs types over time. whether these changes are due to adaptive shs processes or to different key activities within the three stages requires further investigation. regarding the transfer between types, the explorative assumption was made that students who receive help from their supervisor that they find useful in the first stage need less help later on (hypothesis 3). sixty-two percent of the factual supervisor-focused students (t1) change to the class of independents (t2). thus, hypothesis 3 can be largely confirmed. it indicates that these students feel that they are prepared to implement the paper concept based on a productive concept stage with the supervisor. an effective concept and precise aim specifications facilitate targeted work afterwards (mccardle, webster, haffey & hadwin, 2017; ritschl, weigl & stamm, 2016), which in turn benefits independent work. thus, we also assumed that without a broad discussion on or clarification of the concept, the students might need help at a later stage (hypothesis 4). seventy-six percent of students from the independents (t1) changed to the factual supervisor-focused and motivational family-focused (t2) type of shs, confirming hypothesis 4. these students seem to have many open factual issues and apparently need help with motivational issues in the implementation stage. the assumption is that the concept was not specifically thought through or that certain questions will only arise during the implementation stage. questions that are not clarified at the concept development stage can reappear in the implementation stage. another reason could be a relatively large amount of prior knowledge, which can result in the perceived need for help at the concept development stage (t1) being underestimated by overestimating the understanding of the learning material (scardamalia & bereiter, 1992). we also expected students who need help throughout all the phases with different issues and therefore ask different contact persons for help (hypothesis 5). students assigned to the motivational family-focused during the implementation stage switch 100% to the factual and motivational family-focused for the final stage. this large group of students is apparently dependent on help from their families not only with motivational issues but also with factual issues, especially timetable and organisation of the work. hence, hypothesis 5 can be partially confirmed. although these students continuously seek help, the contact persons do not vary considerably: these students strongly involve their families to resolve challenges. however, a conclusion that can be derived based on the results is that issues with content simply no longer exist and that these queries must have been resolved with the supervisor beforehand. this group does not turn to the supervisor during either the implementation stage or the final stage; thus, a performance goal orientation should also be considered, as these students may not intend to embarrass themselves in front of their assessing supervisor (karabenick, 2003; ryan et al., 2001). the third aim of the present research was to examine the extent to which the shs types found per stage are important for academic achievement (q3). against the background of previous studies (butler, 1998, 2006; finney et al., 2018; karabenick, 2003; karabenick & knapp, 1988, 1991; nadler, 1998; nelson-le gall, 1981; ryan et al., 2005), we had assumed that the shs types differ in their academic performance (hypothesis 6). this hypothesis could not be confirmed, because no significant differences between the shs types per stage can be identified. the students examined here were apparently able to adequately assess their need for help and then seek and receive help in such a way that they received a favourable grade on the paper. however, it is another question whether this help-seeking behaviour enhances the student's understanding in the area for which help was sought. this condition would require closer examination of individual help-seeking/help-giving interactions and behaviour and knowledge in future, similar situations. 4.2 conclusion for theory and practice it should be noted for shs theory that shs types can significantly change during task stages and thus within a learning process as well. we have found some shs types that are addressed by the types discovered in previous investigations. nevertheless, through the identification of the thematic focus of shs, other shs types emerge, such as factual supervisor-focused and motivational family-focused. this inference denotes that future investigations should pay more attention to hs as an issue-specific process. furthermore, shs avoidance should not be stated unless the need for help has also been considered. this group of students could be quickly misunderstood and mistakenly classified as a 'risk group'. therefore, shs types should be interpreted with caution because they may also substantially change depending on the context. in the context of preparing an academic paper, the students apparently rarely turn to their paper supervisors for help with motivational issues. nevertheless, paper supervisors in particular could teach suitable motivational regulatory strategies to encourage their students to continue their work (dignath & büttner, 2008). in this case, the supervisors must first be aware that their students are struggling with motivational difficulties. social interactions concerning not only content but also motivational issues should be intensified (wolters, 2003). however, this process requires a strong basis of trust between supervisor and student, a relationship that would also allow students to reveal their weaknesses or difficulties (butler & neuman, 1995; newman, 2000; ryan & pintrich, 1997). in addition, 62% of the students who were in the factual supervisor-focused group at the concept development stage switched to the independents. this result illustrates the relevance of close support during the development of the concept for independent work during the implementation stage, which is also described in the coaching literature (e.g. ulmi et al., 2017). most of the students in this study seem to know how to achieve desired grades through help-seeking processes if they have a perceived need for help. nonetheless, for the help-givers, this help should be provided in a way that enables students to help themselves in similar situations (e.g. by developing appropriate motivational regulation strategies). in other words, helpers such as parents, peers and supervisors should not primarily focus on the students' good grades but should instead direct their support toward sustainability and lifelong learning (eu council, 2002). 4.3 limitations and suggestions for further research directions in this study, we adopted a longitudinal, person-centred perspective on shs types outside the classroom context and integrated important aspects such as perceived need for help, specific issue areas, and the related sources of help. we could therefore show that shs types can vary based on the different issue areas encountered. however, the data evaluated in this study are based on self-reports regarding shs. measurements that are more objective could probably provide deeper and/or extended insights into the shs processes. this aspect especially relates to measurement of the frequency of seeking help. notably, our response options reflected the individually perceived frequency of help seeking and not objective frequency. the same applies to assessing the need for help. these aspects should be considered in future studies. although the recorded academic achievements were not based on self-reports, the students showed a relatively high average grade, which may have caused insignificant differences between the shs types. a comparison of extreme groups and the resulting shs types could lead to extended results. the backgrounds of the individual class members must also be further analysed to improve the understanding of the shs strategies of the students in these classes; the reason is that individual characteristics such as gender (e.g. nadler, 1998) and goal orientation (e.g. butler, 2006; ryan et al., 2001) can affect shs behaviour. the influence of the parents' support on the shs types is an important aspect that is lacking in research up to now. as newman (2000) was able to confirm that experiences from the parental home can influence help seeking, this area should also be integrated into further analyses of the shs types. finally, how different types of shs can influence performance levels, as well as how performance is influenced by switching between classes, should be taken into consideration. the outcomes clearly indicate that without the inclusion of analyses relating to family relationships and the cooperation/relationship between the students and the paper supervisor, only vague assumptions can be made about the reasons for the changes. this factor must be taken into account in further analyses. the present analyses are related to the context of producing a compulsory school leaving certificate paper (matura thesis) at the end of the upper secondary school. further analyses in the same or a similar context would be needed to strengthen these results, which could particularly show that shs should be considered context-specifically. keypoints students who do not seek help are not necessarily social help-seeking (shs) avoiders, but can also be students who work independently (i.e. students who do not perceive a need for help). within a long-term challenging task with various key activities, students can switch between shs types. for the final grade, the shs type per task stage is unimportant in this context. references aleven, v., stahl, e., schworm, s., fischer, f., & wallace, r. (2003). help seeking and help design in interactive learning environments. review of educational research, 73, 277-320. https://doi.org/10.3102/00346543073003277 asparouhov, t., & muthén, b. o. (2014). variable-specific entropy contribution. retrieved from https://www.statmodel.com/download/univariateentropy.pdf backhaus, n., & tuor, r. (2008). leitfaden für wissenschaftliches arbeiten, 7. überarbeitete und ergänzte auflage [guidelines for scientific work, 7th revised and supplemented edition]. zurich, switzerland: department of geography, university of zurich. https://doi.org/10.5167/uzh-10134 boldero, j., & fallon, b. j. (1995). adolescent help seeking: what do they get help for and from whom? journal of adolescence, 18, 193-209. https://doi.org/10.1006/jado.1995.1013 bonati, p., & hadorn, r. (2009). maturaund andere selbständige arbeiten betreuen. ein handbuch für lehrpersonen und dozierende. 2., überarbeitete und erweiterte auflage [supervising matura and other independent work. a manual for teachers and lecturers. 2nd, revised and extended edition]. bern, switzerland: hep verlag. butler, r. (1998). determinants of help seeking: relations between perceived reasons for classroom help-avoiding and help-seeking behaviors in an experimental context. journal of educational psychology, 90(4), 630-644. https://doi.org/10.1037/0022-0663.90.4.630 butler, r. (2006). an achievement goal perspective on student help seeking and teacher help giving in the classroom: theory, research, and educational implications. in s. a. karabenick & r. s. newman (eds.), help seeking in academic settings: goals, groups, and contexts (pp. 17-34). new york, ny: erlbaum. butler, r., & neuman, o. (1995). effects of task and ego achievement goals on help-seeking behaviors and attitudes. journal of educational psychology, 87, 261-271. https://doi.org/10.1037/0022-0663.87.2.261 dignath, c., & büttner, g. (2008). components of fostering self-regulated learning among students: a meta-analysis on intervention studies at primary and secondary school level. metacognition learning, 3, 231-264. https://doi.org/10.1007/s11409-008-9029-x elliot, a. j., & church, m. a. (1997). a hierarchical model of approach and avoidance achievement motivation. journal of personality and social psychology, 72(1), 218-232. https://doi.org/10.1037/0022-3514.72.1.218 elliot, a. j., & mcgregor, h. a. (2001). a 2 x 2 achievement goal framework. journal of personality and social psychology, 80, 501-519. https://doi.org/10.1037/0022-3514.80.3.501 ellis, s. (1997). strategy choice in sociocultural context. developmental review, 17, 490-524. https://doi.org/10.1006/drev.1997.0444 eu council. (2002). council resolution of 27 june 2002 on lifelong learning. official journal of the european communities, 9. retrieved from https://op.europa.eu/en/publication-detail/-/publication/0bf0f197-5b35-4a97-9612-19674583cb5b fallon, b. j., & bowles, t. (1999). adolescent help-seeking for major and minor problems. australian journal of psychology, 51(1), 12-18. https://doi.org/10.1080/00049539908255329 finch, h. w., & bronk, k. c. (2011). conducting confirmatory latent class analysis using mplus. structural equation modeling: a multidisciplinary journal, 18(1), 132-151. https://doi.org/10.1080/10705511.2011.532732 finney, s. j., barry, c. l., horst, s. j., & johnston, m. m. (2018). exploring profiles of academic help seeking: a mixture modeling approach. learning and individual differences, 61, 158-171. https://doi.org/10.1016/j.lindif.2017.11.011 geiser, c. (2011). datenanalyse mit mplus: eine anwendungsorientierte einführung [data analysis with mplus: a practical introduction]. wiesbaden, germany: vs verlag für sozialwissenschaften. https://doi.org/10.1007/978-3-531-92042-9 huber, c., husfeldt, v., lehmann, l., & quesel, c. (2008). projektteil d2: die qualität von maturaarbeiten in der schweiz [project part d2: the quality of matura work in switzerland.]. in f. eberle, k. gehrer, b. jaggi, m. kottnau, m. oepke, c. pflüger, c. huber, v. husfeldt, l. lehmann, & c. quesel (eds.), evaluation der maturitätsreform 1995 (evamar). schlussbericht zur phase ii (pp. 277-352). bern, switzerland: edi, sbf. retrieved from http://edudoc.ch/record/29677/files/web_evamar-komplett.pdf huber, c., lehmann, l., & husfeldt, v. (2011). unterschiedliche rahmenbedingungen bei der realisierung von maturaarbeiten [differing framework conditions for the realisation of matura theses]. revue suisse des sciences de l’éducation, 33(3), 443-460. retrieved from https://www.pedocs.de/volltexte/2015/10122/pdf/szbw_2011_3_huber_ua_unterschiedliche_rahmenbedingungen.pdf järvelä, s. (2011). how does help seeking help? new prospects in a variety of contexts. learning and instruction, 21, 297-299. https://doi.org/10.1016/j.learninstruc.2010.07.006 karabenick, s. a. (1998). strategic help-seeking: implications for learning and teaching. mahwah, nj: erlbaum. karabenick, s. a. (2003). seeking help in large college classes: a person-centered approach. contemporary educational psychology, 28, 37-58. https://doi.org/10.1016/s0361-476x(02)00012-7 karabenick, s. a., & berger, j.-l. (2013). help seeking as a self-regulated learning strategy. in h. bembenutty, t. j. cleary, & a. kitsantas (eds.), applications of self-regulated learning across diverse disciplines: a tribute to barry j. zimmerman (pp. 237-261). charlotte, nc: information age publishing. karabenick, s. a., & gonida, e. n. (2018). academic help seeking as a self-regulated learning strategy: current issues, future directions. in d. h. schunk & j. a. greene (eds.), handbook of self-regulation of learning and performance (pp. 421-433). new york, ny: routlege. karabenick, s. a., & knapp, j. r. (1988). help seeking and the need for academic assistance. journal of educational psychology, 80(3), 406-408. https://doi.org/10.1037/0022-0663.80.3.406 karabenick, s. a., & knapp, j. r. (1991). relationship of academic help seeking to the use of learning strategies and other instrumental achievement behavior in college students. journal of educational psychology, 83(2), 221-230. https://doi.org/10.1037/0022-0663.83.2.221 karabenick, s. a., & newman, r. s. (2006).help seeking in academic settings: goals, groups and contexts. mahwah, n.j.: lawrence erlbaum associates. karabenick, s. a., & newman, r. s. (2010). seeking help as an adaptive response to learning difficulties: person, situation, and developmental influences. in s. järvela (ed.), social and emotional aspects of learning (pp. 244-250). kidlington, oxford, uk: elsevier. https://doi.org/10.1016/b978-0-08-044894-7.00610-2 magnusson, d. (1998). the logic and implications of a person-oriented approach. in r. b. cairns, l. r. bergman, & j. kagan (eds.), methods and models for studying the individual: essays in honor of marian radke-yarrow (pp. 33-64). thousand oaks, ca: sage. makara, k. a., & karabenick, s. a. (2013). characterizing sources of academic help in the age of expanding educational technology: a new conceptual framework. in s. a. karabenick & m. puustinen (eds.), advances in help-seeking research and applications: the role of emerging technologies (pp. 37-72). charlotte, nc: information age publishing. martens, j. (2003). statistische datenanalyse mit spss für windows [statistical data analysis with spss for windows]. munich, germany: oldenbourg wissenschaftsverlag. https://doi.org/10.1515/9783486815085 mccardle, l., webster, e. a., haffey, a., & hadwin, a. f. (2017). examining students’ self-set goals for self-regulated learning: goal properties and patterns. studies in higher education, 42(11), 2153-2169. https://doi.org/10.1080/03075079.2015.1135117 muthén, b. o., & asparouhov, t. (2011, 27. july). lta in mplus: transition probabilities influenced by covariates. mplus web notes: no. 13. retrieved from http://www.statmodel.com/examples/ltawebnote.pdf muthén, l. k., & muthén, b. o. (1998-2017). mplus user’s guide (8th ed.). los angeles, ca: muthén & muthén. retrieved from https://www.statmodel.com/download/usersguide/mplususerguidever_8.pdf nadler, a. (1998). relationship, esteem, and achievement perspectives on autonomous and dependent help seeking. in s. karabenick (ed.), strategic help seeking: implications for learning and teaching (pp. 61-93). mahwah, new jersey: erlbaum. nagin, d. s. (2005). group-based modeling of development. cambridge, ma: harvard press. https://doi.org/10.4159/9780674041318 nelson-le gall, s. (1981). help-seeking: an understudied problem-solving skill in children. developmental review, 1(224-246). https://doi.org/10.1016/0273-2297(81)90019-8 nelson-le gall, s. (1985). help-seeking behavior in learning.review of research in education , 12, 55-90. https://doi.org/10.3102/0091732x012001055 nelson le gall, s., gumerman, r. a., & scott-jones, d. (1983). instrumental help-seeking and everyday problem-solving: a developmental perspective. in b. m. de-paulo, a. nadler, & j. d. fisher (eds.), new directions in helping (pp. 265-281). new york, ny: academic press. newman, r. s. (2000). social influences on the development of children’s adaptive help-seeking: the role of parents, teachers, and peers. developmental review, 20, 350-404. https://doi.org/10.1006/drev.1999.0502 nylund, k. l. (2007). latent transition analysis: modeling extensions and an application to peer victimization. retrieved from http://www.statmodel.com/download/nylund%20dissertation%20updated1.pdf pintrich, p. r., & zusho, a. (2002). the development of academic self-regulation: the role of cognitive and motivational factors. in a. wigfield & j. s. eccles (eds.), development of achievement motivation (pp. 249-284). san diego, ca: academic press. https://doi.org/10.1016/b978-012750053-9/50012-7 rickwood, d. j. (1995). the effectiveness of seeking help for coping with personal problems in late adolescence. journal of youth and adolescence, 24, 685-703. https://doi.org/10.1007/bf01536951 ritschl, v., weigl, r., & stamm, t. (2016). wissenschaftliches arbeiten und schreiben. verstehen, anwenden, nutzen für die praxis [academic work and writing. comprehension, application and benefit for practice]. berlin, germany: springer. https://doi.org/10.1007/978-3-662-49908-5 robins, r. w., john, o. p., & caspi, a. (1998). the typological approach to studying personality. in r. b. cairns, l. r. bergman, & j. kagan (eds.), methods and models for studying the individual: essays in honor of marian radke-yarrow (pp. 135-160). thousand oaks, ca: sage. rost, j. (2006). latent class analysis. in f. petermann & m. eid (eds.), handbuch der psychologischen diagnostik [manual of psychological diagnostics]. göttingen, germany: hogrefe verlag. ryan, a. m., patrick, h., & shim, s.-o. (2005). differential profiles of students identified by their teacher as having avoidant, appropriate, or dependent help-seeking tendencies in the classroom. journal of educational psychology, 97(2), 275-285. https://doi.org/10.1037/0022-0663.97.2.275 ryan, a. m., & pintrich, p. r. (1997). ‘‘should i ask for help?’’ the role of motivation and attitudes in adolescents’ help seeking in math class. journal of educational psychology, 89, 329-341. https://doi.org/10.1037/0022-0663.89.2.329 ryan, a. m., pintrich, p. r., & midgley, c. (2001). avoiding seeking help in the classroom: who and why? educational psychology review, 13(2), 93-114. https://doi.org/10.1023/a:1009013420053 ryan, a. m., & shim, s. s. (2012). changes in help seeking from peers during early adolescence: associations with changes in achievement and perceptions of teachers. journal of educational psychology, 104 (4), 1122-1134. https://doi.org/10.1037/a0027696 ryan, a. m., & shin, h. (2011). help-seeking tendencies during early adolescence: an examination of motivational correlates and consequences for achievement. learning and instruction, 21, 247-256. https://doi.org/10.1016/j.learninstruc.2010.07.003 scardamalia, m., & bereiter, c. (1992). text-based and knowledge-based questioning by children. cognition and instruction, 9(3), 177-199. https://doi.org/10.1207/s1532690xci0903_1 schenke, k., lam, a. c., conley, a. m., & karabenick, s. a. (2015). adolescents’ help seeking in mathematics classrooms: relations between achievement and perceived classroom environmental influences over one school year. contemporary educational psychology, 41, 133-146. https://doi.org/10.1016/j.cedpsych.2015.01.003 schworm, s. (2018). lernen in computerbasierten lernumgebungen: instruktionale unterstützungsmöglichkeiten [learning in computer-based learning environments: instructional support possibilities]. in m. heilemann, h. stöger, & a. ziegler (eds.), lernen im internet (pp. 93-112). berlin, germany: lit verlag. schworm, s., & fischer, f. (2006). academic help seeking. in h. mandl & h. f. friedrich (eds.), handbuch lernstrategien (pp. 282-239). göttingen, germany: hogrefe. sheskin, d. j. (2011). handbook of parametric and nonparametric statistical procedures. boca raton, fl: chapman & hall/crc press. specht, j., luhmann, m., & geiser, c. (2014). on the consistency of personality types across adulthood: latent profile analyses in two large-scale panel studies. journal of personality and social psychology, 107, 540–556. https://doi.org/10.1037/a0036863 stroebe, w., hewstone, m., codol, j.-p., & stephenson, g. m. (2013). sozialpsychologie: eine einführung. [social psychology: an introduction]. heidelberg, germany: springer verlag. swiss federal council, & edk. (1995). verordnung des bundesrates/reglement der edk über die anerkennung von gymnasialen maturitätsausweisen (mar) vom 16. januar/15. februar 1995 [ordinance of the federal council/regulation of the edk on the recognition of matura certificates (mar) of 16 january/15 february 1995]. bern, switzerland: schweizerischer bundesrat/edk. retrieved from https://edudoc.educa.ch/static/web/aktuell/medienmitt/vo_mar_1995_d.pdf tinsley, h. e. a., de st. aubin, t., & brown, m. (1982). college students' help-seeking preferences. journal of counselling psychology, 29, 523-533. https://doi.org/10.1037/0022-0167.29.5.523 ulmi, m., bürki, g., verhein, a., & marti, m. (2017).textdiagnose und schreibberatung. fachund qualifizierungsarbeiten begleiten [text diagnosis and writing advice: accompanying specialized and qualification work] (2nd ed.). berlin, germany: barbara budrich. wolters, c. a. (2003). regulation of motivation: evaluating an underemphasized aspect of self-regulated learning. educational psychologist, 38(4), 189-205. https://doi.org/10.1207/s15326985ep3804_1 wolters, c. a., pintrich, p. r., & karabenick, s. a. (2003).assessing academic self-regulated learning. paper presented at conference on indicators of positive development: definitions, measures, and prospective validity, washington dc. retrieved from https://www.researchgate.net/profile/stuart_karabenick/publication/225229608_assessing_academic_self-regulated_learning/links/5416daec0cf2bb7347db788a/assessing-academic-self-regulated-learning.pdf zimmerman, b. j. (2002). becoming a self-regulated learner: an overview. theory into practice, 41(2). https://doi.org/10.1207/s15430421tip4102_2 zimmerman, b. j., & moylan, a. r. (2009). where metacognition and motivation intersect. in d. j. hacker, j. dunlosky, & a. c. graesser (eds.), handbook of metacognition in education. new york, ny: routledge. codepen knoop publication frontline learning research vol.8 no. 4 (2020) 37 51 issn 2295-3159 how teachers integrate dashboards into their feedback practices carolien knoop-van campena& inge molenaara abehavioural science institute, radboud university, the netherlands article received 12 march 2020 / revised 20 may / accepted 18 june / available online 13 july abstract in technology empowered classrooms teachers receive real-time data about students’ performance and progress on teacher dashboards. dashboards have the potential to enhance teachers’ feedback practices and complement human-prompted feedback that is initiated by teachers themselves or students asking questions. however, such enhancement requires teachers to integrate dashboards into their professional routines. how teachers shift between dashboardand human-prompted feedback could be indicative of this integration. we therefore examined in 65 k-12 lessons: i) differences between humanand dashboard-prompted feedback; ii) how teachers alternated between humanand dashboard-prompted feedback (distribution patterns); and iii) how these distribution patterns were associated with the given feedback type: task, process, personal, metacognitive, and social feedback. the three sources of feedback resulted in different types of feedback: teacher-prompted feedback was predominantly personal and student-prompted feedback mostly resulted in task feedback, whereas dashboard-prompted feedback was equally likely to be task, process, or personal feedback. we found two distribution patterns of dashboard-prompted feedback within a lesson: either given in one sequence together (blocked pattern) or alternated with studentand teacher-prompted feedback (mixed pattern). the distribution pattern affected the type of dashboard-prompted feedback given. in blocked patterns, dashboard-prompted feedback was mostly personal, whereas in mixed patterns task feedback was most prevalent. hence, both sources of feedback instigation as well as the distribution of dashboard-prompted feedback affected the type of feedback given by teachers. moreover, when teachers advanced the integration of dashboard-prompted feedback in their professional routines as indicated by mixed patterns, more effective types of feedback were given. keywords: teacher dashboards; feedback; adaptive learning technologies info corresponding author: e-mail: c.knoop-vancampen@pwo.ru.nl doi: https://doi.org/10.14786/flr.v8i4.641 1. introduction with the growing use of educational technologies, teacher dashboards are increasingly available in classrooms. while students are practising with educational technologies, dashboards provide teachers with concurrent information about students’ performance, pace, and progress (van leeuwen, janssen, erkens, & brekelmans, 2014; molenaar & van schaik, 2016). teachers can use dashboards with additional information to improve their feedback practices with dashboard-prompted feedback. this complements human-prompted feedback, which teachers initiate themselves and/or give in response to students’ questions. so far, most studies concerning teacher dashboards have primarily focused on how teachers understand dashboard information and translate this into action (molenaar & knoop-van campen, 2019; verbert, et al., 2014). to the best of our knowledge, how the source of feedback, e.g. humanversus dashboard-prompted feedback, is related to the type of feedback given has not yet been investigated. nor is it known how teachers alternate between humanand dashboard-prompted feedback during teaching and whether this distribution pattern is related to the feedback given. in this paper, we postulate that dashboard-prompted feedback is given alongside human-prompted feedback and that the way teachers alternate between dashboardand human-prompted feedback could be indicative of how well they have incorporated dashboards into their professional routines. we therefore examined: i) differences between humanand dashboard-prompted feedback; ii) how teachers alternated between humanand dashboard-prompted feedback (distribution patterns); and iii) how these distribution patterns were associated with the type of feedback given. by investigating all feedback and not only dashboard-prompted feedback, this paper takes a broad perspective on how dashboards are integrated into teachers’ feedback practices. 1.1 teacher feedback practices providing feedback in classroom situations is a complex professional task (roelofs & sanders, 2007). decisions on how to help students during learning are based on teachers’ pedagogical knowledge base, which consists of information on students’ knowledge and abilities, teachers’ perceptions of their students, teachers’ content knowledge combined with more general knowledge and beliefs about effective pedagogical practices (meijer, verloop, & driel, 2001; roelofs & sanders, 2007). this pedagogical knowledge base therefore combines content, pedagogy and learners’ characteristics in such a way that teachers can take effective pedagogical actions (gudmuindsdottir & shulman, 1987). the practical impact of teachers’ professional knowledge base is visible in their professional routines, i.e., “patterns and routines of action, interaction, and sense-making” (ballet, & kelchtermans, 2009, p. 1153). such routines are indicative of the way teachers apply their pedagogical knowledge base and also extend to feedback provided by teachers to their students. in order for teachers to provide appropriate feedback, it is important to identify students’ current level of performance and knowledge and target feedback at identified needs of the students (wood, brunner, & ross, 1976). in the classroom context, feedback can be defined as information provided by the teacher regarding aspects of the students’ performance or behaviour (hattie, & timperley, 2007). five types of feedback are often distinguished: process, metacognitive, task, social, and personal feedback (hattie, & timperley, 2007; keuvelaar-van den bergh, 2013). process feedback gives information about students’ progress towards learning goals (hattie, & timperley, 2007), while metacognitive feedback helps students to control and monitor their learning (de jager, jansen, & reezigt, 2005). these two types of feedback are the most effective types of feedback to improve learning as they increase students’ self-regulation and strategic handling (hattie, 2012). next, task feedback informs students about the state of their performance and helps them to reflect on their current understanding (butler, & winne, 1995; hattie, & timperley, 2007). as task feedback directly supports task execution, it is also supportive for learning (keuvelaar-van den bergh, 2013). social feedback helps students to adequately collaborate with other students and is found to be important in collaborative learning settings to enhance learning outcomes (keuvelaar-van den bergh, 2013). personal feedback, which supports students to improve their behaviour during learning, is considered less effective as it does not explain how student behaviour is related to task or process elements and therefore entails little direction to improve learning (shute, 2008). teachers are thus not only challenged to diagnose when students are in need of feedback, but also to select the appropriate and effective type of feedback accordingly. even though much has been written about feedback, less is known about which types of feedback teachers actually give during lessons (bennett, 2011; voerman, meijer, korthagen, & simons, 2012). we found only two studies that investigated teachers’ actual feedback practices in classrooms. in these studies, observations were performed to identify the feedback types teachers gave in class. one study observed 32 teachers in primary education and showed that they mostly gave task and process feedback, whereas metacognitive, social, and personal feedback were hardly used (keuvelaar-van den bergh, 2013). this indicated that in primary education one of the most effective types of feedback, namely metacognitive feedback, is scarce. a second study (voerman, meijer, korthagen, & simons, 2012), in which 78 teachers in secondary education were observed, found that these teachers mostly provided students with non-specific feedback, for example providing personal encouragement such as “well done” and this did not give students directions to improve learning. additionally, they also noticed that only 7% of all feedback given was process feedback. this is problematic, as process feedback is helpful for students to improve their learning. the authors concluded that teachers seldom provided effective feedback during classroom activities (voerman, meijer, korthagen, & simons, 2012). combined, these two studies indicate that the type of feedback provided in classrooms is not always effective and that there are still considerable gains to be achieved with regards to teachers' feedback practices. 1.2 how teacher dashboards augment feedback practices in addition to human-prompted feedback given by teachers on their own initiative (teacher-prompted feedback) or in response to students’ questions (student-prompted feedback), feedback can also be elicited from information teachers view on their dashboards. teacher dashboards are increasingly used in k-12 education. these dashboards provide additional information on students’ performance, pace and progress and, as such, can be viewed as tools to augment teachers’ pedagogical knowledge base (holstein, mclaren, & aleven, 2017; molenaar & knoop-van campen, 2019). during learning, software captures real-time data on learner performance which is immediately displayed to the teachers on their laptop or computer screen (dashboards). as such, dashboards provide teachers with organized visualizations giving information about individual students or the whole class. class dashboards often provide an overview of all of the students’ correct and incorrect answers on problems. individual dashboards mostly show individual students’ progress on different learning goals and the development of their knowledge and skills, while predictive analytics are also presented. however, merely providing this information on dashboards is not enough to impact teachers’ feedback practices. even though professional routines are flexible and develop continuously (lacourse, 2011), actively and intentionally changing these professional routines by integrating new tools, such as dashboards, involves several stages of implementation. the learning analytics process model identifies four stages which teachers have to go through before dashboard data can impact their teaching practices (verbert, et al., 2014). first, teachers need to become consciously aware of data on the dashboards and build understanding as to when they can use this information during teaching (awareness stage). second, teachers have to be able to formulate questions for the data to answer (reflection stage). for example, “how can i see when a student needs help and which feedback is most appropriate?” third, teachers have to analyse the data to answer these questions (sense-making stage). for example, “lia makes a lot of mistakes; she needs additional help”. fourth and last, teachers need to determine which response to the data is appropriate and fits their analysis of the situation best (impact stage). for example, “lia does not seem to understand how to simplify mixed fractions, i should explain to her that she needs to extract the whole numbers first”. for dashboard data to be converted into feedback, teachers need to enact all stages of the learning analytics process model. during this process, they transform data into meaningful feedback actions to improve their teaching (molenaar & knoop-van campen, 2019). the learning analytics process model described above provides a theoretical model to understand how teachers translate data into action. however, how teachers apply this in their real-life professional routines remains unclear. there are only limited empirical insights into the actual impact of dashboards on teachers’ feedback practices in classrooms (van leeuwen, janssen, erkens, & brekelmans, 2014). some initial evidence has been found for the enactment of the learning analytics process model in practice; teacher awareness of dashboards was found to positively influence the feedback they gave to students (molenaar & knoop-van campen, 2017a). when teachers consulted dashboards more often, they activated more and also more diverse knowledge in their pedagogical knowledge base. furthermore, teachers who activated more diverse knowledge also gave more and more different types of feedback to students (molenaar & knoop-van campen, 2019). other studies have also provided evidence that dashboards changed teachers’ feedback practices. teachers tend to provide more feedback when they are supported by dashboards compared to situations without dashboards (van leeuwen, janssen, erkens, & brekelmans, 2014). in addition, regarding feedback allocation, one study showed that teachers who received real-time notifications about student performance directed their attention to low performing students more than they did without dashboards (martinez-maldonado, clayphan, yacef, & kay, 2014). this in turn led to improvements in these students’ performance. another study indicated that teachers also gave more feedback to good students who were struggling (knoop-van campen, wise, & molenaar, submitted.). this study also showed that dashboard-prompted feedback given to low performing students entailed equal amounts of task and process feedback while human-prompted feedback to this group consisted mostly of task feedback. lastly, an augmented reality dashboard that indicated which students needed additional support while practising in an intelligent tutor system caused teachers to spend more time on students who showed poor productivity in learning, which consequently had a positive impact on their learning (holstein, hong, tegene, mclaren, & aleven, 2018). thus, there is initial evidence that dashboard-prompted feedback is profoundly different from human-prompted feedback and can enhance learning outcomes. dashboards not only increased the amount of feedback given, but also elicited more effective types of feedback. in these studies, however, there was great variation in the frequency of dashboard-use and the type of feedback given during lessons (e.g. martinez-maldonado, clayphan, yacef, & kay, 2014; molenaar & knoop-van campen, 2017a, 2019). given the positive benefits of teacher dashboards on students’ learning, it is important to examine how this variation can be explained. with the learning analytics process model and the importance of enacting all stages in mind, a possible explanation may be the level at which teachers incorporated dashboards into their professional routines. the way teachers alternate between humanand dashboard-prompted feedback in their daily classroom activities may reflect how well they have integrated dashboards into their professional routines and this may be related to how well they enact the stages of the learning analytics process model. properly enacting the stages of the model enables teachers to optimally translate the dashboard data into practice. better integration of the dashboard into their feedback practices could therefore be followed by more effective types of feedback (i.e., more specific and more responsive feedback) after consulting the dashboard. 1.3 present study there is initial evidence that dashboards have a positive impact on teachers’ feedback practices, but only when they effectively integrate the dashboard information into their professional routines. previous studies have not specifically addressed dashboard-prompted feedback, as differences between dashboardand human-prompted feedback and how teachers integrate dashboards into their professional routines have not been investigated. therefore, we examined i) differences between humanand dashboard-prompted feedback; ii) how teachers alternated between humanand dashboard-prompted feedback (distribution patterns); and iii) how these distribution patterns were associated with the type of feedback given. we expected that dashboard-prompted feedback would elicit more effective types of feedback (e.g., task and process feedback) than human-prompted feedback. regarding the integration of dashboard-feedback into lessons, we expected to see two distribution patterns. first, when teachers have not (or not yet) fully integrated dashboards into their professional routines, they will likely use them only during a particular phase of the lesson. we hypothesize that this may result in a blocked pattern with dashboard-prompted feedback in one part of the lesson and human-prompted feedback in the other parts of the lesson. in contrast, when teachers become more proficient in using dashboards and have incorporated them into their professional routines, we expect more flexible dashboard use during the lesson. this may result in patterns displaying a mix of humanand dashboard prompted feedback (mixed patterns), in which teachers use teacher-, studentand dashboard-prompted feedback interchangeably. better integration of dashboards into teachers’ professional routines is expected to result in more effective types of feedback being given after consulting the dashboards. 2. method 2.1 participants in total, 65 lessons were observed: 45 teachers were observed of whom 20 were observed twice. all lessons were 50 minutes long and taught in year group 2 (8-year-old students) to year group 6 (12-year-old students). lessons were arithmetic or spelling lessons, dealing with topics on the regular school curriculum. the teachers were mostly female (75%), between 20 and 65 years old spread evenly across the age range and with a corresponding range of teaching experience between 2 and 30 years. adaptive learning technology was used in these classrooms on a daily basis. while students worked on problems in the adaptive learning technology, real-time data was shown on the teacher dashboards. teachers had between 1 and 3 years of experience in working with this technology. 2.2 materials 2.2.1 adaptive learning technology the adaptive learning technology (alt) used in this study called ‘snappet’ , ran on tablet computers and is widely used for arithmetic and spelling across schools in the netherlands (molenaar & knoop-van campen, 2016). the arithmetic problems in the alt were comparable to those done by students in regular classrooms. the alt offered both adaptive and non-adaptive problems. non-adaptive problems were pre-selected for a particular topic and all students in the class received the same non-adaptive problems in a lesson. the adaptive problems adjusted to the skills of the individual student. a derivative of the elo algorithm adapted problems to the current knowledge level of the student (klinkenberg, straatemeier, & van der maas, 2011). the algorithm worked with a student’s knowledge score; the representation of a student’s current level of knowledge on a particular topic. the knowledge score was calculated based on all problems that a student had worked on. every problem in the system had a difficulty level which was automatically generated and updated by the system based on all of the student’s answers (klinkenberg, straatemeier, & van der maas, 2011). based on the student’s knowledge level, the alt selected the next practice problem. the problem was selected in such a way that the student had a 75% probability of answering correctly. 2.2.2 dashboards teachers viewed a visualization of students’ data on the dashboard (see figure 1). the software captures real-time data on learner performance which was immediately displayed to the teachers on dashboards. the dashboard showed information on problems students had worked on. after a student’s name, it indicated how many problems students had solved (progress) and whether the problems were answered correctly (performance). the circles indicated problems answered. green indicated a correct answer, red an incorrect response and combined green with red circles indicated a correct response on the second attempt. a blue open circle indicated the current problem. the first section of the dashboard dealt with non-adaptive problems that were part of the lesson, the second section indicated adaptive problems with the topic the students were working on in the heading above the circles. finally, the progress indictor (human icon in front of students’ names) showed which students were making progress (green icon), not making progress (red icon) or were currently unknown (grey icon). the system defined these aspects based on an algorithm and visualized it with color-coding of these icons. figure. 1. teacher dashboard (anonymized) 2.2.3 the observations the classroom observation app (molenaar & knoop-van campen, 2017b) was used to track the sources and feedback types. observations were performed by trained student assistants using the observation app. these observers were instructed to code all feedback actions during a lesson. for each feedback action, the source (teacher, student, dashboard) and the type (task, process, personal, metacognitive, and social feedback; see table 1) were coded (hattie, & timperley, 2007; keuvelaar-van den bergh, 2013). teachers accessed the dashboards on their computer screen, laptop, or tablet. computers and laptops were situated in such a way that they were accessible to the teacher during lessons. table 1. types of feedback 2.3 data analyses to examine differences between humanand dashboard-prompted feedback, we ran a chi-square with bonferroni column proportion comparisons. the column proportions tests were used to determine the relative ordering of categories of the columns (type of feedback) in terms of the category proportions of the rows (source of feedback). to examine how teachers alternated between human and dashboard-prompted feedback, we plotted feedback actions within one lesson on three distinct levels (teacher, student, dashboard) on the y-axis and all feedback actions on the x-axis (see figure 4). a grounded approach was used to determine the codes on a subset of the sample. the subset of plots was inspected and this comparison lead to three categories: no dashboard-prompted feedback, blocked distribution pattern, or mixed distribution pattern (see figure 4 in the results section). the remaining plots were then coded by two independent coders, with a very high cohen's kappa (k = .93). where the two coders disagreed, codes were discussed until agreement was reached. chi-square analysis with bonferroni column proportions were used to understand how these distribution patterns were associated with the type of feedback given. 3. results 3.1 descriptives in 65 lessons, 3,410 feedback actions were observed. the source was recorded for 3,330 actions (80 missing) and the type was recorded for 3,237 actions (173 missing). overall, teacher-prompted feedback was most prevalent (1,869 times (56%)), followed by student-prompted feedback (867 times (26%)) and dashboard-prompted feedback (594 times (18%)) (see figure 2a). regarding all feedback actions, task feedback was given most often (997 times (31%)) closely followed by process feedback (980 times (30%)) and personal feedback (894 times (28%)). both metacognitive feedback (199 times (6%)) and social feedback (167 times (5%)) were less frequent (see figure 2b). figure 2a. feedback source figure 2b. type of feedback 3.2 associations between source and type of feedback first, we examined the association between source and feedback type. there was a significant association, χ2 (8, n = 3157) = 466.81, p < .001 (see figure 3). the most likely feedback types differed depending on the source of the feedback. below we further specify the results for teacher-, studentand dashboard-prompted feedback. bonferroni column proportions were used to determine how the type of feedback was associated with the source of feedback. when teachers triggered feedback themselves (teacher-prompted feedback), personal feedback (n = 692, 37%) was most frequently given followed by process (n = 477, 26%) and task feedback (n = 410, 22%). social (n = 146, 8%) and metacognitive feedback (n = 126, 7%) were less frequently given. bonferroni column proportions indicated a significant difference between all the feedback types in how often each feedback type was given. this means that personal feedback was most likely to follow when teachers prompted feedback, followed by process, task, social and metacognitive feedback. when teachers reacted to students’ questions (student-prompted feedback), task feedback (n = 413, 48%) was most frequently given, followed by process feedback (n = 354, 41%). personal feedback (n = 56, 7%), metacognitive (n = 33, 4%) and social feedback (n = 6, 1%) were less frequent. bonferroni column proportions indicated that task and process feedback were equally likely to occur after students’ questions. both task and process were more likely to be given in response to students’ questions than personal, social, and metacognitive feedback. metacognitive feedback was more likely to be given than social feedback but less likely than the other types of feedback. when teachers provided feedback after dashboard consultation (dashboard-prompted feedback), task feedback (n = 160, 36%) was most frequently given, followed by personal (n = 125, 28%) and process feedback (n = 124, 28%). again metacognitive (n = 31, 7%) and social feedback (n = 4, 1%) were less frequent. bonferroni column proportions indicated that except for social feedback, which was less likely to be given, the other four types of feedback were equally likely. figure 3. feedback type for teacher-, studentand dashboard-prompted feedback 3.3 patterns of dashboard-prompted feedback second, we examined how teachers alternated between humanand dashboard-prompted feedback. thirteen lessons (20%) did not show any dashboard-prompted feedback (see figure 4a for an example: there was no feedback prompted by the dashboard). furthermore, we found two different distribution patterns: fifteen lessons (23%) showed a blocked distribution pattern (see figure 4b for an example: dashboard prompted-feedback occurred in one or more chunks during the lesson), and thirty-seven lessons (57%) showed a mixed pattern where dashboard feedback was alternated with teacher and student feedback (see figure 4c for an example: the feedback source interchanged during the lesson). each figure shows all feedback actions in that specific lesson sequentially on the x-axis and in three rows dashboard-, student-, and teacher-prompted feedback on the y-axis. figure 4a. no dashboard-prompted feedback figure 4b. blocked pattern figure 4c. mixed pattern 3.4 dashboard-prompted feedback third, we investigated how the two distribution patterns (blocked and mixed) were associated with the type of dashboard-prompted feedback given. there was a significant association between the distribution pattern and the feedback type, χ2 (4, n = 438) = 33.56, p < .001 (see figure 5). the most likely feedback types differed between the two distribution patterns. in lessons with a block pattern, most dashboard-prompted feedback was personal (n = 38, 36%), process (n = 26, 25%), or task feedback (n = 23, 22%). some metacognitive (n = 18, 17%) and no social feedback (n = 0, 0%) was given. bonferroni column proportions indicated that personal feedback was equally likely to be given as process, but more likely than task or metacognitive feedback. metacognitive feedback was less likely to be given than task and process feedback, but just as likely as social feedback. in lessons with a mixed pattern, most dashboard-prompted feedback was task (n = 135, 41%), process (n = 95, 29%), and personal feedback (n = 87, 26%). some metacognitive (n = 12, 4%) and social feedback (n = 4, 1%) was given. bonferroni column proportions indicated that all types of feedback, except for social feedback, were equally likely to be given. social feedback was less likely to be given than the other types, but just as likely as metacognitive feedback. to conclude, during blocked pattern lessons personal feedback was most likely to be given, whereas in mixed pattern lessons task feedback was most frequent. in figure 5, we can observe a reversed pattern of task, process, and personal feedback in the blocked versus the mixed patterns. figure 5. dashboard-prompted feedback per distribution pattern 4. discussion dashboards have the potential to enhance teachers’ feedback practices and complement human-prompted feedback that is initiated by teachers themselves or students asking questions. however, such enhancement requires teachers to integrate dashboards into their professional routines. how teachers shift between dashboardand human-prompted feedback could be indicative of this integration. therefore, we examined i) differences between humanand dashboard-prompted feedback; ii) how teachers alternated between humanand dashboard-prompted feedback (distribution patterns); and iii) how these distribution patterns were associated with the type of feedback given. results will foster new understanding on how dashboards are integrated into teachers’ feedback practices. 4.1 how teachers integrate dashboards into their feedback practices regarding the first research question, results indicated that, in line with our expectation, feedback source and feedback type were related. teacher-prompted feedback was predominantly personal and student-prompted feedback mostly resulted in task feedback, whereas dashboard-prompted feedback was equally likely to be task, process, or personal feedback. this indicates that dashboards induce more diverse feedback practices. moreover, as teacher-prompted feedback constitutes mostly personal feedback which is less effective, we postulate that dashboard-prompted feedback may be more efficient. moreover, dashboards stimulated the less frequently used process feedback, which is known to be positively associated with learning (hattie, 2012). whereas van leeuwen, janssen, erkens, and brekelmans (2014) showed that dashboards changed the number of teachers’ feedback actions, we found that dashboards also changed the quality of feedback, by increasing the number of effective feedback practices. secondly, we investigated how teachers alternated between human and dashboard-prompted feedback and whether so-called distribution patterns of dashboard-prompted feedback could be observed. the two expected distribution patterns were found. there were lessons in which dashboard-prompted feedback occurred in one part of the lesson and human-prompted feedback in another part of the lesson (blocked pattern). there were also lessons in which teachers used teacher-, studentand dashboard-prompted feedback interchangeably and where dashboard-prompted feedback was mixed with human-prompted feedback throughout the whole lesson (mixed pattern). these distribution patterns indicate that there was not only variation in the extent to which teachers used the dashboard, but also in how they integrated dashboards into their professional routines. additionally, it is important to note that there was also a group of teachers (20%) who did not give any dashboard-prompted feedback. these teachers did not integrate the dashboards into their feedback practices. thirdly, we examined how the observed patterns of dashboard-prompted feedback were associated with feedback type. in line with our expectations, the type of dashboard-prompted feedback was found to be affected by the distribution pattern it occurred in. dashboard-prompted feedback in blocked patterns was mostly personal feedback, while in mixed patterns task feedback was most prevalent. ergo, higher levels of alternation of dashboard-prompted feedback with other feedback sources, leads to more task feedback. we postulate that different patterns might be indicative of differential development of teachers’ professional routines with regards to the analysis and application of dashboard information. even though all teachers who use teacher dashboards enact the learning analytics process (verbert, et al., 2014), the achieved depth and effectiveness of their awareness, reflection, sense-making, and implementation in their feedback actions differs between teachers. the patterns we observed visualized these differences in implementation and showed that along with integration of dashboard-feedback among their other feedback, teachers also gave more efficient types of feedback. teachers who gave dashboard-prompted feedback in blocked patterns mostly gave personal feedback, which resembles the main feedback type also initiated by themselves. this may indicate that teachers recognized a specific student was in need of support, but only acted on a general level providing personal feedback. still, teachers showing mixed patterns have developed their analytic skills to integrate the dashboard information within their existing knowledge base and consequently are able to follow up with task-related support (butler & winne, 1995; hattie & timperly, 2007). hence, this informs a hypothesis that teachers go through different stages in learning how to integrate dashboards into their daily practices and that distribution patterns may be indicative of their current level of integration. future research should further explore the relation between distribution patterns of dashboard-prompted feedback in lessons and teacher proficiency in these feedback practices. 4.2 limitations the limitation of this study was that it was a naturalistic study of teachers’ feedback practices in lessons. there was no strict study design and hence cause and effect relations could not be detected. however, the study did provide ecologically validated data on how teachers in real life use and integrate dashboard data, which cannot be simulated in a lab setting. to understand teachers’ feedback actions in response to dashboards, it is vital to take into account their pedagogical knowledge base and thus investigate their professional routines in their class with children they know and work with on a daily basis. 4.3 future research future research could investigate how the integration of dashboard-prompted feedback develops over time. as more integrated use of dashboards seems to relate to positive changes in feedback type, a training intervention could be envisaged, in which teachers would be taught to enhance their analytical abilities. this could support understanding and interpretation of the information on the dashboard and could enhance teachers’ capacities to integrate dashboard information into their existing pedagogical knowledge base. future research in turn could investigate whether this enhances teachers’ feedback practices. it would also be useful to examine whether particular teacher characteristics, for example teachers' experience with the dashboard or their analytical skills, relate to the integration of dashboard-prompted feedback.   4.4 practical implications practical implications of this study are fourfold. first, we now know that dashboards elicit more diverse feedback practices compared to human-prompted feedback. hence, dashboards seem to sustain rich feedback practices especially with respect to inducing more process feedback. second, we learned that different sources of feedback are related to different types of feedback. in professional training the function of different sources of feedback could be highlighted. third, mixed patterns are more inductive to task feedback, hence teacher training could focus on how teachers integrate dashboard-prompted feedback into their overall feedback practices. fourth, even though dashboard-prompted feedback is more diverse than human-prompted, both social and metacognitive feedback are underrepresented. social feedback is less inherent to the context studied, but metacognitive feedback could be beneficial for students. especially as young learners’ metacognitive skill have not yet fully developed (veenman et al. 2006), it is important that teachers support the acquisition of these skills by providing feedback. hence, dashboard designers should consider how to include information that can elicit metacognitive feedback. 4.5 conclusions to conclude, we showed that dashboard-prompted feedback is more diverse than feedback provided on teachers’ own initiative and in response to students’ questions. the way dashboard-prompted feedback is distributed in a lesson affected the feedback type teachers gave. this distribution may be indicative of teachers’ professional routines regarding dashboard usage and dashboard-prompted feedback and more specifically be indicative of the extent to which dashboards are integrated into educational practice. in technology empowered classrooms, dashboards thus provide the opportunity to optimize teachers’ feedback practices when integrated into their professional routines. keypoints feedback initiated after dashboard consultation is more diverse compared to teacher-prompted and student prompted feedback teachers gave dashboard-prompted feedback in one block or they mixed humanand dashboard-prompted feedback the mixed pattern was associated with more task feedback, compared to mostly personal feedback in a blocked pattern the advanced integration of dashboards into teachers’ professional routines positively impacts feedback acknowledgments this work was supported by the nro doorbraak project awarded to inge molenaar (grant number: 405-15-823). the authors wish to thank the students of the radboud university for collecting the data, all the schools and teachers who participated in this study, and dr. j.a. houwman for proofreading the manuscript. references ballet, k., & kelchtermans, g. (2009). struggling with workload: primary teachers’ experience of intensification. teaching and teacher education, 25(8), 1150-1157. doi.10.1016/j.tate.2009.02.012 bennett, r. e. (2011). formative assessment: a critical review.assessment in education: principles, policy & practice, 18(1), 5-25. doi.10.1080/0969594x.2010.513678 butler, d. l., & winne, p. h. (1995). feedback and self-regulated learning: a theoretical synthesis. review of educational research, 65(3), 245-281. doi.10.3102/00346543065003245 de jager, b., jansen, m., & reezigt, g. (2005). the development of metacognition in primary school learning environments. school effectiveness and school improvement, 16(2), 179-196. doi. 10.1080/09243450500114181 gudmundsdottir, s., & shulman, l. (1987). pedagogical content knowledge in social studies. scandinavian journal of educational research, 31(2), 59-70. doi.10.1080/0031383870310201 hattie, j. & timperley, h. (2007). the power of feedback. review of educational research, 77(1), 81-112. doi.10.3102/003465430298487 hattie, j. (2012). visible learning for teachers: maximizing impact on learning. routledge. doi.10.4324/9780203181522 holstein, k., hong, g., tegene, m., mclaren, b. m., & aleven, v. (2018, march). the classroom as a dashboard: co-designing wearable cognitive augmentation for k-12 teachers. in proceedings of the 8th international conference on learning analytics and knowledge (pp. 79-88). doi.10.1145/3170358.3170377 holstein, k., mclaren, b. m., & aleven, v. (2017, march). intelligent tutors as teachers' aides: exploring teacher needs for real-time analytics in blended classrooms. in proceedings of the seventh international learning analytics & knowledge conference (pp. 257-266). acm. doi.10.1145/3027385.3027451 keuvelaar-van den bergh, l. (2013). teacher feedback during active learning: the development and evaluation of a professional development programme. klinkenberg, s., straatemeier, m., & van der maas, h. l. (2011). computer adaptive practice of maths ability using a new item response model for on the fly ability and difficulty estimation. computers & education, 57(2), 1813-1824. doi.10.1016/j.compedu.2011.02.003 knoop-van campen, c.a.n., wise, a.f., & molenaar, i. (subm.). the equalizing effect of teacher dashboards on feedback in a k-12 classroom. lacourse, f. (2011). an element of practical knowledge in education: professional routines. mcgill journal of education / revue des sciences de l'éducation de mcgill , 46 (1), 73–90. doi.10.7202/1005670ar leeuwen, a. van, janssen, j., erkens, g., & brekelmans, m. (2014). supporting teachers in guiding collaborating students: effects of learning analytics in cscl. computers & education, 79, 28-39. doi.10.1016/j.compedu.2014.07.007 martinez-maldonado, r., clayphan, a., yacef, k., & kay, j. (2014). mtfeedback: providing notifications to enhance teacher awareness of small group work in the classroom. ieee transactions on learning technologies, 8(2), 187-200. doi.10.1109/tlt.2014.2365027. molenaar, i., & knoop-van campen, c. a. n. (2016, april). learning analytics in practice: the effects of adaptive educational technology snappet on students' arithmetic skills. in proceedings of the sixth international conference on learning analytics & knowledge (pp. 538-539). doi.10.1145/2883851.2883892 molenaar, i., & knoop-van campen, c.a.n. (2017a, september). teacher dashboards in practice: usage and impact. in european conference on technology enhanced learning (pp. 125-138). springer, cham. doi.10.1007/978-3-319-66610-5_10 molenaar, i., & knoop-van campen, c.a.n. (2017b, september). how teachers differ in using dashboards: the classroom observation app. presented at the workshop “multimodal learning analytics across (physical and digital) spaces” at european conference on technology enhanced learning. springer, cham.. springer, cham. molenaar, i., & knoop-van campen, c.a.n. (2019). how teachers make dashboard information actionable. ieee transactions on learning technologies, 12(3), 347-355. doi.10.1109/tlt.2018.2851585. molenaar, i., & schaik, a. van (2016) a methodology to investigate classroom usage of educational technologies on tablets. in: aufenanger, s., bastian, j. (eds.) tablets in schule und unterricht. forschungsergebnisse zum einsatz digitaler medien , pp. 87–116. springer, wiesbaden. doi.10.1007/978-3-658-13809-7_5 roelofs, e., & sanders, p. (2007). towards a framework for assessing teacher competence. european journal of vocational training, 40(1), 123-139. shute, v. j. (2008). focus on formative feedback. review of educational research, 78(1), 153-189. doi.10.3102/0034654307313795 schwartz, r. m. (2005). decisions, decisions: responding to primary students during guided reading. the reading teacher, 58 (5), 436-443. doi.10.1598/rt.58.5.3 verbert, k., govaerts, s., duval, e., santos, j. l., assche, f., parra, g., & klerkx, j. (2014). learning dashboards: an overview and future research opportunities. personal and ubiquitous computing, 18(6), 1499-1514. doi. 10.1007/s00779-013-0751-2 verloop, n., van driel, j., & meijer, p. (2001). teacher knowledge and the knowledge base of teaching. international journal of educational research, 35(5), 441-461. doi.10.1016/s0883-0355(02)00003-4 voerman, l., meijer, p. c., korthagen, f. a., & simons, r. j. (2012). types and frequencies of feedback interventions in classroom interaction in secondary education. teaching and teacher education, 28 (8), 1107-1115. doi.10.1016/j.tate.2012.06.006 wood, d., bruner, j. s., & ross, g. (1976). the role of tutoring in problem solving. journal of child psychology and psychiatry, 17(2), 89-100. doi.10.1111/j.1469-7610.1976.tb00381.x van gasse publication frontline learning research vol.7 no. 2 (2019) 40 56 issn 2295-3159 the effect of formal team meetings on teachers’ informal data use interactions roos van gasse a auniversity of antwerp, belgium article received 2 january 2019 / revised 19 march / accepted 18 april / available online 8 may abstract in recent years, the emphasis on interaction in data use has grown because of its potential to support individual teachers. however, in practice, teachers do not appear to interact widely in their use of data, either formally or informally. to gain knowledge of how sustainable data use interactions can be facilitated, this study investigated how formal data use in teams of teachers affects the teachers’ informal interactive data use. a survey provided insight into 72 teachers’ perceptions of data use discussion, interpretation, diagnosis and action at formal team meetings. subsequently, social network analysis of seven teacher informal data use networks revealed that teachers with more positive perceptions about formal data use become more active in their informal data use network. within the problem diagnosis phase, this tendency is to generalize across the participating teams. the results of this study imply that, particularly to define problems and formulate actions based on pupil learning outcome data, it is necessary to ensure strong connections between teachers in formal groupings in order to affect their informal interactive behaviour. keywords: data use; informal interactions, formal groupings; collaboration info corresponding author: roos.vangasse@uantwerpen.be doi: 10.14786/flr.v7i2.443 1. introduction to date, teachers are increasingly stimulated to use data (e.g. student data) to learn about and improve their practice. this emphasis on data use in education originated from the belief that data use objectivises educational decisions and contributes to more effective (changes in) instructional practices. the idea is that the systematic analysis and interpretation of different types of data can lead to school improvement (campbell & levin, 2008; carlson, borman, & robinson, 2011). teachers often experience difficulties in the process of transforming data into knowledge and action (datnow & hubbard, 2016; hubbard, datnow, & pruyn, 2014; jimerson, 2014; wayman, midgley, & stingfield, 2007). therefore, the international literature has pointed to the important role of teacher interactions in data use. interactions have the potential to provide teachers struggling to use data appropriately with the necessary support to accomplish the complex translation of data into decisions and actions (bertrand & marsh, 2015). moreover, the belief has grown that data use interactions create an environment for teachers’ professional development (vanhoof & schildkamp, 2014). however, the current situation regarding interactive data use among teachers gives cause for pessimism, for two reasons. first, the frequency of interactions matters for teacher learning (penuel, sun, frank, & gallagher, 2012). yet, data use interactions are often limited in comparison with other forms of professional interaction, either in formally-established groups or on an informal basis (farley-ripple & buttram, 2015; hubers, moolenaar, schildkamp, daly, handelzalts, & pieters, 2017; keuning, van geel, visscher, fox, & moolenaar, 2016). second, interdependence among teachers facilitates learning. this implies that teachers interact from shared values and goals and that there is collective responsibility for pupils’ learning (horn & little, 2010; moolenaar, sleegers, & daly, 2012; stoll, bolam, mcmahon, wallace, & thomas, 2006). nevertheless, if data use interactions occur, teachers are not likely to share responsibility with colleagues in data use, with the result that brief exchanges of information take place rather than powerful learning activities (van gasse, vanlommel, vanhoof, & van petegem, 2016; 2017). the aforementioned issues mean that the value of learning outcomes from data use interactions must be put into question (van gasse et al., 2016). despite the great emphasis on teacher interactions in data use, there is a need for research to establish how sustainable and effective data use interactions are cultivated. an opportunity in this regard may lie in the interrelation between formal and informal interactions. after all, it has been known that teachers seek stability in terms of the number of colleagues that belong to their personal networks in schools, particularly when it comes to data use (van gasse, vanlommel, vanhoof, & van petegem, 2017; farley-ripple & buttram, 2015). therefore, knowing each other and working together in formal data use settings may be an important facilitator for teachers’ informal interactions. for example, research has shown that teachers involved in formal subgroups are more likely to interact with those colleagues on an informal basis (daly, moolenaar, bolivar, & burke, 2010; meredith, van den noortgate, struyve, gielen, & kyndt, 2017). however, what has remained underexplored up to now in the context of data use is whether it is simply the involvement in formal interactions that determines teachers’ informal interactions or whether what happens within those formal interactions can also contribute to teachers’ informal interactions. in other words, the question arises whether it is being familiar with colleagues that affects teachers’ informal data use interactions or the degree to which they feel that formal occasions facilitate the proper use of data. to fill this lacuna, the following research questions will guide this paper: 1. to what extent do teachers perceive that proper data use is accomplished at formal team meetings? 2. how do teachers’ perceptions of data use at formal team meetings affect their informal interaction-seeking behaviour with colleagues from the formally-constructed team? 2. conceptual framework this conceptual framework will first provide broader information and a theory of data use. subsequently, it will describe (the merits of) teacher interactions in the context of data use and outline the relationship between formal and informal interactions found in the literature. 2.1 data use and data data use is a way to manage processes within the school. the aim is to map school processes, to ensure that these processes are in line with school-wide goals and to use data to improve these processes (barrezeele, 2012; schildkamp & kuyper, 2010). therefore, many types of qualitative and quantitative data can be used (hulpia, valcke & verhaeghe, 2004; schildkamp & kuyper, 2010). data use is a somewhat simplistic linguistic merger of ‘data’ and ‘use’. effective data use is not only about ‘data’ and about ‘use’. it is a complex and sequential process in which data are transformed into information and knowledge (coburn & turner, 2011; marsh, 2012). therefore, data users need to run through different sub-processes to interrupt teachers’ tendency to jump from data to decisions (schildkamp et al., 2016). a lot of research distinguishes the phases of data discussion, analysis, interpretation and action (gummer & mandinach, 2015; marsh, 2012; schildkamp et al., 2016). however, teachers often struggle with the translation of data to classroom interventions (gummer & mandinach, 2015; datnow & hubbard, 2016). therefore, we explicitly insert a phase of problem diagnosis in our conceptualisation of data use. in this study, the sequence of data use is one of discussion, interpretation, diagnosis and action (verhaeghe et al., 2010). this means that data first needs to be read and discussed. subsequently, data must be interpreted correctly. next, potential causes and explanations are hypothesised and checked in the diagnosis phase. finally, teachers design and implement improvement actions (verhaeghe et al., 2010). the data use sequence appears straightforward in outline. nevertheless, the literature has repeatedly shown that in practice complexity arises because the sequence of activities is often interrupted or teachers return to previous phases (schildkamp et al., 2015; marsh & farrell, 2015). moreover, activities within the different phases of data use cannot be considered identical (schildkamp et al., 2016). it is essential to approach the concept of data use with sufficient precision and to take into account that different phases can imply differences in teacher behaviour. therefore, we will explicitly distinguish between data discussion, interpretation, diagnosis and action in this study. the early research on data use indicated that it is a difficult process for teachers. the sequence (discussion, interpretation, diagnosis and action) includes numerous potential pitfalls. for example, individual teachers become stuck in the interpretation phase of data use or have difficulty ascertaining where exactly pupils’ problems are located when analysing the data (datnow & hubbard, 2016; verhaeghe et al., 2010). therefore, more recently, an increasing emphasis has been laid on teacher interactions in the context of data use. the conviction has grown that teacher interactions are beneficial because they provide a supportive environment in which individual data use struggles can be overcome (bertrand & marsh, 2015; hubers et al., 2017). moreover, interaction in the context of data use has been identified as conducive to a professional learning environment for teachers (vanhoof & schildkamp, 2014). the following sections will further delineate the relationship between formal and informal interactions. thereafter, we will describe how social network theory will be used to explore this connection. 2.2 informal teacher interactions in data use informal teacher interactions are formed based on personal goals. informal interactions can occur very systematically or ad hoc, but always on the initiative of one (or more) teachers without the central and external creation of a common mission (blankenship & ruona, 2009). as a result, informal interactions may become more formal on the initiative of teachers themselves, but may equally remain unstructured and ad hoc. despite the potential of informal interactions for teacher learning (jones & dexter, 2014; kyndt, gijbels, grosemans, & donche, 2016) , there has been limited research on informal interactions in the area of data use. the few studies to report on such interactions have shown that they are fairly limited (farley-ripple & buttram, 2015; van gasse et al., 2017). in addition, the group of colleagues with whom teachers interact appears to remain relatively fixed. teachers’ consult a similar (but smaller) pool of colleagues for data use purposes, compared to for their regular professional activities (farley-ripple & buttram, 2015). this implies that teachers will not turn to specific colleagues with regard to data use (e.g. data use experts), but rather that they maintain a stable network for their different professional activities. this is also illustrated in their approach to the different activities within the data use sequence. across the discussion, interpretation, diagnosis and action stages, teachers involve the same colleagues; however, for the more complex phases (e.g. data use action), fewer colleagues are consulted and deeper interactions are established (van gasse et al., 2017). the facts that teachers seek stability in interactions with colleagues and that deeper interactions are not established with all colleagues imply that informal interactions may be of significant importance for sustainable and effective data use. therefore, further insights are needed into how these interactions can be facilitated and the role of formal interactions in this regard. 2.3 formal teacher interactions in data use contrary to informal interactions, formal interactions are formed in function of a specific organizational goal instead of individual goals. this means interactions for which a common mission to use data is created externally. generally, such interactions result in the construction of pre-structured work groups, in which team members are made responsible for the outcomes, or the division of labour is clearly explicated beforehand (blankenship & ruona, 2009). common examples of such formal interactions are data use interventions that include the creation of a team of educators working on specific school-related problems in order to learn how to use data (ciampa & gallagher, 2016; cosner, 2011; keuning et al., 2016; schildkamp et al., 2016). research has shown marked differences in how such teams work through the different phases of data use, in a sense that some teams reach deeper levels of inquiry than others (schildkamp et al., 2016). moreover, it is not the case that creating such formal group settings automatically leads to highly interactive groups; in fact, levels of interactions in formal data use groups remain limited (hubers et al., 2017; keuning et al., 2016). in addition, the distribution of knowledge from inside formal data use work groups to colleagues outside the teams appears to be scarce (hubers et al., 2017). thus, although formal groups have been widely implemented with the aim of facilitating interactions between teachers beyond those groups and facilitating school-wide data use, such formal interactions seem to fall short of this intention. given the teachers’ wish for stability in data use interactions and the limited number of colleagues with whom they undertake the most intense data use interactions (farley-ripple & buttram, 2015; van gasse et al., 2017), more knowledge is needed about the mechanisms mediating between formal and informal data use interactions. these insights are needed to better facilitate the connection between interventions using formal group settings and teachers’ informal interactive behaviour in data use. research into teacher interactions has already exposed some interconnections between formal and informal interactions. for example, research of meredith and colleagues (2017) showed that being connected to formal subunits can, to some extent, facilitate informal interactions. in the findings of their study, teachers who were connected to the same subunit appeared to interact more often with each other informally. similar findings were reported by penuel et al. (2010), who concluded that teachers’ being connected in a formal structure (by grade level) were more likely to interact with each other. other researchers have argued that formal groupings provide a specific focus for interaction and shape opportunities for interactions (spillane & kim, 2012; spillane, parise, & sherer, 2011). therefore, some formal groupings are considered to have more influence on teacher interactions than personal characteristics (spillane, hopkins, & sweet, 2015). however, the aforementioned studies pay insufficient attention to the impact of what happens within formal meetings on teachers’ informal interactive behaviour. the studies look at formal structures and group compositions to explain why informal interactions take place; yet, in the context of data use, we know that large differences in group processes may occur (schildkamp et al., 2016). moreover, even among teachers who are bounded to the same formal data use groupings, informal interactions can be scarce (van gasse et al., 2017). therefore, greater insight is needed into whether teachers perceive that proper discussion, interpretation, diagnosis and action of data takes place in formal team meetings as this might also explain their informal interactive behaviour within the same constellation of teachers. 3.method 3.1 research context the study was carried out in flanders, the dutch speaking part of belgium. in flanders, schools get autonomy over how they achieve the required educational standards (penninckx, vanhoof, & van petegem, 2011). the government does not impose central exams (oecd, 2014). instead, schools themselves are responsible for developing strategies to meet the flemish standards at the end of secondary education. therefore, the flemish government’s perspective on data use is quite improvement oriented. this is different to countries with a longer though more accountability oriented tradition in data use (e.g. the netherlands, united states, united kingdom). in practice, this implies that flemish schools and teachers often primarily rely on their own data sources (e.g. tests, assignments, observations or portfolios) for data use purposes. in this study, we will report on teachers’ use of pupil learning outcome data because this is an informative source of data for teachers to improve their practices and to evaluate whether or not pupils meet the flemish standards at the end of secondary education. pupil learning outcome data include cognitive outcomes (i.e. linguistic and arithmetic skills) as well as non-cognitive outcomes (i.e. attitudes, and artistic and physical education), and those data can be both quantitative (e.g. class tests) and qualitative (e.g. observations). this conceptualisation of ‘data’ is broader than often-used definitions which refer solely to cognitive output indicators (schildkamp et al., 2012). as such, this study contributes to an enriched conceptualisation of ‘pupil learning outcome data’ that includes both cognitive outcomes (e.g., linguistic and arithmetic skills) and non-cognitive outcomes (e.g., attitudes, and artistic and physical education). the study took place in the context of a project on the assessment of competences (d-pac.be). all ten schools involved in the project were asked to participate in this study. in each school, the target population were all teachers of the pupil group that participated in an assessment of writing competences in the aforementioned project, i.e. the fifth grade of an academic track in economics and languages (16to 17-year-olds). in flanders, these forms of teacher teams are temporary interdisciplinary groupings that are collectively responsible for pupils’ learning. two to three times during the school year, the teams are obliged to discuss the pupils’ learning outcomes in a formal team meeting. in the last team meeting of the year, team members deliberate as to whether or not pupils will successfully complete their year. this study will report both on teachers’ perceptions of data use at those formal team meetings and their informal interactive behaviour in data use within the same team. in order to answer the present research questions, quantitative data analysis has been combined with social network analysis. both types of data were collected in the same online survey. 3.2 social network theory and analysis in this study, interactions will be studied by means of social network analysis. this method draws on social network theory, which uses the position of actors within a network to determine their access to resources (e.g. colleagues’ knowledge and skills in a broad sense) (finnigan & daly, 2012). social network analysis approaches interactions with fine-grained information. it brings together information about both actors in an inter-action, which creates great depth of analysis. there are three general elements present in inter-actions. the first is interaction-seeking behaviour. for example, in a data use interaction, teacher a may ask teacher b for advice. in this case, teacher a is sending a connection (or a tie) to teacher b. this is what in social network analysis is called a sent tie (or outdegree). in reverse, teacher b may also ask the advice of teacher a (or send a connection to a). from teacher a’s perspective, this a received tie (or indegree). if teacher a and teacher b ask each other for advice, both teachers are sending and receiving ties from each other. social network analysis calls them reciprocated ties (borgatti, everett, & johnson, 2013). although the different characteristics of inter-actions are explained by means of advice ties, social network studies have used a range of topics for interaction purposes (e.g. friendship ties, information ties, general professional ties) (e.g. van gasse et al., 2017; daly et al., 2010; moolenaar et al., 2012). social network analysis provides a means of explaining the different types of connections between teachers. for example, perceptions of data use at formal team meetings can be used to explain different types of ties. given the present research questions, this study will explain teachers’ outdegree measures. this measure reflects the number of outgoing interactions or the extent to which they take the initiative in interaction with colleagues (borgatti, mehra, brass & labianca, 2009). in other words, it will provide insight into whether teachers become more active interactors in data use when feeling more positive about formal data use with the same colleagues. in the analysis section, the term sender effects will refer to the effect of teachers’ perceptions of formal data use on their informal interactive behaviour in terms of outdegree (sweet, 2016). for example, a positive sender effect would mean that the more positive teachers are about formal data use, the more likely they are to initiate informal interactions within the same network (i.e. to have a higher outdegree measure). 3.3 participants the data of three teams were excluded from this study because of a response rate lower than the 80% required in social network analysis. the other response rates are shown in table 1. for confidentiality purposes, the team names are fictive. response rates above 80% were reached in all teams, with maximum response rates (100%) in four out of seven teams. the high response rates imply that accurate conclusions can be drawn about the relationship between formal and informal interactions in teachers’ use of pupil learning outcome data. across the teams, 3048 data points provide sufficient statistical power to reveal some general tendencies. table 1: teams' response rates (social network analysis) apart from team mckinley (13 teachers) and team eppingswood (8 teachers), the teams consisted of eleven teachers. the teams were interdisciplinary, which means that teachers teach different subjects to the pupil group. all teachers had a master degree; 60 % were female, 40 % were male. further, a small majority of the participants taught the pupil group more than three course hours per week and the teaching experience of the participant group varied from less than five years to over thirty years. 3.4 instrument both teachers’ perceptions of their formal and informal data use interactions (i.e. the discussion, interpretation, diagnosis and action with regard to pupil learning outcome data) were measured by using an online survey. first, some general information was questioned, such as gender, level of educational attainment or amount of teaching time per week in the specific pupil group (fifth year track economics and languages). the second part of the survey included questions on the formal team meetings. teachers were asked to rate five statements concerning the extent to which they use pupil learning outcome data at the formal team meetings (e.g. ‘together with colleagues, i diagnose problems based on pupil learning outcome data during the formal team meetings’) on a scale from 1 (totally disagree) to 5 (totally agree). the cronbach’s alpha of 0.88 indicated good internal consistency of teachers’ perceptions regarding to the data use at these formal team meetings (see table 2). table 2. descriptive statistics of the formal interaction scale in order to measure teachers’ perceptions of their informal data use interactions, social network questions were included in the survey. for each of the data use phases (i.e. discuss, interpret, diagnose, take action), a social network question was included (e.g. ‘which of the following colleagues do you consult to discuss pupil learning outcome data?’). subsequently, all members of the teacher team were listed and participants indicated which of the listed colleagues they consult for data use discussion, interpretation, diagnosis and action apart from the formal team meetings of the team. 3.5 analyses to answer the first research question, we aggregated teachers’ item scores of the five items on formal data use interactions at team level using spss 22 software. subsequently, descriptive statistics (i.e. average and standard deviation) were calculated for each participating teacher team. as such, we are able to draw team-level conclusions about general perceptions regarding the use of pupil learning outcome data on formal occasions. this was necessary to build a point of reference for the informal interactive behaviour of teams that was analysed in light of the second research question. with regard to the second research question, we first calculated a scale score (i.e. average score) of the five items on formal data use interaction at teacher level. this was needed for being able to attribute relational differences to individual differences and subsequently to team differences. the relation between teachers’ scale scores on the formal interaction scale and their informal interaction seeking behaviour was tested using exponential random graph modelling (ergm). ergm enables researchers to analyse interaction patterns in social networks and to explain specific relationships. this type of analysis predicts the presence of particular relations in the network and, thus, it can be used to assess the predictive value of teachers’ perceptions of interactions on formal data use occasions for their informal behaviour in networks. in doing this, ergm takes into account global network structures as such, ergm accounts for the multilevel effect that occurs in using the level of relationships (within individuals within teams) as the unit of analysis. statnet’s r-based ergm package was used for the analyses (handcock, hunter, butts, goodreau, & morris, 2016). ergms are specified at team level. therefore, per team multiple ergms were specified; one for each phase of the data use cycle. each of those ergms included a single sender effect for teachers’ scores on the formal interactions scale. this means that the model investigated whether higher scores on the formal interactions scale affect the probability of sending relationships (i.e. consulting colleagues). this effect was investigated in the different data use networks of each team (i.e. discussion, interpretation, diagnosis and action). for each ergm, the model with the sender effect included was compared with the baseline model by means of the akaike information criterion (aic). this method was used to evaluate whether informal teacher interactions were better explained by teachers’ perceptions of formal interactions than by chance. to evaluate overall effects in the discussion, interpretation, diagnosis and action networks, a meta-analysis was conducted across the seven teams using the ‘metafor’ package in r. 4.results the structure of this section is aligned with the research questions. therefore, we will first describe teachers’ perceptions regarding data use at formal team meetings. then we will provide insight into the descriptive statistics with regard to teachers informal data use interactions in light of a better understanding of the subsequent analyses. finally, we will present the findings on interrelationships between teachers’ perceptions of data use at formal team meetings and their informal data use interactions. 4.1 formal data use interactions in teacher teams table 3 provides an overview of the descriptive statistics of this study. the statistics with regard to ‘formal data use’ include the aggregated team scores of teachers’ perceptions of the use of pupil learning outcome data (i.e. discussion, interpretation, diagnosis and action) at the obliged formal team meetings (cf. context description). with regard to teachers’ perceptions of formal data use in the team, we find positive to strongly positive perceptions overall. in all teams, teachers report that pupil learning outcome data are discussed and interpreted at formal team meetings, that problems are diagnosed and that appropriate actions are formulated to improve pupils’ learning. across the teams, averages range from moderately positive perceptions (average team northvale = 3.80) to extremely positive perceptions (average team melrose = 4.94). however, in some teams teachers’ perceptions are more disparate than in others. for example, the standard deviation of team melrose indicates that all teachers answered the questions on formal data use similarly (sd = 0.11), whereas the same measure in teams easton and northvale indicates a significantly larger variation between teachers (sd = 0.82 and sd = 1.04 respectively). this means that some teachers were a lot more positive than others about data use on formal occasions in those teams. the ergm analysis will reveal whether such different perceptions also affect teachers’ informal data use interactions. 4.2 informal data use in teacher teams in order to better understand the interrelation between formal and informal data use interactions, table 3 also shows the average number of informal interactions (i.e. interactions independent from the formal team meetings) within the discussion, interpretation, diagnosis and action phase per teacher teams. this is represented by the ‘average degree’ measure, which refers to the number of outgoing relations (or the extent to which teachers consult colleagues), aggregated at team level. we find limited informal activity within the teams across the data use sequence (discussion, interpretation, diagnosis and action). table 3 shows that the maximum average degree, for teams of 11 teachers, occurs in team riverbank’s discussion network (average degree = 4.45). this means that teachers in team riverbank are, on average, connected to 4–5 teachers (out of 10) for data discussion. team riverbank is the most active network of 11 teachers with regard to informal data use interactions. the smaller team eppingswood is interacting to a slightly greater extent compared to the other teams. for example, the average degree of 3.63 in the action network implies that teachers are interacting informally with 3–4 colleagues (out of 7). furthermore, higher average degree numbers are found in team colby, but only for the data discussion and interpretation networks; the established relations in the diagnosis and action networks of team colby are in line with all other teams. therefore, in general, informal data use interactions are scarce, with teachers connected to approximately 2 other teachers for informal data use discussion, interpretation, diagnosis and action. table 3 descriptive statistics per team 4.2 perceptions of formal data use and informal interactive behavior table 4 shows the results of the ergm analyses of the impact of teachers’ perceptions of formal data use on their informal interaction-seeking behaviour in data use. these effects are represented in the ‘sender effects’ columns. the different ergms provide an overview of the effects of teachers’ perceptions on formal interactions across the different phases of the data use sequence. as such, fine-grained conclusions on the effects of perceived formal interactions can be drawn. a first finding of the ergm analyses is that there are some teams in which there are no significant effects from teachers’ perceptions of formal data use on informal interaction-seeking data use behaviour. more specifically, teachers’ informal data use interactions in teams riverbank and melrose cannot be related to their perceptions concerning formal data use in the same team constellations. in these teams, the teachers’ informal data use interactions can be explained equally well by chance. it is remarkable that teams riverbank and melrose in particular show no significant effects from teachers’ perceptions of formal interactions on their interaction-seeking behaviour. as shown in table 3, these are two of the four highest-scoring teams on the formal data use scale, which implies that teachers in those teams are particularly positive about the discussion and interpretation of data at formal team meetings, and about diagnosing problems and designing improvement actions based on data. additionally, the aforementioned teams have smaller standard deviations in our sample, which indicates that the teachers gave similar answers on the formal data use scale. table 4 sender effects of formal data use per team in all other teams, the ergm analyses show at least one effect. in team colby, we find a negative sender effect of formal data use on teachers’ informal interactive behaviour in the interpretation of pupil learning outcome data. in this team, teachers’ reporting higher scores on the formal data use scale are less likely to seek interaction with colleagues for the interpretation of pupil learning outcome data on an informal basis. in team mckinley a significant sender effect is again found for data use interpretation. however, in this team the effect is positive. this implies that, in team mckinley, teachers who are more positive about the discussion and interpretation of pupil learning outcome data, and about diagnosing problems and defining improvement actions are more likely to consult their teammates for informal data interpretation. given this disparity of effects in the interpretation networks of these two teams, no significant generic effects across the teams were found in the meta-analysis. together with the absence of effects in any of the teams’ discussion networks, this implies that no generic effects can be detected in the least complex phases of data use (i.e. discussion and interpretation). this implies that teachers’ perceptions of formal data use explain informal interactions for data discussion and interpretation no better than chance. however, the ergm analyses do show some effects of teachers’ perceptions of formal data use on their informal interaction-seeking behaviour in the more complex phases of data use. moreover, in each of the networks, these effects are positive. in teams mckinley, easton, eppingswood and northvale, teachers reporting higher scores on the formal data use scale are more likely to seek data use interactions with colleagues for diagnosing problems. this means that the more strongly teachers believe that pupil learning outcome data are discussed and interpreted effectively at formal team meetings, and that subsequently diagnosis is carried out and actions are defined, the more likely they are to consult (some of those) colleagues for informal data discussion. with regard to the definition of improvement actions, we find similar effects in teams mckinley, eppingswood and northvale. according to the network data, teachers with higher scores on the formal data use scale hare more likely to consult colleagues on data use actions. in other words, the teachers’ perceptions of data use at formal team meetings affect their informal interaction-seeking behaviour for the definition and implementation of improvement actions for pupils. thus, teachers’ perceptions of data use at formal team meetings seem to matter more for their informal interaction-seeking behaviour in the more complex phases of data use (i.e. discussion and action). the meta-analysis shows that these positive effects can be generalised across the teams for diagnosing problems from pupil learning outcome data, though not for defining and implementing improvement actions. 5. discussion and conclusion in data use, or the discussion and interpretation of data, the diagnosis of problems and the definition and implementation of improvement actions, researchers put a great emphasis on teacher interactions. peer interactions can provide teachers with the support necessary to acquire the complex knowledge and skills needed to transform data into meaningful decisions and actions (hubbard et al., 2014; jimerson, 2014; wayman et al., 2007). given the context, in which changes in teaching practices sometimes require prompt action, informal data use interactions may be of particular importance for teachers. nevertheless, the literature has shown that their informal data use interactions remain limited (farley-ripple & buttram, 2015; van gasse et al., 2017). therefore, it is crucial to understand how those informal data use interactions can be influenced. an important factor may be engagement in formal data use interactions. research has shown that being active in formal structures affects informal interactions (daly et al., 2010; meredith et al., 2017). however, what remained under-researched was how teachers’ perceptions of (data use) activities within formal team meetings influence their informal interactive behaviour. the current study approached this lacuna using a combination of survey questions and social network analysis. as such, insight was provided into teachers’ perceptions of the use of pupil learning outcome data at formal team meetings and the influence of these perceptions on their informal interaction-seeking behaviour. the descriptive statistics showed that, overall, teachers were positive about the discussion and interpretation of pupil learning outcome data at formal team meetings, as well as the diagnosis of problems and definition and implementation of improvement actions. there was notable variation in teachers’ perceptions across the teams, and some variation within them. the findings also showed limited informal data use interaction. on average, teachers were not connected to many of the colleagues with whom data use was carried out at formal team meetings. nevertheless, drawing on the ergm analyses, teachers’ perceptions of formal data use appeared be related to their informal interaction-seeking behaviour in the more complex phases of data use (i.e. diagnosis and action). significant positive effects of teachers’ perceptions of formal data use were found on the extent to which teachers consult colleagues for diagnosing problems and defining and implementing improvement actions using pupil learning outcome data. however, these effects could only be generalised across the teams for the problem diagnosis phase. two findings were particularly remarkable in the results of the ergm analyses. the first is that, in some teams, no significant effects were found. these teams (riverbank and melrose) had high average scores on the formal interactions scale with small standard deviations. this implies that teachers in those teams took a similar, positive stance towards the use of pupil learning outcomes at formal team meetings, making differences in informal behaviour more difficult to explain by this factor. the second remarkable finding, which may represent a step forward in the field of data use interaction research, is that the impact of teachers’ perceptions of formal data use only appeared to have an effect on the more complex phases of informal data use (i.e. diagnosis and action). apparently, a more or less positive stance towards the team’s formal data use is of less importance for teachers’ informal interaction-seeking at the data discussion and interpretation stages. yet for informal interactions at the problem-diagnosis phase, teachers’ perceptions regarding data use at formal team meetings do matter. an explanation may lie in the fact that for data discussion and interpretation, teachers lay slightly less weight on the specific colleagues involved in their interactions. as prior research has shown, teachers use interactions in these phases to build a frame of reference (van gasse et al., 2016). therefore, more colleagues may be suitable to interact with in these phases, and thus perceptions of what is happening at formal team meetings do not seem to play an important role. this is different in the phase of problem diagnosis and, to some extent, the phase of action. within these phases, more specific knowledge is needed. in this regard, the use of pupil learning outcome data can to some extent serve as an example of problem diagnosis based on data. positive perceptions of data use at formal team meetings can provide teachers with evidence that interactions in these phases may add value to the process of data use. additionally, they may develop a sense of who, among their colleagues, would make interesting partners for diagnosing problems in data use. as a result, the impact of teachers’ perceptions of formal interactions on their informal interactive behaviour mainly manifests in the phases where teachers often struggle most (i.e. diagnosis and action). this implies that, when schools and teachers aim to arrive at fruitful data use, a pit of the matter may lie in facilitating teachers in thorough interactions to diagnose problems and formulate actions based on pupil learning outcome data at formal team meetings. in line with previous research on data use interactions, this study showed limited teacher interactions (farley-ripple & buttram, 2015; hubers et al., 2017; keuning et al., 2016). previous research has also shown that teachers seek stability in their interactive data use behaviour. for example, teachers tend to interact with similar colleagues for data use and for their regular professional interactions (farley-ripple & buttram, 2015). furthermore, within the data use sequence (discussion, interpretation, diagnosis and action), teachers tend to select and retain a stable group of interlocutors (van gasse et al., 2017). in this regard, it is remarkable that this study reported a relatively low number of interactions between teachers. the colleagues involved in the social network checklist were all connected via formal data use occasions, thus informal interactions with these colleagues would have contributed to the stability of the teachers’ personal network. furthermore, research has repeatedly shown that data use cannot be considered a straightforward process because of the different sub-phases (gummer & mandinach, 2015; schildkamp et al., 2016). each of these phases require different knowledge and skills of teachers, and to some extent ‘a shift’ can be found between interpreting data and diagnosing problems and then taking actions in the knowledge that is used (gummer & mandinach, 2015). in teachers’ interactive behaviour, such a ‘shift’ has appreciable effects (van gasse et al., 2017). therefore, the finding that teachers’ informal interaction-seeking behaviour may be influenced differently according to the data use phase, confirms and broadens this prior knowledge. in addition to contributing to the field of data use research, this study also broadens understanding of teachers’ interaction-seeking behaviour in general. this study is an extension of earlier work that related formal interactions to informal interaction-seeking behaviour (e.g. meredith et al., 2017; spillane et al., 2015). in those studies, formal groupings were related to interaction-seeking behaviour. the current study, however, adds to this knowledge in two ways. first of all, it is not only the involvement in formal groupings that affect teachers’ informal interaction-seeking behaviour. this study shows the importance of what happens within these formal interactions. thus, next to just being related by means of formal team meetings (e.g. daly et al., 2010), teachers might only interact with their team members on an informal basis when they are positive about what happens at the formal team meetings. the second aspect we learn on interactions in teacher networks is that the type of task may matter for the influence of formal interactions on informal behaviour. in previous work that shed light on the interconnection between formal and informal interactions, this aspect stayed somewhat under surface (e.g. daly et al., 2010; spillane). this study makes clear that this relation might be affected by teachers’ task, because effects were only found when the complexity of the task in front increases. the social network approach taken in this study proved valuable for generating additional insights into how formal data use in team meetings can influence teachers’ informal data use interactions. however, there are some limitations to this study. the first is that the sample size was limited. social network analysis requires intensive data collection because of the high response rates (over 80%) required. this because missing data can have a huge impact on the results. suitable sample sizes are therefore difficult to obtain. however, for generalisable conclusions across the teams, the specificity of each team plays a role. for example, the fact that in teams riverbank and melrose no significant effects were found impacted on the way the findings could be generalised across the teams. this could have been resolved with a larger sample. therefore, to generalise the findings of this study, larger-scale network research will be required. a second limitation of this study relates to the team context that was used. to define the boundaries of teams, we used the criterion of teaching a specific pupil group, which meant that the teachers in the formal grouping taught different subject areas. it is important to note that some teachers may feel closer connections to other formal groupings within schools, such as subject-related teams. therefore, the present results are strongly context-depending based on the geographical situatedness in flanders, but also based on the team boundaries. therefore, choosing another formal grouping in future research may have an impact on the similarity of the findings to those in this study. the last limitation we want to emphasize is the limitations of using one survey-instrument. therefore, we could not fully capture why we arrived at the present research results. this study has generated some important implications for future research and practice. the first relates to the history of the formal team. the type of teams in this study might change about every school year. therefore, the results of this type of analyses might be slightly different when focusing on formal teams with a longer history together (e.g. grade-level teams in the work of, inter alia, meredith et al. (2017) or spillane et al. (2015)). replications of this study in other team contexts is necessary to fully understand how bounded the present results are to the current research context. next, further studies might investigate the reasons that teachers are more interested in interacting with some colleagues rather than others in data use diagnosis and action. if, for example, we assume that teachers learn about each other’s strengths and weaknesses through formal data use interactions, do teachers seek out specific (and similar) colleagues for data use diagnosis and action on that basis? thus, further research into whether, and why, specific teachers become more important in the complex phases of data use is essential for theory and practice. certain information on general characteristics (e.g. age, experience) might be helpful in this regard. in addition, the fact that teachers’ perceptions of formal data use affect their informal interactions in the diagnosis (and action) phase warrants further exploration. more insight is needed into what exactly happens in formal data use interactions, how schools and teacher teams differ in those formal interactions and when and how those interactions affect teachers’ informal interactive behaviour. opportunities in this regard can be found in combinations of social network analysis and qualitative data sources (e.g. observations or in-depth interviews) to further deepen the social network analyses of this study. for the field of data use, this study brings about implications regarding the sustainability of the many interventions that are used to promote data use among practitioners (e.g. ciampa & gallagher, 2016; cosner, 2011; marsh, 2012; schildkamp et al., 2016). such interventions often introduce (new and temporary) formal teams to support schools in the use of data. this study shows that ‘getting to know each other’ is not sufficient for the sustainability of the data use interactions during the interventions. such interventions might only succeed and grow into professional learning communities when sufficient attention is paid to how meaningful the data use activities are for the different team members. therefore, it is important to ensure strong connections between teachers in formal groupings in order to encourage teachers’ informal interactive behaviour. keypoints teachers are generally positive about the extent to which they discuss, interpret, diagnose and take action upon pupil learning outcome data at formal team meetings. teachers generally interact rarely on an informal basis on pupil learning outcome data. teachers who are more positive about their formal data use, interact to a greater extent for problem diagnosis in the context of the use of pupil learning outcome data. strong connections between teachers in formal groupings are essential to encourage teachers’ informal interactive behaviour. references blankenship, s. s., & ruona, w. e. a. (2009). exploring knowledge sharing in social structures: potential contributions to an overall knowledge management strategy. advances in developing human resources, 11(3), 290-306. doi: 10.1177/1523422309338578 bertrand, m., & marsh, j. a. (2015). teachers' sensemaking of data and implications for equity. american educational research journal, 52 (5), 861-893. doi: 10.3102/0002831215599251 borgatti, s. p., everett, m. g., & johnson, j. c. (2013). analyzing social networks. los angeles: sage. borgatti, s. p., mehra, a., brass, d. j., & labianca, g. (2009). network analysis in the social sciences. science, 323, 892-895. campbell, c., & levin, b. (2008). using data to support educational improvement. educational assessment, evaluation and accountability, 21(1), 47-65. doi: 10.1007/s11092-008-9063-x ciampa, k., & gallagher, t. l. (2016). teacher collaborative inquiry in the context of literacy education: examining the effects on teacher self-efficacy, instructional and assessment practices. teachers and teaching, 22(7), 858-878. doi: 10.1080/13540602.2016.1185821 coburn, c. e., & turner, e. o. (2011). research on data use: a framework and analysis. measurement: interdisciplinary research & perspective, 9(4), 173-206. doi: 10.1080/15366367.2011.626729 cosner, s. (2011). supporting the initiation and early development of evicence-based grade-level collaboration in urban elementary schools: key roles and strategies of principals and literacy coordinators. urban education, 46(4), 786-827. doi: 10.1177/0042085911399932 daly, a., moolenaar, n. m., bolivar, j. m., & burke, p. (2011). relationships in reform: the role of teachers’ social networks. journal of educational administration, 48(3), 359-391. doi: 10.1108/09578231011041062 datnow, a., & hubbard, l. (2016). teacher capacity for and beliefs about data-driven decision making: a literature review of international research. journal of educational change, 17(1), 7-28. doi: 10.1007/s10833-015-9264-2 farley-ripple, e. n., & buttram, j. l. (2015). the development of capacity for data use: the role of teacher networks in an elementary school. teachers college record, 117(4), 1-34. finnigan, k. s., & daly, a. j. (2012). mind the gap: organizational learning and improvement in an underperforming urban system. american journal of education, 119, 41-71. gummer, e. s., & mandinach, e. b. (2015). building a conceptual framework for data literacy. teachers college record, 117(4), 1-22. handcock, m. s., hunter, d. r., butts, c. t., goodreau, s. m., & morris, m. (2016). statnet: software tools for the statistical modeling of network data (version 2016.9). horn, i. s., & little, j. w. (2010). attending to problems of practice: routines and resources for professional learning in teachers’ workplace interactions. american educational research journal, 47(1), 181-217. doi: 10.3102/0002831209345158. doi: 10.3102/0002831209345158 hubbard, l., datnow, a., & pruyn, l. (2014). multiple initiatives, multiple challenges: the promise and pitfalls of implementing data. studies in educational evaluation, 42, 54-62. doi: 10.1016/j.stueduc.2013.10.003 hubers, m. d., moolenaar, n. m., schildkamp, k., daly, a., handelzats, a., & pieters, j. m. (2017). share and succeed: the development of knowledge sharing and brokerage in data teams’ network structures. research papers in education, 1-23. doi: 10.1080/02671522.2017.1286682 jimerson, j. b. (2014). thinking about data: exploring the development of mental models for data use among teachers and school leaders. studies in educational evaluation, 42(2014), 5-14. doi: 10.1016/j.stueduc.2013.10.010 jones, w. m. & dexter, s. (2014). how teachers learn: the roles of formal, informal, and independent learning. educational technology research and development, 62(3), 367-384. doi: 10.1007/s11423-014-9337-6 keuning, t., van geel, m., visscher, a., fox, j.p., & moolenaar, n.m. (2016). transformation of schools' social networks during a data-based decision making reform. teachers college record, 118(9), 1-33. kyndt, e., gijbels, d., grosemans, i., & donche, v. (2016). teachers’ everyday professional development: mapping informal learning activities, antecedents, and learning outcomes. review of educational research, 86(4), 1111-1150. doi: 10.3102/0034654315627864 marsh, j. a. (2012). interventions promoting educators’ use of data: research insights and gaps. teachers’ college record, 114, 1-48. marsh, j. a., & farrell, c. c. (2015). how leaders can support teachers with data-driven decision making: a framework for understanding capacity building. educational management administration & leadership, 43(2), 269-289. doi: 10.1177/1741143214537229 meredith c., van den noortgate w., struyve c., gielen s., kyndt e. (2017). information seeking in secondary schools: a multilevel network approach. social networks, 50, 35-45. doi: 10.1016/j.socnet.2017.03.006 moolenaar, n. m., sleegers, p. j. c., & daly, a. j. (2012). teaming up: linking collaboration networks, collective efficacy and student achievement. teaching and teacher education, 28, 251-262. doi: 10.1016/j.tate.2011.10.001 oecd (2014). talis 2013 results: an international perspective on teaching and learning , talis, oecd publishing, http://dx.doi.org/10.1787/9789264196261-en. penninckx, m., vanhoof, j., & van petegem, p. (2011). evaluatie in het vlaamse onderwijs. beleid en praktijk van leerling tot overheid . [evaluation in flemish education. policy and practice from student to government] antwerpen-apeldoorn: garant. penuel, w. r., riel, m., joshi, a., pearlman, l., kim, c. m., & frank, k. a. (2010). the alignment of the informal and formal organizational supports for reform: implications for improving teaching in schools. educational administration quarterly, 46(1), 57-95. doi: 10.1177/1094670509353180 penuel, w. r., sun, m., frank, k. a. & gallagher, h. a. (2012). using social network analysis to study how collegial interactions can augment teacher learning from external professional development. american journal of education, 119(1), 103-136. schildkamp, k., & kuiper, w. (2010). data-informed curriculum reform: which data, what purposes, and promoting and hindering factors. teaching and teacher education, 26(3), 482-496. doi: 10.1016/j.tate.2009.06.007 schildkamp, k., poortman, c. l., & handelzalts, a. (2016). data teams for school improvement. school effectiveness and school improvement, 27(2), 228-254. doi: 10.1080/09243453.2015.1056192 schildkamp, k., rekers-mombarg, l. t. m., & harms, t. j. (2012). student group differences in examination results and utilization for policy and school development. school effectiveness and school improvement, 23(2), 229-255. doi: 10.1080/09243453.2011.652123 spillane, j. p., hopkins, m., & sweet, t. (2015). intra-and inter-school interactions about instruction: exploring the conditions for social capital development. american journal of education, 122(1), 71 -110. spillane, j. p., & kim, c. m. (2012). an exploratory analysis of formal school leaders’positioning in instructional advice and information networks in elementary schools. american journal of education, 119(1), 73-102. doi: 10.1086/667755 spillane, j. p., parise, l. m., & sherer, j. z. (2011). organizational routines as coupling mechanisms policy, school administration, and the technical core. american educational research journal, 48(3), 586-¬619. stoll, l., bolam, r., mcmahon, a., wallace, m., & thomas, s. (2006). professional learning communities: a review of the literature. journal of educational change, 7(4), 221-258. doi: 10.1007/s10833-006-0001-8 sweet, t. m. (2016). social network methods for the educational and psychological sciences. educational psychologist, 51(3-4), 381-394. doi: 10.1080/00461520.2016.1208093 vanhoof, j., & schildkamp, k. (2014). from ‘professional development for data use’ to ‘data use for professional development’. studies in educational evaluation, 42, 1-4. doi: 10.1016/j.stueduc.2014.05.001 van gasse, r., vanlommel, k., vanhoof, j. and van petegem, p. (2016). teacher collaboration on the use of pupil learning outcome data: a rich environment for professional learning? teaching and teacher education, 60, 387-397. doi: 10.1016/j.tate.2016.07.004 van gasse, r., vanlommel, k., vanhoof, j. and van petegem, p. (2017). unravelling data use in teacher teams: how network patterns and interactive learning activities change across different data use phases. teaching and teacher education, 67, 550-560. doi: 10.1016/j.tate.2017.08.002 verhaeghe, g., vanhoof, j., valcke, m., & van petegem, p. (2010). using school performance feedback: perceptions of primary school principals. school effectiveness and school improvement, 21(2), 167-188. doi: 10.1080/09243450903396005 wayman, j. c., midgley, s., & stringfield, s. (2007). leadership for data-based decision making: collaborative educator teams. in a. b. danzig, k. m. borman, b. a. w. jones & w. f. wright (eds.), learner-centered leadership: research, policy and practice (pp. 189-205). new jersey, usa: lawrence erlbaum associates. codepen fryer et nakao frontline learning research vol.8 no. 3 special issue (2020) 10 25 issn 2295-3159 the future of survey self-report: an experiment contrasting likert, vas, slide, and swipe touch interfaces luke k. fryera, kaori nakaob afaculty of education, the university of hong kong, hong kong bseinan gakuin university, fukuoka, japan article received 1 june 2019 / revised 4 december/ accepted 6 december / available online 30 march abstract self-report is a fundamental research tool for the social sciences. despite quantitative surveys being the workhorses of the self-report stable, few researchers question their format—often blindly using some form of labelled categorical scale (likert-type). this study presents a brief review of the current literature examining the efficacy of survey formats, addressing longstanding paper-based concerns and more recent issues raised by computerand mobile-based surveys. an experiment comparing four survey formats on touch-based devices was conducted. differences in means, predictive validity, time to complete and centrality were compared. a range of preliminary findings emphasise the similarities and striking differences between these self-report formats. key conclusions include: a) that the two continuous interfaces (slide & swipe) yielded the most robust data for predictive modelling; b) that future research with touch self-report interfaces can set aside the vas format; c) that researchers seeking to improve on likert-type formats need to focus on user interfaces that are quick/simple to use. implications and future directions for research in this area are discussed. keywords: : likert; vas; slide; self-report; response format; experimental design; mobile; touch interface info corresponding author: email: fryer@hku.hk doi: https://doi.org/10.14786/flr.v8i3.501 1. introduction the present special issue (fryer & dinsmore, 2020) has planted its flag in an unpopular or extremely popular—depending on your perspective—area of research. unpopular because of the nature of self-reported data: if it is qualitative, it lacks external validity, and if it is quantitative, it is ordinal at best. unpopular because it is, after all, just intra-psychic “stuff”, only loosely tied to the observed, interval/ratio construct gold standards (i.e., our own version of physics envy; see howell, et al., 2014). popular because very few researchers in the social sciences can avoid self-report in some form or another. for these reasons, and the fact that year on year humankind collects and analyses more self-report data than ever before, self-report data, and how it is collected, deserves more of our attention. although there are many, many means of collecting these self-reports, surveys/questionnaires are the most ubiquitous. despite the fact that the two most popular formats for surveys have been around for nearly a century (i.e., visual analogue scale vas1, hayes. & patterson, 1921; likert, likert, 1932), very little has been done to improve on them. the scant existing research comparing them has often concluded with a statement equivalent to “same difference”. only during the past two decades has the ground begun to shift under likert and vas formats. computers made slider formats possible, vas easier to implement and, survey data in all formats far easier to obtain. mobile devices have led to a natural expansion in the amount of surveys, but little actual development in how they are conducted. towards this development the current study presents longstanding issues (commonly and rarely addressed) alongside newer factors made prominent by computer and mobile survey interfaces. this short review is complimented by an experimental study comparing labelled categorical scale (lcs; a likert format with no center/neutral point), vas, slider (a sliding bar with labels and numbers) and a new format/interface (swipe; an adaptation/extension of the slide format) for collecting quantitative self-reported, micro-analytic data regarding students’ interest in classroom tasks. 2. background 2.1 the criticisms and critical roles of self-report there are many means of collecting self-reported information, but the most common means of self-report across all fields of human sciences are surveys measuring agreement to a set of statements across a numerical scale of some type (see durik & jenkins, 2020). being the most common type of self-report, surveys also receive the most criticism. these censures generally focus on two critical weaknesses inherent in survey data. the first is the often ordinal (or at least not technically interval) nature of the data itself. the second concern has two related parts, the first is that it is latent and therefore invisible to the senses, the related second part is the data's often tenuous (and generally indirect) connection to the observed world. nearly every researcher working with survey data has received a review of their manuscript pointing to one or both of these concerns as a limitation – if not as a reason for rejection. despite these acknowledged weaknesses, self-reported data are often the only or most direct means (at a large scale) of getting at human psychology. the obvious areas it is critical in assessing are intra-psychic aspects like beliefs, motivations, and emotions. less obvious, but an equally important area where self-reports are essential tools, are processes which are partially evident to the observer, but like an iceberg are mostly submerged: i.e., metacognitive and cognitive strategies. in addition to the broad concerns regarding the fact that these data are "just self-reported", there are a host of other issues less often discussed, and often unresolved, with quantitative survey data. the current study presents a brief review of some of these issues ranging from those that are (a) longstanding and commonly addressed, to (b) longstanding less often discussed, and finally (c) modern issues specific to computer and mobile (touch) interfaces. following this brief "highlight" review supported by recent research from a range of domains, a short experimental study examining four touch interfaces, with four self-report formats, for collecting survey data through mobile phones immediately after classroom experiences will be presented. discussion will seek to tie the review and experiment together, while lighting the way for more understanding, research and general development in this critical, but often unquestioned area of research methods. 2.2 longstanding issues with survey research before engaging with less often addressed issues with survey data, two important concerns commonly addressed through design and analyses should be noted. the first is the "less than interval" nature of survey data (i.e., it might be continuous but who knows what "it" is). this problem is generally addressed along with construct validity and reliability by the use of multiple items and either meanor, preferably, latent-variable analysis. latent-variable analysis is preferred for a range of reasons, of which measurement error is most commonly cited. algorithms such as those natively used by latent software packages like mplus (muthén & muthén, 1998-2015) are purported to ameliorate the stepwise nature of ordinal data, smoothing the distribution that classical statistics relies upon. reliability is supported by scales utilising items with similar content and reliability that can be assessed at a latent level (raykov, 2009), offering flexibility to latent modelling research. 2.3 longstanding less often discussed issues with survey research some of the longstanding, but often left unspoken, issues with survey data include central tendency, ceiling effects, number of appropriate categories, influence of proximal items, and self-report agreement vs. magnitude. central tendencies generally occur when survey respondents over subscribe to middling amounts of agreement and can be related to the use of a non-committal category (foddy, 1994). while central tendency has long been seen as bias, recent bayesian analysis suggests it might actually be a reflection of the probability of surveyed choice (douven, 2018). this is still an unresolved issue and many researchers will no doubt continue to blame scale midpoints as the source of this problem. likert format surveys (voutilainen, et al., 2016) and the survey statements themselves (austin & brunner, 2003) have been linked to ceiling effects. ceiling effects are when a large proportion of survey respondents report the highest possible scale value. like central tendency biases, ceiling effects can affect the normality of data and result in type i errors (austin & brunner, 2003). the number of appropriate categories in survey report formats is one of those issues that all researchers have to face when designing instruments and often results in a best guess. linked with concerns regarding central tendency, these questions also focus on odd vs. even numbers of categories (e.g., adelson & mccoach, 2010). the last issue is that of the difference between agreement with survey labels and the magnitude of that agreement (berger & alwitt, 1996). a two-step approach, with a likert response format followed by a cumulative scale from not very strong to very strong has been suggested as a mechanism for assessing both aspects of respondents' experience (albaum, 1997). while this pairing of self-report has presented robust predictive strength for related variables, this line of research has not been consistently pursued (see durik & jenkins, 2020). 2.4 new issues with survey research four relatively new issues that computer and now mobile interfaces are making central are (a) the use and number of labels and/or ticks on a slider or vas report scale line (no longer focused on explicitly stated categories), (b) precision in selecting the level of self-report, (c) speed in selection, (d) bias due to everything from age to education, and (e) relative non-response to different scale formats. for most of these issues there is only a budding body of research to draw upon. matejka, et al. (2016) is to our knowledge the only in-depth study testing the effect of the number of ticks (on a slide line) on self-report precision and speed. this study indicated that with regard to precision that 11 ticks is superior to five. this study also supported the use of dynamic feedback (a running quantitative score above the moveable slide marker). this addition enhances precision but has a detrimental effect on the speed of self-report. this study also pointed to the benefit of banded coloring of the slider line to signify increments as being superior to ticks alone. bias is a complex area to research and has not to our knowledge been properly investigated with studies supported by experimental design. survey studies have noted apparent biases supporting labelled categorical scale (lcs; i.e., likert-type) interfaces over slide and vas interfaces (voutilainen et al., 2016). this research attributes the benefits of lcs to age (i.e., easier for older and younger respondents) and/or education (i.e., easier for respondents with less education). their very specific supposition regarding bias reflects broad support for lcs over other continuous survey interface formats. 2.5 four formats for self-reporting agreement: lcs, vas, slide, and swipe a considerable number of studies, in a wide range of domains have assessed the relative usefulness of different survey interfaces. the majority have focused specifically on lcs and vas, which have been the predominant self-report formats. on comparing lcs and vas, most studies conclude that they are highly correlated and present similar overall distributions of data (bolognese et al 1990; reed, et al., 2017; vickers, 1999). if we include ease of administration, these studies generally support the use of lcs over other formats (i.e., generally vas). a smaller number of studies comparing lcs and vas have cited similar consistencies between the two-survey format but fallen on the vas side of the fence. these studies often cite the interval nature of the vas data relative to the ordinal nature of lcs data (bishop & herron, 2015). some of these studies also note vas’ robustness to ceiling effects and, in some cases, shorter time to complete when compared to lcs (couper, et al., 2006; voutilainen et al., 2016). lower standard deviation for vas vs. lcs has been reported, but has been difficult to replicate (kuhlmann, et al., 2017). as more surveys go online there has been a related increase in research examining slider interfaces as self-report tools. this research is generally focused on specific aspects of sliders, rather than comparing them to vas or lcs (radio button) interfaces. what little comparative research there is has suggested no significant differences between slider and lcs response formats (e.g., roster, et al., 2015). research has also pointed towards non-response being higher for slider compared to lcs response formats (liu, 2017). research of specific slider related issues such as direction (liu, 2017), suggest that the direction of the labels has no effect on self-report outcomes. the starting values for slider markers have an impact on 101, but not 21 or seven-point scales. however, forcing users to click the scale to start (i.e., no marker initially visible), increases missing data, particularly for 101-point scales (liu & conrad, 2018). the present review provides scant clear direction for continued research in the area of survey responses. the majority of the studies presented have pursued a relatively weak research design (liu and colleagues' programme is a nice exception to this problem) and focused exclusively on longstanding 20th century approaches to survey response formats (lcs and vas). in line with more recent research focused on computer-based surveys and the growing use of sliders, the future of self-report (like almost everything else) is mobile and touch based. it is critical that as our medium for engaging with media changes, that we adapt the ways in which we structure these media. for example, researchers should consider why we often use a touch radio button when something more intuitive and potentially more powerful might be invented. as wetzel and greiff (2018) have called for, future research needs to seek alternative response formats. in the current study we therefore pursued an experimental approach (i.e., random assignment of a three-item survey's scale interface) to testing both well-known and new response formats on mobile touch-based devices. 2.6 an empirical test of four mobile interfaces for survey data collection four survey interfaces were compared: labelled categorical scale (likert-type), vas, slider, and swipe. the first two were included due to their prevalent use across the previous century of research. the third (slider) was included because of its increasing use through computer and now mobile devices. the fourth (swipe; fryer & fryer, 2019) was included to test some new and alternative approaches that touch interfaces afford. swipe is built on a basic slider interface, but is presented on a slope. users “swipe” along the 45-degree angle (up, left to right) to move a ball up the incline. consistent with matejka et al. (2016), this interface integrated dynamic feedback and an approach to banding the intervals between labels in addition to ticks. ticks were presented both for the six labels and at 1/10 increments between the labels. 3. aims, research questions and hypotheses in the current study we aimed to highlight established issues with survey data, some of which are regularly addressed, and others less often discussed. in the current study we also aimed to introduce new questions that survey measurement faces as it integrates with the digital, increasingly mobile age. embracing this mobile era of survey use, the current study concludes with a brief experimental study comparing the four survey interfaces: lcs, vas, slider and swipe. five research questions (rq) were addressed in the current study’s experiment. sufficient prior research existed to support a hypothesis for one of the questions, the lower time to complete for lcs (likert type in most cases) relative to other interface formats. first, we were interested in whether the reliability (cronbach’s alpha) for scales would vary meaningfully across the four self-report interfaces (rq1). second, we aimed to determine whether any mean differences in interest for each of the six tasks separately could be attributed to the four interfaces (rq2). third, we aimed to assess/compare the predictive relationships from (a) prior interest and self-efficacy to the task interest (with each interface) and then (b) from the task interest to future interest in the course and domain (rq3). fourth, we sought to assess and compare the latent structure of the interest constructs measured by each of the four interfaces (rq4). fifth, potential differences in central tendency of the data resulting from each interface were compared, looking for patterns of response bias that might be due to the four interfaces (rq5). finally, we were interested in whether the time to complete the surveys varied meaningfully across the four interfaces. in this case, we hypothesised that likert would be the fastest to complete (hypothesis-1). 4. methods for the interface comparison 4.1 participants, ethics and procedures participants for the current study were postgraduate students (n = 81; female = 38; resulting in 644 responses) from one research intensive university in hong kong. students came from eight of the university's 10 faculties. students were completing a short course in preparation for teaching responsibilities as a part of their degree. the comparison of the four interfaces (survey formats) was undertaken within a broader project examining students' interest in course tasks, the course itself, and the domain of teaching and learning. across the course, participating students responded to short surveys either directly after tasks (task interest) or at the beginning/end of the course (course and domain interest). all surveys were completed during regular class time. students completed the short surveys on their mobile phones by capturing a qr code (embedded in course power points) which directed them to a survey within a custom designed online platform for micro-analytic surveys. the survey interface students engaged with were randomised for each survey qr code, meaning that students had an equal chance of facing any of the four interfaces for each of the six task interest surveys they were asked to complete. in the current study we therefore pursued a within-individual experimental design. as the interfaces were randomised for each of the six surveys, there was no guarantee that students would engage with all four interfaces and even if, by chance, they did, the number could not be even. this means that this study also relied on between student differences as well. for the main component of the current study (a comparison of four quantitative self-report interfaces) a three-item survey designed to assess students' interest in a specific task was utilised: this activity/task is personally meaningful; this activity/task is interesting; i want to learn by doing more activities/tasks like this. the predictive validity for the task scales were tested using regression from a future course interest scale consisting of five items (i.e., this course is personally meaningful; this course is interesting; i want more courses like this; i'm enjoying learning about teaching during this course; this course stimulated my curiosity about teaching) and a domain interest scale consisting of five items (i.e., 1) how much do you know about teaching?; in your spare time, how often have you tried to learn about teaching?; i have spent time learning about teaching on my own. how well does this statement match you?; i'm confident in my knowledge of teaching. how well does this statement match you?; i always have questions about teaching. how well does this statement match you?). all survey items were self-reported across a scale 0-5. labels for task, course and domain (3, 4, 5) asked students to what degree the item matched them specifically (not at all 0 completely 5). domain item 1 used the labels almost nothing 0 almost everything 5. domain item 2 used the labels almost never 0 almost always 5. the task and course scales have demonstrated acceptable reliability and construct validity in several past uses (fryer, et al., 2020; fryer, et al., 2019; fryer, et al., 2017; fryer, et al., 2016). the domain-level, depth of interest scale was developed recently (renninger, & schofield, 2014). it is consistent with current conceptions of individual interest and its development (i.e., renninger & hidi, 2015). in preparation for the current study, ethical approval was sought and obtained from the university’s human research ethics committee (ethics approval #1608028). prior to beginning the study, all students read an overview of the project, were informed that their self-reports would be anonymous and invited to contribute their self-reports to the research project. six students declined to participate in the research after reading the ethics statement and were removed from the current study, resulting in the aforementioned n-size. 4.2 analyses analysis for the empirical component of the current study began with an examination of the overall and interface specific descriptive statistics for the scale means and reliability (rq1). anova were conducted for the four interfaces, overall and task by task (rq2). regressions were conducted from prior domain interest and course self-efficacy predicting task interest for each interface separately; then regressions from task interest predicting course interest in the future was conducted and compared (rq3). factor loading (confirmatory factor analysis) for results from each interface was then compared (rq4). the central tendency, both visual and skew/kurtosis were calculated and reviewed (rq5). finally, an anovas were conducted to compare completion times for each interface (hypothesis #1). 5. results 5.1 descriptive statistics and reliability the overall means for the four interfaces (across all six tasks), for the pre-post measure, their differences by task, and scale reliabilities are presented in table 1. the reliability for each scale and for the task interest scales used with each of the four interfaces were all well above what is commonly suggested as being acceptable (> .70; devellis, 2012). significant differences were observed for the four interfaces across the project as a whole, but at each of the individual tasks assessed for interest no statistically significant differences were found (table 1 & 2). table 1 means, cronbach’s alpha and anova for the four interfaces across all tasks table 2 anova for differences for each task note: task interest means for tasks a-f for each self-report format 4.2 predictive difference by interface regression was used to experimentally test (i.e., random assignment of interface) the prediction from prior domain interest and perceived self-efficacy for the course to the four survey interfaces used for all the six tasks (combined). this test was then followed by regression predicting future interest in the course and domain from students’ interest in the course tasks, again for each of the task interest survey interfaces (table 3). for the prediction from prior domain interest to future tasks, the r2 (.04) was consistent for all but vas, which presented a non-significant (p < .05) relationship. course self-efficacy was significant for all four interfaces presenting the highest r2 the new interface (swipe = .09) and the lowest for vas (.06). task interest predicting future interest in the course (generally strong in past research with these constructs: e.g., fryer, et al., 2019; fryer, et al., 2017; fryer, et al., 2016) resulted in substantially more variance being explained (r2) (slide = .49, swipe = .37, vas = .29, lcs = .28). a similar pattern of relationships resulted for tasks predicting future domain interest (slide = .51, swipe = .37, lcs = .35, vas = .33). table 3 regression findings note: n refers to the number of survey completions 4.3 confirmatory factor analytic loading each item for each interface the cfa loading findings were generally consistent across the four interfaces (table 4). “interesting” generally presented the strongest loading, followed by a desire to reengage and finally perceptions of tasks being personally meaningful. table 4. cfa loading for each item for each survey question format 4.4 time to complete differences by interface the average time to complete the task surveys with each of the four interfaces was calculated and compared (table 5). despite the relatively large mean differences, no statistically significant differences were observed (p < .05). this is likely due to the relatively large standard deviation for the means. table 5 average time to complete task surveys with each of the four interfaces note: n refers to the number of survey completions with the specified format central tendency table 6 presents the distribution for the four tested interfaces. skew and kurtosis for each of the interfaces were within even the strictest heuristics (+1 – -1). the graphical distribution for each survey interface is included in the appendices (figures 1-4). the distribution presented by these charts makes visually clear the inherent differences between the types of data the different interfaces result in. vas presents the most skew and lcs appears to encourage students to choose the same ordinal rank, regardless of question, resulting in large amounts of twos, threes and fours but far fewer scores in between. swipe and slide presented the most normal looking distributions. table 6 distribution for the four interfaces 5. discussion a brief review of the extant research in the area of quantitative survey self-report formats (or interfaces in the current context) was presented. the literature reviewed came from a broad range of fields, with much of it providing scant direction beyond support for lcs (commonly likert in format) due to its ease of administration and in some cases for vas due to the nature of the resulting data (i.e., interval-like). more recent research examining slider formats has resulted in a handful of incremental suggestions for the field (e.g., use of dynamic response and potential of banding rather than ticks on the slide area) which have not yet been meaningfully taken up by the field. some of these findings were integrated into the interface tested alongside lcs, vas and slide, in a touch-driven format tentatively named swipe (an early version of fryer & fryer, 2019). in the short experimental study undertaken, six research questions were addressed. reliability for each of the interfaces was acceptable, with swipe and lcs presenting the highest reliability across the 6 tasks (rq1). no statistically significant mean differences were found for the individual tasks (table 2), but a significant difference across all tasks was observed—albeit with a small r2 (table 1). in this case vas presented the highest overall mean and slide the lowest (rq2). predictive modelling was undertaken – prior self-efficacy for the course and interest in the domain predicting future course interest; task interest predicting future course and domain interest – for each of the four interfaces. the clearest contrast was for task to future course and domain interest, where slide and then swipe presented the strongest relationship (rq3). confirmatory factor analysis followed, focusing on item loading, with results suggesting a consistent pattern of loading across the four interfaces (rq4). central tendency for the responses were examined statistically and reviewed graphically (appendices: figures 1-4). skew and kurtosis were within acceptable boundaries for all four interfaces. graphical representations of the four distributions suggested that the swipe interface presented the most normal distribution (rq5). the time to complete the three-question survey with the four interfaces was compared, indicating that, consistent with our hypothesis, the lcs format was the fastest to complete (but not statistically significant, p < .05) (hypothesis #1). a careful review of the data across the six surveys suggest that the differences between the lcs and the other formats declined precipitously with increased use suggesting a learning effect (i.e., getting used to the new interface) across students’ engagement with the task interest self-reports. 5.1 implications for measurement assuming the sample size was large enough for the experimental nature of the study (i.e., random distribution of the four conditions), two general findings standout. the first is the relative predictive strength of responses with each of the four interfaces. slide, followed by the new interface swipe, stood out as presenting the strongest βs for future interest in the course and domain. given the fact that much of the research with surveys like this will be aimed at predictive modelling, this finding is both alarming and potentially hopeful: alarming as the results suggest that the interface matters and can result in substantial differences; hopeful because it suggests that the slide and swipe (i.e., interactive and touch-based) formats have significant advantages over older formats like lcs and vas. the second is the time difference to complete the survey. marked differences supporting past findings pointing to the ease of lcs over other formats like vas were observed. rather than suggesting, as many previous researchers have, that lcs is therefore preferred due its ease of administration, we suggest that the flexible nature of mobile devices might be channeled to overcome this issue. a careful review of the survey completion time data suggested that the difference between lcs and slide/swipe narrowed substantially with increasing use. more intuitive interfaces for slide/swipe might be developed to close the gap, and animated directions for interacting with the interfaces might also be used to ameliorate this issue. while the skew and kurtosis outcomes for each interface were within acceptable boundaries, the graphical presentation made a clear case for swipe, vas and slide (in that order) as providing a more normal distribution of scores. given the reliance of most of our statistical procedures on such a distribution and the amount of potential data collected with mobile interfaces in the years to come, it seems reasonable to continue to develop continuous self-report interfaces. 6. limitations and future directions despite the experimental design, this study faced a number of limitations that should be addressed by future studies in this area. the first is the learning effect that is apparent for all of the continuous interfaces, but most obvious for the newest version (swipe). the high sds that resulted, clearly affected the study’s power to detect differences between the interfaces which were apparent in the means but were not statistically significant. in this study only participants’ responses to a very short survey were examined, whereas most surveys are much longer. future studies should examine what effect prolonged survey engagement has on different touch interface experiences as well. while more than 600 individual responses across the four interfaces were collected for this study, the actual sample of participants was quite small and very specific. it is important that future studies embrace a broader sample as well as a larger one. implicit in the study's design, analyses were conducted between persons, but participants were represented at multiple time points. this design therefore violates the assumption of independent errors as some of the data is nested within-person. to achieve the sample size necessary for a meaningful experimental test of all four interfaces, this limitation could not be avoided. a future experiment in a more controlled context (rather than a classroom setting) could undertake to obtain a clear counter-balanced sample and avoid this limitation. this was the first published test of the swipe interface (a pilot version). this test suggested both positive (high βs and reliability) and negative findings (high sds and time to complete) for the new interface. future studies from our research programme continue to refine this approach to self-report. the most recent version of the interface (fryer & fryer, 2019) is prefaced by an animated user interface infomercial (to spell out how to interact with it). additional tests comparing swipe with the slide (highest βs) and lcs (fastest time to complete) are being conducted towards fine-tuning this new dynamic, touch-based survey interface. it is critical to note that the present study's questions focus on students' self-reported emotions, beliefs and desires. it is reasonable therefore to constrain the implications of our results to the use of similar types of survey questions. one means of continuing to advance the research presented here would be through pairing think-aloud protocols with survey use (e.g., chauliac, et al., 2020; rogiers, et. al., 2020). this would provide a small window into the user's mind, suggesting how/whether a specific self-report format and touch interfaces interact with the self-report experience and outcome: i.e., send in a spider to catch the fly. an additional important area for investigation is that of surveys which enable the seamless integration of both categorical choice and continuous magnitude. to some degree the swipe interface sought to combine these elements into a single experience. future interfaces might extend this work or separate them into an intuitive two-step process: choose a label and then indicate the strength of your feeling for that category (see durik & jenkins, 2020). 7. conclusions vas and likert-type (lcs) formats are approaching their centenary. at the same time humankind sprints towards touch-based mobile devices as a critical nexus for interacting with and managing its world. self-report is therefore ripe to be improved (disrupted?). the revolution must start with those of us that rely on surveys for research. better, easier measurement means clearer results and more of them. as one baby step towards this revolution, results from the present study suggest that vas might be set aside as an option. it presented no clear benefits over the other interfaces in any of the tests and lacks any clear path to enhancement. in contrast, the present research suggests that continuous and interactive formats (slide & swipe) are a strong base for development in this area. the field is waiting for researchers with a penchant for disruptive improvement. notes: 1. visual analogue scale (vas) is "a testing technique for measuring subjective or behavioral phenomena (as pain or dietary consumption) in which a subject selects from a gradient of alternatives (as from "no pain" to "worst imaginable pain" or from "every day" to "never") arranged in linear fashion". (merriam webster, 2019) 2. a likert scale is a "rating system used in questionnaires, that is designed to measure people’s attitudes, opinions, or perceptions. subjects choose from a range of possible responses to a specific question or statement; responses typically include “strongly agree,” “agree,” “neutral,” “disagree,” and “strongly disagree.” often, the categories of response are coded numerically, in which case the numerical values must be defined for that specific study, such as 1 = strongly agree, 2 = agree, and so on." (britannica, 2019) keypoints the two continuous interactive interfaces (slide & swipe) yielded the most robust data for predictive modelling. future research with touch self-report interfaces can ignore vas formats. researchers seeking to improve on likert-type formats need to focus on ui that are quick and reliable to use. review of the existing research generally suggests that likert-type is superior to vas due to its ease of use. many researchers still maintain that vas formats yield more robust data than likert-type formats acknowledgements we would like to acknowledge the contribution of alex shum for carefully reviewing a previous draft and his overall contribution to this ongoing project. we would also like to acknowledge ada lee and peter lau who were central to collecting the data for this research and the broader programme. references adelson, j. l., & mccoach, d. b. (2010). measuring the mathematical attitudes of elementary students: the effects of a 4-point or 5-point likert-type scale. educational and psychological measurement, 70 (5), 796-807. https://doi.org/10.1177/0013164410366694 albaum, g. (1997). the likert scale revisited. market research society. 39(2), 1-21. https://doi.org/10.1177/147078539703900202 austin, p. c., & brunner, l. j. (2003). type i error inflation in the presence of a ceiling effect. the american statistician, 57(2), 97-104. https://doi.org/10.1198/0003130031450 berger, i., & alwitt, l. f. (1996). attitude conviction: a measure of strength and function. unpublished paper. bishop, p. a., & herron, r. l. (2015). use and misuse of the likert item responses and other ordinal measures. international journal of exercise science, 8(3), 297-302. boognese, j. a., schnitzer, t. j., & ehrich, e. (2003). response relationship of vas and likert scales in osteoarthritis efficacy measurement. osteoarthritis and cartilage, 11(7), 499-507. https://doi.org/10.1016/s1063-4584(03)00082-7 britanica.com (2019). likert definition. retrieved on november 18, 2019 from https://www.britannica.com/topic/likert-scale couper, m. p., tourangeau, r., conrad, f. g., & singer, e. (2006). evaluating the effectiveness of visual analog scales: a web experiment. social science computer review, 24(2), 227-245. https://doi.org/10.1177/0894439305281503 chauliac, m., catrysse, l., gijbels, d., & donche v. (2020). it is all in the surv-eye: can eye tracking data shed light on the internal consistency in self-report questionnaires on cognitive processing strategies? frontline learning research. 8(3), 26 – 39. https://doi.org/10.14786/flr.v8i3.489 devellis, r. f. (2012). scale development: theory and application. new york: sage douven, i. (2018). a bayesian perspective on likert scales and central tendency. psychonomic bulletin & review, 25, 1-9. https://doi.org/10.3758/s13423-017-1344-2 durik, a. m., & jenkins, j. s. (2020). variability in certainty of self-reported interest: implications for theory and research. frontline learning research, 8(2) 86-104. https://doi.org/10.14786/flr.v8i3.491 foddy, w. (1994). constructing questions for interviews and questionnaires: theory and practice in social research. cambridge: cambridge university press. fryer, l. k., thompson, a., nakao, k., howarth, m., & gallacher, a. (2020). supporting self-efficacy beliefs and interest as educational inputs and outcomes: framing ai and human partnered task experience. learning and individual differences. https://doi.org/10.1016/j.lindif.2020.101850 fryer, l. k., & dinsmore d.l. (2020). the promise and pitfalls of self-report: development, research design and analysis issues, and multiple methods. frontline learning research, 8(3), 1–9. https://doi.org/10.14786/flr.v8i3.623 fryer, l. k., nakao, k., & thompson, a. (2019). chatbot learning partners: connecting learning experiences, interest and competence. computers in human behavior, 93, 279-289. https://doi.org/10.1016/j.chb.2018.12.023 fryer, l. k., & fryer, k. (2019).情報処理装置、情報プログラムおよびこれを記録した記録媒体、ならびに情報処理方法.. patent # 6585129 (japan). translation: [dynamic touch based interface for survey self-report; translation of japanese patent title: information processor (information technology equipment), information program and a medium for the recording, and a method of information processing] fryer, l. k., ainley, m., thompson, a., gibson, a., & sherlock, z. (2017). stimulating and sustaining interest in a language course: an experimental comparison of chatbot and human task partners. computers in human behavior, 75, 461-468. https://doi.org/10.1016/j.chb.2017.05.045 fryer, l. k., ainley, m., & thompson, a. (2016). modelling the links between students' interest in a domain, the tasks they experience and their interest in a course: isn't interest what university is all about? learning and individual differences, 50, 157-165. https://doi.org/10.1016/j.lindif.2016.08.011 hayes, m. h., & patterson, d. (1921). experimental development of the graphic rating method. psychological bulletin, 18, 98-107. howell, j. l., collisson, b., & king, k. m. (2014). physics envy: psychologists’ perceptions of psychology and agreement about core concepts. teaching of psychology, 41, 330-334. https://doi.org/10.1177/0098628314549705 jaeschke, r., singer, j., & guyatt, g. h. (1990). a comparison of seven-point and visual analogue scales: data from an andomized trial. controlled clinical trials, 11, 43-51. https://doi.org/10.1016/0197-2456(90)90031-v kuhlmann, t., dantlgraber, m., & reips, u.-d. (2017). investigating measurement equivalence of visual analogue scales and likert-type scales in internet-based personality questionnaires. behavior research methods, 49, 2173-2181. https://doi.org/10.3758/s13428-016-0850-x likert, r. (1932). “a technique for the measurement of attitudes”. archives of psychology, 140, 5-55. liu, m. (2017). labelling and direction of slider questions: results from web survey experiments. international journal of market research, 59, 601-624. https://doi.org/10.2501/ijmr-2017-033 liu, m., & conrad, f. g. (2018). where should i start? on default values for slider questions in web surveys. social science computer review, 37(2), 248-269. https://doi.org/10.1177/0894439318755336 chauliac, m., catrysse, l., gijbels, d. and donce, v. (2020). it is all in the surv-eye: can eye tracking data shed light on the internal consistency in self-report questionnaires on cognitive processing strategies? frontline learning research. 8 (2), 26 – 39. http://doi.org/10.14786/flr.v8i3.489 matejka, j., glueck, m., grossman, t., & fitzmaurice, g. (2016). the effect of visual appearance on the performance of continuous sliders and visual analogue scales. paper presented at the proceedings of the 2016 chi conference on human factors in computing systems. merriam-webster. (2019). visual analogue scale definition. retrieved on november 18, 2019 from https://www.merriam-webster.com/dictionary/likert muthén, l. k., & muthén, b. o. (1998-2015). mplus user's guide. (sixth ed.). los angeles, ca: muthén & muthén. raykov, t. (2009). evaluation of scale reliability for unidimensional measures using latent variable modeling. measurement and evaluation in counseling and development, 42, 223-232. http://doi.org/10.1177/0748175609344096 rogiers, a.; merchie, e. & van keer (2020). opening the black box of students’ text-learning processes: a process mining perspective. frontline learning research. 8(3) 40 – 62. http://doi.org/10.14786/flr.v8i3.527 reed, c. c., wolf, w. a., cotton, c. c., & dellon, e. s. (2017). a visual analogue scale and a likert scale are simple and responsive tools for assessing dysphagia in eosinophilic oesophagitis. alimentary pharmacology & therapeutics, 45, 1443-1448. https://doi.org/10.1111/apt.14061 renninger, k., & hidi, s. (2015). the power of interest for motivation and engagement. new york: routledge. renninger, k., & schofield, l. s. (2014). assessing stem interest as a developmental motivational variable. paper presented at the american educational research association, philadelphia, pa. roster, c. a., lucianetti, l., & albaum, g. (2015). exploring slider vs. categorical response formats in web-based surveys. journal of research practice, 11(1), 1. vickers, a. j. (1999). comparison of an ordinal and a continuous outcome measure of muscle soreness. international journal of technology assessment in health care, 15, 709-716. https://doi.org/10.1017/s0266462399154102 voutilainen, a., pitkäaho, t., kvist, t., & vehviläinen‐julkunen, k. (2016). how to ask about patient satisfaction? the visual analogue scale is less vulnerable to confounding factors and ceiling effect than a symmetric likert scale. journal of advanced nursing, 72, 946-957. https://doi.org/10.1111/jan.12875 wetzel, e., & greiff, s. (2018). the world beyond rating scales: why we should think more carefully about the response format in questionnaires. european journal of psychological assessment, 34 , 1-5. http://doi.org/10.1027/1015-5759/a000469 9.appendices figure 1. distributions for interface for lcs figure 2. distributions for interface for slide figure 3. distributions interface swipe figure 4. distribution for vas interface codepen unlusoy publication frontline learning research vol.8 no. 2 (2020) 109 130 issn 2295-3159 expanding the notion of global learning: turkish-dutch teens’ networked configurations for learning aslı ünlüsoya, mariëtte de haana a utrecht university, the netherlands article received 25 october 2018 / revised 28 february 2020 / accepted 6 march/ available online 7 may abstract digital technology facilitate interactions between learners and resources at a global level. new learner prototypes are therefore proposed, such as the notion of the global learner. in this paper, we argue that these prototypes of global learning often do not account for the variety of ways in which youth use technology and see themselves as learners. we take the example of turkish-dutch youth to show empirically how they represent an alternative for what is often seen as the prototype of what a global learner is. we combine ego-network methodology with in-depth interviews to provide a detailed account of how 25 turkish-dutch teens see themselves as learners, how they make use of technology to pursue their interests, how they reach out to others and media resources, and how they form selves in relation to the values and norms of their (transnational) community. using the notion of ‘learner identity’, the study shows how these teens develop learner identities that are built on specific and culturally informed notions of ‘what a learning subject is’ that challenge the universality of the autonomous subjectivity implied in prototypical notions of the global learner. in addition, the study shows how through digital affordances, unique networked (trans)national connectivities are formed, which are informed by these teens’ specific socio-cultural position. we argue that by acknowledging these alternative ways of what a learning subject is, and how connections are formed, we can proactively incorporate them as useful models of global learning. keywords: connected learning; global learner; learner identity; turkish-dutch teens; ego-network analysis info corresponding author email: a.unlusoy@uu.nl doi: https://doi.org/10.14786/flr.v8i2.423 1. introduction: aim and scope in the learning sciences, new prototypical notions of learning have been put forward that correspond to the possibilities and challenges of the digital era. these notions foreground the informal domain as a space where learning takes place and oppose or challenge traditional models for schooling. for instance, inspired by the possibilities of gathering an endless amount of resources on the internet and connecting with likeminded others to explore these resources, notions such as ‘affinity spaces’ (gee, 2005), ‘connected learning’ (ito, gutiérrez, livingstone, penuel, rhodes, & salen, et al., 2013), or ‘personalized e-learning’ (o'donnell, lawless, sharp, & wade, 2015) have arisen. a similar example of a technology-driven prototypical model of learning is the notion of ‘global learning’. inspired by the possibilities of utilizing technology to facilitate interactions between learners of different cultures, which, in principle, provides learners with the opportunity to develop global perspectives, the notion of global learning has been put forward to inspire educational reform (gibson, rimmington, & landwher-brown, 2008). these concepts have in common that they put the learners’ personal engagement at the centre as well as the learners’ ability to gather (digital) resources based on this personal engagement. as such, they challenge traditional, authority-driven models of learning, in which knowledge distribution by institutions is the norm. at the background of these discussions about new metaphors and models for learning in the digital age, our ambition with this paper is to expand our ideas of what a global learner might be. we do so through showing empirically how turkish-dutch youth develop particular socio-culturally informed ‘learner identities’ as well as unique networked (trans)national connectivities that challenge dominant metaphors of learning in the digital age. as we hope to show, they challenge the image of the autonomous, individualistic self-implied in these ideals of global learning, as well as the idea that connectivity evolves around the agentic efforts of the individual learner. adopting a perspective on learning as socially and culturally situated, this paper argues that such situated perspectives seem to be forgotten with the launching of 21st century notions of learning. therefore, the paper seeks to expand such a perspective into learning in the 21st century and new models for learning. in this paper, we build upon earlier work (de haan, leander, ünlüsoy, & prinsen, 2014, p. 508) in which we argued for a critical reconsideration of “idealized digital connectivities for learning”. in work on these idealized connectivities, the suggestion is made that “people are optimally networked so that resources are equally available, shared and voiced, and participation possibilities are maximized” (p. 508). we have argued that everyday social practices of connectivity reflect a much more nuanced and differentiated reality, based on the idea that ‘connectivities for learning’ are situated over time and socially constructed social practices that are informed by specific cultural norms and values. we have proposed that personal networks as a unit of analysis are a good starting place to explore these nuances and we have coined the term ‘networked configurations for learning’ (ncl) to refer to the idea that connectivities for learning are diverse and socially situated. in this paper, we expand our earlier argument on the specificity of connectivities. first, in this paper we provide a more detailed account of one group of learners, turkish-dutch youth, of which we have gathered more ethnographic data in comparison to the earlier paper. second, we are making use of this sample to also elaborate more extensively on how the notion of ‘what a learner is’ can be socio-cultural-specific. we draw on sinha’s (1999) idea of ‘learner identity’, who has argued that being or knowing how to be a particular kind of learner is not something that is ‘given’ or universal but rather something that is formed in socialization practices associated with particular communities. third, in this paper we elaborate more explicitly on how digital connectivities are part of global-local dynamisms shaped by both migration and digital technology. in particular, we focus on the transformative potential of these mobilities for learning, by showing how ‘to be here and there at the same time’ and how being a member of several normative communities simultaneously provides unique opportunities for learners. the study thus provides an empirical record of what we think of as an ‘a-typical case’ of a 21st century learner. the study documents turkish-dutch immigrants’ use of technological affordances to expand their learning and then asks how their efforts relate to the personalized, individually engaged learner pictured in new prototypical notions of learning. before we present our theoretical take on learning as a cultural and situated phenomenon, and how this relates to notions of connectivity and new technologies, we give a brief overview of how globalization and new technologies have spurred new notions of learning (1.2) as well as how teens from minority backgrounds constitute a good example of how technology is adopted in particular ways, related to the dynamics of migration (1.3). 1.1 global societies and new notions for learning we live in an era that is marked with abundant information and almost constant exposure to it. news headlines, blog, vlog and status updates, tweets, social media feeds, emails and text messages ask for our attention not only as the recipients of the information but also as the distributors, co-creators, and recyclers of it. new information and communication technologies (ict) are widely acknowledged for their role in lowering the threshold of information access for everyone and in enabling new ways to interact. however, these changes are not only dependent on the influence of technologies. how people use these technologies is strongly related to who they are and their social, cultural and material environment. the dynamic interplay between technology and identity eventually also shapes the ways in which people interact, socialize and learn and can create specific socio-technical practices and divides in this respect (hildreth & kimble, 2004). knowledge production and consumption in so-called ‘global’ societies happens at geographically dispersed scales. in globalized information and knowledge societies, individuals are not only part of relatively homogeneous locally based communities, but, at the same time, they are a member of many different, locally and globally dispersed networks, which provides them with unique and tailored possibilities to find knowledge and learn in these networks (farrell, 2006). this idea resonates with the more general concept of networked individualism that addresses how we relate to people in the digital age (rainie & wellman, 2012). rainie and wellman (2012) observe that in the past, personal networks used to be mainly defined by small, densely knit local groups, and communication was primarily face-to-face and location-dependent. now, individuals are much less constrained by geographical boundaries, and even though traditional social spaces defined by, for instance, kinship relationships, neighbourhood and work remain important, they are no longer the only sites for socialization. according to castells (2007), these changes also mean a shift from a more hierarchically structured social system to a more networked and participatory one, which is profoundly transformative for individuals as well as for the foundations of society as we know it. some have argued that this development fundamentally changes the way we learn, while simultaneously causing a greater diversification of the possibility to learn. an example of such work is developed in alignment with the notion and educational ideal of ‘connected learning’ (ito, et al., 2013). connected learning, which is enabled through new digital infrastructures in globalized societies, is defined as learning that is socially embedded, interest-driven, and oriented towards educational, economic, or political opportunity. basically, the premise is that new digital infrastructures and networks allow young people to pursue personal interests or passions, which they, with the support of others, turn into learning opportunities, which again might also lead to academic achievement or civic engagement. the premise is that new technologies enable people to explore and share interests freely and openly. there is a much greater freedom -in comparison to standardized educationin how people invest their time and energy to satisfy their (varied) interests as well as in the actual potential to turn these interests into careers. connected learning has been presented as an ideal of learning in the global society for all, and in opposition to and as an alternative for outdated notions and practices of learning and education (ito, et al., 2013; kumpulainen & sefton-green, 2014). although its idealized form is only available for progressive digital media users typically associated with privileged minorities (ito, et al., 2013), this idea in fact highlights the variation in the lives and learning possibilities of young people. 1.2 new migration and technology: changing opportunities for learning for immigrant youth in particular, teens from minority backgrounds constitute a good example of how technology is adopted in particular ways. for a long time, an important defining aspect of being an immigrant has been the geographical, social and cultural gap between the two ‘homelands’; the one that is left behind and the one of settlement. however, under the influence of new technologies, the image of the “uprooted migrant” is now replaced with the “connected migrant” (diminescu, 2008). new technologies enable a space to be ‘together’ regardless of actual physical locations and enable being here and there simultaneously. the effort to establish new belongings and associations while maintaining the connections with loved ones and acquaintances wherever they may be is now a key part of the migration experience (diminescu, 2008). these network connections can be considered also as paths of information, belonging, support etc., and form important “linguistic and social capital” (lam, 2014, p. 503). more importantly, these new technologies provide immigrant teens with forms of networked capital, which reflects their social, cultural, ethnic, and historical background as well as their material reality. often these networks provide them access to different social spheres that are heterogeneous. these new connectivities and the life worlds they give access to have implications for what it means to learn and socialize. the focus becomes much more on what it means to learn to participate and move through multiple different social spheres as well as on the process of transformation that is necessary to participate in these heterogeneous social and culture spheres and networks (de haan, 2011). although new technologies also provide mainstream youth with these possibilities and challenges, they seem to define immigrant youth in particular. there is a small body of literature that indeed shows that immigrant youth access a variety of different spaces, social networks, which enable as well as challenge their learning in particular ways in comparison with mainstream youth. for instance, lam (2009) observes that as chinese-american teens explore their interests online they use both chinese and english. this enables them to access a distinct range of information and media content, which provides alternative, more empowering spaces for their learning compared to learning at school. likewise, messina dahlberg and bagga-gupta (2014) show how in online communities with multiple ethnic backgrounds, culturally and linguistically hybrid ways for co-constructing and mediating learning are supported, which are different from (monocultural or monolinguistic) institutional learning spaces. below, we will elaborate our argument on how new technologies create particular and situated opportunities for learning, departing from the notion of learning as a situated phenomenon (1.2.1). we argue that both the notion of ‘what a learner is’ (1.2.2) as well as connectivities that are constructed for learning are culturally and socially situated (1.2.3). 1.2.1 the ‘particular’ of learning and the acknowledgement of non-mainstream notions we draw upon sociocultural learning theories, and more specifically on the notion that learning is situated in socio-cultural practice in two different ways. first, learning is situated in the sense that learning is a product of the activity, context, and culture in which it is developed and used (brown, collins & duguid, 1989). it is situated in sociocultural practices precisely because ‘human beings have the need and ability to mediate their interactions with each other and the nonhuman world through culture’ (cole, 1998, p. 291). it cannot be captured by just looking at individuals. learning is distributed among co-participants of communities of learners (lave & wenger, 1998). second, learning is situated in the sense that it involves the appropriation of particular heritages and particular learner identities, and there is variation in how communities guide learners according to culturally informed notions of what learning is (rogoff, 2003). this second position represents a more politically oriented strand of studies, as the issue is often raised that the heritages, identities and culturally informed learning practices of minorities are not always acknowledged in mainstream education (gonzalez & moll, 2002) or in educational theories (rogoff, 2003). this study wants to highlight in particular the second sense of situatedness, while acknowledging the first. 1.2.2. becoming a particular kind of learner: adopting a ‘learner identity’ to foreground the subjectivity of the learner, studies in the sociocultural tradition have argued that becoming a learner also involves developing a version of ‘the self’, which fits the cultural expectations of what is expected from a novice. as sinha (1999) argues: becoming a learner is a situated phenomenon, which requires earlier experience in a particular socio-cultural practice. for instance, learning to recognize the appropriateness of a particular socially organized set up for a teaching learning situation and positioning oneself as a learner in accordance with socially appropriate roles (e.g., teacher and learner positions) is something that requires knowledge and prior experience of how learning is culturally and socially framed. developing human beings are being constructed and positioned in ‘particular and specific kinds of non-discursive practices, in such a way that he or she becomes a learning subject, or self, of the kind required by the culture within which teaching learning situations and opportunities are situate’ (sinha, 1999, p. 33). to elaborate his point, sinha contrasts the often taken-for-granted image of the creative learner with other taken-for-granted images of learners, such as the idea that learners are information-processing subjects. he claims that we often forget that these notions of what a learner is or should be are themselves shaped by normative traditions on learning. when we, for instance, assume learners to be creative, this implies a socio-culturally constructed self that understands him/herself as a creative developing being. the same applies for the idea that learners represent an autonomous self that is operating relatively independent from her/his social environment in terms of motivation, cognition, awareness, judgement and action. in other words, the learning self is not a culturally neutral concept but depends on particular interpretations of how a subject is supposed to grow, relate, identify, know, etc. although the relationship between learning and identity has been addressed in different ways (see for an overview moje & luke, 2009), this particular point is often forgotten. it is partly reflected in the distinction that arnseth & silseth (2013) make when they describe the learning self as both ‘a’ novice, that is, as becoming a central participant of a community that is endowed with a particular (community related) identity, and ‘a particular kind’ of novice, involving all it takes to become a central participant of that community. it is this second issue that we address here. however, evidently, both notions of a learner identity can never be entirely independent as both are embedded in culturally based notions of what membership in a community means. 1.2.3 notions of connectivity and learning not only learner identities are particular and situated, but likewise (online) connectivities that are constructed for learning are defined by socially and culturally informed experiences. following what we described above regarding the unique and tailored possibilities to find knowledge and use connections for learning afforded by technology, we argue that these diversified connectivities are situated in socio-cultural practices. as noted above in section 1.1, we have termed this networked configurations for learning (ncl). as ‘networked individualism’ and ‘connected learning’, ncl focuses on the role of the new technologies and the importance they deem to our increased networking capacity via these ict. however, in the concept of ncl, an argument is developed on how this network capacity matches with the socio-cultural, economic, personal conditions and drives of individuals or groups. moreover, it is used to study how these networks function for learning and allows description of the particular online and offline networked connectivities of diverse socio-cultural groups and the culturally and socially informed experiences for learning these connectivities enable (de haan, leander, ünlüsoy, & prinsen, 2014, p. 532). ncl builds upon the idea that the personal networks and a person’s learning and socialization experiences are directly related to and interdependent with one another. personal networks are the dynamic mechanisms where important everyday learning experiences are situated. configurations of these networks are only partly shaped by new technologies and, as argued earlier, it is essentially people’s social, cultural, ethnic, and historical background and material reality that shape these networks. in this study, we describe how the formations of the networks of turkish-dutch youth inform and shape their learning, while also paying attention to the wider socio-cultural and historical context of these immigrant youth. before we introduce our study, we provide an overview of the literature on turkish-dutch teens in the netherlands, in particular as related to their media use, and how this has been discussed as related to what it means to grow up as a minority youth. 1. 3 turkish-dutch teens the turkish-dutch youth in our study are secondor third-generation immigrants: children of families whose (grand-)fathers were recruited mostly from the rural regions in turkey. they migrated to the netherlands for labour and reunited with their family over the course of eighties and nineties (schneider, crul, & van praag, 2014). although current policies expect minorities to integrate, earlier integration was not facilitated as labour migrants were expected to return to their country, and language and culture maintenance as well as concentrated settlement were supported by the dutch government. this policy is now seen as one of the explanations for the relative segregation of the turkish immigrant community (vedder & virta, 2005; verkuyten, 2001). studies on turkish-dutch adolescents have shown that they are raised in families that are very concerned with transmitting the turkish tradition, history and language, and relationships between adolescents and their parents are highly impacted by what is considered appropriate according to the norms and values in the turkish community. turkish youth also show a strong attachment and self-esteem (related) to the turkish community (verkuyten, 2001). earlier media researchers have reported how media, especially television, is used by turkish families, including youth, to orient themselves towards turkey and that they are also oriented towards homeland media (d’haenens, 2003). moreover, turkish immigrants are documented as less active on the web, e.g., on discussion fora, in comparison to their moroccan peers, the other large immigrant population in the netherlands (ünlüsoy, de haan, leander, & völker, 2013). content analyses showed that the online discussion fora they use generally deal with turkey and turkish culture or identity (d’haenens, 2003). from another perspective, milikowski has pointed out how television watching can also have de-ethnicizing effects on these youth through the comparative lens it offers (2000). there is not much known from the literature on how turkish youth orient themselves on the internet from the perspective of their learning. mostly, the literature that touches upon issues of education and learning deals with the participation of turkish youth in formal schooling and their school success. other literature centres around key factors relevant for public participation such as employment (e.g., crul & schneider, 2010) or issues of identity and well-being (e.g., phalet & hagendoorn, 1996; verkuyten, 2001; vedder & virta, 2005). studies on the informal educational climate in turkish families describe turkish families as being defined by traditional gender division roles, fear of ‘dutchification’ of their children (lindo, 2000), as well as the significant gap between turkish children’s and their parents’ experiences in education (coenen, 2001). additionally, studies have shown that immigrant parents of turkish origin in the netherlands orient themselves towards collective and in-group-serving values in the education of their children (phalet & schönpflug, 2001). for youth in turkey, studies show that they have changed towards more independence, self-respect and autonomy in comparison with their parents under the influence of rapid economic and social change. however, these orientations continue to exist next to a strong orientation towards respect for tradition, obedience, politeness, honour for parents and elders and adherence to social expectations (morsunbul, crocetti, cok, & meeus, 2016). the abovementioned literature, apart from the fact that the studies that report on turkish-dutch immigrant youth are relatively outdated to provide the background for this study, are informative with respect to the challenges youngsters in this community might be facing for their education and learning. nevertheless, it lacks a perspective that considers how the global changes induced by ict and social media have changed the learning opportunities for turkish-dutch youth in the netherlands. we lack knowledge of the impact of new technologies on the learning opportunities of young immigrants such as the turkish-dutch youth in the netherlands. how do the affordances of technology, and the connections and resources it provides, define these teens in who they want to become, and how they learn to become? how do the affordances of these technologies also define the global-local dynamism that characterizes the lives of these immigrant youth? and, in line with the aims and scope of this paper as described above, how can we describe these teens as a particular case of a global and connected learner to meet our ambition of expanding our ideas of what a global learner might be? 2. current study & research questions in line with the aim as outlined above, in this study we ask how turkish-dutch teens perceive themselves (as learners), what the characteristics are of their personal online and offline networks, as well as how these networks enable and inspire them to achieve their learning goals. the specific research questions that guide our analyses are: 1. how can the personal networks of turkish-dutch youth be described in terms of structural characteristics (size, density, clusters) and composition (e.g., homogeneity, geographical spread)? 2. what characterizes turkish-dutch teens as learners? we approach this question by asking how turkish-dutch teens characterize themselves, what their interests and ambitions are, who or what they want to become, and what their view is on how they learn (to become someone)? 3. how do turkish-dutch teens’ networks function for their learning? in line with our goal to understand how new technologies create particular and situated opportunities for learning, we ask the following sub questions. can we distinguish particular interest-driven learning network (sub)clusters? are such networked sub clusters mediated by specific technologies or media resources? how do such sub clusters mediated by technologies enable or put boundaries on the learning of turkish-dutch teens? 3. methodology 3.1 sample and procedure a total of 25 turkish-dutch teens of 13-16-year-old (m = 14.68, sd = 1.03; 14 female participants) were interviewed for this study. the participants were from two inner-city schools in secondary education. the school in rotterdam (n = 12; 6 female) was a preparatory school for vocational university (called havo: hoger algemeen voortgezet onderwijs) and the school in den bosch (n = 13; 6 female) was a lower preparatory school for secondary vocational training (called vmbo: voorbereidend middelbaar beroeps onderwijs). participants who went to the same school knew each other as schoolmates. all participants were born in the netherlands; their families (either parents or grandparents) have migrated to the netherlands for labour. participants were drawn from a largescale survey study on learning, identity and the use of new media; the survey sample was representative of migrant youth age 12-18 in the netherlands, in secondary education. given our interest in personal networks and (online) connectivities we selected the students who had reported online media use (i.e., checking in their online social media account, watching videos and using an instant messaging application) on a regular basis in the earlier survey. through the schools we informed youth and their parents regarding our continued research and that participation was voluntary. the participants were informed that they could withdraw from the interview at any point. none made use of this possibility. the interviews took place in a quiet room in schools, during school hours. they lasted on average 1,5 hours and the students received a voucher for their participation. the interviews were audio-recorded transcribed verbatim. during the transcription process, we found out that 3 interview recordings were corrupted, and one interview was only partially recorded. these 4 cases (3 girls, 1 boy) were excluded from qualitative analyses. the participants were interviewed using a social network interview (sni) technique which revealed information regarding the structure and composition of their online and offline personal networks, see section 3.2. for details. on average, networks consisted of 21.5 contacts (sd = 6.57), varied between 12-35 contacts across the sample, we collected information over 537 network contacts in total. 3.2 instrument and measurements sni is a semi-structured, in-depth interview instrument that is used to gather information and analyse personal networks, also called ego-networks. it consists of two parts. the first part, called the name generator, identified the ‘important people’ in the lives of our participants. we asked the participants to think of important people in their lives, e.g., who they identified with, who were reliable or who they hang out with. we also prompted the participants to think of different spaces (school, neighbourhood, social media, vacations) to help them remember people who might be considered important for their personal network. we used the network analysis programs vennmaker 1.0 and nodexl to collect information and visualize the networks. ego-network data contain demographic information regarding all contacts (called alters) in the participants’ (called ego) networks and information to interpret the relationships between the ego and his or her alters (e.g., how frequently they communicate) (crossley, belotti, edwards, everett, koskinen, & tranmer, 2015). in this study, we collected the following information about each alter: age, gender, location (same household, neighbourhood, city, elsewhere in the netherlands, outside the netherlands, unknown) and level of education. we also collected the alters’ relationship to the ego (immediate family, extended family, friends from school, friends elsewhere, acquaintance), how they communicate with that alter (mainly online, mainly in-person [offline], both onand offline), and whether alters knew each other (i.e., whether they would recognize and talk to each other if they saw each other on the street). the (clustered) position of alters, as related to each other and the respondent, was determined using the harel–koren fast multiscale algorithm, which is one of nodexl’s force-directed algorithms (alters/nodes naturally push away from each other, while edges [relations/connecting lines] bring them closer together). this results in highly connected nodes migrating to the centre, while less connected nodes are pushed to the outside. the ‘groups’ function of nodexl was then used to calculate clusters, which works by aggregating closely interconnected groups of nodes. only when the network visualizations were generated by this software, we progressed to the second part of the interview. we asked the participants if the visualization resembled what they thought their network would look like (e.g., ‘does this network picture and the groups generated represent your network?’). overall, the representations were reported to be accurate, and small differences were discussed in the interviews. the second part of sni covered 1) how teens defined and identified with the different parts of their networks (we asked questions such as “are there people in this network picture that you look up to?”, “who are the people in this network that you spend most of your time with?”, “what do you do together?”); 2) what kind of (online or offline) learning activities they recognized in their network relationships (we asked questions such as “are there people or groups of people in this network with whom you undertake activities in which you want to become better?”); 3) how new technologies played a role in maintaining the network and what role these play for their learning (we asked questions such as “what are some of the things that you became better at (online or offline) over time?”, “how (if at all) did using new technologies made the experience different?”). the interviews were conducted with continuous attention for the personal networks of these teens, and their statements were consistently connected with the visualized personal network maps throughout the interview. prior to starting the interviews, we also checked briefly what the participants’ associations with learning were. participants who strictly thought of school learning were encouraged to think of the concept more broadly (such as how they learned to bike, how they found out about a new app, how they explored different sports or developed a hobby) so that we could come to a shared understanding of the idea of learning. we informed the participants that school-learning examples were okay to mention, but that our study had a broader perspective on learning. 3.3 analyses the first research question: ‘how can the personal networks of turkish-dutch youth be described in terms of structural characteristics (size, density, clusters) and composition (e.g., homogeneity, geographical spread)?’ was answered by analysing the quantifiable characteristics of ego-networks. based on frequencies and averages, we described the general structural and compositional features of networks. variables of ethnic, gender and age homogeneity were created per ego-network by computing the amount of alters who share the same ethnic background, gender or age as the ego. this measurement reveals the proportion of people who are similar to and/or different from the ego, in other words, the relative diversity (or uniformity) in each network. density in each network, that is the proportion of individuals in a network who know each other, was computed to assess how tightly connected each network was. the network characteristics of girls and boys were also compared to each other. section 4.1 describes the results of this analyses. to answer the second research question, ‘what characterizes turkish-dutch teens as learners?’, the transcriptions were first read, with this research question and the respective sub-questions in mind. nvivo software was used to label and analyse the narratives. we paid attention to perception of the self, identity markers, self-descriptions, and in cases where these were present, we pay attention to how these were related to issues of development, becoming and learning. next, we focused on how they defined themselves as a learner, or how they defined striving to be someone (becoming) more generally. section 4.2 and 4.3 describe the results of this analyses. to answer the third research question: ‘what characterizes turkish-dutch teens’ networks as learning networks? and how do their networks function for their learning?’ as well as to answer the respective sub questions, we focused on particular interests, hobbies, and activities that they mentioned, asked if these were represented by particular sub clusters of their networks, if and how these were mediated by particular technologies, in particular when the relations were contacted offline, while also paying attention to the specific location of these network clusters or individual relations. finally, we focused on if and how these sub-clusters enabled or hindered their learning. we start off presenting general network characteristics in 4.1 (e.g., divides in their networks, and what characterizes the people in their networks), and continue with the narratives on their identity as a learner (represented in 4.2. and 4.3), while also connecting these narratives to the network data from 4.1. in sections 4.4 to 4.6, we again combine network data with their narratives on learning when we focus on how their learning happens in particular networked sub-configurations, paying attention to how technology mediates these configurations, and how these function for their learning. for instance, we argue how technology plays a role in creating specific network divides and how this works for their learning, or how technology provides access to particular networks, which then provides entrance to distinctive opportunities to gain information, form opinions, discuss positions and gain new insights. the analyses as a whole must also be read as a commentary on assumptions of models of global learning, especially when the analyses address how global learners can be identified and what kind of connectivities technologies create. 4. findings: networked configurations for learning of turkish-dutch teens 4.1 turkish-dutch teens’ network characteristics: quantitative data the following structural and compositional characteristics of the networks are derived from 25 ego-networks with 537 alters in total. in table 1, we present a detailed overview of the personal networks and how boys’ and girls’ networks compare to each other. there were no significant differences between boys and girls regarding the proportions of different network characteristics. the noteworthy similarities across the networks are highlighted below. as explained in the methods section, networks were generated based on the important relationships of participants. the turkish-dutch participants generated largely family-based, ethnically homogenous personal networks. the networks were densely connected, meaning that most people knew each other. the algorithm we used generally created two clusters, given the interconnectedness of the networks. the clusters created by the algorithm were typically characterized by family versus friends’ relations, or older generations versus peer relations. the participants confirmed the cluster structure; most of the participants divided their networks based on a friends and family sub-cluster. in a few cases, the algorithm created three clusters, which youth identified as family, and two different groups of friends (e.g., from a sports club and school or from the mosque and from school), and in one case virtually all network contacts were connected, resulting in a single cluster. table 1 overview of turkish-dutch youth’s network composition (in %) family members were nearly always the majority in their networks. the personal network with the least amount of family still had 45% (9 out of 20 alters) of family members, and the percentage went up to 80%, with an average of 60.6% family presence in networks. on average, 40.3% of alters were older than the participants and 5.8% were younger; peers were on average 53.9% of all network contacts. network contacts who were mainly contacted online were 14.7% of all network contacts (79 out of 537). these 79 people were nearly exclusively family members who lived in turkey or elsewhere, but outside the netherlands. there were no statistically significant differences in network configurations between boys and girls. for the whole sample, density scores varied between .50 and .98, indicating in the least dense network 50% of contacts knew each other. on average, 88.3% of all contacts were of turkish descent (varied between 55% and 97%). there was a clear preference for hanging out with same-gender peers (70.6% were same-gender); often the only men in turkish girls’ networks and women in boys’ networks were their relatives. nearly a quarter (22.3%) of all network contacts lived outside the netherlands (often in turkey but also in germany, france and belgium), indicating that geographical distances were not preventing them from keeping in touch with their family and friends. the contacts that lived abroad were predominantly family members (89%), friends (10%) and 1 acquaintance (1%). 4.2. perceptions of the (learning) self: wanting to be like them “you become who your parents raised you to be” in both 4.2 and 4.3, we analyse how turkish-dutch youth perceive their ‘learning self’. in line with how we defined the notion of learner identity above, we first concentrate on their notion of ‘self’ in 4.2, while in 4.3 we extend this analysis with a focus on their vision on development and becoming. in both cases, we do so under the assumption that these two notions are highly related. we asked the participants to think about the characteristics, experiences, people, things and interests that made them ‘who they are’. although there were individual differences in the way the responses were formulated, the prominent trend among all participants was their emphasis on and identification with their family and community. this was also clear from their social networks, as we just reported in 4.1, which for a large part consisted of members of the turkish community, mostly family. another sign of this family orientation, as the network pictures illustrate (see figures 2 & 3), was that parents knew (almost) every one of the network contacts of their child. according to the participants, their family relationships, and in some cases relationships with good friends, shaped who they were. in response to who or what made them who they are, the participants often simply stated ‘my parents’, ‘my family’ or tahir (15, m), “without my parents and siblings i am nothing”. emel (13, f) “you become who your parents raised you to be”. simge (16, f) “what i learn at home from my mother and father shapes how i think and how i behave. my friends, they learn from their parents and behave that way…when we are together [with her group of friends] we influence each other too and do the same things together”. these examples illustrate a common understanding among turkish-dutch youth that the development of the self does not so much relate to becoming an independent self but a self that is highly relational, involving their closest relationships (with their parents and friends). these examples show not only that the notion of (being like your) family is a central and essential aspect of turkish-dutch teens’ identity but also that in their discourse on the self, a reference to the collective was always prominent. this was also evident from the fact that youth, when asked to describe themselves, more often referred to community values, such as being “respectful, especially towards older people” and “trustworthy, or honest”, than unique qualities. thus, rather than characteristics that typify an individual, these qualities reflect a community ideal of how one should be and behave. in addition to mentioning family, being turkish was a prominent identity marker in their discourse on the self. this was inferred from a variety of responses to questions such as “with whom do you feel you can be yourself” and “where/when do you feel at home”. “turkish-ness” seemed to represent a ‘comfort-zone’; a ‘place to withdraw’ or a state of feeling particularly at ease and seemed to be related to having a common history, values, and language. other implicit references to their being turkish included speaking turkish at home, especially with parents, but also among friends, going to turkey for vacation and following turkish media. this identification with the turkish community was also reflected in their network structure; 88.3% of all network contacts had a turkish background (see figure 1 below). figure 1. ethnic groups in turkish-dutch teens’ networks. the orientation towards turkey was also encouraged in the family. for example, yildiz’s father explicitly encouraged her to speak turkish more fluently: “my father says ‘you must learn turkish’, he corrects my turkish… his turkish is very good. with my mother, i speak only turkish because her dutch isn’t good” (yildiz, 14, f). the media-diet of the participants was primarily in turkish, and this, too, was sometimes encouraged by their parents. adnan (16, m): “my father comes home from work and he talks about all the news. he must listen to the [turkish] news, read the newspapers and teletext and i’m at home beside him so i listen with him. i talk about turkish politics a lot…”. satellite television and online streaming were accessible for all participants. these technologies gave participants continuous access to turkish media products (i.e., news, series, reality shows) and provided them with a wealth of information and material to understand and define ‘being turkish’ for themselves. 4.2.1 different others as contrasting examples in the diaspora however, as already indicated above, through their social networks, youth were able to contact extended family and friends of family members who live in turkey and in other migration countries (compare table 1, which indicates that 22% of their network contacts are transnational contacts). these transnational contacts, especially the ones from other migration countries, ruptured the relative homogeneity of their models for identification as these family members were socialized in communities that partly hold different values and norms. for example, ahmet (16, m) whose sister’s family lives in germany says “my nephew is very different from me. he is a good person, that’s true, but he’s different…he doesn’t do sport, he sits too much behind the computer and he smokes…when we are there i get along with him and his friends, but up to a certain point…if they say come we’ll go smoke, i won’t…i learn german at school here, but when i am in germany with my nephew i learn more. i understand everything, but i cannot talk very well”. this, and many other examples, show that the turkish diaspora, and the possibility to connect with it through digital technology, brought these teens in contact with other cultural traditions, alternative possible selves and ‘versions’ of being turkish, that serve as extended opportunities for learning and identification. for instance, they provided important language learning opportunities, or a comparative perspective on life between netherlands and other countries of the turkish diaspora in terms of economic chances, school experiences, teenage life, youth cultures, and gender roles. in this sense, their perception of the (learning) self, as grounded in a particular version of communal belonging, seems to be changing through these digitally afforded networks, which allows a more diverse and fragmented identification with their community. figure 3. network of ahmet. 4.3. learner identity: loyalty to the community, hierarchy and learning from role models as a next step in our analyses, we focused on how teens expressed a process of becoming someone, or in other words, how they saw themselves as a learner. to understand turkish-dutch youth’s associations with learning, and what kind of learners they perceived themselves to be, we asked them questions such as ‘what do you associate with the word learning?’, ‘when and with whom do you feel that you learn something?’, or ‘is there something you strive to get better at?’. our findings show that informal learning experiences were often expressed in narratives of ‘becoming a particular kind of person’, while taking someone from their community or family as a model that represented particular values and status. the participants often told us what kind of person they wanted to become, taking an individual as an example. for instance, emel (13, f) mentioned her uncle (who is part of her online network, see her network picture in figure 2) as her role model; “i would like to be exactly like him…when he was young he said to the family that he was going to study and graduate (at) university. he kept following his dream until he achieved it and i want that for myself. he is from elazig and he studied in cambridge”. another example of this was tahir (15, m), who said that his cousin was an inspiration for him because “he has a good life although he did not have much money. he has, how should i say that, he has worked a lot, worked a lot, gave it [money] to his parents to pay for the house […] therefore, later when i have a job, i will also give a part [of my income] to my mother, i also want to take care of my parents.” these role models often share certain characteristics such as being loyal to their family, working hard, starting with very little and achieving their goals despite difficulties. the narratives often highlight these teens’ appreciation for such role models and their desire to become a similar example once it is ‘their turn’ to do so. learning then represented modelling the important others from the community, as well as returning or giving back to the community, rather than seeking out a unique, individual path that distinguishes the individual from other members from that community. figure 2. personal network of emel. age-related hierarchy and status play a significant role in how youth perceive the workings of learning as relational. in the following example, emel (13, f) illuminates this hierarchy by describing herself as a role model for her younger brother: “my younger brother learns a lot from me, that he needs to respect older people, that he needs to follow his dreams…if there is something he doesn’t understand in his schoolwork he also comes to me”. furthermore, age and experience were essential elements and aspects in their vision of how one learns and gains wisdom. when comparing her peers (friends) to the older people in her network, simge (16, f) said: “as you grow older you become more thoughtful and more understanding. that is the obvious difference. a person who is 15, 16 years old is more ‘uzmanlaşmış’ [which means specialized in turkish, referring to the idea of being skilled, but here she means more prone] in making mistakes than, say, a 45-year-old. a 45-year-old [referring to her teacher at the mosque] can know more and is more thoughtful”. these examples make clear that for these youth, learning does not represent a process of making themselves independent from the community, pursuing a unique identity, or following a unique personal trajectory, but on the contrary, learning to become as one ‘ought to be’, to be like important others from the community and to return back value to the community. although these teens see learning as related to the explicit guidance of older generations and accept this guidance as learning, they did not exclude other ways of learning such as peer-learning or experimenting. however, these forms of learning were not foregrounded in their discourse in relation to learning or did not always count or were recognized as learning. 4.4. interest-based activities and networks? the point that turkish-dutch teens hold up collective identities to describe themselves, and do not use identifications that point to an autonomous, unique and individualized self as much, was also clear from their narratives about specific interests or activities that would typify them. hobbies, individual habits or an exclusive personal expertise were rarely mentioned. in most cases, these hobbies would represent more generally appreciated activities for boys or for girls, such as fighting sports for boys and fashion for girls. when asked ‘which activities do you strive to get better at’, boys often responded with sports, games and sometimes also interests such as cars, computers, planes/flying. turkish-dutch boys were mostly keen participants in sports, specifically football and martial arts (e.g., karate, kendo, boxing). they practised these sports often in sport-schools or sport-clubs on an amateur or semi-professional level. sport practices represented relatively unique and individualized learning spaces for them, which was also evident from their social network pictures. for instance, ahmet (16, m), a goal-keeper, stated: “i learn a lot from football … i learn a lot about how i should move, there is a lot of interaction (between coach and other keepers), and we learn to make decisions and logical thinking, especially logical thinking”. ahmet’s network picture (see figure 3) illustrates that sports and gaming are in fact personal spaces that are relatively independent from the rest of his mainly family-based, densely connected network. ahmet’s interest in sports is fostered through two contacts represented in this part of his network: a friend with whom he plays the online game “online soccer manager” and only talks about football-related issues, and his football coach. girls, on the other hand, found informal learning interests rather difficult to pinpoint, but most of them expressed their interest in fashion and spending time together with friends. in contrast to the boys’ enthusiasm for sports, there was very little attention to sports from girls. none of the female participants were actively doing any sports at the time of the interview. additionally, there were no other overlapping interests between boys and girls. for girls, it seemed the social aspect of any given interest was more central than gaining expertise in their field of interest, such as improving their ‘eye for fashion’. in other words, they were ‘just’ interested in fashion because they enjoyed the social side of consulting each other about clothes. thus, we found that for these youth, ‘interests’ were more generally appreciated activities and were not seen as personal. boys sometimes developed relatively unique interest-based networks, mostly related to sports or online gaming, while girls were reluctant to recognize interest-based learning in their favourite activities. 4.5. access to digital media as transformative potential as already illustrated in the example of ahmet in 4.4, turkish-dutch youth used the internet to support their offline interests or activities, such as sport, school or music preferences. for instance, as turkish-dutch boys often were engaged in fight-sports such as karate, taekwondo and boxing, most of these boys also visited youtube to watch fragments of fight choreography (e.g., bruce lee movies), fighting tournaments or street-fight videos. they often searched for information that often would be hosted in turkey or have content related to turkey. for instance, boys who are interested in football would search turkish websites about football, such as fanatik (a turkish sports (online-)newspaper), or they would visit websites that stream turkish television series. however, this media content based in turkey would be shared and discussed in social networks that are transnational and consist of social contacts both based in the netherlands and abroad, mostly in turkey, but also in the turkish diaspora. this is clear from the example of ceylin (16, f) (see figure 4, which shows her social network). in the interview, ceylin mentions how a combination of media network resources has helped her to think more consciously and critically about the social position, rights and demands of ethnic minorities. she explains how she learned that a television series (behzat c., a crime-detective television series) in turkey was cancelled due to, among other issues, bringing up the issue of education in kurdish for kurdish people. the news of cancellation combined with what she knew about kurdish people in turkey through her personal transnational social network triggered the conversation. she explains that her cousin, one of the transnational contacts in her network, informed her that in turkey, he observed that kurdish people were living comfortably similar to how they do and did not have a lower status or have lesser means to maintain their lives “[when her cousin was in izmir] he said that he saw kurdish people, and they were all very rich…those who live in the cities especially are powerful people and have all the means….”. through her transnational social network, she is made aware that kurdish minority status in turkey is not necessarily a problematic one regarding economic means and that they have consumer patterns she can also identify with. see figure 4, which shows 3 of her cousins that inform ceylin about ‘new’ places in turkey she does not know yet and give access to knowledge regarding kurdish minorities in turkey, among other issues. the perspective she gained about kurdish people through her nephew enabled her to also identify with and see them (also) as minorities. through the information provided via her online social network, she started to see the kurdish as minorities similar to herself: “well kurdish people should have their rights [particularly referring to the right of education in native language, which was an issue in the crime-detective television series], [...] i’m here [in the netherlands] a turkish person”. she continues her comparison of her own situation as a turkish minority in the netherlands with the situation of kurdish people in turkey. the combination of watching this turkish television series, hearing the news of cancelling, and having online contact with family in turkey, who have contact with kurdish people in turkey, enabled her to compare the situation of minorities in turkey and in the netherlands. through these different resources and sometimes conflicting stories from these resources, ceylin has learned to see the complexities of minority status and of political rights, including her own situation and that of the kurdish people. figure 4. network of ceylin. 4.6. access to digital media as network boundaries in addition to tapping into the content issued in turkey, turkish-dutch immigrant youth use media specifically tailored for turkish-dutch immigrants. the following example shows how their media use also develops along ethnic lines and social networks and marks divides between turkish-dutch immigrants and their dutch classmates. the first author asks ceylin (16, f) about a radio app for turkish-dutch immigrants. “i: what is taksim fm? turkish music? c: yes. taksim.fm is a radio channel in the netherlands made by turkish people. but it’s turkish, look, [she turns on the radio (on her phone)] but it’s not only music, it’s talk-shows, and they have a website. i don’t remember if there’s a mehmet akif (dj). (…) they talk about the dutch and the news here [meaning in the netherlands] but also about turkey and other stuff. it’s [focused] specifically [on] the things that are interesting for the turkish-dutch young people.” ceylin continues to explain how different media and apps are utilized differently for different ethnic groups. she says: “the dutch people of my age wouldn’t know taksim fm […] one main difference between dutch people and me is that i speak both turkish and dutch and a little english sometimes like “i love you” (giggles). turkish people also write out accents, you know, like the laz messages [on whatsapp], so that’s different with us [referring to her turkish-dutch friends]. but (with) my dutch classmates, well, we use it [whatsapp] for school stuff, because, well, we’re friends but not the best friends, and school is our only common subject. they are not part of my other daily life”. this example shows how specific media applications, media content, language used, and even typography are network-dependent. while texting with her turkish-dutch friends, ceylin uses the turkish language or specific typographical codes associated with the laz language to joke or tune in to themes specifically interesting for turkish-dutch immigrant youth. another example of such a divide is experienced by fatos (15, f). she is a fan of a turkish actor in her favourite drama-series little secrets (turkish: küçük sırlar). she explains “i love the internet. we have satellite tv to watch turkish channels at home, and i watch television there, but if i’m not at home or if i don’t have time at the time of the show, then i stream it from the internet…. she mentions a list of series and talk-shows she follows; when asked which one she likes most, she says: cetin, from küçük sırlar. i’m his fan. so are my friends […]. fatos tells that after school she spends much time chatting with her friends, and one of their favourite subjects is what happens in the series [küçük sırlar], for instance, how the characters dress up and about their expensive lifestyle in istanbul. she also seeks information and other related content (e.g., photos, news) regarding the series and the actor she likes and shares this on her social media account (hyves, a dutch social networking platform active between 2004-2013), where she reports to have approximately 400 contacts. in addition to its entertainment value, this series enables these girls (fatos and her other turkish-dutch friends) a window into life in turkey. however, she is sharing this interest exclusively with her turkish-dutch friends. fatos tells us how she cannot share this topic with one of her best friends, m., and how this creates a boundary between them. the access to the show through satellite tv and the internet creates an information divide between m., who is dutch and who she considers one of her best friends, and her friends who have access to the show. what both of these examples show is that digital (mobile)communication also mediates and re-informs specific network divides. therefore, next to media resources and the networks associated with them, which provide unique learning opportunities in the form of distinctive opportunities to gain information, form opinions, discuss positions and gain new insights, these youth also create clear boundaries in their social networks, which cut them off from other opportunities to learn and socialize. 5. discussion the results show how turkish youth create their own version of a global learner, based on notions of the self and of becoming that are primarily relational and oriented towards the collective. furthermore, afforded by technology, these youth create unique networked relationships for their learning, which are relatively closed for outsiders and are organized around the collectivities of the family and ethnically informed networks. at the same time, in the diaspora, their (transnational) networks are slowly becoming more diverse and fragmented. through contact with different migrant communities, settled in different countries, they are confronted with multiple versions of the ideal self as well as with diversification of socialization ideals and practices. although the network configurations of the participants sometimes reflect existing traditional (e.g., gender-based) boundaries, their networks also provide novel learning opportunities. unique trans-local experiences and corresponding means of reflection are mediated by a combination of the specific configuration of their (transnational) social networks, access to technology and media content. in this discussion, the main point we want to address is how these teens form a specific kind of global learner, which is not covered in all respects by recent prototypical models for learning in the digital era, such as in the concept of the connected learner. before doing so, we first discuss the other main point we want to bring under the attention, namely how our network analysis approach has enabled us to reach the goal of this study: to provide a critical reconsideration of “idealized digital connectivities for learning”. 5.1. how network analysis approach has enabled us to reach the goal of this study ego-network analysis as we have applied in this study, combining the gathering of social network data with in-depth interviewing, is especially adept for exploring the interaction between social structures and certain qualities or processes assigned to individuals and how these influence and shape each other (crossley, et al. 2015). in our case, we were able to map the specific social relationships that youth employ for their learning and understand their experience and perception of the “learning self” in relation to the structural and compositional aspects of their personal communities. in other words, this methodology allowed us to study learning as a networked phenomenon. as such, the approach was particularly useful to comment on models of learning that put connectivity up front. given that with this methodology, we can map the particularity of the connectivities of these youth empirically, it is suited to relate this conceptual work with the empirical record. the combination between quantitative analyses and interpretative work also enables the analysis of underlying paradigms associated with models of connectivity, such as the idea of an autonomous, independent self that is at the centre of the connectivity. moreover, the methodology allows us to study more precisely than with, for instance, interview studies, how learning opportunities and identities are created by, and vice versa create, social capital. ego-network analysis is up to now only used by a limited number of studies to study learning. we hope this study contributes to showing the potential of this approach for the study of learning. 5.2. how these teens form a specific kind of global learner as we have argued before (de haan et al., 2014) the prototypical image of a so called ‘connected learner’ as implied in the connected learning project resonates with a learner that is “highly agentic, driven by individual needs and interests, and pursues his or her learning in individualized and tailored-to-the-need networks”(p. 510). it is important to be specific here to what of the connected learning project we direct our critique. we argue that the initial ideal of highly engaged learners that seek out (online) connections to fulfil their individual interests is itself a culturally informed particular image of a learner. we do not direct our critique to the educational ideal of the project of connected learning which has been presented as an ideal of learning in the global society for all, and in opposition to and as an alternative for outdated notions and practices of learning and education (ito et al., 2013). we think that although connected learning is presented as an inclusive project in which individual learners are stimulated to connect to peers and other collectivities, there is not enough attention for how it was inspired initially by an individualistic learner ideal, based on the idea of unique preferences, networking efforts and independence in formulating their knowledge interests. the turkish-dutch teens in this study are well-connected learners, but they diverge from the ideal implied in this prototypical learner in several critical ways. first, there is very little emphasis on individuality among this group. these teens underline the interdependency and connectedness within their family and community much more than they bring up individual characteristics or interests. the driving force for these teens seems to be establishing interdependence with the family and the turkish(-immigrant) community. second, the learning experiences of turkish-dutch teens can be characterized as conformist or traditional in the sense that they appreciate the guidance from their parents to lead them to what is considered the key values and virtues of their community. to be a good person, it is important to act according to the norms of this community. the role models for a ‘good person’ are often those people who are respected within the family. in this regard, these teens also diverge from the image of a teenager in western middle-class families more generally, who puts less stress on relatedness with their family and much more on their individual agency (kağıtçıbaşı, 2005). furthermore, our data showed that the socialization and learning experiences of these teens are defined by relatively (ethnically) homogeneous, closed and dense social networks and that this is also partly the case for their online networks. in these communities, which are now also extended to the online world, the passing on of traditional values, family bonds, and hierarchical relationships, a focus on the collective and strong gender divisions remain important. this part of our data is in line with the image that was provided in the literature on turkish-dutch immigrant populations and that depicts this group as a relatively gender-segregated community (vedder, 2005), with a strong attachment to turkey (verkuyten, 2001) and a “fear of ‘dutchification’ of their children” (lindo, 2000, p.221). this part of our results would imply that, even given their access to online media, these youth’s learning ecologies seem rather stable and closed towards new influences, which is rather atypical for learning in migration (de haan, 2011). however, our study also revealed that their ncl undergo important changes, related to new possibilities provided by digital media. our results partly confirm earlier studies that the turkish community’s media use is geared towards content from turkey and that turkish-dutch immigrant youth’s online activities are also geared towards turkey in terms of the language used, reference to turkish culture or identity (d’haenens, 2003). this was evident from how turkish-dutch youth ‘plugged in’ media content in their networks that came from media channels based in turkey directed at the turkish community. nevertheless, our study also shows how tendencies described by the phenomenon “networked individualism” (rainie & wellman, 2012) impact these youth. looking at where the online contacts of turkish-dutch youth are located geographically, it was clear that they connect online with people who live relatively close by (in their neighbourhoods and in their cities). however, technology was also used to build networks across spaces, and their networks were defined by particular local-global dynamics. digital media allows the learning of these turkish-dutch youth not only to reach towards turkey but also to other turkish diaspora countries in europe (e.g., germany, belgium, france). these networks provide them with important trans-local learning experiences in terms of access to different languages and life worlds, even if these happen within their extended families. moreover, our data has shown that mobility patterns between turkey and the netherlands allow media content to be reinterpreted in similar ways as milikowski (2000) has argued. as shown by the example of ceylin, through a constant comparison between contexts, particular media-based content is re-weighted and re-interpreted, which provides important new possibilities for learning. certainly, these youth are not only “connected migrants” (diminescu, 2008) who establish new belongings and associations while maintaining the connections with their root community; they are also ‘connected learners’. they use new technologies to create spaces for their learning, regardless of actual physical locations, and give form to new ways of being that allow them to be ‘here and there simultaneously’. this greatly expands their socialization and learning possibilities. the ‘being here and now simultaneously’ has been associated with the notion of deterritorialization and the possibility it allows to develop a critical position by authors such as braidotti (1994). rather than the detachment from particular places in a literal sense, it is the distantiation of conventions and the multi-perspectivity that is seen as enabling the development of a critical position in relation to the “canonical”. in a similar vein, the living within or moving between heterogeneous spaces as well as the need to take distance from the existing cultural paradigms while reconsidering and recreating them has been referred to as the migrant condition (papastergiadis, 2000). the data shows that the trans local social network configurations of these young migrants, also in combination with their mobility patterns, generated particular opportunities for deliberation and reflection, that are related to the particular kinds of both deterritorialization and connectivity these youth experience. 5.3 implications for practice we believe that picturing this kind of a-typical global and connected learner helps us to expand our ideas of what a global learner might be. seeking to uncover the one-sidedness of notions of ‘new’, 21st century learning helps to understand how some might be privileged while others are marginalized. on a more positive note, acknowledging diverse types of connected learners can help to proactively incorporate them as useful models of global learning (doerr, 2017). further, the particular form of the ncl utilized by these turkish-dutch youth might also involve the risk of growing up relatively isolated. therefore, we would plea for more attention be paid to these particular informal learning experiences within the (semi-)formal contexts of learning, such as schools, libraries or community centres. it is important for teens and educators alike to realize how these network configurations are playing a role in shaping who these teens are and how they shape their future opportunities. with this paper, we hope to have contributed to a critical reflection on the particularity of networked connectivities and their impact on the potential diversification of learning and socialization in our societies. with this, we align with the ideal implied in the educational connected learning project that seeks ways to expand patterns that have been found for what might be privileged learners to all learners. however, our contribution turns the way to work towards this ideal around. instead of starting with elite or privileged learner ideals, and expand these to larger populations, our mission is to first expand our knowledge and ideals of the ways in which youths can make use of the possibilities our digital societies offer. it is key that educators and practitioners are partner in this process and are aware of this variation in their work with both majority and minority students. keypoints new prototypical models for learning in the 21st century are grounded in particular culturally informed ideals. learner identities of turkish-dutch teens contradict the autonomous, individualistic learning self, implied in the ideal of the ‘connected learner’. unique networked (trans)national connectivities are formed in response to the interaction of digital affordances and the learners’ specific socio-cultural position. acknowledging diverse types of connected learners can help to proactively incorporate them as useful models of global learning. references arnseth, h. c., & silseth, k. (2013). tracing learning and identity across sites: tensions, connections and transformations in and between everyday and institutional practices. in o. erstad, & j. sefton-green (eds.), identity, community, and learning lives in the digital age (pp. 23-38). cambridge: cambridge university press. braidotti, r. (1994). nomadic subjects. new york: columbia university press. castells, m. (2007). communication, power and counter-power in the network society. international journal of communication, 1, 238-266. retrieved from https://ijoc.org/index.php/ijoc/article/view/104/47 brown, j. s., collins, a., & duguid, p. (1989). situated cognition and the culture of learning. educational researcher, 18, 32–42. coenen, e. (2001). 'word niet zoals wij': de veranderende betekenis van onderwijs bij turkse gezinnen in nederland. amsterdam: het spinhuis. cole, m. (1998). can cultural psychology help us think about diversity? mind, culture, and activity, 5, 291-304. doi:10.1207/s15327884mca0504_4. crossley, n., bellotti, e., edwards, g., everett, m. g., koskinen, j., & tranmer, m. (2015). social network analysis for ego-nets: social network analysis for actor-centred networks. london: sage. crul, m., & schneider, j. (2010). comparative integration context theory: participation and belonging in new diverse european cities. ethnic and racial studies, 33, 1249-1268. doi:10.1080/01419871003624068 de haan, m. (2011). immigrant learning. in k. symms gallagher, r.k. goodyear, d.j. brewer & r. rueda (eds.), urban education: a model for leadership and policy (pp. 328-341) (14 p.). new york: routledge de haan, m., leander, k., ünlüsoy, a., & prinsen, f. (2014). challenging ideals of connected learning: the networked configurations for learning of migrant youth in the netherlands. learning, media and technology, 39, 507-535. doi:10.1080/17439884.2014.964256 d'haenens, l. (2003). ict in multicultural society: the netherlands: a context for sound multiform media policy? international communication gazette, 65, 401-421. doi:10.1177/0016549203654006. hildreth, p. m., & kimble, c. (eds.). (2004). knowledge networks: innovation through communities of practice. igi global. diminescu, d. (2008). the connected migrant: an epistemological manifesto. social science information, 47, 565-579. doi:10.1177/0539018408096447 doerr, n. m. (2017). phantasmagoria of the global learner: unlikely global learners and the hierarchy of learning. learning and teaching, 10, 58–82. doi:10.3167/latiss.2017.100206 farrell, l. (2006). making knowledge wor: literacy & knowledge at work. new york: peter lang. gee, j. p. (2005). semiotic social spaces and affinity spaces: from the age of mythology to today's schools. in d. barton, & k. tusting (eds.), beyond communities of practice: language, power and social context (pp. 214--232). cambridge: cambridge university press. gibson, k., rimmington, g., & landwehr-brown, m. (2008). developing global awareness and responsible world citizenship with global learning. roeper review, 30, 11-23. doi:10.1080/02783190701836270. gonzalez, n., & moll, l. c. (2002). cruzando el puente: building bridges to funds of knowledge. educational policy, 16, 623–641. hirzalla, f., de haan, m., & ünlüsoy, a. (2011). new media use among youth in migration: a survey-based account. wired up technical research report. utrecht university. retrieved from http://www.uu.nl/wiredup/publications.htm ito, m., gutiérrez, k., livingstone, s., penuel, b., rhodes, j., salen, k., ... watkins s. c. (2013). connected learning: an agenda for research and design. digital media and learning research hub: irvine, usa. kağıtçıbaşı, ç. (2005). autonomy and relatedness in cultural context: implications for self and family. journal of cross-cultural psychology, 36, 403-422. kumpulainen, k., & sefton-green, j. (2014). what is connected learning and how to research it? international journal of learning and media, 4, 7-18. doi:10.1162/ijlm_a_00091 lam, w. (2009). literacy and learning across transnational online spaces. e-learning, 6, 303-324. doi:10.2304/elea.2009.6.4.303 lam, w. (2014). literacy and capital in immigrant youths' online networks across countries. learning, media and technology, 39, 488-506. doi:10.1080/17439884.2014.942665 lave, j., & wenger, e. (1991). situated learning: legitimate peripheral participation. cambridge university press: u.k. lindo, f. (2000). does culture explain? understanding differences in school attainment between iberian and turkish youth in the netherlands. in h. vermeulen & j. perlmann, immigrants, schooling and social mobility (1st ed., pp. 206-224). houndmills: macmillan press ltd. messina dahlberg, g., & bagga-gupta, s. (2014) understanding glocal learning spaces: an empirical study of languaging and transmigrant positions in the virtual classroom. learning, media & technology, 39,468-487. doi:10.1080/17439884.2014.931868 milikowski, m. (2000). exploring a model of de-ethnicization. european journal of communication, 15, 443-468. doi:10.1177/0267323100015004001 moje, e. b., & luke, a. (2009). literacy and identity: examining the metaphors in history and contemporary research. reading research quarterly, 44, 415-437. morsunbul, u. crocetti, e., cok, f., & meeus, w. (2016). identity statuses and psychosocial functioning in turkish youth: a person-centered approach, journal of adolescence, 47, 145-155. doi:10.1016/j.adolescence.2015.09.001 o'donnell, e., lawless, s., sharp, m., & wade, v. (2015). a review of personalised e-learning: towards supporting learner diversity. international journal of distance education technologies, 13, 22-47. doi:10.4018/ijdet.2015010102 papastergiadis, n. (2000). the turbulence of migration: globalisation, de-territorialisation, hybridity. blackwell: oxford. phalet, k. & hagendoorn, l. (1996). personal adjustment to acculturative transitions: the turkish experience. international journal of psychology, 31, 131-144. doi:10.1080/002075996401142 phalet, k., & schönpflug, u. (2001). intergenerational transmission of collectivism and achievement values in two acculturation contexts: the case of turkish families in germany and turkish and moroccan families in the netherlands. journal of cross-cultural psychology, 32, 186–201. doi:10.1177/0022022101032002006 rainie, h., & wellman, b. (2012). networked. cambridge, mass.: mit press. rogoff, b. (2003). the cultural nature of human development. oxford: oxford university press. schneider, j., crul, m., & van praag, l. (2014). upward mobility and questions of belonging in migrant families. new diversities, 16, 1–7. sinha, c. (1999). situated selves: learning to be a learner. in j. bliss, r. säljö & p. light (eds.), technological resources for learning (pp. 32-46). oxford: pergamon. ünlüsoy, a., de haan, m.j., leander, k. & völker, b.g.m. (2013). learning potential in youth’s online networks: a multilevel approach. computers and education, 69, 522-533. doi:10.1016/j.compedu.2013.06.007 vedder, p. & virta, e. (2005). language, ethnic identity, and the adaptation of turkish immigrant youth in the netherlands and sweden. international journal of intercultural relations, 29, 317-337. doi:10.1016/j.ijintrel.2005.05.006 verkuyten, m. (2001). global self-esteem, ethnic self-esteem, and family integrity: turkish and dutch early adolescents in the netherlands. international journal of behavioral development, 25, 357-366. doi:10.1080/01650250042000339 microsoft word sasse_publication_approved.docx frontline learning research vol. 9 no. 3 (2021) 31-51 issn 2295-3159 info corresponding author: heide sasse, august-croissant-str., 5, 76829, landau, germany, sasse@uni-landau.de doi: https://doi.org/10.14786/flr.v9i3.723 capturing primary school students’ emotional responses with a sensor wristband heide sasse & miriam leuchter1 1institute for children and youth education, educational sciences, university of koblenz and landau, landau, germany article received 3 september 2020 / article revised 10 january 2021 / accepted 4 march / available online 25 may abstract the emotions experienced by primary school students have both positive and negative effects on learning processes. thus, to better understand learning processes, research should consider emotions during class. standard survey-based methods, such as selfreports, are limited in terms of capturing the detailed trajectories of primary school children’s emotions, as their abilities of self-reporting are developing and still limited. emotions can also be tracked by capturing emotional responses as they occur e.g. from physiological reaction measured with sensor wristbands. this technology generates an emotional responses typology based on continuously captured physiological data, such as skin conductivity and skin temperature. however, such measurement methods need to be validated before being used. the present study thus attempted to validate this instrument with primary school students. we used the bm sensor wristband technology, as its emotional response typology is based on the categorical emotion and homeostasis approach. in our research, we focus on the emotional responses that can be distinguished by the bm typology and that can influence learning processes. these emotional responses are: “joy”, “curiosity”, “attention”, “fear”, “anger” and “passivity”. therefore, we induced emotional responses in primary school children through specifically developed audio-visual stimuli. using logistic mixed effects modelling, we investigated the occurrence of opposing reactions. we observed that primary school children’s reactions to audio-visual stimuli could be differentiated. we conclude that primary school children’s emotional responses, such as “joy”, “curiosity”, “attention”, “fear”, “anger” and “passivity”, can be accurately measured by evaluating physiological data. keywords: emotional responses; primary school children; physiological data; sensor wristband sasse et leuchter 32 | flr 1. introduction emotions are a foundation for cognitive resources, learning strategies, self-regulation, learning performance and learning motivation, thus, emotional experiences are closely related to learning processes (ahmed et al., 2013; tyng et al., 2017). emotions relevant to the generation of knowledge (epistemic emotions) and learning (achievement emotions) include enjoyment, anxiety, anger, boredom and curiosity (pekrun & stephens, 2012; pekrun & linnenbrink-garcia, 2012). although, according to pekrun (2017), attention is not considered as an emotion, it is crucial for fundamental processes of engagement in learning (steinmayr et al., 2010; posner & rothbart, 2005). emotions are mostly surveyed using questionnaires, such as standardized self-report scales (pekrun & bühner, 2014). this measurement approach, which has been used successfully, seems to have its limits when it comes to capturing emotions at the moment they emerge. first, students who are to report on their overall perceived emotional experiences evaluate their emotions retrospectively. in these retrospective analyses, subjects are asked to remember a situation in which they experienced these emotions. such sources are submitted to selective memory. events and considerations that occur between the originally experienced emotion and the remembered emotion can lead to a differing evaluation (moosbrugger & kelava, 2012). in addition, participants often state the sum of their experiences, although fluctuations may occur during the process itself (brandstätter et al., 2018). thus, it is difficult to capture emotions in detail in the course of learning processes (d´mello et al., 2007; lohrmann, 2008). moreover, capturing primary school students’ emotions during learning must take into account children’s limited working memory. hence, it may be even more difficult to measure and classify emotional experiences by classical survey methods, such as questionnaires (turner & trucano, 2014). consequently, we must find a way to capture emotional experiences not only retrospectively but additionally also during learning in order to analyse the fluctuation of emotions more closely (linnenbrink-garcia et al., 2016). therefore, we could use measuring devices that go beyond self-reports (järvela et al., 2010; paris & turner, 1994; winne & perry, 2000). recent emotion research has developed instruments that offer a solution for measuring emotional reactions using physiological data in real time without selfreport bias. various manufacturers offer devices for tracking physiological data, with the possibility to use this information for the recognition of emotional reactions. although these analyses always remain interpretive, such evaluations could contribute to emotion analysis. järvenoja et al. (2018) captured skin conductivity using the empatica e4 wristband (empatica inc., cambridge, ma) to analyse the intensity of emotional responses in cooperative learning situations. the empatica device is focused on providing physiological parameters like skin conductivity or heart rate to indicate arousal changes. however, it does not provide indicators for the interpretation of physiological data in emotional terms like fear or joy, relevant in school achievement research. another device that captures emotional responses through physiological reactions is the sensor wristband from bodymonitor. in contrast to the empatica e4 wristband, the bm sensor wristband offers, according to the manufacturer, an emotion classification of physiological data. thus, we examine emotions that are part of the bm repertoire, which allows to identify high arousal responses such as appetitive arousal (“joy”/“curiosity”) and aversive arousal (“fear”/“anger”), as well as low arousal responses such as deactivation (“passivity”) and homeostasis (“attention”). in addition, the wristband can be adjusted to different wrist sizes of children. in the following, we examine the theoretical foundation of these emotions. in emotion research, two dominant approaches describe the generation of emotions, namely the categorical (ekman, 1999; panksepp, 1982) and the dimensional theory of emotions (barrett, 2006; pekrun, et al., 2014; russell, 1980; scherer, 1999). both approaches see emotional reactions as a result of the brain's appraisal processes of external and internal stimuli. in the dimensional approach, the appraisal process leads to a subjective, positive or negative evaluation of the stimuli and thus, emotional labelling depends on subjective interpretation. sasse et leuchter 33 | flr in the categorical approach, it is assumed that there are separate basic appraisal-response sets (also called emotional responses) that have evolved as functional systems for survival in an evolutionary process. although there is no consensus on how many elementary emotional responses there are (von scheve, 2014), they are supposed to encompass a) high-arousal responses, which have appetitive and aversive aspects, and b) low-arousal responses (lee & lang, 2009; oxendine, 1970). a) high-arousal responses include emotions such as fear, joy, anger and surprise (piórkowska & wrobel, 2017; zillmann, 2008), which are also referred to as basic emotions according to ekman and cordaro (2011). b) according to levenson (2006), low arousal responses include homeostasis, where the body is generally alert and physiological responses are balanced, which can be understood as attention. furthermore, low arousal responses include deactivation, where the central nervous system withdraws when no other emotional responses occur for a period of time, this can be understood as passivity. there is consensus that emotional responses are automatic, continuous neural evaluative processes (izard, 1993; lang & bradley, 2010; ledoux, 1998; ortony & turner, 1990), that unfold in multiple dimensions, namely as changes in physiological data (skin conductivity, skin temperature, heart rate) and musculoskeletal systems (posture, gestures, facial expressions). moreover, these physiological responses are supposed to effect motivational tendencies and subjective feelings (levenson, 2003, 2014). changes in skin conductivity, skin temperature and heartrate can signify a wide variety of neurophysiological processes, such as the level of arousal (high and low) (kołodziej et al., 2019) and the quality of arousal (appetitive or aversive) (hayes et al., 2014). thus, their measurement may give insights into various emotions. the emotional responses that can be captured by the bm sensor wristband via physiological data play an important role in learning research. some students learn successfully, while others have difficulties developing and activating resources to be academically successful (pintrich, 2003) positive predictors of learning are joy, curiosity and attention. joy is one of the basic emotions that arise as a reaction to a pleasant situation (ekman, 1982). joy in learning can enable children to focus their attention on the task at hand and thus promote deep immersion in the learning content (roth, 2011; sturm, 2005). hascher & brandenberger (2018) examined the effectiveness of computer-assisted individualization in teaching mathematics on the performance and emotions of fifth-grade students. these researchers observed that learning progress triggered joy and that joy, in turn, had a positive effect on learning progress. curiosity can be defined as an intrinsic desire to acquire new knowledge and sensory information (engel, 2011). for example, eren & coskun (2016) have shown that curiosity, which reflects an “urge to know”, is significantly linked to university students' learning, commitment, performance goals, and knowledge acquisition and a deeper understanding of the knowledge content. finally, attention is a decisive factor for learning because it enables students to engage with learning content in the first place (brünken & seufert, 2006). jamet et al. (2008) showed that directing attention through the use of typographic highlighting facilitates the learning of vocabulary by undergraduate students. negative predictors of learning are fear, anger and passivity (eren & coskun, 2016; götz et al., 2007; mulryan, 1992). fear is one of the basic emotions that arises when a threat of harm appears, whether physical or psychological, real or imagined (ekman, 1982). in learning, fear can have a negative effect because it is associated with the use of rigid strategies that lose themselves in detail and do not promote creative, independent and holistic learning (götz et al., 2007). students report that of all negative emotions, they experience anxiety most often (pekrun et al., 2002), which is closely related to fear (horwitz, 2013). eysenck et al. (2005) showed that undergraduate students with high anxiety levels had significantly poorer access to their memory systems for task processing and made significantly more mistakes than the group with low anxiety levels. anger is one of the basic emotions that arises when we are prevented from pursuing a goal and/or treated unfairly (ekman, 1982). anger may arise whenever obstacles to learning occur that students perceive as unnecessary or arbitrary, which may be the case if performance requirements are perceived to be too high or if performance assessments are not comprehensible. to summarize, anger in learning can lead to task-irrelevant thinking and hinder interest, intrinsic motivation and self-regulated learning (pekrun, 2018). in a survey with undergraduate students, sasse et leuchter 34 | flr pekrun et al. (2011) showed that anger correlated negatively with intrinsic motivation, elaboration and self-regulation. furthermore, anger correlated negatively with students' overall self-reported learning efforts and academic achievement. finally, passivity is related to off-task behaviour in learning and thus hinders learning (riley et al., 2011). mulryan (1992) showed, in a study of cooperative small group teaching in mathematics with primary school children, that especially low-performing children showed a high degree of passive behaviour in class. according to (sidelinger & booth-butterfield, 2010) engagement in class should be encouraged instead of passivity, in order to increase students' learning success. thus, exploring emotional responses such as appetitive arousal (“joy”/“curiosity”) and aversive arousal (“fear”/“anger”), as well as deactivation (“passivity”) and homeostasis (“attention”) with the bm sensor wristband might give valuable insights into students’ emotions during learning. however, the use of sensor wristbands might be viewed critically (endedijk et al., 2018). this is particularly important if learning processes are to be analysed in authentic environments. for example, teacher`s praise might be a more or less positive stimulus for students depending on teacher, task, social context, personal thoughts, peer behaviour and physical needs. therefore, it is necessary to determine what triggers emotional responses in order to analyse physiological reactions. moreover, movement of wrists or arms might drown the measurement of physiological reactions with young students who tend to move more intensely than adults. since the measurement is based on the function of the sweat glands, physiological differences between adults and children may occur (falk, 1998). accordingly, such devices should be tested in advance with regard to their use and the transferability of the measurement procedure from adults to young children. additionally, testing such devices are a first step towards their implementation in authentic learning settings. therefore, we conducted a study in which we validated the use of the bm sensor wristband on primary school children. 2. rationale of the study in our study, we have targeted at evoking the emotional responses "curiosity", "passivity", "attention", "anger", "joy" and "fear" with different stimuli. we coupled "curiosity" with "passivity", "attention" with "anger" and "joy" with "fear" according to the manufacturer's suggestion in order to check whether the probability of one of the emotional responses occurring during the respective stimulus is higher than during a reference stimulus. furthermore, these comparisons are also be supported by literature. studies on learning situations have shown that learning outcomes are influenced by curiosity and passivity. curiosity promotes learning performance and knowledge exploration (markey & loewenstein, 2014; krapp & prenzel, 2011). passivity, however, prevents a focus on the current learning situation (lewalter & schreyer, 2000; mulryan, 1992). to be able to process learning tasks, attention is essential (lombrozo, 2016; roth, 2011; rösler et al., 2009), but anger can also develop. whenever students feel anger, their attention may decrease (mees, 2020; siyez, 2017). joy and fear occur regularly in everyday school life. götz et al. (2004) showed that students experience different levels of fear and joy during learning. while the performance expectations of some students are positive (joy), other students are afraid of failure (fear) (lohrmann et al., 2011). sasse et leuchter 35 | flr 3. research question can primary school children’s emotional responses be accurately measured using physiological data with the bm sensor wristband? we hypothesize that when comparing two contrasting emotional responses, the probability of occurrence of the induced response is higher than in the non-induced response. 4. method and data 4.1. bm sensor wristband and bm emotion classification based on the categorial emotion theory (levenson, 2011; panksepp & watt, 2011), the bm technology generates a typology of six different emotional responses based on continuously captured physiological data, such as skin conductivity and skin temperature, via a sensor wristband (figure 1). the emotional response typology categories (table 1) include appetitive responses: “joy”, “curiosity” and aversive responses: “fear”, “anger”), homeostasis (“attention”) and deactivation (“passivity”). appetitive and aversive responses (“joy”, “curiosity”, “fear” and “anger”) are typically short-term reactions (ekman, 1999, martinent et al., 2012); therefore, data on the occurrence (“moments”) can be recorded. the homeostasis and deactivation responses (“attention” and “passivity”) are sequential episodes that can be classified after a certain period of time (jänig, 2008). therefore, for “attention” and “passivity” the duration (“periods”) must also be recorded. figure 1. bm sensor wristband. sasse et leuchter 36 | flr table 1 bm typology of emotional responses. sca/stb based emotion response emotional response bm classification perceived as appetitive responses joy curiosity, surprise “joy” “curiosity” “moments” aversive responses fear anger, stress “fear” “anger” “moments” homeostasis attention, vigilance, wellbeing, hedonic pleasure “attention” “moments” and “periods” deactivation passivity, disinterest, shutoff, mental withdrawal, tired, deeply relaxed “passivity” “moments” and “periods” notes. a skin conductance. b skin temperature. 4.2. stimuli validated stimuli, such as those defined by the international affective digitized sounds (iads2) and the international affective picture system (iaps) and the geneva affective picture database (gaped), could not be employed because these stimuli were developed by using responses of the adult population. since the participants in this study were primary school students, the stimuli had to be adapted to their age and stage of development (see table 2). the adapted stimuli were evaluated by 25 adult experts. each expert was presented with a selection of various stimuli on a screen in a quiet room. each example was then rated for its supposed effect on emotional responses in children of primary school age. in addition, 12 children (m = 8.9, sd = 1.34, 67% female) were interviewed regarding their occurring emotions when monitoring the stimuli. in this way, their physical reactions, such as facial expressions and gestures, could also be evaluated. this selection procedure resulted in the stimuli listed in table 2. finding stimuli for anger and fear was particularly challenging. to provoke anger, we had a stimulus in which the children were promised a gift that they ended up not receiving. however, this led to sadness and disappointment rather than anger. it was only when we inserted the repetitive counting task that the students reported perceived anger. to provoke fear, we told the children a scary story with spooky pictures. however, this story generated many laughs and thus rather joy instead of fear. thus, we developed a video clip showing shere khan attacking mowgli and evaluated the reaction by the 25 adult experts and the 12 children. based on their feedback, we decided to use this fear stimulus instead. sasse et leuchter 37 | flr table 2 overview of the stimuli 4.3. sample this study used repetitive measurements and was initially conducted with children of the first to fourth grade children, which corresponds to the four years of primary school in germany (sample 1). subsequently, as a validation and specification of these results, we repeated the study with third and fourth grade children (sample 2). standard procedures were used for recruiting participants: written informed consent, approved by the university institutional review board and the school board, was obtained from the parents of each participant. students were drawn from 6 primary schools in rural germany. participation in the study was voluntary, and the students were able to revoke their consent at any time during the course. bm classification stimuli content and duration “joy” picture and sound of mewling kittens (13 sec) “curiosity” video: magic trick (34 sec) video: magic trick resolution (39 sec) “attention” watching a car draw-by-numbers video: (25 sec) “fear” video: shere khan attacking mowgli (26 sec) “anger” repeated counting task: count correctly to ten (and you will get a treat), feedback: something was wrong “passivity” watching a picture of a light bulb (20 sec) sasse et leuchter 38 | flr sample 1 consisted of 113 children between 5.2 and 11.4 years of age (m = 8.7, sd = 1.12, 58% female). sample 2 consisted of 164 children between 8.1 and 11.2 years of age (m = 9.9, sd = .66, 56.7% female). 4.4. procedure the study was conducted in a classroom in the primary school that the children attended. the students’ desks and chairs were arranged in rows. two workplaces that were separated through blinds were assembled on each table, so that each child was as undisturbed as possible. each workplace consisted of a tablet pc, a bm sensor wristband, and headphones. a group of maximum 10 students each were welcomed into the classroom and assigned to a workplace. next, the sensor wristbands were strapped on. the children were told that they would be playing a computer game with different tasks and that they would receive candy as a reward. before the game was started, the teacher explained the different tasks and ensured that everyone understood. it was also made clear that when the children were finished, they would have to remain quietly at their places until everyone had finished. since the school staff was informed about the implementation of the study, the survey was not disturbed from outside. if a child needed help, the child was spoken to very quietly at the workplace as not to disturb others. this was noted in the study protocol. if a child needed substantial help (more than once), the data was not included in the study (sample 1: 2 cases; sample 2: 0 cases). to start, the children put on headphones and opened the app on the tablet pc. although the children worked in parallel, each worked at their own pace according to how quickly they completed each task. the stimuli were presented in the following order to mix the expected positive and negative emotional responses: 1. “curiosity” (magic trick) 2. “passivity” (light bulb) 3. “fear” (shere khan is attacking mowgli) 4. “joy” (mewling kittens) 5. “anger” (repeated counting) 6. “attention” (car draw-by-numbers) moreover, each stimulus was followed by a flower/balloon/fish draw-by-numbers video with the aim of avoiding spill-over effects over consecutive emotional responses and to re-establish attention and the baseline. according to nakasone et al. (2005), acquiring and analysing physiological data always faces the so-called baseline problem. the baseline problem implies that it is difficult to identify a response (the baseline) to which physiological changes can be compared. therefore, the event-related skin conductance responses (scr values) were first transformed into phi values using the transformation procedure for area correction (lykken, 1972) in order to compare the responses of different subjects. afterwards, the current scr reaction intensity was calculated. these express the average increase in skin conductivity per unit of time (i.e. per second). the data could then be reduced to response peaks (local maxima) in an individual response curve. 4.5. physiological measurement the bm sensor wristband acquired the physiological data. the original wristband has been adapted for usage with children (figure 2). the overall size of the wristband was reduced, as it was originally developed to fit adults’ wrists. in addition, the length of the wristband can be extended to fit wrists of different sizes. a child-friendly fabric was employed for the complete wristband (textile housing and extension). by wearing the wristband, the following parameters were captured at a 10-hz rate: skin conductance, skin temperature, ambient temperature, 3-dimensional acceleration, and pressure force on electrodes. this raw data was stored locally on the sd card of the wristband. sasse et leuchter 39 | flr figure 2. child-compatible bm sensor wristband. 4.6. data preparation by using two different kinds of electronic devices (tablet pc and sensor wristband), two types of data was generated: a) subjective choice data via normal computer-human interaction of the developed application. for each child, the interaction with the tablet pc application was automatically protocolled in a data matrix with one line for each action. each action was time stamped. b) physiological parameter data in tab-delimited ascii format. as the capture rate was 10 hz, the data matrix consisted of one line for each reading, which means that one second of data collection led to 10 observation lines. as children wore the wristband for 45 minutes, one block of emotional response induction was recorded for a duration of 45 minutes, which resulted in 45*60*10 = 27 000 observations. after automated purge of changing contact quality related artefacts, skin conductance and skin temperature were used to reach a classification according to the bm emotional responses classification system. bm automatically applied its algorithm and, using interval wise regressions, provided a dataset where raw physiological reaction data were replaced by ratio-scaled scores of emotional responses, such as “joy”, “curiosity”, fear” and “anger”, as well as binary information regarding “attention” and/or “passivity”, given on a moment-by-moment basis. however, due to the manufacturer, we are not able to present the algorithm. as both the tablet pc application data and wristband parameter data were time-stamped, data could then be merged via corelating timestamps. this approach enabled the analysis of the physiological data by intervals, which were characterized by their concurrent stimulus presentation. if technical errors occurred, for example, if there was an instruction or application error or time stamps were missing, these cases were excluded from the analyses (sample 1:4 cases; sample 2:13 cases). 4.7. methods of analyses to evaluate the validity of the bm emotional responses classification, statistical analyses were performed by contrasting the outcomes of emotional response classification for the inductive stimulus interval with the classification outcome for a reference stimulus interval. hence, we used the following contrasts: magic trick vs. light bulb, focusing on “curiosity” and “passivity” repeated counting task vs. car draw-by-numbers video, focusing on “anger” and “attention” shere khan attacking mowgli video clip vs. mewling kittens image, focusing on “fear” and “joy” sasse et leuchter 40 | flr for the emotional responses “joy”, “curiosity”, “fear” and “anger”, as well as “attention” and “passivity”, binary information regarding occurrence on a moment-by-moment basis was analysed. for “attention” and “passivity”, which were perceived as time periods, the durations of these episodes were analysed. the merged data matrix led to a hierarchically multilevel data set with subject as level 1 and measurement time as level 2. according to singer & willett (2003), this is adequate for this type of multilevel data structure, contrast effects were estimated by running mixed models. in the case of ratioscaled dependent variables (duration of episodes), linear regression mixed models were applied; in the case of binary occurrence of emotional responses, logistic regression mixed models were utilized. the dependent variable was the occurrence of an emotional response (yes/no) and the predictive variable was the respective stimulus. 5. results 5.1. descriptive statistics table 3 provides an overview of the occurrence rates of emotional responses during the selected stimuli that show that the occurrence of emotional responses is rather low with a large standard deviation. focusing on the coupling of the emotional responses suggested by the manufacturer of the bm sensor wristband, the descriptive data supports the examination of the contrasts between "curiosity" and "passivity", "attention" and "anger" as well as "joy" and "fear", as we found that the frequency and duration of each emotional response seems higher during the induced stimulus than during the reference stimulus. table 3 means (m) and standard deviations (sd) for the overall occurrence (moments) and durations of contrasted emotional responses during the selected stimuli stimuli emotional response sample 1 sample 2 sample 1 sample 2 sample 1 sample 2 m (sd) m (sd) m (sd) m (sd) m (sd) m (sd) moments of “anger” moments of “attention” duration of “attention” (sec.) car draw-bynumbers .023 (.170) .054 (.226) .629 (.483) .725 (.447) 15.005 (7.343) 11.767 (7.487) rep. counting .034 (.181) .114 (.318) .507 (.500) .655 (.475) 9.372 (3.605) 6.322 (3.409) moments of “curiosity” moments of “passivity” duration of “passivity” (sec.) light bulb .034 (.227) .048 (.224) .149 (.356) .132 (.339) 10.445 (6.703) 4.276 (3.512) magic trick .063 (.243) .053 (.213) .090 (.325) .061 (.326) 10.195 (7.803) 4.036 (3.161) moments of “fear” moments of “joy” mewling kittens .011 (.105) .031 (.175) .073 (.260) .061 (.240) shere khan .028 (.165) .094 (.291) .049 (.216) .069 (.270) sasse et leuchter 41 | flr 5.2. contrasts between magic trick and image of a light bulb: “curiosity” and “passivity” the first analysed stimuli were magic trick and light bulb (see table 4). we calculated the occurrence of emotional responses such as “curiosity” and “passivity” during the stimulus magic trick compared to the stimulus light bulb. the results showed for both samples, that the prevalence of moments for "curiosity" was higher and "passivity" was lower during the magic trick stimulus compared to the light bulb stimulus. the results for the duration of periods of “passivity” showed that “passivity” periods were shorter during the magic trick stimulus than during the light bulb stimulus. table 4 “curiosity” and “passivity” during watching a magic trick video compared to looking at an image of a light bulb bm emotional response classification prevalence of moments a duration of periods b odd ratio a (sd) b-coefficient b (sd) sample 1 c sample 2 d sample 1 c sample 2 d “curiosity” 1.853** (.211) 1.230* (.085) “passivity” .599*** (.043) .310*** (.047) -2.513*** (.063) -1.360*** (.161) notes. a logistic mixed effect model. b mixed effect model. c 1st 4th graders. d 3rd and 4th graders. *** p<.001, ** p<.01, * p<.05 5.3. contrasts between repeated counting task and car draw-by-numbers: “anger” and “attention” the second analysed stimuli were repeated counting and car draw-by-numbers (see table 5). we calculated the occurrence of emotional responses, such as “anger” and “attention”, during the stimulus repeated counting compared to the stimulus car draw-by-numbers. the results showed for both samples, that the prevalence of moments for "anger" was higher and "attention" was lower during the repeated counting stimulus compared to the car draw-by-numbers stimulus. the results for the duration of periods of “attention” showed that “attention” periods were shorter during the repeated counting stimulus than during the car draw-by-numbers. sasse et leuchter 42 | flr table 5 “anger” and “attention” during repeated counting compared with the viewing of numbers getting connected with strokes to form the picture of a car bm emotional response classification prevalence of moments a duration of periods b odd ratio a (sd) b-coefficient b (sd) sample 1 c sample 2 d sample 1 c sample 2 d “anger” 1.482* (.219) 2.34*** (.940) “attention” .532*** (.037) .63*** (.019) -57.29*** (3.462) -58.3*** (4.810) notes. a logistic mixed effect model. b mixed effect model. c 1st 4th graders. d 3rd and 4th graders. *** p<.001, * p<.05 5.4. contrasts between shere khan attacking mowgli and mewling kittens: “fear” and “joy” the third analysed stimulus was shere khan attacking mowgli and mewling kittens. we calculated the occurrence of emotional responses, such as “fear” and “joy”, during the stimulus shere khan attacking mowgli compared to the stimulus mewling kittens (see table 6). the results showed for sample 1 that the prevalence of moments for "fear" was higher and "joy" was lower during the shere khan attacking mowgli stimulus compared to the mewling kittens stimulus. in sample 2, the results pointed in the same direction, however, the contrast was not significant. “fear” and “joy" were not perceived as episodes; therefore, their duration was not analysed. sasse et leuchter 43 | flr table 6 “fear” and “joy” during watching the scene of shere khan attacking mowgli compared to looking at mewling kittens bm emotional response classification prevalence of moments a odd ratio a (sd) sample 1 b sample 2 c “fear” 2.801*** (.786) 3.25** (1.548) “joy” .394*** (.066) 1.154 (.353) notes. a logistic mixed effect model. b 1st 4th graders. c 3rd and 4th graders. *** p<.001, ** p<.01 6. discussion overall, our results indicate that emotional responses can be analysed validly based on physiological data such as skin conductivity and skin temperature in primary school children. thus, they confirm and expand the findings of bergner et al. (2013) and li et al. (2016), who examined this measurement method on adults. hence, the measurement of physiological data could contribute to a more accurate recording of emotions in young children, which has been highlighted as an important requirement for understanding emotional and motivational processes during learning (turner & trucano, 2014). young children’s developmental constraints, which have an impact on retrospective recording methods such as self-reports (d’mello & graesser, 2007; turner & trucano, 2014), could thus be avoided. therefore, by using the bm system, we hope to have identified a way to go beyond self-reports (järvenoja et al., 2018). the results showed that, the respective induced emotional responses occurred primarily during the corresponding stimulus, however, the values of the descriptive statistics are rather low. considering that these are mean values that also include subjects with a value of 0, i.e. who showed no reaction, the results are plausible. interpreting the results of the descriptive data, it was noticeable that both the induced and the non-induced emotional responses occur during the stimuli. this could be possible because a) spontaneous or non-specific changes also occur in the absence of an identifiable stimulus (boucsein, 2012), and b) stimuli enhance specific emotional responses, but these do not always occur exclusively (kreibig, 2014). accordingly, although our results indicate that the emotional responses occurred primarily during the corresponding stimulus, we cannot be sure that these emotional responses were exclusively based on the stimuli. in both samples, according to the contrasting analysis, emotional responses “curiosity”, “attention”, “fear”, “anger” and “passivity” were high or low, respectively. thus, we assume that our expectations on the emotional induction outcome of different stimuli were met. accordingly, in the light of the studies by kreibig (2010) and immordino-yang & christodoulou (2014), we managed to evaluate emotional responses via physiological data. however, although the expected fear-producing stimulus (shere khan attacking mowgli), induced more “fear” and less “joy” response moments than the joysasse et leuchter 44 | flr producing stimulus (mewling kittens), the value for "joy" was not significant in sample 2. the interpretation of the results of the non-significant value for "joy" could possibly be based on the different age composition. it could be possible that older children experienced similar amounts of “joy” during the two stimuli shere kan attacking mowgli and mewling kittens. possibly, the older children were aware that the film scene was fictional and therefore not really dangerous but also funny. to better investigate the “joy” responses compared to the “fear” responses, a different stimulus should be chosen for older students in future studies. in sum, capturing physiological data and thus drawing conclusions regarding individual emotional experiences during learning can provide a more accurate understanding of the occurrence of different emotional responses and how these can influence learning. if it were possible in future studies to systematically and comprehensively analyse the dynamic interplay of emotional activities with regard to learning processes more precisely, we will be able to better understand the learning processes of children. this research would be especially beneficial if the learning processes of younger children are to be studied, as they still have limited working memory capacity, and their emotional responses are often difficult to capture through conventional self-reports (d’mello & graesser, 2007; turner & trucano, 2014). hence, the analysis of emotional responses during learning using physiological data may help to elucidate why some students learn successfully, while others have difficulties developing the resources to be academically successful (pintrich, 2003). the capture and analyses of emotional responses can provide initial insights into the frequency with which emotions occur during learning. our study can be understood as a first step towards measuring and understanding primary school children’s emotional responses using physiological data. in future studies, the bm system might enable us to investigate the occurrence of emotions and learning processes more closely, as stipulated by pintrich (2003). if we can capture emotional responses via physiological reactions in primary school-age children, we would be able to associate them with learning outcomes. achieving a better understanding of the wide variety of emotions and fundamental processes of engagement in learning that arise during children's learning processes, such as joy, curiosity, attention, fear, anger, or passivity, may help us design learning environments that are more suitable for primary school children and help them improve their learning (rodriguez et al., 2012; shen et al., 2009). thus, in a future study, we want to examine emotional responses during learning in more detail. this research may help to determine appropriate teaching methods to support children's learning processes according to their needs. 7. limitations capturing physiological data results in large, complex and context-specific data sets. the evaluation of the raw data is a considerable challenge (järvenoja et al., 2018). the handling of missing or blurred data using traditional statistical techniques, modern data mining and machine learning techniques has not yet been sufficiently examined on this type of data (azevedo, 2019). in addition, the classification of physiological data as emotional responses, especially in different contexts, is still insufficiently researched and have not been sufficiently empirically validated (harley, 2015). in order to complement standardized survey methods, it is essential to use measurement devices that reliably capture physiological data and meet research standards (e.g., measurement frequency or filtering of artefacts) (gravina et al., 2017). many commercially distributed devices are available, but we trusted the bm sensor wristband, as it has already been used and tested on adults in cooperation with the gesis leibniz institute for the social sciences (li et al., 2016; papastefanou, 2013) generating validated results. however, to achieve full validation of the sensor wristband, it would be necessary to perform the same procedure with other devices available on the market and compare them. besides, data are interpreted based on the algorithms of the devices (lohani et al., 2019). in this study, we note that we sasse et leuchter 45 | flr are not able to present the bm emotional response classification algorithm. this makes the evaluation of the results much more difficult than if it could be presented, which is why we also verified our results via a second sample. in addition, we have tried to be as careful as possible in selecting and evaluating the stimuli. the actual study was based on the measurement of the physiological data of the emotional responses of the students. however, tracking the facial expressions or gestures of the students during the confrontation with the stimuli would certainly have provided further valuable insights into the validity of the measurement. accordingly, our study is only a first step in validating the bm sensor wristband and further validation would be useful, for example by comparing occurring emotional responses with the observation of students' facial expressions and gestures in authentic learning situations. the findings of the present study rely on two samples that consist of children from first to fourth grade. for more specific applications, it might be worthwhile to analyse age-related subgroups separately to control differential stimulus effects. however, this approach would require a larger sample. with the second sample, we have attempted to take a first step in this direction. moreover, these analyses were group-based and findings that hold for groups cannot be generalized to within-person processes (hamaker, 2012). therefore, the classification system might not be reliable enough to be used at the level of the individual students and individual emotions. finally, the analysis was limited to six emotional responses and performed by comparing the physiological data of two contrasting emotional responses and two contrasting stimuli with each other. in addition, no counterbalance was created to the sequence of stimuli. however, students constantly experience a variety of emotions, not only the selected six emotional responses but also others such as hope, pride or shame which cannot be captured with the instrument. if such devices are to be used in authentic learning situations, it will be even more difficult to identify the stimuli that trigger emotional responses in the context of learning, as emotional responses can be based on any causal contexts, such as teacher, task, social context, personal thoughts, peer behaviour and physical needs which cannot be separated (lewis, 2000). moreover, the setup of the current study required a rather passive movement behaviour of the children. in authentic learning environments, children might move more, which could have an impact on skin conductance and thus emotional response data (endedijk et al., 2018). although the bm technology compensates movement artefacts by using simultaneously measured data on electrode contact quality and acceleration (nuñez et al., 2018; resch et al., 2015; zeile et al., 2015), further studies should consider this limitation. accordingly, we cannot assume with certainty that the instrument would function in an authentic learning situation which allows for more movement and where the stimuli are not controlled. thus, the combination with other forms of data collection (e.g. video recordings or questionnaires) seems to be crucial and offers the possibility to expand and further refine the use of emotion measurement procedures in educational research. (harley, 2015). 7.1 conclusions despite these limitations, the results of this study indicate that the bm system is capable of validly measuring emotional responses via physiological data in primary school children. it can capture multiple data sets per second without the selfreport bias, allowing rapid and dynamic transitions of emotional responses to be tracked (harley, 2019). the bm system as a physiological measuring device offers a promising method in the context of learning research in capturing emotion responses at a granular level (harley, 2015). in the future, it might be possible to capture physiological responses and their trajectories during teaching and learning processes and thus explore the reciprocal influence of emotional responses and learning processes. sasse et leuchter 46 | flr key points positive and negative emotions influence learning processes; thus, it is crucial to capture the trajectories of emotional responses during learning. children of primary school age still have limited capacity of working memory; thus, it is difficult to capture emotional responses during learning. we stimulated appetitive and aversive emotional responses in children of primary school age with a specifically developed app. we used this app to validate the bm sensor wristband, capturing emotional responses of primary school children. we demonstrated that we were able to capture emotional responses of primary school children with the bm sensor wristband. acknowledgements the authors thank georgios papastefanou for his consultation on this project and philip koch for his assistance in data collection. we would also like to thank all schools, teachers, parents and students who participated in the research. references azevedo, r., & gašević, d. (2019). analyzing multimodal multichannel data about self-regulated learning with advanced learning technologies: issues and challenges. computers in human behavior, 96, 207210. barrett, l. f. (2006). solving the emotion paradox: categorization and the experience of emotion. personality and social psychology review, 10(1), 20–46. https://doi.org/10.1207/s15327957pspr1001_2 bergner, b. s., exner, j. p., memmel, m., raslan, r., talal, m., taha, d., & zeile, p. (2013). human sensory assessment methods in urban planning–a case study in alexandria. in planning times–you better keep planning or you get in deep water, for the cities they are a-changin'. proceedings of 18th international conference on urban planning, regional development and information society (pp. 407– 417). corp–compentence center of urban and regional planning. boucsein, w. (2012). electrodermal activity. springer science & business media. brandstätter, v., schüler, j., puca, r. m. & lozo, l. (2018). motivation und emotion: allgemeine psychologie für bachelor (springer-lehrbuch) (2nd ed.). springer. brünken, r., & seufert, t. (2006). aufmerksamkeit, lernen, lernstrategien. in h. mandl & h. f. friedrich (eds.), handbuch lernstrategien (pp. 27–37). hogrefe. d'mello, s., & graesser, a. (2007). monitoring affective trajectories during complex learning. proceedings of the annual meeting of the cognitive science society 29(29), 203–208. https://escholarship.org/uc/item/6p18v65q???? ekman, p., & cordaro, d. (2011). what is meant by calling emotions basic. emotion review, 3(4), 364–370. ekman, p. (1999) basic emotions. in: t, dalgleish & m, power (eds.) handbook of cognition and emotion (pp.45–60). wiley. sasse et leuchter 47 | flr ekman, p. (1982). methods for measuring facial action. in k. r. scherer & p. ekman (eds.), handbook of methods in nonverbal behavior research (pp. 45–90). cambridge university press. endedijk, m., hoogeboom, m., groenier, m., de laat, s. & van sas, j. (2018). using sensor technology to capture the structure and content of team interactions in medical emergency teams during stressful moments. frontline learning research, 123–147. https://doi.org/10.14786/flr.v6i3.353 engel, s. (2011). children’s need to know: curiosity in schools. harvard educational review, 81(4), 625– 645. https://doi.org/10.17763/haer.81.4.h054131316473115 eren, a. & coskun, h. (2016). students’ level of boredom, boredom coping strategies, epistemic curiosity, and graded performance. the journal of educational research, 109(6), 574–588. https://doi.org/10.1080/00220671.2014.999364 eysenck, m., payne, s. & derakshan, n. (2005). trait anxiety, visuospatial processing, and working memory. cognition & emotion, 19(8), 1214–1228. https://doi.org/10.1080/02699930500260245 falk, b. (1998). effects of thermal stress during rest and exercise in the paediatric population. sports med. 25, 221–240. https://doi.org/10.2165/00007256-199825040-00002 götz, t., frenzel, a. c., & pekrun, r. (2007). emotionen im lernund leistungskontext. katechistische blätter, 132(1), 13–19. götz, t., pekrun, r., zirngibl, a., jullien, s., kleine, m., vom hofe, r. & blum, w. (2004). leistung und emotionales erleben im fach mathematik. zeitschrift für pädagogische psychologie, 18(3/4), 201–212. https://doi.org/10.1024/1010-0652.18.34.201 gravina, r., alinia, p., ghasemzadeh, h., & fortino, g. (2017). multi-sensor fusion in body sensor networks: state-of-the-art and research challenges. information fusion, 35, 68-80. https://doi.org/10.1016/j.inffus.2016.09.005 hamaker, e. l. (2012). why researchers should think “within-person”: a paradigmatic rationale. in m. r. mehl & t. s. conner (eds.), handbook of research methods for studying daily life (pp. 43-61). guilford. harley, j. m., jarrell, a., & lajoie, s. p. (2019). emotion regulation tendencies, achievement emotions, and physiological arousal in a medical diagnostic reasoning simulation. instructional science, 47(2), 151– 180. https://doi.org/10.1007/s11251-018-09480-z harley, j. m. (2015). measuring emotions: a survey of cutting edge methodologies used in computerbased learning environment research. in s. tettegah & m. gartmeier (eds.), emotions, technology, design, and learning (pp. 89–114). academic press, elsevier. https://doi.org/10.1016/b978-0-12801856-9.00005-0 hascher, t., & brandenberger, c.c. (2018). emotionen und lernen im unterricht. in m. huber & s. krause (eds.), bildung und emotionen (pp. 289–310). springer. hayes, d. j., duncan, n. w., xu, j. & northoff, g. (2014). a comparison of neural responses to appetitive and aversive stimuli in humans and other mammals. neuroscience & biobehavioral reviews, 45, 350–368. https://doi.org/10.1016/j.neubiorev.2014.06.018 horwitz, a. v. (2013). anxiety: a short history. jhu press. immordino-yang, m. h., & christodoulou, j. a. (2014). neuroscientific contributions to understanding and measuring emotions in educational contexts. in r. pekrun, & l. linnenbrink-garcia (eds.), educational psychology handbook series. international handbook of emotions in education (pp. 607–624). taylor & francis / routledge. jamet, e., gavota, m. & quaireau, c. (2008). attention guiding in multimedia learning. learning and instruction, 18(2), 135–145. https://doi.org/10.1016/j.learninstruc.2007.01.011 sasse et leuchter 48 | flr jänig, w. (2008). integrative action of the autonomic nervous system: neurobiology of homeostasis. cambridge university press. järvenoja, h., järvelä, s., törmänen, t., näykki, p., malmberg, j., kurki, k., mykkänen, a. & isohätälä, j. (2018). capturing motivation and emotion regulation during a learning process. frontline learning research, 85–104. https://doi.org/10.14786/flr.v6i3.369 kołodziej, m., tarnowski, p., majkowski, a., & rak, r. j. (2019). electrodermal activity measurements for detection of emotional arousal. bulletin of the polish academy of sciences: technical sciences, 813– 826. https://doi.org/10.24425/bpasts.2019.130190 krapp, a. & prenzel, m. (2011). research on interest in science: theories, methods, and findings. international journal of science education, 33(1), 27–50. https://doi.org/10.1080/09500693.2010.518645 kreibig, s. d. (2010). autonomic nervous system activity in emotion: a review. biological psychology, 84(3), 394–421. https://doi.org/10.1016/j.biopsycho.2010.03.010 kreibig, s. d. (2014). autonomic nervous system measurement of emotion in education and achievement settings. in r. pekrun, & l. linnenbrink-garia (eds.), educational psychology handbook series. international handbook of emotions in education (pp. 625–642). taylor & francis / routledge. lee, s. & lang, a. (2009). discrete emotion and motivation: relative activation in the appetitive and aversive motivational systems as a function of anger, sadness, fear, and joy during televised information campaigns. media psychology, 12(2), 148–170. https://doi.org/10.1080/15213260902849927 levenson, r. w. (2014). the autonomic nervous system and emotion. emotion review, 6(2), 100–112. https://doi.org/10.1177/1754073913512003 levenson, r. w. (2011). basic emotion questions. emotion review, 3(4), 379–386. https://doi.org/10.1177/1754073911410743 levenson, r. w. (2006). blood, sweat, and fears. annals of the new york academy of sciences, 1000(1), 348–366. https://doi.org/10.1196/annals.1280.016 levenson, r. w. (2003). autonomic specificity and emotion. in r. j. davidson, k. r. scherer, & h. h. goldsmith (eds.), series in affective science. handbook of affective sciences (p. 212–224). oxford university press. lewalter, d., & schreyer, i. (2000). entwicklung von interessen und abneigungen-zwei seiten einer medaille?: studie zur entwicklung berufsbezogener abneigungen in der erstausbildung. in u. schiefele, & k.p. wild (eds.), interesse und lernmotivation, untersuchungen zu entwicklung, förderung und wirkung (pp. 53–72). waxmann. lewis, m. (2000). the emergence of human emotions. in m. lewis & j. m. haviland‐ jones (eds.), handbook of emotions ( 2nd ed., pp. 304– 319). guilford press. li, x., hijazi, i., koenig, r., lv, z., zhong, c. & schmitt, g. (2016). assessing essential qualities of urban space with emotional and visual data based on gis technique. isprs international journal of geoinformation, 5(11), 218. https://doi.org/10.3390/ijgi5110218 linnenbrink-garcia, l., patall, e. a. & pekrun, r. (2016). adaptive motivation and emotion in education. policy insights from the behavioral and brain sciences, 3(2), 228–236. https://doi.org/10.1177/2372732216644450 lohani, m., payne, b. r., & strayer, d. l. (2019). a review of psychophysiological measures to assess cognitive states in real-world driving. frontiers in human neuroscience, 13, 57. https://doi.org/10.3389/fnhum.2019.00057 lohrmann, k., haag, l., & götz, t. (2011). dösen bis zum pausengong: langeweile im unterricht; ursachen sasse et leuchter 49 | flr und regulationsstrategien von schülerinnen und schülern. schulverwaltung: zeitschrift für schulleitung und schulaufsicht, 34(4), 113–116. http://nbn-resolving.de/urn:nbn:de:bsz:352-183935 lohrmann, k. (2008). langeweile im unterricht. ergänzende darstellung des forschungsstands: zusammenfassung von einzelstudien. waxmann. http://www.waxmann.com/kat/1896.html lombrozo, t. (2016). explanatory preferences shape learning and inference. trends in cognitive sciences, 20(10), 748–759. https://doi.org/10.1016/j.tics.2016.08.001 markey, a., & loewenstein, g. (2014). curiosity. in r. pekrun, & l. linnenbrink-garia (eds.), educational psychology handbook series. international handbook of emotions in education (pp. 228–245). taylor & francis / routledge. martinent, g., campo, m., & ferrand, c. (2012). a descriptive study of emotional process during competition: nature, frequency, direction, duration and co-occurrence of discrete emotions. psychology of sport and exercise, 13(2), 142–151. https://doi.org/10.1016/j.psychsport.2011.10.006 mees, u. (2020). ärger. in m. a. wirtz (ed.), dorsch – lexikon der psychologie. hogrefe. moosbrugger, h. & kelava, a. (2012). testtheorie und fragebogenkonstruktion (springer-lehrbuch) (2nd ed.). springer. mulryan, c. m. (1992). student passivity during cooperative small groups in mathematics. the journal of educational research, 85(5), 261–273. https://doi.org/10.1080/00220671.1992.9941126 nakasone, a., prendinger, h., & ishizuka, m. (2005). emotion recognition from electromyography and skin conductance. proc. of the 5th international workshop on biosignal interpretation, 219–222. nuñez, j., teixeira, i., silva, a., zeile, p., dekoninck, l. & botteldooren, d. (2018). the influence of noise, vibration, cycle paths, and period of day on stress experienced by cyclists. sustainability, 10(7), 2379. https://doi.org/10.3390/su1007237 oxendine, j. b. (1970). emotional arousal and motor performance. quest, 13(1), 23–32. https://doi.org/10.1080/00336297.1970.10519673 panksepp, j. & watt, d. (2011). what is basic about basic emotions? lasting lessons from affective neuroscience. emotion review, 3(4), 387–396. https://doi.org/10.1177/1754073911410741 panksepp, j. (1982). toward a general psychobiological theory of emotions. behavioral and brain sciences, 5(3), 407–422. https://doi.org/10.1017/s0140525x00012759 papastefanou, g. (2013). experimentelle validierung eines sensor-armbandes zur mobilen messung physiologischer stress-reaktionen. gesis-technical reports, 07, 1–14. https://nbnresolving.org/urn:nbn:de:0168-ssoar-339493 paris, s. g., & turner, j. c. (1994). situated motivation. in p. r. pintrich, d. r. brown, & c. e. weinstein (eds.), student motivation, cognition, and learning: essays in honor of wilbert j. mckeachie (pp. 213–237). lawrence erlbaum associates. pekrun, r. (2018). emotion, lernen und leistung. in m. huber, & s. krause (eds.), bildung und emotion (pp. 215–579). springer. pekrun, r. (2017). emotion and achievement during adolescence. child development perspectives, 11(3), 215–221. https://doi.org/10.1111/cdep.12237 pekrun, r., & bühner, m. (2014). self-report measures of academic emotions. in r. pekrun, & l. linnenbrink-garia (eds.), educational psychology handbook series. international handbook of emotions in education (pp. 561–579). taylor & francis. pekrun, r., & linnenbrink-garcia, l. (2012). academic emotions and student engagement. in s. l. christenson, a. l. reschly & c. wylie (eds.), handbook of research on student engagement (pp. 259– 282). springer. https://doi.org/10.1007/978-1-4614-2018-7_12 sasse et leuchter 50 | flr pekrun, r., & stephens, e. j. (2012). academic emotions. in k. r. harris, s. graham, t. urdan, s. graham, j. m. royer, & m. zeidner (eds.), apa handbooks in psychology®. apa educational psychology handbook, vol. 2. individual differences and cultural and contextual factors (pp. 3–31). american psychological association. pekrun, r., goetz, t., frenzel, a. c., barchfeld, p. & perry, r. p. (2011). measuring emotions in students’ learning and performance: the achievement emotions questionnaire (aeq). contemporary educational psychology, 36(1), 36–48. https://doi.org/10.1016/j.cedpsych.2010.10.002 pekrun, r., goetz, t., titz, w. & perry, r. p. (2002). academic emotions in students’ self-regulated learning and achievement: a program of qualitative and quantitative research. educational psychologist, 37(2), 91–105. https://doi.org/10.1207/s15326985ep3702_4 pintrich, p. r. (2003). a motivational science perspective on the role of student motivation in learning and teaching contexts. journal of educational psychology, 95(4), 667–686. https://doi.org/10.1037/0022-0663.95.4.667 piórkowska, m., & wrobel, m. (2017). basic emotions. in v. zeigler-hill & t. k. shackelford (eds.), encyclopedia of personality and individual differences (pp. 1–6). springer. posner, m. i., & rothbart, m. k. (2005). influencing brain networks: implications for education. trends in cognitive sciences, 9(3), 99–103. https://doi.org/10.1016/j.tics.2005.01.007 resch, b., sudmanns, m., sagl, g., summa, a., zeile, p. & exner, j.-p. (2015). crowdsourcing physiological conditions and subjective emotions by coupling technical and human mobile sensors. gi_forum, 1, 514–524. https://doi.org/10.1553/giscience2015s514 riley, j. l., mckevitt, b. c., shriver, m. d. & allen, k. d. (2011). increasing on-task behavior using teacher attention delivered on a fixed-time schedule. journal of behavioral education, 20(3), 149– 162. https://doi.org/10.1007/s10864-011-9132-y rodriguez, p., ortigosa, a., & carro, r. m. (2012). extracting emotions from texts in e-learning environments. in 2012 sixth international conference on complex, intelligent, and software intensive systems (pp. 887–892). rösler, f., heuer, h., hasselhorn, m. & gold, a. (2009). pädagogische psychologie: erfolgreiches lernen und lehren (kohlhammer standards psychologie) (2nd ed.). kohlhammer. roth, g. (2011). bildung braucht persönlichkeit, wie lernen gelingt (3rd ed.). klett-cotta. russel, w. b. (1980). review of the role of colloidal forces in the rheology of suspensions. journal of rheology, 24(3), 287–317. https://doi.org/10.1122/1.549564 scherer, k. r. (1999). appraisal theory. in t. dalgleish & m. j. power (eds.), handbook of cognition and emotion (pp. 637–663). john wiley & sons ltd. shen, l., wang, m., & shen, r. (2009). affective e-learning: using “emotional” data to improve learning in pervasive learning environment. journal of educational technology & society, 12(2), 176–189. sidelinger, r. j. & booth-butterfield, m. (2010). co-constructing student involvement: an examination of teacher confirmation and student-to-student connectedness in the college classroom. communication education, 59(2), 165–184. https://doi.org/10.1080/03634520903390867 singer, j. d., & willett, j. b. (2003). applied longitudinal data analysis: modeling change and event occurrence. oxford university press. https://doi.org/10.1093/acprof:oso/9780195152968.001.0001 siyez, d. ğ. m. (2017). school-related variables in the dimensions of anger in high school students in turkey. international journal of school & educational psychology, 6(2), 112–123. https://doi.org/10.1080/21683603.2017.1302849 sasse et leuchter 51 | flr steinmayr, r., ziegler, m. & träuble, b. (2010). do intelligence and sustained attention interact in predicting academic achievement? learning and individual differences, 20(1), 14–18. https://doi.org/10.1016/j.lindif.2009.10.009 sturm, w. (2005). aufmerksamkeitsstörungen. hogrefe. turner, j., & trucano, m. (2014). measuring situated emotion. in r. pekrun & l. linnenbrink-garia (eds.), educational psychology handbook series. international handbook of emotions in education (pp. 643– 658). taylor & francis / routledge. winne, p. h., & perry, n. e. (2000). measuring self-regulated learning. in m. boekaerts, p.r. pintrich, & m. zeidner (eds.), handbook of self-regulation (pp. 531–566). academic press. zeile, p., resch, b., dörrzapf, l., exner, j. p., sagl, g., summa, a., & sudmanns, m. (2015). urban emotions–tools of integrating people’s perception into urban planning. in real corp 2015. plan together–right now–overall. from vision to reality for vibrant cities and regions. proceedings of 20th international conference on urban planning, regional development and information society (pp. 905–912). corp–competence center of urban and regional planning. zillmann, d. (2008). emotional arousal theory. the international encyclopedia of communication. https://doi.org/10.1002/9781405186407.wbiece022 microsoft word final_proofs.docx frontline learning research vol. 9 no. 4 (2021) 134 issn 2295-3159 corresponding author: folke j. glastra, department of educational studies, leiden university, the netherlands, email address: glastra@fsw.leidenuniv.nl doi: https://doi.org/10.14786/flr.v9i4.831 valences and sense of personal autonomy with regard to professional development in dutch primary teachers: do decision contexts and age make a difference? folke j. glastra1 & cornelis j. de brabander1,2 1 department of educational studies, leiden university, the netherlands 2 query informatisering, voorhout, the netherlands article received 5 april 2021 / article revised 9 august / accepted 9 september / available online 15 october abstract in this study on motivations concerning professional development (pd) we interviewed 95 primary school teachers in the netherlands. we coded these data using the unified model of task-specific motivation (de brabander & martens, 2014) in different decision contexts concerning who decides about teacher participation in pd: school board, teacher teams, or individual teachers. we analysed the valences that teachers associated with pd activities, their experiences of autonomy, and whether and how these variables were affected by decision context and teacher age. results show that decision contexts relate differently to valences and autonomy experiences. positive autonomy and positive valences increased going from school board to team to individual decision contexts. whereas the literature on effective teacher pd stresses the importance of pd design features, our study is the first to empirically demonstrate the crucial influence of decision contexts. among older teachers, teaching experience informed the selection of pd content to transfer to their classrooms. younger teachers tended to first explore whether pd worked in their classrooms before deciding about adoption. direct applicability emerged as a dominant criterion for evaluating pd. decision context and autonomy regarding pd programmes play important roles in ensuring applicability. our research revealed that the dominance of the direct applicability criterion was not motivated by student benefits alone. it was also based on an attitude of efficiency among primary teachers, reflecting growing work pressures and a general prioritisation of classroom teaching above all other tasks, including pd. keywords: task-specific motivation; professional development; primary teachers; decision contexts; teaching experience glastra et de brabander 2 | f l r 1. introduction stimulated by the transformation of advanced nations into knowledge societies that compete with each other on the strength of their knowledge base, and by international rankings of student performance (pisa, timss, pirls), national educational systems around the world have been faced with demands for higher student performance outcomes (borko, 2004). at the same time, the alleged insufficient quality of teachers and teacher education has been highlighted in national debates (ministerie van ocw, 2012; webb et al, 2004). in many places, these demands for higher standards and performance-based approaches of learning have coincided with the devolvement of educational responsibilities from the central government to the level of the school (board) (ball, 2003). this has prompted the strengthening of external assessments by educational inspectorates, as witnessed in the case of the uk by ofsted (hall & noyes, 2009), and the establishment of a system of internal quality controls for accountability purposes within schools (karsten, 1999). these educational reform programmes, together with developments such as the growing digitalisation, ultimately rely for their realisation on teachers and their teaching practices (borko, 2004). teachers are required to develop new skills, knowledge, and attitudes; in short, “new professionalism” (webb et al., 2004; day et al., 2005). however, research has revealed resistance from teachers to the implications of these reform programmes for their professional autonomy (leithwood et al., 2002; locke et al., 2005; moore et al., 2002). to address demands for reform and to overcome resistance, teacher education had to be improved. moreover, the continuing professional development (pd) of teachers was given a pivotal role in raising standards and attaining educational reform goals (hochberg & desimone, 2010). this had to be pd that is effective. in policy discourse, the hallmark of effective pd is teacher quality enhancement leading to gains in student performance (cf. little & bartlett, 2010). however, experts are not in agreement about what exactly constitutes effective teacher pd (cf. wilson & berne, 1999). 1.1. effective teacher professional development: a short review there is extensive literature on effective teacher pd (borko, 2004; clark et al., 2012; darlinghammond & richardson, 2009; desimone, 2009; garet et al., 2001; penuel et al., 2007). desimone (2009) defines effective pd as learning activities resulting in increased teacher knowledge and skills, improvements in practice, and increases in student achievement. it focuses on content knowledge, makes use of active learning arrangements, is coherent with teacher beliefs and knowledge as well as the external reform and policy context, is not a one-off but has a longer duration and is characterised by collective participation that fosters collaboration. penuel et al. (2007) identify similar aspects of effective pd. they differ from desimone, however, in that they find teaching strategies just as valuable in pd as content knowledge. darling-hammond and richardson (2009) state that effective teacher pd should be part of a school reform effort, which requires collective participation and collaboration to be successful. the above conceptions of effective teacher pd stress the role of design features in determining pd effectiveness. van duzor (2010) concentrates on the actual transfer of pd lessons to classroom practice and points out that teachers will be motivated for transfer when pd is coherent with their own teaching philosophies and fits both their own learning needs and those of their students. this should be achieved by actively integrating teachers’ professional expertise into their pd. pd should constitute teacher learning, rather than teacher training from a deficit perspective. schibeci and hickey (2004) take this argument a step further. like van duzor, they suggest that pd will only pay off if it fits teacher needs and classroom contexts. however, they argue that this cannot be achieved in mandatory pd. only when teachers themselves decide on the content and process of pd will there be teacher learning, contextual relevance, and teacher-driven innovation (emo, 2015). this point of view is also manifest in the “teacher communities of practice” literature (vescio et al., 2008). teachers are motivated when they are responsible for the organisation of their own pd (hawley & valli, glastra et de brabander 3 | f l r 1999; little, 2006). kennedy (2016) concludes that teacher autonomy in pd stimulates the implementation of pd lessons, through increased motivation. 1.2. some corollaries for this investigation besides identifying characteristics of effective pd according to the literature, this short overview of effective teacher pd leads to four observations. first, the allocation of responsibility for choice and organisation of teacher pd emerges as an important factor explaining teacher motivations for engaging in and implementing pd. the issue of the power relations that govern pd participation has rarely been investigated empirically. therefore, the impact of who decides about pd the school board, the team, or the individual teacher on motivations, and teacher autonomy vis-à vis pd programmes, will be analysed in this study. second, most authors in our review seem to agree that pd can only be effective if it aligns with the beliefs and learning needs of teachers. from this point of view, it is remarkable that the literature pays little attention to the issue of teacher career and life cycles (day & gu, 2007; fox et al., 2015; huberman, 1989; maskit, 2011; richter et al., 2011). as working conditions change and teaching experience accrues, learning needs may change. therefore, we will investigate the role that teaching experience plays in motivations for pd. third, the literature focuses on what pd could and should be and much less on what it is. little (1993) has pointed out that teacher pd appears as a patchwork of learning activities, predominantly involving teacher training in the form of short-term, packaged programmes, amounting to a fragmented curriculum (cf. ball & cohen, 1999). this image of pd provision seems valid for the dutch situation too, that constitutes the context of our study (seor, 2012; van veen et al., 2010). in this study, we approach this pd patchwork from the perspective of primary teachers. we define teacher pd broadly, as any activity that teachers engage in, individually or collectively, that is directed at or results in the development of their skills, knowledge, attitudes or practices (cf. evans, 2014). such activities may be formal, non-formal or informal; on-the-job or off-the-job; one-off or long-lasting; provided or selforganised. fourth, the literature on effective teacher pd, even where it argues for the necessity of teacher control, fails to acknowledge that such effectiveness is always to some degree contingent on the restrictions imposed by working relationships in primary education, the relative dominance of pd practices (little, 1993), and teachers’ professional beliefs. the insecurity that teachers experience regarding their contribution to the development of their students, as highlighted in the classic study of the schoolteacher by lortie (1975), may serve as an illustration. such insecurity results in a tendency to look for signs of teaching effect in the present and to seek short-term psychic rewards. it may predispose teachers to prefer interaction with their students over pd. taking such pd motivations as the unquestioned basis for improving pd entails the risk of undermining conditions for both deep, innovative, and ground-breaking learning in pd and for school development. therefore, we propose in this investigation to reflect on the motivational data gathered from the perspective of general attitudes to teaching and pd among teachers, as they are impacted by both the constraints and the potential of their schools as work environments. the above debate constitutes the background for our investigation of (the impact of decision context and teaching experience on) motivation for pd among primary teachers. in this paper, our primary goal is to explore motivational aspects of pd to better understand the perspectives of primary teachers regarding pd. in the following section, we explain the motivation model guiding our research. 1.3. valences and autonomy in the unified model of task-specific motivation our research into teachers’ motivation for pd was guided by the unified model of task-specific motivation (umtm; de brabander & martens, 2014). for a full explanation of the umtm, we refer to glastra et de brabander 4 | f l r the theoretical proposal by de brabander and martens (2014) and empirical explorations of the model by de brabander and martens (2018) and de brabander and glastra (2018, 2021). in this section, we will introduce the concepts of valences and autonomy that form the focus of the current study. the umtm (de brabander & martens, 2014) was intended as an integration of the task-specific aspects of different and partly conflicting theories of motivation, such as the self-determination theory (ryan & deci, 2000, 2020), the (situative) expectancy-value theory of achievement motivation (wigfield & eccles, 2000; eccles & wigfield, 2020), the social-cognitive theory (bandura, 1997; schunk & dibenedetto, 2020), the person-object theory of interest (krapp, 2005), and the theory of planned behaviour (ajzen & fishbein, 2008). furthermore, the umtm also incorporates the approachavoidance distinction (elliot, 2006). the umtm aims to cover all the constructs needed to understand the motivation of a person for an activity in a specific situation at a certain point in time. more stable factors relating to the person or the context also influence motivation, but it is assumed that these factors find expression in task-specific constructs. the definitions of the concepts used in this study are taken from de brabander and martens (2014), using adaptations by de brabander and glastra (2018, 2021). in the umtm, motivation is defined as readiness for action. readiness for action is influenced by a valence appraisal, to which both affective and cognitive valences contribute. valence appraisal represents the overall attractiveness of an activity. affective and cognitive valences are produced by separate and relatively independent systems of behaviour regulation. affective valences are defined as feelings that originate from undertaking an activity. affective valences represent the levels of positive and negative feelings experienced when doing an activity. any action that comes to mind as an action option immediately and unavoidably evokes the corresponding feelings. the reason for or function of such feelings is not necessarily known and is often unknown. cognitive valences are defined as the articulation and valuation of the consequences of performing an activity. cognitive valences are brought about by active reflection of the actor. an activity has multiple consequences, intended and unintended, each of which will have a separate cognitive valence. the quality of cognitive valences obviously depends on the actor’s competence to articulate and value the consequences of an activity. cognitive valences can be broken down into different parts depending on who receives the outcome value. evidently, the first to profit from an activity is almost always the actor themselves. in the context of teacher pd, this would be the teacher, while students and the school are natural candidates for sources of non-personal cognitive valences. affective valences have no non-personal component. only the feelings the person experiences about an activity are relevant. affective valences and cognitive valences are defined as two theoretically independent categories of motives. what is pleasurable to do is generally also recognised as profitable and vice versa, but discrepancies between the two are also possible (“i know it is good for me, but i don’t like it”). affective and cognitive valences are combined on a common scale (valence appraisal) through interaction between different systems of behaviour regulation. in this respect, the umtm adopts the position of the person-object theory of interest (krapp, 2005) and differs from both self-determination theory and expectancy-value theories. in self-determination theory (ryan & deci, 2020) intrinsic motivation (affective component) is opposed to extrinsic motivation (cognitive component), as extrinsic motivation diminishes intrinsic motivation. expectancy-value theories, on the other hand, define affective and cognitive aspects as facets of the same value construct, summing to a total value. the distinction between affective and cognitive valences is combined with the distinction between approach and avoidance motivation (elliot, 2006). this distinction is related to the possibility of valences being positive or negative. positive valences call for an action that promises to realise them. negative valences prompt refraining from an action or initialising a counter-action in order to prevent their realisation. approach and avoidance motivation are also seen as being regulated by different systems (carver, 2006). combining the distinction between positive and negative valences with the glastra et de brabander 5 | f l r distinction between affective and cognitive valences gives rise to four relatively independent motivational components. together, these combine to form a resultant value, which is called valence appraisal, which in turn determines readiness for action. valences are influenced in the first place by their task-specific precursors. the umtm encompasses the concepts of autonomy, feasibility, relatedness, and subjective norm. in this study, we will make use of the concept of autonomy. autonomy refers to the origin of the action: this may be the self, or an internal or external source that is experienced as foreign. in umtm, a higher level of taskspecific autonomy can lead to increased task-specific motivation, because it can affect the level of affective and cognitive valences. the model makes a distinction between individuals and their contexts, which leads to a partition of some components in the model, among them autonomy. the personal facet of autonomy is labelled “sense of personal autonomy”, and is defined as the extent to which a person experiences themselves as the originating force that drives the performance of an activity. the distinction here concerns whether the self feels that it is driving or being driven. the contextual part of autonomy is labelled “perceived freedom of action”, and entails the liberty a person perceives in the action context to decide independently about the choice and execution of action alternatives (cf. reeve et al., 2003). of course, the two aspects are related; their fundamental distinctiveness, however, is revealed by the possibility that one can still experience the self as the driving force in a situation where one has no freedom of action. while autonomy is identified as a general human need in motivation theory (ryan & deci, 2000), it has specific significance in the work of teachers. this relates to the status of teachers as professional employees, who exercise a high level of discretion in the relative isolation of their classrooms. research shows an association between teacher autonomy and engagement, job satisfaction, and emotional exhaustion (pearson & moomaw, 2006; skaalvik & skaalvik, 2014). however, in times of standardisation of national curricula, standardised testing, and doubts about the competence of primary teachers, teacher autonomy as a strong motivational component within the teaching profession is called into question (apple & jungck, 1990; ball, 2003; ballet & kelchtermans, 2009; clark, et al., 2012; gleeson & husbands, 2003; hardy & lingard, 2008; milner, 2013; stevenson et al., 2007). therefore, in our investigation, we pay particular attention to the role of autonomy in pd motivation. 1.4. research concerning decision contexts and participation in pd the debate concerning teacher autonomy manifests itself in pd issues too: should participation be determined by teachers themselves, or by governmental educational reform agendas (borko et al., 2002; starkey et al., 2009). in 2012, the dutch ministry of education (2012, pp. 18–20) saw its educational reform agenda increasingly reflected in school board and teachers’ pd choices. a survey among dutch primary and secondary teachers (onderwijscoöperatie, 2016) reported that decisions about teacher participation in pd were made by school boards in 48% of cases, by teachers in 1% of cases, and by both parties jointly in 36% of cases. according to clark et al. (2012), teachers in canada experienced increasing constraints in determining pd participation. policy-driven, compulsory, and often one-size-fits-all pd programmes met with fierce criticism from these teachers (locke et al., 2005). the current study examines three decision contexts: decisions concerning participation in pd made by the school board, by the team, or by individual teachers. in most cases, school board pd decisions are made for groups or whole teams of teachers, and they are binding. they may be the outcome of local school board decisions, but often they result from government educational policies or inspectorate reports (seor, 2012). team decisions are based on collectively defined learning priorities, often concerning issues that are overarching for (parts) of the school. individual decision contexts may result from teachers’ perceptions of knowledge and skills gaps concerning emerging classroom needs or externally developed knowledge (hashweh, 2003). glastra et de brabander 6 | f l r previous studies have considered teacher motivations for pd (e.g., jansen in de wal et al., 2014; kwakman, 2003; thoonen et al., 2011; van eekelen et al., 2006; vermunt & endedijk, 2011), as well as teacher involvement in school decision-making (scribner et al., 2007; smylie, 1992). however, the impact of different decision contexts on pd motivations has not been studied empirically. schibeci and hickey (2004) and clark et al. (2012) have pointed out the importance of teacher autonomy in pd decisions for teacher motivations. jansen in de wal et al. (2014) concluded that freedom of pd choice for teachers enhances their feelings of autonomy, thus enhancing their motivation for pd; however, decision context was not a condition in their study. research concerning reform-driven pd shows that unless it leaves room for teacher input, it risks being rejected (avidov-ungar, 2016; wallace & priestly, 2011). hargreaves (2003) stresses the detrimental motivational consequences of reform-driven pd that involves “a strong insistence on performance standards and prescribed classroom techniques” (p. 176). nir and bogler (2008) note that local educational authorities are seen by teachers as bodies too remote from day-to-day educational practice to make sensible pd decisions. kennedy (2016) states that mandatory participation in teacher pd makes for much lower learning outcomes in teacher pd. 1.5. research concerning teaching experience and (motivation for) participation in pd the relationship between teaching experience or teacher age and pd has been the subject of a body of research. many of these studies (e.g., philipp & kunter, 2013; richter et al., 2011), make use of the career stages model of huberman et al. (1993). five ideal-typical stages that describe the development of teachers from novice to end-of-career (survival and discovery; stabilisation; experimentation/activism and stocktaking; serenity and conservatism; disengagement) are used in pd research to explain motivation, actual participation, or preference for various types of pd content. maskit (2011) proposes a curvilinear relationship between career phase and attitudes towards external educational reform programmes: in the induction phase and the career-frustration and winddown phases, positive attitudes are lower than in the middle phases of competency building, enthusiasm and growth, and stability. richter et al. (2011) come to similar conclusions. their data show a curvilinear relationship between age and participation in in-service training courses. however, they discovered that older teachers invest less in collaboration and more in reading. the assumption that teachers’ pd participation decreases in the last third of their career (cf. day & gu, 2007) may obscure this move from formal and collaborative to informal and individual pd activities. in addition to examining participation patterns in pd, research has also focused on learning needs and preferred pd subjects in teacher careers. cochran-smith and lytle (1999) conclude that in pd, novice teachers seek new knowledge for practice, in order to become more effective teachers, while experienced teachers value knowledge of practice. kyndt et al. (2016) found that novice teachers (1–3 years) have more motivation for pd, and are more interested in knowledge and skills with direct applicability to their own classroom. more experienced teachers showed lower levels of motivation for pd, and concentrated more on learning new teaching methods and changing teaching and learning beliefs. in summary, some career models concerning teacher pd participation distinguish between novice and more experienced teachers, while others use a three-fold distinction and predict a curvilinear model. these models may reflect participation in formal and collective forms of pd, but risk obscuring participation in other types of pd. these studies are not highly informative when it comes to careerspecific learning needs. we have not found research regarding career stage and pd autonomy, with the exception of clark et al. (2012, p. 138) who note that significantly more senior than junior teachers perceived degradation of their classroom autonomy. the research reviewed does not permit us to specify expectations regarding teaching experience and pd motivations. glastra et de brabander 7 | f l r 1.6. research questions this research is guided by the following research questions: 1. with regard to pd activities among primary teachers, what are the differences and similarities between board/principal, team and individual decision contexts in terms of (positive and negative) affective and cognitive valences (for the school, the student, and the teacher personally) and autonomy (sense of personal autonomy and perceived freedom of action)? 2. what are the differences and similarities in these variables between less and more experienced teachers? 3. how are pd experiences of primary teachers associated with positive and/or negative valences regarding pd? 4. how are pd experiences of primary teachers associated with (the role of) autonomy with respect to pd? 2. method 2.1 design and sample this study was based on a mixed method design. data collection began in 2014. teachers in the sample had a minimum of one year’s teaching experience. a convenience sample of 441 teachers from 54 primary schools in the netherlands completed a questionnaire about motivational aspects of their pd participation in different decision contexts. results of this component of the study have been published elsewhere (de brabander & glastra, 2018). within each participating school, two teachers were approached for interviews, resulting in a sample of 96 teachers. interviewees were recruited on a voluntary basis from among each school’s questionnaire respondents. one teacher with predominantly remedial teaching tasks was removed from the sample. of this sample of 95 interviewees, 81.6 % were female and 19.4% were male; m (sd) age = 41.01 (11.01), with a range of 22 to 63 years; m (sd) teaching experience = 14.09 (9.03) with a range of 2 to 41 years. both age and teaching experience had a bimodal distribution, with modes at 31 and 50 years (age) and 11 and 34 years (teaching experience). gender and age data show a reasonable match with official dutch primary teacher population statistics (m population age = 43.04 years, 83.7% female teachers; stamos, 2013). although sample characteristics showed a reasonable fit with population characteristics, drawing interviewees on a voluntary basis from survey respondents per school may have had the effect of selecting teachers who were particularly motivated to share their pd experience. 2.2. procedure teachers were interviewed at their schools by pre-master students of education and child studies at leiden university. interviewers received training. teachers were told that the interviews would centre on their pd experiences. interviews lasted 30–45 minutes and were recorded. given that interviewing respondents about all three decision contexts would take too much time, each interviewee was pre-assigned two decision contexts, in such a way that the overall frequency of each decision context being probed would be equal. however, some of the respondents turned out not to have relevant experience in one of the decision contexts assigned to them, which necessitated some last-minute adaptations of decision contexts; nevertheless, the three decision contexts were investigated in almost equal measure. the interviews were transcribed verbatim and imported into atlas.ti for further analysis. glastra et de brabander 8 | f l r in some cases, it proved difficult to distinguish between decision contexts; for instance, when board decisions were perceived by individual teachers or teams as their own decisions, or when memories regarding the locus of decision-making were not clear. in each case, we tried to reconstruct which decision context was concerned, but some errors in the categorisation of cases may have arisen. 2.3. instruments in the semi-structured interview, respondents were first asked to describe their pd activities over the last two years and to indicate whether their participation was decided by themselves, their team, or the school board. respondents were then interviewed about a pd activity of their own choice for each of the two decision contexts assigned to them. the interviewer was instructed to ask for pd experiences. this question was meant to yield insights into both cognitive and affective valences. the question “what consequences has the pd activity had?” was specifically meant to trace cognitive valences. interviewers were instructed to probe for both positive and negative consequences and for consequences for teachers personally, for students, and for the school. sense of autonomy was targeted in the question “did you feel that you were directing your own learning in the pd activity?”, and interviewers were instructed to probe for how much room there had been for teachers’ initiatives regarding these pd activities. the decision was taken to use age as a proxy for teaching experience, and to define two age/experience groups using the median as the cut-off point; this decision was prompted by the fact that (a) a division into three teaching experience groups would yield a very uneven distribution (induction n = 5, mid-career n = 75, end-of-career n = 9, unknown n = 6); (b) both age and teaching experience had bimodal distributions; (c) the relation between age and length of teaching experience was quite strong (r = .86); and (d) teaching experience had far more missing values. after removing four interviewees for whom the exact age was missing, the younger group (< 41 years) comprised 45, and the older group (≥ 41 years) comprised 47 teachers. 2.4. analysis interviews were coded using atlas.ti. text fragments were coded for (negative and positive) cognitive valence (for the teacher, the student, and the school), (negative and positive) affective valence, and autonomy. text fragments could be assigned all of the above codes simultaneously. after being trained by the authors, two research assistants coded the interviews. inter-coder discrepancies were discussed with the authors until consensus was reached. these discussions were guided by definitions of umtm concepts and occurred mainly in the initial coding stages. as it proved to be difficult to distinguish between a sense of personal autonomy and perceived freedom of action, given that the former often remained latent or implied in the latter, it was decided to code for both under the code “autonomy”. once all interviews were coded, the coded fragments were analysed by the authors in three steps. the first step in the analysis focused on identifying the presence of codes and combinations of codes across the interviews for each age group and each decision context. each interviewee could add a score of 1 to the number of code occurrences across interviews, irrespective of their frequency within an individual interview. this number, that is, the number of teachers with a specific code, was then transformed to give the proportion of the total number of teachers per age group who were interviewed about a specific decision context (see table 1). in this way, the uneven distribution of age groups and decision contexts was corrected for. co-occurrences indicate semantic associations in statements relating to different codes. co-occurrences of codes, that is, combinations of codes, were treated likewise. for each combination of codes, we counted the number of teachers to whom that combination of codes was assigned, irrespective of the number of text fragments that were actually tagged with that code combination. this number was then transformed to give the proportion of the total number of teachers per age group who were interviewed about a specific decision context. the proportions of occurrences and co-occurrences are presented in tables 1 to 6 in the appendix. with the aid of qgraph glastra et de brabander 9 | f l r (epskamp et al., 2012) these tables were then transformed into code network diagrams (figure 1) to summarise differences and similarities across decision contexts and age groups. in these diagrams, cooccurrence proportions are indicated by lines between codes that differ both in thickness and density. a proportion of 0 is represented by the absence of a line, a proportion of 1 by a thick, black line. proportions between these extremes are represented by lines with different thicknesses and different densities (see also appendix 1, tables 1 to 6). differences in code occurrence proportions are indicated by different circle sizes. the exact formula for rendering the circle size was: vsize = 3 + 10 * proportion of code occurrence. this analysis constitutes an initial, mainly quantitative approach to the interviews. it was also intended to generate research ideas for a further analysis of pd experiences found in the coded text fragments. table 1 distribution of decision contexts across age groups board context team context individual context younger teachers 34 25 28 older teachers 26 34 33 in the second step of the analysis, all interview fragments with the same code were reviewed for differences and similarities in experiential content. the question leading this second analytic step was: what were the experiences associated with positive or negative valences, with affective valences, and with feelings of (lack of) autonomy attached to a pd activity? here, the analysis was not guided by an explicit theory; categorisation followed an inductive, grounded theory approach (strauss & corbin, 1998). categories were formed based on similarities and differences in experiential content guiding associations with pd valences. as a result of this constant comparison of new experiences with those already attributed to categories (glaser & strauss, 1967, p. 102), the coding process necessitated the adaptation of categories, and categories were thus reformulated. in this way, several categories were further developed (including their definitions), and experiences were incorporated into these categories until all experiences in the interviews pertaining to codes had been analysed (see categories and text fragment examples in the appendix). to enhance trustworthiness, the fit of all coded text fragments with category definitions was checked. additionally, the associations of the experiences in each category with decision contexts and age groups were examined to discover specific patterns. finally, an overview of core categories, definitions, and text examples associated with valences and autonomy was produced (see appendix). the third step concerned a reflection on the data thus analysed, from the perspective of teachers’ broadly shared attitudes to pd in the context of teaching as a profession and the working conditions that govern primary education. 3. results in this initial, quantitative part of the analysis, we present codes and code co-occurrences in code networks organised by age group and decision context. in the second part of this section, we will focus on the qualitative analysis of pd experiences in coded text fragments. glastra et de brabander 10 | f l r 3.1. quantitative patterns of codes and co-occurrences table 2 presents an overview of the codes used in the analysis and their meanings. figure 1 depicts the presence of codes and co-occurrences for each age group and decision context. all these patterns constitute interaction effects between age groups and decision contexts. the main differences and similarities between these patterns may be summarised as follows: 1. in the board decision context, there is a difference in the more negative cognitive and affective valences (ncvp, nav) towards pd and a stronger connection between the two for the older teachers compared to the younger teachers (figure 1: on the right-hand side of the board code networks). 2. there are greater positive valences (pcvp, pcvs, pcvl, pav) of team decisions for the younger teachers compared to the older teachers (figure 1: team code networks). 3. the age groups are similar in the individual decision context, except for greater importance of negative personal cognitive valences (ncvp) among the younger age group. together, these last two patterns could point towards a more insecure position of younger teachers, who faced negative consequences of their own pd choices and found protection from negative consequences under team conditions (figure 1: individual code networks). 4. in all decision contexts, there were no big differences between age groups in the importance of autonomy, nor in its associations (figure 1: autonomy in all code networks). 5. for the younger teachers, the individual and team contexts showed similar patterns of strong positive cognitive and affective valences (pcvp, pcvl, pcvs, pav) and associations between these codes, while the board condition showed less strong positive valences and associations (figure 1: left-hand side of code networks in the board context compared to team and individual contexts, for younger teachers). table 2 overview of codes and their meanings aut autonomy pcvp positive cognitive valence personal ncvp negative cognitive valence personal pav positive affective valence pcvs positive cognitive valence for the school ncvs negative cognitive valence for the school nav negative affective valence pcvl positive cognitive valence for the student (learner) ncvl negative cognitive valence for the student (learner) 6. for the older teachers, the individual decision context had the most positive and most strongly associated valences, while the board decision context had the most negative personal cognitive and negative affective valences (figure 1: rightand left-hand side of board code networks compared to individual code networks, for the older age group). 7. for both age groups, the team decision context was the one associated with the least negative consequences and affects (ncvp, nav) (figure 1: right-hand side of team code networks compared to board and individual code networks, for both age groups). 8. for both age groups, positive affective valences (pav) were more strongly connected to personal cognitive valences (pcvp) in the individual decision context compared to board or team decision contexts (figure 1: pav for both age groups in individual compared to team and board contexts). glastra et de brabander 11 | f l r pd). further analysis (figure 2) showed that positive autonomy was far stronger for younger teachers < 41 years ³ 41 years board team individual figure 1. code occurrences and co-occurrences for different combinations of age and decision context (for legend, see table 2) glastra et de brabander 12 | f l r analysing text fragments labelled with the autonomy code, we discovered that these fragments could be subdivided into three different sub-codes: positive autonomy, negative autonomy (the negative valuation of sense of personal autonomy or perceived freedom of action regarding pd), and absence of autonomy (the observation that freedom of action in, or a sense of personal autonomy regarding pd, was absent). further analysis showed that positive autonomy (figure 2) was far stronger for younger teachers from the board (20.6%) through team (32%) to individual (60.7%) decision contexts, compared to older teachers (26.9%; 23.5%; 30.3%). negative autonomy (figure 3) was, at far lower levels, slightly more present among younger teachers, especially in the team (8%) and the individual (7.1%) decision contexts, while older teachers experienced negative autonomy in the board (3.8%) and the individual (3%) decision contexts. this may mean that pd that grants greater autonomy causes more discomfort for younger teachers than older teachers. older teachers perceived the absence of autonomy (figure 4) far more in the team (23.5%) and individual (18.2%) decision contexts than younger teachers (8%, 3.6%). this may mean that younger teachers identify more strongly with team decisions and have a stronger tendency to transpose the feelings of autonomy associated with individual decisions about pd onto the pd activity itself. figure 2. distribution of positive autonomy codes in age groups and decision contexts glastra et de brabander 13 | f l r figure 3. distribution of negative autonomy codes in age groups and decision contexts figure 4. distribution of absence of autonomy codes in age groups and decision contexts 3.2. qualitative patterns in coded text fragments in the following qualitative analysis, we address three questions: (a) what is the nature of primary teachers’ experiences concerning (their participation in) pd that are coded as having positive and negative valences; (b) what is the nature of primary teachers’ statements concerning (the role of) glastra et de brabander 14 | f l r autonomy in pd; and (c) as a reflection on the answers to these questions, which elements of a broadly shared attitude impacting teachers’ valuations of pd can be found in the coded data? to indicate how our findings relate to the interview texts, each finding is labelled with the respondent’s database id number followed by the number of the text fragment. we have also indicated whether the respondent belongs to the older or the younger teacher group by prefixing this reference with “o” or “y” respectively. furthermore, the appendix provides definitions and textual illustrations of the categories presented in the following sections, including citations referred to by respondent/fragment number alone in the following sections. 3.2.1. valences further analysis of coded text fragments yielded several broad categories of experiences associated with positive and negative pd valences. first, the school as a decision and work context, with its workloads, routines, colleagues, and leadership, impacts the ways in which pd is experienced by teachers. second, teachers’ valuation of pd is influenced by the particular consequences for stakeholders, such as the students, the school, and the teachers personally. third, qualities of pd content and design in relation to teachers and their classrooms also informed teachers’ pd valuations. 3.2.1.1. context conditions valuations of particular pd decision contexts were transferred onto valuations of pd itself. “we have had plenty of courses, where you wondered: ‘what am i doing here?’ courses that you had to sit out. okay, the school board decided we have to do this course, we’ll just sit it out.” (y32:16). for many respondents, mandatory participation meant that its contents did not fit with their learning needs or those of their school (o35:19). in the individual context, positive transfer of decision context to pd valences abounded. “the nice thing about it is that you have chosen it yourself, that you’re very interested and therefore you’re very motivated.” (y9:1). particularly in the individual decision context, teachers chose pd programmes focusing on specific problems faced by children in their classrooms, thus addressing their own perceived professional deficits as well as the learning or behavioural problems of these children (o58:3). while individual decisions resulted mostly in positive and board decisions in negative valuations, pd valuations in the team decision context depended more strongly on the qualities of pd itself and the consequences for stakeholders. as an example of the latter, following the pd programme as a team was valued positively by many teachers for reasons of team building, collective engagement, efficient distribution of tasks, and alignment of the course with team learning needs (y66:8; y82:10). fit with workload and work schedule was associated with positive pd valuations. many teachers disliked pd programmes because they were time-consuming or scheduled at particularly busy times. behind these negative valuations were arguments in which teaching time was set off against and prioritised over time spent on pd. “a whole day, for me a waste of time, (…) but for the students as well, because it could have been one day more for them to learn something.” (y36:16) here, teachers clearly demonstrated efficiency reasoning, resulting from the experience of time pressures and teaching as a professional priority. “my appreciation of professional development activities is variable. and that depends on whether it concerns an activity that you need at that moment or an activity that you were not specifically looking for at that moment. and that moment may have to do with the fact that you are in the middle of a school report period or preparing for a christmas musical. it depends greatly on all such things.” (o18:18). the (lack of) support from the school principal for pd participation was also mentioned frequently. “i have a principal who stimulates it very much, who likes it when people can function in .. yes, as if on fire, who acknowledges that then the school benefits in a special way.” (o33:6) 3.2.1.2. stakeholder consequences experiences regarding the consequences of pd for the students, the school, and for teachers themselves impacted pd valuations. here, greater importance was assigned to students and teachers, often in combination, than to schools. glastra et de brabander 15 | f l r students pd was valued positively when it was in the students’ best interest according to the teachers. teachers were satisfied with pd where they felt better equipped to help their students, in the sense of stimulating learning motivations, improving test scores, or diagnosing students’ problems. pd was perceived to better help students when it fit students’ learning needs and attainment levels more precisely or promised to raise student motivations (o92:11) and improve social relations in the classroom. negative valuations of pd were expressed where teachers felt that it lacked the above qualities or harmed students or classroom relations, or where teachers thought that teaching their students was more useful than participating in pd. importantly, teachers referred more frequently to “their” students and their “own” classrooms than to the school as an organisation or to the school’s students in general. the fact that only a small proportion of teachers mentioned negative cognitive valences for students is related to the teachers’ perception that students would often be largely unaware of pd effects in the classroom (y48:7), to the fact that pd was frequently not implemented in the classroom at all, or to the fact that teachers actively filtered negative consequences to protect their students. in the quantitative results, the co-occurrences between positive cognitive valences for teachers and for students is evident. in many statements, teachers presented themselves as the ones passing positive pd learning and inspiration on to the students. “the better the grasp you have of the subject matter, the more freedom you have in dealing with it. therefore, you can offer much more to the children. and i am convinced, in fact, that you cannot teach children just because you know things, but because you are inspired, you really love it yourself.” (o33:3) if pd is to be positively valued, it should have perceptible effects for the teacher and the students. “when you really see the effect—so you’re doing it and you see, you’re pleasantly surprised (…) that you see that it does something for the child and for yourself as well when you see real progress (…). a visible effect, something of a wow factor, what happens now is really cool.” (y9:16). teachers professional growth, addressing knowledge and skills gaps, and becoming more capable of helping students and improving classrooms are experiences that lead to positive valuations of pd. frequently, witnessing pd effects in children appears to be crucial in generating such experiences. in addition, many teachers referred to personal growth and learning as pleasant and rewarding in itself. “i have also looked at what i like. because i like to improve my skills, to challenge myself. and i thought, a bit of reading, spelling, dyslexia: there’s my challenge.” (y3:1). acknowledgement of expertise through pd in school and the perception of pd as a career boost also led to positive pd valuations. a teacher who took an autism course experienced positive consequences. “in fact, i was initially given the cold shoulder in this regard by the school’s in-house autism specialist, but the school principal stood up for me. that felt good.” (o24:19) lack of acknowledgement or the impossibility of putting pd based expertise to use within the school was also often mentioned. here, the school context, often the school principal, plays an important role. school teachers saw pd content as relevant if it addressed issues that they and the principal or the board had defined as urgent. pd that enabled teachers to solve such issues was valued positively. the sense of urgency among teachers was especially great following a negative assessment by the inspectorate of the school’s performance on standardised tests. here, teachers generally indicated that pd participation was necessary and often resulted in removing the threat of school closure. pd effectiveness in these cases was often seen as self-evident, and readiness to participate in pd was without question (y60:1). these cases were an exception to the rule of teachers experiencing a lack of fit of (usually mandatory) pd within board decision contexts. glastra et de brabander 16 | f l r 3.2.1.3. pd qualities immediate applicability of pd most of the teachers stressed the desirability of pd to be immediately applicable in their classroom. “well, the timing was simply very convenient, and you could do something with it straight away, without having to sort out a lot of things.” (o18:2) pd applicability means coherence with prevalent classroom and school situations, as perceived by teachers. a younger teacher stated: “this course referred directly to the group. each teacher could think of their own group: do i have children in my group that could be involved in this? so yes, that’s very practical.” (y65:3) generally, pd lessons that were not put into practice led to negative valuations. “what i disliked was that you had to be patient, that it just took so long before it was implemented. well, of course, you can be deeply serious yourself, but when your colleague isn’t…” (y79:15) time pressures and school routines were said to limit the applicability of pd (y9:12). additionally, some teachers expressed an aversion to “theory”, which served as a generic term for anything not practical in nature (o70:11). in pd, theory should always be closely tied to practice, according to the respondents. many of the pd activities originating in government policies implied (perceived) mandatory application. here, valences depended on judgments about the time investment involved, and on perceived benefits for the students. in general, applicability was attributed to pd in individual and (to a lesser degree) team decision contexts more than in the board contexts. overall, school boards and the ministry were thought to be too far removed from daily school practice to be able to put pd to appropriate use (y23:12). novelty of pd content the novelty of the pd content as perceived by teachers was mostly associated with positive valuations. “sometimes a course disappoints me because there is little new information and because i’ve been in teaching a long time, of course, and i have taken lots of courses.” (o44:1) however, positive valuations were dependent on perceived practicability. when pd served to implement new government policies, teachers often saw it as interfering with tried and tested school routines. “the first time was positive, but then you get courses on ‘action-directed teaching’ and ‘result-oriented teaching’, and then you have to change group plans and you are asked to write down more, and most of all to evaluate; you become increasingly negative, you think, we finally have things in good order, and then we have to start changing everything all over again.” (o61:4). particularly older teachers reported experiencing a policy merry-go-round (o29:6) and often felt that pd programmes did not acknowledge their professional experience. in these cases, (alleged) novelty was valued negatively. pd learning arrangements teachers tended to associate their valuations with certain aspects of pd learning arrangements. for many teachers, pd should not be in a lecture format (o11:4), should facilitate the expression of personal experiences and learning needs (y79:6), and should fit with preferred ways of learning, such as working towards clear learning goals specified from the outset (o18:11) or working in small groups and engaging in direct interaction with peers. “and i always find it valuable when other people from other schools participate, they have totally different input. they do things differently and that’s very valuable.” (o44:1). learning by doing, allowing room for application and experiment (y17:4), and receiving performance feedback were mentioned especially by younger teachers, who were more inclined to adopt innovations and explore whether they worked for students. “i would have appreciated a follow-up session to discuss the problems that you often encounter. not only before you actually work with the method.” (y52:2) in addition, teachers valued the enthusiasm of those delivering pd (y64:2). in the case of board decisions—and in individual and team decisions too, albeit to a lesser degree—prepackaged, standardised pd formats increased impressions of lack of fit of pd with teacher learning needs (o35:16; o43:10). 3.2.1.4. valences and age groups the quantitative analysis revealed some marked differences between younger and older teachers in terms of positive and negative valences in different decision contexts. younger teachers seemed more glastra et de brabander 17 | f l r inclined to accept pd applications at face value, and to discover the consequences of those applications in the classroom through trial and error (y9:16), rather than deciding beforehand about the usefulness of pd. “i take that lesson on board, i will then try it out with the children, looking at whether it helps more if i do it like this. well, when you see that it works, you keep using it.” (y81:7) younger teachers were also more inclined to perceive changes in standardised test scores following student-performancetargeted pd as an indication of their own teaching competence (y60:7). in addition, they made use of pd to critically reflect on their own practice and competencies (y65:4), while most older teachers started from a place of confidence in their teaching skills and classroom experience, and used pd to supplement their knowledge and skills. 3.2.2. autonomy in the interviews, the autonomy code was applied where teachers referred to themselves as steering important aspects of pd or making use of perceived freedom of action related to pd. the coding of sense of personal autonomy proved rather difficult, as it often functioned as a latent theme or was less easily expressed than perceived freedom of action. nonetheless, the analysis elucidated key features in relation to pd that tend to induce feelings of autonomy or perceptions of freedom among teachers. when teachers presented themselves as the central steering force in pd activities they had personally organised such activities (y52:5), instructed pd developers, or composed or selected relevant pd elements, or set personal pd goals beforehand. furthermore, they had decided afterwards which pd knowledge and skills they would and would not bring to their schools, and they actively adapted relevant knowledge and skills to their own teaching practices and school conditions (y48:4). “i think they should show more trust in your professionalism. (…) you will be able to get similar results [to those aimed at by the pd programme] with children if you have an appropriate approach as a teacher. and i do not think that people sitting at their desks know how to do that.” (y7:9) regarding perceived freedom of action, teachers often mentioned pd that offered a choice of content and activities, or that left room for or even demanded their active input. these options and demands ranged from low-level possibilities (mostly) available during the pd, including the chance to ask questions or to influence the pace (o50:5) or time investment involved, to those of a broader, structural nature, such as being able to make a personally or team-appropriate selection of programme elements. further analysis of interview fragments with the autonomy code revealed that when teachers perceived themselves as steering their own learning or when they experienced freedom of action within pd, this was almost always valued as a positive aspect of pd. individual choice of pd was often associated with the perception of positive autonomy in pd. “it was something i had to do for myself, and the school had not really instructed me to do it. so yes, (…) it was in fact ‘my thing’.” (o30:10) pd in board decision contexts could still lead to positive valuations when teachers perceived elements of autonomy in the curriculum or in the way it engaged them in learning (y17:11-21). negative autonomy concerned experiences where teachers felt that the pd gave too much room for personal choice or initiative, thereby jeopardising course coherence. “the course could have been much better. it was rather vague. it was the kind of course where you were very much left to find your own way.” (y76:12) “there was plenty of opportunity for contributions, but in the end, they should add up to something, and that was a bit difficult. it seems easier to me when it is clear that this is it, this is what it should be. there was too much room for participant contributions, and those contributions should fit together.” (y72:14) these statements seem to express a need for provider-centred pd that does not offer too much opportunity for participant contribution. the absence of autonomy was associated with teacher statements where autonomy was generally not linked to either positive or negative valuations. often, this concerned mandatory courses where teachers were obliged to make use of the course content, which was delivered in a fixed format, with no room for self-regulation on the part of the teacher. glastra et de brabander 18 | f l r 3.2.3. broadly shared attitudes a reflection on the data analysis thus far also yielded elements of broadly shared attitudes among the respondents towards pd in the context of teaching as a profession and the working conditions that govern primary education. in many cases when teachers evaluated their pd experiences in negative ways, they made use of compensatory mechanisms, the most frequent of which were, firstly, their generally positive attitude towards learning and professional as well as personal development, and secondly, their professional capacity and willingness to learn even under less than favourable conditions (o18:8). another broadly shared attitude was the view that teaching is an activity that should involve students as directly as possible. in the eyes of many teachers, pd participation often appeared to distract from this professional priority (y36:16). the individual decision context and a degree of autonomy regarding pd proved helpful in increasing the professional priority accorded to pd (see 3.2.2.), but preferably pd should be efficient in delivery and result in practicable knowledge and skills (see 3.2.1.3.). this attitude is reinforced by time pressures experienced by most teachers, who frequently mentioned time-consuming administrative tasks (y9:15). in short, teachers showed a willingness to engage in development, but under the strict condition that pd be efficient and directly applicable. 4. discussion in this investigation, we established code patterns pertaining to affective and cognitive valences and autonomy among primary teachers concerning pd in board, team, and individual decision contexts, across different teacher age groups. we also examined the nature of primary teachers’ experiences in valence and autonomy statements concerning pd. 4.1. the role of decision contexts in pd valuations our research indicates that decision contexts matter: the individual decision context gives rise to the most positive valences for teachers, and the board context is associated with the most negative valences, while the team context appears to shield most against negative valences. the situational relevance and personal fit of pd in terms of learning needs is particularly important for teachers. school boards and the government are perceived as too distant from classrooms and schools to make relevant and appropriate pd decisions in most cases, which is consistent with earlier research findings (clark et al., 2012; locke et al., 2005). teachers see themselves as best positioned to determine the needs of their classrooms and schools (clark et al., 2012). furthermore, autonomy with regard to pd is strong in individual decision contexts and weaker in board contexts. our qualitative analysis revealed that teachers were inclined to project the nature of the decision context onto their experience and valuation of the pd activity. the individual decision context may in and of itself lead to positive valences and to the perception of autonomy regarding pd activities. the reverse holds for the board decision context, where positive valences are low, and teachers perceive a lack of autonomy. nonetheless, the perception of autonomy within pd may limit negative valences even in the case of board decisions. 4.2. the role of age groups in pd valuations age or length of teaching experience also impacts pd valuations. older teachers had more negative cognitive and affective valences towards pd in the board decision context than younger teachers. younger teachers had greater positive valences for pd in team decisions and associated the individual decision context more strongly with negative consequences for themselves. the two age groups were similar in that team decisions shielded them the most from negative consequences. glastra et de brabander 19 | f l r we found expressions of positive and negative autonomy and absence of autonomy. contrary to expectations (clark et al., 2012), younger teachers scored higher on positive autonomy, but also on negative autonomy (figure 2): younger teachers felt that too much autonomy led to incoherence in pd. this ambivalent valuation of autonomy in pd may indicate a division within the younger teacher group. career starters (years 1–6) typically focus on survival and may be insecure about their teaching capabilities (huberman et al., 1993). mid-career teachers (years 7–18; richter et al., 2011) included in our younger group have well-developed teaching and working routines (philipp & kunter, 2013) and appreciate freedom of action in pd. furthermore, an explanation for the stronger positive autonomy in the younger teachers might be that they experience autonomy where freedom of action is local and relatively restricted, whereas older teachers experience autonomy where they can contribute in more substantial ways to pd. clark et al. (2012) have pointed out that experiences of younger and older teachers regarding autonomy in their professional work differ. from this perspective, it is also possible that autonomy experiences regarding pd are impacted by the declining levels of professional autonomy that older teachers experience during their careers, while younger teachers enjoy the freedom of action that their classrooms afford them. this aspect certainly merits further investigation. our qualitative analysis (3.2.1.4.) showed that older teachers used pd primarily to supplement rather than critically reflect on their skills and knowledge. based on their more extensive experience, they acted as “gatekeepers”, inclined to deciding right away regarding what to take and what not to take from pd into their classrooms. in comparison, younger teachers were more inclined to use pd lessons for critical reflection on their teaching. they were also more inclined to suspend judgment and explore whether pd content worked in their classrooms. this resembles a tendency towards experimentation/activism (huberman et al., 1993) characteristic of mid-career teachers, who in this study were included in the younger age group. 4.3. teacher experiences and valences the analysis of experiences of teachers leading to positive and negative pd valuations yielded three broad categories: context conditions, consequences for stakeholders (students, school, and teachers), and pd qualities. concerning context conditions, we found that the nature of the decision context was often projected onto the pd activity and its consequences. this led to positive valuations mostly in the individual context, and negative valuations mostly in the board decision context. concerning stakeholder consequences, students and teachers emerged as more important than schools. above all, to be positively valued, pd needed to be perceived as being in the best interest of students. direct visibility of pd effects in classrooms helped to ensure this. this finding corresponds with lortie’s observation that teachers look for signs of teaching effect in the present and seek short-term psychic rewards (lortie, 1975). the teachers in this study viewed professional growth and an improved capacity to help students as the most important contributions of pd. as for pd qualities, the direct applicability of pd content to classroom practices and learning needs stood out. in general, direct applicability was attributed to pd in individual and (to a lesser degree) team decision contexts, rather than in board decision contexts. pd design features such as interactive learning, room for expressing personal experiences, and alignment with preferred ways of learning led to positive valuations. standardised pd formats enhanced impressions of pd aligning poorly with teacher learning needs. our findings regarding decision contexts, applicability, and pd format are consistent with clark et al. (2012, p. 93; also knight, 2015); however, our investigation points towards the positive influence of autonomy even in mandated pd, and of the team decision context, which produces higher levels of acceptance. unlike kyndt et al. (2016) we found no difference between age groups in terms of the importance accorded to the direct applicability of pd. this may be a result of the fact that kyndt et al. concentrated on informal learning, where autonomy is a defining characteristic and where practicability might be less questionable than in formal types of pd that our interviewees mostly referred to. moreover, kyndt et al. used three years’ experience as the cut-off point between younger and older teachers. glastra et de brabander 20 | f l r there is an important proviso of all of these experiences associated with valuations of pd: many teachers stated that time spent on pd could in many cases be better invested in teaching their students. 4.4. teacher experiences and autonomy our research shows that autonomy plays a multifaceted role in teacher experiences with pd. before the pd activity, autonomy plays a role in influencing pd topics and programming; during pd, it plays a role in active and participative designs; and after pd, it features in the selection and adaptation of pd content for classroom application. when teachers perceived themselves as steering their own learning or when they experienced freedom of action in pd, this was almost always valued as a positive aspect of pd. selecting and adapting pd outcomes was mostly reserved for instances where teachers doubted the fit of pd. in line with the general appreciation of direct applicability, the valuations of adapting pd for classroom purposes were only positive when deemed cost-effective and efficient. van duzor (2011, p.364) points out that “teachers are professionals who use their understanding of innovations, their classrooms, and their personal goals to impact transfer from pd to the classroom.” our research shows that this impact is dependent on the experience of autonomy in pd and a preference for direct applicability. negative autonomy related to experiences where pd gave too much room for personal choice or initiative according to teachers (particularly the younger group), thereby jeopardising course coherence. teachers frequently related the role of autonomy regarding pd to benefits for “their own” classroom or students. where professional autonomy of teachers is confined to their own classroom autonomy (clark et al., 2012), overarching and innovative learning in pd may be at risk. 4.5. effective teacher pd: design features versus decision contexts as to the discussion about effective teacher pd (borko, 2004; clark, et al., 2012; darlinghammond & richardson, 2009; desimone, 2009; garet et al., 2001; penuel et al., 2007), it can be concluded that certain design features (such as interactive learning and active use of teacher experience) are associated with positive valences and experiences of autonomy. however, the overriding importance attached to the direct applicability of pd makes it clear that design features can never determine effective pd independent of situational learning needs, school contexts, and the teachers involved. this underscores van duzor’s (2011) conclusion that the transfer of pd to classrooms cannot be achieved without the transformative potential and situational insight of teachers. moreover, the impact of decision contexts on valuations of pd programmes irrespective of their particular design features, and the association of positive autonomy with teacher influence on content and format of pd provision point more strongly towards decision context as a precondition for effective pd (kennedy, 2016; schibeci & hickey, 2004). the most important background here is the fact that teachers see themselves as best placed to judge what constitutes relevant pd. nir and bogler (2008, p. 384) conclude “that teachers’ satisfaction is likely to prevail when supervisory processes are constructed and designed to serve teachers’ actual needs rather than to meet procedural requirements determined by higher-level bureaucrats, often presenting schools with top-down programmes.” in our study, the confidence of teachers in their own judgment regarding the relevance of pd is not to be taken as an expression of pure professional autonomy, since it also reflects teachers’ short-term efficiency considerations. kennedy et al. (2008, p. 415) note a similar tendency among teachers who “tend to value more formal, structured cpd opportunities. this view would tend to suggest that teachers favour a more managerial conception of professionalism, aligning with a discourse of efficiency and accountability.” in sum, our findings demonstrate that questions of effective pd cannot be fruitfully discussed without recourse to decision contexts that govern pd participation. glastra et de brabander 21 | f l r 4.6. direct applicability and pd efficiency: attitudes of teachers our study revealed several elements of broadly shared attitudes concerning pd. teachers showed a general willingness and sense of professional duty to develop their skills through pd to serve their students. together with the experience of some degree of autonomy regarding pd, this had a compensatory effect in cases of pd that felt inadequate to teachers. however, given a strong preference for teaching duties and direct interaction with their students, pd was frequently seen as being a noncore task which could only be positive under the strict conditions of efficiency and direct applicability. possible selection bias towards teachers with positive pd motivations (see section 2.1) may have manifested in the positive learning orientation of the teachers in our sample. however, it has certainly not translated into uncritical acclaim of pd courses taken. the applicability criterion refers to teachers being able to imagine new skills or knowledge aligning smoothly with their everyday teaching practices (knight, 2015). this implies that applicable pd must show a reasonable fit with existing practices and therefore cannot be (too) inconclusive as to its short-term effects, or (too) innovative. hence, direct applicability seems to be an inherently conservative criterion that embraces small change but keeps more fundamental innovation at bay. shulman (2004) has characterised the conservatism of schools as an elastic cord pulling teachers back to existing practices. here, we see that the general attitudes of teachers themselves can form that cord. the appreciation of direct applicability and efficient delivery of pd was also associated with the increase in administrative tasks. carstensen et al. (1999) have argued that when time is perceived as expansive, people are able to focus on knowledge goals, even when these imply emotional costs or a delay of emotional rewards. however, when time is perceived to be limited, as is the case for teachers, a focus on the present is induced, which is “likely to involve goals related to feeling states, deriving emotional meaning, and experiencing emotional satisfaction” (p. 167). if the applicability and efficiency demands among teachers are indeed (also) driven by this kind of limitedtime perspective and its associated quest for affective need fulfilment (lortie, 1975), this would not be conducive to deep changes in teachers’ knowledge or teaching practices. teacher demands for pd with direct applicability may be an expression of what clark et al. (2012, p. 103) have called survival or “reactive learning”, and may constitute a danger to deep-level and inquisitive learning among teachers (cf. biesta et al., 2015; nir & bogler, 2008). 4.7. concluding remarks future research should aim to shed more light on which forces underlie the dominance of applicability and efficiency. the current study has found evidence of a strong focus on the visible effects of pd lessons among students. according to lortie (1975), this focus results from insecurity among teachers regarding their contribution to student development. this may be one of the immediate forces at play; however, other forces must also be considered, such as the workload of primary teachers, the dominant trend in pd towards short, off-the-shelf, packaged courses, and the general accountability culture with its emphasis on efficiency and effectiveness. therefore, positive valences attached to pd— enhanced by individual choice, the experience of autonomy, efficient delivery, or perceived applicability—can not be considered intrinsically good, and should therefore not be targeted instrumentally to increase pd motivations. our research highlights how teachers’ motivations to engage in pd is shaped by the educational field as a system of rules, interdependencies and power relations between educational agents. given the implied risk of restricting pd to only that which is directly applicable in practice, this impact should be tackled head-on in future research. a fruitful starting point would be pierre bourdieu’s field theory, in particular the concept of habitus, i.e., the feeling for practice that teachers develop while working in the educational field under the influence of the rules and structures of that field (bourdieu, 1990). teachers experience teaching and pd within a field that puts a premium on cost-effectiveness and efficiency and holds schools and teachers accountable for performance outcomes (gewirtz et al., 2019). increasing existing forms of motivation may inadvertently strengthen this logic, including where it will result in undesirable consequences for teaching practices. greater vision regarding the educational field, including teaching practices and the provision of pd, is glastra et de brabander 22 | f l r required to determine the practical consequences of the motivation patterns found in our analysis (cf. biesta et al., 2015). this study has demonstrated that decision contexts and teaching experience will be important building blocks in such an endeavour. keypoints the literature on effective teacher professional development (pd) stresses the importance of design features such as active learning arrangements, coherence with teacher beliefs and the external reform context, and collective participation. our research reveals the crucial influence of a factor overlooked in the literature, namely decision contexts. who decides about teacher participation in professional development, the school board, the team, or the individual teacher, impacts teacher motivation. the main criterion for judging pd relevance is direct applicability to classrooms. individual decision contexts and autonomy regarding pd both play an important role in ensuring such applicability. the dominance of direct applicability for evaluating pd was not motivated by student or school interests alone; it was also based in a logic of efficiency among primary teachers, which reflected growing work pressures and an accountability culture. among older teachers, teaching experience informed the selection of which pd content to transfer to their classrooms. younger teachers were more inclined to first explore whether pd worked in their classrooms before deciding about adoption. references apple, m. w., & jungck, s. (1990). “you don't have to be a teacher to teach this unit”: teaching, technology and gender in the classroom. american educational research journal, 27, 227-251. ajzen, i., & fishbein, m. (2008). scaling and testing multiplicative combinations in the expectancy– value model of attitudes. journal of applied social psychology, 38, 2222-2247. https://doi.org/10.1111/j.1559-1816.2008.00389.x| avidov-ungar, o. (2016). a model of professional development: teachers’ perceptions of their professional development. teachers and teaching, 22(6), 653-669. https://doi.org/10.1080/13540602.2016.1158955 ball, s. j. (2003). the teacher’s soul and the terrors of performativity. journal of education policy, 18, 215-228. https://doi.org/10.1080/0268093022000043065 ball, d. l., & cohen, d. k. (1999). developing practice, developing practitioners: toward a practicebased theory of professional education. teaching as the learning profession: handbook of policy and practice, 1, 3-22. ballet, k., & kelchtermans, g. (2008). workload and willingness to change: disentangling the experience of intensification. journal of curriculum studies, 40, 47-67. https://doi.org/10.1080/00220270701516463 bandura, a. (1997). self-efficacy: the exercise of control. new york: w. h. freeman & co. biesta, g., priestley, m., & robinson, s. (2015). the role of beliefs in teacher agency. teachers and teaching, 21(6), 624-640. https://doi.org/10.1080/13540602.2015.1044325 borko, h. (2004). professional development and teacher learning: mapping the terrain. educational researcher, 33(8), 3-15. borko, h., elliott, r., and uchiyama, k. (2002). professional development: a key to kentucky's educational reform effort. teaching and teacher education, 18, 969-987. https://doi.org/10.1016/s0742-051x(02)00054-9 glastra et de brabander 23 | f l r bourdieu, p. (1990). the logic of practice, translated by richard nice. stanford, calif.: stanford university press. carstensen, l. l., isaacowitz, d. m., & charles, s. t. (1999). taking time seriously: a theory of socioemotional selectivity. american psychologist, 54(3), 165-181. https://doi.org/10.1037/0003-066x.54.3.165 carver, c. s. (2006). approach, avoidance, and the self-regulation of affect and action. motivation and emotion, 30, 105-110. https://doi.org/10.1007/s11031-006-9044-7 clark, r., livingstone, d. w., & smaller, h., eds. (2012). teacher learning and power in the knowledge society. rotterdam, netherlands: sense. cochran-smith, m., & lytle, s. (1999). relationships of knowledge and practice: teacher learning in communities. review of research in education, 24, 249-305. https://doi.org/10.3102/0091732x024001249 darling hammond, l. (2005). teaching as a profession: lessons in teacher preparation and professional development. phi delta kappan, 87(3), 237-240. darling hammond, l, & richardson, n. (2009). teacher learning: what matters? educational leadership, 66(5), 46-53. day, c., elliot, b., & kington, a. (2005) reform, standards and teacher identity: challenges of sustaining commitment. teaching and teacher education, 21, 563-577. https://doi.org/10.1016/j.tate.2005.03.001 day, c, & gu, q. (2007). variations in the conditions for teachers. oxford review of education, 33(4), 423-443. https://doi.org/10.1080/03054980701450746 de brabander, c. j., & glastra, f. j. (2018). testing a unified theory of task-specific motivation: how teachers appraise three professional development activities. frontline learning research, 6, 54–76. https://doi.org/10.14786/flr.v4i2.342 de brabander, c. j., & glastra, f. j. (2021). the unified model of task-specific motivation and teachers’ motivation to learn about teaching and learning supportive modes of ict use. education and information technologies, 26, 393-420. https://doi.org/10.1007/s10639-020-10256-7 de brabander, c. j., & martens, r. l. (2014). towards a unified theory of task-specific motivation. educational research review, 11, 27-44. https://doi.org/10.1016/j.edurev.2013.11.001 de brabander, c. j., & martens, r. l. (2018). empirical exploration of a unified model of task-specific motivation. psychology, 9, 540–560. https://doi.org/10.4236/psych.2018.94033 desimone, l m. (2009). improving impact studies of teachers’ professional development: toward better conceptualizations and measures. educational researcher, 38(3), 181-199. https://doi.org/10.3102/0013189x08331140 devos, g, tuytens, m, & hulpia, h. (2014). teachers’ organizational commitment: examining the mediating effects of distributed leadership. the american journal of education, 20, 205-231. 0195-6744/2014/12002-0003$10.00 eccles, j. s., & wigfield, a. (2020). from expectancy-value theory to situated expectancy-value theory: a developmental, social cognitive, and sociocultural perspective on motivation. contemporary educational psychology, 61. https://doi.org/10.1016/j.cedpsych.2020.101859 elliot, a. j. (2006). the hierarchical model of approach-avoidance motivation. motivation and emotion, 30, 111-116. https://doi.org/10.1007/s11031-006-9028-7 elliot, a. j., & church, m. a. (1997). a hierarchical model of approach and avoidance achievement motivation. journal of personality and social psychology, 72, 218-232. https://doi.org/10.1037/0022-3514.72.1.218 emo, w. (2015). teachers’ motivations for initiating innovations. journal of educational change, 16(2), 171-195. https://doi.org/10.1007/s10833-015-9243-7 epskamp, s., cramer, a. o. j., waldorp, l. j., schmittmann, v. d., & borsboom, d. (2012). qgraph: network visualizations of relationships in psychometric data. journal of statistical software, 48(4), 1-18. https://doi.org/10.18637/jss.v048.i04 glastra et de brabander 24 | f l r evans, l. (2014). leadership for professional development and learning: enhancing our understanding of how teachers develop. cambridge journal of education, 44(2), 179-198. https://doi.org/10.1080/0305764x.2013.860083 fox, r. k., muccio, l. s., white, c. s., & tian, j. (2015). investigating advanced professional learning of early career and experienced teachers through program portfolios. european journal of teacher education, 38(2), 154-179. https://doi.org/10.1080/02619768.2015.1022647 garet, m. s., porter, a. c., desimone, l., birman, b. f., & yoon, k. s. (2001). what makes professional development effective? results from a national sample of teachers. american educational research journal, 38(4), 915-945. https://doi.org/10.3102/00028312038004915 gewirtz, s., maguire, m., neumann, e., & towers, e. (2019). what’s wrong with ‘deliverology’? performance measurement, accountability and quality improvement in english secondary education, journal of education policy, 36(4), 504-529. https://doi.org/10.1080/02680939.2019.1706103 glaser, b. g., & strauss, a. l. (1967). the discovery of grounded theory: strategies for qualitative research. chicago il: aldine. hall, c., & noyes, a. (2009). new regimes of truth: the impact of performative school self evaluation systems on teachers' professional identities. teaching and teacher education, 25(6), 850-856. https://doi.org/10.1016/j.tate.2009.01.008 hardy, i., & lingard, b. (2008). teacher professional development as an effect of policy and practice: a bourdieuian analysis. journal of education policy, 23, 63-80. https://doi.org/10.1080/02680930701754096 hargreaves, a. (2003). teaching in the knowledge society: education in the age of insecurity. new york and london: teachers college press. hashweh, m. z. (2003). teacher accommodative change. teaching and teacher education, 19(4), 421434. https://doi.org/10.1016/s0742-051x(03)00026-x hawley, w., & valli, l. (1999). the essentials of effective professional development: a new consensus. in l. darling-hammond & g. sykes (eds.), teaching as the learning profession: handbook of policy and practice (pp.127-150). san francisco: jossey-bass. hochberg, e. d., & desimone, l. m. (2010). professional development in the accountability context: building capacity to achieve standards. educational psychologist, 45(2), 89-106. https://doi.org/10.1080/00461521003703052 huberman, m. (1989). the professional life cycle of teachers. teachers college record, 91, 31-57. huberman, m., gronauer, m.-m., & marti, j. (1993). the lives of teachers. london: cassell. jansen in de wal, j., den brok, p. j., hooijer, j. g., martens, r. l., & van den beemt, a., (2014). teachers' engagement in professional learning: exploring motivational profiles. learning and individual differences, 36, 27-36. https://doi.org/10.1016/j.lindif.2014.08.001 karsten, s. (1999). neoliberal education reform in the netherlands. comparative education, 35(3), 303317. https://doi.org/10.1080/03050069927847 kennedy, a., christie, d., fraser, c., reid, l., mckinney, s., welsh, m., wilson, a., & griffiths, m. (2008). teacher learning: key informants’ perspectives on teacher learning in scotland. british journal of educational studies, 56(4), 400–419. https://doi.org/10.1111/j.14678527.2008.00416.x kennedy, m. m. (2016). how does professional development improve teaching? review of educational research, 86(4), 945–980. https://doi.org/10.3102/0034654315626800 knight, r. (2015). postgraduate student teachers’ developing conceptions of the place of theory in learning to teach: ‘more important to me now than when i started’. journal of education for teaching, 41(2), 145-160. https://doi.org/10.1080/02607476.2015.1010874 krapp, a. (2005). basic needs and the development of interest and intrinsic motivational orientations. learning and instruction, 15(5), 381-395. doi.org/10.1016/j.learninstruc.2005.07.007 kwakman, k. (2003). factors affecting teachers’ participation in professional learning activities. teaching and teacher education, 19(2), 149-170. https://doi.org/10.1016/s0742051x(02)00101-4 glastra et de brabander 25 | f l r kyndt, e., gijbels, d., grosemans, i., & donche, v. (2016). teachers’ everyday professional development mapping informal learning activities, antecedents, and learning outcomes. review of educational research, 86(4), 1111–1150. https://doi.org/10.3102/0034654315627864 leithwood, k., steinbach, r., & jantzi, d. (2002). school leadership and teachers’ motivation to implement accountability policies. educational administration quarterly, 38(1), 94-119. https://doi.org/10.1177/0013161x02381005 little, j. w. (1993). teachers' professional development in a climate of educational reform. educational evaluation and policy analysis, 15, 129-151. little, j. w. (2006). professional community and professional development in the learning-centered school. arlington, va: education association national. little, j. w., & bartlett, l. (2010). the teacher workforce and problems of educational equity. review of research in education, 34, 285-328. https://doi.org/10.3102/0091732x09356099 locke, t., vulliamy, g., webb, r., et al. (2005). being a ‘professional’ primary school teacher at the beginning of the 21st century: a comparative analysis of primary teacher professionalism in new zealand and england. journal of education policy, 20, 555-581. https://doi.org/10.1080/02680930500221784 lortie, d. (1975). schoolteacher. a sociological study. chicago: university of chicago press. maskit, d. (2011). teachers’ attitudes toward pedagogical changes during various stages of professional development. teaching and teacher education, 27(5), 851-860. https://doi.org/10.1016/j.tate.2011.01.009 meyers, b., meyers, j., & gelzheiser, l. (2001). observing leadership roles in shared decision making: a preliminary analysis of three teams. journal of educational and psychological consultation, 12(4), 277-312. https://doi.org/10.1207/s1532768xjepc1204_01 milner, h. r. (2013). policy reforms and de-professionalization of teaching. boulder, co: national education policy center. retrieved from http://nepc.colorado.edu/publication/policy-reformsdeprofessionalization ministerie van ocw (2012). nota werken in het onderwijs 2012. [memorandum working in the field of education] the hague, the netherlands: ministerie van ocw. moore, a., edwards, g., halpin, d., & george, r. (2002). compliance, resistance and pragmatism: the (re) construction of schoolteacher identities in a period of intensive educational reform. british educational research journal, 28(4), 551-565. https://doi.org/10.1080/0141192022000005823| nir, a., & bogler, r. (2008). the antecedents of teacher satisfaction with professional development programs. teaching and teacher education, 24(2), 377-386. https://doi.org/10.1016/j.tate.2007.03.002 onderwijsraad (2005). kwaliteit en inrichting van de lerarenopleidingen. briefadvies aan de tweede kamer [quality and organisation of teacher education. advice letter to the lower house of parliament]. den haag: onderwijsraad. onderwijscoöperatie (2016). de staat van de leraar 2016. [the state of the teacher 2016] retrieved from http://www.fvov.nl/wp-content/uploads/2016/04/w-20160413-staat_van_de_leraar.pdf pearson, l. c., & moomaw, w. (2006) continuing validation of the teaching autonomy scale. the journal of educational research, 100(1), 44-51, https://doi.org/10.3200/joer.100.1.44-51 peetsma, t. (2000). future time perspective as a predictor of school investment. scandinavian journal of educational research, 44(2), 177-192. https://doi.org/10.1080/713696667 penuel, w. r., fishman, b. j., yamaguchi, r., & gallagher, l. p. (2007). what makes professional development effective? strategies that foster curriculum implementation. american educational research journal, 44, 921–958. https://doi.org/10.3102/0002831207308221 philipp, a., & kunter, m. (2013). how do teachers spend their time? a study on teachers. teaching and teacher education, 35, 1-12. doi.org/10.1016/j.tate.2013.04.014 reay, d. (2015). habitus and the psychosocial: bourdieu with feelings. cambridge journal of education, 45(1), 9-23. https://doi.org/10.1080/0305764x.2014.990420 glastra et de brabander 26 | f l r reeve, j., nix, g., & hamm, d. (2003). testing models of the experience of self-determination in intrinsic motivation and the conundrum of choice. journal of educational psychology, 95, 375– 392. https://doi.org/10.1037/0022-0663.95.2.375 research voor beleid (2011). tussenmeting convenant leerkracht 2011 [intermediate evaluation covenant teacher 2011]. zoetermeer, netherlands: research voor beleid. richter, d., kunter, m., klusmann, u., lüdtke, o., & baumert, j. (2011). professional development across the teaching career: teachers’ uptake of formal and informal learning opportunities. teaching and teacher education, 27(1), 116-126. https://doi.org/10.1016/j.tate.2010.07.008 ryan, r. m., & deci, e. l. (2000). self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. the american psychologist, 55, 68-78. https://doi.org/10.1037/0003-066x.55.1.68 ryan, r. m., & deci, e. l. (2020). intrinsic and extrinsic motivation from a self-determination theory perspective: definitions, theory, practices, and future directions. contemporary educational psychology, 61:101860. https://doi.org/10.1016/j.cedpsych.2020.101860 schibeci, r. a., & hickey, r. l. (2004). dimensions of autonomy: primary teachers' decisions about involvement in science professional development. science education, 88(1), 119-145. https://doi.org/10.1002/sce.10091| schmidt, r. a., & bjork, r. a. (1992). new conceptualizations of practice: common principles in three paradigms suggest new concepts for training. psychological science, 3(4), 207-217. https://doi.org/10.1111/j.1467-9280.1992.tb00029.x scribner, j. p., sawyer, r. k., watson, s. t., & myers, v. l. (2007). teacher teams and distributed leadership: a study of group discourse and collaboration. educational administration quarterly, 43(1), 67-100. https://doi.org/10.1177/0013161x06293631 schunk, d. h., & dibenedetto, m. k. (2020). motivation and social cognitive theory. contemporary educational psychology, 60. https://doi.org/10.1016/j.cedpsych.2019.101832 seor (2012). inventarisatie van lacunes in het opleidingsen scholingsaanbod: eindrapport [inventory of shortcomings in education and training opportunities: final report]. rotterdam, netherlands: seor. shulman, l. (2004). the wisdom of practice. san francisco, ca: jossey bass. skaalvik, e. m., & skaalvik, s. (2014). teacher self-efficacy and perceived autonomy: relations with teacher engagement, job satisfaction, and emotional exhaustion. psychological reports, 114(1), 68-77. https://doi.org/10.2466/14.02.pr0.114k14w0 smylie, m. a. (1992). teacher participation in school decision-making assessing willingness to participate. educational evaluation and policy analysis, 14(1), 53-67. https://doi.org/10.3102/01623737014001053 stamos (2013). werkgelegenheid, in personen, in het basisonderwijs [employment in primary education]. retrieved from http://www.stamos.nl/index.rfx?verb=showitem&item=3.2 starkey, l., yates, a., meyer, l. h., hall, c., taylor, m., stevens, s., & toia, r. (2009). professional development design: embedding educational reform in new zealand. teaching and teacher education, 25, 181-189. https://doi.org/10.1016/j.tate.2008.08.007 stevenson, h., carter, b., & passy, r. (2007) ‘new professionalism’, workforce remodelling and the restructuring of teachers’ work. international journal of leadership for learning, 11(15). retrieved from https://journals.library.ualberta.ca/iejll/index.php/iejll/article/download/670/331/669 strauss, a., & corbin, j. (1998). basics of qualitative research: techniques and procedures for developing grounded theory. thousand oaks: sage swann, m., mcintyre, d., pell, t., hargreaves, l., & cunningham, m. (2010). teachers' conceptions of teacher professionalism in england in 2003 and 2006. british educational research journal, 36, 549 -571. https://doi.org/10.1080/01411920903018083| thoonen, e. e. j., sleegers, p. j. c., oort, f. j., peetsma, t. t. d., & geijsel, f. p. (2011). how to improve teaching practices: the role of teacher motivation, organizational factors, and glastra et de brabander 27 | f l r leadership practices. educational administration quarterly, 47(3), 496-536. https://doi.org/10.1177/0013161x11400185 van eekelen, i. m., vermunt, j. d., & boshuizen, h. p. a. (2006). exploring teachers' will to learn. teaching and teacher education, 22(4), 408–423. https://doi.org/10.1016/j.tate.2005.12.001 van duzor, a. g. (2011). capitalizing on teacher expertise: motivations for contemplating transfer from professional development to the classroom. journal of science education and technology, 20(4), 363-374. https://doi.org/10.1007/s10956-010-9258-z vermunt, j. d., & endedijk, m. d. (2011). patterns in teacher learning in different phases of the professional career. learning and individual differences, 21, 294–302. doi.org/10.1016/j.lindif.2010.11.019 vescio, v., ross, d., & adams, a. (2008). a review of research on the impact of professional learning communities on teaching practice and student learning. teaching and teacher education, 24(1), 80-91. https://doi.org/10.1016/j.tate.2007.01.004 wallace, c. s., & priestley, m. (2011). teacher beliefs and the mediation of curriculum innovation in scotland: a socio‐cultural perspective on professional development and change. journal of curriculum studies, 43(3), 357-381. https://doi.org/10.1080/00220272.2011.563447 webb, r., vulliamy, g., hämäläinen, s., sarja, a., kimonen, e., & nevalainen, r. (2004). a comparative analysis of primary teacher professionalism in england and finland. comparative education, 40(1), 83-107. https://doi.org/10.1080/0305006042000184890 wigfield, a., & eccles, j. s. (2000). expectancy–value theory of achievement motivation. contemporary educational psychology, 25, 68-81. https://doi.org/10.1006/ceps.1999.1015 wilson, s. m., & berne, j. (1999). teacher learning and the acquisition of professional knowledge: an examination of research on contemporary professional development. review of research in education, 24, 173-209. https://doi.org/10.3102/0091732x024001173 glastra et de brabander 28 | f l r appendix 1. tables table 1 number of codes (n) and number (nr) and proportion (pr) of teachers with codes, and proportion of teachers with co-occurrences of codes in the group of younger teachers with whom the board decision context was discussed (n=34) n nr pr aut nav ncvl ncvp ncvs pav pcvl pcvp aut 21 16 0.47 nav 25 13 0.38 0.06 ncvll 7 5 0.15 0.00 0.03 ncvp 46 22 0.65 0.06 0.21 0.03 ncvs 24 15 0.44 0.03 0.12 0.12 0.06 pav 35 20 0.59 0.12 0.06 0.00 0.06 0.00 pcvl 32 20 0.59 0.03 0.03 0.00 0.06 0.00 0.12 pcvp 46 24 0.71 0.09 0.03 0.00 0.06 0.03 0.18 0.32 pcvs 37 19 0.56 0.06 0.00 0.00 0.03 0.12 0.09 0.24 0.15 table 2 number of codes (n) and number (nr) and proportion (pr) of teachers with codes, and proportion of teachers with co-occurrences of codes in the group of older teachers with whom the board decision context was discussed (n=26) n nr pr aut nav ncvl ncvp ncvs pav pcvl pcvp aut 13 13 0.50 nav 21 13 0.50 0.04 ncvll 4 2 0.08 0.00 0.00 ncvp 36 23 0.88 0.04 0.35 0.04 ncvs 10 9 0.35 0.00 0.00 0.00 0.08 pav 30 16 0.62 0.08 0.04 0.00 0.15 0.00 pcvll 19 13 0.50 0.04 0.00 0.00 0.04 0.00 0.08 pcvp 41 19 0.73 0.12 0.12 0.00 0.15 0.00 0.27 0.31 pcvs 25 16 0.62 0.00 0.00 0.00 0.04 0.04 0.08 0.15 0.23 glastra et de brabander 29 | f l r table 3 number of codes (n), number (nr) and proportion (pr) of teachers with codes, and proportion of teachers with co-occurrences of codes in the group of younger teachers with whom the team decision context was discussed (n=25) n nr pr aut nav ncvl ncvp ncvs pav pcvl pcvp aut 18 13 0.52 nav 8 7 0.28 0.08 ncvll 4 4 0.16 0.00 0.00 ncvp 17 12 0.48 0.00 0.12 0.08 ncvs 10 7 0.28 0.04 0.00 0.04 0.08 pav 41 19 0.76 0.04 0.08 0.00 0.00 0.00 pcvll 38 23 0.92 0.04 0.04 0.04 0.00 0.00 0.20 pcvp 55 22 0.88 0.12 0.04 0.00 0.00 0.00 0.32 0.52 pcvs 45 23 0.92 0.08 0.00 0.00 0.00 0.04 0.32 0.40 0.60 table 4 number of codes (n), number (nr) and proportion (pr) of teachers with codes, and proportion of teachers with co-occurrences of codes in the group of older teachers with whom the team decision context was discussed (n=34) n nr pr aut nav ncvl ncvp ncvs pav pcvl pcvp aut 18 18 0.53 nav 10 9 0.26 0.03 ncvll 8 7 0.21 0.00 0.00 ncvp 15 11 0.32 0.03 0.03 0.06 ncvs 17 12 0.35 0.00 0.06 0.12 0.09 pav 43 19 0.56 0.06 0.06 0.03 0.03 0.03 pcvll 39 21 0.62 0.06 0.00 0.00 0.06 0.00 0.09 pcvp 49 21 0.62 0.06 0.06 0.00 0.06 0.00 0.24 0.24 pcvs 44 27 0.79 0.03 0.03 0.03 0.00 0.06 0.15 0.27 0.18 glastra et de brabander 30 | f l r table 5 number of codes (n), number (nr) and proportion (pr) of teachers with codes, and proportion of teachers with co-occurrences of codes in the group of younger teachers with whom the individual decision context was discussed (n=28) n nr pr aut nav ncvl ncvp ncvs pav pcvl pcvp aut 24 20 0.71 nav 2 2 0.07 0.04 ncvll 2 1 0.04 0.04 0.04 ncvp 32 22 0.79 0.07 0.07 0.04 ncvs 8 6 0.21 0.00 0.00 0.00 0.04 pav 60 27 0.96 0.14 0.04 0.00 0.04 0.04 pcvll 44 21 0.75 0.07 0.00 0.00 0.04 0.04 0.29 pcvp 102 28 1.00 0.36 0.00 0.00 0.07 0.07 0.57 0.54 pcvs 46 22 0.79 0.00 0.00 0.00 0.04 0.00 0.25 0.39 0.50 table 6 number of codes (n), number (nr) and proportion (pr) of teachers with codes, and proportion of teachers with co-occurrences of codes in the group of older teachers with whom the individual decision context was discussed (n=33) n nr pr aut nav ncvl ncvp ncvs pav pcvl pcvp aut 18 17 0.52 nav 9 7 0.21 0.03 ncvll 4 3 0.09 0.00 0.00 ncvp 25 15 0.45 0.00 0.03 0.06 ncvs 9 6 0.18 0.00 0.00 0.03 0.03 pav 76 31 0.94 0.09 0.06 0.00 0.12 0.06 pcvll 41 23 0.70 0.03 0.00 0.00 0.03 0.00 0.15 pcvp 96 33 1.00 0.18 0.03 0.03 0.18 0.03 0.61 0.49 pcvs 33 22 0.67 0.06 0.00 0.00 0.03 0.00 0.24 0.24 0.49 glastra et de brabander 31 | f l r 2. citations referred to in the article by respondent number only § 3.2.1.1. o35:19 “i thought those social media laws, that’s not my thing. then i think, but i can do nothing about it, it is a mandatory course, so you can say i don’t want to work with it, but you have to.” o58:3 “it just feels very nice to know how to act, to know that you know, and you can do it and then see the effect. it is just a good feeling, for yourself and the child.” y66:8 regarding a course in cooperative learning “it is very visible throughout the school. it works the same for everyone. so if i teach a different class and i raise my hand they just know what i mean.” y82:10 “it is interesting to work together with colleagues on something, and to know that you’re not alone (…) it inspires you to keep talking to each other and find out things together.” § 3.2.1.2. o92:11 on the consequences of participating in a course “nieuwsbegrip’ (understanding news) “more enthusiasm about reading comprehension, that’s what it has brought about, that children are much more motivated to engage in it.” y48:7 following a negatively valued course: “no, of course, children are not so aware of what happened behind the scenes. i think for instance that remedial teaching was of inferior quality (…) but the children wouldn’t notice.” y60:1 “the inspectorate says ‘you should do it like this’ or ‘you should look at it this way’, but then everybody supports it as a team. it was hard work, mainly because the inspectorate put pressure on it. we were put on orange, and of course you want to get a green ball as a school.” § 3.2.1.3. y9:12 “each two weeks a certain didactic format is central, you should really practice with it and sometimes that’s a … then i am already completely tied to my lesson plan. (…) you have to manage so many things, that cooperative didactics loses out.” o70:11 ‘that course (…) was just a flood of theory and in that case, you go home feeling dissatisfied, because you can do nothing with it.” y23:12 “of school board members make up lots of things behind their desks. and then it is difficult to keep an open eye for practice. so i always kindly invite board members to walk in our shoes for a week and then come up with the same plan.” o29:6 “every year something else comes up. (…) and that’s a shame. you think, okay, we’re in the process of implementing this, and then something new is discovered and we’re spending a whole dayseminar on that again.” o11:4 “you can have a lot of people in that room. and they give a lecture, but i prefer reading the material myself. so i think that’s a waste of time.” y79:6 “the learning process was its strongest point, i think. they always provide a general curriculum that you can personalize. what is important for me differs from what’s important for others around me.“ o18:11 “i would have preferred that the course had stated its goals clearly from the outset. this and this is what you are expected to have mastered at the end of the course.” y17:4 following a course on stimulating literacy and about ways of preparing students for short literacy tests ”i like to experiment, does this work, can i do it and is it fun.” glastra et de brabander 32 | f l r y64:2 “and what is important too, it was taught in an attractive, exciting way by the course leader.” o35:16 “my expectations for the day were not too high. (…) the legal framework for social media, i thought, ‘geez, we’re in for a long story’. she [the pd instructor] hadn’t much choice, did she? since you have to tell it and i don’t think that’s inspiring.” o43:10 “it wasn’t a real education course, it was workshops äll the way (…) you hear the theory for once, you see some examples, you may practice for half an hour in a training centre. and then you go back to your school and do it all by yourself, in between teaching and parent-teacher conferences.” § 3.2.1.4. y9:16 “that you are working with it and you really see that it does something for the child and for yourself as well, when you see real progress.” y60:7 “the course was very intense. but then, when you see the standardised test scores printout and you think, ‘geez, have they improved’, then i think that’s positive. i haven’t spent my time in vain.” y65:4 on the consequences of participating in this pd activity “looking critically at our children here at school. looking critically at yourself. are you...ehm are you observing adequately… ehm where are your weaknesses? are you prejudiced?” § 3.2.2. y52:5 “i was steering my learning process myself in this activity, yes, but that was because i organized it myself.” y48:4 “yes, i had some influence on the course. of course, there was a fixed format: this is the way to make a plan, but h. [course leader] is always very … we make it really fit the school, so i have left out some things and added others and he was okay with all of that.” o50:5 “if i told them, ‘look here, this is way too fast for me’ they said to me ‘come sit next to me’ and ‘this way you can develop a group plan (…) ’ and ‘we have implemented it like this, have you thought about that?’ i got very good guidance.” y17:11-21 on her mandatory master in special educational needs “you had to desal with tests, six tests during the course and in the end you were obliged to pass the exam, that was it, there was some pressure involved, you had to do this and you had to do that, and that’s not my cup of tea (…) and we had to make a test about change processes and you were allowed to choose a case. i thought it would be nice to choose a case that i like and that meets the goals of the test and that would help my team. and this how i came to work on collaborative learning and wrote a complete handbook for each teacher.” § 3.2.3. o18:8 “of course, it always brings you something, it’s what you get out of it yourself. and you always hear new things, that make you think (…). however, whether you really have to sit out a whole afternoon with all the work waiting for you…?” y36:16 “a whole day, for me a waste of time, (…) but for the students as well, because it could have been a day more to learn something for them.” y9:15) “it looks as if i have an administrative job next to my teaching which is in itself a lot of work (…) and you have to pay attention always, because it doesn’t belong to… it’s something extra.” glastra et de brabander 33 | f l r 3. coding categories cognitive valences 3.1. school context definition: the school as a decision and work context with its task loads, work routines and schedules, colleagues, and leadership impacting teacher pd valuations. text examples: see § 3.2.1.1. paper and appendix 2. transfer definition: transfer of (positive or negative) valuations of the locus of decision-making regarding pd participation unto the valuation of ensuing pd experiences per se. fit with workload and work schedule definition: levels of fit of pd with workload and work schedules (positively or negatively) impacting pd valuations. support of the principal for pd participation definition: (lack of) support by school principals in their participation in pd impacting teacher pd valuations. 3.2. stakeholder consequences of pd definition: positive and/or negative consequences of pd participation for stakeholders such as the students, the teachers personally and the school impacting teacher pd valuations. text examples: see § 3.2.1.2. paper and appendix 2. students definition: both students in the classroom of the teacher as those in their school. teachers definition: text fragments must pertain to or include the interviewed teacher personally. school definition: consequences for the school as an organization, for all the schoolteachers, for the school curriculum, school policies, and for (the relations with) the parents of the students. 3.3. qualities of pd content and design definition: content and design and delivery of pd as such or in relation to learning needs and preferences of teachers and their classrooms impacting teacher pd valuations. text examples: see § 3.2.1.3. paper and appendix 2. immediate applicability of pd definition: pd content and/or what teachers learn from pd permits or impedes speedy application in their classroom. novelty of pd content definition: (lack of) earlier experience with pd content impacting teachers pd valuations. pd learning arrangements definition: text fragments referring to ways of distributing, reflecting on and working with pd information and knowledge and connecting it to learning needs and preferences of teachers and to classroom and school realities. glastra et de brabander 34 | f l r 4. coding categories autonomy definition: text fragments in which teachers refer to their steering of important aspects of pd or their making use of perceived freedom of action related to pd. instances of steering can occur in preparation of, during and in selecting what is (not) taken from pd to the school. text examples: see § 3.2. 2. paper and appendix 2. positive autonomy definition: the experience of steering one’s own learning process in pd or of being offered a level of freedom of action within pd is valued positively by the teacher. negative autonomy definition: the experience of steering one’s own learning process in pd or of being offered a level of freedom of action within pd is valued negatively by the teacher. absence of autonomy definition: the feeling of steering one’s own learning process or the experience of freedom of action with regard to pd was absent and this absence was not linked to positive nor negative valuations. codepen rogiers et al frontline learning research vol.8 no. 3 special issue (2020) 40 62 issn 2295-3159 opening the black box of students’ text-learning processes: a process mining perspective amelie rogiersa, emmelien merchiea, & hilde van keera adepartment of educational studies, ghent university, ghent, belgium article received 26 june / revised 11 september / accepted 3 january / available online 30 march abstract the current study uncovers secondary school students’ actual use of text-learning strategies during an individual learning task by means of a concurrent self-reported thinking aloud procedure. think-aloud data of 51 participants with different learning strategy profiles, distinguished based on a retrospective self-report questionnaire (i.e., 15 integrated strategy users, 15 information organizers, 10 mental learners, and 11 limited strategy users), were analysed by means of educational process mining. both the frequency of students’ strategy use, as well as the temporal patterns between these strategies were studied. the process mining results clearly demonstrated differences between the strategy profiles with respect to the frequency of their applied strategies, as well as concerning the temporal sequences wherein strategies were applied throughout the course of students’ text-learning process. the added value of combining both retrospective and concurrent self-report measures of students’ strategies as well as conducting process mining analysis is discussed. keywords: process mining; learner profiles; think-aloud protocol analysis; on-line measures; off-line measures info corresponding author email: amelie.rogiers@ugent.be doi: https://doi.org/10.14786/flr.v8i3.527 1. introduction recently, both educational researchers and practitioners have emphasized the importance of adjusted or personalized curricula wherein both the instructional content and methods are tailored to students’ individual learning needs (deed et al., 2014). this is also recognized by the oecd learning framework 2030 (2018) advocating the importance of learner-oriented teaching and learning. in view of contributing to the evidence-based design of personalized curricula, educational researchers are concerned with both measures and data analysis approaches to fully map and understand individual students’ learning. considering these measures, it is clear that the inclusion of both off-line and on-line instruments for measuring students’ learning is preferable given their complementary properties (veenman, 2011). while off-line measures are administered prospectively or retrospectively to performance on a learning task (e.g., self-report questionnaire data), on-line measures are gathered concurrently during task performance (e.g., think-aloud protocol or verbal self-report data). consequently, while off-line measures enable researchers to uncover learners’ perceptions of which and how often certain strategies are applied during learning, on-line measures additionally enable to map how and when these strategies are actually applied throughout the learning process (i.e., in which sequence strategies are applied or which switches occur between strategies; merchie & van keer, 2014). in this respect, researchers increasingly advocate to combine both measures in view of gaining rich and detailed insight into both students’ perceptions and actual strategic behaviour (bråten & samuelstuen, 2007; veenman, 2005). as to the data analysis approaches for gaining insight into students’ learning processes, researchers call progressively for applying a more person-oriented approach, next to the rather dominant variable-oriented approach focusing primarily on analysing relationships among variables (alexander et al., 2018; fryer & vermunt, 2017). such a person-oriented approach is highly recommended as it emphasizes the study of naturally occurring clusters or profiles in students’ learning (bergman et al., 2003). stemming from a person-oriented approach on students’ text-learning strategies (i.e., strategies to select, organize, condense, and retain text information in a more memorable form; rogiers et al., 2019a; weinstein et al., 2011), previous research already succeeded to identify learning strategy profiles in a large sample of 1,931 secondary school students (rogiers et al., 2019a). four learning strategy profiles, in which students differently combine diverse strategies during text learning, were determined based on a retrospective self-report questionnaire. more particularly, integrated strategy users (isu) were identified as learners with the most preferable profile, as they engaged in the strategic combination of different covert (i.e., non-observable, e.g., elaborating) and overt (i.e., observable, e.g., summarizing), cognitive (e.g., elaborating), and metacognitive (e.g., monitoring) text-learning strategies, and outperformed their peers on a subsequent performance test. the information organizers (io) frequently applied text-noting strategies (i.e., highlighting, summarizing) and reported limited use of mental learning strategies. conversely, mental learners (ml) restricted their repertoire to covert mental learning strategies (i.e., rereading, paraphrasing) without text-noting strategy use. finally, limited strategy users (lsu) were considered as the non-strategic or less preferable profile, as they mainly focused on the frequent application of one single text-learning strategy (i.e., highlighting, rereading) and obtained the lowest performance scores afterwards. these learning strategy profiles were also identified in late elementary education (merchie et al., 2014) and in subsequent samples of secondary school students (rogiers et al., 2018, 2019a). hence, the abovementioned learning strategy profiles were already corroborated several times in different age groups and independent study samples. to date, however, there is a gap in the literature when it comes to research providing insight into the temporal sequences in which certain strategies are applied differently by learners during the course of their text-learning process. in this respect, it is seldom investigated which strategy switches unfold during this process (cromley & wills, 2016). although current theories of (text) learning implicitly or explicitly state to account for what happens during this process, this matter has rarely been tested empirically with sequential analyses of real-time process data (e.g., from think-aloud protocol transcripts; cromley & wills, 2016). in this respect, the present study contributes to the first explicit question regarding self-report data tackled throughout the different contributions in this special issue. more particularly, it is believed that the learning process of a strategic learner can be characterised as cyclical and adaptive. first, from a self-regulated learning (srl) perspective, students’ learning process is considered as a cyclical process, consisting of different phases occurring before, during, and after learning (i.e., forethought, performance, and reflection phase; zimmerman, 2002). these phases are not viewed as linearly structured, but considered dynamic and iterative (panadero, 2017; pintrich, 2000; zimmerman, 2002). second, next to the general comprehensive models of srl, also more domain-specific learning strategy models (i.e., good strategy user model by pressley et al., 1987; model of strategic learning by weinstein et al., 2011; model of domain learning by alexander, 1998) point to the importance of adaptive strategy use, which encompasses engaging deliberately and flexibly in the use of various strategies. rather than following a linear and rigid approach to text learning, strategic learners are believed to undertake learning in an adaptive way, wherein they interactively return to prior learning activities or phases when necessary (alexander & jetton, 2000; mcnamara, ozuru et al., 2007; simpson & nist, 2000; wade et al., 1990). prior research already indicated that high achievers appear to be mostly integrated strategy users, adopting various strategies (merchie et al., 2014; rogiers et al., 2019a) and regulate their learning process more effectively (stoeger et al., 2015). however, it is unclear whether these process statements as put forward in different theoretical models can be grounded empirically and whether and how exactly specific strategy sequences unfold during the course of learners’ text-learning process (cromley & wills, 2016). in this respect, it is often difficult to grasp the cyclical and adaptive nature of learning processes as described in the abovementioned theoretical models by means of retrospective self-report measures after learning occurred. it is therefore necessary to analyse students’ learning process in a more fine-grained way (i.e., occurrence after occurrence) as it unfolds in real time during learning. opening this black box and gaining insight into the cyclical and adaptive nature in which particular sequences and strategies unfold throughout students’ learning process can offer valuable starting points for providing learner-oriented teaching and learning. if certain effective sequences between strategies come to the fore, for example, then not only strategies, but also effective sequences of applied strategies should be taught (e.g., from one learning strategy to another). a promising and emerging technique to gain systematic insight into these sequences and analyse students’ concurrent self-reports (e.g., think-aloud protocols) is educational process mining (epm). the idea behind epm is to discover, monitor, and improve students’ actual learning processes by extracting knowledge from recorded time stamps (bannert et al., 2014). a time stamp refers to the moment wherein the learner is executing or initiating a certain learning activity (e.g., highlighting, rereading). by means of these timestamped activities derived from learners’ observed learning behaviour, compact educational process models are composed (van der aalst, 2011). these process models provide an overview of both learners’ executed activities and the paths that occur between these activities. whereas the activities map the number and frequency of certain applied strategies, the paths represent how, and in which sequences these strategies were adopted throughout the learning process (fluxicon, 2019; van der aalst, 2011). as such, epm enables to visualise students’ learning behaviour and facilitates a thorough understanding of the course of students’ complex real-time learning process. in the context of srl, for example, epm research has shown that university students’ sequences of selfor group-regulatory activities differed among successful and less successful students (e.g., schoor & bannert, 2012). however, the application of epm in educational research is still in its infancy and, to date, the temporal order of students’ applied strategies during task completion has been widely neglected (bannert et al., 2014; reimann, 2007). more in-depth epm analyses can, therefore, yield valuable insights into students’ learning process and can complement off-line measures of students’ applied strategy use. in this respect, it also enables to investigate to which degree retrospective self-report measures accurately reflect students’ actual strategy use that is revealed while concurrently thinking aloud. our study adds to the literature by systematically analysing real-time think-aloud protocol (further referred to as ‘tap’) data from students with different learning strategy profiles who are requested to learn an informative text and by considering strategy sequences by means of epm. further, this study adds to the literature by confronting the frequency of students’ text-learning strategies as measured via concurrent measures on the one hand (i.e., tap) and retrospective measures on the other hand (i.e., a task-specific self-report questionnaire) and study their overlap (rogiers et al., 2019b). 1.1 the present study by means of epm, the current study investigates students’ actual use of text-learning strategies when executing an independent learning task while thinking aloud. in a first step, this study aims to examine the frequency of students’ occurred text-learning strategies depending on their learning strategy profile (rq1). referring to previous research using task-specific self-report questionnaires (merchie et al., 2014; rogiers et al., 2018, 2019a), we hypothesize more frequent verbalisations of various text-learning strategies in integrated strategy users, and less frequent and diverse strategy verbalisations in limited strategy users. in addition, we expect more frequent verbalizations of the application of overt text-noting strategies (e.g., summarizing) in information organizers and a predominant application of covert mental learning strategies (e.g., paraphrasing) in mental learners. in a second step, this study aims to explore temporal patterns in students’ text-learning process based on the sequences in which their strategies are applied (rq2). as the theoretical and empirical literature indicates that particularly effective learners apply diverse strategies in a flexible and systematic way (alexander & jetton, 2000; mcnamara et al., 2007; rogiers et al., 2019a; simpson & nist, 2000; wade et al., 1990; weinstein et al., 2011), we hypothesize a more cyclical use of text-learning strategies in integrated strategy users, including more recursive patterns between their applied strategies. conversely, a more linear and unidirectional text-learning process is expected in limited strategy users. 2. methodology 2.1 participants a think-aloud study was carried out with 51 secondary school students (62.75% seventh and 37.25% eight graders) from 11 schools and 51 classes who were part of a large-scale study (n = 1,931, rogiers et al., 2019a). based on a large-scale cluster analysis, 15 integrated strategy users, 15 information organizers, 10 mental learners, and 11 limited strategy users (n = 51 participants) were identified within the sample of the think-aloud study. the sample consisted of 70.59% girls and 29.41% boys, with an overall mean age of 12.99 years (sd = .69). the majority of the students (87.23%) were native dutch speakers, which is the language of instruction in flanders (the dutch speaking part of belgium). all participants and their parents agreed to participate in the tap administration by means of informed consent. 2.2 instruments and procedure the data collection procedure consisted of several steps. figure 1 provides a visual representation of the data collection procedure. figure 1. chronological representation of the data collection procedure. as can be seen in both last steps of the procedure, a combination of two self-report measures was opted for in the context of the present study, respectively an on-line and concurrent think-aloud measure on the one hand and an off-line, retrospective questionnaire on the other hand. 2.2.1 practice session in thinking aloud following the recommendations of prior research (greene et al., 2011; van someren et al., 1994), a 20-minute practice session in thinking out loud was organised by the researcher to familiarize students with the think-aloud method. this practice session was based on prior research in a similar age group (merchie & van keer, 2014; vandevelde et al., 2005) and consisted of three phases. in a first phase, the researcher thoroughly explained the purpose and procedure of the think-aloud method. second, the researcher modelled thinking aloud during an origami assignment. no learning task was opted for practicing thinking out loud to avoid possible training effects (afflerbach & johnston, 1984; greene et al., 2011). the origami assignment provided ample opportunities for self-regulation. for instance, a step-by-step approach could be followed and there were ample opportunities for students to evaluate or adjust their approach. third, an individual practice phase took place in which students practiced thinking out loud. during this session, students were asked to fold an origami cat while verbalizing their thoughts, feelings, and actions. during the practice session, feedback was provided on students’ verbalisations in view of optimizing their thinking aloud. accordingly, no feedback on students’ approach was provided. the researcher prompted the student to continue verbalizing when (a) meaningful silences or (b) certain nonverbal behaviours took place (i.e., frowning, repeatedly turning the text page, staring; merchie & van keer, 2014; vandevelde et al., 2015), thereby thoughtfully considering the student and situation at hand to avoid the loss of meaningful information about students’ behaviour (e.g., boekaerts & corno, 2005). as prompt, students were consistently given the instruction: “verbalize everything that you are doing or thinking” or “keep thinking aloud”. in this respect, type 1 (verbal content) and type 2 (nonverbal content) verbalizations were encouraged, and type 3 verbalizations were avoided since students were not asked to explain their cognition. consequently, researchers were able to identify spontaneous self-regulatory learning activities (ericsson & simon, 1980; vandevelde et al., 2015). 2.2.2 prior knowledge test as prior knowledge might influence text learning (alexander & jetton, 2000; bråten & samuelstuen, 2004), a prior knowledge test regarding the text topic was administered before the actual learning task. students were asked to write down everything they already knew about the topic. following prior research, the matching of students’ notes to the text content was opted for to score the prior knowledge test (for more information on this procedure, see merchie et al., 2014) the matching of students’ notes to the text content revealed very limited to no prior knowledge regarding the text content (m = 5.66, sd = 2.71; min = 0, max = 24). 2.2.3 learning task since studying in preparation for a classroom test is a regular task in secondary education, students were instructed to study an informative text in the way they would prepare for a test while thinking out loud (fox, 2009; samuelstuen & bråten, 2007). for the learning task, a 442-word informative text was used of which the participants did not study the topic (i.e., chewing gum) as part of their courses. the multi-paragraph text consisted of one title (i.e., chewing gum), four sections and subtitles (i.e., history, production, advantages, and disadvantages), and three pictures. text quality was verified in advance (see rogiers et al., 2019a). in view of encouraging students to plan their work, they were informed to have 50 minutes time for task completion. to enable students to monitor their progress, a clock was provided, but no further time indications were given to prevent the prompted monitoring of time. in line with previous studies (slotte et al., 2001), students were allowed, but not obligated to make notes on scratch paper while studying. during the task completion process, students were observed by the researcher and were only prompted to continue verbalizing when necessary (greene et al., 2011). 2.2.4 task-specific self-report inventory immediately after learning task execution, students completed the ‘text-learning strategies inventory’ (tlsi; merchie et al., 2014). this task-specific questionnaire consists of 37 items, subdivided into nine subscales (see appendix a) to which students respond on a five-point likert-scale (1 = completely disagree, 5 = completely agree). in line with theoretical frameworks (wade et al., 1990; zimmerman, 2002), the tlsi incorporates both cognitive (e.g., paraphrasing) and metacognitive (e.g., monitoring) text-learning strategies, as well as overt (e.g., summarizing) and covert (e.g., paraphrasing) strategies. good model fit results were obtained for this nine-factor model in prior large-scale research (rogiers et al., 2019a). appendix a presents the descriptive statistics and reliability coefficients for the tlsi-subscales. by means of hierarchical and k-means cluster analyses on the tlsi-subscale scores within the larger sample (n = 1,931), students learning strategy profile was determined (for a detailed description, see rogiers et al., 2019a). 2.3 think-aloud coding procedure of learning strategies in a first step, think-aloud sessions were transcribed and coded. as all sessions were audioand videotaped, both students’ verbal and non-verbal behaviour (e.g., highlighting text) was transcribed to increase coding accuracy (veenman, 2011; young, 2005). transcriptions were made by means of a computer program for subtitling videos (i.e., subtitle workshop 4). this program enables researchers to register the start and end time of each verbalisation and action. this is essential in view of conducting epm, as the sequence of strategies is calculated based on their exact time frame. subsequently, transcripts were segmented by one researcher into units of meaning, with one unit referring to a thematically consisted verbalization of a single text-learning activity (scott, 2008; van someren et al., 1994). repeated actions were analysed as separate activities in view of considering the recurrence of different text-learning strategies. as a result, 1,015 minutes of thinking aloud, and 4,107 units of meaning were identified and coded by means of the coding scheme based on prior research of merchie and van keer (2014). this coding scheme is an adapted version of the ‘text-learning strategy protocol’ (tlsp; see table 1), comprising 11 subcategories referring to different text-learning strategies. in line with the self-report questionnaire, the coding scheme reflects both cognitive and metacognitive, as well as overt and covert strategies. mean learning time was 20 minutes (sd = 3.89), with a minimum of 6 and a maximum of 34 minutes. analysis of variance showed no statistically significant differences between the four strategy profiles in terms of their mean learning time, f(3, 50) = 2.437, p = .076). finally, two trained coders independently double-coded 27% of the protocols, resulting in high interrater reliability (krippendorff’s α = .95; hayes & krippendorff, 2007). table 1 coding scheme for analysing students’ learning activities note. tlsp = text-learning strategy protocol. * in accordance to schoor and bannert (2012), this category was excluded from the process mining analysis, as we wanted to concentrate on task-related behaviour. 2.4 process mining analysis on the think-aloud data in a next step, the coded learning activities of each learner profile were analysed separately via process mining using disco (fluxicon, 2019). this software program enables researchers to study the course of students’ actual learning processes by generating process models for each learner profile. in these process models, both (1) the activities performed by the learners (i.e., the executed strategies during text learning), and (2) the paths or connections that occurred between these activities are displayed (fluxicon, 2019; van der aalst, 2011). thus, whereas the activities refer to the extent in which certain text-learning strategies are adopted (i.e., boxes in figures 2-5), the paths visualize the sequence of these performed activities (i.e., arrows in figures 2-5). above these paths, the frequency of each of these sequences is represented. further, both unidirectional paths (→), bidirectional paths (⇆), and loops (↺) are depicted in the process models, indicating that activities have respectively been conducted in consecution, in alternation, or that the same activity was performed several times in succession. in line with prior research in the field of srl (bannert et al., 2014; schoor & bannert, 2012) the fuzzy miner algorithm in disco was used to perform the analysis. this algorithm relies on two metrics (i.e., significance and correlation) to calculate which activities and paths should be included in the process models and which to be excluded (günther & van der aalst, 2007). significance refers to the relative importance of activities and paths, implying that more frequent text-learning activities are retained in the model. correlation is deployed for selecting only paths of closely connected activities (günther & van der aalst, 2007). by means of this algorithm, disco automatically includes strategies and paths that are often conducted by a large group of students in the process model, while less frequent activities and paths or paths that have been seldom conducted by only few students are excluded. to date, however, no specific standards are available on the amount of activities and paths that should be included in the process models. researchers argue that the ideal number of activities and paths strongly depends on the type of research data and questions involved (fluxicon, 2019). in general, the inclusion of as much activities and paths as possible while simultaneously avoiding too complex process models is recommended (fluxicon, 2019). in the current study, the percentages of included activities and paths in students’ process models were carefully deliberated among four experts on text learning and srl. in this respect, 33.33% of the most frequent strategies and the 33.33% most frequent connections between these strategies were included in the analysis. as a result, initially coded categories such as paraphrasing and elaborating (see table 1) were not included in the 33.33% process models (figures 2-5). although these strategies occurred commonly in the group of integrated strategy users and mental learners (table 2), they were adopted by a rather small share of learners compared to the occurrence of the other strategies. put differently, these activities did not belong to the 33.33% most frequent activities conducted at least once by a large group of learners. as a final step, following schoor and bannert (2012) and in view of obtaining split-half-reliability for the generated process models, we repeated the epm analyses for the five most typical individuals of each learning strategy profile. we perceived students as typical integrated strategy users (isu) when high frequencies were found for different text-learning strategies, whereas typical limited strategy users (lsu) were characterized by the dominant application of only one strategy (e.g., highlighting, rereading). for typical information organizers (io) and mental learners (ml), strategies with high frequency were respectively text-noting strategies (e.g., highlighting, summarizing for io) and mental learning strategies (e.g., rereading, rehearsing for ml; rogiers et al., 2019a). the obtained models for the five most typical individuals of each learning strategy profile were very similar to those in figures 2-5. 3. results 3.1 frequency of occurrence of text-learning strategies in different learning strategy profiles’ text-learning process (rq1) in view of the first research question, we examined which text-learning activities were executed by the different learning strategy profiles during actual text learning. table 2 displays the frequency of occurrence of all text-learning strategies included in the process models, as well as the number of students conducting each strategy at least once. one-way analysis of variance was used to test differences between the four learning strategy profiles regarding students’ use of different strategies. additionally, post hoc pairwise tests with bonferroni correction were conducted to investigate these differences in-depth. the analysis revealed significant differences between the four learning strategy profiles (see appendix b for detailed results of the post hoc pairwise comparisons and effect sizes). as can be derived from table 2, the results with regard to students’ cognitive strategy use show that integrated strategy users (isu) generally executed more diverse text-learning strategies than information organizers (io), mental learners (ml), and limited strategy users (lsu). the frequency of occurrence of most strategies was higher in integrated strategy users, as well as the percentage of students that adopted the strategies at least once. the most occurring cognitive strategies for integrated strategy users were summarizing, highlighting, paraphrasing, and elaborating, whereas for limited strategy users highlighting, memorizing, and elaborating were the most frequent strategies. for mental learners, summarizing and highlighting activities seldomly occurred, whereas rereading and rehearsing were frequently coded. in contrast to mental learners, summarizing and highlighting frequently occurred in the information organizers group, in addition to rereading and rehearsing. these results were reflected in the post hoc pairwise test results. as can be derived from appendix b, a statistically significant difference between the four learning strategy profiles was found for the verbalized overt text-learning strategies: summarizing (f(3, 4732) = 85.59, p < .001) particularly in favour of integrated strategy users and information organizers, and highlighting (f(3, 4732) = 43.88, p < .001) in favour of all learning strategy profiles, except for mental learners. regarding the covert text-learning strategies, the results show that limited strategy users more frequently applied memorizing (f(3, 4732) = 4.54, p = .004) than information organizers. further, rereading (f(3, 4732) = 22.89, p < .001) and rehearsing (f(3, 4732) = 144.34, p < .001) were most frequently executed by mental learners and information organizers and less frequent by limited strategy users, whereas paraphrasing (f(3, 4732) = 21.12, p < .001) and elaborating (f(3, 4732) = 5.252, p = .001) were less frequent performed by information organizers. no statistically significant difference between the four profiles was found regarding the execution of initial reading (f(3, 4732) = 2.61, p = 0.05). the results with respect to students’ metacognitive strategy use reveal no statistically significant differences between the four profiles regarding the use of comprehension monitoring activities (f(3, 4732) = 1.10 p = 0.348). in contrast, a statistically significant difference between learning strategy profiles was found with regard to planning (f(3, 4732) = 38.56, p < .001), indicating that planning activities were particularly executed by limited strategy users. in addition, a statistically significant difference with regard to progress monitoring (f(3, 4732) = 15.57, p < .001) reveals that this strategy frequently occurred in integrated and limited strategy users. however, in order to gain insight into the sequences in which these different strategies are adopted throughout students’ learning process, a closer look at the process models is needed (see rq2). table 2 frequency of occurrence of text-learning strategies for each learner profile (n = 51), including absolute frequency and number of students a note. a number of students conducting the strategy at least once. * strategies not included in the process models as they do not belong to the 33.33% most frequent activities that are conducted by a large group of students (see method section and rq2). isu = integrated strategy users, io = information organizers, ml = mental learners, lsu = limited strategy users. 3.2 temporal patterns in the different learner profiles’ text-learning process figures 2-5 display the resulting 33.33% process models for each learning strategy profile. the direction of the arrows represents the order in which the text-learning activities were adopted throughout students’ text-learning process. strategies that took place in the beginning of students’ text-learning process (e.g., planning, initial reading) are depicted at the top of the figures, while strategies that were executed at a later moment or at the end of the learning process (e.g., highlighting, summarizing, memorizing, rereading, rehearsing) are represented at respectively the centre or bottom of the figures. when contrasting the process models of the different learning strategy profiles and focussing on the cognitive and metacognitive strategies that were included in the models, clear differences can be noticed. first, the cognitive activities that were included in students’ process models indicate that integrated strategy users (isu), information organizers (io), and limited strategy users (lsu) applied a combination of both overt (e.g., summarizing and highlighting) and covert strategies (e.g., rehearsing, memorizing, rereading) during text learning, whereas mental learners (ml) exclusively adopted covert strategies. when considering the overt strategies, the process models show that integrated strategy users, information organizers, and limited strategy users frequently applied highlighting, whereas only the models of integrated strategy users and information organizers include summarizing strategies. regarding the covert strategies, the results show that rehearsing was only included in the process models of information organizers and mental learners. second, planning was included as a metacognitive strategy in all process models, while progress monitoring was only included in integrated strategy users’ process model, indicating that – compared to the frequency of the other strategies – a large share of integrated strategy users frequently tracked and controlled their progress throughout their learning process. the same applies for comprehension monitoring, which was only included in limited strategy users’ process model, implying that a large group of limited strategy users actively monitored (a lack of) understanding while processing the text. when studying the sequences in which these different strategies were adopted throughout students’ learning process, differences in the phases of the text-learning process can be identified. first, when focussing on the beginning of students’ learning process (i.e., start symbol in the process model), the results indicate that integrated and limited strategy users initiated their learning process with planning before initially reading the text. in contrast, information organizers and mental learners immediately started reading the text without planning in advance, which is indicated by the unidirectional arrows between initial reading and planning. subsequently, they performed planning after they initially read the text. concerning the strategies conducted during actual text studying, differences between the four learning strategy profiles were found as well. in the group of limited strategy users (figure 5), the unidirectional arrows between the different strategies indicate that planning, initial reading, highlighting, memorizing and rereading were consecutively executed. in addition, the unidirectional connection between initial reading and comprehension monitoring in limited strategy users’ process model denotes that 18% of these strategy users monitored their understanding after reading the text. this strongly differs from integrated strategy users’ process model (figure 2). while highlighting is also preceded here by initial reading, bidirectional paths are found between initial reading and highlighting, as well as between initial reading and summarizing. yet, the arrows connecting the different cognitive strategies, as well as the presence of reciprocal arrows, indicate that integrated strategy users alternately adopted these strategies before they started to memorize the text. further, the position of progress monitoring as rather isolated from the other activities in these strategy users’ process model must be noticed. this position is due to the fact that progress monitoring was applied before and after a wide variety of activities, suggesting that integrated strategy users tracked and controlled their progress throughout the entire learning process. however, since the process model only represents 33.33% of the performed activities, the wide variety of arrows were omitted by the program. when analysing the process model in detail, the large number of arrows between progress monitoring and a diverse set of text-learning strategies can be found. furthermore, it is notable that integrated strategy users (figure 2) alternated strategies before learning (i.e., planning) with activities during learning, which is indicated by the bidirectional arrows connected to students’ planning strategy. more particularly, they considered their planning not only before reading in the pre-learning phase, but throughout the different phases in their learning process (i.e., after reading, highlighting, and memorizing). when taking a closer look at information organizers’ process model (figure 3), many paths are visible, demonstrating that information organizers frequently switched between various strategies throughout their learning process, or frequently resumed previous strategies. this reveals that their text-learning process was rather cyclical organized. especially the strategies ‘summarizing’ and ‘highlighting’ played a prominent role in these learners’ learning process, as can be derived from the large number of incoming and outgoing arrows. for instance, after initially reading the text, the unidirectional arrows indicate that information organizers considered their planning before summarizing. after summarizing, a large share of these learners returned to reading the text or started memorizing or rehearsing the text. a clear bi-directional path is present between memorizing and highlighting activities, indicating that these activities were performed in alternation. further, highlighting was also frequently followed by rehearsing, initial reading, and/or summarizing. remarkable is that after conducting memorizing, summarizing, and highlighting activities, a large share of information organizers returned to initially reading the text. this could imply that these learners started to engage in different text-learning strategies without first reading or fully understanding the study text. although both information organizers’ and mental learners’ learning process was initiated by initial reading and planning, the further course of their learning process clearly differed. while a large share of information organizers started summarizing the text, a large share of mental learners (figure 4) started memorizing the text after planning. further, the recursive loop for memorizing, rehearsing and rereading demonstrates an alternated application of these strategies in mental learners, while the unidirectional arrows show that rehearsing was often followed by rereading and rereading by memorizing. more fine-grained differences in the course of students’ text-learning process can be detected when taking a more detailed look at the direction of the arrows in the process models for each learning strategy profile. the results show that limited strategy users followed a mainly linear structured learning process, as is indicated by the unidirectional arrows in their process model and the absence of any bidirectional paths. in contrast, the other three profiles returned more to prior activities or phases throughout their learning process, which is illustrated by the returning arrows pointing from the bottom to the top. hence, no strict linear, but rather a cyclical approach to learning was followed by these profiles. at last, the arrows leading to ‘stop’ in the process models (i.e., stop symbol in the figures) show which strategies were conducted lastly by the learners. as can be derived from the models, across all learning strategy profiles, most students finished their learning process with memorizing and rehearsing. however, initial reading also occurred as a final activity in some mental learners (20%) and limited strategy users (9%), while this is not the case in the learning process of the other learning strategy profiles. in addition, some integrated strategy users (13%) finished their learning process with reflecting on their progress, while rereading also occurred as final activity in some limited strategy users (9%). figure 2 text-learning process model of integrated strategy users (isu; n = 15), including the frequencies of occurrence and, between brackets, the case frequencies (i.e., the number of students that conducted the activities at least once). the more frequent an activity was performed, the darker it is displayed. the more frequent a path between activities occurred, the thicker the arrow is displayed. figure 3 text-learning process model of information organizers (io; n = 15), including the frequencies of occurrence and, between brackets, the case frequencies (i.e., the number of students that conducted the activities at least once). the more frequent an activity was performed, the darker it is displayed. the more frequent a path between activities occurred, the thicker the arrow is displayed. figure 4 text-learning process model of mental learners (ml; n = 10), including the frequencies of occurrence and, between brackets, the case frequencies (i.e., the number of students that conducted the activities at least once). the more frequent an activity was performed, the darker it is displayed. the more frequent a path between activities occurred, the thicker the arrow is displayed. figure 5 text-learning process model of limited strategy users (lsu; n = 11), including the frequencies of occurrence and, between brackets, the case frequencies (i.e., the number of students that conducted the activities at least once). the more frequent an activity was performed, the darker it is displayed. the more frequent a path between activities occurred, the thicker the arrow is displayed.   4. discussion to date, little is known about the sequences in which certain strategies are applied by different learners during the course of their text-learning process. in this respect, it remains unclear whether a cyclical and flexible approach to learning, as put forward as the most effective in various theoretical models, unfolds in different learning strategy profiles when learning from text. nevertheless, in-depth insight into students’ learning processes enables to be responsive to individuals’ learning needs and avoid ‘one-size-fits-all’ approaches to learning. the purpose of this study therefore was to uncover both the frequency of students’ applied strategies throughout their learning process, as well as the temporal patterns between these text-learning strategies. more particularly, the strategic behaviour of students from four different learning strategy profiles (i.e., integrated strategy users, information organizers, mental learners, and limited strategy users) based on a retrospective self-report questionnaire in a previous study (rogiers et al., 2019a), was further depicted and compared by means of educational process mining (epm) on their think-aloud protocol (tap) data. in this respect, both students’ concurrently and retrospectively measured strategy use was complementary taken into account. the first research question focused merely on the quantity of students’ strategy use by studying the frequency wherein text-learning strategies were executed by the different learning strategy profiles during actual text learning. the results clearly correspond to the findings of rogiers and colleagues (2019a) who determined different learning strategy profiles based on students’ retrospectively self-reported text-learning strategies. the results postulated less diverse learning strategy use for limited strategy users (lsu; e.g., highlighting and rereading), versus more varied overt and covert text-learning strategies for integrated strategy users (isu). similarly, the frequent use of overt text-noting strategies (i.e., highlighting, summarizing) reported by information organizers (io) was reflected in their verbalized learning behaviour. the same applies for mental learners (ml), who both reported and actually applied the frequent use of covert mental learning strategies (e.g., memorizing, rehearsing, paraphrasing). this was also reflected in the strategies included in the different process models (rq2). in this regard, the clusters determined based on students’ retrospective self-report data were largely confirmed by their concurrent tap data. although some research has clearly shown discrepancies between retrospective and concurrent measures of students’ strategic behaviour, our comparison overall shows that both measures enable us to uncover which strategies students do or do not use frequently. given this convergence, the current study provides empirical support for retrospective self-report questionnaires as acceptable alternatives for more timeand labour-intensive measures such as tap (e.g., greene & azevedo, 2009). it must be noticed, however, that retrospective self-reports offer a more general picture of the frequency of students’ strategy use, while tap enable a more fine-grained analysis of students’ learning process, for example by exploring the temporal patterns in which it unfolds. this was particularly elaborated on in response to the second research question by applying epm. with respect to the second research question (i.e., studying temporal patterns in students’ text-learning process based on the sequences in which their strategies are applied), the process models of the learning strategy profiles enabled a qualitative and systematic analysis based on several theoretical models in the field (see introduction section). when overviewing the results regarding the second research question, three major aspects should be noticed. first, the findings revealed that information organizers and mental learners immediately started their learning process with reading the text before considering their planning, whereas limited and integrated strategy users initiated their learning process with planning before they started to read. in addition, planning was strongly interwoven in the different phases of integrated strategy users’ learning process (i.e., before, during and after learning), which was clearly different from the other process models. the connections with planning in integrated strategy users’ process model could indicate that isu adopted a more efficient and systematic study approach (pintrich, 2000; zimmerman, 2002). pressley and colleagues (1987) for instance, consider good strategy users as planful strategy users who think before they act. their plan is not conceived as a linear sequencing of strategies, however, but rather as interacting and integrating with other strategies throughout the learning process. while at a more basic level, learners will develop a single (reading) plan for reading text materials (for the reading task) before learning, more advanced learners additionally develop a profound (action) plan for task execution and learning (desoete, 2007; pressley, 2000), which was the case for integrated strategy users in the current study. second, a remarkable difference between the learning strategy profiles regards the use of monitoring strategies. on the one hand, progress monitoring was included in integrated strategy users’ process model, suggesting that a considerable share of integrated strategy users actively tracked and controlled the quality of their progress and the available time left for task execution (meijer et al., 2006; moos & azevedo, 2009). by applying this progress monitoring strategy, adherence to the plan is stimulated, as well as revisions to comply with the plan (pressley et al., 1987). in this respect, progress monitoring during learning is strongly linked to planning before and during learning. hence, particularly in integrated strategy users’ process model, metacognitive strategies (i.e., planning and progress monitoring) and, by extension, cognitive strategies, mutually interacted. on the other hand, comprehension monitoring was included in the process model of limited strategy users. this strategy refers to control activities directed at the correctness and comprehensiveness of one’s understanding (moos & azevedo, 2009). an indicator of applying this strategy, for example, concerns learners’ noting lack of full understanding, as well as efforts to monitor their understanding after reading the text (veenman et al., 1997). as previous research shows that limited strategy users’ level of reading ability is generally lower compared to the other learning strategy profiles, limited strategy users’ level of reading ability may also have played a role here (rogiers et al., 2019a). finally, the results showed that limited strategy users followed a rather linear sequenced approach, whereas integrated strategy users, mental learners, and particularly information organizers adopted a more cyclical approach to learning as they often repeated or returned to prior activities. compared to the other learning strategy profiles, limited strategy users did not seem to interact with the text as actively and recursively. instead, they confined their study behaviour to highlighting, memorizing, and rereading. contrary, particularly information organizers and integrated strategy users frequently switched between various strategies throughout their learning process or resumed previous strategies. according to important theoretical models concerning successful strategy use (alexander & jetton, 2000; pintrich, 2000; pressley et al., 1987; weinstein et al., 2011; zimmerman, 2002), also the simultaneous use of different strategies is what characterises a good strategy user. as different strategies are executed ever more efficiently in good strategy users, pressley and colleagues (1987) state that in these learners, more short-term capacity is ‘left over’ to adopt other strategies simultaneously and enhance their text learning (pressley et al., 1987). in information organizers’ process model, for instance, a large share of information organizers returned to reading the text or started memorizing or rehearsing the text after summarizing. further, memorizing and highlighting were often performed in alternation and highlighting was also frequently followed by rehearsing, initial reading and/or summarizing. remarkable, however, is that after conducting memorizing, summarizing, and highlighting activities, a large share of information organizers returned to initially reading the text. these paths could imply that information organizers started to engage in different text-learning strategies without first reading or fully understanding the study text. equally, this might indicate that information organizers had the tendency to interrupt their first reading with other activities (wade et al., 1990). this could be due to the fact that they did not initiate their actual learning process by planning this process in advance. in this respect, their text-learning process seems less systematic than, for example, integrated strategy users’ learning process. rather than directly selecting important ideas in the text or starting to summarize, the findings indicate that integrated strategy users read text fragments, deliberate on the importance of the given information and then highlighted or summarized the main ideas. subsequently, integrated strategy users applied their notes as tools to memorize. this might again indicate that these learners adopted a more strategic approach to text learning. however, it is remarkable that integrated strategy users rarely applied rereading or rehearsing strategies during their text-learning process. since they more actively monitored their progress, it might have been the case that they did not consider it necessary or feasible to repeat or rehears the text within the given time span. further, also the recursive loop for memorizing, rehearsing, and rereading in mental learners’ process model demonstrates an alternate application of these strategies. although we must be aware of our small sample size, the initial reading activities as final activities in the process models of mental learners and limited strategy users could imply that some students finished their learning process quite abruptly and did not implement a thoughtful text-learning approach. 4.1 limitations and implications the present study is associated with some strengths and concerns regarding both the measure and data analysis approach used. first, we must be aware of the fact that various self-report measures reflect learning strategy conceptualizations in a different way. retrospective self-report data has shown to be valuable in prior research to provide insight into the frequency and variety of applied overt, covert, cognitive, and metacognitive learning strategies during a learning task (e.g., rogiers et al., 2019a; merchie et al., 2014). however, this particular data provides us with less information on the cyclical and adaptive nature of these processes, characteristics that have been identified in various theoretical models as being essential in strategic learning. tap can be regarded as concurrent self-report measures and are recognised as useful data sources to provide additional insight (dinsmore, 2018; veenman, 2011). more particularly, by the unique combination of concurrent self-report think-aloud data and educational process mining in this study, we were able to shed light on not only the diversity of applied learning strategies, both also on their cyclical and adaptive nature. in this way, epm on students’ concurrent tap really opened the black box and provided in-depth insight into the course of different learners’ actual text-learning process (cromley & wills, 2016; veenman, 2005). in this respect, this study illustrates the complementarity of both retrospective (i.e., task-specific self-report questionnaires) and concurrent self-reports (i.e., tap). this reflection touches upon the first explicit question regarding self-report data tackled throughout the different contributions in this special issue. however, limitations of this study should equally be recognised. a first risk inherent to thinking out loud concerns the incompleteness as automated or unconscious behaviour is not explicitly verbalized (boekaerts & corno, 2005). it is possible that students’ actions and thoughts might have sometimes remained covert, making them difficult to record in the tap. second, students were instructed to report both verbal and nonverbal processes during thinking out loud. to prevent that students’ verbalisations did interfere with their learning process (greene et al., 2011), they were not asked to explain these processes. therefore, tap gave no insight into students’ underlying motives for their executed activities and sequences that occurred in the process models. including retrospective stimulated interviews (schellings & broekkamp, 2011) based on students’ tap could be useful in future research to learn more about the underlying motives of students’ behaviour. as to the data analysis approach, no specific epm guidelines are currently available with respect to the number of activities and paths to be included in the process models. more process mining research in educational settings, as well as exploring epm techniques that rely on different algorithms could therefore contribute to a better understanding of students’ learning process on the one hand and to more evidence-based guidelines for conducting epm on the other hand (bannert et al., 2014). related to this, it is to be recommended as well to engage in more fine-grained coding of the think aloud data in future research in view of considering valences of specific srl processes during students’ learning. positive judgment of learning (e.g., “i am getting this”), for example, can elicit a distinct subsequent srl process than a negative judgment of learning (e.g., “i am so confused with this paragraph”). within the scope of the current study, however, this fine-grained coding was not applied (e.g., both ‘detecting lack of comprehension’ and ‘mentioning awareness of understanding’ were more generally coded as ‘monitoring comprehension’). we therefore make a plea for more fine-grained coding of the distinct subprocesses of particular learning strategies, such as for instance ‘comprehension monitoring’, to enable the study of more detailed subprocesses and their temporal nature. further, prior studies have shown that students adapt their strategies and the effort they spend on studying according to the learning task, their prior domain knowledge, and their learning goals (e.g., boekaerts & niemivirta, 2000; broekkamp & van hout-wolters, 2007). as a result, students may decide to select from their available strategies these strategies that are most appropriate given the assigned learning task, their prior knowledge, and/or the learning goal they set for themselves. in this respect, they might opt, for example, to systematically reread the study text instead of engaging in summarizing and paraphrasing the text (broekkamp & van hout-wolters, 2007). it will therefore be interesting in future research to study students’ strategy use across more and varied learning tasks as well as to investigate the impact of their prior knowledge and their personal learning goals. in this respect, it is to be recommended to also consider other types of prior knowledge tests, such as open questions, multiple choice tests, cloze tests, completion tests, and recognition tests, which also provide valid means of assessment (dochy et al., 1999). with regard to the implications for research, this study extends earlier work by including new possibilities for analysing learning processes by means of epm. this study must be considered as a first important investigation in unravelling patterns in secondary school students’ text-learning processes. educational research is encouraged to fine-tune this type of analysis by the suggestions mentioned above. further, our results encourage data triangulation in future research, preferably combining both off-line (e.g., self-report questionnaires) and on-line measures (e.g., tap) to gain a more accurate portrayal of students’ learning process (rogiers et al., 2019b; veenman, 2005). in view of implications for theory, this study showed that epm can be used to test the cyclical and adaptive use of learning strategies as put forward in different theoretical models. the present study confirmed differences between four learning strategy profiles in secondary school students. these findings carry important implications for educational practice to help and support students to also evolve towards the more adaptive and cyclical use of strategies. the proposed process models provide a detailed picture on students’ text-learning process and could be used as starting points for supporting learner-oriented teaching and learning. keypoints integrated strategy users execute more diverse text-learning strategies than the other learner profiles limited and integrated strategy users initiate their learning process with planning this process, while planning is strongly interwoven in the different phases of integrated strategy users’ learning process progress monitoring is included in integrated strategy users’ process model, comprehension monitoring is included in limited strategy users’ process model limited strategy users follow a rather linear sequenced approach to learning whereas integrated strategy users, mental learners, and particularly information organizers adopt a more cyclical approach to learning references alexander, p. a. (1998). the nature of disciplinary and domain learning: the knowledge, interest, and strategic dimensions of learning from subject matter text. in c. r. hynd (ed.), learning from texts across conceptual domains (pp. 263-287). new york, ny: routledge. alexander, p. a., & jetton, t. l. (2000). learning from text: a multidimensional and developmental perspective. in m. l. kamil, p. b. mosenthal, p. d. pearson, & r. barr (eds.), handbook of reading research (vol. iii, pp. 285-310). mahwah, nj: lawrence erlbaum associates, inc. alexander, p. a., grossnickle, e. m., dumas, d., & hattan, c. (2018). a retrospective and prospective examination of cognitive strategies and academic development: where have we come in twenty-five years? in a. o’donnell (ed.), oxford handbook of educational psychology. oxford, uk: oxford university press. bannert, m., reimann, p., & sonnenberg, c. (2014). process mining techniques for analysing patterns and strategies in students’ self-regulated learning. metacognition and learning, 9(2), 161–185. https://doi.org/10.1007/s11409-013-9107-6 bergman, l. (2001). a person approach to adolescence: some methodological challenges. journal of adolescent research, 16(1), 28–53. https://doi.org/10.1177/0743558401161004 boekaerts, m., & corno, l. (2005). self regulation in the classroom: a perspective on assessment and intervention. applied psychology, 54 (2), 199–231. https://doi.org/10.1111/j.1464-0597.2005.00205.x boekaerts, m., & niemivirta, m. (2000). self-regulated learning: finding a balance between learning goals and ego-protective goals. in m. boekaerts, p. r. pintrich & m. zeidner (eds.), handbook of self-regulation. san diego, ca: academic press. bråten, i., & samuelstuen, m. s. (2004). does the influence of reading purpose on reports of strategic text processing depend on students’ topic knowledge? journal of educational psychology, 96(2), 324-336. https://doi.org/10.1037/0022-0663.96.2.324 båten, i., & samuelstuen, m. s. (2007). measuring strategic processing: comparing task-specific self-reports to traces. metacognition and learning, 2(1), 1-20. https://doi.org/10.1007/s11409-007-9004-y cohen, j. (1977). statistical power analysis for the behavioral sciences. new york, ny: academic press. cromley, j. g., & wills, t. w. (2016). flexible strategy use by students who learn much versus little from text: transitions within think-aloud protocols. journal of research in reading, 39(1), 50-71. https://doi.org/10.1111/1467-9817.12026 deed, c., lesko, t. m., & lovejoy, v. (2014). teacher adaptation to personalized learning spaces. teacher development, 18(3), 369–383. doi:10.1080/13664530.2014.919345 dinsmore, d. l. (2018). strategic processing in education. new york, ny: routledge. dochy, f., segers, m., & buehl, m. m. (1999). the relation between assessment practices and outcomes of studies: the case of research on prior knowledge. review of educational research, 69(2), 145-186. https://doi.org/10.3102/00346543069002145 ericsson, k. a., & simon, h. a. (1980). verbal reports as data. psychological review, 87(3), 215251. https://doi.org/10.1037//0033-295x.87.3.215 fluxicon. (2019). disco user’s guide. retrieved from https://fluxicon.com/disco/files/disco-userguide.pdf fox, e. (2009). the role of reader characteristics in processing and learning from informational text. review of educational research, 79(1), 197-261. https://doi.org/10.3102/0034654308324654 fryer, l. k., & vermunt, j. d. (2017). regulating approaches to learning: testing learning strategy convergences across a year at university. british journal of educational psychology, 88 (1), 21-41. https://doi.org/10.1111/bjep.12169 günther, c., & van der aalst, w. (2007). fuzzy mining: adaptive process simplification based on multi-perspective metrics. in g. alonso, p. dadam, & m. rosemann (eds.), international conference on business process management (bpm 2007) (pp. 328–343). berlin, germany: springer. greene, j. a., robertson, j., & croker costa, l.-j. (2011). assessing self-regulated learning using think-aloud methods. in b. j. zimmerman & d. h. schunk (eds.), handbook of self-regulation of learning and performance (pp. 313–328). new york, ny: routledge. mcnamara, d. s., ozuru, y., best, r., & o'reilly, t. (2007). the 4-pronged comprehension strategy framework. in d. s. mcnamara (ed.), reading comprehension strategies: theories, interventions, and technologies (pp. 465-496): new york, ny: erlbaum merchie, e., & van keer, h. (2014). learning from text in late elementary education. comparing think-aloud protocols with self-reports. procedia social and behavioral sciences, 112(2013), 489–496. https://doi.org/10.1016/j.sbspro.2014.01.1193 merchie, e., van keer, h., & vandevelde, s. (2014). development of the text-learning strategies inventory: assessing and profiling learning from texts in fifth and sixth grade. journal of psychoeducational assessment, 32(6), 1-15. https://doi.org/10.1177/0734282914525155 organisation for economic cooperation, development [oecd]. (2006). personalizing education. paris: oecd publishing. panadero, e. (2017). a review of self-regulated learning: six models and four directions for research. frontiers in psychology, 8(422), 1–28. https://doi.org/10.3389/fpsyg.2017.00422 pintrich, p. r. (2004). a conceptual framework for assessing motivation and self-regulated learning in college students. educational psychology review, 16(4), 385–407. https://doi.org/10.1007/s10648-004-0006-x pressley, m., borkowski, j. g., & schneider, w. (1987). cognitive strategies: good strategy users coordinate metacognition and knowledge. annals of child development, 4, 89-129. reimann, p. (2007). time is precious: why process analysis is essential for cscl (and can also help to bridge between experimental and descriptive methods). in c. chinn, g. erkens & s. puntambekar (eds.), mice, minds, and society. proceedings of the computer-supported collaborative learning conference (cscl 2007) (pp. 598–607). new brunswick, nj: international society of the learning sciences. rogiers, a., merchie, e., & van keer, h. (2018). fostering text-learning strategies in secondary education through explicit strategy-instruction. paper presented at the international conference of the earli sig 2, freiburg, germany, 27-29 august, 2018. rogiers, a., merchie, e., & van keer, h. (2019a). learner profiles in secondary education: occurrence and relationship with performance and student characteristics. the journal of educational research, 112 (3), 385-396. https://doi.org/10.1080/00220671.2018.1538093 rogiers, a., merchie, e., & van keer, h. (2019b). what they say is what they do? comparing task-specific self-reports, think-aloud protocols, and study traces for measuring secondary school students’ text-learning strategies. european journal of psychology of education, 1-18 . https://doi.org/10.1007/s10212-019-00429-5 samuelstuen, m. s., & bråten, i. (2007). examining the validity of self-reports on scales measuring students' strategic processing. british journal of educational psychology, 77(2), 351-378. https://doi.org/10.1348/000709906x106147 schellings, g. l. m., & broekkamp, h. (2011). signaling task awareness in think-aloud protocols from students selecting relevant information from text. metacognition and learning, 6, 65–82. https://doi.org/10.1007/s11409-010-9067-z slotte, v., lonka, k., & lindblom-ylanne, s. (2001). study-strategy use in learning from text. does gender make any difference? instructional science, 29(3), 255-272. https://doi.org/10.1023/a:1017574300304 schoor, c., & bannert, m. (2012). exploring regulatory processes during a computer-supported collaborative learning task using process mining. computers in human behavior, 28(4), 1321–1331. https://doi.org/10.1016/j.chb.2012.02.016 scott, d. b. (2008). assessing text processing: a comparison of four methods. journal of literacy research, 40, 290-316. doi:10.1080/10862960802502162 simpson, m. l., & nist, s. l. (2000). an update on strategic learning: it's more than textbook reading strategies. journal of adolescent & adult literacy, 43(6), 528-541. stoeger, h., fleischmann, s., obergriesser, s. (2015). self-regulated learning (srl) and the gifted learner in primary school: the theoretical basis and empirical findings on a research program dedicated to ensuring that all students learn to regulate their own learning. asia pacific education review, 16, 257-267. https://doi.org/ 10.1007/s12564-015-9376-7 veenman, m. v. j. (2005). the assessment of metacognitive skills: what can be learned from multi-method designs? in c. artett & b. moschner (eds.), ledrnstrategien und metakognition. implikationen für forschung und praxis (pp. 77-99). münster: waxmann. veenman, m. v. j. (2011). alternative assessment of strategy use with self-report instruments: a discussion. metacognition and learning, 6(2), 205–211. https://doi.org/10.1007/s11409-011-9080-x veenman, m. v., elshout, j. j., & meijer, j. (1997). the generality vs. domain-specificity of metacognitive skills in novice learning across domains. learning and instruction, 7(2), 187–209. https://doi.org/10.1016/s0959-4752(96)00025-4 wade, s. e., trathen, w., & schraw, g. (1990). an analysis of spontaneous study strategies. reading research quarterly, 25(2), 147-166. https://doi.org/10.2307/747599 weinstein, c. e., jung, j., & acee, t.w. (2011). learning strategies in v. g. aukrust (ed.), learning and cognition in education (pp. 137-143). oxford, uk: elsevier limited. van der aalst. (2011). process mining: discovery, conformance and enhancement of business processes. new york, ny: springer. van someren, m. w., barnard, y. f., & sandberg, j. a. c. (1994). the think aloud method. a practical guide to modelling cognitive processes. london, uk: academic press. zimmerman, b. j. (2002). becoming a self-regulated learner: an overview. theory into practice, 41(2), 64–71. young, k. a. (2005). direct from the source: the value of ‘think-aloud’ data in understanding learning. journal of educational enquiry, 6 (1), 19-33. appendices appendix a descriptive statistics and reliability coefficients of the different text-learning strategies inventory subscales note. tlsi = text-learning strategies inventory. cronbach’s α is based on the total sample of 1,931 students wherein learner profiles were determined. appendix b results of the post hoc pairwise comparisons between the four learning profiles on the different coding categories note. *p ≤ .05, ** p ≤ .01, ***p ≤ .001. to interpret effect sizes, cohen’s benchmarks for the social sciences apply (i.e., small effect size: d = 0.2, medium effect size: d = 0.5, and large effect size: d = 0.8; cohen, 1977). isu = integrated strategy users, io = information organizers, ml = mental learners, lsu = limited strategy users. frontline learning research vol.7 no. 3 (2019) 91 118 issn 2295-3159 measuring academic learning and exam self-efficacy at admission to university and its relation to first-year attrition: an irt-based multi-program validity study tine nielsen a, ida sophie friderichsen a bjarke tarpgaard hartkopf b a university of copenhagen, denmark b the danish evaluation institute, denmark article received 4 june / revised 13 august / accepted 30 august / available online 2 october abstract self-efficacy is associated with both academic performance and attrition in higher education. whether it is possible to measure students’ academic self-efficacy after admission and prior to commencing higher education (i.e. pre-academic self-efficacy) in a valid and reliable way has hardly been studied. aims: 1) to evaluate the construct validity and psychometric properties of two short scales to measure pre-academic learning self-efficacy (pal-se) and pre-academic exam self-efficacy (pae-se) using rasch measurement models, 2) to investigate whether pre-academic self-efficacy was associated with half-year attrition across degree programs and institutions. data consisted of 2686 danish students admitted to nine different university degree programs across two institutions. item analyses showed both scales to be essentially objective and construct valid, however, all items from the pae-se and two from the pal-se were locally dependent. differential item functioning was found for the pal-se relative to degree programs. reliability of the pae-se was .77, and varied for the pal-se from .79 to .86 across degree programs. targeting was good only for the pal-se, thus we proceeded with the pal-se. pal-se was found to be associated with half-year attrition: a difference in pal-se from minimum to maximum was associated with a difference in half-year attrition of approximately 7%. this association was found both in the bivariate model and in the multivariate models with control of degree program, and with control of degree program and individual covariates such as earlier educational achievement and social background variables. results thus also indicate that pal-se has a causal effect on half-year attrition. keywords: pre-academic self-efficacy; construct validity; differential item functioning; graphical loglinear rasch model; attrition corresponding author: tine.nielsen@psy.ku.dk doi: 10.14786/flr.v7i3.503 1. introduction self-efficacy, at the general level, refers to the belief in one’s own capability to plan and perform necessary actions to attain a certain outcome (bandura, 1997). students who feel efficacious when learning or performing a task, participate more readily, work harder, persist longer when they encounter difficulties, and achieve at a higher level of academic performance (schunk & pajares, 2002). self-efficacy is a central and much used construct in many fields, not least in education and mental health research and the combination of the two: the positive correlational relationship between self-efficacy and academic performance is well-established (ferla et al., 2009; luszczynska et al., 2005; zimmerman et al., 1992). low self-efficacy has been found to be associated with increased risk of depression and anxiety (luszczynska, et al., 2005; muris, 2002; tahmassian & jalali moghadam, 2011), and self-efficacy has been found to mitigate the negative effect of stress on life satisfaction (burger & samuel, 2017). self-efficacy, self-esteem and psychological distress were found to be the most important predictors of stress in university students (saleh et al, 2017). recent meta-analyses concluded that self-efficacy is the strongest (non-ability) predictor of grade point average (gpa) in higher education (he) above personality traits, motivation and various learning strategies (bartimote-aufflick et al., 2015; richardson et al., 2012). an ongoing and unresolved discussion in self-efficacy research is whether self-efficacy is a general or a specific construct, or something in between (bandura, 1997; pajares, 1997; scherbaum, cohen-charash, & kern, 2006). thus, there is a plethora of self-efficacy instruments measuring from the altogether general self-efficacy, over the domain-specific, the context-specific, to courseor even task-specific self-efficacy. while the general self-efficacy scale (gses; schwartzer & jerusalem, 1995), or adaptations of it, is probably the most commonly used in health related research, in educational research it is more common to construe self-efficacy as at least domain-specific; e.g. the general academic self-efficacy scale (gase; nielsen et al., 2018) and the physics self-efficacy questionnaire (pseq: lindstrã¸m & sharma, 2011), both for the academic domain and both based on the gse. in between the general and very specific instruments, are a range of sub-domain-, contextand course-specific self-efficacy instruments, e.g. the computer self-efficacy scale (scherer & siddiq, 2015), the psychologist and counsellor self-efficacy scale (watt et al., 2019), the diabetes management self-efficacy scale (williams et al., 2014), and the commonly used specific academic self-efficacy scale from the motivated strategies of learning inventory (mslq; pintrich et al, 1991). given the importance of self-efficacy for educational outcomes, a large body of research has focused on investigating differences in the levels of self-reported self-efficacy for student subgroups, e.g. gender, age, academic discipline, institution and so on. however, as pointed out by nielsen and colleagues (2018), hardly any studies comparing gender groups of students have determined whether the self-efficacy scales they use in fact are measurement invariant and thus whether self-efficacy scores for these groups can be compared without bias. the formal definition of measurement invariance (mi), according to mellenbergh (1989) and meredith (1993), states that mi holds when the probability of observed scale scores depends only on the estimated latent variable for the measure and not any other characteristic of the persons being measured (i.e. relevant demographic and background variables such as gender, age, etc.). in terms of self-efficacy this would mean that the observed raw scores for any students should only be determined by their latent level of self-efficacy and not also on whether they are, for example, male or female, or whether they are entering one higher educational degree program or another. within item response theory, including the subgroup of rasch models, mi is more often formulated as a requirement of absence of differential item functioning (dif), which is defined in exactly the same way as at the item level. thus, no dif is formally defined as that the probability of an item responses should depend only on the estimated latent variable, and not additionally on any background variables (or exogenous variables) (kreiner, 2013). we can then say that items are conditionally independent of the background variable given the estimated latent variable (or person parameter), and thus items function in the same way for all respondents belonging to any of the subgroups defined by the exogenous variable, and the scores will not be biased for some groups of students. avoiding bias caused by dif (i.e. having mi) is important, as some students would otherwise obtain unfair scores compared to other students (i.e. artificially low or high scores obtained due to subgroup membership and not because they have higher or lower self-efficacy). for research purposes this is equally important, because dif, depending on its magnitude, can cause statistical comparisons of student self-efficacy to be confounded by their subgroup status (kreiner, 2007). in the literature search for the current study, we found that none of these looked into the issue of invariant measurement across the subgroups that were compared. a recent meta-analysis suggested student self-efficacy to be heterogeneously associated with gender as the relationship is moderated by academic discipline (huang, 2013). young, wendel, esson and plank (2018) found that self-efficacy levels were higher for men than for women, but that levels of student self-efficacy did not differ according to the included academic disciplines (science, technology, engineering, and mathematics). whether levels of self-efficacy vary across students from different he institutions has been left largely un-investigated. to the authors’ knowledge, only one study has looked into this: ginsborg, kreutz, thomas and williamon (2009) found that the level of general self-efficacy in music performance students from he music conservatories was significantly lower than in university health students from nursing and biomedical science. the predictive power of pre-higher education (pre-he) self-efficacy for academic outcome has hardly been studied internationally: yong (2010) found no discipline-specific differences in general self-efficacy level in a study of 105 pre-university engineering and business students admitted to a private malaysian university. van herpen et al. (2017) found that pre-university academic self-efficacy did not predict academic success in the first year in general (sample sizes were too small to look into single disciplines), when assessed at the time of application to university. aryee (2017) found that high levels of pre-he mathematics self-efficacy were associated with student persistence and the completion of college, while low levels of pre-he mathematics self-efficacy were associated with a diminished likelihood of attaining any degree or credential. the relationship of self-efficacy (measured after the commencement to he) and student persistence and attrition is well studied within the educational research field. a meta-analysis concluded that student self-efficacy beliefs were associated with a variety of academic persistence outcomes (“number of academic terms completedâ€�, “number of items or tasks completedâ€�, and “amount of time spent on the performanceâ€�) (multon, brown & lent, 1991). furthermore, academic self-efficacy in he students has been found to be associated with student attrition: de clercq, galand and frenay (2017) found that low levels of academic self-efficacy in students during their first year at university were associated with different combinations of achievement predictors including an increased risk of attrition. devonport and lane (2006) also found an association between low levels of self-efficacy (on a self-developed academically related instrument) in first year university students, and an increased risk of withdrawal. the relationship between academic self-efficacy and student attrition has also been studied through students’ self-perceived intention to persist in he and several studies have found a positive association between high levels of academic self-efficacy and an increased completion intention (elliot, 2016; lent et al., 2016; thomas, 2014; vuong, brown-welty & tracz, 2010; you, 2018). however, none of the studies investigating self-efficacy either pre-he or after commencement to he looked into the issue of measurement invariance across subgroups such as academic discipline, degree program or institution. in addition, it remains unanswered whether pre-academic self-efficacy can indeed predict outcomes, and specifically attrition. 1.1 the present study the present study seeks to fill the aforementioned gaps in the existing research on academic self-efficacy: 1) to investigate the relationship between pre-academic self-efficacy and first semester attrition in university. in order to do this, we thus needed 2) to obtain a construct valid and psychometrically sound instrument for the measurement of pre-academic self-efficacy, suitable for use before students commence the university degree programs to which they have been admitted; and within the psychometric analysis 3) to investigate whether the instrument is measurement invariant across institutions, degree programs and other relevant subgroup divisions of students. none of these elements has been researched earlier. as we could not identify an appropriate instrument, we, as an initial step, adapted an existing self-efficacy instrument to be an instrument of pre-academic learning and exam self-efficacy. the aims of the present study were then two-fold: firstly, to conduct a thorough investigation of the construct validity and psychometric properties of the instrument, using cutting edge item response theory models; i.e. graphical loglinear rasch models. these are models where both differential items functioning and a lack of local independence between items can be adjusted for, as long as these are uniform, in order to obtain unbiased scores for subgroups, while retaining sufficiency of the summed score. in the psychometric analyses, we included investigations of measurement invariance across both institutions and degree programs and other subgroupings of students, in order to prevent subsequent confounding of the relationship between pre-academic self-efficacy and attrition by institution or degree program. secondly, to employ the instrument in an analysis of the association between the two measures of pre-academic self-efficacy and attrition within the first half-year. more specifically, we investigated the following research questions: rq1: is it possible to obtain valid, reliable and well-targeted measures of pre-academic learning self-efficacy and/or pre-academic exam self-efficacy of undergraduate higher education students after admission and prior to commencing education, across multiple degree programs and institutions? in these analyses, a major focus was measurement invariance in relation to degree programs and institutions, as well as across subgroups of students defined by gender, age, and priority of program at the time of application. for student groups where psychometrically sound and invariant measures of self-efficacy could be established, we also investigated the research question: rq2: how is pre-academic self-efficacy at admission associated with half-year attrition? how is this association between pre-academic self-efficacy and attrition changed when variation between educational programs and individual covariates are taken into account? and how can these results be interpreted in terms of causality? 2. methods 2.1 data the data for the current study was collected by the national evaluation institute (eva) within the framework of the ongoing national longitudinal panel on student attrition encompassing several year cohorts of undergraduate students admitted to higher education (he). this attrition panel is conducted in the population of students admitted to he degree programs through the coordinated admission (kot) in denmark. among students who were admitted to a higher educational program in 2017 a cohort of students participate in a total of four waves of data collection time-wise spanning from admission notice to 13 months after commencing the degree program. student participation was solicited through the national digital mail system [e-boks] in the first wave, where consent was also given. data was collected during two weeks, starting 10 days after students had received their admittance letter and ending before commencement of the he programs in denmark. the first wave for the 2017-cohort included the two scales intended to measure academic learning and exam self-efficacy in relation to the students’ future first semester courses, which are utilized in the current study. sample selection criteria: 1. for validity studies of new measurement scales or adapted versions of existing scales, it is desirable to employ a controlled and maximally comparable data sample across subgroups, in order to be able to identify properly the source of any measurement issues arising. the current sample was thus selected from the first wave of data from the 2017 cohort (c.f. the above) as student cases with complete self-efficacy data which fulfilled three a-priori criteria: 2. students who were admitted to the two largest universities in denmark; the university of copenhagen (ucph) and the university of aarhus (uaa) for the summer start of programs, 3. academic programs existing at both universities, 4. academic programs who admitted at least 200 students in one of the included universities. however, as the initial selection of programs did not include any from the faculty of arts, we decided to include also the largest language and largest “other humanitiesâ€� programs as well. 2.1.2 the current data sample the resulting data sample consisted of complete self-efficacy and background information from, on average, 54% of the students admitted to nine large academic programs (table 1). table 1. distribution of the study sample on institution and degree programs notes. a of the 2777 students from the included programs who responded to the survey, 91 had completely missing self-efficacy data. of these, 56 were from the university of copenhagen and 35 from the university of aarhus, which was not a significant difference across universities (fischer’s exact ï�£2 1,219, df 1, p =0.285). looking further into the pattern of missing data in each of the 18 programs, we found that between 0 and 11 students had missing self-efficacy data, with the highest numbers being in the largest programs. thus, the missing pattern appears to be random. the 91 cases were excluded from the analysis, which was thus done on 2686 students. the university of copenhagen (ucph) sample consisted of 1496 students. their mean age was 21.7 years (sd 4.8), 61.3 % (n = 917) were women, and 92.5 % (n = 1384) were admitted to the academic program that was their first priority when they applied. the university of aarhus (uaa) sample consisted of 1190 students. their mean age was 21.2 years (sd 3.2), 59.5 % (n = 708) were women, and 92.1 % (n = 1096) were admitted to the academic program that was their first priority. 2.2 instrument & administration the self-efficacy scales employed in the current study was adapted from the specific academic learning self-efficacy scale (sal-se) and specific academic exam self-efficacy scale (sae-se) (nielsen et al., 2017). the sal-se and the sae-se scales are subscales of the self-efficacy scale in the motivated strategies for learning questionnaire (mslq; pintrich, smith, garcia, & mckeachie, 1991), translated into danish and with an adapted response scale. the mslq is intended to measure students’ motivational orientation and learning strategies in high school and he (pintrich et al., 1991), and has been widely used and translated into several other languages (credã© & philips, 2011; duncan & mckeachie, 2005). nielsen and colleagues (2017), with samples of psychology students, found that the sal-se and the sae-se scales each fit a rasch model and as such the sum scores of each scale are sufficient statistics for the latent rasch score. it was clearly rejected that the two subscales formed a single unidimensional scale. both the sal-se and the sae-se subscales had sufficient reliability for the investigated samples to be used both in surveys at the group level and to assess the se levels of individual students as intended with the mslq (sal-se r = .87, sae-se r = .89). targeting of the sal-se and the sae-se subscales to the sample of psychology students were not optimal, in the sense that students were located more towards the higher end of the scales while most information could be obtained more towards the lower end of the scales. the targeting of the sal-se subscale was the better of the two. the sal-se and the sae-se scales (nielsen et al., 2017) are course specific measures of self-efficacy. however, it was not possible to refer to specific courses in the current student population, as the students had not yet commenced the degree program that they had been admitted to. accordingly, all items had been adapted so that instead of referring to “this class/courseâ€�, they referred to “the first courses in my programâ€�, while the response categories were not changed (table a in the appendix shows the danish as well as an english version of the items). the adaptation made the subscales somewhat more general measures of learning and exam self-efficacy than the sal-se and sae-se scales, as well as measures of what might be termed pre-academic self-efficacy as the students had not yet entered academia. as a consequence, we have renamed the two scales to the more appropriate pre-academic learning self-efficacy scale (pal-se) and pre-academic exam self-efficacy scale (pae-se). the item texts of the pal-se and the pae-se are provided in both danish and english in the appendix (table a1). all pal-se and pae-se items were administered in the same relative order as in the original mslq (pintrich et al., 1991), and they were mixed in with items from other scales. 2.3 rasch analysis and graphical loglinear rasch models the family of rasch models, which is a subgroup of item response theory models (van der linden & hambleton, 1997) will be used for the item analysis. fit to a rasch model (rasch, 1960) provides optimal measurement properties of the scale in question (kreiner, 2007, 2013; mesbah & kreiner, 2013) including: local independence of items (i.e. responses to a self-efficacy item depends only on the level of self-efficacy and not on responses to any of the other self-efficacy items). absence of differential item functioning (no dif) (i.e. responses to a self-efficacy item depends only on the level of self-efficacy and not on persons’ membership of subgroups such at gender, age, degree program, etc.). optimal reliability, as items are conditionally independent. score sufficiency (i.e. the sum score is a sufficient statistic for the latent self-efficacy score), which is a property only provided by the rasch model. the property of sufficiency is desirable when using the summed raw score of a scale, as it is the usual case with the present self-efficacy scales. thus, fit to the rasch model allows the use of the summed raw score or the estimated person parameters (sometimes called rasch-scores) subsequently. thus, the choice of using the summed raw score or the estimated person parameters has to be made in each case, as it depends on both the purpose for using the score (i.e. for statistical analysis or individual assessment), the length of the scale, the targeting and the reliability of the scale, and so on. one additional factor to take into account in relation to the choice of using sum scores or person parameters is the interpretability of these; the first is easily interpreted in relation to item scores, while the latter is not as it is a logit scale. these considerations extend to the graphical loglinear rasch models described below. specifically, we used the partial credit model (pcm; masters, 1982), which is a generalization of the rasch model for ordinal items – thus this is denoted rasch model (rm) throughout the article. dif will to some degree cause bias in the scores (sum or person parameters) of some groups relative to other groups, thus affecting the validity of the scales and limiting their use in future studies where comparisons are made, as such comparisons would be confounded (kreiner, 2007; nielsen & kreiner, 2013). therefore, adjustment for dif is particularly important in relation to subsequent comparisons of the self-efficacy of students across degree programs and institutions. local dependence between items will not confound subsequent comparisons of subgroups, but will instead inflate reliability, if not taken into account. in order to accommodate these issues, we used graphical loglinear rasch models (gllrms; kreiner & christensen, 2002, 2004, 2007), when fit to the rasch model was rejected, as gllrms allow adjustments for local dependence between items and of differential item functioning (dif or item bias), when these departures are uniform (i.e. are of the same strength at all levels of the latent trait), while retaining close to optimal measurement (kreiner & christensen, 2007). 2.3.1 item analysis overall tests of fit (i.e. tests of global homogeneity by comparison of item parameters in low and high scoring groups) and of no global dif were conducted using andersen’s (1973) conditional likelihood ratio test (clr). effectively, this tests the hypothesis that the item parameters were the same for those students with a low level of self-efficacy and those with a high level of self-efficacy (i.e. overall fit), and the hypothesis that the item parameters were the same for subpopulations of students defined by university, program gender and so on (i.e. no global dif). the fit of individual items was tested by comparing the observed item-rest-score correlations with the expected item-rest-score correlations under the model (i.e. the rm or the specified gllrm) (kreiner & christensen, 2004). the presence of ld and dif in gllrms was tested using two tests; kelderman’s (1984) conditional likelihood ratio test of local independence (i.e. no dif, no ld), and conditional independence tests using partial goodman-kruskal gamma coefficients for the conditional association between item pairs (presence of ld) or between items and exogenous variables (presence of dif) given scores (kreiner & christensen, 2004). specifically, we tested for dif in relation to institution (university of copenhagen, university of aarhus), degree program (medicine, psychology, political science, economy, law, biology, computer science, danish language, history), program priority (1st, 2nd or less), gender (male, female), age groups (20 years or younger, 21 years or older). evidence of overall fit and no global dif found in the overall tests (clr) was rejected if this was not supported by lack of evidence of both ld and dif and evidence of individual item fit. reliability was estimated using hamon and mesbah’s (2002) monte carlo method for reliability in rasch scales, as this method can overcome violations of the assumption of locally independent items, in contrast to cronbach’s alpha which tends to overestimate reliability in such cases. targeting was assessed numerically by two indices as well as graphically, to allow for both graphical and numerical evaluation of targeting (kreiner & christensen, 2013). numerically, we calculated the test information target index as the mean test information divided by the maximum test information for theta, and the root mean squared error (rmse) target index as the minimum standard error of measurement divided by the mean standard error of measurement for theta. both indices should preferably have a value close to one. in addition, we estimated the target of the observed score and the standard error of measurement (sem) of the observed score. graphically, we plotted so-called items maps with the distribution of the item threshold locations against weighted maximum likelihood estimations of the person parameter locations as well as the person parameters for the population (assuming a normal distribution) and the information function. the benjamini-hochberg (1995) procedure was used to adjust for false discovery rate (fdr) due to multiple testing, whenever appropriate. as recommended by cox et al. (1977), we did not apply a critical limit of 5% for p-values as a deterministic decision criterion. instead, we distinguished between weak (p <. 05), moderate (p < .01) and strong (p < .001) evidence against the model. 2.4 pre-academic learning self-efficacy and half-year attrition in order to answer the second research question, the association between the measures of academic self-efficacy and attrition was investigated using the adjusted sum score scales resulting from the item analysis described above. the linear probability model (lpm) was applied using a dichotomous outcome variable indicating dropout within the first half year (1) or not (0) (i.e. the dependent variable). lpm models the probability of dropout as a linear function of covariates, and present the conditional expectation of the outcome (greene, 2011). when we use lpm, ï�¢-coefficients can be interpreted as differences in the probability of attrition in percentage points corresponding to a +1 difference in the independent variables (e.g. pal-se). this straightforward interpretation is one of the reasons why the linear probability model has been an increasingly popular choice for modelling binary outcomes (breen et al, 2018: 4.12). the pal-se variable has been rescaled to 0-1 for the lpm analysis, in order to make findings more intuitive. three models where applied to answer the question about how pre-academic self-efficacy relates to attrition: 1) a bivariate model including academic self-efficacy and half-year attrition. this is in order to investigate whether pal-se can be used to predict differences in half-year attrition without taking any other covariates into account. 2) a model including academic self-efficacy, half-year attrition and a nominal variable indicating educational program as a covariate. this was in order to investigate whether a correlation between pal-se and attrition could simply be due to differences in the type of students between programs that also relates to both pal-se scores and dropout, regardless of student’s educational experience. this is to answer the question whether pal-se can be used to predict differences in half-year attrition if differences between programs are taken into account. from an institutional point of view, this can be valuable information although it does not imply causality. 3) a model including academic self-efficacy, half-year attrition, a nominal variable indicating educational program and a range of individual covariates as for example earlier educational attainment (gpa) and a range of social background variables (parents education, income etc.).we thus investigated the correlation between pal-se and attrition in a stratified comparison comparing students with different values of pal-se who are otherwise as similar as possible. this was to get an indication as to whether a correlation between pal-se and attrition could be explained by differences in background variables or whether it alternatively could be interpreted as a causal effect of pal-se on attrition. 2.5 software all item analyses by rm and gllrms were conducted in digram (kreiner, 2003; kreiner & nielsen, 2013), while r was used for some of the graphical illustrations. the analysis of the effect of pre-academic learning self-efficacy on half-year attrition was conducted in stata. 3. results the pre-academic learning self-efficacy subscale (pal-se) did not fit the rasch model, nor did the pre-academic exam self-efficacy subscale (pal-se). however, for both subscales it was possible to establish fit to a gllrm, though of different complexity (figure 1). global tests-of-fit and dif for the two subscales are presented in table 2, while item fit statistics are provided in table 3. the pal-se subscale fitted a gllrm with strong local dependence between items 1 (i'm certain i can understand the most difficult material presented in the readings for the first courses in my program) and 3 (i'm confident i can understand the most complex material presented in the first courses in my program), as well as very weak local dependence between items 2 (i'm confident i can understand the basic concepts taught in the first courses in my program) and 4 (i'm certain i can master the skills being taught in the first courses in my program). in addition, item 2 (see above) was found to function differentially in relation to the degree programs. this meant that the student responses to item 2 differed systematically in how likely they were to feel described as future students by the statement dependent on the degree program they were admitted to, no matter their level of pre-academic learning self-efficacy. the pae-se subscale also fitted a gllrm, but in this case all items were locally dependent. the strongest local dependence was found between items 2 (i'm confident i can do an excellent job on the assignments and tests in the first courses in my program) and 3 (i expect to do well in the first courses in my program), while the remaining instances of local dependence were both weak. no evidence of dif was found in relation to the pae-se scale. figure 1. the final graphical loglinear rasch models for the pre-academic learning self-efficacy subscale (left side) and the pre-academic exam self-efficacy subscale (right side). notes. correlations are partial goodman & kruskal’s gamma coefficients. in the gllrm for the pal-se scale no gamma coefficient is available to describe the dif of item two relative to degree program, as this is a nominal variable the chi-square value is shown instead. similarly, in the gllrm for the pae-se scale the correlation between degree program and the pae-se score is so low (.01) because the association is not an ordered one across the nine nominal degree programs. table 2. global tests-of-fit and differential item function for the pre-academic exam self-efficacy and the pre-academic learning self-efficacy subscales for the total sample notes. pal-se: pre-academic learning self-efficacy; pae-se: pre-academic exam self-efficacy; rm: rasch model; gllrm: graphical loglinear rasch model; clr: conditional likelihood ratio; df: degrees of freedom; p: p-value; dif: differential item function. global homogeneity test compares items parameters in approximately equal-sized groups of high and low scoring parents. the critical limits for the p-values after adjusting for false discovery rate were:  5% and 1% limit unaltered. 5% limit p = .0083, 1% limit p = .0017. athe gllrm for the pae-se subscale assumed that some items pairs are locally dependent (items 3 and 4, and items 3 and 2, and items 2 and 1). bthe gllrm for the pal-se subscale assumed that items 1 and 3 are locally dependent and that item 2 functions differentially relative to degree program. table 3. item fit statistics for the pre-academic exam self-efficacy and the pre-academic learning self-efficacy subscales under the respective rm and the gllrms notes. pal-se: pre-academic learning self-efficacy; pae-se: pre-academic exam self-efficacy; ï�§ = goodman & kruskal’s gamma coefficients; rm: rasch model; gllrm: graphical loglinear rasch model. the critical limits for the p-values after adjusting for false discovery rate were:  5% limit unaltered, 1% limit p = .0092.  5% limit p = .0375, 1% limit p = .0058.  5% limit p = .0197, 1% limit p =.0023. a the gllrm for the pae-se subscale assumes that some items pairs are locally dependent (items 3 and 4, and items 3 and 2, and items 2 and 1), while the gllrm for the pal-se subscale assumes that items 1 and 3 are locally dependent and that item 2 functions differentially relative to degree program. 3.1 effect of degree program dif a single item in the pal-se scale suffered from dif relative to degree program. thus, to be able to use either the summed scale scores or the estimated person parameters in subsequent statistical analysis, the dif first has to be taken into account by adjusting scores accordingly. in the appendix, we provide conversion tables providing both the necessary information for converting the summed scale scores to estimated person parameters, the estimated person parameters for all the different subgroups affected by the dif, and dif-adjusted scale scores for these subgroups where appropriate (table a3 parts 1 to 3 for the pal-se scale, and table a4 for the pae-se scale). using the summed and the dif-equated scores for the pal-se scale it is possible to investigate to which extent the degree program dif would confound subsequent statistical comparison, if not adjusted for. the results of such an overall comparison is shown in table 4. table 4. comparison of observed and dif-adjusted mean pre-academic learning self-efficacy scores in the degree program subgroups affected by differential item functioning notes. se: standard error. overall differences in observed mean scores (x2 (8) = 59.1, p < .001). overall differences in adjusted mean scores (x2 (8) = 52.1, p = .001). in addition to the overall tests of difference in the observed mean pal-se scores and the dif-adjusted mean pal-se scores of students admitted to the nine different degree programs, we conducted post-hoc step-wise analyses of pairwise collapsibility of the nine nominal degree program categories (see supplement file for details). in the case of the observed scores, the analysis of pairwise collapsibility resulted in four groups of degree programs with significantly different mean pal-se scores: medicine, political science & biology (mean 12.68, se .087); psychology, law & danish language (mean 12.34, se .089); economy (mean 13.14, se .156); and finally computer science & history (13.65, se .162). the pairwise collapsibility analysis of the dif-adjusted scores resulted in only two groups of degree programs with significantly different mean pal-se scores: medicine, political science, biology, psychology, law & danish language (mean 12.54, se .063); economy, computer science & history (13.41, se .114). 3.2 targeting and reliability targeting of the pal-se and pae-se scales was very different (table 5). targeting of the pal-se scale was considered good for all students, though it varied slightly across subgroups of students admitted to the nine different degree programs (ranging from 73% to 81% of the maximum information obtained for students admitted to biology and economy, respectively). targeting of the pae-se scale was the same for all students and not good (52% of the maximum information obtained). the differences in targeting are also illustrated by the item maps showing the relative locations of person parameters and item thresholds along the two theta scales as well as the test information curves (figure a1 in the appendix file). the item map for the pae-se scale shows that even though person parameters and item thresholds are both spread out along the entire scale, maximum information is located at the lower end of the scale where hardly any students are located, and there is much less information where most of the students are concentrated. for the pal-se scale, the separate item maps for students admitted to the nine different degree programs also show that person parameters and item thresholds are both spread out along, more or less, the entire scale. however, for the pal-se scale, the maximum information is found closer to more of the student locations, and there is much more information where most of the students are located than was the case with the pae-se scale. the reliability of the pae-se scale (all students) was .77, while the reliability of the pal-se scale varied slightly across subgroups of students admitted to the nine different degree programs; from .79 for students admitted to biology and economy, and .86 for students admitted to computer science (weighted average for all students .81) (table 5). thus, by conventional standards, the reliability of both scales was at acceptable levels. table 5. targeting and reliability of the pre-academic exam self-efficacy and the pre-academic learning self-efficacy subscales notes. pal-se: pre-academic learning self-efficacy; pae-se: pre-academic exam self-efficacy; ti: test information; rmse: the root mean squared error of the estimated theta score; sem: the standard error of measurement of the observed score; r: reliability. a. targeting and reliability is provided for groups defined by dif variables. 3.3 pre-academic learning self-efficacy and half-year attrition based on the results of the item analyses of the pal-se and the pae-se scales (rq1), particularly the marked difference in the targeting of the two (table 5 and figure a1 in the appendix), we decided to proceed to rq2 (how is academic self-efficacy at admission associated with half-year attrition) using only the pal-se scale. we report the result of the regression analysis using the dif-adjusted sum score of the pal-se, as there were no noteworthy differences in the results using the estimated person parameters, and the results using the adjusted sum score are more readily interpreted as they relate directly to the instrument. the analysis showed a significant association between pal-se and half-year attrition; the higher students scored on pal-se the lower was the probability for them to drop out within the first half year in the degree program (figure 2 and table a5 in the appendix). the association was even stronger when program was included in the model as a covariate, and could thus not be explained by differences in the type of students who were admitted to the different programs. this leads to the conclusion that pal-se can predict significant differences in students’ probability to drop out within the first half year. more precisely the probability of dropping out within the first half year was 7.2% lower for students with the highest possible level of pal-se compared to students with the lowest possible level of pal-se, when degree program was included as a covariate. this is considered a substantial difference, considering the overall dropout percentage in the sample was 3.9 % during the first half year. adding individual controls did not change the correlation substantially. hence, the correlation between pal-se and attrition also could not be explained by differences in the included individual covariates indicating that pal-se has a causal effect on attrition. the fact that the correlation did not change substantially when we added covariates further supports a causal interpretation (heckman, humphries et al. 2016). the design is however not a strong design in terms of potential selection bias stemming from unobservable heterogeneity. overall, in terms of causality, the results thus indicate that pal-se does have an effect on student attrition, although especially the size of this effect is associated with uncertainty. figure 2. results from regression analysis by linear probability model: association between pre-academic learning self-efficacy on half-year attrition in models with different covariates included. notes. the estimated effect corresponds to a difference in pal-se from minimum to maximum, since pal-se is measured on a scale ranging from 0-1. we use the adjusted raw score from the item analysis to measure pal-se. individual controls included as covariates: 1) priority of program in application 2) sex 3) age 4) grade point average from high school [danish gymnasium] 5) parents highest completed educational level 6) parents income 7) ethnicity 8).n: bivariate; 2.340, n: with program as covariate; 2.340, n: with program and individual covariates: 2.195. 4. discussion and future directions this aim of the study was firstly to investigate the construct validity and psychometric properties of the proposed pal-se and pae-se scales in a large sample of university students just admitted to the same nine degree programs in two large universities, using rasch models. secondly, and depending on the results of the psychometric part of the study, the aim was to investigate how self-efficacy before commencement of degree program was associated with half-year attrition. with regard to the first aim, the construct validity of the pal-se and pae-se scale was to a large extent confirmed, though with different psychometric properties. thus, neither scale fit a rasch model. the pal-se scale was found to be, psychometrically speaking, the better of the two scales. we found evidence of differential item functioning relative to degree program, but it was possible to address this by adjusting the score. we did not find any evidence for differential items functioning for the pae-se scale, only evidence of local dependence among items, this we also found to some degree in the pal-se scale. the main reason that we find the pal-se scale to be the better of the two scales are the differences in targeting; the targeting of the palse scale was good, while the targeting of the pae-se scale was poor. the items in both scales are spread out across the latent scale, as are the person estimates, however the level of information in the pae-se scale is substantially lower than in the pal-se scale. thus, the items in the pal-se scale provide more information on the students in the study sample than it is the case with the pae-se items. furthermore, the variations in the obtained information with the pal-se scale for students from the different degree programs vary very little, even though the admission gpa across degree programs range from “all admittedâ€� (i.e. no admission gpa) to 11.5 (the highest is 12). this finding adds substantially to the field, as no previous attempts have been made to investigate the construct validity and psychometric properties of an instrument for pre-academic self-efficacy using the framework of graphical loglinear rasch models. when considering that the scales are intended to measure learning and exam self-efficacy, respectively, after admission but before commencement of university studies, we find the difference in targeting to be further evidence of construct validity. all things being equal, it is reasonable to assume that the pal-se items should be better suited to the study population, as the students probably have a more well-developed idea about their future learning capacity, as they have applied to a program most likely of their interest and have obtained the admission gpa necessary to be admitted, while they do not have as well-developed a sense of the exams in the university system, and thus the items carry less information for this student population. another entirely new finding is that we have established that it is indeed possible to obtain invariant measures of both pre-academic learning self-efficacy and pre-academic exam self-efficacy for students admitted to a wide range of university degree programs and at two different universities, when making the proper adjustments. this new finding will prove valuable in future studies on self-efficacy, whether pre-academic self-efficacy or academic self-efficacy measured during studies, as it paves the way for more studies assuring the application of unbiased self-efficacy scores. the lpm regression analysis adds to the current evidence that self-efficacy is strongly correlated to educational outcomes, such as attrition. it more specifically points out that this is also the case, even though we measured the level of self-efficacy before the students started on the university program and as a measure that did not relate to a specific course but to the upcoming courses in the first semester as a whole. van herpen and colleagues (2017), using an adaptation of the self-efficacy scale of the mslq (pintrich et al., 1991) where they changed items from being course-specific to referring in general to “the first year of studyâ€�, found no association between pre-university self-efficacy and academic success. the instrument they used was the most similar to ours, but they treated it as a single scale, did not investigate its measurement properties, and their sample size was too small to take academic discipline into account in the analysis. in a study on discipline-specific differences in general self-efficacy prior to commencing university, yong (2010) found no such differences with small samples of future engineering and business students. the only study we located, which investigates pre-higher education self-efficacy and attrition was that of aryee’s (2017), which found that high levels of pre-higher education mathematics self-efficacy were associated with completion of college, while low levels were associated with a diminished likelihood of attaining a degree. our findings are thus an expansion of the very scarce research on pre-higher education self-efficacy, and the even scarcer research on its association with attrition, as it represents the first multi degree program study on this relationship. furthermore, our findings are in line with the research where self-efficacy is measured during degree programs, typically during the first semester, as they find that low self-efficacy is associated with increased attrition, increased risk of withdrawal, decreased intention to persist, and decreased intention to complete (de clercq, galand & frenay, 2017; elliot, 2016; lent et al., 2016; thomas, 2014; vuong, brown-welty & tracz, 2010; you, 2018). even though we still need further evidence relating to the causal link between pre-academic learning self-efficacy and attrition, as we do not have a very strong design in terms of causality, the results do indicate that pal-se has a causal effect on attrition. the results are promising, since self-efficacy to a large extend is regarded as context specific and malleable, thus providing an opportunity for teachers and other educational professionals to promote students’ self-efficacy in order for them to achieve higher learning outcomes and reduce attrition. the current study is the first to show that even prior to commencing the degree program students admitted to university are learning-wise self-efficacious with regard to the upcoming courses and that their level of pre-academic learning self-efficacy is associated with attrition. we therefore call for three lines of further studies. firstly, studies to replicate the findings with a stronger causal design. secondly, applied studies to show how end how much teachers and other educational professionals are able to affect student’s self-efficacy positively. thirdly, controlled intervention studies showing the effect of pedagogical strategies to enhance self-efficacy on attrition, intention to persist and complete degree programs. keypoints short scales measuring pre-academic learning self-efficacy (pal-se) and pre-academic exam self-efficacy (pae-se) were developed from existing scales. validation data was drawn from a national sample of students admitted to, but prior to starting at, university, using a-priory defined criteria. using graphical loglinear rasch models the pal-se and pae-se were found to be construct valid. the pal-se was used for investigating the relationship between pre-academic self-efficacy and half-year attrition. results from restricted models indicate that pre-academic learning self-efficacy has a substantial effect on half-year attrition. acknowledgments the authors would like to thank pedro henrique ribeiro santiago for providing us with the r-code for the items maps preliminary to its release. references andersen, e. b. (1973). a goodness of fit test for the rasch model. psychometrika, 38(1), 123–140. doi: 10.1007/bf02291180 aryee, m. (2017). college students’ persistence and degree completion in science, technology, engineering, and mathematics (stem): the role of non-cognitive attributes of self-efficacy, outcome expectations, and interest, (doctoral dissertation) . available from: proquest dissertations and theses databases. (umi no. 10264673) bandura, a. (1997). self-efficacy. the exercise of control. new york, ny: freeman. doi: 10.5860/choice.35-1826 bartimote-aufflick, k., bridgeman, a., walker, r., sharma, m., & smith, l. (2015). the study, evaluation, and improvement of university student self-efficacy. studies in higher education, (ahead-of-print), 1-25. doi: 10.1080/03075079.2014.999319 benjamini, y., & hochberg, y. (1995). controlling the false discovery rate: a practical and powerful approach to multiple testing. journal of the royal statistical society. series b (methodological) , 57(1), 289–300. doi: 10.1111/j.2517-6161.1995.tb02031.x breen, r.; karlson, k. b. & holm, a. (2018). interpreting and understanding logits, probits, and other nonlinear probability models. annual review of sociology, 44, 4.1–4.16. doi: 10.1146/annurev-soc-073117-041429 burger, k. & samuel, r. (2017). the role of perceived stress and self-efficacy in young people’s life satisfaction: a longitudinal study. journal of youth and adolescence, 46, 78–90. doi: 10.1007/s10964-016-0608-x cox, d. r.; spjã¸tvoll, e.; johansen, s.; van zwet, w. r.; bithell, j. f. & barndorff-nielsen, o. (1977). the role of significance tests [with discussion and reply]. scandinavian journal of statistics, 4, 49–70. credã©, m., & phillips, l. a. (2011). a meta-analytic review of the motivated strategies for learning questionnaire. learning and individual differences, 21(4), 337–346. doi: 10.1016/j.lindif.2011.03.002 de clercq, m.; galand, b. & frenay, m. (2017). transition from high school to university: a person-centered approach to academic achievement. european journal of psychology of education, 32(1), 39-59. doi: 10.1007/s10212-016-0298-5 devonport, t. j. & lane, a. m. (2006). relationships between self-efficacy, coping and student retention.social behavior and personality: an international journal, 34(2), 127-138. doi: 10.2224/sbp.2006.34.2.127 duncan, t. g. & mckeachie, w. j. (2005). the making of the motivated strategies for learning questionnaire. educational psychologist, 40(2), 117–128. doi: 10.1207/s15326985ep4002_6. elliott, d. c. (2016). the impact of self beliefs on post-secondary transitions: the moderating effects of institutional selectivity. higher education, 71(3), 415-431. doi: 10.1007/s10734-015-9913-7 ferla, j.; valcke, m. & cai, y. (2009). academic self-efficacy and academic self-concept: reconsidering structural relationships. learning and individual differences, 19(4), 499-505. doi: 10.1016/j.lindif.2009.05.004 ginsborg, j.; kreutz, g.; thomas, m. & williamon, a. (2009). healthy behaviours in music and non-music performance students. health education, 109(3), 242-258. doi: 10.1108/09654280910955575 greene, w. h. (2011). econometric analysis. seventh edition. upper saddle river, prentice hall. hamon, a. & mesbah, m. (2002). questionnaire reliability under the rasch model. in: mesbah, m.; cole, b. f. & lee, m. t. (eds.). statistical methods for quality of life studies. dordrecht: kluwer academic publishers, pp. 155-68. doi: 10.1007/978-1-4757-3625-0_13 heckman, j.j.; humphries, j.e. and veramendi, g. (2016). returns to education: the causal effects of education on earnings, health and smoking. nber working paper, may 2016 (22291). huang, c. (2013). gender differences in academic self-efficacy: a meta-analysis. european journal of psychology of education, 28(1), 1–35. doi: 10.1007/s10212-011-0097-y kelderman, h. (1984). loglinear rasch model tests. psychometrika, 49, 223-245. doi: 10.1007/bf02294174 kreiner, s. (2003). introduction to digram. copenhagen: department of biostatistics, university of copenhagen. kreiner, s. (2007). validity and objectivity. reflections on the role and nature of rasch models. nordic psychology, 59, 268-298. doi: 10.1027/1901-2276.59.3.268 kreiner, s. (2013). the rasch model for dichotomous items. in christensen, k. b.; kreiner, s. & mesbah, m. (eds.) rasch models in health. london: iste ltd, wiley, pp. 5–26. doi: 10.1002/9781118574454.ch1 kreiner, s. & christensen, k.b. (2002). graphical rasch models. in: mesbah, m.; cole, b. f. & lee, m. t. (eds.) statistical methods for quality of life studies. dordrecht: kluwer academic publishers, pp. 187–203. doi: 10.1007/978-1-4757-3625-0_15 kreiner, s. & christensen, k. b. (2004). analysis of local dependence and multidimensionality in graphical loglinear rasch models. communication in statistics – theory and methods, 33(6), 1239–1276, doi: 10.1081/sta-120030148 kreiner, s. & christensen, k. b. (2007). validity and objectivity in health-related scales: analysis by graphical loglinear rasch models. in von davier, m. & carstensen, c. h. (eds.) multivariate and mixture distribution rasch models, new york, springer, pp. 329-346. doi: 10.1007/978-0-387-49839-3_21 kreiner, s. & christensen, k. b. (2013). person parameter estimation and measurement in rasch models. in christensen, k. b.; kreiner, s. & mesbah, m. (eds.) rasch models in health. london: iste ltd, wiley, pp. 63–78. doi: 10.1002/9781118574454.ch4 kreiner, s. & nielsen, t. (2013). item analysis in digram 3.04. part i: guided tours. research report 2013/06 . university of copenhagen, department of public health. lent, r. w.; miller, m. j.; smith, p. e.; watford, b. a.; lim, r. h. & hui, k. (2016). social cognitive predictors of academic persistence and performance in engineering: applicability across gender and race/ethnicity. journal of vocational behavior, 94, 79-88. doi: 10.1016/j.jvb.2016.02.012 lindstrã¸m, c. & sharma, m. d. (2011). self-efficacy of first year university physics students: do gender and prior formal instruction in physics matter? international journal of innovation in science and mathematics education (formerly cal-laborate international), 19 (2), 1–19. luszczynska, a.; gutiã©rrez-doã±a, b. & schwarzer, r. (2005). general self-efficacy in various domains of human functioning: evidence from five countries. international journal of psychology, 40(2), 80-89. doi: 10.1080/00207590444000041 masters, g. n. (1982). a rasch model for partial credit scoring. psychometrika, 47, 149-174, doi: 10.1007/bf02296272 mellenbergh, g. j. (1989). item bias and item response theory. international journal of educational research, 13, 127-143. doi: 10.1016/0883-0355(89)90002-5 meredith, w. (1993). psychometrika, 58(4), 525-543. doi:10.1007/bf02294825 mesbah, m. & kreiner, s. (2013). the rasch model for ordered polytomous items. in christensen, k. b.; kreiner, s. & mesbah, m. (eds.) rasch models in health. london: iste ltd, wiley, pp. 27-42. doi: 10.1002/9781118574454.ch2 multon, k. d.; brown, s. d. & lent, r. w. (1991). relation of self-efficacy beliefs to academic outcomes: a meta-analytic investigation. journal of counseling psychology, 38(1), 30-38.doi: 10.1037/0022-0167.38.1.30 muris, p. (2002). relationships between self-efficacy and symptoms of anxiety disorders and depression in a normal adolescent sample. personality and individual differences, 32(2), 337–348. doi: 10.1016/s01918869(01)00027-7 nielsen, t.; dammeyer, j.; vang, m. l. & makransky, g. (2018). gender fairness in self-efficacy? a rasch-based validity study of the general academic self-efficacy scale (gase). scandinavian journal of educational research, 62(5), 664-681. doi: 10.1080/00313831.2017.1306796 nielsen, t. & kreiner, s. (2013). improving items that do not fit the rasch model: exemplified with the physical functioning scale of the sf-36. annales de l’i.s.u.p. publications de l’institut de statistique de l’universitã© de paris , numero special, 57(1-2), 91-108. nielsen, t.; makransky, g.; vang, m. l. & dammeyer, j. (2017). how specific is specific self-efficacy? a construct validity study using rasch measurement models. studies in educational evaluation, 57, 87-97. doi: 10.1016/j.stueduc.2017.04.003 pajares, f. (1997). current directions in self-efficacy research. in m. maehr & p. r. pintrich (eds.), advances in motivation and achievement (vol. 10). greenwich, ct: jai press, pp. 1–49. pintrich, p.r.; smith, d.a.f.; garcia, t. & mckeachie, w. j. (1991). a manual for the use of the motivated strategies for learning questionnaire (mslq). technical report no. 91-8-004. the regents of the university of michigan. rasch, g. (1960). probabilistic models for some intelligence and attainment tests. copenhagen, danish institute for educational research. richardson, m.; abraham, c. & bond, r. (2012). psychological correlates of university students' academic performance: a systematic review and meta-analysis. psychological bulletin, 138(2), 353-387. doi: 10.1037/a0026838. saleh d.; camart n. & romo l. (2017). predictors of stress in college students. frontiers in psychology, 8(19). doi: 10.3389/fpsyg.2017.00019 scherbaum, c. a.; cohen-charash, y. & kern, m. j. (2006). measuring general self-efficacy: a comparison of three measures using item response theory. educational and psychological measurement, 66(6), 1047–1063. doi:10. 1177/0013164406288171 scherer, r. & siddiq, f. (2015). revisiting teachers’ computer self-efficacy: a differentiated view on gender differences. computers in human behavior, 53, 48–57. doi:10.1016/j.chb.2015.06.038 schunk, d. & pajares, f. (2002). the development of academic self-efficacy. in a. wigfield, & j. eccles (eds.), development of achievement motivation. san diego: academic press, pp. 15-31. doi: 10.1016/b978-012750053-9/50003-6 schwarzer, r. & jerusalem, m. (1995). generalized self-efficacy scale. measures in health psychology: a user’s portfolio. causal and control beliefs, 1, 35–37. tahmassian, k. & jalali moghadam, n. (2011). relationship between self-efficacy and symptoms of anxiety, depression, worry and social avoidance in a normal sample of students. iranian journal of psychiatry and behavioral sciences, 5(2), 91–98. thomas, d. (2014). factors that influence college completion intention of undergraduate students. the asia-pacific education researcher, 23(2), 225-235. doi: 10.1007/s40299-013-0099-4 van der linden, w. j. & hambleton, r. k. (1997). handbook of modern item response theory. springer-verlag, new york. doi: 10.1007/978-1-4757-2691-6 van herpen, s. g.; meeuwisse, m.; hofman, w. a.; severiens, s. e. & arends, l. r. (2017). early predictors of first-year academic success at university: pre-university effort, pre-university self-efficacy, and pre-university reasons for attending university. educational research and evaluation, 23(1-2), 52-72. doi: 10.1080/13803611.2017.1301261 vuong, m.; brown-welty, s. & tracz, s. (2010). the effects of self-efficacy on academic success of first-generation college sophomore students. journal of college student development, 51(1), 50-64. doi: 10.1353/csd.0.0109 watt, h. m. g.; ehrich, j.; stewart, s.e.; snell, t.; bucich, m.; jacobs, n.; furlonger, b. & english, d. (2019). development of the psychologist and counsellor self-efficacy scale.higher education, skills and work-based learning, (early cite). doi: 10.1108/heswbl-07-2018-0069 williams, b. w.; kessler, h. a. & williams, m. v. (2014). relationship among practice change, motivation, and selfâ€�efficacy.journal of continuing education in the health professions, 34(s1), 5-10. doi: 10.1002/chp.21235. yong, f. l. (2010). a study on the self-efficacy and expectancy for success of pre-university students. european journal of social sciences, 13(4), 514-524. young, a. m., wendel, p. j., esson, j. m., & plank, k. m. (2018). motivational decline and recovery in higher education stem courses. international journal of science education, 40(9), 1016-1033. doi: 10.1080/09500693.2018.1460773 zimmerman, b. j.; bandura, a. & martinez-pons, m. (1992). self-motivation for academic attainment: the role of self-efficacy beliefs and personal goal setting. american educational research journal, 29(3), 663-676. doi: 10.2307/1163261 appendix the self-efficacy scales table a1. items and response scales of the danish and english versions of the pre-academic learning self-efficacy (pal-se) and pre-academic exam self-efficacy (pae-se) scales adapted from nielsen et al. (2018) notes. question prompt for the items was: how well do the following statements describe you as a future student? items were administered mixed with items from other scales. here they are shown by subscale, not the order they were administered in.   post-hoc step-wise analyses of pairwise collapsibility of the nine nominal degree program categories in the stepwise analyses, the two categories with the largest categories were collapsed if the p-value was larger than the critical p-value. critical p-values were determined while controlling the false discovery rate at 0.05 relative to the total number of tests performed at the end of each step, using the benjamini-hochberg procedure (benjamini & hochberg, 1995; kreiner & nielsen, 2013). the steps and decisions in the analysis of the observed mean pal-se scores and the dif-adjusted mean pal-se scores are shown in table s-ã†, while the collapsed groups of non-different degree programs are provided in the results section. table a2. steps and decisions in the analyses of pairwise collapsibility notes : 1 = medicine, 2 = psychology, 3 = political science, 4 = economy, 5 = law, 6 = biology, 7 = computer science, 8 = danish language, 9 = history. conversion tables table a3 – part 1. weighted maximum likelihood estimates of person parameters and dif-equating of the sum score for the pre-academic learning self-efficacy scale table a3 – part 2. weighted maximum likelihood estimates of person parameters and dif-equating of the sum score for the pre-academic learning self-efficacy scale table a3-part 3. weighted maximum likelihood estimates of person parameters and dif-equating of the sum score for the pre-academic learning self-efficacy scale notes. p.param = person parameter. sem = standard error of measurement table a4. weighted maximum likelihood estimates of person parameters for the pre-academic exam self-efficacy scale notes. p.param = person parameter. sem = standard error of measurement item maps figure a1. item maps with distributions of person parameter locations and information curve above item threshold locations. notes. person parameters are weighted maximum likelihood estimates and illustrate the distribution of these for the study sample (black bars above the line) and for the population under the assumption of normality (grey bars above the line), as well as the information curve, relative to the distribution of the item thresholds (black bars below the line). for the pal-se scale item maps are shown for students admitted to each of the nine degree programs, as the scale functioned differentially relative to degree program (top nine graphs). for the pae-se scale a single item map is shown for all students (bottom graph). codepen iaconelli frontline learning research vol.8 no. 3 special issue (2020) 104 125 issn 2295-3159 insufficient effort responding in surveys assessing self-regulated learning: nuisance or fatal flaw? ryan iaconelli a, & christopher a. woltersa adennis learning center, department of educational studies, the ohio state university, usa article received 18 june 2019 / revised 9 october / accepted 30 october / available online 30 march abstract despite concerns about their validity, self-report surveys remain the primary data collection method in the research of self-regulated learning (srl). to address some of these concerns, we took a data set comprised of college students’ self-reported beliefs and behaviours related to srl, assessed across three surveys, and examined it for instances of a specific threat to validity, insufficient effort responding (ier; huang, curran, keeny, poposki, & deshon, 2012). using four validated indicators of ier, we found the rate of ier to vary between 12-16%. critically, while we found that students characterised as inattentive and attentive differed in some basic descriptive statistics, the inclusion of inattentive students within the data set did not alter more substantial inferences or conclusions drawn from the data. this study provides the first direct examination of the impact of respondents’ attention on the validity of srl data generated from self-report surveys. keywords: insufficient effort responding; self-regulated learning; self-report; validity info corresponding author email: iaconelli.1@osu.edu doi: https://doi.org/10.14786/flr.v8i3.521 1. introduction 1.1 use of srs in studying self-regulated learning researchers have used models of self-regulated learning (srl) to understand engagement, learning, and achievement in academic contexts from preschool through college (perry et al., 2018; pintrich & zusho, 2007; usher & schunk, 2018; winne & hadwin, 2008). models of srl posit that students can plan, monitor, control, and reflect upon their own thoughts, behaviors, and motivation related to their learning (panadero, 2017). engagement in srl requires that students feel both motivated and efficacious to enact these sub-processes (pajares, 2007; pintrich & zusho, 2007). efforts to use srl as a basis for developing instructional policies and practices designed to improve students’ academic success is an accepted goal among educators (cleary & zimmerman, 2004; dignath & buttner, 2008; schunk & zimmerman, 1998). achieving critical goals with regard to these two efforts is in no small part dependent upon the availability of sound methods for assessing srl (winnie & perry, 2000; wolters & won, 2018). the need for sound assessment of srl has spawned the development of many methods (azevedo et al.,2018; winne & perry, 2000; zimmerman, 2008). for instance, observing students’ behaviours within the classroom, recording traces of their thinking or behaviour when engaged in academic tasks, and reports by teachers or parents have all been used to assess students’ srl (biswas, baker, & paquette, 2018; cleary & callan, 2018). despite the promising increase in the diversity of assessments, self-report surveys (srs) remain the most common method used to assess srl (winne & perry, 2000). we use the term srs to describe any type of questionnaire or survey in which respondents are presented with a question or statement and asked to provide a response, either retrospectively or concurrently, based on their own beliefs, attitudes, or behaviours. although they offer many advantages (butler, 2002; mccardle & hadwin, 2015; wolters & won, 2018), criticisms of srs, including fundamental questions regarding the validity of the data they produce (karabenick & zusho, 2011; schellings & van hout-wolters, 2011; winne & jamieson-noel, 2003). developments in the manner under which srs can be completed, such as the increased use of unsupervised internet-based administrations (e.g. qualtrics, redcap, amazon mechanical turk), raises concerns about the evidence for validity from data produced when students complete online srs for educational research. responding appropriately to items on an srs is a function of a complex, multi-step process. to authentically respond to an item, the respondent must read and understand the item, search their memory for relevant information, integrate any activated memories into a coherent answer, match this answer to one of the available response options, and finally decide whether to select that or some other response option (duckworth & yeager, 2015). at any point in this process, lack of motivated engagement or inattention on the part of the respondent is a threat to the validity of individual items as well as the overall data produced. hence, it is important to carefully evaluate potential threats to the validity of the data produced from these srs (wolters & won, 2018). we address this need by evaluating college students’ responses to three online srs designed to assess factors associated with their srl, motivation, and academic success for issues related to inattentive responding. 1.2 insufficient effort responding threats to validity that result from respondents’ lack of cognitive effort, inattention, or motivation when completing surveys have been examined under several different names, including careless responding (meade & craig, 2012), insufficient effort responding (ier; huang et al., 2012), and low-quality data (desimone & harms, 2017). we adopt ier as our preferred term and it is defined as “a response set in which the respondent answers a survey measure with low or little motivation to comply with survey instruction, correctly interpret item content, and provide accurate responses” (huang et al., 2012, p. 100). in other words, ier occurs when a respondent does not provide the necessary cognitive effort required to go through the multi-step process needed to produce data that appropriately represents the underlying construct. the ier framework and the methods used to identify it encompass physical, cognitive, and motivational disengagement from the survey, all of which threaten its’ validity. ier may occur for many reasons. for instance, a respondent may be fatigued, distracted by their surroundings, or motivated to complete the srs as fast as possible and without effort, either because they are forced to take the survey or because they are taking it solely for some promised compensation, like money or extra credit in a course (johnson, 2005). mischievous responding, in which a respondent purposefully provides systematically invalid responses (e.g. selecting the same response for all items) can be considered a form of ier because, even though they are providing some effort, it is not directed at interpreting and responding to items appropriately. socially desirable responding and other unintentional biases that may alter a person’s response patterns do not fall under the ier umbrella because these respondents are working to read, understand, and answer items appropriately. researchers have appreciated the threats to validity represented by the underlying causes of ier for some time (beach, 1989; nichols et al., 1989). it was not until recently though that researchers began to investigate more vigorously the propensity and importance of ier as a detriment to validity. two key findings have emerged from this work. one, the exact level of ier within a data set varies from survey to survey and sample to sample, but generally about 10% of respondents are identified as having engaged in some form of ier (meade & craig, 2012; maniaci & rogge, 2014). two, while this proportion may not seem noteworthy, the inclusion of even a small percentage of inattentive responders, as low as five percent, can cause spurious relationships among otherwise uncorrelated measures to become significant, mask otherwise significant relationships between variables, lead to inflated mean-level scores on latent constructs, and alter effect size differences between groups (huang et al., 2017). this line of research raises serious concerns for those who rely on srs as a primary means of data collection, srl researchers included. there are reasons however to question the extent to which this previous research is applicable to srl, primary among them is that most of this research has derived from assessments of personality and other more stable traits. the more dynamic and changeable nature of motivation and srl constructs, compared to personality constructs, may result in different manifestations of ier. additionally, much of the research on ier has utilized samples of undergraduates drawn from participant pools (e.g. dunn et al., 2016; huang et al., 2012; meade & craig, 2012) that presumably had little intrinsic motivation for completing the srs. in contrast, our sample consists of students who anticipated using the results of their srs for meaningful diagnostic purposes as part of a course assignment. thus, in addition to extending ier research into a new field, we also extend this research into a presumably more motivated sample, which should produce different manifestations of ier. 1.2.1 methods for identifying ier to identify instances of ier, researchers can call upon both proactive and reactive indicators. proactive methods are based on the inclusion of “check items” throughout an srs (huang et al., 2012; meade & craig, 2012; huang et al., 2015; maniaci & rogge, 2014; bowling et al., 2016; dunn et al., 2016) that provide a way to determine if respondents are reading each item carefully. some check items direct respondents to a certain answer (e.g. “mark strongly agree for this item”), while others are statements that no respondent should agree with (e.g. “my birthday is february 30”). finally, some researchers (e.g. huang et al., 2012) have simply asked respondents how much effort they put forth when completing items. although proactive measures have proven useful, the focus of this study was on the use of reactive methods to detect ier. reactive methods refer to a variety of post-hoc statistical analyses used to identify ier. with the increased use of technology to administer surveys, using total survey response time, which programs like qualtrics record automatically, has become an easy and effective means of assessing ier (huang et al., 2012; meade & craig, 2012; maniaci & rogge, 2014; bowling et al., 2016; dunn et al., 2016). the basic assumption of this approach is that those who spend very little time completing the survey are not fully engaged in the various cognitive processes necessary to respond to items as intended. for instance, respondents may skim items rather than carefully read them or may quickly select a response without deliberate recollection and reasonable consideration of the events that should inform their response. hence, survey completion times that are extremely short are highly suggestive of ier. a family of statistical analyses designed to assess the consistency of one’s responses are another set of common reactive methods for identifying ier. as one example, even-odd consistency (huang et al., 2012; meade & craig, 2012; johnson, 2005) is a measure of individual reliability that is generated by dividing a survey into two parts (traditionally based on odd and even-numbered items) and calculating the correlation between the two parts. the underlying assumption of this method is that alternate items from unidimensional scales should be strongly correlated. another common individual reliability indicator is referred to as psychometric synonyms/antonyms (curran, 2016). this indicator is computed by identifying the items with the strongest bivariate correlations within a data set and comparing individual respondents’ correlations on these items. semantic synonyms/antonyms are used in the same manner, except that the items used in the calculation of this index are decided on an a priori basis, based on item content. for each of these indices, respondents who exhibit unusually weak correlations for the set of relevant pairs of items are thought to have engaged in ier. this conclusion is based on the assumption that the atypical correlations are a function of not reading items carefully enough, answering items randomly, or utilizing another response strategy that falls substantially short of full engagement in the response process, a process which provides increased evidence for the validity of the data (duckworth & yeager, 2015; winne & perry, 2000). another reactive method of assessing ier is founded on statistical analyses designed to assess the variability of an individual’s responses within a survey and includes indices such as long-string analysis (costa & mccrae, 2008) and individual response variability (irv; dunn et al., 2016). the underlying assumption of these indices is that, on surveys with multiple scales assessing individual difference constructs, individuals should be responding with some degree of variability. that is, when individuals respond to a great many items in a row using the same response(s) options (e.g. answering 5 to many items or answering 4,5,4,5,4), especially across scales that do not assess the same trait, they are likely not paying attention to the items or are actively looking to avoid expending cognitive effort. curran (2016) provides an excellent review of the methods mentioned here, as well as many others that have been used to identify ier. 1.3 the present study because of the overreliance on srs to assess and understand srl, it is important that researchers can trust that respondents address items with the care and attention necessary to provide data for valid conclusions. identifying instances of ier within data sets that contain items assessing the motivational and strategic aspects of srl is one way that researchers can empirically evaluate their data and in part verify the validity of conclusions they draw regarding srl. further, we expand on recent work to identify ier by examining it amongst a group of college students who completed three srs intended to assess various aspects of motivation and srl, over the course of one academic semester. a key contribution of the present research lies in our ability examine ier in a single sample of participants, across multiple surveys measuring different constructs. from an srl perspective, this study provides a direct examination of the possibility that students are not providing the necessary cognitive effort needed when answering items designed to assess their motivation and srl, which threatens any conclusions using data generated by srs. we pursued the following research questions: (1) how prevalent is ier within a sample of college students completing srs that assess constructs associated with srl? (2) does ier manifest itself in a consistent way across different survey administrations within the same sample? (3) does students’ engagement in ier alter the results and conclusions drawn from srs assessing the motivational and strategic aspects of srl? to answer these questions, we utilized four established reactive indices of ier to: a) identify the percentage of students engaging in ier within each survey administration, as well as across survey administrations; b) examine the relationship of these indices across survey administrations; and c) identify students who engaged in ier and test whether their inclusion within the data set impacted basic univariate and multivariate statistics for motivational and strategic aspects of srl. 2. method 2.1 participants participants were students (n = 297) at a large public university in the united states who indicated their ethnicity as white (n = 194, 64%), african-american (n = 41, 13%), hispanic (n = 22, 7%), asian (n = 15, 5%), and other/multiple (n = 29, 11%). our sample included a majority of students who were in their first or second year of college (n = 177, 57%), had an average age of 20.4 years (sd = 4.1) and included more males (n = 168, 55%). 2.2 procedures participants were recruited from 16 sections of a three-credit, letter-graded, semester-long elective course designed to improve students’ srl, and as a result, their overall academic success. as part of their assigned work for the course, students responded to four online srs. the initial survey solicited information about how students learned about the course, their reasons for taking the course, and their knowledge about other academic outreach resources available through the university. the three remaining surveys, from which the data for the present study are drawn, were designed to assess various dispositions, beliefs, attitudes, and behaviours associated with engagement, learning, and academic success. three-hundred and five students provided informed consent that allowed for the use of their course data for research, but eight (2.6%) of those students were missing data to an extent that precluded them from inclusion in the present study. procedures for each of the three relevant surveys were similar. within the course’s learning management website, students were provided a short description of the topic and purpose of the survey and concomitant assignment along with a hyperlink. when students clicked the link, a new browser window appeared displaying the first page of the particular qualtrics-based survey. for the most part, students accessed and completed these surveys outside of the regular class period, at a time and place of their own choosing. as a final step of each survey, students were provided a “score report” that included a short description of their own mean scores for the relevant scales. except for the final survey, these reports served as the basis for a personal reflection assignment and for in-class discussions, which factored into the calculation of the overall course grade. hence, students’ motivation for completing each survey was likely derived from both their interest in obtaining personal insights regarding their own motivation and strategic behaviour as well as the connection to their course grade. 2.3 measures the three relevant surveys (hereafter referred to as week 2 survey, week 6 survey, and week 14 survey for the weeks that they were assigned during the 15-week semester), each started with a short set of directions (121-171 words) and a few items (e.g. university id numbers), that later could be used to link them to a particular student. the initial directions identified the general topic of the survey (e.g. motivation), assured the students that there were no right or wrong answers, and provided information about how to ensure they received credit for the associated assignment. the remainder of each survey was organized into sections, each of which began with a few sentences that identified the more specific topic (e.g. self-confidence) it covered, instructed the students to read each item carefully before responding, and reminded students to be honest in their answers. for most items, students were presented with a statement and asked to indicate the extent to which it applied to themselves using a 5-point response scale ranging from strongly disagree to strongly agree. the self-efficacy for self-regulated learning items, the lone exception to this format, were answered using a 5-point response scale ranging from not confident at all to very confident. the week 2 survey contained 77 items, organized into seven sections that assessed dispositions, an array of motivational beliefs, and various attitudes and behaviours related to time management and procrastination. the week 6 survey consisted of 43 items, organized into four sections, which assessed students’ reported use of various cognitive, metacognitive, and motivational strategies. the week 14 survey was designed as a follow-up that would allow students to consider changes in some of their beliefs and behaviours. this final survey included 24 items from the week 2 survey and 18 items from the week 6 survey, organized into five sections. see table 1 for a full list of the constructs assessed on each survey, as well as appendix a for brief descriptions of these constructs and sample items. table 1 description of surveys note: items on several of the scales were modified slightly in order to better reflect the particular context and/or level of specificity. in some cases, additional original items were also included. srl = self-regulated learning 2.3.1 ier indices students’ responses to each survey were used to compute four indices designed to assess the extent to which they engaged in ier. these indices were the primary measures used to address our research questions. each index was computed for each survey, based on the items within that particular survey. 2.3.1.1 response time students’ total response time for each survey was computed from start (when the “begin” button was clicked) to stop (when the final “submit” button was clicked). this time was automatically and surreptitiously recorded via the qualtrics software. the total response time then, included all the time that participants used to read the directions, the time taken to read, think about, and respond to all of the items, and any time taken to read the report of their scores on the variables measured in the survey. the total response time also included any time in which the students had the survey open and active but were not actively working to complete it (e.g. were distracted). in our effort to identify students who engaged in ier, we considered only excessively short response times. an exact minimal amount of time considered necessary for properly responding to any particular survey is difficult to determine. huang et al. (2012) recommended a cutoff score of two seconds per item as a means of detecting ier. however, unlike surveys examined in much of the recent ier research, our surveys contained not only items and response options, but also directions and descriptions of each variable being assessed. to account for the additional time needed to read and consider all of this material, we elected to use a method of determining a cutoff score based upon expected reading rates. carver (1982) found that the average college student can read and comprehend about 300 words per minute (wpm). based on this rate, we computed the minimal amount of time that a student might be expected to spend completing each survey and used it to establish distinct cutoff scores for each (see table 2). 2.3.1.2 individual response variability (irv) the irv index (dunn et al., 2016) was computed by calculating a standard deviation for all items on a specific survey for each individual participant. the irv index is sensitive to both long-string responses and less obvious forms of ier evidenced by low variability among responses (e.g. answering 4,5,4,5,4,5). values of the irv index could range from zero to 2.5 (the highest maximum standard deviation given our response options). lower values on the irv represent less variability among a participant’s responses and are interpreted as greater engagement in ier. that is, when respondents use a restricted response set like in the example above, their irv index score will be lower and thus indicative of ier. however, no established cutoff score for ier detection exists for this index. we elected to use 2 sds below the mean irv score for all participants as our cutoff criteria. based on the assumption of a normal distribution, this cutoff value is rather conservative, as less than three percent of scores would be expected to fall below this value. 2.3.1.3 psychometric synonyms (ps) the ps index was computed as a respondent’s average bivariate correlation for the set of items found to have the strongest correlations for the total sample (e.g. “i constantly worry about how little time i have for completing my assignments” and “i stress a lot about not having the time i need for my coursework”) . the criteria for identifying the set of item pairs to use in the computation of the ps index is not absolute. some have suggested identifying a particular number of item-pairs with the strongest bivariate correlations (johnson, 2005), while others have suggested using a specific threshold for the magnitude of the correlations used in this computation (|.60|; meade & craig, 2012). given the length of our surveys and the recommendation that items not be repeated when considering what item pairs to use in this computation (curran, 2016), we computed a ps index using the 10 non-repeating pairs on each survey with the strongest, positive bivariate correlations for the sample. the ps index could range from -1 to 1 with lower scores interpreted as increased engagement in ier. however, because there is no consensus as to what value unequivocally indicates that students have engaged in ier we again elected to use 2 sds below the mean value for the sample as a cutoff value. 2.3.1.4 even-odd consistency (e-o) the e-o index (huang et al., 2012; johnson, 2005; meade & craig, 2012) is computed as a participant’s correlation for the set of even and odd item-pairs that assess the same underlying construct (e.g., grit item #1 & grit item #2). the e-o index was computed using 37 item-pairs for the week 2 survey, 21 item-pairs for the week 6 survey, and 19 item-pairs for the week 14 survey. the range of the e-o index is also between -1 and 1, with lower scores interpreted as reflecting greater engagement in ier. although a specific cutoff value for this index has been suggested (.30 by jackson, 1977, as cited in johnson, 2005), this value has not been used consistently in other ier research. thus, to maintain consistency with the irv and ps indices, we used 2 sds below the sample mean for this index as the cutoff value to categorize students as engaging in ier. 3. results table 2 indices, cutoff scores, and number of participants identified as engaging in ier note. the id column represents the number of participants identified as engaging in ier for each method. the total unique id row represents the number of individual students that engaged in ier as indicated by at least one method. it was possible (and happened to be) that a student could be identified as engaging in ier by more than one index. rt = response time. irv = individual response variability. ps = psychometric synonyms. e-o = even-odd consistency 3.1 prevalence of ier our first research question concerned the prevalence of ier exhibited by students who completed each of the three relevant srs. as a first step in addressing this question, table 2 provides the specific cutoff scores used to identify ier, the average score for each index, as well as how many students were identified as engaging in ier for each survey, based on each of the four indices utilized. based on at least one of the indices, 12.8% of students who completed the week 2 survey, 12.5% of students who completed the week 6 survey, and 15.7% of students who completed the week 14 survey were identified as having engaged in ier. overall, out of 827 survey administrations, 15.1% displayed some evidence of ier. 3.2 consistency of ier 3.2.1 index consistency our second research question concerned the consistency of ier across survey administrations. we conceptualized consistency in two ways. first, we considered the consistency of students’ values for each particular index across the three surveys. regarding the consistency of index values across survey administrations, table 3 displays mixed evidence. most apparent, the values of the irv index showed moderate to strong correlation with one another (rs > .43). these correlations indicate that students were somewhat consistent in whether their responses were tightly centred around a particular response option (e.g. the midpoint) or whether they tended to be more varied in the response options they selected across surveys. students’ general tendency to respond to the selected item pairs in a way that was consistent with the overall sample (i.e., the ps index) was neither exceptionally strong nor consistent. table 3 correlations among ier index values note. rt = response time. irv = individual response variability. ps = psychometric synonyms. e-o = even-odd consistency. correlations in bold indicate the relationship between corresponding indices across surveys. * p < .05. ** p < .01. e-o values between any of the surveys were not correlated even though, similar to the ps index, this index is a method of assessing individual response reliability. as well, evidence of consistency in the amount of time it took students to complete each survey was weak and inconsistent. that is, students who exhibited longer (or shorter) response times on one survey were no more or less likely to exhibit longer (or shorter) response times when completing the other surveys. 3.2.2 person consistency as a second method of assessing consistency, we also considered whether the ier behaviour is something that students consistently engage in or is more of a one-off, situational behaviour. we examined this by looking at students who were identified as engaging in ier, by at least one index on one survey, and seeing if they were identified as engaging ier on one of the other survey administrations. of the 297 students in our sample, just one was identified as engaging in ier on all three surveys. further, only 27 students (9%) were identified as engaging in ier on two of the three surveys. of those students who were identified as engaging in ier on more than one srs, 64% were identified by the same index on the surveys in which they were deemed inattentive. these results suggest that engagement in ier is not a consistent behaviour and, therefore, may likely be more dependent upon situational factors, such as fatigue or being distracted. when students do engage in ier repeatedly though, they tend to be inattentive in the same manner in which they had previously been inattentive. 3.3 impact of ier on motivation and srl variables our third and most critical research question concerned whether students’ engagement in ier had an appreciable impact on basic analyses that involved the srl and motivation variables assessed by each survey. we evaluated whether the inclusion (or exclusion) of students categorized as engaging in some form of ier in the sample altered key psychometric and descriptive properties of the motivation and srl variables and would thus alter how this data would be interpreted. further, we compared groups of students categorized as attentive and inattentive for each survey. table 4 differences between the total sample, attentive, and inattentive students on srl variables note. n = week 2 survey, week 6 survey, week 14 survey; srl = self-regulated learning. a significant mean level difference between attentive and inattentive (p < .001). b significant cronbach’s alpha difference between attentive and inattentive (p < .05). 3.3.1 internal consistency of srl scales table 4 displays cronbach’s alpha for each of the substantive measures of srl and motivation across the three surveys for the total sample (both attentive and inattentive students) as well as for attentive and inattentive (i.e. were identified as engaging in ier) students separately. we conducted feldt’s tests (feldt et al., 1987; diedenhofen & musch, 2016) to evaluate the observed differences in the reliability of the srl scales computed for these three groups. most notably, there was no statistical difference in the reliabilities computed for the total and attentive students. that is, the internal consistency of the srl and motivation scales remained essentially the same, regardless of whether inattentive students were or were not included in the computations. in contrast, some differences in the internal consistency of these variables was observed when comparing attentive and inattentive students. on the week 2 survey, compared to inattentive students, reliabilities for the attentive students were significantly higher (p < .05) for four of the eight scales assessed (utility value, time pressure, intentional delay, and procrastination). in contrast, on the week 6 survey, inattentive students had higher internal consistency for one of the four scales assessed (cognitive strategies). on the week 14 survey, attentive students had significantly higher consistency on two of the five scales assessed (time pressure and motivation regulation). 3.3.2 mean response for srl scales table 4 also displays means and standard deviations for the total sample, attentive, and inattentive students for each of the srl and motivation scales on each of the three surveys. based on independent t-tests, we found no statistical differences in the means for any of the srl or motivation scales when comparing total and attentive students (ps > .16). again, the inclusion of inattentive students did not meaningfully change the computation of these basic descriptive statistics. in contrast, there were several observed differences in mean scores for attentive and inattentive students. on the week 2 survey, attentive and inattentive students showed significantly different means (p < .001) on utility value, self-efficacy for srl, and time management (mattentive > minattentive), as well as intentional delay and procrastination (minattentive > mattentive). no significant differences in means were found for the scales on the week 6 survey. on the week 14 survey, inattentive students reported significantly higher time pressure than did attentive students. the observed significant differences between attentive and inattentive students on the srl and motivation scales across all surveys was medium to large, based on computed effect sizes (hedges’ g > .57). < i> 3.3.3 relations among srl scales using fisher’s r to z transformation, we examined differences between the total sample, attentive, and inattentive students in their patterns of correlations among the substantive srl and motivation measures. consistent with the findings on internal consistency and means, there were no significant differences in the correlations between srl variables when comparing total and attentive students (zs > |.65|). put differently, the removal of students who were identified as not providing the necessary cognitive effort needed to complete the srs did not impact the bivariate relations between the srl and motivation variables. again, we did find several differences between attentive and inattentive students regarding their correlations between srl and motivation scales. on the week 2 survey, attentive and inattentive students had significantly different (p < .05) correlations between self-efficacy for srl and grit, time management and procrastination, and grit and intentional delay (stronger correlations for attentive students), as well as between intentional delay and mindset and utility value and time pressure (stronger correlations for inattentive students). on the week 6 survey, inattentive students had significantly stronger correlations between cognitive strategies and environmental management, as well as motivation regulation and metacognitive strategies. on week 14 survey, the only significant different in correlations between attentive and inattentive students was between motivation regulation and environmental management, in which inattentive students displayed a stronger correlation. 4. discussion despite their acknowledged limitations, srs have been and remain, the most common form of assessment used to investigate srl (winne & perry, 2000; wolters & won, 2018). adding to this overall trend, the increasing use of online surveys continues to add to the concerns about the validity of the srs data used to understand srl (karabenick & zusho, 2011; schellings & van hout-wolters, 2011; winne & jamieson-noel, 2003). one major threat to validity that is consistently highlighted when considering these types of assessments is that participants do not engage in the necessary cognitive processes required to provide responses that are a valid representation of their true beliefs, attitudes, and behaviours (duckworth & yeager, 2015; schwartz & oyserman, 2001). respondents that are inattentive or provide insufficient effort when responding to survey items are seen as corrupting influences on the resulting data that ultimately lead to invalid conclusions (huang et al., 2015; curran, 2016). our primary goal was to evaluate whether responses from inattentive students are prevalent enough to degrade the quality of data produced from srs designed to assess key aspects of srl and thus alter any inferences made from it. we pursued this goal by addressing three related questions about college students’ engagement in ier on three srs designed to assess multiple factors associated with srl. in the remainder of this section, we discuss findings with regard to these questions, identify paths for future research, and make suggestions to researchers about ways they can assess the quality of data derived from online srs. 4.1 how common is ier? our first objective was to evaluate the extent to which college students engaged in ier when responding to srs designed to assess motivational and strategic aspects of srl. across three surveys administered over the course of a semester, we found a rate of ier (about 15%) using four common post hoc indices of ier that were similar to prior ier research (meade & craig, 2012; maniaci & roegge, 2014; bowling et al., 2016; desimone & harms, 2017). much of the prior research has been conducted using personality or job-related surveys, typically on instruments with over 100 items. our findings, therefore, support an extension of this work by demonstrating similar levels of engagement in ier when surveys are shorter and designed to assess constructs more central to the study of srl. hence, the theoretical content of the items may not have a strong influence on how likely respondents are to engage in ier. perhaps even more remarkable, this consistency was found even when students were expected to have a greater personal investment in responding to items in a careful and attentive manner. students were reminded repeatedly that the information they provided would be used for a class assignment and that it, in turn, would serve to support their positive growth and development as a student. while we assume that we were working with a more motivated sample, future research should manipulate the personal investment respondents have for responding to items and test whether this leads to differences in ier behaviours. our findings also exposed interesting differences based on the methods used to identify students who engaged in ier. across survey administrations, ier was most likely to be identified based on students’ score on the e-o index, followed closely by their response time. the use of the e-o consistency index, a measure of a respondent’s consistency when answering items on the same scale, is primarily a check on random responding to conceptually similar items. the comparatively high proportion of students identified by this method suggests that the type of ier most common in srl research may be more nefarious than simply speeding through the srs or repeatedly answering with the same response option. this finding highlights the need to use more sophisticated indices to evaluate ier, rather than simply “eye-balling” srs data for overt instances of inattentive responding. of the reactive methods used to identify ier, response time is arguably the most objective measure because it is founded on a less debatable premise. in particular, this method is based on the straightforward assumption that there is some minimum amount of time needed to process written text and provide a response. based on carver's (1982) estimation for college students, and unlike previous studies, we used reading rate to determine an ier cutoff score for response time. it provides the advantage of considering all of the reading that accompanies taking an srs (directions, explanations of items), factors that are unaccounted for when using huang et al.’s (2012) recommendation of two seconds per item. even with the advantages of using reading rate, this method of determining a cutoff score should be viewed as conservative because it does not directly consider the amount of time that a respondent would need to thoughtfully reflect upon and answer the item fewer students were identified as engaging in ier using the ps and irv indices. this may be in part due to the length of the surveys we examined, as well as how the ier cutoff was determined. previous studies that have used the irv index identified a certain percentage of students (10%) with the lowest irv index score to be inattentive (dunn et al., 2016; desimone & harms, 2017). rather than predetermine the exact percent of students who must be engaged in ier, we chose instead to use a cutoff value (-2 sd) that is commonly used as a criterion for identifying extreme outliers in a normal distribution of scores. this criterion may underestimate who should be categorized as engaging in ier. however, as the irv index is a relatively new method of detecting ier, more work needs to be done in order to understand its proper utilization as a method of detecting ier. in particular, evaluation of the best criteria to use for determining which students are engaged in ier would be useful. the length of our surveys made finding items to use in calculating the ps index difficult, especially given the recommendation that items only be used once (curran, 2016). consequently, in order to create a reliable coefficient, we were forced to rely on only the highest, non-repeating item-pair correlations, some of which had somewhat modest correlations. the length and nature of our surveys also made it so that we could not compute a psychometric antonym index, which is typically used in conjunction with the ps index. it may be that these indices are not useful measures of ier when assessing multiple distinct constructs using scales with a relatively low number of items as is typical within the research examining srl. put differently, the ps index may prove more valuable for identifying ier when using scales with larger sets of items to assess cohesive underlying constructs. overall, these findings provide additional support for the recommendation that researchers utilize multiple indices to identify respondents who engage in ier (huang et al., 2012; meade & craig, 2012). within each survey, we found that the various indicators of ier had, at best, weak to moderate positive correlations with one another. further, we found that very few students were identified as engaging in ier by more than one index on any particular survey. in fact, of the 125 students who were identified as being inattentive, only 14 were flagged by two or more indices within a particular survey. overall, these findings are in line with the assumption that inattention or insufficient effort may be manifested in a variety of ways (curran, 2016; desimone & harms, 2017; huang et al., 2012) and therefore any single method of detecting ier will fall short of identifying all the participants who engage in ier. 4.2 how consistent is ier across surveys and time? our second objective was to examine the consistency of students’ engagement in ier across survey administrations. that is, we sought preliminary evidence of whether students’ engagement in ier was more or less stable across the three surveys. as a first check on this issue, we found that the pattern of correlations between the same ier indices across different surveys was inconsistent. for example, the irv indices across survey administrations were moderately correlated, whereas the other indices showed a much lower level of consistency. the most immediate implication of these findings is that students do not engage in certain forms of ier on a consistent basis. as well, this pattern of findings has implications for the potential causes of students’ engagement in ier. if stable individual differences played a dominant role, one would expect that those students who, for example, had very quick response times on the week 2 survey would also display very quick response times on the week 6 and week 14 surveys. in contrast, greater variability suggests that students’ engagement in the various forms of ier may be the result of situational factors that are more likely to change between surveys. further corroboration for the importance of situational influences on inattentive survey behaviour comes from our evaluation of whether particular students or groups of students were more likely to be identified as engaging in ier. the “recidivism rate” of ier was very low; only one student was identified as engaging in ier on each of the three surveys and less than 10% of students were two-time offenders. further, we found that no specific group of students, be those based on sex, ethnicity/race, year in school, or academic probation status, were more likely than another to be identified as engaging in ier. one more general interpretation of our findings, therefore, is that students’ ier behaviour is influenced more strongly by situational features linked to a particular survey rather than by more stable demographic or individual characteristics, such as personality variables (cf. dunn et al., 2016). further, this conclusion suggests that the best index to use in evaluating the presence of ier should be tied to expectations about the situational factors and the type of unwanted behaviour they are likely to promote. our findings regarding the consistency of ier are slightly discordant with bowling et al. (study 1; 2016), who also studied ier consistency across different srs administrations. noteworthy differences in the population studied, the constructs assessed, and the nature of the repeated srs administration make direct comparisons of these studies difficult to reconcile. despite this, both studies suggest that ier consistency is a function of the type of survey administered, the sample answering the survey, and perhaps most importantly, situational factors (e.g. a transient environmental distraction) that influence the attention and effort participants provide when answering survey items. 4.3 does students’ engagement in ier impact conclusions about srl? finally, and most critically, our findings indicate that including data from students who were deemed inattentive during the assessment process did not dramatically alter the results of some basic quantitative analyses. we compared cronbach’s alphas, means, standard deviations, and correlations of the substantive motivation and srl measures in each survey for the total sample (attentive and inattentive students) with those for only the attentive students. across all survey administrations, no statistically significant differences emerged. that is, the inclusion of the data from students identified as inattentive did not appear to corrupt basic analyses computed for the whole sample that are fundamental to studies of motivation and srl. participants who engage in ier, therefore, may add “noise” to the overall set of data (see below) but their presence does not appear to substantially corrupt the overall “signal” when considering these fundamental statistics. despite this overall lack of corruption, however, it would be inaccurate to say that inattentive and attentive students provided equivalent data. we found an array of significant mean-level, reliability, and correlational differences when comparing the attentive and inattentive students. a closer examination of these findings does not expose a simple or obvious pattern. in some instances, attentive students displayed higher means, reliability, and correlation coefficients, whereas in others, this pattern was reversed. attentive students displayed higher values for the more “desirable” motivational and srl constructs, such as self-efficacy for srl and time management, while also displaying lower values for less adaptive constructs such as procrastination. this pattern does not hold for all variables however, as inattentive students displayed higher mean levels of motivation regulation and higher internal consistency in their use of cognitive strategies than did attentive students. in sum, we found clear evidence that the methods we used to evaluate ier identified some students who provided atypical response patterns that resulted in differences in some fundamental descriptive properties of the data. the way in which inattentive responding influences assessment of the underlying constructs, however, was not straightforward and needs further investigation. in spite of the clear response set differences between attentive and inattentive students, our finding that the presence of data contributed by these inattentive students did not substantially degrade the quality of data collected or observed relations when assessing srl should provide some relief to researchers. that is, our findings indicate that interpretations of past srl research based on srs may be relatively sound, in spite of the likelihood that there are instances of ier within the relevant data. further, our findings suggest that elaborate screening techniques, such as latent profile analysis (shukla & konold, 2018) or lengthy infrequency scales (maniaci & rogge, 2014) may not be necessary to use when trying to ensure the validity of students’ self-reported motivation and use of self-regulated learning strategies. 4.4 limitations as with any exploratory research, there are several limitations to this study. the first limitation is the attrition of students over the course of the semester. the context in which our surveys were administered, as part of the coursework for a college elective course, meant that our access to students was dependent upon their continued enrollment and participation in the course. the week 14 survey, administered at the end of the semester, had the fewest number of participants, which is likely due to a decrease in enrollment and participation over the course of the semester. it is possible that we were unable to examine a sub-group of students who are more likely to engage in ier, those students who dropped the course or simply did not take the survey. a related limitation comes from our total sample size, just under 300 students. this modest sample size prevents strong conclusions about the generalizability of our results. rather, our results should be viewed as an initial step to more explicitly evaluate the potential limitations of srs data in srl research. the second limitation relates to the ier indices used. the post hoc nature of this analysis precluded us from using proactive methods of identifying ier, which have been shown to be useful in this line of research (huang et al., 2012; meade & craig, 2012; maniaci & rogge, 2014; huang et al., 2015). the use of response time as an indicator of ier is limited to only identifying respondents who answer too quickly. as of now, there is no agreed upon way to assess whether respondents who take too long to answer a survey are engaging in ier. finally, our cutoff scores (300 wpm for response time and -2 sds for the irv, ps, e-o indices) have not been previously used to identify ier. we chose to utilize these cutoffs with the hope of providing a quick and simple metric to determine ier, without the use of elaborate statistical techniques, so that detecting ier may become a standard part of the data-cleaning process, such as looking for outliers or missing data. it is possible however that our simple cutoff scores reduce the complexity inherent in identifying ier. 4.5 future research and implications for srl researchers our findings point to a number of additional lines of research that should be pursued in order to better understand ier and the conditions under which it may inhibit researchers’ abilities to draw valid conclusions from their data. ultimately, it would be beneficial for researchers to routinely compute and report a small set of easily understood indices, including both proactive and reactive indices, which would provide a ready metric regarding the extent to which ier is an issue within any particular set of data. based on our measures and results, we see the e-o index and response time as easily computable, effective indices for detecting ier in srl-related data sets. additional research is necessary to determine the usefulness of the ps index and irv index in these types of data sets. beyond these indicators of ier, the implication that situational forces seem to play a more prominent role in the extent to which a respondent will engage in ier supports the need for more work to better understand the conditions that lead to more and less attentive responding by participants. along these lines, we also suggest that researchers should provide more detailed descriptions of the conditions under which respondents’ complete srs, such as we have done in this paper (see section 2.2 procedures), so that others can evaluate the extent that ier may be an issue. for example, providing information regarding specific instructions given to respondents, any incentives that respondents have to answer the srs, and whether the srs is completed in the presence of a researcher or without supervision, are several factors that may influence the likelihood of ier being present with collected data. in this way, typical issues such as fatigue, interest in content, distractions, along with more subtle factors such as surveillance and expected feedback can be better investigated for their impact on participants’ response behaviours. it is also worth noting that the work investigating ier needs to encompass both experimental and more “naturalistic” designs. experimental studies in which researchers purposefully manipulate aspects of the assessment process in order to consider their effects on participants’ response behaviours is essential for establishing causal connections. at the same time, studies of ier when participants complete research surveys under more typical and less controlled conditions (e.g., our sample) also are necessary for more ecologically valid conclusions. 5. conclusion although they remain a popular choice for researchers, there are a number of important limitations that threaten the validity of using srs to investigate srl (karabenick & zusho, 2011; schellings & van hout-wolters, 2011; winne & jamieson-noel, 2003). in the present study, we focused on evaluating just one of these critiques, that students’ inattention or insufficient effort while completing the items on srs will substantially reduce the integrity of the resulting data and, therefore, its usefulness for investigating srl. our findings lead to two overall insights regarding this critique. on the one hand, we found evidence that a notable proportion of students engaged in ier and, as a result, produced data with some basic statistical properties that were inconsistent with those produced by students who appeared to complete the surveys more thoughtfully. on the other hand, we also found evidence that the irregular response patterns or “noise” contributed by the students who engaged in ier did not corrupt the data to an extent that basic “signals” or statistical properties were lost or debased. in sum, researchers examining srl should likely consider ier more as a nuisance that should be reduced whenever and in as many ways as possible, rather than as fatal flaw that precludes the use of srs as a viable methodology. keypoints the prevalence of ier across three surveys assessing aspects of srl was about 15% attentive and inattentive students provided data that was significantly different from one another data from inattentive students did not degrade descriptive statistics computed for whole sample separate detection methods identified different students as engaging in ier participants’ engagement in ier was a function of situational influences more than individual differences references azevedo, r., taub, m., & mudrick, n.v. (2018). using multi-channel trace data to infer and foster self-regulated learning between humans and advanced learning technologies. in d. schunk & greene, j.a (eds.), handbook of self-regulation of learning and performance (2nd ed., pp. 254-270). new york, ny: routledge. bandura, a. (2006). guide for constructing self-efficacy scales. in self-efficacy beliefs of adolescents (pp. 307–337). https://doi.org/10.1017/cbo9781107415324.004 beach, d. a. (1989). identifying the random responder.journal of psychology: interdisciplinary and applied, 123(1), 101–103. https://doi.org/10.1080/00223980.1989.10542966 biswas, g., baker, r,, & paquette, l. (2018). data mining methods for assessing self-regulated learning. in d. h. schunk & j. a. greeene (eds.),handbook of self-regulation of learning and performance (2nd ed., pp. 388 403). new york: routledge. bowling, n. a., huang, j. l., bragg, c. b., khazon, s., liu, m., & blackmore, c. e. (2016). who cares and who is careless? insufficient effort responding as a reflection of respondent personality. journal of personality and social psychology, 111(2), 218-229. https://doi.org/10.1037/pspp0000085 butler, d. l. (2002). qualitative approaches to self regulated learning: contributions and challenges. educational psychologist, 37, 59–63. https://doi.org/10.1207/s15326985ep3701 carver, r.p. (1992). reading rate: theory, practice, and practical mplications. journal of reading, 36(2), 84-95. choi, j. n., & moran, s. v. (2010). why not procrastinate? development and validation of a new active procrastination scale why not procrastinate. the journal of social psychology, 149, 37–41. https://doi.org/10.3200/socp.149.2.195-212 cleary, t. j, & callan, g. (2018). assessing self-regulated learning using microanalytic methods. in d. h. schunk & j. a. greene (eds.), handbook of self-regulation of learning and performance (2nd ed.). new york: routledge. cleary, t. j., & zimmerman, b. j. (2004). self-regulation empowerment program: a school-based program to enhance self-regulated and self-motivated cycles of student learning. psychology in the schools, 41(5), 537–550. https://doi.org/10.1002/pits.10177 costa, p. t., & mccrae, r. r. (2008). the revised neo personality inventory (neo-pi-r). in the sage handbook of personality theory and assessment: volume 2 personality measurement and testing (pp. 179–198). london: sage publications ltd. https://doi.org/10.4135/9781849200479.n9 curran, p. g. (2016). methods for the detection of carelessly invalid responses in survey data. journal of experimental social psychology, 66, 4–19. https://doi.org/10.1016/j.jesp.2015.07.006 desimone, j. a., & harms, p. d. (2017). dirty data: the effects of screening respondents who provide low-quality data in survey research. journal of business and psychology, 1–19. https://doi.org/10.1007/s10869-017-9514-9 diedenhofen, b., & musch, j. (2016). cocron : a web interface and r package for the statistical comparison of cronbach ’ s alpha coefficients. international journal of internet science, 11(1), 51–60. http://www.ijis.net/ijis11_1/ijis11_1_diedenhofen_and_musch.pdf dignath, c., & büttner, g. (2008). components of fostering self-regulated learning among students. a meta-analysis on intervention studies at primary and secondary school level. metacognition and learning, 3(3), 231–264. https://doi.org/10.1007/s11409-008-9029-x duckworth, a. l., & quinn, p. d. (2009). development and validation of the short grit scale ( grit – s ). journal of personality assessment, 3891, 166–174. https://doi.org/10.1080/00223890802634290 duckworth, a. l., & yeager, d. s. (2015). measurement matters: assessing personal qualities other than cognitive ability for educational purposes. educational researcher, 44 (4), 237–251. https://doi.org/10.3102/0013189x15584327 dunn, a. m., heggestad, e. d., shanock, l. r., & theilgard, n. (2016). intra-individual response variability as an indicator of insufficient effort responding: comparison to other indicators and relationships with individual differences. journal of business and psychology, 1–17. https://doi.org/10.1007/s10869-016-9479-0 dweck, c. (2012). mindset: how you can fulfil your potential. london: robinson. feldt, l. s., woodruff, d. j., & salih, f. a. (1987). statistical inference for coefficient alpha. applied psychological measurement, 11(1), 93–103. https://doi.org/10.1177/014662168701100107 huang, j. l., curran, p. g., keeney, j., poposki, e. m., & deshon, r. p. (2012). detecting and deterring insufficient effort responding to surveys. journal of business and psychology, 27(1), 99–114. https://doi.org/10.1007/s10869-011-9231-8 huang, j. l., liu, m., & bowling, n. a. (2015). insufficient effort responding: examining an insidious confound in survey data. journal of applied psychology, 100(3), 828–845. https://doi.org/10.1037/a0038510 hulleman, c. s., durik, a. m., schweigert, s. a., & harackiewicz, j. m. (2008). task values, achievement goals, and interest: an integrative analysis. journal of educational psychology, 100(2), 398–416. https://doi.org/10.1037/0022-0663.100.2.398 johnson, j. a. (2005). ascertaining the validity of individual protocols from web-based personality inventories. journal of research in personality, 39, 103–129. https://doi.org/10.1016/j.jrp.2004.09.009 karabenick, s. a., & zusho, a. (2011). examining approaches to research on self-regulated learning: conceptual and methodological considerations. metacognition and learning, 10 (1), 151–163. https://doi.org/10.1007/s11409-015-9137-3 macan, t. h. (1994). time management: test of a process model. journal of applied psychology, 79(3), 381–391. https://doi.org/10.1037/0021-9010.79.3.381 maniaci, m. r., & rogge, r. d. (2014). caring about carelessness: participant inattention and its effects on research. journal of research in personality, 48(1), 61–83. https://doi.org/10.1016/j.jrp.2013.09.008 mccardle, l., & hadwin, a. f. (2015). using multiple, contextualized data sources to measure learners’ perceptions of their self-regulated learning. metacognition and learning, 10(1), 43–75. https://doi.org/10.1007/s11409-014-9132-0 mckibben, w. b., & silvia, p. j. (2017). evaluating the distorting effects of inattentive responding and social desirability on self-report scales in creativity and the arts. journal of creative behavior, 51(1), 57–69. https://doi.org/10.1002/jocb.86 meade, a. w., & craig, s. b. (2012). identifying careless responses in survey data. psychological methods, 17(3), 437–455. https://doi.org/10.1037/a0028085 nichols, d. s., greene, r. l., & schmolck, p. (1989). criteria for assessing inconsistent patterns of item endorsement on the mmpi: rationale, development, and empirical trials. journal of clinical psychology, 45(2), 239–250. https://doi.org/10.1002/1097-4679(198903)45:2<239::aid-jclp2270450210>3.0.co;2-1 panadero, e. (2017). a review of self-regulated learning: six models and four directions for research. frontiers in psychology, 8(apr), 1–28. https://doi.org/10.3389/fpsyg.2017.00422 pajares, f. (2007). motivational role of self-efficacy beliefs in self-regulated learning. in b. j. zimmerman, & d. h. schunk (eds .), motivation and self-regulated learning: theory, research, and applications (pp. 111-140). new york: erlbaum. perry, n. e., hutchinson, l. r., yee, n., & maatta, e. (2018). advances in understanding young children’s self-regulation of learning. in d. h. schunk & j. a. greeene (eds.), handbook of self-regulation of learning and performance (2nd ed.). new york: routledge. pintrich, p. r., smith, d. a. f., garcia, t., & mckeachie, w. j. (1993). reliability and predictive validity of the motivated strategies for learning suestionnaire (mslq). educational and psychological measurement, 53(3), 801–813. https://doi.org/10.1177/0013164493053003024 pintrich, p. r., & zusho, a. (2007). students’ motivation and self-regulated learning in the college classroom. in r. p. perry & j. c. smart (eds.), the scholarship of teaching and learning in higher education: an evidence based perspective (pp. 731–810). new york: springer. schellings, g., & van hout-wolters, b. (2011). measuring strategy use with self-report instruments: theoretical and empirical considerations. metacognition and learning, 6(2), 83–90. https://doi.org/10.1007/s11409-011-9081-9 schunk, d. h., & zimmerman, b. j. (1998). self-regulated learning: from teaching to self-reflective practice . psychological science. guilford press. schwarz, n., & oyserman, d. (2001). asking questions about behaviour. american journal of evaluation, 22(2), 127–160. https://doi.org/10.1177/109821400102200202 shukla, k., & konold, t. (2018). a two-step latent profile method for identifying invalid respondents in self-reported survey data. journal of experimental education, pp. 1–16. https://doi.org/10.1080/00220973.2017.1315713 tuckman, b. w. (1991). the development and concurrent validity of the procrastination scale. educational and psychological measurement, 51(2), 473–480. https://doi.org/10.1177/0013164491512022 usher, e., & schunk, d. h. (2018). social cognitive thoretical perspective of self-regulation. in d. h. schunk & j. a. greene (eds.), handbook of self-regulation of learning and performance (2nd ed.). new york: routledge. winne, p. h., & hadwin, a. f. (2008). the weave of motivation and self-regulated learning. in d. h. schunk & b. j. zimmerman (eds.), motivation and self-regulated learning: theory, research, and applications (pp. 297–314). mahwah, nj: erlbaum associates. winne, p. h., & jamieson-noel, d. (2003). self-regulating studying by objectives for learning: students’ reports compared to a model. contemporary educational psychology, 28 (3), 259–276. https://doi.org/10.1016/s0361-476x(02)00041-3 winne, p. h., & perry, n. e. (2000). measuring self-regulated learning. in handbook of self-regulation (pp. 531–566). elsevier. https://doi.org/10.1016/b978-012109890-2/50045-7 wolters, c. a., & benzon, m. b. (2013). assessing and predicting college students use of strategies for the self-regulation of motivation. journal of experimental education, 81(2), 199–221. https://doi.org/10.1080/00220973.2012.699901 wolters, c. a., & won, s. (2018). validity and the use of self-report questionnaires to assess self-regulated learning. in d. h. schunk & j. a. greeene (eds.), handbook of self-regulation of learning and performance (2nd ed.). new york: routledge. zimmerman, b. j. (2008). investigating self-regulation and motivation: historical background, methodological developments, and future prospects. american educational research journal, 45(1), 166–183. https://doi.org/10.3102/0002831207312909 appendix a items were repeated to create week 14 henritus publication frontline learning research vol.5 no. 3 (2025) 1 28 issn 2295-3159 university students’ emotional states during virtual learning eija henritius1, markku s. hannula1, reito visajaani salonen1, panu erästö 2 & erika löfström1 1 university of helsinki, finland 2aalto university, finland article received 30 december 2023 / article revised 13 march 2025 / accepted 13 may 2025/ available online 26 may 2025 abstract this research examines students' emotional states during a virtual course at a finnish university. the mixed methods study drew on students' experienced emotions, perceived course value, perceived control, and open-ended descriptions related to their emotions. the sample consisted of 85 university students. data were collected at nine measurement points during a half-semester foundation course in statistics. through latent profile analysis (lpa), we identified five distinct learner profiles described as the “average”, “struggling”, “thriving”, “victorious”, and “determined”, and analysed how they differ based on students’ gender, form of course implementation, previous attempts at the same course, and performance. the longitudinal design revealed distinct study experiences amongst the five profiles and pinpointed that most of the challenges took place in the middle of the course. the qualitative analysis of open responses identified different explanations students gave for the changes. the multiple measurement points bring forth emotional fluctuation, which is missed if only pre-post measures are used. the multi-measurement point approach, with the identification of emotion profiles and qualitative accounts on student experiences makes a novel contribution to the field by showing how emotions fluctuate in various profiles. keywords: emotion; control-value theory; latent profile analysis; virtual learning info corresponding author email: eija.henritius@vantaa.fi doi: https://doi.org/10.14786/flr.v13i3.1423 1. introduction the widespread use of digital tools and environments in higher education presents us with new perspectives and forms of learning that benefit from different technologies. the current generations of students have grown up with virtual environments and expect these to be available for learning (tan et al., 2021). by virtual learning in this research, we understand learning that takes place in a completely virtual context or in a blended context, in which virtual and face-to-face learning alternates. researchers studying virtual learning environments have begun to acknowledge the role of student emotions and their study experiences in such learning environments (tan et al., 2021). research in education, neuroscience, and psychology have revealed that emotions are important in learning (seli et al., 2016; tyng et al., 2017; um et al., 2012). in education, there has been specific interest in how emotions influence students’ learning processes, experiences of learning, and achievement (daniels et al., 2009; lu et al., 2023; pekrun & linnenbrink-garcia, 2012). our research adds to this literature by examining the dynamics of emotional experiences in a virtual learning environment and, secondly, examining qualitatively student experiences related to these dynamics. positive emotions have been shown to be associated with higher learning performance; negative emotions with poor learning performance (ge, 2021; jarrell et al., 2017). students who have positive experiences are more likely to re-enrol in virtual courses again (wang & newlin, 2000). students’ experiences and learning-related emotions play an important role in their behavior, performance, motivation, and the quality of their learning. challenges can make students feel fear or anxiety and drop out of a course, while success in an exam may lead to relief and boost the motivation to finish a course (diaz-espinoza, 2017; goetz et al., 2003; pekrun, 2006; pekrun & linnenbrink-garcia, 2014). emotional factors have been discussed as one reason for high drop-out rates in virtual learning (im, 2007; rowe, 2006). at the same time, students' self-efficacy (zajacova et al., 2005) and commitment (human-vogel & rabe, 2015) are related to students’ academic performance and capacity to complete a course successfully (bandura, 1991, 1977). when studying affective and motivational factors of learning in an online mathematics course, self-efficacy appeared to be a significant individual predictor of student achievement (kim et al., 2014). previous studies on students’ emotions in learning have used more variable-oriented (e.g., gender or achievement) than individual-oriented approaches (ganotice et al., 2016; jarrell et al., 2017; tamin et al., 2011) and typically investigated how emotions (e.g., duration, frequency, intensity, and valence) affect student performance. the variability or the fluctuation of emotional states has been less studied (henritius et al. 2019; harley et al., 2016; li et al., 2021). however, emotions may vary during the course, for example, because of failing or succeeding in a task (jarrell et al., 2017). a student may have ups and downs while studying a course and individuals differ in how they cope with failure (jarrell et al., 2017). moreover, repeating an academic course may have emotional implications (lewis, 2020). this research investigates the emotional states of university students during a virtual learning course. the study examines students' self-reported perceived control (self-efficacy), perceived value (commitment), and experienced emotions, identifying five distinct learner profiles. the article commences with an overview of the theoretical framework, emphasizing the control-value theory (cvt), which has been specifically developed to describe academic emotions related to learning (pekrun, 2006). cvt offers a comprehensive and integrated approach to the study of emotions, making it particularly useful in mixed-methods research such as ours (pekrun et al., 2010). this research utilizes a mixed methods approach, integrating both quantitative and qualitative data to provide a comprehensive analysis of these experiences. subsequently, the article discusses the methodology employed in the study, encompassing the mixed methods approach and data collection procedures. the results section elucidates the findings of the latent profile analysis, identifying the five learner profiles and their characteristics. this is followed by a qualitative analysis of students' open-ended responses, elucidating the reasons behind their emotional experiences. 1.1 control-value theory and academic emotions the control-value theory (cvt) is based on attribution and expectancy value approaches, and it identifies students' emotions related to their learning performance (pekrun, 2006; pekrun et al., 2010; pekrun, 2024). according to the theory, two types of cognitive-motivational factors, namely perceived control (e.g., expectations, self-efficacy, perceived competence, and perceived control over outcomes), and perceived value (e.g., perceived importance of success, intrinsic interest, utility value, and attainment value), are the main sources of achievement of emotions (pekrun, 2006; pekrun, 2024). this entails in practice that students experience pleasure when they are interested and confident; pride when they value the outcome and have a sense of control, and anxiety when they feel the outcome is important but do not experience sufficient control to avoid failure. the cvt concept perceived control includes self-efficacy (bandura, 1997), which makes research on self-efficacy highly relevant to cvt. on the other hand, human-vogel & rabes’ (2015) and pillai & williams’ (2004) research on commitment is relevant to values. environmental factors, such as student support and atmosphere, can influence students' evaluations of their control and values, which in turn affects their emotions and academic learning performance. for example, in a supportive learning environment where students receive positive feedback from peers and teachers, their feelings of control and value can be strengthened. this supportive atmosphere can lead to positive emotions, such as joy and pride, which promote better learning performance. conversely, a negative environment with little support can diminish students' sense of control and value, leading to negative emotions such as anxiety and frustration, which can hinder learning performance. cvt highlights the importance of environmental factors in shaping students' evaluations of their control and values, which subsequently influence their emotions and learning achievement (pekrun, 2006). the theory provides an integrative framework for analyzing the antecedents and effects of emotions experienced in achievement and academic settings. emotions are not only an outcome of the process, but they also influence attention, thoughts, actions, and learning performance (pekrun, 2024). empirical research has confirmed that positive activating emotions (e.g. enjoyment) are beneficial for learning, while negative deactivating emotions (e.g. boredom) have detrimental effects (pekrun, 2024). on the other hand, effects of activating negative emotions (e.g. anxiety) and deactivating positive emotions (e.g. relief) suggest more complex relation to performance and the empirical evidence is mixed (pekrun, 2024). emotion theories and empirical evidence suggest that the positive emotions promote the more inductive, bottom-up thinking while the negative emotions would promote the more deductive, top-down thinking (forgas 2008). hence, positive emotions seem to better facilitate creative processes, while the negative emotions would facilitate reliable memory retrieval and performance of routines (pekrun & stephens, 2010). while moderate anxiety may thus help in some cognitive tasks, more intense anxiety seems to be exclusively detrimental for learning – likely because attention is directed towards worries, overloading working memory (ashcraft and krause, 2007; rubinsten and tannock, 2010) furthermore, cvt emphasizes also the multiplicative effects of control and value appraisals on emotions. according to the theory, both control and value influence emotions in a multiplicative manner: positive emotions arise when both factors are high, while negative emotions occur when control and perceived values are low. this means that perceived value moderates the relationship between perceived control and emotions. cvt highlights the mutual reinforcement of control and value: when both factors are present to a sufficient degree, they reinforce each other, intensifying the emotional reaction (pekrun, 2006). 1.2. person oriented approaches in emotion research patterns in individual emotional experiences have been identified in prior research, often using a latent profile analysis (lpa) (e.g., orri, 2017; wang et al., 2021; wang et al., 2023; wytykowska et al., 2022). the person-centered approach offers a holistic and parsimonious way to study affective personality dimensions (orri, 2017). a longitudinal study on affective personality profiles and unique patterns in emotional experiences showed that profiles were consistent over time (orri, 2017). the study showed that women tended to experience slightly stronger negative emotions and slightly weaker positive emotions compared to men. the study identified three latent personality profiles using the affective neuroscience personality scales anps: seeking, caring, playfulness, fear, anger, and sadness. these profiles were consistent across time and genders, with some variations in the intensity of emotions experienced by women and men. associations between profiles and emotion regulation skills measures (e.g., emotional intelligence) offered concurrent validity evidence. another study (wang et al., 2021) identified students' differences in self-efficacy (perceived control) operationalized as students’ confidence in english-as-a-foreign-language skills with a person-centered approach and comparing students in distinct self-efficacy profiles on language proficiency and academic emotions. the results identified three groups representing low, medium, and high self-efficacy levels. students in the low and medium self-efficacy groups showed differences in most measures of academic emotions but not in language proficiency. the third studied university students’ psychological development during the covid-19 outbreak. three profiles were identified as having high, moderate, and low adaptation. the students with high adaptation possessed a more positive self-efficacy (perceived control) belief and demonstrated lower levels of anxiety. in contrast, the students with low adaptation possessed a less positive self-efficacy beliefs and demonstrated higher levels of anxiety (wang et al., 2023). 1.3 the influence of prior attempts, form of course, and gender in the learning experience prior attempts to complete a course, the form of the course, and gender have been identified as influencing the learning experience. students’ previous knowledge of the subject has been found to have a significant effect on learners' achievements (hailikari et al., 2007; 2008). for example, learners with more extensive academic preparation tend to have better academic success (kurlaender & howell, 2012), and conversely, inaccurate knowledge can hinder future development (ambrose, et. al., 2010). on the other hand, research has also shown that there is not always strong evidence for the positive effect of repeated practice intervals on learning. for example, a study by arnold (2017) found that while some students benefit from repeated practice, others do not show substantial improvement, indicating that the effectiveness of repetition can vary depending on individual learning styles and contexts. additionally, repeating an academic course appears to have emotional consequences in the form of grief and loss (lewis, 2020). furthermore, it has been shown that the pass rates of students on their second attempt are significantly lower than of those students enrolled in their first attempt (snead et al., 2022). student performance in a virtual course can be significantly different compared to their performance in a traditional course (schoenfeld-tacher et al., 2001) and the research shows that course design, time management problems, commitment (cvt: perceived value), lack of peers, and familiarity with technologies are the general issues that arise and have an impact on students' emotional experiences in virtual learning (akojie, 2019; allan, 2012; blackmon & major, 2012; howland & moore, 2010; song et al., 2004). furthermore, research has identified that student achievement is related to their satisfaction and correlated with their anger, boredom, and enjoyment (henritius et al. 2019; baturay, 2011; chaparropeláez et al., 2013; kim et al., 2014; lin et al., 2015). yet another aspect that has been identified in prior research as influencing emotions in virtual learning, is gender. research has indicated that female students exhibited stronger self-regulation, which led to their significantly more positive virtual learning outcomes than those experienced by males (alghamdi et al., 2020). however, previous studies have also shown gender differences related to student anxiety (naghavi & redzuan, 2011). female students have more emotional distress, which has negative effects on students' performance (cassady & johnson, 2002; hembree, 1988; chin et al.,2017). 1.4 emotional states during virtual learning research has captured changes in emotions or the relationship between emotions and personally meaningful situations during learning (anwar et al., 2023; d’mello, 2013; strain et al., 2011). however, we found only two studies (hilliard et al., 2019; madsgaard et al., 2022) that have explored emotions in virtual learning environments through multiple measurement points throughout the learning process. emotions are generally studied retrospectively rather than in situ (eteläpelto et al., 2018). in research on students’ self-reported experiences of emotions during simulation-based education (madsgaard et al., 2022), several emotions were identified, including anxiety, engagement, eagerness, derailment, dreadfulness, fear, nervousness, and stress. fear was found to be related to performance pressure and the desire to complete the simulation; dread was related to the pressure to manage technical equipment. students shared the experience of feeling frustrated because of not being as well prepared as they had hoped, or when they felt uncertain about what would happen during the scenario. during the simulation students felt tired, exhausted, and overwhelmed. after the simulation, some mentioned that they had a good feeling because they had managed to handle the situation (madsgaard et al., 2022). in a second study that investigated students’ emotional experiences during an online, collaborative group project (hilliard et al., 2019), self-report data about the experienced emotions and their causes were gathered using a structured diary at six points of time during the group activity. findings revealed that learners experienced a range of pleasant and unpleasant emotions before, during, and after the collaborative activity. pleasant emotions were associated with completing the project (e.g., satisfaction) or working with others (e.g., enjoyment). many unpleasant emotions stemmed from self-beliefs (e.g., anxiety) or students not communicating or participating (frustration, disappointment, and anxiety). a lack of guidance and support predominantly from the course content and the tutor were a source of unpleasant emotion for some learners. collaborating with others, workload, task difficulty, and assessment timing were commonly cited reasons for unpleasant emotions. to understand the variability of emotions, this study examined the change in students' emotions, perceived value of the course, and feeling of control at several measurement points during a virtual course. through open-ended questions, we also gained more detailed information about students' experiences. furthermore, we examined how previous course attempts, gender, and course format were related to students' learning experience and achievements. we used cvt as a framework for understanding students' academic emotions. the following research questions were addressed: 1. how do student experiences change in terms of their perceived control and value of the course, and achievement emotions over the duration of a virtual course? 2. how are students’ perceived control and value of the course, and achievement emotions related to prior course completion attempts, form of course implementation (blended vs. completely virtual), gender, and performance in a virtual course? 3. what reasons do students give for their experienced achievement emotions, perceived control, and value of the course? 2. method 2.1 participants and context the research was conducted in a half-semester (approximately eight weeks) foundation-level statistics course at a large public research university in finland. the virtual course had two forms of course implementation: blended learning and completely virtual learning (named “fast track”). the difference between the two options was a weekly face-to-face group meeting for the blended course participants. the course was intended for degree program students and had 450 participants. eighty-five students volunteered to participate in this study (65 completely virtual and 20 blended learning students). seventy-two of those students participated in the course exam. basic descriptive statistics are presented in table 1. table 1 basic descriptive statistics of background variables for the entire dataset (n=85 students) note. min= minimum value; max= maximum value; mo= mode; mdn=median. note. course credits: 0=0 credits, 1= 1-29 credits, 2= 30-6 credits, 3= 61-120 credits, 4= more than 120 credits. note. one person out of 85 students did not have a finnish matriculation examination. note. best exam results: only those students who have completed the course exam are considered. 2.2 ethical considerations participation in research was voluntary and based on informed consent. consent was indicated on a form. the students’ decisions to participate or not to do so had no influence upon the treatment they received at the university or their grading. students had the right to cancel their participation in the research at any time without notice. the researchers followed the guidelines of the finnish national board on research integrity (tenk, 2019). in finland, an ethics review is required when research involves intervention in the physical integrity of research participants; deviates from the principle of informed consent; involves participants under the age of 15 being studied without parental consent; exposes participants to exceptionally strong stimuli; risks causing long-term mental harm beyond that encountered in normal life; or signifies a security risk to subjects (tenk, 2019, p.19). none of these conditions materialized in this study, and consequently an ethics review was not required. 2.3 procedure students received a short weekly online questionnaire nine times (measurements m0-m8) during the course. the first survey point (m0) was before the beginning of the weekly online tasks and the last was at the end of the course (m8). all material, literature, and general information was published on the online platform including the weekly questionnaires. the researchers only processed anonymized material without individual identifiers. we included only complete data, i.e. those students who had responded to all the surveys. 2.4 measures based on cvt, we used three questions to measure students’ emotions during the course, namely “i have a good feeling about the online work of the course” (cvt: achievement emotion), “i cope well with the course's online tasks” (cvt: perceived control), and “i am committed to completing the course” (cvt: perceived value). for each item, a 5-point likert-type scale (1 = strongly disagree; 5 = strongly agree) was used. we chose to use single items for several reasons. firstly, we aimed to minimize the burden on students by keeping the survey concise, especially given the frequency of the weekly measurements. this approach aligns with the literature that supports the use of single items in psychological research (allen et al., 2022). secondly, while we acknowledge that established and differentiated instruments like the aeq by pekrun et al., (2011) exist for measuring performance emotions, our study's design required a more streamlined approach. the single item "i have a good feeling about the online work of the course" was intended to capture a general sense of the students' emotional state without overwhelming them with lengthy questionnaires. we recognize that this may simplify the measurement of emotions, but it was a necessary compromise to ensure high response rates and participant engagement. additionally, the collected qualitative data provided deeper insights into the students' emotional experiences, complementing the single-item quantitative measures. this mixed-methods approach allowed us to balance the need for brevity with the richness of qualitative descriptions. we have acknowledged the limitations of using single items in our manuscript and discussed how this choice may have influenced the reliability of our findings. we believe that despite these limitations, the single items provided valuable insights into the students' emotional states during the course. the survey also included one open-ended item in which the participants could describe in their own words their emotions and the reasons for these, namely “justify your answer briefly”, which was part of the weekly survey. in addition, we obtained background information regarding gender, form of course implementation, previous attempts and exam scores. 2.5 analyses and tests we applied person-centered latent profile analysis (spurk et al., 2020) on the weekly collected data on student experiences during the course. all data processing and statistical analysis were performed with the rstudio statistical computing software (version 4.4.2; r core team 2020) and the latent profile analysis (lpa) was conducted using the tidylpa package in r (rosenberg et al., 2018). a latent profile represents a subgroup of individuals who share a pattern of responses on a set of variables (lanza, 2016). based on statistical model fit indices (nylund et al., 2007; lo et al., 2001), we selected a model with five profiles. while some fit indices are likely to be inflated because of the repeated measurement, the information is valid for finding the best from alternative models. as a result of the latent profile analysis (lpa) the akaike information criterion (aic) (4790), and bayesian information criterion (bic) (5196) had the smallest values and entropy was 0.987 (akaike, 1987; schwarz, 1978), which can be seen as an indicator of an optimal number of profiles modelled from the data. from one profile model to four profile models, the bic dropped rapidly (6186-5429-5374-5308), furthermore there was a clear drop in bic value when moving from four profiles to five profiles (5308-5196). similarly, the sabic value decreased correspondingly to the bic value (sclove, 1987). from a five-profile model to a six-profile model only one individual was added, and the smallest profile size was considered to be too small (marsh et al., 2009). for five profiles the bootstrap likelihood ratio test (blrt) was statistically significant at p<.01, indicating this to be an appropriate number of profiles. (see table 2). table 2 results of lpa with 5-profile model we chose lpa (latent profile analysis) even though it is not the most typical analysis to use in similar studies it can even be considered a bit naive approach (den teuling et al., 2021). on the other hand, it answers the questions posed, and the analysis converges even in our small sample. moreover, e.g. the growth curves offer a distinctly different perspective compared to the assumption that the points form an overall picture of various respondents. we decided that although a 3-profile model might be statistically stronger, the 5-profile model is qualitatively justified and provides more insights into students' experiences. this model is an important part of our research, even if it does not fully meet statistical criteria. using multiple variables in profile analysis is justified because it gives a more accurate picture of the process and adds information for identifying profiles. however, this also increases the risk of random deviation, making it important to examine profiles qualitatively. furthermore, we chose to use a profile under the 5% threshold due to the nature of our research. our study allows us to capture valuable nuances and variations. this approach provides deeper insights into the phenomena under investigation, which might not be fully captured through a purely quantitative lens. we believe that showcasing this variation is crucial for a comprehensive understanding of the research context, even if it falls below the typical quantitative threshold. this approach aligns with our research objectives and enhances the richness of our findings. to find out statistically significant differences between profiles we conducted chi-square tests to determine whether there were any differences in gender, kruskal-wallis h test for best exam results differences (where the normal distribution of the sample means was not possible), and a one-way analysis of variance (p < .05) in terms of students’ perceived control and value, and emotion differences between profiles followed by a post hoc tukey hsd test. students justified their 5-point likert-type scale answers (concerning emotions, commitment, and self-efficacy) with open-ended answers. the open-ended responses were coded using an inductive approach and systematically collected into a spreadsheet. emotions (such as anger, fear, etc.) and relevant aspects like student motivation were identified separately for each measurement point. the first author coded all data and then took selected coding to be discussed with the second and the last author, who discussed the coding until reaching consensus. we included several data excerpts in the results section to allow the reader to judge our interpretations. in some cases, the justifications could contain additional information. for example, when seeking a justification for the given value of the satisfaction question, the students could also describe their general emotional state or perceived control or value at the same time. the justifications for the answers were analysed in relation to the corresponding question, and in addition, any other comments that the student had made were extracted from the answers. in each profile, the key issues were identified at each measurement point. these key issues included the difficulty of course assignments and student motivation. a summary was then made of the issues most frequently highlighted in each profile, based on cumulative observations from all measurement points. additionally, all emotions that emerged from the data (e.g., anxiety, depression, frustration) were identified at each measurement point. for each identified emotion, student responses were recorded. as all student responses were given in the context of studying this course, we interpreted all emotions as achievement emotions. this was supported by the content of the open responses. from the qualitative analysis, we identified three recurring themes: workload, motivation, and other emotional experiences. each student profile is described through these three themes and chronologically evolving from the first to the last measurement point. 3. results we explored the evolution and variation in students’ achievement emotions, perceived value, and perceived control, identifying five distinct profiles. the responses to the open-ended questions shed light on what the students perceived as the reasons for their experiences. 3.1 profiles identified based on a latent profile analysis, students’ perceived control (pc), perceived value (pv), and achievement emotions experienced (ae) over the duration of a virtual learning course resulted in five profiles. while two of the profiles had very few students (four and nine), the dynamics of their learning experiences seemed to be sufficiently different from other profiles. hence, we decided to include them in the qualitative analysis. in qualitative analysis exceptional cases can be illuminating (donaldson et al., 2013). the profiles with distinct experiential paths were named as follows: the “average”, the “victorious”, the “struggling”, the “thriving”, and the “determined” (see table 3). table 3 the evolution of students' experiences by profile note. ae= achievement emotion, pc=perceived control, pv=perceived value. there was a significant effect of profile on ae, pc, and pv at the p < .05 level. the post hoc tukey hsd test showed significant differences in ae and pc values (p  .001) between profiles in the first half of the course (at the measuring points 2–4), and in pv at the measuring points 4–7 (see additional online resource appendix 1, table 11). in the "determined" profile the values are at their lowest at early measurement points (mp 2 and 3), and in the profiles "average", "struggling" and "victorious" at mp 4 or 5 (see fig.1) (see additional online resource appendix 1, table 11 and appendix 2, table 12). in all profiles, ae and pc values (m and sd) stayed lower than pv during the whole course (see table 4). in the "determined" profile, it appears that a decrease in ae is followed by a decrease in pc, and in the "struggling" profile, pv, ae, and pc decrease or increase in similar patterns (see fig.1). ae and sc values stayed mainly lower in the "determined" profile and in the "struggling" profiles all the values (ae, pc, pv) stayed lower than in the other profiles during the whole course (see fig.1). table 4 students’ achievement emotions, perceived control and values m and sd values note. ae=achievement emotion, pc=perceived control, pv=perceived value. figure 1. student’s experiences in achievement emotion, perceived control, and perceived value by profile. note. values based on weekly surveys results grouped by the mp (measuring point) 0-8. 3.2 prior course completion attempts, form of course implementation, gender, and performance a chi-square test showed that the gender differences between the profiles were statistically significant: χ² (4, 2) = 18.153, p < .001. however, it is worth noting that three cells (30%) had an expected count of less than 5 when the minimum expected count was 1.74. in the “average”, “determined”, “struggling”, and ”victorious” profiles there were more female than male students (see fig.2, table 5). figure 2. gender distribution (n=85 students) by profile. note. profiles: 1: “average” (n=26), 2: “struggling” (n=9), 3: “thriving” (n=21), 4: “victorious” (n=25), and 5: “determined” (n=4). furthermore, there were differences between the profiles based on form of course implementation (blended, fast track): χ² (4, 2) = 11.755, p < .05. in every profile, more students chose the blended form of course implementation than fast track. in the “thriving” (43%) and “victorious” (32%) profiles, the proportion of students who had chosen fast track was the highest (see fig.3, table 5). however, the results should be interpreted with caution as four cells (40%) had an expected count of less than 5 when the minimum expected count was .94. figure 3. form of course implementation (n=85 students) by profile. note. profiles: 1: “average” (n=26), 2: “struggling” (n=9), 3: “thriving” (n=21), 4: “victorious” (n=25), and 5: “determined” (n=4). in the “struggling” profile, there were more students that had previously attempted the course without an approved grade than in the other profiles (see fig. 4, table 5). the relationship between the number of previous attempts and the profiles could not be determined, as the number of students who had previous attempts at the course was very small (0-4 students per profile). figure 4. previous attempts (n=85 students) by profile. note profiles: 1: “average” (n=26), 2: “struggling” (n=9), 3: “thriving” (n=21), 4: “victorious” (n=25), and 5: “determined” (n=4). when comparing the kurskal-wallis h test’ best exam results of different profiles, we first removed from the data 13 students who had not taken the exam. the main test identified significant differences between profiles (p < .05), but none of the pairwise comparisons had significant differences. the students in the “victorious” and “thriving” profiles had the best exam results (see fig. 5, table 5). figure 5. the average exam grade points (n=72 students) by profile. note students who did not participate in the exam have been removed. note profiles: 1: “average” (n=26), 2: “struggling” (n=9), 3: “thriving” (n=21), 4: “victorious” (n=25), and 5: “determined” (n=4). table 5 summary: the distribution of gender, previous attempts, performance, results, and form of implementation 3.3 students’ self-reported emotions in each weekly survey, students were asked to indicate the reasons for the emotions they experienced “i have a good feeling about the online work of the course” (cvt: achievement emotion), “i cope well with the course's online tasks” (cvt: perceived control), and “i am committed to completing the course” (cvt: perceived value). from the qualitative analysis, we identified common themes: perceived feeling of challenge (e.g., with weekly tasks and related materials), workload / time management, motivation , and other emotional experiences. each student profile is described through these themes and chronologically evolved from the first to the last measurement point (1-8). the initial survey before the start of the course (mp 0) was excluded from open-ended results. in the first profile, the “average”, all the values (achievement emotions, perceived control and perceived value) remained at an average level throughout the course and were at their lowest at measuring point five (mp 5) (see fig. 6). figure 6. evolution of students’ experiences (scores 1-5) in the “average” profile in different measurement points (mp 0-8). commonly, in open-ended responses, students in this profile described their perceived feeling of challenge to be easier at the beginning (mp 1-2) and increasing in challenge in the middle of the course (mp 4-5). students mentioned challenges concerning time management and weekly tasks. furthermore, in the profile, low motivation and negative emotions dominated over the high and positive ones in terms of occurrence in the data (see table 6). students’ motivation and emotions were mainly related to weekly tasks. twenty-six students (96%) verbally described their views in open-ended questions at least at one measurement point. table 6 profile the “average”: identified themes in students’ open-ended descriptions note. profile: the “average” (n=26). students described their perceived feeling of challenge e.g., regarding time management and difficulty of weekly task package: ”i haven't spent enough time doing the tasks” (mp 4);“the package was insanely difficult, and the material was not very helpful” (mp 5); “hopefully things won’t get any harder, otherwise i’ll be lost” (mp 5), and “it’s not very rewarding to do tasks every week on a sunday night just before the assignments close” (mp 7). students reported their perceived level of motivation e.g., regarding weekly task package: "the tasks are going well; the course task system feels motivating rather than punitive. the materials are of high quality" (mp 1); “p4 [weekly task package 4] really hard and lowers motivation” (mp 4);” one weaker week, and motivation declined”;” this week’s tasks were very difficult, and i didn’t understand many things, which reduced my motivation” (mp 5), and “online assignments were difficult, which reduced my level of motivation” (mp 6). students mentioned various emotions e.g., regarding exam and weekly task package: “the exam is coming, and i don’t think i’ve internalized the course matters at any level necessary to succeed in the exam” and “more difficult than the previous packages, begins to distress” and “more difficult than the previous packages, begins to distress” (mp 4) and further at the end of the course: “i invest more in the course, so i don’t stress as much”; ”it feels like everything essential has to be studied online, which is frustrating”; “the assignments made me feel good”; “i am pretty depressed, and i don't get anything done” (mp 7), and “towards the end, my own enthusiasm may have diminished a little. i'm a little disappointed with that” (mp 8). in the second profile, the “struggling”, students’ all values decreased until the fifth measuring point and increased slightly before the end of the course (mp 7–8) (see fig. 7). figure 7. evolution of students’ experiences (scores 1-5) in the “struggling” profile in different measuring points (mp 0-8). commonly, in open-ended responses students in this profile described perceived time management challenges in relation to the workload. they also mentioned challenges that they had faced but overcome in the weekly assignments. the most challenges were experienced at the fifth measurement point. furthermore, in this profile, motivation was high, and emotions negative in the beginning, perceived motivation was lower in the middle and began to improve towards the end in terms of occurrence in the data (see table 7). students’ motivation and emotions were mainly related to weekly task difficulties and the upcoming exam. eight students (89%) described their views verbally at least at one measuring point. table 7 profile the “struggling”: identified themes in students’ open-ended descriptions of their feelings. note. profile: the “struggling” (n=9) note. the table shows the number of students who have given an answer under the theme at each measurement point note. the frequency (f) indicates how many students have given any answer related to the theme overall students described their perceived feeling of challenge e.g., regarding weekly task package and need of help: “i was successful with the tasks, but the time limits put pressure” (mp 2); “i got most of it done on my own, but i need help“ (mp 5) and “i don't cope very well” (mp 5). students reported their perceived level of motivation e.g. regarding weekly task package: "i was quite successful with this week's tasks, and therefore i gained motivation from the package" (mp 1); "if the topics remain equally interesting and not too difficult, i am motivated. however, if the topics become significantly more challenging or less interesting, my motivation can easily decrease”; ”the grade i get from the assignments motivates me to improve my performance or maintain a good average” (mp 2) and further in the middle of the course: “motivation is still quite poor" (mp4); “motivation is weak” (mp5); and finally in the end of the course: “motivation is lost “(mp6); "i gained energy and motivation from the well-executed package”; “motivation completely lost”, and “i want to study for the exam and get a good one, and i am motivated at the moment” (mp 8). students felt various emotions e.g., regarding exam, weekly tasks: “now it feels good to work, but in the back of my head, there is a thought: when you can no longer keep up”; “in the [beginning], i cursed and got frustrated” (mp 1); “the exam is scary” (mp 2); and in the end of the course: "overall, it has gone quite well on average, and the assignments left a good feeling."; and “frustrating and tired of the rigidity of the electronic system”, and "the exam is frightening" (mp 8). in the third profile, the “thriving”, students’ all values stayed high during the course (see fig. 8). figure 8. evolution of students’ experiences (scores 1-5) in the “thriving” profile in different measuring points (mp 0-8). commonly, in open-ended responses students in this profile did not perceive any major inconveniences; they experienced the course as interesting, meaningful, or useful, and they had good experiences studying online: “the things in the course are interesting and online work is right for me” (mp 1); “important things for working life” (mp 1), and “the tasks are convenient to answer online, and the materials are good” (mp 2). students did not raise issues related to motivation or emotions much during the course (see table 8). fourteen students (67%) described their views verbally at least at one measuring point. table 8 profile the “thriving”: identified themes in students’ open-ended descriptions to their feelings note. profiles: the “thriving” (n=21). note. the table shows the number of students who have given an answer under the theme at each measurement point. note. the frequency (f) indicates how many students have given any answer related to the theme overall. in the fourth profile, the “victorious”, students’ perceived value stayed high, and both achievement emotion and perceived control slightly lower throughout the whole course (see fig. 9). figure 9. evolution of students’ experiences (scores 1-5) in the “victorious” profile in different measuring points (mp 0-8). commonly, in open-ended responses, students in this profile, experienced no major challenges during the course except for time management challenges in the middle of the course. for the students, the course was interesting, meaningful, or useful and they had good experiences studying online, just as in the profile “thriving”. furthermore, in this profile, some students mentioned high motivation and negative emotions in the beginning and more positive emotions in the end of the course in terms of occurrence in the data (see table 9). at the beginning, some students were especially upset about the rigorous assessment practices and strict evaluation practices. nineteen students (76%) described their views verbally at least at one measuring point. table 9 profile the “victorious”: identified themes in students’ open-ended descriptions to their feelings note. profile: the “victorious” (n=25). note. the table shows the number of students who have given an answer under the theme at each measurement point. note. the frequency (f) indicates how many students have given any answer related to the theme overall. students described their perceived feeling of challenge e.g., regarding to weekly task package, workload / time management and online work: “the time for weekly exercises was short” (mp 2); “nice to do tasks, but it takes a lot of time”; “online work is very good for my busy schedule” (mp 3) and “online work is a good way to learn these things when one has to practice” (mp 3), and “useful information that i will definitely need in the future” (mp 4). students reported their perceived level of motivation e.g., regarding completing the course and time management: “the topics of the course are important, which further motivates you to complete the course well”; “the motivation to complete the course well is high”; "i am ahead of schedule, which increases my motivation and reduces stress about the course" (mp 2), and “the course is a compulsory part of the degree, and i have already spent a considerable amount of my free time on it. i am indeed motivated to complete the course and committed to it” (mp 3). students felt various emotions to e.g., regarding the penalty points for the weekly tasks: ”it got a little annoying when careless mistakes were sanctioned so harshly” (mp 2). "it is a bit frustrating how unforgiving the review process is; typos and similar errors are penalized too harshly" (mp 3); “good feeling when i got through!” (mp 5) and at the end of the course; “good feelings excluding harsh minus points for task pack delay” (mp 7); “otherwise, a good feeling, but the past two worst weeks are annoying because they decrease total points” (mp 8). in the fifth profile, the “determined”, students’ perceived value was very high throughout the course when the achievement emotion and perceived control stayed on an average level and close to each other. the perceived control decreased more sharply at the beginning (mp 2) and remained near average throughout the rest of the course. the emotions decreased more slowly and increased higher than perceived control towards the end of the course (mp 6) (see fig. 10). figure 10. evolution of students’ experiences (scores 1-5) in the “determined” profile in different measuring points (mp 0-8). commonly, students pushed through even though they had a hard time. some students reported that one week of good performance influenced their motivation to complete the course. during the course, the students experienced periods in which they had no difficulty and, on the other hand, periods in which they experienced difficulties. furthermore, in this profile low motivation dominated in the end of the course and negative emotions appeared in measuring points three and six in terms of occurrence in the data (see table 10). students’ motivation was related to weekly tasks and emotion to exam and quizzes. three students (75%) described their views verbally at least at one measuring point. table 10 profile “determined”: identified themes in students’ open-ended descriptions to their feelings (n=4) note. profile: the “determined” (n=4) note. the table shows the number of students who have given an answer under the theme at each measurement point. note. the frequency (f) indicates how many students have given any answer related to the theme overall. students described their perceived feeling of challenge e.g., regarding difficulties with weekly tasks: “a lot goes wrong” (mp 3); “i haven't been able to do b-tasks properly and it worries me” (mp 4) and “the examples are not given in advance, so everything must be applied on the basis of theory. the points obtained for the tasks have been declining” (mp 5). students reported their perceived level of motivation regarding weekly tasks: “this week's tasks went better, which increases motivation” (mp 4); "completing task sets takes a lot of time at once and concentration wanes. doing them is very demotivating, so they are not always done as well as they could be” (mp 6), and “requires far too much independent learning, completing the course is very unmotivating when they don't actually teach how to calculate tasks” (mp 8). students felt some emotions e.g., regarding exam and quizzes: “some of the tasks feel awkward and i don’t know how to cope with the exam” (mp 3); “i would learn things much better through worked-out examples. quizzes are mostly annoying and even though they prepare for the exam, they don’t seem to help me to internalize things through them” (mp 6). 4. discussion this study examined university students’ changing emotional states during a virtual foundation course in statistics. we adopted a person-centered approach and used student profiles of perceived control, perceived value, and achievement emotions to identify how these profiles differed in terms of students’ performance, previous attempts, gender, and form of course. this approach aligns with previous research by orri (2017), which also utilized a person-centered approach to study profiles and emotion regulation. we identified five distinct profiles of student experiences during the course, namely the “average”, the “determined”, the “struggling”, the “thriving”, and the “victorious” profiles. with person-centered profiles we were able to explore the association between the variables of interest at the level of a typical individual in each profile. for example, the “average” students had motivational and time management challenges and a characteristic of the “determined” profile was that students pushed through the whole course despite their challenges. statistically significant differences between profiles were found in the first half of the course with achievement emotions and perceived control. during the latter half of the course, however, profiles had significant differences in their perceived course value (see additional online resource appendix 1, table 11). perceived control, and perceived value stayed high throughout the course among students who found the course interesting or useful and who mentioned that they liked virtual learning as a learning method. with several measurement points we were able to distinguish variation during the course, which would easily go unnoticed unless specifically surveyed. indeed, most of the challenges were experienced in the middle of the course when the weekly tasks got more challenging, which couldn’t be detected if there were only measurement points at the beginning and the end of the course. all the students who valued the course had more stability in their emotions and overall, more positive emotions and a better course performance. students experienced joy and pride when they valued the course outcome and felt in control. they felt enjoyment when they were interested and confident. "most importantly, when students had problems with a weekly task, changes in their emotional reactions and perceived control sometimes were misaligned. for example, the “victorious” students reached the lowest point in their emotions at measurement point four, yet their perceived control continued to decline until measurement point five. moreover, the “average” students’ perceived control and value showed a declining trend at the third measurement point, while their emotions value showed decline only at the measurement point four. this lack of synchrony in these changes does not align with the assumed model in cvt, where emotion is seen as an outcome of values and control. further longitudinal studies should examine with more robust measurements whether this observation of non-synchrony can be repeated. secondly, we found significant differences according to gender, and differences according to form of course implementation. in the “average” (62%), the “determined” (100%), and the “struggling” (67%) profiles, the females were the majority with lower outcomes (see table 5). the result corresponds to previous findings indicating that female students have a higher prevalence for emotional distress, which often can have a negative impact on exam performance (cassady & johnson, 2002; hembree, 1988). females have been shown to exhibit higher levels of test anxiety leading to lower course grades (chin et al., 2017). at the same time, there were more men in highly successful profiles, namely the “thriving” (81%) and “victorious” (72%) profiles. both of these profiles had more fast-track (totally virtual learning) students. prior research suggests that students who do not perceive challenges in their learning or performance are more prone to prefer virtual courses than their peers who are struggling with learning and prefer face-to-face courses (mclaren, 2004; paden 2006; roach, 2002, schoenfeld-tacher et al., 2001). thirdly, we were interested in the reasons students gave for the emotions experienced, perceived control, and perceived value in virtual learning. we found that students in the “average” and “struggling” profiles lost their sense of a good feeling and had less perceived control and value when they experienced lack of time, and the tasks became more difficult. in the “average” profile, some students felt anxious or disappointed regarding their success with the weekly tasks or exam and among the “struggling” profile some students felt tired and frustrated and scared about the exam. yet, they also had good feelings when successfully completing the tasks. in the “determined” profile students kept their perceived value throughout the whole course as they pushed through despite difficulties. similarly, task and time management problems and perceived value (commitment) are known to have an impact on students' emotional experiences in virtual learning (akojie, 2019; allan, 2012; blackmon & major, 2012; hilliard et al., 2019; howland & moore, 2010; song et al., 2004). fear has been related to performance pressure and the desire to complete the task while students also have been proud of themselves and have had a good feeling after handling the task well (madsgaard et al., 2022). students in the “thriving”, and “victorious” profiles had high perceived control and values, and good feeling, throughout the course. they enjoyed virtual studying and found the course interesting, meaningful and useful. an interesting difference between these profiles was that although both the “thriving” and “victorious” mainly had good feelings and experiences throughout the course, the “victorious” were at the same time upset about the rigorous or strict practices, harsh minus points, or task delays. task delays caused time pressure, and similar experiences were found in connection with time management in a previous study in which students felt frustrated because of not being as well prepared as they had hoped (madsgaard et al., 2022). students in the “thriving”' and the “victorious” profiles did not experience difficulties studying; they even enjoyed the course, had good motivation, and experienced success. they had high perceived control (e.g., self-efficacy, perceived competence) and high perceived value (e.g., importance of success, intrinsic interest. students’ difficulties with motivation were related either to performance with weekly tasks (the “average” and “determined” profiles) or to course success (the “struggling”). these findings are in line with cvt (pekrun, 2006); interested and confident students feel they value the outcome and have perceived control, and anxiety when they feel the outcome is important, but they don’t have enough perceived control to avoid failure. we regard exposing the fluctuation of emotions as a key contribution of this research. this fluctuation gets easily overlooked when emotions are investigated once or twice during a learning process. 4.1 limitations the research included only one course, in which case some of the results may be specific to this course. however, such basic courses in statistics are common at many universities. the data came from a small convenience sample with approximately 20% of the students in the course participating in the research. additionally, 13 students who did not attend the exam were removed from the exam data set. some of the profiles contained a small number of students, which impeded statistical analyses. the results cannot be generalised. nevertheless, the results indicate variation among students’ experiences over the duration of a course. further person-oriented research with bigger samples is needed to identify typical emotional-motivational trajectories and student profiles. furthermore, we acknowledge the limitations of using single items for our measurements, which compromises their reliability. however, this shortcoming is counterweighted by using nine measurement points and including qualitative data for more interpretative analysis. one limitation of this study is the potential inflation of fit indices due to repeated measurements. this phenomenon occurs because each data point is treated as independent, despite being part of repeated measures, which can lead to an overly optimistic assessment of model fit. in the context of latent profile analysis (lpa), this inflation complicates the accurate evaluation of the quality of the identified profiles. the use of multiple variables in profile analysis further complicates the interpretation of fit indices. the small sample size also poses a challenge, as it may not allow for a proper cluster analysis. to mitigate these issues, we have included a detailed explanation of the potential inflation of entropy and its impact on our findings. we have also provided the analysis code as an appendix (see appendix 3) to ensure transparency and allow for verification. furthermore, we have emphasized the qualitative differences that support the selection of the five-profile model over the three-profile model, despite the potential inflation of fit indices. our study did not utilize latent class growth analysis (lcga), which can provide a more sophisticated approach for analysing growth trajectories over time (sözer & kahraman, 2023). instead, we opted for latent profile analysis (lpa) due to our focus on situational emotions and the combination of quantitative and qualitative data. furthermore, we were unable to include all fit indices because the statistical package we used, tidylpa, does not support some specific indices. consequently, we recommend that future research consider employing lcga and are willing to provide the data to researchers upon request for further or additional analyses. we have checked the replication of the best log likelihood in our latent profile analysis. specifically, we tested whether the order of the data affects the results and found that it does not. this indicates that the best fit was established consistently, regardless of the data order. however, we acknowledge that our statistical package, tidylpa, has limitations in handling multiple sets of random start values. therefore, we recommend future research to consider using more advanced statistical methods to ensure the robustness of the findings. 5. conclusions this study has highlighted the need to consider a temporal dimension in emotion research. the aim of the study was to provide a better understanding of university students’ dynamically changing academic emotions, perceived control, and perceived value with a person-centered and individual-centered approach. with the help of the profile analysis, we were able to identify different types of students’ experiences in a virtual course. we were able to focus on the fluctuation of students' emotions; something, which alongside a person-centered longitudinal approach has been researched to a lesser extent. using several frequent measurement points, we were able to identify how changes in emotions in some profiles tended to precede changes in perceived control. consequently, emotions can be an important indicator of learning, and above all, of a potential trajectory. students' challenges during the course may easily stay invisible and the same challenges can be caused by various reasons. the teacher is better able to direct right and timely support to students when they understand how emotions affect learning, and that the role of emotions is different at different points in the course. it is also important to understand students' experiences of virtual learning for the benefit of higher education, because for example, students who have experienced positive experiences and emotions are more likely to re-register for virtual courses (wang & newlin, 2000). it can also reveal effective practices, student perceptions, and areas of satisfaction for virtual learning course designers. because the study included 85 students and some profiles included a few students, future research would benefit from larger sample sizes to determine how common similar findings are in online courses. future research may also benefit from comparing emotional fluctuations in virtual/blended and face-to-face learning environments. keypoints multiple measurement points revealed fluctuation in students’ emotions and variation in their online learning experiences, which would have gone undetected with fewer measurement points. five distinct emotion profiles were identified, namely the “average”, “struggling”, “thriving”, “victorious”, and “determined”. achievement emotion and perceived control tended to follow similar patterns whereas perceived value was at times independent of these. a qualitative analysis of the open-ended questions revealed that time management was often related to challenges reflected in perceived control and achievement emotions. references akaike, h. (1987). factor analysis and aic. psychometrika, 52 , 317–332. https://doi.org/10.1007/bf02294359 akojie, p., entrekin, f., bacon, d., & kanai, t. (2019). qualitative meta-data analysis: perceptions and experiences of online doctoral students. american journal of qualitative research, 3(1), 117–135. https://doi.org/10.29333/ajqr/5814 alghamdi, a., karpinski, a. c., lepp, a., & barkley, j. (2020). online and face-to-face classroom multitasking and academic performance: moderated mediation with self-efficacy for self-regulated learning and gender. computers in human behavior , 102, 214–222. https://doi.org/10.1016/j.chb.2019.08.018 allan, b. (2007). time to learn? e-learners' experiences of time in virtual learning communities. management learning, 38(5), 557–572. https://doi.org/10.1177/1350507607083207 allen, m. s., iliescu, d., & greiff, s. (2022). single item measures in psychological science: a call to action. european journal of psychological assessment, 38 (1), 1–5. https://doi.org/10.1027/1015-5759/a000699 ambrose, s. a., bridges, m. w., dipietro, m., lovett, m. c., & norman, m. k. (2010). how learning works: seven research-based principles for smart teaching . john wiley & sons. anwar, a., rehman, i. u., nasralla, m. m., khattak, s. b. a., & khilji, n. (2023). emotions matter: a systematic review and meta-analysis of the detection and classification of students’ emotions in stem during online learning. education sciences, 13(9), 914. https://doi.org/10.3390/educsci13090914 arnold, i. (2017). resitting or compensating a failed examination: does it affect subsequent results? assessment & evaluation in higher education , 42(7), 1103–1117. https://doi.org/10.1080/02602938.2016.1233520 ashcraft, m. h. & krause, j. a. (2007). working memory, math performance, and math anxiety. psychonomic bulletin & review 14 , 243–248. https://doi.org/10.3758/bf03194059 bandura, a. (1991). social cognitive theory of self-regulation. organizational behavior and human decision processes , 50(2), 248–287. https://doi.org/10.1016/0749-5978(91)90022-l bandura a. (1997). self-efficacy: the exercise of control.freeman baturay, m. h. (2011). relationships among sense of classroom community, perceived cognitive learning and satisfaction of students at an e-learning course. interactive learning environments, 19(5), 563–575. https://doi.org/10.1080/10494821003644029 blackmon, s. j., & major, c. (2012). student experiences in online courses. a qualitative research synthesis. quarterly review of distance education , 13(2). https://www.cu.edu/doc/student-experiences-online-classesqual-study.pdf cassady, j. c., & johnson, r. e. (2002). cognitive test anxiety and academic performance. contemporary educational psychology, 27 (2), 270–295. https://doi.org/10.1006/ceps.2001.1094 chaparro-peláez, j., iglesias-pradas, s., pascual-miguel, f. j., & hernández-garcía, á. (2013). factors affecting perceived learning of engineering students in problem based learning supported by business simulation. interactive learning environments, 21(3), 244–262. https://doi.org/10.1080/10494820.2011.554181 chin, e. c., williams, m. w., taylor, j. e., & harvey, s. t. (2017). the influence of negative affect on test anxiety and academic performance: an examination of the tripartite model of emotions. learning and individual differences , 54, 1–8. https://doi.org/10.1016/j.lindif.2017.01.002 daniels, l. m., stupnisky, r. h., pekrun, r., haynes, t. l., perry, r. p., & newall, n. e. (2009). a longitudinal analysis of achievement goals: from affective antecedents to emotional effects and achievement outcomes. journal of educational psychology, 101(4), 948. https://doi.org/10.1037/a0016096 diaz-espinoza, z. j. (2017). i’m here for a reason: motivational factors of first-generation latino males to attend college. university of tennessee. https://trace.tennessee.edu/cgi/viewcontent.cgi?article=5669&context=utk_graddiss den teuling, n. g. p., pauws, s. c., & van den heuvel, e. r. (2021). a comparison of methods for clustering longitudinal data with slowly changing trends. communications in statistics simulation and computation, 52(3), 621–648. https://doi.org/10.1080/03610918.2020.1861464 d'mello, s. (2013). a selective meta-analysis on the relative incidence of discrete affective states during learning with technology. journal of educational psychology , 105(4), 1082. https://doi.org/10.1037/a0032674 donaldson, j. a., ching, j. m. r., & tan, a. t. m. (2013). going extreme: systematically selecting extreme cases for study through qualitative methods. paper presented at the american political science association annual meeting 2013, august 29 september 1, chicago . available at: https://ink.library.smu.edu.sg/soss_research/1509 eteläpelto, a., kykyri, v. l., penttonen, m., hökkä, p., paloniemi, s., vähäsantanen, k., ... & lappalainen, v. (2018). a multi-componential methodology for exploring emotions in learning: using self-reports, behaviour registration, and physiological indicators as complementary data. frontline learning research, 6(3). https://doi.org/10.14786/flr.v6i3.379 forgas, j.p. (2008). affect and cognition. perspectives on psychological science, 3 (2), 94–101. https://doi.org/10.1111/j.1745-6916.2008.00067.x ganotice jr, f. a., datu, j. a. d., & king, r. b. (2016). which emotional profiles exhibit the best learning outcomes? a person-centered analysis of students’ academic emotions. school psychology international , 37(5), 498–518. https://doi.org/10.1177/0143034316660147 ge, x. (2021). emotion matters for academic success: implications of the article by jarrell, harley, lajoie, and naismith (2017) for creating nurturing and supportive learning environments to help students manage their emotions. educational technology research and development, 69(1), 67–70. https://doi.org/10.1007/s11423-020-09925-8 goetz, t., zirngibl, a., pekrun, r., & hall, n. (2003). emotions, learning and achievement from an educational-psychological perspective. in p. mayring & c. von rhoeneck (eds.). learning emotions: the influence of affective factors on classroom learning (pp. 9–28). peter lang. hailikari, t., nevgi, a., & komulainen, e. (2008). academic self‐beliefs and prior knowledge as predictors of student achievement in mathematics: a structural model. educational psychology, 28 (1), 59–71. https://doi.org/10.1080/01443410701413753 hailikari, t., nevgi, a., & lindblom-ylänne, s. (2007). exploring alternative ways of assessing prior knowledge, its components and their relation to student achievement: a mathematics based case study. studies in educational evaluation , 33(3-4), 320–337. https://doi.org/10.1016/j.stueduc.2007.07.007 harley, j. m. (2016). measuring emotions: a survey of cutting edge methodologies used in computer-based learning environment research. emotions, technology, design, and learning , 89–114. https://doi.org/10.1016/b978-0-12-801856-9.00005-0 hembree, r. (1988). correlates, causes, effects, and treatment of test anxiety. review of educational research, 58(1), 47–77. https://doi.org/10.3102/00346543058001047 henritius, e., löfström, e., & hannula, m. s. (2019). university students’ emotions in virtual learning: a review of empirical research in the 21st century. british journal of educational technology, 50 (1), 80-100. https://doi.org/10.1111/bjet.12699 hilliard, j., kear, k., donelan, h., & heaney, c. (2019). exploring the emotions of distance learning students in an assessed, online, collaborative project. in eden conference proceedings (no. 1, pp. 260–268). https://www.ceeol.com/search/article-detail?id=846956 howland, j. l., & moore, j. l. (2002). student perceptions as distance learners in internet-based courses. distance education, 23 (2), 183–195. https://doi.org/10.1080/0158791022000009196 human-vogel, s., & rabe, p. (2015). measuring self-differentiation and academic commitment in university students: a case study of education and engineering students. south african journal of psychology, 45 (1), 60–70. http://dx.doi.org/10.1177/0081246314548808 im, y-w. (2007). a substantial study on the relationship between students' variables and dropout in cyber university. journal of the korean association of information education, 11 (2), 205–219. https://koreascience.kr/article/jako200732056742348.page jarrell, a., harley, j.m., lajoie, s., & naismith, l. (2017). success, failure and emotions: examining the relationship between performance feedback and emotions in diagnostic reasoning. education technology research and development, 65, 1263–1284. https://doi.org/10.1007/s11423-017-9521-6 kim, c., & hodges, c. b. (2012). effects of an emotion control treatment on academic emotions, motivation and achievement in an online mathematics course. instructional science, 40, 173–192. https://doi.org/10.1007/s11251-011-9165-6 kim, c., park, s. w., & cozart, j. (2014). affective and motivational factors of learning in online mathematics courses. british journal of educational technology , 45(1), 171–185. https://doi.org/10.1111/j.1467-8535.2012.01382.x kurlaender, m., & howell, j. s. (2012). academic preparation for college: evidence on the importance of academic rigor in high school. background paper of thecollege board advocacy & policy center. https://files.eric.ed.gov/fulltext/ed541982.pdf lanza, s. t., & cooper, b. r. (2016). latent class analysis for developmental research. child development perspectives, 10 (1), 59–64. https://doi.org/10.1111/cdep.12163 lewis, l. s. (2020). nursing students who fail and repeat courses: a scoping review. nurse educator, 45(1), 30–34. https://doi.org/10.1097/nne.0000000000000667 li, s., zheng, j., & lajoie, s. p. (2021). the frequency of emotions and emotion variability in self-regulated learning: what matters to task performance? frontline learning research, 9(4), 76–91. https://doi.org/10.14786/flr.v9i4.901 lin, h. c. k., chao, c. j., & huang, t. c. (2015). from a perspective on foreign language learning anxiety to develop an affective tutoring system. educational technology research and development, 63 , 727–747. https://doi.org/10.1007/s11423-015-9385-6 lo, y., mendell, n. r., & rubin, d. b. (2001). testing the number of components in a normal mixture. biometrika, 88(3), 767–778. https://doi.org/10.1093/biomet/88.3.767 lu, g., xie, k., & liu, q. (2023). an experience-sampling study of between-and within-individual predictors of emotional engagement in blended learning. learning and individual differences, 107, 102348. https://doi.org/10.1016/j.lindif.2023.102348 madsgaard, a., røykenes, k., smith-strøm, h., & kvernenes, m. (2022). the affective component of learning in simulation-based education–facilitators’ strategies to establish psychological safety and accommodate nursing students’ emotions. bmc nursing, 21(1), 91. https://doi.org/10.1186/s12912-022-00869-3 marsh, h. w., lüdtke, o., trautwein, u., & morin, a. j. (2009). classical latent profile analysis of academic self-concept dimensions: synergy of person-and variable-centered approaches to theoretical models of self-concept. structural equation modeling: a multidisciplinary journal , 16(2), 191–225. https://doi.org/10.1080/10705510902751010 mclaren, c. h. (2004). a comparison of student persistence and performance in online and classroom business statistics experiences. decision sciences journal of innovative education , 2(1), 1–10. https://doi.org/10.1111/j.0011-7315.2004.00015.x naghavi, f., redzuan, m., asgari, a., & mirza, m. (2012). gender differences and construct of the early adolescent’s emotional intelligence. life science journal, 9(2), 124–128. http://www.dx.doi.org/10.7537/marslsj090212.21 nylund, k. l., asparouhov, t., & muthén, b. o. (2007). deciding on the number of classes in latent class analysis and growth mixture modeling: a monte carlo simulation study. structural equation modeling: a multidisciplinary journal , 14(4), 535–569. https://doi.org/10.1080/10705510701575396 orri, m., pingault, j. b., rouquette, a., lalanne, c., falissard, b., herba, c., ... & berthoz, s. (2017). identifying affective personality profiles: a latent profile analysis of the affective neuroscience personality scales. scientific reports, 7(1), 4548. https://www.nature.com/articles/s41598-017-04738-x paden, r. r. (2006). a comparison of student achievement and retention in an introductory math course delivered in online, face-to-face, and blended modalities (doctoral dissertation, capella university). https://www.proquest.com/openview/90bd37d4a83aa271ebb557a4df32591a/1?pq-origsite=gscholar&cbl=18750&diss=y pekrun, r. (2006). the control-value theory of achievement emotions: assumptions, corollaries, and implications for educational search and practice. educational psychology review,18(4), 315–341. https://doi.org/10.1007/s10648-006-9029-9 pekrun, r. (2024). control-value theory: from achievement emotion to a general theory of human emotions. educational psychology review, 36 , 83. https://doi.org/10.1007/s10648-024-09909-7 pekrun, r., goetz, t., daniels, l. m., stupnisky, r. h., & perry, r. p. (2010). boredom in achievement settings: exploring control–value antecedents and performance outcomes of a neglected emotion. journal of educational psychology , 102(3), 531. https://doi.org/10.1037/a0019243 pekrun, r., goetz, t., frenzel, a. c., barchfeld, p., & perry, r. p. (2011). measuring emotions in students’ learning and performance: the achievement emotions questionnaire (aeq), contemporary educational psychology, 36 (1), 36–48. https://doi.org/10.1016/j.cedpsych.2010.10.002 pekrun, r., goetz, t., titz, w., & perry, r. p. (2002). academic emotions in students' self-regulated learning and achievement: a program of qualitative and quantitative research. educational psychologist, 37(2), 91–105. https://doi.org/10.1207/s15326985ep3702_4 pekrun, r., & linnenbrink-garcia, l. (2012). academic emotions and student engagement. in a. l. reschly & s. l. christenson (eds.) handbook of research on student engagement (pp. 259–282). springer. https://doi.org/10.1007/978-1-4614-2018-7_12 pekrun, r. & stephens, e. j. (2010). achievement emotions: a control value approach. social and personality psychology compass, 4(4), 238–255. https://doi.org/10.1111/j.1751-9004.2010.00259.x pillai, r., & williams, e. a. (2004). transformational leadership, self‐efficacy, group cohesiveness, commitment, and performance. journal of organizational change management , 17(2), 144–159. https://doi.org/10.1108/09534810410530584 roach, r. (2002). staying connected. black issues in higher education, 19 (18), 22–25. https://www.proquest.com/openview/06825bcb3ba5a9c4533ec0d8e32faf2d/1.pdf?pq-origsite=gscholar&cbl=27805 rosenberg jm, beymer pn, anderson dj, van lissa cj, schmidt ja (2018). “tidylpa: an r package to easily carry out latent profile analysis (lpa) using open-source or commercial software.” journal of open source software , 3(30), 978. doi:10.21105/joss.00978 , rowe, j. (2006). non-defining leadership. kybernetes, 35(10), 1528–1537. https://doi.org/10.1108/03684920610688568 rubinsten, o. & tannock, r. (2010). mathematics anxiety in children with developmental dyscalculia. behavioral and brain functions, 6(46). https://doi.org/10.1186/1744-9081-6-46 schoenfeld-tacher, r., mcconnell, s., & graham, m. (2001). do no harm—a comparison of the effects of on-line vs. traditional delivery media on a science course. journal of science education and technology, 10 , 257–265. https://doi.org/10.1023/a:1016690600795 rstudio team (2020). rstudio: integrated development for r. rstudio, pbc, boston, ma url http://www.rstudio.com/ . schwarz, g. (1978). estimating the dimension of a model. the annals of statistics, 6(2), 461–464. https://doi.org/10.1214/aos/1176344136 sclove, s. l. (1987). application of model-selection criteria to some problems in multivariate analysis. psychometrika, 52(3), 333–343. https://doi.org/10.1007/bf02294360 seli, p., wammes, j. d., risko, e. f., & smilek, d. (2016). on the relation between motivation and retention in educational contexts: the role of intentional and unintentional mind wandering. psychonomic bulletin & review , 23, 1280–1287. https://doi.org/10.3758/s13423-015-0979-0 snead, s. l., walker, l., & loch, b. (2022). are we failing the repeating students? characteristics associated with students who repeat first-year university mathematics. international journal of mathematical education in science and technology , 53(1), 227–239. https://doi.org/10.1080/0020739x.2021.1961899 song, l., singleton, e. s., hill, j. r., & koh, m. h. (2004). improving online learning: student perceptions of useful and challenging characteristics. the internet and higher education, 7(1), 59–70. https://doi.org/10.1016/j.iheduc.2003.11.003 spurk, d., hirschi, a., wang, m., valero, d. & kauffeld, s. (2020). latent profile analysis: a review and “how to” guide of its application within vocational behavior research. journal of vocational behavior, 120, p. 103445. https://doi.org/10.1016/j.jvb.2020.103445 strain, a. c., & d ‘mello, s. k. (2011). emotion regulation during learning. in artificial intelligence in education: 15th international conference, aied 2011, auckland, new zealand, june 28–july 2011 15 (pp. 566-568). springer berlin heidelberg. http://dx.doi.org/10.1007/978-3-642-21869-9_103 sözer boz, e., & kahraman, n. (2023). latent trajectories of subjective well-being: an application of latent growth curve and latent class growth modeling. international journal of contemporary educational research , 10(2), 411–423. https://doi.org/10.52380/ijcer.2023.10.2.308 tamin, r. z., tamin, a. z., & marzuki, p. f. (2011). performance based contract application opportunity and challenges in indonesian national roads management. procedia engineering, 14, 851-858. https://doi.org/10.1016/j.proeng.2011.07.108 tan, j., mao, j., jiang, y., & gao, m. (2021). the influence of academic emotions on learning effects: a systematic review. international journal of environmental research and public health , 18(18), 9678. https://doi.org/10.3390/ijerph18189678 tenk, (2019). national advisory board on research ethics (2019). ethical principles of research in the humanities and social and behavioural sciences and proposals for ethical review. available online at: http://www.tenk.fi/sites/tenk.fi/files/ethicalprinciples.pdf tyng, c. m., amin, h. u., saad, m. n., & malik, a. s. (2017). the influences of emotion on learning and memory. frontiers in psychology , 1454. https://doi.org/10.3389/fpsyg.2017.01454 um, e., plass, j. l., hayward, e. o., & homer, b. d. (2012). emotional design in multimedia learning. journal of educational psychology, 104(2), 485. https://doi.org/10.1037/a0026609 wang, a. y., & newlin, m. h. (2000). characteristics of students who enroll and succeed in psychology web-based classes. journal of educational psychology , 92(1), 137. https://doi.org/10.1037/0022-0663.92.1.137 wang, c., teng, m. f., & liu, s. (2023). psychosocial profiles of university students’ emotional adjustment, perceived social support, self-efficacy belief, and foreign language anxiety during covid-19. educational and developmental psychologist , 40(1), 51–62. https://doi.org/10.1080/20590776.2021.2012085 wang, y., shen, b., & yu, x. (2021). a latent profile analysis of efl learners’ self-efficacy: associations with academic emotions and language proficiency. system, 103, 102633. https://doi.org/10.1016/j.system.2021.102633 wytykowska, a., fajkowska, m., & skimina, e. (2022). taking a person-centered approach to cognitive emotion regulation: a latent-profile analysis of temperament and anxiety, and depression types. personality and individual differences , 189, 111489. https://doi.org/10.1016/j.paid.2021.111489 zajacova, a., lynch, s. m., & espenshade, t. j. (2005). self-efficacy, stress, and academic success in college. research in higher education , 46, 677–706. https://doi.org/10.1007/s11162-004-4139-z codepen compagnoni frontline learning research vol.8 no. 2 (2020) 131 152 issn 2295-3159 “i’m the best! or am i?”: academic self-concepts and self-regulation in kindergarten miriam compagnonia, kelsey marie losennob auniversity of zurich, switzerland bmcgill university, canada article received 15 january 2020/ revised 8 april/ accepted 25 april / available online 25 may abstract in this paper, we examined how kindergarteners’ self-evaluation biases are related to behavioural self-regulation (sr) and learning goal orientation (go). according to educational research and practice, fostering high and optimistic academic self-concepts promotes the setting of challenging goals and initiates effective behavioural sr processes. however, research on metacognition states that it is a match between academic self-concept and abilities that provides the optimal conditions for behavioural sr and a learning go. there is theoretical and empirical evidence in favour of both positions, yet the correlates of self-evaluative tendencies may differ with children’s different levels of achievement, which are rarely considered. this cross-sectional study used response surface analysis, an innovative research methodology capable of assessing the complex interaction of academic self-concept and academic abilities on the behavioural sr and go of 147 kindergarten children (m = 6.47 years, sd = 0.39 years). polynomial regression models were used to test the presence of a fit pattern in empirical data and offer a new perspective on the interaction of academic self-concept and academic abilities. results showed that a fit is generally associated with better behavioural sr and a learning go but that correlates of academic self-concept differ with different achievement levels and outcome measures. this study extends current knowledge, as it offers important insights on how to conceptualise and pursue questions regarding self-concepts and behavioural sr. at an applied level, the findings indicate that interventions with kindergarteners that target sr should take the interactions between self-evaluation biases and ability level into account. keywords: self-concept, behavioural self-regulation, goal-orientation, kindergarten, response surface analysis info corresponding authoremail: mcompagnoni@ife.uzh.ch doi: https://doi.org/10.14786/flr.v8i2.605 1. introduction supporting students in their acquisition of positive self-concepts is generally accepted as a pedagogical goal (dickhäuser, 2006; hellmich, 2011), but the strong positive self-image of kindergarteners still raises questions for researchers and teachers on whether and how to deal with it. to date, the educational research literature supports two contrasting positions on the role of academic self-concepts in supporting students’ behavioural self-regulation (sr) and achievement (bouffard & narciss, 2011; praetorius et al., 2016). specifically, educational research often describes a positively biased self-concept as desirable for promoting the setting of challenging goals and initiating effective behavioural sr processes (bouffard et al., 2006; dupeyrat et al., 2011; taylor et al., 2000). however, research on self-regulated learning (srl) and metacognition highlights the importance of non-biased self-concepts. a fit between students’ self-concepts and an external criterion of academic abilities suggests that students effectively engage in metacognition, which supports children in setting appropriately challenging learning goals and becoming effective self-regulated learners (butler & winne, 1995; roebers et al., 2012). there is empirical evidence in support of both positions. the contradictory findings have been discussed concerning differences in terms used (bouffard & narciss, 2011), outcome measures (pinxten et al., 2010), measurement (bouffard et al., 2006; praetorius et al., 2016), and costs and benefits (butler, 2011; destan & roebers, 2015; narciss et al., 2011). the widely accepted reciprocal effects model (marsh & martin, 2011) posits that self-concept and achievement mutually reinforce each other through sr mechanisms, yet the intermediating sr mechanisms remain understudied. additionally, in the instructional context a differentiated view is needed, especially with research on young children, since the correlates and consequences of self-evaluative tendencies may differ at different achievement levels (butler, 2011). in kindergarten, where self-concepts tend to be overoptimistic, the relation with early sr as an important predictor of school success has yet to be investigated (blair & raver, 2015; perry et al., 2017). with the lack of a well-developed theoretical framework, there is little insight into whether kindergarteners’ self-evaluative tendencies towards a positive bias (positive bias hypothesis), a fit (fit hypothesis) or differential correlates at different academic ability levels (differential hypothesis) are associated with better learning behaviour. although the differential hypothesis may be integral to resolving the inconsistencies in our other hypotheses, we propose an exploratory approach that allows us to describe and test fit patterns in empirical data. we therefore employed response surface analysis, an innovative research methodology that computes and visualises polynomial models (humberg et al., 2018; schönbrodt et al., 2018) to compare the three different hypotheses (humberg et al., 2017) and to take into account some shortcomings of the previous methods employed in self-concept research. as such, the present study examined the interaction of kindergarteners’ academic self-concepts and an external criterion of academic ability in explaining differences in their goal orientation (go) and behavioural sr. before describing our study, we outline previous theoretical and empirical research on self-concepts, behavioural sr, and learning go. 1.1 self-concepts in kindergarten the academic ability self-concept is an individual’s cognitive representation of their academic abilities (dickhäuser, 2006). efklides (2011) describes self-concepts as trait-like characteristic at the person level that interact with their competences, motivation, affect, volition and metacognition. as such, self-concepts set goal-directed top-down and bottom-up sr processes in motion (e.g. behavioural sr) and are closely linked to metacognition (i.e. cognition about cognition; flavell, 1979). metacognition can be differentiated into metacognitive knowledge, metacognitive experience and metacognitive skills. metacognitive knowledge is understood as a person’s awareness of their strengths and weaknesses and is often conceptualised similarly to self-concepts. in contrast, metacognitive experience is conceptualised as feelings and judgments about cognition, and it provides feedback on children’s ability self-concept during learning tasks (efklides, 2011). for example, children who have a positive self-concept of their mathematic abilities might choose to play a difficult numbers game, use different strategies and expect to succeed in this task. therefore, self-concepts are understood as an antecedent of sr processes and as motivational beliefs with great significance for academic learning (e.g. schunk & green, 2018). children experiencing failure may adjust their self-concept in dependency on that metacognitive experience. therefore, a person’s development of their self-concept is also the result of metacognitive experiences, knowledge and skills during the self-regulation process. in kindergarteners, self-concepts tend to be biased towards overconfidence; self-concepts then become more accurate throughout primary school (arens et al., 2016; hasselhorn, 2005; jacobs et al., 2002; lipko et al., 2012). according to hasselhorn (2005), when asked to list the three highest-performing children in their class, 95% of kindergarteners name themselves. plausible reasons for this overconfidence are their poorly developed metacognition, the lack of opportunity for social comparison, limited formal feedback within the learning environments and praise for accomplishing easy tasks (hasselhorn, 2005; hellmich, 2011). although kindergarteners’ self-concepts are biased towards overconfidence, they are related to academic ability (cimeli et al., 2013; marsh et al., 2002). findings demonstrate that young children are more capable and accurate in gauging their own abilities than previously believed (whitebread et al., 2007), especially when assessed with a domain-specific measure (cimeli et al., 2013) or on the same scale with the same reference standards (müller et al., 2015). like adults’ self-concepts, kindergarteners’ self-concepts are composed of several domain-specific facets (shavelson et al., 1976). for example, in swiss kindergarteners a self-concept for combined academic abilities in math and language can be distinguished from social and play-based self-concepts (cimeli et al., 2013). this reflects that in switzerland, domain-specific, formal instruction and assessment in reading and writing do not typically begin before primary school. rather, kindergartens emphasize free play in an open learning environment, allowing children to plan what to play and where, and for how long and with whom (hauser, 2013). this setting gives children opportunities to employ and improve their sr (timmons et al., 2016); given their relation to metacognition, motivation and academic ability, self-concepts may play an important role in early sr (cimeli et al., 2013; marsh et al., 2002). since few studies have examined the correlates of self-concepts in this age group (butler, 2011), there is a paucity of research examining the role of self-concepts in explaining early sr. 1.2 early self-regulation young learners’ early sr is a known predictor of school adjustment and academic success (blair & raver, 2015; gestsdottir et al., 2014; moffitt et al., 2013). research on sr—and srl as extension of sr during learning—emphasizes metacognition, motivation, affect, volition and cognition as key components to regulate and control behaviour (for a discussion on different models of sr and srl, see efklides, 2011; panadero, 2017). children are seen as agents of their own sr processes, through which they set goals and engage in learning tasks, monitor and evaluate their cognition, behaviour and learning outcomes and reflect on themselves as learners, which includes updating their self-concepts (efklides, 2011). based on efklides’ model (2011), interactions between personal characteristics (e.g. self-concepts, motivation, ability) guide sr processes. for kindergarteners, behavioural sr refers to the child’s abilities of “focusing and maintaining attention on tasks, following instructions, and inhibiting inappropriate actions” (sektnan et al., 2010, p. 466). these abilities are a result of the child’s inhibitory control, working memory and attention flexibility, which are known as basic executive functions: a family of top-down, domain-general skills that guide and control thought and behaviour and are implicit in sr (diamond, 2016; garner, 2009; perry et al., 2017). to train them, repeated practice and increased challenge to executive functions is mandatory (diamond & lee, 2011). being motivated to engage in challenging learning tasks to increase competencies and acquire or master new skills reflects a learning go (dweck & leggett; 1988), which is sometimes termed a mastery approach go (for a discussion on the different terms, see zimmerman & schunk, 2008). a learning go is associated with children’s incremental motivational framework (compagnoni et al., 2019), which indicates that learning requires time and effort (perry et al., 2019). for example, a kindergartener with a learning go might choose to play a new difficult game with numbers rather than replay a familiar game, even though it requires attention and persistence and success is not guaranteed. in contrast, a performance orientation is associated with engaging in easy tasks that one can master quickly with minimal effort; it conveys the desire to achieve success and outperform others (bakadorova & raufelder, 2020). a learning go seems to be a hallmark for training srl, since students with well-developed sr tend to engage and persist in challenging tasks (hutchinson, 2013; perry, 2013). in the autonomous, open learning environments in swiss kindergartens this raises the question as to the role of self-concepts. who is more likely to be learning goal oriented, therefore engaged in challenging tasks and show better behavioural sr: children with positively biased self-concepts or children with more congruent self-views? and does it differ for high or low achievers? 1.3 self-concept and self-regulation given that early sr is a predictor of successful school adjustment and learning (butler, 2011), and because researchers agree that self-concept can influence one’s sr through interactions with go and cognitive resources at a personal-level (efklides, 2011), it is crucial to gain insight into the interaction of self-concepts with self-regulation for learning in young children (diamond, 2016; efklides, 2011; perry et al., 2017). researchers recognise two effects regarding self-concepts: a self-enhancement effect, which suggests that positively biased self-concepts have a positive effect on the development of abilities (positive bias hypothesis) through self-affirmation or sr mechanisms (taylor & brown, 1988; valentine et al., 2004), and a skill-development effect, which suggests that with increasing metacognitive development, self-concepts adapt to abilities (roebers et al., 2012), resulting in an increased fit, which is related to better sr (fit hypothesis). studies on these two effects are often conducted in longitudinal panel studies, where both self-concept and achievement are measured, but the mediating sr mechanisms are neglected. based on these studies, most researchers assume a reciprocal relationship, where self-concept and achievement mutually reinforce each other (arens et al., 2016; marsh & martin, 2011; pinxten et al., 2010), although the skill-development effect is more pronounced (helmke & vanaken, 1995; muijs, 1997; praetorius et al., 2016). for example, praetorius et al. (2016) reported no or only a small self-enhancement effect at the start of primary school, suggesting that self-concept might only have a motivational effect when the environment is challenging and new. in the instructional context, it is especially important to tease these two effects apart; kindergarten teachers need to know how (and whether) to deal with their pupils’ strong positively biased self-concepts. as such, it is important to explore and understand how these competing effects can be disentangled and examine whether there is a unique relation with behavioural sr and go. positive bias hypothesis. self-concept theory (marsh & martin, 2011) posits that positive beliefs act as an internal resource that fuels motivational-emotional learning and initiates effective behavioural sr processes (bouffard & narciss, 2011). studies with school children and adolescents demonstrate a link between positively biased self-concepts and higher intrinsic motivation (bouffard et al., 2003; dupeyrat et al., 2011), higher interest (gonida & leondari, 2011), higher effort to persist and maintain motivation (dermitzaki et al., 2009), higher expectations of success (dickhäuser, 2006) and higher learning gains (shin et al., 2007). longitudinal findings during adolescence indicate that the effect of academic self-concepts on achievement is partially mediated by learning go (bakadorova & raufelder, 2020). this pattern suggests that positive beliefs promote the setting of challenging goals and the employment of effective learning behaviours despite set-backs, and they may ultimately support adaptive sr. regarding go, cost and benefits of a positive bias are discussed (butler, 2011). gonida and leondari (2011) reported that a positive bias is related to a higher mastery, performance-approach and performance-avoidance go and an orientation towards pleasing significant others. these findings suggest that children with a positively biased self-concept are more likely to choose challenging learning tasks over easy learning tasks, but the motivation may be externally sourced or may be the desire to outperform others. fit hypothesis. research in self-regulated learning highlights the importance of a fit between self-concept and abilities. as students gain control over their cognitive abilities, they are better able to form and evaluate representations of their abilities (metacognition), which is central to adaptive sr (destan & roebers, 2015; diamond, 2016). children showing a fit between their self-concept and abilities are assumed to be able to adapt to task conditions and set challenging learning goals and are therefore more likely to adopt and train effective behavioural sr (destan & roebers, 2015; flavell, 1979; perry et al., 2017). as studies have demonstrated that students’ self-concept is informed and refined by their metacognitive processes (diamond, 2016; efklides & tsiora, 2002), positively and negatively biased self-concepts may be both the cause and result of poor sr abilities and may signal a developmental lag in metacognition. researchers commonly agree that negatively biased self-concepts are related to only costs and no benefits (bouffard et al., 2003; gonida & leondari, 2011). however, positive biases may also come at a cost, with overconfidence being described as a powerful cognitive bias that leads to poor monitoring and control abilities (destan & roebers, 2015), self-handicapping (young-hoon et al., 2010), a performance go, and avoidance of challenging tasks as a means of performing well and preserving one’s positive self-views (butler, 2011; dupeyrat et al., 2011). although there is little evidence that children’s accurate self-views have more benefits than costs compared to a positive bias (butler, 2011), researchers examining dyadic effects point out that an optimal fit might not be exactly on the line of numerical congruence (schönbrodt et al., 2018), and a slightly positively biased self-concept might better support behavioural sr. differentiation hypothesis. since a bias or a fit in academic self-concept may represent different meanings for high and low achievers (butler, 2011; marsh et al., 2002), we propose a third hypothesis that may be integral to resolving the inconsistencies in the literature on children’s self-concept in relation to behavioural sr and go. some authors suggest that positively biased self-concepts act as motivational boost in challenging or threatening situations (praetorius et al., 2016), similar to the preference for downward comparison in threatening situations described by early social comparison theory (guyer & vaughan-johnston, 2018). if we assume that low-achieving children see the academic environment as challenging, we would expect that the effects of the positive bias position would be especially applicable. although a positive bias might lead to more persistence for children with low academic ability levels, it may also lead to a performance go, since children might prefer easy tasks that they have already mastered so as to keep their positively biased self-concept (dweck, 2017). therefore, it may act as a self-defence mechanism against poor performance (loveless, 2006). a negative bias in children with low academic ability levels is likely to be related to helpless behavioural patterns and disengagement and should therefore be consistently related to maladaptive behaviour (eckert et al., 2006). some researchers argue that for competent children, a negative bias might also have positive motivational consequences, resulting in better monitoring and control abilities (destan & roebers, 2015), more effort and investment in deep learning strategies (blanton et al., 1999), and―based on social comparison theory―upward comparison as a need for self-improvement (collins, 1996). for children with higher academic ability levels, a fit between self-concept and ability is associated naturally with higher abilities and should be more conducive to adjustment and motivation (butler, 2011). therefore, it is not the absolute level of the self-concept in an inter-individual comparison that is relevant for correlations with behavioural sr and go but rather an intra-individual approach in which self-concept and ability level are calculated from a person-centred perspective. longitudinal studies or studies using difference scores might therefore neglect the achievement level or the proposed non-linear relation as proposed in the fit hypothesis. 2. the present study this study aims to disentangle the roles that kindergarteners’ academic self-concepts play in explaining differences in behavioural sr and go, by examining kindergarteners’ self-concept biases with an intra-individual approach and to take the general ability level into account as well as assumed non-linear correlations. three competing hypotheses have been suggested, which all have sound empirical and theoretical bases but have never been tested simultaneously for different sr outcomes: (1) the positive bias hypothesis suggests that a positively biased self-concept is related to better behavioural sr and learning go than a negatively biased self-concept or a fit between self-concept and ability level; (2) the fit hypothesis suggests a fit between self-concept and ability level is related to better behavioural sr and learning go; and (3) the differential hypothesis suggests differential effects for different ability levels. specifically, a fit between self-concept and ability level is the most beneficial for behavioural sr and a learning go for students with medium ability levels. for students with low ability levels, a slightly positive bias is associated with higher behavioural sr than a fit or a negative bias is. lastly, for students with high ability levels, a slightly negative bias is related with higher behavioural sr than a fit or a positive bias is. we proposed that the differential hypothesis may be the most representative of the actual dynamics between these variables. 3. method 3.1 context of the study and participants in switzerland kindergartens are part of the public education system, and 95% of children attend a two-year kindergarten program in their local public schools starting at age 4 or 5 (edk, 2017). classes are comprised of age-mixed children with diverse socioeconomic (ses) backgrounds, ethnicities and first languages. kindergarten education is highly interdisciplinary, and an open learning environment and play are of great importance; children are slowly introduced to domain-specific learning like math and literacy. nineteen kindergartens from different schools (m = 7.4 children per class) in urban and rural areas participated in the study and reflect the demographic composition of the german-speaking part of switzerland. the kindergarteners’ first languages were: 45% swiss german, 10% albanian, 7% serbian/croatian, 5 % turkish, 3% for portuguese, english, german and arabic, respectively, 2% spanish; the rest had other first languages. teachers reported that 72% of the children were of swiss nationality, which matched official data. parents and children’s informed consent to participate was received from 91%, and refusals were unsystematic. the final sample consisted of 147 children (52% girls) in the second year of kindergarten (m = 6.47 years old, sd = 0.39). missing data was due to children who were ill at one of the two measurement times or due to technical failures. 3.2 materials academic abilities and academic self-concept. children provided assessments of their academic abilities (i.e. knowing letters, reading, writing, knowing numbers, arithmetic, and counting) by selecting their rank position out of 9 stickmen in a row that represented their classmates (cimeli et al., 2013). they were told that the stickman on the right represented the classmate with the best abilities in e.g. counting, whereas the stickman on the left represented the classmate with the poorest abilities, whereupon children marked the stickman that was most representative of their own position. this approach counteracts children’s positively biased self-concept ratings (cimeli et al., 2013). teachers provided ratings of their students’ math and literacy abilities on the same scale, to take social comparisons into account (pinxten et al., 2010). raw scores ranged from 1 (poorest in the class) to 9 (best in the class) for academic self-concepts (m = 6.66, sd = 1.75, α = .77), and for teacher ratings of academic abilities (m = 6.19, sd = 2.05 α =.84). residual scores were used to calculate the absolute and relative deviation of children’s self-concept from teacher ratings. residuals represent the part of the self-concept that cannot be explained by the corresponding teacher rating. behavioural self-regulation. the new version of the head-toes-knees-shoulders (htks) measure was used as a direct observational indicator of behavioural sr, as it measures inhibitory control, working memory and attention flexibility (mcclelland et al., 2014). in the htks, the children were asked to play a game where they must do the opposite of what the experimenter says. for example, if the experimenter instructed the students to touch their toes, they had to touch their head instead. the first 10 items included two paired commands (head – toes); the next 10 items added two new paired commands (shoulders – knees); and for the last 10 items, the four commands were paired differently than before. for children to be successful in these tasks, they must focus on the instructions and commands (attention flexibility), remember the paired rules (working memory) and stop a dominant response tendency and replace it with the opposite response (inhibitory control). each of the 30 items was scored as 0 for an incorrect response, 1 for a self-correction, or 2 for a correct response. as instructed in the test manual, the five children who scored very low on the first tasks were not allowed to finish, and scores were adjusted accordingly. total scores on the htks ranged from 12 to 57 points (m = 41.71, sd = 10.35, α = .86), where higher scores indicated higher levels of behavioural sr. learning goal orientation. to assess children’s learning go, a self-report method was used based on the berkeley puppet interview (see figure 1; measelle et al., 1998) with items derived from the motivational framework measures from gunderson et al. (2013). learning goal orientation therefore was assessed as preference for challenging versus easy tasks in service of learning goals. children listened to two elephant puppets on a touchscreen: one elephant expressed a high learning go (e.g. “i prefer to do very hard tasks so i can get better”) and the other a low learning go (e.g. “i prefer to do easy tasks that i’m good at”). children indicated on a 5-point semantic differential scale how well they could identify with one of the puppets. as suggested by marsh et al. (2002), a double binary response strategy was used, where the identification with one puppet (by pressing a button) was always followed by a second probe (“do you totally agree with this puppet, or do you agree only a little?”) to counter the tendency to select endpoints and neglect intermediate points. total scores ranged from 1 to 5 points (m = 3.68, sd = 1.16, α = .88), where higher scores indicated higher levels of learning go. figure 1. learning goal orientation measurement instrument. 3.3 procedure data collection took place in the spring semester of 2017. since the entire assessment took more than 50 minutes per child and would have exceeded their attention span, each kindergarten was visited twice within a period of 2-4 weeks. given the students’ lack of reading and writing skills, the first author and two trained research assistants administered each measure with each child separately during regular classroom hours. order of exposure to each measure was the same for all students, but items on measures of self-concept and go were counterbalanced to control for any effects of order. teachers completed an online questionnaire on the children’s demographics and academic achievement in a session that lasted 4 minutes per child. 3.4 data analytic approach research on self-concept bias often uses difference or residual scores, which represents the part of the self-concept that cannot be explained by the corresponding external criterion. however, this procedure is not recommended, as it does not take linear major effects into account (humberg et al., 2017) and relies on the assumption that the optimal fit is exactly on the line of numerical congruence (schönbrodt, 2016). since linear regression models impose linear constraints on the parameters and would fit our anticipated non-linear data pattern poorly, polynomial regression models were employed to test our hypotheses. the response surface analysis (rsa) package for r (schönbrodt, 2016) allowed for the comparison of different polynomial models using a path modelling approach (humberg et al., 2017). we predicted the impact of the interaction of two predictors collected with comparable scales (academic self-concept and academic ability rating) on an outcome measure (behavioural sr and go), so non-linear effects could be modelled. the different hypothesised models could be expressed as constrained multiple regressions, and data were analysed to detect fit patterns that would confirm one of the competing hypotheses. three models were tested: (1) the rising ridge model, examined if a fit (a ridge) between academic self-concept and ability rating is related to positive behavioural sr and a learning go (fit hypothesis); (2) the shifted rising ridge model, tested if the positive bias hypothesis best describes our data (we assumed additionally a shifted ridge; positive bias hypothesis); and (3) the shifted and rotated rising ridge model, which tested differential effects for high and low academic ability levels (allowing the ridge to rotate; differential hypothesis). the mean level (bm) effect was incorporated in the three models―a linear major effect of the predictors on the outcome, so that the ridge is inclined in a direction that high/high combinations of self-concept and academic ability is related to higher values of the outcome than low/low combinations. the rsa package additionally computed an additive model with linear main effects of the two predictors and an interaction model (ia). for a full review of rsa and associated equations for all models, see humberg et al. (2017) and schönbrodt et al. (2018). 4. results 4.1 preliminary analyses table 1 shows descriptive information and correlational analyses for self-concept, ability ratings, behavioural sr, go, residual scores and covariates measures. for small class sizes, icc scores for the main constructs between 0.11 and 0.08 are considered reasonably small (hox, 2002). power analysis conducted with the variance inflation factor and these sample sizes showed a type i error of .05, and a power of 0.80. residual scores revealed that the more positive the bias, the more learning-oriented the children were, also when gender and age was controlled for (rs = .182, p = .035). there was no significant effect of residual scores on behavioural sr (rs = -.026, p = .762). absolute residuals scores demonstrated that the more incongruent children rated themselves with teacher ratings the higher their behavioural sr score was (rs = -.145, p = .085) and the more learning oriented they were (rs = -.281, p = .001). however, results derived from residuals have serious shortcomings (see above) and should be interpreted carefully. rsa can take non-linear effects and mean level effects into account and compare the different polynomial models. table 1 descriptive statistics and correlations for study variables note. ci = confidence interval; sr = self-regulation; go = goal orientation. pearson correlations are presented above the diagonal, spearman correlations below the diagonal. *p < .05 two-tailed. 4.2 response surface analyses tables 2 and 3 report model indices. as schönbrodt (2016) suggested for model comparison, we focused on the corrected form of the akaike information criterion (aicc), which corrects for a bias when the sample size is small compared to the number of model parameters (the model with the smallest aicc is considered the best model; schönbrodt, 2016). behavioural sr. table 2 shows the model indices for behavioural sr. according to aicc, the best model to predict behavioural sr from self-concepts and teacher ratings was the srrr model (differential hypothesis) with a model weight of 0.29. the δaicc between the srrr model and the less restricted additive and interaction model was < 2, which indicates that they were equally representative and in the range of plausible models. the χ2 test indicated that the srrr, the ia and the additive model were not significantly worse than the full polynomial model and were significantly better than the null model. according to the comparative fit index (cfi), all three models were around the rule of thumb value of .95 and had relatively good fit (hu & bentler, 1999). interpretation of parameter estimates reinforced the notion that both were suitable models to describe the data, but the best fitting model was the srrr, as it had the lowest aic. however, the additive model was the simplest well-fitting model. table 4 (see appendix) shows the regression coefficients b1 to b5 for the full polynomial regression models. the parameter for the mean-level effect, bm, was significantly different from a flat ridge, meaning that children with high/high combination showed better behavioural sr than children with low/low combinations. although the parameter s for the rotation of the ridge as well as a'4 for the fit did not show a significant value, according to the aicc criterion they still added to the quality of the model. table 2 model comparison for behavioural sr. ordered by δaicc note. k = number of parameters; aicc = corrected akaike information criterion; evidence ratio = ratio of model weights of the best model compared to each other model; cfi = comparative fit index; srmr = standardized root-mean-square residual; r2 = variance explained; pf = p value, model compared to full model; pn = p value, model compared to null model; r2adj = adjusted r2. model abbreviations: full = full polynomial model; srrr = shifted and rotated rising ridge model; srr = shifted rising ridge model; rr = rising ridge model; ia = interaction model; am = additive model; null = intercept-only model. *p < .001. since the results can be hard to interpret, we plotted the regression result as a three-dimensional response surface for the srrr model (figure 2) and the additive model (figure 3). the shape of the coloured area showed how behavioural sr depended on the combination of teacher rating and self-concept, with greener areas displaying higher behavioural sr. for the srrr model, the restrictions imposed a mean effect of the levels of the predictors and a maximum ridge. therefore, the area falls off to both sides, but both a rotation and a shift of the ridge is allowed. results showed a significant mean level effect, which in terms of content means that behavioural sr was higher, the higher the self-concept and the teacher ratings. for children with lower ability levels, behavioural sr was at a maximum when children had a slightly positive bias, with medium levels when there was a fit, and with high levels when there was a negative bias. the graph thus described the hypothesis of congruence, with a linear main effect of self-concept and teacher ratings on behavioural sr but with differing effects for low and high achievers. the additive model results showed a significant linear main effect, which in terms of content means that behavioural sr was higher, the higher the predictors, but the effect was based mainly on teacher ratings. figure 2. srrr model for behavioural sr. srrr model = shifted and rotated rising ridge model. blue line: line of congruence (loc). inner black circle contains 50% of data points. greener areas display higher behavioural self-regulation than redder areas. c = parameter for the shift of the ridge; s = parameter for the rotation of the ridge. figure 3. additive model for behavioural sr. blue lines: line of congruence (loc). inner black circle contains 50% of data points. greener areas display higher behavioural self-regulation than redder areas. goal orientation. table 3 shows the model indices for go. according to the cfi, only the interaction model had relatively good fit (hu & bentler, 1999). according to the aicc, the best model to predict go from self-concepts and teacher was also the ia model with a model weight of 0.69. the δaicc between the ia model and the other models was > 2, which indicates that the ia model may fit the data best. inspecting the other fit indices, all other models had a cfi < 0.86. the χ2 lr test indicated that the ia model fit the data significantly better than the null model and was not significantly worse than the full polynomial model. table 3 model comparison for goal orientation. ordered by δaicc note. k = number of parameters; aicc = corrected akaike information criterion; evidence ratio = ratio of model weights of the best model compared to each other model; cfi = comparative fit index; srmr = standardized root-mean-square residual; r2 = variance explained; pf = p value, model compared to full model; pn = p value, model compared to null model; r2adj = adjusted r2. model abbreviations: full = full polynomial model; srrr = shifted and rotated rising ridge model; srr = shifted rising ridge model; rr = rising ridge model; ia = interaction model; am = additive model; null = intercept-only model. *p < .001. the shape of the coloured area in figure 4 depicts how go was related to the combination of teacher judgement and self-concept. the greener the area, the more learning oriented the children described themselves. the restriction of the coefficients of the ia model led to a linear effect of self-concept and teacher ratings on behavioural sr as well as an effect of the interaction. regression coefficients b1 to b5 in table 5 (see appendix) confirmed a significant effect of self-concept as well as the interaction on learning go. in terms of content, this means that the learning go was generally higher, the higher the teacher ratings and the higher the self-concepts. but children with high self-concepts and low academic ability levels (strong positive bias), as well as children with low self-concepts but high academic ability levels (strong negative bias), had a very low learning go and reported a preference for easy tasks that they already master, so as to get a lot right. this is in contrast to children with high/high and low/low combinations, who were very learning oriented. figure 4. interaction model for goal orientation. blue line: line of congruence (loc). inner black circle contains 50% of data points. greener areas display higher learning go than redder areas. 5. discussion the main objective of the study was to determine the role that self-evaluative tendencies in kindergarteners play for behavioural sr and learning go, as inconsistent positions exist in the literature and a well-developed framework for this age group is missing. our review of theoretical and empirical work revealed three positions: first, a positively biased self-concept might act as motivational fuel and lead to a learning go and better behavioural sr (taylor et al., 2000). second, a fit between self-concept and abilities level may signal well-developed metacognition (destan & roebers, 2015), which is known to play a central role in the setting of learning goals and leads to effective behavioural sr (diamond, 2016; perry et al., 2017). and third, a differentiated view suggests that correlates of self-evaluative tendencies might differ at different ability levels and for different outcomes (bouffard & narciss, 2011; butler, 2011). our findings support the hypothesis that the correlates of biased self-concepts in kindergarteners differ for different ability levels. regarding behavioural sr, a fit between self-concept and academic ability level is beneficial for children with average abilities. this is in line with previous research in sr which suggests that a fit signals effective metacognition and ultimately better behavioural sr (butler & winne, 1995; roebers et al., 2012). however, our research extends previous findings, as it demonstrates that for children with high and low academic ability levels, a differentiated view is more appropriate. specifically, children with a high ability level had the best behavioural sr when they had a slightly negatively biased self-concept, which may indicate that it encourages them to exert more effort and to invest in deep learning strategies (blanton et al., 1999). alternatively, children with a low ability level performed best when they adopted a slightly positive bias, which might indicate that positive self-concepts act as a motivational boost to increase effort and persistence in learning (e.g. mastery go; gonida & leondari, 2011). however, the effects were quite small, and teacher ratings of children’s academic abilities were a better predictor of behavioural sr than children’s self-concepts were. these findings suggest that kindergarteners’ self-evaluative tendencies and absolute level of self-concepts only play a minor role in explaining their behavioural sr. future research should continue to pursue lines of questioning that anticipate unique relationships between biased self-concepts and sr based on individual students’ academic ability levels. regarding learning go, our results extend previous research findings, as they reveal differential relations based on academic ability levels. linear effects demonstrated that the higher ability levels were and the more positively biased students’ self-concepts were, the more the students reported a learning go. this is in line with previous educational research that suggests that a positive bias may be desirable for promoting learning go’s (dupeyrat et al., 2011; taylor et al., 2000). however, when taking non-linear effects into account, our results extend previous research as they demonstrate a significant interaction whereby children with strongly biased self-concepts were more likely to report preferring easy learning tasks that they had already mastered over challenging tasks. interestingly, this was true for both children with a low ability level who perceive themselves as being among the best children in class and children with a high ability level who perceive themselves as being among the worst children in their class. this may reflect the fact that children with low academic abilities and high self-concepts engage in easy tasks to perform well and avoid possible failure, protecting their positive self-view in front of others and themselves, and this reflects a performance orientation (dweck & leggett, 1988). in contrast, children with high ability levels and very low self-concepts might also avoid challenging tasks as a means of buffering against failure but clearly not to protect their self-concept. reasons might be a high avoidance orientation, fear of failure, pressure to perform or low self-efficacy. although these children have appropriate behavioural sr abilities in kindergarten, their go may become a hindrance over time if they continue to avoid challenging tasks, since challenge is critical for training sr (diamond & lee, 2011). children with a fit and with high as well as low academic ability levels showed a high learning go, which may reflect well-developed metacognitive abilities. our findings on children with mid-range ability were somewhat inconclusive regarding their go and may reflect a practice at schools whereby children at the extreme ends receive more resources in terms of attention, feedback and support to set challenging goals than children in the middle range do. in sum, early behavioural sr deficiencies are known to be problematic for school transitioning and future learning behaviour (blair & raver, 2015). therefore, research on correlates (e.g. self-concepts) in this age group is required to better understand the processes involved in the development of sr. kindergarteners’ self-concepts are slightly positively biased but are related to teacher ratings of their academic abilities (rs =.34), suggesting that kindergarteners do not form unreasonably biased self-concepts. the educational goal of supporting positive self-concepts is certainly valid if positive self-concepts are enhanced by fostering the underlying abilities (e.g. mathematic abilities). however, positive self-concepts alone do not seem to be a solution to behavioural sr deficiencies, as they too have costs (butler, 2011). a slightly positively biased self-concept might foster a learning orientation, and for low-ability children it seems to be related to better behavioural sr. but for higher achieving children a fit or even a slightly negative bias might set more adaptive sr processes in motion. children with an extreme self-concept misfit in both directions are the most at risk in their development. the use of an innovative approach like rsa offers a different perspective on how to conceptualise and pursue fit patterns regarding self-concept and external criterion in kindergartens in relation to an outcome such as behavioural sr and go and should continue to be adopted moving forward. 6. limitations and implications for research and practice although the present study has several advantages, there are four substantial limitations that should be addressed. first, the correlational nature of this study precludes any claims of causation, and the small sample size and reduced power prevented us from employing a structural equation model. although our use of rsa accounts for the unique non-linear influence of self-evaluative tendencies on an outcome and takes different ability levels into account, the approach neglects the possible influence of other variables (e.g. demographics). in our preliminary analyses, where we used residual scores and computed univariate analyses of variances, the results did not diverge when gender or age were included as control variables. but continued research is needed to gain insight into the directions and weights of paths between kindergartener’s self-concepts, abilities, behavioural sr and go, and possible interactions between variables such as gender, age or ses. second, the teacher ratings of students’ academic abilities used in this study have advantages and disadvantages. with no formal grades given in kindergarten, teacher’s assessment of students’ abilities are valid judgments and a congruence is considered relevant for common objectives (skaalvik & hagtvet, 1990). they also consider social comparison processes, as teachers in this study used their kindergarten group as social reference norm. this approach allowed us to compare children’s and teacher’s ability perceptions and led to a small variance across kindergarten classes. however, they cannot be considered objective measures, and they represent more than a mere reflection of students’ abilities because teacher ratings also take motivational characteristics into account (pinxten et al., 2010). future research should consider the use of both achievement tests and ratings but as separate latent constructs, since they have different psychological meaning (pinxten et al., 2010). third, although the assessment of learning go on a unidimensional scale is acceptable for this age group, future studies should try to capture differentiated gos (e.g. performance/mastery, avoidance/approach) to fully address the correlates and relations between self-concept, behavioural sr and go. in the open learning environment in kindergartens, where task demands are not always clear, differentiating performance and mastery go on two separate scales might be beneficial to gain deeper insights. fourth, given our small sample of 19 typical swiss kindergarten classes and the limited power, questions arise regarding the generalisability of our findings to different classes, schools and school levels. swiss kindergartens emphasize open learning environments, and it may be the case that the interaction of self-concept, ability level and sr is different in more structured environments, where there is less free play and free choice. future research should employ longitudinal designs with multiple variables to assess developmental patterns after the transition from kindergarten to primary school, where the educational setting often changes dramatically (e.g. open learning environments are rare; regular tests of performance and formal feedback are given). although different school types may also play a role, the classroom level is next to the individual level the most important pedagogical unit for explaining cognitive and motivational learning outcomes (wurster & feldhoff, 2019). it is plausible that teachers’ approaches to instruction or differences in classroom climate (e.g. non-threatening teacher feedback, graded work), may lead to differences in the associations between self-evaluation bias and sr in students. as we did not collect data on features of instruction or task design within classrooms and thereby across schools, future research should consider these variables to conduct multilevel analyses to examine differences across kindergartens and the prevalence of different profiles across schools. in practical terms, our results highlight the role of self-concepts in supporting motivation in early childhood. without attention to students’ underlying abilities, high self-concepts are not very meaningful in explaining differences in self-regulation. as such, teachers should mainly support students’ self-concepts by focusing on the improvement of their students’ abilities. our findings suggest that kindergarten teachers should adapt their instructional approaches (e.g. relevant feedback and task-specific experiences that give students information about their abilities and train their metacognition) to help kindergarteners align their self-concept and abilities, as a fit is related to better behavioural sr for most achievement levels. in students with low achievement levels, teachers could attempt to strengthen their self-concepts to be slightly positively biased (e.g. provide positive feedback and support in selecting tasks that are appropriately challenging and supportive of ability and self-concept), as this may result in better achievement and behavioural sr outcomes. further, our findings suggest that teachers should be sensitive to personal characteristics that influence learning go and behavioural sr in the classroom, as these characteristics support learning processes. although higher self-concepts are related to a learning go, which is typically preferred in schools, teachers should be particularly sensitive to strongly biased self-evaluations. that is, children with strong positive as well as strong negative biases report a preference for easy tasks they have already mastered, which may represent a fixed motivational framework where learning is considered something that is preferably quick and easy (dweck, 2017), and this can have negative consequences for srl in the long run (hutchinson, 2013). researchers, educators and policy makers may integrate these findings into their instructional approaches and interventions that target behavioural sr or srl to differentially support students based on their academic ability level. the world is not black or white, and it seems this is also the case with the question of whether a fit or a positive bias better supports learning outcomes. instead of dichotomising, researchers and educators should consider that there are unique benefits and risks associated with distinct self-concept biases for students with different academic ability levels. interventions should be designed with these findings in mind, so that low and high achieving students receive well-tailored support in developing their self-concept―that is, support that moves them towards a view of their academic abilities that best supports their behavioural sr. this study provides researchers and educators with new and interesting insights into the belief systems of kindergarteners and contributes to developing the theoretical framing of self-concepts, behavioural sr, go and ability level in educational psychology research. considering our findings, it seems crucial for educators to provide children with regular feedback as early as kindergarten to emphasise the setting of challenging yet appropriate learning goals and to promote the development of metacognition and consequently srl. keypoints with response surface analysis the article offers an innovative way to measure fit patterns between academic ability self-concepts and academic ability in explaining early sr one of the few studies that examines the role that self-evaluative tendencies of kindergarteners plays in explaining behavioural self-regulation and learning goal orientation there are unique benefits and risks associated with distinct self-evaluative tendencies for kindergarteners with different ability levels when planning and implementing educational interventions that target sr, researchers and practitioners should take young children’s self-evaluation biases into consideration based on the children’s ability levels. references arens, a. k., marsh, h. w., craven, r. g., yeung, a. s., randhawa, e., & hasselhorn, m. (2016). math self-concept in preschool children: structure, achievement relations, and generalizability across gender. early childhood research quarterly, 36, 391-403. https://doi.org/10.1016/j.ecresq.2015.12.024 bakadorova, o. & raufelder, d. (2020) the relationship of school self-concept, goal orientations and achievement during adolescence, self and identity, 19(2), 235-249. https://doi.org/ 10.1080/15298868.2019.1581082 blair, c., & raver, c. c. (2015). school readiness and self-regulation: a developmental psychobiological approach. annual review of psychology, 66, 711-731.https://doi.org/10.1146/annurev-psych-010814-015221 blanton, h., buunk, b. p., gibbons, f. x., & kuyper, h. (1999, mar). when better-than-others compare upward: choice of comparison and comparative evaluation as independent predictors of academic performance. journal of personality and social psychology, 76(3), 420-430. https://doi.org/10.1037/0022-3514.76.3.420 bouffard, t., boisvert, m., & vezeau, c. (2003). the illusion of incompetence and its correlates among elementary school children and their parents. learning and individual differences, 14(1), 31-46. https://doi.org/10.1016/j.lindif.2003.07.001 bouffard, t., & narciss, s. (2011). benefits and risks of positive biases in self-evaluation of academic competence: introduction. international journal of educational research, 50(4), 205-208. https://doi.org/10.1016/j.ijer.2011.08.001 bouffard, t., vezeau, c., & chouinard, r. (2006). l'illusion d'incompétence et les facteurs associés chez l'élève du primaire. revue française de pédagogie, 2(155), 9-20. https://doi.org/10.4000/rfp.61 butler, d. l., & winne, p. h. (1995). feedback and self-regulated learning: a theoretical synthesis. review of educational research, 65(3), 245-281. https://doi.org/10.2307/1170684 butler, r. (2011). are positive illusions about academic competence always adaptive, under all circumstances: new results and future directions. international journal of educational research, 50(4), 251-256. https://doi.org/10.1016/j.ijer.2011.08.006 cimeli, p., neuenschwander, r., roethlisberger, m., & roebers, c. m. (2013). self-concept of children at school entry: mean level, structure, and relations to indicators of early cognitive achievement. zeitschrift für entwicklungspsychologie und pädagogische psychologie , 45(1), 1-13. https://doi.org/10.1026/0049-8637/a000075 collins, r. l. (1996, jan). for better or worse: the impact of upward social comparison on self-evaluations. psychological bulletin, 119(1), 51-69. https://doi.org/10.1037/0033-2909.119.1.51 compagnoni, m., karlen, y., & maag merki, k. (2019). play it safe or play to learn: mindsets and behavioral self-regulation in kindergarten. metacognition and learning. https://doi.org/10.1007/s11409-019-09190-y dermitzaki, i., leondari, a., & goudas, m. (2009). relations between young students' strategic behaviours, domain-specific self-concept, and performance in a problem-solving situation. learning and instruction, 19(2), 144-157. https://doi.org/10.1016/j.learninstruc.2008.03.002 destan, n., & roebers, c. m. (2015). what are the metacognitive costs of young children’s overconfidence? metacognition and learning, 10(3), 347-374. https://doi.org/10.1007/s11409-014-9133-z diamond, a. (2016). why assessing and improving executive functions early in life is critical. in j. a. griffin, p. mccardle, & l. s. freund (eds.), executive function in preschool-age children: integrating measurement, neuro-development and translational research (pp. 11–43). american psychological association. https://doi.org/10.1037/14797-002 diamond, a., & lee, k. (2011). interventions shown to aid executive function development in children 4 to 12 years old. science, 333(6045), 959-964. https://doi.org/10.1126/science.1204529 dickhäuser, o. (2006). fähigkeitsselbstkonzepte. zeitschrift für pädagogische psychologie, 20(1/2), 5-8. https://doi.org/10.1024/1010-0652.20.12.5 dupeyrat, c., escribe, c., huet, n., & regner, i. (2011). positive biases in self-assessment of mathematics competence, achievement goals, and mathematics performance. international journal of educational research, 50(4), 241-250. https://doi.org/10.1016/j.ijer.2011.08.005 dweck, c., & leggett, e. (1988, apr). a social cognitive approach to motivation and personality. psychological review, 95(2), 256-273. https://doi.org/10.1037//0033-295x.95.2.256 dweck, c. s. (2017). the journey to children's mindsets—and beyond. child development perspectives, 11(2). 139-144. https://doi.org/10.1111/cdep.12225 eckert, c., schilling, d., & stiensmeier-pelster, j. (2006). einfluss des fähigkeitsselbstkonzepts auf die intelligenzund konzentrationsleistung. zeitschrift für pädagogische psychologie, 20(1/2), 41-48. https://doi.org/10.1024/1010-0652.20.12.41 edk, the swiss conference of cantonal ministers of education (2017, march). bildungssystem schweiz. http://www.edk.ch/dyn/14798.php efklides, a., & tsiora, a. (2002). metacognitive experiences, self-concept, and self-regulation. psychologia, 45(4), 222-236. https://doi.org/10.2117/psysoc.2002.222 efklides, a. (2011). interactions of metacognition with motivation and affect in self-regulated learning: the masrl model. educational psychologist, 46(1), 6-25. https://doi.org/10.1080/00461520.2011.538645 flavell, j. h. (1979). meta-cognition and cognitive monitoring: new area of cognitive-developmental inquiry. american psychologist, 34(10), 906-911. https://doi.org/doi 10.1037/0003-066x.34.10.906 garner, j. k. (2009, 2009/07/01). conceptualizing the relations between executive functions and self-regulated learning. the journal of psychology, 143(4), 405-426. https://doi.org/10.3200/jrlp.143.4.405-426 gestsdottir, s., von suchodoletz, a., wanless, s. b., hubert, b., guimard, p., birgisdottir, f., gunzenhauser, c., & mcclelland, m. (2014). early behavioral self-regulation, academic achievement, and gender: longitudinal findings from france, germany, and iceland. applied developmental science, 18(2), 90-109. https://doi.org/10.1080/10888691.2014.894870 gonida, e. n., & leondari, a. (2011). patterns of motivation among adolescents with biased and accurate self-efficacy beliefs. international journal of educational research, 50(4), 209-220. https://doi.org/10.1016/j.ijer.2011.08.002 gunderson, e., gripshover, s. j., romero, c., dweck, c., goldin-meadow, s., & levine, s. c. (2013, sep-oct). parent praise to 1to 3-year-olds predicts children's motivational frameworks 5 years later. child development, 84(5), 1526-1541. https://doi.org/10.1111/cdev.12064 guyer j.j., vaughan-johnston t.i. (2018) social comparisons (upward and downward). in: zeigler-hill v., shackelford t. (eds) encyclopedia of personality and individual differences. springer. https://doi.org/10.1007/978-3-319-28099-8_1912-1 hasselhorn, m. (2005). lernen im altersbereich zwischen 4 und 8 jahren: individuelle voraussetzungen, entwicklung, diagnostik und förderung. in t. guldimann & b. hauser (eds.), bildung 4 bis 8-jähriger kinder (pp. 77-88). waxmann. hauser, b. (2013). spielen. frühes lernen in familie, krippe und kindergarten. kohlhammer. http://www.ciando.com/ebook/bid-1473646 hellmich, f. (2011). selbstkonzepte im grundschulalter modelle, empirische ergebnisse, pädagogische konsequenzen . kohlhammer. helmke, a., & vanaken, m. a. g. (1995, dec). the causal ordering of academic achievement and self-concept of ability during elementary school: a longitudinal study. journal of educational psychology, 87(4), 624-637. https://doi.org/10.1037/0022-0663.87.4.624 hox, j. j. (2002). multilevel analysis: techniques and applications. erlbaum. hu, l.-t., & bentler, p. m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives.structural equation modeling: a multidisciplinary journal, 6(1), 1-55, https://doi.org/10.1080/10705519909540118 humberg, s., dufnert, m., schonbrodt, f. d., geukes, k., hutteman, r., van zalk, m. h. w., denissen, j. j. a., nestler, s., & back, m. d. (2018, jul 23). why condition-based regression analysis (cra) is indeed a valid test of self-enhancement effects: a response to krueger et al. (2017). collabra: psychology, 4(1). http://doi.org/10.1525/collabra.137 humberg, s., förster, n., kaiser, j., & schönbrodt, f. d. (2017). konsequenzen akkurater lehrerurteile: response-surface-analyse als statistisches verfahren zur untersuchung von übereinstimmungshypothesen. in a. südkamp & a.-k. praetorius (eds.), diagnostische kompetenz von lehrkräften (pp. 174-200). waxmann. hutchinson, l. r. (2013). young children's engagement in self-regulation at school (unpublished doctoral dissertation). university of british columbia. jacobs, j. e., lanza, s., osgood, d. w., eccles, j. s., & wigfield, a. (2002). changes in children's self-competence and values: gender and domain differences across grades one through twelve. child development, 73(2), 509-527. https://doi.org/10.1111/1467-8624.00421 lipko, a. r., dunlosky, j., lipowski, s. l., & merriman, w. e. (2012, 2012/04/01). young children are not underconfident with practice: the benefit of ignoring a fallible memory heuristic. journal of cognition and development, 13(2), 174-188. https://doi.org/10.1080/15248372.2011.577760 loveless, t. (2006). the 2006 brown center report on american education: how well are american students learning ? brookings institution press. https://www.brookings.edu/wp-content/uploads/2016/06/10education_loveless-1.pdf marsh, h. w., ellis, l. a., & craven, r. g. (2002, may). how do preschool children feel about themselves? unraveling measurement and multidimensional self-concept structure. developmental psychology, 38(3), 376-393. https://doi.org/10.1037//0012-1649.38.3.376 marsh, h. w., & martin, a. j. (2011, mar). academic self-concept and academic achievement: relations and causal ordering. british journal of educational psychology, 81(1), 59-77. https://doi.org/10.1348/000709910x503501 mcclelland, m. m., cameron, c. e., duncan, r., bowles, r. p., acock, a. c., miao, a., & pratt, m. e. (2014). predictors of early growth in academic achievement: the head-toes-knees-shoulders task. frontiers in psychology, 5, 599. https://doi.org/10.3389/fpsyg.2014.00599 measelle, j., ablow, j., cowan, p., & cowan, c. (1998). assessing young children’s views of their academic, social, and emotional lives: an evaluation of the self-perception scales of the berkeley puppet interview. child development, 69(6), 1556-1576. https://doi.org/10.2307/1132132 moffitt, t. e., poulton, r., & caspi, a. (2013, sep-oct). lifelong impact of early self-control: childhood self-discipline predicts adult quality of life. american scientist, 101(5), 352-359. https://doi.org/10.1511/2013.104.352 muijs, r. d. (1997, sep). predictors of academic achievement and academic self-concept: a longitudinal perspective. british journal of educational psychology, 67, 263-277. https://doi.org/10.1111/j.2044-8279.1997.tb01243.x müller, e., wustmann seiler, c., perren, s., & simoni, h. (2015). young children’s self-perceived ability: development, factor structure and initial validation of a self-report instrument for preschoolers. journal of psychopathology and behavioral assessment, 37 (2), 256-273. https://doi.org/10.1007/s10862-014-9447-9 narciss, s., koerndle, h., & dresel, m. (2011). self-evaluation accuracy and satisfaction with performance: are there affective costs or benefits of positive self-evaluation bias? international journal of educational research, 50(4), 230-240. https://doi.org/10.1016/j.ijer.2011.08.004 panadero, e. (2017). a review of self-regulated learning: six models and four directions for research. frontiers in psychology, 8, 422. https://doi.org/10.3389/fpsyg.2017.00422 perry, n. e. (2013). understanding classroom processes that support children’s self-regulation of learning. in d. whitebread, n. mercer, c. howe, & a. tolmie (eds .), self-regulation and dialogue in primary classrooms. british journal of educational psychology monograph series ii: psychological aspects of education current trends , 10, 45-68. perry, n. e., hutchinson, l. r., yee, n., & määttä, e. (2017). advances in understanding young children’s self-regulation of learning. in p. a. alexander, d. h. schunk, & j. a. greene (eds.), handbook of self-regulation of learning and performance (2nd ed., pp. 457-472). routledge handbooks online. https://doi.org/10.4324/9781315697048.ch29 perry, n., lisaingo, s. & ford, l. (2019). understanding the role of motivation in children's self-regulation for learning. in d. whitebread, v. grau, & k. kumpulainen (eds.), the sage handbook of developmental psychology and early childhood education (pp. 517-534). sage publications. https://doi.org/10.4135/9781526470393.n30 pinxten, m., de fraine, b., van damme, j., & d'haenens, e. (2010, dec). causal ordering of academic self-concept and achievement: effects of type of achievement measure. british journal of educational psychology, 80(4), 689-709. https://doi.org/10.1348/000709910x493071 praetorius, a.-k., kastens, c., hartig, j., & lipowsky, f. (2016). haben schüler mit optimistischen selbsteinschätzungen die nase vorn? zeitschrift für entwicklungspsychologie und pädagogische psychologie , 48(1), 14-26. https://doi.org/10.1026/0049-8637/a000140 roebers, c. m., cimeli, p., rothlisberger, m., & neuenschwander, r. (2012). executive functioning, metacognition, and self-perceived competence in elementary school children: an explorative study on their interrelations and their role for school achievement. metacognition and learning, 7(3), 151-173. https://doi.org/10.1007/s11409-012-9089-9 schönbrodt, f. d. (2016). testing fit patterns with polynomial regression models. osf preprints. https://doi.org/https://doi.org/10.31219/osf.io/ndggf schönbrodt, f. d., humberg, s., & nestler, s. (2018). testing similarity effects with dyadic response surface analysis. european journal of personality, 32(6), 627-641. https://doi.org/10.1002/per.2169 schunk, d. h., & greene, j. a. (2018). historical, contemporary, and future perspectives on self-regulated learning and performance. in d. h. schunk & j. a. greene (eds.), handbook of self-regulation of learning and performance (pp. 1–15). routledge/taylor & francis group. https://www.doi.org/10.4324/9781315697048 sektnan, m., mcclelland, m. m., acock, a., & morrison, f. j. (2010, dec). relations between early family risk, children's behavioral regulation, and academic achievement. early childhood research quarterly, 25(4), 464-479. https://doi.org/10.1016/j.ecresq.2010.02.005 shavelson, r. j., hubner, j. j., & stanton, g. c. (1976). self-concept: validation of construct interpretations. review of educational research, 46(3), 407-441. https://www.doi.org/10.3102/00346543046003407 shin, h., bjorklund, d. f., & beck, e. f. (2007). the adaptive nature of children's overestimation in a strategic memory task. cognitive development, 22(2), 197-212. https://doi.org/10.1016/j.cogdev.2006.10.001 skaalvik, e. m., & hagtvet, k. a. (1990). academic achievement and self-concept: an analysis of causal predominance in a developmental perspective. journal of personality and social psychology, 58(2), 292-307. https://doi.org/10.1037/0022-3514.58.2.292 taylor, s. e., & brown, j. d. (1988, mar). illusion and well-being: a social psychological perspective on mental health. psychological bulletin, 103(2), 193-210. https://doi.org/10.1037/0033-2909.103.2.193 taylor, s. e., kemeny, m. e., reed, g. m., bower, j. e., & gruenewald, t. l. (2000, jan). psychological resources, positive illusions, and health. american psychologist, 55(1), 99-109. https://doi.org/10.1037//0003-066x.55.1.99 timmons, k., pelletier, j., & corter, c. (2016, 2016/02/01). understanding children's self-regulation within different classroom contexts. early child development and care, 186(2), 249-267. https://doi.org/10.1080/03004430.2015.1027699 valentine, j. c., dubois, d. l., & cooper, h. (2004). the relation between self-beliefs and academic achievement: a meta-analytic review. educational psychologist, 39(2), 111-133. https://doi.org/10.1207/s15326985ep3902_3 whitebread, d., bingham, s., grau, v., pino pasternak, d., & sangster, c. (2007). development of metacognition and self-regulated learning in young children: role of collaborative and peer-assisted learning. journal of cognitive education and psychology, 6(3), 433-455. https://doi.org/10.1891/194589507787382043 wurster, s. & feldhoff, t. (2019). schulund unterrichtsqualität aus der mehrebenenperspektive: ist die schule oder klasse die relevante pädagogische gestaltungseinheit? zeitschrift für pädagogik, 65(1), 24-39. young-hoon, k., chi-yue, c., & zhimin, z. (2010). know thyself: misperceptions of actual performance undermine achievement motivation, future performance, and subjective well-being. journal of personality and social psychology, 99(3), 395-409. https://doi.org/10.1037/a0020555 zimmerman, b.j., & schunk, d.h. (2008). motivation: an essential dimension of self-regulated learning. in d. h. schunk & b. j. zimmerman (eds.), motivation and self-regulated learning (pp. 1-30). routledge. https://doi.org/10.4324/9780203831076 appendix table 4 regression coefficients b1 to b5 and derived model parameters for the full polynomial model, the shifted and rotated rising ridge (srrr) model and the additive model for behavioural sr note. srrr model = shifted and rotated rising ridge model; regression coefficients b1 b5: b1 = academic self-concept; b2 = ability level; b3 = academic self-concept2; b4 = interaction; b5 = ability level2; c = lateral shift of the ridge; s = rotation of the shift; bm = mean effect; a'4 = curvature orthogonal to the ridge; ci = confidence interval. confidence intervals and p-values are derived from a percentile bootstrap with 10,000 replications. table 5 regression coefficients b1 to b5 and derived model parameters for the full polynomial model and the interaction model (ia) for goal orientation   note. regression coefficients b1 b5: b1 = academic self-concept; b2 = ability level; b3 = academic self-concept2; b4 = interaction; b5 = ability level2; ci = confidence interval. confidence intervals and p-values are derived from a percentile bootstrap with 10,000 replications. frontline learning research, special issue: vol. 13 no. 2 (2025) 102 121 issn 2295-3159 corresponding author: eetu haataja, p.o. box 8000, fi-90014 university of oulu, finland. eetu.haataja@oulu.fi doi: https://doi.org/10.14786/flr.v13i2.1315 a momentary view of engagement in collaborative learning: triangulation through multimodal data eetu haataja12, tiina törmänen1, matthew p. somerville3, jonna malmberg1, hanna järvenoja1& sanna järvelä1 1 learning and educational technology research lab, faculty of education and psychology, university of oulu, oulu, finland 2 research unit of population health, faculty of medicine, university of oulu, oulu, finland 3 department of psychology and human development, ucl institute of education, london, uk article received 30 june 2023 / article revised 1 september 2023 / accepted 20 october 2023 / available online 14 march 2025 abstract despite recognising momentary challenges while learning, collaborative groups do not necessarily regulate and adapt their learning process according to the demands. various online measures have recently been explored to unobtrusively study engagement and adaptation in collaborative learning (cl), as it occurs in the classroom. for example, physiological synchrony derived from electrodermal activity (eda) has been a prominent reflector of momentary engagement in cl. however, how physiological synchrony relates to students’ views about cl, regulation of learning, and performance remains unclear. this study investigates how momentary measures of physiological synchrony, students’ perceived value of cl, and regulation of learning, align and further relate to group performance. the participants were 94 students attending a physics course consisting of four 90-minute lessons and a collaborative exam. each lesson included a cl task. at the beginning and end of each session, students reported their perceived value of cl. students’ eda was recorded to derive physiological synchrony. coregulation (corl) and socially shared regulation (ssrl) were coded from the video. results suggest that when groups show higher physiological synchrony, they perceive their cl as less valuable and tend to perform worse in collaborative exams. it seems that self-reports on the value of cl, rather than physiological synchrony, may better reflect the regulation of cl. interestingly, the association patterns for corl and ssrl differed, as frequent corl was linked to the less valued cl, while ssrl tended towards a positive relation. the study demonstrates the complex and multidimensional role of momentary engagement in cl. keywords: collaborative learning, momentary engagement, socially shared regulation of learning, physiological synchrony mailto:eetu.haataja@oulu.fi haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 103 | flr introduction collaborative learning (cl) requires students' continuous active participation to form a shared understanding of a problem, which makes student engagement a fundamental prerequisite for successful cl to unfold (roschelle & teasley, 1995). recent research has emphasised the momentary and situated nature of classroom engagement (symonds et al., 2021). symonds et al. (2024) define momentary engagement as “the active involvement of an individual in an educational task as the task proceeds across seconds and minutes” (p. 5) involving cognitive, motivational, behavioural, and emotional components. such temporality and multidimensionality make momentary engagement a complex dynamic system that requires self-regulatory control mechanisms, allowing its self-organisation over time (symonds et al., 2024). in cl research, engagement intertwines with self-regulated learning processes, including the active planning, monitoring, and control of learning (järvelä et al., 2016; sinha et al., 2015). more specifically, in cl, students can engage in self-regulation (srl), co-regulation (corl), and socially shared regulation of learning (ssrl) according to their situational needs (hadwin et al., 2018). where srl is more about the individual planning, monitoring, and controlling of one’s learning process, ssrl refers to a type of regulation where group members negotiate, transact, and build their regulation upon each other (järvelä et al., 2021). corl, in turn, holds more of a transitional role, switching and activating either the individual’s srl or the group's ssrl during the learning process. for example, a student in a collaborative group can point out that the task is not progressing as planned, allowing the individual or the group to monitor and control the learning process towards the goal. though momentary engagement can be considered a broader concept than the regulation of learning, in this study, we see the regulation of collaborative learning as an important proxy for momentary engagement. prior studies suggest that students’ momentary engagement, srl, corl, and ssrl are challenging to capture, and multiple methods and data channels are needed to study them (azevedo, 2015; sobocinski et al., 2020). recently, martin et al. (2023) proposed that engagement, motivation, and learning should be studied with an integrated approach and that physiological measures could offer tools to conduct such research. we argue that this approach is particularly relevant to momentary engagement, as it is a temporal multilevel phenomenon manifesting on different grain sizes (symonds et al., 2024). promising results suggest that electrodermal activity (eda) measures and different dimensions of engagement could be positively linked (lee et al., 2019; malmberg et al., 2023; zhang et al., 2018). furthermore, eda measured in collaborative settings allows computation of physiological synchrony: interdependence between group members’ physiological signals. physiological synchrony has gained interest in revealing the dynamics of collaboration and has been considered potentially a relevant condition for the regulation of cl. maybe surprisingly, physiological synchrony derived from eda appears to be higher when the group faces challenges and needs to make an effort to proceed with their task (dindar et al., 2020; malmberg et al., 2019). due to these results, physiological synchrony could be considered a potential measure to indicate momentary engagement during challenging cl events. the current research views students’ subjective appraisals of collaboration as an important prerequisite for momentary engagement in cl. however, momentary engagement can also be shaped by actualised learning activities, such as adaptation to learning challenges through corl and ssrl during the learning process. therefore, we study how observed corl and ssrl relate to students’ situated views about collaboration, physiological activity, and collaborative learning outcomes. very few studies have empirically investigated how different subjective and objective variables relevant to momentary engagement in cl align. this study explores whether subjective and objective measures recorded from natural classroom learning settings align and further relate to cl outcomes. the present study is also essential in empirically demonstrating the complex system nature that momentary engagement (symonds et al., 2024) and cl involve (amon et al., 2019; ouyang et al., 2023). more specifically, it considers the complexity and temporality of momentary engagement by applying novel nonlinear analysis methods (multidimensional recurrence quantification analysis) and acknowledges the haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 104 | flr lesson-to-lesson, within-group variation when analysing the relations. the results contribute to a more comprehensive understanding of the dynamics emerging in cl. this holistic view can inform the creation of interventions and tools that promote momentary engagement and help students regulate their learning effectively in group settings. background momentary engagement and regulation in collaborative learning being actively involved in a task across seconds and minutes is critical for learning, particularly in complex and dynamic learning environments such as cl (isohätälä et al., 2017; järvelä et al., 2016; symonds et al., 2024). in cl, momentary engagement may not necessarily target only the task content itself but also all the social interactions relevant to carrying out the task. engaging in these interactions is a fundamental prerequisite for students to build up their reasoning and argumentation with each other (cohen, 1994) and for supporting momentary engagement in the task (lai, 2021). engagement in learning is inherently multidimensional. this has been acknowledged for individual engagement (fredricks et al., 2004; sinatra et al., 2015), collaborative engagement (rogat et al., 2022; sinha et al., 2015), and recently also within a momentary engagement framework (symonds et al., 2021). rogat et al. (2022) distinguish between behavioural, socioemotional, collaborative, metacognitive, and disciplinary dimensions of engagement in cl. the core of cl is the idea of building a shared understanding of a problem through the contributions of multiple learners (roschelle & teasley, 1995). this requires collaborative engagement, that is, the coordination of task and knowledge construction processes, as well as joint and balanced contributions of group members. however, this also requires sustaining learners' joint participation, attention and focus on the task at hand (i.e., behavioural engagement), which can be challenging in group settings (rogat et al., 2022). in addition to cognitive processes (e.g. attention), momentary engagement is intertwined with motivational and socio-emotional processes (rogat et al., 2022; symonds et al., 2021) that serve as important conditions for students’ willingness to engage in the task at hand (bakhtiar et al., 2017; upadyaya et al., 2021). for example, lai’s (2021) study indicated that interest and utility value are positively associated with students’ behavioural engagement in a group-based flipped learning context. zschocke et al. (2016) studied individuals’ group work appraisal. they found that appraisals of the cognitive benefits of group work (i.e., the value of collaboration) were a significant predictor of positive activating emotions, considered to promote learning. however, momentary engagement in cl can also be pivotal in promoting the perceived value of collaboration, starting a positive self-reinforcing feedback loop supporting future engagement (lai, 2021, malmberg et al., 2023). this type of interplay manifests, for example, when students engage in discussion, reciprocally offering their ideas and perspectives, which builds a more comprehensive and nuanced understanding of the topic (vuopala et al., 2019; zabolotna et al., 2023). this, in turn, can feed back to students perceiving the collaboration as more valuable for their learning and, thus, support their momentary engagement for the continuing learning activities (kreijns et al., 2003). by actively engaging with each other and sharing their thoughts and ideas, learners can simultaneously build a shared sense of purpose and common ground (isohätälä et al., 2020). furthermore, as a result, a positive socioemotional atmosphere and sense of community among learners can help to foster engagement through its socio-emotional dimension (järvenoja & järvelä, 2013; mänty et al., 2020; rogat et al., 2022). from the different dimensions of engagement, the metacognitive dimension appears to be particularly central and strongly connected with other dimensions (rogat et al., 2022). for example, being actively involved with a task does not necessarily mean the collaboration process unfolds productively (sinha et al., 2015). this is to say that the group can be momentarily engaged with the task and still struggle to make progress, solve the problem and learn together. prior research has shown that haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 105 | flr students need to regulate, that is, to plan, monitor and control their learning process to overcome challenges and succeed (haataja, malmberg, et al., 2022; järvelä & hadwin, 2013). regulation demands metacognition, thinking about one’s thinking, adapting and making strategic decisions when facing challenges, all of which ultimately result in better learning (järvelä et al., 2021). according to rogat et al. (2022), group use of regulatory strategies indicates metacognitive engagement. in particular, highquality metacognitive engagement can be characterised by goal-focused, effective ssrl (rogat et al., 2022). that is, the regulation of learning relies on momentary metacognitive engagement (isohätälä et al., 2017; vuorenmaa et al., 2023). for instance, in srl, momentary engagement allows learners to monitor their level of attention and motivation, judge their performance, and adjust their learning strategies accordingly. in corl, momentary engagement enables students to perceive and respond to the actions of their peers, adjust their behaviour to support the group's goals and facilitate the emergence of shared regulation. in ssrl, momentary engagement facilitates the negotiation of shared goals and strategies among group members, the monitoring of obstacles, and the proposals and decisions of shared control. however, the relationship between momentary engagement and regulation is reciprocal: regulation also serves as a means to strategically influence and ensure the self-organisation and maintenance of momentary engagement in the face of challenges (järvelä et al., 2016; symonds et al., 2024). for example, corl can serve as an invitation to the group as a whole to momentarily engage and negotiate their decision-making regarding learning in the form of ssrl (ahola et al., 2023), which, however, does not always happen (törmänen et al., 2021). ssrl, in turn, can mutually support the groups’ engagement in executing the strategic decisions and, importantly, the motivational and socioemotional atmosphere, maintaining the group's engagement (törmänen et al., 2021). there is also preliminary evidence suggesting that corl and ssrl could relate to better group performance (de backer et al., 2020; zheng & huang, 2016) and individual learning achievement (haataja, dindar, et al., 2022; zheng et al., 2017). physiological data reflecting the momentary engagement in collaborative learning though the critical roles of engagement and regulation of learning have been acknowledged for some time, the challenge has been to track engagement in situ (azevedo, 2015). recently martin et al., (2023) proposed that engagement, motivation, and learning should be studied with an integrated approach and that biophysiology could offer a framework for such research. physiological data is particularly interesting for momentary engagement research because it can be traced through seconds and minutes. although physiological data entered the learning sciences quite recently, studies have demonstrated that autonomic nervous system measures such as heart rate (hr) and electrodermal activity (eda) are potential channels for studying engagement (ba & hu, 2023). more specifically, eda reflects the sympathetic branch of the autonomic nervous system, traditionally considered to activate in the fight or flight stress response, signalling the anticipated need to act and engage in a situation at hand (dawson et al., 2017). previously, lee et al. (2019) investigated how eda measures reflect the momentary engagement of students during a maker movement course. they found a moderate correlation between the eda measures and cognitive and behavioural engagement. however, emotional engagement did not correlate meaningfully with eda peaks. the authors suggested that more research would be needed to show the relevancy of eda measures in studying momentary engagement. further, zhang et al. (2021) examined the correlations between eda indicators of emotional and cognitive engagement, selfreported class engagement, and learning outcomes; and how physiological dynamics differ across students and within students over time with classroom events and observable individual engagement. they found moderate positive correlations between individual-level eda features and self-reported attention, emotional valence, and knowledge. the synchrony between the students’ physiological signals is an interesting phenomenon for investigating momentary engagement, particularly in cl. physiological synchrony refers to any haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 106 | flr interconnected or associated activity among physiological signals between individuals (palumbo et al., 2018), often captured by temporal measures to trace the similarity in people's physiological activity patterns, such as eda responses over time. considering that shared joint attention is a fundamental prerequisite for cl, one might hypothesise that better collaboration would be reflected as a similar physiological activity between the students and, therefore, as higher synchrony. however, so far, the empirical results regarding physiological synchrony derived from eda measures seem to be mixed. some studies have linked higher synchrony with higher learning gains (pijeira-díaz et al., 2016) and better collaboration (montague et al., 2014). in contrast, in our studies, synchrony was higher when the students experienced tasks challenging and demanding more mental effort (dindar et al., 2020) and in situations where they discussed challenges in their collaboration (malmberg et al., 2020). other researchers have found similar results. for example, sung et al. (2023) recently studied eda-based physiological synchrony by comparing two pedagogical settings – direct instruction and hands-on learning. they found that some physiological synchrony metrics (ssi and pc) were negatively associated with learning gains. the authors suggest that the relationship between eda and physiological synchrony to learning gains and behavioural engagement is context-dependent. physiological synchrony metrics tended to be higher when instructors and learners worked on the same task. further, previously schneider et al. (2020) found increasing physiological synchrony to be linked with a lack of consensus in a collaborating group. these results align with our recent qualitative case study (malmberg et al., 2023), where increased physiological synchrony was related to moments where shared momentary engagement was demanded from the group members, for example, when they aimed to finalise the task. in summary, though physiological synchrony can potentially also occur in situations where eda is low, preliminary evidence suggests it is likely to increase when the group struggles with the task. based on the accumulated evidence, our preliminary hypothesis is that continuous higher physiological synchrony during cl may reflect challenges in collaboration or the task and shared mental effort when the students are trying to move towards the solution. at the same time, these results might be specific for physiological synchrony derived from eda data, reflecting the sympathetic nervous system activity and a need to act and engage when facing challenges. in the present study, we had a preliminary hypothesis that physiological synchrony derived from eda could signal that the groups were making a shared effort to solve the challenges. especially if such challenges continuously persist, this could be seen as higher average physiological synchrony indices for each session. the other measures related to momentary engagement during the collaborative task could potentially reflect these challenges. for example, since ssrl often involves the group adapting in challenging situations strategically (hadwin et al., 2018), groups lacking such adaptive processes could show higher physiological synchrony on average (mønster et al., 2016). to date, most studies have applied these measures for a single session analysing relations based on between-group differences. this study also considers the within-group variation in synchrony throughout the physics course. this is consistent with the momentary engagement framework, where the temporal aspect is central. eda and physiological synchrony are still novel measures in cl research. therefore, this exploratory study has the potential to contribute to the development and application of novel temporally intensive online measures in learning sciences. aim this study investigated how observed (corl and ssrl), self-reported (perceived value of cl), and physiological measures aligned with each other during cl, how they related to students' momentary engagement, and how these momentary measures related to group performance. the research questions were as follows: rq1. how do students’ perceived value of cl, observed regulation of learning, and physiological synchrony relate to each other? haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 107 | flr rq2. how do the perceived value of cl, observed regulation of learning, and physiological synchrony relate to performance in a collaborative exam? methodology a total of 94 seventh-grade students (∼13 years of age, 58 females, 36 males) attended their first secondary school physics course, focused on the topics of light and sound. participation in the study was voluntary for the students. each 90-minute lesson involved cl tasks where the students worked in groups of two to four, making 30 groups. students had prior experience of learning collaboratively. teachers of the course formed groups intending to make their composition as homogeneous as possible based on prior learning achievement of the students in other stem courses. the groups remained the same throughout the course, apart from normal student absences (e.g., sick leave) occurring in classrooms. based on flipped classroom principles, the students independently studied the upcoming topic in their science textbook before each lesson (järvenoja et al., 2020). at the beginning of each class, the teacher introduced the new topic to the students and ensured that each student had enough knowledge to engage in cl. the introduction was followed by cl tasks to co-construct a more profound and shared understanding of the topic. each collaborative task included hands-on scientific experiments. the content and exercises were designed with assistance from science teachers to ensure that they covered the required subjects and content. in one learning task, for example, the task was to perform experiments on light and sight. the groups were provided with four main themes for investigation (1. illumination, 2. intensity, 3. propagation, and 4. reflection of light). a flashlight, related materials, and instructions were provided to the groups. before the groups started to work on the task, they were asked to briefly discuss the following prompting questions: “what are the collaborative goals for your group?” and “what will you do to achieve your goals?” after the collaborative tasks, the groups were asked to discuss the question: “did you achieve your collaborative goals? why?” this short reflection was followed by a quiz the students took individually. a more detailed context description has been reported in a separate, nonempirical article focused on the study design (järvenoja et al., 2020). before and after each collaborative session, students answered situated self-reported statements regarding their cl. the statement regarding the perceived value of collaboration was, “working in a group helps me to learn”, which the students assessed using a scale of 0-100. the statement was adapted from the contextualised saga instrument (volet, 2001). at the end of the course, students took a collaborative exam. the collaborative exam was codesigned with the physics teachers of the course, and it focused on the topic the teachers considered the most challenging in the course: the refraction of convex and concave lenses. in the exam, student groups first had twenty minutes to explore a simulation of light refraction with a computer and write notes. after that, they were given seven different refraction cases where the distances of the objects from the refracting lenses varied. in their exam answer, the students were asked to draw the lines of refracting light in each case and define the quality, direction, and size of the refracted image with a short text answer. overall, there were 28 items to be answered. physics teachers evaluated the exam using a scale from 4 to 10 with a 0.25-point increment (m = 7.4, sd = 1.2, min = 5.5, max = 9.5). analysis of the physiological data students’ electrodermal activity was measured with shimmer gsr3 + sensors. due to the limited number of sensors, only 27 groups from 30 wore the sensors. the remaining three groups wore another sensor type, which could not be used in further analyses due to differences in the measures. the haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 108 | flr data were first downsampled from 128 to 16 hz in the preprocessing phase to accelerate the analysis. next, based on a visual inspection, the eda data recordings with missing electrode contact were removed from the data. furthermore, a butterworth low-pass filter (frequency 1, order 5) was applied to remove small movement artefacts from the signal (kelsey et al., 2018). ledalab software with continuous decomposition analysis was used to differentiate the rapidly changing phasic signal component from the more slowly reacting tonic component (benedek and kaernbach, 2010). the measures of physiological synchrony aimed to grasp the interdependence in physiology between individuals. in line with previous studies (e.g. mønster et al., 2016), this study used the phasic eda signal component, further downsampled to 4hz, to calculate the synchrony (mendes, 2009). to quantify the physiological synchrony, we used multidimensional recurrence quantification analysis (mdrqa; coco et al., 2021), one of the few suitable methods for more than two signals (wallot, roepstorff, et al., 2016). mdrqa is a nonlinear time series analysis method that assesses patterns of synchrony between two or more time series, which do not need to be stationary. the base of mdrqa statistics is a recurrence plot, which graphically displays the temporal dynamics of a multidimensional phase space of a system. for instance, the recurrence rate (rr%) derived from the recurrence plot indicates how many individual elements between the signals are shared (wallot, roepstorff, et al., 2016). in this case, the plots and resulting rr% index representing synchrony were calculated separately for each group on each session. therefore, the resulting values represent the average level of synchrony for each group in each session. the parameters for running the mdrqa analysis were decided based on suggestions in the rqa literature (wallot, mitkidis, et al., 2016). first, the delay (del) parameter was estimated using the average mutual information function for each individual eda signal. second, the false nearest neighbour function was used for each eda signal to estimate the embedding dimension parameter. with both functions, the first local minimum was determined for each signal and then averaged and rounded up for all the signals. in this case, the resulting values were divided by two, as the signals were embedded together in the mdrqa. this means that only some of the dimensions had to be reconstructed by timedelayed embedding because they were available as separately measured signals (wallot, mitkidis, et al., 2016). as a result, the parameters used were delay = 35 and embedding = 2. the radius parameter was set to 0.2, keeping most rr% values between the suggested 1-10% (wallot & leonardi, 2018). for the rr%, outliers more than two standard deviations from the mean were removed before the statistical analysis. analysis of the video data the groups’ collaborative work was videotaped with 360-degree cameras (90 min per session/week = 225 hours). the video data were segmented into 30-second episodes, from which two coders identified the occurrence of corl, and ssrl in verbal exchanges during task execution. both coders were involved in the refinement of the coding system and engaged in qualitative discussions about specific codes, as the coding scheme was clarified with the researchers. the coding reliability was ensured by selecting 10% of the video data to be coded by both coders. what differentiated corl and ssrl codes from other group interactions was that, in regulation, students had to clearly express the observation of an obstacle or a challenge in their learning process (e.g., a lack of task understanding) and involve a regulatory initiation/control (e.g., a suggestion to reread the task instruction), which led to a strategic change in the groups’ action (e.g., rereading the task instruction together). corl and ssrl codes were mutually exclusive. in corl, no additional strategic content from other group members followed the initiation of regulation, which meant that the verbal interactions of corl did not build upon each other. in contrast, ssrl involved the reciprocally negotiated participation of several group members in regulatory discussion, where the interaction between group members built upon each other and led to strategic changes in the learning process. in practice, when coded as ssrl, the initiation of regulation needed to be followed by additional strategic content or control attempts from other group haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 109 | flr members aiming to solve the challenge (e.g., agreeing or articulating the challenge was not enough). when corl or ssrl targeted emotions and motivation, it could also be coded without a clear obstacle or change in action, including the strategic activities to maintain or strengthen the already favourable motivational and affective conditions. however, mere emotional or motivational expressions were not coded as regulation. instead, regulation needed to be strategic and purposefully targeted to alter the emotional and motivational state of the group with appropriate strategies (see e.g., lobczowski et al., 2021; mänty et al., 2023). the inter-rater reliability of the regulation coding yielded cohen’s kappa of 0.79. the frequencies of identified corl and ssrl for each group and session were counted. statistical analysis only the sessions involving three students were included in further statistical analyses because the number of participants in the interaction was considered to affect group dynamics (cen et al., 2016), which might have influenced the occurrence of corl and ssrl and mdrqa values (wallot et al., 2016). this restriction resulted in data from 26 groups and 78 students involving 82 recordings with 5576 (2788 minutes) collaboration video segments, from which 244 recordings (122 minutes) included corl, and 62 (31 minutes) included ssrl. because there was variation in the collaboration length (m = 34.0 minutes, sd = 8.9) between the sessions, a proportion of corl and ssrl from collaboration for each session was calculated by dividing the corl and ssrl frequency by the overall number of segments in that collaborative session (see table 1.). for the self-reported perceived value of collaboration, group average values were calculated. the resulting data set involved 89 group observations, meaning that 66% of the potential complete data were included in the analysis. missing data is a limitation in the study but also a reality when collecting momentary data in real classroom settings. because the data regarding the first research question involved repeated measures, meaning that the observations were not independent, the relations between variables in the first research question were analysed with repeated measures correlation analysis (rmcorr r-package; bakdash & marusich, 2017). rmcorr considers the non-independence among repeated observations using analysis of covariance to adjust for inter-individual variability. it fits parallel regression lines with varying intercepts for each cluster (e.g. group), making it conceptually close to the multilevel “random intercept only model” (see figure 1b). due to the non-normal distribution of the data, we used bootstrapping (1000 samples) to more robustly estimate the confidence intervals for rmcorr analysis (bakdash & marusich, 2017). for the second research question, we first calculated course average values for each group for each variable. regarding the perceived value of collaboration, preand post-measures were combined to one average. after that, spearman's correlational analyses were conducted to investigate the relationships between the averaged variables and group exam scores. the statistical significance for both research questions was adjusted based on the benjamini–hochberg procedure (1995) to avoid false discoveries due to multiple comparisons. table 1 descriptive statistics of observations averaged for each group on each session m sd min max perceived value of collaboration (pre) 61.49 15.77 19.00 100.00 perceived value of collaboration (post) 61.46 18.24 8.00 100.00 perceived value of collaboration (δ) 0.39 13.31 –49.00 32.67 haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 110 | flr physiological synchrony (rr%) 2.34 3.21 0.09 18.14 socially shared regulation (0–1) 0.01 0.02 0.00 0.23 co-regulation (0–1) 0.03 0.04 0.00 0.23 results alignment of momentary data regarding research question 1 we found (see correlogram in figure 1a.) students’ selfreported perceived value of collaboration after the session to be related to all other variables relevant to the momentary engagement. from these relations, the perceived value of collaboration (post) was moderately negatively linked to physiological synchrony rrm = –.47, ci [–.74, –.03], and corl, rrm = – .4, [–.71, –.03]. in contrast, ssrl tended to relate positively with the perceived value of collaboration, but the finding did not remain statistically significant after the adjustment for multiple tests rrm = .25, [– .07, .59]. figure 1b further demonstrates the importance of session-by-session within-group analysis in the case of the relationship between the perceived value of collaboration (post) and the physiological synchrony, as the intercepts between groups (vertical levels of the coloured lines) differ remarkably. this would imply that just focusing on the between-group differences would not have given a realistic picture of the relationship between the variables. overall, the perceived value of collaboration measured before and after the session followed a similar pattern with other variables. however, for the post-measure, the relations tended to be stronger. also, the relationships between the change in the perceived value of collaboration with other variables followed a similar pattern, though the results did not reach statistical significance. physiological synchrony showed a negative relation with before and after collaboration self-report but no link with observed corl or ssrl. notably, in repeated measure correlational analysis, corl and ssrl were not related to each other rrm = .04, [–.16, .22]. haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 111 | flr figure 1. repeated measures correlations determining the common within-group association for paired variables assessed on multiple lessons for multiple groups (a) and an example of repeated measure correlation plot (b) between the perceived value of collaboration (post) and physiological synchrony. colours represent groups. * p < .05 (adjusted for multiple tests with benjamini–hochberg procedure.) the relationship between momentary measures and group performance regarding research question 2, spearman rank order correlations show (figure 2) that the average physiological synchrony during the course showed a strong to moderate negative association with a group performance in the collaborative exam, r (24) = –.56, p = .005. no other variables relevant to momentary engagement showed a significant association with group performance. it is, however, also notable that socially shared regulation showed signs of a moderate positive correlation with the group exam score without reaching statistical significance, r (24) = .36, p. = .07. in contrast to the sessionbased repeated measures analysis, a course level between group analysis showed a tendency of a positive association between corl and ssrl r (24) = .45, p. = .02, but after the adjustment of multiple comparisons the finding did not remain statistically significant. haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 112 | flr figure 2. a correlogram presenting spearman correlation values between the course average values of each variable with group exam score. * p < .05 (adjusted for multiple tests with benjamini-hochberg procedure.) discussion despite the large body of literature examining cl and engagement in education (e.g. fredricks et al., 2004; järvelä et al., 2016; rogat et al., 2023; sinatra et al., 2015), very few studies have examined how different subjective and objective variables relevant to momentary engagement in cl align with each other. this study supports the view that multiple data channels can contribute to a better understanding of the multidimensional nature of momentary engagement in cl settings (azevedo, 2015). further, the study demonstrates that although groups show some lesson-to-lesson consistency in their measures, it is important to consider the temporal momentary within-group variation when analysing engagement in cl (symonds et al., 2024). one of the major findings of this study is that an important prerequisite of momentary engagement, the students’ perceived value of cl (salmela-aro et al., 2021; upadyaya et al., 2021), relates to other dimensions close to engagement in cl: negatively with physiological synchrony and corl and positively with ssrl. this finding implies that students’ situated subjective view about the value of collaboration reflects the overall picture of engagement in cl and should not be ignored as a measure, especially given that it is relatively easy to capture in classroom settings. temporally, the perceived value measured after the collaboration shows the most robust associations, but the measure before the collaboration seems to follow a similar pattern. this result could imply that students hold expectations about how useful it is to engage in collaboration, which aligns but does not determine the perception after the collaboration. furthermore, the situated self-report seems to benefit from the analysis and interpretation with the within-group lens, as the between-group approach is likely to miss some of the associations. for example, even though the situated self-report reflected the cl process in other measures, it was not related to the performance in the group exam when averaged to betweengroup analysis. haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 113 | flr prior research has emphasised the importance of momentary engagement for srl, corl, and ssrl in cl (järvelä et al., 2016; malmberg et al., 2023). notably, for corl, the association with the momentary perceived value of collaboration was negative. as conceptualised in this study, the function of corl is often to invite the momentary engagement of other members in the group when a challenge is faced, monitored, and when strategic control is needed (hadwin et al., 2018). therefore, constant corl can imply that ssrl and shared engagement to regulation is not reached (vuorenmaa et al., 2023). these results regarding corl might mean that some group members made repeated attempts to involve their peers in shared regulation without a fruitful result. previous research suggests that reasons for some members not engaging in regulation in these situations may be due to motivational (ahola et al., 2023) or emotional factors (törmänen et al., 2021). identifying these types of recurring dynamics in groups could allow timely interventions to move the momentary engagement of a group as a complex dynamic system to new types of states (van eijndhoven et al., 2023) involving, for example, ssrl. correspondingly, repeated occurrence of ssrl implied engagement of multiple group members in the regulation and showed a tentative positive association with the higher perceived value of cl. a pattern aligning with this also occurred concerning group performance, as ssrl showed a tentative positive association with the group exam score, whereas the case was the opposite for corl. though, based on these non-significant results, the relationship between ssrl and the perceived value of collaboration remains speculative, we believe it deserves to be studied further. on the one hand, groups that perceive collaboration as valuable for their learning might be more likely to share their regulatory process through engagement in negotiating strategic actions. on the other hand, if the group can, through ssrl, strategically solve the challenges faced, that might make the students perceive cl as more valuable, promoting a positive feedback loop of momentary engagement as a dynamic system (symonds et al., 2024). future studies could focus on this potential relationship and further temporally see if it is the reciprocally engaging form of regulation, namely ssrl, that promotes student perception about the value of collaboration, or is it the student’s expectation towards the value of collaboration that results in more ssrl? the physiological data used in this study paints a complex picture of the engagement indicators in cl. first, as expected, physiological synchrony between group members does not simply imply that the collaboration is progressing well. in fact, in eda-derived physiological synchrony measures, our results join the prior research findings, suggesting that the case might be the opposite (haataja, malmberg et al., 2022) and higher continuous physiological synchrony could reflect higher effort (dindar et al., 2020) and possible challenges in the group (malmberg et al., 2023). in this study, higher average synchrony during collaboration was related to students' lower self-reported collaboration value and lower group performance, aligning with other studies regarding physiological synchrony (sung et al., 2023; mønster et al., 2016). this negative link could be due to eda reflecting sympathetic nervous system activity, which has been linked to stress. previous research has also described cases where the physiological synchrony in cl arises momentarily, especially when joint efforts and engagement to solve the challenge are needed (malmberg et al., 2023). after the problem is solved, for example, through strategic regulation, the physiological synchrony tends to get back to a lower level (mønster et al., 2016). with dyads, the cycles between low and high levels of synchrony have been found to be linked with outcome measures (schneider et al., 2020). such results, coupled with the current study's findings, suggest that constant high physiological synchrony may indicate the group is engaged in collaboration across seconds and minutes but might struggle with socioemotional or cognitive challenges the group cannot surpass. from a dynamic systems perspective of momentary engagement (symonds et al., 2024), this could mean that the group is stuck in some unproductive attractor state. therefore, these groups could also benefit from targeted regulation support. however, it should be acknowledged that physiological synchrony can also occur during positive events and events unrelated to learning (slovák et al., 2014), and therefore, higher physiological synchrony signalling groups “struggling” could also be partly context-specific. to conclude, it is evident that the relationship between engagement and physiological synchrony is not direct nor straightforward. physiological haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 114 | flr synchrony is one potential temporal indicator that must be combined with other situational and temporal indicators, including the group members' subjective premises and appraisals. limitations and future studies despite the strengths of the current study and its novel approach to examining cl, it also has several limitations. firstly, the small sample size may limit the generalizability of the study's findings. secondly, missing data could have reduced the accuracy of the study's findings. future research should address this by implementing methods to reduce the amount of missing data, such as improving data collection techniques. thirdly, the non-normal data distribution may affect the validity of the statistical analyses. to overcome this limitation, we utilised methods where such an assumption is not central. fourthly, a more fine-grained temporal investigation of engagement is needed. future research should include more frequent and longitudinal measurements of momentary engagement to capture its dynamic nature. this will provide a more comprehensive understanding of how engagement changes over time and what factors might influence this change. finally, the study would benefit from a more intra-individual and situated approach. future research should explore individual engagement experiences in a group and the situational factors that may influence these experiences. conclusion this study supports the view that momentary engagement in cl is a complex multidimensional phenomenon that unfolds over time. this form of engagement appears to be shaped by the interaction of various factors, including interpersonal physiology, situated value appraisals of cl, and the regulation of cl. multiple measures offer a way to explore the (mis)alignment of these dimensions and better understand how the engagement of a group formed by individuals develops over time. physiological data offers a potential channel, and physiological synchrony is a potential index for studying momentary engagement in cl. however, more research is needed to validate the use of these measures for cl support. the study further highlights a need to consider both within-group and between-group variation when studying momentary engagement in authentic cl settings. haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 115 | flr keypoints we explore how different data channels can shed light on the multidimensional nature of momentary engagement in cl. results support the central role of the perceived value of collaboration for momentary engagement in cl. when groups show high physiological synchrony, they perceive their cl as less valuable and tend to perform worse in collaborative exams. frequent corl was linked to less valued cl, while the case tended to be the opposite for ssrl. methodologically, physiological synchrony can reflect collaboration processes in authentic learning situations. acknowledgments this work was supported by the academy of finland [grant numbers, 297686 (hj) and 308809 (jm)] and university of oulu (sj). leaf research infrastructure has been used in the data collection of this study. references ahola, s., malmberg, j., & järvenoja, h. (2023). investigating the relation of higher education students’ situational self-efficacy beliefs to participation in group level regulation of learning during a collaborative task. cogent education, 10(1), 2164241. https://doi.org/10.1080/2331186x.2022.2164241 amon, m. j., vrzakova, h., & d’mello, s. k. (2019). beyond dyadic coordination: multimodal behavioral irregularity in triads predicts facets of collaborative problem solving. cognitive science, 43(10), 1–22. https://doi.org/10.1111/cogs.12787 azevedo, r. (2015). defining and measuring engagement and learning in science: conceptual, theoretical, methodological, and analytical issues. educational psychologist, 50(1), 84–94. https://doi.org/10.1080/00461520.2015.1004069 ba, s., & hu, x. (2023). measuring emotions in education using wearable devices: a systematic review. computers & education, 200, 104797. https://doi.org/10.1016/j.compedu.2023.104797 bakdash, j. z., & marusich, l. r. (2017). repeated measures correlation. frontiers in psychology, 8(mar), 1–13. https://doi.org/10.3389/fpsyg.2017.00456 bakhtiar, a., webster, e. a., & hadwin, a. f. (2018). regulation and socio-emotional interactions in a positive and a negative group climate. metacognition and learning, 13(1), 57–90. https://doi.org/10.1007/s11409-017-9178-x benedek, m., & kaernbach, c. (2010). a continuous measure of phasic electrodermal activity. journal of neuroscience methods, 190(1), 80–91. https://doi.org/10.1016/j.jneumeth.2010.04.028 haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 116 | flr benjamini, y., & hochberg, y. (1995). controlling the false discovery rate: a practical and powerful approach to multiple testing. journal of the royal statistical society. series b (methodological), 57(1), 289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x cen, l., ruta, d., powell, l., hirsch, b., & ng, j. (2016). quantitative approach to collaborative learning: performance prediction, individual assessment, and group composition. international journal of computer-supported collaborative learning, 11(2), 187–225. https://doi.org/10.1007/s11412-016-9234-6 coco, m. i., mønster, d., leonardi, g., dale, r., & wallot, s. (2021). unidimensional and multidimensional methods for recurrence quantification analysis with crqa. the r journal, 13(1), 145–163. https://doi.org/10.32614/rj-2021-062 cohen, e. g. (1994). restructuring the classroom: conditions for productive small groups. review of educational research, 64(1), 1–35. https://doi.org/10.3102/00346543064001001 dawson, m. e., schell, a. m., & filion, d. l. (2017). the electrodermal system. in j. t. cacioppo, l. g. tassinary, & g. g. berntson (eds.), handbook of psychophysiology (pp. 217–243). cambridge university press. https://doi.org/10.1017/9781107415782.010 de backer, l., van keer, h., & valcke, m. (2020). variations in socially shared metacognitive regulation and their relation with university students’ performance. metacognition and learning, 15(2), 233–259. https://doi.org/10.1007/s11409-020-09229-5 van eijndhoven, k., wiltshire, t. j., hałgas, e. a., & gevers, j. m. p. (2023). a methodological framework to study change in team cognition under the dynamical hypothesis. topics in cognitive science. advance online publication. https://doi.org/10.1111/tops.12685 fredricks, j. a., blumenfeld, p. c., & paris, a. h. (2004). school engagement: potential of the concept, state of the evidence. review of educational research, 74(1), 59–109. https://doi.org/10.3102/00346543074001059 haataja, e., dindar, m., malmberg, j., & järvelä, s. (2022). individuals in a group: metacognitive and regulatory predictors of learning achievement in collaborative learning. learning and individual differences, 96(may), 102146. https://doi.org/10.1016/j.lindif.2022.102146 haataja, e., malmberg, j., dindar, m., & järvelä, s. (2022). the pivotal role of monitoring for collaborative problem solving seen in interaction, performance, and interpersonal physiology. metacognition and learning, 17(1), 241–268. https://doi.org/10.1007/s11409-021-09279-3 hadwin, a. f., järvelä, s., & miller, m. (2018). self-regulation, co-regulation and shared regulation in collaborative learning environments. in d. h. schunk & j. a. greene (eds.), handbook of self-regulation of learning and performance (pp. 83–106). routledge. https://doi.org/10.4324/9781315697048-6 isohätälä, j., järvenoja, h., & järvelä, s. (2017). socially shared regulation of learning and participation in social interaction in collaborative learning. international journal of educational research, 81, 11–24. https://doi.org/10.1016/j.ijer.2016.10.006 isohätälä, j., näykki, p., & järvelä, s. (2020). cognitive and socio-emotional interaction in collaborative learning: exploring fluctuations in students’ participation. scandinavian journal of educational research, 64(6), 831–851. https://doi.org/10.1080/00313831.2019.1623310 järvelä, s., & hadwin, a. f. (2013). new frontiers: regulating learning in cscl. educational psychologist, 48(1), 25–39. https://doi.org/10.1080/00461520.2012.748006 järvelä, s., järvenoja, h., malmberg, j., isohätälä, j., & sobocinski, m. (2016). how do types of interaction and phases of self-regulated learning set a stage for collaborative engagement? learning and instruction, 43, 39–51. https://doi.org/10.1016/j.learninstruc.2016.01.005 haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 117 | flr järvelä, s., malmberg, j., sobocinski, m., & kirschner, p. a. (2021). metacognition in collaborative learning. in u. cress (ed.), international handbook of computer-supported collaborative learning (pp. 281–294). springer international publishing. https://doi.org/10.1007/978-3-03065291-3_15 järvenoja, h., malmberg, j., törmänen, t., mänty, k., haataja, e., ahola, s., & järvelä, s. (2020). a collaborative learning design for promoting and analyzing adaptive motivation and emotion regulation in the science classroom. frontiers in education, 5(july). https://doi.org/10.3389/feduc.2020.00111 järvenoja, h., & järvelä, s. (2013). regulating emotions together for motivated collaboration. in m. baker, j. andriessen, & s. järvelä (eds.), affective learning together: social and emotional dimensions of collaborative learning (pp. 162–181). routledge. https://doi.org/10.4324/9780203069684 kelsey, m., akcakaya, m., kleckner, i. r., palumbo, r. v., barrett, l. f., quigley, k. s., & goodwin, m. s. (2018). applications of sparse recovery and dictionary learning to enhance analysis of ambulatory electrodermal activity data. biomedical signal processing and control, 40, 58–70. https://doi.org/10.1016/j.bspc.2017.08.024 kreijns, k., kirschner, p. a., & jochems, w. (2003). identifying the pitfalls for social interaction in computer-supported collaborative learning environments: a review of the research. computers in human behavior, 19(3), 335–353. https://doi.org/10.1016/s07475632(02)00057-2 lai, h.-m. (2021). understanding what determines university students’ behavioral engagement in a group-based flipped learning context. computers & education, 173, 104290. https://doi.org/10.1016/j.compedu.2021.104290 lee, v. r., fischback, l., & cain, r. (2019). a wearables-based approach to detect and identify momentary engagement in afterschool makerspace programs. contemporary educational psychology, 59(july), 101789. https://doi.org/10.1016/j.cedpsych.2019.101789 lobczowski, n. g., lyons, k., greene, j. a., & mclaughlin, j. e. (2021). socioemotional regulation strategies in a project-based learning environment. contemporary educational psychology, 65(march), 101968. https://doi.org/10.1016/j.cedpsych.2021.101968 malmberg, j., haataja, e., törmänen, t., järvenoja, h., zabolotna, k., & järvelä, s. (2023). multimodal measures characterising collaborative groups’ interaction and engagement in learning. in v. kovanovic, r. azevedo, d. c. gibson, & d. lfenthaler (eds.), unobtrusive observations of learning in digital environments: examining behavior, cognition, emotion, metacognition and social processes using learning analytics (pp. 197–216). springer international publishing. https://doi.org/10.1007/978-3-031-30992-2_12 malmberg, j., haataja, e., seppänen, t., & järvelä, s. (2019). are we together or not? the temporal interplay of monitoring, physiological arousal and physiological synchrony during a collaborative exam. international journal of computer-supported collaborative learning, 14(4), 467–490. https://doi.org/10.1007/s11412-019-09311-4 mendes, w. b. (2009). assessing autonomic nervous system activity. in e. harmon-jones & j. s. beer (eds.), methods in social neuroscience (pp. 118–147). guilford press. martin, a. j., malmberg, l.-e., pakarinen, e., mason, l., & mainhard, t. (2023). the potential of biophysiology for understanding motivation, engagement and learning experiences. british journal of educational psychology, 93(s1), 1–9. https://doi.org/10.1111/bjep.12584 montague, e., xu, j., & chiou, e. (2014). shared experiences of technology and trust: an experimental study of physiological compliance between active and passive users in haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 118 | flr technology-mediated collaborative encounters. ieee transactions on human-machine systems, 44(5), 614–624. https://doi.org/10.1109/thms.2014.2325859 mänty, k., järvenoja, h., & törmänen, t. (2023). the sequential composition of collaborative groups’ emotion regulation in negative socio-emotional interactions. european journal of psychology of education, 38(1), 203–224. https://doi.org/10.1007/s10212-021-00589-3 mänty, k., järvenoja, h., & törmänen, t. (2020). socio-emotional interaction in collaborative learning: combining individual emotional experiences and group-level emotion regulation. international journal of educational research, 102(april), 101589. https://doi.org/10.1016/j.ijer.2020.101589 mønster, d., håkonsson, d. d., eskildsen, j. k., & wallot, s. (2016). physiological evidence of interpersonal dynamics in a cooperative production task. physiology & behavior, 156, 24–34. https://doi.org/10.1016/j.physbeh.2016.01.004 ouyang, f., xu, w., & cukurova, m. (2023). an artificial intelligence-driven learning analytics method to examine the collaborative problem-solving process from the complex adaptive systems perspective. international journal of computer-supported collaborative learning, 18(1), 39–66. https://doi.org/10.1007/s11412-023-09387-z palumbo, r. v., marraccini, m. e., weyandt, l. l., wilder-smith, o., mcgee, h. a., liu, s., & goodwin, m. s. (2017). interpersonal autonomic physiology: a systematic review of the literature. personality and social psychology review, 21(2), 99–141. https://doi.org/10.1177/1088868316628405 pijeira-díaz, h. j., drachsler, h., järvelä, s., & kirschner, p. a. (2016). investigating collaborative learning success with physiological coupling indices based on electrodermal activity. proceedings of the sixth international conference on learning analytics & knowledge lak ’16, 64–73. https://doi.org/10.1145/2883851.2883897 rogat, t. k., hmelo-silver, c. e., cheng, b. h., traynor, a., adeoye, t. f., gomoll, a., & downing, b. k. (2022). a multidimensional framework of collaborative groups’ disciplinary engagement. frontline learning research, 10(2), 1–21. https://doi.org/10.14786/flr.v10i2.863 roschelle, j., & teasley, stephanied. d. (1995). the construction of shared knowledge in collaborative problem solving. in c. o’malley (ed.), computer supported collaborative learning (vol. 128, pp. 69–97). springer berlin heidelberg. https://doi.org/10.1007/978-3642-85098-1_5 salmela-aro, k., upadyaya, k., cumsille, p., lavonen, j., avalos, b., & eccles, j. (2021). momentary task-values and expectations predict engagement in science among finnish and chilean secondary school students. international journal of psychology, 56(3), 415–424. https://doi.org/10.1002/ijop.12719 schneider, b., dich, y., & radu, i. (2020). unpacking the relationship between existing and new measures of physiological synchrony and collaborative learning: a mixed methods study. international journal of computer-supported collaborative learning, 15(1), 89–113. https://doi.org/10.1007/s11412-020-09318-2 sinatra, g. m., heddy, b. c., & lombardi, d. (2015). the challenges of defining and measuring student engagement in science. educational psychologist, 50(1), 1–13. https://doi.org/10.1080/00461520.2014.1002924 sinha, s., rogat, t. k., adams-wiggins, k. r., & hmelo-silver, c. e. (2015). collaborative group engagement in a computer-supported inquiry learning environment. international journal of computer-supported collaborative learning, 10(3), 273–307. https://doi.org/10.1007/s11412-015-9218-y haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 119 | flr slovák, p., tennent, p., reeves, s., & fitzpatrick, g. (2014). exploring skin conductance synchronisation in everyday interactions. proceedings of the 8th nordic conference on human-computer interaction fun, fast, foundational nordichi ’14, september 2015, 511–520. https://doi.org/10.1145/2639189.2639206 sobocinski, m., järvelä, s., malmberg, j., dindar, m., isosalo, a., & noponen, k. (2020). how does monitoring set the stage for adaptive regulation or maladaptive behavior in collaborative learning? metacognition and learning, 15(2), 99–127. https://doi.org/10.1007/s11409-02009224-w sung, g., bhinder, h., feng, t., & schneider, b. (2023). stressed or engaged? addressing the mixed significance of physiological activity during constructivist learning. computers & education, 199, 104784. https://doi.org/10.1016/j.compedu.2023.104784 symonds, j. e., kaplan, a., upadyaya, k., aro, k. s., torsney, b. m., skinner, e., & eccles, j. s. (2024). momentary student engagement as a dynamic developmental system. journal of theoretical and philosophical psychology. https://doi.org/10.1037/teo0000288 symonds, j., schreiber, j. b., & torsney, b. m. (2021). silver linings and storm clouds: divergent profiles of student momentary engagement emerge in response to the same task. journal of educational psychology, 113(6), 1192–1207. https://doi.org/10.1037/edu0000605 törmänen, t., järvenoja, h., & mänty, k. (2021). all for one and one for all – how are students’ affective states and group-level emotion regulation interconnected in collaborative learning? international journal of educational research, 109(september), 101861. https://doi.org/10.1016/j.ijer.2021.101861 upadyaya, k., cumsille, p., avalos, b., araneda, s., lavonen, j., & salmela-aro, k. (2021). patterns of situational engagement and task values in science lessons. the journal of educational research, 114(4), 394–403. https://doi.org/10.1080/00220671.2021.1955651 volet, s.e. (2001). significance of cultural and motivational variables on students' appraisals of group work. in f. salili, c.y. chiu, & y.y. hong (eds). student motivation: the culture and context of learning (ch15) (pp. 309-334). kluwer academic / plenum publishers. 10.1007/978-1-4615-1273-8_15 vuopala, e., näykki, p., isohätälä, j., & järvelä, s. (2019). knowledge co-construction activities and task-related monitoring in scripted collaborative learning. learning, culture and social interaction, 21(march), 234–249. https://doi.org/10.1016/j.lcsi.2019.03.011 vuorenmaa, e., järvelä, s., dindar, m., & järvenoja, h. (2023). sequential patterns in social interaction states for regulation in collaborative learning. small group research, 54(4), 512– 550. https://doi.org/10.1177/10464964221137524 wallot, s., & leonardi, g. (2018). analysing multivariate dynamics using cross-recurrence quantification analysis (crqa), diagonal-cross-recurrence profiles (dcrp), and multidimensional recurrence quantification analysis (mdrqa) – a tutorial in r. frontiers in psychology, 9(december), 1–21. https://doi.org/10.3389/fpsyg.2018.02232 wallot, s., mitkidis, p., mcgraw, j. j., & roepstorff, a. (2016). beyond synchrony: joint action in a complex production task reveals beneficial effects of decreased interpersonal synchrony. plos one, 11(12), e0168306. https://doi.org/10.1371/journal.pone.0168306 wallot, s., roepstorff, a., & mønster, d. (2016). multidimensional recurrence quantification analysis (mdrqa) for the analysis of multidimensional time-series: a software implementation in matlab and its application to group-level data in joint action. frontiers in psychology, 7, 1–13. https://doi.org/10.3389/fpsyg.2016.01835 haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 120 | flr zabolotna, k., malmberg, j., & järvenoja, h. (2023). examining the interplay of knowledge construction and group-level regulation in a computer-supported collaborative learning physics task. computers in human behavior, 138, 107494. https://doi.org/10.1016/j.chb.2022.107494 zhang, y., qin, f., liu, b., qi, x., zhao, y., & zhang, d. (2018). wearable neurophysiological recordings in middle-school classroom correlate with students’ academic performance. frontiers in human neuroscience, 12(november), 1–8. https://doi.org/10.3389/fnhum.2018.00457 zheng, l., & huang, r. (2016). the effects of sentiments and co-regulation on group performance in computer supported collaborative learning. internet and higher education, 28, 59–67. https://doi.org/10.1016/j.iheduc.2015.10.001 zheng, l., li, x., & huang, r. (2017). the effect of socially shared regulation approach on learning performance in computer-supported collaborative learning. educational technology & society, 20, 35–46. zschocke, k., wosnitza, m., & bürger, k. (2016). emotions in group work: insights from an appraisal-oriented perspective. european journal of psychology of education, 31(3), 359– 384. https://doi.org/10.1007/s10212-015-0278-1 haataja et al _________________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 121 | flr appendix the coding scheme used in the video data analysis type of interaction description of behavior examples co-regulation of learning (corl) co-regulation of learning was coded when the following occurred: 1) observation of an obstacle to the individual's or the group's learning process 2) regulatory initiation from a group member 3) no additional strategic content from other group members following the initiation 4) strategic change in action. when corl was targeted to emotions and motivation, it could be coded without a clear obstacle or change in action to allow the strategic activities that were aiming to maintain and strengthen the already favourable motivational and affective conditions. [group is stuck on a task] s1: "are we supposed to first calculate how long it takes for it to travel? i don't know!" s2: "should we just start with this other task?" s1: "yes, let's do that!" [group struggles with calculations and asks for help] s1: "what the heck! these are all fractions! our calculations are all screwed then!" s2: "we should ask for help." s1: [raises hand] [s1 is frustrated] s1: we are going to get an f. s2: no we are not! we are getting an a. socially shared regulation of learning (ssrl) socially shared regulation of learning was coded when the following occurred: 1) observation of an obstacle to the group's learning process 2) regulatory initiation from a group member 3) active shared strategic negotiation between at least two group members 4) strategic change in action. when ssrl was targeted to emotions and motivation, it could be coded without a clear obstacle or change in action to allow the strategic activities that were aiming to maintain and strengthen the already favourable motivational and affective conditions. [students struggle to understand task] s1: "this case looks exactly the same as the previous. what on earth does this mean? s2: [raises hand to ask teacher for help] s3 [to teacher]: "we understand nothing about this!" s2 [to teacher]: "how do we know which lens this is?" [students maintain motivational conditions] s1: "it's great that we are all participating." s2: "yes, everyone participates!" s1: "...so then everyone is going to get an a." s3 [laughs]: "yep!" s4 [laughs]: "well, of course." frontline learning research vol. 12 no. 3 (2024) 20 44 issn 2295-3159 corresponding author: andreas gegenfurtner, university of augsburg, universitätsstraße 10, 86159 augsburg, germany, andreas.gegenfurtner@uni-a.de. doi: https://doi.org/10.14786/flr.v12i3.543 horizontal transition of expertise andreas gegenfurtner1 , hans gruber2&3, erno lehtinen3 & roger säljö 4 1 university of augsburg, germany 2 university of regensburg, germany 3 university of turku, finland 4 university of gothenburg, sweden article received 14 july 2019 / article revised 10 august 2024 / accepted 20 august 2024 / available online 3 september 2024 abstract expert performance in a domain is often defined as maximal adaptation to stable task constraints. this definition is useful when analysing the vertical transition when novices become experts. however, many workplaces undergo considerable changes and, thus, task constraints change as well. in this paper a complementary conceptualisation of expertise is offered, one that focuses on expert performance as recurring adaptation to dynamic task constraints. this definition is useful when analysing the horizontal transition when experts adapt to dynamically changing work contexts. using the documentary method, the aim of the present study was to analyze cases of horizontal transitions based on qualitative biographical interview data from five experts reconstructing different types of adaptation to technological change in their domains that they have experienced. implications for studying horizontal transitions at dynamic worksites are discussed. keywords: expertise; expert performance; change; work; documentary method. mailto:andreas.gegenfurtner@uni-a.de https://doi.org/10.14786/flr.v12i3.543 gegenfurtner, gruber, lethinen & säljö 21 | f l r 1. introduction many professions undergo recurring transformations as they face rapid and constant change owing to frequent technological innovations that challenge the way experts work in these domains. these challenges also concern what constitutes expert performance in novel environments (billett et al., 2018; harteis & goller, 2014; palonen et al., 2014; ward et al., 2019), and how experts adapt their practices to technological shifts in a ‘liquid modernity’ (baumann, 2007). notably, boshuizen and van de wiel (2014, p. 71) argue that “[n]ew professions may emerge and the tasks currently undertaken by experts will change requiring learning and gradual or revolutionary adaptations.” if we understand expertise as interdependences between human agency, minds, bodies, and digital technological artifacts (boshuizen et al. 2020; gegenfurtner et al., 2023; gruber & harteis, 2018; lehtinen et al., 2014; säljö, 2019; szulewski et al., 2019), then existing frameworks for understanding expertise and expert performance in dynamically changing contexts need to be reconsidered. such reconsideration is afforded by the notion of horizontal transition of expertise. this notion invites an analysis of expertise in dynamic contexts. the focus on horizontal transition of expertise aims to complement existing research on expertise development in relatively stable and well-structured (burgoyne et al., 2019; de groot, 1965; ericsson, 2018) as well as in relatively fluid and ill-structured domains (hatano & oura, 2003; längler et al., 2017; lehtinen et al., 2020). 1.1 vertical and horizontal transition of expertise research points to two complementary cases of expertise development: one that concerns vertical transition and one that concerns horizontal transition of expertise. figure 1 illustrates these cases. expert performance in the sense of a vertical transition addresses the development from novice to expert as a result of maximal adaptation to stable task constraints. in contrast, expert performance in the sense of a horizontal transition addresses the development and maintenance of expertise as a result of recurring adaptation to dynamic task constraints. figure 1. vertical and horizontal transition of expertise. research analysing the vertical transition of a novice becoming an expert often focuses on how experts and novices differ. this line of research is often labelled the contrastive approach, as experts and novices are compared and contrasted with regard to their performance levels and the processes that are hypothesised to lead to different performance levels. these studies often treat the context or task as being relatively stable or predictable. it is no surprise that studies interested in expert performance had horizontal expertise: recurring adaptation to dynamic task constraints expert novice expert controlled context vertical expertise: maximal adaptation to stable task constraints gegenfurtner, gruber, lethinen & säljö 22 | f l r their origins in the domain of chess, a highly predictable and well-structured domain with fixed rules and standardised performance measures (de groot, 1965). because a highly controlled or ‘representative’ (ericsson, 2018; feltovich et al., 2018) task is essential when comparing novices and experts in any given domain, this line of research often collects data in laboratory-like contexts. the analytic focus is descriptive, interested in revealing how and to what extent participants at varying stages of expertise differ. studies on horizontal transition are interested in how experts adapt to changing contextual affordances. this line of research tends to differ from research examining vertical transition in several respects. for example, the participant focus is largely on the expert, not because novices are uninteresting, but rather because the main interest is in understanding how experts adapt to contextual change. this line of research is often conducted in and around sociomaterial systems that are characterised by contextual dynamics or ‘moving targets’ (hoffman et al., 2017). it is no surprise that studies interested in adaptations to change are often conducted in technology-intensive domains (gruber & harteis, 2018; ivarsson et al., 2016; lehtinen et al., 2014; lehtinen et al., 2020; sellberg et al., 2024) because constant introduction of new technologies implies constant adaptations to novel technologymediated practices (baumann, 2007; gegenfurtner et al., 2009; palonen et al., 2014; säljö, 2022; troshani et al., 2018). to capture those developmental trajectories, this line of research often collects data in the field or ‘in the wild’ (hutchins, 1995). the analytic focus is reconstructive, interested in revealing how experts adapt their practices to novel task affordances (engeström, 2018). 1.2 horizontal transition of expertise compared to other conceptualisations how does the notion of horizontal transition of expertise relate to other conceptualisations? when describing how experts behave in changing contexts, the literature offers several explanatory frameworks, including the notions of adaptive expertise, knowledge encapsulation, polycontextuality, and expert cognitive flexibility. first, adaptive expertise refers to a highly developed conceptual knowledge base that allows experts to invent novel solutions when environmental constraints change (anthony et al., 2015; bohle carbonell et al., 2014; lin et al., 2007; mylopoulos & woods, 2017). in a classic paper, hatano and inagaki (1986) offered the example of a farmer who can effectively deal with contextual covariations such as unusual weather or plant disease. a related example describes an adaptive sushi chef who excels in inventing novel and innovative sushi menus. according to hatano, adaptive expertise includes three main characteristics: a) the ability to explicate the principles underlying task performance, b) the ability to estimate when routine and non-routine task procedures are necessary, and c) the ability to adapt procedures and solution steps when needed (hatano & oura, 2003; lin et al., 2007). a focus of hatano’s adaptive expertise approach is on educating students to become adaptive experts, particularly through ‘built-in randomness’ of the teaching material—that is, variability of tasks and contexts to foster transfer of learned knowledge and skills—and a learning climate that encourages active experimentation and play (hatano & inagaki, 1986). this approach proves useful not only in undergraduate schooling, but also in professional education, for example when educating forensic specialists (mustonen & hakkarainen, 2015), health professionals (pusic et al., 2018), and teachers (männikkö & husu, 2019; suh et al., 2023). hatano focuses on the design of learning environments that can foster adaptive expertise in novices; this focus differs from analyses of horizontal transition, which aim at tracing how experts adapt to novel domain affordances. second, knowledge encapsulation refers to the clustering of lower-level biomedical knowledge structures into higher-level concepts of greater generality (boshuizen & van de wiel, 2014). within the realms of encapsulation theory (boshuizen & schmidt, 1992), the term describes the cognitive restructuring of knowledge as expertise develops, leading to abbreviations in reasoning processes. these abbreviations explain why experts are faster than novices in task completion (rikers et al., 2002; violato & et al., 2018) because experts use encapsulated knowledge concepts. these are further enriched with clinical practice and eventually transformed into illness scripts (jaarsma, 2015; strasser & gruber, 2015). boshuizen and colleagues argue that, for routine cases inside an expert’s domain, expert gegenfurtner, gruber, lethinen & säljö 23 | f l r biomedical knowledge is encapsulated and integrated into clinical knowledge, while for non-routine cases outside an expert’s domain, biomedical knowledge remains easily accessible when the need arises. indeed, encapsulation theory has been used to test how experts solve diagnostic problems outside their medical specialty (rikers et al., 2002), indicating that experts process routine and non-routine clinical case descriptions in qualitatively similar ways. thus, encapsulation theory is useful to test the robustness of expertise (boshuizen & van de wiel, 2014; jaarsma, 2015; violato et al., 2018). while the theory affords a cognitivist analysis of how experts adapt their knowledge-based reasoning processes to dynamic task constraints, to date, encapsulation theory has not been employed to explore how experts adapt their practices when contextual affordances of a domain change. third, polycontextuality refers to multiple work tasks, communities of practice, or activity systems within which experts are simultaneously engaged (engeström, 2018). the concept describes how experts cross the boundaries between parallel activity systems that afford complementary or conflicting participation frameworks. examples of polycontextual situations include multiprofessional health teams in hospitals or interacting work groups in industrial plants (engeström, 2018). the focus of polycontextuality is on parallel, already existing activity systems (dochy et al., 2021). engeström, engeström, and kärkkäinen (1997) convincingly argue that modern work practices are not singular, linear, or stable; instead, “practitioners face the challenge of negotiating and combining ingredients from different contexts to achieve hybrid solutions” as they “operate in and move between multiple, parallel activity systems” (engeström et al., 1997, p. 442). this observation, made decades ago, has not lost its relevance. still, while the theoretical notion of polycontextuality is concerned with how experts move between different activity systems, the notion of horizontal transition is on how an activity system changes and how experts adapt to these changes. so the temporal dimension differs between the two notions: polycontextuality is oriented toward presently existing, multiple activity systems and the manoeuvring of experts between them; horizontal transition is oriented toward the development of an activity system from (past to) present to future and how experts accommodate to this development. finally, expert cognitive flexibility refers to an expert’s extensive and highly differentiated cognitive schematisation (expert flexibility of type one) and to an expert’s ability to override schemadriven processing to engage in more basic kinds of reasoning (expert flexibility of type two) when confronted with very atypical problems that are infrequent and highly unusual in nature (spiro et al., 2019) in their domain. for example, medical experts may diagnose rare variations of a medical condition using their knowledge-based flexibility (feltovich et al., 2018). the essence of cognitive flexibility is in problem solving within one relatively stable domain. spiro and colleagues (2012) note that when facing a complex new case, “one must assemble just those aspects from prior knowledge that will help in the current situation, while discounting those aspects that are less helpful. further, the assembled elements must be meaningfully related to each other, and tailored to the specific content of the case at hand to create what spiro and colleagues call a “schema-of-the-moment” (spiro et al., 2012, p. 119). teaching children how to read is a good example here, according to spiro, because expert teachers will examine each new teaching situation, based on the reading skills of different children, and adapt their teaching strategy flexibly in light of their observations (spiro et al., 2019). the concept of cognitive flexibility represents a cognitivist perspective with a focus on knowledge and memory; this notion was not designed to reconstruct how experts adapt to changes within their domain. instead, it is a highly useful framework to analyse how experts deal with problems and challenges within their domain as when teaching children with differing reading skills or when handling rare medical cases. all these conceptualisations describe how experts flexibly react to and adapt their knowledge and practices to novel or atypical contextual constellations. adaptive expertise, knowledge encapsulation, and expert cognitive flexibility stress the importance of a highly developed individual knowledge base, while polycontextuality stresses the importance of coordinating elements from different contexts to achieve hybrid solutions. what all these conceptualisations have in common is an interest in atypical or uncertain, yet representative, situations in a domain. in contrast, horizontal transition of expertise occurs when the representativeness of a domain changes. as such, this perspective is useful when analysing how experts adapt their knowledge and practices to changes within existing gegenfurtner, gruber, lethinen & säljö 24 | f l r professional fields—which affords an analytic focus not presently found in expertise theories. more specifically, while hatano and inagaki (1986) focus on the design of learning environments that can foster adaptive expertise in novices, this focus differs from analyses of horizontal transitions, which aim at tracing how experts adapt to novel domain affordances. furthermore, while boshuizen and schmidt (1992) afford a cognitivist analysis of how experts adapt their knowledge-based reasoning processes to dynamic task constraints, horizontal transitions of expertise reconstruct how experts adapt their practices when contextual affordances of a domain change. moreover, while engeström (2018) is concerned with how experts move between different activity systems, the notion of horizontal transition is on how an activity system changes and how experts adapt to these changes. so the temporal dimension differs between the two notions: polycontextuality is oriented toward presently existing, multiple activity systems and the manoeuvring of experts between them; horizontal transition is oriented toward the development of an activity system from (past to) present to future and how experts accommodate to this development. in addition, spiro et al. (2012) represent a cognitivist perspective with a focus on knowledge and memory to analyse how experts deal with problems and challenges within their domain while horizontal transitions of expertise put into focus how experts deal with their work when the domain itself changes. in this manuscript, we apply the framework of horizontal transitions of expertise to the domain of medical image diagnosis. 1.3 the context of the present study medical image diagnosis is the interpretation of graphical representations of the human anatomy or its functions (krupinski, 2018). examples of medical images are x-ray and computer tomography (ct) scans in radiology; positron emission tomography (pet) images in nuclear medicine; or microscopic images of tissue samples in clinical pathology. past research has examined how expertise develops in medical image diagnosis (for reviews of this literature, see gruber et al., 2010; krupinski, 2018). a large body of studies on expertise in medicine or medical image diagnosis concerns vertical transition of expertise, with typical research questions such as “how do novices develop?” and “how do experts, intermediates, and novices differ?” specialised areas of modern medicine increasingly rely on digital technologies. the constant development and implementation of novel digital tools invite an analysis of how established expert practices and routines change. in addressing this topic, a pioneering study by rystedt and colleagues (2011) examined how experts interpreted an image produced by what was then a new technology, tomosynthesis, and how experts revised their routine practices of seeing. the re-working of their practices aimed at improving diagnostic accuracy and making their diagnoses accountable. of course, adaptation is not always accomplished easily, and difficulties or problems are frequently encountered (engeström, 2018; palonen et al., 2014). the present study deepens and extends first explorations of horizontal transitions published in earlier manuscripts (gegenfurtner et al., 2009; lehtinen et al., 2020) and contributes to this line of research by analysing the horizontal transition of expertise induced by technological change in three medical specialties: paediatric radiology, nuclear medicine, and clinical pathology. the rationale for choosing different medical specialties was largely informed by an interest in selecting worksites that are faced with constant invention and implementation of new tools and resources. adopting a retrospective, biographical perspective in which experts were invited to reflect on their multi-decade working lives, this study pursued the main research question: how and to what extent have experts adapted to technological change in their domains? to elaborate more deeply on these adaptations, we used the documentary method (bohnsack, 2014; garfinkel, 1967) as an analytical lens to reconstruct sociogenetic types and to answer a set of associated research questions: which technological artifacts were introduced in the domain during the course of the experts’ working lives? what was problematic in terms of work practices and ways of knowing after the new artifacts were introduced? what tools or boundary objects did the experts use to adapt their expertise? how did the adaptation evolve? which factors influenced the adaptation positively or negatively? gegenfurtner, gruber, lethinen & säljö 25 | f l r 2. method 2.1 participants to answer the research questions, five experts were interviewed. four of the experts were the current directors of their units in large university hospitals in western finland while the fifth expert was the recently retired director of his unit. one expert was from paediatric radiology, two experts came from clinical pathology, and two experts represented nuclear medicine. all participants were male, had a mean work experience of 27.68 years (± 6.42), and were on average 55.19 years old (± 9.24). the experts were selected based on several criteria, including their age, their experience, their network centrality, their status within the hospital, and their nomination by peers. these are classic criteria used in studies addressing expertise as a vertical development. since the present study is interested in the participants’ multi-decade professional biographies, this classic set of criteria for addressing vertical transition proved useful also for addressing horizontal transition. participation in the study was voluntary. anonymity and confidentiality were guaranteed for all participants. 2.2 interviews retrospective, semi-structured interviews were used to identify horizontal transition of expertise. the interviews aimed at eliciting episodes from the experts’ multi-decade working lives, with a particular focus on how the experts experienced and reported adapting to technological change in their respective domains. semi-structured qualitative expert interviews are a particularly feasible method here to approach the sociomateriality and hybrid nature of complex expert practices, and how expertise evolves in changing worlds of work (lehtinen, 2022; säljö, 2022; van de wiel, 2017; yardley et al., 2019). van de wiel (2017, p. 101) state: “as jobs, tasks, equipment, and work roles are not stable but continuously develop, experts in the field par excellence may provide valuable perspectives on future developments, innovations and novel problems and how to deal with them”. the interviews unfolded as a dialogue between the interviewee and the interviewer. all experts were interviewed individually and face-to-face. interview duration ranged from 50.55 to 90.35 minutes, with an average length of 73.20 minutes (± 11.30) and a total duration of 6 hours and 23.30 minutes. the interviews were recorded and transcribed verbatim. from the interview material, we selected all recorded talk that was associated with technological change and experts’ adaptions to change, with a total duration of 63.30 minutes of interview talk. other portions of the interviews not selected for analysis included talk about the experts’ families, their education, and their biography and career trajectory. a breakdown of minutes of the selected material per interview is reported in the appendix. 2.3 analysis the interview transcripts were analysed using the documentary method (bohnsack, 2014; nohl, 2010). the documentary method is a method for qualitative data analysis that involves searching for patterns underlying a variety of different realisations of meaning. following garfinkel (1967, p. 78), utterances in the interview transcripts were treated “as the ‘document of’, as ‘pointing to’, as ‘standing on behalf of’ a presupposed underlying pattern. not only is the underlying pattern derived from its individual documentary evidences, but the individual documentary evidences, in their turn, are interpreted on the basis of ‘what is known’ about the underlying pattern. each is used to elaborate the other.” in contrast to other methods for analyzing qualitative interview data, the documentary method aims to reconstruct the interviewees’ own frames of orientation (bohnsack, 2014) which offers an interpretational framework for understanding, in our study, how the clinical directors adapted to technology changes throughout their professional careers. as philipps and mrowczynski (2021, p. 60) observe: “on the one hand, interviewees recapitulate in narrated stories and descriptions how they lived through different biographic events and processes; on the other hand, they often make argumentative or gegenfurtner, gruber, lethinen & säljö 26 | f l r evaluative statements which offer their own interpretations of narrative passages”; the documentary method seeks to reconstruct these frames of orientation experts articulate. in this respect, the documentary method complements other forms of qualitative analyses of expert interviews, such as qualitative content analysis or thematic analysis, which often develop a category scheme, code the interview transcripts based on this scheme, count the occurrence of different codes, and/or visualize interconnections between codes in epistemic network analyses (gabel et al., in press; stahnke & gegenfurtner, in press; szulewski et al., 2019; van de wiel, 2017; white et al., 2018). while such analyses are of course valid, they often seek to contrast experts, intermediates, and novices by adopting the view of vertical transitions. exploring horizontal transitions, the present study addressed the what and how of expert reconstructions following the documentary method (for more details on the method and step-by-step empirical illustrations, see beck, 2021; bohnsack, 2014; nohl, 2010; nohl, 2017; philipps & mrowczynski, 2021). particularly, the interview transcripts were analysed in three stages and six steps. table 1 presents an overview. table 1 stages and steps in the documentary method stages i. the formulating interpretation ii. the reflecting interpretation iii. type formation steps 1. topical structuring 2. detailed formulating interpretation 3. semantic interpretation 4. comparative sequential analysis 5. sensegenetic type formation 6. sociogenetic type formation the first stage of the documentary method involved the formulating interpretation. the aim was to establish the “what” of the interview text—that is, the “actual appearance” (garfinkel, 1967, p. 78) of evidence. specifically, the interview transcripts were segmented chronologically according to the appearance of topics (step 1: topical structuring). for the purpose of the present study, one topic was chosen for deeper analysis: adaptations to technological change. transcript segments describing instances of how the experts adapted to technological change were paraphrased and condensed (step 2: detailed formulating structuring). the second stage of the documentary method involved the reflecting interpretation. the aim was to establish the “how” of the interview text—that is, the “presupposed underlying pattern” (garfinkel, 1967, p. 78). specifically, the sequence of topics from stage 1 was semantically interpreted as argumentation, evaluation, description, and narrative (step 3: semantic interpretation). semantic interpretation was done individually for each interview. for the purpose of the present study, the narratives from step 3 were then compared across the five interviews to analyse the orientation frameworks within which the experts reconstructed their adaptations to technological change (step 4: comparative sequential analysis). finally, the last stage of the documentary method involved type formation. the aim was to generate insights into different types of adaptation to technological change in the form of common conclusions across interviews. specifically, the different orientation frameworks identified in step 4 were abstracted from the interview transcripts and formulated as types in their own right. specifically, the identified types are given a stand-alone meaning that go beyond the individual interviews; this allows considering maximally contrastive types how experts orient themselves when adapting to technological change (step 5: sense-genetic type formation). the interpretation of these sense-genetic types was finalised using further segments of the interview transcripts to establish and reconstruct the social contexts and constellations within which the experts’ adaptations evolved (step 6: sociogenetic type formation). while the sense-genetic type formation illustrates the different frames of orientations within gegenfurtner, gruber, lethinen & säljö 27 | f l r which the experts react and adapt to domain changes, this sense-genetic type formation cannot illuminate how these different types are embedded in different sociomaterial constellations (bohnsack, 2014; nohl, 2017). to unveil these constellations, the sociogenetic type formation draws from additional documentary evidence from other interview parts that illustrate, in our case, what was problematic when new technology was introduced, which tools or boundary objects were used, how adaptations evolved, and which factors influenced the adaptation. 3. results the findings from the documentary method analyses will be presented in three steps. first, we present the outcomes of the formulating interpretation that highlight relevant narratives from the five experts. second, we report the findings of the interpreting formulation to identify technological change in each medical specialty. third and finally, we present and describe the resulting type formation that synthesises adaptations to technological change. figure 2 presents a graphical representation of how the types of adaptation are related to the medical experts’ narratives. 3.1 formulating interpretation the formulating interpretation analysed the topical structure of each interview. the present study focuses on one of the topics: adaptation to technological change. results of the detailed formulating interpretation of utterances associated with adaption to technological change are presented for each of the five interviews in the appendix. 3.2 reflecting interpretation the reflecting interpretation aims at establishing the modus operandi of how the interview narratives describe technological change in the different medical domains. figure 2 provides an overview. in total, eight instances of significant technological change emerged from the five interviews. the following paragraphs elaborate on these changes and enrich them with text fragments taken from the semantic interpretation of each interview. 3.2.1. from x-ray radiography to computer tomography a first change refers to the transition from x-ray radiography to computer tomography in the domain of paediatric radiology. x-ray radiographs are analogous pictures or films that are placed upon a lightbox illuminated from behind to allow viewing the x-ray films with high contrast. in comparison, a computer tomography (ct) digitises the information from the x-ray scanning and presents the pictures on a computer screen with high resolution. radiographs and computer tomographs use the same underlying processes for creating the pictures. the perceived similarity made an adaptation to this change relatively easy according to the participants. nowadays, ct is part of the clinical standard in hospitals worldwide. in his narrative describing the transition to digital radiography (interview 3, 36:25’—36:45’), the expert says: “i have been lucky to see this development quite early. so i got to know this new technology very early. i… i am working in a university hospital. you can learn new tricks every day. that… that… so that makes my life easy, i had that kind of… environment, so i could learn this new trick during my work.” gegenfurtner, gruber, lethinen & säljö 28 | f l r figure 2. representation of the data analysis process. 3.2.2. from old versions of pet to new versions of pet a second change refers to the transition from old versions of positron emission tomography (pet) to newer versions of pet in the domain of nuclear medicine. new hardware and new software were frequently implemented and produced images of increasingly higher resolution. these changes were perceived as making work more convenient. the underlying biophysical processes of image production remained unchanged, so adapting to new versions of pet was relatively easy. as the expert describes it (interview 5, 22:05’—22:15’): “of course, in the beginning, you have to become familiar again with the new program… with the new version. but nowadays… the software is so similar. this is all not a big problem.” adaptation 2 problematic adaptation 3 contingent adaptation 1 successful adaptation 4 failed reflecting interpretation formulating interpretation type formation changes in technology adaptation to change narratives from medical experts change 4 pet/mri change 3 pet/ct change 2 pet change 5 regulations change 6 virtual microscopy change 7 teaching change 8 3d and 4d ct change 1 ct interview 4 clinical pathology 1 interview 1 clinical pathology 2 interview 2 nuclear medicine interview 5 nuclear medicine interview 3 pediatric radiology gegenfurtner, gruber, lethinen & säljö 29 | f l r 3.2.3. from separate images produced by pet and ct to fusion images pet/ct a third change reported refers to the transition from separate pet and ct images to a novel image type that combines pet and ct into one image, pet/ct. these new fusion images are used in the domains of radiology and nuclear medicine. to interpret pet/ct images correctly, experts need to understand how both elements, pet and ct, are produced. this requires crossing the boundaries from one’s “home” domain to a neighbouring domain as the underlying biophysical processes of image production are radically different. these differences complicate an adaptation, or to use the words of an expert (interview 5, 29:20’—29:30’): “and this is a problem with it… ok? that the mode and … the style of how different findings are treated is different.” these differences are also evident in the order of the diagnostic practices. radiologists typically start with interpreting the ct information and continue with interpreting the pet information, while physicians in nuclear medicine “do it just the other way around. i look first at what i can see in the pet and then i look at… where is that what i see in pet… where is that anatomically located and then i look at the surrounding morphology” (narrative from interview 5, 29:35’—29:50’). 3.2.4. from separate images produced by pet and mr to fusion images pet/mr similar to the third change, a fourth change described in the interviews refers to the transition from pet and mr images to a novel image type that combined pet and mri into one image, pet/mri. this new fusion technology is used in nuclear medicine. again, pet/mri crosses boundaries. to interpret pet/mri with high diagnostic accuracy, nuclear medicine experts need to leave their “home” domain and learn from another domain. as the expert remembers (interview 5, 30:10—30:40): “now this was… much more difficult. because we have two imaging technologies produced with very, very different underlying principles… that is in the brain… previously it has been used in neuro… neurosciences. this is exciting. but i had to adapt… really quite… radically different.” 3.2.5. from old regulations to new regulations a fifth change was associated with the introduction of new regulations, particularly in the domain of nuclear medicine. this change illustrates how governing policies will guide and steer adaptation. within the context of pet/ct for example, an expert says (interview 2, 13:15’—13:50’): “when this kind of combination came, it was an immediate question how we can cope with this and … in different countries, this has been solved very differently. i know that… in certain countries, where everything is very tightly regulated… like for example in germany… then… there is a… all is said by some rules or statutes or whatever… law…that a certain education and training is needed to interpret this case… that is all very strictly regulated, everything that you can do.” these regulations were the basis for implementing training programs. at the same time, new regulations are perceived as increasing cost and slowing down the workflow. these perceptions contribute to a critical attitude toward regulations. “they are… seen artificially to bring more safety but in reality there is very little evidence that the safety is improving… it is just you know the authorities think that the control increases the safety of … or make the life of patients in these studies more safe… but this decision is not based on evidence or rational justification, it is just their gut feelings, and then you know… regulators like to regulate” (35:05—35:20). adapting to these new regulations and implementing them is described as a frustrating experience. “it’s frustrating in that sense that… as a scientist you would like to see the evidence for these regulations… but this is not the case.” (37:20—37:25). 3.2.6. from light microscopy to virtual microscopy a sixth change refers to the transition from light microscopy to virtual microscopy in clinical pathology. virtual microscopy integrates microscopy technology and digital technology and digitises slide sets of tissue samples that can be used for clinical purposes. however, as an expert notes, adaptation from light to virtual microscopy is compromised by image quality of the virtual slides (interview 1, 39:00’—39:45’): “there are still differences… and especially the… the focusing and the depth of the focusing… it is something that is that is still lacking or is not as far and developed as in the light microscopy… and it’s an important feature… the samples are… the slides are typically 4 or 5 gegenfurtner, gruber, lethinen & säljö 30 | f l r micrometers thick… and the… the… it’s an important diagnostic aid… you can scan the entire sample especially when you are focusing on cellular… cellular details.” and although virtual microscopy is perceived to have high potential for the future of the domain, there was still resistance to change. “if i would have to move to giving diagnoses on the virtual slides, it would be a problem for me, because i am so used to, i have this 25-year history of looking through the microscope” (interview 1, 57:35’— 58:00’). another expert stated that once he was asking his colleagues what they thought about virtual microscopy or if they use it, the responded that (interview 4, 21:10’—21:15’). “[t]hey did not like it”. for younger generations, an adaptation might be easier because of their digital nativeness, but for an older generation, change is difficult. “people are lazy animals” (interview 1, 71:00’—71:05’). 3.2.7. from analogous teaching environments to digital teaching environments a seventh change refers to the transition from analogous teaching environments to digital teaching environments in the domains of radiology and clinical pathology. clinical pathology has benefitted from the introduction of virtual microscopy. although this change has, as we have pointed out, been associated with some problems for clinical work, it is seen as positive for teaching purposes. it eases the availability of study materials and equipment, as one expert argues (interview 4, 15:35’— 15:50’): “when i was young, every student was given a box filled with about 200 slides… this was the traditional way… and now that you have this virtual… virtual microscopy… teaching is basically based on digital images.” the interviews highlight how suitable virtual microscopy is for teaching, and teachers have been using it with considerable success. in radiology, students can access computer tomographic (ct) scans for training purposes at any time from any place. these digital databanks are useful, “but students should… also have access to a teacher when… in case they need one” (interview 3, 82:40’—82:50’). 3.2.8. from 2d computer tomography to 3d and 4d computer tomography a final change refers to the transition from two-dimensional to threeand four-dimensional computer tomography images. two-dimensional ct images resemble analogous x-ray films, while 3d representations allow zooming in and out of the image and use a set of images as if watching a movie. this change from static 2d pictures to 3d pictures was seen as challenging. even more challenging was the transition to 4d, such as in scans of a pumping heart. some radiologists face problems in adapting their routines to 3d and 4d. as an expert notes (interview 3, 39:05’—39:45’): “some of my colleagues… they look at the 3d data still as if it was 2d… they are going through all the 1000 images… instead of picking up… like… you have to learn new working… practices and a new workflow system… you cannot spend your time looking at all 1000 images carefully but you have to learn a new way how to pick up the information in 3d.” explanations for the failed adaptation included lack of training and lack of standardised procedures. (interview 3, 41:45’—41:50’): “there is a generation thing too… especially younger doctors have… are doing it easier.” gegenfurtner, gruber, lethinen & säljö 31 | f l r table 2 sense-genetic type formation type 1 type 2 type 3 type 4 utterance 1 “the new imaging technology was introduced in our department.” utterance 2 “it was similar to the previous one.” “it used a different physical procedure to create the representations.” “it digitalised the analogical tissue samples.” “it transformed the analogical slides into digital 4d representations.” utterance 3 “it was easy to adapt.” “so we started to collaborate with colleagues in other departments.” “though good for teaching, the perception of depth was problematic.” “but we did not know how to change our old work routines.” 3.3 type formation the stage of type formation aimed at abstracting instances of technological change described and clustered in the reflecting interpretation. type formation included two steps. in a first step, the sense-genetic type formation identified different types of how technological change was reacted to. these instances are interpreted as homologous patterns of adaptation. to illustrate the type formation process, table 2 presents the four types with utterances from the interview narratives. in a second step, the four types of adaptations were further contextualised with additional documentary evidence described in the formulating interpretation. this additional documentary evidence helped to reconstruct the social contexts and constellations associated with which the technology introduced; what was problematic after the new technology had been introduced; which tools or boundary objects were used; how the adaptation evolved; and which factors influenced (promoted or hindered) the adaptation process. ultimately, this analysis resulted in four socio-genetic types of adaptation presented in table 3: (a) successful adaptation, (b) problematic adaptation, (c) contingent adaptation, and (d) failed adaptation. 3.3.1. successful adaptation the type of successful adaptation emerged in narratives related to the transition (a) from x-ray radiography to computer tomography and (b) from old versions to new versions of positron emission tomography. in both cases, the change was associated with lower-level surface changes in image processing, such as an increased resolution and updates of familiar software packages. this perceived surface similarity functioned as a cognitive artifact (in a vygotskian sense) and successfully mediated the transition from the old technology to the new technology. the interview narratives from both cases illustrate that the adaptation was gradually achieved during regular work practice; it was not accompanied by formal training or regulatory policies. similarities with the previous technology were perceived as supporting the adaptation. gegenfurtner, gruber, lethinen & säljö 32 | f l r table 3 sociogenetic types of adaptation successful adaptation problematic adaptation contingent adaptation failed adaptation change 1, 2 3, 4, 5 6, 7 8 1. what was introduced? developments of pet and ct fusion images (pet/ct and pet/mri) ct, virtual microscopy digital radiography 2. what was problematic? lower-level surface changes in image processing different processes underlying image production perception of depth dimensionality of visualisation (2d, 3d, 4d) 3. what tools or boundary objects were used? cognitive artifacts: perceived surface similarity crossdisciplinary collaboration physical and cognitive artifacts: slide similarity work practices 4. how did the adaptation evolve? gradually during regular work practice guided by regulations and formal training as part of teaching practice unguided, no formal training 5. which factors influenced the adaptation? similarity to previous technology curiosity, interest, attitudes, skepticism, frustration digital nativeness, skepticism (lack of) digital nativeness, routine practices 3.3.2. problematic adaptation the types of problematic adaptations that emerged in narratives related to the fusion of positron emission tomography with computer tomography (pet/ct) and magnetic resonance imaging (pet/mri), as well as administrative regulations accompanying these fusion images. what was perceived as being problematic were the different biophysical processes underlying image production, which required cross-disciplinary collaboration with neighbouring specialties. the adaptation process was guided by regulatory bodies that introduced new statutes and control mechanisms, including formal training programs. factors that promoted the adaptation to technological change were curiosity, interest, and a positive attitude toward digital technologies. factors that hindered the adaptation process were scepticism with respect to the added value as well as frustration over the regulations introduced. although the new technologies required great efforts of adaptation, and a level of formal education to cross the boundaries to neighboring disciplines, the adaptation, ultimately, succeeded. 3.3.3. contingent adaptation the type of contingent adaptation emerged in narratives related to the transition from (a) light microscopes to virtual microscopy in clinical pathology, and (b) from 2d to 3d computer tomography in radiology. in the case of pathology, the perception of depth was considered problematic for clinical gegenfurtner, gruber, lethinen & säljö 33 | f l r purposes. the similarity between material and digital slides functioned as boundary object that mediated the adaptation process, a process that was hindered by scepticism and lack of digital nativeness that seemed particularly prevalent in older generations of clinical pathologists. being “sceptic”, which was uttered by the experts, has a connotation of disbelief and unwillingness, while the ground for their hesitations might be more solid. still, virtual microscopy was successfully introduced and adapted to when the technology was framed as a teaching aid in classrooms. thus, the adaptation was contingent on the context in which it was used. virtual microscopy is not replacing light microscopes in clinical work but serves as a resource for medical education and teaching. similarly, in the case of radiology, a shift from 2d to 3d tomography was successful only when ct scans were used for teaching purposes as a resource for students. again, while adaptation in the clinical context failed, adaptation in the teaching context succeeded, so success of the adaptation process was contingent on the (teaching) context. 3.3.4. failed adaptation the type of failed adaptation emerged in narratives related to the transition from twodimensional to threeand four-dimensional computer tomography images in clinical radiology contexts. the increased levels of dimensionality were perceived as being very problematic, so routine practices and work flows that had developed around 2d images proved ineffective for 3d and 4d representations of the human anatomy. adaptation failed because it was unguided by formal training programs or mentored work practices. additional barriers compromising a horizontal transition of expertise included attitudinal and age-related hindrances associated with not being a ‘digital native’. 4. discussion the purpose of the present article has been to use the notion of horizontal transition of expertise for the purpose of empirically examining adaptations to technological change reported by medical experts with decades of work experience. analyses of retrospective, biographical interviews using the documentary method illustrated experiences of a number of technological changes. the way experts reported adapting to these changes differed: while some changes could easily be coped with, other changes were more profound, resulting in problems and even failure in the adaptation process. the analyses also highlight how each horizontal transition of expertise evolved, which tools and boundary objects were used, and which factors promoted or hindered the transition. promoting factors that were mentioned were formal education and training programs as well as digital nativeness and a positive attitude toward technology. hindering factors were the lack of training as well as scepticism and the perception that one was too old to adapt. we should also note that, although digital nativeness was mentioned as a promoting factor, younger professionals are not per se digital natives. what is interesting is that technology can sometimes increase image quality (as was the case for radiology and nuclear medicine) and sometimes decrease image quality (as was the case for pathology): the scanning of actual slides for virtual microscopy can result in blurred images that cannot be adapted like in light microscopes; so clinical pathologists miss the advantage of increased image quality the radiologists and nuclear physicians reported. in summary, the interview data helped to reconstruct how and to what extent experts adapted to the turbulences in their working lives produced by the introduction of novel digitised imaging technologies (boshuizen & van de wiel, 2014; krupinski, 2018; säljö, 2019). conceptually, we offer the notion of horizontal transitions to elaborate on how experts in their domains cope with and adapt to fundamental, disruptive changes and ‘technology shocks’ (rather than minor changes occurring on a more regular basis, assuming that work and life are inherently dynamic). this line of research contributes to the field of expertise research as, to date, a very limited number of studies explore if and how experts adapt when the representativeness of their domain changes. profound gegenfurtner, gruber, lethinen & säljö 34 | f l r adaptions of these kinds are not easily explained within the framework of routine expertise development as the ‘routine-ness’ of the domain itself transforms as a consequence of contextual changes. lehtinen et al. (2020) links expert adaptations to research on conceptual change, particularly to framework theory. such a perspective is certainly useful because experts can use their rich repertoire of semantic and episodic knowledge stored in working memory—knowledge about clinical concepts, patient cases, medical diagnoses, and the underlying principles of imaging technologies— to better understand changes in their domains than novices can. however, as lehtinen et al. (2020, p. 4) argue, the described cases “call into question the role of learning new conceptual knowledge and professional practices as well as the kinds of conceptual change processes that are related to initial learning and the subsequent extension of expertise”, which we describe in detail elsewhere (gegenfurtner et al., 2009; 2017; 2019). there are instances in which experts with a long professional history in one medical arena prefer rigid attempts of using familiar practices to adapt to change, even though these are not effective (cases of failed adaption) while we also see instances of successful adaptions that include the conceptual change of multi-layered professional skills (lehtinen et al. 2020) constituting the hybridity of expertise (säljö, 2019; 2022) documented in the multi-decade biographies that were analysed in the present study. future research may follow lehtinen et al.’s (2020, p. 8) recommendation to use framework theory “for predicting when the changing of professional practices to adapt to new professional situations is relatively easy and when stronger resistance to change can be expected.” this study has implications for future research studying expertise and expert performance that should be noted. a first implication relates to a focus on changing contexts. if we conceptualise expertise as being interdependent between human agency, minds, bodies, and digital tools (billett et al., 2018; engeström, 2018; säljö, 2022), and if we further assume that digital tools frequently change in professions (gekara & thanh nguyen, 2018), then we can adopt a relational perspective on expertise, one that is interested in the recurring adaptations of expert work to dynamic task constraints. such a lens invites an analysis of experts and their professional agency (billett et al., 2018; goller & paloniemi, 2017; harteis & goller, 2014) and how experts orient “toward the future, with people not merely repeating past routines but challenging, reconsidering and reformulating their ideas, projects and plans” (damşa et al., 2017, p. 447). although it is intuitive to assume that experts need to adapt their practices regularly, there is still a paucity of studies addressing expert performance in changing contexts. this line of research can complement work on expert performance in controlled contexts with representative tasks in well-structured domains that remain relatively unchanged. such a relational, ontological approach resonates with a lifeworld perspective of expertise in which expertise is conceptualised as “a continuing process of becoming; never entirely complete, nor achieved once and for all” (dall’alba, 2018, p. 35). from this perspective, vertical and horizontal expertise development can be analysed as two poles of an analytical continuum. a second implication for further inquiry relates to replications in fields outside medicine. the medical arena has been selected here because it is a technology-intensive domain with frequent innovations of digital tools and artifacts (gegenfurtner et al., 2019; rystedt et al., 2011; szulewski et al., 2019; violato et al., 2018; white et al., 2018). but, of course, medicine is not the only domain in which the material affordances of work change. there is good reason to assume that horizontal transition of expertise can be identified in other worksites as well, including, but not limited to, aviation (hutchins, 1995), architectural design (degen et al., 2017), accounting (troshani et al., 2018), teaching (horlenko et al., 2024; keskin et al., 2024; seidel et al., in press), meteorology (hoffman et al., 2017), or pop music (längler et al., 2018). it is an interesting question to study if and to what extent the identified sociogenetic types of adaptation to technological change can be found in other high-technology domains or if these adaptation types are intrinsic to medical specialties. for example, the turn to digital teaching in primary, secondary, and higher education institutes worldwide during the covid 19 pandemic might afford an analysis of the horizontal transitions of teacher expertise. such work is likely most successful if it follows sociological traditions of expertise research in adopting qualitative analysis (gobet, 2018; van de wiel, 2017; white et al., 2018; yard et al., 2019). gegenfurtner, gruber, lethinen & säljö 35 | f l r a third implication for future research relates to conceptual alignments between theories of expertise and theories of transfer. the presented analyses of horizontal transition of expertise touch upon issues associated with transfer of learning. clearly, the identified instances of horizontal transition can also be read as instances of horizontal transfer in the tradition of ‘innovation’ and ‘multicontextuality’ (bohle carbonell et al., 2014; pusic et al., 2018; roig et al., 2024; testers et al., 2015; testers et al., 2024). if we assume that expertise develops in and across changing contexts (gruber & harteis, 2018), then modern workplaces become dynamic sites for invention and reorganisation. as lehtinen and colleagues (2014, p. 213) state: “how one recognises the familiarity or similarity when entering new situations is important (…). if we shift our focus from very explicit experimental situations of typical transfer studies towards everyday situations or long-term learning of complex scientific or professional tasks, how people interpret the situations and recognise new phenomena with the help of their previously constructed mental concepts is far from trivial.” a final implication for future work relates to education and training. the analyses indicated that horizontal transition can be unsuccessful, particularly if experts do not receive formal education, training, or mentoring to adapt their practices. even though the reconstructed narratives seem to suggest that education programs can promote adaptation to technological change, it is of course an empirical question to study how such education programs should be designed, implemented, and monitored to support experts facing technological turbulences in their working lives (gegenfurtner et al., 2019; harteis & goller, 2014; jossberger et al., 2022; lehtinen et al., 2020). addressing these questions is highly important when we seek to understand how experts maintain (gruber & harteis, 2018) or renew (frie et al., 2019) their superior levels of expertise. “the maintenance of expertise is a task which requires a number of people to contribute, both the excellent individual and persons in her or his teams, networks and societies. the expert, with her or his skills and knowledge, continues to work at a high level of performance, even within changing work conditions or societal requirements. the acquired level of expertise, thus, is not an activity at a static level, but the expert has to extend her or his skills and knowledge, which often means to restructure and modify them” (gruber & harteis, 2018, p. 110). this restructuring is often facilitated by significant others or persons in the shadow (längler et al., 2018) that might have potential to function as mediators in managing successful adaptations. what is frontline when examining cases of horizontal transition of expertise? the argument that we developed and empirically supported in this paper is that the conceptual framework of horizontal transition addresses expertise in changing work contexts in ways no other existing theory does. we do not mean to be disrespectful to the conceptual notions of adaptive expertise, knowledge encapsulation, polycontextuality, or expert cognitive flexibility – quite to the contrary: all these theories have, in our view, very elegantly signified the development of expertise. still, it is our belief that when we are interested in how experts adapt to changing task requirements in modern workplaces, the notion of horizontal transition proves useful to understand better the corollaries, contingencies, and consequences of these expert adaption processes. keypoints horizontal transition of expertise is a framework for understanding expertise development when the expert domain evolves. horizontal transition describes recurring adaptations of expertise to dynamically changing task constraints at work. four types of adaptation are reconstructed from biographical interview data: successful, problematic, contingent, and failed adaptation. gegenfurtner, gruber, lethinen & säljö 36 | f l r references anthony, g., hunter, j., & hunter, r. (2015). prospective teachers’ development of adaptive expertise. teaching and teacher education, 49, 108–117. https://doi.org/10.1016/j.tate.2015.03.010 baumann, z. (2007). liquid times: living in an age of uncertainty. polity press. beck, t. (2021). the praxeological sociology of knowledge–an introduction to the documentary method and a sketch of an empirical implementation. in p. j. white, r. tytler, j. p. ferguson, & j. c. clark (eds.), methodological approaches to stem education research (vol. 3, pp. 218– 243). cambridge scholars publishing. billett, s., harteis, c., & gruber, h. (2018). developing occupational expertise through everyday work activities and interactions. in k. a. ericsson, r. r. hoffman, a. kozbelt, & a. m. williams (eds.), cambridge handbook of expertise and expert performance (2nd ed., pp. 105–126). cambridge university press. https://doi.org/10.1017/9781316480748.008 bohle carbonell, k., stalmeijer, r. e., könings, k. d., segers, m., & van merriënboer, j. j. g. (2014). how experts deal with novel situations: a review of adaptive expertise. educational research review, 12, 14–29. https://doi.org/10.1016/j.edurev.2014.03.001 bohnsack, r. (2014). documentary method. in u. flick (ed.), the sage handbook of qualitative data analysis (pp. 217–233). sage. https://doi.org/10.4135/9781446282243 boshuizen, h. p. a., gruber, h., & strasser, j. (2020). knowledge restructuring through case processing: the key to generalise expertise development theory across domains? educational research review, 29, 100310. https://doi.org/10.1016/j.edurev.2020.100310 boshuizen, h. p. a., & schmidt, h. g. (1992). on the role of biomedical knowledge in clinical reasoning by experts, intermediates and novices. cognitive science, 16(2), 153–184. https://doi.org/10.1016/0364-0213(92)90022-m boshuizen, h. p. a., & van de wiel, m. w. j. (2014). expertise development through schooling and work. in a. littlejohn & a. margaryan (eds.), technology-enhanced professional learning: processes, practices, and tools (pp. 71–84). routledge. burgoyne, a. p., nye, c. d., macnamara, b. n., charness, n., & hambrick, d. z. (2019). the impact of domain-specific experience on chess skill: reanalysis of a key study. american journal of psychology, 132(1), 27–38. https://doi.org/10.5406/amerjpsyc.132.1.0027 dall’alba, g. (2018). reframing expertise and its development: a lifeworld perspective. in k. a. ericsson, r. r. hoffman, a. kozbelt, & a. m. williams (eds.), cambridge handbook of expertise and expert performance (2nd ed., pp. 33–39). cambridge university press. https://doi.org/10.1017/9781316480748.003 damşa, c. i., froehlich, d. e., & gegenfurtner, a. (2017). reflections on empirical and methodological accounts of agency at work. in m. goller & s. paloniemi (eds.), agency at work: an agentic perspective on professional learning and development (pp. 445–461). springer. https://doi.org/10.1007/978-3-319-60943-0_22 degen, m., melhuish, c., & rose, g. (2017). producing place atmospheres digitally: architecture, digital visualisation practices and the experiences economy. journal of consumer culture, 17(1), 3–24. https://doi.org/10.1177/1469540515572238 gegenfurtner, gruber, lethinen & säljö 37 | f l r de groot, a. d. (1965). thought and choice in chess. mouton. dochy, f., engeström, y., sannino, a., & van meeuwen, n. (2021). inter-organisational expansive learning at work. in f. dochy, d. gijbels, m. segers, & p. van den bossche (eds.), theories of workplace learning in changing times (2nd ed., pp. 209–231). routledge. https://doi.org/10.4324/9781003187790-12 engeström, y. (2018). expertise in transition: expansive learning in medical work. cambridge university press. engeström, y., engeström, r., & kärkkäinen, m. (1997). the emerging horizontal dimension of practical intelligence: polycontextuality and boundary crossing in complex work activities. in r. j. sternberg & e. l. grigorenko (eds.), intelligence, heredity, and environment (pp. 440–462). cambridge university press. ericsson, k. a. (2018). capturing expert thought with protocol analysis: concurrent verbalizations of thinking during experts’ performance on representative tasks. in k. a. ericsson, r. r. hoffman, a. kozbelt, & a. m. williams (eds.), cambridge handbook of expertise and expert performance (2nd ed., pp. 192–212). cambridge university press. https://doi.org/10.1017/9781316480748.012 feltovich, p. j., prietula, m. j., & ericsson, k. a. (2018). studies of expertise from psychological perspectives: historical foundations and recurrent themes. in k. a. ericsson, r. r. hoffman, a. kozbelt, & a. m. williams (eds.), cambridge handbook of expertise and expert performance (2nd ed., pp. 59–83). cambridge university press. https://doi.org/10.1017/9781316480748.006 frie, l. s., potting, k. c. j. m., sjoer, e., van der heijden, b. i. j. m., & korzilius, h. p. l. m. (2019). how flexperts deal with changing expertise demands: a qualitative study into the processes of expertise renewal. human resource development quarterly, 30(1), 61–79. https://doi.org/10.1002/hrdq.21335 gabel, s., keskin, ö., & gegenfurtner, a. (in press). comparing the effects of a specific task instruction and prompts on pre-service teachers’ noticing of classroom management situations. zeitschrift für erziehungswissenschaft. garfinkel, h. (1967). studies in ethnomethodology. prentice hall. gegenfurtner, a., gruber, h., holzberger, d., keskin, ö., lehtinen, e., seidel, t., stürmer, k., & säljö, r. (2023). towards a cognitive theory of visual expertise: methods of inquiry. in c. damşa, a. rajala, g. ritella, & j. brouwer (eds.), re-theorising learning and research methods in learning research (pp. 146–163). routledge. https://doi.org/10.4324/9781003205838-10 gegenfurtner, a., lehtinen, e., helle, l., nivala, m., svedström, e., & säljö, r. (2019). learning to see like an expert: on the practices of professional vision and visual expertise. international journal of educational research, 98, 280–291. https://doi.org/10.1016/j.ijer.2019.09.003 gegenfurtner, a., lehtinen, e., jarodzka, h., & säljö, r. (2017). effects of eye movement modeling examples on adaptive expertise in medical image diagnosis. computers & education, 113, 212– 225. https://doi.org/10.1016/j.compedu.2017.06.001 gegenfurtner, a., nivala, m., säljö, r., & lehtinen, e. (2009). capturing individual and institutional change: exploring horizontal versus vertical transitions in technology-rich environments. in u. cress, v. dimitrova, & m. specht (eds.), learning in the synergy of multiple disciplines. lecture gegenfurtner, gruber, lethinen & säljö 38 | f l r notes in computer science (pp. 676–681). springer. https://doi.org/10.1007/978-3-642-046360_67 gekara, v. o., & thanh nguyen, v.-x. (2018). new technologies and the transformation of work and skills: a study of computerisation and automation of australian container terminals. new technology, work and employment, 33(3), 219–233. https://doi.org/10.1111/ntwe.12118 gobet, f. (2018). the future of expertise: the need for a multidisciplinary approach. journal of expertise, 1(2), 107–113. goller, m., & paloniemi, s. (eds.). (2017). agency at work: an agentic perspective on professional learning and development. springer. https://doi.org/10.1007/978-3-319-60943-0 gruber, h., & harteis, c. (2018). individual and social influences on professional learning. supporting the acquisition and maintenance of expertise. springer. https://doi.org/10.1007/978-3-31997041-7 gruber, h., jansen, p., marienhagen, j., & altenmüller, e. (2010). adaptations during the acquisition of expertise. talent development & excellence, 2(1), 3–15. harteis, c., & goller, m. (2014). new skills for new jobs: work agency as a necessary condition for successful lifelong learning. in t. halttunen, m. koivisto, & s. billett (eds.), promoting, assessing, recognizing and certifying lifelong learning: international perspectives and practices (pp. 37–56). springer. https://doi.org/10.1007/978-94-017-8694-2_3 hatano, g., & inagaki, k. (1986). two courses of expertise. in h. stevenson, h. asuma, & k. hakuta (eds.), child development and education in japan (pp. 262–272). freeman. hatano, g., & oura, y. (2003). commentary: reconceptualizing school learning using insight from expertise research. educational researcher, 32(8), 26–29. https://doi.org/10.3102/0013189x032008026 hoffman, r. r., ladue, d. s., mogil, m., roebber, p. j., & trafton, j. g. (2017). minding the weather: how expert forecasters think. mit press. horlenko, l., kaminskienė, l., & lehtinen, e. (2024). student self-regulated learning in teacher professional vision: results from combining student self-reports, teacher ratings, and mobile eye tracking in the high school classroom. frontline learning research, 12(2), 51–69. https://doi.org/10.14786/flr.v12i2.1417 hutchins, e. (1995). cognition in the wild. mit press. ivarsson, j., rystedt, h., asplund, s., johnsson, å., & båth, m. (2016). the application of improved, structured and interactive group learning methods in diagnostic radiology. radiation protection dosimetry, 169(1–4), 416–421. https://doi.org/10.1093/rpd/ncv497 jaarsma, t. (2015). expertise development under the microscope: visual problem solving in clinical pathology. open university of the netherlands. jossberger, h., breckwoldt, j., & gruber, h. (2022). promoting expertise through simulation (pets): a conceptual framework. learning and instruction, 82, 1016876. https://doi.org/10.1016/j.learninstruc.2022.101686 gegenfurtner, gruber, lethinen & säljö 39 | f l r keskin, ö., seidel, t., stürmer, k., & gegenfurtner, a. (2024). eye-tracking research on teacher professional vision: a meta-analytic review. educational research review, 42, 100586. https://doi.org/10.1016/j.edurev.2023.100586 krupinski, e. a. (2018). perceptual factors in reading medical images. in e. samei & e. a. krupinski (eds.), handbook of medical image perception and techniques (2nd ed., pp. 95–106). cambridge university press. https://doi.org/10.1017/9781108163781.008 längler, m., nivala, m., & gruber, h. (2018). peers, parents and teachers: a case study on how popular music guitarists perceive support for expertise development from “persons in the shadows”. musicae scientiae, 22(2), 224–243. https://doi.org/10.1177/1029864916684376 lehtinen, e. (2022). how to deal with the complexity in research on workplace learning. in m. goller, e. kyndt, s. paloniemi, & c. damşa (eds.), methods for researching professional learning and development (pp. 619–627). springer. https://doi.org/10.1007/978-3-031-08518-5_28 lehtinen, e., gegenfurtner, a., helle, l., & säljö, r. (2020). conceptual change in the development of visual expertise. international journal of educational research, 100, 101545. https://doi.org/10.1016/j.ijer.2020.101545 lehtinen, e., hakkarainen, k., & palonen, t. (2014). understanding learning for the professions: how theories of learning explain coping with rapid change. in s. billett, c. harteis, & h. gruber (eds.), international handbook of research in professional practice-based learning (pp. 199–224). springer. https://doi.org/10.1007/978-94-017-8902-8_8 lin, x. d., schwartz, d. l., & bransford, j. d. (2007). intercultural adaptive expertise: explicit and implicit lessons from dr. hatano. human development, 50(1), 65–72. https://doi.org/10.1159/000097686 männikkö, i., & husu, j. (2019). examining teachers’ adaptive expertise through personal practical theories. teaching and teacher education, 77, 126–137. https://doi.org/10.1016/j.tate.2018.09.016 mustonen, v., & hakkarainen, k. (2015). tracing two apprentices’ trajectories toward adaptive professional expertise in fingerprint examination. vocations and learning, 8(2), 185–211. https://doi.org/10.1007/s12186-015-9130-7 mylopoulos, m., & woods, n. n. (2017). when i say… adaptive expertise. medical education, 51(7), 685–686. https://doi.org/10.1111/medu.13247 nohl, a.-m. (2010). the documentary interpretation of narrative interviews. in r. bohnsack, n. pfaff, & w. weller (eds.), qualitative analysis and documentary method in international educational research (pp. 195–217). budrich. nohl, a.-m. (2017). interview und dokumentarische methode [interview and documentary method] (5th ed.). springer vs. https://doi.org/10.1007/978-3-658-16080-7 palonen, t., boshuizen, h. p. a., & lehtinen, e. (2014). how expertise is created in emerging professional fields. in s. billett, t. halttunen, & m. koivisto (eds.), promoting, assessing, recognizing and certifying lifelong learning: international perspectives and practices (pp. 131– 150). springer. https://doi.org/10.1007/978-94-017-8694-2_8 gegenfurtner, gruber, lethinen & säljö 40 | f l r philipps, a., & mrowczynski, r. (2021). getting more out of interviews. understanding interviewees’ accounts in relation to their frames of orientation. qualitative research, 21(1), 59–75. https://doi.org/10.1177/1468794119867548 pusic, m. v., santen, s. a., dekhytar, m., poncelet, a. n., roberts, n. k., wilson-delfosse, a. l., & cutrer, w. b. (2018). learning to balance efficiency and innovation for optimal adaptive expertise. medical teacher, 40(8), 820–827. https://doi.org/10.1080/0142159x.2018.1485887 rikers, r. m. j. p., schmidt, h. g., boshuizen, h. p. a., linssen, g. c. m., wesseling, g., & paas, f. g. w. c. (2002). the robustness of medical expertise: clinical case processing by medical experts and subexperts. american journal of psychology, 115(4), 609–629. https://doi.org/10.2307/1423529 roig-ester, h., robalino guerra, p. e., quesada pallarès, c., & gegenfurtner, a. (2024). transfer of learning of new nursing professionals: exploring patterns and the effect of previous working experience. education sciences, 14(1), 52. https://doi.org/10.3390/educsci14010052 rystedt, h., ivarsson, j., asplund, s., johnsson, å. a., & båth, m. (2011). rediscovering radiology. new technologies and remedial action at the worksite. social studies of science, 41(6), 867–891. https://doi.org/10.1177/0306312711423433 säljö, r. (2019). materiality, learning, and cognitive practices: artifacts as instruments of thinking. in t. cerrato-pargman & i. jahnke (eds.), emergent practices and material conditions in learning and teaching with technologies (pp. 21–32). springer. https://doi.org/10.1007/978-3-030-107642_2 säljö, r. (2022). development, ageing and hybrid minds: growth and decline, and ecologies of human functioning in a sociocultural perspective. learning, culture and social interaction, 37, 100465. https://doi.org/10.1016/j.lcsi.2020.100465 seidel, t., kosel, c., böheim, r., gegenfurtner, a., & stürmer, k. (in press). a cognitive model of professional vision and acquisition of visual expertise using video excerpts in the teaching profession. in a. gegenfurtner & r. stahnke (eds.), teacher professional vision: theoretical and methodological advances. routledge. sellberg, c., nordenström, e., & säljö, r. (2024). the development of visual expertise in a virtual environment: a case of maritime pilots in training. frontline learning research, 12(1), 16–33. https://doi.org/10.14786/flr.v12i1.1217 spiro, r. j., feltovich, p. j., gaunt, a., hu, y., klautke, h., cheng, c., et al. (2019). cognitive flexibility theory and the accelerated development of adaptive readiness and adaptive response to novelty. in p. ward, j. m. schraagen, j. gore, & e. m. roth (eds.), the oxford handbook of expertise (pp. 951–976). oxford university press. https://doi.org/10.1093/oxfordhb/9780198795872.013.41 spiro, r. j., morsink, p., & forsyth, b. (2012). point of view: principled pluralism, cognitive flexibility, and new contexts for reading. in r. f. flippo (ed.), reading researchers in search of common ground. the expert study revisited (2nd ed., pp. 118–128). routledge. stahnke, r., & gegenfurtner, a. (in press). beyond analysing frequencies: exploring teacher professional vision with epistemic network analysis of teachers’ think-aloud data. learning and instruction. https://doi.org/10.1016/j.lcsi.2020.100465 gegenfurtner, gruber, lethinen & säljö 41 | f l r strasser, j., & gruber, h. (2015). learning processes in the professional development of mental health counselors: knowledge restructuring and illness script formation. advances in health sciences education, 20(2), 515–530. https://doi.org/10.1007/s10459-014-9545-1 suh, j. k., hand, b., dursun, j. e., lammert, c., & fulmer, g. (2023). characterizing adaptive teaching expertise: teacher profiles based on epistemic orientation and knowledge of epistemic tools. science education, 107(4), 884-911. https://doi.org/10.1002/sce.21796 szulewski, a., braund, h., egan, r., gegenfurtner, a., hall, a. k., howes, d., dagnone, j. d., & van merriënboer, j. j. g. (2019). starting to think like an expert: an analysis of resident cognitive processes during simulation-based resuscitation examinations. annals of emergency medicine, 74(5), 647–659. https://doi.org/10.1016/j.annemergmed.2019.04.002 szulewski, a., braund, h., egan, r., hall, a. k., dagnone, j. d., gegenfurtner, a., & van merriënboer, j. j. g. (2018). through the learner’s lens: eye-tracking augmented debriefing in medical simulation. journal of graduate medical education, 10(3), 340– 341. https://doi.org/10.4300/jgme-d-17-00827.1 testers, l., alijagic, a., brand-gruwel, s., & gegenfurtner, a. (2024). predicting transfer of generic information literacy competencies by non-traditional students to their study and work contexts: a longitudinal perspective. education sciences, 14(2), 117. https://doi.org/10.3390/educsci14020117 testers, l., gegenfurtner, a., & brand-gruwel, s. (2015). motivation to transfer learning to multiple contexts. in l. das, s. brand-gruwel, k. kok, & j. walhout (eds.), the school library rocks: living it, learning it, loving it (pp. 473–487). iasl. troshani, i., janssen, m., lymer, a., & parker, l. d. (2018). digital transformation of business-togovernment reporting: an institutional work perspective. international journal of accounting information systems, 31, 17–36. https://doi.org/10.1016/j.accinf.2018.09.002 van de wiel, m. w. j. (2017). examining expertise by interviews and verbal protocols. frontline learning research, 5(3), 94–122. https://doi.org/10.14786/flr.v5i3.257 violato, c., gao, h., o’brien, m. c., grier, d., & shen, e. (2018). how do physicians become medical experts? a test of three theories: distinct domains, independent influence and encapsulated models. advances in health sciences education, 23(2), 249–263. https://doi.org/10.1007/s10459017-9784-z ward, p., schraagen, j. m., gore, j., roth, e. m., hoffman, r. r., & klein, g. (2019). reflections on the study of expertise and its implications for tomorrow’s world. in p. ward, j. m. schraagen, j. gore, & e. m. roth (eds.), the oxford handbook of expertise (pp. 1193–1213). oxford university press. https://doi.org/10.1093/oxfordhb/9780198795872.013.52 white, m. r., braund, h., howes, d., egan, r., gegenfurtner, a., van merriënboer, j. j. g., & szulewski, a. (2018). getting inside the expert’s head: an analysis of physician cognitive processes during trauma resuscitations. annals of emergency medicine, 72(3), 289–298. https://doi.org/10.1016/j.annemergmed.2018.03.005 yardley, s., mattick, k., & dornan, t. (2019). close-to-practice qualitative research. in p. ward, j. m. schraagen, j. gore, & e. m. roth (eds.), the oxford handbook of expertise (pp. 408–428). oxford university press. https://doi.org/10.1093/oxfordhb/9780198795872.013.18 gegenfurtner, gruber, lethinen & säljö 42 | f l r appendix formulating interpretation of the five expert interviews interview 1 the first interview was performed with an expert in clinical pathology. in the interview, the expert describes the transition from light microscopes to virtual microscopes (minute 38:25 to 39:55). he then elaborates on how virtual microscopy is used for teaching (minute 57:35 to 60:40). finally, the expert describes age-related differences in the adaptation process (minute 67:05 through 71:30). 38:25’ 39:55’ virtual microscopy is introduced in addition to light microscopes. the perception of depth in virtual slides is not as good as in electronic microscopes. 57:35’ 60:40’ virtual microscopy is used in seminars and lectures. it is useful and students like it. however, as a teacher with a 25-year experience in analogical slides, using digital slides is problematic. younger teachers familiar with the interest and digital media will adapt more quickly because they are digital natives. 67:05’ 71:30’ preference for digital media might be a generation thing. people are lazy animals. if you have a technique that you are familiar with and that you like, why not stick to it. a positive mind for digital media will make adaptation easier. interview 2 the second interview was performed with an expert in nuclear medicine. the expert describes how nuclear medicine has embraced technological innovations, particularly pet/ct and pet/mri (minute 6:10 through 18:55). he then describes how the introduction of new imaging technologies is guided and accompanied by newly implemented regulations and bureaucratic decisions (minute 28:50 through 37:55). 06:10’ 18:55’ pet has developed into a fusion technology with pet/ct and pet/mri. this requires collaboration with neighbouring disciplines. adaptation to pet/ct and pet/mri is regulated and organised by law statutes. training programs are developed to support the move from pet to pet/ct. 28:50’ 37:55’ new regulations follow new technologies. adaptation to these regulations is necessary, but also frustrating because the rationale behind regulatory changes does not always seem evident. regulators like to regulate. interview 3 the third interview was performed with an expert in paediatric radiology. first, the expert describes the transition from analogous x-ray radiography to digital computer tomography (ct; minute 35:50 to 37:05). second, he elaborates on new developments, including the introduction of pet/ct and the move from two-dimensional ct (2d) ct scans to three(3d) and four-dimensional (4d) ct scans (minute 37:10 to 38:45). finally, he focuses on difficulties associated with adapting to 3d and 4d ct (minute 38:55 to 46:40). gegenfurtner, gruber, lethinen & säljö 43 | f l r 35:50’ 37:05’ traditional x-ray scans are digitised. ct scans are the new standard. x-ray pictures and 2d ct are very similar. ct has a higher resolution. 37:10’ 38:45’ ct is fused with pet into pet/ct images. training programs exist for pet/ct reading. 3d and 4d ct scans are introduced, but no formal training or mentoring programs exist. 38:55’ 46:40’ adaptation to 3d and 4d scans is not easy for everyone. some colleagues look at 3d and 4d images as they do 2d: one slide after the other. new practices, a new workflow, are not yet developed. there is no training. the new generation adapts more easily because they are digital natives and have developed this way of digital thinking. interview 4 the fourth interview was performed with an expert in clinical pathology. the expert describes how teaching practices in classrooms have changed through the introduction of virtual microscopy (minute 15:25 to 18:35). he then elaborates on sceptical attitudes toward virtual microscopy for clinical practice (minute 20:30 to 23:35). 15:25’ 18:35’ pathology teaching changed from the use of tissue samples stored in boxes to the projection of digitised tissue samples. in lecture halls, virtual microscopy rendered multi-tube electronic microscopes partially obsolete. 20:30’ 23:35’ the transition from using light microscopes to virtual microscopy was slow. in clinical practice, the attitude toward virtual microscopy was negative because the added value of virtual microscopy was not immediately clear. pathological tissue material can be exchanged globally within a few seconds. interview 5 the fifth interview was performed with an expert in nuclear medicine. the expert first describes developments associated with positron emission tomography (pet; minute 21:50 to 23:25). in minutes 23:25 to 24:30 and from 27:00 through 32:25, he describes very recent innovations including the fusion of pet with computer tomography (ct) into pet/ct and with magnetic resonance imaging (mri) into pet/mri. finally, the expert discusses factors that support or hinder adaptations to technological change (minute 64:50 to 73:55). 21:50’ 23:25’ the graphical resolution of pet images increased. software for analysing the images was updated. new versions of pet image and software were similar to old versions. 23:25’ 24:30’ when new imaging technologies are introduced, formal training is implemented to support staff. 27:00’ 28:50’ pet is combined with computer tomography (ct). radiology and nuclear medicine collaborate. ct images are more anatomical than pet. radiology experts start with the ct part, while nuclear medicine experts start with the pet part. this is a new challenge for nuclear medicine. 30:10’ 32:25’ pet is also combined with magnetic resonance imaging (mri). diagnosing with pet/mri will become much more difficult because the underlying principles of generating pet and mr images are very different. the older you become, the more gegenfurtner, gruber, lethinen & säljö 44 | f l r difficult it is to learn new things. curiosity is still high, but also scepticism if technological innovations have any added value. 64:50’ 73:55’ interest, curiosity, and scepticism are associated with novel technologies. if innovations are always rejected, you risk being outdated and losing patients. nuclear medicine is a technological discipline, so people here have a positive attitude toward technology, a fascination. frontline learning research vol. 12 no. 4 (2024) 1 21 issn 2295-3159 corresponding author: z. vermeire, martinus j. langeveldgebouw, heidelberglaan 1, 3584 cs utrecht, the netherlands; z.vermeire@uu.nl doi: https://doi.org/10.14786/flr.v12i4.1325. platformised affinity spaces: learning communities on youtube, twitch and tiktok z. vermeire¹, m.j. de haan¹, j. sefton-green², & s.f. akkerman¹ ¹utrecht university, the netherlands ²deakin university, australia article received 11 july 2023 / article revised 19 september 2024 / accepted 15 october 2024 / available online 15 november 2024 abstract online, informal learning communities bring youth opportunities for learning that schools cannot offer. yet, there are concerns about the impact of social media platforms’ control over online learning. we argue for a re-evaluation of what an ‘online informal learning community’ is by looking at such active communities on three platforms: youtube, twitch and tiktok. we do this by reconsidering gee’s ‘affinity spaces’ and by asking: how can we understand online informal learning communities in the current sociotechnical context? we observed and analysed interactions of six learning communities on youtube, twitch and tiktok. our results show that in today’s platformised online context, gee’s concept of ‘affinity spaces’ should be reconsidered in three ways. first, platforms call for discussion about affinity spaces’ boundaries through the visibility regimes that play a part in access. second, platforms challenge the affinity spaces’ grammar; to maintain a focus on their interest, platforms need to engage with interests provided by platform cultures. third, a more fixated hierarchisation, informed by platforms’ focus on creators, impacts affinity spaces’ social structures. we introduce the concept of ‘platformised affinity space’ as a first step to specific dynamics that platforms introduce to online informal learning communities. we conclude that we only understand these communities when acknowledging how these dynamics are appropriated as well as resisted to achieve community goals. keywords: affinity space; informal learning; social media platforms; learning community; ethnography mailto:z.vermeire@uu.nl vermeire, de haan, sefton-green, & akkerman 2 | f l r 1. introduction although we generally think of school as the main place for learning, young people (increasingly) also learn in communities on social media platforms. on social media platforms they share knowledge and skills around a common interest, such as books, games, history, or coding. we aim to gain a better understanding of the role that social media platforms might play in such communities. to gain such an understanding, we conducted an empirical study of online informal1 learning communities on youtube, twitch and tiktok that sheds a ‘new’2 light on knowledge from the learning sciences and media studies. we draw on work from media studies that raises concerns about the control of social media platforms over youth and learning and the potentially harmful effects of such control (alegre, 2021; decuypere et al., 2021; van dijck et al., 2018; koopman, 2019; sefton-green & pangrazio, 2022). they describe this increasing control of a few big tech companies as ‘platformisation’ (van dijck et al., 2018). from the learning sciences, we build on literature that explores how (online) learning communities have challenged traditional, formal notions of learning. from this body of work, we turn specifically to gee’s concept of learning communities as ‘affinity spaces’, which he introduced in response to the fluidity that online spaces brought to the characteristics of learning communities. we want to explore whether and how gee’s (2005) concept of affinity space requires an update regarding informal learning communities operating in an increasingly platformised context. in particular, since gee introduced his ‘affinity space,’ online spaces have become increasingly controlled by big tech companies (van dijck et al., 2018), increasingly ‘platformised’, which impacted their (experienced) fluidity. we seek to map these changes by empirically researching the role of social media platforms in young people’s online informal learning communities. 1.1 theoretical background to understand this new online context, we use insights from media studies, to analyse how online informal learning communities change when platforms offer the context for young people’s online learning activities. van dijck et al. understand platforms as programmable digital architectures designed to organise interactions between users (2018). platforms collect and process user data for the commercial, surveillance and normative purposes of tech companies (van dijck et al., 2018). in work from media studies, platforms’ algorithms are argued to steer users to produce particular behaviour and knowledge by presenting information in a manner that can be considered political by their classification, sorting, and ranking of data (bucher, 2018). we are interested in exploring whether such control of platform infrastructures over how and to whom information is presented might impact young people’s learning activities in communities online. in other words, we are interested in how information is bound and/or released to them based on algorithmic sorting rooted in platforms’ commercial and surveillant aims. we will refer to this steering of users through the algorithmic management of information as ‘visibility regimes’ here. such ‘regimes of visibility’ and their impact on young people have also raised concerns in popular media, such as the netflix documentary-drama hybrid ‘the social dilemma’ (orlowski, 2020), and academic work, such as shoshana zuboff’s book on ‘surveillance capitalism’ (2019). these argue that platforms manipulate young people to become addicted users in ways that reinforce prejudice and radicalise and divide youth. furthermore, there is a growing concern in media studies and in studies of learning, that automated manipulation of behaviour by platforms is detrimental to youth’s agency and critical thinking potential (alegre, 2021; koopman, 2019; sefton-green & pangrazio, 2022). for example, there are concerns about how users’ control over these platforms in making autonomous (learning) choices is limited by algorithms’ visibility regimes in an attempt to manipulate user behaviour (alegre, 2021; koopman, 2019). a well-known example, mentioned also in the social dilemma, is how algorithms might aim to control a young person’s attention by showing 1 see bronkhorst & akkerman (2016) for our understanding of the relation between formal and informal learning, as a representation of this discussion falls outside of the scope of this paper. 2 we have used quotation marks around the first usage of the word ‘new’ here as we want to acknowledge that knowledge is never truly ‘new’ as it is rooted in a long history of developing technologies, tools, and theories. vermeire, de haan, sefton-green, & akkerman 3 | f l r them increasingly radicalised or extreme content, raising societal concerns on what and how youth are learning online. we wonder how such insights and concerns about the control of platforms translate to the ways in which informal learning communities operate online. how could platforms’ manipulations for instance implicate youth’s control over their learning in such communities? studies so far have mostly theorised these concerns on platforms’ visibility regimes and their impact on young people’s participation in online formal learning communities (sefton-green & pangrazio, 2022; williamson et al., 2022). simultaneously they call for more empirical research to see whether the control of young people over their learning activities and their critical perspectives are really under such a threat on digital platforms both in and beyond explicitly educational environments (seftongreen & pangrazio, 2022; williamson et al., 2022). we partly respond to this call by conducting observations of online informal learning communities to understand whether we need to critically reassess the conceptualisation of ‘learning community’ for the platform context. in other words, by bringing in empirical material on informal learning communities from youtube, twitch and tiktok, we reflect on the epistemological underpinnings on what an online informal learning community is to explore whether there is, in this sociotechnical context, a need for a new conceptualisation of what an online informal learning community is. the relationship between young people, learning communities and digital technologies has previously been explored in academic literature (boyd, 2014; deng et al., 2016; ito et al., 2019; jenkins, 2006; lankshear & knobel, 2007). for example, there is a large body of work on ‘networked learning’ that has provided insights into how learning interactions are shaped when people learn in the context of online, networked technologies (del valle et al., 2020; rehm et al., 2018). such work for instance predicts factors for how learning connections are shaped online on reddit (del valle et al., 2020) or how informal learning networks can be shaped in formal educational contexts to push back against the power of platforms over data (wilson et al., 2023). although the literature on networked learning has traditionally paid limited attention to informal learning and the normative aspects of learning, there has been an increase in (calls for) such work (del valle et al., 2020). we take up this new, still small, strand of studies on networked learning that addresses normative issues, and in particular the normative aspects of platform design for learning (see e.g. gyldendahl jensen et al., 2022; vaessen et al., 2014; wichmand et al., 2023). to further explore the normative issues of learning and the role of technology in learning, we have looked at literature that departs from the assumption that learning is networked and draws attention to how digital technologies might challenge formal education. this work argues that the opportunities offered by digital technologies for young people to learn with others online beyond the boundaries of school, could challenge how schools organise learning activities (jenkins, 2006; lankshear & knobel, 2007; säljö, 2010). for example, young people can learn a language that no one in their town speaks or teaches from an online expert or access resources online that are not available to them in their immediate environment. it has been argued that these technologies expand people’s social networks in ways that were not possible before such technologies existed (rainie & wellman, 2012). simultaneously, academics argue that these informal, ‘networked’, online learning communities challenge traditional learning institutions (akkerman & leijen, 2010; säljö, 2010). for example, it has been argued online informal learning communities challenge the school’s monopoly on knowledge (säljö, 2010). in contrast to such claims, the connected learning paradigm sought to connect these opportunities for online learning to academic, civic and professional opportunities (ito et al., 2019). rather than perceiving digital technology as a ‘threat’ to schools and their public values, these scholars explored the benefits of informal online learning communities for youth and their futures (ito et al., 2019). we continue this conversation started by these authors, adding a unique and fresh perspective by looking at the variety of ways in which these communities address the challenges of learning in the current online context that becomes increasingly platformised. we use the idea of ‘affinity space’ as a starting point to rethink ‘online informal learning communities’ in the platform context. ‘affinity space’ is a concept introduced by gee as a comment on the idea of ‘communities of practice’ (2005). communities of practice are defined by wenger, in whose work the roots for ‘communities of practice’ and gee’s 'affinity space’ lie (in part), as a group of people vermeire, de haan, sefton-green, & akkerman 4 | f l r who come together around a shared problem to learn about it as a joint enterprise (wenger, 1999). they learn through the exchange of experience and expertise between core expert members and peripheral novices, thereby acquiring skills and knowledge (wenger, 1999). communities of practice as a concept inherently linked the social process of engaging in a community with learning, seeing the joint enterprise of sharing experiences and competencies between novices and experts, and the resulting individual and community transformations, as learning (wenger, 1999). our understanding of what constitutes a learning community and learning is consistent with this work and understands learning as a social, transformative process of sharing knowledge, skills and resources between individuals who come together around a common interest, problem or goal (akkerman et al., 2021; wenger, 1999). thus, we assume that all communities that gather around a particular interest (affinity) allow for learning. we understand this both in the sense of becoming more knowledgeable and in a broader, more transformative sense of coming to belong and developing one’s identity. although these conceptual roots of learning (communities) as a socio-historical process have remained similar in later work, influenced by the rise of digital communities, there has been criticism of the concept of communities of practice in terms of its applicability to this ‘new’ context of learning communities (angouri, 2015). gee introduced the idea of ‘affinity space’ as one such critique to mark the more fluid nature of communities in informal, online settings (2005). for example, he argues that the relationships between novices and experts are more fluid than originally discussed in the ‘communities of practice’ literature (gee, 2005). furthermore, he notes that the idea of ‘community’ consistently requires a determination of boundaries, such as who is a member, who ‘belongs’ and who does not, or the physical space in which they meet, yet such boundaries are often difficult to establish (2005). as a solution to this problem gee shifts the focus to what binds various people within a community, which he calls the ‘grammar’ (2005). the ‘grammar’ is the interactions, values, thoughts and practices that constitute the relationships between people to create a community around a shared affinity (2005). we lead this ongoing discussion on the characterisations of learning communities into a new era of digital technologies, in which big tech companies are leaving their marks on public discourse (van dijck et al., 2018), and perhaps also on online informal learning communities. we understand online informal learning communities as communities of practice who informally (out-of-school-context) learn about a shared interest in an online context. following gee (2005), we refer to affinity space as a further specification of this term: we are specifically interested in describing this online informal learning community as an affinity space, in which boundaries are more fluid and belonging is characterised by a shared grammar rather than membership. by moving this discussion about the characterisation of learning communities into a time of platformisation, new questions arise. for example, we might wonder whether gee’s further development of the idea of ‘learning communities’ for the online context may have overemphasised fluidity over boundaries. the discussion about influencers online as celebrities also taking up pedagogical roles introduces for example new dynamics into either the fluidity or bounded relation between expert and novice in online communities (goodyear, 2022; hendry et al., 2022). additionally, the normative control that platforms exert over learning in formal contexts (williamson et al., 2022; decuypere et al., 2021) might call for a re-evaluation of the idea of young people taking control of their learning online, so typically ascribed to online informal learning communities. if digital platforms developed for formal education are already pedagogically impacting young people’s learning, we might assume that this could even go further for online contexts where these platforms are not primarily concerned for the pedagogical consequences of their platforms. accordingly, gee may not have addressed the implications of platformisation in which many of the current online informal learning communities operate. this context of platformisation has been discussed for formal learning contexts, for example how platforms could take away learners’ control over their learning trajectories (alegre, 2021; sefton-green & pangrazio, 2022) or normatively shape learning through platform design (decuypere et al., 2021; perrotta et al., 2021; sefton-green, 2021). how platforms are designed to collect and process user data for their commercial, surveillant and normative aims and whether and how this might affect online informal learning communities, deserves attention. furthermore, an empirical, interdisciplinary study analysing online informal learning communities operating in this platformising context, using the vermeire, de haan, sefton-green, & akkerman 5 | f l r concept of platformisation to describe a changing learning environment from media studies, and affinity space to draw on the ongoing discussion about what constitutes learning communities from the learning sciences, has not yet been done and is needed. this study adds a needed and new understanding of the impact of platformisation on how online informal learning communities operate. in summary, the aim of this study is to ask how we can understand online informal learning communities in a platformised context. to this end, we consider whether there is a need for a new conceptualisation of online informal learning communities, by starting from the question whether ‘affinity space’ is still sufficient to describe these communities in the platform context, and if not, how we need to understand ‘affinity space’ in times of platformisation to capture how informal online learning communities interact and exist on platforms. this study reopens earlier discussions about how to conceptualise informal learning communities in response to new socio-technical contexts. however, more fundamentally, it adds a new and needed update to understanding online informal learning communities, based on empirical data, of how platforms might control the informal learning activities of young people online based on their commercial and surveillant, rather than educational, aims. it is striking such an understanding is still lacking given the attention to the impact of educational platforms on youth’s formal learning (kerssens & dijck, 2021; perrotta et al., 2021). however, if we wish to understand how youth learn across contexts as young people who are not confined to just learning in school, but also online, such an understanding is crucial and still missing. in other words, this article hopes to raise fundamental questions about how the socio-technical context of online informal learning communities has changed and to ask whether perhaps a new conceptualisation of what an online informal learning community is, is perhaps urgent and necessary. 2. method to answer our question, ‘how can we understand online informal learning communities in a platformised online context?’ we used observational data that has been collected in a larger ethnographic research of the videos/streams and comments/chat of six online informal learning communities on the social media platforms youtube, twitch and tiktok. the ‘platformised context’ refers, based on work by van dijck et al. (2018) to how these communities operate in an online space that is increasingly becoming more controlled by a few digital platforms whose interest it is to collect and process user data for surveillant, commercial and normative aims for a few big tech companies. for our data analysis, on which we will elaborate in more detail later, we used critical discourse analysis to distinguish how platforms’ power might create regimes of truth shaping the grammar of online informal learning communities. subsequently, we describe how such shaping might ask for a new conceptualisation of what an online informal learning community is. ‘grammar’ is a term used by gee to describe the interactions, values, thoughts, and practices constituting relationships between people in an affinity space (2005). we draw not only on gee’s ‘grammar’ but also on another concept gee uses to conceptualise ‘affinity spaces’: ‘portal’ (2005). ‘portal’ is a term used by gee to describe access points to an affinity space which exercise power over how interactions around that affinity are taking place, which is why gee argues portals also generate the grammar (2005). 2.1 selection criteria communities to consider differences and similarities in how platforms might shape online informal learning communities, though representativeness was neither the aim nor possible, we have implemented various sampling criteria to capture a diverse set of informal learning communities on platforms. we first selected the platforms, then the communities and lastly creators, which are those people who create streams on twitch, and videos on tiktok, and youtube. vermeire, de haan, sefton-green, & akkerman 6 | f l r 2.1.1 platform selection to be able to compare the different workings of specific platforms and to capture how commonalities between platforms might come to the fore in the learning communities, we included three platforms. to maximise the chance that we had data about young users, the criteria for platforms were that these are popular among youth. publicly available data by the platforms themselves and additional research on the user bases of youtube, twitch and tiktok shows that these are popular platforms among youth who are between 13 and 25 years old (ceci, 2022a, 2022b; hoekstra et al., 2022; twitch, 2021). 2.1.2 community selection to have explicit and visible interactions of how members perceive to be learning in these communities, we selected communities in which learning, though in different degrees, is made explicit as (a part of) the aim of the community. based on theories on communities of practice, we consider these to be learning communities when they state within self-descriptions on their platform or in comments/chat to convene around a shared interest, problem, or practice to jointly expand on their knowledge and skills by engaging with one another. to use gee’s terminology: part of their grammar is a joint enterprise of expanding the knowledge and competences of the people engaging with the shared affinity. based on these criteria, we selected an e-commerce community and a lgbtqi+ vlogging community on youtube, an info-security and a speedrunning community on twitch, and a history and a sustainability community on tiktok. one might wonder how we can compare these selected communities that are so different to understand online informal learning communities. we would argue exactly their difference makes it a valuable selection of communities. what they have in common is that they are learning communities in the sense that they are comprised of people who convene and share knowledge and skills around a shared affinity. furthermore, though this does not constitute a full (if that is even possible) representation of the platforms, we have included two communities per platform to see if we can already identify some commonalities in how online informal learning communities are shaped on the same platform, though further research would be needed to be able to generalise such results. the communities are purposefully selected to differ in their affinities as we are interested in comparing how youth’s learning communities are differently shaped online by looking at how they are learning, not at what they are learning. including different communities centred around different affinities allows us to analyse what is common in how these communities learn in online spaces, even if their affinities differ radically, such as in how they are generally considered to be valuable skills or knowledge for youth’s formal education. as such, we hope to see how these different informal learning communities learn in a platformising online space, where the learning communities are similar in their joint enterprise of learning about a shared interest outside the formal school context, informally, yet differ in what that interest is. 2.1.3 creator selection within each community, we took a sample of creators based on a match with the community’s main interest and whether the creators were embedded and recognisable in the wider community, by for instance having a collaborative video with another creator of the same community. we selected on youtube and twitch two to three creators per community. on tiktok we selected creators who were members of an interest-related ‘house’, which are collaborative accounts of creators on tiktok, e.g., ‘history house’. we cannot and do not want to claim representativeness, as we only looked at these communities via one platform and via selected creators, even though these communities extend beyond these platforms and creators. yet these creators do give a general idea of the community as they are selected based also partly on whether they would be recognisable within the wider community. 2.2 procedure to observe these communities, we drew from traditions of data collection from digital ethnographies and learning ethnographies. we looked at the community’s grammar by observing vermeire, de haan, sefton-green, & akkerman 7 | f l r interactions, understanding content (videos/livestreams) and responses (comments/chat) as the vehicles for interaction with attention for how such interactions are shaped in interaction with the platform environment. on youtube, we focused on the interactions on the watch page, on twitch on the livestream page, and on tiktok on the for you page. we position our data collection within a wider tradition of digital ethnography (hjorth et al., 2017; pink et al., 2015) and more specifically borrow from ethnographies of learning (azevedo, 2013; paradise & de haan, 2009), in an attempt to provide a rich description of the social structures for learning of these affinity spaces, as situated in their digital environment, although we do not claim this is a ‘full’ ethnography. such an approach also aids in describing how portals, platforms as the digital environment, generate social (learning) interactions, the ‘grammar’ of the affinity space. to capture learning interactions, we drew from learning ethnographies method of observation to capture learning interactions, such as observing moments in which experts and novices interact to understand how knowledge is ‘taught’ (paradise & de haan, 2009). in addition, for our observations we wanted to collect data about specifically learning in digital contexts for which we drew from digital ethnographies (hine, 2000; pink et al., 2015) and learning ethnographies stressing the importance of attention for the material environment for social behaviour (hasse, 2014). more concretely, this entailed that we also observed how the platform, as a portal, constituted affinity spaces’ grammar, and more specifically, social structures for learning. to obtain a rich description of the platform as a portal generating the grammar of the community, we looked at the interaction between the community members expressions about the platform’s role in their learning activities and at the platform itself, and how its infrastructure and a history of usage and tradition function as a social structure for learning. in other words, we understand the platform environment not simply as a technological backdrop for learning but as an environment that carries structures and norms about behaviours and practices with it, and as a portal generates the affinity spaces’ grammar. in some ways, youtube, twitch and tiktok can also be understood as an affinity space. that is, affinities might exist at the level of a learning community about e-commerce on youtube, or learning community about lgbtqi+ vlogging on youtube, but also at the level of the platform youtube. people who frequently engage with youtube or tiktok might be able to connect with one another over shared experiences that are tied to the platform. we refer to this as the ‘platform culture’ to distinguish it from the affinities that exist at the level of ‘online informal learning communities’. as we also acknowledge that these two levels are existing in entanglement with one another, we also bring these levels together to see how these levels together shape the identities, social relationships, and interactions of the affinity spaces in ways that might perhaps be platforms specific. to observe the communities, we used an observation schedule to capture the grammar of the community, which, as indicated previously, is a term used by gee to describe the interactions, values, thoughts, and practices constituting relations between people. by choosing for observations, we, however, have to limit ourselves to what is visible on these platforms: comments/chats and videos/streams. our observations thus captured interactions foremostly, in which we focused on three aspects of those interactions that gee mentions as characteristic of ‘affinity spaces’: the boundaries, hierarchies and interactions on how commentors/chatters and creators relate to the affinity. for this research, we had particular attention for learning interactions. we understood the grammar to be about learning when it could be understood as a transformation of an individual or community in relation to knowledge and/or competences about the shared affinity. an example of such a ‘learning interaction’ could be this asynchronous, fictious comment thread (inspired by actual comments in one of the observed youtube communities) in which commentors c1 and c2 aid each other in understanding a term by exchanging information related to their shared affinity of lgbtqi+ vlogging: c1 17-11-2022 09:33: ‘sorry new here, but what does cisgender mean?’ c2 17-11-2022 12:30: ‘no worries! great you ask, cisgender is when you identify with the gender you were identified with at birth.’ c1 18-11-2022 10:30: ‘thank you for explaining, that makes sense!’ vermeire, de haan, sefton-green, & akkerman 8 | f l r we screen recorded these interactions to keep a consistent data set of interactions. we also recorded the number of views, likes, shares, and comments if such numbers were afforded by the platforms, to monitor how certain interactions were received within the community. we took two weeks for observation on each platform as a baseline to start from, to see if we could obtain an understanding of these communities’ ‘grammar’ on the platform. though we used the same observation schedule for each platform, the way in which the platform worked, made observations per platform different. as both youtube and tiktok work asynchronously, it is not possible to observe live interaction in the way that is possible on twitch. however, we can still observe interactions on these platforms, even though members do not immediately respond to one another. we gathered such data especially around the time of a video being posted by a creator on the platform, as we assumed that commentors and creators interact with one another mostly at those moments. we observed on youtube the text in the ‘about’ page and videos and comments that were posted during a two-week observation period by the selected creators. in addition, we observed an additional five popular videos and related comments of those selected creators, as often only one or two videos were posted in those two weeks. on tiktok, we observed the home pages of all included channels and videos and comments that were posted on those channels during a two-week observation period. on twitch we observed the livestreams, including video and chat, of the channels. for most channels we observed livestreams taking place in also a two-week observation period, though as the included streamers streamed sometimes for four hours four times a week or sometimes streamed at the same time, we had to pick which streams to observe. in such cases, we asked streamers which of their streams would be most exemplary of their everyday practice and observed those. one twitch channel was an exception as it was an event-based channel that only had one weekend long event every couple of months. for that channel, we observed an event that took place during the days from friday morning till sunday evening. we asked community leaders which of the parts of that weekend long stream were most exemplary for the identity of the community. in total, we observed on youtube 22 videos and on tiktok 70 videos. we observed approximately 48 hours of livestreams on twitch. 2.2.1 ethics and privacy for this research, we have permission from the faculty of social and behavioural sciences ethics review board of utrecht university (20-0553) and worked in line with their requirements. hyperlinks linking to online content in this paper are either from creators who provided permission for sharing their content and/or from content that is viewed and liked by so many users that we can assume creators and chatters/commentors consider this public content (see also the nesh guidelines for internet research). on twitch we asked the selected channels for permission to join certain streams as the live nature and intimacy of the streams might make it appear as a more private space than people who interact with a youtube or tiktok video. streamers implemented a bot, meaning an automated command activated in chat, like ‘!study’, that would provide information about the research if the command was typed in chat, which moderators and the streamer would do to inform that research was taking place, apart from also announcing it in their discord and title of the stream. as such, people participating in these streams were made aware that there was a researcher observing them, who also had a recognisable username ‘researcher_zowiez0’ and a profile page with more information about the research. they were also given the opportunity to have their chat messages excluded from the analysis by filling in a form shared in chat by the automated command. during our observations on twitch, we interfered minimally, only asking for clarification when certain interactions were unclear due to for instance the usage of abbreviations. on youtube and tiktok, apart from members who were interviewed, users did likely not know observations were taking place as we left very few traces, like comments or videos, to make users aware of our presence, though channel owners were informed observation would take place on their channel. for anonymisation purposes, we have paraphrased english comments and translated dutch (the native language of three authors) youtube and tiktok comments to stick to their original wording as closely as possible while also avoiding that googling a comment might result in identification of a specific commentator. if interactions are translated from dutch to english this is indicated by ‘(t)’. https://www.forskningsetikk.no/en/guidelines/social-sciences-humanities-law-and-theology/a-guide-to-internet-research-ethics/ vermeire, de haan, sefton-green, & akkerman 9 | f l r 2.3 analysis to answer our research question, we did a critical discourse analysis informed by gee’s concepts ‘grammar’ and ‘portal’. more specifically, we looked at how the ‘grammar’ of the selected communities, as affinity spaces, is or is not informed by the platforms on which they operate, and if so, how. we do so understanding that platforms create particular ‘regimes of truth’ (foucault, 1995) through their algorithmic power (bucher, 2018). critical discourse analysis aids in uncovering such underlying truths in the everyday interactions of young people in the selected communities on such platforms, with attention for how those platforms create particular discourses. to focus our critical discourse analysis, we draw on gee’s concepts of ‘portals’ and ‘grammar’ to analyse platforms’ power over regimes of truth and to uncover how community interactions and norms, their grammar, could be informed by such regimes of truth. ‘grammar’ is then a regime of truth that consists of the interactions, values, thoughts, and practices that constitute the relationships between people to create a community around a shared affinity (2005). using a critical discourse analysis informed by these terms aids in recognising patterns in how platforms and affinity spaces create a cultural ‘truth’, particular to these communities on how learning interactions are shaped. we use the concept of ‘portals’ to explain our analytical perspective on platforms as also shaping the social structures, the regimes of truth, for learning. we draw on gee’s argument that ‘portals offer access’ (p. 220), that is provide the socio-material affordances, to an affinity space, and that ‘portals’ are strong generators of the grammar of an affinity space (2005). in other words, a portal is an access point to an affinity space which simultaneously also exercises power over regimes of truth on that platform: over how interactions around that affinity are taking place (2005). for example, portals to the affinity space of the role-playing game dungeons and dragons (d&d) are the player’s handbook that explains the rules of the game, or an online forum about the game, or a d&d game night in a game café. all these portals, gee would argue, provide a means to access the affinity, and through providing that access also generate the grammar of this affinity space as a café with live communication might differently shape the grammar around d&d than an asynchronous forum for instance (2005). in this paper, we analyse how the platforms youtube, tiktok, and twitch as portals offer access to learning communities and simultaneously generate the grammar of that learning community to understand how learning communities are shaped within this particular type of portal. using this concept of the ‘portal’ as a lens allows us to look at how platforms’ regimes of truth shape the identity of the affinity and the social relations and interactions in relation to the socio-material infrastructure of the platform, and its power over interactions. in our analysis, we focused mostly on the grammar of the affinity space when we could understand it as ‘learning’ interactions. as we are also interested in ways to learn that are not generally acknowledged as learning within schools, we took a broad conceptualisation of what such ‘learning’ interactions could be. in line with our introduction, we understand ‘learning’ to be those interactions in which the community or a member indicated transformation, meaningful movement, towards a particular purpose (akkerman et al., 2021) related to that affinity in terms of skills and/or knowledge. because we were interested in analysing ‘online informal learning communities’ as affinity spaces existing on platforms, we were interested in those learning interactions that are related to the affinity, so the particular purpose of the transformation had to be part of acquiring an understanding of the shared affinity. in addition, we wanted to analyse how the platform generates the grammar of these communities. we did this by on the one hand looking at how users reflected on how the platform facilitated access to their affinity in interactions. on the other hand, we looked at how such interactions were situated in the infrastructure of the platform as a portal that also shapes how the grammar, meaning the interactions themselves, were generated. in sum, we did a critical discourse analysis of online informal learning communities as affinity spaces to analyse recurring patterns in: how learning interactions were happening within the selected communities how platforms were reflected upon in interactions as shaping communities’ learning interactions vermeire, de haan, sefton-green, & akkerman 10 | f l r how platforms as portals generated specific learning interactions through their infrastructure 3. results below we describe three themes that exemplify recurring patterns in our data on how according to our observations, in interactions platforms are described to generate the grammar of affinity spaces. we will share these themes based on our analysis of the interactions within the six affinity spaces, which we will firstly introduce below, after which we will present the themes. 3.1 community introductions before delving into the results per platform, we introduce the communities by describing their affinity and how they relate to that affinity on the observed platform. e-commerce the e-commerce community centres around an affinity for e-commerce knowledge and skills. members relate to their affinity on youtube by sharing personal stories of how to become a successful online entrepreneur, showing off luxury lifestyles, and providing tutorials and information. lgbtqi+ vlogging the lgbtqi+ vlogging community centres around everyday life experiences of lgbtqi+ people. members relate to their affinity on youtube by vlogging or commenting directly about being lgbtqi+ and the life experiences that come with that, or by other types of more common kinds of vlog content on youtube, about for instance shopping, make-up, fashion, and other everyday activities. infosec the information security (short: infosec) community shared affinity is about (developing tools to) test and strengthen online security of information. they relate to their affinity on twitch through meeting in livestreams via chat and video in which streamers show relevant skills and information to infosec, and programming more widely. speedrun speedrunning is about the shared affinity for completing games as quickly as possible. the speedrun community relates to their affinity on twitch through either streams in which speedrunners hone their skill or by watching competitions and marathons in which speedrunners show their skill. history the history community has a shared affinity for learning about history and sharing historical knowledge. members relate to their affinity on tiktok by making, watching, and commenting on short videos with historical information. sustainability the sustainability community has the shared affinity for advocating for sustainable behaviour, policy, and practices. they relate to their affinity on tiktok by making, watching, and commenting on videos with sustainability information and climate activism. 3.2 platform’s visibility regimes generate boundaries for affinity spaces our results showed that an important way in which platforms as portals are by communities described to generate the grammar of affinity spaces, is through what we call ‘visibility regimes’: the ways in which platform algorithms are by platform users described as governing the visibility of their interactions and the boundaries of their affinity space. to describe how affinity spaces appropriate and https://www.youtube.com/@joshuakaats/featured https://www.youtube.com/@jessiemaya https://www.twitch.tv/d0nutptr https://www.twitch.tv/j3nnalive/about https://www.tiktok.com/@thebullmoose5845 https://www.tiktok.com/@ecotokcollective vermeire, de haan, sefton-green, & akkerman 11 | f l r resist such regimes, we zoom in on one particularly illustrative example from the lgbtqi+ vlogging community, a video from the channel ‘jessie maya’. jessie is a popular dutch youtuber who vlogs about topics ranging from fashion, food, makeup to anti-gay commercials and being a transwoman. in her video ‘reacting to transgender updates after 6 years…’ (t), which at the moment of observation had received 234,137 views and 1057 comments, jessie reflects on her previous ‘transgender updates’ video series. jessie’s introductory text of the video is telling of the community’s experience with youtube (t): transgender will be in the title, and often youtube intercepts that or it doesn’t get into the algorithm […] this video will probably be demonetised as well […] jessie describes here that youtube does not offer the same visibility and rewards (monetisation) for her content that explicitly deals with her trans identity such as this video, as other content she has made. she describes to experience youtube as controlling whether her content is seen by others based on its topic, which could play a part in her affinity space’s potential popularity, growth or decline on the platform. youtube here is described as controlling the visibility of her community. as was evident from our data, such visibility regimes might also play a role in the learning trajectories of members of an affinity space. we often found that comments by members described that they found the lgbtqi+ community by accident. an example of such a comment, underneath one of alice olsthoorn’s videos (t): i once started following you when you decorated pumps with glitter but these videos [on calling out transphobic comments] are the ones i stayed for, the way you can put people in their place in a peaceful, civilised manner. this might indicate that the content that is related to the core affinity of the lgbtqi+ vlogging community might not be as visible as content that is further removed from their shared affinity. in other words, based on community members’ comments and expressions in videos, they seem to experience that the boundaries, access points and potential growth of their community are managed (and thus limited in this case) by the platform’s visibility regimes. however, the lgbtqi+ vlogging community does not simply abide by youtube’s visibility regimes and the boundaries it imposes on their affinity. they also attempt to resist and appropriate such boundaries. as we will demonstrate below, the lgbtqi+ community in response to youtube’s ‘visibility regimes’, employs activities that they perceive to generate the visibility on the platform that they want in an attempt to resist these regimes. after mentioning the ways in which youtube thwarts her videos, jessie shares the following (t): ‘the last time you all succeeded, because you did a thumbs up, because you spammed comments.’ a quarter of all comments answer her call, such as this comment with 910 likes: ‘okay let’s do this again, put this shit back into the algorithm!’ or this one (t): ‘yt is sooo lame, your content is so important, not just for trans babies who watch you but also to create more awareness among the cis people.’ jessie and her commentors’ interactions indicate an attempt to ‘curate’ youtube’s algorithm to obtain visibility by posting comments, which they mention to believe aids in gaining visibility on the platform. in sum, the latter results show that users’ activities on youtube indicate that youtube might play a part in how visible and potentially popular some affinity spaces can become on their platform. such activities are also clearly seen on tiktok. our data shows how in particular underneath activist videos from the sustainability community, commentors attempt to engage in similar forms of ‘curation’ by commenting ‘boost’ (trying to raise attention) or ‘comment for the algo’ to ‘curate’ the tiktok algorithm into letting their message be more visible and expand the boundaries of their affinity space. moreover, another common comment from our observations of tiktok is: ‘comment to stay on this side of tiktok’. such a comment again indicates a sense of ‘algorithmic curation’ but this time a member appropriates it to see more content produced by their community. sometimes this even works the other way around, where commentors want to let the creator know whether the video is on the ‘right side’ of tiktok’s algorithm: ‘target audience reached’. this comment expresses to the creator that they have vermeire, de haan, sefton-green, & akkerman 12 | f l r reached the ‘right side of tiktok’: the community of people they want to address. such expressions in comments of appropriating and resisting the algorithm to support own and other’s access to the community, reflects a sense of control over the algorithm and over the visibility of their interactions, and the boundary of their community. simultaneously, it shows the potential decentring of their affinity; and how they try to maintain access and interaction with an affinity space through playing with what might help to ‘curate’ the algorithm. portals, by definition, generate growth through access, and create boundaries around affinity spaces. similarly, platforms (as portals) are recurringly reflected upon by communities to introduce manipulations that can be related to their visibility regimes, which, as described in our introduction, are rooted in a complex combination of commercial, normative and surveillant aims that play a part in both content visibility and access, generating an affinity space. within this platformised social media space, members respond to perceived algorithmic control through ‘curation’ to ‘fight’ for the existence and visibility of their community in their interactions. 3.3 platform cultures generate competitive hybridisation as our data shows, a second way in which platforms generate the grammar of affinity spaces, is through specific platform cultures. such platform cultures give rise to what we will refer to as ‘hybridised affinities’ that draw on distinct affinity spaces by combining affinities from both those spaces in created videos. by hybridisation we explicitly refer to how affinities are hybridised. such hybridisation of affinities can be happening by using the different modalities a platform affords users to work with (text, sound, gesture, moving imagery) to refer to different affinities, but here we centre not on this combining of modalities but on the combining of different affinities and how this affects how learning communities operate in these platform spaces. we will argue, by discussing two examples that are exemplary of how this theme is recurring in our data, that ‘hybridising affinities’ can be understood both as a conformity with platform cultures as well as a resistance of such cultures. the example video and interactions that we discuss below come from a video by ecotok, a house of sustainability creators. the video had 39.8k likes, 436 comments and 2598 shares at the moment of observation. the video goes as follows: we see a person standing on a canoe pushing themselves forward through a rice field. a textual overlay states: ‘we can’t drink fossil fuels.’ when it disappears a new overlay state, while the camera turns to overlook the rice fields 180 degrees away from the person on the boat: ‘so why do we prioritize them over water?’ a new overlay then states, while the camera returns to the person on the boat: ‘help us stop the line 3 pipeline in minnesota.’ the camera than moves away from the person in the other direction, showing first the text: ‘link in our bio.’ and then ‘[adult swim]’, using the actual squared brackets as the logo has these. the vano 3000 – vano 300 sound accompanies this video, which is a sound that was originally used by adult swim, a television network, to make short videos with their logo. this video employs a, for a tiktok-user recognisable, trend of mimicking with video, text, and sound the adult swim videos. this time, however, it is not an ‘ad’ for adult swim but a remix of its aesthetics to deliver a different message: they call attention to a petition against an oil pipeline. this video adopts and ‘borrows’ an aspect of a different affinity space: the wider affinity of tiktok culture in which it is common to participate in particular video trends, as this sound and aesthetics that they mimic is originally not used for sustainability content, but for a tiktok trend. other examples of ‘borrowing’ from other affinity spaces are for example the creator nosebled, whose bio says: ‘that one dancing history chick’. this creator makes videos in which she frequently combines her interest in dancing and in history by doing a dance while using textual overlays that share historical events. some of the comments express appreciation for how nosebled combines dancing with history: ‘i’m a fan of learning facts while a person dances #innovativeeducation’. and: ‘i’m just realizing how absurd it is that tiktok used to be cringe cause it was just people dancing.’ by dancing while transferring historical knowledge, nosebled creates a hybridised affinity that speaks both to the (larger) dance affinity space on tiktok, to which the latter comment also refers, as well as to history afficionados. tiktok as a portal for the history and vermeire, de haan, sefton-green, & akkerman 13 | f l r sustainability communities generates affinity spaces in which affinities can hybridise and overlap as indicated by these hybridised affinity interactions. based on these hybridised affinity interactions of the selected communities, platform cultures could be understood as potentially enabling and encouraging the hybridisation of affinities within affinity spaces. on the one hand this can be interpreted as platform cultures encouraging a ‘dilution’ of one’s affinity as part of a popularity contest for views: speaking to two affinity spaces might help creators to generate more attention on the platform. on the other hand, the hybridising of affinities can also be understood as a creative resistance against platform cultures that might make space for one affinity, but not the other. by hybridising their affinity, an affinity space can then aim to enforce space on the platform for their affinity. we can also see this on youtube, where the trans creators draw on larger platform trends, belonging to youtube’s platform culture, such as so-called ‘reaction videos’ to still talk about their own affinity, yet within a way that speaks to the larger platform culture of ‘acceptable’ content. another example is how in the e-commerce community conservative gender ideologies are sometimes wrapped up in inspirational and motivational videos. there are also examples of more direct resistance to how the platform culture, particularly the reward system, works by turning away from it, or expressing clear mocking of a larger platform culture. for instance, during our observations of infosec streamer ash_f0x, they had a subscriber goal on top of their stream for doing a ‘hot tub stream’. ash_f0x explained that this was a joke on stream, in response to so many hot tub streams suddenly appearing on twitch. however, after realising that some of their viewers might not interpret it as a joke, and they did not want to be seen as encouraging people to subscribe to them, they put the subgoal down. twitch’s reward system and how other streamers use hybridised affinities to generate more views and subscribers on twitch, was appropriated and ridiculed by this streamer, clearly distancing themselves from such practices. in summary, hybridised affinity content indicates that platforms might generate specific interactions based on their larger platform culture, which affinity spaces use to create their content, hybridising it with other affinity spaces, to potentially gain more visibility on the platform, such as by using a trend. this also indicates that each platform that an affinity space uses to reach their aims, could differently generate the interactions related to that affinity as befitting the platform culture and the affinity. 3.4 platform hierarchies generate an hierarchisation of expertise furthermore, our observations of interactions in the context of the platform infrastructure showed that platforms intervene in how relations between ‘old timers’ or experts, and ‘newbies’ or novices take shape. in the infrastructure that they offer for interaction, platforms push towards a hierarchisation of expertise relationships in which creators are foregrounded and other users, such as viewers and commentors, are pushed to the background. we will illustrate this by discussing comments from tiktok and youtube communities that are indicative of the idolisation of creators that is widely present in our data. subsequently, we will also discuss how the ethical hacking community negotiates such practices. first, it is key to realise how creators have a focal presence on the social media platforms included in this study when we analyse platforms’ infrastructures designs. youtube, twitch and tiktok are designed so that videos and streams are placed central on the page where the user watches these videos. comments and chats are initially hidden or happening to the side of the video or stream. as such, the focus lies on streams and videos, that are created by creators. furthermore, platforms reward mostly creators: only on twitch can users who do not create video material obtain rewards such as channel points for viewing or badges for participation in a community. only creators can gain a partnership with youtube as symbolised by a small red symbol next to one’s channel name. only creators can gain a ‘verified channel’ on tiktok as symbolised by a small blue checkmark next to one’s channel’s name. only creators on twitch can become twitch affiliates or partners, as also symbolised by a specific badge. when observing these platforms, it quickly became clear who matters the most from the perspective of the platform. in many of the observed comments users also paid respect, vermeire, de haan, sefton-green, & akkerman 14 | f l r thankfulness, admiration towards creators, which reflects a relationship between leaders and followers, resonating elements of fandom culture rather than the more egalitarian expert-novice relationships that gee described for affinity spaces. for example, underneath the video by jessie maya mentioned in the first theme: ‘i have learnt so much about transgenders and the whole community because of you’ (t), and another underneath the history video by nosebled: ‘you’re my fave creator everr, i legit always learn new stuff and the dance talent!!’. we concluded that within these affinity spaces, based on the design and observed interactions, the platform portal, including the platform culture, affords a hierarchy in which creators are generally positioned as the educators and experts, whereas commentators are there to learn from them. the interactions of communities on twitch indicate a negotiating of such hierarchisation though. we take an illustrative example from the infosec community: a segment from an exchange within a stream by d0nut titled ‘resync ‘n chill (working on http client -!study) – rust’. this stream had around 34 viewers at the observed moment, and it took place approximately one hour and 50 minutes into the stream. the viewer sees d0nut’s screen and activities thereon and in a corner a live webcam video of d0nut. next to this screen, viewers can interact with one another and d0nut via a chat box. in this moment, d0nut is trying to solve an issue they are working on with chat while another conversation unfolds on how to begin with programming. we have presented this conversation below in a way that makes it easier to follow, but this is not how it happens on stream, where chat and streamer ‘talk’ simultaneously. people in chat are represented by c and a number instead of their username, ‘@’ is a way to say one is responding to a specific person in chat. c1: ‘i actually want to learn to program, but everyone says it is very difficult’ […] c3: ‘@[c1] getting started is easy. eventually, you get to choose where you want to go. some routes are easier. other routes are harder.’ c6: ‘@[c1] programming is more fun than difficult. it is hard from time to time, but like with any skill, it gets easier and easier soon after.’ c7: ‘what about masscan’ c7: ‘are there full network stacks written in rust?’ d0nut: reads out above chat by c3 ‘absolutely agree, it can totally be easy. […], but uh [following animation pops up of c3 following]. oh [c3] thank you for the follow. yeah, but you hm it totally be easy, uh, uhm. i just say, start the right way. don’t learn programming to learn programming, have a project, have a task, have a thing that you need done and programming is the way you get to that result. that way you are not even worry about it, you are actually just looking things up so you can get it out of your way and get your task or goal accomplished, […]uhm dot dot dot, reads aloud above chat by c6 yup, yeah, it is. […] [c6] is absolutely right after this moment d0nut replies to c7s suggestions by showing examples on their screen. c3 asks for help with a project they want to start, like one d0nut recommends doing to start with programming, d0nut and chat all help c3 with suggestions and advice on how to achieve this project. this example shows that though the streamer takes a central position in guiding the conversation, sometimes chat and sometimes the streamer has the position of information provider. everyone helps one another, regardless of experience. in this example, we can also see how d0nut seamlessly appropriates functionalities of twitch and their usage by the community into the conversation with chat: ‘oh [c3] thank you for the follow.’ this reading out of chat messages by d0nut, the streamer, is a form of recognition for, in this case, c3’s contribution of offering advice to a fellow member of the community. by mentioning the advice of people in chat, d0nut uses their hierarchical position to amplify the voices in chat. the person in chat obtains the opportunity to then also express their appreciation for being included in the stream, by ‘following’ (an interactional act). in these ways, the monetisation rewards that come with reaching a certain number of followers or subscribers on twitch are by these communities used to not create hierarchies but aid their collaborative practices. we have seen vermeire, de haan, sefton-green, & akkerman 15 | f l r recurringly throughout our observed communities, particularly on twitch, how communities attempt to subvert this asymmetry a platform imposes between creators and chatters/commentors by using this asymmetry to invite expertise from the community and by employing monetisation from the platform to benefit not (only) themselves but also their community. d0nut for instance emphasizes the purpose of community for donations in their ‘about’ section: donations help me give back to the community another example is how the benelux speedrunners gathering (bsg) marathon channel uses donations they receive during their marathons to support charities, in their bio: ‘bsg’s donations support mind, 100% of your donations will go straight to them.’ during speedrun streams such collaboration and exchange between distinct roles of expertise is also present, for instance during a minecraft run, streamer buggy expresses that they do not know what a particular building is, and chat tries to help her. at another time buggy explains a particular speedrun trick, taking the position of expert. these examples show how the speedrun and infosec community encourages everyone to collaborate regardless of their experience. to sum up in the words of the bio of the info-security streamer ash_f0x: ‘my personal goal is to learn something new every stream, together with my viewers. […] nobody will be judged here based on a “stupid” question, so…just ask!’ in summary, we have seen that the platform infrastructures are geared towards a hierarchisation between members of affinity spaces, generally positioning creators as the ‘experts.’ however, community interactions indicate that affinity spaces can resist such hierarchisations and appropriate and resist such structures to create more fluid relations of expertise, to match with collaborative cultures, in line with those that have been claimed by gee as typical for affinity spaces. 4. discussion based on these results, in this discussion we will first describe the three ways in which we need to re-evaluate affinity spaces to capture the new platformising context of online informal learning communities, putting forward the term ‘platformised affinity spaces’ as a first step towards a needed new conceptualisation of what an online informal learning community is in times of a platformised web. we then discuss how this raises fundamental questions for media studies and the learning sciences about how the socio-technical context of (online) informal learning communities has changed. first, our results show how we need to re-evaluate how online informal learning communities have been previously conceptualised to respond to the changed context in which they operate. we particularly continue the discussion on learning communities by looking at gee’s development of the conceptualisation of online informal learning communities as ‘affinity spaces’ and how those affinity spaces would be characterised by more fluid boundaries. more specifically, we have demonstrated how platforms, driven by their commercial and surveillance goals, are in comments and videos often reflected on as taking (partial) control over the visibility of, and access to content of online informal learning communities, which has implications for how we can conceptualise online informal learning communities’ boundaries in the context of platforms. our results can be interpreted to reintroduce a discussion about this fluidity of boundaries of learning communities in this platform context. talking again about the boundaries of learning communities might be seen as surprising, especially in the online context, as gee introduced the concept of ‘affinity space’ precisely to address the problems associated with earlier conceptualisations of ‘learning community’ that required a demarcation of boundaries. gee’s work was in line with studies arguing that the boundaries around access to resources and experts, and who has and who does not have knowledge and expertise, would become more permeable due to the networked affordances of online spaces (rainie & wellman, 2012; ünlüsoy et al., 2013). however, our results show that the introduction of platforms as portals to access learning communities creates the need for a re-evaluation of how boundaries play a part in the conceptualisation of online informal learning communities. particularly as these boundaries introduced by platforms might control whether vermeire, de haan, sefton-green, & akkerman 16 | f l r and how young people are able to access (their) learning communities in a way that could be informed by the commercial goals of platforms, rather than young people’s own learning desires. in summary, a new dynamic that needs to be considered when reformulating affinity spaces is how affinity spaces are characterised by how platforms as portals generate and take (partial) control over their boundaries by governing the visibility, access and reach of the affinity space. second, platforms as portals to informal learning communities generate a competitive hybridisation of the community’s affinity with platform cultures that challenge the grammar of these affinity spaces. whereas in work from the learning sciences on informal learning communities the grammar was seen as ‘given’ by the affinities of the learners, in a platformised context affinities have to actively engage with platform cultures, economies and dynamics to continue to exist on these platforms, giving rise, for example, to the examples of hybridised content mentioned above. in other words, a new dynamic that needs to be taken into account when reconsidering the idea of affinity spaces is how the grammar of affinity spaces in platformised contexts is co-defined by engagement with the platform culture and visibility regimes that affect how young people control their own ways of connecting to the affinities they might wish to learn from. third, platform hierarchies generate a new hierarchisation of expertise. our results show that the relationships between novices and experts within learning communities on platforms are not as fluid as gee proposed for affinity spaces. platforms, through their reward system and infrastructure, introduce relatively fixed hierarchies between creators as experts and chatters/commenters as novices, that (although not all communities adopt those hierarchies), position creators as experts on these platforms. although gee perceived expertise, in line with literature on networked and informal online learning, as dispersed in the network and distributed among community members (lankshear & knobel, 2007; säljö, 2010), it can be argued that platforms control to some extend the positionality of novices and experts in their design which generate relatively fixed ‘expert’ roles. in other words, a new dynamic the platformising online context introduces to the concept of ‘affinity space’ is how it is shaped by platforms generating creators as experts. given these latter new dynamics that platforms introduce, we introduce the term ‘platformised affinity spaces’ to allow for the analysis of online informal learning communities with particular attention to how platform infrastructures and cultures shape their boundaries, grammar, and hierarchies. we offer this as a first, not exhaustive, step towards conceptualising informal learning communities operating in todays’ platformised environment, in which we mostly can share insights related to how platforms as portals might consistently play a similar role in how young people’s learning activities are taking place within affinity spaces. we use ‘affinity space’ here rather than online informal learning communities as these communities still share characteristics with gee’s ‘affinity spaces’, such as convening around a shared interest and bonds between members being determined by having a shared interest with a shared grammar. yet, ‘platformised affinity spaces’ is a further needed specification of gee’s concept that aids in a wider understanding of how online informal learning communities operate by adding three core characteristics. young people’s activities in these communities on platforms indicate that: 1. …platforms’ visibility regimes take control of access and create boundaries around online informal learning communities, while their activities also indicate a perceived control over platforms’ algorithmic sorting of content 2. …platform cultures invite hybridisation of interests to learn about them, which can be understood as both a dilution of niche interests and learning about them, as well as a way for eager learners to find a place to learn about niche interests 3. …platforms generate hierarchies for learning rooted in their positioning of creators that succeed in drawing an audience, which entails that the more fluid relation between novices and experts that gee (2005) describes becomes more rigid on platforms, generally positioning popular creators as ‘experts’ for learning vermeire, de haan, sefton-green, & akkerman 17 | f l r we now want to briefly broaden the focus from gee’s affinity space and use the insights from this study to critically examine the assumptions about online informal learning communities of both media studies’ knowledge on platforms and learning sciences’ knowledge about online informal learning communities. as introduced earlier, media studies state that platforms manipulate behaviour and take control over learning to such an extent that we can worry about whether it takes away freedom of thought (alegre, 2021), critical thinking (sefton-green & pangrazio, 2022), and agency (koopman, 2019), skills that scholars argue are key to educating youth to become critical citizens in democratic societies and to participate in communities (sefton-green & pangrazio, 2022). however, our results show that young people’s online activities demonstrate the perceived ability to appropriate and resist the dynamics that platforms generate within their affinity spaces and push back against platform control over their learning. therefore, to argue that they have no agency in their online informal learning communities could be seen as reductive of youth’s online activities. youth’s activities in these communities demonstrate a power to appropriate and resist these platforms in order to create space for the ways they want to organise their learning within their communities, not only having an imagined idea of how the algorithm works, as described by bucher (2017), but also actively curating these algorithms to achieve the goals of their online informal learning communities. by discussing youth as data subjects manipulated and controlled by platforms, they are rendered as passive objects, or as zuboff describes, ‘addicted users’. at the same time, we can be critical of the lack of attention to platform control in the literature on online learning communities. for example, cousin (2005) and ünlüsoy et al. (2021) showed how online spaces afforded access to seemingly infinite connections to resources and experts, connections that sometimes appeared to be accidental. in our results, we also see that access to learning communities on these platforms is in comments and chat sometimes described as random or accidental. if we interpret this ‘accidence’ through the lens of literature on platforms, such coincidence or accidence could be the result of manipulation by the platform which predicts for a young person what they should learn based on the data collected about their previous interactions (alegre, 2021; van dijck et al., 2018; koopman, 2019; sefton-green & pangrazio, 2022). such a form of control by platforms is not necessarily only ‘good’ or ‘bad’, as, for example, young people may be ‘manipulated’ by the platform to deepen their understanding of history through the provision of increasingly in-depth history videos or manipulated to engage with increasingly toxic and radicalised online communities. such manipulation raises concerns about the power of platforms to determine which learning communities are (easily) offered to young people for learning, and which are not. overall, based on our results, we would like media studies and learning research to move forward with nuance regarding the role and control of platforms when researching online communities. it is valuable for these fields to understand that online learning communities are in part controlled by the dynamics that social media platforms introduce, while at the same time these learning communities’ activities reflect a sense of control by the perceived ability to resist and push back against such platform dynamics to achieve their own learning goals. keypoints we introduce the term ‘platformised affinity space’ to describe the ‘new’ dynamics social media platforms introduce for learning communities. we argue that, first, social media platforms re-introduce a discussion about the boundaries of learning communities through their visibility regimes. secondly, platform cultures introduce a practice of hybridising communities’ core interests with other interests that are popular on the platform. thirdly, a more fixated hierarchisation, informed by platforms’ focus on creators, is introduced into learning communities’ social structures. vermeire, de haan, sefton-green, & akkerman 18 | f l r funding this research was conducted in the context of the research project young people’s learning in digital worlds: the alienation and reimagining of education which is supported by het haagsche genootschap. references akkerman, s. f., bakker, a., & penuel, w. r. (2021). relevance of educational research: an ontological conceptualization. educational researcher, 50(6), 416–424. https://doi.org/10.3102/0013189x211028239 akkerman, s., & leijen, ä. (2010). challenges of today: maniness, multivoicedness, and dispersion of knowledge. qwerty, 5(1), 29–43. alegre, s. (2021). protecting freedom of thought in the digital age. centre for international governance innovation. https://www.jstor.org/stable/resrep33463 angouri, j. (2015). online communities and communities of practice. in the routledge handbook of language and digital communication. routledge. azevedo, f. s. (2012). the tailored practice of hobbies and its implication for the design of interest driven learning environments. journal of the learning sciences, 22(3), 462–510. https://doi.org/10.1080/10508406.2012.730082 boyd, danah. (2014). it’s complicated: the social lives of networked teens. yale university press. http://www.jstor.org/stable/j.ctt5vm5gk bronkhorst, l. h., & akkerman, s. f. (2016). at the boundary of school: continuity and discontinuity in learning across contexts. educational research review, 19, 18–35. https://doi.org/10.1016/j.edurev.2016.04.001 bucher, t. (2017). the algorithmic imaginary: exploring the ordinary affects of facebook algorithms. information, communication & society, 20(1), 30–44. https://doi.org/10.1080/1369118x.2016.1154086 bucher, t. (2018). if...then: algorithmic power and politics. oxford university press. https://doi.org/10.1093/oso/9780190493028.001.0001 ceci, l. (2022a, march 15). u.s. youtube reach by age group 2020. statista. https://www.statista.com/statistics/296227/us-youtube-reach-age-gender/ ceci, l. (2022b, april 28). u.s. tiktok users by age 2021. statista. https://www.statista.com/statistics/1095186/tiktok-us-users-age/ cousin, g. (2005). learning from cyberspace. in r. land & s. bayne (eds.), education in cyberspace (pp. 117–130). psychology press. decuypere, m., grimaldi, e., & landri, p. (2021). introduction: critical studies of digital education platforms. critical studies in education, 62(1), 1–16. https://doi.org/10.1080/17508487.2020.1866050 del valle, m. e., gruzd, a., kumar, p., & gilbert, s. (2020). learning in the wild: understanding networked ties in reddit. in n. b. dohn, p. jandrić, t. ryberg, & m. de laat (eds.), mobility, data and learner agency in networked learning (pp. 51–68). springer international publishing. https://doi.org/10.1007/978-3-030-36911-8_4 deng, l., connelly, j., & lau, m. (2016). interest-driven digital practices of secondary students: cases of connected learning. learning, culture and social interaction, 9, 45–54. https://doi.org/10.1016/j.lcsi.2016.01.004 vermeire, de haan, sefton-green, & akkerman 19 | f l r dijck, j. van, poell, t., & waal, m. de. (2018). the platform society: public values in a connective world. oxford university press. foucault, m. (1995). discipline & punish: the birth of the prison (a. sheridan, trans.). vintage books. gee, j. p. (2005). semiotic social spaces and affinity spaces: from the age of mythology to today’s schools. in d. barton & k. tusting (eds.), beyond communities of practice: language power and social context (pp. 214–232). cambridge university press. https://doi.org/10.1017/cbo9780511610554.012 goodyear, v. a. (2022). young people, social media and health: a pedagogical perspective on influencers. in digital wellness, health and fitness influencers. routledge. gyldendahl jensen, c., clausen, n. r., dau, s., bertel, l. b., & ryberg, t. (2022). envisioning scenarios in designs for networked learning: networked learning conference 2022. networked learning 2022. https://www.networkedlearning.aau.dk/digitalassets/1274/1274674_proceedings-for-thethirteenth-international-conference-on-networked-learning_.pdf hasse, c. (2014). an anthropology of learning: on nested frictions in cultural ecologies. springer. hendry, n. a., hartung, c., & welch, r. (2022). health education, social media, and tensions of authenticity in the ‘influencer pedagogy’ of health influencer ashy bines. learning, media and technology, 47(4), 427–439. https://doi.org/10.1080/17439884.2021.2006691 hine, c. m. (2000). virtual ethnography. sage publications ltd. https://doi.org/10.4135/9780857020277 hjorth, l., horst, h. a., galloway, a., & bell, g. (eds.). (2017). the routledge companion to digital ethnography (pp. 21-28). routledge. hoekstra, h., jonker, t., & van der veer, n. (2022). nationale social media onderzoek 2022 [newcom research & consultancy b.v.]. ito, m., martin, c., pfister, r. c., rafalow, m. h., salen, k., & wortman, a. (2019). affinity online: how connection and shared interest fuel learning. nyu press. jenkins, h. (2006). fans, bloggers, and gamers: exploring participatory culture. nyu press. kerssens, n., & dijck, j. van. (2021). the platformization of primary education in the netherlands. learning, media and technology, 46(3), 250–263. https://doi.org/10.1080/17439884.2021.1876725 koopman, c. (2019). how we became our data: a genealogy of the informational person. in how we became our data. university of chicago press. https://doi.org/10.7208/9780226626611 lankshear, c., & knobel, m. (2007). researching new literacies: web 2.0 practices and insider perspectives. e-learning and digital media, 4(3), 224–240. https://doi.org/10.2304/elea.2007.4.3.224 orlowski, j. (director). (2020). the social dilemma [documentary-drama]. netflix. paradise, r., & de haan, m. (2009). responsibility and reciprocity: social organization of mazahua learning practices. anthropology & education quarterly, 40(2), 187-204. https://doi.org/10.1111/j.1548-1492.2009.01035.x perrotta, c., gulson, k. n., williamson, b., & witzenberger, k. (2021). automation, apis and the distributed labour of platform pedagogies in google classroom. critical studies in education, 62(1), 97–113. https://doi.org/10.1080/17508487.2020.1855597 vermeire, de haan, sefton-green, & akkerman 20 | f l r pink, s., horst, h., postill, j., hjorth, l., lewis, t., & tacchi, j. (2015). digital ethnography: principles and practice. sage. rainie, l., & wellman, b. (2012). networked: the new social operating system. mit press. https://doi.org/10.7551/mitpress/8358.001.0001 rehm, m., littlejohn, a., & rienties, b. (2018). does a formal wiki event contribute to the formation of a network of practice? a social capital perspective on the potential for informal learning. interactive learning environments, 26(3), 308–319. https://doi.org/10.1080/10494820.2017.1324495 säljö, r. (2010). digital tools and challenges to institutional traditions of learning: technologies, social memory and the performative nature of learning. journal of computer assisted learning, 26(1), 53–64. sefton-green, j. (2021). towards platform pedagogies: why thinking about digital platforms as pedagogic devices might be useful. discourse: studies in the cultural politics of education, 1–13. https://doi.org/10.1080/01596306.2021.1919999 sefton-green, j., & pangrazio, l. (2022). the death of the educative subject? the limits of criticality under datafication. educational philosophy and theory, 54(12), 2072–2081. https://doi.org/10.1080/00131857.2021.1978072 twitch. (2021). twitch.tv |. twitch advertising. https://twitch.tv/audience/ ünlüsoy, a., de haan, m., leander, k., & volker, b. (2013). learning potential in youth’s online networks: a multilevel approach. computers & education, 68, 522–533. https://doi.org/10.1016/j.compedu.2013.06.007 ünlüsoy, a., leander, k. m., & de haan, m. (2021). rethinking sociocultural notions of learning in the digital era: understanding the affordances of networked platforms. e-learning and digital media, 20427530211032302. https://doi.org/10.1177/20427530211032302 vaessen, m., van den beemt, a., & de laat, m. (2014). networked professional learning: relating the formal and the informal. frontline learning research, 2(2), 56–71. https://doi.org/10.14786/flr.v2i2.92 wenger, e. (1999). communities of practice: learning, meaning, and identity. cambridge university press. wichmand, m., pischetola, m., & dirckinck-holmfeld, l. (2023). how to design for the materialisation of networked learning spaces: a cross-case analysis. in n. b. dohn, j. jaldemark, l.-m. öberg, m. håkansson lindqvist, t. ryberg, & m. de laat (eds.), sustainable networked learning: individual, sociological and design perspectives (pp. 145– 165). springer nature switzerland. https://doi.org/10.1007/978-3-031-42718-3_9 williamson, b., gulson, k. n., perrotta, c., & witzenberger, k. (2022). amazon and the new global connective architectures of education governance. harvard educational review, 92(2), 231–256. https://doi.org/10.17763/1943-5045-92.2.231 wilson, a., hamilton, h., singh, g., & lockley, p. (2023). open is not enough: designing for a networked data commons. in n. b. dohn, j. jaldemark, l.-m. öberg, m. håkansson lindqvist, t. ryberg, & m. de laat (eds.), sustainable networked learning: individual, sociological and design perspectives (pp. 49–66). springer nature switzerland. https://doi.org/10.1007/978-3-031-42718-3_4 zuboff, s. (2019) the age of surveillance capitalism. profile books. vermeire, de haan, sefton-green, & akkerman 21 | f l r appendices questions of observational tool • concept: interest-based community of practice o aim: to capture the interest-based learning community as a community of practice centred on an interest o moment to capture: fixating identity statements/performances for the community as a whole (territorialisation) ▪ fixating activities/performances about community’s positions/positionality • ‘we are not experts at hacking’ ▪ disciplining activities/performances about fixating identity • comment deleted from chat by moderator • concept: learning o aim: to capture moments that could indicate learning is taking place o moment to capture: moments that surpass the here and now (change) ▪ change in position/positionality community • as a community we became twitch partners then because… ▪ change in position/positionality individual member • well-liked comment says that the creator has helped them understand the trans experience better • concept: platform affordances o aim: to capture the affordances of the platform for this community o moment to capture: explicit expressions by the community on how they use the platform ▪ specific features • “the emote is based on the mario burn trick that our streamer demonstrates here.” ▪ platform in general • streamers consistently use the raid function of twitch at the end of a stream to joke or support other streamers. codepen nanu et al frontline learning research vol.8 no. 1 (2020) 56 75 issn 2295-3159 the effect of first school years on mathematical skill profiles cristina nanua, eero laakkonen a minna hannula-sormunen a adepartment of teacher education, center for learning and instruction, university of turku, finland article received 17 may 2019/ revised 24 june / accepted 22 january / available online 25 february 2020 abstract this study investigated the effect of children’s first formal school years on mathematical skill profiles, measured by a variety of arithmetical skills and spontaneous focusing on numerosity (sfon) tasks. by using person-centered approach the aim was to investigate whether the amount of formal schooling is associated with mathematical skills in the same way for all children, or, whether the associations differ according to the children’s mathematical skill profiles. data was analyzed from 652 4–7-year-old children from four european countries with different school entrance ages. a person-centered approach with latent profile regression analyses on four-factor score variables identified six mathematical skill profiles with both qualitative and quantitative differences. the results revealed significant, but small effects of the amount of schooling on mathematical profiles when chronological age and country-specific school entrance age were controlled for. educational implications of the findings emphasize regarding the heterogeneity in children’s mathematical skill profiles and the potentially different effects of starting formal schooling across different profiles. keywords: arithmetical skills; spontaneous focusing on numerosity; amount of schooling; schooling effects; latent profile analyses info corresponding author: minna.hannula-sormunen@utu.fi doi: 10.14786/flr.v8i1.485 1. introduction mathematical skills are among the most influential skills needed for survival and success in modern working life (fitzsimons, 2013), and different aspects of their development have inspired an increased research interest in the last few decades (dowker, 2008). a novel dimension of recent research is the effect of different learning settings on mathematical development (clements & sarama, 2014). here we focus on the question of formal school-entry-age and whether this has an effect on mathematical skill profiles. specifically, by using person-centered approach, the aim is to investigate whether the amount of formal schooling is associated with mathematical skills in the same way for all children, or, whether the associations differ according to the children’s mathematical skill profiles. this is motivated by ambiguous evidence of the effect of school starting age in previous research. on one hand, cross-country studies have shown that children from countries where formal school starts at a younger age demonstrate better mathematical skills than their peers from countries with older school entrance ages (kavkler et al., 2000; luyten, 2006; wolke et al., 2015). developmental data shows that starting formal instruction early has a positive effect on mathematical skills, both in the short term (cliffordson & gustafsson, 2010; herbst & strawiński, 2015) and in the longer term (black, devereux, & salvanes, 2011; melhuish et al., 2008; sylva et al., 2008). on the other hand, within-country studies focusing on special groups of children as well as early interventions with both fading out and catching up effects (clements & sarama, 2014), indicate opposite trends (for a review, datar, 2006) which call for and necessitate further investigation of these effects. for example, altwicker-hámori & köllő (2012) showed that hungarian children from low socioeconomic backgrounds benefit from starting formal school later, which was attributed to effective early math education in preschool. similarly, even though there are overall positive schooling and age effects on cognitive development, younger entrance age was found to have negative effects on the economically and cognitively lowest group of a sample from different states in the united states (datar, 2006). in addition, schooling may have different effects on different quantitative skills other than age (bisanz, morrison, & dunn, 1995). the current paper uses exogenous variation in formal school entrance age policies and a cross-sectional sample from four countries with different school entrance ages to examine the effects of variation in the years of formal education on mathematical skill profiles. this cross-national and cross-sectional sample provides a large variation not only in children’s school entrance age but also in their amount of schooling, which is partly independent of the chronological age and school starting age variation in the sample. this allows for the investigation of these effects on mathematical skill profiles. there is not a linear relation between chronological age and school entrance age, neither between school entrance age and amount of schooling in our 4-age-group sample covering ages from 4 to 7 years. therefore, we can investigate if a larger amount of schooling is associated with better mathematical outcomes in different skill profiles. 1.1 mathematical skills this study aims to investigate the effects of formal education on mathematical skill profiles based on two sets of different aspects of mathematical thinking, namely arithmetical skills and spontaneous focusing on numerosity (sfon) in children from four european countries in which the starting age of formal schooling is 4, 5, 6, or 7 years. arithmetical skills represent typical mathematical skills explicitly taught in the beginning of formal school even though their learning starts well before (e.g. clements & sarama, 2014) while sfon is a measure of how frequently children focus their attention on numerical aspect and use their existing numerical skills in their everyday surroundings, thus a measure of more informal mathematical thinking that has been shown to be important predictor of later success in mathematics (e.g., hannula & lehtinen, 2005; nanu et al., 2018). having these both kinds of measures is important, since it has been shown that children do not learn and use mathematical skills only in formal learning situations, such as mathematics lessons, but also in their play and everyday life outside of explicitly mathematical, adult-guided tasks (ginsburg & seo, 1999; hannula & lehtinen, 2005). also, one important goal of mathematics education is to produce transferable mathematical skills and knowledge, which can be used outside of formal learning contexts (de corte, 1999). for these reasons, including both relatively formal and informal mathematical skills is necessary for broader coverage of developmentally relevant aspects of numeracy. here we refer to informal and formal mathematical skills similarly to ginsburg (1977) who defined informal mathematical knowledge as mathematical skills generally learned before or outside of school, often in spontaneous but meaningful everyday situations including play. it is characterized by the use of nonconventional and even self-invented symbols, strategies, or procedures rather than conventional written symbols or algorithms. formal mathematical knowledge consists of mathematical skills and concepts taught in school and include the use of conventional written numerical notation and written algorithms (ginsburg, 1977). 1.1.1 arithmetical skills in 4to 7-year-old children. before receiving formal instruction, children can read and write some arabic digits (e.g., clements & sarama, 2007; fuson, 1988) and perform basic nonverbal addition and subtraction operations (huttenlocher, jordan, & levine, 1994). after learning to count objects and recite number word sequence, children integrate counting with their initial knowledge of arithmetical operations, which forms the foundation of calculation fluency, an important mathematical skill in primary school (locuniak & jordan, 2008). several studies focusing on symbolic number knowledge recognize digit knowledge as an important predictor or mediator between informal and formal mathematical knowledge (martin, cirino, sharp, & barnes, 2014; purpura, baroody, & lonigan, 2013). formal mathematical learning is considered to start when children go to formal school and when, typically, the basics of both verbal and written arithmetic are systematically taught to children (clements & sarama, 2014). dowker (2005) suggested that the most striking difference between arithmetic in formal and informal contexts is that schoolchildren learn mostly written arithmetic, while outside of school they rely, rather, on culturally-based mental strategies. when starting formal schooling, children are typically taught arabic digits and counting-based strategies to solve addition and subtraction problems, strategies that increase the accuracy of calculation. thus, tests of arithmetical skills, such as basic addition and subtraction as well as digit knowledge, are adequate output measures of mathematical development and so are included in the current study. in order to cover skills developing both in school, preschool, and day-care contexts, including both verbal and written arithmetic measures in addition to digit naming is necessary. 1.1.2 spontaneous focusing on numerosity (sfon). hannula and lehtinen (2005) argued that the tendency to spontaneously notice exact numerosities, which leads to the self-initiated practice of mathematical skills, is particularly important for the development of early mathematical skills. hannula, lepola, and lehtinen, (2010) defined sfon as: a process of spontaneously (i.e., in a self-initiated way not prompted by others) focusing attention on the aspect of the exact number of a set of items or incidents and using of this information in one’s action. sfon tendency indicates the amount of a child’s spontaneous practice in using exact enumeration in her or his natural surroundings. (p. 395) more specifically, hannula and lehtinen (2005) showed that some children spontaneously pay attention to the numerical dimension of the environment and use that numerical information to perform different tasks, while others do not show any self-initiativeness to notice and utilize numbers in their actions. this variation is related to both concurrent and later mathematical skills even after controlling for the general attention and cognitive skills needed for the tasks (batchelor, inglis, & gilmore, 2015; bojorque, torbeyns, hannula-sormunen, van nijlen, & verschaffel, 2016; edens & potter, 2013; hannula et al., 2005, 2007, 2010; hannula-sormunen, lehtinen, & räsänen, 2015; nanu et al., 2018). hannula and lehtinen (2005) showed that sfon tendency can be enhanced at day-care by means of social interaction, and recently, braham, libertus and mccrink (2018) demonstrated that children’s sfon can be promoted in an informal setting where parent-child pairs play shopping at a museum. in line with these findings, the word “spontaneous” does not refer to the developmental origins of this tendency, but only to a momentary self-initiated focusing on the aspect of number. as evidenced by different sfon measures, individual differences in self-initiated practice using numerical skills can occur in different activities, such as copying pictures as well as performing different actions like feeding a bird, selecting the right number of socks for different monster toys, or describing pictures (gray & reeve, 2016; batchelor et al., 2015; hannula et al., 2009; hannula & lehtinen, 2005), thus children do not use and learn mathematical skills only in adult-guided explicitly mathematical tasks. recent research suggested considerable contextual effects and variability across different kinds of sfon measures (batchelor et al., 2015; rathé, torbeyns, hannula-sormunen, & verschaffel, 2016). thus, it is important to use multiple sfon tasks. 1.2 person-centered approach comparative studies on early mathematical skills and the onset of formal education have predominantly used variable-centered approaches that, while valuable in testing relations among variables, are detrimental in translating these relations at the individual level. directly related to current investigation, a previous variable-centered study did not detect any systematic effects of starting age of schooling on a variety of early mathematical skills in a four-country cross-sectional study (batchelor et al., in preparation). a person-centered approach identifies optimal combinations of skill levels, resulting in data-driven grouping of participants without having to use artificially cut-off scores (hickendorff, edelsbrunner, mcmullen, schneider, & trezise, 2017; magnusson, 2003). the identified groups with unique combinations of skill-levels can be utilized as units of analyses, allowing for the investigation of potential relations of profiles with external variables. previous studies using person-centered approaches in early math mostly investigated how the children’s mathematical skill profiles were related to different cognitive skills, as well as specific math measures (e.g., gray & reeve, 2016; hannula-sormunen et al., 2017). to our knowledge, no previous studies have analyzed the effects of formal schooling on early mathematical skill profiles based on both formal and informal set of mathematical skills. these analyses help identify children’s unique mathematical skill profiles across a range of country-specific formal schooling settings, as well as investigate the effects of formal schooling on children with different mathematical skill profiles. what is more, these profile differences can be explored in relation to other relevant aspects such as socioeconomic status and chronological age, both factors that have been sources of individual differences in previous early math studies (altwicker-hámori & köllő, 2012; sprietsma, 2010). one advantage of person-centered approach is the way how it creates educationally relevant information for preparing tailored education for different groups of learners with their unique strengths and weaknesses across wide range of measures (mcmullen & hickendorff, 2018). this analytical strategy allows investigating whether school starting age and amount of schooling have differing effects on different groups of children as represented in the mathematical skill profiles. 1.3 the present study the aim of the present study is to investigate the effects of children’s first years in formal school on their mathematical skill profiles by using a person-centered approach. a cross-sectional sample with 4–7-year-old children came from four western european countries. the countries differ systematically in their starting age for school and, consequently, also in the amount of formal schooling they provide for children of the same age. variability in the years of schooling may be a more sensitive indicator of the effects of formal education than just school entrance age, which reflects on both country-specific differences in formal schooling policies as well as entrance age effects. our research question is as follows: does the amount of schooling have unique effects on children’s mathematical skill profiles when chronological age, school starting age / country, and socioeconomic status, i.e., ses are controlled for? in order to investigate this question, first, mathematical skill profiles including both relatively informal and formal mathematical skills are formed. school entrance age refers to the age children start school based on policies in each country while the amount of schooling is a characteristic of an individual participant, indicating how many years they have been in formal school. because of the wide age range of the participants, children that started school at same age vary in amount of schooling. 2. method 2.1 participants participants were 652 children aged from 4 to 7 years from northern ireland, england, belgium, and finland (ca. 40 per age group in each country), where formal schooling begins at different ages (4, 5, 6, or 7 years, respectively). the study is part of a larger project named international comparison of children’s attention and learning, iccal. schools, preschools, and daycare centers were selected from middle ses neighbourhoods. educators distributed and collected sealed parent questionnaires with ses information, including mothers’ occupations, education, and income. table 1 shows frequencies of participants by school entrance age and years of formal schooling. table 1. number of participants as a function of school entrance age/country, age in months and amount of schooling in years note 1 there were approximately 40 children in each age group (i.e., 4-, 5-, 6-, and 7-year-old groups) in each country, thus in the no formal schooling group, in belgium, there were 42 children in 4year-old and 42 5-year-old children’s group, and in finland 41 4-year-old, 40 5-year-old and 39 6-year-old children’s age group. all countries had obligatory preschool education year before the actual formal school entrance age. permission was obtained from parents and verbal assents from the children to conduct this research. the study followed the guidelines of the ethical advisory board at each institution, and permissions were granted by local school, kindergarten, and daycare administrators. for this study, we included a random selection of equally-sized samples from all countries and age groups so that all groups would be equally represented. according to t-tests, this sample (n = 652) did not differ from the original sample (n = 685) in ses, age, arithmetical skills, or sfon factor-score variables (all p’s > 0.05). a one-way anova revealed a significant age difference between corresponding age group samples in the countries, f(3, 651) = 10.41, p < 0.001, = 0.05. children in northern ireland (m = 68.22, sd = 14.06) and england (m = 66.86, sd = 13.70) were significantly younger in all age groups than those in corresponding age groups in belgium (m = 73.48, sd = 13.69) and finland (m = 73.46, sd = 13.44) because the cutoff date for entry to school in northern ireland and england is later in the calendar year (northern ireland vs. belgium, t[328] = -3.44, p = 0.003, d = -0.38; northern ireland vs. finland, t[320] = -3.42, p = 0.004, d = -0.38; england vs. belgium, t[328] = -4.39, p < 0.001, d = -0.48; england vs. finland, t[320] = -4.36, p < 0.001, d = -0.49). also, there was a significant ses difference between the samples from four countries, f(3,456) = 5.22, p < 0.001, = 0.03. post-hoc comparisons with bonferroni correction showed that ses (see table 2 for ses descriptives) was significantly higher in belgium (m = 0.29, sd = 0.96) compared to northern ireland (m = -0.20, sd = 1.10, t[215] = 3.54, p = 0.002, d = 0.47), england (m = -0.06, sd = 1.02, t[226] = 2.69, p = 0.047, d = 0.35), and finland (m = -0.07, sd = 0.92, t[255] = 3.22, p = 0.015, d = 0.40). 2.2 procedure children were tested during the spring term in march-april, individually in a quiet space in their own school, preschool, or daycare setting by one of ten trained research assistants in two 30-minute testing sessions separated by a short break. four sfon tasks—two imitation sfon tasks, a picture description sfon task, and a memory card game sfon task—were presented in the first session while digit naming, verbal arithmetic, and written arithmetic were assessed in the second session, and all tasks were administered in the same order for all participants. in order to ensure similar testing procedures across the 10 research assistants, group training was carried out before the assessments and videotaped test sessions were checked throughout the data collection period. in between the tasks, children were involved in short physical activities to refocus their attention and maintain engagement levels. the tester sat next to the child when administering the imitation tasks and sat behind the table opposite to the child in all other tasks. during the assessment, children received general praise but no specific feedback. the tester ensured that the testing area did not have any numerical displays that might have prompted the children to focus on numbers or helped them to solve arithmetical problems. to keep the numerical aspect of the study concealed, children, parents, and teachers were told that the study was an international comparison of children’s attention and learning. multilingual, close to native, and speakers of the three languages involved (english, dutch, and finnish) translated and back-translated all task materials carefully to ensure that tasks were identical across countries and languages. 2.2.1 sfon measures bird imitation task (hannula & lehtinen, 2005) materials were a toy parrot, placed on the table in front of the child, and two plates of differently colored glass berries placed in front of the parrot. in the first trial, the tester introduced the materials and said, “watch carefully what i do, and then you do just like i did.” the tester put three berries (two red and one blue), one at a time, into the parrot’s mouth, and these dropped with a bumping sound into the parrot’s stomach, where the child could not see them. next, the child was told, “now you do exactly like i did.” three green and two white berries were used in the second trial, and two yellow and three blue berries were used in the third. the tester filled out structured observation forms to record the child’s behavior for any signs of counting or enumeration during the imitation and memory card game sfon tasks. all of the child’s (a) utterances, including number words (e.g., “i’ll give him two berries”), (b) use of fingers to express numbers, (c) counting acts, such as a whispered number word sequence and indicating acts by fingers, (d) other comments referring either to exact quantities or counting (e.g., “oh, i miscounted them”) or (e) interpretation of the task’s goal as quantitative (e.g., “i gave an exactly right number”) were identified. the child was given a sfon score of 1 if they produced the same numerosity as the tester and/or if they were observed presenting any of the mentioned (a–e) quantifying acts. the maximum score for the task was 3. cronbach’s alpha was 0.74. picture description task (batchelor, inglis, & gilmore, 2015) in each of the three trials, the child had to describe a cartoon picture that contained several elements (objects, people, and animals) that could be enumerated, such as an image of a landscape with two children flying in a hot balloon, three houses in the background, and three clouds. the tester displayed a picture and asked, “what can you see in this picture?” there was no time limit. when the child stopped describing the picture, the tester asked, “is that everything?” before displaying the next picture. responses were audio recorded. the child received a sfon score of 1 if they (a) mentioned cardinal or ordinal number words while describing the picture or (b) stated a number word list (at least two consecutive number words). the maximum score for this sfon task was 3. because the dutch word for “a” is the same as the word for “one,” the number word “one” was excluded to avoid possible language effects. cronbach’s alpha was 0.74. postbox imitation task (hannula & lehtinen, 2005) the materials were a postbox, placed in front of the child on the table, and six piles of ten envelopes, each pile different in color. before starting each trial, two piles of envelopes were placed in front of the postbox. verbal instructions were similar to the bird imitation task. for the first trial, the tester put one orange envelope and two green envelopes, one at a time, into the postbox. then the child was asked to do exactly as the tester had done. for the second trial, two brown and three yellow envelopes, and for the third trial, three blue and two pink envelopes, were used. the scoring criteria was the same as the one used in the bird imitation task. the maximum score for this sfon task was 3. cronbach’s alpha was 0.69. memory card game (developed from hannula, grabner, & lehtinen, 2009) in this task, children were asked to look at a card, memorize it, and, immediately after turning it over, describe the card so that the tester could find a matching card from their pile of cards. there were four sets of cards with photos of an array of everyday items and toys. the child had a pile of cards in front of them at the table and the tester had four cards from each set in their hand. for each trial the tester asked the child to turn over their top card and look at it for five seconds. the tester then asked the child to turn the card facedown and describe what was on the card so that they could find the matching card from the tester’s pile. the child caught a brief flash of the pile to show that the cards were “very similar,” and thus, careful description was required. there was no time limit to respond. when the child stopped, the tester asked, “what else do you remember about this card?” the trial then finished with the tester playfully trying to find the matching card from their pile. there was one practice trial followed by three test trials. responses were audio recorded and transcribed. for each trial, children scored either 0 or 1, depending on whether they focused on numerosity. the scoring was based on audio recordings and the tester’s written observations of signs of counting during the game. they were determined to be focusing on numerosity if their description included any number word utterances referring to the items in the child’s cards and/or if the tester observed any signs of counting or comments about the numerical goal of the task (e.g., “how many red ones were there?”). as with the picture task, the number word “one” was excluded to avoid language effects. children received a total score out of 3 (cronbach’s alpha = 0.74). 2.2.2 arithmetical skills measures digit naming participants were asked to read aloud a series of arabic digits. there were ten blocks with three trials per block (ranging from oneto five-digit numbers). the test was discontinued when a child made a mistake on all three items within a block. children scored 1 point for each correct identification of a digit, resulting in a maximum total score of 30 (cronbach’s alpha = 0.96). verbal arithmetical task verbal arithmetical skills (developed from hannula & lehtinen, 2005) were measured with a task containing addition and subtraction items. in the first part, six addition and subtraction items were illustrated with supporting material consisting of glass sweets (1.5 cm in diameter) and a non-transparent box. for instance, the tester quickly showed four glass sweets in her hand to the child and said, “there are four sweets over here. i put them in this box. then i put three more sweets there. how many sweets are there in the box now?’’ the child was not able to see the end result in the box. this comprehension support was used to make sure that the measure would target arithmetical skills without heavy reliance of verbal comprehension skills. one point was awarded for each correctly solved item, thus, there was a maximum total score of 6. cronbach’s alpha was 0.71. the second part comprised nine items for addition and nine items for subtraction with oneto two-digit numbers. children were asked this type of question: “what do you get when you add two and two together?” one point was awarded for each correctly solved item, thus there was a maximum total score of 18. cronbach’s alpha was 0.94. written arithmetical task the written arithmetical task was developed from aunola, leskinen, lerkkanen, and nurmi’s (2004) research. children were asked to solve a series of written arithmetical problems that consisted of addition and subtraction operations. the problems increased in difficulty, ranging from oneto three-digit addends and subtrahends, and were arranged in four blocks of six items. each subtest was discontinued when a child made two or more mistakes within a block. children scored 1 point for each correct solution, giving a maximum total score out of 48 (cronbach’s alpha = 0.98). 2.3 analytical strategy all analyses were completed using mplus software, version 8 (muthén & muthén, 1998–2015). the full information ‘maximum likelihood’ estimation with ‘robust standard errors’ was used. exploratory factor analyses on raw sum scores of all arithmetical skill measurements and categorical sfon scores were conducted (see appendix a) in order to condense the data, to get the number of factors and to allow items of sub-tests to be represented as the factor scores based on their loadings. the main advantage of a factor score over a summated scale is that it is based on the factor loadings of all variables used in the analysis. cfas were run to confirm the efa solution for the entire sample as well as for all countries and age groups. next, standardized factor scores of confirmatory factor analyses with the factors using whole sample were saved as variables. the variable school starting age / country refers to school starting ages, 4 for northern ireland, 5 for england, 6 for belgium, and 7 for finland. the amount of schooling was calculated as a categorical variable with the following categories: 0 no schooling, 1, 2, 3, and 4 years of formal schooling. using the cfa factor scores saved as variables, latent profile analysis (lpa) was utilized to identify groups of children with homogeneous mathematical skill profiles (i.e., lpa and mixture modelling) (magidson & vermunt, 2002; nylund, asparouhov, & muthén, 2007). lpa is an explorative model-based procedure for identifying the smallest number of meaningful profiles that best group participants within a sample while accounting for the probability of belonging to each latent profile (marsh, lüdtke, trautwein, & morin, 2009). the models were evaluated by using a combination of theoretical interpretability of profiles and statistical indicators—bayesian information criterion (bic), akaike´s information criterion (aic), entropy, vuong lo mendell rubin likelihood ratio test (vlmr), and bootstrap likelihood ratio test (blrt)—to identify the best-fitting model (collins & lanza, 2010). smallest values of bic and aic are indicators of better model fit. the two likelihood ratio tests give a p value which indicates if a k-class model fits the data better compared to a k 1 class model. in order to complement a variable-centered approach, lpa should identify groups that reflect a combination of level and shape differences (marsh et al., 2009). the quality of classification is represented by the average posterior probabilities (close to 1) of being assigned to a specific latent profile and by the entropy value (greater than .8 or closest to 1.0 as acceptable value). next, in order to study the effects of schooling on mathematical skill profiles, lpa with covariates ses, age, school starting age / country, and amount of schooling was run, as this allows for simultaneous lpa and regression of the latent profiles on covariates (nylund-gibson & masyn, 2016), and therefore determines their unique effects on latent profile membership and the mean levels of the profiles. this analytical strategy allows investigating whether covariates have differing effects on different groups of children as represented in the mathematical skill profiles. in addition, direct effects of covariates on factor score variables were tested by adding their effects individually to each factor score variable, as recommended in an lpa model building process with covariates (masyn, 2013). both pathways (the indirect effects via latent variable and the direct effects) through which a covariate could affect latent profile indicators were included in the model (nylund-gibson & masyn, 2016). by allowing direct effects, potential differences in the indicators of the profiles as a function of the covariates are accounted for. finally, likelihood ratio tests were used to evaluate the statistical significance of each covariate’s unique effect on profile membership by adding covariates hierarchically and comparing the fit of the models following collins and lanza’s (2010) procedure., the odds of belonging to a profile as a function of covariates were investigated in the final model. the odds ratios describe the strength of the association of the covariate on the profiles. 3. results 3.1. descriptives and explorative factor analyses descriptives of all variables are presented in table 2. table 2. descriptive statistics of the mathematical skill measures note. ses was a standardized score (based on mothers’ occupation, education, and income) modeled as a formative measure using principal component analysis [e.g., caro & cortés, 2012]). 1a child was missing from the second testing session. 2missing data due to failure in audio recording or comprehension of the task. exploratory factor analyses on raw sum scores of all arithmetical skill measurements and categorical sfon scores were conducted. these showed the best fit for a four-factor solution with factors of (factor 1) arithmetical skills consisting of digit naming, verbal and written arithmetical items, and all sfon items loading along their task type, i.e., (factor 2) sfon imitation, (factor 3) sfon memory, and (factor 4) sfon picture description (see appendix a). next, cfa was run to confirm the efa solution. results showed good model fit, χ² (98) = 164.88, p < .001, cfi = 0.99, rmsea = 0.032. the same model fitted well (hu & bentler, 1999) also for each school starting age / country and for each age group. cfi values for the models of the countries, and also, age groups varied between 0.96 and 1.00, while rmsea values for the models of the countries varied between 0.00 and 0.05 and for the age groups between 0.00 and 0.04. 3.2 latent profile analysis latent profile analysis (lpa) was used to identify groups of children with homogeneous mathematical skill profiles. table 3 illustrates the fit indices for all lpa solutions. although the minimum bic and aic were not found, the bic value becomes more stable at the third and the sixth profile solutions. in addition, vlmr lrt do not improve significantly in the 7-profile solution, thus supporting the 6-profile solution (i.e., lpa and mixture modelling) (magidson & vermunt, 2002; nylund, asparouhov, & muthén, 2007). the 6-profile solution recognizes sub-groups that are both qualitatively and quantitatively differing from each other. importantly, qualitative differences are theoretically supported by previous research showing variability and contextual effects across different sfon and math measures (see rathe et al., 2016), unlike the 3-profile solution, which has only quantitative (i.e. level) differences (marsh, lüdtke, trautwein, & morin, 2009). thus, the 6-profile solution was chosen. after calculating the lpa, it was checked whether the used cfa scores masked possible higher order interactions. for this purpose, the lpa analysis was repeated with scale values that were calculated independently. the structural results of the first lpa were basically confirmed. table 3. fit measures of the 2to 7-latent profile models for mathematical latent factor score variables note. 1 unstable solution. selected solution is marked with bold. 3.3 lpa with covariates first, we investigated the significance of each covariate in relation to profile structure. we constructed separate lpa models for each covariate (ses, age, school starting age / country, or amount of schooling) and compared the fit of each model with the fit of a model without any covariates. based on chi-square difference tests, age (δχ2(8) = 426.77, p < 0.0001), school starting age / country (δχ2(8) = 54.29, p < 0.0001), and amount of schooling (δχ2(8) = 162.64, p < 0.0001) were significantly related to mathematical skill profiles, while ses was not (δχ2(5) = 4.97, p = 0.419). therefore, ses was not included in further analyses. in the following step, we investigated whether the amount of schooling contributes significantly to the prediction of latent profile membership over and above the contribution of age and school starting age / country. multicollinearity was tested for the covariates of age, school starting age / country, and amount of schooling. spearman correlations between covariates ranged from 0.17 for age and school starting age / country to -0.67 for school starting age / country and amount of school. vif was smaller than 10 and tolerance was larger than 0.01, as recommended by menard (1995) and myers (1990) for all the covariates (age: vif = 4.35, tolerance = 0.23; school starting age / country: vif = 5.78, tolerance = 0.17; amount of schooling: vif = 7.65, tolerance = 0.13). thus the analyses could be continued. the effects of amount of schooling were tested by hierarchically adding the covariates and comparing the models with chi-square difference tests (table 4, figure 1). significant direct effects of the covariates on the mathematical skills factor score variables were also included in all models. table 4. statistical significance of adding the covariates hierarchically in the prediction of latent profile membership note. ℓ represents log-likelihood from fit models. statistical significance of each covariate was tested with the likelihood-ratio test. the final latent profile regression model, where all covariates are allowed to have both direct effects on profile membership and significant direct effects on factor score variables, is illustrated in figure 1. these results show that the amount of schooling had a significant, unique effect on the profile structure. also, the amount of schooling had direct effects on arithmetical skills and on sfon imitation when all other covariates were controlled for. chronological age had direct effects on arithmetical skills and on sfon picture description. figure 1. final latent profile regression model (last model from the hierarchical steps). all presented regression coefficients are significant at p < 0.001. school starting age /country refers to the child’s school entrance age in years (4, 5, 6, or 7 years of age), and amount of schooling refers to how many years the child had experienced formal schooling at the time of testing (ranging from 0 to 4 years). age (in months) refers to child’s chronological age at the time of testing. figure 2 illustrates the estimated means of factor score variables for the six mathematical skill profiles when the covariates age, school starting age / country, and amount of schooling are included in the model. the average posterior probabilities of being assigned to a specific latent profile in the final 6-profile model were 0.90, 0.89, 0.96, 0.91, 0.93, and 0.94, respectively, indicating a clear classification. profiles were ordered from lowest to highest based on arithmetical skills and sfon imitation. identified profiles differed both in their mean levels of factor score variables and their shape. there are only factor score mean level differences in arithmetical skill and sfon imitation factors across all profiles, indicating that these two skill factors are highly associated at the person-centered level. not only quantitative level differences, but also qualitative differences appear in profiles 2, 3, 4 and 5 in sfon memory and picture description, suggesting that the tasks requiring verbal responses do not follow the same pattern of individual profile levels as arithmetical and sfon imitation factor scores. the lowest profile 1 had lowest mathematical skills in all factor score variables. the second lowest profile 2 had below average scores in all other mathematical skills except for sfon picture description, in which they had above average scores. the profile 3 had average mathematical skills in all other factors except for their high sfon picture description skill while the children in the profile 4 had average factor scores in all other factors, except for in sfon picture description, in which they had lower than average scores. the second highest profile 5 children had above average factor scores in all other measures, except for in the sfon picture description, in which they were on average. the highest profile 6 children had highest mathematical skills in all factor score variables. figure 2. latent profiles of mathematical factor score variables with covariates age, school starting age / country, and amount of schooling. percentages represent proportion of the sample in each profile. 3.4 odds ratios odds ratios were computed to investigate the change in odds of membership to different profiles as an effect of each covariate (figure 3) when all other covariates were controlled for. we chose the profile 1 – lowest mathematical skills to be the comparison group and compared the odds of other five profiles in covariates to the odds of this group. results show that there was a a significant difference in the odds ratios in age when comparing the profile 1 with all other five profiles and significant difference in the odds ratios of amount of schooling when comparing the profile 1 with the profiles 4 and 5. when age increases with one unit, in comparison to profile 1, the odds of belonging to one of the higher-level profiles are significantly higher. for the amount of schooling, the odds of belonging to profile 4 average mathematical skills but low sfon picture description [b = -1.060, p = .030, or = .347 (95% ci: .133, .903)], and profile 5 – second highest mathematical skills but average sfon picture description [b = -1.061, p = .037, or = .346 (95% ci: .128, .938)], are significantly lower if the amount of schooling increases with one unit in profile 1. due to large variation in the school starting age / country and amount of schooling, profile 6 – highest mathematical skills did not differ significantly from profile 1. for school starting age / country, no significant odds ratio was found. figure 3. odds ratios (or) of age, school starting age / country, and amount of schooling having profile 1 as a reference class. significant or are marked with * and 95% confidence intervals are illustrated by the straight lines. or represent the odds of belonging to a higher profile compared to profile 1 when there is one unit increase in age, school starting age / country, or amount of schooling, when the other two covariates are controlled for. if or = 1, groups have the same odds ratio. 4. discussion in this study, we investigated mathematical skill profiles in relation to their amount of formal schooling. the sample comprised of children who started school within a four-year range in four countries, and thus differed largely also in their years of schooling. particularly, by using person-centered approach, the aim was to investigate whether the amount of formal schooling is associated with mathematical skills in the same way for all children, or, whether the associations differ according to children’s mathematical skill profiles. to build up the profiles, we measured formal mathematical skills, such as digit naming and written and verbal arithmetical skills, with more informal mathematical skill called spontaneous focusing on numerosity (sfon), which is a form of self-initiated mathematical activity leading to differences in mathematical practice (hannula & lehtinen, 2005). having the school entrance age spread across four years offered a unique opportunity to disentangle the effects of country-specific school entrance age, amount of schooling, and chronological age (herbst & strawinski, 2016). our results showed that years of formal schooling had unique effects on mathematical skill profiles before and after the effects of chronological age and country-specific school entrance age were controlled for. ses was not significantly related to math skill profiles. differences in the profile levels and shape illustrate both quantitative and qualitative differences in mathematical skill profiles, which confirms the existence of a nonlinear relation between the investigated early mathematical skills and years of schooling (chew, forte, & reeve, 2016). almost one third of the participants belonged to profiles with variable skill levels across our measurements, thus having: either high arithmetical skills and sfon imitation combined with either low sfon memory card and sfon picture description skills, or low arithmetical skills and sfon imitation combined with high sfon memory card and sfon picture description skills. this is in line with previous research showing task-related variability in sfon (batchelor, inglis, & gilmore, 2015; rathe et al., 2017). the multiple comparisons provided by odds ratios suggest that the relationship between schooling years and mathematical skill profiles is negative when school entrance age and chronological age are controlled for, specifically when comparing the lowest level profile with two higher level profiles (profiles 4 and 5, which have above average or average mathematical skills across nearly all factor score variables), so that children have a higher likelihood of belonging to profile 1, the lowest level profile, if they have experienced a larger amount of schooling. this negative association may first seem counterintuitive, but it is similar to typical unique age-related effects in contrast to composite age-effects (herbst, 2016). chronological age had strong, positive effects in our model showing that older children have better chances for belonging to profiles with higher levels of sub-skills, thus, controlling for its effects on skill profiles is needed when any other effects are investigated. in this respect, our results of the lowest skill profile are similar to datar (2006) who shows that younger entrance age for kindergarten negatively affects the economically and cognitively lowest group of the sample. recognition of the negative association of schooling and specific mathematical skills in a particular sub-group of children adds educationally-relevant knowledge to previous variable-centered research findings, which showed positive educational effects on early numeracy on average level (bojorque et al., 2016; hannula, mattinen & lehtinen, 2005; watts et al., 2017). identifying these kinds of unique associations on sub-populations can be educationally beneficial when dealing with heterogeneous populations (abenavoli, greenberg, & bierman, 2017). along these lines, the generally positive effects of schooling on mathematical skills that have been found (kavkler et al., 2000; luyten, 2006; wolke et al., 2015) are accompanied by recent research showing that for lower achievers, instructional quality is more strongly related to performance than for high achievers (crosnoe et al., 2010; hamre & pianta, 2005). our findings raise important questions concerning, first, the differences in nature and quality of our participants’ early mathematical support in different educational contexts, such as at day care, in preschool or in school, and, second, the differing effects early formal school start may have for some sub-groups of children. over-representation of children with more years of formal schooling in the lowest mathematical skill profile could indicate that formal school math settings as they are provided in the countries which start school early is not optimal for early mathematical development. however, and importantly, this effect is rather small, and it needs replications with larger samples and wider set of mathematical skills studied before conclusions concerning school starting age policies can be made. this study adds to our current knowledge of children’s domain-specific attentional skills in relation to formal mathematical skills and formal education. first, the observed similarity in sfon measures within most of the profiles is in line with previous studies showing that there is a more general sfon tendency component across different sfon measures (gray & reeve, 2016; hannula & lehtinen, 2005; hannula-sormunen et al., 2015). however, the factor structure suggested that sfon might be a multidimensional construct, unlike digit naming, verbal and written arithmetic, which formed only one factor. this is in line with a series of studies showing discrepancies between a relatively new picture description task and the original action-based sfon measures (batchelor et al., 2015; rathé, et al. 2017). interestingly, the variation in the means of different profiles across sfon measures was such that the mean level of imitation tasks followed the same pattern as arithmetical skills, while the qualitative differences in the profiles were in the measures requiring verbal responses. additional research is needed to investigate the potential multidimensional structure of sfon tendency. so far, data has suggested that verbally-based sfon is different from action-based sfon (batchelor et al., 2015; rathé et al., 2017). since sfon picture description task drives qualitative differences between profiles, it would be worth exploring more closely what aspects of this verbally-based task produce these differential effects. whether these differences are related to deeper processing requirements for imitation and memory card game in contrast to sfon picture description with continuously visible stimuli, need to be further investigated in the future. in imitation and memory card game sfon tasks the stimuli disappears from sight before the child performs the acts defining his or her focusing aspect. our study showed for the first time that by adding a verbal sfon measure that requires memory retrieval skills, factor analyses suggested a three-factor structure for sfon while previous studies had been demonstrating differences between verbal and non-verbal tasks (batchelor, inglis & gilmore, 2015; rathé et al., 2017). 4.1. limitations the analyses of the effects of age, school starting age / country, and amount of schooling on mathematical profiles were limited by the cross-sectional design of our study, even though the sample covering different school entrance ages across same-aged children also allows for the exploration of these effects. longitudinal data would be required to explore the developmental trajectories of mathematical profiles, and experimental designs to draw conclusions on causality between covariates and mathematical development. importantly, our study focused on a rather limited set of formal and informal mathematical measures, thus any generalizing to a broader mathematical achievement should be done cautiously. there were no measures of verbal skills, working memory or inhibitory control, which could affect the association of formal schooling and mathematical skills. for instance, welsh et al. (2010) showed that working memory and attention control predicted growth in numeracy skills between 4 and 6 years of age. however, gray and reeve (2016) showed that patterns of strength and weaknesses in mathematical skill profiles in preschoolers are mostly related to domain-specific abilities such as numerical distance effect and sfon, while general cognitive skills (working memory, response inhibition, attention, and vocabulary) do not differ between profiles. in addition, this study investigated only how the amount of formal schooling, in years, and school entrance age is related to mathematical skill profiles while chronological age was controlled for. equally important would be to study the content and quality of mathematics education in different educational settings. this remains to be investigated in future research. we utilized standardized means in the analyses. therefore, the identified profiles can only be compared with the average skill level in the present sample. in addition, age groups in each school starting age / country were rather small and came from middle-ses areas, thus more representative, larger samples as well as measures of other developmentally important and academic skills would be required for drawing conclusions that policy makers could use in deciding about school entrance age policies. however, current study demonstrates how, potentially, person-centered approach with its focus on individual skill profiles can shed light on differing educational effects of school starting age policies. the countries were purposely selected so that participants would differ in school entrance age but would have relatively similar, middle-ses western european backgrounds. previous cross-cultural comparisons of early mathematical skills typically showed similar early mathematical skills when comparing children from countries with similar cultural backgrounds (aunio, korhonen, bashash, & khoshbakht, 2014; van de rijt et al., 2003). our selection of the sample allowed us to identify the effect of formal schooling on mathematical skill profiles. our sample comes from four different countries, three languages, and spreads across four years in age, which all improve generalizability of the profiles of formal and informal mathematical skills. however, we were not able to gather more specific information about the educational settings. 4.2 conclusions and future directions the person-centered approach adopted in this study demonstrated that children’s amount of formal schooling have unique effects on their mathematical skill profiles when controlling for chronological age and school starting age / country. educational implications of the findings emphasize the heterogeneity in mathematical skills across sub-groups of children, as well as the need to address these individual differences in the assessment and support of early mathematical development. is earlier formal mathematics instruction the best option to develop arithmetical skills and sfon tendencies, as well as other mathematical skills, for all children, particularly those with the lowest mathematical skill profile? this is an exciting and important question for future studies focusing on the elements of different early learning environments. these studies should clarify what kind of educational arrangements would optimally support mathematical development in the early years. keypoints latent profile regression model was used to study schooling effects cross-sectional sample from 4 countries with different school entrance ages enabled investigation arithmetical skills and spontaneous focusing on numerosity formed skill profiles amount of schooling has unique, and differing effect on math skill profiles specifically lowest math skill profile was related to negative schooling effect acknowledgements we would like to warmly thank the participating children, schools, preschools, kindergartens and day care centers, and the families of children who took part in the project. we would also like to thank joke torbeyns (ku leuven), sophie batchelor (loughborough university), bert de smedt (ku leuven), victoria simms (ulster university) and jake mcmullen (university of turku) for their generous help in different phases of the project and all the research assistants who collected data for the project, as well as lieven verschaffel (ku leuven) and erno lehtinen (university of turku) for their guidance. the research was supported by funding the academy of finland, awarded to mh-s (278579). references abenavoli, r. m., greenberg, m. t., & bierman, k. l. (2017). identification and validation of school readiness profiles among high-risk kindergartners. early childhood research quarterly, 38, 33–43. doi: 10.1016/j.ecresq.2016.09.001 altwicker-hámori, s., & köllő, j. (2012). whose children gain from starting school later?—evidence from hungary. educational research and evaluation, 18(october), 459–488. doi: 10.1080/13803611.2012.695142 aunio, p., korhonen, j., bashash, l., & khoshbakht, f. (2014). children’s early numeracy in finland and iran. international journal of early years education, 22(4), 423–440. doi: 10.1080/09669760.2014.988208 aunola, k., leskinen, e., lerkkanen, m.-k., & nurmi, j.-e. (2004). developmental dynamics of math performance from preschool to grade 2. journal of educational psychology, 96(4), 699–713. doi: 10.1037/0022-0663.96.4.699 batchelor, s., inglis, m., & gilmore, c. (2015). spontaneous focusing on numerosity and the arithmetic advantage. learning and instruction, 40, 79–88. doi: 10.1016/j.learninstruc.2015.09.005 batchelor, s., torbeyns, j., simms, v., nanu, c., laakkonen, e., de smedt, b., & hannula-sormunen, m. (in preparation). the effect of school starting age on children`s spontaneous focusing on numerosity and mathematical skills. bisanz, j., morrison, f. j., & dunn, m. (1995). effects of age and schooling on the acquisition of elementary quantitative skills. developmental psychology, 31(2), 221-236. doi: 10.1037/0012-1649.31.2.221 black, s. e., devereux, p. j., & salvanes, k. g. (2011). too young to leave the nest? the effects of school starting age. review of economics and statistics, 93(2), 455–467. doi: 10.1162/rest_a_00081 bojorque, g., torbeyns, j., hannula-sormunen, m., van nijlen, d., & verschaffel, l. (2016). development of sfon in ecuadorian kindergartners. european journal of psychology of education, 1–14. doi: 10.1007/s10212-016-0306-9 caro, d. h., & cortés, d. (2012). measuring family socioeconomic status: an illustration using data from pirls 2006. ieri monograph series: issues and methodologies in large-scale assessments, 5, 9–33. doi: 10.1787/9789264091504-en chew, c. s., forte, j. d., & reeve, r. a. (2016). cognitive factors affecting children`s nonsymbolic and symbolic magnitude judgment abilities: a latent profile analysis. journal of experimental child psychology, 152, 173-191. doi: 10.1016/j.jecp.2016.07.001 clements, d. h., & sarama, j. (2007). early childhood mathematics learning. in f. lester (ed.), second handbook of research on mathematics teaching and learning (pp. 461–488). charlotte, nc: information age publishing. clements, d. h., & sarama, j. (2014). the importance of early years. in r. e. slavin (ed.), science, technology & mathematics (stem) (pp. 5-9). thousand oaks, ca: corwin. cliffordson, c., & gustafsson, j.-e. (2010). effects of schooling and age on performance in mathematics and science: a between-grade regression discontinuity design with instrumental variables applied to swedish timss 1995 data. fourth iea international research conference (irc-2010), gothenburg, sweden, 1–13. collins, l. m., & lanza, s. t. (2010). latent class and latent transition analysis. new jersey, nj: john wiley & sons. crosnoe, r., morrison, f., burchinal, m., pianta, r., keating, d., friedman, s. l., … the eunice kennedy shriver national institute of child health and human development early child care research network. (2010). instruction, teacher-student relations, and math achievement trajectories in elementary school. journal of educational psychology, 102(2), 407–417. http://doi.org/10.1037/a0017762 datar, a. (2006). does delaying kindergarten entrance give children a head start? economics of education review, 25, 43-62. retrieved from http://www.sciencedirect.com/science/article/pii/s0272-7757(05)00011-7 de corte, e. (1999). on the road to transfer: an introduction. international journal of educational research, 31, 555–559. doi: 10.1016/s0883-0355(99)00023-3 dowker, a. (2005). individual differences in arithmetic. implications for psychology, neuroscience and education. new york, ny: psychology press. dowker, a. (2008). individual differences in numerical abilities in preschoolers. developmental science, 11(5), 650–654. doi: 10.1111/j.1467-7687.2008.00713.x edens, k. m., & potter, e. f. (2013). an exploratory look at the relationships among math skills, motivational factors and activity choice. early childhood education journal, 41(3), 235–243. doi: 10.1007/s10643-012-0540-y fitzsimons, g. (2013). doing mathematics in the workplace. a brief review of selected literature. adult learning mathematics: an international journal, 8(1), 7-19. fuson, k. (1988). children’s counting and concepts of number. new york, ny: springer verlag. ginsburg, h. p. (1977). children’s arithmetic: the learning process. oxford, england: van nostrand. ginsburg, h. p., & seo, k.-h. (1999). mathematics in children’s thinking. mathematical thinking and learning, 1, 113-129. gray, s. a., & reeve, r. a. (2016). number-specific and general cognitive markers of preschoolers’ math ability profiles. journal of experimental child psychology, 147, 1–21. doi: 10.1016/j.jecp.2016.02.004 hamre, b. k., & pianta, r. c. (2005). can instructional and emotional support in the first grade classroom make a difference for children at risk of school failure? child development, 76(5), 949–967. doi: 10.1111/j.1467-8624.2005.00889.x hannula-sormunen, m. m., lehtinen, e., & räsänen, p. (2015). preschool children’s spontaneous focusing on numerosity, subitizing, and counting skills as predictors of their mathematical performance seven years later at school. mathematical thinking and learning, 17(2–3), 155–177. doi: 10.1080/10986065.2015.1016814 hannula, m. m., grabner, r., & lehtinen, e. (2009). neural correlates of spontaneous focusing on numerosity (sfon) in a 9-year-longitudinal study of children’s mathematical skills, paper presented at the biennial meeting of the european association for research in learning and instruction, amsterdam, the netherlands. hannula, m. m., grabner, r., lehtinen, e., laine, t., parkkola, r., & ansari, d. (2009). neural correlates of spontaneous focusing on numerosity (sfon). neuroimage, 47, s39–s41. hannula, m. m., & lehtinen, e. (2005). spontaneous focusing on numerosity and mathematical skills of young children. learning and instruction, 15(3), 237–256. doi: 10.1016/j.learninstruc.2005.04.005 hannula, m. m., lepola, j., & lehtinen, e. (2010). spontaneous focusing on numerosity as a domain-specific predictor of arithmetical skills. journal of experimental child psychology, 107(4), 394–406. doi: 10.1016/j.jecp.2010.06.004 hannula, m. m., mattinen, a., & lehtinen, e. (2005). does social interaction influence 3-year-old children’s tendency to focus on numerosity? a quasi-experimental study in day care. in l. verschaffel, e. de corte, g. kanselaar, & m. valcke (eds.), powerful environments for promoting deep conceptual and strategic learning (studia paedagogica 41) (pp. 63–80). leuven, the netherlands: leuven university press. hannula-sormunen. m. m., nanu, c. e., laakkonen, e., munck, p., kiuru, n., lehtonen, l. & pipary study group. (2017). early mathematical skill profiles of prematurely and full-term born children. learning and individual differences, 55, 108-119. doi: 10.1016/j.lindif.2017.03.004 hannula, m. m., räsänen, p., & lehtinen, e. (2007). development of counting skills: role of spontaneous focusing on numerosity and subitizing-based enumeration. mathematical thinking and learning, 9(1), 51–57. doi: 10.1207/s15327833mtl0901_4 herbst, m., & strawiński, p. (2015). early effects of an early start: evidence from lowering the school starting age in poland. journal of policy modeling, 38, 256–271. doi: 10.1016/j.jpolmod.2016.01.004 hickendorff, m., edelsbrunner, p. a., mcmullen, j., schneider, m., & trezise, k. (2017). informative tools for characterizing individual differences in learning: latent class, latent profile, and latent transition analysis. learning and individual differences, (november). doi: 10.1016/j.lindif.2017.11.001 hu, l. & bentler, p. (1999). cutoff criteria for fit indices in covariance structure analysis: conventional criteria versus new alternatives. structural equation modeling, 6, 1-55. huttenlocher, j., jordan, n. c., & levine, s. c. (1994). a mental model for early arithmetic. journal of experimental psychology: general, 123(3), 284–296. doi: 10.1037//0096-3445.123.3.284 kavkler, m., tancig, s., magajna, l., aubrey, c., göbel, s. m., watson, s. e., … oeckert. (2000). is early learning really more productive? the effect of school starting age on school and labor market performance. european early childhood education research journal, 8(1659), 45. doi: 10.1177/0956797613516471 kolkman, m., e., kroesbergen, e. h., & leseman, p.p.m, (2013). early numerical development and the role of non-symbolic and symbolic skills. learning and instruction, 25, 95-103. http://dx.doi.org/10.1016/j.learninstruc.2012.12.001 braham, e., libertus, m., mccrink, k. (2018). increasing children’s spontaneous focus on number through guided parent-child interactions in a children’s museum. developmental psychology, 54(8), 1492-1498. doi: 10.1037/dev0000534 locuniak, m. n., & jordan, n. c. (2008). using kindergarten number sense to predict calculation fluency in second grade. journal of learning disabilities, 41(5), 451–459. doi: 10.1177/0022219408321126 luyten, h. (2006). an empirical assessment of the absolute effect of schooling: regression‐discontinuity applied to timss‐95. oxford review of education, 32(3), 397–429. doi: 10.1080/03054980600776589 magidson, j., & vermunt, j. k. (2002). latent class models for clustering: a comparison with k-means. canadian journal of marketing research, 20, 37–44. magnusson, d. (2003). the person approach: concepts, measurement models, and research strategy. new directions for child and adolescent development, 2003(101), 3–23. doi: 10.1002/cd.79 marsh, h. w., lüdtke, o., trautwein, u., & morin, a. j. s. (2009). classical latent profile analysis of academic self-concept dimensions: synergy of personand variable-centered approaches to theoretical models of self-concept. structural equation modeling: a multidisciplinary journal, 16. doi: 10.1080/10705510902751010 martin, r. b., cirino, p. t., sharp, c., & barnes, m. (2014). number and counting skills in kindergarten as predictors of grade 1 mathematical skills. learning and individual differences, 34, 12–23. doi: 10.1016/j.lindif.2014.05.006 masyn, k. e. (2013). latent class analysis with finite mixture modeling. in t. d. little (ed.), the oxford handbook of quantitative methods. new york, ny: oxford university press. melhuish, e., sylva, k., sammons, p., siraj-blatchford, i., taggart, b., phan, m., & malin, a. (2008). preschool influences on mathematics achievement. science, 321, 1161-1162. doi: 10.1126/science.1158808 muthén, l. k., & muthén, b. o. (n.d.). mplus user’s guide (7th ed.). los angeles, ca: muthén & muthén. nanu, c. e., mcmullen, j., munck, p., pipari study group, & hannula-sormunen, m. m. (2018). spontaneous focusing on numerosity in preschool as a predictor of mathematical skills and knowledge in the fifth grade. journal of experimental child psychology, 169, 42-58. doi: 10.1016/j.jecp.2017.12.011 nylund-gibson, k., & masyn, k. e. (2016). covariates and mixture modeling: results of a simulation study exploring the impact of misspecified effects on class enumeration. structural equation modeling: a multidisciplinary journal, 23(6), 782–797. doi: 10.1080/10705511.2016.1221313 nylund, k. l., asparouhov, t., & muthén, b. o. (2007). deciding on the number of classes in latent class analysis and growth mixture modeling: a monte carlo simulation study. structural equation modeling, 14(4), 535–569. doi: 10.1080/10705510701575396 purpura, d. j., baroody, a. j., & lonigan, c. j. (2013). the transition from informal to formal mathematical knowledge: mediation by numeral knowledge. journal of educational psychology, 105(2), 453–464. doi: 10.1037/a0031753 rathé, s., torbeyns, j., hannula-sormunen, m., & verschaffel, l. (2016). kindergartners’ spontaneous focusing on numerosity in relation to their number-related utterances during numerical picture book reading. mathematical thinking and learning, 18(2), 125-141. doi: 10.1080/10986065.2016.1148531 rathé, s., torbeyns, j., de smedt, b., hannula-sormunen, m. m., & verschaffel, l. (2017). verbal and action-based measures of kindergartners’ sfon and their associations with number-related utterances during picture book reading. british journal of educational psychology.. doi: 10.1111/bjep.12201 sprietsma, m. (2010). effect of relative age in the first grade of primary school on long‐term scholastic results: international comparative evidence using pisa 2003. education economics, 18(1), 1–32. doi: 10.1080/09645290802201961 sylva, k., melhuish, e., sammons, p., siraj-blatchford, i. & taggart, b. (2008). effective pre-school and primary education 3-11 project (eppe 3-11)—final report from the primary phase: pre-school, school and family influences on children’s development during key stage, 2(7-11). london, united kingdom: department for children schools and families. van de rijt, b., godfrey, r., aubrey, r., van luit, j. e. h., ghesquière, p., torbeyns, j., … tzouriadou, m. (2003). the development of early numeracy in europe. journal of early childhood research, 1(2), 155–180. doi: 10.1177/1476718x030012002 watts, t. w., clements, d. h., sarama, j., wolfe, c. b., spitler, m. e., & bailey, d. h. (2017). does early mathematics intervention change the processes underlying children’s learning? journal of research on educational effectiveness, 10(1), 96–115. doi: 10.1080/19345747.2016.1204640 welsh, j.a., nix, r.l., blair, c., bierman, k.l., & nelson, k.e. (2010). the development of cognitive skills and gains in academic school readiness for children from low-income families. journal of educational psychology, 102(1), 43-53. doi: 10.1037/a0016738 wolke, d., strauss, v. y. c., johnson, s., gilmore, c., marlow, n., & jaekel, j. (2015). universal gestational age effects on cognitive and basic mathematic processing: 2 cohorts in 2 countries. journal of pediatrics, 166, 1410–1416. doi: 10.1016/j.jpeds.2015.02.065 appendix a fit indices for exploratory factor models of the mathematical skills note. χ² = chi square goodness of fit statistic; df = degrees of freedom; cfi = comparative fit index; tli = tucker lewis index; rmsea = root-mean-square error of approximation; * indicates χ² is statistically significant. summary of exploratory factor loadings codepen van meter publication frontline learning research vol.8 no. 3 special issue (2020) 174 184 issn 2295-3159 commentary: measurement and the study of motivation and strategy use: determining if and when self-report measures are appropriate peggy n. van metera, apennsylvania state university, usa abstract the goal of this special issue is to examine the use of self-report measures in the study of motivation and strategy use. this commentary reviews the articles contained in this special issue to address the primary objective of determining if and when self-report measures contribute to understanding these major constructs involved in self-regulated learning. guided by three central questions, this review highlights some of the major, emergent themes regarding the use of self-report. the issues addressed include attention to evidence for construct validity, the need to consider broad methodological factors in the collection and interpretation of self-report data, and the innovations made possible by modern tools for administering and analyzing self-report measures. conclusions forward a set of conditions for the use of self-report measures, which center on the role of theoretically-driven choices in both the selection of self-report measures and analysis of the data these measures generate. keywords: self-report, self-regulated learning, motivation, strategies, strategy use info corresponding author email: pnv1@psu.edu doi https://doi.org/10.14786/flr.v8i3.631 student learning can be understood through a lens of self-regulation, which explains learning as involving dynamic, cyclic processes that are both selfand goal-directed (zimmerman, 1990). the self-directed component of this definition is central to understanding models of self-regulated learning (srl) because these models posit that learning is influenced by the learner’s own choices and abilities to apply effortful, effective learning processes. as such, the study of motivation and strategy use is a centrepiece in the study of srl. if one is to understand a learner’s choices and abilities, then one must understand the motives and strategic operations on which these rest. one must, that is, be able to answer questions such as, “what factors influence learners’ motivational states?”, “which strategies do learners apply?”, and “how do motivation and strategy use influence learning?”. although each article in this special issue contributes empirical evidence that advances our understanding of motivational and strategic processes and just how we might answer these questions, individually they differ with regard to the specific constructs of interest. vriesema and mccaslin (2020), for example, report on the processes of identity formation in small group learning. moeller et al.(2020) study the emotions and beliefs of interest for class activities while rogiers, et al. (2020) examine how profiles of learning are associated with dynamic strategy use during text learning. altogether then, these articles do give insights into a variety of srl constructs associated with motivation and strategy use and each can be interpreted in the context of the corresponding construct-specific research literature. the focus of this special issue, however, is not on these constructs per se, but rather how these constructs are measured and how they are understood through the lens of that measurement. specifically, the purpose of this special issue is to examine the use of self-report measures and address the challenge of determining “when and if self-report measures can contribute to our collective understanding of theory surrounding constructs.” (fryer & dinsmore, 2020). the articles in this issue represent different ways of answering that call. articles by fryer and nakao (2020), iaconelli and wolters (2020), and chauliac et al. (2020) for example, adopt a measurement approach and focus on factors that can influence the reliability and validity of self-report data. other articles; namely those by durik and jenkins (2020) and moeller et al. (2020); explore methods to enhance the evidentiary value of self-report data. a final grouping of articles sought to establish the need for self-report data by demonstrating the benefits of using these instruments in pursuit of theoretically compelling components of learning. included in this grouping are articles by vriesema and mccaslin (2020), van halem et al. (2020), and rogiers et al. (2020). despite these differences, what unites these articles is shared attention to the set of central questions that drive this special issue. specifically, author teams were tasked with addressing some combination of three questions that concern the utility of self-reports for the study of motivation and strategy use. these central questions ask about (1) the alignment of self-report methodology and theoretical conceptualizations of constructs, (2) the influence of self-report methodology on the interpretation of study results, and (3) the connection between self-report methodology and analytic choices. the articles in this issue present data obtained through particular programs of research and, as such, each article offers some particular view on the answers to these questions. the goal of this commentary is to look across those specifics and offer a more synthetic perspective; one that draws across constructs and methodologies to highlight themes around these questions and draw conclusions about what this body of articles suggests for the future of self-report use. toward that end, my comments will admittedly overlook differences with regard to the specific constructs represented in this set of papers and instead treat each as representative of the set of constructs associated with srl. the remainder of this commentary is organized around the three central questions and addresses some of the major themes that emerged from the articles in this special issue. in what ways do self-report instruments reflect the conceptualization of the constructs suggested in theory related to motivation and strategy use? on the face of it, this is a rather straightforward question about an aspect of construct validity. that is, do the measures align with, and therefore reflect, the theoretical conceptualizations of the constructs (edwards & bagozzi, 2000)? construct validity is critical to the relationship between measurement and theory because it is measurement that provides the operational definition of a construct. whereas theoretical descriptions of a construct may be abstract and difficult to pin down, an operational definition is the specific way in which the construct is measured, including the exact prompts to which participants respond and the ways that data is collected. consequently, if one wants to know what is meant by theoretical terms such as identity formation (vriesema & mccaslin, 2020) or interest (fryer & nakao, 2020), one need only look to how those constructs are operationalized. in this regard, establishing this aspect of construct validity requires three elements (1) a clear articulation of the theoretical conceptualization, (2) a clear articulation around the measurement methodology, and (3) a coherent mapping between the conceptualization and the methodology. efforts toward establishing construct validity can also feed into a cycle of theory and measurement refinement. that is, confidence in the validity of a measure is increased when obtained data behave in theoretically consistent and predictable ways, but innovations in measurement can also reveal evidence of phenomena that stimulate refinement and development of theoretical accounts. the articles in this special issue provide a number of examples of how this form of validity can be established when using self-report measures. most specifically, the authors achieve this by carefully and explicitly defining the constructs of interest in the context of guiding theoretical frameworks, and tying these definitions to the measurement instrument. while several articles provide examples of how this can be done, just two will be presented as illustrations here. first is the study by rogiers et al. (2020), in which they examined the text learning strategies of middle school students. the purpose of this study was to “fully map and understand individual students’ learning” (p. 1) using the research context of students studying to learn from an expository text to engage in this mapping. srl is the theoretical framework that guides this research and, consistent with the definitions used throughout this issue, rogiers et al. defined srl as involving adaptive, flexible strategy use in dynamic, iterative phases. further, rogiers et al. stated that there are individual differences in how learners employ strategies and in their perceptions of this strategy use. most central to the theoretical conceptualization, rogiers et al. also reasoned that these individual differences could provide insight into the dynamic, adaptive ways that learners employ strategies during learning. their use of two different self-report measures follows from this conceptualization. first, participants thought aloud while engaged in the text learning task with the resulting protocols revealing of the dynamic strategic processes employed during the task. second, after reading, participants completed a self-report survey, which queried the task-specific cognitive and metacognitive strategies used during study. this survey measure identified meaningful individual differences and served to group participants into different profiles of strategy use (e.g., integrated strategy user). the value of both forms of self-report data was realized by using the profiles of strategy use to guide interpretation of think aloud data. in brief, consistent with theoretical conceptualizations, rogiers et al. were able to use self-report measures to demonstrate that different types of strategy users employ dynamic srl processes in different ways. a second example of how articles demonstrate the connection between conceptualizations of a construct and self-report measures of that construct can be found in moeller et al.’s (2020) study of situational interest in a college course. this article defined interest as a motivational and emotional state that fluctuates over time, and measured these fluctuations as situational expectancy and task value (i.e., expectancy-value; eccles, & wigfield, 2002). moreover, the authors argued that, at any point in time, these states are a function of (1) stable personal traits, (2) situational personal perception, and (3) objective components of the situation. in order to follow from this theoretical conception then, measures of interest must capture and distinguish all three sources of this variance. the use of a self-report interest survey, which was administered periodically during class, was a logical choice in this context because survey responses allow the capture of individuals’ perceptions. it was the manner in which moeller et al. employed the survey, however, that permitted the strong connection between the theoretical conceptualization of situational interest and the self-report measure. while the reader is referred to that article for a full explanation of the methodology, a short summary here will suffice: course students completed the survey at multiple time points with multiple students intentionally sampled at each time point. this sampling pattern then permitted examination of both objective evaluations (i.e., group means) and personal perceptions (i.e., deviation from the mean). ultimately, the use of the self-report measure was validated when moeller et al. were able to parse the variance in individuals’ time-point interest reports into the three theoretically predicted sources of variance. in sum, the articles in this special issue demonstrate that self-report measures can not only reflect conceptualizations of constructs, they can do so in theoretically compelling ways. these efforts toward construct validity feed the mutually reinforcing cycle of theory development and methodological refinements. rogiers et al.’s finding of relations between profiles based on learners’ perceptions of strategy use and their dynamic application of those strategies, that is, furthers understanding of individual differences and srl. at the same time, moeller et al.’s study advances theory regarding the personal and objective sources of situational interest; a motivational construct central to understanding srl. in the context of this special issue, however, where the challenge is to determine when and if self-report data contributes to the understanding of constructs, there is another layer to the question of how self-report reflects theoretical conceptualizations. specifically, in this context, it is not sufficient to show that some measurement choice is consistent with theoretical definitions or even that the self-report data accounts for some theoretically interesting variance. instead, this task calls on us to consider when and if self-report data provides insight into some phenomenon that is not gained by another measurement approach. in other words, we are challenged to show not only that self-report measures can reflect conceptualizations of motivation and strategy use, but also that some self-report methodology is uniquely suited to doing so. a partial response to this challenge can be obtained by pointing back to the constructs of interest. specifically, when the construct of interest is a learner’s perception of intra-psychic states (e.g., beliefs, motivations), then it is sensible to suggest that the best way to uncover these perceptions is to ask the learner (fryer & nakao, 2020). in addition to this argument, however, articles in this special issue lay out an even more convincing reason for using self-report measures. namely, self-reports are a justifiable measurement tool because data from these measures offer unique explanatory power when it comes to understanding motivation and strategy use. again, two studies from this special issue can be used to illustrate this point. the first example is the study by vriesema and mccaslin (2020), which sought to understand the processes of identity formation in small group settings. guided by a co-regulation theoretical model, the authors collected self-report data on students’ perceptions of how they engaged with the members of their group as well as preand post anxiety and emotional adaption profiles. observational data of group interactions was also collected and analyzed to show the actual interaction pattern that took place in the groups over a series of six lessons. the analysis of this data demonstrates that more is learned about identity formation and co-regulation from both self-report and observational data than from either source alone. while pre-group self-reports of emotional adaptation were predictive of some co-regulation styles, for example, certain co-regulation styles were predictive of post-group emotional adaptation profiles. the value of self-report methodology is also demonstrated in the study by van halem et al. (2020), which shows that data obtained from these measures offers unique explanatory insights. in this study, trace data was collected over a period of eight weeks as students in a college statistics course accessed online course resources. using the theoretical framework of srl, the authors point out that these trace data provide insights into behavioral aspects of learning; these traces are “observable evidence of particular cognitions…where a cognitive process is applied” (p. 3) at the same time, however, these traces do not indicate just what those cognitive processes are. one student, for example, may access some resource because it covers content from a class that was missed while another student may access that same resource because they did not understand the content when it was covered in class. to gain insight into the processes underlying these behavioral traces, van halem et al. had participants complete a self-report survey of srl behaviors (i.e., motivated strategies for learning questionnaire; mslq; pintrich et al. 1993) in the fourth week of the course. at the end of the course, analyses showed that, when both trace and self-report data were included, some mslq sub-scales accounted for variance in grades above and beyond that accounted for by the behavioral data. in short, like the research on identity formation in small groups, this study shows that a self-report measure can explain important aspects of motivation and strategy use that would not be captured in the absence of the measure. in summary, the research teams represented in this special issue demonstrate three specific ways in which self-report instruments reflect theoretical conceptualizations of motivation and strategy use. first, across the set of articles, one can see that self-report instruments and methodologies can operationalize constructs in theoretically consistent ways. second, these measures can generate data that not only behaves in theoretically predictable ways, but also offers refinements to the conceptualization of constructs. third, self-report measures can reflect conceptualizations by capturing patterns and variance in motivation and strategy use that are not obtained through other means. while these answers justify the use of self-report from a conceptual standpoint, they cannot be completely disentangled from more specific methodological choices associated with the use of self-report. the two remaining questions posed by this special issue provide the opportunity to address some of these points. how do the interpretations of self-report data influence interpretations of study findings? this second question, which asks how the interpretations of self-report data influence the interpretations of study findings, is similar to the first question in that it can be understood as addressing an aspect of validity. namely, validity is not determined by some measure itself but rather by the degree to which the interpretations drawn from the scores on that measure can be justified (messick, 1995). in this respect, the interpretations of a study’s findings are valid when the data sources on which those findings are based have been interpreted in valid ways. this logical argument then calls for a particular view on the question that frames this section: to understand how self-report data influences the interpretation of study findings, we must understand the factors that influence the reliability and validity of the data derived from these measures. also similar to the previous section, there are two different perspectives we can take on this question. the first perspective concerns the factors that may influence the reliability and consequently, the validity of scores from self-report measures. this perspective is primarily concerned with potential sources of error in self-report measurement of motivation and strategy use. the second perspective takes a more conceptual view and considers the ways in which the methodologies of collecting self-report data influence the interpretations of that data. this perspective draws attention to the broader theoretical and contextual factors that influence how scores can be interpreted. with respect to the first perspective, several studies in this special issue examine sources of error in self-reports and how those sources can be understood or reduced. one potential source of error, which is examined in the study by fryer and nakao (2020), is the format of the response scales and interfaces used to record participant’s responses. this examination was prompted by the body of work on survey instruments, which suggest that the response scales themselves can impact the nature and reliability of scores. participants in this study were graduate students enrolled in a course on teaching and, throughout the course, these participants responded to surveys assessing their interest in class activities. to examine response formats as a potential source of measurement error, this study had participants complete self-report surveys that asked the same questions, but used four different interfaces: labelled categorical scales, visual analog scales (vas), swipe, and slider. these surveys were administered at six time points throughout the semester so that all participants responded using each of the interfaces and, at any one time point, all four interfaces were used. this design permitted comparisons across the different interfaces to determine if any significant differences in response patterns could be tied to differences in the interface. on the whole, the results suggest that response interfaces are not a significant source of error. each of the measurement methods yielded acceptable levels of reliability and there were no differences in either the mean scores across the measures within the six time points or differences in the factor structures of the measures. although details in the findings lead fryer and nakao (2020) to suggest that the swipe method shows promise and the vas method is the weakest, the totality of the data indicates that scores obtained from each of the response formats and interfaces can be validly interpreted. another potential source of error, one that has been suggested throughout the history of self-report surveys, is insufficient effort on the part of respondents. according to this view, the results of self-report surveys are tainted by participants who either do not put forth the cognitive effort to answer survey questions or bias the results by responding in unserious ways. two articles in this special issue address this concern by examining data related to participants’ survey response patterns. first, as part of a larger study, chauliac et al. (2020) collected eye tracking data while college students responded to a task-specific survey on processing strategies (i.e., inventory of learning styles; vermunt & donche, 2017). the time and frequency of eye fixations on any given question were interpreted as indicators of effort while the consistency of an individual’s within-scale responses were taken as an indicator of within-person reliability. analyses showed a relationship between fixations and response variability wherein participants who spent the most time on an item were also most likely to show only small degrees of variation in item responses. in other words, these participants showed patterns indicative of reliable responding. by contrast, participants who spent the least amount of time were most likely to either select the same categorical response for each scale item or show extreme variability; i.e., poor reliability. chauliac et al.’s (2020) finding are complimented by the research of iaconelli and wolters (2020), which also examined indicators of insufficient effort responding, but extends this work by presenting techniques to detect these participants in large scale data collections. this study involved nearly 300 college students in a course designed to improve their srl and, at three different times in the course, all participants completed self-report surveys tapping into their dispositions, beliefs, and behaviors. consistent with chauliac et al. (2020), iaconelli and wolters posit that the validity of a measure is threatened if respondents exert too little effort while answering questions. further, they posit that these response patterns can be detected by examining indicators of effort (i.e., time) and consistency in response patterns. while the reader is referred to the article itself for details on these indicators, there are three main conclusions relevant here. first, there are some participants who show insufficient effort. but, two, these participants comprise only a small percentage of the total sample and their inclusion, at least in a large data set, does not significantly alter either the mean or the internal consistency of the data set. third, although each of the three surveys had some participants who gave insufficient effort, this did not appear to be a stable individual difference. instead, if a participant did exhibit insufficient effort, this tended to occur on only one of the three surveys. as summarized here, the articles in this issue report evidence that self-report surveys can and do provide reliable indicators of variables associated with motivation and strategy use. neither the response format nor a lack of respondent effort introduced sufficient error variance to question the interpretations that are drawn from these instruments. under these conditions then, we can conclude that study findings based on these self-report measures can be interpreted in the intended ways. the second perspective on this question of how self-reports influence study findings, however, encourages us to look at a broader set of factors that influence the validity of self-report data interpretations. these broader factors include the totality of the context in which the measure is administered including theoretically-motivated methodological factors. to illustrate this, consider the article by durik and jenkins (2020), which ultimately concludes that the person-domain context must be taken into account when interpreting the results from self-report surveys of learner interest. in a pilot study and two experiments, college students responded to self-report surveys that assessed their interest in different domains (e.g., math and psychology) and also indicated the likelihood that they would pursue future learning in these domains (study 1 and 2). the purpose of this research was to explore how self-report can be used to better explain the relationship between interest and behavior and, toward that goal, durik and jenkins had participants also respond to questions gauging their certainty in provided interest ratings. factor analyses showed that interest and certainty comprise separate factors, indicating that respondents are able to distinguish these two beliefs. analyses also showed that the level of certainty moderated the relationship between reported interest and future behavior with the interest-behavior relationship markedly stronger for participants with high levels of certainty. altogether, the data presented in this article shows that, at least in the study of interest, (1) participants’ ability to provide accurate, predictive self-reports depends on how certain they are about these reports and (2) certainty varies with exposure to the domain. considering this in light of the question of how self-report influences the interpretation of study findings, this research highlights the need to attend to the broader context in which the measure is administered; in this case, the context of the person-domain relationship. another methodological factor that emerged as important to the interpretations of study results is the timing of self-report administration. although there are different types of self-reports possible, each requires participants to respond to some query on the basis of their memory for the relevant information (see the articles by chauliac et al. and iaconelli & wolters, this issue for a discussion of these models) and each is administered prospectively, concurrently, or retrospectively. in this respect, self-report responses provide a snapshot in time (durik and jenkins, 2020). yet, because effective learning processes are understood as dynamic, flexible, and adaptive; a challenge to the use of self-report data is the need to show how a snapshot can shed light on active motivational and strategic operations. this special issue presents one possible answer to this: researchers can enhance the validity of interpretations by attending to the timing in which self-report measures are administered and incorporating this timing into the interpretation of study findings. to illustrate this point, consider the study by van halem et al. (2020) in which trace data was collected from students as they accessed online materials throughout an eight-week statistics course. participants also completed a self-report measure of srl in the fourth week (i.e., mslq). as described previously, study results showed that both the trace and self-report data accounted for variance in students’ final course grades. additionally, however, van halem et al., also examined relations between trace and self-report data for each of the eight course weeks and found that the relationships were the strongest in the weeks preceding completion of the self-report survey and weak in the periods thereafter. in short, the timing of the self-report measure influenced the nature of the resulting data and thus, must be incorporated into the interpretation of study findings: self-report can be an accurate snapshot of students’ memories for what they have done in a course but are not necessarily prognosticators of future behavior, at least not with respect to srl as measured by the mslq. the influence of timing in the administration of self-report measures is also demonstrated in the study by rogiers et al. (2020). as explained in the previous section, middle school students in this study thought aloud while reading expository text and, immediately after reading, completed a self-report survey of the strategies used. in this respect, both concurrent and retrospective self-report measures are used with the retrospective survey placed close in time and in direct reference to the just-completed srl event. again, this timing influences how the data can be interpreted. first, because the survey immediately followed the srl event, results can be interpreted as valid representations of learners’ perceptions of their strategy use and; second, concurrent think alouds reveal the pattern in which strategies were used. finally, rogiers et al. were able to use the profiles that emerged from retrospective self-reports to guide data mining and uncover differences in how individuals deploy srl processes. the timing matters here because it is the time-based relationship of the two self-report measures that permit the data and study findings to be interpreted in this way. the two methodological factors covered here, person-domain relations and timing, are just two of the contextual factors addressed in this special issue that should be considered when interpreting study findings. vriesema and mccaslin’s research on identity formation in groups, for example, demonstrates that this development must be understood in the context of the specific group’s dynamics; i.e., individuals are nested in groups. exactly how data is collected should also be taken into consideration. iaconelli and wolters (2020) show this in their examination of insufficient effort responding. recall these authors found that, while insufficient effort responding does occur, these occurrences have a negligible effect on the data set. these authors, however, were careful to point out that the surveys were completed as part of homework assignments in participants’ course on srl. as a result, it is possible that insufficient effort responding was infrequent in this study because participants had a high degree of investment. higher, that is, than one might expect from study participants who complete a survey only to receive course extra credit for study participation (e.g., durik & jenkins, 2020). taken as a whole, the current set of articles provide at least two important insights into the ways that the interpretation of self-report data can and should be used to interpret study findings. first, self-report data sources can be trusted to provide reliable and valid indicators of studied constructs. although there is some error in these measures, evidence culled from these studies provide confidence that this is no greater a problem for self-report measures than other types of data collection methods that rely on human responses. second, the broader contextual and methodological factors of measurement administration matter. although the studies reported here shed light on some of these factors, no doubt there are many more that warrant attention. as a summary though, one can conclude that the methodology around the administration of self-report measures influences the interpretation of resulting data and consequently, influences the interpretation of study findings. how does the use of self-report constrain the analytical choices made with that self-report data? this final question can be understood to specifically address data analytic concerns rather than issues related to construct conceptualizations and the interpretation of findings covered in the first two questions. with these boundaries in mind, the short answer to this question is that the self-report nature of this data does not place constraints on analytic choices above and beyond what must be considered with other data sources; constraints such as scales of measurement, distributions, and floor or ceiling effects. indeed, what is most striking in relation to this question are the creative and innovative ways in which self-report data can be collected and analyzed. of course, it has long been argued that a chief benefit of self-report data is the ability to collect data from large sample sizes and this benefit is only increasing with technological advances in digital delivery systems (fryer & nakao, 2020). beyond this ease, however, the articles in this issue highlight two valuable connections between the use of self-report data and subsequent analytic choices. the first connection that emerged is how advances in both measurement delivery and statistical analytic tools permit self-report data to be collected and analyzed in increasingly creative and sophisticated ways. as authors in these special issue articles have noted, self-report measures have traditionally been delivered in paper-and-pencil form and resulting data treated in aggregated, variable-centered ways: group means attest to some agreed upon (i.e., averaged) descriptor of a construct and scores are treated to traditional forms of comparisons and correlations. by contrast, today’s researcher has access to much more sophisticated tools to deliver measures as well as to parse variance and model data patterns. consider, for example, the study by moeller et al. (2020) that examined both personal and objective contributions to fluctuating states of interest. thus far, this commentary has described this research in terms of the studied construct and empirical findings. an examination of the study methods, however, illustrates how innovation in the delivery and analysis of self-reports is expanding our understanding of motivation and strategy use in srl. specifically, these authors leveraged online delivery mechanisms and innovative experience sampling methods to collect the self-report data from groups of participants in context and over time. once collected, the application of cross-classified multilevel modelling permitted participants’ time-point self-report data to be parsed into the three sources of variance that were predicted by the theoretical framework of situational interest used in this study. in short, moeller et al. were able to apply modern research tools to the collection and analysis of data in a manner that advances theoretical understanding. moeller et al.’s work is not the only illustration of the ways that self-report data can be meaningfully analyzed given the tools currently available. the studies by both iaconelli and wolters (2020) and fryer and nakao (2020), for example, show how the online delivery of self-report measures can yield not only participants’ responses but also indicators of invested effort (i.e., time). others took advantage of recent techniques to detect patterns within data sets and used these methods to mine for dynamic, iterative, srl cycles (rogiers et al., 2020); classify participants according to profiles of individual differences and group co-regulation dynamics (vriesema & mccaslin, 2020); and disentangle the relationships between interest, certainty, and behavior (durik and jenkins, 2020). altogether, this special issue shows that the choice to use self-report data opens the door to a great number of analytic choices, but the most promising of these may be the use of self-report data alongside other, complimentary data sources. as previously summarized, self-report data can have unique explanatory power when combined with other data sources in the study of motivation and strategy use. that previous discussion, however, was narrowly focused on how self-report reflects conceptualizations of theoretical constructs, and did not address methodological and analytic dimensions of this point. with respect to the current question though, one can see significant potential in the use of self-report measures in conjunction with additional measures of motivation and strategy use. for instance, self-report measures can play an important role in mixed methods srl research in which qualitative process data can be combined with quantitative scores derived from self-report surveys. just such an approach is demonstrated in the studies by both rogiers et al. (2020) and vriesema and mccaslin (2020) in which qualitative process data was collected and coded in addition to the administration of self-report surveys. ultimately, these data sources were combined to shed light on how self-reported individual differences related to srl processes in either individual (rogiers et al., 2020) or group (vriesema & mccaslin, 2020) settings. in addition to these two studies, articles in this special issue show other ways of combining self-report survey responses with additional forms of process data such as eye fixations (chauliac et al., 2020), trace data (van halem et al., 2020), and response times (iaconelli & wolters, 2020). in sum, this section can be closed by returning to the answer offered at the opening; namely, the self-report nature of some data does not place constraints on analytic choices above and beyond what must be considered with other data sources. instead, what does constrain both the choice of measures and data analysis methods, are the theoretically-based conceptualizations of the construct and the questions that drive the research. as the articles in this issue show, self-report data can be analyzed in a wide variety of ways with innovations paving the way for breakthroughs in both measurement and theory. certainly, one must be concerned with the psychometric properties of scores and the match between the data set and the assumptions of a particular analysis. beyond these constraints, however, self-report data has, and can continue to be, analyzed in ways that yield relevant insights into the individual differences and processes of motivation and strategy use. conclusion and final remarks the articles in this special issue shed light on a number of theoretical constructs associated with motivation and strategy use, but the main objective of this collection is to examine the self-report methodology used to study these constructs. the task for the articles in this issue, including this commentary, was to use three organizing questions to “determine when and if” (fryer & dinsmore, 2020) self-report measures positively contribute to the study of theoretical srl constructs. this final conclusion section focuses on this task by considering first, the question of “if” self-report measures can contribute followed by the question of when this might be true. the question of “if” self-report measures can be used calls for answers to two relatively straightforward questions: (1) is there evidence that scores on self-report measures can be reliable and valid indicators of motivational and strategic constructs? and (2) is there evidence that self-report measures provide explanatory power in the study of motivational and strategic constructs? across all of the articles in this special issue, the answer to both of these questions is, “yes”. one bit of evidence in support of this affirmative response is found in demonstrations that self-report measures yield scores with acceptable reliability and adequate psychometric properties. iaconelli and wolters (2020), for example, showed that insufficient effort responding had little impact on a full data set and moeller et al. (2020) showed how theoretically-driven analysis can increase the amount of variance explained in self-report data. additional evidence from this set of articles comes from the repeated demonstrations that self-report measures play an important role in capturing and understanding theoretical constructs. in short, this body of research advances our understanding of motivation and strategy use and this is due, in large part, to the use of self-report measures. from the data presented in these studies, we learned about srl phenomena such as situational fluctuations in motivational states, the relationship between group dynamics and identity formation, individual differences in the dynamic application of strategy use, and the role of domain exposure and certainty in understanding the influence of interest on behavior. in short, the research in this special issue supports the conclusion that self-report measures do indeed have an important role to play in the study of motivation and strategy use. and, this is true whether one is focused specifically on psychometric measurement properties or theoretically-driven conceptualizations. the second part of our task, the task of determining “when” self-report measures contribute to the study of motivation and strategy use, raises questions about the conditions under which self-report may or may not be appropriate. indeed, the articles in this special issue raised concerns about several of the limitations of self-report measures. for example, because self-report measures capture a snapshot of a learner’s perceptions, these instruments may be better at explaining a learner’s past than predicting that learner’s future (e.g., van halem et al., 2020). likewise, when self-reports are in the form of surveys, they capture variance associated with motivation and strategy use, but do not effectively capture dynamic aspects of srl (e.g., vriesema & mccaslin, 2020). finally, like any other measure, self-reports are not immune to potential sources of error such as individual differences (moeller et al., 2020; durik and jenkins, 2020) and insufficient effort responding (e.g., chauliac et al., 2020). these limitations notwithstanding, it is possible to draw conclusions about when self-reports are likely to advance the study of srl constructs. specifically, there are three conditions under which self-report measures can be effectively used: (1) when the measure aligns with theoretically-driven conceptualizations of the construct, (2) when measure selection is driven by alignment with theoretically-driven research questions, and (3) when measure administration, data analysis, and results interpretations are grounded in theoretically-driven choices. in sum, self-report measures can contribute to the study of motivation and strategy use when a close coupling of the measure and relevant theory allows for a mutually beneficial cycle of refinement and development. with these recommendations in mind, i will close with one final thought that points out the primary weakness of this commentary; namely, the choice to synthesize across the specific constructs studied in each article and group them under the broad umbrella of srl. while this choice permitted general conclusions to be drawn about the use of self-report in the study of motivation and strategy use, it also meant that attention was not paid to possible construct-measurement interactions. that is to say, interactions in which the conditions for how and when self-report measures are best used vary according to the construct under study. reading this set of papers, for example, raises a number of interesting questions about these possible interactions; e.g., whether the degree of certainty influences the prognostic abilities of self-reported strategy use in the same way that it influences measures of interest, if the rates of insufficient effort responding are consistent across srl constructs, how the methods used to evaluate the effects of classroom activities on interest could be used to evaluate how those activities stimulate strategic processes. despite the lack of attention given here to possible construct-measurement interactions such as these, their exploration offers a direction for future research. carrying out this work has the potential to shed light on not only the use of self-report measurement tools, but also the theoretical conceptualizations in which they are grounded. keypoints self-report measures can accurately and constructively reflect theoretical conceptualizations of srl constructs. data derived from self-report measures can provide reliable and valid indicators of motivation and strategy use. the application of modern research tools to the use of self-reports can lead to breakthroughs in both srl measurement and theory. self-report is most effectively used when it is closely aligned with theory. references chauliac, m., catrysse, l., gijbels, d., & donche, v. (2020). it is all in the surv-eye: can eye tracking data shed light on the internal consistency in self-report questionnaires on cognitive processing strategies? frontline learning research, 8(3), 26–39. https://doi.org/10.14786/flr.v8i3.489 durik, a., & jenkins j. (2020). variability in certainty of self-reported interest: implications for theory and research. frontline learning research, 8(3), 86–104. https://doi.org/10.14786/flr.v8i3.49 eccles, j. s., & wigfield, a. (2002). motivational beliefs, values, and goals. annual review of psychology, 53(1), 109-132. https://doi.org/10.1146/annurev.psych.53.100901.135153 edwards, j. r., & bagozzi, r. p. (2000). on the nature and direction of relationships between constructs and measures. psychological methods, 5(2), 155. https://doi.org/10.1037/1082-989x.5.2.155 fryer, l. k., & dinsmore, d. l. (2020). the promise and pitfalls of self-report: development, research design and analysis issues, and multiple methods. frontline learning research, 8(3), 1–9. https://doi.org/10.14786/flr.v8i3.623 fryer, l. k., & nakao, k. (2020). the future of survey self-report: an experiment contrasting likert, vas, slide, and swipe touch interfaces. frontline learning research, 8(3), 10–25. https://doi.org/10.14786/flr.v8i3.501 iaconelli, r., & wolters, c. a. (2020). insufficient effort responding in surveys assessing self-regulated learning: nuisance or fatal flaw?frontline learning research, 8(3), 105–127. https://doi.org/10.14786/flr.v8i3.521 messick, s. (1995). validity of psychological assessment: validation of inferences from persons' responses and performances as scientific inquiry into score meaning. american psychologist, 50(9), 741. https://doi.org/10.1037/0003-066x.50.9.741 moeller, j., dietrich, j., viljaranta, j., & kracke, b. (2020). disentangling objective characteristics of learning situations from subjective perceptions thereof, using an experience sampling method design. frontline learning research, 8(3), 63–85. https://doi.org/10.14786/flr.v8i3.529 pintrich, p. r., smith, d. a., garcia, t., & mckeachie, w. j. (1993). reliability and predictive validity of the motivated strategies for learning questionnaire (mslq). educational and psychological measurement, 53(3), 801-813. https://doi.org/10.1177/0013164493053003024 rogiers, a.; merchie, e., & van keer, h. (2020). opening the black box of students’ text-learning processes: a process mining perspective. frontline learning research, 8(3), 40–62. https://doi.org/10.14786/flr.v8i3.527 van halem, n., van klaveren, c. p. b. j., drachsler, h., schmitz, m., & cornelisz, i. (2020). tracking patterns in self-regulated learning using students’ self-reports and online trace data. frontline learning research, 8(3), 142-164. https://doi.org/10.14786/flr.v8i3.497 vermunt, j. d., & donche, v. (2017). a learning patterns perspective on student learning in higher education: state of the art and moving forward. educational psychology review, 29(2), 269-299. https://doi.org/10.1007/s10648-017-9414-6 vriesema, c. c., & mccaslin, m. (2020) experience and meaning in small-group contexts: fusing observational and self-report data to capture self and other dynamics. frontline learning research, 8 (3), 128–141. https://doi.org/10.14786/flr.v8i3.493 zimmerman, b. j. (1990) self-regulated learning and academic achievement: an overview, educational psychologist, 25(1), 3-17, https://doi.org/10.1207/s15326985ep2501_2 stahl publication frontline learning research vol.7 no. 3 (2019) 27 63 issn 2295-3159 epistemic beliefs and googling tore ståhl a aarcada university of applied sciences helsinki; university of tampere, finland. article received 26 september 2018 / revised 4 april / accepted 5 july / available online 18 july abstract with the introduction of internet as a source of information, parents have observed youngsters’ tendency to prefer internet as a source, and almost a reluctance to learn in advance since “you can look it up when needed”. questions arise, such as ‘are these phenomena symptoms of changing beliefs about knowledge and learning? is it at all possible to learn on a deeper level simply by looking up the basic facts, without memorizing them?’ within an existing line of investigation, epistemic beliefs have been described as a set of dimensions. although internet-based information and internet as a source of information have been acknowledged, studies so far have not explored how dealing with internet-based information relates to other epistemic beliefs dimensions. to capture how users view internet-based information per se but also in relation to other epistemic beliefs, i suggest three new dimensions, out of which the most crucial is labelled ‘internet reliance’. offloading memory using memory aids is not a new phenomenon but the ‘internet reliance’ dimension indicates that especially internet-reliant users may be confusing external information with personal knowledge, with all the risks it may entail. besides including beliefs about learning, this study also challenges earlier assumptions regarding uncorrelated dimensions. keywords: epistemic beliefs; internet; constructivism; outsourcing knowledge; factor analysis info corresponding author: tore.stahl@arcada.fi doi: 10.14786/flr.v7i3.417 1. introduction and aim of study during the last decade, most people will have heard youngsters respond to a question with the acronyms jfgi or giyf (“just f…g google it” and “google is your friend”, see https://en.wiktionary.org/wiki/jfgi). for most adults, expecting a proper answer, this response was surprising, puzzling and perhaps even offensive. the response is, however, an illustration of the gap between the parent generation’s “you should know this”-view on knowledge, and the young generation’s stance “i’ll look it up when i need it”. with the introduction of easy and ubiquitous access to information over internet, the attitude of looking it up when one needs it became common, especially among frequent internet-users. given that the young generation born after the mid 1980’s grew up surrounded by information and communications technologies (hereafter ict), the interesting question is, has the easy and ubiquitous access to information actually influenced their view on knowledge, knowing and learning? during the first decade of this millennium, the so-called digital natives of the net generation were supposed to hold characteristics such as being constantly on-line, being ict savvy and being at home on social media (e.g. prensky, 2001; siemens, 2005). indeed, the youngsters differ from their parent generation in that they lack a personal history of the time before mobile phones, internet and search engines (gunter, rowlands, & nicholas, 2009, p. 3), not to mention smart phones. large parts of the youngsters within this cohort embrace the opportunities provided by ict, e.g. preferring internet-based information instead of books (cf. osf, 2010; purcell et al., 2012, p. 4). still, several studies have pointed out the heterogeneity within the generation (cf. jones & hosein, 2010; van den beemt, akkerman, & simons, 2011). also among the students participating in the present study, large differences occurred regarding both self-reported ict and media use patterns and performance-based ict skills (ståhl, 2017). within education, the easy and ubiquitous access to information raises concerns about how and upon which information students build their knowledge, since they seem to accept the veracity of on-line information too easily, and lack the skills of thinking critically and synthesizing the information found on-line (purcell et al., 2012, pp. 26-27). the vast popularity of search engines (with covert operating logics) in combination with users’ lacking critique has considerable epistemic implications, as demonstrated in the theoretical work and the studies cited below (section knowledge and information in the internet era). the present study will build upon the above studies that confirm the existence of the jfgi phenomenon. existing self-report instruments for measuring epistemic beliefs are not capable of capturing signs indicating internet-induced changes in the views of knowledge and learning. especially the digital natives’ ways of dealing with knowledge and learning have been described in literature (some examples in section hypothesized dimensions) but so far, this topic has been scarcely approached from an epistemic point of view. this topic calls for empirical investigation, which requires instruments. this paper will describe how the existing dimensions (structure and certainty of knowledge, innate learning ability and omniscient authority) are extended with the new dimensions constructivist approach, internet reliance and learning by dialogue. creating a validated instrument requires more than one round and therefore, the aim of this endeavour is an initial exploration of how new dimensions might contribute to a better description of how today’s higher education learners in an internet-saturated context view knowledge and learning. contemporary research regarding epistemic beliefs largely subscribes to epistemic beliefs being limited to beliefs about knowledge, and not about learning. the present study will deviate from this view by exploring also views about learning. doing so, this study contributes to the discussion by looking beyond the knowledge dimensions of epistemic beliefs, and by describing the connection between beliefs about knowledge and beliefs about learning, a connection that is necessary to illuminate consequences for educational practice. 2. personal knowledge, external information to provide a rationale for the present study, this section will 1) review some studies regarding knowledge, information and epistemic beliefs in the internet era, 2) review epistemic beliefs as a research area, 3) review some arguments regarding learning as part of epistemic beliefs, and 4) discuss why domain specificity and justification of knowledge where omitted from the study at this stage. 2.1 knowledge and information in the internet era george siemens tried to grasp the impact of technology and the decreasing half-life of knowledge by introducing connectivism as a new learning theory for the digital age. he suggested supplementing the existing forms of propositional (knowing-that) and procedural (knowing-how) knowledge with ‘knowing-where’ and ‘knowing-who’, i.e. an understanding of where to find knowledge. according to siemens, since we cannot experience everything or store all knowledge ourselves, we store knowledge in other people and in non-human appliances. the key is connectedness, and the knowledge is distributed (downes, 2007, p. 84; siemens, 2005). connectivism was apparently neither a learning nor a knowledge theory but rather a pedagogical view but still, the connectivist ideas resemble the concept of distributed mind, which suggests that knowledge can reside in people, in tools, and in cultural settings, and that the potential lies in the combination of those (cf. shaffer & clinton, 2006). the results of an experimental study by sparrow and her team suggest that internet has become a kind of extension to our individual memory system. if the net is available, we do not bother to memorize the information itself but rather, where to find the information, as when youngsters respond: “jfgi!” we are becoming increasingly symbiotic with our computer-based tools, growing into interconnected systems that remember less by knowing information than by knowing where to find the information. (sparrow, liu, & wegner, 2011) the concept of the extended mind (clark & chalmers, 1998) suggests that human cognition may extend beyond the brain and include elements from social and technological environments (cf. siemens, 2005). applying the concept to the context of the web opens up for the concept of the web-extended mind, which includes the idea that “… the informational and technological elements of the web can, at least on occasion, constitute part of the material supervenience base for (at least some of) a human agent’s mental states and processes” (smart, 2012, p. 451). the mere existence of the web does not automatically make it part of a person’s extended mind but in addition, three criteria need to be met: the availability criterion, the trust criterion and the accessibility criterion (clark & chalmers, 1998; smart, 2012). considering the development, that has taken place within the web and smart phone contexts since smart wrote his article, we have reason to suspect that users often regard these criteria as met, and too easily incorporate on-line information into their personal body of knowledge: due to internet capable smartphones, the availability and the accessibility criteria are easily met. the problematic part is the trust criterion: on-line information is too easily endorsed and too rarely subject to critical scrutiny (purcell, brenner, & rainie, 2012, pp. 10-11). this is especially problematic since e.g. google made personalized search in 2009 the default option for all users (simpson, 2012, p. 437). the personalization of search results performed by search engines means that the results are tailored to what will probably interest the enquirer, and that those hits that do not fit the enquirer’s profile are ranked down or even omitted. according to thomas simpson (2012), the epistemic significance of search engines lies in their acting as surrogate experts, firstly as they assist the enquirer in finding sources and secondly as they orient the enquirer to supposedly relevant sources of information (the expert role also discussed by fisher, goddu, & keil, 2015, below). the problematic aspect here is that by filtering and ranking the results, the search engine implies a judgment about what is relevant, without the enquirer having neither insight into, nor the possibility to influence the criteria for judgement. as simpson (2012, p. 427) puts it: “… objectivity may require telling enquirers what they do not want to hear, or are not immediately interested in” (my emphasis) (also see hinman, 2008). therefore, simpson regards personalization as an actual threat to objectivity. by leaving out relevant voices, the tailored search results contribute to an epistemic bubble, and the operating logics of search engines combined with the enquirers’ ignorance increases the risk of the enquirer being trapped in an epistemic bubble or even an echo chamber (nguyen, 2018). the complexity of the objectivity problem is illustrated by the findings of purcell, brenner, & rainie (2012): although a majority in their study disapproved search engines collecting information about their searches, 23-29% thought that using the information for personalizing search results was a positive feature (pp. 19-21). further, on average two thirds of the participants believed that the information provided by search engines was fair and unbiased: the younger, the more they relied on search engines’ objectivity (pp. 10-11). a further aspect, illustrating the objectivity problem, is the ritualization described by bhatt & mackenzie (2019), i.e. students’ information seeking practices being largely motivated by adhering to what they call the rules of the game. these rules can be appropriate in the beginning to induce students to the knowledge creation practices of the discipline but when detained too long, they may inhibit the development of students’ information seeking skills and trust in their own capacity to consider the justification of the information they find. in an experimental study, fisher et al. (2015) highlight the risks embedded in ubiquitous access to information, which may blur the boundaries between personal knowledge and external information, thus creating an illusion of possessing personal understanding. further, their results suggest that some individuals tend to regard internet as an expert regardless of domain. these results pose a true challenge for education at all levels, at least if we consider personal and integrated knowledge, instead of loose bits of information, as the objective of education and learning. miller & record (2013) discuss the covert operating logics of search engines and their epistemic implications using a framework building upon a responsibilist account of justified belief. according to this, an epistemically responsible enquirer will aim at having true beliefs and will therefore perform all the necessary actions to collect sufficient evidence to support his belief, such as checking a broad enough range of e.g. web pages and comparing them to other types of sources (cf. bråten, brandmo, & kammerer, 2018). there are, however, three cases where the enquirer may fail to acquire justification for his belief: 1) the enquirer neglects performing a proper search, 2) the enquirer performs a proper enquiry, but the results do not support his belief or 3) the activity to justify his belief is not possible, e.g. due to lack or impracticability of a technology. assuming that an enquirer is literate enough to avoid the first case, he can still fail as in cases 2 and 3. in cases of internet searches the problem is that, due to the covert search logics, the enquirer may not even know that he has failed. he may believe that he has performed a proper search but, due to the search engine’s filtering and ranking, the results may not provide the full picture of facts required to justify or rule out the belief. furthermore, due to the covert operating logic, it is impracticable (case #3) for the enquirer to assess the quality of the set of sources provided by the search engine. as shown above, the past decades’ technological development has induced changes in how individuals acquire information, and blurred the boundaries between personal knowledge and external information. the problem is not about using external memory aids or systems for offloading information (säljö, 2012). as säljö explains, man started developing external symbolic storages and artificial memory systems thousands of years ago, and memory aids such as otto’s physical notebook (smart, 2012) or address books in smartphones are everyday tools used to offload information from our memory. however, there is a risk that (especially young) users not only offload information but perhaps even outsource cognitive processes, since they may lack the epistemic competencies and practices required in this new information ecology (cf. bhatt & mackenzie, 2019; fisher et al., 2015; säljö, 2012; sparrow et al., 2011). to provide a rationale for the approach of this study, the following sections will briefly review 1) epistemic beliefs as a research area, 2) how epistemic beliefs may relate to learning and 3) dimensions and tools for measuring epistemic beliefs. these sections also aim at explaining how this study was delimited and why some aspects, albeit frequently discussed in other studies, were not included in this study. 2.2 epistemic beliefs as a research area william g. perry’s (1970) study of college students’ ideas regarding source and certainty of knowledge is commonly regarded as the starting point for research on epistemic beliefs or personal epistemology over the past decades, epistemic beliefs have been conceptualized in different ways (cf. schraw, 2013). some researchers conceive them as broad and developing stage-like. other researchers conceive them as a set of more or less independent dimensions expressing beliefs about knowledge and learning, marlene schommer (1990; 1993) being the first in this line of research. the term ‘epistemic beliefs’ will be used here since the study will focus on the respondents’ (implicit and unconscious) views of knowledge, not their theories of knowledge or epistemology (cf. kitchener, 2002; hofer, 2008, p. 5). the works during the 1990ies of marlene schommer (1990; 1993, later schommer-aikins) and barbara k. hofer and paul r. pintrich (1997) in developing research around epistemological theories are important to acknowledge. during the first decade of this century, research around epistemic beliefs increased and extended from perry’s original north american, white, elite, male college students context to other age groups and geographical and cultural contexts. for extensive overviews, please see the works by hofer & pintrich (2002), niessen, vermunt, abma, widdershoven, & van der vleuten (2004), debacker, crowson, beesley, thoma, & hestevold (2008) and khine (2008). further, the more recent works by schraw (2013) greene, sandoval, & bråten (2016), bernholt, gruber, & moschner (2017) and knight et al. (2017), out of which the four latter where not yet available at the time for planning this study. domain-specificity and domain differences have been issues throughout the years. the initial assumption, that one’s epistemic beliefs are general across domains, has been questioned and instead, it has been suggested that one can hold different epistemic beliefs, depending on the field of knowledge one is dealing with (muis, bendixen, & haerle, 2006). the longitudinal study by trautwein & lüdtke (2007), albeit focusing on the certainty dimension only, confirmed the hard-soft difference but also that students aiming at certain college programmes differed regarding their beliefs already at the end of their upper secondary education. in their large review, muis et al. (2006) noted that empirical research had been presented in support for both domain general and for domain specific epistemic beliefs respectively, and that they may co-exist and possibly interact. the suggestions by muis et al. were strongly supported by both hofer (2006) and alexander (2006). to conclude, i acknowledge the co-existence of and interaction between domain-general and domain-specific epistemic beliefs. the question regarding domain-generality vs. domain-specificity was, however, not the focus of the present study. the development of self-report instruments for measuring epistemic beliefs has encountered several challenges. in his review article, schraw notes that there has been disagreement about the underlying conceptual structure, and replications of exploratory factor analyses (hereafter efa and cfa will be used for exploratory and confirmatory factor analysis, respectively) have often failed. common problems have been that items load in an unexpected manner often resulting in less factors or another factor structure than anticipated in the underlying conceptual model, too few items loading per factor and the resulting model showing a low explanation score. (schraw, 2013) in the present study, i will subscribe to the line of research that considers the concept of epistemic beliefs as multidimensional. assuming that hitherto described dimension sets are not sufficient to describe epistemic beliefs in the new information ecology, i attempt to introduce some new dimensions. the aim of testing new dimensions required starting on a general level and therefore, the questionnaire items (except for the internet-related items) did not refer to any specific discipline or context (section instrument construction). 2.3 epistemic beliefs and learning alongside with motivation and cognitive styles, the concept of epistemic beliefs is an important factor affecting learning and study success. hofer & pintrich (1997) called for more research to understand how students’ epistemic beliefs may influence learning performance. further, they suggested that the type of learning tasks may shape the students’ epistemic beliefs, as shown later by kienhues, bromme, & stahl (2008). brownlee, walker, lennox, exley, & pearce (2009) approached the topic of epistemic beliefs qualitatively, and their results highlight that first-year students may hold subjectivist or objectivist core beliefs that may decrease their ability to engage in critical thinking, required in higher education. walker et al. (2009) also approached first-year students and identified some students being at risk of having difficulties in higher education due to their naïve beliefs about learning and knowing. there is also evidence suggesting cultural differences. zhang & watkins (2001) observed that chinese students’ cognitive-developmental patterns were the opposite of the patterns observed in the u.s. sample. their results also indicate that epistemic beliefs are not static but developing (cf. kienhues et al., 2008). further, hofer (2008, pp. 11-12) observed differences between japanese and us college students such that us students had more sophisticated beliefs about the factors describing certainty, simplicity, source and justification of knowledge. education is moving towards methods of teaching and learning that often involve using internet-based resources (e.g. the flipped classroom, knewton, 2011). these methods require more self-regulation from part of the student, and e.g. bråten (2008, pp. 369-370) highlights the risk that students with naïve epistemic beliefs may tend to over-reliance towards internet-based resources. regarding the changes in teaching methods, it is worth noting that the teachers’ choices of pedagogical activities and learning settings are also influenced, perhaps unconsciously, by the teacher’s own epistemic beliefs (palmer & marra, 2008, p. 337). an overall awareness regarding epistemic beliefs is called for among teachers at all levels of education. an interesting attempt to support this awareness is the theoretical model between epistemic beliefs and self-regulation suggested by muis (2007), where epistemic beliefs facilitate self-regulation and play a crucial role in all four phases (task definition, goal setting, enactment and evaluation) of the learning process. an example from a constructivist education context (pbl) is the study by otting, zwaal, tempelaar, & gijselaers (2010), where the results showed a connection between conceptions of expert knowledge and traditional conceptions of teaching and learning on one hand, and on the other hand a connection between learning effort and a constructivist conception of teaching and learning. the examples above illustrate that there is much going on within the educational context, most importantly that education is moving from being teacherand subject-centred towards being more studentand learning-centred. the development of the technological structures around ict is increasingly beyond control of the educational system. however, learning analytics is an area where education is actively applying ict: the core characteristic is the generation of high-resolution data about various types of [learning] actions (knight, wise, & chen, 2017), and applying knowledge from multidisciplinary perspectives such as business intelligence, web analytics and data mining for analysis purposes (ferguson, 2012). thus, learning analytics can generate real-time individual and group performance information with potential to support teachers’ decision-making (knight, wise, & chen, 2017). knight et al. (2017) present a novel approach as they explore how students’ epistemic beliefs predict e.g. students search behaviour (traced using learning analytics methods). their results did not show a convincing predictive value, whereas the results by pieschl, stallmann, & bromme (2014) were a bit more encouraging. this issue is further commented in section internet-specific epistemic beliefs. 2.4 dimensions and measurement within the line of investigation that regards epistemic beliefs as multidimensional, self-report instruments have been developed to capture the dimensions of epistemic beliefs. in her original 63-item schommer epistemological questionnaire (seq), schommer (1990; 1998) suggested the dimensions simple knowledge, certain knowledge, innate ability, quick learning and omniscient authority. using efa, schommer managed to extract four but not the omniscient authority dimension. thus, the dimensions described views on both knowledge and learning. several authors (e.g. hofer & pintrich, 1997) have criticized schommer for not performing factor analysis on the 63 original items but using 12 subscale scores (packages) based on those items, as variables. still, schommer’s questionnaire has been the starting point for a large part of later development regarding questionnaire-based instruments (for an overview, please see niessen et al., 2004), out of which the following instruments, besides the seq, were used as reference in the present study: wood & kardasch (2002) developed the epistemological beliefs survey (ebs) containing 38 items, out of which 32 stemmed from or resembled items in the seq, and covering two seq dimensions. schraw, bendixen, & dunkle (2002) developed the epistemic beliefs inventory (ebi) containing 28 items, out of which 17 stemmed from or resembled items in the seq. ebi reflected the same dimensions as the seq. moschner, gruber, & studienstiftungsarbeitsgruppe epi (2005) developed the 43-item fragebogen zur erfassung epistemischer überzeugungen (questionnaire for capturing epistemic beliefs, hereafter fee) containing nine items from seq. fee included the three seq dimensions certainty of knowledge, learning ability and omniscient authority. additionally, the fee proposed five new dimensions labelled social aspects of knowledge, value of knowledge, culture related aspects of knowledge, gender related approaches to knowledge and reflective nature of knowledge. 2.4.1 knowledge, knowing and learning the discussion whether epistemic beliefs should be limited to beliefs about knowledge and knowing, or whether beliefs about learning should be included, has been ongoing throughout the decades. hofer & pintrich (1997) recommended excluding beliefs about learning for the sake of clarity of the concept of epistemic beliefs. instead, they retained certainty and simplicity of knowledge (describing nature of knowledge) and proposed the dimensions source of knowledge and justification for knowing to describe the nature of knowing. schommer introduced an embedded systemic model that included beliefs about ways of knowing, interplaying with beliefs about knowledge and beliefs about learning, i.e. beliefs about knowledge and learning as separate constructs but within the same system (schommer-aikins, 2004). sandoval (2005) warned for conflation of the concepts. although beliefs about knowledge will probably influence one’s beliefs about learning, sandoval proposed that they should be investigated as separate constructs. in a comment to the discussion, elby (2009) suggested that it is too early to decide and therefore, views on learning should at least for the time being be included in the concept of epistemic beliefs for further empirical and theoretical development. for the present study, data were collected regarding beliefs about both knowledge and learning and consequently, the analyses include both aspects. this approach is also supported by previous research presented in the section epistemic beliefs and learning. 2.4.2 internet-specific epistemic beliefs the point of departure for this study, the tendency not to look up information until needed and to rely on internet-based sources, is close to the research regarding internet-specific epistemic beliefs by bråten, strømsø and their teams. in 2005, they developed the internet specific epistemic questionnaire (iseq: bråten, strømsø, & samuelstuen, 2005), which was based on the four dimensions described by hofer & pintrich (1997) and thus omitting learning dimensions. in performing efa, bråten et al. used maximum likelihood (hereafter ml) as extraction method together with an oblique rotation method but did, however, extract only two factors. they labelled the first one general internet epistemology, which included beliefs concerning the certainty and simplicity of internet-based knowledge, as well as beliefs concerning the internet as a source of knowledge, i.e. three dimensions in one factor. the second factor was labelled justification for knowing and described whether internet-based knowledge claims could be accepted without critical evaluation, or should they be critically evaluated using multiple sources, reasoning and prior knowledge. all eighteen iseq items referred to internet and thus, all questions connected explicitly and exclusively to the internet context. further, when reading the iseq general internet epistemology items it seems obvious that they do not actually reflect the certainty or structure of knowledge (cf. corresponding items in table 1) but rather, they mainly express the coverage and availability of information on the internet. thus, the iseq seems to leave questions open about the respondent’s beliefs regarding certainty and simplicity of knowledge in general, about the beliefs regarding other sources of knowledge, and how these beliefs relate to each other; unanswered questions constituting a research gap. in a subsequent study, bråten & strømsø (2006) applied parts of the seq (schommer, 1990), but not the iseq, to explore the connection between epistemic beliefs and internet-based search and communication activities. it turned out e.g. that students who believed in quick learning tend to overlook the importance of critically evaluating web-based resources. in another study, based on 17 out of 18 items in the iseq item set, the authors extracted only three factors (using ml and direct oblimin): certainty and source of knowledge, justification for knowing and structure of knowledge (strømsø & bråten, 2010). the iseq has also been applied in other contexts and for other purposes: karimi (2014), exploring the connection between internet-specific epistemic beliefs and grammar achievement, extracted the same three factors as strømsø & bråten (2010), although with varimax rotation. chiu, liang, & tsai (2013) used a chinese translation of the iseq, and applied an efa method (apparently with oblique rotation) but upon only twelve items. these authors did, however, not extract iseq dimensions as described by bråten et al. (2005) but instead, the four dimensions originally suggested by hofer & pintrich (1997), i.e. certainty, simplicity and source of knowledge and justification for knowing, but using items specifically denoting an internet-based context. kammerer & gerjets (2012) applied iseq to categorize users for comparison, but they only used eight items attributed to the iseq-dimension certainty and source of knowledge and thus, did not test the factor structure proposed in the original iseq. the study by knight et al. (2017) exemplifies a research approach linking epistemic beliefs with log data analytics. they used the iseq in an extensive study to explore whether the two-factor iseq scores could predict e.g. trustworthiness ratings of internet-based sources or traced search behaviour. according to their results, the factor scores did not predict search behaviour, and they had only small predictive value for trustworthiness rating. the approach by knight et al. is interesting and relevant but raises the question: could the connections to search behaviour have turned out differently had they not used the two-factor iseq, where the general internet epistemology factor contains a mix of certainty, structure and source of knowledge? e.g. the results by pieschl et al. (2014), indicate that students’ epistemic beliefs influence how they approach complex tasks. to conclude, epistemic beliefs have been explored also in relation to internet-based information, but the picture is disparate. the studies referred to above, as well as many other studies, suffer from the problems addressed by schraw (2013). the studies published prior to the present data collection (bråten et al., 2005; bråten & strømsø, 2006; strømsø & bråten, 2010) focused on beliefs about internet-based information without actually relating these beliefs to beliefs about knowledge based on other information sources. the studies referred to above also leave the question open, whether internet should be regarded as an authority or knowledge source, or a specific context (cf. grossnickle peterson, alexander, & list, 2017, p. 262). 2.4.3 justification for knowing hofer & pintrich (1997) introduced justification for knowing as a dimension, which was later supported by several researchers. both alexander (2006) and greene, azevedo, & torney-purta (2008) have noted that this dimension is least developed, and that exploring justification is more challenging than exploring other dimensions. this assumption seems well founded considering the complexity of the justification aspect, e.g. in terms of the responsibilist account of justified belief suggested by miller & record (2013) (see section personal knowledge, external information). greene et al. (2008) point out two aspects that are part of the challenge in investigating the justification dimension. first, considering the number of different kinds of justification identified in philosophy, justification as part of the epistemic beliefs model will probably require to be described by multiple factors rather than one single factor. further, greene et al. suggest that a person needs to have a sophisticated ontology of a domain before issues of justification, such as critical thinking, become relevant. this seems congruent both with bloom’s original cognitive process dimensions and especially with the knowledge dimensions described later by krathwohl (2002): issues of justification are probably far more relevant when applying, analysing or evaluating conceptual knowledge than when recalling facts. the above suggestion by greene et al. comes to expression in a recent study by bråten, brandmo & kammerer (2018), where they delimit the context to internet and the domain to that of educational topics within teacher education. their study focuses solely on the justification dimension, approaching it as a three-dimensional concept including justification by authority, justification by multiple sources and justification against prior personal knowledge and reasoning. as a result, they present the validated internet-specific epistemic justification inventory (isej). against the background of the considerations referred above, and the fact that epistemic beliefs in the new information ecology was totally uncharted territory, it seemed appropriate to leave the justification dimension outside this investigation. hence, the fee instrument (moschner et al., 2005) was chosen as a starting point (see section instrument construction). 2.5 research questions capturing all dimensions of epistemic beliefs (or cognition, cf. greene et al., 2008) while at the same time adding and testing new dimensions would be both adventurous and beyond this study. therefore, while acknowledging that epistemic beliefs consist of multiple dimensions developing over time, this study adopts a narrow focus on capturing a snapshot of the participants’ current epistemic beliefs, including beliefs in internet-based information. thus, the justification dimension as well as the topics regarding subject-, domain-, discipline-, cultureor gender-specificity of epistemic beliefs (see e.g. debacker et al., 2008) are beyond the scope of this study. the approach of this study is openly explorative in testing whether it is possible, overall, to extend the existing instruments and their dimension sets with new dimensions of epistemic beliefs, and specifically to capture such ways of relating to knowledge that have become common among frequent internet users during the past decades. further, this study will explore the relation between existing epistemic dimensions and those describing internet-based knowledge and knowing. apart from iseq (bråten et al., 2005), this study does not aim to explore how individuals justify internet-based information, but rather to explore whether and to which extent individuals rely on and prefer internet-based information sources, and how this preference relates to other epistemic dimensions. the investigation is framed in a single research question: (how) can the set of epistemic beliefs dimensions be extended so that it also expresses a googling attitude? the research question is openly phrased since, although research on epistemic beliefs has been going on for some time, the proposed dimensions are on uncharted territory. for the sake of clarity, i will use the term original dimensions for those dimensions described in or stemming from schommer’s seq (1990). hypothesized dimensions will be used to denote suggested dimensions until their existence has been confirmed, after which they are denoted as novel dimensions or scales in the proposed model, which is the endpoint of the present study. 3 material and methods by way of introduction to this section, i provide a rough outline for instrument construction and data collection. the first version of the instrument was created for the data collection in august 2011. the instrument was evaluated so that a revised version was used for the second data collection in august 2012, which resulted in the material being reported here. the usability and validity of data from 2012 is commented in the discussion section. it needs to be noted, that after the current data were collected, new studies describing further development have been published. the present instrument was, naturally, based on instruments that were published and available prior to 2012. 3.1 instrument construction the fee questionnaire developed by moschner et al. (2005) combined experiences from previous instruments and also contained some potentially interesting extensions. therefore, the fee was taken as point of departure for constructing the first version of the on-line survey called ‘me and my knowledge’. a replication of using the fee-specific items was performed on the first data set collected in 2011 (reported in ståhl & mildén, 2017). due to unsuccessful replication, the instrument was revised prior to the 2012 data collection: the five new dimensions suggested in fee were omitted, structure of knowledge items were included as well as some other items, based on item level analysis. in addition, some items describing the hypothesized subscales were reversely phrased. table 1 shows the entire instrument, item descriptives and item associations before and after analyses. since swedish and english are the working languages of the university, the questionnaire was set up in both languages. to ensure comprehensibility, both swedish-speaking domestic and english-speaking international students were involved in read-aloud sessions during instrument construction. an important aspect of the cultural adaptation of the questionnaire was rephrasing the questions into first person present tense, as suggested e.g. by kitchener (2002) and schommer-aikins (2004, p. 23). the main motive was to ensure a first-person perspective: the phrasing should clearly signal that the researchers were interested in knowing what the student herself thinks, not what she thinks that people in general think, or what is socially desirable to think about a topic. during the read-aloud sessions, the students provided valuable feedback acknowledging the need for cultural adaptation and inducing some further rephrasing. overall, the students’ feedback supported the choice to use direct and active wording. the items were consistently generic (not domainor discipline-specific), and the instructions did in no way refer to relating the responses to any specific subject, academic field or context (cf. wood & kardash, 2002, p. 244; muis et al., 2006, p. 25).   table 1 questionnaire items in original and hypothesized dimensions, including item descriptives and item use in the proposed model table footnotes a) the number after 'k' refers to the page number (03-14) in the web questionnaire. the number after 'f' refers to the original fee numbering. b) able learning ability; auth omniscient authority; cert certainty of knowledge; constr constructivist approach; dia learning by dialogue; int internet reliance; struct structure of knowledge c) item phrasing is reverse compared to other items in the same dimension. 3.1.1 previously established dimensions the fee questionnaire (moschner et al., 2005) included the original seq dimensions certainty of knowledge, omniscient authority and learning ability. unfortunately, the dimension structure (or simplicity) of knowledge was excluded from the fee but was included in the 2012 survey being reported here (table 1). justification of knowledge should undoubtedly be a part of the epistemic beliefs dimension set. however, the new students (see section participants and data collection) that were involved as informants could hardly be expected to possess a sophisticated ontology of the domain they were just entering to study (cf. greene et al., 2008). based upon this, upon previously presented considerations (section justification for knowing) and upon the scope of the study, the justification dimension was omitted at this stage. 3.1.2 hypothesized dimensions out of the five new dimensions suggested in the fee, only reflective nature of knowledge was used in this study, and the items associated with it were rephrased to reflect reflective nature of learning. this dimension deals with the learning aspect and was intended to express a reflective stance towards new knowledge. the debate regarding digital natives did not produce an actual definition for digital natives but instead, researchers published different descriptions about how the (supposedly) digital generation acted and behaved (cf. ståhl, 2017). therefore, the dimensions described below were constructed with a starting point in descriptions regarding attitudes towards knowledge and learning, as reported in various studies. the instrument also set out to test whether the suggested attributes could be identified within this sample. the descriptions of connectivism (downes, 2007; siemens, 2005; 2006, pp. 31, 91) together with anderson & balsamo (2008, p. 244) stating that "they treat their affiliation networks as informal delphi groups” have contributed to the items proposed to describe a connectivist approach to learning (hereafter the short forms connectivist approach and constructivist approach will be used). a constructivist approach to learning has yet not been suggested in previous instruments, although some items in the dimension knowledge construction and modification suggested by wood & kardash (2002, p. 250) and the dimension reflective nature of knowledge suggested by moschner et al. (2005) point in this direction. the writings of siemens (2006, pp. 6, 20, 31) have also provided input to the items proposed to describe a constructivist approach. anderson & balsamo (2008, p. 244) described the young generation as ”…knowing and being confident where to find information once they need it”. siemens (2006, p. 31) described deciding what to memorise and choosing what to learn as characteristics in connectivist learning, inspiring the construction of items describing the hypothesized dimension just-in-time learning. reliance on internet is an integral part of the googling mind-set. at the time of planning this research the iseq had been introduced (bråten et al., 2005) but as mentioned above (section internet-specific epistemic beliefs), the iseq items focussed exclusively on internet-based information. thus, the five items concerning internet-based knowledge in the present instrument were generated from literature regarding the so called digital natives and the net generation (prensky, 2001; siemens, 2005; anderson & balsamo, 2008), and their preference for internet sources instead of printed sources (cf. head & eisenberg, 2010; purcell et al., 2012, p. 33). the items where phrased to express how the googling mind-set reflects a reliance in that any information you need can always be found on internet and accordingly, the dimension was labelled internet reliance. siemens (2006, pp. 16, 31, 56, 117) described valuing diversity as a central trait in connectivism, which requires interaction (downes, 2007, p. 78) and also involves exposing oneself to and valuing different opinions, all contributing to the individual learning process. this trait, requiring “… the widest possible spectrum of points of view…” (siemens, 2006, p. 16), can be regarded an expression for both a general scholarly approach and also the epistemic development from realist over absolutist and multiplist to evaluativist understanding (kuhn & weinstock, 2002, p. 124). the present instrument includes four previously described and six hypothesized dimensions, altogether 60 items (table 1). 3.2 participants and data collection the study was part of a university development project with the objective of collecting information about the new students’ mind-sets to develop teaching and learning practices. the university’s board on ethics approved the project research plan, including procedures for data collection, analysis and reporting. data were collected among all new students in august 2011 and 2012 (n = 476/440). since epistemic beliefs can change through intervention (cf. kienhues et al., 2008), it was crucial to get a “snapshot” of the students’ epistemic beliefs by collecting data during the very first week of the semester, before the students were exposed to study subjects or pedagogical influences at the university. data collection was organised during compulsory and scheduled ict level test sessions, where students first completed another survey called ‘ict, media and me’, then the compulsory ict driving license level tests and finally the survey ‘me and my knowledge’. figure 1 on-line questionnaire screenshot the students were introduced to the objectives of the project, and informed orally and in writing that although the ict level tests were compulsory, the surveys were voluntary and did not include any financial or other incentives. due to the survey being an operationalization of the university’s statutory obligation to continuously develop its education, informed consent was registered following a simplified procedure. the students were informed that by (performing the action of) filling in the questionnaire, they express their consent for the data being used for the purposes described in the information sheet and in the description of the scientific research data file as required in the legislation concerning personal data in research (personal data act, 1999). accordingly, the students had the opportunity to withdraw their permission by contacting the researcher by a given date, after which the data set was anonymized. the students were also introduced into the functionality of the questionnaires and informed that support was provided if needed. the survey was presented in an on-line questionnaire using a 6-point likert-type response format (figure 1). when applying the 63-item seq, wood & kardash (2002, p. 244) received student comments indicating respondents’ difficulties in understanding certain items. although some researchers (e.g. martin, 2005, p. 728) discourage the use of ‘don’t know’ options, the scale in this questionnaire was supplemented with two non-substantial options, ‘don’t know’ and ‘don’t understand’. this is partly supported by muis et al. (2006, p. 25), noting that it has not been empirically studied what individuals actually think as they fill out questionnaires. providing both options was especially important when introducing new items, since these options provided information regarding comprehensibility, potentially valuable when considering items to exclude (cf. finch, immekus, & french, 2016, p. 144). further, the non-substantial options were placed on both sides of the substantial options in order not to distort the visual midpoint of the likert-type response format (cf. tourangeau, couper, & conrad, 2004). in survey presentation, it was necessary to prevent fatigue effect and satisficing (cf. cape, 2010), and any effect where question context or order might influence question interpretation (cf. martin, 2005, p. 726; tourangeau et al., 2004). therefore, a progress indicator was included and the items were distributed over twelve pages containing four to six items each, which also improved readability. further, to prevent inter-item influence, each subscale’s items were distributed over different pages (e.g. the page in figure 1 containing items from five subscales) and the survey service was set to randomise item order within each page. 3.3 research data and sample characteristics the present study is based on data collected in 2012, where 371 students chose to complete the survey ‘me and my knowledge’. only those cases containing substantial responses to more than 70% of the items where retained for further analyses (n = 348). the 23 excluded cases had responded only to first-page items and were therefore regarded as dropouts. the complete data set with 371 cases exhibited missing values increasing from 4.2% up to 11.7% on page level, whereas this trend in the 348-case subsample developed from 2.5% to 7.0%. this, together with the dropouts, indicates that most respondents who started the survey also completed it, and that an actual fatigue effect was avoided. on item level, the portion of missing values ranged from 1.4% to 10.1%, where the two certainty of knowledge items k13_2f44 and k14_2f49 (table 1) showed the highest portions of missing values, mostly ‘don’t know’ responses. the highest ‘don’t understand’ portions occurred for three items representing the dimensions just-in-time learning (k04_1), constructivist approach (k07_2) and connectivist approach (k08_5). since the questionnaire applied a likert-type response format producing data on an ordinal scale, it is not meaningful to analyse distribution or assess normality on item level (cf. carifio & perla, 2007) but instead, analysis of the actual scales is postponed to the discussion section. for those calling for an item level analysis it can be mentioned that for each item, the response value ranged over the whole scale (1..6). the items showed a standard deviation between 1.02 and 1.66, a skewness between -1.23 and 0.91 and a kurtosis between -1.26 and 1.38. the criterion of the skewness and kurtosis value being within the range ±1 was met regarding 57 and 55 items, respectively. the shapiro-wilks test suggested non-normal distribution, whereas the kolmogorov-smirnov test suggested normal distribution throughout all items. a visual inspection of histograms, normal q-q plots and box plots showed that the items were approximately normally distributed. for the items showing skewness or kurtosis outside the ±1 range, the deviation was minor and further, the sample size was large enough to reduce a possible detrimental effect (cf. hair, black, babin, & anderson, 2010). based on the aforementioned criteria, the items were considered as normally distributed. the current 348-case subsample holds students from twelve degree programmes, both domestic and international students (86.8% / 13.2%), and a gender distribution holding 66% female students. the age average was 21.7 with 91% being born in 1986-1995. for this study, sample demographics should be reviewed in relation to access to internet resources. internet and publicly available search engines were launched already in the mid 1990’ies and during the following ten years, search engine use was established (http://www.searchenginehistory.com/). 2011-2012 were the very years when internet services, previously available via computers, became truly ubiquitous due to 3g/4g-connected smartphones becoming everyday tools, and finnish net operators offering affordable 3g/4g-subscriptions including generous mobile data. the mobile phone prevalence within both cohorts was close to 100%. smartphone as a concept was not yet established and thus, the corresponding survey item was phrased “my mobile phone is connected to the internet”. from 2011 to 2012, the portion of users across the cohorts having an internet-connected phone increased among domestic students from 48.7 to 81.3% and among international students from 60.7 to 90.6%, within the total cohorts from 50.2 to 82.5%. this corresponds well with the national statistics, according to which 53% of those aged 16-24 had a smartphone in the spring of 2011 (osf, 2011). at the time of data collection, the respondents had been exposed to computers, mobile phones and internet for in average 12, 10 and 9 years respectively. to conclude, the sample can be regarded a rather typical net generation cohort. 3.4 analysis methods fabrigar, wegener, maccallum, & strahan (1999) argue that principal component analysis is not a true method of factor analysis. they recommend the use of maximum likelihood, as later supported by osborne (2014, p. 9) and finch et al. (2016, p. 131). thus, the analysis procedure starts with an efa with ml as extraction method, followed by a validation procedure including efa and cfa on split halves of the sample (fabrigar et al., 1999; fokkema & greiff, 2017; knight et al., 2017; leal-soto & ferrer-urbina, 2017; osborne, 2014, pp. 6, 119-120; tang, 2010). for all statistical tests, a significance level of .05 was used and in efa and cfa procedures, the absolute loading value .32 was used as the threshold when assessing item (non-)loadings and cross-loadings (cf. finch et al., 2016, p. 143). in table and diagram presentations, loadings <.32 are generally not displayed although during efa procedures, low loadings were not suppressed since that may cause loss of valuable information (such as item k14_5, table 3). the spss software package (spss, 2016b) was used for efa procedures, and the cfa procedures were performed using the amos software package (spss, 2016a). to build analysis on true data, missing data values were not imputed, since any kind of replaced or imputed values are, after all, only estimates. this choice was made at the cost of listwise deletion reducing the number of cases in efa, and missing values thwarting the use of modification indices to support refinement in cfa. 4 results in this section, the results are presented together with analyses, since some results inform the subsequent steps. the reasoning behind e.g. item disposal, number of factors and factor labelling will be presented in conjunction with efa on the complete item set. 4.1 original and hypothesized dimensions of epistemic beliefs 4.1.1 replicating original dimensions for replication purposes, efa was first performed on the 27 items associated with the original dimensions. dysfunctional items (zero, low and cross-loading) were stepwise discarded (cf. finch et al., 2016, pp. 143-144). a model based on 18 items showed good fit indices and was interpreted as a successful replication (despite ml extraction and promax rotation). 4.1.2 emerging dimensions the 60-item questionnaire contained 33 items that were associated with six hypothesized dimensions: reflective nature of learning, connectivist approach, just-in-time learning, constructivist approach, internet reliance and valuing diversity (table 1). seeking inspiration from bråten et al. (2005) and trautwein & lüdtke (2007) who analysed only one or two factors, this item subset was initially factor analysed separately in order to identify dysfunctional items. the efa on the hypothesized dimensions was truly exploratory, including different rotation methods, varying the number of extracted factors and stepwise reduction of dysfunctional items (finch et al., 2016, pp. 143-144; osborne, 2014, pp. 17, 30-33). using ml, promax rotation and listwise deletion, the efa resulted in a four-factor model based on 23 items (n=191). the model showed good fit indices (eigenvalues 6.35 .. 1.5, 51% of variance explained; kmo=.864, bartlett's chi-square=1397, df=253, sig.<.000, goodness-of-fit test chi-square=182.7, df=167, sig.=.192) and reflected three of the hypothesized dimensions: constructivist approach, connectivist approach, internet reliance and a fourth one, now labelled learning by dialogue. each item loaded strongly on one factor without cross-loadings. throughout the various models, constructivist approach and connectivist approach correlated strongly, and several of the other factors correlated weakly with each other (table 2). table 2 factor correlation matrix, four-factor model based on 23 new items 4.1.3 an extended set of dimensions since inter-factor correlation occurred in all the previous analyses, efa on the complete item set were performed using ml extraction and promax or oblimin as oblique rotation methods (cf. finch et al., 2016, p. 133; knight et al., 2017; osborne, 2014, pp. 30-33; strømsø & bråten, 2010). the model was stepwise refined by removing low-loading and cross-loading items while simultaneously assessing their conceptual relevance, their communality estimates and their internal consistency within the anticipated scale. during the process, the six reversely phrased items occurring in four dimensions (table 1) were discarded due to dysfunctionality. thus, within all the hypothesized dimensions, the items were unidirectional. the refinement procedure boiled down to a model with seven factors and 26 items that fit the data reasonably well (table 3). both original and hypothesized dimensions appeared distinctly without cross-loadings (except for items k10_7 and k14_5), each dimension loaded on at least 3 items, and 25 out of 26 items loaded (>.32) on the anticipated factor. table 3 the proposed efa model based on 26 items table footnotes ml, promax rotation converged in 16 iterations; listwise deletion, n=195, eigenvalues 4.85 .. 1.01, 59.3% of variance explained; kmo=.782, bartlett's chi-square=1468, df=325, sig. <.000, goodness-of-fit test chi-square=169.8, df=164, sig.=.361 a) item loading not consistent with hypothesized dimension but conceptually coherent. from this model, several items were dropped due to loading weakly or inconsistently with the hypothesized dimension. the retained original items loaded on the same factors as in the eq, except for item k04_4f04 that was originally associated with omniscient authority. the current loading on certainty of knowledge can be regarded as conceptually coherent. the issues regarding which items to discard, how to decide on the number of factors, and how to label the subscales require some comments. as osborne (2014, pp. 17-18) notes, efa is a low-stakes procedure and expressly exploratory. accordingly, i entered the process with 60 items, a hypothesized underlying conceptual model, and used the statistical package (spss, 2016b) to provide suggestions for a factor model. during efa iterations, the dysfunctional items were eventually revealed and discarded. besides varying extraction and rotation, the search for an adequate number of factors included extracting factor sets ranging from two factors below up to two factors above the number suggested by the scree plot elbow (osborne, 2014, p. 18). this method provided valuable information: increasing the number of factors caused related items to split over several factors, whereas reducing factors caused items to pile up on one factor, usually then holding items from dimensions that at the end turned out to correlate. thus, the search for a factor model included weighing of theory, scree plot, item loadings and communalities, eigenvalues, internal consistencies and conceptual considerations. during the efa iterations, the hypothesized dimensions did not turn out quite as anticipated, which is only part of the nature in explorative work (cf. osborne, 2014, p. 17). in most of the explored models, the five hypothesized dimensions boiled down to three (table 3). the dimension learning by dialogue holds items from the suggested dimensions valuing diversity and connectivist approach, whereas the dimension constructivist approach besides its own items also holds items from the hypothesized dimensions valuing diversity and reflective nature of learning. all the three items originally associated to internet reliance consistently loaded on that dimension (table 1). retaining the items k14_5 and k10_7 violates the rule of using only strong, single-loading items and requires commenting. the internal consistency test showed that deleting the item k14_5 entailed a slightly improved alpha value, but at the cost of reducing the factor internet reliance to only two items. further, since the item communality value was reasonably good, the connection to structure of knowledge occurred also in the path diagram, and the cfa indicated that discarding the item impaired fit indices, there was enough support for retaining the item. as expected, the item k10_7 loaded strongly on the structure of knowledge dimension, but surprisingly also on constructivist approach. the item was retained since discarding it would have impaired the structure of knowledge internal consistency considerably. the split half efa suggested single-loading on structure of knowledge, whereas the cfa suggested a connection to constructivist approach and indicated that discarding the item impaired fit indices. three of the original dimensions (except learning ability) correlated weakly with each other. further, constructivist approach correlated strongly (.506) with learning ability and weakly (.282) with learning by dialogue. 4.2 evaluating the extended instrument for the purpose of evaluating the stability of the model presented in table 3, the data set was randomly split into two equal halves a and b, that were subject to efa and cfa, respectively (cf. fokkema & greiff, 2017, p. 401). 4.2.1 exploratory split half an efa was performed with 26 items on the split half a using the same methods as in the initial model. the model arrived at (table 4) did not show a one-to-one correspondence to the initial model (table 3) but resembled it strongly, with 22 out of 26 items loading as anticipated. all the hypothesized dimensions were reflected in the seven factors, although two of the constructivist approach items loaded on the learning ability factor (a). further, two items loaded on unexpected factors (b), and the learning by dialogue item k03_5 caused a heywood case. still, the fit indices suggested that the proposed model, appearing almost similar in both oblimin and promax rotation, fit also the split data set fairly well. table 4 exploratory factor analysis on split half a table footnotes a) item loading not consistent with hypothesized dimension but conceptually coherent b) item loading not consistent with hypothesized dimension, vague conceptual coherence ml, oblimin rotation converged in 12 iterations; listwise deletion n=96, eigenvalues 5.00 .. 1.06, 62.9% of variance explained; kmo=.706; bartlett's chi-square=928, df=325, sig. <.000; goodness-of-fit test chi-square=179.5, df=164, sig.=.193 as in the initial model, the constructivist approach factor correlated strongly with the learning ability factor (.494) and omniscient authority correlated with certainty of knowledge (.363). 4.2.2 confirmatory split half cfa was performed on the same 26 items as the previous efa but on the split half b (n=174) of the data set. in the first step, conceptually irrelevant and low connections between latent variables were removed, which resulted in an initial model with partly insufficient fit indices. assessing model fit and choice of cut-off criteria (in brackets) follow the recommendations by schreiber, nora, stage, barlow, & king (2006) and hooper, coughlan, & mullen (2008). since the data set contained empty cells, it was not possible to utilize the feature where the amos software would provide suggestions for modification (spss, 2016a). instead, model refinement was performed manually, partly following loadings and correlations indicated in the previous efa models (tables 3 and 4), and partly by adding and removing connections based on conceptual considerations in an exploratory manner. thus, some connections between latent variables, although weak, were retained, under the condition that they were conceptually defensible and contributed to improving fit indices. then again, in some cases conceptually defensible connections had to be discarded if their loading value was low (mainly <.32) and retaining them impaired the fit indices. the procedure resulted in a conceptually defensible path diagram (figure 2), similar to the efa models (tables 3 and 4) and reasonable although not perfect fit indices (chi-square/df=1.504, rmsea=.054, tli=.821, cfi=.853, pclose=.260). the item k13_6 caused a minor heywood case (1.01), which was accepted since any attempt to manipulate constraints impaired fit indices. figure 2. simplified cfa path diagram based on 26 items and split half b data set (cut-off criteria in brackets). chi-square/df=1.504 (<2), rmsea=.54 (<.60), tli=.821 (≥.95), cfi=.853 (≥.95), pclose=.260 (≥.05). 5 discussion 5.1 construct validity after the seven dimensions had been identified (table 3), they appeared stable throughout the succeeding analyses. in some cases, items associated with the dimensions constructivist approach and learning ability cross-loaded. this phenomenon is conceptually coherent considering the strong correlation between these factors that, in turn, is possibly due to a latent second-level variable. the internal replications by exploratory and confirmatory factor analyses on randomized split halves provide information speaking in favour of the proposed factor model. since the suggested construct holds good or reasonable fit indices and behaves in a consistent manner throughout the different analyses, it can be regarded as holding initial construct validity. initial meaning here that the present study was only a first attempt to launch the hypothesized dimensions, and further research (with new data and adjusted items) is still required to test the generalizability of the construct (finch et al., 2016, pp. 127-128). further testing should also involve a diverse student population with regards to domains and cultural background. 5.2 content validity in general, the factors reflect both original and hypothesized dimensions. in all efa models (tables 3 and 4) as well as in the cfa path diagram (figure 2), the highest loading on each factor occurred on one of the anticipated items. regarding the original dimensions, it turned out that 12 out of the 27 items reflected the anticipated construct and one item loaded differently than in previous studies (see rightmost column in table 1). within the hypothesized scales, 13 items where included in the novel scales. to answer the question if and to which extent the factors actually describe the dimensions, this section presents comments regarding each dimension in the model (table 3, figure 2). the original dimensions retained their original labels, and the labelling of the novel dimensions is commented in section novel dimensions. choices regarding factor model and number of factors were discussed in section an extended set of dimensions. 5.2.1 original dimensions learning ability three of the four items in the learning ability dimension proved stable across most analyses and models, whereas the item k12_9 was dropped at an early stage. in some efa models, this factor attracted items from the constructivist approach dimension, which is consistent with the strong correlation between these dimensions (table 3, figure 2). omniscient authority in her earliest studies, schommer (1990; 1998) reported this dimension as difficult to capture, whereas schraw et al. (2002, p. 267) and moschner et al. (2005) were able to identify this dimension. in this sample, the authority dimension manifested clearly in all models, and the items k09_3f29, k12_5f42 and k14_7 loaded consistently on the omniscient authority factor. structure of knowledge throughout the analyses, most of the structure of knowledge items loaded as expected. item k12_7 often loaded on the learning ability factor but for this item, a connection to learning is not far-fetched; combining information across sources may express an active stance towards learning, rather than a view of the structure of knowledge. thus, this item may be an example of a phrasing containing something that might be called keyword shifting, where the keyword “combining” is perceived differently: some respondents recognize the active learning approach, whereas other see it as an expression for knowledge as bits and pieces that can be combined or kept isolated. during the refinement process, item k12_7 as well as several other items were discarded, leaving four items to represent this dimension. looking at the discarded vs. retained items (table 1) does, however, raise some questions. it seems unfortunate to discard the items k05_7, k06_8, k07_8 and k12_7, since they indeed express a knowledge aspect, i.e. a very clear stance regarding the structure of knowledge as isolated facts vs. information that can or should be combined into larger entities. then again, the retained items seem to focus very much on the actions and behaviour from part of the teacher, almost like introducing a teaching aspect to epistemic beliefs, besides the knowledge and learning aspects. interestingly, five out of the six discarded items (k05_7, k06_8, k07_8, k09_7 and k12_7) stem from the original eq. certainty of knowledge in most models, items k03_8, k04_2f13 and k13_2f44 loaded as expected on the certainty of knowledge factor, but often accompanied by items k04_4f04 and k05_1f15, originally associated with omniscient authority. the items k04_4f04 and k05_1f15 (discarded) may be examples of items with keyword shifting of another kind. here, the respondent may pay attention either to the “experts/ teachers” as authorities, or rather focus on “same answers / same understanding”, the latter option connecting more to knowledge being certain. this observation shows similarity to the eq subset avoid ambiguity loading on simple (structure of) knowledge instead of certain knowledge (cf. schommer, 1990; wood & kardash, 2002, p. 241). it should be mentioned that the discarded items k13_2f44 and k14_2f49 were the ones to show the highest ‘don't know’ portions. 5.2.2 novel dimensions out of the six hypothesized dimensions, three survived the efa and cfa iterations. constructivist approach and internet reliance were retained and learning by dialogue was introduced as the third dimension. the crucial question is, whether the novel dimensions are defensible. do they reflect the constructs, and are the constructs relevant and credible? constructivist approach to learning the novel dimension constructivist approach to learning did not turn out as anticipated but instead, in the proposed model it holds items also from the hypothesized dimensions valuing diversity and reflective nature of learning (table 1). most of the seven items loading on this factor in the initial model (table 3) proved stable throughout the different analyses. in the split half efa, the items k05_2 and k07_2 loaded on learning ability, which is both conceptually coherent as well as understandable considering the strong correlation between these dimensions. to some extent, the constructivist approach can be regarded as an antithesis to omniscient authority; a naïve stance on the omniscient authority dimension would entail a belief that knowledge is handed down by some authority, which implicitly would exclude the possibility of the individual constructing knowledge herself. however, if these dimensions were opposite to each other, they would also correlate negatively, which was not the case. the explanation may lie therein that the items that were used to operationalize the omniscient authority dimension mainly focus on how the respondent relates to authorities and to the knowledge handed down by them. the items do not actually provide information about to which extent the respondent thinks it is possible to construct knowledge. thus, the constructivist approach dimension can rather be regarded as a supplement to the omniscient authority dimension and furthermore, whereas the omniscient authority dimension expresses a knowledge (source) aspect, the constructivist approach dimension expresses a learning (as construction) aspect. learning by dialogue in the initial efa (table 3), this dimension contained items from the hypothesized dimensions valuing diversity (k03_5, k03_7) and connectivist approach (k04_3), all with strong loadings. in the efa on split half a, the item k03_5 caused a heywood case while both other items loaded weakly and moreover, k04_3 loaded on the omniscient authority factor, which is conceptually questionable (table 4). the cfa path diagram on split half b (figure 2) shows rather weak loadings on this latent variable, but the correlation with constructivist approach is strong, which is conceptually coherent. the discarded items, especially k04_5 and k08_5, that were suggested to describe connectivist approach and valuing diversity, might still be worth testing after rephrasing. despite some instability, probably due to low number of cases in the split halves, this dimension can still be defended since it expresses an aspect not expressed in the previous dimensions, namely learning as a social process where the interaction with others, also those representing divergent opinions, is central. internet reliance internet reliance contains three of the five items originally associated to this dimension. some items associated to just-in-time learning might have been associated to this dimension but were discarded due to instability. the items k12_6 and k13_6 proved stable across the analyses, whereas k14_5 loaded weakly and cross-loaded in the initial model and loaded on certainty of knowledge in the split half efa. in the split half cfa, k13_6 caused a heywood case and k14_5 loaded weakly on this dimension. adding a connection from structure of knowledge to k14_5 improved fit indices. the corresponding loading also occurred in the initial efa model (table 3). the fact that iseq items (bråten et al., 2005) were not included to a larger extent may be surprising. however, a closer look shows that the three items included in the present instrument resemble the iseq general internet epistemology items strongly, and basically cover the same topics. the items in this dimension were presented from a naïve perspective and were slightly skewed to the right, indicating that the respondents were not quite as convinced of internet as the digital natives debate may have suggested. 5.3 correlating dimensions the question whether the dimensions correlate or not has been an issue throughout the years within this line of investigation. in her first studies, schommer (1990; 1998) used only varimax rotation and apparently assumed non-correlating dimensions. one might ask if the idea of a set of “more or less independent dimensions” (schommer, 1990) has created an expectation of the dimensions being uncorrelated? schraw et al. (2002, p. 265) analysed their material using both orthogonal and oblique rotation, but concluded that the factors did not correlate. still, their principal component analyses with varimax rotation revealed a weak positive correlation between the omniscient authority and simplicity/structure of knowledge dimensions (p. 269). then again, wood & kardash (2002, p. 252) found moderate to strong inter-factor correlations using factor analysis. wood & kardash (2002, p. 239) also discourage from limiting exploration to orthogonal rotation methods, since forcing inter-correlated factors into an orthogonal model will cause items to cross-load, and the attempt to find a simple structure will fail. otting et al. (2010) identified a relation between expert knowledge (cf. omniscient authority), certainty of knowledge and traditional conceptions of teaching and learning. accordingly, they also identified a relation between learning effort (cf. learning ability) and constructivist conceptions of teaching and learning (cf. the constructivist approach identified in the present study). thus, since the first explorations in the present study indicated that at least some factors correlate, it was obvious that oblique rotation methods should be used (cf. finch et al., 2016, pp. 133, 142; osborne, 2014, pp. 30-33) to allow the factors to correlate, and as it turned out, they did. throughout the analyses (tables 2, 3, 4 and figure 2), the naïvely oriented original dimensions omniscient authority, structure of knowledge and certainty of knowledge correlated with each other. this is in line with the findings by bråten et al. who merged these dimensions into a factor labelled general internet epistemology, but also raises the question if the general internet epistemology factor (bråten et al., 2005; knight et al., 2017) actually suggests a second-level latent variable? across the novel dimensions, correlations occurred between learning by dialogue and constructivist approach, although surprisingly weak. then again, constructivist approach always correlated strongly to learning ability, which was also confirmed in cfa (figure 2). the internet reliance dimension correlated weakly with certainty of knowledge and structure of knowledge (in cfa only with the latter) which may seem surprising but still coherent when taking a closer look at the single items. believing that you can get almost all information about a subject by googling one or two internet sources, and that they can provide you with a clearer picture (than books), will probably go hand in hand with a belief in knowledge being certain and structured. the overall weak correlations to internet reliance may also suggest that this dimension develops “more or less independently” as schommer (1990) originally suggested. the strong correlation between constructivist approach and learning ability is coherent, since believing in everyone’s ability to learn how to learn is part of the constructivist view where the metacognitive component, the learner’s awareness of her/his own learning, is central. the correlation between constructivist approach and learning by dialogue is also coherent. the constructivist approach regards learning as a process of reasoning and construction, where meaning and interpretation is often negotiated in social settings, in dialogue with other learners, and learning is enriched by multiple views and perspectives. the correlation is lower than anticipated, which may be due to learning by dialogue being represented by only three items. to conclude, limiting the efa to orthogonal rotation methods would have concealed the inter-factor relations reported here and perhaps also forced the items to load on inappropriate factors (cf. osborne, 2014, pp. 30-33). 5.4 methodological considerations 5.4.1 scale considerations data and sample have been partly described in the section research data and sample characteristics. due to elimination of cases with a high portion of non-response, the items used for analysis contained between 91.7 and 98.6% substantial responses. since the purpose was to form subscales, it is appropriate to inspect the characteristics and normality of the subscales (cf. carifio & perla, 2007). table 5 comparison of subscale item means, number of items and subscale internal consistencies in fee (moschner et al., 2005), in hypothesized and in proposed model subscales table footnotes a) means and alpha values are calculated excluding six items with reverse phrasing, cf. table 1 b) containing one item also from omniscient authority c) containing items also from reflective nature of learning and valuing diversity d) containing items from connectivist networking and valuing diversity the internal consistencies were analysed both for the hypothesized subscales (54 unidirectional items, table 1) and for the proposed model subscales (26 items, table 3). as illustrated in table 5, the subscales showed large variations; for two of the novel dimensions, the internal consistency index was acceptable. however, for the subscales certainty of knowledge and learning by dialogue, the alpha values were disappointingly low, although not necessarily poor compared to earlier studies (e.g. schommer, 1993; schraw et al., 2002, pp. 266-267; wood & kardash, 2002, p. 253). however, as wood & kardash (2002, p. 237) point out, a low internal consistency value should not too hastily be taken as a motive to discard a subscale. rather, a low value should encourage increasing the number of items and developing them such that they can more precisely express the respondent’s stance on a specific matter. further, as osborne (2014, p. 105) notes, the alpha values (table 5) rather express properties of the sample than properties of the instrument. carifio & perla (2007) recommend 6-8 items for each factor in efa, and the presented model can be criticized for not reaching up to that recommendation. further, the items within each subscale were unidirectional and thus, the lack of reversely phrased items can be criticized (cf. carifio & perla, 2007). however, comparing items within the hypothesized subscales (table 1, still containing bidirectional items) shows that on average, the sophistically oriented items score higher than naïvely oriented items, which indicates that the items measure accurately. as factor analyses often show (e.g. bråten et al., 2005; chiu et al., 2013; leal-soto & ferrer-urbina, 2017; schraw, 2013), the model arrived at in the exploratory procedure is not necessarily identical with the hypothesized conceptual model, regarding neither item set nor factor structure, as was the case here. the results of an efa are not sufficient to confirm a model (osborne, 2014, pp. 19, 49) and therefore, the model arrived at was subject to an internal replication, i.e. efa and cfa on randomized split half data sets (table 4, figure 2). these analyses largely hold the same factor structure as the initial efa model, thereby confirming it. in both split halves, the same inter-factor correlations as in the initial model recurred, which also applies for the cross-loading items k10_7 and k14_5. 5.4.2 data considerations in addition to methodological issues discussed above, the usability and relevance of the current data set (stemming from 2012) should be assessed against the aim of the study and the research question, while taking into account if and to which extent the past years’ technological development has changed the cognitive operating environment. firstly, the aim of the study, as expressed in the research question, was not a validated version of a new instrument but rather, a first exploration of new epistemic dimensions that might contribute to a more nuanced epistemic profile, especially regarding the googling attitude. for this purpose, the data set proved sufficient. secondly, the googling attitude is highly dependent on access to internet and search engines. as reported earlier (section research data and sample characteristics), not much has changed on that point. by 2012, internet penetration within the sample and in finland had long been close to 100% (osf, 2010; osf, 2011), and the majority of the informants had a long history of internet exposure. after 2012, the width of services over mobile devices has undeniably increased beyond browsers and search engines to various applications, probably inducing use habits that rely even more on ubiquity. it is not far-fetched to assume that users today may be even more prone than in 2012 to consult internet-based sources. consequently, if the current research data can demonstrate even weak signs of a googling attitude, then the data fulfils its purpose and one may assume that a newer set of data would reveal even clearer signs. to conclude, the current data set has served the aim and provided an answer to the research question of the current study as for the current sample. as further elaborated in the concluding section, i did not produce a validated instrument. still, the results corroborate the initial assumption about a connection between the googling attitude and epistemic beliefs and encourage further development along this line. should we choose to regard the current results simply as expressing the 2012 state of affairs, the results will still be relevant for historical comparison. 6 conclusions 6.1 dimensions and constructs in the present study, five novel dimensions were introduced and operationalized in 33 items, based both on literature about so-called digital natives and learning in the digital era as well as empirical observations. three novel dimensions, described by thirteen items, survived the process; constructivist approach, internet reliance and learning by dialogue. the dimension constructivist learning approach appeared as a rather stable dimension, correlating strongly with learning ability and moderately with learning by dialogue. these correlations are conceptually coherent, as is the lack of correlation to omniscient authority. the latter suggests that having a constructivist learning approach does not exclude believing that an omniscient authority can be an important source of knowledge but rather, the constructivist learning approach can be regarded as a learning aspect supplementing omniscient authority, describing a knowledge aspect. learning by dialogue was mainly inspired by the connectivist model suggested by siemens (2005; 2006), but during the analyses a picture emerged, where this dimension mainly deals with learning and construction of knowledge as a social process. just as the dimension constructivist learning approach, learning by dialogue provides a learning aspect not captured by previously described dimensions. internet reliance poses a dimension with a knowledge aspect, not covered by previous instruments, and is probably the dimension that most of all expresses the googling attitude referred to in the introduction. furthermore, it expresses a way of relating to knowledge that has not been possible before. indeed, during the pre-internet era it was possible to offload your memory to books or other external media. however, due to access, time and distance barriers, “looking it up in a book” was not an option of the same range as “looking it up on the net” (cf. fisher et al., 2015). thus, since the introduction of internet, it is in fact possible to refrain from memorizing and instead to offload one’s memory and to rely on finding the information on the net, immediately and once you need it, which is not a problem per se. the problems and risks lie in the confusion of knowledge and information, where the ubiquitous access to information creates the illusion of possessing personal knowledge (fisher et al., 2015). technology developing and becoming more powerful accentuates this problem, when not only information storage but also information processing is outsourced, thereby changing our epistemic practices (säljö, 2012; sparrow et al., 2011). the confusion of knowledge and information can also be viewed from the perspective of cognitive processing as described e.g. in the extended version of bloom’s taxonomy (krathwohl, 2002). if a person is to achieve a deeper level of knowing about a topic, the first level, remembering or ‘knowing-that’, is always a prerequisite for moving on to understanding, applying, analysing, evaluating and creating. in this perspective, the googling attitude suggests a ‘knowing-where’ (siemens, 2006, p. 10), which can be regarded as a stage of external information, possibly preceding remembering. however, not until that external information has been memorized and transformed into a ‘knowing-that’ as part of the personal body of information, it can enable the following levels of knowing. 6.2 epistemic awareness and educational practice muis et al. (2006, p. 42) have drawn our attention to that students should be made aware of their epistemic beliefs, since this awareness may be important for epistemic change. the same challenge has recently been addressed by bhatt & mackenzie (2019) but now with focus on the internet context and digital literacy. thus, epistemic awareness is a component in epistemic competence for both teachers and learners. much of the pedagogical potential of the novel dimensions can be deduced from the cross sea between changing pedagogies and the new learning environments emerging with new ict and media. many teaching methods and learning activities, such as the flipped classroom (cf. knewton, 2011) and pbl (cf. otting et al., 2010), increase the demands on students' self-regulation and their ict and media literacy (cf. muis, 2007; brownlee et al., 2009; walker et al., 2009; bhatt & mackenzie, 2019). thus, if a study programme is built e.g. upon pbl, it is useful to know to which extent the students in a new group actually have a constructivist approach and readiness for learning by dialogue, and how to support students’ self-directedness. should it turn out that many students lack these prerequisites, appropriate interventions can be applied to develop their epistemic mind-sets on these dimensions, thereby improving their academic performance. increased understanding regarding both teachers' and students' epistemic beliefs has been called for (cf. palmer & marra, 2008, p. 345). if the novel dimensions can increase awareness regarding the connection between epistemic beliefs and learning tasks over changes in epistemic beliefs by intervention (cf. kienhues et al., 2008), they have the potential of contributing to instruction and learning strategies that are better aligned to both learning objectives and the learners’ epistemic orientations. due to internationalization and student mobility, classes will increasingly hold students and teachers with diverse cultural backgrounds. thus, if epistemic beliefs are dependent on cultural background as suggested by e.g. zhang & watkins (2001) and hofer (2008, pp. 11-12), then awareness about this connection is increasingly important for the teacher to support and guide the learning processes in a multicultural class with students holding diverse, culturally induced, epistemic orientations. the most crucial finding of this study is the introduction of the internet reliance dimension. identifying students with a naïve stance on this dimension may prove important especially if these students are over-reliant towards internet-based resources (cf. bråten, 2008, pp. 369-370). if so, they are at risk of developing an ever-narrowing worldview and an epistemology of ignorance resulting from the ranked and filtered results provided by search engines (cf. bhatt & mackenzie, 2019; hinman, 2008, p. 73; nguyen, 2018). 6.3 future research the results presented above respond to the openly phrased research question by confirming that it is indeed possible to extend epistemic dimensions so that they also express the googling attitude. this is, however, only part of the answer: the novel dimensions need to be further tested e.g. by exploring whether they show between-groups variations congruent with the googling attitude they are expected to express. a connection between epistemic beliefs and academic performance has been suggested (e.g. aditomo, 2018). if the instrument for measuring epistemic beliefs can be developed to measure more precisely, it will probably have a predictive value in assessing each student’s epistemic competence in relation to study context, and a value for teaching practices in supporting students’ epistemic competencies by appropriate choice of learning activities. on this point, the picture is disparate with both encouraging (pieschl et al., 2014) and discouraging (knight et al., 2017) results and thus, epistemic beliefs as predictors of learning behaviour seems an under-researched area. however, net-based learning environments (lms, vle) having started to include learning analytics features will provide better possibilities to investigate the connection between students’ epistemic dimensions and trace data from authentic learning contexts, i.e. courses. there are indicators suggesting that epistemic beliefs dimensions should be measured on a sufficiently fine-grained level, since coarsely composed dimensions as the general internet epistemology (knight et al., 2017), will blur the picture. the study by trautwein & lüdtke (2007), focusing on the certainty dimension, is an interesting initiative in this line. the recent study by bråten, brandmo & kammerer (2018) expresses what we might call increased granularity: besides focusing only on the justification dimension, they divide it into three sub-dimensions, justification by authority, multiple sources and personal knowledge. these examples, together with earlier replication problems (schraw, 2013) expose a challenging tension: should we measure epistemic beliefs as a set of dimensions or as separate constructs? the proposed model arrived at (table 3) and confirmed by internal replication (table 4 and figure 2) shows fit indices that are not ideal but sufficient to encourage further development. despite deficiencies, the model provides an interesting input to the debate whether epistemic beliefs should include only views on knowledge, or also views on learning. the cfa path diagram (figure 2) provides an illustration to this debate: two groups of latent variables, the upper group describing views on learning, and the lower one describing views on knowledge. it is not far-fetched to imagine two second-level latent variables, influencing views on knowledge and views on learning, respectively (cf. section correlating dimensions). the correlations within the two groups of latent variables, especially the strong correlation between constructivist approach and learning ability, also point in this direction, and exploring second-level latent constructs is a topic for further investigation. topics dealing with the instrument itself include 1) developing the instrument such that each dimension would be represented by more than only three items (cf. carifio & perla, 2007), 2) improving items with low loadings, and 3) exploring the discarded items regarding common features that might have contributed to their dysfunctionality. in addition to these topics, the functionality of the model should be tested by exploring how well the dimensions distinguish different learners. this will be done by exploring if and to which extent dimensional group differences can be identified e.g. across users representing different digital orientations or study domains. the extensions to the epistemic beliefs instrument and the proposed (but not validated) model are, needless to say, only a beginning. considering the twenty years of history with seq and its successors gives an idea of the work that still lies ahead. keypoints the novel dimension internet reliance may help in identifying learners that are over-reliant towards internet-based resources. beliefs about learning contribute to describing one’s epistemic orientation, although they are not regarded as part of the epistemic beliefs concept. although assumed to develop independently, the epistemic beliefs dimensions correlate when using an appropriate rotation method. the novel dimensions contribute to an epistemic awareness and to adapting instruction and learning practices to learners’ epistemic orientations. an increasingly international learning context and multicultural student body requires awareness about culturally induced epistemic orientations. acknowledgments this research was funded by föreningen konstsamfundet, koulutusrahasto and svenska kulturfonden. i am grateful for the support provided by arcada through filip levälahti and all participating students during the data collection process, and for the support from the meda project through matteo stocchetti. i am grateful also for the feedback provided by my supervisors marita mäkelä and eero sormunen, and by anonymous reviewers to earlier drafts of this paper. references aditomo, a. (2018). epistemic beliefs and academic performance across soft and hard disciplines in the first year of college. journal of further and higher education, 42(4), 482-496. doi:10.1080/0309877x.2017.1281892 alexander, p. a. (2006). what would dewey say? channeling dewey on the issue of specificity of epistemic beliefs: a response to muis, bendixen, and haerle (2006). educational psychology review, 18(1), 55-65. doi:10.1007/s10648-006-9002-7 anderson, s., & balsamo, a. (2008). a pedagogy for original synners. in t. mcpherson (ed.), digital young, innovation, and the unexpected (1st ed., pp. 241-259). cambridge, ma: the mit press. doi:10.1162/dmal.9780262633598.241 bernholt, a., gruber, h., & moschner, b. (eds.). (2017). wissen und lernen. wie epistemische überzeugungen schule, universität und arbeitswelt beeinflussen [knowing and learning. the influence of epistemic beliefs on schools, universities and working life]. münster: waxmann verlag. bhatt, i., & mackenzie, a. (2019). just google it! digital literacy and the epistemology of ignorance. teaching in higher education, 24 (3), 302-317. doi:10.1080/13562517.2018.1547276 bråten, i. (2008). personal epistemology, understanding of multiple texts, and learning within internet technologies. in m. s. khine (ed.), knowing, knowledge and beliefs: epistemological studies across diverse cultures (pp. 351-376). dordrecht: springer. doi:10.1007/978-1-4020-6596-5_17 bråten, i., brandmo, c., & kammerer, y. (2018). a validation study of the internet-specific epistemic justification inventory with norwegian preservice teachers. journal of educational computing research, (onlinefirst), 1-24. doi:10.1177/0735633118769438 bråten, i., & strømsø, h. i. (2006). epistemological beliefs, interest, and gender as predictors of internet-based learning activities. computers in human behavior, 22(6), 1027-1042. doi:10.1016/j.chb.2004.03.026 bråten, i., strømsø, h. i., & samuelstuen, m. s. (2005). the relationship between internet-specific epistemological beliefs and learning within internet technologies. journal of educational computing research, 33(2), 141-171. doi:10.2190%2fe763-x0ln-6nmf-cb86 brownlee, j., walker, s., lennox, s., exley, b., & pearce, s. (2009). the first year university experience: using personal epistemology to understand effective learning and teaching in higher education. higher education, 58(5), 599-618. doi:10.1007/s10734-009-9212-2 cape, p. (2010). (2010). questionnaire length, fatigue effects and response quality revisited. paper presented at the re:think 2010: the arf 56th annual convention, ny. retrieved from https://www.surveysampling.com/ carifio, j., & perla, r. j. (2007). ten common misunderstandings, misconceptions, persistent myths and urban legends about likert scales and likert response formats and their antidotes. journal of social sciences, 3(3), 106-116. doi:10.3844/jssp.2007.106.116 chiu, y., liang, j., & tsai, c. (2013). internet-specific epistemic beliefs and self-regulated learning in online academic information searching. metacognition and learning, 8(3), 235-260. doi:10.1007/s11409-013-9103-x clark, a., & chalmers, d. (1998). the extended mind. analysis, 58(1), 7-19. doi:10.1093/analys/58.1.7 debacker, t. k., crowson, h. m., beesley, a. d., thoma, s. j., & hestevold, n. l. (2008). the challenge of measuring epistemic beliefs: an analysis of three self-report instruments. journal of experimental education, 76(3), 281-312. doi:10.3200/jexe.76.3.281-314 downes, s. (2007). an introduction to connective knowledge. paper presented at the media, knowledge and education: exploring new spaces, relations and dynamics in digital media ecologies, ed. t. hug, innsbruck university press, innsbruck, austria, 2007, june 25-26, pp. 77-102. elby, a. (2009). defining personal epistemology: a response to hofer & pintrich (1997) and sandoval (2005). journal of the learning sciences, 18(1), 138-149. doi:10.1080/10508400802581684 fabrigar, l. r., wegener, d. t., maccallum, r. c., & strahan, e. j. (1999). evaluating the use of exploratory factor analysis in psychological research. psychological methods, 4(3), 272-299. ferguson, r. (2012). learning analytics: drivers, developments and challenges. international journal of technology enhanced learning, 4(5), 304-317. doi:10.1504/ijtel.2012.051816 finch, w. h., immekus, j. c., & french, b. f. (2016). applied psychometrics using spss and amos. charlotte, nc: information age publishing inc. fisher, m., goddu, m. k., & keil, f. c. (2015). searching for explanations: how the internet inflates estimates of internal knowledge. journal of experimental psychology, general, 144(3), 674-687. doi:10.1037/xge0000070 fokkema, m., & greiff, s. (2017). how performing pca and cfa on the same data equals trouble. european journal of psychological assessment, 33(6), 399-402. doi:10.1027/1015-5759/a000460 greene, j. a., azevedo, r., & torney-purta, j. (2008). modeling epistemic and ontological cognition: philosophical perspectives and methodological directions. educational psychologist, 43(3), 142-160. doi:10.1080/00461520802178458 greene, j. a., sandoval, w. a., & bråten, i. (eds.). (2016). handbook of epistemic cognition. new york: routledge. doi:10.4324/9781315795225. retrieved from http://ebookcentral.proquest.com/ grossnickle peterson, e., alexander, p. a., & list, a. (2017). the argument for epistemic competence. in a. bernholt, h. gruber & b. moschner (eds.), wissen und lernen. wie epistemische überzeugungen schule, universität und arbeitswelt beeinflussen (pp. 255-270). münster: waxmann verlag. gunter, b., rowlands, i., & nicholas, d. (2009). the google generation. are ict innovations changing information-seeking behaviour? cambridge: chandos publishing. retrieved from http://ebookcentral.proquest.com/ hair, j. f., black, w. c., babin, b. j., & anderson, r. e. (2010). multivariate data analysis: a global perspective (7th ed.). upper saddle river (n.j.): prentice hall. head, a. j., & eisenberg, m. b. (2010). how today’s college students use wikipedia for course-related research. first monday, 15(3). doi:10.5210/fm.v15i3.2830 hinman, l. m. (2008). searching ethics: the role of search engines in the construction and distribution of knowledge. in a. spink, & m. zimmer (eds.), web search multidisciplinary perspectives (pp. 67-76). berlin, heidelberg: springer. doi:10.1007/978-3-540-75829-7_3. retrieved from https://ebookcentral.proquest.com/ hofer, b. k. (2006). beliefs about knowledge and knowing: integrating domain specificity and domain generality: a response to muis, bendixen, and haerle (2006). educational psychology review, 18(1), 67-76. doi:10.1007/s10648-006-9000-9 hofer, b. k. (2008). personal epistemology and culture. in m. s. khine (ed.), knowing, knowledge and beliefs: epistemological studies across diverse cultures (pp. 3-22). dordrecht: springer. doi:10.1007/978-1-4020-6596-5_1 hofer, b. k., & pintrich, p. r. (1997). the development of epistemological theories: beliefs about knowledge and knowing and their relation to learning. review of educational research, 67(1), 88-140. doi:10.2307/1170620 hofer, b. k., & pintrich, p. r. (eds.). (2002). personal epistemology: the psychology of beliefs about knowledge and knowing . mahwah, n.j: l. erlbaum associates. doi:10.4324/9781410604316 hooper, d., coughlan, j., & mullen, m. (2008). structural equation modelling: guidelines for determining model fit. electronic journal of business research methods, 6(1), 53-60. retrieved from http://www.ejbrm.com/volume6/issue1/p53 jones, c., & hosein, a. (2010). profiling university students' use of technology: where is the net generation divide? international journal of technology, knowledge & society, 6 (3), 43-58. doi:10.18848/1832-3669/cgp/v06i03/56097 kammerer, y., & gerjets, p. (2012). effects of search interface and internet-specific epistemic beliefs on source evaluations during web search for medical information: an eye-tracking study. behaviour & information technology, 31(1), 83-97. doi:10.1080/0144929x.2011.599040 karimi, m. n. (2014). efl students' grammar achievement in a hypermedia context: exploring the role of internet-specific personal epistemology. system, 42, 1-11. doi:10.1016/j.system.2013.10.017 khine, m. s. (ed.). (2008). knowing, knowledge and beliefs: epistemological studies across diverse cultures . dordrecht: springer. doi:10.1007/978-1-4020-6596-5. retrieved from https://ebookcentral.proquest.com/ kienhues, d., bromme, r., & stahl, e. (2008). changing epistemological beliefs: the unexpected impact of a short-term intervention. british journal of educational psychology, 78(4), 545-565. doi:10.1348/000709907x268589 kitchener, r. f. (2002). folk epistemology: an introduction. new ideas in psychology, 20(2–3), 89-105. doi:10.1016/s0732-118x(02)00003-x knewton. (2011). the flipped classroom infographic. retrieved from http://www.knewton.com/flipped-classroom/, 03.08.2012 knight, s., rienties, b., littleton, k., mitsui, m., tempelaar, d., & shah, c. (2017). the relationship of (perceived) epistemic cognition to interaction with resources on the internet. computers in human behavior, 73, 507-518. doi:10.1016/j.chb.2017.04.014 knight, s., wise, a. f., & chen, b. (2017). time for change: why learning analytics needs temporal analysis. journal of learning analytics, 4(3), 7–17. doi:10.18608/jla.2017.43.2 krathwohl, d. r. (2002). a revision of bloom's taxonomy: an overview. theory into practice, 41(4), 212-218. doi:10.1207/s15430421tip4104_2 kuhn, d., & weinstock, m. (2002). what is epistemological thinking and why does it matter? in b. k. hofer, & p. r. pintrich (eds.), personal epistemology: the psychology of beliefs about knowledge and knowing (pp. 121-144). mahwah, n.j: l. erlbaum associates. leal-soto, f., & ferrer-urbina, r. (2017). three-factor structure for epistemic belief inventory: a cross-validation study. plos one, 12(3), 1-16. doi:10.1371/journal.pone.0173295 martin, e. (2005). survey questionnaire construction. in k. kempf-leonard (ed.), encyclopedia of social measurement (pp. 723-732). new york: elsevier. doi:10.1016/b0-12-369398-5/00433-3. retrieved from http://www.sciencedirect.com/ miller, b., & record, i. (2013). justified belief in a digital age: on the epistemic implications of secret internet technologies. episteme, 10(02), 117-134. doi:10.1017/epi.2013.11 moschner, b., gruber, h., & studienstiftungsarbeitsgruppe epi. (2005). epistemologische überzeugungen. forschungsbericht nr. 18. regensburg: universität regensburg, lehrstuhl für lehr-lern-forschung. retrieved from https://portal.uni-regensburg.de/48/ muis, k. r. (2007). the role of epistemic beliefs in self-regulated learning. educational psychologist, 42(3), 173-190. doi:10.1080/00461520701416306 muis, k. r., bendixen, l. d., & haerle, f. c. (2006). domain-generality and domain-specificity in personal epistemology research: philosophical and empirical reflections in the development of a theoretical framework. educational psychology review, 18(1), 3-54. doi:10.1007/s10648-006-9003-6 nguyen, c. t. (2018). echo chambers and epistemic bubbles. episteme, (firstview, sept 13, 2018). doi:10.1017/epi.2018.32 niessen, t., vermunt, j., abma, t., widdershoven, g., & van der vleuten, c. (2004). on the nature and form of epistemologies: revealing hidden assumptions through an analysis of instrument design. european journal of school psychology, 2(1-2), 39-64. osborne, j. w. (2014). best practices in exploratory factor analysis. retrieved from http://pareonline.net/ osf. (2010). use of information and communications technology. official statistics finland. retrieved from http://www.stat.fi/til/sutivi/2010/sutivi_2010_2010-10-26_tie_001_en.html osf. (2011). use of information and communications technology by individuals. official statistics finland. retrieved from http://www.stat.fi/til/sutivi/2011/sutivi_2011_2011-11-02_tie_001_en.html otting, h., zwaal, w., tempelaar, d., & gijselaers, w. (2010). the structural relationship between students' epistemological beliefs and conceptions of teaching and learning. studies in higher education, 35(7), 741-760. doi:10.1080/03075070903383203 palmer, b., & marra, r. m. (2008). individual domain-specific epistemologies: implications for educational practice. in m. s. khine (ed.), knowing, knowledge and beliefs: epistemological studies across diverse cultures (pp. 325-350). dordrecht: springer. doi:10.1007/978-1-4020-6596-5_16 perry, w. g. (1970). forms of intellectual and ethical development in the college years: a scheme . new york: holt, rinehart and winston. personal data act of 22.4.1999. retrieved from http://www.finlex.fi/fi/laki/ajantasa/1999/19990523 pieschl, s., stallmann, f., & bromme, r. (2014). high school students' adaptation of task definitions, goals and plans to task complexity the impact of epistemic beliefs. psychological topics, 23(1), 31-52. doi:10.31820/pt prensky, m. (2001). digital natives, digital immigrants part 1. on the horizon, 9(5), 1-6. doi:10.1108/10748120110424816 purcell, k., brenner, j., & rainie, l. (2012). search engine use 2012. washington, dc: pew research center. retrieved from http://www.pewinternet.org/2012/03/09/search-engine-use-2012/ purcell, k., rainie, l., heaps, a., buchanan, j., friedrich, l., jacklin, a., . . . zickuhr, k. (2012). how teens do research in the digital world. washington dc: pew research center. retrieved from http://www.pewinternet.org/2012/11/01/how-teens-do-research-in-the-digital-world/ säljö, r. (2012). literacy, digital literacy and epistemic practices: the co-evolution of hybrid minds and external memory systems. nordic journal of digital literacy, 7(1), 5-19. retrieved from http://www.idunn.no/ts/dk/2012/01/art08 sandoval, w. a. (2005). understanding students' practical epistemologies and their influence on learning through inquiry. science education, 89(4), 634-656. doi:10.1002/sce.20065 schommer, m. (1990). effects of beliefs about the nature of knowledge on comprehension. journal of educational psychology, 82(3), 498-504. doi:10.1037/0022-0663.82.3.498 schommer, m. (1993). epistemological development and academic performance among secondary students. journal of educational psychology, 85 (3), 406-411. doi:10.1037/0022-0663.85.3.406 schommer, m. (1998). the influence of age and education on epistemological beliefs. british journal of educational psychology, 68(4), 551-562. doi:10.1111/j.2044-8279.1998.tb01311.x schommer-aikins, m. (2004). explaining the epistemological belief system: introducing the embedded systemic model and coordinated research approach. educational psychologist, 39(1), 19-29. doi:10.1207/s15326985ep3901_3 schraw, g. (2013). conceptual integration and measurement of epistemological and ontological beliefs in educational research. isrn education, vol. 2013, 1-19. doi:10.1155/2013/327680 schraw, g., bendixen, l., & dunkle, m. e. (2002). development and validation of the epistemic belief inventory (ebi). in b. k. hofer, & p. r. pintrich (eds.), personal epistemology: the psychology of beliefs about knowledge and knowing (pp. 261-275). mahwah, n.j: l. erlbaum associates. schreiber, j. b., nora, a., stage, f. k., barlow, e. a., & king, j. (2006). reporting structural equation modeling and confirmatory factor analysis results: a review. the journal of educational research, 99(6), 323-337. doi:10.3200/joer.99.6.323-338 shaffer, d. w., & clinton, k. a. (2006). toolforthoughts: reexamining thinking in the digital age. mind, culture, and activity, 13(4), 283-300. doi:10.1207/s15327884mca1304_2 siemens, g. (2005). connectivism: a learning theory for the digital age. international journal of instructional technology and distance learning, 2 (1), 3-10. siemens, g. (2006). knowing knowledge, elearnspace. retrieved from http://www.elearnspace.org/ simpson, t. w. (2012). evaluating google as an epistemic tool. metaphilosophy, 43(4), 426-445. doi:10.1111/j.1467-9973.2012.01759.x smart, p. r. (2012). the web-extended mind. metaphilosophy, 43(4), 446-463. doi:10.1111/j.1467-9973.2012.01756.x sparrow, b., liu, j., & wegner, d. m. (2011). google effects on memory: cognitive consequences of having information at our fingertips. science, 333(6043), 776-778. doi:10.1126/science.1207745 spss. (2016a). amos 24.0 [computer software]. chicago, il: spss inc., ibm corporation. spss. (2016b). spss 24.0 [computer software]. chicago, il: spss inc., ibm corporation. ståhl, t. (2017). how ict savvy are digital natives actually? nordic journal of digital literacy, 12(3), 89-108. doi:10.18261/issn.1891-943x-2017-03-04 ståhl, t., & mildén, p. (2017). applying the fee to explore epistemic beliefs among students. in a. bernholt, h. gruber & b. moschner (eds.), wissen und lernen. wie epistemische überzeugungen schule, universität und arbeitswelt beeinflussen (pp. 59-97). münster: waxmann verlag strømsø, h. i., & bråten, i. (2010). the role of personal epistemology in the self-regulation of internet-based learning. metacognition and learning, 5(1), 91-111. doi:10.1007/s11409-009-9043-7 tang, j. (2010). exploratory and confirmatory factor analysis of epistemic beliefs questionnaire about mathematics for chinese junior middle school students. journal of mathematics education, 3(2), 89-105. tourangeau, r., couper, m. p., & conrad, f. (2004). spacing, position, and order: interpretive heuristics for visual features of survey questions. public opinion quarterly, 68(3), 368-393. doi:10.1093/poq/nfh035 trautwein, u., & lüdtke, o. (2007). epistemological beliefs, school achievement, and college major: a large-scale longitudinal study on the impact of certainty beliefs. contemporary educational psychology, 32(3), 348-366. doi:10.1016/j.cedpsych.2005.11.003 van den beemt, a., akkerman, s., & simons, p. r. j. (2011). patterns of interactive media use among contemporary youth. journal of computer assisted learning, 27(2), 103-118. doi:10.1111/j.1365-2729.2010.00384.x walker, s., brownlee, j., lennox, s., exley, b., howells, k., & cocker, f. (2009). understanding first year university students: personal epistemology and learning. teaching education, 20(3), 243-256. doi:10.1080/10476210802559350 wood, p., & kardash, c. (2002). critical elements in the design and analysis of studies of epistemology. in b. k. hofer, & p. r. pintrich (eds.), personal epistemology: the psychology of beliefs about knowledge and knowing (pp. 231-260). mahwah, n.j: l. erlbaum associates. zhang, l., & watkins, d. (2001). cognitive development and student approaches to learning: an investigation of perry's theory with chinese and u.s. university students. higher education, 41(3), 239-261. doi:10.1023/a:1004151226395 codepen 3.jenert frontline learning research special issue vol.9 no.2 (2021) 50 77 issn 2295-3159 the interplay of personal and contextual diversity during the first year at higher education: combining a quantitative and a qualitative approach tobias jenert1 & taiga brahm2 1paderborn university, germany 2university of tübingen, germany article received 13 may 2020/ revised 30 november/ accepted 2 decemeber/ available online 12 march 2021 abstract research on student transition into higher education (he) has taken different theoretical perspectives. first, studies investigated personal variables such as students´ self-efficacy, emotions and motivation regarding the transition from school to he. a second strand of research focused on contextual variables, for instance college effectiveness research. with this paper, we combine both the personal and the contextual approach. we aim to investigate the interaction between personal and contextual diversity during the transition into he, taking into account students’ diversity in particular with regard to gender and individual characteristics, such as self-efficacy. we explored the heterogeneity in students’ personal characteristics by conducting a latent profile analysis (lpa) based on students’ intrinsic motivation, self-efficacy and anxiety before entering higher education. lpa resulted in three distinct profiles, with significant differences in how students perceived the first year. this finding suggests that students’ personal characteristics when entering higher education influence how they experience the study environment. to investigate the interplay between individual and contextual differences in more detail, we conducted a qualitative longitudinal study with 14 first-year students in parallel with the panel survey. we found that individual students react very differently to specific characteristics and events of the first-year environment. our study adds to the growing body of research that aims to grasp the complexity of interactions between individual and contextual differences. specifically, we illustrate how combining quantitative and qualitative methods can provide new insights into person-context interactions. keywords: transition, longitudinal study, latent profile analysis, quantitative-qualitative, longitudinal info corresponding author email: tobias.jenert@uni-paderborn.de doi: https://doi.org/10.14786/flr.v9i2.669 1. introduction for many students, entering he is a decisive moment in their life with implications that reach far beyond merely changing the educational institution, their hometown, or country. often, students in he at the same time stop living with their parents, having to adjust to a lifestyle and to a social environment that is fundamentally different from what they have known before. to some extent, it is even necessary to disconnect from their previous environment, e.g. loosening the ties to parents and friends from home (tinto, 1993). a successful transition requires students to develop an identity and a sense of belonging to the new socio-cultural context of he (perry & allard, 2003). concerning the academic requirements of he, students have to adapt or develop their learning strategies to respond to various challenges, such as greater learner autonomy or higher amounts of content to be mastered (coertjens, donche, maeyer, van daal, & van petegem, 2017; donche, coertjens, & van petegem, 2010; donche, maeyer, coertjens, van daal, & van petegem, 2013). consequently, many students experience the transition into he as a shock that may impede their academic success and even lead to dropping out even though they have the intellectual ability to master the academic requirements (briggs, clark, & hall, 2012; dyson & renk, 2006; gale & parker, 2012; kuh, cruce, shoup, & kinzie, 2008; leese, 2010; tinto, 1993). research on student transition has been mainly conducted from two theoretical perspectives. first, studies investigated personal variables such as students’ self-efficacy, emotions, and motivation regarding the transition from school to he. for example, students with high self-efficacy find it easier to master the challenges of developing their identities as learners within the new educational context, are more motivated to learn, show better performance (fenollar, román, & cuestas, 2007; hsieh, sullivan, & guerra, 2007; lau, liem, & nie, 2008; martin, colmar, davey, & marsh, 2010; prat-sala & redford, 2010). academic emotions, particularly study-related anxiety, affect performance and retention (pekrun, elliot, & maier, 2009; pekrun, goetz, titz, & perry, 2002; pekrun, götz, & perry, 2005; pekrun & linnenbrink-garcia, 2012; villavicencio & bernardo, 2013). both self-efficacy and students’ emotions are closely linked to student motivation, which influences how students approach academic tasks. the second strand of research investigates the transition to he by looking at contextual characteristics of study environments. these contextual factors broadly fall into two categories, academic and social (chapman & pascarella, 1983; tinto, 1993). academic factors include requirements and challenges such as exams as well as resources such as communication with teaching faculty. social factors are the quality of students’ social relationships with each other as well as access to peer networks as resources for coping with challenges (nevill & rhodes, 2004; rocconi, 2011). previous research primarily established the effects of specific characteristics of the academic and social study contexts on student performance and retention, showing e.g. general positive effects of student learning communities (ibid.). in comparison, little is known about how subgroups of students, distinguished by sets of personal characteristics, perceive and interact with features of the academic and social study environment. both the personal and the contextual approach to researching the transition to he point to the importance of acknowledging student diversity when investigating transition processes. depending on their personal prerequisites, students respond differently to he contexts, leading to very individual experiences of the transition. hitherto, personal diversity and contextual diversity have been mostly investigated separately from each other, however, limiting our understanding of the interactions between these two dimensions. recently, researchers are increasingly studying the role of students’ diversity in transition processes, e.g. by using longitudinal research (kyndt et al., 2015; kyndt, donche, van daal, gijbels, & van petegem, 2019) and profile analysis (de clercq, galand, & frenay, 2020; martens & metzger, 2017). likewise, research has addressed contextual diversity at the meso-level, i.e. the design of study environments and its effects on different students’ educational experiences (duchatelet & donche, 2019). such research suggests that different bodies of students may react in very distinctive ways to elements in their study environment. building on social-cognitive theory (bandura, 1989), the research presented in this paper integrates the personal and the contextual approach to develop a more detailed picture of student transition from school to he. we address two main goals: first, we aim to identify subgroups in the first-year population of a swiss business school, using study-related motivation, anxiety, and self-efficacy as grouping variables. second, we aim to specify which contextual aspects of the study environment are particularly relevant for shaping students’ experiences of the transition to he and how students with different personal characteristics interact differently with these contextual features. to achieve these goals, we conducted two studies in parallel: first, a longitudinal panel study investigating the development of students’ self-efficacy, anxiety, and motivation during the first year in he. we found that these constructs developed negatively throughout the first year of study (brahm, jenert, & wagner, 2017). based on the distribution of the data, we assumed that, depending on their individual dispositions, students interact differently with the study environment they encounter. to test this hypothesis and to better understand such diverse developments, we identified subgroups using latent profile analysis. second, in parallel to the quantitative panel survey, we conducted a longitudinal interview study with students sampled from the panel cohort. the interview questions and the qualitative analysis of the interviews were based on the same theoretical framework as the panel study. our results point to various dimensions of student and contextual diversity that influence how and how well students manage the transition process. our research contributes to the scholarly discourse on student transition as it provides a more fine-grained view on the different dimensions of diversity that influence how students experience the transition to he. it complements the growing number of studies using longitudinal and profile analysis (see above). furthermore, our results support the notion that theoretical assumptions on the relationships between contextual and personal variables may need reconsidering when student diversity is taken into account (de clercq et al., 2020). 2. literature review over the last decades, research has produced a plethora of variables, which impact student success and retention in he (schneider & preckel, 2017). our research aims to combine personal and contextual factors influencing the transition to he. this approach is rooted in social-cognitive theory (bandura, 1989) which conceives of human agency as an interaction between personal factors such as self-efficacy, environmental factors such as perceived support or competition, and behavioural factors such as expected outcomes. consequently, in our review of personal variables, we focused on constructs that supposedly influence how students interact with their study environment. following this rationale, we investigated study-related self-efficacy, anxiety, and motivation in our survey-based panel study. from a social-cognitive perspective, all three constructs can be described as personal dispositions that influence how individuals perceive and act on their environment. moreover, all three constructs have already been investigated regarding the transition to he. concerning contextual characteristics, we applied a broader approach to exploit the potential of the in-depth interviews in the qualitative study. research on student transition has distinguished between academic and social integration as two important dimensions of transitioning into he (chapman & pascarella, 1983; tinto, 1993). consequently, we focused our interviews on characteristics of the study environment that related to academic challenges and support as well as to social aspects such as peer interaction. 2.1 personal variables the first personal variable relevant for dealing with the challenges of the transition to he is students’ academic self-efficacy (robbins et al., 2004; talsma, schüz, schwarzer, & norris, 2018). self-efficacy refers to students’ judgements about their capabilities to fulfil the performance expectations across different task activities. there is ample evidence that self-efficacy has positive effects on studying. a systematic review by honicke and broadbent (2016) shows moderate correlations between self-efficacy and academic performance. in a meta-analytic panel analysis, talsma et al. (2018) established a reciprocal relationship between self-efficacy and academic performance. the level of self-efficacy predicted future academic performance, and, at the same time, performance affected the development of self-efficacy. regarding students’ transition processes, a study with 192 students suggested that “high self-efficacy was related to better college adjustment“ (ramos-sánchez & nichols, 2007, p. 6). furthermore, an increasing number of studies points to self-efficacy as an important resilience factor for disadvantaged groups such as so-called non-traditional or at-risk students in he. in a study comparing first-generation to non-first-generation students, aymans and kauffeld (2015) found that for both groups, a higher level of self-efficacy was associated with a reduced dropout risk. research on the development of self-efficacy in he reveals complex interrelations with student characteristics and learning environments. in a longitudinal study with more than 600 students, duchatelet and donche (2019) found that autonomy-supporting learning environments can foster self-efficacy, thus, establishing a link between study context and personal development. overall, findings in the he context are in line with bandura’s (1989) social-cognitive theory which states that self-efficacy can be fostered by experiencing mastery and, at the same time, is a prerequisite for agency. students who feel self-efficacious are more confident about their abilities and may feel less threatened by he contexts. this makes them more agentic and resilient against challenges, improving their performance, and supporting their self-efficacy even further. thus, we consider self-efficacy a key construct for understanding the relations between personal and contextual factors during he transitions. second, in contrast to self-efficacy, students’ anxiety negatively affects their academic performance (mellanby & zimdars, 2011; zeidner, 1998) and their transition to higher education (christie, tett, cree, hounsell, & mccune, 2008). hailikari, kordts-freudinger, and postareff (2014) found that during their first study year students experienced satisfaction and enthusiasm, but they more frequently reported dissatisfaction, confusion, and anxiety. the study also found a positive relationship between the absence of negative emotions and study progress as well as achievement. a recent study with 233 first-year students in new zealand confirms the negative effects of state anxiety on students’ grades and self-efficacy. at the same time, this research found correlations between students’ individual learning strategies and their experiences of anxiety and self-efficacy in exam situations (sotardi & brogt, 2019). further investigation revealed that the relationships between anxiety and performance are influenced by the type of assessment (test or essay assignment) (sotardi, bosch, & brogt, 2020). this finding hints at the interaction between students’ personal and contextual diversity, i.e. a specific kind of anxiety being triggered by specific kinds of exams. much like self-efficacy, anxiety can thus be regarded as a personal factor that is actualized through specific situations. as the transition to he confronts students with many challenging situations such as exams, uncertain performance expectations, or unfamiliar social conventions, anxiety is an important construct for understanding differences in students’ interactions with the study environment. in our quantitative research, we are interested in students’ general levels of anxiety throughout the first study year. finally, we consider motivation to be an important construct for understanding how different students interact with study contexts. intrinsic motivation is associated with proactive student behaviour such as using deep approaches to studying (byrne & flood, 2005). the development of students’ motivation during the transition from school to he has been repeatedly investigated. most studies report negative developments throughout university studies (busse, 2013; jacobs & newstead, 2000; lau et al., 2008; lieberman & remedios, 2007; martin et al., 2010; pan & gauvain, 2012), while few report an increase (e.g. ratelle, guay, larose, & senécal, 2004). in a longitudinal study with measurements towards the end of secondary school and at the beginning of he, kyndt et al. (2015) found a sharp increase in autonomous motivation at the point where students entered university. once in he, however, this growth reduced again significantly. thus, it seems that students’ expectations of what he will be like are more motivating than their actual experiences. martens and metzger (2017) distinguished different subpopulations of students based on their motivation and related them with different elements of the integrated model of learning and action. they found that self-determined motivation was associated with desirable learning behaviour such as persistent goal pursuit, acceptance of responsibility, and experiences of success. other subgroups showed more differentiated profiles. for example, students with anxious learning motivation, mostly trying to avoid negative consequences, scored high in responsibility acceptance, but showed little sensitive coping strategies. the authors state that “the imbalance of high motivation and low intention will be most probably experienced as anxiety” (martens & metzger, 2017, p. 41). a comparison between students of business economics and educational sciences showed differences in the occurrence, the distribution, and the respective patterns of the subgroups between study programs (ibid.). this, again, suggests that motivation is an important construct for understanding diversity among students’ personal characteristics as well as their interaction with study contexts. this is further supported by noyens, donche, coertjens, van daal, and van petegem (2019) who showed that amotivation negatively impacts students’ social integration during their first year at university. 2.2 contextual factors as we have argued above, the extent of self-efficacy, anxiety, and motivation that students exhibit is related to their experience of the study environment. in the qualitative part of our research, we aim to uncover, which concrete features of the first-year study environment shape students’ transition experience. in particular, we strive to understand better, how students with different personal characteristics differ in their perceptions and interactions with these relevant contextual features. following the qualitative paradigm, we kept the interview study more open than the quantitative study. developing the interview questions, however, we used the well-established distinction between academic and social integration during he transition. concerning academic factors, research found interactions between students and faculty to be particularly important for first-year students’ experience, performance, and persistence. in a survey of 530 first-year students, the first impressions of the university staff mattered most for their first-year experience followed by their satisfaction with university life (meehan & howells, 2018). various studies report significant effects of student-faculty interactions on performance indicators such as growth in knowledge and academic adjustment (delaney, 2008; kuh & hu, 2001). kim and lundberg (2016) modeled the relationships between the extent of student-faculty interactions, students’ behavior, and their intellectual development. they found that this interaction fostered students’ engagement in class. this relationship was moderated by the levels of academic self-challenge and sense of belonging. in their qualitative account, cotten and wilson (2006) distinguish between different kinds of student-faculty interactions. they report that, generally, student-faculty interactions outside formal settings are scarce and that students fear negative effects from approaching faculty. this may suggest that students with a lower general level of anxiety and higher self-efficacy may also be able to benefit more from faculty interaction. another factor relating to academic contexts are performance requirements in general. obviously, exams play an important role in this regard (cassady & johnson, 2002). beyond that, qualitative findings indicate that students’ anxiety is caused by “not knowing what is expected” (christie et al., 2008, p. 569), i.e. uncertainty about academic performance expectations. building on these findings, we can hypothesize that students with higher levels of self-efficacy may feel less uncertain and anxious about performance requirements. concerning social aspects, student-peer interactions and particularly learning communities can have positive effects by connecting students with peers (rocconi, 2011; thompson & mazer, 2009). for instance, walsh, larsen & parry (2009) identified peers as the most frequent source of information and support. zander, brouwer, jansen, crayen, and hannover (2018) found that integrating students in organized learning communities also helps them find social support networks. social support from peers may be particularly beneficial in stressful situations, such as the transition to university (ahern et al., 2006; thomas, 2002). however, peer communities can also have segregation effects because students tend to group with others who share similar achievement levels (brouwer, flache, jansen, hofman, & steglich, 2018). also, some students may find it difficult to socialize away from the learning environment (riordan & carey, 2019) which might also be due to the varying responsibilities students face (e.g. balancing work or study-related demands). again, this may be a hint that students profit more or less from peer interactions, depending on how efficacious, anxious, and motivated they are to engage with others. summarizing, there is an increasing number of studies that aim to identify systematic differences within student populations by identifying subgroups defined by some personal variables (e.g. de clercq et al., 2020; duchatelet & donche, 2019; martens & metzger, 2017; sotardi et al., 2020). in this regard, self-efficacy, anxiety, and motivation have proved to be appropriate grouping variables. in contrast, only few studies (e.g. de clercq et al., 2020) have investigated how different subgroups of students perceive their academic and social environment of the first year in he. thus, in our research, we apply a quantitative analysis to identify subgroups of students as defined by their self-efficacy, anxiety, and motivation while the parallel qualitative study helps us to understand better how such personal differences affect students’ interactions with their study environment. 3. the studies: methods, samples and data analysis 3.1 panel study on student transition the research was conducted as a longitudinal study on students’ self-efficacy, anxiety, and motivation during their first year at the university of st. gallen/switzerland. at this university, all students go through the same first-year study program, confronting them with very similar experiences, academic demands, and social environments. we asked all first-year students enrolled during the academic year 2011/2012 to fill in an online questionnaire at three time-points throughout their first year, thereby assessing their transition into university studies. the first data (t1) was collected about one week before the students entered the university (august 2011), the second (t2) in december 2011, and the third (t3) in april 2012 (after they had received the results of their first exams). thus, the time lag between the measurement points is roughly equal. the age, gender, and nationality distribution of the sample reflected the general characteristics of the student population; the sample, therefore, was representative of the first-year students at the university of st. gallen in 2011/2012. the response rate was about 63%, with 820 utilizable questionnaires at t1. the return rate at t3 was approx. 22% (285 questionnaires) of the students registered at t1. this is significantly worse than t1, but not unusually low for online surveys, in particular for a longitudinal study (nulty, 2008; sax, gilmartin, & bryant, 2003). table 1 sample overview for the panel study at each time point, we collected the data with an online questionnaire, using scales on different aspects of motivation and emotions as well as scales on other individual and socio-cultural factors of studying. the questionnaire was administered in german. based on self-determination theory (deci & ryan, 1996), the constructs intrinsic motivation (3 items, sample item: „i work and study for my course of studies because i am interested in the learning content.”) (grätz-tümmers, 2003) and extrinsic motivation (3 items, sample item: „the most important thing to me is having a good grade point average in my studies.”)1 (pintrich, smith, garcia, & mckeachie, 1991) were assessed. furthermore, we included scales on students’ self-efficacy (3 items, sample item: „if i make enough of an effort, i can master the learning content.”), as well as students’ anxiety (sample item: „i am worried about whether i can even manage my studies.”) (pekrun et al., 2005). to account for students’ overall attitude, we also asked them what they thought about the institution (3 items, sample item, i like studying at the university of st. gallen). at each measurement point, we asked the students which grade they strive for. at the third measurement point, we asked the remaining students how satisfied they were with their overall exam results. we used this as an, albeit weak, proxy for their performance. all scales met the expectations concerning psychometric properties (cronbach`s alpha between .712-784) and the assumed factor structure of the instrument has been confirmed (brahm & jenert, 2015). as independent variables, students’ gender, origin, and age were also collected. for data analysis, we used two major procedures: first, subgroups were identified applying latent profile analysis (lpa). this is a “technique[s] for recovering hidden groups in data by obtaining the probability that individuals belong to different groups” (ferguson, moore, & hull, 2019, p. 1). in comparison to traditional cluster analysis, it has multiple advantages (de clercq et al., 2020): “reducing type-1 error, accounting for measurement error, providing more rigorous decision criteria upon the number of profiles to retain, assessing the probability of membership of one participant into each respective profile” (p 4). to determine the latent profile, we used the personal variables (see section 2.1) at t1 to account for students’ diversity at the beginning of their studies. we decided on the number of profiles to select by applying the statistical criteria suggested by masyn (2013) and ferguson et al. (2019) to assess lpa models: the bayesian information criterion (bic), and the lo-mendell-rubin likelihood ratio test (lmr p value), to evaluate the relative fit of the models. the information criterion index bic is a descriptive index and is based on the model likelihood. it takes the complexity of the model into account (i.e., more profiles indicate more complexity) and assesses how well the model fits the data. nylund, asparouhov, and muthén (2007) recommend using the bic index for evaluating the relative model fit. a lower value of the bic indicates a better model fit and is preferred. the lo-mendell-rubin likelihood (lmr) ratio is a statistical test that is not based on chi-square distributions but on a derived distribution and parametric bootstrapping. this test compares the model fit improvement between a model with k profiles and (k-1) profiles. a significant p-value indicates an improvement in the model fit in the k-profile model compared to the (k-1) profile model. finally, entropy is a measure of classification uncertainty which is used less often (ferguson et al., 2019). however, it can be helpful to determine the quality of the delineation of profiles as it measures “how well each lpa model partitions the data into profiles” (ibid., p. 3). values higher than 0.8 indicate a good profile separation (asparouhov & muthen, 2014). it is worthwhile noting that this interpretation is counterintuitive as “lower entropy values actually represent more uncertainty or chaos in the model” (ferguson et al., 2019, p. 3). to interpret the lpa models, profile separation is examined as well (masyn, 2013). profile separation shows to what extent the different profiles are separate from each other. a high profile separation means that a respondent can clearly be assigned to a particular profile. a value of 0.70 or greater is considered acceptable (nagin, 2005). beyond the statistical criteria, the model with the best interpretable and theoretically meaningful solution should be preferred. to analyze the data further, we also used anova and general linear models to detect differences between the profiles. data were analyzed with spss version 24 and mplus 8. 3.2 qualitative longitudinal study the qualitative study was conducted in parallel with the panel survey. we conducted a longitudinal series of interviews with 14 first-year students of the same cohort as the quantitative panel. we used purposeful sampling (patton, 2002), defining selection criteria, to recruit participants. unfortunately, we could not use the t1 survey data to select our sample due to privacy reasons. the aim, though, was to have a varied sample with students who would differ in their self-efficacy, anxiety, and motivation. therefore, we slightly overrepresented women compared to the overall first-year student population. in the years before our study, women’s dropout during the first-year at the university of st. gallen had been significantly higher compared to their male counterparts despite showing the same a-level gpa. this suggested that, generally, women might be different in their non-cognitive personal characteristics as compared to men. second, swiss and non-swiss students were representative of the first-year student population. as non-swiss nationals have to pass an entry test, we supposed that this would likely result in systematic personal differences. third, concerning the familial background, students from both academic as well as non-academic backgrounds were chosen, as non-academic backgrounds are usually associated with lower self-efficacy and higher anxiety (zajacova, lynch, & espenshade, 2005). table 2 provides an overview of the participants and the sampling criteria gender and nationality. table 2 overview of participants the initial interview followed a detailed interview guide (see appendix a). the interviews were aligned to the quantitative survey addressing students’ feelings of self-efficacy, anxiety, and motivation. in contrast to the panel survey, we did not ask about these constructs directly, but rather talked about students’ relationship to the university context regarding both academic demands and social relations. questions, for instance, included: “which kind of students would the university of st. gallen want to develop/support? what does a student need to bring and do to successfully study at the university?” we asked students to talk about concrete experiences since the beginning of their study, respectively since the last interview, which affected their motivation, made them feel secure or anxious. thus, the obtained data linked the development of students’ personal characteristics to their experiences of specific situations in the study environment such as examinations, interactions with teachers and peers, etc. compared to the initial interview, the follow-up meetings were more narrative in nature. after each interview, we produced a short summary that was used to thematically link the interviews. in consequence, as the interview series progressed, each participant developed their individual narration of the first year, highlighting specific experiences and critical situations. we conducted four to five interviews with each participant, with the first interview lasting 45 to 75 minutes, and the follow-ups lasting 20 to 45 minutes. to analyze all 60 interviews, each recording was transcribed verbatim and coded, using the atlas.ti software. a combination of deductive and inductive coding was used (fereday & muir-cochrane, 2006). first, we developed a list of codes, including the personal constructs underlying the quantitative and qualitative study. after the first round of coding, this list was supplemented, with codes representing the contextual factors that influenced students’ first-year experience (see appendix c for the coding scheme). within each interview, two researchers coded sections in parallel to determine intercoder-reliability (ir). during the second round of coding (with the final codebook), adequate ir values were achieved (>.70). each student was analyzed as an individual case with the aim to associate personal developments with each student’s specific perceptions of the study context (e.g., positive or negative valuations of examination situations). 4. results 4.1 quantitative study: differentiating students according to personal variables concerning personal variables, results from the panel study showed a general decline in students´ overall motivation and self-efficacy as well as an increase in study-related anxiety over the first year (table 3). the descriptive statistics (table 3) show that the mean level of the constructs used for the latent profile analysis is high in general (scale values: 1=very low, 6=very high). the distribution of the data is skewed to the left (with one exception) and most variables have a low standard deviation. only anxiety has a higher standard deviation and a more uniform distribution. the different constructs correlate with each other to a medium extent, with higher correlations appearing among the constructs at the different measurement points (see appendix b). table 3 descriptive statistics of the observed constructs at each measurement point se = self-efficacy; anx. = anxiety; im = intrinsic motivation before students begin their studies, they are highly motivated and have a high level of self-efficacy (correlation of r = .24, p < 0.01). nevertheless, some students score high on anxiety despite their high motivation and self-efficacy. in particular, the rather high standard deviation of ‘anxiety’ indicates that there might be different subgroups regarding this aspect. to identify potential subgroups in the student population, we conducted a latent profile analysis based on students’ anxiety, intrinsic motivation, and self-efficacy at the measurement point 1. table 4 shows a summary of the relevant fit indices for two to five latent profiles. table 4 latent profile analysis: fit indices for the different profile solutions based on the fit indices, we opted for the three-profile solution for four reasons: first, it shows the lowest value of bic which we use since bic “has been shown to outperform other indices with more continuous indicators” (ferguson et al., 2019, p. 3). second, the three-profile solution is the last one where the lo, medell, and rubin (lmr) test is still significant, thus, indicating that this solution is better than the more parsimonious 2-class solution. third, this solution results in a reasonable number of students in each profile and is theoretically interpretable. finally, the entropy value of the three-profile solution is above .8, also indicating adequate profile separation. table 5 shows the probability of whether a student is assigned to the appropriate class, providing another indicator that the three-profile solution is the most appropriate for the sample (values around .90 are seen as the relevant threshold). the three latent profiles were labelled accordingly (table 6). table 5 classification probabilities for the most likely latent profile membership (column) by latent profile (row) table 6 descriptive statistics for the three student profiles with mean (standard deviation) se = self-efficacy, anx. = anxiety, im = intrinsic motivation in the following, we will first briefly describe the different profiles, and then we will give further information regarding the composition of these profiles and the differences concerning the relevant correlates. the profile with most students (324 at t1) is characterized by rather high motivation and medium-level anxiety. their self-efficacy, however, is not as high as with the second profile. the second biggest profile, with 221 students at t1, consists of highly motivated and self-confident students showing the highest levels of self-efficacy and intrinsic motivation and, correspondingly, the lowest level of anxiety. the last class (216 students) is characterized by the highest level of anxiety; both intrinsic motivation and, in particular, self-efficacy are rather low. anovas with post hoc bonferroni tests showed that the differences between the profiles are significant even though they are sometimes small. furthermore, the different profiles develop differently over time (table 6). while we will not analyze these differences in detail here, we can state that intrinsic motivation is significantly declining over time in all three profiles (see also brahm et al., 2017) while students’ self-efficacy plummets at the start of their studies, but then increases again from t2 to t3 (after students have completed their first assessment). for anxiety, however, the pattern is not the same for all profiles: while those in profiles 1 and 2 show increased anxiety at t2 which then lowers again at t3 (albeit at different levels), the students who are assigned to profile 3 (least motivated and most anxious) seem to lose their anxiety continuously over the first year. this may be seen as a hint that depending on the profile, students experience different sources of anxiety. we will elaborate on this in the discussion. in a second step, we investigated how the different profiles are composed concerning gender, nationality, and ambitions of the students. regarding gender, we found significant differences between the subgroups (f = 3.41, p < .05): in profile 1, 36% of students were female; in profile 2, females accounted for 29.3% of the students. in profile 3 which is characterized by the largest level of anxiety and the lowest levels of self-efficacy, female students were represented relatively more often with 41.1% (in the overall sample, 35.5% were female). in terms of students’ nationality, there were no significant differences between the major countries reported in the study. the same holds for the age distribution in the different profiles. regarding students’ performance ambition, their expected grades declined over time and there is a significant difference (for t1: f = 53.23, p < 0.00) between the profiles at all three measurement points. the first profile aims for medium grades, while the highly self-confident students expect significantly higher grades, and the least self-efficacious believe that they will receive relatively lower grades. finally, we looked into how the different profiles related to variables that suggest links to contextual aspects in the study environment. we used students’ satisfaction with their grades as a proxy for their achievement. although this can only function as a weak indicator of students’ actual performance, we also found significant differences (f = 5.98, p < 0.01) between the three profiles. interestingly, the highly motivated and medium self-efficacious profile 1 showed the highest satisfaction, followed by the second profile. this could indicate that these students have the most pragmatic and thus realistic idea about their performance, while students in profile 2 tend towards overconfidence. the students allocated to profile 3 were the least satisfied with their performance. furthermore, we investigated students’ attitudes towards the institution and found that the students in profile 2 showed significantly higher values in their attitude (for t1: f = 16.01, p < 0.00) while students in the most anxious profile showed a lower attitude. this pattern did not change over the first year. summarizing, the latent profile analysis hints at a significant diversity regarding students’ personal dispositions when starting in higher education. most interestingly, the profiles show differing patterns regarding important correlates of students’ motivation and anxiety. the quantitative data, however, do not allow for deeper insights on how such personal diversity affects the ways in which students perceive and behave in their study environment. how, for example, do differences in self-efficacy, anxiety and motivation become apparent in concrete situations such as exam preparation or peer interactions? do different students experience similar features of the study environment in different ways? which situations are crucial in shaping different students’ first-year experiences? the qualitative study allowed us to address these questions and develop a better understanding how personal differences are relevant for students’ actual perceptions and practices during their first year in he. 4.2 qualitative study: different students and different perceptions of similar contexts as a first analytical step, we looked at students’ personal diversity, ordering the qualitative sample according to their self-efficacy, anxiety, and motivation, following the logic of the quantitative study. as a second step, we investigated contextual diversity, i.e. whether there were systematic differences in how students perceived and acted on the academic and social environment of their first year at university. 4.2.1 personal diversity six of our 14 students reported rather high levels of self-efficacy and motivation and low levels of anxiety from the beginning. for those students, the overall first-year experience can be described as positive, as none of them ever mentioned doubts about continuing their studies. in contrast, a further six students were rather anxious and easily demotivated. in general, their first-year experience was rather negative, they often voiced doubts about being in the right place. roger is a typical example for this group, blaming himself for struggling with his studies: „yes, i think so, it’s my own fault, because i’ve problems to motivate myself. and when i really start studying, then i’m somehow distracted so fast. yes, that’s actually the problem” (roger, interview 2, line 156). interestingly, one student from this group, chris, changed from the ‘negative’ to the ‘positive’ group, having succeeded in the first series of exams, and most likely increased his self-efficacy from this mastery experience. two students, harry and olivia, cannot be easily attributed to either the ‘positive’ or the ‘negative’ experience group. both reported rather high levels of self-efficacy throughout the first year, i.e. they always felt confident they could master the academic challenges. however, they had severe motivation losses at several points during the first year and repeatedly voiced doubts about belonging to this university and subject. “yes, actually, well, we just had the [semester] break, the phase without lectures and there, now actually a lot has changed. before, i was relatively motivated, i still am, but [before] i was really positive. and slowly, especially during the break, then i was a bit, well, i thought: ‘do you really want to study this and do you really want to belong to these people?’ yes, like that, i reflected a bit. that changed a bit. and actually, i’m still not really sure about it” (harry, interview 2, line 12). overall, when looking at the patterns of self-efficacy, anxiety, and motivation, the findings from the qualitative study tie in nicely with the student subgroups detected with the latent profile analysis. the students that scored high in these constructs generally did well in mastering the challenges of the first year. 4.2.2 contextual diversity as expected, performance expectations in general, and exams in particular, were hugely important for the students. especially towards the end of the first semester, all 14 students talked a lot about the coming exams. “well, of course it’s on my mind that they [the exams] are coming. i know there is the preparation phase, but i know that i need to prepare as thoroughly as possible from now on. generally, i am not badly prepared. i don’t want to sit at it all the time, but also to have a little free time …. but, for example, what i’m working on is law and math because i feel that i will only be good there by practicing a lot. … in management i feel that i can easily hammer in the stuff during the final preparation phase” (daniel, interview 2, line 147-157). taking a closer look at how the students experienced and acted upon performance expectations highlighted differences between the subgroups. daniel is among the students in the ‘positive’ group, and for him, the exams are a source of motivation to engage with those academic subjects that he feels are most difficult for him. he aims to do very well in all subjects and the impending exams make him develop a preparation strategy. he is very focused on academic success and, during exam preparation, even manages his relationships with peers for optimizing his success: “well now, i have agreed with a friend, well he’s my best friend at uni, and i have thought about studying with him. i think he is a good student, well a bright head. and we said that we’ll see that first we reach a sufficiently high level individually and at the end of the preparation phase we’ll sit together if there’s something to discuss in management or economics.” (daniel, interview 3, line 127-132). daniel regards the exams as a challenge that he masters by using all resources accessible to him. in contrast, students in the ‘negative’ group generally regard exams as a threat that causes anxiety and great stress. “in some subjects, for example management, i am quite disoriented. i don’t know how to study in a structured way. it is somehow, it all consists of various components, there are the tutorials, the lectures and then the horror stories from the previous years; that it’s not even enough to learn the slides by heart and what not…” (chris, interview 3, line 73-86). chris’ anxiety coincides with low self-efficacy as well as a lack of suitable learning strategies to come to terms with the challenge of preparing for the exams. furthermore, students in our sample differed concerning the emphasis they put on the academic and social aspects of studying. some students, such as ‘positive daniel’ and ‘negative chris’ focused mostly on the academic aspects of studying. they showed a rather rational stance, drawing motivation from subject-related interest or extrinsic rewards or being anxious about mastering exams in certain subjects. emily is another example of students belonging to this group: “i especially want to do well in my studies. thus, i give my best – as much as possible – to get good grades and yes, just to have a good development, actually, to work on myself. i just think, time will tell in which direction i want to go, but now i just want to make the best possible out of myself, and i think i‘m at the right place for that” (emily, interview 1, line 123). students belonging to this group tended to keep social relations such as close friendships outside the university context, making a distinction between university friends and private friends. students in this group regularly developed their individual study plans relying on their capabilities and not wanting to be distracted by others. in contrast, other students talked about the first year as an essentially social experience. everything about studying had a social connotation, even when related to the academic aspect such as exam preparation. “i think that it’s not bad that the others [i.e. students] are there. it‘s rather like, when i’m sitting at home and do not work for half an hour, i don’t reflect on it. but when i’m sitting in the library and i don’t work for half an hour and i see how the others around me go on working, i somehow compare myself more to them in my head. and it shows you that that one over there, he’s on that page and i’m far from where he is. i think that can also stress you out if you see how the others are doing it. but in general, i think it’s more of a motivation” (mary, interview 1, line 124-131). while for ‘positive mary’, the social environment acts as a source of motivation, ‘negative rebecca’ feels that her uncertainties and anxieties are worsened by her interaction with peers: „i actually always ask somebody: ‘how did you do that’ or ‘how did you study that’. … yes, i’m pretty much influenced by that, because i’m so insecure and don’t know, if i’m doing it right, yes” (rebecca, interview 1, line 158). interestingly, these orientations towards the academic or the social seem to have a huge influence on the students’ first-year experience. emma belongs to the ‘negative’ group and emphasizes the social dimension of studying. she also interpreted academic aspects of studying such as performance expectations and exams in social terms, perceiving them as part of a social relationship between “the university” and “the students”: „the oral exam was a bit, well, it was the first oral exam at the university, and it was somehow strange. well, the focus was actually not on the most important aspects and it actually showed – also when i talked to others –, that it’s about us memorizing, really memorizing everything. and not, yes, i don’t know. that was actually pretty disappointing, because, i don’t know. they asked me such banal stuff. well, i needed to answer such banal questions. … but i also think, that was a statement by the uni for the first year so that you know ‘ok, learn everything by heart, really everything’” (emma, interview 2, line 21-40). the distinction between students who emphasize the academic and those who stress the social nature of studying extends to other characteristics of the first-year study context: in evaluating the first year, students in the ‘academic’ group would focus on the resources or challenges that helped or impeded their academic performance. for example, they valued well-designed learning materials or felt overwhelmed by information that was hard to access online. concerning interactions with faculty, students in the ‘social’ group tended to interpret faculty actions such as jokes, irony, exam questions, etc. as intentional statements. in contrast, the ‘academic’ group saw faculty rather as a part of the learning environment who could be a more or a less supportive resource for their learning processes. as shown above, the same is true for peer relations, which are regarded rather instrumental by students who regard studying as a job, focusing on the academic aspects. essentially, peers are either perceived as a resource for supporting academic performance (for the ‘positive’ students) or as competition (for the ‘negative’ students). for students in the ‘social’ group, developing friendships with likeminded people was what motivated them to continue studying, and what helped them to deal with uncertainties and anxieties. 5. discussion taking both the quantitative and the qualitative studies into account, we found personal and context-related aspects of diversity. integrating the two studies, the results from the panel-based latent profile analysis and the small-sample interview study clearly converge and complement each other. our research targeted two main goals: first, we aimed to identify subgroups within the first-year cohort at university by grouping them according to their self-efficacy, anxiety, and motivation. second, we investigated whether such personal diversity relates to contextual diversity, i.e. differences in how students perceive elements of their study environments. with our results, we contribute to research on student diversity, in particular in the transition to higher education. our study highlights the importance to consider both the microand meso-level when investigating students’ transition into higher education. furthermore, we can identify practical implications concerning the design of first-year study contexts for diverse student populations. 5.1 theoretical implications regarding the first aim, the results from our quantitative study complement and extend previous research on student transition in general, and on profile analyses of first-year students in particular (e.g. de clercq et al., 2020; martens & metzger, 2017). in line with the results reported by martens and metzger (2017), we found profiles that showed consistent patterns across self-efficacy, anxiety, and motivation. as expected, students who were assigned to the highly anxious profile were also least self-confident and least motivated. this corresponds to findings by sotardi and brogt (2019). we also found that the students in the most anxious profile were the least satisfied with their performance in the first assessment period in the first year. as a final relevant variable, we looked into motivation. for all three profiles identified, our data confirm the negative development of students’ motivation throughout the first year and is, thus, in line with many studies in higher education (busse, 2013; jacobs & newstead, 2000; lau et al., 2008; lieberman & remedios, 2007; martin et al., 2010; pan & gauvain, 2012). interestingly, motivation is the only construct showing such a stable decline. in contrast, both self-efficacy and anxiety either remain stable or, after an initial decline, improve again from the second to the third time point (see table 6). unfortunately, we are not able to distinguish between different kinds of motivation as martens and metzger (2017) did. this is clearly an avenue for future research. concerning the links between personal and contextual diversity, our findings are in line with de clercq et al. (2020, p. 8), who state that “students with different entrance profiles reported different perceptions of context, motivation, engagement, and consequently reached different levels of achievement”. in our research, the affiliation to one or another profile was associated with more or less positive attitudes towards the university and with distinct satisfaction with their first assessment results. profiles 1 and 2 show quite similar development patterns over the first year albeit on different absolute levels of the investigated constructs. in contrast, students in the ‘least motivated and most anxious’ profile differ from the other two profiles regarding the development of anxiety. for them, anxiety seems to be a constant companion which declines slightly yet constantly during the first year. in contrast, students in the other profiles become more anxious in anticipation of exams and go back to a rather low level of anxiety after the exam period. again, this ties in with findings by sotardi et al. (2020) who show that students may be anxious about different contextual stimuli (in their case different assessment types). these quantitative findings are a first hint that students in the different profiles react in varying ways to contextual aspects. findings from the qualitative study underpin this assumption, highlighting that students with diverse personal prerequisites perceive performance expectations as well as the social environment very differently. for example, highly motivated and self-efficacious students regard high-performing peers as a source of motivation and a resource for studying which is in line with the general results on peer support (e.g. thomas, 2002; thompson & mazer, 2009). in contrast, students with low motivation and self-efficacy regard these peers as a potential threat and a source of anxiety. while we did not have objective data regarding students’ achievement, we were able to show that those students who were assigned to the most motivated and self-confident group also showed the most ambition regarding their grades, the most positive attitude towards the university, and were most content with their performance. this finding can be seen as a confirmation of the relation between high self-efficacy and college adjustment (ramos-sánchez & nichols, 2007) which is also in alignment with bandura’s (1989) theory indicating that students with high self-efficacy may be able to deal better with challenges during the first year. our results suggest that when student diversity is taken into account, theoretical assumptions of relationships between personal and contextual variables may become less clear than previously thought. for example, our qualitative results show that peer interaction counts among the factors, which different students experience in very different ways. while most studies posit that peer interaction is generally a good thing for first-year students (e.g. rocconi, 2011; walsh et al., 2009; thomas, 2002; thompson & mazer, 2009), our findings suggest that for some student groups and in some situations, being exposed to their peers may have rather negative effects (cf. brouwer et al., 2018; riordan & carey, 2019). the notion that theoretical assumptions need to be differentiated in the light of student diversity ties in nicely with recent studies finding differential effects between personal and contextual characteristics. for example, duchatelet and donche (2019) found that the perception and the effects of autonomy-supportive learning environments differed according to the students’ level and type of motivation. sotardi et al. (2020) distinguish different kinds of test anxiety depending on student and exam characteristics. de clercq et al. (2020) found that the affiliation to a specific student profile was related to strong variations in how students perceived contextual aspects such as course value. as a theoretical implication of their study, they suggest distinguishing between “shared” and “specific” factors that influence student transition (p. 8). shared factors influence students’ achievements independent of their individual profiles in the same ways, while specific factors have different influences on students in different profiles. being aware of such a distinction seems particularly important for research on the effects of specific educational interventions. depending on whether they address shared or specific factors, they may have little or even unintended effects for some student sub-populations. 5.2 limitations of course, the design of our studies has some important limitations, which need to be addressed here. both the quantitative and the qualitative studies have their individual flaws: our research merits the benefits of a longitudinal panel study, but at the same time suffers the high attrition that is well-known for this type of survey, especially with the long intervals between the measurement points. the students who only participated in the first measurement did not differ significantly from those who replied to all three questionnaires. nevertheless, some students who replied before the beginning of their studies might have decided to drop out of the university, thus not being reachable for the survey anymore. future research should attempt to gather information about the students’ whereabouts after the important transition period. also, we did not gather data on the students’ actual academic performance, making it impossible to relate our findings directly to previous studies such as de clercq et al. (2020) who investigated the effects of being affiliated to a specific sub-population on study success. furthermore, the research was conducted at one particular university with a limited range of subjects – typical for a business school. this is both a limitation and a benefit. as the students study the same subjects in the first year, the effects of different study programs or disciplines can be excluded. however, this also means that the results are limited to one university. thus, our findings should ideally be replicated at other institutions and – if possible – with more disciplinary diversity. above all, a more data-based linkage between the quantitative latent profile analysis and the qualitative interview study would strengthen the ties between personal and contextual diversity. unfortunately, we could not analyze the interviewed students’ survey data in a personalized way, as this would have violated the confidentiality guidelines for both studies. an alternative approach to solve this problem could have been to include variables measuring students’ perceptions of aspects of the study environment such as the support structures or the degree of challenge perceived in learning tasks in the survey (kuh, 2009; kuh, kinzie, schuh, & whitt, 2005; zhao & kuh, 2004). this may be a promising route for future survey-based research as the qualitative study identified some specific events that according to the interviews affected the students’ stance towards studying. despite these limitations, this study emphasizes the need to consider student diversity when investigating their transition experiences. it was particularly worthwhile complementing our latent profile analysis with the qualitative results as they illustrated the interplay of personal and contextual characteristics. 5.3 conclusion and practical implications combining a person-centered quantitative panel study with a small-sample qualitative approach emphasizes the social-cognitive lens applied in this research (bandura, 1989). it links students’ personal development to the experiences they make when interacting with their study environment. this interaction is reciprocal as students’ individual characteristics influence how they evaluate certain aspects of the study environment. this becomes apparent in the different orientations of the students in our sample, which highlight how different students perceive similar situations in distinctive and even contradictory ways. this is also visible in the different patterns of our student profiles. in turn, specific events in the study environment influence how students’ motivations and emotions regarding studying develop. this interactional view emphasizes the conception of transition into he as a process (tett, cree, & christie, 2017). rather than regarding students’ developments as something purely personal, we contribute to research on students’ transition by understanding their development as a result of each individual student’s interactions with specific aspects of his or her study context. thus, our study confirms the assumption that it is not only necessary to look into students’ diversity, but rather to investigate how this diversity and the study environment, which students encounter at university, reinforce or hinder each other. this notion may inform future research which should investigate the interactional nature of the developmental process during the transition on a larger scale. from a practical point of view, our study emphasizes the need for more customized support structures during the first year of higher education. these should go beyond the usual distinction of traditional and non-traditional students, but could, for instance, consider different levels of students’ anxiety and confidence. one particularly interesting avenue could be the provision of learning-oriented communities (zander et al., 2018) as these might obviate the challenge for some students who are not inclined to socialize on a more private level. our results, however, imply that such communities should be carefully composed and supported to not exaggerate existing inequalities between students. another important practical implication concerns students’ anxiety, in particular when confronted with their first assessments. in light of the findings of both the quantitative and qualitative study, universities could further support their students by clarifying their expectations regarding the first exam period and by providing different kinds of support measures in advance of the first exams as some students may need more content-oriented support while others may need more learning-focused help. finally, our study may serve as an example for the value of using both quantitative and qualitative research approaches despite the limitations as to how the two studies were connected (see above). as discussed, recent studies that combine personal and contextual variables (e.g. de clercq et al., 2020; duchatelet & donche, 2019; sotardi et al., 2020) show that existing theoretical models regarding the relationships between contextual and personal variables may only be valid for very specific student populations. from an epistemological viewpoint, this could be considered as a limitation of variable-based approaches in general. one way to deal with this challenge is the increasing use of typological analysis such as latent profile analysis. to grasp the full complexity of person-context interactions, however, also more phenomenological approaches supported by qualitative methods seem to be promising (aspers, 2009). in this sense, our qualitative study provides exemplary illustrations for different kinds of students that we had identified through latent profile analysis in the quantitative panel. keypoints transition to higher education is influenced by personal and contextual characteristics our study combines latent profile analysis (lpa) on a large-scale student panel and small-scale qualitative longitudinal interviews the lpa identified three student subgroups based on their anxiety, motivation, and self-efficacy the qualitative study found that depending on personal differences, students perceived and acted differently on specific situations such as exam preparation or peer-interaction. footnotes 1 results for extrinsic motivation showed that it developed much the same as intrinsic motivation, however, on lower levels. as this does not add any extra information for this study, we do not report the results for extrinsic motivation. references ahern, t. c., thomas, j. a., tallent-runnels, m. k., lan, w. y., cooper, s., lu, x., & cyrus, j. (2006). the effect of social grounding on collaboration in a computer-mediated small group discussion. internet and higher education , 9(1), 37–46. https://doi.org/10.1016/j.iheduc.2005.09.008 asparouhov, t., & muthen, b. (2014). auxiliary variables in mixture modeling: using the bch method in mplus to estimate a distal outcome model and an arbitrary secondary model. mplus web notes , 21(2), 1–22. aspers, p. (2009). empirical phenomenology: a qualitative research approach (the cologne seminars). indo-pacific journal of phenomenology , 9(2), 1–12. https://doi.org/10.1080/20797222.2009.11433992 brahm, t., & jenert, t. (2015). on the assessment of attitudes towards studying – development and validation of a questionnaire. learning and individual differences, 43, 233– 242. https://doi.org/10.1016/j.lindif.2015.08.019 brahm, t., jenert, t., & wagner, d. (2017). students’ transition into a business school – a longitudinal study of their motivational development. higher education 73(3), 459–478. https://doi.org/10.1007/s10734-016-0095-8 aymans, s. c., & kauffeld, s. (2015). to leave or not to leave? critical factors for university dropout among first-generation students. zeitschrift für hochschulentwicklung, 10, 23–43. https://doi.org/10.3217/zfhe-10-04/02 bandura, a. (1989). human agency in social cognitive theory. american psychologist , 44(9), 1175–1184. https://doi.org/10.1037/0003-066x.44.9.1175 briggs, a. r. j., clark, j., & hall, i. (2012). building bridges: understanding student transition to university. quality in higher education , 18(1), 3–21. https://doi.org/10.1080/13538322.2011.614468 brouwer, j., flache, a., jansen, e., hofman, a., & steglich, c. (2018). emergent achievement segregation in freshmen learning community networks. higher education , 76(3), 483–500. https://doi.org/10.1007/s10734-017-0221-2 busse, v. (2013). why do first-year students of german lose motivation during their first year at university? studies in higher education , 38(7), 951–971. https://doi.org/10.1080/03075079.2011.602667 byrne, m., & flood, b. (2005). a study of accounting students' motives, expectations and preparedness for higher education. journal of further and higher education , 29(2), 111–124. https://doi.org/10.1080/03098770500103176 cassady, j. c., & johnson, r. e. (2002). cognitive test anxiety and academic performance. contemporary educational psychology , 27(2), 270–295. https://doi.org/10.1006/ceps.2001.1094 chapman, d. w., & pascarella, e. t. (1983). predictors of academic and social integration of college students. research in higher education , 19(3), 295–322. https://doi.org/10.1007/bf00976509 christie, h., tett, l., cree, v. e., hounsell, j., & mccune, v. (2008). ‘a real rollercoaster of confidence and emotions’: learning to be a university student. studies in higher education , 33(5), 567–581. https://doi.org/10.1080/03075070802373040 coertjens, l., donche, v., maeyer, s. de, van daal, t., & van petegem, p. (2017). the growth trend in learning strategies during the transition from secondary to higher education in flanders. higher education , 73(3), 499–518. https://doi.org/10.1007/s10734-016-0093-x cotten, s. r., & wilson, b. (2006). student-faculty interactions: dynamics and determinants. higher education , 51(4), 487–519. https://doi.org/10.1007/s10734-004-1705-4 deci, e. l., & ryan, r. m. (1996). need satisfaction and the self-regulation of learning. learning and individual differences , 8(3), 165–183. https://doi.org/10.1016/s1041-6080(96)90013-8 de clercq, m., galand, b., & frenay, m. (2020). one goal, different pathways: capturing diversity in processes leading to first-year students' achievement. learning and individual differences , 81, 101908. https://doi.org/10.1016/j.lindif.2020.101908 delaney, a. m. (2008). why faculty-student interaction matters in the first year experience. tertiary education and management, 14 (3), 227–241. https://doi.org/10.1080/13583880802228224 donche, v., coertjens, l., & van petegem, p. (2010). learning pattern development throughout higher education: a longitudinal study. learning and individual differences , 20(3), 256–259. https://doi.org/10.1016/j.lindif.2010.02.002 donche, v., maeyer, s. de, coertjens, l., van daal, t., & van petegem, p. (2013). differential use of learning strategies in first-year higher education: the impact of personality, academic motivation, and teaching strategies. british journal of educational psychology , 83(2), 238–251. https://doi.org/10.1111/bjep.12016 duchatelet, d., & donche, v. (2019). fostering self-efficacy and self-regulation in higher education: a matter of autonomy support or academic motivation? higher education research & development , 38(4), 733–747. https://doi.org/10.1080/07294360.2019.1581143 dyson, r., & renk, k. (2006). freshmen adaptation to university life: depressive symptoms, stress, and coping. journal of clinical psychology , 62(10), 1231–1244. https://doi.org/10.1002/jclp.20295 fenollar, p., román, s., & cuestas, p. j. (2007). university students' academic performance: an integrative conceptual framework and empirical analysis. british journal of educational psychology , 77(4), 873–891. https://doi.org/10.1348/000709907x189118 fereday, j., & muir-cochrane, e. (2006). demonstrating rigor using thematic analysis: a hybrid approach of inductive and deductive coding and theme development. international journal of qualitative methods , 5(1), 80–92. https://doi.org/10.1177/160940690600500107 ferguson, s. l., moore, e. w. g., & hull, d. m. (2019). finding latent groups in observed data: a primer on latent profile analysis in mplus for applied researchers. international journal of behavioral development , 1–11. https://doi.org/10.1177/0165025419881721 gale, t., & parker, s. (2012). navigating change: a typology of student transition in higher education. studies in higher education , 39(5), 1–20. https://doi.org/10.1080/03075079.2012.721351 grätz-tümmers, j. (2003). arbeitsprobleme im studium . marburg: philipps-universität. https://doi.org/10.17192/z2004.0128 hailikari, t., kordts-freudinger, r., & postareff, l. (2014). exploring the relationship between university students' emotions and their studying and learning: a mixed-method approach. paper presented at the earli sig higher education conference, leuven. honicke, t., & broadbent, j. (2016). the influence of academic self-efficacy on academic performance: a systematic review. educational research review , 17, 63–84. https://doi.org/10.1016/j.edurev.2015.11.002 hsieh, p.-h., sullivan, j. r., & guerra, n. s. (2007). a closer look at college students: self-efficacy and goal orientation. journal of advanced academics , 18(3), 454–476. https://doi.org/10.4219/jaa-2007-500 jacobs, p. a., & newstead, s. e. (2000). the nature and development of student motivation. british journal of educational psychology , 70(2), 243–254. https://doi.org/10.1348/000709900158119 kim, y. k., & lundberg, c. a. (2016). a structural model of the relationship between student-faculty interaction and cognitive skills development among college students. research in higher education , 57(3), 288–309. https://doi.org/10.1007/s11162-015-9387-6 kuh, g. d. (2009). what student affairs professionals need to know about student engagement. journal of college student development , 50(6), 683–706. https://doi.org/10.1353/csd.0.0099 kuh, g. d., cruce, t. m., shoup, r., & kinzie, j. (2008). unmasking the effects of student engagement on first-year college grades and persistence. the journal of higher education , 79(5), 540–563. https://doi.org/10.1080/00221546.2008.11772116 kuh, g. d., & hu, s. (2001). the effects of student-faculty interaction in the 1990s. the review of higher education , 24(3), 309–332. https://doi.org/10.1353/rhe.2001.0005 kuh, g. d., kinzie, j., schuh, j. h., & whitt, e. j. (2005). assessing conditions to enhance ecuational effectiveness: the inventory for student engagement and success . san francisco: jossey-bass. kyndt, e., coertjens, l., van daal, t., donche, v., gijbels, d., & van petegem, p. (2015). the development of students' motivation in the transition from secondary school to higher education. a longitudinal study. learning and individual differences , 39, 114–123. https://doi.org/10.1016/j.lindif.2015.03.001 kyndt, e., donche, v., van daal, t., gijbels, d., & van petegem, p. (2019). does self-efficacy contribute to the development of students' motivation across the transition from secondary school to higher education? european journal of psychology of education , 34(2), 457–478. https://doi.org/10.1007/s10212-018-0389-6 lau, s., liem, a. d., & nie, y. (2008). taskand self-related pathways to deep learning: the mediating role of achievement goals, classroom attentiveness, and group participation. british journal of educational psychology , 78(4), 639–662. https://doi.org/10.1348/000709907x270261 leese, m. (2010). bridging the gap: supporting student transitions into higher education. journal of further and higher education , 34(2), 239–251. https://doi.org/10.1080/03098771003695494 lieberman, d. a., & remedios, r. (2007). do undergraduates' motives for studying change as they progress through their degrees? british journal of educational psychology , 77(2), 379–395. https://doi.org/10.1348/000709906x157772 martens, t., & metzger, c. (2017). different transitions of learning at university: exploring the heterogeneity of motivational processes. in e. kyndt, v. donche, k. trigwell, & s. lindblom-ylänne (eds.), higher education transitions: theory and research (1st ed., pp. 31–46). london: routledge. martin, a. j., colmar, s. h., davey, l. a., & marsh, h. w. (2010). longitudinal modelling of academic buoyancy and motivation: do the '5cs' hold up over time? british journal of educational psychology , 80(3), 473–496. https://doi.org/10.1348/000709910x486376 masyn, k. (2013). latent class analysis and finite mixture modeling. the oxford handbook of quantitative methods in psychology , 2, 551–611. meehan, c., & howells, k. (2018). 'what really matters to freshers?': evaluation of first year student experience of transition into university. journal of further and higher education , 42(7), 893–907. https://doi.org/10.1080/0309877x.2017.1323194 mellanby, j., & zimdars, a. (2011). trait anxiety and final degree performance at the university of oxford. higher education , 61(4), 357–370. https://doi.org/10.1007/s10734-010-9335-5 nagin, d. s. (2005). group-based modeling of development . london: harvard university press. https://doi.org/10.4159/9780674041318 nevill, a., & rhodes, c. (2004). academic and social integration in higher education: a survey of satisfaction and dissatisfaction within a first-year education studies cohort at a new university. journal of further and higher education , 28(2), 179–193. https://doi.org/10.1080/0309877042000206741 noyens, d., donche, v., coertjens, l., van daal, t., & van petegem, p. (2019). the directional links between students’ academic motivation and social integration during the first year of higher education. european journal of psychology of education , 34(1), 67–86. https://doi.org/10.1007/s10212-017-0365-6 nulty, d. d. (2008). the adequacy of response rates to online and paper surveys: what can be done? assessment & evaluation in higher education , 33(3), 301–314. https://doi.org/10.1080/02602930701293231 nylund, k. l., asparouhov, t., & muthén, b. o. (2007). deciding on the number of classes in latent class analysis and growth mixture modeling: a monte carlo simulation study. structural equation modeling: a multidisciplinary journal , 14(4), 535–569. https://doi.org/10.1080/10705510701575396 pan, y., & gauvain, m. (2012). the continuity of college students’ autonomous learning motivation and its predictors: a three-year longitudinal study. learning and individual differences , 22(1), 92–99. https://doi.org/10.1016/j.lindif.2011.11.010 patton, m. q. (2002). qualitative evaluation and research methods (3rd). london et al.: sage. pekrun, r., elliot, a. j., & maier, m. a. (2009). achievement goals and achievement emotions: testing a model of their joint relations with academic performance. journal of educational psychology , 101(1), 115–135. https://doi.org/10.1037/a0013383 pekrun, r., goetz, t., titz, w., & perry, r. p. (2002). academic emotions in students' self-regulated learning and achievement: a program of qualitative and quantitative research. educational psychologist , 37(2), 91–105. https://doi.org/10.1207/s15326985ep3702_4 pekrun, r., götz, t., & perry, r. p. (2005). academic emotions questionnaire (aeq). user’s manual. pekrun, r., & linnenbrink-garcia, l. (2012). academic emotions and student engagement. in s. l. christenson, a. l. reschly, & c. wylie (eds.), handbook of research on student engagement (pp. 259–282). new york: springer. https://doi.org/10.1007/978-1-4614-2018-7_12 perry, c., & allard, a. (2003). making the connections: transition experiences for first-year education students. journal of educational inquiry , 4(2), 74–89. pintrich, p. r., smith, d. a. f., garcia, t., & mckeachie, w. j. (1991). a manual for the use of the motivated strategies for learning questionnaire (mslq) . ann arbor, mi: national center for research to improve postsecondary teaching and learning. prat-sala, m., & redford, p. (2010). the interplay between motivation, self-efficacy, and approaches to studying. british journal of educational psychology , 80, 283–305. https://doi.org/10.1348/000709909x480563 ramos-sánchez, l., & nichols, l. (2007). self-efficacy of first-generation and non-first-generation college students: the relationship with academic performance and college adjustment. journal of college counseling , 10(1), 6–18. https://doi.org/10.1002/j.2161-1882.2007.tb00002.x ratelle, c. f., guay, f., larose, s., & senécal, c. (2004). family correlates of trajectories of academic motivation during a school transition: a semiparametric group-based approach. journal of educational psychology , 96(4), 743–754. https://doi.org/10.1037/0022-0663.96.4.743 riordan, b. c., & carey, k. b. (2019). wonderland and the rabbit hole: a commentary on university students' alcohol use during first year and the early transition to university. drug and alcohol review , 38(1), 34–41. https://doi.org/10.1111/dar.12877 robbins, s. b., lauver, k., le, h., davis, d., langley, r., & carlstrom, a. (2004). do psychosocial and study skill factors predict college outcomes?: a meta-analysis. psychological bulletin , 130(2), 261–288. https://doi.org/10.1037/0033-2909.130.2.261 rocconi, l. m. (2011). the impact of learning communities on first year students’ growth and development in college. research in higher education , 52(2), 178–193. https://doi.org/10.1007/s11162-010-9190-3 sax, l. j., gilmartin, s. k., & bryant, a. n. (2003). assessing response rates and nonresponse bias in web and paper surveys. research in higher education , 44(4), 409–432. https://doi.org/10.1023/a:1024232915870 schneider, m., & preckel, f. (2017). variables associated with achievement in higher education: a systematic review of meta-analyses. psychological bulletin , 143(6), 565–600. https://doi.org/10.1037/bul0000098 sotardi, v. a., bosch, j., & brogt, e. (2020). multidimensional influences of anxiety and assessment type on task performance. social psychology of education , 23(2), 499–522. https://doi.org/10.1007/s11218-019-09508-3 sotardi, v. a., & brogt, e. (2019). influences of learning strategies on assessment experiences and outcomes during the transition to university. studies in higher education . retrieved from https://www.tandfonline.com/doi/full/10.1080/03075079.2019.1647411 talsma, k., schüz, b., schwarzer, r., & norris, k. (2018). i believe, therefore i achieve (and vice versa): a meta-analytic cross-lagged panel analysis of self-efficacy and academic performance. learning and individual differences , 61, 136–150. https://doi.org/10.1016/j.lindif.2017.11.015 tett, l., cree, v. e., & christie, h. (2017). from further to higher education: transition as an on-going process. higher education , 73(3), 389–406. https://doi.org/10.1007/s10734-016-0101-1 thomas, l. (2002). student retention in higher education: the role of institutional habitus. journal of education policy , 17(4), 423–442. https://doi.org/10.1080/02680930210140257 thompson, b., & mazer, j. p. (2009). college student ratings of student academic support: frequency, importance, and modes of communication. communication education , 58(3), 433–458. https://doi.org/10.1080/03634520902930440 tinto, v. (1993). leaving college: rethinking the causes and cures of student attrition (2. ed., 4. print). chicago, il: university of chicago press. villavicencio, f. t., & bernardo, a. b. i. (2013). positive academic emotions moderate the relationship between self-regulation and academic achievement. british journal of educational psychology , 83(2), 329–340. https://doi.org/10.1111/j.2044-8279.2012.02064.x walsh, c., larsen, c., & parry, d. (2009). academic tutors at the frontline of student support in a cohort of students succeeding in higher education. educational studies , 35(4), 405–424. https://doi.org/10.1080/03055690902876438 zajacova, a., lynch, s., & espenshade, t. (2005). self-efficacy, stress, and academic success in college. research in higher education , 46(6), 677–706. https://doi.org/10.1007/s11162-004-4139-z zander, l., brouwer, j., jansen, e., crayen, c., & hannover, b. (2018). academic self-efficacy, growth mindsets, and university students' integration in academic and social support networks. learning and individual differences , 62, 98–107. https://doi.org/10.1016/j.lindif.2018.01.012 zeidner, m. (1998). test anxiety: the state of the art . new york, ny: plenum press. zhao, c.-m., & kuh, g. d. (2004). adding value: learning communities and student engagement. research in higher education , 45(2), 115–138. https://doi.org/10.1023/b:rihe.0000015692.88534.de appendix a interview guideline initial interview: how do you experience the transition? please let us know how you experienced your first month at the university of st. gallen? appendix b i correlation table of the variables considered in the latent profile analysis appendix c coding scheme codepen durik frontline learning research vol.8 no. 3 special issue (2020) 85 103 issn 2295-3159 variability in certainty of self-reported interest: implications for theory and research amanda m. durik a& jade s. jenkinsa anorthern illinois university, usa article received 15 may 2019 / revised 29 august/ accepted 29 august / available online 30 march abstract these studies examined self-reported interest, how level of interest is related to reported certainty of interest, and whether certainty helps to clarify the relationship between interest and behavior. this research borrows from research on attitudes showing that attitude certainty helps to clarify the relationship between attitudes and behavior. a pilot study examined the relationship between self-reported interest and certainty of interest within four disciplines (math, psychology, biology, and astronomy). these relationships were replicated in math (study 1) and psychology (study 2), and the relationships between interest and behavior were stronger for those with greater certainty. for domains in which participants had sufficient levels of experience and varied levels of interest, curvilinear relationships were found between level of interest and certainty, showing that certainty is higher among individuals who report more extreme (high or low) levels of interest. moreover, self-reported interest predicted behavior more strongly for those with more certainty in their responses. discussion surrounds the theoretical and methodological utility of considering certainty of interest alongside measures of self-reported interest. keywords: self-report; certainty; interest info corresponding author email: adurik@niu.edu doi: https://doi.org/10.14786/flr.v8i3.491 1. background when scanning a classroom, it is hard not to notice variability in student motivation. the students who are attentive, active, listening, and responding can be differentiated from those who are off task and distracted. some of this differentiation can be traced to students’ varying levels of interest in the domain that is being taught. interest is an emotional response to particular stimuli that has both cognitive and motivational features (see reviews by hidi & renninger, 2006; krapp, 2002; prenzel, 1992). it can guide students toward better self-regulation while learning because it helps students focus on content, choose to engage, persevere through challenges, and recall more. interest is conceptualized as both residing in the person over time (individual interest) as well as varying in the moment in response to stimuli in the immediate situation (situational interest). as such, interest can fluctuate in response to processes operating within the person and stimuli present in the situation. individual interest is an enduring tendency to approach and seek learning opportunities in a given domain (ainley & ainley, 2011; deci, 1992; hidi & renninger, 2006; krapp, 2002; prenzel, 1992; renninger, hidi, & krapp, 1992; schiefele, 1991; silvia, 2006). the domain is associated with affective involvement and meaningfulness (schiefele, 1991), which are enhanced by the acquisition and use of knowledge (hidi & renninger, 2006; renninger, 2000; prenzel, 1992). a hallmark of individual interest is willingness to reengage in the domain over time (hidi & renninger, 2006). the measurement of interest has been given considerable attention. options of how to measure interest often include self-report scales and interviews (see discussions by krapp & prenzel, 2011; renninger & hidi, 2011). within those, decisions about which features of interest should be captured vary. for example, within the collection of self-report scales that are available, some include feelings associated with interacting with domain content, the meaning or value associated with the domain, and/or the perceived knowledge or actual knowledge that individuals have stored in the domain (see review by renninger & hidi, 2011). most conceptualizations of interest reflect the idea that interest can change in response to experience and exposure to content, and we argue that research on interest must inherently take a developmental perspective. this leads to challenges in measurement because any measure is a snapshot of a person’s interest at a particular moment. moreover, in the case of closed-ended self-report measures of interest, a given response is that person’s best attempt at quantifying their level of interest at that time. these issues also manifest in the capacity (or lack of, in some cases) for self-report measures of interest to predict behavioral manifestations of interest. behavioral measures of interest often assess whether and how people choose to engage with domain content. not surprisingly, measures of self-reported interest often positively predict behavior, such as free-choice behaviors observed in the laboratory, course taking, retrospective reports of behavior, and intentions to behave in the future (ainley et al., 2002; harackiewicz, durik, barron, linnenbrink-garcia, & tauer, 2008; renninger, 1990; simpkins et al., 2006; wijnia et al., 2014). that said, although correlations between self-reported interest and behavior are often present, they may not be as strong as one might expect. other areas of research have carefully considered why self-reported and behavioral data are not always strongly associated, and social psychologists who study attitudes began working on this issue within their own area in the 1970s. this work toward understanding the relationship between self-reported attitudes and behavior has led to greater clarity surrounding the nature of and research on attitudes. similarly, careful attention to the relationship between self-reported interest and behavior may also help clarify the construct of interest. the challenges encountered by attitude researchers in assessing self-reported attitudes are likely similar to many of the challenges encountered by interest researchers assessing self-reported interest. in both cases, participants are asked to evaluate their responses to a particular class of stimuli or ideas. although it is simple enough for an individual to provide a response on a scale, this deceivingly simple response is the outcome of much more complex processes. it should be noted, however, that we do not argue that interest and attitudes are the same. interest assumes an active process on the part of the individual that propels them toward knowledge acquisition, elaboration, or growth (deci, 1992) whereas an attitude does not necessarily trigger this process and may actually do the opposite. for example, a person who holds a positive attitude toward recycling would believe that recycling is good. a person who holds a negative attitude about recycling would believe that recycling is bad, possibly a waste. whereas the person with a positive attitude may be more open to learning about recycling than the person with a negative attitude, the attitude itself does not motivate learning. in contrast, a person with an interest in recycling would be expected to have learning goals related to recycling (e.g., a desire to learn about the processes related to recycling, which materials can be recycled, and why they should be recycled). attitude researchers have addressed the issue of attitude-behavior consistency in several ways. for example, one approach recognized that other environmental variables such as norms and the opportunity and ability to engage in the behavior were also critical (ajzen, 1991). another approach recognized that attitudes predicted behavior more strongly when people focused on only one side of an attitude (e.g., for or against; glasman & albarracin, 2006). finally, another approach identified that other attitude features contributed to attitude strength, and increased the relationship between attitudes and behavior (krosnick & petty, 1995). borrowing from this last approach, the current research centers on the idea that individuals vary in the extent to which they are certain of their attitudes. certainty refers to the extent to which individuals are confident in their assessment of an attitude as clear and correct (rucker et al., 2014). certain attitudes are held more strongly and are less likely to change (krosnick & petty, 1995; pomerantz et al., 1995). moreover, participants who reported greater certainty of their attitudes, which was measured separately from the attitudes themselves, were more likely to behave in ways that were consistent with their attitudes (e.g., bizer et al., 2006; fazio & zanna, 1978; glasman & albarracıin, 2006; tormala, 2016; tormala & rucker, 2007). as an example, although participants may report varying levels of attitudes toward recycling, those who have more certain and positive attitudes toward recycling would be more likely to actually recycle. just as individuals can report less or more certainty of their attitudes, we theorized that some participants would be more certain of their self-reported interest and others less so. we reasoned that if attitude researchers were able to clarify the relationship between attitudes and behaviors by considering attitude certainty, it was worthwhile to attempt the same for individual interest. this approach, compared to some of the other approaches explored within the attitude literature, was selected because it preserved the assumption that research on interest must consider development. interest changes and people may become more certain of their interest over time. as such, not only might the inclusion of certainty clarify the relationship between interest and behavior, but it may also provide insight into how participants’ awareness of interest changes. specifically, certainty of interest may prove useful in gaining insight into the nature of interest and how individuals come to recognize their interests. drawing again from the attitude literature, the extent to which individuals become more confident or certain of their attitude is related to the extent and valence of prior experiences and the amount of careful thought put toward the object of the attitude (berger, 1992; bizer et al., 2006; fazio & zanna, 1978; glasman & albarracin, 2006; jonas et al., 1997; krishnan & smith, 1998; prislin et al., 1998). prior experience and careful thought toward the object of an attitude have been found to increase certainty, which then contributes to stability of attitudes. the current research examines variability in certainty as a starting point in determining whether similar processes may be operating for interest as they are for attitudes. the first aim of the current research is to examine certainty of interest and how certainty varies with levels of self-reported interest. the second aim was to test whether individuals who are more certain behave in ways that align more closely to their interests. 2. pilot study the purpose of the pilot study was to explore the patterns of association between self-reported interest and certainty of interest among different domains (math, biology, astronomy, and psychology). these domains were chosen because it was expected that participants would vary in their prior experience with each. the pilot sample was composed entirely of advanced psychology students. as such, this population was anticipated to have varying levels of exposure to math and biology (due to compulsory education), high exposure to psychology (as their program of study), and low exposure to astronomy (neither compulsory nor inherently linked to their program of study). this anticipated variability in experience with the different domains may have implications for certainty, and create meaningful comparisons across the domains. we tested whether certainty and interest would be related in a linear or curvilinear fashion, and were especially interested in a curvilinear relationship such that participants who reported more extreme levels of interest (either low or high) may also be more certain of their interest. this pattern was found in prior research on attitudes revealing a curvilinear relationship between certainty and willingness to advocate for an attitudinal position (cheatham & tormala, 2017). moreover, if the relationship was linear, the redundancy in interest and certainty may undermine the utility of considering certainty of interest as separate from interest. 2.1 method design this was a correlational study using a within-participants design in which participants answered questions about their interest and certainty of interest in four domains. 2.1.1 participants the participants were 21 undergraduate students at a mid-sized university in the midwestern united states. they were all in an upper-level psychology course that students typically complete their last year of undergraduate study. they completed the questionnaire in exchange for extra credit. 2.1.2 measures and procedure participants responded to a 4-page, paper-and-pencil survey, in which questions for each domain (math, biology, astronomy, and psychology) were presented on different pages. the order in which participants responded about the different domains was counterbalanced across participants. participants responded to items assessing interest, certainty, and the number of college courses they had completed in the domain as well as other items that are not central to the current research. interest was measured with 6 items that were adapted from those used in prior research and capture both feeling and meaning/value aspects of interest (durik & harackiewicz, 2007). the scale included “i find ___ interesting,” “___is fascinating to me,” “i find ___enjoyable,” “___is a boring subject,” “___just doesn’t appeal to me,” and “i think ___is a meaningful discipline.” wherein the blank spaces were replaced with the domain name. participants rated each item from 1 (strongly disagree) to 7 (strongly agree). cronbach alphas for interest were .94 (math), .90 (biology), .77 (psychology), and .81 (astronomy). the lower reliability observed for psychology is likely due to restriction of range because all participants were highly interested in psychology, which is known to constrain estimates of reliability (nunnally & bernstein, 1994). certainty was assessed with one item, “how sure are you of your attitudes about ___?” and rated from 1 (not at all) to 5 (very much). although only a single item was used, a similar approach has been taken in prior research (gross et al., 1995). participants were also asked, “how many college courses have you taken in ___?” with response options ranging from zero to greater than eleven. this item was included in order to examine whether certainty was related to participants’ prior exposure to the domain. 2.2 results and discussion the data for this study (and both subsequent studies) were analyzed using multiple regression analysis conducted in spss version 25.0. for each domain, certainty was designated as the criterion variable. the measure of interest in each domain was standardized and a squared term was calculated by multiplying the standardized measure by itself. these two variables, the standardized measure of interest and its square, were entered into the regression simultaneously to predict certainty of interest for that domain. for each effect, squared semi-partial correlations are provided as measures of effect size. these denote the portion of total variability in the outcome variable that is uniquely accounted for by a given predictor. 2.2.1 math the analysis predicting certainty of math interest revealed a negative average relationship of interest, t(18) = -2.48, p = .02, b = -0.42, sr2 = .23, and a positive quadratic relationship, t(18) = 2.54, p = .02, b = 0.38, sr2 = .24. the top left panel of figure 1 depicts the relationship, showing that certainty was higher for those reporting either low or high levels of interest, and lower for those reporting more moderate levels of interest. 2.2.2. biology the analysis predicting certainty of biology interest revealed no relationship of interest, t(18) = 0.68, p = .51, b = 0.12, sr2 = .02, but similar to math, yielded a positive quadratic relationship, t(18) = 3.12, p < .01, b = 0.65, sr2 = .35. the bottom left panel of figure 1 depicts the relationship. similar to what was observed in math, participants who reported lower or higher levels of interest in biology also reported greater certainty, compared with those who reported more moderate levels of interest. 2.2.3 astronomy the model used to predict certainty of astronomy interest revealed a different pattern. neither a linear relationship, t(18) = 0.68, p = .50, b = 0.10, sr2 = .02, nor a quadratic relationship, t(18) = -0.18, p = .86, b = -0.02, sr2 < .01, emerged (see top right panel of figure 1). 2.2.4 psychology the regression predicting certainty of psychology interest yielded a positive linear relationship of interest, t(18) = 3.34, p < .01, b = 0.24, sr2 = .12, as well as a negative quadratic relationship, t(18) = -2.56, p = .02, b = -0.19, sr2 = .07. the bottom right panel of figure 1 reveals a different quadratic relationship than was observed for math and biology. among this sample of upper-level psychology students, interest appears to be positively related to certainty, and then levels off at the highest levels of interest and certainty. figure 1. curvilinear relationships tested between interest and certainty for each domain in the pilot study. math and biology (left panels) revealed statistically significant (p < .05) positive quadratic relationships, psychology (bottom right) revealed a significant negative quadratic relationship, and astronomy revealed no relationship (top right). overall these analyses of the relationships among level of interest and certainty reveals several things. first, the relationship between interest and certainty varies considerably by domain, such that negative quadratic relationships are observed in this sample for math and biology, a positive quadratic relationship was observed in this sample for psychology, and no relationship was observed for astronomy. these varied relationships are likely due to both the amount of experience these participants have with each domain, as well as their high interest in psychology due to the fact that participants were sampled from an upper-level psychology class. it seems likely that the quadratic relationships observed in math and biology are representative of most domains in which individuals have sufficient exposure to the domain to report their interest, and assuming that the sample includes the full range of interest. in most cases when participants are asked to report their interest, they have sufficient experience from which to draw conclusions about and be aware of their level of interest (either high or low) in the domain. in contrast, we interpreted the absence of relationships found in the astronomy domain as likely due to participants having had little prior exposure to astronomy. lack of experience likely limited both their level of certainty as well as their extremity of interest in the domain. finally, the results for psychology revealed a different quadratic pattern from the other three and reflects this sample’s high certainty and high interest in psychology. to examine whether certainty did covary with participants’ prior experience, a final analysis was performed in which reports of certainty and the number of college courses reported was aggregated across participants. a correlation was calculated between the average number of courses students reported for each domain and the average level of certainty for each of the four domains. the correlation of the aggregated measures revealed a strong and positive association, r(2) = .99, p < .01, suggesting that the number of reported courses among participants in this sample was strongly and positively associated with level of certainty. one could imagine that a more thorough measure of prior experience, including courses taken in secondary school or informal learning opportunities, would add further insight into this relationship. although the sample size for this pilot study was extremely small and only included participants who were highly interested in psychology, the results were promising enough to explore further. studies 1 and 2 were designed to examine these relationships more in depth within two domains, math and psychology, among participants drawn from a more general population. 3. study 1 study 1 was designed to replicate the results observed in the pilot study in the domain of math. given that certainty of interest is likely to increase as individuals have more exposure to domains, and that students are exposed to years of compulsory math in primary and secondary school, we expected the negative quadratic relationship between interest and certainty that was observed in the pilot study to also emerge in study 1. the second purpose of study 1 was to test the relationship between self-reported interest and behavior for those with lower or higher certainty. we hypothesized that self-reported interest in the domain would predict behavior in the domain positively and more strongly if individuals were more versus less certain of their interest. 3.1 method 3.1.1 design this was a correlational study that took place in a laboratory context. participants’ math interest, certainty of math interest, and math-related behaviors were assessed in a single session. 3.1.2 participants the participants were 138 undergraduate students (54% women) completing an introductory psychology course at a mid-sized university in the midwest united states. they participated in the study for partial course credit. the sample included participants who reported their race or ethnicity as african american (24%), hispanic (16%), asian (9%), caucasian (48%), or as another, unlisted category (3%). 3.1.3 measures and procedure participants were invited into the lab individually and completed the measures using medialab software (jarvis, 2004). first, participants reported their interest in math using five items adapted from prior research (harackiewicz et al. 2008), including “i've always been fascinated by mathematics,” “i'm really excited about learning mathematics,” “i'm really looking forward to learning more about mathematics,” “i think mathematics is an important discipline,” and “i think mathematics is important for me to know.” participants responded to each item from 1 (strongly disagree) to 7 (strongly agree). the internal consistency of the scale items was strong (cronbach’s alpha = .90). participants responded to 6 items designed to assess their certainty of interest (krosnick et al., 1993). these items were, “how certain are you of your feelings toward mathematics?”, “how sure are you that your opinion of mathematics is correct?”, “how firm are your opinions of mathematics?”, “how easily could your opinion of mathematics be changed?” (reversed), “how definite are your views of mathematics?”, and “how convinced are you of your views of mathematics?” from 1 (not at all ___) to 7 (very ___), in which the blank restated the word in the question that was presented in capital letters. the reversed item was omitted because it decreased the internal consistency of the scale, which left 5 items (final cronbach’s alpha = .91). participants also had the opportunity to report their engagement in math-related behaviors. behavioral indicators are often influenced by many factors in a given situation so the three behavioral indicators were combined into a composite after being standardized. one set of items asked participants to reflect on the past two years and indicate whether or not they had voluntarily chosen to engage in 12 activities related to math, including “i have surfed a website about mathematics in my spare time,” “i have voluntarily discussed topics related to mathematics with friends or family,” “i have chosen to join a club related to mathematics,” and “i have spent free time reading a magazine article about mathematics.” the number of behaviors indicated were summed for each participant. two additional behavioral measures occurred during the session. participants were given a set of 10 math-related topics (e.g., “polygons and figures,” “pi,” “statistics,” “quadratic equations”) and asked to mark any about which they would like to receive more information via email (and to provide their contact information to do so). participants were also given the opportunity to watch any of 5 short video clips about math-related topics (e.g., pi, mental math techniques, square roots). the number of topics marked and the number of videos watched were summed and each served as an additional measure of behavior. the standardized scores for all three measures were averaged to obtain the behavior composite (cronbach’s alpha = .69). 3.2 results and discussion an exploratory factor analysis using oblique rotation was conducted to explore whether the measure of certainty was different from that of interest. two eigenvalues over 1 emerged and the pattern matrix showed two factors with fairly simple structure. each item had a loading of at least .74 on its expected factor and no loading over .07 on the unexpected factor. the next analysis focused on testing the relationship between level of interest in math and certainty of math interest. as was done in the pilot study, this was achieved by conducting a multiple regression analysis in which certainty served as the criterion variable, and a standardized measure of interest as well as its square served as the two predictors. this analysis revealed both a positive relationship of interest, t(135) = 4.68, p < .01, b = 0.38, sr2 = .12. as well as a positive quadratic relationship, t(135) = 6.02, p < .01, b = 0.47, sr2 = .20. comparable to the pattern that emerged in the pilot study with regard to math, those with either lower or higher levels of interest in math also reported more certainty. in contrast, those who reported a moderate amount of interest in math reported lower certainty (see figure 2). figure 2. curvilinear relationship observed between interest in math and certainty in study 1. the second analysis focused on whether the relationship between math interest and behavior would be positive and stronger for participants who reported greater certainty. to this end, a second regression analysis was conducted. the criterion variable was the composite measure of behavior and the three predictors included standardized measures of math interest, certainty of math interest, and their product. the analysis yielded a strong positive relationship of interest, t(134) = 4.67, p < .01, b = 0.32, sr2 = .12, and the predicted interaction, t(134) = 2.57, p = .01, b = 0.17, sr2 = .04, indicating that the relationship between interest and behavior varied depending on level of certainty. certainty was not a significant predictor. simple slope analyses were conducted to examine the relationship between interest and behavior separately for participants reporting certainty that was one standard deviation above and below the mean. the relationship between math interest and behavior was significant and positive for those with high certainty, t(134) = 7.12, p < .01, b = 0.49, but not significant for those with low certainty, t(134) = 1.28, p = .20, b = 0.15 (see figure 3). figure 3. interaction depicting different relationships between math interest and behavior for those reporting lower (one sd below the mean) and higher (one sd above the mean) certainty in study 1. these data replicate the pattern observed in the pilot study between math interest and certainty, and also showed that individuals who are more certain of their interest are more likely to behave in ways that are consistent with their levels of interest. in other words, those who are more certain of having lower interest are less likely to engage whereas those who are more certain of having higher interest are especially likely to engage. an interesting picture begins to emerge with regard to participants who are less certain. what is learned from study 1 is that these participants’ behavior is not as closely associated with their interest. 4. study 2 study 1 offered additional evidence that individuals’ self-reported interest and certainty vary in a curvilinear way, and that participants with higher certainty are more likely to behave in ways that are consistent with their level of reported interest. study 2 was designed to test this again but in a different domain, psychology. psychology was chosen for two reasons. first, the pilot study showed that interest in psychology and certainty showed a different relationship (a negative quadratic relationship) in contrast to the positive quadratic relationship observed in study 1 as well as the two other domains examined in the pilot study. we suspected that the observed relationship between interest and certainty for psychology found in the pilot study was due to the sample being composed entirely of highly interested psychology students nearing the completion of their degree. that said, it could instead be due to the domain itself. we wanted to examine the relationship between interest and certainty in psychology, but with a more general sample. second, psychology offered the possibility of testing the ideas related to interest and certainty with regard to individual interest as well as situational interest, given that the students were completing introductory psychology at the time that the data were collected. whereas individual interest is thought of as an enduring person characteristic, situational interest refers to interest that is triggered by cues in the environment (hidi & renninger, 2006; mitchell, 1993; schraw & lehman, 2001). although individual interest and situational interest are often highly correlated, we reasoned that certainty may function differently for these two types of interest. among students enrolled in an introductory psychology class, it was possible to examine overall individual interest in psychology as well as situational interest in the class, and compare how each measure of interest predicted behavior when taking into account level of certainty. we did not make hypotheses about whether one measure of interest (individual or situational) would predict behavior more strongly than the other. 4.1 method 4.1.1 design this was a correlational study in which participants’ interest in psychology, certainty of their interest, and behaviors related to the domain of psychology were assessed. 4.1.2 participants the participants included 142 undergraduate students (55% women) completing an introductory psychology course at a mid-sized university in the midwest united states. they participated in the study for partial course credit. participants in the sample reported their race or ethnicity as african american (20%), hispanic (10%), both african american and hispanic (1%), asian (3%), caucasian (65%), or as another, unlisted category (1%). one person did not respond to the question about race/ethnicity. 4.1.3 measures and procedure there were two measures of interest, both individual and situational. to report individual interest, participants rated whether “psychology is…” “interesting,” “stimulating,” “boring” (reversed), “engaging,” “meaningful,” “worthless” (reversed), and “useful” from 1 (not at all) to 7 (very much; schiefele, 1990). situational interest was measured with eight items reflecting the students’ level of interest in their introduction to psychology course (e.g., “what we are learning in psychology class this semester is fascinating to me” and “we are learning valuable things in psychology class this semester”; linnenbrink-garcia et al., 2010). participants responded to each item on a scale from 1 (strongly disagree) to 7 (strongly agree). the cronbach’s alphas for individual and situational interest were .88 and .95, respectively. participants reported their certainty immediately following both measures of interest. the items that measured certainty were the same as in study 1 and were general in that they did not specify whether participants should report certainty of individual or situational interest. cronbach’s alpha for the certainty measure was .88. at the end of the survey, reports of behavior were measured with two types of items, and the two types were standardized and combined into a composite in the same way as in study 1. in parallel with study 1, participants were asked whether or not they had engaged in 11 psychology-related behaviors in the past two years (e.g., “i have chosen to join a club related to psychology,” “i have spent free time reading a magazine article about psychology”). participants were also given the option of indicating whether they would like to receive information about various psychology-related topics, and if so, to mark their topic choices and provide their email address so this information could be sent. fifteen topics were listed, designed to capture broad areas of psychology (e.g., “how the brain works,” “mental illness,” “how memories form,” and “stereotyping and prejudice”). the total behaviors indicated and the total number of topics selected were both standardized and then averaged to form the composite measure of behavior. given that there were only two types of behaviors, cronbach’s alpha for this measure was modest equaling .48, which may attenuate the relationships observed between interest and behavior. in contrast to study 1, the option to offer participants videos to watch was not possible because study 2 used pencil-and-paper surveys. 4.2 results and discussion 4.2.1 individual interest as in study 1, an exploratory factor analysis was conducted on the certainty and interest items in order to evaluate their structure. the structure was not as clean as in study 1, due to the two interest measures (individual and situational) having items that shared variability. this is not terribly surprising given the similarity in the constructs, methods of measurement, and timing. that said, the certainty items tended to load together and separately from the interest items, again attesting to the uniqueness of the certainty measure as a complement to typical measures of interest. we proceeded with the two separate measures of interest given their conceptual distinction but also recognize that the similarity in their measurement may hinder the ability to see predictive differences across the two measures. as in study 1, the first analysis was designed to test the relationship between level of interest and certainty, which was then followed by an analysis to test the relationship between self-reported interest and behavior, with the addition of certainty as a moderator. a multiple regression model was tested using certainty of interest as the criterion variable, and interest and its square as the predictors. as in study 1, the measures of interest were standardized prior to calculating the squared term. replicating the relationship observed in study 1 with the domain of math, individual interest in psychology had both a linear, t(139) = 5.68, p < .01, b = 0.53, sr2 = .18, and quadratic relationship with certainty, t(139) = 4.00, p < .01, b = 0.22, sr2 = .09. similar to math, and unlike the relationship observed in the pilot study, certainty was higher for those with lower or higher interest in psychology, and lower for those who reported more moderate interest. next, to test whether interest was a stronger predictor of behavior for those with more certain interest, a multiple regression model was tested in which the behavior composite was the criterion variable and the three predictors were the standardized composite measure of individual interest, the standardized measure of certainty, and their product. this analyses revealed a positive relationship between interest and behavior, t(138) = 5.09, p < .01, b = 0.34, sr2 = .10, as well as an interaction, t(138) = 1.99, p < .05, b = 0.12, sr2 = .02 (see figure 4). simple slope tests were conducted to examine the relationship between interest and behavior for those scoring one standard deviation below and above the mean of certainty. these analyses revealed that the relationship between individual interest and behavior was significant and positive for both, but stronger for those with higher certainty, t(138) = 6.04, p < .01, b = 0.46, than for those with lower certainty, t(138) = 2.17, p = .03, b = 0.22. figure 4. interaction depicting different relationships between psychology interest and behavior for those reporting lower (one sd below the mean) and higher (one sd above the mean) certainty in study 2. these relationships replicate the patterns that were observed in study 1. furthermore, the pattern between interest in psychology and certainty that was observed in the pilot study seems to have been due to the sampling of participants for the pilot study and not to the domain. 4.2.2 situational interest the next set of analyses were parallel to those described for individual interest, but instead focused on situational interest. when situational interest in the psychology class was used to predict certainty, both a linear relationship, t(139) = 5.48, p < .01, b = 0.48, sr2 = .17, and a quadratic relationship, t(139) = 3.17, p = .02, b = 0.21, sr2 = .06, emerged. participants who reported more extreme levels of situational interest also reported greater certainty whereas those who reported more moderate levels of situational interest reported less certainty. when situational interest, certainty, and the interaction between them were used to predict the composite measure of behavior, both situational interest, t(138) = 2.71, p < .01, b = 0.19, sr2 = .05, and their interaction were significant, t(138) = 2.29, p = .02, b = 0.15, sr2 = .03. simple slopes were tested and showed that situational interest positively predicted behavior for those with higher certainty, t(138) = 4.10, p < .01, b = 0.34, but did not predict behavior for those with lower certainty, t(138) = 0.40, p = .69, b = 0.04. the results from both individual and situational interest demonstrate that certainty may be useful in better predicting behavior with both types of interest. if anything, the interaction was slightly more pronounced for situational than for individual interest, although this difference was not tested directly. one explanation is that the salience of an ongoing situation may be a vivid motivator of behavior toward (or away from) domain content. however, this pattern is also tied to the nature of the situation assessed here. given that the situation that defined situational interest in this study was in reference to a semester-long, introductory class, situational interest as well as the behavior were in reference to a fairly broad definition of the field in general. other situations that are more narrow (e.g., focused on a particular topic) may not predict behavior as strongly, especially if the behavioral opportunity is more broad or more narrow than the experience of situational interest. 5. general discussion this research lays out some of the complexity inherent in asking individuals to self-report their interest and recognizes that these self-reports may be more or less certain. for domains in which individuals have varied amounts of experiences and sufficient prior exposure (e.g., math, psychology, and biology), there appears to be a curvilinear relationship between self-reported interest and self-reported certainty. individuals who reported more extreme levels of interest (either high or low) tended to report being more certain of their interest. in other words, when individuals were quite interested in a domain, they were sure of the presence of their interest; when individuals were quite disinterested in a domain, they were also sure of their lack of interest. moreover, as predicted, level of interest was a stronger predictor of behavior when certainty was high than when certainty was low. it is not clear from these data whether participants who are less certain (and more moderate in their interest) have truly neutral beliefs about the domain or if they actually have both positive and negative experiences, which are best captured as neutral on this bidirectional scale. 5.1 promises and pitfalls of self-report measures certainty of interest highlights both a promise and a pitfall of using self-report measures to assess interest. this research speaks to two of the three main questions guiding the compilation of this special issue. first, this research relates to the complexity of interpreting self-report data. on the one hand, if a sample has high certainty of their interest, then self-report measures may be easier to interpret. the certainty with which people report their interest may support processes that strengthen associations between interest and various behavioral outcomes. in considering this promise, however, it is also critical to consider the pitfall; when people are not certain, they will still provide a response but that response may not be rooted in as much experience, which could make interpretation difficult. if participants are asked to provide reports of interest in domains in which they have little experience, their reports will be ill-informed. this challenge adds to the challenge of inattentiveness, examined by iaconelli and wolters (2020). the interpretation of self-report data will likely be murky not only when people are responding inattentively, but also when they are attentive but do not have sufficient self-knowledge or experience with which to provide a meaningful response. although the current work centers on the certainty with which individuals can identify their interests, certainty may extend to other types of reports as well, such as certainty in metacognition and strategy use. other authors in this issue (e.g., rogiers et al. 2020, van halem et al.,2020), report how self-reported measures link to trace measures that occur during task engagement, and to subsequent behavior. these relationships may be stronger to the extent that participants are more certain of their self-reported meta-cognition and strategy use. second, this research also highlights constraints of self-report data that have implications for methodology. self-reported domain interest may be less valid when interest is in early phases of development (e.g., renninger & hidi, 2011), and this research suggests that certainty may capture an important element. it may take months and years for individuals to collect information about their response to a given domain in order to have certainty in their level of interest (either high or low). presumably, the development of this certainty will emerge as individuals have experiences with a domain and then think back to them retrospectively (see dinsmore et al., this issue for a discussion of retrospective processes). if interest must be measured among a sample with limited certainty of interest, one approach may be to provide supports for them to know how to respond (e.g., definitions of the domain, a particular experience to reflect on) in order to provide a more valid measure, or to assess interest multiple times during a task (moeller et al., 2020) rather than relying on a global assessment of domain interest. alternatively, researchers may want to consider certainty when selecting domains to study and identifying appropriate samples within those domains. it is noteworthy that the sample in the pilot study had relatively little certainty about their interest in astronomy and very strong certainty about their interest in psychology. given the results of studies 1 and 2, this might also have implications for behavior. if the purpose of a research study is to use interest to predict behavior, then it may be wise to select a domain in which participants have considerable experience in the domain (i.e., math), because interest may predict behavior if the sample as a whole has more experience with the domain. that said, it is also important to have variability in the sample so that it represents the full range of interest, with which to predict behavior. the pilot study included a sample composed entirely of students in their last year of studying psychology. although this sample reported very high certainty, restriction of range in their interest is likely to have limited any observed association between level of interest and behavior, had behavior been assessed. 5.2 implications for theory the observed relationship between interest and certainty may be related to participants’ experiences with domain content, and captures the changing aspect of interest within a developmental trajectory. data from the pilot study showed that the number of classes students took in each domain positively predicted certainty in a linear fashion. prior experience also varied along with the different observed relationships between interest and certainty across the domains. no reliable relationship was detected in the domain of astronomy, likely because participants had such limited experiences in the domain, and the opposite quadratic relationship was detected in the domain of psychology, presumably because students had extremely high interest and certainty. although these fluctuations are consistent with research on attitudes showing that direct experience contributes to attitude certainty (see review by glasman & albarracin, 2006), the nuances of how this occurs is important to consider for understanding how individuals come to recognize and be aware of their interest. one possibility is that the extremity of the emotion prompts awareness of the experience and contributes to certainty. individuals who have relatively intense experiences in a domain—experiences that are either highly interesting or highly distasteful—may come to realize their level of interest and have greater certainty (dutta et al., 1972). in contrast, those who have more mixed or vague reactions may be less aware of their emotional reaction to the domain content, which then leaves them with less certainty of their interest. another possibility is that the clarity or vividness of individuals’ memories contributes to certainty. for example, those who have had multiple, semester-long courses in the domain may have more vivid memories of learning in that domain, which may contribute to greater certainty. along these lines, it may be fruitful to bring research on autobiographical memory into research on interest in order to better understand how individuals recall prior experiences that may inform their report of interest. research suggests that the valence of how an event ends has a disproportionate influence on how the event is remembered (kahneman et al., 1993). research on interest may benefit by building on this foundation from the memory literature in order to better understand not only how interest develops, but whether the timing of experiences impacts how people recall the events that they come to see as foundational to their perception of interest in the domain. 5.3 implications for research certainty may also provide a handle for predicting whose interest may be more or less altered by new experiences with a domain. for example, individuals who are less certain of their interest may be more responsive to situational variables designed to affect interest. less certain individuals may be more open to collecting information through their experiences with domain content and updating their level of interest. if so, situational enhancements designed to foster interest may be more effective for low versus high certain individuals. as such, it may be beneficial to assess certainty of interest in research testing interventions designed to foster interest. the effects of a situational intervention may be positive for a subset of individuals (i.e., those with lower certainty), but this effect may be masked by others who are more certain and therefore less likely to change their level of interest. in a sense, if one thinks of situational interest as being analogous to attitudes in the face of persuasive attempts, then the attitudes literature provides some hints of this possibility as well. for example, tormala (2016) has noted how curiosity is itself a form of “interested uncertainty” which, in some situations, can aid persuasion. specifically, kupor and tormala (2015, study 3) reason that curiosity motivates more thorough processing of the persuasive message and increases the message’s impact. this may suggest that a similar process could explain differential impact of situational enhancements that are designed to foster interest. in general, future directions focused on underlying cognitive and affective processes are warranted and would very much help illuminate how interest may or may not change in response to new experiences. 5.4 differences between attitudes and interest although this research was initiated because of similarity between interest and attitudes, these constructs are not the same, which has implications for how they emerge and function. for example, certainty in the attitude literature has been divided into correctness and clarity (petrocelli et al., 2007). whereas correctness reflects the veraciousness of an attitude in an absolute sense (e.g., efforts to slow global warming is the correct perspective), clarity is the extent to which individuals are confident that they know their own stance (i.e., i know my attitude about efforts to slow global warming). applied to interest, clarity is more relevant than correctness if one assumes a stance of relativism, in which interest in one domain is not more valuable or correct than interest in any other domain. a difference between the attitude and interest literatures also emerges when one considers what is being evaluated when self-assessing an attitude versus self-assessing level of interest. attitude objects are typically outside the person and the question is whether the person agrees with or disagrees with the attitude object. this is in contrast to interest in which the person assesses how they interact with the domain. when people evaluate their interest, they evaluate their personal experiences with it and how they respond to the domain. as such, assessments of an attitude object may be more about the object and less about how the person interacts with the object. for interest, in contrast, the assessment is about one’s own response to the domain. 5.5 limitations finally, although positive relationships between interest and behavior were observed in the domains of math and psychology, the direction of causality is not clear in the present correlational studies. the theoretical model describing the relationship between interest and behavior typically places interest as the motivation (i.e., cause) of behavior; however, it is also worth pointing out that behavior may influence self-reports of interest as well as certainty. research on attitudes has addressed several of these processes (see review by olson & stone, 2005). when individuals are unsure of their attitudes, they may reflect on their past behavior as information that can be used to inform their attitude (e.g., bem, 1967). for example, when asked about global warming, individuals may scan their memories for events in which they chose to engage in pro-environment activities or not. these memories may lead them to decide on a particular self-reported attitude. a similar process may operate with regard to interest, especially when individuals have less certainty of their interest. for example, when participants in the pilot study were asked about their interest in astronomy, the domain in which they had taken the fewest classes, they may have tried to recall relevant memories. those who could generate more positive memories are likely to have rated their interest higher than those who could generate fewer memories, or negative memories. it is also worthwhile to consider how behaviors may affect self-reported interest ratings, which can also explain the observed relationships between interest and behavior found in this set of studies. in studies 1 and 2, participants first reported their interest, then their certainty of interest, and finally responded to behavioral opportunities. it is possible that participants felt pressure to behave in a way that was consistent with their initial reports, which may have strengthened the observed results (olson & stone, 2005). specifically, those who had just reported low or high levels of interest may have felt internal pressure to behave in ways that were consistent with their reports, and this may have been especially strong for those who reported greater certainty. the present research does not address this possibility, but opens up a line of research in which these ideas could be explored. finally, specific features of the studies reported here also warrant caution in drawing broad conclusions. the pilot study examined prior exposure to various domains by collecting information on students’ course experiences. however, experiences in high school courses and extracurricular activities may have also provided opportunities for exposure. inclusion of all these experiences could help paint a fuller picture of the relationship between domain exposure and certainty. furthermore, these studies involve a very narrow sample of individuals—namely, undergraduate students at a single university who are taking a psychology course. for example, it is healthy to question whether similar patterns would be observed among a sample of older adults (e.g., who may have more experience and greater certainty) or younger children (i.e., who may have even less experience than the sample tested in the current research). moreover, these psychology students may have been especially sensitive to environment cues or demand characteristics related to behaving in ways that are more consistent with their stated interest. this tendency would be exaggerated if participants felt that the desired response involved giving higher interest ratings and engaging in more behaviors. these questions cannot be answered with the current data but could be tested in the future. 5.6 concluding thoughts in summary, when interest is self-reported, the report reflects a synthesis of what individuals have available to them at the moment of measurement, and these reports are more certain for some participants than for others. this variation in certainty may provide a lens for better understanding how interest develops and how it is internalized and becomes known to individuals. as with any type of measure, it is critical to interpret the data in light of the assumptions, capacities, limitations, and processes that are relevant at the time of measurement. keypoints self-reported interest varies in certainty across individuals those with greater certainty self-report more extreme interest (high or low) self-reported interest and behaviour correlate more strongly for those with greater certainty references ainley, m., & ainley, j. (2011). a cultural perspective on the structure of student interest in science. international journal of science education, 33, 51-71. https://doi.org/10.1080/09500693.2010.518640 ainley, m., hidi, s., & berndorff, d. (2002). interest, learning, and the psychological processes that mediate their relationship. journal of educational psychology, 94, 545-561. https://doi.org/10.1037//0022-0663.94.3.545 ajzen, i. (1991). the theory of planned behavior. organizational behavior and human decision processes, 50, 179-211. https://doi.org/10.1016/0749-5978(91)90020-t bem, d. j. (1967). self-perception: an alternative interpretation of cognitive dissonance phenomena. psychological review, 74, 183-200. https://doi.org/10.1037/h0024835 berger, i. e. (1999). the influence of advertising frequency on attitude-behavior consistency: a memory based analysis. journal of social behavior & personality, 14, 547–568. bizer, g. y., tormala, z. l., rucker, d. d., & petty, r. e. (2006). memory-based versus on-line processing: implications for attitude strength. journal of experimental social psychology, 42, 646–653. https://doi.org/10.1016/j.jesp.2005.09.002 cheatham, l. b., & tormala, z. l. (2017). the curvilinear relationship between attitude certainty and attitudinal advocacy. personality and social psychology bulletin, 43, 3-16. https://doi.org/10.1177/0146167216673349 deci, e. l. (1992). the relation of interest to the motivation of behavior: a self-determination theory perspective. in k. a. renninger, s. hidi, & a. krapp, a. (eds.),the role of interest in learning and development (pp. 43-70) . hillsdale, nj: lawrence erlbaum. dutta, s., kanungo, r. n; & freibergs, v. (1972). retention of affective material: effects of intensity of affect on retrieval. journal of personality and social psychology, 23, 64–80. https://doi.org/10.1037/h0032790 fazio, r. h., & zanna, m. p. (1978). attitudinal qualities relating to the strength of the attitude– behavior relationship. journal of experimental social psychology, 14, 398–408. https://doi.org/10.1016/0022-1031(78)90035-5 glasman, l. r., & albarracin, d. (2006). forming attitudes that predict future behavior: a meta-analysis of the attitude-behavior relation. psychological bulletin, 132, 778-822. https://doi.org/10.1037/0033-2909.132.5.778 gross, s. r., holtz, r., & miller, n. (1995). attitude certainty. in r. e. petty & j. a. krosnick (eds), attitude strength: antecedents and consequences (pp.215-245). mahwah, nj: lawrence erlbaum. harackiewicz, j. m., durik, a. m., barron, k. e., linnenbrink-garcia, l., & tauer, j. m. (2008). the role of achievement goals in the development of interest: reciprocal relations between achievement goals, interest and performance. journal of educational psychology, 100(1), 105-122. https://doi.org/10.1037/0022-0663.100.1.105 hidi, s. & renninger, k.a. (2006). the four-phase model of interest development. educational psychologist, 41, 111-127. https://doi.org/10.1207/s15326985ep4102_4 iaconelli, r. & wolters c.a. (2020). insufficient effort responding in surveys assessing self-regulated learning: nuisance or fatal flaw? frontline learning research. 8 (3) 104 – 125. https://doi.org/10.14786/flr.v8i3.521 jarvis, w. b. g. (2004). medialab [computer software]. new york: empirisoft. jonas, k., diehl, m., & broemer, p. (1997). effects of attitudinal ambivalence on information processing and attitude-intention consistency. journal of experimental social psychology, 33, 190–210. https://doi.org/10.1006/jesp.1996.1317 kahneman, d., fredrickson, b. l., schreiber, c. a., & redelmeier, d. a. (1993). when more pain is preferred to less: adding a better end. psychological science. 4, 401–405. https://doi.org/10.1111/j.1467-9280.1993.tb00589.x krapp, a. (2002). an educational-psychological theory of interest and its relation to sdt. in. e. l. deci & r. m. ryan (eds.), handbook of self-determination research (pp. 405-427). rochester, ny: university of rochester press. krapp, a., & prenzel, m. (2011). research on interest in science: theories, methods, and findings. international journal of science education, 33, 27-50. https://doi.org/10.1080/09500693.2010.518645 krishnan, h. s., & smith, r. e. (1998). the relative endurance of attitudes, confidence and attitude-behavior consistency: the role of information source and delay. journal of consumer psychology, 7, 273–298. https://doi.org/10.1207/s15327663jcp0703_03 krosnick, j. a., boninger, d. s., chuang, y. c., berent, m. k., & carnot, c. g. (1993). attitude strength: one construct or many related constructs? journal of personality and social psychology, 65, 1132-1151. https://doi.org/ 10.1037/0022-3514.65.6.1132 krosnick, j. a., & petty, r. e. (1995). attitude strength: an overview. in r. e. petty & j. a. krosnick (eds.), attitude strength: antecedents and consequences (pp. 1-24). mahwah, nj: lawrence erlbaum. kupor, d. m., & tormala, z. l. (2015). persuasion, interrupted: the effect of momentary interruptions on message processing and persuasion. journal of consumer research, 42, 300–315. https://doi.org/ 10.1093/jcr/ucv018 linnenbrink-garcia, l., durik, a. m., conley, a. m., barron, k. e., tauer, j. m., karabenick, s. a., & harackiewicz, j. m. (2010). measuring situational interest in academic domains. educational and psychological measurement, 70, 647-671. https://doi.org/10.1177/0013164409355699 mitchell, m. (1993). situational interest: its multifaceted structure in the secondary school mathematics classroom. journal of educational psychology, 85, 424–436. https://doi.org/10.1037/0022-0663.85.3.424 moeller, j., viljaranta, j., kracke, b., & dietrich, j. (2020). disentangling objective characteristics of learning situations from subjective perceptions thereof, using an experience sampling method design. frontline learning research. 8 (3) 63-84. https://doi.org/10.14786/flr.v8i3.529 nunnally, j. c., & bernstein, i. a. (1994). psychometric theory (3rd ed.). new york: mcgraw-hill. olson, j. m., & stone, j. (2005). the influence of behavior on attitudes. in d. albarracin, b. t. johnson, & m. p. zanna (eds.), the handbook of attitudes (pp. 223-271). mahwah, nj: lawrence erlbaum. petrocelli j. v., tormala, z. l., & rucker, d. d. (2007). unpacking attitude certainty: attitude clarity and attitude correctness. journal of personality and social psychology, 92, 30-41. https://doi.org/10.1037/0022-3514.92.1.30 pomerantz, e. m., chaiken, s., & tordesillas, r. s. (1995). attitude strength and resistance processes. journal of personality and social psychology, 69, 408-419. https://doi.org/10.1037/0022-3514.69.3.408 prislin, r., wood, w., & pool, g. j. (1998). structural consistency and the deduction of novel from existing attitudes. journal of experimental social psychology, 34, 66–89. https://doi.org/10.1006/jesp.1997.1343 renninger, k. a. (1990). children’s play interests, representation, and activity. in r. fivush & k. hudson (eds.), knowing and remembering in young children (pp. 127-165). new york: cambridge university press. renninger, k. a. (2000). individual interest and its implications for understanding intrinsic motivation. in c. sansone and j. m. harackiewicz (eds.), intrinsic and extrinsic motivation: the search for optimal motivation and performance (pp. 373-404). san diego, ca: academic press, inc. renninger, k. a., & hidi, s. (2011). revisiting the conceptualization, measurement, and generation of interest. educational psychologist, 46, 168-184. https://doi.org/10.1080/00461520.2011.587723 renninger, k. a., hidi, s., & krapp, a. (1992). the role of interest in learning and development. hillsdale, nj: lawrence erlbaum. renninger, k. a., & su, s. (2012). interest and its development. in r. m. ryan (ed.), the oxford handbook of motivation (pp. 167-187). oxford: oxford university. rogiers, a., merchie, e., & van keer h. (2020). opening the black box of students’ text-learning processes: a process mining perspective. frontline learning research. 8(3), 40–62. https://doi.org/10.14786/flr.v8i3.527 rucker, d. d., tormala, z. l., petty, r. e., & briñol, p. (2014). consumer conviction and commitment: an appraisal-based framework for attitude certainty. journal of consumer psychology, 24 (1), 119-136. https://doi.org/10.1016/j.jcps.2013.07.001 schiefele, u. (1991). interest, learning, and motivation. educational psychologist, 26, 299-323. https://doi.org/10.1080/00461520.1991.9653136 schiefele, u. (1999). interest and learning from text. scientific studies of reading, 3, 257-279. https://doi.org/10.1207/s1532799xssr0303_4 schraw, g., & lehman, s. (2001). situational interest: a review of the literature and directions for future research. educational psychology review, 13, 23-52. https://doi.org/10.1023/a:1009004801455 silvia, p. j. (2005). what is interesting? exploring the appraisal structure of interest. emotion, 5, 89–102. https://doi.org/10.1037/1528-3542.5.1.89 silvia, p. j. (2006). exploring the psychology of interest. new york: oxford university press. silvia, p.j. (2008). appraisal components and emotion traits: examining the appraisal basis of trait curiosity. cognition and emotion, 22, 94–113. https://doi.org/10.1080/02699930701298481 simpkins, s. d., davis-kean, p. e., & eccles, j. s. (2006). math and science motivation: a longitudinal examination of the links between choices and beliefs. developmental psychology, 42, 70-83. https://doi.org/10.1037/0012-1649.42.1.70 tormala, z. l. (2016). the role of certainty (and uncertainty) in attitudes and persuasion. current opinion in psychology, 10, 6-11. https://doi.org/10.1016/j.copsyc.2015.10.017 tormala, z. l., & rucker, d. d. (2007). attitude certainty: a review of past findings and emerging perspectives. social and personality psychology compass, 1, 469-492. https://doi.org/10.1111/j.1751-9004.2007.00025.x van halem, n., van klaveren, c., drachsler h., schmitz, m., & cornelisz, i. (2020). tracking patterns in self-regulated learning using students’ self-reports and online trace data. frontline learning research. 8 (3) 140-163. https://doi.org/10.14786/flr.v8i3.497 wijnia, l., loyens, s. m. m., derous, e., & schmidt, h. g. (2014). do students’ topic interest and tutors’ instructional style matter in problem-based learning? journal of educational psychology, 106, 919-933. https://doi.org/10.1037/a0037119 frontline learning research special issue: vol. 13 no. 2 (2025) 122 -134 issn 2295-3159 corresponding author details: national and kapodistrian university of athens, greece email: ankyriak@ecd.uoa.gr doi: https://doi.org/10.14786/flr.v13i2.1587 commentary: reconceptualizing momentary engagement through the lens of conceptual change learning natassa kyriakopoulou national and kapodistrian university of athens, greece article received 12 september 2024/ article revised 31 october 2024 / accepted 29 november 2024/ available online 14 march 2025 abstract this commentary reviews the five papers featured in this special issue, which foster a cross-disciplinary discussion on momentary engagement (me). the papers represent diverse theoretical perspectives and address key research questions central to understanding students’ me. the commentary approaches each paper through the lens of conceptual change, focusing on the learning processes needed when the information to be acquired is inconsistent with the existing theoretical frameworks. methodological challenges in measuring me within the context of conceptual change are explored, moving beyond traditional acquisition type of learning. the variation in quality and depth of momentary engagement is also discussed, distinguishing between different modes of active learning and engagement. further attention is given to the complex, dual role of factors such as learner characteristics, prior knowledge, and epistemic beliefs in shaping me, particularly in domains requiring radical reorganization of initial beliefs. finally, the potential for constructing an integrated model of me is discussed, in alignment with the holistic approach to me implied by the papers in this issue. the author emphasizes the importance of studying me’s interconnected components within both the individual and the context, employing varied methodologies and accounting for different learning types. the implications of integrating different theoretical frameworks are discussed in relation to developing interventions aimed at enhancing students’ me in the classroom context. keywords: conceptual change; holistic approach; momentary engagement kyriakopoulou special issue: perspectives on momentary engagement and learning situated in classroom contexts 123 | f l r 1. introduction the parable of “the blind men and the elephant”, common across many traditions, illustrates the limits of individual perspectives. each blind man describes only a part of the elephant, without realizing the limited viewpoint. this metaphor highlights the need for multiple perspectives to fully grasp a complex phenomenon and find common ground. in this context, i will discuss the five contributions in this special issue, each addressing students’ momentary engagement (me) from a different theoretical and methodological perspective. these varied perspectives reflect the scope of the earli funded emerging field group integrated model of momentary learning in context (immolic), which aims to synthesize distinct lines of inquiry and identify both differences and useful synergies among various traditions. in this commentary, i will address the significance and potential limitations of various aspects highlighted in each paper focusing on a specific kind of learning, that of conceptual change (cc). emphasis will be placed on a more complex type of cc, where major changes in our knowledge system occur during learning and development, especially following exposure to counter-intuitive science and mathematical concepts. when the information to be acquired contradicts existing beliefs or presuppositions, radical knowledge restructuring is needed (vosniadou, 1994). for example, children struggle to understand the counter-intuitive notion that we live on the outside of a spherical earth, requiring the reorganization of presuppositions such as the earth being flat and stable, unsupported objects fall, or up/down gravity (vosniadou & brewer, 1992). a recent constructivist approach to cc, the framework theory approach, suggests that children form naïve theories before formal science instruction, based on their experiences and cultural context (vosniadou, 2013, 2017; vosniadou et al., 2008). these naïve theories form framework theories, i.e., skeletal structures grounding deep ontological commitments and forming loose but coherent explanatory systems (vosniadou & mason, 2012). because scientific theories differ significantly from the framework theories, learning science requires many conceptual changes involving new modes of knowing, new ways of reasoning and changes in categorization, representation and students’ epistemic beliefs (vosniadou, 2017). unlike posner’s classical approach (posner et al., 1982), cc is not merely the replacement of one theory with another but rather a slow and gradual process during which misconceptions and fragmented conceptions may emerge. indeed, empirical evidence shows that cc can occur without intentional learning, with students changing from intuitive to scientific conceptions without being aware of the change, disrupting the coherence of initial frameworks (vosniadou, 2003). me refers to a student's immediate interaction with a learning task, occurring over brief intervals, in contrast to macro-level engagement, which encompasses extended activities and larger groups. this commentary has a dual focus: to explore me both as a momentary, micro-level phenomenon and as a redefined, holistic construct. at the micro-level, applying cc theory allows us to examine me as a sequence of real-time learning moments, during which students actively confront and attempt to reconcile discrepancies between prior knowledge and new information. in this process, i will also consider the complex interactions among various factors that should be accounted for when measuring me within the cc framework. this perspective offers insights into how learners actively work through conflicts in understanding, shedding light on the pathways of cc. in parallel, at the metalevel, cc theory provides a foundation for reconceptualizing me itself, challenging and expanding current frameworks to develop a more holistic, integrated view. by identifying gaps in existing models of me, i aim to use cc theory to capture me’s complexity as a dynamic, interdependent system. kyriakopoulou special issue: perspectives on momentary engagement and learning situated in classroom contexts 124 | f l r 2. measuring momentary engagement: rethinking the “how” and the “what” in conceptual change type of learning 2.1 methodological reflections: how to measure me if me is conceptualized not simply as the addition of emotional, motivational, cognitive, and behavioral components, but as an integrated meta-construct (fredricks et al., 2004; symonds et al., 2024), it requires measurement approaches that capture this holistic view. haataja and colleagues (this issue) suggest that “multiple data channels can contribute to a better understanding of the multidimensional nature of momentary engagement”. therefore, engagement must be assessed across varying grain sizes (micro to macro) and using multimodal measures to capture its dynamic nature (fredricks & mccolskey, 2012). the five studies explored different aspects and levels of me using diverse methods and measures. haataja and colleagues combined subjective and objective data from natural classroom settings, analyzing cognitive and behavioral me during collaborative learning (cl), with a focus on within-group variations. they used electrodermal activity and self-reports to assess collaboration in real time. renninger and colleagues studied cognitive (executive functions) and behavioral (participation, cooperation) engagement during collaborative math problem-solving, using discourse analysis to track engagement across problem-solving phases. their online collaboration study revealed engagement fluctuations that traditional observations often miss. tang and colleagues employed experience sampling to examine emotional and motivational correlates of optimal learning moments, applying network analysis to explore me relationships. baines and colleagues conducted a multi-method study, assessing engagement at momentary, classroom, and school levels through self-reports, peer questionnaires, and science tests, supplemented by teachers’ ratings of attention and behavior. symonds and colleagues used the oracle tool in a large-scale study involving 92 irish schools to categorize students based on behaviors, social interactions, and proximity to teachers. these methodological tools range from traditional, easy to administer measures (e.g., selfreports) to novel approaches (e.g., physiological measures). regarding the strengths and limitations of the different methods for assessing engagement, it seems important to triangulate engagement through the combination of various measures (fredricks & mccolskey, 2012). the contributions in this issue reveal this potential, as most researchers attempt to integrate different tools. for example, haataja and colleagues found that while physiological synchrony reflected the collaboration process in authentic learning situations, situated subjective self-reports more effectively captured regulatory strategies. they argued that the relationship between physiological synchrony and engagement might not be straightforward, requiring consideration alongside students’ appraisals for a complete understanding. other researchers argued that self-reports, though useful, have limitations such as developmental suitability, interpretation issues, overestimation bias or accuracy concerns (appleton et al, 2006; fulmer & frijters, 2009). they suggested combining self-report data with teacher rating scales (skinner et al., 2009), observational methods although valuable can be susceptible to observer bias, similarly eyetracking and physiological measures need to be complemented with contextual information. for example, in the cc context, mikkilä-erdmann and colleagues (2008) argued that while eye-tracking is useful for revealing cognitive processes during learning from scientific texts, it should be paired with think-aloud methods, interviews and other techniques. triangulating findings from multiple tools ensures more reliable conclusions. when students engage in scientific inquiry, it is also crucial to collect data without disrupting the flow of their activities. the experience sampling method allows real-time data collection in natural contexts, reducing recall bias by prompting students to respond at several points throughout the day (bolger et al., 2003; hektner et al., 2007; inkinen et al., 2020; sinatra et al., 2015). however, the content kyriakopoulou special issue: perspectives on momentary engagement and learning situated in classroom contexts 125 | f l r of these questions and the interpretation of children's responses are important considerations. the same concern applies to the self-reported measures used in the various studies in this issue. for instance, in the cc framework, different methods of questioning, such as forced-choice versus open-ended formats yield different insights. vosniadou and colleagues (2004) showed that different questioning methods affect children’s responses regarding phenomena as the shape of the earth and the day/night cycle, eliciting different forms of knowing and reasoning. open-ended questions require children to generate an explanation (if you walked for many days in a straight line, where would you end up? is there an end or an edge to the earth? would you ever reach the end/edge of the earth? would you fall off that end/edge? why? / why not?) rather than simply recognizing the correct scientific one among alternatives (is the earth round or flat? if it’s round/flat, does it look like a circle or a ball?) (vosniadou & brewer, 1992). different interviewing approaches provide varied insights into how children approach their ideas, become aware of them, reflect on their knowledge, and reveal changes in their underlying beliefs and in the nature of their engagement. renninger and colleagues also measured change and engagement as they unfolded moment-to moment, focusing on collective collaboration outcomes while considering individual variations and different learning patterns. studies in cc learning show that children progress along different pathways, with mental models and conceptual understanding varying significantly (schneider & hardy, 2013; vosniadou & brewer, 1992). schneider and hardy (2013), used latent profile transition analysis, to identify five developmental pathways in third-graders’ understanding of floating and sinking, revealing how prior knowledge influences conceptual development. these findings highlight the importance of considering both qualitative and quantitative individual differences at specific measurement points, as well as changes over time (hickendorff et al., 2018). a microgenetic learning analysis perspective, commonly used in cc and developmental psychology, provides a fine-grained, moment-by-moment view of how engagement evolves. this method detects the small steps learners take throughout the learning process (lee & karmiloff-smith, 2002, parnafes & disessa, 2013; siegler, 2006). the focus is primarily on the child in real time, “understanding the child’s planning, monitoring, self-repairs, and spontaneous comments, charting change as it occurred in the space of a single session” (karmiloff-smith, 2013, p.49). for instance, regarding renninger and colleagues’ research, it would be intriguing to further investigate the specific thinking exhibited by each child and the group in different learning pathways that led to similar performance levels in the final phase of solving the mathematical problem. are there qualitative differences, and where do individual and interindividual variations originate? do these different pathways reveal different levels of understanding and metacognitive reflection? do students engage similarly as they reflect on their initial beliefs and compare them with new concepts? renninger and colleagues assessed both students’ activity during problem-solving and their engagement modes. in the context of cc, it would also be beneficial to examine instances such as cognitive conflicts (e.g., states of disagreement, inconsistencies), analogical reasoning, argumentation type (e.g., negotiation of ideas, discussion of multiple views) and moments of reflection (e.g., self-awareness of cc) (luebeck & bice, 2005). these smaller, incremental, momentary changes could provide deeper insights into students’ conceptual changes and help track their progress in knowledge construction. the complex interaction among various factors should also be considered when measuring engagement. learners’ characteristics, emotional and motivational factors, pre-existing knowledge, epistemic beliefs, and interest in the specific problem situation all play a role. for example, stathopoulou and vosniadou (2007) interviewed two high-achieving physics students with similar grades but different approaches to learning. one demonstrated awareness of his beliefs and integrated new ideas with a meaning making orientation, while the other student was unaware of his beliefs and adopted superficial strategies, such as memorization, with a performance-orientated perspective. their differing epistemic profiles one with a constructivist, physics-related epistemology and the other with a less-constructivist kyriakopoulou special issue: perspectives on momentary engagement and learning situated in classroom contexts 126 | f l r epistemologysignificantly influenced their engagement level. these issues will be further addressed in the second section of the paper and discussed through the prism of cc framework. 2.2 theoretical reflections: modes of engagement measures another important aspect to consider is the quality and depth of engagement examined, which are linked to how active engagement is defined. for instance, a student might be behaviorally engaged but exhibit low interest or weak cognitive or metacognitive engagement (chi & wylie, 2014; renninger & bachrach, 2015). identifying overt behavioral indicators of knowledge change processes is essential for assessing students' engagement during learning (chi & wylie, 2014). in cc learning, simply being “active” may not be sufficient. active behaviors range from merely doing something (e.g., manipulating, underlying, choosing a justification), to actually constructing knowledge, to collaborating in groups in a co-constructive mode (chi & wylie, 2014; vosniadou et al., 2023). chi and colleagues associate three distinct hierarchical modes -active, constructive, and interactivewith "active learning", each implying different underlying processes of knowledge change and resulting in progressively more effective learning (chi et al., 2018; chi & wylie, 2014). conceptual change learning extends beyond simple "active learning" to include a "constructive learning" mode, involving activities such as drawing analogies, asking questions, comparing, selfexplaining, reflecting and monitoring. these activities produce external outputs, such as hypotheses, predictions, and justifications, which introduce new ideas beyond the presented information leading to the creation of new knowledge (chi & wylie, 2014). similarly, teachers’ observable behaviors, like scaffolding prompts, can further promote children's generativity. within this constructive mode, the focus is on the individual but can be further enhanced by an “interactive” mode of engagement between partners, which builds upon and encompasses the constructive mode. chi (2021) emphasizes the concept of co-inference, where partners engage in mutual knowledge construction through dynamic exchanges, enhancing understanding by integrating each other’s contributions (chi & wylie, 2014). many researchers assert that understanding is a social process, that also involves individual mental effort. classroom pair activities should be analyzed for both shared understanding and the individual processing of differing perspectives (inagaki et al., 1998; miyake, 1986). hatano and inagaki (2013) suggest examining collective comprehension activity alongside individual outcomes, addressing both intermental and intramental processes. the papers in this issue explore different modes of active learning and engagement. symonds and colleagues observed students’ behavioral engagement, providing evidence of active engagement but not of constructive or interactive engagement. although the presence or absence of cooperative behavior was recorded, further information about the type of cooperation between students or between students and teachers is not provided. moreover, while a process of self-regulation (important process in constructive teaching) underlies behavioral engagement, it is not explicitly discussed in their paper. in tang and colleagues’ study, students are actively engaged, but there is no explicit evidence of constructive or interactive engagement, as the focus was on situational feelings and experiences. renninger and colleagues implicitly referred to a constructive-interactive mode of engagement highlighting elements such as the use of prior math knowledge, exploration, planning, awareness of multiple perspectives and self-regulation. group cooperation in their study involves sharing new ideas, taking turns in participation, and engaging in discussions. collaboration, in turn, involves shared understanding, negotiations, incorporating diverse viewpoints, and extending thinking. haataja and colleagues’ study also includes aspects of the constructive-interactive mode, such as the co-construction of knowledge, co-regulation, and socially shared regulation of learning, where group members kyriakopoulou special issue: perspectives on momentary engagement and learning situated in classroom contexts 127 | f l r negotiate, reciprocally expressing perspectives and adapt to learning challenges. however, their study lacks a qualitative description of the various verbal exchanges during task execution. additional information of individual engagement experiences within the group may be needed. the work of baines and colleagues introduces another parameter: the role of peer relationships, particularly as expressed within the classroom context. their study showed that measures of classroom peer relations were strongly associated with both observed behavioral engagement and classroom cognitive engagement. indeed, being popular as a work partner or recognized as highly effective at group work, positively predicted both science achievement and increased progress over the year. this finding highlights the importance of examining peer relationships within the context of the co-constructive interactions discussed before. this underscores the importance of simultaneously exploring various factors to analyse the working partners' profile, detect potential developmental differences, and observe how the depth of engagement evolves over time based on the individuals’ profiles. analysis of classroom observations should also include evaluating the types of activities designed by teachers to promote learning and engagement. research reveals that teachers often adopt a transmissive approach to learning, frequently lacking explicit teaching theories (lawson et al., 2023; torsney & symonds, 2019; vosniadou et. al, 2023). vosniadou and colleagues (2023) developed a coding guide to analyse lesson tasks based on the icap (interactive-constructive-active-passive) model of engagement (chi, 2009; chi et al., 2018; chi & wylie, 2014). their findings indicate that most tasks promoted only active engagement, without fostering deeper constructive or interactive learning. most whole-class discussions failed to engage students constructively, as they lacked elements such as critical reflection, comparisons between new and prior concepts, or the transfer of new knowledge across domains. symonds and colleagues (this issue) observed higher levels of behavioral disengagement in larger classes with 22 or more students, emphasizing the importance of smaller pupil-teacher ratios. smaller class sizes enable teachers to interact more effectively with students, which is an important prerequisite for cc learning. lower ability was identified as a risk factor for disengagement in larger classrooms, but not in smaller ones. additionally, smaller classrooms were shown to be particularly beneficial for children from privileged backgrounds, suggesting that class size has a greater impact than socioeconomic status alone. similar studies indicate that students in small schools or those with a communal organization, where teachers create socially supportive and intellectually challenging environments, show higher levels of engagement (fredricks et. al, 2004). inkinen and colleagues (2020) linked certain scientific practices in science teaching, such as developing models and constructing explanations, to optimal learning moments. these practices, central to cc, were linked to deeper situational engagement and were more likely to challenge students, thereby triggering their interest. johnson and sinatra (2013), drawing on the expectancy-value theory framework, demonstrated that perceived task utility enhances engagement by activating prior knowledge and focusing attention on relevant information, thereby supporting cc. these findings underscore a significant challenge in educational contexts, particularly for cc learning. it is important to examine children's me across various levelsindividual, group, and taskrelatedwhile also considering teachers’ beliefs system about learning and teaching. mcneill (2009) showed that teachers’ conceptualization of scientific argumentation significantly affects students’ support and instructional practices. simplified inquiry approaches often lead to weaker learning gains in students’ ability to develop scientific arguments and to explain phenomena using evidence and reasoning. furthermore, alongside evaluating the engagement levels in task design, it is essential to consider the types of self-regulated learning embedded in these tasks, i.e., how students can take control of and regulate their learning, processes that are stimulated by the different modes of engagement (lawson et al., 2023). kyriakopoulou special issue: perspectives on momentary engagement and learning situated in classroom contexts 128 | f l r 3. rethinking momentary engagement beyond traditional acquisition type of learning: the paradigm of conceptual change the papers in this special issue discuss current models that aim to explain me during the learning process, focusing primarily on traditional learning acquisition (e.g. xin tang et al., this issue). however, these models fall short in addressing the complexity of conceptual change learning. this raises an intriguing question: how can we extend these models beyond traditional learning acquisition to contexts involving cc? conceptual change, in domains like science and mathematics, requires deep engagement and intricate cognitive processes (sinatra & pintrich, 2003). understanding the emergence and revision of students’ conceptual frameworks also requires consideration of motivational processes (linnenbrink & pintrich, 2002). however, factors influencing engagement in cc – such as epistemic cognition, participation in scientific practices, misconceptions, emotional responses to specific topics, attitudes toward science, and considerations of gender and identitydiffer from those influencing simpler knowledge enrichment (sinatra, heddy & lombardi, 2015). cordova and colleagues (2014), stress the importance of understanding how these factors interact to influence knowledge reconstruction. while motivational beliefs such as self-efficacy beliefs and learning goalsare linked to cognitive engagement (pintrich & schragben, 1992), recent studies indicate that effective learning characteristics in varied learning contexts may not yield equivalent benefits in cc learning (cordova et al., 2014. pintrich et al., 1993. sinatra et al., 2015). for instance, high interest and high prior knowledge, which typically enhance traditional learning, may hinder cc learning (dole & sinatra, 1998). the challenge in cc lies in encouraging students to critically assess their initial naïve theories and acknowledge inconsistencies when exposed to scientific information. how can students be motivated and interested in a topic where they hold strong, preexisting beliefs? what stimuli can redirect attention to conflicting viewpoints? how do factors like prior knowledge and epistemic beliefs shape motivational and cognitive processes during task engagement? 3.1 learners’ characteristics (interest, skill, challenge) and prior knowledge assessing students’ commitment, strength and coherence of their prior knowledge can predict engagement levels when radical cc is required (dole & sinatra, 1998; jones et al., 2015). strong prior knowledge, coupled with a high commitment to entrenched beliefs, reduce the likelihood of cc, particularly when combined with high topic interest (dole & sinatra, 1998; murphy et al., 2005). several researchers highlighted the importance of examining how interest interacts with other variables, such as epistemic beliefs and misconceptions, rather than studying it in isolation (linnenbrink-garcia et al., 2012; murphy & alexander, 2004). cordova and colleagues (2014) found that students with high selfefficacy and interest but low prior scientific understanding were more likely to achieve cc than those with strong prior knowledge combined with high confidence and interest. similarly, mason and colleagues (2008) highlighted the influence of combining topic interest with epistemic beliefs about scientific knowledge (complex and evolving vs. simple and certain) and text type (refutational vs. traditional text) on cc outcomes. self-efficacy beliefs, reflecting a student’s confidence in their ability to perform a specific task (skill component), can influence cc in opposing ways. high self-efficacy in learning can support cc by fostering confidence in engaging with challenging tasks, such as argumentation and experimentation (pintrich et al., 1993; sinatra, 2005). however, if self-efficacy reinforces confidence in pre-existing conceptions, it may hinder cc by fostering resistance to new ideas. interventions like refutational texts can help engage students by challenging misconceptions and guiding them toward scientifically accepted concepts (cordova et al., 2014; mason et al., 2008). kyriakopoulou special issue: perspectives on momentary engagement and learning situated in classroom contexts 129 | f l r 3.2 epistemic emotions and epistemic beliefs recent research brings in the discussion the role of epistemic emotions that arise when students encounter unexpected or incongruent information and pertain to ongoing knowledge-generating activities (muis et al.,2015; pekrun et al., 2017). contradictory information can evoke surprise, leading to two distinct pathways. firstly, it may trigger curiosity and situational interest, motivating students to actively seek resolution for the inconsistency and engage in a challenging task. on the other hand, confusion may emerge, leading to unsuccessful attempts of resolution, causing students to abandon both the proposition and the task altogether. however, confusion can be beneficial for learning in some cases, particularly when students attempt to resolve it by evaluating its source, adapting their strategies and employing metacognitive monitoring to track their progress (d'mello et al., 2014). in this case, confusion is associated with enhanced, deep metacognitive strategies, more sophisticated self-regulated learning abilities and critical thinking (chevrier et al. 2019). however, this positive dimension is more commonly observed in adult learners (d’mello & graesser, 2012; muis et al., 2015). furthermore, students with constructivist epistemic beliefs, viewing knowledge as complex and evolving, tend to experience more positive epistemic emotions and engage in deep self-regulated learning (chevrier et al., 2019; muis et al., 2015). muis and colleagues (2015) suggest that epistemic beliefs act as antecedents to epistemic emotions through appraisals of epistemic congruity, novelty and complexity and of the attainment of epistemic aims. 3.3 detractors and accelerants the boundaries between detractors and accelerants in cc may not always be clear-cut. various factors influence the extent and the way students engage. interestingly, "negative" factors often perceived as detractors, can sometimes promote cc by encouraging deeper engagement and learning (broughton et al., 2013). for example, students with initial epistemic beliefs may feel surprised or confused hen confronted with information that contradicts their prior knowledge, potentially leading to disengagement. however, students with more constructivist epistemic beliefs may respond to such contradictions differently. negative emotions can foster deep engagement with message processing, while positive emotions may sometimes result in shallower engagement. instructional strategies, such as group discussions and debates, can help transform negative emotions into opportunities for cc (gregoire, 2003; nadelson et al., 2018). 4. can we argue about a multi-perspective approach to understanding momentary engagement? in previous sections, i used the cc paradigm to discuss me at the micro-level, focusing on students’ dynamic learning moments in real time. understanding how micro-level factors -such as individual differences, learner characteristics, cognitive and motivational factors, socio-cultural influences, and classroom settingsshape engagement, and examining their interrelations, can provide a more comprehensive perspective on me. rather than simply identifying correlations, researchers should prioritize exploring how these factors interact in various configurations that may affect me in diverse ways. for example, fredricks and colleagues (2004) propose a pattern-centered analysis that could uncover different configurations of engagement types. potential variations across different types of learning should also be considered. at the microlevel, various factors may play distinct roles in cc learning compared to traditional learning. to better understand the interrelation between engagement and the dynamic processes underlying knowledge revision, researchers should account for variations within specific domains, such as the physical and kyriakopoulou special issue: perspectives on momentary engagement and learning situated in classroom contexts 130 | f l r social sciences, as well as for topics with differing levels of contradiction. in cc research, methods should evolve to assess not only the cognitive components of engagement but also its affective and behavioral components. in turn, the cc paradigm could help inform new perspectives and methodologies for analyzing learning. at the meta-level, cc could contribute to reconceptualizing the concept of engagement itself. the papers in this issue emphasize the need for a fundamental shift in how we conceptualize me. previous research has often examined engagement components in isolation. is this approach sufficient, or should me reconsidered as a holistic, integrated construct? the five papers discussed advocate for a pluralistic approach, incorporating multiple theoretical frameworks to deepen our understanding of the complexities surrounding engagement. from a meta-level perspective, a holistic view of me conceptualizes it as a complex system of interdependent parts that co-act, rather than as a mere accumulation of isolated components. the dynamic systems approach supports this perspective, suggesting that new forms of me emerge from shifts in the relationships between components, resulting to the system’s qualitative development through self-organization (lee & karmiloff-smith, 2002). symonds and colleagues (symonds et al., 2024) further reconceptualize me as a complex, dynamic developmental system. this integrative view emphasizes the co-action of parts and their interactions with the whole, recognizing the dynamic nature of engagement and offering a more nuanced understanding of its structure and processes. expanding this approach to other types of learning, nadelson and colleagues (2018) introduced the dynamic model of cc (dmcc), which addresses all three aspects of engagement -cognitive, affective and behavioral while acknowledging the contextual and situational nature of cc. this holistic model describes cc as a dynamic, non-linear and non-recursive process, characterized by multidirectional interactions sensitive to changes in emotions, behaviors, motivation and contextual factors. it thereby offer pathways for both engagement and disengagement. although the dmcc warrants further empirical study, it presents a promising framework for examining engagement in the context of cc. while a holistic view offers valuable insights, it also poses challenges. how is me conceptualized across different theoretical frameworks, and what are its boundaries? what elements does me truly encompass? how can existing models be expanded to integrate multiple perspectives coherently? addressing these questions requires careful consideration of how different frameworks influence one another, as varying constructs and definitions may hinder conceptual clarity and consistency in measurement (appleton et al., 2008). future research should also explore the implications of these perspectives for both research and educational practice. translating these insights into practical applications can help educators support students more effectively. a holistic understanding of me can guide the development of targeted interventions that address the sequential coaction of motivation, emotion, cognitive engagement and physical actions within a task, rather than focusing exclusively on broad measures of classroom and school engagement. programs such as the professional student program for educational resilience (prosper) by torsney and symonds (2019) exemplify this approach by enhancing me through a focus on both personal and social resources, fostering resilience across physical, motivational, emotional, and cognitive dimensions. evaluating holistic programs like prosper in comparison to those targeting specific engagement components is essential to bridging the gap between research and practice, facilitating the application of holistic engagement models in educational settings. overall, a holistic, dynamic approach to me deepens our understanding of its complexity and paves the way for educational practices that are responsive to students’ diverse needs and learning contexts. kyriakopoulou special issue: perspectives on momentary engagement and learning situated in classroom contexts 131 | f l r keypoints expand the examination of momentary engagement (me) beyond traditional acquisition models to encompass conceptual change learning. address the methodological challenges of capturing a moment-by-moment analysis of me, considering its multidimensional nature. investigate the complex interplay of factors such as prior knowledge, epistemic beliefs, emotions, and learner characteristics in shaping me across different types of learning. differentiate between various modes of active learning and the depth of student engagement. advocate for a holistic approach to me, discussing its challenges and implications for educational practices and interventions. references appleton, j. j., christenson, s. l., & furlong, m. j. (2008). student engagement with school: critical conceptual and methodological issues of the construct. psychology in the schools, 45(5), 369-386. https://doi.org/10.1002/pits.20303 appleton, j. j., christenson, s. l., kim, d., & reschly, a. l. (2006). measuring cognitive and psychological engagement: validation of the student engagement instrument. journal of school psychology, 44(5), 427-445. https://doi.org/10.1016/j.jsp.2006.04.002 bolger, n., davis, a., & rafaeli, e. (2003). diary methods: capturing life as it is lived. annual review of psychology, 54(1), 579-616. https://doi.org/10.1146/annurev.psych.54.101601.145030 broughton, s. h., sinatra, g. m., & nussbaum, e. m. (2013). “pluto has been a planet my whole life!” emotions, attitudes, and conceptual change in elementary students’ learning about pluto’s reclassification. research in science education, 43, 529-550. https://doi.org/10.1007/s11165-011-9274x chevrier, m., muis, k. r., trevors, g. j., pekrun, r., & sinatra, g. m. (2019). exploring the antecedents and consequences of epistemic emotions. learning and instruction, 63, 101209. https://doi.org/10.1016/j.learninstruc.2019.05.006 chi, m. t. (2009). active‐constructive‐interactive: a conceptual framework for differentiating learning activities. topics in cognitive science, 1(1), 73-105. https://doi.org/10.1111/j.1756-8765.2008.01005.x chi, m. t. (2021). translating a theory of active learning: an attempt to close the research‐practice gap in education. topics in cognitive science, 13(3), 441-463. https://doi.org/10.1111/tops.12539 chi, m. t., & wylie, r. (2014). the icap framework: linking cognitive engagement to active learning outcomes. educational psychologist, 49(4), 219-243. https://doi.org/10.1080/00461520.2014.965823 chi, m. t., adams, j., bogusch, e. b., bruchok, c., kang, s., lancaster, m., levy, r. li, n., mceldoon, k. stump, g., wylie, r., xu, d., & yaghmourian, d. l. (2018). translating the icap theory of cognitive engagement into practice. cognitive science, 42(6), 1777-1832. https://doi.org/10.1111/cogs.12626 cordova, j. r., sinatra, g. m., jones, s. h., taasoobshirazi, g., & lombardi, d. (2014). confidence in prior knowledge, self-efficacy, interest and prior knowledge: influences on conceptual change. contemporary educational psychology, 39(2), 164-174. https://doi.org/10.1016/j.cedpsych.2014.03.006 d’mello, s., & graesser, a. (2012). dynamics of affective states during complex learning. learning and instruction, 22(2), 145-157. https://doi.org/10.1016/j.learninstruc.2011.10.001 d’mello, s., lehman, b., pekrun, r., & graesser, a. (2014). confusion can be beneficial for learning. learning and instruction, 29, 153-170. https://doi.org/10.1016/j.learninstruc.2012.05.003 dole, j. a., & sinatra, g. m. (1998). reconceptualizing change in the cognitive construction of knowledge. educational psychologist, 33(2-3), 109-128. https://doi.org/10.1080/00461520.1998.9653294 fredricks, j. a., & mccolskey, w. (2012). the measurement of student engagement: a comparative analysis of various methods and student self-report instruments. in a.l.reschly, & s.l.christenson (eds.), https://doi.org/10.1002/pits.20303 https://doi.org/10.1016/j.jsp.2006.04.002 https://doi.org/10.1146/annurev.psych.54.101601.145030 https://doi.org/10.1007/s11165-011-9274-x https://doi.org/10.1007/s11165-011-9274-x https://doi.org/10.1016/j.learninstruc.2019.05.006 https://doi.org/10.1111/j.1756-8765.2008.01005.x https://doi.org/10.1111/tops.12539 https://doi.org/10.1080/00461520.2014.965823 https://doi.org/10.1111/cogs.12626 https://doi.org/10.1016/j.cedpsych.2014.03.006 https://doi.org/10.1016/j.learninstruc.2011.10.001 https://doi.org/10.1016/j.learninstruc.2012.05.003 https://doi.org/10.1080/00461520.1998.9653294 special issue: perspectives on momentary engagement and learning situated in classroom contexts 132 | f l r handbook of research on student engagement, (pp. 597-616). springer. https://doi.org/10.1007/978-3031-07853-8_29 fredricks, j. a., blumenfeld, p. c., & paris, a. h. (2004). school engagement: potential of the concept, state of the evidence. review of educational research, 74(1), 59-109. https://doi.org/10.3102/00346543074001059 fulmer, s. m., & frijters, j. c. (2009). a review of self-report and alternative approaches in the measurement of student motivation. educational psychology review, 21(3), 219-246. https://doi.org/10.1007/s10648-009-9107-x gregoire, m. (2003). is it a challenge or a threat? a dual-process model of teachers' cognition and appraisal processes during conceptual change. educational psychology review 15, 147–179. https://doi.org/10.1023/a:1023477131081 hektner, j. m., schmidt, j. a., & csikszentmihalyi, m. (2007). experience sampling method: measuring the quality of everyday life. sage. hatano, g., & inagaki, k. (2013). sharing cognition through collective comprehension activity. in d. faulkner, k. littleton, & m. woodhead, learning relationships in the classroom (pp. 276-292). routledge hickendorff, m., edelsbrunner, p. a., mcmullen, j., schneider, m., & trezise, k. (2018). informative tools for characterizing individual differences in learning: latent class, latent profile, and latent transition analysis. learning and individual differences, 66, 4-15. https://doi.org/10.1016/j.lindif.2017.11.001 inagaki, k., hatano, g., & morita, e. (1998). construction of mathematical knowledge through whole-class discussion. learning and instruction, 8(6), 503-526. https://doi.org/10.1016/s0959-4752(98)00032-2 inkinen, j., klager, c., juuti, k., schneider, b., salmela‐aro, k., krajcik, j., & lavonen, j. (2020). high school students' situational engagement associated with scientific practices in designed science learning situations. science education, 104(4), 667-692. https://doi.org/10.1002/sce.21570 johnson, m. l., & sinatra, g. m. (2013). use of task-value instructional inductions for facilitating engagement and conceptual change. contemporary educational psychology, 38(1), 51-63. https://doi.org/10.1016/j.cedpsych.2012.09.003 jones, s. h., johnson, m. l., & campbell, b. d. (2015). hot factors for a cold topic: examining the role of task-value, attention allocation, and engagement on conceptual change. contemporary educational psychology, 42, 62-70. https://doi.org/10.1016/j.cedpsych.2015.04.004 karmiloff-smith, a. (2013). ‘microgenetics’: no single method can elucidate human learning. commentary on parnafes and disessa. human development, 56(1), 47-51. https://doi.org/10.1159/000345541 lawson, m. j., van deur, p., scott, w., stephenson, h., kang, s., wyra, m., darmawan, i., vosniadou, s., murdoch, c., white, e., & graham, l. (2023). the levels of cognitive engagement of lesson tasks designed by teacher education students and their use of knowledge of self-regulated learning in explanations for task design. teaching and teacher education, 125, 104043. https://doi.org/10.1016/j.tate.2023.10404 lee, k., & karmiloff-smith, a. (2002). macro-and microdevelopmental research: assumptions, research strategies, constraints, and utilities. in n. granott & j. parziale (eds.) microdevelopment: transition processes in development and learning, (pp. 243-265), cambridge university press linnenbrink, e. a., & pintrich, p. r. (2002). motivation as an enabler for academic success. school psychology review, 31(3), 313-327. https://doi.org/10.1080/02796015.2002.12086158 linnenbrink-garcia, l., pugh, k. j., koskey, k. l., & stewart, v. c. (2012). developing conceptual understanding of natural selection: the role of interest, efficacy, and basic prior knowledge. the journal of experimental education, 80(1), 45-68. https://doi.org/10.1080/00220973.2011.55949 luebeck, j. l., & bice, l. r. (2005). online discussion as a mechanism of conceptual change among mathematics and science teachers. international journal of e-learning & distance education/revue internationale du e-learning et la formation à distance, 20(2), 21-39. retrieved from https://www.ijede.ca/index.php/jde/article/view/81 mason, l., gava, m., & boldrin, a. (2008). on warm conceptual change: the interplay of text, epistemological beliefs, and topic interest. journal of educational psychology, 100(2), 291. https://psycnet.apa.org/doi/10.1037/0022-0663.100.2.291 https://doi.org/10.1007/978-3-031-07853-8_29 https://doi.org/10.1007/978-3-031-07853-8_29 https://doi.org/10.3102/00346543074001059 https://doi.org/10.1007/s10648-009-9107-x https://doi.org/10.1023/a:1023477131081 https://doi.org/10.1016/j.lindif.2017.11.001 https://doi.org/10.1016/s0959-4752(98)00032-2 https://doi.org/10.1002/sce.21570 https://doi.org/10.1016/j.cedpsych.2012.09.003 https://doi.org/10.1016/j.cedpsych.2015.04.004 https://doi.org/10.1159/000345541 https://doi.org/10.1016/j.tate.2023.10404 https://doi.org/10.1080/02796015.2002.12086158 https://doi.org/10.1080/00220973.2011.55949 https://www.ijede.ca/index.php/jde/article/view/81 https://psycnet.apa.org/doi/10.1037/0022-0663.100.2.291 special issue: perspectives on momentary engagement and learning situated in classroom contexts 133 | f l r mcneill, k. l. (2009). teachers' use of curriculum to support students in writing scientific arguments to explain phenomena. science education, 93(2), 233-268. https://doi.org/10.1002/sce.20294 mikkilä-erdmann, m., penttinen, m., anto, e., & olkinuora, e. (2008). constructing mental models during learning from science text: eye tracking methodology meets conceptual change. in d. ifenthaler, p. pirnay-dummer & j.m. spector (eds.). understanding models for learning and instruction: essays in honor of norbert m. seel, (pp.63-79), springer. https://doi.org/10.1007/978-0-387-76898-4_4 miyake, n. (1986). constructive interaction and the iterative process of understanding. cognitive science, 10(2), 151-177. https://doi.org/10.1016/s0364-0213(86)80002-7 muis, k. r., pekrun, r., sinatra, g. m., azevedo, r., trevors, g., meier, e., & heddy, b. c. (2015). the curious case of climate change: testing a theoretical model of epistemic beliefs, epistemic emotions, and complex learning. learning and instruction, 39, 168-183. https://doi.org/10.1016/j.learninstruc.2015.06.003 murphy, p. k., & alexander, p. a. (2004). persuasion as a dynamic, multidimensional process: an investigation of individual and intraindividual differences. american educational research journal, 41(2), 337-363. https://doi.org/10.3102/00028312041002337 murphy, p. k., holleran, t. a., long, j. f., & zeruth, j. a. (2005). examining the complex roles of motivation and text medium in the persuasion process. contemporary educational psychology, 30(4), 418-438. https://doi.org/10.1016/j.cedpsych.2005.05.001 nadelson, l. s., heddy, b. c., jones, s., taasoobshirazi, g., & johnson, m. (2018). conceptual change in science teaching and learning: introducing the dynamic model of conceptual change. international journal of educational psychology, 7(2), 151-195. https://doi.org/10.17583/ijep.2018.3349 parnafes, o., & disessa, a. a. (2013). microgenetic learning analysis: a methodology for studying knowledge in transition. human development, 56(1), 5-37. https://doi.org/10.1159/000342945 pekrun, r., vogl, e., muis, k. r., & sinatra, g. m. (2017). measuring emotions during epistemic activities: the epistemically-related emotion scales. cognition and emotion, 31(6), 1268-1276. https://doi.org/10.1080/02699931.2016.1204989 pintrich, p.r., & schragben, b. (1992). students' motivational beliefs and their cognitive engagement in classroom tasks. in d. schunk & j. meece (eds.), student perceptions in the classroom: causes and consequences (pp.149-184), routledge pintrich, p. r., marx, r. w., & boyle, r. a. (1993). beyond cold conceptual change: the role of motivational beliefs and classroom contextual factors in the process of conceptual change. review of educational research, 63(2), 167-199. https://doi.org/10.3102/00346543063002167 posner, g. j., strike, k. a., hewson, p. w., & gertzog, w. a. (1982). accommodation of a scientific conception: toward a theory of conceptual change. science education, 66(2), 211-227. renninger, k. a., & bachrach, j. e. (2015). studying triggers for interest and engagement using observational methods. educational psychologist, 50(1), 58-69. https://doi.org/10.1080/00461520.2014.999920 schneider, m., & hardy, i. (2013). profiles of inconsistent knowledge in children’s pathways of conceptual change. developmental psychology, 49(9), 1639-1649 https://psycnet.apa.org/doi/10.1037/a0030976 siegler, r. s. (2006). microgenetic analyses of learning. in d. kuhn, & r. s. siegler (eds.), handbook of child psychology, 2, 464-510. sinatra, g. m. (2005). the" warming trend" in conceptual change research: the legacy of paul r. pintrich. educational psychologist, 40(2), 107-115. https://doi.org/10.1207/s15326985ep4002_5 sinatra, g. m., & pintrich, p. r. (2003). the role of intentions in conceptual change learning. in g., sinatra & p., r., pintrich (eds.). intentional conceptual change, 10-26. routledge. https://doi.org/10.4324/9781410606716 sinatra, g. m., heddy, b. c., & lombardi, d. (2015). the challenges of defining and measuring student engagement in science. educational psychologist, 50(1), 1-13. https://doi.org/10.1080/00461520.2014.1002924 skinner, e. a., kindermann, t. a., & furrer, c. j. (2009). a motivational perspective on engagement and disaffection: conceptualization and assessment of children's behavioral and emotional participation in academic activities in the classroom. educational and psychological measurement, 69(3), 493-525. https://doi.org/10.1177/0013164408323233 https://doi.org/10.1002/sce.20294 https://doi.org/10.1007/978-0-387-76898-4_4 https://doi.org/10.1016/s0364-0213(86)80002-7 https://doi.org/10.1016/j.learninstruc.2015.06.003 https://doi.org/10.3102/00028312041002337 https://doi.org/10.1016/j.cedpsych.2005.05.001 https://doi.org/10.17583/ijep.2018.3349 https://doi.org/10.1159/000342945 https://doi.org/10.1080/02699931.2016.1204989 https://doi.org/10.3102/00346543063002167 https://doi.org/10.1080/00461520.2014.999920 https://psycnet.apa.org/doi/10.1037/a0030976 https://doi.org/10.1207/s15326985ep4002_5 https://doi.org/10.4324/9781410606716 https://doi.org/10.1080/00461520.2014.1002924 https://doi.org/10.1177/0013164408323233 special issue: perspectives on momentary engagement and learning situated in classroom contexts 134 | f l r stathopoulou, c., & vosniadou, s. (2007). exploring the relationship between physics-related epistemological beliefs and physics understanding. contemporary educational psychology, 32(3), 255281. https://doi.org/10.1016/j.cedpsych.2005.12.002 symonds, j. e., kaplan, a., upadyaya, k., aro, k. s., torsney, b. m., skinner, e., & eccles, j. s. (2024). momentary student engagement as a dynamic developmental system. journal of theoretical and philosophical psychology. advance online publication. https://doi.org/10.1037/teo0000288 torsney, b. m., & symonds, j. e. (2019). the professional student program for educational resilience: enhancing momentary engagement in classwork. the journal of educational research, 112(6), 676692. https://doi.org/10.1080/00220671.2019.1687414 vosniadou, s. (1994). capturing and modeling the process of conceptual change. learning and instruction, 4(1), 45-69. https://doi.org/10.1016/0959-4752(94)90018-3 vosniadou, s. (2003). exploring the relationships between conceptual change and intentional learning. in g.m. sinatra & p.r. pintrich (eds) intentional conceptual change (pp. 373-402). routledge. https://doi.org/10.4324/9781410606716 vosniadou, s, (2013). model based reasoning and the learning of counter-intuitive science concepts, infancia y aprendizaje, 36(1), 5-33, https://doi.org/10.1174/021037013804826519 vosniadou, s. (2017). initial and scientific understandings and the problem of conceptual change. in t. g. amin, & o. levrini (eds.) converging perspectives on conceptual change (pp.17-25). routledge. vosniadou, s., & brewer, w. f. (1992). mental models of the earth: a study of conceptual change in childhood. cognitive psychology, 24(4), 535-585. https://doi.org/10.1016/0010-0285(92)90018-w vosniadou, s., & mason, l. (2012). conceptual change induced by instruction: a complex interplay of multiple factors. in k. r. harris, s. graham, t. urdan, s. graham, j. m. royer, & m. zeidner (eds.), apa educational psychology handbook, vol. 2. individual differences and cultural and contextual factors (pp. 221–246). american psychological association. https://doi.org/10.1037/13274-009 vosniadou, s., skopeliti, i., & ikospentaki, k. (2004). modes of knowing and ways of reasoning in elementary astronomy. cognitive development, 19(2), 203-222. https://doi.org/10.1016/j.cogdev.2003.12.002 vosniadou, s., vamvakousi, x., & skopeliti, e., (2008). the framework theory approach to the problem of conceptual change. in s. vosniadou (ed.), international handbook of research on conceptual change (pp.3-34). routledge. vosniadou, s., lawson, m. j., bodner, e., stephenson, h., jeffries, d., & darmawan, i. g. n. (2023). using an extended icap-based coding guide as a framework for the analysis of classroom observations. teaching and teacher education, 128, 104133. https://doi.org/10.1016/j.tate.2023.104133 https://doi.org/10.1016/j.cedpsych.2005.12.002 https://doi.org/10.1037/teo0000288 https://doi.org/10.1080/00220671.2019.1687414 https://doi.org/10.1016/0959-4752(94)90018-3 https://doi.org/10.4324/9781410606716 https://doi.org/10.1174/021037013804826519 https://doi.org/10.1016/0010-0285(92)90018-w https://psycnet.apa.org/doi/10.1037/13274-009 https://doi.org/10.1016/j.cogdev.2003.12.002 https://doi.org/10.1016/j.tate.2023.104133 codepen mouw publication frontline learning research vol.8 no. 6 (2020) 88 113 issn 2295-3159 the differential effect of perspective-taking ability on profiles of cooperative behaviours and learning outcomes jolien m. mouw1, nadira saab2, hannie gijlers3, marian hickendorff4, yolinde van paridon4 & paul van den broek 4 1department of educational sciences, university of groningen, the netherlands 2iclon, graduate school of teaching, leiden university, the netherlands 3department of instructional technology, university of twente, the netherlands 4institute of education and child studies, leiden university, the netherlands article received 30 march 2020/ revised 13 october/ accepted 15 october / available online 12 november abstract the present study aims to provide a systematic understanding of how perspective-taking ability contributes to primary-school students’ cooperative behaviours and learning outcomes. the present study is frontline as we combined person-oriented (e.g., describing patterns of behaviours based on individual characteristics), process-oriented (e.g., examining factors affecting the quality of cooperative behaviours), and effect-oriented (e.g., examining the effect of cooperative learning on individual learning outcomes) analytical approaches within one research framework. in addition, we adhered to the multi-dimensional nature of perspective-taking ability and differentiated between social and cognitive perspective-taking ability while taking into account the contribution of perspective-taking ability at both the individual level and group level (i.e., heterogeneous and homogeneous perspective-taking ability groups) to cooperative behaviour profiles and learning outcomes of primary-school children. based on transcribed episodes of interaction of 115 fifth-grade students, four different profiles of cooperative behaviours were discerned: captains, hard workers, switchers, and passive participants. we found that these profiles are related to perspective taking conceptualized at the group level, but not to individual-level perspective-taking ability. profile membership, cognitive perspective-taking ability, and group-level perspective-taking ability could not predict students’ learning outcomes. social perspective-taking ability and reading comprehension did positively predict learning outcomes. our findings add to existing knowledge as they suggest that the influence of perspective-taking ability on cooperative behaviours and learning outcomes is susceptible to the conceptualization (i.e., cognitive vs. social) and measurement level (i.e., individual vs. group level) of perspective-taking ability. keywords: cooperative learning; perspective-taking ability; group composition; profiles of cooperative behaviours; learning outcomes. info corresponding author email: j.m.mouw@rug.nldoi: https://doi.org/10.14786/flr.v8i6.633 1. introduction many primary school teachers implement cooperative learning activities on a daily base because this educational method can enhance students’ academic performances, motivation, and social skills (e.g., gillies, 2014; johnson & johnson, 2009). the ultimate goal of a cooperative learning activity is to engage students in promotive interaction in which learning can become meaningful. during promotive interaction, group members ideally participate actively, share information, help each other, and reach a common ground. these cooperative processes are often equated with effective cooperative behaviours as they support a group to attain its goals and enable individual and group learning processes to occur. however, not everyone finds it easy to cooperate or knows how to cooperate effectively as the coordination of activities, motivations, and goals is often a difficult task, especially when group members hold diverging opinions or reason from different points of view (webb, 2013). a potential black-box mechanism underlying promotive interaction in which group members engage in effective cooperative behaviours could be perspective taking (johnson, 1975a, 1975b; tjosvold, johnson, & johnson, 1984). for example, in research on adults, it is assumed that perspective taking supports the formation of social bonds and facilitates cognitive and communicative processes and behaviours such as problem solving and conflict resolution underlying promotive interaction (galinsky, ku, & wang, 2005; johnson, 1975a, 1975b). although some research has been conducted on adults’ perspective-taking abilities in the context of cooperative learning, a systematic understanding of how perspective-taking ability contributes to several aspects of primary-school students’ cooperative learning processes is lacking. for example, most studies focus on the relation between perspective taking and broad categories of communicative processes relevant for cooperation such as prosocial behaviours, conflict resolution, and problem solving (e.g., cigala, mori, & fangareggi, 2015; falk & johnson, 1977; tjosvold et al., 1984; trötschel, hüffmeier, loschelder, schwartz, & gollwitzer, 2011). however, no single study has evaluated the relation between perspective-taking ability and specific effective cooperative behaviours such as basic communicative functioning, helping behaviour, and grounding processes. more specifically, it is not clear if (and how) individual differences in perspective-taking ability yield differences in the whole range of cooperative behaviours in which primary-school students engage during face-to-face group work activities, and how this relates to subsequent learning outcomes. in addition, cooperative learning involves individual students working together in groups, and both individual and group-level characteristics (i.e., how students are grouped) shape interaction during group work and affect the quality of learning (e.g., arjava, salovaara, häkkinen, & järvelä, 2007; janssen, cress, erkens, & kirschner, 2013). however, the effects of individual-level and group-level perspective-taking ability on cooperative behaviours and learning outcomes have not yet been evaluated simultaneously. therefore, the present study strives to add to existing knowledge by differentiating between types of perspective-taking abilities and taking into account the contribution of both individual-level and group-level perspective-taking ability (i.e., heterogeneous and homogeneous perspective-taking ability groups) to cooperative behaviours and learning outcomes of primary-school children. below we present our theoretical framework in which we first elaborate on cooperative behaviours that are essential for the establishment of promotive interaction. we then continue by introducing perspective-taking ability and describe how types of individual-level and group-level perspective-taking ability could facilitate promotive interaction during group work and stimulate subsequent learning. we then address some conceptual and methodological challenges that need to be taken into account when evaluating the role of perspective taking during cooperative learning. 1.1 establishing promotive interaction the basic premise of effective cooperative learning is that learning becomes meaningful during promotive interaction between two or more communicative partners who attempt to learn something together (e.g., johnson & johnson, 2009). in promotive interaction, group members create the ideal circumstances for themselves and their peers to learn by participating actively, helping others, supporting each other’s thinking processes, sharing information and resources, and negotiating possible solutions. such a rich learning environment allows for collaborative knowledge construction and can be beneficial for understanding and learning processes (e.g., johnson & johnson, 2009; webb, 2013), provided that students engage in effective cooperative behaviours known to bring about higher individual learning gains (e.g., erkens, jaspers, prangsma, & kanselaar, 2005; gillies, 2014; webb, 2013). below we will discuss various types of effective cooperative behaviours that can be subsumed under three categories: basic communicative functioning, helping behaviour, and grounding processes. the first category of effective cooperative behaviours concerns aspects of basic communicative functioning such as giving compliments, appropriate turn-taking, listening actively (for example, acknowledging a message is received by responding with “hmm hmm”), and encouraging group members to continue their reasoning. if groups of students adhere to these basic rules for communication, a positive climate is created in which all group members feel empowered to contribute (wegerif & mercer, 1996). the second category of effective cooperative processes entails helping behaviours. in specific, it is imperative that students engage in high-quality helping behaviours during peer interaction and ask precise questions and help other group members by giving elaborate and detailed explanations. these processes stimulate cognitive restructuring and enable group members to identify and correct possible flaws or misconceptions in their reasoning (e.g., king, 2002; webb, 2013). most studies indeed have found that high-quality helping behaviour can be conducive to individual learning, whereas low-quality helping behaviours such as asking non-specific questions and giving and receiving simple answers are less beneficial (e.g., webb, farivar, & mastergeorge, 2002; for an exception see mouw, saab, janssen, & vedder, 2019). the third category of effective cooperative processes pertains to grounding behaviours, as all task-related and social activities taking place during group work need to be planned, monitored, and coordinated. only then are groups able to reach a shared frame of reference and co-construct a common language that enables them to communicate, discuss, and appraise individual perspectives (e.g., clark & brennan, 1991; di eugenio, jordan, tomason, & moore, 2000; erkens, 2004; erkens et al., 2005; janssen, erkens, kirschner, & kanselaar, 2012). erkens, jaspers, prangsma, and kanselaar (2005) found that these planning, monitoring, and coordination processes (such as focusing, checking, and argumentation) can be conducive to learning as it helps groups to obtain a common ground and facilitates in-depth processes of collaborative knowledge construction. 1.2 perspective taking: a prerequisite for establishing promotive interaction we posit that fruitful group work also requires the awareness that others may interpret the world or a task at hand from a different point of view (e.g., johnson, 1975b). for example, if cooperating group members are unable to recognize their differences in knowledge, motivations, or expectations, they are at risk of miscommunication, impasses, and even conflict (e.g., pronin, puccio, & ross, 2002; trötschel et al., 2011); all processes that make it difficult to reach a common ground. perspective-taking ability enables co-operators to explore differences in each other’s points of view and helps to understand that resulting disagreements do not necessarily reflect hostile communication, and thus sets the ground for promotive interaction (beersma, bechtoldt, & schouten, 2018; järvelä & häkkinen, 2002). in this study, we conceptualize perspective taking as the ability to (mentally) place yourself in someone else’s shoes, to assess social situations adequately, to understand, infer, and appropriately respond to expectations, thoughts, feelings and emotions of another person, and the motivation to do so (e.g., cigala et al., 2015; davis, 1983). in the context of cooperative learning, johnson (1975b) found that 9-to 10-year-old students’ affective perspective-taking abilities are correlated with their predisposition to cooperate. in addition, some studies have demonstrated that perspective-taking processes can facilitate more positive and effective communication processes. for example, training pre-schoolers’ perspective-taking abilities resulted in the tendency to more often behave in a prosocial manner during interactions with their peers (cigala et al., 2015). the 3to 5-year-old children who received this training more often engaged in helping and sharing behaviours as compared to their peers who did not receive training. in adult samples, it is shown that perspective taking can have a positive impact on group processes. for example, adults with high perspective-taking skills are more capable to match the content of their messages to the recipient in such a way that others better understand what is being said (e.g., johnson, 2015). moreover, when working in adult pairs, more information is brought into a conversation, understood, and remembered when one of the co-operators takes the perspective of the other as a mechanism to overcome conflict (johnson, 1971). in addition, perspective taking enables consensus making, the appropriate use of social and interpersonal skills, supportive and respectful behaviours, trust building, and leadership, and thereby facilitates the establishment of a positive climate in which group members possibly feel safer to contribute (e.g., järvelä & häkkinen, 2002; trötschel et al., 2011). last, perspective taking plays a role in the perceived self-other overlap and thus determines the degree to which group members experience similarities and feel connected to each other (e.g., davis, conklin, smith, & luce, 1996; galinsky et al., 2005). this way, perspective taking can strengthen social bonds, facilitates social coordination, promotes group cohesion, and can tie individuals together in a group (galinsky et al., 2005; galinsky & moskowitz, 2000). 1.3 considerations for evaluating the role of perspective taking during cooperative learning to summarize, it seems that previous studies (primarily focussing on adult samples) have evidenced a positive correlation between perspective taking and aspects of cooperative learning. however, more recent research advocates the notion that perspective taking should be considered and operationalized as a multi-dimensional construct as it appraises a visual, cognitive, and social (e.g., affective) component (cigala et al., 2015). visual perspective taking helps to determine the spatial position of objects and understanding and supports predicting others’ visual and physical experiences (michelon & zachs, 2006). cognitive perspective taking comprises the ability to infer, understand, and reason about another person’s motivations, intentions, and thoughts. social perspective taking strongly relates to empathy and concerns understanding others’ emotional states or experiencing these emotions yourself (cigala et al., 2015). in the context of cooperative learning—which is characterized by the verbal exchange of thoughts, opinions, and ideas within a social-emotional context—cognitive and social perspective-taking processes are of high instrumental purpose. a few studies that touched upon the assumed relation between perspective taking and cooperative processes did distinguish between these types of perspective taking. however, their conclusions pertain to either cognitive or social perspective taking (e.g., falk & johnson, 1977; tjosvold & johnson, 1977), and we are not aware of any study in which the relative contribution of cognitive and social perspective-taking ability to cooperative learning has been evaluated simultaneously. however, such a unidimensional conceptualization of perspective taking is problematic as this provides no insight into a possible differential effect of either type of perspective-taking ability on cooperative processes, even though it is not unlikely that cognitive and social perspective-taking ability relate to different types of cooperative behaviours. for example, based on the study of falk and johnson (1977), one would expect that students with higher cognitive perspective-taking abilities more often engage in task-oriented grounding behaviours such as monitoring, planning, and coordination. moreover, in line with cigala, mori, and fangareggi (2015), a positive relation between high-quality helping behaviour and social perspective taking is most likely. in addition, the few studies describing the relation between perspective taking and cooperative processes have hardly ever extended their findings by testing a possible direct relation between perspective-taking ability and learning outcomes (e.g., järvelä & häkkinen, 2002; johnson, 1975b; tjosvold & johnson, 1977; tjosvold et al., 1984; trötschel et al., 2011). another methodological consideration concerns the fact that cooperative learning involves individual students working together in groups, and both individual characteristics as well as group-level attributes (i.e., how students are grouped) are known to affect the success of a cooperative learning activity (e.g., arjava et al., 2007; cohen, 1994; janssen et al., 2013). therefore, it is important to determine if (and how) individual perspective-taking ability and perspective-taking ability conceptualized as a group characteristic (i.e., group formation based on individual students’ perspective-taking abilities) influence the effectiveness of cooperative learning in terms of cooperative behaviours and learning outcomes. however, until now, the majority of research on the effects of heterogeneity and homogeneity of groups is based on variance in cognitive ability levels (e.g., baer, 2003; cohen, 1994), whereas there is a scarcity of research focusing on the heterogeneity of social skills. one of the few exceptions is the work of falk and johnson (1977), where perspective-taking ability was taken into account at the group level (i.e., grouping based on perspective-taking ability) and instructional level (conditions varying in perspective-taking instructions). the solutions of groups of student nurses who were instructed to take their group members’ perspectives were of a higher quality and contained more creative suggestions as compared to the solutions of students who approached the task more individualistically. falk and johnson also found that students in the perspective-taking condition were more cooperative and engaged in fewer conflicts than the students in the egocentric condition, but across these conditions, heterogeneous and homogeneous groups did not differ in the effectiveness of their communication. however, the authors formed heterogeneous and homogeneous groups based on how similar student nurses ranked the relative importance of items for surviving on the moon. it can be argued whether this instrument gauged perspective-taking ability (i.e., the ability to view the world from another person’s perspective) or merely mapped students’ controversy regarding their views on a certain problem-solving activity. hence, it is not only unclear if students with higher and those with lower cognitive or social perspective-taking abilities engage in different grounding processes, helping behaviours, and basic communicative skills during group work: the same question arises with regard to groups varying in their group members’ perspective-taking abilities (i.e., heterogeneous and homogeneous groups). last, many studies on cooperative learning have exclusively adopted an effect-oriented research approach and delineated which distinct cooperative behaviours positively affect learning (see webb, 2013, for a review; asterhan & schwarz, 2016; webb et al., 2002). however, this does not fully capture what actually happens during peer interaction. co-operators do not use a single behaviour in isolation but instead engage in combinations of cooperative behaviours and activities in interaction with their peers. hence, a more fruitful approach would be to focus on the combinations or patterns of cooperative behaviours students use during a complete episode of group work and capture these patterns in profiles of cooperative behaviours (driskell, driskell, burke, & salas, 2017), for example, by using a latent profile analysis. with this technique, it is possible to examine how perspective-taking ability is related to differences in students’ engagement in patterns of essential cooperative behaviours throughout a cooperative learning activity (i.e., profiles of cooperative behaviours) and to subsequently evaluate how these profiles of cooperative behaviours relate to learning outcomes. 1.4 the present study based on our argumentation presented above, we emphasize the importance of understanding how grouping students could influence the quality of interaction and the extent to which learning is promoted. therefore, our goal is to evaluate the relation between types of individual-level and group-level perspective-taking ability (i.e., heterogeneous and homogeneous groups), patterns of cooperative behaviours, and learning outcomes. we will combine person-oriented (e.g., describing the relation between profiles of cooperative behaviours and individual characteristics), process-oriented (e.g., examining factors affecting the quality of cooperative processes), and effect-oriented (e.g., examining the effect of cooperative learning on individual learning outcomes) research approaches within one framework as we aim to answer the following questions: 1. are profiles of cooperative behaviours related to individual students’ cognitive and social perspective-taking abilities? 2. do students in heterogeneous and homogeneous perspective-taking groups engage in different profiles of cooperative behaviours? 3. how do individualand grouplevel perspective-taking ability and cooperative behaviour profiles contribute to individual students’ cooperative learning outcomes? for the first research question, we follow the work of järvelä and häkkinen (2002) and hypothesize that cooperative behaviour profiles of students with higher perspective-taking abilities can be characterized by engagement in more effective cooperative behaviours (i.e., higher occurrence of basic communicative functioning, helping behaviours, and grounding processes) as compared to students with lower perspective-taking abilities. more specifically, based on the study of falk and johnson (1977), we expect that students with higher cognitive perspective-taking abilities are more often allocated to profiles with higher frequencies of task-oriented behaviours such as the input of new information, monitoring, and planning activities (i.e., grounding activities). furthermore, we expect that students with higher social perspective-taking abilities more often engage in cooperative behaviour profiles with a dominance of high-quality helping behaviours (cigala et al., 2015). regarding the second research question, we expect that students in heterogeneous and homogeneous perspective-taking groups engage in different cooperative behaviours, with a higher prevalence of effective cooperative behaviour profiles in homogeneous groups (i.e., all group members with strong perspective-taking abilities) as compared to heterogeneous groups. regarding the third research question, we expect a positive relation between perspective-taking ability and subsequent learning gains, and that this effect will be particularly apparent for cognitive perspective taking (falk & johnson, 1977; johnson, 1971). furthermore, we expect a positive relationship between cooperative behaviour profiles with higher frequencies of effective behaviours (i.e., basic communicative functioning, grounding behaviour, and high-quality helping behaviour) and learning outcomes. 2. methodology 2.1 participants a total of 120 fifth-grade students (55 girls, mage = 10.58 years, age range: 9.5-11.92 years) from five classes from three primary schools in the netherlands participated in a three-session classroom-based study. we asked parents or caretakers to provide active written consent for their child(ren)’s participation. the ethics committee of the institute of education and child studies at leiden university approved the procedures of this study. 2.2 perspective taking during the first session, we assessed our participants’ cognitive and social perspective-taking abilities using (scales from) three different tests. this way, we were able to encapsulate several distinctive aspects of perspective taking in our assessment, as described by mouw, saab, pat-el, and van den broek (2019). 2.2.1 cognitive perspective taking in this study, we define cognitive perspective taking as the ability to mentally infer, understand, and reason about another person’s motivations, intentions, and thoughts. it pertains to one’s cognitive capacity to process and understand another’s point of view (e.g., cigala et al., 2015). we used two subscales of the interpersonal reactivity index (iri; davis, 1983), namely the fantasy and perspective taking scales, as measures of cognitive perspective taking. participants rated 14 items on a 5-point likert scale (1 = does not describe me well and 5 = describes me very well). an example item is: “i try to understand my friends better by imagining how something looks like from their perspective.” the internal consistency of both seven-item scales is sufficient (fantasy: α = .80; perspective taking: α = .72). cognitive perspective taking entails a second distinct—yet essential—facet, namely theory of mind (cigala et al., 2015; wellman, 2018). therefore, we decided to administer the read the mind in the eyes-test (rme; baron-cohen, wheelwright, hill, raste, & plumb, 2001) as a measure of cognitive perspective taking in addition to the two iri-scales. the child version of the rme measures how well children can attune to the mental state of others (vogindroukas, chelas, & petridis, 2014). for each of the 28 items, children are asked to attribute the relevant state of mind after having been presented with a photograph of a person’s eyes. the rme is a multiple-choice test: participants have to select one of four given mental states, and a point is given for each correct answer. internal consistency of this test is sufficient (α =.69; vogindroukas et al., 2014). 2.2.2 social perspective taking feelings of empathy and experiencing and understanding emotional aspects of social situations (such as emotional states) are indicators of social perspective-taking ability. therefore, we used the empathic concern scale of the iri (davis, 1983) to assess the first aspect of social perspective-taking ability. the empathic concern scale measures emotional reactions and other-oriented feelings of empathy and sympathy. participants rated this scale’s seven items on a 5-point likert scale (1 = does not describe me well and 5 = describes me very well). an example item is: “i often have concerned feelings for people less fortunate than me.” the dutch child version of the empathic concern scale was sufficiently reliable (α = .73). the revised emotion awareness questionnaire (eaq-30r; rieffe, oosterveld, miers, terwocht, & ly, 2008) was used as a second measure to assess children’s social perspective-taking abilities. this 30-item questionnaire gauges how students recognize, understand, think, and feel about others’ and their own emotional states; all of which are aspects of emotional functioning that are essential for emotional processing and reasoning, and thus, for social perspective taking (ruby & decety, 2004). the eaq-30r measures six aspects of emotional functioning (differentiating emotions, verbal sharing, not hiding, bodily awareness, other’s emotions, and analyses of emotions) on a likert scale that ranges from 1 (not true) to 3 (often true). it should be noted that the verbal sharing scale comprises three items and the differentiating emotions scale comprises seven items, whereas all other scales consist of five items. in the current sample, we found that the internal consistency of the six scales ranges between .64 and .78. 2.2.3 group formation because we were curious to find out if group composition based on perspective-taking ability influenced students’ cooperative behaviours and subsequent learning outcomes, we formed groups of three to four students based on their perspective-taking abilities. we formed eight homogeneous and 23 heterogeneous groups. in homogeneous groups, only students with higher perspective-taking abilities (i.e., those who scored above the classroom average on the rme, on a combined iri scale score, and on at least three out of six eaq-scales) were placed together. we did not form low-level homogeneous perspective-taking groups as this often concerned students of whom their teachers indicated that being placed together would lead to tension between students and undesired situations (for example, because of previous conflicts or bullying behaviours). students in heterogeneous groups varied in their perspective-taking abilities (i.e., mixed perspective-taking ability groups in which students with high, low, and average scores were placed together). 2.3 training and cooperative learning task at the start of the second session, an instruction video was shown in which ten-year-old peers modelled both desired and undesired cooperative behaviours. the researchers discussed these fragments with the students and, using a structured questioning technique, six main rules for effective cooperative learning were abstracted. these rules pertained to effective grounding techniques, high-quality helping behaviour, and basic cooperative behaviours such as turn-taking (janssen et al., 2012; johnson & johnson, 2009; webb et al., 2002). the rules, as outlined in table a1 (appendix), were written down on a poster that was displayed in the classroom where students worked on their cooperative assignments. immediately after receiving the cooperative-learning training, students practised these rules and desired behaviours in their groups as they worked on a jigsaw history task facilitating resource interdependence (e.g., ortiz, johnson, & johnson, 1996). the first part of the task was to collaboratively read a general introduction that set the scene for a historical event (i.e., the opening of the first railway line in the netherlands). group members then individually continued with the second part of the task, which was to read a text in which the historical happening was described from one of four protagonist-specific perspectives. because each group member read about the historical event from a different perspective, they each had different information at their disposal. during the third part of the task, students were challenged to discuss unique pieces of information presented in each of the four texts and had to reach a common ground to be able to write a brochure for a history museum. during the third session, groups of students worked a similar jigsaw-task (i.e., learning about a historical event from unique perspectives), but now learned about child labour in the 19th century. the texts (fictional narratives) were written in such a manner that students with a fourth-grade reading comprehension and/or technical reading ability level were able to read them (clib; evers, 2008). the suitability and readability of the texts and the corresponding knowledge test (discussed below) were discussed by a group of cooperative learning researchers and primary-school teachers (n = 4) and subsequently evaluated in a pilot study in which 133 fifth-grade students (72 girls, mage = 10.60 years, age range: 9.83-12.25 years) from seven classes from six dutch schools participated. based on this pilot study, minor adjustments—mostly pertaining to vocabulary—were made to the texts and the task description. 2.4 learning outcomes students filled out a 10-item test after completing the group assignment. this researcher-constructed test was used to assess individual students’ knowledge of the historical event they read about and have discussed during the group assignment. seven open questions and one multiple-choice question assessed factual knowledge or general information incorporated in all four texts. in addition, students had to answer two integration questions for which they had to combine pieces of unique information presented across the four texts. this means that they could only give a partial solution if they merely recalled information they had read themselves instead of using information from other texts as discussed during group work. a point was given for each correct answer, and a total of 11 points could be awarded as one question comprised an ‘a’ and ‘b’ part. the first and fifth author (who had no knowledge of the type of group a student was placed in) marked these tests individually and agreed on 94.78% of all points awarded. the points awarded to this knowledge test served as a proxy measure for student’s understanding of the learning materials, and from here on will be referred to as individual learning outcomes. no measure of group learning outcomes was included in our analyses because the quality of the group product (i.e., the brochure) did not always reflect the quality of the group’s cooperative processes. for example, some groups engaged in high-quality interaction in which they discussed a lot of information, evaluated their perspectives thoroughly, and continuously monitored their strategies. these important processes took up much time and effort; time that could no longer be spent on finishing a high-quality final product. 2.5 cooperative behaviours thirty-one groups were videotaped while working on their cooperative task during the third session. however, due to technical difficulties, only 30 episodes of interaction could be transcribed and used for further analyses. following the works of erkens et al. (2005), janssen, erkens, kirschner, and kanselaar (2012), webb, farivar, and mastergeorge (2002), and wegerif and mercer (1996), the coding scheme included seven types of effective cooperative behaviours to operationalize the three (previously discussed) categories of cooperative behaviours that are considered to be essential for establishing promotive interaction: basic communicative functioning (i.e., encouraging each other), grounding activities (i.e., checking, monitoring, planning, and input of new information), and helping behaviours (i.e., asking specific questions and giving explanations). we added task-related activities such as reading out loud as an eight category of effective behaviours and included three ineffective, but frequently occurring, low-quality helping behaviours (i.e., asking non-specific questions, giving answers, and off-topic behaviours) in our coding scheme. this way, it is possible to examine the whole range of behaviours students engage in during peer interaction. examples of each of the 11 categories of cooperative behaviours are presented in table 1. transcripts were coded in the multiple episode protocol analysis tool (mepa, version 4.10; erkens, 2005). in total, 13783 utterances (mlength transcript = 459.43 utterances, sd = 137.19) were coded independently by two coders (i.e., the first and fifth author). the interrater reliability, calculated on approximately 15% of the transcripts, was satisfactory (κ =.82). 2.6 reading comprehension individual levels of reading comprehension were included in our examination because students worked on a cooperative task highly narrative in nature. the dutch national institute for educational measurement (cito) developed a curriculum-independent test that is used nationally to assess primary-school students’ reading comprehension abilities (cito-test reading comprehension; egberink, janssen, & vermeulen, 2015). students’ reading comprehension ability ranges from a to e levels (a = 25% best scoring students; b = 25% on or just above average; c = 25% below average; and levels d and e together comprise the 25% lowest-scoring students). table 1 scheme for coding cooperative behaviours 2.7 procedure students’ perspective-taking abilities were measured during a first, 30-minute session. based on their perspective-taking skills, we placed students in heterogeneous or homogeneous groups. at the start of the second session, students received a 20-minute training on cooperative behaviours and rules for effective group work were abstracted. immediately after receiving this training, students familiarised themselves with these rules while working on a practice task. this cooperative jigsaw-task asked students to jointly read a general introduction in which a historical event (i.e., the opening of the first railway line) was introduced. students then continued reading a text individually. the group members cooperatively wrote an information brochure based on information offered in all texts. during the third session, groups of students worked a similar jigsaw-task (i.e., learning about a historical event from unique perspectives), but now learned about child labour in the 19th century. all peer interaction taking place during this third session was videotaped. groups of students on average worked on the cooperative task for 35 minutes. an individual knowledge test was administered directly after group work. all measurements and activities took place in students’ own classrooms. 2.8 statistical analyses five students were absent during the third meeting. as a result, our analyses are based on the behavioural data and learning outcomes of 115 students. below, we will first describe how the distinct profiles of cooperative behaviours were discerned. then, we will set forth which statistical approaches were used to answer our research questions. 2.8.1 students’ profiles of cooperative behaviours with regard to the behavioural data, we posit that group members do not use a single behaviour in isolation. instead, they engage in combinations of activities in interaction with their peers. therefore, we followed a person-oriented approach and focused on profiles of cooperative behaviours (i.e., combinations of cooperative behaviours) students used during a whole episode of group work to capture all behaviours occurring during peer interaction. these profiles are used in all subsequent analyses to answer our research questions. to identify distinct profiles of students’ cooperative behaviours we conducted a latent profile analysis (lpa). lpa is a statistical technique that belongs to the family of model-based clustering techniques. it aims to trace the heterogeneity in observed responses to a number of distinct classes or clusters, each with its characteristic response profile (e.g., oberski, 2016). lpa not only captures quantitative individual differences (where individuals are placed on a scale from low to high frequency of behaviours); it can also reveal qualitative individual differences (for a conceptual introduction to the application of this technique in learning research, see hickendorff, edelsbrunner, mcmullen, schneider, & trezise, 2018). the program latent gold 5.1 was used to conduct the lpa on the frequency of the 11 cooperative behaviours described in section 2.5 (vermunt & magidson, 2013). as a first step, the number of profiles needs to be determined. models with increasing numbers of profiles were estimated. based on a combination of statistical fit measures and conceptual appeal the number of profiles was selected, and the profiles were interpreted. 2.8.2 profiles of cooperative behaviours in relation to individualand group-level perspective-taking ability to investigate the relation between students’ profiles of cooperative behaviours and individual-level perspective-taking ability, we ran two univariate anovas with students’ profile membership (obtained by modal assignment; collins & lanza, 2010) as the independent variable and individual-level cognitive and social perspective-taking ability as dependent variables. to this end, we calculated a compound cognitive perspective-taking score by averaging students’ normalized scores on the rme and the iri fantasy and iri perspective-taking scales. likewise, we constructed a social perspective-taking scale score by normalizing and averaging students’ ratings on the six eaq scales and the iri empathy scale. the descriptive statistics of all variables (prior to normalization and scale construction) are presented in table 2. a chi-square test was then used to evaluate the relation between cooperative behaviours and perspective taking at the group level (i.e., heterogeneous and homogeneous group composition based on group members’ perspective-taking abilities). more specifically, we investigated if the prevalence of the discerned profiles of cooperative behaviours differs between students working in heterogeneous groups and those working in homogeneous groups. 2.8.3 predictors of individual learning outcomes for evaluating the contribution of individualand group-level perspective-taking ability, cooperative behaviour profiles, and reading comprehension to cooperative learning outcomes, a multilevel analysis was used to take into account the hierarchical data structure (i.e., students working in groups) when predicting students’ individual learning outcomes. the dependent variable was the learning outcome measure (i.e., individual score on the knowledge test), and the independent variables at the individual level were students’ profile membership (dummy-coded), cognitive perspective taking, social perspective taking, and reading comprehension (dummy-coded with the largest class serving as the baseline category as suggested by field, 2009). the correlations between all individual-level variables can be found in table 3. group-level perspective-taking ability (e.g., heterogeneous and homogeneous composition) was added as a group-level predictor. table 2 descriptive statistics note. iri = interpersonal reactivity index. rme = read the mind in the eyes-test. eaq = emotion awareness questionnaire. table 3 correlations between individual-level variables note: * p < .001. 3. results 3.1 profiles of cooperative behaviours based on the fit indices bic, aic3, and caic (hooper, coughlan, & mullen, 2008), we have scrutinized the solutions with two to four profiles (see table 4). of these, the four-class solution is chosen based on conceptual appeal as one of the profiles in the three-class solution could be further divided into two meaningful subgroups that differed in terms of the occurrence of off-topic behaviours. hence, the addition of a fourth class increased the conceptual interpretation of our data as this allowed us to discern profiles based on the prevalence of less effective behaviours. the class-specific profiles of mean frequencies in which each type of behaviour occurred (i.e., the estimated averages of absolute frequencies) are plotted in figure 1. table 4 model fit indices latent profile analyses note. the lowest fit-indices are set in boldface. bic = bayesian information criterion. aic3 = akaike information criterion 3. caic = consistent akaike information criterion. npar = number of parameters. around 36% of the students displayed a behaviour pattern that we labelled the captains profile. these captains contributed most by asking questions, giving answers, monitoring the process, and engaging in task-related behaviours and, thus, steered the cooperative learning process substantially. however, the captains also frequently engaged in behaviours less beneficial for learning: they were easily distracted, as illustrated by the high occurrence of off-topic behaviours. the behaviour patterns of 22% of students are clustered in a second profile; the hard workers. this profile highly resembles the captains profile, as the hard workers also took the lead in the cooperative process. an important difference between the hard workers and the captains is that the behaviours of students with a hard-worker profile substantially contributed to the group’s shared knowledge base: they most often brought in new information that was used as an input for the brochure and gave most explanations. in addition, hard workers least often engaged in off-topic behaviours and, instead, kept the focus on the task by engaging in planning and task-related activities. almost 30% of the students engaged in cooperative behaviour patterns clustered in the switchers profile. when retrospectively examining the transcripts of the students with a switcher profile, it is apparent that they constantly switched between onand off-topic behaviours. in comparison to students in other profiles, the switchers most often engaged in disruptive behaviours. students in this profile also gave a relatively substantial number of answers, monitored the groups’ processes, and engaged in both planning and task-related activities. however, the frequency with which switchers engaged in these behaviours is lower as compared to the captains and hard workers. the passive-participants profile captures the cooperative behaviours used by 12% of the students. these students did take part in the cooperative activity, but only if other group members asked them to do so. for example, they gave short answers and summarized what they had read (i.e., input of new information), but did so least often as compared to students with other profiles. passive participants did not take the initiative and they only asked a minimal number of (non-)specific questions. in comparison to students in other profiles, passive participants far less often engaged in grounding behaviours such as planning and monitoring. figure 1. estimated average absolute frequencies of types of cooperative behaviours for all profiles. 3.2 profiles of cooperative behaviours in relation to perspective-taking ability two univariate anovas were used to examine if students’ profiles of cooperative behaviours are related to individual-level perspective-taking ability. specifically, we evaluated if students in the four cooperative behaviour profiles differed in their cognitive and/or social perspective-taking abilities. table 5 shows that there was no significant relationship between cognitive perspective-taking ability and cooperative behaviour profiles, f(3, 111) = 1.94, p = .13, nor between social perspective-taking ability and cooperative behaviour profiles, f(3, 111) = 0.15, p = .93. thus, there is no evidence that students in the four profiles differ in their cognitive and/or social perspective-taking abilities. table 5 univariate anovas figure 2. profile prevalence (in %) as a function of group composition. for answering the second research question, whether students in heterogeneous and homogeneous groups (i.e., group-level perspective-taking ability) engaged in different profiles of cooperative behaviours, a chi-square test was used. as can be seen in figure 2, the cooperative behaviour profiles of students in heterogeneous groups indeed differed from those of students in homogeneous groups, χ2(3) = 9.06, p = .03. in specific, 48.4% of the cooperative behaviour patterns of students working in homogeneous groups were classified as a hard-worker profile, whereas this cooperative behaviour profile could only characterize 23.8% of the behavioural patterns of students in heterogeneous groups. in addition, 42.9% of the cooperative behaviour profiles of students working in heterogeneous groups were classified as a captain profile, whereas this classification only applied to 16.1% of the profiles of students working in homogeneous groups. students in homogeneous and students in heterogeneous groups do not differ with regard to the prevalence of switcher and passive-participant profiles. 3.3 predictors of individual learning outcomes we initially performed a multilevel analysis to examine the relation between types of individual-level perspective-taking abilities, group-level perspective taking, reading comprehension, profile membership, and individual learning outcomes. however, the intra-class correlation coefficient of 6% indicated that the variance in students’ learning outcomes residing at the group level was too low to warrant a multilevel analysis. in other words, we found no significant differences between the test scores of students working in heterogeneous groups and those in homogeneous groups and, thus, there is no evidence for a relation between group-level perspective-taking ability (i.e., group composition) and students’ individual learning outcomes. therefore, multiple regression analysis with stepwise deletion was used to evaluate the prediction of individual learning outcomes by cooperative behaviour profiles, cognitive and social perspective taking, and reading comprehension. all variables were entered simultaneously in the baseline regression model, f(16, 98) = 3.94, p < .001, r2 = 0.39. the resulting regression coefficients are presented in table 6. next, all non-significant predictors were eliminated using stepwise deletion to facilitate interpretation of effects. the final model was significant, f(4, 113) = 8.78, p < .001, and explained 23.7% of the variation in individual learning outcomes. as table 7 shows, social perspective-taking ability had a significant positive effect on students’ learning outcomes (b = 0.95, p = .018). this implies that, on average, higher social perspective-taking ability is reflected in higher test scores. finally, all levels of the test measuring reading comprehension significantly predicted individual learning outcomes. in comparison to the 25% of students with the highest level of reading comprehension (a-score), all other students scored significantly lower on the test (with b’s ranging between -1.59 and -2.28). we found no evidence for a relation between cooperative behaviour profiles and individual learning outcomes. table 6 regression coefficients of the baseline regression model predicting individual learning outcomes note. n = 115. ci = confidence interval. *the captains profile (dummy-coded) served as the baseline category. #cito reading comprehension level a served as the baseline category. table 7 regression coefficients of significant predictors of individual learning outcomes note. n = 115. ci = confidence interval. #cito reading comprehension level a served as the baseline category. 4. discussion the general aim of the present study was to provide a systematic understanding of how perspective-taking ability contributes to primary-school students’ cooperative behaviours and learning outcomes. the present study adds to existing knowledge as we differentiate between types of perspective-taking abilities while taking into account the contribution of both individual-level and group-level perspective-taking ability (i.e., heterogeneous and homogeneous perspective-taking ability groups) to cooperative behaviour profiles and individual learning outcomes of primary-school children. below, we will first briefly reflect on the nature and meaning of the discerned profiles of cooperative behaviours. then, we will answer each of our research questions and discuss theoretical and practical implications, limitations, and directions for future research. 4.1 profiles of cooperative behaviours because we wanted to learn if—and how—differences in perspective-taking ability are reflected in the range of cooperative behaviours in which primary-school students engage during a whole episode of group work, we first examined to what extent individual differences in primary-school students’ cooperative behaviours can be captured in a number of distinct profiles. the results of the lpa show that four profiles of cooperative behaviours can be discerned from the current sample’s behavioural data: captains, hard workers, switchers, and passive participants. the captain and hard-worker profiles approximate the effective leadership role described by driskell, driskell, burke, and salas (2017), even though the cooperative behaviour profile of the captains entailed a high occurrence of off-topic behaviours. the hard workers contribute to the quality of group work as they gave most explanations as compared to students whose cooperative behaviours could best be captured in one of the other profiles. this finding only partially confirms our hypothesis that a cooperative behavioural profile with a dominance of high-quality helping behaviours can be discerned. even though we found a difference between the hard workers and all other profiles of cooperative behaviours with regard to the absolute frequencies in which students gave explanations, even the hard workers gave relatively few explanations during group work as compared to all other cooperative behaviours in which they engaged. in addition, we found that the hard workers most often brought in new information serving the group’s construction of knowledge (erkens & janssen, 2008). with regard to the switchers, it is interesting to observe that almost a third of the students working in groups continuously (and seemingly effortlessly) switched between behaviours that can be placed on the opposite ends of a ‘continuum of effective cooperative behaviours’, and thus, are able to correct themselves (or are corrected by their group members) and regain focus as they easily switch back onto task-related behaviours and continue working on the task at hand. the cooperative behaviour profile of the passive participants seems to fit existing notions on social loafing and freeriding, which is a common problem in cooperative learning in which certain group members do not participate or contribute to group work (kreijns, kirschner, & jochems, 2003). however, our findings indicate that these ‘passive participants’ do not completely refrain from the task, but instead, engage minimally. the four profiles were discerned based on the frequency with which each of the cooperative behaviours co-occurred during group work. all students engaged in the whole range of cooperative behaviours, but they differed systematically in the frequency with which each of these behaviours occurred. for example, irrespective of profile membership, students engaged in monitoring activities, but students in the captain profile monitored the group’s process most often in comparison to students in other cooperative behaviour profiles. when we reflect on the nature of the cooperative behaviour profiles as discerned by the lpa, it seems that the profiles provide empirical evidence of the actual occurrence of some of the team-role clusters identified by driskell and colleagues (2017). moreover, the discerned profiles align with the functional team roles that are frequently used in cooperative classrooms. assigning functional roles (i.e., task interdependence) such as manager, planner, or questioner, is an often implemented and effective method to structure cooperative learning as it increases group efficiency and task-related interaction (e.g., strijbos, martens, jochems, & broers, 2007). interestingly, the results of the current study suggest that even if a teacher does not actively impose functional roles, these roles occur naturally during group work as students engaged in more or less role-specific behaviours, possibly as a product of the interplay between different team members. when interpreting all results, we stress that the cooperation with other group members, in the end, determines what is said and which cooperative behaviours are needed, and thus, could steer the distribution of team roles (driskell et al., 2017). future research could usefully explore if students engage in similar patterns of behaviours while working on tasks differing in complexity or when collaborating with different group members. 4.2 profiles of cooperative behaviours in relation to perspective-taking ability with regard to our first research question, ‘are profiles of cooperative behaviours related to students’ individual cognitive and social perspective-taking abilities?’, the results of the anovas provide no evidence for a relation between profile membership and individual-level cognitive or social perspective-taking abilities. the absence of such a relation contradicts the hypotheses we held prior to performing the lpa and challenges findings as reported in previous studies. for example, based on the work of falk and johnson (1977), we expected that students with higher cognitive perspective-taking abilities would take on a profile characterized by higher frequencies of task-oriented behaviours such as checking, planning, monitoring, and input of new information (i.e., grounding behaviours). and following cigala and colleagues (2015), we expected that the cooperative behaviours of students with higher social perspective-taking abilities could be captured in a profile in which high-quality helping behaviours prevail. however, the results of the lpa show that both students allocated to a captain profile and those allocated to a hard worker profile often engage in different types of effective grounding and helping behaviours, and as a result, we could not make a clear-cut distinction between cognitively-oriented and social cooperative behaviour profiles behaviours to evaluate our hypotheses. another explanation for the fact that that the discerned interactive roles students took on during group work could not uniquely be related to either type of individual perspective-taking ability could stem from the group composition used in the current study: students who scored high on cognitive perspective-taking measures were placed together with students with higher social perspective-taking abilities in homogeneous groups, and in heterogeneous groups students varying in perspective-taking ability worked together. we hypothesize that this has led to a situation in which the typical behaviours described in previous studies (e.g., stronger engagement with high-quality helping behaviour if you score high on social perspective-taking measures) possibly are counterbalanced by behaviours in which other group members engaged during group work. for example, it could be that in heterogeneous groups, students with higher social perspective-taking skills refrained from high-quality helping behaviours because none of the cooperating partners with lower perspective-taking abilities engaged in these behaviours. we suppose that more pronounced helping-behaviour and grounding profiles could be discerned if we distinguish between homogeneous cognitive perspective-taking ability groups and homogeneous social perspective-taking ability groups. with regard to our second research question, ‘do students in heterogeneous and homogeneous perspective-taking groups engage in different patterns of cooperative behaviours?’, the empirical evidence suggests that perspective taking conceptualized at the group level and profile membership are related. the cooperative behaviours of students in homogeneous groups (i.e., all group members with strong perspective-taking abilities) were most often classified as a hard-workers profile, whereas the prevalence of the captain profile is higher among students in heterogeneous groups. this is noteworthy because the general assumption is that forming heterogeneous groups augments learning gains as cooperating partners are then stimulated to bring in a variety of skills and knowledge facilitating group discussion (e.g., baer, 2003). however, this assumption stems from research focusing mostly on cognitive ability levels such as mathematical performance or language proficiency (e.g., baer, 2003; cohen, 1994; mouw, saab, janssen, & vedder, 2019). when group composition is based on group members’ social skills such as perspective-taking ability, different processes may come into play. however, a word of caution is advised as a chi-square distribution does not take into account the hierarchical nature of the data. instead, we have disaggregated group-level data to the individual level, leading to an exaggeration of the sample size for the grouping variable which increases the chance of a type 1 error (janssen et al., 2013). future research on the relation between group-level characteristics and profiles of cooperative behaviours should, therefore, include statistical approaches that can deal with nested categorical data. 4.3 predictors of individual learning outcomes the third research question was, ‘how do individualand group-level perspective-taking ability and cooperative behaviour profiles contribute to individual students’ cooperative learning outcomes?’ to control for relevant student characteristics, we included levels of reading comprehension in the regression analysis. lower levels of reading comprehension resulted in lower test scores, which can be explained by the textual nature of the task. with regard to perspective taking, we found that cognitive perspective taking could not predict individual learning outcomes. the absence of a relationship between cognitive perspective taking and individual learning outcomes is remarkable as it contradicts our hypothesis and previous lines of research (falk & johnson, 1977; johnson, 1971). the lack of predictive power of cognitive perspective taking in the context of cooperative learning could be the result of maturational differences between the perspective-taking abilities of the (young) adults participating in previous studies and the primary-school students participating in the current study (epley, morewedge, & keysar, 2004; pons, harris, & de rosnay, 2004). hawk and colleagues (2013) found that scores on the perspective taking and fantasy scales (i.e., scales from in the interpersonal reactivity index that are indicators of cognitive perspective-taking ability) are lower for early adolescents (13-year-olds) than for late adolescents (18-year-olds). hence, our findings could point to a developmental difference in children and adults’ engagement in effective cooperative behaviours as a function of (types of) perspective-taking ability, and thus, warrants further examination. interestingly, our results suggest that social perspective taking does play a role in the context of cooperative learning. more specifically, we found that students with higher social perspective-taking abilities score higher on the individual test tapping into the information that was discussed during group work as compared to peers with lower social perspective-taking abilities. it should be noted that we found no evidence that this result is a function of a (cor)relation between social perspective taking and reading comprehension ability. the results of the current study extend previous studies, such as those of johnson (1975a, 1975b), who reported on a relationship between fourth-grade students’ social perspective-taking abilities and their predisposition to cooperate. our results suggest that the influence of social perspective-taking ability reaches beyond a mere disposition towards group work and possibly even affects the quality of learning. being able to engage in other-oriented feelings of empathy and experiencing, understanding, and appropriately responding to emotional aspects evoked by the social context of cooperation in which a group functions as a collective and social entity (järvenoja & järvelä, 2016), possibly supported adequate group functioning. working together on an assignment not only implies cognitive engagement with (and coordination of) task-related aspects but also requires co-operators to get along with each other for the sake of task completion (e.g., barron, 2003; kreijns et al., 2003). however, being confronted with diverging points of view and strong opinions can lead to cognitive conflict, often accompanied by negative emotional associations or arousal (järvenoja & järvelä, 2016) that impede effective cooperative behaviours such as grounding and helping behaviours. the results of the current study may suggest that even in primary-school settings, social perspective-taking ability could function as a mechanism facilitating groups to overcome these differences on an emotional level. higher social perspective-taking ability enables group members to recognize, control, or regulate their own and others’ situational interpretation of the social context and ensuing emotional states. adopting such an affective—instead of a cognitive—point of view to cope with these emotional experiences might have helped groups to overcome differences, restore their working climate, motivational well-being, and regain task focus (e.g., järvenoja & järvelä, 2009, 2016). the last individual characteristic included in the regression analysis was cooperative behaviour profile membership. unexpectedly, the inclusion of profile membership did not reach statistical significance. this could stem from the fact that the profiles were based on a mixture of effective (i.e., basic communicative functioning, grounding, and high-quality helping behaviours) and ineffective, yet frequently occurring cooperative behaviours (i.e., off-topic behaviours and low-quality helping behaviours). in a similar vein, some of the behaviours were task-related, whereas others pertain to social aspects of cooperation. one direction for future research is to code episodes of interaction along a social dimension and a task-related dimension, for example, corresponding to the sociability and task-orientation dimensions of driskell and colleagues (2017). this way, it is possible to examine if perspective-taking ability can predict profile membership if the discerned cooperative behaviour profiles more specifically represent a social or a task-related dimension. 4.4 conclusion the added value of the present study lies in the fact that discerning profiles of cooperative behaviours more closely reflects what actually takes place during a cooperative learning task. in addition, the four discerned profiles shed light on which behaviours of cooperating primary-school students co-occur in which frequency. furthermore, we found that students allocated to these four profiles of cooperative behaviours do not differ in their cognitive or social perspective-taking abilities, suggesting there is no relation between individual-level perspective-taking ability and cooperative behaviours such as grounding, helping behaviours, and basic communicative functioning. moreover, we now understand that grouping students based on a combination of cognitive and social perspective-taking ability measures does not necessarily have a direct benefit for individual learning outcomes. however, group composition does seem to play a role in the combination of cooperative behaviours in which students engage during group work, and thus, possibly functions as a mechanism facilitating effective interaction processes that are essential for learning. last, our study is one of the first to show a differential effect of types of perspective taking on individual learning as social, but not cognitive, perspective-taking ability positively predicts learning in a cooperative context. all in all, our findings suggest that the relationship between perspective taking, cooperative behaviours, and individual students’ learning outcomes is not as straightforward as previously documented. on the contrary, it seems that this relation is highly susceptible to the level of measurement and the conceptualization of perspective-taking ability. hence, acknowledging perspective taking as a group-level attribute seems to be a fruitful endeavour when the goal is to understand how grouping students influences individual students’ engagement in cooperative behaviours, whereas adopting a multidimensional approach at the individual level by differentiating between cognitive and social perspective-taking ability provides an opportunity to understand the relation between perspective taking and individual students’ cooperative learning outcomes. keypoints patterns of cooperative behaviours of 115 fifth-grade students could be captured in four different profiles. these profiles are related to perspective taking conceptualized at the group level, but not to individual-level perspective-taking ability. profile membership, cognitive perspective-taking ability, and group-level perspective-taking ability could not predict students’ learning outcomes. social perspective-taking ability and reading comprehension did positively predict learning outcomes. the effect of perspective-taking ability on cooperative behaviours and learning outcomes depends on its conceptualization and measurement level. acknowledgments we thank irene enschede, sharyl van harlingen, veerle hoff, annelies keller, lisanne noordzij, mirjam post, sanne samsom, nicole spaans, and inge vreugdenhil for their help with material construction, data collection, and/or transcription. we also thank all participating fifth-grade students and their teachers who have welcomed us in their classrooms. references arjava, m., salovaara, h., häkkinen, p., & järvelä, s. (2007). combining individual and group-level perspectives for studying collaborative knowledge construction in context. learning and instruction, 17, 448–459. doi:10.1016/j.learninstruc.2007.04.003 asterhan, c. s. c., & schwarz, b. b. (2016). argumentation for learning: well-trodden paths and unexplored territories. educational psychologist, 51, 164‒187. doi:10.1080/00461520.2016.1155458 baer, j. (2003). grouping and achievement in cooperative learning. college teaching, 51, 169–175. doi:10.1080/87567550309596434 baron-cohen, s., wheelwright, s., hill, j., raste, y., & plumb, i. (2001). the ‘reading the mind in the eyes’ test revised version: a study with normal adults, and adults with asperger syndrome or high-functioning autism. journal child psychopathology psychiatry, 42, 241‒251. doi:10.1111/1469-7610.00715 barron, b. (2003). when smart groups fail. the journal of the learning sciences, 12, 307–359. doi:10.1207/s15327809jls1203_1 beersma, b., bechtoldt, m. n., & schouten, m. e. (2018). when ignorance is bliss: exploring perspective taking, negative state affect and performance. small group research, 49, 576‒599. doi:10.1177/1046496418775829 cigala, a., mori, a., & fangareggi, f. (2015). learning others’ point of view: perspective taking and prosocial behavior in preschoolers. early child development and care, 185, 1199‒1215. doi:10.1080/03004430.2014.987272 clark, h. h., & brennan, s. (1991). grounding in communication. in l. b. resnick, j. m. levine & s. teasley (eds.), perspectives on socially shared cognition (pp. 127-149). washington, dc: american psychological association. cohen, e. g. (1994). restructuring the classroom: conditions for productive small groups. review of educational research, 64, 1–35. doi:10.3102/00346543064001001 collins, l. m., & lanza, s. t. (2010). latent class and latent transition analysis. hoboken, nj: wiley. davis, m. h. (1983). measuring individual differences in empathy: evidence for a multidimensional approach. journal of personality and social psychology, 44, 113‒126. doi:10.1037/0022-3514.44.1.113 davis, m. h., conklin, l., smith, a., & luce, c. (1996). effect of perspective taking on the cognitive representation of persons: a merging of self and other. journal of personality and social psychology, 70, 713–726. doi:10.1037/0022-3514.70.4.713 di eugenio, b., jordan, p. w., thomason, r. h., & moore, j. d. (2000). the agreement process: an empirical investigation of human–human computer mediated collaborative dialogs. international journal of human–computer studies, 53, 1071–1075. doi:10.1006/ijhc.2000.0428 driskell, t., driskell, j. e., burke, c. s., & salas, e. (2017). team roles: a review and integration. small group research, 48, 482–511. doi:10.1177/1046496417711529 egberink, i. j. l., janssen, n. a. m., & vermeulen, c. s. m. (2015). cotan beoordeling 2015 leerling,en onderwijsvolgsysteem begrijpend lezen [cotan review 2015, evaluation of the system for monitoring educational progress: reading comprehension]. amsterdam, the netherlands: boom test uitgevers. epley, n., morewedge, c. k., & keysar, b. (2004). perspective taking in children and adults: equivalent egocentrism but differential correction. journal of experimental social psychology, 40, 760–768. doi:10.1016/j.jesp.2004.02.002 erkens, g. (2004). the dynamics of coordination in collaboration. in j. van der linden & p. renshaw (eds.), shifting perspectives to learning, instruction, and teaching (pp. 191‒213). dordrecht, the netherlands: kluwer academic publishers. erkens, g. (2005). multiple episode protocol analysis (version 4.10) [computer software]. utrecht: utrecht university. erkens, g., & janssen, j. (2008). automatic coding of dialogue acts in collaboration protocols. international journal of computer-supported collaborative learning, 3 , 447‒470. doi:10.1007/s11412-008-9052-6 erkens, g., jaspers, j., prangsma, m., & kanselaar, g. (2005). coordination processes in computer supported collaborative writing. computers in human behavior, 21, 463‒486. doi:10.1016/j.chb.2004.10.038 evers, g. (2008). programma voor berekening cito leesindex voor het basisonderwijs (p-clib, version 3.0) [computer software]. arnhem, the netherlands: cito. falk, d. r., & johnson, d. w. (1977). the effects of perspective-taking and egocentrism on problem solving in heterogeneous and homogeneous groups. the journal of social psychology, 102, 63‒71. doi:10.1080/00224545.1977.9713241 field, a. (2009). discovering statistics using spss (3 rd ed.). los angeles, ca: sage. galinsky, a. d., ku, g., & wang, c. s. (2005). perspective-taking and self-other overlap: fostering social bonds and facilitating social coordination. group processes and intergroup relations, 8, 109‒124. doi:10.1177/1368430205051060 galinsky, a. d., & moskowitz, g. b. (2000a). perspective-taking: decreasing stereotype expression, stereotype accessibility, and in-group favoritism. journal of personality and social psychology, 78, 708–724. doi:10.1037/0022-3514.78.4.708 gillies, r. m. (2014). developments in cooperative learning: review of research. anales de psicología, 30, 792–801. doi:10.6018/analesps.30.3.201191 hawk, s. t., keijsers, l., branje, s. j. t., van der graaff, j., de wied, m., & meeus, w. (2013). examining the interpersonal reactivity index (iri) among early and late adolescents and their mothers. journal of personality assessment, 95, 96–106. doi:10.1080/00223891.2012.696080 hickendorff, m., edelsbrunner, p. a., mcmullen, j., schneider, m., & trezise, k. (2018). informative tools for characterizing individual differences in learning: latent class, latent profile, and latent transition analysis. learning and individual differences, 66, 4–15. doi:10.1016/j.lindif.2017.11.001 hooper, d., coughlan, j., & mullen, m. r. (2008). structural equation modelling: guidelines for determining model fit . electronic journal of business research methods, 6, 53–60. retrieved from http://www.ejbrm.com/volume6/issue1/p53 janssen, j., cress, u., erkens, g., & kirschner, p. a. (2013). multilevel analysis for the analysis of collaborative learning. in c. e. hmelo-silver, c. a. chinn, c. k. k. chan, & a. m. o’donnell, the international handbook of collaborative learning (pp. 41–56). new york, ny: routledge. janssen, j., erkens, g., kirschner, p. a., & kanselaar, g. (2012). task-related and social regulation during online collaborative learning. metacognition and learning, 7, 25–43. doi:10.1007/s11409-010-9061-5 järvelä, s., & häkkinen, p. (2002). web-based cases in teaching and learning ‒ the quality of discussions and a stage of perspective taking in asynchronous communication. interactive learning environments, 10, 1–22. doi:10.1076/ilee.10.1.1.3613 järvenoja, h., & järvelä, s. (2009). emotion control in collaborative learning situations – do students regulate emotions evoked from social challenges? british journal of educational psychology, 79, 463–481. doi:10.1348/000709909x402811 järvenoja, h., & järvelä, s. (2016). regulating emotions together for motivated collaboration. in m. baker, j. andriessen, & s. järvelä (eds.), affective learning together: social and emotional dimensions of collaborative learning (pp. 162-182). new york, ny: routledge. johnson, d. w. (1971). the effectiveness of role reversal: the actor or the listener. psychological reports, 28, 275‒282. doi:10.2466/pr0.1971.28.1.275 johnson, d. w. (1975a). affective perspective taking and cooperative predisposition. developmental psychology, 11, 869‒870. johnson, d. w. (1975b). cooperativeness and social perspective taking. journal of personality and social psychology, 31, 241‒244. doi:10.1037/h0076285 johnson, d. w. (2015). constructive controversy: theory, research, practice. cambridge, uk: cambridge university press. johnson, d. w., & johnson, r. t. (2009). an educational psychology success story: social interdependence theory and cooperative learning. educational researcher, 38, 365‒379. doi:10.3102/0013189x09339057 king, a. (2002). structuring peer interaction to promote high-level cognitive processing. theory into practice, 41, 33–40. doi:10.1207/s15430421tip4101_6 kreijns, k., kirschner, p. a., & jochems, w. (2003). identifying the pitfalls for social interaction in computer-supported collaborative learning environments: a review of the research. computers in human behavior, 19, 335‒353. doi:10.1016/s0747-5632(02)00057-2 michelon, p., & zachs, j. m. (2006). two kinds of visual perspective taking. perception & psychophysics, 68, 327‒337. doi:10.3758/bf03193680 mouw, j. m., saab, n., janssen, j., & vedder, p. (2019). quality of group interaction, ethnic group composition, and individual mathematical learning gains. social psychology of education, 22, 383–403. doi: 10.1007/s11218-019-09482-w mouw, j. m., saab, n., pat-el, r. j., & van den broek, p. (2019). studentand task-related predictors of primary-school students’ perceptions of cooperative learning activities. pedagogische studiën, 96(2), 98–122. oberski, d. (2016). mixture models: latent profile and latent class analysis. in j. robertson & m. kaptein (eds.), modern statistical methods for hci (pp. 275–287). switzerland: springer international publishing. ortiz, a. e., johnson, d. w., & johnson, r. t. (1996). the effect of positive goal and resource interdependence on individual performance. the journal of social psychology, 136, 243–249. doi:10.1080/00224545.1996.9713998 pons, f., harris, p. l., & de rosnay, m. (2004). emotion comprehension between 3 and 11 years: developmental periods and hierarchical organization. european journal of developmental psychology, 1, 127–152. doi:10.1080/17405620344000022 pronin, e., puccio, c. t., & ross, l. (2002). understanding misunderstanding: social psychological perspectives. in t. gilovich, d. griffin, & d. kahneman (eds.), heuristics and biases: the psychology of intuitive judgment (pp. 636‒665). cambridge, uk: cambridge university press. rieffe, c., oosterveld, p., miers, a. c., terwocht, m. m., & ly, v. (2008). emotion awareness and internalizing symptoms in children and adolescents: the emotion awareness questionnaire revised. personality and individual differences, 45, 756‒761. doi:10.1016/j.paid.2008.08.001 ruby, p., & decety, j. (2004). how would you feel versus how do you think she would feel? a neuroimaging study of perspective-taking with social emotions. journal of cognitive neuroscience, 16, 988‒999. doi:10.1162/0898929041502661 strijbos, j. w., martens, r. l., jochems, w. m., & broers, n. j. (2007). the effect of functional roles on perceived group efficiency during computer-supported collaborative learning: a matter of triangulation. computers in human behavior, 23, 353‒380. doi:10.1016/j.chb.2004.10.016 tjosvold, d., & johnson, d. w. (1977). effects of controversy on cognitive perspective taking. journal of educational psychology, 69, 679‒685. doi:10.1037/0022-0663.69.6.679 tjosvold, d., & johnson, d. w., & johnson, r. t. (1984). influence strategy, perspective-taking, and relationships between highand low-power individuals in cooperative and competitive contexts. the journal of psychology, 116 , 187‒202. doi:10.1080/00223980.1984.9923636 trötschel, r., hüffmeier, h., loschelder, d. d., schwartz, k., & gollwitzer, p. m. (2011). perspective taking as a means to overcome motivational barriers in negotiations: when putting oneself into the opponent’s shoes helps to walk toward agreements. journal of personality and social psychology, 101, 771‒790. doi:10.1037/a0023801 vermunt, j. k., & magidson, j. (2013). technical guide for latent gold 5.0: basic, advanced, and syntax. belmont, ma: statistical innovations inc. vogindroukas, i., chelas, e. n., & petridis, n. e. (2014). reading the mind in the eyes test (children’s version): a comparison study between children with typical development, children with high-functioning autism and typically developed adults. folia phoniatrica et logopaedica, 66, 18‒24. doi:10.1159/000363697 webb, n. m. (2013). information processing approaches to collaborative learning. in c. e. hmelo-silver, c. a. chinn, c. k. k. chan, & a. o’donnell. (eds.), the international handbook of collaborative learning (pp. 19‒40). new york, ny: routledge. webb, n. m., farivar, s. h., & mastergeorge, a. m. (2002). productive helping in cooperative groups. theory into practice, 41, 13‒20. doi:10.1207/s15430421tip4101_3 wegerif, r., & mercer, n. (1996). computers and reasoning through talk in the classroom. language and education, 10, 47–64. doi:10.1080/09500789608666700 wellman, h. m. (2018). theory of mind: the state of the art. european journal of developmental psychology, 15, 728‒755. doi:10.1080/17405629.2018.1435413 appendix 1 table a1 cooperative learning rules microsoft word jones_finalproofs.docx frontline learning research vol. 10 no. 1 (2022) 46 75 issn 2295-3159 corresponding author: cheryl jones, college of health college of science, health, engineering and education, murdoch university, western australia. c.a.jones@murdoch.edu.au doi:https://doi.org/10.14786/flr.v10i1.851 interpersonal affect in groupwork: a comparative case study of two small groups with contrasting group dynamics outcomes cheryl jones1, simone volet1, deborah pino-pasternak2 & olli-pekka heinimäki3 1 murdoch university, australia 2 university of canberra, australia 3 university of turku, finland article received 28 april 2021 / article revised 24 february 2022 / accepted 4 july / available online 2 august 2022 abstract teamwork capabilities are essential for 21st century life, with groupwork emerging as a fruitful context to develop these skills. case studies that explore interpersonal affect dynamics in authentic higher education groupwork settings can highlight collaborative skills development needs. this comparative case-study traced the sociodynamic evolution of two groups of first-year university students to investigate the high collaborative variance outcomes of the two groups, which reported starkly contrasting group dynamics (negative and dysfunctional or positive and collaborative). mixedmethods (video-recorded observations of five groupwork labs over one semester, and group interviews) provided interpersonal affect data as real-time visible behaviours, and the felt experiences and perceptions of participants. the study traced interpersonal affect dynamics in the natural fluctuation of not just task-focused (on-task), but also explicitly relational (off-task) interactions, which revealed their function in both task participation and group dynamics. findings illustrate visible interpersonal affect behaviours that manifested and evolved over time as interactive patterns, and group dynamics outcomes. fine-grained analysis of interactions unveiled interpersonal affect as a collective, evolving process, and the mechanism through which one group started and stayed highly positive and collaborative over the semester. the other group showed a tendency towards splitting to undertake tasks early, leading to low group-level interpersonal attentiveness, and over time, subgroups emerged through interactions both off-task and on-task. the study made visible the pervasive nature of interpersonal affect as enacted through seemingly inconsequential everyday behaviours that supported the relational and task-based needs of groupwork, and those behaviours which impeded collaboration. keywords: group dynamics; interpersonal affect; groupwork; socioemotional interaction; higher education jones, volet, pino-pasternak & heinimäki 47 | f l r 1. introduction while it is generally believed that engaging in groupwork in higher education prepares students for future teamwork (curşeu et al., 2018), the relational realm of their social interactions (i.e., interpersonal dynamics) can be highly challenging (näykki et al., 2014). as bakhtiar et al. (2018) have argued, not only academic performance, but process outcomes such as students’ interpersonal experiences, are an essential part of the collaboration picture given that perceived experiences influence attitudes towards future groupwork. research on groupwork learning processes has traditionally theorised social interaction in terms of the dual function of the cognitive for performing shared learning tasks, and the socioemotional for social (i.e., relational) performance (isohätälä et al., 2019; kreijns et al., 2003). while such distinctions may serve analytical purposes, the interdependent nature of cognitive and socioemotional processes of groupwork are also widely recognised, although the inherently relational aspect has traditionally been a secondary focus for groupwork research (baker et al., 2013). yet, as baker et al. (2013) note, for some students, the social can be particularly salient, and as students often struggle with the social dynamics of groupwork (näykki et al., 2014) there is a need for better understanding these aspects through case studies that closely examine real time interactions in authentic group situations. relying exclusively on post hoc individual self-report data cannot shed light on the social dynamics as manifest through the interdependent actions that unfold between participants. further, the function of affect as inherently interpersonal phenomena in social interaction is often overlooked due to its pervasive and hidden in plain sight nature, yet as barsade and knight’s (2015) review of group affect research has found, it is an important part of understanding the group dynamics puzzle. the present research was grounded in a perspective of interpersonal affect as inherently social (i.e., relational) and dynamically evolving over time (jones et al., 2021; mesquita & boiger, 2014) to examine the starkly contrasting social dynamics outcomes reported by two groups, each with four members, who were first year teacher education students undertaking a mandatory introductory science unit. the aim was to understand the function of interpersonal affect, as visible behavioural phenomena enacted by participants within their moment-to-moment interactions, and how it could explain the contrasting perceptions of the two groups regarding their group dynamics as negative, or positive. the following sections present the conceptual framework that guided the present study, and selected research in social, and educational psychology, that has empirically examined affect phenomena as social and dynamic in groupwork situations using dynamic methods (i.e., observations). 1.1 affect as inherently interpersonal and temporally evolving phenomena in social interaction affect has traditionally been studied as individual (i.e., intrapersonal) phenomena incorporating a range of affective states, such as the preferences, attitudes, moods, affect dispositions, interpersonal stances, and emotions that individuals experience or express (scherer, 2005, p. 704). the american psychological association, for example, defines affect as: any experience of feeling or emotion, ranging from suffering to elation, from the simplest to the most complex sensations of feeling, and from the most normal to the most pathological emotional reactions. often described in terms of positive affect or negative affect, both mood and emotion are considered affective states. (american psychological association, n.d.) a paradigm shift in recent decades “from intrapersonal to interpersonal” perspectives, however, reflects growing recognition of the social nature of emotions (van kleef, 2021, p. 91) and affect phenomena more broadly (kuppens, 2015). hess and hareli (2019), for example, posit that even the most routine of everyday social encounters involve some emotion exchange, which acts as the “communicative signals” (p. 2) that coordinate social interaction (van kleef, 2021), and which are often taken for granted due to their pervasive presence. according to philosopher sheets-johnstone (2009), affect in its most fundamental sense compels avoidance or engagement and can be seen as “responsivity, jones, volet, pino-pasternak & heinimäki 48 | f l r a feature affectively characterizable as interest or aversion, hence as movement toward or away from something in the environment” (p. 376), and described variously in terms of unpleasant-pleasant, goodbad, positive-negative, and so on. the nature of affect as interpersonally manifest and evolving over time in social interaction, is highlighted in mesquita and boiger’s (2014) sociodynamic model of emotions, which describes emotions as arising through interaction with others and serving an important role in cultivating the cohesion of the sociocultural contexts in which they occur. according to mesquita and boiger (2014), affect and social interaction “form one system” (p. 298) such that the affect (i.e., emotions; moods) that arises during social encounters is not reducible to an individual’s experience or expression; it is part of the interpersonal situation as it unfolds in groupwork. this sociodynamic perspective highlights how affect is collectively cocreated (e.g., as group climate, conflict, or mood), as demonstrated in social psychology research that has focused on its visible nature in social contexts. for example, bartel and saavedra (2000) posited that for the relational effect of group mood to manifest it must be communicated in social interaction through visible behaviours. they demonstrated group mood as perceptible phenomena in 70 diverse workgroups using an observation instrument based on affect valence and activation, their observations aligning with participants’ self-reports. barsade’s (2002) experimental study with university students then showed a contagion effect of affect as dynamically evolving in group interactions, in which a confederate enacted bartel and saavedra’s (2000) behavioural indicators. the relational effect of affect phenomena (i.e., as group mood, climate, tone) has subsequently been illustrated in education and workplace contexts (barsade & knight, 2015). as slaby (2016) observed, “relational affect is often more a matter of specific modes of interaction various ways of beingand acting-together in a situation, modes of joint or cocomportment regardless of whether these modes of interaction assume the shape of a specific emotion type or not” (p. 8). based on the above, the construct interpersonal affect was conceptualised in the present study as visible negative or positive behaviours in order to explore their manifestation, and function in the fundamentally relational realm of groupwork. studying the function of affect in social interaction also requires its conceptualization as dynamically manifesting and evolving over time (kuppens, 2015). reviews of the research on affect in groups have highlighted its dynamic temporal nature, such as the way in which affect during early group life has been shown to impact how groups interact and develop going forward (barsade & knight, 2015). observational analyses of group interaction in collaborative learning contexts (e.g., bakhtiar et al., 2018; kwon et al., 2014) have found that socioemotional interactions influenced ongoing group interaction, including group learning processes. for example, observations of group interaction have found that negative socioemotional interactions can influence group learning processes over time, such as reducing task engagement (e.g., näykki et al., 2014). affect can also evolve as temporal interactive patterns. for example, järvenoja et al. (2019) reported temporal interaction patterns in an explorational study of emotional regulation processes, which found groups exhibited three types of challenges – cognitive; emotional and motivational; social context and interaction – that evolved as temporal patterns in the absence of any perceptible collective emotion regulation. more studies of groups’ real time interactions, that shed light on affective processes as they spontaneously arise and unfold are needed, as “collaborative learning is a temporally unfolding process, and as such, can only be captured as a series of interactions emerging over time” (isohätälä et al., 2019, p. 833). furthermore, observations of group interactions often examine episodes, such as socioemotional hotspots, emotion regulation processes, or interactions from one meeting, and studies that trace affect phenomena as they arise moment-to-moment and sequentially unfold over longer periods are also needed to explore how they emerge as collective, group-level relational dynamics. jones, volet, pino-pasternak & heinimäki 49 | f l r 1.2 interpersonal affect and the importance of group dynamics outcomes recent studies on emotion regulation in collaborative learning have been instrumental in highlighting the pervasive nature of affect as innately interpersonal in groupwork, unveiling positive socioemotional behaviours that are important in supporting the quality of group learning processes. for example, widely agreed socioemotional behaviours found to support group learning, and which are innately relational, include providing encouragement (e.g., bakhtiar et al., 2018; isohätälä et al., 2018; järvenoja et al., 2019; kwon et al., 2014; lobczowski et al., 2021), and displaying respect (e.g., bakhtiar et al., 2018; isohätälä et al., 2018) towards one another. conversely, socioemotional behaviours found to hinder group learning processes include undermining, rejecting, or overruling others’ contributions (e.g., bakhtiar et al., 2018; näykki et al., 2014). groupwork often also involves socioemotional challenges such as participants’ anxiety, and frustration (järvenoja et al., 2019), and emotion regulation strategies that can have an unfavourable impact, for example complaining, or venting, which can spread among participants (lobczowski et al., 2021). research has also identified conflict emergence due to inadequate regulation of relational challenges, proving detrimental to task engagement (näykki et al., 2014), and groups also often avoid the critical argumentation needed for problem solving in favour of maintaining positive relations (e.g., isohätälä et al., 2018; sohr et al., 2018). this suggests that participants can struggle balancing task and relational demands of collaboration (näykki et al., 2014). research on emotion regulation processes has thus shown the ubiquitous presence of affect and its important function in the quality of joint learning processes, and their deeply intertwined nature with the fundamental relational realm of groupwork. yet, as garcía et al. (2020) note, it remains the case that typically, “socioemotional interactions are studied in function of the results of the task and not as a phenomenon of interest in itself” (p. 209) and interactions not necessarily oriented to the learning task also warrant attention (järvenoja et al., 2017). thus, there remains much to be understood about the jointly manifested nature of affect phenomena in groups’ social dynamics, and the subsequent impact of interpersonal affect on participants’ subjective experience of these dynamics. the importance of participants subjective experiences of their group dynamics was underscored by bakhtiar et al. (2018, p.59) who argued that “although performance is commonly used as an indicator of productive collaboration, another important indicator is group members’ perceptions of their experience, as these perceptions are carried forward as beliefs and knowledge informing approaches to future collaborative work.” in the present study, following poupore (2018), groups’ alternative negative or positive dynamics were conceptualised as outcomes on the grounds that the group dynamics of each meeting are viewed as micro-outcomes which serve as inputs to subsequent meetings, and critically, the self-reports of participants at semester end. group dynamics can be broadly understood as: the processes, operations, and changes that occur within social groups, which affect patterns of affiliation, communication, conflict, conformity, decision making, influence, leadership, norm formation, and power. the term…emphasizes the power of the fluid, ever-changing forces that characterise interpersonal groups. (american psychological association, n.d.) according to reviews of the group affect literature (e.g., barsade & knight, 2015; knight & eisenkraft, 2015), understanding group dynamics outcomes requires examining affect phenomena. this was a key finding of barsade’s (2002) study, which showed that affect dynamically evolved over the course of a meeting, impacting the group dynamics of the experimental groups. forsyth (2014) describes group dynamics as “the influential actions, processes, and changes that occur” (p. 2). in the present study, interpersonal affect is thus examined as influential actions (i.e., negative or positive behaviours) that unfolded at the micro-temporal level, and which evolved over time into macro-temporal interactive patterns that contributed to the groups’ contrasting dynamics outcomes. as the present study focused primarily on the relational (otherwise known as affective) realm, group interaction was extended beyond traditional task focus to incorporate off-task interactions. kreijns jones, volet, pino-pasternak & heinimäki 50 | f l r et al. (2003) argued that off-task interaction is typically affect-laden, less formal, and a space where people can establish relationships. according to vygotsky (1978), the intersubjectivity that occurs in social interaction is fundamental to human relations (garcía et al., 2020) thus off-task interaction equally has relevance in understanding how the starkly contrasting groups intersubjectively cocreated their social understanding. along this line, barkaoui et al. (2008) adopted a vygotskian perspective for their analysis of off-task interaction, arguing that all interaction, including affective (i.e., relational) interaction off-task, is germane to collaboration. other empirical research has also shown that off-task interaction influences ongoing interaction. for example, in experimental research with university students, pre-meeting small-talk influenced ongoing positive socioemotional interactions (yoerger et al., 2018), and in workplace teams, informal (e.g., sports, weather) chat was found to be infused with interpersonal affect that had a positive relational and task impact (gorse & emmitt, 2009). in the present study, which focused primarily on the relational realm of groups, all of the off-task interactions were therefore conceptualised as innately affective, as described further below in section 2.4 observational data coding. 1.3 the present study this study explored interpersonal affect, of two groups of first-year university students who reported starkly contrasting group dynamics outcomes (negative and dysfunctional; positive and collaborative). the aim was to examine the extent to which the two groups’ contrasting perceived dynamics outcomes could be understood in relation to the visible interpersonal affect that arose and evolved in their ongoing interactions. in the present study, off-task interactions were distinguished from those on-task to enable the exploration of how interpersonal affect in relational talk off-task may influence not only ongoing interpersonal affect, but also participants’ evolving task participation given that learning was after all, the groups’ raison d’être (thus foregrounding relational but not ignoring its task function). the present study therefore traced interpersonal affect, in a phenomenological sense following the participants’ interpersonal affect in social interaction, in its natural fluctuation off-task and on-task. task participation was operationalised as participant/s contributing to the group task, evidenced through nonverbal and verbal communications (isohätälä et al., 2019) or undertaking task functions (e.g., interacting with materials, with or on behalf of the group). absence of task participation, in turn, was apparent by participant/s talking off-task. as groups naturally fluctuate between more or less informal modes, from spontaneous small-talk to task participation, the evolution of interpersonal affect is sequentially interwoven throughout the fabric of these two broad intersecting domains. off-task and on-task interactions comprise the whole social context, expected to provide unique insights into the relational function of interpersonal affect as ontologically unfolding in groupwork. two research questions guided this study: rq1: how does interpersonal affect manifest and evolve over time in the off-task and on-task interactions, of two groups that reported contrasting group dynamics outcomes following their groupwork? rq2: what kind of interpersonal affect phenomena characterise the fluctuation of off-task and on-task interaction? jones, volet, pino-pasternak & heinimäki 51 | f l r 2. methodology 2.1 research design this comparative case study (yin, 2018) explored two small groups that were video-recorded undertaking shared science activities and then interviewed at the end of semester, yielding both external observations and participants’ own perspectives. following näykki et al’s (2014) suggestion, a comparative case study was used to examine moment-to-moment interpersonal affect behaviours that taken together, contributed to participants perceived negative or positive group dynamics outcomes. these kinds of everyday, hidden in plain sight phenomena are often only made apparent through contrast (mills et al, 2012), and could contribute to better understanding interpersonal dynamics given that participants’ perceptions of their group interaction influence their engagement in ongoing interaction. 2.2 participants and context data for the study are a subset of a larger research project conducted within an introductory science unit for first-year teacher education students, in which the students were filmed during five groupwork labs. twenty-two groups, spread across six different lab classes with different teachers, remained intact with their four members attending all classes over the semester (no natural attrition). two case groups were selected as the focal point of this study, based on their starkly contrasting (negative or positive) self-reported group dynamics in their group interviews at semester end. specifically, one group repeatedly expressed highly positive dynamics and an enjoyable experience, while the other reported salient negative events and ongoing interpersonal tensions. the two groups were from the same lab class, hence had the same teacher and lab conditions, limiting confounds potentially associated with different teachers, therefore making them highly suitable for this comparative analysis. group a comprised two females and two males, and group b, three females and one male. each group had one mature-aged student (over 25-years-old), and three under 25-years-old. the students selfselected into groups, however being a first-year unit, typically did not know one another well yet, and tended to form into groups as they were seated when instructed to form groups. students were also asked to stay in their groups for the semester but could discuss with the teacher if they wanted to change. participants were also advised that they could withdraw from the research at any time. approval for the research was provided by the university’s human research ethics committee and conducted in accordance with the national research code of conduct. participants provided written consent for videorecordings and interviews. all names used are pseudonyms. the research context was a science unit aimed to develop first-year student teachers’ knowledge of fundamental concepts in chemistry, earth sciences, and physics, and understanding of scientific inquiry including practical experimental skills of planning and conducting investigations. weekly lab activities consisted of one two-hour class, in which learning tasks were undertaken in small groups, using everyday materials for hands-on experiments. the groups were advised to work together (i.e., not to split but to work as an intact group on their activities). details of the five labs’ science activities can be found in appendix a. 2.3 data sources 2.3.1 video-recorded observations of nine groupwork labs undertaken, five were video-recorded, filming the groups in their initial three weeks working together, then mid-semester, and their final group activity, providing a macrotemporal perspective spanning the semester. the teacher-instructed labs included collective hands-on experimenting followed by group science reasoning. activities included shared planning, jones, volet, pino-pasternak & heinimäki 52 | f l r experimentation, and conceptual reasoning. almost eight hours (479.5 minutes) of video-footage of the two groups in five labs, were coded and analysed (see table 2 for number of coded interactions and breakdown by off-task and on-task). the duration of coded observations for each activity ranged from 31-55 minutes. 2.3.2 group interviews conversation style focus group interviews (approximately one hour with each group separately), were audio-recorded and transcribed. they elicited participants’ feelings and perceptions regarding their groupwork. although all participants were present for all video-recorded labs, one participant in both groups declined the interview. participants were invited to start the conversation with a general question asked to ignite discussion: “what would you like to share [with us] about your experience in the labs?” then followed conversation that was inspired by a video-stimulated recall interview approach (sherin, 2004). videoclips were shown to stimulate informal discussion and directly tap group members’ elucidations of their group interactions, followed by the question, “what would you like to say about this episode?” 2.4 observational data coding the video-recordings were systematically coded using the observer xt behavioural coding software. a coding scheme was developed to exhaustively, and exclusively, parse group interactions into one of 18 discrete codes (see appendix b) and trace visible interpersonal affect as concrete behaviours. (data examples for each code are provided in appendix c). the unit of analysis was a discrete verbal behaviour (i.e., single utterance) or nonverbal behaviour. coding was undertaken at the individual level as each participant could be enacting different behaviours (off-task or on-task) at the same time. the scheme was informed by a review of observational research in education and workplace contexts (e.g., jones et al., 2021), with codes adapted from studies in collaborative learning (e.g., isohätälä et al., 2019; rogat & linnenbrink-garcia, 2011), and kauffeld and lehmann-willenbrock’s (2012) act4teams instrument, which has been extensively validated in workplace and university contexts. while only some on-task behaviours included visible manifestations of affect, with a positive or negative valence, all off-task behaviours were conceptualised as affective. off-task codes were exploratory in nature to tap relational small-talk, humour, and laughter targeted to the group (i.e., positive interpersonal affect), or otherwise non-inclusive behaviours (e.g., whispered side-talk), conceptualised as negative interpersonal affect. the non-affective category (fifth column of appendix b) represents behaviours with no overtly obvious affective valence. the code empty-talk (adapted from kauffeld & lehmann-willenbrock, 2012) was used 103 times of 6,500 coded behaviours (1.6%), therefore was excluded from analysis, as reviewing their occurrence suggested these instances did not impact group interaction. the codes, representing broad interpersonal affect behavioural types, and valence, aimed to tap interpersonal affect as it occurs in the kinds of hidden in plain sight, everyday behaviours of social interaction, what slaby (2016) referred to as the relational affect that is reflected in the ways we behave and act together in a situation, and which do not necessarily always involve expression of a particular emotion. in the off-task categories, for example, codes are generic (e.g., “small-talk”, “humour”, see appendix b). within on-task categories (positive, negative) codes have several behaviours (identified as salient affect behaviours in empirical research as discussed above) grouped together (e.g., “abrupt, curt or rude behaviours; interrupting to over-rule; ignoring”) as the aim was to denote the general type of interpersonal affect behaviours, and valence. for example, “complaining; negative utterances” can be directed to objects, or the task, whereas “criticizing/ running someone down” is clearly directed towards other/s, therefore a different kind of interpersonal affect. jones, volet, pino-pasternak & heinimäki 53 | f l r initial exploratory data analysis by the first author inspired a draft of the coding scheme, which was trialled with a second researcher (fourth author) from a dissimilar sociocultural milieu as we considered that researchers from diverse sociocultural contexts might contribute richly distinct viewpoints on interpersonal affect, and address observer biases. the trial process included joint viewing of video-clips, sharing conceptual and empirical understanding of events (rogat & linnenbrink-garcia, 2013), then an iterative process involving individually coding test data followed by joint meetings, and further independent coding and comparison. following, systematic coding including inter-rater reliability coding (hallgren, 2012), was conducted. to assess inter-rater reliability, a portion of each video was coded by two researchers (1,753 behaviours, 27%). segments for inter-rater coding were randomly selected to comprise a portion from each lab for each group. the overall average interrater reliability produced a cohen’s kappa of κ=.86. by coding category, off-task agreement overall was κ=.92 (positive κ=.90; negative κ=1.0). on-task agreement overall was κ=.85 (negative κ=.82; positive κ=.85). table 1 lists the breakdown of inter-rater agreement by group and lab. disagreements were resolved through discussion and repeated observations, and a small portion (n=43) of highly ambiguous behaviours were coded collaboratively rather than individually (rogat & linnenbrink-garcia, 2013). these were typically in the on-task negative interpersonal affect category, and involved unravelling ambiguous episodes (i.e., several exchanges) that appeared to have some abrupt, curt or rude, behaviours in the context of the interactive flow, such as whether other/s had been deliberately, or inadvertently, ignored in discussion. table 1 inter-rater reliability agreement (cohen’s kappa) week 2 week 3 week 4 week 7 week 12 all labs total group a .85 .92 .81 .90 .90 .89 group b .89 .76 .82 .75 .73 .81 both groups .87 .88 .81 .86 .87 .86 2.5 data analysis 2.5.1 frequency analysis the coded data were exported from observer xt for each group by lab, and their frequencies tabulated and analysed. the data analysis comprised three steps, which reflect the gradual zooming into the data, starting by focusing on the groups’ interactions to identify the extent of off-task and on-task interaction. next, moving to their interpersonal affect within interactions, the coded data were analysed by group, in each lab to develop a picture of the groups’ evolutionary trajectories over the semester. the groups’ evolutionary trajectories were analysed in terms of off-task and on-task interactions, and interpersonal affect, to highlight the emergence of interactive patterns over time, for each group. then, the analysis focused on the breakdown of the visible behaviours that were coded as evidence of interpersonal affect in each group. 2.5.2 qualitative analysis of interpersonal affect in the fluctuation of off-task and on-task interaction the coded observations were then qualitatively analysed in the observer xt by first temporally segmenting each of the ten videos (two groups in five labs), into 30-second segments with brief descriptive labels for an overview of the entire dataset for each group (isohätälä et al., 2018; näykki et al., 2014). the video-recordings were also transcribed in the observer xt. an “elaborated running record” (rogat & linnenbrink-garcia, 2013, p. 105) additionally documented salient nonverbal phenomena (i.e., orientation to other/s, eye-gaze, spatial and material use). common episodes across groups were identified in each lab, which highlighted salient comparative events and evidence of interpersonal affect in the fluctuation of off-task and on-task interactions. jones, volet, pino-pasternak & heinimäki 54 | f l r 2.5.3 qualitative analysis of group interviews the interviews provided a perspective of interpersonal affect as the felt experiences of the participants, and their interpretations of their own and others’ interactions regarding task and relational aspects of their groupwork experience. qualitative content analysis of the focus group interviews (huber, 2020) was conducted after the video-recordings had been coded and fully analysed in the above steps. the analysis was undertaken in two phases. first, content of participants’ talk was explored in terms of: i) members’ own feeling states (negative or positive valence) expressed regarding relational aspects of their groupwork, or; ii) interpersonal perceptions about other/s (negative or positive valence) or about others’ affect state/s; iii) negative or positive comments about the learning tasks; iv) negative or positive comments about task interactions; and v) perceptions of how they got on as a group. the interview transcripts were then explored for any other phenomena that may be insightful regarding the groups’ interpersonal dynamics, such as whether participants exhibited agreement regarding their perceptions. 3. results 3.1 interpersonal affect in off-task and on-task interactions, by valence, and over time (rq1) the findings addressing the first research question are reported in three sub-sections, reflecting the focus of the three data analysis steps: interactions; interpersonal affect; and visible interpersonal affect behaviours. two main group differences emerged, consistent with the self-reports of contrasting negative (group b) and positive (group a) group dynamics. first, interpersonal affect was overall more negative than positive in group b, and highly positive overall in group a; and, secondly, group a’s interactions both off-task and on-task exhibited minimal presence of the side conversations evident in group b. 3.1.1 interactions: by off-task and on-task, and evolution over time the breakdown of off-task and on-task interactions overall is presented in table 2, showing that off-task interactions comprised over 20% of all interactions, of each group. table 2 breakdown of off-task and on-task interactions overall group a frequency (%) group b frequency (%) total frequency (%) total interaction 3,672 (100.0) 2,828 (100.0) 6,500 (100.0) off-task 1,124 (30.6) 628 (22.2) 1,752 (27.0) on-task 2,484 (67.7) 2,161 (76.4) 4,645 (71.5) note: 1. the total of off-task and on-task does not equal 100% as empty talk was excluded from analyses due to evident minimal impact on group interaction and minimal appearance in both groups. 2. all off-task interactions were conceptualised as inherently capturing interpersonal affect. 3. not all on-task interactions exhibited visible affect. jones, volet, pino-pasternak & heinimäki 55 | f l r the presence of off-task interactions was found in both groups in every lab, with a remarkably similar temporal pattern in the frequencies off-task and on-task interaction across groups over the five labs, shown in figure 1. this similarity suggests the presence of common contextual factors (i.e., the task activities) and therefore the need to look beyond task characteristics (casciaro & lobo, 2008) to understand what contributed to the starkly differing valence of the interpersonal affect across groups. figure 1. evolution of off-task and on-task interactions over time. 3.1.2 interpersonal affect: by valence, off-task and on-task, and its evolution over time concerning valence, the two groups were starkly different in their overall interpersonal affect, within both off-task and on-task interactions. as shown in the upper part of table 3, group a exhibited overall 91.0% positive and 9.0% negative, and group b 47.0% positive and 53.0% negative interpersonal affect. these findings align with the two groups’ self-reports of their dynamics at semester end (reported later in section 3.3.3), suggesting that the visible interpersonal affect behaviours identified through the coding (reported in section 3.1.3) contributed to the groups’ contrasting social dynamics outcomes. table 3 breakdown of interpersonal affect by valence overall and within off-task and on-task interactions interpersonal affect by valence group a frequency (%) group b frequency (%) overall (off-task + on-task) positive 1,748 (91.0) 534 (47.0) negative 173 (9.0) 600 (53.0) off-task 1,124 (100.0) 628 (100.0) positive 1,024 (91.1) 224 (35.7) negative 100 (8.9) 404 (64.3) on-task 797 (100.0) 506 (100.0) positive 724 (91.0) 310 (61.3) negative 73 (9.0) 196 (38.7) 63 76,3 79,2 55,7 68 35 22,5 19 41,5 31 0 10 20 30 40 50 60 70 80 90 100 week 2 week 3 week 4 week 7week 12 group a on-task off-task % 78,3 88,4 81 61 73,4 19,4 10,6 18,5 36,3 26 0 10 20 30 40 50 60 70 80 90 100 week 2 week 3 week 4 week 7 week 12 group b % jones, volet, pino-pasternak & heinimäki 56 | f l r the breakdown of interpersonal affect by valence within off-task and on-task interactions is shown in the lower sections of table 3. the findings for group a are particularly striking, being almost identical for off-task and on-task (i.e., around 91.0% positive and 9.0% negative). in contrast, group b display more negative than positive interpersonal affect overall (53%), and somewhat opposite findings across off-task and on-task, specifically, off-task being 35.7% positive and 64.3% negative, and on-task interpersonal affect 61.3% positive and 38.7% negative. the patterns of group b suggest that here too interpersonal affect may traverse off-task and on-task. both groups exhibited more positive than negative interpersonal affect when on-task (see second last row of table 3). however, off-task and on-task interaction could operate at the same time (e.g., member/s in side-talk and other/s on-task) and likewise negative or positive interpersonal affect could coincide, showing that the full picture of how interpersonal affect in off-task and on-task interactions intersected during groupwork is more complex. this is examined qualitatively in rq2. a temporal overview of the breakdown of interpersonal affect by valence within off-task and on-task interactions is shown in figure 2. within on-task interactions the breakdown by valence across labs shows a systematically higher percentage of positive interpersonal affect in group a than in group b but a relatively consistent pattern over time in both positive and negative interpersonal affect across groups. in contrast, within off-task interactions, both groups display noticeable fluxes across the labs, but for group a it is in regard to their positive interpersonal affect while for group b it is in regard to their negative interpersonal affect, with off-task interactions a relatively high source of negative interpersonal affect. of note, however, in the first lab, both groups’ positive interpersonal affect dominated their interaction, especially off-task. yet, group b also started with around 5% negative interpersonal affect off-task and on-task (10.7% combined), which in the second lab rose slightly, and increased off-task thereafter. in contrast, figure 2 shows group a’s negative interpersonal affect was 3.5% for off-task and on-task combined, and remained under 5% over time, with more positive interpersonal affect in both off-task and on-task interactions across labs. this begs the question regarding the interpersonal affect arising in early group life, and its potential function in the divergent dynamics of the groups over time, which is qualitatively explored in rq2. figure 2. evolution of interpersonal affect by valence in off-task and on-task interactions over time. 3.1.3 visible interpersonal affect behaviours: by valence, off-task and on-task, and their evolution over time zooming in to the visible interpersonal affect behaviours of the two groups, figure 3 presents a temporal overview for each group of their positive or negative interpersonal affect behaviours off-task. 0 5 10 15 20 25 30 35 40 45 wk.2 wk.3 wk.4 wk.7 wk.12 group a pos off-task neg off-task pos on-task neg on-task % 0 5 10 15 20 25 30 35 40 45 wk.2 wk.3 wk.4 wk.7 wk.12 group b pos off-task neg off-task pos on-task neg on-task % jones, volet, pino-pasternak & heinimäki 57 | f l r (a full breakdown of the two groups’ positive and negative interpersonal affect behavioural codes offtask and on-task over time is provided in a table in appendix d.) regarding off-task interactions, the most striking group difference was the way in which group a started, and stayed positive over time, compared with group b. figure 3 shows the difference between the two groups off-task (inherently affective relational interactions) regarding side-talk. overall, this comprised half (50.3%) of all off-task interaction of group b, compared to 5.6% in group a. side-talk is off-task chat that innately excludes member/s because its content is not inclusive, or due to its low volume (i.e., whispering), or corporeal positioning (e.g., turned away from other/s). figure 3 shows that side-talk emerged in group b’s first lab (3.2%), increasing steadily over time to peak in week seven (18%). its manifestation and evolution as a pervasive interpersonal affect behaviour over time in group b, and its impact on group dynamics and task participation are examined in rq2. in contrast, figure 3 shows that in their first lab group a exhibited a lot of positive small-talk, involving humour and laughter. small-talk was relationally positive by its characteristics (e.g., content and volume were group-inclusive), and remained comparatively high in group a over time. group b’s positive smalltalk was far less frequent, with group-level off-task humour and laughter consistently lower and decreasing over time. figure 3. evolution of visible positive and negative interpersonal affect behaviours in off-task interactions. regarding on-task interactions, as interpersonal affect was coded into five negative and five positive behaviours in on-task interaction, for clarity of presentation they are shown separately (figures 4 and 5, respectively), for an overview of each group over the semester. scrutinising negative interpersonal affect behaviours within on-task interactions (figure 4) reveals a key intergroup difference: the non-existence of splitting the group in group a. in group b, splitting was apparent in the first lab, but decreased over time. as side-talk showed a temporal increase (figure 3), the possibility of a link between these two behaviours across off-task and on-task interaction is examined in rq2. considering positive interpersonal affect behaviours within on-task interaction (figure 5), across groups (although varying in frequency), efforts in lightening the atmosphere (e.g., task-related humour) featured most, and typically followed a similar pattern. in group a, laughter closely tracked lightening the atmosphere over time, suggesting member/s responding to lightening contributions, such as responding to a task-related joke with laughter. in contrast, in group b although laughter followed lightening at the beginning, it steadily decreased over time, but peaked in the final lab, suggesting there typically was not the same response to lightening contributions as in group a. 0 2 4 6 8 10 12 14 16 18 20 22 24 26 28 30 wk.2 wk.3 wk.4 wk.7 wk.12 group a small talk (pos) humour laughter neg small talk side talk on phone % 0 2 4 6 8 10 12 14 16 18 20 22 24 26 28 30 wk.2 wk.3 wk.4 wk.7 wk.12 group b small talk (pos) humour laughter neg small talk side talk on phone % jones, volet, pino-pasternak & heinimäki 58 | f l r figure 4. evolution of visible negative interpersonal affect behaviours in on-task interactions. figure 5. evolution of visible positive interpersonal affect behaviours in on-task interactions. in summary, an evolutionary perspective of interpersonal affect behaviours indicates that overall, group a, both off-task and on-task, started and remained positive over time. in contrast, group b started somewhat positive but the negative interpersonal affect evident from the beginning appeared to seed, increasing over the semester, most evident off-task. overall, the coding analysis shows that interpersonal affect was pervasive across groups, with its valence transcending off-task and on-task interaction, in both groups. the presence of both off-task and on-task interaction took a remarkably similar pattern across groups over the semester, indicating the similar task conditions, and therefore the need to look beyond the task to explore the groups’ different interpersonal affect trajectories, and contrasting dynamics outcomes. 3.2 manifestation of interpersonal affect in the fluctuation of off-task and on-task interaction (rq2) qualitative analysis explored how the interpersonal affect behaviours identified in the first research question actually manifested in dynamic interactions and contributed to the contrasting group dynamics outcomes reported by participants at semester end. the interplay of off-task interactions in 0 1 2 3 4 5 6 7 8 9 10 wk.2 wk.3 wk.4 wk.7 wk.12 group a complaining / neg talk criticise/ run down abrupt/ curt/ ignore disruptor splitting group % 0 1 2 3 4 5 6 7 8 9 10 wk.2 wk.3 wk.4 wk.7 wk.12 group b complaining / neg talk criticise/ run down abrupt/ curt/ ignore disruptor splitting group % 0 1 2 3 4 5 6 7 8 9 10 wk.2 wk.3 wk.4 wk.7 wk.12 group a inclusive praise/support enthusiasm/int erest lighten atmosphere laughter % 0 1 2 3 4 5 6 7 8 9 10 wk.2 wk.3 wk.4 wk.7 wk.12 group b inclusive praise/support enthusiasm/int erest lighten atmosphere laughter % jones, volet, pino-pasternak & heinimäki 59 | f l r their natural fluctuation with on-task interactions was explored, focusing in particular on the contrast of side-talk (which was over half of all off-task interaction in group b, and just 5.6% in group a). 3.2.1 developing social cohesion early in their first lab, both groups started with a high task focus, with longer episodes of social (offtask) chat occurring late in the lab when students finished the task (e.g., they had cleaned activity materials away, and had stopped discussing their activity outcomes). their first task involved preparation, and observations of two products (see appendix a for task information). the following brief excerpts present each group’s first moments, as they commenced. in group b, one member suggested splitting to manage the task’s two experiments. the teacher, overhearing, instructed that groups undertake the activities together: excerpt 1 group b initial interactions nick alright [standing as teacher finishes explaining activities; glances to nell, then to lisa and abby, who are talking quietly together] lisa [responding] alright, i’ll get the stuff for oobleck nick i’ll try silly slime, i’m no cook! nell well do we want to split in twos: two make the oobleck and two make the silly slime? nick okay, good idea! lisa no [frowning, mouth turned down] nell [responds to lisa] or, make it all together? lisa i don’t want to miss out on making both of them [smiles] teacher [overhearing, tells the group, also reiterating to the class]: no, make it all together group b thus started amicably, as did group a, yet subtle differences were apparent: excerpt 2 group a initial interactions eric alright [standing as teacher finishes explaining] anna let’s go! eric yeah? let’s go grab the stuff [all four stand] anna okay. i’ll grab the playdough eric take that over there so we can just check it [points to lab manual] anna yep eric singing: check yourself before you wreck yourself [as they go together to the materials table] group a commenced similarly to group b with anna saying, “i’ll grab…” but to which eric immediately responded, “so we can…check…”, which subtly adjusts the materials gathering as a collective process. a few minutes later, they returned together with materials for their first product, and subsequently together collected materials for the second activity. in group a, the collective start provided task affordance for group relational development (i.e., social and task cohesion). this was evident in the way that each product’s preparation involved all members working together, with relatively high on-task humour and laughter (see figure 5) that involved all four participants. conversely, group b participants returned separately in dyads a few minutes apart, each with a tray of materials. they commenced working collectively (therefore not coded “split-group”). yet, embedded in their language was an implicit reference to dyadic ownership of product preparation (i.e., the two activities), which they all expressed. for example, nell: “did you guys get the kettle water?” “lisa: no. ours is just normal water;” nick: oh, we needed hot water;” lisa: “shall we do yours first?” jones, volet, pino-pasternak & heinimäki 60 | f l r abby: “are we doing ours yet?” accompanying these kinds of comments, a split-group also emerged occasionally, as lisa quietly discussed on-task with abby, sometimes subtly resisting nick’s contributions. for example, at the beginning of their first lab, nick extended his arm to assist lisa, who was mixing some of the materials she and abby had collected. lisa tells nick that she just needs a spoon, and he immediately dropped his arm. lisa’s manner is not overtly curt, but nor is it inclusive (coded conservatively task na) and these kinds of borderline interactions became the norm in the group (e.g., “yeah wait”, “hang on”, “just read the-!”, “you’re reading the wrong one!”). in the second lab, latent dyadic subgroups emerged again in group b, with nell requesting nick assist gathering materials, while lisa and abby sat chatting off-task for five minutes, making no move to join the task. there was also no attempt to include them, thus the latent dyads of the previous week were tacitly endorsed by all participants, their pattern of commencing labs with off-task and on-task dyads continuing over the semester. in contrast, group a collected materials together, side-talk was usually brief, and interactions were characterised by positive interpersonal affect such as small-talk involving humour and laughter that was group-level (i.e., involved all four members). a qualitative difference between the two groups regarding their off-task interaction, distinguishing side-talk in group b from the small-talk typical of group a, was its low volume, and (generally the same) side-talkers, sometimes turned towards one another exclusively. group b’s split group on-task and side-talk off-task characterised early low social cohesion. this appeared to create procedural confusion with at times two dyads interacting separately, sometimes not knowing what the other was doing or saying. within this context, occasionally other negative interpersonal affect behaviours arose (e.g., highly directive interactions, ignoring) as member/s tried to ascertain what had been done, where they were up to, and so on. in this way the dyadic interactions contributed to the overall group dynamics, not only as non-cohesiveness but also an undertone of tension that occasionally surfaced as the visible negative interpersonal affect behaviours reported in section 3.1. this established the basis for ongoing interactions and highlighted a key group-level relational difference between the two groups, whereby group a for the most part interacted as a group, and group b interacted increasingly in dyads. 3.2.2 interpersonal affect in the evolution of group dynamics and task participation over time the analysis of the interplay of interpersonal affect in the fluctuation of off-task and on-task interaction over the semester highlighted another key difference between the groups in how each group’s interactive dynamics evolved over time. group-level attentiveness to one another was consistently evident in group a. in contrast, in group b, low attentiveness to one another as a group appeared exacerbated by subgroup emergence. the contrast in group-level interpersonal attentiveness is illustrated in the following brief excerpts from week four, in which groups had to plan, conduct, and document an experiment. the first excerpt is characteristic of group b’s communication: excerpt 3 inattentiveness in group b nell so, what’s our hypothesis? [reading aloud from lab book as abby was verbalising a hypothesis, which nell ignores] lisa missy here[signals that abby has a hypothesis] here, just, go on! [encourages abby to continue. nell glances briefly at lisa, then to nick, who is writing] abby briefly laughs [appears shy, quietly spoken, looking down at her writing] abby the more we increase the vinegar the ... [starts reading hypothesis again; nell ignores abby, looks to nick] nick if we increase the vinegar volume the reaction...decrease [abby stops speaking as nick speaks] nell the quicker the reaction rate nick yeah. the reaction rate should quicken jones, volet, pino-pasternak & heinimäki 61 | f l r following these interactions, lisa and abby had a brief, quiet exchange. in excerpt 4 below, group a had discussed and agreed their experiment, then sam suggested an alternative, but without justification. using science reasoning, eric and anna opposed the idea, to which sam remarked “yeah, okay”, looking downwards, and becoming quiet. a few minutes later, eric appeared to take responsibility for group harmony, checking if sam was happy with their decision: excerpt 4 attentiveness in group a eric are you happy with that sam? sam yeah, i just don’t know how you’dit should be alright, it should be alright eric what’s your question? sam how long it will actually be in the air though…to get a good measurement. we can give it a go and then we’ll find out eric … we’ve got the trial sam there’s only one way to find out anyway so, as i say [emphasizing his contribution] anna yeah, let’s just trial then modify here, eric exhibits interpersonal attentiveness, as sam had been quiet and appeared withdrawn as he looked downwards and stopped interacting with the group for almost two minutes (1 minute, 54 seconds). during this time eric continually made task-related jokes (lightening the atmosphere). after asking if sam is happy with the group decision, eric showed further interest in sam’s thoughts: “what’s your question?” following, suzi initiated an off-task relational episode, which appeared to reweave the social fabric of the group. this was evident by all members engaging in the talk and sharing personal information. later, in a similarly challenging episode, another off-task relational conversation followed. the excerpts are characteristic of each group’s myriad, fleeting yet pervasive behaviours of interpersonal affect that together cocreated each group’s social space, illustrating how group a participants routinely exhibited interpersonal attentiveness. conversely, group b participants unintentionally, and deliberately, ignored (i.e., inattentiveness) one another. the analysis revealed the way in which interpersonal attentiveness, highlighted by its consistent presence in group a and its relative absence in group b, was a subtle but relevant form of positive interpersonal affect in the groups’ face-to-face task interactions. the groups’ contrasting interpersonal affect was further emphasized later in the semester in week seven when an off-task peak across groups (e.g., see figure 3) occurred. according to conversation across groups, this appeared due to a combination of two broader contextual factors. first, students had just returned following their first practicum, which permeated off-task conversations. second, the task (electrical circuits) was considered challenging, stated by teachers and students alike. in group b, after initially working as a group, the subgroups emerged with one dyad increasingly off-task and the other on-task, and tensions surfaced (e.g., lisa: “she doesn’t want my help, i’m not smart enough for this”, nell: “what? well, you’re more than welcome to try!”). in contrast, in group a, although there was also uneven task participation with two members doing the lion’s share of making electrical circuits, the group typically engaged together more in small-talk while exploring with the electrical circuits. summarising, the qualitative analysis showed how interpersonal affect behaviours in the interplay of groups’ off-task and on-task interaction in early group life evolved into their diverging relational trajectories (group dynamics) and task participation. specifically, in group b, the early appearance of side-talk off-task and splitting the group on-task, although minimal in the first lab, evolved as an implicit interactive (social) norm, and in contrast, group a started, and stayed positive and intact as a group. 3.3.3 self-report interpretations of groupwork experience the groups’ overall negative or positive interpersonal affect extended into the focus group interviews, and group a expressed being “lucky” regarding their positive experience of their jones, volet, pino-pasternak & heinimäki 62 | f l r groupwork. group b members, complaining about their groupwork experience reported that researchers would see plenty of off-task chat, summarily dismissing the frequent side-talk (e.g., explaining that they “just like talking while doing our work”). overall, group a’s self-reports largely aligned with researchers’ observations. group b members reported negative experiences, which aligned with the observational data, but also displayed lacking awareness regarding their own behaviours in cocreating the group’s social dynamics. group a commenced with a focus on their learning experiences, agreeing that the experiments were fun, but the conceptual reasoning highly challenging. this involved each group discussing and producing a group reasoning statement linking their observations and results of experiments using everyday household materials, with the relevant science concepts. group a participants discussed how their group context supported their science learning through this activity. for example, anna: “i think at the beginning i felt really nervous” but members were “bouncing ideas off each other…i think the confidence came from the groupwork and actually just, having fun.” the others agreed, suggesting that suzi too (absent from the interview) had enjoyed their groupwork. watching video-clips of their final lab, they commented on their task participation, anna reflecting, “it was good that we all contributed…”, eric agreeing: “yeah. you can see that everyone’s really involved…it was good.” they explained how over the semester they engaged everyone by rotating critical task elements including turn taking with leading the conceptual reasoning talk and documenting their joint reasoning statement, sam noting, “i think by the time we started rotating we were all pretty comfortable with each other.” participants discussed their positive interaction off-task, which anna believed had supported her learning: “for me…the contact of the group…we had a little chat and then we got into it…”, adding that it had changed her negative perspective of groupwork. they all reflected that their interactions offtask enabled them to relate well across the age-divide through showing reciprocal interest in one another’s diverse leisure pursuits, thus revealing how they utilised the affordances of off-task chats to bridge their individual differences. the oldest member praised his peers as “champs” in this regard which meant that the group members “were able to relate and talk about things other than just science”, noting awareness of the discord other groups experienced. in contrast, group b participants commenced with “where’s the smart one?” (nell) referring to nick who was absent, then briefly mentioned that they all “hated” the conceptual reasoning and moving swiftly to the relational realm: “it’s hard work being in groups” (nell). lisa and abby discussed how they liked chatting (off-task) as they worked, saying “that’s just what we do” (abby), which they perceived nick disapproved of (what they referred to as “gossiping”) and so ignored them. they suggested nick was too task-focused (nell: he’s like nuh, it has to be all science). yet, their first lab together also shows nick initiating small-talk with the group, which he continued to do over the semester. they were all vocal regarding their perceptions that nick had ignored lisa and abby’s task contributions from the start, lisa repeatedly stating frustration about feeling unheard, while abby commented: “i just gave up saying anything because he didn’t even listen to me. so, i said nothing.” watching video-clips of their final lab elicited further relational dynamics comments, lisa stating: “i was getting so frustrated this day” because nick insisted his boat would be the group boat. however, lisa and nell then explained “we sort of just gave up and we were like, nick, you just make the boat” (on behalf of the group), which reflects more closely what actually occurred. they attributed the negative group dynamics to nick, and ultimately to age difference. while age difference was also present in group a, it was reported as unproblematic (e.g., the above-mentioned comment of the oldest group a member praising his peers as “champs” for how they all engaged as a group both on-task and off-task. the three group b members repeatedly commented, “we were very chilled, all three of us”; “we’re just really relaxed people…” (referring to the high amount of side-talk). yet, they also repeatedly mentioned how “angry” lisa would become during their groupwork, at odds with the “relaxed” comments, and the bickering (e.g., abrupt, curt) between lisa and nell that appeared in all five labs. jones, volet, pino-pasternak & heinimäki 63 | f l r this went unmentioned in the interview, potentially highlighting a limitation of group interviews, where member/s may not be comfortable expressing fully their real feelings and perceptions of their groupwork experience, although abby and nell expressed feeling uncomfortable when lisa became angry (which was attributed to her frustration with nick ignoring her task contributions). in sum, group a’s accounts largely aligned with researchers’ observed salience of interpersonal affect dynamics in the interplay of off-task and on-task interactions. conversely, although group b’s accounts reflected their negative dynamics, their self-reports diverged from researchers’ observations (coding frequencies and fine-grained qualitative analysis), which indicated that all members had contributed to (cocreated) the group’s dynamics. 4. discussion the present study focused on the relational realm of groupwork, emphasizing the important function of interpersonal affect as collectively manifest and dynamically cocreated by all members in the social dynamic of groups. the analysis of interactions confirmed self-reports regarding participants perceived negative or positive group dynamics outcomes, showing that interpersonal affect which arose early in off-task and on-task interactions swiftly became interactive patterns, that shaped task participation, and group dynamics outcomes. this study extends the limited case study research on affect in group interaction as it unfolds in real time, unveiling the microlevel interpersonal affect behaviours that evolve as group patterns, and their function in the collaborative variability that continues to be reported in groupwork (e.g., lobczowski et al., 2021). the coding scheme was instrumental for capturing the frequency, valence, and temporal evolution of interpersonal affect through behaviours manifest in the natural fluctuation of off-task and on-task interaction for a more complete picture of how social dynamics sequentially unfolded (langer-osuna et al., 2020). while group dynamics research has largely relied on static, post hoc self-report methods, this process-oriented study provided a dynamic perspective (vriesema & mccaslin, 2020) that unveiled how in both groups, visible interpersonal affect behaviours comprised a vital piece of the group dynamics puzzle (barsade & knight, 2015). insights afforded through a sociodynamic perspective of affect in the relational realm of groupwork are considered below in terms of key findings, and their implications for groupwork in higher education. the sociodynamic conceptual lens illuminated the innately interpersonal nature of affect as a jointly manifest and pervasive component of group interaction that was irreducible to any one participant (mesquita & boiger, 2014). its perceptible nature (bartel & saavedra, 2000) in the behaviours of participants was traced as dynamically woven through the natural ebb and flow of the groups’ on-task and off-task interactions that cocreated each group’s social space (langer-osuna et al., 2020). the visible nature of interpersonal affect can be viewed through philosopher sheets-johnstone’s (2009) perspective of affect as fundamentally compelling actors’ towards or away from (the group), reflecting the way in which “relational affect” is manifest through interactions that are not always emotion expressions (slaby, 2016). likewise, social cohesion is broadly defined as “the attraction of members to one another and to the group as a whole” (forsyth, 2014, p. 136). importantly, our systematic, microlevel analysis revealed the collective interpersonal affect behaviours that evolved so differently in the two groups, with the opportunity for social cohesion thwarted early in group b despite some positive efforts (evidenced in the coding results). alternatively, social cohesion developed early in group a and was sustained all semester, withstanding inevitable challenges (e.g., during week four, illustrated in excerpt 4). the finding of early interpersonal affect in shaping both groups’ interactive patterns (i.e., negative or positive) over the semester, aligns with previous studies that have identified the influence of early affect in ongoing group processes (e.g., bakhtiar et al., 2018; kwon et al., 2014; näykki et al., 2014), reflecting group development theories regarding the tenuous nature of early group life (braun et jones, volet, pino-pasternak & heinimäki 64 | f l r al., 2020). it highlights the critical role of early interpersonal affect as enacted behavioural phenomena, for the ongoing function of groups. a key finding was the difference between the two groups of early latent subgroup emergence in group b while group a started, and stayed, intact. fine-grained analysis illuminated how in group b subgroups developed through seemingly inconsequential behaviours that solidified into increasingly negative interpersonal affect over time. group dynamics scholars have cautioned the propensity of subgroups for creating tension and conflict (forsyth, 2014), which the present study not only affirmed but unveiled how they actually emerged. the development of subgroups is underexamined, yet their presence has been observed as unhelpful in higher education groupwork. for example, näykki et al’s (2014) case study of group conflict showed participants providing dyadic support for one another, which was not advantageous at group-level, and tensions ultimately diminished the group’s task engagement. in the present study, group b, by their own admission, in the final lab left one member to do the group task alone. conversely, in group a, qualitative analysis revealed that rarely, and briefly, were task functions undertaken dyadically, and then always done in different dyads, which appeared a fruitful way of preventing subgroups from inadvertently developing. this is important for students and educators to be aware of since group tasks often involve some activity dispersion. furthermore, the socially complex dynamic of subgroups, and their consequences also need to be better understood. the off-task and on-task dyads that were present in group b not only reduced all-group task participation but also importantly, decreased participants’ opportunities for improving their social dynamics, and collaboration skills. in group b this appeared to create a kind of spiral effect, not only increasing negative interpersonal affect but also further entrenching the subgroups. the detrimental impact of the subgroups echo collaborative learning literature highlighting the importance of working truly together on a task (e.g., dillenbourg, 1999; summers & volet, 2010). extending the research, which has shown the widespread propensity for students to divide tasks, reducing opportunities for joint engagement (oţoiu et al. 2019), the present study also revealed the important relational implications this can have, including subgroup development off-task. especially in first-year university the opportunity to establish subgroup friendships off-task while working within a group may be enticing but as the present study suggests, can be detrimental for group dynamics and task participation. moreover, the qualitative analysis also showed that side-talk appeared even before tension was evident, signalling its potential in contributing to subgroup emergence also in positive groups, therefore participants need to be aware that seemingly inconsequential side-talk can be counterproductive if frequent and prolonged. indeed, one group a member made a significant contribution to side-talk, which was typically responded to only briefly, preventing its establishment as a relational dynamic, and the potential for subgroup development through off-task chat. alternatively, off-task talk when at whole group-level, enhanced group cohesion (barkaoui et al., 2008). during their interview, group a members themselves attributed the social cohesion they had developed as helpful to what they acknowledged as the challenging task of their group science reasoning. in contrast, at their interview group b participants expressed their aversion to the science reasoning in each lab, which the video-recordings showed at times appeared exacerbated by members not responding to each other’s contributions. this may have fuelled perceptions of this aspect of their groupwork as highly negative (rather than challenging) since strong emotion was also expressed about being ignored. the fine-grained qualitative analysis revealed that another key difference between the two groups was interpersonal attentiveness, with its relative absence in group b a key source of aggravation. as external observers it was relatively easier to discern attentiveness in group a’s responses to one another when coding interactions. in group b, systematic lack of acknowledgement of contributions came from all members, exacerbated by nonverbal behaviours such as low eye-contact, making it difficult to distinguish if participants were deliberately ignored or literally unheard. it appeared a combination of both, stemming from the subgroups and in turn further cementing them. do and schallert’s (2004) study of affect in class discussions found that adult students “tuned out” for numerous reasons, including if discussion was off-track, or to manage negative affect. when side-talk, and non jones, volet, pino-pasternak & heinimäki 65 | f l r responsiveness on-task arose early, group b participants may at times have mentally tuned out to one another. in the literature, interpersonal attentiveness has sometimes been observed as active listening, categorised as a positive socioemotional behaviour (e.g., garcía et al., 2020; isohätälä et al., 2018; rogat & linnenbrink-garcia, 2011). the term interpersonal attentiveness adopted in the present study acknowledges the reciprocal nature of active listening. this is consistent with scherer’s (2005) affect phenomena typology, which includes the interpersonal stances actors adopt in social interaction, such as an “active listening attitude” (garcía et al., 2020, p. 217) displayed through behaviours including eyegaze, nodding, and verbal responses (isohätälä et al., 2018). these were apparent in group a’s high frequency of positive interpersonal affect over the semester (e.g., responses of laughter, reciprocal lightening comments) during task interaction. in group b, an early tendency towards splitting the group, which evolved into subgroup emergence, increased other negative interpersonal affect behaviours, further reducing group-level interpersonal attentiveness. this led to the frustration and anger expressed in the interview, which although reported as stemming from one participant, the analysis revealed that low attentiveness was visible early from all members (i.e., group-level). group a did not exhibit, or report being or feeling unheard. the importance of attentiveness for productive collaboration and positive group dynamics outcomes has been shown in various contexts (e.g., barron, 2003; garcía et al., 2020; rogat & linnenbrink-garcia, 2011; ucan & webb, 2015) and this innately relational (i.e., interpersonal) aspect of groupwork deserves more empirical attention. 5. limitations, future research, and conclusion being an in-depth case study, our sample of participants was necessarily small, but our realtime behavioural data (n=6,500 frequencies) were substantial, capturing a broad range of interpersonal affect behaviours during both off-task and on-task interaction, providing a full picture of group dynamics. examining the contrast groups in five labs over a semester unveiled the wide range of interpersonal affect behaviours that otherwise would not have been revealed as sociodynamically manifest, their evolution as interactive patterns over time, and their function in task participation and the group dynamics outcomes of each group. the explorative case study design means that while the findings cannot be generalised to other groupwork situations, the contrasting nature of the groups contributes to observational studies that help to explain variability in groupwork outcomes through a detailed exploration of a wide range of interpersonal affect behaviours. theoretically, the collective cocreation of interpersonal affect in groupwork, manifest through yet underexplored, taken for granted behaviours may be more pervasive, and influential than is currently understood. although some participants referred to their prior groupwork experiences in the interviews (three in group b, and two in group a) and in video-recordings, the study did not include individual participant data such as prior experience of groupwork. it focused instead on the cocreated, collective nature of interpersonal affect in relational dynamics, given it is now typical in educational contexts and in the workplace, that actors are expected to enter groups with different levels of collaborative experience as well as other individual differences, such as knowledge, attitudes towards groupwork, and goals. however, future research that also includes individual-level background data could provide important insights into how particular individual differences interplay to influence affect, and other group dynamics. related to this point, how affect functions as interpersonal phenomena in socioculturally diverse groupwork settings is an important research area (kuppens et al., 2017; lehmann-willenbrock et al., 2014) of increasing relevance for education, the workplace, and social life more broadly. research on university students’ intercultural social interaction (e.g., ujitani & volet, 2008), for example, has shown that humour expression that is culturally insensitive can result in hurt feelings and misunderstandings. jones, volet, pino-pasternak & heinimäki 66 | f l r individual interviews might also have provided further insight in the present study, as students could be reluctant to fully share their real feelings with their peers. having one participant absent in each group interview is a limitation indicative of “messy” real-life research. the focus group interviews did, however, provide a window into the way in which each group spoke of an absent member, reflecting the contrasting group dynamics, and the way in which three group b members had co-constructed their own social meaning of their group dynamics. the study highlights the value of combining self-report and observations for gaining insight into group dynamics (vriesema & mccaslin, 2020), providing affect data as participants internal feelings and perceptions, and as visibly unfolding sociodynamic phenomena (barsade, 2002). importantly, their combined analysis revealed that participants were, variously, more, or less aware of their own behaviours and how they themselves created their group experiences and outcomes, revealing the extent “perceptions and actual behaviours are related to one another” (lehmann-willenbrock & chiu, 2018, p. 1156). riebe et al. (2016, p. 639) report in their review of higher education teamwork pedagogy that the literature also typically lacks “recognition that students [themselves] have a significant role to play when it comes to the achievement of teamwork learning outcomes.” distinguishing off-task from on-task interactions for analytical purposes confirmed barkaoui et al’s (2008) finding that all interaction is part of collaboration and therefore should be more widely incorporated into research (langer-osuna et al., 2020). the present study also unveiled actual behavioural referents of interpersonal affect as dynamically evolving joint action in groupwork, and perhaps most striking is that even in close proximity face-to-face around their worktable over an entire semester, group b participants complained of feeling unheard. according to ferreira (2021), a key issue for collaborative learning is whether participants are actually able to “understand what it takes to participate in joint action” (p. 1466). this may be better understood, ferreira (2021) proposes, by adopting an embodied perspective that can more deeply incorporate the function of nonverbal phenomena to provide new insights, such as how bodies are utilised in ways that foster or impede collaboration. this could be a fruitful avenue for exploring more deeply the joint nature (barron, 2003) of interpersonal attentiveness and how it develops as a group norm. in this study, we explored interpersonal affect in the relational realm of group interaction in a comparative case study with two small groups of students in the same class, who reported contrasting group dynamics outcomes (negative and dysfunctional; positive and collaborative). systematic coding traced the sequential flow of interpersonal affect in groups’ social interaction as it naturally ebbed and flowed through task focused (on-task) and more informal (off-task) interactions, revealing specific behaviours which arose during early group interaction that were formative for the different relational pathways that each group took over time. the study unveiled seemingly routine, everyday interpersonal affect behaviours (e.g., side-talk; humour; laughter) on a micro time-scale which, taken together, unfolded over the semester as interactive patterns, and group dynamics outcomes. keypoints exploring the interplay of off-task and on-task interactions enabled unique insights into interpersonal affect in groupwork. interpersonal affect emergent in groups’ first meetings, served as affective inputs to subsequent meetings. interpersonal affect behaviours evolved over time into macro-temporal interactive patterns and group process outcomes. combining observations and self-report data revealed variance in participants’ awareness of their own behaviours in co-creating their group dynamics. the study revealed that subgroups can emerge during off-task or on-task interaction, proving detrimental to group cohesion. jones, volet, pino-pasternak & heinimäki 67 | f l r acknowledgments the authors would like to thank the student and teacher participants in the research. the first author was supported by an australian government research training program (rtp) scholarship through the college of science, health, engineering and education, murdoch university, western australia. the work of the second and third authors was supported by the australian research council, under the discovery award (dp150101142), and the fourth author was supported by the turku university foundation (no. 080805). references american psychological association. (n.d.). affect. in apa dictionary of psychology. retrieved 23 january 2020 from https://dictionary.apa.org/affecthttps://dictionary.apa.org/?_ga=2.160359743.368224603.16 09865270-776417432.1606326331 baker, m., andrriessen, j., & järvelä, s. (eds.). (2013). affective learning together: social and emotional dimensions of collaborative learning, pp. 1-30. oxon, uk: routledge. bakhtiar, a., webster, e. a., & hadwin, a. f. (2018). regulation and socio-emotional interactions in a positive and a negative group climate. metacognition and learning, 13(1), 57-90. https://doi.org/10.1007/s11409-017-9178-x barkaoui, k., so, m., & suzuki, w. (2008). is it relevant? the role of off-task talk in collaborative learning. journal of applied linguistics, 5(1), 31-54. https://doi.org/10.1558/japl.v5i1.31 barron, b. (2003). when smart groups fail. journal of the learning sciences, 12(3), 307-359. https://doi.org/10.1207/s15327809jls1203_1 barsade, s. g. (2002). the ripple effect: emotional contagion and its influence on group behavior. administrative science quarterly, 47(4), 644–675. https://doi.org/10.2307/3094912 barsade, s. g., & knight, a. p. (2015). group affect. annual review of organizational psychology and organizational behavior, 2, 21-46. https://doi. org/10.1177/0963721412438352 bartel, c. a., & saavedra, r. (2000). the collective construction of work group moods. administrative science quarterly, 45(2), 197-231. https://doi.org/10.2307/2667070 braun, m. t., kozlowski, s. w., brown, t. a., & deshon, r. p. (2020). exploring the dynamic team cohesion–performance and coordination–performance relationships of newly formed teams. small group research, 51(5), 551-580. https://doi.org/10.1177/1046496420907157 casciaro, t., & lobo, m. s. (2008). when competence is irrelevant: the role of interpersonal affect in task-related ties. administrative science quarterly, 53(4), 655-684. https://doi.org/10.2189/asqu.53.4.655 curşeu, p. l., chappin, m. m., & jansen, r. j. (2018). gender diversity and motivation in collaborative learning groups: the mediating role of group discussion quality. social psychology of education, 21, 289-302. https://doi:10.1007/s11218-017-9419-5 dillenbourg, p. (1999). what do you mean by collaborative learning? in p. dillenbourg, (ed.), collaborative learning: cognitive and computational approaches (pp. 1-19). elsevier. do, s. l., & schallert, d. l. (2004). emotions and classroom talk: toward a model of the role of affect in students’ experiences of classroom discussions. journal of educational psychology, 96(4), 619-634. https://doi.org/10.1037/0022-0663.96.4.619 jones, volet, pino-pasternak & heinimäki 68 | f l r ferreira, j. m. (2021). what if we look at the body? an embodied perspective of collaborative learning. educational psychology review, 33(4), 1455-1473. https://doi.org/10.1007/s10648021-09607-8 forsyth, d. r. (2014). group dynamics (6th ed.). wadsworth cengage learning. garcía, a., olivares, h., simão, l. m., & dominguez, a. l. (2020). socioemotional interactions in collaborative learning: an analysis from the perspective of semiotic cultural psychology. culture & psychology, 1-19. https://doi.org/10.1177/1354067x20976513 gorse, c. a., & emmitt, s. (2009). informal interaction in construction progress meetings. construction management and economics, 27(10), 983-993. https://doi.org/10.1080/01446190903179710 hallgren, k. a. (2012). computing inter-rater reliability for observational data: an overview and tutorial. tutorials in quantitative methods for psychology, 8(1), 23-34. https://doi.org/10.20982/tqmp.08.1.p023 hess, u., & hareli, s. (2019). the emotion-based inferences in context (ebic) model. in u. hess & s. hareli (eds.), the social nature of emotion expression: what emotions can tell us about the world (pp. 1-5). springer. https://doi.org/10.1007/978-3-030-32968-6 huber, m. (2020). video-based content analysis. in m. huber & d. e. froehlich (eds.), analyzing group interactions: a guidebook for qualitative, quantitative and mixed methods (pp. 37–48). routledge. isohätälä, j., näykki, p., & järvelä, s. (2019). cognitive and socio-emotional interaction in collaborative learning: exploring fluctuations in students’ participation. scandinavian journal of educational research, 64(6), 831-851. https://doi.org/10.1080/00313831.2019.1623310 isohätälä, j., näykki, p., järvelä, s., & baker, m. j. (2018). striking a balance: socio-emotional processes during argumentation in collaborative learning interaction. learning, culture, and social interaction, 16, 1-19. https://doi.org/10.1016/j.lcsi.2017.09.003 järvenoja, h., järvelä, s., & malmberg, j. (2017). supporting groups’ emotion and motivation regulation during collaborative learning. learning and instruction, 70, 101090 https://doi.org/10.1016/j.learninstruc.2017.11.004 järvenoja, h., näykki, p., & törmänen, t. (2019). emotional regulation in collaborative learning: when do higher education students activate group level regulation in the face of challenges? studies in higher education, 44(10), 1747–1757. https://doi.org/10.1080/03075079.2019.1665318 jones, c., volet, s., & pino-pasternak, d. (2021). observational research in face-to-face small groupwork: capturing affect as socio-dynamic interpersonal phenomena. small group research, 52(3), 341-376. https://doi.org/10.1177/104649642098592 kauffeld, s., & lehmann-willenbrock, n. (2012). meetings matter: effects of team meetings on team and organizational success. small group research, 43(2), 128-156. https://doi.org/10.1177/1046496411429599 knight, a. p., & eisenkraft, n. (2015). positive is usually good, negative is not always bad: the effects of group affect on social integration and task performance. journal of applied psychology, 100(4), 1214–1227. https://doi.org/10.1037/apl0000006 kreijns, k., kirschner, p. a., & jochems, w. (2003). identifying the pitfalls for social interaction in computer-supported collaborative learning environments: a review of the research. computers in human behavior, 19(3), 335-353. https://doi:10.1016/s0747-5632(02)00057-2 jones, volet, pino-pasternak & heinimäki 69 | f l r kuppens, p. (2015). it’s about time: a special section on affect dynamics. emotion review, 7(4), 297-300. https://doi.org/10.1177/1754073915590947 kuppens, p., tuerlinckx, f., yik, m., koval, p., coosemans, j., zeng, k. j., & russell, j. a. (2017). the relation between valence and arousal in subjective experience varies with personality and culture. journal of personality 85(4), 530-542. https://doi.org/10.1111/jopy.12258 kwon, k., liu, y-h., & johnson, l. p. (2014). group regulation and social-emotional interactions observed in computer supported collaborative learning: comparison between good vs. poor collaborators. computers & education, 78, 185-200. https://doi.org/10.1016/j.compedu.2014.06.004 langer-osuna, j. m., gargroetzi, e., munson, j., & chavez, r. (2020). exploring the role of off-task activity on students’ collaborative dynamics. journal of educational psychology, 112(3), 514532. http://dx.doi.org/10.1037/edu0000464 lehmann-willenbrock, n., allen, j. a., & meinecke, a. l. (2014). observing culture: differences in u.s.-american and german team meeting behaviors. group processes & intergroup relations,17(2) 252-271. https://doi.org/10.1177/1368430213497066 lehmann-willenbrock, n., & chiu, m. m. (2018). igniting and resolving content disagreements during team interactions: a statistical discourse analysis of team dynamics at work. journal of organizational behavior, 39(9), 1142-1162. https://doi.org/10.1002/job.2256 lobczowski, n. g., lyons, k., greene, j. a., & mclaughlin, j. e. (2021). socioemotional regulation strategies in a project-based learning environment. contemporary educational psychology, 65, 101968. https://doi.org/10.1016/j.cedpsych.2021.101968 mesquita, b., & boiger, m. (2014). emotions in context: a sociodynamic model of emotions. emotion review, 6(4), 298-302. https://doi.org/10.1177/1754073914534480 mills, a. j., durepos, g., & wiebe, e. (eds.). (2012). comparative case study. in encyclopedia of case study research (pp. 175-176). sage publications. https://doi.org/10.4135/9781412957397 näykki, p., järvelä, s., kirschner, p. a., & järvenoja, h. (2014). socio-emotional conflict in collaborative learning: a process-oriented case study in a higher education context. international journal of educational research, 68, 1-14. https://doi.org/10.1016/j.ijer.2014.07.001 oțoiu, c., rațiu, l., & rus, c. l. (2019). rivals when we work together: team rivalry effects on performance in collaborative learning groups. administrative sciences, 9(3), 61. https://doi.org/10.3390/admsci9030061 poupore, g. (2018). a complex systems investigation of group work dynamics in l2 interactive tasks. the modern language journal, 102(2), 350-370. https://doi.org/10.1111/modl.12467 riebe, l., girardi, a., & whitsed, c. (2016). a systematic literature review of teamwork pedagogy in higher education. small group research, 47(2), 619-664. https://doi.org/10.1177/1046496416665221 rogat, t. k., and linnenbrink-garcia, l. (2011). socially shared regulation in collaborative groups: an analysis of the interplay between quality of social regulation and group processes. cognition and instruction, 29(4), 375-415. https://doi.org/10.1080/07370008.2011.607930 rogat, t. k., & linnenbrink-garcia, l. (2013). understanding the quality variation of socially shared regulation: a focus on methodology. in m. vauras & s. volet (eds.), interpersonal regulation of learning and motivation: methodological advances (pp. 102-125). routledge. scherer, k. r. (2005). what are emotions? and how can they be measured? social science information, 44(4), 695-729. https://doi.org/10.1177/0539018405058216 jones, volet, pino-pasternak & heinimäki 70 | f l r sherin, m. g. (2004). new perspectives on the role of video in teacher education. in j. brophy (ed.), advances in research on teaching: vol. 10. using video in teacher education (pp. 1–27). elsevier. sheets-johnstone, m. (2009). animation: the fundamental, essential, and properly descriptive concept. continental philosophy review, 42 (3), 375-400. https://doi.org/10.1007/s11007009-9109-x slaby, j. (2016). relational affect. working paper sfb 1171 affective societies 02/16. http://edocs.fuberlin.de/docs/receive/fudocs_series_000000000562 sohr, e. r., gupta a., & elby, a. (2018). taking an escape hatch: managing tension in group discourse. science education, 102(5), 883-916. https://doi.org/10.1002/sce.21448 summers, m., & volet, s. (2010). group work does not necessarily equal collaborative learning: evidence from observations and self-reports. european journal of psychology of education, 25(4), 473–492. https://doi.org/10.1007/s10212-010-0026-5 ucan, s., & webb, m. (2015). social regulation of learning during collaborative inquiry learning in science: how does it emerge and what are its functions? international journal of science education, 37(15), 2503-2532. https://doi.org/10.1080/09500693.2015.1083634 ujitani e, & volet, s. socio-emotional challenges in international education: insight into reciprocal understanding and intercultural relational development. journal of research in international education, 7(3), 279-303. https://doi.org/10.1177/1475240908099975 van kleef, g. a. (2021). comment: moving (further) beyond private experience: on the radicalization of the social approach to emotions and the emancipation of verbal emotional expressions. emotion review, 13(2), 90-94. https://doi.org/10.1177/1754073921991231 vriesema, c. c., & mccaslin, m. (2020). experience and meaning in small-group contexts: fusing observational and self-report data to capture self and other dynamics. frontline learning research, 8(3), 126-139. https://doi.org/10.14786/flr.v8i3.493 vygotsky, l. s. (1978). mind in society: the development of higher psychological processes. harvard university press. yin. r. k. (2018). case study research and applications: design and methods (6th ed.). sage publications inc. yoerger, m., allen. j. a., & crowe. j. (2018). the impact of premeeting talk on group performance. small group research, 49(2) 226-258. https://doi.org/10.1177/1046496417744883 jones, volet, pino-pasternak & heinimäki 71 | f l r appendix a group task and activities in five labs jones, volet, pino-pasternak & heinimäki 72 | f l r appendix b coding scheme for individual-level interpersonal affect behaviours in groupwork jones, volet, pino-pasternak & heinimäki 73 | f l r appendix c coding scheme for individual-level interpersonal affect behaviours in groupwork with data examples behavioural category and code example off-task positive small talk “how does everyone else feel? good?”; “do you have a cold?”; “no i’ve got allergies”; “running late, were you?”; “yeah, i missed the train” humour and joking “it’s like a star wars punishment?”; “you’re gonna be a quality dad, you’ve already got your dad jokes ready!” laughter laughter that is related to the social chat off-task negative small talk negative “i’m just tired. i’m sick. we’re all sick”; “yeah, don’t become teachers you’ll be sick”; “i swear it’s getting worse as the day goes on” side-talk “we’ll have to ask [name] how his surgery went…”; “he seems like a nice guy” [looking at other/s phone] using mobile phone using phone for personal purposes: scrolling; texting; talking on phone on-task positive inclusiveness “now, [name], you can do the honours it you want”; “i think we’ve got this! all over it!”; “everyone else agree with that?” offering praise or support “ooh, look at his prep skills, it’s immaculate!”; “that’s so good [name]!”; “do you want help?”; “oh yeah, good point!” showing enthusiasm, interest “i’m still blasting rockets in the air, it’s still cool!”; “this should be interesting. listen, listen! it’s sizzling!” “that’s awesome! i wonder why it’s flashing like that…” lightening the atmosphere “the hot air from my mouth could keep it up in the air!”; “yeah, we just don’t have the power captain” “oh, she might have madness to her reasoning” laughter laughter that is related to the task focus on-task negative complaining; negative expressions “i’m bored”; “i can’t be bothered!”; “i’m starting to hate this experiment” criticising/running someone down “…your handwriting’s driving me insane”; “no, it does! she’s wrong!”; “we spend half an hour organising what we’re gonna do abrupt, curt or rude behaviour; interrupting to over-rule; ignoring “shut up!”; “noooo, hang on!”; “you’re reading the wrong one!” “oh no. it’s a circuit love. not just to play…there’s a certain way to connect things!” disrupting “you calm down!” [said jokingly to member reading aloud conceptual question, stopping task discussion from proceeding]; “are you going to your class today?” [spoken as member is science explaining, disrupting conceptual reasoning] jones, volet, pino-pasternak & heinimäki 74 | f l r splitting group “well do we want to split in two’s? two make the oobleck and two make the psylli slime?”; “we’ll continue doing this if you guys want to do that” non-affective task interaction [non-affective] “what did you write?”; “okay put green in the middle there and then we create the series circuit”; “it doesn’t do anything. it’s an insulator. it doesn’t conduct” empty talk “what was i gonna say?”; “okay i need to write my name…” jones, volet, pino-pasternak & heinimäki 75 | f l r appendix d breakdown of interpersonal affect behavioural codes by frequency (and %) over time in five labs codepen van halem et al frontline learning research vol.8 no. 3 special issue (2020) 140 163 issn 2295-3159 tracking patterns in self-regulated learning using students’ self-reports and online trace data nicolette van halema, chris van klaverena, hendrik drachsleb, c, marcel schmitzd & ilja cornelisz a avrije universiteit amsterdam, the netherlands bopen universiteit, the netherlands cdipf | leibniz institute for research and information in education, frankfurt, germany dhogeschool zuyd, the netherlands article received 29 may 2019/ revised 26 july / accepted 17 september/ available online 30 march abstract for decades, self-report instruments – which rely heavily on students’ perceptions and beliefs – have been the dominant way of measuring motivation and strategy use. event-based measures based on online trace data arguably has the potential to remove analytical restrictions of self-report measures. the purpose of this study is therefore to triangulate constructs suggested in theory and measured using self-reported data with revealed online traces of learning behaviour. the results show that online trace data of learning behaviour are complementary to self-reports, as they explained a unique proportion of variance in student academic performance. the results also reveal that self-reports explain more variance in online learning behaviour of prior weeks than variance in learning behaviour in succeeding weeks. student motivation is, however, to a lesser extent captured with online trace data, likely because of its covert nature. in that respect, it is of importance to recognize the crucial role of self-reports in capturing student learning holistically. this manuscript is ‘frontline’ in the sense that event-based measurement methodologies with online trace data are relatively unexplored. the comparison with self-report data made in this manuscript sheds new light on the added values of innovative and traditional methods of measuring motivation and strategy use. keywords: self-regulated learning; self-report measures; event-based measures; online trace data info corresponding author email n.van.halem@vu.nl doi 10.14786/flr.v8i3.497 1. introduction motivation and strategy use are core concepts in the literature on learning and instruction. widely known and prominent in contemporary educational psychology is the theory of self-regulated learning (srl), which integrates these constructs in explaining student success. srl can be defined as “an active, constructive process of goal setting and attempting to monitor, regulate, and control cognition, motivation, and behaviour, guided and constrained by goals and the contextual features in the environment” (dinsmore et al., 2008; jupp, 2006; panadero et al., 2016; panadero, 2017; pintrich, 2000, p. 453). srl is an internal process that we cannot directly access, such that proxies are necessary to assess this srl process (boekaerts & corno, 2005). for decades, it has been argued that aptitude-based self-report instruments – which rely heavily on students’ perceptions and beliefs – do not fully capture srl. theories of srl emphasize that each instance of self-regulation is a function of the individual’s dynamic interaction with the learning environment, but few instruments satisfactorily capture such data (boekaerts et al., 2000; efklides, 2011; veenman, 2011; winne & perry 2000). yet, self-reports have remained the dominant way of measuring srl (boekaerts & corno 2005; winne and perry 2000), as the implementation of more time-intensive data collection methods, such as thinking-aloud protocols, event-based self-reports, or observations, are often times not feasible in educational settings. the recent introduction of tracing methods in online learning environments mainly through learning analytics (greller & drachsler, 2012) sparked the development of an alternative event-based measurement method of srl, enabling a form of online observation methods, while influencing the learning process as little as possible (panadero et al., 2016). these measurement methods are spurred by active efforts in the learning analytics community to bridge the gap between learning sciences and data analytics. however, so far learning theories such as srl are seldom used as theoretical basis for the design and evaluation of tracing methods (jivet et al., 2017; jivet et al., 2018). furthermore, there is a dearth of empirical work into the potential of online observation methods to complement self-reports on srl. the purpose of this study is therefore to triangulate constructs suggested in srl theory and measured using aptitude-based self-reported data with revealed online traces of learning behaviour. this study takes place in the context of the implementation of a learning analytics application, called ‘the learning analytics experiment’ in dutch higher education. the main aim of the project was to create an opportunity for institutions, teachers, and students to gain experience with different facets of learning analytics (e.g. privacy, feedback provision, insight in the use of learning materials, etc.). the online traces of students’ learning behaviour were recorded using xapi (or tin can) trackers, which have the potential to track learning experiences and store records of learners’ (e.g. ‘access video’ or ‘receive grade’). the trackers used in this study captured information specifically on the use of learning materials, such that the use of the online learning environment can be compared between students over time in light of different components of srl. this study describes an implementation of the xapi trackers in a mandatory first-year statistics course at a dutch university, during two consecutive academic years. during the implementation, self-reports on motivation and strategy use were collected to triangulate the trace data. this offered a rich case to unpack the aggregated data collected with self-reports and to, vice versa, colour the trace data with self-reports on motivation and strategy use. accordingly, this study aims to shed light on two of the central questions addressed in this special issue: in what ways do self-report instruments reflect the conceptualizations of the constructs suggested in theory related to motivation or strategy use? and: how does the use of self-report constrain the analytical choices made with that self-report data? 2. theoretical framework 2.1 complementarities of self-reported srl measures and online trace data online learning environments are central in today’s higher education, as they not only form a learning portal for a variety of purposefully selected learning resources, but also help to navigate through the course, enable students to be in contact with the instructor and peers, and to engage in various learning activities in a student-led fashion (lust et al., 2012; molenda, 2008). students leave traces when interacting with these online learning environments. trace data can be defined as “observable evidence of particular cognitions that are obtained at points where a cognitive process is applied while completing a task” (howard-rose & winne, 1993, p. 594). a growing body of literature confirms the importance for online learning environments for learning outcomes. gašević et al., 2014), for example, showed that the number of logins, number of operations performed on discussion forums and resources accounted for approximately 21% of the variance in academic performance. this finding is in line with earlier research on the relation between the frequency of visits to an online learning environments and students’ academic performance (coogan et al., 2005; deneui & dodge, 2006; wang & newlin, 2000). models on srl provide a holistic theoretical foundation for the relation between observed student behaviour in online learning environments and academic performance, based on the cognitive, metacognitive, behavioural, motivational, and affective aspects of learning (panadero, 2017). one of the latest meta-analyses on the effect of srl (sitzmann and ely, 2011) shows that the four biggest predictors of learning gains — goal level, persistence, effort, and self-efficacy — have a significant motivational value. longitudinal investigation of the interaction between motivation, strategy use, and the learning environment is, however, scarce (panadero, 2017). there are several reasons to believe that online trace data hold potential for tapping into the process of srl. students (especially in higher education) are the agents in online environment usage: they determine which resources are used and how these resources are used. the effects of the instructional design of a learner’s environment are therefore never deterministic, since the use of the environment depends on the personal goals, motivation, and volition of the student (winne & baker, 2013). usage of online environment can be considered a skill in itself, as it requires a repertoire of learning strategies, confidence, and competences. as lust and colleagues (2012) put it, online environments are only beneficial when learners recognize the learning resources as a learning opportunity for which they are motivated to spent effort and time on. in other words, effective use of online learning environments can be conceptualized as a manifestation of srl. the extent to which an individual student is motivated and able to self-regulate their learning process is thus a prerequisite for effective tool-use (winne & baker, 2013; lust et al., 2012). in addition, and elaborately described by fryer (2017), the relation between motivation, strategy use, and the use of learning environments can be conceptualized as reciprocal. namely, learning activities undertaken in the process phase of learning, as well as the resulting product of learning, feed back to students’ beliefs, attitudes, and ideas around motivation, strategy use, and self-regulation that play a role in the presage stage of learning. online trace data, thus, reflects the dynamic relation between students and their learning experiences over time. 2.2 removing analytical restrictions of inventories on motivation and strategy use with online trace data an event-based measure of srl with online trace data arguably has the potential to remove analytical restrictions of aptitude measures. firstly, and according to the overconfidence effect, unrealistically favourable attitudes that people have towards themselves (taylor & brown, 1988) can impose an upward bias in self-reports of motivation and strategy use. in other words, students might apply less and less effective study strategies than self-reported. zhou and winne (2012) confirm this theory empirically by showing that “trace data”-based measures of student achievement goal orientation were much stronger associated with learning outcomes than with self-reported ones. this is particularly pressing since existing research suggests that learners tend to use ineffective learning strategies (jamieson-noel & winne, 2003), and do not make effective use of available resources to optimize their learning, even in those environments that build on effective learning designs (ellis et al. 2005; lust et al., 2013). as discussed in this special issue, potential reasons are students’ cognitive processing capacity (chauliac et al. 2020), exerted effort (iaconelli & wolters 2020), and student characteristics (vriesema 2020). comparing self-reported data with online traces of learning behaviour taps right into a potential upward bias in self-reports on motivation and strategy use. since online traces of learning behaviour do not suffer from socially desirable responding bias, it is possible that they provide a better approximation of motivation and strategy use. secondly, inventories are restricted as they often predetermine the level of aggregation in the analysis of data on motivation and strategy use (e.g. fixed at student-level), which, as a result, generates student-focussed or aptitude-based measures. this jeopardizes the potential of adapting effectively to individual needs and preferences during a study episode or educational program. alongside self-report instruments that operationalized motivation and self-regulation as event-bound (winne, 2010), event-based trace data of learning behaviour can provide a dynamic insight in how motivation and strategy use does not only vary between students, but also within students over time. this is particularly relevant since previous research show that self-motivational beliefs and strategy use can vary considerably over the course of a study episode, or throughout the educational program (boekaerts et al., 2000; efklides, 2011; veenman, 2011; winne & perry 2000). one example is presented in this special issue by moeller et al. 2020, in their study on this intra-individual variation in self-motivational beliefs. 2.3 status of research on srl combining self-reports and online trace data up until now, empirical studies on srl that use online trace data often aim at identifying different patterns of usage behaviour. this has led to a wide variety of student typologies aiming at a better understanding of the type of learning strategies used by students, looking at the use of information-tools (such as forums, instruction video’s, and interactive mind maps) over the duration of a course (heffner & cohen, 2005; hoskins & van hooff, 2005; huon et al, 2007). examples of typologies are the active and passive users of forums (hoskins & van hooff, 2005); the early-, constant-, and late users (knight, 2010; lust et al., 2012); the low-, average-, and high frequency users (bera & liu, 2006; jiang et al. 2009). research methods in identifying learning strategies inferred from online trace data evolve rapidly. a recent strand of research on trace data adopts data mining techniques in detecting striking, previously unknown, patterns in study tactics and strategy use (han et al., 2011). in general, it remains a challenge to qualify these patterns, typologies, and clusters of study tactics, and to make these insights actionable for instructional design accordingly. so far, studies repeatedly find that usage patterns do explain variance in academic performance (e.g. cho & yoo, 2017; cornelisz & van klaveren, 2018; gašević et al., 2014; han et al., 2011; schmitz et al., 2018), but what type of self-regulation it represents and the quality thereof remains a black box. only a handful of studies focussed on triangulation of self-reports specifically on srl using online trace data of student learning. the few studies that describe a quantitative comparison of self-reports on srl and online trace data show that the relation between these two sources of data is not straightforward. hadwin et al. (2007) used ten relevant items from the motivated strategies for learning questionnaire (mslq) on strategy use to explain students’ learning in gstudy, a web-based learning environment in which students read texts, summarize, and use concept maps in an introductory educational psychology course (n = 188). they clustered students into four groups based on self-reported srl data and found their actual learning patterns in gstudy differed substantially, even among students from the same cluster. kim et al. (2018) made a similar comparison among undergraduates in korea. they analysed online trace data from 284 undergraduate students enrolled in an asynchronous online statistics course. based on self-reports collected with the mslq, students were classified as fully, partially, or not self-regulating. surprisingly, this distinction did not reveal different patterns in online traces of learning. the main difference between the groups was timing in study behaviour, students classified as not self-regulating studied mainly shortly before the examination, which was negatively related to academic performance. a study of guerra et al. (2016) used the achievement-goal questionnaire and online trace data (n = 89). their study shows that students who report a high mastery-approach show a higher level of activity in the online learning environment. there results also suggest that highly motivated students are more sequential in their patterns of navigation, which means students are less likely to follow the suggested order of topics to study. cho & yoo (2017) adopted a different approach and compared precision in prediction of students’ academic performance based on patterns in self-reports based on mslq and patterns in online trace data established with data mining techniques (n = 60). the model based on online trace data provided a more precise prediction. the authors note that this might be partially due to the fact that the trace data provided more variables and was based on a bigger data set. also, the study did not address whether or not it is likely that the self-reports and online trace data actually measured the same constructs. like in the studies described earlier, the inventory is aptitude-based (muis et al., 2007; zimmerman, 2008), whereas the online trace data is event-based. in this special issue, rogiers, merchie, and van keer present one of the first studies in this space that compares event-based self-report data with trace data. overall, the interpretation of online trace data in light of self-reports is not clear-cut, there are many questions left unanswered about the relationships between the constructs measured with an instrument such as the mslq and online trace data collected in different education settings. 3. research questions and hypotheses the aim of this study is to further explore the relationship between measures of srl through self-reports and online trace data, by scrutinizing variance in self-reports and online trace data between and within students. because the online traces of learning behaviour are event-based, our study provides a dynamic insight in how motivation and strategy use vary between and within students over time. the findings are instrumental in guiding innovations in education towards effective personalized learning. accordingly, the following research questions are formulated for this study: 1. to what extent do self-reports explain variance in online trace data of learning behaviour? 2. how stable is the relation between self-reports and online trace data throughout the various weeks of the course? 3. how well do self-reports explain student performance in comparison to online trace data? following fryer (2017), we expect that self-reports on motivation, strategy use, and self-regulation gauge between-student differences in the presage stage and that online trace data gauges the process phase of learning. with respect to the first research question, it can, therefore, be expected that self-reports explain substantial variance in learning behaviour observed through online trace data. furthermore, student differences in the presage stage will likely be affected over time by the feedback loop between the presage, process, and product phase. therefore, it can be expected that the self-reports predict online learning behaviour most precisely at the time the self-reports are administered. with respect to research question two, thus, we expect that the relationship between self-reports and online trace data varies over time. finally, in light of the third research question, we expect online trace data to explain an equal amount or more variation in student performance than self-reports, in line with the findings of cho and yoo (2017). this study adds to the existing literature as follows. firstly, there are only a few studies so far that tapped into the relationship between self-reports and online trace data. given the sensitivity of the usage patterns to the instructional design of an online learning environment (gašević et al., 2016), it is of great importance that a broad range of educational settings are explored with online trace data. the course, instructional design, and educational setting investigated in this study is considered particularly relevant because it is highly representative for other courses in mainstream dutch higher education and since this particular course plays a crucial role in all social science programs in the netherlands. in addition, this study compares multiple cohorts, which yields insight in continuity and change of strategy use with the constant evolvement of course design. secondly, the few studies that have been comparing self-reports and online trace data so far dealt with relatively small sample sizes (cho & yoo, 2017; guerra et al., 2016; hadwin et al., 2007; kim et al., 2018) or did not measure the full range of self-motivational beliefs and strategies with self-reports (hadwin et al., 2007; cho & yoo, 2017). as a result, there is no clarity, yet, on the relevance and actionability of online trace data in comparison to inventories on motivation, strategy use, and self-regulation. 4. methodology 4.1 participants the data used consist of self-reports and detailed log-data on srl, collected among two cohorts of first year students at the faculty of behavioural and movement sciences during a mandatory statistics course that took place between october and december 2016 (n = 435; 94.44% female, mage = 20.60 years, sd = 5.18) and 2017 (n = 489; 78.90% female, mage = 20.88 years, sd = 3.48). the 2017 cohort included international students, as this was the first year that this course was also taught in english. the faculty of behavioural and movement sciences offers the following educational programs: movement sciences, education sciences and behavioural, developmental and clinical psychology. 4.2 course design 4.2.1 online learning environment the online learning environment was available to all students throughout the course and its usage was not mandatory in any sense. in 2016, the online learning environment consisted of a learning management system (blackboard) and a separate online learning tool, called i hate statistics . blackboard is one of the leading commercial lms software packages used by north american and european universities (itmazi & megias, 2005; munoz & duzer, 2005). from the start of the academic year 2017-2018, the institution switched to a new learning management system, called canvas. together with the standard features of lmss, canvas provides advanced options like learning outcomes, peer review, migration tools, e-portfolios, screen sharing and video chat etc. canvas is gaining popularity, hundreds of colleges, universities, and school districts currently use this package (www.instructure.com). across cohorts, the learning management system was structured in a similar fashion. both blackboard and canvas provided students with three types of tools: 1) an information tool (lecture slides, instruction video’s, and general course information), which provided the course content in a structured way; 2) a cognitive tool for self-assessments that enabled students to interact with the subject matter, to assess and to reflect on the learning content; and 3) a communication tool (forums) that enabled students to communicate with peers and instructors (lust et al., 2012). the tools were structured based on the week-topics of the course and were available at all times throughout the course. the learning management system referred students to ‘i hate statistics’, which provided students with an online environment for practicing and studying, where students could engage in lessons, challenges, and self-assessments that were related to the week topics, or to other topics that were available in this environment. each challenge consisted of on average maximal ten questions; yet, the challenge was automatically finished when students correctly answered five questions within the challenge. a unique feature of the environment was that it is built around the statistics course and offered content in a particular week that was similar to the content offered in the lectures and seminars. the xapi tracking method is applicable for all online learning environments (see for example https://xapi.com/). in the context of this study, it was used to gain insight into the use of learning materials in the learning management systems. the teacher created so-called 'recipes' and placed them in blackboard, where the use of information resources and participation in online activities were tracked. these recipes ensured that the desired statistics were generated, as well as information on the type of tool that was used (e.g. slides, challenges, lessons) an action verb (e.g. accessed, received), and a label (e.g. lessons on the chi-square test, lecture slides of week 1). a comparable tracking method was used in i hate statistics, which generated similar data about the use of the learning materials. 4.2.2 instructional design during the eight-week introductory statistics course, students had to attend two lectures each week in which theoretical concepts were addressed. additionally, they had to attend one seminar each week with mandatory attendance in which the assignments and the subject matter were discussed and opportunities for peerand teacher feedback were organized. offline and in the course manual, students were referred to the textbook that the teacher selected as a starting point for the course. use of this textbook is not traced within the online learning environment. the online learning environment provided opportunities for self-assessments. the self-assessments in the learning management systems and i hate statistics were similar in nature and contained multiple-choice questions in which knowledge and comprehension were assessed. the learning management system contained four self-assessments; i hate statistics contained eight self-assessments on the course topics. in the second year of this study, it also gave access to a practice exam, with questions that were representative for the final exam. at the end of the course (week 8) each student was graded based on a final multiple-choice exam and on a research report. 4.3 instruments 4.3.1 self-reports the motivated strategies for learning questionnaire (mslq), a measure developed by pintrich and colleagues (mckeachie et al 1985; pintrich, 1991; pintrich et al., 1987; see also duncan and mckeachie, 2005, for a more in-depth discussion), was used to assess self-reports on srl. the mslq was derived from an extensive body of literature and was one of the first inventory on the quality of student learning that not only included attitudes, motivation, and strategy use, but also self-regulatory strategies (entwistle & mccune, 2004). the mslq was progressive at the time by including the dimension of students’ consciousness about the teaching-learning environment, leading to adaptation of ways to tackle academic work (entwistle & mccune, 2004). several studies argue that there is piecemeal evidence for the scale structure of the mslq (hilpert et al., 2017), yet, to this date, this inventory for students in higher education is still considered relevant in light of the wide range of motivation, affect, strategy use, and self-regulation it covers. appendix a provides a description of the scales, along with a couple of sample items per scale. the mslq was administered once in the seminar of week four, along with several general questions about background variables, such as age and gender. the mslq contained 81 questions of which 31 items determined students’ motivational orientation towards the course and 50 items assessed metacognition. the questions assess the propensity of students to engage in self-regulated learning within the specific context of this course, but, overall, the mslq has been classified as an aptitude measure of self-regulated learning (muis et al., 2007; zimmerman, 2008). students answered with a 7-points likert scale, ranging from ‘not at all true for me’ to ‘very true for me’. the motivational orientation is divided into six subscales: intrinsic goal orientation, extrinsic goal orientation, task value, self-efficacy, control beliefs and test anxiety. metacognition was scored on nine subscales: rehearsal, elaboration, organization, critical thinking, metacognitive self regulation, time and study environment, effort regulation, peer learning and help seeking. a definition per subscale is provided in appendix a. for a complete description of the mslq and each of its subscales we refer to the manual of the msql (pintrich, 1991). reliability coefficients were determined with the cronbach’s alpha (see appendix a). it is important to note that the reliability of the goal orientation sub-scales is poor. scrutinizing the data per item did not point out particular weak items that could be removed from the scale to improve reliability. 4.3.2 online trace data online trace data was obtained as part of the project ‘the learning analytics experiment’. the teacher of the course was actively involved in defining what type of data was collected. the teacher was facilitated to track any learning activity in the learning management system with trackers designed by surfnet and the amsterdam center for learning analytics (acla). the trackers were based on xapi recipes (a set of rules). the teacher defined the set of rules, based on activities, verbs, and labels. for example, ‘formative exam’ (activity) ‘accessed’ (verb) ‘week 3’ (label). after defining a recipe, a html-code was provided that accordingly was embedded in the online learning environment, often in the form of an empty object in the environment. in addition, the designers of the application i hate statistics provided access to the data they collected on each learning activity a student engaged in, as well as the timestamp, the length of the learning activity, the number of questions that a student answered within a lesson or a challenge, and the success rate within a lesson or challenge was logged for each student. after processing the data, variables were selected to be included in the current study, following the work of guerra and colleagues (2016), kim and colleagues (2018), and theobold and colleagues (2018), these variables are described in table 1. the countand timebased variables were aggregated to the week level based on the week-based structure of the course; each week was structured around a particular topic. the variables sequential navigation and distributed learning provide a metric (respectively a ratio and a count measure) on student level, in order to deduce an overall measure of students’ patterns in use of the online learning environment throughout the eight weeks of the course. table 1 interpretation of online trace data in weekand student level variables 4.3.3 academic performance students’ academic performance is measured with the summative course evaluation. at the end of the course students’ final grade was determined based on the results of a multiple-choice exam and the grading of the research report. the exam included 30 four-answer choice questions; ten questions targeted knowledge, ten targeted insight, and the other ten targeted calculations. appendix b provides sample questions of each category. the exam was designed by the course coordinator. in general, all course coordinators are equipped with a training that covers basic knowledge and skills on constructing multiple choice tests. the course coordinator had, furthermore, access to a data base with high-quality questions used in previous years, which – albeit slightly modified – could be used to put together the exam. the exam underwent peer review and psychometric tests, the latter was used in the grading process. final grades were scored on a scale from 1 to 10 with 10 being the highest and with 5.5 as passing threshold. 4.4 procedure students were provided with an informed consent form for both the self-report data and the online trace data. in the 2016 cohort, informed consent for the online trace data was obtained at the beginning of the course, whereas informed consent for the self-report data was obtained during the student survey in week 4 of the course. in the 2017 cohort, students were provided with one informed consent form including permission to use both sources of data at the beginning of the course. 4.5 analysis strategy data preparation procedures were applied prior to the analysis. homogeneity of the data across cohorts was established for motivation and strategy use reported by students, for the variables based on the online trace data, and for academic performance, using levene’s test for homogeneity of variance and independent-sample t-tests. multiple imputation methods were applied to deal with missing data on items of the mslq scales in the following analyses, using the markov chain monte carlo method with a number of 10 iterations, with ibm spss statistics 25. to answer the first research question, the extent to which student self-reports tap motivation and strategy use as reflected in the online trace data online learning behaviour was examined. to that end, explained variance in online trace data of learning behaviour by student self-reports was investigated with ordinary least square regression analyses. total participation in learning activities was used as dependent variable, because it identified between-student and between-week variance and comprised student engagement in both the learning management system and the online practicing environment. after looking at the data on the course level, models were tested on a week level, in order to gauge variability in the relation between online trace data and self-reported data based on the mslq and answer the second research question. to answer the third research question, the relation between self-reports and online trace data was gauged with a third proxy of student motivation and strategy use: student academic performance based on students’ grade on the multiple-choice exam at the end of the course. the proportion explained variance in academic performance by online trace data and self-reports was identified with ordinary least square analyses, separately and combined. in all analyses, students’ gender and age were included as control variables. 5. results 5.1. descriptive statistics in total, 605 students gave active consent to include their data collected during the course for these students could be used in the analyses. the results of the homogeneity tests and descriptive statistics are shown in table 2. both the self-reports and the online trace data differ significantly between the cohorts, likely related to the introduction of the practice exam for the 2017 cohort and the demographical shift based on the inclusion of international students in 2017. for example, the level of extrinsic motivation is significantly higher in the second cohort, as well as the average number of self-assessments per student per week. even though students in the 2017 cohort engaged significantly more in self-assessments, their participation in online learning activities was on average significantly lower and they engaged in online learning activities in six out of eight weeks versus seven in 2016. table 2 tests of homogeneity * α =.05; ** α =.01; *** α =.001 5.3 the relationship between self-reports and online traces of motivation and strategy use over time the variables obtained through self-reports were regressed on online trace-data, both on studentand week-level, using standardized values, revealing how stable relationship between self-reports and online trace data is over time. student controls (age, gender, and cohort) were included in each of the models. the results are presented in table 3. a comparison of the adjusted r2 of the baseline model and the model of the sum of participation over the whole course reveals that students’ self-reports on motivation and strategy use explain about 9% of variance in the resources accessed in the online learning environment. the adjusted r2 of the models across weeks indicate that students’ self-reports seem to explain more variance in participation in the weeks prior to the collection of the self-reports (week 4) than participation in the weeks afterwards. the self-reports on time and study environment management explain a substantial proportion in variance in participation up and including week 4, but this changes in the weeks after the self-reports were collected. self-reports on rehearsal and elaboration strategies emerge as significant predictors of participation in week 3-6 and week 8. figure 1 shows the extent to which predictions of participation based on the online trace data of week 4 and the self-reports are representative for participation in other weeks. actual participation levels reveal that student participation fluctuates over time, with on average a peak in week 2 and week 8. table 3 self-reports regressed on online trace data * α =.05; ** α =.01; *** α =.001 figure 1. predicted participation based on the regression of students’ self-reports on total participation in week 4 and the actual participation across weeks 5.3 explained variance in academic performance by self-reports and online traces of motivation and strategy use the variables based on self-reports and online trace data were regressed on academic performance; using standardized values and in a stepwise fashion, see table 4. student controls (age, gender, and cohort) were included in each of the models. the differences in explained variance between model 1 and model 2 need to be interpreted with caution, as there are substantially fewer variables included in model 2. based on the explained variance of model 3, it can be concluded that there is a substantial unique contribution of the self-report and online trace data-based variables in predicting academic performance. in the final model, there are three scales of the mslq that explain a substantial proportion in academic performance; self-efficacy, elaboration strategies, and effort regulation. four variables based on the online trace data explain substantial variation; total time invested, total participation, self-assessment and distributed learning. the total time invested in the online learning environment is negative related to academic performance, even though the total amount of resources accessed by the student (i.e. total participation) is positively related to academic performance. table 4 self-reports and online trace data regressed on academic performance * α =.05; ** α =.01; *** α =.001 6. conclusion and discussion starting from the central questions of this special issue, this study aimed to provide insight in the ways in which self-reports reflect the conceptualizations of the constructs suggested in theory related to motivation or strategy use. to that end, this study looked into the relationship between measures of srl through aptitude-based self-reports and event-based online trace data. using event-based measurement methods of srl based on online trace data complementary to self-report instruments can be a first step in capturing self-regulation as a function of the individual’s dynamic interaction with the learning environment (boekaerts et al., 2000; efklides, 2011; veenman, 2011; winne & perry 2000). capturing the individual’s dynamic interaction with the learning environment can be instrumental in guiding innovations in education towards effective personalized learning and enabling educators to adapt to individual differences in srl during a study episode or educational program. the results of this study show that self-reports on motivation and strategy use explain 28% of variance in online study behaviour overall versus 18% in academic performance. none of the mslq scales on motivation significantly predicted online learning behaviour, whereas several mslq scales on strategy use did; namely the scales that measured rehearsal strategies, elaboration strategies, and time and study environment management strategies. given the prominent role of personal goals, efficacy, and interest in general models for student learning (fryer, 2017), it is striking that the mslq scales on motivation did not explain variance in online trace data. at the same time, these results are in line with previous research on self-reports of motivation in higher education that does show a strong relation with achievement (e.g. hattie, 2015) but not with study behaviour measured with online trace data (e.g. cho & heron, 2015; zhou & winne, 2012). this could suggest that the motivational aspects of self-regulatory processes are not captured with online trace data, because of their covert nature. in that respect, it is of importance to recognize the crucial role of self-reports in gaining broad insights in srl. at the same time, the results of this study underline the added value of online trace data of learning behaviour. firstly, online trace data of learning behaviour explained a unique proportion of variance in student academic performance. in this study, the number of logins, number of operations performed on discussion forums and resources accounted for approximately 18% of the variance in academic performance. although this is substantially less than the 21% found by gašević and colleagues (2014), it is equal to the amount of variance in academic performance explained by students’ self-reports on motivation and strategy use. as expected and along with the findings of the study of cho & yoo (2017), where online trace data produced an even more precise prediction of academic performance than students’ self-reports, this study provides a strong indication of the potential of online trace data to tap self-regulatory processes, in particular strategy use. the fact that online trace data still predict a unique proportion of variance in academic performance when studied in combination with self-reports might be explained by its potential to bypass a potential upward bias in self-reports of motivation and strategy use due to the overconfidence effect (taylor & brown, 1988) as they do not suffer from socially desirable responding bias. finally, this study shows that online study behaviour and the relation between self-reports and online trace data varies vastly from week to week. students’ self-reports seem to be mainly based on prior learning experiences within the course, as the reports did explain variation in online learning behaviour prior to the collection of self-reports, but substantially less thereafter. this finding is in line with the expectations around the second research question and fits the theoretical lens provided by fryer (2017) about the feedback loop between students and their learning experiences. in light of the second central questions of the special issue addressed in this study, we can conclude that the use of self-reports does constrain the analytical choices made with self-report data to some extent: the online trace data revealed large amount of within student variance. online trace data are therefore an important addition to self-reports in guiding innovations in education towards effective instructional design and personalized learning as they have the potential to give teachers and students real-time insights in strategy use and self-regulation. an avenue for future research is to combine event-based self-reports and online trace data in measuring srl. especially since the results of this study did not identify a relation between online trace data and self-motivational beliefs, potentially due to the weak reliability of some of the mslq scales. furthermore, the identified heterogeneity of the two cohorts in this study warrants further investigation. it is possible that the consent procedure in the first cohort did not yield a representative sample of the students in the first cohort with respect to their propensity to engage in self-regulated learning within the specific context of this course. future research is also needed to investigate how real-time measures of srl through online trace data and self-reports can inform teaching and learning. the lissa project is a good example where learning analytics were used to support the adviser-student dialogue, in a way that helped motivate students, triggered conversation, and provided tools to add personalization, depth, and nuance to the advising session (charleer et al., 2018). similar efforts based on online trace data or event-based self-reports are scarce, but pivotal to design and evaluate personalized interventions in online learning environments that promote effective study behaviour. keypoints self-report measures of strategy use such as management of time and study environment predict participation in online learning activities, although this relationship is not stable across weeks. student self-reports seem to explain more variance in online learning behaviour of prior weeks than variance in learning behaviour in succeeding weeks. self-report measures and online trace data on self-regulated learning are complementary in predicting study success as they both explain a unique proportion of variance in student academic performance. 8. acknowledgements funding: this work is part of the project ‘surfnet learning analytics hoger onderwijs’, supported by the national control unit educational research (nro) (project-id 405-17-851). references bera, s., & liu, m. (2006). cognitive tools, individual differences, and group processing as mediating factors in a hypermedia environment. computers in human behavior, 22, 295-319. http://doi.org/10.1016/j.chb.2004.05.001 boekaerts, m., & corno, l. (2005). self‐regulation in the classroom: a perspective on assessment and intervention. applied psychology, 54, 199-231. http://doi.org/10.1016/j.chb.2004.05.001 boekaerts, m., pintrich, p. r., & zeidner, m. (2000). self-regulation: an introductory overview. in handbook of self-regulation (pp. 1-9). academic press. charleer, s., moere, a. v., klerkx, j., verbert, k., & de laet, t. (2017). learning analytics dashboards to support adviser-student dialogue. ieee transactions on learning technologies, 11, 389-399. http://doi.org/10.1109/tlt.2017.2720670 cho, m. h., & heron, m. l. (2015). self-regulated learning: the role of motivation, emotion, and use of learning strategies in students’ learning experiences in a self-paced online mathematics course. distance education, 36, 80-99. http://doi.org/10.1080/01587919.2015.1019963 cho, m. h., & yoo, j. s. (2017). exploring online students’ self-regulated learning with self-reported surveys and log files: a data mining approach. interactive learning environments, 25, 970-982. http://doi.org/10.1080/10494820.2016.1232278 coogan, j., dancey, c. p., & attree, e. a. (2005). webct: a useful support tool for psychology undergraduates–a q methodological study. psychology learning and teaching, 5, 61-66. http://doi.org/10.2304/plat.2005.5.1.61 cornelisz, i., & van klaveren, c. (2018). student engagement with computerized practising: ability, task value, and difficulty perceptions. journal of computer assisted learning, 34(6), 828-842. http://doi.org/10.1111/jcal.12292 deneui, d. l., & dodge, t. l. (2006). asynchronous learning networks and student outcomes: the utility of online learning components in hybrid courses. journal of instructional psychology, 33, 256-260. dinsmore, d. l., alexander, p. a., & loughlin, s. m. (2008). focusing the conceptual lens on metacognition, self-regulation, and self-regulated learning. educational psychology review, 20, 391-409. http://doi.org/10.1007/s10648-008-9083-6 greller, w. & drachsler, h. (2012). translating learning into numbers: a generic framework for learning analytics. journal of educational technology & society, 15, 42-57. retrieved from http://www.jstor.org/stable/jeductechsoci.15.3.42 chauliac, m., catrysse, l., gijbels, d., & donche, v. (2020). it is all in the surv-eye: can eye tracking data shed light on the internal consistency in self-report questionnaires on cognitive processing strategies? frontline learning research, 8(3), 26–39. http://doi.org/10.14786/flr.v8i3.489 duncan, t. g., & mckeachie, w. j. (2005). the making of the motivated strategies for learning questionnaire. educational psychologist, 40, 117-128. http://doi.org/10.1207/s15326985ep4002_6 efklides, a. (2011). interactions of metacognition with motivation and affect in self-regulated learning: the masrl model. educational psychologist, 46, 6-25. http://doi.org/10.1080/00461520.2011.538645 ellis, r. a., marcus, g., & taylor, r. (2005). learning through inquiry: student difficulties with online course‐based material. journal of computer assisted learning, 21, 239-252. http://doi.org/10.1111/j.1365-2729.2005.00131.x. entwistle, n., & mccune, v. (2004). the conceptual bases of study strategy inventories. educational psychology review, 16, 325-345. http://doi.org/10.1007/s10648-004-0003-0 fryer, l. k. (2017). building bridges: seeking structure and direction for higher education motivated learning strategy models. educational psychology review, 29, 325-344. http://doi.org/10.1007/s10648-017-9405-7 gašević, d., dawson, s., rogers, t., & gašević, d. (2016). learning analytics should not promote one size fits all: the effects of instructional conditions in predicting academic success. the internet and higher education, 28, 68-84. http://doi.org/10.1016/j.iheduc.2015.10.002. gašević, d., kovanovic, v., joksimovic, s., & siemens, g. (2014). where is research on massive open online courses headed? a data analysis of the mooc research initiative. the international review of research in open and distributed learning , 15, 134-176. http://doi.org/10.19173/irrodl.v15i5.1954. guerra, j., hosseini, r., somyurek, s., & brusilovsky, p. (2016). an intelligent interface for learning content: combining an open learner model and social comparison to support self-regulated learning and engagement. in proceedings of the 21st international conference on intelligent user interfaces (pp. 152-163). acm. hadwin, a. f., nesbit, j. c., jamieson-noel, d., code, j., & winne, p. h. (2007). examining trace data to explore self-regulated learning. metacognition and learning, 2(2-3), 107-124. http://doi.org/10.1007/s11409-007-9016-7 han, j., pei, j., & kamber, m. (2011). data mining: concepts and techniques. elsevier. heffner, m., & cohen, s. h. (2005). evaluating student use of web-based course material. journal of instructional psychology, 32, 74-82. hilpert, j. c., stempien, j., van der hoeven kraft, k. j., & husman, j. (2013). evidence for the latent factor structure of the mslq: a new conceptualization of an established questionnaire. sage open, 3, 2158244013510305. hoskins, s. l., & van hooff, j. c. (2005). motivation and ability: which students use online learning and what influence does it have on their achievement? british journal of educational technology, 36, 177-192. http://doi.org/10.1111/j.1467-8535.2005.00451.x. howard-rose, d., & winne, p. h. (1993). measuring component and sets of cognitive processes in self-regulated learning. journal of educational psychology, 85, 591. http://doi.org/10.1111/j.1464-0597.2005.00205.x huon, g., spehar, b., adam, p., & rifkin, w. (2007). resource use and academic performance among first year psychology students. higher education, 53, 1-27. http://doi.org/10.1007/s10734-005-1727-6 iaconelli, r., & wolters, c. a. (2020). insufficient effort responding in surveys assessing self-regulated learning: nuisance or fatal flaw?frontline learning research, 8(3), 105–127. http://doi.org/10.14786/flr.v8i3.521 itmazi, j. a., & megías, m. g. (2005). survey: comparison and evaluation studies of learning content management systems. unpublished manuscript. jamieson-noel, d., & winne, p. h. (2003). comparing self-reports to traces of studying behavior as representations of students' studying and achievement. zeitschrift für pädagogische psychologie/german journal of educational psychology. http://doi.org/10.1024//1010-0652.17.34.159 jivet, i., scheffel, m., drachsler, h., & specht, m. (2017). awareness is not enough. pitfalls of learning analytics dashboards in the educational practice. in é. l., h. d., k. v., j. b., & m. p-s. (eds.), data driven approaches in digital education: 12th european conference on technology enhanced learning, ec-tel 2017, tallinn, estonia, september 12–15, 2017, proceedings (lecture notes in computer science (lncs); vol. 10474). cham: springer international publishing ag. http://doi.org/10.1007/978-3-319-66610-5_7 jivet, i., scheffel, m., specht, m., & drachsler, h. (2018). license to evaluate: preparing learning analytics dashboards for educational practice. in proceedings of the 8th international conference on learning analytics and knowledge (pp. 31-40). acm. http://doi.org/10.1145/3170358.3170421 jiang, l., elen, j., & clarebout, g. (2009). the relationships between learner variables, tool-usage behaviour and performance. computers in human behavior, 25, 501-509. http://doi.org/10.1016/j.chb.2008.12.010 jupp, v. (2006). the sage dictionary of social research methods. sage. kim, d., yoon, m., jo, i. h., & branch, r. m. (2018). learning analytics to support self-regulated learning in asynchronous online courses: a case study at a women's university in south korea. computers & education, 127, 233-251. http://doi.org/10.1016/j.compedu.2018.08.023. knight, j. (2010). distinguishing the learning approaches adopted by undergraduates in their use of online resources. active learning in higher education, 11, 67–76. http://doi.org/10.1177/1469787409355873 lust, g., collazo, n. a. j., elen, j., & clarebout, g. (2012). content management systems: enriched learning opportunities for all? computers in human behavior, 28, 795-808. http://doi.org/10.1016/j.chb.2011.12.009 lust, g., elen, j., & clarebout, g. (2013). regulation of tool-use within a blended course: student differences and performance effects. computers & education, 60, 385-395. http://doi.org/10.1016/j.compedu.2012.09.001 mckeachie, w. j., pintrich, p. r., & lin, y. g. (1985). teaching learning strategies. educational psychologist, 20, 153-160. http://doi.org/10.1207/s15326985ep2003_5 molenda, m. (2008). historical foundations. in m. j. spector, m. d. merrill, j. van merrienboer, & m. p. driscoll (eds.). handbook of research for educational communications and technology (pp. 5–20). routledge. moeller, j., dietrich, j., viljaranta, j., & kracke, b. (2020). disentangling objective characteristics of learning situations from subjective perceptions thereof, using an experience sampling method design. frontline learning research, 8(3), 63–85. http://doi.org/10.14786/flr.v8i3.529 muis, k. r., winne, p. h., & jamieson-noel, d. (2007). using a multitrait-multimethod analysis to examine conceptual similarities of three self-regulated learning inventories. british journal of educational psychology, 77, 177–195. http://doi.org/10.1348/000709905x90876 munoz, kd, & van duzer, j. (2005). blackboard vs. moodle: a comparison of satisfaction with online teaching and learning tools. humboldt state university. panadero, e. (2017). a review of self-regulated learning: six models and four directions for research. frontiers in psychology, 8, 422. http://doi.org/10.3389/fpsyg.2017.00422 panadero, e., klug, j., & järvelä, s. (2016). third wave of measurement in the self-regulated learning field: when measurement and intervention come hand in hand. scandinavian journal of educational research, 60, 723-735. http://doi.org/10.1080/00313831.2015.1066436 pintrich, p. r. (1991). a manual for the use of the motivated strategies for learning questionnaire (mslq). pintrich, p. r. (2000). multiple goals, multiple pathways: the role of goal orientation in learning and achievement. journal of educational psychology, 92, 544. http://doi.org/10.1037/0022-0663.92.3.544 pintrich, p. r., mckeachie, w. j., & lin, y. g. (1987). teaching a course in learning to learn. teaching of psychology, 14, 81-86. http://doi.org/10.1207/s15328023top1402_3 rogiers, a.; merchie, e., & van keer, h. (2020). opening the black box of students’ text-learning processes: a process mining perspective. frontline learning research, 8(3), 40–62. http://doi.org/10.14786/flr.v8i3.527 schmitz, m., scheffel, m., van limbeek, e., van halem, n., cornelisz, i., van klaveren, c., ... & drachsler, h. (2018). investigating the relationships between online activity, learning strategies and grades to create learning analytics-supported learning designs. in european conference on technology enhanced learning (pp. 311-325). springer, cham. taylor, s. e. & brown, j. d. (1988). illusion and well-being: a social psychological perspective on mental health. psychological bulletin , 103, 193–210. http://doi.org/10.1037/0033-2909.103. tock, j. l., & moxley, j. h. (2017). a comprehensive reanalysis of the metacognitive self-regulation scale from the mslq. metacognition and learning, 12, 79-111. http://doi.org/10.1007/s11409-016-9161-y veenman, m. (2011). learning to self-monitor and self-regulate. in r. mayer & p. alexander (eds.), handbook of research on learning and instruction (pp. 197–218). new york: routledge. vriesema, c. c., & mccaslin, m. (2020) experience and meaning in small-group contexts: fusing observational and self-report data to capture self and other dynamics. frontline learning research, 8 (3), 128–141. http://doi.org/10.14786/flr.v8i3.493 wang, a. y., & newlin, m. h. (2000). characteristics of students who enroll and succeed in psychology web-based classes. journal of educational psychology, 92, 137. http://doi.org/10.1037/0022-0663.92.1.137. winne, p. h., & baker, r. s. (2013). the potentials of educational data mining for researching metacognition, motivation and self-regulated learning. journal of educational data mining, 5, 1-8. retrieved from https://jedm.educationaldatamining.org/index.php/jedm/article/view/28 winne, p. h., & perry, n. e. (2000). measuring self-regulated learning. in m. boekaerts, p. r. pintrich, & m. zeidner (eds.), handbook of self-regulation. academic press. zhou, m., & winne, p. h. (2012). modelling academic performance by self-reported versus traced goal orientation. learning and instruction, 22, 413-419. http://doi.org//j.learninstruc.2012.03.004 zimmerman, b. j. (2008). investigating self-regulation and motivation: historical background, methodological developments, and future prospects. american educational research journal, 45, 166–183. http://doi.org/10.3102/0002831207312909 appendix a appendix b example questions from the 2016-2017 exam an example of a knowledge question:* fill in the blanks: the ….i…. test is used for testing the difference in proportions between dependent samples and the ….ii…. test for testing the difference in proportions between small independent samples. a. i: binomial, ii: chi-squared b. i: chi-squared, ii: binomial c. i: mcnemar, ii: fisher’s exact d. i: fisher’s exact, ii: mcnemar an example of an insight question: in order to evaluate differences in study success across different programs offered at the vu, an equal number of students are randomly selected from each program and asked to participate in the research. this is an example of: a. systematic random sampling b. cluster sampling c. stratified random sampling d. multi-stage sampling an example of a calculation question: research results have revealed that intelligence in the netherlands is normally distributed. the mean iq score is 100 with a standard deviation of 15. john has an iq-score of 122.5. what percentage of people in the netherlands will have an iq lower than that of john? a. 3.34 % b. 6.68 % c. 93.32 % d. 96.66 % *the correct answer is provided in bold. microsoft word final_proofs.docx frontline learning research vol. 9 no. 4 (2021) 35-75 issn 2295-3159 info corresponding author: athina koutsianou, department of primary education, university of ioannina, 45110, ioannina, greece. email: a.koutsianou@uoi.gr doi: https://doi.org/10.14786/flr.v9i4.777 unravelling the interplay of primary school teachers’ topicspecific epistemic beliefs and their conceptions of inquiry-based learning in history and science athina koutsianou et anastassios emvalotis department of primary education, university of ioannina, greece article received 27 december 2020 / article revised 6 september 2021 / accepted 16 september / available online 21 october abstract inquiry-based learning remains both an important goal and challenge for primary school teachers within and across different subjects, such as history and science. by addressing primary school teachers, for the first time, as both learners who deal with controversial topics and teachers who have significant teaching experience, this study aims to unravel the interplay between teachers’ topic-specific epistemic beliefs and their conceptions of inquiry-based learning in history and science. fifteen primary school teachers from greece participated in this exploratory study through scenario-based semi-structured interviews. data were analysed through a qualitative content analysis by applying both deductive and inductive approaches. the results of this study revealed the complex nature of teachers’ epistemic beliefs and the necessity of using a nuanced approach to elicit their epistemic belief patterns in the context of working on a task. further, this study revealed an overview of teachers’ conceptions of inquiry-based learning in both history and science, by giving voice to teachers’ thoughts and reasoning. but most importantly, the interplay between the two constructs was unravelled, indicating a complex connection between teachers’ epistemic belief patterns and their conceptions of inquiry-based learning. overall, it could be argued that the more availing the teachers’ epistemic beliefs, the more thoroughly they conceive inquiry-based learning. similarities and differences between history and science were also detected. theoretical, empirical, and educational implications are discussed in an attempt to support primary school teachers involve themselves, and then their students, in an active process of knowing by applying helpful epistemic criteria. keywords: epistemic beliefs; inquiry-based learning; primary school teachers; history; science koutsianou et emvalotis 36 | f l r 1. introduction the increasing challenges of the 21st century require citizens to think critically and make informed decisions about everyday complex and usually controversial issues, such as the consumption of genetically modified food or the use of artificial intelligence products (greene & yu, 2016; sinatra & hofer, 2016). in order to support students in addressing such challenges, there are many calls worldwide for changes in educational systems both in terms of naturaland social-science subjects (e.g., european commission, 2015; ncss, 2010, 2018; ngss lead states, 2013; nrc, 2012). specifically, teachers are asked to engage their students in epistemic practices that experts from different academic domains adopt to produce knowledge: such as interpretation and integration of multiple sources in history; experimentation and argumentation in science (greene & yu, 2016; sinatra, 2016). through such practices, teachers can guide students to construct and evaluate their own and others’ knowledge claims about specific topics (greene & yu, 2016, p. 49). in broader terms, teachers are called to apply inquiry-based learning (ibl) as an instructional approach to engage their students in authentic scientific processes through a set of inquiry phases (e.g., dobber, zwart, tanis, & van oers, 2017; pedaste et al., 2015), which incorporate the epistemic practices above. however, in-service teachers often struggle to implement such practices and foster ibl, due to many internal and external factors, especially primary school teachers who are trained as generalists and asked to teach various subjects, such as mathematics, literacy, science, history, and citizenship (avraamidou, 2017; gillies & nichols, 2015; greene & yu, 2016; levy, thomas, drago, & rex, 2013; van uum, verhoeff, & peeters, 2016). a large body of theory and research indicates that teachers’ beliefs are important internal factors that shape their teaching actions (see buehl & beck, 2015; fives & buehl, 2012, 2016). according to fives and buehl (2017), teachers’ beliefs are conceptualised as part of an integrated multidimensional system that also includes connected conceptions and values that “govern (their) cognitive and external actions” (p. 26). specifically, teachers’ beliefs about knowledge, learning and teaching are considered to be the internal factors which function as filters for new information, frames of salient tasks and guides of their actions (see fives & buehl, 2012, 2016, 2017); and thus, support or hinder them from applying innovative teaching practices and new curriculum standards (feucht, 2011; fives & buehl, 2012, 2016; greene & yu, 2016). teachers’ epistemic beliefs or teachers’ broader epistemic cognition, in particular, are at the forefront of the current literature (e.g., buehl & fives, 2016; feucht, lunn brownlee, & schraw, 2017; lunn brownlee, ferguson, & ryan, 2017; schraw, lunn brownlee, olafson, & vanderveldt, 2017) as one aspect of this complex system, supporting their crucial but also salient role in both learning how to teach and teaching praxis. there is growing evidence to suggest that teachers’ beliefs about the nature of knowledge and knowing (i.e., epistemic beliefs) mediate how they conceive learning and teaching, and further, how they engage in teaching. nevertheless, in-service primary school teachers’ epistemic beliefs have not been explored while they personally deal with the diversity and complexity of knowledge within topics from different academic domains, such as history and science. from another perspective, in order for primary school teachers to foster ibl in the context of such different subjects, a deep understanding of what ibl is and how it can be achieved within the classroom is also required (e.g., ireland, watters, brownlee, & lupton, 2012; levy et al., 2013; van uum et al., 2016; voet & de wever, 2016). but how do in-service primary school teachers conceive ibl within such different subjects as history and science is another underexplored aspect of this complex system. framed within different kinds of literature as epistemic cognition, teachers’ beliefs, science and history education, this study attempts to unravel the interplay between in-service primary school teachers’ topic-specific epistemic beliefs and their conceptions of ibl in history and science education. by addressing primary school teachers, firstly, as learners who deal with two controversial topics and then as teachers who articulate how they conceive ibl within different subjects, this study aims to get greater insight into teachers’ internal factors that might motivate or prevent them from engaging their students in ibl. further, this study addresses important conceptual and measurement issues regarding koutsianou et emvalotis 37 | f l r teachers’ epistemic beliefs by applying an integrative theoretical framework and addressing them as learners in the context of working on a task. 2. theoretical background 2.1. epistemic beliefs: an integrative theoretical framework epistemic cognition and personal epistemology are two umbrella terms that have been used extensively to merge different but also related theoretical constructs (see greene, sandoval, & bråten, 2016b, 2016c; hofer & bendixen, 2012; hofer & pintrich, 1997, 2002). these constructs represent a large and heterogeneous body of theory and research regarding “how people acquire, understand, justify, change, and use knowledge in formal and informal contexts” (greene, sandoval, & bråten, 2016a, p. 1). nowadays, the terms epistemic cognition and epistemic beliefs are the most predominant (sinatra, 2016), though there is no consensus about their definition (hofer, 2016; greene et al., 2016a). epistemic beliefs are mostly described as individuals’ beliefs about the nature of knowledge (i.e., what knowledge is) and knowing (i.e., how one comes to know) (hofer & pintrich, 1997). besides, epistemic cognition is conceptualised as a broader construct encompassing all kinds of explicit and tacit cognitive processes related to epistemic or epistemological matters (chinn, buckland, & samarapungavan, 2011, p. 141). furthermore, it is not conclusive whether epistemic beliefs are part of the broader construct of epistemic cognition or whether they are distinct but related constructs, depending often on the adopted conceptual framework (greene et al., 2016a; greene, yu, & copeland, 2014; hofer, 2016; sinatra, 2016). based primarily on theory and research with students from primary school to university, a proliferation of models of epistemic beliefs and epistemic cognition has emerged, revealing conceptual confusion, but also the developmental path of this interdisciplinary field (see greene et al., 2016b, 2016c). in short, the existing models can be broadly organised into three theoretical approaches or can be seen as a result of their integration: the developmental, multidimensional and contextual (or later situated) approach (see hofer, 2016 for an updated review). moreover, the dominant way of measuring epistemic beliefs and epistemic cognition, self-report instruments, are followed by numerous limitations and criticism (see mason, 2016 for a review), which cast doubt on a large part of the existing research findings (see greene & yu, 2014). another important issue pertains to the degree of specificity of epistemic beliefs and epistemic cognition (hofer, 2016; mason, 2016). although this line of research started by considering epistemic beliefs and epistemic cognition as more domain-general, there is also a proliferation of models and frameworks supporting the domain-, topic-, task-specific or even the situated nature of these constructs based on empirical evidence (e.g., chinn et al., 2011; greene et al., 2016c; hammer & elby, 2002; merk, rosman, muis, kelava, & bohl, 2018; muis, bendixen, & haerle, 2006; sandoval, 2005). framed within this ever-growing literature of learners’ epistemic cognition, a lot of studies have investigated both preand in-service teachers’ epistemic beliefs about teaching knowledge (e.g., buehl & fives, 2009), specific academic domains (e.g., guilfoyle, mccormack, & erduran, 2020; voet & de wever, 2016) or even topics (e.g., feucht, 2011, 2017; merk et al., 2018). in addition, teachers’ epistemic beliefs have frequently been operationalised through aspects of teaching and learning, resulting in a partial overlap with their beliefs about teaching and learning (maggioni, vansledright, & alexander, 2009; muis & foy, 2010; olafson & schraw, 2010). more recently, frameworks have been developed for conceptualising: a) the complex nature of teachers’ epistemic cognition regarding both learning to teach and teaching praxis (e.g., buehl & fives, 2016); and, b) potential mechanisms for changing teachers’ epistemic cognition in action (e.g., feucht et al., 2017; lunn brownlee et al., 2017). however, in-service teachers’ epistemic beliefs in the context of dealing with challenging issues as learners seem to have been overlooked empirically (vanslendright & maggioni, 2016), followed by conceptual confusion and measurement limitations. attempting to address this gap, the present study koutsianou et emvalotis 38 | f l r adopted an integrative theoretical framework to both conceptualise and approach primary school teachers’ epistemic beliefs probed within the context of controversial topics in history and science. in more detail, the integrative framework of epistemic beliefs applied in this study consists of four dimensions (i.e., the certainty and structure of knowledge, indicating the nature of knowledge; and the source of knowledge and justification of knowing, indicating the nature of knowing) which are interrelated in a relatively coherent system; and three perspectives (i.e., absolutism, multiplism and evaluativism) through which people understand the content of each dimension. table 1 presents a thorough description of the applied framework. this framework was developed based on feucht’s (2011, 2017) framework which integrated the dimensional and developmental frameworks proposed by hofer (2001) and kuhn (1999), and by incorporating further elaborated descriptions of these dimensions and perspectives derived from the literature (barzilai & weinstock, 2015; bendixen, winsor, & frazier, 2017; chinn et al., 2011; hofer, 2000; hofer & pintrich, 1997; king & kitchener, 1994; kuhn, iordanou, pease, & wirkala, 2008; kuhn & weinstock, 2002). such a nuanced approach of epistemic beliefs enables the recognition of unique epistemic belief patterns rather than simply categorising teachers’ epistemic beliefs as naïve or sophisticated for each dimension, which is quite a common limitation (feucht, 2011, 2017; mason, 2016; merk et al., 2018). moreover, the terms “less availing” and “more availing” were chosen (muis, 2004) to describe teachers’ beliefs across three perspectives, from absolutism to multiplism and evaluativism in an attempt to avoid making unnecessary assumptions about their value (greene et al., 2016c). koutsianou et emvalotis 39 | f l r table 1 conceptualisation of epistemic beliefs: an integrative framework of 4 dimensions and 3 perspectives dimensions perspectives absolutism multiplism evaluativism certainty of knowledge knowledge as a more or less certain and stable product knowledge about this topic is (absolutely) certain and stable. knowledge about this topic can be obtained with (absolute) certainty at some future point (in temporary uncertainty). (cert*a) knowledge about this topic is (completely) uncertain, constantly changing and evolving. knowledge about this topic cannot be obtained with certainty at any future point. (cert*m) knowledge about this topic is uncertain and evolving/cannot be obtained with (absolute) certainty, but it is possible to improve the degree of certainty at some future point. (cert*e) structure of knowledge knowledge as a more or less abstract and structured product knowledge about this topic is seen as an accumulation of objective and rather independent facts (e.g., concrete data or concepts) that are correct or incorrect in their representation of reality. (stru*a) knowledge about this topic is seen as a set of personal and subjective opinions (i.e., knowledge as an abstraction, not limited to concrete data), that are accountable only to their owners and possibly connected with others. (stru*m) knowledge about this topic is seen as a set of judgments based on evidence (e.g., theories, explanations and interpretations) which are complex and highly interrelated. (stru*e) source of knowledge knowing as a more or less active process of acquiring knowledge knowledge comes from external and objective sources outside the self. knowledge is passively perceived through external sources/is transmitted (i.e., people seeking the source of their knowledge outside themselves). (sour*a) knowledge is generated by human minds within the self. knowledge is actively formed based on internal and subjective sources (i.e., knowledge is filtered through the perceptions of the person making an interpretation; hence, what is known is limited by the perspective of the knower). (sour*m) knowledge is generated by human minds, anchored in standards of what counts as reliable knowledge, outside and within the self. knowledge is actively constructed based on the integration of internal and external sources, usually, in interaction with others. (sour*e) justification of knowing knowing as a process of evaluating knowledge claims knowledge requires no justification since one must only “observe” to know (i.e., firsthand experience). knowledge is justified based on absolute criteria (e.g., right or wrong) or based on authority and experts. (sour*a) knowledge is justified by giving reasons and using evidence, but the criteria are idiosyncratic to the individual (e.g., choosing the arguments and evidence that fit a personal opinion, preference, experience or judgment). (just*m) knowledge is justified based on criteria of argument and evidence, in terms of what is most reasonable or probable based on current evidence; and it is re-evaluated when relevant new evidence, perspectives, or tools of inquiry become available (i.e., based on the rules of inquiry in different domains). (just*e) koutsianou et emvalotis 40 | f l r 2.2. primary school teachers’ conceptions of ibl: an underexplored area according to fives, lacatena, and gerard’s (2015) systematic literature review (from a contentgeneral perspective), there is much inconsistency in how researchers define and approach teachers’ beliefs about (or conceptions of)1 teaching and learning, leading usually to the choice of dichotomising beliefs along a continuum of transmissionist (teacher-centred) to constructivist (student-centred) approaches. however, this categorisation is too broad to capture meaningful nuances, followed by conceptual issues. for instance, even though “constructivism” is a theory of learning according to which students learn by constructing their knowledge (see also windschitl, 2002), it has often been misused as a theory of teaching as well. in their review, fives et al. (2015) concluded that very few studies have focused directly on teachers’ beliefs about learning (e.g., chan, 2011), including within specific academic domains (e.g., tsai, 2002), regardless of their role in teaching practice in conjunction with teachers’ beliefs about teaching. with regard to in-service teachers’ conceptions of ibl, the existing research is very limited and mainly localised in science education (however, see voet & wever, 2016 for an example in history education). according to levy et al. (2013), the definition of ibl may differ even within the same subject, such as science and history education. in general terms, however, ibl can be defined as an educational strategy according to which students are actively engaged in methods and practices similar to those that experts apply to produce knowledge (pedaste et al., 2015), such as experimentation and argumentation (greene & yu, 2016; sinatra, 2016). several studies in science education have focused on how teachers conceive scientific inquiry (e.g., bartos & lederman, 2014; windschitl, 2004) and classroom inquiry (e.g., kang, orgill, & crippen, 2008), or how they implement ibl in action based either on classroom observations (e.g., forbes, biggers, zangori, 2013) or on teachers’ self-reports (e.g., lakin & wallace, 2015) and self-reflections (e.g., gillies & nichols, 2015; seung, park, & yung, 2014). most of these studies though pertain to pre-service teachers or in-service secondary science teachers, and a common way of analysing their conceptions and practices of ibl is based on the detection of essential features proposed by science education standards, including engaging in scientifically oriented questions and formulating explanations based on evidence (e.g., ngss lead states, 2013; nrc, 2012). ireland et al. (2012), on the other hand, avoided using such a predefined coding scheme and terminology and applied a phenomenographic analysis to qualitatively categorise in-service primary school teachers’ reported experiences of teaching for inquiry learning in science. based on teachers’ interviews, the researchers revealed three prevailing conceptions: the experience-centred conception where teachers mainly engage their students in sensory activities; the problem-centred conception where teachers mainly involve their students in solving challenging problems; and the question-centred conception where teachers mainly support their students to ask and answer their questions (pp. 166-169). on the other hand, even though there is an equivalent interest in applying ibl in history education (levy et al., 2013; martell, 2020), research on teachers’ conceptions of ibl is very scarce (see voet & de wever, 2016); especially with regard to primary school teachers, no relevant study was detected in the literature. specifically, voet and de wever (2016) explored in-service secondary history teachers’ beliefs about the nature of history and teaching history, their interplay, and further potential contextual influences in order to gain a greater overview of teachers’ conceptions of ibl. based on teachers’ interviews, the researchers also revealed three availing conceptions, ibl as: understanding by collecting more information about a given topic; evaluating critically more information about a given topic to find the right answer; and, investigating by generating questions within a given topic, analysing information and forming arguments (pp. 61-62). finally, it is worth mentioning levy et al.’s (2013) endeavour to compare and contrast conceptions of ibl in science, history and english education in order to outline cross-disciplinary conceptions of ibl by combining their conceptions with the findings of their previous studies with preor in-service teachers of each domain. in short, they concluded that even though a singular definition may be elusive and conceptions of ibl can vary even within each subject, there are common components. that is, teachers are asked to support their students to ask and answer questions based on data, and to justify koutsianou et emvalotis 41 | f l r explanations about the phenomena under investigation. framed within this underexplored research area, this study attempts to produce an overview of in-service primary school teachers’ conceptions of ibl both in history and science education, in order to offer a new perspective on the existing literature. 2.3. the interplay of teachers’ epistemic beliefs and conceptions of ibl many studies, especially in science education, have explored the relation of teachers’ epistemic beliefs with different constructs situated in teaching practice, including: teachers’ conceptions of teaching and learning in general (e.g., chan, 2011; cheng, chang, tang, & cheng, 2009); their selfreported or applied practices in science teaching (e.g., kang, 2008; tsai, 2007); and both their conceptions and practices in science teaching (e.g., lee & tsai, 2011). in general, these studies suggest that teachers’ epistemic beliefs are connected somehow with their conceptions of teaching and learning, and further, with their teaching practices; even though these relationships are not linear but very complex (for a comprehensive review see buehl & beck, 2015). additionally, a common claim found in this literature is that the more “sophisticated” the teachers’ epistemic beliefs, the more “constructivist” approaches they support, and vice versa. however, such findings are not considered well-documented because of the many conceptual and measurement limitations, which come with the existing literature (see also 2.1. and 2.2. sections). the interplay between teachers’ epistemic beliefs and conceptions of ibl has been directly examined by voet and de wever (2016), through the perspective of history teachers in secondary education. in short, their study revealed that although most history teachers appeared to hold criterialist (i.e., evaluativist) beliefs about the interpretive nature of history, only a few described ibl through the lenses of a full historical inquiry, indicating a misunderstanding of inquiry practices that lie at the core of this domain. however, a much clearer connection was found between teachers’ beliefs about the nature of history and history teaching in general. by adopting a much more detailed approach of epistemic beliefs, and through the perspective of primary school teachers, this study attempts to unravel this interplay within two academic domains and the corresponding primary school subjects, history and science. 3. research questions in this study, the following research questions were addressed: 1. what are participants’ epistemic beliefs while dealing with two controversial topics in history and science? 2. what are participants’ conceptions of ibl in history and science education? 3. are there any differences in participants’ epistemic beliefs and conceptions of ibl between history and science? 4. how are participants’ topic-specific epistemic beliefs related, if at all, with their conceptions of ibl in history and science education? 4. methods 4.1. participants and research design a qualitative research design was adopted to unravel the interplay of primary school teachers’ epistemic beliefs and conceptions of ibl, without aiming to generalise the findings of this study koutsianou et emvalotis 42 | f l r (merriam, 2009). given the exploratory nature of this study, the snowball sampling was applied (creswell, 2015) by setting the prerequisite that participants had taught at least one school year of the 6th grade of primary education (i.e., 11-year-old students). in total, 15 primary school teachers from three different counties in greece agreed to take part in semi-structured interviews with the first author, during the spring semester of the school year 2018-2019. the condition of having taught in the 6th grade was set to ensure that all teachers were familiar with both the historical and scientific topics used in their interviews, as well as aiming to teach history and science through ibl based on the current curricula in greek primary education. after being informed about the research procedure, the confidential use of the collected data and participants’ right to withdraw from the interview at any time, teachers who gave their consent participated in this study voluntarily and anonymously. there was also an attempt to make participants feel comfortable enough to avoid having to reply in a socially desirable way, by explaining that there were no right and wrong answers. in total, eight teachers were female and seven were male, with the mean age being 47.6 years (sd = 5.87 years). twelve teachers had completed a four-year bachelor’s degree programme at university, for teaching all the basic subjects in primary education (e.g., literacy, mathematics, science and history). five of these teachers had also completed master studies in education, one also had a four-year bachelor and a master degree in computer science (teacher 11), and another teacher also had a four-year bachelor degree in pre-school education (teacher 14). one teacher was a master student when this study took place. all six teachers with master studies had some research experience in the context of their thesis. the other three teachers held a degree of a two-year training programme at a pedagogical academy for teaching all the basic subjects in primary education as well. participants’ teaching experience in primary education ranged from 12 to 32 school years, and teaching experience in the 6th grade in particular, ranged from one to 16 school years. their teaching experience ensured that they had a lot of opportunities to reflect on their beliefs about knowledge, teaching and learning in the context of the school classroom. table 2 presents the demographic characteristics of the participants in detail. table 2 demographic characteristics of participants (n = 15) participant gender age level of education teaching experience in the 6th grade total teaching experience teacher 1 m 51 4-year bachelor degree 12 21 teacher 2 f 53 2-year bachelor degree 16 23 teacher 3 m 51 4-year bachelor degree & master student 4 15 teacher 4 f 52 4-year bachelor degree & master degree 10 24 teacher 5 m 49 4-year bachelor degree 11 22 teacher 6 m 52 2-year bachelor degree & master degree 8 26 teacher 7 f 48 4-year bachelor degree 4 21 teacher 8 m 51 2-year bachelor degree 6 23 teacher 9 m 44 4-year bachelor degree 7 18 teacher 10 f 40 4-year bachelor degree 1 16 teacher 11 m 37 two 4-year bachelor degrees & two master degrees 1 12 teacher 12 f 43 4-year bachelor degree & master degree 3 18 teacher 13 f 54 2-year bachelor degree 15 32 teacher 14 f 52 two 4-year bachelor degrees & master degree 14 22 teacher 15 f 37 4-year bachelor degree 3 15 notes. f = female; m = male. age was measured in years and teaching experience in school years (i.e., 10 months). koutsianou et emvalotis 43 | f l r 4.2. interview protocols and procedure semi-structured interviews were conducted to elicit both participants’ epistemic beliefs and conceptions of ibl (cohen, manion, & morrison, 2007; merriam, 2009). focusing on the construct of teachers’ epistemic beliefs, two topics were chosen from the 6th-grade greek textbooks of history and science: a) the governance of ioannis kapodistrias (the first governor of greece during the period 18281831); and b) the orientation of the migratory birds (specifically, the european robins) based on earth’s magnetic field. these topics were selected because they are controversial according to the existing literature and so the condition was necessary in order to trigger teachers’ epistemic beliefs within a task (barzilai & weinstock, 2015; vansledright & maggioni, 2016). for each topic, a scenario was developed by using two authentic historical (kremmidas, 2015; ploumidis, 2015) and scientific studies (holland & helm, 2013; wiltschko, gehring, denzau, nießner, & wiltschko, 2014), which address the same topic differently and conclude in different explanations. within this context, and handling each topic separately, participants were first asked to read both accounts and then answer a set of open-ended questions, aiming to elicit their epistemic beliefs as learners, implicitly and explicitly. for instance, teachers were asked to think about how they would go about making their own decision on each topic (this question aimed to prompt teachers’ beliefs about the source of knowledge). furthermore, they were asked to judge the different perspectives on each topic by mentioning specific criteria (this question aimed to prompt teachers’ beliefs about the justification of knowing). the applied interview protocol was developed based on the adjustment and integration of existing open-ended questions derived from the literature (barzilai & weinstock, 2015; king & kitchener, 1994; kuhn & weinstock, 2002). after finishing the aforementioned task for each topic, participants were asked to give a short description of how they would go about teaching this topic in their classroom with 6th-grade students, by putting more emphasis on the teaching process. in fact, this was a transitional question for shifting their attention from their role as learners (in the topic-specific level) to their role as teachers (in the domain-specific level). focusing on the construct of the conceptions of ibl in history and science education, teachers were also asked to answer a set of indirect and direct open-ended questions. the applied interview protocol was developed based on voet and de wever’s (2016) interview protocol regarding ibl in history education, in addition to certain clarifying questions (e.g., how would you define ibl in history/science education? what is the role of the students and what is your role in a lesson that focuses on ibl within history/science education?). after taking into account differences in the terminology between history and science, the questions were formulated to be equivalent. both scenarios and interview protocols were first pilot tested with greek experts in science and history, and primary school teachers, to further improve their wording and validity. all interviews were conducted by the first author in greek based on the aforementioned scenarios and interview protocols, which are provided translated in english in appendices a1 and a2 respectively. the same procedure was applied twice with every participant for both topics and domains in a counterbalanced order. however, the researcher could deviate from the protocols as needed to help participants elaborate more on their thoughts via clarifying questions. overall, the time duration of interviews ranged from 50.32 to 105.08 minutes per participant (m = 71.32 min., sd = 17.51 min.). 4.3. data analysis all interviews were audio-recorded, transcribed verbatim and coded through qualitative content analysis (hsieh & shannon, 2005; mayring, 2000) by using nvivo 12 software. specifically, the collected data were analysed by applying both deductive and inductive approaches depending on the specific theoretical construct. each interview was divided into codable segments and the segments that captured the meaning of participants’ thoughts were defined as the units of analysis (chi, 1997). the analysis of data regarding teachers’ epistemic beliefs was more deductive since there is a lot of relevant research. this happened by applying as a coding framework the integrative theoretical framework presented in table 1. hence, indicating all the combinations of dimensions and perspectives koutsianou et emvalotis 44 | f l r of epistemic beliefs per topic, 12 codes were generated in nvivo. representative examples of these codes are included in table 3. on the other hand, the analysis of data about teachers’ conceptions of ibl was more inductive since there is less relevant research. this happened by elaborating existing categories from the literature (ireland et al., 2012; voet & de wever, 2016) according to participants’ reasoning and thoughts. table 4 describes these categories and gives representative examples. the first author coded the interview data through a recursive process, by comparing and contrasting participants’ quotations within each code and category of the epistemic beliefs and conceptions of ibl. the researcher attempted to be reflexive in every step of the data analysis to avoid forcing her own ideas about the two theoretical constructs on the data (creswell, 2015; greene & yu, 2014). in order to check the reliability of the results, the second author independently coded eight randomly selected interview transcripts, acting as a peer debriefer for the first author, given that the latter was not involved in the data collection and development of the coding frameworks. the intercoder agreement was 85% for teachers’ epistemic beliefs and 92% for their conceptions of ibl. the different interpretations of the collected data, especially for the epistemic beliefs, were mainly because of the different types of beliefs expressed by the same participant throughout the interview. thus, the two authors presented to each other their arguments behind their decisions on the data analysis, in order to discuss and resolve their discrepancies, resulting in 100% intercoder agreement. koutsianou et emvalotis 45 | f l r table 3 representative examples of teachers’ topic-specific epistemic beliefs in history and science (dimension*perspective) codes example history (scenario topic: the governance of ioannis kapodistrias) cert*a teacher 14: if someone sits down and searches whatever has been written for him, they will safely come to a conclusion. so, from this point of view, yes, but i don’t know if there are sources. ... for me, historians do [know], so over here, this contrast is troubling me. yes, historians do [know]. cert*m teacher 6: no way! it’s too, too difficult. ... there is no certainty in history ... there is an approach, there is no certainty and the approach, how they say it, [it] depends on the school you are in, the historical school you are in ... you capture your own opinion, a historical opinion. there is no certainty. cert*e teacher 4: yes, with certainty yes, but to some extent. absolutely, no. absolutely no, because... perhaps though, these historians who are more concerned with and read the sources, certainly they have a more specialised opinion and can express it. ... i think yes, that complete certainty no. because, is it even possible that they have read all the relevant, all the critical sources that exist? ... but again, wouldn’t the element of subjectivity exist inside? i think that there is always inside the historian. it always exists, no matter how objective they want to be. stru*a teacher 15: [knowledge] consists of the data we have, again as i said before, which are concrete because they are things and events that happened and have passed through written texts and from other sources... i can’t think of anything else. stru*m (see the description of the teacher 1 provided in the 5.3.1. section of this article.) stru*e teacher 8: [knowledge structure] is not simple... in no case it is simple. it is rather a complex process... perhaps more complex than a common reader can afford, let’s say, right? you read the work of a historian about a person for whom you know basic things, for his/her life, and you face a difficulty in understanding how they justified their opinion, how they integrate the facts and how they make it the knowledge that they want to offer. because the historian after their research... i consider the result of their research to be the knowledge that they want to also offer to the public, right? sour*a teacher 15: well, i would try to focus on the data of each source that will help me to acquire a thorough opinion of his life and work, data in order to understand his personality and then ... about how he functioned in his politics, in this way. sour*m teacher 6: i try to approach it as well ... because you can’t have it, the truth especially in this thing, the truth exists, but we try to approach it, but we approach it from our own side, it is not possible; i will write, i have a ‘load’ inside me that i am carrying, that is, i say even if i lived [then] or kapodistrias was my father, again i would have a load, i would write from my own side. in other words, there should be an objective observer who can see and record. yes! ... but again, he/she would have his/her own perspective, right? it would be very difficult. sour*e teacher 9: the events are always specific; the interpretation of the events is the one that differs each time. so, depending on what everyone classifies as the most important, they also express their opinion afterwards. ... basically, i try to see it ... or at least gather information from as many sources as possible, and from as many people as possible, if i know people who are related to the topic. and beyond that, i try to ... judge, to draw my own conclusions, whether something is positive or not, or to what extent it is positive or not, let’s say. just*a teacher 13: the sources would count, i would only count the facts, the sources, i would not place my own opinion ... that is, without being placed, here both [historians] are placed. ... without historian’s opinion. just*m teacher 14: logically, i will also be influenced by the teaching of history for so many years, let’s say, with the one book that we have as a source. also, some reading of mine about kapodistrias and in general, he is one of the political people that i personally like; maybe because he was the first koutsianou et emvalotis 46 | f l r governor, maybe because he really made the first attempt to set up the education and, hence, make a fight for the country to deal with the primary sector. i believe my thinking has a reasonable basis, i believe, subjectively. just*e teacher 3: what plays a role is whether everyone can justify it, their opinion. of course, and the fact that i personally lack a lot of knowledge on this specific topic, maybe i am wrong based on my previous knowledge ... and tend to an opinion which ... is based on what i have as knowledge; and it is possible to be unjust to another opinion. therefore, it is easy to get carried away based on your experiences and knowledge to think that one opinion may be more correct than the other, but the right and fair is, since there is justification on the other hand, as much as we do not like this view, to check it and investigate it in order to see if and to what extent it ... may be to the right direction. ... criteria ... obviously, what i mainly pay attention to is to try to understand whether this [any historical account] can be appropriate at that time. science (scenario topic: the orientation of migratory birds based on the earth’s magnetic field) cert*a teacher 7: of course, it’s not so difficult that scientists cannot help us to certify it. cert*m teacher 15: i find it very difficult ... because some things are ... inexplicable! we cannot determine everything based on physics and laws. ... as i told you before, i think that there are some things that, no matter how close we get, we cannot give a solution and answers to these issues. cert*e teacher 5: nor do i think that is absolute... well ... now i think that ... if this research is repeated in different populations of robins and has results, at some point we will get to the point of talking about ... with greater certainty about the journey. str*a teacher 14: [knowledge consists] clearly of measurable results, that is, here you have to measure, have a result, the experiment must be specific, the experiment is not a theory. you apply, try and end up. hypothesis, test and error, normally the course we have to do in science, so it goes by itself. ... data mainly. and no interpretation of the data; measurement and after, they build the result, i believe. str*m (there was not found a representative example of this dimension*perspective regarding this topic.) str*e teacher 8: [knowledge consists of] too much research and the result of this research, right? and a practical result of this research. ... definitely data ... opinions because ... different people express different opinions. the more trained people are, the better and higher quality products they are based on. their opinions are obviously better, and their interpretations as well, which are their perspicacity from there and beyond ... of everyone or how they will manage all this. ... because the interpretations produce the results obviously, right? you observe a phenomenon, you experiment, but you must definitely interpret it and leave your consignment. this consignment can be theoretical, it can create a thousand ... more questions, but it can also be practical. so, the interpretation, for me, plays a very important role, [the interpretation] of the many big data and the different opinions. sour*a teacher 3: [to make a decision is] very difficult because ... i don’t have the proper knowledge to do it. well, then, i couldn’t take the place of one or the other, but i would just stay in the simple quotation of the research results. sour*m teacher 6: the truth is one, which we try to approach, the same happens in history as well; the truth exists, we try to approach it and we perceive it with our perspective. i don’t know. i can’t answer. the truth exists, we cannot, we try to approach it but with our subjective way. even researchers try to be objective but i don’t know, since there are a lot of studies on this specific topic, i don’t know where they can conclude and how they can support it, i don’t know. ... ι try to approach the topic as much as possible, to have a personal opinion. sour*e teacher 8: i would read them once again. if i leaned towards one, i would look for it a little more and that might have attracted me to look for the other one a little more, and i would probably end up intuitively again and not scientifically. you know, [for] a scientific study, someone who is not a scientist but is a listener or a reader says after all, which one i like and not which one is the best, right? ... but obviously, the groups of scientists, who applied both methods, have their weapons in their arsenal, they are based on many things and since they have an outcome, maybe they are both... they answer a part of the question, which is what? how can they succeed and come and go in the same place, right? they don’t lose their way, they know. koutsianou et emvalotis 47 | f l r just*a teacher 13: depends on what data [the study] will have better, statistical data. ... i would look for more information to see the data. just*m teacher 3: if it is close to what ... i have learned and ... it can affect me more, i would say ... the first opinion, and be able to understand it, right? to bring it at my cognitive level; it would suit me such a plausible [opinion], i would say that it is truer. ... the only criterion i could adopt is this one ... what suits me best is what i have in my mind, that i have experienced, that is, my experiences ... that the magnet with the iron [are attracted], as i said before. just*e teacher 8: in an ideal environment that happens to fall into my hands two scientific, different ... not opposite approaches; since i am interested in the topic, because if i was not interested, i would not read the approaches, right? [i would focus on] how the issues are posed and what results each study has. in this way i would weigh it, that is, yes, the whole project done, the experimental, research [project], but from the first one, these results were emerged, from the second one, maybe less or poor results [were emerged]. obviously, i as a common mind as well, i would go ... i would judge as more well-justified ... the study that had more practical results perhaps? perhaps [the study] that ... opened new ways? in other words, if, from the first study, ... they said that after the experiment with robins, the same experiment done with the canaries let’s say ... and there were found corresponding results, obviously it would cover me more... notes. cert = certainty of knowledge; stru = structure of knowledge; sour = source of knowledge; just = justification of knowing. a = absolutism; m = multiplism; e = evaluativism. the abbreviations, such as cert*a, indicate all the different combinations of the four dimensions and three perspectives of epistemic beliefs: i.e., cert*a indicates the certainty of knowledge and absolutism. koutsianou et emvalotis 48 | f l r table 4 teachers’ conceptions of ibl in history and science education category description participants representative excerpts in history education understanding ibl is seen more as a practical way of gathering information and sources in order to better analyse and understand the given topic. 2, 10, 13, 15 teacher 2: [students] should know how to read the text very well ... because if they cannot read it correctly, they cannot understand it. from there on, they should isolate the facts, the people and listen to what i tell them at school, but also investigate further, that is, search for and bring opinions and information in the classroom. ... [ibl is] that i ask my students to bring the information, to gather all this and then, the main teaching begins, something like that i have in my mind. let them first investigate and, then, all together find the common goal. evaluating ibl is seen more as a critical way of learning how to compare and evaluate different opinions and sources about the same topic in order to realize their subjectivity and develop critical thinking skills. 1, 3, 4, 5, 7, 9 teacher 9: [with reference to the topic of the given scenario, students] should think on both sides that historians support. and then, if they could find out if something is right, which is right, which is wrong, and, basically, where they conclude. in other words, i would like to see children’s conclusions. ... [in ibl] students need to understand that they don’t have to accept whatever they read ... without thinking. the purpose of the school is to create children’s ability of critical thinking, be able to read three different sources for the same topic and choose, think; that is, they should have created a critical way of thinking regarding what they read and from there on ... they are led to knowledge and understanding. investigating ibl is seen more as an active process which resembles problem solving, by generating questions, searching in the literature, analysing information and formulating arguments; in order to explain a historical topic and eventually learn how to proceed in order to answer their own questions. 6, 8, 11, 12, 14 teacher 6: (he mentioned a lot the importance of local history.) you should find new ways, new instructional tools outside the school ... so that they become the young researchers, to search as much as they can. you will give them books, sources of the past as well. ... they should discover some things so that they can connect them with the local, with their own, with what they perceive. ... [regarding a historical topic,] you can go to search what is written in the local newspapers. ... you can make them photographers, ask them to take photographs from the old newspapers, this is a game at the same time and they really like it, right? ... [in ibl] students should investigate and the teacher should be the research assistant. of course, the teacher should bear in mind that he should ask them what to look for, right? (in order to explain his thoughts, he mentioned a lot of questions regarding the local history of his city, such as) what statues, monuments are there in [this city]? for example, i will go again to the [greek] civil war, why are there only statues of one side and not of the other? children can understand this, right? or how is the memory imposed in the space? ... but again, this requires you to leave school. ... or why isn’t there a female figure in a statue? they should wonder ... observe ... koutsianou et emvalotis 49 | f l r search. search here, did you see a female figure? no, there is no female figure or i think there is one now, recently put. you can do such things. in science education “hands-on” experiencing ibl is seen more as hands-on activities, such as conducting experiments or making observations in order to obtain the expected right results (or prove an already known theory) and understand a natural phenomenon. 1, 2, 7, 13, 14, 15 teacher 7: last time we went to a chemistry laboratory and the chemist showed us various chemical reactions of the elements, they were very happy and they really liked it when the colour changed; the alcohol for example, if we throw it in a chemical compound, or if it gets more intense colours. she showed them various such tricks and they really liked them and then, we did them and said them in the classroom. ... inquiry is all these experiments we try to do. “minds-on” experiencing ibl is seen more as minds-on activities, such as conducting experiments, making observations and/or searching in the literature in order to explain a natural phenomenon and become familiar with the scientific way of thinking. 4, 5, 6, 8, 9, 10, 12 teacher 4: we try to make young scientists in the sense that we try to instil in children the logic that in order to come to some conclusions about what is happening around me, first, i should observe something, to pique my interest, and beyond that to do an experiment to see, is this really the case? and then to come to a conclusion, which i will generalise, that since it applies here, it also applies in some other areas. ... we should suspect children, from an early age, in such a way of thinking that... everything around us should have a logical explanation and be based on some, some evidence. investigating ibl is seen more as an active process which resembles problem solving, by generating questions/hypotheses, using methods and tools, conducting experiments, making observations and/or searching in the literature, analysing information, and formulating arguments; in order to explain a natural phenomenon and eventually learn how to proceed in order to answer their own questions. 3, 11 teacher 3: always, in every lesson, i attempt to learn students’ prior knowledge, their concerns, what they would like to learn, what would be of their best interest, their perception of the problems that they pose, how it may have been created, and then, through the existing experiments or their opinion [i attempt] to include them in the lesson. ... in science lessons, they should have the willingness to discover ... to search, experimenting with different materials ... to be involved in the natural world and ... to be able even to reconsider some of their opinions and understand that what we often accept as empirical knowledge may not be true and may not be confirmed by scientific knowledge. to be able to come into a conflict with the data they have in mind. ... [in ibl] students, through their questions, through their willingness to discover, either by finding the appropriate tools or by providing them with the appropriate tools, attempt to form hypotheses, and through experimentation to confirm these hypotheses. koutsianou et emvalotis 50 | f l r 5. results the results of this study are presented in three sections. the first one gives an overview of teachers’ epistemic beliefs which emerged while they were dealing with two controversial topics in history and science. then, the second section describes teachers’ conceptions of ibl both in history and science education. a comparison of teachers’ epistemic beliefs and conceptions of ibl between history and science (as groups of patterns or categories and intra-individually) is included. finally, the third section attempts to capture the interplay of teachers’ epistemic beliefs with their conceptions of ibl within each topic and domain, by presenting three illustrative teacher cases. 5.1. primary school teachers’ epistemic beliefs table 3 includes representative examples of coding teachers’ utterances regarding each dimension and perspective of their topic-specific epistemic beliefs. every participant contributed uniquely to this study by expressing their epistemic beliefs; however, it was possible to detect some patterns in their responses. what follows is a short description of participants’ epistemic belief patterns found within the context of the two controversial topics. 5.1.1. topic-specific epistemic beliefs in history while dealing with the historical topic of the governance of ioannis kapodistrias, most participants articulated a complex set of epistemic beliefs without fitting in one single perspective. their responses revealed some meaningful patterns (see table 5). usually, participants’ beliefs about the certainty and structure of knowledge fell into the same perspective while their beliefs about the source of knowledge and the justification of knowing were rather complementary. specifically, a lot of participants were coded as evaluativists regarding the nature of knowledge about this topic, and as multiplists regarding the nature of knowing, especially the justification of knowing. this finding indicates that even though teachers recognised a rather uncertain and interpretive nature of historical knowledge, and were willing to involve themselves in an active process of knowing, they were not familiar with the epistemic criteria necessary to evaluate different sources. on the other hand, some participants were coded as absolutists regarding the nature of historical knowledge and as multiplists and/or evaluativists regarding the nature of knowing. this finding could indicate a more dogmatic view of historical knowledge as certain and mainly based on concrete data, in combination with teachers’ willingness to form their own opinion by integrating multiple sources, mainly based, however, on their personal preferences. furthermore, a few participants can be seen as typical examples of each perspective in total. 5.1.2. topic-specific epistemic beliefs in science while dealing with the scientific topic, that is the orientation of the migratory birds based on earth’s magnetic field, most participants expressed a rather complex set of epistemic beliefs without fitting into a single perspective. some meaningful patterns were also detected in participants’ responses; however, they were not identical to those found in the historical topic (see table 5). quite often, participants’ beliefs about the certainty and structure of knowledge fell into the same perspective, and their beliefs about the source of knowledge and the justification of knowing were complementary. specifically, a lot of participants were coded as evaluativists regarding the nature of knowledge, but also as absolutists regarding the nature of knowing (or even as multiplists for the justification of knowing). this finding could indicate that teachers were familiar with the evolving and interpretive nature of scientific knowledge; but at the same time, they believed that such topics are very specialised for them to have any involvement and make any decision. besides, if they had to take a perspective, the most common criterion would be what makes sense based on their prior knowledge of the topic and/or logic. a few participants showed a similar pattern with a notable difference regarding their beliefs about the source of knowledge, given that they were coded as evaluativists; but simultaneously, they had difficulty koutsianou et emvalotis 51 | f l r articulating specific epistemic criteria for evaluating different sources. compared with the historical topic, more participants can be seen as typical examples of the absolutist and evaluativist perspectives; but none of the multiplist perspective. 5.1.3. intraindividual comparison of teachers’ topic-specific epistemic beliefs in history and science one third of the participants showed quite similar patterns in their topic-specific epistemic beliefs in history and science. the remaining participants appeared to have more availing epistemic beliefs about the historical topic rather than the scientific one, except for one participant who had the opposite difference (see table 5). table 5 the interplay of teachers’ topic-specific epistemic beliefs and their conceptions of ibl in history and science topic-specific epistemic belief patterns conceptions of ibl participant topic certainty structure source justification teacher 1 i. kapodistrias m m e m evaluating migratory birds a a a m “hands-on” experiencing teacher 2 i. kapodistrias a a a a understanding migratory birds a a a a “hands-on” experiencing teacher 3 i. kapodistrias e a a e evaluating migratory birds e e a m investigating teacher 4 i. kapodistrias e e e m evaluating migratory birds e e a a “minds-on” experiencing teacher 5 i. kapodistrias m e e m evaluating migratory birds e e e m “minds-on” experiencing teacher 6 i. kapodistrias e e m & e m investigating migratory birds e e m m “minds-on” experiencing teacher 7 i. kapodistrias a a e m evaluating migratory birds a a a a “hands-on” experiencing teacher 8 i. kapodistrias e e e m investigating migratory birds e e a & e a & e “minds-on” experiencing teacher 9 i. kapodistrias m e e m evaluating migratory birds e e a m “minds-on” experiencing teacher 10 i. kapodistrias e e m m understanding migratory birds e a a m “minds-on” experiencing teacher 11 i. kapodistrias e e e e investigating migratory birds e e e e investigating teacher 12 i. kapodistrias e e m m investigating migratory birds m a & e a m “minds-on” experiencing teacher 13 i. kapodistrias m e a a understanding migratory birds a a a a “hands-on” experiencing teacher 14 i. kapodistrias a e m m investigating migratory birds a a a a “hands-on” experiencing teacher 15 i. kapodistrias a a a a understanding migratory birds m e a m “hands-on” experiencing notes. a = absolutism; m = multiplism; e = evaluativism. 5.2. primary school teachers’ conceptions of ibl 5.2.1. conceptions of ibl in history education based on the categorisation proposed by voet and de wever (2016) and participants’ interview responses, three qualitatively different ways of conceptualising ibl in history education were identified: koutsianou et emvalotis 52 | f l r understanding, evaluating, and investigating. table 4 presents these three categories by providing excerpts from teachers’ interviews as representative examples of each category. also, a list of participants in each category is included. initially, teachers of the understanding category tended to describe ibl mostly as a practical way of searching and collecting information from additional sources regarding the lesson topic, besides school books. the main goal of ibl is the enhancement of students’ content knowledge and a more practical understanding of the lesson. some teachers mentioned that they should search and collect additional sources to provide students with understandable and simplified information while others described this process by giving to students a more active role. in addition, a few teachers said that critical thinking should be another goal reached through ibl without, however, elaborating further. the teachers of the evaluating category described ibl mostly as a critical way of learning how to search, compare and evaluate different opinions and potentially conflicting information from different sources about the same topic. the main goal of ibl is the development of a more critical way of thinking and the realisation that history always involves subjectivity. students’ active role was frequently supported by these teachers and some of them further implied that through such a process, students would learn how to discern right and wrong information or even choose the right source, in an attempt to form their own opinions. many of these teachers hinted at the importance of using trustworthy sources within the context of ibl. finally, the teachers of the investigating category described ibl as an active process of generating questions, searching in the literature (or through outdoor activities), collecting and analysing information to form arguments and explain a historical event or phenomenon, which resembles a lot the process of solving a problem by asking research questions. despite the central role of asking questions in this category, it is worth mentioning that most teachers referred to it more indirectly by formulating questions independently, while giving examples of historical topics. the main goal of ibl for students is to learn how to investigate historical events and phenomena on their own and construct knowledge by integrating information from different sources. besides, only one teacher highlighted the importance of learning how to argue based on multiple sources, as part of the ibl. hence, it is apparent that there were some qualitative differences among teachers’ conceptions in this category. 5.2.2. conceptions of ibl in science education based on the existing literature (ireland et al., 2012; voet & de wever, 2016) and participants’ interview responses, three qualitatively different ways of conceptualising ibl in science education were also detected: “hands-on” experiencing, “minds-on” experiencing, and investigating. table 4 describes in detail these three categories and presents excerpts from teachers’ interviews as representative examples of each category. a list of participants in each category is also included. firstly, the teachers of the “hands-on” experiencing category defined ibl mostly as sensory activities, including experiments and/or observations. according to their descriptions, experiments are conducted either by the teacher as a demonstration or by students under the teacher’s close supervision in order to find out the expected results and understand the natural phenomena described in their school book more practically. in other words, experimental procedures are used to prove what is already known and has been discovered by scientists. even though science equipment and materials are described as an integral part of the inquiry, the use of simulations could be an alternative means, when an experiment is difficult to be conducted. however, none referred to the use of the literature without being asked directly by the researcher. even then, most teachers referred to the literature as necessary information to be learnt or even the theory that students should know before experimenting. only two teachers replied that searching in the literature for more information could be another kind of inquiry, aiming at understanding (in accordance with the understanding category of ibl in history education), but also mentioned how that would be much less interesting than experiments. the teachers of the “minds-on” experiencing category also defined ibl as sensory activities, such as experiments and/or observations, aiming furthermore to acquire a more scientific way of thinking, koutsianou et emvalotis 53 | f l r while attempting to explain the natural phenomena. based on their descriptions, students should have a very active role since they are supported by their teacher to search and discover how “nature works” from a scientific point of view. however, the mere experience of experimenting is not enough to describe ibl. the use of simulations was also mentioned as an alternative option. in comparison with the previous category, the teachers of this category have also connected ibl with the experiments and observations rather than with the literature. however, after being asked, most of them supported the idea of searching and using multiple sources as a different kind of inquiry to explain a natural phenomenon, although it is not considered so tempting for students. only one teacher further connected ibl with the search of the literature in an attempt to compare and contrast multiple sources and their trustworthiness based on critical thinking. this description can be seen as similar to the evaluating category of the ibl in history education. the teachers of the investigating category were only two and defined ibl more as an active process that resembles the process of solving a problem in order to explain a natural phenomenon, by highlighting the importance of forming hypotheses and/or asking questions. according to their descriptions, students should be very active in the process of inquiry by forming hypotheses; making observations and/or conducting experiments; analysing information to test the hypotheses; and finally, formulating explanations based on evidence. eventually, the goal for students should be to learn how to proceed in order to ask and answer their questions through inquiry. however, argumentation was not mentioned as part of the ibl. in accordance with the two previous categories, teachers also connected ibl with more material-based activities than with the searching in the literature, and further mentioned the alternative use of simulations. after being asked, the description of one teacher was very close to the investigating category of the ibl in history education, as another kind of inquiry in science education. the importance of learning how to think critically and more scientifically through ibl was also highlighted by both teachers. 5.2.3. intraindividual comparison of teachers’ conceptions of ibl in history and science education roughly half the participants fell into the same category regarding their conceptions of ibl in both domains, though the recognised categories in history education are not identical with those in science education. the other half of participants appeared to define ibl in history education more thoroughly than ibl in science education, except for two participants who had the opposite difference (see table 5). 5.3. the interplay between teachers’ epistemic beliefs and their conceptions of ibl unravelling the interplay between teachers’ epistemic beliefs and their conceptions of ibl within each topic and domain was the most challenging and essential part of this study. table 5 gives an overview of participants’ epistemic belief patterns for each topic and of their conceptions of ibl within the corresponding domain. there is an interesting interplay between the two constructs. starting with history, most participants who described ibl as understanding also expressed absolutist epistemic beliefs, either for all four dimensions (teachers 2 and 15) or at least for that of the nature of knowing (teacher 13). furthermore, participants who described ibl as evaluating showed a wide range of epistemic belief patterns, but two interesting tendencies were detected: a) almost all of them expressed evaluativist epistemic beliefs about the source of knowledge and multiplist epistemic beliefs about the justification of knowing (teachers 1, 4, 5, 7, and 9); and b) some of them similarly expressed multiplist epistemic beliefs about the certainty of knowledge (teachers 1, 5, and 9). finally, most participants who described ibl as investigating expressed evaluativist epistemic beliefs, at least regarding the nature and source of knowledge (teachers 6, 8 and 11). however, there were participants who could be seen as exceptions in terms of the aforementioned tendencies (e.g., teachers 3, 10 and 14). continuing with science, roughly all participants who defined ibl in science education as “handson” experiencing expressed absolutist epistemic beliefs (teachers 1, 2, 7, 13 and 14). furthermore, participants who defined ibl as “minds-on” experiencing also showed a range of epistemic belief koutsianou et emvalotis 54 | f l r patterns, but with a common tendency of expressing evaluativist beliefs about the nature of knowledge (teachers 4, 5, 6, 8 and 9). additionally, some of them were also coded as absolutists about the source of knowledge and/or as multiplists about the justification of knowing (teachers 4, 8, and 9). finally, the two participants who defined ibl as investigating articulated evaluativist beliefs about the nature of knowledge, but differed in their beliefs about the nature of knowing: one was quite hesitant to involve himself in an active process of knowing (teacher 3), and the other was entirely coded as evaluativist (teacher 11). nevertheless, there were participants who could be seen as rather exceptions in terms of the aforementioned tendencies (e.g., teachers 3 and 15). in order to further understand the interplay between teachers’ epistemic beliefs and conceptions of ibl, below follows an in-depth analysis of three teacher cases that could represent the aforementioned findings. at the same time, there was an attempt to choose the most illustrative teacher cases of each category of ibl in both history and science education, and also to present the same teacher cases for both domains as a whole in order to make cross-topic and cross-domain comparisons. according to table 1, letter codes in parentheses follow teachers’ quotations to indicate the specific dimension and perspective of their epistemic beliefs. participants addressed the topics in the same order that they are presented below. 5.3.1. teacher 1: ibl as “hands-on” experiencing in science and evaluating in history teacher 1 was 51 years old and he had been teaching at primary schools for 21 years, more than half of them were in the 6th grade. he had completed his 4-year bachelor studies in primary education, but he had no research experience. his interview started with the scientific topic and his first reaction was that “i am thinking that yet, there is nothing certain, it hasn’t been proven. although i converge more on the second study” by further explaining that he had read a lot about birds and had first-hand experience. his answers implied that knowledge comes from outside, including studying the bibliography and having first-hand experience by observing the world (sour*a). although initially, he seemed unwilling to further justify his way of knowing, as something given, later he explained that he would choose a study based on his personal preferences: i would read, i would look to see who these researchers are, what everyone has been dealing with, the bibliography they have read, and i would read it too, i would see afterwards, subjectively of course, with which of the two ... i would go. (just*m) following his line of thought, he was asked about the certainty of knowledge regarding the topic to which he noted: “yes, i believe that in the future, with appropriate experiments and ... observations, [it] can be proven. ... in the future ... [experts] will be able to know.” (cert*a) besides, he struggled to understand the questions about the structure of knowledge. in fact, he insisted that knowledge consists of “the bibliography and the studies that researchers have done” (stru*a), even after giving him the option of choosing among data, opinions or interpretations of data as basic components of knowledge. afterwards, he exemplified the use of experiments as the most important part of school science, which brings this subject very close to scientific research, even if in a simplified way. although he repeatedly mentioned the importance of supporting students to discover knowledge instead of transferring it, his description of inquiry in the class was exclusively focused on hands-on activities in order for students to prove what is already known by the scientists: “inquiry is through the experiments, you give them the materials, you tell them how, what they should do and they see themselves the result ... from what they do, with the instructions.” (“hands-on” experiencing) then, the interview focused on the historical topic, immediately provoking a more active role, thus indicating a common trend set by participants who appeared to feel more comfortable with the historical topic than with the scientific one. obviously, he had already developed an opinion about the topic based on many different (written and oral) sources, but the element of subjectivity in history was expressed at every opportunity, without articulating specific criteria for addressing it, except his personal opinion: koutsianou et emvalotis 55 | f l r as for kapodistrias, look, if you ask me now to tell you my opinion about kapodistrias ... i will speak to you subjectively. i may agree with ... some of what they say here, and with some [i may] disagree. ... i didn’t like some of them there, because i have a different perception, right? the sources, when they refer to a historical event, you can’t ignore them, right? but you can judge them ... set parameters. ... i have read about kapodistrias, you have done that at school ... until you became 20 years old, you had a perception. then, various studies come out and they tell you other things. what you knew based on what you are learning, i believe, it’s what shapes you and somewhere you draw a conclusion. (sour*e, just*m) regarding the nature of knowledge about this topic, his descriptions were indicative of the perspective of multiplism, though he distinguished between the certainty of facts and historians’ interpretations of kapodistrias’ ideology: for his work, whatever he has done, that is, what is mentioned [such as] schools, i do think that we know a few things. now, about how he was ideologically himself, i don’t know. ... because ... the historian who writes also has his own political, social, economic perceptions and writes as he wants. (cert*m) [knowledge consists of] sources. you can draw conclusions from people who lived then and may have known him. ... of course, since you know who wrote them ... [knowledge comes from] some who have written, who have been concerned, right? or from people who may have experienced him in person. let’s not forget the [greek] civil war ... people ... who lived during the civil war and wrote, you will see that one writes differently, the other writes differently, depending on the different political and social beliefs. (stru*m) going to history education, he directly connected what he already revealed about himself with what his students should achieve through this subject: “[students] should search on their own, see and understand that history shows us specific historical events, but it also hides the subjective element, that everyone writes as they wish.” regarding ibl, he emphasized the active role of students by developing their own opinion and judging what they read, based on trustworthy sources: [ibl is] literature searching, and from the internet ... but not from random sites ... that is, maybe with my own guidance with some historians’ names who are really historians ... [students should use sources] which are approved and from there on, through discussion, what everyone believes, to express their opinion, and so forth. that is, to become a historian themselves. ... to understand that the historian will also write their opinion in their own way. (evaluating) in sum, teacher 1 apparently had different epistemic beliefs about the two topics, which are considered quite representative examples of absolutism in science and multiplism in history. furthermore, the way he appeared to deal with knowledge within the context of the two topics was also reflected in his conceptions of ibl in science and history education. 5.3.2. teacher 12: ibl as “minds-on” experiencing in science and investigating in history teacher 12 was 43 years old and she had been teaching at primary schools for 18 years, after finishing her 4-year bachelor studies in primary education. she had taught in the 6th grade for three years. in parallel, she had completed her master studies in the same discipline, during which time she had participated in a qualitative research project. her interview started with the historical topic and initially, she had some difficulty involving herself in the task by asking questions regarding the historical accounts and whether they were appropriate for that age group. this reaction was quite common among participants, given that they expected to be addressed merely as teachers. after finding her role in this task as a learner, she took a more active role, concluding that “both of them could be equally right or wrong ... depending on who is studying them.” when she was asked to describe how she would go about forming her own opinion, she replied in a way indicative of the belief that knowledge is generated from the self: koutsianou et emvalotis 56 | f l r perhaps, i would study ... more sources and texts, and i would have a more complete picture. ... it’s like truth, it has many aspects ... one topic can, each of us can see it from many perspectives depending on our knowledge, level of studies, work and family. i think that each of us ... has a lot of perspectives, so [imagine the scenario] where the perspectives of many are collected. (sour*m) complementary, she expressed multiplist beliefs about the justification of knowing by mentioning subjective criteria that she could apply to judge different sources: i imagine that you study ... you judge and accordingly you come to conclusions. the criteria now of what you think [is] closer to you. let’s say, i see the structure of a text, its vocabulary. something is ... perhaps more or less understandable. the style of the text, right? that is, you find elements that may be closer to you, more understandable ... i don’t know. (just*m) then, she elaborated on her beliefs about the nature of knowledge, indicating the (re)constructionist nature of historical knowledge, based on the existing sources, which may also result in some uncertainty: historians [may know] up to a point, i imagine. because ... we all learn from some sources, right? which are the traces of the past from which we try to draw the plot of the history ... so there certainly are gaps. in other words, [historians] don’t completely know kapodistrias, either as a person or his work as a whole. (cert*e, stru*e) regarding the structure of knowledge, she added that it consists of “a set of things ... an integration of different sources, of data, of an integration in general.” (stru*e) however, she struggled to answer the direct question about the organisation of knowledge, in terms of mentioning quality features. interestingly, when the interview focused on history education, she was even more willing to share her way of thinking and working with students. her definition of ibl was very close to ibl as investigating, even though she did not mention straightforwardly the importance of formulating questions about the past: [ibl] gives the first role to the student, right? to become a researcher, that is, not to hand learning on a plate, your knowledge, tradition ... you give them paths, you guide them and they search and construct knowledge on their own. ... i think that a student learns only when he/she is interested in something. that is yes, i am also interested and want to learn. hence, ibl opens up the way, in essence ... you give them the motivation to learn what they want to learn. (investigating) she continued with the scientific topic, which she characterised as much more difficult to understand compared to the historical one. from the start, she stated that she could not judge which account was right and which was wrong, without being asked, rather influenced by the previous topic. later, she concluded that if she had to make her own decision, she would search “outside” to find what is really the case: i would definitely look for more information. because they both seem just as right or wrong. ... i can’t say they are wrong because i don’t have evidence that they are wrong; or if they are right, of course. i would do my research ... [and] search to see what is really the case. (sour*a) moreover, she struggled to articulate other criteria for evaluating different sources, except for what makes sense based on logic: i imagine that it will be ... okay, i should be more confident, [and] have more knowledge to be able to say, yes, this seems to be more logical ... more right? i don’t know! but i would definitely look for it. i mean, i couldn’t, i can’t judge now which of the two, let’s say. (sour*a, just*m) apparently, she could not involve herself in a process of meaning-making and she expected to find the right answer in the literature. on the other hand, based on the two accounts, she doubted that there will be an answer for this topic in the future: if i judge ... here from the two studies on the same type of bird ... we can see a third [study] that tells us something completely different, right? i don’t know whether we could come up with one koutsianou et emvalotis 57 | f l r opinion that would be considered correct. i imagine that here the same is true as before, that the opinions are more. ... the robin is most expert; if they could tell us, they would tell us! (cert*m) about the structure of knowledge, she was coded as both absolutist and evaluativist due to her evasive answers, which could indicate both her difficulty in articulating her beliefs and addressing such a specialised topic: in order to know things about the robin, you must have studied them ... through information, but also to study them alive as the researchers do here, who experiment... (stru*a) okay, i imagine data at first, right? opinions may now be different and the interpretations. ... i guess that knowledge is based more on something that is proven ... now, interpretations? but that’s how we end up with interpretations, yes. again, i will say that it is an integration of all this! (stru*e) she often noted the similarity of her thoughts with the previous topic, revealing that she addressed knowledge about both topics in similar terms. continuing with the science education, she set several goals for her students, which are illustrative of ibl as “minds-on” experiencing, by further elaborating when she was asked to define ibl: i think that children as physics ... they cannot understand even what the subject includes. ... the only thing that [children] can say is ... that we are doing experiments. but okay, there is a lot behind it, which must ... be discovered by children and based on that, they must acquire appropriate skills. [such as] ... a way of thinking a little different, to understand ... what it is, the so-called natural environment. the basic vocabulary that concerns the natural environment. ... its scientific ... profile. how all this is connected to the children themselves ... okay, whatever we know about our natural world is through experiments, through investigations, right? so, we cannot ignore this information completely from the way children learn. how else will they learn? if they don’t investigate ... or experiment? (“minds-on” experiencing) in sum, teacher 12 showed a rather complex set of epistemic beliefs while dealing with both topics and in parallel, described ibl in both domains with an emphasis on how crucial it is that students learn to think more scientifically by developing important skills through inquiry. although she appeared quite aware of how to organise a lesson through ibl, she struggled to involve herself in an active process of knowing for such specialised topics. 5.3.3. teacher 11: ibl as investigating in both history and science teacher 11 was 37 years old, and was both a computer science teacher in secondary education and a primary school teacher, with master studies in both disciplines. during his master studies, he had participated in both qualitative and quantitative research projects; and at the same time, he had been teaching computer science for the first 7 years of his teaching career and as a primary school teacher for the last 5 years. he had taught in the 6th grade for one school year. his interview started with the historical topic and he immediately involved himself in an active process of meaning-making regarding the two historical accounts. when he was asked how he would go about making his decision on the topic, he replied: we should investigate more sources that present the issue from all sides and then, with critical thinking and emotional distancing, to be able ... to draw conclusions that are as much as possible... objective and based on sources. but everything has to be ... investigated, somehow. (sour*e) when he was asked to mention specific criteria that he could use when it comes to evaluating multiple sources, he focused mainly on how well-justified they are, based on evidence: whether they are justified ... and what is their validity ... for something to be considered valid ... it must be based on sources which have a certain power. and not to ... be a product of someone’s imagination or subjectivity in a bad way. (just*e) koutsianou et emvalotis 58 | f l r focusing on the certainty of knowledge about this topic, he believed that historians should investigate further “by applying specific methodology to clarify, as much as possible, the specific historical period and the specific historical figure”, implying the possibility of improving the degree of certainty about the topic: i think that if there is a fruitful discourse between scientists who have followed certain scientific methods and ... principles, i think that even this opposition and multifaceted consideration of this historical figure ... can lead to useful conclusions, so we can rely on them with a greater percentage of certainty. (cert*e) interestingly, he didn’t face any difficulty with the questions about the structure of knowledge; instead, he gave insightful descriptions mentioning both different kinds and quality features of knowledge, revealing the complex nature of this dimension. knowledge consists of many components ... some of them are about events, others about the results recorded ... which of course are presented sometimes in different ways ... but the knowledge that has to do with the investigation of the causes, ... due to an event that happened or because political decisions had been made, ... the interpretation of [them] in some way... is very much influenced by the researcher’s hypotheses. then, this part is also knowledge, but perhaps it is a knowledge which ‘must be constantly doubted’ and ... needs to be renegotiated in order to ... bring it closer to reality. ... the structure of knowledge is complex, dynamic, that is non-static, constantly changing ... (stru*e) additionally, when he was asked about the source of knowledge for those who investigate this topic, he brought out the interpretive nature of historical knowledge: knowledge is the result of ... the research method [the historians] have chosen, that is, they choose a research method, they process the data of the past, whatever form they take, and their processing and the correct questions that are asked in the past, have, as a result, the extraction of some ... conclusions, which in my opinion are knowledge. (stru*e) then, the interview focused on history education. even though he noted many differences between historical inquiry and school history, he articulated interesting goals that students should meet through this subject. these goals strongly reflected his epistemic beliefs. he stated that students “... should develop their critical knowledge and learn to argue based on the sources and historical events with which they come into contact, and be able to find some relevance to their own everyday life.” finally, his definition of ibl was the most illustrative of the investigating category: [ibl] presupposes the active action of the students ... with the teacher’s guidance, students ... construct knowledge through inquiry, search for sources, analyse information, answer questions on their own and, synthesize and produce knowledge. (investigating) the interview continued with the scientific topic. while he was dealing with the two scientific accounts, he noted that “both of them are attempts to interpret a natural phenomenon, and they are in completely opposite directions; ... [researchers] interpret [the phenomenon] in a completely different way; they attempt to attribute orientation to different mechanisms. ... but both with plausible arguments.” when he was asked to explain his own way of making a decision, he emphasized the importance of being interested in the topic in order to: ... try to have access to ... scientific data and to ... scientific sources in order to see the conclusions of studies that have been done for the specific issue, let’s say, of the orientation, so i would try to make a comparison firstly, an investigation of the conclusions, to have a more ... integrated opinion. (sour*e) when he was asked to name specific criteria that he would adopt to evaluate different studies, he implied that he might choose the study which would further convince him, based on its justification: i would investigate it more, i could ... do an investigation of the bibliography, the research that has been done around it... based on which one would convince me more, let’s say, from what point of koutsianou et emvalotis 59 | f l r view they see it ... every researcher. i could try to identify a ... study that focused somewhere ... in order to see what the conclusions of other studies say about it. (just*e) regarding the certainty of knowledge about this topic, he was quite hesitant to give a straightforward answer, but his description implied the evolving nature of knowledge through experts’ attempts to interpret this phenomenon to a greater extent (cert*e). also, he mentioned quite often the case of the reinterpretation of the given findings in light of new research. about the structure of knowledge, he said that knowledge “consists basically of all the conclusions of ... the studies, of all the findings of the studies, of the analysis of the behaviour that has emerged from the observation of the robins; and perhaps of comparative studies between these findings.” finally, he commented that in correspondence with the historical topic, knowledge “is not simple, it is complex ... dynamic in nature ... changing and can be interpreted in many ways.” (stru*e) comparing scientific inquiry and school science, he acknowledged many more commonalities than in history, mainly because of the frequent use of experiments in the classrooms, following how scientists work methodologically. he elaborated further by describing specific goals that students should meet through this subject, based on the approach of ibl: students can be motivated to construct knowledge in a safe context, where the teacher’s role is guiding. i think it has multiple benefits from many aspects, both concerning the cognitive outcomes and the skills that they develop, collaboration and communication, ... and their familiarisation with the scientific process. and basically, it is knowledge gained by students’ actions, ... they discover, investigate, learn to compare, learn to think critically, be reflective and discuss with themselves. (investigating) in sum, teacher 11 was the only teacher who clearly articulated evaluativist beliefs while dealing with both topics in history and science, and his conceptions of ibl fell into the category investigating for both domains. hence, it seems that his epistemic beliefs are reflected in the way he conceives ibl through the goals he sets for his students and the skills he believes they need to acquire in the context of these subjects. 6. discussion this study aimed to investigate qualitatively, through scenario-based semi-structured interviews, the interplay between primary school teachers’ topic-specific epistemic beliefs and their conceptions of ibl in history and science education. overall, the findings of this study extend prior research from different fields, support the nuanced approach of teachers’ epistemic beliefs and the more inductive analysis of teachers’ conceptions of ibl; and, further delineate the magnitude of their interplay. a discussion of the key findings of this study follows in light of the existing literature by offering theoretical and empirical implications, the study’s limitations, and educational implications. 6.1. a nuanced approach of epistemic beliefs through an integrative framework this study provides evidence about how the complex nature of teachers’ epistemic beliefs emerged while dealing with two controversial topics in history and science; and further, acknowledges the importance of applying a nuanced approach to detect unique epistemic belief patterns rather than simply categorising epistemic beliefs as “naïve” or “sophisticated” per dimension (barzilai & weinstock, 2015; feucht, 2011, 2017; mason, 2016; merk et al., 2018). this study contributes to the existing literature by giving voice, for the first time, to primary school teachers to articulate their epistemic beliefs within a task of two very different topics. additionally, the proposed integrative framework could be seen as a useful tool for further conceptual and empirical research on the construct of epistemic beliefs. koutsianou et emvalotis 60 | f l r specifically, most participants aligned with more than one epistemic perspective with regard to different dimensions while dealing with each topic (feucht, 2011, 2017); however, they could be included in a predominant perspective (barzilai & weinstock, 2015; king & kitchener, 1994). moreover, there were teacher cases who were included in a single perspective for all dimensions and others who could be in transition between two perspectives (barzilai & weinstock, 2015; feucht, 2011, 2017; kuhn cheney, & weinstock, 2000). this study did not aim to strictly categorise participants’ epistemic beliefs by giving frequency data; on the contrary, it attempted to give thorough descriptions of teachers’ epistemic belief patterns in order to challenge and extend the existing frameworks (greene & yu, 2014). hence, a scenario-based approach followed by open-ended questions in the form of semistructured interviews was considered to be the best option to gain full access to teachers’ wording and thoughts. teachers’ epistemic beliefs were probed within the context of a task where they were asked to deal with controversial topics (barzilai & weinstock, 2015; vansledright & maggioni, 2016); in an attempt to reduce any possible difficulty in answering more general questions about academic domains (merk et al., 2018). similar to barzilai and weinstock’s (2015) conclusion, however, in many teacher cases, their domain-specific epistemic beliefs emerged and interfered while dealing with the specific topics, verifying that topicand domain-specific beliefs are not independent (see merk et al., 2018 for a detailed discussion). based on the findings of this study, there are some important remarks about the conceptualisation and operationalisation of this construct that should be made. first, most teachers struggled to understand and answer the direct questions about the structure of knowledge for both topics, and they often asked for further clarification. second, their responses to these questions were very similar to their responses to questions about the source of experts’ knowledge, which they also addressed with some difficulty, although less often. therefore, it seems that primary school teachers may not be wondering what knowledge is and where it resides in their everyday lives; quite reasonably, they had difficulty understanding these questions and finding the appropriate words to describe their thoughts. a different interpretation of this finding could be that the applied questions failed to adequately capture participants’ epistemic beliefs about the structure of knowledge. but there were teacher cases who managed to thoroughly articulate their beliefs (e.g., teacher 11). third, some participants discriminated between different kinds of knowledge, especially in history (e.g., the events from their potential causes), highlighting the complex nature of this dimension and the importance of recognising the different kinds of knowledge (e.g., declarative, procedural, and conceptual) in the conceptualisation and assessment of epistemic beliefs, verifying, therefore, greene and yu’s (2014) conclusions. fourth, teachers’ beliefs about the source of knowledge elicited more indirectly by asking them how they would go about making their own decisions on the topics (king & kitchener, 1994), instead of asking them about the source of knowledge for experts who investigate these issues (barzilai & weinstock, 2015). fifth, teachers’ beliefs about the source of knowledge and the justification of knowing were found to be distinct but also complementary dimensions, pertaining more to teachers’ beliefs about their involvement in knowledge production and their evaluation of different knowledge claims (feucht, 2011, 2017) instead of how experts develop and establish knowledge in their fields. nevertheless, teachers’ difficulty in articulating epistemic criteria for evaluating different accounts could be a clear indication of their unfamiliarity with experts’ practices. but this is only a speculation that needs to be explored. 6.2. primary school teachers’ conceptions of ibl in history and science education this study offers a comprehensive overview of how primary school teachers may conceive ibl in both history and science education. actually, this is the first study that explored primary school teachers’ conceptions of ibl in history and compared them with their conceptions of ibl in science. regarding history education, the findings of this study provide clear evidence for the applicability of the categorisation proposed by voet and de wever (2016) in their study with secondary history teachers. primary school teachers’ conceptions of ibl in history education fell into the following three categories: understanding, evaluating, and investigating. however, the content of these categories was elaborated based on participants’ responses. on the other hand, primary school teachers’ conceptions of ibl in koutsianou et emvalotis 61 | f l r science education did not exactly follow any categorisation of the existing literature. based more on participants’ descriptions of ibl and less on existing coding schemes (ireland et al., 2012; voet & de wever, 2016), three categories were also identified, which were named “hands-on” experiencing, “minds-on” experiencing and investigating. moreover, this study adopted a broader view of ibl in science education by directly asking participants to consider the case of ibl in science as a literature search and to use multiple sources regarding a given topic. hence, within the above three categories, there were also included teachers’ conceptions of ibl in science through this perspective; which were found to resemble a lot the three categories of ibl in history. a general remark is that even though most participants gave interesting and rich descriptions of ibl in both history and science education, very few teachers defined ibl in terms of doing full historical and scientific inquiries within the classroom. similar to the existing literature (e.g., ireland et al., 2012; kang et al., 2008; voet & de wever, 2016), primary school teachers noted very rarely the importance of engaging their students in formulating research questions and argumentation as part of ibl in both history and science. nevertheless, this study endeavoured to explore how primary school teachers understand and conceptualise ibl through their own wording, rather than capturing the frequency of these conceptions among the participants. also, these findings do not indicate whether or not these teachers attempt to foster ibl in their classrooms, even though their school experiences may have affected their descriptions (see voet & de wever, 2016 for a detailed description). besides, according to levy et al. (2013, p. 396), “understanding the process of ibl [in] history [and science] is necessary but not sufficient for using it in the classroom”. 6.3. unravelling the interplay of primary school teachers’ epistemic beliefs and conceptions of ibl this study contributes uniquely to the existing literature, by providing clear indications that primary school teachers’ conceptions of ibl, in history and science education, are aligned well with their epistemic belief patterns per topic. in comparison with voet and de wever’s (2016) findings, with reference to history education, this study found a greater connection between primary school teachers’ epistemic beliefs and their conceptions of ibl, probably due to the much more nuanced approach of teachers’ epistemic beliefs. further, even though comparisons of this study with studies from science education (e.g., chan, 2011; lee & tsai, 2011; tsai, 2002; see 2.2. and 2.3. sections) are difficult to make due to the different theoretical constructs involved, this study provides evidence in support of the common claim that teachers who hold more availing epistemic beliefs, also support more “constructivist approaches”; however, in terms of describing ibl more thoroughly in science education. regarding history, there is evidence that teachers who held predominantly absolutist epistemic beliefs defined ibl mainly as understanding; and, teachers who held predominantly evaluativist epistemic beliefs defined ibl mainly as investigating. besides, the epistemic belief patterns of participants who described ibl as evaluating were much more complex, but all of them held evaluativist beliefs about the source of knowledge and multiplist beliefs about the justification of knowing. overall, it could be argued that the more availing the teachers’ epistemic beliefs, the more thoroughly they conceive ibl in history education. in regards to the evaluating category of ibl, it seems that teachers who maintain an active role in the process of knowing, support corresponding goals for their students in history by focusing on evaluating sources and developing critical thinking skills, even though they adopt subjective epistemic criteria. some teachers of this category maintained a more dogmatic view about the nature of historical knowledge (i.e., it should exclusively consist of data/facts). the latter could be seen as an indication of the greater connection between teachers’ beliefs about the nature of knowing and their conceptions of ibl in history. but this hypothesis needs to be further examined. regarding science, there is evidence that teachers who held predominantly absolutist epistemic beliefs, defined ibl as “hands-on” experiencing. however, teachers who held predominantly evaluativist epistemic beliefs defined ibl either as “minds-on” experiencing or investigating. furthermore, the epistemic belief patterns of participants who defined ibl as “minds-on” experiencing were much more complex, with a common tendency to express evaluativist beliefs about the nature of koutsianou et emvalotis 62 | f l r knowledge, almost in all cases. similarly, many of these teachers held absolutist beliefs about the source of knowledge and multiplist beliefs about the justification of knowing. overall, it could be argued that the more availing the teachers’ epistemic beliefs, the more thoroughly they conceive ibl in science education. however, in the “minds-on” experiencing category, the situation is rather more complex and different. almost all these teachers endorsed the evolving and complex nature of scientific knowledge. however, they mainly described themselves as passive receivers of such specialised knowledge, without articulating other epistemic criteria except for what makes sense based on their prior experiences with magnets. this finding is very interesting because even though these teachers struggled to involve themselves in the process of knowing, they were also quite aware of what ibl is in science education. this might be explained by the more frequent use of ibl, both as a term and instructional goal in science rather than in history education, based on greek primary school textbooks. in conclusion, this study reveals that there is a complex interplay between: a) how primary school teachers view knowledge and engage themselves as learners in the process of knowing; b) how teachers conceptualise ibl in order to engage their students in the process of knowing. additionally, this interplay seems to differ in history and science, but further research is needed to verify these early findings. also, it could be argued that there is a theoretical overlap between the two constructs (i.e., epistemic beliefs and conceptions of ibl) given that inquiry is a way of knowing and, consequently, teachers’ conceptions of ibl could be also addressed as an epistemic construct. even though this is certainly a foundation for further research, the findings of this study provide ample evidence of the need to separately measure teachers’ epistemic beliefs as learners and their conceptions of ibl as teachers; in order to capture in-depth teachers’ internal factors that might motivate or prevent them from engaging their students in ibl. 6.4. limitations several limitations follow this qualitative study and should be mentioned. first, the results of this study cannot be generalised beyond its context and they should be interpreted with caution, due to its sampling method. therefore, further research is needed to replicate the usefulness of the integrative framework of epistemic beliefs and the coding frameworks of teachers’ conceptions of ibl. additionally, more research is needed to learn the frequency of the epistemic belief patterns and teachers’ conceptions of ibl found in the broader population of primary school teachers. second, the rather small sample size in terms of statistical analysis and the exploratory nature of this study did not allow for an examination of the nature of this interplay. thus, future studies could investigate whether this is a causal relationship and further, its potential predictive utility regarding teachers’ decisions and practices in terms of ibl. third, teachers’ epistemic beliefs were explored within two specific topics from history and science; hence, research with additional topics is required to verify and extend these results. fourth, even though this study aimed to elicit teachers’ epistemic beliefs and conceptions of ibl and not their practices, the exclusive use of interviews, even in the context of a task, could raise some reasonable criticism regarding the need for triangulation in both data collection and analysis. future studies could apply additional techniques, such as observations of teachers’ classroom behaviours and/or member checking. 6.5. educational implications the findings of this study raise some important implications for both teacher educators and educational policymakers in greece concerning both preand in-service primary school teachers, and on an international level, by taking into account the different cultural environments. even though this study does not reflect the proportion of greek primary school teachers who fit in the aforementioned epistemic belief patterns and conceptions of ibl, it is somehow worrisome that a very small number of the participants: a) expressed predominantly evaluativist epistemic beliefs; and, b) described ibl, comprehensively. reflecting on these findings, the following crucial question emerged: if one of the koutsianou et emvalotis 63 | f l r main aims of contemporary primary education is to engage students in experts’ epistemic practices from different domains, in order to learn how to make informed decisions in such an epistemically challenging society, then, how can we best prepare teachers to engage their students in ibl within and across such different subjects as history and science? starting with the education of pre-service primary school teachers, it is crucial that they get acquainted with the epistemology underlying different academic domains such as literacy, mathematics, history and science; and furthermore, to get involved in different ways of knowing within and across the corresponding school subjects. therefore, teacher education programmes should take care of including such crucial elements, both explicitly through the content of the courses provided, and implicitly through teacher educators’ practices (greene & yu, 2016; guilfoyle et al., 2020; voet & de wever, 2016; sinatra & hofer, 2016). within such educational contexts, pre-service teachers can be triggered to reflect on their epistemic beliefs and conceptions of ibl, and their interconnections (fives et al., 2015), and acquire determining experiences of planning and implementing ibl in action (ireland et al., 2012; levy et al., 2013; martell, 2020). at the same time, in-service teachers should be supported in similar ways through their participation in professional development programmes, after having taken further care of their more established beliefs and conceptions, due to their long teaching experience (fives & buehl, 2012, 2016). the recent conceptual work on teachers’ epistemic cognition in action while learning to teach, and in teaching action, incorporates many of the elements mentioned above (e.g., buehl & fives, 2016; feucht et al., 2017; hagani & barzilai, 2018; lunn brownlee et al., 2017). footnote: 1. although the interchangeable use of the terms “belief” and “conception” is quite often found in such a large and interdisciplinary research field as teachers’ beliefs, followed by conceptual issues and measurement limitations, the present study: a) used these terms in the theoretical background based on the cited literature each time; and, b) adopted the term “conception” to approach how primary school teachers conceive ibl in history and science education based on the existing literature (ireland et al., 2012; voet & de wever, 2016). acknowledgments: the authors would like to thank lida desikou for acting as a critical reader of this manuscript and for her assistance in transcribing the interview data. this research was financially supported by the general secretariat for research and technology (gsrt) and the hellenic foundation for research and innovation (hfri) under the hfri phd fellowship grant (fellowship number: 231) allocated to the first author. keypoints based on an integrative framework, a nuanced qualitative approach of teachers’ epistemic beliefs was applied. primary school teachers’ epistemic beliefs were elicited while they were dealing with controversial topics both in history and science. important domain differences were detected in primary school teachers’ conceptions of inquiry-based learning. primary school teachers’ epistemic beliefs and conceptions of inquiry-based learning are interconnected in both history and science. koutsianou et emvalotis 64 | f l r references avraamidou, l. (2017). a well-started beginning elementary teacher’s beliefs and practices in relation to reform recommendations about inquiry-based science. cultural studies of science education, 12(2), 331-353. doi: 10.1007/s11422-015-9700-x bartos, s. a., & lederman, n. g. (2014). teachers' knowledge structures for nature of science and scientific inquiry: conceptions and classroom practice. journal of research in science teaching, 51(9), 1150-1184. doi: 10.1002/tea.21168 barzilai, s., & weinstock, m. (2015). measuring epistemic thinking within and across topics: a scenario-based approach. contemporary educational psychology, 42, 141-158. doi: 10.1016/j.cedpsych.2015.06.006 bendixen, l., winsor, d. frazier, r. (2017). exploring bloom’s taxonomy as a bridge to evaluativism: conceptual clarity and implications for learning, teaching, and assessing. in g. schraw, j. lunn brownlee, l. olafson and m. vanderveldt (eds.), teachers’ personal epistemologies: evolving models for informing practice (pp. 191-211). charlotte, nc: information age publishing. buehl, m. m., & beck, j. s. (2015). the relationship between teachers' beliefs and teachers' practices. in h. fives & m. g. gill (eds.), international handbook of research on teachers’ beliefs (pp. 66-84). new york, ny: routledge. buehl, m. m., & fives, h. (2009). exploring teachers' beliefs about teaching knowledge: where does it come from? does it change? the journal of experimental education, 77(4), 367-408. doi: 10.3200/jexe.77.4.367-408 buehl, m. m., & fives, h. (2016). the role of epistemic cognition in teacher learning and praxis. in j. a. greene, w. a. sandoval, & i. bråten (eds.), handbook of epistemic cognition (pp. 247-264). new york, ny: routledge. chan, k.-w. (2011). preservice teacher education students’ epistemological beliefs and conceptions about learning. instructional science, 39(1), 87-108. doi: 10.1007/s11251-009-9101-1 cheng, m. m. h., chan, k.-w., tang, s. y. f., & cheng, a. y. n. (2009). pre-service teacher education students' epistemological beliefs and their conceptions of teaching. teaching and teacher education, 25(2), 319-327. doi: 10.1016/j.tate.2008.09.018 chi, m. t. h. (1997). quantifying qualitative analyses of verbal data: a practical guide. journal of the learning sciences, 6(3), 271-315. doi: 10.1207/s15327809jls0603_1 chinn, c. a., buckland, l. a., & samarapungavan, a. l. a. (2011). expanding the dimensions of epistemic cognition: arguments from philosophy and psychology. educational psychologist, 46(3), 141-167. doi: 10.1080/00461520.2011.587722 cohen, l., manion, l., & morrison, k. (2007). research methods in education (6th ed.). usa: routledge. creswell, j. w. (2015). educational research: planning, conducting, and evaluating quantitative and qualitative research (5th ed.). boston: pearson. dobber, m., zwart, r., tanis, m., & van oers, b. (2017). literature review: the role of the teacher in inquiry-based education. educational research review, 22, 194-214. doi: 10.1016/j.edurev.2017.09.002 european commission. (2015). science education for responsible citizenship. brussels: directorategeneral for research and innovation, science with and for society. http://ec.europa.eu/research/swafs/pdf/pub_science_education/ki-na-26-893-en-n.pdf feucht, f. c. (2011). the epistemic underpinnings of mrs. m's reading lesson on drawing conclusions: a classroom-based research study. in j. brownlee, g. schraw, & d. berthelsen (eds.), personal epistemology and teacher education (pp. 227-245). new york: routledge. feucht, f. c. (2017). the epistemic climate of mrs. m’s science lesson about the woodlands as an ecosystem: a classroom-based research study. in g. schraw, j. lunn brownlee, l. olafson and m. vanderveldt (eds.), teachers’ personal epistemologies: evolving models for informing practice (pp. 55-84). charlotte, nc: information age publishing. koutsianou et emvalotis 65 | f l r feucht, f. c., lunn brownlee, j., & schraw, g. (2017). moving beyond reflection: reflexivity and epistemic cognition in teaching and teacher education. educational psychologist, 52(4), 234241. doi: 10.1080/00461520.2017.1350180 fives, h., & buehl, m. m. (2012). spring cleaning for the "messy" construct of teachers' beliefs: whay are they? which have been examined? what can they tell us? in k. harris, s. graham, t. urdan, c. mccormick, g. m. sinatra, & j. sweller (eds.), apa educational psychology handbook, vol 2: individual differences and cultural and contextual factors (pp. 471-499). washington, dc: american psychological association. fives, h., & buehl, m. m. (2016). teachers’ beliefs, in the context of policy reform. policy insights from the behavioral and brain sciences, 3(1), 114-121. doi: 10.1177/2372732215623554 fives, h., & buehl, m. m. (2017). the functions of beliefs: teachers’ personal epistemology on the pinning block. in g. schraw, j. lunn brownlee, l. olafson and m. vanderveldt (eds.), teachers’ personal epistemologies: evolving models for informing practice (pp. 25-54). charlotte, nc: information age publishing. fives, h., lacatena, n., & gerard, l. (2015). teachers’ beliefs about teaching (and learning). in h. fives & m. g. gill (eds.), international handbook of research on teachers’ beliefs (pp. 249265). new york, ny: routledge. forbes, c. t., biggers, m., & zangori, l. (2013). investigating essential characteristics of scientific practices in elementary science learning environments: the practices of science observation protocol (p-sop). school science and mathematics, 113(4), 180-190. doi: 10.1111/ssm.12014 gillies, r. m., & nichols, k. (2015). how to support primary teachers’ implementation of inquiry: teachers’ reflections on teaching cooperative inquiry-based science. research in science education, 45(2), 171-191. doi: 10.1007/s11165-014-9418-x greene, j. a., sandoval, w. a., & bråten, i. (2016a). an introduction to epistemic cognition. in j. a. greene, w. a. sandoval, & i. bråten (eds.), handbook of epistemic cognition (pp. 1-16). new york, ny: routledge. greene, j. a., sandoval, w. a., & bråten, i. (2016b). handbook of epistemic cognition. new york, ny: routledge. greene, j. a., sandoval, w. a., & bråten, i. (2016c). reflections and future directions. in j. a. greene, w. a. sandoval, & i. bråten (eds.), handbook of epistemic cognition (pp. 495-510). new york, ny: routledge. greene, j. a., & yu, s. b. (2014). modeling and measuring epistemic cognition: a qualitative reinvestigation. contemporary educational psychology, 39(1), 12-28. doi: 10.1016/j.cedpsych.2013.10.002 greene, j. a., & yu, s. b. (2016). educating critical thinkers: the role of epistemic cognition. policy insights from the behavioral and brain sciences, 3(1), 45-53. doi: 10.1177/2372732215622223 greene, j. a., yu, s. b., & copeland, d. z. (2014). measuring critical components of digital literacy and their relationships with learning. computers & education, 76, 55–69. doi: 10.1016/j.compedu.2014.03.008 guilfoyle, l., mccormack, o., & erduran, s. (2020). the "tipping point" for educational research: the role of pre-service science teachers’ epistemic beliefs in evaluating the professional utility of educational research. teaching and teacher education, 90, 103033. doi: 10.1016/j.tate.2020.103033 hagani, s. m. & barzilai, s. (2018). reflecting on epistemic ideals and processes: designing opportunities for teachers’ epistemic growth. in j. kay, & r. luckin (eds.) rethinking learning in the digital age: making the learning sciences count, 13th international conference of the learning sciences (icls) 2018, vol. 2. london, uk: international society of the learning sciences. retrieved from https://doi.dx.org/10.22318/cscl2018.1129 hammer, d., & elby, a. (2002). on the form of a personal epistemology. in b. k. hofer & p. r. pintrich (eds.), personal epistemology: the psychology of beliefs about knowledge and knowing (pp. 169-190). lawrence erlbaum associates, inc. hofer, b. k. (2000). dimensionality and disciplinary differences in personal epistemology. contemporary educational psychology, 25(4), 378-405. doi: 10.1006/ceps.1999.1026 koutsianou et emvalotis 66 | f l r hofer, b. k. (2001). personal epistemology research: implications for learning and teaching. educational psychology review, 13(4), 353-383. doi: 10.1023/a:1011965830686 hofer, b. k. (2016). epistemic cognition as a psychological construct: advancements and challenges. in j. a. greene, w. a. sandoval, & i. bråten (eds.), handbook of epistemic cognition (pp. 1938). new york, ny: routledge. hofer, b. k., & bendixen, l. d. (2012). personal epistemology: theory, research, and future directions. in k. r. harris, s. graham, & t. urdan (eds.), apa educational psychology handbook: theories, constructs, and critical issues (vol. 1, pp. 227-256). washington, dc: american psychological association. hofer, b. k., & pintrich, p. r. (1997). the development of epistemological theories: beliefs about knowledge and knowing and their relation to learning. review of educational research, 67(1), 88-140. doi: 10.3102/00346543067001088 hofer, b. k., & pintrich, p. r. (2002). personal epistemology: the psychology of beliefs about knowledge and knowing: lawrence erlbaum associates, inc. holland, r. a., & helm, b. (2013). a strong magnetic pulse affects the precision of departure direction of naturally migrating adult but not juvenile birds. journal of the royal society interface, 10(81). doi: 10.1098/rsif.2012.1047 hsieh, h.-f., & shannon, s. e. (2005). three approaches to qualitative content analysis. qualitative health research, 15(9), 1277-1288. doi: 10.1177/1049732305276687 ireland, j. e., watters, j. j., brownlee, j., & lupton, m. (2012). elementary teacher’s conceptions of inquiry teaching: messages for teacher development. journal of science teacher education, 23(2), 159-175. doi: 10.1007/s10972-011-9251-2 kang, n.-h. (2008). learning to teach science: personal epistemologies, teaching goals, and practices of teaching. teaching and teacher education, 24(2), 478-498. doi: 10.1016/j.tate.2007.01.002 kang, n.-h., orgill, m., & crippen, k. j. (2008). understanding teachers’ conceptions of classroom inquiry with a teaching scenario survey instrument. journal of science teacher education, 19(4), 337-354. doi: 10.1007/s10972-008-9097-4 king, p. m., & kitchener, k. s. (1994). developing reflective judgment: understanding and promoting intellectual growth and critical thinking in adolescents and adults. san francisco, california: jossey-bass publishers. kremmydas, v. (2015). η διακυβέρνηση καποδίστρια: κοινωνία, πολιτική, ιδεολογία. [the governance of kapodistrias: society, policy, ideology.] in g. georgis (ed.) ο κυβερνήτης ιωάννης καποδίστριας: κριτικές προσεγγίσεις και επιβεβαιώσεις [the governor ioannis kapodistrias: critical approaches and confirmations] (pp. 46-53). athens: kastaniotis. [in greek] kuhn, d. (1999). a developmental model of critical thinking. educational researcher, 28(2), 16-46. retrieved from www.jstor.org/stable/1177186 kuhn, d., cheney, r., & weinstock, m. (2000). the development of epistemological understanding. cognitive development, 15(3), 309-328. doi: 10.1016/s0885-2014(00)00030-7 kuhn, d., iordanou, k., pease, m., & wirkala, c. (2008). beyond control of variables: what needs to develop to achieve skilled scientific thinking? cognitive development, 23(4), 435-451. doi: 10.1016/j.cogdev.2008.09.006 kuhn, d., & weinstock, m. (2002). what is epistemological thinking and why does it matter? in b. k. hofer & p. r. pintrich (eds.), personal epistemology: the psychology of beliefs about knowledge and knowing (pp. 121-144). mahwah, nj: erlbaum. lakin, j. m., & wallace, c. s. (2015). assessing dimensions of inquiry practice by middle school science teachers engaged in a professional development program. journal of science teacher education, 26(2), 139-162. doi: 10.1007/s10972-014-9412-1 lee, m-h., tsai, c.-c. (2011). teachers’ scientific epistemological views, conceptions of teaching science, and their approaches to teaching science: an exploratory study of in-service science teachers in taiwan. in j. brownlee, g. schraw, & d. berthelsen (eds.), personal epistemology and teacher education (pp. 246-262). new york: routledge. koutsianou et emvalotis 67 | f l r levy, b. l. m., thomas, e. e., drago, k., & rex, l. a. (2013). examining studies of inquiry-based learning in three fields of education: sparking generative conversation. journal of teacher education, 64(5), 387-408. doi: 10.1177/0022487113496430 lunn brownlee, j., ferguson, l. e., & ryan, m. (2017). changing teachers' epistemic cognition: a new conceptual framework for epistemic reflexivity. educational psychologist, 52(4), 242-252. doi: 10.1080/00461520.2017.1333430 maggioni, l., vansledright, b., & alexander, p. a. (2009). walking on the borders: a measure of epistemic cognition in history. the journal of experimental education, 77(3), 187-214. doi: 10.3200/jexe.77.3.187-214 martell, c. c. (2020). barriers to inquiry-based instruction: a longitudinal study of history teachers. journal of teacher education, 71(3), 279-291. doi: 10.1177/0022487119841880 mason, l. (2016). psychological perspectives on measuring epistemic cognition. in j. a. greene, w. a. sandoval, & i. bråten (eds.), handbook of epistemic cognition (pp. 375-392). new york, ny: routledge. mayring, p. (2000). qualitative content analysis. forum: qualitative social research, 1(2). doi: 10.17169/fqs-1.2.1089 merk, s., rosman, t., muis, k. r., kelava, a., & bohl, t. (2018). topic specific epistemic beliefs: extending the theory of integrated domains in personal epistemology. learning and instruction, 56, 84-97. doi: 10.1016/j.learninstruc.2018.04.008 merriam, s. b. (2009). qualitative research: a guide to design and implementation. san francisco, ca: jossey-bass. muis, k. r. (2004). personal epistemology and mathematics: a critical review and synthesis of research. review of educational research, 74(3), 317–377. doi: 10.3102/00346543074003317 muis, k. r., bendixen, l. d., & haerle, f. c. (2006). domain-generality and domain-specificity in personal epistemology research: philosophical and empirical reflections in the development of a theoretical framework. educational psychology review, 18(1), 3-54. doi: 10.1007/s10648006-9003-6 muis, k. r., & foy, m. j. (2010). the effects of teachers’ beliefs on elementary students’ beliefs, motivation, and achievement in mathematics. in l. d. bendixen & f. c. feucht (eds.), personal epistemology in the classroom: theory, research and implications for practice. new york, ny: cambridge university press. national council for the social studies (ncss). (2010). national curriculum standards for social studies: a framework for teaching, learning, and assessment (silver spring, md). retrieved from http://www.socialstudies.org/standards national council for the social studies (ncss). (2018). national standards for the preparation of social studies teachers. retrieved from https://www.socialstudies.org/standards/nationalstandards-preparation-social-studies-teachers next generation science standards (ngss) lead states. (2013). next generation science standards: for states, by states. washington, dc: national academies press. national research council (nrc). (2012). a framework for k-12 science education: practices, crosscutting concepts, and core ideas. committee on a conceptual framework for new k-12 science standards. board on science education, division of behavioral and social sciences and education. washington, dc: the national academies press. olafson, l., & schraw, g. (2010). beyond epistemology: assessing teachers’ epistemological and ontological worldviews. in l. d. bendixen & f. c. feucht (eds.), personal epistemology in the classroom: theory, research and implications for practice (pp. 516-551). new york, ny: cambridge university press. pedaste, m., mäeots, m., siiman, l. a., de jong, t., van riesen, s. a. n., kamp, e. t., . . . tsourlidaki, e. (2015). phases of inquiry-based learning: definitions and the inquiry cycle. educational research review, 14, 47-61. doi: 10.1016/j.edurev.2015.02.003 ploumidis, s. g. (2015). το όραμα του ιωάννη καποδίστρια για το ελληνικό έθνος και την κοινωνία. [the vision of ioannis kapodistrias for the greek nation and society.] in g. georgis (ed.) ο κυβερνήτης ιωάννης καποδίστριας: κριτικές προσεγγίσεις και επιβεβαιώσεις [the governor koutsianou et emvalotis 68 | f l r ioannis kapodistrias: critical approaches and confirmations] (pp. 68-87). athens: kastaniotis. [in greek] sandoval, w. a. (2005). understanding students' practical epistemologies and their influence on learning through inquiry. science education, 89(4), 634-656. doi: 10.1002/sce.20065 schraw, g., lunn brownlee, j., olafson, l., & vanderveldt, m. (2017). teachers' personal epistemologies: evolving models for informing practice. charlotte, nc: information age publishing. seung, e., park, s., & jung, j. (2014). exploring preservice elementary teachers’ understanding of the essential features of inquiry-based science teaching using evidence-based reflection. research in science education, 44(4), 507-529. doi: 10.1007/s11165-013-9390-x sinatra, g. m. (2016). thoughts on knowledge about thinking about knowledge. in j. a. greene, w. a. sandoval, & i. bråten (eds.), handbook of epistemic cognition (pp. 479-491). new york, ny: routledge. sinatra, g. m., & hofer, b. k. (2016). public understanding of science. policy insights from the behavioral and brain sciences, 3(2), 245-253. doi: 10.1177/2372732216656870 tsai, c.-c. (2002). nested epistemologies: science teachers' beliefs of teaching, learning and science. international journal of science education, 24(8), 771-783. doi: 10.1080/09500690110049132 tsai, c.-c. (2007). teachers' scientific epistemological views: the coherence with instruction and students' views. science education, 91(2), 222-243. doi: 10.1002/sce.20175 van uum, m. s. j., verhoeff, r. p., & peeters, m. (2016). inquiry-based science education: towards a pedagogical framework for primary school teachers. international journal of science education, 38(3), 450-469. doi: 10.1080/09500693.2016.1147660 vansledright, b., & maggioni, l. (2016). epistemic cognition in history. in j. a. greene, w. a. sandoval, & i. bråten (eds.), handbook of epistemic cognition (pp. 128-146). new york, ny: routledge. voet, m., & de wever, b. (2016). history teachers' conceptions of inquiry-based learning, beliefs about the nature of history, and their relation to the classroom context. teaching and teacher education, 55, 57-67. doi: 10.1016/j.tate.2015.12.008 wiltschko, r., gehring, d., denzau, s., nießner, c., & wiltschko, w. (2014). magnetoreception in birds: ii. behavioural experiments concerning the cryptochrome cycle. the journal of experimental biology, 217(23), 4225–4228. doi: 10.1242/jeb.110981 windschitl, m. (2002). framing constructivism in practice as the negotiation of dilemmas: an analysis of the conceptual, pedagogical, cultural, and political challenges facing teachers. review of educational research, 72(2), 131-175. doi: 10.3102/00346543072002131 windschitl, m. (2004). folk theories of “inquiry:” how preservice teachers reproduce the discourse and practices of an atheoretical scientific method. journal of research in science teaching, 41(5), 481-512. doi: 10.1002/tea.20010 koutsianou et emvalotis 69 | f l r appendix a1: scenarios 1. history scenario (developed based on authentic historical studies) the governance of ioannis kapodistrias the following text developed on the occasion of chapter 17 of the history textbook of the 6th grade, entitled “ioannis kapodistrias and his work” (pp. 138-139), based on the following published secondary sources: 1. ploumidis, s. g. (2015). το όραμα του ιωάννη καποδίστρια για το ελληνικό έθνος και την κοινωνία. [the vision of ioannis kapodistrias for the greek nation and society.] in g. georgis (ed.) ο κυβερνήτης ιωάννης καποδίστριας: κριτικές προσεγγίσεις και επιβεβαιώσεις [the governor ioannis kapodistrias: critical approaches and confirmations] (pp. 68-87). athens: kastaniotis. [in greek] 2. kremmydas, v. (2015). η διακυβέρνηση καποδίστρια: κοινωνία, πολιτική, ιδεολογία. [the governance of kapodistrias: society, policy, ideology.] in g. georgis (ed.) ο κυβερνήτης ιωάννης καποδίστριας: κριτικές προσεγγίσεις και επιβεβαιώσεις [the governor ioannis kapodistrias: critical approaches and confirmations] (pp. 46-53). athens: kastaniotis. [in greek] introduction “in 1827, the third national assembly of troizina elected ioannis kapodistrias governor of greece. kapodistrias arrived in nafplio, the first capital of the greek state, taking over the government of a country that had emerged from a long fight while its inhabitants and especially the refugees were impoverished. the governor attempted to organise the state, improving its administration and economy. in order to achieve his goal, he concentrated all the authority on his face, postponing for two years the convening of the fourth national assembly. [...] the centralist governance of kapodistrias and his conflict with many local interests provoked the dissatisfaction of political groups that reacted to his policy. on september 27, 1831, kapodistrias was assassinated in nafplio, resulting in anarchy in the country.” (excerpt from the greek textbook of history entitled “history of the modern and contemporary world”). ioannis kapodistrias and his work have been the research field for many historians of modern and contemporary greek history. excerpts from the research work of two greek historians follow, explaining the governance of ioannis kapodistrias. account a excerpt from the chapter “the vision of ioannis kapodistrias for the greek nation and society” written by the historian s. g. ploumidis (2015), assistant professor of the national and kapodistrian university of athens. “in the collective consciousness of the greek people, the myth of kapodistrias is made around ‘barbajohn’ and all the symbolism and stereotypes that accompany the name of the father protector (hence, mavromichalaioi are presented as patricides); in kapodistrias, from the time of his arrival in greece, was given the status of messiah: the one who, in times of crisis, would undertake to restore trust and ensure the necessary continuity of the national ‘family’.1 hence, kapodistrias’ views were not coincidental and fragmentary, but together constituted a comprehensive and structured long-term political programme. kapodistrias, despite nepotism and an authoritarian way of exercising his governmental power, had a genuine interest in the popular classes, he abhorred client networks and wanted a small state with a minimal number of civil servants. he clearly detested the ‘insidious and ambitious who had been brought up in the muslim school’ [of administration] and ‘in the school of fanari of constantinople’, that is, the kotzabasids, and their ‘system’ of authority. and he envisioned the ‘rebirth of the people’ (régénération du peuple)2 in a completely different and unconventional way koutsianou et emvalotis 70 | f l r ‘from above’, without the alliance of the traditional pre-revolutionary elites: through the creation of a compact class of smallholder farmers and through hard productive labour. so, the basic belief of the governor was that the people are civilised with hoes and not with bayonets3. however, with his harsh adherence to ground-breaking innovation and his pro-popular authoritarianism, he limited himself we would even say that he self-trapped, in the passive consent of the landless farmers, who did no weigh decisively in the political struggle.4 in addition, the first governor of greece based his expectations and plans on the long term. his long-term social programme needed a long time to mature. but time was not on his side. the premature and sudden violent death of the governor has halted the implementation of this long-term and grandiose social vision.” (pp. 86-87) ________ 1. christina koulouri and christos loukos, τα πρόσωπα του καποδίστρια: ο πρώτος κυβερνήτης της ελλάδας και η νεοελληνική ιδεολογία (1831-1996) [the faces of kapodistrias: the first governor of greece and modern greek ideology (1831-1996)], poreia, athens 1996, p. 28. [in greek] 2. αρχείον ιωάννου καποδίστρια [archive of ioannis kapodistrias], v. x, corfu 1986, pp. 97, 103 (notice sur la situation de la grece, november 6/18, 1830). [in greek] 3. in the same, v. viii, p. 37; dimitris loules, the financial and economic policies of president ioannis capodistrias 1828-1831, ioannina 1985, pp. 111-112. 4. c. m. woodhouse, capodistrias: the founder of greek independence, london 1973, p. 431; loukos, η αντιπολίτευση κατά του κυβερνήτη ιωάννη καποδίστρια 1828-1831 [the opposition against the governor ioannis kapodistrias 1828-1831], athens 1988, pp. 42, 398. [in greek] account b excerpts from the chapter “the governance of kapodistrias: society, policy, ideology” written by the historian v. kremmydas (2015), emeritus professor of the national and kapodistrian university of athens. “from the moment he was informed of his election, io. kapodistrias trembled at the idea that he would govern with a constitution he was determined not to do so, even if he needed to deny the position of governor. but what is the reason for both his specific behaviour and the whole policy? io. kapodistrias was an irreconcilable follower of the socio-political system of the enlightened despair and a fan of the policy of the holy alliance against the revolutionary movements. as a political system, the enlightened despair was a kind of resistance of totalitarianism to the enlightenment and implemented the following programme: favour and order; development of agriculture and trade, and basic education, that is, education for the people. but only the enlightened despot knew what would benefit agriculture, trade and the people. if we pay attention to the actions of the governor, we will find: first instance courts in each province, primary schools in each province as well, liberation of the seas from piracy to facilitate the trade, sale of cheap income, i.e., land, to farmers to develop agriculture. together, frequent tours and contact with the ‘people’, to show the interest of the ‘father’, but also for the history to record the return of his ‘love’ for him a boundless paternalism, if we want it, in today’s terms. this, however, the loving relationship of the bright father with the small people, as much as it can be characterised as pro-people policy, can equally be [characterised] as anti-social; ioannis kapodistrias was not interested in the society and social relations. let me now read you a short excerpt, foreign: ‘all the care of the governor was focused on how to exclude the phanariotes from things and he did not have to put in their position no other but only iakovakis rizos, who showed the greatest ingratitude as phanariot.’1 [...] the governor did not want to have any established power in front of him, even if it was with him; because then it would seem as if he was sharing his power with them.” (pp. 47-48) “the end of kapodistrias governance with his violent end was only the result of a tough conflict at the social, ideological and political level of almost four years; a conflict of power. the end with the koutsianou et emvalotis 71 | f l r assassination of the governor seems like the solution to a drama that was played almost every day.” (p. 52) “kapodistrias did not clash with local authorities; he despised the social relations, as they had been formed in the vortex of the revolution. he organised and exerted his policy in the absence of society.” (p. 53) ____________ 1. anagnostis kontakis, απομνημονεύματα [memoirs], g. tsoukala edition, athens, 1957, p. 71. [in greek] 2. science scenario (developed based on authentic scientific studies) the orientation of migratory birds based on earth’s magnetic field the following text developed on the occasion of the reference of the science textbook of the 6th grade to the migratory birds and how they are oriented themselves based on earth’s magnetic field (thematic unit: electromagnetism, pp. 96-97), based on the following published scientific studies: 1. holland, r. a., & helm, b. (2013). a strong magnetic pulse affects the precision of departure direction of naturally migrating adult but not juvenile birds. journal of the royal society interface, 10(81). doi: 10.1098/rsif.2012.1047 2. wiltschko, r., gehring, d., denzau, s., nießner, c., & wiltschko, w. (2014). magnetoreception in birds: ii. behavioural experiments concerning the cryptochrome cycle. the journal of experimental biology, 217(23), 4225–4228. doi: 10.1242/jeb.110981 introduction “every year, millions of migratory birds travel thousands of miles from one part of the planet to another and when they return, they usually find their old nest again, without anyone showing them the way. this feat of birds has not been fully explained by researchers. it is known, however, that some species of birds, in addition to the position of the sun, the direction of the wind and the sight, also perceive and use the magnetic field of the earth for their orientation. that way, they can continue their journey even at night, when they cannot orient themselves visually.” (excerpt from the greek textbook of science entitled “primary school science: investigate and discover”). an example of migratory birds that scientists often study is the european robin (erithacus rubecula). several studies have been conducted to identify possible mechanisms that allow robins to recognise and use the earth’s magnetic field in order to orient themselves and make their journey. a brief description of two scientific studies conducted with european robins follows, explaining the orientation of migratory birds based on earth’s magnetic field. account a brief description of the study conducted by the researchers r. a. holland and b. helm (2013) of the max planck institute for ornithology and university of konstanz in germany. researchers claim that migratory birds perceive the earth’s magnetic field with the help of a receptor that uses ferromagnetic material (e.g., ferro-magnetite particles found in birds’ beaks) to detect the inclination of the magnetic field. however, this mechanism seems to be used by adult birds that have made at least one migratory journey. experimentally, this mechanism can be disrupted by applying a strong magnetic pulse, which can magnetise the ferromagnetic material in the opposite direction. to test their hypotheses, the researchers experimented with european robins, which they caught and divided into two groups: the experimental group was placed in a special device and received a strong magnetic pulse, while the control group was placed in the same device without receiving the pulse. it is noted that both groups consisted of both adult and juvenile birds before their first migratory journey. after the intervention, robins were released into their natural environment and the researchers monitored koutsianou et emvalotis 72 | f l r them, through radio transmissions that researchers had applied on the robins, until the moment of their departure. the results of this study showed that the orientation of the adult birds of the experimental group was disturbed by the magnetic pulse and in particular, a significant deviation of their orientation was found in comparison with the respective control group. in contrast, the orientation of the juvenile birds on both groups (experimental and control) did not differ significantly after the researchers’ intervention. therefore, the researchers claim that robins have a receptor with magnetic properties, which plays an important role in their orientation during the migratory period; especially for the adult robins that seem to have developed a “magnetic map” after their first journey. account b brief description of the study conducted by the researchers r. wiltschko, d. gehring, s. denzau, c. nießner, and w. wiltschko (2014) of the goethe-universität frankfurt in germany. researchers claim that migratory birds perceive the earth’s magnetic field with the help of a photoreceptor (it is called cryptochrome and is a type of protein), which is found in birds’ eyes, is sensitive to blue light and allows them to receive information about the directions of the magnetic field, through appropriate biochemical reactions. experimentally, this mechanism can be disrupted if the birds are in a closed environment with green light and especially when they have been kept in a dark environment before (because biochemical reactions cannot be completed). to test their hypotheses, the researchers experimented with european robins that they caught and tested their orientation under different types of light (wavelengths) during the migratory period. the researchers applied two different experimental conditions. in the first condition, each bird was kept for an hour in a dark environment and then exposed to different wavelengths (blue, turquoise, and green light). in the second condition, each bird was exposed twice in succession to the same wavelength (blue, turquoise, and green light, respectively), without being kept in a dark environment. the same birds formed the control group, which were exposed only once to the green light, without being kept in a dark environment. in order to record their movements and orientation, the birds were placed in specially designed cages. in the first condition, the results of this study showed that the robins were oriented appropriately for the season when exposed to the blue and turquoise light, but a significant deviation in their orientation was found when exposed to the green light. similar results were found in the second condition, with the difference that only during the second exposure of the birds to the green light, their orientation was significantly affected. finally, in the control group, the robins were properly oriented after their exposure to the green light. therefore, the researchers claim that robins have a photoreceptor that acts as a “magnetic compass” and allows them to receive information about their orientation during the migratory period, through the completion of biochemical reactions. koutsianou et emvalotis 73 | f l r appendix a2: interview protocol(s) in the context of my doctoral dissertation, i study primary school teachers’ views on historical and scientific topics that have been derived from the school subjects of history and science of the 6th grade. i would like to let you know that there are no right and wrong answers; because i am interested in how you think about these topics. i would like to note that your participation in this study is anonymous and the collected information will be exclusively used as research data in the context of this study. could i have your permission to record our conversation? 1. questions about teachers’ demographic characteristics: how many years have you been working as a primary school teacher? how old are you? what studies have you done? have you done postgraduate studies? have you done any teaching (re)education/training? have you participated in any research project, as a researcher? can you determine (approximately) how many years you have taught in the 6th grade? when was the last school year? next, i would like you to read this text carefully (see appendix a1 for each scenario), which will be the context of our conversation today. i would like to remind you that there are no right and wrong answers, and that i am interested in your personal view on the topic that we will address. you have as much time as you want. 2. open-ended questions after reading the scenario in history and science, separately (in a counterbalanced order): (given that the questions for both topics and domains, history and science, were equivalent and in many cases exactly the same, the two interview protocols are presented at the same time, by noting the points where they differ with italics.) what do you think about the alternative views regarding the governance of ioannis kapodistrias/orientation of migratory birds based on earth’s magnetic field? (and/or) what is your opinion about the governance of i. kapodistrias/orientation of migratory birds based on earth’s magnetic field? do you think that these two accounts differ? if so, how? could both accounts be correct? if so, could one of them be more correct than the other? questions for the participants who endorsed a particular point of view on this topic: how did you get to this point of view? where do you base your point of view? something else? can you be sure that your view on this topic is correct? if so, how? if not, why not? do you think that we will ever know for sure? questions for the participants who did not endorse a particular point of view on this topic: could you (ever) say which was the better position? if so, how? if not, why not? what would you do/how would you think in order to decide on this topic? do you think we will ever know for sure which is the better position? if so, how? if not, why not? questions for all participants in your opinion, can there be certainty about this topic? koutsianou et emvalotis 74 | f l r (and/or further) can experts know with certainty the way of i. kapodistrias’ governing/how do migratory birds orient themselves based on earth’s magnetic field? when two people disagree on this topic, do you think that one is right and the other is wrong? if yes, what does the word “correct” mean to you? if not, can you say that one view is somehow better than the other? and why? what does the word “better” mean to you? how is it possible that people have such different views on this topic? in your opinion, is there any connection between the different views on such a topic? how is it possible that experts/researchers disagree on this topic and come to different conclusions? (and/or) the researchers who study this topic seem to come to different conclusions. how do you explain that? could the conclusions of a third historian/group of scientists be different? if so, how? if not, why not? (the wording about “the historian” and “the group of scientists” is random and was used according to the composition of the research groups of the authentic studies.) do you think that it is possible to have multiple different perspectives on this topic and all be correct? how could you personally judge the different perspectives on this topic? by what criteria? in your opinion, what does the knowledge about this topic consist of? (and/or) what does the knowledge about this topic include? what is the structure/organisation of knowledge with regard to this topic? (and/or) how do you imagine the structure/organisation of knowledge with regard to this topic? in your opinion, where does the knowledge come from for those who study this topic? (and/or) what is the source(s) of knowledge for those who study this topic? is there anything else that you would like to add to what we have already discussed? transition from the topic-specific to the domain-specific level: now, please imagine that you are asked to teach the governance of i. kapodistrias/orientation of migratory birds based on earth’s magnetic field in the 6th grade, in the context of the history/science lesson. please, take for granted that you have access to both accounts and everything else you need; and neither the school curriculum nor the context determines how you will approach it. how would you go to teach this topic in the classroom with 6th-grade students? 3. open-ended (indirect and direct) questions about ibl in history and science education: how does school history/science differ from historical/scientific research? are there similarities between school history/science and historical/scientific research? why (not)? do you think that teachers should explain to their students how the knowledge included in their textbooks is produced? (if yes) are you trying to apply this to your classroom? if so, how? do you think that school history/science should make students proficient in applying the reasoning skills that historians/scientists use to investigate the past/nature? why (not)? what should students know and be able to do in the context of school history/science? do you teach such skills in your classroom? if so, how? in your opinion, is inquiry a good approach to teach historical/scientific knowledge and skills? why (not)? (in case of asking the meaning of the word “inquiry”, the same clarification about both history and science was given: i.e., searching, finding, studying and synthesizing multiple sources. regarding school science, teachers were asked to direct answer this question twice, given that they defined inquiry through experiments, before being asked directly about the literature search.) do you use this approach in your own lessons? koutsianou et emvalotis 75 | f l r if yes, please describe how you apply inquiry in your classroom. how would you define ibl in the context of school history/science? where did you learn about ibl in history/science education? what do you consider to be the role of the students and what is your role in a lesson that focuses on ibl within school history/science? do you choose ibl as an approach in the context of school history/science? if applicable: what factors motivate you to foster ibl in your classroom (personal and external)? what difficulties do your students experience when engaging in ibl in the context of school history/science? what difficulties do you encounter when preparing, organising and facilitating ibl activities? microsoft word lamsa_ finalproofs.docx frontline learning research vol. 9 no. 3 (2021) 1-12 issn 2295-3159 info corresponding author: joni lämsä, p.o. box 35, fi-40014, university of jyväskylä, finland, joni.lamsa@jyu.fi doi: https://doi.org/10.14786/flr.v9i3.645 staying at the front line of literature: how can topic modelling help researchers follow recent studies? joni lämsä1, catalina espinoza2, ari tuhkala3, & raija hämäläinen1 1department of education, university of jyväskylä, finland 2center for advanced research in education, university of chile, chile 3finnish institute for educational research, university of jyväskylä, finland article received 23 june 2020 / article revised 20 december / accepted 26 march 2021 / available online 14 april abstract staying at the front line in learning research is challenging because many fields are rapidly developing. one such field is research on the temporal aspects of computer-supported collaborative learning (cscl). to obtain an overview of these fields, systematic literature reviews can capture patterns of existing research. however, conducting systematic literature reviews is time-consuming and do not reveal future developments in the field. this study proposes a machine learning method based on topic modelling that takes articles from a systematic literature review on the temporal aspects of cscl (49 original articles published before 2019) as a starting point to describe the most recent development in this field (52 new articles published between 2019 and 2020). we aimed to explore how to identify new relevant articles in this field and relate the original articles to the new articles. first, we trained the topic model with the results, discussion, and conclusion sections of the original articles, enabling us to correctly identify 74% (n = 17) of new and relevant articles. second, clusterisation of the original and new articles indicated that the field has advanced in its new and relevant articles because the topics concerning the regulation of learning and collaborative knowledge construction related 26 original articles to 10 new articles. new irrelevant studies typically emerged in clusters that did not include any specific topic with a high topic occurrence. our method may provide researchers with resources to follow the patterns in their fields instead of conducting repetitive systematic literature reviews. keywords: automatic content analysis; computer-supported collaborative learning; literature review; temporal analysis; topic model lämsä et al 2 | f l r 1. introduction research in learning sciences has become more interdisciplinary because increasingly complex datasets and methods may require the expertise of computer scientists and signal processors. this interdisciplinary collaboration opens up the possibility of new publication forums in the learning sciences. however, this could also make thorough systematic or thematic literature reviews (see gruber et al., 2020) even more arduous. thus, it would be useful if the vast amount of work done by scholars when conducting systematic literature reviews could be exploited when monitoring how a specific line of research would proceed. if relevant future studies can be automatically identified and related to previous research, this would decrease the need to perform recurring systematic literature reviews on similar topics, thus affording researchers more working hours to advance in their fields. to address these aspirations, we present a machine learning–based method that takes articles from a manual systematic literature review as a starting point to describe the recent developments in the field. we illustrate the potential of our innovative method in the context of research focusing on the temporal analysis of computer-supported collaborative learning (cscl). this field of research forms a particularly promising basis for studying its progress because the studies focusing on the temporal aspects of cscl are increasingly being published and involve interdisciplinary collaboration (e.g., hadwin, 2021; lämsä et al., 2021). in this study, we define the temporal analysis of cscl as analysing the characteristics of events or the interrelations between these events over time. the events may relate to learner interaction, thoughts and ideas developed during the interaction and the use of technological resources to mediate the interaction (see lämsä et al., 2021). a temporal analysis of cscl may benefit both practitioners and researchers by revealing how (not only what) learning occurs in cscl settings (lämsä, 2020), particularly now when covid-19 highlights the need for effective cscl more than ever (järvelä & rosé, 2020). when we manually reviewed the literature focusing on the temporal aspects of cscl (see section 2 and lämsä et al., 2021), we found that the interdisciplinary collaboration in this field has caused challenges regarding the commensurability and comparability of the studies. particularly, the studies seemed to be fragmented in terms of their theoretical frameworks (cf. hew et al., 2019), methodologies, and results and implications. this finding implies that both practitioners and researchers may struggle with staying at the front line concerning the big picture of cscl and its research because of this fragmentation. practitioners may benefit from our method if it can filter applicable research to support them in the design and implementation of research-based cscl innovations. similarly, our method can benefit researchers because it can illustrate whether and how the recent research has contributed to prior studies. we investigate the added value of our method for practitioners and researchers by addressing the following research questions: rq1: how and to what extent can a machine learning–based method be used to identify new relevant articles in the field of manual systematic literature review? rq2: how and to what extent can the machine learning–based method be used to relate new and original articles to each other? 2. methodology when manually reviewing the literature on the temporal aspects of cscl in february 2019 (see lämsä et al., 2021), we carefully selected the search terms concerning temporality, collaborative learning, and computer-supported learning. we used the education resources information center (eric), scopus, and web of science databases and identified 436 articles, of which we manually screened and assessed their eligibility. in this study, we included 49 peer-reviewed journal articles that focused on the temporal analysis of cscl for further analysis (original articles). to find new articles, lämsä et al 3 | f l r we repeated the literature searches with the same search terms and databases in february 2020 as for the original articles. the searches found 88 articles that had been published between february 2019 and 2020. from these 88 articles, we excluded 36 articles, of which 31 were duplicates, three had no full text available, one was a conference proceeding article, and one was already included in the set of the original articles. in the following analyses, we refer to these included 52 peer-reviewed journal articles as a set of new articles. the utilised machine learning–based method was grounded on a natural language processing technique known as topic modelling, which is based on statistical algorithms that find topics in a collection of documents (boyd-graber et al., 2017). these topics are ranked lists of words, where each word has a probability of belonging to a topic (see table 1), or more formally, topics are probability distributions over vocabularies. in the following sections, we describe how the original articles were exploited to build the topic models that, in turn, were used to identify the new relevant articles (rq1) and relate them to the original articles (rq2). figure 1 summarises our procedure. figure 1: procedure for describing original and new articles to address the research questions (rqs) 2.1 extracting and preprocessing text first, we extracted raw text from the original and new articles and removed tables, figures, formulas, bullet points, footnotes, and page numbers. second, we separated the different sections of the articles under the following headings: introduction, theoretical framework, methodology, results, discussion, and conclusion. however, because not all the articles had all of these sections (e.g., an article may have a combined results and discussion section), we decided to combine the sections into three wider sections: (1) introduction and theoretical framework, (2) methodology, and (3) results, discussion, and conclusion. then, we carried out text preprocessing, including common text cleaning, such as transforming text to lowercase and removing symbols and infrequent words. finally, we utilised lämsä et al 4 | f l r the natural language toolkit (bird et al., 2009) to perform word stemming (reducing words to their root form) and common english stop word removal (e.g., the, at, is). 2.2 training topic models we used latent dirichlet allocation (lda) (blei et al., 2003) and the gensim library (rehurek & sohjka, 2010) to train topic models for each section 1–3 of the original articles. the output of training topic models includes both a list of topics and the trained topic model itself. the trained topic model can process new text and measure the presence of the listed topics. we performed a sensitivity analysis based on topic coherence values (provided by the gensim library) to find an appropriate number of topics for each section. as an outcome, we had trained three topic models, one for each section, that all included 17 topics. 2.3 labelling topics from topic models the trained topic models contained a list of topics found in each section. we labelled the topics by analysing the most representative words and utilising expert knowledge from the manual systematic review of the literature. if possible, we labelled the topics based on the theoretical framework to which the most representative words refer. we demonstrate this idea in table 1 using topic models for section 3 as an example, presenting the labels and the 10 most representative words. for example, topic 1 (temporal aspects of cscl) is a generic topic that illustrates a stage in the temporal analysis procedure. namely, researchers code messages of groups of students, after which they analyse the typical sequences of messages. this kind of sequential analysis reveals what kind of messages follow each other in a short temporal context whose duration may be a few messages (the words with italics refer to the 10 most representative words from topic 1). 2.4 obtaining topic occurrence in original articles in lda, articles are represented as lists of topic probabilities; the goal is to find the topic probabilities of a document that are better suited to rebuild the document by randomly selecting words. for example, if an article has a higher topic probability for topic 16 compared with other topics (see table 1), most of the words in the article can be selected from the top of topic 16. we refer to topic probabilities in an article as a topic occurrence to distinguish them from words’ probabilities inside a topic. when we used topic models for sections 1–3, we obtained 51 topic occurrences for each original article (17 topic occurrences for each topic model). 2.5 obtaining topic occurrence in new articles the process used for the new articles was very similar to the one applied to the old articles (figure 1). the only difference was that we directly applied the trained topic models for sections 1–3 to obtain the topic occurrences of the new articles. we illustrate the topic occurrences of original, new relevant, and new irrelevant articles using the topic model for section 3 in figure 2. for each article, some topics have a higher probability than the rest (e.g., topic 16 is more relevant to an original article than to a new irrelevant article; see (a) and (c) in figure 2). therefore, we expect to find semantic similarity between topic occurrences that have shorter distances. lämsä et al 5 | f l r table 1 five topics and the assigned topic labels, including the 10 most representative words from the topic model for section 3 (results, discussion, and conclusion). topic 1: temporal aspects of computersupported collaborative learning topic 7: regulation of learning and learning performance topic 8: regulation of learning topic 11: socially shared metacognitive regulation (ssmr) topic 16: collaborative knowledge construction number model regul ssmr discuss student group learn process group signific perform collabor phase student code focus task thread knowledg group student social studi behaviour analysi ssrl1 share research construct show collabor student differ learn sequenc challeng group inquiri result knowledg differ result note process messag individu discuss data pattern 1socially shared regulation of learning (a) (b) (c) figure 2: the topic occurrence of (a) an original article, (b) a new relevant article, and (c) a new irrelevant article obtained using the topic model for section 3. the distance between (a) and (b) was 0.33, between (a) and (c) was 0.65, and between (b) and (c) was 0.49. to answer rq1, the first and second authors screened and labelled the new 52 articles manually as relevant or irrelevant regarding the analysis of the temporal aspects of cscl. in the first phase, we screened the journal and title of the articles and labelled the studies that did not have a learning or lämsä et al 6 | f l r instructional context as irrelevant (n = 27, e.g., studies from environmental sciences). in the second phase, we also screened the abstract of the articles and labelled the studies that did not focus on cscl and analyse its temporal aspects as irrelevant (n = 2, e.g., a study that focused merely on learning performance). we solved the disagreements between the first and second authors in the common meetings among all the authors. altogether, 23 new articles focused on the analysis of the temporal aspects of cscl (relevant), while 29 articles did not (irrelevant). next, for each topic model, we measured the distance between the corresponding topic occurrences of the new articles and original articles. the shorter the distance between two articles, the more similar the topic occurrences (figure 2). for each new article, we kept the distance to the closest original article (i.e., the most similar because a new relevant article might not be related to every article in the manual systematic literature review). finally, we compared the distances between the relevant and irrelevant articles. we selected the most suitable topic model so that the topic occurrences of the new relevant articles were similar to the ones from the original articles. to answer rq2, we used the articles’ topic occurrences from the previously selected topic model (rq1). we measured the similarity between topic occurrences using the euclidean distance, and we applied hierarchical clustering to find groups of similar articles. we performed the clustering in three levels: the root, two subgroups, and the leaves. the root of the clustering contains all the articles: 52 new articles and 49 original articles. the root was then divided into two subgroups, denoting the greatest distance between the articles belonging to different subgroups. the leaves are groups of articles of varying sizes. we interpreted the clusters by examining the topic occurrences (figure 2) and previously assigned topic labels (table 1). 3. results 3.1 the topic model trained with the results, discussion, and conclusion sections identified new relevant articles most accurately. we identified new relevant articles relating to the temporal aspects of cscl by measuring the distance between a new article and the closest original article. the results showed that for the three topic models, the relevant new articles were closer to the original articles than the irrelevant articles (figure 3). particularly, the topic model for section 3, which we trained with results, discussion, and conclusion sections, gave the best results because the distance between the new relevant articles and the closest original article overlapped the least with the distance between new irrelevant articles and the closest original article [figure 3 (c)]. when we used the topic model for section 3 and the distance of 0.27 as a threshold, 71% of the new articles, which were closer than the threshold, were relevant. those relevant articles represent 74% of the total relevant articles, which minimised the number of irrelevant articles. when we used topic models for sections 1 and 2, the distances between the original articles and new relevant articles overlapped more with new irrelevant articles [figure 3 (a) and (b)]. table 2 summarises our results if the distance of 0.27 is considered for the threshold. lämsä et al 7 | f l r (a) (b) (c) figure 3: boxplots of the distances between relevant and irrelevant new articles and the closest original article separately for (a) topic model for section 1, (b) topic model for section 2, and (c) topic model for section 3. lämsä et al 8 | f l r table 2 the numbers of relevant and irrelevant articles identified and missed when using three topic models topic model (section 1): introduction and theoretical framework topic model (section 2): methodology topic model (section 3): results, discussion, and conclusion relevant identified (true positives) 10 17 17 relevant missed (false negatives) 13 6 6 irrelevant identified as relevant (false positive) 3 14 7 irrelevant identified as irrelevant (true negatives) 26 15 22 precision (proportion of the true positives to the sum of true and false positives) 0.77 0.55 0.71 recall (proportion of the true positives to the sum of the true positives and false negatives) 0.43 0.74 0.74 3.2 a few topics with high topic occurrence relate new relevant to original articles figure 4 shows the outcome of the hierarchical clustering. when interpreting figure 4, based on the cscl theoretical frameworks, a few topics concerning collaborative knowledge construction and regulation of learning relate new relevant articles to original articles. topic 16 (see table 1) relates five new relevant articles to 17 original articles (the leaves with double borders), and these articles mostly belong to a smaller subgroup. topics 7, 8, and 11 (see table 1) relate five new relevant articles to nine original articles (the leaves with bold borders), and these articles belong to a larger subgroup. most of the new irrelevant articles (n = 28) were clustered into three different leaves (figure 4). from this set, 17 articles appeared in the leaves with different topics. moreover, 11 articles appeared in the leaf that included only one new relevant article and one original article. most of the original articles (n = 31) had a topic with a value higher than 0.45. in contrast, the clusters formed by various topics contained articles in which the topic occurrence of the most important topic was less than 0.2, meaning that there were no predominant topics. because topic occurrence is a probability distribution (must sum up to one), topic occurrence is more scattered if no particular topic is more significant [see figure 2 (c)]; this feature clusters together most of the irrelevant articles, but it also mixes irrelevant articles with relevant articles that have several important topics. in our case, 12 new relevant articles emerged in the leaves with different topics. lämsä et al 9 | f l r figure 4: new and original articles’ clustered and associated topics. each leaf has a grey box with the number of new and original articles. the text in the leaves corresponds to the number of the main topics and their labels. lämsä et al 10 | f l r 4. discussion and conclusion when considering some of the most-cited journals in the educational research field (review of educational research and educational research review), systematic literature reviews may ‘shape the future of research and practice’ (murphy et al., 2017, p. 2; alexander, 2020). we showed how a machine learning–based method can be used to identify new relevant articles in the field of the manual systematic literature review (rq1) and how it can relate new and original articles to each other (rq2). this novel method may help to follow the evolution of the ‘big picture’ of the research fields based on multidisciplinary collaboration, such as studies focusing on the temporal analysis of cscl. because these studies are published in various forums and involve different theoretical frameworks, methodological approaches, and results and implications, our methodological innovation may reveal how literature reviews can ‘shape the future of research and practice’. even though our method has potential, there are several limitations and critical issues to consider because the current study was an initial attempt to investigate the potential of topic modelling in the context of staying in the front line of literature. first, instead of an ‘automatic’ method, our method could be called ‘semiautomatic’ (cf. tuhkala et al., 2018). for instance, we extracted the texts of different article sections manually. even though this extraction process could be automatised, we decided to focus on automating more complex phases of our procedure (see figure 1). in the future, we aim to automatise a pipeline in which all the articles that arise from certain search terms can be processed, filtered according to their relevance, and related to original articles. second, because the number of original articles was relatively small, we could apply a heuristic to automatically identify the relevant and irrelevant new articles (rq1; see table 2). our heuristic was based on identifying new relevant articles without including too many irrelevant ones (high precision value) or filtering relevant ones (high recall value; see table 2). in the future, more complex methods can be tested, particularly if there are more articles. third, we trained the topic model only with original articles (figure 1), so all the topics can relate to the temporal aspects of cscl and the analysis of these aspects. thus, there were no topics that could have properly described new irrelevant articles. we will consider training the topic models by using both original and new articles and using the articles of related systematic literature reviews to capture a broader picture of the field. despite these limitations, our innovative method may open up new avenues to follow patterns in the different research fields based on the content of articles, instead of, for example, mere bibliographic information (chen et al., 2020). a recently published editorial of educational research review (gruber et al., 2020, p. 1) highlighted that systematic literature reviews should ‘extend beyond reporting or summarising what has been done in a particular field’. here, we see our method as more complementary than contradictory to researchers’ manual work when using the review approach to address their research questions. namely, topic modelling (figure 1) is an unsupervised method, so it may reveal patterns (or topics; see an example in table 1) from the existing literature to which researchers may not pay attention to. moreover, our method may assist researchers in some timeconsuming tasks when they conduct systematic literature reviews. if considering the preferred reporting items for systematic reviews and meta-analyses (prisma) statement (moher et al., 2009) as an example, machine learning–based methods may help researchers in identifying the relevant articles (rq1) and in screening and assessing the eligibility of the articles based on their relatedness to the research problems of interest (rq2). this kind of assistance would allow for investing more resources in a critical review of the included articles and, thus, scientifically valuable contributions. at the same time, it is as crucial in machine learning–based methods as it is in manual systematic literature reviews that researchers report their decisions transparently throughout the process (cf. our procedure in figure 1 and sections 2.1–2.5). in our research context, the topic model for section 3—which we trained with the results, discussion, and conclusion sections—had the best performance in identifying new relevant articles (rq1); this may be related to the theoretical fragmentation of studies in the educational technology field (hew et al., 2019) because original papers had been published in the journals of both computer sciences lämsä et al 11 | f l r and learning sciences. thus, the topic model for section 1 could not properly separate new relevant and irrelevant articles. moreover, the methods used to analyse the temporal aspects of cscl (e.g., sequential analysis) have been used in many disciplines, so the predictive power of the topic model for section 2 might be affected by this issue. in addition to identifying new relevant articles with moderate accuracy by using the topic model for section 3, we could relate new and old articles based on the few topics present in this topic model (rq2). for example, we found leaves of eight articles (four new relevant and four original articles) and seven articles (five new relevant and two original articles) that seemed to concern the temporal aspects of cscl in the context of the regulation of learning and collaborative knowledge construction, respectively (figure 4). these findings may inform both practitioners and researchers by showing widely used theoretical frameworks and providing a ‘state of the research’ (murphy et al., 2017, p. 5; see figure 4). in the future, when the number of new articles increases, clearer clusters and leaves of original and new articles may emerge. researchers can follow the fluctuation of the rising research topics by monitoring the size of the leaves (figure 4). the increasing number of articles would allow for more focused machine learning–based literature reviews so that the procedure for describing new articles (figure 1) would focus on, for example, a certain theoretical framework through which the temporal aspects of cscl can be analysed. our method could also be applied in completely different research fields if there is an existing systematic literature review from that field, and it is possible to train the topic models based on the included articles in the review (section 2.1). as research fields differ from each other and similar fields may have fundamentally different research traditions, further studies could, for example, investigate how to obtain a topic model (section 2.2) and its essential topics (section 2.3) whose topic occurrences (sections 2.4–2.5) could separate studies with fundamentally different epistemological stances. obtaining topic models and interpreting their essential topics, which researchers can do to address their research aims, require thorough expertise on the research field, in addition to the knowledge and skills to apply machine learning–based methods. key points many research fields on learning sciences are developing rapidly, which makes conducting systematic literature reviews a time-consuming task. we propose an innovative method that uses an existing systematic literature review to describe the recent developments in the field being reviewed. we illustrate the potential of our method using the literature on the temporal analysis of computer-supported collaborative learning. our machine learning–based method identified new relevant articles and related them to the previous literature with moderate accuracy. our method may decrease the need to do recurring systematic literature reviews, giving researchers more working hours to advance their fields. acknowledgements this research was funded by the academy of finland [grant numbers 292466 and 318095, the multidisciplinary research on learning and teaching profiles i and ii of university of jyväskylä]. lämsä et al 12 | f l r references alexander, p. a. (2020). methodological guidance paper: the art and science of quality systematic reviews. review of educational research, 90(1), 6–23. https://doi.org/10.3102/0034654319854352 bird, s., loper e., & klein, e. (2009). natural language processing with python. o’reilly media inc. blei, d. m., ng, a. y., & jordan, m. i. (2003). latent dirichlet allocation. journal of machine learning research, 3, 993–1022. boyd-graber, j. l., hu, y., & mimno, d. (2017). applications of topic models (vol. 11). now publishers incorporated. chen, x., zou, d., & xie, h. (2020). fifty years of british journal of educational technology: a topic modeling based bibliometric perspective. british journal of educational technology, 51(3), 692–708. https://doi.org/10.1111/bjet.12907 gruber, h., hämäläinen, r. h., hickey, d. t., pang, m. f., & pedaste, m. (2020). mission and scope of the journal educational research review. educational research review, 30, 100328. https://doi.org/10.1016/j.edurev.2020.100328 hadwin, a. f. (2021). commentary and future directions: what can multi-modal data reveal about temporal and adaptive processes in self-regulated learning? learning and instruction, 72, 101287. https://doi.org/10.1016/j.learninstruc.2019.101287 hew, k. f., lan, m., tang, y., jia, c., & lo, c. k. (2019), where is the ‘theory’ within the field of educational technology research? british journal of educational technology, 50(3), 956–971. https://doi.org/10.1111/bjet.12770 järvelä, s., & rosé, c. p. (2020). advocating for group interaction in the age of covid-19. international journal of computer-supported collaborative learning, 15(2), 143–147. https://doi.org/10.1007/s11412-02009324-4 lämsä, j. (2020). developing the temporal analysis for computer-supported collaborative learning in the context of scaffolded inquiry [doctoral dissertation, university of jyväskylä]. jyu dissertations, 245. http://urn.fi/urn:isbn:978-951-39-8248-5 lämsä, j., hämäläinen, r., koskinen, p., viiri, j., & lampi, e. (2021). what do we do when we analyse the temporal aspects of computer-supported collaborative learning? a systematic literature review. educational research review, 33, 100387. https://doi.org/10.1016/j.edurev.2021.100387 moher, d., liberati, a., tetzlaff, j., altman, d.g., & the prisma group. (2009). preferred reporting items for systematic reviews and meta-analyses: the prisma statement. plos med, 6(7), 1–6. https://doi.org/10.1136/bmj.b2535 murphy, p. k., knight, s. l., & dowd, a. c. (2017). familiar paths and new directions: inaugural call for manuscripts. review of educational research, 87(1), 3–6. https://doi.org/10.3102/0034654317691764 rehurek, r., & sojka, p. (2010). software framework for topic modelling with large corpora. proceedings of the lrec 2010 workshop on new challenges for nlp frameworks (pp. 45–50). elra. https://doi.org/10.13140/2.1.2393.1847 tuhkala, a., kärkkäinen, t., & nieminen, p. (2018). semi-automatic literature mapping of participatory design studies 2006–2016. in proceedings of the 15th participatory design conference (pp. 1–5). association for computing machinery. https://doi.org/10.1145/3210604.3210621 codepen zhao frontline learning research vol.8 no. 6 (2020) 59 76 issn 2295-3159 regulating distance to the screen while engaging in difficult tasks fang zhao1,robert gaschler 1,wolfgang schnotz2 & inga wagner2 1university of hagen, germany 2 university of koblenz-landau, germany article received 20 april 2020 / revised 31 july/ accepted 2 october / available online 4 november abstract regulation of distance to the screen (i.e., head-to-screen distance, fluctuation of head-to-screen distance) has been proved to reflect the cognitive engagement of the reader. however, it is still not clear (a) whether regulation of distance to the screen can be a potential parameter to infer high cognitive load and (b) whether it can predict the upcoming answer accuracy. configuring tablets or other learning devices in a way that distance to the screen can be analyzed by the learning software is in close reach. the software might use the measure as a person-specific indicator of need for extra scaffolding. in order to better gauge this potential, we analyzed eye-tracking data of children (n = 144, mage = 13 years, sd = 3.2 years) engaging in multimedia learning, as distance to the screen is estimated as a by-product of eye tracking. children were told to maintain a still seated posture while reading and answering questions at three difficulty levels (i.e., easy vs. medium vs. difficult). results yielded that task difficulty influences how well the distance to the screen can be regulated, supporting that regulation of distance to the screen is a promising measure. closer head-to-screen distance and larger fluctuation of head-to-screen distance can reflect that participants are engaging in a challenging task. only large fluctuation of head-to-screen distance can predict the future incorrect answers. the link between distance to the screen and processing of cognitive task can obtrusively embody reader’s cognitive states during system usage, which can support adaptive learning and testing. keywords: distance to the screen; multimedia reading; motor-cognition dual tasking; task difficulty; eye movement info corresponding author: email: fang.zhao@fernuni-hagen.de doi: https://doi.org/10.14786/flr.v8i6.663 1. introduction employing dual-task paradigms, previous research has provided evidence that posture control is associated with cognitive processing (chen et al., 2018; chong et al., 2010; kerr et al., 1985; lajoie et al., 1993; langhanns & müller, 2018; patel & bhatt, 2015; stelmach et al., 1990; woollacott & vandervelde, 2008). according to the multiple resource theory (wickens, 2002), when the cognitive task is demanding, fewer attentional resources can be allocated to postural control, leading to degraded performance in postural control. regulation of distance to the screen is one potential indicator of the quality of postural control (balaban et al., 2004; bonnet et al., 2017; kaakinen et al., 2018; qiu & helbig, 2012). it is thus likely that distance to the screen can indicate high cognitive demand. 1.1 seated posture the links between postural control and cognitive processing motivated our interest in testing the potential association of distance measures and state of cognitive processing. the current study used a still seated posture as the postural task. first, even highly practiced postures and seemingly effortless actions such as sitting require some cognitive processing. the classical setting of the studies on cognitive-postural interference was to ask participants to stand or walk on a balance platform and perform a cognitive task (e.g., donath et al., 2015; maylor & wing, 1996; pellecchia, 2003; schaefer et al., 2015). participants should keep the balance in a range to ensure not to fall. high cognitive demand was reflected by large variance of the set point of the position (e.g., center of pressure) and the actual value of the position. however, a recent study revealed that maintaining a seated posture while performing a cognitive task caused the largest decrements in the cognitive task – compared to relaxed lying and slight movement (langhanns & müller, 2018). in a study with children (igarashi et al., 2016), maintaining the seated posture was degraded when accompanied by a concurrent demanding cognitive task. hence, maintaining the seated posture was selected as the postural control task in the study, as it requires cognitive resources and seating is one of the most common body positions while learning. the research on postural stability (i.e., fluctuation of set point and the current position) observed contradictory findings on how people regulate the postures under demanding cognitive processing. some researchers claimed a significant increased distance deviation of sway from front to back with a concurrent difficult cognitive task (kerr et al., 1985; see a review, woollacott & shumway-cook, 2002). this statement has been supported by a recent study (gaschler et al., 2019), suggesting that increased variability of responses (i.e., where exactly the correct response panel is hit on a touch screen) can be influenced by the high cognitive load in a dual-task paradigm. yet, other researchers reported decreased fluctuation of head-to-screen distance when people conducted a visual search task demanding precision compared to a control visual task (balaban et al., 2004; bonnet et al., 2017; kaakinen et al., 2018). these studies demonstrated that there is a synergy between control of postural and oculomotor behaviors in precise visual search. the reduced postural sway facilitates the efficient control of eye movements (cf. balaban et al., 2004; kaakinen et al., 2018). it is thus worthwhile to investigate the change of the fluctuation of head-to-screen distance when people sit still and perform a cognitive task. in front of the screen, one has to keep the head in such a position that one can see well and one does not tilt back or forth. we hence define the regulation of distance to the screen as the changes or shifts of the set point, which is the distance of the middle of the eyes to the screen. previous studies have shown that the closer head-to-screen distance (i.e., changes of the set point of the posture) can indicate high cognitive engagement or highly focused attention (balaban et al., 2004; bonnet et al., 2017; kaakinen et al., 2018; qiu & helbig, 2012). the head-to-screen distance becomes smaller, when texts segments are task-relevant (kaakinen et al., 2018), and when tasks are challenging (e.g., identifying planes in the warship commander task in balaban et al., 2014; a difficult search task, “where is wally” in bonnet et al., 2017; two-digit addition task in qiu & helbig, 2012). however, the studies did not directly scrutinize whether regulation of distance to the screen can be influenced by task difficulty, and whether regulation of distance to the screen can predict the future answer accuracy. 1.2 multimedia learning the current study used multimedia learning as the cognitive task due to three motives. first, multimedia learning involves the visuospatial sketchpad (cf. baddeley, 1986), which has been shown to interfere with postural control (e.g., chen et al., 2018; vandervelde, woollacott & shumway-cook, 2005). multimedia reading refers to reading with texts and pictures. the integrative processing of texts and pictures is proposed to take place in a verbal channel and in a pictorial channel according to different theories: dual-coding theory (paivio, 1986), cognitive theory of multimedia learning (mayer, 2009), and integrative model of text-picture comprehension (schnotz & bannert, 2003). the theories share the common view that verbal processing in the verbal channel requires constructing the propositional representation of the text by making references of the objects and events being described in the text to the relevant world knowledge (i.e. situational representations, kintsch & van dijk, 1985; mental model, johnson-laird, 1983). the pictorial processing in the pictorial channel requires constructing the mental representations by structure mapping based on analogies between the external and internal depictive representations (gentner, 1989; knauff & johnson-laird, 2002; sims & hegarty, 1997). during multimedia processing, there are continuous interactions between verbal processing and pictorial processing in terms of mental model construction and mental model inspection (schnotz & wagner, 2018; zhao et al., 2014; zhao et al., 2020). dual-task studies have provided evidence that the comprehension of text-picture blended materials draws on more resources in the visuospatial sketchpad than does the comprehension of text-only materials (gyselinck et al., 2000; kruley et al., 1994). potentially, the visuospatial sketchpad is demanded in the perceptive analysis of pictures, and in the basic performance of mapping texts and pictures (cf. gyselinck et al., 2002). given that, pictures convey the meaning based on visual trace (e.g., lines, blob), the visuospatial sketchpad can be responsible for formatting and storing visual and spatial information (cf. gyselinck et al., 2002). a further study provides evidence that the spatial representation rather than the visual aspect plays an essential role in constructing and inspecting the mental models (knauff & johnson-laird, 2002). as mentioned earlier, spatial cognitive tasks interfere significantly with postural regulation (e.g., vandervelde, woollacott & shumway-cook, 2005). compared to the non-spatial cognitive tasks, spatial tasks caused significant decrease in postural control (chen et al., 2018). in order to investigate the cognitive-postural interaction, we employed a multimedia reading task, which relies on visuospatial processing. our usage of multimedia learning material contrasts with previous studies using cognitive tasks in cognitive-postural interference that come in small units, such as arithmetic calculation (igarashi et al., 2016) or backward counting (schaefer et al., 2015) or memorizing words (donath et al., 2015). we opted for a task demanding the extraction and integration of information units from different sources as it can be found in computerized tutorials in and outside the classroom. we aimed at testing whether difficulty of more complex learning tasks involving reading and picture processing has an impact on postural stability. furthermore, displaying multimedia materials in different reading processing can validate the generalization of the link between distance to the screen and cognitive processing. in accordance with mccrudden and schraw (2007), learners initially use a semantic coherence strategy as a default orientation for global understanding, and then switch to relevance-oriented processing after having constructed the initial mental model to meet the task demand. in multimedia learning, zhao, schnotz, wagner and gaschler (2020) have defined the phase of initial task-independent, general coherence-oriented processing as initial mental model construction. the later phase of selective processing (to adapt the mental model according to the question that should be answered) is defined as adaptive mental model specification. there are sequential constraints between both processing modes, as initial mental model construction is required before adaptive mental model specification can occur. in the current study, we analyzed distance to the screen when participants had the multimedia material and the question to be answered available. we differentiated two conditions that were identical with respect to the elements on the screen (question and multimedia material), but differed with respect to prior engagement with the material. in the question-after-material condition, participants had previously been granted the opportunity to process the material without yet knowing what the question would be. in this condition, participants should engage in initial mental model construction followed by adaptive mental model specification once the question is communicated. in the question-with-material condition, the question comes first and the material is displayed later. in this condition, the mental model should be constructed to specifically cover information required by the question. these different reading conditions are considered in the current study in order to generalize the findings on the association of distance measures and cognitive processing to different variants of engaging with multimedia materials. 1.3 adaptive learning the posture control indicator (i.e., distance to the screen) may be used to assess when learners are challenged by the material. in multimedia learning learners can engage with the material for many seconds to minutes. rather than waiting until finally the response might make explicit that scaffolding would have been needed, distance to the screen could be used as an online-measure of processing that might trigger support offers by the computerized system. educational systems have used user-modelling techniques to track user’s behavior and to analyze the learning progress to personalize the human-system interaction (scheiter et al., 2019). for instance, adele (garcia-barrios et al., 2004) used real-time eye tracking of fixations and saccades to establish a user-profiling database. the system can suggest adaptive information in proper media to match the user’s behavior. in the current study, head-to-screen distance is being tracked as a by-product of eye tracking. yet, such measures might be made available in interaction with everyday electronic media. current devices (e.g., eye trackers, nintendo wii, laptops, tablets and smartphones) have cameras to track the distance of users’ faces and the screen (kaltwasser et al., 20017; roig-maimó et al., 2018). as tracking the distance will be possible at low cost with standard hardware, we need to know whether this measure is useful to unobtrusively capture challenges encountered by the learner. if there is a link between distance to the screen and the processing of a cognitive task, distance measures can help to obtrusively track users’ cognitive states during system usage, enabling the system to adapt to users’ current cognitive load. this may be especially useful in children as they are still developing the postural control system (rival et al., 2005; schmid et al., 2005, 2007). increased head and body sway in children compared to adults were observed, when asked to stand on a platform with eyes open or closed (sakaguchi et al., 1994). reading particularly seemed to affect postural control. when children stood on a balance platform, reading aloud and standing still revealed more body sway than a counting backward task (blanchard et al., 2005). under cognitive-postural dual tasking, children compared to young adults had a larger decline in performance on postural control (wide stance vs. tandem romberg stance) while concurrently performing a difficult (visual working memory) task (reilly et al., 2008). the studies revealed that the attentional resources of the concurrent cognitive task could affect the postural control in children. as this work suggests that the detrimental effect of load on regulation of distance to the screen seems to be greater in children than in adults, we were confident to use a sample of children for this study. 1.4 research questions and predictions taken together, this study aimed to scrutinize (a) whether regulation of distance to the screen (i.e., head-to-screen distance and fluctuation of head-to-screen distance) can potentially help to unobtrusively infer whether a learner is currently engaging in a challenging task, (b) whether regulation of distance to the screen can predict the future answer accuracy. if this were the case, one would not need to await the classification of an answer as an error, but could receive extra support during the task. alternatively, adaptive testing (cf. wainer et al., 1990) can be employed by presenting learners tasks tailored from a range of difficulty corresponding to their level of proximal development. two reading conditions were included to validate the relationship of distance to the screen and cognitive processing. the question-after-material condition (material first) can possibly lead to lower head-to-screen distance, due to the selective processing in adaptive mental model specification (cf. kaakinen, 2018). the question-with-material condition (question first) can possibly have large variance of head-to-screen distance over the course of time, due to the involvement of initial mental model construction and adaptive mental model specification. in the current study, children were told to maintain the seated posture while reading multimedia materials at various difficulty levels. an eye tracker was used to detect the distance to the screen. their engagement in demanding tasks and the upcoming response accuracy are expected to be related to head-to-screen distance and fluctuations of head-to-screen distance. accordingly, we proposed the following research questions and predictions. research question 1: can head-to-screen distance and fluctuation of head-to-screen distance allow inferring whether people are engaging in a challenging task? prediction 1: with the increase of task difficulty, the head-to-screen distance decreases. prediction 2: with the increase of task difficulty, the fluctuation of head-to-screen distance increases. research question 2: can head-to-screen distance and fluctuation of head-to-screen distance allow predicting the upcoming answer accuracy? prediction 3: prior to incorrect responses to a question, the head-to-screen distance decreases. prediction 4: prior to incorrect responses to a question, the fluctuation of head-to-screen distance increases. 2. method the data were re-analyzed from a published article (zhao et al., 2020). more details concerning the method can be found in the abovementioned article. 2.1 participants one hundred forty-four secondary school students without any motoric impairments (mage = 13.0 years, sd = 3.2; 72 females, 72 males) participated in the experiment. all participants had normal or correct-to-normal vision. a post-hoc power analysis was conducted using g*power 3.1 (faul et al., 2009) for testing the difference between easy, medium and difficult questions in question-after material and question-with material reading conditions using a repeated-measures anova. results indicated that a sample size of 144 students would allow to detect an effect size of ηp2 = 0.25 at α = 0.05 with a statistical power (1− β) = 0.99. 2.2 experimental design participants were asked to perform a cognitive task, while maintaining a constant seated posture (60 65 cm from the screen). the cognitive task was to read six text-picture materials and answer one question per material. the text-picture blended materials were selected from authentic school textbooks in geography and biology from grades 5 to 8 in germany. we used a within-subject design, in which children were compared with themselves in three difficulty conditions: reading and solving (1) easy questions (2) medium questions (3) difficult questions combined with two reading conditions: question-after material and question-with material. we could thus counterbalance the individual differences, such as vision, interest or prior knowledge. 2.3 cognitive tasks an example of the text-picture materials is presented in figure 1. global understanding and question answering require students to integrate the textual and pictorial information. three questions at easy, medium, and difficult levels were produced based on the text-picture integration requirements (cf. wainer, 1992). an easy item requires only element mappings between text and picture. for example, “which of the following species is able to perceive tones/sounds at 120,000 hertz?” students have to read the text; find out that the hearing range is in blue; search for 120,000 hertz in the picture; find “dolphin and bat”; and find the correct answer, “bat”. a medium-complexity item requires mappings of simple relations. for instance, “which of the following species has the vocal range for producing the lowest tones (human being, 45-year-old/ cat/ dog/ cricket)?” students need to read the text and find the vocal range is in pink. then they should search for the vocal ranges for people (45-year-old), cat, dog, and cricket. then they should compare the ranges with the x-axis and find “dog” is the correct answer. the difficult item requires mappings of more complex relations between the text and picture. for instance: “which of the following four species is able to hear tones below 100 hertz as well as to produce tones below 1000 hertz and above 1500 hertz?” students need to integrate information from the text and information from the picture via color-coding. then they can answer that “cat” is the one who can hear tones below 100 hertz (numbers in blue) and can produce tones below 1000 hertz and above 1500 hertz (numbers in pink). each participant received one out of three questions per text-picture unit according to a latin square so that each worked on easy, medium and difficult items without re-encountering the same material. figure 1. example of a cognitive reading task on the topic auditory ranges (translated from german). for each animal, the numbers above are in pink, which are the vocal range, and the numbers below are in blue, which are the hearing range. to test the effect of cognitive load on the distance to the screen, we implemented two processing variants (see figure 2). in the question-after-material condition, participants first received a text-picture unit without any question. this aimed to stimulate initial mental model construction only, because no specific task was shown. then they were presented the same text-picture unit just seen before, but now together with a delayed question. this condition was meant to stimulate adaptive mental model specification, as participants were expected to engage in adapting their mental model to the requirements of the specific question. additionally, we implemented a question-with-material condition, in which a specific question was presented first. it was followed by the corresponding text-picture unit, while the question remained visible. this reading processing allows participants to combine initial mental model construction and adaptive mental model specification. we only report the comparisons between the last phase in question-after-material and question-with-material conditions that were identical to the elements on the screen but differed with prior engagement with the material. figure 2. reading conditions in the study. this study only focused on the distance to the screen in the last phase in question-after-material condition and question-with-material condition (marked in bolded frames), as item answering was only involved in these phases. 2.4 eye tracking after obtaining informed consent from the parents, each child was tested individually in a lab environment. first, participants’ verbal and spatial intelligence were tested (heller & perleth, 2000). then they were told to sit still in front of the tobii xl60 24-inch eye tracker operating at 60 hz temporal resolution. the text-picture units were self-paced reading tasks. students pressed the space key to turn pages, and pressed the arrow keys (up/down/left/right) to give answers. no feedback to the answers was provided. the accuracy of responses was reported to indicate the performance of the cognitive task. two distance parameters were used to indicate the performance of the postural task. the head-to-screen distance refers to the distance between the middle point of both eyes and the screen. the distances of both eyes to the screen were exported directly from the tobii studio, which is a software provided by the eye tracker company. as the cognitive demands in tasks can have momentary changes as readers proceed through the material (hyönä & niemi, 1990; kaakinen et al., 2018), we examined the time course changes of the head-to-screen distance during a reading task (i.e., from the appearance of the reading material until giving an answer to). specifically, we divided the total fixation duration of each participant for each reading material into five equal intervals. each quintile is 20% of the overall fixations (cf. lindner et al., 2017; zhao et al., 2020). the advantage of using this method is to rule out the effect of individual reading speed. in the current study, the distance to the screen was compared across five time segments of engaging with an item (i.e., the first 20% of recorded data to the last 20% of eye-tracking data). the fluctuation of head-to-screen distance is the standard deviation of the distance to the screen across the (five) quintiles. in equation (1) xi refers to the distance to the screen in one quintile, and μ refers to the mean of the distance to the screen in numbers of quintiles (n). the fluctuation of head-to-screen distance is the square root of the sum of the squared differences of each quintile from the mean. 3. results as a manipulation check, we first report the cognitive load of the difficulty level of items. then we test the influence of item difficulty and future response accuracy on the distance to the screen and on the fluctuation of the distance to the screen (i.e., deviation of distance to the screen across five quintiles). 3.1 validity of task difficulty the cognitive load of the item difficulty was verified by the accuracy rate and spatial distance between successive gaze-points. that is, we tested whether the difficult items demanded indeed more cognitive load compared to easy and medium items. a repeated-measures anova on the proportion of correct answers with 3 (item difficulty: easy vs. medium vs. difficult) × 2 (reading condition: question after material vs. question with material) within-subject factors showed only a significant main effect of item difficulty, f(1.88, 269.11) = 27.20, p < .001, ηp² = .16 (here and elsewhere we applied greenhouse geisser-correction when appropriate). easy items (m = 72.6%, sd = 34.9%) had a higher accuracy rate than medium items (m = 51.4%, sd = 39.2%), t(143) = 5.52, p < .001, d = .46, and difficult items (m = 42.4%, sd = 36.6%), t(143) = 7.48, p < .001, d = .62. there was no significant difference between medium and difficult items, t(143) = 1.92, p = .056, d = .16. no other effect was shown: reading condition, and item difficulty × reading condition, fs < 1, suggesting that presenting questions before or after the materials does not affect response accuracy. shorter saccade length is suggested to indicate higher cognitive load (cf. debue & van de leemput, 2014). we computed the euclidean distance of successive gaze-points per 60 ms, as we used a 60 hz eye tracker. the 3 (item difficulty: easy vs. medium vs. difficult) × 2 (reading condition: question after material vs. question with material) anova on the saccade length revealed a main effect of item difficulty, f(1.89, 270.65) = 4.21, p = .02, ηp² = .03. the paired-samples t-test in two-tails showed that easy items (m = 12.25 pixel, sd = 3.49 pixel, t(143) = 2.40, p = .02, d = .20) had longer saccade lengths than medium items (m = 11.86 pixel, sd = 3.22 pixel) and difficult items (m = 11.83 pixel, sd = 3.25 pixel, t(143) = 2.37, p = .02, d = .20). no difference was found between medium items and difficult items, p = .85. this suggests that the easy items required less attentional resources than medium and difficult items. the main effect of reading condition, f(1, 143) = 9.95, p = .002, ηp² = .07, suggested a longer successive length in the question-with-material condition (m = 12.29 pixel, sd = 3.40 pixel) than in the question-after-material condition (m = 11.67 pixel, sd = 3.57 pixel). this was not surprising, as participants have read the item first in the question-with-material condition. they hence searched and selected task-related information in a top-down manner, varying the attended location quickly over a wide range (cf. theeuwes & belopolsky, 2010). the accuracy rate and saccade length therefore confirmed that the cognitive load of answering easy items was lower than of difficult items. 3.2 task difficulty panel a panel b figure 3. the head-to-screen distance in millimeter in five quintiles in panel a, and the fluctuation of head-to-screen distance in millimeter (averaged within-subjects standard deviation distance of five quintiles) in panel b during question-after-material and question-with-material conditions, when easy, medium and difficult items were performed. 3.2.1. head-to-screen distance the average head-to-screen distance for all participants had a range from 528.5 mm to 692.3 mm (m = 622.7 mm, sd = 24.2 mm). the 3 (item difficulty: easy vs. medium vs. difficult) × 2 (reading condition: question after material vs. question with material) × 5 (quintile) within-subjects repeated-measures anova on head-to-screen distance revealed a significant main effect of item difficulty, f(2, 286) = 7.54, p = .001, ηp² = .05 (see figure 3a). it confirmed prediction 1 that the concurrent requirement of the difficult task led to a decrement in head-to-screen distance. participants tended to be closer to the screen with difficult items (m = 621.1 mm, sd = 24.9 mm) compared to medium items (m = 622.9 mm, sd = 24.9 mm), t(143) = -2.24, p = .03, d = .19, and easy items (m = 624.1 mm, sd = 24.1 mm), t(143) = -3.68, p < .001, d = .31. there was no significant difference between easy and medium items, p = .10. the interaction effect of reading condition × quintile, f(3.17, 453.03) = 3.06, p = .02, ηp² = .02, indicated a larger distance discrepancy across five quintiles when participants read the material only once (question with material). no other effect was revealed: reading condition, f(1, 143) = 3.73, p = .06, ηp² = .03; quintile, f(3.26, 466.71) = 1.48, p = .22, ηp² = .01; item difficulty × reading condition, item difficulty × quintile, reading condition × item difficulty × quintile, fs < 1. 3.2.2. fluctuation of head-to-screen distance the 3 (item difficulty: easy vs. medium vs. difficult) × 2 (reading condition: question after material vs. question with material) anova was performed on fluctuation of head-to-screen, which was the within-subject standard deviation across the five quintiles for each person under each reading condition. a main effect of item difficulty was revealed, f(1.84, 262.53) = 4.67, p = .01, ηp² = .03. participants had larger body sway in the anteroposterior direction for difficult items (m = 6.8 mm, sd = 8.0 mm) compared to easy items (m = 5.2 mm, sd = 4.9 mm), t(143) = -2.81, p = .006, d = .23, which confirmed prediction 2 (see figure 3b). there was no significant difference between easy items and medium items (m = 5.8 mm, sd = 5.4 mm), p = .20, and between medium items and difficult items, p = .07. no other effect was revealed: reading condition, item difficulty × reading condition, fs < 1. 3.3 future response accuracy panel a panel b figure 4. the head-to-screen distance (panel a) and the fluctuation of head-to-screen distance (panel b) during question-after-material and question-with-material conditions, when items were answered correctly and incorrectly. the error bars in panel b are the standard errors of the mean. while in the above analysis of task difficulty, all participants were included as each participant read materials at easy vs. medium vs. difficult levels, this approach was not feasible when predicting future response accuracy. here we used the data from 78 participants out of 144, because they had both correct and incorrect responses in two reading conditions allowing for a within-subjects analysis. we could thus compare how the same participant regulated the head-to-screen distance when s/he answered the questions correctly vs. incorrectly. 3.3.1. head-to-screen distance as shown in figure 4a, 2 (future answer: correct vs. incorrect) × 2 (reading condition: question after material vs. question with material) × 5 (quintile) within-subjects repeated-measures anova on head-to-screen distance showed no significant effect of future answer, f(1, 77) = 1.93, p = .17, ηp² = .02. no other effect was found, quintile, f(2.88, 221.65) = 1.32, p = .27, ηp² = .02; reading condition × quintile, f(3.13, 241.05) = 2.58, p = .052, ηp² = .03; future answer × reading condition × quintile, f(3.06, 235.87) = 1.57, p = .20, ηp² = .02; reading condition, future answer × reading condition, future answer × quintile, fs < 1. the results disconfirmed prediction 3 and indicated that head-to-screen distance cannot differentiate the future correct or incorrect answers. 3.3.2. fluctuation of head-to-screen distance we conducted the 2 (future answer: correct vs. incorrect) × 2 (reading condition: question after material vs. question with material) anova on fluctuation of head-to-screen distance. it revealed a main effect of future answer, f(1, 77) = 9.17, p = .003, ηp² = .11, suggesting larger deviations of distance to the screen when the item will be answered incorrectly (m = 7.3 mm, sd = 8.6 mm) than when the question will be answered correctly (m = 4.7 mm, sd = 3.9 mm). it confirmed prediction 4. no main effect or interaction effect involving reading condition was revealed, fs < 1, indicating that the fluctuation of head-to-screen distance could predict the upcoming response accuracy regardless of different reading conditions. 4. discussion our main goal in this study was to examine whether regulation of distance to the screen (i.e., head-to-screen distance and fluctuation of head-to-screen distance) can be influenced by task difficulty and whether it can predict the upcoming incorrect responses. using a within-subject design, secondary school children were told to solve multimedia reading tasks at easy vs. medium vs. difficult levels while maintaining a seated posture. we expected a deteriorated regulation of distance to the screen, while children performing a difficult task concurrently. 4.1 association of task difficulty and distance to the screen the major finding is that the concurrent demand of difficult tasks led to a deterioration in regulating the distance to the screen, which is in agreement with studies on cognitive-postural control (chong et al., 2010; igarashi et al., 2016; kerr et al., 1985; lajoie et al., 1993; patel & bhatt, 2015; stelmach et al., 1990). this suggests that distance to the screen can be employed to monitor scaffolding demands of learners interacting with electronic devices. moreover, we observed effects of task difficulty even though learners were instructed not to move. this suggests that when learners do not face such a demand, effects of task difficulty on distance regulation might be larger than observed in the current study. future work will have to detail whether effects will be large enough to engage in individual online diagnostic of task engagement or are limited to scenarios where online analysis of group-averaged data can be applied (i.e. whole classroom engaging in one task with each student working on an electronic device). in our setup, children tended to move slightly closer to the screen, while seated still and performing a challenging task (difficult vs. medium vs. easy). the findings suggest that reduced head-to-screen distance can reflect the high cognitive load, replicating results in previous studies (balaban et al., 2004; bonnet et al., 2017; kaakinen et al., 2018; qiu & helbig, 2012). according to kaakinen et al. (2018), the underlying mechanism responsible for closer distance to the screen is the high demand of visual precision. on the one hand, the task is difficult in terms of content. on the other hand, one needs visual precision to better recognize or process the information. previous studies have shown that hard-to-read stimuli (stimuli-background contrast) can cause lack of conflict adaptation (fritz et al., 2015) and stimulus size is closely linked to the properties of early information processing (bayer et al., 2012). it is therefore necessary to approach to the screen with difficult tasks. however, larger body sway (i.e., fluctuation of head-to-screen distance) under high cognitive load contradicts with previous findings (balaban et al., 2004; bonnet et al., 2017; kaakinen et al., 2018; qiu & helbig, 2012). likely, the task employed in the current study distinguishes from these studies: text reading without item-answering (kaakinen et al., 2018), identifying planes in the warship commander task (balaban et al., 2014), visual search task, “where is wally” (bonnet et al., 2017), two-digit addition task (qiu & helbig, 2012). in these tasks, body stability can be beneficial, because the eyes can efficiently locate relevant information. however, the current study used a cognitive task, in which participants should read multimedia materials and answer questions. larger body sway under high cognitive load can be interpreted as the dynamic process of searching and locating the task-relevant sentences during the course of reading. in figure 1, participants only need to map elements via color-coding and compare relatively simple relations (e.g., “which of the following species has the vocal range for producing the lowest tones?”). in contrast, they should map more complex relations between text and picture by extracting task-relevant information (e.g., “which of the following four species is able to hear tones below 100 hertz as well as to produce tones below 1000 hertz and above 1500 hertz?”). accordingly, readers “zoom in” their attention with task-relevant segments and “zoom out” their attention with task-irrelevant segments (cf. kaakinen & hyönä, 2014). the adjustment of their attention was more diverse when the cognitive task was difficult compared to medium and easy. only the fluctuation of head-to-screen distance (rather than mean head-to-screen distance) can predict the upcoming response accuracy. incorrect responses can signal cognitive overload, resulting in increased demands in attentional resources (cf. sweller, 1988). when children are unable to solve a task, their cognitive capacity can be overloaded and they may struggle with the task and feel stressful. according to doumas et al. (2018), increased emotional stress caused by difficult tasks can lead to larger body sway. we hence assume that the detected larger fluctuation of head-to-screen distance for the upcoming incorrect responses was due to the increased cognitive load and emotional stress. furthermore, it is possible that the fluctuation of distance to the screen is more sensitive to individual’s cognitive capacity. specifically, it concerns the difference between each quintile and the mean of the distance to the screen across five quintiles (see equation 1). accordingly, the fluctuation of head-to-screen distance is influenced by task difficulty and can predict the upcoming response accuracy. according to the multiple resource theory (wickens, 1989), two tasks interfere with each other when they share similar attentional resources, as the processing capacity is limited. under dual-task paradigm, many studies provide evidence of cognitive-postural interference (e.g., chen et al., 2018; kerr et al., 1985; patel & bhatt, 2015; stelmach et al., 1990; woollacott & vandervelde, 2008). the studies give clear hints that the mechanisms responsible for regulation of postural control interact with higher-level cognitive systems. in the current study, we found a deterioration in performance on regulation of distance to the screen, when accompanied by a difficult cognitive task. there is possibly a trade-off of attentional resources between the regulation of distance to the screen and the cognitive task. when the cognitive task is attentionally demanding, less attentional resources can be allocated to the regulation of distance to the screen. consequently, a decline in performance on regulation of distance to the screen takes place. our findings support the hypothesis that regulation of distance to the screen can be influenced by task difficulty. furthermore, several studies have demonstrated that the control and regulation of simple postural tasks such as still sitting require attention. langhanns and müller (2018) reported the largest decrements in the cognitive task, while concurrently performing still sitting compared to relaxed lying and slight movement. igarashi et al. (2016) observed the detrimental effect of the regulation of postural control while children concurrently performing difficult cognitive tasks. the present experiment yielded a declined regulation of distance to the screen, when performing a challenging cognitive task (reading and answering difficult vs. medium vs. easy items). these observations suggest that even seemingly effortless actions such as sitting require some cognitive processing. while sitting, ongoing contraction of muscles is necessary to avoid bending the trunk forward (cf., o’sullivan et al., 2002, 2006). other than that, while maintaining the seated posture, unwanted movements should be inhibited, which may trigger higher resource demands than execution of movement (cf. huestegge & koch, 2014). with a concurrent cognitive task, the attention to maintain the seated posture can be reduced. therefore, participants failed to control their seating position. their body moved closer to the screen and the amplitude of the movement in anteroposterior direction became large. nevertheless, the effects of distance measures are rather small, as sitting is a relatively stable posture (igarashi et al., 2016). yet it does not imply that they will be useless for practical purposes. in this eye-tracking experiment, children were instructed to sit still and not to move. thus, we tested this measure under conditions that should have made it difficult to find an effect at all. if students do not receive the instruction “not to move”, mean shift of distance to the screen and fluctuation of distance to the screen might show stronger effects. 4.2 multimedia learning as cognitive task the cognitive task in this study was to read multimedia materials, which is one of the ubiquitous learning scenarios in the classroom. according to the previous studies (gyselinck et al., 2000, 2002; kruley et al., 1994), the visuospatial sketchpad is involved in the integrative processing: perceptive analysis and storage of pictorial information and matching information in texts and pictures. previous research has demonstrated that spatial cognitive tasks interfere more with the regulation of postural control than non-spatial cognitive tasks (chen et al., 2018; fuhrman et al., 2015; kerr et al., 1985; vandervelde, woollacott & shumway-cook, 2005). the comprehension task in the current study is much more complicated, as it involves a multimodal process: verbal processing and pictorial processing (paivio, 1986; mayer, 2009; schnotz & bannert, 2003). nevertheless, we still observed the effect of cognitive task difficulty on the distance parameters. distance to the screen showed a significant interaction of reading condition and quintiles. readers moved closer to the screen over the course of processing the multimedia material to find the answer to the question, which had been presented upfront. whereas distance to the screen varied less over the time quintiles when participants processed question and material after having had a chance to process the multimedia material without a question being posed at first. in the question-after-material condition, readers should already have engaged in initial mental model construction before quintile 1 and now engage in adaptive mental model specification in light of the question now made available (cf. zhao et al., 2020). apparently, this is not accompanied by moving closer to the screen. yet, such a dynamic movement is observed in the question-with-material condition. as the question is presented first, from quintile 1 onwards readers could deal with the multimedia material from the perspective offered by the question. yet, although reading can be selective, some initial mental model construction is still required before the adaptive mental model specification can take place. given the hints that encountering task-relevant segments of the material can lead to a shorter head-to-screen distance (kaakinen & hyönä, 2014), the reduction of distance over time might reflect that (after initial mental model construction) readers in later quintiles become more and more likely to find and identify task-relevant segments of the multimedia material. 4.3 limitations one can argue that the head-to-screen distance can be moderated by other factors, such as prior knowledge, vision (wearing no glasses vs. near-sighted glasses vs. far-sighted glasses), state of puberty or interest. these factors differ always individually as learners differ from each other. hence, these factors cannot be easily controlled by the instructor, when s/he designs the lessons or assignments. however, the task difficulty is relatively easy to control. in the current study, we produced the questions at different levels based on the text-picture requirements (cf. wainer, 1992). by using a within-subjects design with a large sample size, in which each child was working on (1) easy questions (2) medium questions (3) difficult questions, it was possible for us to counterbalance the inter-individual differences. the factors regarding the individual differences could be considered as constituting individually different base lines for head-to-screen distance, from which an individual deviates according to his or her experienced task difficulty. whether the results of the present study can be generalized to other tasks or to other samples remains to be investigated. as children’s cognitive systems are still developing, young adults may have less detrimental dual-task effects while performing postural task concurrently with a cognitive task (e.g., less body sway in rankin et al., 2000). due to the decline in execution and cognitive functions, old adults (66 84 years) tended to have larger body sway than young adults (19 30 years) under demanding cognitive load (stelzel et al., 2017). further research should examine whether the regulation of distance to the screen may be a relevant measure in other age groups as well. as the medium difficulty questions in this study in some comparisons did not differ from easy questions or difficult questions, it might be of special interest to study whether and to what extent the task difficulty with clear differences affects the regulation of distance to the screen. furthermore, while there are high performance instruments to track posture parameters (e.g., nintendo wii), it is relevant to explore the power of simpler indicators which might soon be easily available in interfaces to interact with electronic media. this study employed the eye tracker to track the distance to the screen. yet, smartphones or tablets can also track the distance to the screen via front camera (cf. kaltwasser et al., 2017; roig-maimó et al., 2018). future research should test the sensitivity of distance to the screen with mobile phones or other devices. 4.4 conclusions tentative evidence gave the first glimpse that distance measures might become useful indicators of those learners that are encountering difficulties in self-regulated multimedia learning. with the increase of task difficulty, children cannot manage to maintain the constant distance to the screen well. the fluctuation of head-to-screen distance seems to be the most promising measure, as it not only can be significantly influenced by the task difficulty but also can predict the future response accuracy. this study sheds light also on educational implications. it may be better to allow children to sit freely when they perform difficult tasks on a computer screen (see also, igarashi et al., 2016; langhanns & müller, 2018; reilly et al., 2008). tracking distance to the screen thus might be useful to scaffold self-regulated learning in the long run. with an increasing fluctuation of head-to-screen distance, the computerized learning systems can assess the level of user’s knowledge, detect the current high cognitive load and can give adaptive assistance to facilitate deeper learning. an alternative method is to employ adaptive testing (cf. wainer et al., 1990) with tailored tasks from a range of difficulty corresponding to their level of proximal development. nevertheless, it is important to note that the distance to the screen can be individually different, in terms of prior knowledge or vision, etc. here we only considered variability of distance to the screen within persons. further research should focus more on the effect of individual differences. keypoints keypoints distance to the screen measures are promising to infer reader’s current cognitive state. closer head-to-screen distance can indicate that readers are engaging in a challenging task. larger fluctuation of head-to-screen distance can indicate high cognitive load and can predict upcoming response accuracy. acknowledgments this study is part of the bite project on text-picture integration funded by the german research foundation (grant no: schn 665/3-1, schn 665/6-1, schn 665/6-2) within the special research program ‘competence models for assessing individual learning results and for balancing of educational processes’. we thank robert szwarc for making the graphs. we thank university of hagen for support in publishing this work. references atkinson, r. c., & shiffrin, r. m. (1968). human memory: a proposed system and its control processes. in r. j. sternberg, k. w. spence, & j. t. spence (eds.), the psychology of learning and motivation: advances in research and theory (vol. 2, pp. 89–195). academic press. baddeley, a. (1986). working memory. oxford university press. balaban, c. d., cohn, j., redfern, m. s., prinkey, j., stripling, r., & hoffer, m. (2004). postural control as a probe for cognitive state: exploiting human information processing to enhance performance. international journal of human-computer interaction, 17 (2), 275–286. https://doi.org/10.1207/s15327590ijhc1702_9 bayer, m., sommer, w., & schacht, a. (2012). font size matters—emotion and attention in cortical responses to written words. plos one, 7(5), e36042. https://doi.org/10.1371/journal.pone.0036042 blanchard, y., carey, s., coffey, j., cohen, a., harris, t., michlik, s., & pellecchia, g. l. (2005). the influence of concurrent cognitive tasks on postural sway in children. pediatric physical therapy, 17(3), 189–193. https://doi.org/10.1097/01.pep.0000176578.57147.5d bonnet, c. t., szaffarczyk, s., & baudry, s. (2017). functional synergy between postural and visual behaviors when performing a difficult precise visual task in upright stance. cognitive science, 41(6), 1675–1693. https://doi.org/10.1111/cogs.12420 chen, y., yu, y., niu, r., & liu, y. (2018). selective effects of postural control on spatial vs. nonspatial working memory: a functional near-infrared spectral imaging study. frontiers in human neuroscience, 12, 243. https://doi.org/10.3389/fnhum.2018.00243 chong, r. k., mills, b., dailey, l.-, lane, e., smith, s., & lee, h. (2010). specific interference between a cognitive task and sensory organization for stance balance control in healthy young adults: visuospatial effects. neuropsychologia, 48(9), 2709–2718. https://doi.org/10.1016/j.neuropsychologia.2010.05.018 debue, n., & van de leemput, c. (2014). what does germane load mean? an empirical contribution to the cognitive load theory. frontiers in psychology, 5, 1099. https://doi.org/10.3389/fpsyg.2014.01099 donath, l., roth, r., lichtenstein, e., elliot, c., zahner, l., & faude, o. (2015). jeopardizing christmas: why spoiled kids and a tight schedule could make santa claus fall? gait & posture, 41(3), 745–749. https://doi.org/10.1016/j.gaitpost.2014.12.010 doumas, m., morsanyi, k., & young, w. r. (2018). cognitively and socially induced stress affects postural control. experimental brain research, 236(1), 305–314. https://doi.org/10.1007/s00221-017-5128-8 faul, f., erdfelder, e., buchner, a., & lang, a. g. (2009). statistical power analyses using g*power 3.1: tests for correlation and regression analyses. behaviour research methods, 41(4), 1149–1160. https://doi.org/10.3758/brm.41.4.1149 fritz, j., fischer, r., & dreisbach, g. (2015). the influence of negative stimulus features on conflict adaption: evidence from fluency of processing. frontiers in psychology, 6, 185. https://doi.org/10.3389/fpsyg.2015.00185 fuhrman, s. i., redfern, m. s., jennings, j. r., & furman, j. m. (2015). interference between postural control and spatial vs non-spatial auditory reaction time tasks in older adults.journal of vestibular research: equilibrium & orientation, 25(2), 47–55. https://doi.org/10.3233/ves-150546 garcia-barrios, v., gütl, c., preis, a. m., andrews, k., pivec, m., mödritscher, f., & trummer, c. (2004). adele: a framework for adaptive e-learning through eye tracking. i-know’04, 609–616. gaschler, r., zhao, f., röttger, e., panzer, s., & haider, h. (2019). more than hitting the correct key quickly: spatial variability in touch screen response location under multitasking in the serial reaction time task. experimental psychology, 66(3), 207–220. https://doi.org/10.1027/1618-3169/a000446 gentner, d. (1989). the mechanisms of analogical learning. in s. vosniadou & a. ortony (eds.), similarity and analogical reasoning (pp. 199–241). cambridge university press. gyselinck, v., cornoldi, c., dubois, v., de beni, r., & ehrlich, m.-f. (2002). visuospatial memory and phonological loop in learning from multimedia. applied cognitive psychology, 16(6), 665–685. https://doi.org/10.1002/acp.823 gyselinck, v., ehrlich, m.-f., cornoldi, c., de beni, r., & dubois, v. (2000). visuospatial working memory in learning from multimedia systems. journal of computer assisted learning, 16(2), 166–176. https://doi.org/10.1046/j.1365-2729.2000.00128.x huestegge, l., & koch, i. (2014). when two actions are easier than one: how inhibitory control demands affect response processing. acta psychologica, 151, 230–236. https://doi.org/10.1016/j.actpsy.2014.07.001 hyönä, j., & niemi, p. (1990). eye movements during repeated reading of a text. acta psychologica, 73(3), 259–280. https://doi.org/10.1016/0001-6918(90)90026-c igarashi, g., karashima, c., & hoshiyama, m. (2016). effect of cognitive load on seating posture in children. occupational therapy international, 23(1), 48–56. https://doi.org/10.1002/oti.1405 johnson-laird, p. n. (1983). mental models: towards a cognitive science of language, inference, and consciousness. oxford university press. kaakinen, j. k., ballenghein, u., tissier, g., & baccino, t. (2018). fluctuation in cognitive engagement during reading: evidence from concurrent recordings of postural and eye movements. journal of experimental psychology: learning, memory, and cognition , 44(10), 1671–1677. https://doi.org/10.1037/xlm0000539 kaakinen, j. k., & hyönä, j. (2014). task relevance induces momentary changes in the functional visual field during reading. psychological science, 25(2), 626–632. https://doi.org/10.1177/0956797613512332 kaltwasser, l., moore, k., weinreich, a., & sommer, w. (2017). the influence of emotion type, social value orientation and processing focus on approach-avoidance tendencies to negative dynamic facial expressions. motivation and emotion, 41(4), 532–544. https://doi.org/10.1007/s11031-017-9624-8 kerr, b., condon, s. m., & mcdonald, l. a. (1985). cognitive spatial processing and the regulation of posture. journal of experimental psychology: human perception and performance , 11(5), 617–622. https://doi.org/10.1037/0096-1523.11.5.617 kintsch, w., & van dijk, t. a. (1978). toward a model of text comprehension and production. psychological review, 85(5), 363–394. https://doi.org/10.1037/0033-295x.85.5.363 knauff, m., & johnson-laird, p. n. (2002). visual imagery can impede reasoning. memory & cognition, 30(3), 363–371. https://doi.org/10.3758/bf03194937 kruley, p., sciama, s. c., & glenberg, a. m. (1994). on-line processing of textual illustrations in the visuospatial sketchpad: evidence from dual-task studies. memory & cognition, 22(3), 261–272. https://doi.org/10.3758/bf03200853 lajoie, y., teasdale, n., bard, c., & fleury, m. (1993). attentional demands for static and dynamic equilibrium. experimental brain research, 97(1), 139–144. https://doi.org/10.1007/bf00228824 langhanns, c., & müller, h. (2018). effects of trying ‘not to move’ instruction on cortical load and concurrent cognitive performance. psychological research, 82(1), 167–176. https://doi.org/10.1007/s00426-017-0928-9 lindner, m. a., eitel, a., strobel, b., & köller, o. (2017). identifying processes underlying the multimedia effect in testing: an eye-movement analysis. learning and instruction, 47, 91–102. https://doi.org/10.1016/j.learninstruc.2016.10.007 mayer, r. e. (2009). multimedia learning. cambridge university press. maylor, e. a., & wing, a. m. (1996). age differences in postural stability are increased by additional cognitive demands. the journals of gerontology series b: psychological sciences and social sciences , 51(3), 143–154. https://doi.org/10.1093/geronb/51b.3.p143 mccrudden, m. t., & schraw, g. (2007). relevance and goal-focusing in text processing. educational psychology review, 19, 113–139. https://doi.org/10.1007/s10648-006-9010-7 o’sullivan, p. b., dankaerts, w., burnett, a. f., farrell, g. t., jefford, e., naylor, c. s., & o’sullivan, k. j. (2006). effect of different upright sitting postures on spinal-pelvic curvature and trunk muscle activation in a pain-free population. spine (phila pa 1976), 31(19), e707-712. https://doi.org/10.1097/01.brs.0000234735.98075.50 o’sullivan, p. b., grahamslaw, k. m., kendell, m., lapenskie, s. c., möller, n. e., & richards, k. v. (2002). the effect of different standing and sitting postures on trunk muscle activity in a pain-free population. spine (phila pa 1976), 27(11), 1238–1244. https://doi.org/10.1097/00007632-200206010-00019 paivio, a. (1986). mental representations: a dual coding approach. oxford university press, clarendon press. patel, p. j., & bhatt, t. (2015). attentional demands of perturbation evoked compensatory stepping responses: examining cognitive-motor interference to large magnitude forward perturbations. journal of motor behavior, 47(3), 201–210. https://doi.org/10.1080/00222895.2014.971700 pellecchia, g. l. (2003). postural sway increases with attentional demands of concurrent cognitive task. gait & posture, 18(1), 29–34. https://doi.org/10.1016/s0966-6362(02)00138-8 qiu, j., & helbig, r. (2012). body posture as an indicator of workload in mental work. human factors, 54(4), 626–635. https://doi.org/10.1177/0018720812437275 rankin, j. k., woollacott, m. h., shumway-cook, a., & brown, l. a. (2000). cognitive influence on postural stability: a neuromuscular analysis in young and older adults. the journals of gerontology: series a: biological sciences and medical sciences , 55(3), m112–m119. https://doi.org/10.1093/gerona/55.3.m112 reilly, d. s., van donkelaar, p., saavedra, s., & woollacott, m. h. (2008). interaction between the development of postural control and the executive function of attention. journal of motor behavior, 40(2), 90–102. https://doi.org/10.3200/jmbr.40.2.90-102 rival, c., ceyte, h., & olivier, i. (2005). developmental changes of static standing balance in children. neuroscience letters, 376(2), 133–136. https://doi.org/10.1016/j.neulet.2004.11.042 roig-maimó, m. f., mackenzie, i. s., manresa-yee, c., & varona, j. (2018). head-tracking interfaces on mobile devices: evaluation using fitts’ law and a new multi-directional corner task for small displays. international journal of human-computer studies, 112, 1–15. https://doi.org/10.1016/j.ijhcs.2017.12.003 sakaguchi, m., taguchi, k., miyashita, y., & katsuno, s. (1994). changes with aging in head and center of foot pressure sway in children.international journal of pediatric otorhinolaryngology, 29(2), 101–109. https://doi.org/10.1016/0165-5876(94)90089-2 schaefer, s., schellenbach, m., lindenberger, u., & woollacott, m. h. (2015). walking in high-risk settings: do older adults still prioritize gait when distracted by a cognitive task? experimental brain research, 233(1), 79–88. https://doi.org/10.1007/s00221-014-4093-8 scheiter, k., schubert, c., schüler, a., schmidt, h., zimmermann, g., wassermann, b., krebs, m.-c., & eder, t. (2019). adaptive multimedia: using gaze-contingent instructional guidance to provide personalized processing support. computers & education, 139, 31–47. https://doi.org/10.1016/j.compedu.2019.05.005 schmid, m., conforto, s., lopez, l., & d’alessio, t. (2007). cognitive load affects postural control in children. experimental brain research, 179(3), 375–385. https://doi.org/10.1007/s00221-006-0795-x schmid, m., conforto, s., lopez, l., renzi, p., & d’alessio, t. (2005). the development of postural strategies in children: a factorial design study. journal of neuro engineering and rehabilitation, 2 (1), 29. https://doi.org/10.1186/1743-0003-2-29 schnotz, w., & bannert, m. (2003). construction and interference in learning from multiple representation. learning and instruction, 13(2003), 141–156. https://doi.org/10.1016/s0959-4752(02)00017-8 schnotz, w., & wagner, i. (2018). construction and elaboration of mental models through strategic conjoint processing of text and pictures. journal of educational psychology, 110(6), 850–863. https://doi.org/10.1037/edu0000246 sims, v. k., & hegarty, m. (1997). mental animation in the visuospatial sketchpad: evidence from dual-task studies. memory & cognition , 25(3), 321–332. https://doi.org/10.3758/bf03211288 stelmach, g. e., populin, l., & müller, f. (1990). postural muscle onset and voluntary movement in the elderly. neuroscience letters, 117(1), 188–193. https://doi.org/10.1016/0304-3940(90)90142-v stelzel, c., schauenburg, g., rapp, m. a., heinzel, s., & granacher, u. (2017). age-related interference between the selection of input-output modality mappings and postural control a pilot study. frontiers in psychology, 8, 613. https://doi.org/10.3389/fpsyg.2017.00613 sweller, j. (1988). cognitive load during problem solving: effects on learning. cognitive science, 12(2), 257–285. https://doi.org/10.1207/s15516709cog1202_4 theeuwes, j., & belopolsky, a. (2010). top-down and bottom-up control of visual selection controversies and debate. in v. coltheart, tutorials in visual cognition. psychology press, taylor & francis group. vandervelde, t. j., woollacott, m. h., & shumway-cook, a. (2005). selective utilization of spatial working memory resources during stance posture. neuroreport, 16(7), 773–777. https://doi.org/10.1097/00001756-200505120-00023 wainer, h. (1992). understanding graphs and tables. educational testing service. wainer, h., dorans, n. j., green, b. f., steinberg, l., flaugher, r., mislevy, r. j., & thissen, d. (1990). computerized adaptive testing: a primer. lawrence erlbaum associates publisher. wickens, c. d. (2002). multiple resources and performance prediction. theoretical issues in ergonomics science, 3(2), 159–177. https://doi.org/10.1080/14639220210123806 wickens, c. d., vidulich, m., & sandry-garza, d. (1984). principles of s-c-r compatibility with spatial and verbal tasks: the role of display-control location and voice-interactive display-control interfacing. human factors, 26(5), 533–543. https://doi.org/10.1177/001872088402600505 woollacott, m. h., & shumway-cook, a. (2002). attention and the control of posture and gait: a review of an emerging area of research. gait & posture, 16(1), 1–14. https://doi.org/10.1016/s0966-6362(01)00156-4 woollacott, m. h., & vandervelde, t. j. (2008). non-visual spatial tasks reveal increased interactions with stance postural control. brain research, 1208, 95–102. https://doi.org/10.1016/j.brainres.2008.03.005 zhao, f., schnotz, w., wagner, i., & gaschler, r. (2014). eye tracking indicators of reading approaches in text-picture comprehension. frontline learning research, 2(4), 46–66. https://doi.org/10.14786/flr.v2i4.98 zhao, f., schnotz, w., wagner, i., & gaschler, r. (2020). texts and pictures serve different functions in conjoint mental model construction and adaptation. memory & cognition, 48(1), 69–82. https://doi.org/10.3758/s13421-019-00962-0 codepen discussion publication frontline learning research special issue vol 8, no. 5 (2020) 92 104 issn 2295-3159 moral emotions and moral motivation beyond childhood: discussion to the special issue gertrud nunner-winklera, beate sodianb aludwigs-maximilians-universität münchen, germany doi: https.www.doi.org/10.14786/flr.v8i5.689 1. introduction research on the role of moral emotions in moral judgment, both in hypothetical dilemmata and in real-life moral decision making, has focused on preschool and elementary school age, with few studies spanning a larger age range, into adolescence and adulthood. the present special issue addresses a neglected area, the development of moral emotions and moral motivation in adolescence and adulthood. the focus is on the “happy victimizer phenomenon”, a pattern of emotion attributions to a moral transgressor that has been primarily observed in childhood, but that does not seem to disappear with age. we will begin by briefly reviewing 30 years of developmental research on moral emotion attribution and the “happy victimizer phenomenon”. this review will be followed by a discussion of the present papers. 2. a brief review of the literature 2.1 the happy victimizer phenomenon in childhood while research on early moral development has emphasized, over the last 30 years, that traditional descriptions of the young child as “pre-moral”, unaware of the nature of moral rules, and unable to take an agent’s intentions into account when evaluating his or her actions, were fundamentally wrong, there is reason to believe that young children’s understanding of moral emotions differs in important ways from older children’s and adults’, and that this has consequences for their understanding of moral agency and their moral motivation. in the first systematic study of moral emotion understanding in children, nunner-winkler and sodian (1988) found marked developmental change between the ages of 4 and 8 years, in children’s attributions of emotions to moral wrongdoers: when asked, for instance, how a child who had pushed another child from the swing, would “feel” after having committed this transgression, 74% of 4-year-olds, 40% of 6-year-olds, but only 10% of 8-year-olds attributed positive emotions, rather than moral emotions such as guilt, shame or empathy with the victim to the transgressor. this was not due to young children’s lack of knowledge about moral or empathic emotions in general. they attributed empathic emotions to a bystander who witnessed a harmful act, but they appeared to base their emotion attributions to a transgressor exclusively on the information about this agent’s desires. this developmental trend in emotion attributions, with negative or mixed feelings being attributed to the victimizer by the majority of the children from around the age of 7 years, has proved stable in subsequent research (see arsenio, 2014; krettenauer, malti, & sokol, 2008, for reviews). even with rigorous probing for opposite valence emotions most 4-year-olds continued to anticipate positive emotions in the victimizer (arsenio & kramer, 1992). in contrast, when asked how they would feel themselves if they had committed a transgression, even young children attributed negative emotions more often to themselves than to another person. however, more than 50% of the 5to 6-year-olds still responded with positive emotions when answering for themselves (keller, lourenco, malti, & saalbach, 2003). a recent study by gummerum, lopez-perez, ambrona, rodriguez-cano, dellaria, smith, & wilson (2016) found that executive function (inhibition) and counterfactual reasoning were associated with the happy victimizer phenomenon, indicating that a tendency to respond impulsively and a limited ability to consider hypothetical alternatives may contribute to the happy victimizer response pattern. however, this response pattern can not be reduced to executive demands of the tasks. theoretical explanations of the happy victimizer phenomenon in children have accounted for the sharp developmental trend by interpreting it in the broader context of children’s developing understanding of the mind. from the point of view of children’s reasoning about desires (as part of their theory of mind), the hv emotion attribution pattern indicates an understanding of the subjectivity of desires: while 3-year-olds often make their emotion attributions dependent on the objective valence of an action outcome, 4-year-olds consistently attribute positive emotions to an agent when his desire was fulfilled, and negative emotions when this was not the case (yuill, perner, pearson, peerbhoy, & van den ende, 1996). early cognitive interpretations of the happy victimizer pattern (nunner-winkler & sodian, 1988 ; yuill et al., 1996) have argued that young children, when forming a first understanding of the subjectivity of desires, infer emotional reactions exclusively from the correspondence between desires and action outcomes. only later, they begin to take other relevant determinants of emotions into account, such as the violation of moral norms or the suffering of the victim. different explanations have been proposed for why it is difficult for young children to integrate their knowledge about moral norms with their attribution of emotions to a victimizer. one reason may lie in young children’s failure to understand emotional conflict (arsenio & kramer, 1992): victimizers’ emotional reactions are likely to be deeply ambivalent, with happiness over the satisfaction of one’s wicked desires being intermixed with sadness over/ empathic concern for the victim’s suffering and/or regret and shame with regard to one’s own behaviour. research on children’s concepts of emotions has shown that ambivalent or conflicting emotions are generally conceptualized late in development, only around the age of 8 years (pons & harris, 2005), even though even infants appear to experience emotional ambivalence, for instance when reacting to their mother’s return after a brief separation in the strange situation. similarly, self-evaluative moral emotions (shame and guilt) have been demonstrated in 2-year-olds (e.g., vaish, carpenter, & tomasello, 2016), with effects on prosocial behaviour in the case of guilt. there is a long developmental lag between children’s first experience of complex emotions and emotional conflict and their conceptual understanding of such emotions and their antecedents. other theorists have interpreted children’s developing understanding of moral emotions within the broader framework of developmental change in their understanding of agency (see krettenauer et al., 2008). harris (1989) argued that an understanding of moral emotions requires a shift from seeing people as agents to seeing them as observers of their own agency. the process of internalizing an external audience allows children to view their own and others’ actions from a third-person perspective, and this perspective is necessary to integrate self-evaluative emotions with simple emotional reactions to a desire-action match or mismatch. another proposal (krettenauer et al., 2008; sokol, 2004; sokol & chandler, 2003) focuses on the development of an understanding of autonomous agency and the free will. young children may not understand that an agent who acts based on his desire has a choice to act according to this desire or not. perner (1991) argued that children may progress from a non-representational to a representational understanding of desires. a representational understanding is required to comprehend, for instance, that agents can change their desires, and thus are not helplessly exposed to the frustration of seeing their desired course of action thwarted. such a representational understanding of the mind is reached around the age of 4 years. krettenauer et al. (2008), however, argue that a higher level of understanding of the interpretive mind is necessary to achieve an understanding of autonomous agency that allows for a deliberate decision between fulfilling selfish desires and acting in accordance with moral norms. such an interpretive understanding of agency involves a rudimentary understanding of interpretive frameworks. developmental research on understanding social stereotypes (pillow & henrichon, 1996) and interpretive processes (carpendale & chandler, 1996) has shown that a first understanding of interpretive frameworks emerges around the age of 7 years. an understanding of autonomous agency and the deliberate choice between moral and immoral courses of action is assumed to require a notion of the interpretive mind that enables agents to weigh alternatives based on rational argument. consistent with this theory, strong correlations were found between children’s performance on the happy victimizer task and independent measures of an interpretive theory of mind (sokol, 2004; sokol, chandler, & jones, 2004). these associations could not be accounted for by the more basic ability to represent two different aspects of a situation or event at the same time (sokol, 2004). in contrast, the relation of hv emotion attributions and counterfactual reasoning found by gummerum et al. (2016) is consistent with the idea that conceiving of alternative interpretations of a given set of phenomena is linked to moral emotion understanding via a conception of autonomous agency. 2.2 the happy victimizer pattern in adolescence and adulthood while there is a clear age trend towards increasingly morally oriented emotion attributions in childhood, several studies indicate that the happy victimizer pattern does not disappear in middle childhood (see keller et al., 2003). murgatroyd and robinson (1993) introduced a variation of the original paradigm in which an authority figure witnessed the transgression. under these conditions, the attribution of fear (of sanctions) was more frequent than the attribution of sadness or moral emotions. when the witness mistakenly displayed a positive reaction (e.g., ‘good of you to pick up litter’, when a student crumpled up a classmate’s homework), up to 40% of adolescents and young adults (college students) attributed happy emotions to a transgressor who gained some personal benefit through his transgression (murgatroyd & robinson, 1997). about 30% of young adults failed to attribute any moral emotions, even in an open-ended format, indicating a strong dependency of emotion attributions in adolescence on the witness, rather than on moral norms. without the influence of an onlooker, however, only a very small proportion of adolescents appear to show a “happy victimizer pattern”. a study by krettenauer and eichler (2006), using the standard assessment procedure, found almost no evidence for a happy victimizer pattern in 13to 19-year-olds: less than 10% of the participants anticipated “no bad” feelings following transgressions of varying severity (giving false testimony to get a desired job; absconding from a traffic accident because of being drunk; not returning a found wallet; stealing from a flea market salesman). moreover, intensity of emotion ratings, using a 6-point scale (extending from “not bad” to “extremely bad”), with age became increasingly consistent with moral judgment and justification, as well as with certainty ratings (‘how sure are you about your moral judgment’?), thus indicating an increasing integration of norm orientation and emotion attribution in adolescence. a long-term longitudinal study of moral reasoning and moral motivation was carried out by nunner-winkler (1999; 2009), as part of the project logic (weinert & schneider, 1999; schneider & bullock, 2009), a longitudinal study of originally n=200 3-4 year old children, n=172 participants at age 18 and n=153 at age 23 were presented with three moral conflicts (breaking a promise given to the first customer in a market transaction; lying for a career; keeping a lost good) and one moral dilemma (allowing the blaming of an innocent person to favour one’s own friend). they were asked to specify and justify what they themselves (or the same sex protagonist in the career story) would do and how they (or the protagonist) would feel in the role of the agent and the counter role of the victim of another’s transgression. answers were scored as reflecting moral versus pragmatic concerns using the following criteria: reference to moral principles or harm done vs. personal profit; asymmetries in emotions ascribed to self in the role of agent vs. victim (e.g., happy as transgressor vs. indignant as victim); intensification or mitigation of action decisions or emotion attributions (e.g., “i would never do such a thing” vs. “i’d probably not do that, i assume”). based on the combination of these criteria strength of moral motivation was rated as high, middle or low according to the predominance of moral vs pragmatic aspects expressed. 25.6% of the participants were classified as low and 40.3% as high at the age of 18 years, and 18.3% as low and 47.1% as high at age the age of 23 years. a slightly higher percentage of low moral motivation (35%) was obtained with the same vignettes and the same coding procedure in a cross-sectional study of n=203 15-16-year old german students (drawn from former east and west germany and higher and lower educational tracks (nunner-winkler, meyer-nikele, & wohlrab, 2006). the longitudinal data showed moderate stability of moral motivation between the ages of 18 and 23 years (r=.36). there was, however, no predictive relation between childhood and adolescence: on the contrary, moral motivation at the age of 9 years was negatively correlated with moral motivation at the age of 18 years (r=-.21). this may be partly due to the difference in assessment methods. while the assessment of moral motivation in childhood was exclusively based on moral emotion attribution, an aggregate score of moral reasoning and emotion attribution was used in adolescence and adulthood. alternatively, the low stability of moral emotion attributions in childhood years might reflect differences in speed of children’s socio-cognitive development, whereas in adolescence the relative importance attributed to moral versus non-moral values becomes more influential and value orientations are more liable to change in response to varying circumstances and social context factors. identity formation plays a crucial role in adolescence. the impact processes of identity formation exert on adolescents' moral commitment is convincingly illustrated by findings on the connection between moral motivation and sex role identification (nunner-winkler 2009, nunner-winkler et al., 2006). no difference in strength of moral motivation was found among participants with low gender identification. however, among highly gender-identified participants there were significantly more males with low moral motivation as well as male logic participants who had experienced a decrease in moral motivation between ages 9 and 23. this finding reflects the fact that participants' male stereotypes were predominantly morally aversive (e.g. assertive, unwilling to admit to weakness, shrewd). 2.3 the relation of moral emotion attributions to prosocial and antisocial behaviours if moral emotion attribution is, in fact, a valid indicator of the strength of moral motivation, then there should be a relationship between moral emotion attributions in hypothetical scenarios and real-world prosocial and antisocial behaviours. in the longitudinal study logic, 7-year-olds’ amoral emotion attributions combined with low values in shyness predicted egotistic ruthless behaviour in a lab-assessment of sharing behaviours (asendorpf & nunner-winkler, 1992). further, moral emotion attributions in adolescence and adulthood (but not in childhood) predicted antisocial behaviours at age 23, independently of conscientiousness and agreeableness. interestingly, conscientiousness at age 12 contributed to the development of moral emotion attributions at age 18 and these, in turn, predicted change in conscientiousness at the ages of 18 and 23 years (krettenauer, asendorpf, & nunner-winkler, 2013). similarly, krettenauer and eichler (2006) reported a correlation between adolescents’ strength of self-attributed moral emotions in hypothetical situations and their self-reported real-life delinquent behaviours even when social desirability was controlled for. consistent with these findings, a meta-analysis of 42 cross-sectional studies of moral emotion attribution and pro and/or antisocial behaviours (with 80000 participants aged 4 to 20; malti and krettenauer, 2012) yielded evidence for significant associations between emotion attributions and social behaviour in all age groups under study, not restricted to adolescence. there were small size relations of emotion attribution with prosocial behaviour and moderate size relations of emotion attributions with antisocial behaviour, emphasizing the role of moral emotion attributions such as guilt, regret (or sadness) in morally relevant real-world behaviours. age did not moderate the relation between moral emotion attribution and social behaviour. thus, despite developmental change in children’s understanding of moral emotions, moral emotion attributions appear to reflect at all ages individual differences in moral motivation. self-attributed moral emotions were more strongly related to antisocial behaviour than other-attributed emotions. furthermore, measures of the intensity of moral emotions showed larger effect sizes than mere positive or negative ratings. in sum, theoretical views of the important role of moral emotion attribution in moral development are supported impressively by research on the relation of emotion attribution and social behaviour. 3. the present papers despite its significance for moral choices of real-world importance, moral emotion attribution in adolescence and adulthood is an underresearched area. while most studies have focused on developmental change in childhood, sometimes including an adolescent or adult comparison group, systematic research into moral emotion attribution in adolescence and adulthood is rare. the present papers are the first to specifically address this issue, with a focus on the “happy victimizer” pattern, its frequency, transsituational stability and theoretical significance beyond childhood. 3.1 paper 1 (heinrichs, gutzwiller-helfenfinger, latzko minnameier, & döring) paper 1 “happy victimizing in adolescence and adulthood – empirical findings and further perspectives” broadly investigated the incidence and transsituational stability /variation of the hv pattern in adolescence and adults in four empirical studies. in study 1 a large, representative, standardized sample of 4th, 7th, and 9th grade students (n= 4935) gave written responses to test questions concerning two vignettes on moral transgressions (breaking a contract / keeping money instead of returning it to the owner). the test questions were “what would you (participant) do in this situation? why would you do it? how would you feel?” there was an increase in happy victimizer patterns from 4th to 9th grade, but only a minority of 9th graders (27% / 12%) showed the hv pattern. the responses to the two vignettes were not correlated, i.e., the responses appeared to be situation specific. rule enforcement was apparently not assessed, i.e., we do not know whether participants who said they would transgress saw the rule as personally binding. study 2 introduced a situation of passive moral temptation (getting too much change money) and asked what the protagonist should do (instead of will do). n=331 14-year-old secondary students participated. only 72% enforced the rule, i.e., they said, it was not okay that the protagonist kept the money. 10.9% showed the hv pattern, judging the protagonist’s transgression as wrong while attributing positive emotions to the protagonist. the predominant justification for emotion attributions in all these patterns was hedonism. study 3 addressed moral emotion attributions in adults. n=271 students of economics and business education gave written responses to vignettes about morally relevant decisions in business contexts. participants were asked to decide about a course of action for themselves and to justify their choice, as well as to judge how they would feel. hv patterns emerged in 31% to 49% of the participants. depending on the context, up to half of the participants who chose to transgress were anticipating happy emotions. correlations among the scenarios ranged from .17 to .35, indicating substantial cross-situational variation. there was no assessment of rule validity. study 4 included a fuller assessment of rule understanding and adults’ own decisions relative to a perceived norm. n=233 students judged two situations of temptation. 35% approved of the transgression in the “start-up” / 77% in the “jana” story. emotion attributions showed 29% / 10% positive emotion attributions following transgression choices for self. it is unclear whether these reflect hv patterns in the strict sense, since many of the trangressors may not have acknowledged the rule as personally binding. again, the findings indicate situation specificity of moral emotion attribution. commentary paper 1 clearly demonstrated the incidence of the hv phenomenon beyond childhood. moreover, situational factors influencing hv responses, and variations in response styles emerged. incidence of hv beyond childhood. the findings are clear: hv can still be found in later development. the present studies yielded between 10% and 50% hv-like responses in all age groups, with a tendency for a higher proportion of hv responses in adults than in adolescents. however, developmental trends can not be inferred from these comparisons across studies, since different situational types of transgressions and different questioning procedures were used in the different studies. consistency across situations. the frequency of hv like answers varied considerably between vignettes even when the same age group was tested and the same measurement used. moreover, there was also substantial intra-individual variation across situations. this variation is hard to interpret, since the vignettes used varied in several dimensions. more hv responses might be expected if the transgression arose from a passive temptation than from an intentional plan (e. g. study 3: keeping change vs. tv sale), if the damaged party is a large anonymous company rather than a concrete individual (e. g. study 3 travel cost vs keeping change), if the situation makes it easy to justify the transgression or to put the blame on the victim (e. g. study 1: the first customer had beaten down the price, should have brought the money along in the first place cf. paper 2 disengagement strategies), and if the damage caused is low. moreover, different assessment procedures were used. study 1 asked for a justified action decision and an emotion ascriptions to self (what would you do? why? how would you feel? why?) hv was coded when participants expected to feel good for hedonistic reasons about the profit gained by an immoral action decision. study 2 and study 4 first requested 'deontic judgments' (what should protagonist do? why? how does s/he feel? why), and 'self judgement' (what would you do? why? how would you feel? why?). then they presented a hypothetical wrongdoer and asked for a moral judgment and an emotion ascription ('is it ok or not ok that protagonist transgressed? why? how does protagonist feel? why?). as documented by keller et al. (2003), hv responses are ascribed more often to hypothetical wrongdoers than to self. in contrast, the present study 4 found “significantly more morally appropriate responses to the the perpetrator in the classical hv situation than in the self judgement situation” (p-14). this finding may be due to the fact that many participants judged the transgression to be ok (35% in the start-up, 73% in the “keeping change” situation). as the present authors acknowledge, a positive emotion attribution to a transgressor does not, strictly speaking, conform to the hv pattern, if the rule that was transgressed is not seen as personally binding. given that there is no longer a collective consensus on a hierarchy of values in modern pluralistic societies, people may differ in how much importance they attribute to money, professional standing, moral integrity. thus, the degree of temptation experienced will differ not only across stories but also across individuals. some people attribute very high / very low importance to moral integrity. these persons can be expected to consistently attribute negative / positive emotions to wrongdoers. most people, however, will evaluate the costs or gains of transgressing according to context features and their individual value preferences. thus, moral emotion attribution in adulthood cannot be interpreted independently of adult participants’ evaluation of moral rules and value systems. response patterns. the present authors found several new response types. among participants judging a norm as valid there are happy and unhappy victimizers (transgress and feel good / bad), as well as moralists (conform and feel good / bad). among participants who did not accept the norm as binding they found – depending on the test question used “deontic happy transgressors”, “happy transgressors self”, and “happy transgressors misdeed”. according to the authors these new patterns reflect the openness of the situation in which an intention has yet to be formed, whereas in the classical hv paradigm a moral judgment is requested after the transgression has already been committed (p.10). note however, that the original hv paradigm was just as open with respect to moral judgment which was requested in the situation of temptation. rather, the new response patterns may reflect the fact that the norms presented in the present research were by no means regarded as personally binding by all participants, whereas this was the case for the norms used in the classical hv research with children. the new patterns documented raise the question as to their meaning. it is not clear, for instance, whether unhappy victimizing (uv) indeed stands for true regret or merely serves impression management in the testing situation. further, one may wonder whether happy moralizing versus unhappy moralizing may indicate high versus low moral commitment. research on moral exemplars (persons who all through their lives remained true to their moral convictions despite grave costs incurred) has not reported any special positive feelings about having acted morally right. for them doing what is right was a matter of course. they did, however, mention feelings of regret or guilt about having inadvertently caused harm to their family members (e. g., children suffered when for moral reasons they gave up well paying jobs) (colby & damon,1992). such considerations may also play a role in adults' attribution of mixed emotions (paper 3). it should also be noted that victimizing responses may be affected by the authors’ “deontic judgments”: ‘what should the protagonist do’? these request action recommendations. as has been argued with respect to kohlberg’s description of children’s pre-conventional stage, these questions confound the cognitive and the motivational dimension (nunner-winkler, 1999). thus, participants might interpret ‘should’ in the test question in prudential rather than in moral terms. in other words, they may recommend that the protagonist should do not what is right but what benefits him most. in sum, paper 1 extensively documents the existence of hv responses beyond childhood, consistent with previous research. thus, hv is not a transitional phenomenon of childhood cognitive development and needs to be interpreted in a broader context of moral judgment and moral motivation. further, the studies document high response variability across individuals and situations an innovative finding that can not yet be conclusively interpreted. finally, by describing new response patterns, the studies open a whole set of promising research questions. 3.2 paper 2 (heinrichs, kärner, & reinke) paper 2 “an action-theoretical approach to the happy victimizer pattern – exploring the role of moral disengagement strategies on the way to action” advances an action theoretical approach to moral decision making which focuses on the phase of forming an intention: moral distancing strategies (mds) are conceptualized as cognitive controlor volitional strategies which support an individual’s dealing with an inner conflict or ambivalence. the perceived moral intensity of a situation may impact the process of forming an intention and acting. situations of high moral intensity may evoke the conscious, reflective application of cognitive control strategies, whereas more automatic processes of moral motivation may dominate in situations of low moral intensity. the research questions concerned the extent to which adults apply victimizing strategies in situations of low moral intensity, the situation specificity of victimizing decisions, the application of mds in situations of low moral intensity, and the situation specificity of the application of mds. n=587 university students (economics and teacher students) were tested in hypothetical low intensity moral transgression situations. participants were asked what they would do and how they would feel. justifications for the decisions were coded for mds. an hv pattern was observed in 11% to 19% of the participants. among the other response patterns, the happy moralizer was the most frequent pattern. there was moderate transsituational consistency. mds were analyzed for different types of strategies. some differences emerged descriptively between the mds associated with hv and uv response patterns. commentary paper 2 replicates core results presented in paper 1: hv responding occurs beyond childhood and responses vary across situations and – to a moderate degree – between individuals. besides, paper 2 adds a very interesting expansion in the hv paradigm – the study of mds in situations of moral temptation. the present findings indicate that mds vary across situations but that there are also some person-specific preferences for particular types of mds. the data prompt some speculations which – in addition to the rich research program suggested by the authors – might be followed up in future studies. in criminal law, two types of defenses are distinguished– justification (the act is justifiable or at least not that reprehensible, e.g., served higher moral ends, was not so bad) and excuse (the act is bad but the agent is pardonable, e.g., acted under constraint). the 'travel costs' story draws more and allows for more justifications (mds types 2 and 3 presenting seemingly fair solutions, e.g., paying the friend, belittling the harm affecting a large company). the 'change' story produces more excuses – justifications are hard to come up with (mds types 6 and 7). despite the low numbers involved the finding of differences in mds used by hvs and uvs is suggestive. hvs more often use denial of the wrongness of the act or of own responsibility (travel cost – md 2: trivializing the behaviour, change md type 7 – putting the blame on the victim), uvs, in contrast, (at least implicitly) acknowledge the wrongness of the act but appeal to stressful circumstances (travel cost/ change: md type 6 pointing to own needs). if this finding were robust, uv might indicate higher moral commitment: participants admit having done wrong and hope for lenience. 3.3 paper 3 (gutzwiller-helfenfinger & latzko) paper 3 “happy victimizing in emerging adulthood: reconstruction of a developmental phenomenon?” deepens our understanding of moral emotion attribution in adulthood. n=285 pre-service teachers received a paper & pencil questionnaire consisting of 3 vignettes (change money; the motorbike; lying to a customer). moral rule understanding was assessed by a “deontic” judgment plus justification, extended emotion ascription (choice between positive, negative and mixed emotions, specification of the chosen emotion(s) and justification). finally, a self-judgment was followed by an emotion attribution. the vignettes were followed by an assessment of the “moral self”. the results showed a “pure” hv pattern only in 2.2% of the participants, however, a mixed emotion attribution in roughly 40%/ 49%/27%, depending on the story. most justifications for the mixed emotions hv pattern were hedonistic. several types of hedonistic justifications emerged, the protagonist being “happy” for enjoying a profit, e.g., having money, or the protagonist being “justifiably” happy because the victim was “stupid”. some of the mixed emotion attributions contained a hedonistic justification for the positive, and a moral (e.g., bad conscience) justification for the negative emotion (unfortunately, no detailed analysis of these mixed patterns is provided). the authors provide a further analysis of the positive and negative emotions by levels of complexity, showing that in the case of negative emotions the large majority was unspecific (“bad, uncomfortable”), but roughly 30% of the participants explicitly mentioned a bad conscience. the analysis of the moral self scales revealed that participants categorized as hv ascribed more importance to honesty and truthfulness to themselves than participants categorized as no-hv. commentary the study showed that “pure” hv responses occur only very rarely, when adults are offered the option of emotional ambivalence: most adults who see a moral rule as binding and expect a protagonist to transgress expect him or her to experience mixed emotions. ambivalent emotions can vary in level of complexity and in terms of the recognition of moral concerns. justifications need to be analyzed more deeply. it is possible that participants, who expect emotional ambivalence, may be more strongly morally motivated/ and or more advanced in their moral understanding than participants who expect a single emotional reaction. it is difficult, however, to use the mixed emotion pattern as a straight measure of moral motivation, since the emotion attribution question was originally designed to yield a spontaneous decision for one or the other emotional valence, which is supposed to be an indicator of the strength of moral motivation. another problem of interpretation arises from the substantial proportion of participants who said the rule should be transgressed. these cannot be included in an analysis of hv patterns. the finding points to a problem associated with the use of situations of low moral intensity. if the rule is not seen as binding, then, strictly speaking, the instrument is not suitable for measuring moral motivation. an interesting finding is the increase in mixed hv responses if the 'deontic' judgment is added to the classical hv judgment (tables 1 and 3). the 'should' question and the 'is wrong' question obviously tap different meanings: the first one expresses an action recommendation which may comprise not only moral considerations but also prudential concerns (evaluation of potential costs, risks and gains). the second question unequivocally requests a moral judgment – it is the one that more adequately could be labeled 'deontic'. another interesting finding is the higher importance that hvs, compared to non hvs, ascribed to honesty and truthfulness. they in fact do openly admit to their lack of moral concerns. after all, even young children know that a remorseful wrongdoer is better than a joyful one. conversely, hvs’ straightforwardness suggests it might be worthwhile to more carefully explore the potential influence social desirability might have on moral emotion attributions. 3.4 paper 4 (minnameier) paper 4 “how to explain the happy victimizer in adulthood” addresses the hv phenomenon from the point of view of rational choice theory. n=481 university students (economists vs. teacher students) were tested with versions of the prisoner’s dilemma (wall street game; community game). participants were given a choice between strategies (cooperate vs. defect), asked to explain their choice, rate how they felt about the decision, and explain their feelings. the cooperation strategy was less frequent in economists than in teacher students. there were between 55% (economists) and 31% (teacher students) happy victimizers, and conversely between 36% and 67% happy moralists. a further analysis distinguished between strategic moralists and hv in the strict sense. the strategic moralists reason that the situational constraints of the dilemma made it necessary for them to pursue their own interests. if strategic moralists are separated, there is a proportion of 47% “pure” hv among economics students and of 27% among teacher students. thus, the study demonstrates a high proportion of hv or hv-like patterns in a prisoner’s dilemma situation, with marked intraand interindividual differences, reflecting situational determinants, as well as (possibly) disciplinary variation in strategic thinking and moral judgment. commentary minnameier follows the core claim of rational choice theory: agents will always try to maximize utility by pursuing their interests – be these selfish or other-regarding. according to the author, economists' higher rate of defection results from their realizing that – without social sanctions – the conflict of interests cannot be solved. in such situations it is morally justified that agents pursue their own interests as long as they accept others doing so, too. thus, in line with this neo-kohlbergian “stage 2a” moral principle, defecting is morally acceptable behaviour. however, modern functionally differentiated societies comprise morally neutral domains which allow pursuing one's interests as long as the more global moral/legal framework is respected. pd behaviour belongs to this sphere, is non-moral behaviour. therefore, minnameier is right in claiming that defecting does not necessarily indicate a lack of moral motivation. nevertheless, cooperation might indicate special moral commitment – the willingness to follow moral convictions even when this is not mandatory and no sanctions impend. thus, the ‘naive’ understanding of cooperation that many teachers displayed might not as minnameier asserts reflect a cognitive deficiency but rather “habitual action schemes” (p.4) that might be characteristic of persons with high moral motivation. these would be based on fundamental premises – basic trust, appreciation of equality, and fairness. such moral attitudes are not contrary to rationality as minnameier argues. this claim is supported by game theory (cooperation in the tit-for-tat strategy yields the best results), by analyses of political systems (e. g. an analysis of 20 regional italian governments found efficiency strongly correlated with citizens’ solidarity, tolerance, trust, putnam et al., 1993) and by studies on economic growth (e. g. in 160 countries a strong correlation was found between prosperity and residents’ happiness with absence of corruption: delhey, 2002).in fact, minnameier’s own data confirm the claim that cooperation based on trust is advantageous. on average, economists will end up with 20 euros, given that most other economists will defect whereas teacher students will end up with 50 euros given that most others will cooperate. yet, irrespective of the long-term advantage of cooperation, morally motivated people might appreciate equal-distribution principles as an intrinsic value. after all, the idea of equality is constitutive for a secularized moral understanding. also, those who care about morality will disapprove of unfairness and in pds the greater profit results from exploiting the trust advanced by the partner. 3.5 concluding remarks the present papers contribute importantly to a fuller understanding of the hv phenomenon. despite large variation in the incidence of hv-type responses in adolescents and adults, depending on the task format, all studies documented the classic hv pattern (the participant regards a rule as binding and nevertheless attributes exclusively positive emotions to the transgressor) in a small minority of adults (up to 10%). there was no clear age-trend from adolescence to adulthood. given the association of hv-responding with antisocial action (malti & krettenauer, 2012), this minority may be at risk for deviant behaviours. it is possible that the hv pattern in the strict sense is also a sign of cognitive immaturity in a small minority of adolescents and adults who fail to develop an interpretive theory of mind, including an understanding of humans as autonomous agents, whose actions are subject to interpretation and evaluation. depending on task format and context, up to 50% of hv-like response patterns were observed in adults. situations of low-intensity moral demands that allow for moral distancing strategies, especially situations of competition that are ambiguous with respect to norms of competitive “fair play” versus norms of cooperation, elicit a high proportion of hv-like responses. generally, emotion attribution in situations of norm violation cannot be interpreted without evidence on participants’ construal of the norms in question. if participants do not regard the norm as binding, and justify this judgment with pragmatic reasons, then an emotion attribution that focuses solely on the agent’s desires is to be expected. since there is considerable variation in norm understanding among adolescents and adults, it appears that a compound score of moral motivation as used by nunner-winkler (2009), including both norm evaluation and emotion attribution, may provide a better indicator of moral motivation than emotion attribution per se. with respect to emotion attribution, paper 3 clearly indicates that adults will generally be aware of emotional ambivalence and choose mixed emotions if this option is provided. it is hard to interpret the choice of emotional ambivalence in terms of moral motivation without an in-depth assessment of the emotional qualities anticipated by the participants, and the justification for these attributions. in contrast, the original rationale for using emotion attribution as an indicator of moral motivation was that the spontaneous choice of a specific emotional valence may reflect dispositional tendencies that may also contribute to real-world moral decision making. thus, offering the choice of emotional ambivalence may make emotion attribution invalid as an indicator of moral motivation. the present studies revealed a lack of consistency in moral emotion attributions in adulthood that had not been noted before. responses varied greatly across situations. it appears that the determinants of such diversities are not yet fully understood and deserve further explorations. among the dimensions that should be further explored are the impact of positive vs. negative duties (e.g. breaking a promise vs not returning a lost good), the severity of the harm done (e.g. not returning change vs giving false testimony at court or absconding from an accident); the status of the victim (e.g. personal friend, unknown individual, anonymous organization), the social context of the transgression (action takes place secretly or in public, a (dis)approving authority is or is not present). furthermore, the present authors used a variety of modifications of the interview procedures. the results document the unexpectedly large influence details of the wording of questions and the order of their presentation may have on participants' reactions. further systematic research on methodological variations seems warranted. in the classical hv paradigm, for instance, justified moral judgments were requested in the situation of temptation and emotion attributions after the protagonist had transgressed. in contrast, some recent studies present the transgression before requesting its evaluation or do not request a moral judgment at all, assuming broad agreement on the immorality of the act presented. heinrichs et al. (paper 1, study 2) start off with ‘deontic’ and ‘self judgements’. krettenauer and eichler (2006) added an epistemic judgment. response formats may have effects as well. some studies allowed for open answers, others presented 2 to 6-point likert scales on which specific emotions (happy, ok, scared, sad, good, bad, not bad, a little bad, moderately bad, bad, very bad, extremely bad) had to be rated. maybe most importantly, the present authors have succeeded in linking up the hv phenomenon with other research areas. thus, they have taken up the recent debate on moral disengagement and analysed participants' strategies of vindicating deviant action decisions. and they have connected the study of moral emotions with behavioural economics thus opening up a way of experimentally testing whether emotion attributions to wrongdoers may predict real life behaviour, i.e. to what extent a motivational interpretation of hv patterns is warranted. finally, the present findings raise a number of issues relevant to moral education. if emotion attribution becomes increasingly situation-specific and contextualized in adolescence and adulthood, then what needs to be supported in education is the establishment of contexts that are conducive to the growth of moral motivation. it would, however, be a fatal mistake to focus interventions on moral emotions. such educational efforts might merely promote conformist public reactions. attempts at improving the moral climate in classrooms appear to be more promising. a good way – also suggested by the present authors is the establishment of just communities. studies have shown: in such democratic contexts, participants feel more bound to norms given they freely agreed to them, they feel more responsible for class mates and are more ready to support and help them (oser & althof, 1992; higgins, power, & kohlberg, 1984). acknowledgments beate sodian was supported by dfg grant so 213/27-3. footnotes 1 nunner-winkler & sodian (1988) have recently been labelled as proponents of a purely motivational interpretation of the hv phenomenon (gummerum et al, 2016; krettenauer et al., 2008). this is strange since they initially discussed their findings in terms of emotionand desire-understanding, i.e., with respect to young children’s emotion concepts and theory of mind. there has never been a debate on whether the hv-phenomenon had either cognitive or motivational significance. moral emotion attributions in young children may reflect an immature understanding of the determinants of emotions, and this understanding may undergo conceptual change. at the same time, the level of understanding of moral emotions that a child has may gain motivational significance in real-life moral conflict situations. references arsenio, w. (2014). moral emotion attributions and aggression. in m. killen & j. smetana (eds.). handbook of moral development (2 nd edition) (pp.235-255). new york: psychology press. arsenio, w. f., & kramer, r. (1992). victimizers and their victims: children's conceptions of the mixed emotional consequences of moral transgressions. child development, 63(4), 915-927. asendorpf, j.b. & nunner-winkler, g. (1992). children's moral motive strength and temperamental inhibition reduce their immoral behaviors in real moral conflicts. child development, 61, 1223-1235. carpendale, j. i., & chandler, m. j. (1996). on the distinction between false belief understanding and subscribing to an interpretive theory of mind. child development, 67(4), 1686-1706. colby, a. & damon, w. (1992) some do care. contemporary lives of moral commitment. new york: free press. delhey, j. (2002) korruption in bewerberländern zur europäischen union. soziale welt, 53, 345-366. gummerum, m., lópez-pérez, b., ambrona, t., rodríguez-cano, s., dellaria, g., smith, g., & wilson, e. (2016). children's moral emotion attribution in the happy victimizer task: the role of response format. journal of genetic psychology, 177(1), 1-16. harris, p. l. (1989). children and emotion: the development of psychological understanding . basil blackwell. higgins, a., power, c., kohlberg, l. (1989).the relationship of moral atmosphere to judgements of responsibility. in: w.m. kurtines, & l. gewirtz (eds.) morality, moral behavior, and moral development. basic issues in theory and research . (pp 74106) new york: wiley. keller, m., lourenco, o., malti, t., saalbach, h. (2003).the multifaceted phenomenon of 'happy victimizers': a cross-cultural comparison of moral emotions. british journal of developmental psychology 21, 1‒18. krettenauer, t., asendorpf, j. b., & nunner-winkler, g. (2013). moral emotion attributions and personality traits as long-term predictors of antisocial conduct in early adulthood: findings from a 20-year longitudinal study. international journal of behavioral development, 37(3), 192-201. krettenauer, t., & eichler, d. (2006). adolescents' self‐attributed moral emotions following a moral transgression: relations with delinquency, confidence in moral judgment and age. british journal of developmental psychology, 24(3), 489-506. krettenauer, t., malti, t., sokol, b. w. (2008). the development of moral emotion expectancies and the happy victimizer phenomenon: a critical review and application. european journal of developmental science, 2/3, 221‒235. malti, t., & krettenauer, t. (2013). the relation of moral emotion attributions to prosocial and antisocial behavior: a meta‐analysis. child development, 84(2), 397-412. murgatroyd, s. j., & robinson, e. j. (1993). children's judgements of emotion following moral transgression. international journal of behavioral development, 16(1), 93-111. murgatroyd, s. j., & robinson, e. j. (1997). children's and adults' attributions of emotion to a wrongdoer: the influence of the onlooker's reaction. cognition & emotion, 11(1), 83-101. nunner-winkler, g. (1999) development of moral understanding and moral motivation. in: f.e. weinert, & w. schneider (eds.) individual development from 3 to 12. (pp253-290) cambridge uk: cambridge university press. nunner-winkler, g. (2009). moral motivation from childhood to early adulthood. in: w. schneider & m.bullock (eds.) human development from early childhood to early adulthood. (pp. 91-118) new york: taylor &francis. nunner-winkler, g., meyer-nikele, m. wohlrab, d. (2006). integration durch moral. moralische motivation und ziviltugenden jugendlicher . kap. 4 erhebung und validierung moralischer motivation (pp.65-81). wiesbaden: verlag für sozialwissenschaften. nunner-winkler, g., meyer-nikele, m., & wohlrab, d. (2007) gender and moral motivation. merrill palmer quarterly, 53, 26-52. nunner-winkler, g., & sodian, b. (1988). children's understanding of moral emotions. child development, 59, 1323-1338. oser, f. & althof, w. (1992) moralische selbstbestimmung. stuttgart: klett cotta. perner, j. (1991). understanding the representational mind. cambridge, ma. mit press. pillow, b. h., & henrichon, a. j. (1996). there's more to the picture than meets the eye: young children's difficulty understanding biased interpretation. child development, 67(3), 803-819. pons, f., & harris, p.l. (2005). longitudinal change and longitudinal stability of individual differences in children's emotion understanding. cognition & emotion, 19(8), 1158-1174. putnam, r.d., leonardi. r. & nanett, r.y. (1993) making democracy work.civic traditions in modern italy. princeton, n.j.: princeton up. schneider, w., & bullock, m. (eds.). (2010). human development from early childhood to early adulthood: findings from a 20 year longitudinal study . psychology press. sokol, b. w. (2004 ). children’s conceptions of agency and morality: making sense of the happy victimizer phenomenon . (unpublished doctoral dissertation). university of british columbia, vancouver, canada. sokol, b.w., chandler, m.j. (2003). taking agency seriously in the theories-of-mind enterprise: exploring children's understanding of interpretation and intention.british journal of educational psychology monograph series ii, number 2 ‒ development and motivation, 125‒136. sokol, b. w., chandler, m. j., & jones, c. (2004). from mechanical to autonomous agency: the relationship between children's moral judgments and their developing theories of mind. new directions for child and adolescent development, 2004 (103), 19-36. vaish, a., carpenter, m., & tomasello, m. (2009). sympathy through affective perspective taking and its relation to prosocial behavior in toddlers. developmental psychology, 45(2), 534. weinert, f. e., & schneider, w. (eds.). (1999). individual development from 3 to 12: findings from the munich longitudinal study . cambridge university press. yuill, n., perner, j., pearson, a., peerbhoy, d., & van den ende, j. (1996). children's changing understanding of wicked desires: from objective to subjective and moral. british journal of developmental psychology, 14(4), 457-475. microsoft word flr proofs_rogat.docx frontline learning research vol. 10 no. 2 (2022) 1 21 issn 2295-3159 corresponding author: toni kempler rogat, purdue university, 100 n. university street, west lafayette, in 47906, usa, trogat@purdue.edu doi:https://doi.org/10.14786/flr.v10i2.863 a multidimensional framework of collaborative groups’ disciplinary engagement toni kempler rogat1, cindy e. hmelo-silver2, britte haugan cheng3, anne traynor1, temitope f. adeoye1, andrea gomoll4 & brenda k. downing1 1 purdue university, usa 2 indiana university, usa 3 menlo education research, usa 4 gomoll research & design, usa article received 17 may 2021 / article revised 12 march 2022 / accepted 27 july / available online 9 september 2022 abstract this research is aimed at developing novel theory to advance innovative methods for examining how collaborative groups progress toward productively engaging during classroom activity that integrates disciplinary practices. this work draws on a situative perspective, along with prior framings of individual engagement, to conceptualize engagement as a shared and multidimensional phenomenon. a multidimensional conceptualization affords the study of distinct engagement dimensions, as well as the interrelationships of engagement dimensions that together are productive. development and exploration of an observational rubric evaluating collaborative group disciplinary engagement (gde) is presented, leveraging the benefits of observational methods with a rubric specifying quality ratings, enabling the potential for analyses of larger samples more efficiently than prior approaches, but with similar ability to richly characterize the shared and multidimensional nature of group engagement. mixed-methods analyses, including case illustrations and profile analysis, showcase the synergistic interrelations among engagement dimensions constituting gde. the rubric effectively captured engagement features that could be identified via intensive video analysis, while affording the evaluation of broader claims about group engagement patterns. application of the rubric across curricular contexts, and within and between lessons across a curricular unit, will enable comparative studies that can inform theory about collaborative engagement, as well as instructional design and practice. keywords: engagement; collaborative learning; stem education; observational rubric rogat et al 2 | f l r 1. introduction in response to science standards, which call for students’ integrated understanding of stem content and practices, as part of collaborative activity (e.g., ngss, 2013; forsthuber et al., 2011), this research is aimed at advancing theory and methods toward understanding how collaborative groups come to productively engage in stem activities. in this research, we build from work on individual engagement that advances a multidimensional conceptualization from individual engagement (fredricks, et al., 2004) to the collaborative group context. to accomplish this, we need to consider interpersonal engagement and to account for collective group engagement practices. further, individual perspectives on engagement do not represent recent theoretical advances regarding the social and situated nature of engagement (gresalfi, et al., 2009; ryu & lombardi, 2015). other engagement-related frameworks, including engle and conant’s productive disciplinary engagement (pde) framework, have integrated developments from situative perspectives (danish & gresalfi, 2018; engle & conant, 2002; greeno, 2006; gresalfi, et al., 2009; hand & gresalfi, 2015; hickey, 2003), including assumptions that engagement is co-negotiated in collective interaction, evolves in moment-by-moment interactions, and is contextualized in activity systems. these activity systems are comprised of instructional opportunities that support and constrain engagement, given curriculum materials, teacher scaffolds, tasks, disciplinary content and practices, and interactions among learners (greeno, 2006; shechtman, et al., 2012). productive disciplinary engagement reflects deep-level engagement yielding intellectual progress during authentic disciplinary tasks (engle & conant, 2002) in which students grapple with central domain concepts while participating in the authentic disciplinary practices (duschl, 2008; forman & ford, 2014). engagement, and its interrelated constituent dimensions, are not merely influences on learning, but instead are central to and inseparable from learning (gresalfi, et al., 2009). in this view, invested effort or persistence in the face of challenge, including interpersonal interactions, coordinated activity, and being strategic while making meaningful connections are central to what and how learners come to understand. this earlier work has characterized how shared practices are established in the collective, encompassing teacher-student and whole class negotiation, but with limited focus on the negotiation of norms within the collaborative group. drawing on this multidimensional conceptualization and situated framework of engagement, our goal is to understand collaborative groups’ disciplinary engagement (gde) as being comprised of interrelated, but distinguishable aspects of interaction in group activity. we omitted productive from this descriptor because we wanted to capture the range of variation in quality of disciplinary engagement from none or superficial (i.e., low) to high quality disciplinary engagement that is likely to be productive. we investigate this primary goal within three stem curricula integrating collaboration and disciplinary practices as central design features. we developed and applied an observational rubric to assign quality ratings to explore the engagement profiles of more and less productive groups, and to characterize the synergies among engagement dimensions. building on earlier work (sinha et al, 2015), we delineate five dimensions of group engagement (see table 1). behavioral engagement (be) characterizes the degree to which a group jointly participates and persists on assigned tasks or chooses to go off-task (fredricks, et al., 2004). sustained group participation amongst group members yields potential for building from others’ perspectives, while temporary off-task exchanges can reinvigorate positive interpersonal interactions when returning to task (barron, 2000; langer-osuna, et al., 2020). both socioemotional engagement and collaborative engagement are extensions to individual engagement dimensions, accounting for the interpersonal nature of group engagement (linnenbrink-garcia, et al., 2011). socioemotional engagement (se) characterizes the group’s interpersonal interaction quality and climate, where positive climate involves the negotiation and maintenance of respectful and inclusive interactions, team cohesion, and psychological safety (rogat & linnenbrink-garcia, 2011; rogat & adams-wiggins, 2015). negative socioemotional engagement involves disrespect and competence put-downs, which may result from challenge, conflict, or status differences can derail collaborative engagement (adams-wiggins, 2020; rogat et al 3 | f l r näykki, et al., 2014). research on learning in collaborative groups indicates that positive socioemotional interactions elevate the quality of joint task work (e.g., friendly; supportive; fostering risk-taking) (barron, 2000; kreijns et al., 2002). collaborative engagement (ce) considers groups’ task and conceptual coordination in constructing knowledge, as well as the balance of participation amongst group members in making contributions. high-quality ce undergirds joint knowledge construction accounting for multiple perspectives and promotes the development of a shared problem space (roschelle & teasley, 1995), whereas low-quality ce is characterized by independent task contributions (i.e., low coordination) or a group member’s efforts to control and direct the task (i.e., imbalance). metacognitive engagement (me) describes groups’ use of regulatory strategies, including planning, monitoring, and evaluation (rogat & linnenbrink-garcia, 2011; järvelä, et al., 2016; schoor, et al., 2015). recent findings show that high-quality me is differentiated by effective shared regulation which is goal-focused toward understanding and progress on the task, content, and/or disciplinary practices (rogat & linnenbrink-garcia, 2013; molenaar & chiu, 2014; khosa & volet, 2014), which is supported by, rather than being the sole focus of, regulation of behavior, time, group process, and task completion. finally, disciplinary engagement (de) refers to the nature of the group’s content and disciplinary activity, with high-quality de reflecting connections toward integration of conceptual and disciplinary competencies to solve lesson problems. low-quality de reflects fragmented discussion of content with limited elaboration, and a focus on recall, which may reflect initial understand early in the task or unit or may reflect task or instructional constraints. prior research suggests that high-quality de leads to growth in disciplinary achievement (hmelo-silver et al, 2015). table 1 productive and low-quality indicators collaborative group disciplinary engagement indicators of productive engagement indicators of low-quality engagement behavioral sustained on-task behavior, participation, persistence, and effort, even in the face of challenge primarily off-task behavior, disengagement, and limited focus on shared task work socioemotional respectful, inclusive, cohesive, with a climate characterized by psychological safety disrespectful interactions, exclusion, lacking cohesion and underlying tension reflecting strain to the climate collaborative coordinated and responsive interactions, with balanced participation and diverse perspectives solicited when building and engaging in knowledge co-construction lack of coordination given separate and unrelated contributions, without attempts or willingness to link (i.e., imbalance due to power differential) metacognitive planning, monitoring and evaluation focused on content and/or discipline, and meeting task expectations, aimed toward understanding, improvement, progress, integration, consensus, revisions, and task quality regulation focuses on basic task completion or is ineffective (e.g., unable to cohere around a plan; regulation is not pursued, accepted or ignored; heavy focus on regulating behavior), obstructing task progress disciplinary conceptual and disciplinary connections from prior lessons, across domains, or everyday experiences, with extended elaboration and rationale fragmented and surface-processing of content and practice, with no elaboration or attempts to connect, such as when memorizing or recalling facts or eliciting prior relevant terms when brainstorming 1.1 beyond current methods for studying engagement we operationalize this multidimensional conceptualization in a rubric, enabling the study of distinct engagement dimensions as well as the interrelationships of engagement dimensions that together describe groups’ productive (or unproductive) progress. although extant theory has conceptualized an rogat et al 4 | f l r individual’s engagement as multidimensional (fredricks, et al., 2004), much existing observational research has assessed group engagement narrowly as a single dimension, such as on-task behaviors (hmelo, et al., 1998; lipponen, et al., 2013) or disciplinary engagement (gresalfi & barnes, 2016; koretsky, et al., 2021; mortimer & oliveira de araujo, 2014; sengupta-irving & agarwal, 2017). there has been some qualitative research toward investigating interrelations among two dimensions, including socioemotional and metacognitive engagement, socioemotional and cognitive engagement, and metacognitive and cognitive engagement (isohätälä, et al., 2020; khosa & volet, 2014; rogat & adamswiggins, 2015; rogat & linnenbrink-garcia, 2011). thus, we have limited understanding of the interplay among multiple dimensions. in our own prior research, we posited a multidimensional conceptualization of engagement with some initial exploration of the role of a threshold of engagement practices for on-task behavior and respectful climate to further collaborative engagement toward understanding for differentiating high and low case illustrations (sinha et al, 2015). the observational research that draws on a multidimensional approach tends to focus on single cases in specific curricular and disciplinary contexts (engle & conant, 2002; järvenoja, et al., 2018; sinha et al., 2015). moreover, these approaches rely on intensive and often line-by-line analysis of few cases and thus are not sufficient, alone, for evaluating broader claims concerning how group engagement yields productivity or how engagement fluctuates over time. addressing broader questions would require the use of moderate-to-large samples of many groups and multiple observations per group. we seek to enable analyses of these larger samples by proposing a method that can be applied directly to video or during real-time observations, more efficiently than prior approaches, but with similar ability to richly document the nature and quality of group engagement. empirical study of individual student engagement has assessed engagement as stable, an artifact of self-report surveys capturing of particular moments or retrospective accounts of extended time periods (perry & winne, 2006). further, self-report surveys limit access to information about the nature and quality of interactional processes within the context of particular activities or practices (ryu & lombardi, 2015; vriesema & mccaslin 2020). in alignment with situative views, we prioritize theorizing engagement as involving dynamic change, as groups come to negotiate engagement practices in particular unit phases, on specific tasks and instructional circumstances. capturing engagement as dynamic enables specification of fluctuations and change in group engagement over time, such as during a task within a lesson. moreover, we seek to understand how groups reach productive levels of de, as groups make intellectual progress in their understanding of disciplinary content and/or practices. further, our conceptualization of disciplinary engagement concerns groups’ collective reasoning with domain content knowledge and/or practices authentic to the discipline within specific problem solving, design, or modeling tasks. opportunities for students to engage disciplinarily is important within collaborative exchanges, which offer opportunities for peer-to-peer knowledge codevelopment of concepts, explanation, and critical examination of arguments presented by other group members (roschelle & teasley, 1995). this conceptualization is an extension of prior notions of cognitive engagement as context and discipline-general (fredricks, et al., 2004; pintrich & degroot, 1990) to disciplinary engagement common to stem fields, but instantiated differently within mathematics, science and engineering. toward these ends, we focus on groups’ conceptual work as integrated with discipline-specific reasoning (e.g., modeling, argumentation, design) to generate knowledge needed to solve problems. taken together, we draw on and extend prior theory and methods in several critical ways including 1) developing a multidimensional conceptualization of group engagement, 2) developing a rubric to apply that conceptualization in less time and with less labor intensive analyses, 3) focusing on dynamics of group engagement toward groups’ productive disciplinary engagement, and 4) by comparing patterns of group engagement across disciplinary contexts to explore relationships between patterns of engagement and the contexts in which they emerge. extending these theoretical and methodological precursors, our research team developed and piloted a rubric for describing group gde in three stem curricular contexts. we examine the interrelations among the five engagement dimensions and how, together, these constitute gde. we also illustrate the synergy and mutual influence among dimensions using a case example and a profile rogat et al 5 | f l r analysis with a larger sample drawn from our video corpus. finally, we illustrate how we access discipline-specific patterns of engagement via our multidimensional rubric, drawing on a comparison case from a second curricular context. here, we aim to interrogate, through converging sources of evidence, whether the theoretical framework that we developed and rubric ratings can approximate access to gde gained from intensive qualitative analysis. toward these aims we pose the following research questions: 1. how can a multidimensional conceptualization of engagement account for collaborative group disciplinary activity in the context of particular science, mathematics, and engineering practices? 2. what levels of engagement quality together promote groups’ productive disciplinary engagement? 2. method group disciplinary engagement (gde) is contextualized in collaborative tasks involving modeling, design, and argumentation in middle school math, science, and engineering. we draw on a rich corpus of video data collected from three projects with common features including stem content, disciplinary modeling and argumentation practices, and group work as central to unit goals and what groups came to understand. that is, groups worked together for the majority of lessons in each unit, if not daily, and teachers expected that group work should be the primary mode of activity. within this corpus, the variation in stem domain (science, math and engineering) and the disciplinary practices and curricular features (e.g., technology tools, scaffolds) as contextualized in each curriculum has enriched our theoretical development efforts (koretsky, et al., 2019). the three projects from which collected video was available for rubric development are summarized below. the promoting reasoning and conceptual change in science (praccis) project developed three inquiry units aimed to involve students in scientific practices of evaluating evidence and model fit based on evidence (chinn, et al., 2018). collaborative groups develop, evaluate, and revise explanatory models. for this study, video was selected from the third unit, which focused on evolution and natural selection. video data stemmed from 9 groups in 4 teachers’ classrooms. the school district’s demographics include a student body that is 49% white, 5% black, 34% asian, and 9% hispanic/latino. 5% of students are english language learners. moreover, 14% of students are eligible for free and reduced priced lunch. the simcalc engagement project (sep) leveraged student’s use of multiple dynamic representations of mathematical relationships (i.e., modeling) to support their understanding of rate and proportion in a technology-based instructional unit (roschelle et al., 2010). study videos were drawn from 3 focal groups in 2 teachers’ classrooms. one teacher taught in a school where the student body is 56% hispanic/latino, 12 % asian, 10% white, 9% filipino, 4% black, and 1% native american. in this school, 53% of students are eligible for free and reduced priced lunch and 22% are classified as english language learners. the second teacher taught in a school where the student body is 63% white, 14% asian, 13 % multi-racial, 9% hispanic and 1% filipino. in this school, 1% of students are eligible for free and reduced priced lunch and 1% are classified as english language learners. the human-centered robotics project aimed to inspire youth interest in stem topics to develop robotic technologies in response to people’s needs (gomoll et al., 2018). video data were used for two groups during two curriculum implementations of a stem elective course in one classroom. study videos were drawn from 2 focal groups in one teacher’s classrooms in a rural midwestern us school district that was largely white (~90%), with 39% eligible for free and reduced lunch. the sample consisted of 36 five-minute segments drawn from the larger video corpus of 77 lessons. this included 15 groups, each observed once or twice. in these videos, students were primarily rogat et al 6 | f l r assigned to work as groups of three or four, but occasionally were assigned to work in pairs and then self-organized into groups. video segments were balanced across the three units. we aimed to observe some variation in the sample by including videos from across unit phase and disciplinary practices integrated in tasks. moreover, given our interest in exploring the dynamics of group engagement, we intentionally included some additional segments to observe groups throughout a lesson (i.e., consecutive time segments or segments from later in a lesson). 2.1 measuring collaborative disciplinary engagement we developed a rubric to employ when observing groups collaborating during joint activity with the aim of characterizing their co-negotiated engagement practices. the rubric encompasses five engagement dimensions using primarily 3-point rating scales, with de specified using a 4-point scale. we recognized that the de dimension benefited from an extended scale based on our observations during rubric development, both in the case of limited discourse and non-verbal interactions such as during independent activity in which productive disciplinary engagement could still be happening, and in the case of the high-end of the continuum where elaboration and justification as well as initial and brief conceptual connections are made. beyond the engagement dimensions, the rubric also captures group structure (i.e., whether groups opt to or are assigned to work in pairs, as a full group, and even individually) to characterize fluctuations in how groups are organized over time. initial rubric drafts were informed by a review of extant research on engagement, empirical studies, and self-report items. these initial drafts were also informed by conducting joint analyses of group interactions (n = 4 groups), drawing on the range of expertise of the various project team members when describing observable group engagement along multiple dimensions from across curriculum contexts (jordan & henderson, 1995). iterative revision of this initial framework was informed by a pilot study, expert feedback, and questions during rater training. raters observe group interactions including members’ behaviors, discourse content, and nonverbal behaviors (e.g., gaze, gesture, leaning in toward joint task, spatial closeness) (chi & wylie, 2014). using this rubric, we assign group engagement ratings based the predominant character of group interactions for the majority of the selected segment. quality indicators differentiated behavioral engagement as primarily on-task or off-task joint participation, socioemotional engagement as the respectful, inclusive and cohesive group climate and alleviation of tension and frustration stemming from different perspectives or interpersonal dynamics, collaborative engagement as responsive and coordinated knowledge building, metacognitive engagement as regulation focused on the progress and understanding of content and disciplinary practice, and disciplinary engagement as the integrated conceptual and disciplinary contributions with rationale. across dimensions, we assume that high-level ratings facilitate the likelihood of attaining productive gde, with some potential exceptions of low or moderate ratings also supporting productive gde (e.g., temporary off-task joking (low be) may benefit group cohesion (high se)). rubrics for each engagement dimension can be found in the appendix. 2.2 study procedures in applying the rubric, raters assigned quality ratings for each of the five engagement dimensions, for five-minute time segments of videotaped group work. this time segment choice was based in our own prior research and prior observational studies of engagement (lee & brophy, 1996; sinha, et al., 2015). the time segment also afforded more than a single turn in conversation, with sufficient time for group members to respond and negotiate a task direction and/or understanding. raters independently viewed video segments twice, first to familiarize themselves with the activity and a second time to assign ratings. ratings were applied by three teams, made up of 2 to 3 raters, with each team specializing in one curricular context. the study proceeded in four rating cycles, followed by calculating interrater agreement. in cycles 3 and 4, rater teams had krippendorff’s alpha interrater agreement coefficients of .53 for be, .64 for ce, .59 for se, .30 for me, and .48 for de. lebreton and senter (2008, p. 836) suggest that for chance-corrected interrater agreement coefficients, values of 0 to .30 should be interpreted as “lack of agreement,” .31 to .50 as “weak agreement,” and .51 to .70 as “moderate agreement.” after recording individual ratings and calculating initial interrater agreement, rogat et al 7 | f l r the pair worked together with the project team’s master raters to resolve areas of disagreement and come to consensus; these conversations were structured to clarify discrepancies and produce coding clarifications and exemplars to inform subsequent coding. it is these consensus ratings that we employ in the presented analyses. 2.3 analysis to gain an understanding of relationships between de and each of the other dimensions, we calculated correlations and conducted a profile analysis. given that the five dimensions are theorized to jointly constitute productive group engagement, we would anticipate moderate positive correlations among the dimensions. profile analysis is a variant of multivariate analysis of variance (manova) that allows testing of hypotheses about patterns of means across groups. specifically, if omnibus f tests indicated significant differences in the pattern of engagement dimension means across the high and low de subsamples, we would proceed to univariate analyses of dimension mean differences. we conceptualized the four de rating scale levels as possibly representing two categorically different types of de, relatively low (scale categories 1 and 2) and relatively high (scale categories 3 and 4), rather than as linearly increasing. thus, we identified two subsamples based on the observed level of de and compared profiles of means on the other engagement dimensions across those two subsamples. because fewer than 40% of the student groups were observed more than once over time, we planned to interpret the results only as indicating any cross-sectional relationships that exist among the engagement dimensions. five student group observation cases, a small fraction of the sample, were missing ratings on metacognitive engagement because none was observed during the sampled time period. rather than excluding these cases from the analysis, which would be expected to cause estimation bias unless the missingness was completely at random (e.g., graham, 2009), we implemented multiple imputation for 10 replications. stata was used to compute estimated coefficients and standard errors for each replicated dataset, and to combine the results using rubin’s (1987) rules. a case illustration was purposefully selected from those groups (n = 5) showcasing high-level de of a 4 rating, with assigned dimensional ratings similar to remaining groups with high de ratings. we intentionally drew from two curricular contexts that would require observers to specifically evaluate the disciplinarity of de. we reviewed the video segments and described with rich narrative the central dimensions specified in the rubric. analyses focused on how engagement dimensions worked in synergy to produce high-level disciplinary engagement. 3. results means, standard deviations and correlations among the engagement dimension ratings are presented in table 2. se had the highest mean rating, while ce had the lowest mean ratings. the moderate positive correlations, most of which are significantly greater than zero, suggest positive correspondence among all five dimensions. the high correlation between me and de occurs in all likelihood because the quality of me is important for initiating and/or supporting de. the low correlation between be and se likely indicates that positive socioemotional interactions (se 3 rating) may be evident during off-task exchanges (be 1 or 2 ratings). table 2 correlations among engagement dimension ratings (n = 36) mean sd be ce se me de behavioral engagement 2.44 0.73 1 collaborative engagement 2.14 0.64 .33 1 socioemotional engagement 2.69 0.62 .18 .27 1 rogat et al 8 | f l r metacognitive engagement 2.29 0.74 .74*** .57** .58* 1 disciplinary engagement 2.61† 0.77 .60** .58*** .29 .85*** 1 note. † mean = 2.07, sd = .051 for disciplinary engagement ratings transformed to 3-point scale similar to that employed for the remaining dimensions. * p < .05; ** p < .01; *** p < .001. 3.1 case illustration we conducted a qualitative case analysis to richly characterize a group’s efforts in making intellectual progress and to explore the potential for synergistic interrelations among engagement dimensions in explaining high-quality de. this specific case stems from the simcalc mathematics curricular context. the triad is made up of three girls (pseudonyms: abby, beth, carly) (see figure 1a). as shown in figure 1b, the group is tasked with interpreting a line graph depicting motion of two fictional vehicles over time (i.e., their speed). the assigned ratings indicate some intermittent off-task activity (be=2), with remaining ratings at a high-quality level (se=3, ce=3, me=3, de=4). they were working on recording answers in their individual workbooks, with two open laptops. abby was recently absent and was working to catch up in her workbook. the lesson took place in november and was midway through a 2-week unit during which this group had been working together on a daily basis. students’ comments indicated that they already knew each other before the unit. just prior to the focal time segment, abby raised a question, based on her metacognitive monitoring of the assigned question concerning how to read the graph in question 2a (me)1 (figure 1b). beth and abby’s beginning interactions in making sense of the graph illustrate both girls voicing a common misunderstanding that the line graph represents the physical characteristics of one of the vehicles’ routes rather than its speed and relying on the simple recall of the slope formula. abby’s questions about the meaning of the slope formula (me) as related to the graph, elicited group disciplinary engagement (de) as they mutually monitored their understanding and grappled with interpreting the graph: abby: i don’t get how it [the graph] can describe the motion. like, what does that mean? they’re just lines. it doesn’t say describe the graph, it says to describe the bus and the van. beth: the motion. but it is still like in the graph abby: yeah, but the bus didn’t drive tilted. beth: the motion of the bus. the van had a straighter route, and the bus had a more curved route. abby: i guess that made sense. then we have to write the speed again. beth: and it would be like distance over….distance over time. abby: for one second, does that mean? like what distance and what time? the full group briefly disengaged from the task as abby complimented carly’s hat (be). this elicited carly’s sharing a peer’s previous teasing about it. beth and abby showed their support by giggling along, actively listening, and agreeing that it was silly to dislike her unique hat (i.e., “some people just don’t know your style.”) (se). this temporary off-task exchange (be) was positive in socioemotional interactions and fostered team cohesion (se) amongst the group, which the group subsequently leveraged in the collaborative engagement (ce) which followed. 1 throughout the case we identify evidence of specific engagement dimensions by noting their abbreviations in parenthesis. rogat et al 9 | f l r (a) (b) figure 1. case illustration collaborative group (a) and image of student worksheet (b). abby returned to the task (be) and persisted in raising a question regarding her uncertainty in interpreting the line graph (me). “beth, so i don’t get how they can just be…how can they go the same amount of miles and still go at the same speed? like, it just doesn’t make sense.” beth was responsive to abby’s question, turning toward her and leaning in to view the graph abby was drawing (se), precipitating this collaborative and disciplinary engagement. beth leaned in and explained, “because this one [the bus] is only a little more quicker.” abby further elaborated and built from this explanation (ce) stating, “oh i get it because the bus starts going off faster…so it ...is like it has more slack, like kind of has more time and then they arrive at the same time.” and “because the bus stopped. so they are both going the same speed.” abby continued to negotiate their working understanding of the different speeds of the van and bus while gesturing to her graph (de), “actually, the bus still is going at a faster speed. its speed would be faster than the van. just because it stopped doesn’t mean the speed changed, right?” (me and de). abby further questioned whether their previously employed effective math strategy for using the arrival or end time of the graph to calculate the slope would work for this new representation with vehicles at different speeds. “but, then so i am confused on how i would represent that as two things, fractions, because at the end they are the same” and, “i’m not going to do the end distance and the end time; i am going to do something a little earlier on” (me and de). throughout, beth and carly showed their coordination by voicing agreement, responding “yeah,” ensuring they were looking at the same question in the packet and by prompting abby’s meaning making (i.e., carly “so (inaudible) what do you think of the motion?”) (ce). carly further supported the group’s knowledge construction as she moved her physical position to be proximal to abby and facilitated the interpretation of the graph and the task, further suggesting that the former off-task but positive socioemotional interaction (se) maintained her involvement in the conversation. this case illustration showcases how the five-engagement dimensions interrelated across a fiveminute time segment in ways that afforded intellectual progress at the integration of mathematical practice (generating a model, coordinating the use of a graphic representation) and content (conceptual understanding of slope) (de). the group engagement practices reflect a willingness to persist in the face of uncertainty (be) and metacognitively monitor for understanding concerning how to interpret the graphic representation of speed, the meaning behind the slope formula and different vehicle speeds, as well as whether their former mathematical strategy of using graphic endpoints was still relevant (me), initiated by abby but sustained by the joint group’s high level of engagement. carly and beth’s highquality collaborative engagement was responsive to the metacognitive monitoring and yielded joint knowledge co-construction (ce). the group was mutually respectful and cohesive throughout the exchange, including when off-task behavior was similarly leveraged (se). ultimately, it is our multidimensional conceptualization of the group’s engagement that enabled the examination of synergy among dimensions as facilitating the joint accomplishment of de. rogat et al 10 | f l r 3.2 exploring the disciplinarity of engagement common curriculum features, including incorporation of the authentic disciplinary practices of modeling, argumentation, and design within the curriculum corpus, facilitated questions about how de could be evaluated in ways that were common in three domains, while remaining sensitive to detecting discipline-specificity. during rubric development, collaborative analysis of videos from across curriculum contexts enabled the team to generate descriptions of de that are applicable across disciplinary tasks, involving various disciplinary practices (see appendix). indicators that are specific to each context ground the de rubric in examples that support raters; this is especially important for specifying disciplinary variation. this approach, a broader definition of de with context-specific indicators, allows for comparative analyses across contexts, including patterns of and interrelations between various dimensions of the rubric. for example, observing in one curricular context where frequent student interaction was aligned with disciplinary norms, we noted long periods of independent work eroded team cohesion (se) and ultimately constrained conceptual progress. however, in one segment from the robotics dataset, students spent far more time in independent on-task work (be=3), captured in group structure ratings of ‘individual work,’ with mid or low ratings on all but socioemotional engagement (se=3, ce=1, me=0, de=2), as students each attended to building different components of the robot. in this context and similar disciplinary spaces, longer periods of independent work were common and group cohesion was maintained, marked by brief check-ins by group members (e.g., ‘does this look right to you?’) and intermittent gesture-based collaboration accomplished through physical indicators (e.g., leaning over to examine and/or gazing at each other’s progress without comment). in later segments, the teams’ disciplinary engagement was rated higher when their conversation turned to providing rationales for the ways their robot construction choices did and did not meet stakeholder stated preferences. we intended our final rubric to measure de in a unified manner across contexts. 3.3 profile analysis to examine patterns in the relationships among disciplinary engagement and the other dimensions across the cross-sectional video sample, a profile analysis was conducted. taking group disciplinary engagement ratings of 3 or 4 to indicate high (n=20 cases), and ratings of 1 or 2 to indicate low de (n=16 cases) (see section 2.2 and appendix), we prepared a plot of the mean rating profiles (figure 2). visual inspection indicated the high de group observations tended to have higher ratings across all four co-occurring engagement dimensions. figure 2. profiles of (mean) group ratings on other engagement dimensions when disciplinary engagement (de) is high and low a preliminary multivariate analysis of variance (manova) was used to test the hypothesis of no mean differences between the low and high disciplinary engagement observations on any of the other engagement ratings (in which case proceeding with profile analysis would be unwarranted). the null hypothesis was rejected for the wilks’ lambda omnibus test statistic [f(4, 31) = 5.85, p = 0.001], so we rogat et al 11 | f l r continued with the profile analysis. a test of parallelism indicated that the profiles of mean ratings for the low and high de observations had marginally significant differences in overall shape [f(3, 32) = 2.82, p = 0.063]. a test of “level” or group differences confirmed that the profiles were not coincident; that is, the group ratings on each dimension were not identical [f(1, 34) = 14.81, p < 0.001]. given the significant group difference in overall rating outcomes found in the manova, to identify specific engagement dimension(s) that were the source, we estimated univariate anova models with each engagement dimension as a single outcome, and the high/low de indicator again as a predictor. these “stepdown” analyses suggested significant differences between the high and low de observations for behavioral [t = 3.12, p < 0.01], collaborative [t = 2.35, p < 0.05], and metacognitive [t = 4.36, p < 0.001] engagement, but not socioemotional engagement. the profiles suggest that the other engagement dimensions significantly vary with the quality of de and that these dimensions may interrelate in fostering de. 4. discussion in this research, we used multiple methods to investigate how group productive disciplinary engagement is constituted, using a multidimensional framework and our initial rubrics developed to embody this framework. study findings suggest that the five dimensions of our disciplinary engagement framework and rubrics are positively interrelated, illustrating that interrelationships among dimensions mutually support the high-quality of disciplinary engagement observed among groups during joint activity. we used a case illustration to richly characterize the dynamic and synergistic nature of these dimensions for explaining de, as well as the import of being sensitive to disciplinary specificity of de. ratings assigned to the time segments of group activity provide data complementing the nuance of indepth cases, by allowing for the identification of common patterns across a wide range of groups in varying curricular contexts, as initially demonstrated by the profile analysis. the rubric corresponded to group engagement features that could have been identified via intensive video analysis, while affording the evaluation of broader claims of patterns of group engagement with larger datasets. although the profiles of high and low de group observations suggest that the other engagement dimensions significantly vary with the quality of de, this was not the case for socioemotional engagement (se); results suggested that the high and low de observations were both characterized by high-quality socioemotional interactions. one explanation is that the high se (i.e., 3) rating included a range of quality indicators, primarily reflecting polite and collegial interpersonal interactions, climate norms aligned with working well together, expectations as part of classroom work but also interactions that added to or made efforts to maintain a positive climate (rogat & linnenbrink-garcia, 2011; summers, et al., 2005). the highest form of se was observed during off-task group exchanges, characterized by positive and friendly interactions that could be carried over back to on-task interactions (langer-osuna, et al., 2020). it may be that the high-quality se which differentiates high de group engagement reflects the active negotiation and mutual accountability of a climate facilitative of risk taking and inclusion of diverse ideas, at the upper endpoints for this dimension. these interpersonal dynamics are likely a less regular occurrence, contextualized in particular situations (e.g., newly constituted group, following disagreement or provoked tension), requiring examination of longer time periods and the exploration of dynamics via analyses that were not yet conducted in this initial development work. our work aligns with a situative perspective on learning by investigating collaborative engagement as shared, with the group as the unit of analysis (e.g., barron, 2000; engle & conant, 2002; gresalfi & barnes, 2016). this work contributes a method to the study of group engagement that leverages the benefits of observational methods (vs. self-report) and multi-dimensional conceptualizations of engagement in a framework and rubric. these tools enable descriptions and analyses of group engagement as situated and anchored in disciplinary content and practices, and as trajectories comprised of dynamically interrelated aspects of group activity. one important implication for research is that using this tool, researchers can address critical questions about collaborative learning rogat et al 12 | f l r and engagement that have been inaccessible because of limitations of prior methods, including examination of the complex dynamics of separable but interdependent dimensions of engagement and, in particular, the potential variation in quality within and across dimensions. this work extends collaborative group research that has examined single engagement dimensions, which may lend to a conceptualization of these dimensions as separable and independent rather than as interrelated dimensions that present a more enriched characterization. moreover, these phenomena can be examined across disciplinary learning contexts, across specific tasks, and as a function of time, while remaining sensitive to disciplinary specificity when investigating productive (and non-productive) engagement. our rubric, by being broadly applicable but grounded by discipline-specific indicators, supports theory development about groups’ gde, including implications for instructional design, practice, and collaborative learning. in other research, gomoll et al (2020) used these rubrics as a tool for professional development that the teacher could use as a lens for video analysis and subsequent facilitation of collaborative groups. although this was a small-scale study, future research can build on these implications for teacher professional development. limitations of the current study include the examination of interrelations of dimensions of engagement in single time segments, precluding the examination of engagement as dynamic and fluctuations across multiple segments that make up a group task, which is the ultimate goal of this research program and is part of our current research activities. moreover, the video sample included repeat observations from the same day’s lesson, given our long-term interest in exploring lesson dynamics and short-term interest in testing alternative time intervals for recording ratings, but those were not modeled in the profile analysis given the modest sample size. additionally, this study examines a case with high-quality engagement ratings for a majority of the rubric dimensions, but we do not assume that this pattern exclusively fosters gde (e.g., sometime off-task behavior [lower be] enables students to joke and bond socially [increasing se], which may precede gde); future research will aim to identify and illustrate other such patterns. the rubric presented here is an initial exploration for operationalizing collaborative group disciplinary engagement as five dimensions using quality ratings. however, we faced challenges in obtaining inter-rater agreement prior to consensus meetings, with reliability indices indicating room for improvement particularly for metacognitive engagement (me). we understand these challenges as attributable to (1) rater process, (2) the complexities of group data, and (3) unique challenges presented by the me dimension. first, these data were from secondary sources collected from past projects and although raters were trained to become familiar with the curriculum materials, we anticipate that most users of this rubric would be studying curricular contexts with which they are highly familiar. second, we asked raters to examine interactions among members of the group, which is more complex than observing individuals’ engagement (e.g., lee & brophy, 1996), capturing the nature of collective interaction as students present a contribution and other groupmates’ responses by accepting, ignoring or rejecting with or without rationale (barron, 2000). furthermore, raters observe engagement across 5minutes, with fluctuation in quality typical within that timeframe, and being tasked with selecting the rating that best represents the majority of that time. in our current research with a revised rubric, we have modified the time segment to 2.5 minutes to address the noted disagreement provoked by raters’ varying strategies for synthesizing across data. finally, me proved uniquely challenging as raters were asked to simultaneously evaluate several relevant elements in their rating, including the metacognitive target, the duration, and its quality. in other research employing qualitative analyses, these elements are considered in distinct analytic phases, with an initial step identifying the presence of me in group interaction, and then classifying its target (e.g., rogat & linnenbrink-garcia, 2011; volet, summers & thurman, 2009). building on this exploratory study, we have worked to address challenges specific to me by removing the evaluation of duration and developing more specific indicators for each level of quality. nonetheless, because of the initial low reliability, all ratings used here were consensus ratings of two or more raters. future research might also employ mixed methods for coupling the analysis of dynamic group patterns of engagement, such as through latent profile transition analysis on the five engagement dimensions across time (e.g., nylund-gibson, et al., 2014). this analysis would include an in-depth qualitative analysis of groups exemplifying these engagement trajectories. this convergence of methods rogat et al 13 | f l r would offer an enriched understanding of how specific group processes unfold as trajectories. another recommendation would be to investigate how the framework and rubric may prove valuable in classroom contexts by supporting teachers through professional development about the additional support and resources collaborative groups need, targeted toward specific engagement dimensions. similarly, studies could examine the benefit of proximal feedback provided to teachers to identify groups at varying unit phases in need of monitoring and/or support relevant to specific group processes hindering their collective efforts at making intellectual progress. although there is still work to done, this research demonstrates the importance of considering group disciplinary engagement as a complex and multidimensional phenomenon. keypoints we propose that engagement is a group-level construct that builds from and integrates theory on individual engagement and productive disciplinary engagement. we advance innovative methods for evaluating collaborative group engagement during disciplinary activity as shared, multidimensional and dynamic. we developed and piloted an observational rubric encompassing five engagement dimensions, each with quality ratings. mixed-methods analyses included correlations, profile analyses, and case illustrations. results illustrate the synergistic interrelations among engagement dimensions that together constitute group disciplinary engagement. acknowledgments this research was funded by the united states’ national science foundation (nsf) grants # 1661266 and 1661234. any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the national science foundation. we thank clark chinn and ravit duncan for sharing the praccis video, thesimcalc research group at sri, especially nikki shechtman and patrik lundh for simcalc data and thought partnership. we thank the graduate and undergraduate researchers who were part of our project term who employed the described rubric and provided invaluable feedback, specifically brianna cermak, alex lindsay, mehdi ghahremani, and skye huffman. we also thank megan humberg for helpful comments on an earlier draft. rogat et al 14 | f l r references adams-wiggins, k. r. (2020). whose meanings belong?: marginality and the role of microexclusions in middle school inquiry science. learning, culture and social interaction, 24, 100353. https://doi.org/10.1016/j.lcsi.2019.100353 barron, b. (2000). achieving coordination in collaborative problem-solving groups. journal of the learning sciences, 9, 403-436. https://doi.org/10.1207/s15327809jls0904_2 chi, m.t.h. & wylie, r. (2014). the icap framework: linking cognitive engagement to active learning outcomes. educational psychologist, 49, 219-243. https://doi.org/10.1080/00461520.2014.965823 chinn, c. a., duncan, r. g., & rinehart, r. (2018). epistemic design: design to promote transferable epistemic growth in the praccis project. in e. manalo, y. uesaka, & c. a. chinn (eds.), promoting spontaneous use of learning and reasoning strategies. pp. 242-259. new york: routledge. danish, j. a., & gresalfi, m. (2018). cognitive and sociocultural perspective on learning: tensions and synergy in the cognitive and sociocultural perspective on learning: tensions and synergy in the learning sciences. in f. fischer, c. e. hmelo-silver, s. r. goldman, & p. reimann (eds.), international handbook of the learning sciences (pp. 33-43). new york: routledge. duschl, r. (2008). science education in three-part harmony: balancing conceptual, epistemic, and social learning goals. review of research in education, 32, 268-291. https://doi.org/10.3102/0091732x07309371 engle, r. a., & conant, f. c. (2002). guiding principles for fostering productive disciplinary engagement: explaining an emerging argument in a community of learners classroom. cognition and instruction, 20, 399-483. https://doi.org/10.1207/s1532690xci2004_1 engle, r. a., langer-osuna, j. m., & mckinney de royston, m. (2014). toward a model of influence in persuasive discussions: negotiating quality, authority, privilege, and access within a studentled argument. journal of the learning sciences, 23(2), 245-268. https://doi.org/10.1080/10508406.2014.883979 forsthuber, b., motiejunaite, a., & de almeida coutinho, a. s. (2011). science education in europe: national policies, practices and research. education, audiovisual and culture executive agency, european commission. fredricks, j., blumenfeld, p., & paris, p. (2004). a school engagement potential of the concept and state of the evidence. review of educational research, 74, 59-109. https://doi.org/10.3102/00346543074001059 forman, e. a., & ford, m. j. (2014). authority and accountability in light of disciplinary practices in science. international journal of educational research, 64, 199-210. https://doi.org/10.1016/j.ijer.2013.07.009 fulmer, s.m., & frijters, j.c. (2009). a review of self-report and alternative approaches in the measurement of student motivation. educational psychology review, 21, 219-246. https://doi.org/10.1007/s10648-009-9107-x gomoll, a. s., hillenburg, r., & hmelo-silver, c. e. (2020). “i have never had a pbl like this before”: on viewing, reviewing, and co-design. interdisciplinary journal of problem-based learning. 14(1). https://doi.org/10.14434/ijpbl.v14i1.28802. gomoll, a. s., šabanović, s., tolar, e., hmelo-silver, c. e., francisco, m., lawlor, o. (2018). between the social and the technical: negotiation of human-centered robotics design in a middle school classroom. international journal of social robotics, 10, 309-324. https://doi.org/10.1007/s12369-017-0454-3 graham, j. w. (2009). missing data analysis: making it work in the real world. annual review of psychology, 60, 549–576. 10.1146/annurev.psych.58.110405.085530 rogat et al 15 | f l r greeno, j. g. (2006). learning in activity. in r. k. sawyer (ed.), the cambridge handbook of the learning sciences (pp. 79–96). cambridge: cambridge university press gresalfi, m. s., & barnes, j. (2016). designing feedback in an immersive videogame: supporting student mathematical engagement. educational technology research and development, 64(1), 65-86. https://doi.org/10.1007/s11423-015-9411-8 gresalfi, m., martin,t., hand, v. & greeno, j. (2009). constructing competence: an analysis of student participation in the activity systems of mathematics classrooms. educational studies in mathematics, 70, 49-70. https://doi.org/10.1007/s10649-008-9141-5 hand, v., & gresalfi, m. (2015). the joint accomplishment of identity. educational psychologist, 50, 190-203. http://dx.doi.org/10.1080/00461520.2015.1075401 hickey, d. t. (2003). engaged participation versus marginal nonparticipation: a stridently sociocultural approach to achievement motivation. elementary school journal, 103(4), 401-429. https://doi.org/10.1086/499733 hmelo, c. e., guzdial, m., & turns, j. (1998). computer-support for collaborative learning: learning to support student engagement. journal of interactive learning research, 9, 107-130. hmelo-silver, c. e., eberbach, c., jordan, j., sinha, s., rogat, t. k. (2015, august). trajectories for engaged learning about complex systems. presented at european association for research on learning and instruction biennial conference. limassol, cyprus. isohätälä, j., naykki & jarvela (2020). convergences of joint, positive interactions and regulation in collaborative learning. small group research, 51, 229-264. https://doi.org/10.1177/1046496419867760 järvelä, s., järvenoja, h., malmberg, j., isohätälä, j., & sobocinski, m. (2016). how do types of interaction and phases of self-regulated learning set a stage for collaborative engagement?. learning and instruction, 43, 39-51. http://dx.doi.org/10.1016/j.learninstruc.2016.01.005 järvenoja, h., järvelä, s., törmänen, t., näykki, p., malmberg, j., kurki, k., mykkänen, a. & isohätälä, j. (2018). capturing motivation and emotion regulation during a learning process. frontline learning research, 6, 85-104. https://doi.org/10.14786/flr.v6i3.369 khosa, d. k., & volet, s. e. (2014). productive group engagement in cognitive activity and metacognitive regulation during collaborative learning: can it explain differences in students’ conceptual understanding?. metacognition and learning, 1-21. https://doi.org/10.1007/s11409014-9117-z koretsky, m. d., vauras, m., jones, c., iiskala, t., & volet, s. (2021). productive disciplinary engagement in high-and low-outcome student groups: observations from three collaborative science learning contexts. research in science education, 51, 159-182. https://doi.org/10.1007/s11165-019-9838-8 kreijins, k., kirschner, p.a. & jochems, w. (2002). the sociability of computer-supported collaborative learning environments. journal of education technology and society 5, 8–22. http://www.jstor.org/stable/jeductechsoci.5.1.8 langer-osuna, j. m., gargroetzi, e., munson, j., & chavez, r. (2020). exploring the role of off-task activity on students’ collaborative dynamics. journal of educational psychology, 112(3), 514-532. https://doi.org/10.1037/edu0000464 lebreton, j. m., & senter, j. l. (2008). answers to 20 questions about interrater reliability and interrater agreement. organizational research methods, 11(4), 815–852. https://doi.org/10.1177/1094428106296642 lee, o., & brophy, j. (1996). motivational patterns observed in sixth‐grade science classrooms. journal of research in science teaching, 33, 303-318. https://doi.org/10.1002/ rogat et al 16 | f l r linnenbrink-garcia, l., rogat, t.k. & koskey, k.l. (2011). affect and engagement during small group instruction. contemporary educational psychology, 36, 13-24. https://doi.org/10.1016/j.cedpsych.2010.09.001 lipponen, l., rahikainen, m., lallimo, j., & hakkarainen, k. (2003). patterns of participation and discourse in elementary students’ computer-supported collaborative learning. learning and instruction, 13, 487-509. https://doi.org/10.1016/s0959-4752(02)00042-7 molenaar, i., & chiu, m. m. (2014). dissecting sequences of regulation and cognition: statistical discourse analysis of primary school children’s collaborative learning. metacognition and learning, 1-24. https://doi.org/10.1007/s11409-013-9105-8 mortimer, e. & de araújo, a. (2014). using productive disciplinary engagement and epistemic practices to evaluate a traditional brazilian high school chemistry classroom. international journal of educational research, 64, 156-169. https://doi.org/10.1016/j.ijer.2013.07.004 näykki, p. järvelä, s. kirschner, p.a. & järvenoja, h. (2014). socio-emotional conflict in collaborative learning – a process-oriented case study in a higher education context. international journal of educational research, 68, 1-14. https://doi.org/10.1016/j.ijer.2014.07.001 ngss lead states. (2013). next generation science standards: for states, by states. washington, dc: the national academies press. nylund-gibson, k., grimm, r., quirk, m., & furlong, m. (2014). a latent transition mixture model using the three-step specification. structural equation modeling, 21, 439-454. https://doi.org/10.1080/10705511.2014.915375 perry, n. e., & winne, p. h. (2006). learning from learning kits: gstudy traces of students’ selfregulated engagements with computerized content. educational psychology review, 18(3), 211228. https://doi.org/10.1007/s10648-006-9014-3 pintrich, p.r. & degroot, e.v. (1990).motivational and self-regulated learning components of classroom academic performance. journal of educational psychology, 82, 33-40. https://doi.org/10.1037/0022-0663.82.1.33 rogat, t.k. & adams-wiggins, k.r. (2015). facilitative versus directive other-regulation in collaborative groups: implications for socioemotional interactions. computers in human behavior, 52, 589-600. https://doi.org/10.1016/j.chb.2015.01.026 rogat, t.k. & linnenbrink-garcia, l. (2011). socially shared regulation in collaborative groups: an analysis of the interplay between quality of social regulation and group processes. cognition and instruction, 29, 375-415. https://doi.org/10.1080/07370008.2011.607930 rogat, t.k. & linnenbrink-garcia, l. (2013). understanding the quality variation of socially shared regulation: a focus on methodology. in m. vauras & s. volet (eds.), interpersonal regulation of learning and motivation: methodological advances (pp. 102-125). london: routledge. roschelle, j., shechtman, n., tatar, d., hegedus, s., hopkins, b., empson, s.,... & gallagher, l. p. (2010). integration of technology, curriculum, and professional development for advancing middle school mathematics: three large-scale studies. american educational research journal, 47, 833-878. https://doi.org/10.3102/0002831210367426 roschelle, j., & teasley, s. (1995). the construction of shared knowledge in collaborative problem solving. in c. o’malley (ed.), computer-supported collaborative learning (pp. 69-197). berlin, germany: springer. rubin, d. b. (1987). multiple imputation for nonresponse in surveys. new york: wiley. ryu, s., & lombardi, d. (2015). coding classroom interactions for collective and individual engagement. educational psychologist, 50, 70-83. https://doi.org/ 10.1080/00461520.2014.1001891 rogat et al 17 | f l r sandoval, w. a. (2014). conjecture mapping: an approach to systematic educational design research. journal of the learning sciences, 23, 18-36. https://doi.org/10.1080/10508406.2013.778204 schoor, c., narciss, s., & körndle, h. (2015). regulation during cooperative and collaborative learning: a theory-based review of terms and concepts. educational psychologist, 50, 97119. https://doi.org/10.1080/00461520.2015.1038540 sengupta-irving, t., & agarwal, p. (2017). conceptualizing perseverance in problem solving as collective enterprise. mathematical thinking and learning, 19(2), 115-138. https://doi.org/10.1080/10986065.2017.1295417 shechtman, n., cheng, b. h., lundh, p., & trinidad, g. (2012). unpacking the black box of engagement: cognitive, behavioral, and affective engagement in learning mathematics. in van aalst, j., thompson, k., jacobson, m. j., & reimann, p. (eds.) the future of learning: proceedings of the 10th international conference of the learning sciences (icls 2012) – volume 2, short papers, symposia, and abstracts. international society of the learning sciences: sydney, nsw, australia, pp. 53-56. sinha, s., rogat, t., adams-wiggins, k. r., & hmelo-silver, c. e. (2015). collaborative group engagement in a computer-supported inquiry learning environment. international journal of computer-supported collaborative learning, 10, 273-307. https://doi.org/10.1007/s11412015-9218-y summers, j. j., gorin, j. s., beretvas, s. n., & svinicki, m. d. (2005). evaluating collaborative learning and community. the journal of experimental education, 73(3), 165-188. https://doi.org/10.3200/jexe.73.3.165-188 volet, s., summers, m., & thurman, j. (2009). high-level co-regulation in collaborative learning: how does it emerge and how is it sustained? learning and instruction, 19, 128-143. https://doi.org/10.1016/j.learninstruc.2008.03.001 vriesema, c.c, & mccaslin, m. (2020). experience and meaning in small-group contexts: fusing observational and self-report data to capture self and other dynamics. frontline learning research, 8,136-139. https://doi.org/10.14786/flr.v8i3.493 rogat et al 18 | f l r appendix collaborative disciplinary engagement observational rubric behavioral engagement: group norm can be characterized by on-task engagement, persistence, and effort investment, even in the face of challenge 1: low 2: moderate 3: high group characterized by off-task behavior, with limited on-task activity brief and intermittent on-task activity joking in off-task interactions groupmates may attempt to distract on-task activity group characterized by predominantly on-task activity for a majority of the time, but intermittent off-task activity. group characterized by sustained on-task activity, with brief intermittent off-task activity socioemotional engagement: socioemotional climate is respectful, cohesive, and characterized by psychological safety 1: low 2: moderate 3: high group interactions characterized by negative climate reflective of the following qualities: • disrespectful (put downs, harsh criticism of ideas; grabbing, shoving, pushing you out of the way, shouting) • interactions showcase low cohesion/sense of team when tension and frustration are expressed, it is responded to with disrespect, resistance to difference in perspectives; tension may be sustained. when the group makes mistakes, seek blame of groupmates; criticism. group interactions characterized by mixed climate (indicators of both negative and positive climate are present). tension brings strain to the group climate (although not overtly disrespectful or safe) laughter reflects mild tension group interactions characterized by a positive climate reflective of working well together or promoting high-quality positive climate: • respectful, polite, collegial • encouraging of groupmates/team • climate is comfortable in terms of allowing for risktaking, mistakes as well. • cohesive/team; warmth and caring about one another • good-natured and friendly during off-task interactions (e.g., friendly joking) when tension and frustration are expressed, it is alleviated, responded to with safe climate and respect rogat et al 19 | f l r when the group makes mistakes – respect and sense of team is fostered. collaborative engagement: group norm characterized as coordinated and responsive 1: low 2: moderate 3: high group interactions characterized by lack of coordination with the following qualities: • separate contributions without attempts or an unwillingness to link (i.e., parallel play); contributions may be unrelated. • imbalance in perspectives due to dominant/power differential • no attempts to revisit a groupmate’s previous contribution. reject without (conceptual) rationale ignoring (and not returning to idea) unresponsive to questions repetition of one idea, without modifications to incorporate other’s ideas physicality limited eye contact, turning away to another task, spatial distancing low ratings are assigned when there is no content, practices or assigned task to coordinate around, such as during off-task activity. group interactions characterized by intermittent or mixed interactions with the following qualities: a subset of high-quality indicators are present and/or are inconsistent or limited coordination because first response is taken-up as group response with limited or no discussion, elaboration, modification or checking for agreement group characterized by coordinated interactions, in consistent ways with the following qualities: • students build from and are responsive to ideas • diversity in perspectives are solicited and integrated in ways that are balanced among the group • reject groupmates’ ideas with rationale elaborating, integrating and /or adding on to one another’s contributions responsive to questions, feedback when multiple ideas are voiced or solicited, each is considered efforts to build a group response, consensus, and reconcile across contributions, perspectives, or negotiate; taking up one perspective with rationales physicality coordinated, seamless activity with flow, including nonverbal activity; eye contact; spatial closeness; leaning in, turning toward rogat et al 20 | f l r metacognitive engagement dimension: group norm characterized by socially shared regulation and coregulation focused on content and/or practice, and supported by regulation aimed at maintaining on-task behavior, monitoring of group process, time use, productive emotions, and following task directions. no rating 1: low 2: moderate 3: high no observed regulation group norm characterized by ineffective regulation (low-quality regulation or regulation is not pursued not taken up/ accepted, or it is ignored), obstructing task progress unable to cohere around a common task goal or plan. planning occurring late in the task; repeated return to plan with limited task progress or enactment of task. sustained emphasis on behavioral regulation, distracting other regulation and task engagement. monitoring reveals problems with planning, rather than task. no time remains at the end of the task for evaluation. group norm characterized by effective regulation taken up/accepted, but merely focuses on task completion, task directions, processes. planning and monitoring toward task completion (e.g., checking spelling and formatting; meeting task requirements), but not more. this yields lower quality regulation of the task goal. evaluation is brief, and recognizes completion or meeting minimal task requirements. group norm characterized by effective regulation toward high-quality understanding reflected in the task, group-set goals for understanding. regulation is taken up/accepted. planning and monitoring toward task or group-set goals (as supported by the curriculum) focus on understanding, improvement, progress, integration, consensus, revisions, task quality as exemplified in task expectations evaluation at the end of a task of whether making progress or meeting their goals. rogat et al 21 | f l r disciplinary engagement dimension: content of collaborative talk or physical activity characterized by new contributions aimed at making intellectual progress, involving integrated conceptual and disciplinary activity 1: low 2: moderate-low 3: moderate 4: high group norm characterized by limited to no collaborative content or disciplinary talk and physical activity group disengagement; limited task work independent activity with no content/disciplinary talk or gesture group norm characterized by fragmented talk, with no elaboration or attempts to connect (e.g., restating terms; recall of discrete facts) or focus on content and practices as facts, memorization, recall, or reproduction of practices brainstorming, eliciting prior relevant knowledge, building or independent activity with gesture and physicality, without ontask discourse group norm characterized by content of collaborative talk or physical activity involves some brief elaboration or connections of facts, terms, content and/or practices; elaborative telling brief elaboration of a term brief or initial work toward a connection group norm characterized by content of collaborative talk or physical activity integrates content with practice or content or practice, with rationale/explanation, toward solving lesson/unit problem; intellectual progress group explicitly identifies how their content and/or practice activity generates needed knowledge to solve task/problem synthesis, conceptual connections, connections between content and practice, extended elaboration that informs conceptual development, justifications/rationale note. italics provide example indicators of ratings. an updated version of the rubric is available from the authors. frontline learning research vol. 13 no. 2 (2025) 1 9 issn 2295-3159 corresponding author: ricardo böheim, technical university of munich, tum school of social sciences and technology, department educational sciences, arcisstraße 21, 80333 munich, germany. email: ricardo.boeheim@tum.de. doi: https://doi.org/10.14786/flr.v13i2.1646 perspectives on momentary engagement and learning situated in classroom contexts ricardo böheim1 & jennifer e. symonds2 1 technical university of munich, tum school of social sciences and technology, department educational science, germany 2 social research institute, ioe, ucl's faculty of education and society, university college london, united kingdom article received 29 november 2024 / article revised 24 february 2025 / accepted 8 march 2025/ available online 14 march 2025 abstract in recent years, there has been a strong call for more fine-grained analyses of student engagement to better capture its nature as a situated, momentary phenomenon. this special issue aims to promote cross-disciplinary discussions about the complex processes involved in students’ momentary engagement and learning situated in classroom contexts. momentary engagement is conceptualised as students’ involvement with learning activities over short time intervals. we begin by presenting definitional, conceptual, and methodological reflections on the construct of momentary engagement, highlighting how moment-to-moment analyses can deepen our understanding of how engagement unfolds in complex, dynamic learning environments. next, we discuss the need for a holistic and multidisciplinary perspective to foster an integrative understanding of contexts and conditions under which students engage in academic tasks. finally, we provide a brief overview of the papers in this special issue, emphasising their diverse methodological approaches to capturing students’ momentary engagement and summarising their main results that offer practical insights on supporting engagement. each contribution reflects the efforts of a multidisciplinary team who have studied students’ momentary engagement and learning across various contexts, combining insights and identifying cross-disciplinary synergies in theory and method. the authors integrate perspectives from various fields of research, including motivation, emotion, self-regulation, engagement, social interaction and conceptual change. keywords: momentary engagement; situational engagement, classroom context, classroom behaviour, social interaction böheim & symonds special issue: perspectives on momentary engagement and learning situated in classroom contexts 2 | f l r 1. introduction how can we describe, explain and predict students’ engagement during the execution of learning activities? why are some learners actively engaged while others struggle to stay focused? what can educators do to facilitate meaningful interactions and constructive involvement in academic tasks? educators, policymakers, and researchers are continually striving to find the best possible answers to these questions, aiming to enhance student engagement and improve learning outcomes. these complex and multifaceted questions do not lend themselves to simple solutions because research shows that every moment of every day matters for every student. understanding how engagement unfolds from one moment to the next requires a fine-grained analysis of the specific context and conditions that support or thwart students’ engagement during learning. this special issue presents a set of five empirical studies and one commentary that explore students’ momentary engagement situated in various classroom contexts. factors influencing momentary engagement and learning are numerous and often interconnected, encompassing cognitive, emotional, social, and environmental aspects. to develop a comprehensive understanding of students’ momentary engagement in dynamic classroom environments, it therefore requires an interdisciplinary perspective that draws on insights from different academic disciplines. taking a multidisciplinary approach, the special issue integrates perspectives from different fields of research including motivation, emotion, self-regulation, engagement, social interaction and conceptual change. each contribution represents the efforts of a multidisciplinary team of authors who have studied students’ momentary engagement and learning across various contexts, combining insights from their respective fields and identifying cross-disciplinary synergies in theoretical perspectives and research methodologies. 2. momentary student engagement: definitional, conceptual, and methodological reflections in educational research, engagement has become one of the most prominent constructs because of the accumulating evidence that links engagement to highly valued learning outcomes (fredricks et al., 2004; reschly & christenson, 2022). while researchers widely acknowledge the significance of this construct for achievement and academic progress, there is considerably less consensus regarding its conceptualisation and measurement across various studies (eccles, 2016; skinner & raine, 2022; wong & liem, 2022). at heart, engagement draws on the idea of active behaviour, investment or participation in learning activities and is commonly understood as a multidimensional construct that conceptualises how students think (cognitive engagement), feel (emotional engagement) and act (behavioural engagement) during the execution of academic activities (fredricks, 2022). theories of engagement are often rooted in the motivation literature, emphasizing the importance of external contexts (e.g., family, schools, peers) as well as learners’ internal dispositions and appraisals (e.g., task value, competencies), both of which are assumed to shape students’ engagement and learning (reeve, 2012; reschly & christenson, 2012; wang et al., 2019). engagement can be conceptualised across different timescales (e.g., seconds to years) and at different levels such as schools (e.g., engagement in extracurricular activities), classrooms (e.g., collaborative problem-solving with peers) or specific learning tasks (azevedo, 2015; skinner & raine, 2022; wong & liem, 2022). it is this multilevel complexity within the construct that challenges researchers to find and agree upon a shared construct terminology and definition. one important distinction involves the separation of macrolevel school engagement from microlevel learning engagement as suggested, for example, by the dual component framework of student engagement (wong & liem, 2022). on each level, engagement can occur in different timescales such as moments, days or weeks. recently, researchers have reminded us that the concept of engagement was originally introduced with an emphasis on its situated and momentary nature (eccles, 2016; eccles & wigfield, 2020). however, this focus is not reflected in the academic literature, where research on engagement occurring across momentary time has received relatively little attention thus far (salmelaaro et al., 2021; symonds et al., 2024). in fact, in the current edition of the handbook of research on student engagement (reschly & christenson, 2022), the terms “momentary engagement” or “situated böheim & symonds special issue: perspectives on momentary engagement and learning situated in classroom contexts 3 | f l r engagement” are not mentioned even once. it is therefore understandable that, recently, there has been a strong call for more fine-grained analyses of student engagement to better comprehend its nature as a situated, momentary phenomenon (symonds et al., 2021; symonds et al., 2024). so, what is momentary engagement and why does it matter? momentary engagement describes students’ interactions with learning activities across short time intervals. symonds et al. (2024) suggest classifying microand macrolevel engagement according to aspects of agent, task and time. students’ momentary engagement is located at the microlevel and encompasses the learning experiences of individual students during the execution of a specific task across momentary time, i.e., a child solving a mathematics problem. however, a recent review of the literature revealed that there is no agreed upon definition of momentary or situational engagement among scholars (symonds et al., 2024). possibly due to the challenges of measuring engagement occurring at the momentary time scale, most of the research in the field is focused on students interacting with a broader context such as engagement in subjects, schoolwork, or schooling (i.e., participating in mathematics, school-related tasks, or extracurricular activities). in contrast, studying engagement at a microlevel time scale allows for a deep exploration and understanding of the specific conditions that promote or hinder meaningful engagement and learning. this includes a precise description of the learning context, the nature of the academic task, the learners’ prior experiences and prerequisites, as well as the type of interactions with the teacher or peers. exploring the variance in engagement, longitudinal research has shown considerable variability both within and between students (böheim et al., 2024; martin et al., 2015; patall et al., 2016). it is this variability that supports the need for a moment-to-moment analysis with a detailed exploration of the learning context and the interaction between the task and the individual learner. insights gained from such research hold significant practical value, enabling the design of effective learning tasks that are tailored to the needs of individual students across different contexts. momentary engagement and situated learning share a conceptual connection in that both emphasize the importance of context, task, and social interaction in learning (brown et al., 1989; maclellan, 1996; symonds et al., 2024; wang et al., 2019). situated learning, which has a longer history in educational research (see, brown et al., 1989), posits that the process of learning is deeply embedded in its environment and highlights the role of social interaction and real-world application in knowledge construction (billett, 1996). in contrast, momentary engagement is a relatively new construct that examines transient, fluctuating levels of behavioural, cognitive, and affective involvement at specific moments within a learning activity (symonds et al., 2021; symonds et al., 2024). while research on situated learning faces similar methodological challenges as discussed before, some scholars have examined learning in specific classroom situations. for example, research on actor-oriented transfer explores how learners perceive and construct connections between contexts (lobato, 2003, 2006), while research on expansive framing investigates the application of knowledge across different contexts (engle, 2006; engle et al., 2012). however, although these studies emphasize the role of context in learning, they have a strong focus on knowledge construction and transfer rather than the dynamic, momentary fluctuations of emotion, cognitive focus and behaviour that are central to research on momentary engagement. the use of appropriate methodologies is critical to accurately capture the cognitive, behavioural, and emotional processes that occur while students are engaged in academic tasks (azevedo, 2015). researchers have studied engagement from multiple perspectives (i.e., students, teachers and observers) and with different measures (e.g., self-reports, observations, experience sampling; for a review, see fredricks et al., 2019). contributors from the present special issue triangulate different methods including systematic observation, text logs, videoanalysis, physiological data or self-reports and use multiple informants (peers, teachers, students, researchers) to capture engagement data from different viewpoints. student engagement is measured at the momentary level reflecting on learning situations at the microlevel grain size of time. for example, tang et al. (this issue), used the experience sampling method (esm) to assess students’ situational engagement during a lesson. baines et al. (this issue) and symonds et al. (this issue) conducted systematic classroom observations of student momentary engagement, using time intervals of 10 to 30 seconds. renninger et al. (this issue) analysed moment-toböheim & symonds special issue: perspectives on momentary engagement and learning situated in classroom contexts 4 | f l r moment records of groups of students’ engagement, based on chat logs from the virtual environment in which they were working. and finally, haataja et al. (this issue) captured students’ physiological data to explore momentary engagement in collaborative learning tasks measured via electrodermal activity. the commonality across these measures is that they allow researchers to collect engagement data on a momentary time scale within the scope of well-defined educational contexts. 3. advancing a holistic perspective on momentary engagement among a multidisciplinary network of educational researchers from 2020 to 2023, the european association for research on learning and instruction (earli) and the jacobs foundation funded an emerging field group (efg) titled the integrated model of momentary learning in context (immolic). the group was facilitated by the authors of this paper and the goal of this group was to bring together perspectives on how students learn across seconds and minutes in classrooms. the group was inspired by the development of the momentary engagement construct in educational psychology; work that was prompted in the 2010s when motivation researchers in the us conceptualised different grain sizes of engagement (sinatra et al., 2015; skinner & pitzer, 2012) and began calling for attention to how students engage in their learning across real time, i.e., ‘when the rubber [of the car tyres] hits the road’ (eccles, 2016; kaplan, 2016). a group of european colleagues (symonds, upadyaya, and salmela-aro) began collaborating on the momentary perspective in 2015, developing it into an intervention programme (torsney & symonds, 2019) and building a writing team with us motivation and engagement researchers (kaplan, eccles, and skinner). kaplan had been working on a dynamic systems approach to understanding identity (garner & kaplan, 2019) and the writing team applied the same approach to the momentary engagement construct to develop it into a process-based perspective that highlighted the complexity of learning in the moment. in the momentary engagement perspective, engagement comprises multiple components (motivation, emotion, cognitive action, and physical action) which co-act simultaneously and sequentially as the person interacts with the environment, creating a fluid and responsive dynamic psychological system (symonds et al., 2024). to understand engagement as a momentary dynamic system requires researchers to go beyond static models and concepts that assume individual psychology functions as mechanistic interactions between separate cognitive components. this type of reasoning is built more on current statistical modelling technology and less on what happens in the real world as an individual student engages with work in class. the limitations in the design and application of previous models of motivation and engagement inspired a bid for an earli efg award to support a group of researchers comprising early career and established academics from across higher and lower income countries in europe. the efg’s mission was to advance the field and help future theory building by integrating models and theories of engagement, motivation, emotion, meta-cognition, and conceptual change, into a holistic perspective on how students learn in real time in classroom settings. across the three years of the award, the efg built a model library consisting of a range of established theoretical perspectives, ran an online ‘model club’ where each group member presented a different conceptual perspective once per month, and hosted in-person workshops in ireland and greece which focused on networking, collaboration, and integrating perspectives through collaborative research activity. this special issue is the culmination of that work. 4. different perspectives on momentary engagement and learning this special issue set out to promote cross-disciplinary discussions on the complex process of students’ momentary engagement and learning in classroom contexts. each of the five empirical contributions focuses on a different aspect of momentary and situated engagement, learning and performance. it brings together research from different disciplines that draw on varying theoretical frameworks to conceptualize, measure and study momentary engagement. participating authors come böheim & symonds special issue: perspectives on momentary engagement and learning situated in classroom contexts 5 | f l r from different fields of research on motivation, emotion, self-regulation, engagement, social interaction and conceptual change. integrating research from diverse disciplinary perspectives holds great potential for exploring the contexts and conditions under which students engage in academic tasks, for identifying what keeps them engaged, and for understanding why their engagement may change from one moment to the next. each contribution investigates unique research questions on individual and contextual factors that are assumed to be central for understanding students’ momentary engagement. the results presented across the contributions highlight various factors influencing situational engagement and learning. these include social aspects of context, such as peer relations or group dynamics; aspects of the learning context, like class size; and individual factors, such as self-appraisals, experiences of selfregulation, emotions and motivation. across contributions momentary situated engagement is captured at different time scales and analyses are performed at different levels of granularity. the first article by tang and colleagues examines optimal learning moments (olm; the moments of being highly challenged, skilled and interested) in finnish and us science classrooms, exploring how students’ perception of challenge, skill, and interest in task engagement is shaped by various co-occurring experiences. using experience sampling methodology and network analysis, the study analysed multiple responses from high school students to understand which situational experiences and feelings correlate with olm. the main findings reveal that olm are frequently associated with feelings of concentration, success, and control, meeting expectations, and positive emotions such as enjoyment of the task, while negative emotions like boredom and loneliness are rarely present. the discussion highlights that fostering positive attitudes toward science and encouraging creative activities can enhance situational engagement, while competitive feelings and a sense of pride can be beneficial, provided they are maintained at a moderate level. the study points at the strengths of experience sampling to assess engagement across momentary time combined with collecting data on cooccurring situational learning experiences elicited by the task context. in the next article, baines and colleagues shed light on the relevance of peer relations (both academic and social) for students’ observed momentary on-task and off-task engagement, as well as their link to academic performance in primary schools. using a multi-method approach, including peer and teacher reports, self-assessments, and classroom observations, the study examines different levels of engagement (momentary, classroom, and school). findings indicate that academically focused peer relationships are more strongly linked to momentary engagement and academic success compared to socially oriented peer relations, which are more associated with measures of disengagement on the class and school level. looking at the results and the associations between the different engagement measures, it becomes evident that momentary engagement is empirically and conceptually different from class and school-level engagement reflecting a unique behavioural phenomenon beyond these more general measures of engagement. results draw our attention towards the complexity of peer contexts and their fundamental role in understanding students’ active engagement in classroom learning. moreover, the study calls for more research that is mindful of the nested nature (e.g., school, classroom, learning task) and different time scales (e.g., moments, lessons) associated with the engagement construct. the next study by symonds and colleagues examines how class size influences children's momentary behavioural engagement in irish primary schools. using systematic observations of over 600 children in 121 classrooms, momentary data were collected on students’ on-task and off-task behaviour during regular classroom instruction. multilevel path models showed that smaller classes were associated with higher on-task engagement, while off-task behaviour was higher in larger classes. analyses incorporating various individual and classroom-related factors illustrate the inherent complexity of this relationship. for instance, lower ability was a risk factor for students’ momentary engagement in larger classrooms, whereas momentary disengagement was lower in larger classrooms with a higher proportion of students with special educational needs. the study highlights that students’ momentary engagement is situated within a complex network of contextual factors that can enhance or mitigate their individual effects when considered together. the study by renninger and colleagues investigates how middle school students' phases of problem-solving (exploring, constructing, and checking) relate to their use of executive functions and böheim & symonds special issue: perspectives on momentary engagement and learning situated in classroom contexts 6 | f l r collaborative problem-solving behaviours during moments of math activity. they studied two groups of students who worked collaboratively on open-ended geometry tasks in the virtual math teams environment; each group communicated with one another via a chat tool. for each identified math moment, researchers analysed shared use of a dynamic whiteboard and their chat logs to assess students’ momentary cognitive and behavioural engagement during the phases of problem solving. study findings showed that the student groups varied in their cognitive and behavioural engagement in different phases of problem solving. the results pointed to the benefits of students engaging in exploration in particular, as it was associated with their increased use of working memory and collaboration. the results also suggested that students may need support to collaborate during moments when they are engaging in constructing and checking. finally, haataja and colleagues examine how students' momentary engagement in collaborative learning (cl) is related to self-reported perceptions of cl, video-coded regulation of learning and group performance. using multimodal data, including electrodermal activity and video recordings, the study analysed 94 students collaborating on physics tasks across multiple lessons. students’ physiological synchrony (ps) was used as a proxy for momentary engagement in cl. findings indicate that ps was associated with situated value appraisals of collaboration and exam performance. this study highlights benefits of integrating multiple data channels to gain a comprehensive picture of the context and the complex nature related to the multidimensional phenomenon of momentary engagement. in line with the other papers of this special issue, the authors found considerable variations in their engagement data, calling for more fine-grained analyses of momentary engagement that capture its temporal dynamic from one moment to the next. the special issue concludes with a commentary by kyriakopoulou, who is a well-established scholar in the field of conceptual change research. in her commentary, kyriakopoulou discusses how conceptualisations of momentary engagement differ across studies and how this field of research would benefit from a more integrative, holistic perspective to further expand our understanding of momentary engagement through the lens of conceptual change research. the commentary highlights the multifaceted nature of momentary engagement, particularly when addressing complex learning processes associated with conceptual change. it emphasises the need for a holistic view of momentary engagement that integrates components within both the individual and the context. learner characteristics, prior knowledge, and epistemic beliefs are thought to significantly shape how students engage, particularly in challenging learning contexts. the discussion clarifies that momentary task engagement can involve deep learning processes such as negotiating conflicting prior knowledge and adapting it to new concepts. regarding methodology, it recommends triangulating data from different channels to capture the dynamic and situated nature of momentary engagement. the commentary emphasises that current conceptualisations of momentary engagement often overlook the complex nature of situated engagement, particularly in contexts where students need to revise deeply held beliefs or confront conflicting information. it therefore argues for a shift toward viewing momentary engagement as an interconnected, dynamic system where different components interact, rather than as isolated elements. the commentary closes by encouraging future research to promote a more integrative and holistic perspective on momentary engagement to deepen our understanding of how engagement unfolds in real-time and how it can be better supported in educational learning settings. böheim & symonds special issue: perspectives on momentary engagement and learning situated in classroom contexts 7 | f l r keypoints definitional, conceptual, and methodological reflections on momentary engagement are presented. the need for a holistic and complex perspective on momentary engagement is discussed. the history and activities of the integrated model of momentary learning in context (immolic) emerging field group are explained. a summary of papers in the special issue is given. acknowledgments this work has been funded by an emerging field group grant by earli (european association of research on learning and instruction) and the jacobs foundation, awarded to jennifer symonds and ricardo böheim. references azevedo, r. (2015). defining and measuring engagement and learning in science: conceptual, theoretical, methodological, and analytical issues. educational psychologist, 50(1), 84–94. https://doi.org/10.1080/00461520.2015.1004069 billett, s. (1996). situated learning: bridging sociocultural and cognitive theorising. learning and instruction, 6(3), 263–280. https://doi.org/10.1016/0959-4752(96)00006-0 böheim, r., daumiller, m., & seidel, t. (2024). a longitudinal study of student hand raising: stability and reciprocal dynamics with cognitive elaboration and academic self-concept. journal of educational psychology, 116(2), 297–315. https://doi.org/10.1037/edu0000838 brown, j. s., collins, a., & duguid, p. (1989). situated cognition and the culture of learning. educational researcher, 18(1), 32–42. https://doi.org/10.3102/0013189x018001032 eccles, j. s. (2016). engagement: where to next? learning and instruction, 43, 71–75. https://doi.org/10.1016/j.learninstruc.2016.02.003 eccles, j. s., & wigfield, a. (2020). from expectancy-value theory to situated expectancy-value theory: a developmental, social cognitive, and sociocultural perspective on motivation. contemporary educational psychology, 61, 101859. https://doi.org/10.1016/j.cedpsych.2020.101859 engle, r. a. (2006). framing interactions to foster generative learning: a situative explanation of transfer in a community of learners classroom. journal of the learning sciences, 15(4), 451– 498. https://doi.org/10.1207/s15327809jls1504_2 engle, r. a., lam, d. p., meyer, x. s., & nix, s. e. (2012). how does expansive framing promote transfer? several proposed explanations and a research agenda for investigating them. educational psychologist, 47(3), 215–231. https://doi.org/10.1080/00461520.2012.695678 fredricks, j. a. (2022). the measurement of student engagement: methodological advances and comparison of new self-report instruments. in a. l. reschly & s. l. christenson (eds.), handbook of research on student engagement (2nd ed. 2022, pp. 597–616). springer international publishing; imprint springer. böheim & symonds special issue: perspectives on momentary engagement and learning situated in classroom contexts 8 | f l r fredricks, j. a., blumenfeld, p. c., & paris, a. h. (2004). school engagement: potential of the concept, state of the evidence. review of educational research, 74(1), 59–109. https://doi.org/10.3102/00346543074001059 fredricks, j. a., hofkens, t., & wang, m.‑t. (2019). addressing the challenge of measuring student engagement. in k. a. renninger & s. e. hidi (eds.), cambridge handbook on motivation and learning (pp. 689–712). cambridge university press. garner, j. k., & kaplan, a. (2019). a complex dynamic systems perspective on teacher learning and identity formation: an instrumental case. teachers and teaching, 25(1), 7–33. https://doi.org/10.1080/13540602.2018.1533811 kaplan, a. (2016). research on motivation and achievement: infatuation with constructs and losing sight of the phenomenon. international conference on motivation, thessaloniki, greece. lobato, j. (2003). how design experiments can inform a rethinking of transfer and vice versa. educational researcher, 32(1), 17–20. https://doi.org/10.3102/0013189x032001017 lobato, j. (2006). alternative perspectives on the transfer of learning: history, issues, and challenges for future research. journal of the learning sciences, 15(4), 431–449. https://doi.org/10.1207/s15327809jls1504_1 maclellan, h. (ed.). (1996). situated learning perspectives (1. print). educational technology publications. martin, a. j., papworth, b., ginns, p., malmberg, l.‑e., collie, r. j., & calvo, r. a. (2015). realtime motivation and engagement during a month at school: every moment of every day for every student matters. learning and individual differences, 38, 26–35. https://doi.org/10.1016/j.lindif.2015.01.014 patall, e. a., vasquez, a. c., steingut, r. r., trimble, s. s., & pituch, k. a. (2016). daily interest, engagement, and autonomy support in the high school science classroom. contemporary educational psychology, 46, 180–194. https://doi.org/10.1016/j.cedpsych.2016.06.002 reeve, j. (2012). a self-determination theory perspective on student engagement. in s. l. christenson, a. l. reschly, & c. wylie (eds.), handbook of research on student engagement (pp. 149–172). springer. reschly, a. l., & christenson, s. l. (2012). jingle, jangle, and conceptual haziness: evolution and future directions of the engagement construct. in s. l. christenson, a. l. reschly, & c. wylie (eds.), handbook of research on student engagement (pp. 3–19). springer. reschly, a. l., & christenson, s. l. (eds.). (2022). handbook of research on student engagement (2nd ed. 2022). springer international publishing; imprint springer. https://doi.org/10.1007/978-3-031-07853-8 salmela-aro, k., tang, x., symonds, j., & upadyaya, k. (2021). student engagement in adolescence: a scoping review of longitudinal studies 2010-2020. journal of research on adolescence, 31(2), 256–272. https://doi.org/10.1111/jora.12619 sinatra, g. m., heddy, b. c., & lombardi, d. (2015). the challenges of defining and measuring student engagement in science. educational psychologist, 50(1), 1–13. https://doi.org/10.1080/00461520.2014.1002924 skinner, e. a., & pitzer, j. r. (2012). developmental dynamics of student engagement, coping, and everyday resilience. in s. l. christenson, a. l. reschly, & c. wylie (eds.), handbook of research on student engagement (pp. 21–44). springer. skinner, e. a., & raine, k. e. (2022). unlocking the positive synergy between engagement and motivation. in a. l. reschly & s. l. christenson (eds.), handbook of research on student engagement (2nd ed. 2022, pp. 25–56). springer international publishing; imprint springer. böheim & symonds special issue: perspectives on momentary engagement and learning situated in classroom contexts 9 | f l r symonds, schreiber, j. b., & symonds, j. (2021). silver linings and storm clouds: divergent profiles of student momentary engagement emerge in response to the same task. journal of educational psychology, 113(6), 1192–1207. https://doi.org/10.1037/edu0000605 symonds, j. e., kaplan, a., upadyaya, k., aro, k. s., torsney, b. m., skinner, e., & eccles, j. s. (2024). momentary student engagement as a dynamic developmental system. journal of theoretical and philosophical psychology. advance online publication. https://doi.org/10.1037/teo0000288 torsney, b. m., & symonds, j. e. (2019). the professional student program for educational resilience: enhancing momentary engagement in classwork. the journal of educational research, 112(6), 676–692. https://doi.org/10.1080/00220671.2019.1687414 wang, m.‑t., degol, j. l., & henry, d. a. (2019). an integrative development-in-socioculturalcontext model for children's engagement in learning. the american psychologist, 74(9), 1086–1102. https://doi.org/10.1037/amp0000522 wong, z. y., & liem, g. a. d. (2022). student engagement: current state of the construct, conceptual refinement, and future research directions. educational psychology review, 34(1), 107–138. https://doi.org/10.1007/s10648-021-09628-3 microsoft word dohn_finalproofs.docx frontline learning research vol. 9 no. 3 (2021) 13-30 issn 2295-3159 info corresponding author: nina bonderup dohn, department of design and communication, university of southern denmark. email: nina@sdu.dk doi: https://doi.org/10.14786/flr.v9i3.733 conceptualizing knowledge transfer as transformation and attunement nina bonderup dohn1 1department of design and communication, university of southern denmark, denmark article received 16 october 2020 / article revised 18 march 2021/ accepted 7 april / available online 19 april abstract this article articulates a new theory on the ontology of knowledge transfer. this involves the work of 1) showing that the question “what happens to knowledge in transfer across divergent contexts?” can be made sense of within a situative approach, 2) providing a new conceptualization of situated knowledge, 3) articulating transfer in terms of knowledge transformation and attunement, and 4) putting the issue of learning to transfer knowledge across divergent contexts (back) on the research agenda. the article builds on a view of knowledge as a unity of know-that, know-how, and know-of; which unity forms a practical embodied perspective with which the agent meets the world in interaction. it is argued that knowledge is situatedly realized in attunement to the requirements, possibilities, and restrictions of the concrete situation, as they dynamically unfold. a framework of context levels for analyzing requirements, possibilities, and restrictions (termed “situational characteristics”) is presented. the levels reflect that an activity will always engage with a domain, in a life-setting, taking place within a societal structure, making use of encompassing cultural practices. it is shown how differences in unities of situational characteristics necessitate the transformation of the knowledge perspective in attunement to the situational characteristics of the new context. towards the end, it is pointed out how this conceptualization of knowledge transfer opens for research into designing and teaching for learning to transfer. three recent projects are referenced as an illustration of the approach. keywords: situative approach; knowledge transfer; knowledge transformation; attunement to situational characteristics; learning design bonderup dohn 14 | flr 1. situating the article: the need for conceptualizing knowledge transfer from within a situative approach the aim of this article is to show that the question “what happens to knowledge in transfer across divergent contexts?” can be made sense of within a situative approach. this will be done by providing a conceptualization of knowledge transfer as transformation across contexts in attunement to situational characteristics. this conceptualization constitutes a novel theory of the ontology of knowledge transfer. though the article draws on previous work by me and my colleagues on knowledge and knowledge transformation, it goes a wide step beyond this previous work by articulating what happens ontologically to knowledge in knowledge transfer: knowledge transforms in a concrete realization of a practical embodied perspective. my more encompassing aim is for the article to help put the issue of learning to transfer knowledge across divergent contexts (back) on the research agenda. a significant goal for education would seem to be exactly that: to help students learn to navigate different contexts and adapt their knowledge and skills along the way to the situational requirements, possibilities and restrictions at hand. the world of today is characterized by diversity, frequent change, and globalization. it is a defining feature of many people’s lives that they traverse a range of different contexts and participate in a variety of different practices. this is true, both in what jarvis’ (2007) called a life-wide perspective (across our lives here and now) and what he called a life-long perspective (the temporal stretch of our lives). accordingly, people often have to use knowledge, learned in one context, in new contexts. failure to do so in adaptation to the situational characteristics of those contexts can have serious consequences – for the tasks people undertake (dealt with inadequately), for them personally (feelings of incompetence, lack of self-esteem, diminished sense of fulfilment), for the people they engage with (interacted with in insensitive, incompetent or unacceptable ways) and, in general, for the organizational and societal environment within which they are acting. arguably, the challenge itself is not new, as people have had to traverse contexts in decades and centuries preceding the present one. however, the challenge is magnified, intensified and diversified today: the number of contexts we engage with have multiplied as compared to earlier times; the complexity and heterogeneity of many of the contexts have increased markedly; and the set of social, economic, and environmental sustainability problems we face globally and need to learn to take into consideration across myriads of local situations, are more urgent and decisive today than ever before. given this, the challenge to transfer knowledge in ways that accommodate adequately (personally, ethically, socially, societally, environmentally, organizationally, economically, etc. speaking) to the situation at hand is of unprecedented significance. the educational goal of supporting students’ learning to transfer knowledge has, however, been seriously contested in the last decades. its viability is called into question by research in practice-theory (dohn, 2017; dreyfus, 1979; dreyfus & dreyfus, 1986; schatzki, knorr-cetina, & von savigny, 2001; schön, 1983), situated learning (greeno, 1997, 2011; greeno & gresalfi, 2008; greeno & tmsmtapg, 1998; lave, 1988; lave & wenger, 1991), and distributed cognition (hutchins, 1995; hutchins & klausen, 1996; salomon, 1993). this research combines to show that knowledge is situated, attaining form and content from the context in which it is learnt. upon accepting the situativity of knowledge, transfer of knowledge between contexts appears problematic (greeno, 1997; tuomi-gröhn & engeström, 2003); so much so that the notion has been argued to be unsustainable (carraher & schliemann, 2002), even incoherent (lave, 1988; packer, 2001). teaching for transfer to happen between, for example, school and work seems impossible, at least from within education (tuomi-gröhn & engeström, 2003). of course, not all researchers have accepted the situativity of knowledge. it is not generally accepted by what lobato terms mainstream cognitive approaches: approaches focusing on learners’ information processing in a cognitive architecture comprised of long-term, short-term and sensory memories (lobato, 2012; mayer, 2001; mayer & massa, 2003; reed, 1993; singley & anderson, 1989). researchers within these approaches investigate transfer in constrained experiments (e.g., in a laboratory), focusing on how learners in trial ‘transfer situations’ apply content knowledge, abstract bonderup dohn 15 | flr schemas and problem-solving strategies learned in an initial learning situation. they thus ignore the criticism of the situative approaches that a) the experimental context delimits tasks in a very specific way and that b) there is no independent evidence that the knowledge and transfer allegedly displayed by participants represent knowledge and transfer in other life contexts. some have acknowledged the significance of ‘context’, but treat it as a pre-given, delimited set of circumstances, which pose situational requirements, possibilities and restrictions determinable in advance, and from which learned knowledge can be “decontextualized”. allegedly, representations, solutions, and so forth can then be “mapped” across such sets of pre-given circumstances (reed, 2012). alternatively, ‘context’ is treated as a set of factors which have to be dealt with by the learner in mental representations as part of the process of transfer (nokes-malach & mestre, 2013). these researchers neglect the fundamental situative point that contexts are constituted in the interaction of agents, content, and activity – parts and wholes define each other, much like a rope is made up of threads, none of which run its full length and which for their part are held in place precisely by being part of the rope (barab & roth, 2006; mcdermott, 1993; säljö, 2000; van oers, 1998b). for this reason, situational requirements, possibilities and restrictions are not fully determined or determinable prior to interaction. in particular, as argued by van oers (1998a), the concept of “decontextualization” is highly problematic because it is in effect only a negative qualification, indicating that something doesn’t happen. arguably, van oers overemphasizes the individual’s awareness when he states that “context is always strongly related to a personal (explicit or implicit) definition of a situation of action” (p. 136), but the more general point that context is strongly related to human sense-making in interaction holds true. therefore “decontextualization” would seem to imply “no situation, no action, no meaning at all” (p. 136). instead of focusing on “decontextualization”, knowledge is better approached as a “functional stance on the interaction” in response to the “affordance networks” of the environment (barab & roth, 2006, p. 3). affordance is here understood as relational to the actual embodied capabilities of the agent (dohn, 2009). however, as the concepts of “affordance” and “affordance networks” centre on the possibilities of the situation (positive or negative), they do not fully cover the restrictions and requirements of the situation. furthermore, the concept of “affordances” is fraught with diverging understandings in the literature. for these reasons, i here choose a different terminology and speak of the requirements, possibilities and restrictions of a situation and refer to these collectively as “situational characteristics”.1 somewhat ironically, perhaps, researchers who do accept the situativity of knowledge have generally refrained from investigating how people’s situated knowledge in one context might potentially relate in positive ways to their situated knowledge in other, very different contexts. that is, it is underinvestigated how people’s knowledge transforms across major shifts in context. instead, such researchers have contented themselves with documenting the lack of transfer between such contexts (lave, 1988; nielsen & kvale, 2003; wedege, 1999), or have studied in a much broader way how “boundary crossing” leads to shifts in social relations, participatory roles, identities, accountabilities, and values (akkerman & bakker, 2011; engeström, 2001; thrysøe, 2011; tuomi-gröhn & engeström, 2003; wenger, 1998). studies which do focus on transfer and transformation of knowledge confine themselves to the microlevel of situative interaction. here, researchers perform detailed analyses of, for example, the ways in which mathematical formulas are interpreted and reused across different tasks (lobato, 2012; wagner, 2010). or they look at how teachers can frame situations, roles, and content to 1 as introduced originally by gibson, “the affordances of the environment are what it offers the animal, what it provides or furnishes, either for good or ill.” (gibson, 1986, p. 127, italics in original). in other words, affordances are possibilities for action and interaction with others (including possibilities with negative outcomes). constraints of the situation delimit a person’s possibilities but not all constraints need be directly entailed or clear from the situation’s affordances. for this reason, in the reception of gibson’s view, the concept of “constraints” has often been added in the description of situational characteristics (greeno, 1994; norman, 1988/2002, 1999). however, this combination of “affordances and constraints” still does not capture all situational characteristics, because it does not cover the sense in which a situation may require one to do something, for example, save a drowning child. for this reason, i add “requirements” to “possibilities and restrictions” in the delineation of situational characteristics. bonderup dohn 16 | flr foster students’ knowledge transfer between different curricular lessons, for example, between assignments and instances of group work (engle, 2006; engle, lam, meyer, & nix, 2012). interesting and informative vis-a-vis the development of student understanding as they are, such microlevel studies take place within the same overall institutional context of formal education. it is thus an open question to which extent they can inform the questions of how people transfer and transform their knowledge across overall institutional contexts, and, in particular, whether it is possible educationally to support students in learning to do so. 2. conceptualizing knowledge taking a step back, an important first point to clarify in conceptualizing knowledge is whether it is first and foremost a phenomenon ascribed to persons (and perhaps animals) or to a set of propositions/abstract ideas. do humans have/incorporate/display/enact knowledge – or do books (as representations of the abstract ideas)? popper famously formulated this question as the question of whether knowledge resides in world 2, that is, the human world (where world 1 is the physical world) or world 3, that is, the world of ideas (popper, 1972). popper sided with the latter, arguing that though humans produce the abstract ideas, their production brings the ideas into independent existence as knowledge objects. only in a derivative sense can we, according to popper, ascribe knowledge to persons, namely when they understand the abstract ideas – or, as one might alternatively put it, depending on one’s theoretical preferences: when they develop a conception of, master, interpret, make their own, and so forth the abstract ideas. further, popper’s view can be developed to allow this derivative sense to be used about not only individuals, but also groups, for example, about a scientific community’s “state-of-knowledge” (see e.g., bereiter, 2002). on this developed popperian view, the world 2 understanding of the world 3 abstract ideas may be distributed between the scientific community’s members, rather than any single member understanding all the abstract ideas. there are many problems with this view, not least the obvious one that understood this way, knowledge is – quite literally – not useful, not even for its own sake; it cannot grow, inspire, or challenge; it cannot help us out of ignorance; it cannot bring about any of the processes or circumstances which knowledge is usually held accountable for: abstract ideas residing in a realm of their own cannot do anything or lead anywhere; humans (and perhaps animals) who understand or master knowledge can. in any consideration of why knowledge is important, including considerations concerning pursuing it for its own sake, human understanding is central, not the abstract ideas per se. the point is both ontological and epistemological: ontologically, the point is that the kind of existence which knowledge, in the sense of world 3-abstract-objects, could have, would be deficient as compared to its realization-throughhuman-(and-animal)-understanding in so-called worlds 2 and 1. epistemologically, the point is that abstract ideas always need interpretation (i.e., human or animal understanding) to be applicable. because of problems such as these, i do not agree with this view. on the contrary, on my view, knowledge in its primary sense is a phenomenon ascribed to humans, and only derivatively to books or abstract ideas represented in books (derivatively, because books and ideas may be said to embody that which persons’ knowledge is about). incidentally, on my view, it is in a similar derivate sense that one speaks of knowledge as embodied in material and virtual artefacts. but even if one takes a popperian approach and views talk of a person’s or group’s knowledge as derivative to a primary world-3 sense, the question argued above to be of prime importance remains, that is, how persons and groups make use of knowledge learned in one context in new ones. whether this use of the term ‘knowledge’ is construed as constituting its primary or derivative sense, the question concerns what people do in interaction with each other and the situations they come in. since world 3 abstract ideas can’t do anything, the ontological elevation of abstract ideas does nothing to answer this question. bonderup dohn 17 | flr an apparent middle ground between ascribing knowledge to humans and to abstract ideas is taken by radford (radford, 2013). radford interestingly provides a hegelian articulation of knowledge as abstract potentiality which comes into concrete existence – acquires reality in a singular manifestation – through particularization in a specific activity. in this way, the activity, as a particular, mediates between the singular and the general. radford’s view accommodates – indeed hinges on – the point just made: that abstract objects are ontologically deficient (“pure possibility” as he says, p. 18) and require actualization in a concrete singular to come into existence. however, this comes at the expense of the somewhat enigmatic notion of knowledge as “pure possibility”, existing “in itself” (p. 25), developing culturally through “successive determination” (p. 15) – reminiscent of hegel’s idea of the spirit as an overarching subject (in the continental sense of the word), concretizing through cultural development. my own view aligns with radford’s in stressing realization in concrete actualization as the full ontological mode of existence, but the potentiality that is being realized is not that of a general, abstract subject, but of the individual’s practical embodied perspective, as i shall explain in the next subsection. 2.1. conceptualizing knowledge as situatedly realized within situative approaches, it is common to discard the term ‘knowledge’ in favour of ‘knowing’. this is done to stress a process view over an objectifying one, that is, to specify that ‘to know’ is to do something rather than to possess something (barab & roth, 2006; greeno & tmsmtapg, 1998; sfard, 1998; wenger, 1998). i choose to speak of ‘knowledge’, however, to not seemingly decide on the fate of transfer by choice of words alone: ‘knowing’ easily misleads to a narrow view of knowing-as-doing as singular, discrete actions. such actions can hardly be envisaged as transferable. but if knowing were only discrete actions, it would be a tremendous coincidence that people can sometimes consistently and repeatedly perform the same actions, whether in new or wellknown contexts. conversely, using the term ‘knowledge’ does not preclude an analysis depicting knowledge as realized first and foremost in action. quite the opposite, such analyses have been put forward by a number of practice philosophers, some of whom have inspired the situative approaches (dohn, 2017; dreyfus, 1979; dreyfus & dreyfus, 1986; molander, 1992; schatzki et al., 2001; wittgenstein, 1958, 1969/1979). across differences between these philosophers, their analyses concur in ascribing full ontological existence to knowledge only as actualized in action. that is, they concur in providing analyses that do not reify knowledge, even though they retain the seemingly object-indicating noun2. many of these practice philosophers take an ecological approach (gibson, 1986), stressing the basic reciprocity of individual and environment. often with reference to merleau-ponty’s concept of body-schema (merleau-ponty, 1962), this reciprocity is explained as a pre-reflective correspondence between what we can do in the world (skills incorporated in our body-schema) and the meaning it takes on for us (its affordances for us) (dohn, 2009, 2017; gallagher, 2005; sanders, 1993). based on this work, i have elsewhere explicated knowledge as a unity of propositional knowledge (know-that), practical knowledge (know-how), and experiential knowledge (know-of) and argued that this unity forms a practical embodied perspective with which the agent meets the world in interaction (dohn, 2017). this knowledge perspective (as i also term it) lets the world present itself as meaningful to the person in accordance with the holistic set of interests, needs and capabilities of that person. importantly, not all constituents of this holistic set need be consciously recognized, which is part of what is meant by stating that the perspective is practical and embodied. it is practical and embodied in the further sense that meaningfulness will present itself with immediacy, before reflection, as a practical ‘feel for’ the situation and what needs to be done in it. finally, it is practical and embodied also in the sense that the perspective has its full ontological realization in use, taking on concrete form and content from the 2 incidentally, researchers who prefer to speak of ‘knowing’ would seem to invite the question “yes, you have explained what people do when they know; now could you please explain what knowledge is”, inadvertently leading to a reinstatement of a reifying view of knowledge, precisely because they do not give the objectindicating noun a non-reified interpretation. bonderup dohn 18 | flr specific situation in response to the requirements, possibilities and restrictions posed by the situation. this is where my view aligns with radford’s: in stressing actualization as the full ontological mode of existence. the difference is that the potentiality that is being actualized is not an abstract general possibility, but the individual’s embodied perspective: as bodily beings, we are always in a specific situation in the material world, and in this sense the perspective is always concretely realized – there is no way we as bodily beings cannot have a specific viewpoint on the world, both literally and metaphorically speaking. however, as we move around our environment, the perspective changes (literally and metaphorically), meaning that the perspective is also always the potentiality for being realized as a different viewpoint. in this way, the potentiality is not an abstract form, but rather the possibility for continuous attunement to the situation at hand. this rather abstract articulation of the perspectival character of knowledge begs an illustration with a few examples. my first example is polanyi’s description of the way x-ray pictures look to expert radiologists and to novice students, respectively (polanyi, 1962, p. 101). from the radiologists’ knowledge perspective, the x-ray picture presents itself directly as showing “a rich panorama of significant details”, including physiological variations, scars, and acute signs of illness. the students, for their part, at first do not even see the lungs, but only the ribs. the second example is the similar one supplied by kuhn which concerns the way a bubble chamber photograph looks to the physicist in comparison with what students see: the physicist will see a record of familiar sub-nuclear events whereas students will only see a set of broken lines (kuhn, 1970, p. 111). now, in phrasing these examples in terms of seeing, the practical nature of the perspective is left somewhat implicit. however, ‘seeing’ may well have actionable significance: from the radiologists’ knowledge perspective, x-ray pictures are part of their diagnostic practice and a given picture can bear the significance of calling for specific actions as ways of treating the illness; for example, additional tests, medication, or operation. for the physicist, the significance of bubble chamber photographs resonates with a web of experimental procedures and results inscribed within scientific practice. a third example brings the actionable significance of the knowledge perspective more to the fore: lave and de la rocha have studied how shoppers and weight watchers, respectively, do best-buy calculations in the supermarket and mete out prescribed portions in the kitchen (de la rocha, 1985; lave, 1988). their studies show how the participants’ knowledge perspective of the everyday situation they are in let actions with tangible objects, difference strategies, direct comparison of size, and so forth stand out as the relevant mathematical procedures to undertake. in showing this, both studies further demonstrate how the situation provides form and content to the determination of the shoppers and weight watchers knowledge of mathematical procedures. a common strategy at this point is to distinguish between “generalizable” structures and “surface” (or “accidental” or “particular”) features of a situation. this strategy is most prevalent in cognitivist approaches to transfer (gick & holyoak, 1983; judd, 1908; reed, 1993). however, even within the situative approach, some researchers have made use of the analogous ecological distinction between “invariant” and “variant” elements of a situation (barab & roth, 2006). barab and roth thus claim that facts, concepts, and principles are “invariant structures” which have “cross-contextual value” (p. 3) because they can be used “as tools in other situations at other times” (p. 4). they further claim that “learners [can] interrogate the problem [the task they have been posed] in terms of the invariant and variant aspects” (and that transfer is facilitated by such interrogation) (p. 10). the problem with this approach is that it treats ‘invariant’ and ‘variant’ aspects as independent of each other (and similarly with generalizable structures vs. surface features). this, in turn, leads to a view of knowledge of the aspects as decomposable into independent elements (that can then be recombined in new situations). using lave’s shopper case to illustrate the problem; on the criticized view, arithmetic procedures to calculate unit prices would be invariant aspects; variant ones would be for example, family dietary preferences, layout of the supermarket, and available groceries. however, as i shall argue in greater detail below, the concrete situation decides the role of calculations, how they should be performed, when unit prices are relevant, and even what counts as a unit. in this sense, the invariant aspects are concretely realized and only come into actual being through the variant aspects. bonderup dohn 19 | flr this seriously questions the adequacy of the term ‘invariant’ as applied to concrete situations, and, moreover, highlights that invariant and variant aspects are not independent of one another in their realization. this, it should be noted, is an ontological point concerning the actualization of aspects of a situation. it does not preclude taking a realistic stance to for instance physical mechanisms or to human psychological or sociological traits. one may follow bhaskar (1975) in arguing that such mechanisms and traits are real, understood as potentiality (in a related, but different, sense from the hegelian spirit), and have to be actualized in concrete situations (bhaskar speaks of ‘events’). this actualization is precisely the realization i argue for, where the role which mechanisms and traits get to play in any specific situation is one as concretely realized in unity with other aspects. therefore, the way the mechanisms and traits come into actual being differ due to each situation’s different contextual unity of situational aspects. the corresponding epistemological point is that knowledge is not to be understood as a combination of independent elements. instead, in line with the proposed reciprocity of individual and environment (cf. above), knowledge is situatedly realized as an attunement to and enactment of the concrete unity of the allegedly variant and invariant aspects. this attunement may, but need not, be in part available for conscious reflection, but will, in any case, involve a pre-reflective practical embodied accommodation to the situation, which dynamically upholds the correspondence between body and world (between what we can do in the world and the meaning it has for us, cf. above). this dynamic accommodation may be compared to – indeed is exemplified by – partaking in a dance or a choir or band where skilful participants continuously and unreflectively accommodate to the movements of the others in the specific physical environment. the dynamic accommodation is the actualization of the practical, embodied perspective on the situation which, correspondingly, lets the situation stand out with the meaning it has as a concretely realized unity of situational requirements, possibilities, and restrictions for action. some aspects may be recognizable from other contexts, but they are not related to as invariant, but as situationally concrete, as transformed by the current situation’s contextual unity. furthermore, what aspects will be recognizable from other contexts will depend on the present situation, including persons’ interaction in it, and need not be predictable in advance. 3. a framework of context levels of situational characteristics in this section, i shall present a framework for analysing situational requirements, possibilities, and restrictions (henceforth situational characteristics), previously developed in dohn (2017), hachmann and dohn (2018), and dohn and hansen, s.b. (2020). the framework is based, on the one hand, on the recognition that such situational characteristics analytically pertain to different levels of specificity as regards the activity in question: an activity engages with a domain, in a life-setting, taking place within a societal structure, making use of encompassing cultural practices. on the other hand, the framework takes into account that situational characteristics at these different levels interact. amongst others, the framework has been used to criticize the operationalization of the oecd’s programme for international student assessment (pisa) (dohn, 2007). it has also been used to highlight the incoherent competence demands placed on students when web 2.0 practices are introduced as learning activities within education (dohn, 2009). the framework is inspired by wedege’s distinction between “situation context” and “problem context” (wedege, 1999) and by engle’s distinction between “social context of learning” and “content” (engle et al., 2012). two caveats before proceeding: first, the examples of situational characteristics provided below have been chosen because they are easily recognized. their recognizability may, however, mislead to the impression that they are pre-given and static, that is, exist as concrete invariants across contexts. this is not the case; in every specific situation, their significance and meaning (including a potential negation of them) will be concretely realized in the interaction of the agent(s) in practice. i illustrate this bonderup dohn 20 | flr point with examples below. second, at all levels, some situational characteristics may be implicit, acknowledged in action, rather than explicitly articulated. the framework levels are: 1. the domain level. this level is concerned with the domain or content area; for example, genre theory; linear algebra; nuclear physics; and musical harmonization. examples of situational characteristics at this level could be: fairy tales can include magical happenings, documentary essays cannot; subtracting a negative number equals adding the numerically identical positive number (-( a) = a); calculations involving nuclear phenomena should be performed utilizing quantum mechanics rather than newtonian mechanics. 2. the activity level. this level is concerned with the activity itself; for example, reading a book; individually or collaboratively solving a problem; attending a lecture; writing an entry in wikipedia; and having a discussion with peers. situational characteristics at this level could include: all participants in group work should be allowed to speak; students sit relatively quietly whilst attending lectures; entry writing for wikipedia involves building upon and refining others’ contributions and accepting non-copyrighted “use-and-reuse” of one’s own. a typical word problem in school tasks will delineate a fictional setting in which the problem is supposed to take place – the “problem context” in wedege’s terms (cf. above). contextual requirements of this fictional setting will be specified as part of the problem description, for example, the amount of money available for shopping, or two persons’ opposing views on a subject such as graffiti. it is an activity level situational characteristic that one should take such fictional/problem-story contextual requirements into account in solving the problem. 3. the life-setting level. this level is concerned with the life-setting which frames the activity; for example, shopping for groceries; participating in a class within an educational program; and having leisure time to spend as one pleases. situational characteristics at this level could concern, for example, making do with the actual resources available in the situation; taking food preferences of family members and storage limitations at home into account; acknowledging the teacher as the authority in the classroom; making a self-directed choice as to what one spends leisure time on. this is the level of wedege’s “situation context” and of engle’s “social context of learning”. 4. the societal structure level. this level is concerned with societal organizational structures and institutions which allow or enable the life-setting to exist and the activity to take place within it. examples of societal structures are the organization of learning within a school system or within traditional family-centred apprenticeships with room and board provided, of religious practice through church mediation, and of the distribution of commodities in society through the free market. situational characteristics at this level concern for example, general curriculum demands and national standards; (implicit or explicit) expectations concerning parents’ engagement (or not) in their children’s schooling; formal qualification requirements for certain professional jobs; expectations concerning citizens’ involvement (or not) in religious practice. 5. the cultural practices level. this level is concerned with the cultural tools and ways of behaving which are prevalent in a culture across specific practices and societal structures. examples of cultural practices are dominant communication forms (oral communication in some cultures; reading and writing in others; internet-based communication in many cultures today); the use of money as a medium of exchange; and the production of tools with certain materials (stone in the stone age, iron in the iron age, plastic in modern times). examples of situational characteristics at this level are: contemporary expectations of increased digitalization across societal institutions; preference for oral, written or digitalized communication in different cultures. by ‘domain’ i simply understand ‘what the activity is about’. ‘activity’, similarly, is ‘what participants are engaged in doing’. the scope of both will vary: in the context of a university-level mathematics course, linear systems will be a domain, whereas algebra will be a domain in primary school. likewise, in the context of collaborative problem solving, ‘group discussion’ will be part of the activity, whereas in other situations it may be the activity itself. as regards ‘societal structures’, this bonderup dohn 21 | flr term refers primarily to the state/national level at which stateor nationwide institutions are constituted. however, given the pluralistic nature of today’s societies, two points should be noted about this. firstly, within states and nations, communities and subgroups may have their own societal structures negotiated within or challenging the bounds of the state/national ones. so the scope of this level varies, too. secondly, for individuals, the societal structures at the state/national level need not be appropriated as their own; individuals can be estranged and marginalized by societal structures and experience them as unfair, elite or majority rule. that this can be the case of course only underscores the significance of situational characteristics at this level. similar points apply to the cultural practices level. as indicated, situational characteristics interact. the ones at higher (i.e., less activity-specific) levels frame and delimit the ones at lower levels, determining their specific concretization, and in some cases vice versa. in dohn (2007), i thus show through analysis of the assessment guide to specific pisa test items’ how the situational characteristics of the test situation (level 3: the life-setting level) delimit what is accepted as adequate answers to a question of letter style (level 1: domain level). i also show how the situational characteristics of a couple of non-test situations (level 3) would lead to different evaluations of the provided answers. a further example is provided in yackel and cobb’s (1996) study of the establishment over time of socio-mathematical norms in a second-grade classroom. one of the norms which developed was that students are only justified in contributing a solution to the class discussion when the solution is mathematically different from previously proposed solutions. this norm integrates situational characteristics at • domain level: what counts as mathematically different • activity level: discussion norms • life-setting level: students raise their hands to contribute; teacher regulates the discussion and has the authority to evaluate answers the example illustrates that situational characteristics are not pre-given but develop in the interaction of agents with content and each other in the unfolding of socio-historical place and time. in general, situational characteristics interact to form a concrete unity, and it is in response to this concrete unity that knowledge has its full ontological realization, taking on specific form and content in attunement to it. i illustrate this further in the next section, demonstrating also how different contexts lead to different knowledge realizations. 4. attunement of the practical embodied perspective to the unity of situational characteristics knowledge is learnt in concrete situations with specific unities of situational characteristics. thus, the students in yackel and cobb’s study learn mathematical concepts, principles and arguments (level 1) intertwined with learning • to participate in the specific teacher-questions-students-answer-interaction-format (level 2 discussion norms as framed at level 3 by the classroom) • the respective social positions of teacher authority and student compliance (level 3) • to negotiate these social positions (level 3). because the situational characteristics at different levels interact, decontextualization of knowledge pertaining to the domain level is not possible. any apparent decontextualization will actually just be another concrete realization of knowledge in response to the new situation’s unity of situational characteristics. with inspiration from lave’s (1988) study, this can be illustrated by comparing real bonderup dohn 22 | flr supermarket buys (level 3) with the ones of school word problems (level 2, taking place within the level 3 life-setting of school).3 when shopping for ingredients for an apple pie in the supermarket, situational characteristics at level 3 such as the recipe, storage capacity in the car and at home, family members’ preferences for specific apple brands, for apples and for apple pie over other types of fruit and dessert (pp. 120-121, 162, 166), frame the situational characteristics of the actual choosing of groceries (level 2) and of the use of math for doing so (level 1). attuning to these characteristics involves the shopper’s knowledge at all levels concretizing together into a specific realization of the knowledge perspective in the situation. as a result, at level 1, difference strategies4 (pp. 119-120) or direct comparison of sizes (p. 154) become appropriate uses of math for deciding which apples to buy. the use of math may also be adjusted, for example, what counts as a unit and how units should be added. for instance, when 2 apples are added to 3 apples to make the 4 large ones required by the recipe. alternatively, the brands or prices available may lead the shopper at level 2 to decide to shop elsewhere or make another dessert (p. 166). in contrast, the realization of the knowledge perspective in a school context will let a word problem concerned with the buying of apples for an apple pie stand out as exactly that. this restricts the types of calculation permissible at level 1. difference strategies, direct comparison of size, and the adjustments of units are improper uses of math, unless the word problem specifically asks for it. deciding for another dessert at level 2 is not an option, because the activity is that of solving a math problem, not that of getting dessert. lave suggests “the possibility of an indeterminate number of arithmetics” (lave, 1988, p. 63f) corresponding to diverse situations of arithmetic practice. this suggestion begs the question of why we across all these situations should characterize what is going on as “an arithmetic” rather than something else. the framework i propose supplies an answer: the domain is the same – for example, elementary algebra – across the situations. yet its situational characteristics are concretized by the activity and lifesetting situational characteristics. this provides situation-specific form and content to what works as correct domain procedures in the concrete situation. in the shopping example, the different situational characteristics act to make difference strategies, direct comparison, and unit adjustments workable algebraic procedures in the supermarket, but not in school. transfer of knowledge from either context to the other – to the extent that it takes place at all5 – does not happen as a decontextualization of the algebraic procedures used in the one situation, followed by ‘recontextualization’ in the new situation. rather, the concretized practical embodied perspective of the one situation – with its inherent potentiality of being concretized differently – transform in accommodation to the unity of situational characteristics in the other one, resulting specifically in changes in what works as algebraic procedures. again, the analogy and example of dance is illustrative: dancing is realized in concrete physical and social situations, as an immediate bodily attunement to the unity of situational characteristics. dancing transforms between situations, not as a decontextualization of moves – practising dance moves is a concrete realization of moves in accordance with a unity of situational characteristics as well – but as a dynamic accommodation of movement upholding the pre-reflective correspondence between body and world. the dynamic accommodation is the continuous, concrete realization of the embodied 3 to allow all points to be made with just one example, i have constructed the following dessert example from lave’s observations of several supermarket buys of different groceries. page numbers refer to documentation of the point in question in lave (1988) (involving other groceries than ones for dessert). 4 difference strategies compare differences in price and amount without calculating unit price. an example is “one bag of apples is €3, two bags are €4, so i get a bag extra for just €1.” 5 lave’s research shows that transfer of math knowledge from school to everyday situations is not as direct, widespread and obvious to persons themselves as math educators might wish for. still, a person may transfer knowledge without acknowledging the transfer process; that is, make use of knowledge learned in another context without awareness that it stems from this other context. it may simply be part of the practical embodied perspective with which the person meets the world. lave’s shoppers for instance make use of numbers and basic algebraic procedures of adding and subtracting. these will not in general have been learned in the supermarket, though they need not have been learned in school, either. bonderup dohn 23 | flr perspective on the situation, which as part of the process lets the material and human phenomena present stand out with a significance specific to this situation (e.g., obstacle, an axis to revolve around, a possibility for ‘echoing’ moves, etc.), calling for further moves of the dancer. we do not, of course, always succeed in attuning correctly to the unity of situational characteristics. the practical embodied perspective may concretize in ways which overor underacknowledge the situational characteristics at one or more of the levels. children who write “i do not like apples” as the answer to apple pie word problems in school have under-acknowledged how the situational characteristics at level 3 frame activity and domain. an example of over-acknowledgement of situational characteristics at both levels 3 and 1 is provided in säljö and wyndhamn (1993). here, students, aged 15 and 16, had to determine the necessary stamp values for different letters. one group did this by calculating first the proportional price-per-weight and then linearly increasing stamp prices. they thus over-acknowledged the level 3 situational characteristics of the school context in interaction with level 1 situational characteristics of elementary algebra. this overruled their out-of-school post office knowledge of weight ranges for stamp price calculation, though this knowledge was supposed to have been applied in the problem-solving at level 2. arguably, children learning to read who stop to spell their way through every street sign over-acknowledge the situational characteristics of the reading activity at level 2. they thereby neglect the situational characteristics at both level 3 (requiring them to move on) and level 1 (what the signs say). these latter examples of inadequate attuning to the unity of situational characteristics highlight the need for transformation of the knowledge perspective across different situations. without such transformation, new situations will be met inadequately with the concretized knowledge perspective of other situations. this will result in apparently ‘obstinate’, inflexible, context-insensitive behaviour. one example is found in engeström (2001). here, health professionals working within different contexts had to deal with patient issues cutting across the contexts. their initial approaches were inadequate, precisely because they were rooted in their respective work contexts, rather than transformed to meet the crosscutting unity of situational characteristics. to reiterate: situational characteristics are not pre-given or static but develop in the interaction of agents in practice. further, precisely how persons transform their knowledge across contexts is rarely predictable in any detail. still, both situational characteristics and persons’ transformation of knowledge will often be analysable after the fact. this is exemplified at the microlevel within the context of school in lobato’s and wagner’s research (cf. above). hachmann and dohn (2020) illustrate it for the larger shift, involved in a practicum experience, between the contexts of school and professional practice. a similar analysis is allowed for in engeström’s description of the health professionals’ subsequent negotiation of new ways of addressing the patient issues in question. moreover, it is possible to reflect in advance on what might potentially count as adequate transformations of knowledge between specified contexts. though the actual situational characteristics of these contexts may turn out to differ, such reflections can heighten awareness of the need to attune as well as to types of considerations that may be appropriate in attuning. reflections can be done from within educational contexts. for instance, teachers can point to transformed uses of curricular content in learners’ current out-of-school practices (life-wide perspective) as well as in practices in which they might engage in the future (life-long perspective). such pointers will frame the use of curricular content, potentially with the much more far-reaching significance intended by engle et al. for the concept of framing, but not actually realized in their empirical research on microlevel framing between curricular lessons (engle, 2006; engle et al., 2012). if successful, it may help learners both in effecting actual specific transfer between contexts and in developing dispositions for flexible attunement which may facilitate transfer for them more generally. for the sake of clarity of presentation, i have restricted my examples to contexts which differ at the life-setting level (level 3), but which are found within the same society. this means that the state/national situational characteristics at levels 4 and 5 are much the same, though they may have differing significance at level 3 for the different life-settings, and may be negotiated, lived with and challenged bonderup dohn 24 | flr very differently by specific communities and individuals within the life-settings. however, teachers could point to transformations across contexts which differ at these levels, too. this may be relevant for international students enrolling in a schooling system diverging markedly from the one in their home country (level 4) (cf. carroll & ryan, 2005; kandiko & weyers, 2013). variances for such students may also concern cultural practices (level 5), for example, the degree of digitalization and the value accorded to it. 5. putting the issue of learning to transfer knowledge (back) on the research agenda the context level framework allows asking from within a situative approach what happens to knowledge in transfer across divergent contexts. it makes possible the analysis of how the situativity plays out in different situations as the attunement to diverging unities of interacting situational characteristics. that is, the framework on the one hand allows acknowledging that knowledge is realized situatedly in a unique way for each situation. on the other hand, it permits maintaining that the format of the situativity is the same across instances, namely attunement to a unique unity of situational characteristics developing interactionally. this justifies research questions such as: how do people learn to attune to developing unities of situational characteristics? how do they transform their concretized knowledge perspective from one context to fit the unity of situational characteristics in new ones? moreover, the framework makes it possible from within educational contexts to analyse how curricular domain knowledge is realized differently in different activities inside and outside of school. this raises the question of how students can be facilitated educationally in transforming their curricular knowledge perspective across such contexts. and it invites further research questions concerning the development and testing of learning designs aimed at this facilitation. the viability of addressing such research questions from within a situative approach can be illustrated with a couple of recent projects investigating different ‘learning designs for transfer’. in one such learning design, student teachers in a first language course were asked in groups to teach modules in a local primary school (hachmann, 2020). the modules were on subjects which were part of the student teachers’ own curriculum, for example, dramatic structure. the design idea thus was to require the student teachers to transfer curricular knowledge between the teacher training program and the primary school classroom. the curricular knowledge to be transferred was at the domain level, since the curricula of the primary school class and the teacher training program overlapped thematically, though of course at different levels of academic sophistication. when the student teachers employed specific participation formats from their own course (for example, class-based analysis of a text) in the primary school class, curricular knowledge at the activity level was also involved. knowledge transformation between the two contexts was evident in the adaption of academic terminology, use of visualizations and choice of examples (level 1). similarly, it showed up in the way the class-based analysis was structured (level 2) to take into account classroom culture and the need for classroom management in the primary school class (level 3). research questions here concerned how to design for student teachers’ knowledge transformation between the two contexts of the teacher training program and the primary school; characterization of the resulting knowledge transformation; and the identification of factors that facilitated or hindered it. a particular focus was on how the student teachers were supported in transforming knowledge between the divergent contexts by acts of explicit framing (performed by the educator at the teaching training program) (engle, 2006; engle et al., 2012) of the relevance for the primary school class of concepts (level 1) and activity forms (level 2). the research questions were investigated through a design-based research methodology (amiel & reeves, 2008), utilizing ethnographic observations, video recordings of group work, interviews, and analyses of student-teacher assignments. the project resulted in design principles for ‘learning to transfer’ curricular knowledge, applicable beyond this very specific form of in-course school-practice coupling. bonderup dohn 25 | flr a second project (hansen, j.j. & dohn, 2018, 2020) investigated a learning design of ‘simulated social practices’. this design combined role-play within a university classroom context with projects anchored in student engagements in out-of-school workplace practices. the learning design thus coupled between contexts in which the students participated: on the one hand, a university course in which they all participated, and on the other hand, several workplace contexts in which they participated individually. the aim was to facilitate students in learning to transfer knowledge back and forth between these contexts. more specifically, students had a series of tasks (portfolio and in-class role-play) centred around developing and testing a product intended for (though in most cases not presented to) a specific workplace. the aim was to support the students in transforming their curricular knowledge perspective into workplace knowledge as well as, reciprocally, in drawing on their workplace knowledge perspective to elucidate classroom discussions. within the life-setting of the course (level 3), the workplaces had the role of cases-treated-within-education. therefore, the situational characteristics of the workplace life-setting were subsumed and transformed into situational characteristics at the activity level (level 2) within the course life-setting (level 3). the projects, portfolio and in-class role-play tasks likewise were activities at level 2. in consequence, as compared to what would have been expected in the workplaces, there were increased expectations of academic reflections about the projects and correspondingly less focus on workplace issues such as profit or growth. this was reflected in the students’ assignments where all students made use of curricular domain knowledge (as is to be expected within the context of a course), but only some of the students managed to integrate their workplace knowledge perspective in their case analysis and product design. the status of the project as activitywithin-a-course was also evident in the way role-play in class progressed: roles had to be articulated at the outset; they were taken up in a playful ‘as-if’ manner; particularly distinctive in-role comments by the students prompted laughter (because in-role comments were out-of-character at level 3); and out-ofrole comments addressing the course framing of the activity were recurrent. research questions in this project concerned how students transform and integrate their curricular and workplace knowledge in response to the differing unities of situational characteristics in workplace and course contexts. they also concerned the transformations involved within the course context between the different portfolio and role-play tasks. the research questions were investigated with a design-based research approach, utilizing ethnographic observation, analysis of student assignments, and qualitative questionnaires. a third project focused on facilitating children’s transition from a) day-care to b) a 4-monthstransition-module to c) school, through engaging them in the production of digital artefacts (e.g., photo books) and dialogue about these artefacts (odgaard, 2018, 2019). the basic design idea was that the digital artefacts could be negotiated as boundary objects (wenger, 1998) which connected the contexts of day-care, transition module, and school and facilitated the children’s engagement with their new contexts. the project investigated how children may be facilitated in transforming their experience from daycare to school. transfer was supported through initiating production and dialogue (level 2), focused on topics (level 1) central to the children’s experiences – or presumed by the day-care practitioners to be so. the study illustrated, on the one hand, that activities with the digital artefacts can indeed support children in transforming their experience through highlighting similarities and differences between the contexts. on the other hand, however, it also showed that practitioners and children alike sometimes engage in the production and dialogue of the digital artefacts as somewhat forced activities. the problem is that the societal structure-level (level 4) poses the requirement of ‘establishing continuity for children’ on the life-settings of day-care, transition module, and school (level 3). this then again frames the transition activities (level 2) as obligatory. these examples illustrate the usefulness of the framework for research into learning to transfer, both as concerns the development of design principles and as regards the subsequent analysis of transformation processes involved in transfer across contexts. bonderup dohn 26 | flr 6. in conclusion the aim of this article is to articulate a theory on the ontology of knowledge transfer which makes it possible to conceptualize knowledge transfer across contexts within a situative approach. i have presented an analytical framework of context levels, at which interacting situational characteristics are constituted and dynamically negotiated. i have argued that knowledge takes on concrete form and meaning from the specific situation in response to its unity of situational characteristics. therefore, the transfer of knowledge between contexts requires knowledge to be transformed in attunement to the unity of situational characteristics of the new context. my wider goal in presenting this argument has been to put research into learning to transfer across divergent life contexts (back) on the research agenda: given the framework of context levels, it is possible to some extent to analyse situational characteristics of new situations in advance, though their dynamic nature must be emphasized. it is also possible to point out non-exclusive ways in which transformation can be undertaken to attune to these situational characteristics. in this way, it is possible from within education to support students in achieving transfer and potentially also in developing dispositions for flexibly attuning to different unities of situational characteristics. by way of conclusion, i wish to emphasize two things: firstly, the usefulness of the framework of context levels for researching learning to transfer does not hinge on the specific view of knowledge presented here. other situative conceptualizations could make similar use of the framework. thus, given barab and roth’s (2006) understanding, one could analyse how different situations require flexibly transformed enlisting of a person’s “effectivity sets”. likewise, the framework will provide points of focus for analysing the “encompassing process” which wenger claims that participation is, and with it “knowing” (wenger, 1998, p. 4)6. this, in turn, will enable consideration of how knowing (as participation) may transform across contexts in response to the changed unities of situational characteristics of different “practices of social communities”. secondly, the framework will be useful, not only for research into learning to transfer, but also for designing and teaching for it. thus, a detailed analysis of interacting situational characteristics will help teachers articulate to students how their curricular knowledge perspective must be transformed to be utilized in workplace contexts. tasks which involve such transformation may be developed, as in the first two projects above, facilitating students in actually achieving transfer, not only in becoming aware of its challenges. within a somewhat different area, the framework may help teachers and students alike become aware of differences in situational characteristics between the contexts of school and of students’ out-of-school informal digital practices (greenhow, robelia, & hughes, 2009; knobel & kalman, 2016; lankshear & knobel, 2011). hereby, students may be facilitated in utilizing their familiarity with literacy practices on social media in ways which align with school expectations. teachers may also be supported in designing tasks which help students accomplish the necessary transformation by explicitly highlighting differences in situational characteristics between the contexts of informal practices and of school. potentially, it might even stimulate educational designers to question situational characteristics of the educational system not aligned with those of informal practices, such as the individualist focus of assessment practices. similarly, as indicated, analysis of variances in situational characteristics between school systems in different countries (level 4: societal structure-level) may facilitate international students in transforming their curricular knowledge perspectives in attunement to the unity of situational characteristics of their country of study. correspondingly, it may support their teachers in realizing systematic differences in their students’ concretized knowledge perspectives as compared to their home students, and consequently in posing tasks which help them perform the transformation. 6 “knowing is a matter of participating in the pursuit of [valued] enterprises” where “participation ... refers... to a[n] ... encompassing process of being active participants in the practices of social communities and constructing identities in relation to these communities” (wenger, 1988, p. 4, emphasis in original). bonderup dohn 27 | flr for all these suggested practical uses of the framework, complementing research questions may of course be asked, in collaboration with practitioners through participatory or design-based methods, or in case studies of existing learning designs. keypoints a new ontology of knowledge transfer is articulated. the question “what happens to knowledge in transfer across major context shifts” is shown to make sense also from within a situative approach. situated knowledge is conceptualized as attunement to a unity of situational characteristics, analysable with a five-level context framework. transfer is conceptualized as knowledge transformation in attunement to the new situation’s differing unity of situational characteristics. the article puts the issue of learning to transfer knowledge across divergent contexts (back) on the research agenda. acknowledgments i thank two anonymous reviewers for their insightful comments which helped me strengthen my arguments. research for the article has been supported by the independent research fund denmark, grant no. dff – 4180-00062. references akkerman, s. f., & bakker, a. (2011). boundary crossing and boundary objects. review of educational research, 81(2), 132-169. doi:10.3102/0034654311404435 amiel, t., & reeves, t. c. (2008). design-based research and educational technology: rethinking technology and the research agenda. educational technology & society, 11(4), 29-40. barab, s. a., & roth, w.-m. (2006). curriculum-based ecosystems: supporting knowing from an ecological perspective. educational researcher, 35(5), 3-13. doi:10.3102/0013189x035005003 bereiter, c. (2002). education and mind in the knowledge age. mahwah: l. erlbaum associates. bhaskar, r. (1975). a realist theory of science. new york: routledge. carraher, d., & schliemann, a. (2002). the transfer dilemma. journal of the learning sciences, 11(1), 1-24. doi:10.1207/s15327809jls1101_1 carroll, j., & ryan, j. (2005). teaching international students: improving learning for all. abingdon: routledge. de la rocha, o. (1985). the reorganization of arithmetic practice in the kitchen. anthropology & education quarterly, 16(3), 193-198. dohn, n. b. (2007). knowledge and skills for pisa—assessing the assessment. journal of philosophy of education, 41(1), 1-16. doi:10.1111/j.1467-9752.2007.00542.x dohn, n. b. (2009). affordances revisited: articulating a merleau-pontian view. international journal of computer-supported collaborative learning, 4(2), 151-170. doi:10.1007/s11412-009-9062z dohn, n. b. (2009). web 2.0-mediated competence – implicit educational demands on learners. electronic journal of e-learning 7(1), 111-118. bonderup dohn 28 | flr dohn, n. b. (2017). epistemological concerns querying the learning field from a philosophical point of view. (professorial thesis (dr.phil.)). university of southern denmark, retrieved from http://dohn.sdu.dk/ dohn, n. b., & hachmann, r. (2020). knowledge transformation across changes in situational demands between education and professional practice. in n. b. dohn, s. b. hansen, & j. j. hansen (eds.), designing for situated knowledge transformation (pp. 249-266). abingdon: routledge. dohn, n. b., & hansen, s. b. (2020). context framework for analyzing situated knowledge transformation. in n. b. dohn, s. b. hansen, & j. j. hansen (eds.), designing for situated knowledge transformation (pp. 59-74). abingdon: routledge. dreyfus, h. l. (1979). what computers still can't do. new york: harper & row. dreyfus, h. l., & dreyfus, s. e. (1986). mind over machine. the power of human intuition and expertise in the era of the computer. new york: free press. engeström, y. (2001). expansive learning at work: toward an activity theoretical reconceptualization. journal of education and work, 14(1), 133-156. doi:10.1080/13639080020028747 engle, r. a. (2006). framing interactions to foster generative learning: a situative explanation of transfer in a community of learners classroom. journal of the learning sciences, 15(4), 451498. doi:10.1207/s15327809jls1504_2 engle, r. a., lam, d. p., meyer, x. s., & nix, s. e. (2012). how does expansive framing promote transfer? several proposed explanations and a research agenda for investigating them. educational psychologist, 47(3), 215-231. doi:10.1080/00461520.2012.695678 gallagher, s. (2005). how the body shapes the mind. oxford: clarendon press. gibson, j. j. (1986). the ecological approach to visual perception. hillsdale: lawrence erlbaum associates. gick, m. l., & holyoak, k. j. (1983). schema induction and analogical transfer. cognitive psychology, 15(1), 1-38. doi:10.1016/0010-0285(83)90002-6 greenhow, c., robelia, b., & hughes, j. e. (2009). learning, teaching, and scholarship in a digital age: web 2.0 and classroom research: what path should we take now? educational researcher, 38(4), 246-259. doi:10.3102/0013189x09336671 greeno, j. g. (1994). gibson's affordances. psychological review, 101(2), 336-342. doi:10.1037//0033295x.101.2.336 greeno, j. g. (1997). on claims that answer the wrong questions. educational researcher, 26(1), 5-17. doi:10.3102/0013189x026001005 greeno, j. g. (2011). a situative perspective on cognition and learning in interaction. in t. koschmann (ed.), theories of learning and studies of instructional practice (vol. 1, pp. 41-71). new york: springer. greeno, j. g., & gresalfi, m. s. (2008). opportunities to learn in practice and identity. in p. a. moss, d. c. pullin, j. p. gee, e. h. haertel, & l. j. young (eds.), assessment, equity, and opportunity to learn (pp. 170-199). new york: cambridge university press. greeno, j. g., & tmsmtapg (1998). the situativity of knowing, learning, and research. american psychologist, 53(1), 5-26. doi:10.1037/0003-066x.53.1.5 hachmann, r., & dohn, n. b. (2018). participatory skills for learning in a networked world. in n. b. dohn (ed.), designing for learning in a networked world (pp. 102-119). abingdon: routledge. hachmann, r. (2020). didaktisk design for transformationer af faglig viden: en undersøgelse af lærerstuderendes videnstransformationer på tværs af professionsuddannelse og -praksis. (phd dissertation). university of southern denmark, kolding. hansen, j. j., & dohn, n. b. (2018). design principles for learning in simulated social practices. in n. b. dohn (ed.), designing for learning in a networked world (pp. 214-231). abingdon: routledge. hansen, j. j., & dohn, n. b. (2020). designing for mediational transition and learning through simulation: a task analysis of knowledge transformation. in n. b. dohn, s. b. hansen, & j. j. hansen (eds.), designing for situated knowledge transformation (pp. 267-283). abingdon: routledge. hutchins, e. (1995). cognition in the wild. cambridge, massachusetts: mit press. bonderup dohn 29 | flr hutchins, e., & klausen, t. (1996). distributed cognition in an airline cockpit. in y. engeström & d. middleton (eds.), cognition and communication at work (pp. 15-34). new york: cambridge university press. jarvis, p. (2007). globalization, lifelong learning and the learning society: sociological perspectives. london: routledge. judd, c. (1908). the relation of special training to general intelligence. educational review, 36, 28-42. kandiko, c. b., & weyers, m. (eds.). (2013). the global student experience: an international and comparative analysis (1 ed.). abingdon: routledge. knobel, m., & kalman, j. (eds.). (2016). new literacies and teacher learning: professional development and the digital turn (vol. 74). new york: peter lang. kuhn, t. s. (1970). the structure of scientific revolutions (2 ed.). chicago: the university of chicago press. lankshear, c., & knobel, m. (2011). new literacies: everyday practices and social learning. maidenhead, great britain: open university press. lave, j. (1988). cognition in practice mind, mathematics and culture in everyday life. cambridge: cambridge university press. lave, j., & wenger, e. (1991). situated learning legitimate peripheral participation. new york: cambridge university press. lobato, j. (2012). the actor-oriented transfer perspective and its contributions to educational research and practice. educational psychologist, 47(3), 232-247. doi:10.1080/00461520.2012.693353 mayer, r. e. (2001). multimedia learning. cambridge: cambridge university press. mayer, r. e., & massa, l. j. (2003). three facets of visual and verbal learners: cognitive ability, cognitive style, and learning preference. journal of educational psychology, 95(4), 833-846. doi:10.1037/0022-0663.95.4.833 mcdermott, r. (1993). the acquisition of a child by a learning disability. in j. lave & s. chaiklin (eds.), understanding practice: perspectives on activity and context (pp. 269-305). new york: cambridge university press. merleau-ponty, m. (1962). phenomenology of perception. london: routledge and kegan, paul. molander, b. (1992). tacit knowledge and silenced knowledge: fundamental problems and controversies. in b. göranzon & m. florin (eds.), skill and education: reflection and experience (pp. 9-31). london: springer verlag. nielsen, k., & kvale, s. (eds.). (2003). praktikkens læringslandskab : at lære gennem arbejde. københavn: akademisk forlag. nokes-malach, t. j., & mestre, j. p. (2013). toward a model of transfer as sense-making. educational psychologist, 48(3), 184-207. doi:10.1080/00461520.2013.807556 norman, d. a. (1988/2002). the design of everyday things. new york: basic books. norman, d. a. (1999). affordance, conventions, and design. interactions, 6(3), 38-43. doi:10.1145/301153.301168 odgaard, a. b. (2018). networked learning in children's transition from day-care to school: connections between contexts. paper presented at the 11th international conference on networked learning, zagreb, croatia. odgaard, a. b. (2019). teknologimedierede aktiviteter i børns overgang fra dagtilbud til skole et sociokulturelt informeret og design-baseret studie. (phd dissertation). university of southern denmark, kolding. packer, m. j. (2001). the problem of transfer, and the sociocultural critique of schooling. the journal of the learning sciences, 10(4), 493-514. doi:10.1207/s15327809jls1004new_4 polanyi, m. (1962). personal knowledge: towards a post-critical philosophy (vol. 1158). chicago: university of chicago press. popper, k. r. (1972). objective knowledge: an evolutionary approach. clarendon press. reed, s. k. (1993). a schema-based theory of transfer. in d. k. detterman & r. j. sternberg (eds.), transfer on trial: intelligence, cognition, and instruction (pp. 39-67). norwood, new jersey: ablex. bonderup dohn 30 | flr reed, s. k. (2012). learning by mapping across situations. journal of the learning sciences, 21(3), 353-398. doi:10.1080/10508406.2011.607007 salomon, g. (ed.) (1993). distributed cognitions: psychological and educational considerations. new york: cambridge university press. sanders, j. t. (1993). merleau-ponty, gibson, and the materiality of meaning. man and world, 26(3), 287-302. doi:10.1007/bf01273397 schatzki, t. r., knorr-cetina, k., & von savigny, e. (2001). the practice turn in contemporary theory. london: routledge. schön, d. a. (1983). the reflective practitioner: how professionals think in action (vol. 5126). new york: basic books. sfard, a. (1998). on two metaphors for learning and the dangers of choosing just one. educational researcher, 27(2), 4-13. doi:10.2307/1176193 singley, m. k., & anderson, j. r. f. (1989). the transfer of cognitive skill. cambridge, massachussets: harvard university press. säljö, r. (2000). lärande i praktiken: ett sociokulturellt perspektiv. stockholm: prisma. säljö, r., & wyndhamn, j. (1993). solving everyday problems in the formal setting: an empirical study of the school as context for thought. in s. chaiklin & j. lave (eds.), understanding practice: perspectives on activity and context (pp. 327-342). new york: cambridge university press. thrysøe, l. (2011). at blive og at være sygeplejerske: en undersøgelse af oplevelsen ved at være næsten færdiguddannet og nyuddannet sygeplejerske og interaktionens betydning for deltagelse i praksisfællesskabet. ph.d. dissertation. odense: research unit of nursing, university of southern denmark. tuomi-gröhn, t., & engeström, y. (eds.). (2003). between school and work: new perspectives on transfer and boundary-crossing. oxford: elsevier science ltd. van oers, b. (1998a). the fallacy of decontextualization. mind, culture and activity, 5(2), 135-142. doi:10.1207/s15327884mca0502_7 van oers, b. (1998b). from context to contextualizing. learning and instruction, 8(6), 473-488. doi:10.1016/s0959-4752(98)00031-0 wagner, j. f. (2010). a transfer-in-pieces consideration of the perception of structure in the transfer of learning. journal of the learning sciences, 19(4), 443-479. doi:10.1080/10508406.2010.505138 wedege, t. (1999). to know or not to know – mathematics, that is a question of context. educational studies in mathematics, 39(1-3), 205-227. wenger, e. (1998). communities of practice. new york: cambridge university press. wittgenstein, l. (1958). philosophical investigations (g. e. t. anscombe ed.). oxford: blackwell. wittgenstein, l. (1969/1979). on certainty (r. rhees & g. h. von wright eds.). oxford: blackwell. yackel, e., & cobb, p. (1996). sociomathematical norms, argumentation, and autonomy in mathematics. journal for research in mathematics education, 27(4), 458-477. doi:10.2307/749877 codepen dalsgaard publication frontline learning research vol.8 no.1 (2020) 1 14 issn 2295-3159 reflective mediation: toward a sociocultural conception of situated reflection christian dalsgaarda aaarhus university, denmark article received 15 january 2019 / revised 25 september / accepted 22 january / available online 7 february 2020 abstract the objective of the article is to contribute to the development of a sociocultural conception of situated reflection that can be used in empirical studies of reflection, and that can be utilised in development of educational practices. based on a development of the concept of 'reflective mediation', a conception of reflection is developed from a situated understanding of learning processes. taking a situated approach, the concept of reflective mediation describes how to understand reflection as an integral part of the immediate activities of the individual. a theoretical framework is developed for empirical studies on reflective activities of higher education students. the framework can be utilised by teachers to develop teaching methods in support of reflection in student learning. the concept of reflective mediation is developed from a combination of pragmatism and cultural historical activity theory, and it covers seven categories of learning and reflection processes. the article makes a distinction between three forms of mediation, two forms of empirical reflection, and two forms of theoretical reflection. the article concludes in a discussion of the implications of the theoretical framework for educational research and for teaching practices within higher education. the article is frontline in the sense that it aims at developing a theoretical conception of situated reflection by combining cultural historical activity theory with pragmatism and theories of situated learning. the novelty of the article is to consider levels of human activity as levels of reflection and to introduce the distinction between theoretical and empirical reflection. further, the article provides arguments that awareness of objects and instruments of human activity forms the basis of reflective processes. finally, the article explains how reflection can connect levels of human activity and learning. keywords: reflection; mediation; situated learning; cultural historical activity theory; pragmatism info corresponding author cdalsgaard@tdm.au.dk doi: 10.14786/flr.v8i1.447 1. introduction the article addresses the question: how can we conceptualise processes of reflection? the objective of the article is to contribute to a theoretical understanding of reflection that can be used in empirical studies of reflective activities, and that can be utilised for development of educational practices, specifically within higher education. the aim is to develop a framework that can aid teachers in facilitating, identifying and evaluating activities of reflection of higher education students. more specifically, the article develops a concept of reflection that combines theories of situated learning with pragmatism and cultural historical activity theory. starting from a situated learning approach, processes of reflection are positioned in direct relation to the immediate activities of the learner. most notably, schön's (1983; 1987) concepts of knowing-in-action and reflection-in-action situate learning, knowledge and reflection in practice and connects them directly to ongoing activities. also, suchman's (1987; 1996) concept of situated actions connects knowledge to the very processes of performing activities. similar views are found in lave & wenger (1991), wenger (1998), billet (1996; 2001) and semin & smith (2013) who connect learning to social practices, and also argue that knowledge is bound up in the context and processes of practices. within a frame of these theories, this article raises the question of how to understand reflective processes situated in practice. the aim of the article is to further develop this position that examines reflection as an active part of an ongoing practice. schön's (1987) concept of reflection-in-action forms a basis for the theoretical development of a concept of reflection. however, schön (ibid) does not clarify how reflection takes place, and what the prerequisites are for reflection. how does reflection connect to the ongoing activities of a student, and what makes an activity reflective? in order to address these questions and further develop a situated concept of reflection, the article will combine the pragmatism of dewey's (1916; 1958) concept of reflective experience and schön's (1987) conception of reflection-in-action with an activity theoretical approach to learning activities (leontev 1978; engeström 2014). 2. the situated nature of learning within the research literature on situated learning and situated cognition, different dimensions of the situated nature of learning are described (roth & jornet 2013; compton 2013). at least three different dimensions can be identified within the literature. firstly, learning is situated within sociocultural contexts, which emphasises that learning is rooted in practices that have a cultural and historical origin (wertsch 1998; engeström 2014; 1999; engeström & sannino 2010; 2012; leontev 1978). secondly, a body of literature situates learning within local settings of social practices (lave & wenger 1991; wenger 1998; hutchins 1995;1996; billet 2001; semin & smith 2013). thirdly, learning is also conceived as situated within immediate activities meaning that knowledge exists within the very processes of acting (schön 1987; suchman 1987; fors, bäckström & pink 2013). of the first conception of the nature of situated learning, wertsch (1998; 109) writes that "[...] virtually all human action, be it on the individual or social interactional plane, is socioculturally situated [...].". leontev (1978) and engeström (2014; 1999) position activities of the individual within a collective activity. engeström (1999) brings forth the argument that human activity is culturally and historically situated, which means that learning is a local instance that inherits from practices situated in a culture and with a historical development. the understanding that learning is situated in local settings of social practices has most notably been advocated by lave & wenger (1991) and wenger (1998) who connect learning to communities of practice. they provide a different dimension of the situated nature of learning than engeström (2014; 1999) and leontev (1978), because lave & wenger (1991) situate learning within local practices as they unfold and develop. a study in billet (2001) also shows that expertise of practitioners is bound up in the activities unfolding in local practices. both wenger (1998) and billet (2001) argue that knowledge is developed within these practices and cannot be extracted from the setting or processes of the practices. also, this dimension highlights the socially situated nature of learning; learning stems from social relations and negotiations between individuals within a practice (semin & smith 2013; wenger 1998). finally, the view that learning is situated in immediate processes of acting has mainly been influenced by dewey's (1916; 1958) concept of experience. schön (1983; 1987) draws on the approach of dewey in developing the concepts of knowing-in-action and reflection-in-action. these concepts directly situate knowledge and reflection in practice and connect them to ongoing activities. further, suchman's (1987; 1996) concept of situated actions connects knowledge to the very processes of performing activities. a more recent concept of sensory-emplaced learning by fors, bäckström & pink (2013) continue this line of thinking and expands it by emphasising the embodied aspect of situated learning. while theories of situated learning are well-developed, a concept of situated reflection lacks theoretical foundation. the aim of this article is to develop a conception of reflection that encompasses the different dimensions of situated learning. the common denominator for the different dimensions is the close connection between human activity and learning. 3. situating learning within levels of human activity cultural historical activity theory (chat) emphasises human activity as the central aspect of knowing and learning (vygotsky 1978; leontev 1978; leontyev 1981). the fundamental premise of chat is that the higher psychological functions of the mind have a sociocultural origin (vygotsky 1978; wertsch 1994). leontev’s distinction between levels of human activity and learning and engeström’s continuation of this distinction (leontev 1978; engeström 2014) will provide the theoretical foundation of the framework presented in this article. following leontev (1978), human activity fundamentally relates to the concepts of activities, objects and instruments, which describe the levels of human activity. human activity is object-oriented, meaning that activities are directed at objects (or objectives) in the world. leontev (ibid) argues that an activity is directed towards a motive, actions towards goals and operations towards conditions. for example, a human activity could be museum activity with a motive of preserving and disseminating cultural heritage or maintaining a national identity (human activity can have a composite motive). within this motive exists a range of goals of the individual staff members, including setting up exhibitions, cataloguing artefacts, and doing guided tours. finally, operations are the specific visible "movements" such as moving around artefacts, writing texts, labelling objects, etc. in order to achieve the object of the activity, humans mediate their activity with instruments. engeström (2014) completes the structure of activity with his distinction between levels of instruments (figure 1); motives are mediated by a methodology, goals with models and operations with tools. a methodology for a museum could be communication theory, information theory or educational theory, models include principles for organising and setting up exhibitions and systems for cataloguing, and tools could be exhibition cases, exhibited artefacts, posters, etc. figure 1. levels of human activity. inspired by bateson (1972), engeström (2014) makes a classification of three levels of learning. based on the theory of logical types, bateson (1972) defines five levels of learning; learning 0, i, ii, iii and iv, where learning 0 is a response to a situation, and learning iv would imply a combination of phylogenesis and ontogenesis (bateson 1972; p. 293). bateson defines learning on a higher level as a change in the process of learning at a lower level. lower level learning is a member of the higher level. for example, learning i is a member or an instance of the type learning ii, which classifies learning i. within learning 0, the individual adapts his/her behavior according to stimuli of a situation. learning i is a classification of the stimuli of a situation and entails a change in the rules for learning 0. learning ii is a classification of a pattern within the situation. this pattern can classify other situations of the same type, thus changing learning i. according to bateson, learning ii is learning to learn. engeström uses this classification from bateson to define the use of instruments for operations, actions and an activity, respectively, as different levels of learning; i.e. learning i, ii and iii. learning i involves a change in or development of physical operations – that is, movements – of an individual, and it involves use or development of concrete tools. learning ii is development of actions that can attain a goal of an individual. learning ii, then, involves development of models in the form of concepts, principles or ways of working. in engeström's conception of the levels of learning, learning ii must also result in learning i, since an action is not manifested, before it is performed through operations with tools. finally, learning iii is development of new forms of human activity, which include development of methodologies such as theories or policies for the activity. learning i is always based on a change in conditions; that is, the physical materials and surroundings. learning ii begins with a new goal for human action, and finally, learning iii can only take place, if a new motive for human activity emerges. an historical example of learning iii is the emergence of a motive for public child care that has resulted in day care institutions. the distinction between three dimensions (objects, instruments and activities) and three levels of human activity (motive, goal, conditions, etc.) provides the framework for developing a sociocultural framework for learning and reflection. according to engeström (2014), an activity or action is not developed until a given methodology or model has affected the development of new operations and tools. this means that learning iii requires development of new actions, operations and tools (i.e. learning i). learning ii and iii without learning i will be abstract or theoretical and have no impact on human activity. this article will extend this sociocultural understanding by developing a concept of reflective mediation that aims at describing the processes on each level and explain the relations between the levels. consequently, the framework splits engeström’s levels of learning into a range of different processes of learning and reflection. this means, for instance, that learning iii is accomplished through (sub)processes of mediation and reflection on all three levels. a central question in understanding the processes of learning within an activity theoretical conception is: what are the relations between the levels of activity illustrated in figure 1? and, more specifically, how does the individual make a connection between the levels? according to leontev (1978), the levels cannot be separated, since an activity is carried out through actions and operations, while operations serve actions which again serve an activity. in that sense, there is an interdependence between the three levels, which cannot exist in isolation. engeström (2014) also argues for this interdependence between the levels of human activity; he states that models are developed on the basis of a methodology, and tools are developed on the basis of a model. this means that an operation, for instance, is somehow related to an action. below, the article will present a way of understanding, how the levels of human activity are related. 4. awareness of objects and instruments as a basis for reflection using wartofsky's (1979) concept of model, it can be argued that the levels of human activity can be viewed as levels of awareness of the individual, and that this awareness is a prerequisite for reflection. wartofsky (ibid) defines models as cognitive artefacts that are "representations to ourselves of what we do, of what we want, and of what we hope for" (1979; xv). wartofsky (ibid) describes models as "modes of action" that both embody human purpose and instruments for carrying out purposes. this means that models are parallel to and encompass both objects and instruments of human activity (figure 1). following wartofsky, models can be viewed as the individual's awareness or representation of instruments, and wartofsky (1979; xviii) directly connects models and representational artefacts with human consciousness. the key in wartofsky's understanding of models in relation to reflection is that human consciousness uses models to present itself with its own objects. thus, the instruments and objects of human activity can be viewed as something that individuals can be consciously aware of. as leontev (1978) states, there is no obvious relationship between the levels of human activity. leontev (ibid) writes: “let us assume that the goal remains the same; conditions in which it is assigned, however, change. then it is specifically and only the operational content of the action that changes” (no page number). this means that on one hand, a sequence of operations constitutes an action, but on the other hand, an action is a form of generalisation of the different sequences of operations that are able to accomplish the goal of the action. in other words, although there is an interdependence between the levels of human activity, they also have an existence – at least for the individual – on their own. a goal, for instance, can have a conscious representation for the individual in itself without an awareness of the overall motive. consequently, there is no direct link between the three levels (in figure 1). the objects and instruments on each level represent independent units and concepts that the individual might be aware of. a methodology can be defined as a generalisation of all the models that can be used to carry out the activity. a model can be defined as a generalisation of all the tools that can be used to carry out the action. for example, a pedagogical model of collaborative learning can result in different specific pedagogical methods such as peer-feedback, group work, project work, etc. this is the argument for viewing the levels of human activity as different levels of awareness. it is possible to be aware of objects and instruments on all three levels of activity. for example, an individual can perform practical operations such as playing the keys on a piano. the individual can develop operational expertise in playing fast and putting the right pressure on the keys without being aware of a goal of the action which the operations serve, such as expressing oneself or creating enjoyment. an individual's awareness of objects and instruments on a higher level (models and methodology) makes it possible to consider a larger aspect of the activity and therefore to develop the activity more fundamentally. as engeström (2014) writes, a methodology can develop new forms of actions, and models can be used to develop new forms of operations. similarly, awareness of higher levels of human activity can result in a more fundamental understanding of the activity. for instance, the individual who plays the keys on a piano, uses the music to express feelings, and understands the role of music in society, understands music rooted in all three levels. awareness of a goal or motive means that the individual can understand tools and operations on a higher level. the highest level of learning – what engeström (2014) calls learning iii – can according to engeström (ibid) only be performed by a collective and not by the individual. however, the individual can be aware of motives and methodologies of an overall collective activity and use them to perform individual actions. as an example, a doctor is unable to perform all activities of a heart surgery, but he/she may be aware of the activities of the other doctors and nurses and understand the entire activity of the surgical procedure. 5. empirical and theoretical reflection wartofsky (1968) argues that the individual’s awareness of a conceptual construction of an artefact is what makes reflection possible: "[t]he possibility of reflective examination of the relations between ends and means arises only with the development of a conceptual representation of action" (ibid; 37-38). reflective examination of relations between ends and means is equivalent to reflection of relations between activities, objects and instruments. because there is not direct relation between the levels of human activity, objects and instruments on a lower level of human activity cannot be derived from objects and instruments on a higher level. on the other hand, an instrument on a higher level cannot prescribe concrete solutions on a lower level of activity (leontyev 1981). however, an individual can evaluate whether or not, for instance, a specific operation supports the goal of an action. the individual's awareness of a goal can be used to evaluate operations necessary to perform the action. based on this, in this article, reflection is defined as an evaluation of the consequences of activities in relation to an object or an instrument on a higher level of activity. this implies that there are two kinds of reflection. one form of reflection is evaluation in relation to an object on a higher level of activity, whereas the other form is evaluation in relation to an instrument on a higher level of activity. i term the first kind empirical reflection and the second kind theoretical reflection. this distinction follows vygotsky’s (1986) distinction between spontaneous, everyday concepts and scientific concepts, and davydov’s (1999) development of vygotsky’s concepts in his distinction between empirical and theoretical thinking. dewey (1997) makes a similar distinction between empirical and scientific thinking. according to both davydov (1999) and dewey (1997) empirical thinking is based on common sense and concrete observations of the given situation. this corresponds to the individual’s awareness of objects, which relate to the practical and empirical situation of the individual. on the other hand, according to dewey (1997) theoretical thinking involves an analysis of the situation. this corresponds to an awareness of instruments, which enable the individual to think beyond the conditions of the concrete, empirical situation. empirical reflection is based on the individual’s awareness of objects of human activity. an object constitutes the underlying purpose or objective of human activity, and it forms the basis of the very existence of human activity (leontev 1978). but for the individual, the awareness of an object can be a conception that guides his/her activities. billet (2001) uses the term goal-directed and dewey (1916; 1997) uses the term end-in-view to describe the directed nature of human activity. in the words of dewey (1916), activities have an aim or purpose, and to act with an end-in-view means that humans have a conception or awareness of the result of the actions. this imagined idea of the outcome makes empirical reflection possible. when humans perform an operation, they can reflect upon the consequences of the operation in relation to the object of the action. theoretical reflection is based on the individual’s awareness of instruments of human activity. according to wartofsky (1979) an instrument is at the same time a concrete thing and a conceptual construction which describes a generalisation of future activities. this explains how to understand the relationship between practical, empirical knowledge and conceptual, theoretical knowledge. theoretical concepts are invested with meaning from human activity, because they are used for purposeful, object-oriented activity. this means that theoretical knowledge is not detached from human activity, but is rather relative to the objects of the individual. as brown, collins & duguid (1989) argue, concepts and theories are similar to tools in the sense that they are acquired through use. this means that a concrete tool employed to mediate physical conditions has a theoretical dimension in the same way that a concept is utilised to mediate a goal. in other words, using a chair for "resting" has a similarity with using a constructivist model to "plan a university course", or using concepts from rhetoric to "analyse a speech". the theoretical dimension of a tool is not confined to the given physical thing used as a tool, because other things could potentially be used as the same tool to reach the same object – for instance, using a stump of tree for resting, using an action learning approach in course planning, or applying a narrative approach to speech analysis. thus, a physical thing is not conceptualised by the individual as a specific and concrete entity, but instead as a tool with a general “use”; for instance, "use of a chair for resting" is a theoretical dimension of a tool. further, one and the same physical thing can potentially be used as different tools. the chair can be used for sitting or to stand on when changing a light bulb, thus holding different theoretical dimensions. the levels of awareness as described above take the individual as a point of departure. however, it is important not to dismiss the social nature of human activity (semin & smith 2013). ultimately, leontev (1979) and wartofsky (1979) argue that human activity is collective and involves social interaction. the individual’s awareness of the different levels does not explain the role of social interaction for learning. when the individual is directed at a motive, learning also entails the awareness of the actions of other individuals. the key to the social dimension underlying all learning is the motive behind human activity. a motive stands above the individual in the sense that it constitutes a motive for the collective activity of humans (leontev 1978). understanding a motive requires the individual’s understanding of his/her relations to other individuals performing actions within the activity. consequently, reflection in relation to a motive or a methodology involves reflection on one’s own actions in relation to actions of other individuals. 6. processes of reflective mediation the concept of object-oriented mediated activities is key to understand learning processes within chat, and the concept has been treated thoroughly by many authors (vygotsky 1978; leontev 1978; engeström, 2014; billet 1996). however, the very processes involved in mediated activities are not unfolded within chat. engeström (2009) describes sequences of learning actions on a long term, but not situated, immediate processes. in order to explain the processes of learning and reflection, the article will draw on social theories of learning that argue for the importance of studying and understanding processes of human activities to understand learning (suchman 1987;1996; schön 1987; lave & wenger 1991; lave 1996; chaiklin & lave 1996; salomon 1993). with the concept of situated action, suchman (1987; 1996) argues that human activity is not based on static concepts and preconceived plans that are applied to different situations, but that human activity adapts to the specific circumstances, and unfolds in the very course of action. lave & wenger (1991; lave 1996) have continued this line of thinking in their concept of situated learning, which emphasises that “learning is an integral aspect of activity in and with the world at all times” (lave 1996; 8). similar to this, salomon & perkins (1998) describe learning as a "highly situated activity of participation". schön (1987) relates the situated nature of human activity to knowledge and uses the term knowing-in-action to describe that knowledge relates to actions and cannot be separated from them. knowledge should be understood as knowing, which is a process related to situated activities. related to the levels of human activity (figure 1) this means that instruments in relation to objects are not static conceptions but exist within the very processes of mediation. in order to conceptualise processes of learning and reflection, the article draws on dewey’s (1916; 1958) concept of experience, which describes processes involved in mediated activities. the conceptualisation is also inspired by engeström and sannino (2012) who describe seven actions of expansive learning: questioning, analysing, modelling, examining the developed model, implementing the model, reflecting and consolidating (a new form of practice). dewey (1916; 1958) uses the term reflective experience to describe learning from experience. he divides reflective experience into a number of processes: (i) perplexity, confusion or doubt, (ii) a conjectural anticipation, (iii) careful survey or examination, (iv) elaboration of the tentative hypothesis, and (v) a plan of action and doing something (dewey 1916; 150). these processes can be used to supplement activity theory. together, dewey's (i) and (ii) can be translated into the conscious formation of an object that directs the activities of the individual. dewey's (iii), (iv) and (v) can be used to clarify sub processes of mediation. in the third process, dewey highlights an initial process of examining the opportunities of the situation, including examination of tools that might help reach the object of the activities. according to dewey, the result of this examination is a construction of a hypothesis or an idea for doing (iv). finally, (v) is the actual doing, the performing of activities. from the concept of object-oriented mediated activity, reflection as defined above, and dewey’s reflective experience, it is possible to describe the following elements or processes of learning: 1) object, 2) examination, 3) construction, 4) activities, 5) judgment, 6) reflection. figure 2 illustrates the processes of learning. figure 2. process of learning (and reflection). the basis for learning is an object or aim of the individual’s activity. an object can be on either of the three levels of human activity. this corresponds with dewey's (i) and (ii). further, following dewey (iii), the individual will conduct processes of examination of instruments in order to construct an idea or what dewey (iv) terms an hypothesis of which instruments to use to reach the object. then, the individual acts upon the hypothesis (v), meaning that he/she mediates the object by performing actions. to these processes described by dewey, this article will add judging of an action in relation to an object (on the same level of activity). judging is seen as part of learning, but not reflection. to learn, the individual must know, if the action was "successful" (i.e. mediated the object). thus, mediation requires an awareness of the object and a judgment of consequences of an action in relation to this object. otherwise it is blind action, and learning is not possible. if the instrument does not mediate the object, the individual can repeat the processes and mediate again using other tools or using tools differently. this is a learning process and is not considered reflection in this article. together, elements 1 through 5 describe processes of mediation of objects, i.e. a learning process. this article will add the process of reflection to the understanding of learning processes. the individual can – potentially – reflect on the consequences of the activities in relation to an object or instrument on a higher level of awareness. the fifth process is judgment in relation to the object on the same level of activity, whereas reflection is defined as evaluation on a higher level of awareness. in other words, judgment relates to the very immediate activities that the individual is involved in. for example, within an activity of film production, an editor could continuously judge the flow and precision of the scenes. reflection would involve evaluation of the results in relation to a specific genre, a desired atmosphere, etc. following activity theory, an action can have multiple goals (and similarly an activity can have more than one motive). if, for instance, an action has more than one goal, the individual can then potentially reflect on the different goals of the action; multiple goals would mean that reflection should also be multiple. taken together, the processes of learning and reflection form what i will term reflective mediation, which aims at describing both situated learning and reflection. the concept of reflective mediation will be developed in this article to describe processes of learning on each of the three levels of human activity, and processes of learning that cross the levels (cf. figure 1). the latter entails processes of reflection. the processes of reflective mediation can be described within three fundamental forms of learning: 1. mediation of objects with instruments, 2. empirical reflection on consequences in relation to objects on a higher level of activity, and 3. theoretical reflection on consequences in relation to instruments on a higher level of activity each of the three forms can exist on the different levels of human activity. they are unfolded below. h3> 7. a framework for learning and reflection based on the above definition of processes of learning and reflection, it is now possible to develop a theoretical framework for learning and reflection. from the elements of learning (illustrated in figure 2), mediation is defined as use and judgment of an instrument in relation to an object of an activity. as illustrated in figure 3, mediation can exist on three levels; methodologies are used to mediate a motive, models to mediate goals, and tools to mediate conditions. however, methodologies and models are only theoretical mediations or ”thought experiments”, since they do not influence practice. such mediations would typically exist within higher education in the form of project reports or written assignments. to change practice, however, the instruments must be manifested in tools and operations. figure 3. mediation exists on three levels: 1) mediated operation (bottom), 2) mediated action (middle) and 3) mediated activity (top). mediation results in an understanding of a given instrument in relation to its object. this constitutes learning in its own respect. an example within higher education could be students' utilisation of an analytical method (a model) such as conversational analysis to interpret (action) a dialogue. however, adding reflective processes to mediation would result in a deeper understanding of the given instrument. reflection adds to the understanding of the mediation process by connecting it to another level, and thus widens the understanding. reflective mediation describes the process in which learning occurs with the awareness of a higher level of activity. the distinction above between empirical and theoretical reflection implies that there are two forms of reflective mediation processes. empirical reflective mediation is reflection on consequences in relation to an object on a higher level of activity (figure 4), whereas theoretical reflective mediation is reflection on consequences in relation to an instrument on a higher level of activity (figure 5). figure 4. empirical reflective mediation exists on two levels: 1) empirical reflective operation (bottom) and 2) empirical reflective action (top). figure 5. theoretical reflective mediation exists on two levels: 1) theoretical reflective operation (bottom) and 2) theoretical reflective action (top). based on the distinction between mediation, empirical reflective mediation and theoretical reflective mediation, the developed sociocultural framework contains seven different categories of learning and reflection. below, the seven categories are related to engeström's three levels of learning: learning i mediated operation empirical reflective operation theoretical reflective operation learning ii mediated action empirical reflective action theoretical reflective action learning iii mediated activity each of these seven categories can be characterised as processes of learning in their own respect. mediation and reflective mediation in themselves imply learning, and each category of learning carries an independent understanding of a certain aspect of the activity. ideally, however, learning involves all levels and categories of learning and reflection; meaning that the levels of human activity are connected for the individual. 8. implications for teaching methods this section will discuss how the theoretical framework can be used to support reflection within higher education teaching. based on the framework, an educational practice for reflection should aim at supporting the students' connections between all levels of an activity. there is no hierarchy in the levels that ranks the importance or value of each of the levels, and there is no progression from one level to another. thus, a central question in developing teaching methods is, which categories within the framework come first, and how teaching can be organised to support mediation and reflective mediation that covers the entire framework. since a motive is the foundation of human activity, this would be an obvious starting point. however, the activity mediating the motive is a collective activity of humans (leontev 1978), which means that an individual cannot perform an activity him/herself. in that sense, the motive and activity cannot form the starting point for the individual’s learning. as argued above, an individual can, however, be aware of the overall motive of the collective activity. this suggests that the individual as a starting point should be directed at a goal in order to perform an action. the argument for goals as the starting point for learning is that goals can connect to the life of the student. this is in line with the approach of dewey (1916), who argues that the training of skills should never be detached from a purpose. an object should present itself as a problem for the student in order to hold a potential for learning. the goal should take the form of what can be termed an empirical problem. at the same time, the individual should ideally be aware of an overall motive of the action. in that sense, although mediation and reflective mediation on each level of awareness constitute processes of learning in their own respect, learning activities should always be grounded in a motive and directed at a goal. further, learning should ideally point towards operations in practice. an example could be taken from teaching a university communication studies course. starting points for a course within communication studies could be a goal (empirical problem) of the students to either develop a communication strategy for a company, organise a political campaign, organise a public service information campaign, or create an advertisement for a commercial product. once engaged in an empirical problem with the awareness of a goal, the students can perform actions and operations. in the example from communication studies, the teacher should present students with instruments for achieving the objects. instruments for operations and actions of the students will differ depending on which of the outlined goals students are working on. for instance, whereas developing a communication strategy would involve models of branding and organisational image, creating an advertisement includes models for marketing and target groups. operations of students could involve producing a radio or tv commercial, which would involve writing manuscripts, recording, editing, etc. the empirical problem of creating an advertisement for a product can on the one hand provide the opportunity to perform concrete operations in practice. on the other hand, the empirical problem makes it possible to reflect in relation to motives and methodologies. a motive for developing an advertisement for a company might be creating attention and reaching a large audience, whereas the motive of an information campaign could be to communicate a certain message to a specific target group. both of these objects of student activities could result in production of tv ads, but if students reflect empirically on their ads in relation to the motives, the ads will be very different. teachers should support students in theoretical reflection related to their object-oriented activities. although the motives of the student activities are different, the methodologies can be the same in the form of communication theories. students might reflect on their communication strategy or advertisement in relation to theories from writers such as shannon & weaver, jakobson, bakhtin, etc. thus, the role of communication theory would be as an instrument of theoretical reflection within students' work on an empirical problem. since concepts and theories are dominant within academic contexts in higher education, the theoretical framework of reflective mediation has implications for educational practice of higher education. the implications of the presented framework would be not to teach concepts and theories isolated from an empirical problem. the framework calls for employing concepts and theories as instruments related to object-oriented activities. as a consequence, models and methodologies become secondary to the object of student activities, but on the other hand they are key to theoretical reflection. following the approach of the framework, reflection in educational practice is not a matter of students viewing their own work from a distance, but is rather a matter of embedding reflective processes in their situated activities. to support this as a teacher requires that students are made aware of concepts and theories that they can consciously employ for reflection. to put it in other words, students need tools and objects for reflection, and cannot be asked to just "reflect" on their work. a central point taken from the situated approach of this article is that although reflective processes are key in understanding a given subject area, they are very difficult to assess based on a final product or assignment – for instance, the final advertisement. it is not possible to track down the reflective processes of students. they must be identified within the processes. this calls for making both students and teachers aware of the nature and the importance of such situated processes of judgment and reflection. 9. conclusion the presented framework of seven categories of learning and reflection points towards certain focus areas within teaching methods. however, it is necessary to conduct empirical studies of learning activities to learn more about reflective processes. the presented framework has implications for such future research. research within situated reflection calls for studies that can identify very specific processes of students' activities. such research requires observations of students in action. situated reflection cannot be identified through interviews that look back on activities. students might not be aware of their reflective processes, and thus they must be identified within the situations. further, in order to discover reflective processes, it is necessary to engage in dialogue with the students about their objectives and considerations within their situated activities. further, the framework of this article provides analytical concepts that can guide empirical studies on situated reflective activities. based on the framework, empirical studies on reflective activities could employ the concepts of object, examination and construction, instrument, judgment and reflection to analyse student activities through observational studies. the concept of object can be used to identify the directed nature of students' activities: are students aware of an object, aim or purpose of their activities in their studies, and do they understand what that object, aim or purpose is? a challenge of such studies is that it requires asking students questions during their reflective activities, or asking them to "think aloud". empirical studies could also focus on identifying examination and construction of instruments in student activities: are students actively involved in examining instruments, tools, concepts and theories with the aim of using them to reach the object? such processes would involve students' active engagement with the subject matter, but with an intention of using them for an end. when looking for processes of judgment, empirical studies could examine whether students have a conscious awareness of the object, end or aim of their activities and whether they use this awareness to judge their own activities. finally, studying reflection involves an examination of whether objects (in the form of empirical goals and motives) and instruments (in the form of conceptual and theoretical models and methodologies) are present in student activities. for instance, are communication theories present in students' situated work on creating an advertisement? secondly, are these objects and instruments used actively in students work for empirical and/or theoretical reflection on their activities; do concepts from communication theory influence the advertisement? the main challenge of such research is to make the situated reflective activities visible. examination and construction may be visible in some cases, whereas judgment and reflection are often unspoken. studying reflection in group work may be a way forward, because discussions between students may make visible their joint reflection. in conclusion, the framework can be used as a basis for future studies that examine whether and how students are able to reflect, empirically and theoretically. keypoints reflection is not only a matter of students viewing their work from a distance, but also a matter of embedding reflection in situated activities. reflection requires tools (concepts and theories) and objects (aims and purposes) that students can consciously employ for reflection. concepts, models and theories should not be taught in isolation, but should be employed as instruments related to students' object-oriented activity. learning activities should ideally be directed at a goal, grounded in a motive, and point towards operations in practice. concepts, models and theories become secondary to the object of student activities, but on the other hand they are key to theoretical reflection. references bateson, g. (1972). steps to an ecology of mind. the university of chicago press. billet, s. (1996). situated learning: bridging sociocultural and cognitive theorising, learning and instruction, vol 6(3), pp. 263-280. billet, s. (2001). knowing in practice: re-conceptualising vocational expertise, learning and instruction, 11 (2001), pp. 431-452. brown, j. s., collins, a., & duguid, p. (1989). situated cognition and the culture of learning. educational researcher, vol. 18(1), 32-42. chaiklin, seth, & lave, jean. (1996). understanding practice. perspectives on activity and context. cambridge university press. compton, p. (2013). situated cognition and knowledge acquisition research, international journal of human-computer studies, vol 71(2013), pp. 184–190. davydov, v. v. (1988). problems of developmental teaching (part i). soviet education, august. davydov, v. v. (1999). what is real learning activity? in m. hedegaard, & j. lompscher (eds.), learning activity and development (pp. 123-138). aarhus university press. dewey, j. (1916). democracy and education. the free press. dewey, j. (1938). logic. the theory of inquiry. new york: henry holt and company. dewey, j. (1958). experience and nature. new york: dover publications. dewey, j. (1997). how we think. new york: dover publications. engeström, y. (2014). learning by expanding. an activity-theoretical approach to developmental research. 2nd edition. cambridge university press december 2014. engeström, y. (1999). expansive visibilization of work: an activity-theoretical perspective. computer supported cooperative work (cscw), 8(1), 63-93. engeström, y. (2001). expansive learning at work: toward an activity theoretical reconceptualization. journal of education and work, 14(1), 133-156. engeström, y., & sannino, a. (2010). studies of expansive learning: foundations, findings and future challenges. educational research review, 5(1), 1-24. engeström, y., & sannino, a. (2012). whatever happened to process theories of learning? learning, culture and social interaction, 1 (1), 45 – 56. fors, v., bäckström, å., & pink, s. (2013). multisensory emplaced learning: resituating situated learning in a moving world, mind, culture, and activity, vol 20(2), pp. 170-183. hutchins, edwin (1995). cognition in the wild. london: the mit press. hutchins, edwin (1996). learning to navigate. i: chaiklin, seth and lave, jean (red.). understanding practice. perspectives on activity and context, p. 35-63. cambridge university press. lave, j. (1996). the practice of learning. in s. chaiklin, & j. lave (eds.), understanding practice. perspectives on activity and context (pp. 3-32). cambridge university press. lave, j., & wenger, e. (1991). situated learning: legitimate peripheral participation. cambridge university press. leontev, a. n. (1978). activity, consciousness, and personality. online: http://www.marxists.org/archive/leontev/works/1978/index.htm. leontyev, a. n. (1981). problems of the development of the mind. moscow: progress publishers. roth, w-m. & jornet, a. (2013). situated cognition, wires cogn sci, vol 4, pp. 463–478. doi: 10.1002/wcs.1242. salomon, g. (1993). distributed cognitions. cambridge university press. salomon, g., & perkins, d. n. (1998). individual and social aspects of learning. review of research in education, 1-24. schön, d. a. (1983). the reflective practitioner. ashgate. schön, d. a. (1987). educating the reflective practitioner. san francisco: jossey-bass. semin, g.r. & e.r. (2013). socially situated cognition, social cognition, vol. 31(2), pp. 125–146. suchman, l. a. (1987). plans and situated actions. cambridge university press. suchman, l. a., & trigg, r. h. (1996). artificial intelligence as craftwork. in s. chaiklin, & j. lave (eds.), understanding practice. perspectives on activity and context (pp. 144-178). cambridge university press. vygotsky, l. s. (1978). mind in society. london: harvard university press. vygotsky, l. s. (1986). though and language. cambridge: the mit press. wartofsky, m. w. (1968). conceptual foundations of scientific thought. new york: the macmillan company. wartofsky, m. w. (1979). models: representation and the scientific understanding. d. reidel publishing company. wenger, e. (1998). communities of practice. cambridge university press. wertsch, j. v. (1994). the primacy of mediated action in sociocultural studies. mind, culture, and activity, vol. 1(4), 202-208. bateson, g. (1972). steps to an ecology of mind. the university of chicago press. billet, s. (1996). situated learning: bridging sociocultural and cognitive theorising, learning and instruction, vol 6(3), pp. 263-280. billet, s. (2001). knowing in practice: re-conceptualising vocational expertise, learning and instruction, 11 (2001), pp. 431-452. brown, j. s., collins, a., & duguid, p. (1989). situated cognition and the culture of learning. educational researcher, vol. 18(1), 32-42. chaiklin, seth, & lave, jean. (1996). understanding practice. perspectives on activity and context. cambridge university press. compton, p. (2013). situated cognition and knowledge acquisition research, international journal of human-computer studies, vol 71(2013), pp. 184–190. davydov, v. v. (1988). problems of developmental teaching (part i). soviet education, august. davydov, v. v. (1999). what is real learning activity? in m. hedegaard, & j. lompscher (eds.), learning activity and development (pp. 123-138). aarhus university press. dewey, j. (1916). democracy and education. the free press. dewey, j. (1938). logic. the theory of inquiry. new york: henry holt and company. dewey, j. (1958). experience and nature. new york: dover publications. dewey, j. (1997). how we think. new york: dover publications. engeström, y. (2014). learning by expanding. an activity-theoretical approach to developmental research. 2nd edition. cambridge university press december 2014. engeström, y. (1999). expansive visibilization of work: an activity-theoretical perspective. computer supported cooperative work (cscw), 8(1), 63-93. engeström, y. (2001). expansive learning at work: toward an activity theoretical reconceptualization. journal of education and work, 14(1), 133-156. engeström, y., & sannino, a. (2010). studies of expansive learning: foundations, findings and future challenges. educational research review, 5(1), 1-24. engeström, y., & sannino, a. (2012). whatever happened to process theories of learning? learning, culture and social interaction, 1 (1), 45 – 56. fors, v., bäckström, å., & pink, s. (2013). multisensory emplaced learning: resituating situated learning in a moving world, mind, culture, and activity, vol 20(2), pp. 170-183. hutchins, edwin (1995). cognition in the wild. london: the mit press. hutchins, edwin (1996). learning to navigate. i: chaiklin, seth and lave, jean (red.). understanding practice. perspectives on activity and context, p. 35-63. cambridge university press. lave, j. (1996). the practice of learning. in s. chaiklin, & j. lave (eds.), understanding practice. perspectives on activity and context (pp. 3-32). cambridge university press. lave, j., & wenger, e. (1991). situated learning: legitimate peripheral participation. cambridge university press. leontev, a. n. (1978). activity, consciousness, and personality. online: http://www.marxists.org/archive/leontev/works/1978/index.htm. leontyev, a. n. (1981). problems of the development of the mind. moscow: progress publishers. roth, w-m. & jornet, a. (2013). situated cognition, wires cogn sci, vol 4, pp. 463–478. doi: 10.1002/wcs.1242. salomon, g. (1993). distributed cognitions. cambridge university press. salomon, g., & perkins, d. n. (1998). individual and social aspects of learning. review of research in education, 1-24. schön, d. a. (1983). the reflective practitioner. ashgate. schön, d. a. (1987). educating the reflective practitioner. san francisco: jossey-bass. semin, g.r. & e.r. (2013). socially situated cognition, social cognition, vol. 31(2), pp. 125–146. suchman, l. a. (1987). plans and situated actions. cambridge university press. suchman, l. a., & trigg, r. h. (1996). artificial intelligence as craftwork. in s. chaiklin, & j. lave (eds.), understanding practice. perspectives on activity and context (pp. 144-178). cambridge university press. vygotsky, l. s. (1978). mind in society. london: harvard university press. vygotsky, l. s. (1986). though and language. cambridge: the mit press. wartofsky, m. w. (1968). conceptual foundations of scientific thought. new york: the macmillan company. wartofsky, m. w. (1979). models: representation and the scientific understanding. d. reidel publishing company. wenger, e. (1998). communities of practice. cambridge university press. wertsch, j. v. (1994). the primacy of mediated action in sociocultural studies. mind, culture, and activity, vol. 1(4), 202-208. codepen jansen et al publication frontline learning research vol.8 no. 2 (2020) 35 64 issn 2295-3159 a mixed method approach to studying self-regulated learning in moocs: combining trace data with interviews renée s. jansen,a anouschka van leeuwen,a jeroen janssena & liesbeth kester a adepartment of education, utrecht university, the netherlands article received 9 july 2019/ revised 3 february 2020/ accepted 4 march / available online 9 april abstract to be successful in online education, learners should be able to self-regulate their learning due to the autonomy offered to them. accurate measurement of learners’ self-regulated learning (srl) in online education is necessary to determine which learners are in need of support and how to best offer support. trace data is gathered automatically and unobtrusively during online education, and is therefore considered a valuable source to measure learners’ srl. however, measuring srl with trace data is challenging for two main reasons. first, without information on the how and why of learner behaviour it is difficult to interpret trace data correctly. second, srl activities outside of the online learning environment are not captured in trace data. to address these two challenges, we propose a mixed method approach with a sequential design. such an approach is novel for the measurement of srl. we present a pilot study in which we combined trace data with interview data to analyse learners’ srl in online courses. in the interview, cued retrospective reporting was conducted by presenting learners with visualizations of their trace data. in the second part of the interview, learners’ activities outside of the online course environment were discussed. the results show that the mixed-method approach is indeed a promising approach to address the two described challenges. suggestions for future research are provided, and include methodological considerations such as how to best visualize trace data for cued retrospective recall. keywords: interview; mixed-method research; online education; self-regulated learning; trace data info corresponding author email: r.s.jansen-14@umcutrecht.nl. doi: 10.14786/flr.v8i2.539 1. introduction online learning has increased rapidly over the past years (allen & seaman, 2014, 2016). not only are online learning materials increasingly included in campus-based education, the amount of courses offered fully online has also expanded rapidly. such courses that are offered fully online are called massive open online courses (moocs) which have become increasingly common in higher education. they are accessible to anyone with an internet connection without requirements regarding prior knowledge. in some cases, costs are involved to participate in graded assignments, but access to the learning materials (e.g., videos, readings) is always free. in these online courses, learners are free to decide when, where, and what they study. this increased autonomy in online education requires learners to self-regulate their learning to a greater extent compared to students in traditional campus-based education (azevedo & aleven, 2013; beishuizen & steffens, 2011; broadbent, 2017; hew & cheung, 2014; wang, shannon, & ross, 2013). the measurement, analysis, and support of self-regulated learning (srl) in the context of online education is therefore of high scientific as well as practical relevance. in order to examine students’ srl in online education, researchers have increasingly focused their attention on trace data (van laer & elen, 2018; winne, 2010). in online education, all learners’ interactions with the online course materials (e.g., videos, assignments, forum discussions) are stored. these so-called traces of learner behaviour are stored as time stamped events, providing an overview of all learners activities within the online learning environment. using trace data as an indicator for learners’ srl is considered a promising approach because the data are gathered automatically, unobtrusively, and over longer durations of time (rovers, clarebout, savelberg, de bruin, & van merriënboer, 2019; van laer & elen, 2018). over the past years, the call for fine-grained process measures of srl has strengthened (e.g., winne, 2010) and researchers have therefore started using trace data to measure srl (e.g., cicchinelli et al., 2018; van laer & elen, 2018). using data mining techniques, trace data could be used to automatically measure and support srl, for example by automatically detecting when students show a lack of self-regulation and subsequently offering feedback, suggestions, or other forms of interventions. however, in order to accurately measure students’ srl from trace data, it should be possible to interpret trace data in terms of srl in a reliable and valid way. interpreting learners’ srl based on trace data is challenging for multiple reasons. first, while trace data show what and when learners study, it does not provide information on how and why learners engaged with the learning material the way they did (jovanović, gašević, dawson, pardo, & mirriahi, 2017; min & jingyan, 2017; phillips et al., 2011). second, trace data are limited to capturing learners’ behaviour in the online learning environment only, and thus does not include other learning behaviour, such as browsing websites or making notes on paper. these two reasons combined mean that translating trace data into conclusions about learners’ srl is challenging because of the possibility of misinterpreting certain events or overlooking relevant events. it has therefore been argued that a mixed method approach, in which trace data are combined with other methods to measure srl, is useful and necessary to draw valid conclusions from the trace data (cicchinelli et al., 2018; howard-rose & winne, 1993; jovanović et al., 2017; karabenick & zusho, 2015; reimann, markauskaite, & bannert, 2014). in this methodological paper, our aim is to demonstrate that interview data could be a valuable addition to trace data to measure and analyse learners’ srl in online education. through interviews, researchers are able to examine learners’ reasons behind their activities, as well as to gather an overview of their activities outside of the online learning environment, thus offering the possibility to overcome the challenges outlined above. in this paper, we present a pilot study in which we combined trace data with interview data, and discuss the methodological possibilities and challenges of our approach concerning the aim to understand learners’ srl in online education. 1.1. measuring srl with trace data: affordances and challenges self-regulated learners are defined as motivationally, metacognitively, and behaviourally active in their own learning process (zimmerman, 1986). self-regulated learners engage in a number of activities, including goal setting, planning, monitoring, reflection, attention focusing, time management, environment structuring, and seeking help when needed (panadero, 2017; puustinen & pulkkinen, 2001). thus, the extent to which learners adapt to changes in the task and learning context is a critical component of srl (azevedo & cromley, 2004; hadwin, nesbit, jamieson-noel, code, & winne, 2007). as srl is considered a process, any measurement of srl must take into account changes in learners’ behaviour over time (azevedo et al., 2013; winne, 2010). as outlined above, trace data allow for the capturing of learners’ activities unobtrusively over time at a very fine granularity. trace data provide a reliable measure of when students engaged in the online learning environment and with what materials they engaged, and offers information about temporal and sequential characteristics of activities, making it a readily available and valuable data source to study learners’ srl (cicchinelli et al., 2018; fincham, gasevic, jovanovic, & pardo, 2018; hadwin et al., 2007; kizilcec, pérez-sanagustín, & maldonado, 2017; maldonado-mahauad, pérez-sanagustín, kizilcec, morales, & munoz-gama, 2018; winne, 2014). several authors have used trace data to attempt to locate learner behaviour that is representative of srl activities (van laer & elen, 2018). for example, in the study of kizilcec et al. (2017), srl questionnaire data were coupled to learner behaviour information stored in trace data to identify learner activities related to srl. similarly, min and jingyan (2017) also used trace data as a measure of srl. before their data collection, min and jingyan defined sequences of learner behaviour that, in their view, were indicative of srl. learners whose trace data included sequences indicative of all defined srl activities had greater persistence in the course and achieved higher course grades. these exemplary studies show trace data may be beneficial for studying srl in online education: trace data are gathered unobtrusively, can be analysed automatically for large groups of learners, and may in the future even be used for real time support of learners (van laer & elen, 2018). however, the same studies also demonstrate the challenges of using trace data as a measurement of srl. the most striking problem is that while trace data shed light on leaners’ behaviour in the online learning environment, there is doubt on how to interpret behaviour in terms of srl (cicchinelli et al., 2018; jovanović et al., 2017; min & jingyan, 2017; phillips et al., 2011; rovers et al., 2019). srl activities are for a large part covert in nature; they constitute the regulating activities that shape and guide the observable learning activities (nelson & narens, 1990). the consequence is that interactions and interaction sequences with the learning material are ambiguous; there are usually multiple plausible explanations from the perspective of srl (jovanović et al., 2017; phillips et al., 2011). for example, watching the same video twice in the online environment may indicate that the student found the material hard to understand and is therefore re-watching the video (an indication of comprehension monitoring). it may however also be the case that the student did not remember already watching the video (problems with effort regulation), and therefore watches the video twice. the learning activities captured by trace data thus need to be interpreted before they can be labelled as originating from a student’s self-regulating behaviour, which makes interpretation of the results of trace data analyses complicated (maldonado-mahauad et al., 2018; schraw, 2010). this difficulty to interpret learners’ activities in terms of self-regulating behaviour, may thus also compromise the validity of the results obtained in studies that employ this methodology. more insight is thus needed into the how and why of learners’ activities to understand the reasons underlying learners’ behaviour and their meaning in terms of srl (cicchinelli et al., 2018; jovanović et al., 2017; min & jingyan, 2017; phillips et al., 2011). authors such as kizilcec et al. (2017) therefore make use of questionnaire data to identify learner activities related to srl (see also the section on mixed methods). in the sections below, we argue why we think interviews are more suited for this purpose. a second challenge in measuring srl with trace data is that some srl activities take place outside of the online learning environment (min & jingyan, 2017). winne and jamieson-noel (2002) for instance argued that scrolling of learners through a text-document indicated planning, planning may however also have been (mostly) a mental activity (rovers et al., 2019). veletsianos, reich, and pasquini (2016) studied learners’ activities outside of the online learning environment and identified additional activities in three domains. in each of these domains, srl activities may take place: behaviours at the learners’ workplace (srl activities: note-taking, making a planning, picking the study location), learners’ activities online, but off-platform (srl activity: looking for help by browsing the web), and learning activities in the wider context of their lives (srl activity: dilemma’s in time management due to other priorities). thus, trace data may be helpful in providing insight into students’ learning behaviours, and these behaviours may be interpreted in terms of srl, but there is reason to suspect that trace data does not capture all of students’ srl activity. trace data measurement of srl should thus be supplemented with a measurement method that enables a) understanding the reasons underlying learners’ behaviour, and b) measuring and understanding learners’ srl activities outside of the online learning environment. in the present paper, we propose that combining trace data with interviews, i.e., using a mixed methods approach, could help to solve these two challenges. in the sections below, we reflect on what mixed methods research entails, and then elaborate on the mixed methods approach of measuring srl by combining trace data with interviews. 1.2. measuring srl with mixed methods research in mixed method research, qualitative and quantitative research methods are combined (creswell, 2008; johnson & onwuegbuzie, 2004). mixed method research can be classified on two dimensions: the time order decision and the paradigm emphasis decision (creswell, 2008; johnson & onwuegbuzie, 2004). the time order is either sequential or concurrent depending on whether one method informs the other, or if both methods are used concurrently to gather data. the paradigm emphasis decision is either equal status, if quantitative and qualitative methods are of equal importance, or dominant status, if either the quantitative or qualitative data collection is given more weight. by classifying mixed method research on these two dimensions, four types of designs emerge. the most suitable design depends on the purpose of the mixed method study, which could for example be triangulation, complementarity, or expansion. independent of the design, a mixed method approach will generally provide a more valid measurement of the construct studied than any single method can provide (creswell, 2008; johnson & onwuegbuzie, 2004; mcgrath, martin, & kulka, 1981). in the context of srl, several researchers have employed a mixed methods design in various combinations of time order and paradigm emphasis for various empirical goals (ben-eliyahu & linnenbrink-garcia, 2015). for instance, littlejohn, hood, milligan, and mustain (2016) aimed to obtain more in depth information about five sub-processes of srl that are commonly measured with self-report questionnaires, namely motivation and goal setting, self-efficacy, task strategies, task interest value, and self-satisfaction and evaluation. to do so, learners filled in the questionnaire within a mooc environment, and a selection of questionnaire respondents was later interviewed. the authors thus combined two self-report measurement methods in a sequential mixed method design, where quantitative data was used as input for the qualitative data collection, with paradigm emphasis on the qualitative interview data (creswell, 2008; johnson & onwuegbuzie, 2004). the interview data yielded several insights, for example that learners with high scores on srl were less focused on obtaining the certificate compared to learners with low scores on srl, and more focused on professional development and the relevance of the learning material for their job. other work relevant to the present study is the paper by kizilcec et al. (2017), in which trace data was combined with questionnaire data to identify sequences of learner behaviour correlated to high or low scores on specific srl scales. this is another example of a study combining instruments in a concurrent time order. the authors correlated learners’ scores on the srl questionnaire with specific transitions in behaviour. the results included the finding that learners who reported higher srl skills were more likely to revisit earlier materials instead of starting new materials after completing a part of the course. the authors thus aimed to specify behaviours that correlate with self-reported srl. as stated earlier, one of the main challenges associated with interpreting trace data in terms of srl is that learners’ behaviour is sometimes ambiguous. the study by kizilcec (2017) shows that trace data can indeed be coupled to srl, but does not completely solve this challenge, because multiple srl scales were found to be correlated to the same behaviour, leaving open the question how to interpret specific instances of behaviours. additionally, because their starting points were the scores on the srl questionnaire, not all behaviour found in the trace data was ‘matched’ with one of the srl scales. in terms of time ordering in mixed methods studies, we therefore want to propose a sequential methodology that has the trace data as its starting point. in the present paper, our goal is to build on these earlier mixed methods studies that have aimed to increase our understanding of trace data in terms of srl. we propose to employ a mixed methods design in which trace data are complemented with interview data. interviews allow researchers to explore and understand how people behave and think (alshenqeeti, 2014). in contrast to questionnaires, interviews also allow for follow-up questions that emerge from the dialogue to probe for further information and improve understanding of the learner’s activities (dicicco-bloom & crabtree, 2006). they are therefore suitable to gain insight into learners’ behaviour in a flexible way, both concerning behaviour inside and outside of the online learning environment. while we acknowledge that interviews are not scalable like the use of trace data to measure srl, this pilot study shows that interviewing a selection of mooc students already provides a wealth of information on which srl activities can or cannot be measured. we therefore believe that the effort necessary to interview a range of mooc students is justified as it will help in the development of trace data as a valid and reliable measure of srl. in the section below, we reflect on how interviews could be used to address the two challenges of interpreting trace data. 1.3. combining trace data with interviews the first challenge associated with interpreting trace data in terms of srl concerns the trace data’s ambiguity. to understand why a learner performed a specific action at a specific time, the interview technique of verbal protocols, in which a participant verbalizes his or her thoughts and actions (ericsson & simon, 1993), could be a solution. several types of verbal reporting exist, including concurrent reporting (verbalizing during the task), retrospective reporting (verbalizing after the task), and cued retrospective reporting (verbalizing after the task, induced with a cue) (ericsson & simon, 1993; van gog, paas, van merriënboer, & witte, 2005). with cued retrospective reporting, learners report on their thoughts and activities after the task, but receive a cue to help them remember the process correctly (e.g., eye movements). cued retrospective reporting thereby attempts to minimize errors of omission and fabrication, which may occur with retrospective reporting, and without risking to alter the primary process (i.e., reactive invalidity), which may occur with concurrent reporting (russo, johnson, & stephens, 1989; van gog et al., 2005). learners’ trace data could be visualized and presented to them as a cue during an interview to help the learner remember and reflect on his/her learning process. the learner could be asked about specific activities and transitions in the trace data, and thereby be a tool to understand learners’ behaviour in terms of srl. of course, the numerous events in a learners’ trace data, and therefore also numerous transitions, make it impractical and unrealistic to have learners explain and reflect on all of these events and transitions in the form of cued retrospective reporting. the researcher therefore has to make a selection of specific events and transitions that are presented as a cue during the interview, so that learners can remember and reflect on their learning process. combined with follow-up questions from the researcher, the cues could then be used to help learners understand which activities, and which transitions, the interviewer is informing about. by incorporating elements of cued retrospective reporting in such a way, knowledge about specific activities and transitions can be obtained, while also increasing knowledge on the interpretation of these activities, which in traditional verbal reporting is the sole responsibility of the interviewer (ericsson & simon, 1993). the second challenge associated with interpreting trace data is capturing not only the online srl activities within the online learning environment, but also those outside of it. again, we argue that interviews could be a tool to address this issue. for this specific challenge, the more traditional form of interviews could be employed, in terms of the researcher asking overarching questions about the participants’ learning process without presenting a cue. for this interview format, three types are distinguished: structured, semi-structured, and unstructured interviews (alshenqeeti, 2014; dicicco-bloom & crabtree, 2006). structured interviews are most like a verbal questionnaire and often produce quantitative data (alshenqeeti, 2014; dicicco-bloom & crabtree, 2006). since the current aim is to understand learner behaviour and move beyond sole quantitative data, structured interviews provide too little freedom and they are therefore unsuitable for the current purpose. in contrast, unstructured interviews may provide too much freedom. during unstructured interviews, the questions asked often arise during the interview itself and these sessions are thereby more like guided conversations (dicicco-bloom & crabtree, 2006). this approach is valuable when interviewers want to minimalize their influence on the interviewee or when little is known about the topic at hand. however, when measuring srl, it is important to measure all aspects of the construct. over time, srl has been defined in a number of ways, but all entail roughly the same activities that together form srl activities (jansen, van leeuwen, janssen, kester, & kalz, 2017; puustinen & pulkkinen, 2001). to get a full grasp of learners’ srl, it is worth discussing this predetermined list of srl activities with the learner (e.g., goal setting, time management). in a semi-structured interview, the srl activities can be used as topics for which pre-determined questions are created. a checklist can be used to make sure all topics are discussed (alshenqeeti, 2014). at the same time, a semi-structured interview allows for follow-up questions that emerge from the dialogue to be asked, to probe for further information and improve understanding of the learner’s activities (dicicco-bloom & crabtree, 2006). a semi-structured interview is therefore a suitable method to gather data on learners’ srl activities outside of the online learning environment, by asking them about each srl component. 1.4. the present study we present a pilot study in which we augment trace data with interview data to analyse learners’ srl in online courses. trace data from several online learners were analysed. these learners were later interviewed about the regulation of their learning during the course they were enrolled in. in the interview, cued retrospective reporting was conducted by presenting learners with visualizations of their trace data. in the second part of the interview, learners’ activities outside of the online course environment were discussed in a semi-structured interview format. in the current study, interviews were conducted face to face with learners residing in the same country as the interviewer. however, interviews could also have been conducted online in case learners had resided in different countries, as would often be the case in moocs. we present our methodology for a subset of our data, so that we are able to thoroughly explain our procedures for data collection and interpretation. our aim is therefore to make a methodological contribution to measuring leaners’ srl in online education. we conclude the paper by discussing the benefits and possible improvements of our approach. 2. method 2.1. design a mixed method research study was performed as a methodological illustration of how to measure srl in online education, in which quantitative trace data were analysed in conjunction with qualitative interview data (johnson & onwuegbuzie, 2004). the interview served two main goals. the first goal was to gain a better understanding of learners’ trace data in terms of srl. the second goal was to gain insight into learners’ srl outside of the online learning environment. therefore, the time order of this mixed method study was sequential; the trace data was analysed and then used as input for semi-structured interviews. dominant status (paradigm emphasis decision) was given to the interview data, as we were interested in the additional information that can be extracted from adding interviews to trace data. 2.2. participants four students of the same dutch university were interviewed. at this university, students were offered the opportunity to take a traditional exam, with a fixed time and location, after following a mooc. if they passed the exam, they would be given elective credits. the four interviewed students all successfully took such an exam and thus received elective credits. this form of online education is known as a small private online course (spoc). commonly, spoc students form a small subgroup within the mooc environment. they are for instance offered additional help from a tutor and they can discuss learning with other spoc learners in a separate course forum. such distinctions between spoc learners and regular mooc learners were not present. the educational materials all had to be accessed through edx (an online learning platform). the online learning experience of those studying at the university (current participants, spoc learners) was identical to the learning experience of those not studying at the university (mooc learners). the autonomy offered to spoc learners was therefore also identical to the autonomy offered to mooc learners. while the opportunity to earn elective credits may have impacted the motivation of the spoc students, the courses were not part of the formal curriculum of any of the students and therefore not compulsory. the impact of the credits on course motivation was therefore likely to be limited. all students that registered for an exam in the spring of 2018 were invited to be interviewed. five students indicated they were willing to be interviewed, and gave consent for the analysis of their trace data. one student had taken the exam already in the fall of 2017 and was therefore excluded. the data of four students is therefore used (1 male; mean age = 22.0, range 19-24). the students were interviewed by the first author. participants received a 20 euro gift voucher for a large online warehouse as compensation for their participation. 2.3. trace data analysis all learners’ activities in the online learning environment were automatically stored, leaving a trace of their learning behaviour (hence the term ‘trace data’). this trace data included a time stamped log of all their activities, including videos played and paused, self-test questions answered correctly and incorrectly, and exam questions answered correctly and incorrectly. as the data were time stamped and related to a learners’ user id, the order in which a learner engaged with the learning materials is known. the trace data can be analysed in a number of ways, for instance by looking at absolute or relative frequencies of activities, transitions from one activity to another, or patterns of activities per session or over the entire learning period. depending on the aim of the research, different analyses are suitable. for instance, to identify activities or sequences of activities related to learner achievement, analyses counting activities and sequences would be fitting. currently, the aim is to gain insight in participants’ learning processes and to identify potential srl activities present in their trace data. for this aim, the trace data of participants’ was analysed in a number of ways. the timing of learners’ study sessions was analysed, indicating when and how long a learner studied in the online course environment. furthermore, attention was paid to the frequency of activities and potential skipping of learning activities. finally, the order of learner’s activities in the online learning environment was inspected. to stimulate learner recall of their learning activities, a visualization of the learning process was considered most suitable, since it provides learners with an overview of their learning process and such a visualization can be understood rather easily without prior knowledge of the data gathered. therefore, for each participant, a visualization of the learners’ activities within the online learning environment was created. to aid understanding of these trace data visualizations, an excerpt of the transition diagram of p4 is presented in figure 1. more information about the analysis of the trace data, and the interpretation of the transition diagrams can be found in the sections “analysis” in the method and “trace data interpretation” in the results. figure 1. excerpt of the transition diagram of p4. the transition diagram visualizes the order in which p4 engaged with the learning materials of module 2 and 3. 2.4. interview a semi-structured interview guide was developed with a descriptive/interpretive focus to gain insight into the explanations of learners’ online behaviour and their self-regulated learning outside of the online learning environment (mcintosh & morse, 2015). a visualization of learners’ trace data was used during the interview as a cue for recall of the learning process. the interview guide consisted of four segments and is presented in appendix a. first, the interviewer explained the interview goal and asked several introductory questions. the introductory questions included asking the participant about the title and topic of the spoc the student followed and how the participant had learned about the opportunity to take a mooc for credit. the questions were constructed to learn some basic information about the participant and to create rapport between the interviewer and the interviewee. next, learners’ activities were discussed at a micro level, meaning that learners were asked about their behaviour in the learning environment. the aim was to gain an srl-themed explanation for learners’ trace data. to help learners remember their learning process, a visualization of their trace data was presented to learners as a cue (see figure 1 for an example). this segment thereby had similarities to cued retrospective reporting. however, not all activities and transitions in the visualization were discussed, not only because the learning process was simply too complex to do so, but also because some of the activities were conducted months before the interview. more importantly, the aim of the interview also was not to understand all transitions and activities, but to gain an understanding of learners’ overall learning process, how they managed their learning, and the reasons underlying their learning activities. learners were therefore not only presented the trace data visualization, but were also interviewed about their learning process in general and about the way they engaged with the specific elements of the course (e.g., introductory videos, content videos, self-test questions, exam questions), to better understand with what intent the learners engaged with these materials (e.g., content learning, comprehension check). learners were furthermore interviewed about the timing of their learning sessions (i.e., when they worked on the course), to understand the (ir)regularity of their studying and their time management. third, learners’ activities were discussed at a macro level. learners were interviewed about srl activities not visible in the trace data and their behaviour outside of the learning environment. learners were for instance asked about the location where they studied, potential help-seeking, and their planning. it was made sure that all aspects of srl, as described in the articles of jansen, van leeuwen, janssen, and kester (2017) and puustinnen and pulkinnen (2001) were discussed with the participant at a micro and/or at a macro level. finally, learners were interviewed about any challenges they had encountered during learning, and if they had suggestions for improving the course, especially related to srl. the interviews were conducted in dutch as it was the mother tongue of both the interviewer and all interviewees. the interviews lasted between 1 hour and 1 hour and 15 minutes each. at the start of the interview, participants signed an informed consent for the interview to be audio and video taped. the video camera was directed at the table to record the trace data materials shown to participants and any pointing to specific activities and transitions on these materials (e.g. “during this module”, or “when i moved from here to here”). 2.5. procedure after registering for a spoc exam, students were emailed information about the present research study. they were given the option to supply their contact information if they were willing to be interviewed. students were informed that their trace data would be analysed if they indicated their interest in participating. the trace data of interested students was analysed a) to determine the amount of learners’ activities in the online learning environment, b) to determine the length and amount of learning sessions, and c) to create a transition diagram of students’ learning. the interview was structured by following the interview guide described above. learners could be invited to participate in the interviews only after registering for the spoc exam, and only after agreeing to participate could their trace data be analysed. as trace data analysis itself is time consuming, interviews were scheduled approximately two months after the spoc exams. 2.6. analysis the trace data was inspected to gain knowledge on the timing and length of learners’ sessions, the frequency of their activities, and the order of their learning activities. the activities of each learner were visualized in a transition diagram. while the participants followed different courses, the main elements of all courses were the same. therefore their transition diagrams also contained the same elements. per module, the following activities were visualized in the transition diagram: introduction video, content video (there were usually multiple content videos per week, interactions with either of the videos were aggregated to this label), self-test question correct, self-test question incorrect, exam correct, exam incorrect, and summary video. furthermore, browsing the forum and posting on the forum were visualized in the transition diagram, but on a course level and not on a module level. next, each learner’s transitions were added to his course model. as only a small sample of learners participated in the current study, this was done by hand. furthermore, per learner an overview was made of when the learner engaged with the course. additionally, a description was made of the overall learning process of the student, including any activities or transitions that were considered remarkable or interesting. it was then attempted to distil srl activities from the trace data analyses. the transition diagrams visualizing the learning process of the participants were used during the interviews as a cue for recall (see figure 1). after conducting the interviews, they were first transcribed. next, statements that helped interpret the trace data, especially in terms of srl, and statements that indicated srl activities not visible in the trace data were identified in the transcripts. interview codes were created mostly top-down, based on the srl activities that were also used to structure the interviews. additional codes could be created during coding if the srl activities listed beforehand were insufficient saturation (morse, 2010; strauss & corbin, 1994). the srl statements were then labelled with the trace data or srl activity they provided information about. for each theme, statements of the different participants were then grouped and a description of the information gathered from the interviews was written per theme. the coding scheme can be found in appendix b. 3. results the results are presented in three parts. first, case descriptions of the four participants are given. the case descriptions contain, per participant, information about the course followed, their study intentions, and the timing of their learning sessions. next, we present the trace data of one of the participants (p4), and attempt to interpret the trace data in terms of srl. while the trace data provide an objective overview of the learners’ activities, no definitive conclusions regarding srl activities could be made based on trace data alone for p4, nor for the other participants. the trace data of p4 are used to illustrate the data available and to illustrate the opportunities and challenges of using the data, especially for understanding and interpreting the data in terms of srl. last, the interview findings are presented in which we focus on the additional insights gained from interviewing learners in addition to analysing their trace data. we first indicate in what way the interviews improved our understanding of the trace data. second, we describe the srl activities of learners not visible in the trace data. 3.1. case descriptions the four interviewees followed different spocs. the exams for all spocs took place mid-february. below, more information per participant is provided about the course they followed, their study intentions, and the information obtained from their trace data. 3.1.1. participant 1 (p1). p1 followed the course ‘food access’ which she finalized with a grade of 7.5 (out of 10). she had time to spare next to the regular curriculum which she wanted to utilize in a meaningful manner. she started the spoc out of interest in the topic of the course, and decided on taking the exam after completing a large part of the course. the spoc was her first experience with a fully online course. in total, 1146 activities were logged for p1, distributed over 20 days in a time span of five months. figure 2 shows the timing of the learning sessions. figure 2. timing of learning sessions p1. green bars indicate activities were logged on that day. 3.1.2. participant 2 (p2). p2 followed the course ‘food risks’ which he finalized with a 8.0. the spoc was part of a so-called ‘micromaster’ which contains three spocs which together serve as a replacement for a traditional campus-based course. p2 had already taken the other two courses in the micromaster, and thus started this spoc with the intention to finish the course, pass the exam, and receive the credits. in total, 1407 activities were logged for p2, distributed over 28 days in a time span of six months. figure 3 shows the timing of the learning sessions. figure 3. timing of learning sessions p3. green bars indicate activities were logged on that day. 3.1.3. participant 3 (p3). p3 followed the course ‘food risks’, the same course as p2, which she finalized with a 8.0. p3 was abroad for a month during the study year, and planned on attending this spoc and two others to gain elective credits while away. however, as she had much less time available for studying while abroad, and there were technical difficulties with the internet connection, she did not complete any of the spocs while abroad. after returning home, she completed the spoc on ‘food risks’; the other spocs were dropped. this was her first experience with studying online. in total, 1012 activities were logged for p3, distributed over 22 days in a time span of two months. figure 4 shows the timing of the learning sessions. figure 4. timing of learning sessions p3. green bars indicate activities were logged on that day. 3.1.4. participant 4 (p4). p4 followed the course ‘animal behaviour’, which she finalized with a 8.5. she started the course out of interest in the topic, and only later learned about the opportunity to take an exam and receive elective credits. it was her first experience with studying online. in total, 1296 activities were logged for p4, distributed over 14 days in a time span of 6 months. figure 5 shows the timing of the learning sessions. figure 5. timing of learning sessions p4. green bars indicate activities were logged on that day. 3.2. trace data findings before interviewing the students, their trace data were analysed. for all four participants, the trace data provided a good overview of their learning process, and activities could be identified that might be indicate of srl activities. however, interpretations of the data were ambiguous for all four participants. to illustrate the information gathered from the data, a description of the trace data gathered for p4 is provided, including potential srl interpretations. when analysing the activities stored in the trace data, the trace data for september only showed interactions with the introductory module of the course and the first few videos of module 1 (the first video with actual content). during the sessions in october and november the learner completed all other course materials (videos, readings, exams and assignments). in the two sessions in february, right before the course exam, p4 engaged linearly with the introduction and content videos of all modules. this may indicate strategic studying behaviour. it is however not known if the participant also used other materials to study for the exam. the finding that she completed all materials well before the exam may indicate good time management abilities. in figure 1 (in the method) an excerpt of the transition diagram of p4 was presented, showing the activities of p4 in modules 2 and 3 of the course. the excerpt is exemplary for the entire learning process of p4. p4 engaged with the videos, self-test questions, and exam questions mostly in linear order. however, watching videos was regularly alternated with browsing and posting on the forum. the regular forum interactions of p4 may indicate help seeking. the trace data of the video interaction activities of p4 includes frequent pausing (and then continuing playing) videos, especially with content videos. the trace data furthermore regularly includes seeking a specific time point in the video. pausing the video may indicate note-taking, while seeking in the video may indicate that she noticed a gap in her understanding (comprehension monitoring) or missing notes and that she then attempted to gather the missing knowledge. finally, as is also the case in the excerpt of the transition diagram presented in figure 5, the participant answered most self-test questions correctly. in the few cases that she answered a question incorrectly, p4 first answered the question correctly before continuing with other course materials. however, p4 did not consult the associated course materials before re-answering the question, at least not in a way visible in the trace data. she may have guessed the correct answer, as all self-test questions are multiple choice, or she may have consulted her notes if she made them. this activity order cannot be interpreted based on trace data alone. in sum, the trace data provide information on the learning process of p4. however, understanding the reasons underlying the activities, and interpreting them as the result of srl activities, is not possible without ambiguity remaining. the trace data analysis results in assumptions of srl activities and potential explanations, but doubt on the accuracy of these explanations remains. this was no different when attempting to interpret the trace data of the other three participants. the trace data description of p4 is therefore an illustration of the difficulties in using trace data to measure srl for all participants. in order to measure and understand learners’ srl, more information was necessary for all four participants. below, we describe how the four interviews extended and improved our understanding of learners’ srl activities. 3.3. interview findings in the following sections we present several strands of information which were collected by adding interviews as a measurement method to trace data analyses. while we by no means imply that our results provide a complete description of all potential spoc learners, our results do show the benefit of incorporating interviews, instead of solely analysing trace data to measure learners’ srl. our results can be considered an illustration of the results that can be obtained from the methodological approach advocated in this paper. the findings are structured around two themes: (1) interview findings that improve our understanding of the trace data, and (2) interview findings that teach us about learners’ srl outside of the online learning environment. quotes of the interviewees are incorporated to support our findings. 3.3.1. understanding trace data. in the first part of the interview, participants were questioned about their learning activities in the online learning environment. the insights they provided have led to a better understanding of the how and why of learners’ activities in the online learning environment. linear studying behaviour. the trace data of all participants showed a rather linear approach to learning in the spocs; meaning that the participants – to a large extent worked with the course materials in the same order as it was designed by the course designers. in the interviews, participants explained why they approached their learning in a linear fashion. linear studying was more a habit out of convenience than a conscious decision. as p3 commented “i always just clicked ‘resume course”. p1 said “it seemed logical that they would have a structure for offering materials”, which was also given as a reason for linear studying by p4. p4 furthermore added “by studying linearly, you know what you have done and you don’t need to search for where to continue. you know where you left off and what you still need to do”. linear studying thereby thus had benefits for monitoring progress and made it easier to continue working at a later moment in time. in conclusion, the interviews provided an explanation for why the students followed a linear, regular pattern while studying in the spocs. lack of engagement with the course forum. p4 was the only participant who regularly engaged with the course forum. she introduced herself on the forum, posted her answers to questions when invited to, and asked for help on the forum. the other participants hardly engaged with the course forum, or not at all. when asked about the course forum their first response was that it did not seem useful to them (p1 “i didn’t see any added value in using the forum”); they followed the course for their own learning, and were not interested in the experiences of other learners around the world. the forum however does not solely have a social function; it can also be used for seeking help. all participants indicated that they had had some problems with understanding the course materials. their barriers to asking for help on the forum included not knowing how long it would take to get an answer, finding it too much effort to type out mathematical formulas, and being too stubborn (p2 “i was too stubborn to share my questions, i would then just look back at the video once more”). in sum, the interviews helped to understand the lack of help seeking on the forum indicated by the trace data. the interviews also revealed that visiting the forum and posting questions might have been an effective strategy for learners in some cases. engagement with videos. the trace data showed that the participants watched (almost) all videos. the trace data also showed regular, short pausing of the videos. this likely indicated pausing to take notes, which was confirmed in the interviews, as p1 said “i typed notes while watching the videos, and i paused the video when it went too fast”. the notes were useful for studying for the exam, as all participants indicated using their notes as their main study material for the exam, only supplemented with questions and videos when their notes were not fully clear. the interviews thus supported the hypothesis derived from the trace data: short pauses while watching the videos were caused by students taking notes. the courses contained different types of videos. each module consisted of an introductory video, multiple content videos, and a summary video. participants indicated having different approaches for these different types of videos. they watched the introductory and summary videos “in a more laid back manner” (p3). for instance, “the introductory videos were really an introduction, which meant that i often did not take serious notes” (p2) and “i watched the summary videos, but i didn’t do anything with them. i had just seen all the content, so the summary did not really add anything” (p4). while note-taking was thus a common activity for the content videos, it was much less so for the introductory and summary videos. this difference can also be detected in the trace data, as the content videos were paused more often than the introductory and the summary videos. however, the lack of note-taking does not imply that the videos were not useful for the participants. while some watched the introduction and summary only because they feared they would miss something important if they did not (“i will watch them [the introductory videos] anyway, because maybe they will say something important”, p1), the videos helped participants orient on what was to come and to reflect on the content of the past module. as p4 said about the summary videos “if something was mentioned that i did not remember hearing about, i looked it up in the previous videos. even though i was already putting my stuff away while the summary video was playing”. engagement with questions. participants were clearly aware that answering the self-test questions helped them monitor their comprehension. p4 indicated “i used the questions to check if i had paid attention, and that was almost always the case”. self-regulated learning requires learners to be adaptive in their learning strategy, especially when they face adversity. it is therefore interesting to better understand learners’ strategies when they find questions difficult and/or answer questions incorrectly. these strategies are not (always) visible in the trace data. as p1 explained “questions that i did not answer correctly, or that i found difficult, i wrote down in my notes, together with the correct answer. so that i would understand that information the next time i studied”. p2 had a different strategy. he tried to answer the question. if his answer was incorrect, he first looked back in his notes to work out what went wrong. if that did not answer his question, only then did he go back to the associated video. the trace data thereby thus only showed a part of the strategy of p2 to deal with wrongly answered questions. in sum, the interviews helped to further enlighten the srl processes students were engaged in; they specifically shed light on learners’ engagement with self-test questions and the relation between answering questions and comprehension monitoring. timing of sessions. the timing of learners’ activities (figures 1-4) was diverse; some learners worked on the course for an extended period of time and others only had a few, very active, days. some participants also showed a burst in activity the days before the exam (see figure 1 and 4). the interview provided more information about the timing of learners activities. for instance, p2 spread working on the course materials over a large portion of time. he explained “i knew i had to take this course next to other courses at some point (...) so i better started in september, then i still knew some of the knowledge i gained in the previous two spocs, and that would leave me with the biggest chance to finish in february”. for p1 other obligations, such as campus-based courses, resulted in large stretches of time between working on the course. p3 had planned on studying online abroad, but as she had less time available than expected, studying was postponed until she was back in the netherlands. the trace data show that the first couple of sessions (while abroad) were much shorter in length than the later sessions. the sprint in studying in the two weeks before the exam is thereby visible in the trace data (see figure 3). in these cases, the interviews helped understand learner behaviour and provided explanations for the irregular timing of participants’ learning sessions. associated with the timing of learning sessions are different segments in learners’ studying. for some participants, the trace data showed a clear distinction between completing the course materials and, later in time, studying for the exam. this distinction was visible in the timing of the sessions (paced studying throughout most of the course, and then a few days of high intensity right before the exam), but the ordering of activities also showed a distinction. the participants completed the course materials in a linear fashion. after finishing all materials, they looked back at specific materials: they moved from content video to content video, ignoring the introductory and summary videos, and viewed only some of the questions. the assumption that these behaviours were associated with different segments of the study process could be verified in the interviews. when discussing the exam preparations, p1 commented “i looked at the exams again, to see what kind of questions they asked. and for the things that i did not understood, what i found complicated, for those things i watched the videos again.” here too, the interviews helped to better understand the difference in behaviour of the participating students between the phase where they followed the course and the phase where they were studying for the exam. 3.3.2. srl outside of the online learning environment. the second part of the interview focused on learners’ srl outside of the online learning environment. the interviews thus supplemented the trace data by focusing on activities for which no indicators can be found in the trace data. planning and goal setting. participants’ intentions for taking the spoc differed: while p2 needed the credits associated with the course, p3 was interested in the credits but they were not necessary, and p1 and p4 took the course completely out of general interest and only decided on taking the exam after completing part of the course. however, besides the overall goal of course completion that all participants had sooner or later in the course, they did not set goals. furthermore, they also did not have any plans, or had general plans which they failed to stick to: “i knew i wanted to go through the course linearly, but only when it suited me to work on the course. i did not really have a plan or something (…) i had so much time to finish the course, i just worked on the course when it was convenient” (p1) or “i knew i would not be able to finish a module per week, so i tried to finish half a module per week, but that also did not work and i ended up just studying an hour now and then, without any clear structure”(p2). all participants finished their spoc successfully, suggesting that the courses are attainable also without a clear planning. however, participants’ learning might have benefitted from better planning and goal setting. in any case, trace data alone are insufficient to determine which students are in need of support for planning and goal setting. more information, for instance through interviews, is needed. comprehension monitoring and attention focusing. participants indicated that they engaged in note-taking not only to improve their remembrance of the information presented, but also to support comprehension monitoring and attention focusing. note-taking was thereby an important cognitive activity for participants. as p2 indicated “if you pause the video to take notes, and you don’t know where to start, than that is a moment of reflection; maybe you have to start over or continue the next day”. p1 explained “i easily get distracted, for instance by facebook, causing me to miss information. if i take notes, then i really have to pay attention to identify the main message”. the trace of pausing videos could be associated with taking notes. this trace may thus be interpreted not only as a cognitive activity (i.e., taking notes), but also as a metacognitive activity (e.g., comprehension monitoring) or attention focusing. the interview data allowed for the validation of the interpretation of pausing videos as a self-regulated learning activity. reflection on learning. the trace data indicated that participants often ended a session with an exam and the associated closing video, which may indicate reflection. the interviews provided evidence that participants ended after an exam on purpose “if i was working on the course, then i wanted to complete the whole module” (p4). participants were asked to what extent they reflected on their learning strategies and their progress, for instance at the end of a learning session. participants indicated hardly any reflection: “i think it was just a closing [video] indeed, and then it was simply done” (p3) and “and then [after the closing video] it was just done, than i was allowed to stop” (p4). based on the trace data alone, it could not be established of learners’ engaged in the srl activity reflection; it could only be established that they engaged in activities that may be associated with reflection. in the interview, learners could be directly asked about any reflection on their learning. the interviews provided clear information that the learners did not (consciously) engage in reflection. time management. participants repeatedly indicated problems with time management. participants did not orient themselves on the amount of work per module: “i found it difficult to estimate how long one module would take. so if i planned on doing modules 3 and 4 that day, i had no clue how much time that would cost me”(p3). the lack of orientation led to an underestimation of the amount of work involved: “i don’t know how long i expected that it would take. [on the website] it said that it would take 6 to 8 hours per week. but i thought ‘only videos, that can’t take that long” (p1). in addition, p4 indicated underestimating the time she would need to finish the course, as a result of her prior experience with studying, since “i thought that – as i am already enrolled in university – i know how to study, so i will need a few hours less than they say”. consequence of learners’ lack of planning and insufficient time management, was that they were stressed and pressed for time when finishing their course; “eventually you get stressed and that is a shame, because it is absolutely not necessary” (p3). while the trace data showed bursts in activity in de days before the final exam, this could not be interpreted as problematic. the interview data allows us to better understand learners’ time management and taught us that learners underestimate the amount of time needed. help seeking. participants hardly used the course forum, as already described in the section ‘lack of engagement with the course forum’. during the interviews, it became clear that learners overall hardly ever needed help. when asked if there was material that was hard to comprehend, learners responded with “no, i thought it was all just clearly designed” (p2) and “the theory was not complicated” (p3). there was limited external help seeking, as p3 indicated “i may have googled two terms when preparing for the exam, of which i thought ‘what was this again?”. however, it also became clear that participants sometimes did need help when they were unsure about the answer of an exam question. participants found different solutions, including re-watching videos, or skipping the material under the assumption that one would not need to understand everything in order to pass the course (p3 “i thought, if i understand seven out of eight [modules], then it must be alright.”). it was already indicated in the section on forum engagement that p2 preferred re-watching the videos over posting on the forum, and p4 indicated: “i would look back in my notes, or in the transcript, or in the video, in that order”. in sum, the trace data indicate very little engagement with the course forum, and the inference that learners are not searching for external help appears to be correct. however, learners do occasionally need help, and the interview results show a number of solutions employed by learners to seek help outside the course forum. environment structuring. participants consciously decided on where they studied for their spoc and they made use of the autonomy in study location offered by online education (“i appreciated that i could study everywhere. that you were not bound to a specific location”, p1). p1 and p2 studied in varying locations: “i sometimes worked on the course when i went home to my parents, sometimes in my student dorm room, and sometimes in the university library” (p1) and “i switched between the university library and at home. (…) and if i am in the train, i try to do something useful, so then i also regularly worked on it” (p2). p2 further clarified “if i studied in the train, then i often just watched videos, knowing that i would have to watch them again later (..) or i watched specific videos a second time”. studying in the train was thereby in addition to studying at home or the university library, not a replacement. p3 and p4 deliberately choose to study at home, as p3 remarked “i like having the option to study at home, since i can then make a nice cup of tea when i want to, and i am not surrounded by annoying people the way you are in the university library”. overall, the interview data provided clear information for all participants on where they studied, and the reasons underlying their choice of location. as environment structuring cannot be extracted from trace data, but is part of srl, the interviews helped get a more comprehensive measure of learners self-regulated learning. 4. discussion in this study, we used a mixed method approach combining trace data and interview data to study srl in online education. to be successful in online education, learners must be able to self-regulate their learning due to the autonomy offered to them (azevedo & aleven, 2013; beishuizen & steffens, 2011; broadbent, 2017; hew & cheung, 2014; wang et al., 2013). accurate measurement of learners’ srl is necessary to determine which students are in need of support and how to best offer support. trace data have been used previously to measure learners’ srl, since trace data can be gathered unobtrusively from all learners (e.g., cicchinelli et al., 2018; hadwin et al., 2007; kizilcec et al., 2017; van laer & elen, 2018). we identified two main problems with measuring srl with trace data, which we tried to solve by combining trace data with interviews in a mixed method study. first, the interviews helped to remove ambiguities in the interpretation of the trace data. for instance, all learners indicated that the short pauses in watching the content videos were used for note-taking, and that their viewing behaviour of the introduction and summary videos was markedly different (i.e., more laid-back) than that of the content videos. the interviews were thereby helpful for accurate interpretation of learners’ trace data. using interviews to interpret trace data could be a valuable addition to existing studies that aim to measure srl in online education. for instance, interview data could be useful in studies that correlate srl activities measured with questionnaires to sequences found in trace data, as reported for example in the study by kizilcec et al (2017). in that study, sequences of activities could be related to multiple srl constructs, thus leaving open the option for multiple srl interpretations. in such situations, interviews may be helpful to determine the correct srl interpretation of transitions for individual learners. interviews may also be useful to determine accurate indicators of srl activity in trace data. cicchinelli et al. (2018) defined trace data indicators for several srl activities. the frequency with which these indicators were present in learners’ trace data was used as a measure of learners’ engagement in srl activities during learning. in such situations, in which trace data is used to measure srl activity, it may be useful to conduct interviews before collecting trace data to support the selection of accurate trace data indicators of srl activity. the second problem identified was that learners’ behaviour outside of the learning environment is not captured in trace data. the interviews helped gain insight into learners’ srl activities that occur outside of the online learning environment. learners for instance indicated that they studied in different locations, and some learners’ study strategy also differed based on their study location (e.g., only re-watching videos in the train). the interviews showed that while some srl activities occur both inside and outside of the online learning environment (e.g., comprehension monitoring)., others solely take place outside of the online learning environment (e.g., environment structuring), this finding means that researchers might draw incorrect conclusions about how learners self-regulate their learning if they only rely on measurements of learners’ srl activities within the online learning environment. researchers may for instance conclude that learners who do not take self-test questions do not monitor their learning. however, learners may also engage in note-taking to monitor their comprehension of the key points; behaviour which is not visible in trace data directly. furthermore, the interviews showed that learners who are able to self-regulate their learning inside the online learning environment, may still struggle with srl outside of the online learning environment. it may then incorrectly be assumed that these learners are not in need of srl support. many learners however struggle with time management due to conflicting responsibilities in the rest of their lives (hew & cheung, 2014); an issue that is hard to measure with trace data only. measurement of learners’ srl activities outside of the online learning environment as we have done with interviews is thus important to obtain a complete overview of srl activities and to provide learners with adequate support. using trace data in combination with interview data in a mixed method study thus proved a useful approach for addressing the two identified problems with using trace data as a mono-method to measure learners’ srl. however, interviewing learners is much more time consuming and labour intensive than collecting trace data. the main benefits of trace data, unobtrusive collection and measurement of all learners, no longer apply when combining trace data with interviews. while we currently consider it important to combine trace data with another data source due to the problems described, we also acknowledge that combining trace data with interviews is not attainable at a large scale. fortunately, we also do not consider this necessary, as we see other potential solutions for the identified problems in the future. after repeated multi-method studies in which trace data are combined with interviews for larger, and more diverse samples, it will likely be possible to interpret part of the trace data unambiguously in terms of srl. the trace data variables that can be interpreted reliably in the future may then be used without combining them with another data source. for instance, answering test questions indicated comprehension monitoring for all learners in this pilot. if this interpretation is replicated, then quiz taking may be used as an indicator for the srl activity comprehension monitoring. researchers who are interested in the measurement of the srl behaviours that can be interpreted reliably from trace data, will thus be able to use trace data as a mono-method in the future. furthermore, in studies in which trace data is the only available data source, it is then clear what trace data can be interpreted validly and reliably. hereby, we argue, we will be able to reap the benefits (and simultaneously: be aware of the limitations) of unobtrusive, large-scale measurement of srl with trace data in the future. as learners’ behaviour outside of the online learning environment is not captured in trace data, we will never be able to measure this behaviour with trace data alone. repeated measurement of learners’ behaviour outside of the learning environment could help determine what srl activities are overlooked when solely measuring srl with trace data in online education. if trace data are then used as a mono-method in research, it will be clear what aspects of srl are not measured. additionally, the measurement of learners’ srl activities outside of the learning environment will allow for categorization of these activities if enough learners, from diverse samples, have been interviewed. it is therefore anticipated that a questionnaire to measure learners’ srl activities outside of the learning environment can be developed based on the united interviews findings of multiple studies. while questionnaires measuring learners’ srl outside of the learning environment already exist (e.g., barnard, lan, to, paton, & lai, 2009; jansen et al., 2017), these have been developed based on theoretical knowledge of srl and experiences of learners in offline learning environments. participant input to make sure the right srl behaviours are inquired in the questionnaire is therefore currently missing and would make a worthwhile addition to the questionnaires to increase their validity. while measurement of learners’ srl outside of the learning environment with a questionnaire is not unobtrusive, questionnaire data can be gathered on a large scale much more easily than interview data. 4.1. methodological reflection the results of this study have improved our understanding of learners’ activities in spocs. by focusing the interviews on both participants’ reflections on (1) the trace data itself and (2) their srl outside of the learning environment, we were able to better understand both aspects of learners’ activities. in order to measure both the micro and macro aspects of learners’ activities, we made use of interviews, since they allow asking questions on a range of different granularities. we feel this is valuable, as it provided a richer understanding of learner behaviour. however, as it was the first study of its kind within this domain, a reflection on the methodology used is in order. when combining trace data with interview data, three issues that need specific attention can be identified. the first issue concerns the selection of participants. as is the case in all research, the sample included influences the results found and the generalizability of those results. the interviews conducted in the current study formed a pilot sample to show the benefits of the presented methodology. if researchers want to use our proposed methodology to draw conclusion and add to theory development, they should attempt to obtain a larger sample. the current sample of four interviewees was rather homogeneous in characteristics, yet the interviewees already differed in their reported srl. this diversity indicates that a much larger, and more diverse, sample is necessary to determine valid interpretations of learners’ behaviour. likely, an iterative approach of interviewing and analysing data is necessary to reach data saturation (morse, 2010; strauss & corbin, 1994), especially when one is interested in measuring and understanding the srl of a more diverse group of learners. since trace data provide insight in learners’ learning process, analysis of the trace data of learners may be helpful for purposive sampling (robinson, 2014). trace data may be used to identify a diverse set of learners to be interviewed or to identify specific cases one is interested in, such as learners who quit the course early or who finished the complete course. purposive sampling may thus contribute to reach data saturation. the second methodological issue concerns the time span between the learning activities and the interview. interviewing students early, potentially even during the course they are attending, results in richer data that is likely more reliable than verbal reporting collected later. retrospective reporting may lead to both errors of omission (forgetting srl activities) as well as to errors of fabricating (reporting srl activities that did not take place, or were not as deliberate as they are reported; russo et al., 1989). nevertheless, waiting with selecting interviewees until the course is over as happened in this study also has two advantages. first, learners’ activities in the course are then not influenced by the interview, which may be the case when interviewing during the course (reactive invalidity; russo et al., 1989). second, trace data for the full course can then be used to select an adequate sample. for instance, learners that quit the course cannot be selected early on, as they may still continue learning in the spoc or mooc (much) later. in this pilot, the time between studying and the exam on the one hand, and the interview on the other hand, was two months. this likely hampered learners’ recall of their activities and the reasons underlying their activities. the transition diagrams we made for every participant alleviated this problem to some extent, as they helped interviewees remember their learning activities (i.e., cued retrospective recall). however, we acknowledge that approaching potential interviewees at an earlier point in time would likely have resulted in richer data. the third methodological consideration influencing our findings, is the presentation of trace data to participants. the trace data visualizations are used as a cue for cued retrospective reporting. the formatting of the cue, and the information that is and is not included, likely influences learners’ reporting. researchers should therefore carefully decide on the formatting of their trace data cue and be aware of the consequences of their decisions. trace data visualizations in the form of transition diagrams, as they were used in the current study, helped learners to remember all activities present in the course. they furthermore helped learners to explain both the linearity in their learning behaviour as well as the deviations in this linearity. however, the transition diagrams are limited by the fact that if multiple arrows move out of an activity, one cannot determine which one of the out arrows should be followed at which moment in time. for instance, in figure 5 the learner may have posted on the forum every time after browsing the forum, however, the learner may also have posted only once in response to browsing. this difference is not visible in the static transition diagram. in this study, the listing of all activities was used as supportive information during the interviews. a dynamic transition diagram, in which arrows are added over time, thereby including temporal information, could solve this problem. 4.2. conclusion we have presented a methodological approach to measuring srl in online education by combining trace data with interview data. while there is surely room for improvement, for which we have provided suggestions in this discussion, the results of this study show the potential of this combination of methodologies. currently, trace data are already interpreted in terms of srl (e.g., cicchinelli et al., 2018; fincham et al., 2018; maldonado-mahauad et al., 2018). however, the interview results show that interpretation of trace data is often ambiguous. more research of the kind presented here should be conducted to establish what components of trace data can be validly labelled, and how they should be labelled, and what components of trace data will remain ambiguous (rovers et al., 2019; van laer & elen, 2018). to achieve this aim of enhancing the validity of srl measurement, further research could make use not only of trace data and interview data as presented in the current study, but could incorporate promising data collection methods such as eye-tracking (salmeron, gil, & bråten, 2018; trevors, feyzi-behnagh, azevedo, & bouchet, 2016). next to supporting the interpretation of the trace data, the interview data also provided empirical evidence that not all srl activity is visible in the trace data. such srl activities are overlooked when trace data is used as the sole data source to measure srl. the methodology presented here could be used to reliably establish which aspects of srl can and which aspects of srl cannot be measured with trace data. more research into learners’ srl activities outside of the online learning environment will furthermore enable the development of a questionnaire to measure learners’ srl activities that are not captured in the trace data. hopefully, in the future, we can reliably and validly measure learners srl with unambiguous interpretations of unobtrusively collected trace data from all learners, in combination with large scale measurement of learners’ srl outside of the learning environment with a valid questionnaire. keypoints the self-regulated learning of students in online education is often measured with trace data. the interpretation of trace data is often ambiguous. self-regulated learning activities outside of the learning environment are not captured with trace data. it is shown that a mixed method approach can help resolve the issues surrounding the measurement of srl with trace data. interviews with cued retrospective recall enable correct interpretation of trace data. references allen, i. e., & seaman, j. (2014). grade change: tracking online education in the united states. allen, i. e., & seaman, j. (2016). online report card: tracking online education in the united states . alshenqeeti, h. (2014). interviewing as a data collection method: a critical review. english linguistics research, 3(1). https://doi.org/10.5430/elr.v3n1p39 azevedo, r., & aleven, v. (2013). metacognition and learning technologies: an overview of current interdisciplinary research. in r. azevedo & v. aleven (eds.), international handbook of metacognition and learning technologies (vol. 28, pp. 1–16). new york, ny: springer new york. retrieved from http://link.springer.com/10.1007/978-1-4419-5546-3_1 azevedo, r., & cromley, j. g. (2004). does training on self-regulated learning facilitate students’ learning with hypermedia? journal of educational psychology, 96, 523–535. https://doi.org/10.1037/00220663.96.3.523 barnard, l., lan, w. y., to, y. m., paton, v. o., & lai, s.-l. (2009). measuring self-regulation in online and blended learning environments. the internet and higher education, 12(1), 1–6. https://doi.org/10.1016/j.iheduc.2008.10.005 beishuizen, j., & steffens, k. (2011). a conceptual framework for research on self-regulated learning. in r. carneiro, p. lefrere, k. steffens, & j. underwood (eds.), self-regulated learning in technology enhanced learning environments (pp. 3–19). rotterdam, the netherlands: sense publishers. ben-eliyahu, a., & linnenbrink-garcia, l. (2015). integrating the regulation of affect, behaviour, and cognition into self-regulated learning paradigms among secondary and post-secondary students. metacognition and learning, 10, 15–42. https://doi.org/10.1007/s11409-014-9129-8 broadbent, j. (2017). comparing online and blended learner’s self-regulated learning strategies and academic performance. the internet and higher education, 33, 24–32. https://doi.org/10.1016/j.iheduc.2017.01.004 cicchinelli, a., veas, e., pardo, a., pammer-schindler, v., fessl, a., barreiros, c., & lindstädt, s. (2018). finding traces of self-regulated learning in activity streams (pp. 191–200). acm press. https://doi.org/10.1145/3170358.3170381 creswell, j. w. (2008). mixed method designs. in educational research: planning, conducting, and evaluating quantitative and qualitative research (pp. 551–595). new jersey, nj: pearson education international. dicicco-bloom, b., & crabtree, b. f. (2006). the qualitative research interview. medical education, 40(4), 314–321. https://doi.org/10.1111/j.1365-2929.2006.02418.x ericsson, k. a., & simon, h. a. (1993). protocol analysis. cambridge, ma: mit press. fincham, o. e., gasevic, d. v., jovanovic, j. m., & pardo, a. (2018). from study tactics to learning strategies: an analytical method for extracting interpretable representations. ieee transactions on learning technologies, 1–1. https://doi.org/10.1109/tlt.2018.2823317 hadwin, a. f., nesbit, j. c., jamieson-noel, d., code, j., & winne, p. h. (2007). examining trace data to explore self-regulated learning. metacognition and learning, 2, 107–124. https://doi.org/10.1007/s11409-007-9016-7 hew, k. f., & cheung, w. s. (2014). students’ and instructors’ use of massive open online courses (moocs): motivations and challenges. educational research review, 12, 45–58. https://doi.org/10.1016/j.edurev.2014.05.001 howard-rose, d., & winne, p. h. (1993). measuring component and sets of cognitive processes in selfregulated learning. journal of educational psychology, 85(4), 591–604. https://doi.org/10.1037/0022-0663.85.4.591 jansen, r. s., van leeuwen, a., janssen, j., kester, l., & kalz, m. (2017). validation of the self-regulated online learning questionnaire. journal of computing in higher education, 29, 6–27. https://doi.org/10.1007/s12528-016-9125-x johnson, r. b., & onwuegbuzie, a. j. (2004). mixed methods research: a research paradigm whose time has come. educational researcher, 33(7), 14–26. https://doi.org/10.3102/0013189x033007014 jovanović, j., gašević, d., dawson, s., pardo, a., & mirriahi, n. (2017). learning analytics to unveil learning strategies in a flipped classroom. the internet and higher education, 33, 74–85. https://doi.org/10.1016/j.iheduc.2017.02.001 karabenick, s. a., & zusho, a. (2015). examining approaches to research on self-regulated learning: conceptual and methodological considerations. metacognition and learning, 10(1), 151–163. https://doi.org/10.1007/s11409-015-9137-3 kizilcec, r. f., pérez-sanagustín, m., & maldonado, j. j. (2017). self-regulated learning strategies predict learner behaviour and goal attainment in massive open online courses. computers & education, 104, 18–33. https://doi.org/10.1016/j.compedu.2016.10.001 littlejohn, a., hood, n., milligan, c., & mustain, p. (2016). learning in moocs: motivations and selfregulated learning in moocs. the internet and higher education, 29, 40–48. https://doi.org/10.1016/j.iheduc.2015.12.003 maldonado-mahauad, j., pérez-sanagustín, m., kizilcec, r. f., morales, n., & munoz-gama, j. (2018). mining theory-based patterns from big data: identifying self-regulated learning strategies in massive open online courses. computers in human behaviour, 80, 179–196. https://doi.org/10.1016/j.chb.2017.11.011 mcgrath, j. e., martin, j., & kulka, r. a. (1981). some quasi-rules for making judgment calls in research. american behavioural scientist, 25(2), 211–224. https://doi.org/10.1177/000276428102500206 mcintosh, m. j., & morse, j. m. (2015). situating and constructing diversity in semi-structured interviews. global qualitative nursing research, 2, 233339361559767. https://doi.org/10.1177/2333393615597674 min, l., & jingyan, l. (2017). assessing the effectiveness of self-regulated learning in moocs using macro-level behavioural sequence data. in proceedings of emoocs 2017 (pp. 1–9). madrid, spain. morse, j. m. (2010). sampling in grounded theory. in a. bryant & k. charmaz (eds.), the sage handbook of grounded theory (pp. 229–244). sage publications. nelson, t. o., & narens, l. (1990). metamemory: a theoretical framework and new findings. in psychology of learning and motivation (vol. 26, pp. 125–173). academic press. panadero, e. (2017). a review of self-regulated learning: six models and four directions for research. frontiers in psychology, 8. https://doi.org/10.3389/fpsyg.2017.00422 phillips, r., maor, d., cumming-potvin, w., roberts, p., herrington, j., preston, g., … perry, l. (2011). learning analytics and study behaviour: a pilot study. presented at the australasian society for computers in learning in tertiary education, hobart, tasmania, australia. puustinen, m., & pulkkinen, l. (2001). models of self-regulated learning: a review. scandinavian journal of educational research, 45, 269–286. https://doi.org/10.1080/00313830120074206 reimann, p., markauskaite, l., & bannert, m. (2014). e-research and learning theory: what do sequence and process mining methods contribute?: e-research and learning theory. british journal of educational technology, 45(3), 528–540. https://doi.org/10.1111/bjet.12146 robinson, o. c. (2014). sampling in interview-based qualitative research: a theoretical and practical guide. qualitative research in psychology, 11(1), 25–41. https://doi.org/10.1080/14780887.2013.801543 rovers, s. f. e., clarebout, g., savelberg, h. h. c. m., de bruin, a. b. h., & van merriënboer, j. j. g. (2019). granularity matters: comparing different ways of measuring self-regulated learning. metacognition and learning. https://doi.org/10.1007/s11409-019-09188-6 russo, j. e., johnson, e. j., & stephens, d. l. (1989). the validity of verbal protocols. memory & cognition, 17(6), 759–769. https://doi.org/10.3758/bf03202637 salmeron, l., gil, l., & bråten, i. (2018). using eye-tracking to assess sourcing during multiple document reading: a critical analysis. frontline learning research, 6(3), 105-122. https://doi.org/10.14786/flr.v6i3.368 schraw, g. (2010). measuring self-regulation in computer-based learning environments. educational psychologist, 45(4), 258–266. https://doi.org/10.1080/00461520.2010.515936 strauss, a., & corbin, j. (1994). grounded theory methodology: an overview. in n. k. denzin & y. s. lincoln (eds.) (vol. 17, pp. 273–285). thousand oaks, ca, us: sage publications. trevors, g., feyzi-behnagh, r., azevedo, r., & bouchet, f. (2016). self-regulated learning processes vary as a function of epistemic beliefs and contexts: mixed method evidence from eye tracking and concurrent and retrospective reports. learning and instruction, 42 , 31-46. https://doi.org/10.1016/j.learninstruc.2015.11.003 van gog, t., paas, f., van merriënboer, j. j. g., & witte, p. (2005). uncovering the problem-solving process: cued retrospective reporting versus concurrent and retrospective reporting. journal of experimental psychology: applied, 11(4), 237–244. https://doi.org/10.1037/1076-898x.11.4.237 van laer, s., & elen, j. (2018). towards a methodological framework for sequence analysis in the field of self-regulated learning. frontline learning research, 6(3), 228–249. veletsianos, g., reich, j., & pasquini, l. a. (2016). the life between big data log events: learners strategies to overcome challenges in moocs. aera open, 2(3). https://doi.org/10.1177/2332858416657002 wang, c.-h., shannon, d. m., & ross, m. e. (2013). students’ characteristics, self-regulated learning, technology self-efficacy, and course outcomes in online learning. distance education, 34, 302–323. https://doi.org/10.1080/01587919.2013.835779 winne, p. h. (2010). improving measurements of self-regulated learning. educational psychologist, 45(4), 267–276. https://doi.org/10.1080/00461520.2010.517150 winne, p. h. (2014). issues in researching self-regulated learning as patterns of events. metacognition and learning, 9(2), 229–237. https://doi.org/10.1007/s11409-014-9113-3 winne, p. h., & jamieson-noel, d. (2002). exploring students’ calibration of self reports about study tactics and achievement. contemporary educational psychology, 27(4), 551–572. https://doi.org/10.1016/s0361-476x(02)00006-1 zimmerman, b. j. (1986). becoming a self-regulated learner: which are the key subprocesses? contemporary educational psychology, 11 (4), 307–313. https://doi.org/10.1016/0361 476x(86)90027-5 appendix a: interview guideline 5-10 min start-up introduction / welcome. explanation of the interview goals. we will be talking about your learning experience when studying for the course xxx. we will be talking about how you organized your learning, the way you studied, the barriers you encountered, and potential solutions or improvements. the interview will take approximately 45 minutes to one hour. the interview is video recorded. the recordings will be treated confidentially. general outline. we will first talk about your motivations for taking the course, then we will look at your studying behavior in the online course environment. we will continue with discussing the overall learning process. we will then wrap-up by discussion solutions and improvements. 15 – 20 min micro process activities you as a user engage in on edx while you are learning in the course are stored on the server. for instance, when you click to play a video, when you pause a video, or when you answer a recap or test question. with this data, a flow chart can be created of a learner’s online behavior. << show flow chart of another learner in another course >> transitions from one activity to another are drawn to display the steps the learner went through while learning. << show blanc flow chart of the course the learner took; thus without transitions >> this is the outline of the course you took. explanation of the elements. it is understandable if you do not remember everything in detail. please answer the questions to the best of your ability, but also feel free to indicate that you don’t remember. << present the flow chart with the learner’s transitions indicated >> this is the flow chart showing the way you worked with the course materials. explanation of the general pattern of activities visible in the flow chart. 20 min macro process after discussing your learning process in the online course environment, i would now like to talk to you about your learning for the course in general. 10 min improvement suggestions finally, i want to present you with some suggestions to improve spoc learning. i would like to hear your thoughts about them. pop-up when you log in asking you what your plans are. email once or twice a week with study advice and/or motivational message. videos in the course with study tips; especially for srl. assignments based on how to study (or recap questions about srl). 5 min wrap-up / debriefing explanation of the interview purpose in case the interviewee is interested in that information. explanation of the results that can be provided. allow the interviewee the opportunity to check the interview transcript. appendix b: coding guideline frontline learning research vol. 10 no. 2 (2022) 64 85 issn 2295-3159 corresponding author: mari murtonen, assistentinkatu 5, 20500 turku, finland, mari.murtonen@utu.fi doi:https://doi.org/10.14786/flr.v10i2.1031 university teachers’ focus on students: examining the relationships between visual attention, conceptions of teaching and pedagogical training mari murtonena, erkki antoa, eero laakkonena & henna vilppua auniversity of turku, finland article received 21 january 2022 / revised 9 december 2022/ accepted 9 december 2022 / available online 11 january 2023 abstract teachers’ focus on their students’ learning is considered central in high-quality, studentcentred university teaching. this frontline eye-movement research asks whether teachers’ focus can be observed at the intersection of the visual and conceptual levels. it introduces a novel way to study teachers’ visual attention combined with verbal interpretations, including numerical ratings of the success of teaching when they observe teaching situations. teachers’ visual attention and interpretations were further studied in connection to their prior pedagogical training and teaching experience in years. two short videos depicting teaching during a lecture, including different types of trigger events, were presented to teachers (n = 49) who were asked to think aloud while watching. the first video’s trigger was students becoming bored during a content-focused teaching situation, and the second video’s trigger was the teacher replying in an engaging way to students’ questions in a learning-focused teaching situation. the results showed that pedagogically trained teachers paid more visual attention to the students than did their non-trained colleagues, especially in content-focused teaching situations. teaching experience did not have any effect on visual attention or interpretation in this study. the teachers who paid more visual attention to the students in the content-focused teaching situation noticed in their interpretations that the students were not active, expressed higher learning-facilitating teaching conceptions and gave lower numerical ratings for the teaching situation. in conclusion, pedagogical training seems to promote university teachers’ ability to pay visual attention to students in teaching situations and interpret these situations from the students’ perspective, i.e. focus on student learning. keywords: visual attention, conceptions of teaching, university pedagogical training, facilitation of learning, eye tracking mailto:mari.murtonen@utu.fi murtonen et al 65 | f l r introduction to achieve better learning outcomes, university teachers are expected to focus on their students’ learning to be able to support it instead of merely focusing on delivering the content to them (prosser & trigwell, 2014; vilppu et al., 2019). focusing on students’ learning is often used as a synonym for good teaching that acknowledges and answers students’ needs to foster learning. learning-focused or student-centred teaching is often called for; however, it is not yet very clear what focusing on students’ learning means for teaching. focusing on students’ learning can take place at many levels, starting with the curriculum and building the environment and courses such that they support students’ learning actions (entwistle, 2005). however, what it means in a teaching situation in which a teacher and students are present remains unclear. questions such as how university teachers monitor and gain knowledge about their students during a teaching situation, e.g. a lecture, to guide their teaching actions have remained unanswered. teachers’ readiness to facilitate university students’ learning has been studied in terms of their conceptions of teaching and learning (samuelowicz & bain, 2001) as well as their approaches to teaching (prosser & trigwell, 2014). a relationship between teachers’ and students’ approaches has been found (gibbs & coffey, 2004), indicating that teachers’ approaches do have an effect on their teaching and on their students’ learning. the studies on university teachers’ conceptions and approaches use the methods of self-report questionnaires and interviews, which reveal only some aspects of teaching, such as their intentions (trigwell & prosser, 2004) or their underlying orientations and beliefs (samuelowicz & bain, 2001). more knowledge is needed about how teachers perceive, interpret and make decisions in certain teaching situations (blömeke et al., 2015). in secondary school settings, eye-movement studies have revealed interesting information about expert teachers’ perceptions compared to novices. for example, experts tend to look longer at students (mcintyre et al., 2017) and focus their attention on areas where relevant information is available (wolff et al., 2016). in general, previous studies have claimed that novice teachers are not able to focus on students’ learning as deeply as more experienced teachers (e.g. levin et al., 2009), and they may use only bottom-up visual noticing instead of knowledge-based top-down processes that allow shifting of attention from attention-capturing events to pedagogically meaningful events (e.g. theeuwes, 2000). we lack this type of information concerning university teaching, that is, how more experienced or better educated teachers differ from novices and on what levels, besides conceptions and approaches, differences may occur. the teaching situations in higher education are different from those in secondary school classrooms; thus, research on higher education settings is needed. this study aimed to discover whether focusing on students can be observed at the visual level and whether visual attention is related to different teaching conceptions. by using eye-tracking measurements and retrospective think-aloud, we investigated how university teachers perceive and interpret different kinds of teaching situations. in addition, the effects of prior pedagogical training and teaching experience were studied. we begin the paper by describing university teachers’ expertise requirements and the teaching context and move on to consider what is meant by learning-focused teaching at the university level. we then proceed to the question of focus on the visual level, i.e. visual attention and finally propose video cases as a method to gain deeper insight into university teaching. teachers’ pedagogical expertise in the university context teachers’ professional learning of pedagogy, i.e. the development of their pedagogical expertise, can be understood as a complex process whereby changes in knowledge, orientation and skills pertain to one’s conception of teaching and actions as a teacher (garner & kaplan, 2019). these changes often require a change in the teacher’s identity as well. the teacher’s core identity has traditionally been defined as a subject expert murtonen et al 66 | f l r who transmits subject knowledge to students, while contemporary views of teaching highlight the role of the teacher as a learning process expert who fosters active, self-regulated and collaborated learning in students (vermunt et al., 2017). university teachers are typically highly educated experts in their own subject domain, but they often lack pedagogical qualifications, unlike their colleagues in primary and secondary schools. thus, although university teachers excel in the content knowledge of their own discipline, they may lack pedagogical knowledge. in addition, they may lack pedagogical content knowledge, which refers to how pedagogical knowledge can be implemented in their own disciplinary areas (shulman, 1987). it is problematic that university teachers have no pedagogical training, since, according to expertise research, excelling in one’s own disciplinary subject does not necessarily make one an expert in teaching the subject (e.g. ericsson, 2008). knight (2002) argued that without pedagogical training, it is typical for university teachers to adopt their own teachers’ teaching style, even though they know it might not be the best way to promote student learning. persuading teachers who have been teaching at university for a long time to take part in pedagogical training can be difficult. in addition, teachers with extensive teaching experience can be reluctant to change their teaching conceptions and practices (postareff & nevgi, 2015). in contrast, pedagogical training at the beginning of a university teaching career can be very effective (e.g. vilppu et al., 2019). thus, teachers’ pedagogical expertise levels may vary greatly in the university context. the university teaching environment is unique and different from those of primary and secondary schools. according to doyle (2006), a school classroom situation includes features such as a large quantity of events and tasks taking place multidimensionally and simultaneously (e.g. interruptions and other unpredictable situations that require immediate attention). it also includes a common set of experiences that form a history for the class and have an impact on future events. compared to this, a traditional university lecture could be described as a more unidirectional situation, with the teacher lecturing and usually no surprises occurring; in addition, it likely includes certain norms and traditions concerning the university culture for students about how to behave in a lecture. at the university, students and the place where teaching takes place may be different in every lecture, and the teacher often has neither a common history with the students nor personal contact with them. this creates a unique environment in which the research results from other educational levels cannot be directly applied. focusing on students at the level of conceptions and approaches university teachers’ pedagogical expertise has mostly been studied from the perspective of their conceptions of and approaches to teaching. university teachers’ conceptions of teaching have been found to vary between teaching as facilitating learning and teaching as transmitting knowledge (kember & kwan, 2000; samuelowicz & bain, 2001). the former conception includes the idea that the most important task in teaching is to support students’ learning processes and to create learning environments that ‘scaffold’ learning, whereas the transmission conception indicates that the most important task in teaching is to deliver information to students. teachers’ approaches to their teaching, i.e. the strategies they adopt, have been categorised into learning-focused and content-focused approaches (postareff & lindblom-ylänne, 2008). the learning-focused approach refers to teaching strategies in which the teacher’s aim is to foster students’ deep learning processes by activating their knowledge construction. in contrast, with the content-focused approach, the teacher’s intention is to transmit knowledge to students without attempting to activate them. a rather high correspondence between conceptions and approaches seems to exist. teachers who consider teaching as transmitting knowledge tend to adopt a content-focused approach to teaching, whereas teachers murtonen et al 67 | f l r who view teaching as supporting students in building their own understanding more often adopt a learningfocused approach to teaching (kember & kwan, 2000). teachers’ approaches to teaching have been shown to relate to their students’ approaches to learning (gibbs & coffey, 2004; prosser & trigwell, 2014; uiboleht et al., 2018), indicating that a learning-focused approach to teaching encourages the use of a deep approach to learning. students’ adoption of a deep approach to learning seems to indicate that they will achieve higherquality learning outcomes (uiboleht et al., 2018). thus, teachers’ actions and conceptions seem to have important effects on the success of student learning. focusing on student learning and being able to support their learning process also requires skills other than understanding how learning happens, along with an intention to support it. for example, engaging students in lectures has been shown to be an important medium for focusing on students and fostering their active learning (lonka & ketonen, 2012). many university teachers still use the traditional lecturing approach, with the dominant view of teaching as the transmission of knowledge, probably because they teach the way they were taught (knight, 2002). this unidirectional method of lecturing does not give the lecturer much information about students’ learning. teaching methods that engage students are seen as central in giving teachers information about students’ prior knowledge, goals and motivation in studying and motivating students to learn in more depth (lonka & ketonen, 2012). to apply engaging teaching methods, a teacher needs to be sensitive to students’ nonverbal messages in a teaching situation. to be able to monitor, react to different situations and support students’ learning processes, teachers need pedagogical knowledge and skills that guide their own teaching. according to expertise studies, a skill needs to be deliberately practiced to develop (ericsson, 2008). thus, practicing teaching without deliberate training probably does not help university teachers develop; they need pedagogical training to acquire pedagogical expertise. current pedagogical training for university teachers aims to facilitate conceptions of teaching that enhance learning and learning-focused approaches to teaching, and there is evidence that this training has a positive effect on teachers’ conceptions and approaches to teaching (e.g. postareff et al., 2007; stes & van petegem, 2011). focusing on students at the visual level and noticing important events in addition to university teachers’ focusing on their students’ prior knowledge, intentions, goals and study progress, teachers need visual information about their students during a teaching session to be sensitive to their nonverbal messages concerning their learning. focusing on students at the visual level means paying visual attention to students. eye-movement studies offer information about where viewers focus their attention and how they process classroom situations when observing teaching (wolff et al., 2016). eye movements as such are not sufficient to provide information about teachers’ thinking, since there is only a hypothesis about the connection between eye and mind (e.g. just & carpenter, 1980), but when combined with teachers’ verbal interpretations, they can help us to understand teachers’ thoughts. while focusing means deliberately paying attention to the overall situation, noticing means that one actually perceives an important event when it happens. noticing refers to the ability to focus attention on events that are pertinent to teaching and learning (grub et al., 2020), and knowledge-based reasoning implies the ability to apply knowledge about teaching and learning to interpret these events as well as the ability to draw relevant conclusions. lachner et al. (2016) argued that the skills in both noticing and interpreting are knowledge-based in that teachers’ knowledge guides their attention and interpretation of crucial events. compared to novices, expert teachers possess more extensive, elaborate and coherently organised knowledge structures. through teaching experience, teachers integrate formal professional knowledge with their personal murtonen et al 68 | f l r and practical knowledge, thus strengthening their ability to perform effectively (wolff et al., 2021). the differences between experienced and novice teachers’ visual processing of classroom information have mainly been studied at the primary (e.g. pouta et al., 2020) and secondary school levels (e.g. mcintyre et al., 2017; stahnke & blömeke, 2021; wolff et al., 2016). the level of expertise has been shown to influence the noticing and interpretation of classroom events. for example, expert teachers’ noticing is efficient (mcintyre et al., 2017) and knowledge-based, covers wide areas (wolff et al., 2016) and is focused on students (e.g. van den bogert et al., 2014; stahnke & blömeke, 2021). novices, on the other hand, tend to engage in a more timeconsuming and rather indiscriminate search for information (e.g. wolff et al., 2016). furthermore, with regard to knowledge-based reasoning, novices tend just to describe classroom events, while experts explain and integrate the meaning behind what they see (e.g. wolff et al., 2017). a novice may notice only so-called bottom-up events that capture visual attention, while experts use knowledge-based top-down processes that allow them to shift their attention from attention-capturing events to pedagogically meaningful events (e.g. theeuwes, 2000). utilising videos in studying university teachers’ focus on students during the last decade, video-based assessment has become frequent in both teacher training and teacher training research (dunekacke et al., 2015; gaudin & chaliés, 2015), and many studies focusing on visual processes utilise classroom videos. there are many advantages to using video assessment: it provides a standardised measurement, it is close to the complex reality of pedagogical situations, and, due to this perceived authenticity, it is usually considered motivating and highly accepted by the participants. further, compared to written or still picture cases, videos can integrate both verbal and nonverbal information, such as facial expressions, gestures, movements, postures and even emotional states. scripted videos also enable the inclusion of trigger events (könig et al., 2014), i.e. pedagogically meaningful events, which the teacher should notice in order to be successful in learning-focused teaching. as a research method, video assessments may also avoid problems commonly related to self-report measures, such as those inherent in likert scale questionnaires and interviews, which rely on self-perception and are thus prone to credibility issues (see vilppu et al., 2019). higher education teachers’ focus on students has usually been studied through their conceptions and approaches with fairly traditional self-report measures, such as questionnaires and interviews (e.g. postareff & lindblom-ylänne, 2008; trigwell & prosser, 2004), which do not necessarily measure actual teaching practices, but aims and beliefs concerning them. thus, new methodological perspectives and tools would allow for new knowledge of teachers’ pedagogical expertise. considering the varied and constantly changing university teaching situations and differing backgrounds of university teachers and students, analysing practical teaching situations to obtain general knowledge of how teachers’ teaching actions correspond to their conceptions would be challenging. therefore, viewing and interpreting videotaped teaching situations could offer a methodology for approaching this question. recently, many studies have integrated eye-tracking methodology into video viewing (e.g. wolff et al., 2016; wyss et al., 2020), and thus, have also focused on visual processes. the viewer’s attention is central to how classroom situations are visually processed, where viewers’ eye movements offer insight (wolff et al., 2016). because of the link between where the eyes are gazing and what the mind is engaged with (the eye–mind hypothesis; see just & carpenter, 1980), viewers’ eye fixation patterns can be used to investigate their ongoing mental processes during viewing. however, eye tracking has its limitations, and as such, it does not necessarily relate anything about how the viewer comprehends the scene. thus, eye-movement data require many murtonen et al 69 | f l r inferences about the underlying cognitive processes, since they do not explain why a viewer was looking at certain representations (van gog et al., 2009). to reduce the number of researchers’ inferences, complementary methods, such as concurrent or retrospective reporting, are utilised alongside eye movements. present study this study aimed to extend knowledge about university teachers’ focus on their students by examining whether the focus can be observed at the visual level in addition to the conceptual level. we studied teachers’ conceptions of teaching with regard to their visual attention, noticing the trigger event and rating the success of the observed teaching. furthermore, the effects of pedagogical training and teaching experience on visual attention and teaching conceptions were examined. comparisons were made between pedagogically trained vs. untrained and novice vs. more experienced teachers. the research questions of the study were as follows: 1) to what extent do university teachers pay visual attention to students in comparison with the other two central elements of a lecture, the teacher and the slides, when watching videotaped teaching situations with inbuilt trigger events? 2) do teachers notice inbuilt trigger events in videos by paying attention to students at the intersection of the visual level and teaching conceptions? are these in line with their numerical ratings of the success of teaching situations? 3) how are teachers’ prior pedagogical training and the length of their teaching experience connected with their visual attention to students and their conceptions of teaching? since pedagogical training aims to foster learning-facilitating conception of teaching, we assumed that pedagogically trained teachers’ interpretations of the videos would reflect a stronger learning-facilitation conception of teaching than their untrained colleagues, due to their more sophisticated knowledge base (lachner et al., 2016). the main hypothesis for this study was that pedagogically trained teachers would pay more visual attention to students than their untrained colleagues would, especially in a situation where topdown processes would be needed to shift teachers’ attention to important phenomena (theeuwes, 2000). furthermore, we used numerical ratings as evaluations given by teachers on teaching situations to confirm that our analysis of their interpretation was correct. thus, the ratings needed to be in alignment with the interpretations. the role of teaching experience might be more ambiguous among university teachers than among primary and secondary school teachers, who are all pedagogically trained. work experience alone does not help people develop their expertise, but deliberate practices such as training are needed to gain high-level skills (ericsson, 2008). from this standpoint, we assumed that the length of previous teaching experience would not strongly differentiate teachers in their visual attention and conceptions of teaching but that previous pedagogical training would promote paying more attention to students’ learning both visually and verbally (see figure 1). murtonen et al 70 | f l r figure 1. hypothesised model of the connections between visual attention and conceptions of teaching among university teachers methods participants the target group of the study comprised university teachers and doctoral students who either had or did not yet have teaching tasks at the university. they had all applied for voluntary university pedagogy training (n = 51). the measurements took place in the beginning of the training. watching the video vignettes was part of the training, but the trainees could choose whether they wanted to take part in the study. thus, participation in the study was voluntary, and informed consent was obtained from the participants. ethical approval for the study was granted by the ethics committee for human sciences of the target university. as two members of the target group declined to participate, the response rate was 94%. thus, the number of participants was 49. twenty (42%) participants had earlier pedagogical training, which varied from a university pedagogy course bearing 1 study credit (according to the european credit transfer and accumulation system, ects) to a subject teacher degree bearing 60 ects, whereas the rest had no previous pedagogical training. the participants represented seven faculties. their teaching experience varied, but most of them had been teaching at the university for less than 10 years: nine (19%) had no teaching experience, 14 (29%) had a maximum of 2 years’ teaching experience, 13 (27%) had been teaching for 2 to 5 years, 10 (21%) for 5 to 10 years, and two (4%) had been teaching for over 10 years in at least one course per academic year. information on one participant’s faculty and teaching experience was missing from the data. those doctoral students who had no teaching experience at the university were considered prospective teachers who might be given teaching duties in the near future. due to the consistency of the sample, teachers were divided into novices (n = 23, with no teaching experience or a maximum of 2 years) and more experienced teachers (n = 25, with more than 2 years of teaching experience) for further analysis. apparatus and materials a tobii tx300 eye tracker (tobii technology, inc., falls church, va, usa) was used to collect the participants’ eye movements. the eye-tracking component was integrated into a 23-inch high-resolution monitor, with a maximum resolution of 1920 × 1080 pixels. the eye-tracking camera sampled data binocularly at a rate of 300 hz, with a reported gaze accuracy of 0.4°. to ensure that participants were as comfortable as possible while watching the video vignettes, no supporting chinrest was used, since the eye tracker allowed even large head movements. two custom-made videos were used in the study, with actors as teachers and students. the videos were designed and filmed by the researchers of this paper, who were familiar with the local lecturing culture. both videos shared the same simple layout, and they were filmed from the same angle and from an outsider’s perspective of the classroom. in both videos, there was a scene in which students were sitting on the left, the teacher was standing in the middle and the screen was on the right (see figure 2). the first video was 1 minute 33 seconds in duration, and the second was 1 minute 36 seconds; both depicted a situation in the middle of a lecture. to focus on the targeted constructs, the videos were scripted (könig et al., 2014) by a group of experienced university pedagogy educators and researchers. they aimed to represent typical and realistic university murtonen et al 71 | f l r teaching–learning situations, since the perceived authenticity of the video material was important (seidel et al., 2011). research methodology was chosen as the topic of teaching for both videos, since it was considered to be quite domain-general, neutral and equally understandable for teachers from different disciplines. the videos were filmed from the side to allow the three important elements of the setting (teacher, student and slides) to be clearly visible. only a few students were shown on the video to reduce the number of spontaneous movements that could draw observers’ attention. both videos incorporated a pedagogically interesting situation, a so-called trigger event (see also vilppu et al., 2019). we expected these built-in pedagogical events to trigger certain reactions and interpretations in teachers depending on their conceptions of teaching. both videos were scripted according to the relevant literature (e.g. postareff & lindblom-ylänne, 2008) to describe a content-focused (cfts) and a learning-focused teaching situation (lfts). the first video presented cfts, in which a teacher is lecturing about devising interview questions. she is very focused on transmitting her topic and is not paying any attention to the audience. the students are sitting and looking bored. one is yawning, another is tapping her phone and some are conversing with each other. the trigger in the first video was the teacher totally ignoring the students and their off-task behaviour. the situation is reminiscent of a typical situation requiring classroom management (wolff et al., 2021) and thus noticing that students are not attending to the lecture. the second video presented an lfts, in which a teacher is lecturing about observation as a research method when a student interrupts her with a question concerning the ethics of observation. the teacher thinks a while and then prompts the students to have discussions in pairs for a few minutes. the trigger here was the teacher’s positive reaction to the student’s question, followed by engaging the students instead of directly answering the question by herself. thus, there is space and flexibility for changes in her teaching plan; the teacher sees students as active participants and relies on their ability to find the answer and process the knowledge themselves (postareff & linblom-ylänne, 2008). additionally, the teacher’s positive reaction to the interruption implies a good, safe atmosphere in the seminar room. the order of presentation of the videos was selected due to the assumed priming effect of the lfts video, meaning that after seeing the teacher’s behaviour of engaging the students, the participants would be more likely to notice the missing engagement in the cfts video. procedure the data collection procedure began with an orientation (see table 1). for each participant, the eye tracker was calibrated using 9-point calibration at the start of the data collection. to maximise calibration accuracy, participants were requested to take as comfortable a position as possible to prevent changes in position during the recording. participants sat approximately 60 cm from the screen on a manually adjustable chair. after calibration, they were given instructions regarding the session. the participants were told that they would be watching different lecturing situations and rating them from the viewpoint of teaching and learning. on the second watching of each video, they would be asked to think aloud about their interpretation of the situation. after the instructions, the participants watched a rehearsal video and practiced the think-aloud procedure. table 1. the study’s procedure orientation video viewing questionnaire cfts-video lfts-video murtonen et al 72 | f l r content-focused teaching situation learning-focused teaching situation calibration + instructions + rehearsal video first watch + rating + second watch and simultaneous think-aloud + rating and explanation first watch + rating + second watch and simultaneous think-aloud + rating and explanation background questions after the orientation, the actual data collection started. the participants watched both videos twice in the same order. after the first viewing, they rated the situation from the viewpoints of teaching and learning on a scale from 1 to 5 (1 = very poor, 2 = poor, 3 = moderate, 4 = good and 5 = very good). the second viewing took place immediately after the rating. they received the following prompt: “now, you are asked to watch the previous situation again and simultaneously think aloud about your interpretation of it. explain what is going on from the viewpoint of teaching and learning.” they also had a change to correct their rating. such an open approach was considered advantageous, since it is in no way preconditioned by the researchers and thus purely elicits the viewer’s perspective (kaiser et al., 2015). during the second viewing, the video’s sound was muted so that it would not interfere with the think-aloud process. if there were prolonged silences at the beginning of the viewing, the participants were prompted to verbalise what they were thinking about in the situation. they were allowed to continue their verbalisations even after the video vignette ended. since the think-aloud took place during the second watch, it can be considered retrospective; however, it was conducted without a gaze overlay of the first watch as a cue. the participants’ eye movements were recorded each time they viewed the videos. however, only the eye movements of the first viewing were used in the analyses, since these were considered to represent so-called “pure” viewing; as such, they are comparable to an authentic teaching situation as a one-time event without the possibility of reviewing the situation. after finishing the video viewing, the participants answered a short background questionnaire. due to calibration problems and common problems with eye-tracking data quality, such as data loss (holmqvist et al., 2011), the data of only 41 participants in the cfts video and 40 participants in the lfts video were available for the analyses out of the total of 49 participants. the percentage of gaze samples with at least one eye detected was 82.92 for the cfts video and 82.07 for the lfts video. analysis analysis of visual attention when watching teaching situations the participants’ viewings of the videos were analysed using tobii studio version 3.4.5. (tobii ab, danderyd, sweden). additionally, the numerical data from tobii studio were transferred to the ibm statistical package for the social sciences (spss), version 25 (ibm corp., armonk, ny, usa), which was used for further analyses. the videos were divided into areas of interest (aois), that is, the regions in the stimulus from which the authors were interested in gathering data (holmqvist et al., 2011). since both videos depicted the same scene, the same aois were used on the students, the teacher and the slides (see figure 2). murtonen et al 73 | f l r figure 2. the common aois used in both videos with an example scan path fixations and saccades as eye-tracking parameters are thought to reflect voluntary, overt visual attention (e.g. duchowski, 2007). the intake of visual information from the environment is assumed to happen largely during fixations (kok & jarodzka, 2016), which usually reflect the desire to focus attention on a certain object of interest. thus, fixations were considered useful for identifying where teachers focused their attention. the sum of fixation durations on each aoi was chosen to analyse the visual attention of the participants, i.e. for how long they had watched each aoi. as we were interested in the division of fixation time for each participant between different aois, the sum of fixation durations on each aoi was used to calculate the percentage share of fixation time for each aoi. thus, we would get a viewing profile of each participant (i.e. how much they would, in terms of percentage, fixate on the students, the teacher and the slides). we expected more fixations on the relevant regions to indicate deeper cognitive processing or the importance of a region (e.g. grub et al., 2020). the so-called white space, i.e. the visual attention on areas other than aois, was not considered when calculating the provision of fixations. this decision was based on descriptive statistics showing that the number of white space fixations was minimal. the eye-tracking data were normally distributed, thus enabling the use of independent samples t-tests. analysis of video interpretations the think-aloud protocols were transcribed verbatim and analysed qualitatively using nvivo 12 software (alfasoft ab, göteborg, sweden). the analyses were performed by the second and last authors, who are pedagogically qualified teachers and researchers in the field. theory-based content analysis was used to analyse the interpretations of the triggers. the structure of the coding scheme continuum was derived from the theory of teaching conceptions (e.g. kember & kwan, 2000; samuelowicz & bain, 2001). in the continuum from 1 to 5, 1 represented a strong knowledge-transmission conception and 5 represented a strong learningfacilitation conception of teaching. in the analysis of the lfts video, the scale was skewed towards the knowledge transmission end of the continuum, since critical or knowledge transmission reflecting comments on that video were scarce. the descriptions of each category were based on what emerged from the think murtonen et al 74 | f l r aloud protocols. both the think-aloud during the viewing and the summing up of the ratings after the video viewing were analysed when deciding to which category each participant’s answer belonged. multiple rounds of open coding were conducted to reach the current coding scheme (see table 2). table 2. coding scheme of the video interpretations: teachers’ reactions to the trigger event from the perspective of knowledge-transmission and learning-facilitation conceptions. cfts video (content-focused teaching) trigger: students not attending lfts video (learning-focused teaching) trigger: teacher engaging students category description description 1 = reflecting strong knowledgetransmission conception does not notice the trigger. praising the teaching or focusing on the presentation. interpretation of the trigger from the knowledge-transmission perspective. the teacher performs poorly, since the structure of the teaching suffers from a student’s interruption or the teacher should give a clear answer to the student’s question. 2 = reflecting knowledgetransmission conception notices the trigger but does not suggest that the teacher should react to it. if suggestions for improvement are given, they are related to the presentation (e.g. there should be more pictures in the slides). mere description of the situation without taking a positive stand on teacher’s learningfocused performance or neutral interpretation. 3 = reflecting characteristics of both conceptions notices the trigger and suggests that the teacher should do something (e.g. have a break or somehow get students’ attention), but no clear mentioning of supporting student learning. a superficially positive view of the situation/the teacher’s performance is good (no arguments given). no specific mention of the trigger. 4 = reflecting learning-facilitation conception notices the trigger. suggestions for improvement are related to facilitating students’ learning (e.g. engaging/motivating students, increasing interaction). the teacher’s performance is considered good because she reacts positively to the student’s question (the trigger); mentions facilitation of learning. 5 = reflecting strong learning-facilitation conception strongly notices the trigger. teaching is considered very poor since the students are not learning. suggestions for improvement are related to students’ learning (e.g. engaging students, fostering their own thinking). the interpretation is given clearly from the viewpoint of learning. praises teacher’ reaction to the trigger. mentions students’ knowledge building or the pedagogy behind not answering the student directly. viewpoint of deep learning (noticing that the teacher is changing their original plan to answer students’ needs and interests). in the following, citation examples are presented to illustrate the classes in the coding scheme. participants are referred to as p and the identification code, such as p1. the next example of an interpretation of the cfts video was classified in category 2, reflecting a knowledge-transmission conception, since the participant noticed the trigger of the students not focusing but did not suggest that the teacher should do anything about it. i see that the students are looking quite bored and concentrating on their own business. … i don’t know why. i don’t think it is the style of teaching, but maybe the topic. … i don’t see anything special to criticise about the teacher’s actions; this is very typical university teaching. (p50) murtonen et al 75 | f l r in another example of the cfts video, the participant interpreted the trigger from the viewpoint of learning. this interpretation reflected a strong learning-facilitation conception of teaching (category 5): from the viewpoint of teaching, it clearly seems that, in this situation, the conditions for learning new things are not very good. … only a few students are following the situations and the teacher’s teaching style seems to be one in which her message doesn’t reach the students very well. (p47) the next excerpt from the interpretation of the lfts video was classified as reflecting a knowledgetransmission conception (category 2). the participant just described the situation, but neither indicated whether the teacher performed well nor interpreted the trigger from the viewpoint of teaching and learning: in this scenario, the lecturer was still giving this lecture, but this time, she considered the students’ interest in the topic and made them discuss it as group work. (p8) a second example from the lfts video was classified as reflecting a strong learning-facilitation conception (category 5). in this excerpt, the teacher’s teaching actions, i.e. the trigger, were considered good, since she changed her original lecture plan according to what the students showed interest in. … the teacher was able to compromise her original lecturing plan, and when there was a question, instead of directly answering it, she made the students ponder it and this way they would have a more concrete learning experience. (p21) interrater reliability was calculated for the interpretations of the triggers for 25% of the data using cohen’s weighted kappa. substantial agreement was reached for both videos, indicating fair reliability (cfts video: 66.67%, weighted kappa = 0.739; lfts video: 58.33%, weighted kappa = 0.639). analysis of the connections between targeted concepts spearman’s correlations were utilised to examine the relations between targeted concepts. the rating scale of the cfts video was reversed so that it would be comparable to the ratings of the lfts video. a path analysis was conducted using the mplus software (version 8.4, muthen & muthen, 2019) to portray the possible causal linkages between the target variables to better understand the processes and mechanisms behind the phenomenon. path analysis was chosen because it allows for inferring and testing a sequence of causal links between variables of interest and examining the relationship between multiple predictor and criterion variables simultaneously (barbeau et al., 2019). missing values were handled by employing full information maximum likelihood (fiml) in the model estimations. fiml can handle missing data (mar) in an optimal way (muthén & muthén, 2017). results teachers’ visual attention in teaching situations first, the teachers’ visual attention on both videos was examined. on the cfts video, the pedagogically educated teachers watched statistically significantly more at the students (cohen’s d = .77) and almost statistically significantly less at the teacher (cohen’s d = .64) than their untrained colleagues (table 3). the effect sizes were moderate (cohen, 1988). on the lfts video, the differences pointed in the same direction as on the cfts video but were not statistically significant. there were no statistically significant differences between the teaching experience groups in either video (see table 4). murtonen et al 76 | f l r table 3. percentage share of fixation time for each aoi in teaching situations in the cfts and lfts videos between pedagogically untrained and trained university teachers pedagogical training t(37 or 38) p no (n = 21–22) m, sd yes (n = 17–19) m, sd cfts video (content-focused teaching situation) aoi teacher (%) 33.70, 10.45 26.07, 13.35 2.02 .05 aoi students (%) 42.35, 17.95 56.16, 17.82 -2.44 .02* aoi slides (%) 23.95, 14.57 17.77, 10.18 1.54 .13 lfts video (learning-focused teaching situation) aoi teacher (%) 49.79, 9.49 45.50, 12.27 1.23 .23 aoi students (%) 35.66, 11.94 42.65, 15.49 -1.59 .12 aoi slides (%) 14.55, 7.19 11.85, 7.44 1.14 .26 note. the number of teachers varied in the videos due to missing data (cfts video: 21 untrained and 19 trained teachers; lfts video: 22 untrained and 17 trained teachers). *p < .05 table 4. percentage share of fixation time for each aoi in teaching situations in the cfts and lfts videos between novice and more experienced university teachers teaching experience t (3638) p 0–2 years (n = 17–18) m, sd >2 years (n = 21–23) m, sd cfts video (content-focused teaching situation) aoi teacher (%) 30.74, 11.64 29.26, 13.23 .37 .72 aoi students (%) 49.10, 17.28 49.79, 20.83 -.11 .91 aoi slides (%) 20.16, 9.12 20.94, 15.37 -.20 .84 lfts video (learning-focused teaching situation) aoi teacher (%) 48.68, 10.29 47.17, 11.57 .43 .67 aoi students (%) 39.51, 12.58 38.35, 15.31 .26 .80 aoi slides (%) 11.81, 5.37 14.48, 8.64 -1.17 .25 note. the number of teachers varied in the videos due to missing data (cfts video: 17 novice and 23 more experienced teachers; lfts video: 18 novice and 21 more experienced teachers). teachers’ verbal interpretations of teaching situations and triggers overall, the participants’ interpretations of the lfts video included more notions about learning facilitation (m = 4.08, sd = 1.00) than their interpretations of the cfts video (m = 3.16, sd = 1.11) (figure 3). no significant differences between trained and untrained teachers (cfts video: t(46) =.09, p = .93; lfts video: t(46) = -.67, p = .50), nor in relation to teaching experience (cfts video: t(46) = 1.67, p = .87; lfts video: t(46) = -.70, p = .49), were found concerning the interpretations. we assume that the lfts video was easier for the participants to interpret, since there were more happenings on the video, such as the teacher and the students being actively engaged in collaborative learning processes. thus, the lfts video trigger was able to capture watchers’ attention (bottom-up) and no shifting of attention elsewhere was needed (theeuwes, 2000). the cfts video where the lecturer was unidirectionally lecturing and the trigger was that students were passive resulted in more variation in teachers’ interpretations. murtonen et al 77 | f l r figure 3. the division of classified video interpretations of cfts (content-focused teaching situation) and lfts (learning-focused teaching situation) videos teachers’ numerical ratings of the teaching situations teachers’ ratings of the success of the teaching situation were higher overall concerning the lfts video (m = 4.33, sd = .56) than the cfts video (m = 3.43, sd = .74), showing that the teachers considered the lfts video to illustrate a better situation in terms of teaching and learning. no significant differences in ratings were found between untrained and trained teachers (cfts video: t(46) = .29, p = .77; lfts video: t(46) = .35, p = .73) or novice and more experienced teachers (cfts video: t(46) = -.23, p = .82). a path model of university teachers’ visual attention and interpretations in teaching situations finally, path analysis was conducted to portray the causal linkages between the target constructs. the cfts video was selected for the path analysis because it resulted in more variation in participants’ eye movements as well as their interpretations and ratings; thus, its explanatory power was expected to be stronger. the aoi of students was used as the basis of the model, since noticing students’ passivity was central in the cfts video. the correlations among the studied variables are shown in table 5. table 5. means, standard deviations and correlations (rs) among study variables 1 2 3 4 5 1. visual attention to students 1 .43** .36* .36* .04 2. interpretation 1 .47** -.01 -.04 3. rating 1 .00 .03 4. pedagogical training (1 = no, 2 = yes) 1 .17 5. teaching experience 1 m 49.34 3.16 3.43 sd 18.95 1.11 0.75 *p < .05, **p < .01 murtonen et al 78 | f l r the path analysis is depicted in figure 4. the fit indices indicate that the model fits the data well: χ2(3) = 2.30, p = 0.512, cfi = 1.00, tli = 1.00, srmr = .04, rmsea = .00, 90% ci = [0.00, 0.223]. note. ***p < .001, **p < .01, *p < .05, ns = not significant figure 4. the final structural model with standardised path coefficients (n = 47) in this model, two paths were examined: 1) the effects of pedagogical training on the rating given to the teaching situation and 2) the effects of pedagogical training on video interpretation. the first path did not result in any significant indirect effects. however, in the second path, the indirect effect of visual attention on students as a mediator between pedagogical training and video interpretation was significant: β = 0.231 (p = .019; 95% ci = [.051, .463]). in addition, an almost significant direct negative effect of pedagogical training on video interpretation was found: β = -0.267 (p = .050; 95% ci: [-.505, -.016]), meaning that if the teacher was not able to visually focus on students, she would not produce an accurate verbal interpretation. the relationships proposed in the model explain 16.2% of the variance in visual attention on students, 28.9% of the variance in video interpretation and 24.7% of the rating of the teaching situation. thus, pedagogical training seems to affect teachers’ visual attention to students, which is further associated with video interpretations and video ratings. in other words, pedagogically trained teachers gaze more at the students, and gazing at them further engenders more learning-focused interpretations and aligned ratings of the teaching situations. in addition, there was a small negative direct effect of pedagogical training on verbal interpretation, indicating that pedagogically educated teachers used fewer learning-focused explanations of the situation if they did not pay visual attention to the students. thus, visual attention seems to be central to interpreting students’ learning situations. discussion focusing on students’ learning is a central element in high-quality university teaching (e.g. prosser & trigwell, 2014). previous studies have shown that teachers who express a learning-facilitation conception of teaching, i.e. who consider teaching as supporting students’ learning, more frequently report a learning-focused approach to teaching in practice, while those who consider teaching as transmitting knowledge tend to adopt a content-focused approach in their teaching practices (kember & kwan, 2000). pedagogical training, even a short one, has been shown to enhance teachers’ learning-facilitating conception (vilppu et al., 2019). murtonen et al 79 | f l r while prior studies on university teaching have mainly used self-report questionnaires and interviews, we broadened the scale to include eye-tracking methodology. we investigated whether teachers’ focus on students could be found at the intersection of their teaching conceptions and visual attention. based on previous studies, we hypothesised that pedagogically trained teachers would express more learning-focused views; thus, we expected a connection between previous training and focus on students, on both the visual and the interpretation levels. the analyses of verbal interpretations of the videos revealed both learning-facilitation and knowledge-transmission conceptions in teachers. when analysing only the verbal interpretations of the videos, we found no differences between the pedagogically trained and untrained or novice and more experienced teachers. our results concerning teachers’ visual attention to students showed that pedagogically trained teachers fixated more on the students than their untrained colleagues did. the difference was statistically significant in the content-focused teaching situation (cfts) video, where the students were passive and bored, but not in the learning-focused teaching situation (lfts) video, where the teacher actively engaged students in learning. pedagogically, the situation of the cfts video, in which students were passive, would need teacher attention and intervening action. only a few studies have addressed the recognition of possible situations that need teacher’s action (grub et al., 2020), of which our trigger event in cfts video is an example. we claim that the trained teachers, due to their more elaborated knowledge base (lachner et al., 2016), were more competent in noticing, i.e. paying visual attention to the problematic situation and interpreting it adequately. the cognitive theory of the top-down and bottom-up control of visual attention supports our finding: after the teacher’s teaching actions, i.e. lecturing on the cfts video, had captured watchers’ attention (bottom-up), the trained teachers used their attentive processes to shift their attention elsewhere (top-down) to focus on important things (e.g. theeuwes, 2000). in our case, the learning-facilitating conception helped pedagogically trained teachers shift their attention from the lecturing teacher to students when an action-needing event, i.e. boredom and passivity, occurred. analyses of teachers’ visual attention to the other elements of the lecture, the teachers and the slides, showed no statistically significant differences. however, teachers’ visual attention to the lecturing teacher when viewing the cfts video was statistically almost significant and in the direction of our hypothesis, showing that the untrained teachers paid more attention to the teacher than the pedagogically trained teachers did. previous studies have shown that experts tend to fixate more often and for a longer duration on relevant areas, whereas novices look more frequently at irrelevant areas (grub et al., 2020). similarly, our pedagogically trained teachers paid more visual attention to the students, noticed the trigger event and evaluated the situation in terms of learning-facilitation conception. the untrained teachers probably did not notice students’ boredom as a relevant phenomenon since they did not pay enough visual attention to the students. this is probably because, according to their knowledge-transmitting teaching conception, what the teacher does is most important. in our study, teaching experience measured in teaching years was not connected to visual attention and interpretations. this result contrasts with studies on the lower levels of education (e.g. stahnke & blömeke, 2021), in which experienced teachers differ from novices. there are many reasons for this contrasting result. at lower educational levels, teachers are usually pedagogically trained, unlike in universities, where university teachers may totally lack pedagogical education. thus, to compare the settings, we would need a study in which we have pedagogically trained novice and expert university teachers. our sample comprised mainly novice teachers, so more profound studies including teachers with extensive experience will be needed in the future. the university teaching environment also differs significantly from that of other educational levels; for example, a teacher may not always teach the same students and the place where teaching takes place may murtonen et al 80 | f l r always be different. thus, we argue that the results of other educational levels’ eye-tracking studies cannot be directly applied to higher education. on the other hand, our study is in line with expertise research results that found training to be more important than experience (e.g. ericsson, 2008). thus, it may be that in the university environment, having at least some pedagogical training is more important than having long teaching experience without pedagogical education. the other possible concerns of this study included the rather small sample size for quantitative modelling; this may affect the reliability of the statistical analyses, although it is comparable with other eye-tracking studies (see beach & mcconnel, 2019). on the other hand, in small samples, the effects are often undetected; this might indicate that we have discovered an interesting phenomenon that needs to be confirmed in future studies. the division of teachers into two groups with either less or more than two years of experience can be considered a problematic solution. however, the small sample size and the fact that most of the teachers were novices did not allow many other solutions. another concern was that the order in which the videos were watched was not randomised for the participants. however, we think that the order in which the videos were shown (first the content-focused scenario, then the more appropriate learning-focused scenario) was justified to tap into participants’ conceptions of teaching. showing the more favourable teaching video first could have affected their interpretations in the second video. the videos seemed to differ in their discriminatory power, which proved better for the cfts video. we assume this was because the trigger event was the passivity of the students, which required visual attention to the aoi of students. furthermore, it was subtler than the trigger event in the lfts video, requiring the participants to look at areas other than the most obvious, the teacher, who was talking all the time. in addition, in the eye-tracking data analyses, the visual attention on areas other than aois was not considered, since the number of the so-called white space fixations was minimal. however, since the aois were of different sizes, the participants might have looked at some of the aois accidentally more than others. in future studies, this should be considered in the analyses. in this study, the teachers watched teaching situations on a video, which is different from being in a real teaching situation and looking at their own students. university teachers’ gaze at their own teaching situations needs to be studied in the future, which raises its own methodological questions (cortina et al., 2015). however, using this simple eye-tracking design, we were able to conduct operationalisation and analysis of the data and obtain support for our hypotheses, which will lay the groundwork for more complex studies. university teachers are an interesting group to study, since many of them lack pedagogical training, as opposed to primary and secondary school teachers, who are usually pedagogically qualified. this novel study showed that pedagogical training is important for university teachers to develop their ability to notice important events in lecturing situations. the type of video viewing used in this study appeared to be a suitable instrument for measuring university teachers’ visual attention and related conceptions of teaching. we suggest that video interpretations combined with visual attention reflect teachers’ conceptions of teaching and offer new insights into the area of research, which has traditionally been studied almost entirely using self-reporting instruments (see also vilppu et al., 2019). our findings are very important, meaning that when a trained teacher notices on a visual level that the students need engaging, they may be able to engage them in active learning, which is considered central in high-quality teaching (cf. lonka & ketonen, 2012). in contrast, if a university teacher has no pedagogical training, they may not be competent in noticing situations where students need engaging. this finding proves that visual attention plays a central role in teachers’ ability to focus on students. murtonen et al 81 | f l r key points visual attention combined with verbal interpretations offers a frontline method for studying university teachers’ pedagogical expertise. pedagogically trained teachers paid more visual attention to the students in a situation where the students were bored and did not attend to the lecture. teachers who paid visual attention to important events during teaching were also able to formulate a more accurate verbal interpretation, reflecting a learning-facilitating conception of teaching. previous pedagogical training has explained differences in visual attention and verbal interpretations. teaching experience as measured by the number of years teaching was not connected to visual attention and verbal interpretations. acknowledgments we are thankful to the graduate students and researchers who participated in making the videos and to all the teachers who took part in the study. references barbeau, k., boileau, k., sarr, f., & smith, k. (2019). path analysis in mplus: a tutorial using conceptual model of psychological and behavioural antecedents of bulimic symptoms in young adults. the quantitative methods for psychology, 15(1), 38–53. https://doi.org/10.20982/tqmp.15.1.p038 beach, p. & mcconnel, j. (2019). eye tracking methodology for studying teacher learning: a review of the research. international journal of research & method in education, 42(5), 485–501. https://doi.org/10.1080/1743727x.2018.1496415 blömeke, s., gustafsson, j. e., & shavelson, r. (2015). beyond dichotomies: competence viewed as a continuum. zeitschrift für psychologie, 223, 3–13. https://doi.org/10.1027/2151-2604/a000194 cohen, j. (1988). statistical power analysis for the behavioral sciences (2nd ed.). lawrence erlbaum associates. https://doi.org/10.4324/9780203771587 cortina, k. s., miller, k. f., mckenzie, r., & epstein, a. (2015). where low and high inference data converge: validation of class assessment of mathematics instruction using mobile eye tracking with expert and novice teachers. international journal of science and mathematics education, 13, 389–403. https://doi.org/10.1007/s10763-014-9610-5 doyle, w. (2006). ecological approaches to classroom management. in c. evertson & c. weinstein (eds.), handbook of classroom management: research, practice and contemporary issues (pp. 97–125). lawrence erlbaum associates. https://doi.org/10.20982/tqmp.15.1.p038 https://doi.org/10.1080/1743727x.2018.1496415 https://doi.org/10.1027/2151-2604/a000194 https://doi.org/10.4324/9780203771587 https://doi.org/10.1007/s10763-014-9610-5 murtonen et al 82 | f l r duchowski, a. t. (2007). eye tracking methodology. theory and practice. (2nd ed.). springer. https://doi.org/10.1007/978-1-4471-3750-4 dunekacke, s., jenβen, l., & blömeke, s. (2015). effects of mathematics content knowledge on pre-school teachers’ performance: a video-based assessment of perception and planning abilities in informal learning situations. international journal of science and mathematics education, 13, 267–286. https://doi.org/10.1007/s10763-014-9596-z entwistle, n. (2005). learning outcomes and ways of thinking across contrasting disciplines and settings in higher education. the curriculum journal, 16(1), 67–82. https://doi.org/10.1080/0958517042000336818 ericsson, k. a. (2008). deliberate practice and acquisition of expert performance: a general overview. academic emergency medicine, 15, 988–994. https://doi.org/10.1111/j.1553-2712.2008.00227.x garner, j. k., & kaplan, a. (2019). a complex dynamic systems perspective on teacher learning and identity formation: an instrumental case. teachers and teaching: theory and practice, 25(1), 7–33. https://doi.org/10.1080/13540602.2018.1533811 gaudin, c., & chaliès, s. (2015). video viewing in teacher education and professional development: a literature review. educational research review, 16, 41–67. https://doi.org/10.1016/j.edurev.2015.06.001 gibbs, g., & coffey, m. (2004). the impact of training of university teachers on their teaching skills, their approach to teaching and the approach to learning of their students. active learning in higher education, 5, 87–100. https://doi.org/10.1177/1469787404040463 grub, a-s., biermann, a., & brünken, r. (2020). process-based measurement of professional vision of (prospective) teachers in the field of classroom management. a systematic review. journal for educational research online, 12, 75–102. https://doi.org/10.25656/01:21187 holmqvist, k., nyström, m., andersson, r., dewhurst, r., jarodzka, h., & van de weijer, j. (2011). eye tracking: a comprehensive guide to methods and measures. oxford university press. just, m. a. & carpenter, p. a. (1980). a theory of reading: from eye fixations to comprehension. psychological review, 87(4), 329–354. https://doi.org/10.1037/0033-295x.87.4.329 kaiser, g., busse, a., hoth, j., könig, j., & blömeke, s. (2015). about the complexities of video-based assessments: theoretical and methodological approaches to overcoming shortcomings of research on teachers’ competence. international journal of science and mathematics education, 13(2), 369– 387. https://doi.org/10.1007/s10763-015-9616-7 kember, d., & kwan, k. (2000). lecturer’s approaches to teaching and their relationship to conceptions of good teaching. instructional science, 28, 469–490. https://doi.org/10.1023/a:1026569608656 knight, p. (2002). being a teacher in higher education. society for research into higher education & open university press. https://doi.org/10.1007/978-1-4471-3750-4 https://doi.org/10.1007/s10763-014-9596-z https://doi.org/10.1080/0958517042000336818 https://doi.org/10.1111/j.1553-2712.2008.00227.x https://doi.org/10.1080/13540602.2018.1533811 https://doi.org/10.1016/j.edurev.2015.06.001 https://doi.org/10.1177/1469787404040463 https://doi.org/10.25656/01:21187 https://doi.org/10.1037/0033-295x.87.4.329 https://doi.org/10.1007/s10763-015-9616-7 https://doi.org/10.1023/a:1026569608656 murtonen et al 83 | f l r kok, e. m., & jarodzka, h. (2016). before your very eyes: the value and limitations of eye tracking in medical education. medical education, 51(1), 114–122. https://doi.org/10.1111/medu.13066 könig, j., blömeke, s., klein, p., suhl, u., busse, a., & kaiser, g. (2014). is teachers’ general pedagogical knowledge a premise for noticing and interpreting classroom situations? a video-based assessment approach. teaching and teacher education, 38, 76–88. https://doi.org/10.1016/j.tate.2013.11.004 lachner, a., jarodzka, h., & nückles, m. (2016). what makes an expert teacher? investigating teachers’ professional vision and discourse abilities. instructional science 44(3), 197–203. https://doi.org/10.1007/s11251-016-9376-y levin, d. m., hammer, d., & coffey, j. e. (2009). novice teachers’ attention to student thinking. journal of teacher education, 60(2), 142–154. https://doi.org/10.1177/0022487108330245 lonka, k., & ketonen, e. (2012). how to make a lecture course an engaging learning experience? studies for the learning society, 2-3, 63–74. https://doi.org/10.2478/v10240-012-0006-1 mcintyre, n. a., mainhard, m. t., & klassen, r. m. (2017). are you looking to teach? cultural, temporal and dynamic insights into expert teacher gaze. learning and instruction, 49, 41–53. https://doi.org/10.1016/j.learninstruc.2016.12.005 muthén, l. k., & muthén, b. o. (2017). mplus user’s guide (8th ed.). muthén & muthén. postareff, l., & lindblom-ylänne, s. (2008). variation in teachers’ descriptions of teaching: broadening the understanding of teaching in higher education. learning and instruction, 18, 109–120. https://doi.org/10.1016/j.learninstruc.2007.01.008 postareff, l., & nevgi, a. (2015). development paths of university teachers during a pedagogical development course. educar, 51(1), 37–52. https://doi.org/10.5565/rev/educar.647 postareff, l., nevgi, a., & lindblom-ylänne, s. (2007). the effect of pedagogical training on teaching in higher education. teaching and teacher education, 23, 557–571. https://doi.org/10.1016/j.tate.2006.11.013 pouta, m., lehtinen, e., & palonen, t. (2020). student teacher’ and experienced teachers’ professional vision of students’ understanding of the rational number concept. educational psychology review 33, 109–128. https://doi.org/10.1007/s10648-020-09536-y prosser, m., & trigwell, k. (2014). qualitative variation in approaches to university teaching and learning in large first-year classes. higher education, 67, 783–795. https://doi.org/10.1007/s10734-013-9690-0 samuelowicz, k., & bain, j. d. (2001). revisiting academics’ beliefs about teaching and learning. higher education, 41, 299–325. https://doi.org/10.1023/a:1004130031247 seidel, t., stürmer, k., blomberg, g., kobarg, m., & schwindt, k. (2011). teacher learning from analysis of videotaped classroom situations: does it make a difference whether teachers observe their own teaching or that of others? teaching and teacher education, 27, 259–267. https://doi.org/10.1016/j.tate.2010.08.009 https://doi.org/10.1111/medu.13066 https://doi.org/10.1016/j.tate.2013.11.004 https://doi.org/10.1007/s11251-016-9376-y https://doi.org/10.1177/0022487108330245 https://doi.org/10.2478/v10240-012-0006-1 https://doi.org/10.1016/j.learninstruc.2016.12.005 https://doi.org/10.1016/j.learninstruc.2007.01.008 https://doi.org/10.5565/rev/educar.647 https://doi.org/10.1016/j.tate.2006.11.013 https://doi.org/10.1007/s10648-020-09536-y https://doi.org/10.1007/s10734-013-9690-0 https://doi.org/10.1023/a:1004130031247 https://doi.org/10.1016/j.tate.2010.08.009 murtonen et al 84 | f l r shulman, l. s. (1987). knowledge and teaching: foundations of the new reform. harvard educational review, 57, 1–22. https://doi.org/10.17763/haer.57.1.j463w79r56455411 stahnke, r., & blömeke, s. (2021). novice and expert teachers’ noticing of classroom management in whole-group and partner work activities: evidence from teachers' gaze and identification of events. learning and instruction, 74, 101464. https://doi.org/10.1016/j.learninstruc.2021.101464 stes, a., & van petegem, p. (2011). instructional development for early career academics: an overview of impact. educational research, 53, 459–474. https://doi.org/10.1080/00131881.2011.625156 theeuwes, j., atchley, p., & kramer, a. f. (2000). on the time course of top-down and bottom-up control of visual attention. in s. monsell & j. driver (eds.), control of cognitive processes: attention and performance xviii (pp. 105–124). mit press. trigwell, k., & prosser, m. (2004). development and use of the approaches to teaching inventory. educational psychology review, 16, 409–424. https://doi.org/10.1007/s10648-004-0007-9 uiboleht, k., karm, m., & postareff, l. (2018). the interplay between teachers’ approaches to teaching, students’ approaches to learning and learning outcomes: a qualitative multi-case study. learning environments research, 21, 321–347. https://doi.org/10.1007/s10984-018-9257-1 van den bogert, n., van bruggen, j., kostons, d., & jochems, w. (2014). first steps into understanding teachers’ visual perception of classroom events. teaching and teacher education, 37, 208–216. https://doi.org/10.1016/j.tate.2013.09.001 van gog, t., kester, l., nievelstein, f., giesbers, b., & paas, f. (2009). uncovering cognitive processes: different techniques that can contribute to cognitive load research and instruction. computers in human behavior, 25, 325–331. https://doi.org/10.1016/j.chb.2008.12.021 vermunt, j. d., vrikki, m., warwick, p., & mercer, n. (2017). connecting teacher identity formation to patterns in teacher learning. in d. j. clandinin & j. husu (eds.), the sage handbook of research on teacher education (pp. 143–159). sage publications ltd. vilppu, h., södervik, i., postareff, l., & murtonen, m. (2019). the effect of short online pedagogical training on university teachers’ interpretations of teaching–learning situations. instructional science, 47(6), 679–709. https://doi.org/10.1007/s11251-019-09496-z wolff, c. e., jarodzka, h., & boshuizen, h. p. a. (2017). see and tell: differences between expert and novice teachers’ interpretations of problematic classroom management events. teaching and teacher education, 66, 295–308. https://doi.org/10.1016/j.tate.2017.04.015 wolff, c. e., jarodzka, h., & boshuizen, h. p. a. (2021). classroom management scripts: a theoretical model contrasting expert and novice teachers’ knowledge and awareness of classroom events. educational psychology review, 33, 131–148. https://doi.org/10.1007/s10648-020-09542-0 https://doi.org/10.17763/haer.57.1.j463w79r56455411 https://doi.org/10.1016/j.learninstruc.2021.101464 https://doi.org/10.1080/00131881.2011.625156 https://doi.org/10.1007/s10648-004-0007-9 https://doi.org/10.1007/s10984-018-9257-1 https://doi.org/10.1016/j.tate.2013.09.001 https://doi.org/10.1016/j.chb.2008.12.021 https://doi.org/10.1007/s11251-019-09496-z https://doi.org/10.1016/j.tate.2017.04.015 https://doi.org/10.1007/s10648-020-09542-0 murtonen et al 85 | f l r wolff, c. e., jarodzka, h., van den bogert, n., & boshuizen, h. p. a. (2016). teacher vision: expert and novice teachers’ perception of problematic classroom management scenes. instructional science, 44(3), 243–265. https://doi.org/10.1007/s11251-016-9367-z wyss, c., rosenberger, k. & bührer, w. (2021). student teachers’ and teacher educators’ professional vision: findings from an eye tracking study. educational psychology review, 33, 91–107. https://doi.org/10.1007/s10648-020-09535-z https://doi.org/10.1007/s11251-016-9367-z https://doi.org/10.1007/s10648-020-09535-z article received 21 january 2022 / revised 9 december 2022/ accepted 9 december 2022 / available online 11 january 2023 abstract introduction teachers’ pedagogical expertise in the university context focusing on students at the level of conceptions and approaches focusing on students at the visual level and noticing important events utilising videos in studying university teachers’ focus on students present study methods participants apparatus and materials procedure analysis analysis of visual attention when watching teaching situations analysis of video interpretations analysis of the connections between targeted concepts results teachers’ visual attention in teaching situations teachers’ verbal interpretations of teaching situations and triggers teachers’ numerical ratings of the teaching situations a path model of university teachers’ visual attention and interpretations in teaching situations discussion references codepen 6. balloo frontline learning research frontline learning research special issue vol.9 no.2 (2021) 121 144 issn 2295-3159 a primer on gathering and analysing multi-level quantitative evidence for differential student outcomes in higher education kieran balloo1, naomi e. winstone1 1surrey institute of education, university of surrey, uk article received 18 may 2020/ revised 22 september / accepted 21 october/ available online 12 march abstract a significant challenge currently facing the higher education sector is how to address differential student outcomes in terms of attainment and continuation gaps at various stages of students’ transitions. worryingly, there appears to be a ‘deficit’ discourse among some university staff in which differential outcomes are perceived to be due to student deficiencies. this may be exacerbated by institutional analyses placing an over-emphasis on the presence of the gaps rather than the causes. the purpose of this primer is to provide advice about how institutions can carry out far more nuanced analyses of their institutional data without requiring specialist software or expertise. drawing on a multi-level framework for explaining differential outcomes, we begin with guidance for gathering quantitative data on explanatory factors for attainment and continuation gaps, largely by linking sources of internal data that have not previously been connected. using illustrative examples, we then provide tutorials for how to model explanatory factors employing ibm spss statistics (ibm corp., armonk, ny, usa) to perform and interpret regression and meta-regression analyses of individualand group-level (aggregated) student data, combined with data on microand meso-level factors. we propose that university staff with strategic responsibilities could use these approaches with their institutional data, and the findings could then inform the design of context-specific interventions that focus on changing practices associated with gaps. in doing so, institutions could enhance the evidence-base, raise awareness, and further ‘embed the agenda’ when it comes to understanding potential reasons for differential student outcomes during educational transitions. keywords: diversity; transitions; attainment gaps; continuation gaps; micro, meso and macro level info corresponding author email: k.balloo@surrey.ac.uk doi: https://doi.org/10.14786/flr.v9i2.675 1. introduction recent transitions research has signalled a move away from seeing students as a homogeneous group undergoing a ‘transition process’ to acknowledging the importance of their diversity as they progress through university (gravett, 2019). one aspect of diversity that presents a significant challenge for the higher education sector relates to the presence of differential student outcomes. whilst at university, students from underrepresented backgrounds generally have lower levels of achievement (known as attainment gaps), and are less likely to progress from one year to the next (known as continuation gaps), than their traditional counterparts (ofs, n.d.). differential outcomes have been found for students from a range of backgrounds, with poorer attainment and continuation outcomes for those who are mature, studying part-time, from a lower socioeconomic group, and from a black, asian, and minority ethnic background (woodfield, 2014). these gaps have been identified across the sector internationally, including in the uk, usa, australia, belgium, germany, and denmark (lens & levrau, 2020; mountford-zimdars et al., 2015; singh, 2011; tieben, 2020). despite prior attainment and type of entry qualification only partially accounting for the presence of gaps (broecke & nicholls, 2007), there appears to be a ‘deficit’ discourse among some university staff in which differential outcomes are perceived to be due to student deficiencies (miller, 2016; singh, 2011; stevenson, 2012). this may be exacerbated by institutional analyses placing an over-emphasis on the presence of the gaps themselves. universities routinely collect a range of quantitative data at the point of student enrolment, some of which is then utilised to assess whether their widening participation targets are being met. these data may include demographic details such as gender, age, ethnicity, disabilities, and socioeconomic group, and they are further used to model whether there are differential outcomes in terms of attainment or continuation gaps at various stages of students’ transitions (e.g. end of first year grades, when progressing from the first year to the second year, etc.). analysing these data is necessary for regulatory reporting purposes, but these data alone do not aid teaching staff in their design of interventions for tackling the causes of these gaps. a recent report also highlighted this over-reliance on collecting quantitative data at the expense of engaging with students about their actual experiences (uuk & nus, 2019). since differential student outcomes are quantitatively assessed (jones, 2018), qualitative evidence alone cannot be used to ascertain the factors that explain why there are gaps. institutions have a duty to both monitor and attempt to reduce gaps, so this is likely to be an important target for senior management. however, we argue that any large-scale quantitative analyses of attainment and continuation gaps need to move beyond solely focusing on the presence of the gaps, and characteristics of the students, to take into account the explanatory factors for these differential outcomes. in order for teaching staff to improve the focus and design of interventions addressing gaps, they need to know the precise factors that may impact on these outcomes for their own students, drawing on evidence from their own institutional and disciplinary contexts (mountford-zimdars et al., 2017). thus, quantitative analyses may offer a way forward, but more nuance is likely to be needed when determining which data to analyse and how: universities need to take a more scientific approach to tackling the attainment gap, by gathering and scrutinising data in a far more comprehensive way than they may currently be doing, in order to inform discussions between university leaders, academics, practitioners and students. (uuk & nus, 2019, p. 2) therefore, the purpose of the current primer is to provide accessible guidance on the decisions that institutions should consider making when attempting to draw on evidence for the factors that may explain differential student outcomes. this primer provides suggestions for how explanatory factors can be quantitatively collected, and most importantly, how these data can be analysed using relatively straightforward approaches that do not require access to specialist software or expertise. it provides tools to model potential factors explaining differential outcomes within specific contexts throughout students’ transitions. such models could then be used to stimulate conversations between teaching staff and students around issues that have been identified, and enable the design of context-specific interventions that focus on changing practices that have been confirmed to be associated with gaps. this could also extend the literature on differential outcomes. 2. gathering multi-level quantitative evidence for differential student outcomes in his theory of student departure, tinto (1993) asserts that student persistence is based on how well they integrate into the social and academic systems of their university. this ability to integrate is thought to be influenced by students’ background characteristics, the institutional environment, and their experiences and interactions whilst they are at university. building on the findings of cousin and cuerton’s (2012) attainment gap report, mountford-zimdars et al. (2015) conducted a critical review of the potential causes of differential student outcomes to identify explanatory factors for outcome gaps, which they categorised into: students’ experiences of curriculum practices; relationships between staff and students; social, cultural and economic capital; and psychosocial and identity factors. mountford-zimdars et al. also proposed the use of a multi-level framework for understanding these explanatory factors. at a micro-level, explanatory factors may be related to the individual student and their one-to-one interactions with university staff. at a meso-level, explanatory factors may involve institutional and learning environment factors. at a macro-level, explanatory factors may be due to the wider higher education context and structure. these explanatory factors intersect with the multi-level framework, so they may occur at multiple levels and at various stages throughout students’ transitions. figure 1 shows how these explanatory factors might manifest at the multiple levels of mountford-zimdars et al.’s framework. figure 1. multi-level framework of explanatory factors for differential student outcomes based on the model from mountford-zimdars et al. (2015) although institutions have attempted to tackle differential student outcomes, interventions often simultaneously target both microand meso-level components, so it is difficult for teaching staff to disentangle which actions are reducing gaps and which are unsuccessful (miller, 2016). as a result, research-informed changes to practice may currently be limited. incorporating the model in figure 1 into quantitative analyses may enable a greater understanding of the specific aspects of students’ backgrounds and the university environment that are associated with differential outcomes within specific contexts. mountford-zimdars et al. (2015) made several recommendations for reducing differential student outcomes in their report, and we propose that the current primer might support institutions in attending to some of those suggestions. firstly, this primer provides guidance about how to enhance the evidence-base by drawing on module-level data 1 , linking sources of internal data that have not previously been connected to increase understanding about the institution’s context within the larger national and international picture. the methods should raise awareness among staff with strategic responsibilities, enabling interventions to be informed by evidence that is relevant to the institutional context. finally, this primer supports institutions in ‘embedding the agenda’ through aiding the design of both targeted interventions (i.e. interventions that are directed at specific student groups) and universal interventions (i.e. interventions for all students that may be particularly beneficial for specific student groups). since explanatory factors that occur at the microand meso-levels are perceived to be the institution’s responsibilities (mountford-zimdars et al., 2015), these factors need to be operationalised if they are to be included in quantitative models. micro-level factors occur at an individual level, so it is possible to quantify some of these as variables using validated self-report questionnaires. for example, the university attachment scale (france et al., 2010) includes a subscale for measuring students’ sense of belonging to the university. meso-level variables mostly involve the group-level features of modules and programmes. whilst the potential difficulty of gathering quantitative data of curriculum content has previously been noted (miller, 2016), we suggest that these data could be obtained from a document analysis or audit of module descriptors, module booklets, examiner reports, prospectuses, the vle structure and use, etc. there is likely to be a wealth of data available that could be useful, covering support available to students, types of assessment, extra-curricular activities, departmental policies, and anything else that may be relevant to the disciplinary and institutional context. the use of learning analytics data might also prove useful at both microand meso-levels. since macro-level variables occur at the sectoral level, it may be difficult to model variation in these variables unless data are being shared across institutions, but there is scope for gathering these data. ultimately, each institution needs to identify their own priorities whilst ensuring they align with the discussed theory. they must base their priorities on the relevance of the variables to their particular student groups, the data they have available, and whether their documentation is an accurate depiction of actual practices and how they are received by students. 3. analysing multi-level quantitative evidence for differential student outcomes in this section we provide tutorials for how to perform and interpret analyses of institutional data that do not require any specialist knowledge beyond a basic understanding of statistics. all tutorials utilise ibm spss statistics (ibm corp., armonk, ny, usa) and all spss data sets, syntax files and macros to run and adapt analyses for readers’ own purposes are freely available at https://doi.org/10.5281/zenodo.4115264, so readers only need access to spss software to perform the same analyses or adaptations of these2. because spss syntax has been provided with code for all of the procedural steps of running the necessary statistical tests, we only focus on how to interpret the results of such analyses. the data for all examples in this primer are fictional, and have only been designed to simulate the possible behaviour of institutional data for the purposes of demonstrating how the analytical approaches can be used. no inferences or conclusions should be drawn from the findings of these examples, because the results are not real. we anticipate that readers can use the example data sets as templates and substitute in their own data. 3.1 individual-level data (tutorials 1 and 2) if data for individual students (e.g. their demographic details and possibly questionnaire responses) are accessible for the purposes described in this primer, it is possible to analyse these data using two common regression approaches based on the general linear model (glm). therefore, in tutorials 1 and 2 we provide guidance about how to perform and interpret analyses of example individual-level data in order to model potential explanatory factors as covariates 3 of attainment and continuation gaps. these examples focus on micro-level explanatory factors, because these occur at the individual level. the questions we aim to answer in these examples are: are there attainment/continuation gaps based on the students’ background groups, and can a set of micro-level factors predict attainment/continuation over and above these groups? 3.2 group-level (aggregated) data (tutorials 3 and 4) we anticipate that there may only be access to aggregated data for analysis due to individual-level data being restricted to certain staff or because consent has not been obtained to connect sensitive student data to other variables, so in tutorials 3 and 4 we cover how meta-regression can be used to analyse such data. meta-analysis was designed to synthesise an average effect size and how much it varies across multiple studies (pigott & polanin, 2020). the glm can be extended in meta-analysis to also model the role of study-level covariates in explaining effect sizes (field & gillett, 2010; hedges et al., 2009). this is known as meta-regression. therefore, we propose that differences in data averaged for particular student groups at a module-level can be treated as an effect size, with the overall programmes functioning in the same way that separate studies do in meta-analysis studies. further, mesoand macro-level factors can be treated as study-level covariates because these occur at a module, programme, institution and sector level. by only drawing on average effect sizes, this approach reduces issues around data privacy, because it only requires data aggregated for a whole group rather than for individual students (pigott & polanin, 2020). moreover, beyond these privacy issues, one of the main advantages of using meta-regression in this context is that the attainment and continuation gaps become the direct focus of analyses, so we believe this is a novel approach to analysing these data. thus, in tutorials 3 and 4, we provide guidance about how to perform and interpret meta-regression analyses of example student data aggregated to a group level. as a result, these examples focus on meso-level explanatory factors, although they could also theoretically include macro-level factors if cross-institution data are available. since we are focusing on explaining the gaps themselves, the questions we aim to answer in these examples are: are there attainment/continuation gaps based on the students’ background groups, and can a set of meso-level factors predict these gaps? 3.3 tutorial 1: illustrative example of analysing attainment gap data measured at an individual level tutorial 1 presents an analysis of the individual_level_example.sav spss data set, which was designed to closely approximate what real student data might look like if recorded at an individual level. this hypothetical data set simulates a range of variables for a sample of students (n = 319) at the end of their first year of university. the variables are detailed in table 1 and include: students’ average grades (the outcome variable); a student background group variable; and a set of hypothetical micro-level variables that may account for differential student outcomes (as influenced by theory) to be used as covariates.  table 1 variables included in individual_level_example.sav spss data set the question we aim to answer in this tutorial is: is there an attainment gap based on the students’ background groups, and can a set of micro-level factors predict attainment over and above these groups? for tutorial 1, we performed a linear regression analysis, based on the glm, to determine the relative contribution that each variable might make to explaining variance in students’ grades. linear regression analysis functions by modelling the relationships between the variance of an outcome variable (in this case, average grades) and one or more covariates. multiple models can be compared in terms of their ability to predict variance in the outcome variable. for our first model, we entered the student background group variable on its own to initially determine whether there were group differences in grades (model 1). we then entered prior attainment, entry qualification, and the measure of belonging into the next model (model 2). covariates could also have been entered into the analysis at separate steps, so similar variables could have been grouped together; this is a useful way to compare models. the individual_level_attainment.sps spss syntax file includes the code for all stages of this analysis. in order to assess the model fit, we need to check for any sources of bias in the data (i.e. outliers and influential cases). if we want to generalise the model beyond the sample, we also need to check whether certain assumptions have been met. diagnostic tests for checking sources of bias and whether assumptions have been met have been included in the spss syntax, but they are not discussed further here. figure 2 displays the model summary table from the spss output. figure 2. spss output for the model summary table from linear regression analysis of individual_level_example.sav data set using individual_level_attainment.sps syntax in the first row of the table in figure 2, we can see the proportion of variance in grades (the outcome) that is predicted by model 1, which only includes the student background group variable. if we multiply the adjusted r square value (the proportion of variance explained in the outcome by the model, adjusted to take into account the number of covariates) by 100, this indicates that the group variable alone explains 9.8% of the observed variation in grades. the change statistics section of the table shows that this model is statistically significant (p < .001). the second row shows the model with entry qualification, prior attainment, and belonging covariates added (model 2). model 2 explains 42.5% of the variation in grades and the change in this model from model 1 is significant (f change = 59.14, p < .001). the anova table from the spss output also explains whether each model is a significant fit. we now look at the contribution each individual covariate made to each model by examining the coefficients table from the spss output (figure 3). figure 3. spss output for the coefficients table from linear regression analysis of individual_level_example.sav data set using individual_level_attainment.sps syntax the table in figure 3 shows the individual contribution of each covariate to each model in the analysis. all b-values (labelled unstandardized coefficients b in the spss output) are positive and significant (p < .001), meaning that higher scores on all variables predict higher grades. the b-values can also be used to understand how much change we would expect in grades from a one unit change in these variables based on whatever units the original variables were measured in, whilst holding the effects of other variables constant (when using regression, all coefficients are conditional on the effects of other variables in the respective model). in model 1, the b-value for the constant is 56.643, which represents the average end of first year grades (i.e. 56.64%) for the baseline group (group 0). the b-value for student background group is 5.462, which indicates that students coded as group 1 scored on average 5.46% higher in their end of first year grades than students coded as group 0, demonstrating that an attainment gap is present. in model 2, the b-value of .251 for prior attainment, indicates that a 1% increase in average prior grades predicts a 0.25% increase in average end of first year grades, whilst holding the effects of all other variables constant. it is important to note that the use of beta coefficients (labelled standardized coefficients beta in the spss output) for examining relative importance of covariates is problematic in educational research, because there is likely to be high dependence between covariates (karpen, 2017). therefore, courville and thompson (2001) advocate interpreting beta coefficients alongside zero-order correlations between the covariate and outcome in order to avoid misinterpretations. when covariates are highly linearly related, this is known as multicollinearity, and it could indicate there is evidence of bias, such that the b-values for each covariate may not be valid or accurate with respect to the other covariates. since multicollinearity can still be present even when there are relatively small correlations between covariates (alin, 2010), tolerance and vif values (as found in the collinearity statistics section of the coefficients table in figure 3) should also be checked. although there have been various debates about whether to use strict cut-offs when determining what constitutes a problematic level of collinearity in the data (thompson et al., 2017), using a general rule of thumb, we would not anticipate there being a particular issue with overlap in the variables in this analysis, since tolerance values are close to one and vif values are not much greater than one. this means that each variable appears to be making a unique contribution to the model. in this example, we aimed to determine whether there were attainment gaps based on students’ background groups, and whether a set of micro-level factors could predict attainment over and above these groups. the analysis of this hypothetical data set indicates that students coded as group 1 performed better than students coded as group 0, but variation in attainment can also be explained by students’ entry qualification, prior attainment and their sense of belonging to the university. 3.4 tutorial 2: illustrative example of analysing continuation gap data measured at an individual level tutorial 2 also presents an analysis of the individual_level_example.sav spss data set. however, the continuation variable in the data set was used as the outcome instead of grades (see table 1). we are no longer modelling a linear relationship between covariates and the outcome, because we now have a binary outcome (drop out or progress/complete), so the glm requires a logarithmic transformation to express the non-linear relationship in linear terms. binary logistic regression enables us to predict the odds of students dropping out of university based on a set of categorical and continuous covariates by computing an odds ratio (or) for each covariate. the or is the ratio between the probability of an event occurring and the probability of this event not occurring, and it can be used to determine the proportionate change in odds for the outcome variable based on a change in the covariate. to demonstrate this, table 2 displays the association between student background group and continuation, with a, b, c and d representing the numbers of students in each category. table 2 crosstabulation of student background group by continuation following table 2, an or can be computed to determine the odds that students in one group are more likely to drop out than students in the other group. if we divide the odds of dropout occurring for group 0 (a/b) by the odds of dropout occurring for group 1 (c/d), an or value of less than one means that c/d is larger than a/b, so this would indicate a negative relationship (i.e. a score of 0 on one variable and 1 on the other) between the student background group and continuation, meaning that students coded as group 1 are more likely to drop out (coded as 0). an or value greater than one means that a/b is larger than c/d, so this would indicate a positive relationship, meaning that students coded as group 0 are more likely to drop out. to aid with understanding, we have plotted a hypothetical positive relationship in figure 4. figure 4. visual depiction of a positive relationship between student background group and continuation the question we aim to answer in this tutorial is: is there a continuation gap based on the students’ background groups, and can a set of micro-level factors predict continuation over and above these groups? for tutorial 2, we performed a binary logistic regression analysis to determine the relative contribution that each variable might make to explaining the odds of students dropping out of university or progressing/completing. we began by specifying the models, so in the first model we entered the student background group variable on its own (model 1). we then entered prior attainment, entry qualification, and the measure of belonging into the next model (model 2). the individual_level_continuation.sps spss syntax file includes the code for all stages of this analysis. findings from model specifications are found in figure 5. figure 5. spss output of model summary tables from models 1 (table above) and 2 (table below) in logistic regression analysis of individual_level_example.sav data set using individual_level_continuation.sps syntax figure 5 demonstrates whether each model is a good fit to the data. the model row of the first table indicates that model 1 (which includes the student background group variable only) is a significant fit to the data (p < .001). the block row in the second table indicates that the change from model 1 to model 2 (which adds all covariates) is significant (p < .001). as with tutorial 1, we need to check for any sources of bias in the data (i.e. outliers and influential cases), and whether certain assumptions have been met if we want to generalise the model. diagnostic tests for checking sources of bias and whether assumptions have been met have been included in the spss syntax. we now look at the contribution each individual covariate made to each model by examining the variables in the equation tables from the spss output (figure 6). figure 6. spss output of variables in the equation tables from models 1 (table above) and 2 (table below) in logistic regression analysis of individual_level_example.sav data set using individual_level_continuation.sps syntax the tables in figure 6 show the individual contribution of each covariate to each model in the analysis. all b-values indicate positive relationships and are significant, but the or values (labelled as exp(b) in the spss output) are easier to interpret. in model 1, the or for student background group is 6.449. this value is greater than one, which indicates a positive relationship between student background group and continuation, and this is significant (p < .001). therefore, students coded as group 0 were more likely to drop out (also coded as 0) than students coded as group 1 by a factor of 6.449. in other words, students coded as group 0 were over six times as likely to drop out than students coded as group 1, demonstrating that a continuation gap is present. in model 2, the effect of background group drops to non-significance (p = .123, and the confidence interval, labelled as 95% c.i. for exp(b) in the spss output, also crosses one) and prior attainment is also non-significant (p = .072). entry qualification and sense of belonging are significant (or = 7.785, p < .001 and or = 2.208, p = .018, respectively). this indicates that students with lower scores on the measure of belonging and qualification type a (coded as 0) were more likely to drop out (also coded as 0), assuming the effects of all other variables are held constant. finally, we can check for interdependence between covariates by checking whether covariates are correlated, and whether there is evidence of multicollinearity by running a linear regression analysis (included in the spss syntax) to check tolerance and vif values. in this example, we aimed to determine whether there were continuation gaps based on students’ background groups, and whether a set of micro-level factors could predict continuation over and above these groups. the analysis of this hypothetical data set indicates that students coded as group 0 were more likely to drop out than students coded as group 1. however, this association was not present once taking into account students’ entry qualification and their sense of belonging to the university, which better predicted the odds of students dropping out or progressing/completing. 3.5 tutorial 3: illustrative example of analysing attainment gap data measured at a group level tutorial 3 presents an analysis of the group_level_attainment_example.sav spss data set, which was designed to closely approximate what real aggregated student data might look like when modelling attainment gaps. this hypothetical data set simulates a range of variables for a sample of 1056 modules within 40 programmes. the variables are detailed in table 3 and include: average grades at a module-level for two different groups of students (student background group), the year of study of the module in the programme, and a set of meso-level variables that may account for differential student outcomes (as influenced by theory) to be used as covariates. when modelling attainment gaps at an individual level (tutorial 1), the outcome was students’ grades and the initial covariate entered in the analysis was the student background group variable. extending this to a group level, the difference in average grades between the two groups for a particular module represents an effect size 5 , so this is our outcome variable when using meta-regression. standard meta-analytic techniques are predicated on the assumption that effect sizes are independent (cheung, 2019). however, effect sizes for each module are likely to be dependent because grades for each module will include the same students across a programme, known as correlated effects (hedges et al., 2010; tipton & pustejovsky, 2015). robust variance estimation (rve) can be used in meta-regression analyses to correct for correlated estimates (hedges et al., 2010)6 . meta-analyses using rve with fewer than 40 studies may produce overly narrow confidence intervals (tanner-smith & tipton, 2014), so we recommend that analyses include at least 40 programmes. 7 consideration should also be given to the inclusion of modules with very few students from a particular background group, because effect sizes based on small sample sizes could artificially inflate the size of gaps. one way to check for this is to perform sensitivity analyses, which involves running analyses with and without the data from particular modules to see whether their inclusion appears to bias the results. since we are focusing on explaining the attainment gaps themselves, the question we aim to answer in this tutorial is: are there attainment gaps based on the students’ background groups, and can a set of meso-level factors predict these gaps? for tutorial 3, we performed a meta-regression analysis using rve with correlated effects weights. the background group variable does not need to be entered into the analysis as a separate variable, so no covariates were included in the first model. this intercept-only model therefore tested whether there was a significant attainment gap between the two student groups (model 1), equivalent to the first model in tutorial 1. before entering any of the covariates 8 into the analysis we needed to code whether values of covariates can differ within a programme (i.e. values differ between modules within a programme) or only between programmes (i.e. values are the same for all modules within a programme). in the current example, the attendance variable could only differ between programmes (i.e. all modules within a programme took on the same value), so this variable could be included in the analysis as it is. however, the following covariates had the potential to take on different values across modules within a single programme: year of study of the module in the programme (because each programme included modules across years 1, 2 or 3), number of assessment components, and percentage of independent study time expected of students. therefore, new variables needed to be created in spss to average the scores for these covariates across all modules within a programme (programme-mean values for each variable to model between-programme effects) and to center these covariate scores around the programme-mean (to model within-programme/between-module effects). this is an important step, because it separates out betweenand within-effects to aid the interpretation of covariate effects (tanner-smith & tipton, 2014). we then entered all covariates into model 29. these included: year of study of the module, number of assessment components, the percentage of independent study time, and the attendance variable. we also needed to specify a rho value to represent the estimated size of intercorrelations expected between effect sizes within a programme. rho values range between 0 and 1, with higher values indicating greater dependence. as this value is only an estimate of dependency, we used a conservative rho value of .80. table 3 variables included in group_level_attainment_example.sav spss data set the group_level_attainment.sps spss syntax file includes the code for all stages of this analysis, including the creation of new covariate variables with centered and mean scores. the freely available rve meta-regression spss macro file, robustmeta.sps (tanner-smith & tipton, 2014)10 , should be saved on the computer that will be used for performing analyses prior to running any spss syntax files, and the file path for where this macro is saved needs to be added to the syntax, as noted in the syntax instructions. figure 7 displays the spss output for the model that did not include any covariates. figure 7. spss output of the intercept-only model (model 1) from meta-regression analysis of group_level_attainment_example.sav data set using group_level_attainment.sps syntax in figure 7, the b-value for the constant (labelled as coef in the spss output) is 5.27, which represents the overall effect size (i.e. attainment gap) across all modules, and this is significant (p < .001; significance is labelled as pr > |t| in the spss output). because this value was calculated by deducting the average grade of group 0 from the average grade of group 1, we can ascertain that the positive mean difference indicates that group 1 has a higher average grade (5.27% greater) than group 0, demonstrating that an attainment gap is present. the n level 1 value shows that 1056 effect sizes were included in the analysis and the n level 2 value shows that there were 40 programmes. the average level 1 n value shows that there was an average of 26.40 modules per programme in the analysis. the assumed rho of .80 is the average intercorrelation we expected there to be between module effect sizes. finally, the tau-squared estimate indicates the amount of variance between programmes at the rho level we selected, and the sensitivity analysis shows how this amount of variance would have differed at different rho levels. we can see from this range that the variance level would not have been that different if we expected there to be a smaller or larger intercorrelation between modules. model 2 is displayed in figure 8. figure 8. spss output of model 2 from meta-regression analysis of group_level_attainment_example.sav data set using group_level_attainment.sps syntax figure 8 includes all of the covariates in the meta-regression analysis. the m_ (mean) prefix denotes that values for this variable have been averaged within a programme to allow us to interpret between-programme effects, whereas the c_ (center) prefix denotes that these variables have been centered around the programme-mean to reflect within-programme (or between-module) effects. firstly, year of study is a significant negative covariate within programmes (labelled c_year in the spss output; p < .001), which means that the size of any attainment gaps decreases as students transition through university. the amount of independent study time expected of students is also a significant positive covariate within programmes (labelled c_indstu in the spss output; p = .026), with the size of attainment gaps increasing as the amount of expected independent study time increases. finally, the attendance variable is a significant positive covariate (labelled attend in the spss output; p = .006), so having a policy of attendance contributing to students’ grades (coded as 1) predicts a larger gap in attainment between groups. as with the regression analysis in tutorial 1, the b-values in the coef column can also be used to understand how much change we would expect in the attainment gap from a one unit change in these variables (based on whatever units the original variables were measured in), whilst holding the effects of all other variables constant. in this example, we aimed to determine whether there were attainment gaps based on students’ background groups, and whether a set of meso-level factors could predict these gaps. the analysis of this hypothetical data set indicates that students coded as group 1 performed better than students coded as group 0, and this overall attainment gap can be explained by the year of study of the module, the amount of independent study time expected of students in the module, and whether there is a policy of attendance contributing to students’ grades. 3.6 tutorial 4: illustrative example of analysing continuation gap data measured at a group level tutorial 4 presents an analysis of the group_level_continuation_example.sav spss data set, which was designed to closely approximate what real aggregated student data might look like when modelling continuation gaps. continuation occurs at the year-level (rather than at a module-level) within a programme; students either progress from one year to the next (and then complete their studies) or drop out. therefore, this hypothetical data set simulates a range of variables for a sample of 40 programmes, split by the year of study of the programme. the data set includes the number of students who dropped out of university during a particular year of study for two different groups of students (student background group). it also includes the same set of meso-level variables used as covariates in tutorial 3, instead averaged across all modules for the year of the programme rather than split by module (the variables are detailed in table 4). there are three types of effect size that could be calculated for these data set for use in the meta-regression analysis. the first effect size option is to calculate the risk ratio, which in this example would be the ratio between the risk of dropout for one student group compared to the risk of dropout for the other group. the second effect size option is to calculate the risk difference, which is the difference between these two risks. the final option is to calculate the or, which is what is used in logistic regression analyses. hedges et al. (2009) suggest that ors have statistical properties that often make them the best effect size option with meta-analyses of binary data, so since this statistic was also used when modelling continuation gaps at an individual level (tutorial 2), we have used it with the current data in this tutorial. however, the spss syntax for this example also computes variables for risk ratios and risk differences in case the reader would prefer to use these measures instead. it is also necessary to use the log-transformed or over the raw or to ensure that there is a balance in ratios between groups (hedges et al., 2009). when modelling continuation gaps at an individual level (tutorial 2), the outcome was the continuation variable and the initial covariate entered in the analysis was the student background group variable. extending this to a group level, the log or for the association between these two variables (see table 2) represents an effect size, so this is our outcome variable when using meta-regression.  table 4 variables included in group_level_continuation_example.sav spss data set the group_level_continuation_example.sav spss data set includes the dropout numbers for each year of a programme (i.e. years 1, 2 and 3) at a single point in time. therefore, unlike the data set used in tutorial 3, effect sizes do not include any of the same students. however, whilst effect sizes are independent, there are three separate effect sizes clustered within each programme (one for each year of students’ study). therefore, there may be hierarchical dependence between effect sizes (tipton & pustejovsky, 2015), because although different individuals contribute to each effect size, there may be a shared influence on patterns of continuation within a particular programme (e.g. because the same staff may teach across years on a programme, there may be similarities in curriculum design across a programme, etc.). it is worth noting that hierarchical dependency may also have been present in the data set used for tutorial 3, since each programme was also clustered within a department/school that may have led to shared influences between programmes. however, it has been advised that the analysis should use the weights (correlated or hierarchical) based on the most common type of dependency in the data (tanner-smith et al., 2016; tipton & pustejovsky, 2015). since we are focusing on explaining the continuation gaps themselves, the question we aim to answer in this tutorial is: are there continuation gaps based on the students’ background groups, and can a set of meso-level factors predict these gaps? for tutorial 4, we performed a meta-regression analysis using rve with hierarchical effects weights. the background group variable does not need to be entered into the analysis as a separate variable, so no covariates were included in the first model. this intercept-only model therefore tested the odds that students in one group were more likely to drop out (model 1), equivalent to the first model in tutorial 2. as with tutorial 3, new variables needed to be created to average covariates across all modules within a programme (programme-mean values for each variable to model between-programme effects) and to center these variables around the programme-mean (to model within-programme/between-year effects), so this was done for the number of assessments and percentage of independent study time variables. we then entered year of study, the average number of assessment components, and the average percentage of independent study time into the next model (model 2). the group_level_continuation.sps spss syntax file includes the code for all stages of this analysis, and again, the spss macro file, robustmeta.sps (tanner-smith & tipton, 2014), needs to be saved on the computer that will be used for performing analyses prior to running any spss syntax files, and the file path for where this macro is saved needs to be added to the syntax, as noted in the syntax instructions. figure 9 displays the spss output for the model that did not include any covariates. figure 9. spss output of the intercept-only model (model 1) from meta-regression analysis of group_level_continuation_example.sav data set using group_level_continuation.sps syntax in figure 9, the b-value for the constant is 1.54, which represents the overall log or across all years. this effect size is positive and statistically significant (p < .001). as each or is calculated based on the coding used in table 2, the positive value for the effect size indicates that students coded as group 0 were more likely to drop out than students coded as group 1, demonstrating that a continuation gap is present. we can convert this back to a raw or by using an exponential function (e.g. the excel formula =exp), which gives an or value of 4.67. the or value greater than one indicates that students coded as group 0 were more likely to drop out than students coded as group 1 by a factor of 4.67. in other words, students coded as group 0 were over four times as likely to drop out than students coded as group 1. the n level 1 value shows that 120 effect sizes were analysed, and the n level 2 value shows that there were 40 programmes. the average level 1 n value shows that there was an average of three years of study per programme in the analysis. finally, the tau-squared estimate indicates the amount of variance between programmes, and the omega-square estimate indicates the amount of variance between years within programmes. model 2 is displayed in figure 10. figure 10. spss output of model 2 from meta-regression analysis of group_level_continuation_example.sav data set using group_level_continuation.sps syntax figure 10 includes all of the covariates in the meta-regression analysis. as was done in tutorial 3, variables with an m_ prefix have been averaged within a programme to allow us to interpret between-programme effects, and variables with a c_ prefix have been centered around the programme-mean to reflect within-programme (or between-year) effects. firstly, the number of assessments is a significant positive covariate within programmes (labelled c_assess in the spss output; p = .002). since the effect size has been computed as per the coding in table 2 (with a higher number indicating that students coded as group 0 have a higher chance of dropping out), this means that as the number of assessment components increases, so do the odds of students coded as group 0 dropping out. the amount of independent study time expected of students is also a significant positive covariate both between programmes (labelled m_indstu in the spss output; p = .006) and within programmes (labelled c_indstu in the spss output; p = .024), meaning that as the amount of expected independent study time increases, so do the odds of students coded as group 0 dropping out. if the b-values are converted to ors, they can also be used to understand the likelihood to dropout occurring based on a one unit change in these variables, whilst holding the effects of all other variables constant. in this example, we aimed to determine whether there were continuation gaps based on students’ background groups, and whether a set of meso-level factors could predict these gaps. the analysis of this hypothetical data set indicates that students coded as group 0 were more likely to drop out than students coded as group 1, and this overall continuation gap can be explained by the average number of assessment components and the average amount of independent study time expected of students. 4. cautions the purpose of this primer is not to encourage teaching staff to make large-scale changes to their practice based on evidence from quantitative analyses of institutional data alone. we believe this would be very ill-advised for reasons we will now discuss. most importantly, if data are gathered and analysed in the ways demonstrated in this primer, any associations between covariates and outcomes are not causal, so we cannot claim that these micro and meso-level factors are causing any observed gaps. analyses should also be driven by theory, so variables should not simply be included just because data are available, as this may lead to false positives. of course, there is an opportunity to incorporate variables in analyses in an exploratory sense in order to extend the evidence-base on explanatory factors, but efforts should be made to avoid the “tendency to constantly extend the data inquiry to look at more variables with diminishing returns in terms of understanding” (mountford-zimdars et al., 2015, p. 22). furthermore, an explanatory factor in one area might not be the case in other areas, which limits the generalisability of analyses unless it can be argued that student populations are equivalent and representative. this is why interventions may not always prove to be successful across contexts. there is also the risk of potential biases in interpretations occurring that readers should attempt to avoid. firstly, as noted above, each time a covariate is added to a glm analysis, the coefficients for other covariates change (unless there is zero correlation between covariates, which is unlikely), so each model is context-specific (dunlap & landis, 1998). ors (see tutorials 2 and 4) cannot be compared across models for this same reason (mood, 2010). high interdependence between covariates can also be problematic, so karpen (2017) suggests the use of relative importance analysis, which transforms covariates so they become uncorrelated, as an alternative to linear regression in these situations. tonidandel and lebreton (2015) produced a free web-based tool for performing relative importance analyses. secondly, if readers make the assumption that associations found at the group level apply to the individuals within such groups, they will encounter the ecological fallacy (freedman, 1999). if an analysis of this same data at an individual level leads to different findings, ecological bias has occurred. tutorials 3 and 4 show how analyses of aggregated data can enable the introduction of group-level (i.e. mesoand macro-level) variables, which are not able to be incorporated into the individual-level analyses in tutorials 1 and 2. because mesoand macro-level explanatory factors manifest at the group level, these variables are not simply individual-level variables that have been aggregated. therefore, we were not attempting to draw micro-level conclusions about individuals with aggregated data, so we should not be committing the ecological fallacy, because the group level is the level of interest (schwartz, 1994). however, readers should be careful to avoid simply aggregating individual-level covariates for use with aggregated student data in order to avoid committing this fallacy. finally, we do not see the discussed approaches to gathering and analysing quantitative evidence as replacing rich qualitative research on differential student outcomes. qualitative data may be needed to make sense of what the quantitative findings show. for example, if a factor is strongly associated with the presence of gaps, focus groups could be held with students and teaching staff in areas where that factor is particularly prevalent to learn more about the mechanisms impacting on that factor. as a result of the discussed issues, these analyses alone could never tell the complete story, so we advise that results only be used to highlight specific variables of interest for further investigation or to focus the design of novel interventions. 5. summary and recommendations we anticipate that the guidance presented in this primer can enable universities to take a more proactive approach to monitoring and reducing gaps. gathering data on explanatory factors to which many institutions already have access, and integrating these with existing institutional data on differential outcomes, can only enhance the sophistication of any analyses. to elucidate this point, we can look at the hypothetical findings of the four illustrative examples; although they do not draw on real data, if the findings were real we could conclude that attainment and continuation gaps are present in the data sets and a range of microand meso-factors can account for these gaps, including students’ entry qualification, prior attainment, their sense of belonging to the university, their year of study, the amount of independent study time expected of them, and whether there is a policy of attendance contributing to their grades. in relation to the conceptual model in figure 1, these findings would highlight the specific aspects of curriculum and assessment factors that play a role in differential outcomes, rather than only viewing this in terms of the broad categories of ‘curriculum’ and ‘assessment’. the meta-regression techniques are also novel and parsimonious ways of directly analysing attainment and continuation gaps using aggregated data that avoid issues with data privacy. as a result, we believe this primer will be of particular interest to staff with strategic responsibilities who report on the presence of gaps within their institution. our guidance should raise awareness around the causes of differential outcomes and enable these staff to provide teaching staff with more nuanced evidence that can be used for focusing conversations around issues that have been identified and designing context-specific interventions, whether these are universal or targeted at specific student groups. for example, based on the hypothetical findings from the illustrative examples, teaching staff might decide to speak with students about how they feel about the balance of independent study time in their modules and any policies around attendance contributing to their grades. interventions might then be designed based on these aspects of the curriculum and assessment. this evidence is likely to be relevant to specific modules/programmes/disciplines, so rather than just replicating current interventions that may only be effective within specific contexts, findings from analyses could increase institutions’ understanding about the factors that are salient for their own students. this increased precision in the evidence-base used for designing novel interventions means that there will be less investment in approaches that are not likely to work. implementation of the suggested analyses thus represents better value for money in any initiatives undertaken by institutions. furthermore, this will also assuage deficit assumptions that differential outcomes are due to student deficiencies, which could be used to excuse teaching staff from their own responsibilities for attempting to tackle gaps. although this primer particularly attends to opportunities for exploring meso-level variables, largely because these variables are under the control of the university, there is the potential for analysing cross-institutional data, possibly internationally, to consider the role of quantitative macro-level explanatory factors in differential student outcomes. meta-regression techniques in particular may support these endeavours, since there are likely to be far fewer issues sharing aggregated data between institutions than there would be with individual-level data. additionally, if metrics for attainment and continuation are different between institutions, it is possible to standardise them for comparison using these techniques. cross-institutional comparisons would enable a better understanding of how institution type and context impact on gaps. finally, beyond the practical implications discussed, an important contribution of this primer for researchers is that it provides them with methods to extend the literature on differential outcomes, potentially uncovering new insights about multi-level causes of attainment and continuation gaps. one way researchers could do this is by considering how to operationalise some of the microand meso-level factors in the conceptual model in figure 1 that have not previously been quantitatively measured and analysed. for example, data on research income and research and teaching quality metrics could be used as measures of institutional and disciplinary culture (see boliver, 2015, for an analysis of publicly available data encompassing these variables). this approach may then call attention to the specific facets of the conceptual model that merit closer scrutiny in future research. depending on the representativeness of the student populations investigated, potential future research using the approaches in this primer might also consider how meso-level factors manifest differently based on variation in macro-level factors, and how micro-level factors differ based on variation in meso-level factors. for instance, macro-level factors, such as the structure and hierarchy of an institution could affect the design of its learning environments at the meso-level (e.g. teaching-led universities might prioritise innovative teaching approaches more than research-intensive institutions), so moderating effects between micro-, meso-, and macro-level variables on gaps could be explored. as previously noted, differential student outcomes are quantitatively assessed (jones, 2018), so if researchers adopt the quantitative methods used in this primer, they could reveal more about the explanatory power of specific factors, which could enrich the largely qualitative evidence-base. furthermore, by revealing more about potential predictors of gaps, the methods could also open up new conversations with students, providing fresh avenues for future qualitative research in this area. therefore, this primer offers significant scope for developing theoretical perspectives on differential outcomes beyond current understandings in the literature. it is only by fully exploring potential causes of gaps at all levels that we will fully understand the role of student diversity in educational transitions. keypoints explanatory factors for differential student outcomes occur at micro-, mesoand macro-levels universities can model explanatory factors for differential student outcomes at multiple levels by linking previously unconnected quantitative data attainment/continuation gaps can be modelled with regression analyses of individual-level data combined with data on micro-level explanatory factors attainment/continuation gaps can be modelled with meta-regression analyses of aggregated data combined with data on meso-level explanatory factors universities could use the discussed approaches to design more context-specific interventions for addressing differential student outcomes acknowledgements the authors thank professor chris fife-shaw, dr anesa hosein, nicholas moore and dr james munro for their insightful comments on earlier drafts of this manuscript. footnotes 1 throughout this primer we refer to different aspects of university based on terminology usually used in the uk higher education context (e.g. degree courses are referred to as programmes, subject units are referred to as modules, etc.). 2 all analyses can also be performed using other software packages, such as stata and r, so .csv files (for the data sets) and .txt files (for the syntax) have additionally been provided to enable readers to adapt these files for use with their preferred software. 3 our use of the term covariate is synonymous with predictor/independent variable (or moderator in meta-analyses). 4 binary/dichotomous variables (i.e. variables with only two categories) can be entered into regression models as covariates and treated the same as continuous variables if they are coded as 0 and 1. however, categorical variables with more than two categories need to be converted into a set of binary variables, known as dummy coding, before inclusion. field (2018, pp. 509–516) provides an accessible overview of dummy coding. 5 if different measures of attainment are used (e.g. combining grade percentages with grade point average), or if the measure of attainment is not meaningful (e.g. the attainment metric is specific to a particular institution), the standardised difference in means should be used as the effect size measure instead of the raw difference (hedges et al., 2009); code for computing standardised effect sizes has been included in the spss syntax file. 6 rve can also be performed in r using the robumeta package (fisher et al., 2017; see tanner-smith et al., 2016, for a tutorial), and in stata using the robumeta.ado macro (see tanner-smith & tipton, 2014, for syntax and a tutorial). 7 small-sample adjustments in rve are not available in the spss macro, but they are part of the stata macro and r package for rve (tanner-smith et al., 2016). 8 the degrees of freedom in rve analyses are constrained by the number of studies (in this case, the number of programmes) being analysed, not the number of effect sizes (tanner-smith & tipton, 2014). therefore, readers should be cautious about including too many covariates in analyses with only a limited number of programmes being analysed. 9 as with regression analyses performed on individual-level data, it is possible to check for sources of bias, meeting of assumptions, and issues with multicollinearity. however, more sophisticated software is needed for these tests than the spss macro used for tutorials 3 and 4, so findings should be interpreted with some caution in the absence of these tests. 10 this macro has also been converted into an spss plug-in to enable analyses to be run through the main spss menus. this can be downloaded from https://github.com/ahmaddaryanto/meta_analysis_spss_macros. references alin, a. (2010). multicollinearity.wiley interdisciplinary reviews: computational statistics, 2(3), 370–374. https://doi.org/10.1002/wics.84 boliver, v. (2015). are there distinctive clusters of higher and lower status universities in the uk? oxford review of education, 41(5), 608–627. https://doi.org/10.1080/03054985.2015.1082905 broecke, s., & nicholls, t. (2007). ethnicity and degree attainment (dfes research report rw92). dfes. cheung, m. w. l. (2019). a guide to conducting a meta-analysis with non-independent effect sizes. neuropsychology review, 29 (4), 387–396. https://doi.org/10.1007/s11065-019-09415-6 courville, t., & thompson, b. (2001). use of structure coefficients in published multiple regression articles: β is not enough. educational and psychological measurement, 61(2), 229–248. https://doi.org/10.1177/0013164401612006 cousin, g., & cuerton, d. (2012). disparities in student attainment (disa). higher education academy (hea). https://www.advance-he.ac.uk/knowledge-hub/disparities-student-attainment dunlap, w. p., & landis, r. s. (1998). interpretations of multiple regression borrowed from factor analysis and canonical correlation. the journal of general psychology, 125(4), 397–407. https://doi.org/10.1080/00221309809595345 field, a. p. (2018). discovering statistics using ibm spss statistics (5th ed.). sage. field, a. p., & gillett, r. (2010). how to do a meta-analysis.british journal of mathematical and statistical psychology, 63(3), 665–694. https://doi.org/10.1348/000711010x502733 fisher, z., tipton, e., & zhipeng, h. (2017). robumeta (version 2.0) [computer software]. https://cran.r-project.org/package=robumeta france, m. k., finney, s. j., & swerdzewski, p. (2010). students’ group and member attachment to their university: a construct validity study of the university attachment scale. educational and psychological measurement, 70(3), 440–458. https://doi.org/10.1177/0013164409344510 freedman, d. a. (1999). ecological inference and the ecological fallacy (technical report 549) . https://statistics.berkeley.edu/sites/default/files/tech-reports/549.pdf gravett, k. (2019). troubling transitions and celebrating becomings: from pathway to rhizome. studies in higher education, 1–12. https://doi.org/10.1080/03075079.2019.1691162 hedges, l. v., borenstein, m., higgins, j. p. t., & rothstein, h. r. (2009). introduction to meta-analysis. john wiley & sons, ltd. hedges, l. v., tipton, e., & johnson, m. c. (2010). robust variance estimation in meta-regression with dependent effect size estimates. research synthesis methods, 1(1), 39–65. https://doi.org/10.1002/jrsm.5 jones, s. (2018). expectation vs experience: might transition gaps predict undergraduate students’ outcome gaps? journal of further and higher education, 42(7), 908–921. https://doi.org/10.1080/0309877x.2017.1323195 karpen, s. c. (2017). misuses of regression and ancova in educational research. american journal of pharmaceutical education, 81(8), 84–85. https://doi.org/10.5688/ajpe6501 lens, d., & levrau, f. (2020). can pre-entry characteristics account for the ethnic attainment gap? an analysis of a flemish university. research in higher education, 61(1), 26–50. https://doi.org/10.1007/s11162-019-09554-y miller, m. (2016). the ethnicity attainment gap: literature review . the university of sheffield widening participation research & evaluation unit. https://www.sheffield.ac.uk/polopoly_fs/1.661523!/file/bme_attainment_gap_literature_review_external_-_miriam_miller.pdf mood, c. (2010). logistic regression: why we cannot do what we think we can do, and what we can do about it. european sociological review, 26(1), 67–82. https://doi.org/10.1093/esr/jcp006 mountford-zimdars, a., sabri, d., moore, j., sanders, j., jones, s., & hiagham, l. (2015). causes of differences in student outcomes. higher education funding council for england (hefce). https://webarchive.nationalarchives.gov.uk/20180405123119/http://www.hefce.ac.uk/pubs/rereports/year/2015/diffout/ mountford-zimdars, a., sanders, j., moore, j., sabri, d., jones, s., & higham, l. (2017). what can universities do to support all their students to progress successfully throughout their time at university? perspectives: policy and practice in higher education, 21 (2–3), 101–110. https://doi.org/10.1080/13603108.2016.1203368 ofs. (n.d.). continuation and attainment gaps. https://www.officeforstudents.org.uk/advice-and-guidance/promoting-equal-opportunities/evaluation-and-effective-practice/continuation-and-attainment-gaps/ pigott, t. d., & polanin, j. r. (2020). methodological guidance paper: high-quality meta-analysis in a systematic review. review of educational research, 90(1), 24–46. https://doi.org/10.3102/0034654319877153 schwartz, s. (1994). the fallacy of the ecological fallacy: the potential misuse of a concept and the consequences. american journal of public health, 84(5), 819–824. https://doi.org/10.2105/ajph.84.5.819 singh, g. (2011). black and minority ethnic (bme) students participation in higher education: improving retention and success . higher education academy (hea). https://www.heacademy.ac.uk/system/files/bme_synthesis_final.pdf stevenson, j. (2012). black and minority ethnic student degree retention and attainment. higher education academy (hea). https://www.heacademy.ac.uk/system/files/bme_summit_final_report.pdf tanner-smith, e. e., & tipton, e. (2014). robust variance estimation with dependent effect sizes: practical considerations including a software tutorial in stata and spss. research synthesis methods, 5 (1), 13–30. https://doi.org/10.1002/jrsm.1091 tanner-smith, e. e., tipton, e., & polanin, j. r. (2016). handling complex meta-analytic data structures using robust variance estimates: a tutorial in r. journal of developmental and life-course criminology, 2 (1), 85–112. https://doi.org/10.1007/s40865-016-0026-5 thompson, c. g., kim, r. s., aloe, a. m., & becker, b. j. (2017). extracting the variance inflation factor and other multicollinearity diagnostics from typical regression results. basic and applied social psychology, 39(2), 81–90. https://doi.org/10.1080/01973533.2016.1277529 tieben, n. (2020). non-completion, transfer, and dropout of traditional and non-traditional students in germany. research in higher education, 61(1), 117–141. https://doi.org/10.1007/s11162-019-09553-z tinto, v. (1993). leaving college: rethinking the causes and cures of student attrition (2nd ed.). university of chicago press. tipton, e., & pustejovsky, j. e. (2015). small-sample adjustments for tests of moderators and model fit using robust variance estimation in meta-regression. journal of educational and behavioral statistics, 40(6), 604–634. https://doi.org/10.3102/1076998615606099 tonidandel, s., & lebreton, j. m. (2015). rwa web: a free, comprehensive, web-based, and user-friendly tool for relative weight analyses. journal of business and psychology, 30(2), 207–216. https://doi.org/10.1007/s10869-014-9351-z uuk & nus. (2019). black, asian and minority ethnic student attainment at uk universities: #closingthegap . universities uk & national union of students. https://www.universitiesuk.ac.uk/news/pages/universities-acting-to-close-bame-student-attainment-gap.aspx woodfield, r. (2014). undergraduate retention and attainment across the disciplines. higher education academy (hea). https://www.advance-he.ac.uk/knowledge-hub/undergraduate-retention-and-attainment-across-disciplines codepen pekrun frontline learning research vol.8 no. 3 (2020) 168 194 issn 2295-3159 commentary: self-report is indispensable to assess students’ learning reinhard pekruna auniversity of essex, united kingdom & australian catholic university, sydney, australia abstract self-report is required to assess mental states in nuanced ways. by implication, self-report is indispensable to capture the psychological processes driving human learning, such as learners’ emotions, motivation, strategy use, and metacognition. as shown in the contributions to this special issue, self-report related to learning shows convergent and predictive validity, and there are ways to further strengthen its power. however, self-report is limited to assess conscious contents, lacks temporal resolution, and is subject to response sets and memory biases. as such, it needs to be complemented by alternative measures. future research on self-report should consider not only closed-response quantitative measures but also alternative self-report methodologies, make use of within-person analysis, and investigate the impact of respondents’ emotions on processes and outcomes of self-report assessments. keywords: self-report; emotion; motivation; metacognition; self-regulated learning info corresponding author email: pekrun@lmu.de doi: http://doi.org/10.14786/flr.v8i3.637 introduction self-report is indispensable for any more nuanced assessment of mental states. while it is possible to examine general physiological properties of thought and affect using brain-imaging, and their consequences through performance tests and behavioral observation, assessing the contents and complex cognitive processes involved in human thinking, emotion, and motivation requires self-report. as such, self-report was a primary assessment method in psychology and education from early on, and it continued to be a primary method throughout all developmental phases in the history of these disciplines, even in the prime time of behaviorism early in the 20th century. however, self-report also has limitations. self-report is restricted to processes that are accessible to consciousness; is typically limited to assess contents that can be verbally described; can be subject to various biases; and is always lagging behind the processes it aims to assess, even if only for seconds, which implies that it lacks the temporal resolution needed to capture the real-time dynamics of learning. given these problems, it is important to critically scrutinize the power of self-report methods to capture the constructs they intend to assess, and to develop strategies to improve their validity. the papers in this special issue are excellent examples for both directions. specifically, all eight papers examine the validity of specific self-report instruments relative to proposed distributions of scores and relations with other variables. in addition, two of the papers also explore ways to improve validity. in the following sections, i first address the nature of self-report and its advantages and drawbacks. next, i discuss the advances in analyzing and improving the validity of self-report measures that are represented in the contributions to this special issue. in conclusion, i outline three directions for future research. 1. what is self-report? self-report uses participants’ verbal responses to assess their cognition, emotion, motivation, behavior, or physical state. when thinking about self-report, what often comes to mind first is structured questionnaires measuring some kind of personality trait. however, while structured questionnaires are used frequently, the most commonly employed self-report instrument likely is the clinical interview, which typically has a very different format as compared with closed-response questionnaires. by implication, to judge self-report, it is important to consider that this method can take very different forms. self-report can be structured or unstructured; retrospective or concurrent; oral or written; qualitative of quantitative; one-dimensional or multi-dimensional; paper-and-pencil or online; and can comprise single or multiple items (see pekrun & bühner, 2014, for an overview). as such, self-report not only includes structured multi-item questionnaire scales, but also open-ended interviews, single-item momentary reports, unstructured think-aloud protocols, etc. while all these methods share properties of relying on participants’ ability to self-assess and report about the variables under investigation, they differ vastly in terms of structure, temporal resolution, and metric used. as such, it is important to keep in mind that any findings on the validity of self-report instruments, and on ways to improve it, may be specific to some variant of self-report and not be generalizable to other variants. in the current special issue, all of the eight contributions consider multi-item questionnaire scales using closed formats (i.e., defined items and response options). rogier et al. (2020) additionally included a think-aloud protocol. as such, with this exception, the contributions focus on quantitative, structured self-report measures. such measures are well suited to answer quantitative research questions that are defined a priori. however, they are less suited to answer exploratory questions and to gain a more nuanced picture of respondents’ subjective world of multi-layered thoughts and perceptions, which can transcend researchers’ prior conceptions as represented in closed-format scales. for such purposes, qualitative self-report methods are needed. overall, to make progress in research on learning and instruction, it is often useful to employ a mix of quantitative and qualitative self-report, with qualitative methods used to explore new territory and gauge in-depth explanations, and quantitative method to test theoretical hypotheses in more rigorous ways (see, e.g., pekrun et al., 2002). 2. benefits and drawbacks of self-report self-report has clear advantages. first, in contrast to other types of assessment, self-report allows assessment of all types of psychological processes. observation can assess visible behavior, achievement tests can measure cognitive performance, neuro-imaging the activation of brain areas, and physiological analysis the arousal of peripheral systems, but self-report can be used to assess all of the affective, cognitive, physiological, and behavioral processes that are part of self-regulated learning – all of these processes can be represented in the human mind and can be reported accordingly. second, self-report can render a more differentiated assessment of human thinking than any other method. as such, for a nuanced description of emotions, motivation, and metacognition during learning, self-report is needed. third, self-report is more economical than other methods. self-report may be the only method applicable in some types of studies, such as large-scale student assessments. self-report also has disadvantages. as noted, self-report is limited to the assessment of processes that are accessible to consciousness. responses that cannot be represented mentally need to be assessed with other methods. another important limitation is the use of language (although self-report can also employ non-verbal communication). research has shown that terms describing psychological processes tend to be used in consistent ways across languages (e.g., fontaine et al., 2013), but there can nevertheless be differences in semantic understanding across cultures and learners. by implication, measurement equivalence of self-report instruments across groups should not merely be assumed but needs to be established empirically. furthermore, limitations result from the fact that self-report is under respondents’ control. whereas it may be difficult to alter one’s level of physiological activation, reports about perceived activation can easily be changed. as such, depending on motivation and preferences for response options, self-report can be subject to various response biases, such as social desirability. finally, self-report is also subject to memory biases. this is especially true for retrospective self-report that is administered at a later point in time and requires recollection of information from autobiographical memory, but is also true for state self-report asking respondents how they feel or what they think right now – self-report is lagging behind the phenomena it captures, even if only for seconds. as such, self-report inevitably lacks the temporal resolution needed to examine the real-time dynamics of psychological processes. this is even true for momentary methods such as experience sampling or think-aloud protocols as used in the contributions by moeller et al. (2020) and rogier et al. (2020). even these methods cannot reach the temporal granularity of concurrent physiological or behavioral-observational methodologies. as such, self-report needs to be complemented with other methods for many research purposes. for making progress in research on learning, multi-channel assessments of motivation, emotion, and metacognition including self-report along with observational and physiological methods as well as behavioral trace data are especially promising (azevedo et al., 2018; lajoie et al., in press). 3. examining the validity of self-report measures six of the eight contributions to this special issue focus on examining the validity of (quantitative) self-report measures and developing methods to examine validity. iaconelli and wolters (2020) investigated the impact of insufficient effort in responding on university students’ self-report scores for their beliefs and behaviors during self-regulated learning. self-report was assessed as part of students’ coursework. rates of insufficient effort were low, and reported relations between variables were robust against including students with insufficient effort, suggesting that inattentive responding does not represent a major threat to validity (at least under the situational conditions of the study). rogiers et al. (2020) examined secondary school students’ retrospective self-report of learning strategies used while learning from a text in combination with think-aloud data obtained during the same session. the data from the retrospective self-report were used to classify different types of learners, and the findings show that these types differed systematically in their learning process as assessed through the think-aloud protocol, thus attesting to the convergent validity of these two – very different – types of self-report instruments. extending the perspective beyond individual learning, vriesema and mccaslin (2020) used self-report to assess secondary school students’ general test anxiety, their attitudes towards school, and their behavior and emotions during group work in mathematics. there were clear links between self-reported behavior and emotions related to the group work situation, but less so with the test anxiety measure. these findings are consistent with the specificity matching principle (see, e.g., swann et al., 2007): variables show stronger relations when being matched in terms of situational specificity than when not being matched, and the present results suggest that self-report measures can demonstrate validity when attending to this principle. van halem et al. (2020) used the motivated strategies for learning questionnaire (mslq; pintrich et al., 1991) as well as online trace data to assess undergraduate students’ motivation and learning strategies during a statistics course. there was no direct conceptual match between the mslq and online trace data constructs. nevertheless, there were substantial relations between time investment as assessed by the mslq, on the one hand, and trace data, on the other, thus demonstrating convergent validity of the two types of measures. furthermore, both the mslq scores and the online trace data contributed to explaining students’ course performance, thus supporting predictive validity for both types of measures. moeller et al. (2020) used experience sampling methodology (esm) with a two-item measure of situational interest to capture the developmental dynamics of university students’ interest in a series of lectures over one semester. in contrast to traditional esm designs, a fixed rather than random schedule of assessments was used, which facilitated aggregation of assessments across participants. the findings of cross-classified multilevel analysis show that there was substantial variation of interest scores between students as well as within and between lectures, thus documenting the usefulness of situational self-report scores to decompose these sources of variance. finally, in terms of developing additional methods for testing validity, chauliac et al. (2020) observed university students’ gaze behavior while answering items on a questionnaire assessing habitual use of different cognitive strategies during learning from texts. there were systematic links between number and duration of fixations, on the one hand, and the consistency of answering different items from the same scale, on the other. the findings demonstrate that eyetracking has great potential to examine processes of responding to verbal stimuli as presented in self-report scales, suggesting that this methodology could contribute to examining the validity of scores derived from these scales. taken together, these six contributions attest to the potential validity of self-report in assessing students’ learning. there were clear links (1) between quantitative self-report scores for different constructs, as well as (2) between these scores, on the one hand, and think-aloud protocols, online trace data, and academic performance, on the other. while not all of these links were fully robust and significant, they nevertheless document that self-report continues to be useful in measuring facets of students’ learning.   4. improving validity how can we further improve the validity of self-report measures? two of the contributions address this question. fryer and nakao (2020) examined the impact of type of response scale on levels, reliability, and factorial validity of self-reported task interest and its links with prior and subsequent domain interest in a sample of phd students. their study included two traditional formats (labelled categorical scale and visual analogue scale) as well as two more recent formats (slider and swipe scales). reliability and factorial validity were nearly identical across the formats, and mean scores for interest in different tasks did not show systematic differences either. however, predictive validity for future interest tended to be higher for the slider and swipe versions than for the two traditional formats. this is promising and should stimulate research on how to further optimize response scales and their presentation. durik and jenkins (2020) analyzed the link between undergraduate students’ self-reported interest, their certainty in their answers, and their self-reported behavioral engagement in various subjects. the findings show that interest and certainty were related in a curvilinear fashion in most of the subjects; high certainty was associated with either low or high interest scores. furthermore, the link between interest and behavioral engagement was substantially stronger for students with high certainty in their reported level of either individual or situational interest, and not even significant for students with low certainty in their situational interest. these findings suggest that including certainty ratings can increase the validity of self-report in predicting students’ behavior. as such, although replication is needed, they represent a potential breakthrough in boosting the validity of interest measures. in both of these contributions, it remains open to question how the observed effects can be explained. for the effects of certainty, as noted by durik and jenkins (2020), it seems possible that clear beliefs in the strength or weakness of one’s interest contributes to using interest as a guide for action, in contrast to being unclear about one’s interest which may subject action to situational conditions. research on the origins and outcomes of certainty is needed to examine this possibility. 5. directions for future research 5.1 scoping a broader range of self-report methods the contributions to this special issue focus on self-report methods using written verbal statements as stimuli and closed-format options to respond, with quantitative methods employed to analyze responses. however, as noted earlier, there are various alternative formats that are equally important. each of these formats has its own advantages and disadvantages. specifically, a substantial amount of social science research uses oral formats (specifically, interviews) and open-ended answers, either written or provided orally. to an extent, these formats may be subject to similar biases as written closed-format self-report, including response sets and memory biases. however, there also may be differences, especially in terms of strategies to reduce these problems. for example, motivation to respond in socially desirable ways rather than telling the truth can be reduced by generating trust in interviewees that their data will be kept confidential, and memory biases can be reduced through cognitive interviewing techniques that optimize recall. substantial progress in suitable methods has been made in forensic psychology and research on testimony (see, e.g., bowles & sharman, 2014; brown & lamb, 2015). it would be worth exploring if some of these strategies could be made fruitful for educational research as well. this may be especially important for research on learning in young children (preschool, kindergarten, and the early elementary school years). qualitative self-report methodology using open-ended answers is especially important when exploring new research questions, but also when wanting to understand unexpected or paradoxical findings that can be explored with in-depth interviews. how to best structure questions, analyze answers, and aggregate qualitative self-report findings across studies currently is a field of intense methodological debate (see, e.g., clark, 2016; snelson, 2016). mainstream quantitative research on self-report should attend to these developments, and research is needed on how to better integrate different self-report methods and the resulting evidence (e.g., in terms of convergent parallel, exploratory sequential, or explanatory sequential mixed-method study designs; creswell, 2014; creswell & plano clark, 2011). 5.2 importance of within-person research for validating self-report similar to other types of research in education and psychology, the vast majority of investigations using self-report have relied on between-person study designs, including most of the contributions to this special issue (the moeller et al., 2020, contribution is a notable exception). between-person research is suited to examine individual differences and interindividual relations between variables. however, it is not suited to investigate the within-person psychological functioning that is addressed in theories of students’ motivation, emotion, and self-regulated learning. intraindividual and interindividual correlations are statistically independent, and there is no easy way to infer one from the other, except when conditions of ergodicity hold. these conditions include homogeneity of functional relations across persons and stationarity over time, conditions that are often not met (voelkle et al., 2014). as such, to study motivation, emotion, and strategy use during learning, it is best to examine these processes within persons (murayama et al., 2017). to ensure generalizability, the variation of within-person relations across persons needs to be analyzed – if there is little variation, then relations are generalizable and nomothetic conclusions can be reached. the relevance of within-person research has important implications for the validation of self-report measures. in research on learning, some of these measures pertain to trait-like characteristics of students and are used to gauge between-person differences. for example, measures of trait-like individual interest may be used to assess differences in interest between students. for these measures, it is appropriate to use between-person designs to examine validity. however, whenever the purpose is to assess individual development over time, or personal functioning during learning, then it is more adequate to use within-person designs to validate self-report methods. between-person designs can render misleading conclusions, and resulting findings can underor overestimate validity relative to theories of individual learning. 5.3 the role of emotion and emotion regulation in self-report as human performance more generally, adequately responding to self-report instruments requires both competence and motivation. competence includes being able to understand questions, to retrieve relevant information from long-term memory or current working memory, and to integrate the retrieved information such that a decision about an adequate answer can be reached. current models of self-report largely focus on these cognitive processes, and process-oriented methods to validate self-report items focus on techniques of cognitive validation (castillo-diaz & padilla, 2013; karabenick et al., 2007). motivation includes wishes to veridically answer questions, either to get a valid self-assessment (e.g., in contexts such as career counselling or psychotherapy) or to help researchers in their attempts to understand reality, as well as desires to appear to others or oneself as a socially desirable person. motivation has been examined especially in research on social desirability (see, e.g., gignac, 2013), and there is a long-standing tradition of controlling for desirability in studies of personality.   however, beyond cognition and motivation, it seems likely that emotions also play a critically important role in self-report. emotions are defined as affective responses to personally important events. as such, whenever self-report touches issues that are personally relevant, emotions are likely to be aroused. this can be emotions that are already associated with a given topic in memory, such as anxiety when retrieving recollections of prior exams, or emotions that are generated during the process of reading self-report items. in addition, emotions that are elicited by the task of responding and the social context of the assessment can play a role, such as sympathy for the experimenter administering a questionnaire, social anxiety of disclosing personal information, or anger about redundancy of items in lengthy multi-item instruments. it is reasonable to assume that these emotions can substantially influence self-report responses. this can happen through the influence of emotions on retrieval of information from memory (e.g., in terms of mood-congruent retrieval), on integrating memory information in different ways (e.g., holistically in positive mood and detail-oriented in negative mood), on current motivation to persist in answering questions, and on motivation to answer in specific ways (e.g., according to social desirability when being socially anxious about one’s responses). furthermore, ways to regulate emotions may play a role as well. for example, unpleasant emotions triggered by emotionally negative self-report items may be so strong and aversive that one seeks to downregulate them right away, even before answering the item. as a result, the answer may no longer represent the original emotional response to the item. emotions can contribute to changes of the objects of self-report measurement during the process of measuring them – a phenomenon that can render resulting scores an artefact of the response process. research exploring these possibilities is largely lacking. self-report methodologists could team up with memory researchers, social psychologist, and affective scientists to investigate these possible influences of emotions on self-report. the results could inform psychological and educational measurement in terms of shaping instruments and the social situations of assessment in ways that are both emotionally beneficial and suited to increase validity. 6. conclusion self-report is indispensable for any more fine-grained assessment of mental processes, including students’ motivation, emotions, cognitive strategies, and metacognition during learning. certainly, self-report has limitations in terms of assessing conscious processes only, being subject to biases, and not providing the temporal resolution needed to assess some of these processes. nevertheless, the evidence reported in the contributions to this special issue clearly document that self-report continues to be a valid way to assess processes of learning. to further boost its validity, triangulation of different self-report methods (such as closed items and open-format think-aloud protocols) as well as integration of self-report into multi-channel assessments can be helpful. to make further progress in examining and improving the psychometric quality of self-report methods, it may be useful to consider a broad range of different variants of self-report, to consider the influence of respondents’ emotions on their self-report, and to complement traditional between-person study designs with intraindividual analysis. keypoints self-report is required for a nuanced assessment of mental processes self-report is indispensable to assess learning-related emotions, motivation, meta-cognition, and self-regulation learning-related self-report scales show convergent and predictive validity self-report needs to be amended by alternative measures because it lacks temporal resolution and is subject to response sets and memory biases future research should consider a broader range of self-report methods, within-person analysis, and the impact of emotions on self-report references azevedo, r., taub, m., & mudrick, n.v. (2018). using multi-channel trace data to infer and foster self-regulated learning between humans and advanced learning technologies. in d. schunk & greene, j.a (eds.), handbook of self-regulation of learning and performance (2nd ed., pp. 254-270). routledge. bowles, p. v., & sharman, s. j. (2014). a review of the impact of different types of leading interview questions on child and adult witnesses with intellectual disabilities. psychiatry, psychology and law, 21, 205-217. http://doi.org/10.1080/13218719.2013.803276 brown, d. a., & lamb, m. e. (2015). can children be useful witnesses? it depends on how they are questioned. child development perspectives, 9, 250-255. http://doi.org/10.1111/cdep.12142 castillo-diaz, m., padilla, j.-l. (2013). how cognitive interviewing can provide validity evidence of the response processes to scale items. social indicators research, 114, 963–975. http://doi.org/101007/s11205-012-0184-8 chauliac m., catrysse, l., gijbels, d., & douce, v. (2020). it is all in the surv-eye: can eye tracking data shed light on the internal consistency in self-report questionnaires on cognitive processing strategies? frontline learning research, 8(2), 26-39. http://doi.org/10.14786/flr.v8i3.489 clark, a. m. (2016). why qualitative research needs more and better systematic review. international journal of qualitative methods, 15, 1-3. http://doi.org/10.1177/1609406916672741 creswell, j. w. (2014). a concise introduction to mixed methods research. sage. creswell, j. w., & plano clark, v. l. (2011). designing and conducting mixed methods research (2nd ed.). sage. durik, a., & jenkins, j. (2020). variability in certainty of self-reported interest: implications for theory and research. frontline learning research, 8(2), 86-104. http://doi.org/10.14786/flr.v8i3.491 fontaine, j. j. r., scherer, k. r., & soriano, c. (eds.). (2013). components of emotional meaning: a sourcebook. oxford university press. fryer, l. & nakao k. (2020). the future of survey self-report: an experiment contrasting likert, vas, slide, and swipe touch interfaces. frontline learning research, 8(2), 10-25. http://doi.org/10.14786/flr.v8i3.501 gignac, g. e. (2013). modeling the balanced inventory of desirable responding: evidence in favor of a revised model of socially desirable responding. journal of personality assessment, 95, 645–656. http://doi.org/10.1080/00223891.2013.816717 iaconelli, r., & wolters, c. a. (2020). insufficient effort responding in surveys assessing self-regulated learning: nuisance or fatal flaw? frontline learning research, 8(2), 105-127. http://doi.org/10.14786/flr.v8i3.521 karabenick, s. a., woolley, m. e., friedel, j. m., ammon, b. v., blazevski, j., ree bonney, c., . . . kelly, k. l. (2007). cognitive processing of self-report items in educational research: do they think what we mean? educational psychologist, 42, 139–151. http://doi.org/10.1080/00461520701416231 lajoie, s. p., pekrun, r., azevedo, r., & leighton, j. p. (in press). understanding and measuring emotions in technology-rich learning environments. learning and instruction. moeller, j., viljaranta, j., & kracke, b. & dietrich, j. (2020). disentangling objective characteristics of learning situations from subjective perceptions thereof, using an experience sampling method design. frontline learning research, 8(2), 63-85. http://doi.org/10.14786/flr.v8i3.529 murayama, k., goetz, t., malmberg, l.-e., pekrun, r., tanaka, a., & martin, a. j. (2017). within-person analysis in educational psychology: importance and illustrations. in d. w. putwain & k. smart (eds.), british journal of educational psychology monograph series ii: psychological aspects of education – current trends: the role of competence beliefs in teaching and learning (pp. 71-87). wiley. pekrun, r., & bühner, m. (2014). self-report measures of academic emotions. in r. pekrun & l. linnenbrink-garcia (eds.), international handbook of emotions in education (pp. 561-579). taylor & francis. pekrun, r., goetz, t., titz, w., & perry, r. p. (2002). academic emotions in students’ self-regulated learning and achievement: a program of qualitative and quantitative research. educational psychologist, 37, 91-106. pintrich, p. r., smith, d. a. f., garcia, t., & mckeachie, w. j. (1991). a manual for the use of the motivated strategies for learning questionnaire (mslq) (tech. report no. 91-b-004). board of regents, university of michigan, ann arbor, mi. rogiers, a., merchie, e., & van keer, h. (2020). opening the black box of students’ text-learning processes: a process mining perspective. frontline learning research, 8(2), 40-62. http://doi.org/10.14786/flr.v8i3.527 swann jr, w. b., chang-schneider, c., & mcclarty, k. l. (2007). do people's self-views matter? self-concept and self-esteem in everyday life. american psychologist, 62, 84–94. http://doi.org/10.1037/0003-066x.62.2.84 van halem, n., van klaveren, c., drachsler, h., schmitz, m., & cornelisz, i. (2020). tracking patterns in self-regulated learning using students’ self-reports and online trace data. frontline learning research, 8(2), 142-163. http://doi.org/10.14786/flr.v8i3.497 voelkle, m. c., brose, a., schmiedek, f. & lindenberger, u. (2014). towards a unified framework for the study of between-person and within-person structures: building a bridge between two research paradigms. multivariate behavioral research, 49, 193–213. http://doi.org/10.1080/00273171.2014.889593 vriesema, c. c., & van klaveren, c. (2020). experience and meaning in small-group contexts: fusing observational and self-report data to capture self and other dynamics. frontline learning research, 8(2), 128-141. http://doi.org/10.14786/flr.v8i3.493 frontline learning research vol. 11 no. 2 (2023) 49-77 issn 2295-3159 corresponding author: tiina soini, åkerlundinkatu 5, 33014 tampere university, finland, tiina.soini-ikonen@tuni.fi doi: https://10.14786/flr.v11i2.983 primary and lower secondary students’ learning agency and social support tiina soini1, janne pietarinen2, sanna ulmanen1,, henrika anttila3, & kirsi pyhältö3 1tampere university, finland 2university of eastern finland, finland 3university of helsinki, finland article received 2 november 2021/revised 1 september 2023/accepted 27 september 2023/available online 21 november 2023 abstract making initiatives and having ownership over one’s learning is a key for applying and creating knowledge, acquiring new abilities, actively steering one’s life and the engaging in change in society. understanding the preconditions for such learning should be in the core of designing learning environments and the primary interest of frontline learning research. this study focuses on exploring students’ sense of their learning agency in studying and the role of teacher and peer support in cultivating it. we examined how primary (grades 1-6) and lower secondary school students (grades 7-9) perceive their learning agency (la), its relationship with the experienced teacher and peer support in studying. also, differences between the girls and boys, and schools located in low and high ses neighborhoods was examined. we assessed the structure and level of learning agency by using a new measurement and explorative structural equation modeling (esem). results show that learning agency consists of interdependent elements of motivation to learn, self-efficacy beliefs about learning and strategies for learning in meaning making, problem solving and scaffolding in studying. the experienced learning agency was related to social support experienced in several ways. also differences in learning agency and social support in terms of grade level, gender and ses were detected. results indicate that meaning making especially calls for intentional support from teachers in lower secondary grades and that girls and boys have partly different support needs in terms of cultivating strong sense of learning agency. keywords: learning agency; primary school students; lower secondary school students; peer support; explorative structural equation modeling mailto:tiina.soini-ikonen@tuni.fi https://10.0.57.194/flr.v11i2.983 soini | f l r 50 1. introduction there is an evident need for cultivating learning that enables us to solve complex ill-defined problems, create new knowledge, innovative and actively learn through life. how to enhance such learning has been a central interest of the researchers, educational practitioners, and educational policy makers across the globe for quite some time (griffin, mcgraw & care, 2012; ananiadou & claro, 2009; darling-hammond, flook, cookharvey, barron, & osher, 2019; häkkinen et al, 2016; ministry of education, 2012). for instance, many educational policy documents list the goals of effective learning for future, such as 21st century learning and innovation skills (oecd, 2013; pellegrino, 2017). both educational policy documents and learning science emphasize continuous skillful and active learning (schober et al., 2013; waeytens et al., 2002; vainikainen & hautamäki, 2020). this is reflected in commitment to continuous education in many national and organizational strategies for future (marsick & watkins, 1992; maurer & weiss, 2010). there is also a degree of consensus across the policies, and theorisations that learning allowing knowledge creation and steering one’s own life is more than adoption of existing knowledge and skills. such higher order learning is characterized by active, skillful, and creative problem solving, thinking skills, knowledge creation and recognizing others as resources for learning (e.g., anderson et al. 2001; dwyer et al., 2014; greiff et al., 2013; roseth et al., 2008; vainikainen et al., 2015; see also seminal work by bandura, 1989 and bloom et al., 1956). there is ample amount of evidence that engaging in such learning activities is related to range of positive study attributes, such as intrinsic motivation (vansteenkiste et al., 2009), positive study-related emotions (saariaho et al., 2016; pekrun et al., 2002), and lower risk of suffering study burnout (näykki et al., 2018). the characteristics and scope of such learning have been well documented in the literature, especially research on selfand co-regulated learning has brought understanding of strategies and skills needed in this kind of higher order learning (hadwin, järvelä, & miller, 2011; järvenoja, volet, & järvelä 2013; pintrich, 2004). there is also a growing understanding of complex set of elements such as emotions and motivation that are intertwined with learning regulation (e.g., anttila, pyhältö, pietarinen & soini, 2018; mega, ronconi, & de beni, 2014; schunk & zimmerman, 2008; valle et al., 2003). unfortunately, readiness to engage in higher order learning cannot be taken for granted. it calls for not only cultivating learning skills but also learner’s initiative and intentional effort regards to their learning i.e., learning agency (cf. resnick, 1995). cultivating learning agency of students can be considered as a focal task of school in making sure that all children have the opportunity to build learning agency. hence, it is a focal task for school and preferably should start as early as possible (evans & rosenbaum, 2008). however, the range of studies exploring higher order learning beyond learning skills or single attribute of learning, is still limited. our study tackles the challenge by exploring how primary and lower secondary school students perceive their learning agency in terms of problem solving, meaning making and seeking scaffolding for learning in school and how it is related to the teacher and peer support experienced while studying. differences in the sense of learning agency between girls and boys, and schools located in low and high ses neighborhoods will be also examined. in addition, the article will contribute to the arsenal of methods by providing a novel measure for exploring students’ learning agency. 2. theoretical framework 2.1 student’s learning agency having agency in ones learning entails choosing when, what, and where one is learning, and having adequate skills to do it, independently and in collaboration with other (clarke, howley, resnick, & rose, 2016; scardamalia & bereiter, 2006). agency refers to learner’s initiative and ownership of learning that for example self-regulation of learning assumes (pintrich, 2004). such learning cannot be reduced to a single attribute of the learner or to a particular skill. this is also why it may be challenging to recognize and support it in school soini | f l r 51 environment. a student may have adequate learning skills, even skills of self-regulation, but if they lack the motivation to learn, nothing will happen. further skills and the motivation to learn do not guarantee student’s learning agency if the student does not believe in their abilities to learn. accordingly, learning agency calls for will and power to engage and steer learning and believe that one can do it. by drawing on two influences on agency: social structures and self-appraisals, we claim that learning agency needs to be conceptualized both as a function of social structures of learning environment and capacity of the learner (biesta & tedder, 2007). this does not mean that individual students will or can always control their learning in different contexts and social structures which form a framework of learning agency, however, structures are continuously changed and reshaped by individuals (giddens, 1984; pintrich, 2004). the individual self-appraisals include the learner’s will and intention to act (bandura, 1989), and belief in their ability to influence things and the actual power they have through skills and strategies of action (clarke et al., 2016). based on these conceptualizations we argue that learning agency consists of the interrelated elements of student’s motivation to learn continuously (i want), efficacy beliefs about their learning (i am able), and intentional strategies for facilitating and managing new learning (i can and do) (pyhältö, pietarinen & soini, 2016; soini, pietarinen, toom, & pyhältö, 2015; soini, pietarinen, & pyhältö, 2016). motivational, self-efficacy and skill elements of learning allow learner’s initiative and ownership in terms of learning. as such, these cognitive elements (i.e., motivation, efficacy beliefs and intentional strategies for learning) are just decontextualized preconditions for learning, however, to manifest learning agency they have to be situated and applied in learning (darling-hammond et al., 2019; lave & wenger, 1991). it has been proposed that particularly problem solving (hmelo-silver, 2004), meaning making (coburn, 2001; weick, sutcliffe, & obstfeld, 2005) and seeking scaffolding for learning (cheng et al., 2015; lave & wenger, 1991) are increasingly needed in the future and therefore have high applicability value for students later in their lives. therefore, cultivating the motivation, efficacy and skills in school learning – not just in some subjects but as an overall capability is essential. problem solving, which includes where the identification and problem finding is needed (dillon, 1982) and solution is not immediately obvious (scherer & gustafsson, 2015) allows identifying and exploiting opportunities and taking control over the future. in school learning it refers to learners’ will, and their active efforts to try to process problems encountered in learning and persistently aim to solve them (pintrich, 2004; valle, 2003). meaning making refers to interpreting and making sense of situations, as well as recognizing emotions and identity process related to them (e.g., zittoun & brinkmann, 2012) and maintaining wellbeing (e.g., park, 2011). in school learning this involves learners attempt to understand what the learning task entails, why it is important to master, and what the relevant things are related to it. seeking scaffolding includes activities through which learner actively seeks help and builds supporting social elements to facilitate learning. scaffolding helps to perform tasks beyond current personal resources through the social resource for learning (e.g., hadwin, järvelä, & miller, 2011). in school this means participating in building school’s social learning environment. engagement in such learning throughout school career and cross-curriculum allows an active approach to learning and taking agency over learning at any level of education (kagan, 2005; resnick, 1987). we argue that these are crucial for learning agency in school, and to be able to adopt learning as life orientation. to sum up this implies (see h1) that student’s sense of learning agency is embodied by problem solving, meaning making and seeking scaffolding for learning (darlinghammond et al., 2019; hmelo-silver, 2004; timperley & parr, 2005, weick, sutcliffe, & obstfeld, 2005; cheng et al., 2015; lave & wenger, 1991). each of these comprises motivation, efficacy beliefs, and intentional strategies for facilitating and managing new learning (soini et al., 2016; pietarinen, pyhältö, & soini, 2016; pyhältö et al., 2015; scherer & gustafsson, 2015; valle, 2003). learning agency is a capacity of an individual embedded in and constrained and enabled by individualenvironment dynamics (clarke et al., 2016). hence, the degree of learning agency can vary due to both soini | f l r 52 individual and contextual affordances and constrains. deficits in learning agency may lead to challenges in adopting skills and strategies of learning and put students in unequal positions in terms of learning. for example, research on learning regulation has identified differences between boys and girls. in general, the results have shown that girls tend to outshine boys in selfand co-regulative learning skills in primary and lower secondary school (weis, heikamp, & trommsdorff, 2013; raffaelli, crockett, & shen 2005; duckworth & seligman, 2006; fischer, schult, & hell 2013). in other studies, no gendered differences have been detected in self-regulated learning between boys and girls (wanless et al., 2013). accordingly, further studies are needed to explore whether there are gendered differences in learning agency, potentially contributing to differences in learning outcomes. moreover, previous studies have shown that socio-economic background contributes to student learning (lindfors, minkkinen, rimpelä, & hotulainen, 2018; myllyniemi & kiilakoski, 2018). students who study in schools located in low ses areas are less familiar with engaging in higher order learning activities compared with students in schools located in high ses neighborhoods (størksen et al., 2015; xing, liu, & wang, 2019). a reason for this might be that it is more challenging to sustain a socially supportive learning environment enabling sense of learning agency for students in low ses neighborhoods schools. for instance, it calls for more investment promoting student well-being, taking time to talk about interpersonal problems or showing patience with students’ misbehavior leaving fewer resources to invest in engaging students in higher order learning compared to schools in high ses neighborhoods (bottiani, duran, pas, & bradshaw, 2019). so far, in finland, differences in students’ learning outcomes between the schools located in different neighborhoods have been modest. yet, results of a few studies imply that differences might be increasing (berisha & seppänen, 2017; bernelius, 2013; bernelius & vaattovaara, 2016). we presume that first signals of such development might be first detected in students’ sense of learning agency before they are realized in differences in learning outcomes. based on this, we hypothesise that students’ perceived learning agency is related to the gender and schools’ neighborhood ses and primary school students and girls report a higher sense of learning agency in terms of seeking scaffolding, meaning making, and problem solving than boys and students at lower-secondary school (duckworth & seligman, 2006; fischer, schult, & hell 2013; liu et al., 2016; raffaelli, crockett, & shen 2005; salmela-aro & tynkkynen, 2012; wang & eccles, 2012) (see h3). this further implies that girls (ciarrochi et al., 2017; hombrados-mendieta et al., 2012), primary school students, and students from high ses living area (størksen et al., 2015; xing, liu, & wang, 2019) are likely to report higher levels of social support from teachers and peers compared to boys, lower secondary students or students from low ses living areas (weis, heikamp, & trommsdorff, 2013; raffaelli, crockett, & shen, 2005; duckworth & seligman, 2006; fischer, schult, & hell, 2013; camara et al., 2017; rautanen, soini, pietarinen, & pyhältö, 2022; wentzel et al., 2017) (see h4). 2.2 teacher and peer support for studying – school’s social structure for enhancing learning agency student learning agency is relational and hence highly socially embedded (edwards, 2005; clarke et al., 2016). accordingly, the quality and the quantity of social support for studying is likely to have an impact on the degree of learning agency displayed by the students, and its development. social support refers to the social resources that are perceived as being available (cohen et al., 2000) for studying. in school, such resources typically include (but are not limited to) teachers and peers (kiefer et al., 2015). teachers and peers provide the primary source of both informational support, including feedback, advice, affirmation and problem solving that enables students to cope with study and learning related challenges (e.g., malecki & demaray, 2003; liu et al., 2016; wentzel et al., 2016), and the emotional support comprising care, trust, encouragement, acknowledgement and sense of belonging to the school community (e.g., malecki & demaray, 2003; liu et al., 2016; wentzel et al., 2016). there is a strong body of evidence showing that receiving informational support from teachers enhances students subject matter interest, study engagement, school achievement (ahmed et al., 2010; wentzel, 1998), and enables them to overcome study-related challenges (e.g., liu et al., 2016; wentzel et al., 2017). in turn, emotional support from teachers has been shown to promote students’ soini | f l r 53 positive attitudes to schoolwork and school satisfaction (e.g., liu et al., 2016; ulmanen et al., 2016; wang & eccles, 2012; wentzel et al., 2017; jiang, huebner, & siddall, 2013), as well as to protect students against study burnout, anxiety and stress (salmela-aro et al., 2008 roy, kristensen, groholt, & clench‐aas, 2009; tennant et al., 2015). students who receive emotional support from their teachers are also less likely to engage in risky behavior and more likely to show high levels of academic achievement (flashpohler et al., 2009; gini et al., 2008). results on the impact of peer support are less consistent (e.g., kiefer et al., 2015; kindermann, 2007; skinner et al., 2008; wang & eccles, 2012). while in some studies informational and emotional support from peers has been found to enhance study engagement, school adaptation, perceiving school work valuable, and positive emotions about studying such as joy of learning (ulmanen, soini, pietarinen, & pyhältö, 2016; wang & eccles, 2012 furrer & skinner, 2003; estell & perdue, 2013; kiefer et al., 2015; urdan & schoenfelder, 2006; wentzel et al., 2017; ulmanen, soini, pietarinen & pyhältö, 2014), in other studies a negative relationship between the peer support and study engagement have been detected (liu et al., 2016). it has been suggested that this results from the mediating role of peer group attitudes about studying (wang & eccles, 2012). if peers share negative attitudes about studying, the peer support is likely to increase disruptive behavior and disengagement from studying (sage & kindermann, 1999; ulmanen et al., 2014; ryan, pintrich & midgley, 2001; wang & eccles, 2012), while sharing positive attitudes about studying is likely to increase study engagement. we presume that receiving social support for studying from peers that require the displaying of a positive attitude to studying is likely to be related to increased levels of learning agency among the students. there is also research evidence that social support provided by teachers and peers complement each other (see also rautanen, soini, pietarinen & pyhältö, 2022). prior research has also detected differences in experiences of social support in terms of gender, age, and background of the student. girls have been shown to report more teacher support in their schoolwork and they also experience more peer support in their schoolwork (lam et al., 2012; wentzel et al., 2017), particularly in lower grades (liu et al., 2016). moreover, students from high ses living areas have been shown to report higher levels of social support from both teachers and peers (størksen et al., 2015; xing, liu, & wang, 2019). against this backdrop, we claim that informational and emotional support from teachers and peers is likely to cultivate student’s learning agency in terms of problem solving, meaning making and seeking scaffolding and differences in experienced support influence on learning agency. accordingly, we hypothesise that the sources of support that complement each other, i.e., informational and emotional support for studying from teachers and peers, are positively related to the students’ perceived learning agency in terms of seeking scaffolding, meaning making and problem solving in both grades (see h2) (e.g., liu et al., 2016; rautanen et al., 2022; wenzel et al., 2017) (see h2). 3. aim, research questions and hypothesis the aim of this study was to understand the composition of perceived learning agency (la) among primary and lower-secondary school students, and how it is related to their sense of social support for studying from teachers and peers. the following research questions and hypothesis derived from those are addressed (see also figure 1): rq 1: is the perceived learning agency among students comprising of intertwined composition of motivation, efficacy beliefs, and intentional strategies of learning embodied by problem solving, meaning making and seeking scaffolding for learning? hypothesis 1: a primary and lower-secondary school student’s sense of learning agency is embodied by problem solving, meaning making and seeking scaffolding for learning. soini | f l r 54 rq 2: how are the sources of social support related to the student’s perceived learning agency in terms of seeking scaffolding, meaning making and problem solving in both grades? hypothesis 2: the sources of support that complement each other, i.e., informational and emotional support for studying from teachers and peers, are positively related to the student’s perceived learning agency in terms of seeking scaffolding, meaning making and problem solving in both grades. rq 3: is the students’ perceived learning agency related to the gender, grade, and school’s neighborhood ses? hypothesis 3: the students’ perceived learning agency is related to the gender and school’s neighborhood ses and primary school students and girls report a higher sense of learning agency in terms of seeking scaffolding, meaning making, and problem solving than boys and students at lower-secondary school. rq 4: is the students’ perceived social support related to the gender, grade, and school’s neighborhood ses? hypothesis 4: girls, primary school students, and students from high ses living area are likely to report higher levels of social support from teachers and peers compared to boys, lower secondary students or students from low ses living areas. soini | f l r 55 figure 1. hypothesized esem path model between the student’s leaning agency [i.e., meaning making (mm), seeking scaffolding (sc), and problem solving (ps)], school-related social support from teachers and among peers, gender, and school’s neighbourhood ses. 4. method 4.1 participants and the data collection mm sc ps imm1 imm2 imm3 imm4 imm5 imm6 imm7 imm8 isc1 isc2 isc3 isc4 isc5 ips1 ips2 ips3 ips4 ips5 ips6 gender ses peer support p2 p3 p4 p5 p6 p7 p8 p9 p1 p10 teacher support t3 t4 t5 t6 t7 t8 t9 t10 t2 t11 t1 soini | f l r 56 in finland, comprehensive school includes primary school (grades 1–6) and secondary school (grades 7–9) and nearly all children complete the 9-year long compulsory basic education. the data were collected from two cohorts of students: 4th graders (n=2401, 50% girls, 58% high ses, age 10) and 7th graders (n=1545, 51% girls, 59% high ses, age 13) from 71 comprehensive schools from 245 school classes around finland. the schools selected in the study represented the demographic variation of the schools in finland: they were situated throughout the country and varied in terms of size, location (rural/urban) and the ses level (high/low) (referring to the levels of income, employment, and education of the habitants of the surrounding area of the school). the student population of the schools ranged from 50 to 1255 students. the researchers collected the data during school days in fall 2017. they introduced the study to the students, gave them instructions on how to complete the survey and collected the written surveys from the students. researchers received consent for study participation from the chief officers of the school district, the case schools, and the students and their parents. participation in the study was voluntary and no extra credits were given to the students for participation. furthermore, students were informed that neither their parents nor their teachers would see their answers. the total response rate of the survey was 90% in the primary school cohort and 89% in the secondary school cohort. according to the guidelines of finnish national board on research integrity no ethical review for the study was required (finnish national board on research integrity (tenk)). 4.2 learning agency -survey learning agency (la) was assessed by using new measurement developed by research group. it includes three subscales that measure student’s sense of learning agency in the school context: seeking scaffolding (5 items), meaning making (8 items), and problem solving (6 items) each of them involving motivation to learn, efficacy beliefs and intentional strategies for learning. the subscales were tested and further developed based on threephase piloting procedure. first, six students from grades 4 to 8 were asked to answer the pilot survey and freely comment on the items. several ambiguous items were removed, and conceptual clarifications were made to the items based on discussions with the students. second, the subscales were statistically studied for a sample of 93 fourth graders. the three subscales of learning agency were identified with principal axis factoring (paf) and exploratory factor analysis (efa) using multiple rotation options. the consistency and stability of the cronbach’s alpha reliabilities was also studied in different pfa and efa solutions. finally, the final three factor structure was also validated with the explorative structural equation modeling (esem) with the partial samples of the national school data gathered, including students from 4th, 5th 6th, 7th 8th, and 9th grades. due to this three-phase piloting procedure, the number of originally developed items was compressed to the half within each subscale during the piloting. the subscales were developed to assesses students’ motivation to learn (i want), efficacy beliefs, (i am able), and intentional strategies (i can and do) occurred in the three higher order learning activities. the seeking scaffolding scale measures students’ motivation, efficacy beliefs, and strategies to seek actively help for learning when needed (α = .79 in grade 4 and α = .81 in grade 7). the meaning making scale measures students’ motivation, efficacy beliefs, and strategies to construct, understand and make sense of learning intentionally (α = .86 in grade 4 and α = .86 in grade 7), while, the problem solving scale measures students’ motivation, efficacy beliefs, and strategies to find alternative solutions to problems encountered in learning (α = .82 in grade 4 and α = .86 in grade 7). each scale contains at least one item from each element of learning agency (motivation, efficacy, strategy). the scales were measured using a 7-point likert scale (1 = totally disagree, 7 = totally agree). they are presented in the appendix in table a1. school-related social support was examined by assessing social support from teachers and guardians and among peers (rautanen, soini, pietarinen & pyhältö, 2021). the teacher support scale assessed emotional support (i.e., respect, empathy, and care) and informational support from teachers for studying (α = .94 in grade 4 and α = .95 in grade 7; 11 items, e.g. “my teachers give me encouragement and support”, “i often receive constructive feedback from teachers”, “the teachers are interested in my opinions”). the peer soini | f l r 57 support scale measured emotional and informational support for studying. it comprised items concerning both giving and receiving social support for studying (α = .92 in grade 4 and α = .93 in grade 7; 10 items, e.g., “i have the courage to offer my friends help with their studies”, “i have the courage to ask others for help with my studies”, “my classmates' encouragement inspires me in my studies”). students rated peer and teacher support using a 7-point likert-type scale ranging from 1 (totally disagree) to 7 (totally agree). moreover, gender and the socio-economic characteristics of the living area surrounding each school (and student), i.e., low or high ses living area, were used as background variables. 4.3 analysis we used spss for the preliminary analyses of the data and mplus statistical package (version 8.0, muthén & muthén, 1998–2017) for the further analyses. the amount of missing data in the observed variables typically varied from between 2% and 7%, except for 15–16% for two items. little’s test of missing completely at random (mcar) was significant showing that the data were not missing completely at random (x2 = 26840.708, df = 23967, p=.000). missing data were accounted for in further analysis for using full information maximum likelihood (fiml) to estimate models without imputing the missing data. further, the model parameters were estimated by means of maximum likelihood robust (mlr) estimation, which is robust to the nonnormality of the observed variables (mlr estimator; muthén & muthén, 1998–2017). in the first phase of analyses, we examined the structure of la by exploratory factor analysis (efa) using multiple different rotations. cronbach’s alpha reliabilities were also studied. efa supported the three-factor model that consisted of meaning making, seeking scaffolding, and problem solving factors. all correlations between items were statistically significant and in expected directions (see appendix, table a2). in the second phase of analyses, we validated the factor structure by using explorative structural equation modeling (esem). in esem, items loaded on their main factors, and cross-loadings were “targeted”, but not forced, to be as close to zero as possible with the oblique target rotation procedure (tóth-király, bõthe, rigó, & orosz, 2017). the following guidelines were applied to interpret the magnitude of the factor loadings defined as excellent above 0.71, very good between 0.63 and 0.70, good between 0.55 and 0.62, fair between 0.44 and 0.33, and poor below 0.30 (comrey & lee, 2013). due to the nested data (i.e., schools, classes, and students) the intra-class correlation coefficients (icc) and design effects (deff) were examined using classes clustering variables (see e.g., snijders and bosker, 2012). in deff, values over 2 and in iccs, values over .05 indicate a clustering effect in the data (muthén & satorra, 1995; peugh, 2010). the analysis showed that the data was moderately clustered (i.e., iccs varied between .02–.10 in the whole data and in gender groups and grade levels) and this was considered in the analysis with a complex option if possible (muthén & satorra, 1995; peugh, 2010) 1. after identifying/confirming the final factor structure of students’ learning agency, we tested hypothesized path model between the latent variables of teacher support, peer support, and student’s learning agency (i.e., seeking scaffolding, meaning making and problem solving) (see figure 1). gender and ses were included as observed dichotomous variables in the model. the model fit was evaluated using the following criteria for the goodness-of-fit indices: the root mean square error of approximation (rmsea) below 0.06, the comparative fit index (cfi) and the tucker-lewis index (tli) over 0.90, the standardized root mean squared residual (srmr) below 0.05, and the normed fit index (nfi) over 0.90 (byrne, 2012; hooper et al., 2008; hu & bentler, 1999). 1 complex option was used in the analysis of the data related to fourth graders. in the data related to seventh graders, the size of the data limited the use of the complex option when the number of free parameters exceeded the number of clusters. thus, clustering effect could not be considered in the analysis of the data related to seventh graders. however, comparing the results obtained with and without complex option suggested that there were no significant differences in the result obtained in different ways. soini | f l r 58 5. results 5.1. the student’s sense of learning agency the esem analysis for the correlated three-factor model of learning agency resulted good fit with the data after modifications of two added residual covariances (see table 1.). the hypothesized esem three-factor model worked adequately in both student cohorts (i.e., factor loadings above .30) (see table 1). one item (sc3) cross loaded moderately between the seeking scaffolding and meaning making factors among younger student cohort (i.e., fourth graders). however, item (sc3) had an expected factor load to the meaning making factor in the cohort of seventh graders. table 1 standardized parameter estimates and goodness-of-fit summary for the esem solutions of the learning agency scale separately within 4th graders and 7th graders. 4th graders (n=2392) 7th graders (n=1546) items sc (l) mm (l) ps (l) res.var sc (l) mm (l) ps (l) res.var sc1 .55 .12 .00 .62 .40 .33 -.12 .68 sc2 .74 .02 .01 .44 .69 .08 -.02 .47 sc3 .32 .30 -.02 .71 .45 .23 -.11 .69 sc4 .88 -.08 -.05 .32 .88 -.13 .05 .30 sc5 .79 -.06 .08 .37 .87 -.16 .10 .31 mm1 .00 .61 .03 .60 .02 .57 .10 .57 mm2 .14 .57 -.02 .59 .06 .61 .04 .55 mm3 -.01 .49 .14 .62 .00 .71 .02 .48 mm4 .05 .72 -.09 .54 .14 .46 .05 .65 mm5 -.07 .61 .07 .60 .09 .33 .17 .72 mm6 .09 .54 .09 .56 .12 .53 .10 .53 mm7 .02 .57 .15 .51 -.09 .66 .15 .46 mm8 .02 .69 -.00 .52 .01 .69 .01 .51 ps1 .08 .03 .58 .59 .06 .17 .54 .51 ps2 .00 -.10 .76 .53 -.10 .01 .80 .42 ps3 -.10 .29 .41 .63 -.03 .13 .58 .56 ps4 .00 .14 .57 .53 -.00 .18 .59 .46 ps5 -.03 -.08 .82 .45 -.01 -.17 .87 .44 ps6 .09 -.004 .64 .536 .19 .08 .52 .52 factor correlations 4th grade 7th grade sc mm .53 .54 sc ps .46 .43 ps mm .78 .74 grade n χ² df p cfi tli rmsea 90% ci srmr 4 2392 580.56 117 .000 .96 .94 .04 .04-.04 .02 7 1540 563.20 116 .000 .95 .93 .05 .05-.05 .03 soini | f l r 59 note. esem=exploratory structural equation modeling; sc=seeking scaffolding; mm=meaning making; ps=problem solving; res.var.=residual variance; l=factor loadings, target loadings are in bold; nonsignificant parameters (p≥.05) are italicized; χ2=chi-squared test; df=degrees of freedom; cfi = comparative fit index; tli = tucker–lewis index; rmsea = root mean square error of approximation; 90% ci = 90% confidence interval of the rmsea; srmr = standardized root mean square residual. all correlations between the three elements of the learning agency were statistically significant and in expected directions (see table 1). problem solving and meaning making correlated stronger with each other than with seeking scaffolding in both student cohorts (i.e., 4th and 7th graders). this indicated that the students’ intentions to find alternative solutions to problems encountered in learning is stronger related to the need for understanding and making sense of learning than seeking actively help for learning. the results confirmed the first hypothesis (h1) by showing that students’ sense of learning agency, i.e., capacity to engage in higher order learning activities such as complex problem solving, meaning making and seeking scaffolding for learning at the school is relational, and cannot be reduced to a single attribute of the learner. in other words, students’ motivation, efficacy beliefs, and strategies to construct, understand and make sense of learning intentionally (i.e., meaning making), to find alternative solutions to problems encountered in learning (i.e., problem solving) and seek actively help for learning (i.e., seeking scaffolding) are interrelated but separate constructs, in fourth and seventh graders’ experiences. 5.2 the interrelationship between the sense of learning agency and perceived social support the results showed statistically significant grade, gender, and ses differences between students in sense of learning agency and received social support from teachers and peers (see table 2). due to the large sample size the effect sizes of differences were also studied using cohen’s d besides the p-values from the t-tests (lakens, 2013). primary school students (i.e., 4th graders) reported a statistically significantly higher sense of learning agency than lower-secondary school students (i.e., 7th graders) in terms of enhancing their learning by meaning making, seeking scaffolding and problem solving (cohen’s d ranging from .41 to .57). girls reported higher sense of seeking scaffolding and meaning making than boys, while boys reported slightly higher problem solving than girls. however, as table 2 shows, the gender differences were not as strong as the grade differences (cohen’s d ranging from .08 to .28). students from high ses areas perceived all three elements of learning agency to be slightly higher than students from low ses areas (cohen’s d ranging from .08 to .12). there were also statistically significant grade, gender, and ses differences between students in terms of the perceived social support from teachers and peers (see table 2). primary school students reported statistically significantly higher teacher and peer support than lower-secondary school students (cohen’s d ranging from .41 to .46). girls reported statistically significantly higher teacher support and peer support than boys. however, the gender difference in teacher support was not as strong as in peer support. in turn, the observed statistically significant differences between students from different ses-areas were minor (see cohen’s d ranging from .03 to .09). the analysis further showed that the observed correlations between the explored scales were statistically significant and in the expected directions in both student cohorts (see table 2). the findings confirmed mainly the fourth hypothesis (h4) that girls and students from high ses living areas perceive more social support from teachers and peers. they are also more familiar with these learning activities that contribute to sense of learning agency compared to boys or students from low ses living areas. soini | f l r 60 table 2 means, standard deviations, cohens’ d and observed correlations grade differences 4 th graders (n=2396) 7 th graders (n=1546) m sd m sd cohen’s d seeking scaffolding 5.80 1.10 5.27 1.19 .46 meaning making 4.79 1.19 4.31 1.16 .41 problem solving 5.00 1.14 4.34 1.18 .57 teacher support 5.37 1.22 4.80 1.26 .46 peer support 5.60 1.12 5.13 1.16 .41 gender differences girls (n=1961) boys (n=1932) m sd m sd cohen’s d seeking scaffolding 5.76 1.11 5.44 1.19 .28 meaning making 4.76 1.17 4.46 1.20 .25 problem solving 4.69 1.22 4.79 1.18 .08 teacher support 5.24 1.24 5.05 1.29 .15 peer support 5.66 1.04 5.17 1.22 .43 ses differences low ses (n=1635) high ses (n=2309) m sd m sd cohen’s d seeking scaffolding 5.54 1.20 5.63 1.14 .08 meaning making 4.54 1.24 4.65 1.16 .09 problem solving 4.66 1.22 4.80 1.18 .12 teacher support 5.08 1.28 5.19 1.25 .09 peer support 5.39a 1.19 5.43a 1.14 .03 correlations 1 2 3 4 5 1. seeking scaffolding .54 .41 .55 .53 2. meaning making .54 .70 .50 .56 3. problem solving .44 .69 .44 .48 4. teacher support .55 .50 .48 .59 5. peer support .57 .59 .58 .67 note 1. means within a row sharing the same subscripts are not significantly different at the alpha=.05 level in pairwise t-test. the scales range from 1 to 7. note 2. all observed mean correlations were statistically significant at the p<.01 level. fourth graders are below the diagonal and seventh graders are above diagonal. the tested theoretical esem path model, where experienced social support from teachers and peers were positively correlated with the student’s sense of learning agency, is presented in figure 2. in addition, gender and neighborhood ses have been included as dichotomous covariates in the tested model. according to the several fit indicators, the tested path model fit the data well in both student cohorts (4th graders: n=2363; χ²=3111.91, df=766, p=.000; cfi=.93; tli=.92; rmsea=.04; 90% ci=.04-.04; srmr=.04; 7th graders: n=1528; χ²=2985.06, df=764, p=.000; cfi=.92; tli=.91; rmsea=.04; 90% ci=.04-.05; srmr=.04). the statistically significant p value of the chi-square test is due to the large sample size used in this study. the model supported hypothesis 2 by showing that experienced teacher support and peer support are positively associated to each other (𝛽 = .71/.64) and with meaning making (mm), seeking scaffolding (sc) and problem solving (ps), and the associations differ depending on the source of the support and type of experienced higher order learning that contributes to sense of learning agency (see figure 2). more specifically, the results showed that the perceived peer support had a stronger positive correlation with each soini | f l r 61 of the experienced learning contributing to students’ sense of learning agency compared to perceived teacher support in both student cohorts. the peer support had the strongest positive correlation with problem solving (𝛽 = .59/.51) and the weakest with seeking scaffolding (𝛽 = .40/.38) in both student cohorts (4th graders/7th graders). instead, teacher support had the strongest positive correlation with seeking scaffolding (𝛽 = .29/.32), and the weakest with problem solving (𝛽 = .14/.12) in both student cohorts. however, the perceived teacher support seemed to be especially crucial for lower secondary school students’ meaning making (𝛽 = .17/.33), i.e., experienced motivation, efficacy beliefs, and strategies to construct, understand and make sense of one’s own learning intentionally. soini | f l r 62 figure 2. esem path model between student’s learning agency, [i.e., meaning making (mm), seeking scaffolding (sc), and problem solving (ps)], school-related social support (i.e., teacher and peer support among fourth and seventh graders (note: see table a3 in appendix the detailed results of esem mm r 2 =.47/.46 sc r 2 =.42/.40 ps r 2 =.46/.37 mm1 mm2 mm3 mm4 mm5 mm6 mm7 mm8 sc1 sc2 sc3 sc4 sc5 ps1 ps2 ps3 ps4 ps5 ps6 gender ses .17/.33 .55/.40 .29/.32 .40/.38 .14/12 .59/.51 .15/.28 .04/.08 -.23./-.23 .71/.64 .10/.13 .23/.27 .64/.62 peer support r 2 =.05/.05 p2 p3 p4 p5 p6 p7 p8 p9 p1 p10 teacher support r 2 =.03/.00 t3 t4 t5 t6 t7 t8 t9 t10 t2 t11 t1 ns./.09 ns./.06 -.15/ns. .79/.84 .70/.78 .70/.70 .79/.81 .66/.62 .72/.72 .79/.83 .75/.83 .80/.84 .70/.76 .84/.88 .80/.78 .75/.72 .65/.65 .65/.68 .70/.78 .72/.75 .74/.74 .82/.83 .62/.72 .73/.76 .44/.53 -/.34 .34/.34 -/.45 .07/n s. ns./-.06 soini | f l r 63 corresponding the factor loads of la; standardized model, all parameters are statistically significant at p < 0.05). the results showed that fourth (𝛽 = -.23) and seventh (𝛽 = -.23) grade girls reported higher levels of peer support compared to boys. fourth grade girls (𝛽 = -.15) also reported higher levels of teacher support than boys. further, girls were less familiar with problem solving as an efficient learning than boys in both student cohorts (𝛽 = .15/.28) and, in turn, seventh (𝛽 = -.06) grade girls were more familiar with the perceived meaning making as a way of learning. however, the living area surrounding each school, i.e., low or high ses area, did not correlate with experienced social support and experienced learning activities as much as gender. as an indicative result, fourth graders from a high ses area were more familiar with the perceived problem solving as learning (𝛽 = .04) and teacher support (𝛽 = .07) than students from a low ses area. instead, seventh graders from high ses areas perceived meaning making (𝛽 = .09), seeking scaffolding (𝛽 = .06) and problem solving (𝛽 = .08) activities more familiarly in their learning than students from low ses area. the results also showed that especially experienced school-related teacher and peer support, and also gender and the school’s neighborhood ses were positively correlated with fourth and seventh graders’ motivation, efficacy beliefs, and strategies to make sense of own learning intentionally (i.e., mm, r²= .47/.46), to find alternative solutions to problems encountered in learning (i.e., ps, r²= .46/.37) and seek actively help for learning (i.e., sc, r²= .42/.40). all in all, the findings confirmed partly our hypothesis (h3) that girls and primary school students report a higher sense of learning agency. more precisely, girls reported higher sense of seeking scaffolding and meaning making than boys, while boys reported slightly higher problem solving than girls. in turn, the results confirmed mainly our hypothesis (h4) that girls, primary school students, and students from high ses living areas perceive more social support from teachers and peers. that is, primary school students and girls reported higher teacher support and peer support than boys. however, the gender difference in teacher support was not as strong as in peer support and differences between students from different ses-areas were minor. 6. discussion and conclusion 6.1 theoretical and practical implications we explored fourth and seventh graders’ learning agency realized in the motivation, self-efficacy, and skill they experienced in engaging in problem solving, meaning making and seeking scaffolding for learning. also, interrelations between experienced teacher and peers support, and the sense of learning agency was examined. the results confirmed that the learning agency is a multifaceted construct and that these elements of learning agency i.e., will, ability beliefs and skills in terms of learning, are ingrained in solving problems for learning, constructing meaning by bridging prior understanding to current learning tasks, and identifying relations in learned contents and actively seeking help and social scaffolds for learning. the results imply that to understand higher order learning and support it at school, all the elements of learning agency should be considered. overall, students reported high levels of active help seeking for learning, implying that they actively used the social structure of the learning environment as a resource for learning. however, at the same time, they reported lower levels of problem solving for learning and meaning making, indicating that motives, efficacy beliefs and skills related to these higher order learning activities could be further enhanced. solving problems and making meaning in learning in school also correlated with each other more strongly than with seeking scaffolding for learning. it seems that attempts to make sense of the goal of learning is related to trying persistently and in various ways to solve the problems encountered in learning. this implies that encouraging and supporting students when they ponder on learning tasks and even question their relevance will increase their resilience with the task and is hence a good investment in supporting learning. soini | f l r 64 the study detected differences in learning agency in terms of grade level, gender, and ses. fourth graders reported higher levels of learning agency than seventh graders, which may be explained by students’ decreasing engagement in studying as the school path proceeds (janosz, archambault, morizot, & pagani, 2008; li & lerner, 2011). at the same time, teacher support was more strongly associated with learning agency in lower secondary school. this may imply lower secondary school students’ motivation, efficacy beliefs, and strategies to construct, understand and make sense of own learning intentionally (i.e., meaning making) calls for more intentional support from teachers in lower secondary grades. it could be also argued that supporting higher order learning in adolescence may require more effort – or even new kinds of pedagogical practices that enable students to reconstruct their relationship with learning in school and their role as learners. in term of gender differences, finnish girls usually appear to be better adjusted to school and perform better than boys. however, our findings on the gender differences were less straight forward. while girls reported more active help seeking and attempted to make sense of the learning, boys reported more will, efficacy beliefs and skills of solving problems for learning. accordingly, both girls and boys showed a rather strong sense of learning agency, yet it was manifested differently across the learning activities. this implies that not only learning outcomes, but also experiences of higher order learning is partly gendered in school. however, the differences detected were more complex and nuanced that those previously reported regarding school adjustment (liu et al., 2016; wang & eccles, 2012) and school achievement (bursal, 2017; oecd, 2019,). results imply that girls and boys may have partly different support needs in terms of cultivating a strong sense of learning agency. while girls might benefit from encouragement in problem solving, boys may profit from support in meaning making and seeking scaffolding for learning. the results also showed that a school’s neighborhood ses was modestly related to fourth and seventh graders’ sense of learning agency i.e., motivation, efficacy beliefs, and strategies to make sense of one’s own learning intentionally, to find alternative solutions to problems encountered in learning and actively seeking help for learning. a potential reason for the result might be that the students from high ses areas have more experience with learning activities characterized by the learning agency compared to students from low ses areas. this could also be an early sign of increasing differences in learning outcomes starting with the erosion of students’ sense of learning agency, at its worst resulting in inequality in learning opportunities in finland, if the school cannot act as a buffer against risk factors such as low ses. teacher and peer support were both related to experiences of learning agency in general. however, there were both school grade and gender differences in the support experienced; fourth graders and girls reported higher levels of social support from both peers and teachers than seventh graders and boys. the differences were more significant in terms of support from peers. this result is in line with previous findings that girls experience more peer support, especially from their classmates, than boys (camara et al., 2017; rautanen et al., 2020; wentzel et al., 2017). it seems that learning in school for girls is more strongly embedded in peer interaction compared to boys. girls seem to seek and offer help actively and experience having it. further investigation showed that the peer support was more strongly related to experiencing learning agency than support from teachers. however, students in lower secondary school seemed to benefit more from teacher support in constructing meaning and making sense of learning goals and understanding the links between learning tasks and contents than primary school students. this implies that higher order learning is challenging for adolescents, probably due to other developmental challenges that they are going through, and it needs to be very intentionally supported by teachers. the social support experienced had different roles in facilitating different elements of learning agency. for both student groups, peer support had a stronger association with problem solving than other learning activities. in turn, teacher support had a stronger interrelation with seeking scaffolding for learning. yet, the teacher support was strongly linked to the peer support, and it may be argued that to build a learning environment, facilitating learning agency calls for building opportunities both to receive and give peer support. this further requires teacher effort to scaffold peer support i.e., constructing meaning and solving problems soini | f l r 65 together, and using peers as a resource for learning in their pedagogical practices. especially in lower secondary school, students need teacher support to (re)construct the agentic learner role in school and find balance in seeking peer acceptance and meeting the learning goals. based on our results, the social support seems to play a key role as a resource for enhancing students’ learning agency, but support from a range of sources seems to have distinct roles in enhancing it. moreover, the results indicate that the sources of social support provided by the school environment complement each other. accordingly intentional regulation of individual-environment dynamics in the teacher-student and peer interaction can provide the means for cultivating complementary elements of learning agency. engagement in higher order learning may be viewed as important from several perspectives. constructing strong learning agency in school is likely to have far reaching effects on students’ ability for intentional and meaningful learning throughout their lives. this further is a prerequisite for active citizenship, being able to participate in society, and hence a crucial prerequisite for functional community life. therefore, it may be argued that equality of not just education but the whole society depends on the ability of educational systems to support and facilitate learning agency of all children. we detected some variation in students’ sense of learning agency, such differences may become problems if they reflect or result in differing opportunities to engage in agentic learning in school, having further impact on different educational and career trajectories and even life orientations. moreover, enhancing learning agency among children is also the main strategy for bringing human resources into use in solving the massive global challenges and creating new ways of living. hence, we claim that merely supporting some competencies or strategies of learning is not enough if children are unable to develop a strong sense of learning agency that carries them throughout their life. 6.2 methodological reflections this cross-sectional large-scale two cohort student explored the relationships between students’ sense of learning agency and social support received from teachers and peers in the school context. the novel learning agency (la) -measurement was develop for this purpose. path analysis confirmed a pattern of correlations that explained the students’ sense of learning agency and the social support they received for their own learning in the school context. due to the cross-sectional design, conclusions about the causality between the attributes cannot be made (e.g., lleras, 2005). the study contributed to methodological advancement in the field of learning sciences by introducing a new measure for studying a student’s sense of learning agency across the school path. the validity and reliability of the la survey was adequate. however, construct validation of the learning agency sub-scales, and longitudinal studies across the school systems and country-contexts, are needed to further validate the la measure. reliability and validity of the teacher and peer support scale was also sufficient. the support scale has been validated in prior studies (rautanen et al., 2021). all in all, despite some limitations, our study provides important insights into the factors contributing to students’ sense of learning agency that are emphasized and enhanced in the school curricula and educational system/s. the la can potentially be used to identify early signs of disengagement from higher order learning activities in school preceding the actual decrease in school achievement. key points engaging in higher order learning requires strong sense of learning agency. learning agency consists of will and power to act in terms of learning and believe that one can do it. soini | f l r 66 the novel learning agency (la) -measurement was develop for studying learning agency of primary and lower secondary students. differences in learning agency and social support in terms of grade level, gender and ses were detected. the social support from teachers and peers had different relationships with different elements of learning agency. acknowledgements this work was supported by the academy of finland grant 295022; strategic research council within the academy of finland grant number 352509 and the finnish cultural foundation grant 50201386. references anderson, r. c., graham, m., kennedy, p., nelson, n., stoolmiller, m., baker, s. k., & fien, h. (2019). student agency at the crux: mitigating disengagement in middle and high school. contemporary educational psychology, 56, 205–217. https://doi.org/10.1016/j.cedpsych.2018.12.005 anderson, l.w., krathwohl, d.r., eds. (2001). a taxonomy for learning, teaching, and assessing: a revision of bloom's taxonomy of educational objectives. new york: longman. isbn 978-0-8013-1903-7. ananiadou, k., & claro, m. (2009). 21st century skills and competences for new millennium learners in oecd countries. oecd education working papers no. 41. paris: oecd. https://doi.org/10.1787/19939019 anttila, h., pyhältö, k., pietarinen, j. & soini, k. (2018). socially embedded academic emotions in school. journal of education and learning, 7(3), 87–101. https://doi.org/10.5539/jel.v7n3p87 bandura, a. (1989). human agency in social cognitive theory. american psychologist, 44(9), 1175– 1184. https://doi.org/10.1037/0003-066x.44.9.1175 berisha, a.-k. & seppänen, p. (2017). pupil selection segments urban comprehensive schooling in finland: composition of school classes in pupils’ school performance, gender, and ethnicity. scandinavian journal of educational research, 61(2), 240–254. https://doi.org/10.1080/00313831.2015.1120235 bernelius, v. & vaattovaara, m. (2016). choice and segregation in the ‘most egalitarian’ schools: cumulative decline in urban schools and neighborhoods of helsinki, finland. urban studies, 53(15), 3155–3171. https://doi.org/10.1177/0042098015621441 bernelius, v. (2013). eriytyvät kaupunkikoulut. helsingin peruskoulujen oppilaspohjan erot, perheiden kouluvalinnat ja oppimistuloksiin liittyvät aluevaikutukset osana kaupungin eriytymiskehitystä. helsingin kaupungin tietokeskus, tutkimuksia 2013:1. helsinki: edita prima. bloom, b. s.; engelhart, m. d.; furst, e. j.; hill, w. h.; krathwohl, d. r. (1956). taxonomy of educational objectives: the classification of educational goals. vol. handbook i: cognitive domain. new york: david mckay company https://doi.org/10.1016/j.cedpsych.2018.12.005 https://doi.org/10.1787/19939019 https://doi.org/10.5539/jel.v7n3p87 https://doi.org/10.1037/0003-066x.44.9.1175 https://doi.org/10.1080/00313831.2015.1120235 https://doi.org/10.1177/0042098015621441 soini | f l r 67 bottiani, j. h., duran, c. a. k., pas, e. t., & bradshaw, c. p. (2019). teacher stress and burnout in urban middle schools: associations with job demands, resources, and effective classroom practices. journal of school psychology, 77, 36–51. https://doi.org/10.1016/j.jsp.2019.10.002 brass, n., mckellar, s. e., north, e. a., & ryan, a. m. (2019). early adolescents’ adjustment at school: a fresh look at grade and gender differences. the journal of early adolescence, 39(5), 689–716. https://doi.org/10.1177/0272431618791291 bursal, m. (2017). academic achievement and perceived peer support among turkish students: gender and preschool education impact. international electronic journal of elementary education, 9(3), 599–612. https://www.iejee.com/index.php/iejee/article/view/178 byrne, b. m. (2012). structural equation modeling with mplus: basic concepts, applications, and programming. 3rd ed.; multivariate applications series; routledge, taylor & francis group: new york, ny, usa, 2016; isbn 978-1-138-79702-4. camara, m., bacigalupe, g., & padilla, p. (2017). the role of social support in adolescents: are you helping me or stressing me out? international journal of adolescence and youth, 22(2), 123–136. https://doi.org/10.1080/02673843.2013.875480 cheng, h.n.h., yang, e.f.y., liao, c.c.y., chang, b., huang, y.c.y. & chan, t-w. (2015). scaffold seeking: a reverse design of scaffolding in computer-supported word problem solving. journal of educational computing research. 53(3), 409–435. https://doi.org/10.1177/0735633115601598 clarke, s. n., howley, i., resnick, l., & penstein rosé, c. (2016). student agency to participate in dialogic science discussions. learning, culture and social interaction, 10, 27–39. https://doi.org/10.1016/j.lcsi.2016.01.002 clarke, s.n. (2015). the right to speak. in l.b. resnick, c.s.c. asterhan, & s.n. clarke (eds.), socializing intelligence through academic talk and dialogue. washington, dc: american educational research association. https://www.jstor.org/stable/j.ctt1s474m1.16 coburn, c.e. (2001). collective sensemaking about reading: how teachers mediate reading policy in their professional communities. educational evaluation and policy analysis, 23, (2), 145–170. https://doi.org/10.3102/01623737023002145 cohen, s., underwood, l. g., gottlieb, b. h., & fetzer institute. (2000). social support measurement and intervention: a guide for health and social scientists. new york; oxford: oxford university press. https://doi.org/10.1093/med:psych/9780195126709.001.0001 comrey, a. l., & lee, h. b. (2013). a first course in factor analysis. new york, ny: psychology press. https://doi.org/10.4324/9781315827506 darling-hammond, l., flook, l., cook-harvey, c., barron, b., & osher, d. (2020). implications for educational practice of the science of learning and development. applied developmental science, 24(2), 97–140. https://doi.org/10.1080/10888691.2018.1537791 dillon, j. t. (1982). problem finding and solving. the journal of creative behavior, 16(2), 97–111. https://psycnet.apa.org/doi/10.1002/j.2162-6057.1982.tb00326.x duckworth, a. & seligman, m. (2006). self-discipline gives girls the edge; gender in self-discipline, grades and achievement test scores. journal of educational psychology, 98. 198–208. https://doi.org/10.1037/00220663.98.1.198 https://doi.org/10.1016/j.jsp.2019.10.002 https://doi.org/10.1177/0272431618791291 https://www.iejee.com/index.php/iejee/article/view/178 https://doi.org/10.1080/02673843.2013.875480 https://doi.org/10.1177/0735633115601598 https://doi.org/10.1016/j.lcsi.2016.01.002 https://www.jstor.org/stable/j.ctt1s474m1.16 https://doi.org/10.3102/01623737023002145 https://doi.org/10.1093/med:psych/9780195126709.001.0001 https://doi.org/10.4324/9781315827506 https://doi.org/10.1080/10888691.2018.1537791 https://psycnet.apa.org/doi/10.1002/j.2162-6057.1982.tb00326.x https://psycnet.apa.org/doi/10.1037/0022-0663.98.1.198 https://psycnet.apa.org/doi/10.1037/0022-0663.98.1.198 soini | f l r 68 dwyer, c.p., hogan, m.j. & stewart, i. (2014). an integrated critical thinking framework for the 21st century. thinking skills and creativity, 12, 43–52. https://doi.org/10.1016/j.tsc.2013.12.004 edwards, a. (2005). relational agency: learning to be a resourceful practitioner. international journal of educational research, 43(3), 168–182. https://doi.org/10.1016/j.ijer.2006.06.010 estell, d. b., & perdue, n. h. (2013). social support and behavioral and affective school engagement: the effects of peers, parents, and teachers. psychology in the schools, 50(4), 325–339. https://doi.org/10.1002/pits.21681 evans, g. w., & rosenbaum, j. (2008). self-regulation and the income-achievement gap. early childhood research quarterly, 23(4), 504–514. https://doi.org/10.1016/j.ecresq.2008.07.002 finnish advisory board for research integrity. (2009). ethical principles of research in the humanities and social and behavioural sciences and proposals for ethical review. https://www.tenk.fi/sites/tenk.fi/files/ethicalprinciples.pdf. accessed 12 august 2020 finnish national agency for education. (2017). finnish education in a nutshell [pdf]. retrieved from http://www.oph.fi/download/171176_finnish_education_in_a_nutshell.pdf fischer, f. t., schult, j., & hell, b. (2013). sex-specific differential prediction of college admission tests: a metaanalysis. journal of educational psychology, 105(2), 478–488. https://doi.org/10.1037/a0031956 flaspohler, p. d., elfstrom, j. l., vanderzee, k. l., sink, h. e., & birchmeier, z. (2009). stand by me: the effects of peer and teacher support in mitigating the impact of bullying on quality of life. psychology in the schools, 46(7), 636–649. https://doi.org/10.1002/pits.20404 furrer, c., & skinner, e. (2003). sense of relatedness as a factor in children's academic engagement and performance. journal of educational psychology, 95(1), 148–162. https://doi.org/10.1037/00220663.95.1.148 giddens, a. (1984) the constitution of society. outline of the theory of structuration. university of california press, berkeley. gini, g., pozzoli, t., borghi, f., & franzoni, l. (2008). the role of bystanders in students’ perception of bullying and sense of safety. journal of school psychology, 46(6), 617–638. https://doi.org/10.1016/j.jsp.2008.02.001 greiff, s., wüstenberg, g., molnár, a., fischer, j. & csapó, b. (2013). complex problem solving in educational contexts – something beyond g: cocept, assessmenr, measurement invariance, and construct validity. journal of educational psychology, 105, 364–379. griffin, p., mcgaw, b. & care, e. (eds.), assessment and teaching of 21st century skills (pp. 17–66). new york: springer. hadwin, a. f., järvelä, s., & miller, m. (2011). self-regulated, co-regulated, and socially shared regulation of learning. in b. j. zimmerman & d. h. schunk (eds.), handbook of self-regulation of learning and performance (pp. 65–84). new york, ny: routledge. halinen, i., & järvinen, r. (2008). towards inclusive education: the case of finland. prospects, 38(1), 77–97. https://doi.org/10.1007/s11125-008-9061-2 hansen, y. k., gustafsson, j.e..& rosén, m. 2014. school performance difference and policy variations in finland, norway and sweden. in k. yang hansen, j.-e. gustafsson, m. rosén, s. sulkunen, k. nissinen, p. kupari, https://doi.org/10.1016/j.tsc.2013.12.004 https://doi.org/10.1016/j.ijer.2006.06.010 https://doi.org/10.1002/pits.21681 https://doi.org/10.1016/j.ecresq.2008.07.002 http://www.oph.fi/download/171176_finnish_education_in_a_nutshell.pdf https://doi.org/10.1037/a0031956 https://doi.org/10.1002/pits.20404 https://doi.org/10.1037/0022-%200663.95.1.148 https://doi.org/10.1016/j.jsp.2008.02.001 https://doi.org/10.1007/s11125-008-9061-2 soini | f l r 69 r.f. ólafsson, j.k. björnsson, l.s. grønmo, l. rønberg, j. mejding, i.c. borge, & a. hole. northern lights on timms and pirls 2011. temanord 2014:528, 25–48. häkkinen, p., järvelä, s., mäkitalo-siegl, k., ahonen, a., näykki, p., & valtonen, t. (2017). preparing teacherstudents for twenty-first-century learning practices (prep 21): a framework for enhancing collaborative problem-solving and strategic learning skills. teachers and teaching, theory and practice, 23(1), 25–41. https://doi.org/10.1080/13540602.2016.1203772 hombrados-mendieta, m. i., gomez-jacinto, l., dominguez-fuentes, j. m., garcia-leiva, p. & castro-trave, m. (2012). types of social support provided by parents, teachers, and classmates during adolescence. journal of community psychology, 40(6), 645–664. https://doi.org/10.1002/jcop.20523 hmelo-silver, c.e. (2004). problem-based learning: what and how do students learn? educational psychology review 16, 235–266. https://doi.org/10.1023/b:edpr.0000034022.16470.f3 hu, l., & bentler, p. m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus alternatives. structural equationmodeling: a multidisciplinary journal, 6(1), 1–55. https://doi.org/10.1080/10705519909540118 hooper, d., coughlan, j., & mullen, m. r. (2008). structural equation modelling: guidelines for determining model fit. electronic journal of business research methods, 6(1), 53–60. available online at: www.ejbrm.com janosz, m., archambault, i., morizot, j., & pagani, l. s. (2008). school engagement trajectories and their differential predictive relations to dropout. journal of social issues, 64, 21–40. https://doi.org/10.1111/j.15404560.2008.00546.x järvenoja, h., volet, s., & järvelä, s. (2013). regulation of emotions in socially challenging learning situations: an instrument to measure the adaptive and social nature of the regulation process. educational psychology, 33(1), 31–58. https://doi.org/10.1080/01443410.2012.742334 jiang, x., huebner, e. s., & siddall, j. (2013). a short-term longitudinal study of differential sources of schoolrelated social support and adolescents’ school satisfaction. social indicators research, 114(3), 1073–1086. https://doi.org/10.1007/s11205-012-0190-x kelly, s. (2008). race, social class, and student engagement in middle school english classrooms. social science research, 37, 434–44. https://doi.org/10.1016/j.ssresearch.2007.08.003 kiefer, s., alley, k., & ellerbrock, c. (2015). teacher and peer support for young adolescents’ motivation, engagement and school belonging. rmle online, 38(8), 1-8. https://doi.org/10.1080/19404476.2015.11641184 kindermann, t. a. (2007). effects of naturally-existing peer groups on changes in academic engagement in a cohort of sixth graders. child development, 78, 1186–1203. https://www.jstor.org/stable/4620697 kumpulainen, k., & lankinen, t. (2012). striving for educational equity and excellence: evaluation and assessment in finnish basic education. in h. niemi, a. toom & a. kallioniemi. (eds.), the principles and practices of teaching and learning in finnish schools. sense publishers: rotterdam. lakens d. (2013). calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and anovas. frontiers in psychology, 4, 863. https://doi.org/10.3389/fpsyg.2013.00863 https://doi.org/10.1080/13540602.2016.1203772 https://doi.org/10.1002/jcop.20523 https://doi.org/10.1023/b:edpr.0000034022.16470.f3 https://doi.org/10.1080/10705519909540118 http://www.ejbrm.com/ https://doi.org/10.1111/j.1540-4560.2008.00546.x https://doi.org/10.1111/j.1540-4560.2008.00546.x https://doi.org/10.1080/01443410.2012.742334 https://doi.org/10.1007/s11205-012-0190-x https://doi.org/10.1016/j.ssresearch.2007.08.003 https://doi.org/10.1080/19404476.2015.11641184 https://www.jstor.org/stable/4620697 https://doi.org/10.3389/fpsyg.2013.00863 soini | f l r 70 lam, s. f., wong, b. p. h., yang, h., & liu, y. (2012). understanding student engagement with a contextual model. in s. christenson, a. reschly & c. wylie (eds.), handbook of research on student engagement (pp. 403–419). springer. https://doi.org/10.1007/978-1-4614-2018-7_19 lave j. and wenger e. (1991). situated learning: legitimate peripheral participation. cambridge university press, cambridge. lefstein, a., vedder-weiss, d., tabak, i., & segal, a. (2018). learner agency in scaffolding: the case of coaching teacher leadership. international journal of educational research, 90, 209–222. https://doi.org/10.1016/j.ijer.2017.11.002 li, y., & lerner, r. m. (2011). trajectories of school engagement during adolescence: implications for grades, depression, delinquency, and substance use. developmental psychology, 47(1), 233–247. https://doi.org/10.1037/a0021307 lindfors, p., minkkinen, m., rimpelä, a., & hotulainen, r. (2018). family and school social capital, school burnout and academic achievement: a multilevel longitudinal analysis among finnish pupils. international journal of adolescence and youth, 23(3), 368–381. https://doi.org/10.1080/02673843.2017.1389758 liu, w., mei, j., tian, l., & huebner, e. s. (2016). age and gender differences in the relation between school-related social support and subjective well-being in school among students. social indicators research, 125(3), 1065– 1083. http://dx.doi.org/10.1007/s11205-015-0873-1 lleras, c. (2005). path analysis. encyclopedia of social measurement. new york: academic press. malecki, c. k., & demaray, m. k. (2002). measuring perceived social support: development of the child and adolescent social support scale. psychology in the schools, 39, 1–18. https://doi.org/10.1002/pits.10004 marsick, v. j., & watkins, k. e. (1992). continuous learning in the workplace. adult learning, 3(4), 9–12. https://doi.org/10.1177/104515959200300404 maurer, t. j., & weiss, e. m. (2010). continuous learning skill demands: associations with managerial job content, age, and experience. journal of business psychology, 25, 1–13. https://doi.org/10.1007/s10869-009-9126-0 mega, c., ronconi, l., & de beni, r. (2014). what makes a good student? how emotions, self-regulated learning, and motivation contribute to academic achievement. journal of educational psychology, 106(1), 121–131. https://doi.org/10.1037/a0033546 ministry of education. (2012). tulevaisuuden perusopetus – valtakunnalliset tavoitteet ja tuntijako. [finnish ministry of education. the basic education in the future. objectives and curricular basis]. opetusja kulttuuriministeriön työryhmämuistioita ja selvityksiä 2012. muthén, b., & satorra, a. (1995). complex sample data in structural equation modeling. in p. marsden (ed.), sociological methodology. washington, dc: american sociological association. muthén, l.k. & muthén, b.o. (1998-2017). mplus user’s guide. eighth edition. los angeles, ca: muthén & muthén myllyniemi, s. & kiilakoski, t. (2018). tilasto-osio. in e. pekkarinen & s. myllyniemi (eds.) opin polut ja pientareet. nuorisobarometri 2017 (pp. 9–117). helsinki: opetusja kulttuuriministeriö & nuorisotutkimusseura & valtion nuorisoneuvosto. https://doi.org/10.1007/978-1-4614-2018-7_19 https://doi.org/10.1016/j.ijer.2017.11.002 https://doi.org/10.1037/a0021307 https://doi.org/10.1080/02673843.2017.1389758 http://dx.doi.org/10.1007/s11205-015-0873-1 https://doi.org/10.1002/pits.10004 https://doi.org/10.1177/104515959200300404 https://doi.org/10.1007/s10869-009-9126-0 https://doi.org/10.1037/a0033546 soini | f l r 71 oecd (2013), "the skills needed for the 21st century", in oecd skills outlook 2013: first results from the survey of adult skills, oecd publishing, paris, https://doi.org/10.1787/9789264204256-5-en oecd (2019), pisa 2018 results (volume ii): where all students can succeed, pisa, oecd publishing, paris, https://doi.org/10.1787/b5fd1b8f-en park, c. l. (2011). meaning, coping, and health and well-being. in s. folkman (eds.), the oxford handbook of stress, health, and coping. oxford university press. pekrun, r., goetz, t., titz, w., & perry, r. (2002). academic emotions in students’ self-regulated learning and achievement: a program of qualitative and quantitative research. educational psychologist, 37(2), 91–105. http://doi.org/10.1207/s15326985ep3702_4 pellegrino, j. (2017), "teaching, learning and assessing 21st century skills", in guerriero, s. (eds.), pedagogical knowledge and the changing nature of the teaching profession, oecd publishing, paris, https://doi.org/10.1787/9789264270695-12 peugh, j. l. (2010). a practical guide to multilevel modeling. journal of school psychology, 48(1), 85–112. https://doi.org/10.1016/j.jsp.2009.09.002 pietarinen, j. pyhältö, k. & soini, t. (2016). teacher’s professional agency – a relational approach to teacher learning. learning: research and practice, 2(2). http://doi.org/10.1080/23735082.2016.1181196. pintrich, p. r. (2004). a conceptual framework for assessing motivation and self-regulated learning in college students. educational psychology review, 16(4), 385–407. https://doi.org/10.1007/s10648-004-0006-x pyhältö, k. pietarinen, k. & soini, t. (2015). teachers’ professional agency and learning – from adaption to active modification in the teacher community. teachers and teaching: theory and practice. 21(7), 811–830. https://doi.org10.1080/13540602.2014.995483 raffaelli, m., crockett, l.j. & shen, y-l. (2005). developmental stability and change in self-regulation from childhood to adolescence. the journal of genetic psychology, 166(1), 54–75. https://doi.org/10.3200/gntp.166.1.54-76 rautanen, p., soini, t., pietarinen, j., & pyhälto, k. (2022). dynamics between perceived social support and study engagement among primary school students: a three-year longitudinal survey. social psychology of education, 25, 1481–1505. https://doi.org/10.1007/s11218-022-09734-2 rautanen, p., soini, t., pietarinen, j. & pyhältö, k. (2021). primary school students’ perceived social support in relation to study engagement. european journal of psychology of education, 36, 653–672. https://doi.org/10.1007/s10212-020-00492-3 reeve, j. (2013). how students create motivationally supportive learning environments for themselves: the concept of agentic engagement. journal of educational psychology, 105(3), 579–595. https://doi.org/10.1037/a0032690 reeve, j., nix, g., & hamm, d. (2003). testing models of the experience of self-determination in intrinsic motivation and the conundrum of choice. journal of educational psychology, 95(2), 375–392. https://doi.org/10.1037/0022-0663.95.2.375 resnick, l. b. (1995). from aptitude to effort: a new foundation for our schools. daedalus, 124(4), 55–62. http://www.jstor.org/stable/20027327. resnick, l. b. (1987). education and learning to think. washington, dc: national academy press. https://doi.org/10.1787/9789264204256-5-en https://doi.org/10.1787/b5fd1b8f-en http://doi.org/10.1207/s15326985ep3702_4 https://doi.org/10.1787/9789264270695-12 https://doi.org/10.1016/j.jsp.2009.09.002 http://doi.org/10.1080/23735082.2016.1181196 https://doi.org/10.1007/s10648-004-0006-x https://doi.org10.1080/13540602.2014.995483 https://doi.org/10.3200/gntp.166.1.54-76 https://doi.org/10.1007/s11218-022-09734-2 https://doi.org/10.1007/s10212-020-00492-3 https://psycnet.apa.org/doi/10.1037/a0032690 https://psycnet.apa.org/doi/10.1037/0022-0663.95.2.375 http://www.jstor.org/stable/20027327 soini | f l r 72 roseth, c. j., d. w. johnson, and r. t. johnson. 2008. “promoting early adolescents’ achievement and peer relationships: the effects of cooperative, competitive, and individualistic goal structures.” psychological bulletin 134 (2): 223–246. https://doi.org/10.1037/0033-2909.134.2.223. sage, n. a., & kindermann, t. a. (1999). peer networks, behavior contingencies, and children’s engagement in the classroom. merrill-palmer quarterly, 45(1), 143–171. https://www.jstor.org/stable/23093320 salmela-aro, k., kiuru, n., pietikäinen, m., & jokela, j. (2008). does school matter? the role of school context in adolescents' school-related burnout. european psychologist, 13(1), 12–23. https://doi.org/10.1027/10169040.13.1.12 scardamalia, m., & bereiter, c. (2006). knowledge building: theory, pedagogy, and technology. in r. k. sawyer (ed.), the cambridge handbook of the learning sciences (pp. 97–118). new york, ny: cambridge university press. scherer, r. & gustafsson, j-e. (2015). the relations among openness, perseverance, and performance in creative problem solving: a substantive-methodological approach. thinking skills and creativity, 18, 4–17. https://doi.org/10.1016/j.tsc.2015.04.004 schober, b. , lüftenegger, m. , wagner, p. , finsterwald, m. & spiel, c. (2013). facilitating lifelong learning in school-age learners. european psychologist, 18(2), 114–125. https://doi.org/10.1027/1016-9040/a000129. skinner, e., furrer, c., marchand, g., & kindermann, t. (2008). engagement and disaffection in the classroom: part of a larger motivational dynamic? journal of educational psychology, 100(4), 765-781. https://doi.org/10.1037/a0012840 snijders, t. a. b., & bosker, r. j. (2012). multilevel analysis. an introduction to basic and advanced multilevel modeling (2nd ed.). thousand oaks: sage publications. soini, t., pietarinen, j. & pyhältö, k. (2016). what if teachers learn in the classroom? teacher development, 20(3), https://doi.org/10.1080/13664530.2016.1149511 soini, t., pietarinen, j., toom, a. & pyhältö, k. (2015). what contributes to first year student teacher’s sense of professional agency in the classroom? teachers and teaching: theory and practice, 21(6), 641-659. https://doi.org/10.1080/13540602.2015.1044326. størksen, i., ellingsen, i. t., wanless, s. b., & mcclelland, m. m. (2015). the influence of parental socioeconomic background and gender on self-regulation among 5-year-old children in norway. early education and development, 26(5–6), 663–684. https://doi.org/10.1080/10409289.2014.932238 tennant, j. e., demaray, m. k., malecki, c. k., terry, m. n., clary, m., & elzinga, n. (2015). students' ratings of teacher support and academic and social-emotional wellbeing. school psychology quarterly, 30(4), 494-512. https://doi.org/10.1037/spq0000106 tóth-király, i., bõthe, b., rigó, a., & orosz, g. (2017). an illustration of the exploratory structural equation modeling (esem) framework on the passion scale. frontiers in psychology, 8, 1968. https://doi.org/10.3389/fpsyg.2017.01968 ulmanen, s. soini, t., pietarinen, j. & pyhältö, k. (2016). students’ experiences of the development of emotional engagement. international journal of educational research, 76, 86–96. https://doi.org/10.1016/j.ijer.2016.06.003 https://doi.org/10.1037/0033-2909.134.2.223 https://www.jstor.org/stable/23093320 https://doi.org/10.1027/1016-9040.13.1.12 https://doi.org/10.1027/1016-9040.13.1.12 https://doi.org/10.1016/j.tsc.2015.04.004 https://doi.org/10.1027/1016-9040/a000129 https://psycnet.apa.org/doi/10.1037/a0012840 https://doi.org/10.1080/13664530.2016.1149511 https://doi.org/10.1080/13540602.2015.1044326 https://doi.org/10.1080/10409289.2014.932238 https://psycnet.apa.org/doi/10.1037/spq0000106 https://doi.org/10.3389/fpsyg.2017.01968 https://doi.org/10.1016/j.ijer.2016.06.003 soini | f l r 73 ulmanen, s., soini, t., pietarinen, j. & pyhältö, k. (2014). strategies for academic engagement perceived by finnish sixth and eighth graders. cambridge journal of education, 44(3), 425–443. https://doi.org/10.1080/0305764x.2014.921281 urdan, t., & schoenfelder, e. (2006). classroom effects on student motivation: goal structures, social relationships, and competence beliefs. journal of school psychology, 44(5), 331–349. https://doi.org/10.1016/j.jsp.2006.04.003 valle, a., cabanach, r.g., núñez, j.c., gonza´lez-pienda, j., rodrı´guez, s. & piñeiro, i. (2003). cognitive, motivational, and volitional dimensions of learning: an empirical test of a hypothetical model. research in higher education, 44(5), 557–580. https://doi.org/10.1023/a:1025443325499 vainikainen m-p. & hautamäki j. (2020): three studies on learning to learn in finland: anti-flynn effects 2001– 2017, scandinavian journal of educational research, http://doi.org/10.1080/00313831.2020.1833240. vainikainen, m-p., hautamäki, j., hotulainen, r. & kupiainen, s. (2015). general and specific thinking skills and schooling: preparing the mind to new learning. thinking skills and creativity 18, 53–64. https://doi.org/10.1016/j.tsc.2015.04.006 van roy, b., kristensen, h., groholt, b., & clench-aas, j. (2009). prevalence and characteristics of significant social anxiety in children aged 8–13 years: a norwegian cross-sectional population study. social psychiatry and psychiatric epidemiology, 44(5), 407–415. https://doi.org/10.1007/s00127-008-0445-7 vansteenkiste, m., sierens, e., soenens, b., luyckx, k., & lens, w. (2009). motivational profiles from a selfdetermination perspective: the quality of motivation matters. journal of educational psychology, 101, 671– 688. https://doi.org/10.1037/a0015083 waeytens, k., lens, w., & vandenberghe, r. (2002). learning to learn: teachers’ conceptions of their supporting role. learning and instruction, 12(3), 305–322. https://doi.org/10.1016/s0959-4752(01)00024-x wang, m. t., & eccles, j. s. (2012). social support matters: longitudinal effects of social support on three dimensions of school engagement from middle to high school. child development, 83(3), 877– 895. https://doi.org/10.1111/j.1467-8624.2012.01745.x wanless, s. b., mcclelland, m. m., lan, x., son, s.-h., cameron, c. e., morrison, f. j., chen, f.-m., chen, j.-l., li, s., lee, k., & sung, m. (2013). gender differences in behavioral regulation in four societies: the united states, taiwan, south korea, and china. early childhood research quarterly, 28(3), 621–633. https://doi.org/10.1016/j.ecresq.2013.04.002 weick, k. e., sutcliffe, k. m., & obstfeld, d. (2005). organizing and the process of sensemaking. organization science, 16(4), 409–421. https://doi.org/10.1287/orsc.1050.0133 wentzel, k. r. (1998). social relationships and motivation in middle school: the role of parents, teachers, and peers. journal of educational psychology, 90(2), 202–209. https://doi.org/10.1037/0022-0663.90.2.202 weis, m., heikamp, t., & trommsdorff, g. (2013). gender differences in school achievement: the role of selfregulation. frontiers in psychology, 4, 442–442. https://doi.org/10.3389/fpsyg.2013.00442 wentzel, k. r., russell, s., & baker, s. (2016). emotional support and expectations from parents, teachers, and peers predict adolescent competence at school. journal of educational psychology, 108(2), 242– 255. https://doi.org/10.1037/edu0000049 https://doi.org/10.1080/0305764x.2014.921281 https://doi.org/10.1016/j.jsp.2006.04.003 https://doi.org/10.1023/a:1025443325499 http://doi.org/10.1080/00313831.2020.1833240 https://doi.org/10.1016/j.tsc.2015.04.006 https://doi.org/10.1007/s00127-008-0445-7 https://doi.org/10.1037/a0015083 https://doi.org/10.1016/s0959-4752(01)00024-x https://doi.org/10.1111/j.1467-8624.2012.01745.x https://doi.org/10.1016/j.ecresq.2013.04.002 https://doi.org/10.1287/orsc.1050.0133 https://doi.org/10.1037/0022-0663.90.2.202 https://doi.org/10.3389/fpsyg.2013.00442 https://doi.org/10.1037/edu0000049 soini | f l r 74 wentzel, k. r., muenks, k., mcneish, d., & russell, s. (2017). peer and teacher supports in relation to motivation and effort: a multi-level study. contemporary educational psychology, 49, 32–45. https://doi.org/10.1016/j.cedpsych.2016.11.002 xing, x., liu, x., & wang, m. (2019). parental warmth and harsh discipline as mediators of the relations between family ses and chinese preschooler’s inhibitory control. early childhood research quarterly, 48(3), 237– 245. https://doi.org/10.1016/j.ecresq.2018.12.018 zimmerman, h. t., & weible, j. l. (2018). epistemic agency in an environmental sciences watershed investigation fostered by digital photography. international journal of science education, 40(8), 894–918. https://doi.org/10.1080/09500693.2018.145511 schunk, d.h., & zimmerman b.j. (eds.) (2008). motivation and self-regulated learning: theory, research, and applications (2nd ed.), lawrence erlbaum associates publishers, new york, london. zittoun, t., brinkmann, s. (2012). learning as meaning making. in: seel, n.m. (eds) encyclopedia of the sciences of learning. springer, boston, ma. https://doi.org/10.1007/978-1-4419-1428-6_1851 https://doi.org/10.1016/j.cedpsych.2016.11.002 https://doi.org/10.1016/j.ecresq.2018.12.018 https://doi.org/10.1080/09500693.2018.145511 https://doi.org/10.1007/978-1-4419-1428-6_1851 soini | f l r 75 appendix table a1. students’ learning agency (la) subscales and reliability analysis (translated from finnish) items elements of the la seeking scaffolding, α=.79/.81 (4th/7th graders) sc1 it is important that i tell the teacher when i have problems at school motivation sc2 it is important for me to ask if i don’t understand the task set by the teacher motivation sc3 i feel confident enough to ask for extra help if i experience difficulties in my studies efficacy sc4 i ask for help if i am not able to complete the task, we have been set strategy sc5 i can confidently ask about a topic i don’t understand strategy meaning making, α=.89/.89 (4th/7th graders) mm1 i often look for more information in order to understand the things i am studying strategy mm2 i get interested in a subject when i realize how it connects to something, i have previously learned about motivation mm3 i can independently work out problems that were unresolved during lesson time strategy mm4 i get interested in what i’m studying when i understand how it connects to things in my own life motivation mm5 i learn better when i can find an example from my own life that relates to what i am studying efficacy mm6 i try to think back to what i have previously learned about a topic when i am learning something new strategy mm7 i find it stimulating to think about school work outside of lesson time motivation mm8 i often reflect on what i am learning strategy problem solving, α=.89/.91 (4th/7th graders) ps1 i want to finish even the most difficult homework assignments motivation ps2 i enjoy tasks that require me to solve problems motivation ps3 i often come up with alternative solutions to the tasks i’ve been given strategy ps4 it’s a good idea to consider the different perspectives on topics we are learning about efficacy ps5 i perform best in tasks in which there is a specific solution to a problem efficacy ps6 i am usually able to solve problems related to my studies efficacy soini | f l r 76 table a2. correlations between the items in learning agency subscales in fourth and seventh graders’ cohorts correlations items 1 2 3 4 5 6 7 8 10 11 12 15 16 17 19 20 21 22 25 1 sc1 .48 .39 .40 .38 .32 .23 .16 .21 .15 .27 .36 .32 .33 .32 .26 .30 .27 .27 2 sc2 .53 .37 .60 .59 .32 .22 .21 .25 .22 .34 .30 .38 .33 .37 .29 .33 .25 .31 3 sc3 .36 .33 .42 .44 .22 .16 .15 .24 .18 .29 .27 .25 .32 .32 .20 .29 .25 .30 4 sc4 .46 .61 .36 .71 .28 .18 .23 .27 .21 .36 .29 .30 .30 .31 .26 .36 .24 .27 5 sc5 .46 .58 .38 .67 .30 .19 .26 .28 .23 .37 .26 .30 .28 .28 .24 .34 .26 .32 6 ps1 .25 .29 .26 .24 .29 .55 .45 .53 .46 .49 .42 .43 .44 .32 .29 .43 .46 .39 7 ps2 .25 .25 .25 .19 .23 .46 .50 .49 .61 .45 .40 .38 .38 .31 .30 .37 .46 .37 8 ps3 .22 .19 .19 .14 .20 .35 .38 .53 .46 .42 .44 .43 .37 .33 .34 .37 .41 .35 10 ps4 .25 .23 .25 .23 .31 .44 .40 .47 .50 .57 .43 .41 .48 .36 .37 .46 .47 .42 11 ps5 .19 .23 .26 .20 .26 .43 .56 .40 .47 .46 .30 .31 .35 .33 .32 .35 .39 .35 12 ps6 .27 .28 .26 .24 .34 .43 .42 .38 .52 .49 .40 .36 .42 .37 .39 .45 .41 .40 15 mm1 .29 .27 .29 .23 .27 .34 .33 .40 .36 .33 .31 .42 .54 .35 .30 .42 .46 .40 16 mm2 .31 .35 .31 .32 .33 .35 .30 .35 .37 .33 .34 .41 .49 .46 .41 .48 .45 .44 17 mm3 .27 .25 .30 .22 .28 .36 .31 .38 .40 .36 .37 .42 .35 .36 .33 .49 .51 .49 19 mm4 .30 .32 .32 .26 .31 .33 .32 .31 .37 .34 .35 .42 .50 .37 .61 .40 .42 .39 20 mm5 .21 .22 .28 .19 .24 .29 .32 .37 .40 .37 .38 .39 .37 .34 .51 .40 .35 .34 21 mm6 .29 .32 .31 .31 .33 .35 .34 .35 .39 .39 .42 .41 .45 .40 .41 .41 .45 .53 22 mm7 .27 .30 .35 .27 .32 .42 .42 .37 .42 .43 .37 .42 .40 .44 .46 .41 .46 .61 25 mm8 .28 .31 .33 .26 .31 .35 .32 .37 .40 .38 .36 .41 .40 .45 .43 .40 .50 .56 note. all correlation is significant at the p<0.01 level. fourth graders (n=2392) are below the diagonal and seventh graders (n=1546) above the diagonal. soini | f l r 77 table a3. standardized parameter estimates for the scales of learning agency in path model (esem) consisting of the associations between latent social support variables separately within 4th graders and 7th graders (see figure 2) 4lk (n=2363) 7lk (n=1528) items sc () mm () ps () res.var sc () mm () ps () res.var sc1 .56 .11 .00 .60 .39 .41 -.21 .63 sc2 .74 .03 -.01 .44 .68 .13 -.07 .47 sc3 .32 .34 -.05 .70 .44 .27 -.14 .68 sc4 .87 -.07 -.06 .34 .88 -.15 .06 .32 sc5 .80 -.09 .10 .37 .91 -.22 .13 .28 mm1 -.01 .64 .01 .60 .01 .56 .12 .58 mm2 .13 .58 -.03 .59 .05 .60 .05 .55 mm3 -.01 .52 .12 .63 -.02 .70 .04 .49 mm4 .04 .69 -.05 .55 .12 .47 .06 .65 mm5 -.08 .57 .12 .62 .07 .35 .17 .72 mm6 .07 .56 .08 .56 .10 .52 .13 .54 mm7 .00 .63 .10 .50 -.10 .66 .16 .47 mm8 -.00 .72 -.03 .52 .02 .65 .06 .51 ps1 .07 .03 .57 .60 .05 .19 .53 .51 ps2 -.01 -.03 .69 .56 -.11 .10 .71 .45 ps3 -.11 .27 .43 .62 -.03 .10 .62 .54 ps4 .00 .07 .65 .51 -.01 .15 .64 .44 ps5 -.05 -.03 .76 .49 -.01 -.11 .81 .46 ps6 .08 -.09 .73 .50 .18 .09 .52 .51 note. sc=seeking scaffolding; mm=meaning making; ps=problem solving; res.var.=residual variance; =factor loadings, target loadings are in bold; non-significant parameters (p≥.05) are italicized frontline learning research vol. 13 no.1 (2025) 22 -44 issn 2295-3159 corresponding author: kalypso iordanou, university of central lancashire cyprus, kiordanou@uclan.ac.uk doi: https://doi.org/10.14786/flr.v13i1.1241 supporting integration of multiple source perspectives through dialogic argumentation kalypso iordanou & constantina fotiou university of central lancashire cyprus, cyprus article received 2 february 2023 / article revised 18 september 2024 / accepted 9 february 2025 / available online 18 february 2025 abstract we report a study examining, for the first time, the effectiveness of engagement in dialogic argumentation in relation to its ability to promote integration of multiple source perspectives in an argumentive writing task after reading controversial multiple texts. sixty-four primary school students engaged in a dialog-based intervention aiming to support them to learn to argue. participants’ argument skills have been improved and transferred to a writing task completed after reading novel multiple texts on new, non-intervention, topics. in particular, the experimental group participants showed gains in their ability to integrate multiple source perspectives in an argumentive writing task after reading controversial multiple texts, compared with a control group which engaged in business-as-usual school curriculum. microgenetic data revealed a progressive development of experimental participants’ integration skill throughout their engagement in the argumentive discourse activity. the findings have important educational implications. they show that learning to argue by engaging in dialogic argumentation is a promising pathway for supporting the ability to integrate multiple source perspectives after reading controversial multiple texts. keywords: argumentation; multiple texts; integration; argument skill; multiple source perspectives iordanou & fotiou 23 | f l r 1. introduction individuals are often called to take positions and make decisions on issues of individual interest, such as vaccination, or issues of societal interest, such as immigration and climate change, for which they seek consultation from external sources to form a belief. depending on others, especially experts, in forming beliefs and making decisions is almost inevitable in our era of complexity and hyperspecialization (duncan, chinn, & barzilai, 2018; kienhues, jucks, & bromme, 2020; rabb et al., 2019). the replacement of a single textbook, or newspaper that used to serve as the single main source for learning and being up-to-date by a plethora of sources on the internet which provide, in many cases, different perspectives about an issue makes this task more challenging. therefore, the ability to handle effectively multiple sources appears imperative in our digital age where individuals have access to multiple sources at the click of a button. taking into consideration different perspectives and the available data is an important skill for making the right decision at a particular time, and at the societal level, for avoiding extremism and supporting democracy. yet, individuals struggle to integrate information from alternative perspectives, constructing instead one-sided representations (richter & maier 2017; tarchi & mason, 2020). recent results provided by the organisation for economic co-operation and development (oecd)’s programme for international student assessment (pisa), measuring 15-year-olds’ ability for reading, mathematics and science, revealed that the majority of students (73.7%) can identify the main idea of a single text – reaching level 2 of reading proficiency – but only a mere 1.2% can integrate multiple perspectives from multiple texts, that is expected by skilled readers, reaching the most advanced level (level 6) of reading proficiency (oecd, 2023). an emerging line of interdisciplinary research attempts to understand how individuals make sense of information from varying sources (van meteris et al., 2020). for instance, a fundamental skill when reading multiple texts is to understand the authors’ way of thinking and representing a particular issue, namely identification of source perspective (barzilai & weinstock, 2020). following barzilai and weinstock, we define source perspective, as “the perspective of the authors or organizations who create and communicate information using texts” (p. 5) and source perspective comprehension as “readers’ understanding of authors’ particular ways of thinking and knowing and how these inform authors’ interpretation and representation of the issue at hand” (p. 3). despite the importance of this skill for understanding multiple texts and using effectively the information represented in them, as well as individuals’ limitations in their ability to identify source perspective, our understanding of how to develop it remains limited (wiley et al., 2018). this study focuses on people’s ability to integrate source perspectives from multiple texts into reasoning, with a particular emphasis on how to develop this ability. we examine whether engagement in dialogic argumentive reasoning supports integration of multiple source perspectives in argumentive reasoning. although there is empirical evidence showing that engagement in dialogic argumentation can support the development of two-sided reasoning, that is, taking into consideration opposing views on a topic (felton & herko, 2004; kuhn & crowell, 2011), to the best of our knowledge, there is no evidence showing whether engagement in dialogic reasoning can support integration of source perspectives. the latter refers to the identification of text authors’ particular perspective on an issue when reading a text and incorporation of different authors’ perspectives in argumentive writing, after reading multiple texts on a particular topic. in this work, we examine whether gains acquired after engagement in dialogic argumentation transfer to individuals’ ability to integrate different source perspectives in a writing task after reading multiple texts on a particular issue. https://link.springer.com/article/10.1007/s10212-019-00426-8#ref-cr55 iordanou & fotiou 24 | f l r 2. identifying and integrating multiple source perspectives in multiple text comprehension the ability to identify different views presented in different texts – that is, source perspective – is a fundamental ability for multiple text comprehension. in fact, identification and integration of source perspective constitutes an integral component of theoretical models on multiple-text comprehension. for example, integration of information is one of the five essential steps involved in comprehension of multiple texts in the md-trace (multiple-document task-based relevance assessment and content extraction) model (britt & rouet, 2012). the other steps involve creating a task model with information about the goals of reading and how to achieve this, accessing the need for further information, engaging with the completion of the task product, and evaluating the degree of completion of the task. similarly, integration of information is part of the execution stage of the 3-stage model of the integrated framework of multiple texts (list & alexander, 2019). after the preparation stage, when the reader conceptualizes the objectives of the task, and before the production stage where the reader produces an external product, such as a written essay, is the execution stage. in the latter the reader engages in several cognitive and metacognitive strategies while processing the documents, such as identification, representation, and synthesis. to form multiple source perspectives, one needs to have the ability to develop a metarepresentation of each text, where the representation of a particular phenomenon is seen as the author’s representation involving the particular way that the author interprets and represents the phenomenon (barzilai & weinstock, 2020), rather than an objective reflection of how things are in the external world. the ability to infer and consider the views of an author of an academic text is connected with one’s ability to engage with the text as well as with academic engagement and performance(kim et al., 2018). yet, empirical studies show that individuals of different ages struggle with identification and integration of source perspective. almost half of the individuals examined in different studies and of different age groups were not able to identify contrastive views when reading different sources on a particular topic (barzilai, tzadok, & eshet-alkalai, 2015; coiro, coscarelli, maykel, & forzani, 2015; hobbs & frost, 2003) or when writing integrative reports (list et al., 2019; mateos & solé, 2009). after reading multiple-texts on a particular issue, and if not explicitly prompted to take multiple texts into consideration, individuals tend to rely on a single text when engaged in a writing task (monte-sano & de la paz, 2012; iordanou et al., 2020; stahl et al., 1996). even when prompting does take place, elementary school students do not seem to be able to identify position differences between the texts (paul, stadtler, & bromme, 2019). intervention studies aiming to promote individuals’ ability to identify and integrate multiple source perspectives have shown mixed results (de la paz et al., 2017; monte-sano, 2011). for example, barzilai and ka’adan (2017) reported that although scaffolding integration improved high-school students’ integration performance (effect size: ηp2= .08), they still found it difficult to construct fully justified dual-position arguments and address all differences between accounts. review studies on multiple documents acknowledge the need for further research examining how (i.e., with which activities) to support effective engagement with multiple documents (wiley et al., 2018). this work examines whether engagement in dialogic activity is a promising way for supporting identification and integration of multiple source perspectives after reading multiple texts. iordanou & fotiou 25 | f l r 3. dialogic argumentation and multiple perspectives 3.1 dialogic argumentation and integration of multiple source perspectives from multiple texts: theoretical underpinnings the theoretical rationale underpinning the role of dialogic argumentation in promoting students’ ability to identify different perspectives when reading multiple texts and integrate them into reasoning, derives from the proposed connection between construction and evaluation of arguments, which constitute facets of argumentive reasoning (iordanou, kendeou, & beker, 2016) and the conception that reasoning skills emerge and are developed first on the social plane before become internalized, which, in turn, derives from the sociocognitive and sociocultural theories. starting from the latter, both piaget (1928) and vygotsky (1978) conceived social interaction as the primary means for supporting the development of individual reasoning. engagement in dialogic argumentation in the social sphere supports the development of important meta-level insights of the norms of argumentation, but also of the nature of knowledge (iordanou, 2022; kuhn et al., 2013; rapanta & felton, 2022; chinn et al., 2011). one important epistemic understanding that develops through dialogic argumentation is that there is no single self-evident truth and that multiple interpretations may exist of the same phenomenon as the human mind plays an active role in ascribing meaning to the world (iordanou, 2016a; kuhn et al., 2008). this epistemic understanding of other people’s thinking as represented in multiple accounts is fundamental for integration of multiple source perspectives from multiple texts (kuhn, 2020). based on the theoretical proposal that argument construction, which is evident during dialogic argumentation, and argument evaluation, which is evident during text comprehension, are two different facets of the same argumentive reasoning and are both supported by the same core skills (iordanou, kendeou, & beker, 2016), we would expect that gains developed during dialogic argumentation at the social plane would become internalized and manifest at the individual level in other instances which require argumentive reasoning, such as when reading arguments in the context of one or multiple texts. 3.2 dialogic argumentation interventions the ability to take into consideration multiple, even contradicting views, when one reasons is considered fundamental for skilled reasoning in the reasoning literature (walton, 1999). according to graff (2003), the inclusion of multiple perspectives is what actually gives status and value to an argument itself. a comprehensive line of research on argumentation has offered empirical evidence showing that reasoning skills are amenable to improvement when received direct attention. in particular, engagement in dialogic argumentation appears to be a fruitful way to promote two-sided reasoning (see iordanou & rapanta, 2021, for a review of studies) and reduce my-side bias (felton et al., 2015). for example, students who had extensive practice in dialogic argumentation showed gains in using counterarguments and acknowledging opposing views which transferred from the social to the individual plane when writing an essay on a novel topic (iordanou & kuhn, 2020; kuhn & crowell, 2011; shi et al., 2019). notably, the strategic gains of engagement in dialogic argumentation transferred to new topics within a particular knowledge domain – science (iordanou & constantinou, 2015; iordanou & kuhn, 2020) and social (kuhn et al., 2008) ‒ as well as across knowledge domains (iordanou, 2010). this shows that some form of meta-level understanding develops which is then transferable to a context different from the one that has originally been developed. studies using the microgenetic method, aiming to get some insight into the mechanism behind development of argument skills, found that a meta-strategic understanding of the norms of argumentation is developing and supports development of argument skill (iordanou & constantinou, 2015; kuhn et al., 2008; shi, 2020). besides meta-strategic gains, epistemological gains on the nature of knowledge and process of knowing have also been observed to be the result of extensive engagement in dialogic argumentation (iordanou, 2010, 2016b, 2022; shi, 2020; zavala & kuhn, 2017). yet, the transfer of gains in reasoning after engagement in an argument-based intervention on reading multiple texts and integrating multiple perspectives represented in different texts has not been explored in the argumentation literature. iordanou & fotiou 26 | f l r the present study in the present work, we extend the previous line of research by examining whether engagement in systematic dialogic argumentation can support one’s ability to integrate multiple source perspectives in argumentive reasoning, after reading contrasting multiple texts on a particular topic. our research question was the following: does engagement in dialogic argumentation support the ability to integrate multiple source perspectives from multiple texts? based on the findings of previous research showing that a meta-level understanding of the norms of argumentation and the epistemic nature of knowledge — acknowledging the role of human interpretation and therefore of multiple perspectives on an issue (iordanou, 2022) — develops when engaged in dialogic argumentation, we hypothesize that engagement in dialogic argumentation can be a fruitful means for promoting identification of multiple perspectives in multiple text comprehension and integration of those perspectives when writing an argumentive essay. we examine whether engaging in dialogic argumentation with peers who hold opposing views on a topic, can help individuals to develop the ability to identify opposing views when reading multiple texts on an issue, and integrate those views in a written argumentive task. based on evidence from previous research showing that engagement in dialogic argumentation — in person or through the computer — with peers holding opposing views, supports the development of two-sided thinking (kuhn & udell, 2003; kuhn et al., 2008), we hypothesize that individuals will transfer this ability from writing a two-sided report on the intervention topic to writing a two-sided report on a novel, non-intervention, topic, after reading different texts depicting different perspectives on an issue. previous research has also shown that asking individuals to write an argument is more effective for integrating views from multiple sources than asking them to write a summary (bigot & rouet, 2007; maier & richter, 2016; stadtler et al., 2014), providing further evidence of the potential of engagement in argumentive activities in promoting integration of multiple source perspectives from multiple texts. previous research on reading comprehension that has examined the effectiveness of dialogbased pedagogical practices for promoting text comprehension focused on the effects of text-based discussion (see murphy et al. 2009 for a review) on single text comprehension. the novelty of the present study lies in both the medium used and the dependent variable examined. firstly, we investigated the power of engagement in the activity of dialogic discussion on a controversial topic, independently of a particular text. secondly, we examined the effect of engagement in dialogic activity not on a singletext comprehension, but on multiple-texts comprehension. to the best of our knowledge, this is the first time that dialogic argumentation is examined as a tool for promoting multiple text comprehension, as evident in individuals’ ability to incorporate the multiple perspectives presented in multiple texts in writing an essay. in the present study we asked our participants to engage in dialogic argumentation for 14 sessions before we asked them to write an argumentive essay after reading two different texts on a novel topic, each of which presented a different view on the topic. participants conducted the dialogs electronically via instant-messaging software using tablets. this method, which has extensively been used in previous work, offers the advantage of providing an immediately available record of the discourse that participants can use to reflect on. we used authentic sources in line with recent recommendations to use authentic learning environments and authentic information sources (chinn, barzilai & duncan, 2021). we are interested in examining whether engagement in dialogic argumentative-based intervention can support the development of skills needed in real-life, that is identification of authors’ perspectives when reading authentic texts, found on the web. another group of students, which was assessed at the same time points as experimental condition students, but engaged in business-as-usual school activities, served as a control condition. participants’ integration of multiple source perspectives was assessed at initial and final assessment using an open-ended question instrument (barzilai & ka’adan, 2017; bråten et al., 2014). participants’ integration performance was assessed in two novel topics, one in the same domain as the intervention topic — social science domain — and another one in a different domain from the intervention topic – physical science — to examine far iordanou & fotiou 27 | f l r transfer. furthermore, integration of multiple source perspectives was examined, using the microgenetic method, throughout the intervention — coding all the experimental condition’s dialogs — to identify any possible pattern of development that would enable us to get some insights into the mechanism that supported the development of integration of multiple source perspectives. 4. methodology 4.1 participants the participants of this study were 64 sixth graders (11to 12-year-olds) from four classes. they were recruited from three public primary schools in cyprus. the study was conducted in accordance with the declaration of helsinki and with ethics approval from the cyprus national bioethics committee and the cyprus ministry of education, sport and youth written parental consent was obtained for each child. in addition, all children were informed orally about the study. two classes, from two different schools, served as the experimental condition (34 students; 16 female), and two classes from a third school served as the control condition (30 students; 15 female). the size of the recruited sample exceeded the required sample size of 40, as determined by an a-priori power analysis for repeated measures anova, within-between interaction, (gpower, version 3.1.9.7). power was set to 0.80, α-error to 0.05, and the assumed effect size to η2p = .05, as recent studies documented medium effects of integration performance ‒ interaction between time and group (barzilai & ka’adan, 2016). the participants were from middle-class families and with primarily an average academic achievement, typical of those schools. 4.2 measures during the initial and final assessment phase, the experimental and control condition participants’ argument skill, integration of multiple source perspectives and prior knowledge were assessed at about the same time in the middle of the school year. the initial and final assessment phase was identical for the experimental and control condition (i.e. instructions, materials, time). participants’ argument skills and multiple source perspectives were assessed on non-intervention topics, aiming to assess transfer of intervention gains. the final assessment took place two days after the completion of the intervention. this phase took the same form as the initial assessment phase. all participants were given exactly the same instruments. the only difference with the initial assessment phase was that they worked on a different topic from the one they worked initially for assessing integration performance. for example, if a participant completed an integration performance assessment on sun exposure (science topic) during the initial assessment phase, they would work on the cell phone topic (science topic) for the final assessment phase. the same held for the social topics. the participants in the control condition engaged in the same assessment procedure as those in the experimental condition and at the same time of the year, for both initial and final assessments. 4.2.1 argument skill participants’ argument skill was assessed in writing (iordanou et al., 2019; kuhn et al., 2008) on a non-intervention topic addressing the issue of whether an elderly person’s family or the government should be responsible for the care of elderly people (kuhn, 2017). this measure was used as preand post-test measure. the participants were instructed to write a letter they would send to a local newspaper and asked that they be as convincing as possible. they were provided with nine pieces of evidence, supporting equally both positions, in the form of questions and answers that they could use to support their argument if they wished. an example of a piece of evidence provided was “how much does it cost to pay for the care of an elderly person in a long-term care facility? the average cost for one year at a private long-term care facility in the us is around $50,000. such facilities are not always available, iordanou & fotiou 28 | f l r especially in less developed countries.” they were given as much time as they needed to complete their letters. the letters on average were 62 (sd = 51.35) words long. 4.2.2 coding of participants’ argument skill at initial and final assessment one of the authors and a research assistant, blind to condition and time, segmented and coded the letters that students prepared at initial and final assessment on the transfer topic. the letters were segmented into idea units which consist of a claim and supporting justification. only segments that included a claim and evidence served as the data base for further analyses. if there was no connection between the cited evidence and the claim, the unit was coded as non-functional. if the unit included a claim and a supporting (or weakening) piece of evidence connected to it, it was coded as a functional unit and it was further coded according to the type of function served (m+, m-, o+ and o-), employing the coding system used in previous work to assess students’ argument skill (iordanou et al., 2019; kuhn et al., 2016). inter-rater reliability on segmenting and coding was achieved on a subset of 30% of units, with 90% and 88% agreement, respectively. the research assistant proceeded with segmenting and coding the remaining essays, again blind to condition and time. 4.2.3 integration of multiple source perspectives to assess participants’ integration performance in each domain (social and science) two texts were used. the texts were designed to be similar in the content and structure and were administered, in counter-balanced order, in the initial and final assessment. the two social texts were on 1) bilingualism and its connection to cognitive abilities and 2) grades and their connection to learning. the science texts were on 1) sun exposure and health and 2) cell phones and health. the science texts were adapted from two texts originally published in norwegian newspapers and journals, used by bråten et al. (2013) and bråten et al. (2014). the original texts were in norwegian and were authentic sources from norwegian newspapers and journals. the texts were provided to us in english by the researchers and were translated to greek by us. the texts’ difficulty level was adapted to be suitable for sixth graders. in relation to the cell phone and health topic, the first text was a 527-word text published in a science magazine. it mainly reported on an unpublished review article by an academic and brain surgeon who argues that cell phone use and brain tumours are linked and that radiation from wireless computer networks — which is similar to cell phone radiation —is harmful to our health. the second text was a 536-word text published in a newspaper which argued that those who claim that cell phone use can cause cancer exaggerate (bråten et al., 2014). in relation to the sun exposure and health topic, the first text was a 406-word text published in an online research magazine by a group of educational institutions and showed evidence that exposure to sun can cause skin cancer and we should not sunbathe for obtaining vitamin d; instead, we can take supplements (bråten et al., 2013). the second text was a 410-word text taken from a norwegian conservative daily. it reported on a large-scale longitudinal study conducted in the us which showed that vitamin d can prevent the occurrence of cancer. thus, since sun exposure is the natural means through which one gets this vitamin, the authors of the text recommend a 30-minute daily sun exposure (bråten et al., 2013). the social texts were developed by the authors by adapting and translating authentic texts found on blogs written by academics and professionals after getting all authors’ written consent. the first text on bilingualism and its relation to cognition was a 404-word text adapted and translated from an article written by bialystok (2017), arguing for the benefits of bilingualism in relation to bilinguals’ cognitive skills. the second text was a 419-word text adapted and translated from an article written by chatham (2007), arguing for the possibility that the reported benefits of bilingualism apply to a specific part of the population (those of a higher socioeconomic status). it also claimed a relationship between a higher socioeconomic status and children’s cognitive abilities and stated that at least one study shows no advantage of bilingual children over monolingual ones. regarding the topic of grades and their relation to learning, the first text was a 425-word text adapted from an article written by travis (2017), arguing that grades help learning only when the standards according to which the grades are given are known and are clear to students. the second text was a 417-word text adapted from an article by kohn (2010), iordanou & fotiou 29 | f l r arguing for a world without grades, using as supporting evidence an example of a school that stopped giving grades to their students, who then improved their learning and academic performance. participants’ integration of multiple source perspectives, thereafter, referred to as integration performance, was assessed using an approach that has been developed by rukavina and daneman (1996), bråten et al. (2013) and barzilai and ka’adan (2016). the participants were asked to answer three open-ended questions. the first two questions indirectly required participants to integrate ideas from multiple information sources, assessing if participants integrate information from multiple sources without being prompted. the first question asked participants to explain the relation between key components examined in the texts, e.g., cell phones and health. the second question invited them to state their opinion on the controversial topic in question, e.g., their opinion on whether cell phones harm people’s health. finally, the third question directly requested to compare accounts, and state the differences between the two opposing views over the controversy in question. an example: “there are different views on the relationship between using a cell phone and health. describe important differences between these views.” in other words, the third question directly asked the participants to integrate perspectives from different sources (rukavina & daneman 1996, cited in barzilai & ka’adan, 2016). 4.2.4 coding of participants’ responses in integration of multiple source perspectives instrument at initial and final assessment participants’ responses were coded based on an integration coding scheme employed by barzilai and ka'adan (2016) and bråten and his colleagues (bråten et al., 2013; ferguson et al., 2013), which assesses the extent to which participants present and justify the contradictory positions represented in two texts and explicitly make connections between the two positions (see appendix). the coding scheme rates the extent to which participants presented and justified multiple positions they found in the two texts they had at their disposal and the extent to which they connected those positions. participants could receive up to six points for presenting and justifying the different positions put forward in the texts they read, with end points “0” when no position was presented regarding the inquiry question and “6” when two positions were presented with supporting reasons or explanations for both positions. for the second questions, that asked them to state their opinion on the controversial topic in question, participants’ responses were scored based on whether they included alternative explanations, involving any explanations not necessarily the ones presented in the texts. in addition, they received up to two points for connecting those positions, a total of 8 points per question. in other words, the highest score one could get is eight points per question. the same codes were used for assessing the responses to all three questions. the final integration performance score was based on the sum of the scores on all three questions. two coders – the authors ‒ coded 40% of the data, blind to condition and time, with 88% agreement. the rest of the data were coded by one of the two coders, again blind to condition and time. 4.2.5 coding of participants’ integrative events in the electronic dialogues during the intervention all the experimental condition’s transcripts of the electronic dialogs that took place during the intervention were coded for evidence of integration. all dialogs were first segmented into the minimum idea units that served a specific function in the conversational exchange, such as expressing a simple agreement or providing a counterargument. each idea unit was classified as to whether it included evidence (evidence-based idea unit) or not. evidence-based idea units were further coded as to whether they integrated evidence from multiple sources or not. if they included evidence only from a single source, the idea unit was coded as “single-source.” an example of a “single-source” idea unit coming from personal knowledge is “ we believe that refugees should be accepted based on how difficult the conditions are in their country, because christ taught us to love and help our fellow man regardless of our interests.” if they integrated evidence from different sources that we have provided in the form of q&a evidence cards or from a combination of sources in the q&a card and their personal knowledge, the idea unit was coded as “multiple-source”. evidence that was not included in the sources that we provided was coded as evidence coming from a single source – personal knowledge. an example of an iordanou & fotiou 30 | f l r idea unit that was coded as “multiple-source” is the following “but fraudsters are not only the refugees, but also the locals, for example in the netherlands there were 30 suspects not even confirmed among the thousands (of refugees that) had come. even in france, where there were more refugees acclimatized than the natives, they lived much worse than the french. like yiannis agianis (jean valjean), who lived in unfavorable living conditions.” in that example, the student combined three pieces of evidence. the first two pieces were based on information provided in two different q&a cards — one referring to the netherlands that has identified 30 suspected war criminals among thousands of refugees who entered the country in 2015 and the other referring to a published study’s findings showing that the share of immigrants in the population has no significant impact on crime rates once immigrants' economic circumstances are controlled for in france —, while the third one was based on student’s personal knowledge from victor hugo’s novel “les misérables.” two coders – the authors ‒ blind to time, coded 40% of the data, with interrater agreement, 92%. disagreements were resolved through discussion. the rest of the data were coded by one of the coders. 4.2.6 prior knowledge test before the administration of the texts and the individual argument skill instrument, participants were given multiple-choice questions on each topic to assess their prior knowledge. participants' prior knowledge was used to assess the equivalence of the four topics, because they received two of the texts, one from social domain and one from science domain, at initial assessment and the other two at the final assessment, in a counterbalanced order. all prior knowledge tests, except one, consisted of ten multiplechoice questions—only the test on grades tests consisted of six questions. participants received one point for every right answer. the prior knowledge questions were designed by the researchers except those on the science topics which were formed based on the questions developed by bråten et al. (2013) and bråten et al. (2014) to assess prior knowledge. 4.3 the intervention phase the participants in the experimental condition engaged in a dialogue-based argument curriculum over fourteen 80-minute sessions on the topic of immigration. these took place approximately twice per week over a period of three months. the participants in the control condition were taught about the same topic as part of the business-as-usual school curriculum which the topic was part of. in the school curriculum consideration of multiplicity of perspectives is not a standard practice for students of this age group. the experimental condition participants engaged in a series of electronic dialogs with their peers following the curriculum developed by kuhn et al. (2008) and employed, thereafter, in many studies aiming to promote students’ argument skill (e.g. iordanou et al., 2019; iordanou & kuhn, 2020; shi, 2020). the participants in the experimental condition were introduced to the intervention topic by reading two texts which presented two different views on the criteria of accepting immigrants in one country. an introduction to the two texts posed the following question taken from kuhn (2017, p. 7): “should a nation allow people from other countries to come live in their country based on what they can contribute or how bad life is where they come from?” the two texts were developed by the researchers based on information found on valid sources regarding immigration. one text supported the view that immigrants should be accepted based on how bad life is in their home country; the other text supported that immigrants should be accepted based on what they contribute to the arriving country. the texts were equal in size (around 220 words). the participants were asked to take a position. based on their position, two groups were formed: one group was in favor of the view that immigrants should be accepted based on how bad life is in their home country (need group) and the other was in favour of the position that immigrants should be accepted based on what they can contribute to the arriving country (contribution group). the two groups that were formed were approximately equal. undecided participants were allocated to the group with the fewer participants so as to have close to equal number of students in each group. figure 1 shows the experimental design of the study. iordanou & fotiou 31 | f l r 4.3.1 preparation of arguments in the first session, the participants formed four small groups of 4-6 students sharing the same position and were asked to generate reasons supporting their position and note them on cards. then they were requested to rank these reasons with respect to their strength. adult coaches – the authors and the teacher – acted as facilitators in both classes in the experimental condition by encouraging participation of all group members. the cards prepared in this session remained available to participants during the argument chat sessions that followed. 4.3.2 electronic dialogs the participants in each group (“need group” and “contribution group”) were divided into sameside pairs and the members of each pair remained the same throughout these sessions. the participants engaged in eight electronic dialogs with a sequence of peers from the other group (sessions 2-9), holding an opposing position. these electronic dialogs with peers have the advantage of providing students with the opportunity to get extensive experience in engagement in dialog, unlike classroom-based discussions when many students engaged in the same dialog. dialogues were conducted on tablets, provided by the researchers, via an instant messaging software, in students’ classroom. the transcripts of the dialogues were saved and used later for analysis. the participants were instructed to convince the opposing pair about their position. each pair was further instructed to collaborate with their partner to decide what they wished to say to the opposing pair and, once they were in agreement, to send their response to the opposing pair. two adult coaches (one of the authors and the teacher who received training) provided help with technical issues and reminded pairs to collaborate in responding to what the opposing team was saying. during these sessions, the participants had at their disposal pieces of information and evidence that they could use if they wished to. all evidence provided to them was in the form of question and answer, following the recommendation of iordanou et al. (2019) who showed that this is a more effective way to promote evidence use in argumentation compared to providing information in the context of a traditional text. some questions were developed based on questions in kuhn (2017 pp. 179–182) while others were developed by the authors. the answers provided were based on reliable sources with the source provided under each answer. all in all, eight different sets of three questions-answers were formed. participants received one set in each session which remained available to them in the subsequent sessions. students received evidence supporting their own view (m+, n=7), evidence weakening their own view (m-; n=5), evidence supporting the opposing view (o+; n=7) and evidence weakening the opposing view (o-; n=5). in addition, participants in the last 3 sessions were asked to reflect on the transcript of their dialogue using reflection sheets. one reflection sheet asked participants to reflect on the effectiveness of a counterargument they offered to the opposing side’s argument while the other encouraged them to reflect on a rebuttal they offered to opponents’ counterargument, and in both cases to consider possible improvements. 4.3.3 preparation for the ‘showdown’ in sessions 10-11, participants prepared for the showdown for which they knew they would compete and there would be a winning team. the class was divided into four same-side preparation teams. the teams were given all the reflection sheets that their members had already prepared along with the evidence provided to them during the chat sessions and a printed copy of the transcripts of the dialogs they had. they were then asked first to reflect on all of these and prepare two different kinds of sets of cards. the first set consisted of two cards: other’s argument – counterargument and the other iordanou & fotiou 32 | f l r set consisted of three cards: own argument-counterargument provided by the other side-own rebuttal. each part of the sequence was noted in a different coloured card, providing a visual representation of the sequence. all groups were assisted by three adult coaches in both schools (one teacher and the two researchers). 4.3.4 “showdown” and feedback in the next session (session 12), students had an electronic showdown. working toward the social objective of the showdown, previous work has shown (kuhn et al., 2008) to have motivating and focusing effects on students. all participants supporting the same side of the topic were placed in one room with the students supporting the other side being in another room. then, the participants on each side of the topic were divided into two teams (team a and team b). each team was given 20 minutes to debate on the topic with their corresponding opposing team that was in another room. the two sides communicated through the computer and their dialogue was projected onto a whiteboard. all members collaborated to reach an agreement on the text to be sent to the opposing side. during the first half of the showdown, the a team members debated while the b team members were watching the debate and offered suggestions in writing to the a team members if they wished. at half-time, teams switched roles and the b team members continued the debate. the showdown thus consisted of a single 40-minute electronic dialogue between the two opposing sides. following the electronic showdown (session 13), students received feedback and a winning team was declared. the electronic dialogue produced in the showdown was presented to them in an argument map prepared by the researchers. different columns appeared for each team, with their contributions arranged in order of occurrence from top to bottom. all statements were represented and connected by lines to show their interrelation. different colours were used to label statements as effective, ineffective, or neutral argumentive moves. points were assigned for each counter-argument the students produced and for each piece of evidence they used to support their argument in order to declare the winners. in the last session (session 14), a live, face-to-face showdown was pursued, which participants’ parents were invited to watch, using the same rules described for the electronic showdown above. 4.3.5. fidelity to ensure fidelity of treatment, the first author prepared a detailed intervention protocol, including lesson plans and assessment guidelines, that was provided to the second author and the teachers of the experimental condition. the second author coordinated the implementation of the intervention in both classes in the experimental condition and was present in all sessions. the classroom teachers acted as facilitators, following closely the guidelines set out in the intervention protocol. the first author attended about half of the sessions and had regular meetings with the teachers and the second author before and after each session. in addition, the sessions were video recorded to monitor treatment fidelity. an independent researcher and the first author, who coded the videos, confirmed that the sessions adhered to the intervention protocol. the assessment of the control condition students was pursued by another research assistant, with experience in administering assessment instruments, following the same assessment guidelines described in the protocol and after communication with the first author. iordanou & fotiou 33 | f l r 5 results 5.1 argument skill at initial and final assessment to examine whether there were any statistically significant differences between the experimental and control groups at the outset of the study so as to ensure that participants were equivalent, we compared experimental and control condition participants’ skill in using evidence to weaken others’ position, an advanced argument skill. a kruskal-wallis h test showed that there was no statistically significant difference in argument skill – advanced skill of using evidence to weaken others’ position, o-, ‒ between the experimental and control groups at initial assessment, χ2(1)=.744, p=.388. a glmm using the poisson distribution, examining condition differences over time in argument skill, showed a difference between groups in weaken-other usage in the transfer topic, f(1, 138)=8.032, p=.005. the interaction between group and time was significant, f(1, 138) = 12.506, p=.001. the fixed effect of time was significant, f(1, 138)=15.935, p< .001, as well as the fixed effect of group, f(1, 138) = 8.032, p=.005. students in the experimental condition showed an increase in the number of weaken-other units, from 0.083 (sd=0.050), 95% ci [-.15, .182] to 1.028 (sd=.212), 95% ci [.608, 1.447], while students in the control condition showed a more limited increase, from 0.171 (sd=0.72), 95% ci [.028, .314], to 0.229 (sd=0.101), 95% ci [.028, .429]. 5.2 integration performance at initial and final assessment first, we examined whether there were any statistically significant differences between the experimental and control groups at the outset of the study to ensure that students in the experimental and control conditions performed equivalently. data were also examined for outliers and these were ruled out. a manova comparing conditions, with the integration performance score in the social domain and the science domain as dependent variables, failed to achieve statistically significant difference, f(2, 61) = 0.025, p = .975; wilk's λ = 0.999, ηp2 = .001. because the bilingualism topic and grades topic for the social domain, and the cell phone use topic and sun exposure topic for the science domain, were counterbalanced in the pre-test and post-test, we examined participants’ prior topic knowledge about these four topics in order to assess their equivalence. a manova comparing conditions with the four topic knowledge variables as dependent variables failed to achieve statistically significant difference, f(4, 59) = 1.092, p = .369; wilk's λ = 0.931, ηp2 = .069. experimental and control condition students showed comparable prior knowledge in all the four topics, bilingualism (m=5.529, sd=1.942, and m=5.800, sd=1.648), grades (m=3.088, sd=1.264 and m=3.000, sd=1.232), cell phone use (m=4.441, sd=1.691 and m=3.933, sd=1.530) and sun exposure (m=4.618, sd=1.723 and m=3.967, sd=1.691). no statistically significant difference in prior knowledge scores was observed among the 4 classes which took part in the study, either, f(12, 151) = 1.348, p = .197; wilk's λ = 0.764, ηp2 = .086. a 2 (condition) x 2 (time) repeated-measures analysis of variance (anova) comparing the two conditions was used to assess whether the conditions had differential effects on integration skills. on the social topic, a significant time x condition interaction was observed, f(1, 62)=6.949, p=.011, ηp2 = .101. experimental condition students, as seen in figure 1, doubled their integration score from initial (m=3.882, sd=3.444) to final assessment (m=6.647, sd=4.081), while control condition students showed no significant difference from initial (m=3.733, sd=2.258) to final assessment (m=3.666, sd=2.795). on the science topic, a significant time x condition interaction was also observed, f(1, 62)=10.596, p=.002, ηp2=.146. experimental group participants showed a significant increase in their integration score, from 5.588 (sd=3.276) to 7.971 (sd=4.448) (see figure 2). no significant difference was observed in control group participants, from initial (m=5.633; sd=2.220), to final assessment (m=4.733; sd=2.664). iordanou & fotiou 34 | f l r figure 1. experimental and control condition students’ integration of multiple source perspectives on the social domain, from initial to final assessment figure 2. experimental and control condition students’ integration of multiple source perspectives on the science domain, from initial to final assessment iordanou & fotiou 35 | f l r 5.3 integration during the intervention the microgenetic method was employed to examine the process of change during the intervention. the number of idea units per session was different, ranging from m=7.300 (sd=2.830) to m=14.555 (sd=4.666), because of variations in time available due to school curriculum or technology restrictions (e.g. unexpected internet connectivity issues), therefore, we used percentages to examine possible differences during the intervention. evidence-based idea units ranged from m=4.900 (sd=1.969) to m=9.800 (sd=3.881). the percentage of evidence-based idea units that included integration of evidence from multiple sources, as opposed to using evidence from a single source is depicted in figure 3. as can be seen in figure 3, there was an increasing pattern of integrating multiple sources in students’ arguments from dialog session 1 (m=6.455, sd=9.759) to dialog session 8 (m=17.010, sd=10.777), showing that the integration skill was slowly developing during engagement in argumentation in the context of the intervention. figure 3. percentage of evidence-based idea units which included integration of evidence from multiple sources, throughout the intervention 6 discussion we examined the effectiveness of engagement in dialogic argumentation on its ability to promote integration of multiple source perspectives from multiple texts in argumentive writing. results revealed that participants who engaged in a dialog-based argumentive intervention improved their integration of multiple source perspectives in argumentive writing, whereas participants who engaged in their business-as-usual curriculum did not show any improvement. our findings are consistent with wiley and voss (1999) who found that engagement in argument construction supports better integration of information when writing arguments. what accounts for the experimental group’s gains? what needs to be considered when seeking an explanation for the condition effects is the fact that gains in integration of multiple source perspectives in argumentive writing were confined to those in the experimental condition who engaged in argumentive discussions with peers who hold and supported with arguments an opposing position on 0 5 10 15 20 25 1 2 3 4 5 6 7 8 iordanou & fotiou 36 | f l r the main topic. experimental condition participants showed better multiple source perspective integration compared to both their initial assessment performance and the performance of the control group, who was assessed at the same time points as the experimental group but attended their regular school curriculum. noteworthily, gains in integration of multiple source perspectives in argumentive writing were also transferable. experimental group participants not only showed gains in integration of multiple source perspectives in argumentive writing in a new topic in the same domain that they had their intervention on – the social science domain ‒ they also exhibited far transfer of their integration of multiple source perspectives gains to a different, non-intervention, domain, namely the physical science domain. why did this transfer of integration of multiple source perspectives occur and why did the experimental group show an advantage in this regard? the explanation of what accounts for the gains observed is not obvious, given that participants did not receive any direct instruction on multiple source perspectives. we propose that engagement in argumentive dialog supported experimental group participants to learn something they were then able to apply to a new task, a new topic and a different domain – something such as an understanding that there are alternative perspectives on an issue. although this understanding of recognizing different interpretations of an issue seems simple, it is not a developmental achievement that we should take for granted (iordanou, 2016a; lalonde & chandler, 2002). yet, recognizing alternative positions on an issue is fundamental for multiple-text comprehension (britt & rouet, 2012; kuhn, 2020; list & alexander, 2019). engaging in dialogic argumentation where a contrasting perspective is embodied in a “real” person, as did our experimental condition, may have supported this understanding. when ideas are personally represented, receivers’ thinking about the issue benefit more, probably by emphasizing that there indeed exists a flesh-and-blood other who supports such views (iordanou & kuhn, 2020; mill, 1859/1996). engagement in dialogic argumentation with individuals who hold different positions from one’s own on a particular issue, has the added benefit of providing a personal representation of views, in addition to offering exposure to divergent information which previous research shows impacts epistemic understanding (ferguson & bråten, 2013; ferguson et al., 2013; kienhues et al., 2011). dialogic argumentation provides the “interlocutor” which is missing and needs one to envision when reading multiple documents. indeed, previous research showed that identifying perspectives in informational texts is more challenging than identifying perspectives in everyday social interactions (jucks & bromme, 2011; kim et al., 2018). understanding alternative perspectives on an issue is an important epistemic achievement fundamental for appreciating the diversity and complexity of knowledge (barzilai & weinstock, 2020). the lack of direct measures for assessing students’ epistemic beliefs, which constitutes a limitation of the current study, does not enable us to draw definite conclusions regarding epistemic gains. future research needs to explore this possible interpretation further by measuring students’ epistemic beliefs. also, further work is warranted, using other modes of discussion and topics to examine the generalizability of the suggestive gains observed in this study. our microgenetic data show that the skill of integrating information from multiple sources developed gradually over time while individuals were engaged in dialogic argumentation, providing further evidence of the claim that engagement in dialogic argumentation supports integration skills. the microgenetic data show that during engagement in dialogic argumentation students exhibited a progression in combining evidence from multiple sources in their arguments. this progression is slow and extends over time, suggesting that sustained engagement in dialogic argumentation over time provides facilitative conditions for developing the skill of integrating multiple perspectives. the condition differences observed in the argumentive strategy of using counterarguments, which focuses directly on an other’s position in an effort to weaken it, also supports this interpretation. experimental group participants, but not control group participants, after their engagement in dialogic argumentation exhibited improvements in their ability to use evidence to weaken the other’s position, a finding which is consistent with previous empirical work (iordanou et al., 2019; kuhn & crowell, 2011; maywegpaus et al., 2016) and shows an implicit recognition of the value of paying attention to the other’s opposing position. the findings of the microgenetic study, showing gains in integration from multiple sources during the course of dialogic argumentation, which remained evident in argumentive writing after reading multiple texts in the absence of social support, are in line with the sociocognitive and iordanou & fotiou 37 | f l r sociocultural theories (piaget, 1928; vygotsky, 1978) according to which social interaction facilitates the development of reasoning skills which develop first on the social plane and then they become internalized. the unique contribution of the present work is in providing evidence of the power of engagement in a dialog-based argumentive intervention for promoting individuals’ ability to incorporate multiple perspectives in writing an essay after reading multiple texts on non-intervention topics. appreciating alternative perspectives might have supported students both during the process of reading multiple texts – given that the ability to identify the views presented in different texts as discrepant is fundamental for multiple text understanding (britt & rouet, 2012; kuhn, 2020; list & alexander, 2019) – and while they engaged in the argumentive writing task after reading controversial multiple texts. our findings have important educational implications. they extend previous findings which showed that dealing with conflicting information about an issue support an epistemic understanding of appreciating the imprecise nature of knowledge (kienhues, stadtler, & bromme, 2011). our work shows that in addition to dealing with conflicting information about an issue, engagement in purposeful dialogic argumentation with individuals who represent alternative positions about an issue (iordanou & kuhn, 2020), along with reflection on argumentation, facilitates students’ skill of integrating multiple source perspectives from multiple texts on an issue. students’ engagement in direct debate with one another, rather than using the teacher as the channel through which discourse flows, seems to facilitate the development of students’ argument skills. our microgenetic data suggest that this development is gradual, therefore providing students multiple opportunities in the school curriculum for sustained engagement and practice, over successive occasions, is another condition that needs to be taken into consideration in curriculum development and teaching practice. in a nutshell, the present work shows that engagement in dialogic argumentation is a promising pathway for supporting acknowledgment and integration of multiple source perspectives both when writing essays but also when engaged in argumentive writing after reading multiple controversial texts. keypoints engagement in dialogic argumentation supports the ability to integrate multiple source perspectives from multiple texts. microgenetic data revealed a progressive development of participants’ integration skill throughout their engagement in the argumentive discourse activity. a control group which engaged in business-as-usual school curriculum showed no improvement over time in integrating multiple source perspectives. engagement in dialogic argumentation supported development of participants’ argument skills. the gains observed in integrating multiple source perspectives showed far transfer, to a different domain. acknowledgments this work was funded by a grant from the research and innovation foundation (rif) awarded to kalypso iordanou (grant number: κουλτουρα/βρ-νε/0415/13). iordanou & fotiou 38 | f l r references barzilai, s., & weinstock, m. (2020). beyond trustworthiness: comprehending multiple source perspectives. in p. van meter, a. list, d. lombardi, & p. kendeou (eds.), handbook of learning from multiple representations and perspectives (pp. 123–140). routledge/taylor & francis group. https://doi.org/10.4324/9780429443961-11 barzilai, s., & ka’adan, i. (2017). learning to integrate divergent information sources: the interplay of epistemic cognition and epistemic metacognition. metacognition and learning, 12(2), 193-232. 10.1007/s11409-016-9165-7. barzilai, s., tzadok, e., & eshet-alkalai, y. (2015). sourcing while reading divergent expert accounts: pathways from views of knowing to written argumentation. instructional science, 43(6), 737766. https://doi.org/10.1007/s11251-015-9359-4 bialystok, e. (2017, april). bilingual children can focus better. child and family blog. https://www.childandfamilyblog.com/early-childhood-development/bilingual-children-focus/ bigot, l. l., & rouet, j. f. (2007). the impact of presentation format, task assignment, and prior knowledge on students' comprehension of multiple online documents. journal of literacy research, 39(4), 445-470. https://doi.org/10.1080/10862960701675317 bråten, i., ferguson, l. e., anmarkrud, ø., & strømsø, h. i. (2013). prediction of learning and comprehension when adolescents read multiple texts: the roles of word-level processing, strategic approach, and reading motivation. reading and writing, 26(3), 321348. https://doi.org/10.1007/s11145-012-9371-x bråten, i., ferguson, l. e., strømsø, h. i., anmarkrud, ø. 2014. students working with multiple conflicting documents on a scientific issue: relations between epistemic cognition while reading and sourcing and argumentation in essays. in british journal of educational psychology, 84(1), 58–85. https://doi.org/10.1111/bjep.12005 britt, m. a., & rouet, j. f. (2012). learning with multiple documents: component skills and their acquisition. enhancing the quality of learning: dispositions, instruction, and learning processes, 276-314. chatham, c. (2007). doubting the "bilingual cognitive advantage": just an effect of socio-economic status? science blogs: developing intelligence. https://scienceblogs.com/developingintelligence/2007/10/31/the-bilingual-advantage-may-ac chinn, c. a., barzilai, s., & duncan, r. g. (2021). education for a “post-truth” world: new directions for research and practice. educational researcher, 50(1), 51-60. https://doi.org/10.3102/0013189x20940 chinn, c. a., buckland, l. a., & samarapungavan, a. l. a. (2011). expanding the dimensions of epistemic cognition: arguments from philosophy and psychology. educational psychologist, 46(3), 141–167. https:// doi.org/10.1080/00461520.2011.587722 coiro, j., coscarelli, c., maykel, c., & forzani, e. (2015). investigating criteria that seventh graders use to evaluate the quality of online information. journal of adolescent & adult literacy, 59(3), 287-297. https://doi.org/10.1002/jaal.448 de la paz, s., monte‐sano, c., felton, m., croninger, r., jackson, c., & piantedosi, k. w. (2017). a historical writing apprenticeship for adolescents: integrating disciplinary learning with cognitive strategies. reading research quarterly, 52(1), 31-52. https://doi.org/10.1002/rrq.147 duncan, r. g., chinn, c. a., & barzilai, s. (2018). grasp of evidence: problematizing and expanding the next generation science standards’ conceptualization of evidence. journal of research in science teaching, 55(7), 907-937. https://doi.org/10.1002/tea.21468 faul, f., erdfelder, e., lang, a.-g., & buchner, a. (2007). g*power 3: a flexible statistical power analysis program for the social, behavioral, and biomedical sciences. behavior research methods, 39, 175-191. doi: 10.3758/bf03193146 ferguson, l. e., & bråten, i. (2013). student profiles of knowledge and epistemic beliefs: changes and relations to multiple-text comprehension. learning and instruction, 25, 49-61. https://doi.org/10.1016/j.learninstruc.2012.11.003 https://psycnet.apa.org/doi/10.4324/9780429443961-11 https://psycnet.apa.org/doi/10.1007/s11251-015-9359-4 https://www.childandfamilyblog.com/early-childhood-development/bilingual-children-focus/ https://doi.org/10.1080/10862960701675317 https://psycnet.apa.org/doi/10.1007/s11145-012-9371-x https://doi.org/10.1111/bjep.12005 https://scienceblogs.com/developingintelligence/2007/10/31/the-bilingual-advantage-may-ac https://doi.org/10.3102/0013189x20940683 https://doi.org/10.1002/jaal.448 https://psycnet.apa.org/doi/10.1002/rrq.147 https://doi.org/10.1002/tea.21468 https://doi.org/10.3758/bf03193146 https://psycnet.apa.org/doi/10.1016/j.learninstruc.2012.11.003 iordanou & fotiou 39 | f l r ferguson, l. e., bråten, i., strømsø, h. i., & anmarkrud, ø. (2013). epistemic beliefs and comprehension in the context of reading multiple documents: examining the role of conflict. international journal of educational research, 62, 100-114. https://doi.org/10.1016/j.ijer.2013.07.001 fisher, m., knobe, j., strickland, b., & keil, f. c. (2017). the influence of social interaction on intuitions of objectivity and subjectivity. cognitive science, 41(4), 11191134. https://doi.org/10.1111/cogs.12380 felton, m., crowell, a., & liu, t. (2015). arguing to agree: mitigating my-side bias through consensus-seeking dialogue. written communication, 32(3), 317-331. https://doi.org/10.1177/0741088315590788 felton, m. k., & herko, s. (2004). from dialogue to two-sided argument: scaffolding adolescents' persuasive writing. journal of adolescent & adult literacy, 47(8), 672 683. graff, g. (2003). clueless in academe: how schooling obscures the life of the mind. new haven: yale university press. hobbs, r., & frost, r. (2003). measuring the acquisition of media‐literacy skills. reading research quarterly, 38(3), 330-355. https://doi.org/10.1598/rrq.38.3.2 iordanou, k. (2010). developing argument skills across scientific and social domains. journal of cognition and development. 11(3), 293-327. https://doi.org/10.1080/15248372.2010.485335 iordanou, k. (2016a). from theory of mind to epistemic cognition. a lifespan perspective. frontline learning research, 4(5), 106 – 119. https://doi.org/10.14786/flr.v4i5.252 iordanou, k. (2016b). developing epistemological understanding through argumentation in scientific and social domains. zeitschrift für pädagogische psychologie. 30(2-3), 109119. https://doi.org/10.1024/1010-0652/a000172 iordanou, k. (2022). supporting strategic and meta-strategic development of argument skill: the role of reflection. metacognition and learning, 1-27. https://doi.org/10.1007/s11409-021-09289-1 iordanou, k. & constantinou. c. p. (2015). supporting use of evidence in argumentation through practice in argumentation and reflection in the context of socrates learning environment. science education, 99, 282–311. https://doi.org/10.1002/sce.21152 iordanou, k., kendeou., p., & beker, k. (2016). argumentative reasoning. in w. sandoval, j. greene, & i., bråten. (eds). handbook of epistemic cognition. new york, ny: routledge. https://doi.org/10.4324/9781315795225 iordanou, k., kendeou, p., & zembylas, m. (2020). examining my-side bias during and after reading controversial historical accounts. metacognition and learning, 15, 319–342. https://doi.org/10.1007/s11409-020-09240-w iordanou, k. & kuhn, d. (2020). contemplating the opposition: does a personal touch matter? discourse processes. 57(4), 343-359. doi:10.1080/0163853x.2019.1701918 iordanou, k., kuhn, d., flora matos, yuchen shi, & laura hemberger. (2019). learning by arguing. learning and instruction. 63, 101-207. 101207. https://doi.org/10.1016/j.learninstruc.2019.05.004 iordanou, k., & rapanta, c. (2021). “argue with me”: a method for developing argument skills. frontiers in psychology, 12, 631203. https://doi.org/10.3389/fpsyg.2021.631203 jucks, r., & bromme, r. (2011). perspective taking in computer-mediated instructional communication. journal of media psychology: theories, methods, and applications, 23(4), 192–199. https://doi.org/10.1027/1864-1105/a000056 kienhues, d., jucks, r., & bromme, r. (2020). sealing the gateways for post-truthism: reestablishing the epistemic authority of science. educational psychologist, 55(3), 144-154. https://doi.org/10.1080/00461520.2020.1784012 kienhues, d., stadtler, m., & bromme, r. (2011). dealing with conflicting or consistent medical information on the web: when expert information breeds laypersons' doubts about experts. learning and instruction, 21(2), 193-204. https://doi.org/10.1016/j.learninstruc.2010.02.004 https://psycnet.apa.org/doi/10.1111/cogs.12380 https://psycnet.apa.org/doi/10.1177/0741088315590788 https://psycnet.apa.org/doi/10.1598/rrq.38.3.2 https://doi.org/10.1080/15248372.2010.485335 https://doi.org/10.14786/flr.v4i5.252 https://doi.org/10.1024/1010-0652/a000172 https://doi.org/10.4324/9781315795225 iordanou & fotiou 40 | f l r kim, h. y., larusso, m. d., hsin, l. b., harbaugh, a. g., selman, r. l., & snow, c. e. (2018). social perspective-taking performance: construct, measurement, and relations with academic performance and engagement. journal of applied developmental psychology, 57, 24-41. https://doi.org/10.1016/j.appdev.2018.05.005 kohn, a. (2010, january). getting rid of grades: case studies. alfie kohn. https://www.alfiekohn.org/blogs/getting-rid-grades-case-studies/ kuhn, d., goh, w., iordanou, k., & shaenfield, d. (2008). arguing on the computer: a microgenetic study of developing argument skills in a computer-supported environment. child development, 79(1), 233-234. https://doi.org/10.1111/j.1467-8624.2008.01190.x kuhn, d. (2017). building our best future thinking critically about ourselves and our world teachers edition. new york, ny, usa: wessex press, inc. kuhn, d., & crowell, a. (2011). dialogic argumentation as a vehicle for developing young adolescents’ thinking. psychological science, 22(4), 545–552. https://doi.org/10.1177/0956797611402512 kuhn, d., hemberger, l., & khait, v. (2016). tracing the development of argumentive writing in a discourse-rich context. written communication, 33(1), 92–121. doi:10.1177/0741088315617157 kuhn, d., & iordanou, k. (2022). why do people argue past one another rather than with one another?. in n. ballantyne, & d. dunning (eds.), reason, bias, and inquiry: the crossroads of epistemology and psychology (pp. 324-338). oxford university press. https://doi.org/10.1093/oso/9780197636916.003.0015 kuhn, d., & udell, w. (2003). the development of argument skills. child development, 74(5), 12451260. https://doi.org/10.1111/1467-8624.00605 kuhn, d., zillmer, n., crowell, a., & zavala, j. (2013). developing norms of argumentation: metacognitive, epistemological, and social dimensions of developing argumentive competence. cognition & instruction, 31(4), 456–496. doi:10.1080/07370008.2013.830618 lalonde, c. e., & chandler, m. j. (2002). children's understanding of interpretation. new ideas in psychology, 20(2-3), 163-198. https://doi.org/10.1016/s0732-118x(02)00007-7 list, a., & alexander, p. a. (2019). toward an integrated framework of multiple text use. educational psychologist, 54(1), 20-39. https://doi.org/10.1080/00461520.2018.1505514 list, a., du, h., wang, y., & lee, h. y. (2019). toward a typology of integration: examining the documents model framework. contemporary educational psychology, 58, 228-242. https://doi.org/10.1016/j.cedpsych.2019.03.00 maier, j., & richter, t. (2016). effects of text-belief consistency and reading task on the strategic validation of multiple texts. european journal of psychology f education, 31(4), 479-497. https://doi.org/10.1007/s10212-015-0270-9 mateos, m., & solé, i. (2009). synthesising information from various texts: a study of procedures and products at different educational levels. european journal of psychology of education, 24(4), 435-451. https://doi.org/10.1007/bf03178760 mayweg-paus, e., macagno, f., & kuhn, d. (2016). developing argumentation strategies in electronic dialogs: is modeling effective?. discourse processes, 53(4), 280-297. https://doi.org/10.1080/0163853x.2015.1040323 mill, j. s. (1859/1996). on liberty. in d. wootton (ed.), modern political thought: readings from machiavelli to nietzsche (pp. 605–672). indianapolis: hackett. monte-sano, c. (2011). beyond reading comprehension and summary: learning to read and write in history by focusing on evidence, perspective, and interpretation. curriculum inquiry, 41(2), 212249. https://doi.org/10.1111/j.1467-873x.2011.00547.x monte-sano, c., & de la paz, s. (2012). using writing tasks to elicit adolescents’ historical reasoning. journal of literacy research, 44(3), 273-299. https://doi.org/10.1177/1086296x12450445 murphy, p. k., wilkinson, i. a., soter, a. o., hennessey, m. n., & alexander, j. f. (2009). examining the effects of classroom discussion on students’ comprehension of text: a metaanalysis. journal of educational psychology, 101(3), 740. https://doi.org/10.1037/a0015576 https://www.alfiekohn.org/blogs/getting-rid-grades-case-studies/ https://psycnet.apa.org/doi/10.1093/oso/9780197636916.003.0015 iordanou & fotiou 41 | f l r oecd (2023), pisa 2022 results (volume i): the state of learning and equity in education, pisa, oecd publishing, paris, https://doi.org/10.1787/53f23881-en paul, j., stadtler, m., & bromme, r. (2019). effects of a sourcing prompt and conflicts in reading materials on elementary students’ use of source information. discourse processes, 56(2), 155169. https://doi.org/10.1080/0163853x.2017.1402165 piaget, j. (1928). the child’s conception of the world. london: routledge and kegan paul.rabb, n., fernbach, p. m., & sloman, s. a. (2019). individual representation in a community of knowledge. trends in cognitive sciences, 23(10), 891–902. https://doi.org/10.1016/j.tics.2019.07.011 rabb, n., fernbach, p. m., & sloman, s. a. (2019). individual representation in a community of knowledge. trends in cognitive sciences, 23(10), 891-902. https://doi.org/10.1016/j.tics.2019.07.011 rapanta, c., felton, m. k. (2022). learning to argue through dialogue: a review of instructional approaches. educational psychology review 34, 477–509 https://doi.org/10.1007/s10648-02109637-2 richter, t., & maier, j. (2017). comprehension of multiple documents with conflicting information: a two-step model of validation. educational psychologist, 52(3), 1– 19. https://doi.org/10.1080/00461520.2017.1322968. rukavina, i., & daneman, m. (1996). integration and its effect on acquiring knowledge about competing scientific theories for text. journal of educational psychology, 88(2), 272–287. doi:10.1037/0022–0663.88.2.272. shi, y. (2019). enhancing evidence-based argumentation in a mainland china middle school. contemporary educational psychology, 59, article 101809. https://doi.org/10.1016/j.cedpsych.2019.101809 shi, y. (2020). talk about evidence during argumentation. discourse processes, 57(9), 770–792. https://doi.org/10.1080/0163853x.2020.1777498 shi, y., matos, f., & kuhn, d. (2019). dialog as a bridge to argumentative writing. journal of writing research, 11(1), 107–129. https://doi.org/10.17239/jowr-2019.11.01.04 stadtler, m., scharrer, l., skodzik, t., & bromme, r. (2014). comprehending multiple documents on scientific controversies: effects of reading goals and signaling rhetorical relationships. discourse processes, 51(1-2), 93-116. https://doi.org/10.1080/0163853x.2013.855535 stahl, s. a., hynd, c. r., britton, b. k., mcnish, m. m., & bosquet, d. (1996). what happens when students read multiple source documents in history? reading research quarterly, 31(4), 430– 456. https://doi.org/10.1598/rrq.31.4.5 tarchi, c., & mason, l. (2020). effects of critical thinking on multiple-document comprehension. european journal of psychology of education, 35(2), 289-313. https://doi.org/10.1007/s10212-019-00426-8 travis, t.a. (2017). grading for communication not compensation. trusted. https://trustedschool.org/2017/11/12/grading-for-communication-not-compensation/ van meteris, p., list, a., lombardi, d., & kendeou, p. (2020). handbook of learning from multiple representations and perspectives. routledge. https://doi.org/10.4324/9780429443961 vygotsky, l. s. (1978). mind in society: the development of higher mental psychological processes. cambridge, ma: harvard university press. https://doi.org/10.2307/j.ctvjf9vz4 walton, d. (1999). one-sided arguments: a dialectical analysis of bias. suny press. wiley, j., jaeger, a. j., & griffin, t. d. (2018). effects of instructional conditions on comprehension from multiple sources in history and science. in handbook of multiple source use (pp. 341361). routledge. wiley, j., & voss, j. f. (1999). constructing arguments from multiple sources: tasks that promote understanding and not just memory for text. journal of educational psychology, 91(2), 301– 311. https://doi.org/10.1037/0022-0663.91.2.301 zavala, j., & kuhn, d. (2017). solitary discourse is a productive activity. psychological science, 28(5), 578-586. https://doi.org/10.1177/0956797616689248 https://doi.org/10.1787/53f23881-en https://doi.org/10.1016/j.tics.2019.07.011 https://doi.org/10.1007/s10648-021-09637-2 https://doi.org/10.1007/s10648-021-09637-2 https://doi.org/10.1080/00461520.2017.1322968 https://trustedschool.org/2017/11/12/grading-for-communication-not-compensation/ https://doi.org/10.4324/9780429443961 https://psycnet.apa.org/doi/10.1037/0022-0663.91.2.301 iordanou & fotiou 42 | f l r figure 1 outline of the study design experimental group initial assessment a. prior-knowledge test b. argument skill (non-intervention topic) c. integration of multiple source perspectives (social & science*) intervention session(s) supporting reasons with evidence 1 e-chat 2-6 e-chat & reflection 7-9 preparation for showdown 10-11 e-showdown 12 feedback 13 f2f showdown 14 final assessment a. prior-knowledge test b. argument skill (non-intervention topic) c. integration of multiple source perspectives (social & science*) control group initial assessment a. prior-knowledge test b. argument skill (non-intervention topic) c. integration of multiple source perspectives (social & science*) business-as-usual school curriculum final assessment a. prior-knowledge test b. argument skill (non-intervention topic) c. integration of multiple source perspectives (social & science*) *note: the topics were counter-balanced. iordanou & fotiou 43 | f l r appendix coding scheme of multiple source perspectives code description example(s) from the dataset score presenting and justifying multiple positions 0–6 no position no position is presented regarding the inquiry questions. 1) i don’t know (what to write). 2) young children only know one language and when they grow older, they learn another one. a cognitive ability relates to jobs, financial problems, etc. 0 single position a single position is presented without a supporting reason or an explanation. my opinion is that bilingualism helps our cognitive abilities. 1 single position with own justification a single position is presented with a supporting reason or an explanation. the decline in one’s cognitive abilities that occurs as we age is slower in bilinguals and the symptoms of dementia are delayed for 4–5 years. 2 single position with justification and qualification a single position is presented with a supporting reason or an explanation and a qualification that conditionalizes the position. my opinion is that if someone is exposed to the sun at the right time of the day, from 12:00 till 14:00, and puts on sunscreen with a high sun protection factor, they will be ok; but always in moderation. 3 two positions two positions are presented without reasons or explanations. the first text talks about serious problems (in relation to cell phone use) while the second says that these might be exaggerations. 4 two positions with onesided, justification two positions are presented with a supporting reason or explanation for one position only. based on some studies (which according to my opinion are wrong) bilingual children have more cognitive abilities than monolingual children. another study, however, showed that bilingual children come from rich families who can spend a lot of money on their education. 5 two positions with twosided justification two positions are presented with supporting reasons or explanations for both positions. one side says that the sun is good for us due to the vitamin d that it provide us and due to the fact that we have less chance of getting cancer if we are exposed to the sun. the other side says that the sun is very bad for us because of the uv radiation. in fact, they say that we have more chances of getting cancer if we are exposed to the sun because our skin and internal organs can’t handle it. 6 connecting positions 0-2 iordanou & fotiou 44 | f l r no explicit connection no relation between the positions is explicitly stated; they are not presented as contrastive in any way. the sun causes both illness and health. 0 positions connected positions are explicitly related to each other or compared and contrasted for example with the use of contrastive conjunctions, by making reference to the different sources.1 one view, according to fisher, refers to the fact that various scientists recommend sunbathing so as to obtain vitamin d. he warns people that this is problematic. another view is that those who had high levels of vitamin d in their meals and were active had less chances of getting cancer. the differences between these views are that in the first one fisher warns us that it is not ok to sunbathe so as to get vitamin d while the other view says that those who had higher levels of vitamin d they acquired it through their meals and by having an active lifestyle. 1 positions reconciled positions are reconciled by providing an explanation for the differences between them and/or by drawing a conclusion based on consideration of both positions. one difference is that they will learn from their grades (and their mistakes), but if they don’t receive grades they will not learn from their mistakes. however, if they don’t receive any grade, it will still be helpful because getting feedback is a more helpful strategy and gives better results; however, this holds only if this system is implemented correctly. 2 1 the part in italics is our addition to the coding scheme. this was done so that we make more explicit how we have implemented this criterion. codepen paans et al publication frontline learning research vol.8 no. 1 (2020) 76 95 issn 2295-3159 children’s macro-level navigation patterns in hypermedia and their relation with task structure and learning outcomes cindy paansa, inge molenaar aeliane segers ab& ludo verhoevena abehavioural science institute, radboud university, the netherlands binstructional science, twente university, the netherlands article received 9 april 2019/ revised 30 january / accepted 18 february / available online 26 february abstract this study investigated macro-level navigation patterns in children’s hypermedia learning, and how they related to task structure and learning outcomes. for this purpose, 5th and 6th grade learners performed a hypermedia assignment in which a high (n=57) versus a low (n=54) level of structure was provided. by means of qualitative analyses of their navigation activities, 6 macro-level navigation patterns were distinguished: linear reading, selective reading, video viewing, massed writing, late onset writing, and unpredictable reading. results showed that the linear reading pattern was more frequent in the high structure environment, and that both the high structure environment and the linear reading pattern were associated with the highest quality of the children’s written assignments. navigation patterns and task structure did not clearly predict children’s declarative knowledge gains or knowledge transfer. these findings show that there are multiple ways to navigate through a hypermedia environment, but that these are not all equally successful for learning. moreover, the provided task structure in the environment may affect the occurrence of successful navigation patterns. keywords: navigation activities; task structure; hypermedia; primary education info corresponding author c.paans@psych.ru.nl doi: 10.14786/flr.v8i1.473 1. introduction in order to be able to participate in today’s society, primary school children are taught multiple skills that are deemed crucial in the 21st century (rotherham & willingham, 2010). one of these skills involves the adequate use of online materials that are found on the internet (ananiadou & claro, 2009). children sometimes struggle to learn on the internet, however, and learners may learn and comprehend less in digital environments (delgado, vargas, ackerman, & salmeron, 2018; kong, seo, & zhai, 2018). on top of that, hypermedia environments (such as the internet), which combine hypertext with multimedia, may result in disorientation and distraction (scheiter & gerjets, 2007). providing external support within the environment, by structuring the learning material, can improve learning outcomes, depending on the needs of the learner (goldman, 2009). these findings show that individual variation exists in hypermedia learning outcomes, and is affected by task structure. in addition, these findings show there also exists individual variation in the learning process (i.e., whatever happens during learning). as such, investigating children’s learning processes may help to explain individual differences in learning outcomes and their relation with task structure. this, in turn, can be used to develop adequate support for weaker learners. however, when studying the learning process, for example, by studying how learners navigate through a website, interpreting individual events is problematic (see e.g., salmeron, naumann, garcía, & fajardo, 2017). consequently, a macro-level view of their learning process, which looks at overall patterns of events, may aid interpretation. research that takes a macro-level view is largely lacking, however. hence, it is unclear which macro-level patterns can be distinguished in children. the current study therefore investigated which macro-level navigation patterns could be distinguished in children, and how these patterns were related to task structure and learning outcomes. 1.1 theoretical framework of hypermedia learning hypermedia environments provide learners with the flexibility to navigate through their learning environment in many different ways. an advantage of this, is that learners can adapt their learning process to their own needs (gorissen, kester, brand-gruwel, & martens, 2015; scheiter & gerjets, 2007). however, not all leaners succeed in learning in a hypermedia environment. the hypertext structure may create high cognitive load (destefano & lefevre, 2007), result in disorientation (mcdonald & stevenson, 1996), or result in gaps in comprehension (bezdan, kester, & kirschner, 2013). moreover, because of the inherent flexibility of hypermedia environments, learners need to regulate their own learning (azevedo & cromley, 2004). the field of self-regulated learning (srl) focusses on how children achieve learning goals, by using cognitive activities to study the learning material, metacognitive activities to monitor and control their progress, and motivation to keep engaged with the material (winne & hadwin, 1998; winne & nesbit, 2010). the information processing view of srl (winne & hadwin, 1998) describes srl as a cyclical process in which the learner moves through four phases, namely task definition, goal setting and planning, enacting study tactics and strategies, and metacognitively adapting studying. the learner can move through these phases in a non-linear fashion, and may go through them multiple times within a learning assignment (azevedo, 2009; molenaar & järvelä, 2014). srl, and metacognitive activities in particular, have been related to better learning outcomes (e.g., eilam & aharon, 2003; van der stel & veenman, 2008). meanwhile, however, children often do not spontaneously regulate their learning (de jong & van joolingen, 1998), which may partly explain their difficulties in hypermedia learning. 1.2 supporting hypermedia learning through task structure one way to support the regulation of learning in hypermedia environments is by providing external structure or support. empirically, research has investigated the effect of different levels of task structure in the learning environment on both the learning process and learning outcomes. such research included comparisons between linear and non-linear text (blom, segers, knoors, hermans, & verhoeven, 2018; klois, segers, & verhoeven, 2013; mcdonald & stevenson, 1996), the effect of advance organizers (urakami & krems, 2012), the presence of link suggestions (ignacio madrid, van oostendorp, & puerta melguizo, 2009), the specificity of problem solving goals (künsting, wirth, & paas, 2011), and the specificity of questions (rouet, 2003). these studies showed that providing more structure does not universally lead to better learning outcomes. instead, the provided structure needs to fit the regulatory needs of the learner (goldman, 2009; kalyuga, ayres, chandler, & sweller, 2003). furthermore, while it has been suggested that learners adapt to the requirements of a task, this adaptation does not always lead to better performance (pieschl, stahl, murray, & bromme, 2012). also on a theoretical level, srl research has emphasized the importance of taking into account the learning context when studying the effect of srl on learning outcomes (boekaerts, 1999; greene & azevedo, 2010; zimmerman, 1989). in the information processing model of srl, the learning context is incorporated in the notion of task conditions (winne & hadwin, 1998), which refer to task characteristics such as provided instructional cues, or the available time. these task conditions affect which cognitive and metacognitive activities are performed during learning. as such, both empirically and theoretically, differences in task structure have been suggested to result in differences in the learning process (e.g., bezdan et al., 2013; ignacio madrid et al., 2009; mobrand & spyridakis, 2007). therefore, investigating the learning process may help to understand how task structure affects learning outcomes. 1.3 macro-level navigation patterns as described above, one way to illuminate why children differ in their hypermedia learning outcomes, and how task structure affects these outcomes, is by looking at their learning process (see e.g., goldman, braasch, wiley, graesser, & brodowinska, 2012; malmberg, järvenoja, & järvelä, 2013; salmeron et al., 2017). one aspect of the learning process is how children navigate through the hypermedia environment. log-files provide a record of children’s navigation activities while they occur (greene & azevedo, 2010). such online measurement provides a more in-depth way of mapping the learning process than off-line measures, such as questionnaires, without having to rely on the memory of, or introspection by, the learner. log-files have the advantage over other online measures that they are relatively easy to collect at a large scale, are not intrusive, and do not disrupt the learning process (veenman, 2015; winne, hadwin, & gress, 2010). they may also be more suitable than, for example, think aloud protocols when studying children, because they do not rely on the learner’s ability to verbalize what they are thinking (see also veenman, 2011). while navigation activities may be useful for gaining insight into children’s learning processes, interpreting individual navigation events is not straightforward. for example, a reading time of a web-page may indicate that a child scanned the entire page, or that he or she read part of it more thoroughly (salmeron et al., 2017). in addition, the effect of individual navigation events on learning outcomes may depend on how the learner followed up on the event. for example, if a learner read only half of one page, the effect on learning outcomes may depend on whether or not the learner already knew the content of the page, and on whether he or she reread the page at a later stage. consequently, a more holistic approach (see also reimann, 2009) to investigating navigation activities may help to interpret individual navigation events. as such, looking at macro-level navigation patterns, rather than individual navigation events, may help with the interpretation of log-file data. we define macro-level navigation patterns as the overall approach a learner takes to navigate through a hypermedia assignment. at the macro-level, several studies have shown that different learning processes can be observed in (groups of) learners who learn under different task conditions (bannert, reimann, & sonnenberg, 2014; sobocinski, malmberg, & järvelä, 2017; sonnenberg & bannert, 2015), or who have different learning outcomes (schoor & bannert, 2012). as learners gain experience in a task, they may approach the task differently (see e.g., expertise reversal effect; kalyuga et al., 2003). in addition, their approach may change depending on task conditions (pieschl et al., 2012). hence, instead of being stable across time, navigation patterns are likely to be inherently dynamic and variable, even within one individual. together, these findings suggest that children may move through a hypermedia assignment in multiple ways. it is still unclear, however, which macro-level navigation patterns can be distinguished in children. while research is scarce, some studies have investigated macro-level navigation patterns, by means of cluster analysis. macgregor (1999) found three types of navigation patterns in their study on 7th and 11th grade students who used an instructional hypermedia system. first, sequential studiers accessed information in a sequential manner. they moved slowly and methodically through the pages, with an emphasis on the textual components. second, video viewers primarily watched videos. last, concept connectors access the information more selectively. this selection appeared to be based on prior knowledge. in addition, they alternated between skimming text and reading thoroughly. similarly, lawless and kulikowich (1996) also found three macro-level navigation patterns in their study on hypertext learning in undergraduate students who learned about lyme disease, which they referred to as feature explorers, who invested more time in understanding the environment than learning information; knowledge seekers, who had high scores on the outcome measures; and apathetic hypertext users, who spent little time in the environment, used few special features, and inspected few pages. however, barab, bowdish, and lawless (1997) found four patterns in undergraduates who performed a hypermedia information retrieval task. first, model users chose the simpler information retrieval task and provided the correct solution without deviating to irrelevant pages. second, disenchanted volunteers had low retrieval scores and explored very little of the environment. third, feature explorers used the help screens, watched many videos, and had low retrieval scores. finally, cyber cartographers spent a lot of time in the environment and explored many of the pages. while these studies show some similarities in the clusters they found, there is no complete overlap. one potential reason for this difference is that the measures that were used for extracting the clusters differed between the studies. in fact, many studies that use navigation activities to understand the learning process, do not state reasons for their choice concerning which variables they extract from the log-files. as a consequence, it is unclear which navigation measures reflect the most dominant differences between groups of learners. one solution to this problem is to take a qualitative approach, in which macro-level navigation patterns are clustered based on their visual appearance. this way, the most dominant features of a pattern become salient. this is the approach we took in this study. 1.4 the present study summarizing, investigating children’s navigation activities may help to explain how task structure affects hypermedia learning outcomes. because the meaning of individual navigation events is not self-evident and depends on the context in which they occur, the interpretation of children’s navigation activities may benefit from analysing them in terms of macro-level navigation patterns. furthermore, although various clusters have been suggested for older students, it is unclear which patterns children use during hypermedia assignments. the present study therefore investigated the effect of task structure on macro-level navigation patterns and hypermedia learning outcomes in children. for this purpose, three research questions were examined: rq1: which macro-level navigation patterns can be distinguished in children learning in a hypermedia assignment? rq2: does the prevalence of these navigation patterns depend on task structure? rq3: how do navigation patterns and task structure relate to learning outcomes? 2. method 2.1 participants participants were recruited via a letter to the school administrators. active parental consent was obtained. in total, 111 5th and 6th grade children from five different primary schools in the netherlands participated. of these, 54 were boys, and 56 were girls (one child did not fill in his/her gender). age ranged from 9 to 13 years (m = 11 years, 8.5 months, sd = 9.7 months). participants in each classroom were randomly assigned to one of two conditions of the learning environment (see task structure variation). of the participants, 57 learned in a high structure environment, whereas 54 learned in a low structure environment. participants in the two conditions did not differ in terms of their gender or age, p’s > .10. 2.2 materials 2.2.1 the learning environment in the present study, children learnt about the heart and living a healthy lifestyle by means of a webquest (segers & verhoeven, 2009). they were instructed to write a 300-word text that covered four topics: (1) the anatomy of the heart; (2) an explanation of the circulatory system; (3) the components of blood; and (4) how to keep the heart healthy. the assignment was done in a closed hypermedia environment. nine preselected content pages, with links to four short videos were provided as resources. a schematic overview of such a content page is shown in figure 1. participants could not search the internet for more information. on the left of the screen, a navigation menu was displayed, by means of which the children could access information pages about the assignment, a content overview from which they could access the resources, and a link to the worksheet in which they wrote their assignment. participants could switch between reading the resources and writing. whenever they saved their worksheet, they were informed of how many words they had written, and could return to the last page they accessed prior to the worksheet, without having to navigate through the content overview. figure 1. schematic overview of a content page in the webquest environment 2.2.2 task structure variation two versions of the hypermedia environment were created. both versions contained exactly the same assignment and resources, but one provided high structure, whereas the other provided low structure. the two versions differed in the following three ways. first, the information pages about the assignment differed. in the high structure condition, participants were provided with an introduction page; an assignment page in which the assignment was introduced; a roadmap page that broke the assignment down into smaller steps; a review page in which children could review their work; and a conclusion page. in the low structure condition, the road map page was not provided. instead, their assignment page introduced the assignment, and contained three additional sentences about the topics that had to be included, and on how to navigate the website. second, the layout of the worksheet differed for the two conditions (see figure 2). the high structure condition provided four text blocks with a header, one for each assignment topic. in the low structure condition, it showed only one large text block. finally, the layout of the content overview differed for the two conditions. in the high structure condition, the nine resources were grouped under four headers that roughly corresponded to the four assignment topics. in the low structure condition, these four headers were missing. the order of the resources was the same in both conditions, however. figure 2. screenshot of the worksheet in the structured (left) and unstructured (right) learning environment. 2.2.3 navigation activities in order to measure children’s navigation activities, log-files were collected that registered every hyperlink selection with a time stamp with a one-second precision. in the current study, only information concerning their page views for resource pages and the worksheet was taken into account. their use of the assignment pages was not investigated for two reasons. frist, viewing frequencies of these pages were relatively low. on average, the assignment page was viewed 7.54 times (sd = 8.89; total viewing duration: m = 136.69 seconds; sd = 111.28); the roadmap page (for the high structure condition only) was viewed an average of 5.65 times (sd = 6.77; total viewing duration: m = 78.31 seconds; sd = 54.70); and finally, the review pages were viewed an average of 3.55 (sd = 4.16; total viewing duration: m = 54.29 seconds; sd = 94.49). second, a previous study with a similar set-up showed that the mapping of assignment related pages to metacognitive activities is not straightforward, especially for learners with low learning gain (see also: paans, molenaar, segers, & verhoeven, 2019). moreover, a few participants repeatedly accessed the worksheet via the assignment page or roadmap page, rather than the navigation menu, which further complicated the interpretation of page views for these assignment-related pages. 2.2.4 assignment quality the quality of the assignment, which participants wrote in the worksheet, was assessed. for this purpose, their assignment was rated by means of a predefined list of 50 terms, and explanations of terms or mechanisms. the number of used terms and explanations constituted their final score. of the assignments, 24 (22%) were scored by a second rater. interrater reliability was high (terms: intraclass r = .93; explanations: intraclass r = .93; total scale: intraclass r = .98). 2.2.5 declarative knowledge to measure declarative knowledge, participants had to connect 14 words to explanatory sentences. for example, they had to connect “white blood cell” to “helps fight disease”. they were given the same task before and after the hypermedia assignment, to measure their prior knowledge as well as their knowledge at post-test. the number of items correct constituted their score on the measure. internal consistency of the measure at pre-test was relatively poor (α = .51), due to the fact that most participants had very low prior knowledge. thus, many correct items might be attributed to chance. at post-test, internal consistency was good (α = .83). 2.2.6 knowledge transfer to measure knowledge transfer, participants were asked to design a birthday party that was fun and good for one’s heart. they were asked to name three activities and explain why they were healthy. they received one point for each correct activity and explanation. the maximum obtainable score was therefore six. activities were scored based on their novelty. for example, eating three different fruits only received one point, whereas eating a fruit salad, building a tree hut, and playing hide and seek, counted as three activities. explanations were scored on their accuracy. inaccurate (e.g., candy is healthy because it contains sugar) or non-explanations (e.g., “it is healthy because it is healthy”) received no points; incomplete explanations (e.g., “it is healthy because you have to run a lot”) received half points; and complete explanations (e.g., “fruit is healthy because it contains vitamins”) received a full point. to assess the quality of the scoring, 24 transfer assignments (22%) were scored by a second rater. inter-rater reliability was good (intraclass r = .79). 2.3 procedure this study was part of a larger data collection on the effect of cognitive predictors, such as executive functions, prior knowledge, and non-verbal reasoning on children’s hypermedia learning. only the parts relevant to the present study are described below. in each classroom, the study started with a classroom session. this session started with a short explanation of the study, where we explained that not all children would receive the same layout of the website. next, the children received the declarative knowledge pre-test. after making an example item, they were given 15 minutes to complete the test. they were encouraged to fill in what they knew and to guess the rest. finally, they received a short presentation of the hypermedia assignment. after the classroom session, the children participated in an individual session outside the classroom. this session included the hypermedia assignment, the declarative knowledge post-test, and the transfer task. during the hypermedia assignment, participants first received another short explanation of the hypermedia environment. next, they were given 45 minutes to complete the task, and were reminded thrice of the remaining time. any questions about the content were not answered, but about the use of the computer were. whenever children indicated they were finished before all time had elapsed, they were asked to reread the assignment page and to re-evaluate whether or not they had finished. after the hypermedia assignment, participants were given the declarative knowledge post-test, for which they again received 15 minutes. finally, they were given the transfer task. testing in each classroom took place within the course of two weeks. sessions never followed each other directly, to prevent memory and fatigue effects. after participating, all children received a small gift. no classroom credits were rewarded. 2.4 analyses 2.4.1 qualitative exploration of macro-level navigation patterns in order to investigate the first research question, on which macro-level navigation patterns could be distinguished, a qualitative analysis of the children’s navigation patterns was performed. for this purpose, we took a grounded approach. in the first stage, the type and number of macro-level patterns was established. for this purpose, a line graph was made with time on the x-axis and web page on the y-axis for each individual’s navigation activities. these graphs were then grouped together based on similarities and differences in their appearance (see also figure 3) by the first author. after an initial grouping, the groups were re-evaluated to see whether initial groups were overlapping and could be combined, or whether any additional patterns could be distilled. this led to a final set of macro-level navigation patterns, based on which coding criteria were formulated (see appendix a). finally, to check for grouping errors, all graphs were coded based on the final coding criteria. in the second stage, reliability of the grouping was established by a second rater. this rater was trained on the navigation patterns of a different data collection and a random half of the data presented here. reliability was assessed on the other random half of the current data. inter-rater reliability was sufficient (κ = .73). 2.4.2 assessing cluster differences between conditions in order to investigate the second research questions, on whether the prevalence of the navigation patterns depends on task structure, a χ² analyses was performed. as a post-hoc comparison, column proportions with bonferroni corrections were assessed to establish which clusters differed in their expected frequency between the two conditions. 2.4.3 assessing the effect of macro-level navigation patterns and task structure on learning outcomes in order to investigate the third research question, on how navigation patterns and task structure related to learning outcomes, three steps were taken. first, it was tested whether task structure predicted learning outcomes. for declarative knowledge, an ancova was performed, with pre-test as the covariate. for assignment quality and knowledge transfer, two independent samples t-tests were performed. second, it was investigated to which extent cluster membership predicted learning outcomes. because of the low sample size for some clusters, non-parametric tests were used. a kruskall-wallis test was performed, with cluster membership as the independent variable, and the learning outcomes as the dependent variables. because the kruskall-wallis test does not allow for the use of a covariate, declarative knowledge gain was used as the outcome variable, instead of knowledge at post-test. post-hoc comparisons were done using separate mann-whitney tests. for these post-hoc comparisons, a more conservative significance level of α = .01 was used. finally, it was investigated whether the effect of cluster on the learning outcomes was dependent on task structure. again, because of the low sample size for some clusters, a regular 2x2 anova could not be performed to test the interaction between task structure and cluster membership. therefore, the two conditions were analysed separately with the same method as described above (kruskall-wallis tests for assessing cluster differences, and mann-whitney tests for post-hoc comparisons). 3. results 3.1 qualitative description of macro-level navigation patterns the first research question asked which macro-level navigation patterns can be distinguished in children. from a qualitative analysis of their navigation activities, six macro-level patterns emerged. these are described below. figure 3 shows an example graph for each pattern. in each of the panels of figure 3 (a through f), one of the navigation patterns is shown. the x-axis shows the time and the y-axis shows which page the participant is viewing at that particular time. each page has its own horizontal row, with the bottom nine rows corresponding to the nine content pages in their order of presentation. that is, the bottom row corresponds to the resource page that was presented at the top of the page, the second row from the bottom corresponds to the second, and so forth. the next four rows correspond to the four video pages, and the top row corresponds to the worksheet. solid lines indicate that a participant continuously viewed this page, whereas interrupted lines indicate that he or she switched between pages. 1. linear reading: the first navigation pattern was characterized by a relatively linear reading pattern, where the children read most of the resource pages, and primarily in the order in which they were presented in the content overview. figure 3a illustrates this by the stair-like pattern in the resource pages. some participants, but not all, viewed videos as they encountered them (not shown in the example figure). their writing was interleaved with their reading, which could be seen by frequent switches between the resources and the worksheet. in the figure, this is visible through the interrupted lines in both the resource pages and the worksheet. mostly, their writing started relatively early on. 2. selective reading: in this pattern, children selected up to three resource pages that they used for their writing activities. this can be seen in figure 3b, where the participant spent most of his or her time on three content pages. some participants, but not all, first quickly browsed through the content pages in their order, before starting to write. once they started engaging with their selected resources, their writing was interleaved, with many switches between the resource page and their worksheet, which can be seen by the interrupted line for the worksheet page that runs along the entire x-axis. 3. video viewing: in this pattern, children showed a preference for watching video material over written resources, as could be seen (figure 3c) by the relative sparsity in time spent on written resources, and longer stretches on video pages. for most of these children, their writing was interleaved, with many switches between the videos and the worksheet. 4. massed writing: contrary to the previous three patterns, these children massed their writing activity. their navigation patterns could be characterized by long stretches in the worksheet in which the children did not switch back to the resources. figure 3d illustrates this pattern by showing one shorter and one longer solid line for the worksheet. this solid line indicates that the participant did not switch back to a resource page while writing. 5. late onset writing: this navigation pattern was characterized by a relatively late first engagement with the worksheet. for example, for the participant shown in figure 3e, a first line appears for the worksheet roughly 12 minutes into the assignment. the navigation activities started with browsing or reading through the resources. some children, but not all, did this in a linear fashion. figure 3e shows an example of this linear fashion through the stair-like pattern in the resource pages during the first 10 minutes of the assignment. after some initial time engaging with the resource pages, they started to engage with the worksheet. at this stage, some participants chose two or three resource pages with which they interleaved their writing in the worksheet. others, however, started to navigate in a pattern that was more similar to that of massed writing, with long stretches in the worksheet, and little reference to the resource pages. 6. unpredictable reading: in the last pattern, the children showed a relatively random pattern of navigating. their patterns showed many switches between resources pages, and in a non-linear fashion. no resource page dominated their pattern, nor did it show a clear preference for videos or written resources. writing was predominantly interleaved with viewing the resource pages. figure 3f illustrates this, by showing a relatively diffuse pattern with many switches between the resource and video pages. figure 3. examples for the six navigation patterns: linear reading (a), selective reading (b), video viewing (c), massed writing (d), late onset writing (e), and unpredictable reading (f). the x-axis shows the time throughout the assignment. the y-axis shows the page the learner is viewing at that particular time. 3.2 relation between task structure and macro-level navigation patterns the second research question asked whether the prevalence of the navigation patterns depended on task structure. a comparison of the frequency of occurrence of the six clusters within the two conditions (table 1), showed that cluster membership and condition were statistically non-independent, χ² (5) = 13.43, p = .020. a post-hoc comparison of the column proportions showed that only the frequency of the linear cluster differed for the two conditions at the α = .05 level, indicating that linear reading was more likely in the high structure condition than the low structure condition. table 1 distribution of the cluster over the two conditions 3.3 task structure and learning outcomes the third research question asked how navigation patterns and task structure related to learning outcomes. it was first investigated how task structure predicted learning outcomes. the descriptive statistics of the various learning measures are shown in table 21 and the correlations in table 3. the two conditions did not differ in terms of their declarative knowledge at post-test, when controlling for pre-test scores, f(1, 104) = .06, p = .804. additionally, the two conditions did not differ in terms of knowledge transfer, t(105) = .32, p = .751, but did differ in terms of assignment quality, t(106) = 2.51, p = .014, d = 0.48, ci: [.91; 7.82]. on average, assignment quality was 4.37 points higher in the high structure environment than the low structure environment. table 2 descriptive statistics for learning measures per condition table 3 correlations between the learning outcomes 3.4 macro-level navigation patterns and learning outcomes in the second step to investigate the third research question, it was investigated whether macro-level navigation patterns predicted learning outcomes, and whether the effects were different in the two task conditions. descriptive statistics of the learning outcomes per cluster are displayed in table 4, for each condition separately. first, the main effect of cluster membership on declarative knowledge gain, assignment quality, and knowledge transfer was assessed. overall, clusters did not differ in terms of the resulting declarative knowledge gain, h(5) = 10.16, p = .071, or knowledge transfer, h(5) = 3.63, p = .604. they did differ in assignment quality, however, h(5) = 35.04, p < .001. post-hoc comparisons showed that linear reading resulted in a higher assignment quality than video viewing (z = 4.009, p < .001, r = .60), massed writing (z = 4.483, p < .001, r = .63), and late onset writing (z = 3.864, p < .001, r = .55). selective reading and unpredictable reading also had higher assignment quality than massed writing (selective: z = 2.765, p = .006, r = .50; unpredictable: z = 3.105, p = .002, r = .54). to test the condition*cluster interaction, each condition was analysed separately. in the high structure condition, clusters did not differ in terms of declarative knowledge gain, h(5) = 9.68, p = .085, or knowledge transfer, h(5) = 2.75, p = .739. they did differ in terms of assignment quality, h(5) = 13.27, p = .021. post-hoc comparisons showed that linear reading had higher assignment quality than video viewing (z = 2.832, p = .005, r = .50) and massed writing (z = 3.108, p = .002, r = .56). no other comparisons were significant at the α = .01 level. for the low structure condition, clusters did not differ in terms of knowledge transfer, h(5) = 2.55, p = .768. they did differ, however, in assignment quality, h(5) = 21.78, p = .001, and declarative knowledge gain, h(5) = 11.31, p = .045. linear reading predicted higher assignment quality than video viewing (z = 2.789, p = .005, r = .77), massed writing (z = 2.817, p = .005, r = .63), and late onset writing (z = 3.316, p = .001, r = .74). in terms of declarative knowledge gain, no post-hoc comparisons were significant at the α = .01 level2. table 4 descriptive statistics for learning outcomes by cluster and condition 4. discussion this study investigated the effect of task structure on macro-level navigation patterns and hypermedia learning outcomes in children. six macro-level navigation patterns could be distinguished. results showed that the prevalence of these patterns depended on task structure. both task structure and navigation patterns predicted the quality of the written assignment, but not learning gain. the effect of navigation patterns on assignment quality was similar in both task conditions. the first result was that six macro-level navigation patterns could be distinguished. this is a larger number than reported in earlier studies (e.g., barab et al., 1997; lawless & kulikowich, 1996), probably because we also incorporated the writing process. consequently, we could distinguish between interleaved and massed writing, as well as the onset of writing. there are also similarities with previously found navigation patterns. notably, most similarities are found with the study by macgregor (1999), who’s study was most similar in terms of the used environment and the age of the participants. first, the linear reading pattern showed similarities with the sequential studiers, who went through the resources systematically and spent most time on text pages. second, the video viewing pattern was similar to the video viewers macgregor found. finally, while there is no complete overlap, there appear to be similarities between our selective reading pattern and the previously found concept connectors, as both accessed information selectively, and both scanned some information, while reading other information more thoroughly. the second result was that the provided task structure affected the occurrence of the macro-level navigation patterns. linear reading was far more common in the high structure environment. as such, it appears that the provided structure affected the reading and writing patterns of the participants. this finding is coherent with research on srl, which showed that interventions that aided regulation by means of prompts, changed the learning process of the learners (sonnenberg & bannert, 2015). furthermore, prior research has shown that advance organizers, which provide structure to the learning content, can aid learners navigation and comprehension on the internet (urakami & krems, 2012). similarly, hypertext environments with simpler lay-outs (such as hierarchical hypertexts or linear texts) are easier to navigate by children than environments with a more complex lay-out (such as networked hypertexts; e.g., blom et al., 2018). as such, the provided structure may have decreased any potential disorientation that the children were experiencing. the last result was that both task structure and navigation patterns predicted the quality of the written assignment. the linear reading pattern resulted in the highest quality assignments, as did the high structure environment. since the linear reading pattern was also far more prevalent in the high structure environment than in the low structure environment, it might be possible that the effect of task structure on assignment quality is driven by this difference in macro-level navigation patterns. at the level of individual navigation events, it has already been shown that task structure affected navigation (ignacio madrid et al., 2009; mobrand & spyridakis, 2007), and that navigation affected assignment quality (paans, segers, molenaar, & verhoeven, 2018). this study adds that the relation between task structure, navigation patterns and assignment quality, can also be found at the macro level. the provided task structure may have reduced the cognitive load the learners experienced (see e.g., zumbach & mohraz, 2008), by providing a scaffold for the regulation of their learning process. external regulation, for example, by means of a tutor, has been shown to improve hypertext learning outcomes (azevedo, greene, & moos, 2007). at the same time, the linear reading pattern may have ensured that learners saw more of the content, and in a logical order, which in turn may have resulted in fewer gaps in their comprehension of the learning material (salmeron, kintsch, & cañas, 2006a, 2006b). besides the effects of task structure and navigation patterns on assignment quality, the low structure environment also showed an effect of navigation patterns on declarative knowledge gain. however, post-hoc comparisons showed only marginal effects, even though effect sizes were large. as such, we tentatively suggest that massed writing is likely to be related to the lowest knowledge gains. indeed, in the massed writing pattern, learners spent long stretches of time in the worksheet, without referring back to the resources. consequently, learners may not have spent enough time acquiring new knowledge or challenging misconceptions. these results, however, will need to be replicated in another study. a limitation of the study is the sample size. because no a priori expectation was formulated with regard to the number of clusters, or their frequency of occurrence, no reliable estimate could be made for the required sample size. in the current study, the frequency of some clusters was relatively low, especially when task structure was taken into account. consequently, not all analyses could be performed due to lack of power. a future study may seek to replicate and extend the current findings with a larger sample. a second limitation relates to the difference between the two task conditions. the two conditions differed in three ways: the presence of a roadmap page, the layout of the worksheet, and the presence of guiding headers in the content overview. consequently, it cannot be inferred from the data which of these three aspects drove the differences between the two conditions. future research could investigate whether it was one single aspect, or a combination of these that led to the obtained results. a final limitation relates to the transfer task that was used in this study. no effects on knowledge transfer were found for either cluster membership or task structure. a possible explanation for this lack of findings might be that the transfer task also made a demand on the participants’ writing skills. as such, the measure may have contained construct-irrelevant variance that obscured any potential results. this study used a qualitative method to find clusters of macro-level navigation patterns. future research may seek to replicate and quantify these clusters. first, replication is needed to ascertain whether similar patterns are found in different contexts, and whether the patterns are unidimensional. it is possible that different clusters can be found, or that clusters are overlapping in different task contexts. second, the clusters may be quantified. given the large amount of data that log-files provide, it can be challenging to decide which measures to extract from the log-files to describe the learning process. moreover, a clear consensus regarding the methods of analyses for log-files is still lacking in the field (van laer & elen, 2018). therefore, the qualitative findings of this study could be used to inform the selection of potentially relevant navigation measures. another direction for future research is to investigate the stability of the navigation patterns across time. while this study distinguished six macro-level patterns, we cannot ascertain the extent to which children use the same pattern consistently across multiple assignments. perhaps, children who adapt their macro-level navigation pattern to the task requirements and to their own domain knowledge will have better learning outcomes than those who consistently use the same macro-level pattern. indeed, previous research has shown that learners with high or low prior knowledge have different needs in an assignment (goldman, 2009; kalyuga et al., 2003), and approach tasks differently (lawless, brown, mills, & mayall, 2003). consequently, the provided structure should fit the needed support by the learner in order to be effective (azevedo & hadwin, 2005). this study showed that a difference in task structure affected the macro-level navigation patterns of children and the quality of their written hypermedia assignment. a practical implication of these findings is that care should be taken when designing digital resources for primary school children. a relatively small difference in headers or answering formats may result in different learning strategies, as evidenced by navigation patterns, and even in different learning outcomes. as such, it is advisable to test new applications and websites rigorously, so that optimal learning is safeguarded. while such testing is not always performed, teachers could keep an extra eye on children’s learning behaviours online, to help prevent them getting disoriented or distracted (see also scheiter & gerjets, 2007). our study provides many new directions for research. six macro-level navigation patterns were found that children used to learn in a hypermedia environment. on average, the quality of their written assignment was higher, and a linear reading pattern more prevalent, in a high structure hypermedia environment. indeed, findings indicate that the linear reading pattern was related to a higher quality assignment compared to some other navigation patterns. together these findings show that, while there are multiple ways to navigate through a hypermedia environment, not all ways are equally successful. moreover, the provided structure in the environment may affect the occurrence of successful navigation patterns, and could therefore affect overall learning success in hypermedia. keypoints six macro-level navigation patterns could be distinguished in children's hypermedia learning. a linear reading pattern was more prevalent in a high structure environment. a high structure environment was related to better assignment quality. a linear reading pattern was related to better assignment quality. footnotes 1 some participants had a score of zero on their post-test. leaving these participants out of the analyses yielded similar results. 2 at the α = .05 level, results showed that both linear reading (z = 2.18, p = .029, r = .49) and selective reading (z = 2.42, p = .016, r = .54) had higher knowledge gains than massed writing. references ananiadou, k., & claro, m. (2009). 21st century skills and competences for new millennium learners in oecd countries . oecd education working papers. https://doi.org/10.1787/218525261154 azevedo, r. (2009). theoretical, conceptual, methodological, and instructional issues in research on metacognition and self-regulated learning: a discussion. metacognition and learning, 4(1), 87–95. https://doi.org/10.1007/s11409-009-9035-7 azevedo, r., & cromley, j. g. (2004). does training on self-regulated learning facilitate students’ learning with hypermedia? journal of educational psychology, 96(3), 523–535. https://doi.org/10.1037/0022-0663.96.3.523 azevedo, r., greene, j. a., & moos, d. c. (2007). the effect of a human agent’s external regulation upon college students’ hypermedia learning. metacognition and learning, 2(2–3), 67–87. https://doi.org/10.1007/s11409-007-9014-9 azevedo, r., & hadwin, a. f. (2005). scaffolding self-regulated learning and metacognition implications for the design of computer-based scaffolds. instructional science, 33(5–6), 367–379. https://doi.org/10.1007/s11251-005-1272-9 bannert, m., reimann, p., & sonnenberg, c. (2014). process mining techniques for analysing patterns and strategies in students’ self-regulated learning. metacognition and learning, 9 (2), 161–185. https://doi.org/10.1007/s11409-013-9107-6 barab, s. a., bowdish, b. e., & lawless, k. a. (1997). hypermedia navigation: profiles of hypermedia users. educational technology research and development, 45(3), 23–41. https://doi.org/10.1007/bf02299727 bezdan, e., kester, l., & kirschner, p. a. (2013). the influence of node sequence and extraneous load induced by graphical overviews on hypertext learning. computers in human behavior, 29(3), 870–880. https://doi.org/10.1016/j.chb.2012.12.016 blom, h., segers, e., knoors, h., hermans, d., & verhoeven, l. (2018). comprehension and navigation of networked hypertexts. journal of computer assisted learning, 34(3), 306–314. https://doi.org/10.1111/jcal.12243 boekaerts, m. (1999). self-regulated learing: where we are today. international journal of educational research, 31, 445–457. https://doi.org/10.1016/s0883-0355(99)00014-2 de jong, t., & van joolingen, w. r. (1998). scientific discovery learning with computer simulations of conceptual domains. review of educational research, 68(2), 179–201. https://doi.org/10.3102/00346543068002179 delgado, p., vargas, c., ackerman, r., & salmeron, l. (2018). don’t throw away your printed books: a meta-analysis on the effects of reading media on reading comprehension. educational research review, 25(january), 23–38. https://doi.org/10.1016/j.edurev.2018.09.003 destefano, d., & lefevre, j.-a. a. (2007). cognitive load in hypertext reading: a review. computers in human behavior, 23(3), 1616–1641. https://doi.org/10.1016/j.chb.2005.08.012 eilam, b., & aharon, i. (2003). students’ planning in the process of self-regulated learning. contemporary educational psychology, 28(3), 304–334. https://doi.org/10.1016/s0361-476x(02)00042-5 goldman, s. r. (2009). explorations of relationships among learners, tasks, and learning. learning and instruction, 19(5), 451–454. https://doi.org/10.1016/j.learninstruc.2009.02.006 goldman, s. r., braasch, j. l. g., wiley, j., graesser, a. c., & brodowinska, k. (2012). comprehending and learning from internet sources: processing patterns of better and poorer learners. reading research quarterly, 47(4), 356–381. https://doi.org/10.1002/rrq.027 gorissen, c. j. j., kester, l., brand-gruwel, s., & martens, r. (2015). autonomy supported, learner-controlled or system-controlled learning in hypermedia environments and the influence of academic self-regulation style. interactive learning environments, 23(6), 655–669. https://doi.org/10.1080/10494820.2013.788038 greene, j. a., & azevedo, r. (2010). the measurement of learners’ self-regulated cognitive and metacognitive processes while using computer-based learning environments. educational psychologist, 45(4), 203–209. https://doi.org/10.1080/00461520.2010.515935 ignacio madrid, r., van oostendorp, h., & puerta melguizo, m. c. (2009). the effects of the number of links and navigation support on cognitive load and learning with hypertext: the mediating role of reading order. computers in human behavior, 25(1), 66–75. https://doi.org/10.1016/j.chb.2008.06.005 kalyuga, s., ayres, p., chandler, p., & sweller, j. (2003). the expertise reversal effect. educational psychologist, 38 (1), 23–31. https://doi.org/10.1207/s15326985ep3801_4 klois, s. s., segers, e., & verhoeven, l. (2013). how hypertext fosters children’s knowledge acquisition: the roles of text structure and graphical overview. computers in human behavior, 29(5), 2047–2057. https://doi.org/10.1016/j.chb.2013.03.013 kong, y., seo, y. s., & zhai, l. (2018). comparison of reading performance on screen and on paper: a meta-analysis. computers and education, 123(may), 138–149. https://doi.org/10.1016/j.compedu.2018.05.005 künsting, j., wirth, j., & paas, f. (2011). the goal specificity effect on strategy use and instructional efficiency during computer-based scientific discovery learning. computers and education, 56(3), 668–679. https://doi.org/10.1016/j.compedu.2010.10.009 lawless, k. a., brown, s. w., mills, r., & mayall, h. j. (2003). knowledge, interest, recall and navigation: a look at hypertext processing. journal of literacy research, 35(3), 911–934. https://doi.org/10.1207/s15548430jlr3503_5 lawless, k. a., & kulikowich, j. m. (1996). understanding hypertext navigation through cluster analysis. journal of educational computing research, 14(4), 385–399. https://doi.org/10.2190/dvap-de23-3xmv-9mxh macgregor, s. k. (1999). hypermedia navigation profiles: cognitive characteristics and information processing strategies. journal of educational computing research, 20(2), 189–206. https://doi.org/10.2190/1mec-c0w6-111h-yq6a malmberg, j., järvenoja, h., & järvelä, s. (2013). patterns in elementary school students′ strategic actions in varying learning situations. instructional science, 41(5), 933–954. https://doi.org/10.1007/s11251-012-9262-1 mcdonald, s., & stevenson, r. j. (1996). disorientation in hypertext: the effects of three text structures on navigation performance. applied ergonomics, 27(1), 61–68. https://doi.org/10.1016/0003-6870(95)00073-9 mobrand, k. a., & spyridakis, j. h. (2007). explicitness of local navigational links: comprehension, perceptions of use, and browsing behavior. journal of information science, 33(1), 41–61. https://doi.org/10.1177/0165551506068144 molenaar, i., & järvelä, s. (2014). sequential and temporal characteristics of self and socially regulated learning. metacognition and learning, 9(2), 75–85. https://doi.org/10.1007/s11409-014-9114-2 paans, c., molenaar, i., segers, e., & verhoeven, l. (2019). temporal variation in children’s self-regulated hypermedia learning. computers in human behavior, 96, 246–258. https://doi.org/10.1016/j.chb.2018.04.002 paans, c., segers, e., molenaar, i., & verhoeven, l. (2018). the quality of the assignment matters in hypermedia learning. journal of computer assisted learning, 34(6). https://doi.org/10.1111/jcal.12294 pieschl, s., stahl, e., murray, t., & bromme, r. (2012). is adaptation to task complexity really beneficial for performance? learning and instruction, 22(4), 281–289. https://doi.org/10.1016/j.learninstruc.2011.08.005 reimann, p. (2009). time is precious: variableand event-centred approaches to process analysis in cscl research. international journal of computer-supported collaborative learning , 4(3), 239–257. https://doi.org/10.1007/s11412-009-9070-z rotherham, a. j., & willingham, d. t. (2010). 21st century skills not new, but a worthy challenge. american educator, 34(1), 17–20. https://doi.org/10.1145/1719292.1730970 rouet, j.-f. (2003). what was i looking for? the influence of task specifity and prior knowledge on students’ search strategies in hypertext. interacting with computers, 15, 409–428. https://doi.org/10.1016/s0953-5438(02)00064-4 salmeron, l., kintsch, w., & cañas, j. j. (2006a). coherence or interest as basis for improving hypertext comprehension. information design journal, 14(1), 45–55. https://doi.org/10.1075/idj.14.1.06sal salmeron, l., kintsch, w., & cañas, j. j. (2006b). reading strategies and prior knowledge in learning from hypertext. memory and cognition, 34(5), 1157–1171. https://doi.org/10.3758/bf03193262 salmeron, l., naumann, j., garcía, v., & fajardo, i. (2017). scanning and deep processing of information in hypertext: an eye tracking and cued retrospective think-aloud study. journal of computer assisted learning, 33(3), 222–233. https://doi.org/10.1111/jcal.12152 scheiter, k., & gerjets, p. (2007). learner control in hypermedia environments. educational psychology review, 19(3), 285–307. https://doi.org/10.1007/s10648-007-9046-3 schoor, c., & bannert, m. (2012). exploring regulatory processes during a computer-supported collaborative learning task using process mining. computers in human behavior, 28(4), 1321–1331. https://doi.org/10.1016/j.chb.2012.02.016 segers, e., & verhoeven, l. (2009). learning in a sheltered internet environment: the use of webquests. learning and instruction, 19(5), 423–432. https://doi.org/10.1016/j.learninstruc.2009.02.017 sobocinski, m., malmberg, j., & järvelä, s. (2017). exploring temporal sequences of regulatory phases and associated interactions in lowand high-challenge collaborative learning sessions. metacognition and learning, 12(2), 275–294. https://doi.org/10.1007/s11409-016-9167-5 sonnenberg, c., & bannert, m. (2015). discovering the effects of metacognitive prompts on the sequential structure of srl-processes using process mining techniques. journal of learning analytics, 2(1), 72–100. https://doi.org/10.18608/jla.2015.21.5 urakami, j., & krems, j. f. (2012). how hypertext reading sequences affect understanding of causal and temporal relations in story comprehension. instructional science, 40(2), 277–295. https://doi.org/10.1007/s11251-011-9178-1 van der stel, m., & veenman, m. v. j. (2008). relation between intellectual ability and metacognitive skillfulness as predictors of learning performance of young students performing tasks in different domains. learning and individual differences, 18(1), 128–134. https://doi.org/10.1016/j.lindif.2007.08.003 van laer, s., & elen, j. (2018). towards a methodological framework for sequence analysis in the field of self-regulated learning. frontline learning research, 6(3), 228–249. https://doi.org/10.14786/flr.v6i3.367 veenman, m. v. j. (2011). learning to self-monitor and self-regulate. in r. mayer & p. alexander (eds.), handbook of research on learning and instruction. new york, ny: routledge. veenman, m. v. j. (2015). metacognition. handbook of individual differences in reading, reader, text, and context . https://doi.org/10.4324/9780203075562.ch3 winne, p. h., & hadwin, a. f. (1998). studying as self-regulated learning. in d. j. hacker, j. dunlosky, & a. c. graesser (eds.), metacognition in educational theory and practice (pp. 277–304). mahwah, nj: lawrence erlbaum. winne, p. h., hadwin, a. f., & gress, c. l. z. (2010). the learning kit project: software tools for supporting and researching regulation of collaborative learning. computers in human behavior, 26 (5), 787–793. https://doi.org/10.1016/j.chb.2007.09.009 winne, p. h., & nesbit, j. c. (2010). the psychology of academic achievement. annual review of psychology, 61, 653–678. https://doi.org/10.1146/annurev.psych.093008.100348 zimmerman, b. j. (1989). a social cognitive view of self-regulated academic learning. journal of educational psychology, 81(3), 329–339. https://doi.org/10.1037/0022-0663.81.3.329 zumbach, j., & mohraz, m. (2008). cognitive load in hypermedia reading comprehension: influence of text type and linearity. computers in human behavior, 24(3), 875–887. https://doi.org/10.1016/j.chb.2007.02.015 appendix a. flowchart for coding the macro-level navigation patterns. microsoft word wolgastproofs_aw.docx frontline learning research vol. 10 no. 1 (2022) 76 107 issn 2295-3159 corresponding author's current telephone and e-mail information: anett wolgast, university of applied sciences, lister straße 17, 30163 hannover, germany, telephone: 0049 511 533588 17, e-mail: anett.wolgast@gmail.com https://doi.org/10.14786/flr.v10i1.781 flexible social perspective taking in higher education and the role of contextual cues anett wolgast1 & yvonne barnes-holmes2 1martin-luther-university halle-wittenberg, germany 2ghent university, belgium article received 10 january 2021 / article revised 2 may 2022 / accepted 4 august / available online 19 august 2022 abstract being able to coordinate the perspectives of oneself and others is likely to be helpful in educational contexts. for example, teachers need flexible social perspective taking to understand their own perspectives and those of their students. evidence suggests that reading facilitates social perspective taking because it involves readers coordinating social perspectives. however, there is little evidence on actual flexible perspective taking in educational contexts. in the current research, we assumed that the presence of different spatial, temporal, and social cues with regard to (higher) educational contexts would affect flexible social perspective taking performances of prospective psychologists and teachers. across two different studies, we employed relational frame theory and a within-subject design (n = 44 undergraduate students in study 1 and n = 176 teacher education students in study 2). we analyzed the data by rasch-trees and general linear modeling. the results showed faster responding on flexible spatial and temporal social perspective taking tasks, involving a fictional college course in “english” rather than “statistics” (study 1). in study 2, the results suggested greater accuracy on flexible spatial and temporal social perspective taking tasks involving spatial rather than temporal relations (study 2). the results shed some light on the integration of different approaches for research on understanding the relevance of flexible social perspective taking in educational contexts. flexible spatial and temporal social perspective taking may be of benefit to both students in higher education and teachers in school education. keywords: undergraduate students; teacher education; social perspective taking; higher education; relational frame theory; spatial or temporal relational frames wolgast & barnes-holmes 77 | f l r 1. introduction social interactions and interpersonal understanding are core features of (collaborative) learning in higher education and in education at school. social perspective-taking seems to be especially important for interpersonal understanding (davis, 1980, 1983) and dealing with heterogeneous groups, cultural diversity, and inclusion in educational contexts (wilson et al., 2017). it is perspective-taking with reference to a human(-like) target and their situation that bolsters understanding of others’ behavior (davis, 1983) by stimulating a person to establish a mental representation of a social situation (i.e., mentalizing, engen & singer, 2013). social perspective-taking is an umbrella term for several forms of behavior and underlying processes rather than one specific skill (erle & topolinski, 2015). psychologists and teachers use social perspective taking extensively to respond sociallyappropriate in conjunction with their professional knowledge (gehlbach, brinkworth, et al., 2012; gehlbach, young, et al., 2012). psychologists mainly respond to one client in front of them. in contrast, teachers mainly respond to one or more students of a group in the class in school or higher education. one problem is that teachers have to keep many things in mind including their content, pedagogical-content, and pedagogical-psychological knowledge (gudmundsdottir & shulman, 1987; von aufschnaiter et al., 2015). they have to act according to their lesson preparation, as well as orientate to students’ traits and states, and learning goals planned for that lesson. for example, teachers have to constructively interact with students, and react appropriately when students co-construct their learning environment in the class (damşa, 2014; damşa et al., 2019). thus, teachers have to coordinate among themselves, their planned lesson contents, and the students. this coordination is common but subtle. teachers may instruct students to sit in teams facing one-another to facilitate active peer-to-peer learning rather than sitting in rows facing the teacher (wolgast, tandler, et al., 2020; wolgast & oyserman, 2019). in contrast, other teachers seem not to appreciate that students with their backs facing them see something different than the teachers (wolgast, tandler, et al., 2020; wolgast & oyserman, 2019). who has not experienced a teacher 'standing in the picture' in front of the board for too long? if the teacher takes into account the learners' view of the board or the smartboard, they are engaging in visuospatial social perspective taking and can move swiftly 'out of the picture'. in other words, seeing what students see requires teacher’s visuospatial social perspective taking. there are, however, very few published findings on factors that facilitate or impair a teacher’s consideration of a student’s momentary spatial position based on the student’s perspective in the classroom. this visuospatial social perspective taking required by the teacher is a basic feature that enables them to consider another individual’s angle of view (wolgast, tandler, et al., 2020). this study addresses the identified gap in published findings on factors that facilitate or impair a teacher’s consideration of a student’s momentary spatial position in the classroom. in particular, flexible social perspective taking involving spatial relations is an under-researched branch of (higher) education (wolgast, tandler, et al., 2020; wolgast & oyserman, 2019) although it might facilitate classroom management during teaching. one reason for the under-investigation might be that this form of perspective taking is difficult to test in classrooms in higher education and school. our rationale was that inter-individual differences exist in this type of perspective taking in psychologists and teachers even within relatively homogeneous or heterogeneous social groups (e.g., undergraduate students and teacher education students). moreover, we expected to observe intraindividual and inter-individual differences, depending on the presence of specific contextual cues, using tasks that had been employed in previous research (mchugh et al., 2007). the observation of such differences would contribute to existing research on the relationship between social behavior and different contexts, since these differences in flexibility may represent underlying non-observable processes in educational contexts such as teacher’s classroom management and social support. 1.1 previous research effective classroom management and teacher’s social support (hugener et al., 2009; lipowsky wolgast & barnes-holmes 78 | f l r et al., 2009) are two dimensions of teacher’s instructional quality that might require visuospatial social perspective taking. there are rarely research findings on this possible relationship. however, a consensus view from previous studies is that teacher’s instructional quality may impact students’ learning outcomes (kunter et al., 2013; lipowsky et al., 2009; rjosk et al., 2014). effective classroom management involves, for example, establishing clear classroom procedures, manage transitions between lesson sequences smoothly, keep track of students’ work, adapt their well-planned lessons accordingly, and intervene in students’ inappropriate behavior (evertson, 1989; kounin, 1970; lipowsky et al., 2009). effective classroom management provides space and time for cognitive engagement and learning of students (aloe et al., 2014; korpershoek et al., 2016; long et al., 2019). especially for smooth transitions between lesson sequences, it is valuable to examine prospective teachers’ flexible spatial and temporal social perspective taking. teachers often apply an individual reference norm during classes to support students (lohbeck & freund, 2021; marksteiner et al., 2021; rheinberg, 1977, 2001). when applying the individual reference norm, teachers need to make use of temporal relations to highlight, for example, a student’s learning gains from yesterday to today. teachers’ social support was positively related to their conceptual social perspective taking (gehlbach et al., 2016). conceptual social perspective taking involves seeing and understanding another’s current possible intention, aims, and resulting behavior, and this skillset bolsters our understanding of others’ behavior (gehlbach et al., 2015) by enabling us to mentalize a social situation (engen & singer, 2013). conceptual social perspective taking is necessary for effective communication, cooperation (johnson, 1975; mouw et al., 2020), and socially appropriate behavior (gehlbach, 2004; gehlbach, brinkworth, et al., 2012). conceptual social perspective taking, therefore, requires that one can shift flexibly between one’s own and another’s viewpoint visuospatially or conceptually (i.e., to understand someone else). some research has shown that teachers with high levels of conceptual social perspective taking are more effective as teachers (hyun & marshall, 1997; o’keefe & johnston, 1989) than those with relatively low conceptual social perspective-taking. the former might be more effective because accurately reading their students’ cues allows them to adjust their interactions appropriately (hunt, 1976). in contrast, test anxiety in teacher education students significantly predicted relatively low levels of their conceptual social perspective taking about six months later (wolgast, hille, et al., 2020). there is increasing evidence of statistics test anxiety among university students in social sciences (bourne, 2018; onwuegbuzie, 2004; siew et al., 2019; zeidner, 1991). significant differences further existed in response times on presented neutral (e.g., “plate”) vs threat-related words (e.g., “cancer”), with longer response times related to threatening words (bar-haim et al., 2007, p. 3). this effect is known as threat-related attentional bias and is often investigated in samples with low vs high anxiety scores (bar-haim et al., 2007). however, the relationship between a simple cue to statistics and performance in flexible social perspective-taking tasks has not been investigated in previous research. teaching experiences are often assessed by self-reports within standardized inventories (tschannen-moran & hoy, 2001). teachers’ self-reported behavior from complex classroom situations may be a subjective construction at a cognitive level of their cognitive/socio-emotional experiences in a lesson or, more from a bird’s view at a metacognitive level, it may be a form of self-reflection. it is difficult to disentangle these two possibilities through research on basic cognitive processes in teacher education students or teachers. one approach to explore cognitive rather than metacognitive processes is to present tasks instead of using self-report items to teacher education students or teachers. one possible means of assessing these processes is to activate core flexible spatial and temporal social perspective taking in tasks (barnes-holmes et al., 2004; mchugh et al., 2004a) that are adapted to classroom situations (willisch et al., 2021). in the research field on social perspective taking, there are conceptualizations and findings on linguistic aspects that stimulate either flexible social perspective taking towards another person or the mental focus on oneself (barnes-holmes et al., 2013; mchugh et al., 2004b; scanlon & barnes-holmes, 2013; wolgast et al., 2018; wolgast & barnes-holmes, 2018). for the acquisition of flexible spatial social perspective taking, for example, tasks have been used in which a person was asked to imagine two chairs and then to flexibly shift between their own point of view on the chair and another person’s wolgast & barnes-holmes 79 | f l r view from the other chair. when the person responded to a question about sitting there instead of here, they were able to solve the tasks correctly (barnes-holmes et al., 2004; mchugh et al., 2004a). furthermore, there is evidence that relations exist between flexible spatial and temporal social perspective taking in language and reality (barnes-holmes et al., 2013; mchugh et al., 2004b; scanlon & barnes-holmes, 2013; wolgast et al., 2018; wolgast & barnes-holmes, 2018). in summary, psychologists and teachers need social perspective taking to understand the behavior of people they work with. different forms of social perspective taking may be distinguished. for example, teaching might be facilitated by visuospatial perspective taking when a teacher presents learning materials in such a way that all students can see the materials. teaching and classroom management might further be related to conceptual social perspective taking that may be difficult within limited time frames of classroom situations. alternatively, flexible spatial and temporal social perspective taking might facilitate such demands of classroom management but is difficult to assess in educational contexts with a high degree of objectivity. the overarching aim of the current research is to examine whether contextual cues affect flexible spatial and temporal social perspective taking in undergraduate students and teacher education students. effects of the contextual cues on flexible spatial and temporal social perspective taking would demonstrate its contextual malleability. this malleability would suggest more or less flexible social perspective taking in perceived simple or challenging educational contexts. thus, the malleability would underline the nature of a context-related phenomenon and state in higher education (instead of a trait). the findings from the current research would provide an important foundation for further research on social perspective taking as one mental resource that probably helps in classroom management and other contexts in social professions. we focused on two educational contexts: (1) flexible spatial and temporal social perspective taking in undergraduate students with tasks describing situations in higher education courses, and (2) flexible spatial and temporal social perspective taking in teacher education students with tasks that describe classroom situations in the school. both studies represent basic research for facilitating teaching, classroom management, and improving teaching quality (hugener et al., 2009). established tasks assessing flexible social perspective taking have been constructed under the relational frame theory (barnes-holmes et al., 2004, 2013; hayes et al., 2001). we applied the relational frame theory and tested flexible spatial and temporal social perspective taking with regard to spatial and temporal relations in classes. 1.2 relational frame theory and deictic frames in assessments of flexible social perspective taking given that different social perspectives are represented in text material which has been shown to improve social perspective taking (cigala et al., 2015; montoya-rodríguez et al., 2017; mori & cigala, 2016), the current research attempted to capture the putative flexibility of spatial and temporal social perspective taking in different contexts with different samples. one approach to understanding the relationship between contexts and flexible spatial and temporal social perspective taking has emerged from behavioral psychology, especially from a functional-analytic account of language and cognition, known as relational frame theory (rft, see hayes et al., 2001). for rft, the acquired meanings and functions of words and social cues in a given language emerge as specific patterns of relational responding that include relating i to you. for rft, pronouns such as i, you, and they specify the perspectives they represent. in more technical terms, i is in a relation of coordination with the self, but is in most contexts in a relation of distinction from you, they etc., thus maintaining separations between self and others, and facilitating shifts in perspective. for example, when a speaker says “you”, the listener always responds from the perspective of i, whereas when a speaker says “i”, the listener responds from the perspective of you. for rft, this ability to shift from the perspectives of i and you is not only critical to perspective taking, but central to language itself (gore et al., 2010; hayes et al., 2001; montoya-rodríguez et al., 2017). these distinction relations, for example between i and you, are likely to be fundamental to flexible social perspective taking, because they allow one to distinguish between the perspective of self and the perspective of another, while being wolgast & barnes-holmes 80 | f l r able to adopt either or both in a given context (ballard et al., 1997). the distinction between i and you is fundamental to social perspective taking and known as the theory of mind (baron-cohen et al., 1985; wolgast, 2017; wolgast et al., 2018; wolgast & barnes-holmes, 2018) because it allows the person to distinguish between his or her own perspective and the perspective of another person and to switch between the two perspectives (mchugh et al., 2004a). in the language of rft, flexible social perspective taking is referred to as deictic relational responding and specifically involves: interpersonal relations (i and you), spatial relations (here and there), and temporal relations (now and then). it is important to emphasize that these relations interact with each other to create the complexity that comes to characterize having a perspective on oneself and others. that is, i is almost always coordinated with here and now, whilst you/they is almost always coordinated with there and then (levin et al., 2012). in addition, the pronouns are coordinated when a related word of the pronoun (e.g., a person, place, or time) is used (mchugh et al., 2004b). for rft, these patterns of relational responding get abstracted by the learner through a history of using language in these ways in certain contexts, thus enabling verbally-competent individuals to talk about events that have no actual basis in reality (e.g., talking about the future). there is also evidence that perspective taking is stimulated by the three key linguistic features: pronouns for interpersonal (i and you), spatial (here and there), and temporal relations (now and then, (mchugh et al., 2004a). in social interactions, these relations are linked together so that they produce directional meaning. in the context of social interactions, the directional meaning can stimulate people to take their own perspective or to think in terms of another person's perspective (levin et al., 2012). thus, pronouns stimulate different directions of thought and ideas from different perspectives. pronouns, such as you, there, then, stimulate thinking away from ourselves to other individuals (i.e., external frame of reference). the pronoun i stimulates thinking about oneself and directs thus the attention inwards (i.e., internal frame of reference). figure 1 shows the pronouns and their functional meaning for the internal and external frames of reference. researchers empirically tested these assumptions of rft in various studies (cigala et al., 2015; gore et al., 2010; levin et al., 2012; mchugh et al., 2004b; mori & cigala, 2016). in addition, findings from other research directions fit the assumptions of internal and external frames of reference and their importance for perspective taking (e.g., from cognitive psychology, brunyé et al., 2009; pickering et al., 2012). for example, reading several sentences written in the second person (you) stimulated the external frame of reference in tasks assessing social perspective taking. in contrast, reading the same content written in the first person (i) stimulated the internal frame of reference in tasks assessing social perspective taking (e.g., brunyé et al. 2009). fig. 1 pronouns and their function wolgast & barnes-holmes 81 | f l r there is already evidence from many studies that supports rft’s account of flexible social perspective taking as deictic relational responding (cigala et al., 2015; gore et al., 2010; levin et al., 2012; mchugh et al., 2004b; mori & cigala, 2016). for example, in a rft-based protocol designed to assess flexible social perspective taking, researchers separated out the various types of relational patterns that comprise perspective taking, and distinguished among different levels of relational complexity (mchugh et al., 2004a). consider a typical task that was presented without pictures, where the instruction to a child was: “if i am sitting here on a blue chair and you are sitting there on a black chair”, followed by the two questions “where are you sitting? where am i sitting?” answering the questions correctly was argued to involve responding in accordance with interpersonal relations (i and you) and spatial relations (here and there), and was simple at the level of relational complexity because neither the interpersonal relation nor the spatial relation was reversed. now consider the following task that required responding in accordance with temporal relations: “yesterday you were watching television, today you are reading. if now was then and then was now: what would you be doing then? what would you be doing now?” this task was denoted as a more complex reversed task because the temporal relation is reversed in the phrase “if now was then and then was now”. in a study involving adults, researchers reported greater accuracy on simple vs reversed tasks, but no significant difference was observed between responding to spatial vs temporal relations (mchugh et al., 2004b). 2. the present research inspired by the outlined existing literature on clearly defined phenomena and the complex reality in school, we started at possible roots of considering all students in the classroom and asked whether tiny cues may affect flexible spatial and temporal social perspective taking in different samples in higher education. higher education and school are characterized by different subjects that are often assigned to the mathematical vs verbal domains. these domains are related to many educationalrelevant phenomena (e.g., self-concept, motivation, test anxiety, möller et al., 2009; siew et al., 2019; wolgast, hille, et al., 2020; zeidner, 1991). in germany, prospective teachers at university may choose the subjects (e.g., french and spanish; mathematics and sports) for their later work in a secondary school and have to study mainly these subjects. thus, possible effects of their diverse subject combinations on flexible spatial and temporal social perspective taking with regard to different subjects might be hardly psychometrically controlled. in contrast, all psychology undergraduate students have to study statistics that is obviously part of the mathematical domain. that fact suggested the starting point to examine whether contextual cues to statistics (mathematical domain) vs english (verbal domain) affect undergraduate students’ flexible spatial and temporal social perspective taking, before examining spatial and temporal flexible social perspective taking in teacher education students (i.e., prospective teachers). in this paper, we present therefore two different studies: (1) examining subject-related contextual cues on undergraduate students’ flexible spatial and temporal social perspective taking and (2) examining space and timerelated contextual cues on teacher education students’ flexible spatial and temporal social perspective taking. to our knowledge, there are no published findings on undergraduate students’ and teacher education students’ flexible spatial and temporal social perspective taking in exemplary classroom situations. our specific research questions were as follows: (1) is flexible spatial and temporal social perspective taking influenced by specific contextual cues, such as references to a fictitious course as “english” vs “statistics”? (2) does the involvement of spatial vs temporal relations facilitate different performances on the flexible spatial and temporal social perspective-taking task? thus, this research focused on flexibility in spatial and temporal social perspective taking within an rft conceptual framework. the aims of the two studies were as follows. (1) to investigate flexible spatial and temporal social perspective taking with regard to a fictitious college-based “english” vs “statistics” course with undergraduate students in germany. (2) to explore the role of spatial and temporal relations in flexible wolgast & barnes-holmes 82 | f l r spatial and temporal social perspective-taking tasks involving statements about teaching a fictitious “maths” class presented to teacher education students in germany. our hypothesis was (1) that the presence of “english” vs “statistics” as a contextual cue would result in faster and more accurate flexible spatial and temporal social perspective taking on the contextual cue “english” relative to “statistics”, because “english” would be deemed easier. (2) for study 2, we had no hypothesis. we explored differential performances in flexible spatial and temporal social perspective taking with the involvement of spatial vs temporal relations because we were somewhat uncertain as to the nature of this potential difference. to test our hypothesis, we conducted power analyses using the r package pwr (champely, 2017) and effect sizes taken from previous research. for example, researchers compared response times on tasks involving relational responses to “self”, “other”, and to a photograph as target stimuli (we used a similar format here, see supplemental file a), resulting in an effect of a cohen’s d = .79 in favor of “self” responses (mchugh et al., 2007). in study 1, therefore, the flexible spatial and temporal social perspective-taking task with different contextual cues (“english” vs “statistics”) presented 48 tasks and required at least n = 16 participants to disclose effects (cohen’s d = 0.79, significance level = .05, power = .80, type = “one.sample”) within participants (champely, 2017). study 2 had an exploratory nature with a within-subject design between the flexible spatial and temporal social perspective-taking tasks, such that an a priori power analysis was not indicated. we conducted instead a post hoc power analysis (see 4.2 results). all of the current experimental work had received ethical approval from the relevant committee and was conducted accordingly. 2.1 general procedure each participant engaged with an individual laptop for e-exams. all aspects of the procedure were automated and presented via psychopy (peirce et al., 2019). all participants received instructions for logging in and commencing the study at the same time in studies 1 and 2. they were verbally instructed: “follow the instructions on the screen!” these were as follows: “welcome! in the following, you will read statements and questions. below each question, you will see two words. choose one of these words to answer the question. press “n” to move forward!” prior to the first task, the following instruction appeared: “it will start right away. answer as accurately and quickly as you can! press “n” to start!” on each task, participants emitted a response by pressing the n key (on the left of the keyboard), or m (on the right) to select the response option displayed on that side of the screen. on all tasks, the first statement was centrally displayed in white characters on a dark grey background, with the question below presented in identical format. approx. 2cms below the question were the two response options, also in white characters approx. 10cm apart. on all tasks, “i” as it appeared on-screen referred to the computer’s perspective and “you” referred to the participant’s perspective. on all simple tasks, the correct answer matched the task statement, whereas on all reversed tasks, the correct response involved reversing the target relation (essentially responding incorrectly). the test tasks were presented in random order. a white fixation cross was displayed between each task. participants could not skip any task, nor could they return to a previous task. each task remained on-screen until a response was emitted. no time limit was applied on any flexible spatial and temporal social perspective-taking task. no feedback or consequences for responding were provided on any task. completion of the last task marked the end of the participation. 2.2 materials in the current two studies, we adapted flexible spatial and temporal social perspective-taking tasks from previous research (mchugh et al., 2004b) to contexts familiar to higher education. we adapted the established tasks (mchugh et al., 2004b) to the current studies, based on videos of actual maths classes (rakoczy et al., 2005), situations in maths classes described in the literature (rose, 2018), the german teacher education standards (hrk, 2015; kmk & hrk, 2015), and the qualifications framework for german higher education degrees (bartosch & grygar, 2019; hrk et al., 2011; kmk wolgast & barnes-holmes 83 | f l r & hrk, 2015). the tasks describe different situations where individuals reflect on their competence levels by perceived academical challenges or not and applying the individual reference norm (rheinberg, 1977) to themselves or to others. most of the tasks provide the temporal relations yesterday-today and spatial relations here-there necessary for applying the reference norm except one task which describes i-you relations and spatial relations. for testing effects of contextual cues within the tasks in study 1, we adapted the tasks to the mathematical vs verbal academic domain in higher education with replacing “maths class” by “statistics course” vs “english course” respectively. in this way, we examined flexibility in spatial and temporal social perspective taking in the presence of different cues that specified various contexts. all aspects of the research and its materials were presented in german, but are translated into english for current purposes. the protocol presented in study 1 comprised 48 test tasks, all adapted from previous research (mchugh et al., 2004b). there were 24 tasks that referred to english and 24 almost identical tasks that referred to statistics. 1each set of tasks contained a mix of spatial and temporal relations in simple and reversed form. the reader is strongly advised to consult supplemental file a which contains all 24 statistics tasks, where each question represents one task. the 24 english tasks were identical, but referred to english rather than statistics. each statement shown in supplemental file a was presented twice, with each exposure containing one of the two relevant questions (see supplemental file c for all tasks in german). all aspects of the apparatus from the previous study were identical in study 2. the format of all tasks presented in study 2 was identical to the previous study. in study 2, the questions pertaining to each task contrasted the perspective of i (the computer) with you (participant) or with students (others). study 2 presented a total of 15 tasks, which contained spatial tasks and temporal tasks, all in reversed form. the reader is strongly advised to consult in supplemental file b which contains all 15 tasks, separated by task-type. each task was presented once, accompanied by a single relevant question. 2.3 statistical analyses revelle (2019) provides a tool for analyzing internal consistency of multidimensional constructs by mcdonald’s ω for binary data (r package psych, see also https://www.personalityproject.org/r/html/omega.html for details). we used mcdonald’s ω to measure internal consistency and structure of the tasks used in the studies 1 and 2 (dunn et al., 2014; revelle, 2019). as the responses on the tasks were binomial (correct/incorrect response), we ran a rasch-tree model (strobl et al., 2016) to detect different item functioning of responses and irregular response behavior reflecting low test-taking motivation. the rasch tree has been used in previous research for detecting potentially different item functioning of responses (strobl et al., 2015). responses and response times were included in previous rasch tree analyses for detecting low test-taking motivation in participants (ranger & kuhn, 2017). we applied the rasch-tree model within the r environment (r development core team, 2009) by the r package psychotree (strobl et al., 2016). the r codes can be obtained from the corresponding author. to test mean differences in responses (i.e., mean accuracy) and response times (depending on the cue “english” vs “statistics” in study 1, and depending on the presence of spatial vs temporal relations in study 2), we used general linear models (glms) for repeated measures. 3. study 1 the aim of study 1 was to examine any potential differences between responding to flexible 1we included 24 tasks each for english and statistics after using revelle’s (2018) tool for a two-factor solution that resulted in an ω = .70, suggesting acceptable internal consistency of both factors. wolgast & barnes-holmes 84 | f l r spatial and temporal social perspective-taking tasks that referred to english vs statistics. there is increasing evidence of statistics anxiety among students in social sciences in higher education (bourne, 2018; onwuegbuzie, 2004; siew et al., 2019; zeidner, 1991). thus, we predicted superiority (i.e., higher accuracy, lower latency) in responding to the set of tasks involving the cue english over the set of tasks involving the cue statistics. in other words, our simple question was whether a single cue in one area over another would influence the accuracy or response time of flexible spatial and temporal social perspective taking in that domain? 3.1 participants and setting forty-four undergraduate students (one male and 43 females, mage = 22 years) attended for a university lecture in germany. recruitment was part of a university module in psychology, but was undertaken voluntarily. study 1 was entirely conducted in an e-exam hall at the relevant university. two researchers were present at all times. each participant was randomly assigned to an individual desk (approx. 2m apart), at which they waited until all participants were seated. participants completed the tasks in a mean of 8.37 minutes (sd = 0.30) and waited in their seats until all participants had finished. 3.2 results initially, we excluded all response time outliers with a sd > 1.5 above/below each task mean (semmelmann & weigelt, 2017). this left n = 41 participants. table 1 provides m and sd of response times and accuracy on the flexible spatial and temporal social perspective taking tasks. mean accuracy on english tasks vs statistics was not significant f(1, 40) = 1.15, p = .29 (see table 1 for m and sd). however, the mean latency on statistics was significantly longer than english, f(1, 40) = 160.56, p < .001, = .80, cohen’s d = 1.96 (cohen, 1988), with a large effect (lenhard & lenhard, 2016). mean accuracy on ‘spatial and spatial-temporal tasks’ was .51 vs ‘temporal tasks’ at .50, thus the difference was not significant f(1, 40) = 0.42, p = .52. however, the mean latency on ‘spatial and spatial-temporal tasks’ was 2.07 (see table 1 for sd), and on ‘temporal tasks’ was 1.90, which was significant f(1, 40) = 68.32, p < .001, = .63, d = 1.26. the added product term ‘course’ × ‘relations’ suggested no interaction effect between the tasks including the cue english vs statistics course and the tasks including the cues ‘spatial and spatial-temporal tasks’ vs ‘temporal tasks’ t(40) = .66, p = .51. there was, however, an interaction effect between the latency on the tasks including the cue english vs statistics course and the latency on the tasks including the cues ‘spatial and spatialtemporal tasks’ vs ‘temporal tasks’ t(40) = 4.63, p < .001, d = 0.82. 2 ph 2 ph wolgast & barnes-holmes 85 | f l r table 1 study 1: undergraduate students’ mean response accuracy and response time on flexible social perspective-taking tasks including the contextual cues to an english course, statistics course, ‘spatial, spatial-temporal’ and ‘temporal’ relations study 2: teacher education students’ mean response accuracy and response time on flexible social perspective-taking tasks including the contextual cues to teaching a maths class, reversed spatial and reversed temporal relations note. means in brackets significantly differ from each other (i.e., p < .05). fig. 2 study 1: mean response time on flexible spatial and temporal social perspective taking tasks including cues to courses (english vs statistics) and spatial-temporal/temporal relations with 95% confidence intervals course relations accuracy response time m sd m sd study 1 english .52 .12 1.81 .28 statistics .48 .10 2.15 .26 spatial, spatial-temporal .51 .10 2.07 .27 temporal .49 .11 1.90 .26 english spatial, spatial-temporal .53 .17 1.87 .34 temporal .52 .16 1.74 .27 statistics spatial, spatial-temporal .50 .14 2.26 .26 temporal .47 .14 2.05 .29 study 2 maths reversed spatial .94 .16 2.24 .37 reversed temporal .90 .18 2.21 .30 wolgast & barnes-holmes 86 | f l r we additionally employed a rasch-tree analysis (strobl et al., 2015, 2016) to assess for differential item functioning and low test-taking motivation (ranger & kuhn, 2017), but this was not supported (see figure 3). the rasch-tree analysis including all 48 responses yielded a single node that indicated equivalent item functioning of responses on spatial and temporal relations (strobl et al., 2016). the rasch-tree analysis including the 48 response times as covariates (see ranger & kuhn, 2017, for details) on the corresponding 48 responses yielded again the single node (see figure 3) that indicated no differential test-taking motivation patterns. fig. 3 study 1: the single node from rasch tree analysis (indicating no differential item functioning and no differential test-taking motivation patterns) with 48 end nodes representing estimates of task difficulty (y-axis) for each task (x-axis) 4. study 2 given that performances on spatial vs temporal relations did not appear to vary when these tasks were presented in simple form in study 1, we queried in study 2 whether differences might be recorded when spatial and temporal relation types were presented in reversed form. previous evidence indicates that reversing both types of relation does increase the difficulty of the flexible social perspective taking task, but there is little or no research on whether the observed superiority of spatial over temporal relations remains when each is reversed (mchugh et al., 2004a, 2004b). however, it is important to note in advance that the construction of the tasks and our statistical analyses did not permit this type of strict comparison of relation type, but did allow us some level of interesting comparison in terms of relation type (i.e., flexible spatial vs temporal social perspective taking tasks). 4.1 participants and setting one hundred and seventy-six teacher education students (49 males and 127 females, mage = 24 years) participated in study 2 online via qualtrics for course credit in germany. all participation was undertaken voluntarily. participants completed the tasks in approx. 23 minutes (m = 22.65, sd = 3.25) 4.2 results again, we excluded all response time outliers with a sd > 1.5 above/below each task mean (semmelmann & weigelt, 2017). this left n = 149 participants. note that all tasks were reversed tasks wolgast & barnes-holmes 87 | f l r (see supplemental file b). table 1 presents m and sd for the reversed spatial tasks and reversed temporal tasks, thus the spatial task mean was significantly higher than the temporal task mean, f(1, 148) = 7.93, p = .01, = .05, d = 0.46, with a moderate effect. the mean latency on reversed spatial tasks was 2.24 (sd = 0.37) and on reversed temporal tasks was 2.21 (sd = 0.30), which was non-significant f(1, 148) = 3.28, p = .07. again, we employed a rasch-tree analysis by including all flexible reversed spatial and temporal social perspective taking tasks in a rasch-tree model to assess for differential item functioning and differential test-taking motivation. the analysis yielded one single node that suggested measurement invariant item functioning between the two types of relations (i.e., flexible spatial vs temporal social perspective taking tasks), also including response times as covariates indicating no differential response time patterns (see figure 4). finally, we conducted a post hoc power analysis to obtain information on the power of study 2 which yielded 86% power (n = 176 participants, cohen’s d = 0.46, significance level = .05, type = “one.sample”, champely, 2017). fig. 4 study 2: the single node from rasch tree analysis (indicating no differential item functioning) with 15 end nodes representing estimates of task difficulty (y-axis) for each task (x-axis). 5. general discussion psychologists and teachers need flexible spatial and temporal social perspective taking for fast shifting on spatial and temporal relational perspectives for understanding diverse individual’s and heterogeneous group’s behavior. spatial relational perspectives may direct teacher’s attention to student's positions in the classroom. temporal relational perspectives often direct teacher's attention to actions or reactions at a certain time point. the aim of the current research was to examine whether contextual cues affect flexible spatial and temporal social perspective taking in prospective psychologists (i.e., undergraduate students) and teachers (i.e., teacher education students). we assessed flexible spatial and temporal social perspective taking by established tasks (mchugh et al., 2004b). these tasks include contextual cues to classroom situations and can be presented in higher education classes. we focused on two educational contexts: (1) flexible spatial and temporal social perspective taking in undergraduate students with tasks describing situations in higher education courses, and (2) flexible spatial and temporal social perspective taking in teacher education students with tasks that describe classroom situations in school. this research was basic research for facilitating teaching, classroom management, and improving teaching quality (hugener et al., 2009). 2 ph wolgast & barnes-holmes 88 | f l r 5.1 flexible social perspective taking within the rft in the current two studies, we applied rft (barnes-holmes et al., 2004) to classes in higher education and tested the hypothesis (1) that the presence of “english” vs “statistics” as a contextual cue results in faster and more accurate flexible spatial and temporal social perspective taking on the former cue than the latter cue. then, we followed the exploratory question (2) whether the involvement of spatial vs temporal perspective relations can facilitate different performances on the flexible spatial and temporal social perspective taking tasks? prerequisites for analyzing the collected data are motivation reflected in regular response times and item functioning, that we examined by rasch-tree analyses. we applied rasch-tree analyses to detect different item functioning and low test-taking motivation, and glms in both studies to test the hypothesis (1) in study 1 and answer the explorative question (2) in study 2. in study 1, the single rasch tree node suggested measurement invariant item functioning between tasks involving cues for an “english” course vs a “statistics” course. in case of different item functioning, the rasch tree would yield more than one node (strobl et al., 2016). in study 2, the single node suggested measurement invariant item functioning between tasks involving reversed spatial vs temporal relations. taking response times into account in the rasch tree allowed us to analyze if irregular response processes occurred (e.g., due to low test-taking motivation, see ranger & kuhn, 2017). however, we did not find irregular response processes by the rasch tree analysis in study 1 or study 2. apparently, there is an odd difference in the total response times on the 48 tasks in approx. eight minutes in study 1 vs 15 tasks in approx. 23 minutes in study 2. we can only speculate why the total response times differed in this way. indeed, the tasks presented differed because study 2 involved only reversed perspective relations. undergraduate students’ accuracy (solving probability 48–51%) was at a moderate level in study 1. that is, the tasks seemed to be more difficult for the undergraduate students in study 1 than the teacher education students in study 2 (solving probability 90–94%). the teacher education students responded very accurately, representing high flexible spatial and temporal social perspective taking, even though the tasks with reversed spatial vs temporal relations in study 2 should be more difficult than the mainly simple spatial and temporal relations in study 1. on the other hand, the teacher education students’ total response times (approx. 23 minutes) were fairly long. a person can respond slowly on tasks in order to make as few mistakes as possible or they can respond quickly despite the risk of mistakes. this is called the speed-accuracy trade-off (wickelgren, 1977; zimmerman, 2011). in an ideal world, a person strives for maximum performance on both components. the undergraduate students in study 1 seemed to respond as quickly as possible despite the risk of more mistakes. their response and response time means (see table 1) might suggest rapid guessing, however, the response times differed depending on the contextual cues. the undergraduate students might have focused on the cues without reflecting fully on the situation described in each task, and were thus able to respond quickly (e.g., focusing on i-you and herethere relations). thus, they possibly decided just on these cues as criteria without considering the described context. in study 2, the teacher education students seemed to respond slowly in order to increase their overall accuracy level. they may have connected more with the situation presented in the tasks than the undergraduate students in study 1. thus, they possibly decided on the context as criterion. we examined the effects of using the single words “english” or “statistics” as cues on flexible spatial and temporal social perspective taking tasks, and our results supported our hypothesis in part because the undergraduate students responded more quickly on “english” tasks over “statistics”, although they responded with similar accuracy on the flexible spatial and temporal social perspective taking tasks, including the cue “english” vs “statistics”. the significantly faster responses on tasks involving the cue “english” can be explained with findings from other studies suggesting that reading activates mental representations (o’brien & albrecht, 1992). moreover, the longer response times on the tasks including the cue “statistics” than “english” might result from statistics anxiety even in undergraduate students who already had to pass several statistics courses, according to the previous findings from that field (bourne, 2018; onwuegbuzie, wolgast & barnes-holmes 89 | f l r 2004; onwuegbuzie & seaman, 1995; siew et al., 2019; zeidner, 1991). note, we did not assess statistics anxiety with an established inventory and can only speculate that the undergraduate students already perceived the cue “statistics” as a threat. if so, this threat may affect their response times with longer response latencies related to the cue “statistics” relative to the cue “english” in terms of the thread-related attentional bias (bar-haim et al., 2007). we also examined whether the presentation of spatial vs temporal relations within the task would influence the flexible social perspective-taking performances. the analyses yielded longer latencies on ‘spatial and spatial-temporal’ relations over ‘temporal’ relations. this result might reflect the focus on the cues rather than the context because here-there relations required encoding of several words describing the context while just looking for the cues for answering the question “where…?”. in contrast, yesterday-today (i.e., now-then relations in rft) provided the necessary information to answer the question “when…?” (see supplemental file a). the interaction effect of the cues to the english vs statistics course × the ‘spatial and spatialtemporal’ vs ‘temporal’ relations in the tasks on the corresponding response times showed that the relations moderated the influence of the cues to the english vs statistics course on the corresponding response times. it is an ordinal interaction (loftus, 1978). the main effects are interpretable (as discussed above) and the ordering of the data points corresponding to the levels ‘english’ vs ‘statistics’ of the independent variable ‘course’ depends on the level ‘spatial and spatial-temporal’ vs ‘temporal’ of the independent variable ‘relations’. this finding suggests that the ‘spatial and spatial-temporal’ relations additionally extended the response time on tasks involving the cue “statistics”. the results of study 2 suggested effects of spatial over temporal relations on flexible spatial and temporal social perspective taking performances. the teacher education students responded significantly more accurately on the tasks involving spatial rather than temporal relations. this result suggests the assumption that the teacher education students used another response strategy (e.g., imagining the classroom situation from the perspective of students) to solve the tasks, relative to the undergraduate students. the teacher education students seemed to decide on the context as criterion rather than specific cues and put themselves in the complex context of the described classroom situation. the effects of the contextual cues on flexible spatial and temporal social perspective taking demonstrated its contextual malleability. this malleability suggests more or less flexible spatial and temporal social perspective taking in perceived simple or challenging educational contexts. thus, the malleability underlines the nature of a context-related phenomenon and state in higher education (instead of a trait). the presented findings provide an important foundation for further research on social perspective taking as one mental resource that likely helps in higher education and school (e.g., in classroom management, social support, or peer learning). 5.2 flexible social perspective taking in educational contexts from an educational research perspective, the results indicate that flexible spatial and temporal social perspective taking performances differ between higher education students. the results also indicate different flexible spatial and temporal social perspective taking performances within higher education students from one to another educational situation described in the tasks. observing these differences suggests that the teacher education students (i.e., prospective teachers) each constructed and updated an individual meaning to the situations briefly described in the tasks. some tasks (presented to the prospective teachers) describe teachers who reflect on their competence levels by perceived academic demands or positive experiences with students in the educational space of a maths class (incl. the spatial relation here-there). the spatial relations require to shift between here and there similar to real teaching situations where the teacher has to shift between the subjective point of view here and the students there in an educational space (e.g., classroom). other tasks describe teachers who apply the individual reference norm (rheinberg, 1977) to themselves or to students. these tasks include the temporal relation yesterdaytoday which is necessary for applying the individual reference norm. wolgast & barnes-holmes 90 | f l r constructing an individual meaning and flexible spatial and temporal social perspective taking seemed difficult for several prospective teachers. the prospective teachers might also have difficulties in seeing what real students see (e.g., after transitions between lesson sequences) and if they see the learning materials as the teacher intended. if students do not see the learning material as the teacher intends (wolgast, 2017; wolgast, tandler, et al., 2020; wolgast & oyserman, 2019), they may feel anger and disrupt the lesson. disruptive student behavior is often a challenge for teacher classroom management (lipowsky et al., 2009). in addition, prospective teachers who have difficulties in applying flexible spatial and temporal social perspective taking may have difficulties in applying the individual reference norm to students in a real classroom. the teacher's use of the individual reference norm is important because it was positively related to student academic motivation (klopsch et al., 2022; rheinberg, 2001) and negatively related to cheating in school (marksteiner et al., 2021). teaching situations are complex. frequently used flexible spatial and temporal social perspective taking may facilitate the consideration of all students with diverse learning backgrounds (scanlon & barnes-holmes, 2013; wolgast, tandler, et al., 2020). according to the results presented, contextual events may influence flexible spatial and temporal social perspective taking. whether flexible spatial and temporal social perspective taking can be stimulated in teacher trainings, even under possible influences of contextual events, remains an open question for future research. 5.3 limitations the participants here represent self-selected samples tested in groups (one person at one laptop) or online, thus restricting the generalizability of the findings. without a probabilistic sample, additional variables (e.g., socio-economic background, general intelligence, current distress or anxiety) may well have contributed to any observed variability in the flexible perspective taking performances. as noted above, the gender distribution was not equal in the studies; only one male participated in study 1. indeed, our flexible spatial and temporal social perspective taking tasks were language-based (i.e., statements), in contrast to other tasks used in perspective taking research (erle & topolinski, 2015; janczyk, 2013; wolgast, tandler, et al., 2020; wolgast & oyserman, 2019). nevertheless, the studies contribute to the understanding of the potential relationship between contextual cues and tested social perspective-taking processes (not only self-reports) with regard to educational contexts. 5.4 the implications of the study and findings for educational research and practice the current study provides insights into presumed underlying processes of learning in higher education and student orientation of prospective teachers. based on the findings of this study, students' flexible spatial and temporal social perspective taking varies between them, and within them from one educational context to another. future research might investigate the potential benefits to learners and teachers of establishing flexible social perspective taking. the flexibility of social perspective taking was assessed by tasks, rather than by subjective measures such as self-reports. future research may consider the possible training effects that could be obtained with this type of task in terms of facilitating social perspective taking in real educational situations. teachers may facilitate students’ learning by an orderly classroom atmosphere with few disruptions and discipline problems (lipowsky et al., 2009). ideally, a teacher’s attention flexibly alternates between the teacher's own planned teaching actions and each student, depending on the situation and social priority. the teacher should alternate their attention while applying professional knowledge in order to respond to learners. flexible spatial and temporal social perspective taking may help them to shift between personal classroom experience and students’ situations (scanlon & barnesholmes, 2013). wolgast & barnes-holmes 91 | f l r moreover, flexible spatial and temporal social perspective taking can be used to measure intervention effects of other trainings in educational contexts (e.g., educational simulations, cigala et al., 2015; holmes, 2019; scanlon & barnes-holmes, 2013). various further training approaches including the flexible spatial and temporal social perspective taking tasks are possible with relatively little effort. one training approach relates to teachers' attitudes toward learners with special needs (scanlon & barnes-holmes, 2013). before and after an acceptance and commitment training, teachers' attitudes, self-efficacy expectations, and stress experience were assessed with a questionnaire. at both time points, testing occurred to determine the extent to which teachers engaged in flexible spatial and temporal social perspective taking among learners with different characteristics (scanlon & barnesholmes, 2013). compared to the first measurement, teachers showed statistically significantly more positive attitudes, higher self-efficacy expectations, and flexible spatial and temporal social perspective taking toward learners with special needs after the training. in addition, teachers reported statistically significantly higher stress levels before the training than after (scanlon & barnes-holmes, 2013). accordingly, flexible spatial and temporal social perspective taking presumably facilitates consideration of all students with diverse learning backgrounds. this assumption might be tested in further research. however, the different constructed meanings when reading about situations in a maths class imply differences in the prospective teachers’ flexible spatial and temporal social perspective taking in educational situations. difficulties with one or both forms may hinder learning in higher education, giving appropriate social support in psychological or educational contexts, and in managing groups or classes (scanlon & barnes-holmes, 2013; wolgast, tandler, et al., 2020; wolgast & oyserman, 2019). we recommend to support especially prospective teachers with difficulties in applying flexible spatial and temporal social perspective taking to improve both forms for considering each student’s angle of view in heterogeneous classes and for the equal inclusion of all students. flexible spatial and temporal social perspective taking could be practiced in trainings for (prospective) teachers in educational contexts. discussing their experiences of such trainings might help to stimulate flexible spatial and temporal social perspective taking and the related social orientation to others instead of only focusing oneself. several intervention studies can be used to develop workshops that are particularly suitable for teachers and their consideration of the students’ physical positions in the class (scanlon & barnes-holmes, 2013; wolgast, tandler, et al., 2020; wolgast & oyserman, 2019). teachers need such skills and strategies that subjectively facilitate teaching and prevent exhaustion and burnout. in addition, positive teacher-student interactions with all learners can be expected if all learners feel equally considered and supported by a teacher, for example. this consideration could strengthen positive student-teacher interactions and stabilize a social atmosphere conducive to learning. moreover, flexible spatial and temporal social perspective taking might provide space and time for learners’ co-constructions and relational perspectives in virtual environments (damşa et al., 2019). 5.5 conclusion flexible spatial and temporal social perspective taking appears to be essential to psychologist’s and teacher’s understanding and acceptance of diverse learners. we see the concepts of flexible spatial and temporal social perspective taking as a prerequisite for addressing the learning needs of all learners in the classroom. the impact of the current work involves new findings about the state flexible spatial and temporal social perspective taking and its malleability in higher education. the malleability suggests that flexible spatial and temporal social perspective taking can be stimulated in an intervention. in teacher training, the mental moving away from the own person and approaching perspectives of learners should be trained supra-disciplinarily and subject-didactically. the same applies to further training courses aimed at teachers, school psychologists and further educators. we presented insights in prospective psychologists’ and teachers’ flexible spatial and temporal social perspective taking in higher education. the tasks presented here and further developed tasks of wolgast & barnes-holmes 92 | f l r this type can be implemented in interventions to increase accuracy and reduce latency in social coordination processes in educational contexts. training programs should aim at the fact that psychologists and teachers routinely consider, understand and support students with different and unfamiliar learning preconditions in the classroom. for example, texts with descriptions of the same classroom situation from the perspective of teachers and from the perspective of different students are suitable for this purpose. if teachers also regularly participate in training, they put themselves in the role of learners and experience teaching-learning situations from a different perspective. keypoints in two studies, we employed the relational frame theory, a within-subject design with 220 participants, and analyzed the data by rasch-tree and general linear modeling. the results showed faster responding on flexible social perspective-taking tasks, involving a fictional college course in “english” rather than “statistics” (study 1). participants responded more accurate on flexible social perspective-taking tasks involving spatial rather than temporal relations with regard to a maths class (study 2). the results shed some light on the integration of different approaches for research on flexible social perspective taking and learning in educational contexts. references aloe, a. m., amo, l. c., & shanahan, m. e. (2014). classroom management self-efficacy and burnout: a multivariate meta-analysis. educational psychology review, 26(1), 101–126. https://doi.org/10.1007/s10648-013-9244-0 ballard, d. h., hayhoe, m. m., pook, p. k., & rao, r. p. n. (1997). deictic codes for the embodiment of cognition. behavioral and brain sciences, 20(4), 723–767. https://doi.org/10.1017/s0140525x97001611 bar-haim, y., lamy, d., pergamin, l., bakermans-kranenburg, m. j., & van ijzendoorn, m. h. (2007). threat-related attentional bias in anxious and nonanxious individuals: a meta-analytic study. in psychological bulletin (vol. 133, issue 1, pp. 1–24). american psychological association. https://doi.org/10.1037/0033-2909.133.1.1 barnes-holmes, y., foody, m., barnes-holmes, d., & mchugh, l. (2013). advances in research on deictic relations and perspective-taking. in advances in relational frame theory: research and application. (pp. 127–148). new harbinger publications, inc. barnes-holmes, y., mchugh, l., & barnes-holmes, d. (2004). perspective-taking and theory of mind: a relational frame account. the behavior analyst today, 5(1), 15–25. https://doi.org/10.1037/h0100133 baron-cohen, s., leslie, a. m., & frith, u. (1985). does the autistic child have a “theory of mind” ? cognition, 21(1), 37–46. https://doi.org/10.1016/0010-0277(85)90022-8 bartosch, u., & grygar, a.-k. (2019). hochschulbildung mit kompetenz. eine handreichung zum qualifikationsrahmen für deutsche hochschulabschlüsse (hqr). in handreichung. https://www.fuberlin.de/sites/bologna/dokumente_zur_bolognareform/hqr_handreichung_241019_final_ohne_hrk.pdf bourne, v. j. (2018). exploring statistics anxiety: contrasting mathematical, academic performance and trait psychological predictors. psychology teaching review, 24(1), 35–43. https://files.eric.ed.gov/fulltext/ej1180332.pdf brunyé, t. t., ditman, t., mahoney, c. r., augustyn, j. s., & taylor, h. a. (2009). when you and i share perspectives: pronouns modulate perspective taking during narrative comprehension. psychological wolgast & barnes-holmes 93 | f l r science, 20(1), 27–32. https://doi.org/10.1111/j.1467-9280.2008.02249.x champely, s. (2017). package “pwr” (pp. 1–22). https://github.com/heliosdrm/pwr cigala, a., mori, a., & fangareggi, f. (2015). learning others’ point of view: perspective taking and prosocial behaviour in preschoolers. early child development and care, 185(8), 1199–1215. https://doi.org/10.1080/03004430.2014.987272 cohen, j. (1988). statistical power analysis for the behavioral sciences (2.). erlbaum. damşa, c. (2014). the multi-layered nature of small-group learning: productive interactions in objectoriented collaboration. international journal of computer-supported collaborative learning, 9(3), 247–281. https://doi.org/10.1007/s11412-014-9193-8 damşa, c., nerland, m., & andreadakis, z. e. (2019). an ecological perspective on learner-constructed learning spaces. british journal of educational technology, 50(5), 2075–2089. https://doi.org/10.1111/bjet.12855 davis, m. h. (1980). a multidimensional approach to individual differences in empathy. jsas catalog of selected documents in psychology, 85–104. https://www.uv.es/friasnav/davis_1980.pdf davis, m. h. (1983). measuring individual differences in empathy: evidence for a multidimensional approach. journal of personality and social psychology, 44(1), 113–126. https://doi.org/10.1037/00223514.44.1.113 dunn, t. j., baguley, t., & brunsden, v. (2014). from alpha to omega: a practical solution to the pervasive problem of internal consistency estimation. 399–412. https://doi.org/10.1111/bjop.12046 engen, h. g., & singer, t. (2013). empathy circuits. in current opinion in neurobiology (vol. 23, issue 2, pp. 275–282). https://doi.org/10.1016/j.conb.2012.11.003 erle, t. m., & topolinski, s. (2015). spatial and empathic perspective-taking correlate on a dispositional level. social cognition, 33(3), 187–210. https://doi.org/10.1521/soco.2015.33.3.187 evertson, c. m. (1989). improving elementary classroom management: a school-based training program for beginning the year. the journal of educational research, 83(2), 82–90. https://doi.org/10.1080/00220671.1989.10885935 gehlbach, h. (2004). a new perspective on perspective taking. european science editing, 38(2), 35–37. https://doi.org/10.1023/b gehlbach, h., brinkworth, m. e., hsu, l. m., king, a. m., mcintyre, j., & rogers, t. (2016). creating birds of similar feathers: leveraging similarity to improve teacher-student relationships and academic achievement. journal of educational psychology, 108(3), 342–352. https://doi.org/10.1037/edu0000042 gehlbach, h., brinkworth, m. e., & wang, m. . t. (2012). the social perspective taking process: what motivates individuals to take another’s perspective? teachers college record. https://doi.org/10.1177/016146811211400108 gehlbach, h., marietta, g., king, a. m., karutz, c., bailenson, j. n., & dede, c. (2015). many ways to walk a mile in another’s moccasins: type of social perspective taking and its effect on negotiation outcomes. computers in human behavior, 52, 523–532. https://doi.org/10.1016/j.chb.2014.12.035 gehlbach, h., young, l. v., & roan, l. k. (2012). teaching social perspective taking: how educators might learn from the army. educational psychology, 32(3), 295–309. https://doi.org/10.1080/01443410.2011.652807 gore, n. j., barnes-holmes, y., & murphy, g. (2010). the relationship between intellectual functioning and relational perspective-taking. international journal of psychology and psychological therapy, 10(1), 1–17. https://mural.maynoothuniversity.ie/4982/ gudmundsdottir, s., & shulman, l. (1987). pedagogical content knowledge in social studies. scandinavian journal of educational research, 31(2), 59–70. https://doi.org/10.1080/0031383870310201 hayes, s. c., barnes-holmes, d., & roche, b. (2001). relational frame theory: a post-skinnerian account of language and cognition. springer science & business media. holmes, b.-. (2019). teaching a perspective-taking component skill to children with autism in the natural environment. 2(2), 439–450. https://doi.org/10.1002/jaba.523 hrk, & k. m. k. (2015). lehrerbildung für eine schule der vielfalt. gemeinsame empfehlung von hochschulrektorenkonferenz und kultusministerkonferenz. https://www.kmk.org/fileadmin/veroeffentlichungen_beschluesse/2015/2015_03_12-schule-dervielfalt.pdf hrk, bmbf, & kmk. (2011). deutscher qualifikationsrahmen für hochschulabschlüsse. in report. wolgast & barnes-holmes 94 | f l r hugener, i., pauli, c., reusser, k., lipowsky, f., rakoczy, k., & klieme, e. (2009). teaching patterns and learning quality in swiss and german mathematics lessons. learning and instruction, 19(1), 66–78. https://doi.org/10.1016/j.learninstruc.2008.02.001 janczyk, m. (2013). level 2 perspective taking entails two processes: evidence from prp experiments. journal of experimental psychology: learning memory and cognition, 39(6), 1878–1887. https://doi.org/10.1037/a0033336 johnson, d. w. (1975). cooperativeness and social perspective taking. journal of personality and social psychology, 31(2), 241–244. https://doi.org/10.1037/h0076285 klopsch, b., reschke, k., & sliwka, a. (2022). individual reference norm orientation and motivation: perspectives from germany, finland, and canada. in european perspectives on inclusive education in canada (pp. 203–214). routledge. e-isbn 9781003204572 kmk, & hrk. (2015). educating teachers to embrace diversity. https://www.kmk.org/fileadmin/veroeffentlichungen_beschluesse/2015/2015_03_12-kmk-hrkempfehlung-vielfalt-englisch.pdf korpershoek, h., harms, t., de boer, h., van kuijk, m., & doolaard, s. (2016). a meta-analysis of the effects of classroom management strategies and classroom management programs on students’ academic, behavioral, emotional, and motivational outcomes. review of educational research, 86(3), 643–680. https://doi.org/10.3102/0034654315626799 kounin, j. s. (1970). discipline and group management in classrooms. in discipline and group management in classrooms. holt, rinehart & winston. kunter, m., klusmann, u., baumert, j., richter, d., voss, t., & hachfeld, a. (2013). professional competence of teachers: effects on instructional quality and student development. journal of educational psychology, 105(3), 805–820. https://doi.org/10.1037/a0032583 lenhard, w., & lenhard, a. (2016). calculation of effect sizes. psychometrica. https://doi.org/10.13140/rg.2.2.17823.92329 levin, m. e., hildebrandt, m. j., lillis, j., & hayes, s. c. (2012). the impact of treatment components suggested by the psychological flexibility model: a meta-analysis of laboratory-based component studies. behavior therapy, 43(4), 741–756. https://doi.org/10.1016/j.beth.2012.05.003 lipowsky, f., rakoczy, k., pauli, c., drollinger-vetter, b., klieme, e., & reusser, k. (2009). quality of geometry instruction and its short-term impact on students ’ understanding of the pythagorean theorem. learning and instruction, 19(6), 527–537. https://doi.org/10.1016/j.learninstruc.2008.11.001 loftus, g. r. (1978). on interpretation of interactions. memory & cognition, 6(3), 312–319. https://doi.org/10.3758/bf03197461 lohbeck, a., & freund, p. a. (2021). students’ own and perceived teacher reference norms: how are they interrelated and linked to academic self-concept? educational psychology, 41(5), 640–657. https://doi.org/10.1080/01443410.2020.1746239 long, a. c. j., miller, f. g., & upright, j. j. (2019). classroom management for ethnic–racial minority students: a meta-analysis of single-case design studies. in school psychology (vol. 34, issue 1, pp. 1– 13). educational publishing foundation. https://doi.org/10.1037/spq0000305 marksteiner, t., nishen, a. k., & dickhäuser, o. (2021). students’ perception of teachers’ reference norm orientation and cheating in the classroom. frontiers in psychology, 12, 614199. https://doi.org/10.3389/fpsyg.2021.614199 mchugh, l., barnes-holmes, y., & barnes-holmes, d. (2004a). a relational frame account of the development of complex cognitive phenomena: perspective-taking, false belief understanding, and deception. international journal of psychology and psychological therapy, 4, 303–324. https://mural.maynoothuniversity.ie/401/ mchugh, l., barnes-holmes, y., & barnes-holmes, d. (2004b). perspective-taking as relational responding: a developmental profile. psychological record, 54(1), 115–144. https://doi.org/10.1007/bf03395465 mchugh, l., barnes-holmes, y., barnes-holmes, d., whelan, r., & stewart, i. (2007). knowing me, knowing you: deictic complexity in false-belief understanding. psychological record, 57(4), 533–542. https://doi.org/10.1007/bf03395593 möller, j., pohlmann, b., koller, o., & marsh, h. w. (2009). a meta-analytic path analysis of the internal/external frame of reference model of academic achievement and academic self-concept. review of educational research, 79(3), 1129–1167. https://doi.org/10.3102/0034654309337522 montoya-rodríguez, m. m., molina, f. j., & mchugh, l. (2017). a review of relational frame theory wolgast & barnes-holmes 95 | f l r research into deictic relational responding. psychological record, 67(4), 569–579. https://doi.org/10.1007/s40732-016-0216-x mori, a., & cigala, a. (2016). perspective taking: training procedures in developmentally typical preschoolers. different intervention methods and their effectiveness. educational psychology review, 28(2), 267–294. https://doi.org/10.1007/s10648-015-9306-6 mouw, j., saab, n. ., gijlers, h., hickendorff, m., van paridon, y., & van den broek, p. (2020). the differential effect of perspective-taking ability on profiles of cooperative behaviours and learning outcomes. frontline learning research, 8(6), 88–113. https://doi.org/10.14786/flr.v8i6.633 onwuegbuzie, a. j. (2004). academic procrastination and statistics anxiety. assessment & evaluation in higher education, 29(1), 3–19. https://doi.org/10.1080/0260293042000160384 onwuegbuzie, a. j., & seaman, m. a. (1995). the effect of time constraints and statistics test anxiety on test performance in a statistics course. the journal of experimental education, 63(2), 115–124. https://doi.org/10.1080/00220973.1995.9943816 peirce, j. w., gray, j. r., simpson, s., macaskill, m. r., höchenberger, r., sogo, h., kastman, e., & lindeløv, j. (2019). psychopy2: experiments in behavior made easy. behavior research methods. https://doi.org/10.3758/s13428-01801193-y pickering, m. j., mclean, j. f., & gambi, c. (2012). do addressees adopt the perspective of the speaker? acta psychologica, 141(2), 261–269. https://doi.org/10.1016/j.actpsy.2012.06.001 r development core team. (2009). r: a language and environment for statistical computing [computer software manual]. https://www.r-project.org/ rakoczy, k., buff, a., & lipowski, f. (2005). unterrichtsvideos.ch. http://www.unterrichtsvideos.ch/ ranger, j., & kuhn, j.-t. (2017). detecting unmotivated individuals with a new model-selection approach for rasch models. psychological test and assessment modeling, 59(3), 269–295. https://www.psychologie-aktuell.com/fileadmin/download/ptam/3-2017_20170920/01_ranger.pdf revelle, w. (2019). using r and the psych package to find ω. 1–20. www.personalityproject.org/r/tutorials/howto/omega.tutorial/omega.pdf rheinberg, f. (1977). bezugsnorm-orientierung von schülern der 5. bis 13. klasse bei der leistungsbeurteilung. zeitschrift für entwicklungspsychologie und padagogische psychologie, 9(2), 90–93. https://www.researchgate.net/profile/falko-rheinberg/publication/303248673_bezugsnormorientierung_von_schulern_der_5_bis_13_klasse_bei_der_leistungsbeurteilung/links/5739f19408ae2 98602e36913/bezugsnorm-orientierung-von-schuelern-der-5-bis-13-klasse-bei-derleistungsbeurteilung.pdf rheinberg, f. (2001). teachers reference-norm orientation and student motivation for learning. aera, conf., seattle, 10–14. https://www.researchgate.net/profile/falkorheinberg/publication/301356902_teachers_referencenorm_orientation_and_student_motivation_for_learning/links/57151b8d08aeafcb935d2fc6/teachers -reference-norm-orientation-and-student-motivation-for-learning.pdf rjosk, c., richter, d., hochweber, j., lüdtke, o., klieme, e., & stanat, p. (2014). socioeconomic and language minority classroom composition and individual reading achievement: the mediating role of instructional quality. learning and instruction, 32, 63–72. https://doi.org/10.1016/j.learninstruc.2014.01.007 rose, d. (2018). doing maths: constructing procedures for maths processes. in y. d. k. maton, j.r. martin (ed.), studying science: new insights into knowledge and language in education (pp. 257–286). routledge. https://doi.org/10.13140/rg.2.2.19268.73604 scanlon, g., & barnes-holmes, y. (2013). changing attitudes: supporting teachers in effectively including students with emotional and behavioural difficulties in mainstream education. emotional and behavioural difficulties, 18(4), 374–395. https://doi.org/10.1080/13632752.2013.769710 semmelmann, k., & weigelt, s. (2017). online psychophysics: reaction time effects in cognitive experiments. behavior research methods, 49(4), 1241–1260. https://doi.org/10.3758/s13428-0160783-4 siew, c. s. q., mccartney, m. j., & vitevitch, m. s. (2019). using network science to understand statistics anxiety among college students. scholarship of teaching and learning in psychology, 5(1), 75–89. https://doi.org/10.1037/stl0000133 strobl, c., kopf, j., & zeileis, a. (2015). rasch trees: a new method for detecting differential item functioning in the rasch model. psychometrika, 80(2), 289–316. https://doi.org/10.1007/s11336-013wolgast & barnes-holmes 96 | f l r 9388-3 strobl, c., wickelmaier, f., komboz, b., & kopf, j. (2016). package “psychotree.” https://cran.rproject.org/web/packages/psychotree/index.html tschannen-moran, m., & hoy, a. w. (2001). teacher efficacy: capturing an elusive construct. teaching and teacher education, 17(7), 783–805. https://doi.org/10.1016/s0742-051x(01)00036-1 von aufschnaiter, c., cappell, j., dübbelde, g., ennemoser, m., mayer, j., stiensmeier-pelster, j., sträßer, r., & wolgast, a. (2015). diagnostic competence theoretical considerations concerning a central construct of teacher education. zeitschrift fur padagogik, 61(5). wickelgren, w. a. (1977). speed-accuracy tradeoff. springerreference, 41, 67–85. https://doi.org/10.1007/springerreference_183986 willisch, a., wolgast, a., & donat, m. (2021). skalendokumentation „cyber-bullying unter studierenden. wilson, c. j., soranzo, a., & bertamini, m. (2017). attentional interference is modulated by salience not sentience. acta psychologica, 178, 56–65. wolgast, a. (2017). sämtliche schülerinnen und schüler sehen und berücksichtigen: perspektivübernahme von lehrpersonen. http://de.in-mind.org, 1, 1–14. http://de.in-mind.org/article/saemtlicheschuelerinnen-und-schueler-sehen-und-beruecksichtigen-perspektivuebernahme-von wolgast, a., & barnes-holmes, y. (2018). social perspective taking and metacognition of children: a longitudinal view across the fifth grade of school. humanistic psychologist, 46(1). https://doi.org/10.1037/hum0000077 wolgast, a., barnes-holmes, y., hartmann, u., & decristan, j. (2018). interrelations between perspective taking and reading experience: a longitudinal view on students in the fifth year of school. psychology of language and communication, 22(1). https://doi.org/10.2478/plc-2018-0019 wolgast, a., hille, m., streit, p., & grützemann, w. (2020). does test-anxiety experience impair student teachers’ later tendency to perspective-taking? acta educationis generalis, 10(1), 1–24. wolgast, a., & oyserman, d. (2019). seeing what other people see: accessible cultural mindset affects perspective-taking. culture and brain. https://doi.org/10.1007/s40167-019-00083-0 wolgast, a., tandler, n., harrison, l., & umlauft, s. (2020). adults’ dispositional and situational perspective-taking: a systematic review. educational psychology review, 32(2). https://doi.org/10.1007/s10648-019-09507-y zeidner, m. (1991). statistics and mathematics anxiety in social science students: some interesting parallels. british journal of educational psychology, 61(3), 319–328. https://doi.org/10.1111/j.20448279.1991.tb00989.x zimmerman, m. e. (2011). speed-accuracy tradeoff. in j. s. kreutzer, j. deluca, & b. caplan (eds.), encyclopedia of clinical neuropsychology (p. 2344). springer new york. https://doi.org/10.1007/9780-387-79948-3_1247 wolgast & barnes-holmes 97 | f l r appendix task study 1 all flexible spatial and temporal social perspective-taking tasks presented in study 1 according to task-type spatial i-you tasks simple i-you reversed i-you i am standing here writing statistics hints on the board and you are sitting there doing statistics tasks i am standing here writing statistics hints on the board and you are sitting there doing statistics tasks, if i was you and you were me where are you doing statistics tasks? (there) where am i writing statistics hints? (here) where would i be? (there) where would you be? (here) spatial-temporal i tasks simple here-there reversed here-there today i am writing statistics hints here, yesterday i was agreeing with ideas about statistics problems there today i am writing statistics hints here, yesterday i was agreeing with ideas about statistics problems there, if here was there and there was here where was i writing? (there) where was i agreeing? (here) where was i writing? (here) where was i agreeing? (there) today i am standing here in statistics class writing statistics hints, yesterday i was standing there in statistics class and that was fun, if here was there and there was here where am i standing now? (there) where was i having fun? (here) spatial-temporal you tasks today you are struggling to solve statistics tasks here, yesterday you were doing statistics tasks there where were you struggling to solve statistics tasks? (here) where were you doing statistics tasks? (there) temporal i tasks simple now-then reversed now-then yesterday my work in statistics class was challenging, today i am doing work in statistics that is fun yesterday my work in statistics class was challenging, today i am doing work in statistics that is fun, if yesterday was today and today was yesterday when was my work when was my work when was my work when was my work fun? wolgast & barnes-holmes 98 | f l r challenging? (yesterday) fun? (today) challenging? (today) (yesterday) yesterday i wrote statistics hints, today my work in statistics is easy yesterday i wrote statistics hints, today my work in statistics is easy, if yesterday was today and today was yesterday when was i writing statistics hints? (yesterday) when was my work in statistics easy? (today) when was i writing statistics hints? (today) when was my work in statistics easy? (yesterday) temporal you tasks simple now-then reversed now-then yesterday you were doing statistics tasks, today you are struggling to solve statistics tasks yesterday you were doing statistics tasks, today you are struggling to solve statistics tasks, if yesterday was today and today was yesterday when were you doing statistics tasks? (yesterday) when were you struggling to solve statistics tasks? (today) when were you doing statistics? (today) when were you struggling to solve statistics tasks? (yesterday) task study 2 all flexible reversed spatial and temporal social perspective-taking tasks presented in study 2 according to task-type spatial: reversed here-there tasks i tasks i am sitting here at the window signing in the class book and the students are sitting there writing maths hints, if here was there and there was here i am here with students and the other students are writing there with concentration, if here was there and there was here where was i sitting? (there) where was i? (there) i am going here through the aisle and the students are going there to the models, if here was there and there was here where was i going? (there) other tasks i am standing here writing maths tips on the board and the students are sitting there doing maths tasks, if here was there and there was here where were the students doing maths tasks? (here) i am here drawing a triangle on the board and the student is there and calculates an angle, if here was there and there was here where was the student? (here) i am standing here pointing to the pythagorean theorem in a book, a student is standing there and is bored, wolgast & barnes-holmes 99 | f l r if here was there and there was here where was the student? (here) i am sitting here and holding the book that includes the pythagorean theorem, the student is sitting there holding the book 10cm apart from her eyes, if here was there and there was here where was the student sitting? (here) i am sitting here beside the projector and pointing to a triangle, a student there at the table is painting a triangle in green, if here was there and there was here where was the student? (here) temporal: reversed now-then tasks i tasks yesterday the students were struggling to solve maths tasks and today i am demonstrating the solution, if yesterday was today and today was yesterday when was i demonstrating the solution? (yesterday) yesterday the students solved maths problems in groups and today i am listening carefully to students, if yesterday was today and today was yesterday when was i listening to students? (yesterday) you tasks yesterday you had fun doing maths tasks, today you are looking out the window, if yesterday was today and today was yesterday when had you fun? (today) yesterday you were drawing, today you are often raising your hand, if yesterday was today and today was yesterday when were you drawing? (today) yesterday you were drawing a triangle, today you are struggling to solve the task, if yesterday was today and today was yesterday when were you drawing the triangle? (today) yesterday you were doing maths tasks, today you are struggling to solve maths tasks, if yesterday was today and today was yesterday when were you doing maths tasks? (today) other tasks yesterday the students were doing maths tasks, wolgast & barnes-holmes 100 | f l r today the students are struggling to solve maths tasks, if yesterday was today and today was yesterday when were the students struggling to solve maths tasks? (yesterday) supplemental file c all flexible social perspective-taking tasks used in studies 1–3 study 1: flexible social perspective-taking tasks used in german in randomized order 1 ich bin auf einem hohen kompetenzniveau in statistik, du auf einem niedrigen. auf welchem kompetenzniveau bist du in statistik? niedrig hoch 2 ich bin auf einem hohen kompetenzniveau in statistik, du auf einem niedrigen. auf welchem kompetenzniveau bin ich in statistik? niedrig hoch 3 ich bin auf einem hohen kompetenzniveau in statistik, du auf einem niedrigen. wenn ich du wäre und du wärst ich, auf welchem kompetenzniveau wärst du in statistik? niedrig hoch 4 ich bin auf einem hohen kompetenzniveau in statistik, du auf einem niedrigen. wenn ich du wäre und du wärst ich, auf welchem kompetenzniveau wäre ich in statistik? niedrig hoch 5 ich stehe hier und schreibe statistik-tipps an die tafel und du sitzt dort und bearbeitest statistikaufgaben. wo bearbeitest du statistikaufgaben? hier dort 6 ich stehe hier und schreibe statistik-tipps an die tafel und du sitzt dort und bearbeitest statistikaufgaben. wo schreibe ich statistik-tipps? hier dort 7 ich stehe hier und schreibe statistik-tipps an die tafel und du sitzt dort und bearbeitest statistikaufgaben. wenn ich du wäre und du wärst ich, wo im seminarraum bin ich? hier dort 8 ich stehe hier und schreibe statistik-tipps an die tafel und du sitzt dort und bearbeitest statistikaufgaben. wenn ich du wäre und du wärst ich, wo im seminarraum bist du? hier dort 9 gestern war meine arbeit im statistikkurs eine herausforderung; heute macht meine arbeit in statistik spass. wann war meine arbeit herausfordernd? gestern heute 10 gestern war meine arbeit im statistikkurs eine herausforderung; heute macht meine arbeit in statistik spass. wann machte meine arbeit spass? gestern heute wolgast & barnes-holmes 101 | f l r 11 gestern war meine arbeit im statistikkurs eine herausforderung; heute macht meine arbeit in statistik spass. wenn gestern heute wäre und heute wäre gestern, wann war meine arbeit eine herausforderung? gestern heute 12 gestern war meine arbeit im statistikkurs eine herausforderung; heute macht meine arbeit in statistik spass. wenn gestern heute wäre und heute wäre gestern, wann machte meine arbeit spass? gestern heute 13 gestern hast du statistikaufgaben bearbeitet; heute zögerst du beim lösen von statistikaufgaben. wann hast du statistikaufgaben bearbeitet? gestern heute 14 gestern hast du statistikaufgaben bearbeitet; heute zögerst du beim lösen von statistikaufgaben. wann hast du gezögert, statistikaufgaben zu lösen? gestern heute 15 gestern hast du statistikaufgaben bearbeitet; heute zögerst du beim lösen von statistikaufgaben. wenn gestern heute wäre und heute wäre gestern, wann hast du statistikaufgaben bearbeitet? gestern heute 16 gestern hast du statistikaufgaben bearbeitet; heute zögerst du beim lösen von statistikaufgaben. wenn gestern heute wäre und heute wäre gestern, wann hast du gezögert, statistikaufgaben zu lösen? gestern heute 17 gestern habe ich statistik-tipps angeschrieben, heute fällt mir die arbeit in statistik leicht. wann habe ich statistik-tipps angeschrieben? gestern heute 18 gestern habe ich statistik-tipps angeschrieben, heute fällt mir die arbeit in statistik leicht. wann war meine arbeit in statistik leicht? gestern heute 19 gestern habe ich statistik-tipps angeschrieben, heute fällt mir die arbeit in statistik leicht. wenn gestern heute wäre und heute wäre gestern, wann habe ich statistik-tipps angeschrieben? gestern heute 20 gestern habe ich statistik-tipps angeschrieben, heute fällt mir die arbeit in statistik leicht. wenn gestern heute wäre und heute wäre gestern, wann war meine arbeit in statistik leicht? gestern heute 21 heute schreibe ich hier statistik-tipps, gestern habe ich dort den ideen von zu statistikproblemen zugestimmt. wo habe ich geschrieben? hier dort 22 heute schreibe ich hier statistik-tipps, gestern habe ich dort den ideen von zu statistikproblemen zugestimmt. wo habe ich zugestimmt? hier dort 23 heute schreibe ich hier statistik-tipps, gestern habe ich dort den ideen von zu statistikproblemen zugestimmt. wenn hier dort wäre und dort wäre hier, wo habe ich geschrieben? wolgast & barnes-holmes 102 | f l r hier dort 24 heute schreibe ich hier statistik-tipps, gestern habe ich dort den ideen von zu statistikproblemen zugestimmt. wenn hier dort wäre und dort wäre hier, wo habe ich zugestimmt? hier dort 25 heute zögerst du hier statistikaufgaben zu lösen, gestern hast du dort statistikaufgaben bearbeitet. wo hast du gezögert statistikaufgaben zu lösen? hier dort 26 heute zögerst du hier statistikaufgaben zu lösen, gestern hast du dort statistikaufgaben bearbeitet. wo hast du statistikaufgaben bearbeitet? hier dort 27 heute stehe ich hier im statistikkurs und schreibe statistik-tipps, gestern war ich dort im statistikkurs und es hat spass gemacht. wenn hier dort wäre und dort wäre hier, wo stehe ich jetzt? hier dort 28 heute stehe ich hier im statistikkurs und schreibe statistik-tipps, gestern war ich dort im statistikkurs und es hat spass gemacht. wenn hier dort wäre und dort wäre hier, wo hatte ich spass? hier dort 29 ich bin auf einem hohen kompetenzniveau in englisch, du auf einem niedrigen. auf welchem kompetenzniveau bist du in englisch? niedrig hoch 30 ich bin auf einem hohen kompetenzniveau in englisch, du auf einem niedrigen. auf welchem kompetenzniveau bin ich in englisch? niedrig hoch 31 ich bin auf einem hohen kompetenzniveau in englisch, du auf einem niedrigen. wenn ich du wäre und du wärst ich, auf welchem kompetenzniveau wärst du in englisch? gut schlecht 32 ich bin auf einem hohen kompetenzniveau in englisch, du auf einem niedrigen. wenn ich du wäre und du wärst ich, auf welchem kompetenzniveau wäre ich in englisch? gut schlecht 33 ich stehe hier und schreibe englisch-tipps an die tafel und du sitzt dort und bearbeitest englischaufgaben. wo bearbeitest du englischaufgaben? hier dort 34 ich stehe hier und schreibe englisch-tipps an die tafel und du sitzt dort und bearbeitest englischaufgaben. wo schreibe ich englisch-tipps? hier dort 35 ich stehe hier und schreibe englisch-tipps an die tafel und du sitzt dort und bearbeitest englischaufgaben. wenn ich du wäre und du wärst ich, wo im seminarrraum bin ich? hier dort wolgast & barnes-holmes 103 | f l r 36 ich stehe hier und schreibe englisch-tipps an die tafel und du sitzt dort und bearbeitest englischaufgaben. wenn ich du wäre und du wärst ich, wo im seminarrraum bist du? hier dort 37 gestern war meine arbeit im englischkurs eine herausforderung; heute macht meine arbeit in englisch spass. wann war meine arbeit herausfordernd? gestern heute 38 gestern war meine arbeit im englischkurs eine herausforderung; heute macht meine arbeit in englisch spass. wann machte meine arbeit spass? gestern heute 39 gestern war meine arbeit im englischkurs eine herausforderung; heute macht meine arbeit in englisch spass. wenn gestern heute wäre und heute wäre gestern, wann war meine arbeit eine herausforderung? gestern heute 40 gestern war meine arbeit im englischkurs eine herausforderung; heute macht meine arbeit in englisch spass. wenn gestern heute wäre und heute wäre gestern, wann machte meine arbeit spass? gestern heute 41 gestern hast du englischaufgaben bearbeitet; heute zögerst du beim lösen von englischaufgaben. wann hast du englischaufgaben bearbeitet? gestern heute 42 gestern hast du englischaufgaben bearbeitet; heute zögerst du beim lösen von englischaufgaben. wann hast du gezögert, englischaufgaben zu lösen? gestern heute 43 gestern hast du englischaufgaben bearbeitet; heute zögerst du beim lösen von englischaufgaben. wenn gestern heute wäre und heute wäre gestern, wann hast du englischaufgaben bearbeitet? gestern heute 44 gestern hast du englischaufgaben bearbeitet; heute zögerst du beim lösen von englischaufgaben. wenn gestern heute wäre und heute wäre gestern, wann hast du gezögert, englischaufgaben zu lösen? gestern heute 45 gestern habe ich englisch-tipps angeschrieben, heute fällt mir die arbeit in englisch leicht. wann habe ich englisch-tipps angeschrieben? gestern heute 46 gestern habe ich englisch-tipps angeschrieben, heute fällt mir die arbeit in englisch leicht. wann war meine arbeit in englisch leicht? gestern heute 47 gestern habe ich englisch-tipps angeschrieben, heute fällt mir die arbeit in englisch leicht. wenn gestern heute wäre und heute wäre gestern, wann habe ich englisch-tipps angeschrieben? gestern heute 48 gestern habe ich englisch-tipps angeschrieben, heute fällt mir die arbeit in englisch leicht. wenn gestern heute wäre und heute wäre gestern, wann war meine arbeit in englisch leicht? gestern wolgast & barnes-holmes 104 | f l r heute 49 heute schreibe ich hier englisch-tipps, gestern habe ich dort den ideen von zu englischproblemen zugestimmt. wo habe ich geschrieben? hier dort 50 heute schreibe ich hier englisch-tipps, gestern habe ich dort den ideen von zu englischproblemen zugestimmt. wo habe ich zugestimmt? hier dort 51 heute schreibe ich hier englisch-tipps, gestern habe ich dort den ideen von zu englischproblemen zugestimmt. wenn hier dort wäre und dort wäre hier, wo habe ich geschrieben? hier dort 52 heute schreibe ich hier englisch-tipps, gestern habe ich dort den ideen von zu englischproblemen zugestimmt. wenn hier dort wäre und dort wäre hier, wo habe ich zugestimmt? hier dort 53 heute zögerst du hier englischaufgaben zu lösen, gestern hast du dort englischaufgaben bearbeitet. wo hast du gezögert englischaufgaben zu lösen? hier dort 54 heute zögerst du hier englischaufgaben zu lösen, gestern hast du dort englischaufgaben bearbeitet. wo hast du englischaufgaben bearbeitet? hier dort 55 heute stehe ich hier im englischkurs und schreibe englisch-tipps, gestern war ich dort im englischkurs und es hat spass gemacht. wenn hier dort wäre und dort wäre hier, wo stehe ich jetzt? hier dort 56 heute stehe ich hier im englischkurs und schreibe englisch-tipps, gestern war ich dort im englischkurs und es hat spass gemacht. wenn hier dort wäre und dort wäre hier, wo hatte ich spass? hier dort study 2: flexible social perspective-taking tasks used in german and displayed in randomized order 1 ich stehe hier und schreibe mathetipps an die tafel und schüler/innen sitzen dort und bearbeiten matheaufgaben. wenn ich an der stelle der schüler/innen wäre und die schüler/innen an meiner stelle: wer bearbeitet die matheaufgaben? ich die schüler/innen 2 ich stehe am fenster und unterschreibe im klassenbuch; schüler/innen schreiben mathetipps auf. wenn ich an der stelle der schüler/innen wäre und die schüler/innen an meiner stelle: wer schreibt mathetipps auf? ich die schüler/innen 3 ich lese die sachaufgabe und schüler/innen haben die sachaufgabe gelöst. wenn ich an der stelle der schüler/innen wäre und die schüler/innen an meiner stelle: wer hat die sachaufgabe gelöst? wolgast & barnes-holmes 105 | f l r ich die schüler/innen 4 ich lese die sachaufgabe und der schüler meint, er kann die sachaufgabe nicht. wenn ich an der stelle des schülers wäre und der schüler an meiner stelle: wer liest die sachaufgabe? ich der schüler 5 ich sitze bei einer gruppe, die eine aufgabe gemeinsam löst; eine schülerin malt die aufgabe grün aus. wenn ich an der stelle der schülerin wäre und die schülerin an meiner stelle: wer sitzt bei einer gruppe? ich die schülerin 6 ich löse eine aufgabe und eine schülerin gähnt. wenn ich an der stelle der schülerin wäre und die schülerin an meiner stelle: wer löst die aufgabe? ich die schülerin 7 ich zeige auf ein dreieck und eine schülerin malt ein dreieck grün aus. wenn ich an der stelle der schülerin wäre und die schülerin an meiner stelle: wer zeigt auf das dreieck? ich die schülerin 8 gestern zögerten die schüler/innen matheaufgaben zu lösen und heute demonstriere ich die lösung. wenn gestern heute wäre und heute wäre gestern: wann demonstriere ich die lösung? gestern heute 9 gestern haben die schüler/innen matheaufgaben bearbeitet; heute zögern die schüler/innen matheaufgaben zu lösen. wenn gestern heute wäre und heute wäre gestern: wann haben die schüler/innen gezögert matheaufgaben zu lösen? gestern heute 10 gestern lösten schüler/innen in gruppen matheaufgaben und heute höre ich schüler/innen aufmerksam zu. wenn heute gestern wäre und gestern wäre heute: wann höre ich zu? gestern heute 11 gestern hast du matheaufgaben bearbeitet; heute zögerst du beim lösen von matheaufgaben. wenn heute gestern wäre und gestern wäre heute: wann hast du matheaufgaben bearbeitet? gestern heute 12 gestern hast du das dreieck gezeichnet, heute zögerst du die aufgabe zu lösen. wenn heute gestern wäre und gestern wäre heute: wann hast du das dreieck gezeichnet? gestern heute 13 gestern hast du gemalt; heute meldest du dich oft. wenn heute gestern wäre und gestern wäre heute: wann hast du gemalt? gestern heute 14 gestern machten dir die matheaufgaben spaß; heute schaust du aus dem fenster. wenn heute gestern wäre und gestern wäre heute: wann hattest du spaß? gestern heute 15 ich stehe hier und schreibe mathetipps an die tafel und schüler/innen sitzen dort und bearbeiten matheaufgaben. wenn hier dort wäre und dort wäre hier: wo bearbeiten die schüler/innen matheaufgaben? hier dort wolgast & barnes-holmes 106 | f l r 16 ich sitze hier am fenster und unterschreibe im klassenbuch; schüler/innen schreiben dort an tischen mathetipps auf. wenn hier dort wäre und dort wäre hier: wo sitze ich? hier dort 17 ich bin hier bei schüler/innen und die anderen schüler/innen schreiben dort konzentriert. wenn hier dort wäre und dort wäre hier: wo bin ich? hier dort 18 ich gehe hier durch den gang und die schüler/innen gehen zu den modellen dort. wenn hier dort wäre und dort wäre hier: wo gehe ich? hier dort 19 ich bin hier und zeichne ein dreieck an die tafel und der schüler ist dort und berechnet einen winkel. wenn hier dort wäre und dort wäre hier: wo ist der schüler? hier dort 20 ich stehe hier vorn und zeige im buch auf den satz des pythagoras und der schüler steht dort hinten am platz und langweilt sich. wenn hier dort wäre und dort wäre hier: wo steht der schüler? hier dort 21 ich sitze hier und halte das buch mit dem satz des pythagoras; die schülerin sitzt dort und hält das buch 10cm vor ihren augen. wenn hier dort wäre und dort wäre hier: wo sitzt die schülerin? hier dort 22 ich bin hier am beamer und zeige auf ein dreieck und eine schülerin dort am tisch malt ein dreieck grün aus. wenn hier dort wäre und dort wäre hier: wo ist die schülerin? hier dort study 3: flexible social perspective-taking tasks used in english (displayed in randomized order) 1 i am standing here writing math tips on the board and students are sitting there doing math tasks. where are students doing math tasks? here there 2 i am standing here writing math tips on the board and students are sitting there doing math tasks. where am i writing math tips? here there 3 i am standing here writing math tips and students are sitting there doing math tasks. if i were in the shoes of the students and the students were in my shoes: who would be doing math tasks? i the students 4 i am standing here writing math tips and students are sitting there doing math tasks. if i were in the shoes of the students and the students were in my shoes. who would be writing the math tips? i the students 5 yesterday students were doing math tasks; today, students are hesitating to solve math tasks. if yesterday were today and today were yesterday: when were students doing math tasks? wolgast & barnes-holmes 107 | f l r yesterday today 6 yesterday, students were doing math tasks; today, students are hesitating to solve math tasks. if yesterday were today and today were yesterday: when were students hesitating to solve math tasks? yesterday today study 3: flexible social perspective-taking tasks used in german (displayed in randomized order) 1 ich stehe hier und schreibe mathetipps an die tafel und schüler/innen sitzen dort und bearbeiten matheaufgaben. wo bearbeiten die schüler/innen matheaufgaben? hier dort 2 ich stehe hier und schreibe mathetipps an die tafel und schüler/innen sitzen dort und bearbeiten matheaufgaben. wo schreibe ich mathetipps? hier dort 3 ich stehe hier und schreibe mathetipps an die tafel und schüler/innen sitzen dort und bearbeiten matheaufgaben. wenn ich an der stelle der schüler/innen wäre und die schüler/innen an meiner stelle. wer würde die matheaufgaben bearbeiten? ich die schüler/innen 4 ich stehe hier und schreibe mathetipps an die tafel; schüler/innen sitzen dort und bearbeiten matheaufgaben. wenn ich an der stelle der schüler/innen wäre und die schüler/innen an meiner stelle. wer würde die mathetipps an die tafel schreiben? ich die schüler/innen 5 gestern haben die schüler/innen matheaufgaben bearbeitet; heute zögern die schüler/innen matheaufgaben zu lösen. wenn gestern heute wäre und heute wäre gestern: wann haben die schüler/innen matheaufgaben bearbeitet? gestern heute 6 gestern haben die schüler/innen matheaufgaben bearbeitet; heute zögern die schüler/innen matheaufgaben zu lösen. wenn gestern heute wäre und heute wäre gestern: wann haben die schüler/innen gezögert matheaufgaben zu lösen? gestern heute microsoft word silvola et al_publication.docx frontline learning research vol. 11 no. 2 (2023) 78 98 issn 2295-3159 corresponding author: anni silvola, learning and learning processes research unit, faculty of education and psychology, university of oulu, erkki koiso-kanttilan katu 1, 90570 oulu, finland, anni.silvola@oulu.fi doi: https://10.14786/flr.v11i2.1277 learning analytics for academic paths: student evaluations of two dashboards for study planning and monitoring anni silvolaa, amanda sjöblomb, piia näykkic, egle gedrimienea, hanni muukkonena a university of oulu, finland b aalto university, finland c university of jyväskylä, finland article received 21 april 2023 / revised 27 october 2023 / accepted 27 october 2023 / available online 6 december 2023 abstract an in-depth understanding of student experiences and evaluations of learning analytics dashboards (lads) is needed to develop supportive learning analytics tools. this study investigates how students (n = 140) evaluated two student-facing lads as a support for academic path-level self-regulated learning (srl) through the concrete processes of planning and monitoring studies. aim of the study was to gain new understanding about student perspectives for lad use on academic path-level context. the study specifically focused on the student evaluations of the dashboard support and challenges, and the differences of student evaluations based on their self-efficacy beliefs and resource management strategies. the findings revealed that students evaluated dashboard use helpful for their study planning and monitoring, while the challenge aspects mostly included further information needs and development ideas. students with higher selfefficacy evaluated the dashboards as more helpful for study planning than those with lower self-efficacy, and students with lower help-seeking skills evaluated the dashboards as more helpful for study monitoring than those with higher help-seeking skills. the results indicate that the design of lad can help students to focus on different aspects of study planning and monitoring and that students with different beliefs and capabilities might benefit from different lad designs and use practices. the study provides theoryinformed approach for investigating lad use in academic path-level context and extends current understanding of students as users of lads. keywords: academic paths; learning-analytics; multiple-case study; self-regulated learning; student-facing dashboards silvola et al. 79 | f l r 1. introduction 1.1 background this study investigates how students evaluate two student-facing learning analytics dashboards (lads) as a support for planning and monitoring their studies. lads that are well-aligned with students’ needs hold the potential to support students by providing them actionable insights into their learning (lim et al., 2020). however, what students report as relevant to see and use on dashboards often differs from what current lads typically provide (jivet et al., 2020; viberg et al., 2020). moreover, students’ perceptions and evaluations of the lad feedback influence on their willingness to use such feedback productively (schumacher & ifenthaler, 2018). therefore, there is a need to explore student evaluations and ideas of actionable feedback to develop lad designs and la use practices (ochoa & wise, 2021). this study focuses on academic-path level perspectives of la use, which means that learning is approached beyond the course level, focusing on the longer-term educational practices and processes such as study periods and academic years (e.g., ludvigsen et al., 2011). previous research on learning analytics (la) has broadly focused on course-level lads and their use as a support for srl (viberg et al., 2020). also, studies have focused on the institutional aspects of academic-path level support, such as predicting students’ risk for drop out or providing insights for academic advising (gutiérrez et al., 2018; ifenthaler & yau, 2020). the academic path perspective provides an important window to explore how he students’ capacity as learners and future professionals unfolds over time (e.g., lahn, 2011; vanthornout et al., 2017). academic paths involve many situations where students need to make different educational choices and decisions, including the selection of major and minor subjects, time management inside and outside of class, and eventually, what kind of career they wish to build (khiat, 2019; white, 2015). supporting students on their academic paths with la becomes ever more important for he institutions with increasing student intakes, pressure to graduate in given times and organizing education in distance and hybrid formats (pelletier et al., 2022). to be able to succeed in their learning with the increasingly complex, technology-rich environments of higher education, students need strong self-regulation skills (srl; jansen et al., 2020; järvelä et al., 2011). supporting srl has been one of the major interests in the research of lads (matcha et al., 2019). however, current research has shown gaps in the ways support for srl has been designed (viberg et al., 2020). studies have addressed the imbalance in supporting the phases of srl. while some of the lads focus the support on evaluation and reflection phases, the others emphasize more goal setting and planning phases of srl (heikkinen et al., 2023; jivet et al., 2017; viberg et al., 2020). however, all phases of srl are important for a successful srl process (panadero, 2017). this study aims to deepen the current understanding of students’ perspectives on support provided for the srl process via lads. we investigate how students evaluate the support and challenge aspects of two lads in terms of their srl process. students’ success in self-regulated learning is influenced by multiple factors. students’ motivational beliefs, such as self-efficacy beliefs, and their use of learning strategies have an interdependent relationship with the success of their srl (broadbent, 2017; zimmerman, 2000). students who have strong self-efficacy beliefs are likely to perform higher in academic achievement since they are more willing to approach challenging learning activities, put forth more effort and persist longer at difficult activities, and use srl strategies more effectively compared to those students with weak self-efficacy beliefs (schunk, 1991). resource management strategies describe how students can utilize the resources available for their learning, such as elements in their learning environment, or asking help from teachers or their peers (pintrich et al., 1991). self-efficacy beliefs and resource management strategies can influence how students evaluate, interpret, and are able to utilize feedback provided via lads (aguilar & baek, 2019; rets et al., 2021). in this study, we investigate how students’ self-efficacy beliefs and resource management strategies are associated with their evaluations of the lad support for their srl. this study applies a multiple-case design to explore students’ evaluations of the two lads. the study aims to deepen existing understanding of how students articulate the support and challenge aspects of the two lads in terms of their srl. this will increase contextualized understanding of the user aspects for silvola et al. 80 | f l r lad design in the context of academic paths. furthermore, this study analyzes the association between students’ self-efficacy beliefs and resource management strategies with student evaluations. this will deepen the existing understanding of the individual differences in students’ capacities to productively utilize the feedback provided via lads. 1.2. srl on the academic-path level srl is defined as a continuous, cyclical process through which a student engages in making a strategic effort for learning (winne, 2011). according to the srl model of winne and hadwin (1998), srl takes place through four linked phases 1) task understanding, 2) goal setting, 3) studying tactics, and 4) adaptation. the four phases are open and recursive, which means that they do not always unfold in the same order and that the work done in the earlier phases update the conditions in which students conduct their learning activities during the next cycle of regulation activity (panadero, 2017; winne & hadwin, 1998). this model of srl was selected in this study because on the academic-path level, students need to become aware, for example, of the rules and requirements of the different courses to be able to plan their studies. as compared with other srl models, specific for winne and hadwin’s (1998) model of srl is that the phases of task understanding and goal setting have been separated, which helps to formulate more specific ways of supporting srl (greene & azevedo, 2007). studies of la design and student experiences have discussed the ways lads support the phases of the srl process (heikkinen et al., 2023; viberg et al., 2020). we investigate student perspectives for lad use through the concrete processes of study planning and monitoring processes that contextualize and operationalize the four phases srl process on an academic-path level. this offers a new theoretical approach in a context where la has previously been typically developed in dataand practice-oriented ways (mendez et al., 2021). in finnish higher education, study planning and monitoring processes are guided by the creation of a personal study plan (psp) in the study information system. the purpose of the psp is to help students make their educational choices, select their majorand minor subjects, and help advisors and students monitor and reflect on study progression throughout the academic paths. thus, lads discussed in this study focus on the study planning and monitoring processes with the data related to the psp and study information system. 1.2.1 study planning the first two phases of srl, namely context understanding and goal setting, are in a major role in the study planning process. the construction of the objects is necessary for coordination to take place and create order for the students so they can experience continuity in their academic paths (ludvigsen et al., 2011). in the first phase, 1) understanding the study context, students become familiar with the affordances and constraints of making a personal study plan (psps) in their respective degree programs. students have an opportunity to familiarize themselves with the goals and missions of an institution as well as the rationale and structure of the studies (white, 2015). students formulate an understanding of their upcoming studies to set the overall objectives for conducting their degree. in the original model of winne and hadwin (1998), this is a task-understanding phase that explains how a student generates a perception about what the studying task is and what constraints and resources are in place. in the second phase, 2) goal setting and planning, students set goals for studying and plan how to reach them during study periods and academic years (winne, 2011). students have a chance to reflect on their individual interests, needs, and expectations regarding the available educational possibilities and constraints (white, 2015). in the original model of winne and hadwin (1998), this phase explains how a student generates goals and constructs a plan for addressing a study task. 1.2.2 study monitoring when students monitor their learning, they connect the historical aspects of learning with their present and formulate perceptions of their future academic paths (ludvigsen et al., 2011). in the third phase silvola et al. 81 | f l r of srl, 3) enacting the (study) plan, students use their psps to monitor their study progression in relation to their own and institutional goals, and to make necessary adaptations to their plans (winne, 2011). monitoring is crucial for students to continue reflecting the quality and relevance of the set plan (winne, 2011). on academic-path level students typically focus on coordinating several courses and their learning goals in parallel, managing their time and available individual resources for studying. in the original model of winne and hadwin (1998), this phase is called enacting study tactics and strategies, and it describes how the plan of study tactics that was created is carried out. the fourth phase, 4) adaptation of the (study) plan describes how at chosen points in time, the entire plan may be evaluated. this is an important step to make adaptations such as setting more specific or different types of goals for following semesters, selecting different types of courses, or stretching the time used for their degree. in the original model of winne and hadwin (1998), the last phase is described as a large-scale adaptation, and it focuses on how students evaluate the entire approach that they selected to work on their task based on their overall experience of studying in the first three stages (winne & hadwin, 1998). 1.3. self-efficacy beliefs and resource management strategies self-efficacy is students’ motivational perception of their capabilities to succeed in upcoming learning processes and tasks and to achieve a desired outcome (bandura, 1977; schunk, 1991). self-efficacy beliefs have a reciprocal influence on students’ srl, such as their use of different learning strategies (järvelä et al., 2011). students who have strong self-efficacy beliefs are likely to perform higher in academic achievement since they are more willing to approach challenging learning activities, put forth more effort and persist longer at difficult activities, and use srl strategies more effectively compared to those students with weak self-efficacy beliefs (schunk, 1991). students with higher self-efficacy beliefs are more capable in utilizing feedback to improve their learning strategies and outcomes (li et al., 2022). thus, it is expected in this study that students with higher self-efficacy beliefs provide more positive evaluations of lads to support their srl. knowledge of srl strategies is an influential part of students’ srl competence (dresel et al., 2015). resource management strategies include skills such as time management, and managing study environment, effort regulation such as prioritizing and planning one’s use of time, and help-seeking skills such as how well students engage in support-seeking strategies such as asking for help from others (dresel et al., 2015). these strategies are highlighted beyond the course level when students control and coordinate their parallel learning processes throughout different contexts and assignments. learners who are less fluent in self-regulating their learning, use more intensively actionable feedback about their learning (tempelaar, 2019). studies have also indicated that students with fluent srl skills consider dashboard elements relevant to their learning (jivet et al., 2020). thus, it is expected that students with high resource management strategies provide more positive evaluations of lads than students with low resource management strategies. 2. research aims we ground the analysis of student evaluations with the four phases of srl to explore how students perceive the dashboards supporting their study planning and monitoring on the academic path level. we focus on the student perspectives on lads in authentic contexts through two cases where students were instructed to test the dashboards and to reflect on their experience. this study addresses the following research questions: 1) how did the students evaluate the lads as a support for srl phases? 2) what kind of challenges did the students identify in lads in terms of support for srl phases? 3) how were the students’ self-efficacy beliefs and resource management strategies associated with students’ experienced support or challenges? we set the following hypotheses to study the association of self-efficacy beliefs and resource management strategies with student evaluations of the support and challenge aspects of the dashboards: silvola et al. 82 | f l r hypothesis 1: higher self-efficacy beliefs are associated with more positive evaluations of lads to support the srl on the academic path level. hypothesis 2: higher resource management strategies are associated with more positive evaluations of lads to support the srl on the academic path level. 3. materials and methods 3.1. study context we applied a multiple-case design (yin, 2009) to explore student evaluations of lads. this enabled contextualized exploration of students’ lad evaluations. in finnish higher education, each student creates a personal study plan (psp) that contains all courses and major and minor electives that they plan to include in their academic path. the psp is used as a reflection tool for students and advisors to monitor and reflect the study progression through their psps. the design of the two lads were developed within a multidisciplinary project team including members from seven finnish universities. the project focused on identifying essential aspects of legal, ethical, institutional, and pedagogical design and deployment of la in he institutions. the development process of student-facing lads started by investigating students’ information and support needs (silvola et al., 2021a). after that, the use of registry data was explored, and analysis methods for visualizing data were analyzed and evaluated within the development team. during the process, a multidisciplinary development team defined the criteria for the lad design. at first, the usability and fluency of the use of lads were investigated with students (silvola et al., 2021b; gedrimiene et al., 2022). this study focuses on the use of lads as a support for the study planning and monitoring process, thus providing new evidence about the relevance and usefulness of the designed lads. the lads were developed in two finnish universities in interdisciplinary developer teams. the lads use students’ registry-based data (psps, credits achieved, timeline of achieved credits) and degreestructure information. the goal of both lads was to facilitate the process in which a student creates a psp, sets goals for studies, and monitors the enacting of one’s plans. in this study, one dashboard pilot constituted a case. in the two cases, the participants were higher education students who had a task to plan and monitor their studies with the help of the lad feedback. the cases were different in terms of the dashboard design (analyticsai and powerbi) and the participant group (table 1). the two cases provided a window to understand similarities and differences in students’ evaluations of the two different lad designs. this information is important to further improve academic path-level support for students via lads. table 1 overview of the two cases in the study testing analyticsai powerbi participants (n) 104 36 data used artificial study registry data study registry data la positioning visualizations included directly at the interactive platform visualizations included in the dashboard with no interactive elements 3.1.1 analyticsai analyticsai focused on interactively supporting the psp creation process by providing real-time feedback (figure 1). to help students to create a quality plan, feedback was provided about the 1) number of courses chosen, 2) division of the workload, and 3) parts of the degree included in the plan. la was positioned to provide feedback for students while they were modifying the suggested template of their silvola et al. 83 | f l r studies or when students wanted to see how their studies were progressing. the visualizations filtered information about 1) the institutional requirements for the academic path, 2) the individual goals the students had set for their psp, 3) the student’s prior activity, and 4) the timeline for the studies. the dashboard visualized how many courses students had completed from their psps (figures 2,3), and how their progression speed responded to the graduation time goal they had set. the dashboard included recommendations about the courses that others with similar plans had chosen. the content was divided into three main sections, where the “dashboard” and “studies” sections included information about the degree structure and psp respectively (figure 1). the “timetable” section provided information about students’ course schedules and exams. the “dashboard” section contained an overview of students’ progression according to the created psp and selected degree structure (figure 1). figure 1. study planning view: a planning template that students modify by adding and removing courses. students can overview the structure of the studies, see how their studies divide within study periods, and see their study records to update the psp. figure 2. visualization of student’s planned and completed courses for bachelor's degree. silvola et al. 84 | f l r figure 3. visualization of students’ progression during an academic year, comparison with the planned ects with the completed ects. 3.1.2 powerbi in powerbi the content was divided to visualize the psp and to provide an overview of the current progression in the studies. students could modify their psp based on the provided feedback in the platform separate from powerbi. the visualizations provided information about 1) individual goals that students set for their psps (figure 4), 2) information about a students’ own prior activity, including timelines of study progression (figure 5), 3) the distribution of the grades according to the study field and degree part (figure 3) 4) peer comparison information to support monitoring (figure 6), 4) and an estimation of the graduation time based on the number of completed and required credits for a given degree program. with the help of the dashboard, students could take an overview of their ongoing and upcoming courses selected into their psp during the one academic year (figure 6). figure 4. planning view of powerbi dashboard silvola et al. 85 | f l r figure 5. visualizations of student grades (gedrimiene et al., 2022). on the first row: a) timeline of students’ numerical grades on courses, b) average of numerical grades per year. on the second row: c) distribution of the grades according to study field, d) distribution of grades according to study type. figure 6. overview of student progression based on the psp for the selected year. left: bar plot indicating planned and completed courses including peer completion comparison. right: radar plot with peer grade comparison (gedrimiene et al., 2022). 3.2 participants in case 1, the use of artificial study registry data in the testing enabled participation in the study with different levels of experience from he studies. participants (n = 104) were from five finnish universities (37 male, 67 female) from various fields, such as natural sciences, engineering, ict fields, health sciences, educational sciences, economics, medicine, humanities, and social sciences. the sample consisted of 76 finnish students and 28 international students at various stages of their studies, the age distribution is presented in table 2. in case 2, the use of study registry information in the testing limited the group of participants to those who already had studied in university at least one year and already had completed and planned courses in their study registry. participants (n = 36) were from a finnish university (20 female, 16 male). the sample silvola et al. 86 | f l r consisted of 20 bachelor’s phase and 16 master’s phase students from educational sciences (23 students) and electrical engineering (13 students). the age distribution is presented in table 2. students from educational sciences and electrical engineering were invited to participate in the study so that feedback from students with different educational backgrounds would be obtained. table 2 age of participants age case 1 (n=104) case 2 (n=36) 18-22 53.8% 40.5 % 23-27 29.8 % 32.4 % 28 or older 16.3 % 27% 3.3. data collection research data was collected through an online questionnaire, and the same protocol was used to collect data in both cases. data collection protocol included pre-questionnaire, instructed lad testing, and post-questionnaire. students were introduced to lad use with the instructions included in the protocol. likert-scale questionnaire items measured students’ self-efficacy beliefs and resource management strategies (based on the mslq, pintrich et al., 1991), and open-ended questions measured students’ evaluations of the lads. the pre-questionnaire consisted of items on student background information (age, gender, grade, field) and two modules from the motivated strategies for learning questionnaire (mslq). the first module was self-efficacy beliefs including the scales control of learning beliefs (example question: if i try hard enough, then i will understand the contents in my studies) and self-efficacy beliefs (i believe i will receive excellent grades in my studies). the second module was resource management strategies including the scales effort regulation (even when course materials are dull and uninteresting, i manage to keep working until i finish), help-seeking (when i can't understand the course material, i ask another student for help) and study environment (i have a regular place set aside for studying) and time management strategies (i often find that i don't spend very much time on my studies because of other activities) (table 3, pintrich et al., 1991). the reliability of the mslq items was tested with cronbach’s alpha values for each scale (table 3). the internal consistency values for the questionnaire scales were similar to the original values of pintrich et al. (1991) except for the effort regulation scale, which remained lower (0.47) than the original (0.69). one item was removed from the help-seeking scale to improve internal consistency. in the pre-questionnaire, students evaluated their self-efficacy beliefs and resource management strategies using seven-point likert scale items from 1 (strongly disagree) to 7 (strongly agree). students were asked to think about their university studies at a general level when answering these questions. table 3 descriptive statistics for the selected questionnaire scales (pintrich et al., 1991). scale (likert scale 1–7) mean sd cronbach’s alpha self-efficacy control of learning beliefs 5.29 .918 0.697 self-efficacy 5.24 .915 0.694 resource management strategies effort regulation 4.74 .738 0.474 help-seeking 4.87 .805 0.604 study environment 4.09 .745 0.615 silvola et al. 87 | f l r in the post-questionnaire, student evaluations were collected through three open-ended questions focused on support and challenge aspects that students identified when using the dashboards for study planning and monitoring: 1) what things (in the service) were useful in planning and monitoring your studies? 2) what things were challenging? 3) how would you develop the service? 3.4 data analysis data analysis was conducted in three parts. first, qualitative content analysis was selected as an analytical approach to enable systematic and quantifiable analysis of open-ended answers (chi, 1997). in the second phase, the results of the analysis were quantified. finally, t-tests were conducted to analyze the association with students’ evaluations, self-efficacy beliefs, and resource management strategies (rq3). 3.4.1 qualitative content analysis four phases of srl (winne & hadwin, 1998) were used as main categories to structure the analysis of student evaluations of the la dashboards (table 4). the fourth phase of the model, the adaptation of the plan, was left out of the analysis since the research design did not enable students to evaluate how the dashboards could support adaptation phase. the coding scheme focused on identifying how lads supported study planning and monitoring processes and what kind of challenges students identified in using lads. silvola et al. 88 | f l r table 4 coding scheme for main categories to analyze student evaluations of lads’ support and challenge aspects (white, 2015; winne, 2011; winne & hadwin, 1998) main category category description study planning 1) understanding of the study context the student becomes familiar with the affordances and constraints of choosing their courses and making a psp. they come to understand the institutional rules and expectations for creating a study plan. 2) goal setting the student sets goals for studying and plans how to reach them during study periods and academic years. this enables students to compare their individual goals and interests with available courses. study monitoring 3) enacting the plan the student starts to complete courses according to the set study plan. as studies progress, the student starts to monitor their progression based on two standards: institutional and individual goals for study progression and achievement. open-ended answers from cases 1 and 2 were first analyzed separately. at first, the responses were categorized into support or challenge aspects and then divided into the three phases of srl based on the theory-oriented category descriptions in table 4. support and challenge aspects were analyzed separately. subcategories were formulated based on the themes that arose from the data. this allowed analysis of what kind of aspects students emphasized as supportive or challenging in lads. the unit of analysis was one mentioned theme in the open-ended answer and the length of the unit of analysis varied from one word to one sentence. one open-ended answer could include several units of analysis. no overlapping was allowed between coding categories. the reliability of the coding was ensured by two coders coding 15% of the openended answers. cohen’s kappa was calculated at k = 0.96, which indicates high reliability (salkind, 2010). the data from both cases were combined into one dataset, and all qualitative mainand subcategories were quantified by transforming each response into a discrete variable (1 = response in the category, 0 = no response in the category) for the quantitative analyses. similar mainand subcategories were identified from both cases for further analysis. the categories identified from both cases were time management planning from the goal setting phase (n(0) = 108, n(1) = 32), the enacting-the-plan phase (n(0) = 108, n(1) = 32), and goal setting challenges (n(0) = 82, n(1) = 58). thus, the common categories were selected for further analysis. 3.4.2 quantitative data analysis to analyze the association of self-efficacy beliefs and resource management strategies with the student evaluations of the dashboard support and challenges, welch t-tests were conducted. welch t-test was used, as it is better than a student’s t-test at comparing the means of two independent groups when sample sizes and variances are not equal between the groups. a t-test was conducted between each measured subscale (control of learning beliefs, self-efficacy, effort regulation, help-seeking, and study environment) and such mainand subcategories that were common for both cases (time management planning from goal setting phase, enacting-the-plan phase, and goal setting challenges). silvola et al. 89 | f l r 4 results 4.1. student evaluations of the dashboards’ support for study planning and monitoring 4.1.1. analyticsai student evaluations of analyticsai emphasized the support for the context-understanding and goal setting phases of srl (table 5). table 5 support aspects for study planning and monitoring that student evaluated in analyticsai. main category subcategory ƒ % study planning understanding study context suggested plan template 13 12.5 course information 9 8.7 quality of the plan 7 6.7 goal setting course recommendations 12 11.5 modifiability of the plan 23 22.1 time management planning 21 20.2 study monitoring enacting the plan progression monitoring support 18 17.3 at the context-understanding phase, students found that the dashboard helped them to understand how to start planning their degree by providing a suggested plan template (ƒ = 13): “the suggested degree structure that enables adding and removing courses” and “mandatory contents of the degree available in the suggested plan.” they explained it as a fluent way of planning to start modifying the plan based on the degree requirements and their own interests. also, a tool provided information about the courses (ƒ = 9): “periodic course search seems like an easy way to find more courses”. and some students liked that they got important feedback about whether their plan met the institutional requirements for the degree (ƒ = 7): “seeing how to divide your classes in order to graduate earlier”. many student evaluations of analyticsai were focused on the goal setting phase. students found that the tool supported them by providing insights for workload planning and time management (ƒ = 21): “the view for the academic year was useful since it enabled to plan the timing of the courses and the division of the courses between academic years” and “it was easy to follow the workload of the studies in the planning phase.” students also felt that the course recommendations (ƒ = 12) were supporting their goal setting phase with suggestions of courses that might suit their plans: “it's good that the system already has courses built in. it was easy to choose the courses based on that.”. the modifiability of the plan (ƒ = 23) helped students to set relevant goals: “you got options for creating several study plans and modify it”. students could first try out different plans and see from the visualizations what these plans would look like. at the enacting-the-plan phase, students reported that the views on the tool could support their progression monitoring (ƒ = 18). they described the views that showed their completed courses and the overview about the study progress as important: “possibility to follow my studies in a timely manner” and “it is easy to see past and upcoming studies as a whole and by study periods.” students specifically mentioned that the possibility to see completed and upcoming planned courses in the same view on the front page was useful for them. students identified challenges, support needs, and development ideas related to the contextunderstanding and goal setting phases of srl (table 6). students reported a need for more information silvola et al. 90 | f l r about the study structures (ƒ = 29) to be able to better plan their studies: “maybe making things more fluent for beginners and recommending studies by default could be a good idea because someone who knows what they are doing can easily remove things, while someone who has no idea what they are doing may need some help.” they suggested that the service could provide more information that is typically provided through study guidebooks, such as:” more instructions about the structure of the degree or at least a link for a study guide from which you could easily see the structure.” these were for example pre-requirements of the courses, the rules of how and what major and minor studies should be chosen, information about language studies, and more step-by-step guidance on how to formulate a psp. table 6 challenges students identified when using analyticsai. main category subcategory ƒ % study planning context-understanding challenges information about the study structures 29 27.9 goal setting challenges anticipation of the resource use 23 22.1 recognition of the individual goals 6 5.8 plan modification challenges 34 32.7 study monitoring monitoring challenges 0 0 students suggested that the tool could provide even more support for the anticipation of their resource use as a support for goal setting (ƒ = 27): “to think how many courses i can have a year” and “to see if courses overlap.” some students reported that it was difficult to recognize individual interests (ƒ = 6) and that there were also technical difficulties with modifying the plan (ƒ = 34), such as adding or choosing courses: “(it is difficult) to find the courses for the right period. the calendar could have periods separately, or there could be a whole year view added.” 4.1.2. powerbi powerbi specifically supported the goal setting and enacting-the-plan phases of srl (table 7). students felt that the tool was important in formulating and overviewing their studies. table 7 support aspects students identified in powerbi. main category subcategory ƒ % study planning understanding the study context 0 0 goal setting an overview of studies 12 33.3 time management planning 12 33.3 study monitoring enacting the plan progression monitoring support 27 75 information about progression 13 36.1 adapting the plan 3 8.3 for goal setting phase, students described that powerbi provided them an important overview of their studies (ƒ = 12), such as the number of credits in each semester, seeing course credits separately and in total, and views that helped to get an overview of a planned academic path: “it shows my credits and subjects from my studies, it helps me plan and look at the projection” and “i can see the courses credits separately and in total, which gives a good idea for my plan in addition to having filters and visualizations.” students found that the tool helped them with time management planning, such as workload division in each study period (ƒ = 12). at the enacting-the-plan phase, students described that powerbi supported their progression monitoring (ƒ = 27) by providing important overviews of completed and planned courses. students silvola et al. 91 | f l r mentioned that the views provided them with information about the longer-term course performance and allowed them to reflect on upcoming semesters as well. students liked the possibility of seeing an estimation of their graduation time based on their current progression speed: “the views can help to notify that with this pace graduation will not happen according to the originally planned schedule.” some students explained that these monitoring views also helped them to adapt their upcoming plans based on their current evaluations of how they were doing (ƒ = 3). students found that the dashboard could help them to notice if they were not going to graduate according to their plan and helped in deciding what kinds of actions to take next. the tool provided students with information on their progression (ƒ = 13), especially through the view in which they could compare their own grades with those of their peers, as well as insights into how they were doing with their courses: “it is interesting to see my grades compared to others. it is also interesting to see in picture form how i am progressing in my study plan” and “the information about the performance of other students helps me to better understand my own study (success).” students identified challenges, support needs, and development ideas related to the goal setting and enacting-the-plan phases of srl (table 8). as the sample size was smaller than that in case 1, no subcategories were identified from the challenge aspects in the powerbi case. table 8 challenges students identified in powerbi. main category ƒ % study planning context-understanding challenges 0 0 goal setting challenges 12 33.3 study monitoring monitoring challenges 21 58.3 for the goal setting phase (ƒ = 12), students suggested more clarity in visualizations that helps planning or that it did not provide enough ideas about further opportunities and alternatives. some students described that they did not have an idea about how to modify their study plans: “it does not give me ideas about further opportunities and alternatives” or that the dashboard was not very relevant to their degree: “however, most of the courses in our program are mandatory, so i don't have much space to plan the courses by myself, so i am not sure if i can make good use of this tool as much.” as a challenge for monitoring (ƒ = 21), students found that the tool would provide them with information in more detail, or that the information was not relevant from the aspect of their study field: “it isn't super useful, i think, primarily i personally have a good sense of the progression and courses just by looking at my psp.” some students explained that in their field, many courses are assessed as either passed or failed when some of the visualizations (radar plot) did not show relevant information for them. additionally, students recommended that the dashboard could suggest certain electives to choose from for upcoming studies or provide more predictions of their degree progression. the results of the qualitative content analysis show that both dashboards, analyticsai and powerbi, supported students to analytically reflect on their study planning and monitoring processes. the support and challenge aspects of analyticsai addressed that the dashboard focused especially on the study planning process, while the powerbi visualizations were discussed more from the study monitoring aspects. in the following sections, the results of the qualitative content analysis are described. 4.2 the association of self-efficacy beliefs and resource management strategies with student evaluations according to the results of the welch t-tests, self-efficacy had a significant association with the student evaluations of the dashboard in the goal setting phase (table 9). the results of the t-tests indicate that students who mentioned time management planning as a supportive aspect of the dashboards had significantly higher self-efficacy than those who did not mention it; t(49) = -2.92, p = .005. students who silvola et al. 92 | f l r mentioned challenges in the goal setting phase had significantly lower self-efficacy, t(123) = 2.64, p = .009, compared to the students who did not mention them. from the subscales of resource management strategies, students who identified support aspects for the enacting-the-plan phase indicated significantly lower help-seeking compared to those who did not mention it; t(76) = 1.06, p = .043, g = .641. no other significant differences were identified from the other subscales of resource management strategies (table 9). table 9 t-test results for the categories of time management planning and enacting the plan as supportive aspects of the dashboard. goal setting challenges as challenging aspects. time management planning (goal setting phase of srl) enacting the plan (enacting the plan phase of srl) goal setting challenges (goal setting phase of srl) 0(n=108) 1(n=32) 0(n=98) 1(n=42) 0(n=82) 1(n=58) self-efficacy m = 4.81, sd = 0.89 m = 5.34, sd = 0.92 m = 4.87, sd = 0.93 m = 5.07, sd = 0.89 m = 5.10, sd = 0.90 m = 4.69, sd = 0.90 t(49) = -2.92, p = .005*, g = 0.591a t(81) = -1.22, p = .225, g = 0.218 t(123) = 2.64, p = .009, g = 0.456* effort regulation m = 0.39, sd = 0.92 m = 0.43, sd = 0.9 m = 0.39, sd = 0.91 m = 0.40, sd = 0.95 m = 0.51, sd = 0.86 m = 0.25, sd = 0.93 t(55) = -0.22, p = .825, g = 0.043 t(81) = -0.04, p = .967, g = 0.011 t(117) = 1.66, p = .100, g = 0.292 help seeking m = 2.20, sd = 0.94 m = 2.44, sd = 0.97 m = 2.60, sd = 0.93 m = 2.00, sd = 0.95 m = 2.17, sd = 0.99 m = 2.37, sd = 0.89 t(50) = -1.24, p = .219, g = 0.253 t(76) = 1.06, p = .043, g = 0.641 t(130) = -1.23, p = .223, g = 0.211 study environment m = 0.39, sd = 1.10 m = 0.73, sd = 1.05 m = 0.44, sd = 1.05 m = 0.52, sd = 1.20 m = 0.57, sd = 1.11 m = 0.32, sd = 1.06 t(53) = -1.59, p = .119, g = 0.312 t(69) = -0.35, p = .729, g = 0.073 t(126) = 1.32, p = .190, g = 0.229 note: 0 = students who did not mention the category in their open-ended answers, 1 = students who mentioned the category in their open-ended answers. a = hedges’ g. silvola et al. 93 | f l r 5. discussion in this study, student evaluations of two different lads as support for study planning and monitoring were examined through two cases. the findings of this study showed that both lads provided useful feedback for study planning and monitoring processes, but the lads differed in terms of how support was divided within the three phases of srl. the challenges emerged especially due to the granularity level of the feedback provided. furthermore, students’ evaluations of the lads differed based on their selfefficacy beliefs and help-seeking skills. the study planning process was supported especially by analyticsai. analyticsai provided information for understanding the context of planning the studies. however, the lack of sufficient information for understanding the planning context was raised as a challenge. both analyticsai and powerbi provided support for goal setting phase. specifically, support for planning time management was considered useful in both lads. also, many challenges were reported in the goal setting phase, in particular, related to challenges in editing the plan, and time management planning in the goal setting process. some students reported that it was difficult to set goals for their psp, since it was difficult to recognize their individual goals and interests. this finding could be used to develop both more student-centered information sources of students’ goals for lads, and to develop academic advising practices that utilize such lads, but also guide students to reflect on their personal goals. in powerbi, planning was supported by an overview of studies and views supporting time planning. however, students suggested development of the powerbi in terms of providing more insights into how to adjust their plans based on the provided information. in previous studies of lads, support for planning and organizing studies has been considered important by the students (schumacher & ifenthaler, 2018). furthermore, supporting goal setting phase has been well included in the dashboard designs aiming to support srl processes (viberg et al., 2020). study monitoring process support was highlighted in powerbi. here, students described that the overview of studies -view in particular helped students to visualize how their studies have progressed in relation to their psp, what the current study progression looks like in relation to the estimated time of completion. some students described that the overview made it easy to plan what changes were needed to their own plans. however, this was contradictory with some students criticizing overviews in the planning process for not giving actionable feedback for the planning process. analyticsai also supported monitoring by providing views of the student's progress. for both lads, the challenge was that more detailed information on progress was needed to support monitoring. this is an important finding for developing la design and use in academic path-level studies. in addition, some students reported that the data did not seem as actionable for the students from different study fields. it has been noted that the structure and goals of the degrees vary between fields. in the testing of lads, students identified aspects that could help them to refine thinking about their academic path, with added knowledge for decisions about the next steps and adjusting plans. these findings are in line with the previous research suggesting that lads can help students identify upcoming next steps in their learning and elaborate their thinking about their learning (aguilar et al., 2021; jivet et al., 2020). the students in the two cases evaluated the lads in different phases of their academic paths. in case 1, students were from different grades, but it is expected that many of them had the planning task ahead when the data was collected during the opening event of an academic year. in case 2, students were 3rd and 4th year students, who already have created their psps and may benefit more from the monitoring views. thus, more research would be needed about the support needed for planning and monitoring processes in different phases of academic paths. one important advance of la is to provide personalized feedback to different students (van leeuwen et al., 2022). in this study, the association with students’ evaluations, self-efficacy beliefs, and resource management strategies were examined. students who found that the dashboards helped them with time management planning had higher self-efficacy beliefs than those who did not report such support. furthermore, students who reported challenges with lad in goal setting had lower self-efficacy beliefs than those who did not report challenges. these results support hypothesis 1. in the tested lads, time silvola et al. 94 | f l r management planning was guided with several visualizations, and it emerged as an important support aspect for goal setting phase of srl. based on previous research, setting can be challenging for students with low self-efficacy (zimmerman, 2000). there is a need to consider students' perceptions of self-efficacy in la design so that la tools do not only support students who are already successful in their studies. the findings are in line with rets et al. (2021) study in which students’ self-efficacy beliefs impacted how useful they perceived the feedback provided via dashboards. students who evaluated the dashboards as helpful in study monitoring had lower help-seeking skills. no associations were found with other resource management strategies. therefore, hypothesis 2 is only partially supported. the findings align with the previous research showing how students’ help-seeking skills associate with the perceived usefulness of lads for monitoring learning (aguilar & baek, 2019; jivet et al., 2018). this finding is essential for the development of la to support monitoring. especially in online and hybrid contexts, it can be more challenging for students to activate different learning strategies (broadbent, 2017). on the level of academic paths, it can also sometimes be challenging to identify from students’ perspective, where to ask for help with learning or study-related issues. thus, la might serve an important role on the level of academic paths by providing a channel for students to receive relevant and timely feedback from institutional rules and requirements. the results align with the previous studies addressing that one design does not provide the necessary support for all students, but individual and contextual differences should be acknowledged in lad design (teasley, 2017). using la can improve students’ self-efficacy beliefs and academic performance ( yilmaz, 2022; russell et al., 2020). however, individual differences in students’ capacities to utilize provided feedback productively should be further investigated. 5.1 limitations this study has some limitations that should be considered when interpreting the results. first, the study aimed to identify how student evaluations divide within the phases of srl. however, the research design did not address if the dashboards supported the fourth phase of srl, namely adaptation of the plan. this phase is an essential part of srl process during which students adapt their goals, plans and behaviors in terms of their evaluation and monitoring of their learning. thus, future studies should address how students evaluate such dashboards supporting all four phases of the srl. in the qualitative content analysis, the response rates to open-ended questions remained rather low, especially in case 1. in the t-test analyses, only three categories were identified as comparable between the cases. due to this, only goal setting and monitoring phases of srl were included in the t-tests. the variation in responses to assessing their self-efficacy beliefs and resource management strategies was relatively small, and students overall rated their skills as quite high. these limitations might influence the reliability of the findings. furthermore, the study has some limitations regarding the generalizability of the findings. as the study focused on the participants’ unique characteristics, self-efficacy beliefs, and help-seeking skills, and the student evaluations focused on the two specific lads, the findings may be unique to the selected context and student population, and they may not be generalizable to other educational settings or student populations. 6. conclusions there is a need for la tools that can help students carrying out their studies successfully (ifenthaler & yau, 2020). the results offer two contextualized examples of lads as a support for students’ selfregulation on the level of academic paths. the findings showed that students considered visualizations as a support for planning and monitoring useful and relevant, but they raised multiple information needs, which highlighted the need for more detailed la feedback for academic path-level contexts. this means identifying new indicators and addressing new sources of data to enrich the provided feedback on the academic path level. for example, combining course-level data with existing registry data could provide access to visualize silvola et al. 95 | f l r long-term srl processes and help students to improve their srl skills across the study periods and academic years. in open, hybrid, and online environments, a need for self-regulated learning is increased (jansen et al., 2020). lads can foster students’ srl (viberg et al., 2020). theory-informed, contextualized understanding of student perspectives is important in creating lads that encourage students to take ownership of their studies (ochoa & wise, 2021; stenalt, 2021). this study provided an understanding of how the srl process occurs on the academic path-level. for example, the requirements for context understanding support, and monitoring processes appeared different from the task-specific process of srl. supporting students’ srl in the context of academic paths can help students to improve their academic performance, but it also builds a ground for the development of their competencies as future professionals. the use of student-facing lads requires careful planning and consideration of the practices in which the lads are used. students with low self-efficacy beliefs reported more challenges in goal setting phase than the students with high self-efficacy beliefs. it is important for teachers and academic advisors to acknowledge the possible differences in students’ capacities to utilize provided feedback independently. thus, guiding students’ reflection process of the provided feedback might help different students to benefit from provided visualizations. a better understanding of student experiences on online learning technologies helps to understand how these technologies influence students’ academic outcomes and support their agency (stenalt, 2021). keypoints this study investigates student evaluations of two different lads as a support for academic path-level self-regulated learning (srl). analyticsai focused more on study planning support and powerbi on study monitoring support. high self-efficacy beliefs associated with perceived support for goal setting. low helpseeking skills associated with perceived support for study monitoring. student-facing lads can be helpful for students to steer their academic paths and refine their thinking about their study processes. more research of student experiences on academic path-level lads would be needed. acknowledgments this work was supported by the ministry of education and culture, finland [grant number okm/272/523/2017]. suomen kulttuurirahasto, finnish cultural foundation olvi-säätiö, olvi-foundation silvola et al. 96 | f l r references aguilar, s. j., & baek, c. (2019). motivated information seeking and graph comprehension among college students. acm international conference proceeding series, 280–289. https://doi.org/10.1145/3303772.3303805 bandura, a. (1977). self-efficacy: toward a unifying theory of behavioral change. psychological review, 84(2), 191–215. https://doi.org/10.1037/0033-295x.84.2.191 broadbent, j. (2017). comparing online and blended learner’s self-regulated learning strategies and academic performance. the internet and higher education, 33, 24–32. https://doi.org/10.1016/j.iheduc.2017.01.004 chi, m. t. h. (1997). quantifying qualitative analyses of verbal data: a practical guide. journal of the learning sciences, 6(3), 271–315. https://doi.org/10.1207/s15327809jls0603_1 dresel, m., schmitz, b., schober, b., spiel, c., ziegler, a., engelschalk, t., jöstl, g., klug, j., roth, a., wimmer, b., & steuer, g. (2015). competencies for successful self-regulated learning in higher education: structural model and indications drawn from expert interviews. 40(3), 454–470. https://doi.org/10.1080/03075079.2015.1004236 gedrimienė, e., silvola, a., kokkonen, h., tamminen, s., & muukkonen, h. (2022). addressing students’ needs: development of a learning analytics tool for academic path level regulation. in e. de vries, y. hod & j. ahn (eds). proceedings of the 15th international conference of the learning sciences. https://doi.org/10.1016/j.chb.2014.07.013 greene, j. a., & azevedo, r. (2007). a theoretical review of winne and hadwin’s model of self-regulated learning: new perspectives and directions. review of educational research, 77(3), 334–372. https://doi.org/10.3102/003465430303953 gutiérrez, f., seipp, k., ochoa, x., chiluiza, k., de laet, t., & verbert, k. (2018). lada: a learning analytics dashboard for academic advising. computers in human behavior, 107, 105826. https://doi.org/10.1016/j.chb.2018.12.004 hacker, d.j., dunlosky, j., & graesser, a.c. (eds.). (1998). metacognition in educational theory and practice (1st ed.). routledge. https://doi.org/10.4324/9781410602350 heikkinen, s., saqr, m., malmberg, j., & tedre, m. (2023). supporting self-regulated learning with learning analytics interventions – a systematic literature review. education and information technologies, 28(3), 3059–3088. https://doi.org/10.1007/s10639-022-11281-4/ ifenthaler, d., & yau, j. y. k. (2020). utilising learning analytics to support study success in higher education: a systematic review. educational technology research and development, 68(4), 1961–1990. https://doi.org/10.1007/s11423-020-09788-z jansen, r., jansen, r. s., van leeuwen, a., janssen, j., & kester, l. (2020). a mixed method approach to studying self-regulated learning in moocs: combining trace data with interviews. frontline learning research, 8(2), 35–64. https://doi.org/10.14786/flr.v8i2.539 järvelä, s., hurme, t.-r., & järvenoja, h. (2011). selfregulation and motivation in computersupported collaborative learning. in ludvigsen, s., lund, a., rasmussen, i., & säljö, r. (eds.). learning across sites: new tools, infrastructures and practices (1st ed., pp. 330–346). routledge. https://doi.org/10.4324/9780203847817 jivet, i., scheffel, m., drachsler, h., & specht, m. (2017). awareness is not enough: pitfalls of learning analytics dashboards in the educational practice. lecture notes in computer science (including subseries lecture notes in artificial intelligence and lecture notes in bioinformatics), 10474 lncs, 82–96. https://doi.org/10.1007/978-3-319-66610-5_7 jivet, i., scheffel, m., specht, m., & drachsler, h. (2018). license to evaluate: preparing learning analytics dashboards for educational practice. acm international conference proceeding series, 31–40. https://doi.org/10.1145/3170358.3170421 jivet, i., scheffel, m., schmitz, m., robbers, s., specht, m., & drachsler, h. (2020). from students with love: an empirical study on learner goals, self-regulated learning and sense-making of learning analytics in higher education. internet and higher education, 47. https://doi.org/10.1016/j.iheduc.2020.100758 silvola et al. 97 | f l r yilmaz, f. g. (2022). utilizing learning analytics to support students’ academic self-efficacy and problemsolving skills. asia-pacific education researcher, 31(2), 175–191. https://doi.org/10.1007/s40299-02000548-4/tables/8 khiat, h. (2019). using automated time management enablers to improve self-regulated learning. https://doi.org/10.1177/1469787419866304 kyndt, e., donche, v., trigwell, k., & lindblom-ylänne, s. (eds.). (2017). higher education transitions: theory and research (1st ed.). routledge. https://doi.org/10.4324/9781315617367 lahn, l. c. (2011). profesisonal learning as epistemic trajectories. in s. ludvigsen, a. lund, i. rasmussen, & r. säljö (eds.), learning across sites: new tools, infrastructures and practices (1st ed., pp. 53–68). routledge. lang, c., siemens, g., wise, a., gasevic, d. & merceron, a. (eds.). (2022). handbook of learning analytics (2nd ed.). society for learning analytics research. https://doi.org/10.18608/hla22 li, q., xu, d., baker, r., holton, a., & warschauer, m. (2022). can student-facing analytics improve online students’ effort and success by affecting how they explain the cause of past performance? computers & education, 185, 104517. https://doi.org/10.1016/j.compedu.2022.104517 lim, l. a., dawson, s., gaševic, d., joksimović, s., fudge, a., pardo, a., & gentili, s. (2020). students’ sensemaking of personalised feedback based on learning analytics. australasian journal of educational technology, 36(6), 15–33. https://doi.org/10.14742/ajet.6370 ludvigsen, s., lund, a., rasmussen, i., & säljö, r. (eds.). (2010). learning across sites: new tools, infrastructures and practices (1st ed.). routledge. https://doi.org/10.4324/9780203847817 ludvigsen, s., rasmussen, i., krange, i., moen, a., & middleton, d. (2011). intersecting trajectories of participation: temporality and learning. in s. ludvigsen, a. lund, i. rasmussen, & r. säljö (eds.), learning across sites: new tools, infrastructures and practices (1st editions, pp. 105–121). routledge. https://doi.org/10.4324/9780203847817 matcha, w., ahmad uzir, n., gasevic, d., & pardo, a. (2019). a systematic review of empirical studies on learning analytics dashboards: a self-regulated learning perspective. ieee transactions on learning technologies, 13(2), 226–245. https://doi.org/10.1109/tlt.2019.2916802 mendez, g., galárraga, l., & chiluiza, k. (2021). showing academic performance predictions during term planning: effects on students’ decisions, behaviors, and preferences. in proceedings of the 2021 chi conference on human factors in computing systems, 22, 1–17. https://doi.org/10.1145/3411764.3445718 ochoa, x., & wise, a. f. (2021). supporting the shift to digital with student-centered learning analytics. educational technology research and development, 69(1), 357–361. https://doi.org/10.1007/s11423-02009882-2/metrics panadero, e. (2017). a review of self-regulated learning: six models and four directions for research. frontiers in psychology, 8, 422–422. https://doi.org/10.3389/fpsyg.2017.00422 pelletier, k., mccormack, m., reeves, j., robert, j., arbino, n., al-freih, m., dickson-deane, c., guevara, c., koster, l., sánchez-mendiola, m., skallerup bessette, l., & stine, j. (2022). 2022 educause horizon report | teaching and learning edition | educause library. educause horizon report. https://library.educause.edu/resources/2022/4/2022-educause-horizon-report-teaching-and-learning-edition pintrich, p. r. r., smith, d., garcia, t., & mckeachie, w. (1991). a manual for the use of the motivated strategies for learning questionnaire (mslq). national center for research to improve postsecondary teaching and learning. https://doi.org/ed338122. rets, i., herodotou, c., bayer, v., hlosta, m., & rienties, b. (2021). exploring critical factors of the perceived usefulness of a learning analytics dashboard for distance university students. international journal of educational technology in higher education, 18(1), 1–23. russell, j. e., smith, a., & larsen, r. (2020). elements of success: supporting at-risk student resilience through learning analytics. computers and education, 152. https://doi.org/10.1016/j.compedu.2020.103890 schumacher, c., & ifenthaler, d. (2018). features students really expect from learning analytics. computers in human behavior, 78, 397–407. https://doi.org/10.1016/j.chb.2017.06.030 schunk, d. h. (1991). self-efficacy and academic motivation. educational psychologist, 26(3–4), 207–231. https://doi.org/https://doi.org/10.1207/s15326985ep2603&4_2 schunk, d.h., & greene, j.a. (eds.). (2017). handbook of self-regulation of learning and performance (vols. 1–2). routledge. https://doi.org/10.4324/9781315697048 silvola et al. 98 | f l r silvola, a., näykki, p., kaveri, a., & muukkonen, h. (2021a). expectations for supporting student engagement with learning analytics: an academic path perspective. computers & education, 168, 104192. https://doi.org/10.1016/j.compedu.2021.104192 silvola, a., sjöblom, a., svensk, s., lallimo, j., näykki, p. & muukkonen, h. (2021b, august 23-27). student engagement across academic paths: using la as a support for study planning and monitoring [conference presentation abstract]. the19th biennial earli conference for research on learning and instruction. https://earli.org/assets/files/boa-2021.pdf stenalt, m. h. (2021). digital student agency: approaching agency in digital contexts from a critical perspective. frontline learning research, 9(3), 52–68. https://doi.org/10.14786/flr.v9i3.697 teasley, s. d. (2017). student facing dashboards: one size fits all? technology, knowledge and learning, 22(3), 377–384. https://doi.org/10.1007/s10758-017-9314-3 tempelaar, d. (2019). supporting the less-adaptive student: the role of learning analytics, formative assessment and blended learning. assessment and evaluation in higher education 45(4), 579–593. https://doi.org/10.1080/02602938.2019.1677855 van leeuwen, a., teasley, s. d., & wise, a. f. (2022). teacher and student facing learning analytics. in lang, c., siemens, g., wise, a., gasevic, d. & merceron, a. (eds.). (2022). handbook of learning analytics (2nd ed., pp 130-140). society for learning analytics research. https://doi.org/10.18608/hla22 vanthornout, g., catrysse, l., coertjens, l., gijbels, d., & donche, v. (2017). the development of learning strategies in higher education: impact of gender and prior education. in e. kyndt, v. donche, k. trigwell, & s. lindblom-ylänne (eds.), higher education transitions: theory and practice (1st edition, pp. 1–322). routledge. viberg, o., khalil, m., & baars, m. (2020). self-regulated learning and learning analytics in online learning environments: a review of empirical research. proceedings of the tenth international conference on learning analytics & knowledge, 524–533. https://doi.org/10.1145/3375462.3375483 white, e. r. (2015). academic advising in higher education: a place at the core. the journal of general education, 64(4), 263–277. winne, p. (2011). a cognitive and metacognitive analysis of self-regulated learning. in schunk, d.h., & greene, j.a. (eds.). handbook of self-regulation of learning and performance (1st ed., pp.15–32) routledge. https://doi.org/10.4324/9781315697048 winne, p., & hadwin, a. (1998). studying as self-regulated learning. in d. j. hacker, j. dunlosky, & a. c. graesser (eds.), metacognition in educational theory and practice (pp. 277–304). lawrence erlbaum associates publishers. yin, r. k. (2009). case study research: design and methods. in case study research: design and methods (4th ed). sage publications. zimmerman, b. j. (2000). self-efficacy: an essential motive to learn. contemporary educational psychology, 25(1), 82–91. https://doi.org/10.1006/ceps.1999.1016 frontline learning research special issue: vol. 13 no. 2 (2025) 51 66 issn 2295-3159 corresponding author: jennifer e. symonds, ioe, ucl’s faculty of education and society, london, united kingdom, j.symonds@ucl.ac.uk doi: https://doi.org/10.14786/flr.v13i2.1431 children’s momentary behavioural engagement and class size: a national systematic observation study jennifer e. symonds1, ricardo böheim2, matthew p. somerville1, edward baines1, xin tang3, niamh oeri4, raven rinas5, florian jonas buehler4, gertraud benke6, aisling davies7, seaneen sloan7, dympna devine7& gabriela martinez sainz7 1university college london, ioe, ucl’s faculty of education and society, united kingdom 2technical university of munich, school of social sciences and technology, germany 3shanghai jiao tong university, school of education, china 4university of bern, institute for psychology, switzerland 5universität augsburg, lehrstuhl für psychologie, germany 6university of klagenfurt, institute for teaching and school development, austria 7university college dublin, school of education, dublin, ireland article received 29 january 2024 / article revised 12 october 2024 / accepted 15 october 2024/ available online 14 march 2025 abstract this study used systematic observation to test the direct and moderating effects of class size on children’s momentary behavioural engagement in learning. data were collected with 632 children (50.6% girls) in 121 classrooms in 92 schools recruited into the children’s school lives national cohort study of irish primary schooling. the observational and research classroom learning evaluation (oracle) systematic observation tool was used to observe individual children’s behaviour at 30-seconds intervals across a five-minute period in ordinary lessons of english, mathematics, science and irish. multilevel path models identified that behavioural engagement was higher in smaller classes and behavioural disengagement was higher in larger classes. class size also moderated the impact of several individual differences and classroom composition factors on momentary behavioural engagement. for example, smaller classrooms protected lower ability children from disengaging whereas higher ability children were more likely to stay engaged in larger classes compared to lower ability children. implications for research, practice and policy are discussed. keywords: ability; class size; behavioural engagement; observation. mailto:j.symonds@ucl.ac.uk symonds et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 52 | flr 1. children’s momentary behavioural engagement and class size: a national systematic observation study children’s behavioural engagement in classrooms is a gateway to learning and achievement. more time spent on task allows individual children to acquire knowledge over a longer period, facilitating their academic performance (bryce et al., 2019). scoping research on student engagement has identified that most behavioural engagement measurements are student self-reports of how they tend to behave in academic environments. very few studies have focused on engagement by observing how student behaviour proceeds across seconds and minutes in classroom settings (salmela-aro et al., 2021), a phenomena which we refer to as momentary behavioural engagement. furthermore, studies that have observed momentary behavioural engagement tend to use samples which are fairly small and nonrepresentative, due to the intensive resources required to observe individual children in class. these gaps in the evidence base have facilitated a lack of understanding of what children’s momentary behavioural engagement in classrooms looks like across a population of individuals. adding to the lack of information, there are few investigations of children’s engagement in learning in relation to classroom composition factors such as the average socioeconomic status of children in the classroom. this is because collecting data to achieve enough statistical power to test the impact of factors that vary at the classroom level is resource intensive. even cohort studies that sample at the classroom or school level rarely recruit all children in the class, meaning that classroom composition factors are rarely able to be ascertained. research that finds that gender, age, and cultural background impact student engagement at the individual level (archambault et al., 2013; chiu et al., 2012; montroy et al., 2016). therefore, classroom composition factors that could be important for engagement include the proportion of girls versus boys in the class, the average age of children in the classroom, and the proportion of migrant children in the class. a further important classroom composition factor is class size, which can vary widely within education systems due to differences in local conditions (hanushek & woessmann, 2017). in their comprehensive review of class size research, blatchford and russell (2020) propose that class size impacts student engagement not directly, but through other aspects of classroom context and process. in their class size and process model, they propose that other classroom composition factors, such as the proportion of children with special educational needs in the classroom and the ability level of pupils in the classroom, shape children’s engagement in learning differentially according to class size. however, without any large-scale study to examine these assumptions, this part of the model has never been tested. to address the lack of knowledge on children’s momentary behavioural engagement in a population and how it relates to individual and classroom composition factors, the current study collected systematic observation data with a national sample of children in 121 irish primary classrooms. the novel contributions of the study are the first large-scale systematic observation of children’s behavioural engagement in irish primary schools, and an empirical test of the moderating impact of classroom and individual factors on the relationship between class size and children’s momentary behavioural engagement. 2. momentary behavioural engagement in classrooms behavioural engagement refers to children’s participation, concentration, compliance, persistence, and perseverance in the task at hand (skinner, 2016). when behavioural engagement is studied at a fine-grained level across seconds to minutes, this is referred to as momentary behavioural engagement (symonds et al., 2021). momentary behavioural engagement is conceptualised as part of a broader dynamic system of emotion, motivation, and cognition that converges into a state of dynamic stability when children are focused and engaged with their classwork (symonds et al., 2021). the symonds et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 53 | flr psychological processes underpinning behavioural engagement in classrooms include self-regulation of emotions, impulses, and motor skills, enabling children to manage and modulate their behaviour according to situational demands (e.g., schunk & zimmerman, 2023). research shows that behavioural self-regulation skills are critical not only for a successful school transition but also for academic development (see, e.g., rimm-kaufman & wanless, 2012). positive relations between behavioural regulation have been reported for different academic domains such as literacy and language, mathematics, science, and information and communication technology (edossa et al., 2018; gestsdottir et al., 2014). longitudinal research also suggests that behavioural regulation predicts academic growth in reading (e.g., newman et al., 1998) and mathematical skills (e.g., robinson & mueller, 2014). furthermore, the value to individual children of being able to regulate their behaviour in classrooms interlinks with the development of effective learning strategies, which in turn can positively influence their educational trajectories in childhood and adolescence (e.g., matthews et al., 2009). 3. individual differences in momentary behavioural engagement multiple factors, such as age, gender, and migrant status and cognitive ability are found to affect children’s behavioural regulation and engagement at school. as children age, they tend to develop better self-regulation skills, for instance, executive functions and metacognitive skills notably develop between the ages of three and eight (montroy et al., 2016; roebers, 2017). moreover, girls are typically more advanced in self-regulation skills than boys (montroy et al., 2016) and are more engaged in the classroom (archambault et al., 2013). regarding migrant status, large-scale assessments in oecd and european countries revealed that student engagement is often higher in immigrants than native students (chiu et al., 2012). student engagement is also a predictor for immigrant children’s academic resilience (martin et al., 2022). finally, general cognitive abilities, such as effortful control (yang & lamb, 2014), and attention (pagani et al., 2012), but also more specific cognitive abilities, such as vocabulary and number knowledge (archambault et al., 2013), predict classroom engagement in a broad age range from kindergarten to sixth grade. taken together, a higher age, being female, higher cognitive abilities, and immigrant status seem to promote behavioural regulation and engagement in classrooms. 4. class size and momentary behavioural engagement research on class size tends to focus on cognitive skills such as student achievement (blatchford & russell, 2020). although there is intense interest in this topic, a study of 649 elementary school classrooms across the united states found no relationship between class size and academic test scores (hattie, 2008; hoxby, 2000), signalling that research might be productively directed elsewhere. hence, it is important for studies to look beyond student achievement outcomes towards the impact of class size on children’s psychosocial functioning (blatchford, 2021) such as their emotional and behavioural regulation and engagement. research on class size finds that smaller classrooms allow teachers to interact more frequently with individual students, share resources more evenly, deploy creative tasks, and more easily manage their classrooms (blatchford & russell, 2020). these features of teaching in smaller classes are proposed to facilitate children’s behavioural engagement. however, the benefits of smaller classes do not occur by default. quality teaching is required to translate the potential advantages of smaller class sizes into enhanced educational experiences for children (hanushek & woessmann, 2017). in education systems where class size has sufficient variation for statistical modelling (e.g., 8 to 34 children; hoxby 2000), it is possible to identify systematic links between class size and outcomes including children’s behavioural engagement. however, any observed links must be understood in the context of related classroom factors which might mediate the impact of class size on behavioural outcomes. symonds et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 54 | flr 5. class size, classroom composition factors, and momentary behavioural engagement blatchford and russell (2020) show that the empirical literature on the effect of class size on student achievement is riddled with inconsistencies. they argue that class size is a contextual variable that is generally mediated by many other factors which can also be contextual. in their model, they group these factors into three main categories (p. 263), namely contexts, classroom processes, and effects on teachers and pupils. contexts, which includes class size, additionally includes (a) (available) time, (b) types of students, (c) physical characteristics of the environment, (d) curriculum and assessment arrangements, and (e) interactive contexts as designed by the teacher (e.g. individual work). in their monograph, they discuss at length how each factor mediates class size. their summary model synthesises the factors identified in other research to explain how class size affects children's engagement in learning in one comprehensive framework. a crucial element of the model is the composition of the pupils in the classroom, which we refer to as 'classroom composition factors'. for this factor, blatchford and russell (2020, 229) note: “unfortunately, there is surprisingly little systematic research available on fundamental aspects of the classroom support in place for pupils with send [special educational needs and disabilities].” as we discussed in the introduction, these factors are rarely tested in educational psychology because it is difficult to obtain sufficient variation in classrooms for robust statistical testing. accordingly, there are no studies known to the authors which examine whether key classroom composition factors explain the potential connection between class size and children’s momentary behavioural engagement in learning. this leaves a theoretical and empirical gap that the current study seeks to fill. identification of potential classroom composition factors comes from blatchford and russell’s (2020) review, and from research on the impact of children’s individual differences on behavioural engagement (archambault et al., 2013; chiu et al., 2012; montroy et al., 2016). factors that could interact with class size to impact behavioural engagement include ability (when ability levels become more diverse in larger classrooms), special educational needs (that are more difficult to cater to in larger classrooms with a single classroom teacher), gender (when there are more girls in the class if girls tend to be more engaged), migrant children (when there are more migrant children in classrooms in ireland where migrant children tend to be more engaged), socioeconomic disadvantage (when lowerincome schools face more challenges in promoting children’s engagement in learning) and age (when classrooms are comprised of older children and because of age those children are better behaviourally regulated). in general, blatchford and russell (2020) conclude that the composition of the classroom is of consequence, as different groups of children have disparate needs that require the attention and time of teachers. they observe that, on average, larger classes result in less interaction between each individual student and their teachers, and that teachers experience difficulty in differentiating and providing individual attention in larger classes. this may be compounded in classrooms with wider pupil ability ranges. 6. observation systems for studying momentary behavioural engagement in classrooms while the most widely used approach for studying student engagement has been self-report surveys (fredricks et al., 2016, salmela-aro et al., 2021), observational methods have also been extensively used, particularly when assessing individual behavioural engagement. although this approach may be resource intensive, observational methods are able to capture engagement across time and settings, offer a large degree of flexibility in terms of which engagement behaviours to target, and symonds et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 55 | flr provide detailed descriptive information regarding how engagement processes unfold across a wide range of classroom and school contexts. one approach to studying engagement through observational methods is using standardised rating scales of behaviour. the classroom assessment scoring system (class; hamre & pianta, 2010) is one such example. class requires observers to assess classroom quality by assigning a rating from 1 to 7 on a range of seven-point dimensions, including student engagement. these ratings are typically based on two 20-minute classroom observations. other researchers have used more qualitative approaches to observe engagement in the classroom. rubiedavies and colleagues (2010) examined motivation and engagement in relation to interactions between pupils and adults in the classroom. this approach captured rich, detailed descriptions of the quality of the classroom talk between adult and pupils and the extent to which these interactions promoted student engagement. finally, a time-sampling approach has also been applied by student engagement and motivation researchers. this involves assessing whether pre-specified behaviours (e.g., on-task behaviour) occur within set time intervals. observational data collected using time-sampling can be analysed quantitatively and are thus well suited to studies examining the effects of school and classroom level factors such as class size and school socioeconomic disadvantage. a widely cited example of this approach is the oracle (observational research and classroom learning evaluation) observation system (galton & hargreaves, 2019). the oracle system was originally developed for use in pre-schools in the 1960s and has since been adapted to capture detailed information on children’s behaviours, classroom activities, and teaching, by school transitions researchers working with large scale samples of children in english schools in the 1970s, 1990s, and 2000s (e.g., galton & wilcocks, 1983; hargreaves & galton, 2002). torsney and symonds (2019) used the oracle as their measure of momentary behavioural engagement in a small sample of irish socioeconomically disadvantaged secondary schools, finding that profiles of combined emotional, cognitive, and behavioural engagement were predicted by student gender, irish white ethnicity, academic self-efficacy, and peer support. researchers on the irish longitudinal cohort study of primary schooling, children’s school lives, further adapted the oracle to measure children’s behavioural engagement across a national sample of primary school classrooms in ireland, presenting the first opportunity to examine children’s behavioural engagement at a national level and the associated key individual and school level factors that might promote or inhibit momentary behavioural engagement in classrooms. 7. the current study given the significant resources required to observe children’s momentary behavioural engagement in classrooms and to test the impact of classroom composition effects on engagement, there are theoretical and empirical gaps in our understanding of the factors that could explain the impact of class size on engagement in learning. furthermore, although there have been studies of class size and children’s momentary behavioural engagement in the united kingdom (c.f., blatchford & webster, 2017) there has never been an examination of momentary behavioural engagement across irish schools. to address these knowledge gaps, the current study aims to (i) test the impact of class size on children’s momentary behavioural engagement in learning after accounting for individual and classroom composition differences that may systematically associate with class size, and (ii) examine whether those individual and classroom composition differences impact children’s momentary behavioural engagement as a function of class size. symonds et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 56 | flr 8. methods 8.1 participants and procedures one hundred primary schools were recruited using proportionate stratified random sampling to ensure national representation. classroom teachers selected up to six children in each classroom, balanced in gender, teacher-rated ability (low, middle, or high), and ethnic majority versus minority status. the final sample were 632 children (aged 7 – 9-years) in 121 classrooms in 92 schools. the 632 children were 50.6% female, 88.6% born in ireland, 33.3% low ability, 36.7% middle ability, and 36.7% high ability. trained fieldworkers visited each classroom for one day. during the visit, fieldworkers observed an ordinary lesson of english, maths, irish, or science (random allocation), administered pencil and paper questionnaires with children and collected teacher-on-child questionnaires. fieldwork spanned march to june 2019. the data collection was approved by the university college dublin human research ethics committee. the study was carried out as part of the broader children’s school lives study of primary schooling in ireland. 8.2 measures momentary behavioural engagement. each child was observed at 30-second intervals for five minutes, using the observational and research classroom learning evaluation (oracle) pupil record (galton & hargreaves, 2019). the observation order was determined by children’s last names (sequential) and gender (alternating). children’s main behaviour at each interval was coded as one of five forms of engagement: cooperating alone, cooperating with a friend, cooperating with the teacher, cooperating on routine tasks (e.g., sharpening pencils), and waiting for the teacher; or three forms of disengagement: distracted passive, distracted active, horseplay/disruptive. engagement and disengagement forms were summed to give overarching indicators of on task and off task behaviour. class size was reported by school principals prior to fieldwork and was confirmed by fieldworkers during the school visit via a fieldwork record sheet. class size was used as a continuous variable in the analysis, and as a categorical variable of smaller classes (1) and larger classes (2). the categorical variable was created by splitting the continuous class size variable at its mean of 21.7 children per classroom, so that classrooms of 22 children or above were designated as larger. the mean level split was in line with national irish class size data, where in 3092 primary schools in ireland, the average class size in 2024 was 21 children (department of education, nd). children’s age in years was collected in the child questionnaire. children’s gender (girl = 2, boy = 1) was collected in the child questionnaire and in the teacheron-child questionnaire. the three 7 – 8-year-old children reporting their gender as ‘other’ were coded as missing gender for data protection reasons. children’s migrant status (1 = not born in ireland, 0 = born in ireland) was collected in the child questionnaire. children’s family affluence was collected in the child questionnaire using four items from the family affluence scale (inchley et al. 2018). do you have a mobile phone? does your family own a car? how many computers/laptops does your family own? during the past year, how many times did you travel away on holiday with your family? the items were scored 0 (no/none) to 1 (mobile phone = yes) or 2 (cars, computers, holidays = more than two). the sum of items was used to represent children’s family affluence. symonds et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 57 | flr children’s special educational needs status was collected in the teacher-on-child questionnaire. teachers were asked if each child in the study was impacted by any of the irish national council on special educational needs (sen) categories of need: physical disability, visual or hearing impairment, speech impairment, autism, general learning disability, general learning disability, specific learning disability, emotional or behavioural problem, limited knowledge of the main language of instruction. children identified by teachers as having one or more needs were coded as sen = 1, versus children without identified needs coded as sen = 0. teacher perceptions of child ability (hereafter ‘ability’) were collected by asking teachers to identify whether children lower, average/moderate, or higher in overall academic ability in comparison to other children in the class. ability was scored as lower (1), moderate (2), and higher (3). it is clear that there is some unknown variability in teachers' estimation of students' abilities. as a caveat, it is likely that primary school teachers teaching language studies, mathematics, science and other topics have prioritised different subjects in their judgements. it is also possible that they are biased towards some student groups, with the latter bringing in some unknown systematic bias. classroom composition variables were created by computing the average value within each classroom for each variable of child gender, age, sen, ability, family affluence, and migrant status. 8.3 analysis plan data were processed using ibm spss statistics version 29.0.1.0 and mplus version 8.7. descriptive statistics were calculated for individual and classroom level variables, and all variables were correlated with each other using pearson correlations. next, multilevel models were computed with children clustered in classrooms. dependent variables were on task and off task behaviour. individual scores for each child used as predictors at level 1 were child age, gender, sen, ability, family affluence, and migrant status. aggregate scores for each classroom used as predictors at level 2 were the same factors as those modelled at level 1. class size was included as a predictor at level 2 in the first model for each of on task and off task behaviour. next, multigroup multilevel models were run for smaller and larger classes, using the same predictors except for class size. this method resulted in six multilevel models (1) on task behaviour in the whole sample, (2) on task behaviour in smaller classrooms (3) on task behaviour in larger classrooms, (3) off task behaviour in the whole sample, (4) off task behaviour in smaller classrooms, and (6) off task behaviour in larger classrooms. symonds et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 58 | flr table 1. study variables correlations and descriptive statistics variables 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 1 on task 1 2 off task -.83** 1 3 class size -.19** .17** 1 4 large vs. small class -.15** .14** .81** 1 5 age .02 -.03 -.01 -.04 1 6 girls .08* -.08* .03 .05 -.09* 1 7 sen -.07 .07 -.04 -.08 .09* -.14** 1 8 ability .15** -.18** -.04 -.03 .06 -.02 -.24** 1 9 family affluence .02 -.04 .02 .06 -.01 .01 -.10* .12** 1 10 migrant -.02 .03 -.03 -.04 -.06 -.00 -.03 -.04 -.02 1 11 age in class .06 -.07 -.02 -.07 .55** . .07 .02 .07 -.07 1 12 girls in class .05 -.03 .07 .12*** -.05 .46** -.14** -.00 .01 -.09* -.09* 1 13 sen in class .09* -.09* -.08 -.14*** .08 -.12** .51** .01 -.11* .04 .14** -.26** 1 14 ability in class .02 -.03 -.18** -.12*** .04 -.01 .03 .24** .06 .02 .08 -.01 .05 1 15 family affluence in class .01 -.02 .04 .12*** .07 .01 -.11* .03 .53** -.11** .13** .03 -.22** .12** 1 -.22** 16 migrants in class -.01 .01 -.08 -.08* -.08* -.08* .04 .01 -.11** .51** -.14** -.17** .08 .02 -.22** 1 n 632 632 623 623 627 632 561 622 587 607 632 632 586 626 632 632 minimum 0.0 0.0 2.0 1.0 7.0 1.0 0.0 1.0 3.0 1.0 7.0 1.0 0.0 1.3 5.0 1.0 maximum 10.0 10.0 34.0 2.0 9.0 3.0 1.0 3.0 10.0 2.0 8.7 2.0 0.8 3.0 10.0 1.8 m 7.2 1.8 21.7 1.6 8.1 1.5 0.1 2.0 8.7 1.1 8.1 1.5 0.1 2.0 8.7 1.1 sd 2.6 2.1 6.6 0.4 0.5 0.3 0.8 1.3 0.3 0.2 0.2 0.2 0.2 0.7 0.2 % 56.7 50.6 13.2 11.4 note. *p < .05. **p < .01 symonds et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 59 | flr 9. results the average class size was 21.7 children (sd = 6.6). there was plenty of variance in the smaller versus larger classes variable, with 270 of the 623 children being in smaller classes (43.3%) and 353 children being in larger classes (56.7%). on average, children were observed as having on task behaviour for approximately 70% of the time during the five minutes they were observed (m = 7.2, sd = 2.6), and as having off task behaviour for approximately 20% of the time (m = 1.8, sd = 2.1). for more descriptive statistics and correlations between variables please see table 1. the multilevel models fit the data well, with a rmsea of .02, cfi of .99 and tli of .93 for both whole sample models, and a rmsea of .0, and a cfi and tli of 1 for the multigroup models. the models (tables 2 and 3) identified that on task behaviour was lower in larger classes, and that off task behaviour was higher in larger classes. put simply, more instances of momentary behavioural engagement were observed in smaller classes. class size had the strongest impact on momentary behavioural engagement compared to all other factors in the models. the next most important factor was sen. individual child sen predicted less on task behaviour and more off task behaviour across classrooms, suggesting that individual sen is a risk factor for momentary behavioural engagement. however, classrooms comprising more children with sen predicted less off task behaviour, when the classrooms were larger. ability was also an important factor. individual child ability predicted more on task behaviour and less off task behaviour across classrooms. however, the multigroup models revealed that this effect was mainly manifest in larger classrooms, where higher ability was a protective factor for momentary behavioural engagement, and lower ability was a risk factor for momentary behavioural disengagement. the impact of family affluence was only apparent at the classroom level when measured as a classroom composition factor. here, classrooms with more affluent children predicted less off task behaviour in smaller classrooms, and more off task behaviour in larger classrooms. gender had one effect, with girls being more likely to be on task when gender was modelled at the individual level. age and migrant status had no impact in any of the models. symonds et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 60 | flr table 2. multilevel models for behavioural engagement predictor whole sample smaller classes larger classes b se t p b se t p b se t p level 1 age -0.01 0.05 -0.20 0.841 0.02 0.07 0.24 0.808 -0.03 0.07 -0.35 0.726 girls 0.07 0.04 1.70 0.089 0.09 0.07 1.25 0.212 0.07 0.06 1.29 0.197 sen -0.13 0.04 -3.08 0.002 -0.16 0.07 -2.10 0.036 -0.12 0.05 -2.29 0.022 ability 0.14 0.05 2.81 0.005 0.09 0.09 0.97 0.331 0.17 0.06 2.88 0.004 family affluence 0.01 0.05 0.11 0.916 -0.01 0.06 -0.11 0.917 0.02 0.06 0.36 0.721 migrant -0.02 0.04 -0.52 0.603 0.05 0.06 0.80 0.425 -0.07 0.06 -1.24 0.216 level 2 age in class 0.12 0.11 1.16 0.248 0.20 0.18 1.11 0.266 0.01 0.13 0.09 0.930 girls in class 0.11 0.11 0.97 0.332 0.29 0.18 1.59 0.113 0.01 0.16 0.08 0.936 sen in class 0.30 0.11 2.82 0.005 0.37 0.18 2.12 0.034 0.35 0.17 2.09 0.037 ability in class -0.11 0.10 -1.13 0.260 -0.19 0.18 -1.06 0.289 -0.04 0.16 -0.23 0.817 family affluence in class 0.06 0.13 0.44 0.660 0.43 0.19 2.22 0.026 -0.22 0.16 -1.39 0.165 migrants in class 0.01 0.11 0.09 0.926 -0.05 0.16 -0.34 0.732 0.05 0.16 0.33 0.744 class size -0.43 0.11 -3.95 0.000 note. sen = special educational needs symonds et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 61 | flr table 3. multilevel models for behavioural disengagement predictor whole sample smaller classes larger classes b se t p b se t p b se t p level 1 age 0.02 0.06 0.30 0.763 0.00 0.08 0.01 0.989 0.03 0.08 0.37 0.709 girls -0.09 0.04 -2.10 0.036 -0.13 0.07 -1.79 0.074 -0.07 0.05 -1.38 0.168 sen 0.11 0.05 2.16 0.031 0.18 0.06 2.91 0.004 0.08 0.07 1.12 0.264 ability -0.17 0.05 -3.71 0.000 -0.07 0.07 -1.01 0.313 -0.22 0.06 -3.76 0.000 family affluence -0.01 0.04 -0.23 0.821 0.03 0.06 0.49 0.624 -0.03 0.05 -0.63 0.530 migrant 0.03 0.05 0.72 0.471 -0.03 0.07 -0.38 0.706 0.08 0.06 1.26 0.210 level 2 age in class -0.18 0.14 -1.22 0.222 -0.31 0.19 -1.58 0.114 0.04 0.23 0.17 0.869 girls in class -0.05 0.13 -0.41 0.683 -0.19 0.22 -0.90 0.370 -0.01 0.20 -0.05 0.961 sen in class -0.36 0.11 -3.15 0.002 -0.30 0.18 -1.65 0.098 -0.64 0.20 -3.28 0.001 ability in class 0.15 0.11 1.37 0.172 0.16 0.19 0.86 0.391 0.07 0.17 0.39 0.699 family affluence in class -0.10 0.17 -0.57 0.571 -0.51 0.19 -2.63 0.009 0.47 0.23 2.02 0.043 migrants in class -0.03 0.13 -0.24 0.809 -0.09 0.15 -0.63 0.527 0.12 0.23 0.51 0.610 class size 0.50 0.13 3.97 0.000 note. sen = special educational needs symonds et al 62 | f l r 10. discussion children’s momentary behavioural engagement in classrooms is rarely studied across representative samples of populations of individuals, because of the significant resources needed to capture observational data on children’s behaviour occurring across seconds to minutes in classrooms. similarly, the impact of classroom composition factors on children’s momentary behavioural engagement is understudied given the scarcity of studies with more than a handful of classrooms and that recruit children at the classroom level. the current study addressed these empirical gaps, and approached filling a theoretical gap on how classroom composition factors might help explain the impact of class size on children’s engagement, using an irish national sample of children in primary schools observed every 30-seconds for ten intervals (5-minutes total per child) in lessons of english, irish, science, and mathematics. multilevel models identified that children in smaller classes had higher levels of momentary behavioural engagement and that children in larger classes had higher levels of momentary behavioural disengagement. the models also identified that having sen was a risk factor for individual children’s momentary behavioural engagement, but that larger classrooms with a greater number of children with sen were protective for children’s momentary behavioural disengagement. higher individual ability was a protective factor for children in larger classrooms whereas lower ability was a risk factor for children in larger classrooms. classrooms with more affluent children promoted behavioural engagement if the class size was smaller but promoted behavioural disengagement if the class size was larger. finally, girls were more likely to be behaviourally engaged, but there was no impact of the numbers of boys and girls in the classroom on behavioural engagement. class size effects on student outcomes are documented internationally (blatchford & russell, 2020). our study of a national irish sample identified a consistent relationship between larger classes and lower student behavioural engagement in class. of note, this is the first time that momentary behavioural engagement has been studied large-scale in ireland and related to class size. future research might explore this relationship further, by associating class size with student cognitive and emotional engagement. this would allow the irish context to be compared to other nations, where research from the trends in international mathematics and science study (timss) has identified that student enjoyment of learning is higher in smaller classes in some countries (shen et al., 2019). in this study, children’s behavioural disengagement was lower in larger classes, if those classes had a larger number of children with sen. in ireland, even though there are special schools, the state is committed to inclusive education and most children with sen are educated in mainstream schools alongside their peers without identified sen. each school is assessed to identify the level of need across students and is allocated additional resources including special education teachers and special needs assistant teachers who give support where needed (shevlin & banks, 2021). having a larger number of children with sen in larger classrooms would signal greater need in the school therefore these children may have been receiving support from additional adults in class which may have positively impacted their behavioural engagement. this result may signal the impact of an unmeasured factor, such as additional educational staff supporting children with sen, teachers better differentiating their work, or more careful management of the speed of learning, in classrooms with the highest levels of need. however, we also found that having sen was a risk factor for individual children’s momentary behavioural engagement. this result probably stems from the larger proportion of children with social and emotional behavioural difficulties (sebd) in the sen group, as children with sebd typically have challenges with behavioural regulation (skalická et al., 2015). these self-regulatory challenges are not characteristic of many other categories of need studied, for example, vision and hearing impairments, and general learning difficulty (e.g., dyslexia). future research should separate the categories of need into sen with and without behavioural difficulties for a more fine grained analysis. symonds et al 63 | f l r a further finding was the association between ability and behavioural (dis)engagement in larger classes. our teacher rated ability variable was scored 1 (low ability) to 3 (high ability), therefore the finding indicates that students of lower ability had greater behavioural disengagement and less behavioural engagement in larger classes, and vice versa for children of higher ability. this finding aligns with research from blatchford et al., (2011) where, in a similar sized study of 686 children, the authors found that children with low attainment had lower behavioural engagement in smaller classes, and higher behavioural disengagement in larger classes, and that children with high ability had lower behavioural disengagement in larger classes. put simply, larger class size is a risk factor for behavioural disengagement for lower ability children, whereas higher ability children appear to be less vulnerable to the negative impact of larger classes. possibly, teachers were interacting more with higher ability children in larger classes, as these children may have been better able to respond to the teacher’s cues and the teacher had less time to share their attention with the many other children in the room, indicating a selection effect with more students present. our result about children’s family affluence was interesting—that classrooms with more affluent children encouraged behavioural disengagement if classes were larger, and discouraged behavioural disengagement if they were smaller. likewise, classrooms with less affluent children discouraged behavioural disengagement if they were smaller, and encouraged behavioural disengagement if they were larger. in ireland, schools serving low-income communities are awarded additional resources through the delivering equality of opportunity in schools (deis) programme, which includes reducing class sizes in schools with higher levels of need. the impact of the deis programme on children in the current study could explain these effects of classroom composition of family affluence and behavioural disengagement. finally, we found that girls had lower levels of behavioural disengagement than boys, but that there were no differences between girls and boys in behavioural engagement. this result may be explained by the qualities of behavioural disengagement as measured in the current study as a combination of actively and passively distracted and horseplay. post-hoc analysis of the three behavioural disengagement indicators revealed that boys had higher levels of passive distraction (m = 1.18, sd = 1.56) compared to girls (m = .98, sd = 1.40), whereas their levels of active distraction (boys: m = 0.79, sd = 1.47; girls: m = 0.63, sd = 1.23) and horseplay were reasonably similar (boys: m = 0.03, sd = 0.20; girls: m = 0.02, sd = 0.18). when these three indicators were combined, the analysis detected a larger effect of gender on behavioural disengagement. this result aligns with other research that observes gender differences in children’s ability to self-regulate their behaviour (montroy et al., 2016). this study has a number of limitations. first, although smaller classes have been shown to support higher behavioural engagement, this study does not have the data to demonstrate how this translates into student achievement. further investigation into the impact of class size on momentary student behaviour and its subsequent influence on academic outcomes could provide a more comprehensive understanding. secondly, the study only considered first-generation migration status as the sole migration variable. while this is a significant factor, it may not fully account for the diverse range of migration-related influences on student behaviour. future research should incorporate additional variables, such as second-generation migration status or length of time since migration. while the present study did not yield any significant findings on migrant status, a more comprehensive assessment of migration status may potentially yield different results. thirdly, while the study indicates some significant relationships, for example that children with higher abilities are less affected by larger class sizes, the data do not allow for definitive interpretations. further research is required to determine the probable causes and mechanisms for the observed relationship. this could include investigating whether students in larger classes are, in fact, less affected or if teachers in these classes tend to allocate more attention to certain groups of students, thereby mitigating the impact of class size on academic performance for these students, but not their peers. symonds et al 64 | f l r 11. conclusion the main findings of our study are that behavioural disengagement is higher in larger classes and lower in smaller classes, and that behavioural engagement is higher in smaller classes and lower in larger classes. these findings are robust after accounting for children’s individual differences and classroom composition factors that could otherwise be responsible for individual variation in children’s behavioural (dis)engagement. a clear message for policy makers is to continue supporting schools through enabling smaller pupil-teacher ratios. by enabling smaller class sizes, policy makers can facilitate less disengagement in classrooms which can promote learning and educational equality. keypoints children’s momentary behavioural engagement is higher in smaller classes and lower in larger classes. children with lower ability are at risk for disengaging from learning in larger classrooms, but not in smaller classrooms. being a boy is a risk factor for disengaging from learning in classrooms of all sizes. having sen is a risk factor for disengaging from learning in classrooms of all sizes, possibly because of including sebd in the analysis. classrooms comprising a greater number of affluent children promote engagement in learning only when those classrooms are 20 pupils or smaller. references archambault, i., pagani, l. s., & fitzpatrick, c. (2013). transactional associations between classroom engagement and relations with teachers from first through fourth grade. learning and instruction, 23(1), 1–9. https://doi.org/10.1016/j.learninstruc.2012.09.003 blatchford, p. (2021). rethinking class size. a question and answer session with peter blatchford about a new book and approach to the class size issue. education 3-13, 49(4), 387-397. https://doi.org/10.1080/03004279.2021.1874370 blatchford, p., & webster, r. (2017). oracle to mast: 40 years of observation studies in uk junior school classrooms. in r. maclean (ed.), life in schools and classrooms: past, present and future (pp. 169-185). springer singapore. https://doi.org/10.1007/978-981-10-3654-5_11 blatchford, p., & russell, a. (2020). rethinking class size: the complex story of impact on teaching and learning (p. 328). ucl press. bryce, c. i., bradley, r. h., abry, t., swanson, j., & thompson, m. s. (2019). parents’ and teachers’ academic influences, behavioral engagement, and firstand fifth-grade achievement. school psychology, 34(5), 492-502. https://doi.org/10.1037/spq0000297 chiu, m. m., pong, s. ling, mori, i., & chow, b. w. y. (2012). immigrant students’ emotional and cognitive engagement at school: a multilevel analysis of students in 41 countries. journal of youth and adolescence, 41(11), 1409–1425. https://doi.org/10.1007/s10964-012-9763-x/tables/4 department of education (2024, 12 october). class-size information at individual primary school level. https://www.gov.ie/en/collection/class-size-information-at-individual-primary-schoollevel/ https://doi.org/10.1016/j.learninstruc.2012.09.003 https://doi.org/10.1080/03004279.2021.1874370 https://doi.org/10.1007/978-981-10-3654-5_11 https://doi.org/10.1037/spq0000297 https://doi.org/10.1007/s10964-012-9763-x/tables/4 https://www.gov.ie/en/collection/class-size-information-at-individual-primary-school-level/ https://www.gov.ie/en/collection/class-size-information-at-individual-primary-school-level/ symonds et al 65 | f l r edossa, a. k., schroeders, u., weinert, s., & artelt, c. (2018). the development of emotional and behavioral self-regulation and their effects on academic achievement in childhood. international journal of behavioral development, 42(2), 192–202. https://doi.org/10.1177/0165025416687412 fredricks, j. a., filsecker, m., & lawson, m. a. (2016). student engagement, context, and adjustment: addressing definitional, measurement, and methodological issues. learning and instruction, 43, 1-4. galton, m., & wilcocks, j. (eds.). (1983). moving from the primary classroom. routledge and kegan paul. galton, m., & hargreaves, l. (2019). observational research and classroom learning evaluation (oracle): the pupil record. university of cambridge faculty of education. gestsdottir, s., von suchodoletz, a., wanless, s. b., hubert, b., guimard, p., birgisdottir, f., et al. (2014). early behavioral self-regulation, academic achievement, and gender: longitudinal findings from france, germany, and iceland. applied developmental science, 18, 90-109. https://doi.org/10.1080/10888691.2014.894870 hamre, b. k., & pianta, r. c. (2010). classroom environments and developmental processes: conceptualization and measurement. in handbook of research on schools, schooling and human development (pp. 25-41). routledge. hanushek, e. a., & woessmann, l. (2017). school resources and student achievement: a review of cross-country economic research. in m. rosén, k. y. hansen, & u. wolff (eds.), cognitive abilities and educational outcomes. springer international publishing. hattie, j. (2008). visible learning: a synthesis of over 800 meta-analyses relating to achievement. routledge. hargreaves, l., & galton, m. (2002). transfer from the primary classroom 20 years on. routledge falmer. hoxby, c. m. (2000). the effects of class size on student achievement: new evidence from population variation. the quarterly journal of economics, 115(4), 1239-1285. https://doi.org/10.1162/003355300555060 inchley, j., currie, d., cosma, a. & samdal, o. (2018). health behaviour in school-aged children (hbsc) study protocol: background, methodology and mandatory items for the 2017/18 survey. st andrews: cahru. martin, a. j., burns, e. c., collie, r. j., cutmore, m., macleod, s., & donlevy, v. (2022). the role of engagement in immigrant students’ academic resilience. learning and instruction, 82, 101650. https://doi.org/10.1016/j.learninstruc.2022.101650 matthews, j. m., cameron ponitz, c., & morrison, f. j. (2009). early gender differences in self regulation and academic achievement. journal of educational psychology, 101, 689-704. doi:10.1037/a0014240 montroy, ,j j, bowles, r. p., skibbe, l. e., mcclelland, m. m., & morrison, f. j. (2016). the development of self-regulation across early childhood. developmental psychology, 52(11), 1744–1762. https://doi.org/10.1037/dev0000159.supp newman, j., noel, a., chen, r., & matsopoulos, a. s. (1998). temperament, selected moderating variables and early reading achievement. journal of school psychology, 36(2), 215–232. https://doi.org/10.1016/s0022-4405(98)00006-5 pagani, l. s., fitzpatrick, c., parent, s., pagani, l. s., fitzpatrick, c., & parent, : s. (2012). relating kindergarten attention to subsequent developmental pathways of classroom engagement in elementary school. j abnorm child psychol, 40, 715–725. https://doi.org/10.1007/s10802-011-9605-4 rimm-kaufman, s.e., & wanless, s. b. (2012). self-regulation and academic achievement, in r. pianta, l. justice, barnett, s., sheridan, s.m. (eds.), the handbook of early education (pp. 299-323). new york: guilford publications. robinson, k., & mueller, a. s. (2014). behavioral engagement in learning and math achievement over kindergarten: a contextual analysis. american journal of education, 120, 325-349. https://doi.org/10.1086/675530 roebers, c. m. (2017). executive function and metacognition: towards a unifying framework of cognitive self-regulation. developmental review, 45, 31–51. https://doi.org/10.1016/j.dr.2017.04.001 rubie-davies, c. m., blatchford, p., webster, r., koutsoubou, m., & bassett, p. (2010). enhancing learning? a comparison of teacher and teaching assistant interactions with pupils. school effectiveness and school improvement, 21(4), 429-449. https://doi.org/10.1080/09243453.2010.512800 https://doi.org/10.1177/0165025416687412 https://doi.org/10.1080/10888691.2014.894870 https://doi.org/10.1162/003355300555060 https://doi.org/10.1016/j.learninstruc.2022.101650 https://doi.org/10.1037/dev0000159.supp https://doi.org/10.1016/s0022-4405(98)00006-5 https://doi.org/10.1086/675530 https://doi.org/10.1016/j.dr.2017.04.001 https://doi.org/10.1080/09243453.2010.512800 symonds et al 66 | f l r salmela-aro, k., tang, x., symonds, j., & upadyaya, k. (2021). student engagement in adolescence: a scoping review of longitudinal studies 2010–2020. journal of research on adolescence, 31(2), 256272. https://doi.org/https://doi.org/10.1111/jora.12619 schunk, d. h., & zimmerman, b. j. (eds.). (2023). self-regulation of learning and performance: issues and educational applications. taylor & francis. shen, t., & konstantopoulos, s. (2021). estimating causal effects of class size in secondary education: evidence from timss. research papers in education, 36(5), 507-541. https://doi.org/10.1080/02671522.2019.1697733 shevlin, m., & banks, j. (2021). inclusion at a crossroads: dismantling ireland’s system of special education. education sciences, 11(4), 161. https://www.mdpi.com/2227-7102/11/4/161 skalická, v., stenseng, f., & wichstrøm, l. (2015). reciprocal relations between student–teacher conflict, children’s social skills and externalizing behavior: a three-wave longitudinal study from preschool to third grade. international journal of behavioral development, 39(5), 413-425. https://doi.org/10.1177/0165025415584187 skinner, e. (2016). engagement and disaffection as central to processes of motivational resilience and development. in k. r. wentzel & d. b. miele (eds.), handbook of motivation at school (pp. 145-168). routledge. symonds, j. e., schreiber, j. b., & torsney, b. m. (2021). silver linings and storm clouds: divergent profiles of student momentary engagement emerge in response to the same task. journal of educational psychology, 113(6), 1192-1207. https://doi.org/10.1037/edu0000605 torsney, b. m., & symonds, j. e. (2019). the professional student program for educational resilience: enhancing momentary engagement in classwork. the journal of educational research, 112(6), 676692. https://doi.org/10.1080/00220671.2019.1687414 yang, p. j., & lamb, m. e. (2014). factors influencing classroom behavioral engagement during the first year at school. applied developmental science, 18(4), 189–200. https://doi.org/10.1080/10888691.2014.924710 https://doi.org/https:/doi.org/10.1111/jora.12619 https://doi.org/10.1080/02671522.2019.1697733 https://www.mdpi.com/2227-7102/11/4/161 https://doi.org/10.1177/0165025415584187 https://doi.org/10.1037/edu0000605 https://doi.org/10.1080/00220671.2019.1687414 https://doi.org/10.1080/10888691.2014.924710 codepen moeller et al frontline learning research vol.8 no. 3 special issue (2020) 63 84 issn 2295-3159 disentangling objective characteristics of learning situations from subjective perceptions thereof, using an experience sampling method design julia moeller a, jaana viljarantab, bärbel krackec & julia dietrichc auniversity of leipzig, germany buniversity of eastern finland, joensuu, finland cuniversity of jena, germany article received 25 june/ revised 19 october / accepted 21 october / available online 30 march abstract this article proposes a study design developed to disentangle the objective characteristics of a learning situation from individuals’ subjective perceptions of that situation. the term objective characteristics refers to the agreement across students, whereas subjective perceptions refers to inter-individual heterogeneity. we describe a novel strategy for assessing and disentangling objective situation characteristics and subjective perceptions thereof, propose methods for analysing the resulting data, and illustrate the procedure with an example of a first study using this design to examine situational interest in 155 university students. situational interest was assessed nine times per weekly lecture with three measurement time points per person and a rotated multi-group schedule. assessments took place over the course of an entire semester of ten weeks. one of the advantages of the proposed design is that objective group agreements can be disentangled from subjective deviations from the group’s average at each of the nine measurement time points per weekly lecture. furthermore, the proposed design makes it possible to study the development of both subjective and objective parameters across the time span of one weekly lecture and an entire semester, while the burden for each person is kept relatively low with three beeps per lecture. keywords: subjective self-reports, inter-rater agreement, experience sampling method, momentary motivation. info corresponding author email: julia.moeller@uni-leipzig.de doi: https://doi.org/10.14786/flr.v8i3.529 1. introduction imagine you attend a lecture that you love but that all your classmates seem to hate, dread, or find boring. you just love this lecture of statistics and research methods, because of its exciting implications for epistemology and its answers to the question where knowledge comes from and how much we can(not) trust in what we know. you think this is one of the best, most interesting courses you have ever had, but your fellow students just don’t seem to share your enthusiasm for the philosophy of science or mathematical representation of knowledge. while you express your love for this statistics and methods course, nearly everyone else would rather study “real psychology” or sleep in instead of starting the day with the 8:00 a.m. statistics course. one of your classmates even called you a nerd. while you try to convince everyone that this course is objectively interesting, the other students try to convince you that this is an objectively uninteresting lecture, claiming that “if we all agree it’s boring, it can’t be objectively interesting”. your best friend agrees with the others on that, but, trying to put himself in your shoes, also acknowledges that you have subjective reasons to find that lecture interesting, while also trying to convey to you that your subjective interest just isn’t everyone’s cup of tea. you discuss to what extent the agreement, or average interest, of the class reflects the objective interestingness of that course. the dean, in turn, holds the teaching evaluation in hand when announcing the decision to discontinue your favourite course in the future, citing the average lack of interest of the attendants as evidence for the objective lack of teaching quality, because all other psychology courses got higher student ratings in the questions asking about students’ enthusiasm. you feel unheard and unseen, after all, aren’t you a data point in that statistic the dean holds in hand, too? didn’t your favourite statistics teacher just yesterday teach you about the problem that sometimes individual students or subgroups hide behind the overall trend, so that we need methods to detect and describe these subgroups and deviating individuals? this article presents a novel approach to disentangle and describe both the overall trend in the agreement of a class on the ratings of a learning situation, and the deviations of individual, subjective, perceptions from that overall trend. the methods proposed in this article promise to be insightful for a broad audience, including researchers using the experience sampling method for classroom assessments, educational technology developers looking for methods to provide metrics and visuals concerning student heterogeneity and objective situation characteristics in teacher dashboards and class feedback systems, as well as educators who are interested in situational measures of students’ classroom perceptions, momentary assessments supporting personalised learning, or teacher feedback for social-emotional learning. while we use the example of interest ratings throughout the article, the methods proposed here could also be applied to assess other learning-related classroom perceptions, such as students’ observations of teacher behaviour, or students’ perceptions of the current task being difficult or easy, to name a few. by disentangling the idiosyncratic and commonly shared components of motivational self-reports, this article makes a contribution to this special issue’s first question (“in what ways do self-report instruments reflect the conceptualizations of the constructs suggested in theory related to motivation or strategy use?“). we also address the second question of this special issue by proposing analytics strategies, but rather than focusing on the constraints mentioned in the special issue editorial, we focus on novel avenues for analyses. 1.1 how to assess characteristics of learning situations situational self-report assessments of motivation and emotion are more and more frequently used, thanks to new technology that makes it easier and cheaper than ever to ask participants in real-time via mobile devices about their current activities, as well as their subjective perceptions, feelings, and motivations, pertaining to these currently ongoing activities. the methods used to gather such data are called experience sampling method (esm; e.g., hektner et al., 2007), ambulatory assessments (e.g., fahrenberg, 1996), or ecologically momentary assessments (shiffman et all., 2008). in this article, we use the term esm. compared to the classic retrospective one-time administered self-report questions for motivation and emotions, esm assessments have several advantages: a first advantage is that esm assessments can capture the fluctuating and situation-specific components of motivation and emotions, while common retrospective, one-time administered self-reports do not reveal which aspects of the assessed variables fluctuate or remain stable from one situation to another. it is even possible to disentangle situational determinants (e.g., the exciting learning video used in today’s lecture) from stable personal factors (e.g., this student’s well-developed personal interest in the topic taught today or this students’ general openness to experience), or contextual factors (e.g., the generally monotonous teaching style of this teacher, or the loud noise in that classroom from the construction side next door, which has hampered the students’ attention and motivation for a year now). to disentangle such situational, personal, and contextual influences, esm assessments can be combined with multilevel data analysis that decomposes the variance on the situational (within) level from the variance due to stable inter-individual differences (between level 1) and the variance between in contexts, such as class or school (between level 2; see e.g., dietrich et al., 2017; ketonen et al., 2018). a second advantage is that situational measures have been discussed to be more valid than the retrospective self-reports, because in-the-moment assessments can reduce memory errors (e.g., green et al., 2006; takarangi et al., 2006) and response biases linked to beliefs and stereotypes that are otherwise activated in certain retrospective self-reports (e.g., bieg et al., 2015; goetz et al., 2013). while retrospective measures require participants to mentally aggregate their typical experience across all the situations they can remember, esm data enable the researcher to empirically calculate such an aggregated typical experience as the mean score of the many repeated situational assessments for each person. these advantages of esm measures notwithstanding, they are still self-reports and therefore share many of the shortcomings related to self-reports with the retrospective measures. one of these shortcomings is the problem that esm and other self-report measures capture only the subjective perception and rating of an experience, which do not necessarily reflect how other students would perceive the same situation. for example, if a student indicates a current interest of “4 – very much” on a four-point scale, the reasons for that choice of this response option remain unclear. did this student choose this rating because he/she had a personal interest in this topic, while most other students were utterly bored? or because this was the most captivating topic ever taught, and every student in the class was captivated and would agree? or did the student just affirm being interested because of a very high individual level of trying to appear socially desirable? in classic esm studies, it is very difficult to disentangle these different options, because typically, students are asked about activities at random times, implying that every student has an own individual random survey schedule, so that there is usually no way to determine how other students perceived the exact same situation. to provide a solution for that problem, this study presents a research design that enables researchers to systematically assess groups of students at the same time points, so that inter-individual agreements and subjective deviations from these agreements can be distinguished. the previous literature provides some examples of theories that distinguish between the subjective and objective components of information provided by self-reports in education (e.g., göllner wagner et al., 2018; lüdtke et al., 2009). one such example is the research on interest, particularly the person-object theory of interest (fink, 1991; krapp, 2002; krapp & fink, 1992; prenzel et al., 1986), which distinguishes between objective characteristics of a learning situation and the individual’s subjective perceptions thereof, as well as distinguishing between fluctuating situational and stable personal determinants of the subjective perceptions. as the name person-object theory suggests, interest is expected to result from two main conditional factors, person characteristics and situation characteristics (krapp, 1998). according to this theory, interest emerges in the interaction of a person with particular objects (including concrete objects, such as texts, and more abstract ideas, events, topics, texts, etc.). objects are expected to differ in their likelihood of eliciting situational interest in individuals, depending on their verifiable, observable features. for instance, learning materials are more likely to trigger situational interest if their objective features make them surprising, novel, visually stimulating, and intense for the students. texts are likely to trigger situational interest if they are, for instance, easy to comprehend, cohesive, vivid, if they evoke emotional reactions, and allow for collaborations with others (for overviews, see e.g., krapp et al., 1992). the (objective) interestingness of a situation is then perceived by individuals who differ in their stable, dispositional personal interests (krapp et al., 1992) and their perceptions of the situation (e.g., a given information about the theory of relativity might be new to most students in a class, except for max, who has heard about the topic extensively over dinner from his mother, who is a quantum physics professor). the individuals then feel more or less interested in the current learning situation, depending on their previous dispositional interest in the currently discussed topic (person characteristic) and the learning situations’ objective characteristics. thus, a fluctuating psychological state of more or less situational interest can be observed, which is either an expression of the currently actualised dispositional interest, or the fluctuating reaction to the objectively interesting situation, or a mixture of both (e.g., krapp, 1998). the distinction between objective situation characteristics, student’s subjective perceptions of these objective situation characteristics, and objective person characteristics is depicted in figure 1, based on the person-object theory of interest visualized in krapp (1998). figure 1. a model of the person-object theory of interest based on krapp, (1998). we added the distinction between objective situation characteristics and subjective perceptions thereof. we expect that such models of person-object interactions will be fruitful not only for the understanding of interest, but also for the understanding of other motivational constructs, cognitions, and behaviours. for example, a student’s overall perception of a teacher’s behaviour in a given learning situation may be influenced by objective situation characteristics (e.g., the teacher’s in-fact behaviour), and by the student’s subjective perceptions of the teacher’s objective behaviour, and by stable or fluctuating student characteristics (e.g., the student’s subjective liking of that teacher or the student’s dispositional interest in the subject). likewise, a student’s learning in a given learning situation may be influenced by objective situation characteristics (e.g., the difficulty of the task, compared to tasks previously presented to the same student), by the student’s subjective perceptions of those objective situation characteristics (e.g., the student’s subjective appraisals, self-efficacy), as well as by stable person characteristics (e.g., the student’s intelligence, ability self-concept, perseverance in the face of obstacles). these are just a few of the possible applications of the person-object-logic presented in figure 1, suggesting that this model might be a useful framework for understanding and assessing learning processes in classrooms. please note that we use the term “stable” in the following to refer to aspects that do not change across measurement time points for the duration of an experience sampling method study (typically a few days), meaning aspects that are modelled in multilevel models on the person-level as stable inter-individual difference components. by using the term “stable” in that sense, we do not mean to imply that those empirically stable components cannot develop over longer periods of time, and we do not mean to rule out that personality development takes place. we merely imply that such eventual long-term development is usually not captured or distinguishable by the short-term longitudinal data that we discuss in this article. 1.2 the present research this article has three main goals: first, we introduce a novel design for esm studies that enables researchers to disentangle the objective characteristics of a situation (e.g., the situation’s interestingness, in terms of the inter-subjective agreement of all students) from each person’s subjective perception of that situation (e.g., the subjective interest and deviation from the aforementioned inter-subjective agreement). second, we propose different analytical strategies to analyse data assessed with this novel design and illustrate some of these proposed analyses using data from a first study that has used this novel assessment design to assess situational motivation in 155 university students. third, we discuss a number of possible additional strategies of contrasting the situational self-reports of esm assessments with more objective data, such as behavioural classroom observations or psychophysiological measures of emotion-related data. this is a theoretical article with the main goal to propose a new assessment design, and to discuss its advantages as well as limitations. therefore, our research questions refer to theoretical and methodological rationales, while the empirical results presented in this article merely serve as illustration and example for the methodological discussion, rather than being a centrepiece. thus, the topic examined in the empirical part (students’ situational interest in higher education) is treated as a rather exchangeable example of a construct, which to examine the here proposed method could help. 1.3 research questions rq1: what concepts of objectivity can be applied to disentangle objective characteristics of situations from participants’ subjective perceptions thereof? rq2: how do esm research designs and schedules have to look like in order to capture both the objective situation characteristics and the subjective perceptions thereof? rq3: what analyses are needed to disentangle the objective situation characteristics and the subjective perceptions thereof in data collected with a design proposed under rq2? the research questions are mostly answered theoretically, but with references to an empirical study that will serve as an illustrating example for the proposed methods. this example study is described in the following. 2. methods 2.1 sample and procedure the participants were 155 german university students (51% female; mean age m = 21.77 years, sd = 2.91; range: 19 to 46 years). the participants studied in a teacher education program with the aim to become subject teachers for secondary schools. students provided intensive longitudinal data in the form of esm surveys and were followed over one semester in a weekly lecture with 90-minute lessons (except for lesson 4, which ended after 60 minutes ahead of schedule). the subject of the course was ‘psychological fundamentals of learning’. in each of ten consecutive weeks, students received notifications and questionnaires at fixed schedules, three times during each lesson, consisting of ten situational motivation items. the participants chose whether to respond online with their own smartphone or on paper-and-pencil questionnaires (smartphone: 58–71% participants with a mean of 65% across the ten lessons; paper-and-pencil: 29–42% participants with a mean of 35%). n = 155 students provided valid information on situational measures in at least one lesson. during the data cleaning, we removed responses in the following cases: if the response was given more than 15 minutes after the signal (applies to the time-stamped online responses, not the paper-and-pencil response); if a person reported being present at the lecture but responded online after the lecture had ended; if a person responded to the three surveys immediately after another; and if a person responded with the same value on all ten items. this resulted in the omission of 251 surveys. a total of 2,226 completed esm surveys remained in the analysis sample, which equals 48.94% of the possible full data (three responses per lesson by ten lessons, except for week four, which ended early and therefore included two responses per person only) by 155 participants resulting in 4,495 responses). 2200 of those completed esm surveys had valid responses on at least one of the interest variables used in the analysis for this article and thus appear as the sample size in our mplus output (moeller et al., 2019). although paper-and-pencil surveys were not time-stamped, they were handed out before and collected after each lecture, so that responses on paper-and-pencil forms were only possible during the lecture. the theoretical framework for the data collection originally was eccles’ expectancy-value theory (eccles et al., 1983), according to which expectancies and values of a task are central motivational forces in students’ academic behaviours and learning (eccles & wigfield, 2002). they predict academic choices, persistence, and achievement (e.g., battle & wigfield, 2003; cole et al., 2008; durik et al., 2006). the dataset was also used in previous studies (dietrich et al., 2017; dietrich et al., 2019). these previous studies examined associations of situational expectancies and task values with effort (dietrich et al., 2017), and situational expectancy-value profiles (dietrich et al., 2019). none of these previous papers analysed the cross-classified data structure in this dataset. all data and r and mplus syntaxes for the calculations presented in this article are openly accessible at the open science framework (moeller et al., 2019). 2.2 measures the esm assessment captured situational task values and expectancies with eight items (see dietrich et al., 2017 or https://osf.io/qjkmz/). additionally, situational interest and situational effort were assessed with one item each. the students were instructed to think about the lecture contents of the past couple of minutes and to complete the questionnaire within nine minutes. they were asked “to what extent do the following statements apply to you in the present moment?” and responded on a 4-point likert scale ranging from 1 = does not apply to 4 = fully applies. in the present article, we constructed an averaged composite score, labelled situational interest from two items measuring situational interest (“i am interested in these contents”) and situational intrinsic value (“i like these contents”). while most esm studies assess constructs with single items to keep the burden on the participants as low as possible, we opted to assess each construct with multiple items (see dietrich et al., 2017 or https://osf.io/qjkmz/), based on the idea that the shared variance of multiple indicators is a more reliable indicator of an underlying construct than a single item can be. it could be argued that this approach of using composite scores instead of single items in itself is a contribution to making esm assessments more objective, as it reduces the risk of confounding a construct of interest with the random and unique influences (unique variance) that a single item captures apart from the construct it is supposed to represent (e.g., a momentary slip in attention causing the respondent to click on a wrong response option, or an idiosyncratic misunderstanding by a given participant of a given item in a given situation). in a cross-classified multilevel model with responses (within level, n = 2,200) nested in both individuals (between-individual level, n = 155) and time points (between-time point level, n = 87), the correlations between the two items of the situational interest scale were r = .62 at the within level, r = .93 at the between-individual level, and r = .99 at the between-time point level. 3. results and discussion 3.1 which concepts of objectivity should be applied to disentangle objective characteristics of situations from participants’ subjective perceptions thereof? (rq1) when using the term objectivity, we assume that there are true characteristics of an object (in this study, a learning situation) that influence individuals’ subjective responses in a somewhat systematic way that causes at least partial agreements in the subjective responses of multiple individuals. skipping the important and millennia-long philosophical discussions about whether or not there is a truth and how to define and understand it (for a summary, see e.g., glanzberg, 2018), we pragmatically define objectivity here as the (approximate) agreement of all observers about the characteristics of an object, with the object here meaning the learning situation. this here applied concept of objectivity is based on popper’s claim that “the objectivity of scientific statements lies in the fact that they can be inter-subjectively tested” (1934 [2002], p. 22). while popper refers to scientific statements rather than laymen’s implicit concepts, his idea of objectivity as inter-subjective agreement has been extended: douglas (2011) has named this concept of objectivity the concordant objectivity and defines it as the “simple agreement among multiple observers” (douglas, 2011, p. 32). this idea of objectivity as the inter-subjective agreement about the truth of an object is reflected in the classical test theory, which typically considers assessments to be objective to the degree that all trained assessors come to the same conclusion about an assessed construct in a given individual or population. it is also reflected in the practice of treating a classes’ mean score of averaged individual student responses about perceived teaching quality as a proxy for the objective teaching quality in that classroom (see e.g., göllner, 2018; lüdtke et al., 2009). it is important to keep in mind that this is a rather parsimonious concept of objectivity and that many more have been discussed in the social sciences, including education (e.g., eisner, 1992; fisher, 2000). how can self-reports ever be objective in the sense of concordant objectivity, given that the construct to be assessed (e.g., an emotion) is typically experienced only by the individual who experiences it, which also is the reason why we ask participants to report us their feelings in self-reports? while it is controversially discussed to what degree it is possible to objectively determine how exactly a person feels without relying on this person’s subjective self-report (e.g., barrett, 2018), it is arguably possible and useful to assess the characteristics of a situation that are related to the emotional experiences. for example, since previous research on situational interest suggests that it is largely triggered by objective situation characteristics, such as novel and surprising information being presented, we can expect that multiple individuals agree partially in their situational interest in one learning situation, as long as the information presented to them is equally novel and surprising to each of these individuals. one possibility of deriving information about objective characteristics of a situation from subjective self-reports is to aggregate the multiple self-reports of different individuals about the same situation. in that sense, the objective assessment of, e.g., a situation’s interestingness, would be the agreement of a large enough number of randomly selected individuals about the interest they felt in that situation. in this scenario, we would expect that the group of students tends to report higher situational interest in learning situations presenting students with novel and surprising stimuli, compared to learning situations presenting students with known, unsurprising stimuli, as long as the stimuli in both cases are otherwise similar. the group’s average (inter-subjective agreement) of reported interest in a given situation would thus be an indicator of the objective characteristics (interestingness) of that situation. a limitation of the here-applied definition of objectivity is that the group’s mean score in a situation may be sample specific. if we select only the most interested individuals, then their mean interest in a given situation may be high, not because of the situation being novel and surprising, but because of the generally high interest of the group across all situations. thus, the group mean score in a given situation may reflect the possible interactions between the group’s person characteristics and situation characteristics, rather than indicating the objective situation characteristics alone. 3.2 how does an esm assessment have to look like for it to capture both the objective situation characteristics and the subjective perceptions thereof? (rq2) with the aforementioned definition of objective assessments of situational characteristics through inter-subjective agreements across individual subjective self-reports, we need momentary assessments from multiple individuals in the same situation to aggregate these multiple self-reports to a situation-specific group mean score. what exactly the term same situation means depends on the research question of a given study. for instance, the objective interestingness of a learning situation can be assessed by asking all the students in the same class in the same instant about their current interest in that moment, and then aggregating across these individual responses. if, however, students learn remotely and self-paced with digital learning platforms, then a situation in terms of the research question could either be a certain time point (e.g., tuesdays afternoon, or 24 hours before the final exam), or it could be the individual time point at which each student finishes a given task that is relevant to the research question, or another condition with relevance to the understanding of a digital learning moment. in this article, we use the term situation synonymously with a given time point in a given lecture hall in which all students see and hear the same university lecturer talking in the front of the room and are asked at the same time point about their momentary motivation, as described in dietrich et al. (2017). in order to assess multiple individuals in the same situation, we need to modify the common design of individually randomly timed survey notifications used in many esm studies. for our purpose, we need to assess a large enough group of students at the same time, whereas the common esm schedules typically assesses each student at their own individual random times in order to capture a true random sample of all the experiences that students make during a relevant unit of time, like a school day (e.g., hektner et al., 2007). if we deviate from such truly random and individual schedules in order to examine students’ inter-subjective agreements in one given situation, then the so-collected data may not be a representative sample of everyday life activities. however, there are many research questions that do not require a representative sample of all everyday life activities. for example, research questions referring to specific school subjects, or specific teachers, or specific lectures, require these contexts to be oversampled to ensure a large enough sample of situational assessments in that chosen context. one challenge in the use of not randomly timed esm surveys is the risk of systematic context-specific biases in the assessments: the smaller the range of assessed time points or situations, the larger the likelihood of non-random influences. for example, in a truly randomly timed esm study, we would not expect the results to be influenced by the time of the day, or the students’ distractedness during the last minutes of a class when everyone already packs their things to jump up at the first ring of the school bell, or other influences that are particular to a certain timing, because these influences are expected to cancel out. in contrast, if we decide to assess students only in the last five minutes in each class, for instance because teachers are concerned about interruptions and a no-phone policy during lessons, then we cannot rule out that the timing might have biased the responses in a way that a random survey schedule would not have. in order to reduce the risk of contextual biases in the assessments of multiple participants in the same instants, we need to make sure that at least there are no biases concerning the timing of surveys. that means that whatever the time span relevant to the study, no participants should be surveyed only at the beginning or only at the end of that time span. instead, surveys should be distributed equally across these time spans for all participants. for that purpose, we have developed an assessment design described below for the study of momentary study motivation in a university course across an entire semester (see also figures 2 and 3). the weekly 90-minute lessons of that course are split into nine periods of nine minutes each (not ten minutes, because the participants need at least one minute to answer to the last survey in the lecture and would miss that last notification if it occurred when the end-of-lecture-noise and hectic has already started). to make sure that we have data detailing the motivation across the entire lecture, we assess participants after the first ten minutes, after 19 minutes, after 28 minutes, and so on. we start the first assessment after ten instead of nine minutes, because in the first few minutes, some time tends to get lost on welcoming and waiting for students to calm down. to keep the burden on each individual participant low, each participant is only surveyed three times during the lecture, with a time gap of 27 minutes between each assessment. to assess multiple participants at the same time while still pursuing the aforementioned goals (data across the entire lecture, no participant surveyed more than three times), participants are surveyed in groups, with group a being surveyed after the first 10 minutes, group b being surveyed 19 minutes into the lecture, group c being surveyed 28 minutes into the lecture, and then group a again being surveyed 37 minutes into the lecture, and so on. the same design is then repeated one week later in the same lecture, but with the difference that group b starts the assessments 10 minutes into the lecture, in order to rotate the survey times across all groups, times, and weeks (figure 2). individuals were randomly assigned to groups in a way that ensured a relatively equal sample size of each group. figure 2. example signalling schedule in lesson 1, 2, and 3 (to be rotated in following lessons) 3.3 which analyses are needed to disentangle the objective situation characteristics and the subjective perceptions thereof in data collected with the proposed design? (rq3) 3.3.1 analytical strategy 1: visualising the inter-personal agreement (objective parameter) and the subjective deviation from that agreement (subjective parameter): jittered violin plots to start the analyses of the data gathered with the assessment design proposed above, it is recommendable to get an overview of the distribution of the responses at each measurement time point. to explore how much participants agree or individually deviate from the average rating of the interestingness of a learning situation, we suggest examining the inter-individual distribution of interest ratings for every beep in a given lesson with a jittered violin plot (using the r package ggplot2 with the jitter option). figure 3 shows an example of such a plot. figure 3. inter-individual distributions of situational interest for each beep across the ten lessons, visualised in a jittered violin plot (note that lesson 4 ended after 60 minutes, which is why three violins are missing for the last three measurement time points) the jittered violin plot shows the inter-individual mean score (red dot), the standard deviation (distance between the red dot and one end of the red line), and the inter-individual distribution of interest ratings (black dots) for each measurement time point / beep in each lesson. the belly of the violin is proportional to the number of individuals who agreed on a rating, with a thick belly meaning that many participants chose this value in their rating of their current interest. since many observations (dots) would have overlapped on the interest values of 2, 3, and 4, we used the option to jitter the dots, which “adds a small amount of random variation on the location of each point” (wickham et al., n.d.; wickham, 2016) to prevent overlap and to display the number of the individual responses for each value. the visual inspection of the violin plots indicates whether the assessed group of students tends to agree, as indicated by violin plots with one clearly distinguished belly (e.g., figure 3, lesson 1, 10 minutes), or if the responses are randomly distributed across the possible range of values without any particular group agreement, indicated by a flat plot without any belly (e.g., figure 3, lesson 3, 46 minutes), or if there are two or more distinct groups, each of which agree on their own particular score, indicated by a violin plot with multiple bellies (e.g., figure 3, lesson 4, 55 minutes). the latter case of a mixed distribution can additionally be examined with statistical tests that tell whether the distribution is unimodal, bior multimodal, such as hartigan's dip test statistic for unimodality versus multimodality (hartigan & hartigan, 1985). these tests can be performed with the r package diptest (maechler, 2016) or the bootstrapping procedure determining the number of modes described in efron & tibshirani (1993), which uses the r package bootstrap (tibshirani & leisch, 2019). it should be noted that the concept of concordant objectivity applied in this article requires the existence of a unimodal distribution, meaning a high concentration of responses closely around the mean score. if the distribution is bior multimodal, then there is no reason to assume that the responses reflect one inter-subjective objectivity, and no reason to claim that the inter-individual mean score was an indicator of such concordant objectivity. whether or not responses reflect inter-subjectively objectifiable information likely depends on the construct and needs to be determined with the above-mentioned strategies (number of bellies in violin plots and tests for universus multimodality). as figure 3 shows, the distribution and universus multimodality can also differ from moment to moment, not only from construct to construct. furthermore, göllner et al. (2018) have also pointed out that the individual’s deviation from the group’s mean score does not have to be due to individual rater tendencies or biases, but may represent meaningful information about dyadic experiences. using the example of teaching quality, the authors argue that different students can make different experiences with the same teacher, so that their deviation from the group mean score may reflect such observable differences in different dyadic experiences. this suggests that not only the group average can serve as an indicator of objective situation characteristics, it is furthermore possible that individual students make experiences that other students don’t make, but that other students still would rate the same way if they experienced the same. 3.3.2 analytical strategy 2: parameters for the deviation of a subjective rating from the objective group rating: cross-classified multilevel analyses after getting a visual impression of the degree of inter-rater agreement on the interestingness of a situation, we might want to calculate parameters quantifying the degree of inter-rater agreement and the degree to which each person at each measurement time point deviates from the group average. in particular, we want to get estimates for (1) the individual and time point-specific deviation from (2) the stable person-specific mean over time, and from the (3) the inter-individual group mean which may differ from situation to situation (i.e. from time point to time point). while component three represents the objective situation characteristic in terms of the characterization the participants can agree on (the average rating of that situation), component one represents the subjective situation-specific deviation from that objective rating, i.e., the subjective element. we would like to control both components 1 and 3 for component two, which represents the stable individual deviation from both the objective situation characteristic and the momentary subjective component due to stable response tendencies of that person (e.g., traits). such variance decomposition can be done with a cross-classified multilevel analysis (e.g., beretvas, 2010). this type of statistical model separates the total variance of the scores yit (the scores of the different individuals i at the different time points t) in the three aforementioned components: yit = y1i,t (component one, within time point and within individual) + y2i (component two, between individuals) + y3t (component three, between time points) the cross-classified multilevel model has the advantage that it allows disentangling the personand situation-specific deviation (subjective state component) from the group average (objective state component), while accounting for each person’s stable tendency to deviate systematically from other individuals across all measurement time points (trait component). instead of the more common structure of esm data, with situations nested only in individuals due to the randomness of the time points, the here proposed assessment design results in time points nested in both individuals (y2i), and the groups a, b, and c with their respective measurement times (y3t). time points are crossed with individuals, because each individual appears only once within each measurement time point. to get reliable estimates about the variance components y1i,t, y2i, and y3t, sufficiently big samples of individuals and time points are needed (around n = 50 on each of these levels; chung et al., 2018). the present study design comprises of n = 155 individuals and n = 87 time points. we computed the above-described model to separate the total variance [var(yit) = .403] into the three variance components. the biggest portion of the variance pertained to the subjective deviation from both the situational group average and the stable person-specific trait component. this individual, situation-specific component one showed a variance of var(y1i,t) = .250, which equals 62% of the total variance. second, stable inter-individual differences (traits; component two) accounted for 31% of the variance: var(y2i) = .124, which means that about one third (31%) of the variance is due to stable person-specific response tendencies that differ between individuals. finally, the variance of the objective component three was considerably smaller: var(y3t) = .029, 7% of the total variance. that means in other words that only a small amount of variance was due to changes in the situation-specific group mean score from one moment to another. 3.4 summary this study suggested a novel research design that allows to disentangle the objective characteristics of a situation from participants’ idiosyncratic momentary subjective deviations from those objective situation characteristics. for example, this novel approach allows disentangling the objective interestingness of a situation from a participant’s subjective interest in that moment. the key of this assessment design is the simultaneous assessment of multiple participants in the same situation / time point, which enables researchers to examine to what degree individuals agree on their ratings of a given construct in that given situation, and to what degree individual participants deviate from that group agreement. we proposed several methods to analyse data assessed with this design and to further examine the role of objective versus subjective components, including jittered violin plots displaying the distribution, means and standard deviations of each measure for each measurement time point, and cross-classified multilevel analysis. 3.5 practical implications a ground-breaking advantage of the here proposed assessment design is the fact that it makes feedback to teachers about the objective situation characteristics possible. the group agreement indicating the objective interestingness of a situation for example enables teachers to compare their teaching topics and strategies in terms of how they make their class feel. in common momentary assessments in classes, students are typically asked at random time points, implying that for each assessed situation, there is typically only one answer for one individual student. imagine you were a teacher wanting to know how your new teaching strategy came across to the students, and the researcher tells you: “see, at this time point, ten minutes into your lecture, you introduced the theory of evolution, and mary reported high boredom and low interest”. would you, as the teacher, conclude that the new teaching strategy failed to raise the students’ interests, or would you rather hypothesise about this one student’s idiosyncratic reasons for not being interested, or would you remain clueless as to how to interpret this feedback? with the approach of asking multiple students at the same time suggested in this article, it now becomes possible to tell teachers: “see, at this time point, ten minutes into your lecture, you introduced the theory of evolution, and the average interest reported by your students was high, even though a single student, mary reported high boredom and low interest”. this feedback enables teachers to evaluate the average and the individual perception of their teaching strategies by their students, which we hope will become an important tool for immediate feedback in learning settings. imagine for instance that the feedback occurs in real time and the teacher learns that most students are interested but two students are utterly bored. in that case the teacher could offer optional challenging bonus tasks for the two students who might be underwhelmed by the regular classwork. if the entire class is bored, then the teacher could use activating, engaging teaching strategies by trying to cheer up the class with a joke, increasing the task difficulty for everyone, or adding real-life examples allowing students to see the links between the discussed topic and their own interests. as another option, teachers could use the feedback to analyse and revise their teaching strategies and materials after the course or school year has ended. in our studies, we combined the momentary assessments with videos of the lecture, showing both lecturer and slides, allowing us to link the teaching behaviour and materials to the students’ momentary motivation. these videos, which are yet to be analysed, are meant to help the teacher (and us researchers) understand which behaviours are most, or least, motivating, and which slides should be modified to foster future students’ motivation. obviously, collaborations between researchers, and/or software developers and teachers are needed to realise this possibility, unless the researcher and the teacher is the same person (as in the here presented study on motivation in university lectures). the here presented methodological groundwork needed is only the first step in that direction. a next step would require collaborations in which researchers use these methods to identify students’ individual needs as well as momentary classroom levels of motivation and emotions. systematic collaborations of researchers and teachers are needed to provide teachers with the suitable emotion and motivation measures and assessments, and to provide researchers with the real-time data out of real school classrooms. technology experts are needed for the further development of feasible feedback systems that show the collected data in comprehensible form and real time to students, teachers, and – in the case of underage students – parents. science communication and more research are needed to find out which form of feedback about the assessed motivation and emotions would be most helpful to students, and teachers. 3.6 theoretical implications the combination of the approaches proposed in this article has much potential for the research on motivational heterogeneity of students. the intensive longitudinal data allow for the intra-individual examination of short-term developments (from one measurement time point in a given lesson to the next) and intra-individual long-term development of motivation or emotions (from lesson one to lesson ten). the jittered violin plots can be used to identify particular students as much as they can be used to detect overall trends, like an increase or decrease in the average inter-individual interest from one moment in the lecture to the next. while common experience sampling method approaches provided no information about a students’ deviation from the simultaneously present peer group, the here proposed approach can be used to identify, within any given learning situation, those students who score substantially below the benchmark of interest typical for the simultaneously present peer group. this information potentially makes assessments of learning-related emotions and motivation at the same time more person-specific and more situation-specific. instead of classifying students as generally less interested than their peers (which a classic esm approach can do by examining the person-level mean score), our approach enables researchers or educators to say: “although mary has a tendency of being less interested than her peers in math lessons, you really caught her interest with your most recent novel teaching strategy, which brought mary’s interest even above the level of her peers, as you can see in the last two violin plots (where mary can be marked as a yellow star among the black dots representing her peers)”. thus, we expect that the approach proposed here will make a contribution to personalised learning (e.g., corno, 2008) and tailored interventions for individual students at the exact times when they are in need of motivational and emotional support. common experience sampling approaches seem less useful for these purposes, because they leave open whether a given measurement score reflects the individual’s subjective interest or the situations’ objective interestingness, or, if both, which of those components to what degree. by offering techniques to disentangle the idiosyncratic and commonly shared components of motivational self-reports, this article contributed to this special issue’s first question (“in what ways do self-report instruments reflect the conceptualizations of the constructs suggested in theory related to motivation or strategy use?“) and second question (“how does the use of self-report constrain the analytical choices made with that self-report data?”). in sum, our answers to these questions are that self-reports only capture a person’s perception but can be aggregated to draw conclusions about the perceptions of a group of persons, their agreements and disagreements, about the characteristics of the (learning) situations they perceive. this article’s focus on self-reports of interest in learning settings complements several other articles in this special issue (chauliac et al., & donche; fryer et al.; durik & jenkins). 3.7 limitations of the rotated survey design proposed in this article one limitation of the design suggested in this article is the fact that the results inherit the problems linked to self-report data, including the fact that self-reports are always to a certain degree idiosyncratic, even when they are averaged or when group agreements are disentangled as a separate source of variance. this implies for example that the group mean score, which we described as the indicator of the objective situation characteristics (e.g., the objective interestingness of a situation) can be sample specific. imagine if we selected only the most interested individuals for some reason, then their group mean score (objective component) in a given situation will be high, not because the situation is objectively highly interesting but because we only asked the highly interested individuals. therefore, the objective component of the situational assessments is only objective to the degree to which it can be generalised from the observed sample to a larger population, which is a question for systematic replication studies to examine. if all or many participants in a sample are influenced by similar biases (for instance because we are surveying a group with particularly high social desirability), their agreement (the group average in a given situation) will reflect this joint bias rather than an objective situation characteristic. these are limitations to our definition of objectivity that need to be kept in mind. another limitation is the possible diversity of different activities that students who are present in one classroom might be engaged in. for example, in a class taught with a personalised learning approach, different students might be working on different tasks with different instructions. in some personalised learning settings, students in one classroom wear hearing protection to concentrate and receive their individual tasks on technological devices (e.g., tablets) contingent on their prior tasks completed, achievements, and goals. in such settings, it seems unreasonable to assume that the agreement of all raters on, e.g., their current interest, would reflect the objective interestingness of the learning moment, since it is likely that different students were thinking about different tasks when answering. a third limitation is the requirement of large classes for the design proposed in this article. in order to interpret the distribution and degree of agreement of different raters at any given time point, a reasonably large group is needed. the design proposed in this article was developed for large university lectures, which often involve 200 students or more. cross-classified models require least 50 students and 50 measurement time points in total (chung et al., 2018). per student, at least 10 measurement time points and per measurement time point 10 students responses are needed. however, 10 responses still seem too small of a sample from the standpoint of sampling theory and power considerations, for instance because of the biases that are more likely to affect small samples, compared to larger ones (e.g., creswell & guetterman, 2019; schönbrodt & perugini, 2013). the purpose and planned analysis should drive the sample size planning, because different approaches require different sample sizes. in most school classes, it might be less useful to split the class into three groups of responders with different esm signalling schedules, since many school classes comprise less than 30 students, implying that with the design proposed here, each subgroup at any given time would include no more than ten responses, likely less if school absences, smaller class size, and unwillingness to respond to esm signals are taken into account. a possible solution in reasonably large classes might be to signal all students at the same times (e.g., after 25, 50, and 75 minutes of a 90-minute lecture). as a rule of thumb, we recommend to assess all students at the same time if less than 30 students are present, to avoid biases of the group mean score due to outliers. this reduces the number of measurement time points and consequently offers less insight over short-time changes in students’ experiences over the course of a lesson, while keeping the burden on each individual student the same (three signals per lesson). if this suggestion is implemented and all students in a school class are surveyed at the same time, then we suggest that the teacher could interrupt the lesson for the duration of the survey to allow students to concentrate on the survey and to avoid that students miss any relevant learning information. to further avoid sample biases, particularly in smaller samples, it might be worthwhile considering to assign matched participants to different groups, so that individuals with similar person characteristics can be found in and compared across all groups. however, it seems unlikely that the needed variety and combinations of person characteristics needed for such matching procedures can be found in small samples such as school classes. it is furthermore not possible to rule out that nonresponse to esm signals might be confounded with the constructs being assessed. for example, the most motivated, immersed students might prefer to continue working on their captivating math task and might even miss the esm signal due to their intense concentration. in some personalised learning classrooms, the headsets that students wear to avoid distractions by their classmates make it difficult to raise their attention to esm signals, unless these signals come through the same devices their headsets are attached to, which is not always possible. on the other hand, particularly bored and disengaged students might see no point in responding to the esm surveys. these scenarios of data not missing at random imply that the empirically observed agreement of different students does not necessarily reflect the agreement of all students or the objective situation characteristics, but could itself reflect a biased subsample. it seems possible that constructs and situation characteristics differ in their potential of being perceived in the same way by different students. it might be easier for students to agree on a question referring to the interestingness of a situation, since situational interest partially depends on observable situation characteristics, such as novelty or surprising information (hidi & renninger, 2006), while other constructs may be more person-specific and difficult to observe, such as questions concerning the students’ current feeling of competence or frustration. in part, the here proposed design helps detecting and studying such differences between constructs by quantifying the degree to which students agreed in their agreements on different constructs. nevertheless, it is important to bear in mind that a lack of agreement can have many different sources, either rooted in the construct itself being person-specific, or rooted in individual distractions, individual misunderstandings of items, or individualised instructions. a general limitation of assessments during ongoing lessons is the risk that the interruptions through surveys, however short, may interfere with the students’ attention and learning. this affects all in-the-moment self-reports during classes and consequently most experience sampling method studies conducted in school or university. the here proposed method offers a way out: if it is applied in school, where classes are typically smaller than in university lectures, then all students can be surveyed at the same time and the teacher can stop teaching for the time being. this would still imply an interruption and potential loss of attention, but one that the teacher could afterwards address and try to mitigate, for instance by repeating core messages. in university lectures however, where we suggested to survey different groups of students at different times, it cannot be ruled out that some students might miss an important detail while completing the survey. teachers who are informed about the survey schedule may want to provide the information presented during survey times afterwards in a format that the student can read and repeat after the lecture, to catch up with any potentially missed information. future studies should examine whether brief interruptions by esm surveys interfere with students’ learning in school or university. if surveys and teaching occur simultaneously, it cannot be ruled out that the need to split one’s attention may impair the accuracy / validity of the students’ situational self-report or lead to a selective missing data pattern if students decide not to answer in those learning situations they find most difficult and attention demanding. there is no guarantee that there be only one group agreement on a given question in a given situation. there might be multiple subgroups, each with their own mean score / agreed-upon rating. while this would be easy to detect in the violin plot, it poses a limitation to the idea of using the group mean score as the one and only indicator of objective situation characteristics. if researchers were interested in comparing the groups a, b, and c with each other or make sure that they are comparable, then it would be recommendable to modify the assessment schedule proposed in figure 2 in a way that assesses two groups simultaneously. with the here-proposed schedule, it would be possible to determine whether group a reported generally higher interest than group b or c, across all situations, by nesting situational assessments (level 1) in individuals (level 2) in groups (level 3), while ignoring the clustering in measurement time points. this procedure would reveal how much variance is due to differences between groups. alternatively, a multi-group comparison with parameters constraint to be equal (e.g., asparouhov & muthén, 2012) could be used to test the assumption that groups were comparable in their mean scores, variances, or other, co-variance-based parameters. however, the procedures suggested in this article do not allow to disentangle the group-specific influence from the situation-specific influence in a given measurement time point, meaning if group a scores particularly high in the 37th minute of the class (see figure 2), we do not know for sure if group b would have scored the same in a similar situation. if this information is needed for a research questions or application, we recommend to use planned missing data designs that systematically assess multiple (at least 2) groups at a time, in order to be able to compare them (see enders, 2010). please note that assessing multiple groups at a time either increases the burden and interruptions for participants, if the schedule is kept the same and the number of surveys is increased for individuals, or it implies fewer measurement time points across the lesson, if the number of individual surveys is kept constant. finally, it cannot be ruled out that surveying students’ in their learning situation changes the very process we aim to study (e.g., schmitz & perels, 2011). this should be kept in mind in all experience sampling method studies surveying students in class, as well in studies using introspective self-reports in general. 3.8 directions for future method development we mentioned above that the concept of concordant objectivity employed in this article implies that students assessed in a given situation agree, which in turn implies that their responses should form an uni-modal distribution. however, it is possible that students agree while forming heterogeneous subgroups, leading to a mixture, e.g., bi-modal distribution. for example, the 155 students in our lecture might have consisted of two groups, the ones loving the teacher, and the ones hating the teacher, which might have lead to the bi-modal responses observed in some situations. it is also possible, and apparently was the case in this study, that the form of the distribution of responses varies from moment to moment. it would, for instance, be possible that students form two separate groups in assessing a political statement, with one group of, e.g., conservative students rating the joke as funny and appropriate, and another group of, e.g., liberal, students rating the joke as not funny and inappropriate, or vice versa. such an instance might cause a temporary bimodal distribution, while all other moments in the same lecture might see a uni-modal distribution as long as no politically connoted jokes are made. in moments in which the distribution is multi-modal, then it would be interesting to find out what caused the distribution. understanding the reasons and mechanisms behind heterogeneity in responses in given situations is yet to be examined more systematically in future studies. in addition, the variance of the scores at each time point can be small or large, independent of the form of the distribution. for example, even in a study in which all distributions of scores at all measurement time points were uni-modal, the range of scores and the overall variance of scores could be large or small, and could differ from moment to moment. figure 3 illustrates the size and change in variance between measurement time points in form of the red lines, which represent the standard deviation. importantly, a mixed distribution (multimodal distribution) suggests multiple groups hiding behind an overall trend, which is highly relevant for personalised learning and person-oriented methods. thus, apart of quantification of the variance, additional analyses, such as examinations of distributions and cluster/latent profile analyses could complement the search for reasons and mechanisms behind heterogeneity in responses in given situations. in this study we found that only 7% of the variance was due to changes in the situation-specific group mean score from one moment to another. while this might seem to suggest that it might not matter so much how a university teacher teaches, we would like to offer alternative interpretations and directions for future research: on one hand, we do not know whether seminars or practical courses at university, which allow for more diverse, hands-on learning experiences than lectures, might have differed more strongly in their average motivation from one moment to another. our findings only suggest that the lecture examined in this study was relatively consistent in the average interest it elicited from one moment to another (which oscillated around 3 on a scale from 1 = does not apply to 4 = fully applies). future studies could examine whether the diversity and distinctiveness of learning tasks in a university course can increase the variance due to differences between changes in the situation-specific group mean score from one moment to another. the fact that the largest proportion of variance (65%) was due to the individual, situation-specific component is a strong argument for personalised learning tools and other instruments that help teachers address the motivational heterogeneity they encounter in their university courses and classrooms. individual students’ momentary motivation deviated much from the average motivation in a given moment in this lecture, and figure 1 shows that in most moments, there were very interested as well as rather disinterested students present. while heterogeneity and individualised/personalised learning have been addressed increasingly in the literature on learning and instruction in schools (e.g., banister et al., 2014; bingham et al.,2018), there is still a need to implement personalised learning procedures in university teaching. it should be noted that in this article, we do not make use of the longitudinal nature of the esm data, because that was beyond the article’s main scope, which focused on the distinction between subjective and objective components in self-reports. nevertheless, the here-proposed method also has interesting implications for the longitudinal study of learning and teaching processes. our method allows to study the following longitudinal questions: how does the construct of choice (here: situational interest) change within a lesson, within a person, over 30 minutes? how does the construct of choice change in one session of a lecture, across individuals, over 9 minutes? how does the construct of choice change from one week to another, over the course of a semester, on average across individuals or within individuals? reitzle and dietrich (2019) give an overview of possible longitudinal models that can be used to examine such questions, using the data described in this article and providing corresponding r and mplus scripts. while the methods proposed in this article attempt to contribute to further developments of personalised learning and interventions based on momentary assessments, it is important to keep in mind that much more research is needed to get from assessments to valid and helpful interventions. as bastiaansen et al. (2019) have shown, different teams of researchers can draw very different conclusions about needed interventions from the exact same intensive longitudinal dataset and its intra-individual analyses. teams of software developers, methodologists, and educators will need to work together to identify valid and effective ways to draw conclusions about individual students’ emotional needs for support from data like ours, and about the best ways to deliver appropriate interventions in the appropriate moments. 3.9 future directions to overcoming the general limitations of self-reports because of the general limitations of self-reports, it is important to validate self-report data gained with the here proposed research design by linking them to more objective, observable and behavioural data, such as video-recorded observations of the students’ behaviour or the teacher’s behaviour, psychophysiological data with relevance to emotions and motivation, such as mobile electrodermal resistance assessments or heart rate variability measures, verifiable information about students’ performance (ideally standardised test performance), absenteeism, school dropout and objective information about the students’ demographic background, such as their family’s household income. if only self-report assessments are possible due to organisational or other constraints, then different question formats can help to avoid at least the biases typical to rating scales: for example, emotions could be assessed with both open-ended questions (“please write down here how you currently feel”), which can be linked to rating scales after being automatically analysed with sentiment analyses tools (e.g., silge & robinson, 2017) or manual coding (e.g., moeller et al., 2018). researchers and practitioners around the globe work on methods to gather objective information about participants’ emotions and motivations. for example, there are studies and companies that retrieve information about people’s emotions from their voices (e.g., krothapalli & koolagudi, 2013), countless companies and data scientists analyse texts produced by participants for markers of emotions in so-called sentiment analyses (e.g., altrabsheh et al., 2013), wearable heart rate variability sensors are marketed to researchers and private users with the promise that they will provide objective information about the stress, sleep, recovery, and physical exercise of the wearer (e.g., firstbeat, 2012). multiple sensors are integrated to optimise predictions of behaviour and emotions, and machine learning algorithms help integrate all these data, reaching never before seem accuracies in predicting emotions and behaviour (e.g., carroll et al., 2013). on the other hand, a recently emerging debate has questioned whether the subjective information about personal experiences provided by self-reports can be entirely replaced by objective measures, since e.g., barrett (2018) has suggested that even the presumably objective markers of emotions are to some extent idiosyncratic. for these reasons, we might have to keep asking people for their self-reports if we really want to know how an individual feels in a given situation, since the subjective evaluation, a crucial part of the emotional experience, is not always captured in the observable and behavioural measures. for these reasons, we believe that the research design proposed here will remain a useful tool to examine in the future to what degree a given esm response in a given situation was idiosyncratic and thus a reflection of person-specific characteristics, or in line with the assessments of other students in the same situation. we hope that the indicators of concordant objectivity proposed here can be compared and integrated with other objective measures of emotions and motivation in learning situations in the future, in order to improve predictions of students’ learning and behaviour. there is a large array of constructs that could be assessed in line with the here-proposed person-object logic (figure 1) and schedule for disentangling the subjective and concordant-objective aspects of participants’ situational self-reports. apart from the example of interest discussed throughout this article, the method promises to be insightful also for constructs such as perceived teacher behaviour, students’ rating of teaching quality (see e.g., göllner et al., 2018), or perceived situation or classroom characteristics (e.g., task difficulty, social climate, see e.g., lüdtke et al., 2009). keypoints this methodological contribution proposes a new assessment design for experience sampling method data collections that enables researchers to disentangle objective person characteristics from subjective perceptions thereof. the proposed design makes it possible to study the development of both subjective and objective parameters across the time span of one weekly lecture and an entire semester, while the burden for each person is kept relatively low with three beeps per lecture. different options for corresponding analyses are proposed, including jittered violin plots for visual inspection, tests for universus multi-modality, and cross-classified multilevel models. we discuss implications of the proposed research design for the development of teacher feedback for measures of momentary student emotion and motivation. acknowledgments this research has been supported by a jacobs foundation early career research fellowship awarded to the first author. we thank our reviewers for very thoughtful and appreciative feedback. references altrabsheh, n., gaber, m. m., & cocea, m. (2013). sa-e: sentiment analysis for education. in: r. neves-silva, j. watada, g. philipps-wren, l. c. jain, & r. j. howlett (eds.), intelligent decision technologies (pp. 353 362), amsterdam: ios press. doi: 10.3233/978-1-61499-264-6-353 asparouhov, t., & muthén, b. (2012). multiple group multilevel analysis. mplus web notes: no. 16. retrieved march 5, 2020 from https://www.statmodel.com/examples/webnotes/webnote16.pdf asparouhov, t. & muthén, b. (2019). comparison of models for the analysis of intensive longitudinal data,structural equation modeling: a multidisciplinary journal, 00: 1–23. https://doi.org/10.1080/10705511.2019.1626733 banister, s., reinhart, r., & ross, c. (2014). using digital resources to support personalized learning experiences in k-12 classrooms: the evolution of mobile devices as innovations in schools in northwest ohio. in m. searson & m. ochoa (eds.), proceedings of society for information technology & teacher education international conference 2014 (pp. 2715-2721) . chesapeake, va: association for the advancement of computing in education. retrieved march 5, 2020 from https://www.learntechlib.org/primary/p/131202/ . bastiaansen, j. a., kunkels, y. k., blaauw, f., boker, s. m., ceulemans, e., chen, m., … bringmann, l. f. (2019, march 21). time to get personal? the impact of researchers’ choices on the selection of treatment targets using the experience sampling methodology. preprint retrieved on august 24, 2019 from https://doi.org/10.31234/osf.io/c8vp7 bingham, a. j., pane, j. f., steiner, e. d., & hamilton, l. s. (2018). ahead of the curve: implementation challenges in personalised learning school models. educational policy, 32(3), 454 – 489. https://doi.org/10.1177/0895904816637688 barrett, l. f. (2018). how emotions are made. the secret life of the brain. mariner books: new york. battle, a., & wigfield, a. (2003). college women’s value orientations toward family, career, and graduate school. journal of vocational behavior, 62, 56–75. https://doi.org/10.1016/s0001-8791(02)00037-4 beretvas, s. n. (2010). cross-classified and multiple membership models. in j. j. hox & j. k. roberts (eds.), handbook of advanced multilevel analysis (pp. 313–334). new york, ny: routledge. bieg, m., goetz, t., wolter, i., & hall, n. c. (2015). gender stereotype endorsement differentially predicts girls' and boys' trait-state discrepancy in math anxiety. frontiers in psychology, 6, 1404. https://doi.org/10.3389/fpsyg.2015.01404 carroll, e. a., czerwinski, m., roseway, a., kapoor, a., johns, p., rowan, k., & schraefel, m. c. (2013). food and mood: just-in-time support for emotional eating. 2013 humaine association conference of affective computing and intelligent interaction . geneva, switzerland. chauliac, m; catrysse, l. ; gijbels, d. & donche v. (2020). it is all in the surv-eye: can eye tracking data shed light on the internal consistency in self-report questionnaires on cognitive processing strategies? frontline learning research. 8 (3), 26 – 39. https://doi.org/10.14786/flr.v8i3.489 chung, h., kim, j., park, r., & jean, h. (2018). the impact of sample size in cross-classified multiple membership multilevel models. journal of modern applied statistical methods, 17 (1), article 26. https://doi.org/10.22237/jmasm/1542209860 cole, j. s., bergin, d. a., & whittaker, t. a. (2008). predicting student achievement for low stakes tests with effort and task value. contemporary educational psychology, 33, 609–624. https://doi.org/10.1016/j.cedpsych.2007.10.002 corno, l. (2008). on teaching adaptively. educational psychologist, 43(3), 161–173. https://doi.org/10.1080/00461520802178466 creswell, j. w. & guetterman, t. c. (2019). educational research: planning, conducting, and evaluating quantitative and qualitative research , 6th edition, pearson. dietrich, j., viljaranta, j., moeller, j., & kracke, b. (2017). situational expectancies and task values: associations with students' effort. learning and instruction, 47, 53–64. https://doi.org/10.1016/j.learninstruc.2016.10.009 dietrich, j., moeller, j., guo, j., viljaranta, j., & kracke, b. (2019a). in-the-moment profiles of expectancies, task values, and costs. frontiers in psychology, 10:1662. https://doi.org/10.3389/fpsyg.2019.01662 durik, a. m. & jenkins j. s. (2020). variability in certainty of self-reported interest: implications for theory and research. frontline learning research. 8 (3) 85-103. https://doi.org/10.14786/flr.v8i3.491 douglas, h., (2011). facts, values, and objectivity. in: i. jarvie & j. zamora bonilla (eds.), the sage handbook of philosophy of social science, 513–529, london: sage publications. durik, a. m., vida, m., & eccles, j. s. (2006). task values and ability beliefs as predictors of high school literacy choices: a developmental analysis. journal of educational psychology, 98, 382–393. https://doi.org/10.1037/0022-0663.98.2.382 eccles, j. s., adler, t. f., futterman, r., goff, s. b., kaczala, c. m., meece, j. l., & midgley, c. (1983). expectancies, values, and academic behaviors. in j. t. spence (ed.), achievement and achievement motives (pp.74–146). san francisco, ca: freeman. eccles, j. s., & wigfield, a. (2002). motivational beliefs, values, and goals. annual review of psychology, 53, 109–132. https://doi.org/10.1146/annurev.psych.53.100901.135153 efron, b. and tibshirani, r. (1993) an introduction to the bootstrap. chapman and hall, new york, london. eisner, e. (1992). objectivity in educational research. curriculum inquiry, 22(1), 9-15. https://doi.org/10.1080/03626784.1992.11075389 enders, c. k. (2010). applied missing data analysis. new york, ny: the guilford press. fahrenberg, j. (1996). ambulatory assessment: issues and perspectives. in: fahrenberg, j. & myrtek, m. (eds.). (1996). ambulatory assessment: computer-assisted psychological and psychophysiological methods in monitoring and field studies (pp. 3 – 20) . seattle, wa: hogrefe and huber. university of freiburg i. br., germany fink, b. (1991). interest development as structural change in person-object relationships. in: oppenheimer l., valsiner j. (eds) the origins of action. springer, new york, ny. https://doi.org/10.1007 firstbeat (2012). heart beat based recovery analysis for athletic training. firstbeat whitepapers. retrieved from: http://www.firstbeat.fi/physiology/white-papers fisher w. p. jr. (2000). objectivity in psychosocial measurement: what, why, how. journal of outcome measurement, 4(2), 527-563. fryer, l. k. & nakao k. (2020). the future of survey self-report: an experiment contrasting likert, vas, slide, and swipe touch interfaces. frontline learning research, 8 (3),10-25. https://doi.org/10.14786/flr.v8i3.501 glanzberg, m. (2018). truth. in: edward n. zalta (ed.), the stanford encyclopedia of philosophy (fall 2018 edition), retrieved from https://plato.stanford.edu/archives/fall2018/entries/truth/ göllner, r., wagner, w., eccles, j. s., & trautwein, u. (2018). students’ idiosyncratic perceptions of teaching quality in mathematics: a result of rater tendency alone or an expression of dyadic effects between students and teachers? journal of educational psychology, 110(5), 709–725. https://doi.org/10.1037/edu0000236 goetz, t., bieg, m., lüdtke, o., pekrun, r., & hall, n. c. (2013). do girls really experience more anxiety in mathematics? psychological science, 24(10), 2079-2087. https://doi.org/10.1177/0956797613486989 green, a. s., rafaeli, e., bolger, n., shrout, p. e., & reis, h. t. (2006). paper or plastic? data equivalence in paper and electronic diaries. psychological methods, 11, 87–105. https://doi.org/10.1037/1082-989x.11.1.87 hartigan, j. a., & hartigan, p. m. (1985) the dip test of unimodality. annals of statistics, 13, 70–84. hektner, j. m., schmidt, j. a., & csikszentmihalyi, m. (2007). experience sampling method. measuring the quality of everyday life . thousand oaks, ca, us: sage publications. hidi, s., & renninger, k. a. (2006). the four-phase model of interest development. educational psychologist, 41, 111-127. https://doi.org/10.1207/s15326985ep4102_4 ketonen, e., dietrich, j., moeller, j., salmela-aro, k., & lonka, k. (2018). the influence of autonomous and controlled daily goals on positive and negative emotional states: an experience sampling approach. learning and instruction, 53, 10-20. https://doi.org/10.1016/j.learninstruc.2017.07.003 krapp, a. (1998). entwicklung und förderung von interessen im unterricht [development and promotion of interest in instruction]. psychologie in erziehung und unterricht, 45, 186-203. krapp, a. (2002). structural and dynamic aspects of interest development: theoretical considerations from an ontogenetic perspective. learning and instruction, 12(4), 383-409. https://doi.org/10.1016/s0959-4752(01)00011-1 krapp, a., & fink, b. (1992). the development and function of interests during the critical transition from home to preschool. in k. a. renninger, s. hidi, & a. krapp (eds.), the role of interest in learning and development (pp. 397–429). hillsdale, nj: lawrence erlbaum associates. krapp, a., hidi, s., & renninger, k. a. (1992). interest, learning and development. in k. a. renninger, s. hidi, & a. krapp (eds.), the role of interest in learning and development (pp. 3–25). hillsdale, nj: lawrence erlbaum associates. krothapalli, k. s. & koolagudi, s. g. (2013). emotion recognition using speech features. london: springer lüdtke, o., robitzsch, a., trautwein, u., kunter. m. (2009). assessing the impact of learning environments: how to use student ratings of classroom or school characteristics in multilevel modeling. contemporary educational psychology 34, 120–131. https://doi.org/10.1016/j.cedpsych.2008.12.001 maechler, m. (2016). package ‘diptest’. hartigan's dip test statistic for unimodality – corrected . r package. retrieved march 5, 2020 from https://cran.r-project.org/web/packages/diptest/diptest.pdf moeller, j., dietrich, j., viljaranta, j., & kracke, b. (2019). data, r and mplus codes for disentangling objective characteristics of learning situations from subjective perceptions thereof, using an experience sampling method design. rerieved from https://osf.io/yszvm/. https://doi.org/10.17605/osf.io/yszvm moeller, j., ivcevic, z., white, a., & brackett, m. a. (2018). mixed emotions: network analyses of intra-individual co-occurrences within and across situations. emotion,18(8), 1106-1121. https://doi.org/10.1037/emo0000419 popper, k. r. (1934 [2002]), logik der forschung [the logic of scientific discovery], berlin: akademie verlag. prenzel, m., krapp, a. & schiefele, h. (1986). grundzüge einer pädagogischen interessentheorie [outline of an educational interest theory]. zeitschrift für pädagogik, 32(2), 163-173. reitzle, m. & dietrich, j. (2019). from between-person statstics to within-person dynamics. diskurs kindheitsund jugendforschung, 3-2019, 319-339. https://doi.org/10.3224/diskurs.v14i3.06 schmitz, b. & perels, f. (2011). self-monitoring of self-regulation during math homework behaviour using standardized diaries. metacognition & learning, 6, 255-273. https://doi.org/10.1007/s11409-011-9076-6 schönbrodt, f. d. & perugini, m. (2013). at what sampe size do correlations stabilize? journal of research on personality, 47, 609-612. https://doi.org/10.1016/j.jrp.2013.05.009 shiffman, s., stone, a. a., & hufford, m. r. (2008). ecological momentary assessment. annual review of clinical psychology, 4, 1-32. https://doi.org/10.1146/annurev.clinpsy.3.022806.091415 silge, j. & robinson, d. (2017). text mining with r: a tidy approach. sebastopol, ca: o’reilly takarangi, m. k. t., garry, m., & loftus, e. f. (2006). dear diary, is plastic better than paper? i can’t remember: comment on green, rafaeli, bolger, shrout, and reis (2006). psychological methods, 11, 119 –122. https://doi.org/10.1037/1082-989x.11.1.119 tibshirani, r. & leisch, f. (2019). bootstrap: functions for the book "an introduction to the bootstrap. r package . https://cran.r-project.org/web/packages/bootstrap/index.html wickham, h., chang, w., henry, l., pedersen, t. l., takahashi, k., wilke, c., & woo, k. (n.d.). jittered points. retrieved from: https://ggplot2.tidyverse.org/reference/geom_jitter.html wickham, h. (2016). ggplot2: elegant graphics for data analysis . springer-verlag new york. codepen introduction publication frontline learning research special issue vol.8 no.5 (2020) 1 4 issn 2295-3159 the happy victimizer pattern in adulthood – state of the art and contrasting approaches: introduction to the special issue eveline gutzwiller-helfenfingea, karin heinrichsb auniversity of duisburg-essen, germany buniversity of education upper austria, austria keywords: happy victimizer phenomenon; moral cognitions; moral emotions; adulthood info corresponding author email: eveline.gutzwiller-helfenfinger@uni-due.dedoi: https://doi.org/10.14786/flr.v8i5.681 introduction the happy victimizer phenomenon (hvp) relates to the stable finding that young children attribute positive emotions like happiness to a rule transgressor despite judging the transgression as wrong (arsenio & kramer, 1992; arsenio & lover, 1995; nunner-winkler, 1999; 2012; nunner-winkler & sodian, 1988), whereas older children attribute negative emotions like shame or guilt. various studies suggest that the hvp disappears in the course of (moral) development (for reviews, see for example arsenio, gold, & adams, 2006; krettenauer, malti & sokol, 2008). in the moral developmental literature, various, partly interrelated explanations and interpretations of this phenomenon have been suggested. the classical explanation as offered by nunner-winkler and sodian (1988) implies that moral cognitions, that is, making and justifying moral judgments, evolve before moral motivation. the attribution of positive emotions to a rule transgressor is seen as indicating a lack of moral motivation. based on the assumption that (negative) moral emotions like guilt can be seen as an indicating that the self does not only know a moral rule but also feel committed towards it (malti, gummerum, keller, & buchmann, 2009), the hvp can be interpreted to the effect that a lack of negative moral emotion attributions coincides with a lack of moral commitment. another, transition-oriented explanation (arsenio et al., 2006; lagattuta, 2005; krettenauer et al., 2008) views the hvp as a developmental transition based on a dis-integration of moral rule knowledge and moral motivation (as assessed by moral emotion attributions and justifications thereof): whereas young children already possess appropriate moral rule knowledge, their development of moral motivation, that is, prioritising moral values over hedonistic needs, is delayed. according to this understanding, the hvp would be restricted to (early) childhood. this position, however, is challenged by recent empirical evidence suggesting that the hvp can also be found in adolescence and adulthood and seems even to be widely spread (e.g., heinrichs, minnameier, gutzwiller-helfenfinger, & latzko, 2015; krettenauer, asendorpf, & nunner-winkler, 2013; krettenauer & eichler, 2006; nunner-winkler, 2007). therefore, the question arises whether the hvp actually does disappear in the course of sociomoral development. this question is essential: if the hvp represents a transitional stage affecting all (or at least the vast majority of) children and adolescents, then the occurrence of the hvp in adolescence and adulthood must represent either a developmental delay or even a deviation. first longitudinal findings do not yet offer a clear picture (krettenauer et al., 2013) however, cross-sectional research can be used to address the very basic question whether the patterns consisting of moral judgment, emotion attribution and respective justifications found in adulthood actually do represent the happy victimizer phenomenon as documented in children. moreover, there are indications that patterns of moral decision-making may also differ according to the specific context or situation referred to (e.g., bienengräber, 2011). we may therefore assume that happy victimizing in adulthood represents a phenomenon that is distinct from the phenomenon studied in children. we therefore suggest that in adolescents and adults, it is more appropriate to speak of happy victimizer patterns, that is, patterns of moral reasoning and emotion attributions which, while having some similarities with the hvp on the surface, carry different meanings and represent more complex moral functioning. this special issue presents new ideas to explain patterns of moral decision-making in adolescence and adulthood. in paper 1 by heinrichs, gutzwiller-helfenfinger, latzko, minnameier, and döring, the current state of the art in happy victimizer research is presented, with a focus on the core theoretical and methodological issues involved in trying to disentangle the phenomenon and the pattern. in papers 2, 3, and 4, empirical evidence is presented on three levels, each level being addressed in at least one of the papers. the levels refer to (a) the emergence of the happy victimizer pattern in adolescence and adulthood; (b) the personal determinants of the happy victimizer pattern; and (c) situational variations in the manifestation of the pattern. moreover, each paper represents a specific theoretical perspective towards studying the pattern. each of these approaches has different pedagogical implications. thus, the action-theoretical explanation (heinrichs, kärner, & reinke) assumes that cognitive control strategies, in particular moral disengagement strategies, play an important role in forming an intention (as the first phase in the process of acting) in the case of an ambivalence or tension between cognitive, emotional, and motivational states. forming an intention to act requires a decision on the part of the individual as well as his/her commitment to a specific course of action. the emotion development perspective (gutzwiller-helfenfinger & latzko) is grounded in the expectation that adults who display the specific judgment-attribution-justification pattern differ from adults not displaying it with respect to their justifications of emotion attributions. the cognitive-structural explanation of the hvp (minnameier) postulates that it can be reconstructed as a specific moral judgment structure (cf. minnameier, 2012) which is applied in specific situations that can be modelled game-theoretically. in their comment to this special issue, gertrud nunner-winkler and beate sodian, the researchers first investigating the role of moral motivation in the course of children’s moral development, critically discuss the ideas and approaches presented and evaluate the relative merit of the respective positions in explaining the occurrence of happy victimizing in adolescence and adulthood. gaining deeper insights into the happy victimizer and happy victimizing contributes to a better understanding of children’s, adolescents’ and adults’ social, emotional and moral – what we call “sociomoral” – learning and development. sociomoral literacy, that is, successfully engaging in meaningful, positive and caring relationships is both a prerequisite for and consequence of successful teaching and learning processes at school (malti, häcker, & nakamura, 2009) and in other learning environments and is especially important in a globalized society (latzko & malti, 2010). if teachers are to foster students’ sociomoral competencies in diverse classrooms and schools, where multiple, sometimes divergent and contradictory social, religious, and moral values are present and often clash, they need a deeper understanding of children’s and adolescents’ moral functioning. thus, both recent curricula and professional standards for teachers have made explicit reference to the necessity of fostering positive social relationships, openness and tolerance towards diversity, conflict resolution, and democratic values (e.g., australian institute for teaching and school leadership, 2011; kultusministerkonferenz, 2014), all of which are based on sociomoral competencies. references arsenio, w. f., & kramer, r. (1992). victimizers and their victims: children's conceptions of the mixed emotional consequences of moral transgressions. child development, 63, 915-927. https://doi.org/10.1111/j.1467-8624.1992.tb01671.x arsenio, w. f., & lover, a. (1995). children’s conceptions of socio-moral affect: happy victimizers, mixed emotions, and other expectancies. in m. killen & d. hart (eds.), morality in everyday life (pp. 87-130). new york: cambridge university press. arsenio, w. f., gold, j., & adams, e. (2006). children's conceptions and displays of moral emotions. in m. killen & j. g. smetana (eds.), handbook of moral development (pp. 581-610). new jersey: erlbaum. australian institute for teaching and school leadership. (2011). australian professional standards for teachers. melbourne: education services australia. bienengräber, t. (2011). situierung oder segmentierung? – zur entstehung einer differenzierten moralischen urteilskompetenz [situationism or segmentation? – on the formation of differentiated moral judgment competence]. zeitschrift für berufsund wirtschaftspädagogik, 107(4), 499-519. gutzwiller-helfenfinger, e., & latzko, b. (2020). happy victimizing in emerging adulthood: reconstruction of a developmental phenomenon?frontline learning research, 8(5), 47-69. https://doi.org/10.14786/flr.v8i5.382 heinrichs, k., kärner, t., & reinke, h. (2020). an action-theoretical approach to the ‘happy victimizer’ pattern – exploring the role of moral disengagement strategies on the way to action , frontline learning research, 8(5), 24-46. https://doi.org/10.14786/flr.v8i5.386 heinrichs, k., minnameier, g., gutzwiller-helfenfinger, e. & latzko, b. (2015). „don’t worry, be happy“? – das happy-victimizer-phänomen im berufsund wirtschaftspädagogischen kontext [the happy victimizer phenomenon in a vocational and business educational context]. zeitschrift für berufsund wirtschaftspädagogik, 111(1), 31-55. heinrichs, k., gutzwiller-helfenfinger, e., latzko, b., minnameier, g. & döring, b. (2020). happy-victimizing in adolescence and adulthood – empirical findings and further perspectives, frontline learning research, 8(5), 5-23. https://doi.org/10.14786/flr.v8i5.385 krettenauer, t., & eichler, d. (2006). adolescents' self-attributed emotions following a moral transgression: relations with delinquency, confidence in moral judgment, and age. british journal of developmental psychology, 24, 489-506. https://doi.org/10.1348/026151005x50825 krettenauer, t., asendorpf, j. b., & nunner-winkler, g. (2013). moral emotion attributions and personality traits as long-term predictors of antisocial conduct in early adulthood findings from a 20-year longitudinal study. international journal of behavioral development, 37(3), 192-201. https://doi.org/10.1177/0165025412472409 krettenauer, t., malti, t., & sokol, b. (2008). the development of moral emotions and the happy victimizer phenomenon: a critical review of theory and applications. european journal of developmental science, 2, 221-235. doi: 10.3233/dev-2008-2303 kultusministerkonferenz. (2014). standards für die lehrerbildung: bildungswissenschaften. [standards for teacher education: educational sciences]. berlin: sekretariat der kultusministerkonferenz. lagatutta, k. h. (2005). when you shouldn’t do what you want to do: young children’s understanding of desires, rules, and emotions. child development, 76, 713-733. https://doi.org/10.1111/j.1467-8624.2005.00873.x latzko, b., & malti, t. (eds.) (2010). children’s moral emotions and moral cognition – developmental and educational perspectives. new directions for child and adolescent development, no. 129. https://doi.org/10.1002/cd.v2010:129 malti, t., gummerum, m., keller, m., & buchmann, m. (2009). children’s moral motivation, sympathy, and prosocial behavior. child development, 80, 442-460. https://doi.org/10.1111/j.1467-8624.2009.01271.x malti, t., häcker, t., & nakamura, y. (2009). sozial-emotionales lernen in der schule [socioemotional learning in schools]. zurich, switzerland: pestalozzianum verlag. minnameier, g. (2020). explaining happy victimizing in adulthood – a cognitive and economic approach, frontline learning research, 8 (5), 70-91. https://doi.org/10.14786/flr.v8i5.381 nunner-winkler, g. (2012). moral [morality]. in e. schneider & u. lindenmann (hrsg.), entwicklungspsychologie (s. 521-541). weinheim: beltz. nunner-winkler, g. (2007). development of moral motivation from childhood to early adulthood. journal of moral education, 36(4), 399-414. https://doi.org/10.1080/03057240701687970 nunner-winkler, g. (1999). development of moral understanding and moral motivation. in f. e. weinert & w. schneider (eds.), individual development from 3 to 12 (pp. 253-292). cambridge: cambridge university press. nunner-winkler, g. & sodian, b. (1988). children’s understanding of moral emotions. child development, 59, 1323-1338. doi: 10.2307/1130495 scharlau, körper et karsten frontline learning research vol.7 no. 4 (2019) 25 57 issn 2295-3159 plunging into a world? a novel approach to undergraduates’ metaphors of reading ingrid scharlaua, miriam körbera and andrea karstena apaderborn university, germany article received 19 august 2019/ revised 10 october/ accepted 15 october/ available online 22 november abstract although there is considerable research on and knowledge about students’ con-ceptualizations of learning or academic practices and skills, the variability of these conceptualizations has been consistently neglected. in the present study, we ad-dress this variability in the field of academic reading with the help of a novel ap-proach. drawing on qualitative metaphor analysis, we report a detailed system of students’ conceptual metaphors of reading. our specific methodological approach to identify the structure of these conceptual metaphors allows to analyze subjective agency on a lexical as well as grammatical level. the conceptual metaphors we identified by this method are mar¬kedly variable, although they create an overall impression of medium to low agency, that is a reader who is only weakly active or potent. interrater reliability of the coding system was very good. we also report and analyze the frequency of the conceptual metaphors in a sample of 143 texts written by bachelor students. keywords: metaphor, conceptual metaphor, metaphor analysis, academic reading, transitivity info.corresponding author: ingrid.scharlau@uni-paderborn.de doi: 10.14786/flr.v7i4.559 1. introduction academic reading and writing are challenging activities at the core of learning. many factors contribute to their difficulty: they comprise different lowto high-level skills that have to be learned and used at appropriate stages or points in time and with the appropriate effort, and they presuppose adequate awareness of one’s doing as well as self-regulation of cognition, motivation and behavior. furthermore, texts vary considerably across different disciplines or discourse communities so that different reading or writing practices are necessary. so far, the focus of research has been more on academic writing. however, in higher education and science, reading and writing are tightly intertwined (e.g. in the task of discourse synthesis; spivey 1997, boscolo & mason, 2001, wiley & voss, 1999, o’hara et al., 2002; for a more general view see, e.g., shanahan, 2015) and similarly complex. some problems of academic writing may result from problems with reading, for instance inadequate reading strategies, low reading motivation or inappropriate reading goals. in the present paper, we take a look at a factor that possibly contributes to the problems of academic reading but has rarely been recognized and even less frequently tackled rigorously in recent research – the way in which members of academia talk and think about reading and writing. for instance, a german advice book for students addresses the approach of problem-oriented reading in the following way: “it is concentrated reading, that aims less at forming scientific judgment but at a direct involvement with a thematic field and the utilization of the read material for one’s own argumentation” (krajewski 2015, p. 52; translation is). the words krajewski uses, german “ausbilden” (forming), “auseinan¬der¬¬setzung” (involvement) and “verwendung” (utilization) all imply – and are likely meant to imply – an active person who handles objects. interestingly, though, the sentence does not mention the acting person and does not realize the activities as verbs – there is no student who reads and wants to achieve something, judges, grapples with a field or argues and thereto uses others’ words. to give a different example of such metaphor use, in the fourth edition of the “psychologist’s companion” addressed at undergraduates as well as graduates, the author writes: “pursue the references that your teacher and textbook cite, and pursue the references most frequently cited in these references. by digging into the literature on the topic, you will acquire a deeper understanding of the issues that are the focus of psychological research” (sternberg, 2005, p. 36). there are three prominent metaphors in this quote, pursuing, which evokes the image of reading as a path or motion, digging, drawing on exhausting work with material and evoking a distinct spatial dimension (deep vs. shallow), and acquiring which implies that there are things to be collected. in contrast to the first example, the grammatical form is congruent with the activity: sternberg directly addresses a student who pursues, digs and acquires. metaphors are also present in more academic texts. for instance hermida (2009), drawing on earlier research by bowden and marton (2000), writes: „a surface approach to reading is the tacit acceptance of information contained in the text. […] the deep reader focuses on the author’s message, on the ideas she is trying to convey, the line of argument, and the structure of the argument. the reader makes connections to already known concepts and principles and uses this understanding for problem solving in new contexts“ (hernandez, 2009, p. 21). most of the metaphors in this text are so common that they easily go unrecognized, for instance the idea that texts contain something (information, ideas etc., often depicted as things that can be linked to each other or not), the contrast between surface and depth, the image of the line of an argument. interestingly, in terms of metaphors, the two approaches contrasted by the author seem to be less different than one might think at first sight: in both cases, texts are about things that can be conveyed or taken up and assembled. deep readers actively link these ‘bits of text’ and may construct new meaning, but both they and surface readers act upon the bits. there is another metaphor in the quote that draws its meaning from a different source – negotiating –, but it is mentioned only once. why should metaphors matter? we would argue that it generally matters in higher education how we speak about issues and even more so in educational settings, but there is a more straightforward reason for the importance of metaphor. the cognitive linguists lakoff and johnson (1980) were among the first to claim that all thinking is metaphorically structured (e.g. lakoff & johnson 1980; landau, meier & keefer, 2010). whether one follows this strong (and controversial) thesis or not, conceptual metaphors can be regarded as cognitive tools (gibbs, 1994; kövecses, 2002, 2010) that allow to understand something new or abstract (the so-called target domain) in terms of something known or concrete (the source domain). that is, students trying to understand academic reading may do so with the tools of metaphors – their own, their teachers’, those in textbooks or other sources. accordingly, cognitive linguistic studies have shown that linguistic metaphors, superficial as they might seem at first sight, can give insight into the conceptual systems of speakers. however, the linguistic forms used by speakers of a given community also shape semantic concepts by highlighting certain characteristics and downplaying others. for instance, imagining reading as linking bits of information highlights the cognitive effort of understanding different parts of texts, but downplays the author as well as the text as a whole. “according to cognitive linguists, language not only reflects conceptual structure, but can also give rise to conceptualisation. it appears that the ways in which different languages ‘cut up’ and ‘label’ the world can differentially influence non-linguistic thought and action” (evans & green, 2006, p. 101). this holds true not only for the language of a given society, but also for the way more fine-grained speech communities (like academic discourse communities) speak. for example, the way a computer scientist conceptualizes reading may differ from the way a sociologist does. this is due to both differing experiences with reading in the respective discipline and the way the academic community typically talks about reading. taking this stance, it is likely that metaphors shape thinking about academic reading and writing and thereby may influence writers’ feelings, strategies and behavior. generally, research in the conceptual-metaphor framework has shown that metaphors influence what persons attend to or which information they remember as well as their perception and attitudes (for an overview see landau et al., 2010). one thus can draw the presumption that metaphor has its share not only in clarifying as well as obscuring practices in academia, but also some influence on actual behavior. this opens up a wide range of possibilities for metaphor-related didactic action in the field of teaching academic writing and reading. to give an example, bean (2011, p. 138) suggests to use metaphor games and so-called extended analogies in order to understand academic writing. in these exercises, the students compare writing to familiar experiences or things (“writing is like …” or “journal writing is like …, but formal essay writing is like …”). similarly, in the decoding-the-disciplines framework, a general approach to identifying and overcoming discipline-specific bottlenecks to student learning (e.g., middendorf & shopkow, 2018; pace, 2017), metaphors are used to unpack abstract and coarse descriptions of tasks, for instance “critical reading”. the metaphors are chosen so that they detail what an expert would do in the respective situation. teachers model this behavior including its steps with the help of and in the frame of the metaphor. for instance teachers could model their recursive reading strategies with the help of the metaphor of a spiral, returning repeatedly to the same sentence in a text to evaluate it again and again on the basis of their growing understanding of the text. to sum up, there is evidence that metaphors matter for academic engagement, are widely used in discourse on academic practices and seem to – often inadvertently – shape how students understand these practices and shape their behavior. so far, however, there are few methodologically and theoretically rigorous studies on such metaphors. in the present study, we address this gap in the field of academic reading. we collect a large sample of metaphors of reading from undergraduate students and use several approaches to cluster and understand them. such a system, once established, could be used to attend to inadvertent metaphors in academic discourse or to choose helpful metaphors, but also to study the role of metaphors for academic success, matches or mismatches between the metaphors of staff and students, and to compare different disciplines or describe changes by academic enculturation. 2. approach to metaphors in contrast to earlier studies, our aim is to put the rich content of metaphor into focus in order to take a differentiated view on the topic. as mentioned above, metaphor use and recognition are important cognitive strategies because they allow to understand new or abstract concepts in terms of interactions with the physical world or bodily states, for instance as taking things up or immersing oneself into a book. many studies analyzed those metaphors in terms of a few, often only two or three, underlying broad categories (paulson & armstrong, 2011; paulson & theado, 2015; saban, kocbeker & saban, 2007; wegner & nückles, 2015a, 2015b) – such as learning as uptake of pre-existing ideas, learning as problem solving and learning as development of personality (wegner & nückles, 2016). however, people use many more metaphors than just these, and this variety is not only across persons (and groups such as different student groups or different disciplines and cultures), but may also arise within persons, for instance when talking to different others or in different situations or faced with different tasks. with the current study, we want to identify a wide and as comprehensive as possible range of metaphors that students spontaneously use for academic reading and to explore whether there is indeed substantial heterogeneity in a structural and cognitively relevant sense. straightforward metaphoricity is visible on the lexical level, i.e. in the metaphorically used nouns, verbs and adjectives that are related to aspects of our everyday world and lives and with which people describe more abstract concepts. for instance, depicting reading as immersion into a world would have implications different from reading as opening a treasure chest, for instance in terms of activities, size and ‘thinginess’. the sources will point out, downplay or even conceal different aspects of the target domain; this is what lakoff and johnson (1980) call the hiding and highlighting aspect of metaphors. to give an example, one of the most famous metaphors used in western culture is heraclitus’ saying that no man can step into the same river twice. the german philosopher blumenberg (2012, p. 103) noted that this metaphor highlights that reality is ever-changing, but downplays or even hides that stepping into a river and descending from it imply a bank that does not change and even is the same bank farther downwards the stream. while the bigger part of studies on metaphors in different domains focuses on this lexical aspect, we take a more cognitive-linguistically oriented stance and additionally attend to metaphoric conceptualizations on a grammatical level. in practice, this means that we analyze students’ conceptions of academic reading also with regard to questions such as how much agency is attributed to readers, if the focus lies on the products and outcomes of reading or the processes and activities, and how purpose-oriented the latter are. these conceptualizations are expressed less by the choice of explicit metaphors (that is, which metaphors participants choose to conceptualize reading) but rather through the linguistic constructions in metaphorical utterances (that is, which qualities the specific metaphorical expressions attribute to the actors and the activity of reading). more systematically, our novel approach is to suggest that transitivity – the effectiveness with which an action takes place as it is semantically construed and linguistically expressed (hopper & thompson, 1980) – is an important aspect of the transfer of meaning that is inherent in metaphors, at least those used for human action and cognition. hopper and thompson list 11 components that can be rated as – mostly – high or low or present or not present and that contribute to the overall degree of transitivity. these are the number of participants of an action (two or more – one) kinesis, that is whether the event is an action or not (action – non-action) aspect, i.e. having an end-point or not (telic – atelic) punctuality (punctual – nonpunctual) volitionality (volitional – nonvolitional) the agency of the actor, that is the degree to which actors can affect things or events (high in potency – low in potency) the affectedness of the object of the action (highly affected – not affected) the individuation of the object of the action (highly individuated – not individuated) affirmation (an action happens or does not happen) mode (realis – irrealis) these components contain important information about activities, and we used them to characterize the conceptual metaphors we identified in the material, thus specifying lakoff and johnson’s (1980) rather vague notion of the structure of metaphors. the transitivity status of a metaphor can be determined on two levels. one is the level of conceptual metaphors, focusing on the typical transitivity of a metaphorical image. for example, describing reading as immersing oneself into a world implies little agency whereas reading as combining things into a whole highlights volitionality and purposiveness, relatively independent from the exact expressions the speaker uses to invoke these conceptual metaphors. the choice of a conceptual metaphor is thus likely to feed into the way in which specific metaphoric expressions are construed on a grammatical level. therefore, in many cases, specific utterances comply with the typical transitivity of a metaphorical source domain. however, they may as well be inconsistent with it, for instance if a verb such as forming, which implies high agency, was used with a different actor than the reader, for example describing reading as something forming in one’s mind. this distinction has consequences for analyzing metaphors. a purely grammatical approach to transitivity on the utterance level would attest high or low transitivity to a linguistic expression regardless of who or what the agent of the utterance is. slightly departing from this, we are going to focus on the reader-as-agent. thus, we will speak of high or low transitivity of a metaphor looking at the degree to which the reader as our agent in question is conceptualized as active, purposeful etc. this means to always consider both the level of conceptual metaphor and at the content of the metaphor as well as the grammatical form on the utterance level. determining the transitivity status of a metaphor in this way helps us to see if the reader is depicted as agent (or non-agent) both by the choice of metaphor and the grammatical realization. to sum up, our study addresses students’ conceptualizations of academic reading inherent in their metaphors for this practice. we are breaking new ground in so far as we, firstly, want to identify and preserve the heterogeneity of metaphors as far as possible, and, secondly, choose a structured approach to the analysis of conceptual metaphors based on the agency of the reader. 3. methods participants. participants were 143 students enrolled in bachelor programs (teacher training and arts and humanities) at paderborn university, north-rhine-westphalia, germany, a mid-sized university with approximately 20.000 students many of which are enrolled in teacher training programs. all students gave fully informed written consent. the study was approved by the ethics committee of paderborn university.1 procedure and materials. data were collected in four classes. for the teacher training students (two of the classes), the seminar is part of the bachelor module introduction to education studies. its focus is on reading and analyzing psychological research on motivation. the students aim for high and middle school and business school (gymnasium/gesamtschule, haupt-, real und gesamt¬schule, berufskolleg) and should attend the class in their first or second semester although many do so later in their course of studies. the students in arts and humanities (the other two of the classes) are enrolled in a bachelor program with two subject areas from the humanities or social sciences (e.g., education, german literature, british/american literature, philosophy). here, the seminar is part of a so-called orientation module. the module covers key skills and the seminar topic was an interdisciplinary introduction to academic reading and writing. in the first session of the seminar, we handed out a sheet with the following instruction: “(1) please complete this sentence with the image, analogy, or metaphor which comes to mind first. reading is like ... if you come up with several images, analogies, or metaphors, name these. (2) describe or explain your image or metaphor in a few sentences. for instance you can describe similarities between reading and your image or motivate why you have chosen it.“2 the students completed the sheets in a self-paced manner and had as much time for writing down their answers as they wanted. the instruction explicitly required a source metaphor and an explanation or description, a method that has successfully used by, for instance, paulson and armstrong (2011), saban, kocbeker and saban (2007) or weg¬ner and nückles (2015a, 2015b, 2016). the texts were typed before analysis so that handwriting information was removed. no personal information except for first language(s) was collected. the instruction addressed reading in a generic manner instead of asking for academic or scientific reading. we chose this approach for several reasons. firstly, the metaphors were sampled in reading-intensive seminars at university. contextual cues thus were in favor of academic reading. secondly, there are some drawbacks of more specific instructions that would have to be carefully weighed against generic instructions. for instance talking about academic reading might point the students toward the expectation that there is a difference between kinds of reading that they might not have found relevant in the first place. in a preliminary study in which we asked for metaphors of academic reading and writing, the answers tended towards definitions instead of metaphors, which would undermine the whole approach of metaphor research. thirdly, and most importantly, we do not yet know how generic or specific metaphorical concepts of reading are. we are not aware of any research that addresses this question. it is possible that students have different concepts for different genres (such as scholarly texts, novel, news, etc.) or contexts (such as at home, at work, at university), but it is also possible that they have a generic concept or a hierarchy of generic to specific metaphors (similar to general to specific self-concepts or self-efficacy). finally, many texts and much talk addressing academic practices at german universities are similarly generic and it is exactly the spontaneous metaphors that are activated in such situations that we are interested in. however, we cannot make sure that the generic approach misses some important metaphors. 3.1 analysis of metaphors text examples. in order to convey the ‘flavor’ of the texts, we will first give a few examples. one student, for instance, wrote: “reading is like a wall. like standing in front of a wall and not really knowing exactly what is behind. it's the same with a text/book. only when you look behind the wall, into the text, you know exactly what it is about. if you are inside reading a text/book, then this is also a wall that keeps many other external influences away, if you have sunk yourself completely into the text/book“ (132). this is a text of rather strong metaphoricity, as is the next example that however gives a very different overall impression: „reading is like lying in a hammock and letting your soul ‘dangle’. you can switch off while reading and flee from everyday life. it is exactly the same when you lie in a hammock and think of nothing or sink into the world of thoughts” (123). some writers explicitly used and explained several metaphors, as in the next example. „reading is like immersing yourself in another world. reading is like watching a movie, only more detailed. reading is like home.
reading brings you into a completely new environment. you can leave your everyday life behind you and come to rest. i chose this picture because it entices me to immerse myself in a new story and experience it ‘up close’. i chose the term ‘home’ because reading is something comfortable for me. it feels familiar and gives me a good feeling” (113). some students explored a metaphor somewhat thoroughly: “swimming in the sea. the sea consists on the one hand of unknown depths, but on the other hand also of shallow shore zones. if one swims in the sea, one never knows where the next sandbank comes or where it suddenly gets deeper. this is exactly the same when reading. there are simple, easy and shallow places, but also deeper places. reading immerses you in the unknown” (81). other texts were short and more difficult to interpret and understand. one example is “immersing oneself in another world an adventure one disappears from one’s own reality perceives other feelings and things one feels oneself into other figures, forgets oneself for the blink of an eye exciting” (35). there were also some rather vague and nonimaginative texts, such as: “reading is like learning by heart. i chose this picture because i read the information over and over again while learning by heart. by reading i can keep the information better. if i read out key points aloud, it is even easier for me to keep them” (110). unit of analysis. as the texts were short and dense, we decided to use the text as unit of analysis (wegner & nückles, 2015a, 2015b, used a similar task and made the same decision). an individual text consisted of several phrases, the naming of source metaphor and at least one phrase with a description/explanation. metaphors named in the source and in the description were not necessarily consistent, nor did metaphors in different phrases of the same text. all metaphors were coded independently of whether there was a congruent overall picture in the text of an individual participant. development of the coding system. the main goal of the present paper is to develop a system of conceptual metaphors of reading. we first identified metaphorical expressions in the texts and then assigned them to a conceptual metaphor (e.g. schmitt, 2017). to give an example, one student wrote (metaphorical expressions underlined): “reading is like opening a treasure chest. every book you open brings new knowledge or new insight. at the same time, however, you will also find treasures in novels that can help you escape from everyday life or give you a short holiday. readers can only win and not lose anything.” in many cases, identification of metaphorical concepts from metaphorical expressions was straightforward, as in the following example: “plunging into another world. for me, reading is like immersing oneself in another world, because i can forget everything around me when i read an exciting book.” here, we identified the conceptual metaphor of plunging and the conceptual metaphor of forgetting. these metaphors are congruent, but not conceptually identical. in other cases, identification was less straightforward, as in the following example: “a deluge of information. on the one hand, reading is a deluge of information, because there is always a lot of information in what one reads. this information then patters on you like a deluge.“ this example is more difficult to interpret and was finally identified as being flooded. as it was the single example of reading as being flooded with elements in the sample of our texts, we decided to list it as an idiosyncratic metaphor. having identified possible conceptual metaphors, we used the two approaches from the literature mentioned above to describe them. the first, lexical approach is closely related to the structural approach of lakoff and johnson (1980) and is captured in the method of qualitative metaphor analysis (schmitt, 2017). here, the source domain and its conceptual content were described. in the above examples this would be plunging into something, forgetting something and being flooded. the second approach draws on transitivity analysis (hopper & thompson, 1980). for each conceptual metaphor, we described the transitivity according to nine of the eleven components of transitivity mentioned in the introduction. (mode and aspect were not used because they were fixed as realis/indicative and non-negation by our task: the stem completion “reading is like …” did not offer to answer “is not like” or “might be like”.) plunging, for instance involves only one person. it is rather volitional, but punctual, and there is no individuated object of the action – it affects the person rather than an object. there is no clear telos or end-point. the person is at most medium in potency. overall, thus, transitivity is medium. in the following, we describe transitivity mostly on the level of the conceptual metaphors. in many cases and for the reasons given above, conceptual transitivity was consistent with the grammatical constructions in the texts. if however, the specific utterances did not correspond with conceptual transitivity, for instance because they used the passive voice, an indefinite object such as ‘something’ or no object at all, we included this in the descriptions of the metaphors. transitivity analysis was especially effective in separating different conceptual metaphors that draw on the same source domain. for example, the activity of moving appeared in different forms and with different degrees of transitivity, as strongly volitional motion (for instance walking along a path) or as being moved (for instance riding on a roller coaster). these instances could be differentiated by their transitivity. with the help of these two steps, we identified a large system of possible conceptual metaphors that is presented below. we combined similar metaphors into metaphorical systems that capture the essence of their members in terms of source domain and transitivity. the conceptual metaphors were developed by one of the authors (is) and later discussed in the whole team of authors. while coding, the coders could add further categories. definitions and examples of metaphors were written down in a short manual. then, two independent coders (sb: not involved in developing the code system, mk: involved only in adding missing codes) coded 20 texts and compared and discussed their results. the other 123 texts were coded independently. interpretation and coding were done in german; codes were later translated into english. in some cases, translation was difficult because the semantically most appropriate translation was reflexive in one language and not reflexive in the other, or transitive in one language, but intransitive in the other. for instance, german “in etwas eintauchen” is not reflexive whereas its english equivalent “immersing oneself into sth” is; the german word for “switching off, “abschalten”, can be used intransitively meaning that a person relaxes, but this is not the case in english. because of our focus on transitivity, we always chose the grammatically congruent translation and discuss the exact semanticity of the german wording where necessary. 4. results 4.1 qualitative results: the coding system overall, 66 conceptual metaphors were identified. we bundled similar metaphors into groups with a distinct conceptual characteristic. these groups are handling things, moving in or out, moving around, seeing/opening and being acted upon. in contrast to earlier studies (e.g., paulson & armstrong, 2011; wegner & nückles 2015, 2016), we do not integrate the conceptual metaphors within such a group into one overarching conceptual metaphor ¬– the names of the groups are convenient descriptions, not metaphors. the reason is that there is still considerable variability within each group in terms of agency and conceptualization and we want to preserve this variability. the following paragraphs describe the conceptual groups and their transitivity. the definitions of the single conceptual metaphors can be found in appendix a, each with an example sentence and a short description of their conceptual content and transitivity. furthermore, the most frequent conceptual metaphor within each group is described along with the overall description of the group. 4.1.1. handling things: solidification metaphors this group of conceptual metaphors (see table 1) is defined by an actor who handles things. one common example was taking up, as in the following text (the most frequent conceptual metaphor is discussed below): “reading is breathing knowledge. when you read texts, you always take something up, whether you like it or not” (62). conceptually, such handling of things would imply high transitivity. interestingly though, the specific metaphors we found and their use in the texts was more of medium than of high transitivity, as in the example above in which the uptake is depicted as rather involuntary. table 1 gives an overview of the metaphors and their transitivity aspects. the handling, an action, was often telic and volitional, but in other cases lacked these properties. although an individuated object would suggest itself in the context of handling, it often was not individuated in the texts. for instance, students often wrote that “something” was taken up. if included explicitly, the object often was an abstract concept, such as impressions, ideas or knowledge. although the actor is depicted as high in potency, the object is not necessarily affected – on the contrary: it can be unaffected as in holding, weakly affected as in combining, strongly affected as in processing or even coming into existence as in creating. other persons are conspicuously absent from the descriptions. only two metaphors (receiving and transporting) imply other persons, but these were not mentioned. the action is often extended in time, but can also be punctual. overall, thus, transitivity can be described as medium in these metaphors, finding, receiving and sth forms being three exceptions with very low agency, and creating and processing two exceptions with very high agency. the most frequent metaphor in this group was sth forms. a typical example is: “while reading, images form in the head” (32). this metaphor may not seem a genuine part of the “handling things” system because nothing is handled here, but it fits the solidification aspect which is inherent in many of the metaphors. its transitivity is largely opposite to the other members of this system: instead of a person handling something and changing its features, here the thing changes its features by itself. in very many cases, the object that forms is an image, that is a faintly individuated object. table 1 handling things: solidification metaphors and their transitivity aspects. the sign “+” denotes the presence of an aspect, “–“ its absence. a question mark indicates that the feature is unclear, “±” signals that the feature differed in the texts. if a sign is put into brackets, the feature is present, but not strongly so and may be of limited value for characterizing the metaphor 4.1.2 moving in and out there were two groups of conceptual metaphors that centered around locomotion, mostly of the reading person. we grouped these metaphors into two systems because in some of them, the locomotion goes into something or out of it whereas the second one lacks this distinct moment of in–out. most, though not all moving metaphors are an action, and in most cases, this action is telic (for an overview, see table 2). volitionality, though, varies across the texts – in some cases, the actor seems to be induced by an external cause to move, for instance when leaving the everyday world, sinking, or sinking oneself into a book. also conspicuously, most of the actions are punctual or short, even if their effects are long-lasting, as in plunging. none of the metaphors has a direct object that is acted upon or handled; what is affected is the actor, not an object. conceptually, the moving metaphors often go along with the idea of a world or space and different locations. conspicuously again, other persons are absent, even though often a world is mentioned. overall, thus, transitivity is rather weak in this group. this group contains the most frequent metaphor in our material in which reading is com-pared to plunging into a different world. it is a common expression for reading in german (“in etwas eintauchen”), but it is nevertheless striking that it is also used that frequently in the context of academic reading. a typical example is: “reading is like plunging in another world or situation” (124). plunging is a volitional, telic and punctual action. there is only one participant. her or his potency remains indistinct: the only action mentioned is that one enters (a punctual act), and there is no affectedness of an object – if one wants to regard the world as an object. it is difficult to decide wheth-er the world is individuated because conceptually it indicates “everything”. the actor, however, seems to be affected – the conceptual metaphor implies a shift in mode of existence. transitivity of this con-ceptual metaphor is similar to entering (see below in the appendix). the main difference and the reason that we did not combine these metaphors into a single one is that plunging means immersing oneself completely into something. although water was rarely mentioned in the texts, the german word used in all examples, “tauchen” or “eintauchen”, implies a liquid that surrounds the actor. therefore, the shift in existence mode is rather strong. this metaphor has a distinct spatiality in the sense that the person enters into a world. this world is usually not described in any detail, it is simply a container filled with undefined things. table 2 moving in and out metaphors and their transitivity aspects. the sign “+” denotes the presence of an aspect, “–” its absence. a question mark indicates that the feature is unclear, “±” signals that the feature differed in the texts. if a sign is put into brackets, the feature is present, but not strongly so and may be of limited value for characterizing the metaphor 4.1.3 moving around as mentioned above, this category is the second locomotion category. one prominent dif-ference is that moving around does not imply a movement into or out of a location or space. however, the conceptual metaphors assigned to this category also implied a somewhat higher transitivity than those of the moving in and out group (see table 3). many, but not all conceptual metaphors imply an action. this action often is telic and volitional, but several cases lack both features such as if one inad-vertently meets something or is being led by something. the action is typically extended in time. ob-jects are mostly absent and thus not affected. in contrast to the moving in and out group, there is also less affectedness of the person in this group. others are absent from the metaphor; even in being led, which implies somebody or something that leads, this other is never an author but rather the story. overall, this is also a group of medium transitivity. furthermore, there is considerable var-iability between conceptual metaphors: the action can be highly transitive (as in the case of explor-ing) or passive/low in agency (as in the case of being led or being moved). the most frequent metaphor in this group is traveling. an example is: “a journey in my imagination. […] you can visit fantastic worlds, relive their adventures with historical people or ac-quire new knowledge” (71). this is a common metaphor with a single lexical expression (german noun “reise” or, rarely, the verb “reisen”). the activity is usually of high kinesis and volitional. its telos is not clearly present. it is temporarily extended, that is it takes a long time. usually, there is no object mentioned (“reisen” is intransitive), but other (rare) expressions such as “visiting something” could mention such an object. although traveling might include several participants, others were never mentioned in our material. traveling may imply a change in the actor. (in the example above, this is expressed by a different conceptual metaphor, acquiring). transitivity is medium. in traveling, there is spatiality in the sense of a large, but rather indistinct area and a path, whenever a journey in the sense of a distinct trail is alluded to. table 3 moving around metaphors and their transitivity aspects. the sign “+” denotes the presence of an aspect and “–” its absence. a question mark indicates that the feature is unclear, “±” signals that the feature differed in the texts. if a sign is put into brackets, the feature is present, but not strongly so and may be of limited value for characterizing the metaphor. 4.1.4 being acted upon these metaphors do not necessarily differ in conceptual content from metaphors in other groups; the difference is rather in who is active and who is been acted upon. in all cases, what was highlighted was that the reader is influenced or forced by reading, i.e. presented as the object of an ac-tion. in most cases, the utterances used the passive voice so that the agent of the action was masked. looking from the perspective of the reader-as-agent, these metaphors thus are of very low transitivity (see table 4). to give an example, the most frequent metaphor in this group was being influenced or stimulated. a typical example is: “reading stimulates imagination” (8). this is a rather pale metaphor and in many cases it came with the conventional expression – as in the example above – that imagina-tion is stimulated. agency is exceptionally low, and the reader is presented as someone to whom some-thing happens. the source of the influence was either a book or the act of reading. table 4 metaphors of being acted upon and their transitivity aspects. the sign “+” denotes the presence of an aspect, “–” its absence. a question mark indicates that the feature is unclear, “±” signals that the feature differed in the texts. if a sign is put into brackets, the feature is present, but not strongly so and may be of limited value for characterizing the metaphor 4.1.5 seeing and opening these metaphors (see table 5) are grouped together because in western cultures, seeing is a common metaphor for knowledge and insight, and this conceptual content is invoked by all of them. most of them involve a weakly agentive action of (visual) perception in which there is an actor and an object, but the two do not have a direct contact (as in the handling things group). although opening has only an indirect connection to seeing, it is grouped here because it renders things visible. the action can be telic, volitional and temporally extended, but need not be so if, for in-stance, somebody sud¬denly and unwittingly discovers something. a further remarkable feature is that the object of the action is usually not affected by this action; one exception is imagining in which the objects comes into being. other persons are again absent from the conceptual metaphors, even though the existence of different perspectives may imply them. the most frequent conceptual metaphor in this group was seeing. an example is: “watch-ing movies, only it’s all in your head. by reading novels or other books, for example, images are con-structed in the head/brain of each person that match what they have read” (53). seeing has a spatiality characterized by distance. in our material, seeing was usually presented as something that happens and thus is of low kinesis (there were no instances of observing or scrutinizing which would be more active). it is usually not described as strongly telic or volitional. although seeing can be punctual, many of the examples in our material were extended in time, for instance when reading was compared to watching a movie – as in the example above. the object can be individuated or not, but it is not af-fected by being seen. in our material, objects of seeing were images or films rather than things. some instances depicted the head as a container of images. overall, thus, this conceptual metaphor has a low transitivity. table 5 seeing and opening metaphors and their transitivity aspects. the sign “+” denotes the presence of an aspect, “–” its absence. a question mark indicates that the feature is unclear, “±” signals that the feature differed in the texts. if a sign is put into brackets, the feature is present, but not strongly so and may be of limited value for characterizing the metaphor 4.1.6 other metaphors there were a few distinct conceptual metaphors that did not fit into any of the five groups mentioned above. they are listed in table 6 and again explained in appendix a. except for broaden-ing, all of them were rare in the present material. of course, it is not possible to describe these concep-tual metaphors as a whole, neither in conceptual content nor in transitivity. overall, they match the impression gained from the further metaphors: transitivity varies, but is often low to medium and rare-ly high. the most frequent conceptual metaphor in this group was broadening. an example is: “by reading you personally develop and broaden your horizon”(87). broadening has a distinct spati-ality. transitivity is interesting because approximately half of the cases in our sample came with an active voice (see the example), the other half in the passive form (one’s knowledge is broadened). therefore, kinesis, potency, volitionality and aspect are unclear and marked as variable in the table. furthermore, broadening was always used in a cognitive sense (i.e. for broadening knowledge or horizon) and thus could also have been categorized as a cognitive metaphor. lexically, however, it is more concrete than the cognitive metaphors. note however, that broadening is a common expression in german, as are shutting off, which underlines the relaxing function of reading, and being someone else in the sense of being one of the persons in a text. three conceptual metaphors focus on freedom (freeing, being free and letting/letting go), three others on work and exertion (working, grappling and exerting one-self). one single metaphor, speaking, stresses the existence of others, though these were not men-tioned in the texts. table 6 other metaphors and their transitivity aspects. the sign “+” denotes the presence of an aspect, “–” its absence. a question mark indicates that the feature is unclear, “±” signals that the feature dif-fered in the texts. if a sign is put into brackets, the feature is present, but not strongly so and may be of limited value for characterizing the metaphor 4.1.7 cognitive processes and emotion-centered metaphors perhaps not surprisingly in the academic context, a variety of the texts clustered around cognitive processes. these are not conceptual metaphors in the strict sense, because they usually are as abstract as the concept to be explained by them. because these expressions were so common, we want to report them, well aware that they are analogies rather than metaphors. the same arguments can be put forward for a group of emotion-centered descriptions to which we want to draw attention because most of them have emotional content – which is conspicuously absent from almost all of the other metaphors. the cognitive analogies are listening, perceiving, learning, forgetting, focusing, questioning and experiencing, the emotional analogies are relaxing/enjoying/finding comfort, vacation and empathizing with others. a few texts compared reading explicitly to a negative expe-rience such as stepping barefoot on a lego brick. the emotional analogies are low in transitivity (the effects being on the side of the partici-pants whose agency is low), the cognitive analogies have mixed transitivity, varying from very low in forgetting or dreaming over medium in listening, perceiving and experiencing to reasonably high in questioning and grappling. 4.2 quantitative results the 139 texts (4 descriptions were missing, possibly because the students overlooked the page with “reading is like ...”) were coded by two independent coders one of whom was partly and the other not at all involved in the development of the codes. each identified close to 500 conceptual met-aphors (mk: 490 codings, sb: 477 codings), ranging from zero to eleven metaphors per text. interrater agreement for the 66 categories was 82.42, resulting in interrater reliability or cohen’s κ = .84 (cohen, 1960) which can be regarded as excellent agreement (fleiss & cohen, 1973; wirtz & caspar, 2002). figure 1 gives an overview of the frequency of the different groups. conceptual metaphors of locomotion are the most common. they account for roughly 30% of all codings. seeing/opening and handling things metaphors each cover about 14%. metaphors which put into focus that the person is being acted upon are comparably rare. figure 1: percentages of codings in the different groups. within the groups, there is a very unequal distribution of codings. very few metaphors are used widely; many others are very rare. table 7 gives the frequencies of all categories and figure 2 depicts the percentages of codings for those conceptual metaphors that had at least 10 codings. figure 2: percentages of codings for the more frequent metaphors. colors as in figure 1. the three most frequent metaphors are plunging, traveling and seeing, followed by imagining and being influenced. these metaphors are conceptually very different. plunging and traveling draw on a change of location, one punctual, the other one temporally extended. plunging implies that reading is something that a person only has to kick off. this is fairly similar in escaping, although the activity here is somewhat longer and has more ‘force’. seeing puts the activity of percep-tion into focus. imagining is similar because most texts in which this metaphor was found focused on imagining visual material, but it implies higher transitivity. being influenced is an internally variable category in which the reader is depicted as passive and suffers from what is happening. sth forms implies the emergence of an object independent of the reading person, taking up and receiving the transfer of an object from one location to another. broadening and being in/part of sth underline changes in the reading person, though of different kinds. yet, the metaphors are similar in that all are at most medium in transitivity. except for imagining and sth forms, none of the metaphors implies an effect on the object. with the exception of the same two conceptual metaphors, all metaphors imply an effect on the actor that may be weak and reversible as in seeing and receiving, but also strong as in escaping. most of the metaphors have low aspect (as in receiving) or low volitionality (as in escap-ing) or neither (as in forgetting). the only frequent conceptual metaphors that depict readers as effective and active agent are imagining and traveling. in both cases, however, there is no effect of this agent on the world. table 7 conceptual metaphors of reading together with their frequency 5. discussion focusing on metaphors, the present study addressed one aspect of individual understand-ing of academic reading. in contrast to earlier studies (paulson & armstrong, 2011; paulson & theado, 2015; for studies on other academic practices see saban, kocbeker & saban, 2007; wegner & nückles, 2015a, 2015b, 2016), we identified a large variety of conceptual metaphors. although we could cluster them into a few discernable groups, there is considerable intra-group variability, especially if one takes transitivity, that is, the amount of agency and impact, into account. this main finding agrees well with a study by paulson and armstrong (2011) who ana-lyzed the reading and writing metaphors of college students enrolled in developmental reading and writ-ing classes. they classified the metaphors according to negative/nonnegative responses and with re-spect to literacy as a process vs. literacy as a product. negatively connoted metaphors were rare (ap-proximately on sixth of the metaphors) and product metaphors were almost three times as frequent as process metaphors. furthermore, paulson and armstrong semantically analyzed the metaphorical ex-pressions and identified several different conceptual metaphors, among them college reading is a journey, college reading is a sport, college writing is a gaming activity, and college writ-ing is freedom of expression. overall, the authors underline the variability of conceptualizations. somewhat surprisingly they do not directly target the variability and report only very few conceptual metaphors. this gap is filled by the present study. however, we also identified some rather common features. leaving the emotional and cog-nitive metaphors aside, students’ conceptualizations mainly drew on the concrete actions of handling things, of locomotion or being moved and of seeing and opening. one further group centered around different experiences of being acted upon. although transitivity or the amount of agency in our metaphors varied from extraordinarily low to very high, many metaphors were at most medium in transitivity. the most frequent aspects of low transitivity were the absence of other persons, who were rarely mentioned or implied by the con-ceptual content, the absence of individuated objects and a frequent lack of affectedness of the object. what is more, the reading person was often depicted as affected (strictly linguistically speaking as the object of a rather transitive action by someone or something else). reading repeatedly appeared as a passive uptake of information – a notion that has been characterized as an intuitive everyday notion of reading (christmann & groeben, 1999, p. 145) –, but more often as being affected rather than affecting or acting. this overall picture contrasts with psychological theories of reading (e.g. christmann & groeben, 1999; schaffner, 2009) and educational approaches to supporting academic reading (e.g. bean, chappell & gillam, 2014). according to current models of reading (e.g., kintsch, 1987; kintsch & van dijk, 1978), text comprehension is an active and complex process of meaning-making, that is, an interaction between information in the text and reader knowledge or expectations (even though some reading processes may be regarded as passive resulting in differences in the amount to which readers are seen as active, cook & o’brien, 2019). this interaction is regulated by the reader. such regulation does not necessarily imply conscious control, but it does imply activity, a flexible, cognitive and con-structive interaction. furthermore, different processes are involved in reading, some more bottom-up (that is in-fluenced by the text, for instance its coherence), others more top-down (influenced by the reader’s knowledge, goals or expectations). extraction of meaning, for instance, can be studied at the levels of words, word combinations and sentences as well as the meanings of whole texts. also, reading does not only involve cognitive processes of understanding, but also attention, emotion, motivation and self-regulation of these processes. in terms of the transitivity aspects defined by hopper and thompson (1980), according to cognitive or psychological theories reading has to be seen as an action, consisting of several different activities that have to be controlled and coordinated. it is temporally extended and – taking meaning as its object – has a highly individuated object which is strongly affected by reading. although reading may be automatic, reading of academic texts will be only weakly automatic but rather telic and volition-al with a reader who is high in potency. this is almost the opposite of what we found in the concepts of the students. the same striking contrast between students’ metaphors and scholarly conceptualizations holds true for sociocultural theories of reading. according to ridley (2004), encounters and struggles with academic reading practices belong to crucial literacy and learning experiences for students. in the sociocultural literature on academic literacy, reading is often treated much less explicitly than writing. some of the conceptions related to reading from a sociocultural perspective have thus to be inferred from the respective understanding of writing, which is presented in a more straightforward fashion. perhaps this has to do with the fact that sociocultural theories view writing as the participation in social discourse(s), which already implies a to-and-fro-movement between the various interacting persons as readers-and-writers. it is exactly this interactive view of discourse that is most characteristic for sociocultural conceptualizations, be they explicitly or only implicitly focused on reading. the vygotskian notion of sociogenesis of higher mental functions and the brunerian concept of scaffolding in language learning are at the core of a sociocultural understanding of reading (langer & applebee, 1986). from this per-spective, interaction with adults and “more capable peers” (vygotsky, in cole et al., 1978) is central to grow into the reading practices of a given society and even of particular groups within that society. through legitimate peripheral participation learners become part of their communities (lave, 1991; lave & wenger, 1991) – with regard to academic reading, this is the interaction in class (with faculty and peer students) and in informal peer settings (havnes, 2008). we would expect students to mirror the interactive nature of reading and learning to read in their metaphors, e.g. by depicting the reader as social agent who is in connection to other persons. contrasting more cognitive concepts of reading, from a sociocultural stance readers can be conceptualized both as agentive and as patientive actors, i.e. rather as acting on others (e.g. talking to persons, discussing, arguing or even fighting with them) or as being acted upon (e.g. being guided). in this line of argument, social status and power relations are seen to be crucial to academic literacies (lea & street, 2006; lillis, 2003; lillis & scott, 2007). we would thus expect them to feature in students’ metaphors and especially in their transitivity status. another characteristic of reading from a sociocultural perspective is that it draws not on language in general but on a variety of social languages, depending on which socially situated discours-es the reading activity is part of (gee, 2013). in the realm of academic reading and writing, this fact is most attended to in genre theories (swales, 1990; for an overview of sociocultural studies, see russell, 1997). we would expect students to mirror this fact by providing a differentiated use of metaphors and by distinguishing between metaphors for reading in a variety of genres and social settings. however, we could not find such pondering over the adequacy of metaphors for different types of texts and functions of reading in our material. we have to note, however, that our questions “reading is like …” did not ask for such a differentiation. why is it the case that most metaphors that the students in our sample produced are rather different from or even stand in contrast to the main points of scientific theories of reading? the following answers to this question are speculative in the sense that they only address possible reasons. therefore, they should be regarded as questions for further research. one possibility is that students produced everyday notions of reading based on common expressions for reading in their everyday language. as mentioned above, “plunging into” is not too uncommon an expression associated with reading in the german language. such everyday notions can, as research on preconcep-tions and subjective theories has shown, differ clearly from academic concepts and may exist in addi-tion to them, even if they contradict one another. another possibility is that the students addressed leisurely reading instead of reading of academic texts. assuming that leisurely reading is less active than reading of scientific texts, the difference might then be due to two different concepts. note howev-er that this pertains more to the subjective experience of reading than to the processes driving it; these are the same and similarly constructive, although somewhat less effort or regulation of motivation may be necessary for leisurely reading. a third possibility is that many of the processes involved in read-ing, including the constructive ones, may be intransparent to the reader. once automated, routines such as moving the eyes towards relevant words and leaving out predictable ones or reading more slowly when working through difficult and dense parts of texts may be so inconspicuous that the reader does not recognize their constructive nature. one further interesting idea (we are thankful to a reviewer who pointed this possibility out) is that metaphors may capture aspects of reading that are not that easily observed within scientific models of reading. modern empirical psychology, successful as it is, is ra-ther suspicious of subjective, phenomenological or introspective, approaches to practices such as read-ing. seeing reading as, for instance, plunging or seeing might capture interesting dimensions of reading, for instance the phenomenal experience that the content is there (in seeing) and that context changes a person (as in plunging); the individual might even be much less autonomous than psychological think-ing claims. note, however, that the disagreement to scientific theories is as strong in the case of psy-chological models as in that of sociocultural approaches which aim to be more true to actual experi-ence. our present material was not intended to identify the causes of the discrepancy between meta-phorical concepts and scientific theories of reading, but it reveals this discrepancy as a fascinating area for further research. to turn to possible practical consequences of our study, academic reading and writing are difficult problems for undergraduates, and it appears that even if universities seek to support the stu-dents, these problems are rather persistent. the results of the present study indicate one reason for these difficulties: students’ conceptualizations of reading may differ strongly from how staff or re-searchers understand such practices. besides pointing out this problem, the present study also indi-cates possible lines of action. metaphors can be understood as a variant of analogical reasoning, in which representa-tions/their structures, one better and one less understood, are aligned (e.g., gentner, 1983, 2010). ana-logical reasoning in turn can be used for scaffolding understanding of complex issues such as scientific phenomena – or academic practices such as thinking, writing or reading. for this, one would have to identify the structures of the source domain – for instance the lexical and grammatical aspects de-scribed in the results section – and compare them to academic reading in order to find out which ele-ments match and which match less or not at all.3 such an approach does not imply that there is a correct metaphor of academic reading. sfard (1998), for instance, pointed out that the acquisition metaphor for learning (such as in learn-ing is eating) is often considered a simple and insufficient metaphor, inferior to, for instance, encul-turation or participation metaphors (as in learning is growing or learning is becoming part of something). however, she explicitly warned against such a simplification. metaphors establish perspectives and thus help to look at things in related and necessary, but disparate and maybe even not reconcilable ways. for instance, the seemingly simple acquisition metaphor helps to understand indi-vidual change or transfer and transfer problems better than the apparently more complex participa-tion metaphor. the same could be argued with respect to the reading metaphors. if we take only the most frequent ones, plunging could be seen as simplistic: it implies that reading has only to be kicked off, what is done is only one punctual step; afterwards, no activity is necessary. this does not capture the activity and self-regulation that are necessary for constructing meaning from scientific texts, especially if one is very new to a field or domain, but it accords better with more expert reading in which the pro-cess of meaning construction requires little effort and, taking this perspective, is not simplistic. to give another example, the traveling metaphor highlights that reading may change the person and involve a change of location or context. again, this reveals aspects typical of academic reading – for instance that it is an activity with a long temporal perspective and a practice that is typical for special educational settings at universities and different form ‘ordinary life’, but may not be easily transferable to other communities. on the other hand, it tends to hide the necessary effort on the side of the reader. seeing, finally, hides the reader’s activity as well as the fact that meaning is negotiated and not something that is simply there, but it may single out the distance that is helpful in understanding academic texts. departing from transitivity analysis and in accordance with paulson and armstrong’s claim that „classroom conversations need to directly address [student] conceptualizations of academic litera-cies“ (2011, p. 501), researchers could work actively with and on metaphors that are typically used by students as well as staff. note that we do not recommend directly changing student metaphors. drawing on the understanding that all metaphors highlight some aspects and hide others (lakoff and johnson, 2003) and that academic practices are too complex, variable and negotiated to be understood in a single perspective (sfard, 1998) we suggest to focus on addressing or changing some of their aspects, espe-cially certain transitivity features. for instance, seeing has very low transitivity.¬ one may see some-thing involuntarily, in just a moment and without effort – features that would be incongruent with aca-demic reading. this match would be, however, improved if one moved from seeing to imagining, in which the person creates the imagined content, or observing, which implies more activity that is fur-thermore extended in time. both seeing and observing stress that reading has to be true to the text whereas imagining underlines the constructive activity of the reader. together, these metaphors may illustrate the interplay of bottom-up and top-down processes in reading as well as different and contra-dictory demands on the reader. turning to plunging, this metaphor may be used to underline the dif-ference between academic and everyday reading. however, the impression of being immersed into something can also be true of academic reading, and even though reading is an active and constructive process, it also has an element of immersion that cannot be produced, but only be set off by the reader. to underline the temporal extension and recursiveness of reading as well as effort and agency, the plunging metaphor could be turned into a diving metaphor. turning to the third of the three most frequent metaphors, traveling, it may be used to underline the existence of different discourse communities as well as the fact that students switch back and forth between different discourse com-munities and often do not feel completely at home in any of them – often not even as a legitimate pe-ripheral participant (lave & wenger, 1991). freebody and luke (1990) stress the different roles of the reader, which are all suitable to be developed by extending and re-shaping students’ metaphors. the reader as code-breaker highlights the technical, deciphering and rule-following nature of reading. this reader-role could be developed within students’ metaphors that involve handling of things as well as the entrance into unknown worlds. the role of reader as text participant who infers meaning from background knowledge and constructs and connects knowledge (see the cognitive reading theories discussed above) could be strengthened if academic reading is conceptualized as active handling of things involving shaping or building metaphors common to writing (scharlau, rohlfing & karsten, subm.). this active image of the reader could be even further extended to conceptualize the reader as text user, when the metaphors imply handling of things in real-life and metaphors of the text as a tool to achieve goals in the world. the fourth reader role, the reader as text analyst, can be developed if the reader is seen to be in dia-logue and interaction with other persons, involving metaphors of discourse and dialogue, discussion and argumentation (see the sociocultural reading theories discussed above). academic reading and writing researcher bean (2011, p. 138) reports that students enjoy playing with metaphors when grap-pling with topic content and has described methods to use metaphors in the classroom (see also elbow, 1998). in a longitudinal study, wegner and nückles (2015) concluded that students’ metaphors of learning developed during their first year at university. although in approximately half of their sample the students’ main metaphorical content or source concept did not change, in the other half it did. stu-dents who had used a collecting source (such as eating) at the time of the first assessment were likely to change their metaphors. students who had used metaphors congruent with enculturation such as a flower (growing) or journey (discovering) did not change their metaphors. the latter are more con-sistent with how universities conceptualize learning (see also wegner & nückles, 2013). such a change in the direction of more discourse-appropriate metaphors could be fostered by explicit work on meta-phors. this seems even more important because a few studies have shown that metaphors of learning can influence study practices and success (landau, oyserman, keefer & smith, 2014; ryan, 2001; wegner & nückles, 2016). the present study is, however, only a first step towards a thorough understanding of the role of metaphors as conceptualizations of academic practices. firstly, although the system of concep-tual metaphors is large and likely to cover a large amount of metaphorical descriptions, data sampling should be expanded to graduate students and experienced academics. this is necessary to substantiate the claim that our list indeed captures the range of metaphorical conceptualizations of reading. part of the divergence between our metaphor groups and scientific accounts of academic reading may result from the fact that we sampled metaphors only from undergraduate students. secondly, use of meta-phors may vary across groups, not only beginners and experts, but also members of different dis-course communities or disciplines. several studies have shown that reading practices are very different in different disciplines (e.g. shanahan, shanahan & misischia, 2011; wineburg & reisman, 2015). it might be advisable to compare metaphors at least across different disciplines such as the humanities and sciences. and finally, metaphors may vary according to genre or text type such as reading of sci-entific literature, of textbooks, fiction or news or websites. furthermore, it is yet unclear how reading metaphors are related to other conceptualizations of reading such as conceptions, beliefs or attitudes, for instance the two reading beliefs or epistemolo-gies reported by schraw and bruning (1996), one transmissional and one transactional. according to the transmissional belief, texts or authors convey meaning to readers; in the transactional belief, mean-ing is constructed by the interaction of readers, writers and text. the first view draws on a common metaphor that is used in many different context, conduit (reddy, 1979). transmissional views are present in the conceptual metaphors of taking up, receiving, and transporting metaphors in the present study, and although beliefs and metaphors are different cognitive constructs and stem from different theoretical backgrounds, it would be worthwhile to study the contribution of metaphorical language, thinking and discourse to such beliefs. keypoints students used very different metaphors for describing reading, speaking to the heterogeneity of their understanding of academic reading. students did not associate reading with an activity of high agency or impact. students’ metaphorical understanding of reading does not agree well with either cognitive or socio¬-cultural approaches to reading. acknowledgments we gratefully acknowledge the help of sonia kampel, jana schwede, and anastasia schulz with tran-scribing the students’ texts and developing a precursor of the system of conceptual metaphors, and the help of silvia burkhardt with coding the metaphors. footnotes 1 as we did not plan to analyze the metaphors depending on specific individual characteristics such as age, study se-mester, subject areas, gender etc., and we did not want to induce any expectations about the importance of these features in the students’ minds, we did not record this information. 2 the students were also asked for metaphors of writing and motivation. these will be analyzed in other papers (schar-lau, rohlfing & karsten, subm.). 3 one could also think of further structural elements such as spatiality which, among others, lakoff and johnson (2003) described as an important aspect of metaphors. references bean, j. c. (2011). engaging ideas: the professor’s guide to integrating writing, critical thinking, and active learning in the classroom (rev. 2nd ed.). san francisco, ca: jossey bass. blumenberg, h. (2012). quellen, ströme, eisberge [springs, streams, icebergs]. frankfurt am main: suhrkamp. bowden, j., & marton, f. (2000). the university of learning. london: kogan page. doi: 10.4324/9780203416457 cohen, j. (1960). a coefficient of agreement for nominal scales. educational and psychological measurement, 20, 37–46. doi: 10.1177/001316446002000104 cook, a. e., & o’brien, e. j. (2019). fundamental components of reading comprehension. in j. dunlosky & k. a. rawson (eds.), the cambridge handbook of cognition and education (pp. 237–265). cambridge, uk: cambridge university press. christmann, u. & groeben, n. (1999). psychologie des lesens [psychology of reading]. in b. franzmann, k. hasemann, d. löffler & e. schön (eds.), handbuch lesen [handbook of reading] (pp. 145–223). münchen: saur. doi: 10.1515/9783110961898.145 cole, m., john-steiner, v., scribner s., & souberman, e. (eds.) (1978). l.s. vygotsky – mind in society: the development of higher psychological processes . cambridge: harvard university press. elbow, p. (1998). writing with power: techniques for mastering the writing process . oxford: oxford university press. evans, v. & green, m. (2006). cognitive linguistics: an introduction. edinburgh: edinburgh university press. fleiss, j. l., & cohen, j. (1973). the equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability. educational and psychological measurement, 33, 613–619. doi: 10.1177/001316447303300309 freebody, p., & luke, a. (1990). literacies programs: debates and demands in cultural context. prospect: an australian journal of tesol, 5, 7–16. gee, j. p. (2013). reading as situated language: a sociocognitive perspective. in d. e. alvermann, n. j. unrau & r. b. ruddell (eds.), theoretical models and processes of reading (6th ed.) (pp. 136¬–151). newark, de: international reading association. doi: 10.1598/0710.04 gentner, d. (1983). structure-mapping: a theoretical framework for analogy. cognitive science, 7, 155–170. doi: 10.1207/s15516709cog0702_3 gentner, d. (2010). bootstrapping the mind: analogical processes and symbol systems. cognitive science, 34, 752–775. doi: 10.1111/j.1551-6709.2010.01114.x gibbs, r. w. (1994). the poetics of mind. cambridge: cambridge university press. gorzycki, m., howard, p., allen, d., desa, g., & rosegard, e. (2016). an exploration of academic reading proficiency at the university level: a cross-sectional study of 848 undergraduates . literacy research and instruction, 52, 142–162. graff, g. & birkenstein, c. (2014). “they say/i say”: the moves that matter in academic writing (3rd ed.). new york: norton. havnes, a. (2008). peer‐mediated learning beyond the curriculum. studies in higher education, 33, 193–204. doi: 10.1080/03075070801916344 hermida, j. (june 14, 2009). the importance of teaching academic reading skills in first-year university courses. doi: 10.2139/ssrn.1419247 hopper, p. j., & thompson, s. a. (1980). transitivity in grammar and discourse. language, 56, 251–299. doi: 10.2307/413757 kintsch, w. (1986). learning from text. cognition and instruction, 3, 87–108. doi: 10.1207/s1532690xci0302_1 kintsch, w., & van dijk, t. (1978). toward a model of text comprehension and production. psychological review, 85, 363–394. doi: 10.1037/0033-295x.85.5.363 kövecses, z. (2002). metaphor: a practical introduction. oxford: oxford university press. kövecses, z. (2010). a new look at metaphorical creativity in cognitive linguistics. cognitive linguistics, 21, 663–697. doi: 10.1515/cogl.2010.021 krajewski, m. (2015). lesen schreiben denken: zur wissenschaftlichen abschlussarbeit in 7 schritten [reading writing thinking: in seven steps to a scientific thesis]. köln: böhlau verlag. lakoff, g., & johnson, m. (1980). metaphors we live by. chicago: university of chicago press. doi: 10.7208/chicago/9780226470993.001.0001 landau, m. j., meier, b. p., & keefer, l. a. (2010). a metaphor-enriched social cognition. psychological bulletin, 136, 1045–1067. doi: 10.1037/a0020970. landau, m. j., oyserman, d., keefer, l. a., & smith, g. c. (2014). the college journey and academic engagement: how metaphor use enhances identity-based motivation. journal of personality and social psychology, 106, 679–698. doi: 10.1037/a0036414 langer, j. a., & applebee, a. n. (1986). reading and writing instruction: toward a theory of teaching and learning. review of research in education, 13, 171–194. doi: 10.2307/1167222 lave, j. (1991). situated learning in communities of practice. in l. b. resnick, j. m. levine & s. d. teasley (eds.), perspectives on socially shared cognition (pp. 63–82). washington dc: apa books. doi: 10.1037/10096-003 lave, j., & wenger, e. (1991). situated learning: legitimate peripheral participation. cambridge: cambridge university press. doi: 10.1017/cbo9780511815355 lea, m. r., & street, b. v. (2006). the „academic literacies“ model: theory and applications. theory into practice, 145, 368–377. doi: 10.1207/s15430421tip4504_11 lillis, t. (2003). student writing as ‚academic literacies’: drawing on bakhtin to move from critique to design. language & education, 17, 192–207. doi: 10.1080/09500780308666848 lillis, t., & scott, m. (2007). defining academic literacies research: issues of epistemology, ideology and strategy. journal of applied linguistics, 4, 5–32. doi: 10.1558/japl.v4i1.5 löfström, e., nevgi, a., wegner, e., & karm, m. (2015). images in research on teaching and learning in higher education. in j. huisman & m. tight (eds.), theory and method in higher education research, volume 1 (pp. 191–212). bigley, uk: emerald group publishing limited. doi: 10.1108/s2056-375220150000001009 marton, f., & säljö, r. (1976). on qualitative differences in learning: i—outcome and process. british journal of educational psychology, 46, 4–11. doi: 10.1111/j.2044-8279.1976.tb02980.x middendorf, j., & shopkow, l. (2018). overcoming student learning bottlenecks: decode the critical thinking of your discipline . sterling, va: stylus. pace, d. (2017). the decoding the disciplines paradigm: seven steps to increased student learning. bloomington ia: indiana university press. doi: 10.2307/j.ctt2005z1w paulson, e. j., & armstrong, s. l. (2011). mountains and pit bulls: students’ metaphors for college transitional reading and writing. journal of adolescent & adult literacy, 54, 494–503. doi: 10.1598/ja al.54.7.3 paulson, e. j., & theado, c. k. (2015). location agency in the classroom: a metaphor analysis of teacher talk in a college developmental reading class. classroom discourse, 6, 1–19. doi: 10.1080/19463014.2914.888360 prior, p. (2005). a sociocultural theory of writing. in c. a. macarthur, s. graham & j. fitzgerald (eds.), the handbook of writing research (pp. 54-66). new york: guilford press. reddy, m. j. (1979). the conduit metaphor: a case of frame conflict in our language about language. in a. ortony (ed.), metaphor and thought (pp. 284–310). cambridge: cambridge university press. doi: 10.1017/cbo9781139173865.012 ridley, d. (2004). puzzling experiences in higher education: critical moments for conversation. studies in higher education, 29, 91–107. doi: 10.1080/1234567032000164895 russell, d. r. (1997). writing and genre in higher education and workplaces: a review of studies that use cultural-historical activity theory. mind, culture, and activity, 4, 224–237. doi: 10.1207/s15327884mca0404_2 ryan, m. p. (2001). conceptual models of lecture learning: guiding metaphors and model-appropriate note-taking practices. reading psychology, 22, 289–312. doi: 10.1080/02702710127638 saban, a., kocbeker, b.n. & saban, a. (2007). prospective teachers’ conceptions of teaching and learning revealed through metaphor analysis. learning and instruction, 17, 123–139. doi: 10.1016/j.learninstruc.2007.01.003 schaffner, e. (2009). determinanten des leseverstehens [determinants of reading comprehension]. in w. lenhard & w. schneider (eds.), diagnostik und förderung des leseverständnisses [diagnostic and support of reading comprehension] (pp. 19–44). göttingen: hogrefe. scharlau, i., & karsten, a. (in prep.). heterogeneity and development in undergraduates’ metaphors of writing . manuscript in preparation. scharlau, i., rohlfing, k. j., & karsten, a. (subm.).building, emptying out, or dreaming? action structures and space in students’ metaphors of academic writing. schraw, g., & bruning, r. (1996). readers’ implicit models of reading. reading research quarterly, 31, 290–305. doi: 10.1598/rrq.31.3.4 sfard, a. (1998). on two metaphors of learning and the dangers of choosing just one. educational researcher, 27, 4–13. doi: 10.3102/0013189x027002004 shanahan, t. (2015). relationships between reading and writing development. in c. a. mcarthur, s. graham, & j. fitzgerald (eds.), handbook of writing research (2nd ed.) (pp. 194–210). new york: the guilford press. shanahan, c., shanahan, t., & misischia, c. (2011). analysis of expert readers in three disciplines: history, mathematics, and chemistry. journal of literacy research, 43, 393–429. doi: 10.1177/1086296x11424071 steen, g. (2011). the contemporary theory of metaphor now new and improved! review of cognitive linguistics, 9, 26–64. doi: 10.1075/bct.56.03ste sternberg, r. j. (2005). the psychologist’s companion: a guide to scientific writing for students and researchers (4th ed.). cambridge: cambridge university press. swales, j. m. (1990). genre analysis: english in academic and research settings. cambridge: cambridge university press. wegner, e., & nückles, m. (2013). kompetenzerwerb oder enkulturation? lehrende und ihre metaphern des lernens [competence achievement or enculturation? university staff and their metaphors of learning]. zeitschrift für hochschulentwicklung, 8, 15–29. doi: 10.3217/zfhe-8-01/04 wegner, e., & nückles, m. (2015a). from eating to discovering: how metaphors of learning change during students' enculturation. zeitschrift für hochschulentwicklung, 10, 145–166. doi: 10.3217/zfhe-10-04/08 wegner, e., & nückles, m. (2015b). knowledge acquisition or participation in communities of practice? academics’ metaphors of teaching and learning at the university. studies in higher education, 38, 624–643. doi: 10.1080/03075079.2013.842213 wegner, e., & nückles, m. (2016). training the brain or tending a garden? students' metaphors of learning predict self-reported learning patterns. frontline learning research, 3(4), 95–109. doi: 10.14786/flr.v3i4.212 wineburg, s., & reisman, a. (2015). disciplinary literacy in history: a toolkit for digital citizenship. journal of adolescent & adult literacy, 58 (8), 636–639. doi: 10.1002/jaal.410 wirtz, m., & caspar, f. (2002). beurteilerübereinstimmung und beurteilerreliabilität [interrater agreement and interrater reliability]. göttingen: hogrefe. appendix: the conceptual metaphors in this appendix, the conceptual metaphors we identified in the material are described and present-ed with an example. the coding handbook (which is largely overlapping with the following descrip-tions, but contains some further instructions and is written in german) can be requested from the first author. 1. handling things taking up/incorporating sth example: “taking up knowledge, ideas and opinions, which one could not develop by one's own effort” (92). in this metaphor, reading is described as taking up things or elements which are then inside the person. neither the things nor the actor are described as being strongly affected by taking up. the action is volitional, but neither strongly telic nor temporally extended. in our examples, the object was not strongly individuated – it was either “something”, “information” or “knowledge”. creating sth example: “during reading, images are created in your head and you can expand and decorate them with your own imagination. if you read a book, you can create your own universe in which you can escape and escape from the stress of everyday life” (140). by creating, things come into existence. therefore, this conceptual metaphor has very high transi-tivity. creating is volitional and telic, and it can be more or less punctual. the most often named ob-jects were images and films. the example above is interesting because it contains one instance of high and one of low transi-tivity: images “are created” in the head, but the person can also “create your own universe”. however, the passive version was rare in the texts. combining/linking/assembling example: “assembling. all information and characters must be processed and assembled” (2). this metaphor is special in that it involves the idea of pre-existing things that have to be assem-bled in order to form a new thing or collection. the actor acts on this combination rather than the orig-inal things that remain unchanged. the action is volitional and telic. individuation of objects was judged as medium because the students often talked of assembling “information”. collecting example: “a mosaic.
while reading, one collects many 'information particles'. from these parti-cles gradually a larger picture emerges” (47). as in combining, things are important for this metaphor, but transitivity is lower because the combination of things in not in focus. actor and objects (which are always many) appear rather inde-pendent and the objects are often not strongly individuated – the “particles” in the example above are an exception. volitionality and telos are medium – one can collect information more or less purpose-fully. the act is extended in time because it is typically repeated. finding example: “opening a treasure chest. every book you open brings new knowledge or new in-sight. at the same time, however, you will also find treasures in novels that can help you escape from everyday life or give you a short holiday” (9). in finding, transitivity is rather low; finding rather happens to the person than that s/he does it, so that it lacks volitionality and telos. the action is punctual, the object is individuated, but not affected by the action. holding/storing example: “the brain tries […] to store the content of what it reads and through this ‘film’ in the head the content can be reproduced even after some time” (30). the focus of this metaphor is on not changing things, and the agency of the actor appears in the ability to keep them as they are by holding or storing them. it is a rather common writing metaphor (scharlau et al., subm.), but much less common in the field of reading. transitivity is medium: hold-ing/storing is volitional, telic and nonpunctual, but it does not affect the object – indeed is not meant to affect it. processing example: “the difficulty in reading is to be able to absorb, process and then apply information in order to later draw comparisons or to be able to deal critically with what has been read. this is an art that is sometimes very difficult to implement, but you can learn a lot from it” (116). processing (german verarbeiten, from arbeiten, to work, working) is a common metaphor in german, also used for writing. it was rare for reading. in terms of spatiality, it always needs a thing. in terms of transitivity, processing indicates a potent actor who does something in a telic, volitional and mostly nonpunctual manner. the object is individuated and highly affected – changed – by the action. receiving example: “when reading you get a lot of input in a short time” (77). in receiving, the actor is barely active. the tiny action that is necessary for receiving an object is rarely mentioned in our example, so that transitivity is very low. although receiving (as transporting) implies an active instance besides the reader, this actor, human or otherwise, was rarely mentioned in our material. this is an illustrative aspect of the absence of others even from those metaphors that im-ply them. objects are present, but often neither individuated as the “lot of input” above nor affected by the action that takes place. transporting example: “scientific books mainly want to impart new knowledge to you” (130). transporting is an exception in this system because the metaphor was always used in the passive form (something is conveyed to me) or with texts as an actor (see above). in the everyday events it alludes to, it is similar to receiving, distinctly evoking the impression of something being transported without being affected. it can be regarded as an instance of the conduit metaphor described by reddy (1971): communication – in our case written communication – means carrying messages from one to another person. transitivity is very low, because of the passive expressions used, even lower than in receiving. sth forms see above in the description of the group 2. moving in or out plunging see above in the description of the group escaping example: “an escape from reality. at first sight, the term ‘escape’ seems to have a rather negative connotation, but in this sense it is seen as a liberating ‘way out’ of reality, which is in part exhausting or (personally) stressful” (13). this conceptual metaphor implies a strong and quick action and some necessity of acting because escape or flight imply something negative that is avoided by the action. the reader may have noted that the metaphorical descriptions are overall rather neutral. this conceptual metaphor is one of the few exceptions. what is left behind is not individuated (reality or everyday life) and not affected. overall, this is a metaphor of mixed transitivity – the actors can do something and have a goal, but with an effect only on themselves, and they are at least partly pressured to act. being in sth example: “as soon as some pages are read, one is integrated into the events of the book” (5). this is a rather pale conceptual metaphor. agency is very low – some telos and volitionality, but no objects and no effect, often even no action, as in the example above–, but spatiality is characteristic of the moving in/out system. putting oneself into example: “when you read (no matter what kind of text it is), you put yourself in the exact posi-tion that the text describes or the topic it deals with” (124). lexically, this metaphor is similar to being in sth. it has higher transitivity, because the actor does something, even if with an effect only on themselves. transitivity furthermore differs from other members of this system because the verb is reflexive: the actor appears both as the subject and the object. sinking oneself into example: “when reading, one delves into another world/situation” (13). “sich vertiefen”, the german expression, is a common metaphor for reading. there is a volitional and telic action; its extendedness in time is indistinct. instead of the object (which is usually a world or a book/text), the actor is affected, but only temporarily so. besides the spatial aspect of “in/into”, this metaphor has a further distinct spatial dimension, “deep”. note that this is a category that is difficult to translate. the german expression in all texts is „sich in etwas vertiefen“ which is a reflexive verb and implies the vertical dimension (“tief” = deep). the english translation we chose unfortunately overlaps with the conceptual metaphor sinking (german “versinken” oder “abtauchen) explained below. sinking example: “[…] sink into the world of thoughts” (123) this conceptual metaphor depicts a very passive actor who lets things happen. if there is an ob-ject, it appears as a container or world, as in the example above. besides the low transitivity, the char-acteristic aspect of this conceptual metaphor is its spatiality (sinking downwards). entering example: “entering another, unknown world” (5). entering is a volitional, telic and punctual action. the object of action is no thing, but rather a space or world which is neither individuated nor affected. the actor does not seem to be affected ei-ther although the conceptual metaphor implies some shift in mode of existence. there is some, but not high potency. leaving example: “reading brings you into a completely new environment. you can leave your everyday life behind you and come to rest” (113). leaving puts the distance to something into focus. however, the object is only left behind, not affected. although as conceptual metaphors, leaving and entering are similarly in transitivity, the ex-ample above is typical for our text material because leaving is depicted as something that affects the actor who is relieved from something. the activity has some volitionality and telic aspect. often, a whole (everyday) life or world is left behind – though only temporarily. 3. moving around traveling see above in the description of the group exploring example: “reading is like exploring.
through reading, new things are absorbed or perceived. this means that something unknown can be explored through reading” (109). the metaphor of exploring was mostly used in connection with travel and is therefore placed in the moving category. the activity is telic, volitional and nonpunctual. different from what one may have expected, the students did not mention things or perspectives to be explored but left the object rather vague, as in the “something unknown” in the example above. in our texts, objects were not in-dividuated. exploring usually does not affect the object –it should rather leave it unaffected. the actor is not changed either. overall, thus, transitivity is medium. meeting example: “like on a journey, i encounter things new, familiar, different and adventurous” (86). the action of meeting can be volitional or telic (as in meeting friends), but can also be something that happens to an actor punctually and that s/he did not intend to do. both variants were present in the material. neither actor no object are affected, so that this conceptual metaphor is of low transitivity. one of the two examples included others (meet persons). moving around example: “reading is like wandering aimlessly to an unknown destination” (115). this conceptual metaphor integrates different metaphorical descriptions (walking, wandering, moving). one instance of swimming was also coded as moving around. they are similar in that they depict a person on a long trajectory and were therefore grouped into a single category. volitionality and aspect vary in our texts – the example above shows this in a conspicuous contradiction: aimless motion towards a goal. the action itself is nonpunctual. different from the moving into/out of system, the actor is not affected either; they just change their location. following example: “one forgets the original "effort" of reading and only follows the story” (56). this is a very rare and almost dead metaphor, as it only appears in the common phrase of “fol-lowing a story”. although one cannot follow without an action, this action is presented as being con-trolled and maybe even initiated by an instance different from the actor. being led example: “reading is like immersing yourself in another world, because books and texts always lead you into different situations” (46). transitivity high, but agency very weak in being led, although it implies: there is no volitionali-ty and no goal. the activity (which is not caused by the person) is nonpunctual. the conceptual meta-phor is rather similar to following, but even lower in transitivity. backing away example: “when you read, you can completely hide the 'real' world around you. you can back away from the problems of everyday life and concentrate on this ‘other world’, where there are no limits to your imagination” (17). in transitivity, this metaphor is very similar to escaping, but it is more neutral in its emotional aspect. volitionality and aspect are usually low and the action is punctual. the is an object (often eve-ryday life or reality) is not a grammatical object (backing away from something), and it is not affected by the action. receding into the background example: “like on a journey, my current location and my surroundings step into the background. the further i get in the text, the further away they are” (86). conceptually, this metaphor is similar to backing away – it results in a distance between the per-son and an thing or event. in transitivity, however, it is different because it is the object (for instance everyday life or sorrows) are active whereas the person does not do anything. being moved example: “be promoted to another world (if the reading is exciting!)” (54). this is the conceptual metaphor with the lowest transitivity. the actors do not act at all, they are being acted upon. in this aspect, the conceptual metaphor could have been equally placed in the system “being acted upon”, but we decided to place it into the “moving around” system because of its distinct spatiality. moving example: “riding a bike. once you can ride a bike, you never forget it. it's difficult to think back to the time when you couldn't read yet and had difficulties with it” (31). this is a metaphor with very little conceptual content, but it was so frequent that we include it in the system. it is in between a metaphor and an explanation. reading is described as something learned to such a degree that it is automatic. the two examples used were cycling and driving a car. 4. being acted upon being influenced or stimulated see group description above sth awaits me example: “for me reading is like an adventure, because you never know what will await you […]” (82). this is a category of low transitivity because action is not on the side of the person. the meta-phor’s transitivity is low in all aspects. being enchained example: “if you read a good book, you are completely bound to the created world” (50). this metaphor has a conventional expression in german (“von einem buch gefesselt sein”, being captivated by a book). its transitivity is very low although there is some aspect of potency implied: being enchained implies that one cannot do what one would usually do. being confron¬ted with example: “[…] i am […] confronted with the author's emotional world, motives, and goals in a scientific text” (46). the semantic difference of being confronted with to encountering is small because in both cases a person (involuntarily) meets an object which poses a difficulty for their action, but there is a grammat-ical difference because the present metaphor is always passive. transitivity is again very low. being forced example: “the immersion in water. once you start reading something, you can't concentrate on anything else, just like you can't stop swimming in the water to do something else” (18). in this metaphor, the person is forced to act against their will. the force is usually left vague as in the present example; it is neither a person nor a thing. 5. seeing and opening seeing see description of the group above imagining example: “if you read, only your own imagination plays a role. everyone imagines the narrated world differently” (50). imagining mainly differs from seeing in that the objects seen (the images) are created by the person. therefore, transitivity is higher in imagining. this is also one of the few metaphors in which there is an object which is affected by the action. notably, though, in most cases, the texts left the ob-ject of imagining vague or abstract, as in the above example (“the narrated world”). taking perspectives example: “while reading you get a lot of input in a short time. some information is simply nec-essary to understand or clarify the whole thing. others prompt new perspectives or thoughts” (77). transitivity of this metaphor is comparable to that of seeing with the exception of slightly higher kinesis, potency and punctuality. as in the example above, the taking of perspectives may also be of low agency because these new perspectives are enforced on the person by the text. furthermore, this is one of the very few metaphors in which others were present or at least implied: one takes a differ-ent perspective than other persons. as in the example above, the others with the different perspective are often the actors themselves at a different point in time. this metaphor has a distinct spatiality in that it evokes the idea of a large area and a certain stance which can be shifted or enlarged. one might regard the horizon or the perspective as a container–indeed lakoff and johnson (2003, chapter 3), do this when discussing seeing. blinding out example: “when you read, you blind out everything else and focus strongly on what is being told” (19). in blinding out, the actor volitionally, but temporarily, blinds out things or even the whole world or everyday life in order to not be affected by them. the objects are not affected by this action. transitivity still is comparably high. the german expression for this metaphor is “etwas ausblenden” which has a clearly visual source domain and an object. in this, it is different from a related metaphor, shutting off (see below, other metaphors). opening example: “opening a treasure chest. every book you open brings a new knowledge or new in-sight. at the same time, however, you will also find treasures in novels that can help you escape from everyday life or give you a short holiday” (9). the action of opening up is not always one of the actor; possibilities may open up themselves. if there is an action on the part of the reader, it is volitional, telic and punctual. the object may be indi-viduated but it is not affected by the action; thus, the actor is not high in potency. in several of the few instances, the actors are affected by the action in the sense that they gain new insights. discovering example: “discovering and getting to know new worlds.
the similarities lie in the fact that you always get new information during reading” (100). although discovering depends on a person doing something, it is not a strongly transitive action. it is neither telic nor strongly volitional nor extended in time. the actor has low potency. the object is individuated but not affected by the action. 6. other metaphors broadening see description of the group above shutting off example: “reading helps me to relax and switch off” (1). this expression is common in german. shutting off is an action, but with little transitivity, alt-hough it is volitional and telic. its effect (which is on the actor) is reversible. it can be used transitively (e.g. shutting off one’s thoughts) as well as intransitively, and both usages appeared in our texts. freeing example: “letting one’s imagination run wild” (80). this is a rare and variable metaphor, mostly characterized by the lexical expression “free” (the german original of the example above is “seiner fantasie freien lauf lassen”, “freien” = “free”). it implies a weakly agentive action in which the actor volitionally and purposefully releases something from a former state. this something is not substantially changed, and neither is the actor. others are absent. being free example: “you can […] focus on this ‘other world’, where there are no limits to your imagina-tion” (17). as in the example above, this metaphor often came without mentioning an action. the reader is free, even though it remains unclear from what or to which aim. there is almost no transitivity – no action, no object and thus no affectedness. the source domain, however, implies a future high potency of the actor (they could do what they want to). letting go example: “let go, drop yourself and immerse yourself in another world” (56). letting go is a punctual act that can, but needs not, be telic or volitional. although the german verb is transitive and implies an object, this object was not mentioned in the metaphors. even if men-tioned, the object would not have been affected. the actor is low in potency and the action is rather the absence of action. working example: “working autonomously or exploring other points of view” (51). this metaphor depicts the agent as rather high in potency. interestingly, the object which the ac-tors work on was never mentioned in the (few) examples deciphering example: “reading is like deciphering a dna sequence” (137). deciphering (german “entschlüsseln” is partly similar to opening, but in contrast to the latter, it is usually an extended action and it implies a change if not in the object in the relationship between the actor and the object: after deciphering, one understands the nature of the object. the action is telic and volitional and has an individuated object. no others are present. exerting oneself example: “reading is like diving into the depths of the sea and trying to explore and discover them.
fighting through the waves […]” (106). in exerting, the actor does something, usually nonpunctually and with a strong effect. the action itself was usually not depicted as strongly voluntarily and has no desired end state (telos). others are absent. overall, transitivity is mixed – in some aspects, there is high transitivity (kinesis, temporal extension, affectedness), in others low. speaking example: “reading is a form of communication, not with spoken words, but with written words” (33). this is one of the very few conceptual metaphors that imply others – the authors with whom communication takes place. transitivity is rather high, except for the absence of an object. 7. cognitive and emotion-centered metaphors experiencing, perceiving and listening these are three distinct processes, but similar in transitivity. activity is at most medium in kine-sis, without a clearly present telos or will. the actor does not have potency, and the object is usually unaffected (though the actor may be). transitivity is as low as for seeing. dealing/grappling with something apart from its cognitive ‘flavor’, this metaphor is very similar to the solidification group: there is an actor who does something in a telic, volitional and nonpunctual manner, but without any effects on the object of action. the german origin (“sich auseinandersetzen”) is reflective, which underlines possible effects on the actor. questioning typical for cognitive metaphors, there is an agent, but no individuated object the activity has some telos and volitionality and is nonpunctual. affectedness is on the side of the actor. learning learning is an activity. again, what is affected is the actor, not the object. as in the example above, the object of learning was rarely individuated. furthermore, the examples depicted learning as something that happens rather than a telic or volitional activity. forgetting the forgetting metaphor is a very common german expression in the reading context. usually the object that is forgotten is daily life or the everyday world. although forgetting can be seen as voli-tional, it is not telic, and it is reversible. the object is not affected, and the actor is only temporarily affected. focusing transitivity is medium: some volitionality and extension in time, also some object, but no affect-edness of this object. what is characteristic of this metaphor is the necessary effort on the side of the person. dreaming dreaming is not of high kinesis, it is rather something that happens. it is usually not strongly telic and barely volitional, but extended in time. there is no individuated object, rather dreaming is con-cerned with a whole story. relaxing/enjoying/finding comfort this category, the only frequent one, has a positive and recreational emotional experience at its core. transitivity is low because this positive experience is provided and not created. the event is nonpunctual and may be sought for volitionally and purposefully. others and objects were absent from the descriptions. negative emotions in a few cases, negative experiences were explicitly associated with reading such as stepping on a lego brick or a cave where you cannot see the end. these experience were of course not intended but rather endured. vacation this is the single clear metaphor in this system; the other categories often are analogies rather than metaphors. vacation differs from traveling because it was always explained as a relaxing and rec-reational experience, similar to the relaxing category mentioned above. it was assigned a separate cate-gory because it stresses the change of state more than relaxing/enjoying/finding comfort. it was always expressed as a noun and not a verb. being someone else this is a variable category focusing on the activity of being (in) someone else while reading. while the effect of the action is on the reader and not on an object, the action is usually telic, volitional and extended and the actor has at least some potency. here, others are present, but in a very specific form, not as agents with whom one can interact but as a kind of shell into which one slips. empathizing this metaphor is very similar to the last one; we assigned it a separate category because the focus was on feeling in empathizing, that is on experiencing the feelings of the person. furthermore, there is less transitivity because the actor is depicted as passive and suffering from external influences. frontline learning research vol. 12 no. 2 (2024) 99 112 issn 2295-3159 corresponding autor: leandro de brasi, av. francisco salazar, 01145, temuco, chile, leandro.debrasi@ufrontera.cl doi: 10.14786/flr.v12i2.1247 employing the intellectual virtues to better understand argumentation interventions in education gabriel fortes1 , leandro de brasi2, michael baumtrog3 1 universidad alberto hurtado, chile 2 universidad de la frontera, chile 3 toronto metropolitan university, canada article received 7 february 2023 / article revised 21 february 2024 / accepted 5 august 2024 / available online 14 august 2024 abstract argumentation-based classroom interventions are a growing alternative for stimulating conceptual learning, thinking, and communicative skills. however, not all classroom argumentation is desired, nor does every argumentation design lead students to develop their abilities and understanding. in the educational literature, productive argumentation has been associated particularly with deliberation due to the design properties that deliberative practices demand from students, such as collaborating towards a goal, revising one’s own opinion, listening to others, and changing their minds when it is necessary to arrive at a collective decision or problem resolution. we contend that what makes deliberation productive is not argumentation in itself, but how a certain type of design scaffolds students into virtuous-like behavior, which can be the enabling condition for productive argumentation in classroom activities. through the exploration of three cases of classroom argumentation and discussion experiences, we hypothesize that virtuous-like behavior may serve as an enabling condition for each intervention. in particular, intellectually humble behaviors could be scaffolded within these interventions because all three create the proper environment for students to revise their own positions, listen carefully to others, and change their minds in light of appropriate reasoning or new evidence. employing an intellectual virtues framework, advances our understanding of how to design classroom environments for productive argumentation. this paper thus presents a novel and pioneering approach to understanding argumentation in the classroom by incorporating the concept of intellectual virtues, bridging the gap between virtues and traditional research, and offering fresh perspectives on the field. keywords: argumentation; intellectual virtues; intellectual humility; educational design mailto:leandro.debrasi@ufrontera.cl fortes, de brasi & michael baumtrog 100 | flr 1. introduction argumentation practices are growing as an alternative for promoting better epistemic subjects (that is, subjects whose epistemic performance is ameliorated) through enabling a better understanding of different (and divergent) concepts or opinions (ryu & sandoval, 2012), using evidence to support ideas (kuhn & modrek, 2022), a better calibration of one’s own opinions (leitão, 2000), and better dispositions towards dialogue with others in general (crowell & kuhn, 2014). we aim to discuss the implications of the relationship between this literature and contemporary work on intellectual virtues, understood as the study of the dispositions necessary for the flourishing of good epistemic subjects.1 argumentation interventions in education aim to understand and foster the impacts that these classroom activities have on skills like argument identification, production, and assessment, widely known as argumentative competences. improving argumentative competencies is important because they are often associated with improved overall knowledge construction and cognitive development (kuhn & udell, 2003), and most often in terms of the degree of subject-matter learning and argument structure internalization (jonassen & kim, 2010). these interventions are also seen as key to developing interactional and social skills in students (mercer, 2009), such as dealing with opposing views, successfully participating in group work, and improving efforts at rational persuasion. importing insights from work on the intellectual virtues, understood as acquired cognitive habits that avoid related deficits and excesses, could be key to informing and improving argumentation in education interventions. intellectual virtues are a good candidate to help us understand and evaluate the cognitive habits that are often said to be by-products of argumentative interaction, namely: intellectual humility, open-mindedness, curiosity, fair-mindedness, and autonomy (to name a few). porter et al., (2020; 2022), for example, show that under proper classroom conditions intellectual humility might be fostered by changing teachers’ instructional patterns, especially when it aims at enabling the mastery of achievement goals (understanding more of a topic) rather than performance goals (being recognized for knowing more or being rewarded for it). one key point in her studies is that intellectual humility growth was predictive of next year achievements for students, showing that changes in classroom design might be key to developing intellectually virtuous students. there is, however, ongoing discussion regarding the directional relationship between the cultivation of intellectual virtues and argumentative practice. this discussion has been taking place at least since siegel distinguished between the “reasons conception” of critical thinking, which highlights critical thinking and argumentation skills, and what he calls “the critical spirit”, or, “certain attitudes, dispositions, habits of mind, and character traits” (siegel 1988 p. 39) that amount to a willingness to be a better critical thinker rather than merely having the ability to do so. more recently, the debate has been addressed in terms of whether we should “argue to learn” or “learn to argue” (fortes et al., 2022). on one view, if we learn to argue, then our aim is the creation of these intellectual virtues, such as open-mindedness and humility. if we argue to learn, then these traits are prerequisites, not byproducts.2 in our view, virtues are too often expected to emerge in students as a natural byproduct of the acquisition of critical thinking skills with a deprioritization (or removal of) habituation in the ‘critical spirit’, but when students are only taught argumentative skills such as ‘spot the fallacy’ (blair 2023), there is little reason to believe that the virtues needed for arguing well will emerge. in what follows, we a make the case that rather than focus on intellectual virtues as outcomes of argumentation in education interventions (or hope for them to emerge from interventions), we should instead help develop intellectual virtues as early as possible because they are often unrecognized, though crucially important contributors to the rise of productive argumentation. in other words, we understand cultivating intellectual virtues as a theoretical pre-condition for enabling productive argumentation (as discussed further below) in the classroom3. 1 a virtue is here understood as consisting of subject-held attitudes and dispositions that “perfect” a natural human faculty, or correct for proneness to dysfunction and error in certain situations (roberts & wood, 2007, p.59). an intellectual virtue consists, roughly, of attitudes and dispositions for good and productive thinking (ritchhart, 2002, pp.18-31). intellectual virtues are typically acquired, although the virtuous subject needn’t be responsible for possessing them (battaly, 2019). 2thank you to an anonymous reviewer for pressing this point. 3 the exploration of intra-personal variables within the realm of collaborative or dialogic learning has its roots in several significant works. for example, wegerif (2010) delved into the effects of a dialogic space on children, emphasizing the need for them to momentarily set aside their judgments when confronted with others’ opinions. similarly, rapanta introduced the concept of aporia (2019), suggesting that argumentation becomes productive fortes, de brasi & michael baumtrog 101 | flr although a broad notion of the term pre-condition (also taken as enabling conditions), we use it here as it is used in the psychological and behavioral literature where it refers to creating the adequate environment for an intervention to have the desired effect(s). in this sense, we show that a confounding variable in argumentation in education interventions is that productive interventions are those considered as promoting intellectual virtues without addressing them explicitly; in other words, argumentation designs that scaffold students to behave in a virtuous-like manner seem to have a greater chance of success (even if this is not the single predictive factor of productive classroom discourse). we thus argue that most of what is commonly seen as a by-product of high-quality argumentation in education interventions could instead be seen as a prerequisite for arguing well. in this way, our discussion somewhat preempts the “argue to learn” or “learn to argue” debate in that we encourage the cultivation of intellectual virtues prior to engaging either of these approaches in earnest. as cohen (2007) puts it “a good argument is one that has been conducted virtuously” (p. 1), which implies the acquisition of these virtues before the argument(ation) occurs. we will make our argument in three steps: 1) we introduce what the intellectual virtues are and focus on intellectual humility in particular; 2) we advance our case that some intellectual virtues, such as intellectual humility, are pre-conditions for productive argumentation; and 3) we discuss how this idea could help us reinterpret and advance the research agenda for argumentation in education. 2. how intellectual virtues are related to argumentation the intellectual virtues have a long history in philosophical inquiry (fowers, et al., 2021). we are interested in the dispositions that hookway (2003, cited in lepock, 2011) identifies as the “higher level” intellectual virtues,4 which are described as “inquiry-regulating traits of intellectual character”, such as conscientiousness, perseverance, and open-mindedness. despite a long history of philosophical investigation, the development of empirical virtue epistemology research within or outside educational contexts has emerged only recently (baehr, 2013). one important insight from this research is that at least some intellectual virtues participate in the regulation of the confidence we have in our epistemic capacities and the epistemic openness of one toward others. for example, the development of a sense of virtuous intellectual humility and the ability to critically evaluate others’ opinions are considered essential for effective participation in argumentative spaces because they are engaged and critical modes of involvement with this activity. moreover, cohen (2007) points out that the pursuit of high-quality argumentation leads to outcomes such as understanding, truth-seeking, and a love of learning, thereby initiating the discussion of virtue theories in argumentation studies. importantly, aberdein (2010; 2016) establishes a taxonomy of intellectual virtues and vices related to argumentation. in his view, virtuous arguers must show a willingness to engage critically and respectfully with other arguers and maintain good intentions to solve a dispute over a topic or decision. of interest to education research on argumentation, however, is the extent to which applying virtue epistemology to argumentation studies in education might help us advance our understanding of the best conditions for argumentation in schools. among other benefits, doing so will help educators prepare their students for life outside the classroom. researchers in education believe that argumentation (and especially deliberation) promotes the type of discourse that sustains the practices required for political participation in a globalized world and that allow for managing democratic societies' challenges (rapanta et al., 2020). this includes learning to properly gauge the level of confidence in your own knowledge, weigh the different evidence against a topic of interest, and have an open mind so as to allow yourself to build knowledge through exposure to conflicting ideas. these dimensions can all be captured by intellectual virtues (kidd, 2015). in particular, intellectual humility has recently attracted when it evokes feelings of bewilderment or surprise. while both wegerif and rapanta have explored the conditions necessary for creating a productive interactional space for discussion, the concept of intellectual virtues adds another layer to this discourse. for instance, intellectual humility is distinct from the “ego-suspension" discussed in previous works. intellectual humility signifies an active ego measuring itself and others, not a passive or suspended one. this means that students should not only be aware of their positions but also critically evaluate both their own and others' stances from the outset. 4 as opposed to ‘low level’ knowledge generating faculties, such as perception, memory, or deduction. see also roberts & wood, 2007. fortes, de brasi & michael baumtrog 102 | flr the interest of some empirical research (fowers et al., 2021), especially in education (baehr, 2013) and in student argumentation (godfrey & erduran, 2021). 2.1 shared characteristics necessary to argumentation and virtuous cognitive habits while many philosophical accounts of intellectual virtue emphasize how intellectually virtuous motivations drive people toward truth-seeking and knowledge (rothenfluch, 2015), we focus on the shared cognitive processes of virtuous argumentation and cognitive habits conducive to good epistemic outcomes. these processes are taken to be indissociable from argumentative interaction (leitão, 2000) and would lead to knowledge gains and skill development for the participants in a discussion (kuhn & halpern, 2022). we argue that this is only ideally the case, occurring only when cognitive and discursive processes in argumentative interaction are shared with cognitive and discursive processes of intellectually virtuous-like habits. shared, here, means that the sociocognitive processes necessary for productive argumentation are closely related to what an intellectually virtuous person would exhibit in a discussion, such as metacognitive awareness, perspective taking, and understanding. in the psychological and educational literature, metacognition is seen as the core cognitive process related to argumentation discourse because, while arguing, students are asked to revise their understanding of the topic under discussion, assess the quality of their justification, and position themselves using the argumentative products of this cognitive process (leitão, 2000; kuhn & halpern, 2022). in argumentative interaction, participants must assess their knowledge when using it to justify a position, challenge the others’ standpoint(s), and revise the grounds on which they warrant their positions. however, metacognition in argumentation also seems to be central in stimulating student engagement in discussion (kuhn & modrek, 2021) rather than only as a competence for argument production. perspective-taking is another cognitive process central to argumentation, which points to two distinct but related mental operations: considering different epistemic standpoints and recognizing others as epistemic subjects. the former refers to the content of argumentative discussions, the latter to the people engaged in discussion. the content delivered in elementary educational settings is seldom presented as controversial or open to debate even though more than one perspective must be at play for an argumentative interaction to take place (baumtrog, 2018; casey, 2020, cf. larrauri pertierra, 2022). this is a key transformation involved in creating argumentative classrooms: asking students to recognize the existence of different and divergent positions over the same topic, which in turn improves the epistemic quality of the classroom environment (schwarz et al., 2011; fancout, 2022). as cohen (2019) argues, virtuous argumentation must be composed by both the expression of sound arguments (the most traditional account of argumentation in education) but also by arguers as good listeners, in the sense of listening to arguments both critically and charitably. moreover, improved argumentation in education interventions would be conducive to a greater understanding of the topics being discussed in any given classroom. ryu and sadoval (2012) show that the children they observed during argumentation interventions improved their capacity to understand different scientific topics because they became more aware of how (and why) pieces of evidence supported scientific conclusions. in contrast, children that only repeated what teachers said as true knowledge did not improve their understanding. kuhn and modrek (2022) point out that not only is thinking of the evidence needed to support a conclusion a complex mental operation that requires the coordination of differing pieces of information, but that it also shows the nuanced understanding children have of a topic, especially when contextualized evidence is being discussed. it should be noted, however, that all the above-mentioned studies refer to prompted or guided argumentation models. thus, it seems essential that for argumentation in education to be productive it must be framed in a certain way. although this framing may not itself constitute “virtuous argumentation,” it at least supports the notion that when people behave as if (or are guided to) virtuous argumentation, some improvement is observed. moreover, this is compatible with contemporary psychological and pedagogical notions of scaffolding developmental skills, such as the notions of a “zone of proximal development” and “internalization” in the socio-cultural tradition (for the discussion of critical thinking skills scaffolding, see wass et al., 2011). most of what we have discussed thus far broadly relates to intellectual virtues. the characteristics of productive argumentation we have pointed to are: acknowledging the limits of one’s and others’ knowledge, revising one’s perspective (in the sense of maintaining flexibility and changing positions), and generally, being open to a diversity of opinions. taken together, these three characteristics seem very similar to intellectual humility (ballantyne, 2021; de brasi, 2020), which leads us to ask, is intellectual humility a pre-condition for productive argumentation? fortes, de brasi & michael baumtrog 103 | flr 2.2 intellectual humility as central to productive argumentation according to sally jackson (2015), social practices of reasonableness have been moving forward through argumentation design since their invention (in the sense of innovations of intentional designs). jackson's reflections on argumentation design outline how reasoning, rhetoric, and dialogue are integral components of the construction of norms and values, taken as part of innovation enabled through design changes in argumentation settings, rules, or goals. argumentation design provides a framework for understanding how culturally shared situations guide people to create arguments, identify and respond to challenges, and reach agreement. in educational contexts, productive argumentation design (andriessen & schwarz, 2009) is a broad term used to refer to 1) when students express and challenge each other’s points of view through arguments, 2) participants of argumentative interaction show improved knowledge or skills throughout an argumentative intervention, 3) participants constructively collaborate to achieve a common goal through reasoned procedures, and 4) everyone has the chance to contribute to a discussion. moreover, productive argumentation is the result of careful design where a debatable topic is raised and students are motivated to present (and challenge) a given position through strong reasoning (de macedo et al., 2019). in this sense, productive argumentation design occurs when students are motivated to express their opinion, listen to challenges, and then through reasoned argumentation come to a fair conclusion on a given topic. through this type of interaction, students develop thinking skills and a better understanding of the topic itself (leitão, 2000; asterhan, 2013). productive argumentation is also sometimes called deliberative argumentation and thought of as a subtype of argumentative interaction used in the classroom that relies much more on cooperation than competition (felton, et al., 2022), which hase been proven to be much more productive in engaging students in democratic practices such as partisanship division reduction (mcavoy & mcavoy, 2021) and conceptual learning (asterhan, 2018). the exact extent to which cooperative models should be prioritized over competitive models in education is, however, still a debatable topic as competitive and adversarial argumentation seems to maintain an important place in society, and so schools should also enable students to navigate these modes of social reasoning. we aim to provide a general account of virtue epistemology in argumentation design, while focusing on a particular case in close proximity to argumentation in education, namely, intellectual humility. we understand intellectual humility as having a self-directed component, which is concerned primarily with the regulation of confidence we have in our own epistemic goods and capacities, and an other-directed component, which is concerned primarily with one’s epistemic openness to others so as to improve one’s own epistemic situation (krumrei-mancuso & rouse, 2016; leary, 2018; porter & schumann, 2018; tangney 2000).5 accordingly, a key element of intellectual humility concerns the accurate assessment of one’s epistemic self (a component captured in many philosophical accounts of intellectual humility; e.g., kidd, 2015; whitcomb et al., 2017). the intellectually humble person neither overnor under-estimates herself. in particular, intellectual humility reduces intellectual arrogance (without creating an underappreciation of oneself) by promoting a doubting attitude owing to the recognition of our fallibility (due to, say, biases and prejudices) and our limitations (due to, say, our finite cognitive power and time). having said that, intellectual humility also entails an other-directed component, which includes a disposition to change and make up one’s mind by taking others’ opinions into account. moreover, owning one’s limitations certainly entails recognizing one’s positional disadvantages (in time and space) in relation to others (others, located otherwise, can have knowledge about past events and regarding other places that i cannot have; see toole (2021) for a recent overview). similarly, given that we live in societies with hyper-specialized knowledge, which distribute the acquisition of knowledge differently among different people, recognizing one’s own limitations entails recognizing others’ strengths. more generally, recognizing one’s own limitations goes hand in hand with the recognition of our inevitable epistemic dependability on others. as maura priest puts it, “intellectually humble agents recognize that epistemic excellence is rarely (if ever) acquired on one’s own” (2017, p.476). so, this dimension of humility makes clear how it can help one depend epistemically 5 moreover, both the selfand other-directed components are present in folk theories of intellectual humility, particularly in that they view intellectually humble people as open-minded (samuelson et al., 2015). further, porter and schumann (2018), investigating the respect and openness of intellectually humble people to opposing views, re-ran their analyses to examine whether the self-directed or the other-directed component was driving the effects, and the general pattern of results remained the same when using one or the other. that being said, although many researchers agree that humility involves both components, there is more agreement with regard to the exact nature of the self-directed one (i.e., involving an accurate view of the self) than the other-directed one (davis & hook, 2014; reis et al., 2018). fortes, de brasi & michael baumtrog 104 | flr on others in certain circumstances. as vrinda dalmiya puts it, intellectual humility “consists in such ‘other regard’, in spite of, and, in fact because of, a realistic ‘self-regard’” (2016, p.119). it is within this second other-directed component that open-mindedness plays a crucial role in intellectual humility (cf. wright et al. 2018), since it minimally involves the capacity to detect and the disposition to make sense of and take seriously the merits of distinct cognitive standpoints (cf. baehr, 2011; kwong, 2016; riggs, 2016). thus, open-mindedness is here understood as an essential element of this other-directed component of intellectual humility.6 moreover, given that the open-minded person is disposed to give new ideas serious consideration, it is crucial that they listen widely and carefully (cf. dewey, 1986, p.136).7 that is to say, they must pay attention (and not merely remain silent when someone speaks) to those who have from slightly to drastically different viewpoints; or, as we put it, they must listen widely.8 and, to be fair to others’ viewpoints, they must put their cognitive effort toward appropriately grasping others’ perspectives, even if they do not initially seem to make much sense; or, as we put it, they must listen carefully. 3. intellectual virtues are pre-conditions for a productive argumentative interaction within empirical educational research, productive classroom discourse is characterized by collaborative decision making, collective goal setting, collective consensus, and diversity of participation, rather than competitive, individual, and homogenous argumentation (asterhan, 2018; felton et al., 2022). this type of productive argumentation has been called deliberation or collaborative argumentation (felton & crowell, 2022). although this use of the term is much looser than what argumentation theorists have been proposing as deliberation, this is an important starting point; classroom argumentation can be productive or not depending on the way it is framed. deliberation can lead to epistemically good outcomes (see, e.g., mendelberg, 2002), but whether it does so depends partly on certain structural and personal conditions holding (de brasi, 2021). in fact, if we want to avoid certain deliberative distortions, such as domination (the overall group opinion aligning toward the views of the socially privileged members; sanders, 1997), we need the deliberators to instantiate certain intellectual virtues such as humility. in particular, for deliberation to involve a back-and-forth of reasons that can allow the better (less error-prone) reasons to prevail, it must be set within an interactional communicative exchange. importantly, this process of giving and taking reasons includes responding to the reasons others have for their views and against one’s own. in this sense, deliberation is a reciprocal process, where reasons are not only introduced by the different parties but also responded to. for this to be the case, it is not only important to give voice to different viewpoints but also to listen to them. therefore, as much research concerning the dynamics of groups has shown, the question of how to produce deliberative spaces that avoid the effects of domination and other deliberative distortions and produce positive effects on arguers, is still an open problem. in this sense, it is important in education to deal with communication and information as central axes of argumentation interventions. moreover, this seems to indicate that providing a proper space for creating the adequate disposition to argue is an unspoken pre-condition of productive argumentation. if not all argumentation leads to good epistemic outcomes, when argumentation’s goals change, it is not necessarily that argumentation changes from competitive or collaborative, but that the collective disposition to engage in arguing as a means of achieving something emerges. in this sense, we argue that the predictive factor leading to productive argumentation is its pre-conditions, and one key factor is framing argumentation in a way such that students must behave in a virtuous-like way (that is, they must display the conduct 6 understanding open-mindedness as a component of intellectual humility easily explains why people often display the two together. as has been shown, intellectual humility is predictive of open-mindedness (e.g., krumreimancuso & rouse, 2016; porter & schumann, 2018; see also davis et al., 2016), which provides evidence that it is intellectual humility, as opposed to general humility, that is predictive of open-mindedness. this is not to suggest, however, that there is no other way of explaining such an association (see, e.g., spiegel, 2012). 7 there are other forms of communicative receptivity aside from this aural one, but for the sake of simplicity we here focus on listening. 8 virtuous open-mindedness does not require every novel idea to be given serious consideration, since one can have adequate reasons against the competence and/or sincerity of the other. as dewey (1986, p.136) says: “while it is hospitality to new themes, facts, ideas, questions, it is not the kind of hospitality that would be indicated by hanging out a sign: ‘come right in; there is nobody at home.’” see also baumtrog (2016). fortes, de brasi & michael baumtrog 105 | flr a virtuous person would), even if they don't actually possess the stable disposition to do so and even if it doesn’t necessarily lead to being virtuous after all. 3.1 vices in argumentative design for classroom activities as stated before, aberdein (2007; 2010; 2016) has proposed a taxonomy of different classes of argumentative virtues and vices. he argues that having the right disposition is key for a high-quality argumentative interaction (to improve the products and processes of argumentation). most research has centered on understanding the cognitive and discursive dimensions of argumentation in the classroom (kuhn & halpern, 2022). however, some research points out that classroom dispositions toward argumentation are related to goal setting (asterhan, 2018), where students with mastery goals (achieving more understanding, for example) relate to productive argumentation, in contrast to performance goals (such as, seeking recognition as a strong arguer). in our experience with teacher training in argumentation theory and argumentation pedagogy (fortes et al., 2021) teachers tell us that they think argumentation is difficult because students are uncooperative, unmotivated, and are too focused on the “teacher’s right answer” rather than on coming to a solution together. we can assume that schools are actually oriented toward vice education, especially through practices oriented towards obedience, passivity, and indifference. our interpretation of what teachers are saying is that students are much more used to vicious argumentative interactions where they are not invited to think freely and autonomously, where they feel tricked into giving answers that they think teachers expect, where servility is rewarded, and inquisitiveness is punished. although we haven’t systematically gathered this data, a common reaction observed in our experience after a teacher tries to design and enact argumentative activities in the classroom is awe and surprise (memis et al., 2022). when teachers adapt tasks to be engaging, topics to be stimulating, and interactions to be meaningful in an argumentative manner (especially in deliberative settings) they see unexpected changes, such as a change in students’ dispositions towards knowledge construction, not only “better argument production”. we interpret this change not just as a development of thinking skills, but designing proper interaction enables (or sets the preconditions) for students to present themselves as virtuous-like arguers, and more often than not they demonstrate their capacity for virtuous argumentation. in this sense, we do not think we are facing a problem of a “lack of” skills or competence, but rather a problem regarding communicating what is really expected from students in these situations and how to frame it. 4. three educational interventions where intellectual virtues are pre-conditions to a productive argumentative setting as non-exhaustive examples of intellectual virtues being used as, in our view, pre-conditions for argumentation interventions in school contexts, we will present three well-known (and highly cited) cases that have been key to establishing argumentation as a productive mode of discourse for learning subject-matter concepts and developing thinking skills. these examples are not based on an exhaustive literature review, but on the fact that they are all very well-known interventions that prompt argumentation (or dialogue, or discussion) in classrooms with the aim of fostering learning and development in different domains. the first and second both have a shared theoretical background: the first is centered on dialogue and student interaction, while the second emphasizes argumentation and teacher development. the third, on the other hand, concentrates on argumentation design and the format of higher education debates. another reason for these specific selections is that none refers directly to intellectual virtues or people’s dispositions. first, is the notorious exploratory talk exercise (mercer & dawes, 2008). although not traditionally considered argumentative discourse, exploratory talk, in contrast to cumulative and disputative talk (mercer, 2008), is the type of discourse where people engage critically and respectfully with each other. in this sense, exploratory talk is the type of interaction between peers in which they reflect upon each other’s contribution(s) before taking a stance on the matter being discussed. on the other hand, cumulative talk asks students only to accept and agree with each other without much reasoning or discussion, and disputative talk occurs when students disagree too much but each mind their own business without cooperating or reviewing one’s own position. through this description it already looks like we are talking about a continuum of dispositions for talk (and knowledge construction discussions) with two bad extremes and a middle ground that makes it productive. fortes, de brasi & michael baumtrog 106 | flr moreover, we think the exploratory talk exercise functions as a good case of framing the discussion between students because for the talk to be exploratory some conditions must be controlled. for example, one key element for achieving exploratory talk (and internalizing it as an “intra-mental” process) is the establishment of an obedience to ground rules. these rules are a set of guidelines that students must follow if the conditions for good dialogue are to be met. such rules include “listen actively”, “ask questions”, “share relevant information”, “challenge ideas”, “give reasons for challenges”, “build on previous contributions”, “encourage everyone to talk”, “ideas and opinions are to be respected”, “construct an atmosphere of trust”, “embrace shared purpose”, and “the group should seek agreement” — among others. from an intellectual virtue standpoint, we can hypothesize students are behaving virtuously due to how the interaction is being designed beforehand, such as being led to understand the limits of their own knowledge (revise and listen) and to be open-minded (consider both sides), all included in intellectual humility, as well as other virtues such as displaying courage (defy and defend) and respecting each other at all times. this is an example (of several) places where we can argue that intellectual virtue (and intellectual humility in particular) is a precondition for successful argumentation interventions that helps us better understand how the design of these cognitive habits (how to cultivate them) could be developed and internalized throughout educational interventions. second, episteme, a successful project founded in england (ruthven, et al., 2011) that designs classroom activities for promoting learning in physical science and mathematics. this project has impacted many researchers’ projects both within and outside the country where it was developed (ruthven et al., 2017). the intervention is based on changing the discursive setting in stem classrooms so that students engage more critically with the subject-matter and with each other. the episteme model is based on many ideas, such as the exploratory talk method mentioned above, but also on whole-class discussion studies, and particularly in effective teaching practices studies. ruthven and colleagues propose a three-phase intervention: exploration, codification, and consolidation. the first, exploration, is aimed at fostering the inquiry and examination of different perspectives regarding how to understand (and solve) a task in discussion. the second, codification, refers to a more normative take on mathematical-scientific concepts that students must understand to solve a problem. the third, consolidation, refers to a step where students become more autonomous while engaging with the related learning task. these authors (ruthven, et al., 2011; 2017) often emphasize that each phase requires teachers to play a different role, and that their ability to do so is central for the project to be productive. at first, teachers are working to foster imagination, perspective taking, and the understanding of a given problem. in this sense, we argue that teachers are instructed to guide students to engage with each other (and with the subject-matter) with openmindedness as well as intellectual curiosity and inquisitiveness in a productive way. we assume that this means regulating the deficits and excess of each intellectual virtue while the exploration phase is happening. for example, preventing students from passively accepting every idea (excess of open-mindedness) or encouraging them to look at the subject matter in new, but not redundant ways, rather than giving up early or dragging on unnecessarily long (finding the mean of curiosity). in the second stage, codification, teachers are asked to guide students to the correct concepts in physics and mathematics. we could identify this as truth-seeking and, importantly, true understanding of a subject-matter before proceeding to problem-solving. in the third phase, consolidation, teachers are asked to give space for students to more independently engage with the task at hand, in this case, providing space for intellectual autonomy (thinking for oneself). to us, these are all pre-conditions (not consequences) of the intervention, and it seems that teachers are being instructed to guide students (and themselves) in between excesses and deficits in each step of the intervention. last, and perhaps most directly argumentative, is the classroom adaptation of the critical debate model (fuentes, 2011) created by selma leitão (leitão, de chiaro & ortiz, 2016). although the previous examples were discourse and dialogue-based interventions, both focused on individual student gains. we thus also wanted to present the case of an intervention where collective organization (and collective gains), such as teamwork, strategic thinking, and competitive discourse (something rather disputed in the learning sciences community), were central to it. the original version of the critical debate model refers to the pragma-dialectically inspired inter-school debate tournament held in chile. the inter-school debate version already attempted to address issues now considered vices of argumentation, such as an unwillingness to listen to others and an unwillingness to change one’s own mind, both common in traditional debate models such as the parliamentary model used in tournaments (fuentes, 2011). in the critical debate model, the pragma-dialectical phases of a critical discussion were proposed as a method to regulate each step of the debate, allowing for teams to change their minds in light of new evidence, and to respect each step of the discussion as part of a procedure of good thinking. in leitão’s version for fortes, de brasi & michael baumtrog 107 | flr classrooms, she centers the role of in-group and out-group argumentation as key to fostering the skills of appreciating the limits of one’s own knowledge, providing strong arguments for both sides in a discussion, and intensively experiencing perspective taking throughout a semester (all central to intellectual humility). we argue that these changes refer to changes in the disposition to engage in argumentative interaction, not just changes in argument construction (identification, production, and evaluation). most of the results so far point to the fact that students show a greater “willingness to see the other side” and “willingness to revise one’s own opinion” than change or construct more complex argumentation schemes (ramírez, souza & leitão, 2013; de macedo, ramírez & leitão, 2019). in this sense, we think both interventions regulate the levels of engagement with content and with other points of view, making students more willing to concede a group-position, revise their strategies based on group agreements (not always on argument quality), and develop a “love to argue” kind of feeling. although this last possibility has not been scientifically investigated, it is a common refrain that those who work with this model hear as feedback. we thus think the aforementioned interventions help make the case that designed virtue-like behavior leads to productive argumentation because it scaffolds students to the correct disposition for engaging in collective reasoning. these interventions all share idea that through changes in students’ discourse we can promote the development of thinking skills and content learning because of a special properties argumentation has, such as, metacognitive engagement (kuhn 2022), knowledge revision (de chiaro & leitão, 2005), and the necessity for producing rebuttals (leitão, 2000). in our view, and in concordance with current literature, not all argumentation can promote these types of gains, especially because the argumentation is not being properly designed and mediated in classroom contexts (andriessen & schwarz, 2009). most research points to the fact that deliberative (rather than competitive) argumentation is more productive for the classroom (felton et al., 2022) because it establishes common goals, requires participation from the most diverse array of people possible, and consensusseeking — even if suboptimal— provides a good frame for deciding a possible solution (if the best solution is not feasible). again, but only when properly designed and mediated and when these conditions are met, we propose teachers use the lens of intellectual virtues to become informed of some of the central traits and dispositions required for productive argumentation. this is something already beginning to rise in goal achievement in argumentation (asterhan, 2018) and intellectual humility (godfrey & erduran, 2021) research. 4.1 argumentation pedagogies as a step towards human flourishing so far, argumentation-driven pedagogies have centered their research agendas and interventions on changing the discursive settings of classrooms without paying much attention to students’ dispositions for productive argumentation. consequently, it seems that classroom practices and curriculum materials have been too focused on content and skill development at the expense of investigating children as developing virtuous epistemic subjects (especially considering them as epistemic subjects outside the classroom context). without explicitly addressing the formation of virtuous habits, argumentation in education interventions will remain stagnant on specific curricular content learning or specific classroom interactions. moreover, teacher training should be addressing how to help students understand and think about the role of their dispositions in their learning trajectories. although the three classroom experiences discussed above stimulate different types of virtuous-like behavior, they all aim at creating an environment for the revision of one’s own opinion, listening to others’ points of view, and changing their mind in light of good evidence or reasoning. on our reading, this means that intellectual humility plays a central role in making these classroom experiences productive for learning conceptual knowledge, and developing thinking and communication skills. as joshi (2016) argues, argumentation pedagogies have the chance to model values for the next generations. in this sense, argumentation settings can provide students with truth-seeking methods, democratic values of collective problem solving, and demonstrations regarding how students can live among a diversity of opinions. however, the research agenda on argumentative-driven education has yet to reach it. 5. conclusion recognizing the limits of self-knowledge (and of one’s social group’s knowledge), understanding the same problem from multiple sides, and integrating knowledge that comes in different forms are widely accepted as intellectual virtues (baehr, 2013). in addition, rorty (1996) claims that the application of these principles to deliberative contexts should be called dialogic virtue, in that it does not refer to the individual principle of personal fortes, de brasi & michael baumtrog 108 | flr cultivation, but to the cultivation of a social practice that promotes a collective social good only possible through virtuous encounters in dialogue. in this paper, it is crucial to highlight that our work is fundamentally a theoretical exploration. we concentrate on presenting a hypothesis concerning potential confounding variables, with an emphasis on virtue-like behavior instantiated in argumentation and dialogue-driven classroom setting. though we've taken steps to address and discuss the limitations of our paper, we acknowledge the necessity of more extensive empirical investigation. a deeper understanding of individual and group-level dispositions towards argumentation is essential. it is our belief that such investigations could pave the way for a more effective promotion of epistemic goods within classroom settings. incorporating intellectual virtues is an important way of approaching argumentation. the social cultivation of these cognitive and interactional habits promote the search for truth and better knowledge construction, among other goals. cohen (2007) says that “there is more to our cognitive lives than knowing” (p.4), and this is also true for our educational system. there is more to education than knowledge accumulation. employing the intellectual virtues framework might help educators go beyond focusing on what is learned through argumentation (and how much) and to a better understanding of the types of arguers we are helping to flourish in our classrooms. as we have argued, intellectual virtues, and especially intellectual humility, play a central role in enabling the appropriate dispositions toward “productive argumentation” in the sense that they/it guides students to engage with each other by having them give their best justification for a position, but also to being open to revising and changing their point of view in light of strong challenges, with an emphasis on conceding when overpowering reasons are present, but nothing more or less. as baumtrog (2016) argues, it is to enable the willingness to be rationally persuaded. if we can think of argumentation interventions in these terms, this could lead to a renewal in educational studies where intellectual virtues are taught as a pre-condition for the development of argumentation skills in general. democratic citizenship in the 21st century may well depend on using school spaces to cultivate the habit of deliberation as a decision-making process that is transferred to participation in social and civic life (mendelberg, 2002). bringing intellectual virtues to the forefront of the discussion is important because it provides a way to initiate new generations into argumentation and deliberation more seriously (somin, 2010) and effectively. it should therefore be considered an important educational objective, but it requires effort and depends on certain conditions for its achievement. keypoints classroom argumentation is a growing alternative for promoting conceptual learning, thinking, and communicative skills. not all argumentation in the classroom is productive and depends on the design of the argumentation. virtuous-like behavior is an enabling condition for productive argumentation in the classroom. an intellectual virtues framework is employed to advance our understanding of how to design classrooms for productive argumentation. incorporating the concept of intellectual virtues offers fresh perspectives on the field of productive classroom discourse. acknowledgments this research has been partly funded by the fondecyt project #1210724 (anid) and the proyecto piloto de la unversidad alberto hurtado. we, the authors of this paper, declare that we have no conflicts of interest to disclose. fortes, de brasi & michael baumtrog 109 | flr references aberdein, andrew. 2007. virtue argumentation. in proceedings of the sixth conference of the international society for the study of argumentation (vol. 1, pp. 15-19). amsterdam: sic sat. aberdein, andrew. 2010. virtue in argument. argumentation, 24(2), 165-179. https://doi.org/10.1007/s10503009-9160-0 aberdein, andrew. 2016. the vices of argument. topoi, 35(2), 413-422. https://doi.org/10.1007/s11245-0159346-z andriessen, jerry, & baruch schwarz. 2009. argumentative design. in muller-miriza & anne-nelly perretclermont (eds) argumentation and education. springer, boston, ma. 145-174. asterhan, christa. 2013. epistemic and interpersonal dimensions of peer argumentation. affective learning together. new york, ny: routledge, advances in learning & instruction series. 251-271. asterhan, christa. 2018. exploring enablers and inhibitors of productive peer argumentation: the role of individual achievement goals and of gender. contemporary educational psychology, 54, 66-78. https://doi.org/10.1016/j.cedpsych.2018.05.002 baehr, jason. 2013. educating for intellectual virtues: from theory to practice. in ben kotzee education and the growth of knowledge: perspectives from social and virtue epistemology, 106-123. john wiley & sons, ltd, nj, usa. ballantyne, nathan. 2021. recent work on intellectual humility: a philosopher’s perspective. the journal of positive psychology, 1-21. https://doi.org/10.1080/17439760.2021.1940252 battaly, heather. 2019. a third kind of intellectual virtue: personalism. in battaly, h. (ed.), routledge handbook of virtue epistemology. london: routledge, 115-126. baumtrog, michael. d. 2016. the willingness to be rationally persuaded. in argumentation, objectivity and bias. proceedings of the 11th international conference of the ontario society for the study of argumentation (ossa), may 18–21, 2016, eds. patrick bondy and laura benacquista. windsor, on: ossa. baumtrog, michael. d. 2018. reasoning and arguing, dialectically and dialogically, among individual and multiple participants. argumentation 32, 77-98. https://doi.org/10.1007/s10503-017-9420-3 blair, j.a. 2023. teaching the fallacies. argumentation 37, 247–251. https://doi.org/10.1007/s10503-023-09604x casey, john. 2020. adversariality and argumentation. informal logic, 40(1), 77-108. https://doi.org/10.22329/il.v40i1.5969 cohen, daniel h. 2007. virtue epistemology and argumentation theory. ossa conference archive. 29. cohen, daniel h. 2019. argumentative virtues as conduits for reason’s causal efficacy: why the practice of giving reasons requires that we practice hearing reasons. topoi, 38(4), 711-718. https://doi.org/10.1007/s11245015-9364-x crowell, amanda, & deanna kuhn 2014. developing dialogic argumentation skills: a 3-year intervention study. journal of cognition and development, 15(2), 363-381. https://doi.org/10.1080/15248372.2012.725187 dalmiya, vrinda. 2016. caring to know. oxford: oxford university press. davis don, kenneth rice, stacey mcelroy, cirleen deblaere, elise choe, daryl r. van tongeren & joshua n. hook. 2015. distinguishing intellectual humility and general humility. the journal of positive psychology, 11(3), 215-224. https://doi.org/10.1080/17439760.2015.1048818 davis, don. e., & joshua hook. 2014. humility, religion, and spirituality: an endpiece. journal of psychology and theology, 42(1), 111-117. https://doi.org/10.1037/rel0000111 de brasi, leandro. 2020. argumentative deliberation and the development of intellectual humility and autonomy in the classroom. cogency 12, 1, 13-37. http://doi.org/ 10.32995/cogency.v12i1.339 de brasi, leandro. 2021. deliberation. palgrave encyclopedia of the possible. de chiaro, sylvia, & selma leitão. 2005. o papel do professor na construção discursiva da argumentação em sala de aula. psicologia: reflexão e crítica, 18, 350-357. https://doi.org/10.1590/s0102-79722005000300009 dewey, john, & richard rorty. (2008). the later works of john dewey, volume 8, 1925-1953: 1933, essays and how we think. everett, jim a., zach ingbretsen., fierry cushman, & mina cikara. 2017. deliberation erodes cooperative behavior—even towards competitive out-groups, even when using a control condition, and even when eliminating selection bias. journal of experimental social psychology, 73, 76-81. https://doi.org/10.1016/j.jesp.2017.06.014 https://doi.org/10.1016/j.cedpsych.2018.05.002 https://doi.org/10.1080/17439760.2021.1940252 https://doi.org/10.1007/s10503-023-09604-x https://doi.org/10.1007/s10503-023-09604-x https://doi.org/10.1080/15248372.2012.725187 https://doi.org/10.1080/17439760.2015.1048818 https://psycnet.apa.org/doi/10.1037/rel0000111 https://revistaschilenas.uchile.cl/handle/2250/10.32995/cogency.v12i1.339 https://doi.org/10.1590/s0102-79722005000300009 https://doi.org/10.1016/j.jesp.2017.06.014 fortes, de brasi & michael baumtrog 110 | flr fancourt, nigel, & liam guilfoyle. 2022. interdisciplinary perspective-taking within argumentation: students’ strategies across science and religious education. journal of religious education, 70(1), 1-23. https://doi.org/10.1007/s40839-021-00143-9 felton, mark, & amanda crowell. 2022. argumentation as a collaborative enterprise: a study of dialogic purpose and dialectical relevance in novice and experienced arguers. informal logic, 42(1), 171-202. felton, mark, amanda crowell, merce garcia-mila, & constanza villarroel. 2022. capturing deliberative argument: an analytic coding scheme for studying argumentative dialogue and its benefits for learning. learning, culture and social interaction, 36, 100350. https://doi.org/10.1016/j.lcsi.2019.100350 fortes, g., guzmán, v. and larrain, a., 2022. studying argumentation and education in south america: what has been advanced and what lies ahead. argumentation and advocacy, 58(3-4), pp.266-280. https://doi.org/10.1080/10511431.2022.2138174 fortes, g., larraín, a., & gómez, m. (2020). design of a teacher training program for the development of pedagogical knowledge on argumentation content. cogency, 12(2), 169-205. https://doi.org/10.32995/cogency.v12i2.365 fowers, blaine, jason carroll, nathan leonhardt, & bradford cokelet. 2021. the emerging science of virtue. perspectives on psychological science, 16(1), 118-147. https://doi.org/10.1177/1745691620924473 fuentes, claudio b. 2011. elementos para o desenho de um modelo de debate crítico na escola. in selma leitão & maria c. damianovic(eds). argumentação na escola: o conhecimento em construção. (1.ed.). são paulo: pontes editores. 225-249. godfrey, hayden, & sibel erduran. 2021. argumentation and intellectual humility: a theoretical synthesis and an empirical study about students’ warrants. research in science and technological education, 1-22. https://doi.org/10.1080/02635143.2021.2006622 jackson, sally. 2015. design thinking in argumentation theory and practice. argumentation, 29(3), 243-263. https://doi.org/10.1007/s10503-015-9353-7 jonassen, david, & bosung kim. 2010. arguing to learn and learning to argue: design justifications and guidelines. educational technology research and development, 58(4), 439-457. https://doi.org/10.1007/s11423-009-9143-8 joshi, parag. 2016. argumentation in democratic education: the crucial role of values. theory into practice, 55(4), 279-286. https://doi.org/10.1080/00405841.2016.1208066 kidd, i. j. 2015. educating for intellectual humility. in jason baehr (ed) intellectual virtues and education (pp. 54-70). routledge, new york. krumrei-mancuso, elizabeth j., & steven v. rouse. 2016. the development and validation of the comprehensive intellectual humility scale.journal of personality assessment 98.2: 209-221. https://doi.org/10.1080/00223891.2015.1068174 kuhn, deanna, & anahid s. modrek. 2021. mere exposure to dialogic framing enriches argumentive thinking. applied cognitive psychology 35.5: 1349-1355. https://doi.org/10.1002/acp.3862 kuhn, deanna, & anahid s. modrek. 2022. choose your evidence. science and education, 31(1), 21-31. https://doi.org/10.1007/s11191-021-00209-y kuhn, deanna, & mariel halpern. 2022. how might argumentation research inform discourse-based social studies education?. the social studies, 1-7. https://doi.org/10.1080/00377996.2022.2053832 kuhn, deanna, & wadyia udell. 2003. the development of argument skills. child development, 74(5), 12451260. https://doi.org/10.1111/1467-8624.00605 kuhn, deanna. 2022. metacognition matters in many ways. educational psychologist, 57(2), 73-86. https://doi.org/10.1080/00461520.2021.1988603 kwong, jack. (2016). open-mindedness as a critical virtue. topoi, 35(2), 403-411. https://doi.org/10.1007/s11245-015-9317-4 larrauri pertierra, iñaki x. 2022. adversariality in argumentation: shortcomings of minimal adversariality and a possible reconstruction. argumentation, 36, 17–34. https://doi.org/10.1007/s10503-021-09553-3 leary, mark. 2018. the psychology of intellectual humility. john templeton foundation, 3. leitão, selma. 2000. the potential of argument in knowledge building. human development, 43(6), 332-360. https://doi.org/10.1159/000022695 leitão, selma., sylvia de chiaro & maribel i. c. ortiz. 2016. el debate crítico: un recurso de construcción del conocimiento en el aula. textos de didáctica de la lengua y la literatura, (73), 26-33. lepock, christopher. 2011. unifying the intellectual virtues. philosophy and phenomenological research, 83(1), 106-128. https://doi.org/10.1111/j.1933-1592.2010.00425.x https://doi.org/10.1016/j.lcsi.2019.100350 https://doi.org/10.1080/10511431.2022.2138174 https://doi.org/10.1080/02635143.2021.2006622 https://revistaschilenas.uchile.cl/handle/2250/10.32995/cogency.v12i2.365 https://doi.org/10.1177/1745691620924473 https://doi.org/10.1080/02635143.2021.2006622 https://doi.org/10.1080/00405841.2016.1208066 https://doi.org/10.1080/00223891.2015.1068174 https://doi.org/10.1002/acp.3862 https://doi.org/10.1080/00377996.2022.2053832 https://doi.org/10.1111/1467-8624.00605 https://doi.org/10.1080/00461520.2021.1988603 https://doi.org/10.1159/000022695 fortes, de brasi & michael baumtrog 111 | flr memis, e. k., akkas, b. n. ç., sönmez, e., & öz, m. (2022). argumentation-based inquiry practices from the perspective of teachers receiving and implementing argumentation training. international journal of progressive education, 18(2), 325-340. https://doi.org/10.29329/ijpe.2022.431.21 mendelberg, tali. 2002. the deliberative citizen: theory and evidence. political decision making, deliberation and participation, 6(1), 151-193. mercer, neil, & linda dawes. 2008. the value of exploratory talk. in neil mercer and linda dawes exploring talk in school. sage publications, london, uk. 55-71 mercer, neil. 2008. three kinds of talk. thinking together resources, university of cambridge. mercer, neil. 2009. developing argumentation: lessons learned in the primary school. in nathalie muller-mirza & anne-nelly perret-clermont, argumentation and education. springer, boston, ma. 177-194 porter, tenele., diego c. molina, michell lucas, catherine oberle, & kali trzesniewski. 2022. classroom environment predicts changes in expressed intellectual humility. contemporary educational psychology, 102081. https://doi.org/10.1016/j.cedpsych.2022.102081 porter, tenele., karina schumann, diana selmeczy, & kali trzesniewski. 2020. intellectual humility predicts mastery behaviors when learning. learning and individual differences, 80, 101888. https://doi.org/10.1016/j.lindif.2020.101888 porter, tenelle., & karina schumann. 2018. intellectual humility and openness to the opposing view. self and identity, 17(2), 139-162. https://doi.org/10.1080/15298868.2017.1361861 priest, maura. 2017. intellectual humility: an interpersonal theory. ergo, an open access journal of philosophy, 4. https://doi.org/10.3998/ergo.12405314.0004.016 ramírez, nancy, dayse souza, & selma leitão. 2013. desarrollo de habilidades argumentativas en la enseñanzaaprendizaje de contenidos curriculares. cogency–journal of reasoning and argumentation, 5(2), 107134. rapanta, chrysi, maria vrikki, & maria evagorou. 2021. preparing culturally literate citizens through dialogue and argumentation: rethinking citizenship education. the curriculum journal, 32(3), 475-494. https://doi.org/10.1002/curj.95 rapanta, chrysi. 2019. bewilderment as a pragmatic ingredient of teacher-student dialogic interactions. studia paedagogica 24.4: 45-61. https://doi.org/10.5817/sp2019-4-2 reis, harry t, karisa y lee, stephanie d. o'keefe, & margaret s. clark. 2018. perceived partner responsiveness promotes intellectual humility. journal of experimental social psychology, 79, 21-33. https://doi.org/10.1016/j.jesp.2018.05.006 riggs, wayne. 2016. open-mindedness, insight, & understanding. in jason baehr (ed.) intellectual virtues in education: essays in applied virtue epistemology, london: routledge ritchhart, ron. 2002. intellectual character: what it is, why it matters, and how to get it. san francisco: josseybass. roberts, robert, & jay wood. 2007. intellectual virtues. oxford, oup, 340p. rorty, amelie oksenberg. 1996. from exasperating virtues to civic virtues. american philosophical quarterly, 33(3), 303-314. rothenfluch, sruthi. 2015. virtue epistemology and tacit cognitive processes in high-grade knowledge. philosophical explorations, 18(3), 393-405. https://doi.org/10.1080/13869795.2015.1042020 ruthven, kenneth, mercer, neil, keith taber, paula guardia, riika hofmann, sonia ilie, stefanie luthman & fran riga. 2017. a research-informed dialogic-teaching approach to early secondary school mathematics and science: the pedagogical design and field trial of the episteme intervention. research papers in education, 32(1), 18-40. https://doi.org/10.1080/02671522.2015.1129642 ruthven, kenneth, riika hofmann, christine howe, stefanie luthman, mercer, neil, & keith taber. 2011. the episteme pedagogical approach: essentials, rationales and challenges. proceedings of the british society for research into learning mathematics, 31(3), 131-136. ryu, suna, & william a. sandoval. 2012. improvements to elementary children's epistemic understanding from sustained argumentation. science education, 96(3), 488-526. https://doi.org/10.1002/sce.21006 samuelson, peter. l., & ian m. church. 2015. when cognition turns vicious: heuristics and biases in light of virtue epistemology. philosophical psychology, 28(8), 1095-1113. https://doi.org/10.1080/09515089.2014.904197 samuelson, peter. l., mathew j. jarvinen, thomas b. paulus, ian m. church, sam a. hardy, & justin l. barrett. 2015. implicit theories of intellectual virtues and vices: a focus on intellectual humility. the journal of positive psychology, 10(5), 389-406. https://doi.org/10.1080/17439760.2014.967802 schkade, d., cass r. sunstein, & reid hastie. 2010. when deliberation produces extremism. critical review, 22(2-3), 227-252. https://doi.org/10.1080/08913811.2010.508634 https://doi.org/10.29329/ijpe.2022.431.21 https://doi.org/10.1016/j.cedpsych.2022.102081 https://doi.org/10.1016/j.lindif.2020.101888 https://doi.org/10.1080/15298868.2017.1361861 https://doi.org/10.3998/ergo.12405314.0004.016 https://doi.org/10.1002/curj.95 https://doi.org/10.5817/sp2019-4-2 https://doi.org/10.1016/j.jesp.2018.05.006 https://doi.org/10.1080/13869795.2015.1042020 https://doi.org/10.1080/02671522.2015.1129642 https://doi.org/10.1002/sce.21006 https://doi.org/10.1080/09515089.2014.904197 https://doi.org/10.1080/17439760.2014.967802 https://doi.org/10.1080/08913811.2010.508634 fortes, de brasi & michael baumtrog 112 | flr schwarz, baruch, yaron schur, haim pensso, & naama tayer. 2011. perspective taking and synchronous argumentation for learning the day/night cycle. international journal of computer-supported collaborative learning, 6(1), 113-138. https://doi.org/10.1007/s11412-010-9100-x siegel, h., 2013. educating reason. routledge. somin, ilya. 2010. deliberative democracy and political ignorance. critical review, 22(2-3), 253-279. spiegel, james. s. (2012). open-mindedness and intellectual humility. theory and research in education, 10(1), 27-38. https://doi.org/10.1177/1477878512437472 tangney, june p. 2000. humility: theoretical perspectives, empirical findings and directions for future research. journal of social and clinical psychology, 19(1), 70. https://doi.org/10.1521/jscp.2000.19.1.70 toole, briana. 2021. recent work in standpoint epistemology. analysis, 81 (2) pp. 338–350, https://doi.org/10.1093/analys/anab026 wass, rob, tony harland, & alisson mercer. 2011. scaffolding critical thinking in the zone of proximal development. higher education research and development, 30(3), 317-328. https://doi.org/10.1080/07294360.2010.489237 wegerif, rupert. 2000. mind expanding: teaching for thinking and creativity in primary education: teaching for thinking and creativity in primary education. mcgraw-hill education (uk). whitcomb, dennis, heather battaly, jason baehr, & daniel howard-snyder. (2017). intellectual humility: owning our limitations. philosophy and phenomenological research, 94(3). https://doi.og/10.1111/phpr.12228 wright, jennifer. c., tomas nadelhoffer, linda thomson ross, & walter sinnott-armstrong. (2018). be it ever so humble: proposing a dual-dimension account and measurement of humility. self and identity, 17(1), 92-125. https://doi.org/10.1080/15298868.2017.1327454 https://doi.org/10.1177/1477878512437472 https://doi.org/10.1093/analys/anab026 https://doi.org/10.1080/07294360.2010.489237 https://doi.og/10.1111/phpr.12228 https://doi.org/10.1080/15298868.2017.1327454 frontline learning research vol. 11 no. 1 2023 40 56 issn 2295-3159 corresponding author: susanne schmidt, faculty 03, chair of business and economics education, johannes gutenberg-university mainz, jakob-welder-weg 9 55128 mainz, germany. susanne.schmidt@uni-mainz.de doi:https://doi.org/10.14786/flr.v11i1.885 modeling and measuring domain-specific quantitative reasoning in higher education business and economics susanne schmidt1, olga zlatkin-troitschanskaia1 & richard j. shavelson2 1johannes gutenberg-university mainz, germany 2stanford university, usa article received 8 june 2022/ revised 30 january 2023/ accepted 1 february 2023/ available online 22 march 2023 abstract quantitative reasoning is considered a crucial prerequisite for acquiring domainspecific expertise in higher education. to ascertain whether students are developing quantitative reasoning, validly assessing its development over the course of their studies is required. however, when measuring quantitative reasoning in an academic study program, it is often confounded with other skills. following a situated approach, we focus on quantitative reasoning in the domain of business and economics and define domain-specific quantitative reasoning primarily as a skill and capacity that allows for reasoned thinking regarding numbers, arithmetic operations, graph analyses, and patterns in real-world business and economics tasks, leading to problem solving. as many studies demonstrate, well-established instruments for assessing business and economics knowledge like the test of understanding college economics (tuce) and the examen general para el egreso de la licenciatura (egel) contain items that require domain-specific quantitative reasoning skills. in this study, we follow a new approach and assume that assessing business and economics knowledge offers the opportunity to extract domain-specific quantitative reasoning as the skill for handling quantitative data in domain-specific tasks. we present an approach where quantitative reasoning – embedded in existing measurements from tuce and egel tasks – will be empirically extracted. hereby, we reveal that items tapping domain-specific quantitative reasoning constitute an empirically separable factor within a confirmatory factor analysis and that this factor (domain-specific quantitative reasoning) can be validly and reliably measured using existing knowledge assessments. this novel methodological approach, which is based on obtaining information on students’ quantitative reasoning skills using existing domain-specific tests, offers a practical alternative to broad test batteries for assessing students’ learning outcomes in higher education. keywords: quantitative reasoning; confirmatory factor analysis; domain-specific learning; higher education; business and economics. mailto:susanne.schmidt@uni-mainz.de 41 | f l r 1. introduction within many study domains, quantitative reasoning is a required skill for scientific reasoning and arguing. furthermore, in study domains or subjects with a strong quantitative focus, quantitative reasoning is not only a required generic skill in terms of developing and understanding scientific arguments, but it is also necessary to understand domain-specific concepts, such as the supply-demand-function in economics or amortization plans in business. though quantitative reasoning is not necessarily explicitly taught in higher education classes, it is nonetheless part of learning and applying quantitative operations within business and economics tasks. while there are many studies and assessments that measure general quantitative reasoning (for an overview, roohr et al., 2014), there is a lack of research addressing the development of quantitative reasoning in specific domains, including the domain of business and economics. while business and economics outcome measures do not explicitly measure quantitative reasoning, they contain numerous items that demand reasoning quantitatively. the question, then, is: can quantitative reasoning be isolated from test items found on business and economics outcome measures? if it can, a separate measure of quantitative reasoning would not be necessary to track students’ development of quantitative reasoning. especially in the context of the so-called bologna reform in europe with its increasing modularization, the number of examinations and assessments in higher education has increased significantly. therefore, the research question arises as to what information on students’ quantitative reasoning ability can be obtained using the existing domain-specific tests to avoid the practically less suitable use of broad test batteries in higher education. to this end, we introduce a novel approach by isolating an assessment to measure quantitative reasoning from existing tests instead of developing a new test for this purpose. using confirmatory factor analysis (cfa), we show that a combination of questions from existing business and economics assessments provide a valid and reliable measure of quantitative reasoning. in the domain of business and economics, there are several validated and internationally established tests used to assess knowledge-related competences. for instance, there are the test of understanding college economics (tuce) (walstad et al., 2007) and the examen general para el egreso de la licenciatura (egel) (ceneval, 2011), which are standardized instruments adapted and validated in different language versions so that they can be used for assessing higher education students in different countries. the tuce, originally developed in the us, has been adapted and validated for german, japanese, korean and many more languages and higher education contexts (walstad et al., 2007; yamaoka et al., 2010). for germany, there is also a valid adaption of egel that has been used to assess knowledge in business administration (zlatkin-troitschanskaia et al., 2014). since these instruments are validated for measuring knowledge in business and economics, it is unclear whether they can also provide a reliable and valid tool for the assessment of quantitative reasoning as one subfacet of overall business and economics competence within this domain. to define the assessment design for quantitative reasoning out of the existing business and economics tests, we follow mislevy & haertel’s (2006) evidence-centered design and focus on the following three steps: (1) define the construct to be assessed (quantitative reasoning); here, within the domain of business and economics; (2) provide theoretical and empirical evidence as to whether quantitative reasoning-related test items from egel and tuce fit the construct definition of quantitative reasoning (see section 3.1); (3) collect and analyze data to investigate if test items align empirically with the construct definition; and finally, (4) draw a conclusion from (1) through (3). 42 | f l r the aim of this study, then, is to explore if we can isolate, conceptually, quantitative reasoning items on the tuce and egel and bring item-response data to bear on the claim that the subset of items actually measure quantitative reasoning. in the following, we briefly review current conceptual and empirical research on quantitative reasoning and its measurement (section 2). in section 3, we develop four hypotheses related to the overarching research question driving this study—whether a subset of the egel and tuce items can be used to reliably and validly assess the underlying and implicitly measured construct of quantitative reasoning in business and economics (in accordance with aera et al., 2014). in sections 4 and 5, we conduct conceptual and empirical analyses of the items to explore the claim that a subset of them measures quantitative reasoning. we conclude that we can empirically identify a reliable and valid subset of items that conceptually measure quantitative reasoning (and verbal reasoning) in existing business and economics knowledge tests (section 6). 2. conceptual and assessment background 2.1 quantitative reasoning as a key student learning outcome there is a growing consensus that effective learning and citizenship in the 21st century requires college graduates to be ‘quantitatively literate’, that is, to be able to think and reason quantitatively when the situation demands it (shavelson, 2008; ball, 2003; madison, 2009; nrc, 2012). universities are beginning to recognize the need for such quantitative competencies and consider them essential student learning outcomes (slos) (lusardi & wallace, 2013). for instance, in a study among the member institutions of the american association of colleges and universities (aac&u), 71% of the colleges and universities identified the acquisition of quantitative reasoning as a central aim of learning in higher education (hart research associates, 2009). similarly, the national leadership council for liberal education and america’s promise (leap) named quantitative reasoning as one of the essential slos of the new global century (aac&u, 2008). quantitative reasoning is a component of tertiary education as it is one of four key slos (the others being: writing, critical thinking, and information and technological literacy) (davidson & mckinney, 2001). quantitative reasoning is considered more than the ability to perform rough calculations. it is an essential competence and crucial prerequisite for acquiring domain-specific and generic knowledge and skills in higher education. in comparison to mathematics as a particular discipline, quantitative reasoning can be considered a generic skill and a way of thinking that requires dealing with complex, real-world, everyday challenges involving quantities and their different kinds of representations in different disciplines (davidson & mckinney, 2001). following a situated theoretical approach (shavelson, 2008), we assume that this generic skill can be manifested differently in varying domain-specific contexts. in this study, we focus on quantitative reasoning in the domain of business & economics (b&e) and define domain-specific quantitative reasoning (dsqr) primarily as a skill and capacity that allows for reasoned thinking regarding numbers, arithmetic operations, graph analyses, and patterns in real-world business and economics tasks, leading to problem-solving. quantitative reasoning is embedded in the hierarchical construct of cognitive outcomes. the hierarchical nature has five levels, which represent skills with a higher domain-specificity on lower levels and more generic skills on higher levels (shavelson & huang, 2003; figure 1). 43 | f l r figure 1. framework for cognitive outcomes (shavelson & huang, 2003, p. 14). 2.2 assessments of quantitative reasoning as a generic skill a number of instruments have been used to assess quantitative reasoning as a generic skill such as the quantitative reasoning for college science (quarcs) test (follette et al., 2017), the cla+ with the scientific and quantitative reasoning test (sqr) (zahner, 2013), the quantitative reasoning questions from the. graduate record examination (gre) and the heighten quantitative reasoning test (ets, 2016) (for an overview, roohr et al., 2014). currently, however, teaching in higher education does not focus on developing quantitative reasoning in an explicit way, or on assessing it (rocconi et al., 2013). this might be the reason why tests for quantitative reasoning are rarely used in higher education research and practice. rather the focus is more on the assessment of domain-specific competences. this said, domain-specific quantitative reasoning is seldom assessed although claimed to be an important outcome of a program of study. as if to do so requires a separate test from what is usually used, quantitative reasoning is unlikely to be assessed. however, the separate assessment of quantitative reasoning might not be necessary to measure the level and development of quantitative reasoning throughout undergraduate or graduate studies. if domainspecific knowledge tests can also provide a source for reliable and valid measurement of quantitative reasoning within a domain (o’neill & flynn, 2013), it could offer a practicable approach for higher education. 2.3 assessments of quantitative reasoning in the business and economics domain students’ knowledge and skills are commonly assessed with domain-specific competence tests in various fields of study (physics, engineering, psychology, business and economics; for an overview of research on domainspecific competences in different domains, piacc study oecd, 2013, or kokohs program, zlatkintroitschanskaia et al., 2017). following elrod (2014), we suspect that quantitative reasoning is a component embedded in domain-specific tests. therefore, separating quantitative reasoning loaded items within business 44 | f l r and economics tests seems both reasonable and feasible. furthermore, by measuring quantitative reasoning within domain-specific competence tests, teachers and students receive direct feedback on how quantitative reasoning contributes to solving domain-specific problems. indeed, in the domain of business and economics, we found no assessments of quantitative reasoning. the assessment of quantitative reasoning for indexing individual student skills or the effectiveness of curricula remains primarily a local practice (gaze et al., 2014). elrod (2014) assumes that one concern regarding quantitative reasoning assessment is the perception of quantitative reasoning as another outcome to assess aside from all the other regular tests and assessments and for which teachers or researchers may need to create a completely new assessment strategy. therefore, providing a valid measure of quantitative reasoning to be assessed within a domain-specific knowledge test, would open up new opportunities for researchers and practitioners. hereby, quantitative reasoning can be considered a ‘sub-dimension’ in existing assessment instruments. in particular, the assessment of domain-specific business and economics knowledge and understanding offers the opportunity to extract quantitative reasoning as the skill to handle quantitative data and numbers in existing test items. when assessing knowledge in business and economics, the tuce (walstad et al., 2007) and egel (ceneval, 2011) contain items where quantitative reasoning skills are necessary, as brückner and colleagues (2015b) demonstrate. although these tests deal with standardized assessments with a multiplechoice (mc) format, students must complete items with domain-specific tasks with real-life economic questions or problems. this study is based on an assessment of content knowledge in the domain of business and economics with items from the tuce and egel (section 4). we claim that items from the tuce and egel can be classified based on whether quantitative reasoning or non-quantitative verbal reasoning was required to answer a question correctly (brückner et al., 2015a for separating the tuce items into quantitative reasoning and verbal reasoning). we define verbal reasoning as an ability that allows for reasoned thinking without numbers, arithmetic operations and patterns in business and economics tasks. further, we can assume that spatial reasoning may link the two domains of quantitative reasoning and verbal reasoning. however, since there were very few tasks featuring graphs and diagrams in the two assessments considered here, it is not possible to model a third dimension with sr. therefore, the assumption of whether sr presents an empirically separable construct from quantitative reasoning cannot be verified in this study. by differentiating quantitative reasoning and verbal reasoning we assume that the way students deal with numerical or verbal content within a task influences the nature and difficulty of the solution. we suspected that some items would demand a preponderance of quantitative reasoning and some items verbal reasoning. by doing so, we focus on convergent and discriminant validity, including the quantitative reasoning’s relationship to other variables. consequently, we analyzed multiple-choice items from the tuce and egel to identify quantitative and verbal content demands. the sorting process is described in section 4. 3. research questions and hypotheses this study evaluates whether (1) it is possible to (a) identify subsets of items that conceptually measure quantitative reasoning in business and economicscontent knowledge tests, and (b) if this conceptual distinction can be empirically supported and distinguished from other achievement items of a verbal nature (verbal reasoning) (research question 1: internal construct validity). further, (2) whether the resulting scores for quantitative reasoning provide valid measures regarding external criteria for the underlying construct (research question 2: convergent and discriminant validity). we follow aera et al.’s (2014) validation criteria, with particular focus on the criterion of internal structure— the extent to which the empirical structure of the test supports the conceptual structure. more specifically, we assume that quantitative reasoning and verbal reasoning subtext scores are highly correlated but empirically separable, each with high internal consistency (hypothesis a). 45 | f l r in addition, if evidence supports a business and economics quantitative reasoning interpretation, we need to support this claim by showing that the quantitative reasoning score correlates, as expected, with a specific domain and with additional external variables.1 only few such studies have been conducted and they show a relationship between quantitative reasoning and socio-demographic factors (brückner et al., 2015b; tiffin et al., 2014). in particular, male test takers have been found to perform better on numeracy tasks than female test takers (owen, 2012; williams et al., 1992). this finding indicates that gender effects might differ between assessments of different components of the business and economics achievement construct, that is, between quantitative reasoning and verbal reasoning (yamaoka et al., 2010). in economics, higher levels of economicsrelated quantitative reasoning have been reported for male students, while higher levels of verbal reasoning have been reported for female students. in an introductory course in economics at a us university ballard and johnson (2004) found that male students had higher numeracy scores than female students. moreover, brückner et al. (2015a) found that both quantitative reasoning and verbal reasoning were higher for male than for female students. the same is evident in general for business and economics achievement scores: male students generally perform better than female students (brückner et al., 2015a; happ et al., 2018). however, the gender-specific differences were larger, on average, for quantitative reasoning scores than for verbal reasoning scores. according to these findings, we expect male students to outperform female students on a business and economics quantitative reasoning subtest (hypothesis b). moreover, migration background has been shown to impact generic skills, and quantitative reasoning in particular. in studies in europe, students with a migration background score lower, on average, on tests in numerically oriented subdomains of business and economics such as economics (zlatkin-troitschanskaia et al., 2015), and finance (förster et al., 2015). similar results have been found in the u.s. regarding ethnicity and race. for instance, bleske-rechek and browne (2014) have shown a gap between ethnic groups on both the gre vr and quantitative reasoning average scores. furthermore, white examinees’ verbal reasoning scores fall, on average, a full standard deviation above black minority examinees’ scores, and a half standard deviation higher than examinees from other underrepresented groups. consequently, we expect students with a recent migration background, that is students with a least one parent not of german origin, have lower test results, on average, in tasks with quantitative reasoning demands than students without migration background (hypothesis c). the educational background prior to higher education also influences generic skills such as quantitative reasoning. in particular, the school leaving grade (gpa) is considered a valid indicator of a person’s general academic ability (schuler et al., 1990) as well as a significant predictor of students’ academic performance in a domain. for instance, kuncel and colleagues (2010) showed in their meta-analysis that graduates’ gpa positively correlates with their performance in quantitative reasoning tasks (r=0.23) and verbal reasoning tasks (r=0.29) on the gre test. findings based on tests that assess students’ content knowledge in subdomains of business and economics such as accounting (byrne & flood, 2008; fritsch et al., 2015), finance (förster et al., 2015), and macroeconomics (zlatkin-troitschanskaia et al., 2015) indicate that a correlation of this kind between the gpa and domain-specific test results also exists in the business and economics domain. consequently, we expect there to be a positive relationship between school leaving grades (gpa) and tests that demand quantitative reasoning in business and economics (hypothesis d). furthermore, students’ previous subject-related knowledge acquired through learning processes prior to university is highly important in the acquisition of domain-specific knowledge at university (alexander & jetton, 2003; anderson, 2005; happ et al., 2018). for the domain of business and economics, subject-related knowledge can be acquired in different ways prior to university. in germany, many students acquire their higher education entrance qualification at vocational schools that offer advanced courses in business and economics indicating that students have prior content knowledge in business and economics when they enter university. passing advanced courses in business and economics at vocational schools or completing commercial vocational training is associated with a higher level of content knowledge in business and economics subdomains (for economics, brückner et al., 2015b; for accounting, fritsch et al., 2015; for finance, förster et al., 2015). however, this prior knowledge should only have a strong effect when acquiring (or predicting) domain-specific knowledge. as quantitative reasoning and verbal reasoning are generic skills, 46 | f l r there should be little correlation between prior knowledge of business and economics and performance in these dimensions if the test is an appropriate measure for quantitative reasoning and verbal reasoning. it is assumed that students who have pursued advanced courses at vocational schools or who have completed commercial vocational training do not perform significantly better than students who do not have any prior education in business and economics in tasks that explicitly demand quantitative reasoning (hypothesis e). 4. methods and study design 4.1 quantitative reasoning and verbal reasoning items in business and economics tests students’ business and economics content knowledge and understanding were measured in the project wiwikom2 (zlatkin-troitschanskaia et al., 2014). in this project, the tuce and egel were adapted to the german language and higher education context and comprehensively validated (aera et al., 2014; itc, 2005). the following analyses refer to the subtests in the areas of accounting and finance (16 items each from egel) and microeconomics (30 items from tuce), where we use all test items from these areas. each item consists of an item stem and four response options with one correct answer. using our definitions of quantitative reasoning and verbal reasoning we sorted the 30 tuce items and 32 egel items into these two categories (for quantitative reasoning and verbal reasoning example items from the microeconomics part, figure 2). based on the differentiation of whether a task contains numerical properties that can indicate students’ quantitative reasoning skills as described in shavelson et al. (2019), the following analyses assume a dichotomous differentiation between items with numerical content (quantitative reasoning) and items without numerical content (verbal reasoning). figure 2. two tuce items from the dimension microeconomics (walstad et al., 2007). following and expanding upon brückner et al. (2015a), who already classified tuce items into quantitative reasoning and verbal reasoning items, we noted whether numerical operations were contained in the case descriptions for both instruments, tuce and egel. items that contained numerical content and thus required students to apply mainly their mathematical abilities were classified as quantitative reasoning (for the quantitative reasoning item in sunshine city, one local ice cream company operates in a competitive labor market and product market. it can hire workers for $45 a day and sell ice cream cones for $1.00 each. the table below shows the relationship between the number of workers hired and the number of ice cream cones produced and sold. number of workers hired number of ice cream cones sold 4 5 6 7 8 340 400 450 490 520 as long as the company stays in business, how many workers will it hire to maximize profits or minimize losses? a. 5 b. 6 c. 7 d. 8 verbal reasoning item many u.s. interstate highways are crowded with traffic, but tolls are not collected even when the highways are crowded. which of the following is true about this no-toll policy? a. it is efficient because interstates are needed to transport goods. b. it is efficient because there is no cost of using the interstate once it is built. c. it is inefficient because each person’s use of the interstate adds to the congestion. d. it is inefficient because collecting tolls would increase government revenues, allowing other taxes to be decreased. 47 | f l r microeconomics part of tuce, following brückner et al., 2015a), while items dealing with purely verbally described definitions, concepts or conceptual systems were classified as verbal reasoning (table 1). table 1 distribution of quantitative reasoning (qr) and verbal reasoning (vr) items from the three content-domains of the tuce and egel content-domain qr vr total microeconomics 4 26 30 accounting 13 3 16 finance 11 4 15 total 28 33 61 like brückner and colleagues (2015a), we achieved full congruence among the four different raters in the classification of the items into quantitative reasoning and verbal reasoning subsets in our study presented here (an interrater agreement of cohens kappa=1.0; p=.000). therefore, in response to research question 1a, we conclude that it is conceptually feasible to classify test items in a business and economics knowledge test into the categories quantitative reasoning and verbal reasoning. based on this conclusion, we examined next whether the conceptual distinction between quantitative reasoning and verbal reasoning items can be empirically supported (rq1b). 4.2 sample data were collected in the summer semester 2015 using the abovementioned subtests from tuce and egel. the test was administered as a paper-pencil test in a booklet design (frey et al., 2009) using different sets of items from tuce and egel within the booklets. the test booklets were randomly distributed among the participants. the sample included 1,492 students from 27 universities and 13 universities of applied science throughout germany. the institutions involved are a representative sample. at these universities, all beginning students enrolled in master’s degrees in business and economics were invited to participate in this study in the context of introductory courses that all beginning students in these degree courses have to attend; i.e., at each university, all students enrolled in a business and economics master’s degree were assessed at once in the context of a compulsory introductory lecture. the survey was carried out on site by trained test leaders. to encourage the students to participate in this study, every participant received €5 as well as individual feedback on the test results. since participation was voluntary, the possibility of skewed representativeness at the student level cannot be excluded. however, the distribution of descriptive characteristics such as gender and age does not indicate any significant biases compared to overall student population in germany. the composition of the sample in terms of the predictor variables we used is presented in table 2. 48 | f l r table 2 distribution of the sample according to the predictors of domain-specific quantitative reasoning sample n=1,492, frequency (%) predictor yes no gender, male 783 (52.5) 708 (47.5) migration background 400 (26.8) 1,087 (72.9) advanced courses in business & economics 422 (28.3) 1,059 (71.0) vocational training 299 (20.0) 1,190 (79.8) the two indicators – ‘attended an advanced course in business and economics at commercial upper secondary school’ and ‘completed vocational training’–were used as a proxy for prior knowledge in business and economics and operationalized as two dummy-coded variables in the following analysis (section 5). the high school leaving grade (gpa) was used as an indicator of students’ general academic performance (mean=2.241, s.d=0.081 on a 5-point scale with 1 being the highest performance and 5 the lowest). 5. analyses and results this study evaluates whether (1) it is possible to (a) identify subsets of items that conceptually measure quantitative reasoning in business and economics content knowledge tests, and (b) if this conceptual distinction can be empirically supported and distinguished from other achievement items of a verbal nature (verbal reasoning). further, (2) whether the resulting scores for quantitative reasoning provide valid measures regarding external criteria for the underlying construct (research question 2: convergent and discriminant validity). as rq 1a was answered to the affirmative above, attention now turns to questions 1b and 2 and the corresponding hypotheses. 5.1 quantitative reasoning and verbal reasoning subtext scores are highly correlated but empirically separable, each with high internal consistency (hypothesis a) we address this hypothesis by testing the fit of data to alternative models using confirmatory factor analysis (cfa) with the statistical package, mplus 7.3 (muthén & muthén, 2012). all requirements for the structural equation model’s calculation were determined and were confirmed (bagozzi & yi, 1988). model-1 posits two factors corresponding to quantitative reasoning and verbal reasoning. model-2 posits a general reasoning factor combining quantitative reasoning and verbal reasoning. we used a maximum likelihood estimator with robust standard errors (labeled mlr in mplus) to take into account that item responses were dichotomous (0,1). due to the combination of booklet design and dichotomous items, the selection was limited to maximum likelihood estimators. in mplus, usually a weighted least squares estimator (wlsmv) is used for categorical variables. furthermore, as the booklet design inevitably causes missing data patterns that must be taken into account, the modelling options are limited to cfa using mplus’ options. we chose the mlr estimator, as this estimator enabled us to conduct a robust chisquare difference test. to examine the difference between the two cfa models, the 𝜒² difference test was performed. if the 𝜒² value is significant, the more restrictive model fits the data significantly worse than the general model (bagozzi & yi, 1988). the results are presented in table 3. 49 | f l r table 3 model fit of the calculated cfa models model χ² (df) p correction factor χ²/df rmsea aic bic 1 two-factor cfa model 2114.602 (1648) <.001 1.0130 1.28 0.014 51319.703 52312.150 2 one-factor cfa model 2167.981 (1649) <.001 1.0135 1.31 0.015 51371.791 52358.931 overall, both models showed a good fit to the data. the disattenuated correlation between quantitative reasoning and verbal reasoning in model model-1 was 0.76 (p=.000) suggesting that while high, quantitative reasoning and verbal reasoning could be interpreted separately. to decide which model fit the data better, we calculated a 𝜒² difference test for the empirical comparison of both models, including the models’ correction factors in the comparison test formula. as a calculation of the difference test with mplus 7.3 was not possible due to the applied ‘maximum likelihood’ estimator, we had to manually calculate the value. to this end, we applied the satorra-bentler scaled chi-square test statistic (satorra & bentler, 2010). for this purpose, the onefactor model-2 is defined as the constrained model, the two-factor model-1 is defined as the freely estimated model, and the 𝜒² test of the difference of the 𝜒² values of the model (column 2 in table 3) is used to conclude whether the reduction of the 𝜒² value in the freely estimated model is significant and whether it fits the data better than the constrained model. conducting the 𝜒² difference test for the mlr estimated models resulted in a scaled 𝜒² value of the differences of 280.5 with one degree of freedom (df=1). considering the 𝜒² statistic and its distribution, the 𝜒² difference test showed a p-value of 0.000 (δ𝝌𝟐=280.5; δ𝑑𝑓=1). model-1 therefore had a significantly better fit than model-2. aic and bic values can be used as additional indicators in model comparisons. a smaller value signifies a better data fit (schreiber et al., 2006). the model comparison shows that model-1 has lower aic and bic values and is therefore preferable to model-2. thus, we were able to isolate a quantitative reasoning score consistent with brückner et al. (2015a). moreover, the reliabilities for the general reasoning (one factor) and the quantitative reasoning and verbal reasoning scores (factor reliability, bagozzi & yi, 1988) for quantitative reasoning and verbal reasoning scores is acceptable with 0.70 for quantitative reasoning and 0.75 for verbal reasoning). hence, quantitative reasoning and verbal reasoning can be interpreted separately. this means, it is possible to measure business and economics quantitative reasoning from scores on a general knowledge test. thus, hypothesis a is supported. 5.2 the resulting scores for quantitative reasoning provide valid measures regarding external criteria for the underlying construct (hypotheses b-e) we then addressed research question 2: convergent and discriminant validity testing: hypotheses b to e.3 we focused on the correlation of individual quantitative reasoning scores and mean comparisons with other variables in a nomological network (section 3). we analyzed whether the pattern of correlations of other variables with quantitative reasoning is what would be expected based on previous research (reviewed above) and whether it supports our hypotheses. the following variables were used in this correlational analysis: (1) gender (0=female, 1=male; expected males to score higher than females on quantitative reasoning and vice versa on verbal reasoning), (2) migration background (0=no migration background,1=migration background; expected no migrant background to perform higher on both, especially verbal reasoning ), (3) school leaving grade (1=excellent, 2=good, 3=sufficient, 4=acceptable; expected lower numbers result in higher performance in both quantitative reasoning and verbal reasoning), 50 | f l r (4) advanced courses attended (0=advanced course in business and economics, 1=no advanced course in business and economics; expected to have no significant influence neither on quantitative reasoning nor on verbal reasoning), and (5) completion of a commercial vocational training (0=commercial vocational training, 1=no commercial vocational training; expected to have no significant influence neither on quantitative reasoning nor on verbal reasoning). to test whether these convergent or discriminant external variables supported our expectations they were regressed on the quantitative reasoning and verbal reasoning measures. to this end, model 1 from section 4 was extended by adding the 5 variables as predictors to the model. because of the booklet design, this has the advantage that student abilities are not required to be estimated explicitly as, for instance, sum scores, but are estimated within the multiple indicators multiple causes (mimic) regression model. this reduces biases because difficulty and discrimination parameters are taken into account. all calculations were again performed with mplus 7.3 (muthén & muthén, 2012). the results of the latent regression model are presented in table 4. table 4 regression of individual variables on quantitative reasoning (qr) and verbal reasoning (vr) in the chosen business & economics sub-domains (n=1,445) variable qr vr constant 0.788*** 0.377*** gender (male students) 0.091*** 0.080*** migration background -0.066*** -0.064*** school leaving grade (gpa) -0.053*** -0.057*** no advanced course in business & economics -0.013 -0.011 no commercial vocational training -0.031** -0.006 r² 0.182 0.171 note. *=p-value≤.1; **=p-value≤.05; ***=p-value≤.01 regarding hypothesis b that male students perform better than female students, (which is based on all our previous studies in the domain of business and economics, brückner et al., 2015a), the results meet our expectations for quantitative reasoning but not for verbal reasoning. thus, hypothesis b is only partially supported. hypothesis c that students with a migration background perform worse than students without a migration background is supported since students with a migration background have lower scores in both quantitative reasoning and verbal reasoning. similarly, students with better school leaving grades (with 1 high and 5 low) have a higher score in quantitative reasoning and verbal reasoning and hypothesis d on the relationship between school leaving grades and test performance in quantitative reasoning items is supported. in contrast to our expectations regarding hypothesis e, that prior knowledge in business and economics acquired in advanced courses at vocational schools or in commercial training does not necessarily lead to better performance in tasks that demand quantitative reasoning, we identified a significant impact of the gpa on quantitative reasoning. however, there are no significant effects on the verbal reasoning and no significant effects from attending advanced classes in economics on verbal reasoning and quantitative reasoning. 51 | f l r moreover, the relationship between quantitative reasoning and vocational training is less strong than the relationship with all other variables, indicating that prior knowledge could be considered a discriminating criterion for quantitative reasoning. thus, hypothesis e is only partially supported. 6. discussion and conclusion quantitative reasoning is considered an essential outcome of higher education. while often not directly measured in college, existing tests in business and economics, for example, contain enough quantitative reasoning items to estimate this capacity. this study empirically identifies subsets of items that conceptually measure in business and economics knowledge tests (see research question 1). the analysis confirms that the subset of quantitative reasoning items can be empirically distinguished from verbal reasoning items (as suggested in hypothesis a). overall, the internal construct validity of quantitative reasoning was further supported in this study, in line with previous research reported by brückner et al. (2016). in terms of convergent and discriminant validity, the analyses indicate that the resulting quantitative reasoning scores correlate, as expected, with the external criteria focused on in this paper (research question 2). more specifically, male students perform better than female students on a business and economics quantitative reasoning subtest (as suggested in hypothesis b). however, male students also outperform female students on verbal reasoning tasks. further analyses are therefore necessary to determine the underlying reasons, e.g., whether these differences become manifest due to the quantitative nature of these tasks, or whether it is rather a general domain-specific effect in the tasks that is decisive here. students with a migration background have lower scores in a business and economics quantitative reasoning subtest than students without migration backgrounds (as suggested in hypothesis c). however, this difference became evident in a verbal reasoning subset as well, and will therefore require further investigation in future research. furthermore, we have found a relationship between school leaving grades (gpa) and scores in a business and economics quantitative reasoning and verbal reasoning subtests (as suggested in hypothesis d). this finding is in line with other studies, which show correlations between generic skills like quantitative reasoning and verbal reasoning, and scores in domain-specific tasks. finally, our results indicate that students who have pursued advanced courses in business and economics do not perform significantly better than students who have no prior education in business and economics in a quantitative reasoning (or verbal reasoning) subtests (as suggested in hypothesis e). however, students who have completed commercial vocational training outperform students without vocational training on quantitative reasoning but not verbal reasoning. although this effect is small, this finding contradicts our assumption. however, numerous studies show similar weak correlations between generic skills like quantitative reasoning and domain-specific knowledge, and such relations are also plausible in context of the development of cognitive outcomes (see figure 1). to summarize, the present analyses indicate that it is possible to use a knowledge test in the domain of business and economics and identify a reliable and valid subset of items that conceptually measure quantitative reasoning. in terms of construct validity of quantitative reasoning as an indirect measure out of a domainspecific knowledge test, the findings show evidence to support these resulting scores as a valid measure for quantitative reasoning in business and economics. while it is possible to get valid a measure of quantitative reasoning with a domain-specific knowledge test if the domain deals with numeric properties or quantitative features in its contents, the number of existing test items was not equally distributed between both quantitative reasoning and verbal reasoning. furthermore, when it comes to explaining differences between students’ test performances in terms of the two factors quantitative reasoning and verbal reasoning, there might be other, e.g. more general domain-specific or taskrelated effects in play which were not controlled or discovered here. for instance, in another study on the tuce, the linguistic properties of the 60 test items used can explain up to 25% of performance without 52 | f l r considering any other attributes such as gender or prior knowledge (mehler et al., 2018). neither these nor other features of the tuce tasks, particularly when comparing quantitative reasoning and verbal reasoning tasks, have been investigated so far in the research of quantitative reasoning in a specific domain. in future studies, therefore, we should examine whether and to what extent quantitative reasoning and verbal reasoning tasks differ in for instance linguistic features. the empirical differentiability and the significance of spatial reasoning (sr) in relation to verbal reasoning and quantitative reasoning should also be researched using suitable test instruments. here, performance assessments in simulating more complex realistic scenarios show particularly interesting potential (shavelson et al., 2019). the correlation between quantitative reasoning and thinking and understanding in business and economics needs to be examined in a much more detailed and differentiated manner. future studies should assess the role of quantitative reasoning using separate quantitative reasoning tests as external criteria of quantitative reasoning to validate the factor-based quantitative reasoning test scores, their relationship to domain-specific content knowledge (e.g., final grades in a bachelor’s degree program), and their incremental predictive validity. in this context, another important step would be to examine in more detail to what extent these skills are generic or to what extent they also encompass domain-specific components, which is still a fundamental underresearched question. despite these limitations, this study supports the crucial role of quantitative reasoning in solving business and economics tasks. in terms of implications for educational practice, this skill needs more curricular and instructional attention in developing (domain-specific) expertise in higher education. teaching quantitative reasoning should be anchored more deeply in economic education to reduce the substantial deficits in corresponding skills among students as shown in other studies (e.g., brückner et al., 2016). in this context, an objective, reliable, and valid assessment of students’ quantitative reasoning development provides a necessary basis for various diagnostic and instructional purposes in and outside of higher education. understanding how students learn and implement quantitative reasoning to solve domain-specific tasks, and how these skills develop throughout their academic studies can help educational practitioners to develop (more) effective tailored instruction to promote the development of quantitative reasoning among students. this may improve students’ learning outcomes and domain-specific performance. the newly developed methodological approach presented in this paper, which is based on gaining information on students’ quantitative reasoning using existing domainspecific tests, offers a practical alternative to broad timeand resource-intensive test batteries for valid measuring students’ learning outcomes in higher education. 53 | f l r acknowledgments we would like to thank the two reviewers who provided constructive feedback and helpful guidance in the revision of this manuscript. notes 1 although there is extensive research regarding general cognitive abilities and their correlates (carroll, 1993), this and related work does not take into account domain-specificity in thought processes when it comes to quantitative reasoning. 2 anonymized is the acronym for the title ‘modeling and measuring competencies in business and economics among students and graduates by adapting and further developing existing american and mexican measuring instruments (tuce/ egel). for further information, https://www.blogs.unimainz.de/fb03-wiwi-competence-1/. 3 to further evaluate the validity of our interpretation of quantitative reasoning test scores, we examined and reported on test content and student response processes in previous studies (following aera et al., 2014). the validity criterion ‘test content’ was important in adapting tuce and egel items to the german context. content analyses in the form of curricular analyses, expert interviews, and online ratings (zlatkintroitschanskaia et al., 2014) provide support for the claim that the test also measures skills demanding quantitative reasoning. evidence from ‘think aloud’ or ‘cognitive interviews’ with students supports ‘response processes’ claims. the findings of the quantitative analyses presented here were also confirmed in think aloud interviews with the test takers (brückner & pellegrino, 2016). key points this research demonstrates how students’ quantitative reasoning (qr), considered a fundamental facet of 21st century skills, can be validly measured using existing domain-specific tests to avoid the practically less suitable use of broad test batteries in education. two well-established standardized knowledge tests from the domain of business and economics (b&e) are used to conceptually isolate quantitative reasoning embedded in domain-specific test tasks. item-response data and confirmatory factor analysis are used to see if these tasks actually measure quantitative reasoning in a valid and reliable way. https://www.blogs.uni-mainz.de/fb03-wiwi-competence-1/ https://www.blogs.uni-mainz.de/fb03-wiwi-competence-1/ 54 | f l r references aera (american education research association), apa (american psychological association), & ncme (national council on measurement in education). (2014). standards for educational and psychological testing. aera. alexander, p. a., & jetton, t. l. (2003). learning from traditional and alternative texts: new conceptualization for an information age. in a. c. graesser, m. a. gernsbacher & s. r. goldman (eds.), handbook of discourse processes (pp. 199–241). lawrence erlbaum associates. anderson, j. r. (2005). cognitive psychology and its implications (6th ed.). worth. association of american colleges and universities (aac&u). (2008). college learning for the new global century. https://secure.aacu.org/aacu/pdf/globalcentury_execsum_3.pdf bagozzi, r. p., & yi, y. (1988). on the evaluation of structural equation models. journal of the academy of marketing science, 16(1), 74–94. https://doi.org/10.1007/bf02723327 ball, d. l. (2003). mathematical proficiency for all students: toward a strategic research and development program in mathematics education. rand mathematics study panel. ballard, c. l., & johnson, m. f. (2004). basic math skills and performance in an introductory economics class. the journal of economic education, 35(1), 3–23. https://doi.org/10.3200/jece.35.1.3-23 bleske-rechek, a., & browne, k. (2014). trends in gre scores and graduate enrollments by gender and ethnicity. intelligence, 46, 25–34. https://doi.org/10.1016/j.intell.2014.05.005 brückner, s., förster, m., zlatkin-troitschanskaia, o., happ, r., walstad, w.b., yamaoka, m., & asano, t. (2015a). gender effects in assessment of economic knowledge and understanding: differences among undergraduate business and economics students in germany, japan, and the united states. peabody journal of education, 90(4), 503–518. https://doi.org/10.1080/0161956x.2015.1068079 brückner, s., förster, m., zlatkin-troitschanskaia, o., & walstad, w. b. (2015b). effects of prior economic education, native language, and gender on economic knowledge of first-year students in higher education. a comparative study between germany and the usa. studies in higher education, 40(3), 437–453. https://doi.org/10.1080/03075079.2015.1004235 brückner, s., & pellegrino, j. w. (2016). integrating the analysis of mental operations into multilevel models to validate an assessment of higher education students’ competency in business and economics. journal of educational measurement, 53(3), 293–312. https://doi.org/10.1111/jedm.12113 byrne, m., & flood, b. (2008). examining the relationship among background variables and academic performance of first year accounting students at an irish university. journal of accounting education, 26(4), 202–212. https://doi.org/10.1016/j.jaccedu.2009.02.001 carroll, j. b. (1993). human cognitive abilities: a survey of factor-analytic studies. cambridge university press. https://doi.org/10.1017/cbo9780511571312 ceneval (centro nacional de evaluación para la educación superior). (2011). examen general para el egreso de la licenciatura en administración. egel-admon. ceneval. davidson, m., & mckinney, g. g. r. (2001). quantitative reasoning: an overview. dialogue, 8, 1–5. ets (educational testing service). (2016). gre: guide to the use of scores 2016-17. https://www.ets.org/s/gre/pdf/gre_guide.pdf elrod, s. (2014). quantitative reasoning: the next "across the curriculum" movement. peer review, 16(3), 4–8. https://search.proquest.com/docview/1698319962?accountid=14632 follette, k., buxner, s., dokter, e., mccarthy, d., vezino b., brock, l., & prather, e. (2017). the quantitative reasoning for college science (quarcs) assessment 2: demographic, academic and attitudinal variables as predictors of quantitative ability. numeracy, 10(1), 1–33. https://doi.org/10.5038/19364660.10.1.5 förster, m., brückner, s., & zlatkin-troitschanskaia, o. (2015). assessing the financial knowledge of university students in germany. empirical research in vocational education and training, 7(6), 1– 20. https://doi.org/10.1186/s40461-015-0017-5 frey, a., hartig, j., & rupp, a. a. (2009). an ncme instructional module on booklet designs in large-scale assessments of student achievement: theory and practice. educational measurement: issues and practice, 28(3), 39–53. https://doi.org/10.1111/j.1745-3992.2009.00154.x https://doi.org/10.3200/jece.35.1.3-23 55 | f l r fritsch, s., berger, s, seifried, j., bouley, f., wuttke, e., schnick-vollmer, k., & schmitz, b. (2015). the impact of university teacher training on prospective teachers’ ck and pck – a comparison between austria and germany. empirical research in vocational education and training, 7(4), 1–20. https://doi.rog/10.1186/s40461-015-0014-8 gaze, e. c., montgomery, a., kilic-bahi, s., leoni, d., misener, l., & taylor, c. (2014). towards developing a quantitative literacy/reasoning assessment instrument. numeracy, 7(2), 4. https://doi.org/10.5038/1936-4660.7.2.4 happ, r., zlatkin-troitschanskaia, o., & förster, m. (2018). how prior economic education influences beginning university students’ knowledge of economics. empirical research in vocational education and training, 10(5), 1–20. https://doi.org/0.1186/s40461-018-0066-7 hart research associates. (2009). learning and assessment: trends in undergraduate education. a survey among members of the association of american colleges and universities. http://www.aacu.org/membership/documents/2009membersurvey_part1.pdf itc (international test commission). (2005). itc guidelines for translating and adapting tests. https://www.intestcom.org/files/guideline_test_adaptation.pdf kuncel, n. r., wee, s., serafin, l., & hezlett, s. a. (2010). the validity of the graduate record examination for master’s and doctoral programs: a meta-analytic investigation. educational and psychological measurement, 70(2), 340–352. https://doi.org/10.1177/0013164409344508 lusardi, a., & wallace, d. (2013). financial literacy and quantitative reasoning in the high school and college classroom. numeracy, 6(2), article 1. https://doi.org/10.5038/1936-4660.6.2.1 madison, b. l. (2009). all the more reason for qr across the curriculum. numeracy, 2(1), article 1. https://doi.org/10.5038/1936-4660.2.1.1 mehler, a., zlatkin-troitschanskaia, o., hemati, w., molerov, d., lücking, a., & schmidt, s. (2018). integrating computational linguistic analysis of multilingual learning data and educational measurement approaches to explore student learning in higher education. in o. zlatkintroitschanskaia, g. wittum & a. dengel (eds.), positive learning in the age of information (pp. 145– 193). springer vs. https://doi.org/10.1007/978-3-658-19567-0_10 mislevy, r. j., & haertel, g. d. (2006). implications of evidence-centered design for educational testing. educational measurement: issues and practice, 25(4), 6–20. https://doi.org/10.1111/j.17453992.2006.00075.x muthén, l. k., & muthén, b. o. (2012). mplus user’s guide (7th ed.). muthén & muthén. nrc (national research council). (2012). education for life and work: developing transferable knowledge and skills in the 21st century. https://nap.nationalacademies.org/catalog/13398/education-for-life-andwork-developing-transferable-knowledge-and-skills oecd (organisation for economic co-operation and development). (2013). the survey of adult skills: reader’s companion. https://www.oecd.org/skills/piaac/skills%20(vol%202)reader%20companion--v7%20ebook%20(press%20quality)-29%20oct%200213.pdf o’neill, p. b., & flynn, d. t. (2013). another curriculum requirement? quantitative reasoning in economics: some first steps. american journal of business education, 6(3), 339–346. https://doi.org/10.19030/ajbe.v6i3.7814 owen, a. l. (2012). student characteristics, behavior, and performance in economics classes. in g. m. hoyt & k. m. mcgoldrick (eds.), international handbook on teaching and learning economics (pp. 341– 350). edward elgar. rocconi, l. m., lambert, a. d., mccormick, a. c., & sarraf, s. a. (2013). making college count: an examination of quantitative reasoning activities in higher education. numeracy, 6(2), 1–20. https://doi.org/10.5038/1936-4660.6.2.10 roohr, k. c., graf, e. a., & liu, o. l. (2014). assessing quantitative literacy in higher education: an overview of existing research and assessments with recommendations for next-generation assessment. ets research report series, 2014(2), 1–26. https://doi.org/10.1002/ets2.12024 satorra, a., & bentler, p. m. (2010). ensuring positiveness of the scaled difference chi-square test statistic. psychometrika, 75(2), 24–248. https://doi.org/10.1007/s11336-009-9135-y 56 | f l r schreiber, j. b., nora, a., stage, f. c., barlow, e. a., & king, j. (2006). reporting structural equation modeling and confirmatory factor analysis results. a review. the journal of educational research, 99(6), 323–338. https://doi.org/10.3200/joer.99.6.323-338 schuler, h., funke, u., & baron-boldt, j. (1990). predictive validity of high-school grades: a meta-analysis. applied psychology: an international review, 39(1), 89–103. https://doi.org/10.1111/j.14640597.1990.tb01039.x shavelson, r. j. (2008). reflections on quantitative reasoning: an assessment perspective. in b. l. madison & l. a. steen (eds.), calculation vs. context. quantitative literacy and its implications for teacher education (pp. 27–44). mathematical association of america. shavelson, r. j., & huang, l. (2003). responding responsibly to the frenzy to assess learning in higher education. change. the magazin of higher learning, 35(1), 10–19. https://doi.org/10.1080/00091380309604739 shavelson, r. j., marino, j. p., zlatkin-troitschanskaia, o., & schmidt, s. (2019). reflections on the assessment of quantitative reasoning. in b. l. madison & l. a. steen (eds.), calculation vs. context: quantitative literacy and its implications for teacher education. mathematical association of america. tiffin, p. a., mclachlan, j. c., webster, l., & nicholson, s. (2014). comparison of the sensitivity of the ukcat and a levels to sociodemographic characteristics: a national study. bmc medical education, 14(7), 1–12. https://doi.org/10.1186/1472-6920-14-7 walstad, w. b., watts, m. w., & rebeck, k. (2007). test of understanding in college economics: examiner’s manual (4th ed.). national council on economic education. williams, m. l., waldauer, c., & duggal, v.g. (1992). gender differences in economic knowledge: an extension of the analysis. the journal of economic education, 23(3), 219–231. https://doi.org/10.1080/00220485.1992.10844756 yamaoka, m., walstad, w. b., watts, m. w., asana, t., & abe, s. (eds.). (2010). comparative studies on economic education in asia-pacific region. shumpusha. zahner, d. (2013). reliability and validity of cla +. council for aid to education. zlatkin-troitschanskaia, o., förster, m., brückner, s., & happ, r. (2014). insights from a german assessment of business and economics competence. in h. coates (ed.), higher education learning outcomes assessment (pp. 175–200). peter lang. zlatkin-troitschanskaia, o., förster, m., schmidt, s., brückner, s., & beck, k. (2015). erwerb wirtschaftswissenschaftlicher fachkompetenz im studium. eine mehrebenenanalytische betrachtung von hochschulischen und individuellen einflussfaktoren [acquisition of economic competence over the course of studies. a multilevel consideration of academic and individual determinants]. in s. blömeke & o. zlatkin-troitschanskaia (eds.), kompetenzen von studierenden (pp. 116–134). beltz juventa. https://doi.org/10.25656/01:15506 zlatkin-troitschanskaia, o., shavelson, r. j., & pant, h. a. (2017). assessment of learning outcomes in higher education – international comparisons and perspectives. in c. secolsky & b. denison (eds.), handbook on measurement, assessment and evaluation in higher education (2nd ed.). routledge. 1. introduction 2. conceptual and assessment background 2.1 quantitative reasoning as a key student learning outcome 2.2 assessments of quantitative reasoning as a generic skill 2.3 assessments of quantitative reasoning in the business and economics domain 3. research questions and hypotheses 4. methods and study design 4.1 quantitative reasoning and verbal reasoning items in business and economics tests 4.2 sample 5. analyses and results 5.1 quantitative reasoning and verbal reasoning subtext scores are highly correlated but empirically separable, each with high internal consistency (hypothesis a) 5.2 the resulting scores for quantitative reasoning provide valid measures regarding external criteria for the underlying construct (hypotheses b-e) 6. discussion and conclusion notes editorial

https://s25.postimg.cc/ai1akxrpb/flr_groot_logo.png

frontline learning research vol.6 no. 3 (2018) 1 5
issn 2295-3159

the journey to proficiency: exploring new objective methodologies to capture the process of learning and professional development

christian harteis, ellen kok, halszka jarodzka

introduction

over the last decades, educational research established different foci on learning – all of them aiming at understanding how learning takes place and how its outcomes can be improved by instruction. they can be distinguished regarding the object of observation: (a) there is research focusing input to learning processes, particular discussing content and teacher or learner characteristics influencing learning; (b) other research approaches exist that focus learning processes themselves, (c) finally there is research focusing learning outcomes. what varied over the time is the emphasis of these foci. for example, the current interest on international comparisons of educational systems (i.e. timss, isglu, pisa) addresses research with focus on learning outcomes.

within the last two decades, the focus of research on learning and instruction shifted from an emphasis on the outcomes back to the processes that underlie learning. however, in contrast to process research between 1960ies and 1980ies that investigated school-based classroom teaching (shulman, 1986), current process research spreads across the entire lifetime and across all skill levels, from kids studying a textbook for 20 minutes to professionals developing extraordinary expertise over years. how varying these fields may seem, their mutual aim is to understand the mental structures and their changes through learning including social, motivational, and emotional aspects influencing learning processes. however, for decades it was extremely challenging to capture these processes in a meaningful manner.

recent development of software methodology and hardware technology opened fascinating opportunities for educational research. on the one hand, the development of data analysis methodology, such as machine learning or data sequence comparison and string detection, is reaching a level so that it can be applied to approach new research questions within educational science. on the other hand, a variety of sensors made various measurements originating from highly specialized fundamental research (e.g., electroencephalography, cardiovascular measures, infrared eye-tracking) seemingly accessible and applicable. for instance, in the early 2000’s companies started building ‘plug & play’ eye trackers with ready-to-use analysis software that claimed to guide its users intuitively. more recently, relatively cheap wearable devices for measuring brain activity (eeg) and electrodermal activity have become available. it is the combination of decreasing prices for and sizes of sensor systems with increasing usability of operating and analysis software and the development of novel, easier to apply analysis methodologies that reduced inhibition threshold for application the area of education.

hence, many researchers started utilizing these methodologies in a wide area of educational research. however, quickly it turned out that neither was the use of the hardware as easy as the sellers claimed, nor was the analysis of the data as straightforward. all online measures of learning create a kind of data that is not comparable with traditional qualitative or quantitative empirical data. many online measures collect longitudinal data in a frequency of milliseconds and, thus, generate thousands of data-values. while researchers found many opportunities these measures offer, they also faced many challenges. these comprise a variety of problems, e.g. detecting meaningful events in high-frequency measures, combining process measures of different granularities, synchronizing measures, capturing the sequential nature of learning processes and defining reasonable frequencies for statistically analyzing skewed, multilevel data sets. what keeps happening, though, is that researchers face the same issues or similar problems on and on as they cannot get easily access to the progress made by others in their field. due to this lack of exchange, researchers often have to re-invent the wheel. on top of that, online measures of learning operate on a granularity that does not easily match with the predictions that can be made from our current theoretical models of learning and expertise development. online measures provide data on a very micro level of learning whereas theoretical models usually address the macro level of development. thus, researchers that use process measures are in need of ways to exchange thoughts, not just about methodological issues, but also how methodological choices relate to theoretical models.

hence, researchers started activities to establish more or less formal groups which aim at sharing their experiences (e.g., one of these is the earli sig 27 ‘online processes of learning’). however, also within already existing communities the interest for these measures rose, such as in the earli sig 14 ‘learning and professional development’. we argue that this exchange is crucial for meaningful and fruitful further development within educational research, in particular as the usage of such techniques is growing. just as an example, the methodology of eye tracking is extremely growing with over 700.000 google scholar hits in the past decade (~370.000 in 1998 – 2007, ~ 66.000 in 1988 – 1997, and less than 30.000 publications before 1987).

therefore, it is important to go beyond purely talking about experiences with process measures. what is needed is explorative, methodology-focused research in order to initiate negotiation within the scientific community of educational researchers on how these novel approaches of data-collection and data-analysis contribute to the communities’ state of knowledge on learning and instruction. thus far, the challenges of using process measures go by unnoticed as it is hardly possible to discuss them in traditional empirical study papers, which focus on knowledge structures and their changes through learning instead of on methodological developments. the current special issue in frontline learning research is a first step to fill this gap.

about this special issue

this special issue aims at exploring possibilities of using process data on learning in different contexts and critically discussing exactly these methods with respect to their explanatory power for learning and expertise development, and for gaining insight on its underlying processes operating within cognitive structures. these contributions present experiences in applying new methodologies and put the findings up for discussion. the goal of all contributions is to reflect the strengths and limitations of their measures and to provide a statement on how informative their data can be for researching learning. it addresses the broad readership within the earli community. the idea for this special issue resulted from two well attended and highly appreciated sig-invited symposia (sig 14 and sig 27) for the earli conference at tampere, finland in 2017. a public call for contributions provided additional contributions and we invited three discussants to reflect upon the articles. now, more than one year after this conference, this special issue provides a broad set of papers reporting and reflecting selected approaches of online measures of learning processes. we intentionally considered contributions from various fields and domains of educational research. the earli community represents educational research covering the entire life-span and applying laboratory conditions as well as conducting field research. all contributors were encouraged to discuss carefully if and how the selected methodological approaches and data can be informative not only for the context of the respective paper but for the educational research community in general. hence, we hope that everybody can find novel, interesting and fruitful information within this special issue.

the methods discussed in this special issue cover neuroscience topics, such as the possibilities of neuroscience for education (van atteveldt et al., this issue), eeg (scharinger, this issue). it also discusses different applications of eye tracking, such as machine learning to analyze eye tracking data (garcia moreno-esteva et al., this issue; harteis et al., this issue), its limitation in investigating web search processes (salmeron et al., this issue), or the challenge of combining it with musical performance (puurtinnen et al, this issue). further contributions discuss the use of physiological data, such as skin conductance (eteläpelto et al.; nokelainen et al.; both this issue) and combine it with further measures to study collaboration (hoogeboom et al., this issue), logfile analyses to study blended learning (van laer & ellen, this issue), or even prosodic analyses of conversations in classrooms (hämäläinen et al., this issue). several contributions use specifically multimodal aspects to study collaboration (hoogeboom et al., this issue), observational data for analyzing teachers’ behavior (donker et al., this issue), or self-regulated learning (järvenoja et al, this issue).

overarching themes

taken together, the contributions to this special issue discuss opportunities and limitations of process measures, their combination with each other, or their combination with conventional measure for analyzing learning processes. they reveal the following overarching challenges.

objective data

one important benefit of process measures as presented here, is that they do not rely on self-reports, and, thus, can be considered as objective data. we must keep in mind, though, that their interpretation remains subjective to the experimenter. furthermore, these methods open opportunities to gather data about unconscious regulation processes, whereas self-reports necessarily provide access only about what participants are aware of. it is important to keep in mind, however, that sometimes subjective data is more appropriate for a given research question. a clear understanding of what type of data would fit a certain research question remains the crucial challenge of utilizing online measures.

multimodal data

as several contributions to this special issue reveal, one important development is the increased use of multimodal designs. such designs aim to gather a broader understanding of real-world learning situations by using different types of data (e.g., a combination of psychophysiological data, video data and data from sociometric badges). severe challenges hereby are that the added value from each of those types of data for the research question needs to be clear, that the data is often collected at different levels (e.g., some data was collected at the team level and other data was collected at the individual level), as well as at different sampling frequencies (e.g., eye-tracking data can be measured at a level of 500 hz, skin conductance levels are measured at 4 hz, whereas team effectiveness is only measured once). synchronization of data can be challenging, both on a practical level (i.e., data files should use the same time representation) as well as on a sampling level (i.e., data from different participants need to be synchronized).

an important aspect of many of these process measures is that they are potentially personal data (i.e., they may be traced back to one specific person). this is particularly the case if several data sets are collected from one participant. in such a case, the new gdpr regulations come into play (https://eugdpr.org). we should also keep in mind that even if a process data set is not yet easily traceable back to one person, it may easily become so with the fast development of machine learning techniques in near future.

analysis

those large data-sets also call for an improvement in our techniques of analysis, as the current conventional statistics do not allow for a non-biased analysis of large numbers of variables at the same time. furthermore, data is often averaged or summed up over time, so the temporal order of data might get lost, while the process of interest could be reflected in the temporal order. data mining and machine learning techniques are one important option that apply complex algorithms based on different mathematical models than inferential statistics. they provide novel opportunities of revealing hidden patterns within huge data sets. hence, they allow for testing the predictive value of sets of variables for certain outcome measures and, thus, make it possible to quantify and statistically test which online measures predict the outcome.

at the same time, a careful theory-driven decision as to what variables should reflect the processes of interest is still critical, and the fact that large numbers of variables are available should not tempt us to simply report large numbers of analyses, and cherry-pick the interesting (significant) effects for discussion.

ecological validity

many of the authors argue that their designs allow for collecting data in the field (instead of laboratory environments), and, as such, increases ecological validity. the underlying assumption is that ecological validity, the extent to which the study approximates ‘the real world’, predicts external validity, the extent to which the study generalizes to ‘the real world’. in particular because many of the measures do not disturb natural task performance. however, studies within this special issue showcase that ecological validity does not necessarily translate into external validity (e.g., fluctuations in skin conductance under ‘natural’ conditions might also reflect body posture or movement instead of mental effort). it depends on the context and the research interest if and how fuzzy data appear acceptable or precise data are required. we can find argumentations that a compromised quality of, e.g., easy-to-use eeg apparatus is acceptable in the context of the new opportunities for neuroimaging in ‘real-world’ settings, and we can find argumentations that low data quality and the confounding effects that could occur in ecological valid environments may occlude real effects and, thus, compromise external validity.

hence, it is often useful to start from hypotheses generated by laboratory research and investigate these in real-world settings. however, the opposite could also provide insights: taking important but tentative findings from field studies and bringing these into the lab (including, where possible, the real-life complexity) to understand the mechanisms involved in more detail.

more ecologically valid testing environments are mostly useful when appropriate analysis techniques are available. the fixation-related eeg frequency band power analysis that scharinger introduced (in this special issue), for example, is an analysis technique that makes it possible to investigate multimedia learning environments using eeg, where previous eeg research on reading required the presentation of single words instead of free reading tasks.

coupling of high-level theoretical models to fine-grained data streams

common theoretical concepts of learning and development describe processes that usually last (much) longer than milliseconds. online measures, however, bear the particular quality to provide data on a very high resolution. it is important to keep the different granularity in mind when developing research questions and designing research settings. there may be theoretical frameworks that require less precision than others: understanding visual expertise and pattern recognition, on the one hand, focuses detailed phenomena of physical behavior that may remain under the surface of consciousness. this case requires quite a precise coupling of theories of vision with the data-streams. on the other hand, investigating the importance of emotions for learning may allow more fuzziness, as long as the duration of an emotional state is considered less important for learning processes than the pure occurrence of an emotion. hence, there is neither ‘the’ challenge of coupling theories with data, nor is there the ‘one fits for all’ solution. in the developing field of researching learning with process data the full breadth of opportunities can be found. what concretely is to be considered crucially depends on necessary theoretical decisions.

conclusions and outlook

all contributions have put their findings up for discussion and reflected on strengths and limitations of the measures applied in their studies. as such, this special issue provides a valuable resource for any researcher who already works with process measures or start working with process measures. based on the contributions, we derived suggestions on how to implement new methods and technologies to our applied field of educational science in a meaningful way.

choosing a method

when thinking about using a new method or technology for research, we always must as ourselves, why we need it. is it to address an otherwise not possible to address hypothesis? is it to explore thus far hidden processes? or is it rather to simply try out a fancy, new technology that was thrown at us? whichever the answer may be, we need to be clear about it. based on the experiences gathered in this special issue, we strongly advice to always consider how this new methodology or technology will help to approach the research question. moreover, we recommend to be cautious when being drawn towards new gadgets out of pure curiosity and we advise to stay always as low tech as possible and as high tech as necessary.

implementing a methodology when researchers start to use a new method, it is critical that they understand where the method comes from, what its history is (i.e., in which fields is it already successfully applied and how?). this often helps to understand why certain approaches were chosen and decisions were made. how is this technique currently used in its ‘home’ domain? how are the experiments set up? what are the analysis techniques? even though many of the ‘new’ methods are (relatively) new to educational research, often there is a large body of research available in other domains, that can inform the researchers. it is important to get to know this field to make sure that no huge mistakes are made because you do not know the field, as this might results in invalid data recording or analysis. we suggest to always begin by cooperating with an expert from the original field.

analysis

since computers and software developed similar rapidly further, it is now possible to utilize calculator power and software algorithms for completely new procedures of data analysis (e.g., big data, data-mining). since we, as educational researchers, might not have the expected background in statistics, computer science or data-science to execute those kind of procedures, it is central that we collaborate with researchers and practitioners from other fields. the combination of different types of expertise is central for progress in the use of process measures.

interpretation

an important challenge is how to incorporate these new methodologies that measure fine grained processes with our theories that make statements on more macro levels. how to derive predictions from our theories to these measures and how to make meaningful statements for our theories from these empirical findings? these questions should not only guide our choice of research methodology, but also challenge us to further develop and specify existing theories and frameworks.

drawing conclusions

to conclude, we hope that this special issue provides a starting point for more methodological papers in educational sciences, which critically discuss the application of (new) technological approaches and process measures, including validity information for that (set of) measure as well as practical advice for their use.

this field is developing rapidly, so we also tried to realize this special issue quite swift in order to contribute to the start of the scientific discourse on those issues. we are aware though, that the development just started and we are far away from fully understanding potential and limitations of these new kind of data and measurements. hence, we hope to see more methodological publications discussing new ways of capturing learning processes in the future!

acknowledgements

the guest editors would like to thank the reviewers of the special issue.

references

shulman, l. s. (1986). paradigms and research programs in the study of teaching. in m. wittrock (eds.), handbook of research on teaching (pp. 3-36). new york: macmillan.

frontline learning research vol. 12 no. 3 (2024) 69 98 issn 2295-3159 konstantin vinokic, rostocker str. 6, 60323 frankfurt am main, germany, k.vinokic@dipf.de (alternative email address: vinokic@gmx.de) doi: https://doi.org/10.14786/flr.v12i3.51421 the underlying cognitive processes of thin slices judgments on teaching quality konstantin vinokic 1, 2, lukas begrich 2, mareike kunter 2 & susanne kuger 3 1 institute for educational analysis baden-württemberg (ibbw), germany 2 dipf, germany 3 german youth institute (dji),germany article received 29 december 2023 / article revised 19 september 2024 / accepted 2 october 2024 / available online 25th october 2024 abstract thin slices ratings (i.e., ratings based on first impressions) have yielded intriguingly accurate results in various domains. among other, researcher have applied the thin slices technique to assess instructional quality, showing that teacher-student interactions can be reliably inferred by just very short snippets of classroom instruction. the accuracy of thin slices ratings is often explained by dual process theories of social cognition, whereby system 1 refers to an intuitive and fast way of processing, while system 2 denotes a more reflective and analytical way of processing. system 1 is considered the cognitive foundation of thin slices ratings. the central aim of the present study was to understand the underlying cognitive processes shaping the impression formation of thin slices raters of teaching quality. therefore, an unconventional and innovative research design was required to gain insights into the cognitive “black box” of thin slices raters by examining their verbal data. in an exploratory mixed method research design, we set up cognitive laboratories with two different rating situations. in a thin slices rating situation, participants rated instructional quality based on short classroom videos (30 seconds). participants in a long-video rating situation rated instructional quality based on longer classroom videos (10 minutes). we collected, coded and statistically analyzed participants’ verbal reports regarding their rating processes. the findings suggest that thin slices ratings evolve primarily based on typical processes of system 1 and not on those of system 2. for instance, thin slices ratings are associative and tend to be rather negative than positive. moreover, an initially formed impression tends to remain stable and is resistant to alteration. ratings of instructional quality based on longer videos rely on both cognitive systems, with system 2 possibly modifying an initial judgment. thus, our study does not only explain the cognitive processes underlying the thin slices ratings, but additionally provides valuable insights into the processes occurring in conventional rating settings. keywords: thin slices technique; instructional research; dual process theories of social cognition; first impressions; rater cognition vinokic, begrich, kunter & kuger 70 | f l r video-based analysis of teaching quality is an important source of data collection in educational research (de freitas, 2015; hannafin et al., 2010; petko et al., 2003). the production of classroom videos, the rater training programs, and the observational rating procedure of video lessons are effortful, time-consuming and costly (gargani & strong, 2014; kane & staiger, 2012; murphy & hall, 2021). thus, research might benefit from a more economical rating approach that is less resource-demanding and still produces reliable and valid data. the so-called thin slices technique (i.e., ratings based on first impressions) could be such a less effortful and more economical solution and might extend the repertoire of data collection methods for instructional research (ambady et al., 2000; murphy & hall, 2021). in a seminal work on first impressions, asch (1946) systematically examined the cognitive processes that underlie impression formation. although individuals can form accurate impressions of faces within 100 milliseconds (willis & todorov, 2006), first impressions are typically considered to develop over the course of five minutes (e.g., wood, 2014). the thin slices technique investigates the accuracy of judgments based on first impressions. a thin slice is defined as an excerpt of dynamic behavior less than five minutes long (ambady et al., 2000). studies, applying the thin slices technique in the context of teaching, have shown that it can yield accurate results in terms of reliability and validity (ambady & rosenthal, 1993; babad, 2005; begrich et al., 2021; sokolovic et al., 2021; vinokic et al., 2024). in contrast to the ample evidence that thin slices ratings work, there is far less research on why they work so well. a potential explanation for the surprising accuracy of thin slices judgments might be a universal human cognitive principle that is described by dual process theories of social cognition. dual process theories distinguish between two cognitive systems of information processing. system 1 operates fast, autonomously and intuitively, whereas system 2 functions slowly, reflectively, analytically and consciously (kahneman, 2003, 2011; gawronski et al., 2024). theoretically, it has been argued that thin slices ratings operate with cognitive processes associated with system 1 (wood, 2014), however this claim has only rarely been examined empirically (e.g., ambady, 2010). consequently, the present study examines whether participants’ cognitive processes during their thin slices ratings of instructional quality are indeed in accordance with system 1 of dual process theories. in order to gain insights into the black box of the underlying cognitive processes of thin slices ratings of teaching quality, the study pursued a rather unconventional and innovative approach. we developed cognitive laboratories with an exploratory, mixed method research design, using think aloud protocols and guided interviews. to date, nothing is known about the cognitive processes shaping the impression formation of thin slices raters assessing teaching quality. the present study aims to explore various cognitive mechanisms, including cognitive biases, strategies as well as the roles of intuition and associations in the context of rapid information processing when assessing teaching quality. additionally, the results may have implications for understanding the judgment formation of conventional raters of teaching quality. since typical cognitive processes associated with first impressions also appear to affect the judgment formation of conventional raters, the results of the present study provide valuable insights to improve the partially insufficient reliability of conventional raters (praetorius et al., 2012; white & ronfeldt, 2024). 1. theoretical background 1.1 the thin slices technique day-to-day impressions and judgments about others are often formed rapidly, consciously and intuitively based on very few information. the so-called thin slices research technique was developed vinokic, begrich, kunter & kuger 71 | f l r to test the accuracy of ratings based on scarce information. thin slices judgments rely on first impressions of observers who are not interacting with target persons (ambady et al., 2000; wood, 2014). a thin slice as a brief excerpt of dynamic information sampled from the behavioral stream less than five minutes long. the source of information might be audio only, video only or audiovisual (ambady et al., 2000). research has demonstrated thin slices judgments of naive or untrained raters to be highly accurate in terms of reliability (raters agree on their judgments) and validity (ratings of judges correlate with external criteria) (begrich et al., 2020; fowler et al., 2009; tackett et al., 2015; pretsch et al., 2013). the thin slices technique had been applied in various research fields such as social psychology (jung, 2016), personality psychology (holleran et al., 2009) or in the clinical context (rimondini et al., 2019). for example, lambert et al. (2014) found that raters who had watched 3-4-minute videos of unknown couples were able to correctly identify those persons who had cheated on their partner. visser and matthews (2005) showed that ratings of observers who had watched 30-second clips of call center operators’ nonverbal behavior could successfully predict the operators’ job performances as rated by managers and customers. borkenau et al. (2004) applied the thin slices technique to examine the accuracy of big five personality traits and intelligence. the ratings were based on 3-minute videos that showed target persons engaging in various activities. they found significant correlation between thin slices judgments and the outcomes on standardized intelligence tests as well as between thin slices judgments and personality reports of familiar persons. the thin slices technique has also been applied to educational research. ambady and rosenthal (1993) found significant correlations between students’ end-of-semester evaluation of teachers and thin slices ratings (30 seconds) of teachers’ personality by naive raters. even shorter slices (6 seconds, 15 seconds) were strongly related to the criterion variables. babad (2005) demonstrated that high school students could significantly predict differential behavior of unfamiliar teachers toward highand lowachieving students based on 10-second silent clips. in the context of early childhood education and care (ecec), trained thin slices raters could reliably and validly assess the quality of interaction between teachers and children even though the teachers were wearing facial masks (sokolovic et al., 2021; vinokic et al., 2024). begrich et al. (2017, 2020, 2021) applied the thin slices technique in order to assess instructional quality based on 30-seconds classroom video clips. they examined the accuracy of thin slices ratings given by naive, untrained raters, demonstrating that thin slices ratings correlated highly among raters (reliability), correlated substantially with the ratings of trained observers based on full lessons (convergent validity), captured distinctively three different dimensions of instructional quality (construct validity) and predicted students’ outcomes (predictive validity). overall, the body of thin slices research indicates that trained and untrained raters highly agree in their judgments, and these judgments correlate with external criteria. this raises the question of how such highly accurate judgments are formed on the basis of scarce information. what are the underlying cognitive processes of thin slices ratings responsible for this remarkable accuracy? 1.2 dual process theory to explain the accuracy of thin slices judgments in cognitive psychology, dual process theories try to explain social cognition with two systems: system 1 and system 2 (evans, 2006, 2019; kahneman, 2011; milli et al., 2021; stanovich & west, 2000). system 1 is supposed to operate automatically, quickly, emotionally and associatively, with no or little effort, and without a sense of voluntary control (kahneman, 2011). it operates unconsciously, rapidly and intuitively with high capacity (stanovich, 2009), relying on processes that are considered to be evolutionarily old and more directly tuned to ancient reproductive goals (evans, 2008; pennycook, 2017; stanovich, 2009). system 1 processes operate autonomously and holistically, do not require working memory and are thought to be domain-specific (bellini-leite, 2018; evans & stanovich, 2013; stanovich & west, 2000; nisbett et al., 2001). from an evolutionary perspective, first impressions are easily pulled to the negative end of an evaluative dimension. a negative impression may be formed on the basis of very few information and a positive impression may require a greater amount of information (ambady & skowronski, 2008; nesse, 2005). according to the smoke detector principle, the fitness vinokic, begrich, kunter & kuger 72 | f l r costs for a false alarm (i.e., an incorrect negative first impression) are lower than a missed alarm in case of danger (i.e., an incorrect positive first impression; nesse, 2005). in contrast, system 2 describes a conscious, slow, deliberate, analytical and reflective way of processing (evans, 2006; 2008). system 2 processes are considered to be evolutionary young, domaingeneral, capacity-limited and rule-based (pennycook, 2017). system 2 requires attention, is often associated with concentration and reasoning, and allocates attention to demanding, effortful activities, such as complex computation (kahneman, 2011). system 2 operates in a rule-based manner, compares objects across several attributes, makes deliberate choices between options (kahneman, 2011), and encompasses the processes of analytic intelligence (stanovich & west, 2000). the terminology system 1 and system 2 has become popular; however, different terminologies are debated. stanovich et al. (2014) and evans (2018; 2019) propose to refer to these two systems as type 1 and type 2. bellini-leite (2018) discusses various terminologies, like systems, types, clusters or modes. besides the most prominent dual process theory, namely the default-interventionists dual process theory, other accounts exist, such as the parallel dual process theory or the unisystem model. the default-interventionists dual process theory, advocated by e.g., kahnemann (2011) and evans (2006, 2019), proposes that system 1 is the default and can be overridden by system 2. system 2 might take over and is able to overrule the freewheeling associations and impulses of system 1 (kahneman, 2011). how these two systems precisely interact is still under debate. although mugg (2015) does not support the default-interventionist approach, he outlines the literature and claims that type 1 and type 2 processes are generated at different times: at first, type 1 and then, under certain conditions, type 2 processes are activated. type 2 processing intervenes on type 1 responses rather than generating an independent response on its own (mugg, 2015). overriding system 1 cognitions is more likely to occur in persons with higher analytic intelligence because they are more prone to produce responses that are epistemically and instrumentally rational (stanovich, 2009). the parallel dual process theory claims two reasoning systems operating in parallel and competing against each other, with system 1 being faster than system 2 (mugg, 2015; sloman, 1996). stanovich (2009) instead proposes a tripartite mind differentiating system 2 into the algorithmic and reflective mind and referring to system 1 as the autonomous mind. in contrast to dual process theories, one-system accounts or unisystem models postulate only one reasoning system, which operates along a continuum (de neys, 2021; keren & schul, 2009; kruglanski & gigerenzer, 2011; mugg, 2015). although impressions based on system 1 processes can be biased (for an overview see wood, 2014), the surprisingly accurate thin slices ratings are thought to rely predominantly on system 1 processes due to the speed of judgment formation and the limited involvement of cognitive resources (ambady et al., 2000; murphy & hall, 2021; wood, 2014). however, to our knowledge, there are only a few studies (ambady, 2010) that have explicitly and globally examined this assumption. this raises the question of whether the underlying cognitive processes of thin slices ratings of teaching quality feature characteristics that are typical for system 1? 1.3 instructional quality there is wide consensus that instructional quality is a key determinant of students’ achievement (decristan et al., 2016; hattie, 2023). klieme described a three-dimensional structure of instructional quality and its relevance for explaining students’ outcomes (fauth et al., 2024; klieme et al., 2001, 2009; kunter et al., 2005; trautwein et al., 2022). these so-called three basic dimensions of instructional quality are well established, especially in german-speaking countries (kuger et al., 2017; praetorius et al., 2018). the first dimension, classroom management, refers to the teacher’s ability to sustain an orderly and functioning classroom setting, which involves aspects such as classroom discipline, effective handling of disruptions and clarity of rules. further key features of classroom management are smooth transitions between tasks and effective time-on-task learning (decristan et al., 2016; doyle, 2006; kuger et al., 2017; marzano & marzano, 2003). the second basic dimension, cognitive activation, emphasizes the teacher’s ability to engage students in higher-order thinking by encouraging them to vinokic, begrich, kunter & kuger 73 | f l r actively process and reflect on the learning material, rather than passively receive information. cognitively challenging questions can create cognitive conflicts, which can lead to a deeper understanding of concepts (baumert et al., 2010; decristan et al., 2016; lipowsky et al., 2009). the third basic dimension, constructive support, refers to a teacher's supportive behavior, such as caring about individual needs, stimulating personalized learning, motivating students, and providing constructive feedback (decristan et al., 2022; kuger et al., 2017; kunter & voss, 2011; praetorius et al., 2018). instructional quality is usually assessed either by surveys taken from teachers and students or by trained external observers, applying a complex coding system (janik & seidel, 2013; kunter & baumert, 2006; pianta & hamre, 2009). however, it is doubtful whether student ratings provide accurate results due to students’ lack of didactic knowledge, their involvement in the instructional process, and confounding factors such as teacher popularity (aleamoni, 1999; clausen, 2002). however, recent empirical evidence indicates that students can validly assess instructional quality, albeit depending on the dimension (göllner et al., 2021). teachers’ self-reports have shown low agreement with other approaches of assessment as well as low predictive validity regarding student outcomes (clausen, 2002; desimone et al., 2010; wagner et al., 2016). in the light of these constraints, ratings from external observers based on classroom videos are often seen as the most valid way to receive information about instructional processes (helmke, 2014; pianta & hamre, 2009; praetorius et al., 2012). however, ratings from external observers are highly resource-demanding in terms of money and time (helmke, 2014; janik & seidel, 2013). against this backdrop of the high costs of conventional video rating procedures, begrich et al. (2017, 2020, 2021) examined whether the economical thin slices technique can yield accurate results in assessing instructional quality. begrich and colleagues (2017, 2020, 2021) found evidence that thin slices raters highly agree in their judgments and that thin slices ratings can even predict students’ shortterm learning outcomes. therefore, the thin slices technique might bare the potential to become an economic complement to the so far established approaches of data collection in instructional research. 2. research aims the thin slices technique is based on first impressions of untrained raters. it was successfully applied in various domains and proved to be very accurate in terms of reliability and validity (e.g., lambert et al., 2014; visser & matthews, 2005). moreover, thin slices ratings appear to deliver sound results in the context of instructional research (begrich et al., 2017, 2020, 2021). the intriguing precision of thin slices ratings seems to be counterintuitive, given the complex and contextualized nature of teaching (berliner, 2005). thus, not surprisingly, researchers have challenged the validity of thin slices ratings of instructional quality. however, arguing from the perspective of the dual process theory of social cognition (evans, 2008; stanovich et al., 2014), it is yet conceivable that first impressions of teaching quality may convey useful information. thin slices ratings presumably rely on system 1 processes, i.e., intuitive, holistic and automatized cognitive processes that seem to be able to sort and evaluate complex information rapidly, as shown in many other areas of human interaction (murphy & hall, 2021; wood, 2014). the activation of system 1 has been named as a reason for the surprisingly high accuracy of thin slices ratings (ambady et al., 2000; wood, 2014). however, this claim has only rarely been explicitly and globally examined empirically (ambady, 2010). the present study thus explores a potential link between cognitive processes underlying thin slices judgments of instructional quality and processes of system 1, as described in dual process theories of social cognition (e.g., ambady & skowronski, 2008; kahneman, 2011; stanovich 2009). in detail, the study focuses on the following two research questions (rqs). research question 1: do the reported mental processes underlying thin slices ratings of instructional quality resemble typical system 1 functioning? of particular interest is whether typical vinokic, begrich, kunter & kuger 74 | f l r system 1 characteristics can be identified in the judgmental processes reported by thin slices raters. based on the literature, we expect system 1 to serve as the underlying foundation of thin slices ratings. research question 2: are the cognitive processes underlying thin slices ratings not only similar to typical system 1 processes, but also substantially dissimilar to typical system 2 processes? beyond finding proof that system 1 processes are the foundation of thin slices ratings, a stronger proof of concept would involve discriminant evidence showing the dissimilarity between cognitive processes during thin slices ratings and typical system 2 processes. we expect to find less evidence for typical processes of system 2 in the verbal reports of thin slices raters than in the verbal reports of raters whose judgments are based on a traditional systematic video rating approach (i.e., relying on more information). 3. method 3.1 overview in the present study, a mixed method research design was pursued (mckim, 2017; schoonenboom & johnson, 2017; tashakkori & creswell, 2007) by conducting cognitive laboratories (hyytinen et al., 2014; leighton, 2017) with two different rating situations. in one rating situation, participants assessed instructional quality based on 30-second classroom videos, which is considered a typical thin slices rating situation (slices). in the other rating situation, participants assessed instructional quality based on 10-minute classroom videos (longs; see section 3.2). we obtained qualitative data through think aloud protocols and guided interviews and analyzed them by applying the coding method for qualitative research (saldaña, 2009). we combined an inductive and deductive coding approach. the coding aimed to identify patterns in the verbal data indicating underlying cognitive processes. the qualitative data were quantified by generating frequency counts of the codes, enabling us to detect differences between the two rating situations (kawulich, 2004; sandelowski et al., 2009). 3.2 design and procedure prior to the main study, we conducted trials with eight volunteers. during the trial, we tested and refined the technical procedure and the interview. simultaneously, we initiated the analysis process by trying to identify patterns in the data. due to the global covid-19 pandemic, we conducted the trial and the main study completely online in one-to-one sessions. the internet connection was good or acceptable. the study consisted of three successive phases (fig. 1). first, participants watched the videos and completed the rating questionnaire according to their rating situation (slices or longs). during this phase, each participant was videotaped (see 3.2.1). second, we confronted participants with their recording (split-screen video) and asked them to think aloud (see 3.2.2). third, a guided interview was conducted (see 3.2.2). we audio-recorded the think aloud protocols and the guided interviews, giving us two relevant sources of data: the think aloud protocols and the guided interviews. the audio recordings were transcribed according to the system of transcription developed by dresing and pehl (2012). vinokic, begrich, kunter & kuger 75 | f l r 3.2.1 the rating situations two different rating situations were implemented (fig. 1): a typical thin slices rating situation and a typical conventional rating situation. in the typical thin slices rating situation (slices), participants watched 10 classroom videos, each lasting 30 seconds, featuring 10 different teachers (fig. 1). after each video, participants completed a rating questionnaire on classroom quality. the six-point likert scale interactive questionnaire is based on begrich et al. (2017) and consisted of six items representing the three basic dimensions of instructional quality (praetorius et al., 2018). the first of these ten classroom videos served as a trial. since we intended to set up a typical thin slices situation, participants were instructed to rely on their first impressions while rating the teachers’ behavior, and they were given only 30 seconds to fill in the rating questionnaire (e.g., begrich et al., 2017, 2020, 2021). in the longs, participants watched three classroom videos of 10 minutes in length each. after each video, they completed the identical rating questionnaire (fig. 1). to approximate the rating procedure of a typical classroom observation study (hardy et al., 2011; desi-consortium, 2008), participants were instructed to think about their answers before rating the teachers’ behavior and they were given as much time as they needed. while watching the classroom videos, participants were video recorded in a split-screen set-up (recording both the participant as well as the presented classroom videos) for five minutes. this video was used as confrontational video for the think aloud protocols. 3.2.2 generating verbal data after participants had finished all teacher ratings, the think aloud protocols were generated. participants watched the split-screen confrontational video of themselves and the classroom video (fig. 1). we asked participants to think aloud by reporting how their impressions and judgments emerged and how their opinion evolved. in order not to disturb the rating process, we conducted the think aloud protocols afterward. the guided interview consisted of twelve questions (appendix a), for example: what was going on in your head while you were watching the videos and completing the rating questionnaires? how did you arrive at your judgment? 3.3 stimulusmaterial the video material was edited from the german igel-study (hardy et al., 2011). the igelstudy explored third graders knowledge development regarding floating and sinking. in a standardized science lesson, teachers were provided with prepared materials and a script. in the slices, a video snippet of ten seconds was randomly edited from each third of the lesson (see ambady et al., 2000). these three video snippets of ten seconds in length were pasted consecutively for the final thin slice. hence, the thin slice for each teacher was 30 seconds in length and consisted of three snippets of ten seconds each. however, in all sampled snippets, the teacher was required to be clearly visible. since begrich et al. vinokic, begrich, kunter & kuger 76 | f l r (2017, 2020, 2021) applied this sampling strategy and demonstrated the accuracy of thin slices ratings of teaching quality, we applied this sampling strategy in the present study as well. research showed that slices from different phases (beginning, middle, end) of the full footage correlated strongly with each other, indicating good interchangeability of the slices (hall et al., 2019). in the longs, a time stamp was generated randomly from which a 10-minute video was edited. all teachers featured in the videos were female. 3.4 rater sample and randomized assignment the rater sample consisted of 20 bachelor and master students of psychology (two males). ten raters were randomly assigned to the two groups. begrich et al. (2017) obtained with a smaller number of thin slices raters (n = 9) robust results. the age of the participants ranged from 18 to 32 years, averaging 24.8 years (sd = 4.2). participants were recruited via social media. raters were blind to the specific aim of the study and were unaware of the group they were assigned to or the existence of two groups. all participants received 20€ as a reward. 3.5 data collection the study and data collection were conducted in accordance with the general data protection regulation of the european union and the german federal data protection act (bdsg) to which participants agreed with their signature on a letter of consent. according to data protection specifications, the split-screen confrontational videos were deleted under the eye of the participants immediately after the think aloud protocols were generated (see 3.2.1). we utilized various software tools in the process of data collection and data processing: cisco webex served as the video conference tool, and vlc media player was used for presenting the videos, which were in mp4 format. participants were video-recorded with 1.8.3.0 screenpresso pro (2020), and their verbal reports were audiorecorded with 2.4.2 audacity® recording and editing software (2020). for data analysis, we worked with maxqda version 18.2.5. the statistical analyses were conducted in spss version 22 and in rstudio version 1.3.1093. 4. coding and analysis data analysis resulted in three types of information: (1) codes for text segments of think aloud protocols and guided interviews, (2) the frequency of certain terms used in the verbal material (lexical search), and (3) the amount of verbal data produced by the participants (word count). 4.1 development of codes codes were developed inductively as well as deductively. only a few codes were conceived theory-driven before data collection. most of the codes were developed data-driven (table 1). a middleorder approach between inductive and deductive coding was applied in the data coding process, meaning we coded the data while having the research aims and relevant theories of social cognition in mind (saldaña, 2009). from the various terminologies and accounts of dual process theories in the literature, we adopted the system 1 and system 2 terminology (see 1.2). the inferential character of the codes ranged from being high-inferential, requiring the coder to interpret the data, to rather low-inferential, meaning that the coder did not have to abstract or interpret the data. most of the codes were lowinferential. all codes and their modes of generation are listed in table 1 (classification of codes) and table 2 (codebook). the codes “system 1” and “system 2” refer to processes of system 1 and system 2, vinokic, begrich, kunter & kuger 77 | f l r respectively (evans, 2008; stanovich et al., 2014). the code “own school days” denotes participants’ autobiographical memories of their time at school. due to its associative character, this code serves as an indicator of system 1 (evans & stanovich, 2013; stanovich, 2009). the code “emotion” is associated with processes of system 1 (evans, 2008; kahneman, 2011). the code “halo effect” is a marker of system 1, indicating evidence for the halo effect (kahneman, 2011). the code “system 2 overrides system 1” refers to the ability of system 2 to override processes of system 1. the code “detail” is a marker of processes of system 2 because system 2 allocates attention for a detailed and specific way of processing (kahneman, 2011). we assume that the underlying cognitive processes of the code “comparison” include deliberate operations as well as the involvement of working memory (evans, 2008; evans & stanovich, 2013). further, kahneman (2011) explains that system 2 is able to compare objects on several aspects. the code “evaluation” is an indicator of system 2. we consider the underlying cognitive processes of an evaluation to be deliberative, involving the use of working memory (evans, 2008; evans & stanovich, 2013). we consider the two codes “keep s.th. in mind” and “itemsituation” as representing cognitive strategies. both are indicators for system 2 processing (stanovich, 2009). the three codes “insufficient information”, “easiness of judgment” and “difficulty of judgment” are not indicating a specific kind of cognitive processes. however, they provide information about participants’ ability to verbalize their cognitive processes while undergoing the thin slices procedure. 4.2 coder training and intercoder reliability based on the coding scheme, the codebook was developed as a manual for the coders. cues from the coder training were used to increase distinctness of the codes and comprehensibility of the coding manual. table 2 provides an excerpt of the codebook. the so-called unitization problem (campbell et al., 2013; krippendorff, 2004) was solved by segmenting the data into meaningful conceptual breaks. two student assistants were trained as coders in 12 sessions over six weeks (goodell et al., 2016). audio records from the trial were transcribed for the training of the coders. during the process of training, transcripts from the trial were scored independently by the two coders and subsequently compared and discussed. the two coders scored the data blindly in terms of not knowing to what rater group a transcript belonged. all training sessions took place via video call. table 1 classification of codes code inference cognitive system literature system 1 deductive system 1 evans (2008); stanovich et al. (2014); wood (2013) own school days inductive system 1 evans and stanovich (2013); stanovich (2009) emotion deductive system 1 evans (2008); kahneman (2011) halo effect deductive system 1 kahneman (2011) system 2 deductive system 2 evans (2008); stanovich et al. (2014); wood (2013) system 2 overrides system 1 inductive system 2 kahneman (2011); mugg (2015), stanovich (2009) detail inductive system 2 kahneman (2011) comparison inductive system 2 kahneman (2011) evaluation inductive system 2 evans (2008); evans and stanovich (2013) strategy: keep s.th. in mind inductive system 2 stanovich (2009) strategy: item situation inductive system 2 stanovich (2009) focusa: teacher inductive not defined focusa: student(s) inductive not defined focusa: situation & objects inductive not defined insufficient information inductive difficulty of judgment inductive easiness of judgment inductive note. the second column from the left (inference) denotes whether the code was generated inductively or deductively. in the fourth column (far right), relevant literature is listed. the third column indicates to which cognitive system the code is categorized. vinokic, begrich, kunter & kuger 78 | f l r to calculate the intercoder reliability, three out of 20 transcripts (15% of the data) from the main study were independently double-scored by both coders, as recommended by o’connor and joffe (2020). with kappa of .74 over all three transcripts (.71, .73., and .76 for the first, second, and third transcript, respectively), a good intercoder reliability was achieved (brennan & prediger, 1981; rädiker & kuckartz, 2019). subsequently, the two coders reached a discursive agreement on the final versions of the three transcripts they had initially scored independently. finally, one coder scored the remaining 17 transcripts. the scoring process was supervised, discussed and guided by the first author in order to ensure adherence to quality standards. table 2 codebook code definition/description examples system 1 statements of participants reflecting system 1 of dual system accounts of social cognition. impression formation is fast, intuitive, associative, emotionally, automatic and unconscious (evans, 2008; stanovich et al., 2014; wood, 2013). „deciding was hard because everything happened so quick that i could not think about it. i had to decide intuitively guided by my emotions.” “the first impression was by the guts. it was the feeling the teacher gave me.” own school days participants talk about autobiographic memories and experiences of their own school days back in the days when they were students. they mention their own teachers, instructions or classes. „… when i was a student…” “during my school days…” “my teacher in elementary school…” “my instruction in elementary school…” emotion participants report feelings, emotions or affects triggered by the video. these emotions may occur during the video or during the questionnaire. the emotion needs to be self-referenced. this does not include emotions ascribed to the students (e.g. “the students are laughing”) or the teacher. „it was very exciting watching these kids…” “the pupils were so cute…” “it was nice seeing that…” halo-effect a positive or negative aspect of the teacher influences other aspects. this also incorporates sympathy or antipathy. we also include remarks of participants when they are trying to avoid the halo-effect. „the permanent admonishments of the teacher overshadowed pretty much.” although the teacher was pretty unappealing, i tried to judge him fairly.” system 2 statements of participants reflecting system 2 of dual system accounts of social cognition. this refers to i.a. analytic, reflective, rational, deliberate, effortful, complex thinking (evans, 2008; stanovich et al., 2014; wood, 2013). “i was thinking about my decision, and my judgments were based on a pattern.” “i was thinking about it and analyzed it.” system 2 overrides system 1 at first system 1 is active and then over time system 2 comes into play. this is indicated for instance when 1.) participants state that they changed the tick in the questionnaire, 2.) when participants mention a gain in knowledge, or 3.) when participants changed their mind/opinion. „at first, i went by my gut but then the longer the video was i’ve found more and more examples, which made me change my mind.” “at first, i thought that the teacher was bored and not motivated, but then when i was thinking about it, vinokic, begrich, kunter & kuger 79 | f l r i realized that the teacher was not that bad, and i changed my mind.” “at first, i went by the guts, but then my deliberate thinking was also involved.” detail a detail is a precisely contoured unity within a larger context. this includes for example: • a concrete action/behavior (e.g. smile, gesture, statements/quotes) in a specific moment, • clothes, haircuts, pictures, writings on the blackboard, • participants mentioning that they were able to name details, • names of students „through the window, i saw an apartment block.” “in fact, i was looking out for details.” “the teacher gave a student a yellow card because he was talking.” “i think, it was very rude that the teacher made ssshhh, when the boy laughed loudly.” “there was this moment when a girl turned around and…” comparison participants compare (aspects of) teachers or videos. this includes the comparative and superlative. participants use a video as anchor in order to compare. „i always compared the teachers. i realized that the second teacher was in in comparison to the first teacher much more… “ evaluation an evaluation can be situated on a continuum from good to bad and refers to actions, persons or situations. hence, evaluations could be ranked in an order. an evaluation contains critique or praise and would potentially evoke an emotion or affect in the evaluated person. evaluative adjectives are included. “the teacher didn’t do it so well…” “she wasn’t a good teacher.” “she seemed to be very a very nice person.” “the whole group did work very well together and the atmosphere was positive.” strategy keep s.th. in mind participants trying to keep items of the questionnaire in mind while watching the video. they use the information of the questionnaire as guide while watching the video. „i tried to memorize the items and watched so the videos.” “i kept the items in mind so i knew whereon i had to look for.” strategy item-situation 1. while completing the questionnaire participants trying actively to think back to the video. 2. while watching the video participants actively tried to memorize aspects or situations of the video in order to answer the questionnaire. „i watched the video and thought what could be helpful for answering the questionnaire.” “i tried to memorize issues of the video which could be helpful in order to answer the questions” focusa teacher participants focus (commenting/talking) on the teacher. participants trying to infer from hints of teachers’ behavior to answer the questionnaire. this did also include reports about nonverbal cues like gesture, posture, facial expression, tone of voice. „the teacher interrupted one student that’s why i answered in the questionnaire …” focusa student(s) participants focus (commenting/talking) on a/the student(s). participants trying to infer from hints of students’ behavior to answer the questionnaire. „many students put up their hands. that’s why i thought that the teacher is able to involve all pupils.” focusa situation & objects participants focus (commenting/talking) on the situation or objects. participants trying to infer from hints of the situation or objects to answer the questionnaire. this includes remarks about the atmosphere. „the posters on the wall seemed to be very tidy. that is why i thought that …” vinokic, begrich, kunter & kuger 80 | f l r insufficient information participants claimed that 1. the video (field of view) did not contain relevant information about specific traits/items, 2.the video was too short. the participants mention that they would have liked to see more. they state that they did not know something or could not detect anything within the context of the video. „in the video, i could not see whether the teacher was nice.” the video was so short that i could not figure out was it was all about.” “no concrete situation was displayed about the first items…” difficulty or easiness of judgments participants mention that judgments or items were difficult or easy to answer. we also included complaints about the likert scale being too narrow. “answering the first item was pretty hard.” “judging teachers based only on so extreme short videos was very hard.” “this is a tough question.” “some items were quite easy to answer.” afocus of attention 4.3 descriptive analysis based on namey et al. (2008), code frequency lists were generated by counting the number of codes (table 3). the code frequency lists were developed with two different approaches: counting all codes1 (sum of codes across all individuals for each rater group) and counting all individuals2 (sum of individuals a code was ascribed to at least once for each rater group). to compare group means between the two rater groups, we calculated t-tests or welch’s t-tests for the sum of codes across all individuals2 or the exact test of fisher for the sum of individuals3 a code was ascribed to at least once (namey et al., 2008). moreover, for the sum of codes, the minimum of applications of a code for a person and the maximum of applications of a code for a person as well as the standard deviations are presented in table 33. 1 table 3 sum of codes: absolute (third column) 2 table 3 sum of individuals (column far right) 3 table 3 sum of codes: min.-max. (standard deviation) (fourth column) vinokic, begrich, kunter & kuger 81 | f l r 4.4 lexical search a lexical search was conducted, using the built-in lexical search function in maxqda. signal words or lexical phrases were categorized as either referring to processes of system 1 or processes of system 2 (table 4). based on the literature, the terms “quick/fast”, “intuitive” and “automatic” signal system 1 processing (evans, 2008; wood, 2014), whereas “criterion” and “reflect(ed)” are associated with processes of system 2 (evans, 2008; stanovich, 2009). additionally, more signal words were generated data-driven in an exploratory mode. the word “warmth” was interpreted as an indicator of emotional processing, denoting system 1. the word “atmosphere” is considered a holistic approach to information processing, indicating system 1. the three search terms “fidget/fidgety”, “noisy” and “disturbances” could not clearly be allocated to one of the two cognitive system. for each rater group, the number of signal words or lexical phrases was counted, and the relevant findings are listed in table 4. table 3 results of the frequency counts code condition asum of codes: absolute sum of codes: min.-max. (standard deviation) bcorrected sum of codes: relation to words csum of individuals system 1 slices 25 1-6 (1.72) 0.19 10 longs 21 1-6 (1.73) 0.11 10 own school days slices 13 0-3 (1.06) 0.10 8 longs 2* 0-1 (0.44) 0.01 2* emotion slices 31 1-7 (2.03) 0.23 10 longs 64 0-16 (5.95) 0.32 9 halo-effect slices 26 0-6 (1.65) 0.20 9 longs 19 0-5 (1.52) 0.10 9 system 2 slices 5 0-3 (0.97) 0.04 3 longs 22** 0-5 (1.48) 0.11 9* system 2 overrides system 1 slices 4 0-3 (0.97) 0.03 2 longs 19** 0-3 (0.99) 0.10 9** detail slices 7 0-4 (1.34) 0.05 3 longs 39** 1-10 (3.49) 0.20 10** comparison slices 8 0-2 (0.79) 0.06 6 longs 40* 2-9 (2.16) 0.20 10 evaluation slices 69 3-11 (2.73) 0.52 10 longs 148** 6-28 (7.33) 0.75 10 keep s.th. in mind slices 9 0-4 (1.45) 0.07 4 longs 12 0-3 (1.23) 0.06 6 item-situation slices 5 0-3 (0.97) 0.04 3 longs 13 0-3 (1.16) 0.07 7 teacher slices 36 1-7 (1.90) 0.27 9 longs 31 1-5 (1.45) 0.16 9 student slices 13 0-3 (0.95) 0.10 8 longs 15 0-5 (1.43) 0.08 8 situation & objects slices 9 0-3 (0.99) 0.07 6 longs 7 0-2 (0.68) 0.04 6 insufficient information slices 30 0-8 (2.63) 0.23 9 longs 6* 0-4 (1.27) 0.03 3* difficulty of judgment slices 35 1-7 (1.90) 0.26 10 longs 29 1-6 (1.45) 0.15 10 easiness of judgment slices 8 0-2 (0.63) 0.06 7 longs 8 0-2 (0.79) 0.04 6 *test comparing the two rater groups is significant at the 0.05 level (two-tailed). **test comparing the two rater groups is significant at the 0.01 level (two-tailed). a t-test or welch’s t-test (sum of codes across all individuals for each rater group and standard deviation) b number of codes divided by number of words in total and multiplied by 100 c exact test of fisher (sum of individuals a code was ascribed to at least once for each rater group) vinokic, begrich, kunter & kuger 82 | f l r 4.5 word count and relative occurrence of codes an overall word count of participants’ answers was conducted to examine in which rater group more words were uttered—whether participants in the slices uttered more (or fewer) words than participants in the longs. hence, the number of words for each participant in both rating situation was summed up. in addition, we wanted to prevent any distorting effects due to possible differences in the amount of words per answer. therefore, the frequency of codes was related to the total number of words for both rater groups. we divided the number of codes in total by the number of words in total for both rating situations and multiplied it by 100 (table 3)4. all irrelevant communication content, such as interview questions, researcher comments, and off-topic remarks, were removed from the data, leaving only the participants’ responses for analysis. 4.6 further analysis for some codes, the scored text segments were analyzed in depth to uncover further hidden underlying patterns in the data. we did this for the code “evaluation” by examining whether the evaluations were positive, neutral, or negative. further, we analyzed the text segments of the code “system 2 overrides system 1” of participants in the longs. we examined whether an initial first impression changed in the course of judgment formation entirely or only slightly. moreover, we also examined whether a change in participants’ judgments was abrupt or whether the change evolved gradually. finally, some codes were deleted due to the scarcity of their application (cope, 2010; miles & huberman, 1994). 5. results the current study explored the cognitive mechanisms underlying thin slices ratings of instructional quality. since no previous study has examined this phenomenon, we employed both inductive and deductive analysis approaches. the verbal data were analyzed qualitatively, resulting in a set of codes that indicated cognitive mechanisms in relation to impression formation. the codes were then statistically analyzed to compare the two rating situations. 4 table 3 corrected sum of codes: relation to words (fifth column) table 4 lexical search keyword thin slices situation long video situation cognitive system automatic 0 4 1 intuitive 2 1 1 quick/fast 16 5 1 atmosphere 14 3 1 warmth 4 1 1 criterion 3 0 2 reflect(-ed) 3 4 2 fidget/fidgety 11 3 noisy 19 0 disturbances 22 20 note. absolute number of appearances for each condition. vinokic, begrich, kunter & kuger 83 | f l r 5.1 preliminary analysis in preliminary analyses, we tested whether our overall design worked by examining whether raters were able to report about their judgment formation processes. all participants in the thin slices rating situation (slices) could answer all questions. in total, the verbal data of the ten participants in the slices consisted of 13.277 words, with an average of 1.328 words per participant (sd = 419 words). in comparison, participants in the long video situation (longs) answered with a total of 19.773 words, with an average of 1.977 words per participant (sd = 937 words). the code “easiness of judgment” revealed no substantial differences between the groups. the code “difficulty of judgment” showed that participants in the slices mentioned somewhat more often (0.26 vs. 0.15) difficulties than participants in the longs. however, all ten participants in the longs did this to a certain extent as well (table 3). moreover, significantly more participants in the slices (9 vs. 3) complained significantly more often (30 vs. 6) about the scarcity of information than participants in the longs, which was indicated by the code “insufficient information” (table 3). 5.2 system 1 processing as the foundation of thin slices ratings (research question 1) to address research question 1, we examined the verbal data for evidence of typical system 1 processes, expecting to find more evidence for system 1 processes in the verbal reports of participants undergoing the slices. descriptive results are presented in table 3. with respect to the sum of codes (table 3; third column), all seven significant test results indicated group differences as expected in our research questions. concerning the sum of individuals (table 3; column far right), all five significant test results indicated differences between the two groups as expected. the direct coding of supposed system 1 processes is indicated with the code “system 1”. in relation to words, evidence of system 1 was found more often in the slices (0.19) than in the longs (0.11), meaning that per 100 words the code “system 1” was applied 0.19 in the slices and only 0.11 times in the longs (table 3). however, the difference between the conditions was not statistically significant. in the verbal data of all 20 participants, both, in slices as well as in the longs, direct evidence for system 1 was found, meaning that all participants operated in system 1 at some stage. the finding that thin slices raters operated in system 1, as did participants in the longs, is predominantly in accordance with our hypothesis. the code “own school days” is an indicator of system 1. the code was more prevalent in the slices (table 3). not only did participants in the slices refer more often to their own school biography (13 vs. 2), but also more participants did so in the slices than in the longs (8 vs. 2; difference tests were significant for “sum of codes” and “sum of participants”). this result is in line with our expectations. the code “emotion” is an indicator for processes of system 1. all 10 participants in the slices and nine participants in the longs reported at least once an emotion. however, in relation to words, participants in the slices reported fewer emotions (0.23) than participants in the longs (0.32), yet the difference was not statistically significant (table 3). further, we searched all text segments, which were scored with the code “emotion” whether the expressed emotions were positive or negative. yet, no interesting results distinguishing the two rater groups emerged. summarizing, participants in both rater groups reported an emotion similarly often. this result is not in line with our expectations. the code “halo effect” is an indicator for processes of system 1. in the comments of nine participants in both conditions evidence for the halo effect was found. in relation to words, more evidence for the halo effect was found in the slices (0.20) than in the longs (0.10), but the difference was not statistically significant (table 3). summarizing the results, the evidence suggest that thin slices raters are prone to the halo effect—but not exclusively. evidence for the halo effect was found slightly more often, though not significantly, in the verbal reports of participants in the slices. this means that this finding rather supports our research expectations. we assume that the code “evaluation” is associated with processes of system 2 (see also 5.3). we analyzed whether participants’ evaluations were positive, neutral, or negative by conducting a onevinokic, begrich, kunter & kuger 84 | f l r way anova. evaluations in the slices were negative 37 times (m = 3.7, sd = 1.70), neutral 14 times (m = 1.4, sd = 1.17), and positive 19 times (m = 1.9, sd = 1.52), whereby the differences between negative, neutral, and positive evaluations were statistically significant [f(2, 27) = 6.65, p = .004]. in the longs, exactly the opposite occurred: more evaluations were positive (86 times) than negative (45 times) or neutral (25 times). likewise, the difference was statistically significant [f(2, 27) = 9.02, p = .001]. in sum, the evidence suggests that thin slices raters’ evaluations were significantly more often negative than positive. the three codes “teacher”, “student(s)” and “situation & objects”, indicating participants’ focus of attention, are not associated with system 1 or system 2 functioning (table 1). no significant results were found distinguishing the two rater groups (table 3). consequently, these codes are not further discussed. the lexical search revealed that the terms “atmosphere” and “quick/fast”, which we consider as indicators of system 1 occurred considerably more frequently in the slices in comparison to the longs (table 4), even though participants in the slices produced on average fewer words than participants in the longs. these results support our claim that processes of system 1 are the foundation of thin slices ratings. 5.3 are processes of system 2 dissimilar to cognitive processes of thin slices ratings (research question 2)? in order to analyze whether the underlying cognitive processes of thin slices ratings are sufficiently dissimilar to those of system 2, we examined the participants’ verbal data for evidence of typical processes of system 2 and whether such evidence occurred more frequently in the longs. the direct code “system 2” refers to processes of system 2 (table 2). significantly less participants in the slices (3 vs. 9) operated significantly less often in system 2 (5 vs. 22) than participants in the longs (table 3). moreover, in relation to words, participants operated less often in system 2 in the slices (0.04) in comparison to the longs (0.11; table 3). in sum, these findings support our research expectation. the code “system 2 overrides system 1” is associated with cognitive operations of system 2 (table 2). significantly less participants in the slices (2 vs. 9) changed their impression significantly less often (4 vs. 19 times) than participants in the longs (table 3). further, taken the relative occurrence (table 3) into account, it seems that participants in the slices (0.03) changed a first impression less often than participants in the longs (0.10). we analyzed the verbal data regarding the code “system 2 overrides system 1” only for the longs in depth (see section 4.5) in order to analyze whether an initial first impression changed in the course of impression formation entirely or only slightly. we found that only once a participant changed his/her impression about the teacher entirely, whereas eight times an impression was corrected or refined only slightly. moreover, we wanted to determine whether the change of the initial first impression was sudden and abrupt or whether it gradually evolved. we found no evidence in the data that the change was abrupt or sudden, but in six cases the final judgment evolved gradually. summarizing the results, in contrast to raters relying on more information (i.e., 10-min videos), it seems that thin slices raters of instructional quality only rarely change an initial first impression. the code “detail” can be considered as an indicator of system 2 processing. in the slices, significantly less participants (3 vs. 10) remembered significantly less details (7 vs. 39) than participants in the longs. controlling for possible distorting effects due the length of the answers, participants reports were less detailed in the slices (0.05 vs. 0.20; table 3). in sum, it seems that reports of participants in the slices are less detailed, which aligns with our research expectation. the code “comparison” is an indicator of system 2. in the slices, less participants (6 vs. 10) compared significantly less often (8 vs. 40) the given information than participants in the longs (table 3). correcting for the number of words, participants in the slices compared the information less often vinokic, begrich, kunter & kuger 85 | f l r (0.06 vs. 0.20) than in the longs. in sum, it seems that participants in the slices less frequently engage in comparing information, which we consider as supporting evidence for our research question. the code “evaluation” is an indicator of system 2. in the slices as well as in the longs every single participant evaluated the given information in some way. in total, participants in the slices evaluated the information significantly less often (69 vs. 148) than participants in the longs. correcting for the number of words, participants in the slices evaluated the information less often (0.52 vs. 0.75). in line with the research expectation, participants in the slices evaluate the given information less often. the two codes “keep s.th. in mind” and “item-situation” are representing cognitive strategies and both are indicators for system 2 processing. no statistically significant differences were found between the groups (table 3). however, slightly more evidence for cognitive strategies were found in participants’ data in the longs. in sum, we claim to have found some evidence for the presence of cognitive strategies in judgment formation of thin slices raters. this result is not quite in accordance with our research expectation. 6. discussion assessing instructional quality in schools and in ecec is very costly and labor-intensive (murphy & hall, 2021). therefore, a more economical and yet accurate approach would be desirable. previous research has demonstrated that raters, relying on minimal information, can accurately assess teaching quality in schools and ecec (ambady & rosenthal, 1993; begrich et al., 2021; sokolovic et al., 2021; vinokic et al., 2024). the present study examined the underlying cognitive processes of ratings based on first impressions of instructional quality (i.e., thin slices ratings). to our knowledge, no empirical evidence has been documented that directly addresses this issue. based on the literature of dual process theory, we expected to find evidence of typical system 1 processes and little to no evidence of system 2 processing in the verbal data of thin slices raters. gaining insights into the cognitive mechanisms shaping impression formation in thin slices raters may help to optimize the practical application of the thin slices technique as a trustful measurement method to complement the established repertoire of measurement techniques in the field of teaching quality. in the following section, the discussion focuses on embedding this study’s results in the current literature on dual process theory, classroom observation methodology and thin slices research. initially, we will discuss whether participants are able to retrospectively verbalize their cognitive processes. subsequently, we will discuss whether typical processes of system 1 are the foundation of thin slices ratings (rq1) and whether thin slices ratings are dissimilar to typical processes of system 2 (rq2). are participants in thin slices rating situations of instructional quality able to report about their judgmental processes at all? only a few studies have examined the level of awareness associated with first impressions, whereby the evidence is mixed (ames et al., 2010; biesanz et al., 2011). given that participants are able to report about their judgmental processes, are the verbal reports expedient and useful for drawing conclusions about the underlying cognitive processes of thin slices raters? various pieces of evidence suggest that thin slices raters are indeed able to report about their judgmental processes. both thin slices raters as well as raters relying on more information equally complained about the difficulty of the tasks, indicating that the two different rating situations did not seem to influence the perceived difficulty of assessing instructional quality or verbalizing mental processes. the answers of thin slices raters differed significantly in many aspects from the answers of participants relying on more information (i.e., 10 minutes). the fact that statistically significant results were found between the two rater groups may be interpreted as a hint of different cognitive processes at work. the presence of clear and statistically significant patterns distinguishing the two rater groups supports the idea of different processes of social perception. moreover, the significant patterns distinguishing the two groups are in line with the current body of literature about dual process theories. this allows the conclusion that thin slices raters are indeed able to report about their underlying cognitive processes, and that these processes are indeed different from participants in the long video situation. or, to put it differently: if system 1 vinokic, begrich, kunter & kuger 86 | f l r functioning would not manifest in the verbal data, our analysis probably would not have detected statistically significant patterns. 6.1 system 1 processing as the foundation of thin slices ratings all participants in both groups seem to have relied equally on system 1. however, thin slices judgments rely to a lesser degree on system 2 functioning compared to judgments based on more information (i.e., 10-minutes videos; longs). the following evidence suggests that processes of systems 1 are the foundation of thin slices ratings. a defining feature of system 1 processes is autonomy. the execution of autonomous processes tends to be associative (evans & stanovich, 2013; sloman, 1996). thin slices raters more frequently referred to biographic information in terms of associations with their earlier school life (e.g., their elementary school teachers). in comparison, participants exposed to more information had far fewer associations with their own school days. associations are considered a typical feature of system 1 processing. this supports the assumption that system 1 is active while undergoing a typical thin slices rating situation. the fact that these associations are related to teaching or school is not trivial since it may have consequences for rating results received via the thin slices technique. positive or negative biographic memories respectively individual experiences may impact thin slices ratings of instructional quality. further investigations of a potential effect of individual experiences on the accuracy of thin slices ratings are necessary. as expected, several of the codes that reflected system 1 processing occurred particularly often in the thin slices situation, yet some unexpected results emerged. we expected more emotions to occur in judgments based on first impressions because emotional processing is explicitly linked to system 1 (evans, 2008; kahneman, 2011). however, we found the opposite: more emotions were detected in judgments based on more information instead of in judgments based on first impressions. a particularity of the german language might contribute at least partly to this finding, though. the german word for feeling (“gefühl”) can carry two semantically different meanings: one refers to an emotion or a feeling and the other refers to an assumption or a hunch. hence, the validity of this code may be doubted. the halo effect is a type of cognitive bias in which one aspect of a person influences the judgment of other aspects—a bias that is seen as typical for system 1 (e.g., kahneman, 2011). in our study, the majority of thin slices raters indeed exhibited the halo effect, indicating a superficial, holistic cognitive approach. however, raters exposed to more information were also prone to the halo effect, albeit to a somewhat lesser degree, suggesting that they, too, possibly processed information in system 1. in sum, it seems that thin slices raters of instructional quality are prone to the halo effect, but not notably more prone compared to judgments based on more information. all thin slices raters evaluated the information given in the video. from an evolutionary perspective, first impressions tend to be easily pulled to the negative end of the evaluative dimension. the fitness costs of an incorrect negative first impression are potentially lower than the fitness costs of an incorrect positive first impression (ambady & skowronski, 2008; nesse, 2005). a strong negative interpersonal impression may be formed on the basis of very little information, and a positive interpersonal impression may require a greater amount of informational input (ambady & skowronski, 2008). the fact that significantly more evaluations were negative than positive in verbal reports of thin slices raters is in line with ambady and skowronski (2008), as well as the fact that we found significantly more positive evaluations than negative in judgments based on more information. in sum, we claim to have found robust evidence that thin slices raters’ evaluations are rather negatively than positively. the results of the lexical search indicated that participants assessing instructional quality based on their first impression might operate holistically. nisbett et al. (2001) define holistic thought as involving an orientation to the context or the field as a whole, claiming that system 1 operates holistically. the word “atmosphere” (table 4) was used more frequently by thin slices raters than by vinokic, begrich, kunter & kuger 87 | f l r raters relying on more information. in the context of assessing classroom videos, we interpret the meaning of the word “atmosphere” as referring to the situation globally and as a whole, implying that “atmosphere” is a marker of system 1. in sum, we claim to have found some evidence for holistic information processing in thin slices judgments. in summary, assessing instructional quality solely based on first impressions (i.e., thin slices ratings) appears to be associatively, rather negatively than positively, prone to the halo effect, and probably holistically. according to the literature of dual process theories of social cognition, these results can be considered as typical processes of system 1. accordingly, system 1 appears to be the underlying cognitive system of thin slices ratings of instructional quality. therefore, research question 1 can be considered confirmed. 6.2 are processes of system 2 dissimilar to cognitive processes of thin slices ratings? all participants appeared to rely equally on system 1. however, thin slices judgments seem to rely to a much lesser degree on system 2 functioning than judgments based on more information (i.e., 10-minute videos), which seem to involve both systems—system 1 as well as system 2. the data lead to the conclusion that social cognition is initially dominated by system 1 processes, with system 2 being activated after a while. it appears that system 1 processes constitute the essential first phase of social cognition. these findings align with mugg (2015) who claims that type-1 and type-2 processes are generated at different times: at first type-1 is activated and then, under some circumstances, type-2 processes start to work. type-2 processing occasionally overrides or intervenes on type-1 responses rather than producing responses all on its own (mugg, 2015). moreover, our results suggest that raters with access to more information (i.e., 10-minute videos) altered their initial impression much more frequently than thin slices raters. we found no evidence in the data that this change was abrupt or sudden. instead, we found various hints of evidence that the judgment evolved gradually (see 5.1). based on the results, we may claim that the transition from system 1 processing to system 2 processing is not abrupt; rather, we can understand it as a gradual, with system 2 progressively taking over. our results do not provide an explanation of when system 2 is activated; however, according to our design, it appears to occur around or after 30 seconds, but within ten minutes. furthermore, we discovered evidence that system 2 processing did not entirely nullify or override system 1 judgments but rather shaped, modified or refined them. only one participant mentioned once that she/he entirely changed the initial first impression after a while. however, considerably more evidence was detected that the initial first impression was only slightly or to a minor extent modified. therefore, in hindsight, a more suitable name for the code would have been “system 2 modifies system 1”. in sum, we claim to have found evidence that system 2 is not the foundation of thin slices ratings, or only to a substantially lesser degree compared to judgments based on more information. kahneman (2011) posits that more detailed and specific processing of information is a feature of system 2. deeper cognitive processing should enhance the encoding and recall of details. considering that thin slices judgments do not rely on deep elaboration (i.e., system 2), thin slices raters should report less details than participants exposed to more information (10-minute videos). our study confirms this assumption. however, the research design may have confounding effects on the results. the amount of information provided, such as the length of the video, rather than the depth of processing, may have caused this finding. it is possible that longer videos simply lead to more details being memorized and subsequently recalled. we scrutinized the body of literature to determine which cognitive system generates comparisons. system 2 can compare objects based on several attributes (kahneman, 2011). participants exposed to more information made considerably more comparisons than thin slices raters. consequently, this finding appears to be evidence for the dissimilarity of typical processes of system 2 and the cognitive processes underlying thin slices ratings. vinokic, begrich, kunter & kuger 88 | f l r based on what we know about system 2 functioning, we presumed that an evaluation can be considered as a deliberate act involving the working memory and can therefore be subsumed to the processes of system 2 (evans, 2008; evans & stanovich et al., 2014). given that system 2 processes are not at activate in thin slices judgments, we did not expect thin slices raters to evaluate the given information. drawing on the evidence, thin slices raters evaluated the information to a lesser extent compared to raters with access to more information, highlighting a clear distinction between cognitive processes underlying thin slices ratings and processes of system 2. in the present study, a cognitive strategy was operationalized in a broader sense. two codes, “keep s.th. in mind” and “item-situation” (table 1 and 2), represent cognitive strategies. they were generated inductively based on the data and are very narrowly defined in comparison to what is considered a cognitive strategy in the literature. however, they summarized how participants proceeded and how they tried to solve the problem. the results indicate that the two codes occurred slightly more often in judgments based on more information than in judgments based on first impressions. the code “keep s.th. in mind” refers to the strategy of memorizing the rating items before watching the next classroom video. hence, it is not clear whether this code is truly a valid indicator for a cognitive strategy because some cognitive processes occurred already before the rating situation started (fig. 1). the code “item-situation” refers to participants trying to actively retrieving contents of the video while filling in the rating questionnaire. less thin slices raters tried to actively recall content from the video while filling in the rating questionnaire in contrast to participants observing 10-minute videos. in sum, these findings are partly in line with stanovich (2009) because he exclusively attributes strategic operation as being part of system 2 processes. in summary, we claim that research question 2 is confirmed, as the cognitive processes underlying thin slices ratings of teaching quality seem to be dissimilar from typical processes of system 2. based on our data, we posit that thin slices judgments derive very little from analytic or reflective processes. instead, thin slices judgments predominantly rely on system 1 processes. in comparison, the judgments of participants assessing teaching quality based on 10-minute video clips appear to be initially dominated by system 1 processes, but over time, these judgments are gradually refined and modified with the involvement of system 2. 7. limitations and future directions the present studies pursued a rather innovative and unconventional approach. to obtain some humble glimpses into the black box of the human mind, we set up two rating situations by implementing a typical thin slices situation and by approximating the rating procedure of a typical classroom observation study (hardy et al., 2011; desi-consortium, 2008). the aim was to detect the cognitive processes underlying thin slices ratings of teaching quality. although the results seem to be in line with the body of literature, some limitations need to be pointed out. in total, the rater sample consisted of 20 psychology undergraduates. while the sample size was relatively small, it was deemed adequate for our study. (see begrich et al., 2017). however, further research with a larger sample could carve out some cognitive processes or other phenomena more precisely, resulting in a refinement of the codes. moreover, a larger sample could contribute to a more robust interpretation and generalization of the results. in particular conducting the study with a different rater sample would be interesting. for instance, thin slices raters mentioned frequently associations to their own school days. these associations had the same content (e.g., school, teachers, classes) as the stimulus material (classroom videos). only undergraduate psychology students were invited as raters in the present study. conceivably, this rater population has rather positive associations to their own schooling. a different rater population, with rather negative associations to their own schooling, would potentially produce different results. further research should consist of a more diverse sample of raters. in a current study, we examine the accuracy of diverse thin slices rater samples with varying levels of expertise (children, undergraduates, teacher trainees, educational experts and adults with unfortunate educational careers). vinokic, begrich, kunter & kuger 89 | f l r in retrospective think aloud protocols, verbalization problems might arise, whereas concurrent think aloud protocols might negatively impact task performance (van den haak, 2003). to avoid interference with the rating process, the think aloud protocols were generated retrospectively. therefore, we used the confrontational video as a stimulation for participants. in both rating situations (slices and longs), retrospective think aloud protocols were used alike. the results should not be influenced by the retrospective think aloud protocols because they were applied in both rater groups. in order to explore the cognitive processes underlying thin slices ratings, we decided to develop an innovative, unconventional and risk-taking approach. to obtain meaningful results, we designed a typical thin slices rating situation and intended to approximate a conventional observer rating situation (e.g., hardy et al., 2011; desi-consortium, 2008). various parameters of the research design could have been varied, or even a completely different design could have been set up. in the thin slices rating situation, the time for completing the questionnaire was restricted to 30 seconds. young, educated adults, who had time to study the items before the experiment started, can answer six simple items, recurring ten times for each teacher, in 30 seconds. our previous thin slices research indicated that thin slices raters do not need more than 30 seconds to complete the rating of six simple items. further, this is sustained by the very low rate of missing data in the rating items (0.7 %). moreover, inferring from survey research (yan & tourangeau, 2008), we conclude that the answering time for six items should not exceed 30 seconds. 8. theoretical and practical implication untrained raters without teaching experience or didactic knowledge are able to assess instructional quality reliably and validly solely on the basis of 30-seconds classroom videos (begrich et al., 2017, 2020, 2021). researcher applying the thin slices technique to assess instructional quality often encounter skepticism and disbelief from fellow colleagues from the scientific community concerning the accuracy of thin slices ratings. how is such accuracy possible without analyzing and reflecting properly? the present study may contribute a piece to the puzzle on how thin slices ratings can yield accurate results without actively analyzing longer classroom videos (e.g., 10 minutes, 45 minutes or even longer). the fact that thin slices raters cannot actively analyze, reflect or think about the classroom videos is perhaps the clue to the solution. thin slices raters rely on a different, highly powerful cognitive system of information processing: system 1 of dual process theories of social cognition (kahneman, 2011; stanovich et al. 2014). conventional raters assessing full-length classroom videos rely predominantly on system 2 (and probably to a minor extent on system 1). strictly speaking, the fact that thin slices raters do not have enough time to carefully analyze the given information is actually not a drawback but a benefit because different cognitive processes are dominating judgment formation in comparison to conventional ratings. this implies that thin slices ratings could become a promising alternative to other forms of ratings. our study provides novel impulses for two different research strands: (a) for those who want to use or have already used the thin slices technique, the results provide insights into the cognitive processes involved in judgment formation and potential biases, and (b) for those working within the conventional paradigm of systematic video ratings of instructional processes. in the present study, the judgments of participants in a rating situation approximating a conventional observer study (10-minute videos) also appear to be influenced by system 1. what conclusion can be drawn from this finding for conventional observer studies (e.g., raters assessing instructional quality based on full-length videos) and their training programs? does the involvement of system 1 in conventional rating situations play an important yet underestimated role in judgment formation? can raters be trained to be vigilant about the influence of an initial first impression? should conventional raters be instructed to refrain from letting their judgments be guided by their first impressions? could this result in a diminished influence of system 1 processes, causing raters to operate predominantly in system 2? however, it is questionable whether processes of system 1 can actually be suppressed deliberately. does the intentional attempt to exclude or to reduce the influence of system 1 processes in conventional ratings of teaching quality worsen the accuracy of these ratings. further research is needed to unravel the interplay and distinct vinokic, begrich, kunter & kuger 90 | f l r influences of system 1 and system 2 on impression formation in conventional rating situations of instructional quality. keypoints thin slices ratings of teaching quality, based on 30-second classroom videos, appear to rely on typical cognitive processes of system 1 of dual process theories of social cognition. thin slices ratings of teaching quality are associative and tend to be rather negative than positive. ratings of teaching quality based on 10-minute classroom videos rely on both system 1 and system 2 of dual process theories of social cognition, with system 2 possibly modifying an initial judgment. references aleamoni, l. m. (1999). student rating myths versus research facts from 1924 to 1998. journal of personnel evaluation in education, 13(2), 153–156. https://doi.org/10.1023/a:1008168421283 ambady, n. (2010). the perils of pondering: intuition and thin slice judgments. psychological inquiry, 21(4), 271–278. https://doi.org/ 10.1037/0022-3514.64.3.431 ambady, n., bernieri, f. j., & richeson, j. a. (2000). toward a histology of social behavior: judgmental accuracy from thin slices of the behavioral stream. advances in experimental social psychology, 32, 201–271. https://doi.org/10.1016/s0065-2601(00)80006-4 ambady, n., & rosenthal, r. (1993). half a minute: predicting teacher evaluations from thin slices of nonverbal behavior and physical attractiveness. journal of personality and social psychology, 64(3), 431–441. https://doi.org/10.1037/0022-3514.64.3.431 ambady, n., & skowronski, j. j. (eds.). (2008). first impressions. guilford press. ames, d. r., kammrath, l. k., suppes, a., & bolger, n. (2010). not so fast: the (not-quite-complete) dissociation between accuracy and confidence in thin-slice impressions. personality and social psychology bulletin, 36(2), 264–277. https://doi.org/10.1177/0146167209354519 asch, s. e. (1946). forming impressions of personality. the journal of abnormal and social psychology, 41(3), 258–290. https://doi.org/10.1037/h0055756 audacity team. (2020). audacity recording and editing software (version 2.4.2) [computer software]. https://www.audacity.de/ babad, e. (2005). guessing teachers’ differential treatment of highand low-achievers from thin slices of their public lecturing behavior. journal of nonverbal behavior, 29(2), 125–134. https://doi.org/10.1007/s10919-005-2744-y baumert, j., kunter, m., blum, w., brunner, m., voss, t., jordan, a., klusmann, u., krauss, s., neubrand, m., & tsai, y.-m. (2010). teachers’ mathematical knowledge, cognitive activation vinokic, begrich, kunter & kuger 91 | f l r in the classroom, and student progress. american educational research journal, 47(1), 133– 180. https://doi.org/10.3102/0002831209345157 begrich, l., fauth, b., & kunter, m. (2020). who sees the most? differences in students’ and educational research experts’ first impressions of classroom instruction. social psychology of education, 23(3), 673–699. https://doi.org/10.1007/s11218-020-09554-2 begrich, l., fauth, b., kunter, m., & klieme, e. (2017). wie informativ ist der erste eindruck? das thin-slices-verfahren zur videobasierten erfassung des unterrichts [how informative is the first impression? the thin slices technique as video-based assessment of teaching quality]. zeitschrift für erziehungswissenschaft, 20(1), 23–47. https://doi.org/10.1007/s11618-0170730-x begrich, l., kuger, s., klieme, e., & kunter, m. (2021). at a first glance – how reliable and valid is the thin slices technique to assess instructional quality? learning and instruction, 74, 101466. https://doi.org/10.1016/j.learninstruc.2021.101466 bellini-leite, s. c. (2018). dual process theory: systems, types, minds, modes, kinds or metaphors? a critical review. review of philosophy and psychology, 9(2), 213–225. https://doi.org/10.1007/s13164-017-0376-x berliner, d. c. (2005). the near impossibility of testing for teacher quality. journal of teacher education, 56(3), 205–213. https://doi.org/10.1177/0022487105275904 biesanz, j. c., human, l. j., paquin, a.-c., chan, m., parisotto, k. l., sarracino, j., & gillis, r. l. (2011). do we know when our impressions of others are valid? evidence for realistic accuracy awareness in first impressions of personality. social psychological and personality science, 2(5), 452–459. https://doi.org/10.1177/1948550610397211 borkenau, p., mauer, n., riemann, r., spinath, f. m., & angleitner, a. (2004). thin slices of behavior as cues of personality and intelligence. journal of personality and social psychology, 86(4), 599–614. https://doi.org/10.1037/0022-3514.86.4.599 brennan, r. l., & prediger, d. j. (1981). coefficient kappa: some uses, misuses, and alternatives. educational and psychological measurement, 41(3), 687–699. https://doi.org/10.1177/001316448104100307 campbell, j. l., quincy, c., osserman, j., & pedersen, o. k. (2013). coding in-depth semistructured interviews: problems of unitization and intercoder reliability and agreement. sociological methods & research, 42(3), 294–320. https://doi.org/10.1177/0049124113500475 clausen, m. (2002). unterrichtsqualität: eine frage der perspektive? [quality of instruction: a matter of perspective?]. waxmann. cope, m. (2010). coding qualitative data. qualitative research methods in human geography, 3, 281– 294. decristan, j., kunter, m., & fauth, b. (2022). die bedeutung individueller merkmale und konstruktiver unterstützung der lehrkraft für die soziale integration von schülerinnen und schülern im mathematikunterricht der sekundarstufe [the relevance of individual characteristics and teacher‘s constructive support for students‘ social integration in mathematics instruction in secondary education]. zeitschrift für pädagogische psychologie, 36(1–2), 85–100. https://doi.org/10.1024/1010-0652/a000329 vinokic, begrich, kunter & kuger 92 | f l r decristan, j., kunter, m., fauth, b., büttner, g., hardy, i., & hertel, s. (2016). what role does instructional quality play for elementary school children’s science competence? a focus on students at risk. journal for educational research online, 8(1), 66–89. https://doi.org/10.25656/01:12032 de freitas, e. (2015). the moving image in education research: reassembling the body in classroom video data. international journal of qualitative studies in education, 29(4), 553–572. https://doi.org/10.1080/09518398.2015.1077402 de neys, w. (2021). on dual-and single-process models of thinking. perspectives on psychological science, 16(6), 1412–1427. https://doi.org/10.1177/1745691620964172 desi-konsortium (eds.). (2008). unterricht und kompetenzerwerb in deutsch und englisch: ergebnisse der desi-studie [teaching and competency acquistion in german and english: results of the desi-study]. beltz. desimone, l. m., smith, t. m., & frisvold, d. e. (2010). survey measures of classroom instruction: comparing student and teacher reports. educational policy, 24(2), 267–329. https://doi.org/10.1177/0895904808330173 doyle, w. (2006). ecological approaches to classroom management. in c. m. evertson & c. s. weinstein (eds.), handbook of classroom management: research, practice, and contemporary issues (pp. 97–125). lawrence erlbaum associates publishers. dresing, t., & pehl, t. (2012). praxisbuch transkription: regelsysteme, software und anleitungen für qualitative forscherinnen (4th ed.) [practical guide to transcription: rule systems, software, and instructions for qualitative researchers]. eigenverlag. evans, j. s. b. t. (2006). dual system theories of cognition: some issues. proceedings of the annual meeting of the cognitive science society, (28), 202–207. evans, j. s. b. t. (2008). dual-processing accounts of reasoning, judgment, and social cognition. annual review of psychology, 59(1), 255–278. https://doi.org/10.1146/annurev.psych.59.103006.093629 evans, j. s. b. t. (2018). dual process theory: perspectives and problems. in w. de neys (ed.), current issues in thinking and reasoning. dual process theory 2.0 (pp. 137–155). routledge. evans, j. s. b. t. (2019). reflections on reflection: the nature and function of type 2 processes in dualprocess theories of reasoning. thinking & reasoning, 25(4), 383–415. https://doi.org/10.1080/13546783.2019.1623071 evans, j. st. b. t., & stanovich, k. e. (2013). dual-process theories of higher cognition: advancing the debate. perspectives on psychological science, 8(3), 223–241. https://doi.org/10.1177/1745691612460685 fauth, b., herbein, e., & maier, j. l. (2024). beobachtungsmanual zum unterrichtsfeedbackbogen tiefenstrukturen [observation manual for the classroom feedback questionnaire on deepstructures]. institut für bildungsanalysen baden-württemberg. fowler, k. a., lilienfeld, s. o., & patrick, c. j. (2009). detecting psychopathy from thin slices of behavior. psychological assessment, 21(1), 68–78. https://doi.org/10.1037/a0014938 vinokic, begrich, kunter & kuger 93 | f l r gargani, j., & strong, m. (2014). can we identify a successful teacher better, faster, and cheaper? evidence for innovating teacher observation systems. journal of teacher education, 65(5), 389–401. https://doi.org/10.1177/0022487114542519 gawronski, b., luke, d. m., & creighton, l. a. (2024). dual-process theories. in d. e. carlston, k. hugenberg, & k. l. johnson (eds.), the oxford handbook of social cognition (2nd ed., pp. 319–353). oxford university press. göllner, r., fauth, b., & wagner, w. (2021). student ratings of teaching quality dimensions: empirical findings and future directions. in w. rollett, h. bijlsma, & s. röhl (eds.), student feedback on teaching in schools (pp. 111–122). springer international publishing. https://doi.org/10.1007/978-3-030-75150-0_7 goodell, l. s., stage, v. c., & cooke, n. k. (2016). practical qualitative research strategies: training interviewers and coders. journal of nutrition education and behavior, 48(8), 578–585. https://doi.org/10.1016/j.jneb.2016.06.001 hall, j. a., horgan, t. g., & murphy, n. a. (2019). nonverbal communication. annual review of psychology, 70(1), 271–294. https://doi.org/10.1146/annurev-psych-010418-103145 hannafin, m. j., shepherd, c. e., & polly, d. (2010). video assessment of classroom teaching practices: lessons learned, problems and issues. educational technology, 59(1), 32–37. hardy, i., hertel, s., kunter, m., klieme, e., warwas, j., büttner, g., & lühken, a. (2011). adaptive lerngelegenheiten in der grundschule. merkmale, methodisch-didaktische schwerpunktsetzungen und erforderliche lehrerkompetenzen [adaptive learning opportunities in elementary school. characteristics, methodological-didactic prioritization and required teacher competences]. zeitschrift für pädagogik, 57(6), 819–833. https://doi.org/10.25656/01:8783 hattie, j. (2023). visible learning: the sequel: a synthesis of over 2,100 meta-analyses relating to achievement. routledge. helmke, a. (2014). unterrichtsqualität und lehrerprofessionalität: diagnose, evaluation und verbesserung des unterrichts [teaching quality and teacher professionalism: diagnosis, evaluation, and improvement of instruction]. klett/kallmeyer. holleran, s. e., mehl, m. r., & levitt, s. (2009). eavesdropping on social life: the accuracy of stranger ratings of daily behavior from thin slices of natural conversations. journal of research in personality, 43(4), 660–672. https://doi.org/10.1016/j.jrp.2009.03.017 hyytinen, h., holma, k., toom, a., & shavelson, r. j., & lindblom-ylänne, s. (2014). the complex relationship between students’ critical thinking and epistemological beliefs in the context of problem solving. frontline learning research, 2(5), 1–25. https://doi.org/10.14786/flr.v2i4.124 janik, t., & seidel, t. (eds.). (2013). the power of video studies in investigating teaching and learning in the classroom. waxman. jung, m. f. (2016). coupling interactions and performance: predicting team performance from thin slices of conflict. acm transactions on computer-human interaction, 23(3), 1–32. https://doi.org/10.1145/2753767 vinokic, begrich, kunter & kuger 94 | f l r kahneman, d. (2003). a perspective on judgment and choice: mapping bounded rationality. american psychologist, 58(9), 697–720. https://doi.org/10.1037/0003-066x.58.9.697 kahneman, d. (2011). thinking fast and slow. macmillan. kane, t. j., & staiger, d. o. (2012). gathering feedback for teaching: combining high-quality observations with student surveys and achievement gains. research paper. met project. bill & melinda gates foundation. kawulich, b. b. (2004). data analysis techniques in qualitative research. journal of research in education, 14(1), 96–113. keren, g., & schul, y. (2009). two is not always better than one: a critical evaluation of two-system theories. perspectives on psychological science, 4(6), 533–550. https://doi.org/10.1111/j.17456924.2009.01164.x klieme, e., pauli, c., & reusser, k. (2009). the pythagoras study: investigating effects of teaching and learning in swiss and german mathematics classrooms. in j. tomáš & t. seidel (eds.), the power of video studies in investigating teaching and learning in the classroom (pp. 137– 160). waxmann. krippendorff, k. (2004). content analysis: an introduction to its methodology (2nd ed.). sage. kruglanski, a. w., & gigerenzer, g. (2011). intuitive and deliberate judgments are based on common principles. psychological review, 118(1), 97–109. https://doi.org/10.1037/a0020762 kuger, s., klieme, e., lüdtke, o., schiepe-tiska, a., & reiss, k. (2017). mathematikunterricht und schülerleistung in der sekundarstufe: zur validität von schülerbefragungen in schulleistungsstudien [mathematics instruction and student achievement in secondary education: about the validity of student’s survey and school achievement studies]. zeitschrift für erziehungswissenschaft, 20(2), 61–98. https://doi.org/10.1007/s11618-017-0750-6 kunter, m., & baumert, j. (2006). who is the expert? construct and criteria validity of student and teacher ratings of instruction. learning environments research, 9(3), 231–251. https://doi.org/10.1007/s10984-006-9015-7 kunter, m., brunner, m., baumert, j., klusmann, u., krauss, s., blum, w., jordan, a., & neubrand, m. (2005). der mathematikunterricht der pisa-schülerinnen und -schüler: schulformunterschiede in der unterrichtsqualität [mathematics instruction of the pisastudents]. zeitschrift für erziehungswissenschaft, 8(4), 502–520. https://doi.org/10.1007/s11618-005-0156-8 kunter, m., & voss, t. (2011). das modell der unterrichtsqualität in coactiv: eine multikriteriale analyse [the model of instructional quality in coactiv: a multi-criterial analysis]. in m. kunter, j. baumert, w. blum, u. klusmann, s. krauss, & m. neubrand (eds.), professionelle kompetenz von lehrkräften—ergebnisse des forschungs programms coactiv (pp. 85–113). waxmann. lambert, n. m., mulder, s., & fincham, f. (2014). thin slices of infidelity: determining whether observers can pick out cheaters from a video clip interaction and what tips them off. personal relationships, 21(4), 612–619. https://doi.org/10.1111/pere.12052 vinokic, begrich, kunter & kuger 95 | f l r learnpulse sas (2020). screenpresso pro (version 1.8.3.0) [computer software]. https://www.screenpresso.com/de/ leighton, j. p. (2017). using think-aloud interviews and cognitive labs in educational research. oxford university press. https://doi.org/10.1093/9780199372904.001.0001 lipowsky, f., rakoczy, k., pauli, c., drollinger-vetter, b., klieme, e., & reusser, k. (2009). quality of geometry instruction and its short-term impact on students’ understanding of the pythagorean theorem. learning and instruction, 19(6), 527–537. https://doi.org/10.1016/j.learninstruc.2008.11.001 marzano, r. j., & marzano, j. s. (2003). the key to classroom management. educational leadership, 61(1), 6–13. mckim, c. a. (2017). the value of mixed methods research: a mixed methods study. journal of mixed methods research, 11(2), 202–222. https://doi.org/10.1177/1558689815607096 miles, m. b., & huberman, a. m. (1994). qualitative data analysis: an expanded sourcebook. sage. milli, s., lieder, f., & griffiths, t. l. (2021). a rational reinterpretation of dual-process theories. cognition, 217, 104881. https://doi.org/10.1016/j.cognition.2021.104881 mugg, j. (2015). two minded creatures and dual-process theory. journal of cognition and neuroethics, 3(3), 87–112. murphy, n. a., & hall, j. a. (2021). capturing behavior in small doses: a review of comparative research in evaluating thin slices for behavioral measurement. frontiers in psychology, 12, 667326. https://doi.org/10.3389/fpsyg.2021.667326 namey, e., guest, g., thairu, l., & johnson, l. (2008). data reduction techniques for large qualitative data sets. in g. guest, & k. macqueen (eds.), handbook for team-based qualitative research, (pp. 137–161). rowman & littlefield. nesse, r. m. (2005). natural selection and the regulation of defenses. evolution and human behavior, 26(1), 88–105. https://doi.org/10.1016/j.evolhumbehav.2004.08.002 nisbett, r. e., peng, k., choi, i., & norenzayan, a. (2001). culture and systems of thought: holistic versus analytic cognition. psychological review, 108(2), 291–310. https://doi.org/10.1037/0033-295x.108.2.291 o’connor, c., & joffe, h. (2020). intercoder reliability in qualitative research: debates and practical guidelines. international journal of qualitative methods, 19, 1–13. https://doi.org/10.1177/1609406919899220 pennycook, g. (2017). a perspective on the theoretical foundation of dual process models. in w. de neys (ed.), dual process theory 2.0 (1st ed., pp. 5–27). routledge. https://doi.org/10.4324/9781315204550-2 petko, d., waldis, m., pauli, c., & reusser, k. (2003). methodologische überlegungen zur videogestützten forschung in der mathematikdidaktik: ansätze der timss 1999 video studie und ihrer schweizerischen erweiterung [methodological considerations about video-based research in mathematics didactic]. zentralblatt für didaktik der mathematik, 35(6), 265–280. https://doi.org/10.1007/bf02656691 vinokic, begrich, kunter & kuger 96 | f l r pianta, r. c., & hamre, b. k. (2009). conceptualization, measurement, and improvement of classroom processes: standardized observation can leverage capacity. educational researcher, 38(2), 109–119. https://doi.org/10.3102/0013189x09332374 praetorius, a.-k., klieme, e., herbert, b., & pinger, p. (2018). generic dimensions of teaching quality: the german framework of three basic dimensions. zdm – mathematics education, 50(3), 407– 426. https://doi.org/10.1007/s11858-018-0918-4 praetorius, a.-k., lenske, g., & helmke, a. (2012). observer ratings of instructional quality: do they fulfill what they promise? learning and instruction, 22(6), 387–400. https://doi.org/10.1016/j.learninstruc.2012.03.002 pretsch, j., flunger, b., heckmann, n., & schmitt, m. (2013). done in 60 s? inferring teachers’ subjective well-being from thin slices of nonverbal behavior. social psychology of education, 16(3), 421–434. https://doi.org/10.1007/s11218-013-9223-9 rädiker, s., & kuckartz, u. (2019). analyse qualitativer daten mit maxqda: text, audio und video [analysis of qualitative data with maxqda: text, audio, and video]. springer fachmedien wiesbaden. https://doi.org/10.1007/978-3-658-22095-2 rimondini, m., mazzi, m. a., busch, i. m., & bensing, j. (2019). you only have one chance for a first impression! impact of patients' first impression on the global quality assessment of doctors' communication approach. health communication, 34(12), 1413–1422. https://doi.org/10.1080/10410236.2018.1495159 ritchie, s. j., & tucker-drob, e. m. (2018). how much does education improve intelligence? a metaanalysis. psychological science, 29(8), 1358–1369. https://doi.org/10.1177/0956797618774253 saldaña, j. (2009). the coding manual for qualitative researchers. sage. sandelowski, m., voils, c. i., & knafl, g. (2009). on quantitizing. journal of mixed methods research, 3(3), 208–222. https://doi.org/10.1177/1558689809334210 schensual, j. j., lecompte, m. d., hess; g. a., nastasi, b. k., berg, m. j., williamson, l., brecher, j. & glasser, r. (1999). using ethnographic data: interventions, public programming and public policy (vol. 7). altamira. schoonenboom, j., & johnson, r. b. (2017). how to construct a mixed methods research design. kzfss kölner zeitschrift für soziologie und sozialpsychologie, 69(s2), 107–131. https://doi.org/10.1007/s11577-017-0454-1 sloman, s. a. (1996). the empirical case for two systems of reasoning. psychological bulletin, 119(1), 3–22. sokolovic, n., brunsek, a., rodrigues, m., borairi, s., jenkins, j. m., & perlman, m. (2021). assessing quality quickly: validation of the responsive interactions for learning – educator (rifl-ed.) measure. early education and development, 33(6), 1061–1076. https://doi.org/10.1080/10409289.2021.1922851 stanovich, k. e. (2009). distinguishing the reflective, algorithmic, and autonomous minds: is it time for a tri-process theory? in j. evans & k. frankish (eds.), two minds: dual processes and vinokic, begrich, kunter & kuger 97 | f l r beyond (1st ed., pp. 55–88). oxford university press. https://doi.org/10.1093/acprof:oso/9780199230167.003.0003 stanovich, k. e., & west, r. f. (2000). individual differences in reasoning: implications for the rationality debate? behavioral and brain sciences, 23(5), 645–665. https://doi.org/10.1017/s0140525x00003435 stanovich, k. e., west, r. f., & toplak, m. e. (2014). rationality, intelligence, and the defining features of type 1 and type 2 processing. in j. w. sherman, b. gawronski & y. trope (eds.), dualprocess theories of the social mind (pp. 80–91). the guilford press. tackett, j. l., herzhoff, k., kushner, s. c., & rule, n. (2015). thin slices of child personality: perceptual, situational, and behavioral contributions. journal of personality and social psychology, 110(1), 150–66. https://doi.org/10.1037/pspp0000044 tashakkori, a., & creswell, j. w. (2007). editorial: the new era of mixed methods. journal of mixed methods research, 1(1), 3–7. https://doi.org/10.1177/2345678906293042 trautwein, u., sliwka, a., & dehmel, a. (2022). grundlagen für einen wirksamen unterricht. reihe wirksamer unterricht band 1 [basics of effective teaching. series on effective teaching, vol. 1]. institut für bildungsanalysen baden-württemberg. van den haak, m., de jong, m., & jan schellens, p. (2003). retrospective vs. concurrent think-aloud protocols: testing the usability of an online library catalogue. behaviour & information technology, 22(5), 339–351. https://doi.org/10.1080/0044929031000 vinokic, k., baron, f., kunter, m., linberg, a., begrich, l., & kuger, s. (2024). using the thin slices technique to assess interactional quality in early childhood education and care settings. frontiers in education, 9, 1368503. https://doi.org/10.3389/feduc.2024.1368503 visser, d., & matthews, j. d. l. (2005). the power of non-verbal communication: predicting job performance by means of thin slices of non-verbal behaviour. south african journal of psychology, 35(2), 362–383. https://doi.org/10.1177/008124630503500212 wagner, w., göllner, r., werth, s., voss, t., schmitz, b., & trautwein, u. (2016). student and teacher ratings of instructional quality: consistency of ratings over time, agreement, and predictive power. journal of educational psychology, 108(5), 705–721. https://doi.org/10.1037/edu0000075 white, m., & ronfeldt, m. (2024). monitoring rater quality in observational systems: issues due to unreliable estimates of rater quality. educational assessment, 29(2), 124–146. https://doi.org/10.1080/10627197.2024.2354311 willis, j., & todorov, a. (2006). first impressions: making up your mind after a 100-ms exposure to a face. psychological science, 17(7), 592-598. https://doi.org/10.1111/j.1467-9280.2006.01750.x wood, t. j. (2014). exploring the role of first impressions in rater-based assessments. advances in health sciences education, 19(3), 409–427. https://doi.org/10.1007/s10459-013-9453-9 yan, t., & tourangeau, r. (2008). fast times and easy questions: the effects of age, experience and question complexity on web survey response times. applied cognitive psychology: the official journal of the society for applied research in memory and cognition, 22(1), 51–68. https://doi.org/10.1002/acp.1331 vinokic, begrich, kunter & kuger 98 | f l r appendix the questions of the guided interview 1. at first, please tell us with regard to all videos and all questionnaires: what was going on in your head while you were watching the videos and completing the questionnaires? 2. how did you arrive at your judgment? how did you proceed? 3. what was particularly difficult or easy? 4. how did your decision evolve? was it rather reflective or was it rather a gut decision? 5. did some impressions influence your ratings in particular? 6. what was crucial for your decision? was it rather the video or rather the questionnaire or both? 7. now let us talk about the teachers: was something noticeable about the teachers, which was specifically remarkable? 8. what about the audio content? how important was it? 9. on which sector of the screen did you focus? 10. did you pay attention to the gesture, posture, voice or clothes of the teacher? if so, how did it influence your ratings? 11. did you notice a smile? if so, how did it influence your decision? 12. in how far was the attitude of the students, the atmosphere or the condition of the classroom crucial for your ratings? frontline learning research vol. 12 no. 3 (2024) 45 68 issn 2295-3159 correspondence concerning this article should be addressed to kristiina räihä, centre for university teaching and learning (hype), faculty of educational sciences, university of helsinki (p.o. box 9, 00014 university of helsinki, finland). email: kristiina.raiha@helsinki.fi doi: https://doi.org/10.14786/flr.v12i3.1415 effects of study-integrated well-being course intervention for different burnout and engagement profiles of university students kristiina räihä 1, nina katajavuori 1, kimmo vehkalahti 2& henna asikainen 1 1 centre for university teaching and learning (hype), university of helsinki, finland 2 centre for social data science, university of helsinki, finland article received 15 december 2023 / article revised 15 august 2024 / accepted 3 october 2024 / available online 10 october 2024 abstract as the increase in university students’ mental health problems poses significant challenges to education, new research-based ways to support students’ well-being in higher education are urgently needed. acceptance and commitment therapy (act) has been shown to be effective in enhancing different aspects of well-being in a variety of contexts. although study burnout has been shown to have detrimental impacts on students’ well-being and studying, not much is known about the person-oriented features of burnout risk in relation to act-based intervention outcomes. the aim of this study was to compare the effects of an online act-based course on different aspects of university students’ well-being and study ability in latent groups of students that had different levels of study burnout and engagement at the beginning of the intervention. the results of latent profile analysis (lpa) showed that students (n=352) represented four profiles at the beginning of the course: indifferent (37.8%), engaged (29%), engaged-inefficacious (13.6%), and burned-out (19.6%). the results of mixed anova with repeated measures showed that psychological flexibility, well-being, study engagement and organised studying increased, and study burnout risk decreased in the whole sample of courseparticipating students. the changes in students’ exhaustion, well-being, psychological flexibility, and organised studying did not differ between the four burnout and engagement profiles. the profiles differed in the changes of cynicism, inadequacy and engagement. the results of this study provide new knowledge of the person-oriented features of study burnout and indicate that act-based course interventions can be effective in enhancing different profiles representing university students’ well-being. keywords: study burnout; engagement; well-being; psychological flexibility; organised studying; university students mailto:kristiina.raiha@helsinki.fi räihä, katajavuori, vehkalahti & asikainen 46 | f l r 1. introduction higher education students’ mental health problems have been a growing concern throughout recent decade (auerbach et al., 2019), and during the covid-19 pandemic, even more challenges for students’ psychological well-being were reported (unesco, 2020). one key factor threatening students' well-being is study burnout, which has been shown to have detrimental effects also on students’ academic aspirations (salmela-aro & upadyaya, 2017), achievement (madigan & curran, 2021), and study engagement (salmela-aro & upadyaya, 2014; schaufeli, et al., 2002). alongside psychological and behavioural disadvantages, burnout has also been shown to be related to higher prevalence of physical illness in the general population (honkonen et al., 2006), and poorer physical health in university students (pagnin, & de queiroz, 2015). given that burnout during studies has also been shown to be related to mental health issues, such as depressive symptoms (salmela-aro et al., 2009), developing study-integrated methods to mitigate the risk of study burnout is crucial. different kinds of behaviour therapy interventions have been shown to be effective in reducing students’ burnout but studied mainly with students of secondary and tertiary levels of education with initially high levels of burnout (madigan et al., 2023). in recent decades, psychological flexibility rooted in a cognitive behavioural therapy intervention acceptance and commitment therapy (act) has been demonstrated to be a crucial factor for well-being, health, and performance across various contexts (hayes et al., 2006; bond et al., 2013; kashdan & rottenberg, 2010; räsänen et al., 2016). earlier studies of interventions aiming to enhance students’ psychological flexibility have for example shown to reduce students’ burnout risk (frögéli et al., 2016; räihä et al., 2024), and to enhance students’ well-being (howell & passmore, 2019), organised studying (katajavuori et al., 2021; räihä et al., 2024), and study engagement (grégoire et al., 2018), but there has not been prior person-oriented research comparing the effects of interventions on different study burnout and engagement profiles of university students. thus, to gain new knowledge of ways to prevent study burnout risk and to further develop study-integrated wellbeing-enhancing interventions that are useful to different burnout profiles representing university students, more research is needed on the person-oriented features of burnout risk (salmela-aro & read, 2017), and differences in relation to intervention outcomes. in this study, we seek to face this gap in knowledge by examining university students’ study burnout and engagement profiles, and the differences in their well-being, psychological flexibility, and organised studying after participating in an act-based course intervention aiming to enhance students’ psychological flexibility and organised study skills. 2. theoretical background 2.1 university students’ well-being, burnout, and engagement as well-being is a multi-faceted construct with many definitions and traditions (dodge et al., 2012), university students’ well-being can be conceptualised in many ways. according to keyes’s (2002) essential model of well-being, mental health manifests through various positive aspects of emotional, psychological, and social well-being, and is not indicated merely by the absence of mental illbeing. emotional well-being can be viewed as presence of positive affect, absence of negative affect and perceived satisfaction with life (keyes, 2002), as psychological well-being (ryff, 1989), and social well-being (keyes, 1998) refer to how individuals see themselves functioning and thriving in their personal and social life. university students’ higher levels of mental well-being have been shown to be associated for example with lower levels of academic stress (sarasjärvi et al., 2022). longitudinal studyrelated stress can lead to depletion of students’ well-being in the form of study burnout, which can be described as an overall construct of three dimensions of exhaustion, cynicism, and inadequacy (maslach et al., 2001; salmela-aro, et al., 2009; 2022; schaufeli et al., 2002). exhaustion refers to feelings of strain and fatigue that follows overly burdensome study demands (salmela-aro & read, 2017; schaufeli et al., 2002). when faced with long-lasting exhaustion, one might react with emotional detachment and räihä, katajavuori, vehkalahti & asikainen 47 | f l r by lowering the value of studies to them. this second dimension cynicism is manifested in an indifferent stance towards academic work, a loss of interest, and decreased feelings of meaningfulness in studying (salmela-aro & read, 2017; schaufeli et al., 2002). finally, feelings of inadequacy and a lack of studyrelated efficacy refers to diminished feelings of achievement and productivity in one’s studies (maslach et al., 2001; schaufeli et al., 2002). previous studies in higher education have shown that, emotional exhaustion (ríos‐risquez et al., 2018), and cynicism (kachel et al., 2020; pagnin & de queiroz, 2015) are related to lower levels of students’ psychological well-being. in addition, study-related exhaustion, cynicism, and inadequacy have been found to be related to depressive symptoms (salmela-aro et al., 2009), reduced academic achievement (asikainen et al., 2022; madigan & curran, 2021; salmela-aro & upadyaya, 2017), and to have detrimental effect on students’ study engagement (salmela-aro & upadyaya, 2014). engagement, which was originally defined as a positive fulfilling work-related state of mind, refers to energy, dedication, and absorption (schaufeli et al., 2002). in study context, engagement can be seen as a motivational process, and to have an emotional dimension referring to a positive approach and attitude towards studying (energy), cognitive dimension referring to perceiving schoolwork as meaningful (dedication), and behavioural dimension referring to concentration on and absorption in studying (salmela-aro & upadyaya, 2012; salmela-aro et al., 2022). study engagement has been found to predict students’ emotional well-being (salmela-aro & upadyaya, 2014), and to be related to academic success in higher education (ketonen et al., 2016). the multidimensional constructs of study burnout and engagement are related to each other, as well as different demands and resources (bakker et al., 2023). personal resources, such as socioemotional skills, adaptive coping strategies and performance capacity can promote students’ engagement and prevent development of burnout risk (salmela-aro et al., 2022; schaufeli & taris, 2014). for example, organised studying, referring to students’ ability to self-regulate and manage time and the tasks related to everyday studying (entwistle & mccune, 2004), has been shown to be related to university students’ lower study burnout risk (räihä et al., 2024). furthermore, students’ better abilities to self-regulate studying have been shown to be related to academic success (pérez-gonzález et al., 2022), and higher levels of emotional, psychological, and social well-being (davis & hadwin, 2021; howell, 2009). earlier studies have shown that acceptance and commitment therapy (act) based interventions can improve all three components of well-being – emotional, psychological, and social well-being (fledderus et al., 2010; wersebe et al., 2018). the overall aim of act is to increase participants psychological flexibility, that refers to adaptive ability to be aware and open to diverse experiences, and to persist in behaviour that serves individuals valued objectives (hayes et al., 2012). psychological flexibility represents thus a personal resource that has been found to be an important factor for well-being and performance in several contexts (bond et al., 2013; hayes et al. 2006; 2012; kashdan & rottenberg, 2010; onwezen et al., 2014), including higher education (hailikari et al., 2022). psychological flexibility has been traditionally conceptualised with six processes: acceptance (ability to experience or accept psychological events without creating the behavioural harm of trying to avoid experiences), defusion (ability to notice the act of thinking), self-as-context (ability to take perspective of self as opposite of rigid conceptualisation of self), present moment awareness (ability to direct attention flexibly to the present moment), values (ability to acknowledge one’s values), and committed action (ability to engage and commit to valued behaviour) (hayes et al., 2006). these six processes can be further conceptualised as three pillars or dimensions of psychological flexibility: openness to experience (acceptance and defusion), behavioural awareness (self-as-context and present moment awareness), and valued action (values and committed action) (hayes et al., 2011). previous studies in higher education have shown, that act-based interventions decrease students’ perceived stress (katajavuori et al., 2021; räsänen et al., 2016; viskovich & pakenham, 2020) and burnout risk (frögéli et al., 2016; räihä et al., 2024) and increase study engagement (grégoire et al., 2018). in their systematic review and initial meta-analysis of the role of act-based interventions in promoting university students’ well-being, howell and passmore (2019) found a small but significant pooled effect size on well-being of the five studies chosen for the analysis. as the research of effects of act-based interventions on study burnout is still quite scarce, there are yet no systematic reviews. however, in a recent systematic review of act for professional staff burnout, towey-swift and others räihä, katajavuori, vehkalahti & asikainen 48 | f l r (2023) found that an act-based intervention reduced burnout symptoms of employees in nine of the 14 chosen studies. to summarise, psychological flexibility alongside organised studying can be seen as psychological resources, that can play a role in supporting students’ adaptive capacity to answer different kind of demands during studies thus preventing the development of burnout risk, alongside promoting study engagement, and different dimensions of students’ well-being. 2.2 person-oriented features of study burnout and engagement in the earlier burnout studies, the dimensions of engagement were seen to represent opposite of the dimensions of burnout (maslach et al., 2001; salmela-aro & upadyaya, 2012; schaufeli et al., 2002). however, later studies have shown that engagement in studies does not necessarily mean that student has low study burnout risk (salmela-aro et al., 2016). as variable-centred approaches describe associations between variables, person-oriented approaches applying research aims to form patterns, clusters, or groups based on the differences among individuals in how variables are related to each other (laursen & hoff, 2006). research that applies a person-oriented approach has shown that, when burnout is analysed together with other well-being constructs, different components are accentuated in different profiles (mäkikangas & kinnunen, 2016). in their person-oriented research of school burnout and engagement of high school students and young adults (n=979), tuominen-soini and salmela-aro (2014) found four profiles: engaged (44%), engaged-exhausted (28%), cynical (14%), and burned-out (14%). later salmela-aro and reid (2017) found similar profiles – engaged (44%), engaged-exhausted (30%), inefficacious (19%), and burned-out (7%) – in a person-oriented study of higher education students (n=12 394). as both elevated levels of engagement and different dimensions of study burnout can coexist in different profiles, this indicates that engagement and burnout are not merely contradictory phenomena and underlines the need to study them together (salmela-aro et al., 2016). research of employees participating in a mindfulness, acceptance, and value-based (mav) intervention has shown that different profiles of burnout and mindfulness can differ in terms of changes in burnout risk, mindfulness skills (kinnunen et al., 2019), and well-being (kinnunen et al., 2020). to our knowledge, there has not been prior person-oriented research comparing the effects of act-based interventions on different study burnout and engagement profiles. the lack of previous research, and the notion, that study burnout profiles enable a more versatile picture of university students’ burnout risk, underline the need to study latent subgroups of students who might benefit differently from the study-integrated act-based intervention according to the level of study burnout risk and engagement at the beginning of the intervention. 3. present study in this study, we aim to explore the changes in study engagement, burnout risk, well-being, psychological flexibility, and organised studying of different burnout and engagement profiles representing university students, who participate in a study-integrated act-based intervention course. first, we utilise a person-oriented approach to find out if what kind of latent groups students represent at the beginning of the course, and second compare the changes in outcomes at the end of the course between these profiles. two research questions were formulated. rq 1: what kinds of study burnout and engagement profiles do students represent at the beginning of the course? based on previous burnout profile studies with students, we hypothesised that students represent profiles that differ in the levels of exhaustion, cynicism, inadequacy, and study engagement (salmela-aro et al., 2016; salmela-aro & read, 2017). rq 2: what kinds of changes in students’ study burnout, engagement, organised studying, psychological flexibility, and well-being are found at the end of the course intervention, and do the changes differ between the profiles? to our knowledge, there has not been prior person-oriented research comparing the results of act-based interventions on different burnout profiles of students, but räihä, katajavuori, vehkalahti & asikainen 49 | f l r based on research with employees, we set a hypothesis, that the profiles differ in terms of changes in burnout risk and well-being (kinnunen et al., 2019; kinnunen et al., 2020). 4. methods 4.1 participants and the procedure the intervention was conducted as an optional online course (3 ects) that was offered for all students at the university of helsinki. the overall aim of the course was to improve participating students’ ability to identify individual factors related to their well-being and studying, and to learn to apply various psychological skills, such as the processes of psychological flexibility and skills related to organised studying. the online course included eight weeks of act-based exercises alongside time management, and study skills training that took place in online learning environment (moodle) and zoom (see table 1). during the course, students were provided an opportunity to meet the teachers of the course in online meetings in zoom at the beginning, at the middle (week 4), and at the end of the course (week 7). otherwise, the instructions related to the course were provided in the learning environment in text and video formats. alongside individual exercises, students participated in online group meetings, where they were given tasks and prompts to reflect together on the weekly themes. students were required to submit their individual and group learning assignments to the learning environment weekly. the first weeks of the course represented introduction to the content areas of the course, especially the concept of psychological flexibility, and values. students were also given a time management task right at the beginning of the course, in which they followed their time usage for a week and observed for example the workload related to different activities. during these weeks students were for example given exercises, where they reflected the key values in their lives and set a goal for themselves for the course. the objective of weeks three, four and six was for the students to learn, practise and apply psychological skills related to key elements of psychological flexibility: openness to experiences (acceptance of thoughts and feelings, cognitive defusion) and behavioural awareness (present moment awareness, self-as-context). week five consisted of content and exercises related to organised studying, study techniques and the role of health behaviours such as sleep, nutrition and exercising, and aimed for deeper understanding of these concepts in relation to study ability. finally, weeks seven and eight assembled and further deepened the themes of the weeks by reflective exercises related to how to commit to value-based behaviour. in these last weeks of the course students were also given a task to write a 2–4 page-long learning report, in which they were asked to reflect on the meaning and effects of the course on their well-being and study ability, and to give anonymous and constructive peer feedback of two other reports. during the course students were also asked to answer surveys concerning different aspects of their well-being, and study ability at the beginning and at the end of the course and were given general feedback of the results. a further description of the course protocol can be found in an article by asikainen and katajavuori (2021). the online course was organised two times: in autumn 2020 and spring 2021. participants were recruited through convenience sampling via university study program leaders, teachers, university email lists, and social media. the course was available to students from all study programs and stages of study. students were allowed to select their preferred course implementation. if student did not have a preference regarding the timing of the course implementation, they were randomly assigned to either the autumn or spring course. furthermore, to enable small group work, the participants were randomly assigned to groups of 5-6 students. informed consent was obtained from all participants in the online learning environment (moodle). a total of 631 students signed up for the courses, of which 431 gave consent to participate in the study. of the consent-given participants, 363 (84%) completed the course (90,6% female, mean age=27.50, sd= 7.61, median=25.11). räihä, katajavuori, vehkalahti & asikainen 50 | f l r table 1 description of the course modules and content theme assignments & feedback week 1 introduction to the course time management task. completing self-assessment measures. results and feedback of the self-assessments. week 2 values exercises related to acknowledging values. week 3 mindfulness exercises related to recovery, mindful presence, and concentration. individual assignments and online peer discussion. week 4 cognitive defusion exercises related to identifying and differentiating thoughts. individual assignments and online peer discussion. week 5 study skills exercises related to organised study skills. individual assignments and online peer discussion. week 6 acceptance and self-compassion exercises related to facing and accepting difficult or unpleasant thoughts and differentiating self from thoughts. individual assignments and online peer discussion. week 7 committed action exercises related to committing to a value-based behaviour. individual assignments and online peer discussion. reflective learning report. week 8 summary of the course completing self-assessment measures. results and feedback of the selfassessments. peer feedback of learning report. 4.2 measures students responded to questionnaires at the beginning and the end of the intervention course. the dimensions of study-related burnout risk – exhaustion (e.g., ‘i feel overwhelmed by studying’) (exh cronbach’s alpha at the beginning of the course α1 = .78, cronbach’s alpha at the end of the course α2 = .80), cynicism (e.g., ‘i feel a lack of motivation in studying and often think of giving up’) (cyn α1= .87 , α2 = .89), and inadequacy (e.g., ‘i often have feelings of inadequacy when studying’) (inad α1 = .70, α2 =.68) – were measured on a scale of 1–6 (1=completely disagree, 6=completely agree) with the nine-item study burnout inventory (sbi-9) (salmela-aro & read, 2017). study engagement was measured on a scale of 1–6 (1=completely disagree, 6=completely agree) with the nine-item schoolwork engagement inventory (eda) consisting of three dimensions: energy (e.g., ‘when i get up in the morning, i look forward to studying’) (ene α1 = .82, α2 =.85), dedication (e.g., ‘i find studying full of meaning and purpose’) (ded α1 = .88, α2 = .86), and absorption (e.g., ‘time flies when i am studying’) (abs α1 = .81, α2 = .82) (salmela-aro & upadyaya, 2012). the measurement of organised studying was done on a scale of 1–5 (1=completely disagree, 5=completely agree) with the questions of howulearn (α1 = .75, α2 = .76) relating to organised studying (e.g., ‘i organise my study time carefully to make the best use of it’) (parpala & lindblom-ylänne, 2012; modified from the alsi, entwistle et. al 2003). psychological flexibility was measured on a scale of 0–6 (0=strongly disagree, 6=strongly agree) with the 23-item compact, which includes three subscales: openness to experience (oe α1 = .87, α2 = .89) (e.g., ‘i try to stay busy to keep thoughts or feelings from coming’), behavioural awareness (ba α1 = .79, α2 =.84) (e.g., ‘i find it difficult to stay focused on what is happening in the present’), and valued action (va α1 = .87, α2 = .89) (e.g., ‘i can identify the things that really matter to me in life and pursue them’) (francis et al., 2016). the measurement of well-being was done on a scale of 0–5 (0=never, 5=every day) with the 14-item mental health continuum – short form (mhc-sf), which has subscales of emotional wellräihä, katajavuori, vehkalahti & asikainen 51 | f l r being (ewb α1 = .83, α2 = .83) (e.g., ‘during the past month, how often did you feel satisfied with life?’), social well-being (swb α1 =.75, α2 = .83) (e.g., ‘during the past month, how often did you feel that you had something important to contribute to society?’), and psychological well-being (pwb α1 = .78, α2 = .83) (e.g., ‘during the past month, how often did you feel that you had experiences that challenged you to grow and become a better person?’) (keyes, 2002). 4.3 preliminary analysis on eight students’ responses to the questionnaires, there were multiple (>5) missing values for items concerning well-being and study burnout, and these participants were excluded from the data. of the remaining 352 students, 14 had a maximum of one missing value per measure. because no systematic pattern of missingness was detected, missing values were replaced with the individuals’ means of the measures. confirmatory factor analysis (cfa) models with maximum likelihood estimation with robust standard errors and a satorra-bentler-scaled test statistic (mlm) were completed as a preliminary analysis to verify the factor structure of the variables. the fit of the cfa models was based on the comparative fit index (cfi), which indicates a good fit with values above .95, root mean square error of approximation (rmsea) that indexes a good fit with value below .06, and standardised root mean square residual (srmr) that indexes a good fit with values below .08 (hu & bentler, 1999). the results of the cfa showed that the six-factor model of study burnout risk and engagement fit the data (χ2 = 267.619 df = 120, p <.001, cfi = .96, rmsea = .063 (90% c.i. [.053, .074]), srmr = .047). a one-factor model of organised studying was also found to be good (χ2 = 3.906, df = 6, p <.001, cfi = .994, rmsea = .054 (90% c.i. [.000, .133]), srmr = .020). the fit for three-dimensional model of well-being fit the data quite poorly (χ2 = 265.838, df = 74, p <.001, cfi = .880, rmsea = .094 (90% c.i. [.082, .106]), srmr = .062). when examining the modification indices (mi) of the model, we discovered, that the sources of poor fit were related to the items concerning factor that represented social well-being. the covariance between social actualisation representing item 6. (“our society is a good place, or is becoming a better place, for all people”) and social coherence representing item 8. (“the way our society works makes sense to you”) had mi 44.64 with an expected parameter change (epc) of 0.46. in addition, social contribution representing item 4. “you had something important to contribute to society” and social integration representing item 5. “you belonged to a community (like a social group, or your neighbourhood)” had mi 26.47 with an epc of .45. freeing the covariances between these similar variables (6. & 8., and 4. & 5.) resulted in acceptable fit (χ2 = 215.059, df = 72, p <.001, cfi = .911, rmsea = .082 (90% c.i. [.069, .094]), srmr = 0.056). because mhc-sf is a widely used measure of well-being and has demonstrated good psychometric properties across various age groups and nations (iasiello et al., 2022; lamers et al., 2011), and the internal consistency of the variables were good, we decided to use the measure, although the model did not fit the data in an ideal way. the results of cfa concerning three-dimensional model of psychological flexibility of the measure compact indicated also in poor fit (χ2 = 692.868, df = 227, p <.001, cfi = .845, rmsea = .083 (90% c.i. [.076, .090]), srmr = .081). when examining the modification indices (mi) to identify potential areas for improving model fit, we identified pairs of items with high modification indices that had highly similar content. the covariance between item 10. (“i behave in line with my personal values”) and item 21. (“my values are really reflected in my behavior”) representing dimension valued action (va) had mi 67.57 with an expected parameter change (epc) of 0.34. in addition, item 13. (“i am willing to fully experience whatever thoughts, feelings and sensations come up for me, without trying to change or defend against them”) and item 22. (“i can take thoughts and feelings as they come, without attempting to control or avoid them”) of the dimension openness to experiences (oe) had mi 68.45 with an epc of 0.76. freeing the covariances of these highly similar variables resulted in acceptable fit (χ2 = 673.967, df = 225, p <.001, cfi = 0.883, rmsea = 0.072 (90% c.i. [0.065, 0.079]), srmr = 0.078). as this adjustment of the three-factor structure of the items of compact had an räihä, katajavuori, vehkalahti & asikainen 52 | f l r acceptable fit in the data, and the internal consistency of the variables were good, we decided to use the measure. the person-oriented analysis of study-related burnout risk and engagement was conducted using latent profile analysis (lpa). lpa classifies individuals, rather than variables, into homogeneous subpopulations based on underlying classes (collins & lanza, 2009) and gives an opportunity to study intra-individual differences between latent groups of students (hickendorff et al., 2018). lpa was conducted with tidylpa (rosenberg, 2018) on the standardised means of study-related burnout (exh, cyn, inad) and engagement (ene, ded, abs) at the beginning of the intervention course. five fit indices that enable comparison between different models and decision making regarding the number of underlying profiles were used to compare the profile solutions: the akaike information criterion (aic; akaike, 1987), bayesian information criterion (bic; schwarz, 1978), sabic (sample-size-adjusted bic; sclove, 1987), bootstrap likelihood ratio test (blrt; mclachlan,1987), and entropy measure of classification uncertainty (celeux & soromenho, 1996). the solution that best fit the data in accordance with these indicators – and that was also considered reasonable in terms of number of persons in the profiles – and previous research of study-related burnout and engagement was chosen as the final latent profile model. the effect of time, group, and time × group interaction effects were explored by mixed design anova with repeated measures, with profile as the between factor, and bonferroni adjustment for multiple comparisons with alpha level of .05. partial eta squared (η2 p) was used as the effect size. to compare the differences in changes between the profiles, between measurements (1-2) change indicating variable was formed, and the post-hoc comparisons were analysed using dunnett’s t3 test. the data analysis was undertaken by rstudio (version 2023.03.0+386 ‘cherry blossom’) and spss (version 29). räihä, katajavuori, vehkalahti & asikainen 53 | f l r table 2 descriptive statistics n 352 age mean (sd) 27.55 (7.72) age median 25.11 female n (%) 318 (90.3%) male n (%) 33 (9.4%) mean 1 (sd) α1 mean 2 (sd) α2 study burnout risk (sbi-9) exhaustion (exh) 3.51 (1.10) .78 3.14 (1.13) .80 cynicism (cyn) 2.79 (1.37) .87 2.49 (1.32) .89 inadequacy (inad) 4.20 (1.30) .70 3.77 (1.34) .68 study engagement (eda) energy (ene) 3.48 (1.00) .82 3.64 (1.09) .85 dedication (ded) 4.39 (1.01) .88 4.49 (1.02) .86 absorption (abs) 3.48 (1.05) .81 3.70 (1.07) .82 howulearn organised studying (org) 3.08 (.92) .75 3.42 (87) .76 well-being (mhc-sf) emotional well-being (ewb) 3.55 (.92) .83 3.73 (.91) .83 social well-being (swb) 2.70 (.99) .75 2.99 (1.05) .83 psychological well-being (pwb) 3.34 (.88) .78 3.64 (.89) .83 psychological flexibility (compact) openness to experiences (oe) 3.20 (1.23) .87 3.76 (1.21) .89 behavioural awareness (ba) 3.31 (1.22) .79 3.59 (1.26) .84 valued action (va) 4.16 (.95) .87 4.53 (.86) .89 note. mean 1 = mean at the beginning of the course, mean 2 = mean at the end of the course, sd = standard deviation, α1 = cronbach’s alpha of the measure at the first measurement time, α2 = cronbach’s alpha of the measure at the end of the course. 5 results 5.1 the profile solution the first aim of the study was to examine students’ study-related burnout and engagement profiles to find out if there are differences in the changes between the profiles. the different fit indexes of the lpa favoured solutions with six profiles (aic = 4921.90, bic = 5103.50, entropy = .85) (see table 3.). considering the previous findings, the theory of study burnout and engagement (tuominensoini & salmela-aro, 2014; salmela-aro & read, 2017), and the number of students in the profiles, the four-profile solution (aic = 5065.73, bic =5193.23, entropy = .83) was chosen to represent the best fit. the four profiles were named: 1. indifferent, 2. engaged, 3. engaged-inefficacious, and 4. burned-out (see figure 1). in the indifferent profile (n=133, 37.8%), the beginning mean of cynicism was slightly higher, and means of energy, dedication, absorptions slightly lower than the average of all students. in this profile, organised studying, and the dimensions of well-being, and psychological räihä, katajavuori, vehkalahti & asikainen 54 | f l r flexibility were also close to average (see table 4). in the engaged profile (n=102, 29%), students’ energy, dedication, and absorption were higher, and exhaustion, cynicism, and inadequacy lower than the average of all students. in this profile, the beginning means of well-being, psychological flexibility, and organised studying were the highest (see table 4). the engaged-inefficacious profile (n=48, 13.6%) represented students with slightly below average cynicism and high means of energy, dedication, absorption, exhaustion, and inadequacy. the beginning means of well-being, organised studying, and psychological flexibility were close to the average of all students, except for behavioural awareness, which had a slightly lower beginning mean than the average of the whole sample of students. in the burned-out profile (n=69, 19.6%), students had high means of exhaustion, cynicism, and inadequacy and low means of energy, dedication, and absorption. in this profile, the means of well-being, psychological flexibility, and organised study skills were the lowest of the profiles. table 3 fit statistics for the profile solutions profile n aic bic sabic entropy blrt blrt p prob. min-max group sizes 1 6011.59 6057.95 6019.88 1 1 352 2 5331.99 5405.40 5345.13 .87 693.60 .010 .95-.97 195, 157 3 5136.01 5236.47 5153.98 .83 209.98 .010 .91-.94 151, 128, 73 4 5065.73 5193.23 5088.54 .83 84.29 .010 .83-.94 133, 102, 48, 69 5 5002.73 5157.27 5030.38 .81 76.99 .010 .77-.94 85, 88, 75, 44, 60 6 4921.90 5103.50 4954.39 .85 94.83 .010 .84-.94 85, 88, 73, 7, 26,73 note. aic = akaike information criterion; bic = bayesian information criterion; sabic = samplesize-adjusted bic; entropy = a measure of classification uncertainty; blrt: bootstrapped likelihood test; blrt-p: p-value for the bootstrapped likelihood ratio test; prob. min-max = minimum and maximum of the diagonal of the average latent class probabilities for most likely class memberships, by assigned class. figure 1. standardised means of the dimensions of study burnout and engagement at the beginning of the course intervention in the four-profile solution räihä, katajavuori, vehkalahti & asikainen 55 | f l r table 4 means and standard deviations of the profiles at the beginning and at the end of the course indifferent (n=133) engaged (n=102) engagedinefficacious (n=48) burned-out (n=69) m1 (sd) m2 (sd) m1 (sd) m2 (sd) m1 (sd) m2 (sd) m1 (sd) m2 (sd) exh 3.49 (.94) 3.16 (.96) 2.70 (.81) 2.44 (.84) 4.23 (.83) 3.81 (1.10) 4.28 (1.05) 3.68 (1.23) cyn 3.00 (1.01) 2.66 (1.08) 1.50 (.54) 1.42 (.59) 2.53 (1.26) 2.41 (1.29) 4.49 (.83) 3.78 (1.27) inad 4.27 (.79) 3.87 (1.04) 2.69 (.80) 2.56 (1.04) 5.39 (.59) 4.51 (1.00) 5.45 (.60) 4.86 (1.02) ene 3.27 (.53) 3.44 (.79) 4.42 (.54) 4.45 (.80) 4.10 (.51) 4.08 (.96) 2.09 (.62) 2.52 (.91) ded 4.13 (.62) 4.24 (.84) 5.23 (.55) 5.24 (.66) 5.10 (.56) 4.92 (.79) 3.14 (.84) 3.55 (.94) abs 3.21 (.61) 3.54 (.82) 4.23 (.74) 4.32 (.87) 4.37 (.82) 4.31 (.84) 2.29 (.81) 2.68 (1.00) org 2.99 (.83) 3.34 (.78) 3.69 (.73) 3.94 (.65) 3.02 (.78) 3.46 (.79) 2.37 (.86) 2.80 (.91) ewb 3.51 (.81) 3.67 (.80) 4.05 (.65) 4.25 (.54) 3.65 (.88) 3.74 (.76) 2.81 (.99) 3.07 (1.16) swb 2.66 (.90) 2.85 (.93) 3.19 (.88) 3.60 (.81) 2.66 (.98) 3.00 (.99) 2.09 (.97) 2.34 (1.17) pwb 3.27 (.76) 3.54 (.78) 3.88 (.68) 4.17 (.55) 3.34 (.94) 3.67 (.84) 2.69 (.85) 3.05 (1.06) oe 3.08 (1.08) 3.60 (1.04) 3.92 (1.13) 4.53 (.95) 3.19 (1.24) 3.65 (1.26) 2.38 (1.02) 3.00 (1.23) ba 3.24 (1.17) 3.48 (1.26) 3.83 (1.11) 4.26 (1.06) 2.99 (1.34) 3.40 (1.18) 2.88 (1.12) 2.94 (1.15) va 4.05 (.82) 4.43 (.75) 4.79 (.63) 5.07 (.56) 4.10 (.97) 4.46 (.90) 3.48 (1.00) 3.99 (.97) note. exh = exhaustion, cyn = cynicism, inad = inadequacy, ene = energy, ded = dedication, abs = absorption, org = organised studying, ewb = emotional well-being, swb = social wellbeing, pwb = psychological well-being, oe = openness to experience, ba = behavioural awareness, va = valued action, m1 = mean at the beginning of the course, m2 = mean at the end of the course, sd = standard deviation. räihä, katajavuori, vehkalahti & asikainen 56 | f l r figure 2. mean changes in the profiles between the first and second measurement time note. exh = exhaustion, cyn = cynicism, inad = inadequacy, ene = energy, ded = dedication, abs = absorption, org = organised studying, ewb = emotional well-being, swb = social wellbeing, pwb = psychological well-being, oe = openness to experience, ba = behavioural awareness, va = valued action. 5.2 changes within and between the profiles the second aim of this study was to explore the differences in the changes between the beginning and end measurements between and within the profiles. when studying the changes at the end of the intervention course, the results of mixed anova with repeated measures indicated, that the main effect of time was significant on all observed variables: study burnout decreased, and engagement, organised studying, well-being, and psychological flexibility increased statistically significantly in the whole sample of students (see table 5). the largest effect size was found in the increase of openness to experiences (η2 p=.28). no statistically significant time x group effects were to be seen in the changes of exhaustion, organised studying, dimensions of well-being or psychological flexibility the profiles did differ statistically significantly in the changes of cynicism (f(3, 348) = 8.08, p <.001), inadequacy (f(3, 348) = 7.86, p <.001), energy (f(3, 348) = 4.48, p = .004), absorption (f(3, 348) = 5.40, p <.001) and dedication (f(3, 348) = 6.52, p <.001). the largest effect size of the interaction of profile and time was found in changes of cynicism (η2 p=.07). the post-hoc test indicated that the change in cynicism was statistically significantly different between indifferent and engaged (p=.028), engaged and burned-out (p <.001), and engagedinefficacious and burned-out (p=.019) profiles. the pairwise comparisons of the first and second measurement within the profiles showed, that cynicism decreased statistically significantly in the indifferent (f(1, 348) = 20.72, p <.001, η2 p=.06) and burned-out profiles (f(1, 348) = 45.31, p <.001, η2 p=.12). cynicism decreased slightly also in engaged and engaged-inefficacious profiles, but the changes were not statistically significant. according to post-hoc test, changes in inadequacy differed statistically significantly between the profiles indifferent and engaged-inefficacious (p=.014), engaged räihä, katajavuori, vehkalahti & asikainen 57 | f l r and engaged-inefficacious (p <.001), and engaged and burned-out (p=.008). the pairwise comparisons of the first and second measurement within the profiles showed, that inadequacy decreased statistically significantly in the indifferent (f(1, 348) = 24.61, p <.001, η2 p=.07), engaged-inefficacious (f(1, 348) = 42.03, p <.001, η2 p=.11), and burned-out profiles (f(1, 348) =27.19, p <.001, η2 p=.07). the decrease in inadequacy was not statistically significant in the engaged profile, and was greater in the engagedinefficacious profile, than in the indifferent profile. changes in energy differed statistically significantly between the profiles engaged and burnedout (p=.012) and engaged-inefficacious and burned-out (p=.019). the pairwise comparisons of the first and second measurement within the profiles showed, that energy increased statistically significantly in the indifferent (f(1, 348) = 6.64, p=.010, η2 p=.02) and burned-out (f(1, 348) = 20.24, p <.001, η2 p=.06) profiles. energy stayed the same in the engaged, and decreased in the engaged-inefficacious profile, but the changes were not statistically significant. changes in dedication differed statistically significantly between the profiles engaged and burned-out (p=.015) and engaged-inefficacious and burned-out (p <.001). the pairwise comparisons of the first and second measurement within the profiles showed, that dedication increased statistically significantly in the burned-out profile (f(1, 348) = 20.13, p<.001, η2 p=.06). dedication increased marginally in the indifferent and engaged profile, and decreased in the engaged-inefficacious profile, but the changes were not statistically significant. according to post-hoc test, changes in absorption differed statistically significantly between the profiles engaged-inefficacious and indifferent (p=.023), and burned-out and engaged-inefficacious (p=.014). the pairwise comparisons of the first and second measurement within the profiles showed, that absorption increased statistically significantly in indifferent (f(1, 348) = 25.25, p <.001, η2 p=.07) and burned-out (f(1, 348) = 19.36, p <.001, η2 p=.05) profiles. there was a marginal and statistically insignificant increase in absorption in the engaged and decrease in engaged-inefficacious profile. please see figure 2 for a visualisation of the mean changes within profiles. räihä, katajavuori, vehkalahti & asikainen 58 | f l r table 5 mixed anova with repeated measures and the post-hoc comparisons of the mean changes between the profiles note. exh = exhaustion, cyn = cynicism, inad = inadequacy, ene = energy, ded = dedication, abs = absorption, org = organised studying, ewb = emotional well-being, swb = social well-being, pwb = psychological well-being, oe = openness to experience, ba = behavioural awareness, va = valued action, η2 p = partial eta squared. group = profile, a dunnett t3 post-hoc comparisons of mean changes between the profiles, profile 1 = indifferent, 2 = engaged, 3 = engaged-inefficacious, 4 = burned-out, significance level *p < .05, **p < .01, ***p < .001 between subjects within-subjects effects group time time x group df f p η2 p error df f p η2 p df f p η2 p error (time) post-hoc comparisons a exh 3 50.79 <.001 .30 348 1 60.67 <.001 .15 3 2.24 .083 .02 348 cyn 3 131.97 <.001 .53 348 1 38.59 <.001 .10 3 8.08 <.001 .07 348 1≠2*; 2≠4***; 3≠4* inad 3 187.42 <.001 .62 348 1 86.16 <.001 .20 3 7.86 <.001 .06 348 1≠3*; 2≠3***; 2≠4** ene 3 196.46 <.001 .63 348 1 11.30 <.001 .03 3 4.48 .004 .04 348 4≠2,3*** ded 3 145.48 <.001 .56 348 1 4.21 .041 .01 3 6.52 <.001 .05 348 2≠4*; 3≠4*** abs 3 109.50 <.001 .49 348 1 19.32 <.001 .05 3 5.40 .001 .04 348 1≠3*; 3≠4* org 3 42.94 <.001 .27 348 1 79.43 <.001 .19 3 1.27 .285 .01 348 ewb 3 35.44 <.001 .23 348 1 24.72 <.001 .07 3 .78 .507 .01 348 swb 3 26.57 <.001 .19 348 1 48.85 <.001 .12 3 1.75 .157 .02 348 pwb 3 37.34 <.001 .24 348 1 64.54 <.001 .16 3 .32 .812 .01 348 oe 3 34.17 <.001 .22 348 1 114.39 <.001 .28 3 .48 .698 .00 348 ba 3 18.93 <.001 .14 348 1 20.53 <.001 .06 3 1.90 .129 .02 348 va 3 41.91 <.001 .27 348 1 73.80 <.001 .18 3 1.20 .310 .01 348 räihä, katajavuori, vehkalahti & asikainen 59 | f l r 6. discussion the aim of this study was to compare the effects of a study-integrated act-based course intervention on different aspects of university students’ well-being and study ability in latent groups of students that had different levels of study burnout and engagement at the beginning of the course. as we applied a person-oriented approach to explore the different dimensions of students’ study burnout and engagement at the beginning of the course, four profiles were found: indifferent, engaged, engagedinefficacious, and burned-out. when studying the changes at the end of the course, the results of mixed anova with repeated measures indicated, that the main effect of time was significant on all observed variables: over the entire course, students' burnout risk decreased significantly, while their engagement, organised studying, well-being, and psychological flexibility increased significantly. when examining the changes across different profiles, it was found that the changes in exhaustion, psychological flexibility, well-being, and organised studying were similar across all profiles. however, differences were observed between the profiles in terms of changes in cynicism, feelings of inadequacy, and engagement (including energy, dedication, and absorption). the found four profiles were in line with earlier findings of university students’ engagement and burnout profiles (salmela-aro & reid, 2017). the effects of course level changes were also in line with results of previous studies, that have showed that act-based interventions can reduce students’ burnout risk and increase study engagement, organised studying, psychological flexibility, and wellbeing (frögéli et al., 2016; grégoire et al., 2018; howell & passmore, 2019; katajavuori et al., 2021; räihä et al., 2024). since no previous research has applied a person-oriented approach to act-based interventions with students, this study provides new knowledge about the varying effects of these interventions on students’ representing different profiles of burnout and engagement. considering the differences in the changes of study burnout and engagement, the burned-out profile had the most to gain in the beginning of the course and benefitted the most according to the decrease of cynicism. respectively, in the engaged profile, the means of the dimensions of burnout were below and engagement above the average already at the beginning of the course. in the engaged profile study engagement stayed close the same at the end of the course, but there was a significant reduction in the exhaustion also in this profile. in the indifferent profile students’ study engagement increased and burnout risk decreased, largest decrease being in the mean of inadequacy. an interesting combination of features was to be seen in the engaged-inefficacious profile, which represented students that were both engaged and strained from studies. in this profile, the mean of inadequacy decreased the most, alongside decrease in exhaustion. furthermore, there was a small decrease in the dimensions of study engagement. although the decreases in engagement were not statistically significant, the finding is interesting, as engaged-inefficacious was the only profile in which engagement decreased. the reduction in engagement together with significant increase in the dimensions of well-being, psychological flexibility, and organised studying in the engaged-inefficacious profile could reflect the previous notions that study engagement has also its ‘dark side’ (salmela-aro et al., 2016; salmela-aro & reid, 2017). these changes may also suggest that, as a result of increased psychological flexibility, students in this profile have become more aware of their values and behavior, enabling them to better utilise personally relevant resources that help alleviate the strain of studying. as this study was done in autumn 2020 and spring 2021 during the covid-19 pandemic, distance learning and social distancing as contextual features might have added the discrepancy of students’ recourses and demands on some of the students. in a 2021 study of the finnish institute for health and welfare, 70% of students in higher education reported that the challenges with their studying had increased during the pandemic, and almost half of the students had experienced an increase in the workload required for their studies (parikka et al., 2021). however, academic workload has also been shown to be positively related not only to burnout risk but also to study engagement (olson et al., 2023), and not all students’ experienced strain from distance learning, as some students reported that studying became easier during the pandemic (asikainen & katajavuori, 2023; parikka et al., 2021). these differences can be related to different contextual factors such as social support (tindle et al., 2022), but also to different competencies to regulate study-related activities (parpala et al., 2021), and the level of räihä, katajavuori, vehkalahti & asikainen 60 | f l r students’ psychological flexibility (asikainen & katajavuori, 2023). the findings of this present study that organised studying, psychological flexibility, and different aspects of well-being increased alongside decrease in exhaustion in all burnout and engagement profiles indicate effectiveness of actbased course to enhance different kind of students’ study abilities alongside well-being. 6.1 implications for practice the findings of this study show promise of study-integrated act-based courses as an effective way to support students’ well-being and personal resources, such as psychological flexibility and organised studying. the opportunity to enhance psychological flexibility together with study-related skills can help students, for example, to balance the strain from the demands of studying, and to make use of available resources by reducing avoidant behaviour (tindle et al., 2022). as there are also previous findings that act-based online interventions can reduce students’ distress (levin et al., 2017; räsänen et al., 2016), a course that combines practicing skills of psychological flexibility and organised studying can be beneficial for a variety of participants – both for students that report elevated levels of study burnout risk and for the ones who are doing well – to prevent future distress. the importance of the findings of this study, that showed increases in students’ emotional, psychological, and social wellbeing in all the profiles are further emphasised, with the notion that higher levels of well-being can also work as mental illbeing preventing factor (keyes, 2002). therefore, aiming to promote students’ wellbeing related factors alongside targeting the well-being depleting aspects such as study burnout is highly important. this also indicates that providing this kind of course intervention as early as possible in the studies could be important from the point of view of preventing future challenges. 6.2 limitations and future directions although the results of this study provide new knowledge of the effectiveness of act-based interventions for different burnout and engagement profiles representing university students, several limitations should be mentioned. first, the study sample was quite homogeneous, as the participants of the intervention course were mainly young females. to enhance generalisability, future research should aim to investigate the effects of the course using randomised and representative samples. second, as the data was collected during covid-19 pandemic, this could have influenced the results of the study in several ways. higher education students overall mental well-being has been found to be lower during the covid-19 pandemic, than before the outbreak (sarasjärvi et al., 2022), alongside higher levels of study burnout risk, and lower levels of engagement (salmela-aro et al., 2022b). additionally, given the nature of the course, it is reasonable to assume that students experiencing study-related burden were particularly inclined to enrol. however, as 84% of the consent-given participants completed the course, the dropout rate (16%) was not higher than the mean dropout rates in act-interventions (15,8%) (ong et al., 2018). the online delivery of the course made it possible to offer the course to all willing students, alongside providing peer-support and opportunities to connect with other students’ during social distancing in the form of online group discussions. as lack of interaction and emotional support has been shown to be associated with university students’ negative mental health trajectories (elmer et al., 2020), and to mediate the relation between burnout and psychological well-being (rehman et al., 2020), further studies should deepen the knowledge related to the quality of the peer groups in relation to the effects of the course. third, as the fit of the multidimensional measures of psychological flexibility and well-being were found relatively poor in this sample, the results concerning the changes in different dimensions of these measures should be interpreted with caution. a poor fit of the 23-item compact used in this study has been reported also in other studies (hsu et al., 2023; tynan et al., 2022; trindade et al., 2022), and resulted in validation of shorter 18-, 15-, and 10-item versions of the measure. therefore, a validation of a shorted version of compact in a representative finnish sample would be advisable. in this data, the poor fit of the widely used measure of well-being mhc-sf was found to be related räihä, katajavuori, vehkalahti & asikainen 61 | f l r especially to the items that represented social well-being. as social well-being refers to “the appraisal of one's circumstance and functioning in society” (keyes, 1998), students’ evaluations of social actualisation, coherence, contribution, and integration could have been affected by the different contextual factors related to covid-19 pandemic, such as uncertainty, social distancing and distance learning that occurred during the times of the measurements. furthermore, the significance of this study's findings, which demonstrated a significant increase in students' openness to experiences across all profiles, is emphasised by the recognition that openness to experiences, along with the desire and effort to grow, are fundamental components of social actualization (keyes, 1998). this also underscores the need for further research regarding longitudinal trajectories of social well-being among course participating students. fourth limitation, that should be mentioned is the sample size of the data. for example, in their review of lpa applying studies, spurk and others (2020) conclude, that based on past research a sample size around 500 seems reasonable to detect a correct number of latent profiles. however, in a simulation study of tein and others (2013) it was found that a sample size of 1000 did not show higher statistical power than the conditions with sample size of 250 or 500. nevertheless, as the size of some of the identified profiles can be insufficient for additional analyses, further person-oriented studies of actbased interventions should aim for larger sample sizes. the fifth limitation that should be mentioned is that, even though by applying person-oriented approaches it is possible to reach underlying phenomena of different subgroups of students (bergman & lundh, 2015), the research of changes in these subgroups does not reach the variety of individual changes, or lack of changes, which occure during the course. the results of this study also indicated differences in the changes of study engagement and burnout risk between the profiles. even though these differences were related to the changes in the engaged profile with low levels of burnout risk and high levels of engagement, which stayed the same at the end of the intervention course, the changes in the engaged-inefficacious profile give an indication of the importance of identifying students’ different needs at the beginning, and during of the intervention. further studies should apply qualitative data to deepen the knowledge of different outcomes and related factors. furthermore, as these changes are bind not only to students’ individual needs but also the context, further intraand inter-individual process-based research is important (hayes et al., 2022). as changes in psychological flexibility and well-being can take time (kinnunen et al., 2019), there is also a need to study different longitudinal trajectories of the changes. as it has been shown that continuity of practice is related to higher benefits with different burnout profiles representing act-based intervention participants (kinnunen et al., 2019), further research of factors related to learning and intensity of practicing could also deepen the knowledge of the relationship of changes in students’ behaviour, and well-being during and after the course. 7. conclusions the results of this study support the notion that study burnout and engagement, as multidimensional constructs, are not simply opposite ends of a continuum but rather distinct phenomena that can vary in emphasis across different dimensions. the results also indicated that a study-integrated act-based online course can be effective in reducing exhaustion, and enhancing different dimensions of psychological flexibility, well-being, and organised studying of different study burnout and engagement representing latent groups of university students. based on these beneficial effects regardless of the initial burnout risk, it can be concluded, that applying and further developing actbased course interventions more widely as part of university studies could be an effective way to support student well-being – even before there are problems. further research is needed on the long-term effects of the intervention course on students’ well-being and on the individual and qualitative differences in students’ behaviour changes and learning during the intervention course. räihä, katajavuori, vehkalahti & asikainen 62 | f l r ethical approval the research has been conducted in accordance with the declaration of helsinki guidelines and followed finnish national board on research integrity's ethical principles for research with human participants. informed consent was obtained from all individual participants included in the study. key points university students participated in a study-integrated acceptance and commitment (act) based intervention course. students (n=352) represented four different study burnout and engagement profiles at the beginning of the course. well-being, psychological flexibility, and organised studying increased, while exhaustion decreased across all profiles. the profiles differed in the changes observed in cynicism, inadequacy, and engagement. act-based courses can be an effective way to enhance university students’ well-being and study ability, regardless of their initial burnout level. references akaike, h. (1987). factor analysis and aic. psychometrika 52, 317–332. https://doi.org/10.1007/bf02294359 asikainen, h., & katajavuori, n. (2021). development of a web-based intervention course to promote students' well-being and studying in universities: protocol for an experimental study design. jmir research protocols, 10(3), article 23613. https://doi.org/10.2196/23613 asikainen, h., & katajavuori, n. (2023). exhausting and difficult or easy: the association between psychological flexibility and study related burnout and experiences of studying during the pandemic. frontiers in education, 8, article p1215549. https://doi.org/10.3389/feduc.2023.1215549 asikainen, h., nieminen, j. h., häsä, j., & katajavuori, n. (2022). university students’ interest and burnout profiles and their relation to approaches to learning and achievement. learning and individual differences, 93. https://doi.org/10.1016/j.lindif.2021.102105 auerbach r.p., mortier p., bruffaerts, r., alonso, j., benjet, c., cuijpers, p., demyttenaere, k., ebert, d.d., green, j.g., hasking, p., murray, e., nock, m.k., pinder-amaker, s., sampson, n.a., stein, d.j., vilagut g., zaslavsky, a.m., & kessler, r.c. (2018). who world mental health surveys international college student project: prevalence and distribution of mental disorders. journal of abnormal psychology, 127(7), 623-638. https://doi.org/10.1037/abn0000362 bakker, a. b., demerouti, e., & sanz-vergel, a. (2023). job demands–resources theory: ten years later. annual review of organizational psychology and organizational behavior, 10, 25-53. https://doi.org/10.1146/annurev-orgpsych-120920-053933 bergman, l. r., & lundh, l.-g. (2015). introduction: the person-oriented approach: roots and roads to the future. journal for person-oriented research, 1(1–2), 1–6. https://doi.org/10.17505/jpor.2015.01 https://doi.org/10.1007/bf02294359 https://doi.org/10.2196/23613 https://doi.org/10.3389/feduc.2023.1215549 https://doi.org/10.1016/j.lindif.2021.102105 https://doi.org/10.1037/abn0000362 https://doi.org/10.1146/annurev-orgpsych-120920-053933 https://doi.org/10.17505/jpor.2015.01 räihä, katajavuori, vehkalahti & asikainen 63 | f l r bond, f. w., lloyd, j., & guenole, n. (2013). the work-related acceptance and action questionnaire: initial psychometric findings and their implications for measuring psychological flexibility in specific contexts. journal of occupational and organizational psychology, 86(3), 331–347. https://doi.org/10.1111/joop.12001 celeux, g., & soromenho, g. (1996). an entropy criterion for assessing the number of clusters in a mixture model. journal of classification, 13, 195–212. https://doi.org/10.1007/bf01246098 collins, l. m., & lanza, s. t. (2009). latent class and latent transition analysis: with applications in the social, behavioral, and health sciences (vol. 718). john wiley & sons. davis, s. k., & hadwin, a. f. (2021). exploring differences in psychological well-being and self-regulated learning in university student success. frontline learning research, 9(1), 30-43. https://doi.org/10.14786/flr.v9i1.581 dodge, r., daly, a. p., huyton, j., & sanders, l. d. (2012). the challenge of defining wellbeing. international journal of wellbeing, 2(3), 222–235. https://doi.org/10.5502/ijw.v2i3.4 elmer, t., mepham, k., & stadtfeld, c. (2020). students under lockdown: comparisons of students’ social networks and mental health before and during the covid-19 crisis in switzerland. plos one, 15(7), e0236337. https://doi.org/10.1371/journal.pone.0236337 entwistle, n., & mccune, v. (2004). the conceptual bases of study strategy inventories. educational psychology review, 16(4), 325-345. https://doi.org/10.1007/s10648-004-0003-0 entwistle, n., mccune, v. & hounsell, j. (2003). investigating ways of enhancing university teaching-learning environments: measuring students’ approaches to studying and perceptions of teaching. in e. de corte, l. verschaffel, n. entwistle, & j. van merriënboer (eds.), powerful learning environments: unravelling basic components and dimensions (pp. 89-107). oxford: elsevier science. fledderus, m., bohlmeijer, e. t., smit, f., & westerhof, g. j. (2010). mental health promotion as a new goal in public mental health care: a randomized controlled trial of an intervention enhancing psychological flexibility. american journal of public health, 100(12), 2372-2372. https://doi.org/10.2105/ajph.2010.196196 francis, a. w., dawson, d. l., & golijani-moghaddam, n. (2016). the development and validation of the comprehensive assessment of acceptance and commitment therapy processes (compact). journal of contextual behavioral science, 5(3), 134-145. https://doi.org/10.1016/j.jcbs.2016.05.003 frögéli, e., djordjevic, a., rudman, a., livheim, f., & gustavsson, p. (2016). a randomized controlled pilot trial of acceptance and commitment training (act) for preventing stress-related ill health among future nurses. anxiety, stress and coping, 29(2), 202–218. http://dx.doi.org/10.1080/10615806.2015.1025765 grégoire, s., lachance, l., bouffard, t., & dionne, f. (2018). the use of acceptance and commitment therapy to promote mental health and school engagement in university students: a multisite randomized controlled trial. behavior therapy, 49(3), 360–372. https://doi.org/10.1016/j.beth.2017.10.003 hailikari, t., nieminen, j., & asikainen, h. (2022). the ability of psychological flexibility to predict study success and its relations to cognitive attributional strategies and academic emotions. educational psychology, 42(5), 626–643. https://doi.org/10.1080/01443410.2022.2059652 hayes, s. c., ciarrochi, j., hofmann, s. g., chin, f., & sahdra, b. (2022). evolving an idionomic approach to processes of change: towards a unified personalized science of human improvement. behaviour research and therapy, 156, article 104155. https://doi.org/10.1016/j.brat.2022.104155 hayes, s. c., luoma, j. b., bond, f. w., masuda, a., & lillis, j. (2006). acceptance and commitment therapy: model, processes, and outcomes. behaviour research and therapy, 44(1), 1 –25. https://doi.org/10.1016/j.brat.2005.06.006 https://doi.org/10.1111/joop.12001 https://doi.org/10.1007/bf01246098 https://doi.org/10.14786/flr.v9i1.581 https://doi.org/10.5502/ijw.v2i3.4 https://doi.org/10.1371/journal.pone.0236337 https://doi.org/10.1007/s10648-004-0003-0 https://doi.org/10.2105/ajph.2010.196196 https://doi.org/10.1016/j.jcbs.2016.05.003 http://dx.doi.org/10.1080/10615806.2015.1025765 https://doi.org/10.1016/j.beth.2017.10.003 https://doi.org/10.1080/01443410.2022.2059652 https://doi.org/10.1016/j.brat.2022.104155 https://doi.org/10.1016/j.brat.2005.06.006 räihä, katajavuori, vehkalahti & asikainen 64 | f l r hayes, s. c., pistorello, j., & levin, m. e. (2012). acceptance and commitment therapy as a unified model of behavior change. the counseling psychologist, 40(7), 976–1002. https://doi.org/10.1177/0011000012460836 hayes, s.c., villatte, m., levin, m., hildebrandt, m. (2011). open, aware, and active: contextual approaches as an emerging trend in the behavioral and cognitive therapies. annual review of clinical psychology, 7(1), 141-68. https://dx.doi.org/10.1146/annurev-clinpsy-032210-104449 hickendorff, m., edelsbrunner, p. a., mcmullen, j., schneider, m., & trezise, k. (2018). informative tools for characterizing individual differences in learning: latent class, latent profile, and latent transition analysis. learning and individual differences, 66, 4-15. https://doi.org/10.1016/j.lindif.2017.11.001 honkonen, t., ahola, k., pertovaara, m., isometsä, e., kalimo, r., nykyri, e., aromaa, a. & lönnqvist, j. (2006). the association between burnout and physical illness in the general population—results from the finnish health 2000 study. journal of psychosomatic research, 61(1), 59-66. https://doi.org/10.1016/j.jpsychores.2005.10.002 howell, a. j. (2009). flourishing: achievement-related correlates of students’ well-being. the journal of positive psychology, 4(1), 1–13. https://doi.org/10.1080/17439760802043459 howell, a. j., & passmore, h. a. (2019). acceptance and commitment training (act) as a positive psychological intervention: a systematic review and initial meta-analysis regarding act’s role in wellbeing promotion among university students. journal of happiness studies, 20, 1995-2010. https://doi.org/10.1007/s10902-018-0027-7 hsu, t., hoffman, l., & thomas, e. (2023). confirmatory measurement modeling and longitudinal invariance of the compact-15: a short-form assessment of psychological flexibility. psychological assessment, 35, 430-442. https://doi.org/10.1037/pas0001214 hu, l. t., & bentler, p. m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives. structural equation modeling: a multidisciplinary journal, 6(1), 1-55. https://doi.org/10.1080/10705519909540118 iasiello, m., van agteren, j., schotanus-dijkstra, m., lo, l., fassnacht, d. b., & westerhof, g. j. (2022). assessing mental wellbeing using the mental health continuum—short form: a systematic review and meta-analytic structural equation modelling. clinical psychology: science and practice, 29(4), 442– 456. https://doi.org/10.1037/cps0000074 kachel, t., huber, a., strecker, c., höge, t., & höfer, s. (2020). development of cynicism in medical students: exploring the role of signature character strengths and well-being. frontiers in psychology, 11, 328. https://doi.org/10.3389/fpsyg.2020.00328 kashdan, t. b., & rottenberg, j. (2010). psychological flexibility as a fundamental aspect of health. clinical psychology review, 30(4), 865–878. https://doi.org/10.1016/j.cpr.2010.03.001 katajavuori, n., vehkalahti, k., & asikainen, h. (2021). promoting university students’ well-being and studying with an acceptance and commitment therapy (act)-based intervention. current psychology, 42, 4900–4912. https://doi.org/10.1007/s12144-021-01837-x ketonen, e. e., haarala-muhonen, a., hirsto, l., hänninen, j. j., wähälä, k., & lonka, k. (2016). am i in the right place? academic engagement and study success during the first years at university. learning and individual differences, 51, 141–148. https://doi.org/10.1016/j.lindif.2016.08.017 keyes, c. l. m. (1998). social well-being. social psychology quarterly, 61(2), 121–140. https://doi.org/10.2307/2787065 keyes, c.l.m. (2002). the mental health continuum: from languishing to flourishing in life. journal of health and social research, 43(2), 207-222. https://doi.org/10.2307/3090197 https://doi.org/10.1177/0011000012460836 https://dx.doi.org/10.1146/annurev-clinpsy-032210-104449 https://doi.org/10.1016/j.lindif.2017.11.001 https://doi.org/10.1016/j.jpsychores.2005.10.002 https://doi.org/10.1080/17439760802043459 https://doi.org/10.1007/s10902-018-0027-7 https://doi.org/10.1037/pas0001214 https://doi.org/10.1080/10705519909540118 https://doi.org/10.1037/cps0000074 https://doi.org/10.3389/fpsyg.2020.00328 https://doi.org/10.1016/j.cpr.2010.03.001 https://doi.org/10.1007/s12144-021-01837-x https://doi.org/10.1016/j.lindif.2016.08.017 https://doi.org/10.2307/2787065 https://doi.org/10.2307/3090197 räihä, katajavuori, vehkalahti & asikainen 65 | f l r kinnunen, s. m., puolakanaho, a., mäkikangas, a., tolvanen, a., & lappalainen, r. (2020). does a mindfulness-, acceptance-, and value-based intervention for burnout have long-term effects on different levels of subjective well-being? international journal of stress management, 27(1), 82–87. https://dx.doi.org/10.1037/str0000132 kinnunen, s. m., puolakanaho, a., tolvanen, a., mäkikangas, a., & lappalainen, r. (2019). does mindfulness-, acceptance-, and value-based intervention alleviate burnout? a person-centered approach. international journal of stress management, 26(1), 89–101. https://doi.org/10.1037/str0000095 kunttu, k., pesonen, t., & saari, j. (2016). student health survey 2016: a national survey among finnish university students. research publications of the finnish student health service, 48. https://www.yths.fi/app/uploads/2020/03/kott_2016_eng.pdf lamers, s. m., westerhof, g. j., bohlmeijer, e. t., ten klooster, p. m., & keyes, c. l. (2011). evaluating the psychometric properties of the mental health continuum‐short form (mhc‐sf). journal of clinical psychology, 67(1), 99-110. https://doi.org/10.1002/jclp.20741 laursen, b., & hoff, e. (2006). person-centered and variable-centered approaches to longitudinal data author. merrill-palmer quarterly, 52(3), 377–389. https://dx.doi.org/10.1353/mpq.2006.0029 levin, m. e., haeger, j. a., pierce, b. g., & twohig, m. p. (2017). web-based acceptance and commitment therapy for mental health problems in college students: a randomized controlled trial. behavior modification, 41(1), 141-162. https://doi.org/10.1177/0145445516659645 madigan, d. j., & curran, t. (2021). does burnout affect academic achievement? a meta-analysis of over 100,000 students. educational psychology review, 33(2), 387–405. https://doi.org/10.1007/s10648020-09533-1 madigan, d. j., kim, l. e., & glandorf, h. l. (2023). interventions to reduce burnout in students: a systematic review and meta-analysis. european journal of psychology of education, 1-27. https://doi.org/10.1007/s10212-023-00731-3 mäkikangas, a., & kinnunen, u. (2016). the person-oriented approach to burnout: a systematic review. burnout research, 3(1), 11–23. https://doi.org/10.1016/j.burn.2015.12.002 maslach, c., schaufeli, w. b., & leiter, m. p. (2001). job burnout. annual review of psychology, 52, 397– 422. https://doi.org/10.1146/annurev.psych.52.1.397 mclachlan, g. j. (1987). on bootstrapping the likelihood ratio test statistic for the number of components in a normal mixture. journal of the royal statistical society, 36(3), 318–324. https://doi.org/10.2307/2347790 olson n, oberhoffer-fritz r, reiner b, schulz t (2023). study related factors associated with study engagement and student burnout among german university students. frontiers in public health, 11. https://doi.org/10.3389/fpubh.2023.1168264 ong, c. w., lee, e. b., & twohig, m. p. (2018). a meta-analysis of dropout rates in acceptance and commitment therapy. behaviour research and therapy, 104, 14-33. https://doi.org/10.1016/j.brat.2018.02.004 onwezen, m. c., van veldhoven, m. j. p. m., & biron, m. (2014). the role of psychological flexibility in the demands–exhaustion–performance relationship. european journal of work and organizational psychology, 23(2), 163–176. https://doi.org/10.1080/1359432x.2012.742242 pagnin, d., & de queiroz, v. (2015). influence of burnout and sleep difficulties on the quality of life among medical students. springerplus, 4, 1-7. https://doi.org/10.1186/s40064-015-1477-6 parikka s, holm n, ikonen j, koskela t, kilpeläinen h & lundqvist a (2021). the finnish student health and wellbeing survey kott. kott 2021 survey's weighted distributions https://www.terveytemme.fi/kott/taulukot/painotetut_jakaumat_en.html https://dx.doi.org/10.1037/str0000132 https://doi.org/10.1037/str0000095 https://www.yths.fi/app/uploads/2020/03/kott_2016_eng.pdf https://doi.org/10.1002/jclp.20741 https://dx.doi.org/10.1353/mpq.2006.0029 https://doi.org/10.1177/0145445516659645 https://doi.org/10.1007/s10648-020-09533-1 https://doi.org/10.1007/s10648-020-09533-1 https://doi.org/10.1007/s10212-023-00731-3 https://doi.org/10.1016/j.burn.2015.12.002 https://doi.org/10.1146/annurev.psych.52.1.397 https://doi.org/10.2307/2347790 https://doi.org/10.3389/fpubh.2023.1168264 https://doi.org/10.1016/j.brat.2018.02.004 https://doi.org/10.1080/1359432x.2012.742242 https://doi.org/10.1186/s40064-015-1477-6 https://www.terveytemme.fi/kott/taulukot/painotetut_jakaumat_en.html räihä, katajavuori, vehkalahti & asikainen 66 | f l r parpala, a., & lindblom-ylänne, s. (2012). using a research instrument for developing quality at the university. quality in higher education, 18(3), 313-328. https://doi.org/10.1080/13538322.2012.733493 parpala, a., katajavuori, n., haarala-muhonen, a., & asikainen, h. (2021). how did students with different learning profiles experience ‘normal’ and online teaching situation during covid-19 spring? social sciences, 10(9), 337. https://dx.doi.org/10.3390/socsci10090337 pérez-gonzález, j. c., filella, g., soldevila, a., faiad, y., & sanchez-ruiz, m. j. (2022). integrating selfregulated learning and individual differences in the prediction of university academic achievement across a three-year-long degree. metacognition and learning, 17(3), 1141-1165. https://doi.org/10.1007/s11409-022-09315-w räihä, k., katajavuori, n., vehkalahti, k., huotilainen, m., & asikainen, h. (2024). university students’ stress and burnout risk: results of an act-based online-course using self-assessments and hrvmeasurements. current psychology, 1-14. https://doi.org/10.1007/s12144-024-05800-4 räsänen, p., lappalainen, p., muotka, j., tolvanen, a., & lappalainen, r. (2016). an online guided act intervention for enhancing the psychological wellbeing of university students: a randomized controlled clinical trial. behaviour research and therapy, 78, 30–42. https://doi.org/10.1016/j.brat.2016.01.001 rehman, a. u., bhuttah, t. m., & you, x. (2020). linking burnout to psychological well-being: the mediating role of social support and learning motivation. psychology research and behavior management, 545-554. https://doi.org/10.2147/prbm.s250961 ríos‐risquez, m. i., garcía‐izquierdo, m., sabuco‐tebar, e. d. l. á., carrillo‐garcia, c., & solano‐ruiz, c. (2018). connections between academic burnout, resilience, and psychological well‐being in nursing students: a longitudinal study. journal of advanced nursing, 74(12), 2777-2784. https://doi.org/10.1111/jan.13794 rosenberg, j. m., beymer, p. n., anderson, d. j., van lissa, c. j., & schmidt, j. a. (2019). tidylpa: an r package to easily carry out latent profile analysis (lpa) using open-source or commercial software. journal of open source software, 3(30), 978. https://doi.org/10.21105/joss.00978 ryff, c.d. (1998) happiness is everything, or is it? explorations on the meaning of psychological well-being. journal of personality and social psychology, 57(6):1069-1081. https://doi.org/10.1037/00223514.57.6.1069 salmela-aro, k., & read, s. (2017). study engagement and burnout profiles among finnish higher education students. burnout research, 7, 21–28. https://doi.org/10.1016/j.burn.2017.11.001 salmela-aro, k., & upadaya, k. (2012). the schoolwork engagement inventory: energy, dedication, and absorption (eda). european journal of psychological assessment, 28(1), 60–67. https://doi.org/10.1027/1015-5759/a000091 salmela-aro, k., & upadyaya, k. (2014). school burnout and engagement in the context of demands resources model. british journal of educational psychology, 84(1), 137–151. https://doi.org/10.1111/bjep.12018 salmela-aro, k., & upadyaya, k. (2017). co-development of educational aspirations and academic burnout from adolescence to adulthood in finland. research in human development, 14(2), 106–121. https://doi.org/10.1080/15427609.2017.1305809 salmela-aro, k., moeller, j., schneider, b., spicer, j., & lavonen, j. (2016). integrating the light and dark sides of student engagement using person-oriented and situation-specific approaches. learning and instruction, 43, 61–70. https://doi.org/10.1016/j.learninstruc.2016.01.001 salmela-aro, k., savolainen, h., & holopainen, l. (2009). depressive symptoms and school burnout during adolescence: evidence from two cross-lagged longitudinal studies. journal of youth and adolescence, 38(10), 1316–1327. https://doi.org/10.1007/s10964-008-9334-3 https://doi.org/10.1080/13538322.2012.733493 https://dx.doi.org/10.3390/socsci10090337 https://doi.org/10.1007/s11409-022-09315-w https://doi.org/10.1007/s12144-024-05800-4 https://doi.org/10.1016/j.brat.2016.01.001 https://doi.org/10.2147/prbm.s250961 https://doi.org/10.1111/jan.13794 https://doi.org/10.21105/joss.00978 https://doi.org/10.1037/0022-3514.57.6.1069 https://doi.org/10.1037/0022-3514.57.6.1069 https://doi.org/10.1016/j.burn.2017.11.001 https://doi.org/10.1027/1015-5759/a000091 https://doi.org/10.1111/bjep.12018 https://doi.org/10.1080/15427609.2017.1305809 https://doi.org/10.1016/j.learninstruc.2016.01.001 https://doi.org/10.1007/s10964-008-9334-3 räihä, katajavuori, vehkalahti & asikainen 67 | f l r salmela-aro, k., tang, x., upadyaya, k. (2022). study demands-resources model of student engagement and burnout. in: reschly, (pp. 77-93) a.l., christenson, s.l. (eds) handbook of research on student engagement. springer, cham. https://doi.org/10.1007/978-3-031-07853-8_4 salmela-aro, k., upadyaya, k., ronkainen, i., & hietajärvi, l. (2022b). study burnout and engagement during covid-19 among university students: the role of demands, resources, and psychological needs. journal of happiness studies, 23(6), 2685-2702. https://doi.org/10.1007/s10902-022-00518-1 sarasjärvi, k. k., vuolanto, p. h., solin, p. c. m., appelqvist-schmidlechner, k. l., tamminen, n. m., elovainio, m., & therman, s. (2022). subjective mental well-being among higher education students in finland during the first wave of covid-19. scandinavian journal of public health, 50(6), 765–771. https://doi.org/10.1177/14034948221075433 schaufeli, w. b., martínez, i. m., pinto, a. m., salanova, m., & barker, a. b. (2002). burnout and engagement in university students a cross-national study. journal of cross-cultural psychology, 33(5), 464–481. https://doi.org/10.1177/0022022102033005003 schaufeli, w.b., & taris, t.w. (2014). a critical review of the job demands-resources model: implications for improving work and health. (pp. 43-68). in: g.f. bauer, & o. hämmig (eds.) bridging occupational, organizational and public health. a transdisciplinary approach. springer. https://doi.org/10.1007/978-94-007-5640-3_4 sclove l. s. (1987). application of model-selection criteria to some problems in multivariate analysis. psychometrika, 52, 333-343. https://doi.org/10.1007/bf02294360 spurk, d., hirschi, a., wang, m., valero, d., & kauffeld, s. (2020). latent profile analysis: a review and “how to” guide of its application within vocational behavior research. journal of vocational behavior, 120, 103445. https://doi.org/10.1016/j.jvb.2020.103445 tein, j. y., coxe, s., & cham, h. (2013). statistical power to detect the correct number of classes in latent profile analysis. structural equation modeling: a multidisciplinary journal, 20(4), 640-657. https://doi.org/10.1080/10705511.2013.824781 tindle, r., hemi, a., & moustafa, a. a. (2022). social support, psychological flexibility and coping mediate the association between covid-19 related stress exposure and psychological distress. scientific reports, 12(1), 8688. https://doi.org/10.1038/s41598-022-12262-w towey-swift, k. d., lauvrud, c., & whittington, r. (2023). acceptance and commitment therapy (act) for professional staff burnout: a systematic review and narrative synthesis of controlled trials. journal of mental health, 32(2), 452-464. https://doi.org/10.1080/09638237.2021.2022628 trindade, i. a., vagos, p., moreira, h., fernandes, d. v., & tyndall, i. (2022). further validation of the 18item portuguese compact scale using a multi-sample design: confirmatory factor analysis and correlates of psychological flexibility. journal of contextual behavioral science, 25, 1-9. https://doi.org/10.1016/j.jcbs.2022.06.003 tuominen-soini, h., & salmela-aro, k. (2014). schoolwork engagement and burnout among finnish high school students and young adults: profiles, progressions, and educational outcomes. developmental psychology, 50(3), 649–662. https://doi.org/10.1037/a0033898 tynan, m., afari, n., dochat, c., gasperi, m., roesch, s., & herbert, m. s. (2022). confirmatory factor analysis of the comprehensive assessment of acceptance and commitment therapy (compact) in active-duty military personnel. journal of contextual behavioral science, 25, 115-121. https://doi.org/10.1016/j.jcbs.2022.07.001 unesco. (2020). nurturing the social and emotional wellbeing of children and young people during crises. unesco covid-19 education response. https://unesdoc.unesco.org/ark:/48223/pf000037327 https://doi.org/10.1007/978-3-031-07853-8_4 https://doi.org/10.1007/s10902-022-00518-1 https://doi.org/10.1177/14034948221075433 https://doi.org/10.1177/0022022102033005003 https://doi.org/10.1007/978-94-007-5640-3_4 https://doi.org/10.1007/bf02294360 https://doi.org/10.1016/j.jvb.2020.103445 https://doi.org/10.1080/10705511.2013.824781 https://doi.org/10.1038/s41598-022-12262-w https://doi.org/10.1080/09638237.2021.2022628 https://doi.org/10.1016/j.jcbs.2022.06.003 https://doi.org/10.1037/a0033898 https://doi.org/10.1016/j.jcbs.2022.07.001 https://unesdoc.unesco.org/ark:/48223/pf000037327 räihä, katajavuori, vehkalahti & asikainen 68 | f l r viskovich, s., & pakenham, k.i. (2020). randomized controlled trial of a web-based acceptance and commitment therapy (act) program to promote mental health in university students. journal of clinical psychology, 76(6), 929-951. https://doi.org/10.1002/jclp.22848 wersebe, h., lieb, r., meyer, a. h., hofer, p., & gloster, a. t. (2018). the link between stress, well-being, and psychological flexibility during an acceptance and commitment therapy self-help intervention. international journal of clinical and health psychology, 18(1), 60-68. https://doi.org/10.1016/j.ijchp.2017.09.002 https://doi.org/10.1002/jclp.22848 https://doi.org/10.1016/j.ijchp.2017.09.002 1. introduction 2. theoretical background 2.1 university students’ well-being, burnout, and engagement 2.2 person-oriented features of study burnout and engagement 3. present study 4. methods 4.1 participants and the procedure table 1 description of the course modules and content 4.2 measures 4.3 preliminary analysis table 2 descriptive statistics 5 results 5.1 the profile solution table 3 fit statistics for the profile solutions figure 1. standardised means of the dimensions of study burnout and engagement at the beginning of the course intervention in the four-profile solution table 4 means and standard deviations of the profiles at the beginning and at the end of the course figure 2. mean changes in the profiles between the first and second measurement time 5.2 changes within and between the profiles table 5 mixed anova with repeated measures and the post-hoc comparisons of the mean changes between the profiles 6. discussion 6.1 implications for practice 6.2 limitations and future directions 7. conclusions ethical approval references codepen editorial frontline learning research vol.8 no. 3 special issue (2020) 1 9 issn 2295-3159 the promise and pitfalls of self-report: development, research design and analysis issues, and multiple methods. luke k. fryera, daniel l. dinsmoreb athe university of hong kong, hong kong b university of north florida, usa abstract as a prelude to this special issue on the promise and pitfalls of self-report, this article addresses three issues critical to its current and future use. the development of self-report is framed in vertical (improvement) and horizontal (diversification) terms, making clear the role of both paths for continued innovation. the ongoing centrality of research design and analysis in ensuring that self-reported data is employed effectively is reviewed. finally, the synergistic use of multiple methods is discussed. this article concludes with an overview of the si's contributions and a summary of the si's answers to its three central questions: a) in what ways do self-report instruments reflect the conceptualizations of the constructs suggested in theory related to motivation or strategy use? b) how does the use of self-report constrain the analytical choices made with that self-report data? c) how do the interpretations of self-report data influence interpretations of study findings? keywords: self-report, multiple methods, vertical and horizonal development, research design and analyses corresponding email: fryer@hku.hk doi: https://doi.org/10.14786/flr.v8i3.623 this si’s mission while self-report measures are ubiquitous across and often central to educational research, they are also often denigrated for a range of reasons. for instance, the reliability of the measures and the validity of the resultant score interpretations are often called into question (e.g., veenman et al 2006). this has led to calls for moratoria on the use of self-report in some corners. however, rather than discarding or ignoring data generated from self-report measures of cognitive processing, motivation, emotions and beliefs, research is needed to determine when and if self-report measures can contribute to our collective understanding of theory surrounding these constructs. for example, relying solely on self-report to study regulatory processes has contributed little to our understanding of self-regulated learning (dinsmore et al 2008), however, in other instances self-report may be the only viable manner in which to unearth covert constructions, such as self-efficacy (e.g., zimmerman, 2000). this special issue examines the accuracy of interpretations and conclusions drawn from self-reports regarding individuals’ metacognitive and cognitive processing, affect and beliefs, and the analytic choices made. these questions are addressed by an international group of experts examining these constructs from different theoretical and analytical perspectives. the current special issue as whole brings three issues that are often noted, but rarely specifically discussed into focus: lateral versus horizontal development of measurement approaches, the critical role of research design and analyses, and the complex role of utilizing multiple measurement methods. 2. lateral and vertical innovation: both are critical the first topic to be addressed in this special issue is how self-report approaches are advancing both laterally (i.e., improving current methods) and horizontally (i.e., developing new methods). as an analogy, a vibrant city in the late 19th and early 20th century was bustling with horses, horse-drawn trams, and cars that coexisted along the city thoroughfares. figuring out the best way to get across the city is dependent upon many factors – such as the persons wealth or when they are trying to make their transit. similarly, the research literature is replete with different vehicles to transport oneself from point a to point b – namely how to measure the processes described in this special issue to better understand theory, and ultimately, student learning. as with the city, the best mode to measure these constructs depends on many factors. however, unlike the advances in transportation technology, the advances in and pressure to modernize self-report methods has been weak at best. this advancement of methods (or lack thereof) for measuring latent constructs critical to educational research (self-report included) can be framed by horizontal versus vertical conceptions of development. this is a well-established framework for understanding growth (e.g., economic innovation; bondarev, & greiner, 2019) and change (e.g., natural selection; lawrence, 2005) in a broad array of fields. horizontal growth refers to innovating towards entirely new approaches, while vertical growth refers to refining and enhancing current methods. this framework fits the current era as educational research is flooded with new (horizontal) means of measuring students’ cognitive processing (e.g., eye tracking; chaulic et al. 2020) and meta-cognitive processing (e.g., trace data; rogiers, et al., 2020). this horizontal drive for innovation of measurement continues to push into complex areas such as emotion (facial recognition; chiu, et al., 2019; dingle, et al., 2016; skin conductance, järvenoja, et al. 2018; lehikoinen et al., 2019) and motivation (neuroscience; hidi, 2016; mayer, 2017). the considerable momentum behind this drive for alternatives to self-report measurement arise in large part with a longstanding dissatisfaction with their intra-psychic nature and the general lack of lateral development in these measures. given the attention that the horizontal development of these measures has garnered and the lack of lateral development, this special issue addresses the many areas of lateral development that are possible. in other words, this special issue forges new inroads towards further development of self-report measurements across a range of processes. not only do these empirical pieces suggest lateral development is possible, each starts to take us down this journey and provides evidence that these journeys are likely to be fruitful. from empirical multimethod studies (e.g., van halen et al., 2020; rogiers, et al. 2020) to the theoretically-rich commentaries (pekrun, 2020; van meter, 2020; winne, 2020; ), this special issue suggests that self-report measures are a unique, valuable – and therefore irreplaceable – source of information about many critical aspects of the learning processes under study here. clearly, the conclusions drawn from these analyses show that self-report remains critical in our understanding of learning and that educational researchers need to push harder for constant lateral innovation such as these. it safe to say, however, that many of these researchers have not felt this obligation. as noted in fryer & nakao’s (2020) contribution, the primary self-report instrument for the majority of educational research is a tool invented in the 1920s. on paper or smartphones, a likert scale is a likert scale. we can and should be struggling to improve on the tools of our trade – as is already being done widely across the technology industry (e.g., google; lawless & biel, 2020). lateral advances in self-report can take many forms, a few of which are presented in this special issue. addressing perceived weaknesses in the format by either adding to it (durik & jenkins, 2020) or changing it (fryer & nakao, 2020) are both rungs in the ladder up toward vertical innovation. addressing how self-report tools are used and their results analysed (moller, et al) is another means of climbing further up those rungs. a scan of leading journal suggests that the latter approach to lateral innovation is expanding (e.g., gillet, et al., 2019; yuen, et al., 2019), while the former – i.e., analysis – is almost unknown. both are necessary if measurement of constructs critical to education (self-report or otherwise) is to continue to improve. 3. not compounding self-report error: it is just common sense the limitations sections of educational research articles are replete with apologies. there are three apologies that in our experience vie for most prevalent: a) the self-reported nature of the data, b) the cross-sectional nature of the research design and, hinging on the first two, c) the rigor of the analyses. it is time researchers stopped apologising for the first when it is necessary and useful, and did something about the latter two. the weaknesses of self-report have been well known for some time. as the present special issue has confirmed, however, self-report also has its unique strengths – often left unmentioned in critiques of self-report. self-report is an important part of research in many areas (e.g., motivation), but should not be the only measurement tool. the supplement of self-report with other observed measures has the potential to improve our understanding of the complex interrelations between the variables described in this special issue and learning. this balanced approach to measurement should bring an end to self-report’s inclusion in limitations sections. the unresolved issue for the field is that many researchers continue to compound the weaknesses of self-report with cross-sectional designs and inappropriate analyses. these are at least partially linked, as authors, struggling to get published, are desperate to use popular analytical methods. the best example of this mis-use is structural equation modelling (sem) with cross-sectional data sets. the often small sample sizes being utilised means that researchers are forced to use mean-based sem rather than latent-based sem (for an introductory overview of each and their differences see kline, 2011). this pairing of common analytical mis-steps, compounds the inherent weaknesses of self-report: like building a tall, narrow building in an earthquake prone area. these two research design issues amplify the self-report concerns cited previously in separate but connected ways. first, educational research is generally seeking to explain learning or learning related processes; processes which are by their very nature developmental and require longitudinal examinations of that development. research using snapshots of the learning experience can make meaningful contributions to educational research, but only if the limitations of these static designs are taken seriously and appropriate analyses employed. ginns et al (2017) is an example of exactly this kind of theoretically robust, carefully structured cross-sectional research. second, self-report within educational research is generally employed to measure latent constructs. it therefore makes sense to employ analyses that treat self-report data as though it were representing latent constructs. this seems especially pertinent to survey self-report which generally measures constructs with multiple items. for fine grained analysis of subtle aspects of the learning process or interventions having nudge like effects (small but meaningful over time), the error imbibed by averaging across multiple survey items is a serious, and too often ignored issue. novice readers and those skimming through articles, are prone to conflate the often mean-based and latent-based sem. path analysis, in addition to its inherent mistreatment of latent variables, prevent fully forward analysis due to a lack of degrees of freedom, resulting in the picking and choosing of regressive connections. this additional author-induced source of error is akin to the file-drawer bias (i.e., you don’t get the whole picture). readers of articles that utilize mean-based sem are too often presented with a cropped picture, only showing connections which support the researcher’s aims. what is commonly referred to as path analyses is just one example of how self-reported data can be mishandled, and lead to exacerbating their inherent weaknesses. it is important to restate that cross-sectional designs can make a limited contribution to research, but researchers need to acknowledge their limitations and not draw conclusions that their data does not support. researchers seeking to make a strong contribution to an area of educational research where self-reported measures are an important part of quality research design should strive to employ designs that can capture developmental processes and analyses that recognise latent for what it is: unseen. for a detailed and balanced discussion of this issue, we encourage a careful review of martin (2011). 4. the promise of multiple methods as discussed previously, many of the papers in this special issue use multiple measures to present a more complete picture of the complex interrelations between constructs. however, we should be careful to distinguish between the aims of engaging in this process and the analyses of these multiple measures, which has often been referred to as triangulation (godfroid & spino, 2015). we see three potential paths here: using measurements to identify the same aspect or aspects of a constructs, using measurements to identify complementary aspects of a construct, or both. we offer an analogy here to help the reader better understand these paths. the first path, identifying the same aspect would be akin to using multiple measurements of sound to identify the pitches (how high or low a note sounds) of the notes in a melody (i.e., the part of a tune you might hum). one might use a well-trained ear and an electronic tuner to do this. if both the listener’s ear and the tuner are accurate, they should agree on the pitches – maybe the tune starts out and ends on “middle c”. similarly, when examining one of the constructs in this special issue, say metacognition, this first path would be akin to saying that our multiple measurements are indeed measuring the same aspect of metacognition – that they should agree. this approach would often be analysed using a multi-method multi-trait analysis (mtmm; c.f., campbell & fiske, 1959). here, we expect the same techniques used to measure the same construct to “agree” more often than those techniques used to measure different constructs. for example, if retrospective self-report and physiological measurements are used to measure reading comprehension and mathematical achievement, the self-report and physiological measurements of the same construct (e.g., reading comprehension) should correlate more closely than the two self-report measurements of the two different constructs. the latter would be an example of a methods effect, while the former would demonstrate that both measurements are measuring the same aspect or construct. however, the melody of a piece of music is often not the only aspect of a musical composition. in a symphony, for instance, the melody is often accompanied by other lines of music as well (which would be harder for the novice to hum). thus, it might be necessary to use different techniques to identify the sounds present. while a well-trained ear would be able to pick out the chords composed by these multiple musical lines, a simple tuner would not. this issue gets even more complex when thinking about a bach fugue for example, that layers multiple melodies and countermelodies to create a rich tapestry of sound (c.f., j. s. bach’s toccata and fugue in d minor, bwv 565). this more complex conceptualization describes the second path here – are we using multiple measurements to better describe the rich symphony of a process at play? in other words, are there multiple aspects of a particular construct that some measurements are better at tapping than others? for instance, if an mtmm analysis demonstrated that two different measurements of the same construct did not correlate well, does that mean they are inaccurate or does that mean that the construct under investigation is multi-faceted in the same way that bach’s fugues are multi-faceted? the third path – and the one that we recommend – is considering both of these routes as these multiple measurements are considered. in other words, when do our measurements measure the same aspect of a construct and when do they measure different aspects of a construct. for example, although strategy use is considered a construct within a domain, different aspects of that strategy use (i.e., quantity, quality, and conditional use) have been demonstrated to be related to learning in different ways (dinsmore, 2017). thus, how can we operationalize our theoretical conceptions of strategy use in meaningful ways to build and use theory? this is particularly important as we think about the development (e.g., changes) of these processes as they unfold over time. like our tocatta and fugue in d minor example, it is quite possible that we begin with a simple melody, but then morph into a more complex interweaving of voices as the development of the piece progresses. 5. empirical contributions to explore how we can improve self-report measurements or use them in concert with other measurements, eight empirical studies were conducted. these studies examined the validity of score interpretations and future of self-report measurements. these studies each addressed at least two of the special issue’s three central questions: 1. in what ways do self-report instruments reflect the conceptualizations of the constructs suggested in theory related to motivation or strategy use? 2. how does the use of self-report constrain the analytical choices made with that self-report data? 3. how do the interpretations of self-report data influence interpretations of study findings? durik and jenkins’s (2020) test of the role of certainty with self-report surveys. they build on literature connecting attitude to behaviour, seeking a new perspective on the relationship between interest and behaviour. this paper tests the relationship between level of interest and certainty of that self-report. this is then extended to examination of the connections between certainty and related behaviour. durik and jenkin’s is a rare attempt at vertical innovation with interesting preliminary implications for survey methods and interest research theory. this research needs to be followed up with different participants and variations on their research design. chauliac et al. (2020) employed eye-tracking to assess the cognitive processes participants undertake while completing a quantitative questionnaire. they aimed to establish linkages between participants eye movements and their questionnaire answering behaviour. this research yielded no simple answers but lays a foundation for further research into the processes underlying questionnaire response behaviours – namely, in helping to figure out if the questionnaires and eye movements were measuring similar or different aspects of those underlying reading processes. making a case for the multimethod approaches that recognise both the value and weakness of self-report, vriesema and mccaslin (2020), bring survey self-report and observations together in their article. their results suggest that there is clear alignment between self-report and classroom observations of student groups at ages as young as grade three. their findings support the use of self-report as part of robust research design for a broad range of ages. rogiers et al. (2020) employed think aloud protocols to further explore person centered survey self-report findings regarding secondary school students' text-learning strategies. results from this combination of retrospective and concurrent approach to self-report pointed to the validity of self-reports. the latter approach provided an additional, nuanced, often ignored perspective on the frequency and sequence of students’ strategies. this article reviews how this pair of self-report methods offers researchers a unique bifocal perspective on student learning experience. iaconelli and wolters (2020) address an area of survey research which is often noticed but rarely engaged with: insufficient effort responding to surveys. this research tests whether “insufficient effort responding” (ier) to survey question is a meaningful threat to survey data validity. as important as their findings, which point to ier as more nuisance than threat, are their recommendations for survey research when reporting their findings. toward vertical innovation of survey self-report in this mobile age, fryer and nakao (2020) present an experimental test of four survey formats (likert, visual analogue scale, slide, and swipe). a series of analyses on the resulting data set encourage more work with continuous formats like slide and swipe. nearly a century on from the inception of formats like likert and vas, the authors suggest it is time for researchers to look up and embrace our touch-based future. van halem et al. (2020) presented a study triangulating survey self-reports of self-regulated learning with online traces of students learning behaviours. they confirm that aptitude-based self-reports cannot accurately capture complex srl alone. their findings suggest that self-report measures and online data regarding srl are complementary in predicting students’ study success. results demonstrate that both perspectives explain a unique proportion of students’ academic performance. moeller et al. (2020) aimed to make students’ course feedback more meaningful to instructors. they did this through research design and analyses that separate the broader learning situation from the individual’s reported experience. this separation makes it possible to track the subjective and objective development of learning experiences across a course. in addition to demonstrating how their methods might support individualised learning, moeller et al.’s study raised the critical role of multiple methods and including objective measures in self-report centered research. 6. the commentaries in addition to these empirical contributions, three international experts have weighed in on how this special issue's articles make substantive contributions to the extant research literature. each focus on at least two of the three guiding questions for the special issue. winne (2020) takes a conceptual approach by focusing on what self-report data are. while there is much discussion in the literature about our conceptualizations of constructs, there is much less discussion about how we conceptualize the measurements themselves. winne tackles this thorny issue. winne argues – and we agree – that without a better conceptualization of self-reports, there is little evidence that participants can get better at responding to them. in turn, the better participants are at responding to these types of measurements, the better the interpretations of these data will be. pekrun (2020) makes the case for the importance of self-report data, and like winne touches on what they are. pekrun extends this discussion by focusing primarily on how to improve the validity of the score interpretations of self-report data. finally, van meter (2020) tackles all three questions. at the heart of her commentary is a deep dive into when self-reports are useful and how they can be leveraged to best help us build and refine theory. her theoretically-driven set of conditions for when and how to use self-report offer both younger and more experienced researchers alike a useful framework to guide their choices of self-report measurements. 7. implications of the special issue this brings us full circle back to how the empirical and commentary articles have together addressed this special issue’s focal points. the eight empirical contributions stretched across the theoretical domains of self-regulation, interest and cognitive processing strategies, but still presented a coherent picture of the validity and future of self-report. 1. in what ways do self-report instruments reflect the conceptualizations of the constructs suggested in theory related to motivation or strategy use? a common theme across rogiers et al (2020), van halem et al. (2020) and vriesema et al. (2020) is that the retrospective survey self-report of attitudes and dispositions are an important often unique part of understanding future learning experience and outcomes. however, for more comprehensive, dynamic conceptualisations to be drawn, additional online measurement is critical. this online measurement might be self-report (tap) or observed (trace or observations), both perspectives have the potential to expand our understanding of students’ strategies and motivations for learning. the answer to the si’s question is therefore that instruments do matter, and the path toward more robust conceptualizations is multimethod research designs. any questions about whether those methods should be self-report or not can be set aside. 2. how does the use of self-report constrain the analytical choices made with that self-report data? the wide variety of contributions to this special issue demonstrate it is the broader question of research design that determines analytical choices as much or more than how self-report is used. experimental, repeated measures, variable/person-centered analyses and an array of mixed methods arrangements exemplify the full range of analytical tools available to researchers. researchers are strongly encouraged to focus less on well-known issues with self-report, and instead look to the designs they are embedded within and analyses employed. 3. how do the interpretations of self-report data influence interpretations of study findings? while each of this special issue's articles addresses this question in some form, the syntheses of the three commentaries address it best. all three of these commentaries point out the need to clearly understand the core construct (e.g., pekrun, 2020), the measurement itself (winne, 2020), and the conditional nature of their use (van meter, 2020). it is critical that the interpretations of self-report data are situated within theoretical frameworks of the core constructs that are being measured. for instance, interpretations of self-report data around interest (e.g., durik & jenkins, 2020) will be qualitatively different than those around feedback (moeller et al., 2020). in other words, the self-report itself must change to allow better interpretations; even survey formats must remain flexible to innovation (e.g., fryer & nakao, 2020). 8. concluding thoughts the impetus of this special issue was borne out of our frustration as younger scholars with the tools used to study the covert, complex processes at the heart of this special issue. having seen self-report used poorly and hearing the calls from scholars at all stages of their career calling for a moratorium on self-report, we wanted to expand the discussion beyond simply deciding whether we should forge ahead with the same old tools in the same old way or abandon them all together. rather, we wanted a vehicle in which scholars could reflect on the way these tools are used and use them more appropriately. we were fortunate to have a number of scholars willing to contribute high-quality empirical studies to this effort. the three commentaries then provided excellent avenues for extending these conversations and hopefully spurring more deep conversations about these thorny issues. we hope that readers of this special issue will be as satisfied with the result as we were in helping to curate them. references bondarev, a., & greiner, a. (2019). endogenous growth and structural change through vertical and horizontal innovations. macroeconomic dynamics, 23, 52-79. https://doi.org/10.1017/s1365100516001115 campbell, d. t., & fiske, d. w. (1959). convergent and discriminant validation by the multitrait-multimethod matrix. psychological bulletin, 56, 81-105. chauliac, m; catrysse, l. ; gijbels, d. & donche v. (2020). it is all in the surv-eye: can eye tracking data shed light on the internal consistency in self-report questionnaires on cognitive processing strategies? frontline learning research. 8 (3), 26 – 39. https://doi.org/10.14786/flr.v8i3.48 chiu, m. h., liaw, h. l., yu, y. r., & chou, c. c. (2019). facial micro‐expression states as an indicator for conceptual change in students' understanding of air pressure and boiling points. british journal of educational technology, 50, 469-480. https://doi.org/10.1111/bjet.12597 dingle, g. a., hodges, j., & kunde, a. (2016). tuned in emotion regulation program using music listening: effectiveness for adolescents in educational settings. frontiers in psychology, 7, 859. https://doi.org/10.3389/fpsyg.2016.00859 dinsmore, d. l. (2017). towards a dynamic, multidimensional model of strategic processing. educational psychology review, 29, 235-268. https://doi.org/10.1007/s10648-017-9407-5 dinsmore, d. l., alexander, p. a., & loughlin, s. m. (2008). focusing the conceptual lens on metacognition, self-regulation, and self-regulated learning. educational psychology review, 20, 391-409. https://doi.org/10.1007/s10648-008-9083-6 durik, a. m. & jenkins j. s. (2020). variability in certainty of self-reported interest: implications for theory and research. frontline learning research. 8 (3) 85-103. https://doi.org/10.14786/flr.v8i3.491 fyer, l. & nakao k. (2020). the future of survey self-report: an experiment contrasting likert, vas, slide, and swipe touch interfaces. frontline learning research, 8 (3),10-25. https://doi.org/10.14786/flr.v8i3.501 gillet, n., morin, a. j., huyghebaert, t., burger, l., maillot, a., poulin, a., & tricard, e. (2019). university students' need satisfaction trajectories: a growth mixture analysis. learning and instruction, 60, 275-285. https://doi.org/10.1016/j.learninstruc.2017.11.003 ginns, p., martin, a. j., & papworth, b. (2018). student learning in australian high schools: contrasting personological and contextual variables in a longitudinal structural model. learning and individual differences, 64, 83-93. https://doi.org/10.1016/j.lindif.2018.03.007 godfroid, a., & spino, l. a. (2015). reconceptualizing reactivity of think‐alouds and eye tracking: absence of evidence is not evidence of absence. language learning, 65, 896-928. https://doi.org/10.1111/lang.12136 hidi, s. (2016). revisiting the role of rewards in motivation and learning: implications of neuroscientific research. educational psychology review, 28(1), 61-93. https://doi.org/10.1007/s10648-015-9307-5 iaconelli, r. & wolters c.a. (2020). insufficient effort responding in surveys assessing self-regulated learning: nuisance or fatal flaw? frontline learning research. 8 (3) 104 – 125. https://doi.org/10.14786/flr.v8i3.521 lawless, k. a., & riel, j. (2020). exploring the utilization of the big data revolution as a methodology for exploring learning strategy in educational environments. in d.l. dinsmore, l. k. fryer, & m. m. parkinson (eds.), handbook of strategies and strategic processing, (pp.296-316). new york: routledge. lawrence, j. g. (2005). horizontal and vertical gene transfer: the life history of pathogens. contributions to microbiology, 12, 255-271. kline, r. b. (2011). principles and practices of structural equation modeling (3 ed.). new york: guilford press. martin, a.j. (2011). prescriptive statements and educational practice: what can structural equation modeling (sem) offer? educational psychology review. 23. 235-244. https://doi.org/10.1007/s10648-011-9160-0 mayer, r. e. (2017). how can brain research inform academic learning and instruction? educational psychology review, 29(4), 835-846. https://doi.org/10.1007/s10648-016-9391-1 moeller, j. ;viljaranta, j.; kracke, b. & dietrich, j. (2020). disentangling objective characteristics of learning situations from subjective perceptions thereof, using an experience sampling method design. frontline learning research, 8(3), 63-84. https://doi.org/10.14786/flr.v8i3.529 pekrun, r. (2020). self-report is indispensable to assess students’ learning. frontline learning research, 8 (3), 185–193. https://doi.org/10.14786/flr.v8i3.627 rogiers, a.; merchie, e. & van keer h. (2020). opening the black box of students’ text-learning processes: a process mining perspective. frontline learning research, 8(3), 40 – 62. https://doi.org/10.14786/flr.v8i3.527 van halem, n., van klaveren, c., drachsler h., schmitz, m., & cornelisz, i. (2020). tracking patterns in self-regulated learning using students’ self-reports and online trace data. frontline learning research, 8(3) 140-163; https://doi.org/10.14786/flr.v8i3.497 van meter, p. (2020) measurement and the study of motivation and strategy use: determining if and when self-report measures are appropriate. frontline learning research, 8(3), 174–184. https://doi.org/10.14786/flr.v8i3.631. veenman, m. v., van hout-wolters, b. h., & afflerbach, p. (2006). metacognition and learning: conceptual and methodological considerations. metacognition and learning, 1, 3-14. https://doi.org/10.1007/s11409-006-6893-0 vriesema, c.c., & mccaslin, m. (2020) experience and meaning in small-group contexts: fusing observational and self-report data to capture self and other dynamics. frontline learning research, 8 (3), 126-139. https://doi.org/10.14786/flr.v8i3.493 winne, p. (2020) a proposed remedy for grievances about self-report methodologies. frontline learning research. 8 (3) 164 -173. https://doi.org/10.14786/flr.v8i3.625 yuen, a. h., cheng, m., & chan, f. h. (2019). student satisfaction with learning management systems: a growth model of belief and use. british journal of educational technology, 50(5), 2520-2535. https://doi.org/10.1111/bjet.12830 zimmerman, b. j. (2000). self-efficacy: an essential motive to learn. contemporary educational psychology, 25, 82-91. https://doi.org/10.1006/ceps.1999.1016 frontline learning research vol. 12 no. 2 (2024) 51 69 issn 2295-3159 corresponding author: kateryna horlenko, education academy, vytautas magnus university, lithuania, email address: kateryna.horlenko@vdu.lt doi: https://doi.org/10.14786/flr.v12i2.1417 student self-regulated learning in teacher professional vision: results from combining student self-reports, teacher ratings, and mobile eye tracking in the high school classroom kateryna horlenko1, lina kaminskienė1& erno lehtinen1,2 1 vytautas magnus university, lithuania 2 university of turku, finland article received 18 december 2023 / article revised 15 may 2024 / accepted 14 june / available online 27 june abstract teacher professional vision as a concept is gaining importance in research on teaching, and recently models for studying teacher professional vision and student self-regulated learning (srl) have been proposed. there are interview and video intervention studies investigating teacher professional vision for srl, but no real-life classroom research so far. this study investigated the role of student srl behaviour, as it was reported by students themselves and teachers, in teacher attention distribution as part of teacher professional vision. ten teachers and their 158 students at high school level in lithuania took part in the research. the first step of the study resulted in identifying four student srl-profiles, which differed based on student level of srl and the extent to which teacher and student assessments coincided: mixed lower-regulated, mixed higher-regulated, systematic lower-regulated, systematic higher-regulated. the profiles demonstrated only a partial overlap in teacher and student judgement of student srl. the second step of the study explored whether scores of students’ srl from student and teacher reports were related to teachers’ distribution of visual attention in one lesson. the results showed that only one teacher rating scale of student information-seeking behaviour had a slight correlation with teacher attention. the results imply rather bottom-up trends in teacher attention to students in the classroom when it comes to srl. besides, the study results highlight the not directly observable nature of srl processes and imply a difficulty for teachers to assess student srl. keywords: self-regulated learning; self-report; teacher rating; cluster analysis; mobile eye tracking horlenko, kaminskienė & lehtinen 52 | f l r 1. introduction self-regulated learning (srl) has been associated with higher student motivation and academic achievement (zimmerman, 2001; zimmerman & kitsantas, 2007; theobald, 2021) and is part of the european lifelong learning framework (sala et al., 2020). teachers can incorporate promotion of srl in their classroom instruction by teaching the learning strategies to students or structuring the learning environment in an autonomy-supportive way (dignath & veenman, 2020). at the same time, students differ in their experience with and practice of srl (cleary et al., 2021; heirweg et al., 2019). thus, in order to provide support adequately, teachers need to be aware of students’ srl (dignath & sprenger, 2020). however, it is unclear to what extent teachers consider srl-related factors when giving guidance and feedback to students directly in the classroom. in the classroom, teachers do not formally assess students’ srl skills but rely on the cues from students by paying attention to students in the classroom. teacher attention in the classroom can be explained through the concept of professional vision that comprises teacher’s ability to dynamically notice and interpret classroom events that are relevant for student learning (van es & sherin, 2002; seidel & stürmer, 2014). recently, models of professional vision specifically for teaching and assessing srl have been proposed (michalsky, 2014; greene, 2021). teachers need to recognise students’ eventual needs for support and provide support accordingly (van de pol et al., 2010). understanding the variations in students’ current levels of srl is important, as teachers’ srl-related instructions that do not align with students’ needs can be counterproductive (vermunt & verloop, 1999). in this sense, teachers need to employ both their srl knowledge to assess their students’ srl involvement long-term and noticing skills to assess students’ current support needs in each lesson. this makes it relevant to consider the interplay between the top-down and bottom-up noticing trends within teacher professional vision. previous observational studies repeatedly showed that teachers provided limited srl strategy instruction (see dignath & veenman, 2021 for an overview). however, these analyses were done at the classroom level without considering the individual students. another line of research on scaffolding interaction evidenced how teachers support students in applying learning strategies in small groups (kajamies et al., 2017; salo et al., 2022) or one-to-one tutoring situations (abdoulaye, 2003). this study attempts to bring together the two perspectives and to examine if there are any overlaps between student and teacher reports of srl, whether any regularities can be drawn from these perspectives via a personoriented analysis, as well as to consider teacher attention distribution to individual students in relation to student srl in the real-life classroom. 1.2 self-regulated learning and related student behaviour srl is a concept that views learning as a process in which students actively regulate own thinking, emotional responses and behaviour, in order to complete academic tasks and improve continuously (alexander et al., 2018). several models have been proposed to explain srl, where srl is often considered as a multidimensional and cyclical process, in which students apply strategies to self-regulate and reach task goals (panadero, 2017; zeidner & stoeger, 2019; puustinen & pulkkinen, 2001). multidimensionality implies that when approaching academic tasks, effective learners not only engage in cognitive processing (e.g., perceiving, problem-solving, remembering), but also practice metacognitive monitoring and control, as well as regulate emotions that arise in the process. emotional regulation is closely related to the motivational processes, such as finding interest in tasks, volitional control, and attributing failure or success to effort or ability (pintrich & de groot, 1990). the cyclical nature of srl implies that, ideally, when regulating their learning, students go through certain phases (pintrich, 2004; winne, 2017; zimmerman, 2002). for example, zimmerman (2002) described three phases: forethought, performance and self-evaluation. in the forethought phase, learners identify the task goals, plan for reaching them, as well as activate motivational beliefs (e.g., task value) to engage with the task. during the performance phase, students cognitively work on the task and metacognitively monitor their progress in relation to the identified task goals. in the self-evaluation phase, learners reflect horlenko, kaminskienė & lehtinen 53 | f l r on the result of the task and own performance, identify reasons for reaching or not reaching the goal, and how this can be improved further. finally, to go through the above phases, students need to apply strategies to solve tasks (i.e., domain-specific cognitive strategies, winne, 2017), to plan, monitor and reflect on activities (metacognitive strategies, veenman & van cleef, 2019), to acknowledge and change ones’ motivational beliefs and to control emotions (motivational strategies, pintrich & de groot, 1990). research shows that it is challenging for students to regulate own learning. some students can engage in maladaptive regulatory behaviours, such as not applying learning strategies, keeping notes and study environment disorganised, procrastinating, or self-handicapping (bembenutty, 2011), thus lacking selfregulatory skills (martínez-fernández et al., 2024). hence, students need help in initiating and practicing srl. in the classroom, this also may lead to a situation where students with different levels of srl skills require different kinds of srl support from teachers (zimmerman, 2013; vermunt & verloop, 1999; callan et al., 2022). students with lacking srl skills can become lost in the learning activities that require much autonomy, while students who are used to srl may perceive additional support as over-teaching (vermunt & verloop, 1999; van de pol et al., 2010; peeters et al., 2016). in other words, teachers need to provide srl instruction in an adaptive, calibrated way so that different students can benefit from it (corno, 2008). the first step in planning optimal support is diagnosing students’ current level of knowledge and srl (van de pol et al., 2010; kajamies, 2017). 1.3 student and teacher assessment of self-regulated learning measuring srl is generally challenging (veenman et al., 2006; boekaerts & corno, 2005). understanding student srl is important both for students and their teachers. there are mainly two approaches to measuring srl: (1) as an aptitude, in a generalised way as experienced over time, reported at single time point, e.g., in a questionnaire, or (2) via directly following the process of learning, captured over a certain period e.g., with a think-aloud protocol (winne & perry, 2000; panadero et al., 2016). van hout-wolters (2000) made a similar distinction into offline and online measures. although process-based and multimodal measures are introduced (panadero et al., 2016), much evidence in the field of student srl is collected via self-reports, as they can be used in combination with other methods. besides, self-report measures tend to be informative for assessing global self-regulation rather than specific strategy use (rovers et al., 2019). considering the practical sides of measuring srl, rigorous process-based assessment methods are not feasible in the school settings, so teachers and school psychologists would benefit from student self-reporting of srl (cleary, 2006). student srl self-reports have been used in variable-based and person-oriented analyses. application of srl strategies and general perception of own srl generally correlate with achievement (see credé & phillips, 2011 for an overview). however, as srl is a dynamic process and students differ in their degree of practicing srl, person-oriented analyses can be useful in distinguishing student srlprofiles for intervening adaptively. in the study by heirweg et al. (2019) with primary school students, clustering based on self-report questionnaires yielded four student profiles (i.e., active learners with high quantity motivation, active learners with high quality motivation, passive learners with low quantity motivation, passive learners with low quality motivation), while clustering based on think-aloud protocols revealed only two profiles (i.e., low and high srl learners). at the secondary school level, similarly, connecting srl and motivation, ng (2016) distinguished four profiles of student procrastination and self-regulation: active procrastinator, active self-regulator, passive self-regulator, and passive procrastinator. abar and loken (2010) distinguished between high srl, low srl, and average srl students, with high srl students reporting high levels of mastery orientation while the low self-regulation group related more to avoidant goal orientation. martínez-fernández et al. (2024) showed that students who practiced self-regulatory behaviours were more satisfied with autonomous learning environments. cleary et al. (2021) found relationships between student reported srl and perceived school connectedness and support, identifying high srl students with high levels of support, low srl students who felt supported, solid srl students with low support, and very low srl students with low support. in another study, clustering student learning patterns based on the process trace data in the online environment were linked to different srl needs of students: clusters with high srl horlenko, kaminskienė & lehtinen 54 | f l r indicators of task accuracy and knowledge development related to low support needs, and groups with low srl indicators required more support (dijkstra et al., 2023). although there may be discrepancies between student srl self-report and the process-based data from their actions during learning (heirweg et al., 2019; winne & jamieson-noel, 2002), student profiles based on self-reporting can help researchers and teachers to recognise the variety of student support needs. teachers are assumed to be able to make judgements about student srl based on the daily school activities in a process-based manner, such as observing a student performing a task (winne and perry, 2000). when it comes to formal assessment methods, teachers are familiar with srl self-report measures more than with process-based measures (michalsky, 2017). teacher ratings of student srl behaviour can be used as an additional measure of student srl. zimmerman and martinez-pons (1988) used a teacher rating scale of student srl based on zimmerman’s srl model and found that student reports of using srl strategies in structured interviews correlated with teacher ratings (r=.70). more recently, cleary and colleagues (2021) applied teacher ratings to validate student srl self-reports. in their study, teacher rating correlated with student reports of mathematics interest (r =.32), maladaptive regulatory behaviours (r =−.41), and test taking strategies (negatively worded, r =−.42). the overlap between student self-reports and teacher rating of student involvement in regulatory behaviours, such as planning, self-monitoring, organising environment for learning and seeking help has not been widely studied. some of the srl-related behaviours, like persistence, seeking help and feedback, being organised can be directly observed. other more strategic processes of srl related to cognition and metacognition are more challenging to spot. this also aligns with the findings that teachers usually are not trained to pay attention to learning processes and events that are indicative of srl (callan & shim, 2019; dignath & sprenger, 2021). therefore, teachers need to develop the capacity to notice indicators of srl behaviour of students: whether and how students involve in the activities of planning, monitoring, and self-evaluating performance on a task (de vries et al., 2022). this can be described as part of teacher professional vision. 1.4 teacher professional vision for self-regulated learning professional vision (goodwin, 1994) is a link between the professional knowledge and its application in particular situations. for teachers, it stands for noticing and interpreting key classroom events and interactions (van es & sherin, 2002; seidel, & stürmer, 2014). the mechanism for this is based on (1) selective attention to moments that are important for learning, e.g., changes in students’ understanding, (2) teacher’s pedagogical and content knowledge, connecting specific classroom interactions to the broad educational principles, and (3) using one’s knowledge about the specific context to explain the interactions (van es & sherin, 2002). the latter emphasises the impact of each class conditions, student characteristics and behaviour on teacher’s decisions during instruction, including support for srl. michalsky (2014) put forward a conceptual model that combined teacher professional vision with the framework of srl instruction by dignath & büttner (2008). greene (2021) included teacher professional vision to the set of factors that promote teacher’s support of students’ srl in the classroom, along with teacher’s epistemic beliefs, teacher’s own self-regulation capacity and overall teaching competence. teachers’ srl knowledge and teaching experience are shown to be predictors of teacher-reported srl support in the classroom (callan et al., 2022). at the same time, studies in germany and the usa showed that teachers’ understanding of srl did not align with the academic conceptualisation of srl, which also hindered their assessment of student srl (dignath & sprenger, 2020; callan & shim, 2019). teacher’s prior conceptual knowledge about srl also plays a crucial role for the noticing component of professional vision for srl: pre-service teachers were able to distinguish between different types of strategy instruction in episodes of classroom videos after taking a course specialised on srl in teaching (michalsky, 2014), while teachers without a specialised srl training were not able to correctly recognise strategy instruction, regardless of their level of expertise (michalsky, 2021a). michalsky (2021b) incorporated the professional vision lens into an intervention for scaffolding pre-service teachers’ capacity to teach metacognitive and strategic knowledge to students. their results showed that those pre-service teachers who reflected on both teachers’ and horlenko, kaminskienė & lehtinen 55 | f l r students’ behaviours in the video-based learning materials improved the skill for teaching strategies to students, which also led to better student outcomes. such effect was not observed in pre-service teachers whose reflections on classroom videos focused only on teachers’ behaviours. this study highlighted the importance of analysing student behaviours in teachers’ capacity to promote srl. hence, professional vision for srl brings the discussion on srl to the practical domain: it is not only important for teachers to know about srl and how to incorporate it in their teaching, but also how to notice indications of srl in their students. 1.5 teacher visual attention to students in the classroom teacher visual attention in the lesson can be considered as part of the noticing component within teacher professional vision (seidel et al., 2021; chaudhuri, 2023). visual attention is a prerequisite of noticing important events or student characteristics. both screen-based and mobile eye-tracking methods have been helpful in studying teacher noticing and visual attention (grub et al., 2020). eye movement events, such as fixations and saccades, are quantified or examined in scanpaths to follow the focus of teacher visual attention when observing classroom videos, or directly in the process of teaching, captured by the mobile eye tracker (minarikova et al., 2021). the number and duration of fixations on different targets in the classroom are often used as indicators of teacher visual attention in research. fixations are periods of time when the eye is relatively still and acquires new information from the environment, they are the basis of visual attention (holmqvist et al., 2011; duchowski, 2007). another possibility is the visit metric that comprises all fixations in an area of interest (aoi) between the gaze entry and exit of this aoi, with at least one fixation in the aoi (telgmann & müller, 2023; maatta et al., 2021). it is a less sensitive visual attention parameter that represents teacher’s “look” at a student. screen-based studies on teacher professional vision, especially within teacher expertise research, revealed how teachers’ general knowledge relates to their visual processing of classroom scenes. expert teachers tend to pay more attention to students rather than other areas in the classroom (van den bogert et al. 2014; wolff et al., 2016). when assessing student learning profiles, expert teachers monitor more students and show recurring scanning patterns on students, as well as judge student learning dispositions more accurately compared to novices (kosel et al., 2021). besides, experienced teachers notice a higher number of student behavioural cues it the lesson, such as hand-raising, while actively monitoring all students (kosel et al., 2023). in standardised and simulated teaching situations, pre-service teachers tended to focus more on actively participating students and avoid quiet, uninterested, or disrupting ones (goldberg et al., 2021). besides, experienced teachers can recognise subtle changes in students’ engagement and variation in the quality of answers (seidel et al., 2021). in the real-life classroom, teachers work with students over extended periods of time, building context-specific knowledge about their students as individuals. mobile eye tracking research directly in the classroom allows capturing teacher visual attention in the authentic uncontrolled classroom conditions (pouta et al., 2021; huang et al., 2021; mcintyre et al. 2017; mcintyre et al. 2019; cortina et al., 2015). another advantage of mobile eye-tracking research is the possibility to relate teacher visual attention to student-specific information, such as student achievement, learning needs and behaviours. dessus et al. (2016) considered student achievement level and teacher ratings of student self-regulatory behaviours as factor that affected teacher’s gaze allocation between students. the study concluded that more experienced teachers were more likely to distribute gaze based on the student characteristics, but the gaze and student characteristic association was rather weak. at the same time, smidekova et al. (2020) found no association between teacher visual attention and achievement level across several lessons of the same teacher. in the study by chaudhuri et al. (2022), teacher’s fixation counts on the students correlated positively with the amount of teacher-reported individual support to students, and negatively with student scores on math and literacy tests. thus, eye-tracking research shows possible associations between teacher visual attention and student-related characteristics. horlenko, kaminskienė & lehtinen 56 | f l r 1.6 research questions the goal of this study is to investigate the noticing component of teacher professional vision in relation to student srl. in the first step of the study, we aim to identify the degree of agreement between student self-report and teacher rating of student srl. in the second step, we focus on the association between teacher visual attention and student srl (represented as student self-report, teacher rating, and identified joint srl-profile). for the purposes of examining teacher visual attention in relation to student srl in the classroom as part of the noticing component of professional vision, both student self-reports and teacher ratings could be applied. on the one hand, student self-reports have been extensively used in the research on srl. on the other hand, teacher ratings of student srl represent teachers’ judgments of student srl based on observing student learning over time and are part of teacher context-specific knowledge. previous research shows some degree of association between teacher ratings of student learning-related characteristics and teacher visual attention (dessus et al., 2016; chaudhuri et al., 2022). furthermore, previous studies that correlated student srl self-report with teacher ratings did not focus specifically on the relation between student and teacher assessment of student planning, self-monitoring, and helpseeking strategies, in addition to maladaptive regulatory strategies. thus, the research question that explores the extent of the relationship between the student-reported and teacher-rated measures has been formulated: rq1a to what extent do student self-reports and teacher ratings of student srl coincide? a potential mismatch between the student self-report and teacher rating can signify students’ misestimation of own srl (winne & jamieson-noel, 2002), or a need for calibration in teacher’s understanding of students srl involvement. to our knowledge, no previous studies explored the relationship between teachers’ and students’ assessment of student srl in person-oriented analyses. to draw a picture of different student subgroups that form srl-profiles based on combining measures from student and teacher perspectives, the research question was formulated: rq1b which student srl-profiles can be identified based on student self-report and teacher rating of student srl? further, to investigate teacher professional vision in relation to student srl directly in the classroom, we examine whether the amount of teacher visual attention as part of teacher noticing is related to student srl. we examine teacher visual attention in relation to student srl as reported by students, as well as teachers: rq2a is there an association between teacher visual attention and student srl (self-reported and teacher rated)? as some variation between teacher rating and student self-report of srl can be expected, resulting in student srl-profiles, it is also important to consider this variation in relation to teacher visual attention in the classroom: rq2b is there a difference in teacher visual attention distribution between the identified student srl-profiles? 2. methods 2.1 research design, participants, and procedure participants in this study were 10 (female n=8) teachers and their students (n=158) at the high school level in lithuania. the participating teachers and their students were recruited from the university partner schools network following convenience sampling strategy. the participating classes were 9th and 10th grades, with students of 15 – 16 years of age. teachers taught different subjects, such as english, mathematics, biology, physics and lithuanian. teachers’ work experience varied from 2 to 22 years horlenko, kaminskienė & lehtinen 57 | f l r (m=8.2; sd=6.95), all of the teachers had at least a bachelor’s degree and teacher qualification, meeting the minimum state requirements. the signed informed consent forms to participate in the study were collected from teachers, students, and student parents (guardians). the first author attended two lessons of each teacher. in the first lesson, the teachers and students were informed about the study, filled in questionnaires, and were familiarised with the eye-tracking equipment. in the second lesson, the teacher was asked to teach the lesson as usual while wearing the eye-tracking glasses. before the start of the lesson, the researcher helped the teacher to put on the glasses and performed the one-point calibration (tobii pro ab, 2021a). the teacher was instructed to not move the glasses during the recording time. the researcher was present in each lesson that was recorded. the recording length was on average 39 minutes. tobii pro glasses 3 eye tracker was used in all classes. this is a mobile eye tracker resembling usual glasses with a front-looking camera (resolution 1920 × 1080 at 25 fps), a microphone, an eyetracking system for both eyes (two eye cameras and eight infrared illuminators per eye), and a recording unit connected via cable to the glasses frame. the tracker captured eye movement at 100 hz sampling rate with accuracy of 0.6°. the system was operated wirelessly from researcher’s computer (tobii pro ab, 2021a). 2.2 measures 2.2.1 questionnaires student self-report of practicing self-regulated learning. students filled in self-regulation strategy inventory – self-report (srsi-sr; cleary, 2006). it is a 28-item questionnaire with 7-point likert scale (almost never to almost always), divided into three subscales. the initial internal reliability of the questionnaire reported in its validation study was α=.92, subscales ranging from .72 to .88 (cleary, 2006). this questionnaire has been selected because it focuses on both overt and strategic student behaviours associated with srl and has a corresponding validated teacher rating scale (srsi-tr, see below). the subscale managing behaviour and environment included 12 items, with acceptable alpha of .80 in the present sample. this subscale aimed to capture how often students reported self-regulated behaviours such as organising time and environment when studying (“i make a schedule to help me organise my study time”) and strategic behaviours (“i tell myself exactly what i want to accomplish before studying”). subscale seeking and learning information included 8 items (α=.60 in the present sample) with items focusing on students’ help seeking behaviours (“i ask my teacher questions when i do not understand something”). finally, subscale maladaptive regulatory behaviours included 8 items (α =.62) and elicited reports of low regulatory behaviours (“i wait to the last minute to start studying for upcoming tests”). teacher rating of student self-regulated learning. teachers were asked to fill in self-regulation strategy inventory – teacher rating scale (srsi-tr; cleary, & callan, 2014) about each student. the initial questionnaire included 13 items, one item about student attendance of extra consultations was excluded as such consultations were not a common practice at participants’ schools. the questionnaire used 5-point likert scale (almost never to almost always). the original scale was unidimensional, however, for the purposes of this study, exploratory factor analysis was conducted and identified two subscales corresponding to the student questionnaires: managing behaviour and motivation (7 items, α=.96) and seeking and learning information (5 items, α=.93). these instruments are less widely used than other srl questionnaires (tise et al., 2019) and have not been translated into lithuanian previously. the items were translated by a professional translator into lithuanian, and then back into english, the meaning of the items was found to be preserved. 2.2.2 eye-tracking measures visit metric in tobii pro lab software was used to describe teachers’ eye movement in relation to areas of interest (aoi) in the classroom. visit is defined as “all the data between the start of the first fixation inside and aoi to the end of the last fixation in the same aoi” (tobii ab, 2022, p. 124). visit horlenko, kaminskienė & lehtinen 58 | f l r metric was used in previous eye-tracking studies in the classroom (smidekova et al., 2020). two measures based on the visit metric were used: number of visits (or visit count) and visit duration measured in seconds (total and average duration per aoi in a time interval). 2.3 data analysis 2.3.1 eye-tracking data processing and coding the mobile eye tracker yielded a video recording of the lesson from the teacher’s perspective with gaze overlay and audio. this recording was used for coding teacher gaze. the videos with gaze overlay were analysed in tobii pro lab software (version 1.194, tobii ab, 2022). tobii i-vt attention filter has been used, as it is designed for differentiating fixations in dynamic recording conditions. thus, according to filter settings, eye-tracking data points above the velocity threshold of 100 degrees/second and minimum length of 60 milliseconds were classified as fixations (tobii ab, 2022). the first author coded each fixation according to the aoi it was in, the aois included (previously used in mcintyre et al., 2019 and muhonen et al., 2020): student (face and body), student material (worksheet, book, hands with pens during writing), board (white/black/smartboard and projector screen), teacher material (lesson plans, notes, books, teacher’s computer screen), other (non-instructional targets like windows). moments, when the gaze cursor was outside of the screen were coded as unsampled, this code comprised from 0.3 to 5.2 percent of all fixations across teachers. three teachers had logistical situations in the lesson, when they were checking attendance or solving technical problems with the computer, so these intervals were excluded from analyses (5 min 6 sec in total). to ensure the reliability of the coding procedure, pre-defined rules for identifying aois were followed (similar procedure to chaudhuri, 2023). fixation codes were used for calculating visit metrics in the software. the present study focused on teacher’s gaze at students, so the codes student and student material were combined into the overall student metric, as student material areas were relevant for learning situations in the lesson, but not emphasised in the research questions. 2.3.2 statistical analyses to examine the association between student self-report and teacher rating, pearson correlation and k-means cluster analyses were applied. the relation between srl scores and teacher attention indicators were examined with pearson correlations and kruskal-wallis h test. all analyses were performed in ibm spss (v. 28.0.1.1). standardised scores were used for all analyses as student and teacher questionnaires used different scales. 3. results the descriptive information about the analysed variables is presented in table 1. table 1 descriptives m (sd) min. max. skewness kurtosis questionnaire data sr: managing behaviour and environment 4.48(.99) 1.17 6.67 -.346 .460 sr: seeking and learning information 4.66(.93) 1.25 6.63 -.461 .586 sr: maladaptive regulatory behaviours 2.96(.84) 1.00 4.88 .187 -.604 tr: managing behaviour and motivation 3.67(1.07) 1.00 5.00 -.701 -.336 tr: seeking and learning information 3.62(1.11) 1.00 5.00 -.874 -.075 eye movement data number of visits 92.5(69.3) 9.00 448.00 1.663 4.044 total visit duration (s) 59.9(43.3) 2.92 304.00 1.888 4.729 average visit duration (s) 0.64(0.30) .25 1.80 1.962 4.482 note: sr – student self-report, 7-point likert scale; tr – teacher rating, 5-point likert scale. horlenko, kaminskienė & lehtinen 59 | f l r 3.1 to what extent do student self-reports and teacher ratings of student srl coincide? a correlation analysis was performed to explore the relationships between student self-report and teacher rating subscales (table 2). it showed that scores correlated within student self-reports and teacher ratings, but not between these two perspectives, except for student-reported maladaptive regulatory behaviours scale that was slightly negatively correlated with teacher ratings (r=−.275). there was a marginally significant correlation between student-reported seeking and learning information behaviours and teacher-rated managing behaviour and motivation scale (r=.133, p<.1). table 2 pearson correlations between student self-report and teacher rating subscales tr: managing behaviour and motivation tr: seeking and learning information sr: managing behaviour and environment .056 .066 sr: seeking and learning information .133† .110 sr: maladaptive regulatory behaviours -.275** -.228* note: sr – student self-report; tr – teacher rating, * p<.001, ** p<.005, † p < .1. 3.2 which student srl-profiles can be identified based on student self-report and teacher rating of student srl? four student srl-profiles were identified through k-means cluster analysis (table 3; fig. 1). in the initial analyses, different numbers of clusters were considered in an iterative process, and the 4cluster solution was then selected as the most informative. the difference between the profiles appeared based on the extent to which teacher’s rating coincided with student report: two profiles where students’ self-reports and teacher’s ratings were in the same direction (systematic higher-regulated and systematic lower-regulated profiles), and two profiles where either students’ scores were higher than teacher’s (mixed lower-regulated) or students’ scores were lower than teacher’s (mixed higherregulated), with the latter being the largest group (student n=72). figure 1. identified student srl-profiles. 0,32 0,18 0,27 -1,34 -1,32 -0,43 -0,31 0,24 0,39 0,47 0,88 0,9 -0.79 0,57 0,47 -1,32 -1,72 0,73 -1,11 -1,27 -2 -1,5 -1 -0,5 0 0,5 1 1,5 sr: managing behaviour and environment sr: seeking and learning information sr: maladaptive regulatory behaviours tr: manging behaviour and motivation tr: seeking and learning information (1) mixed lower-regulated n=29 (2) mixed higher-regulated n=72 (3) systematic higher-regulated n=44 (4) systematic lower-regulated n=13 horlenko, kaminskienė & lehtinen 60 | f l r the mixed student profiles identified in the cluster analysis above show that students may report both self-regulated and maladaptive behaviours, while correlation analyses show a negative correlation between these two types of behaviour. hence, the person-oriented cluster analysis has provided a more fine-grained picture of self-regulation than the variable-oriented correlation analysis, demonstrating that self-regulatory behaviours are not dichotomous, as students report self-regulatory behaviours along with maladaptive ones. table 3 student srl-profiles based on standardised mean scores of student self-report and teacher rating subscales student profile (1) mixed lower-regulated (2) mixed higher-regulated (3) systematic higher-regulated (4) systematic lower-regulated n=29 n=72 n=44 n=13 sr: managing behaviour and environment .32 -.43 .88 -1.30 sr: seeking and learning information .18 -.31 .90 -1.72 sr: maladaptive regulatory behaviours .27 a .24 a -.79 .73 a tr: manging behaviour and motivation -1.34 d .39 b .57 b -1.11 d tr: seeking and learning information -1.32 e .47 c .47 c -1.27 e note: sr – student self-report; tr – teacher rating a – no significant difference between profile 1, profile 2 and profile 4; profile 3 significantly different from the rest b, c – no significant difference between profile 2 and profile 3 d, e no significant difference between profile 1 and profile 4 the anova procedure with post-hoc scheffe’s test showed that not all mean scores were significantly different across profiles (see notes in table 3). from students’ perspective, all profiles are significantly different in the mean scores of student-reported managing behaviour and environment and seeking and learning information scales. noticeably, positive student-reported maladaptive regulatory behaviours scores in profiles 1, 2 and 4 are not significantly different across these three profiles, but the negative score in profile 3 (systematic higher-regulated) is significantly different from the other three (p<.001). from teachers’ perspective, two groups are distinct: higher-srl students (including both mixed and systematic, profiles 2 and 3 combined) and lower-srl students (including mixed and systematic, profiles 1 and 4 combined), as teacher rating mean scores within each pair are not significantly different. it can be observed that teachers’ positive ratings are close to average even for the systematic higher regulated profile, while the negative ratings are strong, being lower than average by more than 1 sd. the two mixed groups combined (n=101) are larger than systematic ones (n=57), indicating that teacher and student assessments tend to not coincide. 3.3 is there an association between teacher visual attention and student srl (self-reported and teacher rated)? the correlation analysis showed mostly no association between the teacher gaze indicators and student srl scales, except the teacher rating scale seeking and learning information, as seen in table 4. the number of the gaze visits and the total visit duration per student showed slight positive correlations (r=.222 and r=.172 respectively), indicating a connection between teacher’s rating of a student as someone who frequently asks questions in class and the amount of teacher’s attention to that student. horlenko, kaminskienė & lehtinen 61 | f l r table 4 pearson correlations between teacher attention indicators per student and student srl teacher attention indicators number of visits total visit duration (s) average visit duration (s) sr: managing behaviour and environment .001 .033 .040 sr: seeking and learning information -.064 .001 .106 sr: maladaptive regulatory behaviours -.011 -.002 .044 tr: manging behaviour and motivation .102 .037 -.098 tr: seeking and learning information .222** .172* -.027 note: sr – student self-report; tr – teacher rating, ** p < .001, * p < .05 3.4 is there a difference in teacher visual attention distribution between the identified student srl-profiles? non-parametric kruskal-wallis h test was used to check whether there were differences in teacher visual attention indicators between the identified student profiles, resulting in no significant differences for the number of visits (h(3)=.873, p=.83), total visit duration (h(3)=1.40, p=.70), and average visit duration (h(3)=1.08, p=.78) between the profiles. 4. discussion this study focused on the role of students’ srl-related behaviour in teacher professional vision in terms of how teachers assessed student srl, and whether this related to teacher’s visual attention distribution on students during the lesson. it was found that, first, teacher ratings of srl behaviour differed from student self-reports, shown both through correlations and person-oriented analyses. second, analysis of the mobile eye-tracking data showed that teachers’ gaze visit count and total visit duration were moderately associated with teacher rating scale that described students’ tendency to ask questions and seek help in the lessons. the first research question addressed the degree of agreement between student and teacher assessment of srl, including what kinds of student srl-profiles could be identified based on srl reporting from student and teacher perspectives. the results showed that generally, there was a small overlap between the two assessment perspectives, demonstrated by most students having a mixed profile. this finding is somewhat in alignment with previous research reporting teachers’ difficulty to differentiate between student ability and achievement (lavrijsen & verschueren, 2020) or to accurately rate student well-being (urhahne & zhu, 2015). furthermore, teachers’ judgments of student achievement tend to be more accurate than judgments of student motivation and engagement (kaiser et al., 2013). however, the student questionnaire included items that related not only to classroom learning, but also to homework and preparing for tests, while teachers could rely only on student behaviour at school as the basis for rating, which could lead to the differences in judging students’ srl behaviour. reports of maladaptive behaviours appear to be the most distinctive in the analyses. two systematic student profiles, i.e., where teacher and student srl assessment were in the same direction, show the highest (for the lower-regulated profiles) and the lowest (for the higher-regulated profiles) scores of the maladaptive behaviours. besides, the latter was the only student report subscale that correlated with teacher rating. this is consistent with the study by cleary et al. (2006), where the present teacher rating instrument was initially used and negatively correlated with the student-reported maladaptive regulatory behaviours subscale (r=−.41). this may also indicate that the maladaptive student horlenko, kaminskienė & lehtinen 62 | f l r behaviours are more salient for teachers than the strategic ones, thus teachers tend to notice the former. this is also similar to the research showing that teachers partly rely on the off-task behaviour to diagnose students’ srl (dignath & sprenger, 2020). the second research question investigated whether teachers’ visual attention distribution and reports of student srl were related. the only association identified was with seeking and learning information subscale of teacher rating, while no significant correlations were found with the other teacher rating subscale (managing behaviour and motivation) or any of the student self-report subscales. this finding also shows that teacher’s perspective, this time in form of attention allocation to students in the classroom, has a connection to a fairly salient student trait – the tendency to seek information in the classroom – rather than to the relatively covert cognitive and metacognitive regulatory behaviours. the previous mobile eye-tracking classroom studies showed mixed results for the relationship between the amount of teacher gaze and student characteristics: no association with the student achievement level (smidekova et al., 2020), a weak association with student self-regulated behaviour for experienced teachers (dessus et al., 2016), a moderate association with student academic skills (chaudhuri et al., 2022). all these studies reported results from the primary school classrooms. the present study focused on the high school classrooms, where teachers may have different expectations to students and how students demonstrate needs for support. the higher the level of education, the more teachers tend to focus on the study content rather than on the learners (oolbekkink‐ marchand et al., 2007). moreover, srl involvement may be more subtle than student achievement and academic skills. considering the above discrepancy between teacher and student srl assessment, teachers may not reason about students in the srl-related categories when teaching the lesson, especially with the previous studies showing that teachers generally have limited conceptual knowledge of srl processes (dignath & sprenger, 2020; callan & shim, 2019). thus, on the one hand, there may be other student behaviours rather than their usual srl-related practices that attract teacher attention at any given moment in the lesson. also, it is known that more factors, such as student verbal participation (muhonen et al., 2020), teacher movement in the classroom (huang et al., 2023), student position in the classroom (smidekova et al., 2020), and instructional format (stahnke & blömeke, 2021) impact teacher visual attention allocation in the classroom. on the other hand, teachers may be carrying out their lessons as planned, aiming for distributing attention between all students, following more of a top-bottom perspective. there is evidence that experienced teachers can accurately recognise student disengaged behaviour and gaze more on students who appear uninterested or struggling (seidel et al., 2021). still, even if teachers have the knowledge about their individual students’ learning, they may choose not to concentrate on student differences in the lesson and have strategies to address those differences in a more long-term perspective, not captured in the recorded lesson. it is important to note that teachers in the present study taught different subjects. the previous studies show that teachers have different priorities depending on the subject taught, which is reflected in visual behaviours. stahnke & friesen (2023) reported that expert biology teachers focused more on looking for strategies in the classroom management, especially for organising activities and setting up the classroom effectively, while mathematics teachers were concentrated on managing student behaviour and ensuring student participation, which led to variations in attention allocation. in addition, teachers had more frequent and longer fixations on student and student material in the literacy classes comparing to the math classes (huang, 2018). thus, the nature of the subject taught might have created an additional variation in teachers’ attention to students in general, as well as to particular students. finally, the teacher gaze distribution at the classroom level may not be representative of noticing specifically for srl considering the various factors that influence teacher visual attention in the classroom. as the eye movement measures represent both voluntary and involuntary overt attention (duchowski, 2007), triangulation with other data sources, such as verbal reports, can be useful. for example, retrospective stimulated recalls with teachers guided by the lesson recording may shed light on teachers’ rationale behind the visual attention to students and the instructional intentions. thus, more contextualised, and possibly qualitative analyses are needed to connect the instruction, student srlprofiles, student behavioural cues and teacher noticing. nevertheless, by taking the novel approach of incorporating multiple measures and perspectives, this study demonstrates the potential of combining horlenko, kaminskienė & lehtinen 63 | f l r the questionnaire data with the eye-tracking measures to discern the areas of student srl that are more and less noticeable for teachers. overall, theoretically, this study contributes to the development of the evolving field of teacher professional vision for srl (greene, 2021; michalsky, 2014). methodologically, it innovatively combines the data on student srl from students and teachers, as well as the process-based measures of teacher visual attention from the authentic classrooms. 5. limitations, methodological considerations, and implications the present study had several limitations. first, a rather limited sample of 10 teachers can be sensitive to the variability in teacher gaze behaviour and lesson settings and does not provide a possibility to generalise findings. besides, the student overt behaviours in the classroom during lesson recording that had a direct effect on the teacher gaze were not coded, thus the future studies on srl in the classroom could utilise additional ratings of student behaviours, also to explore the relationship between the overt engagement cues and the reported srl. besides, as srl can be both a situational and a long-term process, longitudinal analyses could reveal trends of teacher’s recognition of student srl practices and traits. another limitation is related to questionnaire reliability. two of the student questionnaire subscales had marginal reliability of .60 and .62, which is lower than in the original validation study. this could be due to the change of context where the questionnaire was administered, as it was developed for the usa context and this study took place in lithuania. other studies that used the inventory outside of the usa also reported lower reliabilities (for example, α=.66 in israel reported by madjar et al., 2011). there are methodological considerations related to the authentic classroom conditions of the study. the real-world classroom provided a high ecological validity (jarodzka et al., 2017), still the high variability in the school subjects, teacher behaviours, lesson durations and different lesson phases could influence the results. furthermore, the mobile eye-tracking technology is an innovative tool for data collection, but it is important to note that identification of fixations as attention points of the participant largely depends on the event detection algorithm of the data analysis software. the visit metric used for reporting gaze in the present study was considered informative in the classroom settings, as it captured how many times a teacher entered the student aoi, i.e., looked at the student, rather than made individual fixations on the student. at the same time, the visit metric calculated by the software included not only fixations inside one aoi, but also saccades and blinks (tobii ab, 2022). this is natural for looking in the real-world settings, however, it reduces the possibility to compare results with other classroom mobile eye-tracking research. this study adds to the line of research on teaching for srl, carrying implications for teacher education and further research. the first implication is the need for developing teacher professional vision for srl, and with this bringing the student srl-related behaviours to the focus of teachers, both in regard to teacher selective attention as a bottom-up process, and knowledge about srl for building the top-down perspective. this study has shown that teachers are more prone to noticing salient and maladaptive behaviours of students, highlighting the need for teachers to learn to notice the variations in the regulatory student behaviours and the cognitive, metacognitive, and emotional learning strategy application. one of the steps to promote this can be the development of teachers’ conceptual knowledge on srl (karlen et al., 2020) and using srl as an additional lens for interpreting student behaviours, including maladaptive or disruptive ones. this means that teachers would need to consider student participation in the lessons beyond behavioural engagement as being on-task and following the rules, to the cognitive engagement (fredricks et al., 2004). to this end, interventions can be designed to foster teacher professional vision development for srl for in-service teachers, similarly to the training of the pre-service teachers reported by michalsky (2021). another possibility is to investigate teacher misconceptions about the concept of srl and srl of students (vosniadou et al., 2020) drawing on teachers’ classroom experiences, e.g., following action research approaches. further research on horlenko, kaminskienė & lehtinen 64 | f l r assessing and teaching srl in the authentic classroom conditions could focus on capturing moments in student participation related to srl, as well as teachers’ activities and verbalisations aimed at srl support, and how those interconnect with student learning. keypoints using teacher and student perspectives for assessing srl highlighted the subtleness and complexity of student srl. according to srl-profiles, students can use both self-regulatory and maladaptive strategies, which challenges teachers in assessing student srl. teacher ratings and eye movement show that teachers mostly notice student help-seeking in the classroom. acknowledgments we would like to thank the participating teachers and students. references abar, b., & loken, e. (2010). self-regulated learning and self-directed study in a pre-college sample. learning and individual differences, 20(1), 25–29. https://doi.org/10.1016/j.lindif.2009.09.002 abdoulaye, i. (2003). a study of teacher-student interactions during reading in one-to-one literacy tutoring sessions. [doctoral dissertation, the university of arizona]. retrieved from: http://hdl.handle.net/10150/280396. alexander, p. a., schunk, d. h. & greene, j. a. (eds.). (2017). handbook of self-regulation of learning and performance. routledge. https://doi.org/10.4324/9781315697048 bembenutty, h. (2011). meaningful and maladaptive homework practices: the role of self-efficacy and selfregulation. journal of advanced academics, 22(3), 448–473. https://doi.org/10.1177/1932202x1102200304 callan, g. l., & shim, s. s. (2019). how teachers define and identify self-regulated learning. the teacher educator, 54(3), 295–312. https://doi.org/10.1080/08878730.2019.1609640 callan, g., longhurst, d., shim, s., & ariotti, a. (2022). identifying and predicting teachers’ use of practices that support srl. psychology in the schools, 59(11), 2327–2344. https://doi.org/10.1002/pits.22712 chaudhuri, s. (2023). teachers’ visual focus of attention and related factors in grade 1 classrooms: teacher stress, students’ academic skills and teacher–student relationships. [doctoral dissertation, university of jyväskylä]. retrieved from: https://jyx.jyu.fi/bitstream/handle/123456789/92117/978-951-399859-2_vaitos15122023.pdf?sequence=1&isallowed=y chaudhuri, s., muhonen, h., pakarinen, e., & lerkkanen, m.-k. (2022). teachers’ visual focus of attention in relation to students’ basic academic skills and teachers’ individual support for students: an eyetracking study. learning and individual differences, 98, 102179. https://doi.org/10.1016/j.lindif.2022.102179 cleary, t. j. (2006). the development and validation of the self-regulation strategy inventory—self-report. journal of school psychology, 44(4), 307–322. https://doi.org/10.1016/j.jsp.2006.05.002 cleary, t. j., & callan, g. l. (2014). student self-regulated learning in an urban high school: predictive validity and relations between teacher ratings and student self-reports. journal of psychoeducational assessment, 32(4), 295–305. https://doi.org/10.1177/0734282913507653 https://doi.org/10.1016/j.lindif.2009.09.002 http://hdl.handle.net/10150/280396 https://doi.org/10.4324/9781315697048 https://doi.org/10.1177/1932202x1102200304 https://doi.org/10.1080/08878730.2019.1609640 https://doi.org/10.1002/pits.22712 https://jyx.jyu.fi/bitstream/handle/123456789/92117/978-951-39-9859-2_vaitos15122023.pdf?sequence=1&isallowed=y https://jyx.jyu.fi/bitstream/handle/123456789/92117/978-951-39-9859-2_vaitos15122023.pdf?sequence=1&isallowed=y https://doi.org/10.1016/j.lindif.2022.102179 https://doi.org/10.1016/j.jsp.2006.05.002 https://doi.org/10.1177/0734282913507653 horlenko, kaminskienė & lehtinen 65 | f l r cleary, t. j., slemp, j., & pawlo, e. r. (2021). linking student self‐regulated learning profiles to achievement and engagement in mathematics. psychology in the schools, 58(3), 443–457. https://doi.org/10.1002/pits.22456 corno, l. (2008). on teaching adaptively. educational psychologist 43, 161–173. https://doi.org/10.1080/00461520802178466 cortina, k. s., miller, k. f., mckenzie, r., & epstein, a. (2015). where low and high inference data converge: validation of class assessment of mathematics instruction using mobile eye tracking with expert and novice teachers. international journal of science and mathematics education, 13(2), 389–403. https://doi.org/10.1007/s10763-014-9610-5 credé, m., & phillips, l. a. (2011). a meta-analytic review of the motivated strategies for learning questionnaire. learning and individual differences, 21(4), 337– 346. https://doi.org/10.1016/j.lindif.2011.03.002 de vries, j. a., dimosthenous, a., schildkamp, k., & visscher, a. j. (2023). the impact of an assessment for learning teacher professional development program on students’ metacognition. school effectiveness and school improvement, 34(1), 109–129. https://doi.org/10.1080/09243453.2022.2116461 dessus, p., cosnefroy, o., & luengo, v. (2016). “keep your eyes on ’em all!”: a mobile eye-tracking analysis of teachers’ sensitivity to students. in k. verbert, m. sharples, & t. klobučar (eds.), adaptive and adaptable learning (vol. 9891, pp. 72–84). springer international publishing. https://doi.org/10.1007/978-3-319-45153-4_6 dignath, c., & büttner, g. (2008). components of fostering self-regulated learning among students. a metaanalysis on intervention studies at primary and secondary school level. metacognition and learning, 3(3), 231–264. https://doi.org/10.1007/s11409-008-9029-x dignath, c., & sprenger, l. (2020). can you only diagnose what you know? the relation between teachers’ self-regulation of learning concepts and their assessment of students’ self-regulation. frontiers in education, 5, 585683. https://doi.org/10.3389/feduc.2020.585683 dignath, c., & veenman, m. v. j. (2021). the role of direct strategy instruction and indirect activation of selfregulated learning—evidence from classroom observation studies. educational psychology review, 33(2), 489–533. https://doi.org/10.1007/s10648-020-09534-0 dijkstra, s. h. e., hinne, m., segers, e., & molenaar, i. (2023). clustering children’s learning behaviour to identify self-regulated learning support needs. computers in human behavior, 145, 107754. https://doi.org/10.1016/j.chb.2023.107754 duchowski, a. t. (2007). eye tracking methodology: theory and practice (2nd ed.). springer. fredricks, j. a., blumenfeld, p. c., & paris, a. h. (2004). school engagement: potential of the concept, state of the evidence. review of educational research, 74(1), 59–109. https://doi.org/10.3102/00346543074001059 goldberg, p., sümer, ö., stürmer, k., wagner, w., göllner, r., gerjets, p., kasneci, e., & trautwein, u. (2021). attentive or not? toward a machine learning approach to assessing students’ visible engagement in classroom instruction. educational psychology review, 33(1), 27–49. https://doi.org/10.1007/s10648-019-09514-z goodwin, c. (1994). professional vision. american anthropologist, 96(3), 606–633. http://www.jstor.org/stable/682303 greene, j. a. (2021). teacher support for metacognition and self-regulated learning: a compelling story and a prototypical model. metacognition and learning, 16(3), 651–666. https://doi.org/10.1007/s11409021-09283-7 grub, a.-s., biermann, a., & brünken, r. (2020). process-based measurement of professional vision of (prospective) teachers in the field of classroom management: a systematic review. journal for educational rresearch online 12(3), 75–102. https://doi.org/10.25656/01:21187 heirweg, s., de smul, m., devos, g., & van keer, h. (2019). profiling upper primary school students’ selfregulated learning through self-report questionnaires and think-aloud protocol analysis. learning and individual differences, 70, 155–168. https://doi.org/10.1016/j.lindif.2019.02.001 holmqvist, k., nyström, m., andersson, r., dewhurst, r., jarodzka, h., & van de weijer, j. (2011). eye tracking: a comprehensive guide to methods and measures. oxford university press. huang, y. (2018). learning from teacher’s eye movement: expertise, subject matter and video modeling [doctoral dissertation, university of michigan]. retrieved from: https://doi.org/10.1002/pits.22456 https://doi.org/10.1080/00461520802178466 https://doi.org/10.1007/s10763-014-9610-5 https://psycnet.apa.org/doi/10.1016/j.lindif.2011.03.002 https://doi.org/10.1080/09243453.2022.2116461 https://doi.org/10.1007/978-3-319-45153-4_6 https://doi.org/10.1007/s11409-008-9029-x https://doi.org/10.3389/feduc.2020.585683 https://doi.org/10.1007/s10648-020-09534-0 https://doi.org/10.1016/j.chb.2023.107754 https://doi.org/10.3102/00346543074001059 https://doi.org/10.1007/s10648-019-09514-z http://www.jstor.org/stable/682303 https://doi.org/10.1007/s11409-021-09283-7 https://doi.org/10.1007/s11409-021-09283-7 https://doi.org/10.25656/01:21187 https://doi.org/10.1016/j.lindif.2019.02.001 horlenko, kaminskienė & lehtinen 66 | f l r https://deepblue.lib.umich.edu/bitstream/handle/2027.42/145853/yizhen%20h_1.pdf?sequence=1&is allowed=y huang, y., miller, k. f., cortina, k. s., & richter, d. (2021). teachers’ professional vision in action: comparing expert and novice teacher’s real-life eye movements in the classroom. zeitschrift für pädagogische psychologie, 1–18. https://doi.org/10.1024/1010-0652/a000313 huang, y., richter, e., kleickmann, t., scheiter, k., & richter, d. (2023). body in motion, attention in focus: a virtual reality study on teachers’ movement patterns and noticing. computers & education, 206, 104912. https://doi.org/10.1016/j.compedu.2023.104912 jarodzka, h., holmqvist, k., & gruber, h. (2017). eye tracking in educational science: theoretical frameworks and research agendas. journal of eye movement research, 10(1). https://doi.org/10.16910/jemr.10.1.3 kaiser, j., retelsdorf, j., südkamp, a., & möller, j. (2013). achievement and engagement: how student characteristics influence teacher judgments. learning and instruction, 28, 73–84. https://doi.org/10.1016/j.learninstruc.2013.06.001 kajamies, a. (2017). towards optimal scaffolding of low achievers’ learning: combining intertwined, dynamic, and multi-domain perspectives. [doctoral dissertation, university of turku]. retrieved from: https://research.utu.fi/converis/portal/detail/publication/23505793?lang=fi_fi karlen, y., hertel, s., & hirt, c. n. (2020). teachers’ professional competences in self-regulated learning: an approach to integrate teachers’ competences as self-regulated learners and as agents of selfregulated learning in a holistic manner. frontiers in education, 5. https://doi.org/10.3389/feduc.2020 .00159 kosel, c., böheim, r., schnitzler, k., holzberger, d., pfeffer, j., bannert, m., & seidel, t. (2023). keeping track in classroom discourse: comparing in-service and pre-service teachers’ visual attention to students’ hand-raising behavior. teaching and teacher education, 128, 104142. https://doi.org/10.1016/j.tate.2023.104142 kosel, c., holzberger, d., & seidel, t. (2021). identifying expert and novice visual scanpath patterns and their relationship to assessing learning-relevant student characteristics. frontiers in education, 5. https://doi.org/10.3389/feduc.2020.612175 lavrijsen, j., & verschueren, k. (2020). student characteristics affecting the recognition of high cognitive ability by teachers and peers. learning and individual differences, 78, 101820. https://doi.org/10.1016/j.lindif.2019.101820 maatta, o., mcintyre, n., palomäki, j., hannula, m. s., scheinin, p., & ihantola, p. (2021). students in sight: using mobile eye-tracking to investigate mathematics teachers’ gaze behaviour during task instruction-giving. frontline learning research, 9(4), 92–115. https://doi.org/10.14786/flr.v9i4.965 madjar, n., kaplan, a., & weinstock, m. (2011). clarifying mastery-avoidance goals in high school: distinguishing between intrapersonal and task-based standards of competence. contemporary educational psychology, 36, 268–279. https://doi.org/10.1016/j.cedpsych.2011.03.003 martínez-fernández, j. r., noguera-fructuoso, i., ciraso-calí, a., & vega-martínez, a. (2024). an exploratory study of university students’ regulation profiles and satisfaction with flipped classrooms [estudio exploratorio sobre los perfiles de regulación y la satisfacción con el aula invertida en estudiantes universitarios]. revista española de pedagogía, 82 (287), 111–124. https://doi.org/10.22550/2174-0909.3931 mcintyre, n. a., jarodzka, h., & klassen, r. m. (2019). capturing teacher priorities: using real-world eyetracking to investigate expert teacher priorities across two cultures. learning and instruction, 60, 215– 224. https://doi.org/10.1016/j.learninstruc.2017.12.003 mcintyre, n. a., klassen, r. m., & mainhard, m. t. (2017). are you looking to teach? cultural, temporal and dynamic insights into expert teacher gaze. learning and instruction, 49, 41–53. https://doi.org/10.1016/j.learninstruc.2016.12.005 michalsky, t. (2014). developing the srl-pv assessment scheme: preservice teachers’ professional vision for teaching self-regulated learning. studies in educational evaluation, 43, 214–229. https://doi.org/10.1016/j.stueduc.2014.05.003 https://doi.org/10.1024/1010-0652/a000313 https://doi.org/10.1016/j.compedu.2023.104912 https://doi.org/10.16910/jemr.10.1.3 https://doi.org/10.1016/j.learninstruc.2013.06.001 https://research.utu.fi/converis/portal/detail/publication/23505793?lang=fi_fi https://doi.org/10.3389/feduc.2020.00159 https://doi.org/10.3389/feduc.2020.00159 https://doi.org/10.1016/j.tate.2023.104142 https://doi.org/10.3389/feduc.2020.612175 https://doi.org/10.1016/j.lindif.2019.101820 https://doi.org/10.14786/flr.v9i4.965 https://doi.org/10.1016/j.cedpsych.2011.03.003 https://doi.org/10.22550/2174-0909.3931 https://doi.org/10.1016/j.learninstruc.2017.12.003 https://doi.org/10.1016/j.learninstruc.2016.12.005 https://doi.org/10.1016/j.stueduc.2014.05.003 horlenko, kaminskienė & lehtinen 67 | f l r michalsky, t. (2017). what teachers know and do about assessing students’ self-regulated learning. teachers college record: the voice of scholarship in education, 119(13), 1–16. https://doi.org/10.1177/016146811711901313 michalsky, t. (2021). preservice and inservice teachers’ noticing of explicit instruction for self-regulated learning strategies. frontiers in psychology, 12, 630197. https://doi.org/10.3389/fpsyg.2021.630197 michalsky, t. (2021b). integrating video analysis of teacher and student behaviors to promote preservice teachers’ teaching meta-strategic knowledge. metacognition and learning, 16(3), 595–622. https://doi.org/10.1007/s11409-020-09251-7 minarikova, e., smidekova, z., janik, m., & holmqvist, k. (2021). teachers’ professional vision: teachers’ gaze during the act of teaching and after the event. frontiers in education, 6, 716579. https://doi.org/10.3389/feduc.2021.716579 muhonen, h., pakarinen, e., rasku-puttonen, h., & lerkkanen, m.-k. (2020). dialogue through the eyes: exploring teachers’ focus of attention during educational dialogue. international journal of educational research, 102, 101607. https://doi.org/10.1016/j.ijer.2020.101607 ng, b. (2016). towards lifelong learning: identifying learner profiles on procrastination and self-regulation. new waves educational research & development, 19(1), 41–54. panadero, e. (2017). a review of self-regulated learning: six models and four directions for research. frontiers in psychology, 8, 422. https://doi.org/10.3389/fpsyg.2017.00422 panadero, e., klug, j., & järvelä, s. (2016). third wave of measurement in the self-regulated learning field: when measurement and intervention come hand in hand. scandinavian journal of educational research, 60(6), 723–735. https://doi.org/10.1080/00313831.2015.1066436 peeters, j., de backer, f., kindekens, a., triquet, k., & lombaerts, k. (2016). teacher differences in promoting students’ self-regulated learning: exploring the role of student characteristics. learning and individual differences, 52, 88–96. https://doi.org/10.1016/j.lindif.2016.10.014 pintrich, p. r. (2004). a conceptual framework for assessing motivation and self-regulated learning in college students. educational psychology review, 16(4), 385–407. https://doi.org/10.1007/s10648-004-0006x pintrich, p. r., & de groot, e. v. (1990). motivational and self-regulated learning components of classroom academic performance. journal of educational psychology, 82(1), 33–40. https://doi.org/10.1037/0022-0663.82.1.33 pouta, m., lehtinen, e., & palonen, t. (2021). student teachers’ and experienced teachers’ professional vision of students’ understanding of the rational number concept. educational psychology review, 33(1), 109–128. https://doi.org/10.1007/s10648-020-09536-y puustinen, m., & pulkkinen, l. (2001). models of self-regulated learning: a review. scandinavian journal of educational research, 45(3), 269–286. https://doi.org/10.1080/00313830120074206 rovers, s. f. e., clarebout, g., savelberg, h. h. c. m., de bruin, a. b. h., & van merriënboer, j. j. g. (2019). granularity matters: comparing different ways of measuring self-regulated learning. metacognition and learning, 14(1), 1–19. https://doi.org/10.1007/s11409-019-09188-6 sala, a., punie, y., garkov, v., cabrera giraldez, m. (2020). lifecomp: the european framework for personal, social and learning to learn key competence. publications office of the european union. https://doi.org/10.2760/302967 salo, a.-e., vauras, m., hiltunen, m., & kajamies, a. (2022). long-term intervention of at-risk elementary students’ socio-motivational and reading comprehension competencies: video-based case studies of emotional support in teacher–dyad and dyadic interactions. learning, culture and social interaction, 34, 100631. https://doi.org/10.1016/j.lcsi.2022.100631 seidel, t., & stürmer, k. (2014). modeling and measuring the structure of professional vision in preservice teachers. american educational research journal, 51(4), 739–771. https://doi.org/10.3102/0002831214531321 seidel, t., schnitzler, k., kosel, c., stürmer, k., & holzberger, d. (2021). student characteristics in the eyes of teachers: differences between novice and expert teachers in judgment accuracy, observed behavioral cues, and gaze. educational psychology review, 33(1), 69–89. https://doi.org/10.1007/s10648-02009532-2 smidekova, z., janik, m., minarikova, e., & holmqvist, k. (2020). teachers’ gaze over space and time in a real-world classroom. journal of eye movement research, 13(4). https://doi.org/10.16910/jemr.13.4.1 https://doi.org/10.1177/016146811711901313 https://doi.org/10.3389/fpsyg.2021.630197 https://doi.org/10.1007/s11409-020-09251-7 https://doi.org/10.3389/feduc.2021.716579 https://doi.org/10.1016/j.ijer.2020.101607 https://doi.org/10.3389/fpsyg.2017.00422 https://doi.org/10.1080/00313831.2015.1066436 https://doi.org/10.1016/j.lindif.2016.10.014 https://doi.org/10.1007/s10648-004-0006-x https://doi.org/10.1007/s10648-004-0006-x https://doi.org/10.1037/0022-0663.82.1.33 https://doi.org/10.1007/s10648-020-09536-y https://doi.org/10.1080/00313830120074206 https://doi.org/10.1007/s11409-019-09188-6 https://doi.org/10.2760/302967 https://doi.org/10.1016/j.lcsi.2022.100631 https://doi.org/10.3102/0002831214531321 https://doi.org/10.1007/s10648-020-09532-2 https://doi.org/10.1007/s10648-020-09532-2 https://doi.org/10.16910/jemr.13.4.1 horlenko, kaminskienė & lehtinen 68 | f l r stahnke, r., & blömeke, s. (2021). novice and expert teachers’ noticing of classroom management in whole-group and partner work activities: evidence from teachers’ gaze and identification of events. learning and instruction, 74, 101464. https://doi.org/10.1016/j.learninstruc.2021.101464 stahnke, r., & friesen, m. (2023). the subject matters for the professional vision of classroom management: an exploratory study with biology and mathematics expert teachers. frontiers in education, 8, 1253459. https://doi.org/10.3389/feduc.2023.1253459 telgmann, l., & müller, k. (2023). training & prompting pre-service teachers’ noticing in a standardized classroom simulation – a mobile eye-tracking study. frontiers in education, 8, 1266800. https://doi.org/10.3389/feduc.2023.1266800 theobald, m. (2021). self-regulated learning training programs enhance university students’ academic performance, self-regulated learning strategies, and motivation: a meta-analysis. contemporary educational psychology, 66, 101976. https://doi.org/10.1016/j.cedpsych.2021.101976 tise, j., follmer, d. and sperling, r. (2019). a review of the self-regulation strategy inventory—self-report (srsi-sr). psychology, 10, 305–319. https://doi.org/10.4236/psych.2019.103022 tobii ab (2022). tobii pro lab user manual (version 1.194). tobii ab, danderyd, sweden. tobii pro ab (2021a). pro glasses 3 product description (version 1.6). tobii ab, danderyd, sweden. tobii pro ab (2021b). tobii pro lab (version 1.181.37603 x64) [computer software]. tobii ab, danderyd, sweden urhahne, d., & zhu, m. (2015). accuracy of teachers’ judgments of students’ subjective well-being. learning and individual differences, 43, 226–232. https://doi.org/10.1016/j.lindif.2015.08.007 van de pol, j., volman, m., & beishuizen, j. (2010). scaffolding in teacher–student interaction: a decade of research. educational psychology review, 22(3), 271–296. https://doi.org/10.1007/s10648-0109127-6 van den bogert, n., van bruggen, j., kostons, d., & jochems, w. (2014). first steps into understanding teachers’ visual perception of classroom events. teaching and teacher education, 37, 208–216. https://doi.org/10.1016/j.tate.2013.09.001 van es, e. a., and sherin, m. g. (2002). learning to notice: scaffolding new teachers’ interpretations of classroom interactions. journal of information technology for teacher education, 10(4). van hout-wolters, b. (2000). assessing active self-directed learning. in r.-j. simons, j. van der linden, & t. duffy (eds.), new learning (pp. 83–99). springer. https://doi.org/10.1007/0-306-47614-2_5 veenman, m. v. j., & van cleef, d. (2019). measuring metacognitive skills for mathematics: students’ selfreports versus on-line assessment methods. zdm, 51(4), 691–701. https://doi.org/10.1007/s11858018-1006-5 vermunt, j. d., & verloop, n. (1999). congruence and friction between learning and teaching. learning and instruction, 9(3), 257–280. https://doi.org/10.1016/s0959-4752(98)00028-0 vosniadou, s., lawson, m. j., wyra, m., van deur, p., jeffries, d., & i gusti ngurah, d. (2020). pre-service teachers’ beliefs about learning and teaching and about the self-regulation of learning: a conceptual change perspective. international journal of educational research, 99, 101495. https://doi.org/10.1016/j.ijer.2019.101495 winne, p. h. (2017). cognition and metacognition within self-regulated learning. in d. h. schunk & j. a. greene (eds.), handbook of self-regulation of learning and performance (pp. 36–48). routledge/taylor & francis group. https://doi.org/10.4324/9781315697048-3 winne, p. h., & jamieson-noel, d. l. (2002). exploring students’ calibration of self-reports about study tactics and achievement. contemporary educational psychology, 28, 259-276. https://doi.org/10.1016/s0361-476x(02)00006-1 winne, p. h., & perry, n. e. (2000). measuring self-regulated learning. in m. boekaerts, p. r. pintrich & m. zeidner (eds.), handbook of self-regulation (pp. 531–566). academic press. https://doi.org/10.1016/b978-012109890-2/50045-7 wolff, c. e., jarodzka, h., van den bogert, n., & boshuizen, h. p. a. (2016). teacher vision: expert and novice teachers’ perception of problematic classroom management scenes. instructional science, 44(3), 243–265. https://doi.org/10.1007/s11251-016-9367-z zeidner, m., & stoeger, h. (2019). self-regulated learning (srl): a guide for the perplexed. high ability studies, 30(1–2), 9–51. https://doi.org/10.1080/13598139.2019.1589369 https://doi.org/10.1016/j.learninstruc.2021.101464 https://doi.org/10.3389/feduc.2023.1253459 https://doi.org/10.3389/feduc.2023.1266800 https://doi.org/10.1016/j.cedpsych.2021.101976 https://doi.org/10.4236/psych.2019.103022 https://doi.org/10.1016/j.lindif.2015.08.007 https://doi.org/10.1007/s10648-010-9127-6 https://doi.org/10.1007/s10648-010-9127-6 https://doi.org/10.1016/j.tate.2013.09.001 https://doi.org/10.1007/0-306-47614-2_5 https://doi.org/10.1007/s11858-018-1006-5 https://doi.org/10.1007/s11858-018-1006-5 https://doi.org/10.1016/s0959-4752(98)00028-0 https://doi.org/10.1016/j.ijer.2019.101495 https://psycnet.apa.org/doi/10.4324/9781315697048-3 https://doi.org/10.1016/s0361-476x(02)00006-1 https://psycnet.apa.org/doi/10.1016/b978-012109890-2/50045-7 https://doi.org/10.1007/s11251-016-9367-z https://doi.org/10.1080/13598139.2019.1589369 horlenko, kaminskienė & lehtinen 69 | f l r zimmerman, b. (2001). achieving academic excellence: a self-regulatory perspective. in ferrari, m. (ed.). (2001). the pursuit of excellence through education (1st ed.). routledge. https://doi.org/10.4324/9781410604088 zimmerman, b. j. (2002). becoming a self-regulated learner: an overview. theory into practice, 41(2), 64– 70. zimmerman, b. j. (2013). from cognitive modeling to self-regulation: a social cognitive career path. educational psychologist, 48(3), 135–147. https://doi.org/10.1080/00461520.2013.794676 zimmerman, b. j. (2015). self-regulated learning: theories, measures, and outcomes. in international encyclopedia of the social & behavioral sciences, pp. 541–546. elsevier. https://doi.org/10.1016/b978-0-08-097086-8.26060-1 zimmerman, b. j., & martinez-pons, m. (1988). construct validation of a strategy model of student selfregulated learning. journal of educational psychology, 80(3), 284–290. https://doi.org/10.1037/00220663.80.3.284 zimmerman, b., & kitsantas, a. (2007). reliability and validity of self-efficacy for learning form (self) scores of college students. zeitschrift für psychologie / journal of psychology, 215(3), 157–163. https://doi.org/10.1027/0044-3409.215.3.157 https://doi.org/10.4324/9781410604088 https://psycnet.apa.org/doi/10.1080/00461520.2013.794676 https://doi.org/10.1016/b978-0-08-097086-8.26060-1 https://psycnet.apa.org/doi/10.1037/0022-0663.80.3.284 https://psycnet.apa.org/doi/10.1037/0022-0663.80.3.284 https://doi.org/10.1027/0044-3409.215.3.157 microsoft word maagmerki_publication.docx frontline learning research vol. 10 no. 1 (2022) 1 24 issn 2295-3159 corresponding author: katharina maag merki, university of zurich, institute of education, freiestrasse 36, 8032 zürich, switzerland, kmaag@ife.uzh.ch; doi https://doi.org/10.14786/flr.v10i1.911 preconditions of teachers’ collaborative practice: new insights based on time-sampling data katharina maag merki, urs grob, beat rechsteiner, ariane rickenbacher & andrea wullschleger 1 1university of zurich, switzerland article received 16 july 2021 / article revised 8 february 2022 / accepted 23 march / available online 13 april abstract previous findings on the preconditions of teachers’ collaboration are inconsistent. this might be related to the research methods used to assess the teachers’ collaborative practice. retrospective assessments by self-report on a relatively general level prevail. the validity of these self-reports is limited, however. in contrast, time-sampling methods have the potential to investigate collaborative practice specifically and longitudinally as a day-to-day process over time validly. but to date, no research on collaborative activities in schools based on time-sampling methods is available. in this study, we extended the current state of research by analysing the variability and preconditions of teachers’ collaboration at four secondary schools over three weeks based on time-sampling data collected by a newly developed online practice log. recorded were collaborative activities outside of teaching with a focus on administrative and organisational tasks and on school subject-specific tasks. the results revealed that teachers’ collaborative activities varied significantly between weekdays, showing a linear decrease from monday to friday, regardless of the content of collaboration. collaboration that focused on administrative-organisational tasks seemed to be quite stable over the weeks and was hardly influenced by teachers’ individual characteristics. instead, collaborative activities that focused on school subject-specific tasks varied significantly between weeks; moreover, they were influenced by teachers’ leadership role and gender. the results indicate that rather stable routinised patterns of day-to-day collaboration over the weeks decrease the influence of teachers’ individual characteristics. hence, by collecting data that is closer to content-specific day-to-day collaborative activities, time-sampling methods can be seen as a driver for new insights. keywords: teacher collaboration; time-sampling method; online practice log maag merki et al. 2 | f l r 1. introduction collaborative activities in schools are teachers’ activities that are performed with other teachers and professionals. collaboration is assumed to be important for improving the quality of schools and student learning (antoniou et al., 2015; creemers & kyriakides, 2008; goddard & goddard, 2007). however, research is inconsistent, and it is not clear which factors influence teachers’ collaboration (hargreaves, 2019; vangrieken et al., 2015; vangrieken et al., 2017). this is problematic, as to foster teachers’ collaboration in schools it is important to better understand the preconditions. one reason for this might be related to the research methods implemented to quantitatively assess teachers’ collaboration. retrospective assessment by self-report on a relatively general level prevails (lecat et al., 2020; vangrieken et al., 2015). self-report assessments are less valid for identifying collaboration and its preconditions than other methods such as time-sampling, which allows investigation of collaborative practice as a day-to-day process over time (ohly et al., 2010; reis & gable, 2000). accordingly, analysing teachers’ content-specific day-to-day collaborative activities might increase the understanding of collaboration. in this study, we address this research deficit by applying time-sampling methods for the first time in research on teachers’ collaboration. based on a newly developed online tool, we analysed teachers’ content-specific day-to-day collaborative activities outside of teaching over three weeks. each day, the teachers identified their activities (e.g. reflecting upon individual lessons) and reported if they performed the task cooperatively (see section 4.2). we chose this method based on its potential for education research (zirkel et al., 2015) and on previous studies on capturing principals’ day-to-day practice (camburn et al., 2010; sebastian et al., 2018; spillane & zuberi, 2009) or teachers’ day-to-day classroom practices (adams et al., 2017; elliott et al., 2014; glennie et al., 2017; kurz et al., 2014; rowan & correnti, 2009). those studies support the assumption that a more differentiated view of content-specific day-to-day collaborative activities could also systematically extend previous results on teachers’ collaboration. the benefit can be identified in at least two areas: first, it becomes possible to analyse the variability of teachers’ practice and to learn how stable collaboration is. this increases the understanding of whether collaborative practice can be interpreted as a professional learning community’s organisational routine (sherer & spillane, 2011; spillane et al., 2016) that has the potential to foster student learning (lomos et al., 2011). second, due to the higher validity of the data, the identification of preconditions of collaborative practice is more valid. accordingly, time-sampling methods are considered promising also regarding practical benefits, such as for optimising professional development programs. the aim of this study was to analyse: (a) the variability of teachers’ collaboration outside of teaching on school subject-specific tasks and on administrative and organisational tasks, respectively, over 15 weekdays, and (b) the preconditions of these teachers’ collaborative activities, considering individual and contextual factors as well as the temporal structure of collaborative practice. 2. collaborative activities in schools 2.1 conceptual clarifications systematic literature reviews reveal a variety of concepts, such as professional learning communities (stoll et al., 2006), teacher communities (vangrieken et al., 2017), team learning (decuyper et al., 2010), or teacher cooperation (gräsel et al., 2006). although there are some overlaps between the concepts, they vary in some important aspects. it is therefore important to clarify the concept used in this study. maag merki et al. 3 | f l r first, we follow vangrieken et al.’s (2015) definition of collaboration as an “umbrella term” (p. 23). accordingly, collaborative activities in schools are teachers’ activities that are performed with other teachers and professionals and can vary in terms of quantity and depth of interactions (decuyper et al., 2010; gräsel et al., 2006; havnes, 2009; schippers et al., 2007; yang et al., 2018). in our study, we focus on all activities outside of teaching that were collaboratively performed regardless of the quality of the interactions. second, collaboration in our study is not restricted to teachers but also includes interaction with school leaders and other professionals within and beyond the school (mitchell & sackney, 2011). third, we focus on formal and informal settings, as professional learning is embedded in teachers’ and school leaders’ daily informal work (kyndt et al., 2016; lecat et al., 2020; mitchell & sackney, 2011; van gasse, 2019). fourth, collaboration in schools is related to specific tasks. referring to a typology by vangrieken et al. (2013), possible tasks are management tasks, instruction, innovation and school reform or learning of teacher teams. besides these primary tasks (james et al., 2007), to build schools as learning communities (mitchell & sackney, 2011), “material or practical tasks” (vangrieken et al., 2013, p. 90) must be realised. therefore, we analyse not only school subject-specific collaboration (primary tasks) but also collaboration on administrative and organisational tasks. from a theoretical perspective, collaboration can be understood as a dynamic process that varies over time. it is situated in a complex and rather loosely coupled educational system (weick, 1976) and is strongly embedded in the social context of the schools. further, as empirical studies reveal, collaboration is influenced by multiple dimensions that are intertwined and facilitate or hinder collaboration. accordingly, to identify preconditions of collaboration, a distinction can be made between personal, structural, group, process, and organisational characteristics (kyndt et al., 2016; vangrieken et al., 2015). in this study, we aim to analyse multiple preconditions of collaboration, emphasising individual and school factors. as a new approach, we extend previous research by analysing the effect of the temporal structure on collaboration in schools. 2.2. preconditions of collaboration: previous research 2.2.1 temporal structure of teachers’ collaborative activities there is no research available on the temporal structure of day-to-day teachers’ collaboration in schools. however, there are some studies based on time-sampling or log data that investigated the variability of teachers’ and school leaders’ professional activities. for instance, vannest and parker (2010) analysed classroom time use of special education teachers over nine weeks by time-sampling methods and found that teachers varied considerably in their overall time expenditures and regarding several specific activity types. differences in the variation over time depended significantly on activity type. however, no linear patterns of change could be identified. similarly, rowan and correnti (2009) found an enormous variability of instructional practices from day to day between teachers and between schools. going further, sebastian et al. (2018) shed light on how factors on the school level influenced variation in school principals’ activities. parts of the differences between school activities could be attributed to school factors, the school’s performance level and size and type of school. school context factors tended to predict mostly hour-to-hour variation in principals’ practice (e.g. morning or afternoon). day of the week was important for only one predictor: whether planning occurred on a friday. this activity was positively related to the school’s achievement level. hence, the studies consistently found that there is substantial variability in teachers’ and school leaders’ activities over time (e.g. within days between hours, between days), between school staff, and partially between schools. further, variation is dependent on type and content of activities; however, no consistent pattern of variability has been found. maag merki et al. 4 | f l r 2.2.2 preconditions of teachers’ collaborative activities on the individual level research on individual characteristics reveals no clear picture of the preconditions of teachers’ collaborative practices, as previous analyses are lacking, and existing results are inconsistent. this is particularly true for gender, professional experience, working hours, and teachers’ formal positions, whereas the effect of teachers’ interest seems to be more consistent. a literature review by vangrieken et al. (2015) documented no gender-specific effects. further, kyndt et al. (2016) pointed to only two studies that identified gender as an antecedent of informal teacher learning. an international comparison study (vieluf et al., 2012) and social network analyses (moolenaar, daly, sleegers, et al., 2014) identified gender-specific effects only to some degree, and if effects were identified, they were not consistent. findings on effects of teachers’ age and professional experience are also inconsistent (kyndt et al., 2016). if differences in collaborative practice were found, they were sometimes favourable to novice teachers and sometimes to experienced teachers. for instance, older teachers seem to be asked for advice more often on school subject matters, whereas younger teachers tend to be asked for advice on innovative instructional methods (geeraerts et al., 2018). further, spillane et al. (2018) found that teachers with better performance development were not consulted more often by their colleagues but that they themselves asked others for advice and exchanged ideas with others more frequently. however, formal positions, particularly school principals or teachers having a formal middle leader role, seem to affect collaborative practices (bryant et al., 2020; spillane & kim, 2012; vangrieken et al., 2015; vangrieken et al., 2017). but a general effect is not apparent (moolenaar, daly, cornelissen, et al., 2014), and the effects are dependent on the task of the collaboration (wullschleger et al., 2019). further, differences between part-time and full-time leaders can be identified (spillane & kim, 2012). this is also true for teachers. moolenaar, daly, sleegers, et al. (2014) found that part-time teachers were contacted more often to discuss their own work than full-time teachers. as time is an important prerequisite of teachers’ collaboration (vangrieken et al., 2015), it is argued that full-time teachers have less time to collaborate and that part-time teachers need to compensate their absence from school by getting in touch with other teachers (moolenaar, daly, cornelissen, et al., 2014). also, wullschleger et al. (2019) found that social networks in schools were influenced by the working hours status of the teachers, but the results differed between schools. at one school, there was a homophilyeffect (full-time and part-time teachers collaborated more with other full-time or part-time-teachers, respectively), but at the other school, working hours did not influence collaborative practice. as for teachers’ interest and willingness to collaborate or previous experience in teaming, a clearer picture emerges: they seem to be decisive for teachers’ collaboration (vangrieken et al., 2015; vangrieken et al., 2017). accordingly, teachers’ interest in getting to know what best practice is fosters collaborative knowledge seeking in school teams (mitchell & sackney, 2011). more research is needed, therefore, and time-sampling methods could generate more insights into the complex individual conditions of teachers’ collaboration. 2.2.3 preconditions of teachers’ collaborative activities on the context level besides individual characteristics, it seems that schools affect teacher collaboration, although the effects might be smaller than effects on the individual level. camburn and won han (2017) found small differences between schools in teachers’ reflective practice, which is corroborated by vieluf et al. (2012) and moolenaar, daly, sleegers, et al. (2014), who found that variance in collaboration within schools is greater than variance between schools. nevertheless, school forms and school types (richter & pant, 2016; steinert et al., 2008) or structural aspects of schools (wullschleger et al., 2019) are related to the intensity of teachers’ collaboration in schools. moreover, focusing on school characteristics, vangrieken et al. (2015) reported that various school factors, such as leader support or a school’s culture that supports teaming, have an impact on teachers’ collaboration. also, several school improvement studies pointed to differences between schools in collaborative practice (bryk et al., 2010; hallinger & heck, 2011; muijs et al., 2004). accordingly, it is important to analyse school differences in our study. maag merki et al. 5 | f l r 3. research questions and hypotheses the empirical evidence on the preconditions of collaborative practice is not coherent, and there is a lack of studies that identify the time structure of teachers’ collaborative activities. as previous results on teachers’ and school leaders’ day-to-day activities suggest, the understanding of preconditions of teachers’ collaborative activities could be significantly extended if teachers’ day-to-day collaboration was analysed by time-sampling methods. accordingly, the goal of this study is to identify the variability and preconditions of teachers’ collaborative practice by capturing teachers’ daily practice over three weeks considering different contents of collaboration. to reduce heterogeneity in teachers’ collaborative practice, this study focuses on teachers’ collaborative activities that are performed outside of teaching and investigates two research questions. due to the inconsistency of results and the incompleteness of previous research, the hypotheses are analysed only exploratorily. 3.1 to what extent do collaborative activities outside of teaching vary over the weekdays and over the three weeks (rq1)? h 1a: we expect to find differences between weekdays regarding collaborative activities (rowan & correnti, 2009; sebastian et al., 2018; vannest & parker, 2010). however, due to inconsistent patterns of variabilities between the weekdays in previous research, it is not clear whether the differences follow a linear trend over the weekdays. differential effects in terms of content are not assumed, as in previous research the variation between days was substantial for all analysed types of activities. h 1b: we assume effects of the week on teachers’ collaboration on school subject-specific tasks. we argue that collaborative activities outside of teaching on school subject-specific tasks might differ due to specific demands that vary in intensity across weeks and the school year (for example, professional development training) and due to the only seldom practiced highly intensive, reflexive learning activities of teachers (e.g. camburn & won han, 2017). additionally, it is assumed that for collaboration on administrative and organisational tasks, schools have implemented quite stable schedules and routines that help to structure daily business (feldman & pentland, 2003; sherer & spillane, 2011). consequently, the three weeks are not expected to differ regarding collaboration on administrative and organisational tasks. 3.2 to what extent are teachers’ collaborative activities outside of teaching influenced by schools and by individual factors (rq2)? h 2a: in both content areas of cooperation, we expect to find differences between schools; however, the differences might be rather small (rowan & correnti, 2009; sebastian et al., 2018). h 2b: regarding gender effects, previous research is scarce and inconsistent (moolenaar, daly, sleegers, et al., 2014; kyndt et al., 2016). this is also true for teachers’ work experience at schools (kyndt et al. 2016; geeraerts et al., 2018; spillane et al., 2018) and working hours (moolenaar, daly, sleegers, et al., 2014; wullschleger et al., 2019). therefore, it is not possible to state a clear hypothesis. h 2c: a positive effect of school leaders’ positions on teachers’ cooperation might be seen for cooperative activities on school-subject matters, as school leaders play important roles in building schools’ improvement capacity (mitchell & sackney, 2011) and professional capital (hargreaves & fullan, 2012). however, collaboration on administrative and organisational tasks is probably influenced more by formal structures. accordingly, we do not expect to find a higher level of daily collaboration of school leaders in schools. h 2d: we expect to find a positive relationship between interest in searching for new knowledge and teachers’ collaborative activities (vangrieken et al., 2015; vangrieken et al., 2017). however, we maag merki et al. 6 | f l r assume that this might be particularly true for collaboration on school-subject tasks, as collaboration on administrative and organisational tasks might be regulated more by general guidelines. 4. methodology 4.1 sample we investigated teachers’ collaborative practice in four lower secondary schools (isced 2; 13 to 15-year-old students) in three cantons in the german-speaking part of switzerland. in lower secondary schools, teachers’ collaboration is mandatory. in addition, special education teachers and school social workers complement the regular teaching staff, which also makes collaborative practice necessary. as a basis for school selection, we analysed the social context of the schools, as empirical studies have found that the social context influences school improvement processes and collaborative practice (bryk et al., 2010; muijs et al., 2004; vieluf et al., 2012). it was decisive that the variance between the schools be particularly large. we included two schools in an urban area with a higher vs. lower proportion of foreigners (18.1% vs. 27.4%) (see table 1: schools 1 and 4). in contrast, one school in a rural area was selected (school 2) with a similar proportion of foreigners as the first urban school 1 but with a higher proportion of residents with a university degree than in the first urban school (39% vs. 33%). finally, school 3 differed from the other schools in that it was also located in a suburban area with a similar proportion of foreigners as schools 1 and 2 but with a rapidly growing community and economy in recent years and a particularly high proportion of residents with a university degree (43%). all schools participated voluntarily in this study. of the total population of 105 teachers, 81 participated in the time-sampling study (see table 1). the response rate of 77.1% was high and did not vary between schools (χ2 = 1.12, df = 3, p = .77). most of the teachers who did not participate in the study were on maternity leave. women made up 56.3% of all teachers, which largely corresponds to the distribution in lower secondary schools in the german-speaking part of switzerland. the workload per week of almost two thirds of the teachers (64.9%) was at least 80%, only 21.6% had a workload of less than 60%. the average length of service was 14.6 years (sd = 9.2). moreover, many of the teachers had been working at their school for many years (m = 10.2, sd = 8.2). we did not find any significant school differences in workload or experience. teachers having a leadership position made up 37.0% (n = 30) of the participants, acting as principal or middle leader. before the time sampling started, the teachers had to fill in a questionnaire that assessed individual characteristics. maag merki et al. 7 | f l r table 1 composition of the sample schools populati on teachers response rate time-sampling sub-study community residents school context proportion of foreigners among the community residentsa proportion of residents with tertiary educationb school 1 24 21 (87.5%) 20,077 urban/ regional centre 18.1% 33.0% school 2 23 15 (65.2%) 3,600 rural 15.1% 39.3% school 3 30 23 (76.7%) 8,800 suburban 17.1% 43.4% school 4 28 22 (78.6%) 6,400 urban/ regional centre 27.4% 27.8% total 105 81 (77.1%) a https://www.atlas.bfs.admin.ch/maps/13/de/13559_90_89_70/21857.html [retrieved 1 april, 2021] b https://www.bfs.admin.ch/bfs/de/home/statistiken/bildung-wissenschaft/erhebungen/sba.html [retrieved 1 april, 2021] 4.2 methods to assess collaborative activities: time sampling for three seven-day weeks between fall and christmas 2017, activities were assessed using an online practice log that teachers filled in at the end of each workday (including weekend days, if work had been done) (see figure 1). there was a week’s break between each of the three daily log weeks to reduce teachers’ workload, resulting in a sample of three discontinuous weeks over a five-week period. every day at 5 p.m. teachers received a text message or e-mail with the prompt to log their activities for the day. at the end of the three weeks, we conducted interviews with teachers and the school leaders at each school. the interviews revealed that the teachers had no problems filling in their log. initial analyses point to validity of the practice log (maag merki et al., 2021). figure 1. research design the online practice log was structured by the following three steps: 1. if the teachers worked on a day, they responded to the following: “you are involved in different activities in your school life. please state for each activity which category you assign it to (e.g. teaching)”. they were then asked to identify all their daily activities based on a catalogue of four main categories and 15 subcategories (multiple references were possible). the categories cover the core maag merki et al. 8 | f l r responsibilities of teachers as stated in the laws. for the analyses of the research questions, only the 11 teachers’ activities that took place outside of teaching (see table 2), were included.1 2. the teachers were asked to indicate whether they had performed each activity by themselves or together with others. several response categories were available for selection: e.g. alone, with the school leader, with other individual teacher(s) in the same school, with special needs teacher(s). again, multiple references were possible. 3. after having filled in the log for all daily activities, the teachers rated whether the day had been beneficial for teaching and learning as well as for team and school improvement. additionally, they rated their day in terms of overall stress. however, as this part of the data is not analysed here, we will not go into detail (see maag merki et al., 2021). table 2 main categories and subcategories of the online practice log (only teachers’ activities outside of teaching) main categories subcategories professional development and instruction reflecting upon and further developing individual lessons exchange on school subject-specific questions attending school-internal and -external professional development training studying specialist literature individual feedback (e.g. sitting in on classes) taking part in supervision/intervision team improvement design and further development of teams/work groups school improvement participating in quality management and development (e.g. evaluation, school projects, organisation development) taking part in school conference meetings realisation of tasks for the school administration and organisation exchange on administrative and organisational questions 4.3 methods to assess teachers’ individual characteristics a standardised online questionnaire assessed these teachers’ individual characteristics: gender (0 = female and 1 = male); work experience: how many years have you been working at schools in any kind of role?; and workload (% of a full time equivalent): what is your workload at this school in the current school year? (0: < 80%, 1: ≥ 80%). two scales investigating teachers’ interest in searching for new knowledge (mitchell & sackney, 2011) were included. an internal search interest scale (6 items, cronbach’s alpha = .78; onedimensional) assessed to what extent teachers had a substantial interest in learning how effective their teaching really is, e.g. “please state what you (…) would absolutely like to know for your professional daily routine: absolutely knowing why certain teaching practices do not work well in your own class”. an external search interest scale (6 items, cronbach’s alpha = .67; two-dimensional) included teachers’ substantial interest in ascertaining strategies used by other teachers to successfully promote 1 not included in the analyses were activities such as teaching lessons, talking with students and legal guardians outside of class, or class preparation and follow-up activities such as grading, assessing the competencies of the students. maag merki et al. 9 | f l r students. this scale was two-dimensional: the first dimension was interest in expert knowledge, and the second dimension was interest in the experiences of other teachers, e.g. “please state what you (…) would absolutely like to know for your professional daily routine: absolutely knowing how other teachers teach.” to assess the leadership role of the teachers, we used comprehensive information on the leadership role of every teacher. not only their formal leadership position but also their possible role leading institutionalized working groups at the school (e.g. teachers responsible for steering school improvement processes) were considered (0 = no leadership role, 1 = school leader or teacher with leading role regarding school development). 4.4 data structure and final data base to analyse the complex data structure, we relied on a five-step approach: 1. we analysed all single activities reported during 21 days to get the full picture. in total, we identified 2,642 activities, of which nearly half (44.4%, n = 1,174 activities) were performed outside of teaching. just over half of these activities outside of teaching were performed collaboratively (52.4%, n = 615). on average, teachers reported m = 14.49 (sd = 11.13) activities outside of teaching over the entire data collection period. of these, on average, m = 7.59 (sd = 6.28) activities per teacher were performed collaboratively. 2. we decided to concentrate our analyses on collaborative activities performed during the week from monday to friday, because we considered patterns of practice to be related mostly to weekdays. this led to a small reduction from 615 activities to 598 activities (reduction of 2.8%). across the n = 598 collaborative activities from monday to friday, n = 836 aspects of a collaborative practice were addressed (see table 3). most often, teachers engaged in exchange on administrative and organisational tasks (in 343 collaborative activities outside of teaching; 57.4% of all collaborative activities), followed by collaboration on school subject-specific tasks (31.4%) and on the design and further development of teams/work groups (18.7%). maag merki et al. 10 | f l r table 3 frequency of collaborative activities outside of teaching aspect of collaboration number of activities outside of teaching % of all collaborative activities (n = 598)a a. professional development of teachers and instruction reflecting upon and further developing individual lessons 49 8.2 exchange on school subject-specific questions 188 31.4 attending school-internal and -external professional development training 28 4.7 studying specialist literature 3 0.5 individual feedback (e.g. sitting in on classes) 14 2.3 taking part in supervision/intervision 1 0.2 b. team and school improvement design and further development of teams/work groups 112 18.7 c. school improvement participating in quality management and development (e.g. evaluation, school projects, organisation development) 25 4.2 taking part in school conference meetings 34 5.7 realisation of tasks for the school 39 6.5 d. administration and organisation exchange on administrative and organisational questions 343 57.4 total 836 139.8b a total number of collaborative activities on weekdays (nweekdays = 15) b the sum of all percentages is > 100%, as one activity can have multiple cooperative properties. 3. we concentrated our analyses on collaborative activities that focused on specific aspects only. as teachers had the possibility to address not only one aspect per activity but several (on average 1.40 aspects per activity), it was important to distinguish those activities that focused on school subjectspecific tasks only (see main categories a, b, c in table 3) or on administrative and organisational tasks only (see main category d in table 3). as table 4 depicts, 43.0% of all collaborative activities (n = 257) focused on school subject-specific tasks only, whereas 33.4% of all collaborative activities focused on administrative and organisational tasks only (n = 200). maag merki et al. 11 | f l r table 4 frequency of collaborative activities outside of teaching focusing on school subject-specific tasks only, on administrative and organisational tasks only, and on both types of task simultaneously number of collaborative activities % of all collaborative activities collaborative activities outside of teaching that focused on school subject-specific tasks only 257 43.0% collaborative activities outside of teaching that focused on administrative and organisational tasks only 200 33.4% collaborative activities outside of teaching that focused on both school subject-specific and administrative and organisational tasks 141 23.6% total 598 100% 4. all collaborative activities outside of teaching (n = 598) were aggregated in a binary format on the day level (0 = activity not performed that day, 1 = activity performed that day2). this resulted in 417 (person-) days with any kind of collaborative activities outside of teaching (see table 5). of the 417 (person-) days, 115 (person-) days comprised at least one collaborative activity on administrative and organisational tasks only, 137 (person-) days comprised at least one collaborative activity on school subject-specific tasks only. in sum, 203 (person-) days (48.7%) comprised at least one collaborative activity on school subject-specific tasks only and 181 (person-) days (43.4%) comprised at least one collaborative activity on administrative and organisational tasks only. this resulted in a two-level data structure with level 1 = (person-) days nested in persons on level 2. an additional level (activities within a day) could not be considered, because the number of activities of the persons per day was too low with an average of 1.4 activities per day. 5. finally, two dependent binary variables were calculated, one for collaborative activities outside of teaching with a focus on school subject-specific tasks only and one for collaborative activities outside of teaching with a focus on administrative and organisational tasks only. “1” means that on a day, at least one collaborative activity outside of teaching focused on school subject-specific tasks only (n = 203) and at least one collaborative activity outside of teaching focused on administrative and organisational tasks only (n = 181), respectively. concerning identifying the contrast (= “0”), we aimed to identify preconditions of collaborative practice outside of teaching with a specific single focus. accordingly, as a contrast, we identified activities outside of teaching: (a) that the teachers performed collaboratively but with the contrasting focus, (b) that the teachers performed collaboratively but with a mixed focus, and (c) that the teachers performed alone. this resulted with regard to the dependent variable “collaborative activities that focused on administrative and organisational tasks only” in n = 426 (person-) days and with regard to the dependent variable “collaborative activities that focused on school subject-specific only” in n = 404 (person-) days. in total 607 (person-) days were calculated. 2 in the context of the analysis of collaborative activities, a day record for each teacher was counted valid, if the teacher entered any type of activity outside of teaching in the daily log. for each teacher, days with entries of only teaching-related activities were excluded. maag merki et al. 12 | f l r table 5 distribution of the (person-) days in terms of occurrence of collaborative activity outside of teaching (person-) days with at least one collaborative activity outside of teaching on school subject-specific tasks only 0a 1a total (person-) days with at least one collaborative activity on administrative and organisational tasks only 0a 99 137 236 (56.6%) 1a 115 66 181 (43.4%) total 214 (51.3%) 203 (48.7%) 417 (100%) a 0 = collaborative activity outside of teaching not performed on a day, 1 = collaborative activity outside of teaching activity performed on a day 4.5 analyses we analysed the data by two-level generalised estimating equations (gee) in ibm spss statistics 26 with the daily activities on level 1 and teachers’ characteristics on level 2. in contrast to fixedand random-effects hierarchical models, which account for intrasubject correlation through explicit parameterization, gee instead uses information about the nature of the intracluster dependence to recover more precise estimates of the standard errors (ziegler, 2011). gee was chosen over hierarchical modelling because the given data structure did not meet the sample size-related prerequisites for explicit hierarchical modelling and because of its strengths in handling binary dependent variables. for rq 1, to analyse the temporal structure of the two kinds of collaborative activities, gee with a logistic link function was applied to the two binary outcome variables, with weekdays and weeks as predictors, including the interaction between weekdays and weeks. as previous studies did not show a clear picture in terms of linearity (sebastian et al., 2018; vannest & parker, 2010), two models were estimated: weekdays as a linear covariate (monday = 1 to friday = 5) and weekdays as a categorical factor. for rq 2, we first examined whether there were any significant bivariate (product-moment and point-serial) correlations between the different individual characteristics and the mean percentage of (person-) days of the collaborative activities. moreover, the bivariate correlations (point-serial and phi) correlations among the individual factors were calculated. secondly, to consider the hierarchical data structure of the occurrence of collaborative activities nested in persons as well as multivariate dependencies among the predictor variables, again a gee model with a logistic link function was applied to the two binary outcome variables with gender, work experience, leadership function of the teacher, and the two scales internal and external search interest as predictors on level 2. additionally, school was included as a categorical predictor (4 levels). maag merki et al. 13 | f l r 5. findings 5.1 variation of collaborative activities over the weekdays and over the three weeks (rq 1) the frequency of collaborative activities that focused on administrative and organisational tasks only generally seemed to be highest at the beginning of each of the three weeks and to decline during the week (see figure 2). this was supported by the gee analyses documented in table 6. when including the day of the week as a numerical covariate (column 1), there was a highly significant negative effect (c2 = 13.63, df = 1, p = .000, or = 0.640 for week 5 and non-significant deviations thereof in the two other weeks). this indicated a continuous decline of the frequency of administrative and organisational collaborative activities over the week, irrespective of the week: the mean differences between the three weeks remained within random variation, as did the interaction between weekdays and weeks. when alternatively introducing the weekday as a categorical factor (see table 6, column 2), the results did not change substantially. the activities still varied significantly over the days of the week, and there were no differences between the weeks and no interaction between weekday and week. table 6 weekday and week effects on collaborative activities (significant effects shown in bold) daily collaborative activities with focus on administrative and organisational tasks only daily collaborative activities with focus on school subject-specific tasks only weekdays as linear covariate (monday = 1 to friday = 5) weekdays as categorical factor weekdays as linear covariate (monday = 1 to friday = 5) weekdays as categorical factor weekdays c2 = 13.63a df = 1, p = .000 c2 = 15.33 df = 4, p = .004 c2 = 3.80b df = 1, p = .051 c2 = 9.99 df = 4, p = .041 week c2 = 2.24, df = 2 p = .326 c2 = 0.57, df = 2, p = .752 c2 = 8.43 df = 2, p = .015 c2 = 0.60 df = 2, p = .741 weekday*week c2 = 3.32, df = 2, p = .190 c2 = 10.58 df = 8, p = .227 c2 = 15.44 df = 2, p = .000 c2 = 19.47 df = 8, p = .013 a or = 0.640 per weekday (reference: week 5). b or = 0.734 per weekday (reference: week 5). maag merki et al. 14 | f l r figure 2. percentage of (person-)days with collaborative activities that focused on administrative and organisational tasks only, by weekday and week. the pattern looked different for the frequency of collaborative activities that focused on school subject-specific tasks only (see figure 3 and table 6, columns 3 and 4). as with the administrative and organisational collaborative activities, in weeks 3 and 5 there was a similar steady decline during the course of the week. in week 1, however, the level of school subject-specific collaborative activities was lower on monday, and on friday it was clearly higher. this visual interaction between weeks and weekdays is mirrored in the quantitative analyses in table 6, column 3. possibly due to the diverging activity pattern of week 1, there was, by a thin margin, no significant general linear effect of the weekday. in contrast, the week made a difference, as the interaction between weekday and week was highly significant. when alternatively introducing weekday as categorical factor (table 6, column 4), the effect of weekdays on collaborative daily activities became significant, whereas the effect of weeks was no longer significant. however, the interaction between weekdays and weeks remained significant. figure 3. percentage of (person-)days with collaborative activities that focused on school subjectspecific tasks only, by weekday and week. 5.2 preconditions of teachers’ collaborative activities (rq 2) to assess individual and contextual preconditions, two gee models were run. table 7 shows the results (bivariate correlations are presented in the appendix). regarding the reported daily frequency of collaborative activities that focused on administrative and organisational tasks only, there were no significant effects of the individual factors. when controlling for individual factors, the fixed effect for 0,00 0,10 0,20 0,30 0,40 0,50 0,60 mo tu we th fr pe rc en ta ge o f in di vi du al d ay s week 1 week 3 week 5 0,00 0,10 0,20 0,30 0,40 0,50 0,60 mo tu we th fr pe rc en ta ge o f in di vi du al d ay s week 1 week 3 week 5 maag merki et al. 15 | f l r schools, with four levels, did not add, overall, a significant amount to the prediction of administrative and organisational collaborative activities. however, there were two school contrasts out of a total of six that turned out to be significant (p < .05) when controlling for the individual factors: teachers at school 4 had, on average, more frequent occurrences of administrative and organisational collaborative activities than teachers at school 3 (odds ratio = 2.31) and school 2 (odds ratio = 2.06).3 for daily occurrences of collaborative activities that focused on school subject-specific tasks only, two of the individual factors, gender and leadership, turned out to be predictive, with male teachers being less involved (odds ratio = 0.62, p < .05) and teachers with leadership roles being more involved (odds ratio = 2.18, p < .001) in school subject-specific collaboration. when controlling for all individual factors, there was no general school effect, but again, two school contrasts regarding the mean occurrence of subject-specific collaborative activities were significant: teachers at school 3 showed a higher mean frequency of subject-specific collaborative activities than teachers at school 2 (odds ratio = 2.40, p < .05) and at school 4 (odds ratio = 1.97, p < .05). table 7 prediction of occurrence of daily administrative and organisational and of school subject-specific collaborative activities daily collaborative activities with focus on administrative and organisational tasks only daily collaborative activities with focus on school subject-specific tasks only gender (0 = f, 1 = m) c2 = 0.16, df = 1, p = .692, or = 0.91 c2 = 4.09, df = 1, p = .043, or = 0.62 work experience (standardised) c2 = 1.59, df = 1, p = .208, or = 0.83 c2 = 1.04, df = 1, p = .307, or = 1.14 teachers’ workload (0 = <80%, 1 = ≥80%) c2 = 0.03, df = 1, p = .856, or = 1.06 c2 = 0.13, df = 1, p = .721, or = 1.10 leadership function (0 = no, 1 = yes) c2 = 0.09, df = 1, p = .763, or = 1.08 c2 = 12.27, df = 1, p = .000, or = 2.18 internal interest in searching for knowledge (standardised) c2 = 0.19, df = 1, p = .660, or = 1.09 c2 = 0.43, df = 1, p = .511, or = 0.90 external interest in searching for knowledge (standardised) c2 = 0.03, df = 1, p = .856, or = 0.965 c2 = 1.451, df = 1, p = .228, or = 1.19 school (ref. = school 4) c2 = 7.70, df = 3, p = .053; school 4 > school 3 (c2 = 6.06, df = 1, p = .014, or = 2.31) school 4 > school 2 (c2 = 4.58, df = 1, p = .032, or = 2.06) c2 = 5.82, df = 3, p = .121; school 3 > school 2 (c2 = 5.03, df = 1, p = .025, or = 2.40) school 3 > school 4 (c2 = 3.99, df = 1, p = .046, or = 1.97) 3 due to the double relative character of odds ratios (a ratio of two ratios) and their dependency on base probabilities, an interpretation in terms of (normed) effect sizes is not possible (best & wolf, 2012). maag merki et al. 16 | f l r 6. discussion due to the lack of studies analysing the temporal structure of teacher collaboration and due to inconsistent results in previous research, the aim was to exploratorily analyse the variability and preconditions of teachers’ collaboration outside of teaching. we implemented a newly developed online practice log over 3 weeks. 6.1 variation of collaborative activities outside of teaching (rq 1) collaborative practice outside of teaching varies significantly between the days of the week, regardless of the content of the collaboration. this supports previous research (rowan & correnti, 2009; sebastian et al., 2018; vannest & parker, 2010). further, there is a linear decrease over the weekdays from monday to friday for both collaboration contents, although the linearity is stronger for collaboration on administrative and organisational tasks than for collaboration on school subjectspecific tasks. teachers might start a new week by collaboratively organising the upcoming week and discussing the issues at hand. as the week proceeds, the frequency of collaboration decreases, be it due to proceeding based on a division of work, or a decrease in perceived need. another explanation could be related to teachers’ perceived autonomy (vangrieken & kyndt, 2020). teachers’ perceived autonomy is crucial for collaboration. it might be that to fulfil all tasks, teachers’ need for autonomy is greater towards the end of the week than at the beginning of the week. due to the different profiles of teachers’ perceived autonomy need (vangrieken & kyndt, 2020), it would be interesting to analyse if the decrease depends on these profiles. it could be that the decline of the collaborative practice over the week in the profile ‘autonomous collaborative’ (which reported most collaboration), is less strong then in the other profiles. further, it would be interesting to learn if the collaboration activities performed at the beginning of the week also qualitatively outperform those on the other days (e.g. decuyper et al., 2010; yang et al., 2018). if the quality also varies, it would be important not only to analyse whether the variable frequency of collaborative practice over the weekdays fits the need of fulfilling the tasks even at the end of the week but also to think about fostering the quality of collaborative practice throughout the week, for instance by external support (gutierez, 2015; camburn & won han, 2017). as there were no differences between weeks, collaboration that focuses on administrativeorganisational tasks only seems to be quite stable, as we expected in h1b. at best, these patterns could be interpreted as routines (feldman & pentland, 2003; sherer & spillane, 2011; spillane et al., 2016) that help teachers deal with challenges. routines can be interpreted as a resource for stabilizing the work. however, if the routines do not fit the requirements for successful further development of school processes, they might be barriers rather than drivers for sustainable school improvement. in contrast, collaborative activities that focus on school subject-specific tasks only vary significantly between the weeks. this is also in line with our hypothesis h1b and indicates that this type of collaboration does not strictly follow a routinized pattern but that extraordinary events affect the modus of activities in one particular week, for instance highly intensive and reflexive learning activities (e.g. camburn & won han, 2017). as figure 3 shows, it is particularly week 1 that differs from the other weeks. further analyses revealed that it is mainly school 1 where teachers’ collaborative activities are mostly visible on friday of the first week; on precisely that friday a school-internal professional development training took place. training courses in schools are not performed every week, so these activities peak every now and then. the new insights into the variability of collaborative practice in schools have important practical implications for school improvement. they could help school leaders, as the driving force for school improvement (bryk et al., 2010), to organise collaborative practice dependent on teachers’ capacity and the content-specific need to collaborate. this targeted collaboration has potential to increase the efficiency of collaboration in professional learning communities (stoll & louis, 2007; vescio et al., 2008). maag merki et al. 17 | f l r 6.2 preconditions of collaborative activities outside of teaching (rq 2) although in the literature many predictors were identified, in our study the collaborative activities are hardly influenced by teachers’ individual or school characteristics. however, there are differential patterns of effects dependent on the area of collaboration. in line with h2a (sebastian et al.; 2018; rowan & correnti, 2009), this study reveals only weak school effects. for teachers’ collaborative practice on both administrative-organisational and school subject-specific tasks, there was no overall school effect, but two significant contrasts between schools. school 4 outperformed schools 2 and 3 with a higher level of collaborative practice that focused on administrative and organisational tasks. school 4 is the school with the highest proportion of foreigners and the lowest proportion of residents with tertiary education in the community (see table 1). we consider this socially challenging context to be relevant, as a denser coordination of administrative and organisational work could help students and parents to deal successfully with the demands of the school. for instance, in schools with a higher proportion of migrant children, teachers must coordinate on information materials more often and prepare together for meetings with the parents or with social authorities (muijs et al., 2004). in schools with a higher proportion of children with a higher socioeconomic and non-migration background, on the other hand, the parents have more knowledge about the educational system, as the cultural fit between schools and families is higher (kramer, 2014). as a result, the need for collaboration on administrative and organisational tasks in these schools might be lower. to corroborate this hypothesis, however, it would be important to increase the size of the school sample. individual characteristics turn out to be more relevant for collaboration on school subjectspecific than on administrative-organisational tasks. we argued in h2c and h2d that collaboration on administrative and organisational tasks might be prescribed more by school law and formal structures. accordingly, as we did not find any significant effect of the tested individual characteristics, it seems that there is only little freedom for teachers to decide whether they want to collaborate on administrative and organisational tasks or not. remarkably, also teachers’ interest in searching for new knowledge, internal or external, is not associated with a higher level of teachers’ day-to-day collaborative activities, not even on school subject-specific tasks. this result is clearly contrary to our expectations in h2d. one explanation could be again that the legal regulations in switzerland oblige all teachers to collaborate outside of teaching with other teachers. accordingly, variation is rather small. another reason might be that the rather stable pattern of day-to-day collaboration could be interpreted as a professional learning community practice with established cooperation-related norms (sleegers et al., 2013; stoll & louis, 2007; vescio et al., 2008). also in this case, teachers’ individual interests would be less influential, as schools’ organisation, which is an important facilitating factor for teachers’ collaboration (vangrieken et al., 2015), largely frames teachers’ activities. this result has important practical implications, as teachers differ substantially in their motivation and attitudes towards collaboration (e.g. vangrieken & kyndt, 2020). if collaboration belongs to the professional profile of teachers, it probably makes it easier for teachers to collaborate. regarding collaboration on school subject-specific tasks, we found a most significant effect of leadership function and, somewhat less pronounced, of gender. teachers who have specific responsibilities in leading the school or a subgroup of teachers are more often involved in collaborative practices in terms of reflecting and developing the quality of schooling (see h2c). this supports those studies that found empirical evidence for school leaders to have a higher probability to be involved in discussions about work-related issues than teachers without such a role (moolenaar, daly, sleegers, et al., 2014; spillane & kim, 2012). this result could be interpreted through the lens of professional learning communities (sleegers et al., 2013; stoll & louis, 2007; vescio et al., 2008), as they provide important routines to foster school improvement (spillane et al., 2016). moreover, concepts of distributed or shared leadership (leithwood et al., 2020; spillane & mertz, 2015) could be a reason why school leaders are more involved in collaboration on school subject-specific tasks. the importance of maag merki et al. 18 | f l r leadership for realising processes and activities in schools is also supported by theoretical models such as building learning-community capacity (mitchell & sackney, 2011). further, our results support those studies that indicate a higher frequency of collaboration by women (vieluf et al., 2012). however, this only relates to forms of cooperation that are more freely chosen (subject-specific) and not those forms that are more regulated (administrative and organisational). going beyond self-reports on a general level, time-sampling data can capture not only the frequency but also the variability and stability of teacher collaboration on a daily level. as the analysed temporal structure of the collaborative activities helps to better capture the degree of variability between teachers (ohly et al., 2010), the relevance of the analysed preconditions can be better estimated. if the results regarding the preconditions of collaboration as a routinized pattern of school improvement persist in further studies, they are most relevant for fostering teacher collaboration in schools. the most efficient way, then, would be to build professional learning communities (sleegers et al., 2013; stoll & louis, 2007; vescio et al., 2008) rather than to convince teachers to strive for more collaboration selfresponsibly. 6.3. conclusion and limitations taken together, our results provide evidence that the collaborative practices in both areas follow a rather standardised pattern over the days and weeks and are more strongly related to school characteristics and formal responsibilities than to teachers’ individual characteristics. additionally, differential effects dependent on the area of collaboration are visible. whereas schools are relevant for teachers’ collaboration on administrative and organisational tasks, school leadership roles of teachers affect collaboration on school subject-specific tasks. in line with the highly different content-specific pattern of teachers’ or school leaders’ work (rowan & correnti, 2009; sebastian et al., 2018), predictors of teachers’ collaborative activities vary in terms of the content of the collaboration. however, in previous theoretical frameworks and in empirical literature reviews this result has not been sufficiently recognised up to now. it therefore seems important to consider these differential effects in the theoretical models more consistently. further, it would be important to analyse the extent to which the different patterns of collaboration influence school development and student learning. from a learning sciences perspective, the goal of teachers’ collaboration is to build human capital, to develop the quality of teaching, and finally to enhance student learning (hargreaves & fullan, 2012). again, previous research is inconsistent in this regard, and a more performance-based analysis could add substantially to the existing body of research (kyndt et al., 2016; vangrieken et al., 2015; vangrieken et al., 2017). accordingly, time-sampling methods are convincing in identifying variation in collaboration practice but also school-specific patterns of collaboration and their preconditions. for the identification of predictors for teachers’ collaborative activities outside of teaching, the time-sampling method is particularly suitable, as the bias of remembering the frequency and content of the collaborative practices of teachers is much smaller, although self-reported in nature as well, than with standard questionnaires where teachers have to refer to a long period of time (ohly et al., 2010). nevertheless, it must be kept in mind that the results may also be influenced by the selection of the weeks: if we had chosen another week for our data collection or included more schools, the results could have been different. this is particularly true for activities that are known to be only rarely performed (e.g. camburn & won han, 2017), for instance getting individual feedback by sitting in on classes (see table 3). they could have been identified more often if we had chosen other weeks. further, the length and positioning of the time periods examined are crucial regarding obtaining valid information on teachers’ collaborative practice. although our results are in line with previous results that show that deep-level teacher collaboration, which demands high intensity, critical discussion, and reflection or introspection, is seldom practiced (e.g. camburn & won han, 2017; muckenthaler et al., 2020; oecd, 2016), an increase in the number maag merki et al. 19 | f l r of weeks might lead to a more generalisable picture of collaborative practice and could reduce idiosyncratic results. however, this is dependent on the research question and on the motivation of teachers to fill in an online practice log over a longer period (ohly et al., 2010; vannest & parker, 2010). further, it might be important to sample weeks during the whole school year and not only in the first part, as we did in this study. another important limitation of this study is that it was not possible to capture the quality of the collaborative activities and with whom the teachers collaborated. theoretical models on school effectiveness and school improvement suggest that the quality is an influencing factor to explain differences between the effects of interventions (creemers & kyriakides, 2008). further, the duration of the activities analysed was not specified. as we looked only at daily occurrences of activities (yes/no), activities of a very short duration and activities engaged in for a longer time were treated in the same way. therefore, future studies should consider the duration of the activities. additionally, we were only able to differentiate two areas of collaboration, whereas research on teachers’ and school leaders’ (rowan & correnti, 2009; sebastian et al., 2018; vannest & parker, 2010) activities suggest that there is quite substantial intra-individual variation in terms of the content of activities. therefore, our analyses could be extended by differentiating the collaboration contents further in future studies. due to the sample size, this was not possible in this study. accordingly, it would be fruitful to repeat this study with a much larger representative sample of schools that allows multilevel analyses. this would also be beneficial for analysing the preconditions in greater detail, considering also structural, group and process characteristics (kyndt et al., 2016; vangrieken et al., 2015). finally, it is important to remember that also time-sampling data are self-reported data with some limitations, although with a reduced bias in remembering the performed activities (glennie et al., 2017; moeller et al., 2020; ohly et al., 2010). hence, this approach seems to have potential not only for analysing school leaders’ and teachers’ practices but also school leaders’ and teachers’ collaborative activities. keypoints teachers’ collaborative activities decrease linearly from monday to friday. collaboration on administrative-organisational tasks is quite stable and varied only in dependency on the school context. in contrast, collaborative activities on school subject-specific tasks varied between the weeks and by teachers’ school-related leadership role and gender. routinised patterns of day-to-day collaboration may decrease the influence of teachers’ individual characteristics. time-sampling methods are a driver for new insights into the content-specific day-to-day collaboration of teachers. maag merki et al. 20 | f l r references adams, e. l., carrier, s. j., minotue, j., porter, s. r., mceachin, a., walkowiak, t. a., & zulli, r. a. (2017). the development and validation of the instructional practices log in science: a measure of k-5 science instruction. international journal of science education, 39(3), 335-357. https://doi.org/10.1080/09500693.2017.1282183 antoniou, p., kyriakides, l., & creemers, b. p. m. (2015). the dynamic integrated approach to teacher professional development: rationale and main characteristics. teacher development, 19(4), 535552. https://doi.org/10.1080/13664530.2015.1079550 best, h., & wolf, c. (2012). modellvergleich und ergebnisinterpretation in logit-und probit-regressionen. kzfss kölner zeitschrift für soziologie und sozialpsychologie, 64(2), 377-395. https://doi.org/10.1007/s11577-012-0167-4 bryant, d. a., lun, y., & adams, a. (2020). how middle leaders support in-service teachers’ on-site professional learning. international journal of educational research, 100(101530). https://doi.org/10.1016/j.ijer.2019.101530 bryk, a. s., bender sebring, p., allensworth, e., luppescu, s., & easton, j. q. (2010). organizing schools for improvement. lessons from chicago. university of chicago press. camburn, e. m., spillane, j. p., & sebastian, j. (2010). assessing the utility of a daily log for measuring principal leadership practice. educational administration quarterly, 46(5), 707-737. https://doi.org/10.1177/0013161x10377345 camburn, e. m., & won han, s. (2017). teachers’ professional learning experiences and their engagement in reflective practice: a replication study. school effectiveness and school improvement, 28(4), 527554. https://doi.org/10.1080/09243453.2017.1302968 creemers, b. p. m., & kyriakides, l. (2008). the dynamics of educational effectiveness. a contribution to policy, practice and theory in contemporary schools. routledge. decuyper, s., dochy, f., & van den bossche, p. (2010). grasping the dynamic complexity of team learning: an integrative model for effective team learning in organisations. educational research review, 5, 111-133. https://doi.org/10.1016/j.edurev.2010.02.002 elliott, s. n., roach, a., t., & kurz, a. (2014). evaluating and advancing the effective teaching of special educators with a dynamic instructional practices portfolio. assessment for effective intervention, 39(2), 83-98. https://doi.org/10.1177/1534508413511491 feldman, m. s., & pentland, b. t. (2003). reconceptualizing organizational routines as a source of flexibility and change. administrative science quarterly, 48, 94-118. https://doi.org/10.2307/3556620 geeraerts, k., tynjälä, p., & heikkinen, h. l. t. (2018). inter-generational learning of teachers: what and how do teachers learn from older and younger colleagues? european journal of teacher education, 41(4), 479-495. https://doi.org/10.1080/02619768.2018.1448781 glennie, e. j., charles, k. j., & rice, o. n. (2017). teacher logs: a tool for gaining a comprehensive understanding of classroom practices. science educator, 25(2), 88-96. goddard, y. l., & goddard, r. d. (2007). a theoretical and empirical investigation of teacher collaboration for school improvement and student achievement in public elementary schools. teachers college record, 109(4), 877-896. https://doi.org/10.1177/016146810710900401 gräsel, c., fussangel, k., & parchmann, i. (2006). lerngemeinschaft in der lehrerfortbildung. kooperationserfahrungen und -überzeugungen von lehrkräften. zeitschrift für erziehungswissenschaft, 9(4), 545-561. https://doi.org/10.25656/01:4453 gutierez, s. b. (2015). teachers’ reflective practice in lesson study: a tool for improving instructional practice. alberta journal of educational research, 63(3), 314-328. https://doi.org/10.11575/ajer.v61i3.56087 hallinger, p., & heck, r. h. (2011). exploring the journey of school improvement: classifying and analyzing patterns of change in school improvement processes and learning outcomes. school effectiveness and school improvement, 22(1), 1-27. https://doi.org/10.1080/09243453.2010.536322 hargreaves, a. (2019). teacher collaboration: 30 years of research on its nature, forms, limitations and effects. teacher and teaching, 25(5), 603-621. https://doi.org/10.1080/13540602.2019.1639499 hargreaves, a., & fullan, m. (2012). professional capital. transforming teaching in every school. teachers college press. maag merki et al. 21 | f l r havnes, a. (2009). talk, planning and decision‐making in interdisciplinary teacher teams: a case study. teachers and teaching: theory and practice, 15(1), 155-176. https://doi.org/10.1080/13540600802661360 james, c. r., dunning, g., connolly, m., & elliott, t. (2007). collaborative practice: a model of successful working in schools. journal of educational administration, 45(5), 541-555. https://doi.org/10.1108/09578230710778187 kramer, r.-t. (2014). kulturelle passung und schülerhabitus. zur bedeutung der schule für transformationsprozesse des habitus. in w. helsper & r.-t. kramer (eds.), schülerhabitus (pp. 183-202). springer vs. kurz, a., elliott, s. n., kettler, r. j., & yel, n. (2014). assessing students' opportunity to learn the intended curriculum using an online teacher log: initial validity evidence. educational assessment, 19(3), 159-184. https://doi.org/10.1080/10627197.2014.934606 kyndt, e., gijbels, d., grosemans, i., & donche, v. (2016). teachers’ everyday professional development: mapping informal learning activities, antecedents, and learning outcomes. review of educational research, 86(4), 1111-1150. https://doi.org/10.3102/0034654315627864 lecat, a., spaltman, y., beausaert, s., raemdonck, i., & kyndt, e. (2020). two decennia of research on teachers’ informal learning: a literature review on definitions and measures. educational research review, 30, 1-15. https://doi.org/10.1016/j.edurev.2020.100324 leithwood, k., harris, a., & hopkins, d. (2020). seven strong claims about successful school leadership revisited. school leadership & management, 40(1), 5-22. https://doi.org/10.1080/13632434.2019.1596077 lomos, c., hofman, r. h., & bosker, r. j. (2011). the relationship between departments as professional communities and student achievement in secondary schools. teaching and teacher education, 27(4), 722–731. https://doi.org/10.1016/j.tate.2010.12.003 maag merki, k., grob, u., rechsteiner, b., wullschleger, a., schori, n., & rickenbacher, a. (2021). regulation activities of teachers in secondary schools: development of a theoretical framework and exploratory analyses in four secondary schools based on time sampling data. in a. oude groote beverborg, t. feldhoff, k. maag merki, & f. radisch (eds.), concept and design developments in school improvement research. longitudinal, multilevel and mixed methods and their relevance for educational accountability (pp. 257-301). springer international publishing. https://link.springer.com/book/10.1007/978-3-030-69345-9 mitchell, c., & sackney, l. (2011). profound improvement: building learning-community capacity on living-system principles. routledge. moeller, j., viljaranta, j., kracke, b., & dietrich, j. (2020). disentangling objective characteristics of learning situations from subjective perceptions thereof, using an experience sampling method design. frontline learning research, 8(3), 63-84. https://doi.org/10.14786/flr.v8i3 moolenaar, n. m., daly, a. j., cornelissen, f., liou, y.-h., caillier, s., riordan, r., wilson, k., & cohen, n. a. (2014). linked to innovation: shaping an innovative climate through network intentionality and educators' social network position. journal of educational change, 15(2), 99-123. https://doi.org/10.1007/s10833-014-9230-4 moolenaar, n. m., daly, a. j., sleegers, p., j., & karsten, s. (2014). social forces in school teams. in d. zandvliet, p. brok den, t. mainhard, & j. van tartwijk (eds.), interpersonal relationships in education: from theory to practice (pp. 159-181). sensepublishers. muckenthaler, m., tillmann, t., weiss, s., & kiel, e. (2020). teacher collaboration as a core objective of school development. school effectiveness and school improvement, 31(3), 486-504. https://doi.org/10.1080/09243453.2020.1747501 muijs, d., harris, a., chapman, c., stoll, l., & russ, j. (2004). improving schools in socioeconomically disadvantaged areas: a review of research evidence. school effectiveness and school improvement, 15(2), 149-175. https://doi.org/10.1076/sesi.15.2.149.30433 oecd. (2016). supporting teacher professionalism: insights from talis 2013. oecd. http://dx.doi.org/10.1787/9789264248601-en ohly, s., sonnentag, s., niessen, c., & zapf, d. (2010). diary studies in organizational research: an introduction and some practical recommendations. journal of personnel psychology, 9(2), 79-93. https://doi.org/10.1027/1866-5888/a000009 maag merki et al. 22 | f l r reis, h. t., & gable, s. l. (2000). event-sampling and other methods for studying everyday experience. in h. t. reis & c. m. judd (eds.), handbook of research methods in social and personality psychology (pp. 190-222). cambridge university press. richter, d., & pant, h. a. (2016). lehrerkooperation in deutschland: eine studie zu kooperativen arbeitsbeziehungen bei lehrkräften der sekundarstufe i. bertelsmann stiftung, robert bosch stiftung, stiftung mercator, deutsche telekom stiftung. rowan, b., & correnti, r. (2009). studying reading instruction with teacher logs: lessons from the study of instructional improvement. educational researcher, 38(2), 120-131. https://doi.org/10.3102/0013189x09332375 schippers, m. c., den hartog, d. n., & koopman, p. l. (2007). reflexivity in teams: a measure and correlates. applied psychology: an international review, 56(2), 189-201. https://doi.org/10.1111/j.1464-0597.2006.00250.x sebastian, j., comburn, e. m., & spillane, j. p. (2018). portraits of principal practice: time allocation and social principal work. educational administration quarterly, 54(1), 47-84. https://doi.org/10.1177/0013161x17720978 sherer, j. z., & spillane, j. p. (2011). constancy and change in work practice in schools: the role of organizational routines. teachers college record, 113(3), 611-657. https://doi.org/10.1177/016146811111300302 sleegers, p. j. c., den broke, p., verbiest, e., moolenaar, n. m., & daly, a. j. (2013). toward conceptual clarity. a multidimensional, multilevel model of professional learning communities in dutch elementary schools. the elementary school journal, 114(1), 118-137. https://doi.org/10.1086/671063 spillane, j. p., & kim, c. m. (2012). an exploratory analysis of formal school leaders' positioning in instructional advice and information networks in elementary schools. american journal of education, 119(1), 73-102. https://www.jstor.org/stable/10.1086/667755 spillane, j. p., & mertz, k. (2015). distributed leadership. oxford bibliographies. https://doi.org/10.1093/obo/9780199756810-0123 spillane, j. p., shirrell, m., & adhikari, s. (2018). constructing “experts” among peers: educational infrastructure, test data, and teachers’ interactions about teaching. educational evaluation and policy analysis, 40(4), 586-612. https://doi.org/10.3102/0162373718785764 spillane, j. p., shirrell, m., & hopkins, m. (2016). designing and deploying a professional learning community (plc) organizational routine: bureaucratic and collegial arrangements in tandem. les dossiers des sciences de l’éducation, 35, 97-122. https://doi.org/10.4000/dse.1283 spillane, j. p., & zuberi, a. (2009). designing and piloting a leadership daily practice log: using logs to study the practice of leadership. educational administration quarterly, 45(3), 375-423. https://doi.org/10.1177/0013161x08329290 steinert, b., hartig, j., & klieme, e. (2008). institutionelle bedingungen sprachlicher kompetenzen. in desi-konsortium (ed.), unterricht und kompetenzerwerb in deutsch und englisch: ergebnisse der desi-studie (pp. 411-450). beltz. stoll, l., bolam, r., mcmahon, a., wallace, m., & thomas, s. (2006). professional learning communities: a review of the literature. journal of educational change, 7, 221-258. https://doi.org/10.1007/s10833-006-0001-8 stoll, l., & louis, k. s. (2007). professional learning communities: divergence, depth and dilemmas. open university press. van gasse, r. (2019). the effect of formal team meetings on teachers' informal data use. frontline learning research, 7(2), 40-56. https://doi.org/10.14786/flr.v7i2.443 vangrieken, k., dochy, f., raes, e., & kyndt, e. (2013). team entitativity and teacher teams in schools: towards a typology. frontline learning research, 2, 86-98. https://doi.org/10.14786/flr.v1i2.23 vangrieken, k., dochy, f., raes, e., & kyndt, e. (2015). teacher collaboration: a systematic review. educational research review, 15, 17-40. https://doi.org/10.1016/j.edurev.2015.04.002 vangrieken, k., & kyndt, e. (2020). the teacher as an island? a mixed method study on the relationship between autonomy and collaboration. european journal of psychology of education, 35, 177-204. https://doi.org/10.1007/s10212-019-00420-0 maag merki et al. 23 | f l r vangrieken, k., meredith, c., packer, t., & kyndt, e. (2017). teacher communities as a context for professional development: a systematic review. teacher and teacher education, 61, 47-59. https://doi.org/10.1016/j.tate.2016.10.001 vannest, k. j., & parker, r. i. (2010). measuring time: the stability of special education teacher time use. journal of special education, 44(2), 94-106. https://doi.org/10.1177/0022466908329826 vescio, v., ross, d., & adams, a. (2008). a review of research on the impact of professional learning communities on teaching practice and student learning. teaching and teacher education, 24, 80-91. https://doi.org/https://doi.org/10.1016/j.tate.2007.01.004 vieluf, s., kaplan, d., klieme, e., & bayer, s. (2012). teaching practices and pedagogical innovation: evidence from talis. oecd. https://doi.org/10.1787/23129638 weick, k. e. (1976). educational organizations as loosely coupled systems. administrative science quarterly, 21, 1-19. https://doi.org/10.2307/2391875 wullschleger, a., maag merki, k., rechsteiner, b., & rickenbacher, a. (2019). kooperation von lehrpersonen im hinblick auf schulentwicklung: forschungsstand und forschungsperspektiven am beispiel von sozialen netzwerkanalysen. in u. steffens & p. posch (eds.), lehrerprofessionalität und schulqualität: grundlagen der qualität von schule 4 (pp. 259-286). waxmann. yang, r., zhang, z., & yu, s. (2018). understanding teacher collaboration processes from a complexity theory perspective: a case study of a chinese secondary school. teachers and teaching, 24(5), 520537. https://doi.org/10.1080/13540602.2018.1447458 ziegler, a. (2011). generalized estimating equations. springer. zirkel, s., garcia, j. a., & murphy, m. c. (2015). experience-sampling research methods and their potential for education research. educational researcher, 44(1), 7-16. https://doi.org/10.3102/0013189x14566879 maag merki et al. 24 | f l r appendix table x bivariate correlationsa among individual factors and of individual factors with the occurrence of daily administrative and organisational and of school subject-specific collaborative activities on the person level (n = 74/81b) (significant effects shown in bold) 2 3 4 5 6 daily collaborative activities with focus on administrative and organisational tasks only (n = 81) daily collaborative activities with focus on school subject-specific tasks only (n = 81) 1 gender (0 = f, 1 = m) (n = 74) -.023 (.848) .111 (.348) .093 (.430) .013 (.912) .051 (.666) -.076 (.519) -.187 (.110) 2 work experience (standardised) (n = 74) --.241 (.039) .002 (.998) -.087 (.462) -.078 (.508) -0.163 (.166) .000 (1.000) 3 teachers’ workload (0=<80%, 1=≥80%) (n = 74) -.011 (.926) -.043 (.714) -.103 (.383) .072 (.541) -.111 (.345) 4 leadership function (0 = no, 1 = yes) (n = 81) --.092 (.443) -.122 (.300) .155 (.166) .362 (.001) 5 internal interest in searching for knowledge (standardised) (n = 74) -.725 (.000) .021 (.859) .051 (.664) 6 external interest in searching for knowledge (standardised) (n = 74) -.023 (.848) -.023 (.845) a correlation coefficients are of type product-moment for two continuous variables, point-serial correlation for a continuous and a binary variable, and phi correlation for two binary variables. b the valid n for the individual factors, and therefore all correlations based on these, was 74 (with the exception of leadership function); the n = 7 cases with missing values did not differ significantly from those without missings regarding the mean administrative-organisational and school subject-specific collaborative activities (both p > .05). frontline learning research vol. 13 no. 1 (2025) 45 75 issn 2295-3159 corresponding author: joona moberg, department of teacher education, university of turku, finland, jvmobe@utu.fi, doi: https://doi.org/10.14786/flr.v13i1.1589 analyzing teachers’ scripts from teachers’ reflections after they tried to encourage students’ flexible mathematical thinking joona moberg1, minna hannula-sormunen1, markus hähkiöniemi2 & erno lehtinen3 1 department of teacher education, university of turku, finland 2 department of teacher education, university of jyväskylä, finland 3 department of teacher education, university of turku, finland, vytautus magnus university, lithuania article received 13 september 2024/ article revised 21 december 2024 / accepted 14 march 2025/ available online 1 april 2025 abstract teachers play a key role in promoting flexible mathematical thinking in society. there is a growing need to develop better methods for both preand in-service teacher training, but not enough is so far understood about what knowledge and skills teachers use in practical teaching situations. the means for investigating this are few. a new analytic framework was developed using abductive content analysis to investigate signs of script restructuring and construction as they appear in teachers' written reflections reporting their experiences in applying novel methods in their classrooms. scripts are mental knowledge structures combining formal professional knowledge and the knowledge teachers use in practical situations with representations, assessments, and predictions of different classroom events. scripts enable teachers to make (rapid) snap decisions to structure their teaching and manage classrooms to facilitate students' attention towards objectives, activities, and information that support learning. in this multiple case qualitative study, six teachers enrolled on the “flexible and adaptive arithmetic skills in primary school” course, part of the joma (towards flexible mathematics) in-service training program, teaching assignment and end-of-course reflections were investigated in depth. the goal was to advance the application of script theory to the study of teachers' actions and thinking as they engage in teaching intended to promote flexible mathematical thinking. the results suggest that signs of script restructuring and construction can be investigated post-hoc from textual accounts, scripts may have a considerable influence on teachers’ actions and thinking, and by engaging in teaching practice in real-life settings and reflecting on these accumulating experiences, processes leading to script development may be initiated. the results suggest that the analytic framework developed is functional and robust, paving the way for future investigations with larger samples. this study provided a more profound understanding of how online in-service education can support teachers to develop scripts supporting their competences to teach mathematical flexibility. keywords: teaching scripts, mathematics teaching, flexibility, teachers’ competences, inservice training. mailto:jvmobe@utu.fi https://doi.org/10.14786/flr.v13i1.1589 moberg et al 46 | f l r 1. introduction teachers have a key role in promoting mathematical thinking across societies (council of the european union, 2018; ministry of education and culture, finland, 2023; organization for economic co-operation and development, 2023). research has emphasized the need to promote flexible mathematical thinking (hickendorff et al., 2022; verschaffel, 2024), which means having a more profound understanding and awareness of mathematical concepts, relations, strategies, and representations. this enables the consideration of several alternative strategies or ideas and the adaptive use of those that correspond to the situation at hand (hickendorff et al., 2022; verschaffel, 2024). however, teachers’ competencies to promote this thinking may vary widely. teachers' understanding of the subject may be limited; they may teach math in a highly procedural and restricted fashion and hold beliefs that inhibit their ability to promote students' flexible mathematical thinking (blömeke et al., 2011; brunner & star, 2024; kaiser & blömeke, 2013; star et al., 2015). in finland, many aspiring preservice primary school teachers have weaknesses in basic arithmetic skills and lack a more profound conceptual and structural understanding of mathematical principles and procedures. many also think of math as a difficult, boring, and rule-based subject mastered largely through learning by rote and memorization (ohtonen et al., 2023; tossavainen & leppäaho, 2018). there is a growing need to develop better methods for both preand in-service teacher training to advance teachers' understanding and appreciation of mathematics and their ability to promote mathematical thinking. in recent decades, a lot of information and models have been proposed for the competencies that math teachers need in order to promote their students' mathematical understanding as well as their strategic, creative, and adaptive use (e.g., ball et al., 2005; döhrmann et al., 2011; shulman, 1986; stigler & miller, 2018). there is a lack of well-developed theoretical models of the knowledge which teachers use in complex classroom situations where events take place quickly and require fast (re)actions from teachers (blömeke et al., 2015; star et al., 2015). in this study, we explore how the script theory, which was first developed in the field of medicine to provide insight into how doctors established and used mental knowledge structures to assess patients’ illnesses using cues and event representations (charlin et al., 2000), could be applied to study teachers' knowledge structures. specifically, we investigate how teachers’ formal professional knowledge and the knowledge teachers use in practical situations were affected when they participated in intensive in-service training aiming to advance their understanding of flexible mathematical thinking and competencies to promote it. scripts are mental knowledge structures combining the formal professional and practical knowledge needed for high-quality teaching in a readily accessible form. scripts influence what the teachers notice as they teach, how the noticed events are presented in the teachers’ minds, and how the teachers interpret and respond to these events. scripts provide teachers with a framework to make rapid assessments and decisions to manage their classes and structure their teaching (wolff et al., 2021). there are currently no empirical studies that have explored teachers’ scripts as they reflect their practical teaching experiences. this is the first scientific inquiry into how the theory can be applied to study scripts of teachers who participated in in-service training aiming to promote their competence in teaching flexible mathematical thinking. the goal is to investigate signs of script restructuring and construction as they appear in teachers’ written teaching assignmentand end-of-course reflections. these reflections were provided during the course flexible and adaptive arithmetic skills in primary school (2022–2023). the course was a part of the joma (towards flexible mathematics) program, the first large-scale online in-service program implemented in finland intended to support early childhood education, pre-primary, primary, lowersecondary, and high school teachers to develop competencies to teach flexible mathematical thinking. this study furthers script theory from conceptual towards practical, paving the way for future investigations into how scripts affect teachers’ actions and thinking as they aim to teach flexible mathematical thinking and how in-service programs like joma could encourage teachers to engage in script restructuring and construction. moberg et al 47 | f l r 1.1 teachers’ formal professional knowledge base for advancing flexible mathematical thinking research has increasingly shown that differences in intelligence (e.g. general cognitive ability) or initial skill and knowledge levels play a smaller role in students' becoming flexible mathematical thinkers than do contextual, sociocultural, and demographic factors as well as students' and teachers' beliefs and mindsets (boaler & dweck, 2015; boaler et al., 2021; star et al., 2015; verschaffel, 2024). becoming a flexible mathematical thinker is a process that takes time and effort but nevertheless one that all students can engage in and improve at if provided with high-quality education by skilled and competent teachers (brunner & verschaffel, 2024; newton et al., 2010; rittle-johnson et al., 2012). flexible mathematical thinkers have a more profound conceptual understanding of mathematical principles and procedures (hickendorff et al., 2022), a stronger sense of mathematical agency, and more positive beliefs about their self-efficacy (bui et al., 2023). they focus on mathematical aspects not only in specific math tasks but also spontaneously in non-mathematical task contexts (mcmullen et al., 2020), have more adaptive knowledge about numbers and symbols (mcmullen et al., 2016), and are more inclined to approach math with curiosity and positivity (boaler & dweck, 2015; boaler et al., 2021). mathematically flexible thinkers also possess the necessary conceptual and procedural knowledge (e.g., about mathematical strategies) enabling them to apply mathematics regardless of the context in a way that suits them best (hong, et al., 2023). research suggests that competent teachers in advancing mathematical flexibility act proactively and strategically to prevent disruptive events and turning students' attention towards objectives, activities, and information that support flexible mathematical thinking (depaepe et al., 2020). they provide their students with information about mathematical building blocks (e.g. concepts, procedures, mathematical focusing aspects), challenge them to think about when and how to use these, and encourage students to engage in mathematical exploration and problem-solving in educational settings and real-life situations (hickendorff et al., 2022; hong et al., 2023; lehtinen et al., 2017; star et al., 2015). they model mathematics using a variety of means (e.g., drawing, diagrams, counting sticks) to illustrate mathematical concepts or procedures thoroughly while also clearly covering their meaning (boaler et al., 2016; boaler et al., 2021). competent teachers aim to cultivate a sense of community, where making mistakes is acceptable, students can share their thinking openly and engage in profound discussions about math and its applications (e.g., talk about the relationship between fractions and decimals, compare multiple strategies) (hong et al., 2023; rittle-johnson et al., 2020). hence, to promote flexible mathematical thinking, teachers need to have a general understanding of class management, be familiar with the content covered in math classes, and know how to create clear and well-formulated learning opportunities that enable students to engage with mathematical contents effectively. in this study, we use shulman's (1986) classic conceptualization of teachers’ knowledge to differentiate between these forms of knowing. according to shulman (1986), math teachers’ formal professional knowledge base consists of mathematical content knowledge (ck), mathematics pedagogical content knowledge (pck), and general pedagogical knowledge (gpk) supplemented by personal factors (e.g., traits, beliefs). mathematical content knowledge refers to knowledge about math facts, theories, principles, constructs etc., that are essential elements of mathematics needed to understand the subject and master its application in practice (blömeke, 2017). general pedagogical knowledge consists of the know-how about the general means, practices, and strategies teachers can use to manage their classrooms and organize their teaching (e.g. ways of promoting fruitful student-student relationships (döhrmann et al., 2012). pedagogical content knowledge is information integrating the mathematical content and pedagogical knowledge needed to teach mathematics effectively and accessibly (star, 2023). for instance, knowledge about how to apply teaching methods and assessment formats specifically designed to support students' understanding of the mathematical content as well as their engagement and agency, like number talks and math walks (english et al., 2010; parker & humphreys, 2018), qualifies as pedagogical content knowledge. moberg et al 48 | f l r 1.2 teachers’ scripts for teaching flexible mathematical thinking: merging formal professional and practical knowledge teaching, like clinical medicine and -psychology and the practice of law, is a highly social and dynamic profession where contextual factors, such as students' prior knowledge of the taught subjects, national standards and curriculum as well as the physical learning environments have a major impact on teachers' agency and ability to perform (stigler & miller, 2018). although a lot can be learned about these and how they affect teaching formally (e.g., through books, lectures, etc.,) such knowledge may remain disconnected from practice. formally acquired knowledge lacks information about many subtle cues of things and events that teachers encounter when they work in classrooms with their students (verloop et al., 2001; wolff et al., 2021). most crucially, it is detached from the teacher's subjective world of experience, such as his or her view of class events and actually using pedagogical means, tools, and approaches to support students' learning. these subjective experiences may invoke different affective responses, and they influence how teachers apply formally learned professional knowledge in practice (van dijk et al., 2022; verloop et al., 2001; wolff et al., 2021). teachers' formal professional knowledge can be seen as equivalent to doctors' formal biomedical knowledge (e.g., theoretical information about the human cardiovascular system), while knowledge teachers use in practical situations is the counterpart of doctors’ clinical knowledge (e.g., practice-based knowledge about symptoms associated with different diseases, potential treatments, and prognoses of their effects (boshuizen et al., 2012; charlin et al., 2000; jarodzka et al., 2013; schmidt & boshuizen, 1993). this practical knowledge is used by the teachers in concrete situations to structure teaching, manage the classroom, and support learning as different events occur (van dijk et al., 2022: verloop et al., 2001; wolff et al., 2015). it includes teachers' understanding of enabling conditions (background factors determining the likelihood of different events occurring) and teachers' mental representations of perceived events (wolff et al., 2021). it also includes processes and predictions affiliated with the events presented, for instance, assessments of what would happen if auxiliary questioning were used to tackle undesirable events (e.g., students having difficulties initiating small group discussions about the content of the lesson during revision). teachers accumulate practical knowledge through actual teaching experiences or by investigating the teaching of others e.g., by analyzing teaching cases (wolff et al., 2015; wolff et al., 2016; wolff et al., 2017, 2021). doctors do the same when they analyze clinical cases and practice medicine (boshuizen et al., 2012; charlin et al., 2000; jarodzka et al., 2013; schmidt & boshuizen, 1993), while lawyers accumulate practical knowledge by examining legal cases and practicing law (boshuizen et al., 2020). in clinical medicine, encapsulation is a process that leads to the intertwining of biomedical knowledge and practical clinical knowledge into illness scripts (jarodzka et al., 2013; schmidt & boshuizen, 1993). these scripts provide clinicians with a framework to draw upon to make quick assessments and act decisively, for instance, to effectively diagnose illnesses and treat their patients (boshuizen et al., 2012; charlin et al., 2000). it has been suggested that teachers' formal professional knowledge and the knowledge teachers use in practical situations may also combine by the means of encapsulation leading to the formation of classroom management and teaching scripts. in these mental knowledge structures, individual pieces of information are organized into scripts (wolff et al., 2021). for example, information about a specific environmental signal is linked to assessments of its meaning and consequences as well as potential means to tackle it providing the teacher with a fleshed-out model to quickly assess what is happening and respond effectively (wolff et al., 2021). script theory has been used mainly in the medical field to investigate doctors’ expertise and its development (e.g., boshuizen et al., 2012; charlin et al., 2000; jarodzka et al., 2013). however, a few attempts have been made to expand the theory to investigate expertise in other domains, like psychology, business, and law (boshuizen et al., 2020). recently, wolff et al. (2021) expanded the script theory to the realm of education and theorized how it could be used to investigate teachers' teaching. moberg et al 49 | f l r looking at the work achieved in these fields we present a detailed theoretical model of the processes teachers engage in to restructure and construct scripts. the model (see figure 1) has been constructed by carefully analyzing the previous work by wolff et al. (2021) and synthesizing the core ideas presented in the script literature and the research conducted particularly in the medical field (e.g., boshuizen et al., 2012; charlin et al., 2000; jarodzka et al., 2013; schmidt & boshuizen, 1993). our aim is to describe the various processes associated with script restructuring and construction and how these processes are related. this illustrates how scripts are formed and how teachers' script formation could be supported. our theoretical model consists of three distinct processes: formation of event representations (teacher's mental construction of the moment), perception making (focusing of attention to points of interest), and inferring of explanations (making mental actions of interpretation e.g., about why things are happening in class) (figure 1). these may be initiated by accumulating experiences which are affected by teachers’ formal professional knowledge base and may lead to script restructuring and construction. we will next elaborate on these processes and how reflection may support them. moberg et al 50 | f l r figure 1 the processes leading to script restructuring and construction through the encapsulation of formal professional knowledge and knowledge teachers use in practical situations . as teachers engage with students in class, different events manifest, for instance, students may perpetrate disruptive behavior. learning environments are complex, with many factors, some of which are highly context-dependent (e.g., physical space, students’ prior knowledge) and others less so (teachers' views and beliefs, societal educational goals, etc.) (stigler & miller, 2018). an almost endless number of events may occur as teachers teach, multiple events may happen simultaneously, and their boundaries may be blurred (when one event ends and another starts?) (wolff et al., 2021). also, not all teachers necessarily view the same events similarly as they may manifest differently in the teachers’ minds. moberg et al 51 | f l r when different events occur, teachers form event representations (figure 1). these are teachers’ mental constructions of the events occurring and are formed based on information teachers obtain about their surroundings by using their senses (wolff et al., 2021). a large body of sensory information is constantly acquired by the teacher. most of it is irrelevant for constructing a mental presentation of what is actually going on. to construct an event representation, the incoming information needs to be processed (wolff et al., 2021). perception works on two levels: internally and externally. externally, perception guides how teachers notice behavior and determines what events and associated cues they notice during teaching, influencing the nature and quality of incoming sensory information. internally, perception determines to what signals of the sensory information attention is paid, affecting the interpretations teachers can make about occurring events (wolff et al., 2021). we refer to these mental actions of interpretation as explanations which determine if the teacher knows what is happening in class and, more importantly, why things might be happening (e.g., why students did not engage in number talk and what role the teacher's actions may have played in this) (figure 1). teachers use their explanations to make assessments and predictions about the underlying reasons for events, their consequences, and the effects of responses that could be used to address different events (wolff et al., 2021). in the medical field, doctors engage in a similar process of reasoning and hypothesis formation when seeking to determine what disease a patient is suffering from, what its root causes are, how the disease can be treated, and what consequences different solutions may have (boshuizen et al., 2012; charlin et al., 2000; jarodzka et al., 2013). classroom management and teaching scripts contain teachers’ internal representations of past events (i.e., signals and (potential) sources of problems associated with the event), constructed by directing attention to aspects of the environment and (sensory) information obtained. they include assessments and predictions (process) about the reasons behind presented events (enabling conditions), their outcomes, and the results different responses may have if used to address the events (wolff et al., 2021). there is a bidirectional relationship between constructed event representations, perceptions made, and explanations inferred. when event representations are formed, they begin to influence teachers' perception, affecting what external signals and cues they observe and to what information obtained through the senses attention is paid. as teachers infer explanations using their event representations, their interpretations begin to guide perception. hence, teachers’ future event representations are influenced by previously inferred explanations (wolff et al., 2015; wolff et al., 2016; wolff et al., 2017, 2021). it is crucial to note that event representations, perception, and explanations are also affected by teachers' personal orientations (e.g., views, and attitudes) and formal professional and practical knowledge accumulated through experience (figure 1; wolff et al., 2021). for instance, when new information is absorbed into teachers’ formal professional knowledge base, it may begin to influence their behavior, thinking, and attention as they engage in teaching. this engagement leads to the accumulation of experiences yielding practical knowledge potentially impacted by the changes in the teacher’s formal professional knowledge base (figure 1). through experience teachers’ formal professional knowledge and knowledge used in practical situations is structured into scripts. the structured knowledge is also associated with information about the (potentially) occurring event’s enabling conditions, process, and consequences. hence, when an event matching an established script is encountered it is activated without much conscious effort leading to action (wolff et al., 2021). if an event that does not match an existing script is encountered during teaching, this may cause teachers to form new event representations using their perception and to infer explanations, for instance, to make assessments and predictions associated with that event. this can lead to the reform of an existing script or to the construction of a new script to meet the demands of the novel event encountered. such are the procedures of script restructuring and construction leading to merging (by encapsulation) of teachers' formal professional knowledge and knowledge teachers use in practical situations (figure 1; wolff et al., 2021). moberg et al 52 | f l r 1.3. investigating teachers’ scripts in the joma in-service training context an important contribution of script theory for analyzing the impact of training that aims to enhance teachers' expertise in supporting flexible mathematical thinking is that merely acquiring formal professional knowledge does not suffice to promote expertise (wolff et al., 2021). the training should encourage integrating (through encapsulation) formal professional and practical knowledge into teaching scripts. this can be achieved by providing teachers with formal professional knowledge about flexible mathematical thinking and its promotion as well as practical teaching experiences that enable teachers to explore the use of formal knowledge in practice (wolff et al., 2021). online in-service programs present a promising accessible, scalable, and cost-effective solution to provide teachers with this knowledge, to implement it in practice, and to reflect on their implementation experiences (clements & sarama, 2021; darling-hammond et al., 2017; mulcahy et al., 2021). the joma program (joustavaan matematiikkaan – towards flexible mathematics) consists of professional development courses that incorporate features providing a promising arena for investigating teachers’ scripts. the program has been operating since 2018 and consists of an introductory course designed for all participants and courses aimed at teachers working at different levels of the education system. it is finland's first online in-service training program intended to support teachers’ competences to enhance their ability to teach flexible mathematical thinking. joma program provides participants with up-to-date researchbased knowledge and practical methods for promoting students’ flexible mathematical thinking and instructions for how to implement these in classrooms, pre-school, or daycare. course materials include video recorded lectures, discussions and feedback by experts and other participants as well as various exercises, questionnaires, and readings. discussion forums allow open discussion with course instructors and other participants. the present study investigates the experiences of teachers who participated voluntarily in the program's flexible and adaptive arithmetic skills in primary school course. teachers’ scripts may have been impacted by the course in several ways: by providing teachers with formal professional and practical knowledge about flexible mathematical thinking and how to teach it, enabling teachers to practice and gain experience of teaching flexible mathematical thinking, and giving them the opportunity to reflect on these experiences. firstly, expert video interviews, scientific writings by academics, professorial summaries, and recorded classroom situation examples were used to present formal conceptual knowledge about flexible mathematical thinking and teaching. guidelines and examples showed teachers how to integrate the "theory" of teaching flexible mathematical thinking into practice. these may have prompted teachers to assimilate new formal professional knowledge leading them in the future to view, assess, and interpret events differently as they teach (jarodzka et al., 2013; wolff et al., 2021). secondly, the teachers were provided with opportunities to gain teaching practice and experience by conducting five expert-designed teaching assignments. in these, the teachers experimented with number talks, which are short, open, and explorative discussions during which students and teachers jointly investigate mathematical concepts, relations, phenomena, etc. (parker & humphreys, 2018). they can be used, for instance, to compare multiple strategies or build a class-community supportive of idea sharing and exploration (parker & humphreys, 2018; rittle-johnson et al., 2020). the teachers also designed and implemented math walks, trips to different environments, during which the students were encouraged to pay attention to math in their surroundings (shapes, patterns, numbers, relations, etc.) (english et al., 2010). these provide opportunities to practice everyday mathematical problem-solving and focusing on mathematical aspects (mcmullen et al., 2020). during the assignments, the teachers tested multiple math games with their students, like the number navigation game. this game enables students to practice applying arithmetic operations to solve calculations using fractions, percentages, and decimals promoting adaptive number knowledge and rational number sense (brezovszky et al., 2019; bui et al., 2022). the teachers also designed new or modified verbal math tasks presented in math textbooks. using research and practice-based guidelines, the teachers moberg et al 53 | f l r worked to create tasks with sufficiently challenging problems that enabled a step-by-step inquiry with time and thought as well as creative problem-solving (e.g., sequencing the tasks components by drawing) to promote flexible mathematical thinking (verschaffel et al., 2020; vicente et al., 2022). finally, the teachers explored concretizing mathematical operations and concepts using various means of illustration, like drawing and geometric shapes, but also creative means, for instance, rhythmic clapping and music. these provide multiple avenues to investigate mathematical operations, concepts, and their relations, making mathematics education more memorable, functional, and engaging (boaler et al., 2016; boaler, 2019). the expert-designed teaching assignments bridged the gap between theory and practice. they provided the teachers with practical experience of using means, tools, and practices incorporating flexible mathematical thinking teaching principles with their students. considering earlier script studies (e.g., boshuizen et al., 2012; boshuizen, gruber, & strasser, 2020; si, 2022) such experiences may have prompted teachers to construct event representations and infer interpretations (e.g., assessments and predictions associated with the teaching experiences) on their basis. this may have prompted teachers to engage in procedures of script restructuring and construction to form scripts to meet the demands of novel events they may have encountered while implementing teaching assignments. thirdly, expert-designed teaching assignments and end-of-course reflections encouraged the teachers to consider the formal professional and practical knowledge provided in relation to their pre-and course time experiences and personal orientations (e.g., values, beliefs). by writing a short report on the course platform after the execution of teaching assignments, the teachers had the opportunity to explore their experiences thoughtfully and with consideration. the teachers described, for instance, how they had implemented different means and practices, what they had observed in their classes, what their explanations for these observations were, and what they did in response to support students’ flexible mathematical thinking. in the end-of-course reflection task, instead of focusing on a particular teaching assignment, the teachers were prompted to delve reflectively into their course experience as a whole. the teachers considered, for instance, what they had done to teach flexible mathematical thinking, thought about their skills and abilities to teach it, and stated which practices and methods they found especially effective for its promotion during the course. according to studies conducted in medicine (e.g., boshuizen et al., 2012; jarodzka et al., 2013; si, 2022), psychology, business, and law (boshuizen et al., 2020) as well as education (wolff et al., 2015; wolff et al., 2016; wolff et al., 2017, 2021) such reflective assignments may have influenced teachers’ scripts. they may have prompted teachers to reflect on their teaching and the associated experiences in tandem with the formal professional knowledge the course provided and personal orientations. such reflections may have acted as a catalyst for script restructuring and construction, for instance, causing a teacher to pay more attention in the future to cues relevant for learning leading to sharper event representations enabling a teacher to make more profound interpretations of what was going on in class and why. to conclude, the teachers’ teaching assignment and end-of-course reflections during the flexible and adaptive arithmetic skills in primary school course are likely from the teachers’ perspective to contain novel information about their course-time experiences, thoughts, and actions. hence, they provide a data-set unique in its breadth and depth for investigating teachers’ scripts as they aim to teach flexible mathematical thinking. 2. research objectives and questions the goal of this study is to provide in-depth information on how signs of script restructuring and construction manifest in the joma course participant-teachers’ reflections. hence, the participating teachers’ reflections on teaching assignments designed to promote flexible mathematical thinking and end-of-course reflections where they consider their course time experiences as a whole are investigated. also, from these reflections the associations between signs of script restructuring and construction and formal professional knowledge base are investigated as well as between experiences yielded by the joma course and indications moberg et al 54 | f l r of teachers’ engaging in script formation. these aims further the understanding of how script theory can be used to explore the processes teachers engage in to construct knowledge structures leading to in-class actions affecting teachers’ ability to teach flexible mathematical thinking. additionally, the study shows how signs of script restructuring and construction can be investigated using an abductive qualitative approach laying the foundations for future inquiries. the main research questions are: 1) how do signs of script restructuring and construction manifest in primary school teachers’ reflections, where they consider the experiences and knowledge accumulated as well as the thoughts occurring and actions taken during the course designed to enhance their flexible mathematical thinking teaching competencies? 2) how are the formal professional and practical knowledge as well as experiences yielded by joma course associated with signs of teachers’ script restructuring and construction based on their reflections? 3. data and methods 3.1. participants in the present study, the qualitative data consists of teachers’ expert-designed teaching assignments and end-of-course reflections. six teachers (n = 6) were selected from among the 226 participants who enrolled on the flexible and adaptive arithmetic skills in primary school course 2022−2023 and gave their consent to the use of these for research purposes. the course was aimed at primary school teachers who in finland teach grades 1-6 and work with children typically aged 7-13. the selected teachers are referred to in this article by the aliases sanni (age range 20−29), julia (age range 40−46), anna (age range 20−29), daniel (age range 40−49), heikki (age range 50−59) and alexander (age range 30−39). these teachers were selected because their reflections were particularly rich and detailed. they convey a finetuned image of the teachers' view of the moment as they implemented teaching assignments (what was observed, how students were engaged, what pedagogical methods were used, etc.). they also contain statements, questions, and comments rich in the interpretation of in-class events (e.g. why things unfolded in class as they did), teaching methods (e.g. what effects they had on students’ thinking and learning), and reflective critical thinking (e.g., what the teacher would do differently in future to promote students’ flexible mathematical thinking). in addition, the teachers represent a nicely diverse spectrum of age and teaching experience. the decision to select these six teachers as objects of inquiry to answer the research questions was preceded by a stage of preliminary investigation, during which systematic sampling was used to select 10% of teaching assignment and end-of-course reflections provided by the 226 course enrollees. these were analyzed sentence by sentence to see if a coding framework could be developed to identify signs of teachers' scripts as they engaged in teaching intended to foster flexible mathematical thinking. this was necessary for identifying teachers whose reflections provided insights into teaching scripts embedded in the wealth and depth of the joma data. 3.2. qualitative analysis there are currently no methods for investigating scripts from written reflections. hence, we set out to see if an analytic framework could be developed for exploring signs of script construction from reflective textual accounts. the first author relied heavily on script theory as a starting point. abductive content analysis, a common theory-guided qualitative research method was used due to its well-proven functionality for inquiring in-depth about research subjects’ subjective experiences and internal thoughts (cohen et al., 2018). guided by the established theoretical framework (script theory and studies, shulman’s model, research about flexible mathematical thinking and its teaching), the researcher immersed himself in the data to identify signs (e.g., words, phrases, and statements) indicative of essential script components (event representations, moberg et al 55 | f l r perceptions, and explanations). since formal professional knowledge and accumulated experiences have important roles in script formation and alteration (e.g., boshuizen et al., 2012; jarodzka et al., 2013), references to these were also searched for during the analysis to investigate how they relate to teachers’ scripts in the reflections. signs of event representation construction were identified by sorting out reflections where the teachers presented their views of events or situations that had occurred as they taught. these reflections needed to contain information about cues and signals the teachers paid attention to during teaching and include descriptions of what the teachers did “in class” and what transpired as a result of their actions. also, as event representations are formed by teachers on the basis of information they obtain and attend to using perception, teachers accounts containing signs of event representations construction had to include references to the teachers' points of attention i.e., their perceptions. we concede that assessing the accuracy of the external perceptions the teachers constructed as they implemented teaching assignments from reflective accounts is not feasible. this would have required monitoring and documenting what really transpired in the classrooms as the teachers taught. still, any descriptions of detail about what the teachers thought had transpired as they implemented different teaching assignments, for instance, about students’ actions, thinking, and emotions, indicate that the teachers’ attention must have been focused on these as the events occurred. moreover, if the teachers were recalling such things they must have been deemed worthy of recounting, suggesting the teachers’ attention was also attuned to them during the reflection process. in the teachers’ reflections there are a lot of references to the content the teachers covered, the pedagogical methods they used, and signs of the interpretations the teachers inferred during the teaching assignments. based on the aforementioned definition of perception, the teachers must have oriented their attention to these “in the moment” of teaching and during the reflection process. however, since the formal professional knowledge base and explanations categories covered such factors, they were not coded separately as signs of making perceptions. instead, the focus was kept on students. we wanted to comprehend how the teachers’ perceived students, for instance, when they implemented particular teaching methods or made interpretations about their effectiveness (e.g., how the students engaged with each other, if their learning was impacted and how). we also had an interest in exploring how the teachers perceived learning environments as they taught (e.g., classrooms, outdoor spaces, public spaces like the library). it is well established that perceptions of the environment impact teachers’ acting and thinking, influencing their scripts in multiple ways (e.g., the quality and content of event representations and explanations inferred from these) (wolff et al., 2015; wolff et al., 2016; wolff et al., 2017, 2021). when developing the coding framework, it was discovered that the teachers' reflections contained hardly any signs of the teachers making perceptions associated with learning environments that afforded insight into how they might have influenced the teachers' actions and thinking. hence, this avenue was not further pursued. based on the theoretical frame, teachers’ statements, questions, and comments rich in inferred explanations (e.g., assessment, and predictions associated with various events) had to be accompanied by references to different events, methods, tools, approaches, or actors (students, parents, co-workers, etc.) to give them context and meaning (boshuizen et al., 2020; charlin et al., 2000; wolff et al., 2021). hence, signs of teachers constructing event representations, making perceptions, and referencing formal professional knowledge base and accumulated experiences were first identified and then assessed to determine if they included indications of the teachers' mental striving towards inferring explanations. the correctness of these explanations was not of interest, rather the depth and breadth. to be counted as signs of inferring explanations, the teachers had to describe what they thought about the reasons behind classroom events and situations, student behavior and thinking, or affects related to different tools, practices, approaches, etc., used to teach flexible mathematical thinking. moberg et al 56 | f l r guided by shulman’s (1986) conceptualization and the existing research on teachers' professional knowledge (blömeke, 2017; döhrmann et al., 2012; star, 2023), references to teachers’ formal professional knowledge base were identified and sorted according to whether they contained information about content, general pedagogical, or pedagogical content knowledge. many of the references identified were not so profound, but rather assertive in nature. for instance, a teacher might have reported having explored fractions in class through joint discussion using geometric shapes without providing specifics as to exactly how this was done or to what affect. also, the teachers’ accounts of exploring mathematical content by pedagogical means, approaches, tools, etc., might not include any indications that the teachers directed their attention to the environment to make perceptions or sought to infer explanatory accounts about the reasons, consequences, or solutions related to the situation. identifying references to teachers’ formal professional knowledge base was nevertheless valuable. it enabled us to establish the context for many of the teachers’ remarks. for example, when a teacher was describing using number talk to support students’ fundamental understanding about the concept of equality, we were able to code what mathematical content was covered and how. if such a description contained indications that the teachers’ engaged in script restructuring and construction processes (e.g., made perceptions, inferred explanations), these were coded separately. hence, the depth and meaningfulness of the teachers’ accounts of formal professional knowledge base was determined by the references to script restructuring and construction processes that may or may not have accompanied them. we were also interested in investigating how the teachers’ course experiences were associated with the signs of script restructuring and construction according to their reflections. therefore, the teachers’ reflections were scrutinized for phrases and paragraphs containing indications of changes like statements “i now know understanding math profoundly enables adaptivity” or “i have started to use number talks daily in my class”. references to experiences pre-dating course participation were not categorized as accumulated experiences resulting from the joma course. as indications of change related to course time experiences were discovered, associated references to script restructuring and construction processes and formal professional knowledge base were also coded and grouped. this enabled us to explore how the teachers’ accumulated experiences were related to manifestations of script restructuring and construction as the teachers reflected on teaching flexible mathematical thinking. during the analysis, it became apparent that, in addition to reflecting on experiences that occurred during the course, the teachers consistently referred in their accounts to the time and experiences preceding the course. for instance, they compared pre-, during-, and post-course teaching experiences and described impacts on their perceptual behavior by comparing what they had done previously and aimed to focus on after the course. also, some teachers talked about being unable to implement expert-designed teaching assignments. the reasons included that the teachers were currently on maternity leave or working part-time or as special needs educators, so they did not have classes of their own. in their reflections, these teachers often described teaching experiments they had conducted before the course, using methods included in the expert-designed teaching assignments, like number talks and walks. given these findings, a time dimension was included in the coding framework to try and distinguish between references referring to pre-, during-, and post-course time. the first author completed multiple cycles of reading and coding establishing numerous main, subcodes, and sub-subcodes. this process of structuring and differentiation led to the emergence of three main and seven subcategories deriving firmly from the established theoretical framework. the nature and content of the established categories was jointly discussed by the co-authors on multiple occasions. coloring was used to differentiate units of analysis indicative of information relating to different main and subcategories. the final version of the coding framework developed consisted of three main codes, seven subcodes, and three time codes, including one for coding references where the time dimension could not be inferred. instructions were written with multiple examples of when and how to use main and subcodes in tandem with moberg et al 57 | f l r the time codes. also, coded references were included as illustrative examples for using the codes. a simplified version of the coding framework developed is presented in table 1. for the full framework, see appendix 1. moberg et al 58 | f l r main category references to formal professional knowledge base (fpkb) subcategory mathematical content knowledge (mck) references to math facts, theories, principles, constructs etc. e.g., references to geometry, arithmetic, equality, number line, unit conversion, etc. general pedagogical knowledge (gpk) references to general means, practices, and strategies for classroom management and organization (e.g., ways of forming fruitful student-student relationships). e.g., accounts of engaging in teacher-teacher collaboration, promoting action-based learning, etc. mathematical pedagogical content knowledge (mpck) references to integration of mathematical content-, pedagogical-, and didactical knowledge (e.g., descriptions of using math walks and number talks as a part of teaching). e.g., references to teaching math through number talks, engaging in mathematical modelling with students (using geometric shaped, blocks, etc.), incorporation math play or exploration in teaching, etc. main category signs of restructuring and constructing classroom management and teaching scripts (cmts) subcategory event representations (er) descriptions of events occurring during the teaching assignments written with details of what cues and signals the teachers paid attention (e.g., how students engaged with the task) that also include descriptions of what the teachers did “in class” and what transpired as a result of their actions. no previous presentation (npp) accounts indicating that no previous event presentation had been formed (e.g., participant had never conducted a math walk). perception (p) references to the teachers’ points of attention as they implemented the teaching assignments (e.g., students’ interactions or engagement with the assignments). e.g., students' strategies for solving geometric puzzles are described, students' emotions evoked by the expert-designed tasks are highlighted. explanations (e) statements, questions, and comments rich in interpretation (e.g., what effects number talk had on students’ thinking, why students did not engage with the assignment, what the teacher would do differently in future and why). e.g., the teacher estimates students' prejudices affected their interest in math games, the effectiveness of math walks is evaluated by mirroring situational observations with student feedback. main category references to accumulated experiences (ae) subcategory accumulated experiences (ae) references indicating that course-time experiences affected the teachers’ thinking and teaching (e.g., adoption of new teaching methods, new perspectives gained). e.g., the teacher reports having adopted number talk into personal teaching arsenal, the teacher’s prejudices towards mathematics textbooks have decreased during the course. time pre-course (1) course-time (2) post-course (un) table 1 a simplified version of the developed coding framework. moberg et al 59 | f l r the following is an extract by alexander coded using the framework. colors refer to different sub-categories in this example. the students quickly got into the game, once we started playing together from the same points and had a joint discussion about how to play the game (2_cmts-p students’ engagement with learning tasks/content) (2_fpkbgpk game based learning) (2_cmts-ep-is) (2_fpkb-gpk collaboration, teacher with students) (2_fpkb-gpk joint reflection) this example shows how a single sentence could contain references to signs of multiple script restructuring and construction processes. in alexander's reflection, he describes his event representation (er) formed while engaging with students in a game-based learning session. this event took place during the course. hence, time code 2 precedes other codes used. in the analytical framework, game-based learning is a part of teachers' formal professional knowledge base (fpkb), specifically categorized as a general pedagogical knowledge (gpk) based way of promoting learning. alexander talks about playing the game with his students, an action also categorized as a part of teachers' general pedagogical arsenal (gpk collaboration, teacher with students). from alexander's description, it is possible to infer signs that he paid attention to how his students engaged in gaming and the actions supporting this. hence, his perception (p) was aimed at students' engagement with learning tasks/content. this signals that alexander engaged in a process of perception-making associated with restructuring and constructing classroom management and teaching scripts (cmts) during the implementation of the expert-designed teaching assignment. also, alexander used joint discussion, a general pedagogical approach, to support this engagement. hence, the extract also received codes 2_cmts-p students’ engagement with learning tasks/content and 2_fpkb-gpk joint reflection. the preliminary investigation on the selected 10% of the data yielded by the 226 course enrollees was conducted rigorously, reflectively, and self-critically to derive results representing signs of the teachers' scripts restructuring and construction and their relations to formal professional knowledge and accumulated experiences as truthfully as possible. the results were discussed collectively on multiple occasions among all co-authors. based on these discussions, the coding framework was developed, and the data were analyzed repeatedly until a level of consensus was reached. also, the findings were presented to the members of xxxx research group. the members provided additional comments and suggestions on the coding framework. the preliminary investigation culminated in writing the coding framework presented in appendix 1. next, six teachers, whose reflections were rich and detailed, were selected to show how the coding framework could be used to explore teachers' script restructuring and construction signs in relation to references of formal professional knowledge and accumulated experiences. the first author chose the teachers included in the preliminary investigation best suited for this. these choices were discussed together with all co-authors. he delved into the dataset to find and code the missing teaching assignment and end-of-course reflections for a few of the teachers. these were not included in the preliminary investigation because of the systematic sampling. complete profiles for the selected teachers were constructed which included all their teaching assignments and end-of-course reflections coded using the established framework. the following section is devoted to exploring these findings. 4. results we begin by describing and presenting in detail with reference examples how signs of script restructuring and construction (event representations construction, making perceptions, and inferring explanations) appeared in the selected teachers’ reflections. next, we report how references to formal professional knowledge and accumulated experiences manifested in the reflections. we conclude by exploring how references to formal professional knowledge and accumulated experiences were associated with signs of script restructuring and construction in the teachers’ reflections. 4.1 signs of script restructuring and construction moberg et al 60 | f l r in the six teachers’ reflections signs of constructing event representations were descriptive accounts containing information about the teachers’ experiences, observations, and interpretations as they implemented teaching assignments (what was observed, how students were engaged, what pedagogical methods were used etc.). they conveyed these teachers’ recounted views of specific events and situations they had experienced as they engaged in teaching to promote flexible mathematical thinking. these were accompanied by references signaling what perceptions the teachers made during these events and situations. this makes sense as event representations are mental presentations of what is going on and why, constructed on the basis of perception. the following account by sanni shows how signs of constructing event representations and making perceptions appeared together in the teachers’ reflections. in this assignment, teachers were instructed to explore using number talks to investigate equivalence by adding arithmetic sentences on both sides of the equal sign to maintain the equivalence. i tried the number talk method with third graders. i deliberately carried out the experiment very much following teija's instructions on the video. we haven't tried anything similar with the class before. of course, we have previously explored a few different counting options, but in such detail and duration, the matter has not been discussed. like in teija's video, i drew ___ = 25–5. the students first looked at the equation with their mouths open, and when i told them that this would be discussed together for the whole hour, they were really amazed. i introduced them to the seesaw and because of the s2 students [finnish as second language], a lot of time had to be spent to foster a conceptual understanding of it. at first, the students mainly came up with additions and subtractions [as solutions], and only after a while was i able to guide them to think, for example, of using multiplication and division. this was a very eye-opening experience i realized as a result of this experience that rushing doesn't benefit anyone, but on the other hand, for some students, concentrating on one task seemed to be really challenging and they were unable to focus on thinking about the solutions to the task. there are a few interesting points that warrant attention in the extract. first, sanni indicated she had not used the number talk method before with the class. this lack of experience is accompanied by a reference to exploring only “a few different counting options”. these accounts indicate that sanni has not previously formed an event representation, which would correspond to the situation described in the rest of the extract. the novelty of the situation sets a promising premise for sanni to engage in a new event representation construction by making perceptions and inferring explanations as things evolve. indeed, sanni’s depictions of students’ reactions throughout the account are detailed, showing that her perception was tuned towards the students as she was forming her event representation. there are also signs of her inferring explanations on the basis of her experience as she reports realizing that "rushing doesn't benefit anyone". sanni’s remark at the end that “this was a very eye-opening experience" even suggests that her previously held notions have been impacted by the experience, leaving the door open to the beginnings of a script development. yet sanni also shows signs of doubt. she concludes her recollection of the event representation by drawing attention to her perception that some students’ concentration was challenged by the task, and they seemed unable to focus on coming up with solutions. of all the processes associated with restructuring and constructing scripts, accounts where the teachers made explicit their inferred explanations gave the most profound information about the reasons, hypotheses, and assumptions the teachers’ related to different events they encountered. they provided a glimpse into what was going on in the informants’ minds at the level of interpretation (e.g., about their evaluations, reflections, and thoughts). to illustrate, we analyze the next quotation by alexander. the game [number navigation] immediately inspired students to play. there was a desire [among the students] for a little more personality and personalization options when it came to the selection and modification of characters. the students quickly got into the game, once we started playing together from the same stage and discussed how to play the game. at some point, it was difficult [for the students] to perceive the starting and target locations [in the game], so they could have been moberg et al 61 | f l r visually highlighted in a slightly different way. i also wondered if it would have been easier if the "calculator" [tool for giving answers in the game] had been on the right side instead of on the left. maybe we were left missing that there was a little bit of storytelling in the game too? yet the game’s idea was good, and the students liked it. i believe this is a good game for both math lessons and as a short waiting exercise. in the quotation, there are signs that alexander constructed an event representation on the basis of his perceptions and formed explanations about the gaming method he tested with students. alexander is recalling his mental representation of the gaming event and describes how the students responded and interacted with the game leading him to state that the students had some difficulties with it. this signals that his perception was attuned to the students’ engagement during the teaching assignment. next alexander suggest how to develop the game and even introduces the idea that it could be used not just in math classes but more generally as a “short waiting exercise”. alexander's assessments, development ideas, and suggestions are explanations of the game’s functionality and potential. he seems to have in part inferred these on the basis of his perceptions of how the students responded and engaged with the game. these signals about alexander’s explanations and perceptions indicate that this positive experience may have left him open to incorporating the game, which is designed to enhance students’ adaptive number knowledge and rational number sense, into his teaching as a new pedagogical approach. hence, with time and more practice, alexander could form scripts into which formal professional knowledge about the number navigation game has been integrated with knowledge of using it in practical situations. this teaching assignment at least started to provide him with practical experiences needed for this. as the two previous examples by sanni and alexander illustrated, often teachers seemed to infer explanations as a response to certain events, like the difficulties the students were perceived to have faced during a task. indeed, accounts of teaching events that were challenging included more signs of forming explanations than those that seemed to go smoothly according to the six teachers’ reflections where they recalled the mental event representations they had constructed. yet there were cases where the reflective assignments themselves seemed to act as triggers that prompted the teachers to offer explanations unrelated to any specific teaching situations or events. the following account, in which heikki considers why crafting tasks for students of all ages might be a potential way to support learning, demonstrates this. in my opinion, you can craft more interesting and challenging tasks for all ages. even with a very small toddler, you can try to make a tower out of different blocks/items and see how the tower is created and how high it is or even a really long line of small cars can be built that stretches from the front door to the back. what’s important is that the child grasps the idea and that it’s comprehended in a way that’s meaningful to the child. for a long time, my respect for educational resource providers was high, meaning that i had to learn to do things by myself and work together with the students aiming to promote their learning. vygotsky's theory of the zone of proximal development has been an important part of my approach to learning. unlike alexander, heikki's explanations for why and when task crafting works are not based on teaching experiences yielded by the course. yet heikki’s ideas are signaling that the reflective assignment led him to engage in mental effort and make explicit his explanations for why using toys, blocks, or various items to craft tasks could be an inspiring and engaging way to support student learning. the notions at the end, where heikki says this is something he has had to learn by himself and that vygotsky's theory has guided his approach, strengthens the argument that he has arrived at these explanations through past experiences and accumulated knowledge. moberg et al 62 | f l r 4.3. references to the formal professional knowledge base and accumulated experiences nine times out of ten signs of script restructuring and construction were accompanied by references to the formal professional knowledge base and its subcategories. references to content knowledge provided information on the subject matter the selected teachers covered with their students. often these also indicated what mathematical concepts they related to a particular content or a teaching method. the teachers used general pedagogical and pedagogical content knowledge to manage the classroom and teach the subject matter covered. for instance, when anna tried the number navigation game with third graders, she recalled being able to support students through verbalization and questioning: playing the game and using different strategies seemed difficult for [students] at first. although they were skilled counters, the game was not easy at first. as a teacher, i guided the gaming situation through verbalization and by asking auxiliary questions that supported the students in understanding how to use different strategies. when the students learned to play the game, all went smoothly, and the game seemed to motivate the students. as she constructed her event representation, anna proceeded to assess the situation considering students' mathematical skills and initial difficulties with the game on the basis of her perceptions. anna’s assessment of the situation seems to have led her to conclude that verbalization and auxiliary questioning were viable means of supporting “the students in understanding how to use different strategies” at that moment. as a result, anna used them to structure her teaching and manage the situation. this caused her to feel successful in supporting students’ learning and engagement based on her following observations: the students learned to play the game, and they had seemingly smooth and motivating experiences. to consider utilizing the aforementioned general pedagogical approaches, anna must have incorporated knowledge about them into her professional knowledge base. however, formal professional knowledge may remain static and unusable if not accompanied by knowledge of how and when to use it in practical situations. this knowledge is formed on the basis of experiences and assessments, predictions, and interpretations associated with these. through experience, the processes associated with script construction and restructuring (event representations construction, making perceptions, and inferring explanations) may be initiated potentially starting formal professional and practical knowledge to combine and be structured (through encapsulation) in teachers’ minds into scripts. hence, accumulating experiences is essential for the formation of scripts. to illustrate, here is another example from anna, where she reflects on her number talk implementation experience: we started by looking at different numbers and calculations and discussing what these brought to mind. the students verbalized that the numbers and calculations were "easy" and "difficult". next, we took a closer look at the "difficult" numbers and calculations and pondered together if something could be done to make them nicer and easier to calculate. at first, the students couldn't figure out ways to do this. through joint discussion and auxiliary questions, the breaking down of numbers and calculations slowly started, and the students internalized the idea. i intend to continue these number talk discussions, as the students already began to realize that calculations can be solved using different strategies, which turns them into a more easily calculable form. what another great idea to be used in daily work! in anna’s remarks there are hints that formal professional knowledge about the number talk method has started integrating with practical knowledge of using it because of her experience. anna’s recollection of her event representation signals the teaching assignment enabled her to gain hands-on practice with the number talk method introduced during the course, make assessments of the evolving situation using her perception, and elaborate on the impact of the method on students' thinking and actions. her explanations about the impacts on students’ learning at the end indicate she believes this experience improved students' comprehension of how multiple strategies can be used to solve calculations. this suggests she may have made a connection that number talks can be used to promote flexible mathematical thinking by enabling students moberg et al 63 | f l r to understand that with different strategies numbers and calculations can be turned “into a more easily calculable form” which may be simpler to solve. anna’s mental representation of what was happening and why as she used the method, which took shape on the basis of her observations and assessments of the situation, seems to reinforce this notion. anna’s beliefs about the potential of the method and intention to use it more present in the quotation were also evident in her end-of-course reflections: i feel that the course increased my understanding of flexible mathematical thinking and problem-solving and gave me many concrete means of teaching mathematics next year, i want to focus more on using number talks and foster my students' understanding that mathematics is so much more than book assignments and lessons at school. i will try to add a problem-solving lesson to the weekly program and start almost every lesson with an assignment that supports students' flexible mathematical thinking, like a short math discussion. i feel that i have become enlightened about the potential of mathematics, and i am really excited about it. thanks to both the educators and my course mates. it's been an amazing journey! :) here again anna’s reactions are quite strong. she uses exclamation marks, words, and expressions that intensely indicate her willingness to change her established thinking and acting patterns because of her course-time experiences. descriptions of accumulated experiences like anna’s, which included references to formal professional knowledge and signs of constructing event representations, making perceptions, and inferring explanations, contained the strongest indications that the six teachers’ scripts were affected by their experiences. yet, even when the formal professional knowledge base was referenced and signs of most script restructuring and construction processes were found, hints of knowledge encapsulation leading to script development did not always manifest in the teachers’ accounts. the next example by julia shows this well: i tried the number navigation game in my small class. the students were fifth graders, and some had mathematics as a differentiated [the content and objectives of the teaching have been modified to match the student's limited skills and competencies] subject. it quickly became apparent that the game was too challenging for most students. even understanding the concept of the game was difficult for [the students] because it was so multi-dimensional. also, the lack of immediate performance awards, like those that students are used to from previous games, reduced the students' interest in the game. julia describes how the game was too difficult for the students, some of whom had difficulties in mathematics. she constructed an event representation considering why this was the case explaining how multidimensionality made the game complicated and that lack of immediate rewards, like those the students were accustomed to, reduced their interest. yet julia’s perceptions about the students and explanations aiming to infer reasons for their disengagement do not result in clear indications that she altered her behavior or thinking, or that she is more inclined to do so in the future because of the experience. 4.4. formal professional knowledge, accumulated experiences, and script restructuring and construction: further explorations encouragingly, in the six teachers’ reflections, there are multiple accounts suggesting that the joma course teaching assignments and reflective tasks gave the teachers opportunities to experiment with new teaching methods and take time to contemplate these experiences. such experiences could lead to the beginnings of knowledge encapsulation, an essential part of script restructuring and construction, of which we already saw hopeful signs in anna’s case. to demonstrate this further, we look at a few more examples. the next one is from daniel, who explored fractions with second graders: we did a lot of practical things: used geometric shapes and coloring, investigated the meaning of sequences and their parts using blocks and by dividing students into groups, and above all, there was a lot of number talk and a focus on correct terms. for example, "the moberg et al 64 | f l r denominator names the fraction, it tells how many parts the whole is divided into". my intention was to bring fractions into students' everyday lives and their world of experience. afterwards, it seemed that this goal was reached by many of the students. that is, they understood what fractions really are about. it remains to be seen what they have really internalized and how well they actually learned the next time we discuss the topic. in his teaching assignment reflection, daniel describes using multiple means (e.g., geometric shapes, student grouping) to “bring fractions into students' everyday lives and their world of experience.”. he is guided by an overarching goal of enhancing students’ understanding of “what fractions really are about”. daniel’s remarks allude that his efforts are based on pedagogical content knowledge, since he seems to think that modeling fractions using multiple means connects them to students’ life-experiences, supporting them to develop a more profound understanding of fractions. daniel also emphasizes the importance of using the “correct terms” as they explored fractions. this suggests that daniel considers terminological correctness essential in mathematics and is keen on focusing on it. interestingly, though daniel believes his actions lead many of the students towards a more profound understanding of fractions, he also expresses some doubt. this signals that his experience did not entirely convince him that his actions lead to his goals. daniel does not exactly state why he feels this way. for instance, did he make observations using his perception leading to such doubts as he taught. still, his slight skepticism hints that this single experience, the event representation he formed, and the assessments and predictions daniel associated with the experience on the basis of his representation, need further testing and validation. this could mean that more teaching practice may be needed for daniel to encapsulate formal professional and practical knowledge about teaching fractions with the aforementioned various illustrative means into script(s) aimed at supporting students’ profound fractional understanding as well as using the associated mathematical terms correctly. daniel seems to be keen to do this in the future according to his end-of-course reflections: during the course, i noticed that over the years i have reduced the role of practical learning in my teaching for several reasons, though i have always considered it important. now, i am going to return, for example, to completely practical math lessons. one thing that's clearly changed in my thinking is my attitude to solving word problems. until now, i have insisted on a solution marked with numbers, for instance, for students to receive full marks in exams. in the future, i will be more willing to accept a pictorial or verbal explanation for a task if it reveals the way the student was thinking and how the solution was reached. at the beginning of the extract, daniel shows increased awareness of how his teaching has evolved over the years to be less practical. according to his words, this is something he has realized during the joma course. daniel also says he will “return, for example, to completely practical math lessons”. daniel’s reflections on his course-time experiences strongly suggest that his awareness has increased, and he is more committed to teaching math more practically in the future. in the latter part of the extract, daniel also states that his views have shifted to be “more willing to accept a pictorial or verbal explanation for a task if it reveals the way the student was thinking and how the solution was reached.”. considering his previous remarks about exploring fractions practically to promote his students’ better understanding of them, daniel’s end-of-course reflections suggest that he wants to pursue these aims in the future. moreover, daniel shows that he is inclined to do so using pedagogically solid ways proven to promote students’ flexible mathematical thinking: making mathematics education more practical, accepting different types of answers, and emphasizing thinking over correct solutions. it is also worth considering if daniel’s hesitation after implementing the teaching assignment could have something to do with his teaching having become less practical over the years. daniel may have become less familiar with methods combining content and pedagogical knowledge to teach fractions practically. could this partly be why daniel felt unsure if he was able to achieve the goals he had set for his teaching? if so, daniel’s reflections suggest that his course-time experiences have not deterred him further. they may actually have increased his confidence to pursue actions yielding more teaching practice. if daniel could follow these moberg et al 65 | f l r impulses which the joma course seems to have initiated, this could lead him to gain more experience enabling him to form event representations, make perceptions, and infer explanations. these processes could support daniel in forming scripts into which knowledge used in practical situations and formal professional knowledge has been encapsulated enabling him to teach math more practically in the future. here we look at one more example where alexander describes his creative approach to revising division with his students: i implemented a game-based session that included expressionist elements with my fifth graders centered on divisions (12 students were present). we had practiced divisions previously and i decided to use this session to review things we had been covering with the students. we established teams, and each team had to make a small play/illustrative presentation of division calculation. the other team had to guess what the team's division was and solve the calculation. for each correct answer, the teams always got a point. the game was played on a best-out-of-five principle. the teams could use role-playing clothes, music, etc., in their plays/representations. in other words, they could be creative. the tasks were, for example, 34.50 euros divided by 4, 76 meters divided by 10 meters, etc. so, the actors also had to be careful about what the calculation was actually about. this was fun for the students. the fifth graders already have quite a lot of trust in each other, so they dare to engage in such activities and express themselves. all the tasks were solved except for one. a lot of thinking was required to create the plays/representations, understand what the divisions were, and solve the calculations. but at least the math became grounded and concrete. in the teaching assignment reflection, alexander describes how he used less conventional means of teaching mathematics to support students’ creativity and to promote their engagement in revising division. it seems alexander also aimed to set up a real-life premise for the students’ plays and representations by combining both team’s calculations into specific topics (e.g., dividing money or lengths). by making divisions and their use more “grounded and concrete” alexander perhaps hoped to support the students’ understanding that division is something that is needed and encountered in everyday life to advance their flexible mathematical thinking. the experience appears to have been successful on many fronts. according to alexander's recollection of his event representation and signs included in the extract of his perceptions, the students were having fun, dared to plunge into expressing their creativity, and had to think a lot to create, understand, and solve divisions presented using the aforementioned creative means. when contrasted with alexander’s end-of-course reflections, his decisions and subsequent actions to emphasize creativity, selfexpression, and students’ enjoyment, were based on a continuum of experiences preceding his participation in the joma course: i am a classroom and special education teacher by training, so many things presented during the course were familiar. however, it is important to keep revising and updating my knowledge. thirty-two years as a teacher have given me experience and [this] is not my first math course. through malaty [george malaty], varga nemenyi, other functional courses, and by utilizing ekapeli [graphogame], my mathematics teaching has evolved, and my understanding of learning mathematics has improved. what i think would be important for a young teacher is precisely flexibility in teaching mathematics. one's thinking should not be "locked-in", nor should teaching be textbook-centered drill-and-practice aimed at execution. assessment should support learning, not just be affirmative. the math culture of the classroom must be created and made joyful and suitable for everyone. alexander’s notion that “one's thinking should not be "locked-in", nor should teaching be textbookcentered drill-and-practice aimed at execution” indicates his appreciation for flexibility in teaching mathematics. also, by highlighting that assessment needs to be less affirmative to support learning alexander is implying that he is attuned towards using more constructive forms of evaluation. in fact, looking back at his earlier reflection, there is an argument to be made that alexander aimed partly at this as he implemented the moberg et al 66 | f l r game-based session. having already practiced divisions with his students, he made a conscious decision to use the session “to review things we had been covering with the students”. is this one way for alexander to assess students’ comprehension of divisions that is practical, functional, and elaborative in nature? finally, the statement at the end about the math culture of the classroom and the need to make it joyful and suitable for all students, suggests that alexander believes in the power of positivity and joy when it comes to teaching mathematics. this is made quite evident and concrete by his creative approach to revising divisions. the signs of script restructuring and construction (e.g., alexander’s perceptions about his students, the causes and effects of the creative teaching methods he used) and references to formal professional knowledge (e.g., grounding mathematics to make it more concreate, promoting students agency through creativity means) contained in both alexander’s reflections are alluding that he may have formed scripts encapsulating formal professional and practical knowledge enabling him aim at “practicing what he preaches”: to create a positive and accepting classroom culture, to support student agency through joy and practical teaching, and to nurture thinking over fast performance. at least the joma course offered him additional practice needed to engage in script restructuring and construction processes needed to form such scripts. alexander himself recognizes the importance of this by stating that “it is important to keep revising and updating my knowledge" despite feeling that he already has a solid experience and knowledge base because of his long teaching career and systematic efforts to develop his competencies. 5. discussion and conclusions this study, which is the first scientific inquiry aimed at analyzing signs of script formation from teachers' reflective textual accounts, aimed to highlight the connections that exist between different processes of script restructuring and construction. also, references to formal professional knowledge base and accumulated experiences were investigated from selected teachers’ reflections. this enabled us to explore how references to formal professional knowledge and accumulated experiences were associated with signs of script restructuring and construction in teachers' reflections. the results of this study indicate that signs of teachers constructing event representations, making perceptions, and inferring explanations could be identified from the teachers' textual accounts using the abductively developed coding framework. as expected based on our theoretical framework (e.g., boshuizen et al., 2012; jarodzka et al., 2013; wolff et al., 2021), the six teachers, whose written accounts were studied, recalled event representations they had constructed as they reflected on the teaching assignments, which offered the teachers practical experiences of teaching with methods designed to foster flexible mathematical thinking. also, as teachers recalled their event representations signs about the perceptions they had made when implementing the course assignments appeared. this is in line with our predictions, since teachers construct their mental representations of events using their perception (wolff et al., 2021). signs of explanations the teachers inferred, for instance, about what was happening and why when they were teaching, appeared frequently in the teachers’ reflections. the teachers’ explanations offered insight into the formal professional and practical knowledge they had mustered, and into the nature of the experiences the course provided the teachers. taken together, these insights enabled us to construct a picture of how accumulating experiences and formal professional knowledge may be related to teachers engaging in script restructuring and construction processes in the joma context. based on our analysis, when the studied teachers were given the opportunity to explore by teaching the use of different pedagogical tools, approaches, and methods and had freedom and time to reflect on these teaching experiences, the teachers showed indications of knowledge structuring. their recollections of constructed event representations and signs regarding their perceptions might be accompanied by detailed accounts to inferred explanations. these could lead the teachers to discover new perspectives (e.g., about themselves or their students), assess the functionality of different tools, approaches, and methods, and make predictions about their usability in teaching. also, according to the teachers’ reflections, teachers’ course-time moberg et al 67 | f l r experiences did frequently foster enthusiasm and commitment in them to use these in the future to promote their students’ flexible mathematical thinking. these findings suggest that accumulating practical teaching experiences can support teachers to engage in script restructuring and construction processes and that reflecting on these experiences may strengthen their effects. hence, practical teaching experiences may be capable of initiating the encapsulation of formal professional (e.g., about the importance of mathematical modeling with various means) with practical knowledge (e.g., how to conduct number talks in an engaging and thoughtprovoking manner). with time, more practice, and reflection it is possible that these initial steps of encapsulation will continue to grow and lead to more comprehensive knowledge structuring in the form of teaching scripts. however, gaining more experience may not in itself always be enough to initiate script restructuring and construction potentially leading to knowledge encapsulation in scripts. many of the teachers’ reflections about course-time experiences remained “superficial” (did not include profound and detailed signs about the teachers constructing event representations, making perception, or inferring explanations). accumulating experiences must encourage teachers to form event representations, make perceptions, and, most importantly, provoke mental effort that can yield insightful and thought-provoking interpretations of these experiences (boshuizen et al., 2012; jarodzka et al., 2013; wolff et al., 2021). when the teachers’ reflections included hints of these processes, indications that the teachers’ experiences had encouraged script restructuring and construction, possibly leading to knowledge encapsulation, were observed. this highlights just how important it is that the accumulated experiences are of good quality. the findings of this first explorative study are encouraging. firstly, they indicate that it is possible to investigate signs of script restructuring and construction post-hoc from teachers’ textual reflections using a coding framework developed abductively. the results obtained with the framework aligned with some of the findings of earlier script studies (boshuizen et al., 2012; boshuizen et al., 2020; si, 2022; wolff et al., 2021) conducted in multiple fields with methods designed to elicit in-situational thinking and behavior reaffirm notions that scripts are not simple perception-response mental structures. they are amalgamations of teachers’ experiences, formal professional and practical knowledge as well as personal orientations, and script development may well require continued practice, experiences, and conscious deliberation (wolff et al., 2021). the teachers whose reflections we have studied only had initial opportunities to gain practical and reflective experiences related to flexible mathematical thinking and its teaching during the joma course. yet we were able to identify signals of script restructuring and construction and references to formal professional knowledge and accumulating experiences that implied they had formed event representations based on their perceptions and inferred explanations leading to assessments and predictions (e.g., about events, tools, approaches, etc., encountered). these signs and references as well as the positivity and commitment teachers expressed in their reflections about their experiences give us hope that the course as a whole may have catalyzed the teachers to work to gain more practical and reflective experience of using pedagogically solid and expertly designed tools, methods, approaches, etc., they were introduced to in order to promote their students’ flexible mathematical thinking. if this is the case, the teachers could be more prepared to embark on a path where they gain additional formal professional knowledge about flexible mathematical thinking and knowledge of how to promote it in practical situations that may be encapsulated into scripts. limitations and further studies the study aimed at testing the applicability of script theory in analyzing impacts of an in-service course on teachers’ knowledge structures with data which has certain limitations. from reflective accounts it’s not possible to be certain that the studied teachers’ recollections of their event representations and perceptions match what actually happened as they taught. the possibility must be acknowledged that certain things may not have been recalled accurately. it is known that humans have a tendency to cognitive biases when recalling past events and to make interpretations of these that could impair the content and accuracy of the teachers’ reflections (pinker, 2021). human memory has also been shown to be malleable by contextual and subjectrelated factors, which at their worst, can even lead to false memories (loftus, 2005). still, research also shows moberg et al 68 | f l r that recollections of past events can be far more accurate than might be expected and indicate what the subjects’ points of interest were as these occurred (bainbridge et al., 2019; diamond et al., 2020). most crucially, subjectivity is actually an essential part of scripts. mental representations formed on the basis of perceptions and the explanations associated with these representations are meant to present teachers’ subjective views of various classroom events and situations (wolff et al., 2021). teachers' scripts may not and need not be based on the most realistic depictions of different events. this does not diminish their great influence on teachers' actions and thinking impacting their ability to manage and structure teaching to support flexible mathematical thinking. despite the limitations related to the data used in this study, signs of script restructuring and construction could be identified with a comfortable degree of certainty using the novel analytical coding framework. yet the functionality of the framework needs to be further investigated to assess the validity and reliability of the method. in the future, studies conducted with a more quantitative orientation with larger sample sizes and multiple coders could be used. investigating interrater reliability between coders by counting kappa values for each established category is recommended. the causality between the signs of teachers' engaging in processes associated with restructuring and constructing scripts and accumulating experiences yielded by the joma course cannot be investigated with cross-sectional studies like this one. this would require longitudinal follow up and intervention studies, which would benefit from incorporating additional research methods capable of inferring teachers’ acting and thinking “in the moment” (e.g., wolff et al., 2015; wolff et al., 2016; wolff et al., 2017). in the future, the findings of this inquiry could be used to design such studies by providing information about what content, assignments, and methods are likely to prompt teachers to engage in script restructuring and construction processes (e.g., guidelines and examples showing how to integrate the "theory" into practice, practical teaching assignments, reflection tasks, etc.). also, the findings of this study provide researchers with information about which cues and signals are potentially indicative of signs of script restructuring and construction and beginnings of knowledge encapsulation into scripts. this could help future researchers to direct their attention to those aspects which are informative of teachers’ scripts. it was also beyond the scope of this inquiry to delve deeply into qualitative differences between the joma course enrollees’ scripts from the perspective of teaching flexible mathematical thinking. future studies should examine if teachers' scripts differ in their ability to promote this by contrasting more the exact content of teachers' scripts with knowledge about pedagogical practices shown to support students' flexible mathematical thinking. through this, categories or groups might be formed to differentiate scripts on the basis of their potential to support or hinder teachers as they aim to teach flexible mathematical thinking and to highlight differences between expert and novice teachers in this regard. the information provided by such inquiries would promote a more profound understanding of the mental knowledge structures that enhance or inhibit teachers' teaching. this knowledge could be used to further develop in-service programs, like joma, to support teachers in altering their scripts promoting their ability to teach flexible mathematical thinking. keypoints a novel analytical framework for investigating signs of script restructuring and construction was developed the application of script theory to the study of teachers' teaching competencies was advanced moberg et al 69 | f l r references ball, d. l., hill, h.c., & bass, h. (2005). who knows mathematics well enough to teach third grade, and how can we decide? american educator 29, (1): 14-17, 20-22, 43-46 blömeke, s. (2017). modelling teachers’ professional competence as a multi-dimensional construct. in guerriero, s. (ed.), pedagogical knowledge and the changing nature of the teaching profession, oecd publishing, pp. 119–135 https://doi.org/10.1787/9789264270695-7-en blömeke, s., gustafsson, j.-e., & shavelson, r. j. (2015). beyond dichotomies: competence viewed as a continuum. zeitschrift für psychologie, 223(1), 3–13. https://doi.org/10.1027/2151-2604/a000194 blömeke, s., suhl, u., & kaiser, g. (2011). teacher education effectiveness: quality and equity of future primary teachers’ mathematics and mathematics pedagogical content knowledge. journal of teacher education, 62(2), 154–171. https://doi.org/10.1177/0022487110386798 boaler, j. (2019). developing mathematical mindsets: the need to interact with numbers flexibly and conceptually. american educator, 42(4), 28-. boaler, j., dieckmann, j. a., lamar, t., leshin, m., selbach-allen, m., & pérez-núñez, g. (2021). the transformative impact of a mathematical mindset experience taught at scale. frontiers in education (lausanne), 6. https://doi.org/10.3389/feduc.2021.784393 boaler, j., & dweck, c. (2015). mathematical mindsets: unleashing students’ potential through creative math, inspiring messages and innovative teaching (first edition.). john wiley & sons, incorporated. boaler j, chen l, williams c, & cordero m. (2016). seeing as understanding: the importance of visual mathematics for our brain and learning. journal of applied & computational mathematics, 5(5), doi: 10.4172/2168-9679.1000325 boshuizen, h. p. a., gruber, h., & strasser, j. (2020). knowledge restructuring through case processing: the key to generalise expertise development theory across domains? educational research review, 29, 100310-. https://doi.org/10.1016/j.edurev.2020.100310 boshuizen, h. p. a., van de wiel, m. w. j., & schmidt, h. g. (2012). what and how advanced medical students learn from reasoning through multiple cases. instructional science, 40(5), 755–768. https://doi.org/10.1007/s11251-012-9211-z bainbridge, w. a., hall, e. h. & baker, c. i. (2019). drawings of real-world scenes during free recall reveal detailed object and spatial information in memory. nat commun 10, 5. https://doi.org/10.1038/s41467018-07830-6 brezovszky, b., mcmullen, j., veermans, k., hannula-sormunen, m. m., rodríguez-aflecht, g., pongsakdi, n., & lehtinen, e. (2019). effects of a mathematics game-based learning environment on primary school students’ adaptive number knowledge. computer & education, 128, 63-74. https://doi.org/10.1016/j.compedu.2018.09.011 brunner, e., & star, j. r. (2024). the quality of mathematics teaching from a mathematics educational perspective: what do we actually know and which questions are still open? zdm, 56(5), 775–787. https://doi.org/10.1007/s11858-024-01600-z bui, p., hannula-sormunen, m. m., brezovszky, b., lehtinen, e., & mcmullen, j. (2023). promoting adaptive number knowledge through deliberate practice in the number navigation game. in kiili, k., antti, k., de rosa, f., dindar, m., kickmeier-rust, m., & bellotti, f., games and learning alliance: 11th international conference, gala 2022, tampere, finland, november 30 december 2, 2022, proceedings (1st ed., vol. 13647), springer international publishing, pp. 127–136. bui, p., pongsakdi, n., mcmullen, j., lehtinen, e., & hannula-sormunen, m. m. (2023). a systematic review of mindset interventions in mathematics classrooms: what works and what does not? educational research review, 40, 100554-. https://doi.org/10.1016/j.edurev.2023.100554 clements, d. h., & sarama, j. (2021). sustainable, scalable professional development in early mathematics: strategies, evaluation, and tools. in li, y., howe, r. e., lewis, w. j., madden, j. j. (eds), developing mathematical proficiency for elementary instruction (pp. 221–238), springer international publishing charlin, b., tardif, j., & boshuizen, h. p. (2000). scripts and medical diagnostic knowledge: theory and applications for clinical reasoning instruction and research. academic medicine, 75(2), 182–190. https://doi.org/10.1097/00001888-200002000-00020 https://doi.org/10.1787/9789264270695-7-en https://doi.org/10.1027/2151-2604/a000194 https://doi.org/10.1177/0022487110386798 https://doi.org/10.3389/feduc.2021.784393 doi:%2010.4172/2168-9679.1000325 doi:%2010.4172/2168-9679.1000325 https://doi.org/10.1016/j.edurev.2020.100310 https://doi.org/10.1007/s11251-012-9211-z https://doi.org/10.1038/s41467-018-07830-6 https://doi.org/10.1038/s41467-018-07830-6 https://doi.org/10.1016/j.compedu.2018.09.011 https://doi.org/10.1007/s11858-024-01600-z https://doi.org/10.1016/j.edurev.2023.100554 https://doi.org/10.1097/00001888-200002000-00020 moberg et al 70 | f l r cohen, l., manion, l., & morrison, k. (2018). research methods in education (eighth edition., vol. 1). routledge. https://doi.org/10.4324/9781315456539 council of the european union. (2018). council recommendation on key competences for lifelong learning. general secretariat of the council. darling-hammond, l., hyler, m. e., & gardner, m. (2017). effective teacher professional development. learning policy institute. https://doi.org/10.54300/122.311. depaepe, f., verschaffel, l., & star, j. (2020). expertise in developing students’ expertise in mathematics: bridging teachers’ professional knowledge and instructional quality. zdm, 52(2), 179–192. https://doi.org/10.1007/s11858-020-01148-8 diamond, n. b., armson, m. j., & levine, b. (2020). the truth is out there: accuracy in recall of verifiable real-world events. psychological science, 31(12), 1544-1556. https://doi.org/10.1177/0956797620954812 van dijk, e. e., geertsema, j., van der schaaf, m. f., van tartwijk, j., & kluijtmans, m. (2023). connecting academics’ disciplinary knowledge to their professional development as university teachers: a conceptual analysis of teacher expertise and teacher knowledge. higher education, 86(4), 969–984. https://doi.org/10.1007/s10734-022-00953-2 döhrmann, m., kaiser, g., & blömeke, s. (2012). the conceptualisation of mathematics competencies in the international teacher education study teds-m. zdm, 44(3), 325–340. https://doi.org/10.1007/s11858012-0432-z english, l. d., humble, s., & barnes, v. e. (2010). trailblazers. teaching children mathematics, 16(7), 402–409. http://www.jstor.org/stable/41199504 ericsson, k. a. (2018). the differential influence of experience, practice and deliberate practice on the development of superior individual performance of experts. in ericsson, k. a., hoffman, r. r., kozbelt, a., hoffman, r. r., & williams, a. m. (eds.), the cambridge handbook of expertise and expert performance (2nd ed.), cambridge university press, pp. 745-769 gravemeijer, k., stephan, m., julie, c., lin, f-l., & ohtani, m. (2017). what mathematics education may prepare students for the society of the future? international journal of science and mathematics education, 15, 105–123. https://doi.org/10.1007/s10763-017-9814-6 hickendorff, m., mcmullen, j., & verschaffel, l. (2022). mathematical flexibility: theoretical, methodological, and educational considerations. journal of numerical cognition, 8(3), 326–334. https://doi.org/10.5964/jnc.10085 hong, w., star, j. r., liu, r.-d., jiang, r., & fu, x. (2023). a systematic review of mathematical flexibility: concepts, measurements, and related research. educational psychology review, 35(4), 104-. https://doi.org/10.1007/s10648-023-09825-2 jarodzka, h., boshuizen, h. p. a., kirschner, p. a., & lanzer, p. (2013). cognitive skills in medicine. in catheter-based cardiovascular interventions (pp. 69–86). springer berlin heidelberg. https://doi.org/10.1007/978-3-642-27676-7_7 kaiser, g., & blömeke, s. (2013). learning from the eastern and the western debate: the case of mathematics teacher education. zdm, 45(1), 7–19. https://doi.org/10.1007/s11858-013-0490-x lehtinen, e., brezovszky, b., rodríguez-aflecht, g., lehtinen, h., hannula-sormunen, m. m., mcmullen, j., pongsakdi, n., veermans, k., & jaakkola, t. (n.d.). number navigation game (nng): design principles and game description. in describing and studying domain-specific serious games (pp. 45– 61). springer international publishing. https://doi.org/10.1007/978-3-319-20276-1_4 lehtinen, e., hannula-sormunen, m., mcmullen, j., & gruber, h. (2017). cultivating mathematical skills: from drill-and-practice to deliberate practice. zdm, 49(4), 625–636. https://doi.org/10.1007/s11858 017-0856-6 loftus e. f. (2005). planting misinformation in the human mind: a 30-year investigation of the malleability of memory. learning & memory (cold spring harbor, n.y.), 12(4), 361–366. https://doi.org/10.1101/lm.94705 newton, k. j., star, j. r., & lynch, k. (2010). understanding the development of flexibility in struggling algebra students. mathematical thinking and learning, 12(4), 282–305. https://doi.org/10.1080/10986065.2010.482150 https://doi.org/10.4324/9781315456539 https://doi.org/10.54300/122.311 https://doi.org/10.1007/s11858-020-01148-8 https://doi.org/10.1177/0956797620954812 https://doi.org/10.1007/s10734-022-00953-2 https://doi.org/10.1007/s11858-012-0432-z https://doi.org/10.1007/s11858-012-0432-z http://www.jstor.org/stable/41199504 https://doi.org/10.1007/s10763-017-9814-6 https://doi.org/10.5964/jnc.10085 https://doi.org/10.1007/s10648-023-09825-2 https://doi.org/10.1007/978-3-642-27676-7_7 https://doi.org/10.1007/s11858-013-0490-x https://doi.org/10.1007/978-3-319-20276-1_4 https://doi.org/10.1007/s11858-%20017-0856-6 https://doi.org/10.1007/s11858-%20017-0856-6 https://doi.org/10.1101/lm.94705 https://doi.org/10.1080/10986065.2010.482150 moberg et al 71 | f l r meyer, a., kleinknecht, m., & richter, d. (2023). what makes online professional development effective? the effect of quality characteristics on teachers’ satisfaction and changes in their professional practices. computers and education, 200, 104805-. https://doi.org/10.1016/j.compedu.2023.104805 mcmullen, j., brezovszky, b., rodríguez-afecht, g., pongsakdi, n., hannula-sormunen, m. m., & lehtinen, e. (2016). adaptive number knowledge: exploring the foundations of adaptivity with wholenumber arithmetic. learning and individual differences, 47, 172–181. https://doi.org/10. 1016/j.lindif.2016.02.007 mcmullen, j., chan, j. y.-c., mazzocco, m. m. m., & hannula-sormunen, m. m. (2019). spontaneous mathematical focusing tendencies in mathematical development and education. in norton, a., & alibali, m. w. (eds.), constructing number (pp. 69–86). springer international publishing. mcmullen, j., verschaffel, l., & hannula-sormunen, m. m. (2020). spontaneous mathematical focusing tendencies in mathematical development. mathematical thinking and learning, 22(4), 249–257. https://doi.org/10.1080/10986065.2020.1818466 ministry of education and culture, finland. (2023). finnish national stem strategy and action plan: experts in natural sciences, technology and mathematics in support of society's welfare and growth. http://urn.fi/urn:isbn:978-952-263-733-8 mulcahy, c., day hess, c. a., clements, d. h., ernst, j. r., pan, s. e., mazzocco, m. m. m., & sarama, j. (2021). supporting young children’s development of executive function through early mathematics. policy insights from the behavioral and brain sciences, 8(2), 192–199. https://doi.org/10.1177/23727322211033005 organisation for economic co-operation and development. (2023). pisa 2022 results (volume i): the state of learning and equity in education. https://doi.org/10.1787/53f23881-en ohtonen piuva, a., mårtensson, e., röj-lindberg, a.-s., & braskén, m. (2023). finnish pre-service teachers’ basic mathematical skills: a comparison between 2008 and 2020. fmsera journal. https://journal.fi/fmsera/article/view/127787 parker, r., & humphreys, c. (2018). digging deeper: making number talks matter even more, grades 310 (1st ed.). routledge. https://doi.org/10.4324/9781032681016 pinker, s. (2021). rationality: what it is, why it seems scarce, why it matters. viking. rittle-johnson, b., star, j. r., & durkin, k. (2020). how can cognitive-science research help improve education? the case of comparing multiple strategies to improve mathematics learning and teaching. current directions in psychological science: a journal of the american psychological society, 29(6), 599–609. https://doi.org/10.1177/0963721420969365 rittle-johnson, b., star, j. r., & durkin, k. (2012). developing procedural flexibility: are novices prepared to learn from comparing procedures? british journal of educational psychology, 82(3), 436–455. https://doi.org/10.1111/j.2044-8279.2011.02037.x si, j. (2022). strategies for developing pre-clinical medical students’ clinical reasoning based on illness script formation: a systematic review. korean journal of medical education, 34(1), 49–61. https://doi.org/10.3946/kjme.2022.219 star, j. r. (2023). revisiting the origin of, and reflections on the future of, pedagogical content knowledge. asian journal for mathematics education, 2(2), 147-160. https://doi.org/10.1177/27527263231175885 star, j. r., newton, k., pollack, c., kokka, k., rittle-johnson, b., & durkin, k. (2015). student, teacher, and instructional characteristics related to students’ gains in flexibility. contemporary educational psychology, 41, 198–208. https://doi.org/10.1016/j.cedpsych.2015.03.001 stigler, j. w, & miller, k. f. (2018). expertise and expert performance in teaching. in ericsson, k. a., hoffman, r. r., kozbelt, a., hoffman, r. r., & williams, a. m. (eds.), the cambridge handbook of expertise and expert performance (2nd ed.), cambridge university press, pp. 431-452 shulman l. s. (1986). those who understand: knowledge growth in teaching. educational researcher, 15(2), 4–14. https://doi.org/10.3102/0013189x015002004 tossavainen, t., & leppäaho, h. (2018). matematiikan opettajien ja opettajaksi opiskelevien matemaattisesta osaamisesta [about the mathematical competence of mathematics teachers and student teachers]. in joutsenlahti, j., räsänen, p., & silfverberg, h. (eds.), matematiikan opetus ja oppiminen, niilo mäki instituutti, pp. 294-. https://doi.org/10.1016/j.compedu.2023.104805 https://doi.org/10.%201016/j.lindif.2016.02.007 https://doi.org/10.%201016/j.lindif.2016.02.007 https://doi.org/10.1080/10986065.2020.1818466 http://urn.fi/urn:isbn:978-952-263-733-8 https://doi.org/10.1177/23727322211033005 https://doi.org/10.1787/53f23881-en https://journal.fi/fmsera/article/view/127787 https://doi.org/10.4324/9781032681016 https://doi.org/10.1177/0963721420969365 https://doi.org/10.1111/j.2044-8279.2011.02037.x https://doi.org/10.3946/kjme.2022.219 https://doi.org/10.1177/27527263231175885 https://doi.org/10.1016/j.cedpsych.2015.03.001 https://doi.org/10.3102/0013189x015002004 moberg et al 72 | f l r verloop, n., van driel, j., & meijer, p. (2001). teacher knowledge and the knowledge base of teaching. international journal of educational research, 35(5), 441–461. https://doi.org/10.1016/s08830355(02)00003-4 verschaffel, l. (2024). strategy flexibility in mathematics. zdm, 56(1), 115–126. https://doi.org/10.1007/s11858-023-01491-6 verschaffel, l., schukajlow, s., star, j., & van dooren, w. (2020). word problems in mathematics education: a survey. zdm, 52(1), 1–16. https://doi.org/10.1007/s11858-020-01130-4 vicente, s., verschaffel, l., sánchez, r., & múñez, d. (2022). arithmetic word problem solving. analysis of singaporean and spanish textbooks. educational studies in mathematics, 111(3), 375–397. https://doi.org/10.1007/s10649-022-10169-x wolff, c. e., van den bogert, n., jarodzka, h., & boshuizen, h. p. a. (2015). keeping an eye on learning: differences between expert and novice teachers’ representation of classroom management events. journal of teacher education, 66(1), 68–85. https://doi.org/10.1177/0022487114549810 wolff, c. e., jarodzka, h., van den bogert, n., & boshuizen, h. p. a. (2016). teacher vision: expert and novice teachers’ perception of problematic classroom management scenes. instructional science, 44(3), 243–265. https://doi.org/10.1007/s11251-016-9367-z wolff, c. e., jarodzka, h., & boshuizen, h. p. a. (2017). see and tell: differences between expert and novice teachers’ interpretations of problematic classroom management events. teaching and teacher education, 66, 295–308. https://doi.org/10.1016/j.tate.2017.04.015 wolff, c. e., jarodzka, h., & boshuizen, h. p. a. (2021). classroom management scripts: a theoretical model contrasting expert and novice teachers’ knowledge and awareness of classroom events. educational psychology review, 33(1), 131–148. https://doi.org/10.1007/s10648-020-09542-0 https://doi.org/10.1016/s0883-0355(02)00003-4 https://doi.org/10.1016/s0883-0355(02)00003-4 https://doi.org/10.1007/s11858-023-01491-6 https://doi.org/10.1007/s11858-020-01130-4 https://doi.org/10.1007/s10649-022-10169-x https://doi.org/10.1177/0022487114549810 https://doi.org/10.1007/s11251-016-9367-z https://doi.org/10.1016/j.tate.2017.04.015 https://doi.org/10.1007/s10648-020-09542-0 corresponding author: joona moberg, department of teacher education, university of turku, finland, jvmobe@utu.fi, doi: https://doi.org/10.14786/flr.v13i1.1589 appendix mailto:jvmobe@utu.fi https://doi.org/10.14786/flr.v13i1.1589 moberg et al 74 | f l r moberg et al 75 | f l r guzman publication frontline learning research vol.13 no. 3 (2025) 29 52 issn 2295-3159 yupana inka tawa pukllay arithmetic eye tracking analysis: novices rosario guzman-jimenez1, dhavit prem1, alvaro saldívar1 & eduardo alejandro escotto-córdova 2 1 university of lima, peru 2 national autonomous university of mexico fes-zaragoza, mexico article received 12 august 2024 / article revised 17 february 2025 / accepted 15 may 2025/ available online 27 may 2025 abstract the concept of number emerges from the interaction of psychological, behavioral, and material elements of numerical cognition, collapsing the distinction between "abstract" and "concrete." this dual nature is evident in the inca numerical system, where tools like the yupana integrate abstract numerical concepts with concrete materials. the yupana inka tawa pukllay (yitp), a peruvian arithmetic method, enhances mathematical and visual-spatial skills through tile-based board games. while effective with children, its impact on university students is unexplored. this research used eye tracking to study gaze and attention during yitp operations, comparing novices and experts. eight university students and two experts participated, with eye-tracking data and scatter plot (dispersion plot) analyses collected using tobii pro glasses. the study introduced the variation ratio tokens (vrt) metric to assess visual attention efficiency, showing significant improvements in vrt dispersion and attention during the arithmetic learning process. these findings suggest yitp's potential in higher education for improving cognitive processes and arithmetic performance, laying a foundation for future research and innovative educational practices. this work establishes a foundation for cross-cultural cognitive studies and innovative stem education approaches leveraging ancestral knowledge systems. keywords: inca number system, inca abacus, arithmetic, pattern recognition, eye tracking info corresponding author: email: rguzman@ulima.edu.pe doi: doi: https://doi.org/10.14786/flr.v13i3.1569 introduction spatial thinking can be defined as a set of cognitive skills that enable us to organise, reason and mentally manipulate both real and imagined spaces. these skills encompass the capacity to reason about shape, size, orientation, direction, and trajectory of objects, the relationships among them, the mental visualisation of objects and their relationships, and reasoning about the spatial and temporal relationships of objects (gagnier et al., 2022; thayaseelan et al., 2024). the meta-analysis conducted by uttal et al. (2013) draws the conclusion that spatial skills are highly malleable, that spatial training is durable and transferable, and that it confers benefits to young children. the yupana represents a valuable educational resource, facilitating the integration of spatial reasoning with mathematical learning. its ability to facilitate the manipulation of abstract concepts through concrete materials renders it an efficacious instrument in the field of mathematics education, promoting a more interactive and visual approach to learning. the link between spatial thinking and mathematical learning can be more firmly established by drawing on the findings of cognitive psychology (mix & cheng, 2012). spatial reasoning can be defined as a way of action within the spatial world, encompassing the localization of objects, the relationships between them, and the perspective of their position from different viewpoints (uttal et al., 2013). the inca abacus, also known as the yupana, is a traditional device that facilitates arithmetic operations through a visual and spatial format. the ability for subitization and spatial thinking appear to be essential cognitive abilities for the effective utilisation of the abacus (cui et al., 2024; wang, 2020), such as the yupana (guzman-jimenez et al., 2023). these abilities, in conjunction with a profound comprehension of the decimal numerical system, enabled the incas to undertake intricate calculations with a seemingly straightforward instrument. it is erroneous to assume that numbers are purely mental constructs; they are inextricably linked to material and cultural factors. it is imperative to study these aspects to gain a comprehensive understanding of the nature and evolution of numerical concepts. subitization constitutes the foundation for the development of numerical sense, while the abacus serves as a visual tool that facilitates comprehension of mathematical concepts. mental calculation represents the practical application of memory, subitization and spatial visualization skills. the yupana inka tawa pukllay (yitp) method developed by author dhavit prem and yupanki association (prem d., 2016) is an innovative didactic arithmetic resource based on semiotic alternation (the use of alternative signs to communicate the same concept) that proposes the resolution of arithmetic operations through the recognition of visopraxic patterns (escotto-cordova & sanchez ruiz, 2018). in its serious game version, the yitp method was incorporated into a tablet and delivered to twelve bilingual primary school children (spanish-quechua) in a rural peruvian community. the children learned autonomously, without teachers, during the period of the global pandemic caused by the sars-cov-2 virus (guzman-jimenez et al., 2023). these findings indicate the potential of the yitp method as an educational instrument in the teaching and learning of arithmetic in a context of pandemic-related restrictions. given its visopraxic nature, the objective is to investigate the teaching-learning process of this method, utilising more precise visual process measurement tools, such as eye tracking, within a university learning context. the term 'visual attention' is used to describe the capacity of the human visual system to selectively process only those areas of visual scenes that are deemed relevant (borji et al., 2019). the capture of visual attention data is presented in the form of a heatmap, also referred to as an attention map. this represents an aggregation of gaze data over time, typically comprising the number of fixations or the duration of fixations made by a user. the resulting map is usually colour-coded, with green indicating areas of least interest and red indicating areas of greatest interest or hotspots. the colour-coded information is then superimposed on the original stimulus, and the data can be accumulated from all participants. to prevent the display from appearing sparse and to ensure a smooth map, a gaussian filter is applied to the fixation areas (duchowski et al., 2002). the conventional yitp methodology was primarily concerned with attaining assessment outcomes, thereby neglecting to incorporate insights into the cognitive processes underlying student solution-building. eye tracking enables the capture of this dynamic process in real time, as well as the observation of non-visible aspects such as approaches and strategies, which may not be reflected in the final response (van der weijden et al., 2018). furthermore, it provides insights into the internal processes occurring during the cognitive process, such as the reliance on visualisation and mental representations in mathematical abstract concepts. (hartmann et al., 2016). given that not all cognitive processes are objectively accessible, eye movements offer insights into cognitive activities that may be challenging or impossible to directly observe and communicate. (ott et al., 2018). this research aims to provide information on the gaze dispersion and attention zones involved in solving arithmetic operations obtained with eye tracking technology, with a view to gaining a deeper understanding of the learning and problem-solving processes employed by students (novices) in comparison to experts (authors of the yitp method). understanding the dynamics of yitp provides educators with a novel pedagogical instrument that facilitates students' problem-solving abilities by encouraging the development of their own strategies and algorithms. furthermore, the method enables educators to identify the difficulties students encounter when attempting to solve mathematical problems, thereby facilitating the development of targeted interventions designed to enhance their mathematical abilities. the present research has the potential to make a significant contribution to the field of mathematics education by investigating the cognitive processes involved in solving mathematical problems using the yitp method. the findings of this study may provide insight into these processes, thereby facilitating the development of novel pedagogical approaches that can empower students to excel in mathematics. the present study investigates the utilisation of the yitp method in arithmetic operations, employing the analysis of eye movement patterns through heat maps. the research question was as follows: does the application of the yitp method result in statistically significant differences in individual heatmaps when comparing the initial and final evaluations of the novice group? this question seeks to identify alterations in visual attention patterns resulting from the acquisition of the yitp method. it is of great importance to gain insight into the way undergraduate students learn to solve problems with the yitp method, particularly considering the prevailing educational context. to address this need, it is essential to explore innovative research methods, such as eye tracking through heat maps. the maps provide a comprehensive representation of the locations and modes of attention employed by students during the learning process, offering insights into their cognitive strategies and areas of difficulty. by employing this technology in the investigation of the inca abacus, patterns of attention can be identified and a deeper understanding of student interactions with this traditional mathematical instrument can be attained. the integration of eye tracking analysis with the study of the yitp method has the potential to enhance teaching effectiveness and optimise learning outcomes in university environments. 2. theoretical insight into yitp 2.1 representation of yitp numbers the sequence of dots on the first row at the base of the yupana board represents units, while the immediate superior row represents tens, then hundreds, thousands, and finally ten thousand (figure 2a). a token positioned on any given square will assume the value of the point(s) indicated on that square. on each row, the representation of the digit from 0 to 9 is achieved by placing the minimum quantity of tokens needed for that row. in case of digit 0, a row without tokens is used. the numerical values are inscribed from top to bottom, representing all digits in a continuous sequence (figure 2b). it should be noted that additional rows can be included on the top of a yupana according to the requirements of the quantity of digits, but for the purposes of this experimentation, the five-row yupana has been employed. the columns are represented by the set c = {5, 3, 2, 1}, as indicated by the numbers at the bottom of the board. similarly, the rows are represented by the set r = {1, 2, 3, 4, 5}, while the tokens are represented by the set t[r,c] = {0, 1}. each of the squares [r,c] are assigned one or no token, according to the aforementioned sets, where 1 represents one token and 0 represents no token. the representation of a number is in accordance with the formula (1). when the addends are placed on the yupana for an arithmetic addition (figure 2c), the value of all tokens represents the final numerical value. however, the resulting value cannot be read directly from the tokens on the board until yitp simplifications have been performed. it is therefore essential to identify the patterns and execute the predefined moves, which do not alter the initial value represented but rather simplify it, taking it to its most readable and optimal representation (appendix 1). a) columns & rows values b) representation of number 78309 c) setup of sum 3826 + 2974 figure 2: yitp numerical values representations the concept of an algorithm may be defined as a finite sequence of well-defined instructions that, when followed, produce a defined outcome (kowalski, 1979). consequently, each strategy developed by each participant may be seen as a set of well-defined instructions that they must follow to finally read the defined outcome (appendix 1). 2.2. yupana inka tawa pukllay yitp method as demonstrated by dhavit prem et al., (2022), yitp enables the performance of arithmetic operations through the recognition of patterns formed by tokens on squares and their predefined moves. yitp comprises three types of moves: a) basic moves, which reduce the number of pieces on a square (see appendix 1), b) expansion moves, which enlarge the pieces on the board when necessary, and c) advanced or compound moves, which are the simplification of two or more basic or expansion moves into a single move. this method eliminates the need for conventional mental arithmetic, while encouraging the development of different strategies comparable to those used in chess (prem., 2016). this approach diverges from traditional indo-arabic mathematics, which was developed by brahmagupta (590 ad) (bhattacharyya, 2011; ram & ramakalyani, 2022). the execution of yitp arithmetic operations is not constrained by a fixed sequence. the method allows for the implementation of multiple strategies of simplifying operations, either sequentially or in parallel, making it a parallel computational method. to illustrate, in the case of yupana, the sum of numbers must first be represented on the abacus (figure 2c), and then the operations (movements) proposed by the method can be performed as shown in appendix 1. the kamachiq challenge, which comprises several tokens distributed randomly on the yupana board, is often beneficial for novice players. this is because it provides them with a greater opportunity to identify a greater number of yitp patterns and perform compound plays, while simultaneously enhancing their own solution strategies until an optimal representation of a number is achieved (solution) (prem, 2018). in the case of an arithmetic sum (see figure 2), the total value of all the tokens represents the final numerical value (figure 2c). nevertheless, the value in question cannot be read directly from the tokens on the board until the requisite yitp simplifications have been performed. it is therefore crucial to identify the patterns and execute the predefined moves, which do not alter the initial value represented but simplify it, thereby bringing it to its readable and optimal representation. the academic interest in eye tracking technology in higher education lies in its ability to analyze students' visual and cognitive behavior, providing valuable data to improve teaching and learning. this tool allows for the optimization of educational material design by identifying areas of confusion or disinterest, and assessing skills such as critical online reasoning, differentiating between systematic and heuristic approaches (kunz et al., 2024). additionally, it monitors cognitive load in real time, helping to adapt teaching strategies to improve retention (šola et al., 2024). in environments such as project-based learning (pjbl), it reveals how students focus on critical areas, guiding the development of more effective resources (marlina & yunas, 2024). it also enhances the usability of educational platforms, such as learning management systems, by creating more intuitive interfaces (gu & paracha, 2023). together, eye tracking transforms education by offering data-driven insights to personalize and improve the learning experience. the yupana, used by the incas for arithmetic calculations, is a visual and material semiotic system where pieces on a board represent numerical values. their arrangement is linked to arithmetic operations, allowing for numerical transformations through spatial relationships. unlike abstract notation systems, the yupana externalizes cognitive processes, facilitating calculation in a concrete and manipulative way. additionally, it enables the visualization of multiple solutions, and individuals can create their own solution sequences based on established movement patterns. furthermore, the integrative approach of yitp incorporates historical, cultural, and linguistic aspects of quechua, fostering interdisciplinary studies and promoting innovative solutions in more advanced educational contexts. the properties of yitp were validated through a study conducted with rural children during the covid-19 pandemic, who learned the method through self-directed learning. the results revealed that (a) the children learned in a very short time, (b) they improved digit reading accuracy on the first attempt, (c) they increased their digit reading speed, and (d) they achieved a high percentage of correct readings of numbers containing at least one zero. these improvements in arithmetic accuracy, speed, and autonomy were facilitated by pattern recognition and a playful approach (guzman-jimenez et al., 2023). additionally, follow-up qualitative assessments, based on participant surveys, showed a significant improvement in their perception of mathematics after learning and applying the yitp method. importantly, eye-tracking provided insights into the cognitive processing involved. this study opens a new line of exploratory research on the yitp method in higher education. 3. method this study employs a mixed-methods approach (hayes, 1978; cooper, 1993; arsalidou & pascual-leone, 2016) to investigate visual and cognitive processes during learning with the yitp method within an educational program. this methodology enables a comprehensive analysis of how participants and experts process visual information and actively construct knowledge using the yitp method, facilitating an integrated understanding of learning dynamics. eye-tracking is the primary quantitative tool employed in this study, with metrics such as fixation duration, saccade frequency, and gaze patterns being captured before and after yitp implementation. these data facilitate the evaluation of changes in visual attention and processing efficiency, thereby offering insights into learning outcomes and the development of expertise. qualitative methods—namely, participant observations and semi-structured interviews—are employed to delve deeper into the knowledge construction process and the contextual factors influencing visual behavior. the present study addresses four key research gaps identified in prior literature (gegenfurtner, 2011; dogusoy-taylan & cagiltay, 2014; ooms, 2014): 1) the application of visual expertise findings to educational settings, 2) the examination of contextual moderators of visual behavior, 3) the optimization of visual tool design for varying expertise levels, and 4) the exploration of individual and cultural influences on the development of visual expertise. integrating eye-tracking data with qualitative insights allows for a robust evaluation of the yitp method's effectiveness in fostering visual and cognitive proficiency. this approach establishes a comprehensive foundation for analyzing expertise development across diverse learning environments, representing a significant contribution to the field. 3.1. participants the study included eight university students (four male and four female, aged 18–19) in their first semester at a private university, enrolled in humanities programs (table 1). all were taking an introductory research course, which motivated their voluntary participation. the participants met the following criteria: 1) normal or corrected-to-normal vision (including color perception). 2) no prior experience with the inca yupana or eye-tracking technologies. 3) no significant visual impairments or neurological conditions. the participants, predominantly from middle-class backgrounds, received no monetary compensation but gained academic benefits by contributing to the research. although their knowledge of the yupana’s potential benefits was limited, their curiosity about this cultural artifact led them to participate, seeking to enhance their self-efficacy in mathematics table 1 subjects distribution by sex and major note: all participants were 18 years old except p7, who was 19. p → study subject. all students attended an introductory session in the neuroscience laboratory, where they provided informed consent. the program consisted of three phases: week 1: an expert delivered a lesson on the yitp, covering number representation, basic movements, addition, and an introduction to subtraction. a second expert provided individualized reinforcement. each participant solved a four-digit addition problem, which was video-recorded. notably, numerical representation in the indo-arabic system on the yupana board required the use of 27 tokens distributed across the board (in both test 1 and test 2). week 2: reinforcement of previous content and introduction to subtraction. a whatsapp group was created to facilitate learning between sessions. week 3: participants performed arithmetic addition operations and the "kamachiq challenge," which involved randomly generating an operation with 50 tokens on the inca yupana board. for this study, only addition operations were considered. 3.2. materials and instruments during the experiment, the following software and tools were used: tobii pro 3 eye-tracking glasses and tobii pro lab software. the eye-tracking glasses were calibrated and connected via wi-fi to the laptop, enabling the collection of a substantial amount of raw data on high-frequency eye movements (sundstedt & garro, 2022). tobii pro lab facilitated the generation of heatmaps in jpeg formats, suitable for presentations or publications. these heatmaps were later processed using fiji software, which is based on the java imagej extension for digital image processing. additionally, the csv files generated by the tobii pro glasses 3 provided a comprehensive dataset, including gaze positions, pupil measurements, and event markers, which were crucial for analyzing eye-tracking data. this structured format allowed for further analysis and visualization of user behavior and interactions with the yitp board stimuli using statistical software. to conduct the experiment, video cameras were positioned in the four corners of the neuroscience laboratory (neurolab), alongside a video conferencing camera and a laptop. this configuration enabled the neuroscientist to monitor the experiments in real-time and record semi-structured interviews for later review by neuroscience experts. the sessions utilized a yupana board with magnetic tokens, as well as cardboard yupanas and corn seeds (used as tokens), as illustrated in figure 1. the yupana magnetic board was placed on an adjustable podium with a table stand to ease the arrangement of the magnetic tokens. the height of the podium was modified to ensure a perpendicular view, minimizing muscle strain for the student solving exercises in front of the board. in addition, a whiteboard was positioned nearby to facilitate the verification of arithmetic operations. figure 1: neuro lab: a) workshop,b) tobii glasses and c) tobii pro lab software. 3.3. metrics eye-tracking data often do not meet the assumptions of normality, particularly in small samples, with fixation metrics commonly exhibiting skewed distributions or outliers. the study, conducted with a small sample size of 10 participants, uses the wilcoxon test (allows for the comparison of two related samples when the data do not follow a normal distribution) to compare tests 1 and 2, and the friedman test (provides a nonparametric alternative to repeated measures anova, offering robustness to non-normality and making it well-suited for this type of data (conover, 1999; duchowski, 2002) to compare tests 1, 2, and 3. these non-parametric tests were integrated as complementary analyses, allowing for a deeper understanding of the data and helping to mitigate the limitations associated with the small sample size. although the results are preliminary due to the sample size, this multi-method approach provides a solid foundation for future research with larger samples. furthermore, specific metrics were included, such as seconds per move (s/move), which measures the relationship between the time spent and the number of token moves when solving the proposed exercises, and the token variation ratio (vrt), calculated by dividing the standard deviation (sd) of heatmap areas by the time in seconds (s) and multiplying by the number of tokens (t) at the start (vrt = t x sd/s). the vrt was obtained through heatmap processing and the fiji software (kerkhoff et al., 2022), and is divided into two types: vrt dispersion (green areas, reflecting dispersion efficiency) and vrt attention (red areas, reflecting attention efficiency). for the analysis, the wilcoxon test was used for tests 1 and 2, and the friedman test for test 3. the robustness of the friedman test to non-normality makes it suitable for this data, while the additional analyses provide a deeper understanding and mitigate the limitations associated with the small sample size. although the findings are interpreted with caution, this multi-method approach strengthens the validity of the results, laying the groundwork for future research with larger samples. 3.4. hypotheses to assess the impact of learning the yp arithmetic operations on visual attention, it is important to examine how novices' eye-tracking patterns might change before and after the learning process. eye-tracking data, particularly regarding dispersion and attention, can provide valuable insights into cognitive and perceptual shifts during the learning experience. based on this, the following hypothesis is proposed: h0: there are no significant differences in eye-tracking patterns (dispersion and attention) in novices before and after learning the yitp arithmetic operations. h1: there are significant differences in eye-tracking patterns (dispersion and attention) in novices before and after learning the yitp arithmetic operations an important aspect of learning is how novices' cognitive processes evolve as they gain experience. in the context of eye-tracking patterns, it is valuable to explore whether the visual attention and gaze dispersion of novices move closer to those observed in experts as they progress. this leads to the following hypothesis: h0: the ratios of eye-tracking patterns (attention and dispersion) in novices do not tend to converge towards those of experts. h1: the ratios of eye-tracking patterns (attention and dispersion) in novices tend to converge towards those of experts. 4. results 4.1 novices heat maps & experts references a quantitative comparison, based on the visual analysis of the heat maps, between novices and experts, and a before and after for each novice, reveals a pattern whereby the extent of the continuously coloured area in the heat maps is reduced and becomes more localized in specific isolated points, similar to small "islands" (see figure 3). in the case of the experts, the localizable points, or "islands of attention," are the dominant feature of the pattern, although there are points on the board where it extends and connects, forming "connected beach paths" (see appendix 2 for each participant). for clarity in the figures, green represents dispersion, red represents focused attention, and yellow represents borderline areas. this color coding helps distinguish between regions of broad visual exploration, focused attention, and transitional zones, respectively. figure 3: participant p2’s tests 1 & 2 and experts (27 tokens) solving addition: 5988 + 6417 + 4332 + 8675 to comprehensively evaluate the participants' knowledge gained during the experiment, another operation, known as "kamachiq challenge" (see figure 4), was tested. as detailed in the figure 5, this operation employed 50 tokens, in contrast to the 27 tokens used in earlier experiment stages. operation kamachiq requires the proficient application of learned moves for a comprehensive evaluation of mastery of the yitp method. the implementation of this operation facilitated the acquisition of insights into the way participants transferred their newly acquired skills to a more complex and demanding context. this allowed for a more comprehensive evaluation of their progress and competency in the yitp (see figure 6). figure 4: experts heat maps references (50 tokens) figure 5: participants heat maps (50 tokens) note: participant 7’s heatmaps were unavailable due to an eye-tracking device issue 4.2. variation ratio tokens (vrt) in table 2, the dispersion zone (green, p = 0.01) shows a statistically significant relationship, suggesting that participants distribute their gaze across the board, searching for a movement pattern they can perform rather than focusing on a specific area due to the nature of the game. the attention zone (red, p = 0.08), while not statistically significant at the conventional 95% confidence level, is significant at a 90% confidence level. given the small sample size, this result suggests a potential trend toward statistical significance, indicating that participants may exhibit a high concentration of visual attention in this zone. however, these marginal results should be interpreted with caution and warrant further investigation. one interpretation is that there is no single solution sequence, with each participant developing their own movement sequence or creating their own moves. this aligns with the authors' description of the yitp method and their empirical observations over 10 years. additionally, the relationship between the dispersion zone (p = 0.01) and the attention zone (p = 0.08) should be noted, as both reflect complementary aspects of the visual search process on the board. it is recommended to conduct further studies with a larger sample size or apply alternative metrics to explore the attention zone more robustly. future studies should also consider the visual context of the task (e.g., board design, distractors) to draw stronger conclusions about these patterns of ocular behavior. to investigate changes in performance across learning stages, a friedman rank sum test was conducted. the study included three stages: stage a (test 1, conducted during the workshop with 27 tokens), stage b (test 2, administered two weeks later with 27 tokens), and stage c (test 3, conducted in kamachiq with increased complexity and 50 tokens). the results revealed a statistically significant difference in performance across the three stages (p-value = 0.004) at a significance level of 0.05, with a chi-squared value of 11.143 indicating a moderate to strong effect size. however, the attention vrt analysis showed a p-value of 0.04, which is statistically significant at the 0.05 level but marginal if using a 0.10 threshold. data from participant p7 were excluded from the analysis due to missing records. the degrees of freedom (df) for the friedman test were 2, corresponding to the three stages. these findings are summarized in figure 6. additionally, participants with the exception of p5 and p6, exhibited an increase in their vrt from test 1 to test 2, as well as from test 2 to test 3 (when playing the kamachiq with 50 tokens). this indicates that an augmented number of tokens enhances the probability of identifying patterns that can be simplified through the utilisation of composite moves, thereby facilitating the generation of a greater number of solution strategies. consequently, there is a greater search for patterns in a dispersed manner (i.e., a greater degree of dispersion) and less time spent on this process. in the case of p5, the vrt remains equal to that of the additional evaluation. in contrast, in the case of p6, the vrt decreases, although it remains higher than the vrt shown during the learning process. the vrt values demonstrate an increase across the tests, indicating a greater variability in attention patterns as the tasks become more complex (basic to advanced) and require a greater number of tokens (27 to 50). table 2: participants (p) vrt dispersion, vrt attention for three tests and experts references note: p→ participant. t→ tests in table 3 both women and men showed an increase in vrt for dispersion and attention across all three tests. this suggests that both groups devoted more attention to exploring different areas of the board (dispersion), while maintaining a certain level of concentration on a particular area (attention) as the trials progressed. while both groups showed similar trends, the vrt results suggest that both groups' attention became more dispersed and more focused at the same time as the tests progressed. although there appears to be a similar pattern, with men scoring slightly higher than women, further analysis is needed to confirm whether there are significant differences between them. table 3: sex differences in vrt dispersion and attention inter-participant across tests (1, 2 & 3) the data for test 1 to test 2 and test 2 to test 3 were examined separately for female and male participants (see table 3). in contrast, for male participants, the vrt dispersion increased by approximately 48.79% from test 1 to test 2, and by approximately 22.47% from test 2 to test 3. the vrt attention experienced a significant increase of approximately 79.30% from test 1 to test 2. this was followed by a smaller increase of approximately 15.64% from test 2 to test 3 (see figure 6). figure 6: a) vrt dispersion and b) vrt attention intra participant analysis (test 1 & test 2) 5. discussion novices before and after yitp the findings of the present study indicate that learning the yitp method results in substantial alterations to both the visual and dispersion patterns of attention observed in the three tests conducted with novice participants. as anticipated based on the heat maps (blascheck et al., 2017), the proportion or variance ratio of dispersion zones (green zones) demonstrated an increase for all novice participants. consequently, this suggests a tendency for the dispersion patterns of novices to resemble those of experts over time. this also indicates that novice participants were able to expand the range of search patterns and identify the relevant information in a more efficient manner. the number of fixation zones (red zones) increased for participants 5, 6, 7 and 8, while a decrease was observed for participants 2, 3 and 4. to gain further insight into the cognitive processes and perceptions associated with learning arithmetic with yitp, open-ended interviews were conducted with each participant (see appendix 3). these findings lend support to the hypothesis that training in the yitp method fosters more focused and efficient attention in tasks pertaining to the operation of addition under the inca numeric system. the results indicate that, following the learning, after test 1 and test 2, participants are capable of inhibiting responses to distracting stimuli and directing their attentional resources in a more selective manner towards relevant information. additionally, the participants demonstrate enhanced visuospatial and subitizing skills, which are essential for the execution of movements on the yitp method. this increase in the vrt dispersion also suggests that novices progressively expand their search zones for new strategies while adopting a more rhizomatic approach (without a predetermined order) rather than adhering to a rigid, linear procedure, as is characteristic of indo-arabic mathematics. furthermore, after test 3, it is observed that the dispersion zones of novices tend to resemble those of experts, becoming increasingly scattered. the processing of numerical information is contingent upon visual attention, as evidenced by studies that have compared the performance of abacus experts and non-experts (lo & andrews, 2022). abacus experts employ visuospatial strategies for digit recall and mental calculations, in contrast to the linguistic strategies utilised by non-experts. the results of the research indicate that mental abacus calculation training improves visual image processing skills, which may influence how experts focus their attention on numerical tasks. furthermore, practice with the abacus results in the automation of the decoding of the information represented on the abacus, thereby making this process more efficient (srinivasan, 2018). this abacus training may elucidate the mechanisms by which mental abacus users are able to group abacus beads into columns and represent the abacus within the boundaries of visual working memory (frank & barner, 2012, cited by srinivasan, 2018). although the abacus and inca yupana board are not identical, the design of this one makes use of the principles of the former, facilitating the manipulation of tokens and the spatial location and subitization of the ones. furthermore, it requires memorising predefined patterns of movements in order to apply the most appropriate movements according to the strategy of each participant in search of the answer to the operation posed in the yupana cardboard. it should be noted, however, that the present study focused on a group of eight novice participants and the yitp method for only addition arithmetic operation. about significant differences in vrt ratio across tests before and after learning the yitp method in the novice group. attention analysis revealed a marginally significant difference between vrt values before and after learning (w=15, p=0.08298), suggesting a potential change in measurement location after learning. furthermore, it was observed that as the vrt value increased, there was a corresponding increase in the dispersion of the subject's attention on the yupana board. the mean vtr dispersion per test was 17.52 (test 1), 25 (test 2) and 30.55 (test 3), indicating a wider exploration of the yupana board. while the mean vtr attention scores for each test were 7.14 (test 1), 9.68 (test 2) and 11.46 (test 3), respectively, and did not exhibit a concentration on specific areas of the board, an increase in the value was also observed, indicating a tendency for the participants to focus their attention on specific areas of the screen as they progressed. this greater focus, accompanied by an increase in the precision of eye movements, suggests an adaptation and learning process. it is noteworthy that the experts exhibited notably elevated values for both vrt dispersion. the results indicate that they were in exploration mode, actively seeking information on the board. expert 2 displayed an even more intense exploration mode. similarly, both exhibited a high degree of concentration in specific areas of the board, indicating an efficient and directed visual strategy. the number system and mathematical practices of inca culture, as well as the specific symbolic representations and notations used by them, may have engaged different brain regions or required the recruitment of additional neural resources during mathematical operations compared to modern symbolic systems. this has also been observed in maya culture, particularly in its number system and mathematical practices, as well as in the specific symbolic representations and notations used by the maya (richeson, 1933; nickerson, 1988) subitization of small quantities (1-5 points) on the yupana board, for example, is likely associated with the approximate number system (ans) (dehaene, 1997, 2011; ansari, 2008), which represents an ancient evolutionary system for approximating numerical magnitudes. heatmaps may indicate brief fixations on clusters of points, reflecting a rapid approximation of quantities by the ans, likely following the initial subitizing. the representations in question can currently be used as semiotic alternations in the yitp method (guzmán-jiménez et al., 2023). it should be noted that tests 1 and 2 involved 27 tokens, with the addition arithmetic of four 4-digit numbers, and that this was the same in both tests. test 3, however, used 50 tokens from the kamachiq challenge. the results indicate that as novice participants become more familiar with the yitp method, their attention becomes more focused on relevant areas and there is a greater exploration of the entire board in search of potential pattern moves. following the culmination of the tests, based on the open-ended interviews, it was evident that participants (p) exhibited the activation of a multitude of cognitive processes throughout their interaction with the yitp. in accordance with the theoretical frameworks put forth by piaget (1970). bruner (1974), ausubel et al., (1978) and dehaene et al., (2004, 2005), the participants demonstrated a focused attention on the pertinent elements of the board, which underscores the concentration required to effectively manipulate the tokens. the perception of numerical and symbolic information was of great consequence, with the yitp being perceived as a game with unambiguous rules (p5, p6). memory was identified as a fundamental component in the recall of rules, actions, and mathematical knowledge (p1, p2, p3, p6, p8). furthermore, the formation of mental models pertaining to mathematical operations was discerned (p6), which is pivotal for comprehension. problem-solving skills were evidenced by the utilisation of existing knowledge to identify solutions (p8). furthermore, the participants demonstrated metacognition by reflecting on their learning process, identifying their strengths and weaknesses, and evaluating their personal strategies (p3, p4, p5, p6, p8). this suggests that yitp encourages active learning that engages multiple cognitive processes, from initial perception to metacognitive reflection. furthermore, participants reported positive experiences associated with the utilisation of the yitp, including an increase in enjoyment, self-confidence, self-efficacy, and curiosity. a number of participants drew attention to the playful aspect of the yitp, likening it to a game (p1, p3, p4). furthermore, an enhancement in comfort and the capacity to err without trepidation was observed (p3). some participants indicated an increase in mathematical self-esteem due to the utilisation of the yitp, comparing it favourably with traditional methodologies (p2). ultimately, the potential of the yitp to cultivate interest in incan culture was emphasised (p8). participants expressed positive experiences related to the practice of the yitp method, reporting an increase in enjoyment, self-confidence, self-efficacy and curiosity. several participants highlighted the playful nature of the yitp method, comparing it to a game (p1, p3, p4). in addition, an increase in comfort and the skill to make mistakes without fear was mentioned (p3). some participants reported an increase in mathematical self-esteem due to the practice of the yitp, comparing it favourably with traditional methods (p2). finally, the potential of the yitp method to generate interest in incan culture was highlighted (p8). chang et al., (2022) highlight the influence of cultural factors on mathematical performance.yitp, by connecting with cultural identity, can increase student motivation and performance. moreover, the findings of the present study revealed discrepancies in eye movement patterns between men and women during yitp learning. however, both the dispersion and attention vtr exhibited low values and minimal increases across the three tests. however, previous research (yuan et al., 2019; yang et al., 2022) has indicated that there are gender differences in spatial abilities and that women may have superior verbal and information processing skills. in agreement with the findings of kaczkurkin et al. (2019), we advocate for further investigation into these cognitive discrepancies. novices and experts the comparison between the inca yupana, chess, and the abacus is particularly insightful, as it highlights shared elements related to spatial representation, as well as visual and motor processing, especially in terms of differences in vtr between experts and novices in yitp. however, a more in-depth analysis is necessary considering previous studies on eye movement patterns and spatial representations. both chess and the yupana board utilize a spatial arrangement to represent information. studies on chess (sheridan & reingold, 2014; ribeiro da silva junior et al., 2018) have shown that experts develop highly efficient visual search patterns to identify relevant information. it is likely that prolonged practice with the yitp method also leads to the development of specific visual search patterns, although further research is needed to confirm this. additionally, subitizing and visual processing play a crucial role in both the abacus and the inca yupana board. however, the yupana board combines elements of subitising with a more complex spatial structure, similar to chess. this suggests that yitp boards can train both subitising and spatial processing skills. furthermore, previous studies (duchowski et al., 2002; sheridan & reingold, 2014; silva et al., 2022) have shown that experts develop more efficient eye patterns than novices. expert yupana users also tend to show characteristic eye patterns, such as lower dispersion and longer fixation durations on relevant areas. experts showed significantly increased values for both vrt dispersions. the results suggest that they were in scanning mode, actively searching for information on the screen. expert 2 showed an even more intense scanning mode. both also showed a high degree of concentration on specific areas of the board, indicating an efficient and focused visual strategy. the aim of the present study was to evaluate the differences in visual attention patterns during the teaching-learning process, measured by eye tracking, between groups of novices and experts in the yitp method as well as the changes in these patterns before and after learning the method. the results obtained support the hypotheses raised and confirm significant differences in the areas of dispersion and ocular attention between the tests carried out before and after learning yitp method in the group of novices. these results suggest that learning the yitp method influences the way in which participants direct their visual attention when solving arithmetic problems on the yupana board. significant differences were found in the patterns of visual attention between the groups of experts and novices in the yitp method. the experts showed greater efficiency in the use of visual attention, characterised by less dispersion and greater concentration in relevant areas. these results support the idea that experience and mastery of a method influence the cognitive processes underlying problem solving, and it is seen that the strengthening process of attention and the patterns of dispersion are closer to experts. the results of this study have important implications for an initial exploration of the cognitive processes involved in learning and solving mathematical problems using the yitp method. identifying differences in the patterns of ocular attention between novices and experts contributes to the development of more effective teaching and training strategies. similarly, the findings regarding changes in visual attention after learning the yitp method provide valuable information for the design of educational interventions that promote the development of cognitive skills related to problem solving. 6. conclusions this exploratory study provides evidence that the yupana inka tawa pukllay (yitp) method enhances arithmetic learning in university students by strengthening visuospatial and spatial skills. eye-tracking data reveal that, after applying the method, novice students exhibit reduced gaze dispersion and improved focus on key patterns, suggesting more efficient processing in basic operations with natural numbers involved in yupana manipulation. furthermore, their visual fixation patterns progressively align with those of experts, confirming that yitp facilitates the adoption of new strategies, such as rapid identification of relevant elements and effective use of spatial visualization. however, due to the small sample size, the novelty of the arithmetic system, and the observed results, further research with larger sample sizes is recommended to enable statistical comparisons and generalize findings across all arithmetic operations. qualitative findings complement these results, indicating that students gain confidence in handling abstract concepts through the method’s concrete visual aids. they also perceive it as a motivating, game-like tool—factors that facilitate the teaching-learning process (playful motivation factors that encourage learning, particularly in mathematics). this effect is more significant because participants were enrolled in non-mathematical degree programs, suggesting that the method’s playful nature may help reduce student resistance. thus, yitp emerges as an innovative pedagogical tool that enables personalized problem-solving approaches and helps educators identify specific student difficulties. this study supports yitp as a valuable resource for modern mathematics education, illustrating how ancestral tools like the yupana can be adapted to contemporary academic contexts. its visual and interactive approach not only optimizes arithmetic instruction but also fosters more inclusive and effective learning environments. these findings open new avenues for integrating cognitive and cultural techniques in higher education, particularly in fields requiring spatial and logical-mathematical competence. additionally, to ensure concurrent validity, it is advisable to use eeg as an objective complementary tool for assessing changes in brain activity dynamics, especially given the limited subject group in this study. the results support the idea that yitp produces numerical outcomes, which aligns with findings from other studies, such as those on semiotic alternation in mathematics. acknowledgments we would like to express our sincere gratitude to the students of the general studies program at the university of lima, who participated in this study in a voluntary and invaluable capacity. this research project received institutional support from the directorate of general studies and the neuroscience laboratory, under the direction of the dean and her staff. the research institute of the university of lima must be acknowledged for its pivotal role in facilitating the research that recovers ancestral knowledge and the use of emerging technology. the yupanki association also warrants recognition for its invaluable academic collaboration that has been instrumental in propelling this line of research forward. references ansari, d. (2008). effects of development and enculturation on number representation in the brain. nature reviews neuroscience, 9 (4), 278-291. https://doi.org/10.1038/nrn2334 arsalidou, m., & pascual-leone, j. (2016). constructivist developmental theory is needed in developmental neuroscience. npj science learning, 1 , 16016.https://doi.org/10.1038/npjscilearn.2016.16 ausubel, d. p., novak, j. d., & hanesian, h. (1978). educational psychology: a cognitive view .(2nd ed.). holt, rinehart and winston. bhattacharyya, r. k. (2011). brahmagupta: the ancient indian mathematician. in b. yadav & m. mohan (eds.), ancient indian leaps into mathematics (pp. 185–192). birkhäuser boston.https://doi.org/10.1007/978-0-8176-4695-0_12 blascheck, t., kurzhals, k., raschke, m., burch, m., weiskopf, d., & ertl, t. (2017). visualization of eye tracking data: a taxonomy and survey. computer graphics forum, 36(1), 260–284.https://doi.org/10.1111/cgf.13079 borji, a., cheng, m., hou, q., jiang, h., & li, j. (2019). salient object detection: a survey. computational visual media, 5(2), 117–150.https://doi.org/10.1007/s41095-019-0149-9 bruner, j. (1974). toward a theory of instruction. harvard university press. chang, t. t., chen, n. f., & fan, y. t. (2022). uncovering sex/gender differences of arithmetic in the human brain: insights from fmri studies. brain and behavior, 12(10), e2775.https://doi.org/10.1002/brb3.2775 cooper, p. a. (1993). paradigm shifts in designed instruction: from behaviorism to cognitivism to constructivism. educational technology, 33 (5), 12-19. cui, z., hu, y., wang, x., li, c., liu, z., cui, z., & zhou, x. (2024). form perception is a cognitive correlate of the relation between subitizing ability and math performance. cognitive processing, 1-11.https://doi.org/10.1007/s10339-024-01175-3 dehaene, s. (2011). the number sense: how the mind creates mathematics . oxford university press. dehaene, s., piazza, m., pinel, p., & cohen, l. (2005). three parietal circuits for number processing. in j. i. d. campbell (ed.), handbook of mathematical cognition (pp. 433–453). psychology press. dehaene, s., molko, n., cohen, l., & wilson, a. j. (2004). arithmetic and the brain. current opinion in neurobiology, 14(2), 218-224.https://doi.org/10.1016/j.conb.2004.03.008 dehaene, s. (1997). the number sense: how the mind creates mathematics . new york, ny: oxford university press. dogusoy-taylan, b., & cagiltay, k. (2014). cognitive analysis of experts’ and novices’ concept mapping processes: an eye tracking study. computers in human behavior, 36, 82-93. https://doi.org/10.1016/j.chb.2014.03.036 duchowski, a. t., medlin, e., cournia, n., gramopadhye, a., melloy, b., & nair, s. (2002). 3d eye movement analysis for vr visual inspection training. in proceedings of the 2002 symposium on eye tracking research & applications (pp. 103–110). association for computing machinery.https://doi.org/10.1145/507072.507094 escotto-cordova, a., & sanchez ruiz, j. g. (2018). recursos semióticos en la enseñanza de las matemáticas . universidad nacional autónoma de méxico, facultad de estudios superiores zaragoza. gagnier, k. m., holochwost, s. j., & fisher, k. r. (2022). spatial thinking in science, technology, engineering, and mathematics: elementary teachers' beliefs, perceptions, and self‐efficacy. journal of research in science teaching, 59 (1), 95-126.https://doi.org/10.1002/tea.21722 gegenfurtner, a., lehtinen, e., & säljö, r. (2011). expertise differences in the comprehension of visualizations: a meta-analysis of eye-tracking research in professional domains. educational psychology review, 23 , 523-552. https://doi.org/10.1007/s10648-011-9174-7 gu, y., & paracha, s. (2023, november). when eyes tell a story… an eye-tracking approach towards creating a fit-for-purpose learning management system for higher education. in 2023 ieee international conference on development and learning (icdl) (pp. 306-311). ieee. https://doi.org/10.1109/icdl55364.2023.10364450 guzman-jimenez, r., dhavit-prem, saldívar, a., & escotto-córdova, a. (2023). semiotic alternations with the yupana inca tawa pukllay in the gamified learning of numbers at a rural peruvian school. educational technology & society, 26 (1), 79–94. hartmann, m., mast, f. w., & fischer, m. h. (2016). counting is a spatial process: evidence from eye movements. psychological research, 80 (3), 399-409. https://doi.org/10.1007/s00426-015-0722 hayes, p. j. (1978). cognitivism as a paradigm. behavioral and brain sciences, 1 (2), 238–239. https://doi.org/10.1017/s0140525x00074231 kaczkurkin, a., raznahan, a., & satterthwaite, t. (2019). diferencias sexuales en el cerebro en desarrollo: conocimientos de la neuroimagen multimodal. neuropsychopharmacology, 44(1), 71–85.https://doi.org/10.1038/s41386-018-0111-z kowalski, r. (1979). algorithm = logic + control. communications of the acm, 22 (7), 424–436. https://doi.org/10.1145/359131.359136 kerkhoff, y., wedepohl, s., nie, c., ahmadi, v., haag, r., & behrends, s. (2022). a fast open-source fiji-macro to quantify virus infection and transfection on single-cell level by fluorescence microscopy. methodsx, 9 , 101834.https://doi.org/10.1016/j.mex.2022.101834 kunz, a. k., zlatkin-troitschanskaia, o., schmidt, s., nagel, m. t., & brückner, s. (2024). investigation of students' use of online information in higher education using eye tracking. smart learning environments, 11 (1), 44. https://doi.org/10.1186/s40561-024-00300-3 lo, s., & andrews, s. (2022). the effects of mental abacus expertise on working memory, mental representations, and calculation strategies used for two-digit hindu-arabic numbers. journal of numerical cognition, 8(1), 89–122.https://doi.org/10.5964/jnc.8073 marlina, y., & yunas, m. f. (2024). evaluation of digital-english project based learning (pjbl) model through eye tracking analysis in higher education. edutec: journal of education and technology, 7(3), 401–411. mix, k. s., & cheng, y. l. (2012). the relation between space and math: developmental and educational implications. in advances in child development and behavior, 42 , 197–243.https://doi.org/10.1016/b978-0-12-394388-0.00006-x nickerson, r. s. (1988). counting, computing, and the representation of numbers. human factors: the journal of the human factors and ergonomics society, 30 (2), 181–199.https://doi.org/10.1177/001872088803000206 ooms, k., de maeyer, p., & fack, v. (2014). study of the attentive behavior of novice and expert map users using eye tracking. cartography and geographic information science, 41 (1), 37–54. https://doi.org/10.1080/15230406.2013.860255 ott, n., brünken, r., vogel, m., & malone, s. (2018). multiple symbolic representations: the combination of formula and text supports problem solving in the mathematical field of propositional logic. learning and instruction, 58 , 88–105.https://doi.org/10.1016/j.learninstruc.2018.04.010 piaget, j. (1970). piaget’s theory. in p. h. mussen (ed.), carmichael’s manual of child psychology (vol. 1, pp. 703–732). wiley. prem, d. (2016). yupana inka decodificando la matemática inka: método tawa pukllay. asociación yupanki. prem, d. (2018). hatun yupana qellqa: método tawa pukllay. asociación yupanki. prem, d., guzman-jimenez, r., sotomayor, f., & saldivar, a. (2022). tawa pukllay proof: new method for solving arithmetic operations with the inca yupana using pattern recognition and parallelism. in 2022 international conference on frontiers of artificial intelligence and machine learning (faiml) (pp. 209–218).https://doi.org/10.1109/faiml57028.2022.00048 ram, s. s., & ramakalyani, v. (eds.). (2022). history and development of mathematics in india: proceedings of the annual conference on history and development of mathematics, 2018 . national mission for manuscripts. richeson, a. w. (1933). the number system of the mayas. the american mathematical monthly, 40 (9), 542–546.https://doi.org/10.1080/00029890.1933.11987486 ribeiro da silva junior, l., henrique goncalves cesar, f., theoto rocha, f., & eduardo thomaz, c. (2018). a combined eye-tracking and eeg analysis on chess moves. ieee latin america transactions, 16(5), 1288–1297.https://doi.org/10.1109/tla.2018.8407099 sheridan, h., & reingold, e. m. (2014). expert vs. novice differences in the detection of relevant information during a chess game: evidence from eye movements. frontiers in psychology, 5, article 941.https://doi.org/10.3389/fpsyg.2014.00941 silva, a., afonso, j., sampaio, a., pimenta, n., lima, r., castro, h., & murawska-ciałowicz, e. (2022). differences in visual search behavior between expert and novice team sports athletes: a systematic review with meta-analysis. international journal of environmental research and public health, 19 (12), 7172.https://doi.org/10.3389/fpsyg.2022.1001066 šola, h. m., qureshi, f. h., & khawaja, s. (2024). ai eye-tracking technology: a new era in managing cognitive loads for online learners. education sciences, 14(9), 933. https://doi.org/10.3390/educsci14090933 srinivasan, m., wagner, k., frank, m. c., & barner, d. (2018). the role of design and training in artifact expertise: the case of the abacus and visual attention. cognitive science, 42(suppl 3), 757–782.https://doi.org/10.1111/cogs.12611 sundstedt, v., & garro, v. (2022). a systematic review of visualization techniques and analysis tools for eye-tracking in 3d environments. frontiers in neuroergonomics, 3 , 910019.https://doi.org/10.3389/fnrgo.2022.910019 thayaseelan, k., zhai, y., li, s., & liu, x. (2024). revalidating a measurement instrument of spatial thinking ability for junior and high school students. disciplinary and interdisciplinary science education research, 6 (1), 3.https://doi.org/10.1186/s43031-024-00095-8 uttal, d. h., meadow, n. g., tipton, e., hand, l. l., alden, a. r., warren, c., & newcombe, n. s. (2013). the malleability of spatial skills: a meta-analysis of training studies. psychological bulletin, 139(2), 352.https://doi.org/10.1037/a0028446 van der weijden, f. a., kamphorst, e., willemsen, r. h., kroesbergen, e. h., & van hoogmoed, a. h. (2018). strategy use on bounded and unbounded number lines in typically developing adults and adults with dyscalculia: an eye-tracking study. journal of numerical cognition, 4(2), 337–359.https://doi.org/10.5964/jnc.v4i2.115 wang, c. (2020). a review of the effects of abacus training on cognitive functions and neural systems in humans. frontiers in neuroscience, 14 , 913.https://doi.org/10.3389/fnins.2020.00913 yang, c. c., totzek, j. f., lepage, m., & lavigne, k. m. (2022). normative sex differences in cognition and morphometric brain connectivity: evidence from 30,000+ uk biobank participants. biorxiv.https://doi.org/10.1101/2022.10.12.511938 yuan, l., kong, f., luo, y., zeng, s., lan, j., & you, x. (2019). gender differences in large-scale and small-scale spatial ability: a systematic review based on behavioral and neuroimaging research. frontiers in behavioral neuroscience, 13 , 128.https://doi.org/10.3389/fnbeh.2019.00128 appendix 1: yitp moves all these moves are part of the method yupana inka tawa pukllay (yitp), developed by author dhavit prem and the yupanki association (prem, d., 2016). short open movement. this move is made when square 2 has more than one piece. it consists of taking half of the tokens from box 2 to box 1 and the other half to box 3. if the number of tokens is odd, one token is left in box 2, and with the remaining tokens proceed as in the case of the even number of tokens (see figure 2). figure 2 “short open” movement. yitp figure 3 “long open” movement. yitp pisqa movement (birth of new line). this move is made when square 5 has more than one piece. it consists of taking half of the pieces from square 5 to square 1 of the row immediately above and removing the other half from the board. if the number of tokens in box 5 is odd, 1 token is left in box 5, and with the remaining tokens proceed as in the case of the even number of tokens (see figure 4). figure 4 “pisqa” movement yitp kikin movement (box 1, mnemonic: “2 in 1 = 1 in 2”; “3 in 1 = 1 in 3”; “5 in 1 = 1 in 5”) : this movement consists of replacing a group of 5 pieces of box 1 for 1 token in box 5; replace a group of 3 checkers in box 1 with a 1 checker in box 3 and replace 2 checkers in box 1 with 1 checker in box 2, whichever is the case (see figure 5). figure 5 “kikin” movement yitp pichana movement (boxes 1 and 2): this movement is performed in two cases, (see figure 6): pichana (1 and 2) this movement consists of moving as many pieces from box 1 to box 3 as there are pieces in box 2 of the same row and then removing that number of pieces from the board from box 2. pichana (2 and 3) this movement consists of moving as many pieces from square 2 to square 5 as there are pieces in square 3 of the same row and then removing that number of pieces from the board from square 3. figure 6 pichana movement yitp appendix 2: heatmaps test 1 and test 2 appendix 3: interviews p1 (♀ ): solved maintaining an order in columns starting with the ”pisqa” move (5 dots column) and ending with ”kikin” (1 dot column). she used her own algorithm, which was evidenced by observing the heat map in the fixation areas. of figure 3. she correctly solved all 7 additional exercises. its average addition resolution time is 1.13 min. no advanced moves were observed. test 2 took 9 more moves to solve the exercise, but 20 seconds less. even though she has used more moves, she has decreased her solution time by 20%. time per move (t/m-yitp) has reduced by 46%. p2 (♂ ): solved maintaining order by columns (basic moves). on test 2 she correctly solved 3 of the 7 addition exercises and her average time for solving additions was 1.04 min. no advanced moves were observed. test 2 took 4 more moves to solve the exercise, but 14 seconds less, so the m/ryitp was reduced by 4% p3 (♀ ): solved maintaining order by columns (basic moves). on test 2 she correctly solved 6 of the 7 additional exercises and his average solution time was 1.01 min. no composite moves were observed. on test 2 she used two moves less than in test 1 and it took her 36 seconds less to solve the exercise (29.3% less). t/m-yitp fell by 22.6%. a notable increase is observed in the vrt dispersion with a tendency towards the patterns observed in experts. p4 (♂ ): solved by maintaining order from units to ten thousands (bottom to top). on test 2 he correctly solved 6 of 7 addition exercises and her average time for solving additions is 51 s. he performed an advanced move, ”hatun pichana”. on test 2 he reduced 8 moves and did it 14 seconds faster. although the t/m-yitp increased, the number of final moves was reduced, which responds to the participant’s use of composite moves. a difference has been seen in the heat points, greater dispersion in the evaluation. he showed great confidence when solving the exercise. p5 (♀ ): solved by maintaining order from units to tens of thousands (from bottom to top). on test 2 she was not able to solve any sum correctly, she knows the moves, but she made mistakes in one move per game and her average time for solving sums is 1:10 min. no composite moves were observed. test 2 needed 6 more moves, but it was 2:23 minutes faster. the t/m-yitp was accelerated by 70%. heat zones remain similar with a slight increase in dispersion. p6 (♂ ): solved maintaining order from units to tens of thousands (from bottom to top). on test 2 he solved 5 of 7 sums correctly and his average time for solving sums was 1:22 min. no composite moves were observed. test 2 needed 2 fewer moves than in test 1 and did it 1:03 minutes faster. heat zones remain similar. t/m-yitp fell by 39%. p7 (♀ ): solved maintaining order, starting from thousands and going down to units. on test 2 she solved 1 of 8 sums correctly. she is confused about one of the basic moves: ”kikin”, which leads to bad results. she does the other moves well and her average time for solving sums is 1:38 min. one advanced move was observed: ”sonqo”. test 2 took 4 moves less and 6 seconds longer to solve the exercise. t/m-yitp rose 26%. p8 (♂ ): solved by maintaining order kimsa, kikin, pisqa, iskay. in the evaluation, he solved 3 of 7 sums correctly and his average time for solving sums is 1:13 minutes. no composite moves were observed. test 2 took 8 moves less and 1:46 min less. the t/m-yitp ratio decreased by 26% experts, they prioritise advanced moves. if there are none, experts seek to create their own composite moves. if not, basic moves are performed. it is observed that in the kamachiq both experts use different algorithms evidenced in the heat map (see figure 3c and 3d). the average time for solving sums was between 21-31 s. t/m-yitp is similar for both experts. both experts correctly solved all 7 sums. frontline learning research vol. 11 no. 1 (2023) 57 93 issn 2295-3159 corresponding author: emilie prast, faculty of social and behavioral sciences, institute of education and child studies, leiden university, wassenaarseweg 52, 2333ak leiden, the netherlands. e.j.prast@fsw.leidenuniv.nl. doi: https://doi.org/10.14786/flr.v11i1.1079 what do students think about differentiation and within-class achievement grouping? emilie j. prasta, kim stroeta, arnout koornneefa, & tom f. wilderjansbcde aeducational sciences program group, institute of education and child studies, faculty of social and behavioral sciences, leiden university, leiden, the netherlands bmethodology and statistics research unit, institute of psychology, faculty of social and behavioral sciences, leiden university, leiden, the netherlands cleids universitair medisch centrum (lumc), leiden institute for brain and cognition (libc), leiden, the netherlands dresearch group of quantitative psychology and individual differences, faculty of psychology and educational sciences, ku leuven, leuven, belgium edepartment of clinical psychology, faculty of behavioural and human movement sciences, vrije universiteit amsterdam, amsterdam, the netherlands article received 7 april 2022/ revised 19 february 2023/ accepted 22 february 2023/ available online 23 march 2023 abstract differentiation and achievement grouping are frequently implemented practices to adapt education to students’ varying educational needs based on achievement level. potential didactical and socioemotional advantages and disadvantages of these practices have been discussed in the literature. however, little is known about the perspective of students themselves. this study examined how dutch students (n = 428) perceived differentiation and within-class homogeneous achievement grouping in primary mathematics education, with attention for potential differences between students of diverse achievement levels. students of grades 1, 3 and 5 completed a questionnaire about various differentiated mathematics activities and (if applicable) within-class achievement grouping. in line with the didactical perspective on differentiation, extended instruction and less difficult tasks were appreciated most by low-achieving students whereas more difficult tasks were appreciated most by high-achieving students. students of all achievement groups had largely positive attitudes about achievement grouping and about their own achievement group. however, some differences between achievement groups were found, with less favourable results for students placed in low achievement groups. students’ responses to open-ended questions provided additional insights into the reasons behind students’ evaluations of differentiation and achievement grouping. differences between grade levels were also explored. keywords: differentiation; ability grouping; student perspective; mixed methods; mathematics education. mailto:e.j.prast@fsw.leidenuniv.nl 58 | f l r 1. introduction many teachers strive to adapt education to their students’ diverse educational needs by implementing differentiation (‘an approach by which teaching is varied and adapted to match students’ abilities using systematic procedures for academic progress monitoring and data-based decision-making.’; roy, guay, & valois, 2013, p.1187). however, differentiation is a controversial topic in the literature, particularly when it is organised by grouping students of a similar achievement level. researchers have discussed potential didactical and socioemotional advantages and disadvantages of various types of differentiation and achievement grouping (e.g., campbell, 2021; francis et al., 2017; marks, 2013; mcgillicuddy & devine, 2020; tieso, 2003; tomlinson et al., 2003; van geel et al., 2018). in this debate, students’ voices have not often been heard. given that the aim of differentiation is to adapt education to students’ needs, it is important to examine whether students themselves perceive the adaptations as successful in meeting their educational and socioemotional needs. therefore, this study investigates what students think about differentiation and within-class achievement grouping in primary mathematics education. 1.1 differentiation and achievement grouping differentiation based on students’ current academic achievement level (also called readiness-based or cognitive differentiation) entails two related processes: (1) monitoring students’ progress to determine their current achievement level and educational needs, and (2) adapting learning goals, instruction and practice to students’ current level of knowledge and skills and their corresponding educational needs (prast et al., 2015; roy et al., 2013). differentiation may be convergent or divergent (blok, 2004). in convergent differentiation, all students work towards the same goals, but the way in which students reach these goals is differentiated (e.g., with additional instruction). in divergent differentiation, students of different achievement levels work towards different learning goals. one frequently used way to organise differentiation is to group students based on achievement level. such groups may be homogeneous (similar achievement) or heterogeneous (mixed achievement), and within-class or between-class (see tieso, 2003 for an overview of grouping practices). this paper focuses on within-class homogeneous grouping, that is: subgroups of students with a similar achievement level within a class which includes a broad range of achievement levels. we use the term achievement grouping rather than ability grouping since recent guidelines for grouping and differentiation do not assume that students have a fixed ability level and instead emphasise that grouping arrangements should be flexible and responsive to changes in students’ educational needs (prast et al., 2015; tomlinson et al., 2003; van geel et al., 2018). achievement grouping can be used to differentiate instruction (e.g., additional instruction for subgroups with similar instructional needs) and practice (e.g., with tiered tasks for low-achieving, average-achieving and high-achieving students) (prast et al., 2015). this paper does not only examine students’ views on the grouping itself, but also on differentiated mathematics activities which may or may not take place in achievement groups. 1.2 mathematics education in the netherlands since the implementation of differentiation relies heavily on domain-specific pedagogical content knowledge (prast et al., 2015; van geel et al., 2018; vogt & rogalla, 2009), teachers’ implementation and students’ perceptions of differentiation are likely to be domain-specific. to study students’ perceptions of differentiation and achievement grouping in sufficient depth, this study focuses on one domain and context, namely primary mathematics education in the netherlands. this is a relevant context, because of the increased focus on data-based decision making and, accordingly, on progress monitoring and instructional adaptations in the subject of mathematics in the netherlands over the past decade (prast et al., 2018; van geel et al., 2016, 2018; visscher, 2015). a recent review (prast & hickendorff, in press) about the implementation of differentiation in primary mathematics education in the netherlands indicated that most teachers differentiate instruction and practice based on students’ achievement level at least to some extent. a mathematics lesson typically starts with a whole-class instruction. subsequently, additional instruction is often provided to low-achieving students, whereas practice tasks are frequently differentiated at three levels. 59 | f l r in the lower grades, differentiation is largely convergent since students work towards the same learning goals, but from grade 4 onwards the learning goals may also be differentiated (expertgroep doorlopende leerlijnen, 2008). differentiation is frequently organised using within-class achievement groups (i.e., subgroup instruction and tiered tasks), but alternatives such as individualised differentiation of the practice tasks using software are also used (prast & hickendorff, in press). 1.3 perspectives on differentiation and achievement grouping various theoretical perspectives on differentiation and achievement grouping can be taken. in the current study, we focus on two perspectives: a didactical perspective (concerning the teaching and learning of mathematical content, with a focus on cognitive processes) and a socioemotional perspective (concerning the social and emotional processes that may be involved when differentiated activities and achievement grouping are used). we formulated our hypotheses based on these two perspectives. while other perspectives might also be taken (e.g., a sociological perspective concerning implications of differentiation and achievement grouping at a societal level), we feel that these two perspectives are most relevant in the context of the current study, because they are closely related to students’ daily experiences with differentiation and achievement grouping (and therefore, relevant and understandable for students). 1.4 a didactical perspective on differentiation and achievement grouping from a didactical perspective, the rationale for readiness-based differentiation is that adapting instruction and practice to students’ current skill level enhances learning (tomlinson et al., 2003). according to this view, learning tasks should be at a moderate difficulty level in relation to a student’s current skills (csikszentmihalyi, 1990; murray & arroyo, 2002). when tasks are too easy, this may result in boredom and withdrawal, while confronting a student with tasks that are too difficult may lead to frustration and anxiety (csikszentmihalyi, 1990; murray & arroyo, 2002). when tasks are designed to be just within reach based on the skill level of the student, this may enhance students’ motivation and achievement (arroyo et al., 2014; csikszentmihalyi, 1990). more generally, aptitude-treatment interaction theory predicts that students need different instructional treatments, dependent upon their aptitude (readiness for learning based on current achievement level) (cronbach & snow, 1977; kalyuga, 2007). for example, direct explicit instruction may be highly effective for students with low prior knowledge but not for students with high prior knowledge (kalyuga, 2007; kirschner et al., 2006). in the literature, research relating differentiation and achievement grouping to student achievement is described. for example, a meta-analysis (deunk et al., 2018) in primary school found positive effects of interventions in which software was used to assist teachers in implementing differentiation, by continuously monitoring students’ achievement level and providing (suggestions for) differentiated instruction and practice. programmes in which differentiation was part of a broader school reform also had positive effects. within-class grouping had no overall effect on achievement (for all students together), but there was a negative effect on the achievement of students placed in low achievement groups. however, the effects of achievement grouping were difficult to interpret, because the original studies provided little information on whether and how instruction and practice were differentiated in the achievement groups. a more recent study comparing within-class grouping and whole-class teaching in the uk (jerrim, 2021) did not find clear evidence for effects of achievement grouping on student achievement. from a didactical perspective, achievement grouping is merely an organisational format that can be used to implement differentiation, provided that the groups are actually used to adapt instruction to students’ needs. however, achievement grouping may also have negative didactical consequences, for example if the grouping arrangements do not correspond accurately to students’ current achievement level (e.g., because the groups are insufficiently flexible) or if the quality of differentiation is limited (e.g., with insufficiently challenging learning materials for lower achievement groups). more generally, if the learning goals and tasks are differentiated, low-achieving students may not get the opportunity to reach the same learning goals as their high-achieving peers (divergent differentiation; blok, 2004; see also hart, 1992). 60 | f l r 1.5 a socioemotional perspective on achievement grouping from a socioemotional perspective, achievement grouping is not just a format to organise differentiation, but an educational approach that may affect socioemotional processes within a class. first, achievement grouping may affect social comparison processes, with potential effects on students’ academic self-concept. qualitative case studies have indicated that, even when neutral names are used for the achievement groups, primary school students are largely aware of the hierarchical grouping structure, especially in the higher grades (eder, 1983; gripton, 2020; marks, 2013; mcgillicuddy & devine, 2020). campbell (2021) describes two possible mechanisms for effects of achievement grouping on academic self-concept: labelling effects and reference group effects. in the case of labelling effects, students would internalise the achievement label belonging to their achievement group, with positive effects on the self-concept of students placed in high achievement groups and negative effects on the self-concept of students placed in low achievement groups. in the case of reference group effects, students would start to compare themselves to the other students placed in their achievement group rather than to the whole class, with positive effects on the self-concept of students placed in low achievement groups and negative effects on the self-concept of students placed in high achievement groups (the big-fish-little-pond effect; marsh, 1984, 1987). in a large-scale study about the effects of within-class grouping on students’ self-concept, campbell (2021) found more evidence for labelling effects than for reference group effects. in contrast, jerrim (2021) found no effects of within-class achievement grouping compared to whole-class teaching on students’ self-concept. more generally, the use of achievement groups may affect the social dynamics within a class. qualitative case studies have provided indications that placement in a high achievement group may be associated with a higher social status than placement in a low achievement group (marks, 2013; mcgillicuddy & devine, 2020). for example, students placed in high achievement groups have been described by their peers as ‘smart’, ‘good’ or ‘liked’, whereas students placed in low achievement groups have been described as ‘dumb’, ‘bad’ or ‘not liked’ (mcgillicuddy & devine, 2020). in a study (hargreaves et al., 2021) about peer relations of students placed in low achievement groups (including both within-class and between-class grouping systems), there was generally little evidence that troubles in peer relations were related to placement in a low achievement group. however, in some cases, students experienced feelings of exclusion that were related to their low achievement status. in a study including various types of between-class and within-class achievement grouping, students placed in low and high achievement groups reported achievement-related teasing (hallam et al., 2004). besides potential effects on peer interactions, achievement grouping might also affect teacher-student interactions. for example, teachers may implicitly or explicitly display different expectations about students placed in low versus high achievement groups (mcgillicuddy & devine, 2020; rubie-davies, 2014; van den bergh, 2018). taken together, this socioemotional perspective indicates that within-class achievement grouping may affect various socioemotional processes in the classroom, which may be experienced differently by students placed in low, average or high achievement groups. however, as described above, the direction of effects is not always clear and empirical research about the socioemotional aspects of within-class achievement grouping is relatively scarce. 1.6 students’ perspective on differentiation and achievement grouping when researching differentiation it makes sense to include students’ perspective, since students’ motivation and engagement will be shaped by their perceptions. little is known about students’ views on differentiation and within-class achievement grouping. since the goal of differentiation is to adapt education to students’ educational needs, it is relevant to know whether students feel that their educational needs are met by the adaptations made by teachers. besides, students can be an important source of information regarding potential socioemotional side-effects of differentiation and achievement grouping. questions may be raised regarding the validity of student perceptions as an indicator of the “best” practice in terms of student outcomes such as achievement or motivation: we cannot expect students to oversee all implications of within-class differentiation and achievement grouping. accordingly, the goal of this study is not to give a 61 | f l r complete overview of all potential effects, but to zoom in on one perspective that has not received much attention so far; that of the students themselves. previous research about this topic is scarce, and mostly focused on the grouping itself rather than on specific differentiation practices. first, there are studies comparing different types of grouping (e.g., between-class, within-class, whole-class teaching). such studies have reported negative experiences with between-class achievement grouping (boaler et al., 2000), and have presented mixed-achievement classes as a favourable alternative (hallam et al., 2004; tereshchenko et al., 2019). however, these studies did not focus on the differentiation practices that might be implemented within mixed-achievement classes. second, the smallscale qualitative studies that we have described in the previous section (eder, 1983; gripton, 2020; marks, 2013; mcgillicuddy & devine, 2020) provided insights into students’ experiences with within-class achievement grouping, with a focus on social-emotional aspects related to the grouping. these studies did not directly ask students about their preferences regarding grouping or differentiation. such studies are very scarce, but a third line of studies did ask students directly about their preferences regarding adaptations for students with special educational needs included in general education classrooms. generally, students with and without special educational needs had positive attitudes towards many of these adaptations, although they wanted everybody to have the same homework (vaughn et al., 1995; vaughn, schumm, niarhos, & daugherty, 1993; vaughn, schumm, niarhos, & gordon, 1993). these three lines of research have provided initial indications that students’ experiences or preferences may differ depending on their achievement level, although the direction of effects is not always consistent across studies. for example, one study found that low-achieving students had the most positive attitudes towards mixed-achievement classes, while highachieving students also perceived disadvantages such as a lack of challenge (tereshchenko et al., 2019). in contrast, another study found that low-achieving students tended to prefer homogeneous grouping whereas high-achieving students tended to prefer heterogeneous grouping (vaughn et al., 1995). 1.7 the current study: research questions and hypotheses the originality of the current study lies in the following. we zoomed in on students’ perspective on differentiation and grouping practices within primary school classes. first, we made the abstract concept of differentiation more concrete by asking students about various mathematics activities and relating their evaluations to students’ scores on a standardised mathematics achievement test. second, we investigated students’ opinions on within-class grouping using a mixed-methods approach combining quantitative ratings with qualitative reasons. throughout the study, we had attention for both didactical and socioemotional considerations (rather than focusing on either), and for the perspectives of students of diverse achievement levels and grade levels. the first aim of this study was to investigate whether students’ evaluations of various mathematics activities are dependent upon their achievement level (regardless of the use of achievement grouping in their class). within the didactical perspective on differentiation, the idea of aptitude-treatment interactions is central: students are supposed to have different educational needs depending on their current achievement level (cronbach & snow, 1977; prast et al., 2015; tomlinson et al., 2003). for example, the same activity may be appropriately challenging for some students and too difficult or too easy for other students. note that previous research on aptitude-treatment interactions has typically focused on the outcome of student achievement, whereas students’ perceived frequency, liking and learning from activities are the outcome variables in this study. thus, the first research question was: (1) do different students evaluate various mathematics activities differently, depending on an interaction between the type of activity and the achievement level of the student? in accordance with guidelines for differentiation (prast et al., 2015), we made a distinction between general activities for all students (whole-class instruction, working at mathematics tasks independently, working at mathematics tasks together), activities intended to serve the educational needs of low-achieving students (less difficult tasks and additional instruction in a subgroup or individually), and activities intended for high-achieving students (more difficult tasks and additional instruction about these enrichment tasks, in a subgroup or individually). based on the didactical perspective 62 | f l r on differentiation, we expected that students’ perceptions of these activities would interact with their achievement level. first, we expected that the frequency of activities as perceived by students would be dependent on achievement level. this would be in line with previous teacher self-report and observational studies indicating that many dutch teachers adapt instruction and practice activities to the achievement level of their students, for example by providing additional instruction to low-achieving students and more challenging tasks to high-achieving students (prast & hickendorff, in press). second, we expected that students’ reported liking of and learning from activities would be dependent upon students’ achievement level. based on the idea of aptitude-treatment interactions, the most probable direction of such an interaction effect would be that activities intended for low-achieving students (such as less difficult tasks) are evaluated more positively by low-achieving students whereas activities intended for high-achieving students (such as more difficult tasks) are evaluated more positively by high-achieving students. however, given the more critical views on differentiation that have also been described in the literature (e.g., hart, 1992), as well as the innovative character of this study, it remains to be seen whether these interaction effects are indeed present and whether the effects are in the hypothesised direction. the second aim of this study was to investigate students’ perceptions of within-class achievement grouping in primary school, with attention for potential differences between students placed in low, average and high achievement groups. we were not only interested in quantitative evaluations but also in the reasons behind students’ evaluations. this led to the following research questions: (2a) how do students placed in withinclass achievement groups evaluate their own achievement group and achievement grouping in general and do these evaluations differ between students placed in low, average and high achievement groups? (2b) which reasons do students provide for their evaluations? based on the indications for potentially different experiences of students placed in low, average and high achievement groups in the literature reviewed above, we expected that students’ evaluations would differ between achievement groups. however, given the scarce and inconsistent previous findings on student perceptions of within-class achievement grouping, we did not make specific predictions regarding the direction of those effects. since we expected that students’ reasons behind the quantitative evaluations might include socioemotional as well as didactical considerations (since grouping is typically used to differentiate tasks or instruction), we asked questions that probed both of these aspects. 2. method 2.1 participants and procedure data were collected in the fall of 2018 in the context of the research project ‘differentiation and motivation in primary mathematics education’, which was approved by the local ethics committee (project number ecpw-2018/210). fifty classes from 18 primary schools in the netherlands participated. after obtaining active informed consent from teachers and students, data were collected by students in the final year of academic teacher training, mostly at the school where they also did a teaching internship. the schools were diverse in terms of school size, location, and pedagogical-didactical school characteristics (e.g., public schools, schools with a religious background, montessori, etc.). we recruited one class of grades 1, 3 and 5 (in which students are typically 6-7, 8-9, and 10-11 years old) in each school to have a spread in grade levels while retaining a substantial number of classes per grade. in multigrade classes (nine classes, 18%), only students from the grades selected for our study participated. if a class had multiple teachers (common since 67% of teachers worked part-time), the teacher who most often taught that class participated. the average class size was 23 students (range 13 – 34, including students who did not participate in the research). teachers had an average of fifteen years of teaching experience (range 0 – 42 years). most teachers (n = 40, 80%) were female, reflecting the general dutch population of primary school teachers. in the context of the overarching research project, the participating teachers were interviewed and completed a questionnaire about their differentiation and achievement grouping practices. this yielded the following background information which we provide to assist in interpreting the findings of the current study. thirtytwo teachers (64%) reported that the use of achievement groups was fully integrated in their mathematics 63 | f l r teaching routine. these teachers would typically start a lesson with a whole-class instruction, followed by independent practice at three difficulty levels as provided by the curricular method. simultaneously, extended instruction would be provided to a subgroup of low-achieving students. another fourteen teachers (28%) reported to use achievement groups partly. these teachers would for example provide extended instruction to a subgroup of students who needed it, but would either provide little differentiation in the tasks or differentiate tasks in a different way, for example using software. four teachers (8%) did not or hardly work with achievement groups. the use of achievement groups was within-class, with one exception (in one school, mathematics was taught in separate classes for low-achieving and average-achieving students). of the teachers using achievement groups (partly or fully), fifteen teachers (30%) indicated to create or update grouping arrangements approximately every two to six weeks based on students’ scores on the end-ofchapter tests from the mathematics textbook. another eleven teachers (22%) reported to make new grouping arrangements twice per year based on the results of a standardised mathematics achievement test. eight teachers (16%) indicated to work with flexible groups, created per lesson or per week based on the teachers’ observations, educational software or students’ own view on whether they needed additional instruction. the remaining teachers created new groups 3 to 4 times per year (6 teachers, 12%), did not change the groups (1 teacher, 2%), created grouping arrangements in a different way (3 teachers, 6%) or had missing responses (2 teachers, 4%). across the various methods of grouping, some teachers indicated that the grouping arrangements could be adapted per lesson based on students’ needs, and that other sources of information such as students’ daily work were also used. table 1 overview of student characteristics in the full sample characteristic n grade level grade 1 45 grade 3 199 grade 5 184 gender boy 214 girl 212 missing 2 achievement level on standardised test a i (highest) 97 ii 57 iii 60 iv 55 v (lowest) 41 missing: test scores not available 118 within-class achievement group b high 108 average 93 low 54 not placed in single within-class achievement group during past 3 weeks 173 a see section 2.2.1. b see section 2.2.2. students not placed in a single within-class achievement group include students who switched between achievement groups during the past three weeks, students placed in between-class achievement groups, and students from classes without achievement grouping. 64 | f l r for the current study, the following data were collected: a student questionnaire, student achievement group placement and student achievement on a standardised mathematics achievement test. the grouping and achievement data were collected from the teacher. the student questionnaire was administered during school hours (maximum duration: 45 minutes). in grades 3 and 5, this questionnaire was administered to all students for whom informed consent had been obtained (n = 383). after an explanation and practice of the answering format, students completed the questionnaire independently under supervision of the research assistant. in grade 1, the same questionnaire was administered individually due to the students’ young age (typically six years). the research assistant read the questions out loud, after which the student could point to the answer (when applicable, see section 2.2.3) or say his or her answer, which was written down by the research assistant. since this individual administration was too resource-intensive to include all students of grade 1, we randomly selected one low-achieving, one average-achieving and one high-achieving student from the students with informed consent in each class (n = 45). thus, the total sample consisted of 428 students with a mean age of 8 years (range 5 – 12 years). an overview of student characteristics is provided in table 1. 2.2 measures 2.2.1 mathematics achievement test mathematics achievement was measured with the nationally administered cito mathematics achievement tests, of which the validity and reliability have been demonstrated (janssen, verhelst, et al., 2005; koerhuis & keuning, 2011). various grade level versions of the test are available, including a version for kindergarteners (janssen, scheltens, et al., 2005; koerhuis, 2010). each grade level version covers multiple mathematics domains, appropriate for the grade level of the students (janssen, verhelst, et al., 2005; koerhuis & keuning, 2011). if available (administration of the test is not mandatory), the most recent test scores obtained at the end of the previous schoolyear were collected from the teacher. to ensure the comparability of scores across various grade-level versions of the test, we used the achievement level scores which reflect students’ achievement level relative to a nationally representative sample: i = 80th – 100th percentile, ii = 60th – 80th percentile, iii = 40th – 60th percentile, iv = 20th – 40th percentile, v = 0 – 20th percentile. in the analyses, these scores were recoded (and centered on the middle group) such that the highest value represents the highest achievement level (v = -2, iv = -1, iii = 0, ii = 1, i = 2). 2.2.2 achievement group placement if teachers used within-class achievement grouping, teachers were asked to indicate for each participating student in which group(s) the student had been placed during the past three weeks. the answering options were low, average, high and other (e.g., when students had switched between groups). we asked for a period of three weeks because this was long enough to experience relatively stable placement in an achievement group, but not so long that most of the students in the sample would have changed groups within that period. since comparisons between students placed in low, average and high achievement groups might be confounded when students had switched between groups, only students placed in a single within-class achievement group during the past three weeks were included in the analyses about achievement grouping. while students’ achievement group placement was generally related to their achievement on the mathematics achievement test, this correspondence was not perfect (see appendix 1 in the supplementary materials). 2.2.3 student questionnaire about differentiated activities and achievement grouping the student questionnaire was developed for this study by the first and second author, based on a model for differentiation in mathematics that is frequently implemented in the netherlands (prast et al., 2015). prior to large-scale administration, a small-scale pilot was conducted by administering it one-on-one to two students of grades 1, 3 and 5 to check whether students understood the questions and the answering format. the first part of the questionnaire asked students about nine mathematics activities (based on prast et al., 2015) representing three categories of activities: general activities (whole-class instruction, working on tasks independently, and working on tasks together), differentiated activities intended for low-achieving students (working on less difficult tasks, extended instruction in a subgroup, and individual extended instruction) and 65 | f l r differentiated activities intended for high-achieving students (working on enrichment tasks, subgroup instruction about enrichment tasks, and individual instruction about enrichment tasks). based on the teacher interview, the names of the activities as used in the students’ own class were used (e.g., if enrichment tasks were called “mathematics tigers”, that term would be used rather than “more difficult tasks”). for each of the nine activities, students were asked how often they were engaged in this activity, how much they liked this activity, and how much they learned from it (for a total of 27 questions). the answering format was a fivepoint scale represented by dots as shown in figure 1 (adapted from park et al., 2016). students had to indicate the dot that corresponded to their answer: the smallest dot corresponded to the smallest magnitude (e.g., never receiving whole-class instruction) and the largest dot corresponded to the largest magnitude (e.g., receiving whole-class instruction every lesson). do you ever get whole-class instruction? never always figure 1. sample item with answering format adapted from park et al. (2016). the second part of the questionnaire (consisting of 4 closed-ended and 5 open-ended questions) was only administered to students whose teachers used achievement groups. students were asked how much they liked to be in their achievement group (answered on the dot-likert scale described above), what they liked (openended) and did not like (open-ended) about being in that achievement group, how much they learned from being in that achievement group (dot-likert scale), why they learned much or little in that achievement group (open-ended), whether and why they would rather be in a different achievement group (yes or no followed by an open-ended explanation), and whether and why they would prefer a system without achievement groups (yes or no followed by an open-ended explanation). note that there is no strict separation between the differentiated activities (part 1) and achievement grouping (part 2): in many classes, within-class achievement grouping was used to organise the differentiated activities. however, the first part focused on differentiated activities, regardless of whether achievement groups were used in that class, while the second part focused on students’ perceptions of achievement grouping (only if relatively fixed achievement groups were used in their class). 2.2.4 analyses data were analysed in two parts, corresponding to the questionnaire and our research questions. part 1 focused on students’ ratings of the various activities, relative to their achievement level as measured with the cito mathematics achievement test. thus, students from classes without relatively fixed achievement groups were also included in these analyses, given that differentiated activities could also be organised in different ways, whereas students for whom cito test scores were not available were excluded from these analyses. part 2 focused on students’ perceptions regarding achievement grouping and therefore included only students who had been placed in a single within-class achievement group. data were analysed in r with multilevel models to take into account the nesting of students within classes and schools. to enhance readability, we focus on the most important steps here while additional statistical details are provided in the supplementary materials (see appendix 2). to answer the first research question (do different students evaluate various mathematics activities differently, depending on an interaction between the type of activity and the achievement level of the student?), we analysed whether there was a significant interaction effect between the type of activity (e.g. 66 | f l r whole-class instruction, independent work, etc.) and students’ achievement level in predicting students’ selfreported frequency of being engaged with the activities, students’ liking of the activities and students’ perceived learning from the activities. for each of these three outcome variables separately, we estimated four-level regression models with activity ratings (e.g., liking of whole-class instruction, independent work, etc.) nested in students (i.e., repeated measures), who were nested in classes nested within schools. to evaluate the significance of the interaction effect, we compared the fit of a full model including main effects of activity and achievement as well as an interaction between these variables to the fit of a reduced model without the interaction effect using a likelihood ration test (lrt). a significant lrt indicates that the full model fits the data significantly better. significant interaction effects were followed up with post-hoc tests evaluating the effect of achievement on the ratings of each activity. finally, potential interactions with grade level were explored (see section 3.2.1). as described in the introduction, we expected that the interaction effects would be significant. to answer the second research question, we performed two types of analyses: quantitative analyses to answer question 2a (how do students placed in within-class achievement groups evaluate their own achievement group and achievement grouping in general and do these evaluations differ between students placed in low, average and high achievement groups?), and qualitative analyses to answer question 2b (which reasons do students provide for their evaluations?). the closed-ended questions were analysed using multilevel models, herewith taking the clustering of students within classes and schools into account. achievement group was specified as a predictor of the outcome variable of interest: liking of achievement group, learning from achievement group, wanting to be in a different achievement group or preferring to work without achievement groups. as the latter two outcomes are dichotomous, a multilevel version of logistic regression was used. likelihood ratio tests were used to determine whether this model fitted the data significantly better than a model without achievement group as a predictor. as described in the introduction, we hypothesised that there would be differences between students placed in low, average and high achievement groups, but did not have specific hypotheses regarding the direction of effects. the analyses of the open-ended questions were exploratory and intended to give more meaning to the quantitative results. an inductive approach, in which students’ answers rather than theoretical expectations formed the starting point, was taken (linneberg & korsgaard, 2019). based on an initial review of students’ answers, the first author developed various lower-order codes (about 10 – 20 codes per question) which specified (aspects of) answers that were given by multiple students. since several codes were related to similar themes, these lower-order codes were then classified as belonging to one of the following higherorder themes, which were largely similar across questions: (1) answers about independent work and its difficulty level (2) answers about (social interactions within) the achievement group (3) answers related to instruction and the teacher (4) answers about learning and understanding and (5) general and other answers (mostly unspecific, e.g. “i just like it”). the coding scheme which was thus developed by the first author was discussed with the second author and revised accordingly. a random sample of 50 cases (21% of the 238 cases with achievement grouping data) was coded by both authors to determine interrater agreement. cohen’s kappa for the lower-order codes ranged between .70 and .86 for the various questions, and percentage agreement between 74.0% and 87.0%, indicating fair to good interrater agreement. the most frequent reason for non-agreement was that one of the authors had coded a statement with an unspecific code (i.e., general or other), whereas the other author had coded it with a specific code. after reviewing the cases of non-agreement, the second author understood the choices the first author had made and mostly agreed with them. the full sample was coded by the first author. 3. results 3.1 part 1: students’ perceptions of differentiated activities the analyses for part 1 included 310 students who provided data on both the questionnaire and the achievement test (for some analyses, n is slightly smaller due to missing data on single items in the questionnaire). students’ achievement test scores were distributed as follows: i (highest achievement) = 97 67 | f l r students, ii = 57 students, iii = 60 students, iv = 55 students, v = 41 students. in accordance with our sampling procedure, most students were in grade 3 (n = 139) and 5 (n = 150). only n = 21 students of grade 1 were included in these analyses, since achievement data of the previous year (when they were still in kindergarten) were not available for many students (see also section 3.1.1). descriptive statistics and results from the model building stage can be found in the supplementary materials (appendix 2). in line with our hypothesis, there were significant interaction effects of activity by achievement for all outcome variables. that is, the full models including the interaction term had a significantly better fit than reduced models without this interaction term and this was the case for students’ ratings of the frequency (χ2 (8) = 182.36, p <.001), liking (χ2 (8) = 161.75, p <.001), and amount of learning from these activities (χ2 (8) = 95.74, p <.001). these interaction effects are visualised in figure 2, which displays the estimated means based on the full models for students’ self-reported frequency (left column), liking (middle column), and learning (right column) of activities, split by achievement level (for visual clarity, the achievement levels ii and iv are not displayed, but these follow the same linear regression). the figure shows that some activities were rated more highly by low-achieving students compared to high-achieving students, whereas this pattern was reversed for other activities. for example, easier tasks (printed in red) were rated more highly by lowachieving students (red dots), whereas enrichment tasks (printed in blue) were rated more highly by highachieving students (blue squares). in the following paragraphs, these interaction effects are interpreted further based on post-hoc tests evaluating the significance of the effect of students’ achievement level on students’ ratings for each activity. first, we consider students’ reported frequency of being engaged in the activities. the general activities of whole-class instruction and working together were reported equally frequently by students of all achievement levels. however, high-achieving students reported to work independently more often (although this difference seems small in the figure, it was significant). as hypothesised, the activities intended for lowachieving students (less difficult tasks, extended instruction in a subgroup, and individual extended instruction) were reported more frequently by low-achieving students. regarding the activities for highachieving students, high-achieving students reported to work on enrichment tasks more frequently, as hypothesised. however, high-achieving students did not report to receive more instruction about enrichment tasks, either in a subgroup or individually: these activities were generally reported infrequently, regardless of the achievement level of the students. second, we consider students’ liking of the activities. regarding the general activities, whole-class instruction was liked somewhat more by low-achieving students, independent work was liked somewhat more by high-achieving students, whereas working together was appreciated equally by students of all achievement levels. as expected, the activities intended for low-achieving students (less difficult tasks, extended instruction in a subgroup, and individual extended instruction) were liked more by low-achieving students. regarding the activities for high-achieving students, enrichment tasks were liked more by highachieving students, in line with our hypothesis. however, students’ liking of instruction about enrichment tasks (in a subgroup or individually) did not depend on students’ achievement level. this should be viewed in light of the relatively low reported frequency of instruction about enrichment tasks. third, we consider students’ reported amount of learning from the activities. students’ reported learning from the general activities (whole-class instruction, independent work and working together) was not dependent upon their achievement level. as expected, low-achieving students reported to learn more from the activities intended for low-achieving students (less difficult tasks, extended instruction in a subgroup, and individual extended instruction) compared to high achieving students. note, however, that even the lowestachieving students rated working on easier tasks lower than the general activity of working independently (as can be seen in figure 2). similar to the results regarding frequency and liking, high-achieving students did report to learn more from enrichment tasks, but did not report to learn more from instruction about enrichment tasks compared to students of other achievement levels. 68 | f l r figure 2. estimated means of student-reported frequency, liking and learning from activities split by achievement level based on the final multilevel models. general activities are printed in black, activities intended for low-achieving students are printed in red, and activities intended for high-achieving students are printed in blue. error bars represent 95% confidence intervals. 3.1.1 differences between grade levels in students’ perceptions of differentiated activities we explored whether the above results varied between grade levels, by adding grade level as a variable to the analyses and examining whether there were three-way interactions of activity by achievement level by grade level. again, the significance of the interactions was determined using likelihood ratio tests comparing the model with three-way interaction to the model without three-way interaction. the three-way interaction was significant for students’ reported frequency of activities (χ2 (16) = 65.20, p <.001) and learning from activities (χ2 (16) = 39.55, p = 0.001), but not significant for liking of activities (χ2 (16) = 25.473, p = 0.062). follow-up analyses split by grade level indicated that the activity by achievement level interaction was not significant in any of the grade 1 analyses. however, this result should be interpreted with caution given the small sample size in grade 1.1 within grades 3 and 5, all activity by achievement level interactions were significant. 1 we repeated the analyses with teachers’ estimation of students’ achievement level rather than the standardised achievement test as an independent variable, thereby increasing the sample size to n = 40. this yielded similar results, namely no significant activity by achievement level interaction effects in grade 1. 69 | f l r figure 3. estimated means of student-reported frequency, liking and learning from activities split by achievement level and grade level based on the final multilevel models. general activities are printed in black, activities intended for lowachieving students are printed in red, and activities intended for high-achieving students are printed in blue. error bars represent 95% confidence intervals. these similarities and differences between grade levels are illustrated in figure 3. in the left column representing grade 1, students’ ratings of the frequency (upper row), liking (middle row), and learning (lower row) of activities have broad confidence intervals which mostly overlap between students of diverse achievement levels. this illustrates that there were no significant differences between students of diverse achievement levels, although the pattern of ratings for liking does suggest systematic differences in the expected direction. in grades 3 (middle column) and 5 (right column), the pattern of effects was similar to the overall analyses, and the differences between low-achieving and high-achieving students tended to become more pronounced between grades 3 and 5. for example, regarding enrichment tasks, it can be seen 70 | f l r that the pattern of higher ratings by higher-achieving students was already present in grade 3 (for frequency, liking and learning of the activities), but these differences became more pronounced in grade 5, as indicated by the larger distance between the scores of low-achieving students and high-achieving students. for easier tasks and extended instruction in a subgroup, the differences between low-achieving students and highachieving students also seemed to increase (partially) between grade 3 and grade 5. however, for individual extended instruction, the difference between achievement levels did not seem to increase between grade 3 and grade 5, nor for the remaining activities (general activities and instruction about enrichment tasks), for which no (big) differences between achievement levels were present in the total sample. summing up, students’ ratings of activities tended to become more strongly related to students’ achievement level in higher grades. 3.2 part 2: students’ perceptions of achievement grouping the analyses of part 2 included 240 students who had been placed in the same within-class achievement group during the past three weeks (low: n = 52, average: n = 87, high: n = 101). for an overview of all codes that resulted from the qualitative analyses and their frequency, see appendix 3 in the supplementary materials. 3.2.1 students’ liking of their achievement group first, we asked students how much they liked to be in their achievement group. the likelihood ratio test indicated that the model with achievement group as a predictor of liking did not fit the data significantly better than the reduced model without achievement group as a predictor (χ2 (2) = 5.21, p =.074). the raw means indicated that students placed in low (m = 3.94, sd = 1.24), average (m = 4.16, sd = 1.26) and high achievement groups (m = 4.35, sd = 0.98) generally liked to be in their achievement group. while the raw means increased with achievement level, the likelihood ratio test indicated that these differences were not significant. most responses to the open-ended question what students liked about their achievement group were related to the higher-order theme of independent work and its difficulty level (see table a8). while comments about this theme were made frequently by students from all achievement groups, the content of the answers differed between achievement groups. many students from high achievement groups mentioned that they liked challenges and difficult tasks (“because if i have this work, it’s a challenge”). in contrast, students from low achievement groups tended to appreciate that tasks were not too difficult, and some also mentioned that they liked to have fewer tasks, enabling them to finish their work. students from average achievement groups mentioned that the difficulty level was appropriate (not too difficult and not too easy) and sometimes explicitly mentioned that “it matches my level”. a second frequently mentioned higher-order theme comprised comments related to the achievement group itself, including its members and the social interactions between them. across achievement groups, students made positive comments about the members of the group (“the children in this group are kind”) and about being able to help each other (“because if you don’t know something the other children can help you”). a few students from average and high achievement groups made explicit comments about liking to be in a higher group: “then i think sometimes that i am a bit smarter together with the other children and that gives me confidence”. the remaining answers were related to the higher-order themes of learning and understanding (e.g., “then you learn more of it”), to instruction and the teacher (e.g., “that you get more help”), or belonged to the category of general and other answers (e.g., “i just like it"). upon the question what students did not like about their achievement group, fewer than half of the students mentioned a specific aspect that they did not like (see table a9). in fact, many students replied that they liked everything, although this was mentioned somewhat less frequently by students from low achievement groups. most of the specific answers were related to the higher-order themes of independent work and its difficulty and the achievement group. negative aspects mentioned by students placed in low achievement groups included too easy tasks, boredom and wanting to be in a higher group. negative aspects mentioned by 71 | f l r students placed in high achievement groups included needing to work hard or fast, distraction (by other students but also by the teacher explaining to another group), difficult work, and stress. for students placed in average achievement groups, some of the answers resembled those of students in low achievement groups (e.g., too easy, wanting to be in a higher group) whereas others were similar to those of students in high achievement groups (e.g., too difficult, needing to work too hard). relatively few students made negative comments related to the higher-order themes of learning and understanding (e.g., “i don’t learn so much and i want to get better”) and instruction and the teacher (e.g., “you get additional explanation when you understand it already”). 3.2.2 students’ learning in their achievement group second, we asked students how much they learned in their achievement group. the likelihood ratio test indicated a significant effect of achievement group (χ2 (2) = 7.17, p =.028), indicating that students’ perceived amount of learning differed between achievement groups. students from high achievement groups (m = 4.40, sd = 0.96) perceived to learn more than students from average (m = 4.21, sd = 1.03) and low (m = 3.98, sd = 1.13) achievement groups. note, however, that all these means are relatively high on a fivepoint scale. students’ responses to the question why they learned much or little in that achievement group were frequently related to the higher-order theme of independent work and its difficulty (see table a10). this theme was most prominent for students from high achievement groups, who mentioned frequently that they learned much because of the higher difficulty level of the tasks: “you learn much because you also get more difficult sums”. when students from low and average achievement groups referred to this theme, their answers were more mixed: some mentioned an appropriate difficulty level as a reason for learning much, whereas others indicated that they did not learn so much because of an inappropriate difficulty level (too easy or too difficult). the question why students learned much or little also provoked relatively many comments related to the higher-order theme of learning and understanding. students from average and high achievement groups tended to explain why they learned or understood more (“you learn new goals every time”), whereas students placed in low achievement groups referred to learning and understanding both positively (“i learn much because i understand it better now”), and negatively (“because i don’t understand it at all”). the higher-order theme of instruction and the teacher was mentioned by students across achievement groups, mostly in a positive way (“because the teacher gives you more explanations and does more sums with you”). some students also provided answers related to the higher-order theme of the achievement group, which were similar to the answers described as reasons for liking or not liking the achievement group. again, a substantial proportion (about one-third) of the comments was classified as general or other (“i learn much because i learn much from it”). 3.2.3 students’ preference for an achievement group third, we asked students whether they would prefer to be in a different achievement group. the likelihood ratio test indicated a significant main effect of achievement group (χ2 (2) = 24.68, p <.001). as can be seen in table 2, about half of the students currently placed in a low achievement group would prefer to be in a different achievement group, compared to only 11% of students currently placed in a high achievement group. for students from low achievement groups, reasons for wanting to stay in the same achievement group included appropriate difficulty in the current group), and positive comments about the group members and interaction in the current group (see table a11). in contrast, other students placed in low achievement groups wanted to move to another group because they wanted more difficult tasks, enrichment tasks, or more challenge, because they thought they would learn more or get better at mathematics in another group, because of the members of the desired group or for the sake of being in a higher group. for students placed in high achievement groups, reasons for wanting to stay in the same group were mainly related to appreciating the current difficulty level of the tasks and specifically enrichment tasks (“enrichment tasks are fun, so i want to keep those”), although a few students wanted to move to a different group because they thought that the current material was too difficult. for students from average achievement groups, comments 72 | f l r related to the higher-order theme of tasks and their difficulty level were among the most frequent reasons (besides general comments) to want to stay in the same group, but also to switch achievement groups. table 2 students’ preference for being in another achievement group current achievement group preference for other group wish to stay in current group low (n = 52) 25 (48.1%) 27 (51.9%) average (n = 84) 30 (35.7%) 54 (64.3%) high (n = 97) 11 (11.3%) 86 (88.7%) total (n = 233) 66 (28.3%) 167 (71.7%) 3.2.4 students’ preference for working with or without achievement groups finally, we asked students whether they would prefer to work without achievement groups. the likelihood ratio test indicated no significant main effect of achievement group (χ2 (2) = 1.25, p = 0.535). across achievement groups, about 75% of students wanted to retain the achievement groups (see table 3). the most frequent reasons for wanting to retain achievement groups included general positive comments about the current grouping system, as well as between-student differences or appropriate difficulty in the current system (see table a12). some students explicitly described differences between students as an argument for differentiation: “well, if some students find it difficult and others find it hard and everybody does the same, i just don’t think it’s handy”. reasons for preferring to work without groups included the opportunity to learn more in a system without achievement groups and general negative comments about achievement groups. a few students explicitly mentioned equality, stating that they would like everybody to get the same tasks or to be equal. while all of these reasons for and against achievement grouping were mentioned by students of all achievement groups, the general tendency was that students from high achievement groups relatively frequently mentioned an appropriate difficulty level and the opportunity to learn more as reasons for retaining the grouping system. table 3 students’ preference for working with or without achievement groups achievement group without achievement groups with achievement groups low (n = 51) 15 (29.4%) 36 (70.6%) average (n = 84) 18 (21.4%) 66 (87.6%) high (n = 99) 22 (22.2%) 77 (77.8%) total (n = 234) 55 (23.5%) 179 (76.5%) 3.2.5 differences between grade levels in students’ perceptions of achievement grouping to explore potential differences between grade levels in students’ answers to the closed-ended questions, interactions with grade level were added to the analyses. none of these interactions were significant, indicating that the results were similar across grade levels. an extensive analysis of between-grade level differences in the open-ended answers is beyond the scope of this paper. we did explore the relative frequency of the answering categories across grade levels and found that students in higher grades tended to give relatively more specific answers (i.e., fewer general and other answers) than students in lower grades. in addition, relatively more of the answers of students in higher grades tended to be related to tasks and difficulty. in students’ explanations of why they would prefer to work with or without groups, most of the answers referring to either equality as an argument for no grouping or between-student differences as an argument for grouping were made by students from the highest grade. 73 | f l r 4. discussion while potential didactical and socioemotional advantages and disadvantages of differentiation and achievement grouping have been discussed in the literature, few studies have asked the opinion of students themselves. the current study extends the literature by exploring students’ perspective on differentiated activities and within-class achievement grouping, with attention for potential differences between students of diverse achievement levels. 4.1 students’ perceptions of differentiated activities our first research question was whether different students evaluate various types of mathematics activities differently, depending on their achievement level. as hypothesised, there were significant interactions between the type of activity and students’ achievement level for all three outcome variables: perceived frequency, liking and learning of the activities. in line with guidelines for differentiation and with previous studies in which teachers reported their use of differentiation strategies (prast et al., 2015; roy et al., 2013; van geel et al., 2018), low-achieving students perceived to receive extended instruction and less difficult tasks more frequently whereas high-achieving students worked at more difficult tasks more frequently. the infrequent occurrence of specific instruction for high-achieving students is not recommended, but also corresponds with previous findings (inspectorate of education, 2019; prast & hickendorff, in press). regarding students’ liking and learning of the various activities, we found that activities intended for lowachieving students such as less difficult tasks and extended instruction were rated more highly by lowachieving students, whereas more difficult tasks were rated more highly by high-achieving students. these are examples of perceived aptitude-treatment interactions (cronbach & snow, 1977; kalyuga, 2007). however, the following observations should be kept in mind. first, scores for general activities such as whole-class instruction were also high across achievement groups. second, students’ liking and learning from activities seemed to be related to students’ reported frequency of engaging in these activities. this might imply that students simply like, and perceive to learn from, activities to which they are used. nevertheless, if students’ experiences with activities would have been strongly negative, it seems unlikely that a higher frequency of an unpleasant activity would increase students’ liking of that activity. third, students generally reported to learn less from less difficult tasks, although this was less pronounced for lowachieving students than for high-achieving students. this is related to the issue of convergent versus divergent differentiation (blok, 2004). in the higher grades of primary school, the tasks in the lowest tier of many mathematics textbooks lead towards lower end-of-school learning goals than the tasks in the highest tier (expertgroep doorlopende leerlijnen, 2008). thus, it may be true that students in the higher grades learn less from less difficult tasks in the sense of covering less content (regardless of the degree of understanding of that content). finally, while these results are largely in line with our hypotheses based on the didactical perspective, students’ ratings of liking and learning from activities may also have been influenced by socioemotional factors. these are discussed in the following section. 4.2 students’ perceptions of achievement grouping our second research question was how students of diverse achievement levels evaluate their own achievement group and achievement grouping in general. our results provide partial support for the hypothesis that students’ perceptions of achievement grouping would differ between students placed in low, average and high achievement groups. generally, the average scores for liking and learning from one’s own achievement group were quite high. students’ liking of their own achievement group did not differ significantly across groups, but students’ perceived degree of learning did vary between achievement groups, with lower scores for students placed in low achievement groups. overall, about 70% of the students were satisfied with their achievement group placement, which is more than has been reported in previous studies about between-class grouping (boaler et al., 2000; hallam et al., 2004). however, this question revealed the most pronounced differences between achievement groups: around 50% of students placed in low achievement groups would prefer to be in a different group, compared to only 10% of students placed in high achievement groups. nevertheless, around 75% of students across achievement groups wanted to retain the grouping system (although this could also reflect a general desire for things to stay as they are; hallam et al. 74 | f l r (2004) also found that most students did not want to change anything about the grouping practices in their school). taken together, these results suggest quite positive attitudes towards grouping in general, but somewhat less positive experiences with placement in a low achievement group compared to a high achievement group. in addition to these quantitative ratings, we investigated which reasons students provided for their evaluations. as expected, students’ answers to the open-ended questions included socioemotional as well as didactical considerations, although didactical considerations seemed to be more prominent. students clearly evaluated the use of achievement grouping in relation to the use of differentiated activities. in line with the didactical perspective on differentiation, many students mentioned didactical advantages including the appropriate amount and difficulty level of independent work (mentioned by students of all achievement groups), challenge (mainly mentioned by students placed in high achievement groups) and the possibility to get additional instruction and to understand the material better (mainly mentioned by students placed in low and average achievement groups). however, some students also mentioned didactical disadvantages. some students from low achievement groups perceived the work as too easy, did not appreciate additional instruction, wanted more challenge or thought that they would learn more in a higher group. in contrast, some students from high achievement groups felt that the material was too difficult or that they needed to work too fast, which was sometimes stressful. some students from average achievement groups made similar comments in both directions (too difficult or too easy). ideally, differentiation should ensure that the tasks and instruction are appropriately challenging for the students in each achievement group (prast et al., 2015). students’ answers indicate that, while many students perceived the difficulty level as appropriate, other students did not. moreover, challenge seemed to be viewed by many students as something belonging exclusively to the high achievement group. enrichment tasks were highly valued by students from high achievement groups, but some students from low and average achieving groups also expressed the desire to work on enrichment tasks. compared to previous studies on between-class achievement grouping (boaler et al., 2000; hallam et al., 2004), the students in our sample generally seemed to be more positive about the didactical advantages of achievement grouping, but some of the perceived disadvantages resembled those mentioned in previous studies (e.g., a lack of challenge in low achievement groups, needing to work too fast in high achievement groups). as expected, students’ answers also included socioemotional considerations, but these were not always related to the socioemotional perspective on achievement grouping as described in the introduction (i.e., based on social comparisons of achievement level). many students made comments about their achievement group that did not seem to be related directly to the achievement level of that group, such as (dis-)liking the members of the group or having positive or negative interactions within the group. this importance of peer interactions in general in students’ perceptions of schooling echoes previous findings by hargreaves et al. (2021). partly in line with previous studies (eder, 1983; marks, 2013; mcgillicuddy & devine, 2020), we found some indications for social comparisons based on achievement group placement. a few students mentioned that they liked to know their own level, or the level of the other students. students occasionally mentioned the fact that their group was low (also: “bad”) or high (also: “the best” or “smart”) as a negative or positive aspect of being in that group. while such comments were relatively infrequent, they support the idea that within-class achievement groups may strengthen social comparison processes by making students more aware of whose achievement is low, average or high compared to the class average (i.e., labelling effects; campbell, 2021). based on students’ spontaneous answers to our open-ended questions, explicit teasing or stigmatisation based on achievement group placement did not seem to play a major role in the current sample. of course, these findings do not exclude the possibility of implicit stigmatisation or social status associated with achievement group placement (see marks, 2013; mcgillicuddy & devine, 2020; van den bergh, 2018). taken together, students’ answers provide some support for potential socioemotional sideeffects of within-class achievement grouping, including social comparisons of achievement level, but do not indicate pervasive negative social effects of placement in a low achievement group or of achievement grouping in general. this might be partly explained by differences between countries in the way in which 75 | f l r achievement grouping is implemented. since the achievement groups in this study were typically used only part of the time (besides whole-class activities) and were relatively flexible, it could be that this reduced the potential negative socioemotional effects of achievement grouping (education endowment foundation, 2018). if grouping arrangements are sufficiently flexible to respond to students’ current achievement level and corresponding educational needs, as recommended (prast et al., 2015), students might evaluate them more positively than when students perceive to be stuck in a (low) achievement group. note, however, that the degree of flexibility of the grouping arrangements differed substantially between teachers in the current sample (see section 2.1). 4.3 differences between grade levels we also explored whether students’ views on differentiation and achievement grouping differed between grade levels. we emphasise that our findings in grade 1 should be viewed as exploratory, given the small sample size. by and large, there seemed to be a trend towards more pronounced opinions in the higher grades. students’ reported liking and learning from activities were more strongly related to student achievement level in higher grades. the quantitative ratings of achievement grouping were similar across grades, but students in higher grades gave relatively more specific answers to the open-ended questions. this may have several reasons. first, due to maturation, older students may have been better able to express their opinions. this is likely to have affected the open-ended questions more strongly than the closed-ended questions. second, older students may have developed more pronounced opinions about differentiation due to more experience with differentiation. third, through socialisation, older students may have endorsed the values of an educational system which assumes that lessons should be adapted to between-student differences in achievement level (raveaud, 2005). in future research, it would be interesting to follow a group of students longitudinally from grade 1 onwards to examine how students’ views on differentiation and achievement grouping develop over time. 4.4 limitations, conclusions and implications students’ perceptions of differentiation and achievement grouping were central to this study. we do not claim that students always know what is best for them in terms of learning outcomes (see kirschner & van merriënboer, 2013). nevertheless, we feel that it is important to consider what students themselves think about the degree to which differentiation is meeting their educational needs, even if it would only be to explain teachers’ choices better to students if students would fundamentally disagree with the approach taken. this did not seem to be the case: by and large, students had quite positive attitudes towards differentiation. with a self-report questionnaire, there is always a risk of socially desirable answers. however, our general impression is that students responded quite frankly, maybe also due to their young age (e.g., “i don’t understand a shit of it”). our method of data collection offered several advantages. by asking students to quantitively rate specific mathematics activities (and relating this to students’ achievement level in the analyses), we could investigate the complex construct of differentiation in a way that was easy to understand for students as well as relatively quick and standardised. this enabled data collection on a larger scale and, therefore, provided more opportunities to quantify and generalise differences (or similarities) between students of diverse achievement levels than is typically possible in small-scale qualitative studies. by combining the quantitative ratings of activities and achievement grouping with open-ended questions, we also gained insights into students’ reasons behind their quantitative evaluations, although small-scale qualitative studies can study this in more depth. this study did not examine whether the way in which differentiation and achievement grouping were implemented affected students’ perceptions. for example, the quality of differentiation i.e., the degree to which adaptations are carefully matched to students’ educational needs may also affect students’ perceptions. in the current study, students’ achievement group placement did not always correspond with their achievement on the standardised achievement test administered at the end of the previous schoolyear. 76 | f l r while teachers may have created the achievement groups based on other and more recent achievement information (e.g., curriculum-based tests, daily mathematics work), it might also mean that some students were placed in an achievement group that was not appropriate for their achievement level. this might partially explain why some students perceived the work in their achievement group as too easy or too difficult. in addition, the flexibility of the grouping arrangements might affect students’ perceptions (education endowment foundation, 2018). finally, the use of multigrade classes may have implications for teachers’ practices (for example, if teachers create achievement groups within each grade level, they will need to divide their attention over more achievement groups than a teacher teaching a single-grade class) which may in turn affect students’ perceptions of differentiation and achievement grouping. these would be interesting issues to explore in future research. this study focused on differentiation in a specific context, namely primary mathematics education in the netherlands. due to substantial differences between countries and content areas in the traditions and practices of differentiation and achievement grouping, these results may not be directly generalisable to other countries or other content areas. in the netherlands, for example, differentiation practices seem to be somewhat similar for reading, in the sense that teachers might offer more or less difficult reading materials or more or less instruction to different subgroups of students, based on their achievement level (although the way in which instruction and practice are adapted to the needs of low-achieving or high-achieving students may be qualitatively different in reading compared to mathematics, because of the different content area and belonging didactical models). however, for other subjects such as science, teachers seem to use different approaches to differentiation (slim et al., 2022). compared to other countries, the achievement grouping practices in the netherlands may be relatively flexible, which might partially explain the relatively positive evaluations compared to previous studies (e.g., eder, 1983; gripton, 2020; marks, 2013; mcgillicuddy & devine, 2020). future research could not only study the generalisability of the findings across contexts, but could also use naturally occurring differences in the implementation of differentiation and achievement grouping between various countries and domains to investigate how these differences in implementation might affect students’ perceptions. from the current findings in the context of primary mathematics education in the netherlands, we conclude that students had largely positive attitudes about differentiation and achievement grouping. students appreciated it when the amount and difficulty of tasks and instruction were adapted to their current achievement level, and did not like either too easy or too difficult work. while the majority of students across achievement groups wanted to retain the achievement grouping system and reported high liking of their achievement group, students placed in low achievement groups reported to learn less from their group and more often had the desire to be in a different group. didactical considerations such as wanting to learn more or wanting to be challenged seemed to be more prominent in students’ reasoning than socioemotional considerations such as the social status associated with an achievement group. our findings have the following implications. many students displayed positive attitudes to learning: they liked to learn more, did not like to be distracted by other students, and wanted to be challenged. many students specifically mentioned that they liked or wanted to have enrichment tasks. to retain this positive attitude towards learning, we think that it would be helpful to encourage rather than discourage students who want to try more difficult tasks. this relates to the topic of self-regulation, which has been receiving increasing attention in the differentiation literature: ideally, students should be able to co-decide (in collaboration with the teacher) whether they need additional instruction and which tasks they should do (van geel et al., 2018). this might reduce negative experiences with work that is perceived as too hard or too easy. future research could also examine ways in which the perceived benefits of adapting education to students’ achievement level can be retained, while reducing potential socioemotional or didactical disadvantages of placement in a low achievement group. this might include ways to adapt instruction and practice to students’ educational needs more flexibly based on students’ current understanding of specific mathematical content (perhaps, also using adaptive educational software to assist the teacher in making these 77 | f l r choices), as well as variation of grouping arrangements (for example by using heterogeneous groups in situations where students of diverse achievement levels can learn from each other). in such future research, the perspective of students themselves should not be overlooked. keypoints  students’ voices have not often been heard in the debate on differentiation and achievement grouping  in this mixed-methods study, primary school students (n = 428) evaluated differentiated activities and within-class achievement grouping  didactical and socioemotional aspects, and potential differences between students of diverse achievement levels were considered  students had largely positive views on differentiated activities and achievement grouping  but students also perceived disadvantages of placement in a low achievement group acknowledgments we thank all students and teachers who were involved in this study, as well as the reviewers of this manuscript, for their valuable contributions. references arroyo, i., woolf, b. p., burelson, w., muldner, k., rai, d., & tai, m. (2014). a multimedia adaptive tutoring system for mathematics that addresses cognition, metacognition and affect. international journal of artificial intelligence in education, 24(4), 387–426. https://doi.org/10.1007/s40593-0140023-y bates, d., maechler, m., bolker, b., & walker, s. (2015). fitting linear mixed-effects models using lme4. journal of statistical software, 67(1), 1–48. https://doi.org/10.18637/jss.v067.i01. blok, h. (2004). adaptief onderwijs: betekenis en effectiviteit [adaptive education: meaning and effectivity]. pedagogische studiën, 81(1), 5–27. boaler, j., wiliam, d., & brown, m. (2000). students’ experiences of ability grouping disaffection, polarisation and the construction of failure. british educational research journal, 26(5), 631–648. https://doi.org/10.1080/713651583 campbell, t. (2021). in-class ‘ability’-grouping, teacher judgements and children’s mathematics selfconcept: evidence from primary-aged girls and boys in the uk millennium cohort study. cambridge journal of education. https://doi.org/10.1080/0305764x.2021.1877619 cronbach, l. j., & snow, r. e. (1977). aptitudes and instructional methods: a handbook for research on interactions. irvington. csikszentmihalyi, m. (1990). flow: the psychology of optimal experience. harper perennial. deunk, m. i., smale-jacobse, a. e., de boer, h., doolaard, s., & bosker, r. j. (2018). effective 78 | f l r differentiation practices:a systematic review and meta-analysis of studies on the cognitive effects of differentiation practices in primary education. educational research review, 24, 31–54. https://doi.org/10.1016/j.edurev.2018.02.002 eder, d. (1983). ability grouping and students’ academic self-concepts: a case study. the elementary school journal, 84, 149–161. https://doi.org/10.2307/1001307 education endowment foundation. (2018). sutton trust-education endowment foundation teaching and learning toolkit. education endowment foundation. https://educationendowmentfoundation.org.uk/resources/teaching-learning-toolkit expertgroep doorlopende leerlijnen. (2008). over de drempels met rekenen: consolideren, onderhouden, gebruiken en verdiepen. expertgroep doorlopende leerlijnen. faber, j. m., & visscher, a. j. (2016). de effecten van snappet: effecten van een adaptief onderwijsplatform op leerresultaten en motivatie van leerlingen. [the effects of snappet: effects of an adaptive educational platform on student achievement and motivation]. universiteit twente. francis, b., connolly, p., archer, l., hodgen, j., mazenod, a., pepper, d., sloan, s., taylor, b., tereshchenko, a., & travers, m. c. (2017). attainment grouping as self-fulfilling prophesy? a mixed methods exploration of self confidence and set level among year 7 students. international journal of educational research, 86, 96–108. https://doi.org/10.1016/j.ijer.2017.09.001 gripton, c. (2020). children’s lived experiences of ‘ability’ in the key stage one classroom: life on the ‘tricky table.’ cambridge journal of education, 50(5), 559–578. https://doi.org/10.1080/0305764x.2020.1745149 hallam, s., ireson, j., & davies, j. (2004). primary pupils’ experiences of different types of grouping in school. british educational research journal, 30(4), 515–533. https://doi.org/10.1080/0141192042000237211 hargreaves, e., buchanan, d., & quick, l. (2021). “look at them! they all have friends and not me”: the role of peer relationships in schooling from the perspective of primary children designated as “lowerattaining.” https://doi.org/10.1080/00131911.2021.1882942. https://doi.org/10.1080/00131911.2021.1882942 hart, s. (1992). differentiation. part of the problem or part of the solution? the curriculum journal, 3(2), 131–142. https://doi.org/10.1080/0958517920030203 inspectorate of education. (2019). reken-en wiskundeonderwijs aan potentieel hoogpresterende leerlingen [mathematics education for potentially high-achieving students]. inspectie van het onderwijs. janssen, j., scheltens, f., & kraemer, j. m. (2005). rekenen-wiskunde groep 3-8: handleidingen [mathematics test grade 1-6: manuals]. cito. janssen, j., verhelst, n., engelen, r., & scheltens, f. (2005). wetenschappelijke verantwoording van de toetsen lovs rekenen-wiskunde voor groep 3 tot en met 8 [scientific justification of the mathematics tests for grade 1 through 6]. cito. jerrim, j. (2021). the association between within-class grouping and children’s achievement in mathematics during year 2, year 5 and year 9. school choices report. education endowment foundation. https://educationendowmentfoundation.org.uk/public/files/within_class_grouping_report_-_final.pdf kalyuga, s. (2007). expertise reversal effect and its implications for learner-tailored instruction. educational 79 | f l r psychology review, 19(4), 509–539. https://doi.org/10.1007/s10648-007-9054-3 kirschner, p. a., sweller, j., & clark, r. e. (2006). why minimal guidance during instruction does not work: an analysis of the failure of constructivist, discovery, problem-based, experiential, and inquirybased teaching. educational psychologist, 41(2), 75–86. https://doi.org/10.1207/s15326985ep4102_1 kirschner, p. a., & van merriënboer, j. g. (2013). do learners really know best? urban legends in education. educational psychologist, 48, 169–183. https://doi.org/10.1080/00461520.2013.804395 koerhuis, i. (2010). rekenen voor kleuters [mathematics for kindergarteners]. cito. koerhuis, i. ., & keuning, j. (2011). wetenschappelijke verantwoording van de toetsen rekenen voor kleuters voor groep 1 en 2 [scientific justification of the tests mathematics for kindergarteners]. cito. lenth, r. (2020). emmeans: estimated marginal means, aka least-squares means. r package version 1.5 2.1. https://cran.r-project.org/package=emmeans linneberg, m. s., & korsgaard, s. (2019). coding qualitative data: a synthesis guiding the novice. qualitative research journal, 19(3), 259–270. https://doi.org/10.1108/qrj-12-2018-0012/full/xml marks, r. (2013). “the blue table means you don’t have a clue”: the persistence of fixed-ability thinking and practices in primary mathematics in english schools. forum, 55(1), 31. https://doi.org/10.2304/forum.2013.55.1.31 marsh, h. w. (1984). self-concept, social comparison, and ability grouping: a reply to kulik and kulik. in source: american educational research journal (vol. 21, issue 4). winter. marsh, h. w. (1987). the big-fish-little-pond effect on academic self-concept [article]. journal of educational psychology, 79(3), 280–295. https://doi.org/10.1037//0022-0663.79.3.280 mcgillicuddy, d., & devine, d. (2020). ‘you feel ashamed that you are not in the higher group’— children’s psychosocial response to ability grouping in primary school. british educational research journal, 46(3), 553–573. https://doi.org/10.1002/berj.3595 murray, t., & arroyo, i. (2002). toward measuring and maintaining the zone of proximal development in adaptive instructional systems. lncs, 2363, 749–758. park, d., tsukayama, e., gunderson, e. a., levine, s. c., & beilock, s. l. (2016). young children’s motivational frameworks and math achievement: relation to teacher-reported instructional practices, but not teacher theory of intelligence. journal of educational psychology, 108(3), 300–313. https://doi.org/10.1037/edu0000064 prast, e.j. & hickendorff, m. (in press). how do dutch teachers implement differentiation in primary mathematics education? in: r. maulana, m. helms-lorenz, & r.m. klassen (eds.). effective teaching around the world: theoretical, empirical, methodological and practical insights. springer. prast, e. j., van de weijer-bergsma, e., kroesbergen, e. h., & van luit, j. e. h. (2015). readiness-based differentiation in primary school mathematics: expert recommendations and teacher self-assessment. frontline learning research, 3(2), 90–116. https://doi.org/10.14768/flr.v3i2.163 prast, e. j., van de weijer-bergsma, e., kroesbergen, e. h., & van luit, j. e. h. (2018). differentiated instruction in primary mathematics: effects of teacher professional development on student achievement. learning and instruction, 54. https://doi.org/10.1016/j.learninstruc.2018.01.009 80 | f l r raveaud, m. (2005). hares, tortoises and the social construction of the pupil: differentiated learning in french and english primary schools. british educational research journal, 31(4), 459–479. https://doi.org/10.1080/01411920500148697 roy, a., guay, f., & valois, p. (2013). teaching to address diverse learning needs: development and validation of a differentiated instruction scale. international journal of inclusive education, 17(11), 1186–1204. https://doi.org/10.1080/13603116.2012.743604 rubie-davies, c. m. (2014). becoming a high expectation teacher: raising the bar. routledge. singer, j. d., & willett, j. b. (2003). applied longitudinal data analysis: modeling change and event occurrence. oxford university press. slim, t., van schaik, j., hotze, a. & raijmakers, m. (2022, july). differentiatie in het wetenschap & technologie-onderwijs: overtuiging en praktijk van aankomendeen expertleerkrachten. [differentiation in science & technology education: attitudes and practices of pre-service teachers and expert teachers]. poster presented at the onderwijs research dagen [educational research days], hasselt, belgium. tereshchenko, a., francis, b., archer, l., hodgen, j., mazenod, a., taylor, b., pepper, d., & travers, m. c. (2019). learners’ attitudes to mixed-attainment grouping: examining the views of students of high, middle and low attainment. research papers in education, 34(4). https://doi.org/10.1080/02671522.2018.1452962 tieso, c. l. (2003). ability grouping is not just tracking anymore. roeper review, 26(1), 29–36. tomlinson, c. a., brighton, c., hertberg, h., callahan, c. m., moon, t. r., brimijoin, k., conover, l. a., & reynolds, t. (2003). differentiating instruction in response to student readiness, interest, and learning profile in academically diverse classrooms: a review of literature. journal for the education of the gifted, 27(2–3), 119–145. https://doi.org/10.1177/016235320302700203 van den bergh, l. (2018). waarderen van diversiteit in het onderwijs [valuing diversity in education]. fontys opleidingscentrum speciale onderwijszorg. van geel, m., keuning, t., frèrejean, j., dolmans, d., van merriënboer, j., & visscher, a. j. (2018). capturing the complexity of differentiated instruction. school effectiveness and school improvement, 30(1), 51–67. https://doi.org/10.1080/09243453.2018.1539013 van geel, m., keuning, t., visscher, a. j., & fox, j. p. (2016). assessing the effects of a school-wide data-based decision-making intervention on student achievement growth in primary schools. american educational research journal, 53(2), 360–394. https://doi.org/10.3102/0002831216637346 vaughn, s., schumm, j. s., klingner, j., & saumell, l. (1995). students’ views of instructional practices: implications for inclusion. learning disability quarterly, 18(3), 236–248. https://doi.org/10.2307/1511045 vaughn, s., schumm, j. s., niarhos, f. j., & daugherty, t. (1993). what do students think when teachers make adaptations? teaching and teacher education, 9(1), 107–118. https://doi.org/10.1016/0742051x(93)90018-c vaughn, s., schumm, j. s., niarhos, f. j., & gordon, j. (1993). students’ perceptions of two hypothetical teachers’ instructional adaptations for low achievers. the elementary school journal, 94(1), 87–102. 81 | f l r visscher, a. j. (2015). over de zin van opbrengstgericht(er) werken in het onderwijs [about the value of (more) data-based decision making in education]. gion. vogt, f., & rogalla, m. (2009). developing adaptive teaching competency through coaching. teaching and teacher education, 25(8), 1051–1060. https://doi.org/10.1016/j.tate.2009.04.002 wickham, h. (2016). ggplot2: elegant graphics for data analysis. springer verlag new york. supplementary materials overview of appendices: appendix 1: correspondence between achievement test scores and achievement group placement appendix 2: additional information about the analyses in part 1 appendix 3: additional information about the qualitative analyses in part 2 appendix 1: correspondence between achievement test scores and achievement group placement table a1 correspondence between achievement test scores and achievement group placement achievement test score low achievement group average achievement group high achievement group i (highest) 2 14 59 ii 3 22 21 iii 13 30 4 iv 17 10 9 v (lowest) 17 4 6 table a1 displays the (imperfect) correspondence between students’ achievement group placement and their scores on the achievement test, is based on 231 students who had been placed in a single within-class achievement group during the past three weeks and for whom achievement test scores were available. note that the achievement test scores were collected at the end of the previous school year. reasons for noncorrespondence may include the assignment to achievement groups based on other measures and more recent sources of information (e.g., curriculum-based tests, students’ responses during mathematics lessons). 82 | f l r appendix 2: additional information about the analyses in part 1 descriptive statistics of the outcome variables are provided in tables a2 (frequency), a3 (liking), and a4 (learning). table a2 means and standard deviations of student-reported frequency of activities split by achievement level achievement level v (lowest) iv iii ii i (highest) overall activity m sd m sd m sd m sd m sd m sd whole-class instruction 4.22 1.01 3.96 0.96 4.07 0.99 3.89 0.99 3.83 1.16 3.96 1.05 working independently 4.02 0.94 3.80 1.21 3.95 1.11 4.23 0.93 4.28 0.86 4.09 1.01 working together 3.66 1.22 3.18 1.11 3.05 1.06 3.16 1.00 3.20 1.21 3.22 1.14 easier tasks 3.41 1.41 2.65 1.39 2.70 1.43 2.14 1.44 2.25 1.56 2.54 1.51 subgroup extended instruction 3.20 1.36 3.05 1.39 2.63 1.40 1.98 1.09 1.86 1.10 2.42 1.36 individual extended instruction 3.10 1.32 2.64 1.21 2.61 1.35 2.30 1.24 2.04 1.31 2.44 1.33 enrichment tasks 2.58 1.65 2.64 1.60 2.85 1.62 3.62 1.57 4.09 1.30 3.31 1.64 subgroup enrichment instruction 2.80 1.56 2.22 1.34 2.15 1.44 2.26 1.32 2.45 1.34 2.36 1.39 individual enrichment instruction 2.62 1.61 2.18 1.45 2.19 1.41 2.37 1.37 2.42 1.37 2.35 1.42 overall 3.30 1.46 2.93 1.43 2.91 1.47 2.88 1.47 2.93 1.53 note. students rated the frequency of activities on a five-point scale (5 = highest). 83 | f l r table a3 means and standard deviations of student-reported liking of activities split by achievement level achievement level v (lowest) iv iii ii i (highest) overall activity m sd m sd m sd m sd m sd m sd whole-class instruction 3.80 1.19 3.58 1.27 3.27 1.16 3.30 1.15 3.22 1.27 3.39 1.23 working independently 3.44 1.43 3.45 1.58 3.53 1.33 3.89 1.32 3.78 1.27 3.65 1.38 working together 4.20 1.17 4.05 1.37 4.30 1.01 3.98 1.17 3.89 1.37 4.05 1.25 easier tasks 3.80 1.38 3.25 1.57 3.10 1.54 2.46 1.66 2.38 1.58 2.88 1.63 subgroup extended instruction 3.88 1.36 3.27 1.46 3.29 1.50 2.61 1.50 2.35 1.27 2.95 1.50 individual extended instruction 3.50 1.55 2.84 1.51 2.58 1.42 2.45 1.33 2.34 1.30 2.65 1.44 enrichment tasks 2.52 1.57 3.20 1.59 3.10 1.64 4.11 1.30 4.15 1.25 3.56 1.56 subgroup enrichment instruction 3.15 1.67 3.02 1.50 2.83 1.53 3.16 1.33 3.05 1.35 3.04 1.45 individual enrichment instruction 2.88 1.62 2.89 1.52 2.41 1.43 2.88 1.36 2.74 1.40 2.75 1.45 overall 3.47 1.52 3.29 1.52 3.16 1.49 3.20 1.48 3.10 1.49 note. students rated their liking of activities on a five-point scale (5 = highest). 84 | f l r table a4 means and standard deviations of student-reported learning from activities split by achievement level achievement level v (lowest) iv iii ii i (highest) overall activity m sd m sd m sd m sd m sd m sd whole-class instruction 4.15 1.06 3.93 1.26 3.48 1.30 3.75 1.18 3.70 1.11 3.77 1.19 working independently 3.85 1.20 3.84 1.23 4.18 0.91 4.23 0.87 4.10 0.88 4.06 1.00 working together 4.10 1.09 3.80 1.21 3.92 1.06 3.74 1.19 3.82 1.10 3.86 1.13 easier tasks 3.27 1.45 2.89 1.50 2.69 1.51 2.09 1.34 2.20 1.43 2.54 1.50 subgroup extended instruction 4.17 1.18 4.09 1.14 3.72 1.20 3.33 1.50 3.43 1.37 3.69 1.33 individual extended instruction 3.95 1.20 4.00 1.22 3.71 1.30 3.51 1.35 3.57 1.37 3.71 1.31 enrichment tasks 3.33 1.61 4.24 1.04 4.12 1.14 4.52 1.04 4.57 0.66 4.25 1.12 subgroup enrichment instruction 3.74 1.52 3.69 1.36 3.73 1.28 3.72 1.24 3.82 1.28 3.75 1.31 individual enrichment instruction 3.65 1.53 3.71 1.33 3.49 1.32 3.68 1.31 3.65 1.31 3.64 1.34 overall 3.80 1.35 3.80 1.30 3.67 1.29 3.62 1.38 3.65 1.33 note. students rated their learning from activities on a five-point scale (5 = highest). analytical strategy data were analysed in r with multilevel models to take into account the nested data structure (e.g., scores of students within a class/school being correlated). for all analyses, in order to take the dependencies in the data due to the nested data structure into account, we used random intercepts at the various levels (i.e., student, class and school). as a four-level random intercept model is already quite complex for our data, we decided not to include random slopes in the model. the models were fitted with the package lme4 (bates et al., 2015), post-hoc analyses were conducted with the package emmeans (lenth, 2020), and figures were plotted with ggplot2 (wickham, 2016). effect coding was used. effect coding differs from dummy 85 | f l r coding in the sense that weights other than 0 and 1 (i.e., standard dummy coding) can be assigned to the various levels of a categorical variable, which facilitates the interpretation of the fixed effects. first, we estimated an empty model – also called unconditional means model (singer & willett, 2003) to investigate the amount of variance at the various levels. second, main effects of activity and achievement level were added to the model. third, the interaction between activity and achievement level was added. to evaluate the significance of main effects and interaction effects, likelihood ratio tests (lrt) were used to compare the fit of the full model (i.e., including the effect of interest) to the fit of a reduced model without that main effect or interaction. empty models the empty models indicated that by far the most variance was at the level of the various activities rated by the same students (i.e., repeated measures, see table a5). the variance at the student level was somewhat larger for the degree to which students perceived to learn from activities (13.1%) than for the other outcome variables. the amount of variance at the class and school level was quite small (0.5 – 2.2%), but these levels were retained in the analyses anyway to correct for any clustering effects at these levels. table a5 distribution of the outcome variance across the different levels in the data level frequency liking learning activity (within-student) 92.7% 90.6% 84.8% student 4.0% 5.8% 13.1% class 1.4% 2.2% 1.6% school 1.9% 1.4% 0.5% likelihood ratio tests comparing model fit table a6 provides an overview of the results of the likelihood ratio tests comparing model fit. as can be seen in the table, the model including the interaction between activity and achievement level had the best fit compared to a reduced model for all outcome variables. table a6 outcomes of likelihood ratio tests comparing model fit frequency liking learning model comparison χ2 df p χ2 df p χ2 df p main effect of activity added to empty model 621.79 8 <.001 248.03 8 <.001 326.19 8 <.001 main effect of achievement added to previous model 4.55 1 0.03 6.36 1 0.01 2.05 1 0.152 interaction effect added to previous model 182.36 8 <.001 161.75 8 <.001 95.74 8 <.001 86 | f l r post-hoc tests for the interaction effects table a7 provides an overview of the post-hoc tests for the interaction effect. a significant effect means that students’ achievement level predicts their ratings for that activity. since a score of 1 on the achievement tests reflects the highest achievement and 5 the lowest, a positive value for the effect indicates a negative effect of achievement level (i.e., activity ratings are higher for low-achieving students) whereas a negative value indicates a positive effect of achievement level (i.e., activity ratings are higher for high-achieving students). table a7 post-hoc tests of the interaction effect: the effect of achievement on students’ reported frequency, liking and learning for each activity separately activity and outcome variable 𝛽 se df t p whole-class instruction frequency -0.07 0.05 2491 -1.50 0.13 liking -0.12 0.06 2514 -2.22 0.03 learning -0.08 0.05 2126 -1.62 0.11 working independently frequency 0.11 0.05 2488 2.16 0.03 liking 0.12 0.06 2511 2.15 0.03 learning 0.08 0.05 2126 1.64 0.10 working together frequency -0.06 0.05 2491 -1.23 0.22 liking -0.07 0.06 2511 -1.22 0.22 learning -0.04 0.05 2131 -0.86 0.39 easier tasks frequency -0.25 0.05 2488 -4.88 <.001 liking -0.34 0.06 2512 6.11 <.001 learning -0.27 0.05 2142 5.08 <.001 subgroup extended instruction frequency -0.37 0.05 2488 -7.23 <.001 liking -0.36 0.06 2515 -6.43 <.001 learning -0.11 0.05 2142 5.08 <.001 individual extended instruction frequency -0.23 0.05 2496 -4.63 <.001 liking -0.23 0.06 2522 -4.16 <.001 learning -0.11 0.05 2148 -2.31 <.001 enrichment tasks frequency 0.43 0.05 2496 8.40 <.001 liking 0.41 0.06 2521 7.37 <.001 learning 0.25 0.05 2142 5.08 <.001 subgroup enrichment instruction frequency -0.02 0.05 2495 -0.44 0.66 liking 0.01 0.06 2517 0.27 0.79 learning 0.03 0.05 2156 0.61 0.54 individual enrichment instruction frequency 0.01 0.05 2494 0.15 0.88 liking -0.01 0.06 2520 -0.15 0.88 learning 0.01 0.05 2137 0.15 0.88 87 | f l r appendix 3: additional information about the qualitative analyses in part 2 tables a8 through a11 provide an overview of the answering categories (lower-order codes organised by higher-order themes) for each question, as well as the number of times these categories were mentioned by students placed in low, average and high achievement groups. table a8 students’ responses to the question: what do you like about being in your achievement group? frequency of comments a answering category example low (n = 59) average (n = 97) high (n = 110) total (n = 266) answers about work / difficulty 17 (28.8%) 32 (33.0%) 54 (49.1%) 103 (38.7%) appropriate difficulty you don’t need to do too difficult or too easy stuff 8 (13.6%) 15 (15.5%) 29 (26.4%) 52 (19.6%) challenge because if i have this work, it’s a challenge 1 (1.7%) 2 (2.1%) 13 (11.8%) 16 (6.0%) working more / doing many tasks so you can start to work immediately 3 (5.1%) 6 (6.2%) 4 (3.6%) 13 (4.9%) tasks / activities are fun you get nice tasks 1 (1.7%) 1 (1.0%) 7 (6.4%) 9 (3.4%) fewer tasks and/or finish earlier i need to do less and then i finish earlier 4 (6.8%) 3 (3.1%) 1 (0.9%) 8 (3.0%) my level it matches my level 0 (0.0%) 5 (5.2%) 0 (0.0%) 5 (1.9%) answers about group 17 (28.8%) 15 (15.5%) 23 (20.9%) 55 (20.7%) positive about group members / dynamics the children in this group are kind 8 (13.6%) 7 (7.2%) 13 (11.8%) 28 (10.5%) helping each other / working together because if you don’t know something the other children can help you 5 (8.5%) 4 (4.1%) 5 (4.6%) 14 (5.3%) concentrate / no distraction because it’s a quiet group and i can concentrate better 4 (6.8%) 2 (2.1%) 2 (1.8%) 8 (3.0%) group rank: being in a high(er) / smart(er) group then i think sometimes that i am a bit smarter together with the other children and that gives me confidence 0 (0.0%) 2 (2.1%) 3 (2.7%) 5 (1.9%) answers about learning 7 (11.9%) 12 (12.4%) 11 (10.0%) 30 (11.3%) learn more / faster / new things then you learn more of it 4 (6.8%) 8 (8.3%) 9 (8.2%) 21 (7.9%) understanding better then you understand better 2 (3.4%) 3 (3.1%) 0 (0.0%) 5 (1.9%) getting smarter / better at math / higher level then you get super smart 1 (1.5%) 1 (1.0%) 2 (1.8%) 4 (1.5%) answers about instruction / teacher 5 (8.5%) 11 (11.3%) 8 (7.3%) 24 (9.0%) (more) explanation / that you get more help 4 7 3 14 88 | f l r instruction / help (6.8%) (7.2%) (2.7%) (5.3%) positive about teacher the teacher thinks of a nice way 0 (0.0%) 2 (2.1%) 4 (3.6%) 6 (2.3%) no (additional) instruction / explanation you don’t need to listen to another explanation 1 (1.7%) 2 (2.1%) 1 (0.9%) 4 (1.5%) general and other answers 13 (22.0%) 27 (27.8%) 14 (12.7%) 54 (20.3%) unspecific / don’t know / other i just like it 10 (17.0%) 17 (17.5%) 9 (8.2%) 36 (13.5%) nothing is nice / negative comment i don’t like it at all 3 (5.1%) 7 (7.2%) 3 (2.7%) 13 (4.9%) everything is nice everything about this group is nice 0 (0.0%) 3 (3.1%) 2 (1.8%) 5 (1.9%) a number and percentage of comments within that achievement group. the total number of comments exceeds the number of students since some answers belonged to two categories. table a9 students’ responses to the question: what don't you like about being in your achievement group? frequency of comments a answering category example low (n = 55) average (n = 89) high (n = 105) total (n = 249) answers about work / difficulty 9 (16.4%) 15 (16.9%) 19 (18.1%) 43 (17.3%) needing to work much / longer / fast you need to finish quickly 1 (1.8%) 3 (3.4%) 11 (10.5%) 15 (6.0%) inappropriate difficulty: too hard sometimes it’s difficult 1 (1.8%) 5 (5.6%) 6 (5.7%) 12 (4.8%) boring / takes a long time because it takes a super long time. i find it super boring. 2 (3.6%) 4 (4.5%) 1 (1.0%) 7 (2.8%) inappropriate difficulty: too easy because it’s too easy now 4 (7.3%) 1 (1.1%) 1 (1.0%) 6 (2.4%) want more challenge because i like to be challenged 1 (1.8%) 2 (2.3%) 0 (0.0%) 3 (1.2%) answers about group 9 (16.4%) 8 (9.0%) 20 (19.1%) 37 (14.9%) distraction i am distracted when the others talk 4 (7.3%) 3 (3.4%) 12 (11.4%) 19 (7.6%) negative about group members / dynamics we quarrel sometimes 2 (3.6%) 2 (2.3%) 4 (3.8%) 8 (3.2%) group rank: want to be in higher group / not nice to be in low group of course, i would rather be in the plus-group [= highest group] 3 (5.5%) 3 (3.4%) 0 (0.0%) 6 (2.4%) stress about high achievement group the stress 0 (0.0%) 0 (0.0%) 4 (3.8%) 4 (1.6%) answers about learning 3 (5.4%) 5 (5.6%) 1 (1.0%) 9 (3.6%) not understanding when i don’t understand 1 (1.8%) 3 (3.4%) 1 (1.0%) 5 (2.0%) 89 | f l r learning less / not much i don’t learn so much and i want to get better 1 (1.8%) 1 (1.1%) 0 (0.0%) 2 (0.8%) being (called) bad / stupid it’s not nice to be so bad 1 (1.8%) 1 (1.1%) 0 (0.0%) 2 (0.8%) answers about instruction / teacher 2 (3.6%) 2 (2.3%) 4 (3.8%) 8 (3.2%) negative about instruction / teacher you get additional explanation when you understand it already 2 (3.6%) 2 (2.3%) 4 (3.8%) 8 (3.2%) general and other answers 32 (58.2%) 59 (66.3%) 61 (58.1%) 152 (61.0%) everything is nice / positive comment there’s nothing about this group that i don’t like 12 (21.8%) 34 (38.2%) 41 (39.1%) 87 (35.0%) unspecific / don’t know / other i just don’t like it 20 (36.4%) 25 (28.0%) 20 (19.0%) 65 (26.1%) a number and percentage of comments within that achievement group. the total number of comments exceeds the number of students since some answers belonged to two categories. table a10 students’ responses to the question: why do you learn much or little of being in your achievement group? frequency of comments a answering category example low (n = 52) average (n = 90) high (n = 104) total (n = 246) answers about work / difficulty 9 (17.3%) 17 (18.9%) 42 (40.4%) 68 (27.6%) difficult (positive) / challenging you learn much because you also get more difficult sums 2 (3.9%) 4 (4.4%) 30 (28.9%) 36 (14.6%) working independently i can do it by myself 1 (1.9%) 4 (4.4%) 3 (2.9%) 8 (3.3%) appropriate for my level it’s my level 0 (0.0%) 4 (4.4%) 3 (2.9%) 7 (2.9%) medium or varying difficulty sometimes it’s easy and sometimes it’s difficult 2 (3.9%) 2 (2.2%) 2 (1.9%) 6 (2.4%) too easy / not challenging / want more challenge now it is sometimes a bit too easy 1 (1.9%) 3 (3.3%) 1 (1.0%) 5 (2.0%) too difficult …and i don’t learn much because it’s sometimes difficult 1 (1.9%) 0 (0.0%) 2 (1.9%) 3 (1.2%) easy/not too difficult because the plus-group [highest group] would be too difficult 2 (3.9%) 0 (0.0%) 1 (1.0%) 3 (1.2%) answers about learning 12 (23.1%) 13 (14.4%) 12 (11.5%) 37 (15.0%) learn more / specific math content / getting more sums you learn new goals every time 3 (5.8%) 7 (7.8%) 6 (5.8%) 16 (6.5%) understand more / better i learn much because i understand it better now 3 (5.8%) 5 (5.6%) 3 (2.9%) 11 (4.5%) 90 | f l r get smarter / better at math because i get smarter 1 (1.9%) 1 (1.1%) 2 (1.9%) 4 (1.6%) not understanding because i don’t understand it at all 3 (5.8%) 0 (0.0%) 1 (1.0%) 4 (1.6%) learn less / fewer tasks because you do get fewer sums 2 (3.9%) 0 (0.0%) 0 (0.0%) 2 (0.8%) answers about instruction / teacher 7 (13.5%) 17 (18.9%) 12 (11.5%) 36 (14.6%) positive about instruction / teacher because the teacher gives you more explanations and does more sums with you 7 (13.5%) 15 (16.7%) 12 (11.5%) 34 (13.8%) negative about instruction / teacher i learn more from myself because the teachers confuse me 0 (0.0%) 2 (2.2%) 0 (0.0%) 2 (0.8%) answers about group 5 (9.6%) 7 (7.8%) 7 (6.7%) 19 (7.7%) helping each other / working together when you work together you also learn from it 3 (5.8%) 1 (1.1%) 3 (2.9%) 7 (2.9%) distraction i learn a bit less because sometimes children talk and draw attention 2 (3.9%) 1 (1.1%) 3 (2.9%) 6 (2.4%) positive about group members / dynamics because it’s nice in a nice group 0 (0.0%) 3 (3.3%) 0 (0.0%) 3 (1.2%) concentrate / no distraction quiet so more concentration 0 (0.0%) 2 (2.2%) 1 (1.0%) 3 (1.2%) general and other answers 19 (36.5%) 35 (38.9%) 31 (29.8%) 86 (35.0%) unspecific / don’t know / other i learn much because i learn much from it 17 (32.7%) 34 (37.8%) 29 (27.9%) 80 (32.5%) fun because i think it’s fun 2 (3.9%) 1 (1.1%) 2 (1.9%) 6 (2.4%) a number and percentage of comments within that achievement group. the total number of comments exceeds the number of students since some answers belonged to two categories. table a11 students’ responses to the question: why would you (not) prefer to be in a different achievement group? frequency of comments a answering category example low (n = 54) average (n = 87) high (n = 106) total (n =247) reasons for preferring to stay in the same group answers about work / difficulty 6 (11.1%) 15 (17.2%) 36 (34.0%) 57 (23.1%) difficulty [appropriate in current group] because i don’t do too difficult or too easy work 5 (9.3%) 9 (10.3%) 22 (20.8%) 36 (14.6%) tasks or activities in the current group are fun / want to keep enrichment tasks are fun, so i want to keep those 1 (1.9%) 3 (3.5%) 5 (4.7%) 9 (3.6%) 91 | f l r enrichment tasks my level because this is my level 0 (0.0%) 3 (3.5%) 3 (2.8%) 6 (2.4%) challenge because otherwise i don’t have enough challenge 0 (0.0%) 0 (0.0%) 6 (5.7%) 6 (2.4%) answers about group 6 (11.1%) 3 (3.5%) 6 (5.7%) 15 (6.1%) group members / dynamics / no distraction they are kind and not loud 5 (9.3%) 3 (3.5%) 5 (4.7%) 12 (4.9%) group rank: being in a high(er) / smart(er) group it’s the best group 1 (1.9%) 0 (0.0%) 1 (0.9%) 2 (0.8%) answers about learning 1 (1.9%) 4 (4.6%) 5 (4.7%) 10 (4.1%) learn more / faster / new things because i learn most in this group 1 (1.9%) 4 (4.6%) 3 (2.8%) 8 (3.2%) getting smarter / better at math because i get smarter 0 (0.0%) 0 (0.0%) 2 (1.9%) 2 (0.8%) answers about instruction / teacher 0 (0.0%) 2 (2.3%) 4 (3.8%) 6 (2.4%) positive about instruction / teacher because now i get the explanation that i need 0 (0.0%) 2 (2.3%) 4 (3.8%) 6 (2.4%) general and other answers 14 (25.0%) 34 (39.1%) 44 (41.5%) 92 (37.2%) this group is nice/fun because it’s nice in this group 6 (11.1%) 23 (26.4%) 20 (18.9%) 49 (19.8%) unspecific / don’t know / other i find it hard to explain 8 (14.8%) 11 (12.6%) 24 (22.6%) 43 (17.4%) reasons for preferring to be in a different group answers about work / difficulty 6 (11.1%) 8 (9.2%) 4 (3.8%) 18 (7.3%) difficulty [would be more appropriate in another group] more difficult sums: because i like that 3 (5.6%) 4 (4.6%) 4 (3.8%) 11 (4.5%) tasks or activities in the other group are fun / want to get enrichment tasks because i want enrichment tasks 2 (3.7%) 2 (2.3%) 0 (0.0%) 4 (1.6%) challenge more challenge 1 (1.9%) 2 (2.3%) 0 (0.0%) 3 (1.2%) answers about group 7 (13.0%) 5 (5.8%) 2 (1.9%) 14 (5.7%) group members / dynamics /distraction because i have many friends in the other group 5 (9.3%) 2 (2.3%) 1 (0.9%) 8 (3.2%) group rank: being in a high(er) / because i want to be in a higher group 2 (3.7%) 3 (3.5%) 1 (0.9%) 6 (2.4%) 92 | f l r smart(er) group answers about learning 5 (9.3%) 2 (2.3%) 0 (0.0%) 7 (2.8%) learn more / faster / new things because i learn much 1 (1.9%) 0 (0.0%) 0 (0.0%) 1 (0.4%) getting smarter / better at math / higher level then i think i will get better at math 4 (7.4%) 2 (2.3%) 0 (0.0%) 6 (2.4%) answers about instruction / teacher 1 (1.9%) 1 (1.2%) 0 (0.0%) 2 (0.8%) instruction / teacher then the teacher can help me when i find it difficult 1 (1.9%) 1 (1.2%) 0 (0.0%) 2 (0.8%) general and other answers 8 (14.8%) 13 (14.9%) 5 (4.7%) 26 (10.5%) other group is nice/fun or current group is not nice it’s not as much fun as another group 4 (7.4%) 2 (2.3%) 1 (0.9%) 7 (2.8%) unspecific / don’t know / other i can’t really explain 4 (7.4%) 11 (12.6%) 4 (3.8%) 19 (7.7%) a number and percentage of comments within that achievement group. the total number of comments exceeds the number of students since some answers belonged to two categories. table a12 students’ responses to the question: why would you prefer to work with or without achievement groups? frequency of comments a answering category example low (n = 52) average (n = 87) high (n = 102) total (n = 241) reasons for preferring to retain groups 37 (71.2%) 67 (77.0%) 80 (78.4%) 184 (76.4%) positive about groups / like it as it is because i like to work in a group 13 (25.0%) 29 (33.3%) 18 (17.7%) 60 (24.9%) between-student differences / appropriate difficulty because if you all work at the same level some people don’t get explanation and others don’t get challenge 6 (11.5%) 14 (16.1%) 24 (23.5%) 44 (18.3%) learning / working better with groups because you can learn better like this 2 (3.9%) 8 (9.2%) 13 (12.8%) 23 (9.5%) possibility to get more instruction i like it when you can choose whether you want explanation 4 (7.7%) 2 (2.3%) 3 (2.9%) 9 (3.7%) know your level it’s much more fun when everybody knows in which star [=group] they are 1 (1.9%) 1 (1.5%) 1 (1.0%) 3 (1.2%) unspecific / don’t know / other i don’t know how without [groups] 11 (21.2%) 13 (14.9%) 21 (20.6%) 45 (18.7%) reasons for preferring to work without groups 15 (28.9%) 20 (23.0%) 22 (21.6%) 57 (23.7%) learning / working better then you can do all tasks and then you get smarter 4 (7.7%) 7 (8.0%) 4 (3.9%) 15 (6.2%) 93 | f l r without groups negative about groups it’s so complicated now with all those groups 2 (3.9%) 2 (2.3%) 4 (3.9%) 8 (3.3%) equality, everybody should be/do the same because i think that everybody should get the same tasks 2 (3.9%) 3 (3.5%) 1 (1.0%) 6 (2.5%) unspecific / don’t know / other because i want that 7 (13.5%) 8 (9.2%) 13 (12.8%) 28 (11.6%) a number and percentage of comments within that achievement group. the total number of comments exceeds the number of students since some answers belonged to two categories. frontline learning research vol. 10 no. 2 (2022) 22 44 issn 2295-3159 new materialist network approaches in science education: a method to construct network data from video miikka turkkila, jari lavonen, katariina salmela-aro & kalle juuti university of helsinki, finland article received 15 september 2021 / article revised 19 april 2022 / accepted 14 september 2022/ available online 18 october 2022 abstract lately, new materialism has been proposed as a theoretical framework to better understand material-dialogic relationships in learning, and concurrently network analysis has emerged as a method in science education research. this paper explores how to include materiality in network analysis and reports the development of a method to construct network data from video. the approaches, 1) information flow, 2) material semantic and 3) material engagement, were identified based on the literature on network analysis and new materialism in science education. the method was applied and further improved with a video segment from an upper secondary school physics lesson. the example networks from the video segment show that network analysis is a potential research method within the materialist framework and that the method allows studies into the material and dialogic relationships that emerge when students are engaged in investigations in school. keywords: upper secondary school physics; project-based learning; video-based research corresponding author: miikka turkkila, p.o. box 9 00014 university of helsinki, miikka.turkkila@helsinki.fi doi: https://doi.org/10.14786/flr.v10i2.949 mailto:miikka.turkkila@helsinki.fi turkkila et al 23 | f l r 1. introduction science education research has been criticized for its tendency to ignore material culture and how it is theoretically centred in social culture and especially in discursive practices (milne & scantelebury, 2019; hetherington et al., 2018). in recent years, new materialism has been seen as a potential theoretical approach to start answering this criticism. although there are different orientations of new materialism, it is quite widely agreed that the privileged role of humans is diminished and the role and meaning of materials is emphasised (gamble et al., 2019). current research engaging in new materialism has studied, for example, science teachers’ thoughts about the role of materials (hetherington & wegerif, 2018) and the use of microblogging to support students’ conceptual development (cook, warwick, vrikki, major, & wegerif, 2019). there has also been theoretical work to understand better the role of materials in science learning by hetherington et al. (2018), who propose a theoretical framework based on barad’s agential realism and a bakhtinian dialogic pedagogy. for science education, new materialism implies that, in addition to the students, materials such as laboratory equipment and instruments themselves play an active role in constructing knowledge. moreover, knowledge is not independent of the context, such as the investigations and the material world in which it was constructed. this active role of materials has not been adequately theorised even though there has been research into how and why students should engage in investigative activities (hetherington et al., 2018). although different research orientations have included materiality to a varied degree, typically the instruments are seen as passive tools that when used correctly will produce the expected or correct values. the instruments are given higher authority than the students who use them, and thus the instruments are taken for granted (milne, 2019). for example, project-based learning (pbl) in science sees instruments in conjunction with cognitive tools (krajcik & shin, 2014). the cognitive tools enable students to achieve learning goals that are otherwise unattainable. for example, using data loggers and computer-based laboratory tools, computers function as cognitive tools that show graphical representations of the data. these representations then allow students to see patterns in the data, providing insight into the phenomena students are investigating and enabling knowledge construction. in the new materialist view, computers are seen as active participants in knowledge construction. a computer can modify or even process information or data; for example, the raw data acquired by a sensor from the phenomena is processed and displayed in graphical form. therefore, the computer offers information in a more comprehensible form to the students (actors) who are affected and can construct new knowledge of the phenomena. thus, the computer has become non-human actor in knowledge construction. this kind of interconnected or networked nature of knowledge construction is a feature of new materialist theories, especially in the work of bruno latour (latour & woolgar, 1979; latour, 2005). the actor-network theory (ant) aims to explain social by tracing associations of actor-networks that include human and non-human actors (latour, 2005). in education, ant has been used to understand curriculum or policy development (fenwick & edwards, 2017) and has also been employed in educational reform studies (nespor, 2002; fenwick, 2011). we focus here on students’ investigations and collaboration. consequently, we use a different conceptualisation of a network than the actor-networks of ant. networks based on graph theory are used to analyse and visualise the collection of connected elements, such as communication or social systems (barabási, 2012). in science education research, it has been proposed that these network-based approaches complement other research methods (koponen & mäntylä, 2020) and they have been used to study, for example, classroom interactions (e.g., bokhove, 2018) and classroom discussions (e.g., bruun, lindahl, & linder, 2019). in education, networks are applied mainly through social network analysis (sna). although the two network approaches, sna and ant, differ in their theoretical background and applications, both could benefit from the ideas of the other (vicsek et al., 2016). in particular, the sna researcher should develop methods that include non-human actors in the networks alongside human actors. turkkila et al 24 | f l r in this article, we explore how to include materiality in network analysis and develop a method to construct network data from video within the new materialist perspective in order to understand better the role of material and dialogic relationships in science learning. students’ investigations provide a rich setting to develop a method applying sophisticated network analysis techniques. we begin with a brief review concerning the way in which networks have been used in science education research, and examine what possibilities there are to include materiality in different network approaches. additionally, an authentic video sample from an upper secondary school physics lesson is used to develop new materialist network approaches. 2. network analysis in science education research a network is a collection of nodes (also called vertices or points) that have pairwise connections known as edges (also called links or arcs) that represent a real-world system. the nodes can represent, for example, computers, humans or cities, and the edges represent internet connections, social relationships or railway sections, respectively. depending on what the network represents, the edges may be directional, showing the direction of the connections or edges can have weight to indicate the strength of the connection. weighted directed networks are also possible. examples of basic directed and weighted networks are shown in figure 1. figure 1. examples of a directed simple network (left) and a weighted network (right). the analysis of networks is based on local or global measures characterising the whole network or distinct nodes or edges. one common network measure in social sciences is the different centralities (knoke & yang, 2008), like the degree centrality, which refers to the number of edges a node has. for example, node e is the most central with a node degree of four in figure 1. the more edges a node has, the more central and thus more relevant it is. even though centrality measures are often used, networks can also be analysed using network motifs or roles. network motifs are considered the building blocks of networks (milo, shen-orr, itzkovitz & kashtan, 2002). for example, in figure 1 nodes c, d and g form one triadic motif, and in the directed network the role for nodes c and g is the source, whereas d node’s role is the sink. zweig (2016) argues that for a network analysis to be purposeful there should be flow within the network. this flow originates from local exchange of, for example, used goods, money or e-mails that are transferred or replicated between nodes. in different types of networks this flow is implied to be information; for example, computers connected with the internet or other communication networks. similarly, social networks enable exchange of information between people. however, whereas e-mails turkkila et al 25 | f l r are often sent to multiple recipients, confidential gossip is told to one person at a time, and therefore the flow process is different in the respective networks. understanding the flow process within the network is important in order to choose the right analysis methods, as different centrality measures are appropriate for different types of flow (borgatti, 2005). in education research, network analysis is often understood as the analysis of social networks. social network analysis (sna) is the study of social connections and structures using networks (wasserman, 1994). in sna the nodes represent people, and the edges represent social connections. these connections can be constructed from social media, questionnaires or otherwise reported relationships or interactions. for example, sna has been used to study online written communications (e.g., martínez, dimitriadis, rubia, gómez, & de la fuente, 2003; turkkila & lommi, 2020) and classroom interactions. these interactions have been defined differently, but usually they are connected to student discourse. the types of interactions include each individual utterance between actors (bokhove, 2018), taking turns in argumentation (gonzález-howard, 2019), or discussion in joint work (dou & zwolak, 2019). in addition to analysing social networks, science education research networks have also been used to analyse texts or the content of discussion. for example, bruun et al. (2019) applied semantic networks to analyse annotated speech as text to characterise the whole content of the discussion. this approach adopts the same techniques used in the network analysis of textbooks (yun & park, 2018) and students’ written answers (wagner et al., 2020). similar, but distinct from semantic networks, are conceptual networks that study subject-specific concepts found in texts or utterances (caballero et al. 2020) or in concepts maps (koponen & nousiainen, 2019). there is also epistemic network analysis (ena), which considers epistemic elements coded from text or speech as nodes instead of single words or concepts (shaffer at al., 2009). co-occurrence of the epistemic elements are mapped temporally to investigate, for example, how students share and improve ideas (oshima, oshima & saruwatari, 2020) and how student discourse supports the development of scientific practices (bressler et al., 2019). 2.1 including the material new materialism sees humans as only one type of agent that never acts in isolation of each other (bennett, 2010). in science classrooms or laboratories there are other non-human actors with whom humans act together. thus, the network could comprise these human and non-human actors instead of solely humans. if a network of students can be characterised by the exchange of information, it would be possible to consider laboratory equipment and digital tools as actors capable of providing information (e.g., a graph on a computer screen) to the students. similarly, the measuring apparatus provides students with qualitative information about the phenomenon they are studying (assuming that the phenomenon is visible). thus, a network could be constructed where, for example, computers and other devices that are used to measure and observe phenomena would be nodes alongside students and there would be information exchange between these nodes that would constitute directed edges within the network. if the students are engaged with materials during their discussions, these materials should appear within the student dialogue and therefore within a network constructed from annotated dialogue. however, how the materials appear within the annotations might not be straightforward, and combination with content analysis might be needed. an approach based on semantic networks rather than conceptual networks should be more applicable as the material references may or may not be connected to (physics) concepts. in any case, the network would consist of words relating to the topic of students’ investigations, science concepts or materials as nodes and semantic connection between words are represented by edges. this kind of material semantic network could show how materials are semantically linked or embedded within student dialogues. an additional approach is inspired by assemblages of the material-dialogic framework (see of hetherington et al., 2018). these assemblages are continually formed, dissolved and re-formed by agents engaging materially and dialogically. engaging with materials is seen as embodied actions, like turkkila et al 26 | f l r pointing or manipulating an object, and hence bringing the material directly into the dialogue. similarly, students engage with each other verbally, and through these engagements spoken dialogue and materials are linked. thus, it would be possible to construct a network from these material and verbal engagements, where engagements are represented by the edges in a network of agents. in conclusion, we have identified three possibilities for new materialist network approaches: information flow, material-semantic and material engagement networks. in the next section we explore how these approaches can be applied with a video sample of students’ investigations. 3. constructing material network data from video typically, in video-based research, video file or recording itself is not considered as data but is instead a source of data. how to define that data is a key challenge in video analysis (erickson, 2012). our aim is to construct network data from video; the information used for this construction is student (and teacher) dialogue and their embodied action with each other and the materials evident within the video. we begin with an outline of the method to give general overviews of the process. afterwards we provide a detailed account of the video selection and specific descriptions for each network approach. 3.1 overview of the data processing the flowchart in figure 2 shows an outline of the process. after selecting a video segment for analysis, precise annotations of student and teacher talk are made. next, descriptions of students’ embodied actions are added in brackets next to the annotations. similarly, students’ referencing materials are described and added in brackets into the annotations. the steps up to this point are made within the annotation software when referencing back to the video is needed. turkkila et al 27 | f l r figure 2. flowchart of the data processing. next, the annotations with the description of student actions and references are exported from the annotation software as a spreadsheet. this spreadsheet is then copied onto three separate spreadsheets for the different networks. the individual spreadsheets are used to code information exchange, conduct the manual text processing and code material and verbal engagements. these are then used with network analysis software to construct the network data. 3.2 video selection and annotations the material for constructing the network data were obtained from videos collected as part of a larger research project in which project-based learning units of newtonian mechanics were co-designed turkkila et al 28 | f l r with in-service teachers (juuti et al., 2021; schneider et al., 2020). the pbl units provided rich sources for student activities in which computers and laboratory equipment could play an active role in the knowledge construction. altogether, six focus groups of three to four students from five classrooms of two separate upper secondary schools participated in the study. the activities of each focus group were recorded with two small video cameras during the pbl units, which were six to seven lessons long. during the unit, students planned and investigated the motions of different objects, explored what causes changes in the motion, made models for the objects’ motion with constant speed and acceleration, and constructed explanations. video selection and analysis adapts the three-level approach of video-based research proposed by ash (2007). at each level, videos are divided into shorter segments that are analysed in more detail. here the first level constituted discrete phases of instructional activities that were observed and coded in real time by the researcher operating the video cameras. the raw video material (32 hours) was cut according to the coded phases of instructional activities, resulting in 16 individual cases of student investigations. one such case (duration of 13 mins) was selected for the analysis. the selection was based on video and audio quality, the clarity of student actions and lack of distractions for the students. the selected case is from the early stage of the pbl unit where three students, s1, s2 and s3, investigate the motion of a car on a track using computer-based data loggers and an ultrasound sensor. a sketch of the data collection situation is shown in figure 3. after data collection, the students (s1 and s2) beside the apparatus (a) moved back next to student s3 by the computer (c). figure 3. 3d-sketch showing students’ positions while they measure the speed of the car. while discussing the results, all the students are behind the computer c (yellow). the measuring apparatus, a, consisting of the track (red), the car (blue) and the ultrasound sensor (green). the selected case was further divided into shorter, distinct segments of investigation subtasks identified from the video. these tasks were, for example, preparation, collecting data and interpreting data that were inductively identified from the video. there were clear transitions between these segments where the students, for example, agreed to move on from collecting data to interpreting results. descriptions of each segment are given in table 1. turkkila et al 29 | f l r table 1 segments of distinct student activities during students’ investigations segment label of the segment description of the segment. length number of utterances 1 preparation student s3, who owns the laptop, stays with the computer while the two other students move next to the apparatus (i.e., the track). student s2 takes the ultrasonic sensor and holds it at the far end of the track. student s1 takes the car and prepares to project it down the track it for the measurement. 24 s 15 2 collecting data the students measure the speed of the car three times as student s3 is not satisfied with the graphs from the first two attempts. the apparatus and the computer provide qualitative and quantitative information of the speed of the car. 1min 14 s 27 3 interpreting data the students discuss and interpret the graph. they ponder what to do next. the computer provides information of the measurement. 1min 17 s 43 4 teacher support the students ask the teacher how to analyse the graph. the teacher also gives additional instructions. the computer provides information of the measurement. 50 s 20 5 analysing data the students analyse the graph and discuss what the results mean and what number in the results box indicates the slope. the students discuss what they should analyse. the computer provides information of the measurement. 2min 36 s 60 6 teacher support 2 the students ask the teacher which results are the slope and what else they should analyse. the teacher gives brief instructions. the computer provides information on the measurement. 39 s 20 7 collecting data 2 the students take new measurements with the same set up and positions as in the first measurement. the second attempt is successful with the car clearly being projected faster than before. the apparatus and the computer provide qualitative and quantitative information on the speed of the car. 47 s 23 8 constructing an explanation the students construct an explanation after analysing the new graph and comparing the slopes of the graphs from the first and second measurements. the students arrive at the conclusion that with a higher speed the slope of the graph is greater. the computer provides information on the measurement. 2min 28 s 37 turkkila et al 30 | f l r the third segment, interpreting data, was chosen for initial trials for constructing the network data before analysing the other segments. the segment was chosen so that it was not too short, had the fewest number of outside distractions, and seemed to contain all the necessary student actions and engagements, as detailed in sections 3.3.-3.5. all three approaches (sections 3.3-3.5) to build the network data required annotations of the dialogue. additionally, gestures and embodied actions were needed to interpret information exchange, the materials referenced and material engagements. for example, talk includes many demonstrative pronouns (i.e., this, that) that cannot be deciphered without gestures or embodied actions. the student talk was annotated using elan annotation software (2019). next, different embodied actions and gestures were described next to the annotation in brackets, and finally the meanings of demonstrative pronouns were added in brackets next to the pronouns. an excerpt of the annotations is shown in table 2. table 2 excerpt of annotations from segment 3 interpreting data. student s2 was present but did not participate in the discussion. student s1 student s2 student s3 1 that [part of the graph] is the first [part of the car movement] one 2 over here, here [part of the graph] it [the car], like, slows down [points at the computer screen] 3 all look at the computer screen in silence 4 so, that [part of the graph] is it [part of the car movement] [points at the computer screen] 5 at this point [points at the computer screen] when it [the car] hits that end [points at the end of the track] it [the car] slows down 6 yeah 7 but then it [the car] starts to accelerate again when it [the car] returns in practice, the process of constructing the networks was iterative, namely, looking at the annotated data, going back to the video data and making notes and observations of aspects that needed to be included for constructing the network data. after the initial networks of segment 3 were completed, the practice of constructing the network data was applied to the other segments and revised where necessary. 3.3 information flow information flow can be used to characterise social or communication networks (zweig, 2016). for example, people connected through social media can transmit information between each other or pass information from one to another. similarly, online communication assumes an exchange of information. however, what this information is, is rarely explicitly stated. even if information is not explicitly defined, in written online communications building the network is straightforward. for example, nodes represent people and edges correspond to the messages between people. however, face-to-face communication differs from online communications. face-to face communication includes gestures and other non-verbal communication that are not possible in online communication. spoken utterances are also generally shorter and more indecisive than online turkkila et al 31 | f l r messages. additionally, dialogue contains small utterances that maintain the socially expected structure of the dialogue (heritage, 1984). these are, for example, small agreements (yes, yeah). thus, not all utterances contain relevant information for knowledge construction. additionally, demonstrative pronouns with gestures like pointing can indicate information exchange that is not clear from the text alone. lastly, student discussion might contain irrelevant talk. for example, the students might comment on some outside distraction. the network based on information flow represents the exchange of information by having the students as nodes and edges. to build this network, consideration is first needed when to include an edge, meaning when information is exchanged between the actors. information might have different definitions in different fields of science. in our case, meaningful information is connected to knowledge construction. while conducting the investigation, students need conceptual knowledge about the phenomena they are investigating, but they also need procedural knowledge concerning how to conduct the experiment. thus, information is exchanged when the student’s utterance, or utterance in conjunction with embodied actions, contains information related to the phenomena they are investigating (e.g., the first utterance of student 1 in table 1), or how to investigate that phenomenon and the receiving person is present and not distracted. information exchange from the material world is more challenging, as materials continuously provide information. however, continuously tracking this information is impractical and therefore only clear instances of humans looking at the materials like the computer screen or the measuring apparatus are considered. additionally, as the ultrasound sensor measures the distance of the car as a function of time, it transmits this information to the computer. figure 4. the process of constructing an information flow network. the annotated dialogue with the embodied actions was used to code the information flow into a table on a from-to basis. an example of this, with corresponding network visualisation, is shown in figure 4. each row in the table is a directed edge between nodes (from left to right). the table, and thus the networks, may have multiple edges between them and therefore it is possible to construct weighted networks. 3.4 material semantic a semantic network consists of words and the semantic connections between these words. semantic connections can be found after text pre-processing that removes the structural features of the text. the pre-processing includes removing punctuation, converting capitals to lower case, and removing words that are used for the structure of the sentence rather than to carry meaning. the material semantic network was built on the ideas of thematic discourse network analysis introduced by bruun et al. (2019). the thematic discourse network analysis used an iterative process that combines network analysis and qualitative discourse analysis to study students’ group discussions. qualitative discourse turkkila et al 32 | f l r analysis was used to revise initial semantic networks with the following steps: grammatical reductions to reduce all words referring to the same concept/meaning to one word, combine synonyms and treat slang words as synonyms, treat phrases as a single word and remove indefinite pronouns that do not carry meaning. in addition to synonyms, we considered homonyms as they convey a different meaning even though the word has the same spelling. for example, the finnish word ‘aika’ means ‘time’ or ‘quite’. thus, it was changed to a different synonym with the meaning ‘quite’ so as not to combine these two meanings with the same node. during the dialogue, students refer to materials using demonstrative pronouns. for example, the utterance “here it slows down” refers to the graph (“here”) and to the car on the track (“it”). therefore, to include the material aspects of the dialogue, all demonstrative pronouns that referred to some material aspect or physics knowledge were changed to those references. in practice, to build the material semantic network from the annotated speech, we considered the following steps for text processing: 1. replace demonstrative pronouns with the references 2. grammatical reductions 3. homonyms, synonyms and slang (spoken expressions) 4. remove irrelevant utterances (e.g., reacting to outside distraction) 5. remove punctuation and structural words (computationally). after processing, the text network was built. each distinct word is a node, and a directed edge is added between node a and node b when word b follows word a. figure 5 visualises this process and shows the difference between the network straight from the text and after text processing. figure 5. example of a network without text processing (above) and a material semantic network constructed after text processing. for the complete material semantic network, all unique words or connected phrases in the utterances are added as nodes. then, edges are added between nodes from consecutive words within the utterances. here again, there can be multiple connections between the same nodes as two consecutive words can be present in many different utterances and thus the network can also be constructed as a weighted network. 3.5 material engagement the material engagement network is built around the notion of intra-acting assemblages of the material-dialogic framework of hetherington et al. (2018). as a network representation, an assemblage turkkila et al 33 | f l r would consist of human (i.e., students and teacher) and non-human (i.e., computer and apparatus) actors as nodes. gradually, the assemblages form the material engagement network when the assemblages are formed and re-formed through the students’ verbal and material engagements. thus, the network has the students (and teacher), the computer and the apparatus as nodes just like the information flow network. the engagements are represented, however, by undirected edges as the engagement is shared between the actors without any meaningful direction for the engagement. human actors engage with non-human actors through embodied actions or references. in doing so, they form an edge between human and non-human nodes. similarly, an edge between two human nodes is formed with verbal engagements when a student reacts to another student’s utterance, giving additional voice to the dialogue. figure 6 shows an example of material and verbal engagements. the first utterance of student s1 contains material engagements with the computer and the measuring apparatus that the student points at and references. next, student s3 verbally engages with s1 by agreeing and inviting student s1 to continue. student s1 accepts the invitation and verbally engages with student s3, and again with the apparatus, by referring to the car. figure 6. the process of constructing an information material engagement network. the table with the material engagements is a collection of edges between nodes, as in the information flow network. here there is no direction for the edges, but there may be multiple edges and therefore it is possible to construct weighted networks. 4. constructed networks three different types of networks were constructed for each eight segments (table 1) from the video segment of the students’ investigations. all networks were constructed as weighted networks and node strengths were calculated for each node in each network. the node strength is the sum of the edge weights connected to the node (barrat, barthélemy, pastor-satorras, & vespignani, 2004). the networks were constructed and visualised from the tabulated connections and the processed text using python and the graph-tool module (peixoto, 2014). the 24 visualisations of the networks are shown in figures 7-10. for the information flow and the material engagement networks, the nodes are labelled by the actors: s1, s2, s3 are the students, t is the teacher, c is the computer the students used, and a is the apparatus that includes equipment used for the investigations (i.e., the track, the car and the ultrasound sensor). the nodes in the material semantic networks represent individual words or references, with the strongest (highest strength) nodes labelled with the corresponding words translated from finnish. the weight of each edge is represented by the edge width. similarly, node size is determined by the node’s strength. the layout for the turkkila et al 34 | f l r information flow and the material semantic networks is similar, with nodes arranged into a circular pattern with the same order throughout. the layout of the material semantic networks is based on the fruchterman-reingold spring-block layout (fruchterman & reingold, 1991). figure 7. visualisation of the different networks from segments 1 and 2. turkkila et al 35 | f l r figure 8. visualisation of the different networks from segments 3 and 4. turkkila et al 36 | f l r figure 9. visualisation of the different networks from segments 5 and 6. turkkila et al 37 | f l r figure 10. visualisation of the different networks from segments 7 and 8. the information flow networks show how information was exchanged locally between the different actors. information was received from the measuring apparatus and the computer and exchanged between humans, creating an overall flow of information within the network. collecting data shows clearly that student s3 receives information both from the apparatus and from the computer, whereas the other students only receive information from the apparatus. the student s3 also passes on turkkila et al 38 | f l r the information from the computer screen to the other students. it seems that it could be beneficial for the student to see the phenomena on the measuring apparatus simultaneously with the graph on the computer screen as this allows more opportunities to exchange information and engage with other actors resulting in the strong nodes for student s3. the material semantic networks show the semantic content of the student dialogue and reveal the strongest semantic connections. the material semantic networks of preparation, collecting data 2 and constructing an explanation are separated into two or more components that are not connected with each other. this results from the dialogue having utterances that do not share semantic connections with each other. for example, during the preparation the students did not semantically connect the measuring apparatus and the computer. however, these were connected while doing the measurement, as can be seen in the corresponding semantic network of collecting data in figure 7. in the other two cases, it is possible that the small components separated from the main body of the network are single utterances that while task related are not taken into part of the discussion as a whole. it is also possible that collecting data for the second time, the students did not need to discuss about conducting the measurement as each student’s role was the same as before, and thus there was little or no need to consider who does what and how. similarly, in the last segment the smaller component is a single utterance about scaling the graph that is not an important part of interpreting the data and thus does not semantically connect with the discussion about the slope of the graph and the motion of the car. when the students discuss the results, the main content seems to be the connection between the graph on the computer screen and the car, but not the physics concepts relating to the two. even in the final segment when students are able to connect the slope of the graph with the speed of the car, they do not use the word ‘speed’, instead, they use the phrase ‘goes faster’. this shows that even though the students can connect the phenomena with the measured results, they do not yet use the appropriate physics terminology. the material engagement network shows what the students and the teacher were engaging with during the investigations. during interpreting the data, student s3 is more strongly engaged with the computer than the other two students. overall, the engagement between students s3 and s1 is strongest, but in analysing the data the engagement is strongest between students s3 and s2. in the final segment, student s1 is not engaged at all, and the only engagement is between students s2, s3 and the computer when these two students are discussing the results of the analysis. in general, the teacher, when present, has a very clear role. the teacher receives information from the computer screen and gives information to the students. the material semantic networks are quite small and simple, and this probably results from the segments being short and there not being much dialogue, as the students are mainly listening to the teacher. however, the node representing the graph on the computer screen is strongest, showing that the semantic content in these situations concerns the graph. the engagement of the teacher is mainly with the computer and the student using the computer (student s3). thus, when the teacher is present, he/she receives information about the graph, engages with the student using the computer and gives information. in practice, the teacher sees what stage the students are at and gives guidance for continuation. preparation has the simplest networks of all the segments. it is possible that the network approach is not suited to this kind of student activity, but more likely the students did not need to do much preparing as the measuring apparatus was set up by the teacher previously and there was not much to do for the students other than to conduct the measurement. if they had needed to plan the investigation, and build the apparatus to conduct the measurement more precisely, the networks might be very different. lastly, information flow and material engagement networks in figures 7 to 10 are similar, but the flow network shows the direction of information, so it represents the students’ activity differently. for example, in the final segment, constructing explanation, student s1 is not engaged with other actors yet is receiving information. thus, student s1 is a passive participant in the final discussion about the analysed results, showing that the student is not part of the intra-acting assemblage turkkila et al 39 | f l r 5. discussion currently, science education research within a new materialist frame is still in its early stages. theoretical work has been carried out formulating the material-dialogic approach (hetherington et al., 2018) and some empirical studies have applied the approach (e.g., cook et al., 2019; hetherington & wegerif, 2018). milne and scantlebury (2019) have discussed materiality and material practices more broadly in science education. the growing use and interest of network analyses in science education research led us to explore how network analysis could be applied in authentic classroom learning where the role of materials, such as laboratory equipment, is acknowledged. the main aim of this study was to develop and demonstrate methods for constructing network data from video, using the materialist perspective to take into consideration the materials present in science investigation situations. instead of resulting in a single network approach, we identified three possibilities that were then developed with the video sample from the physics lesson. the resulting networks show clearly that network approaches can be used to study the learning processes where students are engaged in investigations connected with materials for learning science knowledge and skills. for example, in this case the most worthwhile activities seem to be when students are analysing the data or discussing results. then the students are exchanging information and engaging with the materials and each other. this results in strong (large) nodes with heavy (wide) reciprocal edges. moreover, the semantic networks are more complex than at other times. in order to better understand the role of materials in learning, we grounded our approach to network analysis and a new materialistic frame. the resulting method resembles interaction analysis (ia) that investigates the interactions between humans and the objects around them. ia considers human talk and non-verbal communication and aims to identify practices, problems and solutions (jordan & henderson, 1995). ia even has a material perspective as it sees interactions situated in the material world and considers how humans use artefacts and technology. however, ia is interested in the achievement of social order and “how people make sense of each other’s actions” (jordan & henderson, 1995, p.41). this is where the methods diverge. whereas both methods focus on interaction, we construct network data appearing through interactive actions of different actors as seen in the video, in contrast to trying to understand the moment-to-moment construction and maintenance of social order. this allows the use of several network analysis methods to quantitatively investigate different aspects of collaborative learning with a materialist perspective. thus, the research aim, to use cases and data constructed from video, are different between the methods. 5.1 limitations and possible improvements constructing the information flow network involves uncertainty, as coding the information exchange is challenging in face-to-face communication. there is always some need for interpretation, even if the student talk is taken per se. for example, interpreting student embodied actions or material references can be challenging. moreover, there is no distinction between different information or the amount of information per utterance, therefore edge weight does not necessarily represent the quality or quantity of information. similarly, material entities continuously transmit information, making identifying and selecting instances of information exchange challenging. a more detailed coding framework and analysis protocol could improve the interpretation, but it would result in more manual work. additionally, it should be noted that the strength of the nodes cannot be directly interpreted as more flow or a stronger engagement as the longer video segments contain more information exchange and engagements; thus, they the nodes appear stronger in the network. it would be possible to normalise turkkila et al 40 | f l r the networks with the number of connections or length of time to gain more comparable results. however, at least for information flow, information content and thus the amount of information differs from utterance to utterance, and this should also be taken into account if edge weight is to be equated one-to-one with information. this would require a strict coding manual for information and would also require additional work and time spent on the coding. this additional work might not significantly improve the results, as the structures and patterns of the information exchange are more compelling than the exact amount of exchanged information. we considered information relating both to conceptual and procedural knowledge to be relevant information. it could be possible to separate and build networks relating to both types of knowledge that might reveal additional details of information exchange. however, the two knowledge types are not clearly distinguishable from each other, and relating the knowledge types to information within a single utterance might be unfeasible. however, the material semantic networks can distinguish the overall tendency for the type of knowledge. for example, while collecting the data, the strongest nodes are motion detector, end-of-the-track, and measurements, indicating that the students are discussing how to conduct the measurement using the motion detector at the end of the track. conversely, while interpreting the data the strongest nodes are graph and car, showing that the semantic content is the connection between these two. when constructing material semantic networks during the text pre-processing, different steps might be easier to do manually than computationally and vice versa, depending on the tools available and on the language of the actors. for example, speech-to-text automation could help with the annotations as so-called first pass transcripts (moore, 2015) before the manual work of identifying speakers and coding the embodied actions and material references. similarly, eye-tracking technology would enable coding actors to gaze precisely and would help in identifying instances of receiving information from the materials. this would also allow identification of different parts of materials as nodes like, for example, the parts of a graph or the components of the measuring apparatus. concerning the aims of our study, eye tracking was not applied as we wanted the method to cause as little disruption to teaching and learning as possible, and the technology was not mature enough for straightforward deployment. similarly, text-to-speech was seen as unnecessary for the relatively short video segments, especially as the technology in finnish was not reliable enough at the time. however, in the future, with a larger research setting, use of one or both methods could be a major improvement. there are also some practical technical limitations to this method. for example, even with two video cameras not all student actions were clearly visible as they can be obstructed by other students or material objects. similarly, student speech can, at times, be indistinct or ambiguous. however, these challenges are more about video-based research in general and can be addressed by using good practices and appropriate equipment (derry et al., 2010). here the video material was collected from an authentic classroom setting, where there are always trade-offs with camera placements (erickson, 2012). moreover, there is always some degree of external sounds, noises and other distractions in authentic classrooms that affect the video and audio quality. a more controlled setting (e.g., a teaching laboratory) would provide clearer video material that could be used with the automated techniques mentioned above. 5.2 recommendations for future research the networks can be interpreted visually, but they could also be analysed more explicitly with the many different tools that network analysis offers. the analyses could show, for example, the roles in the information exchange, individual contributions to overall semantic networks, or what intra-acting assemblages students are part of. thus, different network approaches and analyses could be used to answer different research questions concerning science learning. for example, combining background information or preand post-test scores would allow one to correlate learning with the roles of information exchange and explain how prior knowledge affects the roles or what effect the roles have on learning. similarly, it could be asked how prior knowledge or experiences affect contributions to the material semantic network. prior experiences with investigations might determine engagements with materials and it could then be asked who uses the materials during investigations and why. turkkila et al 41 | f l r of the individual network approaches, the material engagement network approach seems promising applied within the material-dialogic framework of hetherington et al. (2018). in essence, the networks are a collection of intra-acting assemblages of the framework. these assemblages could be equated with network motifs. it would be possible to compute different motifs and thus ascertain what kind of assemblages are present. computing the network roles from the information flow network could show each individual’s role in the information flow. network roles are directed motifs and include, for example, source, sink and relay (mcdonnell et al., 2014). finding the roles of each actor (human and non-human) could reveal the students’ abilities and the impact of material elements. for example, a student who is a source or who relays information should be more crucial in the knowledge construction of the group. currently, the material semantic network does not differentiate the individual actor’s contributions. for example, the material semantic networks of teacher support are built almost entirely from teacher talk. however, this talk is not explicitly represented in the network. additionally, aggregating the material semantic network into one network and using multidimensional network analysis methods, it would be possible to separate each student’s and teacher’s contribution to the overall material semantic network. understanding each student’s contribution is an important part of the evaluation. moreover, understanding how individual students contribute to the overall knowledge could be used to guide formative assessment practices. additionally, understanding what or how teachers contribute to students’ knowledge could show the role the teacher has while supporting the students. the advantage of network analysis is that it provides an opportunity to analyse large and complex datasets. therefore, the future applicability of our method depends on its scalability. although including materiality requires manual coding, the ongoing development of text-to-speech automation promises easier initial steps for constructing the network data and thus makes constructing larger datasets easier. similarly, machine learning technologies could be used to automate part or all of the manual coding required. here we have demonstrated the steps required to generate the network data from video material. currently, some of these steps were automated and more can be automated in future. thus, scaling the construction of network data should be possible in subsequent studies. finally, a targeted collection of video material could be used. for comparative study, for example, it would be possible to first measure background variables and use that information to select different kinds of groups for video recording. this would limit the amount of video material and ensure high contrast between analysed cases. 6. concluding remarks in this study, we introduced a method to build a network data of human and non-human actors using video. using a network analysis review and a new materialist theoretical background, we identified three possible approaches to constructing network data. additionally, a real-world video was used to refine the approaches. the three approaches are the information flow network, the material semantic network and the material engagement network. each network has its benefits and drawbacks and ultimately an approach combining two or all three networks might be preferred. however, we have shown here that it is possible to build network data from video within the new material perspective. the production of these networks allows the use of network analysis methods to investigate aspects of collaborative learning where the role of materials has been taken into consideration. this work responds to the demand of including materiality in science education research (milne & scantlebury, 2019), contributes new materialist methods to research in education and expands possible methods for a material-dialogic framework. in conclusion, the proposed method allows studies into the material and dialogic relationships that emerge when students are engaged in investigations in school. turkkila et al 42 | f l r keypoints  new materialism and network-based approaches have gained interest in educational research.  the aim was to combine these two approaches into new materialist network approaches to study students’ investigations.  three possible approaches were identified and developed with a video sample from an upper secondary school.  the approaches are: 1) information flow, 2) material semantic and 3) material engagement.  the developed approaches are applicable for future research. acknowledgments this work was supported by the ministry of education and culture, teacher education forum (pi kalle juuti) and the academy of finland (grant number 298323, pi katariina salmela-aro). open access funded by helsinki university library references ash, d. (2007). using video data to capture discontinuous science meaning making in non-school settings. in video research in the learning sciences (pp. 221–240). routledge. https://doi.org/10.4324/9780203877258 barabási, a.-l. (2011). the network takeover. nature physics, 8 (1), 14–16. https://doi.org/10.1038/nphys2188 barrat, a., barthélemy, m., pastor-satorras, r., & vespignani, a. (2004). the architecture of complex weighted networks. proceedings of the national academy of sciences of the united states of america, 101(11), 3747–3752. https://doi.org/10.1073/pnas.0400087101 bennett, j. (2010). vibrant matter: a political ecology of things. durham: duke university press. https://doi.org/10.2307/j.ctv111jh6w bokhove, c. (2018). exploring classroom interaction with dynamic social network analysis. international journal of research and method in education, 41(1), 17–37. https://doi.org/10.1080/1743727x.2016.1192116 borgatti, s. p. (2005). centrality and network flow. social networks, 27 (1), 55–71. https://doi.org/10.1016/j.socnet.2004.11.008 bressler, d. m., bodzin, a. m., eagan, b., & tabatabai, s. (2019). using epistemic network analysis to examine discourse and scientific practice during a collaborative game. journal of science education and technology, 28(5), 553–566. https://doi.org/10.1007/s10956-019-09786-8 bruun, j., lindahl, m., & linder, c. (2019). network analysis and qualitative discourse analysis of a classroom group discussion. international journal of research and method in education, 42(3), 317–339. https://doi.org/10.1080/1743727x.2018.1496414 caballero, d., pikkarainen, t., araya, r., viiri, j., & ... (2020). conceptual network of teachers’ talk: automatic analysis and quantitative measures. fmsera journal, 3(1), 18–31. retrieved from https://journal.fi/fmsera/article/view/79630 https://doi.org/10.4324/9780203877258 https://doi.org/10.1038/nphys2188 https://doi.org/10.1073/pnas.0400087101 https://doi.org/10.2307/j.ctv111jh6w https://doi.org/10.1080/1743727x.2016.1192116 https://doi.org/10.1016/j.socnet.2004.11.008 https://doi.org/10.1007/s10956-019-09786-8 https://doi.org/10.1080/1743727x.2018.1496414 https://journal.fi/fmsera/article/view/79630 turkkila et al 43 | f l r cook, v., warwick, p., vrikki, m., major, l., & wegerif, r. (2019). developing material-dialogic space in geography learning and teaching: combining a dialogic pedagogy with the use of a microblogging tool. thinking skills and creativity, 31, 217-231. https://doi.org/10.1016/j.tsc.2018.12.005 derry, s. j., pea, r. d., barron, b., engle, r. a., erickson, f., goldman, r., hall, r., koschmann, t., lemke, j. l., sherin, m. g., & sherin, b. l. (2010). conducting video research in the learning sciences: guidance on selection, analysis, technology, and ethics. journal of the learning sciences, 19 (1), 3–53. https://doi.org/10.1080/10508400903452884 dou, r., & zwolak, j. p. (2019). practitioner’s guide to social network analysis: examining physics anxiety in an active-learning setting. physical review physics education research, 15 (2), 20105. https://doi.org/10.1103/physrevphyseducres.15.020105 elan (version 5.4) [computer software.] (2019). nijmegen: max planck institute for psycholinguistics, the language archive. retrieved from https://archive.mpi.nl/tla/elan erickson, f. (2012). definition and analysis of data from videotape: some research procedures and their rationales. in green, j.l., green, j., camilli, g., camilli, g., elmore, p.b., & elmore, p. (eds.), handbook of complementary methods in education research, (3rd ed.). routledge. 177–191. https://doi.org/10.4324/9780203874769 fenwick, t. (2011). reading educational reform with actor network theory: fluid spaces, otherings, and ambivalences. educational philosophy and theory, 43(suppl. 1), 114–134. https://doi.org/10.1111/j.1469-5812.2009.00609.x fenwick, t., & edwards, r. (2010). actor–network theory in education. https://doi.org/10.4324/9780203849088 fruchterman, t. m., & reingold, e. m. (1991). graph drawing by force-directed placement. software: practice and experience, 21 (11), 1129–1164. https://doi.org/10.1002/spe.4380211102 gamble, c. n., hanan, j. s., & nail, t. (2019). what is new materialism? angelaki: journal of the theoretical humanities, 24 (6), 111–134. https://doi.org/10.1080/0969725x.2019.1684704 gonzález-howard, m. (2019). exploring the utility of social network analysis for visualizing interactions during argumentation discussions. science education, 103 (3), 503–528. https://doi.org/10.1002/sce.21505 heritage, j. (1984). garfinkel and ethnomethodology. cambridge: polity press. hetherington, l., hardman, m., noakes, j., & wegerif, r. (2018). making the case for a material-dialogic approach to science education. studies in science education, 54 (2), 141–176. https://doi.org/10.1080/03057267.2019.1598036 hetherington, l., & wegerif, r. (2018). developing a material-dialogic approach to pedagogy to guide science teacher education. journal of education for teaching: jet, 44 (1), 27–43. https://doi.org/10.1080/02607476.2018.1422611 jordan, b., & henderson, a. (1995). interaction analysis: foundations and practice. journal of the learning sciences, 4(1), 39–103. https://doi.org/10.1207/s15327809jls0401_2 juuti, k., lavonen, j., salonen, v., salmela-aro, k., schneider, b., & krajcik, j. (2021). a teacher– researcher partnership for professional learning: co-designing project-based learning units to increase student engagement in science classes. journal of science teacher education, 32:6, 625-641. https://doi.org/10.1080/1046560x.2021.1872207 knoke, d., & yang, s. (2008). social network analysis. los angeles, ca; london: sage publications, inc. https://dx.doi.org/10.4135/9781412985864 koponen, i. t., & mäntylä, t. (2020). editorial: networks applied in science education research. education sciences, 10 (5), 142. https://doi.org/10.3390/educsci10050142 koponen, i. t., & nousiainen, m. (2019). pre-service teachers’ knowledge of relational structure of physics concepts: finding key concepts of electricity and magnetism. education sciences, 9(1), [18]. https://doi.org/10.3390/educsci9010018 krajcik, j. s., & shin, n. (2014). project-based learning. in the cambridge handbook of the learning sciences (pp. 275–297). cambridge: cambridge university press. https://doi.org/10.1017/cbo9781139519526 https://doi.org/10.1016/j.tsc.2018.12.005 https://doi.org/10.1080/10508400903452884 https://doi.org/10.1103/physrevphyseducres.15.020105 https://archive.mpi.nl/tla/elan https://doi.org/10.4324/9780203874769 https://doi.org/10.1111/j.1469-5812.2009.00609.x https://doi.org/10.4324/9780203849088 https://doi.org/10.1002/spe.4380211102 https://doi.org/10.1080/0969725x.2019.1684704 https://doi.org/10.1002/sce.21505 https://doi.org/10.1080/03057267.2019.1598036 https://doi.org/10.1080/02607476.2018.1422611 https://doi.org/10.1207/s15327809jls0401_2 https://doi.org/10.1080/1046560x.2021.1872207 https://dx.doi.org/10.4135/9781412985864 https://doi.org/10.3390/educsci10050142 https://doi.org/10.3390/educsci9010018 https://doi.org/10.1017/cbo9781139519526 turkkila et al 44 | f l r latour, b. (2005). reassembling the social: an introduction to actor-network-theory. oxford; new york: oxford university press. https://doi.org/2027/heb32135.0001.001 latour, b., & woolgar, s. (1986). laboratory life: the construction of scientific facts. princeton: princeton university press. https://doi.org/10.2307/j.ctt32bbxc martínez, a., dimitriadis, y., rubia, b., gómez, e., & de la fuente, p. (2003). combining qualitative evaluation and social network analysis for the study of classroom social interactions. computers and education, 41 (4), 353–368. https://doi.org/10.1016/j.compedu.2003.06.001 mcdonnell, m.d., yaveroglu, ö. n., schmerl, b. a., iannella, n. & ward, l. m. (2014).motif-role fingerprints: the building-blocks of motifs, clustering-coefficients and transitivities in directed networks. plos one, 9 (12). https://doi.org/10.1371/journal.pone.0114503 milo, r., shen-orr, s., itzkovitz, s., & kashtan, n. (2002). network motif: simple building blocks of complex networks. science, 298 (5594), 824–827. https://doi.org/10.1126/science.298.5594.824 milne, c. (2019). the materiality of scientific instruments and why it might matter to science education. in c. milne & k. scantlebury (eds.), material practice and materiality: too long ignored in science education (pp. 9–23). cham, switzerland: springer. https://doi.org/10.1007/978-3-030-01974-7_2 milne, & scantlebury, k. (2019). material practice and materiality: too long ignored in science education. cham, switzerland: springer. https://doi.org/10.1007/978-3-030-01974-7 moore, r. j. (2015). automated transcription and conversation analysis. research on language and social interaction, 48(3), 253–270. https://doi.org/10.1080/08351813.2015.1058600 nespor, j. (2002). networks and contexts of freedom. journal of educational change, 3(ccl), 365–382. https://doi.org/10.1023/a:1021281913741 oshima, j., oshima, r., & saruwatari, s. (2020). analysis of students’ ideas and conceptual artifacts in knowledge-building discourse. british journal of educational technology, 51(4), 1308–1321. https://doi.org/10.1111/bjet.12961 peixoto, t. p. (2014). the graph-tool python library. figshare. https://doi.org/10.6084/m9.figshare.1164194 schneider, b., krajcik, j., lavonen, j., & salmela-aro, k. (2020). learning science: the value of crafting engagement in science environments. yale university press. https://doi.org/10.12987/9780300252736 shaffer, d. w., hatfield, d., svarovsky, g. n., nash, p., nulty, a., bagley, e., … mislevy, r. (2009). epistemic network analysis: a prototype for 21st-century assessment of learning. international journal of learning and media, 1(2), 33–53. https://doi.org/10.1162/ijlm.2009.0013 turkkila, m., & lommi, h. (2020). student participation in online content-related discussion and its relation to students’ background knowledge. education sciences, 10 (4), 106. https://doi.org/10.3390/educsci10040106 vicsek, l., király, g., & kónya, h. (2016). networks in the social sciences. corvinus journal of sociology and social policy, 7(2). https://doi.org/10.14267/cjssp.2016.02.04 wagner, s., kok, k., & priemer, b. (2020). measuring characteristics of explanations with element maps. education sciences, 10 (2). https://doi.org/10.3390/educsci10020036 wasserman, s. (1994). social network analysis: methods and applications. cambridge: cambridge university press. https://doi.org/10.1017/cbo9780511815478 yun, e., & park, y. (2018). extraction of scientific semantic networks from science textbooks and comparison with science teachers’ spoken language by text network analysis. international journal of science education, 40 (17), 2118–2136. https://doi.org/10.1080/09500693.2018.1521536 zweig, k. a. (2016). network analysis literacy: a practical approach to the analysis of networks. vienna: springer. https://doi.org/10.1007/978-3-7091-0741-6 https://doi.org/2027/heb32135.0001.001 https://doi.org/10.2307/j.ctt32bbxc https://doi.org/10.1016/j.compedu.2003.06.001 https://doi.org/10.1371/journal.pone.0114503 https://doi.org/10.1126/science.298.5594.824 https://doi.org/10.1007/978-3-030-01974-7_2 https://doi.org/10.1007/978-3-030-01974-7 https://doi.org/10.1080/08351813.2015.1058600 https://doi.org/10.1111/bjet.12961 https://doi.org/10.6084/m9.figshare.1164194 https://doi.org/10.12987/9780300252736 https://doi.org/10.1162/ijlm.2009.0013 https://doi.org/10.3390/educsci10040106 https://doi.org/10.14267/cjssp.2016.02.04 https://doi.org/10.3390/educsci10020036 https://doi.org/10.1017/cbo9780511815478 https://doi.org/10.1080/09500693.2018.1521536 https://doi.org/10.1007/978-3-7091-0741-6 abstract 1. introduction 2. network analysis in science education research 2.1 including the material 3. constructing material network data from video 3.1 overview of the data processing 3.2 video selection and annotations 3.3 information flow 3.4 material semantic 3.5 material engagement 4. constructed networks 5. discussion 5.1 limitations and possible improvements 5.2 recommendations for future research 6. concluding remarks keypoints acknowledgments references readspeaker® docreader™ click here if you are not being automatically redirected. frontline learning research vol. 12 no. 1 (2024) 1 -15 issn 2295-3159 corresponding author: andreas lachner, eberhard karls university tübingen, germany, wilhelmstraße 31, d-72074 tübingen, germany; andreas.lachner@uni-tuebingen.de. doi: https://doi.org/10.14786/flr.v12i1.1179 towards an integrated perspective of teachers’ technology integration: a preliminary model and future research directions andreas lachner, iris backfisch, & ulrike franke1 1 eberhard karls university tübingen, germany article received 26 september 2022 / article revised 17 june 2023 / accepted 5 december 2023 / available online 11 january 2024 abstract technology integration is regarded as a crucial and complex endeavour to enhance students’ learning and prepare them to participate in a digital society. although the research landscape on teachers’ technology integration is vivid and stimulating, an analytical model which synthesises different strands of research to model antecedents (i.e., teachers’ professional competences), processes and outcomes of technology integration in an integrated manner is missing. that said, previous research was often rather product-oriented and ignored potential effects on students’ learning processes and their achievement. to fill this gap, in this paper, we outline a preliminary model, the tpti-model (teachers’ professional competence for technology integration), in which we deliberately link different research perspectives on teachers’ professional competences, professional vision and students’ learning (processes) to model technology integration during teaching. based on the preliminary tpti-model, we propose future research directions, which may allow to gain a better understanding of the teacherand student-related conditions as well as processes of technology integration and their effects on students’ learning. keywords: technology integration, professional competence, professional vision, technology-enhanced teaching, teaching quality mailto:andreas.lachner@uni-tuebingen.de https://doi.org/10.14786/flr.v12i1.1179 lachner, backfisch & franke 2 | f l r 1. introduction the digital transformation is one of the drastic challenges of the 21st century. as such, educational systems are required to continuously prepare students for a digitised society. for instance, in addition to other factors such as school administration or the availability of infrastructure, there is large consensus that teachers play a pivotal role in technology integration and preparing students for a digitally shaped future (backfisch et al., 2020; sailer et al., 2021). despite the central role of technology integration for schooling, however, empirical research demonstrated that in many educational systems, teachers rarely adopt technology into teaching or only to substitute previous teaching processes (backfisch, lachner et al., 2021; fraillon et al., 2020; sailer et al., 2021). for instance, in the international computer and information literacy study (icils, fraillon et al., 2020), teachers indicated that they use technology primarily for substituting conventional teaching activities such as presenting information to students. more innovative adoptions of technology were reported to a less pronounced extent (see also antonietti et al., 2023; fütterer, hoch, et al., 2023, for recent findings). from a research perspective, these findings are crucial as they pose the demanding questions, a) which boundary conditions determine teachers’ technology integration, b) which teaching processes may account for an effective technology integration, and c) whether and how technology integration may contribute to students’ learning processes and their achievement. the current research landscape on teachers’ technology integration is very vivid and stimulating, as it both identified generic as well as specific boundary conditions of technology integration, such as professional knowledge, teacher motivation, and teacher beliefs (backfisch, lachner et al., 2021; ertmer et al., 2012; mishra & koehler, 2006; ottenbreit-leftwich et al., 2010; scherer et al., 2017, 2019). these boundary conditions, however, were rarely put into context of each other, and thus have been investigated mostly in a fragmented manner (see backfisch et al., 2020, for a quasi-experimental approach tackling cognitive and motivational pre-requisites). a second limitation is that researchers often adopted a product-oriented perspective, as they most exclusively investigated only whether but not how teachers integrated technology (see backfisch, lachner et al., 2021; bibi & khan, 2017; pierson, 2001, for more processoriented approaches). thus, it is largely an open question which teaching processes accounted for teachers’ different qualities of technology integration. more importantly, the link between the quality of technology integration and the initiation of students’ learning processes and achievement is largely missing. this type of research on granular processes is needed as it allows to model how the quality of technology integration develops and allow recommendations on how technology integration can adequately be fostered in teacher education. against this background, we propose a preliminary model (see figure 1), the tpti-model (teachers’ professional competence for technology integration) which aims to integrate antecedents (i.e., teachers’ professional competences of technology integration) as well as provide suggestions for a cognitive perspective on a teacher’s processes during technology integration that is linked to students’ learning processes and achievement. to accomplish these goals, we applied generic models of teachers’ professional competence and professional vision (goodwin, 1994; jarodzka et al., 2021; kunter et al., 2013; loewenberg ball et al., 2008; seidel & stürmer, 2014; van es & sherin, 2002; wolff et al., 2021) to model (meta-)cognitive teaching processes (i.e., noticing, reasoning, acting while integrating technology) which may be responsible for successful technology integration. by adopting the concept of teaching quality, we further linked these “teacher variables” to “student variables” to draw conceptual conclusions regarding effects of technology integration processes on learning processes. to this end, we see this paper as potential food for thought to stimulate integrative research in the context of technology integration and provide additional perspectives on how to foster teachers’ continuing professional development regarding technology integration (fütterer, scherer et al., 2023; könig & mulder, 2014; vokatis & zhang, 2016). furthermore, this integrative perspective may also help typologise the different constructs of technology integration and prevent potential jingle-jangle fallacies (marsh et al., 2019), which are erroneous assumptions that either two different constructs are the same, because the same term is used across research contexts (jingle-fallacy), or two constructs are lachner, backfisch & franke 3 | f l r supposed to be different, as they are labelled differently, but bear the identical construct (jangle-fallacies, see 2.1). we want to clearly note that the current scope of this paper was to synthesise previously isolated research strands and propose a potential proposal for future research that builds on an integrated multimethod approach to model and measure antecedents, processes, and products of technology integration. to this end, the empirical validation of the preliminary tpti-model has yet been missing in this paper. rather, we see the proposed tpti-model as a potential analytical starting point, 1) to deliberately link different perspectives on teachers’ technology integration and, more importantly, 2) to take a look forward regarding potential future areas in the research field and to understand the effects of underlying processes. that said, as models clearly aim to reduce real phenomena, and therefore warrant a parsimonious use of assumptions, our preliminary tpti-model takes a distinct perspective regarding teachers’ technology integration focusing on schooling. in this paper, we therefore primarily draw on formal learning processes at school, which preliminarily focus on subject-matter knowledge and skill acquisition. nevertheless, we admit that the integration of other perspectives of learning, such as enculturation and participation, could be stimulating for future model iterations (sfard, 1998; see also dishon, 2022; wegner & nückles, 2015). at this stage of the model, however, a comprehensive consideration of these different perspectives and their delineation would go beyond this paper. 2. the need for an integrative model of teachers’ technology integration in this section, we outline the scientific need for an integrated model of teachers’ technology integration, both from the perspective of potential boundary conditions of technology integration as well as process-models of technology integration (niederhauser & lindstrom, 2018). 2.1 the need for integrating professional competence as antecedents of technology integration in the context of technology integration, many different models and frameworks exist which accentuate different conditions of teachers’ professional competence (niederhauser & lindstrom, 2018). first, most models emphasise teachers’ cognitive pre-requisites as crucial boundary conditions. for instance, the prominent tpack-framework (mishra & koehler, 2006) highlights the critical role of professional knowledge for technology integration. tpack is rooted in general frameworks, such as the one by shulman (1986), who proposed three knowledge components for professional teaching (see also baumert et al., 2010; hill et al., 2005; kunter et al., 2013 for empirical applications): a) content knowledge (ck) constitutes teachers’ subject-specific knowledge of the to-be-taught contents; b) pedagogical knowledge (pk) is operationalised as generic knowledge regarding the realisation of powerful teaching strategies to support students’ learning (baumert et al., 2010; voss et al., 2011); and c) pedagogical content knowledge (pck) is the intersection of content knowledge and pedagogical knowledge describing, content-specific teaching strategies and knowledge about students’ (mis)conceptions (baumert et al., 2010; hill et al., 2005; shulman, 1986, 1987). mishra and koehler (2006) added technological knowledge (tk) as a further component in their framework which refers to knowledge about the functionalities and applications of technologies, which resulted in further intersections: a) technological pedagogical knowledge (tpk) as a generic dimension of technology understanding that support students’ learning across contents (koehler & mishra, 2009; scherer et al., 2017), b) technological content knowledge (tck), as a dimension to apply technologies within a certain domain, and c) technological pedagogical content knowledge (tpack) as a content-specific dimension to apply technologies for subject-matter teaching (koehler & mishra, 2009). in the last decade, several researchers suggested extensions to the tpack-framework. for instance, angeli and valanides (2009) proposed a perspective on the development of tpack and juxtaposed two different developmental processes (transformative versus integrative view). whereas the transformative view suggests that tpack is a separate and distinctive knowledge structure, which is developed over time via deliberate lachner, backfisch & franke 4 | f l r practice from other teacher knowledge structures, the integrative view suggests that tpack is not a distinct knowledge structure. instead, tpack may be spontaneously constructed on the fly by integrating separated knowledge structures during the act of technology integration. relatedly, in addition to these developmental processes, mishra (2019) proposed a contextual knowledge component which highlights situational knowledge about the instructional context (backfisch, lachner et al., 2021; brianza et al., 2022; dishon, 2022; lachner et al., 2019; turner & meyer, 2000), such as knowledge about how a school is functioning, or how organisational change can be enhanced to realise technology integration (see angeli & valanides, 2009, for related concepts). contrarily, technology-acceptance models (tam, see scherer et al., 2019; teo, 2011) emphasise the role of teacher motivation as a fundamental basis to accept and integrate technology for teaching. in the tam-model, the perceived usefulness and perceived ease-of-use constitute important motivational boundary conditions to successfully integrate technology (teo, 2011). in this regard, scherer et al (2019) synthesised findings from 114 studies which examined the relations between teacher motivation (i.e., perceived usefulness and ease of use of educational technologies) and the intention to use these technologies. meta-analytic structural equation modelling showed that the perceived usefulness of educational technologies was the strongest predictor for the intention to use technologies. relatedly, psychology-related approaches adapted generic models of expectancy-value theories (e.g., backfisch, lachner et al., 2021; taimalu & luik, 2019; wozney et al., 2006) towards technology integration. similarly, these models emphasise the role of perceived utility of technology integration and the teachers’ self-efficacy as important boundary conditions. however, whereas tam only proposes indirect effects of self-efficacy on teachers’ intention to use technologies via perceived utility, expectancy-value theories imply direct effects of both self-efficacy and the perceived utility on the use of technologies. to investigate these differential assumptions of tam and expectancy value theories, backfisch, scherer et al. (2021) analysed survey data of n = 524 in-service teachers who taught in fully technology-equipped schools, in which all students had their own tablet device for learning. findings from structural equation modelling showed that both self-efficacy and perceived utility had direct and indirect effects on the frequency of technology use for different classroom scenarios. therefore, the two theoretical perspectives should not be considered exclusive of each other but should be integrated to inform research and practitioners on the impact of teacher motivation on technology integration. a related strand of research regards teacher beliefs and attitudes towards technology integration as further boundary conditions of technology integration (e.g., ertmer et al., 2012; farjon et al., 2019; kim et al., 2013; wilson, 2023). for example, farjon et al. (2019) investigated the relationship between teachers’ attitudes towards technology in general, in education, and attitudes towards the integration of technologies and their perceived level of technology integration (n = 398). the analyses showed that these attitudes were the strongest predictor for their technology integration, even more important than self-assessed skills and available tools. in these studies, the concept of teacher beliefs and teacher motivation has been used relatively inconsistently and comprises relatively stable beliefs for instance about one-self (e.g., self-efficacy) or about the value of technologies for teaching (park & ertmer, 2008) or also about generic epistemological beliefs (kim et al., 2013). for example, tondeur et al. (2017) used the term ‘pedagogical beliefs’ referring to the generic perceptions, premises, or propositions about teaching and learning in their qualitative synthesis. their analysis of 14 selected studies showed that the relationship between pedagogical beliefs and technology use should be considered as bi-directional. this finding indicates that technology-use can be an enabler to change teaching approaches and associated pedagogical beliefs as well as the other way around, that certain pedagogical beliefs, such as constructivist beliefs, can be enablers for technology integration. one desiderate regarding the antecedents of technology integration which emerges due to a plethora of different models is that these models emphasise different aspects of professional competences in isolation (e.g., knowledge versus motivation versus beliefs). likewise, the different concepts are not exclusive enough so that, for instance, beliefs and attitudes may constitute motivational constructs or vice versa. this is not only a constraint of research on technology integration but is rather regarded as a research desiderate in general motivation psychology (see murayama, 2021). similarly, lachner, backfisch & franke 5 | f l r the overly excessive use of self-reports in teacher’s technology integration for assessing professional knowledge may rather reflect the current level of self-efficacy regarding technology integration (see backfisch et al., 2020; lachner et al., 2021). overall, these research practices may have contributed to the well-known phenomenon of jangle-fallacies (see gonzalez et al., 2021), as almost identical constructs (e.g., tpack, self-efficacy, beliefs) are different because they are labelled differently. these caveats require an integrative perspective of teachers’ professional competence regarding technology integration (baier et al., 2019; baumert et al., 2010; kunter et al., 2013, for examples on generic teaching). commonly, professional competence is a multifaceted construct, which refers to clusters of cognitive and motivational pre-requisites but also metacognitive and self-regulatory conditions, which are to-date an underspecified area in the context of technology integration. these sub-facets are regarded to highly depend on each other and are at least conceptually treated in an integrated manner. within these frameworks, professional competence is regarded to be learnable and malleable and a primary aim of teacher education. 2.2 the need for integrating a process-oriented perspective of technology integration another caveat of previous conceptualisations of technology integration is that they were relatively product-oriented and considered technology integration as the “endpoint” of applying technology, and thus, conceptualised the underlying processes of technology integration only to a limited extent. for instance, the substitution augmentation modification redefinition (samr)-model (puentedura, 2006; see also blundell et al., 2022, for a scoping review), conceptualises technology integration as a four-level hierarchical development from technology integration as a simple substitution of previous analogous approaches with no functional change, to redefinition, allowing teachers to apply technology for solving new tasks which could not be realised without technology (see also hughes et al., 2006, for a related model). these approaches provide a sensible lens for analysing whether teachers are capable to exploit the potential technology (backfisch, lachner et al., 2021). at the same time, these approaches are less capable to model the underlying learning and teaching processes (hamilton et al., 2016), emerging during technology integration. therefore, it is important to consider whether and how technology can be integrated to enhance teaching quality, for instance, by providing challenging learning activities (i.e., cognitive activation), by supporting students' learning processes (i.e., supportive climate), and by enabling efficient classroom management (baier et al., 2019; fauth et al., 2014; hugener et al., 2009; kunter et al., 2013). that said, although previous technology integration models, such as samr, propose a developmental hierarchy of technology integration, the cognitive and developmental conditions and its effects on the process of professional technology integration are underspecified (backfisch et al., 2020). a valuable approach could thus be to adopt and integrate process-oriented perspectives to model the underlying processes of teaching behaviour in the context of technology integration. for instance, early research in the context of teacher expertise (e.g., berliner, 1986; borko & livingston, 1989; leinhardt & greeno, 1986, see also jarodzka et al., 2021; lachner et al., 2016, 2019, for recent approaches) adopted the expert-novice paradigm and contrasted the underlying behaviour processes of expert and novice teachers by using granular methods from cognitive psychology, such as think aloud protocols, clinical interviews, and concept mapping. these classic studies can be regarded as precursors of current research of professional vision, which explicitly focusses on the attentive processes of teacher behaviour predominantly during classroom instruction (e.g., jarodzka et al., 2021; lachner et al., 2016; seidel & stürmer, 2014; van es & sherin, 2002). professional vision is regarded as the backbone of teacher behaviour, comprising the subprocesses to notice and interpret crucial events and interactions in the classroom as a prerequisite for enacting teaching. interestingly, in his seminal ethnographic studies, goodwin (1994) brought up a broader definition of professional vision, as it not only was restricted to on-the-fly events such as students’ classroom interactions, but any professional practice, in which attention is a crucial prerequisite to realise complex practices by (collaboratively) applying tools, technologies, and artefacts. as such, in its original meaning, professional vision approaches may also be suited to describe the underlying processes of professional practices such as technology integration. these practices require the attentive processes including noticing and reasoning of the potential of educational technologies to successfully lachner, backfisch & franke 6 | f l r integrate them in the classroom for enhancing teaching quality. considering these (meta-)cognitive processes could allow to model a process-oriented perspective on technology integration. 2.3 the need for relating technology integration with teaching quality and student achievement a final desiderate of previous conceptualisations of technology integration was that it considered the underlying teaching and learning processes only to a limited extent, as most of the research relied on quantitative indicators of technology integration. for instance, in the context of technology integration, a central indicator is the mere frequency of technology integration during teaching (see e.g., backfisch, scherer et al., 2021; fraillon et al., 2020). although these findings give indications regarding the general use and saturation of technology in classrooms, these studies do not allow to provide insights into the quality of technology integration (see backfisch, lachner et al., 2021; scherer et al., 2019; voogt et al., 2013, for a critical consideration). furthermore, research on tam mainly focusses on the teachers’ general acceptance and intention to use technology as a precursor for technology application. however, no assertions can be made whether the integration of these technologies contributed to effective or efficient classroom teaching. to this end, recent research has started to adopt generic indicators of teaching quality in the context of technology integration, for instance, to measure the quality of lesson plans (backfisch et al., 2020; schmid et al., 2021) or retrospective reports of teaching behaviour (backfisch, lachner et al., 2021). these models operationalised high levels of technology integration in terms of a) cognitive activation by presenting students with challenging technology-mediated activities, b) supportive climate, in which students receive adequate support to master their goals, and c) efficient classroom management to maximise time on task. however, whether these indicators really constitute high-quality of technology integration is an open issue. relatedly, another limitation regards the fact that learning processes and associated outcomes were not related to the quality of technology integration. as technology integration can only be seen as an opportunity for student learning during classroom instruction, it is an open issue whether students use this potential for their learning. applying conceptual models which explicitly model the process of teaching as input, and students’ learning processes and outcomes as output which are mediated via the realised quality of technology integration, may therefore constitute a conceptual lens to better understand the complex interactions between teaching and technology-mediated learning processes (see kunter et al., 2013 for applications in generic teaching). 3. an integrative conceptual model of teachers’ technology integration based on generic models of teachers’ professional competences (kunter et al., 2013), professional vision (goodwin, 1994; van es & sherin, 2002), and technology integration (hughes et al., 2006; puentedura. 2006), we developed a model of teachers’ professional competence for technology integration (tpti-model, see figure 1). at the core of the tpti-model, we propose that technology integration should be viewed as an enabler to enhance the process of teaching and therefore to increase teaching quality. as indicated by research on generic teaching quality, it can be assumed that overall teaching quality is not a stable construct, but rather depends on the situational context in which teaching takes place (dishon, 2022). for instance, the availability of infrastructure (tablets, internet access) and time constraints could determine how technology can be implemented in the classroom to foster individual or collaborative learning processes (see backfisch, lachner et al., 2021, for empirical evidence). at the same time, dynamic characteristics of the class and students could also influence teaching quality (zitzmann et al., 2022). at the level of teaching processes, we consider teachers’ professional competences for technology integration, as multi-dimensional antecedents comprising professional technology-related knowledge, motivational orientations, belief systems, and the level of self-regulation as core facets of lachner, backfisch & franke 7 | f l r professional competence (baier et al., 2019; kunter et al., 2013). the level of self-regulation is so far an underspecified facet of professional competence in the context of technology integration and regards teachers’ cognitive, metacognitive, and affective-motivational abilities to monitor and control their teaching behaviours and habits (kunter et al., 2013). these pre-requisites are regarded to interact with each other and serve as a crucial basis for the effective integration of technology. the second part of the model accentuates teachers’ underlying processes during technology integration, which are dependent on the particular level of professional competence (seidel & stürmer, 2014). following common conceptions of professional vision (jarodzka et al., 2021; seidel & stürmer, 2014; van es & sherin, 2002), in the tpti-model, we model technology integration as an iterative and reciprocal three-step process. first, teachers have to notice the distinct potential of technology for improving teaching quality (e.g., multi-modality; manipulation of objects, visualisation). that said, noticing describes whether teachers are able to pay attention to the potential of technology that critically influence teaching and learning in classrooms, and as such, affect student learning in a positive or negative sense. second, teachers are required to adequately reason about the use of technologies. thus, teachers are required to apply evidence-based principles of technology-based teaching and learning to critically reflect on the integration of technology (lachner et al., 2016; seidel & stürmer, 2014). these principles can both be derived from generic and subject-specific evidence. third, teachers have to implement educational technologies in the classroom. throughout these integration processes, teachers have to continuously monitor the underlying processes to regulate their current processes of technology integration. effectively implementing technologies during classroom teaching requires a considerable amount of deliberate teaching experience (e.g., backfisch et al., 2020; meschede et al., 2017). this teaching experience allows to automate technology integration routines and procedures, as well as organise knowledge around encountered teaching cases and experiences which may result in more elaborated and coherently organized knowledge structures (krauss et al., 2008; lachner et al., 2016; pauli & reusser, 2003; putnam, 1987; wolff et al., 2021). putnam (1987) called this specific type of knowledge representations “curriculum scripts” (see also wolff et al., 2021). curriculum scripts refer to generalised knowledge structures a teacher has about distinct types of frequently encountered situations and teaching problems as well as their solutions. as such, curriculum scripts are higher-order knowledge structures which integrate episodic and professional knowledge and allow teachers to rapidly recognise meaningful patterns for technology integration and to make informed and flexible teaching decisions (lachner et al., 2016; putnam, 1987; wolff et al., 2021). again, these three-step processes are considered to be highly constrained by the situational context (turner & meyer, 2000), as, for instance, the availability of infrastructure and tools, but also the involved students with their prerequisites may constrain the activation of different curriculum scripts. we argue that these processes are highly dependent on each other, and can result in cyclical and reciprocal loops of noticing, reasoning, and acting processes, if teachers realise distinct metacognitive self-regulation strategies. lachner, backfisch & franke 8 | f l r figure 1: the tpti-model of teachers’ professional competences for technology integration at the level of learning processes, the provided technology-supported learning activities are seen as opportunities to learn that have to be utilised by students. dependent on students’ prerequisites, they can use these opportunities during teaching to realise germane processes of self-regulated learning, such as cognitive, metacognitive, and motivational-affective learning strategies (nückles et al., 2020; weinstein & mayer, 1986). on the cognitive level, core cognitive processes include organisation and elaboration strategies (see nückles et al., 2020; weinstein & mayer, 1986). organisation strategies help students identify main concepts during teaching and establish relations between the to-be-learned concepts and structure the learning content in a meaningful way. elaboration strategies help students integrate the previously encountered information into their prior knowledge, for instance, by drawing analogies or making examples. both organisation and elaboration strategies have been discussed to enhance students’ meaningful learning. on the metacognitive level, students’ monitoring (of one’s own understanding) and regulation of learning behaviour are regarded as further important processes to successfully enact meaningful cognitive processes. on a motivational level, metacognitive strategies not only regard cognitively oriented strategies, but also the monitoring and regulation of current motivational states, such as boredom or anxiety, that help to maintain successful learning. the realisation of deep-level learning processes should contribute to student achievement. as for professional competence, we argue that these processes are highly interwoven and depend on each other. 4. directions for future research in this article, we proposed a preliminary model (fig. 1) which may guide future research for modelling and measuring teachers’ technology integration. therefore, the main goal of future research could be to empirically test the proposed relationships. in a first attempt of such a research proposal, the development of measures of professional competence, professional vision, technology integration, and student achievement would be a central lachner, backfisch & franke 9 | f l r requirement. objective assessments will be crucial in improving our understanding of the causal relationships of the proposed antecedents, potential processes, and outcomes of technology integration. the construction of objective test instruments is currently at the beginning (see baier & kunter, 2020; drummond & sweeney, 2017; lachner et al., 2019, 2021 for exceptions), as most of the previous studies in the context of technology integration relied on self-reported knowledge assessments (schmidt et al., 2009; e.g., “i can select technologies to use in my classroom that enhance what i teach, how i teach and what students learn.”). yet, it is still an open issue whether self-reports may validly capture the availability of professional knowledge. that is, even when these measures are substantially correlated, the level of judged performance may be higher or lower than the actual performance and prone to distinct cognitive biases. these biases may be even more pronounced for less experienced teachers, as they may tend to over-estimate their professional knowledge (kruger & dunning, 1999; see aesaert et al., 2017). first attempts to measure teachers’ technology-related professional knowledge have been made. for instance, lachner et al. (2019) developed a tpk-test using multiple-choice questions for measuring conceptual and situational tpk (see also baier & kunter, 2020, for related approaches). the authors demonstrated that differences in tpk worked as a function of teacher expertise. moreover, in another study, conceptual tpk was related to teachers’ self-reported task differentiation during the covid-19 pandemic (könig et al., 2020). together, the tests are a crucial basis to test relationships in the tptimodel, however, still the psychometric properties have to be increased to obtain reliable judgments of teachers’ professional knowledge. that said, measures that validly capture the level of technology integration and teaching quality in situ are needed (see fütterer et al., 2022; hammer et al., 2021 for first attempts) to investigate how professional knowledge affects technology integration and student achievement. to reduce the complexity during testing these interwoven relationships, we argue for a stepwise strategy that allows iterative refinements of the tpti-model. in a second and related attempt, we argue to test underlying cognitive and metacognitive processes during technology integration. for instance, think aloud protocols and video analyses in combination with cued retrospective reporting and/or eye-tracking (see maatta et al., 2021; wolff et al., 2016 for recent applications in generic teaching) could help to investigate whether and how these processes of noticing, reasoning, and acting may be responsible for different qualities of technology integration. for instance, eye-tracking could allow to analyse the attentive processes during lesson planning and help to understand which features teachers are focussing on when selecting technologies during teaching. additionally, mobile eye-tracking systems in combination with video analyses could allow to simultaneously trace the attentive processes during technology integration in the classroom both from a student and a teacher perspective (see jarodzka et al., 2021). given that the situational context plays a decisive role during technology integration, it makes sense to generalise the obtained findings to other teaching contexts (e.g., subject, class, and student characteristics), for instance by adopting a many-classes approach (fyfe et al., 2021; see also lachner et al., 2021 for recent applications). many-classes approaches allow to explore patterns such as relations among antecedents and processes of technology integration across a variety of class contexts and educational implementations, subject areas, and students. as such, this approach allows to understand whether distinct relations are generalisable and/or whether findings are determined by distinct situational contexts (fyfe et al., 2021). a third attempt would be to closely investigate the interplay of the facets of teachers’ professional competence. so far, these relationships have been investigated mainly in isolation. however, there is increasing agreement in educational psychology and teacher education that not only the availability of professional knowledge is important, but putting them in combination with motivational orientations, beliefs, and self-regulation abilities seems crucial (kunter et al., 2013). that said, increasingly more research has been showing that not a single variable may account for different outcomes, such as teaching quality, but rather their combinations. these person-centred approaches find more and more ways in research on professional competence (e.g., holzberger et al., 2019; thommen et al., 2021), but have seldomly been used in the context of teachers’ technology integration. lachner, backfisch & franke 10 | f l r a fourth and final attempt would be to close the loop within the tpti-model and relate the teaching and learning processes during technology integration towards student achievement. investigating those relationships among teaching and learning processes during technology integration would help to uncover whether and how technology integration can contribute to students’ learning and achievement in authentic classroom settings, which to date is still an open question in educational research. as for the processes, process analyses by means of advanced and sophisticated technologies such as eye-tracking or video analyses in combination with test data could help trace the effects of technology integration on students’ learning (goldberg et al., 2021; haataja et al., 2021). that said, student achievement should not only be investigated from a cognitive perspective, but also include motivational orientations, belief systems, as well as self-regulation, as a multidimensional construct. 5. conclusions and discussion although there is increased interest from researchers and practitioners in how technology can be successfully integrated in the classroom, our knowledge of when and why technology integration promotes learning is still limited. in this theoretical contribution, we aimed to take a step toward filling this knowledge gap by proposing the tpti-model, a preliminary model of teachers’ technology integration, synthesising models of professional competence, professional vision, and the associated learning and teaching processes during technology integration. one of the clear limitations is that most of the relationships within the model have not yet been tested. we hope that this model may mark a stimulating research road map that can help move this relatively young but promising field of research forward to successfully delineate the underlying conditions of successful technology integration. that said, in future iterations of the model, it could be fruitful to consider additional forms of (informal) learning, such as problem solving, collaborative learning, or creative learning, and explore whether and how the tpti-model may generalise to the development of such competences. in line with these suggestions, whether and how the tpti-model may transfer to different educational stages, such as higher education, is another open issue (see sailer et al., 2021). together, we hope that the tpti-model may provide a starting point for integrative research on the process of technology integration and guide ways to systematically analyse effective strategies for technology integration. keypoints technology integration was predominantly treated in an isolated manner. we provide a comprehensive model which explicitly links teaching and learning processes during technology integration. therefore, we bridge research from professional competence and professional vision. additionally, we propose directions for future research derived from the conceptual model. methodological advancements in the context of technology integration are addressed. acknowledgments the research reported in this article was supported by the federal ministry of education and research in germany (bmbf) under contract number 01ja2009. lachner, backfisch & franke 11 | f l r references aesaert, k., voogt, j., kuiper, e., & van braak, j. (2017). accuracy and bias of ict self-efficacy: an empirical study into students’ over-and underestimation of their ict competences. computers in human behavior, 75, 92-102. https://doi.org/10.1016/j.chb.2017.05.010 angeli, c., & valanides, n. (2009). epistemological and methodological issues for the conceptualization, development, and assessment of ict-tpck: advances in technological pedagogical content knowledge (tpck). computers & education, 52(1), 154–168. https://doi.org/10.1016/j.compedu.2008.07.006 antonietti, c., schmitz, m. l., consoli, t., cattaneo, a., gonon, p., & petko, d. (2023). development and validation of the icap technology scale to measure how teachers integrate technology into learning activities. computers & education, 192, 104648. https://doi.org/10.1016/j.compedu.2022.104648 backfisch, i., lachner, a., hische, c., loose, f., & scheiter, k. (2020). professional knowledge or motivation? investigating the role of teachers’ expertise on the quality of technology-enhanced lesson plans. learning and instruction, 66, 101300. doi: 10.1016/j.learninstruc.2019.101300 backfisch, i., lachner, a., stürmer, k., & scheiter, k. (2021). variability of teachers’ technology integration in the classroom: a matter of utility! computers and education, 166, 104159. https://doi.org/10.1016/j.compedu.2021.104159 backfisch, i., scherer, r., siddiq, f., lachner, a., & scheiter, k. (2021). teachers’ technology use for teaching: comparing two explanatory mechanisms. teaching and teacher education, 104, 103390. https://doi.org/10.1016/j.tate.2021.103390 baier, f., decker, a.-t., voss, t., kleickmann, t., klusmann, u., & kunter, m. (2019). what makes a good teacher? the relative importance of mathematics teachers’ cognitive ability, personality, knowledge, beliefs, and motivation for instructional quality. british journal of educational psychology, 89, 767– 786. https://doi.org/10.1111/bjep.12256 baier, f., & kunter, m. (2020). construction and validation of a test to assess (pre-service) teachers' technological pedagogical knowledge (tpk). studies in educational evaluation, 67, 100936. https://doi.org/10.1016/j.stueduc.2020.100936 baumert, j., kunter, m., blum, w., brunner, m., voss, t., jordan, a., klusmann, u., krauss, s., neubrand, m., & tsai, y.-m. (2010). teachers’ mathematical knowledge, cognitive activation in the classroom, and student progress. american educational research journal, 47(1), 133-180. https://doi.org/10.3102/0002831209345157 berliner, d. c. (1986). in pursuit of the expert pedagogue. educational researcher, 15(7), 5-13. http://www.jstor.org/stable/1175505 bibi, s., & khan, s. h. (2017). tpack in action: a study of a teacher educator’s thoughts when planning to use ict. australasian journal of educational technology, 33(4). https://doi.org/10.14742/ajet.3071 blundell, c. n., mukherjee, m., & nykvist, s. (2022). a scoping review of the application of the samr model in research. computers and education open, 3, 100093.https://doi.org/10.1016/j.caeo.2022.100093 borko, h., & livingston, c. (1989). cognition and improvisation: differences in mathematics instruction by expert and novice teachers. american educational research journal, 26(4), 473-498. https://doi.org/ 10.3102/00028312026004473 brianza, e., schmid, m., tondeur, j., & petko, d. (2022, april). investigating contextual knowledge within tpack: how has it been done empirically so far?. in society for information technology & teacher education international conference (pp. 2204-2212). association for the advancement of computing in education (aace). dishon, g. (2022). what kind of revolution? thinking and rethinking educational technologies in the time of covid-19. journal of the learning sciences, 31(3), 458-476. https://doi.org/10.1080/10508406.2021.2008395 drummond, a., & sweeney, t. (2017). can an objective measure of technological pedagogical content knowledge (tpack) supplement existing tpack measures?. british journal of educational technology, 48(4), 928-939. https://doi.org/10.1111/bjet.12473 ertmer, p. a., ottenbreit-leftwich, a. t., sadik, o., sendurur, e., & sendurur, p. (2012). teacher beliefs and technology integration practices: a critical relationship. computers & education, 59(2), 423-435. https://doi.org/10.1016/j.compedu.2012.02.001 https://doi.org/10.1016/j.chb.2017.05.010 https://psycnet.apa.org/doi/10.1016/j.compedu.2008.07.006 https://doi.org/10.1016/j.compedu.2021.104159 https://doi.org/10.1016/j.tate.2021.103390 https://doi.org/10.1111/bjep.12256 http://www.jstor.org/stable/1175505 https://doi.org/10.1016/j.caeo.2022.100093 https://doi.org/10.1080/10508406.2021.2008395 lachner, backfisch & franke 12 | f l r farjon, d., smits, a., voogt, j., farjon, d., smits, a., & voogt, j. (2019). technology integration of preservice teachers explained by attitudes and beliefs, competency, access, and experience. computers & education, 130, 81–93. https://doi.org/10.1016/j.compedu.2018.11.010 fauth, b., decristan, j., rieser, s., klieme, e., & buttner, g. (2014). student ratings of teaching quality in primary school: dimensions and prediction of student outcomes. learning and instruction, 29, 1–9. https://doi.org/10.1016/j.learninstruc.2013.07.001. fraillon, j., ainley, j., schulz, w., friedman, t., & duckworth, d. (2020). preparing for life in a digital world: iea international computer and information literacy study 2018 international report (p. 297). springer. fütterer, t., scheiter, k., cheng, x., & stürmer, k. (2022). quality beats frequency? investigating students’ effort in learning when introducing educational technology in classrooms. contemporary educational psychology, 69, 102042. https://doi.org/10.1016/j.cedpsych.2022.102042 fütterer, t., scherer, r., scheiter, k., stürmer, k., & lachner, a. (2023). will, skill or conscientiousness: what predicts teachers’ intention to participate in technology-related professional development? computers & education, 198, 104756. https://doi.org/10.1016/j.compedu.2023.104756 fyfe, e. r., de leeuw, j. r., carvalho, p. f., goldstone, r. l., sherman, j., admiraal, d., ... & motz, b. a. (2021). manyclasses 1: assessing the generalizable effect of immediate feedback versus delayed feedback across many college classes. advances in methods and practices in psychological science, 4(3), 25152459211027575. https://doi.org/10.1177/25152459211027575 goldberg, p., sümer, ö., stürmer, k., wagner, w., göllner, r., gerjets, p., kasneci, e., & trautwein, u. (2021). attentive or not? toward a machine learning approach to assessing students’ visible engagement in classroom instruction. educational psychology review, 33, 27-49. https://doi.org/10.1007/s10648-01909514-z gonzalez, o., mackinnon, d. p., & muniz, f. b. (2021). extrinsic convergent validity evidence to prevent jingle and jangle fallacies. multivariate behavioral research, 56(1), 3-19. https://doi.org/10.1080/00273171.2019.1707061 goodwin, c. (1994). professional vision. american anthropologist, 96(3), 606–633. https://doi.org/10.1525/aa.1994.96.3.02a00100 hamilton, e. r., rosenberg, j. m., & akcaoglu, m. (2016). the substitution augmentation modification redefinition (samr) model: a critical review and suggestions for its use. techtrends, 60(5), 433-441. https://doi.org/10.1007/s11528-016-0091-y hammer, m., göllner, r., scheiter, k., fauth, b., & stürmer, k. (2021). for whom do tablets make a difference? examining student profiles and perceptions of instruction with tablets. computers & education, 104147. https://dx.doi.org/10.1016/j.compedu.2021.104147 haataja, e., salonen, v., laine, a., toivanen, m., & hannula, m. s. (2021). the relation between teacherstudent eye contact and teachers’ interpersonal behavior during group work: a multiple-person gazetracking case study in secondary mathematics education. educational psychology review, 33(1), 51-67. https://doi.org/10.1007/s10648-020-09538-w hill, h. c., rowan, b., & ball, d. l. (2005). effects of teachers’ mathematical knowledge for teaching on student achievement. american educational research journal, 42(2), 371-406. https://doi.org/ 10.3102/00028312042002371 holzberger, d., praetorius, a. k., seidel, t., & kunter, m. (2019). identifying effective teachers: the relation between teaching profiles and students’ development in achievement and enjoyment. european journal of psychology of education, 34(4), 801-823. https://doi.org/10.1007/s10212-018-00410-8 hugener, i., pauli, c., reusser, k., lipowsky, f., rakoczy, k., & klieme, e. (2009). teaching patterns and learning quality in swiss and german mathematics lessons. learning and instruction, 19, 66–78. https://doi.org/10.1016/j.learninstruc.2008.02.001. hughes, j., thomas, r., & scharber, c. (2006). assessing technology integration: the rat–replacement, amplification, and transformation-framework. society for information technology & teacher education international conference (pp. 1616–1620). association for the advancement of computing in education (aace). https://doi.org/10.1016/j.compedu.2018.11.010 https://doi.org/10.1016/j.cedpsych.2022.102042 https://doi.org/10.1177%2f25152459211027575 https://doi.org/10.1080/00273171.2019.1707061 https://psycnet.apa.org/doi/10.1525/aa.1994.96.3.02a00100 https://dx.doi.org/10.1016/j.compedu.2021.104147 https://doi.org/10.1007/s10212-018-00410-8 lachner, backfisch & franke 13 | f l r jarodzka, h., skuballa, i., & gruber, h. (2021). eye-tracking in educational practice: investigating visual perception underlying teaching and learning in the classroom. educational psychology review, 33(1), 110. https://doi.org/10.1007/s10648-020-09565-7 kim, c., kim, m. k., lee, c., spector, j. m., & demeester, k. (2013). teacher beliefs and technology integration. teaching and teacher education, 29, 76-85. https://doi.org/10.1016/j.tate.2012.08.005 koehler, m. j., & mishra, p. (2009). what is technological pedagogical content knowledge? contemporary issues in technology and teacher education (cite), 9(1), 60-70. https://citejournal.org/volume-9/issue1-09/general/what-is-technological-pedagogicalcontent-knowledge könig, c., & mulder, r. h. (2014). a change in perspective – teacher education as an open system. frontline learning research, 2(5), 26–45. https://doi.org/10.14786/flr.v2i4.109 könig, j., jäger-biela, d. j., & glutsch, n. (2020). adapting to online teaching during covid-19 school closure: teacher education and teacher competence effects among early career teachers in germany. european journal of teacher education, 43(4), 608-622. https://doi.org/10.1080/02619768.2020.1809650 krauss, s., brunner, m., kunter, m., baumert, j., blum, w., neubrand, m., & jordan, a. (2008). pedagogical content knowledge and content knowledge of secondary mathematics teachers. journal of educational psychology, 100(3), 716. https://doi.org/10.1037/0022-0663.100.3.716 kruger, j., & dunning, d. (1999). unskilled and unaware of it: how difficulties in recognizing one's own incompetence lead to inflated self-assessments. journal of personality and social psychology, 77(6), 1121–1134. https://doi.org/10.1037/0022-3514.77.6.1121 kunter, m., klusmann, u., baumert, j., richter, d., voss, t., & hachfeld, a. (2013). professional competence of teachers: effects on instructional quality and student development. journal of educational psychology, 105(3), 805–820. https://doi.org/10.1037/a0032583 lachner, a., backfisch, i., & stürmer, k. (2019). a test-based approach of modeling and measuring technological pedagogical knowledge. computers & education, 142, 103645. https://doi.org/10.1016/j.compedu.2019.103645 lachner, a., fabian, a., franke, u., preiß, j., jacob, l., führer, c., küchler, u., paravicini, w., randler, t., & thomas, p. (2021). fostering pre-service teachers’ technological pedagogical content knowledge (tpack): a quasi-experimental field study. computers & education, 174, 104304. https://doi.org/10.1016/j.compedu.2021.104304 lachner, a., jarodzka, h., & nückles, m. (2016). what makes an expert teacher? investigating teachers' professional vision and discourse abilities. instructional science, 44(3), 197-203. doi:10.1007/s11251016-9376-y leinhardt, g., & greeno, j. g. (1986). the cognitive skill of teaching. journal of educational psychology, 78(2), 75. https://doi.org/10.1037/0022-0663.78.2.75 loewenberg ball, d., thames, m. h., & phelps, g. (2008). content knowledge for teaching: what makes it special? journal of teacher education, 59(5), 389–407. https://doi.org/10.1177/0022487108324554 maatta, o., mcintyre, n., palomäki, j., hannula, m. s., scheinin, p., & ihantola, p. (2021). students in sight: using mobile eye-tracking to investigate mathematics teachers’ gaze behaviour during task instructiongiving. frontline learning research, 9(4), 92–115. https://doi.org/10.14786/flr.v9i4.965 marsh, h. w., pekrun, r., parker, p. d., murayama, k., guo, j., dicke, t., & arens, a. k. (2019). the murky distinction between self-concept and self-efficacy: beware of lurking jingle-jangle fallacies. journal of educational psychology, 111(2), 331–353. https://doi.org/10.1037/edu0000281 meschede, n., fiebranz, a., möller, k., & steffensky, m. (2017). teachers’ professional vision, pedagogical content knowledge and beliefs: on its relation and differences between student and certified teachers. teaching and teacher education, 66, 158–170. https://doi.org/10.1016/j.tate.2017.04.010 mishra, p. (2019). considering contextual knowledge: the tpack diagram gets an upgrade. journal of digital learning in teacher education, 35(2), 76-78. https://doi.org/10.1080/21532974.2019.1588611 mishra, p., & koehler, m. j. (2006). technological pedagogical content knowledge: a framework for teacher knowledge. teachers college record, 108(6), 1017–1054. https://doi.org/10.1111/j.14679620.2006.00684.x https://doi.org/10.1016/j.tate.2012.08.005 https://psycnet.apa.org/doi/10.1037/0022-3514.77.6.1121 https://psycnet.apa.org/doi/10.1037/a0032583 https://doi.org/10.1016/j.compedu.2021.104304 https://psycnet.apa.org/doi/10.1037/edu0000281 https://doi.org/10.1016/j.tate.2017.04.010 https://psycnet.apa.org/doi/10.1111/j.1467-9620.2006.00684.x https://psycnet.apa.org/doi/10.1111/j.1467-9620.2006.00684.x lachner, backfisch & franke 14 | f l r murayama, k. (2021). motivation resides only in our language, not in our mental processes. in m. bong, s. kim, and j. reeve (eds). motivation science: controversies and insights. oxford university press. niederhauser, d. s., & lindstrom, d. l. (2018). instructional technology integration models and frameworks: diffusion, competencies, attitudes, and dispositions. in j. voogt, g. knezek, r. christensen & k. w. lai (eds.), second handbook of information technology in primary and secondary education. springer international publishing, 1–21. nückles, m., roelle, j., glogger-frey, i., waldeyer, j., & renkl, a. (2020). the self-regulation-view in writingto-learn: using journal writing to optimize cognitive load in self-regulated learning. educational psychology review, 32(4), 1089-1126. https://doi.org/10.1007/s10648-020-09541-1 ottenbreit-leftwich, a. t., glazewski, k. d., newby, t. j., & ertmer, p. a. (2010). teacher value beliefs associated with using technology: addressing professional and student needs. computers & education, 55(3), 1321-1335. https://doi.org/10.1016/j.compedu.2010.06.002 park, s. h., & ertmer, p. a. (2008). examining barriers in technology‐enhanced problem‐based learning: using a performance support systems approach. british journal of educational technology, 39(4), 631-643. https://doi.org/10.1111/j.1467-8535.2008.00858.x pauli, c., & reusser, k. (2003). unterrichtsskripts im schweizerischen und im deutschen mathematikunterricht. unterrichtswissenschaft, 31(3), 238-272. https://doi.org/10.25656/01:6779 pierson, m. e. (2001). technology integration practice as a function of pedagogical expertise. journal of research on computing in education, 33(4), 413-430. https://doi.org/10.1080/08886504.2001.10782325 puentedura, r. (2006). transformation, technology, and education [blog post]. retrieved from http://hippasus.com/resources/tte/. putnam, r. t. (1987). structuring and adjusting content for students: a study of live and simulated tutoring of addition. american educational research journal, 24(1), 13-48. https://doi.org/10.3102/00028312024001013 sailer, m., schultz-pernice, f., & fischer, f. (2021). contextual facilitators for learning activities involving technology in higher education: the c♭-model. computers in human behavior, 121, 106794. https://doi.org/10.1016/j.chb.2021.106794 scherer, r., siddiq, f., & tondeur, j. (2019). the technology acceptance model (tam): a meta-analytic structural equation modeling approach to explaining teachers’ adoption of digital technology in education. computers & education, 128. https://doi.org/10.1016/j.compedu.2018.09.009 scherer, r., tondeur, j., & siddiq, f. (2017). on the quest for validity: testing the factor structure and measurement invariance of the technology-dimensions in the technological, pedagogical, and content knowledge (tpack) model. computers & education, 112, 1-17. https://doi.org/10.1016/j.compedu.2017.04.012 schmid, m., brianza, e., & petko, d. (2021). self-reported technological pedagogical content knowledge (tpack) of pre-service teachers in relation to digital technology use in lesson plans. computers in human behavior, 115, 106586. https://doi.org/10.1016/j.chb.2020.106586 schmidt, d. a., baran, e., thompson, a. d., mishra, p., koehler, m. j., & shin, t. s. (2009). technological pedagogical content knowledge (tpack) the development and validation of an assessment instrument for preservice teachers. journal of research on technology in education, 42(2). https://doi.org/10.1080/15391523.2009.10782544 seidel, t., & stürmer, k. (2014). modeling and measuring the structure of professional vision in preservice teachers. american educational research journal, 51(4), 739–771. https://doi.org/10.3102/0002831214531321 sfard, a. (1998). on two metaphors for learning and the dangers of choosing just one. educational researcher, 27(2), 4-13. https://doi.org/10.2307/1176193 shulman, l. (1987). knowledge and teaching: foundations of the new reform. harvard educational review, 57(1), 1-23. https://doi.org/10.17763/haer.57.1.j463w79r56455411 shulman, l. s. (1986). those who understand: knowledge growth in teaching. educational researcher, 15(2), 4–14. https://doi.org/10.3102/0013189x015002004 https://doi.org/10.1016/j.compedu.2010.06.002 https://doi.org/10.1016/j.compedu.2017.04.012 https://doi.org/10.1016/j.chb.2020.106586 https://doi.org/10.1080/15391523.2009.10782544 https://doi.org/10.3102/0002831214531321 https://doi.org/10.17763/haer.57.1.j463w79r56455411 lachner, backfisch & franke 15 | f l r taimalu, m., & luik, p. (2019). the impact of beliefs and knowledge on the integration of technology among teacher educators: a path analysis. teaching and teacher education, 79, 101–110. https://doi.org/10.1016/j.tate.2018.12.012 teo, t. (2011). factors influencing teachers' intention to use technology: model development and test. computers & education, 57(4), 2432e2440. https://doi.org/10.1016/j.compedu.2011.06.008 thommen, d., sieber, v., grob, u., & praetorius, a. k. (2021). teachers’ motivational profiles and their longitudinal associations with teaching quality. learning and instruction, 76, 101514. https://doi.org/10.1016/j.learninstruc.2021.101514 tondeur, j., van braak, j., ertmer, p. a., & ottenbreit-leftwich, a. (2017). understanding the relationship between teachers’ pedagogical beliefs and technology use in education: a systematic review of qualitative evidence. educational technology research and development, 65(3), 555–575. https://doi.org/10.1007/s11423-016-9481-2 turner, j. c., & meyer, d. k. (2000). studying and understanding the instructional contexts of classrooms: using our past to forge our future. educational psychologist, 35(2), 69–85. https://doi.org/10.1207/s15326985ep3502_2 van es, e. a., & sherin, m. g. (2002). learning to notice: scaffolding new teachers’ interpretations of classroom interactions. journal of technology and teacher education, 10(4), 571-596. https://www.learntechlib.org/primary/p/9171/. vokatis, b., & zhang, j. (2016). the professional identity of three innovative teachers engaging in sustained knowledge building using technology. frontline learning research, 4(1), 58–77. https://doi.org/10.14786/flr.v4i1.223 voogt, j., fisser, p., pareja roblin, n., tondeur, j., & van braak, j. (2013). technological pedagogical content knowledge–a review of the literature. journal of computer assisted learning, 29(2), 109-121. https://doi.org/10.1111/j.1365-2729.2012.00487.x voss, t., kunter, m., & baumert, j. (2011). assessing teacher candidates' general pedagogical/psychological knowledge: test construction and validation. journal of educational psychology, 103(4), 952–969. https://doi.org/10.1037/a0025125 wegner, e., & nückles, m. (2015). training the brain or tending a garden? students' metaphors of learning predict self-reported learning patterns. frontline learning research, 3(4), 95-109. http://dx.doi.org/10.14786/flr.v3i4.212 weinstein, c., & mayer, r. (1986). the teaching of learning strategies. in: m. wittrock (ed.), handbook of research on teaching, macmillan, 315-327. wilson, m. l. (2023). the impact of technology integration courses on preservice teacher attitudes and beliefs: a meta-analysis of teacher education research from 2007–2017. journal of research on technology in education, 55(2), 252-280. https://doi.org/10.1080/15391523.2021.1950085 wolff, c. e., jarodzka, h., & boshuizen, h. (2021). classroom management scripts: a theoretical model contrasting expert and novice teachers’ knowledge and awareness of classroom events. educational psychology review, 33(1), 131-148. https://doi.org/10.1007/s10648-020-09542-0 wolff, c. e., jarodzka, h., van den bogert, n., & boshuizen, h. (2016). teacher vision: expert and novice teachers’ perception of problematic classroom management scenes. instructional science, 44(3), 243-265. https://doi.org/10.1007/s11251-016-9367-z wozney, l., venkatesh, v., & abrami, p. (2006). implementing computer technologies: teachers' perceptions and practices. journal of technology and teacher education, 14(1). retrieved september 22, 2022 from https://www.learntechlib.org/primary/p/5437/. zitzmann, s., wagner, w., hecht, m., helm, c., fischer, c., bardach, l., & göllner, r. (2022). how many classes and students should ideally be sampled when assessing the role of classroom climate via student ratings on a limited budget? an optimal design perspective. educational psychology review, 34 , 511– 536. https://doi.org/10.1007/s10648-021-09635-4 https://doi.org/10.1016/j.tate.2018.12.012 https://doi.org/10.1016/j.learninstruc.2021.101514 https://psycnet.apa.org/doi/10.1207/s15326985ep3502_2 https://psycnet.apa.org/doi/10.1037/a0025125 https://doi.org/10.1007/s11251-016-9367-z https://www.learntechlib.org/primary/p/5437/ https://doi.org/10.1007/s10648-021-09635-4 frontline learning research vol. 12 no. 3 (2024) 1 19 issn 2295-3159 corresponding author: larike bronkhorst, utrecht university, the netherlands l.h.bronkhorst@uu.nl doi: https://doi.org/10.14786/flr.v12i3.1381 learning to change the world: dis/continuity in learning across climate activism and life-wide contexts larike h. bronkhorst, marlon renes, dagmar bon, noah rookmaaker, anne van leest utrecht university, the netherlands article received 4 october 2024 / article revised 4 july 2024 / accepted 21 july/ available online 3 september 2024 abstract feeling the urgency of the climate crisis and judging current societal (re)action insufficient, young adults increasingly engage in climate activism. while individual learning is not the objective of climate activism, research has documented that young adults learn in climate activism movements. this study traces young adults’ learning across climate activism and different life-wide contexts, explicating dis/continuities in learning. content-analysis of interviews with twelve self-defined climate activists indicates that in and across climate activism and other life-contexts young adults report a) learning about the climate, activism, intersectionality, democracy and system structures, b) learning to organize, socialize and take perspective(s), while c) progressively expressing who and how they want to be(come). young adults described experiencing discontinuities between the context of their climate activism and other contexts such as education, friends and family, and their efforts to re-establish continuities are an important part of their learning. when young adults experience discontinuity across contexts structurally, they keep their climate activism to themselves and/or disengage from education, among others. making space in education more explicitly for sharing and shaping what matters to youth seems desirable. keywords: climate activism; dissent; boundary crossing; dis/continuity; learning across contexts mailto:l.h.bronkhorst@uu.nl https://doi.org/10.14786/flr.v12i3.1381 bronkhorst et al 2 | f l r 1. organizing better futures young adults are deeply concerned about their future on a planet that is progressively becoming uninhabitable (marquardt, 2020; marks et al., 2021; rajala, et al., 2023). young adults increasingly denounce climate injustices by engaging in activism: activity with the deliberate intent to change public awareness and government policies via dissent (o’brien et al., 2018). as the public opposition to official or commonly shared perspectives with the intent to change them, dissent in activism can take different shapes, including a protest marches, blocking a road or occupying a building and arguably the most well-known example of climate activism, school strikes. this study understands activism as one of the means available for young adults, as well as researchers, to engage in public debate(s) and create better futures (rajala, et al., 2023). from a learning perspective, activism entails “organizing possible futures" (uttamchandani, 2021) and constitutes learning. existing research indicates that through climate activism, young adults learn a variety of interpersonal and communication skills (curnow & jurow, 2021), and develop their identities (ginwright & james, 2002). as climate activism is still relatively controversial in european societies, young people must shape their activism mostly outside of existing structures (mcgimpsey, rousell, & howard, 2023). at the same time, climate activism plausibly differs epistemically and normatively from other life-wide contexts young adults participate in (malafaia, 2022; neas, et al., 2022). epistemically, previous research has found that in education climate change can be addressed ‘objectively’, as an abstract phenomenon damaging ecosystems in a distant future, benefitting from individual actions as suggested by professionals, if actionable solutions are addressed at all (neas, 2023; verlie & flynn, 2022). in contrast, in climate activism movements the climate crisis is understood as an urgent social and political problem affecting humans and more-than-humans, requiring urgent collective action, wherein young adults take the lead (biswas & mattheis, 2022). normatively, climate activism uniquely brings to the fore that all rules are historically negotiated and need to be renegotiated to be more just for all those inhabiting our planet, and considers civil disobedience justified to reach this end (e.g., mattheis, 2022). given these differences between climate activism and other life-wide contexts young adults participate in, young adults have to navigate between their position and participation across contexts daily. navigating these differences across contexts is challenging, but also carries learning potential (akkerman & bakker, 2011; bronkhorst & akkerman, 2016). to gain a comprehensive understanding of learning through activism, it is imperative to not isolate learning to the context of activism, but to also consider the learning resulting from dis/connections between activism and young adults’ wider lives (bronkhorst & akkerman, 2023; curnow & jurow, 2021; neas et al., 2022). this study advances such an understanding, by deliberately tracing learning through activism across all young adults’ contexts of participation. the findings concerning dis/connections to education can inform and inspire educational professionals seeking to connect to what matters to young people. 1.1 ways of learning to change the world when research attention turns to learning that takes place within social movements, the scope and scale of inquiry expands dramatically. it becomes difficult to see learning, as we usually conceive it. (erickson, 2021; 151). feeling the urgency of the climate crisis, young adults increasingly oppose existing structures by means of activism (mcgimpsey, rousell & howard, 2023). school strikes have arguably become the most well-known climate activism young people engage in, with youngsters publicly dissenting to their mandatory presence in education (belotti et al., 2022; kowasch et al., 2022). but young people’s climate activism is more diverse and can be categorized as dutiful, disruptive and/or dangerous, as well as combinations of these (o’brien, et al., 2018). when activism is voiced within existing or newly created institutional spaces, young adults resist the status quo, but essentially adhere to its script. such dutiful dissent takes shape via (new) political movements, green community activities, and stakeholder bronkhorst et al 3 | f l r meetings, for example. activism that seeks to challenge and change the existing system is referred to as disruptive activism, employing protest marches, occupying buildings, and blocking roads, to name a few examples. these examples immediately illustrate that this type of activism can include civil disobedience, or the public, non-violent, conscientious, sincerely motivated acts contrary to, but within limits of fidelity to laws (rawls, 1971 as cited in mattheis, 2022). a last type of activism is not easily recognized as such, as it aims to lead by example. so-called dangerous activism is meant to show (the viability of) other ways to organize society (o’brien, et al., 2018). for instance, by creating alternative communities and degrowth initiatives, to name a few examples. while learning in a strict sense is not the objective of climate activism, that young adults learn when trying to change the world has been documented (e.g., kirshner, 2007; hayward, 2012; verlie & flynn, 2022). through activism young adults have been observed to become more informed about social and ecological justice (mitchell, 2007), reflect on social consciousness (schlitz et al., 2010), develop social skills (kirschner, 2007), organizational skills (belotti et al., 2022) and gain increased understanding of citizenship (crouzé et al., 2023) and of social and system change (hayward, 2012). learning in climate activism is already visible when adopting a ‘conventional’ ahistorical and decontextualized lens (curnow & jurow, 2021), focussing on the expertise individuals ‘acquire’ (see sfard, 1998 for an elaboration of the acquisition learning metaphor). seeing how learning through activism takes place by collective changing and shaping of what is not yet there (curnow & jurow, 2021; uttamchandani, 2021), studying learning through activism benefits from, if not necessitates, adopting a broader understanding of learning (see also bronkhorst & akkerman, 2023). typically, authors stress the situated, dynamic, collective and/or expansive nature of learning in activism (e.g., erickson, 2021; kluttz & walter, 2018). in the special issue on social movements in the journal of the learning sciences curnow and jurow (2021) synthesize diverse views by conceptualizing learning “as possibility, as shaped by relational, spatial, and natural structures and as improvised in and around structures as resources for creative resistance and action” (p. 14-15). possibility herein refers to not just determining how to go about, but also figuring out what the object of their actions is and ought to be (merry, 2020; rajala, et al., 2023), authoring their learning. not explicitly present in these advancements in theorizing learning through climate activism is horizontal learning (engeström & sannino, 2021) by boundary crossing (akkerman & bakker, 2011; bronkhorst & akkerman, 2016). while vertical learning concerns deepening of knowledge or intensifying of participation, horizontal learning constitutes developing alternative understandings or alternatives for action, through participation in or collaboration between different contexts. by tracing the young adults’ experiences in engaging with climate activism across life-wide contexts, this study acknowledges the unique context activism provides for learning, but positions it within young adults’ wider lives, taking a person-centered perspective (akkerman & bakker, 2019). the advantage of a horizontal learning perspective is that is documents learning through activism wherever and whenever it takes place (bronkhorst & akkerman, 2023). 1.2 dis/continuities in climate activism [a] comprehensive understanding of learning in contemporary societies entails researching individuals, who are socially, culturally, and academically unique, participate in their own set of practices both in and outside of education, and face and shape undecided futures (bronkhorst & akkerman, 2023;262) horizontal perspectives on learning acknowledge that individuals participate in various lifewide contexts on a daily basis (akkerman & bakker, 2019). contexts can be understood as culturally and historically informed, progressively (re)created practices, worlds, or activity systems (engeström & sannino, 2021). contexts differ in their specific, local and routinised ways of organizing, acting, talking, and relating (bronkhorst & akkerman, 2016). climate activism can be seen as a particular context, distinguished from many other contexts in its purpose of (radically) changing existing structures, its distributed, emergent organization, and its acceptance of civil disobedience. the ways of organizing, acting, talking, and relating inherent to climate activist movements likely differ from what is ‘common bronkhorst et al 4 | f l r practice’ in other daily life contexts, such as education and work (e.g., malafaia, 2022) – contexts which in turn also differ from each other. young adults can experience such social-cultural differences between contexts as discontinuities (akkerman & bakker, 2011): temporary or structural disruptions to ongoing learning processes that necessitate shifting in ways of (inter-)acting, positioning, and being. for instance, neas (2023) described experiences of discontinuity when the climate crisis was portrayed as a problem for the distant future at school, whereas in climate activist movements the urgency of the crisis was foregrounded. malafaia (2022) documented opposition to climate activism by family members, manifesting in condescending indifference, but also insults that posed significant daily challenges to young climate activist. next to being experientially tough, discontinuity can lead to disengagement and even drop-out of formal education when discontinuity is structural and/or widespread (bronkhorst & akkerman, 2016). yet, what boundary crossing theory uniquely explicates, is how even such discontinuity does hold learning potential. the learning potential resides in the alternative understandings and alternative modes of action emerging when continuity across contexts is (re-)established (bronkhorst & akkerman, 2016). young adults can (re-)establish continuity by connecting; making sense, translating, integrating and/or introducing elements from one context into another (akkerman & bakker, 2011). as a result, an individual’s understanding of the climate crisis results from integrating (social) media, books, documentaries, personal experiences as well as education (spiteri, 2024). connections across contexts are supported by actors and (boundary) objects (akkerman & bakker, 2011; bronkhorst & akkerman, 2016), such as teachers who choose to support climate activism, even when this involves skipping school. a growing body of literature recognizes four dialogical mechanisms triggered in boundary crossing across different practices (akkerman & bakker, 2011), persons, positions (akkerman & bruining, 2016) and perspectives (wansink et al., 2023): identification, reflection, coordination, transformation. the identification and reflection mechanism focus on learning as sense-making (i.e., alternative understandings) and the coordination and transformation mechanism focus on learning as acting (i.e., alternatives for action). the identification mechanism illustrates the alternative understandings triggered by seeing a practice, person, position and/or perspective in light of another and identifying (poignant) differences. for instance, young adults realize differences in discourse about the climate, in sacrifices different people are willing to make and/or in self-positioning across different (activist) contexts. a step further, the reflection mechanism concerns alternative understandings triggered by making and taking the perspective of another practices, person, and/or position. for instance, reconsidering your understanding by taking the perspective of a government ‘failing to take action’, or reconsidering own ease with police arrests from a parental perspective. the coordination mechanism triggers alternatives for action by efficiently connecting and/or smoothly alternating between different practices, persons, positions, and perspectives. for instance, adopting an activism name to come into activist mode (and also to avoid (police) identification), or scheduling protests on a school free day. the final transformation mechanism involves crafting new, hybridized practices, persons (i.e. brokers), positions and/or perspectives, providing alternatives for action. for instance, shifting to a ‘green’ cafeteria, without single use cups and limited meat options, or identifying as activiststudent/teacher/researcher. while the coordination, reflection and transformation constitute learning by re-establish continuity across practices, the identification mechanism maintains discontinuity, by its explication of differences. 1.3 the current study by not only considering what young adults learn in climate activist movements but also across activism and other life-wide contexts, this study broadens our understanding of learning through climate activism. such understanding will benefit theorizing learning as of and for worldmaking (curnow & jurow, 2021; power et al., 2023), situated as well as distributed (melero & gil-jaurena, 2018), individual bronkhorst et al 5 | f l r as well as collective (kluttz & walker, 2018), and accumulative as well as expansive in nature (melero, & gil-jaurena, 2019). practically, the findings of this study offer insight to all those concerned with deliberately connecting to what matters to young people, including the ways young people choose to express themselves. our research question is: what and how do young adults learn through experiencing dis/continuities across climate activism and life-wide contexts? the ‘what’ and ‘how’ in this research question are deliberately open-ended (see bronkhorst & akkerman, 2023 for an argumentative grammar) to foster new understandings, although sensitized by our theoretical framework. the admittedly scarce available literature and common sense would lead to an expectation that young adults experience continuity in the cause of activism, as recognition for the necessity of climate action is becoming increasingly widespread. yet, young adults likely experience discontinuity in the way they try to achieve their cause, especially with more disruptive and dangerous climate actions, as school and other contexts ascribe to different norms (e.g., neas, 2023; mcgimpsey et al., 2023). 2. methodology the methodological approach for this study is aligned state of the literature, the object of study and the researchers’ positionality, culminating in exploratory approach with a critical stance (rajala et al., 2023). the research was approved by the ethical committee of the faculty of social and behavioral sciences at utrecht university under number 23-0126. 2.1 participants twelve young adults aged 19 to 25 voluntarily participated in this study (see table 1). invitations to participate were initially spread via the social media of diverse climate activist movements. snowballing (i.e., asking participants for potential participants who met the criteria, naderifar, et al., 2017) was additionally used to recruit a diverse sample in terms of age, gender, and interpretation and expression of climate activism (i.e., dutiful, disruptive and/or dangerous), as we postulated that experiences of dis/continuity are related to choices made within climate activism. bronkhorst et al 6 | f l r table 1. participants selfselected pseudonym age gender occupation (work / study program) activist since part of climate activism organizations characterization of dissent pepper 19 cisgender female philosophy (s) 2019 xra, end fossil occupy disruptive dangerous boas 23 cisgender male computer science (s) 2018 xr, code red dutiful disruptive jeroen 23 cisgender male communication technology (s) 2021 fossil free higher education dutiful disruptive pieter 20 cisgender male 1 2020 dutiful anne 24 transgender female writer for television (w) 2019 xr, fff, milieudefensieb dutiful disruptive paardenbloem 25 cisgender female data-management for ecological consultancy (w) 2016 dutiful disruptive fiets 25 cisgender male sustainability sciences (s) 2021 xr, end fossil occupy dutiful / disruptive madelief 24 cisgender female energy politics at the dutch embassy (w) 2020 dutiful olaf 24 cisgender male gap year 2019 xr, young climate movement, greenpeace dutiful disruptive jochem 25 cisgender male consultant sustainability (w) 2019 xr, fff dutiful / disruptive august 19 cisgender female cultural anthropology (s) 2022 xr, milieudefensieb, amelisweerd niet geasfalteerdc dutiful / disruptive lois 22 cisgender female applied biology (s) 2019 dutiful note. 1 this young adult opted to keep his study program confidential to avoid identification. xra = extinction rebellion; b translates as: environment defense; ctranslates as: parc without tarmac; fffd = friday for futures table 1 details participant characteristics, including their self-selected pseudonyms, age in years, gender identification, main occupation (in terms of their study program or description of their job), what climate organizations they are affiliated with (if any), and when they started with climate activism. some of the pseudonyms may be unfamiliar to an international audience, but we decided to maintain them, as they are self-selected activist names that the young adults volunteered as pseudonyms. last, based on their own descriptions, we typified their climate activism as dutiful (i.e., dissenting by bronkhorst et al 7 | f l r reforming from within), disruptive (i.e. dissenting by opposing the status quo, including civil disobedience), and/or dangerous (i.e. dissenting by proposing alternatives). all participants gave active consent for their interviews and received a small pocket with flower seeds as a reward. 2.2 interviews an interview protocol was created for the purposes of the study, based on the literature. although interviews were semi-structured, the questions are deliberately open-ended to allow for a natural flow of conversation, as well as unexpected answers (bronkhorst & akkerman, 2023). table 2 details the interview questions, follow-through questions and probes, in relation to the literature that inspired them. while the focus of this study lies with learning through experiencing dis/continuities, we deliberately also included questions that resonate with different conceptualizations of learning to gain a comprehensive understanding of learning. the interview scheme was piloted, resulting in adjustments to the wording and order of the questions. all interviews took place during spring of 2023 and were conducted at a place convenient for and selected by the participant, resulting in four interviews being conducted online. interviews were conducted in dutch in pairs or triads of interviewers (i.e. the second, third and/or fourth authors) and lasted about 70 minutes. 2.3 analyses interviews were transcribed verbatim and sent back to participants for member-checking; all participants verified their transcripts. after segmentation of all parts that were not relevant for answering the research question (e.g., addressing the logistics of the study; social talk), the within-case analysis (ayres et al., 2003) started with open coding, sensitized by our theoretical framework (i.e., dis/continuity across contexts, identification, coordination, reflection, transformation). in the process of open coding, the second, third and fourth author collaboratively read and re-read the accounts (birks et al., 2019), identifying all the fragments in which participants referred to (implicit) learning through activism. for each initial code, the authors developed overarching themes by iteratively regrouping the labelled fragments in meaningful ways (i.e., axial coding), as is reflected our thematic organization of what young adults learn through activism (see table 3). in the process of selective coding, overarching themes were related to and renamed according to the literature, particularly with respect to relevant life-wide contexts (informed by engeström et al, 2022; phelan, et al., 1991) and dialogical boundary crossing learning mechanisms (akkerman & bakker, 2011; akkerman & bruining, 2016; wansink et al., 2023). subsequently, for the across-case analysis (ayres et al., 2003) each participant was explored in more detail to unravel intra-individual patterns (c.f. a person-centered perspective, akkerman & bakker, 2019). here, we were particularly interested in identifying possible patterns between (re-)establishing continuity and/or discontinuity across different life-wide contexts. to that end, visualisations of each participant were made, locating dis/continuity to specific transitions between activism and family, education, work, friends and/or (student) home contexts. these visualisations were subsequently compared systematically. the data-collection and analyses decisions were documented and justified using methodological literature in an audit trail (akkerman et al., 2008). this audit trail was evaluated by an external auditor with broad expertise in educational research, who was not affiliated with this project. the visibility, comprehensibility and acceptability of all analytical procedures was assessed positively. the auditor did recommend clarifying analytical choices in the method description and results, resulting in presenting all manifestations of dis/continuity in table 4. overall, the auditor concluded that all findings are grounded in the data. bronkhorst et al 8 | f l r table 2. interview protocol topic central question follow-through, in order suitable to the conversation based on background of young adult can you introduce yourself to me? probe for age, education, job, hobbies, family and living conditions, gender identification, pseudonym akkerman & bakker, 2019; bronkhorst & akkerman, 2023 climate activism can you tell me (more) about that your climate activism? manifestation of activism, origin of/trigger for the activism, goals, and motivation o’brien et al., 2018 learning in climate activism (how) have you changed in how you engage with activism? probe for examples, perspectives of others engeström & sannino, 2021 learning in climate activism – conscious would you say that you learn through your activism? probe for examples, perspectives of others, deliberate efforts at learning, regulation of learning by movements kirshner, 2007 climate activism across contexts do people around you know about your activism? probe for anyone else phelan et al., 1991 ; engeström et al., 2022 dis/continuity across contexts for everyone mentioned previously: how do they perceive your activism? if not previously mentioned, explicitly discuss education, home, work, friends, and family akkerman & bakker, 2011; bronkhorst & akkerman, 2016 discontinuity for significant others who don’t know about activism: (what) would you like them to know about your activism? probe for reasons akkerman & bakker, 2011; bronkhorst & akkerman, 2016 reflection is there anything else that is important to know about your activism? possibility to revisit earlier answers 3. findings all young adults engage in climate activism because of their belief that without change, a viable future is impossible, which, coupled with their wish to take responsibility, results in their dissent. young adults define the goal of their activism – in line with the literature – as “to move against the established order and to want to change things” (jeroen). in describing themselves as activist, their emphasis in activity lies in creating awareness: “the climate activist is someone who chooses to break out of their bronkhorst et al 9 | f l r normal routine to raise awareness of climate issues among other people” (jochem). olaf stresses how becoming and being a climate activist need not involving othering: "by becoming an activist yourself and then also immediately taking part in actions…which others don’t want to get near to, namely civil disobedience… [this] did open my eyes to that: ‘you can do just do this’. it is not some kind of offshoot of homo sapiens, the activist, but it is people like you and me." concrete actions that young adults undertake vary, in line with the variety stressed by the typology of dissent mentioned in the literature (see o’brien et al., 2018): “i naturally think climate activism is just really ridiculously broad. to me it can really include petitions, marches or well such disruptive actions that extinction rebellion does, things like eco-sabotage for example, that all falls under climate activism” (pepper). multiple young adults emphasize the importance of caring for each other in activism and “having each other’s backs”, resulting in climate activism “bringing more peace than stress or fear”. six young adults mention explicitly that they do not engage in activism to ‘learn’: “i don’t exactly do it to develop myself more or learn more. it’s more because of the experience. it is a fun and exciting idea that you learn new things from new experiences” (fiets). in what follows, the what and how of young adults’ learning through activism is identified and detailed. 3.1 what is learned through activism learning through activism varies across young adults, although there appear to be some commonalities. all young adults describe learning about different topics related to climate activism. all young adults except one describe to also learn to organize, socialize, and/or formulate perspectives. nine young adults also describe learning processes typically considered as identity development, in terms of learning to express and reflect on oneself (i.e. being and becoming). table 3 summarizes the commonalities in learning through activism reported by the participating young adults, including quotes. although eight participants reported to have attended mandatory trainings offered by activist movements, the young adults contend that most of their learning through activism was not regulated. instead, for the young adults involved, learning emerged over time by collaboratively working on instigating change: “it just feels that i only have an influence when i work with people. and as a result we have a good time together, accomplish things together, set up concepts together, move forward with our plans together.” (fiets) in line with the literature on heterogenous coalitions (engeström & sannino, 2023) in their descriptions of the activism the young adults describe a decentralized organization, and in turn an openness for emerging action(s): “if you and a group of others want to organize an activity, you can. the group makes a plan, and a set-up for the activity and often also the theme behind it” (boas). in young adults’ accounts fluid and limitedly bounded arrangements of activism comes to the fore, identifying the overarching object of climate action as providing the glue: “i’m sure i don’t agree with all, or some of the people at christian climate action, but we are fighting for the same cause by the same means” (jochem). similarly, jeroen’s descriptions of arrests during action echoes this collectivity: “then they [i.e. the police] have an individual, but they can’t shut down the movement”. while climate activism might hence not be considered a context in the sense of well-defined space or single community, all young adults do speak of their activism in unified terms: “then you step outside your bubble again and you notice that there are also people who really don’t give a fuck about the climate.” (paardenbloem). paardenbloem’s observation about people outside her ‘bubble’ invites us to explore how learning through activism involves dis/continuity across contexts. bronkhorst et al 10 | f l r table 3. learning through climate activism learning example quote reported by (n) about (12) activism “you also learn a lot about activism then, and that’s kind of a conscious choice to some extent. by participating and getting involved, you also learn to understand it.” (pieter) pepper, pieter, fiets, olaf, august (5) intersectionalit y “where i used to be focused specifically on climate problems, now i do have a much broader view. say, the asylum crisis and housing crisis and all those kinds of things, that because they’re all so related.” (jochem) pepper, boas, jeroen, anne, jochem, august, lois (7) systems “i think i’ve just become much more aware at all of how the system works and how it can be harmful.” (pepper) pepper, jeroen, jochem (3) climate “certainly, also about the facets that you are less aware of in everyday life, such as climate justice, or how it affects people in the global south. that knowledge will come anyway.” (anne) pieter, anne, jochem, august, lois (5) democracy and justice system “i especially learned a lot about the political process and climate politics, i think, through my activism, also about how democracy works.” (madelief) pepper, pieter, anne, paardenbloem, fiets, madelief, august (7) to (11) organize “through climate activism, i got involved in organizing activities and organizing just taught me a lot. just really professional experience.” (madelief) pepper, jeroen, paardenbloem, fiets, madelief, anne (6) socialize “i think in a way you kind of learn social skills. it’s always a very nice atmosphere and it’s very acceptable to just approach people you don’t know and start a conversation with them.” (boas) pepper, boas, jeroen, pieter, anne, paardenbloem, madelief, olaf (8) formulate perspective(s) “in the beginning you could say i was very radical, really straight from hard facts, but that doesn’t work at all. so yes, what does work is trying to explain to people how things can be done differently without suggesting that they should do that.” (jeroen) pepper, jeroen, pieter, anne, fiets, madelief, olaf, jochem, august (9) be (express oneself) “so in that respect i feel i have learned to stand a bit stronger and to have more of an opinion.” (august) jeroen, pieter, paardenbloem, madelief, olaf, anne, august (7) become (reflect on oneself) “in fact, i have come to realize more and more that it is so important to me that [climate activism] is part of my identity.” (madelief) anne, fiets, madelief, august (4) bronkhorst et al 11 | f l r 3.2 how learning emerges across contexts all young adults reported (re-establishing) continuity, particularly across activism and friends and family (see also table 4). young adults mention feeling able to talk about the climate and their concrete activism plans in other contexts (i.e., the learning mechanism coordination, akkerman & bakker, 2011). jochem shares: “we really talk about it actively. they really ask questions about it, which gives me the idea that they find it very interesting”. continuity is also reflected in young adults recognizing interest in and appreciation for their climate activism in significant others around them (i.e. reflection, akkerman & bakker, 2011): “i also give regular updates in the family [app] group when i am somewhere, and i always get a response to that as well. often with a lot of positive comments” (boas). young adults also identified such shared interest in their upbringing. for instance, olaf reported: “i've always heard a lot about climate change and sustainability. my parents, especially my father, he kind of instilled that from an early age”. some parents have also been involved in some kind of (climate) activism themselves. the young adults also report to have space to pursue activism in other contexts (i.e., transformation of practices, akkerman & bakker, 2011). for instance, in education they appreciate space in dedicating assignments climate (activism), by writing a paper or blog article about a related topic. they also appreciate (significant) others, such as friends, family and even teachers, joining activist outings. for instance, pepper mentions how highway roadblocks are becoming a family activity: “my mother, my little brother and my niece and my grandmother have all joined me in protests of xr so far”. the continuity is reflected in extended learning about topics (e.g., the climate and politics) across activism, work and study contexts, and young adults’ understanding of themselves (i.e. transformation of perspectives and positions, akkerman & bakker, 2011; wansink et al., 2023). anne describes knowing herself better: "this is how i react to stress [..] my perception of things, my behavior, changes because of that. that's all a lot of discovery about myself and now i'm trying to apply that in how i engage in activism". young adults explicitly mention continuity in learning about particular topics (e.g., relating their study program and social media to their understanding of sustainability), but it can also be recognized in their discourse. for instance, pieter’s description of activism echoes his predating knowledge of organizational structures: the [activist] movements don't legally exist. they don't exist on paper; they don't have registrations as associations. you can also see that in the way those movements approach and organize things: very democratic and flat and not really with functions and hierarchy. moreover, some young adults also infer how their activism has impacted others: “i've convinced a lot of people in my immediate environment, to at least be a little more conscious. and i've created awareness around me by talking to friends and saying i'm going to do this” (boas). all participants but one also learn trough experiencing discontinuity (see also table 4). a first manifestation of discontinuity is when climate activism or even the climate is ‘absent’ in other contexts: “my studies are very much focused on computer science, focusing on computers, how it works to write a program. that's just a very different field. and that's very hard to deal with then. the climate just doesn't come into play” (boas). young adults also describe not feeling able to share or pursue the topic of climate nor climate activism in other contexts, recognizing no or overtly negative responses to climate activism, or people from other contexts not being open to changing their behavior. for instance, jeroen shares about his parents: “although sometimes it does feel limited, because i would like to do much more myself. ideally, i wouldn't go to a regular supermarket anymore either, but of course i can't force my parents to go to an ecoplaza1”. 1 sustainable supermarket bronkhorst et al 12 | f l r a few young adults express overt appreciation for the learning potential of discontinuity, particularly in discussions: “but also with people with different opinions, there is usually room for that as well, and i also find it interesting to talk about my political views on this as well” (pieter). however, most young adults’ responses to an experienced lack of space for their activism (i.e. discontinuity) include disengagement and withdrawal from formal and informal contexts (e.g., calling in sick to go to a protest; ending friendships) and keeping their activism to themselves. lois shares how she avoids or cuts off conversation about activism, and how she has also given up on talking to the farming side of her family: “it’s too far apart to get on the same page. we have too little knowledge of each other's fields to get in a direction”. table 4. manifestations of dis/continuity across climate activism and family, education, home, friends, and work context dis/continuity manifestation reported by (n) f am il y continuity (possibility for) sharing climate actions all (12) understanding and/or appreciation for climate actions pepper, boas, jeroen, pieter, anne, fiets, jochem, august, lois (9). climate education and/or activism was part of upbringing pepper, anne, olaf (3). family joins climate actions pepper, paardenbloem, august (3). discontinuity family disagrees with or fears for arrests boas, fiets, jochem, august (4). deliberately not sharing to keep the peace pepper, anne, paardenbloem, lois (4). (financial, physical) limits to climate activism jeroen, jochem, lois (3). s ch o o l / ed u ca ti o n continuity climate (activism) is or can be a part of assignments pepper, boas, jeroen, pieter, paardenbloem, madelief, august, lois (9). sharing perspective on climate activism pepper, pieter, lois (3). study peers and teachers join climate actions pepper, jeroen, olaf (3). discontinuity study-peers express disapproval of climate activism pepper, fiets, olaf, lois (4). climate (activism) is not a topic pepper, boas, jeroen (3). (s tu d en t) h o m e continuity space for climate activism anne, paardenbloem, fiets (3). discontinuity climate activism is not discussed with roommates pieter, anne (2). f ri en d s continuity friends join climate actions pepper, boas, jeroen, pieter, anne, paardenbloem, fiets, olaf, jochem, august (10). friends express support and/or appreciation for climate activism pepper, boas, jeroen, pieter, anne, fiets, jochem, august (8). bronkhorst et al 13 | f l r talking about climate activism is an option pepper, boas, jeroen, paardenbloem, olaf, jochem, august, lois (8). discontinuity friends disagree with ways of climate activism boas, jeroen, pieter, olaf, jochem, august (6). friends are not as engaged in climate activism boas, fiets, madelief, lois (4). w o rk (i n cl . in te rn sh ip ) continuity working on a climate issue jeroen, madelief, olaf, lois (4). taking to colleagues about climate activism anne, paardenbloem, olaf, jochem (4). discontinuity deliberate caution with sharing jeroen, paardenbloem, lois (3). such experiences of discontinuity can be incidental and to some extent fleeting, but also structural. august mentions irritation about his mothers’ cautions against getting arrested: "in itself, i like the fact that she points this out to me, that i'm not going to think [getting arrested] is normal ..., but when she keeps saying 'don't get arrested,' that's also a little irritating". attending to these discontinuities actually appears crucial for young adults’ learning to take-perspective(s), to socialize, and to progressively express and reflect on who they want to be(come). this learning extending across activism, school, and work, is described as a continuous careful treading by madelief: “how can i stay critical without shortchanging my colleagues, or myself, [and] missing opportunities?”. we should note that five young adults mention preferring discontinuity or some difference and distance between activism and their wider lives (see also intended discontinuity, bronkhorst & akkerman, 2016). for example, they report needing or preferring to keep work and private life separate from their climate activism and/or focus solely on their study. jeroen, for instance, wants to avoid a possible impact on how he will be seen as co-worker. maintaining discontinuity also helps young adults to not always be preoccupied with the climate, which supports their mental health, and their ability to maintain a diverse and/or longstanding group of friends, which would otherwise become difficult. the cross-case analysis brought interesting and at first glance counterintuitive findings to the fore. namely that the only participant whose activism also takes shape in actions aiming to “completely destroy the system” actually experiences continuity across all daily life contexts. he only reports incidental discontinuity at his educational program, where peers are mainly preoccupied with other things. in turn, young adults merely engaging in what could be seen as dutiful dissent, do not experience continuity across all life-wide contexts. in contrast, they report feeling challenged by the discontinuity experienced with friends and family, and at work. madelief aptly typifies these encounters: "so i also feel lonely sometimes, because within my circles of friends there are a lot of people who are environmentally conscious, but activist on the level that i am, most are not”. 4. discussion feeling the urgency of the climate crisis, young adults increasingly dissent the status quo by means of climate activism. climate activism has been found to be a context wherein learning emerges. yet, thus far, research into (climate) activism has documented the expansive, collaborative learning taking place in activist movements (e.g., belotti et al., 2022). this explorative study aimed to advance a more comprehensive understanding of learning through activism. we gained a life-wide picture of young adults’ learning to change the world by tracing what and how young adults learn across activism and different life-wide contexts, explicating (un)intended dis/continuities in learning. bronkhorst et al 14 | f l r content-analysis of twelve interviews with self-defined climate activists echoes activism as a rich context for learning, which table 3 details with examples. specifically, first, through climate activism all young adults describe to learn about the various topics they relate to their climate activism, including the climate, activism, intersectionality, democracy and justice and system structures. although the young adults primarily ascribe these topics to engaging in activism, their descriptions indicate an integration across contexts, signalled by comparisons and describing a broadened, or alternative understanding of these topics. second, all but one young adult also describes learning how to organize, socialize and take perspective(s) by engaging in activism and relating those experiences to other lifewide participations. particularly, the caring or ‘regenerative’ social conventions in activism contrasts with socializing elsewhere and invites invoking these conventions elsewhere (see also rowe & ormond, 2023). also, the precarious position of activism in mainstream society coupled with young adults’ wish to convince others to act (see also kutlaca et al., 2020), creates a unique opportunity to learn to take and learn from taking perspectives in socializing with others. last, in engaging with activism in and across contexts, nine young adults also report to learn to progressively express who and how they want to be(come). we should note that most of these learning outcomes are highly valued and sought after, but not easily realized in ‘mainstream’ education (crouzé et al., 2023; neas, 2023; rajala, et al., 2023). even citizenship education, specifically designed to encourage active citizenship early on, has difficulty engaging students in learning about the democratic means of authoring their own futures (merry, 2020). it could read as ironic that climate activism, typically interpreted as misbehaviour (mattheis, 2023), does manage to trigger such learning. while these findings about young adults’ learning through climate activism are mostly in line with the existing literature on learning through activism (e.g., kirshner, 2007; hayward, 2012; mitchell, 2007; schlitz, vieten & miller, 2010; verlie & flynn, 2022), the added value of this paper lies in tracing that learning across activism and other life-contexts. across contexts different dis/continuities were reported (see table 4 for all manifestations of dis/continuity). the different dialogical learning mechanisms (akkerman & bakker, 2011; akkerman & bruining, 2016; wansink et al., 2023) can be identified in young adults’ reports of discontinuity, via the identification mechanism and (re-)established continuity via coordination, reflection, and/or transformation mechanisms. complementing existing research that documented adolescents’ discontinuities between school and activism (malafaia, 2022; neas, et al., 2022), the young adults in this study also reported (re-)establishing continuity between their activism and their education by dedicating assignments (including internships) to their activism. plausibly, higher education affords more degrees of freedom to (re-)establish continuity across contexts (bronkhorst & akkerman, 2016). temporary discontinuities and efforts to (re-)establish continuity appeared crucial to young adults’ learning through activism. yet, our appreciation for the learning potential of discontinuity should not be mistaken for a lack of recognition of what young adults experience when they feel they cannot share what matters most to them (i.e. their climate activism). when young adults experience overt resistance to or structural denial of the climate crisis at home, at work, at their educational program and/or with friends, they not only recognize a challenge to their world view, but to who they are. notably, in some cases, the young adults resigned or choose to maintain disconnections between activism and their wider lives. this coping by keeping activism to themselves and/or structurally disengaging from other contexts (i.e., intended discontinuity, bronkhorst & akkerman, 2016) likely constrains their learning and their participation in other life-wide contexts, including school (bronkhorst et al., 2013). in line with boundary crossing theory (akkerman & bakker, 2011; bronkhorst & akkerman, 2016; 2023), each young adult’s experiences were unique, as their activism and relevant life-wide contexts differed, as did their experiences of dis/continuity. for instance, friends having a different perspective on the climate crisis could be experienced as a struggle, or as a cherished the space for discussion, re-establishing continuity by means of the coordination mechanism. particularly, the young adult who engaged in disruptive and dangerous dissent – and thus used means of activism and civil disobedience that are likely far from what is considered acceptable in mainstream society experienced bronkhorst et al 15 | f l r continuity across every life-wide context. this appears related to the life-worlds sought out and created by this young adult, who has opted to end friendships with non-activists. in contrast, young adults engaging in dutiful dissent within the confinements of the law reported more discontinuity, as they engaged with more ‘diverse’ peer-groups in terms of their stance towards climate activism and climate change. these findings highlight the importance of more person-centered research (akkerman & bakker, 2019; bronkhorst & akkerman, 2023), seeking to understand each individual as unique in their life-wide participations and prospects. yet, at the same time these findings remind us of that learning in progressive activist movements bears resemblance with processes in regressive movements (erickson, 2021). information bubbles, echo chambers and retreat from ‘mainstream’ society are risks we could and perhaps should avoid by continuing to seek open dialogue in education. 4.1 limitations and further research the findings should be interpreted considering the following limitations. first, findings are based on a small sample, which may be biased by the snowball method used for sampling (naderifar et al., 2022). perhaps partly as a result, only one participant was involved in dissent that can be typified as dangerous and a relatively large portion of young adults were involved in (different branches of) extinction rebellion. although dangerous dissent is also least common and extinction rebellion can best be seen as a heterogeneous coalition in itself (engeström & sannino, 2021), future research should seek to increase (the diversity in) their sample. particularly interesting would be to explore younger climate activists, who are not (yet) allowed to vote and thus have limited legal means to change the system from within, to explore the degrees of freedom in their schooling to (re-)establish continuity. comparing the learning reported by activists seeking climate justice and, for instance, intersectional justice or peace, would also be an interesting avenue for further research. second, this study should be understood as uncovering but the tip of the iceberg. learning through activism across contexts is dispersed, multifaceted, multi-voiced and complex, and therefore challenging to document. erickson (2021) compared locating learning in activism to finding the character waldo (or "wally" in england) in the drawings of different gatherings in the waldo books. to trace learning across contexts, in this study we relied on young adults’ self-reports via interviews, following other articles studying learning across contexts in which this method was informative (see bronkhorst & akkerman, 2023; engeström, et al., 2022), especially when adopting a person-centered lens (akkerman & bakker, 2019). however, it is likely that young adults kept part of their activism to themselves, as sharing information about past and/or future civil disobedience can be risky, and activist are trained in keeping parts of their activity secret. a next step could be to include participant observations in different settings and/or interview actors from different life-wide contexts. by including check-ins with the young adults at multiple time points (as also suggested in bronkhorst & akkerman, 2023) future research could uncover to what extent and in what ways learning develops over time, along with changes in the individual young adults’ lives, their climate activist movements and society at large. 4.2 implications the following implications should be read as our position on the role of our research in and for world-making (power et al., 2023; rajala, et al., 2023). our findings call attention to the varied and rich learning taking place through activism. our findings illustrate how learning in climate activism takes shape by negotiating, dissenting, and recreating – abilities most urgent societal challenges necessitate (rajala, et al., 2023), but education does not always foreground. as such, while educators and parents may disagree with some of the means of activism, especially civil disobedience, our findings encourage all of us to acknowledge climate activism, in all its forms, as potentially educative. particularly climate and citizenship education (merry, 2020; crouzé et al., 2023) could benefit from capitalizing on these inherent qualities of climate activism. bronkhorst et al 16 | f l r in this respect, it is worth emphasizing that learning is seen by young adults themselves as a ‘side-effect’ of activism. we know from previous research that contexts can lose their value for young adults when learning is foregrounded (i.e., when informal contexts become ‘educationalized’, see bronkhorst & akkerman, 2016). as such, our documentation and valuation of this side-effect of activism should not be seen as an invitation to lay out an educational structure over activist movements if that would be possible. instead, it should be read as an invitation to dialogue with young adults about their engagement in climate activism, as well as other life-wide contexts wherein learning might not be immediately recognizable. such dialogue likely opens up possibilities for creating meaningful connections, adding to the relevance of education (bronkhorst & akkerman, 2016; verlie & flynn, 2022). in doing so, language to talk about learning is crucial, as learning and/or teaching in activism may not be immediately recognizable to a traditional lens (erickson, 2021). apart from mandatory trainings before disruptive actions, learning in activism is not so much regulated or orchestrated upfront, but instead emerges in and across heterogeneous coactions of activity (engeström & sannino, 2021). our findings also call for action. we concur that: [...] there is a need to look beyond these high profile activities to understand youth concerns and responses to environmental concerns. for most young people, striking is an occasional form of high-profile activism, and yet around the world, young people – including those who may not have access to an organised strike – are deeply concerned about environmental hazards and degradation and the inaction of political leaders on these concerns.” (walker, 2020;11). our findings suggest that young adults could benefit from our support in navigating discontinuities in their daily lives. particularly, the paradox in their accounts of desperately wanting to change the world for the better, but not wanting to burden others in their immediate environment, carefully considering what and when they share what matters most to them. seeing the urgency of the climate crisis, it is vital to take climate activism out of the ‘doghouse’ and welcome young adults to share everything they do for our planet. this may also prevent young adults from disengaging from lifewide contexts, including education, when they are unable to resolve structural discontinuity. all in all, more explicit space in education and beyond for sharing and shaping what matters to youth seems desirable. key points a learning across contexts perspective is adopted to advance our understanding of learning through climate activism by (re-)establishing continuity across contexts, young adults learn about the climate, activism, politics, intersectionality, democracy, and power structures young adults also learn to organize, socialize, take perspective(s), progressively expressing who and how they want to be(come) incidental discontinuity across contexts can trigger learning structural discontinuity can lead to withdrawal and disengagement acknowledgments we are grateful to all young adults who volunteered to participate in this research for their dedication. bronkhorst et al 17 | f l r references akkerman, s., admiraal, w., brekelmans, m., & oost, h. (2008). auditing quality of research in social sciences. quality & quantity, 42(2), 257–274. https://doi.org/10.1007/s11135-0069044-4 akkerman, s., & bakker, a. (2011). boundary crossing and boundary objects. review of educational research, 81(2), 132–169. https://doi.org/10.3102/0034654311404435 akkerman, s. f., & bakker, a. (2019). persons pursuing multiple objects of interest in multiple contexts. european journal of psychology of education, 34(1), 1-24. https://doi.org/10.1007/s10212-018-0400-2 ayres, l., kavanaugh, k., & knafl, k. a. (2003). within-case and across-case approaches to qualitative data analysis. qualitative health research, 13(6), 871–883. https://doi.org/10.1177/1049732303013006008 belotti, f., donato, s., bussoletti, a., & comunello, f. (2022). youth activism for climate on and beyond social media: insights from fridaysforfuture-rome. the international journal of press/politics, 27(3), 718–737. https://doi.org/10.1177/19401612211072776 birks, m., hoare, k., & mills, j. (2019). grounded theory: the faqs. international journal of qualitative methods, 18, 160940691988253. https://doi.org/10.1177/1609406919882535 biswas, t., & mattheis, n. (2022). strikingly educational: a childist perspective on children’s civil disobedience for climate justice. educational philosophy and theory, 54(2), 145-157. https://doi.org/10.1080/00131857.2021.1880390 bronkhorst, l. h., & akkerman, s. f. (2016). at the boundary of school: continuity and discontinuity in learning across contexts. educational research review, 19, 18–35. https://doi.org/10.1016/j.edurev.2016.04.001 bronkhorst, l.h. & akkerman, s.f. (2023). chapter 16: hybridizing in and out-of-school learning. in c. damsa, a. rajala, g. ritella and j. brower (eds): re-theorizing learning and research methods in learning research pp. 259-275. https://doi.org/10.4324/9781003205838-17 bronkhorst, l. h., koster, b., meijer, p. c., woldman, n., & vermunt, j. d. (2014). exploring student teachers' resistance to teacher education pedagogies. teaching and teacher education, 40, 73-82. https://doi.org/10.1016/j.tate.2014.02.001 crouzé, r., godard, l., & meurs, p. (2023). learning democracy through activism: the global climate strike movement and belgian youth’s democratic experience in times of environmental emergency. social movement studies, 23(1), 56–71. https://doi.org/10.1080/14742837.2023.2184792 curnow, j., & jurow, a. s. (2021). learning in and for collective action. the journal of the learning sciences, 30(1), 14–26. https://doi.org/10.1080/10508406.2021.1880189 engeström, y., rantavuori, p., ruutu, p., & tapola-haapala, m. (2022). the hybridisation of young adults’ worlds as a source of developmental tensions: a study of discursive manifestations of contradictions. educational review, 1–22. https://doi.org/10.1080/00131911.2022.2033704 engeström, y., & sannino, a. (2021). from mediated actions to heterogenous coalitions: four generations of activity-theoretical studies of work and learning. mind, culture, and activity, 28(1), 4-23. https://doi.org/10.1080/10749039.2020.1806328 erickson, f. (2021). waldo, blind inquirers, and an elephant: locating learning in social movements. journal of the learning sciences, 30(1), 151–161. https://doi.org/10.1080/10508406.2021.1880190 ginwright, s., & james, t. (2002). from assets to agents of change: social justice, organizing, and youth development. new directions for youth development, 2002(96), 27-46. https://doi.org/10.1002/yd.25 kirshner, b. (2007). introduction. american behavioral scientist, 51(3), 367–379. https://doi.org/10.1177/0002764207306065 hayward, b. (2012). children, citizenship and environment: nurturing a democratic imagination in a changing world. routledge. https://doi.org/10.4324/9780203106839 https://doi.org/10.1007/s11135-006-9044-4 https://doi.org/10.1007/s11135-006-9044-4 https://doi.org/10.3102/0034654311404435 https://doi.org/10.1007/s10212-018-0400-2 https://doi.org/10.1177/1049732303013006008 https://doi.org/10.1177/19401612211072776 https://doi.org/10.1177/1609406919882535 https://doi.org/10.1080/00131857.2021.1880390 https://doi.org/10.1016/j.edurev.2016.04.001 https://doi.org/10.4324/9781003205838-17 https://doi.org/10.1016/j.tate.2014.02.001 https://doi.org/10.1080/14742837.2023.2184792 https://doi.org/10.1080/10508406.2021.1880189 https://doi.org/10.1080/00131911.2022.2033704 https://doi.org/10.1080/10749039.2020.1806328 https://doi.org/10.1080/10508406.2021.1880190 https://doi.org/10.1002/yd.25 https://doi.org/10.1177/0002764207306065 https://doi.org/10.4324/9780203106839 bronkhorst et al 18 | f l r kirshner, b. (2007). introduction: youth activism as a context for learning and development. american behavioral scientist, 51(3), 367-379. https://doi.org/10.1177/0002764207306065 kluttz, j., & walter, p. (2018). conceptualizing learning in the climate justice movement. adult education quarterly, 68(2), 91-107. https://doi.org/10.1177/0741713617751043 kowasch, m., & reis, p. g. n. k. k. (2021). climate youth activism initiatives: motivations and aims, and the potential to integrate climate activism into esd and transformative learning. sustainability, 13(21), 11581. https://doi.org/10.3390/su132111581 kutlaca, m., van zomeren, m., & epstude, k. (2020). friends or foes? how activists and non-activists perceive and evaluate each other. plos one, 15(4), e0230918. https://doi.org/10.1371/journal.pone.0230918 malafaia, c. (2022). ‘missing school isn’t the end of the world (actually, it might prevent it)’: climate activists resisting adult power, repurposing privileges and reframing education. ethnography and education, 17(4), 421–440. https://doi.org/10.1080/17457823.2022.2123248 mattheis, n. (2022). unruly kids? conceptualizing and defending youth disobedience. european journal of political theory, 21(3), 466-490. https://doi.org/10.1177/1474885120918371 marquardt, j. u. (2020). fridays for future’s disruptive potential: an inconvenient youth between moderate and radical ideas. frontiers in communication, 5. https://doi.org/10.3389/fcomm.2020.00048 mcgimpsey, i., rousell, d., & howard, f. (2023). a double bind: youth activism, climate change, and education. educational review, 75(1), 1–8. https://doi.org/10.1080/00131911.2022.2119021 merry, m. s. (2020). can schools teach citizenship? discourse: studies in the cultural politics of education, 41(1), 124-138. https://doi.org/10.1080/01596306.2018.1488242 mitchell, t. d. (2007). critical service-learning as social justice education: a case study of the citizen scholars program. equity & excellence in education 40(2), 101-112. https://doi.org/10.1080/10665680701228797 naderifar, m., goli, h., & ghaljaie, f. (2017). snowball sampling: a purposeful method of sampling in qualitative research. strides in development of medical education, 14(3). https://doi.org/10.5812/sdme.67670 neas, s. (2023). narratives and impacts of formal climate education experienced by young climate activists. environmental education research, 29(12), 1832-1848. https://doi.org/10.1111/chso.12846 neas, s., ward, a., & bowman, b. (2022). young people's climate activism: a review of the literature. frontiers in political science, 4, 940876. https://doi.org/10.3389/fpos.2022.940876 o’brien, k. f. a. r., selboe, e., & hayward, b. (2018). exploring youth activism on climate change: dutiful, disruptive, and dangerous dissent. ecology and society, 23(3). https://doi.org/10.5751/es-10287-230342 phelan, p., davidson, a. l., & cao, h. t. (1991). students' multiple worlds: negotiating the boundaries of family, peer, and school cultures. anthropology & education quarterly, 22(3), 224-250. https://doi.org/10.1525/aeq.1991.22.3.05x1051k power, a., zittoun, t., akkerman, s., wagoner, b., cabra, m., cornish, f., hawlina, h., heasman, b., mahendran, k., psaltis, c., rajala, a. & gillespie, a. (2023). social psychology of and for world-making. personality and social psychology review. https://doi.org/10.1177/10888683221145756 rajala, a., jornet, a.g., & accioly, i. (2023). chapter 19: utopian methodologies to address the social and ecological crises through educational research. in c. damsa, a. rajala, g. ritella and j.brower (eds): re-theorizing learning and research methods in learning research pp. 298312. https://doi.org/10.4324/9781003205838-20 rowe, t., & ormond, m. (2023). holding space for climate justice? urgency and ‘regenerative cultures’ in extinction rebellion netherlands. geoforum, 146, 103868. https://doi.org/10.1016/j.geoforum.2023.103868 sfard, a. (1998). on two metaphors for learning and the dangers of choosing just one. educational researcher, 27(2), 4-13. https://doi.org/10.3102/0013189x027002004 https://doi.org/10.1177/0002764207306065 https://doi.org/10.1177/0741713617751043 https://doi.org/10.3390/su132111581 https://doi.org/10.1371/journal.pone.0230918 https://doi.org/10.1080/17457823.2022.2123248 https://doi.org/10.1177/1474885120918371 https://doi.org/10.3389/fcomm.2020.00048 https://doi.org/10.1080/00131911.2022.2119021 https://doi.org/10.1080/01596306.2018.1488242 https://doi.org/10.1080/10665680701228797 https://doi.org/10.5812/sdme.67670 https://doi.org/10.1111/chso.12846 https://doi.org/10.3389/fpos.2022.940876 https://doi.org/10.5751/es-10287-230342 https://psycnet.apa.org/doi/10.1525/aeq.1991.22.3.05x1051k https://doi.org/10.1177/10888683221145756 https://doi.org/10.4324/9781003205838-20 https://doi.org/10.1016/j.geoforum.2023.103868 https://doi.org/10.3102/0013189x027002004 bronkhorst et al 19 | f l r schlitz, m. m., vieten, c., & miller, e. m. (2010). worldview transformation and the development of social consciousness. journal of consciousness studies, 17(7-8), 18-36. spiteri, j. (2024). pre-service ecec teachers’ conceptions of climate change: a community funds of knowledge and identity approach. education 3-13, 1-15. https://doi.org/10.1080/03004279.2024.2322999 uttamchandani, s. (2021). educational intimacy: learning, prefiguration, and relationships in an lgbtq+ youth group’s advocacy efforts. the journal of the learning sciences, 30(1), 52– 75. https://doi.org/10.1080/10508406.2020.1821202 verlie, b., & flynn, a. (2022). school strike for climate: a reckoning for education. australian journal of environmental education, 38(1), 1-12. https://doi.org/10.1017/aee.2022.5 walker, c. (2020). uneven solidarity: the school strikes for climate in global and intergenerational perspective. sustainable earth, 3, 1-13. https://doi.org/10.1186/s42055-020-00024-3 wansink, b. g. j., timmer, j., & bronkhorst, l. h. (2023). navigating multiple perspectives in discussing controversial topics: boundary crossing in the classroom. education sciences, 13(9), 938. https://doi.org/10.3390/educsci13090938 https://doi.org/10.1080/03004279.2024.2322999 https://doi.org/10.1080/10508406.2020.1821202 https://doi.org/10.1017/aee.2022.5 https://doi.org/10.1186/s42055-020-00024-3 https://doi.org/10.3390/educsci13090938 frontline learning research vol. 13 no. 1 (2025) 76 83 issn 2295-3159 corresponding author: lukefryer@yahoo.com; the university of hong kong, faculty of education, talic. doi: https://doi.org/10.14786/flr.v13i1.1523 a psychological platform for genai and human co-piloting in education luke k. fryer the university of hong kong, hong kong article received 16 june 2024/ article revised 12 march 2025 / accepted 19 march 2025/ available online 4 april 2025 abstract genai (generative artificial intelligence) will have a growing role within formal education. what should that role be? how do we treat genais as an opportunity to enhance and reenergise teaching and learning? this theoretical article suggests that answers to these questions should start with our foundational psychological theories about what students need to function and develop well in educational environments. the article outlines how psychological needs theory, focusing on students' basic psychological needs for competence and relatedness might be a path forward. teacher behavior supporting these psychological needs (i.e., involvement and structure), which have established relationships with learning outcomes, are used as a base for discussing the potential roles of human and ai instructors. a co-piloting model that draws on the strengths of each instructor is framed and suggested as a possible way forward for research and practice in this area. further research is needed to continue to develop this initial theoretical framing of co-piloting and assess if and how it can contribute to herald a brighter future for students across educational levels and contexts. keywords: genai; psychological needs; co-piloting; structure; involvement fryer 77 | f l r 1. introduction as genais surge into nearly every area of human endeavour (kaplan, 2024), the question has become what genais cannot do, not what they can (jiang et al., 2022). one presumed area of eventual adoption is education (chen et al., 2022). on its face, this premise seems reasonable given the fact that genais have already demonstrated that they are able to manage more information than any living human (wenzlaff et al., 2022). they could become tireless teachers that could answer every query a student might have. isn't this the teacher that humankind has waited for? this is most likely a question that humankind will be considering for decades to come. while instructional frameworks have been posited for employing ai in education (e.g., holstein et al., 2020), from the perspective presented in this article, they tend to miss a critical step in answering two questions: a. how do humans learn and then b. how teaching/teachers best support the learning process? humans are remarkable creatures, with substantial ability to exceed their physiological limitations, adapting to environments far beyond their intended evolutionary settings (moran, 2022). however, there are fundamental constraints on both humankind's physiology and psychology that must be taken into consideration when seeking to support students in effective learning (sweller & chandler, 1991). this article begins by introducing some longstanding psychological theories which provide direction for addressing psychological constraints on learning and thus have the potential to support educators as they seek to integrate rapidly developing genai agents with the complexities of classroom learning. co-piloting is put forth as a practical frame for discussing the application of psychological theory to the issues put forward. this is followed by the application of two theoretical perspectives (competence and relatedness psychological needs) on learning and student-teacher classroom interaction. the article concludes with some preliminary suggestions for how ai and human-teachers might work together effectively (i.e., co-pilot classroom learning) to ensure students the world over reap the benefits of the coming age of ai. 2. background 2.1 constraints on learning and guiding psychology theory to make clear what is meant by a constraint on learning, a biological constraint is a good place to start. an example of a biological-psychological constraint that ai instructors might be in a strong position to address is cognitive load. cognitive load theory's (sweller & chandler, 1991) main and well-established contention is that humankind’s cognitive architecture has limited short-term memory space. furthermore, this limited capacity can be further constrained by the way in which information is presented for learning. this has a plethora of implications for how instructional materials are organised and presented (e.g., sweller, 1994). cognitive load is an issue that a potential ai instructor would definitely need to address to be successful. this constraint, however, is not insurmountable. an ai might learn and apply a set of rules for how and how much material to present to an individual. a psychological constraint that might be a bridge too far, and the focus of this commentary, is that of our modern theories of psychological needs (e.g., baumeister & leary, 1995; deci & ryan, 2000; dweck, 2017). many researchers from a broad array of areas (carver & scheier, 1982; mischel & shoda, 1955) have supported the theory that there are a set of basic psychological needs that should be satisfied for human beings to function well and, in the case of children, develop to their fullest capacity. while many psychological needs have been put forth, basic psychological needs are built in, not developed over time (e.g., autonomy; dweck, 2017). different models for basic needs have been suggested with self-determination theory's arrangement taking centre stage for much of the past four decades (ryan & deci, 2017): a. autonomy, b. relatedness, c. competence. self-determination theory is focused on the first, which is central to their overarching theory of fryer 78 | f l r wellbeing through internal regulation of motivation. the latter two and other psychological needs (e.g., selfesteem and trust) have received less attention (patall et al., 2024). a more recent arrangement of psychological needs of different types has provided a balanced and inclusive perspective on this issue (i.e., dweck, 2017). for dweck, psychological needs for competence and relatedness (i.e., acceptance) stand at the heart of psychological needs which establishes a partial venn diagram with self-determination theory's widely used model. not included by dweck is the psychological need for autonomy, which is still important but positioned as an amalgamation of other psychological needs which develop with maturity. the dweck-sdt overlap focuses our attention on relatedness and competence, which have been well-researched and are strongly associated with elements of teaching (garcía-rodríguez et al., 2023; patall et al., 2024). 2.2 competence and structure corresponding to students' need for competence is what is called structure (skinner, 1995). structure refers to the ways in which teachers, and the learning environment they create, informs students about their developing competence (e.g., feedback). it also refers to how teachers provide opportunities for students to enhance (e.g., work on materials at the edge of their growing knowledge in a domain of study) and demonstrate (e.g., tests, projects, presentations to both experience mastery and note gaps in their understanding) their competence to themselves and others (skinner, 1995). this aspect of teaching is in many ways an area of obvious strength for a potential genai instructor. genais have an increasing ability to quickly and effectively construct teaching-learning materials across a number of domains: e.g., formative/summative assessments, practice materials, and a variety of classroom individual and group learning activities (holmes & miao, 2023). genai's increasing ability to efficiently develop precise curricula which can offer students' information about their learning progress and opportunity to practice would be extremely time consuming for a human teacher. furthermore, curricula provided by a human teacher are generally one-size fits all. curricula organised and delivered by an genai instructor could be increasingly individualised as the course of instruction progressed, providing personalised structure. while this basic psychological need seems potentially addressed by genai, perhaps in many ways better than a human instructor, the second psychological need raised is a more complex task for genai, current and future. 2.3 relatedness and involvement the second basic need is relatedness, which supported by teaching behaviours, is referred to as involvement. involvement refers to teaching behaviours that support students in feeling more connected to the teacher and learning environment more broadly. involvement more precisely includes teachers' expressions of caring and affection. it also refers to teachers providing resources and being physically/emotionally available/accessible to students (skinner et al., 1998). with much of the research around basic psychological needs being undertaken within selfdetermination theory, the relentless focus has been on the psychological need for autonomy, which is central to much of these researchers' theorising, with competence and then relatedness receiving far less relative attention. this does not make good practical (or theoretical) sense, given that the beliefs related to need for competence (e.g., self-efficacy) are some of the highest correlates of achievement (richardson et al., 2012). furthermore, a preeminent and widely cited study (skinner & belmont, 1993) made clear that if involvement (i.e., teaching that meets the need for relatedness) was controlled for, that students' perception of autonomysupport had no statistically significant impact on the learning behaviors and emotions. in fact, involvement, in addition to supporting future motivations for learning and learning behaviours, also increased teachers' autonomy-support and structure. a more recent meta-meta-analysis has confirmed the central role of need for relatedness, emphasising its broad and pervasive impact on supporting students' motivation to learn and achieve (jansen et al., 2022). fryer 79 | f l r 2.4 challenges for genai and involvement clearly, involvement is important for student learning. this is due to its paired, direct impact on student behaviours and motivation, and on teacher behaviours which support other psychological needs (structure and autonomy-support; skinner & belmont, 1993). it might be the most important part of what many teachers do. a question for ai researchers that arises from this contention is therefore: is involvement an aspect of teaching that an ai instructor is situated to offer at least as well as human teachers? to begin to answer this question, the components of involvement must be carefully reflected upon. two complex components that might be considered are: a. caring and affection or warmth from the teacher, b. physical and emotional availability/accessibility. can an ai currently or in the near future convincingly express caring and affection? even in its recent instantiations (e.g., chatgpt, gemini, claude etc...), it is conceivable that genai might express this aspect of involvement. the question is how effectively and for how long. it is easy to imagine how the veil of warmth might not be convincing to some students. it is also reasonable to see how an initial veil of involvement might be inadvertently and irretrievably shattered by a poor-quality exchange that informs the student that s/he is being taught by an algorithm. both weaknesses likely contributed to the failure of early efforts to use ai as learning partners in classrooms (e.g., fryer et al., 2017) and more broadly (okonkwo & ad-ibijola, 2021). of course, ai actually expressing warmth is likely only on our distant technological horizon (e.g., de togni et al, 2021). the second component of involvement that should be considered is physical and emotional availability/accessibility of the teacher (skinner, 1995). as neither of these are conceivable in a reasonable timeframe, the question again is whether genai can currently or in the near future convincingly convey these qualities. emotional availability much like caring and affection might be conveyed with increasing effectiveness as ai continues to develop. however, both types of availability might have to wait for a combination of level four (or five) ai (see openai's framework; e.g., velu, d, 2024) and life-like robotics to be possible. 2.5 technical issues for interested developers at least for the moment, involvement is the crucial component of teaching that ai will have substantial difficulty bridging. unlike structure, which genai instructors could effectively provide, involvement will be something that they can only fake. the quality of this facsimile will be a barrier to genai instructors' substantive effectiveness. this makes it an area that developers should be actively seeking to improve and rigorously test. the first fact that developers need to face is that much like structure being about personalising learning experiences based on individual competences, involvement will also need to be personalised to be effective. unlike structure, however, the genai teacher cannot simply give the student a test to estimate the necessary involvement. some clues regarding age-appropriate instructor involvement might be extracted from what we know about human psychological development (e.g., flavell, 1977; lerner, 2001). other clues might be elicited through adept interaction with students. this will rely on experimentation and the development of robust and adaptive communication skills. in addition to trying different instructional strategies, ais will need to gather emotional signals from humans, many of which are not oral. ais are already being trained for this (jaiswal et al., 2020) but integrating these signals with human developmental theory and oral interaction will be a challenge just to create a facsimile of ai instructor involvement. 2.6 how co-piloting might actually work while novel sounding, co-piloting is an academic idea which has been decades in development. the idea of human-computer partnerships has a long history within computer science. early frameworks discussed technical augmentation (engelbart & english, 1968) or a more fluid conceptual symbiosis between "men and electronic computers" (licklider, 1960; p.4). after two decades of technological development, early conceptions of shared/adjustable autonomy of software agents matured (e.g., bradshaw, 1997). in the most fryer 80 | f l r recent decade, there has been an increasing focus on human-centred ai, which seeks a balance between human and ai control while optimally supporting "humans` self-efficacy, creativity and responsibility" (shneiderman, 2020; p.495), with a strong focus on human-ai collaboration in workplaces (wilson & daugherty, 2020). the term co-piloting has recently been popularised by github copilot which is a popular, increasingly indispensable ai tool for programmers (wikipedia, n.d.). education took a parallel route, while building on the same original aspirations of humans being augmented or working in symbiosis with computers. intelligent tutoring systems (for a review see nawan, 1990) and a wide range of increasingly animated pedagogical agents (dehn & van muken, 2000) were developed and tested. with the increasing power of these systems (luckin et al., 2016) and then the eventual addition of genai (chatgpt3 in 2020), supports for learning blossomed across a broad range of education domains and contexts: through games like minecraft for programming skills, as support for classroom teachers' general materials development, or additional tutoring/practice for students (microsoft, 2025, march 5). together this development comes together as co-piloting in education which might be summed up by four central components. these components partially overlap with workplace conceptions of co-piloting and the areas of teaching-learning which genai is well positioned to contribute to. these are a. collaboration (augmenting and blending with teachers; luckin et al., 2016) and b. efficiency (automating tasks and saving teachers' time) are the most straightforward; c. building on intelligent tutoring systems decades-long effort to personalise learning, meeting students' at their knowledge level; d. probably most practical, co-piloting in education means building more and better feedback into the learning experience. this final component comes both through students' direct and indirect access to information. the indirect route arises from teachers gaining a better understanding of students' knowledge development. teachers are then in a stronger position to talk not just about where students are and need to go, but also how to get there (hattie, 2012). co-piloting has not to our knowledge been the focus of research in educational contexts. it is, however, a worthy avenue for discussion because, as noted to this point, ais have the propensity to become better at providing personalised structure than many human classroom teachers ever could (joshi, et al., 2021). however, structure's impact on student learning is further enhanced by involvement (i.e., support for relatedness needs; skinner et al., 1998), which will be provided by teachers. such a co-piloted classroom can readily be imagined. lessons are created by the ai; reviewed and perhaps adjusted by the teacher. these lessons would be individualised in parts by the genai to ensure all students are working in their proximal zone of development. at the same time, the teacher will ensure that important parts of these individualised lessons would intersect across students to foster sufficient social learning experiences (ensuring opportunities for relatedness needs to be met). teachers would keep abreast of students` progress through visualised analytics and ensure students know that teachers understand and care about their progress, that they are available to discuss difficulties in groups or one-on-one. at the same time, teachers would dedicate a substantive amount of classroom time to small and large group interactions, listening and building relationships with and between students. if teachers were left to focus on involvement, it is possible that they might even meaningfully improve the quality of the involvement they provide students. teacher education and the research that supports involvement as a teaching skill could home in on how involvement is best delivered to different students, at different ages and stages of development. in this way, co-piloting could, potentially, dramatically improve the teaching profession. 3. limitations and future directions given the short format of this article, its scope was necessarily narrow, focusing on the introduction of co-piloting between a human teacher and genai agent, aiming to improve students' learning experiences. a specific psychological approach was taken in this introductory effort, but many other (or additional) theories might have been applied to this ai-human educational interface. further theoretical work to a. elaborate the fryer 81 | f l r presently proposed perspective of psychological needs (e.g., dweck, 2017; ryan & deci, 2017), b. extend thinking to other well-established connected theories of education (e.g., skinner, et al, 2022) and c. diversify our understanding of this human and digital educational merger are necessary and called for (e.g., malgilchrist, 2021). furthermore, empirical research (experimental and longitudinal) is critical to test the ideas put forth here and other theorising that must follow. these tests need to be sensitive to students (e.g., students` culture and prior knowledge) and context (e.g., level and aims of education) if they are to be meaningful. ai are unlikely to match all educational needs; understanding where they are useful and where they are not will be critical (e.g., see dinsmore & fryer, 2025). finally, questions of structure and relatedness support by current and future ai raise ethical questions regarding students` understanding of who or what they are learning from. this issue has both expertise and emotional valences that will need to be addressed by research going forward. 4. conclusions educational technology has always had a gap between the technology available and the psychologically framed research that examines best teaching and learning practices (means, 2022). there is a danger, as ais rapidly develop and begin to look like an answer to many of our educational questions, that this gap turns into a gulf. before this happens, we should recognise what ai is likely to excel at and what it can only hope to be a facsimile of. ai needs to be properly harnessed to ensure it drives better learning experiences and outcomes. psychological needs theory (specifically needs for relatedness and competence) is a firm theoretical foundation upon which to develop co-piloting instruction. through such paired instruction (i.e., copiloting), ai and human might each maximise their respective strengths to the benefits of students for the generations to come. keypoints generative ai will have a rapidly growing role within formal education when considering what this role will look like, students' psychological needs must be considered we suggest a co-piloting arrangement with human and genai teacher addressing involvement and structure respectively a proposed co-piloting framework is presented, focusing on four components: a. collaboration, b. efficiency, c. personalisation, d. more and better feedback references baumeister, r. f., & leary, m. r. (1995). the need to belong: desire for interpersonal attachments as a fundamental human motivation. psychological bulletin, 117, 497–529. http://dx.doi.org/10.1037/00332909.117 .3.497 bradshaw, j. m. (1997). software agents. mit press. carver, c. s., & scheier, m. f. (1982). control theory: a useful conceptual framework for personalitysocial, clinical, and health psychology. psychological bulletin, 92, 111–135. http://dx.doi.org/10.1037/0033-2909 .92.1.111 chen, x., zou, d., xie, h., cheng, g., & liu, c. (2022). two decades of artificial intelligence ineducation. educational technology & society, 25(1), 28-47. de togni, g., erikainen, s., chan, s., & cunningham-burley, s. (2021). what makes ai ‘intelligent’ and ‘caring’? exploring affect and relationality across three sites of intelligence and care. social science & medicine, 277, 113874. https://doi.org/10.1016/j.socscime.2021.113874dinsmore, d. & fryer, l. k. https://doi.org/10.1016/j.socscimed.2021.113874 fryer 82 | f l r (2025, march 5). what does current genai actually mean for student learning? learning and individual differences http://dx.doi.org/osf.io/f8z56_v1 dinsmore, d. & fryer, l. k. (pre-print). what does current genai actually mean for student learning? https://doi.org/osf.io/f8z56_v1 dweck, c. s. (2017). from needs to goals and representations: foundations for a unified theory of motivation, personality, and development. psychological review, 124(6), 689–719. https://doi.org/10.1037/rev0000082 engelbart, d. c., & english, w. k. (1968). a research center for augmenting human intellect. proceedings of the december 9-11, 1968, fall joint computer conference, part i on afips ’68 (fall, part i), 395. https://doi.org/10.1145/1476589.1476645 flavell, j. h. (1977). cognitive development. prentice-hall. fryer, l. k., ainley, m., thompson, a., gibson, a., & sherlock, z. (2017). stimulating and sustaining interest in a language course: an experimental comparison of chatbot and human task partners. computers in human behavior, 75, 461–468. https://doi.org/10.1016/j.chb.2017.05.045 garcía-rodríguez, l., iriarte redín, c., & reparaz abaitua, c. (2023). teacher-student attachment relationship, variables associated, and measurement: a systematic review. educational research review, 38, 100488. https://doi.org/10.1016/j.edurev.2022.100488 hattie, j. (2012). know thy impact. educational leadership holmes, w., & miao, f. (2023). guidance for generative ai in education and research. unesco publishing. holstein, k., aleven, v., & rummel, n. (2020). a conceptual framework for human–ai hybrid adaptivity in education. in i. i. bittencourt, m. cukurova, k. muldner, r. luckin, & e. millán (eds.), artificial intelligence in education (vol. 12163, pp. 240–254). springer international publishing. https://doi.org/10.1007/978-3-030-52237-7_20 jiang, y., li, x., luo, h. et al. quo vadis artificial intelligence? discov artif intell 2, 4 (2022). https://doi.org/10.1007/s44163-022-00022-8 jaiswal, a., raju, a. k., & deb, s. (2020, june). facial emotion detection using deep learning. in 2020 international conference for emerging technology (incet) (pp. 1-5). ieee. jansen, t., meyer, j., wigfield, a., & möller, j. (2022). which student and instructional variables are most strongly related to academic motivation in k-12 education? a systematic review of meta-analyses. psychological bulletin, 148(1-2), 1–26. https://doi.org/10.1037/bul0000354 joshi, s., rambola, r. k., & churi, p. (2021). evaluating artificial intelligence in education for next generation. in journal of physics: conference series (vol. 1714, no. 1, p. 012039). iop publishing. kaplan, j. (2024). generative artificial intelligence: what everyone needs to know. oxford university press. lerner, r. m. (2001). concepts and theories of human development. psychology press. licklider, j. c. (1960). man-computer symbiosis. ire transactions on human factors in electronics, (1), 411. luckin, r., holmes, w., griffiths, m., & corcier, l. b. (2016). intelligence unleashed: an argument for ai in education. pearson. macgilchrist, f. (2021). theories of postdigital heterogeneity: implications for research on education and datafication. postdigital science and education, 3(3), 660–667. https://doi.org/10.1007/s42438-02100232-w means, b. (2022). making insights from educational psychology and educational technology research more useful for practice. educational psychologist, 57(3), 226–230. https://doi.org/10.1080/00461520.2022.2061974 microsoft (n.d.). https://www.microsoft.com/en-us/education/blog/2025/01/delivering-greater-impact-withcopilot-and-the-power-of-agents/ retrieved 5 march, 2025. mischel, w., & shoda, y. (1995). a cognitive-affective system theory of personality: reconceptualizing situations, dispositions, dynamics, and invariance in personality structure. psychological review, 102, 246 268. http://dx.doi.org/10.1037/0033-295x.102.2.246 moran, e. f. (2022). human adaptability: an introduction to ecological anthropology. routledge. https://doi.org/10.1037/rev0000082 https://doi.org/10.1016/j.chb.2017.05.045 https://doi.org/10.1016/j.edurev.2022.100488 https://doi.org/10.1007/978-3-030-52237-7_20 https://doi.org/10.1007/s42438-021-00232-w https://doi.org/10.1007/s42438-021-00232-w https://doi.org/10.1080/00461520.2022.2061974 fryer 83 | f l r okonkwo, c.w., & ade-ibijola, a. (2021). genais applications in education: a systematic review. computers and education: artificial intelligence, 2, 100033. https://doi.org/10.1016/j.caeai.2021.100033 patall, e. a., yates, n., lee, j., chen, m., bhat, b. h., lee, k., beretvas, s. n., lin, s., man yang, s., jacobson, n. g., harris, e., & hanson, d. j. (2024). a meta-analysis of teachers’ provision of structure in the classroom and students’ academic competence beliefs, engagement, and achievement. educational psychologist, 59(1), 42–70. https://doi.org/10.1080/00461520.2023.2274104 richardson, m., abraham, c., & bond, r. (2012). psychological correlates of university students’ academic performance: a systematic review and meta-analysis. psychological bulletin, 138(2), 353–387. https://doi.org/10.1037/a0026838 ryan, r. m. & deci, e. l. (2017). self-determination theory: basic psychological needs in motivation, development, and wellness. the guilford press, new york. skinner, e. a., & belmont, m. j. (1993). motivation in the classroom: reciprocal effects of teacher behavior and student engagement across the school year. journal of educational psychology, 85(4), 571–581. https://doi.org/10.1037/0022-0663.85.4.571 skinner, e. (1995). perceived control, motivation, & coping. sage publications, inc. https://doi.org/10.4135/9781483327198 skinner, e. a., zimmer-gembeck, m. j., connell, j. p., eccles, j. s., & wellborn, j. g. (1998). individual differences and the development of perceived control. monographs of the society for research in child development, 63(2/3). https://doi.org/10.2307/1166220 skinner, e. a., kindermann, t. a., vollet, j. w., & rickert, n. p. (2022). complex social ecologies and the development of academic motivation. educational psychology review, 34(4), 2129–2165. https://doi.org/10.1007/s10648-022-09714-0 sweller, j., & chandler, p. (1991). evidence for cognitive load theory. cognition and instruction, 8(4), 351362. https://doi.org/10.1207/s1532690xci0804_5 sweller, j. (1994). cognitive load theory, learning difficulty, and instructional design. learning and instruction, 4(4), 295–312. https://doi.org/10.1016/0959-4752(94)90003-5 tontini, g. e., & neumann, h. (2021). artificial intelligence: thinking outside the box. best practice & research clinical gastroenterology, 52, 101720.https://doi.org/10.1016/j.bpg.2020.101720 velu, d. (2024). ai's 5-level framework to agi. medium. downloaded 7 march 2025. https://medium.com/@dheeren.velu/ais-s-5-level-framework-to-agi-2d0ef4880f95 wenzlaff, k. and spaeth, s. (2022). smarter than humans? validating how openai’s chatgpt model explains crowdfunding, alternative finance and community finance http://dx.doi.org/10.2139/ssrn.4302443 wikipedia (n.d.). github. https://en.wikipedia.org/wiki/github retrieved 5 march, 2025. https://doi.org/10.1080/00461520.2023.2274104 https://doi.org/10.1037/a0026838 https://doi.org/10.1037/0022-0663.85.4.571 https://doi.org/10.4135/9781483327198 https://doi.org/10.2307/1166220 https://doi.org/10.1007/s10648-022-09714-0 https://psycnet.apa.org/doi/10.1207/s1532690xci0804_5 https://doi.org/10.1016/0959-4752(94)90003-5 https://doi.org/10.1016/j.bpg.2020.101720 https://dx.doi.org/10.2139/ssrn.4302443 frontline learning research special issue: vol. 13 no. 2 (2025) 67 101 issn 2295-3159 corresponding author: k. ann renninger, 500 college ave. swarthmore, pa 19081, usa, krennin1@swarthmore.edu. doi: https://doi.org/10.14786/flr.v13i2.1371 exploring math moments: middle-schoolers’ phases of problemsolving, executive functions in practice, and collaborative problem solving k. ann renninger1, gertraud benke2, ricardo böheim3, julien corven4, maria consuelo de dios1, maeve r. hogan1, moe htet kyaw1, ana g. michels1, marina nakayama1, pablo e. torres5, helena werneck1, & feven yared1 1swarthmore college, usa 2university of klagenfurt, austria 3technical university of munich, germany 4illinois state university, usa 5facultad de educación, pontificia universidad católica de chile article received 27 september 2023 / article revised 18 october 2024 / accepted 24 october 2024/ available online 14 march 2025 abstract collaborative problem solving (cps) has been shown to both engage and benefit students’ learning of mathematics. however, there is evidence that group work is not always easy to facilitate, in part because educators lack details about learners’ engagement during group work: the processes of problem solving involved, and how these are engaged. in this exploratory study, we focused on these processes in the moments of related math activity, or math moments, engaged by two groups of interested, urban, middle-school aged students during four sessions of work in the virtual math teams (vmt) environment. we examined three phases of their problem solving: exploring, constructing, and checking. in addition, to further describe the students’ cognitive and behavioral engagement, we considered both the process of students' use of executive functions (ef), during problem solving, termed executive functions in practice (efp), as well as the stage of cps (participation, cooperation, and collaboration), during phases of problem solving. we learned that the relation between each phase of problem solving, categories of efp, and stages of cps vary; for example, the problem-solving phase of exploring was found to have a more positive effect on efp and cps than either constructing or checking. implications for educational practice, and next steps for related research are described. keywords: momentary engagement, problem solving, executive functions, collaborative problem solving mailto:krennin1@swarthmore.edu renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 68| f l r 1. introduction studies have shown that, compared to individual activity, students’ collaboration during mathematical problem-solving increases their opportunities for engagement (e.g., webb et al., 2019) and can support students to develop their identities as learners of mathematics (e.g., featherstone et al., 2011). however, following a review of 66 qualitative and quantitative studies, van leeuwen and janssen (2019) concluded that teachers’ abilities to adjust the assistance they offer their students during collaborative activities vary, and their effectiveness, in turn, significantly impacts whether student learning occurs. to understand how collaborative problem solving (cps) might be optimized in classroom instruction, detail is needed about students’ engagement in phases of problem solving during moments when they are working with mathematics. momentary engagement refers to individuals’ activity, or participation, during a brief period of time, and has primarily been discussed as situation-specific (nolen et al., 2015; symonds et al., 2021). dietrich et al. (2022) pointed to the importance of further examining such experiences to consider complex change. investigation of momentary engagement in group-based contexts, in particular, could extend understanding of the role of group work in supporting students’ attention to tasks (e.g., hmelosilver et al. 2018; pollastri, et al., 2013), enhancing reasoning (e.g., barron, 2003), and promoting the development of a collective working memory that may increase their capacity for problem solving (kirschner et al., 2018; van den bossche, et al. 2011; zambrano et al., 2019). in this article, we report on findings from an exploratory study of two groups of urban middle school students’ work with online collaborative problem solving in the virtual math teams (vmt, https://vmt.mathematicalthinking.org/) environment (stahl, 2013). these data are publicly available and de-identified. the two groups each worked with the same four sessions of open-ended dynamic geometry problems, and their sustained engagement across the sessions indicated that all participants had a developed interest in the problem solving they were doing (renninger et al., manuscript in preparation; renninger & hidi, 2016). to study cognitive and behavioral engagement during phases of the students' problem solving (exploring, constructing, and checking), we investigated the students’ use of core categories of executive functions (ef; working memory, cognitive flexibility, inhibitory control) which we describe as executive functions in practice (efp), their cognitive engagement, as well as their behavioral engagement in each stage of cps (participation, cooperation, collaboration). 1.1 momentary engagement and mathematics studies of engagement in academic tasks have primarily focused on three types of engagement: cognitive, behavioral, and affective (e.g., skinner & pitzer, 2012; see fredricks et al., 2004; fredricks & mccolskey, 2012 for reviews). research has shown that these types of engagement co-occur and are malleable. as pohl (2020) explained, students who are completing math tasks benefit from being cognitively engaged as this positions them to recall math concepts and to use these in planning strategies, work on solving the problem, monitor their progress, and evaluate their correctness. moreover, research has shown that when students are sufficiently behaviorally engaged to complete tasks and participate in their mathematics class, their cognitive engagement is effectively supported (dong, et al., 2020), and they are able to maintain engagement (cook, et. al., 2020). within-person fluctuations in engagement levels have also been observed, leading a number of researchers to call for more situated and momentspecific analyses (e.g., dietrich et al., 2022; nolen, 2020; rogat et al., 2022; salmela-aro et al., 2021; symonds et al., 2019, 2021). given these findings, the study of momentary engagement in the context of collaborative activity in mathematics may be particularly important. 1.1.1 mathematical sense-making and practices mathematical sense-making is central to the process of problem solving in mathematics, as it describes individuals' developing understanding and flexible use of mathematical concepts (e.g., harel & sowder, 2005; schoenfeld, 1992/2016). the development of mathematical sense-making is a renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 69| f l r cognitive process that requires students to work with new information and make connections to what they already know (e.g., kasmer & kim, 2011; rau & matthews, 2017); it also is a skill that is developed and enhanced through collaboration in specific contexts, such as problem solving (bonotto, 2005; gerson 2008; kelton et al., 2018; stahl, 2013). problem solving has been variously described as a set of steps, or phases, that begins with understanding the problem and progresses through a sequence that includes making a plan, carrying it out, and checking (polya, 1945; schoenfeld, 1992/2016). stahl (2013) clarified that the process is not so linear when individuals are working collaboratively with open-ended, inquiry problems that allow exploration. in the collaborative context, the process involves discovery which includes explorative dragging (exploring), experimental construction (constructing), and determination of dependencies (checking)—phases of problem solving that have been shown to benefit conceptual understanding. boaler and selling (2017), for example, reported that students in mathematics classrooms that have explicitly prioritized the exploration of strategies as a component of problem solving have a deeper understanding of mathematics content, as well as more positive feelings about mathematics, than students in classrooms that do not encourage exploration. studies also have provided evidence that checking work following problem solving is associated with higher levels of student performance and conceptual understanding (e.g., eshuis et al., 2019; kuhn et al., 2020; zhang et al. 2021). 1.1.2 how mathematics practices are engaged, efp and cps study of students’ use of cognitive and behavioral engagement with phases of problem solving has the potential to provide detail about how students engage in collaborative mathematics activities. research has addressed the outcomes of efs, as individual categories of behavior, and as a composite description of engagement (e.g., mann et al., 2017; younger et al., 2023) and cps (e.g., andrews-todd & forsyth, 2020) individually. to the best of our knowledge, however, no one has either investigated the use of executive functions (efs) and/or stages of cps during students’ behavioral engagement in phases of problem solving. in the present study, we use efp to describe the process of students’ use of efs in naturally occurring settings such as collaborative group work. whereas efs are usually studied in controlled settings with standardized tasks (bailey et al., 2018, chan et al., 2008; mccoy, 2019), in our investigation we are focused on the process of students’ cognitive and behavioral engagement with problem solving; therefore, we review efs as they are activated and practiced in an ecologically valid context. in general, three core efs are used to describe individuals’ cognitive engagement with tasks: working memory, cognitive flexibility, and inhibitory control (diamond, 2013). working memory refers to the ability to recall and manipulate relevant information to work with tasks (bailey et al., 2018; diamond, 2013; radvansky & copeland, 2006). cognitive flexibility describes the ability to adjust one’s behavior or thoughts to changed circumstances such as unexpected failure, or opportunity (diamond, 2013; jacques & zelazo, 2001). inhibitory control pertains to the ability to resist the impulse to respond to a situational demand and to instead engage in a more appropriate but subdominant response (letang et al., 2021; veraska et al., 2020). studies of efs have shown that they are critical for both achievement and learning (caviola et al., 2020; jose et al., 2020; long et al., 2011; mann et al., 2017; skaguerlund et al., 2019). in mathematics, research has pointed to positive associations between math achievement and core efs (e.g., brookman-byrne et al., 2018; clark et al., 2010; swanson & beebe-frankenberger, 2004; see also cragg & gilmore, 2014 for a review of the literature on ef and mathematics). specifically, in solving a mathematical problem, working memory allows students to recall and apply previously learned knowledge (karpike, 2012; peng et al., 2016); cognitive flexibility allows students to sort through different solutions and choose the most efficient strategy (huizinga et al., 2014; yeniad, et al., 2013); renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 70| f l r and inhibitory control allows students to focus their attention on the task and ignore distractions (bishara & kaplan, 2022; ponitz et al., 2009). different efs also may be uniquely utilized in one or another part of problem solving. for example, viterbori et al. (2017) found that when students are working on multi-step problems, working memory significantly predicted problem solving accuracy, presumably because it allowed students to make use of correct mathematical information. viterbori et al. also reported that students may need to call upon cognitive flexibility to shift between multiple representations of a problem. in other work, lee et al. (2009) suggested that inhibitory control may be particularly necessary for students’ understanding when a problem includes irrelevant information that needs to be ignored. by contrast, collaborative problem solving (cps) describes stages in the process of two or more individuals working together to achieve a shared goal: participation, cooperation, and collaboration. drawing upon the a3c framework, referring to attendance, coordination, cooperation, and collaboration (jeong et al., 2017) and self-regulated learning theory (hadwin et al., 2011), we describe cps stages as reflecting the extent to which students’ behavioral engagement is influenced by other students in their group. participation refers to individual behaviors in which a student may act according to their own goals and methods, but their goals and methods are dependent on and shaped by other people; cooperation refers to group behaviors in which a common goal is explicitly established, but the methods to achieve it are not joint; collaboration refers to group behaviors in which goals and methods are shared and jointly enacted peer-to-peer interactions during group problem solving provide a basis for knowledge construction that results in shared group cognition that is not present when an individual works alone (stahl, 2013; van den bossche, 2011; zambrano, et al., 2019). students’ behavioral engagement when working with a group has been shown to differ from individual problem solving. for example, sun et al. (2022) found that fifth grade students’ behavioral engagement was highest during collaborative work and lowest during independent work associated with direct instruction. furthermore, outcomes of group problem solving often differ from individual problem solving as shown by mohammadhasani and asadi’s (2020) study of students' completion of mathematics problems in online collaborative groups. they reported that those who completed the problems collaboratively experienced greater learning gains than those who worked on them individually. these findings are consistent with other studies demonstrating that students working in groups on academic tasks outperform those working individually on the same tasks (e.g., eshuis et al., 2019; sankaranarayanan et al., 2021). 1.2 current study in this study, we explored the process of two groups’ momentary engagement during four sessions of online collaborative problem solving in the vmt environment. our goal was to detail the process of the students’ problem solving during group work, considering variations in cognitive and behavioral engagement among students who were interested and on task. we selected the groups for study based on the following criteria as assessed in a prior study (renninger et al., manuscript in preparation): high levels of interest in mathematics (all participants in each group studied, remained on task throughout each of the problem-solving sessions; renninger & hidi, 2016), similar number of math moments (see section 2.3.1), but differences in their demonstrated levels of collaboration (for the methodological determination, see 2.3.4). specifically, group 1 was less collaborative than group 2. although there were different numbers of students in the two groups selected for analysis, data from prior study suggested that differences in the number of participants in student groups did not influence the observable dynamics of groups. study of two groups that had high levels of interest allowed study of high levels of cognitive and behavioral engagement during group work; study of groups that had approximately the same number of moments of math-related activity, or math moments ensured that the structure of the sessions were similar for both groups. these two criteria allowed us to focus on potential differences between the renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 71| f l r groups related to the quality of the students’ collaboration. we expected that these data could provide a preliminary mapping of the engagement of students in this age groups’ online collaborative problem solving and could help to address what teachers need to know to effectively support their students to work collaboratively. many studies of engagement measure two time points, usually before and after task completion; however, as siegler (1998) pointed out, study of any type of change ideally assesses and measures change while it is occurring. given findings such as symonds et al.’s (2021) indicating that momentary engagement varies within individuals, methods of measurement should ideally assess student engagement across the entire process of problem solving, and, in the group context, should account for co-negotiated engagement processes (rogat et al., 2022). here, the chat and replayer functions of vmt provided us with moment-to-moment records of the students’ cognitive and behavioral engagement. we focused on three features: (a) the phases of problem solving engaged – exploring, constructing, and checking; (b) the students’ use of core efp – working memory, cognitive flexibility, and inhibitory control during phases of problem solving providing information about cognitive engagement; and (c) each stage of cps – participation, cooperation, collaboration– during phases of problem solving, which provided information about their behavioral engagement. two research questions were addressed: rq 1: what is the relative proportion of each phase of problem solving overall, and is the distribution of phases similar by session? do these proportions vary for two groups with different levels of collaboration? rq 2: how do groups engage with different phases of problem solving? are there differences in the cognitive and behavioral engagement of two groups with different levels of collaboration? specifically, what is the relation between each phase of problem solving and the efp and cps of each group? is there change across sessions? 2. methods 2.1 participants participants were two groups of middle school students from the same urban school who were enrolled in an after-school program. group 1 had three members and group 2 had four members; their participation in the group was anonymous and their data were de-identified. 2.2 vmt learning environment the vmt environment is an online, multi-user version of geogebra that includes a shared workspace and a chat tool to the side of the screen (see figure 1). each group of students worked in a separate vmt room. group members communicated with each other through the chat, which allowed them to type and submit messages that were displayed to all of them. only one student was able to interact with the shared space at a time, and did so by clicking on a “take control” button. while one student was "in control" others made suggestions through the chat. the same groups of students worked together on each of the four sessions of geometry topics (see table 1). the first three sessions each lasted for an hour. during the fourth session, the students worked for more than two hours, mid-way through the session, they took a break. the problems on which the students worked and the instructions were the same for each group in each session, and prompted student use of the chat to discuss their work. renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 72| f l r 2.3 problems the problems for each of the sessions were rich, open-ended geometry problems that were specifically designed to scaffold support for developing skills in “collaborative and mathematical discourse, exploring dependencies, geometric construction, analytic explanation and domain content” (p. 163, stahl, 2013). problem content for each session was distinct and was anchored in the preliminary standards for high-school geometry described in the common core state standards initiative (2011). thus, the students were asked to engage in working with alternate representations and dependencies, including congruence, symmetry, and rigid transformations. problems for each session were sequenced to support the students’ developing levels of skill for understanding and working with proof, consistent with the van hiele levels for geometric reasoning (see devilliers, 2003, as cited in stahl, 2013). problems first focused the students on noticing and wondering (ray-riek, 2013), and followed this with encouragement to describe what they were doing and their justification for this in their chat-based explanations of their work. 2.4 coding and data reduction 2.4.1 math moments math moments refer to chunks of related and sequential mathematics activity during open-ended problem solving. in the present study, identifying the moments when students were working with mathematics (as opposed to socializing or asking procedural questions) during each session of problem solving provided the context for studying the student groups’ cognitive and behavioral engagement during phases of problem solving. as such, math moments were not identified based on screen or tab in this vmt context. rather, the group members set the direction of the math moments that they engaged, and we aggregated these based on the group members’ steps in their discovery process as they worked on the open-ended problem solving of each session. thus, while the sessions of problem solving were comparable to each other in the opportunities they afforded, individual math moments might not be due to their having different foci, lengths, etc. for this reason, we aggregated these data for analyses at the session level. math moments were consensually identified by two researchers for each group (hill, 2012; see tables 2 and 3 for a description of the content and duration of the math moments engaged by each group’ across sessions). most of the moments for each session for each group were math moments. while it was expected that the groups would vary in their work with each of the problems, the math moments identified for each included similar content (e.g., constructing an equilateral triangle, or discussing the meaning of constrained points), and/or often reflected the prompts that the problem provided. the consensual reliability check was employed to ensure that a consistent set of considerations informed the identification of math moments for each group, and between groups (e.g., the decision that multiple attempts at the same construction were described as multiple moments). figure 1. screenshot of vmt activity. renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 73| f l r note. the screenshot is of topic 2, equilateral triangles. table 1 session topic descriptions renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 74| f l r topic 1 introduction to vmt objective: understand how to construct objects and create dependencies session goals: • use construction tools (i.e. point tool, compass tool, etc.) to construct points, segments, and figures • construct appropriately independent and dependent points • perform “drag tests” (moving points of a given figure) to test dependencies topic 2 equilateral triangles objective: understand the properties of triangles session goals: • construct an equilateral triangle using two circles with a common radius • generate sets of dependent relationships of a given figure with an inscribed triangle • converse about the different properties of triangles topic 3 perpendicular bisectors objective: understand the properties of perpendicular lines and perpendicular bisectors session goals: • construct a perpendicular bisector of the radius of a circle • construct a perpendicular bisector through an arbitrary point on a line • construct a parallel line tool topic 4 inscribed polygons objective: understand geometric proportions session goals: • construct an equilateral triangle inscribed within another proportional equilateral triangle • construct a square inscribed within another proportional square • construct a hexagon inscribed within another proportional hexagon 2.4.2 phases of problem solving as described in table 4, three phases of problem solving may characterize math moments: exploring, constructing, or checking (e.g., polya, 1945; salminen-saari et al., 2021; schoenfeld, 1992/2016; stahl, 2013). although more formal conceptualizations of problem solving may point to these phases occurring sequentially, as stahl (2013) noted, groups working with problems in vmt sessions may engage in these phases in any order. for example, a group may enter the constructing phase and begin to construct figures before the group has had a chance to explore the problem, or the renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 75| f l r group may skip the checking phase entirely after they have completed their constructions. two researchers coded the phases of problem solving associated with each math moment, following which an independent researcher coded 20% of the data drawn at random and conducted a reliability check; reliability was substantial, k= .67 (landis & koch, 1977). 2.4.3 efp group members' efp were coded at the individual student level for each math moment. all students were coded for categories of behaviors associated with inhibitory control, working memory, and cognitive flexibility (see table 5). as shown in table 5, the coding rubric for assessing each efp was derived from the existing literature on the relevant ef and consisted of multiple items. by definition, the coding of efp involved (a) reviewing the students’ process during problem solving, (b) required identifying the opportunities to use efp created by the group and afforded by the activity's design, and was followed by (c) consideration of to what extent individuals took advantage of opportunities. coding of efp was undertaken in three steps. first, two researchers reviewed each group’s work using the replayer focusing on the identified math moments for the groups. second, the researchers had an initial meeting to discuss and agree on opportunities that were identified for group participation (e.g., to code working memory, the researchers needed to agree about whether there is information that the students need to recall to complete a task). we only scored items describing behaviors for which students had opportunities to engage in. when students did not have an opportunity to use efp, we coded this as nonapplicable. third, the researchers conducted an independent assessment of each student’s efp during math moments using the questions described in table 5. for each question, students were scored on a fourpoint scale consisting of the values 1 (fully exhibited described behaviors), 0.5 (had some but not all behaviors), -0.5 (had minimal number of associated behaviors), and -1 (did not exhibit the behaviors). scores ranged from -1 to 1, with a 1 indicating that the student took full advantage of the opportunities afforded in that moment to exhibit a given ef, and a -1 indicating that a participant took no advantage of those opportunities. as described above, an exception was made for instances in which no opportunity was present (n/a). this affected cognitive flexibility scores the most, possibly because it is a higher order ef (diamond & ling, 2016) more than other efp, leading to fewer moments in which cognitive flexibility was coded. finally, an average score for each type of efp (working memory, cognitive flexibility, and inhibitory control) was calculated for each individual and for the group. the average efp score was an aggregate of working memory, cognitive flexibility, and inhibitory control average scores for the individual and the group. amount of each type of efp was calculated for each student. to confirm reliability, two raters independently rated students’ use of efp. all coding was established through discussion; scores were reviewed and revised following hill (2012) until 100% agreement was achieved. renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 76| f l r table 2 group 1’s math moment content and duration, sessions 1-4 session 1: introduction session 2: equilateral triangles session 3: perpendicular bisectors session 4: inscribed polygons 1. segment a construction (03:18) 1. segment de construction (06:29) 1. segment construction (03:06) 1. polygon drag and discussion (17:31) 2. dragging points a and b (03:13) 2. triangle def construction (04:47) 2. circle constructions (02:57) 2. outer triangle construction (07:16) 3. various segment construction (02:34) 3. firsts attempt at construction of circles a and b (06:27) 3. segment construction (02:51) 3. first attempt at inner triangle construction (12:07) 4. discussion of figure (01:48) 4. second attempt at construction of circles a and b (06:39) 4. first attempt at line j construction (04:18) 4. inner triangle drag test (09:32) 5. experimenting with text (11:23) 5. point c construction s(03:23) 5. second attempt at line j construction (04:02) 5. tab hints (00:22) 6. various segment construction (01:38) 6. triangle abc construction (02:44) 6. line j drag test (06:01) 6. inner triangle drag test (03:48) 7. circle constructions (02:09) 7. discussion of impressions (03:25) 7. line construction (02:52) 7. exploring compass tool (16:31)a 8. discussion of figure (01:21) 8. discussion of constrainment (03:11) 8. circle constructions (01:28) 8. second attempt at inner triangle construction (25:43) 9. angle constructions (03:25) 9. discussion of segment lengths (01:58) 9. line constructions (03:45) 9. reflection (09:42) 10. experimenting with grid (03:14) 10. discussion of angles (00:15) 10. perpendicular line drag test (00:49) 10. discussion of dependencies (06:19) 11. polygon constructions (02:34) 11. discussion of unsure relationships (03:26) 11. circle constructions (03:08) 11. testing dependencies (02:57) 12. intersection constructions (09:02) 12. discussion of triangle types (02:30) 12. line constructions (05:03) 12. planning outer square (04:00) 13. segment and dependencies (02:48) 13. discussion of unsure relationships (03:41) 13. dragging lines (02:07) 13. outer square construction (17:46) 14. dragging triangles (03:38) 14. creating tools (08:49) 14. planning inner square (03:48) 15. testing tools (05:59) 15. inner square construction (06:01) 16. reflection (03:27) renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 77| f l r note. duration is reported in minutes and seconds based on time stamp. a students took a break and then resumed work. table 3 group 2’s math moment content and duration, sessions 1-4 session 1: introduction session 2: equilateral triangles session 3: perpendicular bisectors session 4: inscribed polygons 1. planning problem (05:57) 1. triangle def construction (07:29) 1. segment ij construction (03:06) 1. discussion of impressions (07:15) 2. exploring angles (04:06) 2. circle g construction (02:03) 2. circle i and j constructions (03:30) 2. first attempt at outer triangle construction (02:37) 3. shape constructions (02:38) 3. first attempt at circle h construction (03:06) 3. segment kl construction without perpendicularity (06:20) 3. first attempt at inner triangle construction (01:39) 4. experimenting with display (09:21) 4. 2nd attempt at circle h construction (05:12) 4. segment kl construction with perpendicularity (03:06) 4. testing construction (06:47) 5. segment constructions (04:19) 5. circle j and l construction (05:26) 5. mn line construction (10:09) 5. second attempt at outer triangle construction (32:47)a 6. dragging figures (05:51) 6. triangle jkl construction (02:08) 6. discussion of perpendicular bisectors (04:49) 6. discussion of compass tool (23:37) 7. experimenting with circles (03:54) 7. circle k and j construction (02:37) 7. testing perpendicularity (16:44) 7. second attempt at inner triangle construction (13:34) 8. experimenting with compass tool (07:08) 8. triangle jkl construction (01:14) 8. creating new tool (04:21) 8. discussion of impressions (05:42) 9. discussion of figures (02:41) 9. discussion of figures (05:29) 9. testing new tool (11:09) 9. first attempt at outer square construction (07:03) 10. line segment constructions (04:31) 10. discussion of impressions (05:00) 10. second attempt at outer square construction (04:22) 11. discussion of constrainment (03:02) 11. third attempt at outer square construction (04:37) renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 78| f l r 12. discussion of segment lengths (03:18) 12. fourth attempt at outer square construction (14:08) 13. discussion of angles (08:08) 14. discussion of triangle types (02:55) 15. discussion of triangle changes (03:47) note. duration is reported in minutes and seconds based on time stamp. a students took a break and then resumed work renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 79| f l r table 4 phases of mathematical problem solving during math moments, definitions and examples phase definition example exploring examining characteristics and dependencies among the geometric elements either to understand the problem or plan how to solve it multiple students drag a point in a pre-constructed triangle; others observe their peers’ dragging and comment on how the proportions and size of the triangle is affected constructing purposefully adding or modifying geometric elements in order advance problem solving by creating new figures or tools the group constructs a triangle in a way that its sides are equilateral checking exploring, moving, turning, flipping, or resizing existing elements in order to assert their relationship the group drags all vertices of an equilateral triangle to see if it stays equilateral note. math moments may also include hybrids of the different phases of problem solving or discussion. 2.4.4 cps group members’ cps was coded using a three-step process that paralleled that described for coding and scoring efp. cps was coded at the individual level for behaviors associated with participation, and at the group level for behaviors associated with cooperation and collaboration (see table 6). although the groups were drawn for analysis based on differences in levels of collaboration, here we studied each stage of each group’s cps to understand their collaboration during phases of problem solving, and in relation to efp during phases of problem solving. as depicted in table 6, the coding rubric for each stage of cps was derived from the existing literature and included multiple items. reliability of all coding was established through discussion; scores were reviewed and revised following hill (2012) until 100% agreement was achieved. renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 80| f l r table 5 rubric for scoring categories of core executive functions in practice (efp) category item examples working memory does the student bring up and use mathematical concepts, ideas, practices, or tools that haven’t been mentioned in the session? (e.g., karpike, 2012; radvansky & copeland, 2006) the student brings up a new term that is relevant in a discussion, such as using the term “right triangle” to identify a triangle that has a right angle. does the student bring up or maintain attention to and use mathematical concepts, ideas, practices, or tools that have been mentioned in the session? (e.g., bailey et al., 2018; gathercole, et al., 2006) the student is chatting with another student and utilizes ideas mentioned in the directions, such as talking about the length of different segments of a triangle. does the student remember how to use relevant features of the tool? (e.g., diamond & ling, 2016; gathercole, et al., 2006) the student uses the circle-making tool correctly in order to complete a step or solve a problem, or the student guides another student on how to use the tool. cognitive flexibility does the student recognize that a shift in approach or perspective might be needed? (e.g., laureiro-martínez & brusoni, 2018; huizinga et al., 2014) when a method of construction does not yield the expected results, such as the construction of an acute triangle when the goal was a right triangle, the student asks their group members if there is another way to construct the figure. does the student switch perspectives, frames, or strategies due to an external perspective? (e.g., laureiro-martínez & brusoni, 2018; huizinga et al., 2014) the student changes their method of construction after another student points out that the current method is not suitable, such as learning from a peer to anchor points to lines due to a peer’s suggestion after initially making improper intersections. does the student move beyond a strategy that is not working? (e.g., diamond & ling, 2016; huizinga et al., 2014; jacques & zelazo, 2001) when the current method of solving the problem doesn’t yield the expected construction, the student suggests and/or tries another method, such as using the radii of circles to measure length after visual line estimations were unsuccessful. inhibitory control does the student exhibit signs of semantic inhibition by inhibiting past conceptions of an idea to engage in problem solving? (e.g., cervera-crespo & gonzalez-alvarez, 2017; chan et al., 2008) the student purposely ignores or rejects the urge to create a construction based on their visual understanding of geometry (e.g., by stopping a group member from copying the example), which is an indication that they realize that their construction may not be correct even if it “looks” right. does the student follow and/or revisit norms for group practice? (e.g., diamond & ling, 2016) the student responds when questions or statements are addressed to them specifically, such as releasing control when another student asks to take control. renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 81| f l r does the student exhibit signs of response inhibition by controlling a motor impulse to instead act more appropriately? (e.g., chan et al., 2008; verbruggen & logan, 2008) when a construction is not working as expected, the student tries different, unfamiliar tools to resolve the problem with the construction rather than using the same, familiar tool again. throughout the session, does the student attend to the task? (e.g., diamond, 2013; chan et al., 2008) all of the student’s activity is relevant to the content. table 6 rubric for scoring stages of collaborative problem solving (cps) level of analysis, stage item and sample sources examples individual level participation does the student make use of the following resources: directions, the group members, past problems/parts of the current problem, math knowledge, tools? (melzner et al., 2020; su et al., 2018) to construct a figure, the student looks to the directions for steps, asks group members for confirmation, uses their math knowledge to choose the appropriate tool, and uses the tool to construct the figure by employing a previously established method. does the student attend to and notice things that need further investigation or explanation? (su et al., 2018; zhang et al., 2021) the student attends to the construction being made by another student and communicates in the chat what they notice. does the student work through their goals? options: no follow through; follow through, no consideration for goals; follow through, incomplete/failed; follow through, complete; modified goal/new attempt (straub & rummel, 2021; zhang et al., 2021) the student follows through with steps they or their group have identified. does the student recognize (is the student aware) that another perspective or approach is being us ed/suggested? (straub & rummel, 2021; zhang et al., 2021) the student, while in control, implements suggestions from other students and changes their original construction to fit their peers’ suggestions. throughout the entire session, is the student’s engagement with the program balanced? (e.g,. melzner et al, 2020; straub & rummel, 2021) the student participates in most math moments and has an equal balance of chat and tool use overall. group level cooperation does the group define their goals? (e.g., jeong, et al., 2017; kuhn et al., 2020) students all agree on a specific step to do before a student actually attempts the step. does the group share new ideas, including brainstorms they may be uncertain about? (e.g., marek et al, 2015; mercer & sams, 2006) students bring up different ways or methods to solve a problem. renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 82| f l r do students maintain positive communication and turn-taking? (e.g., eshuis et al., 2019; mercer & sams, 2006) students make sure that everyone has had a chance to take control and, if not, encourages others to participate. do students engage in discussion during and/or after problem solving beyond naming the directions? (e.g., phelps & damon, 1989) a student uses the chat function to ask another student a question, and that student responds and answers the question. collaboration does the group execute their goals together? (e.g., andrews-todd & forsyth, 2020; jeong, et al. 2017; marek et al., 2015) when constructing a figure, each person takes control to construct some part of the figure. is there evidence of shared understanding? (e.g., kuhn et al., 2020; mercer & sams, 2006) when reflecting on the problem after an attempt, students discuss and agree on their satisfaction with their attempt. if challenges and alternatives are raised, are they pursued and negotiated? (e.g., andrews-todd & forsyth, 2020, marek et al, 2015) a student disagrees with a statement made by another student, and there is a discussion about the disagreement. does the group incorporate viewpoints from each other when necessary? (e.g., kuhn et al., 2020; mercer & sams, 2006) a student makes a suggestion for a specific construction, and the student in control makes the suggested construction. is there evidence of extension of thinking? (e.g., eshuis et al, 2019; jeong, et al., 2017; mercer & sams, 2006) students build off on each other’s ideas by introducing new perspectives. does the group provide justification or give mathematical reasoning for each other? (e.g., phelps & damon, 1989; vandenberg et al., 2021) when a question is posed by a student, other students answer the question and attempt to explain their answers. throughout the activity, can continuity be identified among the group? (e.g., eshuis et al, 2019; mercer & sams, 2006) throughout the entire session, students consistently build on and add to each other’s constructions and discussions. renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 83| f l r 2.5 analysis strategy data analyzed for each group’s cognitive and behavioral engagement with each session of problem solving includes information about the students' activity in the shared workspace, as well as their written contributions in the chat. observed information about student participation is reported first, followed by a description of analyses of the problem solving of each group. to answer rq1, we first used wald z tests to compare the proportion of phases in each group’s total number of math moments. we used wald z to directly compare each proportion as some moments were coded in more than one phase. to consider the distribution of phases between sessions, bar graphs were developed to show patterns in the data. in rq2, we explored how each group engaged during math moments; we studied both their cognitive (efp) and behavioral (cps) engagement. we divided our analyses into three parts. the first set of analyses focused on the scores for aggregated sessions to observe general behavior. to understand the importance of each phase of problem solving on participants' engagement, we used bar graphs to compare efp and cps scores in each of the three phases. for the second and third parts, we ran multivariable regression analyses. each analysis had a particular ef (average efp, working memory, inhibitory control, or cognitive flexibility) or cps (participation, cooperation, or collaboration) indicator score, averaged across students in a group for each moment, as the outcome variable. in analyzing phases, comparison always assessed one phase against the other two phases. the initial model (model 1) included the phase and the group as binary predictors, whereas model 2 (when appropriate) also included a phase-group interaction term. we looked for statistically significant coefficients for the phase predictor, the group predictor, and the phase-group interaction predictor. a significant phase coefficient suggested that the outcome differed between moments in that phase compared to moments not in that phase. a significant group coefficient indicated that the two groups differed on that outcome on average. a significant phase-group interaction suggested that the effect of phase on that outcome differed between the two groups. because the within-group correlations of observations may result in downwardly biased standard errors in an ordinary linear regression (moulton, 1986), we employed sandwich estimators for the standard errors to correct for this possibility. we conducted one follow-up analysis on the effect of phase on collaboration using just the group 1 data, given that group 2 was selected because they evidenced more collaborative behaviors than group 1, to determine whether any particular phase had a more significant effect on this outcome. finally, for the fourth part, we used line graphs to further consider change across sessions for each group’s efp and cps scores during each phase of problem solving. we looked for patterns of increase and/or decrease in each outcome across sessions. 3. results and discussion a total of 8-16 math moments were identified in each session of problem solving (see tables 2 and 3; see table 4 for an explanation and illustration of each phase of problem solving). moments in which students were disengaged or were not actively working on problem solving (e.g., learning to use one of the vmt tools) were not analyzed. although both groups worked for the same amount of time on each session, the number of math moments within problem solving sessions varied by group and by session, but not to a statistically significant degree. we report and discuss results by research question. renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 84| f l r 3.1 rq 1: what is the relative proportion of each phase of problem solving overall, and is the distribution of phases similar by session? do these proportions vary for two groups with different levels of collaboration? we begin addressing rq1 by overviewing findings from analyses of aggregated data from all four sessions and both groups. we examined the relative proportion of the three phases of problem solving, as well as the between-group differences in these proportions. 3.1.1 overall findings as shown in figure 2, both groups spent most of their time constructing. although we identified very few moments that included more than one phase of problem solving, these moments were counted in the analyses of all relevant phases. figure 2. relative proportions of phases of problem solving 3.1.2 between-group comparisons for both groups, we identified similar numbers of math moments corresponding to each phase of problem solving (see table 7). as such, it appeared that across the two groups, the frequency of engagement in each phase of problem solving was approximately the same. in other words, phase and group were not interacting. we suggest, however, this does not necessarily mean that the two groups were engaged in problem solving in the same way. they could potentially vary in how they engaged and the strategies they employed, given that the groups were purposefully selected for study based on differences in their collaboration. we addressed this question in rq2, by analyzing students’ efp and cps. table 7 groups’ math moments by phase phase group 1 group 2 z p n % of moments n % of moments renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 85| f l r exploring 12 20.7 12 26.1 -0.649 0.516 constructing 31 53.4 31 67.4 -1.439 0.150 checking 12 20.7 5 10.9 1.345 0.177 3.1.3 analyses by session overall, for both groups, the relative proportions of all phases of problem solving fluctuated considerably (figure 3). we included the data from the first session in our analyses as it was the students’ orientation to engaging with the vmt. similar to symonds et al.’s (2021) results, this meant that students' engagement in different sessions varied. we conjectured that this might also indicate that the phases engaged were influenced by the context (the topic, the prompts of the activity; opportunities created during the groups' engagement). indeed, collaborative contexts have been found to offer learners support for engagement that enables the development of cognitive, as well as social skills (romero-lópez et al., 2020; sun et al., 2022). figure 3. proportion of phases in each session. 3.2 rq 2: how do the two groups engage with the different phases of problem solving? specifically, what is the relation between each phase of problem solving and the efp and cps of each group? do these relations differ for the two groups? is there change across sessions? having examined the frequencies of each group’s engagement in different phases of problem solving in rq1, we then turned to considering how each of these phases related to students' cognitive and behavioral engagement both as individuals and as a group. rq2 examined how exploring, constructing, and checking were associated with participants' efp and cps scores. we began this investigation with analyses that used aggregated data from all sessions and both groups. an average efp score, representing the mean of the three efp, was calculated. then, both overall and across sessions, we addressed how each group engaged efp and cps during each phase of problem solving. 3.2.1 phases as shown in figure 4, different patterns were identified in students’ efp scores during each of the three phases of problem solving. we found that average efp, which combined the three categories of efp, was highest during exploring (0.85). the same pattern was true with each of the core efp (working memory, cognitive flexibility, inhibitory control): scores were higher in moments of exploring. scores were similar 0,00% 10,00% 20,00% 30,00% 40,00% 50,00% 60,00% 70,00% 80,00% 90,00% 100,00% session 1 session 2 session 3 session 4 exploring constructing checking renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 86| f l r during constructing and checking, but on average, as indicated by average efp, constructing was associated with slightly higher average efp than checking, and specifically with higher working memory and inhibitory control than checking. nevertheless, cognitive flexibility was higher during checking than during constructing. we also observed that, across all phases, working memory was always the highest-scoring efp. inhibitory control was the second highest overall, and cognitive flexibility was the lowest. figure 4. efp scores by phase of problem solving. cps scores also varied by phase, as shown in figure 5. overall, we observed that the three cps scores were high and did not differ substantially during moments of exploring; collaboration – the most developed stage of cps – was especially high in this phase. math moments that included exploring had the highest cps scores; exploring also was the phase in which the three stages of cps were more balanced. figure 5. cps scores by phase of problem solving. 0,85 0,75 0,71 0,95 0,82 0,77 0,84 0,63 0,67 0,75 0,75 0,67 0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1 exploring constructing checking ef p s co re efp total working memory cognitive flexibility inhibitory control 0,81 0,76 0,61 0,7 0,69 0,81 0,7 0,34 0,42 0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1 exploring constructing checking c p s sc o re participation cooperation collaboration renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 87| f l r analyses of each individual stage of cps revealed additional patterns. participation and cooperation scores were relatively similar for each of the three phases of problem solving, whereas collaboration was higher during exploring and much lower during construction and checking. these results appeared to indicate that the phase of students’ exploring afforded more opportunities for full behavioral engagement in cps practices than either of the two other phases of problem solving. 3.2.2 effect of phases in the next analyses, we examined whether the differences observed were significant. we assessed the effect of phases on efp and cps using regression (table 8). as shown in table 8, moments of exploring had significantly higher efp and cps than non-exploring moments, specifically for the indicators of average efp, working memory, participation, and collaboration. these findings complemented the observations of the previous section, suggesting that efp and cps scores were highest in moments of exploring. table 8 regressions, between group analyses, efp and cps model 1 model 2 effect predictors estimate se estimate se exploring average efp intercept .669*** .033 .647*** .036 exploring .129*** .038 .235*** .060 group .085* .038 .129** .045 interaction -.192* .076 r2 .039 .051ꜝ working memory intercept .698*** .050 .675*** .055 exploring .159*** .047 .269** .078 group .167** .053 .213*** .066 interaction -.199* .095 r2 .046 .053 cognitive flexibility intercept .566*** .075 .546*** .080 exploring .280** .097 .435*** .082 group -.013 .093 .029 .109 interaction -.241 .164 r2 .026† .030 inhibitory control n.s. participation renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 88| f l r intercept .709*** .029 .704*** .032 exploring .078* .033 .106 .055 group .100** .032 .111** .039 interaction -.050 .067 r2 .039 .040 cooperation intercept .676*** .027 .681*** .030 exploring .043 .028 .017 .055 group .211*** .029 .200*** .035 interaction -.047 .035 r2 .134 .135 collaboration intercept .302*** .032 .282*** .036 exploring .316*** .036 .414*** .048 group .255*** .043 .296*** .054 interaction -.177* .070 r2 .183 .190 constructing average efp intercept .696*** .038 .713*** .046 constructing -.001 .039 -.031 .062 group .092* .038 .053 .060 interaction .064 .078 r2 .016† .018 working memory intercept .742*** .054 .728*** .067 constructing -.02 .054 .005 .092 group .178*** .054 .209*** .076 interaction -.051 .106 r2 .029 .030 cognitive flexibility n.s. inhibitory control n.s renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 89| f l r participation intercept .721*** .036 .690*** .046 constructing .009 .033 .066 .056 group .103** .032 .174*** .052 interaction -.118 .066 r2 .028 .037 cooperation intercept .671*** .033 .677*** .041 constructing .025 .032 .013 .052 group .210*** .030 .195*** .053 interaction .024 .063 r2 .132 .132 collaboration intercept .394*** .044 .404*** .052 constructing -.049 .047 -.068 .065 group .279*** .045 .255*** .078 interaction .039 .095 r2 .096 .096 checking average efp intercept .700*** .031 .715*** .033 checking -.018 .056 -.093 .083 group .090* .038 .061 .041 interaction .195* .096 r2 .016† .025 working memory intercept .736*** .048 .748*** .050 checking -.026 .074 -.081 .117 group .173** .054 .152* .060 interaction .144 .124 r2 .029 .032 cognitive flexibility n.s. renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 90| f l r inhibitory control n.s participation intercept .741*** .026 .757*** .028 checking -.074 .054 -.150 .080 group .097** .032 .068* .034 interaction .200* .094 r2 .035 .048ꜝ cooperation intercept .656*** .027 .653*** .029 checking .138*** .035 .152** .056 group .227*** .029 .232*** .033 interaction -.038 .058 r2 .159 .159 collaboration intercept .336*** .031 .353**** .033 checking .154* .061 .070 .094 group .287*** .043 .255*** .048 interaction .217* .103 r2 .108 .115 note. we considered model 2 only when the change in r2 was significant, denoted by ꜝ. se = standard error. *** p < .001; ** p < .01; * p < .05. all estimated coefficients are unstandardized. in cases where model 1 is not a significant improvement over the null model (denoted by † in r2), but a predictor significantly differs from 0, we include the model in the table. n.s. indicates a model that does not significantly differ from the null model and has no statistically significant predictors. we also found a group-phase interaction for the effect of exploring on average efp. it appeared that because the group-exploring coefficient in model 2 was negative, exploring had less of an effect on average efp for group 2 than group 1. in fact, despite the positive group coefficient in model 2, group 1 had higher average efp in moments of exploring than group 2. one possible explanation for this finding was that group 2 just had consistently higher average efp, and we were seeing a ceiling effect. regardless, the model results suggested that exploring could have a strong and positive effect on average efp. in contrast, we found that moments of checking evidenced significantly higher cps than nonchecking moments, specifically for cooperation and collaboration in model 1. interestingly, in model 1 for participation, checking was not a significant predictor, although group was. however, model 2, which included the interaction term, was a significant improvement over model 1; in that model, the checking coefficient was significant and negative, whereas the group-checking interaction coefficient was significant and positive, and the group coefficient was not significant. model 2 suggests that both groups experienced similar levels of participation in moments other than checking, but group 1 experienced less participation in checking moments than group 2. one possible explanation for this finding is that group 2, which was selected for study because they evidenced stronger collaboration than group 1, may have had broader participation in moments of checking, whereas group 1 may have had uneven participation in moments of checking. renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 91| f l r in summary, the presence of exploring appears to strengthen students’ efp and cps scores, particularly collaboration, more than other phases. this finding points to the possible benefit of the development of collaborative skills and of having time and opportunities to explore. another potential takeaway is related to the finding that checking was an important predictor of cps stages, but not efp scores. this finding could suggest that moments of checking bring students together to work either in parallel or in coordination. however, the group-checking interaction for participation and the significant group coefficients for cooperation and collaboration suggest that groups also may operate differently in moments of checking. 3.2.3 comparing groups' efp and cps across phases for average efp, working memory, participation, cooperation, and collaboration, we observed significant and positive group coefficients across multiple phase models, suggesting that group 2 scored consistently higher overall on these indicators than group 1. in general, we did not observe significant phasegroup interactions in the models, suggesting that both groups were affected equally by each phase. however, for two models, average efp in exploring and participation in checking, we observed phase-group interactions, as discussed in the previous section. for these two indicators, we did not observe the same patterns or phase effects for both groups. specifically, exploring had a substantially larger effect on average efp for group 1 compared to group 2, whereas checking had a negative effect on participation for group 1 and a small positive effect on participation for group 2. 3.2.4 relative collaboration between phases recall that the key distinction between the two groups in our study was that group 2 evidenced substantially higher collaboration than group 1 overall. as stated above, both exploring and checking had significant effects on collaboration scores. however, we were curious about the relative importance of phase on collaboration, and, based on inspection of the data, we were concerned that a ceiling effect on group 2’s collaboration scores could obfuscate relationships by limiting the potential variance. therefore, we conducted a regression analysis for collaboration with just the group 1 data and all three phases as predictors. in this model (table 9), all three phases had statistically significant and positive coefficients. however, the 95% confidence interval for the exploring coefficient did not overlap with the 95% confidence interval for the constructing or checking coefficient. we concluded that the exploring phase was associated with significantly higher levels of collaboration than either the constructing or checking phases for group 1. additionally, the r2 value for this model was .288, indicating that phases explained 28.8% of the variance in group 1’s collaboration scores. table 9 regression, group 1 collaboration by phase 95% ci outcome predictors estimate se ll ul collaboration intercept -.072 .073 -.215 .072 exploring .768*** .094 .582 .953 constructing .368*** .077 .216 .520 checking .403*** .082 .241 .565 r2 .228 note. se = standard error; ci = confidence interval; ll = lower limit; ul = upper limit. *** p < .001; ** p < .01; * p < .05. all estimate coefficients are unstandardized. renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 92| f l r 3.2.5 across sessions comparison to further explore patterns of change, line graphs were employed to compare both groups' efp and cps scores over time (see figures 6 and 7). given the small number of sessions, we have limited ability to make conclusive statements explaining patterns. however, we could make a few observations based on the figures and our analysis of these data. figure 6. average efp scores across sessions for group 1 and group 2. figure 7. average collaboration scores of group 1 and group 2 during constructing moments for four vmt sessions. 0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1 session 1 session 2 session 3 session 4 to ta l e fp s co re s group 1 group 2 0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1 session 1 session 2 session 3 session 4 w o rk in g m em o ry group 1 group 2 0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1 session 1 session 2 session 3 session 4 c o gn it iv e fl ex ib ili ty group 1 group 2 0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1 session 1 session 2 session 3 session 4 in h ib it o ry c o n tr o l group 1 group 2 renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 93| f l r we first note that the two groups differed in exploring and checking by session. group 1 had no moments of exploring in sessions 3 and 4, while group 2 engaged in this phase in all sessions. no group engaged in checking in session 1, and group 1 engaged in checking in session 2, but not group 2, and both groups engaged in similar amounts of checking in sessions 3 and 4. in what follows, we focus on constructing, as it was not only the most common phase but also the only phase in which both groups were involved across all four sessions. as shown in figure 6, during constructing moments, the efp scores of both groups followed similar patterns of change across sessions, although the scores of group 2 in each session were generally higher than those of group 1. we also note that the two groups generally followed similar patterns of change in their efp and that none of the observed patterns are linear, nor do they show a clear increase or decrease over time; both groups' efp fluctuated from session to session. review of the line graphs for each of the stages of cps during constructing moments showed similar patterns of change; figure 7 provides an example for collaboration. as expected, collaboration scores were consistently higher for group 2 than for group 1. we call attention to the similarity of the groups' patterns of engagement and increases in their collaboration scores over time. these results suggest that, with time, the groups were becoming more collaborative. this analysis does not differentiate between the effect of time and the effect of mathematical topic. each group engaged in the problem solving sessions in the same order, so time and topic were confounded. therefore, it was not possible to determine whether the findings were a function of change over time or characteristics specific to each session. while spending time working together may progressively increase students’ efp and cps, the groups may also be influenced by the opportunities provided by the design of each activity. the fact that we saw similar patterns of variation may suggest that there was some shared feature or structure of the activity that was guiding the students’ behavior. 4. general discussion 0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1 session 1 session 2 session 3 session 4 group 1 group 2 renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 94| f l r we undertook this study to consider the process of middle school students’ momentary engagement during phases of collaborative mathematical problem solving. although students have been repeatedly found to enjoy and benefit from opportunities to work together on problem solving (e.g., featherstone et al., 2011; webb et al., 2019), van leeuwen and janssen’s (2019) review of collaborative activity and learning showed that collaboration does not always result in learning and can be difficult to facilitate. in our study design, we sought insight that could inform teachers about the cognitive and behavioral engagement of students’ online collaborative mathematics activity, and purposefully examined moments during which the students were engaged in mathematics. we selected for study two groups of students who in prior study had been identified as having high levels of interest in working collaboratively online with mathematics problems—students in both groups were continuously engaged in working with their peers on math across the four sessions of problem solving. studying interested youth enabled us to explore the potential for middle-schoolers’ productive engagement during the math moments of their work together. furthermore, because we studied students’ online collaboration in the vmt environment, we were able to examine their moment-to-moment work in the workspace, as well as their interactions in the chat, which allowed study of fluctuations that differs from those possible had our analyses focused on the outcomes of cognitive or behavioral engagement, faceto-face, or even using video footage. from rq1, we learned that constructing was the most frequently engaged phase of problem solving compared to the other phases of problem solving. however, rq 2 showed that exploring was associated with higher efp and cps scores, providing corroboration for boaler and selling’s (2017) results and suggesting that students who were encouraged to explore use of strategies in their work with mathematics developed deeper understanding, as well as positive feelings, compared to students who did not receive the same support. we also noted that scores for collaboration in particular were higher when the math moments involved exploring. while findings from rq 1 suggested that the frequency of the cognitive and behavioral engagement for each group in each phase of problem solving was approximately the same, we learned from rq 2 that how the groups engaged the phases varied. given that we selected these two groups for study because the students had been identified as interested in the collaborative math sessions, it is not surprising that our results indicated that in all phases of problem solving, across all sessions, students in both groups maintained high and continuous levels of participation (renninger & hidi, 2016). that the two groups also differed in their level of collaboration was expected because they were selected for study based on this information. however, group 2 also consistently scored higher than group 1 in cooperation in all phases of problem solving. although prior study has shown that collaboration develops through working with others on problem solving (bonotto 2005; gerson, 2008; kelton et al., 2018; stahl, 2013), and its importance for the development of ef has been noted (pollastri et al., 2013), we did not know how the processes associated with efp might unfold. our results showed that this effect most likely originates from higher efp engagement during exploring. moreover, our findings related to the stage of collaboration are particularly interesting. in each phase of problem solving, both groups of students had similar patterns of increase in their collaboration scores across the problem solving sessions. thus, although group 2 was more collaborative in general, and we do not know how details of each task affected collaboration behaviors per se, we also saw that both groups became more collaborative the more they worked together. these results are consistent with literature showing how groups become better at collaborating over time as they start to share mental models that make them more efficient and effective even when problems demand higher mental power (van den bossche, et al., 2011; zambrano, et al., 2019). we also observed that despite differences in the efp and cps scores of the two groups, their scores fluctuated in similar ways across sessions during moments of constructing. this finding leads us to wonder if there is some shared feature or structure of the problems (e.g., prompts for discussion, mathematics topic) that was guiding the groups’ behaviors (see lieber & graulich, 2020), as both groups of students received the same instructions and tasks in each problem-solving session. renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 95| f l r 4.1 implications for theory and practice our findings contribute to discussions of momentary engagement as both situation specific (e.g. nolen et al., 2015; symonds et al., 2021) and complex (dietrich, et al., 2022); they also provide an extension of the existing literature. focused study of the process of two groups of interested middle school students’ engagement during the math moments of their phases of their problem solving reveals relatively similar fluctuations and also highlights differences in how they are engaging. study of the students’ use of efp and stages of cps, moreover, provides insight into what might be expected of students at this age as they engage in group work that involves open-ended problem solving. these results confirm that the collaborative context of group work promotes attention to the task (e.g., hmelo-silver et al., 2018; pohl, 2020), enhances reasoning (e.g., barron, 2003), and promotes use of working memory (e.g., kirschner et al., 2018); they also underscore the importance of considering how students are engaging in this context. in addition, our findings suggest the potentially essential contribution of task features such as prompts to discuss math in mediating student cognitive and behavioral engagement. these data show that students vary in their cognitive and behavioral engagement in different phases of problem solving. they further point to the benefit of student group engagement in the phase of exploring, in particular, as exploring was associated with increased use of working memory and collaboration. our results also suggest that students may need support to collaborate during moments that include constructing and checking. as such, it may be critical to support teachers to attend to what a group is doing moment to moment, and specifically to variations in students’ behavior during different phases of problem solving. 4.2 limitations and future directions future study with additional student groups, who vary in their level of interest in mathematics, as well as by age, and for whom demographic information is available is clearly warranted. moreover, while for present purposes we aggregated study of moments of collaboration, a more qualitative exploration would provide a rich description of the context from which collaboration emerges. in addition to systematically examining the role of task features as determinants of fluctuations in momentary engagement during collaborative problem solving, we also suggest the utility for practitioners of additional analyses of the components of each efp as well as those of each stage of cps (represented by the questions used for assessment in our coding scheme rubric). analyses such as these would provide additional detail and insight about which behaviors account for and are contributing to how student groups are engaging in problem solving. keypoints this study is the first to report on the process of students' executive functions and collaborative problem-solving during phases of problem solving. the virtual math teams environment enabled assessments of students' moment-to-moment engagement in phases of problem-solving. study findings highlight the importance of assessing both what students are doing during problem solving, as well as how they are engaging. the relations between each phase of problem solving, categories of efp, and stages of cps vary. the problem-solving phase of exploring has a positive effect on use of executive functions and collaborative problem solving. renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 96| f l r acknowledgments the authors gratefully acknowledge jane huynh’s editorial assistance in preparing this manuscript for publication. we also are most appreciative of support for our collaboration from an earli (european association of research on learning and instruction) emerging field group grant and the jacobs foundation, awarded to jennifer symonds and ricardo böheim, as well funding for project research from the ef+math program of the advanced education research and development fund (aerdf) to k. ann renninger. opinions expressed in this article are those of the authors and do not necessarily represent views of the ef+math program or aerdf. references andrews-todd, j., & forsyth, c. m. (2020). exploring social and cognitive dimensions of collaborative problem solving in an open online simulation-based task. computers in human behavior, 104, 105759. https://doi.org/10.1016/j.chb.2018.10.025 bailey, b. a., andrzejewski, s. k., greif, s. m., svingos, a. m., & heaton, s. c. (2018). the role of executive functioning and academic achievement in the academic self-concept of children and adolescents referred for neuropsychological assessment. children, 5(7), 83. https://doi.org/10.3390/children5070083 barron, b. (2003). when smart groups fail. the journal of the learning sciences, 12(3), 307–359. https://doi.org/10.1207/s15327809jls1203_1 bishara, s., & kaplan, s. (2022). inhibitory control, self-efficacy, and mathematics achievements in students with learning disabilities. international journal of disability, development, and education, 69(3), 868– 887. https://doi.org/10.1080/1034912x.2021.1925878 boaler, j., & selling, s. k. (2017). psychological imprisonment or intellectual freedom? a longitudinal study of contrasting school mathematics approaches and their impact on adults' lives. journal for research in mathematics education, 48(1), 78–105. https://doi.org/10.5951/jresematheduc.48.1.0078 bonotto, c. (2005). how informal out-of-school mathematics can help students make sense of formal inschool mathematics: the case of multiplying by decimal numbers. mathematical thinking and learning, 7(4), 313–344. https://doi.org/10.1207/s15327833mtl0704_3 brookman-byrne, a., mareschal, d., tolmie, a. k., & dumontheil, i. (2018). inhibitory control and counterintuitive science and maths reasoning in adolescence. plos one, 13(6), e0198973-e0198973. https://doi.org/10.1371/journal.pone.0198973 caviola, s., colling, l. j., mammarella, i. c., & szűcs, d. (2020). predictors of mathematics in primary school: magnitude comparison, verbal and spatial working memory measures. developmental science, 23(6), e12957. https://doi.org/10.1111/desc.12957 cervera-crespo, t., & gonzález-alvarez, j. (2017). age and semantic inhibition measured by the hayling task: a meta-analysis. archives of clinical neuropsychology, 32(2), 198–214. https://doi.org/10.1093/arclin/acw088 chan, r. c., shum, d., toulopoulou, t., & chen, e. y. (2008). assessment of executive functions: review of instruments and identification of critical issues. archives of clinical neuropsychology, 23(2), 201– 216. https://doi.org/10.1016/j.acn.2007.08.010 clark, c. a., pritchard, v. e., & woodward, l. j. (2010). preschool executive functioning abilities predict early mathematics achievement. developmental psychology, 46(5), 1176–1191. https://doi.org/10.1037/a0019672 common core state standards initiative (ccssi) (2011). high school- geometry. common core state standards for mathematics, https://www.thecorestandards.org/math/content/hsg/ cook, c. r., thayer, a. j., fiat, a., & sullivan, m. (2020). interventions to enhance affective engagement. in a. l. reschly, a. j. pohl, & s. l. christenson (eds.), student engagement: effective academic, renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 97| f l r behavioral, cognitive, and affective interventions at school (pp. 203–237). springer cham. https://doi.org/10.1007/978-3-030-37285-9_12 cragg, l., & gilmore, c. (2014). skills underlying mathematics: the role of executive function in the development of mathematics proficiency. trends in neuroscience and education, 3(2), 63–68. https://doi.org/10.1016/j.tine.2013.12.00 devilliers, m. (2003). rethinking proof with the geometer’s sketchpad. emeryville, ca: key curriculum press. diamond, a. (2013). executive functions. annual review of psychology, 64, 135–168. https://doi.org/10.1146/annurev-psych-113011-143750 diamond, a., & ling, d. s. (2016). conclusions about interventions, programs, and approaches for improving executive functions that appear justified and those that, despite much hype, do not. developmental cognitive neuroscience, 18, 34–48. https://doi.org/10.1016/j.dcn.2015.11.005 dietrich, j., schmiedek, f., & moeller, j. (2022). academic motivation and emotions are experienced in learning situations, so let’s study them. introduction to the special issue. learning and instruction, 81, 101623. https://doi.org/10.1016/j.learninstruc.2022.101623 dong, a., jong, m. s. y., & king, r. b. (2020). how does prior knowledge influence learning engagement? the mediating roles of cognitive load and help-seeking. frontiers in psychology, 11, 591203-591203. https://doi.org/10.3389/fpsyg.2020.591203 eshuis, e. h., ter vrugte, j., anjewierden, a., bollen, l., sikken, j., & de jong, t. (2019). improving the quality of vocational students’ collaboration and knowledge acquisition through instruction and joint reflection. international journal of computer-supported collaborative learning, 14, 53–76. https://doi.org/10.1007/s11412-019-09296-0 featherstone, h., crespo, s., jilk, l. m., oslund, j. a., parks, a. n., & wood, m. b. (2011). smarter together! collaboration and equity in the elementary math classroom. national council of teachers of mathematics. fredricks, j. a., & mccolskey, w. (2012). the measurement of student engagement: a comparative analysis of various methods and student self-report instruments. in s. l. christenson, a. l. reschly, & c. wylie (eds.), handbook of research on student engagement (pp. 763–782). springer. https://doi.org/10.1007/978-1-4614-2018-7_37 fredricks, j. a., blumenfeld, p. c., & paris, a. h. (2004). school engagement: potential of the concept, state of the evidence. review of educational research, 74(1), 59–109. https://doi.org/10.3102/00346543074001059 gathercole, s. e., lamont, e., & alloway, t. p. (2006). working memory in the classroom. in s.j. pickering (ed.), working memory and education (pp. 219–240). academic press. https://doi.org/10.1016/b978012554465-8/50010-7 gerson, h. (2008). david's understanding of functions and periodicity. school science and mathematics, 108(1), 28–8. https://doi.org/10.1111/j.1949-8594.2008.tb17937.x hadwin, a. f., järvelä, s., & miller, m. (2011). self-regulated, co-regulated, and socially shared regulation of learning. in b. j. zimmerman, & d. h. schunk (eds.), handbook of self-regulation of learning and performance (pp. 65-84). routledge. harel, g., & sowder, l. (2005). advanced mathematical thinking at any age: its nature and development. mathematical thinking and learning, 7(1), 27–50. https://doi.org/10.1207/s15327833mtl0701_3 hill, c. e. (ed.) (2012). consensual qualitative research: a practical resource for investigating social science phenomena. american psychological association. hmelo-silver, c. e., kapur, m., & hamstra, m. (2018). learning through problem solving. in f. fischer, c. e. hmelo-silver, s. r. goldman, & p. reimann (eds.), international handbook of the learning sciences (pp. 210–220). routledge. huizinga, m., smidts, d. p., & ridderinkhof, k. r. (2014). change of mind: cognitive flexibility in the classroom. perspectives on language and literacy, 40(2), 31–35. jacques, s., & zelazo, p. (2001). the flexible item selection task (fist): a measure of executive function in preschoolers. developmental neuropsychology, 20(3), 573–591. https://doi.org/10.1207/875656401753549807 renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 98| f l r jeong, h., cress, u., moskaliuk, j., & kimmerle, j. (2017). joint interactions in large online knowledge communities: the a3c framework. international journal of computer-supported collaborative learning, 12, 133–151. https://doi.org/10.1007/s11412-017-9256-8 jose, r. g., samuel, a. s., & isabel, m. m. (2020). neuropsychology of executive functions in patients with focal lesion in the prefrontal cortex: a systematic review. brain and cognition, 146, 105633. https://doi.org/10.1016/j.bandc.2020.105633 karpicke, j. d. (2012). retrieval-based learning: active retrieval promotes meaningful learning. current directions in psychological science, 21(3), 157–163. https://doi.org/10.1007/s10648-012-9202-2 kasmer, l., & kim, o. k. (2011). using prediction to promote mathematical understanding and reasoning. school science and mathematics, 111(1), 20–33. https://doi.org/10.1111/j.1949-8594.2010.00056.x kelton, m. l., ma, j. y., rawlings, c., rhodehamel, b., saraniero, p., & nemirovsky, r. (2018). family meshworks: children’s geographies and collective ambulatory sense-making in an immersive mathematics exhibition. children's geographies, 16(5), 543–557. https://doi.org/10.1080/14733285.2018.1495314 kirschner, p. a., sweller, j., kirschner, f., & zambrano, j. r. (2018). from cognitive load theory to collaborative cognitive load theory. international journal of computer-supported collaborative learning, 13(2), 213–233. https://doi.org/10.1007/s11412-018-9277-y kuhn, d., capon, n., & lai, h. (2020). talking about group (but not individual) process aids group performance. international journal of computer-supported collaborative learning, 15(2), 179–192. https://doi.org/10.1007/s11412-020-09321-7 landis, j. r., & koch, g. g. (1977) the measurement of observer agreement for categorical data. biometrics, 33(1), 159–174. https://doi.org/10.2307/2529310 laureiro‐martínez, d., & brusoni, s. (2018). cognitive flexibility and adaptive decision‐making: evidence from a laboratory study of expert decision makers. strategic management journal, 39(4), 1031–1058. https://doi.org/10.1002/smj.2774 lee, k., ng, e. l., & ng, s. f. (2009). the contributions of working memory and executive functioning to problem representation and solution generation in algebraic word problems. journal of educational psychology, 101(2), 373–387. https://doi.org/10.1037/a0013843 letang, m., citron, p., garbarg‐chenon, j., houdé, o., & borst, g. (2021). bridging the gap between the lab and the classroom: an online citizen scientific research project with teachers aiming at improving inhibitory control of school‐age children. mind, brain and education, 15(1), 122–128. https://doi.org/10.1111/mbe.12272 lieber, l., & graulich, n. (2020). thinking in alternatives–a task design for challenging students’ problemsolving approaches in organic chemistry. journal of chemical education, 97(10), 3731–3738. https://doi.org/10.1021/acs.jchemed.0c00248 long, b., spencer-smith, m. m., jacobs, r., mackay, m., leventer, r., barnes, c., & anderson, v. (2011). executive function following child stroke: the impact of lesion location. journal of child neurology, 26(3), 279-287. https://doi.org/10.1177/0883073810380049 mann, t. d., hund, a. m., hesson‐mcinnis, m. s., & roman, z. j. (2017). pathways to school readiness: executive functioning predicts academic and social–emotional aspects of school readiness. mind, brain, and education, 11(1), 21–31. https://doi.org/10.1111/mbe.12134 marek, l. i., brock, d.-j. p., & savla, j. (2015). evaluating collaboration for effectiveness: conceptualization and measurement. the american journal of evaluation, 36(1), 67–85. https://doi.org/10.1177/1098214014531068 mccoy, d. c. (2019). measuring young children’s executive function and self-regulation in classrooms and other real-world settings. clinical child and family psychology review, 22(1), 63–74. https://doi.org/10.1007/s10567-019-00285-1 melzner, n., greisel, m., dresel, m., & kollar, i. (2020). regulating self-organized collaborative learning: the importance of homogeneous problem perception, immediacy and intensity of strategy use. international journal of computer-supported collaborative learning, 15(2), 149–177. https://doi.org/10.1007/s11412-020-09323-5 mercer, n., & sams, c. (2006). teaching children how to use language to solve maths problems. language and education, 20(6), 507–528. https://doi.org/10.2167/le678.0 renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 99| f l r mohammadhasani, n., & asadi, s. (2020). the investigation of the effect of computer supported collaborative learning (cscl) environment and dynamic mathematics software on trigonometric problem solving skill. technology of education journal, 14(4), 867–875. https://doi.org/10.22061/tej.2020.5964.2312 moulton, b. r. (1986). random group effects and the precision of regression estimates. journal of econometrics, 32(3), 385–397. https://doi.org/10.1016/0304-4076(86)90021-7 nolen, s. b. (2020). a situative turn in the conversation on motivation theories. contemporary educational psychology, 61, 101866. https://doi.org/10.1016/j.cedpsych.2020.101866 nolen, s. b., horn, i. s., & ward, c. j. (2015). situating motivation. educational psychologist, 50(3), 234– 247. https://doi.org/10.1080/00461520.2015.1075399 peng, p., namkung, j., barnes, m., & sun, c. (2016). a meta-analysis of mathematics and working memory: moderating effects of working memory domain, type of mathematics skill, and sample characteristics. journal of educational psychology, 108(4), 455–473. https://doi.org/10.1037/edu0000079 phelps, e., & damon, w. (1989). problem solving with equals: peer collaboration as a context for learning mathematics and spatial concepts. journal of educational psychology, 81(4), 639–646. https://doi.org/10.1037/0022-0663.81.4.639 pohl, a. j. (2020). strategies and interventions for promoting cognitive engagement. in a. l. reschly, a. j. pohl, & s. l. christenson (eds.), student engagement: effective academic, behavioral, cognitive, and affective interventions at school (pp. 253-280). springer cham. https://doi.org/10.1007/978-3-03037285-9_14 pollastri, a. r., epstein, l. d., heath, g. h., & ablon, j. s. (2013). the collaborative problem solving approach: outcomes across settings. harvard review of psychiatry, 21(4), 188–199. https://pubmed.ncbi.nlm.nih.gov/24651507/ polya, g. (1945). how to solve it. princeton university press. ponitz, c. c., mcclelland, m. m., matthews, j. s., & morrison, f. j. (2009). a structured observation of behavioral self-regulation and its contribution to kindergarten outcomes. developmental psychology, 45(3), 605–619. https://doi.org/10.1037/a0015365 radvansky, g. a., & copeland, d. e. (2006). memory retrieval and interference: working memory issues. journal of memory and language, 55(1), 33–46. https://doi.org/10.1016/j.jml.2006.02.001 rau, m. a., & matthews, p. g. (2017). how to make ‘more’ better? principles for effective use of multiple representations to enhance students’ learning about fractions. zdm mathematics education, 49, 531– 544. https://doi.org/10.1007/s11858-017-0846-8 ray-riek, m. (2013). powerful problem solving: activities for sense-making with the mathematical practices. portsmouth, nh: heineman. renningr, k. a., corven, j., de dios, m.c., hogan, m. r., kyaw, m.h., michels, a.g., nakayama, m., werneck, h., & yared, f. (manuscript in preparation). collaborative problem solving and executive functions in middle-schoolers’ work in the virtual math teams environment. renninger, k. a., & hidi, s. e. (2016). the power of interest for motivation and engagement. routledge. rogat, t., hmelo-silver, c., cheng, b., traynor, a., adeoye, t., gomoll, a., & downing, b. (2022). a multidimensional framework of collaborative groups’ disciplinary engagement. frontline learning research, 10(2), 1–21. https://eric.ed.gov/?id=ej1369028 romero-lópez, m., pichardo, m. c., bembibre-serrano, j., & garcía-berbén, t. (2020). promoting social competence in preschool with an executive functions program conducted by teachers. sustainability, 12(11), 4408. https://doi.org/10.3390/su12114408 salmela‐aro, k., upadyaya, k., cumsille, p., lavonen, j., avalos, b., & eccles, j. (2021). momentary taskvalues and expectations predict engagement in science among finnish and chilean secondary school students. international journal of psychology, 56(3), 415–424. https://doi.org/10.1002/ijop.12719 salminen-saari j. f. a., garcia moreno-esteva, e., haataja, e., toivanen, m., hannula, m. s., & laine, a. (2021). phases of collaborative mathematical problem solving and joint attention: a case study utilizing mobile gaze tracking. zdm math education, 53(4), 771–784. https://doi.org/10.1007/s11858-02101280-z). sankaranarayanan, r., kwon, k., & cho, y. (2021). exploring the differences between individuals and groups during the problem-solving process: the collective working-memory effect and the role of renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 100| f l r collaborative interactions. journal of interactive learning research, 32(1), 43–66. https://psycnet.apa.org/record/2021-80316-002 schoenfeld, a. h. (1992). on paradigms and methods: what do you do when the ones you know don't do what you want them to? issues in the analysis of data in the form of videotapes. the journal of the learning sciences, 2(2), 179-214. https://doi.org/10.1207/s15327809jls0202_3 schoenfeld, a. h. (2016). learning to think mathematically: problem solving, metacognition, and sense making in mathematics (reprint). journal of education, 196(2), 1–38. https://doi.org/10.1177/002205741619600202 siegler, r. s. (1998). emerging minds. oxford university press. skaguerlund, k., bolt, t., nomi, j. s., skagenholt, m., västfjäll, d., träff, u., & uddin, l. q. (2019). disentangling mathematics from executive functions by investigating unique functional connectivity patterns predictive of mathematics ability. journal of cognitive neuroscience, 31(4), 560–573. https://doi.org/10.1162/jocn_a_01367 skinner, e. a., & pitzer, j. r. (2012). developmental dynamics of student engagement, coping, and everyday resilience. in s. l. christenson, a. l. reschly, & c. wylie (eds.), handbook of research on student engagement (pp. 21-44). springer. https://doi.org/10.1007/978-1-4614-2018-7_2 stahl, g. (2013). translating euclid: designing a human-centered mathematics. springer cham. https://doi.org/10.1007/978-3-031-02200-5 straub, s., & rummel, n. (2021). promoting regulation of equal participation in online collaboration by combining a group awareness tool and adaptive prompts. but does it even matter? international journal of computer-supported collaborative learning, 16(3), 67–104. https://doi.org/10.1007/s11412-02109340-y su, y., li, y., hu, h., & rosé, c. p. (2018). exploring college english language learners’ self and social regulation of learning during wiki-supported collaborative reading activities. international journal of computer-supported collaborative learning, 13(1), 35–60. https://doi.org/10.1007/s11412-018-9269-y sun, j., anderson, r. c., lin, t. j., morris, j. a., miller, b. w., ma, s., nguyen-jaheil, k. t., & scott, t. (2022). children’s engagement during collaborative learning and direct instruction through the lens of participant structure. contemporary educational psychology, 69, 102061. https://doi.org/10.1016/j.cedpsych.2022.102061 swanson, h. l., & beebe-frankenberger, m. (2004). the relationship between working memory and mathematical problem solving in children at risk and not at risk for serious math difficulties. journal of educational psychology, 96(3), 471–491. https://doi.org/10.1037/0022-0663.96.3.471 symonds, j. e., kaplan, a., upadyaya, k., salmela-aro, k., torsney, b., skinner, e. & eccles, j. s. (2021). momentary engagement as a complex dynamic system. psyarxiv. https://doi.org/10.31234/osf.io/fuy7p symonds, j. e., schreiber, j. b., & torsney, b. m. (2019). silver linings and storm clouds: divergent profiles of student momentary engagement emerge in response to the same task. journal of educational psychology, 113(6), 1192–1207. https://doi.org/10.1037/edu0000605 van den bossche, p., gijselaers, w., segers, m., woltjer, g., & kirschner, p. (2011). team learning: building shared mental models. instructional science, 39(3), 283–301. https://doi.org/10.1007/s11251010-9128-3 van leeuwen, a., & janssen, j. (2019). a systematic review of teacher guidance during collaborative learning in primary and secondary education. educational research review, 27, 71-89. https://doi.org/10.1016/j.edurev.2019.02.001 vandenberg, j., zakaria, z., tsan, j., iwanski, a., lynch, c., boyer, k. e., & wiebe, e. (2021). prompting collaborative and exploratory discourse: an epistemic network analysis study. international journal of computer-supported collaborative learning, 16(3), 339–366. https://doi.org/10.1007/s11412-02109349-3 veraksa, a., bukhalenkova, d., & almazova, o. (2020). executive functions and quality of classroom interactions in kindergarten among 5-6-year-old children. frontiers in psychology, 11, 603776–603776. https://doi.org/10.3389/fpsyg.2020.603776 renninger et al __________________________________________________________________________ special issue: perspectives on momentary engagement and learning situated in classroom contexts 101| f l r verbruggen, f., & logan, g. d. (2008). automatic and controlled response inhibition: associative learning in the go/no-go and stop-signal paradigms. journal of experimental psychology, 137(4), 649–672. https://doi.org/10.1037/a0013170 viterbori, p., traverso, l., & usai, m. c. (2017). the role of executive function in arithmetic problemsolving processes: a study of third graders. journal of cognition and development, 18(5), 595–616. https://doi.org/10.1080/15248372.2017.1392307 webb, n. m., franke, m. l., ing, m., turrou, a. c., johnson, n. c., & zimmerman, j. (2019). teacher practices that promote productive dialogue and learning in mathematics classrooms. international journal of educational research, 97, 176–186. https://doi.org/10.1016/j.ijer.2017.07.009 yeniad, n., malda, m., mesman, j., van ijzendoorn, m. h., & pieper, s. (2013). shifting ability predicts math and reading performance in children: a meta-analytical study. learning and individual differences, 23, 1-9. https://doi.org/10.1016/j.lindif.2012.10.004 younger, j., o'laughlin, k., anguera, j., bunge, s., ferrer, e., hoeft, f. mccandliss, b., mishra, j., rosenberg-lee, m., gazzaley, a., & uncapher, m. (2023). better together: novel methods for measuring and modeling development of executive function diversity while accounting for unity. frontiers in human neuroscience, 17, https://doi.org/10.3389/fnhum.2023.1195013 zambrano, j., kirschner, f., sweller, j., & kirschner, p. a. (2019). effects of group experience and information distribution on collaborative learning. instructional science, 47(5), 531–550. https://doi.org/10.1007/s11251-019-09495-0 zhang, s., chen, j., wen, y., chen, h., gao, q., & wang, q. (2021). capturing regulatory patterns in online collaborative learning: a network analytic approach. international journal of computer-supported collaborative learning, 16(1), 37–66. https://doi.org/10.1007/s11412-021-09339-5 codepen discussion frontline learning research special issue vol.9 no.2 (2021) 170-178 issn 2295-3159 discussion to the special issue supporting the transition of a diversity of students: developing the “whole student” during and beyond their time at higher education. jacques van der meer1 1university of otago, new zeeland info corresponding author email: jacques.vandermeer@otago.ac.nz doi: https://doi.org/10.14786/flr.v9i2.785 introduction apart from the current covid19 context, the higher education sectors across the world have been faced with major challenges over the last few decades (auerbach et al., 2018; haggis, 2004), including increased numbers and diversity. considering the many challenges in higher education, especially the rise of students’ mental health issues, i am strongly convinced that education sectors, but in particular the higher education sector, have a societal responsibility to not just focus on students as learners of knowledge and/or professional skills, but to support them in being developed as “whole students”. all these challenges also raise a need for research into the broader context to identify how we can better support the diverse student population as they transition into higher education, but also how to prepare them for a positive experience during and beyond their time in higher education. overall, it can be said that the contributions to this special issue beneficially addressed some of the main foci to widening the perspectives on diversity related to the transition into higher education. the contribution came from different european countries, including belgium, germany, italy, sweden, switzerland, the netherlands and the united kingdom. de clercq et al. (in this special issue) indicated that environmental characteristics, such as distinctiveness of countries, is often overlooked in research. in this discussion article, therefore, some particular references will also be made to a specific country, new zealand. this may be of interest and relevant for the particular questions raised in this special issue as focusing on student diversity in educational contexts has been considered important for some time in this country. aoteraroa new zealand is a country in the south pacific colonised by europeans in the 19th century. in the second part of the 20th century, the focus across the new zealand education sectors, including higher education, started to develop beyond just a european perspective, and started to focus more on recognition of student diversity. initially, the main focus was on the indigenous population, the māori people. in the last few decades of the 20th century, the focus was extended to the pacific island people, many of whom migrated to new zealand from a wide range of different islands in the south pacific. in the 21st century, the focus on culturally and linguistically diverse (cald) groups was further extended, and over the last decade also because of the increase of refugees from the middle east and asia. providing some insights from the other end of the world, in quite a different and de-colonised ex-european nation may help european (and other) countries to reflect on their own approaches. the whole student considering the main focus of this special issue on the diversity of students’ transitioning, the importance of considering multiple aspects and system levels that may impact on this, and the many challenges in the higher education sectors, the approach of developing the “whole student” and related research approaches can provide a way to bring all this together. making an argument that developing the “whole student” is beneficial not only for a diversity of students, but also for the higher education institutions and overall society, i believe is very important. however, this is a very complex and broad area that could cover hundreds of pages. in this discussion article, i will touch on some of the key facets related to this, with particular references to the articles in this special issue. hopefully this may prompt more interest and research, and broader research approaches in this area. there are different ways to describe “whole student development”. i propose describing this as: the development of students’ intellectual, emotional, social and well-being capacities, aimed at preparing them for academic success, employability, civic-mindedness and a positive resilient holistic wellbeing future. holistic wellbeing can be described in many ways. i suggest that this could include physical, mental, social and spiritual aspects. spirituality could be considered as having a sense that life has meaning and purpose (even when confronted with adversity) and a sense of identity, self-awareness and connectedness with others, nature and the universe. in the following sections, i will relate the whole student with the results of the special issue, in particular regarding its widened perspective of diversity, its contribution to the 21st century learning and teaching approach and its integration of multi-level and longitudinal research approaches to address the transition to higher education. an expanded notion of diversity to consider the whole student considering the whole student asks for an extended notion of diversity in the literature on higher education. diversity has to be considered at multiple levels, not just related to student backgrounds and characteristics (micro level), but also at the institutional (meso) level and the broader socio-cultural and national (macro) level. in their article, baloo et al. (2021) point out that we need to move away from a deficit approach of just focusing on student backgrounds when considering academic achievement. they clarify that only by fully exploring students’ differential outcomes at all levels, we can more comprehensively understand the role of student diversity in educational transitions. messina dahlberg et al. (2021) argue that we have to move away from a deficit approach, and recognise competencies of students, in particular refugee students. again, messina dahlberg et al. (2021) emphasize the importance and support of diversity at different levels. at a micro level, they say it is very important that first-year students are supported to develop a greater sense of their identity, and to feel that they matter. in their transition and during their first year in higher education, to feel that they matter helps them to develop a sense of belonging, which contributes to their overall wellbeing. messina dahlberg et al. also point out that diversity does not just reside in a single individual, but is also situated at a meso (institutional) level. they clarify that an explicit policy focused on equity and diversity is important. as an example of this, in their equity and diversity policy, the university of otago articulate that it “is committed to equity and diversity and seeks to provide an accessible, inclusive, respectful and welcoming environment in which all students and staff are supported towards achieving their full potential”. (university of otago, n.d.). that policy also further indicates the range of diversity. considering the new zealand context, ethnicity comes first, with a particular focus on the importance of the role of the indigenous māori population. the pacific island population is also specifically mentioned. then other groups are mentioned: students and staff with disability and/or impairment; students who are first in their family to attend university; lgbtiq students and staff; students from low socio-economic backgrounds; students and staff from migrant and/or refugee backgrounds and those whose first language is not english, and women where there are barriers to access and/or success. messina dahlberg et al. (2021) also discuss the importance of the role of the universities in the rites of passage of diverse students as they enter university, especially students of minority groups such as migrant and refugee students. at the university of otago, in the formal academic convocation (welcoming ceremony) of all first-year students (in the local stadium at the beginning of orientation week), the diverse student groups are explicitly mentioned and it is emphasized that the university will actively support them all in their diversity. this starts the diverse students’ population sense that they all matter. at a macro level, ideally there are relevant national policies that entice or force higher education institutions to develop equity and diversity. this is the case in new zealand with a wide range of policies. developing a greater focus on 21st century life-long active learning teaching approaches a greater focus on life-long and active learning can contribute significantly to the whole student, in particular their transition into their first year and academic and social integration because it enhances their peer interaction and development of friendships. braxton et al. (2000) focused in particular on these benefits for students: “active learning course practices may directly influence social integration and indirectly affect subsequent institutional commitment and student departure decisions” (braxton et al., 2000, p.572). in their article in this special issue, willems et al. made a clear argument that a focus on learning strategies may be important. they point at the relevance of learning activities that lead to meaningful learning and more in-depth understanding of the learning content. active learning approaches indeed can enhance “deep learning” (haack, 2008). they also emphasize self-regulation. although their study found that students with deep learning approaches and good self-regulation at the end of secondary education did not necessarily do better in their academic adjustment, i would argue that deep learning and self-regulation may need to be developed explicitly for specific higher education contexts. active learning approaches by higher education teachers may contribute to this. self-efficacy can also be considered important in developing life-long learning. willems et al. (2021) showed in their research that students’ self-efficacy was positively associated with academic adjustment. it had a positive impact on academic achievement, both directly and indirectly, through academic adjustment. they also found that it had a direct positive impact on academic achievement both for students who were enrolled in more academically focused programmes and those who were enrolled in vocational/professional courses. the importance and the role of self-efficacy was discussed in many other studies in this special issue. with their multifactorial approach, de clercq et al. (2021) found that self-efficacy beliefs was one of the strongest predictors of academic achievement. in their research bohndick et al. (2021) also identified the importance of self-efficacy beliefs. in their focus on how differences in the perception of first-year requirements can be explained, they showed that it mostly depended on the individual factors such as self-efficacy and volition. jenert and brahm (2021) found in their quantitative study a general decline in students´ overall motivation and self-efficacy as well as an increase in study-related anxiety over the first year. however, some students scored high on anxiety despite their high motivation and self-efficacy. they also considered these factors in their qualitative study and found that students who scored high in these constructs generally did well in mastering the challenges of the first year. supporting first-year and ongoing transition through development of learning power and learning disposition skills supporting first-year students in their transition to higher education can be done through many different approaches, and there has been quite some research focussed on this. in this special issue, various authors focused on different aspects. van der zanden et al. (2021) concluded that secondary teachers’ practices related to students’ social and emotional adjustment across the transition to university, but not to their academic achievement. and they articulated that teachers in secondary education might play a pivotal role in preparing students for university and that their teaching may have a long-term impact on first-year students’ social and emotional adjustment. bohndick et al. (2021) also make a point for the role of universal interventions to support students. they argue that individual diversity needs to be recognised but institutions need to be careful not to stigmatize particular social groups of students. focusing on students’ experience and response to academic requirements, they advocate for general transparency with regard to the specific demands of the different disciplines and study programmes in order to enable all students to better navigate the first-year challenges. de clercq et al. (2021) focused on differences between institutions or programs that could result in different higher education adjustment experiences. they also investigated the individual and contextual factors that impact on academic achievement. their findings support the theoretical assumption of the three-level framework of higher education. they show that micro-level characteristics of the transition must be considered but it is necessary to include meso and macro influences/factors. it is important to consider both immediate and more distal environmental settings and how they interact with individual characteristics. finally, they found that contextual differences predicted about 15 percent of academic achievement, but that this was different depending on students’ study programs. willems et al. (2021) also considered differences between academic programmes and professional programmes. they found that learning strategies and motivational variables at the end of secondary education were more predictive for students’ first-year adjustment in the academic programmes than in professional programmes. they further showed that academic adjustment in the first semester influenced academic achievement more in professional programmes than in academic programmes, even when controlling for students’ prior education. as they clearly articulate, students who are more self-regulated are able to actively steer their own learning processes through activities such as planning tasks, monitoring progress, and diagnosing problems. considering all the articles, it is clear that there is general agreement that supporting students in their transition into higher education is important and that various issues need to be addressed, including, for example, self-efficacy, motivation, self-regulation and emotional adjustment. i would also argue that this would support students to be effective learners not only in their first year, but depending on the approach, could also benefit the remainder of their time in higher education and beyond. enhancing peer interaction in the light of research on the important role of social integration in the transition process of first-years students, jenert and brahm, included that aspect in their study. peer interaction is also closely connected to the issue of diversity. through interacting with other diverse students, they can learn more about diversity, but also learn more about their own identity, and gain other benefits, such as improved communication between students in culturally diverse classes (keating et al, 2020). therefore, developing the whole student could benefit from a greater focus in curricula on developing students inter-relational engagement with other students, and their intercultural competency. this could be done for example by including more group activities during class (all classes, whether lectures, tutorials, workshops, labs etc.) and team/group projects outside of class. this may benefit students’ development of appreciation of diversity, inter-cultural communication, and overall sense of connectedness and sense of belonging (deardoff, 2004, 2006, 2009). einfalt (2020) also specifically identified that it promotes students’ transition. what was really interesting is that the results from jenert and brahm’s study suggest that when student diversity is taken into account, relationships between personal and contextual variables may not necessarily be predictable. their qualitative results, for example, showed that peer interaction counted among the factors which were experienced differently by different students. this is also reflected in studies showing that students from more collectivist cultures may benefit more greatly from peer interaction and collaboration. this was also found in my analysis of data from all new zealand universities related to students’ engagement. it was found that the ethnic groups of māori and pacific island students did indeed appreciate this more than other ethnic groups, such as new zealand european students (van der meer, 2011). integrated multi-level longitudinal research in many articles in this special issue, research related to diversity of students’ first-year transition reported that ideally, we need to take into account a wide range of variables from the micro, meso and macro levels, and also to cover multi-dimensional aspects. in their study, bohndick et al. (2021) clearly went beyond student characteristics and also considered social and organisational diversity factors, which they proved had an impact. it is important to gain holistic knowledge of the complex wide variety of potential factors that may impact on the broad range of diverse student groups in their transition into university. accordingly, many research approaches need to be used, both quantitative, qualitative and mixed-methods. jenert and brahm’s study (2021) is a good example. their study reflects the complexity of interactions between individual and contextual differences. their mixed method approach provides some good new insights into person-contextual interactions. one of their interesting findings was that there were participant profiles with consistent patterns across self-efficacy, anxiety, and motivation and that students who were assigned to the highly anxious profile were the least self-confident and least motivated. concerning the links between personal and contextual diversity, their findings showed that students with different profiles at the point of entry coped very differently with the challenges that they encountered in the transition to higher education. baloo et al. (2021) also made a great contribution regarding the consideration of different methodologies. they identified that there are significant challenges in how to address differential student outcomes in terms of academic achievements and first-year transition. they argue that the negative focus of many studies on differential outcomes is based on students’ deficiencies rather than broader contextual factors. so they provide ideas for far more nuanced analyses of institutional data that can enable universities to take a more proactive approach to monitoring and reducing differences in student success. for example, they point at the importance of interventions to be informed by evidence and identifying whether universal interventions are beneficial for all or may be particularly beneficial for specific student groups. exploring meso-level variables, in particular, is important, because these variables are under the control of the university. these provide the potential for analysing cross-institutional data. this is also an argument i made in the analysis of data from all new zealand universities with regards to the various aspects and benefits of student engagement (van der meer, 2011). baloo et al. (2021) suggested to use different regression techniques. it provides the options to consider what predicts particular outcomes, whether these are predictors related to students before they start university or once they are at university. at the university of otago, we made an effort to use multiple regression approaches (both linear and binary logistic regressions) to consider the impact of a first-year intervention on students’ first-year academic achievement, first-year retention and also degree completion. we used variables from different levels to find out what predicted the outcomes (van der meer et al., 2017). the intervention we studied was the peer assisted study sessions programme (pass). in this co-curricular programme, second or third year undergraduate students facilitate study sessions focused on supporting first-year students to develop study skills, problem solving and working together with other students in parallel to particular courses so that students can learn these skills in a relevant context for them. but the focus is absolutely not on re-teaching content. this programme is based on the us programme called supplemental instruction (si) which was developed in the 1970s with a particular focus on supporting the african american minority who finally were allowed to enter universities as a positive result of the civil rights movement. a lot of research has been done since to study the effectiveness of the si and pass interventions, including a systematic review of world-wide research (dawson et al., 2014). si was introduced in other countries as a result of massification and related diversity of first-year students. the name pass, rather than si, is used in most countries other than the us. drawing on effectiveness of pass/si research (dawson et al., 2014), we moved away at otago from just testing whether involvement in the pass programme had positive benefits or not. instead we included a ‘dose-response’ approach. this concept, based on medical studies, means that we assessed whether the number of pass sessions attended related to better outcomes. we also controlled for many other contextual factors, for example, students’ socio-economic background (which in new zealand is partly measured by the secondary school students attended). so, many variables related to the different levels, micro level (student characteristics), meso level (university based student support intervention) and macro level (socio-economic context). the results demonstrated that pass participation indeed impacted on students’ first-year academic achievement, first-year retention, and degree completion (within 6 years), and that the number of pass sessions attended predicted the level of positive outcomes. and the data also demonstrated that pass participation contributes to students’ success over and above their academic ability as measured by their secondary school results. one of the main reasons that participation likely did not just support their level of achievement in the courses to which their pass programme was related, but also contributed to retention and especially degree completion, is that a great focus of the programme is to help first-year students develop a range of academic skills, and connectedness with other students. in other words, it actively supports academic and social integration which lots of research indicate is of key importance to student success (see e.g. braxton et al., 2000; grillo & leist, 2013; terrion & daoust, 2011). conclusion based on all the articles in this special issue, and this discussion article, it can be argued that higher education institutions should ideally focus more on the broader perspective of how to support the diversity of students transitioning in their first year in higher education, and move away from fragmented research and deficit approaches. the focus of universities ideally would be to develop the “whole student” in order to support students’ overall wellbeing and development of skills both for their time at university and beyond. considering the universal increase of mental health issues of higher education students over the last decade or so, the major current and future challenging contexts for the younger generations, and the rapidly changing work environment, this is of enhanced importance. this is especially important if universities consider they have a definite responsibility to contribute to their current and future (g)local societies. in consequence, national governments should perhaps consider political initiatives to make this a legal requirement and provide related financial funding. in new zealand, the tertiary education strategy published recently in november 2020, the current government has included that “wellbeing is fundamentally entwined with learning, and needs to be a goal through all parts of our education system” (nz ministry of education, n.d.) to achieve many of the goals related to the whole student development, it is also important for governments to invest in broadly skilled educators so that they can meet the diverse and changing needs of students transitioning into higher education. furthermore ongoing longitudinal research evidence needs to be used to inform effective teaching and interventions to support the academic and wellbeing success of the diversity of students. in the nz 2020 tertiary education strategy, all this is explicitly articulated, including that, “new zealand needs an education and training system that prepares learners/ākonga for a changing world and the future of work”, and also the need to “develop staff capabilities to support teaching and learning practices that value languages, cultures and identities” (nz ministry of education, n.d.). developing the whole student also contributes to the united nations objective of the sustainable development goals. that is to ensure that “all learners, by 2030, acquire the knowledge and skills needed to promote sustainable development, including, among others, through education for sustainable development and sustainable lifestyles, human rights, gender equality, promotion of a culture of peace and non-violence, global citizenship and appreciation of cultural diversity and culture’s contribution to sustainable development” (cited in sala et al., 2020, p.13). to achieve these goals, it is important to promote well-being, provide inclusive and equitable quality education and promote lifelong learning skills and opportunities (bebbington & unerman, 2018). the recently developed lifecomp framework created by the european union (sala et al., 2020) can play an important part in supporting the different education sectors to make positive changes in their education contexts, including the higher education context. citations to this special issue balloo, k. et winstone, n.e. (2021). a primer on gathering and analysing multi-level quantitative evidence for differential student outcomes in higher education, frontline learning research, 9(2), 121-144. https://doi.org/10.14786/flr.v9i2.675 bohndick c., bosse, e. jänsch v.k. et barnat, m. (2021). how different diversity factors affect the perception of first-year requirements in higher education, frontline learning research, 9 (2),78-95. https://doi.org/10.14786/flr.v9i2.667 de clerq, m., galand, b., hospel, v. et frenay m. (2021). bridging contextual and individual factors of academic achievement: a multi-level analysis of diversity in the transition to higher education,frontline learning research, 9(2), 96-120. https://doi.org/10.14786/flr.v9i2.671 de clerq m., jansen e., brahm t.et bosse e. (2021), from micro to macro: widening the investigation of diversity in the transition to higher education frontline learning research, 9(2), 1-8. https://doi.org/10.14786/flr.v9i2.783 jenert, t. et brahm, t. (2021). the interplay of personal and contextual diversity during the first year at higher education: combining a quantitative and a qualitative approach, frontline learning research, 9(2),50-77.https://doi.org/10.14786/flr.v9i2.669 messina dahlberg, g., vigmo, s. et surian a. (2021). widening participation? (re)searching institutional pathways in higher education for migrant students the cases of sweden and italy, frontline learning research, 9(2), 145-169. https://doi.org/10.14786/flr.v9i2.655 van der zanden, p.j.a.c., denessen, e., cillessen a.h.n. et meijer, p.c. (2021). relationships between teacher practices in secondary education and first-year students’ adjustment and academic achievement, frontline learning research, 9(2), 9-27. https://doi.org/10.14786/flr.v9i2.665 willems, j., van daal, t., van petegem, p., coertjens, l et donche, v. (2021). predicting freshmen’s academic adjustment and subsequent achievement: differences between academic and professional higher education contexts, frontline learning research, 9(2), 28-49. https://doi.org/10.14786/flr.v9i2.647 references auerbach, r. p., mortier, p., bruffaerts, r., alonso, j., benjet, c., cuijpers, p., demyttenaere, k., ebert, d. d., green, j. g., hasking, p., murray, e., nock, m. k., pinder-amaker, s., sampson, n. a., stein, d. j., vilagut, g., zaslavsky, a. m., kessler, r. c., & who wmh-ics collaborators (2018). who world mental health surveys international college student project: prevalence and distribution of mental disorders. journal of abnormal psychology, 127(7), 623-638. bebbington, j., & unerman, j. (2018). achieving the united nations sustainable development goals. accounting, auditing & accountability journal, 31(1), . braxton, j. m., milem, j. f., & sullivan, a. s. (2000). the influence of active learning on the college student departure process: toward a revision of tinto's theory. the journal of higher education, 71 (5), 569-590. caena, f., (2019) developing a european framework for the personal, social & learning to learn key competence (lifecomp). literature review & analysis of frameworks. luxembourg: publications office of the european union. dawson, p., van der meer, j., skalicky, j., & cowley, k. (2014). on the effectiveness of supplemental instruction: a systematic review of supplemental instruction and peer-assisted study sessions literature between 2001 and 2010. review of educational research, 84(4), 609-639. deardorff, d. k. (2006). identification and assessment of intercultural competence as a student outcome of internationalization. journal of studies in international education, 10(3), 241-266. deardorff, d. k. (2009). exploring interculturally competent teaching in social sciences classrooms. enhancing learning in the social sciences, 2(1), 1-18. deardorff, d. k. (2004). in search of intercultural competence. international educator, 13(2), 13-15. einfalt, j. (2020). let’s talk about transcultural learning: using peer-to-peer interaction to promote transition and intercultural competency in university students. journal of academic language & learning 14(2), 20-39. grillo, m. c., & leist, c. w. (2013). academic support as a predictor of retention to graduation: new insights on the role of tutoring, learning assistance, and supplemental instruction. journal of college student retention: research, theory & practice, 15 (3), 387-408. haack, k. (2008). un studies and the curriculum as active learning tool. international studies perspectives, 9(4), 395-410. haggis, t. (2004). meaning, identity and ‘motivation’: expanding what matters in understanding learning in higher education?. studies in higher education, 29(3), 335-352. keating, m, rixon, a., & perenyi, a. (2020). deepening a sense of belonging: a las faculty collaboration to build inclusive teaching. journal of academic language & learning 14(2), 40-56. nz ministry of education. (n.d.). the statement of national education and learning priorities (nelp) and the tertiary education strategy (tes). retrieved from https://www.education.govt.nz/our-work/overall-strategies-and-policies/the-statement-of-national-education-and-learning-priorities-nelp-and-the-tertiary-education-strategy-tes sala, a., punie, y., garkov, v., & cabrera giraldez, m. (2020). lifecomp: the european framework for personal, social and learning to learn key competence . luxembourg: publications office of the european union. retrieved from: https://publications.jrc.ec.europa.eu/repository/bitstream/jrc120911/lcreport_290620-online.pdf terrion, j. l., & daoust, j. l. (2011). assessing the impact of supplemental instruction on the retention of undergraduate students after controlling for motivation. journal of college student retention: research, theory & practice, 13 (3), 311-327. university of otago, (n.d.). equity and diversity policy kaupapa here ararau tōkeke. retrieved from: https://www.otago.ac.nz/administration/policies/otago666398.html van der meer, j. (2011). māori and pasifika students’ academic engagement: what can institutions learn from the ausse data? in a. radloff (ed.), student engagement in new zealand’s universities (pp. 1-10). melbourne: australian council for educational research. van der meer, j., wass, r., scott, s., & kokaua, j. (2017). entry characteristics and participation in a peer learning program as predictors of first-year students’ achievement, retention, and degree completion. aera open, 3 (3). auerbach, r. p., mortier, p., bruffaerts, r., alonso, j., benjet, c., cuijpers, p., demyttenaere, k., ebert, d. d., green, j. g., hasking, p., murray, e., nock, m. k., pinder-amaker, s., sampson, n. a., stein, d. j., vilagut, g., zaslavsky, a. m., kessler, r. c., & who wmh-ics collaborators (2018). who world mental health surveys international college student project: prevalence and distribution of mental disorders. journal of abnormal psychology, 127(7), 623-638. bebbington, j., & unerman, j. (2018). achieving the united nations sustainable development goals. accounting, auditing & accountability journal, 31(1), . braxton, j. m., milem, j. f., & sullivan, a. s. (2000). the influence of active learning on the college student departure process: toward a revision of tinto's theory. the journal of higher education, 71 (5), 569-590. caena, f., (2019) developing a european framework for the personal, social & learning to learn key competence (lifecomp). literature review & analysis of frameworks. luxembourg: publications office of the european union. dawson, p., van der meer, j., skalicky, j., & cowley, k. (2014). on the effectiveness of supplemental instruction: a systematic review of supplemental instruction and peer-assisted study sessions literature between 2001 and 2010. review of educational research, 84(4), 609-639. deardorff, d. k. (2006). identification and assessment of intercultural competence as a student outcome of internationalization. journal of studies in international education, 10(3), 241-266. deardorff, d. k. (2009). exploring interculturally competent teaching in social sciences classrooms. enhancing learning in the social sciences, 2(1), 1-18. deardorff, d. k. (2004). in search of intercultural competence. international educator, 13(2), 13-15. einfalt, j. (2020). let’s talk about transcultural learning: using peer-to-peer interaction to promote transition and intercultural competency in university students. journal of academic language & learning 14(2), 20-39. grillo, m. c., & leist, c. w. (2013). academic support as a predictor of retention to graduation: new insights on the role of tutoring, learning assistance, and supplemental instruction. journal of college student retention: research, theory & practice, 15 (3), 387-408. haack, k. (2008). un studies and the curriculum as active learning tool. international studies perspectives, 9(4), 395-410. haggis, t. (2004). meaning, identity and ‘motivation’: expanding what matters in understanding learning in higher education?. studies in higher education, 29(3), 335-352. keating, m, rixon, a., & perenyi, a. (2020). deepening a sense of belonging: a las faculty collaboration to build inclusive teaching. journal of academic language & learning 14(2), 40-56. nz ministry of education. (n.d.). the statement of national education and learning priorities (nelp) and the tertiary education strategy (tes). retrieved from https://www.education.govt.nz/our-work/overall-strategies-and-policies/the-statement-of-national-education-and-learning-priorities-nelp-and-the-tertiary-education-strategy-tes sala, a., punie, y., garkov, v., & cabrera giraldez, m. (2020). lifecomp: the european framework for personal, social and learning to learn key competence . luxembourg: publications office of the european union. retrieved from: https://publications.jrc.ec.europa.eu/repository/bitstream/jrc120911/lcreport_290620-online.pdf terrion, j. l., & daoust, j. l. (2011). assessing the impact of supplemental instruction on the retention of undergraduate students after controlling for motivation. journal of college student retention: research, theory & practice, 13 (3), 311-327. university of otago, (n.d.). equity and diversity policy kaupapa here ararau tōkeke. retrieved from: https://www.otago.ac.nz/administration/policies/otago666398.html van der meer, j. (2011). māori and pasifika students’ academic engagement: what can institutions learn from the ausse data? in a. radloff (ed.), student engagement in new zealand’s universities (pp. 1-10). melbourne: australian council for educational research. van der meer, j., wass, r., scott, s., & kokaua, j. (2017). entry characteristics and participation in a peer learning program as predictors of first-year students’ achievement, retention, and degree completion. aera open, 3 (3). frontline learning research vol. 11 no. 2 (2023) 1-30 issn 2295-3159 corresponding author: lonneke boels, utrecht university, princetonplein 5, 3584 cc utrecht, the netherlands, l.b.m.m.boels@uu.nl doi: https://doi.org/10.14786/flr.v11i2.1139 assessing students’ interpretations of histograms before and after interpreting dotplots: a gaze-based machine learning analysis lonneke boels1,2*, alex lyford3*, arthur bakker1 & paul drijvers1 1utrecht university, the netherlands 2university of applied science utrecht, the netherlands 3middlebury college, middlebury, vt, usa *these authors share first authorship article received 25 july 2022/ revised 22 june 2023/ accepted 27 june 2023/ available online 5 september 2023 abstract many students persistently misinterpret histograms. literature suggests that having students solve dotplot items may prepare for interpreting histograms, as interpreting dotplots can help students realize that the statistical variable is presented on the horizontal axis. in this study, we explore a special case of this suggestion, namely, how students’ histogram interpretations alter during an assessment. the research question is: in what way do secondary school students’ histogram interpretations change after solving dotplot items? two histogram items were solved before solving dotplot items and two after. students were asked to estimate or compare arithmetic means. students’ gaze data, answers, and cued retrospective verbal reports were collected. we used students’ gaze data on four histogram items as inputs for a machine learning algorithm (mla; random forest). results show that the mla can quite accurately classify whether students’ gaze data belonged to an item solved before or after the dotplot items. moreover, the direction (e.g., almost vertical) and length of students’ saccades were different on the before and after items. these changes can indicate a change in strategies. a plausible explanation is that solving dotplot items creates readiness for learning and that reflecting on the solution strategy during recall then brings new insights. this study has implications for assessments and homework. novel in the study is its use of spatial gaze data and its use of an mla for finding differences in gazes that are relevant for changes in students’ task-specific strategies. keywords: statistics education; histogram and dotplot; eye-tracking; random forest; practice effect mailto:l.b.m.m.boels@uu.nl https://doi.org/10.14786/flr.v11i2.1139 boels et al. | f l r 2 1. introduction statistical literacy includes “people’s ability to interpret and critically evaluate statistical information, datarelated arguments […], which they may encounter in diverse contexts, and when relevant” (gal, 2002, p. 4; emphasis in original). as data consumers, citizens should be able to correctly interpret various graphical displays. this is particularly important in this era of vague and fake news “that place interpretive and evaluative demands on a reader or viewer” (gal & geiger, 2022, p. 2). in this study, we specifically focus on the graphical representation of histograms. histograms can reveal particular aspects of the distribution of the data often hidden in other graphs (e.g., pastore et al., 2017). furthermore, as histograms are ubiquitous in research and education, they need to be learned (cf. garfield & ben-zvi, 2008). for example, searching for ‘histogram’ in google scholar resulted in more than 3.2 million hits (june 13, 2023). therefore, the guidelines for assessment and instruction in statistics education ii (gaise ii) for all grades up to grade 12 contain several examples of histograms and dotplots (for levels a, b, and c), with levels b and c roughly corresponding to middle and high school (bargagliotti et al., 2020). moreover, some alternatives for histograms, such as boxplots, are even more complex (e.g., bakker et al., 2004). however, many people persistently misinterpret histograms (e.g., cooper, 2018; kaplan, 2014). for example, bakker (2004a) found that secondary school students (grades 7–8) considered the individual heights of bars in a histogram to be the heights of individual people, rather than aggregations of data. students’ conceptual difficulties with histograms are well documented (e.g., boels, bakker, van dooren & drijvers, 2019), but it is unclear how to support students in learning to interpret histograms. figure 1. example of dotplot item17 that required students to compare two datasets regarding their mean. several studies suggest that having students solve dotplot items can scaffold this learning (e.g., delmas & liu, 2005; garfield & ben-zvi, 2008; makar & confrey, 2004). in most of these studies, students’ answers and verbal reports were the main sources of information. dotplots have the advantage that they show all individual data points as well as their distribution (figures 1, 2). in addition, the absence of a vertical scale in dotplots can turn students’ attention toward the horizontal scale, which is where the variable is presented in both graphs. however, little is known about whether solving dotplot items allows students to become aware of aspects of graph representation and statistical variables that are useful for interpreting histograms. the aim of this study is, therefore, to explore how solving dotplot items influences secondary school students’ thinking on a detailed level when they interpret histograms. our overall research question is: in what way do secondary school students’ histogram interpretations change after solving dotplot items? in the theoretical background section, we will specify this overall question with three sub-questions. as we elaborate further in the theoretical background section, gaze data can reveal students’ strategies in real-time, and in more detail, compared to concurrent thinking aloud (verbal reports) and without the risk of influencing the thinking process (van gog et al., 2005; van gog & jarodzka, 2013). we use students’ gaze data when solving four items with histograms before and after solving similar items with dotplots, as well as their answers on these items. the four histogram items were taken from a larger sequence with 25 digital items in total. furthermore, we examined transcripts from stimulated recall (lyle, 2003) verbal reports about boels et al. | f l r 3 students’ strategies (for more details see section 3.3.2). in the next section, we elaborate on difficulties with histograms and dotplots and discuss how gaze data can be used. figure 2. example of a dotplot (left) and a histogram (right) depicting the same distribution. note. the dotplot was part of item16. the histogram was part of item05, not further discussed here (for more details, see boels, bakker, et al., 2022). 2. theoretical background 2.1 review of statistics education literature in this section, we review statistics education literature on the problem (many students persistently misinterpret histograms), a gap in this literature (the variation in results on students’ interpretations of dotplots), and graphs that are suggested for supporting students in learning to interpret histograms. figure 3. an example of a histogram. note. the measured variable (weight) is along the horizontal axis. the weights of 67 packages (sum of frequencies) are depicted in this histogram. the arithmetic mean weight is 3.3 kg. 2.1.1 histograms are persistently misinterpreted many people persistently misinterpret histograms (e.g., cohen, 1996; setiawan & sukoco, 2021). researchers and teachers think that there is no difference between bar graphs and histograms (e.g., clayden & croft, 1990; tiefenbruck, 2007). dabos (2014) found that some college teachers did not see when students incorrectly counted the number of bars in a histogram to get the total frequency instead of adding the bars’ heights. firstyear university students in educational sciences had difficulties finding or interpreting the mean, median, boels et al. | f l r 4 variation, and skewness in histograms (lem et al., 2013). college students interpreted the horizontal salary scale in a histogram as a timescale (meletiou, 2000). middle school students used unequal intervals in a histogram with frequency on the vertical axis—instead of density—hence, not correcting the frequencies for unequal bin widths (mcgatha et al., 2002). other middle school students thought that bars in histograms are connected for easier comparison (e.g., capraro et al., 2005). students in grades 6–12 answered histogram items 17% to 53% correctly on average (whitaker & jacobbe, 2017). many students mistakenly took bars’ heights as the measured value. such students possibly think that only nine packages are depicted in the histogram in figure 3 (the number of bars) instead of 67 (the actual number). 2.1.2 dotplots are not always correctly interpreted generally, dotplots are interpreted better than histograms (e.g., delmas et al., 2005), although stacked dotplots (in the early days also called line plots; e.g., tiefenbruck, 2007) might still confuse students (e.g., lyford, 2017). lem et al. (2013) found that university students understood dotplots slightly better than histograms (on average, 55% correct responses for dotplots versus 51% for histograms). however, in that study, two dotplot items scored worse. university students taking introductory statistics explored variability and standard deviation through a kind of stacked dotplots (delmas & liu, 2005). most of these students did not fully understand how standard deviation was related to the distribution of data in a histogram. figure 4. example of a double dotplot. note. this was item14 in the original sequence of 25 digital items. the question was: ‘which postal worker delivers the heaviest packages on average?’ with three answer options: (a) frans delivers the heaviest packages on average, (b) angela delivers the heaviest packages on average, and (c) the mean weight for both is approximately the same. the correct answer here is (c). a local instruction theory in statistics education suggests that dotplots are suitable for supporting students’ learning of distribution and variability in data represented in histograms (e.g., bakker & gravemeijer, 2004; garfield, 2002). garfield & ben-zvi (2008) stated: “studies [that] suggest a sequence of activities that leads students from […] dotplots […] to histograms” can support students in “developing the concept of distribution as an entity” (p. 175). an advantage of dotplots over histograms is that dotplots show the distribution of data in a disaggregated form. in addition, dotplots can draw students’ attention to the variable being depicted along the horizontal axis—similar to histograms—, as dotplots have only this axis. a possible disadvantage of dotplots for teaching students to interpret histograms (aggregated data) is that dotplots might invite them to see the data as individual cases (konold et al., 2015) instead of looking at aggregated measures (including arithmetic mean). one explanation for dotplots sometimes being misinterpreted is that students do not understand where the measured values are depicted in stacked dotplots due to the countable height of the stack. hence, students confuse frequency—the height of the bar or stack—with the measured value (e.g., cooper & shore, 2008; boels et al. | f l r 5 cooper 2018, kaplan et al., 2014), similar to histograms. for stacked dotplots, a vertical axis is possible but not necessary, which might induce this same height misinterpretation (e.g., lyford, 2017). some of the previously described persistent misinterpretations with histograms are related to the data in a histogram; more specifically:  where the measured variable is depicted—most often along the horizontal axis,  how many variables are measured—only one,  which variable is measured—the one along the horizontal axis (in a regular histogram). in non-stacked dotplots, no variable is represented in the vertical direction (see figures 1, 2, and 4). therefore, non-stacked dotplots have the potential to raise students’ awareness of the variable (here: weight) being depicted along the horizontal axis in dotplots and histograms. nevertheless, abstract dotplots without context or numbers along the horizontal axis can lead to misinterpretations of variability (kaplan et al., 2014). the same may hold for axis titles that are missing (such as in lem et al., 2013). thus, we use context, an axis title, and numbers in our dotplots. we conjecture that these non-stacked (‘messy’) dotplots support students’ histogram interpretations. therefore, as stated in the introduction, our overall research question is: in what way do secondary school students’ histogram interpretations change after solving dotplot items? 2.2 review of literature on eye-tracking in education in this section, we review what is already known from gaze data in education and what measures are most suitable for our aim. in addition, we elaborate on how gaze data can be connected to students’ strategies. we end each section with a sub-question. 2.2.1 use of spatial gaze measures to reveal students’ strategies for interpreting histograms the use of gaze data for studying learning is not new (e.g., strohmaier et al., 2020). for example, garcia moreno-esteva et al. (2018) and khalil (2005) studied students’ visual cognitive behaviors on statistical graphs. a main advantage of eye-tracking “is that it can provide detailed information about the time-course of processing” (kaakinen, 2021, p. 170). most studies neglect this level of detail by using gaze data measures that are temporal (e.g., total fixation duration, reaction times), count (fixation count, number of saccades between relevant or irrelevant parts of the stimuli), or both (e.g., kaakinen, 2021; lai et al., 2013). traditional time measures, for example, can hide visual scanning patterns (goldberg & helfman, 2010). a similar argumentation can be made for count measures such as the percent of fixations on specific parts of the screen (godau et al., 2014). spatial measures, such as a sequence of areas of interest (aois, e.g., garcia moreno-esteva et al., 2018, 2020) can disclose the kind of detailed information kaakinen (2021) refers to. spatial measures, such as scanpaths, seem better suited for providing detailed information about students’ thinking (hyönä, 2010). dewhurst et al. (2018) were one of the first who studied (simplified) scanpaths using vectors in (scene) viewing tasks. their vectors include the direction and magnitude of saccades. in a previous study, we qualitatively analyzed students’ scanpath patterns (sequence of fixations and saccades) when students estimated the mean from histograms (boels, bakker & drijvers, 2019). after qualitatively coding 300 videos with students’ gazes and verbal reports of 25 students in that study, we found that the perceptual form of students’ scanpath patterns within one aoi—the graph area—was most relevant for students’ task-specific strategies on these items, see figure 5. this perceptual form can be captured by the direction (angle) and magnitude (length) of students’ saccades. in that study, we found several scanpath patterns that were indicative of students’ task-specific strategies. all patterns were found on the graph area only. in one pattern, the perceptual form of that pattern was identified as vertical if successive saccades on the graph area were vertical and roughly aligned with each other (figure 5). this vertical scanpath pattern indicates that this student (correctly) tried to find the balancing point of the graph as an estimation of the mean. another scanpath pattern was a horizontal gaze pattern indicating that this student (incorrectly) tried to make all bars equally high which resulted in the mean of the frequencies instead of the mean weight. in total, five different scanpath patterns were found for students estimating and comparing means of histograms, each related to a specific strategy (boels, bakker, et al., 2022). other aois did not emerge as relevant to these students’ task-specific strategies. boels et al. | f l r 6 figure 5. example of a vertical scanpath on item20. note. circles indicate fixations (positions on the screen where students look longer), and thin lines between the circles indicate saccades (fast transitions between two fixations). a vertical line segment—indicating a scanpath—is superimposed for the reader’s convenience. a scanpath is a sequence of fixations and saccades. “a fixation is a period of time during which a specific part of [the computer screen] is looked at and thereby projected to a relatively constant location on the retina. this is operationalized as a relatively still gaze position in the eye-tracker signal implemented using the [tobii] algorithm.” (hessels et al., 2018, p. 22). (the figure has been translated into english.) as the overall research question for this study indicates, we want to explore in what way secondary school students learn from dotplot items. given that the scanpath patterns on the graph area indicate students’ strategies, we examine differences in these patterns on histogram items before and after students solved items with dotplots. we only address the main differences, those being differences relevant to students’ task-specific strategies. the first sub-question for the present study is, therefore: 1) what are the main differences in students’ gaze patterns on histogram items before and after solving dotplot items? 2.2.2 connecting gaze data to students’ strategies although scanpaths can reveal students’ strategies on a detailed level, there is no simple relation between eye movements and strategies (e.g., orquin & holmqvist, 2017; russo, 2010) as not every eye movement is part of a task-specific strategy (e.g., schindler & lilienthal, 2019). therefore, it is often needed to also ask at least some students what approach they took to solve the items. instead of concurrent think-aloud protocols, recalls (retrospective reports) are preferred for complex items (e.g., van gog et al., 2005) as concurrent thinking aloud may influence both eye movements and students’ thinking (van gog & jarodzka, 2013). the disadvantage of such retrospective think-aloud reports, however, is that students may have forgotten their strategy after completing all items. this risk can be reduced by having students look back at their eye movements (e.g., guan et al., 2006; kragten et al., 2015; van gog et al., 2005). therefore, in the stimulated recall (cued retrospective reports), we individually cued each student with their own gazes. how we did that, is explained in the data collection section. the second sub-question for the present study is: 2) what indications can be found in students’ verbalizations during stimulated recall that changes in their approaches to histograms occurred? 2.3 learning from a series of items: the practice effect students learning from a sequence of tasks is known as the practice, test-retest, or retesting effect in assessment theories (e.g., heilbronner et al., 2010; scharfen et al., 2018). the practice effect refers to improved performance (often scores or answers) due to repeated assessment with the same or similar, equally difficult items. the time interval between two assessments can be very short—5 or 10 minutes—to find such an effect (e.g., catron, 1978; falleti et al., 2006). the practice effect was found for several general cognitive function boels et al. | f l r 7 assessments for items that required memorization (e.g., of numbers), change detection (e.g., of changed colors between two items), and matching (e.g., what parts of items are alike). in addition, familiarity with test requirements can cause differences between the test and retest results (e.g., falleti et al., 2006) and reduce anxiety (e.g., catron, 1978). furthermore, regression to the mean can cause extreme results—high and low performance scores—to come closer to the mean, resulting in both underand overestimation of improvement (e.g., temkin et al., 1999). for achievement or knowledge tests, such as formative assessments in secondary education, the practice effect is also associated with actual or true learning as opposed to most cognitive tests, for example, iq tests, for which learning is unlikely to occur (e.g., lievens et al., 2007; scharfen et al., 2018). lumsden suggested that the practice effect can also be found within a sequence of items (1976). in addition, gaze data have been used to examine the practice effect (e.g., guerra-carrillo & bunge, 2018; płomecka et al., 2020). although this is not the focus of our study, to the best of our knowledge, our study is the first that looks at a within-a-sequence-of-items practice effect. most research investigating the practice effect uses scores on standardized tests (e.g., in this meta-analysis: hinton-bayre, 2010). however, standardized tests often lack instructional relevance (e.g., hohn, 1992). practitioners, such as mathematics teachers, are more interested in knowing whether students learn from a lowstake sequence of items. moreover, teachers are interested in students’ strategies, hence “gaining qualitative insight into student understanding” (bennett, 2011, p. 6). in this study, we, therefore, examine students’ changes in strategies during solving items as an indication of potential learning. to exclude several other possible influencing factors—such as peers’ or teachers’ interventions—we use items from one sequence of items with statistical graphs. for some items, students verbally reported their answer (estimation of the arithmetic mean), while for other items, they chose one of three answer options (comparing means, figure 4). if a change in students’ strategies occurred toward a correct instead of an incorrect strategy, we would also expect a difference in students’ answers, including answer correctness. therefore, the third sub-question for this research is: 3) what are the differences in students’ answers on histogram items before and after solving dotplot items? 2.4 rationale for using a machine learning algorithm for very small data sets or very short sequences of tasks, the first sub-research question could theoretically be answered through the careful, manual study of gaze data. our study, however, seeks to use machine learning to both augment the effectiveness of identifying differences in students’ gaze patterns between items and to identify these differences at a scale that would be impractical to do by hand. to build an analytical model of the gazes, a non-ml approach could be used. the ones we tried (e.g., logistic regression) performed relatively poorly (see also lyford & boels, 2022). instead, we use supervised learning, a subset of mlas that use training data and pattern recognition to predict a well-defined output (friedman et al., 2001). in particular, the present study uses the random forests algorithm (breiman, 2001), which will allow us to effectively and efficiently identify systemic differences in gazes between our two hundred student-item pairings. these random forests can not only be efficiently trained and used to identify patterns in students’ gaze data, but they are also likely to identify systematic differences in gaze data that are unnoticeable upon manual inspection (james et al., 2013). in addition, through assessing the importance of specific variables (figure 15), random forests allow for some interpretability so that researchers can better understand what some of the differences in gazes might be (e.g., proportionally more vertical instead of more horizontal gazes could indicate a change from an incorrect to a correct strategy), and postulate about possible mechanisms. 3. materials and methods details on participants, the eye-tracking method, and two items (item02 and item11) were reported previously in a qualitative study (boels, bakker, et al., 2022). two items were used previously in a machine learning analysis (item02 and item20; boels, garcia moreno-esteva, et al., accepted) but with a different aim, namely, to examine how a machine learning algorithm (mla) could identify students that used a correct or incorrect strategy—for solving the item—purely based on their gaze data on the graph area of this item. for the reader’s convenience, we here summarize all information relevant to the present study. boels et al. | f l r 8 3.1 participants: pre-university track students grades 10–12 participants were 50 grades 10–12 pre-university track students from a dutch public secondary school [15– 19 years old; mean = 16.31 years]; 23 males, 27 females (more details in table 1). in the netherlands, secondary school students are in a pre-vocational, pre-college, or pre-university track. generally speaking, the pre-university track implies mostly high-performing students. all participants had statistics in their mathematics curriculum. each student individually solved the items in a separate room in their school. participation was voluntary; permission from the utrecht university ethical committee was obtained, and informed consent was signed. participants received a small gift for their participation. table 1 grade level and age of participants. one participant did not provide details on grade, and another one did not provide age (see also boels, bakker, et al., 2022) grade number of participants age number of participants 10 20 15 12 11 17 16 19 12 12 17 10 unknown 1 18 7 total 50 19 1 unknown 1 total 50 note. due to legislation, data on ethnicity cannot be collected. in the netherlands, there is hardly any difference between public and private schools, nor between city, suburban, and rural schools. private schools are rare. 3.2 materials: histogram and dotplot items requiring comparing and estimating means 3.2.1 estimating and comparing arithmetic means reveals students’ knowledge to reveal strengths and flaws in students’ knowledge about data in graphs such as histograms, gal (1995) advises asking students to compute or estimate means from data in graphs. we, therefore, designed histogram and dotplot items that required students to estimate the arithmetic mean. this mean can be estimated from a histogram and dotplot by finding the equilibrium or balancing point from the graph (e.g., mokros & russell, 1995; o’dell, 2012; see also figure 5). in statistics, the mean is often used for comparing the data for two groups (e.g., gal, 1995; konold & pollatsek, 2002). we, therefore, added items for which a comparison of means was needed (see figure 4 for an example of a dotplot item that was used between the before and after histogram items). 3.2.2 four histogram items—dotplots items in between in the present study, we analyze students’ gaze data on four items from a sequence of items on a computer screen with several statistical graphs (dataset1). we chose two item pairs from this sequence that are suitable for analysis with a machine learning algorithm, as these items are very similar, see figure 6 (dataset1, 5). two items were given before a sequence of six dotplot items, and the other two afterward. note that the after items are mirrored versions of the before items. the first pair of items we examine—before item02 and after item20—require students to estimate the mean of the data in one histogram. we will henceforth refer to these as ‘single-histogram’ items. the question for both items was: what is, approximately, the mean weight of the packages that [anton/mo] delivers? the second pair of items—before item11 and after item21—require students to compare the mean of the data in two histograms. we will henceforth refer to these as ‘doublehistogram’ items. the question for both items was: which postal worker delivers the heaviest packages on average? for each item, three answer options were given: (a) [ellen/elizabeth] delivers the heaviest packages on average, (b) [titia/monsif] delivers the heaviest packages on average, and (c) the mean weights for both are approximately the same. the correct answer for both items is (c). boels et al. | f l r 9 figure 6. graphs of single-histogram items (left) and double-histogram items (middle and right) in the before (top row) and after (bottom row) versions. note. translated into english and numbering added. the numbering of the items (e.g., item11) refers to the numbering in the original sequence of 25 digital items (dataset1, boels, bakker, et al., 2022). each after item (bottom row) is a mirrored version of the before item (top row). six of the items between the items before and after were non-stacked dotplots that were specifically designed to scaffold students (items13–18 from the original data collection, e.g., figures 1, 2, and 4, dataset1, 5). as boels et al. | f l r 10 described in the theoretical background section, we used dotplots to draw students’ attention to specific features of the histograms that are important but might have been misunderstood. 3.3 data collection methods: eye-tracking, stimulated recall data from a previous qualitative study is re-used for this study (boels, bakker, et al., 2022). data collection included students’ answers on each item, xand y-coordinates of gaze data on the items through an eye tracker, and stimulated recall verbal reports. the collection of the gaze data and stimulated recall are briefly described in the following section. for more details, interested readers are referred to the original study. 3.3.1 data collection with an eye tracker a tobii xii-60 (sampling rate: 60 hz) was placed on an hp-probook-6360b laptop between the 13-inch screen (refresh rate 59 hz) and the keyboard, see figure 7. participants used a chin rest. gaze data were recorded and processed with tobii studio software version 3.4.5 (tobii, n.d.). tobii software’s calibration procedure consisted of a 9-point calibration. as this software has no built-in validation procedure, we included a validation screen in the set-up at the beginning, after each item, and at the end (more details in dataset1). we collected the raw data of the eye movements on each item (e.g., xand y-coordinates of the eyes on the screen for each time stamp as well as to which aoi these coordinates belong, see dataset1, 5, and figure 8) through the tobii software. we also collected students’ answers (verbal answers for single graph items, clicked multiple-choice option for double graph items). figure 7. set-up of the experiment. note. for each participant, a chin rest was used. the eye tracker was placed on the laptop (see the red oval below the screen, copied into the figure for the convenience of the reader). figure 8. aois of before item11 (left) and item02 (right). note. left: the graph area consists of the yellow and green areas named it11b_graph_l_ellen and it11b_graph_r_titia. right: the graph area is the light blue area named it2_graph. no students were excluded from the data set, as the data loss per trial (averaged over all 50 participants) and the data loss per participant (averaged over all 25 items from the original dataset1) were below the exclusion point (34% or more). the mean accuracy is 56.6 pixels (1.16°) with the highest accuracy on the most relevant part for our study: the graph area (middle of the screen; 13.4 pixels or 0.27°). the average precision (0.58°) is boels et al. | f l r 11 considered good. more details on accuracy, precision, and the eye tracker can be found in dataset3 and boels, bakker, et al. (2022) in line with advice from holmqvist et al. (2022). the design of the complete sequence of items (25 items in total), including files used in the tobii studio software, aoi sizes, and output, are available from a data repository (dataset1). 3.3.2 data collection through stimulated recall verbal reports stimulated recall (lyle, 2003) is also known as “cued retrospective reporting” (van gog et al., 2005, p. 273). it is called retrospective “own-perspective video think-aloud with eye-tracking” (mcintyre, 2022, p.4) when used with a head-mounted eye tracker. the first part of the verbal reports consisted of cued retrospective thinkaloud. this means that students watched videos of their own gazes laid over the items, while they explained their thinking when they solved the items. the stimulated recall took place after the students had solved all items of the sequence of 25 items (dataset 1). during the second part of the verbal reports, clarifying questions were asked such as why they stated that their previously given answer was incorrect. in this second part, participants were also confronted with inconsistencies in their reports, such as differences between the answer given during recall and the answer during item solving. time constraints influenced how many items could be questioned when students reported verbally. during this stimulated recall, we illuminated the location where students looked—through a kind of spotlight—and made the rest of the graph darker (see also boels, bakker, et al., 2022). we preferred this method over having students look back at their fixations (e.g., red dots) for two reasons. first, it prevents students from making different eye movements when looking back—and describing the corresponding strategy—instead of the strategy they initially used. second, this makes visible the exact information that the learner has looked at, instead of the information being covered by, for example, a red dot (the fixation; e.g., jarodzka et al., 2013). 3.4 data analysis through a machine learning algorithm we used different methods for analyzing our data. for the first sub-question about differences in gaze data, we analyzed our data through a machine learning algorithm (mla). for the second sub-question about changes in students’ strategies, we coded transcripts of verbal reports (for the codebooks see, boels, bakker, et al., 2022). for the third sub-question about students’ answers, we explored changes in answers and answer correctness. in the remainder of this section, we elaborate on the analysis with a machine learning algorithm. studies usually only report on successful approaches. as a result, other researchers keep reinventing the wheel. for the first sub-question, we, therefore, decided to report both the mla approaches we tried: our failed attempt to use time metrics as inputs for the mla and a successful approach with spatial metrics. before applying a machine learning algorithm, we first wanted to get a better understanding of the underlying data. therefore, first, we plotted a graph using the time metric total fixation time per aoi (also known as total dwell time or total fixation duration). next, we used this same time metric as input for training our mla. this approach failed to produce an accurate mla. moreover, although this time metric is commonly used, recent literature strongly advises against using total dwell time (orquin & holmqvist, 2017). second, we examined saccade directions and magnitudes (spatial metrics). finally, using these spatial metrics as inputs for our mla was successful, which is in line with the results of previous studies (boels, bakker, et al., 2022; boels, garcia moreno-esteva, et al., accepted). in the next section, we first describe how the mla we used (random forest) works for those not familiar with mlas and wanting to roughly understand these. next, we describe how we applied the mla in a failing approach using total dwell time (3.4.2), and in a successful approach using saccade direction and magnitude (3.4.3). 3.4.1 gaze-data analysis through a machine learning algorithm (random forest) a machine learning algorithm (mla) learns from input data without explicitly being programmed to use certain characteristics of the data. supervised learning algorithms (see figure 9) are a subset of machine learning algorithms whose training data contain known output values—in our case whether a student’s gaze data belonged to a before or after item. supervised mlas are broadly used for pattern recognition and for making predictions (friedman et al., 2001). specifically, our work focuses on the use of random forests to identify whether student gaze patterns change substantially between similar items across our sequence of items. boels et al. | f l r 12 figure 9. training and identification cycles of a supervised mla. a random forest is a combination of many small decision trees (also known as classification trees). these trees are constructed in unique ways (friedman et al., 2001). the basic structure of any single decision tree in a random forest begins with a central node and two branches (rokach & maimon, 2008, see figure 10). then, a tree-building algorithm (we use cart, breiman et al., 1984/2017) iterates through all variables in the training set to identify one that can split the data as homogeneously as possible. in our case, this involves identifying a variable—typically a saccade magnitude (length)—whose presence indicates belonging to one class, and whose general absence indicates belonging to another class. in short, variables that differentiate between output classes are selected for use in the tree-building process, and variables that do not help differentiate between output classes are not selected, and the cart algorithm determines how to best employ the useful variables. this variable selection process repeats until a stopping criterion—typically a maximum tree depth (in figure 10, this depth is one) or a minimum node size (number of squares in figure 10)—is reached. to learn more about decision trees, see, for example, breiman et al. (1984/2017). figure 10 shows an example of a possible split in a tree that uses the direction and magnitude of a saccade. figure 10. example of a decision tree. ese means east-south-east direction of the saccade. the uniqueness of decision trees in a random forest is that each tree is only given access to a random subset of the data sampled with replacement and each split in the tree is only given access to a random subset of the variables. that is, each tree only uses a random sample of participants’ data—with the possibility of sampling the same participant’s data multiple times—and each split in the tree uses a small subset of the total number of variables. the exact size of each sample is part of the tuning process and our final values can be seen in the supplementary r-code (dataset5). allowing each tree to be built on only a subset of data and variables typically leads to a worse-performing tree than if all data were available (dzeroski & zenko, 2004). however, the risk of building one (the best) tree only, is that this tree might work perfectly for exactly the given data set, but not on data sets that are similar but slightly different. this is called overfitting. building many trees using independently sampled data and variables stops any individual tree from drastically overfitting the training data and leads to trees that are relatively uncorrelated (hansen & salamon, 1990). these uncorrelated trees are then used together in an ensemble to make classifications, known as the random forest. each tree predicts the boels et al. | f l r 13 class of the given data—in our case whether the user is seeing the item for the first or second time—and then the votes are totaled. whichever class receives the most votes is the resulting classification of the random forest (breiman, 2001). this technique of simultaneously combining multiple machine learning algorithms—the trees in a random forest—together is known as ensemble learning. this approach is effective since the combined knowledge of many algorithms is often more accurate than any single algorithm (dzeroski & zenko, 2004). here, we train our ensemble using gaze data as inputs and a binary output indicating whether the user is seeing the item for the first time or the second time. we use the randomforest package in r (liaw & wiener, 2002) to implement our random forest. our final, fully-tuned model utilized the following hyperparameters: 1,000 trees, 5 variables considered at each split, a minimum node size of 1, and a maximum tree depth of 5. we identified these optimal hyperparameters using a grid search of size 3^5 (we tried all combinations of three different values for each hyperparameter). our nested resampling scheme utilized both an outer resampling and an inner resampling of 5-fold cross-validation. the reported best hyperparameters are the average values used across each of our outer resamplings. we likewise evaluated our model using 5-fold cross-validation. to ensure no students’ data were part of both the training and testing data when evaluating our model, we split the data into five groups, each group containing 10 students. we then used the 10 students’ data (yielding a total of 20 student-item pairings) as testing data, and trained our random forest on the remaining 40 students’ data (80 student-item pairings). this process was repeated five times until all students’ data have been separately used as training and testing data. the software used for this data analysis is rstudio (rrid:scr_000432; scicrunch registry). the full reproducible code to build our random forest as used in rstudio as well as the processed data are available through a data repository (dataset 5). the original data can be found in dataset 1. 3.4.2 a failed mla—using dwell time on aois as described at the beginning of section 3.4, we first plotted the data before we analyzed it with an mla. this plotting is considered part of the data analysis, as the plots can provide indications of what features might be relevant as inputs for the mla. differences in where and how long participants looked—fixated—were explored over time throughout each of the four items of interest. as an example, figure 11 shows the fixations for two selected, archetypical participants, l01 and l05, who progressed very differently through the same item (here, the double-histogram item11). the x-axis, time, has been rescaled from 0 to 1 so fixations could be compared between participants who spent different amounts of time on each item. a time of 0.5, for example, indicates the time at which the given participant is halfway through completing item11. in this figure, points are jittered (shifted a slight amount in a random direction) to better display the density of points in a given aoi at a given time. participant l01, like many participants, fixated on several different aois throughout their time working on item11, often moving back and forth between the graphing area and the corresponding axis. participant l05, however, spent most of their time fixating on the graph area of both the left and right graphs—stopping briefly to look at the right graph’s vertical axis and label after having spent a considerable amount of time looking at the graphing area. in addition to these two main archetypes, the remaining gaze patterns varied widely between each of the four items and between participants on a given item. to quantify the differences between student approaches to the before and after items, we began by identifying features (variables) for training our series of random forest models. if the random forest algorithm can consistently differentiate between gaze data from the before and after item in each pairing, then some combination of features must exist that is more prevalent in one item when compared to the other, indicating a difference in gaze patterns between the paired items. for each of the two item pairings (pairing item02 and item20, pairing item11 and item21), we began by treating each participant-item combination as our unit of observation, yielding a total of 200 data points (50 participants’ gaze patterns across two items in each of two pairings). for each data point, we calculated the proportional time spent in each aoi. we identified the path each participant took through the aois and converted this information into features for our random forest model. boels et al. | f l r 14 figure 11. distribution of gazes of participants l01(left) and l05 (right) over aois for item11. note. the horizontal axis shows time, which is rescaled from 0 to 1 for each participant individually. in short, this initial approach was unsuccessful. to prevent readers from scrolling back and forth, we provide a short description of these results here. there were no discernable differences at the individual level between each pairing of before and after items. participants spent roughly the same proportion of time looking at each of the aois when they saw the before items as when they saw the corresponding after item. though the order in which participants progressed through each of the aois differed between before and after items, no discernable pattern emerged, and the correspondingly trained random forest algorithms were unable to accurately predict whether a participant was viewing a before item or an after item in a given pairing. we, therefore, do not further elaborate on this approach in the results section. 3.4.3 a successful approach—exploring saccade direction and magnitude based on previous qualitative work (boels, bakker, et al., 2022) we then used directional movements— saccades—first by, again, visually investigating whether differences appeared in saccade patterns between before items and after items. we noticed a clear difference in the pattern of saccades due to the mirrored orientation of the otherwise-identical graphs in item11 and item21. thus, our subsequent analysis focused on mirrored versions of the after items, item21 and item20, so that the graph area is made identical to their before counterparts. in other words, we took the gaze coordinates for the mirrored after items and adjusted them to match the corresponding coordinate of the unmirrored before items. without this un-mirroring, the random forest algorithm could have easily differentiated between gaze data from the before and after items in each pair. figure 12 shows the patterns of saccades for the same two selected archetypical participants—l01 and l05—on one particular pairing, item11 and mirrored-item21. in this figure, all saccades are centered at the origin and radiate outward based on the direction and magnitude of the saccade. only saccades of magnitudes greater than 200 pixels are shown since these are the saccades used in our final model. saccades of less than 200 pixels were generally eye movements that are not indicative of students moving from one fixation point to another. given the size of the graph areas (width 489 pixels, height 313 pixels for double graph items, 600 x 335 for single graph items), the maximum possible saccade magnitude is 687 pixels (diagonal). the maximum speed boels et al. | f l r 15 for human saccades is approximately 700 degrees per second (fuchs, 1967; oohira et al., 1981) which is 570 pixels per 16.7 ms (a sampling rate of 60 hz equals to one sample every 16.7 ms). therefore, the maximum for long saccades was set at 600 pixels; longer saccades were considered to be artifacts. furthermore, we removed ‘saccades’ smaller than 50 pixels. given the accuracy of the eye tracker (mean 13.4 pixels in the center of the screen and mean 56.6 pixels over all measures), we consider these small ‘saccades’ part of fixations or noise. although this meant removing more than ninety percent of the measurements on the graph area, the accuracy of our mla became slightly better. we defined the beginning of a saccade to be a movement with a velocity greater than 50 pixels per 16.7ms. we defined the end of the saccade as any two consecutive 16.7ms windows where the participant’s gaze had not moved more than 50 pixels. points of fixation were determined by averaging the xand y-pixel values of gazes in between two saccades, and each saccade’s direction and magnitude were calculated between these points of fixation. figure 12. saccades of magnitude 200 pixels or more of participants l01 and l05. note. students’ saccades on the graph area only after mirroring. the number of medium and long saccades decreased from item11 to item21 for these two participants. all saccades are shifted to the origin. figure 13 shows all saccades of a magnitude of 200 pixels or more for each of the four items. there is a discernable difference in the number of vertically oriented saccades between before and after items, especially in the item02 and item20 pairing. the item11 and item21 pairing also shows differences in the orientations of many horizontally facing saccades. notably, there are several more northwestand southeast-facing saccades in item11 and more northeast and southwest-facing saccades in item21 (after mirroring). we examine differences in students’ gaze patterns on histogram items. if the random forest algorithm can consistently differentiate between gaze data from the before and after item in each pairing, then there must be some combination of features (variables) that is more prevalent in one item when compared to the other, indicating a difference in gaze patterns between the paired items. to construct our random forest model, we tried different sets of features—placing each saccade into mutually exclusive bins depending on the direction, magnitude, and phase of the saccade, regardless of the point of origin. we tested two different directional schemes, two magnitude schemes, and three phase adjustment schemes, yielding a total of twelve boels et al. | f l r 16 combinations. here, a phase adjustment is an angle (in radians) at which the direction bins are shifted, where 0 radians is equivalent to 0 degrees in mathematics—a saccade pointed eastward—and pi radians is equivalent to 180 degrees—a saccade pointed westward. table 2 shows the details of each scheme. the direction of each saccade (with no phase shift) was linked to a specific compass rose direction (figure 14). the direction ese, for example, aligns with the direction [7π/4, 2π). figure 13. saccades of magnitude 200 pixels or more of all participants on item11 and item21 (doublehistograms, top) as well as item02 and item20 (single-histogram, bottom). note. notice the difference in the density of students’ saccade directions between the before items (left) and after items (right). we calculated accuracy, sensitivity, and specificity for each combination, seeding our random forest algorithm to obtain reproducible results. we treated the before items as the positive case, meaning that accuracy is defined as the total proportion of times the algorithm correctly identified whether the given data point belonged to the before item or the after item in the pairing. sensitivity is the proportion of the data points the model identified as belonging to the before item that actually belonged to the before item. specificity is the proportion of after-item data points that were correctly identified as belonging to the after item. each metric was calculated using 5-fold cross-validation. we categorized the features into bins (see table 2) for two reasons. first, we wanted to extract the variable importance metrics from our random forest in a way that might better inform us why the model was differentiating so well. more specifically, we wanted to know which direction or magnitude of a saccade is more present in one item’s data and not the other’s. if we would have used continuous features, this interpretation would have been much more convoluted to human beings. second, we also tried a continuous features model. that model performed worse. this might be because saccades that are close in direction and boels et al. | f l r 17 magnitude are functionally identical. in other words, perhaps a saccade of 25 degrees and a saccade of 5 degrees both imply that a student is scanning from left to right, and the difference in angles is either an artifact of data error or a meaningless difference between fixation points. table 2 bins that are used to categorize each saccade for use in our mla feature number categories category name direction 1 [0, π/2); [π/2, π); [π, 3π/2); [3π/2, 2π) ne, nw, sw, se (without phase adjustment; see figure 14) direction 2 [0, π/4); [π/4, π/2); [π/2, 3π/4); [3π/4, π); [π, 5π/4); [5π/4, 3π/2); [3π/2, 7π/4); [7π/4, 2π) ene, nne, nnw, wnw, wsw, ssw, sse, ese (without phase adjustment) magnitude 1 [50, 300); [300, 600) medium, long magnitude 2 [50, 100); [100, 200); [200, 400); [400, 600) very short, short, medium, long phase adjustment 1 0 phase adjustment 2 π/8 phase adjustment 3 π/3 note. for each feature (variable) the specific categories are indicated by a number. in total, there are two times two times three equals twelve possible combinations. 4. results our overall research question is: in what way do secondary school students’ histogram interpretations change after solving dotplot items? in this section, we answer this question by answering the following three subquestions: 1) what are the main differences in students’ gaze patterns on histogram items before and after solving dotplot items? 2) what indications can be found in students’ verbalizations during stimulated recall that changes in their approaches to histograms occurred? 3) what are the differences in students’ answers on histogram items before and after solving dotplot items? 4.1 main changes in students’ gaze patterns on histograms the first sub-question is answered by using a random forest model. the twelve combinations of direction, magnitude, and phase schemes (table 2) yielded accuracies, sensitivities, and specificities that varied between 55% and 88% (table 3). the standard deviations for each performance metric are reported in parenthesis using 100 resamples. the most accurate combination for both pairings was direction 2, magnitude 2, and phase adjustment 1, which corresponded to the most granular direction and magnitude bins and no phase adjustment. the details of each combination can be seen in table 2. this best combination yielded a remarkably high 77% accuracy for the single-histogram items (table 3) and 86% accuracy for the double-histogram items (table 4). we note that accuracy alone can be potentially misleading—in our study attributing scanpath patterns randomly would yield 50% accuracy. therefore, an accuracy near 75% is considered good in this low-stake situation, since it would show gains in accuracy well above random guessing. in addition, accuracy needs to be judged together with sensitivity and specificity for which values near 75% are also considered good. note that this performance may increase if we were to exclude some students for whom we had sparse data for a specific item (e.g., due to data loss for that specific item). to better understand which features were driving the accuracy of these models, we calculated the importance of each variable (figure 15) for the best models for each pairing. these plots show the estimated average decrease in accuracy if the given feature was removed from the data set completely. for example, if the number of ese short saccades was unknown to the random forest, the model for the single-histogram item would see boels et al. | f l r 18 a 15% decrease in absolute accuracy. these metrics are calculated using out-of-bag error estimates (a bootstrapping method for measuring the prediction error for the random forest where a variable is left out of the resampling ‘bag’). the ese, wnw, and wsw saccades were most important to both models’ accuracies, all being almost horizontal saccades. almost vertical saccades (e.g., nnw) were less important. table 3 model efficacy on item02 versus item20 (single-histogram items) direction magnitude phase accuracy (sd) sensitivity (sd) specificity (sd) 1 1 1 0.62 (0.024) 0.60 (0.021) 0.64 (0.027) 1 1 2 0.55 (0.025) 0.56 (0.024) 0.54 (0.028) 1 1 3 0.66 (0.018) 0.60 (0.029) 0.72 (0.021) 1 2 1 0.68 (0.014) 0.66 (0.020) 0.70 (0.018) 1 2 2 0.64 (0.016) 0.60 (0.020) 0.68 (0.026) 1 2 3 0.61 (0.011) 0.62 (0.013) 0.60 (0.017) 2 1 1 0.73 (0.008) 0.72 (0.009) 0.74 (0.014) 2 1 2 0.66 (0.016) 0.62 (0.022) 0.70 (0.024) 2 1 3 0.69 (0.008) 0.66 (0.010) 0.72 (0.014) 2 2 1 0.77 (0.015) 0.70 (0.018) 0.84 (0.021) 2 2 2 0.70 (0.016) 0.68 (0.023) 0.72 (0.024) 2 2 3 0.72 (0.014) 0.66 (0.019) 0.78 (0.022) note. sd = standard deviation table 4 model efficacy on item11 versus item21 (double-histogram items) direction magnitude phase accuracy (sd) sensitivity (sd) specificity (sd) 1 1 1 0.73 (0.006) 0.82 (0.009) 0.64 (0.010) 1 1 2 0.68 (0.010) 0.76 (0.015) 0.60 (0.012) 1 1 3 0.75 (0.011) 0.76 (0.019) 0.74 (0.010) 1 2 1 0.73 (0.012) 0.84 (0.017) 0.62 (0.025) 1 2 2 0.69 (0.012) 0.72 (0.009) 0.66 (0.018) 1 2 3 0.73 (0.011) 0.76 (0.016) 0.70 (0.017) 2 1 1 0.85 (0.007) 0.90 (0.014) 0.80 (0.008) 2 1 2 0.80 (0.009) 0.80 (0.004) 0.80 (0.018) 2 1 3 0.81 (0.007) 0.80 (0.010) 0.82 (0.009) 2 2 1 0.86 (0.013) 0.88 (0.009) 0.84 (0.023) 2 2 2 0.77 (0.011) 0.76 (0.017) 0.78 (0.016) 2 2 3 0.82 (0.007) 0.80 (0.013) 0.84 (0.006) note. sd = standard deviation figure 14. compass rose. boels et al. | f l r 19 phase 1—no phase adjustment—yielded the best model on average. direction and magnitude had moderate effects, with more refined bins yielding better-performing models. both pairings showed a wider variation in accuracies, with low accuracy for more broad categorization schemes and remarkably high accuracy for more refined categorization schemes. figure 15. plots showing the importance of variables for single-histogram mla-model (left) and doublehistogram mla-model (right). 4.2 students’ post-activity verbal descriptions of approaches to histogram items the machine learning analysis indicated that there are differences in the gazes of students. to investigate whether these changes reflect a change in students’ approaches, we analyzed students’ stimulated recall verbal reports (see dataset1 for the items). through qualitative coding of the excerpts of these verbal reports, we answer the second sub-question: what indications can be found in the verbalizations during stimulated recall that changes in students’ approaches to histograms occurred? some students learned from the sequence of items, as we can see for example from the transcript of studentl20 for single-histogram item02: l20: i think i did 18 + 8 + 5 ∙ 2 + 2 ∙ 4. and then divided all that by nine. […] r1: and then you first came up with nine kilos and later you changed it to six kilos. […] l20: it is not right. r1: and why is it not right? l20: it must be somewhere near the three. r1: why? l20: because there’s a lot less than one kilogram, and relatively a lot two kilograms. and then after that, it really expands to nine kilograms but those are all very small numbers. so, then you end up with three. r1: yes. so now that you look at it again you think: i should have given a completely different answer? l20: yes. what this transcript also shows is that this understanding of how to estimate the mean from a histogram took place sometime after single-histogram item02, but it is not clear when exactly this understanding occurred. for some students, it occurred after (at least one item of) the second series of histogram items, as the following boels et al. | f l r 20 student excerpt shows. this student reflects on the chosen approach in (left-skewed, single histogram) item19 in the stimulated recall: l01: the mean will be about between five and nine because there are a lot of values [measured weights] there. and then around seven because that’s a little bit more to the left to zero from the middle between five and nine. r1: would you like to look at your eye movements again? [l01 looks back eye movements] r1: […] you gave the answer ten. and that’s where you looked. l01: ten? [sounds astonished] r1: yes, look at your eye movements again. [l01 looks back eye movements] l01: oh yes, in that case, i misread the axes. for twenty-six students, there was (almost) no room for improvement because they already gave answers within or close to the answer range during the before sequence of single-histogram items. another four to twelve students seem to have learned specifically from the dotplot items. for example, studentl16 answered seven (instead of 2.7) for single-histogram item02 but starts giving answers within or very close to the answer range for all following single dotplot items, and continues with these correct answers for the second series of single-histogram items after the dotplot items. during the recall, this student first describes a correct strategy for finding the mean from item02 which is not in line with the given answer: l16: yes, again looking at frequency and weight, and then we see that the one occurred very often and the further you get [to the right] actually the less [frequency]. so, then the mean goes much more to the one than to the high numbers. r1: yes, you then said about seven. l16: yes, then i looked at it wrong again. then i got weight and frequency flipped again. 4.3 differences in students’ answers to histogram items at first glance, there seems to be no real difference in answer correctness between item02 and item20 (table 5). nevertheless, there are two indications in students’ answers that students learned between item02 and item20. first, the answer range chosen for correct answers impacts answer correctness and was set the same for all items. in this study, students seem to prefer whole and half numbers. enlarging the answer range to include the next whole or half numbers would result in (a non-significant) improvement in answer correctness (see note table 5). answer correctness is, therefore, quite sensitive to researchers’ choices. hence, changes in students’ answers are a better indicator of students’ learning potential. second, differences between students’ answers and the actual mean are much lower for item20 compared to item02. we calculated the difference between the actual mean and the estimated mean (= mdiff). mdiff is, as expected, lower for the after item. we explored if this difference was significant through a one-tailed pairedt-test, as we expected that dotplot items would support students in correctly estimating the mean from histograms in the after items. the assumptions for a paired t-test, such as the unimodality and rough symmetry of the paired differences, were checked and met. the results for before item02 (mdiff = 1.1, sd = 1.8) compared to after item20 (mdiff = 0.4, sd = 1.7) indicate that it is possible that dotplots improve students’ performance on the after item, t(49)= -1.7, p = 0.0469 < 0.05. we consider this (and the next) p-value significant in the way the statistician fisher intended: “in the old-fashioned sense: worthy of a second look” (nuzzo, 2014, pp. 150– 151). the 95%-confidence interval for the differences in mdiff is <-inf, -0.014] and cohen’s d measure for effect size is 0.40. altogether, this points toward an improvement in answers. note that, on the one hand, effect sizes tend to be larger in researcher-made tests compared to general (standardized) tests as well as in studies with small sample sizes. on the other hand, (very) short interventions often have lower effect sizes (e.g., bakker et al., 2019). furthermore, although students’ answers’ correctness improved for the double-histogram item after the dotplot items, this improvement is not significant, as p = 0.3428 > 0.05 (e.g., mcnemar, 1947). boels et al. | f l r 21 table 5 answers that are given by the students, n = 50 it em a ct u al m ea n a n sw er r an g e co rr ec t an sw er s* a v er ag e o f g iv en an sw er s d if fe re n ce b et w ee n st u d en ts ’ an sw er s – ac tu al m ea n n u m b er o f st u d en ts co rr ec t p er ce n ta g e o f st u d en ts c o rr ec t it em answer options (number of students with this answer) p er ce n ta g e o f st u d en ts c o rr ec t item02 2.7 1.6–3.8** 3.8 1.1 19 38% item11 ellen (2) titia (30) same (18) 36% item20 6.3 5.2–7.4*** 6.7 0.4 17 34% item21 elizabeth (26) monsif (1) same (23) 45% note. *experts were also asked to answer these items. based on these results as well as students’ preference for whole numbers, the answer range was set to +/-1.1 for all items. **if answers 1.5 and 4 had been included, 27 students (54%) would have answered correctly. ***if answers 5 and 7.5 had been included, 31 students (62%) would have answered correctly. correct answers are in bold. instead of attributing the smaller mdiff —for single histogram items–—to the solving of dotplot items, one alternative explanation is that the mean of item20 (m = 6.3) compared to item02 (m = 2.7) is closer to the mean of the frequencies (m = 4.9 for both). nevertheless, we do not expect that this is the case, as for another left skewed item of this sequence of items (item06, not further reported here, see boels, garcia moreno-esteva, et al., accepted) the mean of the frequencies (m = 7.1) was also close to the mean of this item (m = 5.7), but the difference between actual and students’ mean (mdiff = 1.2) was similar to before item02. to further exclude this alternative explanation, we suggest enlarging the difference between the actual mean and the mean of frequencies by adding more data (e.g., 50 packages) to the graphs. this number of added packages should not be too high to avoid students guessing from the size of the numbers what the weights are, as may have played a role in an item with sat scores according to kaplan et al. (2014). 5. conclusions and discussion in this study, we answer the main research question of in what way secondary school students’ histogram interpretations change after solving dotplot items. more specifically, we look at students’ estimations and comparisons of means from histograms. we expected that solving dotplot items would focus students’ attention on the measured variable (weight) being depicted along the horizontal axis. in turn, that would invite students to estimate the mean of the weights (along the horizontal axis) instead of the mean of the frequencies (along the vertical axis) in the histograms. we examined three indications that taken together can suggest detailed-level changes in students’ histograms interpretations: a change in students’ gaze patterns, a shift in students’ strategy for solving the histogram items, and an improvement in students’ answers. if the changes are for the better, the relevance of knowing them is that they could underpin the learning potential of using dotplot items before solving histogram items—a hypothesis put forward by researchers in statistics education. for the first indicator—a change in students’ eye movements—we looked at differences in students’ gaze or scanpath patterns on the graph area through a machine learning algorithm. two main differences—on the student level—between scanpath patterns on before and after items were found. first, there were proportionally fewer horizontal directions (ese/wnw) in the gaze patterns on the after items than on the before items. second, proportionally more vertical directions (nnw/nne) were found in the after items. a horizontal gaze pattern is associated with an incorrect strategy while a vertical gaze pattern is associated with a correct strategy. our best implementations of random forests were able to accurately classify (roughly 80% of the instances) whether it was the first (before) or second (after) time a participant had seen an item in one of two pairings. what we can attest to, is a significant, discernable difference in the way participants looked at the before items and when viewing the mirrored after versions, even after accounting for mirroring in the graphs. as we could not identify other confounding factors, it seems reasonable to conclude that our findings exhibit evidence that students changed the way they approached an item when seeing its mirrored version later in this sequence of boels et al. | f l r 22 items. the results of our mla are in line with the results of previous studies (boels, bakker, et al., 2022; boels, garcia moreno-esteva, et al., accepted). we cannot be certain whether the observed differences in gazes indicate a change in strategies. although scanpaths can disclose students’ strategies at a detailed level, the relationship between eye movements and strategies is task-dependent (e.g., orquin & holmqvist, 2017; russo, 2010). in addition, not every eye movement is part of a task-specific strategy (e.g., schindler & lilienthal, 2019). therefore, other data—such as our second and third indicators—are often needed to support or refute conjectures about the association between scanpath patterns and strategies. the second indicator of changes in students’ histograms interpretations—a shift in students’ strategy for solving the items—was evaluated by coding students’ stimulated recall verbal reports. the excerpts provide evidence that at least some students changed their strategies, from an incorrect approach for estimating and comparing means from histograms to a correct approach, during or after solving the dotplot items. the third indicator—improvement in students’ answers—was explored through both answer correctness, the difference between students’ estimation of the mean and the actual mean for single-histogram items, and the changes in students’ answers on the double-histogram multiple-choice items. answer correctness did not change significantly on either item type. nevertheless, the difference in students’ estimation of the mean compared to the actual mean was significantly smaller for the after item compared to the before item. we use ‘significantly’ here in the sense fisher intended: worthy of further investigation. data collection with new and more participants—from the same population (dutch grades 10–12 pre-university track students)—is needed to investigate the hypothesis that this difference becomes smaller, that there is a change in multiple-choice answers, and that both are due to solving the dotplot items. the three indicators taken together suggest that at least some students changed their strategy during or after solving the sequence of dotplot items. a change in gaze behavior was observable through our machine learning analysis with random forests. depending on how learning is defined, this change could point to a learning effect of solving dotplot items. interpreting the results, we abductively arrived at the following explanations for our results. first, the change toward proportionally more vertical gazes on the after items is in line with the conjecture that the absence of a vertical scale in dotplots can turn students’ attention toward the horizontal scale which is where the variable is presented in both histograms and dotplots. these students possibly figured out that the mean can be estimated from the measured values along the horizontal axis. however, we cannot rule out that factors other than solving dotplot items could have contributed to this change. second, we consider the most likely explanation for the mixed results that solving dotplot items promoted readiness for learning (church & goldin-meadow, 1986) about histograms. having students reflect on their previous strategy while they were cued with their own gazes during retrospective verbal reporting then might have given them new insights. after solving the dotplot items, histogram items seem to lie within the region of sensitivity for learning, hence within students’ zone of proximal development (vygotsky, 1978). it is possible that the questions asked by the researcher (an adult), which were intended to figure out how students solved the items, unintentionally stimulated students’ thinking by asking them to explain—hence, reflect on— their strategies. further research is needed to check this explanation. an alternative explanation for the results would be that other items after the second series of histogram items induced students’ thinking. although we cannot exclude this alternative, we regard this to be less likely. further discussing the results, we note that this study is novel in the following ways. first, to the best of our knowledge, our study is the first in education that combined a quantitative analysis of the scanpath patterns found in spatial gaze data with insights from a previous qualitative study about what part of the scanpath pattern is relevant for students’ strategies (namely, the scanpath on the graph area only). the use of qualitative insights contributes to the validity of the study while the quantitative approach through machine learning analysis contributes to the reliability of it. most eye-tracking studies that use spatial measures investigate the sequence of aois (garcia moreno-esteva et al., 2020) and the same holds for those combining it with mlas (e.g., garcia moreno-esteva et al., 2018). instead, we used vectors (i.e., direction and magnitude) of saccades. studies in education that utilize vectors are rare (e.g., dewhurst et al., 2018). second, novel is the use of an boels et al. | f l r 23 mla for finding differences in gazes that are relevant to changes in students’ task-specific strategies between tasks. our study has several limiting factors. first, many of the participants’ gaze data contained data loss. although data loss is normal due to blinking or looking away from the screen, some data loss could be avoided by preexcluding participants who wear glasses, contact lenses, or mascara. in addition, an eye tracker could be used that is better in catching gazes from people with epicanthic folds (almond eyes). as we aimed for a naturalistic setting, we did not exclude any of such participants. in addition, for some participants, we had sparse data. some of these participants spent only a few seconds looking at a given item. this made predictions and training more challenging. most of these participants appeared to parse the graph and answer the corresponding question(s) in a rapid but reasonable manner, although one participant appeared to scan the graph and answer the question in such a rapid way that it is unlikely that they had time to fully understand what the graph was depicting. since no participants’ data were removed, some amount of data cleaning and removal of outlier participants would likely increase the accuracy of our random forests, although our data collection scheme does not allow us to know with certainty why a certain participant’s gaze data were sparse for a particular item. a second limiting factor was that we restricted our final analysis to the graph area of each item, excluding aois such as the axes labels and the graph title. the inclusion of these aois yielded more noise and worse results, but further work might investigate the possibility of productively including them. third, answer correctness and students’ strategies correspond only to a limited extent. finally, and most importantly, our sample size—50 participants and 2 items yielding 100 participant-item-pairings—is relatively small for machine learning and statistical analysis. our results indicate strong evidence of a change in gaze patterns between the before and after items, but more data are needed to generalize these findings appropriately. a theoretical contribution of this study is that having students solve non-stacked (‘messy’) dotplot items can create readiness for learning histograms. a reflection phase seems to be needed to make use of the knowledge obtained. we speculate that this partly explains the results from the literature on dotplots (e.g., garfield & ben-zvi, 2008; lyford, 2017). another reason for these results, we believe, is that only non-stacked dotplots contribute to students’ understanding of where the measured value is. stacked dotplots already contain an information reduction step (the binning) that could lead to similar misinterpretations as for histograms (see boels, bakker, van dooren & drijvers, 2019). we, therefore, advise investigating whether stacked dotplots need to be avoided in secondary education. as a first methodological implication, our study shows how an mla in combination with eye-tracking data can be used to reveal phenomena that are of interest to researchers of education. future use could include interpreting graphical representations in biology, physics, economics, and geography. by choosing features (variables) that are relevant to the phenomena of interest (here: students’ strategies for solving a histogram item) and meaningful to the researchers, an mla can give insights into subtle, detailed-level differences in students’ strategies that are hard to detect through other research methods, such as time measures in eyetracking research or qualitative analysis of gaze data by researchers. a second methodological implication is that there seems to be a practice effect within a sequence of items for at least some students, in line with suggestions from lumsden (1976). a practice effect refers to improved performance after ‘practicing’ (i.e., repeatedly solving similar or the same items). this is also important for judging the validity of summative assessments. more research is needed to confirm this within-a-test practice effect. catron (1978) found an effect of item types on an iq test. for example, he showed that the development of a strategy in strategic items improved performance on retesting. the present study does not consider an effect of item type. further research is needed to find out whether and for what items the order and type of items influence the within-a-test practice effect. what are the possible implications of our findings? the ml approach is generalizable to other sequences of items or any instance when a user may wish to classify eye-tracking data into one of many discrete categories. further analysis is needed to correlate the number of saccades of specific directions and magnitudes with particular viewing strategies. in other words, does the presence of certain features (variables, such as horizontal or vertical saccades) indicate students taking a particular strategy, and if so, is this strategy more common when viewing a before item as opposed to an after item? moreover, we think our ml approach can also be boels et al. | f l r 24 used when researchers want to know whether solving x (a question about a graph or image) changes the way students solve y (a question about a different type of graph or image). for testing a future hypothesis that students’ estimations of the mean from histograms become closer to the actual mean after solving dotplot items, we suggest making a sequence of 24 items: eight histogram items (improved versions of the existing items from the original sequence with more data in them as well as one extra left skewed single histogram and one extra double histogram), eight dotplot items (all items containing the same data as the first eight histograms) and then again eight histogram items (all mirrored versions of the first eight histogram items). to check a future hypothesis that giving students stacked dotplots is a less effective way to scaffold them, a variant of this design could be made with stacked dotplots only, instead of non-stacked dotplots (with the stacks in between two values on the horizontal scale; all stacked dotplots contain the same data as the first eight histograms). to check a future hypothesis that the reflection phase is important, a variant with and without stimulated recall verbal reports could be conducted followed by another series of histogram items. in all these variants, machine learning analysis can support this hypothesis testing. for practitioners, insight into what students learn from doing a sequence of items is also relevant for homework and formative assessment, in particular, if no feedback is given—which is quite a common situation (e.g., when there is less student-teacher interaction). the observed differences in gaze patterns together with the other evidence in this study, suggest that a sequence of items can create readiness for learning, but a teacher may still be needed to ensure that students reach their full potential. key points spatial gaze data on one aoi were used as inputs for a machine learning algorithm (mla) the high accuracy, specificity, and sensitivity of the mla indicate changes in students’ strategies due to detailed-level learning potential students’ answers’ correctness did not improve significantly, but post-activity stimulated recall verbal reports suggest learning research on detailed-level learning during an assessment is still in its infancy acknowledgments this research is funded with a doctoral grant for teachers from the dutch research council (nwo), number 023.007.023 awarded to lonneke boels. any opinions, findings, or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the dutch research council. we thank the following people for their contributions to this study. nathalie kuijpers for checking the document on style and english, ciera lamb for proofreading a previous version for american english, anna shvarts for assisting during the last day of data collection, wim van dooren for his contribution to the design of the eyetracking study, rutmer ebbes for his contribution to the pilot eye-tracking study (boels et al., 2018), aline boels for the programming of the html-files containing the items, juri boels for transcribing almost all verbal reports, iljo boels for exporting gaze plots and heatmaps, willem den boer for writing the macros for processing the eye-tracking data. furthermore, lb thanks all the people organizing and contributing to the eye-tracking seminars of the uu, especially ellen kok, margot van wermeskerken, roy hessels, ignace hooge, and jos jaspers. lb also thanks the faculty of social and behavioral sciences for lending the laptop and tobii-xii-60 eye tracker for this research. conflict of interest the authors declare that there is no potential conflict of interest. data availability the datasets analyzed for this study can be found in the dataversenl repository. dataset1: https://doi.org/10.34894/wekaye dataset3: https://doi.org/10.34894/7kneoh dataset5: https://doi.org/10.34894/plcebc https://doi.org/10.34894/wekaye https://doi.org/10.34894/7kneoh https://doi.org/10.34894/plcebc boels et al. 25 | f l r authors’ contributions ab, pd, and lb contributed to the conception and design of the study. the data were collected and cleaned by lb. ml analysis and statistical tests were conducted by al. lb wrote the first draft of the introduction, materials and methods, and results section except for most of the texts on the ml analysis and results. the latter was initially written by al. results and discussion were initially written by lb and al. all authors contributed to manuscript revision, and finally read and approved the submitted version. references allmond, s. & makar, k. (2014). from hat plots to box plots in tinkerplots: supporting students to write conclusions which account for variability in data. in k. makar, b. de sousa, & r. gould (eds.), sustainability in statistics education. proceedings of the ninth international conference on teaching statistics. https://iase-web.org/icots/9/proceedings/pdfs/icots9_2e1_allmond.pdf?1405041584 bakker, a. (2004a). design research in statistics education: on symbolizing and computer tools [doctoral dissertation, utrecht university]. https://dspace.library.uu.nl/handle/1874/893 bakker, a. (2004b). reasoning about shape as a pattern in variability. statistics education research journal, 3(2), 64–83. http://iase-web.org/documents/serj/serj3(2)_bakker.pdf?1402525004 bakker, a., biehler, r., & konold, c. (2004). should young students learn about box plots? in g. burrill & m. camden (eds.), curricular development in statistics education: international association for statistical education 2004 roundtable, (pp. 163–173). international statistical institute. https://iaseweb.org/documents/papers/rt2004/4.2_bakker_etal.pdf?1402524988 bakker, a., cai, j., english, l. kaiser, g., mesa, v., & van dooren, w. (2019). beyond small, medium, or large: points of consideration when interpreting effect sizes. educational studies in mathematics, 102, 1–8. https://doi.org/10.1007/s10649-019-09908-4 bargagliotti, a., franklin, c., arnold, p., gould, r., johnson, s., perez, l., & spangler, d. (2020). pre-k-12 guidelines for assessment and instruction in statistics education (gaise) report ii. american statistical association and national council of teachers of mathematics. https://www.amstat.org/asa/files/pdfs/gaise/gaiseiiprek-12_full.pdf bennett, r. e. (2011) formative assessment: a critical review. assessment in education: principles, policy & practice, 18(1), 5–25, https://doi.org/10.1080/0969594x.2010.513678 ben-zvi, d. & garfield, j. (2004). research on reasoning about variability: a forward. statistics education research journal, 3(2), 4–6. http://iase-web.org/documents/serj/serj3(2)_forward.pdf?1402525004 boels, l., bakker, a., van dooren, w., & drijvers, p. (2019). conceptual difficulties when interpreting histograms: a review. educational research review, 28, article 100291. https://doi.org/10.1016/j.edurev.2019.100291 boels, l., bakker, a., & drijvers, p. (2019). eye tracking secondary school students’ strategies when interpreting statistical graphs. in m. graven, h. venkat, a.a. essien, & p. vale (eds.), proceedings of the forty-third conference of the international group for the psychology of mathematics education, 2, (pp. 113–120). https://www.igpme.org/publications/current-proceedings/ boels, l., bakker, a., van dooren, w., & drijvers, p. (2022). secondary school students’ strategies when interpreting histograms and case-value plots: an eye-tracking study. [manuscript submitted for publication] freudenthal institute, utrecht university. boels, l., ebbes, r. bakker, a., van dooren, w., & drijvers, p. (2018). revealing conceptual difficulties when interpreting histograms: an eye-tracking study. invited paper, refereed. in m. a. sorto, a. white, & l. guyot (eds.), looking back, looking forward. proceedings of the tenth international conference on teaching statistics, (pp. 1–4). https://iase-web.org/icots/10/proceedings/pdfs/icots10_8e2.pdf https://iase-web.org/icots/9/proceedings/pdfs/icots9_2e1_allmond.pdf?1405041584 https://dspace.library.uu.nl/handle/1874/893 http://iase-web.org/documents/serj/serj3(2)_bakker.pdf?1402525004 https://doi.org/10.1007/s10649-019-09908-4 https://www.amstat.org/asa/files/pdfs/gaise/gaiseiiprek-12_full.pdf https://doi.org/10.1080/0969594x.2010.513678 http://iase-web.org/documents/serj/serj3(2)_forward.pdf?1402525004 https://doi.org/10.1016/j.edurev.2019.100291 https://iase-web.org/icots/10/proceedings/pdfs/icots10_8e2.pdf boels et al. 26 | f l r boels, l., garcia moreno-esteva, e., bakker, a., & drijvers, p. (accepted). automated gaze-based identification of students’ strategies in histogram tasks through an interpretable model and a machine learning algorithm. international journal of artificial intelligence in education. breiman, l., friedman, j. h., olshen, r. a., & stone, c. j. (2017). classification and regression trees. routledge. (original work published 1984) https://doi.org/10.1201/9781315139470 breiman, l. (2001). random forests. machine learning, 45(1), 5–32. https://doi.org/10.1023/a:1010933404324 capraro, m. m., kulm, g., & capraro, r. m. (2005). middle grades: misconceptions in statistical thinking. school science & mathematics, 105(4), 165–174. https://doi.org/10.1111/j.1949-8594.2005.tb18156.x catron, d. w. (1978). immediate test-retest changes in wais scores among college males. psychological reports, 43(1), 279–290. https://doi.org/10.2466/pr0.1978.43.1.279 church, r.b., & goldin-meadow, s. (1986). the mismatch between gesture and speech as an index of transitional knowledge. cognition, 23(1), 43–71. https://doi.org/10.1016/0010-0277(86)90053-3 clayden, a., & croft, m. (1990). statistical consultation—who’s the expert? annals of mathematics and /artificial intelligence, 2, 65–75. https://doi.org/10.1007/bf01530997 cohen, s. (1996). identifying impediments to learning probability and statistics from an assessment of instructional software. journal of educational and behavioral statistics, 21(1), 35–54. https://doi.org/10.2307/1165254 cooper, l. l. (2018). assessing students’ understanding of variability in graphical representations that share the common attribute of bars. journal of statistics education, 26(2), 110–124. https://doi.org/10.1080/10691898.2018.1473060 cooper, l. l., & shore, f. s. (2008). students’ misconceptions in interpreting center and variability of data represented via histograms and stem-and-leaf plots. journal of statistics education, 16(2). https://doi.org/10.1080/10691898.2008.11889559 dabos, m. (2014). a glimpse of two year college instructors’ understanding of variation in histograms. in k. makar, b. de sousa, and r. gould (eds.). sustainability in statistics education. proceedings of the ninth international conference on teaching statistics (pp. 1–4). https://icots.info/9/proceedings/pdfs/icots9_c150_dabos.pdf delmas, r., garfield, j., & ooms, a. (2005). using assessment items to study students’ difficulty reading and interpreting graphical representations of distributions. proceedings of the fourth international research forum on statistical reasoning, literacy, and reasoning. university of auckland. https://www.causeweb.org/cause/archive/artist/articles/srtl4_artist.pdf delmas, r., & liu, y. (2005). exploring students’ conceptions of the standard deviation. statistics education research journal, 4(1), 55–82. http://iaseweb.org/documents/serj/serj4(1)_delmas_liu.pdf?1402525005 dewhurst, r., foulsham, t., jarodzka, h., johansson, r., holmqvist, k., & nyström, m. (2018). how task demands influence scanpath similarity in a sequential number-search task. vision research, 149, 9–23. https://doi.org/10.1016/j.visres.2018.05.006 dzeroski, s., & zenko, b. (2004). is combining classifiers with stacking better than selecting the best one? machine learning, 54(3), 255–273. https://doi.org/10.1023/b:mach.0000015881.36452.6e falleti, m. g., maruff, p., collie, a. & darby, d. g. (2006). practice effects associated with the repeated assessment of cognitive function using the cogstate battery at 10-minute, one week and one month testretest intervals, journal of clinical and experimental neuropsychology, 28(7), 1095–1112. https://doi.org/10.1080/13803390500205718 https://doi.org/10.1201/9781315139470 https://doi.org/10.1023/a:1010933404324 https://doi.org/10.1111/j.1949-8594.2005.tb18156.x https://doi.org/10.2466/pr0.1978.43.1.279 https://doi.org/10.1016/0010-0277(86)90053-3 https://doi.org/10.1007/bf01530997 https://doi.org/10.2307/1165254 https://doi.org/10.1080/10691898.2018.1473060 https://doi.org/10.1080/10691898.2008.11889559 https://icots.info/9/proceedings/pdfs/icots9_c150_dabos.pdf https://www.causeweb.org/cause/archive/artist/articles/srtl4_artist.pdf https://doi.org/10.1016/j.visres.2018.05.006 https://doi.org/10.1023/b:mach.0000015881.36452.6e https://doi.org/10.1080/13803390500205718 boels et al. 27 | f l r friedman, j., hastie, t., & tibshirani, r. (2001). the elements of statistical learning. springer series in statistics. https://link.springer.com/book/10.1007/978-0-387-21606-5 fuchs, a. f. (1967). saccadic and smooth pursuit eye movements in the monkey. the journal of physiology, 191(3), 609–631. https://doi.org/10.1113/jphysiol.1967.sp008271 gal, i. (1995). statistical tools and statistical literacy: the case of the average. teaching statistics, 17(3), 97– 99. https://doi.org/10.1111/j.1467-9639.1995.tb00720.x gal, i. (2002). adults’ statistical literacy: meanings, components, responsibilities. international statistical review, 70(1), 1–25. https://iase-web.org/documents/intstatreview/02.gal.pdf gal, i & geiger, v. (2022). welcome to the era of vague news: a study of the demands of statistical and mathematical products in the covid‑19 pandemic media. educational studies in mathematics, 111, 5– 28. https://doi.org/10.1007/s10649-022-10151-7 garcia moreno-esteva, e., kervinen, a., hannula, m. s., & uitto, a. (2020). scanning signatures: a graph theoretical model to represent visual scanning processes and a proof of concept study in biology education. education sciences, 10(5), article 141. https://doi.org/10.3390/educsci10050141 garcia moreno-esteva, e., white, s. l. j., wood, j. m., & black, a. a. (2018). application of mathematical and machine learning techniques to analyse eye tracking data enabling better understanding of children’s visual cognitive behaviours. frontline learning research, 6(3), 72–84. https://doi.org/10.14786/flr.v6i3.365 garfield, j. (2002). histogram sorting. statistics teaching and resource library (star). https://amser.org/index.php?p=amser--resourceframe&resourceid=8554 garfield, j. b., & ben-zvi, d. (2008). learning to reason about distribution. in j. garfield & d. ben-zvi (eds). developing students’ statistical reasoning: connecting research and teaching practice (pp. 165– 186). springer. https://link.springer.com/content/pdf/10.1007/978-1-4020-8383-9_8.pdf godau, c. haider, h., hansen, s., schubert, t., frensch, p. a., gaschler, r. (2014). spontaneously spotting and applying shortcuts in arithmetic—a primary school perspective on expertise. frontiers in psychology – cognition, 5, 1664–1078, article e556. https://doi.org/10.3389/fpsyg.2014.00556 goldberg, j. h., & helfman, j. i. (2010). comparing information graphics: a critical look at eye tracking. proceedings of the 3rd beliv’10 workshop: beyond time and errors: novel evaluation methods for information visualization (pp. 71–78). https://doi.org/10.1145/2110192.2110203 guan, z., lee, s., cuddihy, e., & ramey, j. (2006). the validity of the stimulated retrospective think-aloud method as measured by eye tracking. proceedings of the sigchi conference on human factors in computing systems (pp. 1253–1262). association for computing machinery. https://doi.org/10.1145/1124772.1124961 guerra-carrillo, b. c., & bunge, s.a. (2018). eye gaze patterns reveal how reasoning skills improve with experience. npj science of learning, 3, article 18. https://doi.org/10.1038/s41539-018-0035-8 hansen, l. k., & salamon, p. (1990). neural network ensembles. ieee transactions on pattern analysis and machine intelligence, 12, 993–1001. https://doi.org/10.1109/34.58871 heilbronner, r. l., sweet, j. j., attix, d. k., krull, k. r., henry, g. k., & hart, r. p. (2010). official position of the american academy of clinical neuropsychology on serial neuropsychological assessments: the utility and challenges of repeat test administrations in clinical and forensic contexts, the clinical neuropsychologist, 24(8), 1267–1278. https://doi.org/10.1080/13854046.2010.526785 hessels, r. s., niehorster, d. c., nyström, m., andersson, r., & hooge, i. t. c. (2018). is the eyemovement field confused about fixations and saccades? a survey among 124 researchers. royal society open science, 5(8), 1–23. https://doi.org/10.1098/rsos.180502 https://link.springer.com/book/10.1007/978-0-387-21606-5 https://doi.org/10.1113/jphysiol.1967.sp008271 https://doi.org/10.1111/j.1467-9639.1995.tb00720.x https://iase-web.org/documents/intstatreview/02.gal.pdf https://doi.org/10.1007/s10649-022-10151-7 https://doi.org/10.3390/educsci10050141 https://amser.org/index.php?p=amser--resourceframe&resourceid=8554 https://link.springer.com/content/pdf/10.1007/978-1-4020-8383-9_8.pdf https://doi.org/10.3389/fpsyg.2014.00556 https://doi.org/10.1145/2110192.2110203 https://doi.org/10.1145/1124772.1124961 https://doi.org/10.1038/s41539-018-0035-8 https://doi.org/10.1109/34.58871 https://doi.org/10.1080/13854046.2010.526785 https://doi.org/10.1098/rsos.180502 boels et al. 28 | f l r hinton-bayre, a. d. (2010). deriving reliable change statistics from test–retest normative data: comparison of models and mathematical expressions. archives of clinical neuropsychology, 25(3), 244–256. https://doi.org/10.1093/arclin/acq008 hohn, r. w. (1992). an analysis of the components of curriculum-based assessment [doctoral dissertation. university of denver]. https://www.proquest.com/openview/abb4ea5179410900d6e2af9a7473f0e9/1?pqorigsite=gscholar&cbl=18750&diss=y holmqvist, k., örbom, s. l., hooge, i. t. c., niehorster, d. c., alexander, r. g., andersson, r., benjamins, j. s., blignaut, p., brouwer, a-m., chuang, l. l., dalrymple, k. a., drieghe, d., dunn, m. j., ettinger, u., fiedler, s., foulsham, t., van der geest, j. n., witzner hansen, d., hutton, s., … hessels, r. s. (2023). eye tracking: empirical foundations for a minimal reporting guideline. behavior research methods, 55, 364–416. https://doi.org/10.3758/s13428-021-01762-8 hyönä, j. (2010). the use of eye movements in the study of multimedia learning. learning and instruction, 20(2), 172–176. https://doi.org/10.1016/j.learninstruc.2009.02.013 james, g., witten, d., hastie, t., & tibshirani, r. (2013). an introduction to statistical learning. springer texts in statistics, 112. springer. jarodzka, h., van gog, t., dorr, m., scheiter, k., & gerjets, p. (2013). learning to see: guiding students’ attention via a model’s eye movements fosters learning. learning and instruction, 25, 62–70. https://doi.org/10.1016/j.learninstruc.2012.11.004 kaakinen, j.k. (2021). what can eye movements tell us about visual perception processes in classroom contexts? commentary on a special issue. educational psychology review, 33, 169–179. https://doi.org/10.1007/s10648-020-09573-7 kaplan, j. j., gabrosek, j. g., curtiss, p., & malone, c. (2014). investigating student understanding of histograms. journal of statistics education, 22(2), 1–30. https://doi.org/10.1080/10691898.2014.11889701 khalil, k. a. i. (2005). expert-novice differences: visual and verbal responses in a two-group comparison task [master’s thesis, university of massachusetts]. https://scholarworks.umass.edu/theses/2428 konold, c., higgins, t., russell, s. j., & khalil, k. (2015). data seen through different lenses. educational studies in mathematics, 88(3), 305–325. https://doi.org/10.1007/s10649-013-9529-8 konold, c., & pollatsek, a. (2002). data analysis as the search for signals in noisy processes. journal for research in mathematics education, 33(4), 259–289. https://www.jstor.org/stable/749741 kragten, m., admiraal, w., & rijlaarsdam, g. (2015). students’ learning activities while studying biological process diagrams. international journal of science education, 37(12), 1915–1937. https://doi.org/10.1080/09500693.2015.1057775 lai, m., tsai, m., yang, f., hsu, c., liu, t., lee, s. w., lee, m., chiou, g., liang, j., & tsai, c. (2013). a review of using eye-tracking technology in exploring learning from 2000 to 2012. educational research review, 10, 90–115. https://doi.org/10.1016/j.edurev.2013.10.001 lem, s., onghena, p., verschaffel, l., & van dooren, w. (2013). on the misinterpretation of histograms and box plots. educational psychology, 33(2), 155–174. https://doi.org/10.1080/01443410.2012.674006 liaw a, wiener m (2002). classification and regression by randomforest. r news, 2(3), 18– 22. https://cran.r-project.org/doc/rnews/. lievens, f., reeve, c. l., & heggestad, e. d. (2007). an examination of psychometric bias due to retesting on cognitive ability tests in selection settings. journal of applied psychology, 92(6), 1672–1682. https://ink.library.smu.edu.sg/lkcsb_research/5693 https://doi.org/10.1093/arclin/acq008 https://www.proquest.com/openview/abb4ea5179410900d6e2af9a7473f0e9/1?pq-origsite=gscholar&cbl=18750&diss=y https://www.proquest.com/openview/abb4ea5179410900d6e2af9a7473f0e9/1?pq-origsite=gscholar&cbl=18750&diss=y https://doi.org/10.3758/s13428-021-01762-8 https://doi.org/10.1016/j.learninstruc.2009.02.013 https://doi.org/10.1016/j.learninstruc.2012.11.004 https://doi.org/10.1007/s10648-020-09573-7 https://doi.org/10.1080/10691898.2014.11889701 https://scholarworks.umass.edu/theses/2428 https://doi.org/10.1007/s10649-013-9529-8 https://www.jstor.org/stable/749741 https://doi.org/10.1080/09500693.2015.1057775 https://doi.org/10.1016/j.edurev.2013.10.001 https://doi.org/10.1080/01443410.2012.674006 https://cran.r-project.org/doc/rnews/ https://ink.library.smu.edu.sg/lkcsb_research/5693 boels et al. 29 | f l r lumsden, j. (1976). test theory. annual review of psychology, 27(1), 251–280. https://doi.org/10.1146/annurev.ps.27.020176.001343 lyford, a. j. (2017). investigating undergraduate student understanding of graphical displays of quantitative data through machine learning algorithms [doctoral dissertation, university of georgia]. https://iase-web.org/documents/dissertations/17.alexanderlyford.dissertation.pdf lyford, a., & boels, l. (2022). using machine learning to understand students’ gaze patterns on graphing tasks. invited paper: refereed. in s. a. peters, l. zapata-cardona, f. bonafini, & a. fan (eds.), bridging the gap: empowering & educating today’s learners in statistics. proceedings of the eleventh international conference on teaching statistics (pp. 1–6). isi/iase. https://doi.org/10.52041/iase.icots11.t8d2 lyle, j. (2003). stimulated recall: a report on its use in naturalistic research. british educational research journal, 29, 861–878. https://www.jstor.org/stable/1502138 makar, k., & confrey, j. (2004). secondary teachers’ statistical reasoning in comparing two groups. in d. ben-zvi & j. garfield (eds.), the challenge of developing statistical literacy, reasoning and thinking (pp. 353–374). springer. https://rdcu.be/dii3p mcgatha, m., cobb, p., & mcclain, k. (2002). an analysis of students’ initial statistical understandings: developing a conjectured learning trajectory. the journal of mathematical behavior, 21(3), 339–355. https://doi.org/10.1016/s0732-3123(02)00133-5 mcintyre, n. a., draycott, b., & wolff, c. e. (2022). keeping track of expert teachers: comparing the affordances of think-aloud elicited by two different video perspectives. learning and instruction, 80, article 101563. https://doi.org/10.1016/j.learninstruc.2021.101563 mcnemar, q. (1947). note on the sampling error of the difference between correlated proportions or percentages. psychometrika, 12(2), 153–157. https://doi.org/10.1007/bf02295996 meletiou, m. (2000). students’ understanding of variation: an untapped well in statistical reasoning [doctoral dissertation, university of texas]. http://iaseweb.org/documents/dissertations/00.meletiou.dissertation.pdf mokros, j., & russell, s. j. (1995). children’s concepts of average and representativeness. journal for research in mathematics education, 26(1), 20–39. https://doi.org/10.2307/749226 nuzzo, r. (2014). scientific method: statistical errors. nature, 506, 150–152. https://doi.org/10.1038/506150a o’dell, r.s. (2012). the mean as balance point. mathematics teaching in the middle school, 18(3), 148– 155. https://www.jstor.org/stable/10.5951/mathteacmiddscho.18.3.0148 oohira, a., okamoto, m., & ozawa, t. (1981). 正常人の衝動性眼球運動最大速度について [peak velocity of normal human saccadic eye movements (author’s translation)]. 日限会誌 [journal of the japanese society of ophthalmology], 85(11), 2001–2007. https://pubmed.ncbi.nlm.nih.gov/7337121/ or https://www.researchgate.net/publication/15862222_peak_velocity_of_normal_human_saccadic_eye_m ovements_author%27s_transl orquin, j. l., & holmqvist, k. (2017). threats to the validity of eye-movement research in psychology. behavior research methods, 50(4), 1645–1656. https://doi.org/10.3758/s13428-017-0998-z płomecka, m.b. barańczuk-turska, z., pfeiffer, c., & langer, n. (2020). aging effects and test–retest reliability of inhibitory control for saccadic eye movements. eneuro, 7(5). https://doi.org/10.1523/eneuro.0459-19.2020 rokach, l., & maimon, o. (2008). data mining with decision trees: theory and applications. toh tuck link: world scientific. https://doi.org/10.1142/9097 https://doi.org/10.1146/annurev.ps.27.020176.001343 https://iase-web.org/documents/dissertations/17.alexanderlyford.dissertation.pdf https://doi.org/10.52041/iase.icots11.t8d2 https://www.jstor.org/stable/1502138 https://rdcu.be/dii3p https://doi.org/10.1016/s0732-3123(02)00133-5 https://doi.org/10.1016/j.learninstruc.2021.101563 https://doi.org/10.1007/bf02295996 http://iase-web.org/documents/dissertations/00.meletiou.dissertation.pdf http://iase-web.org/documents/dissertations/00.meletiou.dissertation.pdf https://doi.org/10.2307/749226 https://www.jstor.org/stable/10.5951/mathteacmiddscho.18.3.0148 https://www.researchgate.net/publication/15862222_peak_velocity_of_normal_human_saccadic_eye_movements_author%27s_transl https://www.researchgate.net/publication/15862222_peak_velocity_of_normal_human_saccadic_eye_movements_author%27s_transl https://doi.org/10.3758/s13428-017-0998-z https://doi.org/10.1523/eneuro.0459-19.2020 https://doi.org/10.1142/9097 boels et al. 30 | f l r rstudio. scicrunch registry. https://posit.co/download/rstudio-desktop/ russo, j. e. (2010). eye fixations as a process trace. in m. schulte-mecklenbeck, a. kühberger, and r. ranyard (eds.), handbook of process tracing methods for decision research (pp. 43–64). psychology press. https://www.researchgate.net/publication/285189001_eye_fixations_as_a_process_trace scharfen, j., jansen, k. & holling, h. (2018). retest effects in working memory capacity tests: a metaanalysis. psychonomic bulletin & review, 25, 2175–2199. https://doi.org/10.3758/s13423-018-1461-6 schindler, m., & lilienthal, a. j. (2019). domain-specific interpretation of eye-tracking data: towards a refined use of the eye-mind hypothesis for the field of geometry. educational studies in mathematics, 101, 123–139. https://doi.org/10.1007/s10649-019-9878-z setiawan, e. p., & sukoco, h. (2021). exploring first year university students’ statistical literacy: a case on describing and visualizing data. journal on mathematics education, 12(3), 427–448. https://ejournal.unsri.ac.id/index.php/jme/article/view/13202 strohmaier, a. r., mackay, k. j., obersteiner, a., & reiss, k. m. (2020). eye-tracking methodology in mathematics education research: a systematic literature review. educational studies in mathematics, 104, 147–200. https://doi.org/10.1007/s10649-020-09948-1 temkin, n. r., heaton, r. k., grant, i., & dikmen, s. s. (1999). detecting significant change in neuropsychological test performance: a comparison of four models. journal of the international neuropsychological society, 5(4), 357–369. https://doi.org/10.1017/s1355617799544068 tiefenbruck, b. f. (2007). elementary teachers conceptions of graphical representations of categorical data [doctoral dissertation. university of minnesota]. https://conservancy.umn.edu/handle/11299/91699 tobii (n.d.). tobii studio. users’ manual. version 3.4.5. https://www.tobiipro.com/siteassets/tobii-pro/usermanuals/tobii-pro-studio-user-manual.pdf/?v=3.4.5 van gog, t., & jarodzka, h. (2013). eye tracking as a tool to study and enhance cognitive and metacognitive processes in computer-based learning environments. in r. azevedo, and v. aleven (eds.), international handbook of metacognition and learning technologies (pp. 143–156). springer. https://doi.org/10.1007/978-1-4419-5546-3_10 van gog, t., paas, f., van merriënboer, j. j., & witte, p. (2005). uncovering the problem-solving process: cued retrospective reporting versus concurrent and retrospective reporting. journal of experimental psychology: applied, 11(4), 237–244. https://doi.org/10.1037/1076-898x.11.4.237 vygotsky, l. s. (1978). mind in society: the development of higher psychological processes. harvard university press. https://www.hup.harvard.edu/catalog.php?isbn=9780674576292 https://posit.co/download/rstudio-desktop/ https://www.researchgate.net/publication/285189001_eye_fixations_as_a_process_trace https://doi.org/10.3758/s13423-018-1461-6 https://doi.org/10.1007/s10649-019-9878-z https://ejournal.unsri.ac.id/index.php/jme/article/view/13202 https://doi.org/10.1007/s10649-020-09948-1 https://doi.org/10.1017/s1355617799544068 https://conservancy.umn.edu/handle/11299/91699 https://www.tobiipro.com/siteassets/tobii-pro/user-manuals/tobii-pro-studio-user-manual.pdf/?v=3.4.5 https://www.tobiipro.com/siteassets/tobii-pro/user-manuals/tobii-pro-studio-user-manual.pdf/?v=3.4.5 https://doi.org/10.1007/978-1-4419-5546-3_10 https://doi.org/10.1037/1076-898x.11.4.237 https://www.hup.harvard.edu/catalog.php?isbn=9780674576292 1. introduction 2. theoretical background 2.1 review of statistics education literature 2.1.1 histograms are persistently misinterpreted 2.1.2 dotplots are not always correctly interpreted 2.2 review of literature on eye-tracking in education 2.2.1 use of spatial gaze measures to reveal students’ strategies for interpreting histograms 2.2.2 connecting gaze data to students’ strategies 2.3 learning from a series of items: the practice effect 2.4 rationale for using a machine learning algorithm 3. materials and methods 3.1 participants: pre-university track students grades 10–12 3.2 materials: histogram and dotplot items requiring comparing and estimating means 3.2.1 estimating and comparing arithmetic means reveals students’ knowledge 3.2.2 four histogram items—dotplots items in between 3.3 data collection methods: eye-tracking, stimulated recall 3.3.1 data collection with an eye tracker 3.3.2 data collection through stimulated recall verbal reports 3.4 data analysis through a machine learning algorithm 3.4.1 gaze-data analysis through a machine learning algorithm (random forest) 3.4.2 a failed mla—using dwell time on aois 3.4.3 a successful approach—exploring saccade direction and magnitude 4. results 4.1 main changes in students’ gaze patterns on histograms 4.2 students’ post-activity verbal descriptions of approaches to histogram items 4.3 differences in students’ answers to histogram items 5. conclusions and discussion key points acknowledgments conflict of interest data availability authors’ contributions references frontline learning research vol. 12 no. 2 (2024) 1 27 issn 2295-3159 corresponding author: max kusters, kolffpad 1, 2333 bn, leiden, the netherlands, m.c.j.kusters@iclon.leidenuniv.nl. doi: https://doi.org/10.14786/flr.v12i2.1419 developing scenarios for exploring teacher agency in universities: a multimethod study max kusters1, arjen de vetten1, wilfried admiraal2, & roeland van der rijst1 1 iclon graduate school of teaching, leiden university, p.o. box 905, leiden, 2300 ax, the netherlands 2 centre for the study of professions, oslo metropolitan university, po box 4 st. olavs plass, n-0130, oslo, norway article received 21 december 2023 / article revised 16 april 2024 / accepted 15 may / available online 13 june abstract lecturers who are actively engaged in shaping their teaching and teaching practices demonstrate agency. teacher agency has increasingly been described as a key factor in educational development at universities. lecturers are expected to innovatively develop courses and continuously improve their teaching practices to respond to, for example, student needs and labor market demands. in this multimethod study, we examined the process of developing and validating scenarios for measuring teacher agency in universities. we conducted four studies to create 23 scenarios that capture the complex nature of teacher agency. first, we interviewed university lecturers to identify bumpy moments in their teaching practice, and so found scenarios based on real-life experiences. then, we employed two expert panels, to evaluate and refine the scenarios, which enhanced their validity. finally, we used a pilot study to standardize the data collection procedures. our multimethod study has established reliability by triangulating methods and researchers, involving multiple stakeholders, and providing detailed descriptions of the research process. this project holds implications for research and practice. the scenarios can be used in professional academic development programs for the collection of research data and to promote self-reflection, peer consultation activities, and professional growth and agency among university lecturers. keywords: teacher agency; university; qualitative scenario construction; vignettes; professional development mailto:m.c.j.kusters@iclon.leidenuniv.nl https://doi.org/10.14786/flr.v12i2 kusters et al 2 | f l r introduction the importance of teacher agency for effective teaching practices has long been recognized (aspburymiyanishi, 2022; priestley et al., 2012; pyhältö et al., 2013). although most research on teacher agency relates to primary and secondary education (cong-lem, 2021), in recent years it has also become clear that teacher agency has a critical impact on teaching and learning at universities (kusters et al., 2023; vähäsantanen et al., 2020). at its core, teacher agency implies that lecturers shape their responsiveness to problematic situations (biesta & tedder, 2007). in an ecological approach, teacher agency entails that this ‘responsiveness’ is informed by lecturers' professional histories, oriented towards future objectives and aspirations, and enacted in concrete situations (biesta & tedder, 2007; priestley et al., 2015). it is both constrained and supported by relational, structural, and material resources available to actors (priestley et al., 2015). teacher agency is required when situations call for a considered solution. this means that a solution may not be imminent, but more options could work; it requires agency to consider which option is best given the broader purpose of the practice in which they work (leijen et al., 2019; priestley et al., 2015). teacher agency is therefore not seen as static, but rather as a dynamic and context-dependent concept (jenkins, 2019; kusters et al., 2023), which constantly changes in response to various factors, such as educational environments, policy changes, and personal beliefs, values, and goals. it evolves in and adapts to the existing context (biesta et al., 2015). this ongoing evolution and adaptation to contexts (or: temporality) raise significant concerns regarding the accurate methodologies to measure teacher agency. developing a tool for measuring teacher agency, therefore, brings challenges such as being able to do justice to the dynamic character and temporality that teacher agency entails. previous methodologies, despite their utility, have demonstrated shortcomings in adequately addressing these aspects. to illustrate, teacher agency can be measured through self-report tools (ghiasvand et al., 2023; leijen et al., 2021; vähäsantanen et al., 2019) as well as classroom observation methods (hiver & whitehead, 2018). questionnaires measuring self-reported agency have the disadvantage of lacking context, and therefore not capturing aspects of considerations about participants’ responses to the items. although classroom observations provide valuable insights about actual behavior in a real-life context, they often take up time and resources and cannot capture the full scope of teacher agency in different, multiple contexts. moreover, lecturers’ reflections cannot be collected by mere observation, which means additional data collection is needed. alternatively, methodologies focused on utilizing real-life stimuli are suitable to effectively incorporate context at a given time. because of our goal to develop scenario-based research that focuses on eliciting teacher agency, our methodological approach is based on the results observed in research using reallife stimuli (for examples see sannino & engeström, 2016; yang, 2021, 2022). recognizing the effectiveness of such studies, there is a need for tools that provide real-life teaching experiences to measure teacher agency in universities. to that end, we developed scenarios (i.e., short descriptions of situations that demand teacher agency) based on real-life teaching experiences at universities. in this project, we aim to present scenarios, developed in a trustworthy way and shaping internally valid conditions for measuring teacher agency in universities. defining teacher agency teacher agency within an ecological approach arises from the intricate interplay between individuals' abilities and environmental conditions. this underscores the importance of an instrument that not only measures lecturers' abilities but also elements from the past, future, and available resources in the present that affect the specific ecologies in which lecturers work (priestley et al., 2015). approaching teacher agency from an ecological perspective also recognizes agency’s fluidity over time. given the conditions under which agency emerges and the implications for measuring teacher agency, the challenge is to address the context in which lecturers work (emirbayer & mische, 1998). universities provide a diverse range of environments, with their own educational beliefs and values for lecturers (gibson, kusters et al 3 | f l r 1986). real-life teaching scenarios allow the exploration of teacher agency in various contexts and leave room for considerations (e.g., which affordances are perceived and/or why one solution is preferred over another) in the decision-making process. to understand lecturers’ perceived affordances and considerations we have tried to represent real-life teaching situations through written scenarios. scenario definition and purpose in methodological contexts, scenarios are descriptive representations of specific situations designed to simulate real events or problems (hughes & huby, 2004; jeffries & maeder, 2005; jenkins et al., 2010). scenarios can have different forms, such as image, video, audio, or written narrative, which research participants are requested to comment on (hughes & huby, 2002). written scenarios are most effective when they range from 50 to 200 words (jeffries & maeder, 2005). they can revolve around people, situations, or events. finch (1987) highlights the importance of employing scenarios in research in order to thoroughly investigate the central phenomenon under examination. by eliciting responses and encouraging discussion valuable information is gathered, providing insight into participants' beliefs, values, judgements, and attitudes. scenarios are also valuable tools for gathering in-depth insights into participants’ nuanced thoughts on varying topics (simon & tierney, 2011; torres, 2009). although in the literature the term vignette is often used (skilling & stylianides, 2020), we here intentionally use the term scenario. vignettes evoke the idea of a snapshot or static image of a situation, but our goal in constructing the scenarios was to represent vivid, authentic situations. research on teacher agency is suited to a scenario study, because context-specific scenarios fit the context-dependent nature of teacher agency. they provide a powerful tool to explore and understand how lecturers demonstrate agency within different university teaching practices, and this approach ultimately contributes to better informed and contextually relevant educational policies and practices. the main consideration when developing compelling scenarios is whether they genuinely reflect the complex and varied educational learning situations that lecturers encounter in their professional roles (stravakou & lozgka, 2018). this type of study also prompts an examination of the reliability of the responses and decisions derived from these scenarios (gould, 1996) as indicators of teacher agency among lecturers. aim of the project the aim of the current project, consisting of four studies, was to develop and test a real-life instrument to measure teacher agency from an ecological perspective. as rushton & bird (2023) pointed out, elements of teacher agency are nonlinearly intertwined and thus do not act in isolation but rather interact with each other in situations, which led us to provide context-rich descriptions of situations for measuring agency. this type of research emphasizes the collection of in-depth, contextual data through methods to understand the complex and dynamic nature of human behavior, interactions, and structures within specific social and cultural contexts (lave & wenger, 1991). in this methodological paper we report on the development of scenarios by which to capture the dynamic nature of teacher agency. the essential challenge for the current project here lies in ensuring that the scenario content accurately represents real-life teaching situations, particularly in terms of internal validity, an aspect that has often been overlooked in prior scenario-based research (hughes & huby, 2004). to address this challenge we relied on a framework that contains the critical elements needed for comprehensive scenario development introduced by skilling and stylianides (2020). the three key elements include the conception of the content, design of the scenarios, and administration of the protocol (skilling & stylianides, 2020). moreover, in the design of the four studies we aimed for maximum trustworthiness of the research process. to that end we used four critical criteria as discussed by guba (1981): credibility refers to the believability of the research findings; dependability is the consistency and stability of the research findings over time and in different circumstances; confirmability relates to the objectivity and neutrality of the research, ensuring that the researcher’s biases or perspectives do not influence findings; and finally, transformability relates to the extent to which qualitative research findings can be applied or generalized to other contexts or settings. to best meet the above criteria, we had to consider which themes of the scenarios are of interest to university kusters et al 4 | f l r lecturers, which writing format contributes most to recognition of the scenarios for participants, how experts assessed the quality of the scenarios, and which scenarios can best be used to make teacher agency measurable. to this end, in the inductive phase of this research we collected specific observations and data by which to identify general patterns and theories. subsequently, in the deductive phase we tested the scenarios empirically. this paper is structured around four separate but interconnected studies (figure 1). these studies were intended together to contribute to answering the overarching research question: in what ways can representative scenarios be developed to measure teacher agency in universities? the following sub-questions for each study were addressed: study 1: which teaching scenario themes are representative of university teaching? study 2: should valid scenarios for teacher agency be written in the firstor third-person perspective, and in open-ended or closed-ended format? study 3: how do experts evaluate the likelihood of specific scenarios actually eliciting teacher agency? study 4: to what extent do lecturers prefer particular scenarios, and which scenarios are useful for eliciting multiple solutions? figure 1: four studies for developing scenarios the first three studies were intended to construct the scenarios inductively, attempting to do justice to the key elements of conception and design (skilling & stylianides, 2020). in study 1, interviews were conducted to determine the themes and content of the scenarios. in study 2 and study 3, expert panels determined writing style and perspective and the alignment of content with the research goal, respectively. study 4 was a deductive pilot study intended to test the scenarios on lecturers, and this aim corresponded to the key element of administration (skilling & stylianides, 2020). below, we present the four studies. the institute’s ethical review committee (iclon-irec 2021-02) gave approval for our research. in addition, participants provided consent for each study. participants consisted of phd candidates, post-doctoral fellows, assistants, associates, and full professors. any names displayed are kusters et al 5 | f l r fictitious so as to ensure participants’ privacy. the four studies all had different participants. table 1 lists the participants in each study. table 1 characteristics of participants and occupational status study 1 study 2 study 3 study 4 total n 28 37 13 30 108 nationality • dutch • norwegian • finnish • other 23 3 2 0 26 0 0 11 10 1 1 1 26 0 0 4 85 4 3 16 gender • female • male 13 15 28 9 9 4 17 13 67 41 university 7 1 9 9 12* position • phd candidate • lecturer • professor 0 27 1 22 12 3 2 6 5 0 24 6 24 69 15 *note. participants came from 12 different universities in total. there is overlap between the universities used in the different studies. study 1 study 1 focused on identifying key themes in which the complexity of teaching practice calls for agentic manifestations. to accomplish this, we conducted interviews with lecturers. the central question in study 1 was: which teaching scenario themes are representative of university teaching? method participants interviews were conducted with 28 lecturers from seven research-intensive universities in the netherlands, norway, and finland (table 1). to obtain a representative sample of the population, lecturers from various domains were approached via a brief introductory email explaining the purpose of the project. later they received a letter with more information, also on the procedure of the project, along with an agreement to participate for them to sign. the authors recruited the lecturers through their professional networks, aiming to investigate bumpy moments for lecturers with both teaching and research duties. only lecturers with a phd were included because we wanted to exclude casual teaching staff, often hired temporarily for teaching assignments only. all participants had explicit teaching duties in their employment contracts along with their research tasks. they all had more than five years of teaching experience, teaching in small-group settings as well as giving lectures. during the interviews, participants were specifically asked about teaching context, i.e., working group or lecture. data collection data collection consisted of online interviews, each lasting approximately 60 minutes, with participants able to choose between english or dutch as their preferred language of communication. in the interviews, participants were asked to recall situations from their teaching practice in which they had to make decisions: “could you please indicate those moments when you acted in a particular way, and with hindsight feel that you could just as well have acted differently?” the term bumpy moment (romano, 2006; van kan et kusters et al 6 | f l r al., 2010) was employed to describe these challenging and dilemma-laden situations. the bumpy moment referred not to a situation where lecturers were unable to act, but to a situation that, in hindsight, could have provided several legitimate and competing courses of action. all interviews were recorded, allowing for later reviewing and in-depth analysis. each interview was transcribed verbatim for data analysis and coding. to ensure a comprehensive understanding of real-life teaching situations in the university context, participants were requested to contemplate their bumpy moments before the interview. they were asked to share at least two of these moments and supply brief descriptions, which helped the interviewer to prepare properly. in cases when participants did not provide their bumpy moments in advance, they were given some time to compile and identify specific incidents for discussion. throughout the interview, the interviewer asked questions in order to clarify and delve deeper into the details of each moment, including circumstances, people involved, location, and underlying considerations. this approach made it easier to collect a diverse and extensive range of themes, thus providing deeper insights as a basis for developing scenario storylines. analysis the transcribed interviews were analyzed by the principal investigator using atlas.ti, and following a systematic approach to ensure a comprehensive examination. interim and final results were discussed in consultation with the other authors. initially, the researchers selected all relevant parts of the transcripts on the basis of the guiding questions, i.e., whether there was room for choices within the given scenario and whether it was possible for a teacher to change the situation. having multiple options to act in teaching settings is a prerequisite given the definition of the ecological approach (priestley et al., 2015). interview fragments that met both criteria were then selected for further study using thematic content analysis (silverman, 2020). overarching themes were created using all the relevant fragments; the bumpy moments could then be subdivided according to these themes. findings the interviews were intended to uncover the topics on which scenarios needed to be based. each interviewee came up with an average of 2.6 bumpy moments, which resulted in 74 bumpy moments in total. the following fragment illustrates a code concerning “student preparation”. james stated: the question is always: should you assume that students have done the preparation that you have actually prescribed for them? and does that mean that you then stick to that or that the moment you find out, that that preparation hasn’t taken place sufficiently, that you then adjust on the spot and go off. always the choice of -back to what they should have studied? and that’s always a trade flexibility on the one hand, on the other hand adapting to the information needs of the student, for which of course there is also a lot to be said. but then with that you send the signal: it’s okay not to be prepared, because they’re starting back at the beginning anyway. important choice you keep running into. in this phase, we tried to list the content, context, and recurring themes in the data. overarching content and similar elements (e.g., other fragments about “student preparation,” similar to the example above) found in multiple fragments formed the basis for the subsequent coding process. in this case, 24 codes were generated based on the overarching and similar content identified. although some codes occurred more frequently than others (see appendix a), each code served as a theme. to sum up, in study 1 we examined how to develop scenarios that reflected a realistic representation of teaching practice at universities. we chose an emic perspective to understand bumpy moments from the point of view of university lecturers in the field. we interviewed professionals in the field to get a realistic representation of the bumpy moments experienced by lecturers. this first part of our research provided valuable kusters et al 7 | f l r insights into the complexity of teaching practice at universities. through interviews and qualitative thematic content analysis we discovered a diverse array of 24 different themes on which to focus the scenarios. study 2 building on the insights from the first study, study 2 focused on the question of narrative perspective and design of the scenarios. the goal was to ensure that the scenarios were structured and written in a way that contributed to the lecturers' identification with the scenario. the literature does not provide a theoretical basis for determining these fundamental design principles (skilling & stylianides, 2020). first, we had to decide whether the scenarios should be written from a firstor third-person perspective. second, as skilling and stylianides (2020) note in their framework, researchers must determine whether scenarios have an open or closed ending, depending on the goal. therefore, in study 2, we also determined the appropriate endings for the scenarios based on the research aims, as our objective was to accurately reflect real-life teaching practice within the scenarios. the guiding question here was: should valid scenarios for teacher agency be written from a firstor third-person perspective, and in an open-ended or closed-ended format? method participants in study 2, 37 researchers with backgrounds in educational science, ranging from phd candidates to full professors, participated as an expert panel (table 1). all panel members had varying experiences with quantitative and qualitative research methods and were given information by the first author during a research group meeting. participants were aware that participation in the panel was voluntary. data collection for study 2 and study 3 we used comparative judgement (cj). cj has been proposed as an assessment technique that can produce consistent results (pollitt, 2012). the cj method relies on comparisons rather than absolute judgements. for this, panelists are shown pairs of draft scenarios, also called appearances, and asked to decide which of the two best suited to the topic being assessed. researchers can then derive a scale value from these assessments. using multiple assessors helps reduce individual bias and contributes to a more balanced and objective assessment. for this purpose, we used the comproved software (see https://comproved.com/en/comparing-tool/) to rank panelists’ judgements, by comparing two scenarios at a time and judging which one best fit the leading question. cj is based on the bradley-terry-luce model (btl model; bradley & terry, 1952; luce, 1959), a statistical model that provides a framework for analyzing cj data. the model assumes that each item (draft scenario) has an underlying latent quality or preference score, and that the probability of choosing one item over another depends on the difference between their scores. the btl model is formulated as follows: where xij = 1 if appearance j is assumed to be superior to appearance i, and vi and vj are the estimated ability values, in logit scores, of the respective appearances (verhavert et al., 2018). the comproved tool uses multiple assessors to rank scenarios according to certain evaluation criteria. each assessor gets to see the same set of scenarios but in different and randomized combinations. in the end the tool consolidates all input received. the result is a quality scale ranking the scenarios, with their corresponding ability scores, in order of preference. the reliability of the scale is expressed in the scale separation reliability (ssr), where ssr serves as a measure of the internal consistency of the assessment kusters et al 8 | f l r results. a study by verhavert et al. (2018) has demonstrated that the ssr can be used to estimate both interrater reliability and split-half reliability. for study 2, we selected the four most common themes from the 24 resulting from study 1, and wrote four scenarios based on these themes. we used these to determine the best perspective for each scenario, and whether it should have an openor closed-ending format. an open ending means that the scenario ends with: “so i knew i had to come up with a solution.” a closed ending means that a solution was already given in the scenario. each of the four scenarios was presented in four different ways: 1) first-person perspective with an open ending; 2) first-person perspective with closed ending; 3) third-person perspective with an open ending; and 4) third-person perspective with closed ending (for an example, see appendix b). thus, the dataset contained 16 documents (draft scenarios) in all. in comproved two randomly selected draft scenarios were presented each time, from which the participant was asked to choose the one they were best able to identify with. participants also indicated why they preferred one scenario over another, which allowed us to collect qualitative data and so gain insight into the participants' considerations. participants made four comparisons per person (verhavert et al., 2019), resulting in an ssr of .59, which can be considered low. comproved also showed “misfit judges”, i.e., assessors that diverged from the average. in this dataset three misfit judges were detected. by deleting these three assessors, we reached an acceptable ssr of .70. however, although with this modification slight changes in the ranking occurred, the main result did not change. therefore, we decided to retain the misfit judges because their answers to the open questions gave valuable insights into the reasons behind their choices. analysis the analysis focused on distinguishing the different forms, aimed at being able to empathize with the scenario. the most popular appearance form (i.e., the scenarios most likely to be selected as favorites) was included in the next study. the additional qualitative data explained why participants identified better with one form over the other. findings four scenarios were presented to the panel (n = 37) in four forms. in each case the panelists were asked to choose which of two scenarios they most identified with as participants. from this, we constructed a ranking (table 2). the graph for the rankings (figure 2) shows the extent to which a particular scenario is likely to win over another scenario. the results indicate that the top seven scenarios were all written in the first-person perspective. although the ranking regarding the closings of the scenarios was less clear, the qualitative data were compelling – they indicated that an open-ended scenario, i.e., one without a given solution, worked best to promote recognizability. it turned out that if a solution was given panelists began to evaluate the solution instead of the rest of the scenario, and thought about whether they could identify with the solution rather than with the bumpy moment. emma, for instance, mentioned the following: the solution given is totally improbable, something i would never go for. so, i do not identify with this. kusters et al 9 | f l r table 2 ranking of scenarios rank scenario perspective closing 1 f 1st perspective open 2 i 1st perspective closed 3 n 1st perspective open 4 e 1st perspective closed 5 j 1st perspective open 6 b 1st perspective open 7 m 1st perspective closed 8 o 3rd perspective closed 9 p 3rd perspective open 10 k 3rd perspective closed 11 g 3rd perspective closed 12 l 3rd perspective open 13 d 3rd perspective open 14 c 3rd perspective closed 15 h 3rd perspective open 16 a 1st perspective closed note. no. of assessors: n = 37 we attribute the fact that scenario a, ranked 16th, was written in the first person but finished last to the implausibility of the solution described. the qualitative data showed that panelists rarely chose this scenario because of its improbably worded ending. the content of this scenario is the same as scenario c which finished 14th. figure 2 study 2: ranking on the basis of ability scores kusters et al 10 | f l r to sum up, in study 2, educational researchers were involved in a panel to determine the ideal perspectives and endings of the scenarios. once we learned how best to write the scenarios in view of the main goal of recognizability, we were able to use experts in our research field to assess the content of the scenarios in study 3. study 3 after the most appropriate perspectives and closings for the scenarios had been established, the third study then involved asking participants to rank the scenarios with respect to their purpose of eliciting teacher agency. for study 3, all 24 themes were developed into scenarios written from the first-person perspective and with open endings. the central question underlying this third study was: how do experts evaluate the likelihood of specific scenarios actually eliciting teacher agency? method participants in study 3, an expert panel consisting of 13 educationalists specializing in teacher agency and/or higher education was assembled (table 1). the panelists were selected from the netherlands, norway, finland, and hungary. the composition of the panel was expressly intended to focus on expertise and perspectives relevant to research on teacher agency manifestations in universities. data collection as in study 2, we used cj to determine the content of the scenarios. the panel was presented with all 24 first-person, open-ended scenarios, requiring the experts to make the comparisons in the same way as the participants in the second study. as in study 2, experts were also asked to explain why they preferred one scenario over another. these qualitative data played a central role in study 3. rather than relying solely on quantitative measures or numerical rankings, we sought to capture the nuanced perspectives and reasoning behind the experts’ choices. the assignment for the panelists was to choose which of two scenarios was best for eliciting manifestations of teacher agency, based on the following definition by priestley et al. (2015): (…) teachers achieve agency when they are able to choose between different options in any given situation and are able to judge which option is most desirable in light of the wider purposes of the practice in and through which they act. agency is not present if there are no options for actions or if the teacher simply follows routinized patterns of habitual behavior with no consideration of alternatives. (p. 141) in this case, the participants made 20 comparisons per person because there were more scenarios and fewer assessors than in study 2 (verhavert et al., 2019). the ssr in this dataset was .70. analysis panelists chose which of two scenarios they thought most appropriate for eliciting teacher agency, and so indicated why one scenario they considered more suitable than the other. the main idea was to find the most appropriate scenario of two by comparing the chances of success or a favorable outcome. a positive score (above 0) suggests a likelihood of success greater than 50%, meaning that one scenario is more likely to be successful than the other. conversely, a score below 0 indicates a likelihood of success lower than 50%, making the scenario less likely to succeed than the other. the greater the deviation from 0, whether positive or negative, the stronger the indication of success or failure in the scenario. additionally, we analyzed the qualitative data to uncover why certain scenarios were favored over others. in the analysis of the open answers we compared the views of different experts on the suitability of the scenarios to elicit teacher agency, focusing on similarities and contradictions. this approach provided insight into the diversity of opinions. the kusters et al 11 | f l r information gathered was used to create a comprehensive synthesis of the experts' key points and conclusions, leading to a better understanding of why scenarios are suitable for eliciting teacher agency. findings the ranking resulting from the analyses (figure 3) shows the ability score, indicating which scenario is most likely to win over another scenario. figure 3 study 3: ranking on the basis of ability scores the ability scores (also converted to odds and probability scores, see table 3) show only slight variations. in particular, figure 3 shows that the ability scores of scenarios 4 through 21 (inside the green rectangle in table 3) were relatively close to each other, implying that the likelihood of eliciting teacher agency from these scenarios was nearly equal, i.e., the scenarios tested have similar qualities. -3,00 -2,00 -1,00 0,00 1,00 2,00 3,00 coherence between courses balancing personal mentoring with professional responsibilities controversial topics in course workload pedagogical choices view on teaching unmotivated students students differ in their pre-knowledge student expectations unexpected students' questions change of job position implement changes in course personal problems of student offended by joke teaching quality of colleague interaction in online teaching unprepared students in lecture overwhelmed as a new lecturer educational re-design university rules lecture preparation administrative tasks in thesis supervision teaching space technical issues kusters et al 12 | f l r table 3 ability, odds, and probability scores rank scenario ability (logit) odds probability 1 coherence between courses 2.27 9.7 0.91 2 balancing professional responsibilities 1.52 4.6 0.82 3 controversial topics in course 1.47 4.3 0.81 4 workload 0.84 2.3 0.70 5 pedagogical choices 0.78 2.2 0.69 6 view on teaching 0.57 1.8 0.64 7 unmotivated students 0.51 1.7 0.62 8 students differ in their pre-knowledge 0.46 1.6 0.61 9 student expectations 0.42 1.5 0.60 10 students’ unexpected questions 0.40 1.5 0.60 11 change of job position 0.09 1.1 0.52 12 implement changes in course -0.11 0.9 0.47 13 personal problems of student -0.30 0.7 0.43 14 offended by joke -0.40 0.7 0.40 15 teaching quality of colleague -0.41 0.7 0.40 16 interaction in online teaching -0.47 0.6 0.38 17 unprepared students in lecture -0.49 0.6 0.38 18 overwhelmed as a new lecturer -0.59 0.6 0.36 19 educational re-design -0.60 0.5 0.35 20 university rules -0.61 0.5 0.35 21 lecture preparation -0.78 0.5 0.32 22 administrative tasks in thesis supervision -0.92 0.4 0.28 23 teaching space -1.12 0.3 0.25 24 technical issues -2.42 0.1 0.08 the qualitative data provided insight into how the scenarios could be improved, for instance the importance of a clear description of the bumpy moment. to illustrate, one of the panelists said: the person appreciates the freedom – that seems to be the most important part – “sometimes it feels that everyone is working on their own island.” according to this formulation, this seems to be only a minor problem. this panelist indicated that the bumpy moment “lecturers are working on their own island” could be phrased more sharply. in this way we were able to tighten up several scenarios that, as evidenced by the experts’ responses, sometimes did not adequately reflect the bumpy moment. panelists were unanimous only in the case of the teaching space scenario; as became clear from their explanations regarding the open-ended questions, this scenario did not elicit teacher agency. for this reason, we eliminated it. for the remaining 23 scenarios the experts gave arguments as to why one scenario was superior to the other. we used the qualitative data to further sharpen up some of the scenarios; in two cases, we specified the bumpy moment more clearly and in two cases we modified the title. the final 23 scenarios can be found in appendix c. to sum up, study 3 was designed to shed light on expert opinions and the factors influencing their judgements on the potential of scenarios to elicit teacher agency. the combination of the rating system and qualitative data collection allowed for a robust exploration of the expert opinions, ultimately contributing to a more complete understanding of the research question. in particular, the qualitative data provided insight into kusters et al 13 | f l r the reasons why participants believed certain scenarios were appropriate. we were also able to use the input from the data to modify scenarios if, for example, it appeared that the bumpy moment was unclear, the wording was not exact enough, or the title did not seem to cover the main message. this validation method resulted in 23 scenarios, which we tested in a pilot study. study 4 finally, study 4 was a pilot study deductively testing the scenarios created in the inductive phase. the primary focus here was to obtain insight into which scenarios resonated most with actual participants, and to determine the operational feasibility and usability of these scenarios for research purposes. a scenario must, first, be recognizable to a participant, and second, elicit multiple solutions if it is to be an effective tool for providing a differentiated indication of teacher agency. this criterion corresponds with the definition of teacher agency, which states that lecturers can choose between different options in a situation and can assess which option is the most desirable (priestley et al., 2015). in this fourth study we tested the scenario set using two quality criteria: which scenarios participants prefer, and whether scenarios elicit multiple responses. the central question in study 4 was: to what extent do lecturers prefer particular scenarios, and which scenarios are useful for eliciting multiple solutions? method to examine the usability of the scenarios, think-aloud sessions were needed to reveal whether participants were able to devise multiple solutions, thus indicating teacher agency. if for a specific scenario participants came up with multiple solutions, we assumed that the answer to that scenario was not obvious. non-obviousness is a prerequisite for an effective scenario, because teacher agency can only be achieved if there are multiple options for action (priestley et al., 2015). participants the think-aloud sessions were conducted with 30 lecturers from 9 research-intensive universities in the netherlands. these were tenured lecturers with phd degrees. their roles involved both teaching and research responsibilities. they were responsible for at least delivering lectures or instructional sessions to students, and engaged in research activities. all lecturers had more than five years of university teaching experience. data collection data were collected in two stages. first, prior to the think-aloud sessions participants were asked to select five scenarios they wanted to address based on recognizability. participants received the 23 scenarios in random order to avoid order bias, prior to the sessions. we then collected the participants' selections. second, during the think-aloud sessions, the lecturers were asked questions about possible solutions to a specific scenario and corresponding considerations. each session, monitored on site by the principal investigator, lasted a maximum of 60 minutes, in which an average of 3.3 scenarios could be covered. participants received the scenarios printed out and laminated on an a6 sheet. the participant read one of the chosen scenarios aloud, each scenario ending with the phrase, “so i knew i had to come up with a solution.” at that point the thinking aloud began, in which the participant had to think of possible solutions to address this scenario. the researcher only asked clarifying questions about whether other solutions were conceivable. when no further solutions were found, a new scenario was presented. the think-aloud sessions were recorded for later transcription and analysis. analysis the analysis consisted of two parts. first, we counted how often participants put each scenario in their top five, in order to determine their preferences for the scenarios. because we did not want to demand too kusters et al 14 | f l r much time or concentration from the participants, we kept to one-hour sessions and for that reason also kept track of which scenarios had been covered during the session. second, we wanted to gain insight into participants’ repertoire of solutions, and hence calculated the average number of solutions they came up with. findings preference for scenarios to determine which scenarios resonated most with the participants, we analyzed the frequency with which each scenario was chosen in the participants' selections (table 4). table 4 frequency choices, coverage, and average solutions scenario frequency chosen (%) frequency covered average solutions 1. students' personal problems 16 (10.8) 13 3.8 2. workload 15 (10.0) 9 2.6 3. coherence between courses 12 (8.0) 7 2.4 4. technical issues 12 (8.0) 6 3.7 5. student expectations 12 (8.0) 5 4.3 6. students differ in prior knowledge 10 (6.7) 7 3.0 7. unprepared students in lecture 9 (6.0) 7 2.9 8. students’ views on teaching quality of colleague 9 (6.0) 5 3.2 9. educational re-design 8 (5.3) 7 2.7 10. implementing changes in course 8 (5.3) 4 3.0 11. students’ unexpected questions 6 (4.0) 2 4.0 12. balancing professional responsibilities 5 (3.3) 4 3.3 13. lecture preparation 4 (2.7) 4 3.5 14. view on teaching 4 (2.7) 4 3.5 15. interaction in online teaching 4 (2.7) 3 3.0 16. administrative tasks in thesis supervision 3 (2.0) 3 2.7 17. controversial topics in course 3 (2.0) 2 4.5 18. being overwhelmed as a new lecturer 3 (2.0) 1 2.0 19. taking offense at joke 2 (1.3) 2 3.0 20. unmotivated students 2 (1.3) 1 4.0 21. change of job position 1 (0.7) 1 5.0 22. pedagogical choices 1 (0.7) 1 4.0 23. university rules 1 (0.7) 1 3.0 total 150 99 3.4 note. five scenarios for each of the participants (n =30); percentage is number of times chosen relative to the total of 150 choices the findings indicate that participants had different preferences for the scenarios presented. students' personal problems and workload, with popularities of 10.8% and 10.0% respectively, emerge as favorites among all options, with the note that although they are the most preferred choices, only approximately one in ten participants selected one of the two. these findings shed light on the diversity of interests and priorities kusters et al 15 | f l r among the participants in our study, and highlight the need for a flexible and adaptive approach when these scenarios are used to elicit manifestations of teacher agency. repertoire of solutions our second aim was to investigate if the scenarios did indeed elicit multiple solutions. to achieve this, we quantified the solutions provided by each individual for each scenario. the average solutions per scenario are shown in the last column of table 4. the data show that there is not one scenario in which the solution is evident, so that all the scenarios seem suitable for eliciting teacher agency. the results indicate that when faced with scenarios, a lecturer is likely to generate an average of two or more solutions. this means a participant must make an informed choice between at least two possible solutions. because we left the choice of scenarios to the participants in this pilot, not all scenarios could be tested equally well; nevertheless, we can say that study 4 has shown that all scenarios proved their purpose at least once. below, we provide a fragment of a think-aloud session to show how we measured what we understood by multiple solutions: scenario: student expectations as a lecturer, i notice that students increasingly expect individual feedback and guidance. students’ expectations seem to have changed since my own student days, and this has put additional pressure on me as a lecturer. i appreciate that students value feedback and guidance, i understand that this is a crucial part of their learning process, and i do my best to meet their expectations. but one day, i felt completely overwhelmed by the large number of emails i received from students who asked for feedback or a one-on-one meeting. so i knew i had to come up with a solution. below we show olivia’s response, with her solutions underlined to clarify how we analyzed them: “yeah, so there are options out there that you can consider. you can organize feedback in different ways. you can try to [1] put less pressure on individual students for feedback and instead focus on providing collective feedback that benefits everyone. another approach is to [2] narrow down your feedback focus. for instance, you could delve deep into specific parts of students’ work and provide more general feedback for the rest. this way, you can frame your feedback more efficiently. this is a viable choice. alternatively, you can [3] incorporate other types of feedback, such as peer feedback. this way, students can give feedback to each other, reducing the feedback burden on you. also, you can [4] communicate with students, explaining the limitations of time you have as a teacher to fully meet all their needs. you can clarify that due to the available time, it’s challenging to address every aspect perfectly. this helps them understand that you’re doing your best within the given constraints. these are the options i see in front of me.” in this example, olivia read the scenario aloud and discussed the preferred options for dealing with this problem by thinking aloud. because there are several affordances – four in this case (see underlining) – and it is unclear which option is best, we can deduce that this scenario does elicit agentic manifestations. the participant can then determine which option is the best in the given situation. to sum up, study 4 allowed us to examine whether the scenarios we developed and implemented proved useful. our analysis focused on the breadth and purpose of this method: are we able to develop scenarios that reflect educational practice and on which lecturers' decision space is tested? the review provided valuable insights into the effectiveness of these scenarios and shed light on their applicability and impact on practice. although not all scenarios could be tested more than once, simply because of the participants' freedom of choice and the limited time during the sessions, the results of this fourth study indicate that the scenarios constructed did lead to manifestations of teacher agency. we arrived at this conclusion because every scenario that was tested elicited multiple solutions. therefore, we can assume that the procedure for developing scenarios as described and applied here produces valuable and usable scenarios by which to measure indications of teacher agency. kusters et al 16 | f l r discussion this paper focuses on the main research question: in what ways can representative scenarios be developed to measure teacher agency in universities? to answer this question, we described four studies, each designed to help with an aspect of developing valid scenarios in a trustworthy way (see figure 1). our investigation was aimed at developing and evaluating scenarios designed to elicit teacher agency in the context of university teaching practice. in the course of the four studies we developed 23 scenarios, suitable for eliciting multiple affordances that seem to capture the nature of teacher agency better than other traditional measurement tools. reviewing prior research made it particularly clear that measurement tools mostly deliver generic statements about agency and cannot adequately incorporate context. moreover, in questionnaire studies there is little or no room for reflection on and consideration of particular answers. for example, vähäsantanen et al. (2019) note that it is important to consider how professional agency dimensions may vary over time and in different contexts. our project builds on this observation by creating context-specific and qualitative scenarios that allow a deeper exploration of these varied contexts, as suggested by skilling and stylianides (2020). leijen et al. (2021) developed a robust conceptual framework for their instrument, but the generic items of their survey raise the question of whether more specific formulations would yield similar results, as responses to the items may depend on the context considered by the participant. we addressed this issue by developing our instrument for specific situations based on real-life experiences (hughes & huby, 2004). to better understand the context of the responses given, and so better measure the agency of teacher educators (a subset of university lecturers), ghiasvand et al. (2023) suggest using qualitative tools such as focus group interviews and reflective journal writing. our research incorporates this qualitative approach and provides contextual frames in the scenarios. in a vignette study by louws et al. (2020), the authors note uncertainty about the influence of situational characteristics on school leaders’ choice of leadership instruments. the scenarios we developed allow us to explore nuanced relationships between context and decision-making (cf. skilling & stylianides, 2020), providing a more comprehensive perspective on agency than offered in previous studies. considering all this, we argue that the measurement tools currently available, both qualitative and quantitative, cannot fully capture the dynamic nature of teacher agency due to the lack of real-life context dependencies. studies on teacher agency can only provide relevant findings if the teaching-specific context is explicitly embedded in the design of the study. only by embedding context can researchers develop a deep understanding of the reasons why teachers do or do not demonstrate agency in a meaningful situation (engeström, 2011). the set of scenarios developed in this project could be utilized as real-life stimuli, as demonstrated in intervention studies aimed at measuring teacher agency (cf. yang, 2021). this means that the set of 23 scenarios can indeed be used in follow-up studies. our project contributes to the existing literature by introducing real-life teaching contexts as a metaphorical kaleidoscope for research on teacher agency in universities. reflections on the inductive and deductive research processes it was essential to develop qualitatively strong and internally valid scenarios to be used for follow-up research (hughes & huby, 2004). although developing scenarios for research is not new, it is often only the procedural and practical aspects of using scenarios that are described. however, in this way the theoretical frameworks underlying the phenomena being studied, and the research paradigm related to the content of the scenario material, are neglected (hughes & huby, 2004; skilling & stylianides, 2020). therefore, as promised in the introduction, we will now reflect briefly on the way our four studies contributed to trustworthy scenarios that meet guba’s (1981) criteria. during the first three studies, constituting the inductive phase, we explicitly focused on increasing credibility, dependability, and confirmability (guba, 1981) of our research. credibility (the degree of kusters et al 17 | f l r believability of the findings) was enhanced by employing a multimethod approach to gather data from various sources, such as interviews and expert panels. by adopting this approach, we aimed to represent lecturers' experiences and views in a trustworthy way. study 1 stands out as particularly solid because scenarios were derived directly from real-life situations, which reinforced the authenticity of our results. moreover, this credibility is emphasized in study 4, demonstrating the usability of the tested scenarios. regarding high dependability (the consistency and stability of the research process and findings over time), expert panels provided input on the design and methodology. in addition, regular consultation between the authors helped maintain consistency in the interpretation of data, and the application of research methods ensured a consistent and trustworthy research process. these steps led to our scenarios being designed in such a way that each respondent received similar information, which promoted consistency and stability in the presentation of situations and ensured that participants considered similar contexts when responding to the scenarios. to ensure confirmability (the objectivity and neutrality of the research findings), multiple researchers were involved in the analysis and interpretation of the data during the inductive phase. this collaborative approach helped minimize individual biases and promoted an objective understanding of the findings. finally, the pilot (study 4) was intended to ensure increasing transformability (the extent to which research findings can be applied). ensuring this factor involved examining how broadly the research conclusions could be used in or adapted to different situations or people. because the pilot was deductive in nature this allowed us to gain insight into the generalizability of the conclusions in order to assess usability among participants. limitations caution should be taken in the interpretation of our findings because of two limitations. the first is that we used a convenience sample in study 4, which may limit the external generalizability of our findings. the sample may not adequately represent the population because it consisted of lecturers who voluntarily chose to participate and are actively involved in educational innovation and improvement. this sampling may have biased our findings, as not all university lecturers are similarly involved in teaching. for example, we used committed lecturers who are highly involved in educational innovation and, therefore, have been able to think about improving university teaching more and could come up with more solutions than less engaged lecturers. this selection may have led to an overestimation of the number of solutions. second, the follow-up questioning in study 4 may have been a limitation. we used a strict think-aloud protocol, in which free association was vital to getting as accurate solutions to the scenarios as possible. at the same time, we also wanted to challenge lecturers to look beyond their initial ideas. therefore, we asked, “can you think of more solutions?” these types of follow-up questions may lead to more solutions but hinder free association. therefore, the extent to which the solutions proposed are still agentic needs to be elucidated. follow-up research should investigate whether the solutions lecturers come up with in response to a followup question are still relevant. implications and conclusions we tested the scenarios and found that they did serve their purpose for research on eliciting teacher agency. besides research purposes, the scenarios can be used for professional lecturer development programs for academic growth and in job interviews. for example, scenarios can be used in a card-sorting game-based activity during teacher training, as an individual reflection activity, as a start for a collegial consultation round, or as a collaborative activity to develop lecturers’ repertoire of solutions. reflecting on teaching scenarios together with colleagues or potential teaching staff creates a fertile environment in which experiences are shared, underlying values and beliefs come to the surface, insights are deepened, and collective knowledge is built. research consistently emphasizes the benefits of self-reflection on teaching behavior (van beveren et al., 2018). using the diverse set of scenarios as conversation starters – rather than as a measurement tool as we intended – can serve as a way to contribute to deep introspection and promote self-awareness and growth. conversations or reflections through scenarios can reveal relationships and structures in this way, which is important in developing agency (priestley et al., 2015; rushton & bird, 2023). kusters et al 18 | f l r moreover, these scenarios appear to fit ideally within the scholarship of teaching and learning (sotl) framework. sotl refers to a systematic approach built on reflection on and publication of the educational process, in order to improve the quality of education and contribute to the knowledge base of effective teaching practices (kreber & cranton, 2000). within sotl, the scenarios can be considered not only tools for personal reflection but also valuable sources for generating knowledge about effective teaching practices (gilpin & liston, 2009). for example, lecturers can use the scenarios to critically analyze their teaching practices, share insights with colleagues, and jointly explore new approaches to teaching. this process contributes to lecturers’ individual professional development and the broader community of educational professionals. to conclude, in this paper we have described how teacher agency can be measured and how to make manifestations of agency visible and accessible. the scenarios listed reveal considerations regarding actions, informed decisions, and the lecturers’ decision space. this contributes to faculty professionalization for the purpose of an engaged, innovative teaching staff within universities. empowering faculty members to reflect on their actions, and conduct peer reviews and supervision contributes to forming a faculty community focused on engagement, continuous development, and innovation in university teaching. kusters et al 19 | f l r references aspbury-miyanishi, e. (2022). the affordances beyond what one does: reconceptualizing teacher agency with heidegger and ecological psychology. teaching and teacher education, 113, 103662. https://doi.org/10.1016/j.tate.2022.103662 biesta, g., priestley, m., & robinson, s. (2015). the role of beliefs in teacher agency. teachers and teaching: theory and practice, 21(6), 624–640. https://doi.org/10.1080/13540602.2015.1044325 biesta, g., & tedder, m. (2007). agency and learning in the lifecourse: towards an ecological perspective. studies in the education of adults, 39(2), 132-149. https://doi.org/10.1080/02660830.2007.11661545 bradley, r. a., & terry, m. e. (1952). rank analysis of incomplete block designs: i. the method of paired comparisons. biometrika, 39(3/4), 324. https://doi.org/10.2307/2334029 comproved. (2022, september 5). assess quick and fair with the comparing tool comproved. https://comproved.com/en/comparing-tool/ cong-lem, n. (2021). teacher agency: a systematic review of international literature. issues in educational research, 31(3), 718-738. http://www.iier.org.au/iier31/cong-lem.pdf emirbayer, m., & mische, a. (1998). what is agency?. american journal of sociology, 103(4), 9621023. https://doi.org/10.1086/231294 engeström, y. (2011). from design experiments to formative interventions. theory & psychology, 21(5), 598-628. https://doi.org/10.1177/0959354311419252 engeström, y., kajamaa, a., & nummijoki, j. (2015). double stimulation in everyday work: critical encounters between home care workers and their elderly clients. learning, culture and social interaction, 4, 48-61. https://doi.org/10.1016/j.lcsi.2014.07.005 finch, j. (1987). the vignette technique in survey research. sociology, 21(1), 105–114. https://doi.org/10.1177/0038038587021001008 ghiasvand, f., jahanbakhsh, a. a., & sharifpour, p. (2023). designing and validating an assessment agency questionnaire for efl teachers: an ecological perspective. language testing in asia, 13(1). https://doi.org/10.1186/s40468-023-00255-z gibson, j. j. (1986). the ecological approach to visual perception. houghton mifflin. giddens, a. (1984). the constitution of society: outline of the theory of structuration. polity press. gilpin, l. s., & liston, d. d. (2009). transformative education in the scholarship of teaching and learning: an analysis of sotl literature. international journal for the scholarship of teaching and learning, 3(2). https://doi.org/10.20429/ijsotl.2009.030211 gould, d. (1996). using vignettes to collect data for nursing research studies: how valid are the findings? journal of clinical nursing, 5(4), 207–212. https://doi.org/10.1111/j.13652702.1996.tb00253.x guba, e. g. (1981). criteria for assessing the trustworthiness of naturalistic inquiries. educational technology research and development, 29(2). https://doi.org/10.1007/bf02766777 hiver, p., & whitehead, g. e. k. (2018). sites of struggle: classroom practice and the complex dynamic entanglement of language teacher agency and identity. system, 79, 70–80. https://doi.org/10.1016/j.system.2018.04.015 hughes, r., & huby, m. (2002). the application of vignettes in social and nursing research. journal of advanced nursing, 37(4), 382–386. https://doi.org/10.1046/j.13652648.2002.02100.x hughes, r., & huby, m. (2004). the construction and interpretation of vignettes in social research. social work and social sciences review, 11(1), 36-51. https://doi.org/10.1921/swssr.v11i1.428 https://doi.org/10.1016/j.tate.2022.103662 https://doi.org/10.1080/13540602.2015.1044325 https://doi.org/10.1080/02660830.2007.11661545 https://doi.org/10.2307/2334029 https://comproved.com/en/comparing-tool/ http://www.iier.org.au/iier31/cong-lem.pdf https://doi.org/10.1086/231294 https://doi.org/10.1177/0959354311419252 https://doi.org/10.1016/j.lcsi.2014.07.005 https://doi.org/10.1177/0038038587021001008 https://doi.org/10.1186/s40468-023-00255-z https://doi.org/10.20429/ijsotl.2009.030211 https://doi.org/10.1111/j.1365-2702.1996.tb00253.x https://doi.org/10.1111/j.1365-2702.1996.tb00253.x https://doi.org/10.1007/bf02766777 https://doi.org/10.1016/j.system.2018.04.015 https://doi.org/10.1046/j.1365-2648.2002.02100.x https://doi.org/10.1046/j.1365-2648.2002.02100.x https://doi.org/10.1921/swssr.v11i1.428 kusters et al 20 | f l r jenkins, g. (2019). teacher agency: the effects of active and passive responses to curriculum change. the australian educational researcher, 47(1), 167–181. https://doi.org/10.1007/s13384-019-00334-2 jeffries, c., & maeder, d. w. (2005). using vignettes to build and assess teacher understanding of instructional strategies. the professional educator, 27, 17– 28. https://files.eric.ed.gov/fulltext/ej728478.pdf kreber, c., & cranton, p. (2000). exploring the scholarship of teaching. the journal of higher education, 71(4), 476–495. https://doi.org/10.1080/00221546.2000.11778846 kusters, m., van der rijst, r., de vetten, a., & admiraal, w. (2023). university lecturers as change agents: how do they perceive their professional agency?. teaching and teacher education, 127, 104097. https://doi.org/10.1016/j.tate.2023.104097 lave, j., & wenger, e. (1991). situated learning: legitimate peripheral participation. cambridge university press. https://doi.org/10.1017/cbo9780511815355 leijen, ä., pedaste, m., & baucal, a. (2021). assessing student teachers’ agency and using it for predicting commitment to teaching. european journal of teacher education, 45(5), 600–616. https://doi.org/10.1080/02619768.2021.1889507 leijen, ä., pedaste, m., & lepp, l. (2019). teacher agency following the ecological model: how it is achieved and how it could be strengthened by different types of reflection. british journal of educational studies, 68(3), 295–310. https://doi.org/10.1080/00071005.2019.1672855 louws, m., zwart, r., zuiker, i., meijer, p. c., oolbekkink-marchand, h., schaap, h., & van der want, a. (2020). exploring school leaders’ dilemmas in response to tensions related to teacher professional agency. professional development in education, 46(4), 691–710. https://doi.org/10.1080/19415257.2020.1787203 luce r. d. (1959). individual choice behavior: a theoretical analysis. new york, ny: john wiley. pollitt a. (2012). comparative judgement for assessment. international journal of technology and design education, 22, 157-170. http://doi.org/10.1007/s10798-011-9189-x priestley, m., biesta, g., & robinson, s. (2015). teacher agency: an ecological approach. https://doi.org/10.5040/9781474219426 priestley, m., edwards, r., priestley, a., & miller, k. (2012). teacher agency in curriculum making: agents of change and spaces for manoeuvre. curriculum inquiry, 42(2), 191–214. https://doi.org/10.1111/j.1467-873x.2012.00588.x pyhältö, k., pietarinen, j., & soini, t. (2013). comprehensive school teachers’ professional agency in large-scale educational change. journal of educational change, 15(3), 303–325. https://doi.org/10.1007/s10833-013-9215-8 romano, m. e. (2006). “bumpy moments” in teaching: reflections from practicing teachers. teaching and teacher education, 22(8), 973–985. https://doi.org/10.1016/j.tate.2006.04.019 rushton, e. a., & bird, a. (2023). space as a lens for teacher agency: a case study of three beginning teachers in england, uk. the curriculum journal. https://doi.org/10.1002/curj.224 sannino, a., & engeström, y. (2016). relational agency, double stimulation and the object of activity: an intervention study in a primary school. in working relationally in and across practices: cultural-historical approaches to collaboration (pp. 58-77). cambridge university press. https://doi.org/10.1017/9781316275184.004 silverman, d. (2006). interpreting qualitative data: methods for analyzing talk, text and interaction. sage. simon, m., & tierney, r. (2011). use of vignettes in educational research on sensitive teaching functions such as assessment. in international congress for school effectiveness and https://doi.org/10.1007/s13384-019-00334-2 https://files.eric.ed.gov/fulltext/ej728478.pdf https://doi.org/10.1080/00221546.2000.11778846 https://doi.org/10.1016/j.tate.2023.104097 https://doi.org/10.1017/cbo9780511815355 https://doi.org/10.1080/02619768.2021.1889507 https://doi.org/10.1080/00071005.2019.1672855 https://doi.org/10.1080/19415257.2020.1787203 http://doi.org/10.1007/s10798-011-9189-x https://doi.org/10.5040/9781474219426 https://doi.org/10.1111/j.1467-873x.2012.00588.x https://doi.org/10.1007/s10833-013-9215-8 https://doi.org/10.1016/j.tate.2006.04.019 https://doi.org/10.1002/curj.224 https://doi.org/10.1017/9781316275184.004 kusters et al 21 | f l r improvement (pp. 1e13) (limassol, cyprus). http://www.icsei.net/icsei2011/full%20papers/0153.pdf. skilling, k., & stylianides, g. j. (2020). using vignettes in educational research: a framework for vignette construction. international journal of research & method in education, 43(5), 541– 556. https://doi.org/10.1080/1743727x.2019.1704243 stravakou, p., & lozgka, e. (2018). vignettes in qualitative educational research: investigating greek school principals’ values. the qualitative report, 23(5), 1188–1207. https://doi.org/10.46743/2160-3715/2018.3358 vähäsantanen, k., paloniemi, s., räikkönen, e., & hökkä, p. (2020). professional agency in a university context: academic freedom and fetters. teaching and teacher education, 89, 103000. https://doi.org/10.1016/j.tate.2019.103000 van beveren, l., roets, g., buysse, a., & rutten, k. (2018). we all reflect, but why? a systematic review of the purposes of reflection in higher education in social and behavioral sciences. educational research review, 24, 1-9. https://doi.org/10.1016/j.edurev.2018.01.002 van kan, c. a., ponte, p., & verloop, n. (2010). how to conduct research on the inherent moral significance of teaching: a phenomenological elaboration of the standard repertory grid application. teaching and teacher education, 26(8), 1553–1562. https://doi.org/10.1016/j.tate.2010.06.007 verhavert, s., bouwer, r., donche, v., & de maeyer, s. (2019). a meta-analysis on the reliability of comparative judgement. assessment in education: principles, policy & practice, 26(5), 541–562. https://doi.org/10.1080/0969594x.2019.1602027 verhavert, s., de maeyer, s., donche, v., & coertjens, l. (2018). scale separation reliability: what does it mean in the context of comparative judgment? applied psychological measurement, 42(6), 428–445. https://doi.org/10.1177/0146621617748321 yang, h. (2021). epistemic agency, a double-stimulation, and video-based learning: a formative intervention study in language teacher education. system, 96, 102401. https://doi.org/10.1016/j.system.2020.102401 yang, h. (2022). 4 developing languages pre-service teachers’ epistemic agency in using technology in languages teaching. in d. banegas, e. edwards & l. villacañas de castro (ed.), professional development through teacher research: stories from language teacher educators (pp. 51-71). bristol, blue ridge summit: multilingual matters. https://doi.org/10.21832/9781788927727-007 https://doi.org/10.1080/1743727x.2019.1704243 https://doi.org/10.46743/2160-3715/2018.3358 https://doi.org/10.1016/j.tate.2019.103000 https://doi.org/10.1016/j.edurev.2018.01.002 https://doi.org/10.1016/j.tate.2010.06.007 https://doi.org/10.1080/0969594x.2019.1602027 https://doi.org/10.1177/0146621617748321 https://doi.org/10.1016/j.system.2020.102401 https://doi.org/10.21832/9781788927727-007 kusters et al 22 | f l r appendices appendix a. all codes and how often they are applied. code applied 1 dealing with rules and requirements of university 7 2 difference in preknowledge 7 3 insufficient preparation of students 6 4 passive unmotivated students 5 5 dealing with conflicts with colleagues and management 4 6 no suitable teaching space or organizational preconditions 4 7 dealing with complaints and unexpected questions from students 4 8 dealing with workload 4 9 too little time for teaching preparation 3 10 technical problems 3 11 managing student expectations 3 12 administration 2 13 student hurt by joke 2 14 being thrown in at the deep end 2 15 dealing with controversial topics in college 2 16 dealing with individual (problems) student 2 17 not enough seats for students 2 18 too little time for educational innovation 2 19 much freedom per individual but little coherence between courses 2 20 implement changes in course 2 21 difference in teaching drive between colleagues 2 22 balancing workload and offer personal guidance 2 23 didactic choices 2 24 change of job position 1 kusters et al 23 | f l r appendix b. four different forms of one scenario. firstperson thirdperson open-ended title: unprepared students in lecture the third lecture of my course was about to begin. as usual, students had to read an article to prepare for the lecture. during the lecture, there came little response to my questions. only the few students who always actively participate made any attempt to answer my questions. when i tried to involve others in the discussion, i didn’t succeed. i asked who had read the article i had sent in advance, but it turned out that only a handful of students had prepared as instructed. i realized that most students had not prepared the lecture, so i knew i had to come up with a solution. title: unprepared students in lecture the third lecture of dr. johnson’s course was about to begin. as usual, students had to read an article to prepare for the lecture. during the lecture, there came little response to dr. johnson’s questions. only the few students who always actively participate made any attempt to answer dr. johnson’s questions. when dr. johnson tried to involve others in the discussion, he didn’t succeed. dr. johnson asked who had read the article he had sent in advance, but it turned out that only a handful of students had prepared as instructed. dr. johnson realized that most students had not prepared the lecture, so he knew he had to come up with a solution. completed title: unprepared students in lecture the third lecture of my course was about to begin. as usual, students had to read an article to prepare for the lecture. during the lecture, there came little response to my questions. only the few students who always actively participate made any attempt to answer my questions. when i tried to involve others in the discussion, i didn’t succeed. i asked who had read the article i had sent in advance, but it turned out that only a handful of students had prepared as instructed. i realized that most students had not prepared the lecture and decided to use the situation as a learning opportunity by suggesting discussion groups to encourage future preparation and simplifying the content to re-engage the entire class. title: unprepared students in lecture the third lecture of dr. johnson’s course was about to begin. as usual, students had to read an article to prepare for the lecture. during the lecture, there came little response to dr. johnson’s questions. only the few students who always actively participate made any attempt to answer dr. johnson’s questions. when dr. johnson tried to involve others in the discussion, he didn’t succeed. dr. johnson asked who had read the article he had sent in advance, but it turned out that only a handful of students had prepared as instructed. dr. johnson realized that most students had not prepared the lecture and decided to use the situation as a learning opportunity by suggesting discussion groups to encourage future preparation and simplifying the content to re-engage the entire class. kusters et al 24 | f l r appendix c. the final 23 scenarios. title: students’ personal problems i’m a lecturer at this university. during one of my courses, i noticed that one of my students had been struggling in class for some time. the student was not meeting the deadlines, seemed unmotivated, and hardly interacted with other students. i had a conversation with the student, and she told me that things were going badly at home and that was why she couldn’t keep her attention in class. i knew i had to come up with a solution. title: workload as a lecturer, i experience constant pressure to teach, conduct research, publish articles, and attend conferences. it feels like a constant struggle to get everything done within the tight deadlines set by the university. i am aware that this pressure is a result of both the high standards set for university teachers and our own passion for the job. it feels like there are never enough hours in a day, and at times i feel overwhelmed. i know i have to come up with a solution. title: coherence between courses one of the things i appreciate most about teaching at the university is the freedom i have to develop my own courses and shape my teaching in the way i feel best. i take pride in developing courses that can inspire and challenge my students, and i enjoy creating a unique learning experience for them. at the same time, sometimes it seems my colleagues and i are all working on our own “island.” it seems that this individual freedom comes at the cost of a lack of coherence and consistency across the curriculum. so i knew i had to come up with a solution. title: technical issues well in time, i arrived in class to get everything ready for my lecture. however, the computer was very slow in starting up and eventually crashed completely. while my students entered the class, i tried to restart the computer. my heart rate accelerated as i watched the clock tick away and realized how much time we were losing. i realized how dependent we are on technology nowadays! i knew i had to come up with a solution. title: student expectations as a lecturer, i notice that students increasingly expect individual feedback and guidance. students’ expectations seem to have changed since my own student days, and this puts additional pressure on me as a lecturer. i appreciate that students value feedback and guidance. i understand that this is a crucial part of their learning process, and i do my best to meet their expectations. but one day, i felt completely overwhelmed by the large number of emails i received from students who asked for feedback or a one-on-one meeting. so i knew i had to come up with a solution. title: students differ in prior knowledge i walked into the classroom, ready to start my subject’s introductory lecture. it was the first lecture of the new academic year. first, i did a small recap of the basics, assuming these were still familiar to the students, and then introduced some new definitions and concepts. however, as i looked around the room, i noticed that many of my students struggled to keep up with the content. some of them just stared blankly at their notes; others flipped desperately through their textbooks; only a few students were able to keep up. to verify whether students indeed struggled to keep up, i asked which of the students were familiar with the basic concepts. it turned out that there was a large variety with regard to their prior knowledge. for some, the new concepts were easy to understand, but for others, even kusters et al 25 | f l r the basic information was completely new. when i realized these differences in students’ prior knowledge, i knew i had to come up with a solution. title: unprepared students in lecture the third lecture of my course was about to begin. as usual, students had to read an article to prepare for the lecture. during the lecture, there was little response to my questions. only the few students who always actively participated made any attempt to answer my questions. when i tried to involve others in the discussion, i didn’t succeed. i asked who had read the article i had sent in advance, but it turned out that only a handful of students had prepared as instructed. i realized that most students had not prepared for the lecture, so i knew i had to come up with a solution. title: students’ views on teaching quality of colleague as a lecturer, i see it as my responsibility to put time and energy into the preparation of my classes and constantly look for ways to improve and innovate my teaching. however, i am told by students from other groups that their instructor is often poorly prepared, cannot provide appropriate answers to questions, and has already failed to meet the promised review deadline a couple of times. i hesitate to bring this up with my colleague: on the one hand, i do not think it is my responsibility, but on the other hand, my professionalism tells me that students have the right to a quality education. so i know i have to come up with a solution. title: educational re-design as a lecturer, i’m ambitious to keep improving my teaching and engage my students more, but constraints from the curriculum and my department hinder what i can achieve. there is no real incentive to improve teaching, nor are proper facilities provided. still, i see it as part of my job to constantly look for ways to improve my teaching, so i know i have to come up with a solution. title: implementing changes in course i’m teaching a new course in which i want to experiment with some new teaching methods. however, the university requires me to send in the course description and grading procedures far in advance because of the long and tedious procedures for creating the study guide. since i am currently busy teaching other courses, i feel hampered in implementing these innovations. on the one hand, i see it as my duty to further develop my teaching; on the other hand, i feel that my current teaching also deserves full attention. i know i have to come up with a solution. title: students’ unexpected questions the students were in their seats, and i welcomed everyone. before i actually began the lecture, a student raised her hand and asked a question unrelated to the lecture’s topic. however, when i listened to the question, i found it an interesting question nonetheless, and, as i appreciated the idea that the student asked the question, i wanted to accommodate the student. i knew i had to come up with a solution. title: balancing professional responsibilities i value the personal mentoring of my students, as i believe personal attention contributes significantly to their success and well-being. i enjoy taking the time to have one-on-one meetings with my students and provide individual feedback, but i find that this becomes increasingly difficult as my student numbers increase. i find myself in a tight spot and have to make choices between my mentoring role and other duties, such as teaching and research. i struggle with this balance, so i know i have to come up with a solution. kusters et al 26 | f l r title: lecture preparation at the last moment, i had taken over a lecture from a colleague. when i started with the lecture, i soon realized that important slides were missing. i felt like i was thrown into the deep sea without a life jacket. i hadn’t had enough time to prepare the lecture because my colleague hadn’t saved all the slides. this frustrated me because i knew how important it was to provide students with a wellstructured and organized lecture. i knew i had to come up with a solution. title: view on teaching it was the first time my colleague and i taught a particular current course. to prepare for the course, we divided the topics among ourselves. beforehand, i was very excited to teach the course with my colleague. however, i discovered that we held completely different views on what good teaching entails. conflicts arose over things like whether class attendance would be compulsory, the amount of feedback we would give, the method of grading, etc. these were long and exhausting debates; discussions ran high, and we struggled to understand each other. i knew i had to come up with a solution. title: interaction in online teaching i stared at my computer screen in frustration as i tried to lead an online discussion with my students. i found it hard to feel the same energy and connection as i did in the physical classroom. i missed the spontaneous conversations, body language, and in-person conversations with my students. i felt isolated and uncomfortable in this new environment. still, i didn’t want to give up because i felt i should be able to teach my classes properly even in this situation. so, i knew i had to come up with a solution. title: administrative tasks in thesis supervision as a lecturer, i find the administrative hassle surrounding thesis supervision particularly timeconsuming and frustrating. it feels like it never ends. i have to fill out all kinds of forms, keep track of deadlines, prepare reports, answer countless emails, and attend meetings. it seems more like an administrative job than supervising students. i would like to spend more time giving feedback and guidance to students instead of being stuck in a bureaucratic system. it is time for a more efficient way of working. i know i have to come up with a solution. title: controversial topics in course i taught a course that included some highly contested and controversial topics. i knew that these topics could lead to heated debates and even division among the students in class. although i felt i was usually well-able to lead class discussions about sensitive issues, it seemed to become increasingly difficult to maintain an atmosphere of respect and understanding in class. what i feared did indeed happen: students felt attacked, and discussions got out of hand. i knew i had to come up with a solution. title: being overwhelmed as a new lecturer as a starting lecturer, i felt overwhelmed. the first period of teaching felt like a big pandemonium, full of challenges, like supervising students, preparing and delivering courses, and doing that in an inspiring way. although i understood that it is normal to experience these challenges and i still had to build routines, it also felt like i was thrown into the deep end, and i had no idea where to start or how to manage all of these tasks. i knew i had to come up with a solution. title: taking offense at joke kusters et al 27 | f l r a student came up to me after the lecture. he said he felt offended by a joke another student had made during class. i felt uncomfortable because i was not aware of this situation, but i knew it was important to create a safe learning environment where students feel safe and free to express themselves. i took the student’s concerns seriously because he was genuinely upset. i knew i had to come up with a solution. title: unmotivated students i have been teaching at this university for several years now and have encountered many difficult students, but i had never experienced a class like this one before. many students seemed uninterested in the material. some students were sleeping; others were looking at their phones or talking to each other. when i asked who was interested in the subject, only a few hands went up. when i realized that the subject did not interest students at all, i knew i had to come up with a solution. title: change of job position i have accepted the position of educational director and find it difficult to be the manager of former colleagues. i feel i am in an awkward position and don’t really know how to handle this situation. it occurs to me that my attitude toward my former co-workers has changed and that i am struggling to make choices that affect them. i want to find a way in my new role and find the right balance between collegiality and leadership. still, i notice from the side of colleagues and myself that there is friction because of my new role. i know i have to come up with a solution. title: pedagogical choices as a lecturer, i feel i need to continuously develop. this means that i have to use new pedagogical approaches. but i’m reluctant to change my pedagogical choices too drastically, since i am afraid that the students will not understand it and therefore will not perform as well. i still want to keep looking for ways to improve my teaching and challenge my students without disadvantaging them through poor teaching, but traditional ways of teaching feel more comfortable and may be safer. i know i have to come up with a solution. title: university rules i’m an enthusiastic and dedicated lecturer, but i feel restricted by the rules and requirements of the university where i work. i have to follow a strict protocol for deviating from an exam date. currently, a good student of mine is unable to take the exam due to personal circumstances. i want to accommodate her by offering another date, but because of all the rules of the university, she has to take the retake. neither of us wants that, so i know i have to come up with a solution. frontline learning research vol. 12 no. 2 (2024) 70 -98 issn 2295-3159 corresponding author: tuulikki ukkonen-mikkola, faculty of educational sciences university of helsinki, siltavuorenpenger 5, 00170 helsinki, finland, email address:tuulikki.ukkonen-mikkola@helsinki.fi doi: https://doi.org/10.14786/flr.v12i2.1153 with sensitive eyes: ecec teachers’ visual gaze and related reflections on pedagogical actions in toddler groups using eyetracking glasses tuulikki ukkonen-mikkola1,2, susanna isotalo1, saswati chaudhuri1, jenni salminen1, olli merjovaara1, carita lindén1 & niina rutanen1. 1 university of jyväskylä, finland 2 university of helsinki, finland article received 15 august 2022 / article revised 6 june 2024 / accepted 8 july 2024 / available online 22 july 2024 abstract this study explored early childhood education and care (ecec) teachers’ visual gaze and related reflections on pedagogical actions during pedagogical activities in groups of children under three years of age in finland. the data were collected from play and teacher-guided activities using mobile eye-tracking glasses, the retrospective thinking aloud (rta) method, and semi-structured interviews. the results showed that even though the teachers were surprised about some aspects of the visual gaze metrics, they reflected on and gave reasons for their visual gazes on children. when observing gaze data from play, teachers explained the high amount of gaze by citing children’s particular needs. when observing gaze data from guided activities, teachers reflected on children’s unpredictable behavior and noted that the children’s need for support in concentration was linked to more gazes by the teacher. the findings showed that both during play and guided activities, children seeking a gaze and the position of children in the classroom influenced the number of teachers’ gazes. in the teachers’ explanations of their visual gaze and related pedagogical actions, five categories were identified: protection; physical and emotional availability, teaching and learning; facilitation; and initiatives. this explorative study showed that teachers utilize their knowledge concerning children’s individuality, development, and learning when they explain their decisions concerning their visual gaze and pedagogical activities with toddlers. the use of mobile eye-tracking technology is relatively new; therefore, its applications to ecec are pioneering for the development of the field in relation to the practices and research of toddlers’ groups and groups with older children in ecec. keywords: toddlers, teachers’ gaze, eye-tracking glasses, visual gaze metrics, knowledge-based reasoning, teacher reflection and explanations mailto:tuulikki.ukkonen-mikkola@helsinki.fi ukkonen-mikkola et al 71 | f l r 1. introduction a child’s day in early childhood education and care (ecec) is characterized by diverse pedagogical activities (guedes et al., 2020). the teachers and their pedagogical actions in the ecec groups are central in guiding these activities based on their values, knowledge, and perceptions of the children (lunn browlee et al., 2016). thus, at the heart of these pedagogical activities is teacher–child interaction, which includes both verbal and nonverbal communication, such as body language, gestures, and visual gaze (bae, 2009; jamison et al., 2014; salminen et al., 2021b). currently, not much is known about ecec teachers’ visual gaze while interacting with children in authentic ecec groups. even though most of the eye-tracking research conducted in educational settings has not been directly related to ecec, it has provided a starting point for designing the present study. for this reason, we will briefly mention eye-tracking studies conducted in elementary and high school settings. elsewhere in educational research, teachers’ visual gaze has been explored using mobile eye-tracking technology. presently, mobile eye-tracking technology is increasingly being used to study teachers’ visual gaze behaviors in various classroom settings in relation to their level of work experience and their knowledge of teaching and learning (dessus et al., 2016; goldberg et al., 2021; huang et al., 2021; mcintyre & foulsham, 2018; shinoda et al., 2021). furthermore, muhonen et al. (2020) utilized mobile eye-tracking glasses to study primary and preschool teachers’ distribution of visual attention during highand lowquality educational dialogue. in the ecec field, ishibashi et al. (2020) and sadamatsu (2022) studied the relationship between ecec staff members’ professional experience and visual gaze using eyetracking glasses. moreover, isotalo et al.’s (2024) study focused on ecec teachers’ gaze behavior during various phases of pedagogical interaction with under three-year-old children. these studies have raised awareness of teachers’ visual gaze behaviors in school classrooms and, more recently, in ecec and on teachers’ professional vision in noticing and knowledge-based reasoning of the noticed situations (seidel & stürmer, 2014). this study aimed to fill the gap in knowledge that exists concerning teachers’ visual gaze in ecec contexts that are different from school classrooms. our research also included, for the first time that we are aware of, ecec teachers’ reflections and explanations of their gaze and related pedagogical actions on the basis of their own visual gaze data. in this study, we concentrated on ecec teachers’ visual gaze during pedagogical activities in groups of children under three years old in finland. we referred to our child participants as “toddlers,” a term used to refer to 18to 36-month-old children (colson & dworkin, 1997). our aim was, first, to investigate ecec teachers’ visual gaze on children and their related reflections based on visual gaze metrics. second, we explored teachers’ explanations based on their own visual gaze and related pedagogical actions with children in toddler groups. accordingly, teachers’ visual gaze was investigated using eye-tracking recordings as well as both gaze metrics and verbal data (see mcintyre et al., 2020). the data were collected using tobii pro glasses 2 eye-tracking glasses, the retrospective think-aloud (rta) method, and semi-structured interviews with teachers. during rta, teachers explained their actions based on what they looked at and described their reasoning in a review of their work and decision process in the classroom (häggström et al., 2015). the findings from the present study will contribute to raising awareness of using eye-tracking technology with teachers in ecec and evoke teachers’ reflections on their visual gaze and related pedagogical actions. 1.1. teacher´s pedagogical actions and interaction in toddler groups the changing situations with under three-year-olds in ecec groups require teacher’s pedagogical awareness and sensitivity; in other words, teacher’s “pedagogical glasses” are needed throughout the day. pedagogy in ecec can be interpreted as the actions, such as decisions and methods, of the teacher. pedagogy is based on the strategies the teacher uses to bring children’s competencies, skills, and ideas to the center of their educational actions (pramling samuelsson & asplund carlsson, 2008). the knowledge about learning, curriculum, theoretical basis of educational science, and children’s development and learning guide teachers’ pedagogical actions and decisions (helavaara robertson et al., 2015; ranta et al., 2023). thus, pedagogical actions with children include planning but ukkonen-mikkola et al 72 | f l r also teachers’ capacity for evaluation and reflection of the practices (kangas et al., 2021; pihlaja & holst, 2013). supportive and high-quality teacher–child interactions are known to promote children’s emotional, social, and cognitive development (organisation for economic co-operation and development, 2018; slot et al., 2018). teachers in ecec can stimulate children’s thinking, reasoning, and language skills by asking questions and giving feedback (guedes et al., 2020; la paro et al., 2012; thomason & la paro, 2009), but they also support socio-emotional skills, such as children’s selfregulation (salminen et al., 2021a). in addition to teachers’ high sensitivity toward children’s initiations and expressions, teachers’ ability to effectively monitor a group of children is measured as an indicator of high quality. studies on preschool and toddlers’ groups have revealed that interactions, and the quality of interactions between teachers and children vary across different activities during a day (booren et al., 2012; guedes et al., 2020; whittaker et al., 2018). for example, a portuguese study showed that when teachers were engaged in several practical demands, such as serving food during meals, interactions were less warm and sensitive than during free play and early academic skills (guedes et al., 2020). based on this previous knowledge of the diversities in teacher‒child interaction in activities during the day in ecec, in this study, we chose to explore two different activities: play and guided activities. play is a key activity for young children and a driving force for their development (vygotsky, 1967). teachers’ role in play involves observing children’s cues, aligning with their initiations, and being emotionally available to support children’s contributions to play (hakkarainen et al., 2013). in the present study, the play situations consisted of the teacher’s choice of timing and toys as a structure, but the children were free to play using their own ideas and creativity in the presence of the teacher. we also chose guided activities that were more teacher-directed. in general, the teacher sets the goals, plans and leads the guided activities in ecec to support the development of early academic skills, such as science, language, or math-oriented activities (ranta et al., 2023). these activities are often implemented creatively through music, art, and physical education. teacher–child interactional processes include nonverbal and verbal communication. in terms of nonverbal communication, body language, gestures, and visual gaze (bae, 2009; jamison et al., 2014), as well as touch, play a meaningful role, particularly with toddlers (johansson et al., 2021). the teacher’s physical presence as well as lap (availability of lap, having children on the lap) has gained attention in research (hännikäinen, 2015; lucas revilla et al., 2022). touch is also used intentionally, as teachers believe that it benefits children’s emotional development, is useful when building relations with the children, and supports learning (johansson et al., 2021). it is also used to control the child. in addition, establishing eye contact with children is one way for teachers to encourage desired behaviors and establish strong relationships (hietanen et al., 2008; ledbury et al., 2004). there is also a growing amount of research underlining the active role and capacity of infants and toddlers to be active participants in interactions as they communicate through smiles, cries, and other emotional strategies (salomon et al., 2017). teachers’ orientation toward these cues provides one important way for the child to be understood and acknowledged (pursi, 2019), thus foregrounding teachers’ visual gaze in interactions. 1.2. teachers’ visual gaze in educational research in the teaching profession, it is important for teachers to constantly adjust their visual gaze based on multiple events that occur simultaneously in the classroom. previous research in the school context showed that teachers face immediacy, unpredictability, multidimensionality, and simultaneity in complex classroom environments (doyle, 1980). teachers’ visual gaze behavior has been popularly studied under the broader concept of teacher professional vision, which has been investigated based on teachers’ ability to notice classroom events and provide knowledge-based reasoning for the noticed classroom events (seidel & stürmer, 2014). knowledge-based reasoning means teachers’ ability to utilize their knowledge to explain the situations they notice. previous research has shown that teachers’ knowledge-based reasoning consists of the three qualitative domains: description, explanation, and prediction (seidel et al., 2011; seidel & stürmer, 2014; sherin & van es, 2009). explanation refers to ukkonen-mikkola et al 73 | f l r the ability of teachers to use what they know to reason the situation (seidel & stürmer, 2014). in this study, we examined teachers’ explanations to determine what kind of knowledge-based reasoning they used when explaining their gaze and pedagogical actions during play and guided activities in toddler groups. additionally, we gained information about ecec teachers’ noticing through their gaze metrics. teachers explained classroom events through rta and interviews, which were implemented while watching eye-tracking videos, and they had the opportunity to link the noticed situation to broader knowledge of ecec pedagogy and expertise. in ecec groups, teachers use their visual gaze to notice children in their nonverbal interactions with them. in addition, teachers need to provide selective visual attention toward children to monitor, guide, and interact during pedagogical activities. in terms of the present study, visual gaze was defined as the total amount of time in terms of fixation durations teachers looked at the children and the number of visual fixations that occurred when teachers looked at the children (as used in previous studies of chaudhuri et al., 2022a, 2022b; isotalo et al., 2024). accordingly, teachers need to use their visual gaze toward children in ecec during both play and guided activities to observe children’s social, cognitive, and motor skills’ development and ensure their safety (kangas et al., 2021). an earlier study in ecec with eye-tracking glasses showed that experienced ecec staff looked less often at children’s faces than at other areas during snack time (ishibashi et al., 2020). moreover, the experienced ecec teachers gazed at children more frequently but for shorter times when monitoring play situations (sadamatsu, 2022). in a study by isotalo et al. (2024), it was shown that the phase of the pedagogical interaction and structure of the activity have an influence on ecec teachers’ gaze behavior by increasing visual gaze toward children in child-initiated interaction and on materials during teacher-initiated interaction. moreover, during structured music activities, teachers’ gaze was usually more on children, but during play, it was more on materials. another eye-tracking study in secondary school classrooms showed that teachers initiated eye contact more with students when they gave directions and that students initiated eye contact more when the teacher showed affection (haataja et al., 2020). previous research has also shown that teachers may show dominance by increasing eye contact with students during questioning and show friendliness toward students by increasing eye contact while lecturing (mcintyre et al., 2020). in the application of a theoretical model called “classroom management scripts,” it was indicated that teachers use their visual gaze to improve their situational awareness of classroom events that occur during lessons to observe, detect, and trace classroom events in relation to children’s learning and development (wolff et al., 2020). previous studies in schools have shown that teachers will gaze longer at children who face challenges in concentrating on tasks (seidel et al., 2021). children may also seek attention from teachers through their behaviors. in addition, students showing more interactive and disruptive behavior have been shown to receive more attention from teachers in terms of visual gazes (goldberg et al., 2021). since little is known about teachers’ visual gaze behavior in ecec settings, the present study examined teachers’ visual gaze in terms of their visual fixation metrics, their reflections on this, and their explanations based on their own visual gaze and related pedagogical actions in toddler groups. 1.3 aims of the study and research questions the aim of this study was to investigate teachers’ visual gaze and their pedagogical actions during pedagogical activities in ecec groups with children under three years old. in ecec, children are exposed to different interactional experiences during the day, which led us to select two different pedagogical activities (play and guided activity) where the gaze data were recorded. accordingly, the present study examined and explored teachers’ visual gaze on children, their reflections on the fixation gaze metrics, and their explanations based on their own visual gaze and related pedagogical actions in toddler groups. we addressed the following research questions: 1. what do teachers’ visual gaze metrics inform us about their visual gaze on children in play and guided activities, and how do teachers reflect on their visual gaze metrics? ukkonen-mikkola et al 74 | f l r 2. how do teachers explain their visual gaze on children and their pedagogical actions during play and guided activities in ecec groups with toddlers? 2. methodology 2.1 approach and the design of the study in the present study, a case study design has been implemented to obtain an in-depth understanding of three ecec teachers’ visual gaze while conducting pedagogical activities in the classroom in the form of play and guided activities. the true essence of a case study lies in detailed investigation of a small sample focusing on specific aspects of a teacher’s experience within a classroom (tight, 2010). accordingly, while exploring an individual ecec teacher’s case, firstly, a positivist approach was taken to investigate the quantifiable measure of their visual gaze using eye-tracking technology and secondly, an interpretive approach was taken wherein each teachers’ subjective reasoning/explanation of their visual gaze was examined. both the positivist and interpretive approaches provided a wholesome understanding of the way teachers support children in ecec. however, it is important to note that the case study approach is merely descriptive (flyvbjerg, 2006) and should not be used in making strong generalizations and conclusions on teaching and learning in ecec groups in general. 2.2 ethical issues this project followed the ethical guidelines of the national advisory board on research ethics in finland to maintain good scientific practice and conform to the applicable regulations for personal data use in scientific studies (finnish advisory board on research integrity, 2012). the research proposal received approval from the ethics committee of the university of jyväskylä, where particular attention was paid to the ethical assessment of the methodological approach and eye-tracking technologies used with young children. it was made clear to the study participants that their participation was completely voluntary. permission was sought from the municipality, participating teachers and guardians of the children who gave their written informed consent to participate in the study. as the children’s assent was negotiated throughout the process, their assent was observed by remaining sensitive to their nonverbal and verbal communication during the data collection (rutanen et al., 2023). the anonymity of the participants was ensured, and pseudonyms were used in reporting. particular care was paid during the reporting phase to eliminate any details that would make the cases identifiable. 2.3 piloting recordings and orienting the children and teachers to eye-tracking glasses utilizing mobile eye-tracking glasses to collect data of the teachers' visual gaze in a toddlers’ group is a fairly new research design; thus, we designed the research project one step at a time. the research team had previously established relationships with one finnish ecec center, the teachers, and the director of the ecec unit. before the data collection, the eye-tracking glasses were piloted in a preschool classroom (6-year-olds). many technical details (such as calibration, connecting the equipment, and recording the video) and requirements concerning the circumstances were learned during piloting, such as that the curtains should be closed to prevent direct sunlight from distorting the recorded video to guarantee good quality of the data. before actual data collection, two researchers visited the selected toddler groups and introduced eye tracking glasses to the groups’ teachers and children. the children were very curious about the glasses, and the research team answered the children’s questions and showed the eye-tracking glasses to them. to get participants (teachers and children) familiar with the glasses and to minimize possible uncomfortable feelings regarding data collection and glasses, the teachers got the chance to test eyeukkonen-mikkola et al 75 | f l r tracking glasses in the presence of participating children before the actual data collection. finally, researchers asked permission from children to come after a few days to their group with the “funny glasses.” 2.4 research participants, data collection and data the case study teachers, rose, joanna, and maria were from the same ecec center and conducted their pedagogical activities in their own groups of toddlers, while wearing eye-tracking glasses. the participating teachers had completed their bachelor’s degree in early childhood education at a finnish university. there was variation in the working experience (from one year and three months to four years) among the participating teachers; the average work experience was two and a half years. the age of the children who participated in this study varied from 29 to 36 months, with an average of 31 months. in one small group, there were from two to four children at a time (table 1) and these small groups were formed by the teacher. the data collection was done in two phases in fall 2021. in the first phase, each teacher was asked to implement one play activity and one guided activity with small groups of children. during eyetracking video recordings, teachers were given the freedom to decide what these activities would contain and in what situation they would be implemented. in addition, teachers chose the toys that were given to children to play with during play activities and children were free to engage in play with these toys as and if they wished. it is possible that teachers’ reflections and explanations can overlap between play and guided activities as these activities may not be often different from one another. more specific information about the content of the guided and play activities can be found in table 1. tobii pro glasses 2 were used to record eye-tracking videos of the participating teachers during these activities. tobii pro glasses enable capturing participants’ eye-movement data in authentic environments since, the glasses are mobile and allow participant to move their head and body freely. additionally, these eye trackers enable multimodal data analysis by capturing not only the eye movement data, but also video and audio data of participants’ actions. before the recording of videos, two trained researchers calibrated the eyetracking glasses using one-point calibration. the researchers asked the teacher to look at three set points in the room (such as the door, shelf, curtain, etc.) at the beginning of the video recording to verify that their gaze met the three determined points (tobii ab, 2018). the researchers confirmed that teachers felt comfortable while wearing eye-tracking glasses during the data collection. to our surprise, the children paid little attention to the glasses and did not try to touch or reach for them. the length of the eye-tracking video recordings varied from 12 to 32 minutes; the video data totaled 127 minutes, 32 seconds. the length of the individual recording depended on the length of the activity. during the recording of the shortest, a 12-minute recording, technical difficulties affected the duration of the video. the gaze sample percentage indicates the total percentage of time when at least one or both eyes were detected during the eye-tracking recording duration. in this study, as shown in table 1, all eye-tracking video recordings above the 70% gaze sample percentage were considered (chaudhuri et al., 2022a, 2022b, 2024). the teachers, children, activities, duration of videos, and the quality of the gaze sample are described in table 1. the eye tracking video data was imported to tobii pro lab analysis software for further analysis. prior to the analyses, each eye-tracking video recording was coded using a coding criterion. in tobii pro lab analysis software, it is advised to select a suitable filter which is best for filtering eye movements from participants whose visual gaze movements were recorded from real-world settings wherein they moved around freely. accordingly, for our eye-tracking video recordings, the i-vt (attention) filter settings were best suited to filter out fixations from teachers’ eye movement data. next, an experienced coder manually mapped fixations onto the areas of interest (aois) where the teacher looked. these areas of interest were pre-determined. after coding the eyetracking video recordings, fixation metrics were exported for further analyses of the case studies of 3 example teachers in our study. in the second phase of the data collection, the rta method and interviews were used wherein teachers watched their own eye-tracking video recordings and reflected on their actions. for the rta, teachers were given instructions orally and these instructions can be found from appendix c. due to ukkonen-mikkola et al 76 | f l r the limited number of previous studies using eye-tracking in ecec, our exploratory research design involved using both rta and interviews. in doing so, we were able to gain deeper insight from the teacher about the teaching situations, their experiences and reflections on using the glasses. rta is a valuable method for identifying teacher’s pedagogical expertise and observing how professionals operate from moment to moment (fox et al., 2011). earlier studies have shown that eye-tracking recordings can be used as part of rta (mcintyre et al., 2019; muhonen et al., 2022). the rta was followed by a short semi-structured interview wherein teachers were asked questions in relation to their experience concerning the eye-tracking video recording. more specific information of the interview protocol can be found in appendix c. during the interviews, teachers were asked if they would like to add more reflections concerning the eye-tracking video recordings. additionally, bar graphs obtained from eye-tracking video analysis were used as stimuli to elicit responses from the teachers. both the rta and interviews were recorded using a voice recorder and recorded data totaled 222 minutes and 16 seconds. interview data was transcribed for further analysis ukkonen-mikkola et al 77 | f l r table 1 teachers, children, activities, duration of videos, and quality of gaze sample content of the activity activities in lesson duration of eyetracking video (in minutes) gaze sample percent teacher: rose children: alma, elisabeth, betty & jasmin play: lego play teacher and children are playing and discussing together. 16.18 94 % teacher: rose children: alma, elisabeth & jasmin guided situation: music teacher and children sing and make hand gestures. 15.31 93 % teacher: joanna children: vincent & adam play: train track play teacher and children are discussing, designing, and building train tracks together. 19.46 88 % teacher: joanna children: vincent & mathias guided situation: fingerpainting teacher was guiding the children’s painting. children were free to choose what they would like to paint. 31.45 91 % teacher: maria children: oliver & emil play: magic sand play teacher gave examples and discussed with children how to mold sand into different shapes. 32.02 84 % teacher: maria children: oliver & emil guided situation: music teacher was using song cards. the teacher and children were singing and playing together. 12.10 74 % ukkonen-mikkola et al 78 | f l r 2.5 analysis in the present study, both the teachers’ eye gaze and verbal data were utilized (see mcintyre et al., 2019). teachers’ visual gaze and gaze metrics were indicated by eye-tracking metrics in terms of fixations, and the verbal data from rtas were used to explain and justify the teachers’ visual gaze and their pedagogical actions in the toddler group. the eye-tracking video recordings were analyzed using tobii pro lab v. 1.171 (tobii ab, danderyd, sweden). there were two eye-tracking video recordings from each teacher wherein one was showing a play activity, and the other was showing a guided activity led by the teacher. the data analysis was done in several steps, and both videos were analyzed using the same steps. in the first step, areas of interest (aois), which are targets in the surroundings on which the teacher focused during play and guided activities, were identified from the eye-tracking video recordings (holmqvist et al., 2015). in the present study, the aois were targets (such as children, and teaching materials) toward which the teacher focused their visual attention during pedagogical activities (as seen previously in chaudhuri et al., 2022a, 2022b; muhonen et al., 2022). in eye-tracking, fixations are defined as the duration of time when eyes remain relatively still and input of new information takes place from the immediate environment by selectively focusing on targets (rayner et al., 2009). the second step involved manually mapping the teachers’ eye gaze in terms of fixations onto the specific aois (set as stationary pictures in the tobii pro lab), as shown by a red circle on the eye-tracking video recording. in the third step, after manual mapping of teachers’ eye-gazes on the respective aois, the eye-tracking metrics related to teachers’ fixations were obtained in total duration of fixations, average duration of fixations, and number of fixations from the software. the fourth step involved developing visual representations in the form of bar graphs from the teachers’ eye-tracking metrics, such as total fixation duration, average fixation duration, and number of fixations using microsoft excel (microsoft, redmond, wa, usa). in this study, we utilized teachers’ gazes in terms of fixations toward children because we were interested in the amount of time teachers gaze at children during teacher-child interactions. the other gazes, for example, toward toys, walls, or furniture were left out of this study. inter-coder reliability was checked by double coding 20% of the videos from the entire dataset. double coding agreement ranged from 75% to 99.8%, with an average of 91.8%. the rta data and interviews were analyzed using a qualitative approach. the transcribed data were reduced, categorized, and classified (holloway, 2011). the analysis followed abductive reasoning, and previous theoretical concepts partially guided this analysis, but the categories were created inductively (cohen et al., 2007). the focus was on the reflections and explanations given to their visual gaze and their actions in the pedagogical activities recorded. the analytical process followed the principles of triangulation (flick, 2004). two researchers conducted preliminary analysis and categories from the qualitative data. after that, we discussed the preliminary findings among the entire research team. this process involved individual and shared interpretations and reflections, leading to the identification of the final categories of teachers' explanations of their visual gaze and their pedagogical actions during play and guided activities. 3. findings 3.1 teachers’ visual gaze metrics and related reflections from bar graphs the answer to our first research question is based on the coded teachers’ visual gaze data and teachers’ reflections when they saw the visual gaze metrics (bar graphs) related to the eye-tracking videos. moreover, teachers’ answers to semi structured interview questions were utilized on this section. in the following results subsections, we present the table of coded teacher’s eye-tracking data related to play and guided activities (tables 2 and 3) and reflections of each teacher concerning the fixation-related eye-tracking metrics, such as total duration of fixation, average duration of fixation, and total number ukkonen-mikkola et al 79 | f l r of fixations. examples of bar graphs that was showed to teachers, can find from appendix a. in the end of this section table 4 is presenting the main findings of this first research question. 3.1.1 teachers’ visual gaze metrics from play and their reflections on them this result subsection discusses both the eye-tracking metrics on teachers’ visual gaze during play activity and their own reflections after seeing the bar graphs concerning these activities. in terms of descriptive statistics, the mean total duration of fixation in play was for rose (m=109,640; sd=14,730), for joanna (m=243,598; sd=69,821) and for maria (m=139,638; sd=62,217). table 2 shows each teacher’s total duration of fixation, average duration of fixation, and total number of fixations directed toward each participant child in the activity. a visual representation of the variation in three teacher’s visual gaze metrics in terms of total fixation duration and number of fixations are shown in appendix b, namely figures b1, b2, and b3. these figures add to the in-depth understanding of each teacher’s visual gaze in play and guided scenarios. table 2 fixation metrics concerning teachers’ visual gaze during play teacher child total duration of fixations (ms) average duration of fixation (ms) total number of fixations rose alma 116,831 526 222 elisabeth 88,688 490 181 betty 122,248 497 246 jasmin 110,794 486 228 joanna vincent 292,970 490 598 adam 194,227 482 403 maria oliver 183,633 220 834 emil 95,644 205 467 note. ms = milliseconds. as can be seen from table 2, teachers visual gaze fixation metrics in terms of their eye-tracking data varied between the individual children during play activities. next, teachers explained and reflected on these variations during play activities when they read their own bar graphs that were plotted using their visual gaze data (example of bar graphs can be find from appendix a). rose’s reflections on her visual gaze during the play activity teacher rose specifically noted that the duration of fixations on betty was higher than to other children in activity. she explained that betty had some challenges in understanding speech, so she needed stable eye contact when communicating. this increased the total duration and total number of fixations on betty. data on table 2 seconds roses explanations. total duration and number of fixations were higher with betty. moreover, with alma the average duration of fixations was higher, and it seems that the total duration of the fixations of betty consisted of shorter fixations. however, only elisabeths total duration of fixations was lower than the mean total duration of fixations. joanna’s reflections on her visual gaze during the play activity from the visual gaze metrics, teacher joanna first noted and was surprised by the number of and total duration of fixations on vincent. to joanna, this communicated that vincent was seeking contact. joanna also noticed fewer fixations on adam and commented on him being self-directed and needing less guidance, focus of attention, and gazes. as from table 2 can be seen, the data supports joanna's explanations. the difference between the total duration of fixations between vincent and adam was ukkonen-mikkola et al 80 | f l r quite high but the average duration of fixations was almost the same. this indicates that the lengths of teacher’s gazes in general were similar toward both children, but vincent got more gazes. maria’s reflections on her visual gaze during the play activity after seeing the visual gaze metrics, teacher maria explained that she had longer discussions with oliver because oliver was more communicative, and this increased the total duration of fixations on him. maria tried to contact emil, but he did respond that much. maria also wondered whether her position being next to the table that they were playing on had affected the number of fixations on children. she had to turn her head to see children playing next to her. as a result, in this situation, when oliver was more active in communication and sought more of her attention, this position and oliver’s activeness increased the total number of fixations on oliver. the data presented in the table 2 supports marias' explanations. the total duration and number of fixations for oliver was higher than for emil. moreover, emils total duration of fixations was under the mean total duration of fixations. 3.1.2 teachers’ visual gaze metrics from guided activities and their reflections on them in this subsection, we present the eye-tracking metrics and teachers’ reflections on the bar graphs concerning guided activities. in terms of descriptive statistics, the mean total duration of fixation in guided activities was for rose (m=181,335; sd=88,899), for joanna (m=348,087; sd=45,836) and for maria (m=34,970; sd=4,876). table 3 presents each teacher’s total duration of fixation, average duration of fixation, and total number of fixations directed toward each participant child in the activity. table 3 fixation metrics concerning teachers’ visual gaze during guided activities teacher child total duration of fixations (ms) average duration of fixation (ms) total number of fixations rose alma 182,974 556 329 elisabeth 91,627 475 193 jasmin 269,404 549 491 joanna vincent 380,498 516 737 mathias 315,676 557 567 maria oliver 38,418 179 215 emil 31,522 178 177 note. ms = milliseconds. as it can be seen from table 3, teachers gaze fixation metrics varied between the individual children. when showing bar graphs to the teachers, made using gaze metrics from table 3, they gave some explanations to their visual gaze in guided activity. rose’s reflections on her visual gaze during the guided activity teacher rose explained the large number of fixations focused on jasmin on the basis of her physical location. children were sitting in line, and jasmin was sitting in front of her (smidekova et al., 2020). in turn, elizabeth, who participated in and concentrated on music sessions, actively got fewer gazes than the other children. rose explained that she knew that she could trust her participation without a need to actively look toward her. when looking at table 3, the data seconds these explanations. jasmin got most fixations and also total duration of fixations was highest on her. in addition, for both jasmin and alma the total duration of fixations was higher than the mean total duration of fixations, indicating that in general rose gazed towards them more than towards elisabeth. ukkonen-mikkola et al 81 | f l r joanna’s reflections on her visual gaze during the guided activity teacher joanna explained that mathias was concentrating the whole time on an ongoing painting and did not need so much encouragement. she wondered if this might have reduced the total number of fixations on him. in addition, joanna mentioned that when she looked at mathias, she knew that he had something important to say or that he needed some care. this increased the average duration of fixations on mathias. data in table 3 seconds the explanations of joanna. the difference in the total duration of fixations between vincent and mathias was quite high. the average duration of fixations also supports teachers' explanation by it being higher with mathias, indicating that the total duration of vincent’s fixations consisted of shorter fixations. moreover, for joanna it was surprising how in both situations (play and guided activity), vincent had more fixations than the other children (adam and mathias) because in her interpretation, the other children were both more self-directed. she wondered if her relationship with vincent might have affected this. she wondered if vincent might have sought attention because he was used to getting it from her. maria´s reflections on her visual gaze during the guided activity teacher maria pointed out that oliver had had a bad day, and he got more gazes because of exceptional behavior. in this situation, maria explained that oliver needed support for concentrating and gazes were a means to that end. data shows in table 3, that the difference in the total duration between the children was quite low and the average duration of fixation was almost the same. in addition, a higher number of fixations supports marias explanation of oliver getting more gazes. moreover, towards oliver, the total duration of fixation was higher than the mean total duration of fixations. table 4 summary of findings from teachers’ visual gaze metrics and related reflections from bar graphs teachers’ pedagogical activities teacher findings from teachers’ visual gaze metrics and related reflections from bar graphs play activity rose • longer fixation durations were on betty than other children because she had challenges in understanding speech and needed stable eye contact when communicating. joanna • more fixations were on vincent as he was seeking more contact. • fewer fixations were on adam as he was self-directed and needed less guidance and attention. maria • more fixation durations were on oliver as he was more communicative, and the teacher’s position was closer to him. guided activity rose • more fixations were on jasmin as the teacher was physically closer. • fewer fixations were in elizabeth as she concentrated better. joanna • more total fixation duration was on vincent as he was seeking more contact. • more average fixation duration was on mathias as he received visual attention whenever necessary. maria • more fixation duration was on oliver as he showed exceptional behavior and needed support for better concentration. ukkonen-mikkola et al 82 | f l r 3.2 teachers’ explanations of their visual gaze and their pedagogical actions during play and guided activities in toddler groups to address the second research question, the rta and interviews were analyzed qualitatively. in this results subsection, we describe the categories and provide concrete examples from the rta and interview data. we identified categories regarding teachers’ noticing certain children and their explanations of pedagogical actions in relation to the noticed children. the categories were protection, physical and emotional availability, teaching and learning, facilitation, and initiatives. in this section, teachers’ reflections from play and guided activities are discussed together as the five identified categories characterize the explanations on both play and guided activities. in the end of this section, table 5 is presenting the main findings of this second research question. 3.2.1 teachers protecting children from harm and maintaining safety teachers’ explanations that were categorized under protection were linked to the physical safety of the children and protective caring of a child with particular needs in the situation. teachers mentioned the need to observe children so that they would not hurt or cause harm to themselves or others in any way; thus, observation was linked to prevention of harm. rose recounted in her interview that she had to restrict jasmin from running because she had been ill, and the movement could cause her to cough: well at first, i still tried to get them to play with the legos, and i knew that one of the children had recently been ill, so i didn’t want her to run yet and start coughing again based on her parents’ request. (teacher rose, interview) in this example, the focus is on the healthcare of the child; thus, the observations and lengthy gazes toward certain children were explained as important to be able to guarantee physical wellbeing of the child. 3.2.2 teachers’ physical and emotional availability to respond to children’s needs our second category consisted of explanations in which the teachers emphasized them being physically and emotionally present for the children. in this category, teachers reflected on touch and bodily interaction linked to gaze. emotional availability also occurred in situations where teachers reacted to children’s nonverbal cues. also, verbalized sensitivity was recognized. for example, rose mentioned that jasmin needed consolation and got more attention when she came into the room crying. rose also took jasmin onto her lap to reassure her. furthermore, maria explained that it was important to her that the children felt noticed and emotionally safe during the play. for example, she used touch to make sure that emil knew she was speaking to him when emil was playing further apart from maria. this way she used touch to make the child feel welcome and maria also intentionally switched her position to be closer to emil. maria was very sensitive to the children’s body language. she got signals about oliver’s need to go to the bathroom. oliver was stepping around and crossed his legs and hands. so, maria inquired if he needed to go to the toilet: as the children are toilet training, and they are now both without diapers, i need to remind them at times. (teacher maria, rta) 3.2.3 teaching and learning in pedagogical activities our third category concerning teachers’ explanations of their visual gaze on certain children and teachers’ associated behavior can be linked to teaching and learning. these explanations show how teachers’ gazes on children are linked to pedagogical planning, observations, and solutions. in addition, guidance for learning was characteristic of this category. teachers particularly mentioned the ukkonen-mikkola et al 83 | f l r importance of supporting children’s language skills, consideration of children’s individuality, and consideration of children’s interests. in addition to observing children’s interests, teachers verbalized and narrated actions and objects to children. rose emphasized the importance of teachers speaking aloud to support the development of children’s language. in addition, joanna mentioned the value of naming directions and sizes, such as up, down, uphill, downhill, small, and big, during a free play: especially in the 0–3-year-old children's group, it is very important to name things in the surrounding world. as the children are still in different phases of learning to speak, and they cannot yet name things, i find it very important for the teacher to do so. that is, then, how the children can also learn new words and concepts. (teacher rose, interview) the teachers often observed and recognized the children’s cues and showed flexibility in their pedagogical activities. for example, joanna noticed that the art sessions had developed more through children’s actions than by her original plan. during play, maria also took ideas from children and praised them for their ideas. this way, teachers enhanced their participation (e.g., ukkonen-mikkola & fonsén, 2018). 3.2.4 facilitation supporting child's persistent engagement on activity the fourth category that characterizes diverse explanations for teachers’ visual gaze and related pedagogical actions is facilitation. in this category, explanations focus on teachers’ gazes toward certain children, and teachers’ pedagogical actions that are associated with motivating the children to continue, and give guidance, ideas, and examples to children. these situations provided evidence of teachers’ precise observation skills and willingness to engage children during pedagogical activities. when reflecting on the eye-tracking videos, joanna explained that she utilized her prior knowledge about vincent’s individual characteristics in the given situation, wherein vincent’s behavior showed that he was not concentrating on the task and seemed distracted (sleepy and demotivated). after observing these situations, joanna encouraged vincent to return to the task using verbal communication (e.g., salminen et al., 2021b). teachers paid attention to how they tended to give examples of the way a pedagogical activity should be performed. when observing the videos, joanna also noticed that vincent, at some point, was becoming exhausted and was not that interested anymore in pedagogical activity. joanna pointed out that when she noticed that vincent needed motivation, she tried to give guidance and some new ideas in relation to what he could paint. to encourage vincent to continue, joanna used verbal encouragement and reinforcement to motivate him in the given situation. maria noted her gaze toward oliver and underlined that when she noticed that oliver could not concentrate on the pedagogical activity, she directed him to alternative tasks in order to ease his disruptive behavior, linking the gaze to classroom management: still, it’s interesting to notice that you can somehow steer him [oliver] to take the christmas gnome [a toy] into his lap or ask him to get the gnome. (teacher maria, rta) 3.2.5 interaction initiatives between teacher and children the fifth category of teachers’ explanations of their visual gaze and pedagogical activities were initiatives. the initiatives are linked to interaction between the teacher and the children. teachers reflected on their visual gaze and emphasized that children were seeking opportunities to interact with teachers using verbal and nonverbal cues. when observing the eye-tracking videos, rose explained how she paid attention to children. she took the children’s verbal initiatives into account, and when they suggested something, she immediately integrated the verbal initiatives into the pedagogical activities. additionally, rose stressed that it is very important to react warmly to children’s initiatives. maria also noticed that oliver sought confirmation from her by constantly asking her to see what he had done: ukkonen-mikkola et al 84 | f l r i noticed the need to be kind of reassuring there, and i repeated what he thought he had done. i know that he needs a lot of reassurance in a positive way. (teacher maria, rta) furthermore, the teachers noticed the children’s nonverbal initiatives. when observing the videos, joanna reflected on attention seeking. for instance, first, vincent tried to establish eye contact with joanna. second, he made bodily gestures to seek visual attention from joanna. furthermore, when reflecting the eye-tracking videos, she interpreted vincent’s behavior as getting tired or bored with the ongoing action. this interpretation is based on joanna’s prior knowledge of vincent. in some cases, there were also initiatives that were difficult for teachers to interpret. joanna noticed that she did not quite understand what mathias was pointing at and that mathias was getting frustrated because of it: this older child (vincent) is clearly seeking my attention by acting very small. (teacher joanna, rta) additionally, teachers reflected on their own initiatives. these initiatives were invitations, comments, enquiries, and encouragement of the children. rose made the initiative by inviting a child to lego play. joanna mentioned that she has a very conversational way of interacting, and she asked children a lot of questions: [i have a] fairly talkative way to do everything with the kids. i talk a lot, and at times i know that the kids are also used to me talking and if they communicate with me, especially the other kid in my small group, you can notice that he [vincent] seeks, for example, eye contact and knows that i chat and ask questions. (teacher joanna, rta) table 5 summary of findings from teachers’ explanations of their visual gaze and related pedagogical actions during play and guided activities in ecec categories examples of teachers’ explanations of visual gaze and pedagogical actions during play and guided activities 1. teachers protecting children from harm and maintaining safety • physical safety of the children • protective caring for a child with particular needs • prevention of harm 2. physical and emotional availability: teachers’ physical and emotional availability to respond to children’s needs • touch and bodily interaction linked to gaze • reaction to nonverbal cues • verbal sensitivity in terms of consoling a crying child • physical closeness to a child 3. teachers’ teaching and learning in pedagogical activities • pedagogical planning, observations, and solutions • providing guidance for learning • supporting children’s language skills and consideration of children’s individuality and interests • observing children’s interests, teachers verbalized and narrated actions and objects to children • speaking aloud to support the development of children’s language 4. facilitation supporting the child’s persistent engagement in activity • motivating the children to continue • giving guidance, ideas, and examples to children ukkonen-mikkola et al 85 | f l r 5. interaction initiatives between teacher and children • observing and responding to children’s verbal and nonverbal cues • integrating verbal initiatives into pedagogical activities • reacting warmly to children’s initiatives • providing confirmation and validation to children 4. discussion 4.1. summary and discussion of the findings the aim of this study was to examine ecec teachers’ visual gaze and pedagogical actions during pedagogical activities (play and guided activities) in groups of children under three years old in finland. first, the present study investigated ecec teacher’s visual gaze metrics in terms of children and how teachers reflected on their own fixation durations from these metrics when they saw the bar graphs obtained from eye-tracking video analyses. second, the aim was to investigate how teachers explained their visual gaze toward children and their pedagogical actions during play and guided activities. the results showed that when the ecec teachers saw bar graphs of their visual gaze, they were surprised at the differences in their visual gaze in the durations of fixations among children in the pedagogical activities. however, the ecec teachers were able to give rational reasons for their visual gaze. our findings indicate that the teachers attributed the high duration of visual gaze during play to the children’s individual needs. this is in line with a previous study wherein teachers focused their visual attention longer on students based on individual support needs (chaudhuri et al., 2022a; seidel et al., 2021). furthermore, teachers attributed a high duration of visual gaze during guided activities to children who showed unpredictable behavior and needed support to concentrate on the given task (van den bogert et al., 2014). this is also in accordance with previous research, which showed that teachers look more at students who are off task or exhibit unpredictable behavior than students who participate actively in the classroom (shinoda et al., 2021). in addition, the teachers gave a longer visual gaze in terms of fixation duration to students requiring reassurance and emotional support. this agrees with previous research showing that ecec teachers need to provide emotional support to toddlers in order to reduce unpredictable behavior, such as tantrums, and encourage task-related behavior engagement (shafer et al., 2022) as well as to offer emotional support and enable a safe and supportive relationship to be built for children’s communication and diverse expressions (salomon et al., 2017). the findings showed that during both play and guided activities, children who sought contact got more gazes. these findings align with previous studies showing that teachers give more visual attention to students when they show interactive or disruptive behavior (goldberg et al., 2021). in addition, we found that the physical position of children in the room and their distance to the teacher influenced the frequency and duration of the teacher’s visual gaze. this is partially in line with previous research showing that students who are closer to the teacher may get more gazes than others (smidekova et al., 2020). in our second results subsection, we described how teachers explained their visual gaze toward children and their related pedagogical actions during play and guided activities. five categories regarding teachers’ explanations were identified. the categories were protection, physical and emotional availability, teaching and learning, facilitation, and initiatives. in toddler groups, teachers prioritized the protection of the physical health and safety of children in the classroom. teachers observed the safety of children by being alert with their visual gaze. these findings agree with earlier ukkonen-mikkola et al 86 | f l r studies that found that the protection of safety is an integral part of activities with children (gonzalezmena, 2002; kettukangas, 2017). in addition, ecec teachers showed physical and emotional availability through verbal and nonverbal approaches. earlier studies have shown that teachers often use touch and verbal sensitivity toward children in ecec (hännikäinen, 2013; thomason & la paro, 2009). this could also be in response to the children’s behavior showing sadness and loneliness, wherein the teacher communicates emotional availability by reassuring the child. this kind of emotional availability has been recognized as a key aspect of teachers’ expertise in earlier studies (harkoma et al., 2021). in these two categories, the knowledge of children’s individual health, physical development, and interaction was utilized. according to our results, teachers select their pedagogical activities based on their conceptions of teaching and learning principles. they focus on developing children’s linguistic skills and consider their interests and individuality. the teachers remain flexible with their plans for the day and incorporate the children’s interests during the lessons (la paro et al., 2012.; ukkonen-mikkola & fonsén, 2018). these findings are in line with earlier studies, which showed that children’s participation is an essential part of ecec pedagogy (kangas et al., 2021). in addition, teachers need to model instructions for children to follow and learn. we found that teachers are aware of the importance of speech in daily activities and that children’s learning is constant throughout the whole day in all activities and interaction situations. this result supports the idea that pedagogy in ecec is implemented throughout the whole day (lämsä, 2021). in this category, teacher’s conceptions of teaching and knowledge of a child’s cognitive development and individuality have been taken into account. teachers facilitate a child's persistent engagement in activity and offer guidance when children face challenges while concentrating. children often need encouragement and verbal guidance from the teacher to maintain their attention toward and concentration on a given task (salminen et al., 2021b). teachers could also provide this guidance using their visual gaze, for example, by establishing eye contact. this finding is in line with earlier studies (haataja et al., 2020). finally, teachers need to consider children’s verbal and nonverbal initiatives during lessons. since children in ecec groups are quite young, they often practice their social skills by initiating communication and interacting with their teachers and peers. these nonverbal cues have also been identified in earlier studies (pursi, 2019; salomon et al., 2017). moreover, teachers need to invite children to participate in discussions, comment on children’s tasks and inquire about children’s interests and feelings. these teachers’ initiatives are an essential part of ecec pedagogy and their interactions with children (clark, 2005; ranta et al. 2023). in these two categories, the teachers utilized their prior knowledge related to individual children. 4.2. conclusions, limitations, and future studies our findings indicate that ecec teachers utilize their prior knowledge concerning children’s individuality, development, learning, and interaction when they explain their visual gaze and decisions concerning the pedagogical actions with children. these explanations based on prior knowledge about the child and the child group in general are in line with what could be interpreted as one aspect in knowledge-based reasoning (seidel et al., 2011; seidel & stürmer, 2014; sherin & van es, 2009). however, in this study, we did not intend to explore the concept of knowledge-based reasoning any further but focused on the explanations given to gaze metrics and pedagogical actions in a more general manner. according to our findings, knowledge-based reasoning can broaden teachers’ professional vision, in other words, teachers’ ability to notice activities in toddler groups and reasoning for their pedagogical actions (see seidel & stürmer, 2014). further, we can state that the teachers’ gaze behavior has some similar aspects from ecec groups to school classrooms, including secondary school. overall, in our study, the teachers’ reflections and explanations suggest that the teachers approached these activities not only as situations that included both intentional teaching but also the education and care of children, as they are more broadly understood. through their pedagogical actions, it was possible to interpret that their approach to ukkonen-mikkola et al 87 | f l r pedagogy was holistic, in line with the finnish curricula framework for early childhood education. in the finnish framework, the teaching, learning, and caring for children are implemented together throughout the whole day. this interpretation is in harmony with the finnish ecec curriculum (finnish national agency for education, 2022; lämsä, 2021). there are some limitations to this study. one is that the research participants (teachers) were familiar with one of the researchers even before the study was conducted. social relationships between informants and researchers can affect the objectivity of the research (atkins & wallace, 2012). however, as there was a group of researchers, it was possible to reflect jointly on the interpretations from diverse perspectives and critically form an outsider’s perspective, as not all were familiar with the setting. the other possible limitation was the small number of research participants; therefore, the results cannot be generalized extensively (lincoln & guba, 2000). moreover, teachers participating in the study were all in the early stages of their career, therefore comparing teacher s’ visual gaze based on their experience was not relevant. moreover, children who participated in individual data collections varied based on who was present in ecec that day when data were collected. therefore, comparisons of the teacher’s visual gaze between the activities (play and guided) and between the children were not possible. additionally, in this study children’s background information was not collected. however, in future studies it is recommended to collect data regarding children’s social habits and relationship with teacher to gain deeper insights related to ecec teachers’ visual gaze. regardless of these limitations, our study contributes to educational research especially as it utilized eye-tracking technology in ecec research. this was one of the first exploratory studies in ecec settings in finland, and our endeavor was to give an in-depth insight to the way teachers use their visual gaze to reflect on their pedagogical activities with children. with these exploratory and descriptive findings, we contributed to the existing literature of teachers’ professional vision in ecec settings. methodologically data triangulation, where eye-tracking metrics, rta, and interviews are combined, is a useful design for analysis. it also strengthens the validity of the results (flick, 2004). the benefits of data triangulation were clearly evident in this study, as the qualitative data were paramount in understanding the reasons behind the observed gaze behavior. previous research has shown that visual gaze metrics obtained from eye-tracking video data may not be meaningful on their own unless they are combined with other data sources. for example, van den bogert et al. (2014) combined teachers’ verbal reports or rta explaining their observations along with their visual gaze data. in the present study, teachers’ interview data and rta provided justifications related to the variations of teachers’ visual gaze on the children during their pedagogical actions. in future research, it would be advisable to use data sources other than interviews. for example, physiological data, children’s classroom experiences, etc., could be used for improved data triangulation. the practical implication of our study is that it is essential to understand the role of the teacher’s gaze in the interaction and pedagogical actions in authentic ecec activities with toddlers. this understanding can support teachers in their reflections on their focus of attention and pedagogical actions during interaction situations. eye-tracking metrics can be used to facilitate teachers’ reasoning in their pedagogy and in this way enable the development of teachers’ understanding of their pedagogical actions. these findings can also be utilized in ecec teacher training to enhance the understanding of the importance of noticing, reasoning, sensitive interaction, and teachers’ pedagogical actions with toddlers. in addition, the study reveals the possibilities that eye-tracking technologies offer to researchers for studying interactions and teacher’s pedagogical expertise and awareness of their visual gaze and pedagogical actions toward children in ecec groups. this study opens several future research directions. in future research a bigger dataset would enable to group children based on their age, gender, interests, and to investigate how these factors may influence ecec teachers’ visual gaze and makes possible also the comparison of teachers´ gaze behavior between activities and children. additionally, extending the research to activities during the ukkonen-mikkola et al 88 | f l r entire day including outdoor activities is worth exploring. a more explicit study of nonverbal interaction between teachers and toddlers with eye-tracking technology would be a useful goal for future research. in addition, it is possible to utilize eye-tracking metrics to identify the differences between novice and expert teachers’ gazes as well as the gazes' concerning interactions between teachers or between teachers and parents. furthermore, including the gazes of toddlers would be a useful aspect of research. using gaze metrics linked to child development and developmental psychology would provide a new and relevant research field. key points • although teachers were surprised about some aspects of the visual gaze metrics, they explained their gazes on children. • children’s particular needs, unpredictable behavior, need for support to concentrate, gaze seeking, and the position of children increased teachers’ gazes toward children. • five categories concerning explanations of teachers’ visual gazes and pedagogical actions were identified. • teachers used knowledge-based reasoning when explaining their actions. • eye-tracking technologies can be utilized when studying teacher’s pedagogical actions and reflections in ecec. acknowledgments this study was supported by the university of jyväskylä, department of education. we are grateful for all the ecec teachers, children, families, and ecec centers that participated in this study. references atkins, l., & wallace, s. (2012). qualitative research in education. sage publishing. https://doi.org/10.4135/9781473957602 bae, b. (2009). children’s right to participate – challenges in everyday interaction. european early childhood education research journal, 17(3), 391–406. https://doi.org/10.1080/13502930903101594 booren, l. m., downer, d., & vitiello, v. e. (2012). observations of children’s interactions with teachers, peers, and tasks across preschool classroom activity settings. early education and development, 23(4), 517–538. https://doi.org/10.1016/j.appdev.2019.101100 chaudhuri, s., pakarinen, e., muhonen, h., & lerkkanen, m. k. (2024). association between the teacher–student relationship and teacher visual focus of attention in grade 1: student task avoidance and gender as moderators. educational psychology, 1–19. https://doi.org/10.1080/01443410.2024.2346104 chaudhuri, s., muhonen, h., pakarinen, e., & lerkkanen, m.-k. (2022a). teachers’ focus of attention in first-grade classrooms: exploring teachers experiencing less and more stress using mobile eye-tracking. scandinavian journal of educational research, 66(6), 1076– 1092. https://doi.org/10.1080/00313831.2021.1958374 chaudhuri, s., muhonen, h., pakarinen, e., & lerkkanen, m.-k. (2022b). teachers' visual focus of attention in relation to students' basic academic skills and teachers' individual support for students: an eye-tracking study. learning and individual differences, 98, article 102179. https://doi.org/10.1016/j.lindif.2022.102179 https://doi.org/10.4135/9781473957602 https://doi.org/10.1080/13502930903101594 https://doi.org/10.1016/j.appdev.2019.101100 https://doi.org/10.1016/j.appdev.2019.101100 ukkonen-mikkola et al 89 | f l r clark, a. (2005). listening to and involving young children: a review of research and practice. early child development and care, 175(6), 489–505. https://doi.org/10.1080/03004430500131288 cohen, l., manion, l., & morrison, k. (2007). research methods in education (6th ed.). routledge. colson, e. r., & dworkin, p. h. (1997). toddler development. pediatrics in review, 18(8), 255– 259. https://depts.washington.edu/dbpeds/toddlerdvt.pdf dessus, p., cosnefroy, o., & luengo, v. (2016). “keep your eyes on ‘em all!”: a mobile eyetracking analysis of teachers’ sensitivity to students. adaptive and adaptable learning, 9891, 72–84. https://doi.org/10.1007/978-3-319-45153-4_6 doyle, w. (1980). classroom management. kappa delta pi (indiana, united states), 1–35. doi:https://files.eric.ed.gov/fulltext/ed206567.pdf finnish advisory board on research integrity (2012). http://www.tenk.fi/sites/tenk.fi/files/htk_ohje_2012.pdf finnish national agency for education. (2022). national core curriculum for early childhood education and care (regulation oph-700-2022). finnish national agency for education. flick, u. (2004). triangulation in qualitative research. in u. flick, e. kardoff, & i. steinke (eds.), companion to qualitative research (pp. 178–183). sage publishing. flyvbjerg, b. (2006). five misunderstandings about case-study research. qualitative inquiry, 12(2), 219–245. https://doi.org/10.1177/1077800405284363 fox, m. c., ericsson, k. a., & best, r. (2011). do procedures for verbal reporting of thinking have to be reactive? a meta-analysis and recommendations for best reporting methods. psychological bulletin, 137(2), 316–344. https://doi.org/10/c7kr67 guedes, c., cadima, j., aguiar, t. aguiar, c., & barata, c. (2020). activity settings in toddler classrooms and quality of group and individual interactions. journal of applied developmental psychology, 67, article 101100. https://doi.org/10.1016/j.appdev.2019.101100 goldberg, p., schwerter, j., seidel, t., muller, k., & sturmer, k. (2021). how does learners’ behavior attract preservice teachers’ attention during teaching? teaching and teacher education, 97, article 103213. https://doi.org/10.1016/j.tate.2020.103213 gonzalez-mena, j. (2002). infant/toddler caregiving: a guide to routines. cde press. haataja, e., salonen, v., laine, a., toivanen, m., & hannula, m. s. (2020). the relation between teacher-student eye contact and teachers’ interpersonal behavior during group work: a multiple-person gaze-tracking case study in secondary mathematics education. educational psychology review, 33(1), 57–67. https://doi.org/10.1007/s10648-020-09538-w hakkarainen, p., brėdikytė, m., jakkula, k., & munter, h. (2013). adult play guidance and children’s play development in a narrative play-world. european early childhood education research journal, 21(2), 213–225. https://doi.org/10.1080/1350293x.2013.789189 harkoma, s. m., sajaniemi, n. k., suhonen, e., & saha, m. (2021). impact of pedagogical intervention on early childhood professionals’ emotional availability to children with different temperament characteristics. european early childhood education research journal, 29(2), 183–205. https://doi.org/10.1080/1350293x.2021.1895264 helavaara robertson, l., kinos, j., barbour, n., pukk, m., & rosqvist, l. (2015). child-initiated pedagogies in finland, estonia and england: exploring young children’s views on decisions. early child development and care, 185(11-12), 1815–1827. https://doi.org/10.1080/03004430.2015.1028392 hietanen, j., leppänen, j., peltola, m., linna-aho, k., & ruuhiala, h. (2008). seeing direct and averted gaze activates the approach–avoidance motivational brain systems. neuropsychologia, 46(9), 2423–2430. https://doi.org/10.1016/j.neuropsychologia.2008.02.029 holloway, i. (2011). being a qualitative researcher. qualitative health research, 21(7), 968–975. https://doi.org/10.1177/1049732310395607 holmqvist, k., nyström, m., andersson, r., dewhurst, r., jarodzka, h., & deweijer, j. (2015). eye tracking: a comprehensive guide to methods and measures. oxford university press. https://files.eric.ed.gov/fulltext/ed206567.pdf https://doi.org/10/c7kr67 https://doi.org/10.1007/s10648-020-09538-w https://doi.org/10.1080/1350293x.2021.1895264 https://doi.org/10.1016/j.neuropsychologia.2008.02.029 https://doi.org/10.1177/1049732310395607 ukkonen-mikkola et al 90 | f l r huang, y., miller, k. f., cortina, k. s., & richter, d. (2021). teachers’ professional vision in action: comparing expert and novice teacher’s real-life eye movements in the classroom. zeitschrift für pädagogische psychologie, 37(1-2), 122–139. https://doi.org/10.1024/10100652/a000313 häggström, c., englund, m., & lindroos, o. (2015). examining the gaze behaviors of harvester operators: an eye-tracking study. international journal of forest engineering, 26(2), 96–113. https://doi.org/10/f3ptw6 hännikäinen, m. (2013). varhaiskasvatus pienten lasten päiväkotiryhmissä: hoitoa, kasvatusta vai opetusta? [early childhood education in the young children’s day care groups: care, education or teaching?]. in k. karila & l. lipponen (eds.), varhaiskasvatuksen pedagogiikka [the pedagogy of early childhood education] (pp. 28–56). vastapaino. hännikäinen, m. (2015). the teacher’s lap – a site of emotional well-being for the younger children in day-care groups. early child development and care, 185(5), 752–765. https://doi.org/10.1080/03004430.2014.957690 ishibashi, m., takahashi, m., & nozawa, s. (2020). relationship between ecec (early childhood education and care) staff’s years of professional experience and gaze patterns: a study using wearable eye trackers. cognitive studies: bulletin of the japanese cognitive science society, 27(4), 540–53. https://10.11225/cs.2020.019. isotalo, s., ukkonen-mikkola, t., lämsä, j. & rutanen, n. (2024). early childhood education and care teachers’ gaze behavior across pedagogical episodes in toddler groups in finland. international journal of early childhood. https://doi.org/10.1007/s13158-023-003876 jamison, k., cabell, s., locasale-crouch, j., hamre, b., & pianta, r. (2014). class–infant: an observational measure for assessing teacher–infant interactions in center-based childcare. early education and development, 25(4), 553–572. https://doi.org/10.1080/10409289.2013.822239 johansson, c., åberg, m., & hedlin, m. (2021) touch the children, or please don’t – preschool teachers’ approach to touch. scandinavian journal of educational research, 65(2), 288–301. https://doi.org/10.1080/00313831.2019.1705893 kangas, j., ukkonen-mikkola, t., harju-luukkainen, h., ranta, s., chydenius, h., lahdenperä, j., neitola, m., kinos, j., sajaniemi, n., & ruokonen, i. (2021). understanding different approaches to ece pedagogy through tensions. education sciences, 11(12), article 790. https://doi.org/10.3390/educsci11120790 kettukangas, t. (2017). perustoiminnot-käsite varhaiskasvatuksessa [the concept of basicactivities in ecec]. [doctoral dissertation, university of eastern finland savonlinna]. publications of the university of eastern finland dissertations in education, humanities, and theology. la paro, k. m., hamre, b. k., & pianta, r. (2012). classroom assessment scoring system toddler. paul h. brookes publishing co. ledbury, r., white, i., & darn, s. (2004). the importance of eye contact in the classroom. the internet tesl journal, 10(8), 11–21. http://iteslj.org/techniques/darn-eyecontact.html lincoln, y. s., & guba, e. g. (2000). the only generalization: there is no generalization. in r. gomm, m. hammersley, & p. foster (eds.), case study method key issues, key texts (pp. 27– 44). sage publishing. lucas revilla, y., rutanen, n., harju, k., sevon, e., & raittila, r. (2022). relational approach to infant-teacher lap interactions during the transition from home to early childhood education and care. journal of early childhood education research, 11(1), 253-271. https://journal.fi/jecer/article/view/114018 lunn browlee, j., schraw, g., walker, s., & ryan, m. (2016). changes in preservice teachers’ personal epistemologies. in j. greene, b. sandoval, & i. braten (eds.), handbook of epistemic cognition (pp. 300–325). routledge. https://doi.org/10.1024/1010-0652/a000313 https://doi.org/10.1024/1010-0652/a000313 ukkonen-mikkola et al 91 | f l r lämsä, t. (2021). kokopäiväpedagogiikka ja sen kehittäminen varhaiskasvatuksessa [the whole day pedagogy and its development in ecec]. kasvatus ja aika, 15(2), 79–86. https://doi.org/10.33350/ka.102527 mcintyre, n. a., draycott, b., & wolff, c. e. (2021). keeping track of expert teachers: comparing the affordances of think-aloud elicited by two different video perspectives. learning and instruction, 80, article 101563. https://doi.org/10.1016/j.learninstruc.2021.101563 mcintyre, n. a., & foulsham, t. (2018). scanpath analysis of expertise and culture in teacher gaze in real-world classrooms. instructional science, 46(3), 435–455. https://doi.org/10.1007/s11251-017-9445-x mcintyre, n. a., jarodzka, h., & klassen, r. m. (2019). capturing teacher priorities: using realworld eye-tracking to investigate expert teacher priorities across two cultures. learning and instruction, 60, 215–224. https://doi.org/10.1016/j.learninstruc.2017.12.003 mcintyre, n. a., & mainhard, t. m. (2020). looking to relate: teacher gaze and culture in student-rated teacher interpersonal behaviour. social psychology of education, 23(2), 411– 431. https://doi.org/10.1007/s11218-019-09541-2 muhonen, h., pakarinen, e., & lerkkanen, m-k. (2022). professional vision of grade 1 teachers experiencing different levels of work-related stress. teaching and teacher education, 110, article 103585. https://doi.org/10. 1016/j.tate.2021.103585 muhonen, h., pakarinen, e., rasku-puttonen, h., & lerkkanen, m.-k. (2020). dialogue through eyes: exploring teachers’ focus of attention during educational dialogue. international journal of educational research, 102, article 101607. https://doi.org/10.1016/j.ijer.2020.101607 organisation for economic co-operation and development. (2018). engaging young children: lessons from research about quality in early childhood education and care. starting strong. oecd publishing. https://doi.org/10.1787/9789264085145-en pihlaja, p., & holst, t. (2013). how reflective are teachers? a study of kindergarten teachers´ and special teachers´ levels in reflection in day care. scandinavian journal of educational research, 57(2), 182–198. https://doi.org/10.1080/00313831.2011.628691 pramling samuelsson, i. & asplund carlsson, m. (2008). the playing learning child: towards a pedagogy of early childhood. scandinavian journal of educational research, 52(6), 623–641, https://doi.org/10.1080/00313830802497265 pursi, a. (2019). play in adult-child interaction: institutional multi-party interaction and pedagogical practice in a toddler classroom. learning, culture and social interaction, 21, 36– 150, https://doi.org/10.1016/j.lcsi.2019.02.014. ranta, s., kangas, j., harju-luukkainen, h., ukkonen-mikkola, t., neitola, m., kinos, j., sajaniemi, n., & kuusisto, a. (2023). teachers’ pedagogical competence in finnish early childhood education. a narrative literature review. education sciences, 13(8), article 791. https://doi.org/10.3390/educsci13080791 rayner, k. (2009). the 35th sir frederick bartlett lecture: eye movements and attention in reading, scene perception, and visual search. quarterly journal of experimental psychology, 62(8), 1457–1506. https://doi.org/10.1080/17470210902816461 rutanen, n., raittila, r., harju, k., lucas revilla, y., & hännikäinen, m. (2023). negotiating ethics-in-action in a long-term research relationship with a young child. human arenas, 6(2), 386-403. https://doi.org/10.1007/s42087-021-00216-z sadamatsu, j. (2022). experienced nursery teachers gaze longer at children during play than do novice teachers: an eye-tracking study. asia pacific education review 24, 577–589. https://doi.org/10.1007/s12564-022-09818-w salminen, j., guedes, c., lerkkanen, m.-k., pakarinen, e., & cadima, j. (2021a). teacher–child interaction quality and children’s self-regulation in toddler classrooms in portugal and finland. infant and child development, 30(3). article e2222. https://doi.org/10.1002/icd.2222 https://doi.org/10.1016/j.learninstruc.2021.101563 https://doi.org/10.1016/j.learninstruc.2017.12.003 https://doi.org/10.1007/s11218-019-09541-2 https://doi.org/10.1016/j.ijer.2020.101607 https://doi.org/10.1787/9789264085145-en ukkonen-mikkola et al 92 | f l r salminen, j., muhonen, h. & lerkkanen, m.-k. (2021b). scaffolding patterns of dialogic exchange in toddler classrooms. learning, culture and social interaction, 28, article 100489. https://doi.org/10.1016/j.lcsi.2020.100489 salomon, a., sumsion, j., & harrison, j. (2017). infants draw on ‘emotional capital’ in early childhood education contexts: a new paradigm. contemporary issues in early childhood, 18(4), 362–374. https://doi.org/10.1177/1463949117742771 seidel, t., schnitzler, k., kosel, c., stürmer, k., & holzberger, d. (2021). student characteristics in the eyes of teachers: differences between novice and expert teachers in judgment accuracy, observed behavioral cues, and gaze. educational psychology review, 33, 69–89. https://doi.org/10.1007/s10648-020-09532-2 seidel, t., & stürmer, k. (2014). modeling and measuring the structure of professional vision in preservice teachers. american educational research journal, 51(4), 739–771. https://doi.org/10.3102/0002831214531321 seidel, t., stürmer, k., blomberg, g., kobarg, m. & schwindt, k. (2011). teacher learning from analysis of videotaped classroom situations: does it make a difference whether teachers observe their own teaching or that of others? teaching and teacher education, 27(2), 259– 267. https://doi.org/10.1016/j.tate.2010.08.009. sherin m. g. & van es, e. a. (2009). effects of video club participation on teachers' professional vision. journal of teacher education, 60(20), 20–37. https://doi.org/10.1177/0022487108328155 shinoda, h., yamamoto, t., & imai-matsumura, k. (2021). teachers’ visual processing of children’s off-task behaviors in class: a comparison between teachers and student teachers. plos one, 16(11). article e0259410. https://doi.org/10.1371/journal.pone.0259410 slot, p. l., bleses, d., justice, l. m., markussen-brown, j., & højen, a. (2018). structural and process quality of danish preschools: direct and indirect associations with children’s growth in language and preliteracy skills. early education and development, 29(4), 581–602. https://doi.org/10.1080/10409289.2018.1452494 smidekova, z., janik, m., minarikova, e., & holmqvist, k. (2020). teachers’ gaze over space and time in a real-world classroom. journal of eye movement research, 13(4). https://doi.org/10.16910/jemr.13.4.1 thomason, a. c., & la paro, k. m. (2009). measuring the quality of teacher–child interactions in toddler child care. early education and development, 20, 285–304. https://doi.org/10.1080/10409280902773351 tight, m. (2010). the curious case of case study: a viewpoint. international journal of social research methodology, 13(4), 329–339. https://doi.org/10.1080/13645570903187181 tobii ab. (2018). tobii pro glasses 2 product description manual (v 1.95). https://www.tobiipro.com/siteassets/tobii-pro/product-descriptions/tobii-pro-glasses-2-product van den bogert, n., van bruggen, j., kostons, d., & jochems, w. (2014). first steps into understanding teachers’ visual perception of classroom events. teaching and teacher education, 37, 208–216. https://doi.org/10.1016/j.tate.2013.09.001 vygotsky, l. s. (1967). play and its role in the mental development of the child. soviet psychology, 5(3), 6–18. https://doi.org/10.2753/rpo1061-040505036 ukkonen-mikkola, t., & fonsén, e. (2018). researching finnish early childhood teachers’ pedagogical work using layder’s research map. australasian journal of early childhood, 43(4), 48–56. https://doi.org/10.23965/ajec.43.4.06 whittaker, j. e. v., williford, a. p., carter, l. m., vitiello, v. e., & hatfield, b. e. (2018). using a standardized task to assess the quality of teacher–child dyadic interactions in preschool. early education and development, 29(2), 266–287. https://doi.org/10.1080/10409289.2017.1387960 wolff, c. e., jarodzka, h., & boshuizen, h. p. a. (2020). classroom management scripts: a theoretical model contrasting expert and novice teachers’ knowledge and awareness of https://doi.org/10.1177/1463949117742771 https://doi.org/10.1007/s10648-020-09532-2 https://doi.org/10.3102/0002831214531321 https://doi.org/10.1371/journal.pone.0259410 https://doi.org/10.1080/10409289.2018.1452494 https://doi.org/10.16910/jemr.13.4.1 https://doi.org/10.1080/10409280902773351 https://www.tobiipro.com/siteassets/tobii-pro/product-descriptions/tobii-pro-glasses-2-producthttps://doi.org/10.1016/j.tate.2013.09.001 ukkonen-mikkola et al 93 | f l r classroom events. educational psychology review, 33(1), 131–148. https://doi.org/10.1007/s10648-020-09542-0 https://doi.org/10.1007/s10648-020-09542-0 ukkonen-mikkola et al 94 | f l r appendixa in appendix-a, figure a1 shows the example of one case study teacher, rose’s visual gaze metrics. after recording the eye-tracking videos, teacher rose’s visual gaze was coded and thereafter visually represented in the form of bar graphs. bar graphs a1.1, a1.2, and a1.3 showed rose’s visual gaze metrics on children during play activities and bar graphs a1.4, a1.5, and a1.6 showed rose’s visual gaze metrics on children during guided activities. these bar graphs were shown to rose in order to evoke reflections related to her visual gaze on children. figure a1. example of bar graphs shown to teacher rose to evoke their reflections related to their visual gaze. appendixb a1.1 a1.2 a1.3 a1.4 a1.5 a1.6 ukkonen-mikkola et al 95 | f l r in appendix-b, figure b1, b2 and b3 showed how teacher rose, joanna, and maria’s visual gaze on children varied in play and guided activities. in the case of teacher rose, one child, betty, was not present in the guided activity. additionally, in the case of teacher joanna, adam was present in the play but not in guided activity, whereas mathias was present only in guided and not in play activity. in the case of all three teachers, it is visible that teachers’ visual gaze duration is not equal and varies from child to child. these variations are explained by teachers in sections 3.1 and 3.2 of the present study. figure b1. teacher rose’s visual gaze on children in play and guided activities ukkonen-mikkola et al 96 | f l r figure b2. teacher joanna’s visual gaze on children in play and guided activities ukkonen-mikkola et al 97 | f l r figure b3. teacher maria’s visual gaze on children in play and guided activities ukkonen-mikkola et al 98 | f l r appendixc in appendix-c, the instructions for the retrospective think aloud (rta) given prior to showing the eye tracking recording to the participant are presented. instructions given to the ecec teachers by the researchers: “here is a recording of your eye movements and activity recorded by the eye movement camera. i'll show it to you now. watch the video and tell me what you thought during this teaching and guidance situation. explain why you acted in that way. if you want to stop the recording, press the space key. to continue, press the space key again. if you do not speak for a long time, i will remind you by saying, ‘you can continue speaking.’” while showing gaze metrics in the form of bar graphs to the teachers, the following question was asked to evoke their reflections related to their pedagogical actions from their visual gaze: “what kind of thoughts do these diagrams evoke for you?” the list of questions asked during semi-structured interview. from the list are excluded those questions, which are used in the other forthcoming studies. 1. “how did you experience the video recording situation?” 2. “were there any surprises while watching the eye tracking video? if so, what kind?” 3. “was there anything in the eye tracking video that would make you change your actions?” these questions and seeing the bar graphs inspired teachers to ponder explanations of their visual gaze towards certain children. they also referred to their own or children's actions during pedagogical activities, to reason their visual gaze. frontline learning research vol. 12 no. 4 (2024) 113 137 issn 2295-3159 corresponding author: tessa consoli, university of zurich, switzerland , email: tessa.consoli@ife.uzh.ch doi: https://doi.org/10.14786/flr.v12i4.1555 different media education approaches predict distinct aspects of digital citizenship tessa consoli university of zurich, switzerland article received 23 july 2024 / article revised 9 december 2024 / accepted 10 december 2024 / available online 23 december 2024 abstract based on a national survey of 8,915 students, this study examines the prevalence of different types of media education practices in swiss upper secondary schools and their relationship with students’ online civic engagement and respectful online behavior. descriptive statistics reveal that schools place little emphasis on media education and that current practices focus primarily on protective media education approaches. results from multi-level hierarchical multiple regression analysis show that expressive media activities positively predict students’ online civic engagement but not respectful online behavior, whereas addressing topics such as the ‘dangers of the internet’ predicts respectful online behavior but not online civic engagement. these findings underscore the challenge of promoting digital citizenship education, which simultaneously encourages civic engagement and respectful behavior. additionally, there is a significant negative correlation between online civic engagement and respectful behavior, suggesting that digital citizenship frameworks encompassing both dimensions are more prescriptive than empirical in nature and that diverse digital citizenship profiles may exist, with some being more respectful butt less civically engaged, and others being more civically engaged but less respectful. keywords: media education, digital citizenship education, online civic engagement, online political participation, online respectful behavior, netiquette consoli 114 | f l r 1. introduction research on youth political engagement shows two contrasting trends. a prominent trend shown by many studies from switzerland (gfs.bern, 2018, 2020, 2023), germany (hurrelmann et al., 2019), and other countries is that the political engagement of young people in institutionalized forms of politics is low or even decreasing. in switzerland, for example, young people are less likely to vote than their counterparts in other european countries (grassi et al., 2024) and they are less likely to vote than older people, which is often seen as a problem for democracy (wittwer, 2015). furthermore, about 30% of young people also do not have the right to participate in voting because they do not possess swiss nationality (bundesamt für statistik, 2023). however, several scholars have pointed out that young people have a different understanding of what the political dimension is, and that they are often involved in "unconventional" forms of politics: the so-called participatory politics (weiss, 2020; parth et al., 2020; jenkins, 2009; kahne et al., 2016). these forms of politics are peer-based, nonhierarchical, interactive, and independent of elite-driven institutions and can help shift cultural and political understandings and create pressure for change (jenkins, 2009; kahne et al., 2016). national surveys show that these unconventional forms of politics are also present in switzerland and seem to have an impact on the mobilisation of young people, although their relevance and acceptance among young people seems to have declined slightly in recent years (gfs.bern, 2018; gfs.bern, 2022). digital media play a crucial role in participatory politics by making it easier for citizens to join political campaigns, mobilize their contacts, participate in fundraising, sign online petitions, and engage in political debates. for this reason, several scholars consider school-based media education an opportunity to foster greater civic and political engagement among young people (mihailidis & thevenin, 2013). however, only a few large-scale survey studies have investigated the relationship between school-based media education practices and students’ online civic engagement (bowyer & kahne, 2020; martens & hobbs, 2015). a large-scale panel study by kahne and bowyer (2020) shows that providing students with opportunities to learn about the creation and sharing of digital content has a significant impact on students’ online civic engagement. nonetheless, not much is known about the situation in switzerland and whether these findings also apply in this particular context. furthermore, existing studies on the impact of media education often focus either on a critical or a creative-expressive approach to media education (botturi, 2019). to the author’s knowledge, no study has compared the potential of different types of media education practices at schools in promoting students’ online civic engagement. moreover, existing studies often focus only on online civic engagement, leaving out relevant digital citizenship dimensions, such as online respectful behavior (jones & mitchell, 2016). therefore, this study investigates the extent to which swiss upper secondary school students have the opportunity to engage in different types of school-based media education practices and whether different types of school-based media education practices are related to students’ online civic engagement and respectful behavior. this study aims to contribute to the understanding of the complex mechanisms involved in young people’s development of digital citizenship (choi, 2016; jones & mitchell, 2016; krutka & carpenter, 2017). by identifying school-based media education practices that are likely to be particularly effective in promoting young people’s digital citizenship practices, this study may also inform practitioners and policymakers interested in implementing school digital citizenship education practices. 1.1 online civic engagement and digital citizenship in the last few decades, the possibilities of political and civic participation have expanded beyond traditional forms, such as voting, demonstrating, or militating in a political party (choi, 2016; isin & ruppert, 2020; theocharis, 2015; theocharis & van deth, 2018a, 2018b). digital technologies have especially opened up a landscape of new participation options, such as signing online petitions, mobilizing social networks, coordinating online and offline activities and protests, engaging in online discussion, disseminating political content online, and participating in online campaigns. many terms have been proposed to describe this new landscape of phenomena. jenkins et al. (2016) developed the concept of «participatory politics» to describe “interactive, peer-based acts through which individuals and groups seek to exert both voice and influence on issues of public concern” (p. 41). the authors argue that these forms of participation often operate outside traditional and hierarchical political power structures. consoli 115 | f l r theocharis (2015) introduced the term “digitally networked participation,” which can be understood as a networked media–based personalized action that is carried out by individual citizens with the intent to display their own mobilization and activate their social networks in order to raise awareness about, or exert social and political pressures for the solution of, a social or political problem (p. 6). he criticized initial scholarly hesitance to recognize digital acts as meaningful political participation and argued that digitally networked participation should be recognized as a legitimate form of political participation and included in surveys and studies on democratic engagement. another term that has emerged to describe new ways to engage in society in the digital age is “digital citizenship.” many different understandings of the term exist (chen et al., 2021; choi, 2016; heath, 2018; isin & ruppert, 2020; krutka & carpenter, 2017). one of the most well-known conceptualizations of the term, which was popularized by the international society for technology in education (iste), is that of ribble and bailey (international society for technology in education, 2007), who defined digital citizenship as the “norms of appropriate, responsible behavior with regard to technology use”. ribble and colleagues (international society for technology in education, 2011; ribble, 2012; ribble & miller, 2013) identified nine areas of digital citizenship: digital etiquette (i.e., netiquette), digital access, digital law, digital communication, digital literacy, digital commerce, digital rights and responsibility, digital safety, and digital health and welfare. according to this conceptualization of digital citizenship, digital citizenship education aims primarily to teach students how to respect and protect themselves and others in online spaces. recently, various authors (heath, 2018; m. johnson, 2015; krutka & carpenter, 2017) have criticized this narrow conceptualization of digital citizenship, arguing that although promoting safety and respectful online behavior are essential aspects of digital citizenship education, they are insufficient to encourage students’ active participation in digital spaces and that this conceptualization is not fully aligned with the notion of democratic citizenship. according to them, education should not only make students safe netizens who follow netiquette rules but also prepare them to use digital tools to promote democratic engagement and address social justice issues. in line with this recommendation, his study adopts the perspective of jones and mitchell (2016), who suggested that digital citizenship education may focus on fostering two dimensions: respectful online behavior and online civic engagement. despite increasing opportunities for civic engagement in the digital age, only a small proportion of young people appear to be politically engaged online (kahne & bowyer, 2019; keating & melis, 2017; schulz et al., 2023). furthermore, the skills and competencies needed to engage in civic and political activities through digital media are unequally distributed among young people, and factors such as gender and class appear to lead to different patterns of online civic engagement (grasso & smith, 2022; grasso & giugni, 2022; jones & mitchell, 2016;van deursen & van dijk, 2014). in addition, digital media also exposes young citizens to potential dangers such as hate speech, misinformation, fake news, and echo chambers (rhodes, 2022). for these reasons, several scholars have argued that pedagogical interventions may be needed to provide young people with equal opportunities to learn the skills and competencies needed to participate safely in online spaces and society (jenkins, 2009; kahne et al., 2016; international society for technology in education, 2011; ribble & miller, 2013). 1.2 media education approaches and goals in the evolving landscape of media education, scholars and practitioners have developed various approaches (hobbs, 1999; süss et al., 2008). these approaches typically fall into categories such as protective/preventive, critical, creative/expressive, functional, and civic (see botturi, 2019), each serving distinct educational goals. the protective/preventive approach is primarily concerned with safeguarding vulnerable populations, particularly children and teenagers, from the potential risks and harms of media use. this includes risks such as exposure to inappropriate content, cyberbullying, and privacy breaches. protective media education aims to equip learners with the knowledge and skills to recognize and manage risks effectively. this approach often consoli 116 | f l r highlights the necessity of limiting access to the media. livingstone (2014), for example, investigated how children navigate risks on social media and discussed the strategy of age boundaries and other measures to ensure children’s safety. the critical approach to media education emphasizes the development of the knowledge and competencies needed to critically analyze media content and the impact of media on society. this involves understanding how digital media work (e.g., algorithms that influence what appears in the feed) and influence behaviors through techniques such as framing and agenda-setting. educators employing this approach encourage learners to question the power structures and ideological contexts within which media operate, thereby fostering discerning media consumption. the ultimate goal is to cultivate an audience capable of reading beyond surface-level content and recognizing underlying biases and structures of power. buckingham’s (2019) ‘the media education manifesto’ emphasizes this approach. the creative-expressive approach (highlighted, for example, in burn & durran, 2007) emphasizes using media as a tool for personal and creative expression. it integrates media production into the learning process, allowing students to use different digital tools and platforms to create their own content. this perspective is based on the idea that creating media can help learners to gain a deeper understanding of how media work, while also developing creative and critical thinking skill (banerjee & greene, 2006). the functional media education approach focuses on teaching students practical skills and competencies to effectively use and navigate digital media (botturi, 2019). the goal is to prepare individuals to operate competently in a media-saturated world, ensuring that they can use tools and technologies responsibly and effectively. lastly, the civic approach to media education addresses the role of media in fostering civic engagement and democratic participation (see, for example, mihailidis & thevenin, 2013). it underscores the importance of the media in promoting public discourse, political participation, and community mobilization. through this lens, media education is not just about consumption but also about using media to participate in society. educators encourage learners to use media platforms to voice their opinions, engage in debates, and mobilize for collective action, thus strengthening democratic processes. this media education approach is often seen as part of digital citizenship education (choi, 2016; kahne et al., 2016; krutka & carpenter, 2017). different media education approaches might also be combined. each of these approaches offers valuable perspectives on the diverse objectives of media education. 1.3 the relationship between school-based media education practices and young people’s online civic engagement and respectful behavior a large-scale panel study by kahne and bowyer (2020) shows that approximately 11% of young people are engaged in online political activities. the study also found that giving students opportunities to learn about producing and distributing digital content can significantly enhance their engagement in online political activities. however, the study neither compared the impact of different types of learning opportunities nor examined other dimensions of digital citizenship, such as students’ respectful online behavior (jones & mitchell, 2016). the international civic and citizenship education study (iccs) (2023) considered students’ engagement with political or social issues using digital media (operationalized as interacting with and sharing social media posts) and found that across all participating countries, digital media was seldom used for civic engagement. the iccs also asked teachers to report the opportunities students had to learn about responsible internet use (e.g., privacy, source reliability, and social media) and to perform different activities with digital media (online research, making presentations, class or work, and online posting to support actions about the environment). however, the relationship between these school-based digital media practices and students’ engagement with political or social issues using digital media was not investigated. furthermore, the study did not examine other dimensions of digital citizenship, and no data were collected in switzerland. consoli 117 | f l r studies from other disciplines have examined the relationship between different forms of internet use (e.g., engagement with user-generated content and informational, interactional, and creative use of the internet) and young people’s civic and political engagement without considering school-based activities. östman (2012) found that young people’s involvement in user-generated content (i.e., audience experience characterized by expressivity, performance, and collaboration) is positively related to offline and online political participation but negatively related to political knowledge (operationalized as fact knowledge on institutional politics). this study suggests (by adopting the terminology of bennett et al., 2009) that there might be two different profiles of civic engagement among young people: the ‘dutiful’ citizens (characterized by high levels of news consumption, informational internet use, and political knowledge) and the ‘self-actualizing’ citizens (characterized by high engagement in political expressive behavior). similar results were obtained by ekström and östman (2015), who found that informational and interactional internet use is indirectly related to political participation, while creative online engagement (not necessarily in political activities) is a positive predictor of online and offline political participation and a negative predictor of political knowledge (also operationalized here as fact knowledge on institutional politics). scholars have raised concerns that young people’s engagement in social media and online political activities could detach them from offline and perhaps more effective civic engagement. however, many studies and meta-analyses have proved that this is not the case (boulianne, 2015; boulianne & theocharis, 2020; schulz et al., 2023). however, these studies and meta-analyses mainly consider cross-sectional design data, leaving unanswered questions about the nature of the relationship (e.g., the direction of causality). a recent meta-analysis of repeated-wave panel studies (oser & boulianne, 2020) found evidence of a long-term reinforcing effect (i.e., engaging in political activities increases media consumption and online civic engagement), highlighting the risk of social media use in exacerbating inequalities in political participation rather than mitigating them by mobilising new population groups. examining the relationship between different school-based media education practices and civic engagement is particularly intriguing in this context. the fact that students’ exposure to media literacy education affects students’ engagement in online civic activities (kahne & bowyer, 2019; martens & hobbs, 2015) suggests that schools could potentially mitigate the risks of participatory inequality. from a theoretical perspective, gotlieb and sarge (2021) developed a theoretical model to explain how engagement in user-generated content (ugc) could foster civic readiness among young people. following this model, the characteristics of involvement in ugc (i.e., expressivity, performativity, and collaboration) provide opportunities for both the development of civic skills (i.e., communication, organization, and collective decision-making) and the satisfaction of self-determination needs (deci & ryan, 1985; ryan & deci, 2017), which in turn lead to sustained intrinsic civic motivation. the authors suggest that this model offers practical insights for designing interventions for nurturing active, long-term civic involvement among the youth and that integrating media literacy education emphasizing ugc could even empower resource-poor and marginalized youth. this approach would prepare them for unconventional and conventional political participation, helping them perceive their involvement as self-driven rather than obligatory. many intervention studies have investigated the impact of school-based educational interventions on preventing cyberbullying and promoting respectful online behavior. a systematic review (hutson et al., 2018) found that interventions based on communication and social skills, empathy training, coping skills, and digital citizenship (conceptualized in terms of responsible and respectful online behavior) were effective. few largescale survey studies have analyzed these issues. a survey study by park et al. (2014) found that the more students use the internet and social media, the more likely they are to experience online bullying, victimization, and witnessing. however, the study did not consider the impact of school-based activities. to the author’s knowledge, no study has examined how different media education practices are related to respectful online behavior. consoli 118 | f l r 1.4 research questions given the lack of large-scale studies on the relationship between different types of school media education practices and students’ online digital citizenship practices, as well as the lack of data on the current situation in switzerland, this study addresses the following research questions: a) what kind of school-based media education practices do students encounter in swiss upper secondary schools? b) how do different types of media education practices at school predict students’ online civic engagement and respectful behavior by considering relevant students’ background variables? 2. methodology 2.1 context in the swiss educational landscape, post-compulsory education encompasses roughly two-thirds of students pursuing vocational training, which typically combines on-the-job training with school-based education. the remaining third enrolls in general or specialized baccalaureate programs that are exclusively school-based (bundesamt für statistik, 2024). students who participated in pisa 2022 reported using digital media at school between one and two hours a day, which is well below the oecd average (oecd, 2023). the curricula in both vocational and general baccalaureate programs commonly underline the importance of enhancing students’ computer and information literacy, recognizing them as a critical cross-disciplinary objective. many curricula also incorporate computer science as a distinct academic subject (edk, 2017). in compulsory education curricula, special attention is also given to promoting students’ active civic participation as an objective of general education or education for sustainable development (ciip, 2011; d-edk, 2014). in most cases, these objectives are also included in the curricula of upper secondary schools, although they are sometimes vaguely formulated, and most of them show didactical shortcomings (waldis & ziegler, 2019). switzerland’s political system is a paradigm of semi-direct democracy. unlike representative democracies, where elected officials alone make decisions on behalf of their constituents, swiss semi-direct democracy allows individuals to have a more direct impact on laws and policies through referendums and popular initiatives. this system requires citizens to vote on various issues several times per year. despite these possibilities, young people in switzerland participate much less in voting than older generations (wittwer, 2015). according to a national study on student political participation (gfs.bern, 2018), around 30% of young people aged 15-25 declared that they were very likely to engage in online political activities (such as sharing political content on social media or following politicians on social media). another study found that 22.1% of swiss young adults aged 18-38 (women 17.8%, men 27.8%) discuss or exchange opinions about politics on social media sites, and 12.6% of them (women 9%, men 17.4%) have joined or created a political group on social media (grasso & smith, 2022). these results show that online political engagement is moderately widespread among swiss adolescents and that gender influences it, but nothing is known about the relationship between school activities and the development of this type of engagement. 2.2 sample this research draws on data from a comprehensive national survey investigating the digital transformation of upper secondary schools in switzerland (petko et al, 2022; antonietti et al. 2023). data were collected anonymously and on a voluntary basis through an online questionnaire conducted in two phases: between september and november 2021 in the canton of zurich, and then between may and july 2022 in other cantons. participation was sought from all upper secondary schools in switzerland through email invitations sent to school administrators, who then distributed the survey to their students in the penultimate school year. consoli 119 | f l r approximately 20% of the targeted schools participated in the study. during the survey period, face-to-face classes were held normally in switzerland without pandemic-related restrictions on classroom activities. the sample used in this study consisted of 8,915 students in the penultimate school year of upper secondary school (mean age = 18.15 years, sd = 3.62; median = 17; modus = 17; 48.3% male, 48.3% female, 3.5% other/non-binary) from 108 schools. of these participants, 56.1% were enrolled in vocational programs, 37% in general baccalaureate programs, and 6.9% in specialized baccalaureate schools. the sample consisted predominantly of students from the german-speaking region of switzerland (80.1%), followed by those from the french-speaking (11.4%) and italian-speaking (8.5%) regions of switzerland, reflecting a slight underrepresentation of vocational students and french-speaking students in relation to the national population. in the data cleaning and processing phase, responses were excluded based on the following criteria: time to complete (less than two standard deviations from the mean), incorrect school code entries, and nonsecond-year status. 2.3 measures 2.3.1 frequency of computer use at school to assess the frequency of computer use at school, students were asked how often they used computers or tablets for learning purposes at school. the questions were based on a questionnaire from the european commission (2019). a five-point likert scale for responses, ranging from 1 (‘never or almost never’) to 5 (‘several times a day’), was employed. recognizing that the students were unlikely to use both a tablet and a computer regularly, a new composite variable was created. this variable captured the higher frequency of use between the two devices (tablet or computer) for each student. this composite variable was used in the analyses. 2.3.2 media activities at school the questionnaire developed for the national study asked students how often they used digital technologies to engage in 29 activities. this study specifically examines activities that can foster the development of digital engagement literacy, as termed by kahne and bowyer (2019). these activities can also be called ‘engagement with user-generated content’ and are characterized by being expressive, performative, and collaborative (gotlieb & sarge, 2021; östman, 2012). specifically, we selected items that asked students how often they had the opportunity to use social media, develop and design online content, create video or audio productions, collaborate online with other learners, and publish their work on the internet at school. although these activities are not political in nature, there is empirical evidence that they can promote learners’ civic engagement both online and offline (östman, 2012). the answer option for the items was a five-point likert scale ranging from 1 (‘never ‘) to 5 (‘very often ‘). 2.3.3 media education topics addressed at school as it is crucial to teach not only how to use digital technologies (for learning and other purposes) but also how they work (guggemos & seufert, 2021; schmitz et al., 2024) and the impact they have on society (brinda et al., 2020), the survey asked students whether they had the opportunity to address the following topics at school: the impact of digital technologies on the economy and work, the impact of digital technologies on democracy, the impact of digital technologies on the environment, the impact of digital technologies on health (physical and mental), the impact of digital technologies on social relationships (friendships, family, etc.), the reliability of online information, the dangers of the internet (e.g., scams, data security). the response options were dichotomous (yes/no). 2.3.4 online civic engagement to investigate students’ online civic engagement, this study developed a short scale inspired by the internet political activism sub-scale developed by choi et al. (2017) and jones and mitchell’s (2016) online civic engagement sub-scale. the short scale inquired whether the students (1) had used the internet to make a difference on a political or social level, (2) participate in online groups that engage with political or social consoli 120 | f l r issues, or (3) regularly publish posts on political or social issues on the internet. responses were measured on a five-point likert scale, ranging from 1 (‘completely disagree’) to 5 (‘completely agree’). both exploratory and confirmatory factor analyses confirmed the unidimensionality of the scale. the scale demonstrated a cronbach’s alpha of 0.794, indicating a good internal consistency. the three variables measured by the scale were combined into a single average score variable for additional analysis. 2.3.5 respectful online behavior students’ engagement in respectful online behavior was assessed using a short scale inspired by jones and mitchell’s (2016) internet online respect sub-scale. the scale asked students whether they (1) made sure that the pictures they posted or shared of other people would not embarrass them or get them into trouble, (2) made sure that they would not later regret the things they said and posted online, and (3) participated in offending arguments and interactions online. responses were measured on a five-point likert scale, ranging from 1 (‘completely disagree’) to 5 (‘completely agree’). both exploratory and confirmatory factor analyses provided empirical evidence for the scale’s unidimensionality. the scale showed a cronbach’s alpha of 0.834, indicating an excellent internal consistency. the three aspects of the scale were combined into a single mean score variable for further analysis. 2.3.6 students ‘background variables this study considered the following background variables (for an overview of the relationship between young people’s background variables and their civic engagement, see schulz et al., 2023): a) gender: students were asked to identify their gender as male, female, or other. this variable was subsequently transformed into two dummy variables. b) language spoken at home: students were asked about the primary language they spoke at home. this variable was recoded dichotomously, where ‘1’ indicates a swiss national language and ‘0’ represents any other language. c) type of schooling: the analysis differentiated between students attending a full-time school program (specialized and general baccalaureate schools), who received code ‘0’, and students involved in a vocational education and training program, who received code ‘1’. d) parental education level: the education levels of the students’ mothers and fathers were considered (“low”, “middle” or “high”). this variable was transformed into two dummy variables. e) language region of the attended school: the three national language regions represented in the sample (german-speaking, french-speaking, and italian-speaking) were considered. this variable was transformed into two dummy variables. 2.4 descriptive statistics a simple weighting procedure was applied to the data to correct over-representation and underrepresentation of certain sub-populations in the sample and thus make the data more representative. the weighting procedure was performed with spss 26 and considered school type (vocational, general baccalaureate, and specialized baccalaureate) and language region (german-speaking part, french-speaking part, and italian-speaking part). the weights were not applied to further multilevel multiple-regression analysis, as unweighted estimates are considered more unbiased and consistent (winship & radbill, 1994). 2.5 hierarchical multilevel multiple regression analysis recognizing the nested structure of the data (schielzeth & nakagawa, 2013), with students grouped within schools, this study employed a hierarchical multilevel multiple regression approach with random intercepts for schools. as it was not possible to take the class level into account in the study, and given that the values reported in school-based digital media practices might be influenced by (a) students’ affiliation to different schools, (b) different study programs, (c) different classes, and (d) individual preferences and perceptions, the study utilizes a grand-mean-centering approach, which allows for a simultaneous investigation consoli 121 | f l r of contextual and individual effects without modeling them separately (enders & tofighi, 2007; paccagnella, 2006; wu & wooldridge, 2005). compared to a simple multiple regression approach, this approach allows the intercepts to vary between schools and estimates the percentage of variation due to students’ different school affiliations. residual normality was assessed using two graphical diagnostic approaches: inspection of quantilequantile (q-q) plots and residual histograms. observations with standardized residuals greater than three standard deviations from the mean were excluded. in concrete, 32 observations in the multilevel regression analysis with online civic engagement as the output variable and 84 observations in the analysis with engagement in respectful online behavior as the output variable were excluded. the analysis was conducted utilizing r and jamovi (the jamovi project, 2022; r core team, 2021). the following r packages and jamovi modules were used for the analysis: psych (revelle, 2019), car (fox & weisberg, 2020) and gamlj (gallucci, 2019). missing data were handled by the default listwise exclusion method embedded within the software, and three schools with fewer than five participants were excluded from the sample. the marginal r² provided an estimate of the variance explained by the fixed factors alone, while the conditional r² offered insight into the total variance explained by both fixed and random factors (nakagawa & schielzeth, 2013). a likelihood radio model testing was also performed to assess whether the more complex models fit the data significantly better than the simpler one (peugh, 2010). for each dependent variable (online civic engagement and respectful online behavior), three models were constructed: 1. null model (model 0): this baseline model, without any predictors, aimed to partition the variance in the dependent variable into within-group (individual-level) and between-group (school-level) components. it provides a reference point for assessing the variance explained by subsequent models. 2. background variables model (model 1): incorporating only background variables (gender identity, language spoken at home, type of schooling, parental education level, and language region), this model assessed the influence of these factors on the dependent variables. 3. full model (model 2): this full model, which includes the background and school digital media practices variables, was designed to measure the predictive impact of all independent variables on online civic engagement and respectful online behavior. the comparison of r2 values between model 1 and model 2 revealed the additional variance in dependent variables explained by the media school-based activities variables alone. 3. results 3.1 descriptive statistics the descriptive statistics (table 1) of the main variables used in this study showed that the frequency of computer use had a mean of 3.74 (sd = 1.31) on a 5-point likert scale, indicating a moderate to high level of computer use in class. despite this high level of computer use, the data showed that students engaged in only a few media activities that had the potential to promote digital engagement literacy. using social media was the most frequently performed activity, with a mean of 3.11 (sd = 1.55). this was followed by online collaboration (m = 2.5; sd = 1.38). however, the standard deviation of these variables suggests a wide range of responses, which indicated that these activities had not yet been implemented homogeneously in schools. all other media activities had low mean scores and lower standard deviations, suggesting that they were rarely carried out and that the students’ responses were more consistent. publishing one’s own work online had the lowest mean of 1.62 (sd = 1.05), reflecting the very limited opportunity for students to engage in this activity at school. creating a video or audio production had a mean of 1.81 (sd = 1.08), and developing and designing online content had a mean of 1.91 (sd = 1.2). consoli 122 | f l r table 1 also displays the percentages of students who had studied a particular media education topic at school. these variables were categorical, with responses indicating either the presence or absence of the topic in their education; thus, no mean or standard deviation was reported. the most frequently discussed topic at school was the dangers of the internet (65.5% yes). the second most frequently discussed topic was the reliability of online information (55.5% yes). the correlation between the two dependent variables was also calculated and showed a negative and significant, albeit weak, correlation with each other (r = -0.1, p < .001). table 1 descriptive statistics of predictors and dependent variables computed with weighted data. 3.2 multilevel multiple regression analysis: prediction of online civic engagement table 2 displays the model fits and estimates of all three models. model 0 served as a baseline, including only random intercepts without fixed effects. it yielded a marginal r2 of 0 and a conditional r2 of 0.009, indicating minimal explained variance. the low intra-class correlation (icc) of 0.009 indicates that only 0.9% of the variance in online civic engagement is due to school-level factors. model 1 introduced the students’ background variables. this model showed a significant improvement over the baseline, with a χ² of 144.266 (df = 10, p < .001), suggesting that the added predictors significantly improved the model fit. the marginal r2 increased to 0.025, and the conditional r2 to 0.028. the decrease in aic and bic also indicated a better model fit than model 0. model 2 extended model 1 by adding two predictors: the media activities students had the opportunity to perform at school and the media education topics they addressed at school. this model explained significantly more variance than model 1, with a χ² of 423.888 (df = 13, p < .001). the model’s marginal r2 was 0.090, and the conditional r2 was 0.095, suggesting that model 2 explained approximately 9.5% of the variance in online civic engagement. the lowest aic and bic also indicated that including these predictors enhanced the model’s fit. in terms of fixed effects, identifying as ‘female’ (compared to students identifying as male) and attending general education (compared to students attending vocational education) were significantly associated with lower levels of online civic engagement (β = 0.083, se = 0.028, p < .01; β = -0.155, se = 0.038, p < .001), while identifying as ‘other’/’non-binary’ and higher maternal education were associated with higher estimates (β = 0.544, se = 0.075, p < .001; β = 0.126, se = 0.039, p < .001). n m sem sd min max yes no α media activities at school frequency of computer use 8915 3.74 0.001 1.312 1 5 publishing own work online 8915 1.62 0.001 1.049 1 5 collaborating online 8915 2.50 0.001 1.380 1 5 developing and designing online content 8915 1.91 0.001 1.199 1 5 using social media 8915 3.11 0.002 1.554 1 5 creating a video or audio production 8915 1.81 0.001 1.079 1 5 media education topics addressed at school 8915 impact on economics and work 8915 34.9% 65.1% impact on democracy 8915 15.1% 84.9% impact on the environment 8915 38.1% 61.9% impact on health (physical and mental) 8915 41.8% 58.2% impact on social relationships 8915 39.9% 60.1% reliability of online information 8915 55.5% 44.5% dangers of the internet (e.g. data security) 8915 65.5% 34.5% online civic engagement 8915 2.05 0.001 1.112 1 5 0.794 respectful online behavior 8915 4.16 0.001 0.973 1 5 0.834 note. m = mean, sem = standard error of mean, sd = standard deviation, min = minimum, max = maximum, α = cronbach’s alpha consoli 123 | f l r regarding the school-based media activity variables, the results showed that the simple frequency of computer use in class was not significantly associated with online civic engagement. by contrast, all other media activities aligned with a creative or expressive media education approach were positively and significantly associated with higher levels of internet political activism. among the media education topics addressed at school, the impact of digital technologies on democracy stood out with a substantial positive effect (β = 0.168, se = 0.039, p < .001). the impact on the environment also showed a significant positive relationship (β = 0.080, se = 0.032, p < .05). other topics did not show significant associations with online civic engagement. graphical analysis of the residual histogram and the q-q plot suggested that the residuals were approximately normally distributed. although deviations from a normal distribution of residuals should not have a large impact in larger samples (bohrnstedt & carter, 1971; schmidt & finan, 2018; vasu & elmore, 1975), some caution is still required when interpreting the results. consoli 124 | f l r table 2 model fit and model estimates for the multilevel analysis with online civic engagement as the dependent variable model 0 model 1 model 2 estimate se estimate se estimate se model fit aic 18428.1 18303.8 17904 bic 18454.4 18443.6 18209.8 loglikel. -9214.1 -9165.1 -8991.5 r-squared marginal 0 0.025 0.090 r-squared conditional 0.010 0.030 0.096 n. par 2 12 25 χ² 98.039 347.221 df 10 13 p <.001 <.001 reference model 0 model 1 random components (intercept) sd 0.106 0.078 0.087 σ 0.011 0.006 0.00757 icc 0.010 fixed effects (intercept) 2.02 *** 0.019 2.120 *** 0.042 2.116 *** 0.046 gender female -0.078 ** 0.028 -0.083 ** 0.028 other 0.525 *** 0.077 0.538 *** 0.075 language at home national language -0.116 *** 0.034 -0.049 0.033 schooling type general education -0.182 *** 0.035 -0.155 *** 0.036 education mother middle 0.075 0.041 0.064 0.040 higher 0.155 *** 0.040 0.128 *** 0.039 education father middle -0.049 0.043 -0.079 0.042 higher 0.054 0.040 0.017 0.039 language region french 0.087 0.053 0.078 0.056 italian -0.053 0.061 -0.011 0.063 media activities at school frequency of computer use -0.018 0.011 publishing own work online 0.113 *** 0.014 collaborating online 0.032 ** 0.012 developing and designing online content 0.079 *** 0.013 using social media 0.044 *** 0.009 creating a video or audio production 0.090 *** 0.014 media education topics addressed at school impact on economics and work 0.016 0.031 impact on democracy 0.168 *** 0.039 impact on the environment 0.080 * 0.032 impact on health (physical and mental) 0.006 0.032 impact on social relationships -0.040 0.032 reliability of online information -0.053 0.031 dangers of the internet -0.057 0.031 note. n = 6160, groups: schoolcode. * p < .05, ** p < .01, *** p < .001 consoli 125 | f l r 3.3 multilevel multiple regression analysis: prediction of respectful online behavior table 3 presents the model fits and estimates derived from the multilevel analysis. model 0, which included only the random intercepts, provided a baseline for comparison with more complex models. in this initial model, icc suggested that 3.0% of the variance in respectful online behavior could be attributed to school-level factors. the marginal r2 of 0 and the conditional r2 of 0.03 indicate that the model explained none or minimal variance in the dependent variable. in model 1, which incorporated students’ backgrounds, a substantial improvement in model fit was observed, as evidenced by a decrease in the aic and bic, and an increase in marginal (0.103) and conditional (0.111) r2 values. this model also significantly deviated from the null model (χ² = 593.992, df = 10, p < .001). model 2 further expanded upon model 1 by including school-based media education practices. this final model showed the best fit with the lowest aic and bic values and an increase in explained variance, with marginal r2 at 0.121 and conditional r2 at 0.127. model 2 significantly deviated from model 1 (χ² = 119.368, df = 12, p < .001). despite the significant deviation, it is interesting to note in this case that the added variance of model 2 (compared to model 1) is much smaller than in the case of civic online engagement. in terms of fixed effects, gender differences were significant, with students identifying as ‘female’ showing a significant positive association (β = 0.412, p < .001) and students identifying as ‘other’/’non-binary’ showing a negative association (β = -0.505, p < .001) with respectful online behavior (compared to students identifying as ‘male’). speaking one of the three national languages at home was positively associated with the outcome (β = 0.131, p < .001) compared to speaking another language, as was attending a general education type of school (β = 0.127, p < .001) compared to attending a vocational school. interestingly, media activities at school had mixed effects. the increased frequency of general computer use at school and collaborating online during lessons was positively and significantly correlated with respectful online behavior (β = 0.063, p < .001; β = 0.024, p < .05), while having the possibility to use social media (β = -0.047, p < .001), create videos or audio productions (β = -0.045, p < .05), and publish one’s own work online (β = -0.041, p < .05) were negatively and significantly associated. finally, among the media education topics addressed at school, the ‘dangers of the internet’ had a significant positive impact on respectful online behavior (β = 0.111, p < .001). similarly, the ‘reliability of online information’ was associated with a positive outcome (β = 0.070, p < .01). an examination of the residual histogram and the q-q plot indicated slight leptokurtosis and some minor asymmetries. despite these minor deviations, the residuals largely conform to normality, which is generally acceptable in large-sample statistical analyses (bohrnstedt & carter, 1971; schmidt & finan, 2018; vasu & elmore, 1975). however, interpreting these results also requires caution. consoli 126 | f l r table 3 model fit and model estimates for the multilevel analysis with respectful online behavior as the dependent variable. model 0 model 1 model 2 estimate se estimate se estimate se model fit aic 15359.202 14785.2 14691.8 bic 15379.35 14872.5 14866.5 loglikel. -7676.601 -7379.6 -7319.9 r-squared marginal 0 0.103 0.121 r-squared conditional 0.030 0.111 0.127 n. par 2 12 25 χ² 593.992 119.368 df 10 13 p <.001 <.001 reference model 0 model1 random components (intercept) sd 0.147 0.076 0.004 σ 0.217 0.006 0.007 icc 0.030 fixed effects (intercept) 4.29 *** 0.020 3.946 *** 0.033 3.970 *** 0.033 gender female 0.409 *** 0.022 0.412 *** 0.022 other -0.527 *** 0.059 -0.505 *** 0.059 language at home national language 0.148 *** 0.026 0.131 *** 0.026 schooling type general education 0.145 *** 0.029 0.127 *** 0.028 education mother middle -0.038 0.032 -0.035 0.031 higher -0.049 0.030 -0.046 0.030 education father middle -0.020 0.033 -0.030 0.033 higher -0.017 0.031 -0.027 0.031 language region french -0.007 0.044 -0.024 0.043 italian -0.052 0.049 -0.046 0.048 media activities at school frequency of computer use 0.063 *** 0.034 publishing own work online -0.041 * 0.015 collaborating online 0.024 * 0.009 developing and designing online content -0.006 0.014 using social media -0.047 *** 0.007 creating a video or audio production -0.045 * 0.015 media education topics addressed at school impact on economics and work 0.007 0.024 impact on democracy -0.049 0.030 impact on the environment 0.043 0.025 impact on health (physical and mental) 0.015 0.025 impact on social relationships 0.017 0.025 reliability of online information 0.070 ** 0.024 dangers of the internet 0.111 *** 0.024 note. n = 6108, groups: schoolcode. * p < .05, ** p < .01, *** p < .001 consoli 127 | f l r 4. limitations there are a number of limitations to this study. although the sample size is large and several measures have been taken to ensure a reasonable degree of representativeness of the data, the data do not perfectly represent the demographic landscape, background variables were missing many entries, and self-selection mechanisms in data collection could not be excluded entirely. caution should, therefore, be exercised in generalizing these results. furthermore, when looking at other countries, it is crucial to recognize that differences in political context and education systems may affect the results. the cross-sectional design of this study limited the inference of causality from the relationships examined. although it is reasonable to assume that students have limited influence over the curricula and the proposed classroom activities, it is crucial to consider that participatory teaching styles may encourage studentinitiated discussions and activities in certain situations. nonetheless, kahne and bowyer’s (2019) longitudinal study provided some initial empirical evidence that digital media activities in schools influence students’ online civic engagement. future research should further explore the nature of this relationship. moreover, this study did not account for several critical variables considered key predictors of (online) civic engagement and respectful behavior, such as political interest and parents’ level of activism (as done, for example, by ekström & östman, 2015; kahne & bowyer, 2019; östman, 2012 ). future studies should also consider these variables to obtain more accurate and detailed insights into the mechanisms that lead young people to develop their (online) civic or political engagement. furthermore, due to the different organizational structures in the schools, it was not possible to collect data regarding the class level. this limitation prevented the inclusion of the class level in the multilevel multiple regression analysis, which could have provided more profound insights into the relationships examined. data reliability is another concern in this study, as selfreported measures potentially subject to social desirability bias were employed, particularly regarding the assessment of online respectful behavior. furthermore, only a few items were used to measure online civic engagement and respectful behavior. the focus on highly active forms of civic engagement, such as content creation, might exclude other less active but still relevant aspects of online civic engagement, such as simply liking or sharing political posts, which were, for example, considered in the iccs (schulz et al., 2023). studies with more detailed or differently focused operationalizations of online engagement and respectful behavior might yield different conclusions. moreover, the methodology employed in this study limited the capacity to draw significant conclusions about the qualities of the studied activities and behaviors. future qualitative or mixed-methods studies could, for example, explore the concrete activities engaged in by students in their civic engagement (soep, 2014) and whether respectful behavior is based on students’ deep-seated ethical beliefs or is simply a matter of blind obedience and adherence to digital etiquette norms. another shortcoming of this study is that the operationalization of online civic engagement focuses only on expressive activities, such as publishing posts. scholars have pointed out that it is crucial to consider activities such as conscientiously boycotting specific platforms as negative forms of online civic engagement (casemajor et al., 2015; krutka et al., 2021; lutz & hoffmann, 2017). future studies could also consider these forms of online civic engagement. despite these limitations, the strengths of the study include its high ecological validity, having collected data from everyday real-world settings, and its large sample size. this study is also unique in comparing a wide variety of school-based digital media practices and considering two different output variables. consoli 128 | f l r 5. discussions the study shows that swiss upper secondary schools put little emphasis on media education and digital citizenship education. although the average use of computers in class ranged from moderate to high, the students encountered few activities aligned with a creative-expressive approach to media education that would allow them to develop their digital engagement literacy and, consequently, their online civic engagement (ekström & östman, 2015; gotlieb & sarge, 2021; kahne & bowyer, 2019). these findings are consistent with those of antonietti et al. (2023; 2024), who found that digital media in swiss upper secondary schools are mainly used to support teacher-centred, transmissive teaching styles rather than constructivist, studentcentred approaches, where students are encouraged to actively construct their knowledge or work on creative and collaborative projects. the limited prevalence of these activities represents a missed opportunity to promote online (civic) engagement and engaging learning opportunities for young people. regarding the media education topics addressed in school, our findings reveal that topics related to a protective/preventive media education approach (i.e., the dangers of the internet) are the most prevalent in swiss upper secondary schools, followed by topics related to a critical approach (i.e., teaching students to evaluate the reliability of online information). however, although these activities and approaches are widespread, many students still do not encounter these topics. for example, 44.5% of students reported that they did not address the topic of online information reliability; this proportion of students is highly problematic, considering the importance of media in shaping public opinion and given the higher rate of “news deprivation” (i.e., the underconsumption of professionally produced news in line with quality standards) among young people in switzerland (eisenegger & vogler, 2022). additionally, the results show that addressing the impact of digital technologies on democracy, the economy, and the environment is uncommon in upper secondary schools. notably, the impact of digital technologies on democracy, a topic that aligns with a critical or civic media education approach, was encountered by only a minority of students (15.6%). these findings are concerning, as they indicate that civic education, if students experience it, does not adequately consider developments related to digital technologies, and that media education, if students experience it, does not consider the political impact of digital technologies. interestingly, schmitz et al. (2024) found that teachers reported addressing media education topics on class much more frequently than what was perceived by students. these findings underline the importance of considering the students' perspective and may suggest that teachers only address these topics in some classes. it is also important to note that the high standard deviations of the media education activity variables indicated that the students’ responses varied widely. these results may be due to the fact that in swiss upper secondary schools, both media education and civic education are mostly considered cross-curricular areas that should be integrated into all subjects, but whose concrete implementation is often treated very vaguely in the curricula and thus left to the initiative of individual teachers (waldis & ziegler, 2019). regarding the prediction of online civic engagement, the analysis revealed that only 0.9% of the variance in online civic engagement is attributable to school-level differences, suggesting that the vast majority of the variance (99.1%) is due to factors within schools (e.g., class-level factors and teacher variables) or other variables not considered by model 0. this suggests that schools did not significantly promote online civic engagement at the institutional level. if students have the opportunity to experience digital citizenship education, it is likely to occur through the initiative of individual teachers or the students themselves. gender and school type emerged as significant predictors of online civic engagement. identifying as ‘female’ and attending general or specialized baccalaureate education were associated with lower levels of online civic engagement, whereas identifying as ‘other’/’non-binary’ and higher maternal education correlated with higher levels. the greater online civic engagement among vocational students may suggest that these programs’ more practical and hands-on orientation could promote a more ‘actualizing’ form of civic engagement (bennett et al., 2009). by contrast, general and specialized baccalaureate programs prepare students for higher education and for assuming societal responsibilities and may, therefore, promote a more ‘dutiful’ form of civic engagement (bennett et al., 2009). the increased online engagement of students who identify as ‘other’/’non-binary’ may reflect their use of online platforms to explore and express their identity, consoli 129 | f l r to connect with and seek support from the lgbtqia+ community, and to advocate for their rights (fisher et al., 2024; mcconnell et al., 2017). the fact that students who identified as female tended to be less politically engaged online than male students resonates with the findings of several other studies (grasso & smith, 2022; östman, 2012; schulz et al., 2023). these findings—as well as scholarly work on gender and digital citizenship—highlight the importance of adopting gender-sensitive digital citizenship education strategies (heath, 2018; henry et al., 2022; m. johnson, 2015). regarding media activities in classrooms, the general frequency of computer use at school showed no significant association with civic engagement, contrary to creative and expressive media practices (e.g., publishing work online, online collaboration, and developing online content), which significantly and positively predicted online civic engagement. these results align with those of khane and bowyer (2019) and the theoretical insights by gottlieb and segre (2021), underscoring the importance of engaging students in user-generated content to foster their (online) civic engagement. the results of this study also highlight both opportunities and challenges of online civic engagement: on the one hand, digital participation seems to be an opportunity for some minorities (e.g., students who identify with the gender “other”); on the other hand, the results show that inequalities in the real world (e.g., the impact of gender and mother’s educational level) are also reflected in these digital environments (grasso & smith, 2022; grasso & giugni, 2022; van deursen & van dijk, 2014). regarding the prediction of respectful online behavior, the icc indicated that 3% of the variance in respectful online behavior could be attributed to differences between schools, although the specific contributing factors remain unclear. this explained variance might be influenced partly by the composition of the student body, such as gender proportions and other demographic characteristics. further investigations are needed to understand how school-level dynamics influence students’ online behavior. gender again played a significant role in respectful online behavior, with females exhibiting more respect. this finding is consistent with prior research (jones & mitchell, 2016) suggesting that socialization processes influence online behaviors. females showed significantly more respectful behavior online compared to males. by contrast, students identifying as ‘other’/’non-binary’ tended to exhibit less respectful behavior, possibly due to higher exposure to online conflicts, such as cyberbullying or hate speech (mcconnell et al., 2017) and, consequently, a higher possibility of becoming involved in these conflicts (walrave & heirman, 2011). concerning young women and respectful behavior, scholars (heath, 2018; m. johnson, 2015) have highlighted how digital citizenship education, which focuses primarily on respectful behavior and safety rather than empowerment and justice-oriented participation, may perpetuate gender stereotypes and inequalities. the norms around ‘appropriate’ and ‘responsible’ online behavior may inadvertently silence women and girls, who are often socialized to be non-confrontational and seen as more vulnerable and in need of protection than young men. this could limit their (online) civic engagement, thus reinforcing existing gender dynamics and potentially exacerbating digital inequities. both studies (heath, 2018; m. johnson, 2015) emphasized the need for digital citizenship frameworks and education approaches that encourage critical thinking, inclusivity, empowerment, and active justice-oriented participation. speaking one of the national languages at home and attending a general or specialized baccalaureate type of school, compared to not speaking a national language at home and attending a vocational school, showed a positive association with respectful online behavior. these results point to the complex interplay between the educational environment, socio-economic background, and possible experiences of discrimination or racism that could influence online interaction. future studies on these aspects are needed to obtain more interpretable results. the results also highlighted that various media activities affect online respectful behavior differently. while the frequency of computer use and online collaboration were positively and significantly associated with respectful behavior, creative and expressive activities (e.g., using social media and creating videos) significantly and negatively predicted its variance. again, it can be hypothesized that students who use social consoli 130 | f l r media more frequently and who are given the opportunity to express themselves online through digital creations are more likely to encounter harm, discrimination, and hate speech (livingstone & smith, 2014; mcconnell et al., 2017; park et al., 2014) and, therefore, more likely to become involved in it (walrave & heirman, 2011). to address this issue, teachers and practitioners could, for example, address the topic of hate speech in their lessons and teach strategies to counter and reduce hate speech effectively through empathybased counterspeech (hangartner et al., 2021). to sustainably promote respectful behavior, approaches addressing values such as respect, tolerance, openness, empathy, curiosity, and non-violence are crucial (berkowitz & bier, 2004; 2014; kundu, 2021; nucci, 2001). the positive and meaningful relationship between collaborative online activities and online behavior is noteworthy and suggests, in line with findings from other studies, that collaborative activities can be an opportunity to promote cooperative, pro-social, and respectful behavior (d. w. johnson & johnson, 2009). the results also highlight the importance of educational content that addresses ‘dangers of the internet’ and ‘reliability of online information’, which significantly and positively predicted some variance in respectful online behavior. this suggests that increased awareness of online risks and a deeper understanding of media content and mechanisms can lead to more thoughtful online engagement. however, as highlighted by livingstone (2014), it is also important to consider the opportunities that young people encounter online. in digital spaces, as in the real world, opportunities and risks often go hand in hand. overall, the findings of this study highlight the challenges that teachers and practitioners face in promoting online civic engagement and respectful behavior among students. while media-based expressive activities are essential for promoting students’ engagement in online (civic) activities, they can lead to disrespectful interactions if not properly facilitated. at the same time, a protective approach to media education that focuses on online dangers and risks seems to promote respectful online behavior among students but may also prevent them from becoming highly engaged online. therefore, researchers, teachers, policymakers, and practitioners are challenged to find approaches that simultaneously promote respectful behavior and civic engagement online. furthermore, the fact that the additional variance explained by model 2 is greater in the prediction of online civic engagement and that the variance explained by background variables is greater in the prediction of respectful online behavior suggests that school-based media education practices are more likely to have an impact on civic engagement than on respectful behavior. finally, the findings stimulate reflection on the concept of digital citizenship, highlighting its normative rather than empirical nature. the observed negative correlation between online civic engagement and respectful online behavior may be due to different profiles among young people, with some being more respectful but less engaged, while others are more engaged but less respectful (similar to what was theorized by bennett et al., 2009). future studies should aim to empirically identify these different profiles of digital citizenship and explore interventions that can balance respect and engagement in digital spaces. by considering different types of media education activities as predictors (simple frequency of computer use, different digital media activities that have the potential to promote online engagement, and media education topics covered at school) and two dimensions of digital citizenship as outcome variables (online civic engagement and respectful behaviour), this study contributes to the existing body of research on this topic (kahne & bowyer, 2019; schulz et al., 2023) by providing a more nuanced perspective on the complex mechanisms underlying young people's development of digital citizenship. in particular, it provides valuable insights into the challenges faced by practitioners in promoting digital citizenship in schools. consoli 131 | f l r keypoints swiss upper secondary schools place little emphasis on digital citizenship education current media education practices focus primarily on protective or critical media education approaches creative school-based digital media activities predict online civic engagement but not respectful behavior addressing topics such as ‘dangers of the internet’ at school predicts respectful online behavior but not online civic engagement. online civic engagement and respectful online behavior are negatively correlated gender, language spoken at home, type of schooling and parental education significantly affect students' digital citizenship practices acknowledgements this study was founded by the swiss national science foundation (snsf) and related to project grant number 187277. there are no conflicts of interest. this study was non-interventional, and it complies with the ethical standards of the swiss academy of social sciences and humanities. i want to thank the digitrasii (digital transformation in upper secondary schools) team for their valuable contributions to the data collection process, including dominik petko, philipp gonon, alberto cattaneo, chiara antonietti and maria-luisa schmitz. i am also profoundly grateful to the participating schools the administrators, teachers and students for their generous time, cooperation and openness throughout the study. i would also like to thank monika waldis for her insightful feedback on an early version of this study, discussed during a symposium at the 2023 annual conference of the swiss society for research in education. 6. references antonietti, c., schmitz, m. l., consoli, t., cattaneo, a., gonon, p., & petko, d. (2023). development and validation of the icap technology scale to measure how teachers integrate technology into learning activities. computers & education, 192, 104648. https://doi.org/10.1016/j.compedu.2022.104648 antonietti, c., consoli, t., schmitz, m. l., cattaneo, a., gonon, p., & petko, d. (2024). digital constructivists, activators or presenters? different profiles of technology integration among swiss upper secondary school teachers. computers & education, 105225. https://doi.org/10.1016/j.compedu.2024.105225 banerjee, s. c., & greene, k. (2006). analysis versus production: adolescent cognitive and attitudinal responses to antismoking interventions. journal of communication, 56(4), 773–794. https://doi.org/10.1111/j.1460-2466.2006.00319.x bennett, w. l., wells, c., & rank, a. (2009). young citizens and civic learning: two paradigms of citizenship in the digital age. citizenship studies, 13(2), 105–120. https://doi.org/10.1080/13621020902731116 berkowitz, m. w., & bier, m. c. (2004). research-based character education. the annals of the american academy of political and social science, 591(1), 72–85. https://doi.org/10.1177/0002716203260082 bohrnstedt, g. w., & carter, t. m. (1971). robustness in regression analysis. sociological methodology, 3, 118. https://doi.org/10.2307/270820 https://doi.org/10.1016/j.compedu.2022.104648 https://doi.org/10.1016/j.compedu.2024.105225 https://doi.org/10.1111/j.1460-2466.2006.00319.x https://doi.org/10.1080/13621020902731116 https://doi.org/10.1177/0002716203260082 https://doi.org/10.2307/270820 consoli 132 | f l r botturi, l. (2019). digital and media literacy in pre-service teacher education: a case study from switzerland. nordic journal of digital literacy, 14(3-4), 147–163. https://doi.org/10.18261/issn.1891-943x-2019-03-04-05 boulianne, s. (2015). social media use and participation: a meta-analysis of current research. information, communication & society, 18(5), 524–538. https://doi.org/10.1080/1369118x.2015.1008542 boulianne, s., & theocharis, y. (2020). young people, digital media, and engagement: a meta-analysis of research. social science computer review, 38(2), 111–127. https://doi.org/10.1177/0894439318814190 bowyer, b., & kahne, j. (2020). the digital dimensions of civic education: assessing the effects of learning opportunities. journal of applied developmental psychology, 69, 101162. https://doi.org/10.1016/j.appdev.2020.101162 brinda, t., brüggen, n., diethelm, i., knaus, t., kommer, s., kopf, c., missomelius, p., leschke, r., tilemann, f., & weich, a. (2020). frankfurt-dreieck zur bildung in der digital vernetzten welt. ein interdisziplinäres modell. kopaed. https://doi.org/10.25656/01:22117 bundesamt für statistik. (2023). bevölkerung nach migrationsstatus. retrived december 17, 2024, from https://www.bfs.admin.ch/bfs/de/home/statistiken/bevoelkerung/migration-integration/nachmigrationsstatuts.html bundesamt für statistik. (2024) sekundarstufe ii: statistik der lernenden. retrived january 2, 2024 from https://www.bfs.admin.ch/bfs/de/home/statistiken/bildung-wissenschaft/personenausbildung/sekundarstufe-ii.html burn, a., & durran, j. (2007). media literacy in schools: practice, production and progression. sage publications. https://doi.org/10.4135/9781446213629 casemajor, n., couture, s., delfin, m., goerzen, m., & delfanti, a. (2015). non-participation in digital media: toward a framework of mediated political action. media, culture & society, 37(6), 850–866. https://doi.org/10.1177/0163443715584098 chen, l. l., mirpuri, s., rao, n., & law, n. (2021). conceptualization and measurement of digital citizenship across disciplines. educational research review, 33, 100379. https://doi.org/10.1016/j.edurev.2021.100379 choi, m. (2016). a concept analysis of digital citizenship for democratic citizenship education in the internet age. theory & research in social education, 44(4), 565–607. https://doi.org/10.1080/00933104.2016.1210549 ciip. (2011). plan d’études romand. retrieved january 2, 2024 from https://portail.ciip.ch/per/domains deci, e. l., & ryan, r. m. (1985). intrinsic motivation and self-determination in human behavior. berlin: springer science & business media. https://doi.org/10.1007/978-1-4899-2271-7 d-edk. (2014). lehrplan 21: grundlagen. retrived january 2, 2024 from https://vfe.lehrplan.ch/container/v_fe_grundlagen.pdf edk. (2017). rahmenlehrplan für die maturitätsschulen: informatik. retrived december 17, 2024 from https://edudoc.ch/record/131917/files/rlp_inf_2017_d.pdf eisenegger, m., & vogler, d. (2022). hauptbefunde: zunahme der news-deprivation mit negativen folgen für den demokratischen prozess. in: forschungszentrum öffentllichkeit und gesellschaft (fög), jahrbuch qualität der medien 2022. basel: schwabe verlag 1-18. https://doi.org/10.5167/uzh224549 ekström, m., & östman, j. (2015). information, interaction, and creative production. communication research, 42(6), 796–818. https://doi.org/10.1177/0093650213476295 https://doi.org/10.18261/issn.1891-943x-2019-03-04-05 https://doi.org/10.1080/1369118x.2015.1008542 https://doi.org/10.1177/0894439318814190 https://doi.org/10.1016/j.appdev.2020.101162 https://doi.org/10.25656/01:22117 https://www.bfs.admin.ch/bfs/de/home/statistiken/bevoelkerung/migration-integration/nach-migrationsstatuts.html https://www.bfs.admin.ch/bfs/de/home/statistiken/bevoelkerung/migration-integration/nach-migrationsstatuts.html https://www.bfs.admin.ch/bfs/de/home/statistiken/bildung-wissenschaft/personen-ausbildung/sekundarstufe-ii.html https://www.bfs.admin.ch/bfs/de/home/statistiken/bildung-wissenschaft/personen-ausbildung/sekundarstufe-ii.html https://doi.org/10.4135/9781446213629 https://doi.org/10.1177/0163443715584098 https://doi.org/10.1016/j.edurev.2021.100379 https://doi.org/10.1080/00933104.2016.1210549 https://portail.ciip.ch/per/domains https://doi.org/10.1007/978-1-4899-2271-7 https://v-fe.lehrplan.ch/container/v_fe_grundlagen.pdf https://v-fe.lehrplan.ch/container/v_fe_grundlagen.pdf https://edudoc.ch/record/131917/files/rlp_inf_2017_d.pdf https://doi.org/10.5167/uzh-224549 https://doi.org/10.5167/uzh-224549 https://doi.org/10.1177/0093650213476295 consoli 133 | f l r enders, c. k., & tofighi, d. (2007). centering predictor variables in cross-sectional multilevel models: a new look at an old issue. psychological methods, 12(2), 121–138. https://doi.org/10.1037/1082989x.12.2.121 european commission. (2019). 2nd survey of schools: ict in education.: technical report. https://data.europa.eu/doi/10.2759/035445 fisher, c. b., tao, x., & ford, m. (2024). social media: a double-edged sword for lgbtq+ youth. computers in human behavior, 156, 108194. https://doi.org/10.1016/j.chb.2024.108194 fox, j., & weisberg, s. (2020). car . https://doi.org/10.32614/cran gallucci, m. (2019). gamlj. https://gamlj.github.io/ gfs.bern. (2018). easyvote-politikmonitor 2018: problem alltagsbezug. https://cockpit.gfsbern.ch/de/cockpit/easyvote-politikmonitor-2018/ gfs.bern. (2020). easyvote-politikmonitor 2020: krisen und bewegungen aktivieren schweizer jugend. https://cockpit.gfsbern.ch/de/cockpit/easyvote-politikmonitor-2020-2/ gfs.bern. (2023). easyvote-politikmonitor 2023: die politisierungswelle der letzten jahren flacht zunehmend ab. https://dsj.ch/wp-content/uploads/2024/04/dsj_jugend_und_politikmonitor_2023_schlussbericht.pdf gotlieb, m. r., & sarge, m. a. (2021). civic learning and self-determination: a model of user-generated content and civic readiness among actualizing citizens. communication theory, 31(1), 127–149. https://doi.org/10.1093/ct/qtaa032 grassi, e. f. g., portos, m., & felicetti, a. (2024). young people's attitudes towards democracy and political participation: evidence from a cross-european study. government and opposition, 59(2), 582-604. https://doi.org/10.1017/gov.2023.16 grasso, m., & giugni, m. (2022). intra-generational inequalities in young people’s political participation in europe: the impact of social class on youth political engagement. politics, 42(1), 13-38. https://doi.org/10.1177/02633957211031742 grasso, m., & smith, k. (2022). gender inequalities in political participation and political engagement among young people in europe: are young women less politically engaged than young men?. politics, 42(1), 39-57. https://doi.org/10.1177/02633957211028813 guggemos, j., & seufert, s. (2021). teaching with and teaching about technology – evidence for professional development of in-service teachers. computers in human behavior, 115, 106613. https://doi.org/10.1016/j.chb.2020.106613 hangartner, d., gennaro, g., alasiri, s., bahrich, n., bornhoft, a., boucher, j., demirci, b. b., derksen, l., hall, a., jochum, m., munoz, m. m., richter, m., vogel, f., wittwer, s., wüthrich, f., gilardi, f., & donnay, k. (2021). empathy-based counterspeech can reduce racist hate speech in a social media field experiment. proceedings of the national academy of sciences of the united states of america, 118(50). https://doi.org/10.1073/pnas.2116310118 heath, m. k. (2018). what kind of (digital) citizen? a between-studies analysis of research and teaching for democracy. international journal of information and learning technology, 35(5), 342–365. https://doi.org/10.1108/ijilt-06-2018-0067 henry, n., vasil, s., & witt, a. (2022). digital citizenship in a global society: a feminist approach. feminist media studies, 22(8), 1972–1989. https://doi.org/10.1080/14680777.2021.1937269 hobbs, r. (1999). the seven great debates in the media literacy movement. retrieved december 2017, 2024 from https://eric.ed.gov/?id=ed439454 hurrelmann, k., quenzel, g., schneekloth, u., leven, i., albert, m., utzmann, h., & wolfert, s. (2019). jugend 2019: eine generation meldet sich zu wort. 18. shell-jugendstudie. hamburg. https://doi.org/10.1037/1082-989x.12.2.121 https://doi.org/10.1037/1082-989x.12.2.121 https://data.europa.eu/doi/10.2759/035445 https://doi.org/10.1016/j.chb.2024.108194 https://doi.org/10.32614/cran https://gamlj.github.io/ https://cockpit.gfsbern.ch/de/cockpit/easyvote-politikmonitor-2018/ https://cockpit.gfsbern.ch/de/cockpit/easyvote-politikmonitor-2020-2/ https://dsj.ch/wp-content/uploads/2024/04/dsj_jugend-_und_politikmonitor_2023_schlussbericht.pdf https://dsj.ch/wp-content/uploads/2024/04/dsj_jugend-_und_politikmonitor_2023_schlussbericht.pdf https://doi.org/10.1093/ct/qtaa032 https://doi.org/10.1017/gov.2023.16 https://doi.org/10.1177/02633957211031742 https://doi.org/10.1177/02633957211028813 https://doi.org/10.1016/j.chb.2020.106613 https://doi.org/10.1073/pnas.2116310118 https://doi.org/10.1108/ijilt-06-2018-0067 https://doi.org/10.1080/14680777.2021.1937269 https://eric.ed.gov/?id=ed439454 consoli 134 | f l r hutson, e., kelly, s., & militello, l. k. (2018). systematic review of cyberbullying interventions for youth and parents with implications for evidence-based practice. worldviews on evidence-based nursing, 15(1), 72–79. https://doi.org/10.1111/wvn.12257 international society for technology in education (2007). digital citizenship in school. mike ribble and gerald baley. international society for technology in education (2011). digital citizenship in schools (2nd ed.). mike ribble. isin, e. f., & ruppert, e. (2020). being digital citizens (2nd ed.). rowman & littlefield. the jamovi project. (2022). jamovi (version 2.3). https://www.jamovi.org/ jenkins, h. (2009). confronting the challenges of participatory culture: media education for the 21st century. cambridge: the mit press. http://library.oapen.org/handle/20.500.12657/26083 jenkins, h., shresthova, s., gamber-thompson, l., kligler-vilenchik, n., & zimmerman, a. m. (2016). by any media necessary: the new youth activism. connected youth and digital futures. new york university press. https://doi.org/10.18574/nyu/9781479829712.001.0001 johnson, d. w., & johnson, r. t. (2009). an educational psychology success story: social interdependence theory and cooperative learning. educational researcher, 38(5), 365–379. https://doi.org/10.3102/0013189x09339057 johnson, m. (2015). chapter xiii. digital literacy and digital citizenship: approaches to girls’ online experiences. in j. bailey & v. steeves (eds.), egirls, ecitizens. les presses de l’université d’ottawa. retrived december 17, 2024, from https://books.openedition.org/uop/520 jones, l. m., & mitchell, k. j. (2016). defining and measuring youth digital citizenship. new media & society, 18(9), 2063–2079. https://doi.org/10.1177/1461444815577797 kahne, j., & bowyer, b. (2019). can media literacy education increase digital engagement in politics? learning, media and technology, 44(2), 211–224. https://doi.org/10.1080/17439884.2019.1601108 kahne, j., hodgin, e., & eidman-aadahl, e. (2016). redesigning civic education for the digital age: participatory politics and the pursuit of democratic engagement. theory & research in social education, 44(1), 1–35. https://doi.org/10.1080/00933104.2015.1132646 keating, a., & melis, g. (2017). social media and youth political engagement: preaching to the converted or providing a new voice for youth? the british journal of politics and international relations, 4(19), 877–894. https://doi.org/10.1177/1369148117718461 krutka, d. g., & carpenter, j. p. (2017). digital citizenship in the curriculum. educational leadership, 75(3), 50–55. retrived december 17, 2024 from https://static1.squarespace.com/static/60cf7350af4a9a05e12ced8f/t/622e73630696395c7c9fd73c/164 7211364240/digitalcitizenshipinthecurriculum+krutkadg+carpenterjp+edleadership+2017.pdf krutka, d. g., smits, r. m., & willhelm, t. a. (2021). don’t be evil: should we use google in schools? techtrends, 65(4), 421–431. https://doi.org/10.1007/s11528-021-00599-4 kundu, v. (2021). integrating nonviolent communication in pedagogies of media literacy education. in d. frau-meigs, s. kotilainen, m. pathak-shelat, m. hoechsmann, & s. r. poyntz (eds.), global handbooks in media and communication research. the handbook of media education research (pp. 141–154). wiley blackwell. https://doi.org/10.1002/9781119166900.ch11 livingstone, s. (2014). developing social media literacy: how children learn to interpret risky opportunities on social network sites. communications, 39(3), 283–303. https://doi.org/10.1515/commun-20140113 livingstone, s., & smith, p. k. (2014). annual research review: harms experienced by child users of online and mobile technologies: the nature, prevalence and management of sexual and aggressive risks in https://doi.org/10.1111/wvn.12257 https://www.jamovi.org/ http://library.oapen.org/handle/20.500.12657/26083 https://doi.org/10.18574/nyu/9781479829712.001.0001 https://doi.org/10.3102/0013189x09339057 https://books.openedition.org/uop/520 https://doi.org/10.1177/1461444815577797 https://doi.org/10.1080/17439884.2019.1601108 https://doi.org/10.1080/00933104.2015.1132646 https://doi.org/10.1177/1369148117718461 https://static1.squarespace.com/static/60cf7350af4a9a05e12ced8f/t/622e73630696395c7c9fd73c/1647211364240/digitalcitizenshipinthecurriculum+krutkadg+carpenterjp+edleadership+2017.pdf https://static1.squarespace.com/static/60cf7350af4a9a05e12ced8f/t/622e73630696395c7c9fd73c/1647211364240/digitalcitizenshipinthecurriculum+krutkadg+carpenterjp+edleadership+2017.pdf https://doi.org/10.1007/s11528-021-00599-4 https://doi.org/10.1002/9781119166900.ch11 https://doi.org/10.1515/commun-2014-0113 https://doi.org/10.1515/commun-2014-0113 consoli 135 | f l r the digital age. journal of child psychology and psychiatry, and allied disciplines, 55(6), 635–654. https://doi.org/10.1111/jcpp.12197 lutz, c., & hoffmann, c. p. (2017). the dark side of online participation: exploring non-, passive and negative participation. information, communication & society, 20(6), 876–897. https://doi.org/10.1080/1369118x.2017.1293129 martens, h., & hobbs, r. (2015). how media literacy supports civic engagement in a digital age. atlantic journal of communication, 23(2), 120–137. https://doi.org/10.1080/15456870.2014.961636 mcconnell, e. a., clifford, a., korpak, a. k., phillips, g., & birkett, m. (2017). identity, victimization, and support: facebook experiences and mental health among lgbtq youth. computers in human behavior, 76, 237–244. https://doi.org/10.1016/j.chb.2017.07.026 mihailidis, p., & thevenin, b. (2013). media literacy as a core competency for engaged citizenship in participatory democracy. american behavioral scientist, 57(11), 1611–1622. https://doi.org/10.1177/0002764213489015 nakagawa, s., & schielzeth, h. (2013). a general and simple method for obtaining r2 from generalized linear mixed‐effects models. methods in ecology and evolution, 4(2), 133–142. https://doi.org/10.1111/j.2041-210x.2012.00261.x nucci, l. p. (2001). education in the moral domain (1. publ). cambridge univ. press. https://doi.org/10.1017/cbo9780511605987 oecd. (2023). pisa 2022 results (volume ii): learning during and from disruption. pisa, oecd publishing. https://doi.org/10.1787/a97db61c-en oser, j., & boulianne, s. (2020). reinforcement effects between digital media use and political participation: a meta-analysis of repeated-wave panel data. public opinion quarterly, 84(s1), 355–365. https://doi.org/10.1093/poq/nfaa017 östman, j. (2012). information, expression, participation: how involvement in usergenerated content relates to democratic engagement among young people. new media & society, 14(6), 1004–1021. https://doi.org/10.1177/1461444812438212 paccagnella, o. (2006). centering or not centering in multilevel models? the role of the group mean and the assessment of group effects. evaluation review, 30(1), 66–85. https://doi.org/10.1177/0193841x05275649 park, s., na, e.‑y., & kim, e. (2014). the relationship between online activities, netiquette and cyberbullying, children and youth services review, 42, 74–81. https://doi.org/10.1016/j.childyouth.2014.04.002 parth, a. m., weiss, j., firat, r., & eberhardt, m. (2020). “how dare you!”—the influence of fridays for future on the political attitudes of young adults. frontiers in political science, 2, 611139. https://doi.org/10.3389/fpos.2020.611139 petko, d., antonietti, c., schmitz, m. l., consoli, t., gonon, p., & cattaneo, a. (2022). digitale transformation der sekundarstufe ii: erste ergebnisse einer repräsentativen bestandsaufnahme in der schweiz. gymnasium helveticum, 76(5), 20-21. https://doi.org/10.5167/uzh-223928 r core team. (2021). r (version 4.1) . https://cran.r-project.org/ revelle, w. (2019). psych . https://cran.r-project.org/web/packages/psych/index.html ribble, m. (2012). digital citizenship for educational change. kappa delta pi record, 48(4), 148–151. https://doi.org/10.1080/00228958.2012.734015 ribble, m., & miller, t. n. (2013). educational leadership in an online world: connecting students to technology responsibly, safely, and ethically. journal of asynchronous learning networks, 17(1), 137-145. https://doi.org/10.1111/jcpp.12197 https://doi.org/10.1080/1369118x.2017.1293129 https://doi.org/10.1080/15456870.2014.961636 https://doi.org/10.1016/j.chb.2017.07.026 https://doi.org/10.1177/0002764213489015 https://doi.org/10.1111/j.2041-210x.2012.00261.x https://doi.org/10.1017/cbo9780511605987 https://doi.org/10.1787/a97db61c-en https://doi.org/10.1093/poq/nfaa017 https://doi.org/10.1177/1461444812438212 https://doi.org/10.1177/0193841x05275649 https://doi.org/10.1016/j.childyouth.2014.04.002 https://doi.org/10.3389/fpos.2020.611139 https://doi.org/10.5167/uzh-223928 https://cran.r-project.org/ https://cran.r-project.org/web/packages/psych/index.html https://doi.org/10.1080/00228958.2012.734015 consoli 136 | f l r rhodes, s. c. (2022). filter bubbles, echo chambers, and fake news: how social media conditions individuals to be less critical of political misinformation. political communication, 39(1), 1-22. https://doi.org/10.1080/10584609.2021.1910887 ryan, r. m., & deci, e. l. (2017). self-determination theory: basic psychological needs in motivation, development, and wellness. the guilford press. https://doi.org/10.1521/978.14625/28806 schielzeth, h., & nakagawa, s. (2013). nested by design: model fitting and interpretation in a mixed model era. methods in ecology and evolution, 4(1), 14–24. https://doi.org/10.1111/j.2041210x.2012.00251.x schmitz, m. l., consoli, t., antonietti, c., cattaneo, a., gonon, p., & petko, d. (2024). why do some teachers teach media literacy while others do not? exploring predictors along the “will, skill, tool, pedagogy” model. computers in human behavior, 151, 108004. https://doi.org/10.1016/j.chb.2023.108004 schmidt, a. f., & finan, c. (2018). linear regression and the normality assumption. journal of clinical epidemiology, 98, 146–151. https://doi.org/10.1016/j.jclinepi.2017.12.006 schulz, w., ainley, j., fraillon, j., losito, b., agrusti, g., valeria, d., & friedman, t. (2023). education for citizenship in times of global challenge. iea international civic and citizenship education study, 2024-02. retrieved december 17, 2024 from https://www.iea.nl/sites/default/files/2023-12/iccs2022-international-report.pdf soep, e. (2014). participatory politics: next-generation tactics to remake public spheres. the john d. and catherine t. macarthur foundation reports on digital media and learning. the mit press. retrieved december 17, 2024 from http://library.oapen.org/handle/20.500.12657/26056 süss, d., lampert, c., & wijnen, c. w. (2008). medienpädagogik: ein studienbuch zur einführung. studienbücher zur kommunikationsund medienwissenschaft. vs verlag für sozialwissenschaften. https://doi.org/10.1007/978-3-531-92142-6 theocharis, y. (2015). the conceptualization of digitally networked participation. social media + society, 1(2), 205630511561014. https://doi.org/10.1177/2056305115610140 theocharis, y., & van deth, j. w. (2018a). the continuous expansion of citizen participation: a new taxonomy. european political science review, 10(1), 139–163. https://doi.org/10.1017/s1755773916000230 theocharis, y., & van deth, j. w. (2018b). political participation in a changing world: conceptual and empirical challenges in the study of citizen engagement. routledge. https://doi.org/10.4324/9780203728673 van deursen, a. j., & van dijk, j. a. (2014). the digital divide shifts to differences in usage. new media & society, 16(3), 507–526. https://doi.org/10.1177/1461444813487959 vasu, e. s., & elmore, p. b. (1975). the effect of multicollinearity and the violation of the assumption of normality on the testing of hypotheses in regression analysis. retrieved december 17, 2024 from https://eric.ed.gov/?id=ed106341 waldis, m., & ziegler, b. (2019). politische bildung in der halbdirekten demokratie der schweiz. in n. braun binder, l. p. feld, p. m. huber, k. poier, & f. wittreck (eds.), jahrbuch für direkte demokratie 2018 (pp. 42–66). nomos verlagsgesellschaft mbh & co. kg. https://doi.org/10.5771/9783748904557-42 walrave, m., & heirman, w. (2011). cyberbullying: predicting victimization and perpetration. children & society, 25(1), 59–72. https://doi.org/10.1111/j.1099-0860.2009.00260.x weiss, j. (2020). what is youth political participation? literature review on youth political participation and political attitudes. frontiers in political science, 2, 1. https://doi.org/10.3389/fpos.2020.00001 https://doi.org/10.1080/10584609.2021.1910887 https://doi.org/10.1521/978.14625/28806 https://doi.org/10.1111/j.2041-210x.2012.00251.x https://doi.org/10.1111/j.2041-210x.2012.00251.x https://doi.org/10.1016/j.chb.2023.108004 https://doi.org/10.1016/j.jclinepi.2017.12.006 https://www.iea.nl/sites/default/files/2023-12/iccs-2022-international-report.pdf https://www.iea.nl/sites/default/files/2023-12/iccs-2022-international-report.pdf http://library.oapen.org/handle/20.500.12657/26056 https://doi.org/10.1007/978-3-531-92142-6 https://doi.org/10.1177/2056305115610140 https://doi.org/10.1017/s1755773916000230 https://doi.org/10.4324/9780203728673 https://doi.org/10.1177/1461444813487959 https://eric.ed.gov/?id=ed106341 https://doi.org/10.5771/9783748904557-42 https://doi.org/10.1111/j.1099-0860.2009.00260.x https://doi.org/10.3389/fpos.2020.00001 consoli 137 | f l r winship, c., & radbill, l. (1994). sampling weights and regression analysis. sociological methods & research, 23(2), 230–257. https://doi.org/10.1177/0049124194023002004 wittwer, s. (2015). politische partizipation von kindern und jugendlichen in der schweiz. schweizerische arbeitsgemeinschaft der jugendverbände wu, y.‑w. b., & wooldridge, p. j. (2005). the impact of centering first-level predictors on individual and contextual effects in multilevel data analysis. nursing research, 54(3), 212. https://doi.org/10.1177/0049124194023002004 frontline learning research vol. 12 no. 4 (2024) 85 -112 issn 2295-3159 corresponding author: isabelle krummenacher, abteilung für schulund unterrichtsforschung, institut für erziehungswissenschaft, bern isabelle.krummenacher@unibe.ch doi: https://doi.org/10.14786/flr.v12i4.1211 navigating the paradox between professional challenges and teacher well-being: the role of strategies within the resilience process isabelle krummenacher1, tina hascher1 caroline mansfield2, susan beltman3julia mori1 & irene guidon1, 1 universität bern, switzerland 2 edith cowan university, australia 3 curtin university, australia article received 23 december 2022 / article revised 22 february 2024 / accepted 3 december 2024 / available online 9 december 2024 abstract teaching is an immensely complex profession that often requires managing multiple and varied professional challenges. despite these challenges, teachers tend to report moderate to high levels of well-being. this qualitative study explored this paradox by investigating the professional challenges reported and the coping strategies swiss teachers use to support their well-being. relational problem-solving was identified as a commonly used strategy when professional challenges occur. the unique contribution of this study lies in its further elucidation of the complex relationship between teachers’ professional challenges and well-being embedded in a resilience process. the results of this study provide implications for understanding how to support teachers’ resilience and to design interventions to enhance teachers’ well-being. keywords: teacher professional challenges; teacher well-being; teacher resilience; teachers; teacher strategies mailto:isabelle.krummenacher@unibe.ch krummenacher et al 86 | f l r 1. introduction teacher shortages, teacher stress, and a high attrition rate have become critical areas of concern and a focus of research because of the important societal role of teachers (borman & dowling, 2008; viac & fraser, 2020). it is well known that teaching is an immensely complex profession, particularly when job demands are high and in challenging situations. nevertheless, it is also well known that many teachers feel well and thrive in their work and, despite high job demands, maintain their well-being, passion, engagement, and commitment (hascher & waber, 2021). how can this paradox between concurrent challenges and well-being be explained? research has shown that one reason why teachers succeed and maintain their well-being despite high demands is teacher resilience (e.g. li et al., 2019). thus, resilience—defined as the process of drawing on a range of resources to navigate challenges and restore or improve well-being (ungar, 2012)—seems to be a critical factor in teachers’ professional lives. it has been found that the two dynamic and multidimensional constructs—teacher well-being and resilience—are used in various disciplines and research areas; they are often studied together (hascher et al., 2021). to better understand the relationship between resilience and well-being, the aligning wellbeing and resilience in education (aware) model was developed (hascher et al., 2021). the aware model incorporates challenges and personal and contextual resources (e.g. self-efficacy beliefs as a personal resource and support by principals as a contextual resource) and illustrates how the resilience process restores well-being. within this model, strategies are given a crucial role in supporting individuals in responding to professional challenges. however, the model does not explain how teachers cope with professional challenges or which coping strategies are applied and valuable to maintain their well-being. the frontline contribution of this study is thus (a) to explore the paradox of concurrent teacher professional challenges and teacher well-being; (b) to better understand the complex association between resilience and well-being; (c) to investigate empirically the role of strategies within a resilience process; and (d) to shed further light on some of the more nuanced strategies and processes that support well-being in teachers’ professional life. 1.1 professional challenges in teachers’ work teaching is considered a highly demanding profession, and professional challenges can be manifold (avidov-ungar, 2018). for example, teachers can face professional challenges relating to students, colleagues, time stress, or workloads (mansfield et al., 2014). research on teachers’ professional challenges can contribute to a better understanding of the profession and the requirements to succeed as a teacher. in their study, kitching et al. (2009, p. 43) suggested that frequent “little issues” are more relevant to teachers’ commitment, motivation, satisfaction, stress, burnout, and attrition than infrequent and severe issues. professional challenges can trigger the resilience process in teachers. these professional challenges can be structured within a social-ecological framework at the personal, relational, and organisational levels, because individuals develop, live, and act within numerous internally and externally interacting systems, as demonstrated by research on teacher resilience (masten, 2014). professional challenges at the organisational level include high workloads, limited time, scarce support (flores, 2006), high responsibility (johnson et al., 2014), the pressures of societal expectations (schelvis et al., 2014), and policy changes (gu & day, 2013). gu (2018) has suggested the governmental policy revisions can increase teachers’ work stress, accountability responsibilities, and complexity. at the relational level, difficult relationships with students, parents, or coworkers can lead to unpleasant experiences and professional challenges for teachers (marzano, 2003; papatraianou et al., 2018). challenging professional relationships may involve conflicts, lack of cooperation, or communication problems, which can increase stress and reduce teacher job satisfaction (beltman et al., 2022; mansfield et al., 2014). at the personal level, professional challenges include low self-efficacy (kitching et al., 2009), poor health (day & gu, 2010), and negative emotions (veronese et al., 2018). a disparity between professional expectations and actual practices (flores & day, 2006) or reluctance to seek help (fantilli & mcdougall, 2009), as well as inadequate social and emotional krummenacher et al 87 | f l r competencies, may also result in classroom relationship issues and, ultimately, professional challenges (cefai & cavioni, 2014). 1.2 teacher well-being over the past 20 years, teacher well-being has risen to the top of scientific and political agendas due to teacher shortages, the high incidence of mental and physical health problems, and the significance of teacher well-being for student academic progress and school outcomes (viac & fraser, 2020). as shown by hascher and waber (2021), teacher well-being is used as an all-encompassing concept to describe a variety of dimensions, including different aspects of burnout, stress, emotions, motivation, and health factors, as well as their interactions. they identified five different research fields (a) psychology of well-being, (b) positive psychology, (c) psychology and work organisation, (d) teacher well-being, and (e) health science (p. 6). the specific definition of teacher well-being frequently relates to the disciplinary approach of the studies. in the present study, we define teacher well-being as a positive imbalance—that is, the prevalence of positive dimensions (e.g. enjoyment of teaching, selfefficacy beliefs) over the negative dimensions (e.g. worries, physical discomfort) (hascher & hagenauer, 2011; hascher & waber, 2021). the more pronounced the contrast between positive and negative qualities, the deeper the sense of well-being. recent research has identified numerous factors that contribute to teacher well-being (for an overview, see hascher & waber, 2021), and the quantity of research is increasing. however, except during the pandemic, few studies have aimed to elucidate the processes that lead to maintaining or restoring well-being when facing challenges. primarily, cross-sectional studies have demonstrated the relationships between teacher well-being and age, gender, tenure, and concepts like self-efficacy or resilience (e.g. jones et al., 2019). studies have also confirmed the role of social support for teacher well-being (e.g. viac & fraser, 2020). professional challenges, however, may be of specific importance for their impact on teacher well-being, because they urge teachers to engage actively and react to a situation. the benefits of investigating teacher well-being have been demonstrated in prior research. for example, studies have pointed to the association of teacher well-being with more effective teaching (duckworth et al., 2009) and improved student performance (klusmann et al., 2016). also, it has been confirmed that teacher well-being promotes students’ well-being (harding et al., 2019), which can, in turn, support their academic outcomes. understanding the process that maintains well-being despite professional challenges thus seems critical. 1.3 teacher resilience as with teacher well-being, teacher resilience is a vital construct associated with teachers’ commitment, engagement, and job satisfaction (day & gu, 2014). resilience can be considered a dynamic process or an outcome that is the product of a person’s engagement with his or her environment over time. it is demonstrated by how individuals respond to challenging events (mansfield et al., 2012). teacher resilience is characterised by various positive aspects such as job satisfaction, engagement, teaching effectiveness, well-being, and a sense of professional identity (e.g. day & gu, 2014). mansfield et al. (2016) note that resilient teachers could draw on individual resources such as self-belief and optimism (e.g. day & gu, 2014) and contextual resources, such as support from colleagues (ainsworth & oldfield, 2019) while managing professional challenges. recent research, however, has illuminated the complex and context-dependent nature of resilience and the crucial role of the school context in fostering or inhibiting it (ainsworth & oldfield, 2019). although teacher resilience has been considered from various perspectives and disciplines, a common understanding can be identified regarding the definition of resilient teachers. ungar’s (2012) social-ecological perspective has been used to understand teacher resilience (mansfield et al., 2016) and, krummenacher et al 88 | f l r in particular, the interplay of both individual and multi-contextual resources in the resilience process while managing professional challenges. however, it is still necessary to comprehend how strategies are used for dealing with professional challenges and how this contributes to maintaining or restoring teacher well-being. 1.4 teacher strategies research on teacher resilience has highlighted the significance of coping strategies when teachers encounter challenging situations (parker & martin, 2009). coping strategies are “fundamental human adaptive processes” that assist teachers in managing stress and negative emotions (zimmergembeck & skinner, 2016, p. 2). typically, coping strategies focus on reducing extreme negative emotions or modifying the stressful situation that triggered them (bonanno & burton, 2013; lazarus & folkman, 1991). lazarus and folkman (1991) identified two primary coping strategies for handling difficult situations—emotion-focused and problem-focused. emotion-focused strategies are reactive and prove beneficial when control over the stressor is limited, while problem-focused strategies prove more successful in addressing the source of stress (lewis & frydenberg, 2002). prior research has indicated that problem-focused coping (e.g. planning or active coping) is positively connected with work engagement, self-efficacy for teaching, and job satisfaction (parker et al., 2012; parker & martin, 2009). in a recent study, beltman and poulton (2019) explored the strategies identified by 73 teachers and classified them into waiting (e.g. taking a breath or waiting for the next day), assessing (e.g. viewing from a distance or referring to existing literature), problem-solving (e.g. talking with colleagues), and being proactive (e.g. deciding not to take work home or pursuing hobbies). given the growing concerns regarding the challenges in the teacher profession, further research is needed to understand how teachers apply coping strategies when faced with challenging situations, as this might help to elucidate how strategies contribute to teacher well-being within a resilience process. the aware model provides a comprehensive framework to further explore the relationship between teachers’ professional challenges and their strategies. 1.5 integrating teacher resilience and well-being: the aware model although the terms well-being and resilience are often used in tandem and sometimes interchangeably, the conceptualisation of these constructs needs closer attention. one of the reasons is the dynamic, complex, and multidimensional nature of both constructs and a need for more explanation of how they are related to each other (hascher et al., 2021, p. 417). the recently introduced aware model outlines the resilience process, framed by challenges and resources at the contextual and individual levels (see figure 1). figure 1. the aware model, showing the relationship between resilience and well-being (hascher et al., 2021, p. 422). krummenacher et al 89 | f l r according to hascher et al. (2021), the resilience process (b) is initiated by an event that may have a detrimental effect on teacher well-being (see b1 in figure 1). in our paper, we refer to this event as a professional challenge. regarding teachers’ resilience process, both single and mild types of stressful or unpleasant experiences—as well as longer-term stressful events like burnout or leaving the field—are possible professional challenges. the model’s next component, the appraisal process, refers to teachers’ assessment of the event (see b2 in figure 1). if individuals appraise a professional challenge as negative, stressful, or threatening, they believe that the harm of the event is likely to affect their wellbeing, which activates the resilience process (hascher et al., 2021). in the resilience process, teachers employ various strategies to cope with this negative, stressful, or threatening professional challenge (see b3 in figure 1). the strategy selection and application lead to an outcome (see b4 in figure 1) that reflects the strategy’s effectiveness. this outcome is cognitively and emotionally evaluated a second time by reflecting on the negative/stressful/threatening professional challenge (see b5 in figure 1) and, if the evaluation is neutral or positive, leads to the restoration of teacher well-being. if the evaluation is negative, the steps of selecting and applying strategies is repeated (see b6 in figure 1). this evaluation process, which, according to the lazarus model (lazarus & folkman, 1984), consists of a primary and secondary evaluation, is crucial for the individual’s perception of his or her ability to overcome similar professional challenges using the same strategies (hascher et al., 2021, pp. 426–427). while research has shown many challenges that teachers may face (e.g. gu & day, 2013), few studies have specifically addressed teachers’ strategies in the face of individually relevant professional challenges. 1.6 the present study in the aware model, identifying professional challenges and strategies supporting teacher well-being are vital for understanding how the resilience process contributes to teacher well-being. it is assumed that the individual cognitive and affective interpretation of an event, here defined as a professional challenge, is relevant for the activation of the resilience process. within this process, negative appraisals of, for example, student behaviour as disturbing, collegial exchange as less supportive, or school management as less effective, as well as negative emotions such as anger, frustration, discontentment, or anxiety are expected to activate the selection of strategies that help to overcome these challenges and to restore well-being. thus, knowledge about this critical part of the resilience process—that is, the interplay of professional challenges and teachers’ strategies—could advance the research field. previous research has shown that the teaching profession is under pressure due to the high responsibility for students, multiple demanding tasks that go beyond teaching, high workload and time pressure, and societal expectations such as inclusion or digital education (fernández-batanero et al., 2021; oecd, 2019). however, there is still a lack of knowledge of how teachers cope with these professional challenges (keller-schneider et al., 2020) and how they restore or maintain their well-being in the face of these challenges. while some quantitative research has successfully identified factors negatively or positively associated with teacher well-being (e.g. skaalvik & skaalvik, 2018), this qualitative study aims to illuminate the interrelatedness of individually relevant professional challenges and strategies within the resilience process. by considering teachers’ professional challenges and strategies, this study offers a subjective perspective on the situations of the teaching profession that affect teacher well-being and call for resilience. it helps to clarify how teachers cope with professional challenges in restoring their wellbeing. the insights into how teachers manage to function at work can inform schools on how to support teachers better and guide teacher education programmes in preparing future teachers. moreover, knowledge about professional challenges and coping strategies contributes to translating the theoretical aware model into empirical research and increases our understanding of the resilience process. accordingly, our study was guided by the following research questions (rqs): krummenacher et al 90 | f l r rq 1: what professional challenges do teachers encounter? rq 2: what strategies do teachers employ to restore or improve their well-being in response to professional challenges? embedded into the aware-model (hascher et al., 2021), we expect answers to these research questions to contribute to an understanding of the resilience process—which helps teachers to succeed and thrive in their profession—and, thus, add new knowledge to the theory of teacher well-being and resilience. 2. methods 2.1 participants and procedure study participants were n = 29 swiss teachers (28 females) of compulsory schooling (grades 1–9) who were invited for an online interview (45 to 60 minutes) between april and may 2021. participants were selected through snowball sampling. in a quantitative check, they proved to be high on the well-being scale except for one person. throughout the survey, schools were open for the entire school year, in contrast to the prior year, when, in switzerland, most schools were closed from march 16 to may 11, 2020. the study period was explicitly chosen as the pandemic was having less of an impact on schools. recognising that the pandemic could be challenging, we did not want to limit responses to this specific challenge and asked about challenges and corresponding strategies in teachers’ daily lives. this type of questioning may have resulted in some comments about the pandemic but should not limit the variety of responses. nearly half (48%) of the sample were early career teachers (5 years of experience or less), 32% had 6–15 years of teaching, and 20% had more than 16 years of experience. more than half of teachers (56%) were under 30, 32% were aged 31–50, and 12 % were older than 51. the majority of participants (56%) worked full-time. interviews were recorded and subsequently transcribed verbatim. participation was entirely voluntary, and participants consented to the audio recording. all personally identifiable information was anonymised. 2.2 instruments a semi-structured interview protocol (see appendix a) allowed us to delve deeply into the participants’ perspectives and concerns during the resilience process (creswell, 2002; hascher et al., 2021). the interview questions explored teacher well-being, resilience, and coping strategies for professional challenges. the questionnaire guide encouraged the participants to respond to the questions in light of their everyday professional life. we investigated the strategies used to cope with a challenging event or situation, the impact of these strategies on teachers’ well-being, and aligned teachers’ reported experiences with the aware model. the data used for this study were drawn from responses to the following questions. (professional) challenges • what (professional) challenges have you experienced? • in which phase of your professional life did this challenging event/situation happen? • please reflect on what you think the reasons for this event/situation were. • how did you feel? what emotions/feelings did you experience? • what role did students, colleagues, school administration, and parents play? strategies • what did you do about the challenging event/situation you experienced? • what personal resources and which of your character strengths helped you to deal with this event/situation? krummenacher et al 91 | f l r • which strategies helped you to cope with this event/situation? • what other resources and supporting factors did you find helpful? 2.3 data analysis to explore the strategies used in challenging situations, we conducted a comprehensive analysis of the data in three phases. in the first phase, we transcribed the interviews and used kuckartz’s (2018) qualitative content analysis method to develop a coding scheme. this resulted in the isolation of 71 challenges and 143 strategies, which were analysed separately to gain an overall understanding of their prevalence. codes emerged throughout the analysis to avoid predetermined frameworks (hewitt-taylor, 2001). maxqda software facilitated the coding process, and a codebook with clear definitions and characteristics of the categories was created. intercoder reliability assessment using cohen’s kappa was conducted, with discrepancies resolved through discussion and consensus. the analysis in this phase was inductive, leading to the identification of 9 first-order categories of professional challenges and 12 first-order categories of strategies (oancea & punch, 2014). in the second phase, we analysed the coded data to establish higher-order categories. similar concepts were grouped together (using a coding unit from 30 to 60 words), and we referred to relevant literature, including the work of beltman and poulton (2019; beltman et al., 2022), to identify these higher-order categories. consequently, we identified three higher-order categories for challenges. table 1 illustrates the challenging events and situations teachers encountered. further subcoding was performed from the three main codes—organisational, relational, and personal—and are shown below with anchor examples. participant identification is indicated in parentheses—that is, p11 indicates an exact quote from participant number 11. table 1 professional challenges: main and subcategories, anchor examples organisational professional challenges school system critique “(…) or children for whom you can’t find a place, or you notice in the whole system that this child is simply not in good hands and (…) yes, what do you do then? and these are already such questions or moments of stress, which are close to me and bother me, which i also take home with me.” [p10] class constellations “and also with the individualisation, i think that’s a good thing, but certain things are just somehow not possible, because you have so many students, because it’s a wild to and fro, because it’s difficult to always keep track of the learning level of each child. “[p22] reforms, new curriculum “(…) the current development with the curriculum 21, where i am really partly overwhelmed (…) everything that would be required, where i get the impression that we can no longer manage that as a school.” [p23] krummenacher et al 92 | f l r relational professional challenges concerns with colleagues “i have had to listen to a lot of very terrible comments about myself, which have nothing to do with my everyday work. badmouthing, blasphemies in a nasty way.” [p12] concerns with students “when a child refuses to do the tasks in class, that’s difficult.”[p26] concerns with parents “in my second year of working i already knew that one parent was rather difficult and then i had the parent meeting, and it didn’t go on that long, then this father was yelling at me. and i, as a young teacher, was there, and i didn’t know how to react to it. “[p20] personal professional challenges physical “i was very stressed because i knew ‘i have to go to school and i have to do this and that and oh no, there are still seven students who don’t have an apprenticeship’ (…) and i realised i come home, and i can’t sleep, or i can’t sleep very well.” [p06] workload “all the corrections, tests, meetings and so on (…) i don’t know where to start.” [p05] emotional “for me it is particularly stressful or particularly challenging when i notice that everything is slipping away from me, as if emotionally (…) when the daily routine with the children is simply too turbulent. when i don’t really believe in myself.” [p10] in the next step, the reported coping strategies were categorised. we discovered emotional strategies as found by beltman and poulton (2019). we subsequently classified them into four overarching categories, namely waiting, assessing, problem-solving, and being proactive. as all of the examples in the category problem-solving were related to activating social resources, this category was named relational problem solving (table 2). we then revisited the data to further differentiate categories related to interactional relationships at school and outside of school, as well as cooperative relationships at schools. table 2 provides anchor examples of the differentiated subcategories. throughout the analysis, an iterative process involving three researchers was employed to enhance reliability and credibility. regular discussions among two independent coders ensured consistency and consensus. the findings of each phase of the analysis are presented in the following section. krummenacher et al 93 | f l r table 2 strategies: main and subcategories, anchor examples waiting taking a deep breath “something can always happen (…) so you should just be open and take a breath, not panic right away.” [p29] stepping away from the situation “…and sleep on it once and (…) then look at it again the next day. maybe then you look at it with a completely different view. [p19] assessing looking at the bigger picture “yes, i was glad to realise that the problem was not me, but this woman. the fact that she herself had problems—understanding that helped me. [p08] taking someone else’s perspective “maybe i also have a bit of empathy and can respond empathetically to the parents. i think that then my goodwill towards the child also comes across.” [p27] optimism “(…) looking back, it was very enriching for me. this overpromotion was a challenge that strengthened me. so just as the situation was negative, it was also positive afterwards.” [p18] literature “in a book (…) i read tips on how to deal with difficult situations.” [p07] relational problem-solving instrumental social action “there was a school counsellor; i called him and also explained the situation (…) he explained to me that this is quite normal (…) and then he dealt with the parents and told them clearly (…) then it never happened again.” [p20] asking friends and family “yes, definitely the exchange with colleagues or family. they have a more neutral view of it because they are not in the middle of the situation. that’s why they have a different krummenacher et al 94 | f l r perspective and can come up with different solutions. yes, and so they can give you tips.” [p03] collaboration “i think that it makes a lot of difference that i am not alone in the classroom, because there is still the remedial teacher who sits down with me and says: what is our next step? [p07] being pro-active downtime “so, for me, it’s very important that i separate school and home. i also have a school phone where parents can reach me and which i rarely take home….” [p27] good preparation “if i know that i have parent–teacher meetings in two weeks, i prepare things for class in advance.” [p11] reducing workload “i have a correction station where the kids can correct themselves.” [p07] 3. results the following sections describe the essential findings and most significant aspects obtained from teachers’ responses to the interview questions in three sections. the first section demonstrates the reported professional challenges. the second section describes the strategies applied afterwards to overcome the professional challenge. the third section takes a holistic view of the resilience process through two illustrative cases. 3.1 reported professional challenges participants in this study faced various professional challenges regardless of how long they had been teachers. the professional challenges exhibited significant variation in duration and intensity, ranging from brief encounters to prolonged experiences lasting several months. daily weaker professional challenges and highly stressful and emotionally taxing challenges were described. the three levels of challenges are illustrated in table 3. the stated challenging situations or events were grouped into three primary levels—organisational (f=13), relational (f=43), and personal (f=15). despite the variety of professional challenges, which highlight the complex and multifaceted nature of the teaching profession, we found that the most common professional challenges were relational and equally distributed between concerns with parents, colleagues, and students. the challenges are outlined in order of frequency. krummenacher et al 95 | f l r table 3 reported professional challenges participants (01–29)1 challenges f relational 42 concerns with parents 15 concerns with students 15 concerns with colleagues 12 organisational 13 school system critique 6 class composition 5 reforms, new curriculum 2 personal 16 physical 5 emotional 5 workload 6 note: (f) relates to the overall frequencies of mentions. the intensity of the shades represents the frequency of mentions per participant. 1the columns correspond to the 29 participants. professional challenges at the relational level were reported most frequently. responses were coded into subcategories based on the groups of interaction partners. reported relational challenges included concerns with colleagues (f=12), with students (f=15), and with parents (f=15). teachers agreed that relational situations had a profound impact on their daily professional lives. they noted unresolved conflicts and lack of support from colleagues. the reported professional challenges with students were related to managing diverse student needs and behaviours and difficulties in engaging students effectively. teachers also expressed frustration with disruptive student behaviour and students’ lack of motivation. regarding professional challenges with parents, participants described feeling judged and scrutinised by parents. they expressed a need for better communication and partnership with parents to foster a supportive learning environment. at the personal level, physical issues (f=5), increased workload (f=5), and emotional challenges (f=6), such as low self-efficacy, were reported as professional challenges. teachers commented on the physical challenges by pointing out that they have insomnia and tinnitus and struggle with headaches krummenacher et al 96 | f l r more often than average. they felt burdened with the many different tasks and with meeting the different requirements. they also worried about a subjective lack of self-efficacy regarding professional tasks such as digital literacy, heterogeneity or parent cooperation. struggles with sleeping patterns, thought circles, and “not being able to let go” were factors reported on a personal level. professional challenges revealed a sense of overwhelm and the need for additional support. a small group experienced difficulties in work–life balance. professional challenges at the organisational level were mentioned least often. among the organisational challenges, critique of the existing school system (f=6), class composition (f=5), and issues with reforms and the new curriculum (f=2) were reported as challenging factors. the challenges related to reforms and the new curriculum highlighted the inherent difficulties in implementing changes within the educational system. participants expressed a need for adequate training and professional development to navigate such changes successfully. teachers seem concerned with the school system when children do not seem to “fit the existing system”. the critique of the existing school system and concerns about class compositions indicated a desire for more inclusive and flexible approaches to education. participants expressed a need for tailored support and resources to address the diverse needs of their students effectively. critique of the system was closely linked to the reported professional challenges related to class composition that include student heterogeneity, social integration, and adaptive teaching. 3.2 reported strategies the strategies that teachers reported could be assigned to four categories (see table 4)—waiting (f=9), assessing (f=24), problem-solving (f=88), and pro-active (f=19). among the various coping strategies, problem-solving strategies capitalising on social support emerged as consistently employed by the majority of participants, hence the category name of “relational problem solving”. krummenacher et al 97 | f l r table 4 reported coping strategies participants (01–29)1 strategies f relational problem-solving 88 instrumental social action 45 asking friends and family 23 collaboration 20 assessing 24 looking at the bigger picture 6 taking someone else’s perspective 8 optimism 5 literature, further reading 5 being proactive 22 downtime 12 good preparation 7 reducing workload 3 waiting 9 taking a deep breath 6 stepping away from the situation 3 note: (f) relates to the overall frequencies. the intensity of the shades represents the frequency of mentions per participant. 1the columns correspond to the 29 participants. the most prevalent strategy was relational problem-solving, which included three subcategories: instrumental social action, asking friends and family, and cooperation. instrumental social action (f=45) involved actively discussing professional challenges and seeking professional support. for example, the teachers’ lounge was recognised as a supportive environment where teachers could share their problems and seek advice from colleagues (f=15). teachers also shared their professional challenges more exclusively with a trusted colleague, followed by an in-depth discussion about potential solutions (f=14). teachers also sought advice from mentors (f=8) and school social workers (f=8); these professionals provided strategic guidance that facilitated the planning of subsequent actions. beyond the school environment, teachers found support through friends and family (f=23). these out-of-school relationships provide a network that helps teachers reporting about professional krummenacher et al 98 | f l r challenges in a non-professional context. in talking with friends and family, teachers primarily sought emotional support in dealing with professional challenges. collaboration included exchanging materials (f=5) and team teaching (f=5). the active cooperation of a second teacher implies sharing resources and responsibility for coping with professional challenges. team activities such as co-planning (f=14) and initiating team meetings addressing prevalent professional challenges (f=6) were also mentioned. assessing was the second most frequent coping strategy. the 24 reported situations of assessing were evenly distributed across four subcategories (see table 4). taking someone else’s perspective emerged as a prominent aspect of assessing strategies, with an equal number of statements addressing the perspectives of students and parents. in the case of optimism, the strategy of seeing the positive in a situation was reported. also, strategies of professional development such as further training (f=3) and consulting literature (f=2) were mentioned. the proactive coping strategy was divided into three categories, with downtime (f=12) being the most mentioned strategy. participants engaged in sports and allotted time dedicated to themselves. good preparation (f=7) was mentioned as a crucial individual strategy in which teachers planned the lessons and their free time ahead. the explicit mention of reducing their workload (f=3) further demonstrated the teachers’ proactive efforts to manage their responsibilities effectively. the least often reported strategy was waiting. despite the relatively low frequency of mentions, it is noteworthy that teachers reported remaining calm in a moment of heightened emotion as helpful, such as taking a deep breath (f=6) and stepping away from the situation (f=3). 3.3 understanding the resilience process through two illustrative cases our data analysis so far has revealed that relational professional challenges and problem-solving were reported most frequently among the participants. in the following section, we provide a more comprehensive understanding of how relational challenges negatively affect teacher well-being and how relational coping strategies contribute to restoring teacher well-being. two in-depth cases help to elucidate the resilience processes depicted in the aware model (see figure 1). both cases centre on a negative appraisal of a professional challenge involving students. however, they differ regarding the coping strategies employed and the outcomes achieved. in the first case, the teacher proactively sought help from professionals within the school and engaged in open communication with parents, which ultimately resulted in restoring the teacher’s well-being. in the second case, the initially applied strategies failed to restore the teacher’s well-being. however, through reflection and adjustment (illustrated as a loop in the aware model), the teacher identified a new strategy that ultimately led to an improvement in her condition, thus signifying an evolving journey wherein the resilience process remains incomplete. 3.3.1 a resilience process based on successful instrumental social action and collaboration participant 18 (pseudonym laura) began working at her current school as a firstand secondgrade teacher in 2015 after travelling and substituting. she was in her 8th year of teaching and taught students from grades 1–6. laura described having weeks where she felt “at ease” and how she loved the flexibility of being a teacher, as she was able to integrate hobbies into her daily activities whenever the weather was great. when a week was more challenging, she explained that she had to work “on weekends in cases of emergency”. however, having a good work–life balance has always been important to her. laura encountered a challenging situation (b1) with two boys with “behavioural issues” who “made a pact.” laura noticed “how much power they can have together” and worried about their impact on the class. she realised she needed to break the negative spiral to avoid unforeseen consequences. she appraised the professional challenge (b2c) as “extremely demanding” and realised that it kept “occupying her thoughts”. as a first strategy (b3), she talked to “school krummenacher et al 99 | f l r administrators, school social workers, the parents of students and the curative teacher”. she also assessed her part and reminded herself of tips she had learned from a former coach of autistic children in her class. she realised that her strategies had a positive outcome (b4) and that the children responded to the measures she had undertaken. looking back on the situation, laura described it as a highly enriching experience. she highlighted the effort between herself and the children as strengthening their relationship. laura noted that the children became aware that submitting to the measures implemented resulted in more freedom and the opportunity to engage in particular activities. in her perception, the situation had both negative and positive aspects. in retrospect, she appraised the situation (b5) as follows: looking back, it was very enriching for me. i was able to work with the children, strengthen the relationship, and (…) they became aware that it makes sense to submit to it because they have much more freedom or can also do special things. so, just as the situation was negative, it was also positive afterwards. 3.3.2 a resilience process with a loop based on unsuccessful instrumental social action but successfully asking friends and family participant 12 (pseudonym lily) had been working for 5.5 years and had experience in the third and fourth grades, as well as the fifth and sixth grades. her desired grade level would be fifth and sixth grade, but she took on a third-grade class in the summer. she had a workload of 30 lessons, up to 100%. lily described her daily routine as teaching in the morning and the afternoon and going home between 18:30 and 19:00 each evening. if she could take a break, it never lasted longer than 3 minutes, which she described as “not so pleasant”. for her, during teacher training, “a wonderful world (was) shown, which is not true and does not correspond to reality.” lily would have liked to see “all the teachers pulling together” at her current school and not “just looking out for themselves.” this became evident when she described it as “always the same people doing something” and “the same people doing nothing.” she described the professional challenge (b1) when “children were left to run wild”, and she was “not supported” in this challenging situation. the school administration “bullied” her and “picked on her personally and privately.” she appraised the situation (b2) as “very difficult”; she “almost burnt out because of it”. she then went on sick leave because she had “such severe cramps in all her muscles” and “could no longer switch off from her job”. she had to put up with “terrible comments about her person” when she left. she had the impression (b2c) of being “held responsible” for the situation. in this situation, lily had “felt helpless.” she described not knowing where to “get help.” on advice, she went to an outside agency (b3). however, this could have been of “little help” to her. they told her that she had “no chance to take action against such a school management” and was advised to “leave the school” (b4). she had hoped for help to learn “how one can defend oneself against such a system” and was “disappointed” (b5). afterwards, she did not know how to “get out of this situation” and felt that “nobody helped her” (b5c). it was only when her partner also advised her several times to quit that she “decided to leave” (b3 loop) to “help herself that way.” in reflecting on the professional challenges, lily stated that she was happy to have gone through them because they taught her valuable lessons on maintaining her wellbeing. she expressed that leaving school was not a sign of giving up (b5a) but rather a way of protecting herself and caring for her well-being. 4. discussion teacher well-being and resilience are topical in educational research, given the high demands placed on teachers and the teacher shortage. empirical studies are needed to further our understanding of the relationship between teachers’ resilience and well-being. this study investigated teachers’ professional challenges (rq1) and analyse which coping strategies teachers use to maintain or restore krummenacher et al 100 | f l r their well-being (rq2). in doing so, it aimed to advance our theoretical understanding of the resilience process within the newly introduced aware model (hascher et al., 2021). teachers in this study faced a variety of professional challenges at the organisational (e.g. school system, class composition) and personal levels (e.g. low self-efficacy, physical issues), with relational challenges being the most prevalent (rq1). this aligns with current research, as social challenges such as lack of recognition or support of principals, parent complaints, and challenging behaviour among students have repeatedly been described in the educational literature (castro et al., 2010; gu & day, 2013). challenging relationships with parents, students, and fellow teachers may cause negative experiences for teachers (beltman et al., 2019) and impede well-being (hascher & waber, 2021). professional challenges may be isolated and highly intense, such as an argument with a colleague or a parent, or less intense and sustained over time, such as the continuous disruptive behaviour of a student (gu, 2014). the variety and complexity of the case studies also illustrate the role of subjective interpretation of the interplay of individual competencies and work conditions, which leads to the appraisal of a professional challenge as positive, neutral, or negative. these appraisals broaden our understanding of the beginning of the resilience process as described in the aware model (hascher et al., 2021). regarding strategies to sustain well-being (rq2), teachers reported that relational problemsolving was the most frequent coping strategy, which indicates that relationships are perceived as both constraining and enabling (papatraianou et al., 2018). teachers reported using supportive interactional relationships in the school context with colleagues or school principals (honingh & hooge, 2014) administration and relationships outside of school (e.g. with family or friends) to seek advice or take action with their help to deal with professional challenges. this aligns with prior research showing that positive relationships with students, colleagues, and other school community members play a crucial role in boosting teacher work satisfaction and well-being (hascher & waber, 2021; rueger et al., 2016; skaalvik & skaalvik, 2011). positive interactions with students and colleagues can also be viewed as job resources in the job demands resources (jd-r) model, the “good things” at work that have inherent motivational quality and induce positive energy at work, leading to better outcomes (schaufeli, 2017, p. 121). more importantly, our results helped to identify an additional form of relational problem-solving that we describe as collaboration. collaboration related to a professional challenge can be seen as a more advanced form of a coping strategy that goes beyond seeking advice or sharing one’s feelings with others and is beneficial in managing stress and promoting well-being (akbari & eghtesadi, 2017; kar et al., 2021). instead, collaboration as a coping strategy enrols others into problem-solving. when teachers initiate collaboration in solving professional challenges, they recognise the systemic character of their challenges and aim to share their resilience process, as shown in the aware model. this mode of collaboration appears to be unique compared to established or regular collaboration (e.g. de jong et al., 2019; little, 1990; vangrieken et al., 2015), as it can be interpreted as a collective reaction to an individually perceived professional challenge. this likely reflects a teacher’s belief that education in school calls for sharing and collaboration instead of singular action (liu & benoliel, 2022). in line with ungar (2012), we advocate recognising the role of contextual and social environment. within the resilience literature, “flocking” is presented as a coping mechanism similar to collaboration (ebersöhn, 2012, p. 30). it underscores the group’s joint response when confronted with professional challenges by sharing their resources and support networks. like collaboration, which centres on working jointly to reach a mutual goal, flocking involves individuals uniting to surmount obstacles. it accentuates the strength of unity and collective action, thus demonstrating that when individuals band together, they can successfully navigate professional challenges to maintain well-being. in an educational institution, fostering a flocking culture would encourage educators to depend on each other, share resources, and collaboratively address professional challenges. along with relational problem solving, teachers reported employing assessing (e.g. finding the positive in a situation), being proactive (e.g. using downtime), and waiting (e.g. staying calm) as strategies. assessing involves strategies that allow individuals to evaluate the situation thoroughly, krummenacher et al 101 | f l r including reappraising the situation from different perspectives (beltman & poulton, 2019). maintaining a sense of purpose and self-efficacy while taking momentary breaks to reassess professional challenges also supports the resilience process illustrated in the aware model, as it changes the view of how a challenge is perceived. the proactive strategy aligns with the concept of self-care as an essential response among teachers (schussler et al., 2018). as beltman and poulton (2019) showed, teachers actively implemented strategies to manage future instances of heightened emotions effectively, including engaging in hobbies or physical exercise. accordingly, teachers aim to enhance their capacity to handle professional challenges that promote the resilience process in the aware model. consistent with sutton’s (2004) observation that well-prepared teachers experience fewer issues during lessons, participants in our study emphasised the significance of high-quality lesson preparation in preventing or mitigating professional challenges. although rarely mentioned, waiting as a strategy supports the idea that a resilience process needs emotional regulation. engaging in deep breathing and creating space for oneself can activate the sympathetic nervous system and facilitate recognising and labelling emotions, which enables teachers to respond appropriately to professional challenges (sharp & jennings, 2016). accordingly, mindfulness training might be helpful to enrich teachers’ strategies in reducing stress and improving well-being (beshai et al., 2016; sheppes et al., 2011) drawing a connection between professional challenges (rq1) and strategies (rq2), it becomes evident that social interactions serve a dual purpose. social interactions are frequently cited as the primary professional challenge, yet simultaneously, they offer a valuable resource for coping strategies. our research highlights that these social elements form a strategy for managing professional challenges. given that teaching is inherently a social profession and given, too, the crucial role of social belonging (ryan & deci, 2000), our findings underscore the importance of social competencies for thriving in this field. the two illustrative cases selected elucidate the relationship between professional challenges and strategies and a better understanding of the resilience process, as shown in the aware model. both cases are strongly related to the double role of social relatedness—that is, social interactions as professional challenges and for social support as a coping strategy—in restoring teacher well-being (e.g. aelterman et al., 2007). although the two cases differ regarding the resilience process and its outcomes, they are both represented in the aware model. in both cases, the teacher’s negative appraisal of a challenging situation with students sets the stage for the subsequent events (grayson & alvarez, 2008). both cases illustrate how a resilience process may be initiated through issues with students and how teachers apply instrumental social action as a coping strategy. in both cases, the retrospective appraisal of the situation as enriching suggests personal growth, which indicates the potential for challenging situations to foster professional development (mansfield et al., 2012). the two cases also demonstrate the variability of the resilience process within the aware model. in the first case, laura’s instrumental social action to seek help from various professionals and engage with parents demonstrates the role of relational strategies within the resilience process. she invited school team members and parents to discuss the challenging student behaviour and to search for a shared solution. this enabled collaboration in managing the professional challenges and expanded the perception of the issue as an individual professional challenge that needs collaboration inside and outside school. her commitment to maintaining work–life balance may further support her well-being. laura’s case also shows how the school context contributes to a successful resilience process (ainsworth & oldfield, 2019), as school members and parents were willing to share the responsibility for laura’s professional challenge. the second case demonstrates the negative consequences of lacking support and guidance from colleagues and school administration; it highlights how an unsupportive environment can hamper the resilience process and worsen the impact of professional challenges on teacher well-being (mansfield, 2021). the process leading to leaving school can be understood through lily’s experience and the theoretical framework of stress and coping. despite employing relational problem-solving strategies such as seeking support from school administration, colleagues, and the student’s parents, her wellbeing was not restored. this lack of support potentially frustrated her need for relatedness (ryan & deci, 2001), which led her to take leave and seek external advice, both of which were unsuccessful. as a final krummenacher et al 102 | f l r strategy, she left the profession to protect herself and regain her well-being (mansfield, 2021). lily’s perception of the challenge and her evaluation of her competence (keller-schneider et al., 2020) led to stress and dissatisfaction. as mcclelland (1998) suggests, a sense of competence helps form a selfimage capable of meeting requirements. when this self-image and perceived competence are undermined, it can lead to stress and, if unaddressed, could result in leaving the profession. providing adequate professional support and addressing these factors can thus potentially prevent or delay the point of leaving school. 4.1 limitations and strength of this study the present study has several limitations that lead to suggested avenues for future research. first, the findings of our study refer to a particular cultural group, swiss teachers. future research could profit from findings on maintaining well-being within the resilience process among teachers in other cultural groups. furthermore, future studies could control for teacher characteristics that may affect the relationship between well-being and resilience, such as teacher gender, migration background, professional years, or school stage (primary vs secondary education). in our study, almost half of the participants were early career teachers. while it is possible to track the resilience process, it does not necessarily mean that well-being is synonymous with staying in the profession. regarding the aware model, it would be expedient to investigate the underlying mechanisms in appraising the effectiveness of the chosen coping strategy. it would also be beneficial to investigate teachers’ strategies for coping with challenging situations in depth. we did not investigate professional challenges and strategies during the resilience process, as our analysis was conducted retrospectively. however, examining their implications within the context of the resilience process would be an intriguing aspect to explore in future studies. more detailed elaboration of the literature on teacher strategies could contribute to understanding the resilience process. furthermore, longitudinal research on strategies could contribute to the understanding of possible patterns of associations between challenges and strategies. our research has several strengths. first, the present study is one of the first to analyse the wellbeing and resilience processes in alliance. we investigated how teachers manage professional challenges and which challenges they face. we also summarised teachers’ strategies for challenging situations or events. moreover, we applied a qualitative approach using interviews to allow teachers to express themselves and reflect upon their reported past experiences. second, implementing the aware model may have led to a more sophisticated understanding of the relationship between well-being and resilience processes. in line with prior studies (e.g. johnson & down, 2013), the findings in our study have confirmed that teacher resilience plays a vital role in developing or restoring teacher well-being. a unique insight from our study reveals the intricate interplay of these social challenges and strategies in the day-to-day interactions of teachers with students, parents, and colleagues within the resilience process. by delving into the challenges faced by teachers, our research not only enriches our understanding but also offers practical implications for teacher education programmes and ongoing training initiatives. empowering both students and teachers with evidence-based strategies can be instrumental in navigating these challenges effectively. furthermore, fostering a collaborative environment within teacher education and training programmes can instil a sense of belief and camaraderie among teachers, potentially cultivating protective strategies early in their careers. krummenacher et al 103 | f l r keypoints teachers face a variety of professional challenges relational challenges were the most common reported challenges among teachers relational problem-solving was the most frequently reported strategy social relationships serve both as professional challenges as well as strategies teachers aim at actively maintaining their well-being in the face of challenges as illustrated in the aware model references aelterman, a., engels, n., van petegem, k., & verhaeghe, j. p. (2007). the well‐being of teachers in flanders: the importance of a supportive school culture. educational studies, 33(3), 285-297. https://doi.org/10.1080/03055690701423085 ainsworth, s., & oldfield, j. (2019). quantifying teacher resilience: context matters. teaching and teacher education, 82, 117–128. https://doi.org/10.1016/j.tate.2019.03.012 akbari, r., & eghtesadi, a. r. (2017). burnout coping strategies among iranian efl teachers. applied research on english language, 6(2), 179–192. https://doi.org/10.22108/are.2017.21346 avidov-ungar, o., & forkosh-baruch, a. (2018). professional identity of teacher educators in the digital era in light of demands of pedagogical innovation. teaching and teacher education, 73, 183-191. https://doi.org/10.1016/j.tate.2018.03.017 beltman, s., dobson, m. r., mansfield, c. f., & jay, j. (2019). “the thing that keeps me going”: educator resilience in early learning settings. international journal of early years education, 28(4), 303–318. https://doi.org/10.1080/09669760.2019.1605885 beltman, s., hascher, t., & mansfield, c. (2022). in the midst of a pandemic: australian teachers talk about their well-being. zeitschrift für psychologie, 230(3), 253–263. https://doi.org/10.1027/2151-2604/a000502 beltman, s., mansfield, c., & price, a. (2011). thriving not just surviving: a review of research on teacher resilience. educational research review, 6(3), 185–207. https://doi.org/10.1016/j.edurev.2011.09.001 beltman, s., & poulton, e. (2019). “take a step back”: teacher strategies for managing heightened emotions. the australian educational researcher, 46(4), 661–679. https://doi.org/10.1007/s13384-019-00339-x beshai, s., mcalpine, l., weare, k., & kuyken, w. (2016). a non-randomised feasibility trial assessing the efficacy of a mindfulness-based intervention for teachers to reduce stress and improve wellbeing. mindfulness, 7, 198-208. https://doi.org/10.1007/s12671-015-0436-1 bonanno, g. a., & burton, c. l. (2013). regulatory flexibility: an individual differences perspective on coping and emotion regulation. perspectives on psychological science: a journal of the association for psychological science, 8(6), 591–612. https://doi.org/10.1177/1745691613504116 borman, g. d., & dowling, n. m. (2008). teacher attrition and retention: a meta-analytic and narrative review of the research. review of educational research, 78(3), 367–409. https://doi.org/10.3102/0034654308321455 castro, a. j., kelly, j., & shih, m. (2010). resilience strategies for new teachers in high-needs areas. teaching and teacher education, 26(3), 622–629. https://doi.org/10.1016/j.tate.2009.09.010 https://doi.org/10.1080/03055690701423085 https://doi.org/10.1016/j.tate.2019.03.012 https://doi.org/10.22108/are.2017.21346 https://doi.org/10.1016/j.tate.2018.03.017 https://doi.org/10.1080/09669760.2019.1605885 https://doi.org/10.1027/2151-2604/a000502 https://doi.org/10.1016/j.edurev.2011.09.001 https://doi.org/10.1007/s13384-019-00339-x https://psycnet.apa.org/doi/10.1007/s12671-015-0436-1 https://doi.org/10.1177/1745691613504116 https://doi.org/10.3102/0034654308321455 https://psycnet.apa.org/doi/10.1016/j.tate.2009.09.010 krummenacher et al 104 | f l r cefai, c., & cavioni, v. (2014). from neurasthenia to eudaimonia: teachers’ well-being and resilience. in c. cefai, & v. cavioni (eds.), social and emotional education in primary school: integrating theory and research into practice (pp.133-148). new york; ny: springer science + business media. https://doi.org/10.1007/978-1-4614-8752-4. creswell, j. w. (2002). educational research. planning, conducting, and evaluating quantitative and qualitative research. london: pearson education. day, c., & q. gu. (2010). the new lives of teachers. routledge. day, c., & q. gu. (2014). resilient teachers, resilient schools: building and sustaining quality in testing times. routledge. de jong, l., meirink, j., & admiraal, w. (2019). school-based teacher collaboration: different learning opportunities across various contexts. teaching and teacher education, 86, 1–12. https://doi.org/10.1016/j.tate.2019.102925 duckworth, a. l., quinn, p. d., & seligman, m. e. (2009). positive predictors of teacher effectiveness. the journal of positive psychology, 4(6), 540–547. https://doi.org/10.1080/17439760903157232 ebersöhn, l. (2012). adding ‘flock’to ‘fight and flight’: a honeycomb of resilience where supply of relationships meets demand for support. journal of psychology in africa, 22(1), 29-42. https://doi.org/10.1080/14330237.2012.10874518 fantilli, r. d., & mcdougall, d. e. (2009). a study of novice teachers: challenges and supports in the first years. teaching and teacher education, 25(6), 814–825. https://doi.org/10.1016/j.tate.2009.02.021 fernández-batanero, j. m., román-graván, p., reyes-rebollo, m. m., & montenegro-rueda, m. (2021). impact of educational technology on teacher stress and anxiety: a literature review. international journal of environmental research and public health, 18(2), 548. https://doi.org/10.3390/ijerph18020548 flores, m. a. (2006). being a novice teacher in two different settings: struggles, continuities, and discontinuities. teachers college record, 108(10), 2021–2052. https://doi.org/10.1111/j.1467-9620.2006.00773.x flores, m.a. (2018). teacher resilience in adverse contexts: issues of professionalism and professional identity. in: wosnitza, m., peixoto, f., beltman, s., mansfield, c.f. (eds) resilience in education. springer, cham. https://doi.org/10.1007/978-3-319-76690-4_10 flores, m. a., & day, c. (2006). contexts which shape and reshape new teachers’ identities: a multi-perspective study. teaching and teacher education, 22(2), 219-232. https://doi.org/10.1016/j.tate.2005.09.002 grayson, j. l., & alvarez, h. k. (2008). school climate factors relating to teacher burnout: a mediator model. teaching and teacher education, 24(5), 1349-1363. https://doi.org/10.1016/j.tate.2007.06.005 gu, q. (2014). the role of relational resilience in teachers’ career-long commitment and effectiveness. teachers and teaching, 20(5), 502–529. https://doi.org/10.1080/13540602.2014.937961 gu, q. (2018). (re)conceptualizing teacher resilience: a social-ecological approach to understanding teachers’ professional worlds. in: wosnitza, m., peixoto, f., beltman, s., mansfield, c.f. (eds) resilience in education. springer, cham. https://doi.org/10.1007/978-3-319-76690-4_2 gu, q., & day, c. (2007). teachers resilience: a necessary condition for effectiveness. teaching and teacher education, 23(8), 1302–1316. https://doi.org/10.1016/j.tate.2006.06.006 gu, q., & day, c. (2013). challenges to teacher resilience: conditions count. british educational research journal, 39 (1), 22–44. https://doi.org/10.1080/01411926.2011.623152 hagenauer, g., hascher, t., & volet, s. e. (2015). teacher emotions in the classroom: associations with students’ engagement, classroom discipline and the interpersonal teacher-student relationship. european journal of psychology of education, 30, 385–403. https://doi.org/10.1007/s10212-015-0250-0 https://doi.org/10.1007/978-1-4614-8752-4 https://doi.org/10.1016/j.tate.2019.102925 https://doi.org/10.1080/17439760903157232 https://doi.org/10.1080/14330237.2012.10874518 https://doi.org/10.1016/j.tate.2009.02.021 https://doi.org/10.3390/ijerph18020548 https://doi.org/10.1111/j.1467-9620.2006.00773.x https://doi.org/10.1007/978-3-319-76690-4_10 https://doi.org/10.1016/j.tate.2005.09.002 https://doi.org/10.1016/j.tate.2007.06.005 https://doi.org/10.1080/13540602.2014.937961 https://doi.org/10.1007/978-3-319-76690-4_2 https://doi.org/10.1016/j.tate.2006.06.006 https://doi.org/10.1080/01411926.2011.623152 https://doi.org/10.1007/s10212-015-0250-0 krummenacher et al 105 | f l r harding, s., morris, r., gunnell, d., ford, t., hollingworth, w., tilling, k., evans, r., bell, s., grey, j., brockman, r., campbell, r., araya, r., murphy, s., & kidger, j. (2019). is teachers’ mental health and wellbeing associated with students’ mental health and wellbeing? journal of affective disorders, 242, 180–187. https://doi.org/10.1016/j.jad.2018.08.080 hascher, t., beltman, s., & mansfield, c. (2021). teacher wellbeing and resilience: towards an integrative model. educational research, 63(4), 416–439. https://doi.org/10.1080/00131881.2021.1980416 hascher, t., & hagenauer, g. (2011). wohlbefinden und emotionen in der schule als zentrale elemente des schulerfolgs unter der perspektive geschlechtsspezifischer ungleichheiten. in a. hadjar (hrsg.), geschlechtsspezifische bildungsungleichheiten. (s. 285-308). vs verlag. hascher, t., & waber, j. (2021). teacher well-being: a systematic review of the research literature from the year 2000–2019. educational research review, 34, 100411. https://doi.org/10.1016/j.edurev.2021.100411 hewitt-taylor j. (2001). use of constant comparative analysis in qualitative research. nursing standard (royal college of nursing (great britain) : 1987), 15(42), 39–42. https://doi.org/10.7748/ns2001.07.15.42.39.c3052 honingh, m., & hooge, e. (2014). the effect of school-leader support and participation in decision making on teacher collaboration in dutch primary and secondary schools. educational management administration & leadership, 42(1), 75-98. https://doi.org/10.1177/1741143213499256 johnson, b., & down, b. (2013). critically re-conceptualising early career teacher resilience. discourse: studies in the cultural politics of education, 34(5), 703–715. https://doi.org/10.1080/01596306.2013.728365 jones, c., hadley, f., waniganayake, m., & johnstone, m. (2019). find your tribe! early childhood educators defining and identifying key factors that support their workplace wellbeing. australasian journal of early childhood, 44(4), 326–338. https://doi.org/10.1177/183693911987 kar, n., kar, b., & kar, s. (2021). stress and coping during covid-19 pandemic: result of an online survey. psychiatry research, 295, 113598. https://doi.org/10.1016/j.psychres.2020.113598 keller-schneider, m., yeung, a. s., & zhong, h. f. (2020). supporting teachers’ sense of competence: effects of perceived challenges and coping strategies. in l. a. caudle (ed.), teachers and teaching: global practices, challenges, and prospects (pp. 269-289). nova science publishers. kitching, k., morgan, m., & o’leary, m. (2009). it’s the little things: exploring the importance of commonplace events for early‐career teachers’ motivation. teachers and teaching, 15(1), 43–58. https://doi.org/10.1080/13540600802661311 klusmann, u., richter, d., & lüdtke, o. (2016). teachers’ emotional exhaustion is negatively related to students’ achievement: evidence from a large-scale assessment study. journal of educational psychology, 108(8), 1193–1203. https://doi.org/10.1037/edu0000125 kuckartz, u. (2018). qualitative inhaltsanalyse. methoden, praxis, computerunterstützung (4., überar. aufl.). beltz juventa. lazarus, r. s., & folkman, s. (1984). stress, appraisal and coping. springer. lazarus, r. s., & folkman, s. (1991). the concept of coping. in a. monat & r. s. lazarus (eds.), stress and coping: an anthology (pp. 189–206). columbia university press. (reprinted from “stress, appraisal, and coping,” new york: springer publishing company, inc, 1984) lewis, r., & frydenberg, e. (2004). adolescents least able to cope: how do they respond to their stresses?. british journal of guidance & counselling, 32(1), 25-37. https://doi.org/10.1080/03069880310001648094 li, q., q. gu, & he, w. 2019. resilience of chinese teachers: why perceived work conditions and relational trust matter. measurement: interdisciplinary research and perspectives 17 (3): 143–159. https://doi.org/10.1080/15366367.2019.1588593 https://doi.org/10.1016/j.jad.2018.08.080 https://doi.org/10.1080/00131881.2021.1980416 https://doi.org/10.1016/j.edurev.2021.100411 https://doi.org/10.1177/1741143213499256 https://doi.org/10.1080/01596306.2013.728365 https://doi.org/10.1177/1836939119870906 https://doi.org/10.1016/j.psychres.2020.113598 https://doi.org/10.1080/13540600802661311 https://doi.org/10.1037/edu0000125 https://doi.org/10.1080/03069880310001648094 https://doi.org/10.1080/15366367.2019.1588593 krummenacher et al 106 | f l r little, j. w. (1990). the persistence of privacy: autonomy and initiative in teachers’ professional relations. teachers college record, 91(4), 509–536. liu, y., & benoliel, p. (2022). national context, school factors, and individual teacher characteristics: which matters most for teacher collaboration? teaching and teacher education, 120, 103885. https://doi.org/10.1016/j.tate.2022.103885 mansfield, c. f. (2021). cultivating teacher resilience: international approaches, applications and impact. springer. https://doi.org/10.1007/978-981-15-5963-1 mansfield, c., beltman, s., & price, a. (2014). ‘i’m coming back again!’the resilience process of early career teachers. teachers and teaching, 20(5), 547-567. https://doi.org/10.1080/13540602.2014.937958 mansfield, c. f., beltman, s., broadley, t. & weatherby-fell, n. (2016). building resilience in teacher education: an evidenced informed framework. teaching and teacher education, 54, 77–87. https://doi.org/10.1016/j.tate.2015.11.016 mansfield, c. f., beltman, s., price, a. & mcconney, a. (2012). “don’t sweat the small stuff:” understanding teacher resilience at the chalkface. teaching and teacher education, 28(3), 357–367. https://doi.org/10.1016/j.tate.2011.11.001 marzano, r. j. (2003). what works in schools: translating research into action. association for supervision and curriculum development masten, a. s. (2014). global perspectives on resilience in children and youth. child development, 85(1), 6–20. https://doi.org/10.1111/cdev.12205. mcclelland, d. c. (1998). identifying competencies with behavioral-event interviews. psychological science, 9(5), 331-339. https://doi.org/10.1111/1467-9280.00065 oecd (2019), talis 2018 results (volume i): teachers and school leaders as lifelong learners. oecd publishing, paris, https://doi.org/10.1787/1d0bc92a-en papatraianou, l. h., strangeways, a., beltman, s., & schuberg barnes, e. (2018). beginning teacher resilience in remote australia: a place-based perspective. teachers and teaching, 24(8), 893–914. https://doi.org/10.1080/13540602.2018.1508430 parker, p. d., & martin, a. j. (2009). coping and buoyancy in the workplace: understanding their effects on teachers’ work-related well-being and engagement. teaching and teacher education, 25(1), 68-75. https://doi.org/10.1016/j.tate.2008.06.009 parker, p. d., martin, a. j., colmar, s., & liem, g. a. (2012). teachers’ workplace well-being: exploring a process model of goal orientation, coping behavior, engagement, and burnout. teaching and teacher education, 28(4), 503-513. https://doi.org/10.1016/j.tate.2012.01.001 oancea, a. e., & punch, k. f. (2014). introduction to research methods in education. introduction to research methods in education, 1-448. rueger, s. y., malecki, c. k., pyun, y., aycock, c., & coyle, s. (2016). a meta-analytic review of the association between perceived social support and depression in childhood and adolescence. psychological bulletin, 142(10), 1017–1067. https ://doi.org/10.1037/bul0000058 ryan, r. m., & deci, e. l. (2000). self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. american psychologist, 55, 68–78. https://doi.org/10.1037110003-066x.55.1.68 ryan, r. m., & deci, e. l. (2001). on happiness and human potentials: a review of research on hedonic and eudaimonic well-being. annual review of psychology, 52(1), 141–166. https://doi.org/10.1146/annurev.psych.52.1.141 https://doi.org/10.1016/j.tate.2022.103885 https://doi.org/10.1007/978-981-15-5963-1 https://doi.org/10.1080/13540602.2014.937958 https://doi.org/10.1016/j.tate.2015.11.016 https://doi.org/10.1016/j.tate.2011.11.001 https://doi.org/10.1111/cdev.12205 https://doi.org/10.1111/1467-9280.00065 https://doi.org/10.1787/1d0bc92a-en https://doi.org/10.1080/13540602.2018.1508430 https://doi.org/10.1016/j.tate.2008.06.009 https://doi.org/10.1016/j.tate.2012.01.001 https://doi.org/10.1037/bul0000058 https://doi.org/10.1037110003-066x.55.1.68 https://doi.org/10.1146/annurev.psych.52.1.141 krummenacher et al 107 | f l r schaufeli, w. b. (2017). applying the job demands-resources model: a ‘how to’ guide to measuring and tackling work engagement and burnout. organizational dynamics, 46(2), 120–132. https://doi.org/10.1016/j.orgdyn.2017.04.008 schelvis, r. m., zwetsloot, g. i., bos, e. h., & wiezer, n. m. (2014). exploring teacher and school resilience as a new perspective to solve persistent problems in the educational sector. teachers and teaching, 20(5), 622–637. https://doi.org/10.1080/13540602.2014.937962 schussler, d. l., greenberg, m., deweese, a., rasheed, d., demauro, a., jennings, p. a., & brown, j. (2018). stress and release: case studies of teacher resilience following a mindfulness-based intervention. american journal of education, 125(1), 1-28. https://doi.org/10.1086/699808 sharp, j.e., jennings, p.a. strengthening teacher presence through mindfulness: what educators say about the cultivating awareness and resilience in education (care) program. mindfulness 7, 209–218 (2016). https://doi.org/10.1007/s12671-015-0474-8 segovia, f., moore, j. l., linnville, s. e., & hoyt, r. e. (2015). optimism predicts positive health in repatriated prisoners of war. psychological trauma: theory, research, practice, and policy, 7(3), 222–228. https://doi.org/10.1037/a0037902 sheppes, g., scheibe, s., suri, g., & gross, j. j. (2011). emotion-regulation choice. psychological science, 22(11), 1391–1396. https://doi.org/10.1177/0956797611418350 skaalvik, e. m., & skaalvik, s. (2011). teacher job satisfaction and motivation to leave the teaching profession: relations with school context, feeling of belonging, and emotional exhaustion. teaching and teacher education, 27(6), 1029–1038. https://doi.org/10.1016/j.tate.2011.04.001 sutton, r. e. (2004). emotional regulation goals and strategies of teachers. social psychology of education, 7, 379– 398. https://doi.org/10.1007/s11218-004-4229-y ungar, m. (2012). social ecologies and their contribution to resilience. in: ungar, m. (eds) the social ecology of resilience. springer, new york, ny. https://doi.org/10.1007/978-1-4614-0586-3_2 vangrieken, k., dochy, f., raes, e. & kyndt, e. (2015). teacher collaboration: a systematic review. educational research review, 15, 17–40. https://doi.org/10.1016/j.edurev.2015.04.002 veronese, g., pepe, a., dagdukee, j., & yaghi, s. (2018). teaching in conflict settings: dimensions of subjective wellbeing in arab teachers living in israel and palestine. international journal of educational development, 61, 16–26. https://doi.org/10.1016/j.ijedudev.2017.11.009 viac, c., & fraser, p. (2020). teachers’well-being: a framework for data collection and analysis. oecd education working papers no. 213. https://doi.org/10.1787/c36fc9d3-en zimmer‐gembeck, m.j., & skinner, e.a. (2016). the development of coping: implications for psychopathology and resilience. development and psychopathology, 1-61. https://doi.org/10.1002/9781119125556.devpsy410 https://doi.org/10.1016/j.orgdyn.2017.04.008 https://doi.org/10.1080/13540602.2014.937962 https://doi.org/10.1086/699808 https://doi.org/10.1007/s12671-015-0474-8 https://psycnet.apa.org/doi/10.1037/a0037902 https://doi.org/10.1177/0956797611418350 https://doi.org/10.1016/j.tate.2011.04.001 https://doi.org/10.1007/s11218-004-4229-y https://doi.org/10.1007/978-1-4614-0586-3_2 https://doi.org/10.1016/j.edurev.2015.04.002 https://doi.org/10.1016/j.ijedudev.2017.11.009 https://doi.org/10.1787/c36fc9d3-en https://doi.org/10.1002/9781119125556.devpsy410 krummenacher et al 108 | f l r appendix a a semi-structured interview protocol warm-up (introductory questions, small talk) block: warm-up i would like to get to know your daily work routine a little better. please tell me what a typical working day looks like for you. from morning to evening. what do you do during the day? please refer to the time before the corona pandemic. content aspects: specific questions: follow-up questions: level (preschool, primary, lower secondary, upper secondary, timeout class) • what grades do you teach? • how long have you been teaching at this level? • can you tell me any other things? • is there anything else? • do you have an example so i can picture it clearer? • can you describe it in more detail? • what do you mean by that? typical workday • what does the morning and afternoon work look like? • how many teaching hours do you normally have per day/week? • when do you prepare for class? • how many breaks do you have and how long do they take? main part (topics to answer the research questions) theme 1: teacher career block: reflection on teaching career let’s now look back at your teaching career. what were the most important stages in your development as a teacher? please briefly describe the most important stages. content aspects: specific questions: follow-up questions: education/teaching experience • where did you complete your training? • where did you gain your first teaching experience? • what further training have you done? • do you have an example so i can picture it clearer? • can you describe it in more detail? • what do you mean by that? professional passion • if you had the chance to start over again, would you still choose a career in teaching? • (if no) what profession would you like to choose and why? krummenacher et al 109 | f l r theme 2: teacher well-being block 1: school level (positive and negative aspects) now we will talk about the school where you teach. could you please tell me: what makes your school stand out? what are the strengths that make your school stand out? what need for improvement do you see? content aspects: specific questions: follow-up questions: positive and negative aspects of the context (challenging events/situations) • how would you describe your school? • what is the reputation of your school, in your opinion? • what professional development opportunities/educational resources do you have? • how is your relationship with your colleagues? do the teachers in your school support each other? do you have enough collegial support? • how is the school climate? how is the school culture? • what would you like to see differently? • can you tell me any other things? • is there anything else? • do you have an example of this so i can imagine it more concretely? • can you describe it in more detail? • what do you mean by that? block 2: class level (lessons, students) content aspects: specific questions: follow-up questions: contextual factors/class level teaching • how do you feel about teaching? how much do you like teaching? • are there differences in terms of classes or subjects? • what is your teaching philosophy? • what is most important to you while teaching? • how would you rate the classroom climate? • what facilitates teaching in the class(es)? • what makes teaching in this class/these classes challenging? • what would you like to see differently? students • do you have an example so i can picture it clearer? • can you describe it in more detail? • what do you mean by that? krummenacher et al 110 | f l r • could you describe your class to me a little bit? what is the class climate like? • how do you get along with your students? • how is the interaction with students in the class? how is the relationship with the students? • what can be said about the class composition? • what are the relationships with parents like? • what would you like to see differently? block 3: individual level (teachers) could you please explain: what personal characteristics are most important to your work as a teacher? what strengths characterize you as a teacher? in which areas do you see a need for development? content aspects: specific questions: follow-up questions: personal factors/ individual level (level of the teacher) • as a teacher, what do you enjoy most of all? which aspects of your work make you happy? • (if positive aspects(s) were mentioned) you said yes, that (answer: an aspect that makes you happy) makes you happy. what helps you maintain this? • are there personal aspects that make your teaching career difficult? • what role does your school play in this? • what would you have liked to have done differently? • we all know that the teaching profession can be stressful, and rest is essential. how do you organize your rest? • do you have an example so i can picture it clearer? • can you describe it in more detail? • what do you mean by that? theme 3: resilience block 1: challenging/negative event/situation and coping strategies working as teachers is often associated with challenges. please now select a particularly challenging situation or event. describe this situation/event. content aspects: specific questions: follow-up questions: challenging event/situation • what challenging event or situation did you experience? • at what stage of your professional life did this happen? • thinking back, what do you think were the reasons for this situation/event? • how complex was the challenge/event? • do you have an example so i can picture it clearer? • can you describe it in more detail? • what do you mean by that? krummenacher et al 111 | f l r • how stressful was the challenge/event? • what was particularly difficult for you? • how did you feel? what emotions/feelings did you experience? • what role did students, colleagues, school administration, and parents play in this? coping strategies • what did you do about it? • how helpful was what you did? why? • what personal resources, which of your character strengths helped you deal with this situation? • where did you experience difficulties? what didn't work out so well? why? • what would have been other options/strategies? why did you not choose them? • what other resources, and supporting factors did you find useful? theme 4: prevention & intervention block 1: prevention what should be paid attention to in teacher education so that young teachers can successfully enter the profession? what should be paid attention to in teacher education so young teachers feel comfortable in the profession? content aspects: specific questions: follow-up questions: prevention (promoting well-being and resilience in teacher education) • what should change in teacher education so that teachers feel ready to teach and begin a teaching career? • what can help teachers do well in their work lives? • what can contribute to teachers’ wellbeing? • what could/should schools contribute? • do you have an example so i can picture it clearer? • can you describe it in more detail? • what do you mean by that? block 2: intervention what do you think would help make teachers happy to go to work? content aspects: specific questions: follow-up questions: intervention (measures; promoting teachers’ well-being and resilience) • what does a school need to consider to make teachers feel good? • what could a school change to help teachers cope better with challenges? • what structures need to be adapted? • what could improve job satisfaction? • do you have an example so i can picture it clearer? • can you describe it in more detail? • what do you mean by that? krummenacher et al 112 | f l r conclusion block: conclusion in conclusion to your reflections: what would you advise a teacher to do if she finds herself in a difficult situation and does not know what to do? content aspects: specific questions: follow-up questions: what is your conclusion? • which strategies are helpful currently? whom can you trust? • how should you behave to bring about an actual change? • what can you do yourself to change the situation? • can you tell me any other things? • do you have an example so i can picture it clearer? • can you describe it in more detail? • what do you mean by that? block: completeness is there anything else you would like to mention that is important to you but has not come up here in the interview? content aspects: specific questions: follow-up questions: is there anything else you would like to elaborate on? • can you tell me any other things? • is there anything else? • what else? frontline learning research vol. 11 no. 2 (2023) 31-48 issn 2295-3159 corresponding author: david john, idp-educational technology, indian institute of technology bombay, powai, mumbai, maharashtra 400076, india, write2john@iitb.ac.in doi: https://10.14786/flr.v11i2.1165 rethinking pedagogical use of eye trackers for visual problems with eye gaze interpretation tasks david john1 and ritayan mitra1 1indian institute of technology bombay, india article received 31 august 2022/ revised 23 february 2023/ accepted 14 september 2023 / available online 15 november 2023 abstract eye tracking technology enables the visualisation of a problem solver's eye movement while working on a problem. the eye movement of experts has been used to draw attention to expert problem solving processes in a bid to teach procedural skills to learners. such affordances appear as eye movement modelling examples (emme) in the literature. this work intends to further this line of work by suggesting how eye gaze data can not only guide attention but also scaffold learning through constructive engagement with the problem solving process of another human. inferring the models’ problem solving process, be it that of an expert or novice, from their eye gaze display would require a learner to make interpretations that are rooted in the knowledge elements relevant to such problem solving. such tasks, if designed properly, are expected to probe or foster a deeper understanding of a topic as their solutions would require not only following the expert gaze to learn a particular skill, but also interpreting the solution process as evident from the gaze pattern of an expert or even of a novice. this position paper presents a case for such tasks, which we call eye gaze interpretation (egi) tasks. we start with the theoretical background of these tasks, followed by a conceptual example and representation to elucidate the concept of egi tasks. thereafter, we discuss design considerations and pedagogical affordances, using a domain-specific (chemistry) spectral graph problem. finally, we explore the possibilities and constraints of egi tasks in various fields that require visual representations for problem solving. keywords: emme, eye tracking, visual problems, perceptual learning mailto:write2john@iitb.ac.in https://10.0.57.194/flr.v11i2.1165 john & mitra 32 | flr 1. introduction in the last decade, eye movement capture has expanded its utility from being a research method for studying cognitive processes (rayner, 1992, 2009) to a tool for instructional design (chiu, 2016; kok & jarodzka, 2017). eye gaze patterns of successful problem solvers can contain meaningful task relevant patterns (thomas & lleras, 2007; grant & spivey, 2003) and the viewing of such gaze patterns has been shown to improve problem solving performance (litchfield & ball, 2011; litchfield et al 2010). the affordance of eye gaze displays to make attention observable has since been explored for designing multimedia instruction. one of the earliest such implementations were the eye movement modelling examples (emme) for procedural problem solving tasks (van gog et al., 2009). an emme is an instructional video where a pointer – whose location is fixed by gaze data – is used to show exactly where the expert is looking while they are solving a problem, with the goal of teaching that skill. such modelling, that is, observation of expert problem solving is intended to help learners achieve joint attention with the model (usually an expert) and thereby improve information processing. multiple studies have shown that emmes improve learner attention, i.e., resulting in faster, longer, and more frequent attention to task relevant elements, which in turn leads to increased learning performance (for a meta-analysis, see xie et al., 2021). outside of modelling examples, the use of eye gaze displays to provide attentional guidance can be seen in general multimedia instruction, where it’s used to perform the role of a laser pointer for establishing joint attention with the presenter (sung, g., feng, t., & schneider, b., 2021; d'angelo, s., & schneider, b., 2021). additionally, it was expected that eye gaze displays would enhance modelling by allowing a learner to observe problem-solving processes that are covert or absent in verbal explanations or actions of the expert. this aspect has been successfully demonstrated with emmes designed for nonprocedural classification problems (jarodzka et al., 2010, 2012, 2013) and multimedia comprehension tasks where learners modelled expert behaviour of text-image integration (kerbs et al., 2019, mason et al., 2015, 2017). however, for procedural problems there is little evidence that emmes offer any benefit beyond attentional guidance (chisari et al., 2020, van marlen et al., 2016, noord, s. v., 2016, van gog et al., 2009). as concluded by such studies, this could either mean that eye gaze displays were made redundant by expert verbalisations, or indicate the absence of processes that can be learned solely by observing an eye gaze display. this raises the question of whether there is more to what eye gaze displays can offer or if it is simply a matter of narrowing down the conditions under which emmes are effective (tunga & cagiltay, 2023). we believe there exists untapped potential in eye gaze displays if the affordance of only making attention observable is extended to the affordance of making cognitive processes interpretable. this affordance originates from the eye mind hypothesis (just & carpenter, 1984, 1980) according to which there is an average to good association between a person’s gaze and conscious thought. while it has been suggested that the eye-mind hypothesis is weak and limited for certain tasks (anderson et al., 2004), recent research on eye gaze displays highlights their ability to make the intentions and processes of a person interpretable (emhardt et al., 2020; foulsham & lock, 2015; van wermeskerken et al., 2018; zelinsky et al., 2013). furthermore, the domain-specific validity of the eye-mind hypothesis indicates that even if the association is not universally true, there could be specific domains (e.g. geometry problems) where it is sufficiently valid (schindler & lilienthal, 2019). therefore, in domains and problems where eye gaze displays are reasonably interpretable, the comprehension of a problem solver's eye gaze display (hereafter referred to as the ‘model’ and used in an expertise-agnostic sense) whilst performing a task has the potential to become a meaningful meta-task for another problem solver (hereafter referred to as the ‘learner’). this possibility arises from the fact that the interpretation of an eye gaze display is firmly rooted in the contextual information intrinsic to the given problem-solving task (kok & jarodzka, 2017). in this position paper, we explore and demonstrate the possibility of creating such meta-tasks, or eye gaze interpretation (egi) tasks, for a spectral graph problem. the egi task will use a model’s eye gaze display to pose questions about the problem solving processes of the model. therefore, the egi problem is in essence a problem about a problem, i.e., a meta-problem and we create such tasks by problematizing the problem solving process as stated in the title. we will simplify the domain-specificity of the problem to the extent possible for our general audience. we discuss a pilot demonstration of creating such egi tasks using the eye gaze data of two models solving one spectral graph problem. we will create seven egi tasks and discuss their pedagogical possibilities. section 2 discusses the theoretical framing of egi tasks. john & mitra 33 | flr section 3 provides a conceptual framework for such tasks. section 4 discusses the experimental setup used for the development of egi tasks. section 5 discusses the eye gaze data and the egi tasks generated thereof. section 6 discusses the scope of the tasks, followed by section 7, which discusses the limitations of the pilot demonstration and outlines future work. 2. theoretical framing of egi tasks the theoretical positioning of egi tasks is best understood in terms of the potential perceptual processes they elicit, such as what information receives attention, how they are organised, and interpreted. perceptual learning can be defined as ‘an increase in the ability to extract information from the environment, as a result of experience and practice with stimulation coming from it’ (gibson, 1969). even though perceptual learning applies to various problem solving scenarios, it is most significant to visual problem solving tasks, from which much of the early evidence for perceptual learning has been obtained (kellman & massey, 2013). the acquisition of expertise in visual problem solving tasks is best seen as perceptual learning that comprises the acquisition of various perceptual skills. this fact is well known in domains such as medical education or radiology where visual problems are of great significance and perceptual learning is achieved through repeated practice (alexander et al., 2020; guegan et al., 2021). one part of perceptual learning is the low level perceptual processes such as the identification of visual objects, symbols, or space (watanabe & sasaki, 2015, dosher & lu 2017). as low level perceptual skills are mostly acquired through experience, novice learners benefit from perceptual support. for low level perceptual skills, much like the emmes, the egi tasks also give perceptual support to novice learners. the perceptual support from eye gaze displays can reduce the demand for low level perceptual processes by drawing attention to relevant visual objects and improving visual search. from a cognitive load standpoint alone, emmes can be as good or even superior to egi tasks in this regard by virtue of it being an example based learning approach (merrienboer, 2013; renkl, 2014). however, an egi task adds to this experience because it goes beyond scaffolding of low level perceptual processing and gives structure to the visual information being processed. in particular, the major benefits of egi tasks are expected from the leveraging of high level perceptual processes like ‘noticing’. noticing is defined as the process of actively selecting and interpreting relevant information in the environment, assuming that any given situation contains an infinite amount of information to be perceived (van es & sherin, 2002). it is important to note that noticing is not limited to being able to identify visually salient features or being able to pay attention to important elements of a problem. instead, noticing is a perceptual process where a learner is able to notice elements of deeper concepts associated with a problem. recent studies that attribute improved learning performance to naturally occurring instances of noticing in socio-cultural practices (lobato et al., 2012) and the design of experimental tasks that support noticing (chase et al., 2019) make it a significant perceptual process that allows learning to take place. the proposed egi tasks can be expected to elicit noticing in learners about the conceptual relationships between various perceptual objects and the conceptual objects within the context of a problem solving situation. the eye gaze data of the egi task is expected to draw attention to key elements of problem-solving and the task itself is supposed to make learners notice how the elements are related within the context of the problem. it is this latter affordance that makes egi tasks novel and distinct from emmes. the possibility of egi tasks to offer perceptual support, and make learners notice interrelated domain concepts resemble the two complementary mechanisms used by reiser (2004) to characterise learning in scaffolded problem solving environments. the first mechanism is that of structuring which allows a learner to interact with the problem in a systematic and meaningful manner without being overwhelmed by its complexity. the second mechanism is the problematization of critical domain concepts so that learners can spend sufficient attention engaging with such concepts. in a similar fashion, we can expect egi tasks to support learning on two fronts. first, as a perceptual scaffold which reduces the perceptual complexity of the problem and structures the visual problem in terms of the problem solving paths or states of the model. for example, even in graph comprehension problems (mitra et al., 2017), where learners have to synthesise information from the different axes and the problem statement, learners can fail to effectively select and organise relevant information. the eye gaze present in an egi task will reduce the problem complexity by highlighting and sequencing the relevant information. this could provide john & mitra 34 | flr significant scaffolding for tasks where the mere identification of visual objects or patterns may itself be challenging such as in ecg graph signatures or seismic charts. second, by asking questions that require noticing of the underlying domain concepts, the learner is challenged to think beyond what is required from the original problem on which the egi task is based. for example, in a study by mitra et al. (2017), successful learners solving a graph comprehension problem were found to dwell on the ‘conversion factor’ longer than others. an egi task designed for this problem would inquire about why successful learners spend more time on the conversion factor, prompting the learner to truly recognise the role of the ‘conversion factor’ in solving the problem. finally, the unique pedagogical potential of egi lies in the ‘meta’ nature of the task. the learner is expected to interpret the thoughts of the model by using their knowledge of the task and the eye gaze display of the model. while metacognition is a term used for cognition of one’s own cognition, flavell’s (1979) original definition happens to be more inclusive and one that we intend to use here. flavell (1979) proposed that metacognitive knowledge extended beyond the person as “the person category encompasses everything that you could come to believe about the nature of yourself and other people as cognitive processors.” instructional designs that foster metacognition tend to focus on learners being metacognitive about their own cognition and is a coveted learning outcome in any type of learning (berardi-coletta et al., 1995; mayer, 1998; rickey & stacy, 2000; schoenfeld, 2016). therefore, egi tasks, which are essentially meta-tasks, could nudge learners to be metacognitive about the problem-solving processes of others, which is expected to improve learning outcomes as well. 3. conceptualization of egi tasks 3.1 conceptual examples of egi tasks we describe the concept of an egi task using four hypothetical gaze displays (m1, m2, m3, and m4) for a geometry problem of parallel lines. the geometry problem (shown in fig. 1) consists of identifying if the two red lines are parallel based on the intersecting angles they make with a common transversal. fig. 1 shows the hypothetical eye gaze pattern (represented as sequentially numbered circles for fixations and arrows for saccades) of four models, who arrived at the right answer to the geometry problem (i.e., the red lines are not parallel). a plausible egi task for these eye gaze displays would be to interpret the problemsolving logic of the model (i.e., the axioms and theorems applied by the model), assuming that the model solved the problem successfully. we shall use this egi task as an example to understand how different models map to the aforementioned conceptual representation of egi tasks. a. m1 b. m2 john & mitra 35 | flr c. m3 d. m4 figure 1. hypothetical eye gaze displays of four models (m1, m2, m3, and m4) solving a geometry problem. (a) m1: higher process interpretability and lower task relevant knowledge. (b) m2: higher process interpretability and higher task relevant knowledge. (c) m3: lower process interpretability and lower task relevant knowledge. (d) m4: lower process interpretability and higher task relevant knowledge. the hypothetical gaze models vary in terms of the interpretability of both the eye gaze display and the model's problem-solving process. interpretability of a gaze display can be understood as the degree to which the problem-solving process is reflected in the gaze display. in this context, interpretability is higher in m1 and m2 compared to m3 and m4, respectively. the reduced interpretability of the process in m3 and m4 results from the presence of stray gaze patterns that lack meaningful relevance to the tasks at hand. the problem solving process of m1 is similar to m3 and m2 is similar to m4. interpreting each of these models requires a different level of task relevant knowledge that can be characterised based on the axioms and theorems of line geometry as discussed below. the relevant knowledge associated with the egi task can be grouped into three levels, which are theorems of intersecting lines, axioms of parallel lines, and combined insights from these theorems and axioms. at the first level, we have theorems of intersecting lines, consisting of the vertical angles theorem and the adjacent angles theorem (vertically opposite angles are equal and adjacent angles are supplementary). at the second level, we have axioms of parallel lines, consisting of the corresponding angle axiom, alternate angle axiom and interior angle axiom (corresponding angles are equal, alternate angles are equal, and interior angles are supplementary). in addition to this, the third level consists of the insights about the intersecting lines that are obtained from combining these axioms and theorems and are beyond the immediate application of either of them. hence, we can operationalize the task relevant knowledge axis for the egi task to have 3 levels, in the increasing order of theorems of intersecting lines, axioms of parallel lines and insights from a combination of the theorems and axioms. the problem solving process of m1 and m3 can be interpreted as a sequential application of the domain knowledge components (specifically alternate angle axiom and adjacent angle theorem). the sequence of eye gaze fixations and the information present in those areas are collectively interpreted to infer the problem solving process. however, interpreting the problem solving process of m2 and m4 will require knowledge beyond that of the individual axioms and theorems, i.e., the two given angles should add to 180o; demanding a synthesis of the theorem of intersecting lines and the axiom of parallel lines. a learner who can make an accurate interpretation can be said to have displayed a comprehensive understanding of the respective theorems and axioms. 3.2 conceptual representation of egi tasks now we shall conceptualise egi tasks in terms of the general factors that are involved in interpreting an eye gaze display. interpretation of an egi task is dependent on the interpretability of eye gaze display and the learner’s task relevant knowledge (which includes domain concepts and task understanding). hence, all egi tasks can be conceptually represented in terms of two mutually orthogonal factors; interpretability john & mitra 36 | flr and task relevant knowledge (as shown in fig. 2). the x-axis in the figure represents the amount of task relevant knowledge of a learner solving the egi task. the y-axis represents the interpretability of the eye gaze display that is being used to create an egi task. we conceive the interpretability of a gaze display as being based on the capture of several task-related attributes and the solution processes. a highly interpretable gaze display would have minimal erratic or task-irrelevant eye movements and would be mostly free of task-independent information. the y-axis intercept at (a) denotes the minimum level of interpretability required to create an egi task, as beneath this level, the model processes pertaining to the task are not interpretable from the eye gaze display. the y-axis intercept at (b) denotes the level of interpretability beyond which an egi task is trivial, as there is no need for any interpretation as the process is directly observable. while it's uncommon for a gaze display of a problem-solving task to fully convey the model's process without any interpretation, such gaze displays are highly suitable for emmes when it comes to communicating the model's processes to learners. this is because learners can readily understand these gaze patterns without significant interpretation. these gaze patterns are ideal for emmes, especially when gaze data is the sole source of process information. for instance, studies by mason et al. in 2015 and 2016 demonstrated that students extracted critical aspects of the model's behaviour from eye movements alone, without any external explanations or prompts. in addition to this, we can also conceive an x-axis intercept at (c), to represent the minimum amount of task relevant knowledge required to engage with a non-trivial egi task, and a y-axis intercept (d) to denote the total amount of task relevant knowledge needed to solve the egi task. all possible instances of learners solving egi tasks generated from a specific problem solving task can be denoted by the space bounding the four intercepts (a, b, c, and d). we hypothesise that the ability to interpret a gaze display with lower interpretability will depend on their task relevant knowledge. we expect the higher levels of task relevant knowledge to aid in the abstraction of process information and ignoring task irrelevant gaze patterns. as a corollary, the egi tasks of gaze displays with high interpretability are suited for learners with lower task-related knowledge. the green triangle has been used to conceptualise this inverse relationship. it denotes a set of egi tasks that is suitable (solvable) for various learners in the expert-novice spectrum, possessing different amounts of task relevant knowledge. the triangle is only representative of the inverse relationship and does not imply a linear nature of any kind. figure 2. a conceptual representation of egi tasks. (a) minimum level of interpretability required to frame an egi task. (b) eye gaze display becomes trivial to interpret (c) minimum amount of task relevant john & mitra 37 | flr knowledge required to solve an egi task. (d) the total task relevant knowledge associated with a task. m1, m2, m3, and m4 represent the conceptual examples of egi tasks discussed in section 3.2. 4. experimental setup for developing egi tasks 4.1 nature of the problem we chose spectral graph problems to develop egi tasks. a spectral graph problem can be seen as a type of puzzle where scientists use a graph to figure out the arrangement of atoms in a substance. the graph shows peaks and valleys (fig. 4), each representing a frequency of radio waves absorbed and emitted by the atoms. by analysing the position and shape of the peaks, scientists can deduce the structure of the substance. it's like solving a puzzle where each peak represents a clue that needs to be put together to reveal the final picture. this technique is widely used in chemistry to determine the molecular structure of substances. besides the domain-specific relevance, we chose spectral problems to create egi tasks as its solution requires continual visual processing and the visual information has good spatial spread. 4.2 study design for collecting gaze data and creating egi tasks our study consisted of two chemistry postgraduate students who solved four spectral graph problems in succession. they were provided with a graphic tablet to scribble or make markings while solving the problem. the eye gaze data of the participants were captured using a tobii pro-x120 eye tracker (tobii pro ab, 2014) and imotions software (imotions, 2021). the study design consisted of collecting eye gaze models, creating egi tasks from interpretable segments of the eye gaze models and then having the models retrospectively interpret each other’s gaze patterns to evaluate the validity of the interpretable segments identified and the corresponding egi tasks (fig. 3). the study was approved by the local institutional ethics committee (no. iitb-iec/2019/012) and both the participants provided informed written consent. figure 3. study design prior to solving the spectral graph problems, participants engaged in a priming activity that featured a quiz to facilitate active recall of relevant domain concepts. they were also granted access to cheat sheets pertinent to solving the spectral problems. the priming activity was done so that the participants understood the nature of the problem solving activity and got accustomed to using the graphic tablet in the problem solving interface. the spectral graphs were obtained from the open access repository spectral zoo (muzyka, 2021) and were sequenced in increasing order of difficulty. even though the participants attempted 4 spectral problems, the egi tasks from the first problem (q1) alone are presented in this paper. as this is an exploratory position paper, this truncation of data is unproblematic. the eye gaze displays (video of eye gaze overlaid on the problem solving interface) were inspected to identify interpretable eye john & mitra 38 | flr gaze segments and were used to frame egi tasks. the participants solved the egi tasks on a later date to partly validate that the gaze segments are interpretable and have meaningful interpretations. 5. creation of egi tasks in this section, we discuss the egi tasks created from the eye gaze data of models p1 and p2 solving the spectral graph problem. we briefly describe the spectral graph problem and then describe the problem solving process of both models along with the eye movement displays that were used to create the egi tasks. lastly, we list the egi tasks generated and discuss their pedagogical implications. 5.1 spectral graph problem: nmr spectra of 1-nitro-propane the nmr spectral graph problem is a key technique in organic chemistry used to determine the arrangement of atoms in a substance. decoding a spectral graph requires the study of a graph that displays peaks and valleys, each representing different frequencies of radio waves. each peak and valley of a spectral graph corresponds to a piece of information that must be assembled to deduce the molecular structure of an organic compound. a spectral graph (in fig. 4) is solved by piecing together the inferences made from the position (i.e. the x coordinate) and the shape (i.e. the number of vertical lines making a peak) of the signals (marked a, b, and c in fig. 4), to deduce the molecular structure. the position of a signal is read from its relative x-coordinate, which is indicative of the electronegative nature of the corresponding component of the molecule. the shape of a signal refers to the splitting pattern (no. of peaks in each signal) and the relative height of each peak. the shape gives information regarding the hydrogen atoms present in the substance. for a detailed and accurate description of the spectral graph and the solution to the specific spectral graph please refer to the appendix. figure 4. a spectral graph. image from spectral zoo (muzyka, 2021) 5.2 description of eye gaze models here we describe the eye gaze patterns of models p1 and p2 using scan path diagrams (as in fig. 5 & fig. 6). the scan path diagram consists of gaze fixation bubbles connected by lines of gaze saccades. the numerical values on each fixation bubble denote the fixation sequence. the different signals in the spectral graph shall be referred to as signals a, b, and c as described in fig. 4. gaze model representations for longer segments can become crowded. therefore, we recommend that readers refer to the eye gaze display videos of p1 and p2 (john, 2022a, 2022b). also, note that the static scan paths are presented only for the sake of describing the eye gaze models to the reader as egi tasks would rarely use static scan paths. instead short video showing the temporal evolution of scan paths is a precondition for such tasks. 5.2.1 eye gaze pattern of p1 p1 was successful in solving the spectral graph and arrived at the solution within 45 seconds of inspecting the spectra. p1 started inspecting the spectra from the left side. he fixated on the signal a, and then at the signal b. his gaze then moved past the signal c. this was followed by fixating on the molecular formula john & mitra 39 | flr of the compound. the scan path plot corresponding to p1’s initial inspection is shown in fig. 5a. this was followed by fixating on the signal c, and a quick scan of individual signals before writing the solution to the spectra on the white space on the right (fig. 5b, right panel). the scan path up to the writing of the solution is shown in fig. 5b. p1 spent a total of 38 seconds to arrive at the final answer. the eye gaze display of p1 can be accessed from (john, 2022a). (a) (b) figure 5. scan path of p1 relevant to egi tasks. the three signals have been marked as a, b and c for ease of readability. 5.2.2 eye gaze pattern of p2 p2 was unsuccessful in solving the spectral graph problem. the model began by carefully inspecting and decoding the splitting patterns of each signal, following the corresponding scan path. (fig. 6a). p2 then created four partial structures (fig. 7a, 7b, 7c and 7d) before settling on the wrong answer (fig. 7e). partial structures are partly formed answers made by p2 at different instances. p2 had wrongly marked signal b as a septet (i.e. he identified signal b to have 7 splits) during his initial inspection of the spectra and corrected it to a sextet (i.e. a signal of six splits) when working on the first partial structure (fig. 7a). the gaze pattern pertaining to this can be seen in fig. 7b. p2 deleted it and proceeded to create the next partial structure (fig. 7b) which was similar to the earlier one and was very close to the correct solution. the gaze pattern of this segment (fig. 6c) shows p2 making adjustments to the partial structure based on the information from the entire spectra. just before he abandoned the third partial structure (fig. 7c), with gaze focused on signal a and the nitrogen atom on the molecular formula (fig. 6d). p2 then moved on to another partial structure (fig. 7c), this time without paying much attention to the spectral information (fig. 6e), and realised that it was very far off from the solution (moved on from it within 10 seconds). further, he tried one more partial structure (fig. 7d) by focusing primarily on signal c and signal b areas of the spectra as depicted in (fig. 6f). then he arrived at this final structure (fig. 7e) focused on the previously partial structure (fig. 7d) as depicted in (fig. 6g) p2 spent a total of 8 minutes 35 seconds to arrive at the final answer. the eye gaze displays of p2 can be accessed from (john, 2022b) john & mitra 40 | flr (a) (b) (c) (d) (e) (f) (g) figure 6. scan paths of p2 that are relevant for creating egi tasks. a. scan path from the initial inspection of spectral graph. b. first partial structure. c. second partial structure. d. second partial structure (end segment) e. third partial structure. f. fourth partial structure. g. final answer. the three signals have been marked as a, b, and c for ease of readability. john & mitra 41 | flr (a) (b) (c) (d) (e) figure 7. partial structures formed by p2 at different instances. a. first partial structure. b. second partial structure. c. third partial structure. d. fourth partial structure. e. final answer. 5.3 egi tasks created from eye gaze data of p1 and p2 we were able to create two broad egi tasks (t1 and t2 in table 1) from the eye gaze of models p1 and p2. these tasks require the entire gaze data of the respective models to be interpreted. p1 was a successful model that used the minimum amount of information from the spectral graph (the no. of signals and the no. of carbon atoms) and hence the egi task was to deduce his problem solving logic (t1). in the case of p2 who was unsuccessful, the egi task was to uncover the reason for this failure (t2). interpreting the reason for p2 is possible because the gaze patterns reveal the segments of spectral information that he used at different stages of problem solving and pertaining to different partial structures. in addition to these broad egi tasks, we were also able to create egi tasks that are specific to particular gaze segments of model p2 (see tasks t2a, t2b, t2c, and t2d in table 1) and an egi task comparing the gaze segments of initial inspection of p1 and p2 (see tasks t1a in table 1). these tasks could be used to aid a learner who is not able to solve the broad egi tasks and can be framed as objective multiple-choice questions as shown in table 1. since broad egi task t2 requires the learner to inspect a long gaze model, the specific egi tasks (t2a, t2b, t2c, and t2d) could guide them to notice specific instances, make interpretations and then collectively arrive at the solution. the solutions to all the egi tasks were validated by both models, by independently arriving at the same interpretations as the author. both the models performed a retrospective evaluation of either of their gaze patterns guided by the egi tasks. the models attempted the egi tasks in the same sequence (t1, t1a, t2a, t2b, t2c, t2d, t2), where the solution of egi task t2 was aided by first answering t2a, t2b, t2c and t2d. next, we discuss the pedagogical value of the seven egi tasks created in this study. table 1 egi tasks created from gaze models of models p1 and p2 egi task solution t1: infer the problem solving logic of p1 from his eye gaze model. p1 identified that there are three signals. then p1 realised that the no. of signals is the same as the no. of carbons. the model hence did not search for further information from the spectral graph and arrived at the final answer. t1a: how is the problem solving approach of p1 different from p2? infer based on their initial inspection of the spectral graph? p2 looked at each individual signal and started to unpack the splitting pattern information of each signal john & mitra 42 | flr from left to right using the n+1 rule1, unlike p1 who broadly inspected the entire spectral graph and realised that the no. of signals is the same as the no. of carbons. t2: identify the misconception or confusion that resulted in p2 failing to solve the problem from his eye gaze model. the model appears to be confused with applying the n+1 rule and integration factor1. the n+1 rule applies to the adjacent carbon and not to the carbon corresponding to the signal. the spectral information inspected by p2 at the beginning when the partial answer was close to the correct answer (fig. 7a, and fig. 7b) and the subsequent modifications made by p2 indicate the confusion. t2a what made p2 identify his mistake of having wrongly labelled signal b as a septet (having 7 lines) instead of a sextet (having 6 lines)? what triggered p2 to make this correction? a. revisiting the spectral graph and close inspection of 𝛿 (2.1) b. hydrogen insufficiency2 c. violation of n+1 rule. t2b: interpret the major domain concept that is embodied in the eye gaze of the given segment. a. de-shielding electronegativity2 b. n+1 rule c. integration factor d. valency of carbon t2c: identify the difference between the problem solving process used by p2 in arriving at the 3rd partial structure (fig. 7d) with respect to the other partial structures (fig. 7b and fig. 7c) a. random guessing b. accounting for the remaining hydrogen atom. c. n+1 rule t2d: identify the nmr spectral information and the corresponding domain concept that p2 never considered while solving the problem. a. de-shielding electronegativity b. n+1 rule c. integration factor2 the egi tasks t1 and t1a allow a learner to notice the hierarchy of information that is relevant to solving the nmr problem. t1a contrasts the problem solving approach of model p2 with the successful model p1, where p2 has failed to realise the basic information present in the no. of signals and no. of carbons. this was the most elementary piece of information needed to arrive at the correct solution. these egi tasks may also be helpful in realising the need for being systematic when considering the spectral information and choosing the appropriate strategy suitable to the problem. it’s also interesting to note that regardless of the model’s problem solving success or failure, the eye gaze display can still yield valuable insights about the problem solving process and hence be suitable for egi task generation. for example, even though t2b was created from the unsuccessful model p2, it contained an interpretable gaze segment that clearly captured the elements pertaining to a domain concept (i.e. deshielding electronegativity). the egi tasks generated from model p2 (i.e. t2, t2a, t2c and t2d) also reveal to a learner major misconceptions or confusions that could arise in the process of problem solving. they also highlight the inefficiency of problem solving strategies (random assembly of atoms to satisfy 1 detailed descriptions of the rules and concepts related to the spectral graph are provided in the appendix. 2 egi tasks t2a, t2b, t2c and t2d being mcqs their solutions are indicated in bold text. john & mitra 43 | flr the molecular formula) that a learner might opt for when not able to combine two pieces of spectral information. this attribute is somewhat underutilised in conventional instructional approaches, as they are primarily designed to support novices to learn from successful (or expert) models alone. egi tasks t1a and t2c could allow learners to compare problem solving processes and the contrasting cases and thereby understand varied perspectives of different problem solvers. such perspective taking might be particularly valuable if the learner is an instructor or teacher who is expected to provide feedback or scaffold other learners. hence, egi tasks seem to indicate a novel possibility for engaging learners with deeper and tacit components associated with the problem solving processes. as demonstrated, they can allow reflection on cognitive processes of problem-solving that go beyond their own and those of their immediate peers. 6. scope of egi tasks this work is an early-stage conceptualization and demonstration of egi tasks that explore their pedagogical potential in visual problem solving. the pilot demonstration shows that egi tasks can allow learners to interact with various elements of the problem solving processes such as choice of problem solving strategy, the influence of misconceptions or lack of concept clarity, and consequences of inefficient problem solving tactics. these elements extend beyond the domain knowledge components typically encountered during conventional learning activities involving instruction or problem solving. consequently, egi tasks could be regarded as a promising tool to foster a deeper understanding of the interactions among domain concepts, knowledge of problem-solving strategies, and misconceptions. our conceptualization of egi tasks for a specific problem solving task is based on the interpretability of gaze display and the learner’s task relevant knowledge. however, when considering egi tasks across the space of all possible visual problems, a key determinant of the nature and value of an egi task is going to be the type of visual representation that is required to solve the root problem. hence, a brief description of the variations that exist amongst visual representation tasks in the stem fields and their relevance vis-avis egi task creation can set the stage for understanding the true scope of egi tasks. figure 8. mapping visual problems based on the entropy of visual objects and their procedural nature. the citations correspond to visual problem solving tasks for which emme studies exist. we have mapped common visual problems studied in stem fields, focusing on two parameters we consider critical in the context of eye movements and visual problem-solving (as shown in fig. 8). while the above mapping is by no means exhaustive, it aids in the discussion of egi tasks beyond the limited scope of the examples presented in this paper. first, we consider the entropy of visual objects in the visual representation corresponding to the problem solving task. we borrow the term entropy from shannon’s information theory (1948) as a measure of the amount of information present in a visual representation being used in a problem. in tasks with higher entropy of visual objects, the requirements for perceptual processing can be assumed to be larger. for such tasks identifying or noticing a salient visual object itself john & mitra 44 | flr can become critical to the problem solving process, and hence lead to eye gaze patterns that are less interpretable. this is so because the model's perception of visual objects can be ambiguous and not interpretable from the eye gaze data alone. second, we consider a binary classification based on the procedural nature of the task (procedural vs. non-procedural); where the procedural tasks are those that require some set of transformations to be performed on the information present in the visual form. this classification is also present in the emme literature and has been found to be a differentiator in the effectiveness of emmes, with them being ineffective for non-procedural tasks (van marlen et al., 2016; xie et al., 2021). they both have different advantages in the context of designing egi tasks. the benefit of using non-procedural tasks is that most (if not all) of the problem solving process is perceptual and hence a significant amount of the process is captured within the gaze patterns. the benefit of using procedural tasks is that they create partial solutions or artefacts during the problem solving process; which if captured (through screen recording or graphic inputs etc.) can provide more contextual information for gaze interpretation (as seen with the partial structures of molecules in the above case study). 7. limitations and future work this work is an early stage conceptualisation and therefore has some inherent limitations. one major limitation is the absence of validation of these tasks with actual learners. thus, future work should characterise learner interactions and performance with egi tasks to validate the hypothesised benefits. another limitation is that we have not assessed the efficiency and effectiveness of egi tasks in terms of time and resource requirements compared to alternative methods like direct instruction or guided practice for supporting learners' problem-solving. moreover, our choice of spectral graph problems to create egi tasks was guided by their inherent characteristic of having task relevant information evenly distributed across the display. this characteristic reduces the likelihood of violating the eye-mind assumption due to factors like para-foveal vision or ambiguity about the attended information at any given moment. nevertheless, it's important to acknowledge that the creation of egi tasks may be susceptible to the limited validity of the eye-mind hypothesis (anderson, 2004; kok & jarodzka, 2017). this is a valid concern and warrants further investigation. however, we are fairly optimistic about the applicability of a reasonably strong eye-mind hypothesis in domain-specific problems like the spectral graph problem, as has been shown by some researchers (schindler & lilienthal, 2019; wu & liu 2022). there is, however, a need to improve the methodological reliability of identifying interpretable segments within a broader dataset. in addition to the pedagogical possibilities discussed in this paper, we also anticipate an assessment potential from our conceptualisation (fig. 1) of the inverse relation between the interpretability of an eye gaze display and the task relevant knowledge needed to interpret such a display. such assessment, if possible, could probe knowledge elements beyond domain concepts that are tested in conventional assessments, like that of strategies and latent concepts related to a problem solving task. this possibility shall be explored in future work, with a battery of validated and reliable egi tasks. key points pedagogical possibilities that exploit the affordance of eye gaze displays to make cognitive processes interpretable are beginning to be explored. the proposed eye gaze interpretation (egi) tasks are one way to leverage this affordance and potentially foster a deeper understanding of problem solving in learners. egi tasks require learners to interpret different aspects of a model's problem solving process (such as the problem solving strategy) from interpretable eye gaze displays and the task relevant knowledge (i.e. primarily the domain knowledge relevant to the task). as learners need to infer the reasoning of the problem solver, an egi task could potentially nudge the learner to be metacognitive about the problem solving processes. we have provided a theoretical background for, a conceptual representation of and a set of egi tasks in the position paper. future study directions and challenges have been discussed. the egi tasks indicated the potential of such tasks to reach deeper knowledge elements than traditional questions related to such problems. john & mitra 45 | flr acknowledgements we acknowledge the role of two anonymous reviewers and the editor-in-chief, prof. nina b. dohn, for providing detailed feedback which greatly improved the quality of the manuscript. we also extend our gratitude to mr. amrit pal singh for help with proofreading of the manuscript. this work was supported by an iit bombay internal research grant, rd/0517-irccsh0-001, to ritayan mitra. references alexander, r. g., waite, s., macknik, s. l., & martinez-conde, s. (2020). what do radiologists look for? advances and limitations of perceptual learning in radiologic search. journal of vision, 20(10), 17-17. https://doi.org/10.1167/jov.20.10.17 anderson, j. r., bothell, d., & douglass, s. (2004). eye movements do not reflect retrieval processes: limits of the eye-mind hypothesis. psychological science, 15(4), 225-231. https://doi.org/10.1111/j.09567976.2004.00656.x berardi-coletta, b., buyer, l. s., dominowski, r. l., & rellinger, e. r. (1995). metacognition and problem solving: a process-oriented approach. journal of experimental psychology: learning, memory, and cognition, 21(1), 205. https://doi.org/10.1037/0278-7393.21.1.205 chase, c. c., malkiewich, l., & s kumar, a. (2019). learning to notice science concepts in engineering activities and transfer situations. science education, 103(2), 440-471. https://doi.org/10.1002/sce.21496 chiu, m.-h. (ed.). (2016). science education research and practices in taiwan. springer singapore. https://doi.org/10.1007/978-981-287-472-6 dosher, b., & lu, z. l. (2017). visual perceptual learning and models. annual review of vision science, 3, 343. emhardt, s. n., wermeskerken, m., scheiter, k., & gog, t. (2020). inferring task performance and confidence from displays of eye movements. applied cognitive psychology, 34(6), 1430–1443. https://doi.org/10.1002/acp.3721 flavell, j. h. (1979). metacognition and cognitive monitoring: a new area of cognitive–developmental inquiry. american psychologist, 34(10), 906. https://doi.org/10.1037/0003-066x.34.10.906 foulsham, t., & lock, m. (2015). how the eyes tell lies: social gaze during a preference task. cognitive science, 39(7), 1704–1726. https://doi.org/10.1111/cogs.12211 gibson, e. j. (1969). principles of perceptual learning and development. https://doi.org/10.2307/1572721 guegan, s., steichen, o., & soria, a. (2021, march). literature review of perceptual learning modules in medical education: what can we conclude regarding dermatology?. in annales de dermatologie et de vénéréologie (vol. 148, no. 1, pp. 16-22). elsevier masson. https://doi.org/10.1016/j.annder.2020.01.023 helle, l. (2017). prospects and pitfalls in combining eye-tracking data and verbal reports. frontline learning research, 5(3), 1-12. https://doi.org/10.14786/flr.v5i3.254 imotions (9.1), imotions a/s, copenhagen, denmark, (2021). jarodzka, h., scheiter, k., gerjets, p., & van gog, t. (2010). in the eyes of the beholder: how experts and novices interpret dynamic stimuli. learning and instruction, 20(2), 146–154. https://doi.org/10.1016/j.learninstruc.2009.02.019 jarodzka, h., van gog, t., dorr, m., scheiter, k., & gerjets, p. (2013). learning to see: guiding students’ attention via a model’s eye movements fosters learning. learning and instruction, 25, 62–70. https://doi.org/10.1016/j.learninstruc.2012.11.004 john, david [egi tasks]. (2022, may 13). p1nitropropane [video]. youtube. https://www.youtube.com/watch?v=wfcferb0v2k https://doi.org/10.1167/jov.20.10.17 https://doi.org/10.1111/j.0956-7976.2004.00656.x https://doi.org/10.1111/j.0956-7976.2004.00656.x https://doi.org/10.1037/0278-7393.21.1.205 https://doi.org/10.1002/sce.21496 https://doi.org/10.1007/978-981-287-472-6 https://doi.org/10.1002/acp.3721 https://doi.org/10.1037/0003-066x.34.10.906 https://doi.org/10.1111/cogs.12211 https://doi.org/10.2307/1572721 https://doi.org/10.1016/j.annder.2020.01.023 https://doi.org/10.14786/flr.v5i3.254 https://doi.org/10.1016/j.learninstruc.2009.02.019 https://doi.org/10.1016/j.learninstruc.2012.11.004 https://www.youtube.com/watch?v=wfcferb0v2k john & mitra 46 | flr john, david [egi tasks]. (2022, may 13). p2 nitropropane [video]. youtube. https://www.youtube.com/watch?v=u5zwduver-c just, m. a., & carpenter, p. a. (n.d.). a theory of reading: from eye fixations to comprehension. 26. https://doi.org/10.4324/9780429505379-8 kellman, p. j., & massey, c. m. (2013). perceptual learning, cognition, and expertise. in psychology of learning and motivation (vol. 58, pp. 117-165). academic press. https://doi.org/10.1016/b978-0-12407237-4.00004-9 kok, e. m., & jarodzka, h. (2017). before your very eyes: the value and limitations of eye tracking in medical education. medical education, 51(1), 114–122. https://doi.org/10.1111/medu.13066 krebs, m. c., schüler, a., & scheiter, k. (2019). just follow my eyes: the influence of model-observer similarity on eye movement modeling examples. learning and instruction, 61, 126137. https://doi.org/10.1016/j.learninstruc.2018.10.005 litchfield, d., & ball, l. j. (2011). rapid communication: using another’s gaze as an explicit aid to insight problem solving. quarterly journal of experimental psychology, 64(4), 649–656. https://doi.org/10.1080/17470218.2011.558628 lobato, j. (2012). the actor-oriented transfer perspective and its contributions to educational research and practice. educational psychologist, 47(3), 232-247. https://doi.org/10.1080/00461520.2012.693353 mason, l., pluchino, p., & tornatora, m. c. (2015). eye-movement modeling of integrative reading of an illustrated text: effects on processing and learning. contemporary educational psychology, 41, 172–187. https://doi.org/10.1016/j.cedpsych.2015.01.004 mason, l., pluchino, p., & tornatora, m. c. (2016). using eye‐tracking technology as an indirect instruction tool to improve text and picture processing and learning. british journal of educational technology, 47(6), 1083-1095. https://doi.org/10.1080/00461520.2012.693353 mason, l., scheiter, k., & tornatora, m. c. (2017). using eye movements to model the sequence of textpicture processing for multimedia comprehension: using eye movements to model. journal of computer assisted learning, 33(5), 443–460. https://doi.org/10.1111/jcal.12191 mayer, r. e. (1998). cognitive, metacognitive, and motivational aspects of problem solving. instructional science, 26(1-2), 49-63. https://doi.org/10.1007/978-94-017-2243-8_5 mitra, r., mcneal, k. s., & bondell, h. d. (2017). pupillary response to complex interdependent tasks: a cognitive-load theory perspective. behavior research methods, 49, 19051919. https://doi.org/10.3758/s13428-016-0833-y muzyka, j. l. (2021). spectral zoo: combined spectroscopy practice problems for organic chemistry. 12. rayner, k. (ed.). (1992). eye movements and visual cognition: scene perception and reading. springer new york. https://doi.org/10.1007/978-1-4612-2852-3 rayner, k. (2009). the 35th sir frederick bartlett lecture: eye movements and attention in reading, scene perception, and visual search. quarterly journal of experimental psychology, 62(8), 1457–1506. https://doi.org/10.1080/17470210902816461 reiser, b. j. (2004). scaffolding complex learning: the mechanisms of structuring and problematizing student work. in the journal of the learning sciences (pp. 273-304). psychology press. https://doi.org/10.4324/9780203764411-2 renkl, a. (2014). toward an instructionally oriented theory of example based learning. cognitive science, 38(1), 1-37. https://doi.org/10.1111/cogs.12086 rickey, d., & stacy, a. m. (2000). the role of metacognition in learning chemistry. journal of chemical education, 77(7), 915. https://doi.org/10.1021/ed077p915 https://www.youtube.com/watch?v=u5zwduver-c https://doi.org/10.4324/9780429505379-8 https://doi.org/10.1016/b978-0-12-407237-4.00004-9 https://doi.org/10.1016/b978-0-12-407237-4.00004-9 https://doi.org/10.1111/medu.13066 https://doi.org/10.1016/j.learninstruc.2018.10.005 https://doi.org/10.1080/17470218.2011.558628 https://doi.org/10.1080/00461520.2012.693353 https://doi.org/10.1016/j.cedpsych.2015.01.004 https://doi.org/10.1080/00461520.2012.693353 https://doi.org/10.1111/jcal.12191 https://doi.org/10.1007/978-94-017-2243-8_5 https://doi.org/10.3758/s13428-016-0833-y https://doi.org/10.1007/978-1-4612-2852-3 https://doi.org/10.1080/17470210902816461 https://doi.org/10.4324/9780203764411-2 https://doi.org/10.1111/cogs.12086 https://doi.org/10.1021/ed077p915 john & mitra 47 | flr schoenfeld, a. h. (2016). learning to think mathematically: problem solving, metacognition, and sense making in mathematics (reprint). journal of education, 196(2), 138. https://doi.org/10.1177/002205741619600202 schindler, m., & lilienthal, a. j. (2019). domain-specific interpretation of eye tracking data: towards a refined use of the eye-mind hypothesis for the field of geometry. educational studies in mathematics, 101, 123-139. https://doi.org/10.1007/s10649-019-9878-z shannon, c. e. (1948). a mathematical theory of communication. the bell system technical journal, 27(3), 379-423. https://doi.org/10.1002/j.1538-7305.1948.tb01338.x thomas, l. e., & lleras, a. (2007). moving eyes and moving thought: on the spatial compatibility between eye movements and cognition. psychonomic bulletin & review, 14(4), 663–668. https://doi.org/10.3758/bf03196818 tobii pro, a. b. (2014). tobii pro lab. computer software. http://www. tobiipro. com. tunga, y., & cagiltay, k. (2023). looking through the model’s eye: a systematic review of eye movement modeling example studies. education and information technologies, 127. https://doi.org/10.1007/s10639-022-11569-5 underwood, g., & everatt, j. (1992). the role of eye movements in reading: some limitations of the eye-mind assumption. in advances in psychology (vol. 88, pp. 111–169). elsevier. https://doi.org/10.1016/s0166-4115(08)61744-6 van es, e. a., & sherin, m. g. (2002). learning to notice: scaffolding new teachers’ interpretations of classroom interactions. journal of technology and teacher education, 10(4), 571596. https://doi.org/10.1016/j.tate.2006.11.005 van gog, t., jarodzka, h., scheiter, k., gerjets, p., & paas, f. (2009). attention guidance during example study via the model’s eye movements. computers in human behavior, 25(3), 785–791. https://doi.org/10.1016/j.chb.2009.02.007 van marlen, t., van wermeskerken, m., jarodzka, h., & van gog, t. (2016). showing a model’s eye movements in examples does not improve learning of problem-solving tasks. computers in human behavior, 65, 448–459. https://doi.org/10.1016/j.chb.2016.08.041 van merrienboer, j. j. (2013). perspectives on problem solving and instruction. computers & education, 64, 153-160. https://doi.org/10.1016/j.compedu.2012.11.025 van wermeskerken, m., litchfield, d., & van gog, t. (2018). what am i looking at? interpreting dynamic and static gaze displays. cognitive science, 42(1), 220–252. https://doi.org/10.1111/cogs.12484 xie, h., zhao, t., deng, s., peng, j., wang, f., & zhou, z. (2021). using eye movement modelling examples to guide visual attention and foster cognitive performance: a meta analysis. journal of computer assisted learning, 37(4), 1194–1206. https://doi.org/10.1111/jcal.12568 watanabe, t., & sasaki, y. (2015). perceptual learning: toward a comprehensive theory. annual review of psychology, 66, 197. https://doi.org/10.1146/annurev-psych-010814-015214 wu, c. j., & liu, c. y. (2022). refined use of the eye-mind hypothesis for scientific argumentation using multiple representations. instructional science, 50(4), 551-569. https://doi.org/10.1007/s11251-02209581-w zelinsky, g. j., peng, y., & samaras, d. (2013). eye can read your mind: decoding gaze fixations to reveal categorical search targets. journal of vision, 13(14), 10–10. https://doi.org/10.1167/13.14.10 https://doi.org/10.1177/002205741619600202 https://doi.org/10.1007/s10649-019-9878-z https://doi.org/10.1002/j.1538-7305.1948.tb01338.x https://doi.org/10.3758/bf03196818 https://doi.org/10.1007/s10639-022-11569-5 https://doi.org/10.1016/s0166-4115(08)61744-6 https://doi.org/10.1016/j.tate.2006.11.005 https://doi.org/10.1016/j.chb.2009.02.007 https://doi.org/10.1016/j.chb.2016.08.041 https://doi.org/10.1016/j.compedu.2012.11.025 https://doi.org/10.1111/cogs.12484 https://doi.org/10.1111/jcal.12568 https://doi.org/10.1146/annurev-psych-010814-015214 https://doi.org/10.1007/s11251-022-09581-w https://doi.org/10.1007/s11251-022-09581-w https://doi.org/10.1167/13.14.10 john & mitra 48 | flr appendix nmr spectral graph problem nmr spectroscopy is an analytical technique used in chemistry to determine the arrangement of atoms in a substance. the technique involves analysing a graph called an nmr spectrum that shows peaks and valleys representing the different frequencies of radio waves absorbed and emitted by the atoms in the substance. this spectral graph problem requires analysing the position and shape of the signals in the spectrum to deduce the molecular structure. the position of a signal is read from its relative x-coordinate, which is indicative of the electronegative nature of the corresponding component of the molecule. the shape of a signal refers to the splitting pattern and relative height of each peak. interpreting the pattern of peaks and valleys is like solving a puzzle, where each signal represents a piece of information that needs to be put together to solve the molecular structure. solution to the spectral graph of 1-nitro-propane inferring the molecular structure of a compound (in this case c3h7no2) from its h+nmr spectra (shown in fig. 9) requires the participant to piece together three types of information from the spectral graph and they are as follows: ● the x coordinate of each unique signal is the chemical shift value measured in ppm. each signal can be represented symbolically as 𝛿 (chemical shift in ppm). each signal corresponds to a group of chemically equivalent hydrogen atoms that is present in the compound. ● the splitting pattern of each signal, which is inferred from the no. of hydrogen proximal to the carbon (or chemically equivalent group of hydrogen atoms) that corresponds to the signal. ● the integration factor of each signal (change in the integral values before and after the signal), which is inferred from the red integral line running from left to right. the ratio between the integral values of different signals is used to estimate the integration factor of each signal; which reveals the no. of hydrogen atoms associated with the corresponding signal. figure 9. nmr spectra of 1-nitropropane. image from spectral zoo (muzyka, 2021) each signal is referred to by ‘𝛿’ and the chemical shift value corresponding to the signal. the nmr spectra of 1-nitropropane (fig. 3) have three spectral signals, which can be represented using their chemical shift values as 𝛿 (4.4), 𝛿 (2.1) and 𝛿 (1.1) from left to right. the three unique spectral signals indicate the presence of three groups of chemically equivalent hydrogen. the ratios of the integral factors indicate that 𝛿 (4.4) and 𝛿 (2.1) have equal no. of hydrogen atoms while 𝛿 (1.1) correspond to larger no. of hydrogen atoms, with the integration factors revealing that 𝛿 (1.1) correspond to three hydrogen atoms and the other two signals correspond to two hydrogen atoms each. the splitting patterns of 𝛿 (4.4) and (1.1) can be seen to have three peaks (referred to as triplets), while 𝛿 (2.1) has six peaks (referred to as a sextet). the splitting pattern can be reasoned with a simplified version of the “n+1 rule”, according to which a signal with n+1 splits then the carbon corresponding to that signal will have “n” neighbouring hydrogen atoms. it is important to note that while the n+1 rule informs about the hydrogen in adjacent carbons, the integration factor corresponds to the no. of hydrogen in the carbon corresponding to a given signal. 1. introduction 2. theoretical framing of egi tasks 3. conceptualization of egi tasks 3.1 conceptual examples of egi tasks 3.2 conceptual representation of egi tasks 4. experimental setup for developing egi tasks 4.1 nature of the problem 4.2 study design for collecting gaze data and creating egi tasks 5. creation of egi tasks 5.1 spectral graph problem: nmr spectra of 1-nitro-propane 5.2 description of eye gaze models 5.2.1 eye gaze pattern of p1 5.2.2 eye gaze pattern of p2 5.3 egi tasks created from eye gaze data of p1 and p2 6. scope of egi tasks 7. limitations and future work frontline learning research special issue: vol. 13 no. 2 (2025) 10 26 issn 2295-3159 *corresponding author: xin tang, xin.tang@sjtu.edu.cn ; barbara schneider, bschneid@msu.edu doi: https://doi.org/10.14786/flr.v13i2.1313 optimal learning moments in finnish and us science classrooms: a psychological network analysis approach xin tang1,2*, i-chien chen3+, jari lavonen2, barbara schneider4*, joseph krajcik5, katariina salmela-aro2 1school of education, shanghai jiao tong university 2faculty of educational sciences, university of helsinki 3department of social and policy sciences, yuan ze university 4college of education and department of sociology, michigan state university 5create for stem institute, college of education and natural sciences, michigan state university + i-chien chen did this research when she was at msu article received 27 june 2023 / article revised 16 february 2024 / accepted 17 june 2024/ available online 14 march 2025 abstract engagement can be situative and, when it occurs, a number of experiences will co-occur. the present study examined the co-occurred experiences of optimal learning moments (olm), a type of situated engagement, using the novel network analysis and including data from two countries: finland and the us. both samples were from high schools and were measured using the experience sampling method. the finnish sample consisted of 282 students (age = 15-16) and was assessed in science lessons only. the us sample consisted of 533 students at the same age. co-occurrence network analysis showed that, when olm occurred, feelings of concentration, success, in control, and meeting self and others’ expectations appeared frequently. these results were highly consistent between finnish and us science classrooms. further analysis found optimal learning moments were mutually reinforced by the creative experiences, feelings of competitiveness and pride, and the attitudes toward science practices. as a result, an updated optimal learning moment framework was proposed to understand its enhancers, detractors, accelerants, and outcomes in science learning situations. this provides new theoretical accounts regarding the cooccurring experiences of optimal learning moments. keywords: optimal learning moments; situated engagement; network analysis; science learning mailto:xin.tang@sjtu.edu.cn mailto:bschneid@msu.edu tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 11 | f l r 1. introduction optimal learning moments (olm), or also being named as situational engagement (j. inkinen et al., 2019, 2020), refer to the experienced engagement of an activity that an individual feels interested, challenged and capable of achieving it (schneider et al., 2016). the concept of optimal learning moment is built upon the idea of “flow” defined by csikszentmihalyi (1990) as situation-specific instances when an individual is deeply engaged in a task that time loses its temporal boundaries and human needs are suspended. during these moments, an individual experiences higher than average levels of challenge and skill; but neither challenge nor skill overtakes the other, meaning that a particular task is within the boundaries of mastery and not overwhelmingly difficult in terms of one’s current skills. in the learning context, schneider et al. (2016) added a new component – interest– to challenge and skill, to constitute the optimal experience of learning in that a student must have a positive preand postdispositional affection, i.e., interest, towards the learning objects. thus, unlike flow which only constitutes the elements of challenge and skill, the optimal learning moments are operationalized as high levels of challenge, skill, and interest in task engagement (schneider et al., 2016). research has shown that when students reported more olm they reported more feelings of enjoyment, success and happiness, and fewer feelings of confusion, stress and anxiety, and better attitudes towards science (schneider et al., 2016), and higher course grades in physics (hendolin, 2016). however, previous studies have at least two limitations. first, though the relationships among olm, and positive and negative emotions (e.g., enjoyment, happiness, boredom, confusion) have been examined previously (schneider et al., 2016), other important experiences such as social-related emotions (e.g., feelings of cooperative, or competitive), persistent behaviors (duckworth et al., 2007; tang et al., 2019), and creative practices (e.g., exploring) are missing. in other words, a comprehensive understanding of olm-related experiences is still needed. second, the common approach for olmrelated experiences is the pairwise correlation (hendolin, 2016; schneider et al., 2016; except some regression analyses, e.g., j. inkinen et al., 2019, 2020), which may limit us to understand the olm from a holistic perspective. thus, the aim of this study is to extend the understanding of the optimal learning moments by having more co-existed experiences. to fill this purpose, a novel approachnetwork analysis was applied to capture a holistic landscape of optimal learning moments. moreover, cross-country consistencies were examined to find preliminary cross-validation evidence of the findings. 1.1 optimal learning moments and the co-occurred experiences while being in optimal learning moments requires the simultaneous feelings of being skilled, challenged, and interested, there are several experiences that may promote or hinder optimal learning moments (schneider et al., 2016). broadly speaking, three types of experiences are important for the optimal learning moments. there are enhancers, including enjoyment, successful, happy, confidence and being active (pekrun, 2006; schneider et al., 2016). those are positive emotions or affective experiences that co-exist with optimal learning moment. there are also detractors, such as boredom and confusion as indicators of negative emotions (m. inkinen et al., 2014; pekrun, 2006). the last category is the accelerants, including feelings of stressed and anxious, that might positively or negatively associate with optimal learning moments depending on their level (schneider et al., 2016). it has been shown that some extent of stress and anxiety might stimulate learning but not excessive amounts of them (de anda et al., 2000). tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 12 | f l r figure 1. the conceptual framework for optimal learning moments. copied from schneider et al. (2016) with permission. despite that the above-mentioned emotions were studied with optimal learning moments (schneider et al., 2016), a few more important emotions and experiences are still worth considering, and they are yet to be examined. one important type of emotions is the social-related affective experiences, such as loneliness, cooperation, competition, and social expectations (van kleef, 2009). given that social-related affective experiences are diverse, their relationships with optimal learning moments should be diverse as well. for instance, the feeling of loneliness (hawkley & cacioppo, 2010), a common indicator of negative emotion may reduce the chance to observe optimal learning moments. in contrast, the feeling of competition, may act as an accelerant since it can alter the interpretation of emotions (van doorn et al., 2012). however, as a whole, there is lack of research on the relationships between optimal learning moments and social-related feelings and experiences. another less-addressed experiences are the feelings and behaviors of persistence (duckworth et al., 2007; tang et al., 2019). according to the flow theory (csikszentmihalyi, 1990), the flow experience is engaging but also shown as a transient state, and thus the persistence of flow experience is less known. now since the optimal learning moments add an interest component to skill and challenge, whether they are accompanied by persistent experiences are unknown. a final missing knowledge is creative practices and experiences. flow has been found closely relating to creative experiences (csikszentmihalyi, 1997), however, again, whether the optimal learning moments co-exist with creative practices and experiences is largely unknown. 1.2 psychological network analysis approach to find the relationships among optimal learning moments and various emotions and feelings, this study applied psychological network analysis (borsboom et al., 2021; tang, lee, et al., 2022). tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 13 | f l r unlike social network analysis, which primarily examines the dynamics of individuals, psychological network analysis focuses on the intricate dynamics of variables. in the present study, these variables are multiple emotional states (e.g., happy, confidence, confused, anxious) that can be present at the same moment/ or situation. as a holistic approach, in psychological network analysis, relationships (known as the edges/ties in network terminology) between any two variables (known as the nodes) are accounted for other variables in the network. thus, it provides the conditional relationships among the variables. since networks can present the interdependencies among variables, they offer several important contributions. first, psychological network analysis helps to understand the constructs from a systematic perspective since any given connection between two nodes is conditioned by the existence of other nodes. second, once a psychological network has been constructed, the most significant node(s) can be identified to see which one plays a critical role in the network. this, on the one hand, informs the relative importance of variables, and on the other hand gives implications for the most desirable intervention targets. moreover, the closeness among variables within a psychological network can be further examined to see which set of variables functions together. this is particularly useful when multidimensional constructs are examined. third, it allows systematic cross-network comparisons among attributes (e.g., classroom, school, country). most commonly, psychological network analysis has been conducted based on zero-order or first-order correlations (borsboom et al., 2021; christensen et al., 2020). however, the present study chose the co-occurrence psychological network analysis, which aims to show how often (or rarely) two variables co-occur at high levels. this is not only better in line with the typical data pre-process for the optimal learning moments (j. inkinen et al., 2019, 2020; schneider et al., 2016), but also can have several strengths over correlation-based network analysis (moeller et al., 2018; trampe et al., 2015). first, co-occurrence analysis avoids misinterpretation of correlations. typically, when interpreting high positive correlations between two variables (e.g., a & b), researchers conclude that variable a is “high” when variable b is “high”, even though, in reality, both a and b might be rated at a low level on the original scale. because correlations only denote how the ratings of two variables are aligned consistently, how these two variables occur together at a high level is not necessarily revealed in correlation analysis. second, a frequent co-occurrence may occur even if two variables correlate negatively, which could only be detected using co-occurrence analysis. thus, the co-occurrence psychological network analysis can shed unique insights into the feelings and experiences that co-exist with optimal learning moments. 1.3 the present study in sum, this study aimed to understand the co-existed experiences of optimal learning moments, a certain type of situated engagement including situated interest, skill, and challenge. in addition, we used a network analysis approach and included sample from two countries (i.e., finland and the us). as it has been reviewed above, except for those outlined in the original model (schneider, et al., 2016), no specific hypotheses can be proposed regarding the olm, social-related affective experiences, feelings of persistence, and creative experiences. 2. method 2.1 participants this study used samples from finland and the us. both samples’ situational feelings and experiences were obtained using experience sampling method (esm) questionnaires delivered via smartphones. the first sample consisted of 282 first-year upper secondary school students from nine classes in three different schools in helsinki, finland. data were collected from the fall of 2018 to the spring of 2019. the students participated in a project-based learning (pbl; krajcik & shin, 2014) module that tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 14 | f l r consisted of six lessons. this module focused on the basics of newtonian mechanics and aimed to engage students in solving real-world scientific problems by exploring and participating in collaborative inquiry. in each lesson (about 75 minutes), the esm questionnaire alerted the students three times: at the beginning, middle, and end of the lesson. therefore, each student received 18 signals from the smartphone. overall, the finnish sample data comprised 3882 responses. the second sample comprised 533 us high school students from 28 classes in 22 schools in the school year of 2018 to 2019. the topic of the lessons was the same as in the previous data-collection. in this sample, we used data collected during science lessons. the smartphones were programmed to alert the students randomly 6-8 times per day (at least 3-4 times when they had science lessons) over a period of 4 days. in total, the data comprised 3234 responses/situations (average response per person during the science lesson was 6.02 signals). 2.2 measures in both samples, the students were signaled to answer an esm questionnaire in which they first had to indicate what kinds of activities (e.g., listening, discussion) they were engaged in. then they were asked to report their emotions, feelings, and behaviors when they received the signal. we selected those experiences that were assessed in both two countries to compare their networks. three components of optimal learning moments were assessed by answering: a) are you interested in what you did? (interest); b) did you feel skilled in what you did? (skill); c) was your work challenging? (challenge). other example questions on emotions, feelings, and behaviors includes: did you feel happy (happy)? how well did you focus (concentrate)? did you cope with your work (control)? did your performance meet the expectations of others (other-expect)? a full list of items can be found in the appendix. all the items were rated on a scale of 1 (not at all) to 4 (very much). 2.3 analysis approach to perform co-occurrence analysis, we first dichotomized all our situational experiences at the scale midpoint. that is, the scores of 3 and 4 were re-coded as 1, and the rest (i.e., 0 and 1) will be recoded as 0. a relative index of edge weight was also calculated to present the most frequent experiences that co-existing with the optimal learning moments (tang, renninger, et al., 2022). to do so, the counts of the variable paired with the optimal learning moments were divided by the total counts of the optimal learning moments 𝐾𝑖𝑗 𝐾𝑖 . in this way, we will know the likelihood that a variable was co-occurring with the optimal learning moments. moreover, community detection algorithm was applied to identify the close correlates for optimal learning moments. the louvain community detection algorithm was employed as it has shown better performance than the walktrap algorithm (see suggestions from christensen, golino, & silvia, 2020). the analyses of co-occurrence networks were conducted using r-package igraph (csardi & nepusz, 2006). finally, we conducted network comparison tests (nct; van borkulo et al., 2017) to understand the similarities and differences between two countries’ networks in terms of pairs with optimal learning moments. both network structure invariance tests (test m; a test of connection strength matrix) and global connectivity invariance tests (test s; a test of weighted sum of absolute connections) were performed. the tests were done by using r-package network comparison test (van borkulo et al., 2017). we also tested the individual edge differences between networks in terms of pairs with optimal learning moments. in other words, the connections between optimal learning moments and other nodes were compared. given that multiple comparisons were performed in this step, and p-values were adjusted using benjamini-hochberg method (thissen et al., 2002). tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 15 | f l r 3. results 3.1 co-occurrences with optimal learning moments the results of co-occurrences for optimal learning moments can be seen in table 1 for finnish students and table 2 for us students. in finnish science classrooms, optimal learning moments appeared 475 times. given that we received 3882 responses in total, this means the chance to observe optimal learning moments in finnish science classrooms was about 12% (475/3882 = 0.1223). when optimal learning moments occurred, feelings of concentration (edge=432; 90.95%), enjoyment (edge=422; 88.84%), success (edge=415; 87.37%), meeting self-expectations (edge=414; 87.16%), meeting others’ expectations (edge=411; 86.53%), feelings of control (edge=411; 86.53%) were the top five cooccurring experiences. the results also showed that the general within-level correlations between optimal learning moments and other experiences were low (max = 0.17), although most were significant. the co-occurrence of optimal learning moments with social-related feelings were salient. when optimal learning moments happened, it was unlikely to be lonely (8.21%). the feeling of being cooperative was also high (85.05%), though was not among the top 5 list. the feeling of competitive co-occurred with optimal learning moments moderately (39.58%), implying that it may serve as an accelerant factor. the co-occurrence likelihoods of creative experiences (i.e., practices of exploring ideas, using imagination, finding new solutions) with optimal learning moments were moderate (58.95%-68.84%). in the us science classrooms, optimal learning moments appeared 534 times. given that we received 3234 responses in total, this means the chance to observe optimal learning moments in us science classrooms was about 16% (534/3234 = 0.1665). when optimal learning moments occurred, feelings of success (edge=527; 91.81%), concentration (edge=513; 89.37%), meeting self-expectations (edge=496; 86.41%), meeting others’ expectations (edge=477; 83.10%), feelings of control (edge=465; 81.01%) were the top five co-occurring experiences. the results also showed that the general withinlevel correlations between optimal learning moments and other experiences were low but still significant (max = 0.21). compared to the correlation table using finnish data, the within-level correlations between optimal learning moments and other experiences were relatively higher, particularly for the top 10 emotional experiences. the co-occurrence of optimal learning moments with social-related feelings was salient. when optimal learning moments happened, it was unlikely to be bored (42.33%). the feeling of cooperative was also high (77.00%). the feeling of competitive co-occurred with optimal learning moments moderately (42.58%), implying that it may serve as an accelerant factor. the co-occurrence likelihoods of creative experiences (i.e., using imagination, finding solutions) with optimal learning moments were moderate (58.71%-67.25%). 3.2 network community while the previous co-occurrence analysis identified the uni-directional information on the pair connection with optimal learning moments. that is, the likelihood of occurrence when optimal learning moments occurred 1 . the communities among co-occurrence networks indicated the feelings and experiences that are mutually closely appeared (i.e., the most closely associated variable groups). each variable belongs to a certain community. variables belonging to the different community mean they are not closely related. with this information, we can then infer the positions of variables in the olm framework. according to the finnish results (see figure 2), creative experiences and feelings of 1 for a variable x, its occurrence likelihood when olm occurred does not equal to the likelihood of observing olm when x occurred. tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 16 | f l r competitiveness and pride were in the same community with the optimal learning moments. positive emotions, such as happiness, excitement, and enjoyment, tended to cluster together. the same happened for negative emotions (e.g., feelings of anxious, lonely, stress, bored, confused, give up) as a cluster. in the us science classrooms (see figure 3), optimal learning moments formed the group with the attitudes toward scientific practices (i.e., feeling that the practices is important to self and to the future). as in the finnish science classrooms, positive and active experiences (e.g., feelings of happy, excited, competitive, proud, confident, active) were in the same group, whereas negative experiences (e.g., feelings of anxious, lonely, stress, bored, confused, give up) were together. 3.3 network comparison finally, we examined the equivalence of the two countries’ networks by conducting the network comparison tests based on the pearson’s correlations (see full results in the online supplementary materials, https://osf.io/3db4r/?view_only=3e79a5389ab04fd7bd599d723052be7e). the global strength of the networks (i.e., the global degree of connectivity among variables) was insignificant (test statistic s = 5.102; finnish network strength = 67.269, the us’s network strength = 62.166; p = .12). the network invariance test that aimed to examine the differences in paired edges was significant (test statistic m =1.022, p <.001). further checking of the individual edges showed that they lied on the pair of olm and concentration, and of olm and future importance of science practice. the us students, in comparison to finnish students, reported more feelings of concentration and future importance when reporting optimal learning moments. https://osf.io/3db4r/?view_only=3e79a5389ab04fd7bd599d723052be7e tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 17 | f l r table 1 co-occurrences of olm-emotion and motivation pairs in the finnish sample rank node1 node2 edge weight % of all edges % of olm edges1 as reference situationallevel correlation person-level correlation olm 1 concentrate 432 5.39% 90.95% 0.09*** 0.37*** 2 enjoy 422 5.27% 88.84% 0.12*** 0.50*** 3 success 415 5.18% 87.37% 0.11*** 0.41*** 4 self expect 414 5.17% 87.16% 0.09*** 0.31*** 5 control 411 5.13% 86.53% 0.06*** 0.30*** 5 other expect 411 5.13% 86.53% 0.08*** 0.34*** 7 happy 409 5.11% 86.11% 0.11*** 0.42*** 8 cooperative 404 5.04% 85.05% 0.09*** 0.30*** 9 active 399 4.98% 84.00% 0.14*** 0.45*** 10 important you 393 4.91% 82.74% 0.03 0.48*** 11 excited 390 4.87% 82.11% 0.10*** 0.49*** 12 confident 385 4.81% 81.05% 0.11*** 0.45*** 13 important future 363 4.53% 76.42% 0.08*** 0.42*** 14 time 332 4.14% 69.89% 0.13*** 0.41*** 15 exploring 327 4.08% 68.84% 0.17*** 0.49*** 16 imagination 318 3.97% 66.95% 0.17*** 0.45*** 17 solutions 280 3.50% 58.95% 0.14*** 0.45*** 18 proud 276 3.45% 58.11% 0.16*** 0.49*** 19 competitive 188 2.35% 39.58% 0.10*** 0.40*** 20 confused 132 1.65% 27.79% 0.04** 0.06 21 stress 130 1.62% 27.37% -0.01 -0.03 22 give up 103 1.29% 21.68% 0.07*** 0.04 23 bored 89 1.11% 18.74% -0.09*** -0.23*** 24 anxious 73 0.91% 15.37% 0.01 0.01 25 lonely 39 0.49% 8.21% 0.02 -0.01 note. 1number of olm self-edges is 475; ** p < .01, *** p < .001 tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 18 | f l r table 2. co-occurrences of olm-emotion and motivation pairs in the us sample. rank node1 node2 edge weight % of all edges % of olm edges1 as reference situationallevel correlation person-level correlation olm 1 success 527 5.47% 91.81% 0.15*** 0.30*** 2 concentrate 513 5.33% 89.37% 0.21*** 0.47*** 3 self expect 496 5.15% 86.41% 0.16*** 0.24*** 4 other expect 477 4.95% 83.10% 0.15*** 0.24*** 5 control 465 4.83% 81.01% 0.15*** 0.33*** 6 enjoy 455 4.72% 79.27% 0.20*** 0.53*** 7 happy 451 4.68% 78.57% 0.17*** 0.43*** 8 cooperative 442 4.59% 77.00% 0.12*** 0.38*** 9 important you 440 4.57% 76.66% 0.21*** 0.55*** 10 confident 419 4.35% 73.00% 0.15*** 0.40*** 11 time 419 4.35% 73.00% 0.17*** 0.54*** 12 solutions 386 4.01% 67.25% 0.14*** 0.49*** 13 excited 377 3.91% 65.68% 0.17*** 0.50*** 14 exploring 371 3.85% 64.63% 0.13*** 0.49*** 15 active 362 3.76% 63.07% 0.16*** 0.45*** 16 proud 360 3.74% 62.72% 0.18*** 0.49*** 17 important future 342 3.55% 59.58% 0.19*** 0.53*** 18 imagination 337 3.50% 58.71% 0.17*** 0.48*** 19 competitive 244 2.53% 42.51% 0.15*** 0.45*** 20 bored 243 2.52% 42.33% -0.09*** -0.12** 21 stress 227 2.36% 39.55% 0.02 0.07 22 anxious 217 2.25% 37.80% 0.06** 0.25*** 23 confused 198 2.06% 34.49% 0.01 0.15*** 24 give up 163 1.69% 28.40% 0.03 0.13** 25 lonely 127 1.32% 22.13% 0.01 0.12** note. 1number of olm self-edges is 574; ** p < .01, *** p < .001 tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 19 | f l r figure 2. co-occurrences networks with communities among finnish sample. notes. nodes under same colours are in the same community; community 1: olm, competitive, proud, imagination, solutions, exploring; community 2: happy, excited, enjoy; community 3: anxious, lonely, stress, bored, confused, give up; community 4: cooperative; community 5: confident, success; community 6: active; community 7: concentrate; community 8: control; community 9: important you, important future; community 10: other expect, self expect; community 11: time. ties between nodes are conditional correlations between them. tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 20 | f l r figure 3. co-occurrences networks with communities among us sample. notes. nodes under same colours are in the same community; community 1: olm, important you, important future; community 2: happy, excited, competitive, proud, confident, active; community 3: anxious, lonely, stress, bored, confused, give up; community 4: cooperative; community 5: concentrate; community 6: enjoy, time.; community 7: control, success, other expect, self expect; community 8: imagination, solutions, exploring. ties between nodes are conditional correlations between them. tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 21 | f l r 4. discussion while optimal learning moments (olm) were examined with positive and negative emotions and experiences, there was a lack of comprehensive knowledge of it; particularly regarding its relationships with social-related experiences (e.g., cooperation, proud), feeling of persistence (e.g., not giving up), and creative experiences (e.g., using imagination). the present study thus filled those research gaps by including more experiences and by using two countries’ samples with a holistic analysis approach. we found that when students indicated that they were in olm, they were most likely concentrating on what they were doing, experiencing positive emotions (e.g., enjoyment, happiness) and success, feeling in control, doing things important to them, and fulfilling both their own and others’ expectations. moreover, the students in olm reported rarely being bored, confused, lonely, or anxious, and were unlikely to give up on what they were doing. however, we also found that social-related feelings of competitiveness and proud should be better understood as olm’s accelerants; creative experiences were enhancers. in sum, our findings corroborate with those of schneider et al. (2016) and expand the model with more experiences. we thus propose an updated olm model in science learning situations (see figure 4). figure 4. an updated optimal learning moments model in science classrooms. 4.1 outcomes of olm using co-occurrence network analysis, we found that when olm occurred, it was high likely to observe the experiences of concentration, efficacy, enjoyment, happiness, and success. these findings align with the speculations of schneider et al. (2016) and the findings of flow theory (csikszentmihalyi, 1990). when students reported high-level challenge, skill, and interest, they were most likely in the zone of flow. as a result, they were completely engaged in the situational task, thus forgetting the time and feeling the pleasure and inner peace. different from schneider et al. (2016), those positive experiences were placed in the outcomes box not in the enhancers box as we did not observe olm and positive experiences in the same network community. however, this did not mean the positive experience will optimal learning moments most frequent outcomes: -concentration -success -efficacy (feeling in control) -meeting own and others’ expectations -enjoyment, happiness detractors: -bored -confused -lonely -anxious -stress -thoughts of giving up enhancers: -attitude toward science -creative experiences (e.g., exploring, imagination) accelerants: -competitive -proud tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 22 | f l r not facilitate olm. in the long run, as students engaged in more positive experiences in those situations, they were more likely to seek them for the olm. 4.2 enhancers of olm the network community analysis tended to suggest that olm enhancers were attitudes toward science (e.g., importance of science for self and the future) and creative experiences (e.g., using imagination, exploring multiple solutions). that is, when students found the science learning situations were important for themselves and for their future, they were in turn most likely engaged in the optimal learning moments. previous intervention studies (e.g., hulleman et al., 2010) showed that interest can be piqued by improving the utility value (i.e., instrumental importance). thus, it is not surprising to see olm was enhanced when students saw the value of science. additionally, we found that olm was enhanced when students were doing creative activities, such as exploring tasks, using imaginations, or finding new solutions. creative activities are typically challenging (runco & jaeger, 2012); this, in turn, facilitates the challenge component of olm and enhances the chance to observe the olm as a whole. 4.3 detractors of olm in line with schneider et al. (2016), the present study found that the feelings of boredom and confusion rarely co-existed with olm, thus implying that they were detractors of olm. since olm was co-occurred with many positive experiences, it is not surprising to see negative emotions like boredom and confusion show opposite direction. this finding extended to a negative social-related emotion: loneliness. when people feel lonely, they tend not to be very active (pels & kleinert, 2016), thus are less likely to seek for challenges. the same goes for the thoughts of giving up. when this happens, people stop taking on new situations, this thus will hinder the development of olm. however, different from schneider et al. (2016), the present study found anxious and stressed experiences should be better understood as detractors rather than accelerants given their low cooccurrence with olm. this may due to the fact that olm co-existed with positive emotions (e.g., enjoyment, happiness), thus it is very unlikely to observe anxious and stressed emotions at the same time. 4.4 accelerants of olm we found that the feelings of competitiveness and pride, on the one hand, were in the same group as olm through the network community analysis. on the other hand, the co-occurred chance of them was moderate (40%-60%) when olm occurred. that means, they have complex relationships with olm, and thus can be regarded as accelerants. this is particularly evident among finnish high school students. being accelerants means their associations with olm can be either positive or negative depending on their level of strength. previous studies found that competitiveness, as a part of social pressure, can stimulate motivation and performance when it is at the benign level (tauer & harackiewicz, 2004). however, when the level of competitiveness is too high, it will most likely hurt learning (murayama & elliot, 2012). similarly, the feeling of pride can also be detrimental to learning when it is at an excess level (tracy & robins, 2007). tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 23 | f l r 4.5 implications and limitations using experience sampling method, this study examined a certain type of situated engagement, i.e., optimal learning moments (olm), that is formed as high levels of interest, skill, and challenge in learning situations. we took a further step by including additional emotions and experiences, using data from two countries, and applying the holistic psychological network approach. the findings extend the understanding of olm and thus provide a broader framework for olm in science learning situations. this, as a consequence, can shed important theoretical and practical implications to the field. first, an updated olm framework was proposed on the basis of schneider et al. (2016). this updated framework enriched the understandings of olm’s enhancers, detractors, accelerants, and outcomes. more importantly, their relationships have been elucidated in the model. second, our findings implied that we should enhance students’ attitudes toward science and provide creative experiences in order to facilitate optimal learning moments in science lessons. when students feel science is important to them, they are more willing to take challenging on tasks and put more effort into the task for further engagement. third, we also found that we can harness the power of competitiveness and pride to accelerate olm, however, they should be utilized with caution and at the benign level. social-related emotions and experiences (e.g., social/peer pressure) can be beneficial for learning and engagement, however, they should be enhanced under control. last, the study also showed that network analysis can serve as a powerful tool to understand the complex associations among a group of variables. it can shed lights on the bi-directional and unidirectional influences among them. despite the novel findings and contributions, some limitations of the study still need to be considered. first, the emotions or feelings were measured using self-reported items. although selfreport is still a vital way to check the mental experiences, it is also subject to reporting bias. second, it is still challenging to make causal inference using cross-sectional data, even though psychological network analysis can partly help to make the causal inferences (lee et al., 2022). thus, the causal interpretation of our findings should be approached with caution. keypoints optimal learning moments as situated engagement lead to joyful and successful experiences positive attitudes and creative experiences facilitate optimal learning moments negative experiences like boredom and anxiety deplete optimal learning moments competitiveness and pride can increase or decrease optimal learning moments depending on situations acknowledgments the work has been supported by the research council of finland (no. 340794; 345117 and 336138) and national science foundation under pire no.1450756. the opinions expressed here are those of the authors and do not represent the views of the funding agency. tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 24 | f l r references borsboom, d., deserno, m. k., rhemtulla, m., epskamp, s., fried, e. i., mcnally, r. j., robinaugh, d. j., perugini, m., dalege, j., costantini, g., isvoranu, a.-m., wysocki, a. c., van borkulo, c. d., van bork, r., & waldorp, l. j. (2021). network analysis of multivariate data in psychological science. nature reviews methods primers, 1(1), 58. https://doi.org/10.1038/s43586-021-00055-w christensen, a. p., golino, h., & silvia, p. j. (2020). a psychometric network perspective on the validity and validation of personality trait questionnaires. european journal of personality, per.2265. https://doi.org/10.1002/per.2265 csardi, g., & nepusz, t. (2006). the igraph software package for complex network research. interjournal, complex sy, 1695. http://igraph.org csikszentmihalyi, m. (1990). flow: the psychology of optimal experience. harper and row. csikszentmihalyi, m. (1997). flow and the psychology of discovery and invention. harperperennial, new york. https://doi.org/10.1037/e586602011-001 duckworth, a. l., peterson, c., matthews, m. d., & kelly, d. r. (2007). grit: perseverance and passion for long-term goals. journal of personality and social psychology, 92(6), 1087–1101. https://doi.org/10.1037/0022-3514.92.6.1087 hawkley, l. c., & cacioppo, j. t. (2010). loneliness matters: a theoretical and empirical review of consequences and mechanisms. annals of behavioral medicine, 40(2), 218–227. https://doi.org/10.1007/s12160-010-9210-8 hendolin, i. (2016). exploring optimal learning moments at tutorial sessions. 2016 physics education research conference proceedings, 144–147. https://doi.org/10.1119/perc.2016.pr.031 hulleman, c. s., godes, o., hendricks, b. l., & ... (2010). enhancing interest and performance with a utility value intervention. journal of …. https://psycnet.apa.org/record/2010-21220-001 inkinen, j., klager, c., juuti, k., schneider, b., salmela‐aro, k., krajcik, j., & lavonen, j. (2020). high school students’ situational engagement associated with scientific practices in designed science learning situations. science education, 104(4), 667–692. https://doi.org/10.1002/sce.21570 inkinen, j., klager, c., schneider, b., juuti, k., krajcik, j., lavonen, j., & salmela-aro, k. (2019). science classroom activities and student situational engagement. international journal of science education, 41(3), 316–329. https://doi.org/10.1080/09500693.2018.1549372 inkinen, m., lonka, k., hakkarainen, k., muukkonen, h., litmanen, t., & salmela-aro, k. (2014). the interface between core affects and the challenge–skill relationship. journal of happiness studies, 15(4), 891–913. https://doi.org/10.1007/s10902-013-9455-6 krajcik, j. s., & shin, n. (2014). project-based learning. in r. k. sawyer (ed.), the cambridge handbook of the learning sciences (pp. 275–297). cambridge university press. https://doi.org/10.1017/cbo9781139519526.018 moeller, j., ivcevic, z., brackett, m. a., & white, a. e. (2018). mixed emotions: network analyses of intra-individual co-occurrences within and across situations. emotion, 18(8), 1106–1121. https://doi.org/10.1037/emo0000419 tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 25 | f l r murayama, k., & elliot, a. j. (2012). the competition–performance relation: a meta-analytic review and test of the opposing processes model of competition and performance. psychological bulletin, 138(6), 1035–1070. https://doi.org/10.1037/a0028324 pekrun, r. (2006). the control-value theory of achievement emotions: assumptions, corollaries, and implications for educational research and practice. educational psychology review, 18(4), 315– 341. https://doi.org/10.1007/s10648-006-9029-9 pels, f., & kleinert, j. (2016). loneliness and physical activity: a systematic review. international review of sport and exercise psychology, 9(1), 231–260. https://doi.org/10.1080/1750984x.2016.1177849 runco, m. a., & jaeger, g. j. (2012). the standard definition of creativity. creativity research journal, 24(1), 92–96. https://doi.org/10.1080/10400419.2012.650092 schneider, b., krajcik, j., lavonen, j., salmela-aro, k., broda, m., spicer, j., bruner, j., moeller, j., linnansaari, j., juuti, k., & viljaranta, j. (2016). investigating optimal learning moments in u.s. and finnish science classes. journal of research in science teaching, 53(3), 400–421. https://doi.org/10.1002/tea.21306 tang, x., lee, h. r., wan, s., gaspard, h., & salmela-aro, k. (2022). situating expectancies and subjective task values across grade levels, domains, and countries: a network approach. aera open, 8, 233285842211171. https://doi.org/10.1177/23328584221117168 tang, x., renninger, k. a., hidi, s. e., murayama, k., lavonen, j., & salmela-aro, k. (2022). the differences and similarities between curiosity and interest: meta-analysis and network analyses. learning and instruction, 80, 101628. https://doi.org/10.1016/j.learninstruc.2022.101628 tang, x., wang, m.-t., guo, j., & salmela-aro, k. (2019). building grit: the longitudinal pathways between mindset, commitment, grit, and academic outcomes. journal of youth and adolescence, 48(5), 850–863. https://doi.org/10.1007/s10964-019-00998-0 tauer, j. m., & harackiewicz, j. m. (2004). the effects of cooperation and competition on intrinsic motivation and performance. journal of personality and social psychology, 86(6), 849–861. https://doi.org/10.1037/0022-3514.86.6.849 thissen, d., steinberg, l., & kuang, d. (2002). quick and easy implementation of the benjaminihochberg procedure for controlling the false positive rate in multiple comparisons. journal of educational and behavioral statistics, 27(1), 77–83. https://doi.org/10.3102/10769986027001077 tracy, j. l., & robins, r. w. (2007). emerging insights into the nature and function of pride. current directions in psychological science, 16(3), 147–150. https://doi.org/10.1111/j.14678721.2007.00493.x trampe, d., quoidbach, j., & taquet, m. (2015). emotions in everyday life. plos one, 10(12), e0145450. https://doi.org/10.1371/journal.pone.0145450 van borkulo, c. d., boschloo, l., kossakowski, j. j., tio, p., schoevers, r. a., borsboom, d., & waldorp, l. j. (2017). comparing network structures on three aspects: a permutation test. journal of statistical software. https://doi.org/10.13140/rg.2.2.29455.38569 van doorn, e. a., heerdink, m. w., & van kleef, g. a. (2012). emotion and the construal of social situations: inferences of cooperation versus competition from expressions of anger, happiness, tang et al special issue: perspectives on momentary engagement and learning situated in classroom contexts 26 | f l r and disappointment. cognition & emotion, 26(3), 442–461. https://doi.org/10.1080/02699931.2011.648174 van kleef, g. a. (2009). how emotions regulate social life. current directions in psychological science, 18(3), 184–188. https://doi.org/10.1111/j.1467-8721.2009.01633.x appendix experiences measured in the study what do you feel and think about the activity you did (1 not at all ….. 4 very much) • did you feel happy? • did you feel excited? • did you feel anxious? • did you feel competitive? • did you feel lonely? • did you feel stress? • did you feel proud? • did you feel cooperative? • did you feel bored? • did you feel confident? • did you feel confused? • did you feel active? • were you interested in what you did? (interest) • did you feel skilled in what you did? (skill) • was your work challenging? (challenge) • did you feel that you wanted to give up? (give up) • how well did you focus? (concentrate) • did you like what you did? (enjoy) • did you manage your work? (control) • did you succeed? (success) • was it what you did important to you? (important to you) • was it what you did important for your future? (important to future) • did your performance meet the expectations of others? (other expect) • did you do follow your own expectations? (self expect) • were you immersed in what you did not notice the passage of time? (time) • while working… i used my imagination (imagination) • while working… solving problems with multiple answers (solutions) • while working… i tried different solutions for exploring (exploring) notes. bolded and underlined ones are newly added experiences in addition to those in schneider et al. (2016). frontline learning research vol. 12 no. 1 (2024) 16 33 issn 2295-3159 corresponding author: charlott sellberg, sem sælands vei 7, 0371 oslo, norway. charlott.sellberg@iped.uio.no doi: https://doi.org/10.14786/flr.v12i1.1217 the development of visual expertise in a virtual environment: a case of maritime pilots in training charlott sellberg1, elin nordenström2 & roger säljö2 1 university of oslo, norway 2 university of gothenburg, sweden article received 4 january 2023 / article revised 9 october / accepted 10 december / available online 12 january abstract this study connects to an ongoing discussion about the limits and affordances of simulators as realistic and relevant contexts for professional learning, in this case in the development of visual expertise. earlier studies of simulator-based maritime pilot training conclude that there are risks associated with so-called negative skills transfer due to a lack of photorealism in simulator environments. the aim of this study is to carefully examine how visual expertise develops in and through training in a simulated environment. through a practice-based approach to the development of visual expertise, and by using qualitative interaction analysis of video recorded training sessions, the analytical focus is directed towards maritime pilot trainees’ talk about imperfections and inconsistencies in the virtual environment during exercises in a high-fidelity bridge simulator. considering the multi-layered nature of the maritime pilot’s visual expertise, findings show that the maritime pilots in training noticed and adapted to the specific methodological and technological challenges when manoeuvring a simulated vessel. during such reflection-in-action, they also commented on and explored the differences between, navigating in a simulator, on the one hand, and, on the other hand, navigating on board a ship. instead of concluding that there is a risk for negative skills transfer that follows from the differences between the two contexts of navigating, we argue that the challenges introduced by representations encountered when training in a virtual environment may add to the expertise of the trainees and lead to enriched conceptual, methodological, and technical knowledge regarding the specificities of visually demanding and ambiguous navigation situations. in this way, this study contributes to advance our understanding of learning in virtual environments to the frontline of learning research. keywords: visual expertise, simulator fidelity, skills transfer, conceptual change mailto:charlott.sellberg@iped.uio.no https://doi.org/10.14786/flr.v12i1.1217 sellberg, nordenström, & säljö 17 | f l r 1. introduction a maritime pilot is an expert navigator, specialised to support maritime officers to manoeuvre their vessel in a challenging maritime territory. through the maritime pilot's intimate knowledge of the fairway, and the experience of manoeuvring many different types of vessels, the maritime pilot contributes to ensuring that maritime and environmental safety can be maintained when vessels operate in pilotage-obliged water (see, e.g., lützhöft & nyce, 2006). in addition to skills in ship manoeuvring, navigation, and seamanship, the ability to interact with various types of technologies, cultures and crews is also required by pilots, as each ship is unique in terms of equipment and instruments. central to the expertise of the maritime pilot is the skilled perception, interpretation and evaluation of the domain's visual materials (see gegenfurtner et al., 2019). building on the pioneering work on “professional vision” by goodwin (1994), visual expertise is described as the “superior performance of professionals in processing domain-specific visual information” (gegenfurtner et al., 2022, p. 3). this ability is particularly important in visually intense domains, such as medicine, aviation, and, as in this case, maritime navigation, where the skilled perception, interpretation, and evaluation of critical visual elements by professionals is crucial. traditionally, maritime pilot training has been based on apprenticeship on board ships, where the skills and competencies specified above have been developed through years of experience travelling a specific territory. contemporary maritime pilot training instead combines periods of apprenticeship with training on board miniature model ships as well as simulator-based training. while these novel methods are claimed to offer unique and, in some respects, better training opportunities, for example by including exercises of handling risky and unusual situations in safe environments (kim et al., 2021), there are also concerns expressed about how well such simulated activities make efficient learning possible. studies conducted in various safety-critical domains warn that lack of fidelity, i.e., resemblance with the real environment, and other shortcomings in the simulator, may lead to trainees learning how to deal with the inaccurate conditions of the simulator environment rather than the conditions that apply in their future work settings, so-called negative skills transfer (see e.g., hontvedt, 2015; petersen et al. 2022; shul et al., 2019; taber, 2013). in a study of simulation-based training of helicopter underwater escape performance, for instance, taber (2013) stresses that “a specific skill practiced in a simulator may not be effective in a real helicopter if for example the actual helicopter windows are different than those used in simulation” (p. 184). in peterson et al. (2022), a treatment group of novice surgeons used a virtual reality (vr) vitreoretinal simulator for pre-training of basic surgical skills. overall, the vr training did not cause any significant effect on the performance curve in comparison to traditional pre-training, neither in a positive nor negative sense. however, as noted by peterson et al. (2022), the control group performed better than the treatment group in one of the investigated modules. some frustration was observed among the participants in the treatment group when moving from simple to more complex tasks, which is interpreted as a possible indication of negative skills transfer. shul et al. (2019) take the debate a step further, arguing that negative skills transfer connected to lack of anatomic accuracy and realistic tissue behaviours of urological simulators used in training was the reason behind high numbers of injuries and infections resulting from catheterisation in real clinical practice. in the context of maritime pilot training, hontvedt (2015) shows how imperfections and inconsistencies in the visual lookout of the simulator environment came in conflict with the trainees’ professional vision and forced them to change their way of working in order to adapt to the shortcomings of the simulator. in line with these results, hontvedt (2015) warns that less experienced trainees might adopt inaccurate work practices due to poor photorealism. hontvedt (2015) argues that “lack of fidelity may harm the logic of the actual work task”, and that the training participants were shifting their focus “from performing within a simulated work environment to simply manipulating the simulated model” (p. 82). following this, hontvedt (2015) emphasises the importance of proper instructional guidance to maintain focus on the learning objectives of the simulation exercises. while such a warning and recommendation can be considered justified and worthy of consideration, especially in relation to training of inexperienced trainees, it should be noted that studies sellberg, nordenström, & säljö 18 | f l r explicitly concerned with the development of visual expertise in various professional domains have shown that the general idea of skills transfer ‒ whether positive or negative ‒ might be a too simplistic way of analysing professional learning in technological work settings (e.g., gegenfurtner et al., 2009; lehtinen et al., 2020; nivala et al., 2012). using the metaphor of different layers of conceptual change for professional learning in a biomedical setting, lehtinen et al. (2020) show how experts learn new ways of conceptualising the unfamiliar conditions of working with new visualisation technologies for doing diagnosis. in this case, learning goes beyond simply transferring knowledge between work settings. rather, it is a matter of adapting to new methods and working conventions, while the basic principle, the biomedical concept these experts need to understand, remained stable across work settings. in this process, lehtinen et al. (2020) found that the experts spent significant time analysing the relevant aspects of unfamiliar conditions in order to adapt their familiar working methods to the new visualisation technology. furthermore, studies of simulation-based training in various domains show that discrepancies between the simulator environment and the work setting, if properly addressed, may provide opportunities for fruitful discussions and learning rather than posing a risk of participants adopting incorrect work methods (e.g., hindmarsh et al., 2014; hontvedt & øvergård, 2020; rystedt & sjöblom, 2012; sellberg, 2017). sellberg (2017), for example, explores how the embodied activity of ship handling is trained in high-fidelity navigation simulators that, while mimicking many of the features of the bridge of a real ship, are lacking kinaesthetic and proprioceptive feedback and thereby not simulating the sense of moving in an authentic way. as shown by the study, these built-in inconsistencies of the simulator environment provided for instructional opportunities and thus facilitated rather than hindered the maritime students’ learning of ship sense. similarly, hindmarsh et al. (2014) in a study of simulation-based clinical skills training for undergraduate dental students demonstrate how differences between the simulated environment and the authentic work environment provided for teachable moments where the instructors got the opportunity to highlight and explain the appropriate performance of and the rationale behind specific occupational procedures. more specifically, this was done by the instructors noting problematic performance by the students occasioned by deficiencies or constraints in the simulator and then intervening, presenting the students with what weeks (1985, 1996) has termed ‘a contrasting pair’: a description of the incorrect performance of the procedure coupled with a presentation of the preferred way of performing it. in the present study, a similar way of contrasting non-preferred conduct with a preferred version, as hindmarsh et al. (2014) and weeks (1985; 1996) put it, has been observed in training exercises for maritime pilot trainees in a high-fidelity bridge simulator. unlike what was the case in the aforementioned studies, however, it is the training participants themselves who identify problems and present a form of contrasting pairs, thereby transforming situations where visual deficiencies/peculiarities in the simulator environment prevent them from adopting established ways of working into learnable moments. through a practice-based perspective on the development of visual expertise, and by taking the multi-layered nature of professional skills into account, this study seeks to further explore such situations where issues of simulator fidelity are addressed. thus, directing the analytical focus towards maritime pilot trainees’ communication about visual discrepancies between the simulator environment and the “real” setting, the aim is to carefully examine how visual expertise is developed in and through training in a simulated environment. the following research questions are in focus: a) how do maritime pilot trainees identify and handle visual imperfections and inconsistencies in the simulator environment? b) what are the implications of adapting to such shortcomings in the simulator environment for the development of visual expertise? with the analytical attention directed towards naturally occurring dialogues between trainees, the research design builds on a videography approach (knoblauch & schnettler, 2012). following this approach, the study draws on focused ethnography at a scandinavian simulator centre. video records from a course on advanced ship handling in maritime pilot training have been used to conduct qualitative interaction analyses (luff & heath, 2019). sellberg, nordenström, & säljö 19 | f l r before we proceed and present selected episodes of the video recordings intended to shed light on the phenomenon under study, a section that briefly outlines the differences between the dominant neuropsychological and cognitive perspectives on the development of visual expertise and the practice-based perspective adopted in the current study will be presented. 2. a practice-based perspective on the development of visual expertise the study of visual expertise has a long tradition in science, involving research fields such as cognitive neuroscience, cognitive psychology, and the learning sciences (e.g., boucheix, 2017; gegenfurtner & van merriënboer, 2017; gegenfurtner et al., 2022). whilst the unit-of-analysis in cognitive neuroscience is the neurophysiological activity of the individual (e.g., gegenfurtner et al., 2017), cognitive psychology focusses attention to the details of the individual’s eye movements and/or verbal reports of cognitive processes during visually intensive tasks (e.g., helle, 2017; jarodzka & boshuizen, 2017). at these levels, findings show that experts display stronger activation patterns in the brain for areas associated with encoding and storing visual objects and events, indicating a nonconscious and stimulus-driven indexical relation between a visual element from their work setting and the corresponding mental representation (gegenfurtner & van merriënboer, 2017). at a behavioural level, visual perception develops with experience, from slow search-to-find modes towards more efficient, holistic modes of seeing. moreover, beyond being faster in their visual search, experts exhibit a higher rate of accuracy when decisions are made based on visual elements in their work setting (gegenfurtner & van merriënboer, 2017). there are also findings which show that experts, in comparison to novices, are better at verbalising the perceptual features of visual elements of their work setting, and they are also better at conceptualising their perceptual activities (gegenfurtner & van merriënboer, 2017). from these findings, we can arrive at an understanding of the development of visual expertise as a process that starts with visual objects and events being stored in the cerebral cortex and next, the perceptual-conceptual relation develops (gegenfurtner & van merriënboer, 2017). these illustrations of the neuropsychological correlates of visual activities are interesting in their own right. however, in order to develop teaching and learning we need a research agenda of how to study the development of visual expertise, to understand how the emergence of “a good eye” is interactionally accomplished and socially organised within a professional field. in his seminal study on professional vision, goodwin (1994) shows that visualisation practices evolve through an increasing coordination with the requirements and expectations of a profession. professional vision is described as the discursive practices that “become the insignia of a profession’s craft: the theories, artifacts, and bodies of expertise that distinguish it from other professions” (goodwin, 1994, p. 606). while goodwin’s seminal work on professional vision suggests a research program for the study of concrete practices connected to visuality, recent studies have started to map out the conceptual changes and transitions of expertise involved in learning and developing visual expertise when working with new visualisation technologies (gegenfurtner et al., 2019; nivala et al., 2012; lehtinen et al., 2020). lehtinen et al. (2020) describe how theories of conceptual change initially were developed to explain the difficulties students face when attempting to understand scientific concepts, often in school settings. in short, the theory is based on the notion of different belief systems, commonly developed through everyday experiences and prior learning in the school system, and the possibility that innate predispositions come into conflict with scientific theories. consequently, there might be resistance to learning new scientific concepts or problems of coping with novel task demands. in lehtinen et al. (2020), the focus was on horizontal conceptual changes, i.e., “increased conceptual knowledge and possible conceptual change” (p. 6) when medical experts developed their diagnostics skills through learning to interpret visualisations from novel x-ray technologies used within their field. they found that the experts could not easily transfer their existing skills to the novel technologies, instead they had to learn new diagnostic methods for being able to make sense of the new visualisations. however, all experts were familiar with the basic scientific principles underlying both medical visualisations under sellberg, nordenström, & säljö 20 | f l r study. thus, the struggles of re-learning were connected with coping with method-specific knowledge and professional practices, rather than linked to the underlying scientific concepts. in all, all of the experts in lehtinen et al. (2020) were able to adapt their method-specific and professional practice to work with the new technology. moreover, the changed working methods led the experts to develop more advanced visual perceptions, including enriched scientific knowledge of anatomy, detailed technical knowledge of the colour nuances seen in x-rays, as well as new professional practices of diagnosing. this is explained in lehtinen et al (2020) as a consequence of “the multi-layer nature of professional skills” (p. 8), involving a system of various layers of conceptual, methodological, and technical knowledge. in the present study, aiming at examining how visual expertise develops in and through training in a simulated environment, the metaphor of multiple layers of professional skills, and the notion of conceptual change, serves as starting points in order to advance our understanding of how the complex professional practice of maritime piloting can be taught and learned in a virtual environment. in the next section of the article, the empirical case in terms of setting, participants, method, data, and analytical approach will be presented. 3. the empirical case maritime pilot training is a one-year specialisation program, organised by national maritime administrations and undertaken by master mariners with extensive experience of working as marine officers. the training program consists of three parts. first, there is an introduction period where the applicant goes through a recruitment process for a probationary employment as maritime pilot. after admission to the program, there is a package of basic courses, addressing a wide range of professional aspects of working as a maritime pilot, including legal and administrative course content, simulatorbased training of teamwork skills as well as advanced navigation and manoeuvring training. in the last part, the maritime pilot trainee undertakes on-the-job training and is subjected to a simulator-based test of competence. when passing this test, the maritime pilot trainee becomes a certified maritime pilot, and the probationary employment transitions into a permanent position. figure 1. maritime pilots in training approaching port in a full mission bridge simulator. the course under study takes place in the second part of training and focuses on advanced ship handling. during data collection, the course gathered six maritime pilot trainees for a week of intense training guided by two well-experienced maritime pilots serving as instructors. the trainees as well as the instructors in our study are all male, and they all have given their written, informed consent to participate in the study. since the trainees all are experienced master mariners, ages vary from trainees who are in their late twenties to those who are in their late forties. in general, the older trainees have gained more work experience, approximately 20 years, while the younger trainees have gained more experience in simulator-based training from their more recent master mariner education. the training activities during this study take place at the simulator centre, which is equipped with a full mission vts sellberg, nordenström, & säljö 21 | f l r simulator as well as four high-fidelity full mission bridge simulators. the full mission bridge simulators mimic a ship’s bridge with high accuracy in terms of navigation equipment. the simulators are provided with state-of-the-art technologies used on board ships, such as radars, automated plotting and tracking aids, electronic chart displays, gyroand magnetic compasses, rudder angle and rate-of-turn indicators, echo sounder and different steering devices. there are also screens with projections of the marine environment, representing views from the front windows, the bridge wings, as well as the rear view (figure 1). in these full-mission bridge simulators, the maritime pilot trainees train in pairs of two when manoeuvring vessels in port areas, shallow and narrow waters, and during tugboat manoeuvring, operations that are highly demanding both for the ship and its crew. 3.1. method, data, and analytical approach videography can be described as a focused ethnography, where observations of social interaction are documented by video recordings (knoblauch & schnettler, 2012). the approach is dedicated to the study of naturally occurring phenomena, that is, instead of designing or controlling the learning activities under study, the purpose is to capture the everyday social interaction that normally takes place in an instructional or any other setting. before filming started, the first and second author made two visits to the simulator centre during the autumn of 2021. during the first visit, the aim was to introduce ourselves and get a tour of the premises, as well as to discuss a suitable course for our interest in the development of visual expertise in simulated environments. later on, a second visit was made with the aim of making a detailed plan for collecting video data. in november 2021, the video recorded data were gathered by the first and second author during one week of training. handycam® cameras on tripods were placed in all three bridge operation simulators, providing mainly a view of the participants’ work on the bridge panel and the front window (see figure 1). another handycam® camera was placed in the adjacent briefing/debriefing room where the participants gather before and after each simulated scenario. in order to ensure satisfactory audio uptake from the collaborative discussions before and after each session in the simulator, a shotgun microphone was placed in the briefing/debriefing room. the collected data capture four full days of simulator-training and include all stages of training: from the pre-simulation introduction (the so-called briefing), through the simulated scenario to the post-simulation debriefing. this covers the entire training process in the course, and in sum approximately 130 hours of video recorded simulator-based training was documented. the video records of simulations form the basis for qualitative interaction analyses. in qualitative interaction analysis, the unit-of-analysis consists of verbal utterances, bodily conduct, and interaction with material and digital objects, observable through turns of talk between participants (luff & heath, 2019). in the first step of analysis, the video recordings were reviewed and catalogued in order to obtain an overview of the entire data corpus (heath et al., 2010). in the next step of the analysis, the catalogue of simulations was revisited with a focus on identifying trainees’ talk about imperfections and inconsistencies in the simulator environment. in all, 30 episodes on this theme were identified and categorised according to the type of “glitch” discussed by the trainees. categories of imperfections and inconsistencies noticed by participants include 1) frozen screens in need of restart (n=7), 2) malfunctioning instruments (n=4), 3) lack of proprioceptive/kinaesthetic feedback (n=3), 4) restricted visibility from bridge wings and aft window (n=11) and 5) lack of depth perception (n=5). it is noteworthy that categories 3-5 mostly occurred during the first two days of training and were associated with tasks where manoeuvring needed to be done with high precision, for example, going to quay. for this study, with its focus on visual expertise, categories 4 and 5 were assessed as most relevant. three episodes from category 4 were selected for further analysis. the episodes were transcribed with attention to verbal utterances and their intonation, as well as to relevant nonverbal behaviours such as gestures and gaze shifts (see table 1 for transcript notations). two of these episodes are presented in the article text (analysis). sellberg, nordenström, & säljö 22 | f l r table 1. notation system used for transcription notation meaning word underlined words or syllables are delivered with emphasis ° degree signs indicate a noticeable quieter utterance than the surrounding speech : colon(s) indicates prolongation of a syllable h the letter h indicates audible aspiration > < greater-than, less-than signs indicate a noticeable faster deliverance of an utterance than the surrounding speech (0.5) numbers in parentheses indicate pause lengths (.) a period between parentheses indicates a micro pause [ ] brackets surrounding words or syllables indicate overlapping speech – a dash sign indicates an unfinished word = equal signs indicate cut offs in speech ? a questions mark indicates a rising intonation xxx three x-letters mark inaudible speech ((action)) double parentheses and italicised letters are used to separate talk from bodily actions new line a new line marks a new turn of talk or action ­ an arrow marks the beginning of a bodily action 4. analysis in our data corpus, two maritime pilots in training work together as a bridge team, consisting of a maritime pilot and the captain of the ship. during the different exercises in the course, the trainees are taking turns working with each other on the bridge. episode 1 presented in section 4.1 concerns a scenario taking place during the second day of simulator training in the course. during the briefing, the trainees were given the task of manoeuvring a cargo vessel in the harbour of hong kong, where they are going to quay under good traffic and weather conditions. the overall aim of the exercise is that the trainees should be able to dock their vessel by using four manoeuvres. this, in turn, will require a clear and strategic plan as well as full control over the ship’s position and angle to quay, the ship’s pivot point and manoeuvring speed. figure 2 visualises the simulated vessel’s manoeuvring actions halfway through the scenario. first, the simulated vessel is going to the specified terminal. second, the trainees are positioning their ship in order to be ready for the third manoeuvre, i.e., going in reverse to quay. the fourth position, not displayed in the visualisation, is to come to quay in position to dock the vessel. hence, the trainees will have to keep a close eye on both the quay and the vessel lying at anchor. figure 2. the purple track is a visualisation of the trainees manoeuvring actions halfway through the scenario. sellberg, nordenström, & säljö 23 | f l r 4.1. collaboratively making sense of professionally relevant visual materials in episode 1, dan and matt are preparing to undertake a scenario where they are docking at the terminal at hong kong harbor. docking at this or any terminal is a manoeuvring task that requires continuous visual lookout of the outside surroundings in all directions, forward, aft, port, and starboard in order to determine the distance to the terminal and other structures, e.g., other vessels. in order to achieve this, a pilot has several navigational aids at hand. on board a real vessel, the bridge wing is an extended platform located on the sides of the ship’s bridge, typically near the outer edges. this structure serves as an extension of the main bridge area and provides additional vantage points for the ship's officers to observe the surroundings, providing an unobstructed view of the ship's sides and forward areas. in addition, this scenario involves going in reverse to quay, which makes the lookout through the aft window critical in order to determine distance to quay and the other vessel anchored at the terminal. electronic navigational aids, such as radar equipment and electronic navigational charts provide information about the vessel's surroundings, and they are particularly valuable under conditions of restricted visibility, for instance due to fog, snow or heavy rain. dan and matt are both experienced master mariners who each has approximately 20 years of working experience as captains. both can thus in a sense be regarded as experts with a "superior ability to interpret and analyze situations as well as solve problems typical of their [... area] of expertise" (lehtinen et al., 2020, p. 1). at the same time, however, they are novices in the sense that they are just entering their training as maritime pilots, which involves learning to handle new and more advanced ship operations and mastering new technologies. as part of this training, they are also gaining experience in simulation-based training in an advanced high-fidelity bridge simulator, something these two trainees have limited experience of from their master mariner training in the 1990’s. they are thus faced with the challenge of developing different layers of their expertise (lehtinen et al., 2020) in that they must simultaneously learn new profession-specific skills and to manage the functions of a more technically advanced simulator. in the scenario, dan takes on the role of pilot responsible for navigating the ship and matt takes on the role of captain. while the first part of the episode (1a) is reproduced to facilitate understanding of the context, our analytical focus is put on the second part (1b). episode 1a 01 dan eh:: bara tänker här nu lite ((suckar)) eh:: just thinking a bit here ((sighs)) ­ ((rubs his tempels)) 02 (3.5) 03 jag kommer liksom puttra framåt här i i’ll somehow chug forward here in ­ ((moves hand up and down in a horizontal position)) 04 matt >m< >m< 05 dan och sen vill man ju få liksom en kick uppåt ((visslar)) and then one wanna get like a kick upwards((whistles)) ­ ((pulls his hand towards his body with a jerking movement)) 06 och liksom sikta in n’ sort of aim in ­ ((moves both hands in to show the turn)) 07 så att man får den här vinkeln so one gets this angle 08 å sen börja backa n’ then start to reverse ­ sellberg, nordenström, & säljö 24 | f l r ((moves his hands towards his body)) 09 vid nåt läge at some position ­ ((turns his whole body towards bridge aft)) 10 matt °ja° °yes° in (1a), we can see how dan, acting as the pilot, is starting to lay out a plan for how to approach the manoeuvring task. the episode starts with dan saying that he is “just thinking a bit here” while rubbing his templates and sighing (line 01), signalling that he finds the situation challenging. the overall objective of the training, to present scenarios that pose challenges for the trainees, thus seems to have been achieved. furthermore, dan’s utterance and non-verbal actions can be seen as responsive to the instructors’ request for the trainees to think-out-loud while conducting the scenarios, i.e., to say what they do and why, an important element when fostering collaborative learning on the bridge (hontvedt & arnseth, 2013). dan continues to think-out-loud for the remainder of (1a), describing what manoeuvres he plans to perform and what he wants to achieve with these manoeuvres (lines 03, 05-09), which is met with minimal and affirmative responses from matt (lines 04, 10). in (1b), however, there is a shift in the activity when dan switches from reporting on planned manoeuvres to orienting to a potentially problematic situation caused by limitations in the visual lookout of the simulator, inviting matt to collaboratively make sense of the available visual materials. episode 1b 11 dan hur vet jag när jag har rätt vinkel då? how do i know when i’ve got the right angle then? ­ ((put his hand on his forehead)) 12 kan man could one ­ ((points towards the bow)) 13 då skulle man vilja se rakt bak ut egentligen then one would like to look straight aft actually ­ ((turns around and points towards the aft window)) 14 om man börjar backa upp på den så if one starts to back up on that one then ­ ((looks at screen showing port side bridge wing)) 15 matt ja men sägsäg bara åt mig hurvad du vill se så löser jag yeah but telljust tell me howwhat you want to see n’ i’ll fix 16 den biten med kameran that part with the camera ­ ((points towards the bridge wing screen)) 17 dan jag tänker ((host)) liksom i verkligheten här hade man ju velat titta i think ((cough)) like in reality here one would have wanted to look ­ ((turns around and looks aft)) 18 bakåt ju aft then 19 liksom attja men där är man (xxx) like towell but there one is (xxx) ­ ((points aft)) 20 ja då börjar jag backa nu yeah then i’ll start going back now 21 å så vill man hålla den hävningen sellberg, nordenström, & säljö 25 | f l r n’ then one wanna keep that heave 22 matt °ja° °yes° 23 dan men det är klart det går att göra samma sak med ecdis:en där nu but of course the same thing could be done with the ecdis there now ­ ((points at the ecdis)) 24 om man ska lita på den liksom if one should trust it somehow 25 matt ja (.) och även den där yes (.) and even that one ­ ((leans over the bridge panel and adjusts the conning display)) 26 här ser du juhär ser man ju precis allting here you can seehere one can see absolutely everything 27 dan ja precis där ser yeah exactly there we see ­ ((leans forward and looks at the conning display)) 28 ja då kanske vi ska ha yeah maybe we should have it 29 lite ut-zoomat så a little zoomed out like that 30 matt ja yes 31 dan ja så man kan se (xxx) yeah so one can see (xxx) ­ ((stands upright again)) 32 matt °ja° °yes° in (1b), we can see how dan initially addresses a problem related to a deficiency in the visual functionality of the simulator: in the simulator it is not possible to get a visual overview of the quay by looking out through the aft window which would be possible on a real ship. he states that to ensure the correct angle of the ship in the accomplishment of the manoeuvring task (line 11) “one would like to look straight aft", while pointing towards the aft window with his left arm (line 13). that this is the established way of the profession to get a visual overview of the quay rather than a personal preference of dan is suggested by his use of the generic third person pronoun “one” (sw. “man”) rather than the first-person pronoun 'i' (sw. “jag”) which is used initially (line 11). matt’s response (lines 15-16) indicates that he understands and is prepared to accept and adapt to the shortcomings of the simulator identified by dan. rather than commenting further on how the task should be performed in a "real" work setting, matt presents an alternative way of performing it which is adapted to the functionality of the simulator environment: if dan specifies what he needs to see, matt will provide visual access to that with the help of the camera. the camera matt refers to is adjustable and provides the representations that can be seen on one of the screens in front of the trainees (see figure 1 and 3) showing the view from the port side bridge wing. compared to a real ship where you get a visual overview of the surroundings by standing on the bridge wing, the visual representation from the bridge wing shown on this screen in the simulator is quite limited. furthermore, as dan initially continues to maintain, this would not be the preferred approach “in reality” to get a visual lookout and determine the current position of the ship. he argues that “one would have wanted to look aft” (lines 17-19), while turning around and looking aft. what dan has presented so far could be seen as one part of “a contrasting pair” (weeks, 1985, 1996), the conversational device mentioned in section 1: a description of the preferred way of performing a particular visually demanding manoeuvring task (lines 13, 17-19). what he produces next, after having announced that he will begin reversing the ship (lines 20-21), could be seen as the other part of the contrasting pair: a description of an alternative non-preferred way of performing the same sellberg, nordenström, & säljö 26 | f l r task (lines 23-24). as stated by dan, it would also be possible to use the electronic chart, displaying navigational information such as ship positions in real-time, to obtain the necessary visual information. as we have explained earlier, the option of using this instrument, instead of relying on the visual lookout, could become relevant also in a “real” situation when visibility is restricted by, for example, fog or heavy rain. using the ecdis would thus not be an incorrect approach, but as is clear from dan’s subsequent comment on line 24, “if one should trust it somehow”, trust in this instrument is not to be taken for granted in the current situation. without going into depth on the topic of trust in automation, it is of significance to point to the safety culture in maritime navigation that is prescribing visual lookout in combination with a range of digital instruments to triangulate information (see, e.g., lützhöft & dekker, 2002). at this point, matt presents yet another alternative for gathering visual information on the bridge, one that could help to achieve triangulation: they could use the so-called conning display, an instrument providing an integrated overview of the situation during manoeuvres, including course, speed, depth and rate of turn, with which matt claims that “one can see absolutely everything” (line 26). dan agrees (line 27), and they proceed to reason about how to adjust the instrument to get the visual information needed, e.g., distances to the quay, as well as to other ships, to complete the task (lines 2832). consensus thereby seems to have been reached that ecdis and the conning display will be used to deal with the problem of poor visual lookout through the aft window initially presented by dan. figure 3. dan and matt exploring different methods for gathering visual information in the simulator. in this episode, we can see how the restrictions of the simulator open up for the trainees to collaboratively explore different navigational methods for completing a visually demanding manoeuvring task. on board a seagoing vessel and with the weather conditions simulated in the current scenario, the participants would likely only have used the standard navigational method initially presented by dan. however, in the simulator they are compelled to collaboratively identify and attempt alternative strategies to complete the task. in this case, the limitations and affordances of the simulated environment thus occasion explorations that might lead to an extension of the participants' visual sellberg, nordenström, & säljö 27 | f l r expertise (lehtinen et al., 2020). however, as seen in our next episode, trainees might also hesitate to adapt their working methods to the visual limitations of the simulator environment. 4.2. drawing on differences in prior experiences for solving the task in this episode, bill and tim are training together in one of the other full-mission bridge simulators, performing the same scenario as matt and dan in the previous example. bill is taking the role as pilot and tim as captain of the vessel. while they both are novices in their roles as maritime pilots, they have different experiences from working as master mariners as well as from training in a virtual environment. tim, the younger of the two, has quite recently graduated from a master mariner program and has extensive experience in training in simulated environments. bill has many years of experience of working at sea but is to be considered a novice when it comes to working in a simulation environment. thus, to use the words of lehtinen et al. (2020), while bill has “acquired a high level of expertise in one specific field [he] must extend the scope of [his] expertise into new fields, such as new technologies that offer opportunities to enrich [his] repertoires of tools and alternative methods of dealing with the work objects” (p. 6) in the current scenario. simultaneously, tim’s experience of both the real-world setting and the simulator environment becomes a valuable resource for solving the task at hand. episode 2 01 bill nu ska vi se (.) siktar på den mittersta kranen där borta just nu då now let’s see (.) aiming at the middle crane over there right now then 02 tim >okej< >okay< 03 bill så:: får vi se hur de so:: we’ll see how it ((adjusting eyeglasses)) 04 de e ju de (.) när man gör den manövern och kommer in så it's like this (.) when one does this manoeuver n’ come in like this ­ ((moves left arm in an angle towards starboard)) 05 vill man ju gärna hänga på bryggvingen så att man ser one really wants to hang on the bridge wing so that one sees ­ ((points towards starboard)) 06 allt va everything right 07 tim ((points towards a screen on portside)) ja men då har du ju yeah but then you have 08 bill ja precis men den känns inte så bra haha((hostar)) yeah precisely but it doesn’t feel that good haha((coughs)) 09 tim ja (.) men vi kan peka den framåt om du vill? yeah (.) but we can point it frontwards if you want? ­ ((leans forward towards the screen)) 10 bill jaja jo menfast jag menarde e lättareom [man yeahyeah well butthough i meanit's easier if [one 11 tim [>ja< [>yeah< 12 bill kan se ordentligt så att säga can see properly so to say 13 nu ska vi se (.) nu börjar vi närma oss henne där now let’s see (.) now we’re approaching her there ­ ((leans forward looking at the radar)) 14 så jag lägger stopp i maskin då so i’ll put a stop to the machine then sellberg, nordenström, & säljö 28 | f l r ­ ((moves the lever)) 15 tim nu är vi ungefär mitt på när du gör det här now we’re about halfway through when you do this 16 bill ja (.) jag stoppar (.) får vi se vad som händer yes (.) i’ll stop (.) we’ll see what happens 17 nu får du en liten vägledning för nästa körning också now you get a little guidance for the next run too the episode begins with bill presenting a plan for how to perform the manoeuvre to dock at the terminal: he will use one of the cranes stationed in the port (see figure 4) as a visual reference (line 01). after a prompt “okay” from tim (line 02), bill then begins to produce what can be heard as a caveat to the successful execution of this plan: “so:: we’ll see how it-” (03). however, the utterance is aborted, and he instead proceeds to explain the preferred way of performing the manoeuvre to maintain a proper visual lookout (lines 04-06): “it's like this (.) when one does this manoeuvre n’ come in like this one really wants to hang on the bridge wing so that one sees everything right”. while producing the utterance, he points towards the location where the starboard bridge wing, i.e., the extended platform located on the side of the ship’s bridge, would be on a real ship, thereby showing tim where he would have liked to position himself in a situation at sea. note here the similarities to (1b) which also begins with the pilot, using the generic pronoun “one”, presenting the approach for performing the manoeuvre that would have been preferred in a “real” situation to get a good visual overview. however, unlike in (1b) the trainees in the current episode do not as easily reach consensus on how to proceed. similar to dan and matt, bill and tim are engaging in collaboratively exploring different methods for completing the task and they reason about the limits and affordances of the simulated environment in comparison to working on a ship’s bridge. their willingness to adapt their way of working to the simulator's functions differs, however. as demonstrated by the remainder of (2), bill, the more experienced master mariner, seems hesitant to adapt his working methods to the simulated environment. tim, on the other hand, seems more willing to explore its possibilities. repeatedly directing bill’s attention towards the available visualisation technologies, he argues for a way of solving the task at hand adapted to the simulator environment. in response to bill's initially stated preference for how to get a visual lookout (lines 04-06), tim points towards a screen on the port side where the view from the bridge wings can be represented (figure 3), saying “yeah but then you have-” (line 07), thereby directing bill’s attention to one of the available resources in the simulator that could compensate for its visual limitations. bill however cuts off the utterance, first delivering a token of agreement “yeah precisely” but then proceeding to claim that “but it doesn’t feel that good” followed by a short laugh (line 08), thus rejecting tim’s suggestion of the alternative method of gathering visual information. tim responds by presenting an additional suggestion, explaining how the intended visualisation could be adjusted and thus used to provide the visual information they would need for solving the task at hand (line 09): “but we can point it frontwards if you want?”. bill however maintains his sceptic stance towards making use of the limited visual lookout through the digital visualisation, insisting that it is easier to perform the manoeuvre “if one can see properly so to say” (lines 10, 12). up to this point, we have seen how the trainees negotiate which strategies to use to obtain the necessary visual information to solve a task without reaching consensus. tim, who has more experience of working in a simulated environment, took a leading role in exploring the technical functionalities of the simulator, while bill, relying on his many years of experience as a master mariner, sought to maintain procedures which would have been preferable in an on board setting. similar to the medical professionals observed by lehtinen et al (2020), who trained to diagnose patient cases using for them new and unfamiliar imaging technologies, bill has so far shown a preference for applying methods which are familiar to him in the new situation, even though these methods might not be the most efficient for this situation. however, given the collaborative reasoning that takes place, we can still assume that the in-situ attention to the shortcomings of the simulator contributed to promoting the visual expertise sellberg, nordenström, & säljö 29 | f l r of both participants. by making the shortcomings a shared topic of discussion, both participants gain access to the discrepancies between a simulated and an on-board perspective on a situation. figure 4. tim pointing towards the screen on portside where the view from the bridge wing is represented. in line 14, bill says “now let’s see” and leans over the bridge panel to look at the radar, stating that “we’re approaching this one there”. moving his hand to the lever and putting it to a stop, he explains “i’ll put a stop to the machine then”. here, bill shows that he will stop the vessel, and hence the scenario before completion, since it’s time for a scheduled break. tim responds with an assessment of the situation “now we’re about halfway through when you do this” (line 15), which is ratified by bill in the next turn with a “yes” before he continues to say, “i’ll stop (.) we’ll see what happens”. here bill repeats that he will stop, and that the outcome of the actions taken still is uncertain, showing that he is quite hesitant of what the correct action would be at this point. however, in line 16, bill acknowledges that tim will “get a little guidance for the next run”, referring to the next session after the break, where tim will act as pilot in the same scenario. hence, their reasoning about the limits and affordances of the simulated environment during this session serves as a starting point for further exploration in the next run. 4.4. analytical findings the findings presented above are in line with those in hontvedt’s study (2015), where the pilots repeatedly criticised the fact that navigation tasks in the simulator needed to be carried out with electronic equipment instead of through a visual lookout. however, while hontvedt sees this as a potential risk for negative skills transfer, our results suggest that the shift in navigational methods in the simulator seems to present opportunities for professional learning. an important element of this argumentation is that the participants themselves in their activities notice and attend to the differences between navigating in a simulator and on board a ship, respectively. the differences thus trigger reflection and problem-solving. in addition, we can see how the trainees continuously connect their manoeuvring in the simulator to their professional experiences of ship handling on board real vessels (see wiig et al. 2018). while previous studies highlight the need for instructors to facilitate discussions with trainees to avoid pitfalls in training due to a lack of fidelity (e.g., sellberg 2017; hindmarsh et al. 2014), the maritime pilots in training in our materials spontaneously connect the simulated practices to their professional practice without the support of an instructor. extensive experiences of ship handling sellberg, nordenström, & säljö 30 | f l r serve as resources for identifying and commenting on the limitations of the simulated environment. in other words, the participants are not constrained by the simulation as a fixed environment, rather they entertain and test hypotheses about differences and similarities between the two settings. put differently, the participants do not learn by passively subordinating their decisions on how to navigate to the design of the simulated environment, rather they mobilise their professional experiences as resources for sensemaking and for reflecting and commenting on what characterises navigation in the two situations. furthermore, when taking the multilayered nature of visual expertise into account in our analysis, it is important to consider the different dimensions of this task. rather than viewing the inconsistencies in the simulator as in conflict with the trainees’ professional vision, we can see that at the conceptual level, the calculations for determining distance, speed and turn ratio in ship handling are the same in the simulated model as in a real vessel (see lehtinen et al., 2020). however, the methods used for gathering information, i.e., using instruments rather than relying on visual lookout, are different in the simulator, where the pilot in training needs to make use of several digital navigation aids. as a result, the task involves attending to different representations for understanding manoeuvring and movement during ship handling. for instance, rather than sensing the movements of the ship and receiving proprioceptive feedback, the trainees need to interpret how the ship moves through rather abstract representations such as numerical values and graphs available through the instruments (see sellberg, 2017). in previous research, such differences in resources for interpretation between tasks have led experts to develop more advanced visual perceptions, enriched scientific understandings, detailed technical knowledge and new professional practices (lehtinen et al., 2020). 5. conclusion and discussion in this study, examining how visual expertise develops in and through simulator-based training, the metaphor of multiple layers of professional skills and the notion of conceptual change serves as starting points to advance our understanding of how the complex professional practice of maritime piloting can be taught and learned through experiences generated in virtual environments. our detailed analysis of the talk of trainees during training, and our close examination of how they handle imperfections and inconsistencies in-situ, show how ship handling in a simulator environment is a different activity than ship handling on board a seagoing vessel. in the simulator, maritime pilots in training make use of a variety of navigational instruments to compensate for, and to adapt to, the shortcomings of the visual lookout in the simulator. these findings are in line with previous studies that warn for negative skills transfer due to the lack of photorealism in simulated environments (hontvedt, 2015). however, our findings show how the trainees articulate and conceptualise the differences between simulations and work on board a seagoing vessel in ways that support the development of visual expertise (see lehtinen et al., 2020). in other words, in our materials, the discussions of the trainees about the imperfections of the simulation show that they have learned about such differences, and that their evaluations of information and judgements about how to act are grounded in conceptual control over what differs between the scenarios in the simulator, and what happens on a bridge at sea. on some occasions, they also articulate these differences on the basis of their maritime experiences at sea. instead of warning for negative skills transfer, we argue that the challenges of training in a virtual environment might lead to enriched conceptual, methodological, and technical knowledge and considerations in the context of visually demanding and complex tasks and broaden the insights of participants of how representations relate to the world. however, for these positive training outcomes to emerge, we want to stress that inconsistencies between the simulator and the work on board a ship need to be reflected on during and/or after training. one way to ensure this is to systematically facilitate reflection on these matters in the post-simulation debriefing that follows the simulated scenario. it is also important to point out that these results might not be directly applicable to novices training in a simulated environment, as they have limited experience of the working context and thus might have difficulties in noticing inconsistencies between the simulator and the working environment sellberg, nordenström, & säljö 31 | f l r in the first place. hence, for novices the simulator instructors’ dedicated work to monitor their activities in the simulator and explain the inconsistencies when they occur is essential to avoid pitfalls in training (sellberg, 2017). however, if discussed and reflected on, inconsistencies in the simulator environment may provide powerful opportunities for professional learning also for novices (hindmarsh et al., 2014; hontvedt & øvergård, 2020; rystedt & sjöblom, 2012). additionally, while we often see novices and experts as representing opposite ends of a spectrum of knowledge and skills, our study shows that the distinction between novices and experts is multifaceted, contingent and non-linear. today, training in simulators is an integral part of educational programs that prepare trainees for professions with high standards of safety, in settings such as healthcare, aviation, and maritime navigation. in this study, we have taken seriously the concerns raised with respect to risks of inducing negative skills transfer when making use of simulators in training. as a general message, it is important to make all participants aware of the fact that simulators can never be realistic in all senses of this term, but neither are all ships and their equipment identical. simulators have other affordances than those that apply to real life situations, and this is their strength, providing a context for development of expertise through deliberate practice. our study contributes by providing a detailed analysis of simulator-based training as it is practically accomplished in maritime pilot education, thereby advancing our understanding of simulation as a tool for professional learning. as a result, our study shows how and why simulation training and training on board ships mutually support the advancement of the trainees’ visual expertise in their learning trajectory towards mastery of maritime skills. finally, we argue that the trainees’ ability to handle inconsistencies and imperfections in the simulator is closely related to their prior experience of both training contexts. hence, learning to simulate is essential in professional education that aims to prepare trainees for work in safety critical domains. keypoints training in simulators is an integral part of educational programs that prepare trainees for professions with high standards of safety. to articulate and conceptualise the differences between simulations and work on board a seagoing vessel are key to the development of visual expertise essential in navigation. the challenges encountered during training in a virtual environment may lead to enriched conceptual, methodological, and technical knowledge in the development of visual expertise. simulation training and training on board ships are activities that mutually support each other by advancing the trainees’ visual expertise through exposure to slightly different, but professionally relevant, situations. acknowledgments this study is part of the project “evaluation of eye-tracking as support in simulator training for maritime pilots” financed by the swedish transport administration between 2020-2023. the authors would like to express their warmest gratitude to the maritime pilots in training and their instructors who participated in the study. we are also grateful towards members of the sociocultural and dialogical studies (sds) seminar at university of gothenburg for insightful discussions on an early draft of the manuscript and the audience at aera for valuable comments on the submitted study at the annual meeting in chicago april 2023. sellberg, nordenström, & säljö 32 | f l r references bassetti, c. (2021). the tacit dimension of expertise: professional vision at work in airport security. discourse studies, 23(5), 597-615. https://doi.org/10.1177/14614456211020141 boucheix, j.-m. (2017). the interplay between methodologies, tasks and visualisation formats in the study of visual expertise. frontline learning research, 5(3), 155–166. https://doi.org/10.14786/flr.v5i3.311 comi, a., jaradat, s., & whyte, j. (2019). constructing shared professional vision in design work: the role of visual objects and their material mediation. design studies, 64, 90-123. https://doi.org/10.1016/j.destud.2019.06.003 garfinkel, h. (2002). ethnomethodology's program: working out durkheim's aphorism. rowman & littlefield publishers. gegenfurtner, a., gruber, h., holzberger, d., keskin, ö., lehtinen, e., seidel, t., stürmer, k., & säljö, r. (2022). towards a cognitive theory of visual expertise: methods of inquiry. in c. damsa, a. rajala, g. ritella & j. brouwer (eds.), re-theorizing learning and research methods in learning research. routledge. gegenfurtner, a., kok, e., van geel, k., de bruin, a., jarodzka, h., szulewski, a., & van merriënboer, j. j. (2017). the challenges of studying visual expertise in medical image diagnosis. medical education, 51(1), 97-104. https://doi.org/10.1111/medu.13205 gegenfurtner, a., nivala, m., säljö, r., & lehtinen, e. (2009). capturing individual and institutional change: exploring horizontal versus vertical transitions in technology-rich environments. in u. cress, v. dimitrova & m. specht (eds.), learning in the synergy of multiple disciplines (pp. 676681). springer. gegenfurtner, a., & van merriënboer, j. j. g. (2017). methodologies for studying visual expertise. frontline learning research, 5(3), 1–13. https://doi.org/10.14786/flr.v5i3.316 gegenfurtner, a., lehtinen, e., helle, l., nivala, m., svedström, e., & säljö, r. (2019). learning to see like an expert: on the practices of professional vision and visual expertise. international journal of educational research, 98, 280-291. https://doi.org/10.1016/j.ijer.2019.09.003 goodwin, c. (1994). professional vision. american anthropologist, 96(3), 606–633. doi:10.1525/aa.1994.96.3.02a00100 heath c., hindmarsh, j. & luff, p. (2010). video in qualitative research: analysing social interaction in everyday life. sage publications ltd. helle, l. (2017). prospects and pitfalls in combining eye-tracking data and verbal reports. frontline learning research, 5(3), 1-12. https://doi.org/10.14786/flr.v5i3.254 hindmarsh, j., hyland, l., & banerjee, a. (2014). work to make simulation work: ‘realism’, instructional correction and the body in training. discourse studies, 16(2), 247-269. https://doi.org/10.1177/1461445613514670 hontvedt, m., & arnseth, h. c. (2013). on the bridge to learn: analysing the social organization of nautical instruction in a ship simulator. international journal of computer-supported collaborative learning, 8, 89-112. https://doi.org/10.1007/s11412-013-9166-3 hontvedt, m. (2015). professional vision in simulated environments—examining professional maritime pilots' performance of work tasks in a full-mission ship simulator. learning, culture and social interaction, 7, 71-84 https://doi.org/10.1016/j.lcsi.2015.07.003 hontvedt, m., & øvergård, k. i. (2020). simulations at work—a framework for configuring simulation fidelity with training objectives. computer supported cooperative work, 29, 85-113. (cscw), https://doi.org/10.1007/s10606-019-09367-8 ivarsson, j. (2017). visual expertise as embodied practice. frontline learning research, 5(3), 123– 138. https://doi.org/10.14786/flr.v5i3.253 jarodzka, h., & boshuizen, h. p. (2017). unboxing the black box of visual expertise in medicine. frontline learning research, 5(3), 167–183. https://doi.org/10.14786/flr.v5i3.332 https://doi.org/10.1177/14614456211020141 https://doi.org/10.14786/flr.v5i3.311 https://doi.org/10.1016/j.destud.2019.06.003 https://doi.org/10.1111/medu.13205 https://doi.org/10.14786/flr.v5i3.316 https://doi.org/10.1016/j.ijer.2019.09.003 https://doi.org/10.14786/flr.v5i3.254 https://doi.org/10.1177/1461445613514670 https://doi.org/10.1007/s11412-013-9166-3 https://doi.org/10.1016/j.lcsi.2015.07.003 https://doi.org/10.1007/s10606-019-09367-8 https://doi.org/10.14786/flr.v5i3.253 https://doi.org/10.14786/flr.v5i3.332 sellberg, nordenström, & säljö 33 | f l r kim, t. e., sharma, a., bustgaard, m., gyldensten, w. c., nymoen, o. k., tusher, h. m., & nazir, s. (2021). the continuum of simulator-based maritime training and education. wmu journal of maritime affairs, 20(2), 135-150. https://doi.org/10.1007/s13437-021-00242-2 knoblauch, h., & schnettler, b. (2012). videography: analysing video data as a ‘focused’ ethnographic and hermeneutical exercise, 12(3), 334-356. qualitative research. https://doi.org/10.1177%2f1468794111436147 lehtinen, e., gegenfurtner, a., helle, l., & säljö, r. (2020). conceptual change in the development of visual expertise. international journal of educational research, 100, 101545. https://doi.org/10.1016/j.ijer.2020.101545 luff, p. k., & heath, c. (2019). visible objects of concern: issues and challenges for workplace ethnographies in complex environments. organization, 26(4), 578-597. https://doi.org/10.1177/1350508419828578 lützhöft, m. h., & nyce, j. m. (2006). piloting by heart and by chart. the journal of navigation, 59(2), 221-237. https://doi.org/10.1017/s0373463306003663 lützhöft, m. h., & dekker, s. w. (2002). on your watch: automation on the bridge. the journal of navigation, 55(1), 83-96. doi:10.1017/s0373463301001588 lymer, g. (2009). demonstrating professional vision: the work of critique in architectural education. mind, culture, and activity, 16(2), 145-171. https://doi.org/10.1080/10749030802590580 nivala, m., rystedt, h., säljö, r., kronqvist, p., & lehtinen, e. (2012). interactive visual tools as triggers of collaborative reasoning in entry-level pathology. international journal of computersupported collaborative learning, 7(4), 499-518. https://doi.org/10.1007/s11412-012-9153-0 petersen, s.b., vestergaard, a.h., thomsen, a.s.s., konge, l., cour, m.l., grauslund, j. & vergmann, a.s. (2022). pretraining of basic skills on a virtual reality vitreoretinal simulator: a waste of time. acta ophthalmol. 100(5). https://doi.org/10.1111/aos.15039 popova, k. (2018). ethnomethodological studies of visuality. ethnographic studies, 15, 23-37. rystedt, h., & sjöblom, b. (2012). realism, authenticity, and learning in healthcare simulations: rules of relevance and irrelevance as interactive achievements. instructional science, 40, 785-798. https://doi.org/10.1007/s11251-012-9213-x schul, a., gong, a., & sweet, r. (2019). mp35-13 development and validation of a high-fidelity urethral catheter simulator. the journal ofurology. https://doi.org/10.1097/01.ju.0000556003.94745.db sellberg, c. (2017). representing and enacting movement: the body as an instructional resource in a simulator-based environment. education and information technologies, 22, 2311–2332. https://doi.org/10.1007/s10639-016-9546-1 sellberg, c., & lundin, m. (2017). demonstrating professional intersubjectivity: the instructor's work in simulator-based learning environments. learning, culture and social interaction, 13, 60-74. https://doi.org/10.1016/j.lcsi.2017.02.003 taber, m. j. (2013). crash attenuating seats: effects on helicopter underwater escape performance. safety science, 57, 179-186. https://doi.org/10.1016/j.ssci.2013.02.007 weeks, p. a. (1985). error-correction techniques and sequences in instructional settings: toward a comparative framework. human studies, 8, 195-233. https://www.jstor.org/stable/20008946 weeks, p. (1996). a rehearsal of a beethoven passage: an analysis of correction talk. research on language and social interaction, 29(3), 247-290. doi: 10.1207/s15327973rlsi29033 wiig, c., silseth, k., & erstad, o. (2018). creating intercontextuality in students learning trajectories. opportunities and difficulties. language and education, 32(1), 43–59. https://doi.org/10.1080/09500782.2017.1367799 https://doi.org/10.1007/s13437-021-00242-2 https://doi.org/10.1177%2f1468794111436147 https://doi.org/10.1016/j.ijer.2020.101545 https://doi.org/10.1177/1350508419828578 https://doi.org/10.1017/s0373463306003663 https://doi.org/10.1080/10749030802590580 https://doi.org/10.1007/s11412-012-9153-0 https://doi.org/10.1111/aos.15039 https://doi.org/10.1007/s11251-012-9213-x https://doi.org/10.1097/01.ju.0000556003.94745.db https://doi.org/10.1007/s10639-016-9546-1 https://doi.org/10.1016/j.lcsi.2017.02.003 https://doi.org/10.1016/j.ssci.2013.02.007 https://www.jstor.org/stable/20008946 https://doi.org/10.1080/09500782.2017.1367799 microsoft word jiangfinalproofs.docx frontline learning research vol. 9 no. 3 (2021) 69 95 issn 2295-3159 info corresponding author: juming jiang, graduate school of psychology of doshisha university, japan email address: jiangjuming@live.cn. doi: https://doi.org/10.14786/flr.v9i3.751 moderating effects of individual differences in causality orientation on relationships between reward, choice, perceived competence, and intrinsic motivation juming jiang1, misaki kusamoto2, & ayumi tanaka2 1 graduate school of psychology of doshisha university, japana 2 faculty of psychology of doshisha university, japanb article received 20 november 2020 / article revised 24 june 2021/ accepted 25 june / available online 12 august abstract this study examined whether individual differences in causality orientation moderate the effects of monetary reward and choice on perceived competence, which affect intrinsic motivation in turn. causality orientation refers to an individual’s tendency to experience or interpret events in a social context in a specific way and to behave accordingly, and is directly related to the development of intrinsic motivation. we randomly assigned 103 undergraduate students to one of four conditions: reward (reward vs. no reward) × choice (choice vs. no choice). participants were presented with puzzles to solve in the experimenters’ presence, and they were free to continue working on it once the experimenters left the room. we measured the time spent on solving puzzles when participants were free to choose other activities, and used self-reported feedback on task enjoyment as an index for intrinsic motivation. we also measured perceived competence as a mediator. task enjoyment was unaffected by reward in participants with high autonomy orientation, but dropped significantly in participants with low autonomy orientation. choice over task increased perceived competence in participants with high autonomy orientation, but lowered perceived competence in the case of low autonomy orientation. we found no significant effects for time spent on puzzles. the present study contributes to current understandings of the causes of performance differences in various settings. keywords: monetary reward; choice; intrinsic motivation; causality orientation. jiang, kusamoto, & tanaka 70 | f l r 1. introduction enhancing and sustaining individuals’ motivation have been a long-standing issue in applied settings such as education and the workplace. intrinsic motivation, where an activity is undertaken for its inherent satisfaction, rather than some separable consequence, is considered the most important and desirable type of motivation (deci & ryan, 2000). not only does intrinsic motivation strongly relate to individuals’ engagement and performance, but it is also significant for improving physical and mental health (huang et al., 2016; li et al., 2015; putra et al., 2017). intrinsic motivation can be significantly influenced by external environmental factors, such as the provision of monetary reward and choice (cerasoli et al., 2014; patall et al., 2008). the effect of the external environment on intrinsic motivation differs according to individual differences, including culture (e.g., iyengar & lepper, 2000; wang & guthrie, 2004), age (e.g., catania & randall, 2013; lepper et al., 2005), gender (e.g., omansky et al., 2016, skaalvik & skaalvik, 2004), self-concept (e.g., bong & clark, 1999; khalaila, 2015), and need for achievement (covington & müeller, 2001) or achievement goals (e.g., elliot & hulleman, 2017; rawsthorne & elliot, 1999). thus, the joint function of environmental factors and internal psychological processes should be considered to facilitate intrinsic motivation. researchers have paid far less attention to an important individual factor directly related to the development of intrinsic motivation: causality orientation. causality orientation is an individual tendency to perceive and organize motivationally relevant information in a specific way (deci & ryan, 1985, 2000). the present research aimed to elucidate how the effect of monetary reward and choice on intrinsic motivation is moderated by individual causality orientation. 1.1 basic psychological needs and causality orientations. self-determination theory (sdt; deci & ryan, 2000; ryan & deci, 2017) states that the needs for autonomy, competence, and relatedness are universal, innate, and essential for intrinsic motivation and well-being. need for autonomy is the psychological need to be the origin of one’s own behaviour. need for competence is the psychological need to interact effectively with one’s environment and have opportunities to express one’s abilities. need for relatedness is the psychological need associated with experiencing a sense of belonging and connectedness to others within a social context. when people are more successful at satisfying these three basic psychological needs, they exhibit more intrinsic motivation and internalization of motivation, and integrate cultural values and regulations, resulting in greater behavioural effectiveness and psychological well-being. however, when needs frustration occurs, there is diminished autonomous motivation, along with fragmentation, defensiveness, and rigidity, which results in ill-being (ryan et al., 2016). even though needs satisfaction and frustration have proximal effects across individuals, persistently experiencing differences in contextual supports and deprivation of basic psychological needs can lead to significant individual differences in how people orient to their environments over time, especially with regard to motivation. (ryan & deci, 2017). according to sdt, people can learn to focus more on certain affordances, rewards, or pressures, and develop characteristic approaches to regulating their emotions and behaviours. this concept of individual difference in sdt is described as causality orientation (deci & ryan, 1985, 2000). causality orientations reflect people's propensities to orient to different motivationally relevant aspects of situations, especially with respect to whether the individuals will exercise autonomy, attend to controls, or fear non-contingent reactions to their initiations and behaviours (mcadams & pals, 2006). deci and ryan (1985, 2000; see also ryan & deci, 2017) sort causality orientations into three broad classes: autonomous, controlled, and impersonal. when autonomy orientation dominates, individuals seek opportunities for self-determination and choice; they tend to interpret their existing situations as more autonomy-promoting and organise their actions based on personal goals and interests, rather than controls and constraints. when the controlled orientation dominates, people experience their social context in terms of rewards and social pressures relating to values or interests. people high on jiang, kusamoto, & tanaka 71 | f l r controlled orientation tend to use external and introjected styles of regulation, and have a low level of intrinsic motivation. when the impersonal orientation dominates, people tend to see their environment as uncontrollable. the impersonal orientation develops as people experience a considerable degree of unpredictable thwarting of their basic psychological needs, leaving them feeling non-autonomous, ineffective, and anxious. people high on impersonal orientation often foster amotivation or a sense of lack of self-control, which leaves them unable to master or take command of themselves or situations. 1.2 effect of monetary reward on intrinsic motivation one powerful external environmental factor thought to have tremendous impact on intrinsic motivation is monetary reward. the ongoing debate over how external monetary rewards influence intrinsic motivation has continued since the 1970s. on the one hand, some researchers claim that rewards can be used to increase intrinsic motivation and performance (e.g., byron & khazanchi, 2012; hendijani et al., 2016; jovanovic & matejevic, 2014). specifically, expected tangible rewards have been found to increase the time of free-choice behaviour and performance on tasks in which individuals have low initial interest; on high-interest tasks, self-reported task enjoyment is positively influenced when rewards signify perceived competence (cameron, 2001; cameron, banko, & pierce, 2001). on the other hand, some researchers believe that the positive influence of extrinsic rewards, especially monetary rewards, on intrinsic motivation is not only temporary but also inherently harmful (e.g. deci, 1971; ma et al., 2014; murayama et al., 2010; warneken & tomasello, 2008). they argue that external rewards are likely to generate an external perceived locus of causality for the task, leading individuals to perceive their actions as being controlled by external contingency rather than personal agency, thus lowering intrinsic motivation. recent neural evidence has also demonstrated the undermining effect of monetary reward on intrinsic motivation (ma et al., 2014; murayama et al., 2010). 1.3 the moderating effect of individual differences between reward and intrinsic motivation people with high autonomy orientation tend to perceive their actions as self-determined, and define their success according to the satisfaction of the three basic psychological needs (deci & ryan, 1985, 2000). thus, causality orientation could be understood as an intrapersonal bias that moderates the positive or negative effects of environmental factors on intrinsic motivation. for example, it is possible that people who are autonomy-oriented tend to interpret rewards as having an informational function, that is, as signifying their competence at the task, leading to higher intrinsic motivation towards the task; people who are controlled-oriented usually interpret rewards as controlling. accordingly, hagger and chatzisarantis (2011) found the undermining effect of monetary reward on intrinsic motivation to be observed only in controlled-oriented participants, and no such negative effect was observed for participants dominated by autonomy orientation. however, whether perceived competence mediates the interaction effect of autonomy orientation and reward on intrinsic motivation remains unclear. based on previous results and theoretical assumptions, we hypothesized that for people with high autonomy orientation, the provision of monetary reward will increase their perceived competence and lead to higher intrinsic motivation. as for people with high controlled orientation, reward will decrease their intrinsic motivation. further, we assumed that reward would have no significant influence on perceived competence and intrinsic motivation of people with high levels of impersonal orientation. this is because people with high impersonal orientation often foster amotivation, which makes them insensitive to the effects of external environmental factors. 1.4 effect of choice on intrinsic motivation in the field of motivation, another external environmental factor thought to significantly influence intrinsic motivation is choice. choice has been suggested to increase people’s perception of jiang, kusamoto, & tanaka 72 | f l r self-determination with greater opportunity to learn about and exercise their ability to manipulate, compare, and analyse information (decharms 1968; deci & ryan 1985; jellison & harvey, 1973). thus, providing choice may support a person’s experience of autonomy and sense of competence, in turn leading to higher intrinsic motivation and better performance outcomes (e.g., how et al., 2013. patall et al., 2010; ryan & deci, 2000). despite a great deal of theory and research suggesting that choice is a powerful motivator of behaviour, several studies find that choice may have a negative effect on adaptive motivation and performance outcomes when there are too many options (iyengar & lepper, 2000; patall, 2012), or when the options exceed the individual’s abilities (ryan & deci, 2000; skinner & belmont, 1993). baumeister et al. (1998) proposed that all acts of choice or self-control are effortful and draw on a limited resource that can be depleted, analogous to a source of energy or strength. making choices as one form of self-regulation thus can result in a state of fatigue called ‘ego-depletion’, where the individual experiences a decreased capacity for initiating activity, making choices, or further selfregulate. 1.5 the moderating effect of individual differences between choice and intrinsic motivation although there is no direct evidence showing that individual differences in causality orientation could moderate the effect of choice on perceived competence and intrinsic motivation, some preliminary studies provide sufficient grounds for an assumption. previous research showed that in individualistic cultures, personal agency, independence, and autonomy may be central to one’s self-concept, whereas in more collectivistic cultures, agency may have much less importance (markus & kitayama, 1991). similarly, iyengar and lepper (1999) found that intrinsic motivation was enhanced most for caucasian americans when they were making a personal choice, but for asian americans, intrinsic motivation was enhanced most when trusted authority figures or peers made choices for them. however, we would like to argue that it is not cultural differences that moderate the effect of choice, but rather individual differences in causality orientation, shaped by the overall external environment. it is possible that individuals with high autonomy orientation view choice as a chance to demonstrate their competence, leading to higher intrinsic motivation, whereas people with high controlled orientation are more likely to experience ego depletion when facing choice, leading to a decrease of intrinsic motivation towards activities. in line with this assumption, mouratidis et al (2011) found that students with high relative autonomous motivation were significantly more interested in a class when its overall climate prioritized choice provision, compared to those with low relative autonomous motivation. based on the theoretical assumptions and results of previous research, we hypothesized that the provision of choice will improve the intrinsic motivation of people with high autonomy orientation by satisfying their need for autonomy and improving their competence. choice will have no positive effect on the intrinsic motivation and perceived competence of people with a high controlled orientation as such people can be motivated when their autonomy is supported; however, their motivation might be strongly undermined when their autonomy is thwarted. for the people with high impersonal orientation, we assumed that the enhancing effect of choice on intrinsic motivation will not occur, and that choice will have no positive effect on their competence. this is because people with high impersonal orientation often foster amotivation and are hardly affected by external environmental factors. 1.6 interaction effect of reward and choice on intrinsic motivation in real-world contexts, the simultaneous provision of multiple motivators is rather common. however, the interaction between reward and choice has rarely been examined; we only found two studies focusing on this issue. cohend (1974) found that the effect of choice was essentially zero when a reward external to the choice manipulation was provided, compared to when participants chose the reward they would receive, or when no reward was involved. on the other hand, marinak (2004) suggested that if individuals have some control over the reward, it is not perceived as controlling, and jiang, kusamoto, & tanaka 73 | f l r the positive effect of choice on motivation remains. however, no research has yet examined how reward and choice affect competence, and whether the interaction effect of reward and choice would be moderated by individual differences. based on the results and theoretical assumptions of previous research, we hypothesized that for people with high autonomy orientation, the simultaneous provision of reward and choice will increase their perceived competence, leading to even higher intrinsic motivation. as for people with high controlled orientation, reward and choice will not increase their competence, but strongly decrease their intrinsic motivation. we also assumed that reward and choice would have no significant influence on perceived competence and intrinsic motivation of people with high impersonal orientation. 1.7 present study this study examined the moderating effect of causality orientations on the effect of monetary rewards and choice on perceived competence, and whether it affects intrinsic motivation. we further aimed to explore how reward and choice affect competence and intrinsic motivation when provided simultaneously, and whether it will be moderated by causality orientation. our three hypotheses are presented in figure 1. figure 1. hypotheses of the moderating effect of causality orientation on the relationships between reward, choice, perceived competence, and intrinsic motivation. jiang, kusamoto, & tanaka 74 | f l r note. ao = autonomy orientation, co = controlled orientation, io = impersonal orientation. h1: for people with high autonomy orientation, we hypothesised that reward and choice will increase their perceived competence, which will in turn improve their intrinsic motivation. the positive effects on their competence and intrinsic motivation will be stronger when providing reward and choice simultaneously. h2: for people with high controlled orientation, the provision of reward and choice will not affect their perceived competence and decrease their intrinsic motivation. these negative effects on their intrinsic motivation will be stronger when providing reward and choice simultaneously. h3: for people with high impersonal orientation, reward and choice were hypothesised to have no effect on their perceived competence and intrinsic motivation, whether provided separately or together. 2. method 2.1 participants participants comprised 103 undergraduate students (41 male, 62 female, mage = 19.562, sd = 1.101) taking an optional psychology class at a private university in japan. additional course credits were given to the students who participated in the study. data were collected from october to november 2017. informed consent was obtained from all participants and the experiment was authorized by the ethics review committee of the university (16061, ‘the effect of individual differences on intrinsic motivation’). 2.2 task we used the solid puzzle (soma cube) used by deci (1971) as the task of the present study. the puzzle contains seven pieces, all of which are in different shapes. in the present study, the puzzle configurations we asked the participants to replicate were constructed with four pieces; participants needed to figure out which four out of the seven pieces were to be used to solve each puzzle. before the main experiment, we conducted a pretest to examine whether the task was interesting for university students. a total of 21 university students (7 male and 14 female, mage = 20.841, sd = 1.281) enrolled in the pretest. participants were asked to replicate three out of nine puzzle configurations, which we used in the main experiment. after they had finished solving all the puzzles, their intrinsic motivation towards the task was measured with three items from the intrinsic motivation inventory (see measures). participants’ mean score on this scale was 6.016 and the standard deviation (sd) was 0.711. a wilcoxon rank sum test showed that the mean score was significantly higher than the median of the scale at p < .001, indicating that the participants perceived the task as interesting. 2.3 experimental design this study adopted the free-choice paradigm of deci (1971) to examine the effects of rewards and choice on intrinsic motivation. participants of the main experiment were randomly assigned to one of four conditions: reward (reward vs. no reward) × choice (choice vs. no choice). reward in the present research was 100 jpy (about 0.9 us dollars) for each time the participants solved a puzzle, and choice in the present research was provided through the freedom of solving puzzles in the order preferred by the participants. jiang, kusamoto, & tanaka 75 | f l r 2.4 procedure figure 2. the experimental process flow in the present research. as shown in figure 2, participants’ causality orientations were measured before the experiment on the day on which they were recruited; the experiment contained three periods in total. at time 1, all participants jiang, kusamoto, & tanaka 76 | f l r were presented with three printed illustrations of puzzle configurations and asked to replicate all the puzzles as quickly as possible. no information about a reward was mentioned during this period, and all participants were asked by the experimenters to solve puzzles in a specific order. at time 2, all participants were presented with three new printed illustrations of puzzle configurations. participants assigned to the reward conditions were informed that they would receive 100 jpy after solving each puzzle. participants in the choice groups were free to solve the puzzles in their preferred order, whereas participants assigned to the no choice conditions were asked by the experimenters to solve the puzzles in a specific order. at time 3, all participants were given three new printed illustrations of puzzle configurations again, but none of the participants received monetary reward. participants in the reward conditions were told that they would not be paid for solving puzzles in this third period because there was only enough money to pay them in one period. all the participants were asked by the experimenters to solve the puzzles in a specific order. after the participants had solved the puzzles at each period, they were asked to complete a brief measure on their perceptions of the task, including task enjoyment and perceived competence. the experimenters then excused themselves from the laboratory for 300 seconds. in the first and second period, the experimenters told the participants that they needed to fetch more configurations for the next period; in the third period, the experimenters told the participants that they had to fetch another questionnaire covering their task persistence and interest in other puzzles. just before leaving the room, the experimenters provided the participants a few magazines and an exercise book that contained more illustrations of puzzle configurations. the experimenters then casually informed the participants that they could read those magazines, continue with the puzzles, or engage in any other activity except for leaving the room, till the experimenters returned. participants' activities in the experimenters' absence were monitored by the experimenters using a hidden video camera. in the end, participants were debriefed about the true purpose of the experiment and the use of hidden camera. 2.5 measures 2.5.1 causality orientation autonomy, controlled, and impersonal orientations were measured with the japanese version (tanaka & sakurai, 1995) of the general causality orientations scale (gcos; deci & ryan, 1985). the gcos comprises 12 written vignettes for which participants rate each of the three possible responses, relating to autonomy orientation, control orientation, and impersonal orientation, respectively. for example, one of the vignettes asks: ‘you have just received the results of a test you took and discovered that you did very poorly. your initial reaction is likely to be…’ the autonomy-oriented response is, ‘i wonder how i did so poorly’ (feeling disappointed) (cronbach's α = .704); the controloriented response is, ‘that stupid test does not show anything’ (feeling angry) (α = .704); and the impersonal-oriented response is, ’i cannot do anything right’ (feeling sad) (α = .633). notice that the original gcos scores range from 1 to 7, but the scores in the japanese version ranged from 1 to 4; in order to measure participants' causality orientations as accurately as measured by the original scale, participants in this study provided responses ranging from 1 (very unlikely) to 5 (very likely). 2.5.2 intrinsic motivation participants' intrinsic motivation was measured using the time spent on free-choice behaviour and self-reported task enjoyment, both of which are generally used to operationalize intrinsic motivation (e.g., hendijani et al., 2016, ng, 2018, woolley, & fishbach, 2018, for meta-analysis, see cerasoli et al., 2014). during each of three free-choice periods when the experimenters left the room, the time spent by the participants on the puzzles was considered as the behavioural index for intrinsic motivation towards the task. each free-choice period lasted 300 seconds. notice that the configurations presented jiang, kusamoto, & tanaka 77 | f l r to the participants during the free-choice period were different from the ones presented to them during the tasks, as the former were constructed with all seven pieces of the soma cube puzzle instead of four. three items from the intrinsic motivation inventory (ryan, 1982) were used as a self-report measure of task enjoyment (e.g. ‘i enjoyed doing this task’). this measure has been used in several studies in japan (e.g., kakinuma et al., 2020; mogami et al., 2011). responses were codified via a 7point likert scale ranging from 1 (not at all) to 7 (very much). the scale's cronbach's α values were .894, .883, and .879 at times 1, 2, and 3, respectively. 2.5.3 perceived competence participants' perceived competence in the task was measured using three items from the selfdeveloped scale used in houlfort et al.’s (2002) study, with responses rated on a 7-point likert scale ranging from 1 (not at all) to 7 (very much) (e.g. ‘i am good at this task’). this scale's cronbach's α values were .937, .902, and .917 at times 1, 2, and 3, respectively. we carried out an english-tojapanese translation and back-translation procedure for this scale. 3. results 3.1 descriptive statistics three participants were excluded from the final analysis because they did not solve at least one puzzle during the experiment1. table 1 presents the means and sds of all measures across different study conditions. the mean for autonomy orientation was higher than that for controlled orientation and impersonal orientation, and 83% of the participants had a higher autonomy orientation than controlled or impersonal orientation. two-way anova results revealed no significant main effect or interaction effect of reward and choice on the participants' intrinsic motivation measures, perceived competence at time 1, and causality orientation, indicating that participants were randomly assigned to each of the four conditions. table 2 presents the correlations between all the variables. 1 we did not find any sample characteristics that differed for excluded participants after conducting t-tests for all measured variables. jiang, kusamoto, & tanaka 78 | f l r table 1 means (standard deviations) of all measures choice no choice variables reward (n = 24) no reward (n = 26) reward (n = 24) no reward (n = 26) causality orientation autonomy orientationa 3.740 (.480) 3.670 (.594) 3.758 (.501) 3.873 (.355) controlled orientationa 3.049 (.539) 2.985 (.557) 3.008 (.460) 3.004 (.352) impersonal orientationa 2.911 (.382) 3.239 (.509) 3.179 (.616) 3.068 (.395) perceived competencec [t1] 3.667 (1.504) 4.093 (1.847) 3.597 (1.500) 3.654 (1.413) perceived competencec [t2] 4.250 (1.275) 4.707 (1.532) 4.236 (1.261) 3.949 (1.278) perceived competencec [t3] 4.611 (1.399) 4.653 (1.568) 4.264 (.997) 4.179 (1.437) intrinsic motivation time spent on puzzlesb [t1] 193.916 (131.480) 231.308 (107.187) 209.542 (127.750) 243.769 (90.309) time spent on puzzlesb [t2] 151.208 (141.370) 174.885 (137.450) 224.708 (132.533) 220.269 (124.369) time spent on puzzlesb [t3] 145.083 (144.690) 207.000 (116.758) 210.625 (133.292) 186.346 (138.094) task enjoymentc [t1] 5.236 (1.202) 5.693 (1.273) 5.819 (.927) 5.846 (.870) task enjoymentc [t2] 5.361 (1.016) 5.707 (1.207) 5.833 (.983) 5.769 (.831) task enjoymentc [t3] 5.319 (1.144) 5.947 (1.061) 5.847 (.901) 5.808 (.839) note. apossible range = 1–5 points; bpossible range = 0–300 s; cpossible range = 1–7 points. jiang, kusamoto, & tanaka 79 | f l r table 2 correlations between all variables (n = 100) variables 1 2 3 4 5 6 7 8 9 10 11 causality orientation 1. autonomy orientation 2. controlled orientation .427** 3. impersonal orientation -.146 -.054 4. perceived competence [t1] .328** .418*** -.249* 5. perceived competence [t2] .232* .398*** -.235* .732*** 6. perceived competence [t3] .213 .336** -.288** .660*** .844*** intrinsic motivation 7. time spent on puzzles [t1] .131 -.001 .149 -.010 .034 .130 8. time spent on puzzles [t2] .039 -.035 .183 -.049 -.053 -.008 .763*** 9. time spent on puzzles [t3] .062 -.063 .242* .066 .119 .066 .574*** .656*** 10. task enjoyment [t1] .187 .203 -.175 .360*** .313** .348*** .235* .184 .255* 11. task enjoyment [t2] .247* .147 -.111 .322*** .352** .315** .245* .290** .390*** .823*** 12. task enjoyment [t3] .191 .080 -.070 .285** .312** .289** .316** .341** .474*** .789*** .852*** note. reward = 1; no reward = 0. choice = 1; no choice = 0. *p < .05, ** p < .01, ***p < .001 jiang, kusamoto, & tanaka 80 | f l r 3.2 moderating effects of causality orientation following the landmark study of deci (1971), we focused on the difference between time 1 and time 3 to examine the interaction effects between causality orientation and the contextual factors of reward and choice on participants' intrinsic motivation2. in the present study, we conducted a series of multiple regressions3 using causality orientation, reward, and task choice as predictor variables; perceived competence, time spent on puzzles, and task enjoyment at time 3 as outcome variables; and controlling the objective variables of time 1. we also examined the interaction effect between causality orientation and reward, and causality orientation and task choice on outcome variables. reward vs. no reward and choice vs. no choice conditions were dummy coded as 1 and 0, respectively. all interaction product terms were constructed using dummy codes and other predictor variables, which were centred (aiken et al., 1991). for each model, all possible interaction product terms were included in initial analyses, but insignificant interactions were trimmed from the final models. all results reported below are from the final, trimmed models; participants' gender, age, and the number of solved puzzles at time 3 were controlled in each model. table 3 shows the results of multiple regression concerning perceived competence at time 3. no significant main effect was found, but the interaction effect between autonomy orientation and choice on perceived competence was significant (β = .341, p = .010). table 4 shows the multiple regression results for time spent on puzzles at time 3. the main effects of causality orientation, reward, and choice, and the interaction effects between causality orientation and reward, and causality orientation and task choice were not significant (all ps > .05). table 5 shows the results of multiple regression concerning task enjoyment at time 3. no significant main effect was found in these results, but the interaction effect between autonomy orientation and reward on task enjoyment was significant (β = .211, p = .028). next, we carried out simple slope analyses to further explore these significant interaction effects. 2 results of time 2 are reported in the appendix section. 3 in the present study we did not use sem due to our small sample size. according to kline (2011), a typical sample size in studies where sem is used is about 10 cases per parameter. each model we built has at least 45 parameters, thus requiring about 450 cases. with this in mind, we conducted multiple regression analyses instead. jiang, kusamoto, & tanaka 81 | f l r table 3 multiple regression results concerning perceived competence at time 3 perceived competence [t3] step 1 step 2 step 3 b se β p 95% ci of b b se β p 95% ci of b b se β p 95% ci of b causality orientation autonomy orientation -.077 .260 -.027 .769 [-.595, .442] -.064 .266 -.022 .809 [-.595, .466] -.164 .248 -.051 .514 [-.614, .345] controlled orientation .175 .278 .059 .532 [-.380, .730] .171 .282 .058 .546 [-.391, .734] .130 .271 .044 .634 [-.411, .670] impersonal orientation -.240 .238 -.085 .318 [-.715, .235] -.229 .242 -.081 .348 [-.711, .254] -.333 .235 -.118 .162 [-.803, .137] perceived competence [t1] .546 .087 .606 .000 [ .372, .721] .542 .088 .601 .000 [ .366, .718] .524 .085 .581 .000 [ .354, .693] reward .020 .235 .007 .933 [-.448, .488] .010 .225 .004 .964 [-.439, .459] task choice .235 .225 .085 .300 [-.214, .684] .251 .216 .090 .250 [-.180, .682] autonomy orientation*choice 1.262 .475 .341 .010 [.315, 2.209] r2 .545*** .553*** .594*** δ r2 .007 .043* note. insignificant interaction effects were removed from the final models. gender, age, and number of solved puzzles at time 3 were included as control variables, but results are omitted for brevity. δ r2 shows the change in r2 when adding interaction effects into the model. given δr2 in step 3 at .043, the effect size of f2 was calculated to be .092 (cohen, 2013). jiang, kusamoto, & tanaka 82 | f l r table 4 multiple regression results for time spent on puzzles at time 3 time spent on puzzles [t3] step 1 step 2 b se β p 95% ci of b b se β p 95% ci of b causality orientation autonomy orientation 12.857 30.260 .046 .672 [-47.465, 73.180] 13.140 30.582 .047 .660 [-47.853, 74.133] controlled orientation -24.070 30.940 -.085 .439 [-85.747, 37.607] -23.843 30.996 -.084 .433 [-85.664, 37.977] impersonal orientation 42.328 27.927 .155 .134 [-13.343, 97.999] 41.466 27.972 .152 .143 [-14.321, 97.254] time spent on puzzles [t1] .645 .124 .536 .000 [.398, .892] .636 .126 .126 .000 [.385, .887] reward 4.163 27.292 .016 .879 [-50.270, 58.595] task choice -37.861 25.931 -.141 .149 [-89.578, 13.856] r2 .342 .362 δ r2 .020 note. insignificant interaction effects were removed from the final models. gender, age, and number of solved puzzles at time 3 were included as control variables, but the results are omitted for brevity. δ r2 shows the change in r2 when adding interaction effects into the model. jiang, kusamoto, & tanaka 83 | f l r table 5 multiple regression results for task enjoyment at time 3 task enjoyment [t3] step 1 step 2 step 3 b se β p 95% ci of b b se β p 95% ci of b b se β p 95% ci of b causality orientation autonomy orientation .231 .161 .110 .154 [-.089, .551] .197 .164 .094 .232 [-.129, .524] -.084 .202 -.040 .680 [-.488, .320] controlled orientation -.291 .166 -.138 .078 [-.629, .034] -.275 .168 -.128 .106 [-.610, .060] -.335 .165 -.156 .053 [-.665,-.005] impersonal orientation .177 .147 .086 .231 [-.115, .469] .154 .149 .074 .306 [-.143, .450] .193 .146 .094 .190 [-.098, .484] task enjoyment [t1] .819 .075 .813 .000 [ .669, .969] .805 .078 .799 .000 [ .650, .960] .829 .076 .823 .000 [ .677, .981] reward -.175 .147 -.086 .238 [-.468, .118] -.163 .143 -.080 .259 [-.448, .123] task choice -.016 .142 -.008 .910 [-.300, .267] -.009 .138 -.005 .946 [ .285, .266] autonomy orientation*reward .655 .291 .211 .028 [.074, 1.235] r2 .668 .674 .697 δ r2 .007 .022 note. insignificant interaction effects were removed from the final models. gender, age, and number of solved puzzles at time 3 were included as control variables, but the results are omitted for brevity. δ r2 shows the change in r2 when adding interaction effects into the model. given δr2 in step 3 at .022, the effect size of f2 was calculated to be .076 (cohen, 2013). jiang, kusamoto, & tanaka 84 | f l r figure 3 shows the interaction effect of autonomy orientation and task choice on perceived competence. the results of simple slope analysis revealed that the regression of perceived competence on task choice for people with high autonomy orientation was significant (b = .852, p = .010), indicating that freedom to choose the order in which the participants solve the puzzles increased the perceived competence of participants with high autonomy orientation. the regression of perceived competence on task choice for those with low autonomy orientation was not significant (b = -.035, p = .263). figure 3. interaction effect between task choice and autonomy orientation on perceived competence at time 3. note. the no choice and choice conditions were coded as 0 and 1, respectively. figure 4. confidence band of the interaction effects of task choice and autonomy orientation (ao) on perceived competence at time 3. note. the standard deviation of centred autonomy orientation was 0.49 and the possible range of centred autonomy orientation was from -1.31 to 1.02. moderator: autonomy orientation m ar gi na l e ffe ct o f t as k ch oi ce bchoice+bao:choiceaoi = 0 rejected jiang, kusamoto, & tanaka 85 | f l r we also calculated the confidence band of interaction effects following the method of johnson (2019). as depicted in figure 4, the regression of perceived competence on task choice for people with high autonomy orientation was significant when the centred score of autonomy orientation was higher than .17. simple slope analysis showed that the regression of perceived competence on task choice for those with low autonomy orientation was not significant when the score of autonomy orientation was 1 sd (.49) less than its mean value; however, the confidence band of interaction effects indicated that the regression of perceived competence on task choice for those with low autonomy orientation was significant when the score of autonomy orientation was lower than -1.11, which is about 2.26 sds less than the mean. importantly, -1.11 was within the possible range of centred autonomy orientation (from2.76 to 1.24), which means that this result was also substantially significant. figure 5 shows the interaction effect of autonomy orientation and reward on task enjoyment. simple slope analysis revealed the regression of task enjoyment on reward for individuals with low autonomy orientation to be significant (b = -.489, p = .015), indicating that presenting a monetary reward led to lower task enjoyment in participants with low autonomy orientation than when no reward was presented. the regression of task enjoyment on reward for those with high autonomy orientation was not significant (b = .145, p = .478). figure 5. interaction effect between reward and autonomy orientation on task enjoyment at time 3. note. the no reward and reward conditions were coded as 0 and 1, respectively. figure 6 shows the confidence band of the interaction effects of autonomy orientation and reward on task enjoyment at time 3. the regression of task enjoyment on reward for individuals with low autonomy orientation was significant when the centred score of autonomy orientation was lower than -.20. jiang, kusamoto, & tanaka 86 | f l r figure 6. confidence band of interaction effects of reward and autonomy orientation (ao) on task enjoyment at time 3. note. the standard deviation of centred autonomy orientation was .49, and the possible range of centred autonomy orientation was -1.31 to 1.02. 4. discussion this study examined the moderating effect of causality orientations on the effect of monetary rewards and choice on perceived competence, and whether it affected intrinsic motivation. we further aimed to explore how reward and choice would affect competence and intrinsic motivation when provided simultaneously, and whether it would be moderated by causality orientation. our first hypothesis was partially supported as we found a significant interaction effect of choice and autonomy orientation on perceived competence, although it did not affect intrinsic motivation. participants with high autonomy orientation had higher perceived competence when given the choice over their task, whereas participants with low autonomy orientation had lower perceived competence. these results are in line with the findings reported by mouratidis et al. (2011) on the interaction between daily fluctuation in need satisfaction and trait-level differences in need satisfaction. that is, just as the daily need satisfaction was higher among individuals who are good at satisfying their needs, our study found that the provision of choice had a stronger effect on perceived competence for individuals who function more autonomously. moreover, participants in the present research completed most of the puzzle tasks (m = 8.360, sd = 0.990, nine in total). considering that people with high autonomy orientation tend to reflect their success in fulfilling their basic psychological needs (ryan & deci, 2017), they might interpret having choice over a task as an opportunity to demonstrate competence, and thus perceived themselves as competent after solving the puzzle tasks. in contrast, people with low autonomy orientation may experience frustration when facing a choice, leading to a decrease in their perceived competence. in the current study, the effect of reward on intrinsic motivation was moderated by autonomy orientation, although the effect was not mediated by perceived competence. reward had no significant moderator: autonomy orientation breward+bao:rewardaoi = 0 rejected m ar gi na l e ffe ct o f r ew ar d (b re w ar d+ b a o: re w ar da o i ) jiang, kusamoto, & tanaka 87 | f l r effect on task enjoyment of people with high autonomy orientation. in contrast, the enjoyment of participants with low autonomy orientation decreased significantly after receiving the reward, partially supporting our first hypothesis that the negative effect of reward on intrinsic motivation will not occur in people with high autonomy orientation. this result is also in line with the findings of hagger and chatzisarantis (2011), who stated that people with low autonomy orientation tend to focus on the controlling aspect of the reward, attributing their actions to this external environmental factor instead of to their own will, which consequently suppresses their intrinsic motivation towards the task. although the positive correlation between time spent on puzzles and task enjoyment in each time period were all significant, we found no significant interaction effect of autonomy orientation and reward on the behavioural index of intrinsic motivation. this might be because during the free choice periods, participants were given new illustrations of puzzle configurations that required the use of all seven pieces of the soma cube to complete. we increased the difficulty to prevent participants from getting used to the puzzle and yield to boredom; however, this change in difficulty might have confounded with the effect of our manipulation of reward and the effect of autonomy orientation on the behaviour index of intrinsic motivation. we did not find any significant interaction effect between reward and choice on competence and intrinsic motivation; autonomy orientation did not moderate interaction effects. our study followed the methods of deci (1971), who aimed to verify whether the external contingency of monetary reward can motivate people only as long as it exists. we adapted this method to examine whether the effect of choice on intrinsic motivation could last even after we withdraw the provision of that choice. our results showed that the effect of choice was powerful enough to affect the participants' competence even after we stopped providing that choice. nevertheless, the procedure of the present research still needs some modification, which will be discussed in the limitation section. our second hypothesis was not supported, as the results showed no significant interaction effect between reward, choice, and controlled orientation on intrinsic motivation. these non-significant results could be attributed to the relatively high autonomy orientation and low controlled orientation in most of our participants, such that the effect of controlled orientation was weak even after we controlled the effect of autonomy orientation. this made it difficult to infer how individuals with high controlled orientation would react to reward, despite our dimensional analysis of causality orientation. controlled orientation is suggested to be associated with type-a personality (deci & ryan, 1985, 2000), which is a pattern of behaviours and emotions with an excessive emphasis on competition, aggression, impatience, and hostility (friedman & booth-kewley, (1987). controlled orientation is also highly related to public self-consciousness, and problems with gambling and alcohol use. a high controlled orientation is linked to motivation and persistence, but this type of motivation is non-optimal and predictive of poorer well-being than an autonomy orientation (ryan & deci, 2017). the interaction effects between reward, choice, and impersonal orientation on both perceived competence and intrinsic motivation were not significant. while these results are consistent with our third hypothesis, it is difficult to determine whether the third hypothesis was supported or not, due to potential type ii error caused by the imbalance in causality orientations among our participants. the negative correlation between impersonal orientation and perceived competence in all three time periods were significant. according to deci and ryan (1985, 2000), people with high impersonal orientation tend to see themselves as incompetent and unable to master situations, which makes them generally difficult to motivate. consequently, they develop a pervasive sense of incompetence that leaves them vulnerable to failure experiences, depressive symptoms, social anxiety, low self-esteem, hostility, and a procrastinatory approach to decision making. more studies on how to motivate people with high impersonal orientation are necessary. however, it might be more important to understand how to create an environment that consistently satisfies students' psychological needs such that they can develop a strong autonomy orientation instead of an impersonal orientation. jiang, kusamoto, & tanaka 88 | f l r 4.1 limitations our study provides preliminary support for the moderating effect of causality orientation on the relationship between environmental conditions and intrinsic motivation; however, some limitations should be considered. first, we did not conduct a power analysis or other sample size justification prior to participant recruitment. future studies should use the effect sizes of the present research to calculate a sufficient sample size for achieving adequate power. second, the distribution of causality orientations in our sample was unbalanced. our methodology could be improved by screening participants for high and low causality orientations before randomization into conditions. future studies could also recruit participants from different classes, schools, regions, or even populations that are not university students. third, the paradigm used to examine the effect of choice in the present study needs to be modified. a simple design of two levels (choice vs no choice) with a single step might be the proper way to achieve our initial goal. fourth, the artificial setting in which the study was conducted may cast doubt on its ecological validity, especially as it focused on intrinsic motivation. 4.2 conclusion tailoring motivation according to individual differences might be difficult to execute in reallife contexts. as it is virtually impossible for one teacher to teach dozens of students in different ways in the traditional education system, or creating a collaborative environment that fits every employees' motivational needs. however, the recent dramatic development of technology use in different contexts can facilitate the development of more robust resources for everyone, for instance, by employing personalized learning systems grounded in research on learner differences to adapt to each student's needs and characteristics. future studies could also explore training on motivation self-regulation in relation to individual contingencies, as an alternative approach to developing personalized support system. funding this work was supported by jsps kakenhi (grant number 16k11978, 16h06406). keypoints the negative effect of reward on intrinsic motivation only occurred in people with low autonomy orientation. providing choice enhanced perceived competence of people with high autonomy orientation but undermined it in people with low autonomy orientation. impersonal orientation was negatively related to perceived competence. jiang, kusamoto, & tanaka 89 | f l r references aiken, l. s., west, s. g., & reno, r. r. (1991). multiple regression: testing and interpreting interactions. sage. baumeister, r. f., bratslavsky, e., muraven, m., & tice, d. m. (1998). ego-depletion: is the active self a limited resource? journal of personality and social psychology, 74, 1252–1265. https://doi.org/10.1037/0022-3514.74.5.1252 black, a. e., & deci, e. l. (2000). the effects of instructors' autonomy support and students' autonomous motivation on learning organic chemistry: a self-determination theory perspective. science education, 84, 740–756. https://doi.org/10.1002/1098-237x(200011)84:6%3c740::aidsce4%3e3.0.co;2-3. bong, m., & clark, r. e. (1999). comparison between self-concept and self-efficacy in academic motivation research. educational psychologist, 34, 139-153. https://doi.org/10.1207/s15326985ep3403_1 brandstätter, v., job, v., & schulze, b. (2016). motivational incongruence and well-being at the workplace: person-job fit, job burnout, and physical symptoms. frontiers in psychology, 7, 1153. https://doi.org/10.3389/fpsyg.2016.01153 *byron, k., & khazanchi, s. (2012). rewards and creative performance: a meta-analytic test of theoretically derived hypotheses. psychological bulletin, 138, 809–830. https://doi.org/10.1037/a0027652 cameron, j. (2001). negative effects of reward on intrinsic motivation—a limited phenomenon: comment on deci, koestner, and ryan (2001). review of educational research, 71, 29-42. https://doi.org/10.3102/00346543071001029 cameron, j., pierce, w. d., banko, k. m., & gear, a. (2005). achievement-based rewards and intrinsic motivation: a test of cognitive mediators. journal of educational psychology, 97, 641–655. https://doi.org/10.1037/0022-0663.97.4.641. catania, g., & randall, r. (2013). the relationship between age and intrinsic and extrinsic motivation in workers in a maltese cultural context. international journal of arts & sciences, 6, 31-45. *cerasoli, c. p., nicklin, j. m., & ford, m. t. (2014). intrinsic motivation and extrinsic incentives jointly predict performance: a 40-year meta-analysis. psychological bulletin, 140, 980. https://doi.org/10.1037/a0035661 cohen, d. s. (1974). the effects of task choice, monetary, and verbal reward on intrinsic motivation: a closer look at deci’s cognitive evaluation theory. unpublished doctoral dissertation, ohio state university. cohen, j. (1988). statistical power analysis for the behavioral sciences. hillsdale, new jersey: lawrence erlbaum associates. covington, m. v., & müeller, k. j. (2001). intrinsic versus extrinsic motivation: an approach/avoidance reformulation. educational psychology review, 13, 157–176. https://doi.org/10.1023/a:1009009219144 decharms, richard (1968), personal causation. new york: academic press. deci, e. l. (1971). effects of externally mediated rewards on intrinsic motivation. journal of personality and social psychology, 18, 105–115. https://doi.org/10.1037/h0030644 deci, e. l. (1980). the psychology of self-determination. dc heath and company. deci, e. l., & ryan, r. m. (1985). the general causality orientations scale: self-determination in personality. journal of research in personality, 19, 109–134. https://doi.org/10.1016/00926566(85)90023-6 deci, e. l., & ryan, r. m. (2000). the "what" and "why" of goal pursuits: human needs and the selfdetermination of behavior. psychological inquiry, 11, 227–268. https://doi.org/10.1207/s15327965pli1104_01 deci, e. l., koestner, r., & ryan, r. m. (2001). extrinsic rewards and intrinsic motivation in education: reconsidered once again. review of educational research, 71, 1–27. https://doi.org/10.3102/00346543071001001 elliot, a. j., & hulleman, c. s. (2017). achievement goals. in a. j. elliot, c. s. dweck, & d. s. yeager jiang, kusamoto, & tanaka 90 | f l r (eds.), handbook of competence and motivation: theory and application (p. 43–60). the guilford press. gilakjani, a. p. (2012). a match or mismatch between learning styles of the learners and teaching styles of the teachers. international journal of modern education and computer science, 4, 51–60. https://doi.org/10.5815/ijmecs.2012.11.05 hagger, m. s., & chatzisarantis, n. l. d. (2011). causality orientations moderate the undermining effect of rewards on intrinsic motivation. journal of experimental social psychology, 47, 485–489. https://doi.org/10.1016/j.jesp.2010.10.010 hendijani, r., bischak, d. p., arvai, j., & dugar, s. (2016). intrinsic motivation, external reward, and their effect on overall motivation and performance. human performance, 29, 251–274. https://doi.org/10.1080/08959285.2016.1157595 huang, y., lv, w., & wu, j. (2016). relationship between intrinsic motivation and undergraduate students’ depression and stress: the moderating effect of interpersonal conflict. psychological reports, 119, 527-538. doi: 10.1177/0033294116661512 houlfort, n., koestner, r., joussemet, m., nantel-vivier, a., & lekes, n. (2002). the impact of performance-contingent rewards on perceived autonomy and competence. motivation and emotion, 26, 279–295. how, y. m., whipp, p., dimmock, j., & jackson, b. (2013). the effects of choice on autonomous motivation, perceived autonomy support, and physical activity levels in high school physical education. journal of teaching in physical education, 32, 131-148. doi: 10.1123/jtpe.32.2.131 jellison, j. m., & harvey, j. h. (1973). determinants of perceived choice and the relationship between perceived choice and perceived competence. journal of personality and social psychology, 28, 376382. https://doi.org/10.1037/h0035110 johnson, p. e. (2019). using rockchalk for regression analysis. https://mirrors.nics.utk.edu/cran/web/packages/rockchalk/vignettes/rockchalk.pdf jovanovic, d., & matejevic, m. (2014). relationship between rewards and intrinsic motivation for learning–researches review. procedia-social and behavioral sciences, 149, 456–460. https://doi.org/10.1016/j.sbspro.2014.08.287 kakinuma, k., nakai, m., hada, y., kizawa, m., & tanaka, a. (2020). praise affects the “praiser”: effects of ability-focused vs. effort-focused praise on motivation. the journal of experimental education, 1-22. https://doi.org/10.1080/00220973.2020.1799313 khalaila, r. (2015). the relationship between academic self-concept, intrinsic motivation, test anxiety, and academic achievement among nursing students: mediating and moderating effects. nurse education today, 35, 432-438. https://doi.org/10.1016/j.nedt.2014.11.001 kozlowski, s. w. (ed.). (2012). the oxford handbook of organizational psychology (vol. 1). oxford university press. lepper, m. r., corpus, j. h., & iyengar, s. s. (2005). intrinsic and extrinsic motivational orientations in the classroom: age differences and academic correlates. journal of educational psychology, 97, 184-196. https://doi.org/10.1037/0022-0663.97.2.184 li, y., wei, f., ren, s., & di, y. (2015). locus of control, psychological empowerment and intrinsic motivation relation to performance. journal of managerial psychology, 30, 422-438. https://doi.org/10.1108/jmp-10-2012-0318 iyengar, s. s., & lepper, m. r. (1999). rethinking the value of choice: a cultural perspective on intrinsic motivation. journal of personality and social psychology, 76, 349-366. https://doi.org/10.1037/0022-3514.76.3.349 iyengar, s. s., & lepper, m. r. (2000). when choice is demotivating: can one desire too much of a good thing? journal of personality and social psychology, 79, 995-1006. https://doi.org/10.1037/00223514.79.6.995 ng, b. (2018). the neuroscience of growth mindset and intrinsic motivation. brain sciences, 8, 20-30. marinak, b. a. (2004). the effects of reward proximity and choice of reward on the reading motivation of third-grade students. unpublished doctoral dissertation, university of maryland, college park. mcadams, d. p., & pals, j. l. (2006). a new big five: fundamental principles for an integrative science of personality. american psychologist, 61, 204–217. https://doi.org/10.1037/0003-066x.61.3.204 jiang, kusamoto, & tanaka 91 | f l r mouratidis, a., vansteenkiste, m., sideridis, g., & lens, w. (2011). vitality and interest–enjoyment as a function of class-to-class variation in need-supportive teaching and pupils' autonomous motivation. journal of educational psychology, 103, 353–366. https://doi.org/10.1037/a0022773 mistree, f., panchal, j. h., schaefer, d., allen, j. k., haroon, s., & siddique, z. (2014). personalized engineering education for the twenty-first century. in m. gosper, & d. ifenthaler (eds,), curriculum models for the 21st century (pp. 91–111). springer. murayama, k., matsumoto, m., izuma, k., & matsumoto, k. (2010). neural basis of the undermining effect of monetary reward on intrinsic motivation. proceedings of the national academy of sciences, 107(49), 20911–20916. https://doi.org/10.1073/pnas.1013305107 omansky, r., eatough, e. m., & fila, m. j. (2016). illegitimate tasks as an impediment to job satisfaction and intrinsic motivation: moderated mediation effects of gender and effort-reward imbalance. frontiers in psychology, 7, 1818. https://doi.org/10.3389/fpsyg.2016.01818 patall, e. a. (2012). the motivational complexity of choosing: a review of theory and research. in r. m. ryan (ed.), the oxford handbook of human motivation. usa: oxford university press. *patall, e. a., cooper, h., & robinson, j. c. (2008). the effects of choice on intrinsic motivation and related outcomes: a meta-analysis of research findings. psychological bulletin, 134, 270–300. https://doi.org/10.1037/0033-2909.134.2.270 putra, e. d., cho, s., & liu, j. (2017). extrinsic and intrinsic motivation on work engagement in the hospitality industry: test of motivation crowding theory. tourism and hospitality research, 17, 228-241. https://doi.org/10.1177/1467358415613393 patall, e. a., cooper, h., & wynn, s. r. (2010). the effectiveness and relative importance of choice in the classroom. journal of educational psychology, 102, 896-915. https://doi.org/10.1037/a0019545 *rawsthorne, l. j., & elliot, a. j. (1999). achievement goals and intrinsic motivation: a meta-analytic review. personality and social psychology review, 3, 326-344. https://doi.org/10.1207/s15327957pspr0304_3 ryan, r. m. (1982). control and information in the intrapersonal sphere: an extension of cognitive evaluation theory. journal of personality and social psychology, 43, 450–461. https://doi.org/10.1037/0022-3514.43.3.450 ryan, r. m., & deci, e. l. (2006). self‐regulation and the problem of human autonomy: does psychology need choice, self‐determination, and will? journal of personality, 74, 1557–1586. https://doi.org/10.1111/j.1467-6494.2006.00420.x ryan, r. m., & deci, e. l. (2017). self-determination theory: basic psychological needs in motivation, development, and wellness. guilford press. skaalvik, s., & skaalvik, e. m. (2004). gender differences in math and verbal self-concept, performance expectations, and motivation. sex roles, 50, 241-252. https://doi.org/10.1023/b:sers.0000015555.40976.e6 skinner, e. a., & belmont, m. j. (1993). motivation in the classroom: reciprocal effects of teacher behavior and student engagement across the school year. journal of educational psychology, 85, 571-581. https://doi.org/10.1037/0022-0663.85.4.571 tanaka, h., & sakurai, s. (1995). construction of the japanese version of the general causality orientations scale. nara university of education academic repository, 31, 177–184. wang, j. h. y., & guthrie, j. t. (2004). modeling the effects of intrinsic motivation, extrinsic motivation, amount of reading, and past reading achievement on text comprehension between us and chinese students. reading research quarterly, 39, 162-186. doi: 10.1598/rrq.39.2.2 warneken, f., & tomasello, m. (2008). extrinsic rewards undermine altruistic tendencies in 20-montholds. developmental psychology, 44, 1785. https://doi.org/10.1037/a0013860 woolley, k., & fishbach, a. (2018). it’s about time: earlier rewards increase intrinsic motivation. journal of personality and social psychology, 114, 877-890. https://doi.org/10.1037/pspa0000116 jiang, kusamoto, & tanaka 92 | f l r appendix we conducted a series of multiple regressions using causality orientation, reward, and task choice as predictor variables; and time spent on puzzles, task enjoyment, and perceived competence of time 2 as outcome variables, while controlling the objective variables of time 1. we also examined the interaction effect between causality orientation and reward, and causality orientation and task choice on outcome variables. all results reported below are from the final, trimmed models; participants' gender, age, and the number of solved puzzles at time 2 were controlled in each model. table a-1 shows the results of multiple regression concerning perceived competence at time 2. no significant main effect and interaction effects were found. table a-2 shows the results of multiple regression concerning time spent on puzzles at time 2. the main effect of task choice was significant (β = -.194, p = .016). this result indicated that compared to those participants who were not able to choose the order of solving puzzles, the participants who were able to choose the order of solving puzzles during time 2 spent significantly less time on the puzzles during the free-choice period that followed. the main effect of causality orientation and reward, and interaction effects between causality orientation and reward, and causality orientation and task choice were not significant. table a-3 shows the results of multiple regression concerning task enjoyment at time 2. no significant main effect and interaction effects were found. jiang, kusamoto, & tanaka 93 | f l r table a-1 multiple regression results concerning perceived competence at time 2. perceived competence [t2] step 1 step 2 b s e β p 95% ci b s e β p 95% ci causality orientations autonomy orientation .127 . 235 .045 . 590 [.512, .723] .154 . 241 .055 . 526 [-.635, .328] controlled orientation . 464 . 258 .161 . 076 [.050, .978] .49 7 . 264 .173 .06 4 [-.030, 1.023] impersonal orientation .109 . 215 .040 . 613 [.539, .320] .127 . 219 .046 . 566 [-.564, .311] competence [t1] . 573 . 075 .652 . 000 [ .424, .723] . 567 . 076 .645 . 000 [ .416, .719] reward .152 . 218 .056 . 486 [-.586, .282] task choice . 139 . 204 .051 . 497 [-.268, .547] r2 .607 .613 δ r2 .006 note. insignificant interaction effects were removed from the final models. gender, age, and number of solved puzzles at time 2 were included as control variables, but results are omitted for brevity. δ r2 shows the change in r2 when adding interaction effects into the model. jiang, kusamoto, & tanaka 94 | f l r table a-2 multiple regression results concerning time spent on puzzles at time 2. time spent on puzzles [t2] step 1 step 2 b s e β p 95% ci of b b s e β p 95% ci of b causality orientations autonomy orientation 11.360 2 5.948 .041 . 663 [-63.086, 40.366] 6.178 2 5.074 .022 . 806 [-56.187, 43.831] controlled orientation 8.060 2 6.540 .028 . 762 [-60.967, 44.847] -12.986 2 5.705 .045 . 615 [-64.254, 38.282] impersonal orientation 1 5.791 2 3.696 .057 . 507 [-31.447, 63.028] 1 6.672 2 2.705 .061 . 465 [-28.611, 61.954] time spent on puzzles [t1] . 898 .109 .738 . 000 [.681, 1.114] .9 17 .107 .754 . 000 [.704, 1.130] reward 3 4.274 2 3.210 .126 . 144 [-12.016, 80.564] task choice 52.562 2 1.249 .194 . 016 [-94.942, 10.183] r2 .526 .587 δ r2 .054 note. insignificant interaction effects were removed from the final models. gender, age, and number of solved puzzles at time 2 were included as control variables, but the results are omitted for brevity. δ r2 shows the change in r2 when adding interaction effects into the model jiang, kusamoto, & tanaka 95 | f l r table a-3 multiple regression results concerning task enjoyment at time 2. task enjoyment [t2] step 1 step 2 b s e β p 95% ci b s e β p 95% ci causality orientations autonomy orientation .272 . 147 .130 . 068 [-.020, .564] .263 . 150 .125 . 084 [-.036, .562] controlled orientation .154 . 154 .071 . 321 [-.460, .153] .138 . 158 .064 . 385 [-.452, .176] impersonal orientation .127 . 133 .062 . 343 [-.138, .393] .112 . 136 .054 . 415 [-.160, .383] task enjoyment [t1] .843 . 065 .836 . 000 [ .713, .973] .831 . 067 .824 . 000 [ .697, .964] reward .035 . 137 .017 . 801 [-.308, .239] task choice .138 . 129 .068 . 289 [-.396, .120] r2 .723 .727 δ r2 .005 note. insignificant interaction effects were removed from the final models. gender, age, and number of solved puzzles at time 2 were included as control variables, but the results are omitted for brevity. δ r2 shows the change in r2 when adding interaction effects into the model. frontline learning research vol. 13 no.1 (2025) 1 21 issn 2295-3159 corresponding author: christian g. k. hahn, institute of educational sciences, university of leipzig, marschnerstraße 31, 04109 leipzig, germany, christian.hahn@uni-leipzig.de doi: https://doi.org/10.14786/flr.v13i1.1225 language-dependent knowledge acquisition: mechanisms underlying language-switching costs in arithmetic fact learning christian g. k. hahn1, henrik saalbach1, clemens brunner2 & roland h. grabner2 1 institute of educational sciences, leipzig university,germany 2 institute of psychology, university of graz, austria article received 13 january 2023 / article revised 30 may 2024 / accepted 10 december 2024 / available online 29 january 2025 abstract within the research on bilingual learning, first studies have revealed that content learned in one language is retrieved more slowly when participants have to switch language from instruction to testing (i.e., language-switching costs, lsc). these costs are attributed to language-dependent knowledge representations. however, the cognitive mechanisms underlying lsc are still largely unknown. we investigated these mechanisms by using strategy as well as translation self-reports and by analysing oscillatory parameters in the electroencephalogram (eeg). thirty-six university students learned arithmetic facts of three different operations over four days either in english or in german. afterwards, they were tested in both languages with concurrent assessments of self-reports and electrophysiological activity. as expected, lsc in response latencies were observed in all arithmetic tasks. more importantly, analyses of self-reports and eeg revealed that both translation processes and calculation procedures contribute to lsc, with translation processes being the main cognitive mechanism underlying lsc. these results corroborate previous findings of languagedependent knowledge representations in arithmetic fact learning and shed new light on the cognitive mechanisms underlying lsc and possible educational consequences. keywords: bilingual learning; language-switching; arithmetic; self-reports; electroencephalograph mailto:christian.hahn@uni-leipzig.de hahn, saalbach, brunner & grabner 2 | f l r 1. introduction speaking a second language is advantageous for various reasons (baker, 2011). one common approach to foster second language learning is content and language integrated learning (clil). in clil, “a language other than the students’ mother tongue is used as a medium of instruction” (daltonpuffer, 2007, p. 1). nowadays, almost all european countries offer programs with non-language classes being taught in a foreign language (eacea, eurydice & eurostat, 2012). within the german school landscape, for example, clil tracks are often introduced in grades six or seven, in which one or two school subjects (e.g., such as geography) are taught in a foreign language (wolff, 2011). in this vein, educators hope to kill two birds with one stone: learning the subject content as well as a foreign language simultaneously. it is far from surprising that this concept of teaching is gaining more and more popularity, especially in a time when foreign language competencies are essential in the job market. it is an unresolved question, however, whether clil programs may negatively affect the learning of the subject content (baker, 2011; pérez-cañado, 2012). negative effects of clil may arise when the acquired knowledge is stored in the language of instruction (the second language) and, therefore, is not (as) easily accessible in another language (the mother tongue). in fact, there is substantial evidence suggesting that some types of knowledge are stored in a language-dependent way and that language-switching from instruction to retrieval produces performance impairments. these performance impairments are referred to as language-switching costs (lsc) and are typically reflected in longer response latencies or lower accuracy. lsc can be evaluated in experimental training studies in which participants first had to learn new information in one language (training phase), and afterwards, were required to retrieve or apply this knowledge in both the language of instruction and another language (test phase). the comparison of test performance in both languages reveals whether lsc emerge for certain types of knowledge. lsc have been found in different domains (for autobiographic knowledge see marian & neisser (2000); for non-numerical knowledge see marian & fausey, 2006) but have been most intensively studied in the field of arithmetic learning (spelke & tsivkin, 2001; dehaene, molko, cohen & wilson, 2004; venkatraman, siong, chee & ansari, 2006; grabner, saalbach & eckstein, 2012; saalbach, eckstein, andri, hobi & grabner, 2013; hahn, saalbach & grabner, 2017; volmer, grabner & saalbach, 2018). spelke and tsivkin (2001), for example, examined lsc in a russian-english bilingual sample for exact (e.g., “what is the sum of fifty-four and forty-eight?”) and approximate (e.g., “estimate the approximate cube root of twenty-nine!”) calculations. participants were trained on arithmetic equations with written number words either in russian or in english and were then tested with a verification task in both languages. while no lsc were found for approximate arithmetic, suggesting that this type of knowledge is language-independent, response latencies were significantly longer when the language of testing differed from the language of instruction in exact arithmetic. thus, this study provided strong evidence that numerical fact knowledge, which is relevant for exact calculation, is language-dependent. further studies corroborated this finding by showing lsc for arithmetic fact knowledge in different operations (multiplication and subtraction: grabner et al., 2012; exact base-7 addition: venkatraman et al., 2006; artificial facts: hahn et al., 2017) and for different language combinations (russian-english: spelke & tsivkin, 2001; italian-german: grabner et al., 2012; germanfrench: saalbach et al., 2012; german-english: hahn et al., 2017). in addition, it has been shown that these lsc emerge also in auditory stimuli (instead of written number words; hahn et al., 2017) and affect the application of fact knowledge in new and more complex task contexts (volmer et al., 2018). to date, research on lsc in bilingual learning settings has mainly focused on the appearance of lsc but not on the underlying cognitive mechanisms. understanding the mechanisms behind lsc is not only of interest to cognitive theories of language-dependent information processing and memory (e.g., gentner & goldin-meadow, 2003; malt & wolff, 2010), but also of practical relevance, since it might help to prevent lsc within clil. in the domain of arithmetic, there are at least two general possibilities about the underlying mechanisms of lsc. on the one hand, lsc may emerge due to the translation of the knowledge stored in the language of instruction into the language of retrieval or hahn, saalbach, brunner & grabner 3 | f l r application. for instance, when the arithmetic fact “13 x 8 = 104” is stored in english but needs to be applied in german, the fact could be first retrieved in english and then translated to german. on the other hand, they may result from additional calculation processes in the test language. in the example above, the performance impairment could result from the need to calculate (parts of) the arithmetic problem in german. since both general possibilities are compatible with the observed performance impairments during language switching, analyses of response latencies and solution rates are not informative regarding the underlying cognitive mechanisms. one approach to gain further insights into them is to use neurophysiological data as has been done in two functional magnetic resonance imaging (fmri) studies. venkatraman et al. (2006) trained 20 english-chinese bilinguals on base-7 additions (exact number task, e.g., “one-four add three-six”) and percentage value estimations (approximate number task, e.g., “forty-four percent of seventy”) over a period of five days. half of the participants were trained in chinese, half in english. during the fmri test session, participants had to perform the trained tasks in both languages. in contrast to spelke and tsivkin (2001), lsc in response latencies were found for both types of tasks. at the neurophysiological level, lsc were associated with stronger activation in taskdependent networks of brain regions. in the exact number task, additional activation occurred in language-related networks, suggesting that either the equation, the solution or both needed to be translated in order to retrieve the answer from memory. in the approximate number task, stronger activation was found in brain regions associated with magnitude processing and calculation, suggesting a greater effort for participants to solve problems in the untrained language. in the second fmri study on this topic, grabner et al. (2012) administered only an exact number task requiring the acquisition of arithmetic fact knowledge. twenty-nine german-italian bilinguals underwent a four-day training session of complex multiplication and subtraction problems. behavioural results again revealed lsc for response latencies. in contrast to venkatraman et al. (2006), the neurophysiological analyses showed increased activation during language switching in the brain regions associated with magnitude processing and calculation. therefore, it was argued that lsc might be due to additional numerical processing rather than language translation. in sum, both fmri studies found increased activation in the language-switching condition but were inconsistent regarding the involved brain networks. therefore, they do not draw a conclusive picture on the mechanisms behind lsc in arithmetic. furthermore, both studies used visual stimuli in the form of written number words, which do not represent an ecologically valid learning material (hahn et al., 2017). an alternative way to examine the underlying mechanisms of lsc are self-reports. self-reports have a long tradition in research on arithmetic and are typically used to assess the problem-solving strategy that is applied to solve a given arithmetic problem (e.g., lefevre, sadesky & bisanz, 1996; campbell & xue, 2001; imbo & vandierendonck, 2007; grabner & de smedt, 2011; vanbinst, ghesquiere & de smedt, 2012, cf. kirk & ashcraft, 2001; smith-chant & lefevre, 2003). strategy can be defined as a ”procedure or set of procedures for achieving a higher-level goal or task” (lemaire & reder, 1999, p. 365). in general, arithmetic problems can be solved either by procedural strategies such as counting (e.g. 8 + 2 = 8 + 1 + 1 = 10) or transformation (e.g. 6 x 12 = 6 x 10 + 6 x 2 = 72), or by direct retrieval of the stored solution from memory (e.g., 6 x 7 = 42). retrieval strategies are common in single-digit multiplications, which were often memorized by rote in school and for which an arithmetic fact network was built up over several years (e.g., imbo & vandierendonck, 2007; grabner & de smedt, 2011). but also, after repeated practice of problems typically solved through procedures (such as two-digit subtraction problems), solutions to these problems can be stored in declarative memory (e.g., grabner & de smedt, 2012; hahn et al., 2019). procedural strategies, in contrast, are used whenever the solution cannot be retrieved because of problem size (e.g., in two-digit multiplications) or operation (e.g., subtraction facts are typically not stored in a fact network; ischebeck, zamarian, siedentopf, koppelstätter, benke, felber & delazer (2006). thus, by means of trial-by-trial strategy self-reports it can be examined whether more procedural (calculation) processes take place when language-switching is required compared to when not. in addition to the problem-solving strategy, participants could report whether translation processes were involved in problem-solving. interestingly, venkatraman et al. (2006) reported that about two-thirds of the participants mentioned to have thought hahn, saalbach, brunner & grabner 4 | f l r occasionally in the language of training while performing tasks in the language-switching condition. unfortunately, there was no systematic acquisition of these comments. translation self-reports could directly address the question of whether lsc are due to translation processes. however, to the best of our knowledge, such translation self-reports have not been used in previous research on lsc. finally, insights into the mechanisms underlying lsc can also be obtained by manipulating the fact learning task. a limitation of most previous studies on lsc in arithmetic lies in the requirement to learn new facts through practicing more complex arithmetic problems, such as two-digit times one-digit multiplications (e.g., grabner et al., 2012; volmer et al., 2018; hahn et al., 2017). even though these facts can be expected to be retrieved from memory after multiple days of training, lsc can still derive either from translation or from (additional) calculation processes. therefore, hahn et al. (2017) introduced a “pure” fact learning condition consisting of artificial arithmetic facts. specifically, they required participants to learn artificial facts (e.g. 17 box 2 = 93) in addition to complex multiplication (e.g. 16 x 4 = 64) and subtraction (e.g. 52 – 9 = 43) facts. since the solutions to the artificial problems cannot be calculated but have to be memorized by rote, lsc can be traced back only to translation processes. thus, comparing the size of lsc across pure and typical fact learning, the impact of translation processes can be estimated. in hahn et al. (2017), lsc were found in all three operations, and there was no difference in their extent between artificial problems and multiplication problems, suggesting a similar mechanism in these two operations. 1.1 the present study the main aim of the present study is to provide further insights into the mechanisms underlying lsc in arithmetic fact learning. to this end, we administered an experimental training design with artificial, multiplication, and subtraction problems using auditory stimuli, similar to hahn et. al. (2017). extending previous studies, participants had to provide two kinds of trial-by-trial self-reports, one on the problem-solving strategy and one on the use of translation processes. moreover, we included electroencephalography (eeg) to complement the information from the strategy self-reports. here we focus on oscillatory eeg activity in the theta band (event-related synchronization, ers), which has turned out to be sensitive in distinguishing arithmetic procedures and fact retrieval (e.g., grabner & de smedt, 2011, 2012; tschentscher & hauk, 2016). in particular, the application of procedures has turned out to be accompanied by lower theta ers than the application of fact retrieval. thus, if calculation procedures contribute to lsc, this should be reflected in a difference in theta eeg activity between the switching and the no-switching condition. we predicted to find longer response latencies for fact learning when the language of training differs from the language of testing, independently of the arithmetic task, while no lsc were expected for accuracy rates (hypothesis 1). with respect to the cognitive mechanism underlying lsc, two hypotheses were tested: in case that lsc are caused by additional translation processes, a higher frequency of self-reported translated trials for problems in the switching condition compared the noswitching condition should emerge across all three arithmetic tasks (hypothesis 2a). alternatively, if lsc are caused by additional calculation procedures, in multiplication and subtraction problems a higher frequency of self-reported procedural strategy use in the switching compared to the no-switching condition can be expected (hypothesis 2b). accordingly, if calculation procedures underlie lsc in multiplication and subtraction, lower theta ers can be expected in the switching compared to the noswitching condition (hypothesis 3). finally, we explored correlations between individual differences in lsc and the control variables we assessed, i.e., l2 vocabulary knowledge, intelligence, and arithmetic competencies. hahn, saalbach, brunner & grabner 5 | f l r 2. methods 2.1 participants the study included 47 right-handed adult students. eleven participants had to be excluded from analysis: four participants due to missing one training session, three due to technical incidents during the test session, and four due to strong eeg artifacts throughout the test session. the final sample consisted of 36 participants, aged between 20 and 28 years (m = 23.0, sd = 2.1). participants were randomly assigned to either a german (l1) or english (l2) training group. all participants studied english linguistics, had german as mother-tongue, and received their previous math education in german. they gave written informed consent and were paid for their participation. the study was approved by the local ethics committee. 2.2 experimental stimuli the study included 18 problems: 6 artificial, 6 multiplication, and 6 subtraction problems. artificial problems were two-digit, and one-digit numbers connected via an arbitrary symbol (“box”) and a two-digit solution (00 box 0 = 00). these solutions were different from results of any existing arithmetic operation and needed to be memorized by rote. multiplication problems were two-digit times one-digit problems with two-digit solutions (00 x 0 = 00). subtraction problems were two-digit minus two-digit problems with two-digit solutions (00 – 00 = 00). in all sessions, the problems were presented auditorily to the participants via a loudspeaker, designed with the text-to-speech software of voice reader studio 15 (linguatec, 2015). all stimuli (i.e., the whole equation) had the same length (i.e., 1850 milliseconds). 2.3 additional measures 2.3.1 english vocabulary knowledge participants’ english (l2) vocabulary knowledge was assessed by administering the online version of the lextale. the lextale has been developed to account for the increasing need in experimental studies to assess vocabulary knowledge of english as a second language within a short time scale (lemhöfer & broersma, 2012). in this test, participants have to indicate whether presented words are existing english words or not. such yes/no tests have been found to be valid measures of l2 vocabulary knowledge (mochida & harrington, 2006). lemhöfer and broersma (2012) were further able to show that the lextale was a better predictor than commonly used self-ratings for vocabulary knowledge, and lextale scores have a substantial correlation with common measures of general english proficiency (i.e., quick placement test (qpt; syndicate, u.c.l.e. (2001)) and test of english for international communication (toeic; schmitt, 2005)). the lextale consists of 60 items (40 words, 20 non-words). non-words are orthographically correct and pronounceable but represent strings without meaning. furthermore, we added a second short test for vocabulary knowledge, the dialang (huhta, luoma, oscarson, sajavaara, takala & teasdale, 2002). similarly, the dialang placement test includes 75 words that need to be marked as existing or non-existing in the english language. in contrast to the lextale, answers can be corrected once marked, because all words appear on the same screen. scores for both tests were averaged to create the final score for l2 vocabulary knowledge. the two tests were strongly correlated (r = .80; p < .001). 2.3.1 arithmetic fluency since the present study was conducted in the field of arithmetic, all participants were tested on their arithmetic fluency using the french kit (french, ekstrom & price, 1963). in this paper-and-pencil test, participants have to solve as many arithmetic problems as possible within a given time period. for each page, the time limit was two minutes. all subtests consist of two pages. the first subtest contains 60 three-term addition problems with multi-digit addends (e.g., 50 + 42 + 15 = ...), the second subtest 60 multi-digit division problems per page (e.g., 56 : 8 = ...), the third subtest six alternating rows of 10 multi-digit subtraction and multiplication problems per page (e.g. 42 – 17 = ..., and 62 x 6 = ...), and the hahn, saalbach, brunner & grabner 6 | f l r fourth subtest 60 multi-digit addition and subtraction problems with a suggested answer (e.g., 22 + 29 = 41) that had to be verified. the final score for arithmetic fluency is calculated as the total number of correctly solved problems. 2.3.2 general intelligence participants’ intelligence profiles were assessed by using the short version of the berlin intelligence structure test (bis-4; jäger, süß & beauducel, 1997). this test includes 15 tasks drawing on three content components of intelligence (numerical, figural, and verbal) and four operational abilities (processing speed, memory, reasoning, and creativity). the overall duration of the test is 45 minutes. the raw scores of the individual tests are aggregated to an iq score for general intelligence. 2.4 procedure the study consisted of five sessions on consecutive days: four training sessions and one test session. all sessions took place at the department of psychology of the university of göttingen, germany. training session 1 and the test session were administered in an eeg lab, while training sessions 2, 3, and 4 took place in a computer lab. during the four-day training, participants had to learn the 18 arithmetic problems either in german (l1) or in english (l2). in training session 1 as well as the test session, participants’ brain activity was recorded by means of eeg, and the applied strategies were assessed with self-reports as described below. 2.4.1 training session 1 training session 1 started with the instruction of the training program as well as an introduction to eeg recording. for later artifact removal (see below), we recorded the eeg during three minutes of eye movements, in which participants were instructed (via visual cues on the display) to roll their eyes, blink, move them up or down, or just keep their eyes open or closed. then the experimental task was presented in three blocks. within each block, there was only one type of task (i.e., mul, sub, art), with each of the 6 problems presented six times (not in succession). the order of the blocks was counterbalanced over the sample and all four training sessions. as depicted in figure 1, each trial started with a fixation point for two seconds. then, the problem was presented auditorily via loudspeakers either in english or in german, depending on the training group. participants had to orally give the answer to the problem as fast as possible in the instructed language. the response time was collected by a voice key. timeout was set to 8.15 seconds after stimulus presentation (i.e. 10 seconds minus 1.85 seconds stimulus length). the examiner – seated outside the eeg cabin – typed in the given answer, after which the participants received visual feedback on the screen (i.e., a red screen for an incorrect answer and a green screen for a correct answer), followed by the correct answer presented again via the loudspeaker. the next slide asked for the strategy the participant had used to answer the problem (strategy report). using a button response box, participants indicated whether they used (a) fact retrieval (e.g., knowing the answer from memory without any type of calculation), (b) a procedural strategy (e.g., calculating the answer), or (c) any other strategy (e.g., guessing the answer). these strategy reports have been used and validated to assess strategy use in arithmetic in several studies before (campbell & xue, 2011; grabner & de smedt, 2011; lefevre et al., 1996). the timeout for the report was set to five seconds. the next trial started after an inter-trial interval of two seconds. notably, since the participants could not know the solutions to the artificial problems in the first training session, all artificial problems were presented together with the solution twice. thereafter, participants had to solve these problems on their own, similar to multiplication and subtraction problems. the first session took between 30 and 40 minutes, depending on the individual speed of each participant. hahn, saalbach, brunner & grabner 7 | f l r figure 1. schematic time course of a) training session 1, b) training session 2, 3, and 4, and c) test session. r = reference interval; a = activation interval. 2.4.2 training sessions 2, 3, and 4 over the next three consecutive days, there were three additional training sessions to learn the 18 problems. each session had a duration between 25 and 35 minutes. in these sessions, no eeg was recorded. before the training, participants were again given instructions on how to proceed during the session. as with training session 1, the three task blocks were counterbalanced across the participants. the fixation point lasted for two seconds, and the problems were presented via headphones. furthermore, participants were instructed to press the enter key as soon as they had the answer in mind. this was used as an alternative measure of response time to the voice-key in sessions 1 and 5. afterwards, they were asked to enter the solution using a numerical keypad. then, participants received corrective visual feedback (correct or incorrect) followed by the correct solution presented auditorily. in these training sessions, no strategy reports were collected. after the four training sessions each problem had been repeated 24 times. this number of trials is in line with previous studies to make sure that participants had sufficiently learned the answer to each problem (hahn et al., 2017; grabner & de smedt, 2012). 2.4.2 eeg test session in the test session on day 5, the problems were presented in both languages, requiring languageswitching or not. after completing the eye-movement eeg as described before (training session 1), all differences to training session 1 were explained to the participants before starting with the test session. first, participants did not receive feedback to their responses. further, participants completed six blocks, including both english and german problems. within each block, the three operations and the two languages were randomly mixed. similar to training session 1, participants had to indicate immediately after giving the answer which strategy they used to answer the problem (strategy report). the timeout was five seconds. participants were then asked whether they translated any numbers during problem solving (i.e., by pressing either button 1 or 2 on a response box). we refer to this as translation report. the timeout was again set to five seconds. hahn, saalbach, brunner & grabner 8 | f l r 2.5 data analysis 2.5.1 behavioural data acquisition and analysis accuracies and response latencies for correctly solved trials were analysed with anovas. trials with voice-key errors in training session 1 and the test session were excluded from analyses. before the main analyses, we tested whether the two training groups (training in l1 vs. l2) differed in l2 vocabulary knowledge, intelligence, and arithmetic fluency, using t-tests for independent samples. for the training data, the anova included the two within-subject factors arithmetic task (artificial vs. multiplication vs. subtraction) and training day (day1 vs. day2 vs. day3 vs. day4). the testing data anova comprised the two within-subject factors arithmetic task and language switching (noswitching vs. switching). a potential impact of the training group (english vs. german) was analysed by means of t-tests for independent samples on the observed lsc. in case of violation of the sphericity assumption (mauchly’s test), degrees of freedom were corrected using greenhouse-geisser estimates of sphericity. all post-hoc tests were conducted using bonferroni adjusted alpha levels. for the analyses of strategy and translation reports, we conducted mixed repeated-measure anovas, including the within-subject factors arithmetic task and language switching. for these analyses, the distributions of strategy and translation reports were calculated for correctly solved trials, i.e. frequencies for the three strategies (retrieval vs. procedure vs. other) and the two options for the translation report (no vs. yes). effect sizes are presented as cohen’s d or partial eta-squared (ηp²). 2.5.2 eeg data acquisition and analysis eeg was recorded from 64 scalp electrodes with a biosemi activetwo system (biosemi, amsterdam, the netherlands). three additional electrodes recorded ocular activity (electrooculogram, eog); two placed horizontally at the outer canthi of both eyes, and the third above the nasion between the inner canthi of both eyes. both eeg and eog signals were sampled at 256 hz. eeg data analysis focused on oscillatory brain activity and was conducted using mne python (gramfort, luessi, larson, engemann, strohmeier, brodbeck, goj, jas, brooks, parkkonen & hämäläinen, 2013) as well as custom python scripts. after manually marking artifact segments (segments with excessive muscle activity) and removing bad channels (channels with excessive amount of noise or channels with flat power spectral density), we re-referenced the data to the average of all remaining channels. next, we removed ocular activity using a regression-based approach with coefficients calculated from the eye movement eeg session (gratton, coles & donchin, 1983). using the clean data, we computed band power in the theta band (4–7 hz) for each epoch (that is, we filtered the continuous data with a fir filter with suitable filter characteristics and squared the resulting values). similar to previous studies (e.g., de smedt , grabner & studer, 2009; grabner & de smedt, 2011), we quantified task-related changes in theta eeg activity by computing event-related synchronization (ers), i.e., the percentage increase in theta power during task processing (an activation period) compared to a pre-stimulus reference period. specifically, within each epoch, we computed the median theta band power within the reference interval (r) between –2.75 seconds to –0.25 seconds before stimulus onset and the median within the activation interval (a) between 1.85 seconds (end of stimulus presentation) until voice onset. then, based on the median across epochs for both r and a, we computed theta ers using the formula: ers (%) = (a / r – 1) · 100% (pfurtscheller & lopes da silva, 1999). for statistical analyses, we averaged all channels per hemisphere and computed the logarithm to make the data more normal. we performed a repeated-measures anova with factors arithmetic task, language switching, and hemisphere (left vs. right). hahn, saalbach, brunner & grabner 9 | f l r 3. results table 1 summarizes the individual characteristics of the participants, separately for the two training groups. there were no significant differences between the german and the english training group in vocabulary knowledge of l2, general intelligence, or arithmetic fluency. table 1 mean scores (standard errors) for the german and english training group (n=18 for each group) measure german training (l1) english training (l2) p vocabulary knowledge l2 (%) 80.4 (3.1) 85.9 (2.2) .16 general intelligence (iq) 94.8 (2.0) 97.1 (1.6) .37 arithmetic fluency (raw score) 128.0 (9.5) 128.8 (5.5) .94 3.1 training data training data for response latencies and accuracies are displayed in figure 2. in both measures, performance improved significantly over the training. for rt, there was a strong main effect of training day (f(2.14, 74.84) = 97.68, p < .001, ηp² = .74), with significant decreases for each consecutive day (all ps < .001). in addition, there was a main effect of arithmetic task (f(2, 70) = 31.35, p < .001, ηp² = .47). artificial problems were solved faster than multiplication problems (1713 ms vs. 2065 ms; t(35) = -4.20, p < .001, d = 0.48), but more slowly than subtraction problems (1713 ms vs. 1357 ms; t(35) = 3.87, p < .001, d = 0.61), and multiplication more slowly than subtraction problems (2065 ms vs. 1357 ms; t(35) = 7.68, p < .001, d = 1.08). there was an interaction between training day and arithmetic task (f(2.92, 102.16) = 22.19, p < .001, ηp² = .39), revealing strong training effects for multiplication problems in the first two trainings sessions, in contrast to substantially smaller training effects for artificial and subtraction problems. for accuracies, there was a main effect of training day (f(1.72, 60.12) = 119.49, p < .001, ηp² = .77), with significant increases for each consecutive day (all ps < .001). further, there was a main effect of arithmetic task (f(1.34, 46.92) = 37.28, p < .001, ηp² = .52). artificial problems were solved less accurately than multiplications (80.1% vs. 90.1%; t(35) = -5.19, p < .001, d = 1.03) as well as subtractions (80.1% vs. 95.0%; t(35) = -7.12, p < .001, d = 1.56), while multiplications were solved less accurately than subtractions (90.1% vs. 95.0; t(35) = -4.50, p < .001, d = 0.96). further, there was a significant interaction between training day and arithmetic task. (f(2.60, 91.03) = 32.16, p < .001, ηp² = .48), attributable to the strong training effects for artificial problems, in contrast to substantially smaller training effects for subtractions and multiplications, having already high accuracy on day 1. hahn, saalbach, brunner & grabner 10 | f l r a) b) figure 2. training data for reaction time (a) and accuracy (b). error bars indicate the standard error (se). separate lines represent the three different tasks. art = artificial problems, mul = multiplication problems, sub = subtraction problems 3.2 test session: performance data 3.2.1 language switching costs descriptive statistics of accuracies and response latencies in the three arithmetic tasks and two switching conditions are shown in table 2. 0 1000 2000 3000 4000 day 1 day 2 day 3 day 4 r es po ns e tim e (m s) training art mul sub 0 25 50 75 100 day 1 day 2 day 3 day 4 a cc ur ac y (% ) training art mul sub hahn, saalbach, brunner & grabner 11 | f l r table 2 mean response latency (for correctly solved trials) in milliseconds (top rows) and accuracy rates in percentage correct (bottom rows) as a function of arithmetic task and switching condition. standard errors are given in parentheses. lsc were only observed for response latencies. artificial multiplication subtraction response latency in milliseconds no language switching 1492 (82) 1493 (100) 1203 (80) language switching 1613 (84) 1638 (97) 1352 (97) difference 121 145 149 accuracy in percentage correct no language switching 94.0 (1.8) 93.2 (0.9) 96.3 (0.8) language switching 91.9 (2.1) 93.3 (0.9) 95.5 (0.9) difference -2,1 -0,1 0,8 hypothesis 1: we predicted to find longer response latencies for fact learning when the language of training differs from the language of testing, independently of the arithmetic task, while no lsc were expected for accuracy rates. in line with hypothesis 1, there was a strong main effect of language switching across trials at test on response latencies (f(1, 35) = 22.93, p < .001, ηp² = .40), showing that problems in the noswitching condition were solved faster (1396 ms) than problems in the switching condition (1534 ms). in addition, there was a significant main effect of arithmetic task (f(1.59, 55.62) = 8.19, p = .002, ηp² = .19). post-hoc analyses revealed that subtraction problems (1278 ms) were solved faster than artificial (1552 ms) and multiplication problems (1565 ms; ps < .006). all other effects were not significant (all ps > .85). an additional t-test revealed that the two training groups (english vs. german) differed in the lsc in response latencies (t(34) = .6.07, p = .019; means of 127 vs. 151 ms, respectively). for accuracy, there was no main effect of language switching (f(1, 35) = 2.64, p = .11, ηp² = .07), and none of the other effects was significant (all ps > .17). 3.3 test session: strategy and translation reports figure 3 displays the distribution of self-reported (a) translation use and (b) procedural strategy use across operations. since the frequency of trials within the strategy category “other” was very low (< 2.5%), these trials were excluded from further analyses. hahn, saalbach, brunner & grabner 12 | f l r a) b) figure 3. distribution of self-reports during the test session for a) translation processes and b) procedural strategies. error bars indicate the standard error (se). hypothesis 2a: in case that lsc are caused by additional translation processes, a higher frequency of self-reported translated trials for problems in the switching condition compared the no-switching condition should emerge across all three arithmetic tasks. in line with the hypothesis 2a, the repeated measures anova on translation reports showed a main effect of language switching (f(1, 35) = 68.52, p < .001, ηp² = .66), indicating that the frequency of translation use was higher in the switching condition (46.45%) compared to the no-switching condition (4.24%). further, there was a main effect of arithmetic task (f(2, 70) = 7.04, p = .002, ηp² = .17). post-hoc pairwise comparisons revealed that the frequency of translation use was higher for artificial (27.43%) compared to subtraction problems (21.96%), as well as higher for multiplication (26.66%) compared to subtraction problems (ps < .02). finally, there was an interaction of arithmetic task and language switching (f(2, 70) = 5.19, p = .008, ηp² = .13). post-hoc t-tests showed that the frequency of using translation during switching was lower for subtraction (40.30%) compared to artificial (50.66%; p = .005, d = .33) and multiplication problems (48.39%; p = .002, d = .25). as validation of the translation reports, we conducted an additional analysis of the response latencies. 0 10 20 30 40 50 60 no switching switching u se o f t ra ns la tio na l ( % ) condition art mul sub 0 5 10 15 20 25 30 no switching switchingu se o f p ro ce du ra l s tra te gi es (% ) condition art mul sub hahn, saalbach, brunner & grabner 13 | f l r overall, response latencies in trials without reported translation were significantly shorter (1557 ms) than in trials with reported translation (2029 ms; t(31) = --6.66, p < .001, d = .78)1. hypothesis 2b: alternatively, if lsc are caused by additional calculation procedures, in multiplication and subtraction problems a higher frequency of self-reported procedural strategy use in the switching compared to the no-switching condition can be expected. the repeated measures anova on strategy reports revealed a main effect of language switching (f(1, 35) = 14.38, p < .001, ηp² = .29), indicating that the frequency for procedural strategy use was higher in the switching condition (12.19%) compared to the no switching condition (9.57%). further, there was a main effect of arithmetic task (f(2, 70) = 19.52, p < .001, ηp² = .36). post-hoc pairwise comparison revealed a higher frequency of procedural strategy use for multiplication (14.15%) and subtraction (18.49%) compared to artificial problems (0%; ps < .001). no other effects were significant (ps > .10). as validation of the strategy reports, we conducted another additional analysis of response latencies. this revealed that trials in which retrieval strategies were reported were solved significantly faster compared to procedural strategies (1445 ms vs. 2308 ms; t(22) = -5.52, p < .001, d = 1.24)2. 3.3.1 relative importance of translation and procedural strategies for lsc since switching effects were found to be associated with both self-reported translation and procedural strategy use, we conducted an additional analysis to evaluate their relative importance for lsc. specifically, we conducted a multiple regression analysis in which we used individual differences in lsc as dependent variable and switching scores of translation reports (percentage of translation: switching – no-switching) and strategy reports (percentage of procedures: switching – no-switching) as independent variables. for response latencies, the regression model explained 22.4% of the variance in lsc (r² = .22, f(2, 33) = 4.77, p = .02). the translation report score was a significant predictor (ß = .48, p = .005), whereas the strategy report score was unrelated to lsc (ß = -.01, p = .93). hence, the more participants used translation processes in the switching (compared to the no-switching) condition, the higher were the lsc. in contrast, despite the fact that participants used significantly more procedural strategies during the switching condition and procedural strategies had significantly longer response latencies, this factor did not predict lsc regarding response latencies. the same analysis was conducted for accuracies. the regression model showed no explanatory value for the prediction of lsc (r² = .10, f(2, 33) = 1.90, p = .17). 3.3.2 patterns of strategy and translation reports to provide a more fine-grained picture of the use of strategy and translation processes for lsc, figure 4 displays a descriptive overview of the different self-report combinations for the three operations. 14 participants have been excluded from analyses, solving >10 trials per condition 213 participants have been excluded from analyses, solving >10 trials per condition hahn, saalbach, brunner & grabner 14 | f l r figure 4. descriptive overview of the different self-report combinations for the three operations. in the no-switching condition on the left, there is a high rate of self-reported retrieval, summing up to 100 % for artificial problems and to around 85-90 % for multiplication and subtraction problems. virtually all problems were reported to be solved without any translation. in the switching condition, the rate of retrieval remains the same for artificial problems and is slightly decreased for multiplication and subtraction problems (see also figure 3). in the latter problems, procedures occur both with and without self-reported translations. the first case may reflect that the problem itself is translated into the trained language, calculated there, and then the solution is translated back into the untrained language. the second case may indicate that a procedure is applied without any translation, i.e., in the untrained language. within the self-reported retrieval and translated trials, it is likely that the solution has been retrieved in the trained language and then translated into the untrained language. here, a slightly higher percentage was observed for multiplication compared to subtraction problems. unfortunately, no inference statistics for performance data can be calculated for the different self-report combinations as there are too few participants (< 10) with at least 10 correctly solved trials for each strategy combination. 3.4 exploratory analyses of individual differences finally, we examined relations between lsc in response latencies and individual differences in the assessed control variables. the results are shown in table 3. in none of the control variables, a significant association with lsc was observed. 0 10 20 30 40 50 60 70 80 90 100 retrieval procedure retrieval procedure retrieval procedure retrieval procedure translation no translation translation no translation no switching switching pe rc en ta ge o f c or re ct tr ia ls art mul sub hahn, saalbach, brunner & grabner 15 | f l r table 3 descriptive statistics and intercorrelations between language switching costs (lsc) and vocabulary knowledge in english (l2), general intelligence (iq) as well as french kit (math fluency) variable n m sd r 1. lsc 36 110 197 2. vocabulary knowledge 36 83 11 -.20 3. general intelligence (iq) 35 97 12 -.16 4. french kit 36 128 32 -.04 3.5 test session: eeg data hypothesis 3: if calculation procedures underlie lsc in multiplication and subtraction, lower eeg theta ers can be expected in the switching compared to the no-switching condition. table 4 lists mean theta ers values for all combinations of conditions and arithmetic tasks. the repeated-measures anova revealed a significant main effect of language switching (f(1, 35) = 6.86, p < .01, ηp² = .16). as expected, the switching condition was associated with a significantly lower theta ers than the no-switching condition (18.9% vs. 16.1%; d = 0.10). furthermore, the interaction between language switching and hemisphere was significant (f = 4.43, p = .043, ηp² = .11. whereas ers was about the same for the two hemispheres in the no-switching condition (left: 18.7%, right: 19.2%), the left hemisphere showed lower ers (15.0%) compared to the right hemisphere (17.2%) in the switching condition. however, no interaction between arithmetic task and language switching emerged (f = 1.13, p = .32, ηp² = .03). table 4 reveals that the switching effect (in terms of cohen’s d) is descriptively largest in the subtraction, followed by the multiplication and the artificial condition, the latter with a close to zero effect size. table 4 mean ± standard error of theta ers/erd (in %) as well as cohen’s d (the effect size of the difference between switching and no switching) for all combinations of conditions and operations artificial multiplication subtraction switching 17.8% ± 1.3% 16.5% ± 1.2% 14.0% ± 1.2% no switching 18.5% ± 1.2% 19.2% ± 1.2% 19.1% ± 1.3% d 0.56 2.25 4.25 4. discussion the aim of the present study was to provide further insights into the mechanisms underlying lsc in arithmetic fact learning by using trial-by-trial self-reports that were complemented by eeg data. bilingual adult students were trained on four consecutive days to learn 18 problems of three different operations (artificial problems, multiplications, subtractions) in either german (l1) or english (l2). on the fifth day, all participants were tested on the arithmetic problems in both languages. hahn, saalbach, brunner & grabner 16 | f l r we found clear-cut lsc across all three operations for response latencies, thus confirming our first hypothesis. specifically, participants required more time to solve the problems of all three operations in the language-switching condition than in the no-switching condition. this finding replicates the results of previous studies on arithmetic fact learning in a more “natural” task context as auditory stimuli presentation was combined with a voice key for oral responses. previous research either collected data using visual stimuli and keyboard responses (e.g., grabner et al. 2012; saalbach et al., 2013) or auditory stimuli and keyboard responses (hahn et al., 2017). no lsc were observed for accuracy rates. this is in line with previous findings revealing that lsc do not emerge in accuracy when participants are given sufficient time to respond (hahn et al., 2017). in the present study, participants had an even more generous time frame to answer in each trial (i.e., 13 seconds with an average response latency < 2 seconds), which may have led to a ceiling effect in accuracy. the present study was also novel in that it is the first in which self-reports were used to uncover the cognitive mechanisms underlying lsc. in line with our expectations, participants not only indicated to use more translation processes in the language-switching (compared to the no-switching) condition (hypothesis 2a), but also reported to having applied more procedural strategies (hypothesis 2b). thus, both hypotheses were confirmed. even though the latter finding suggests that lsc can be explained by additional numerical processing, in particular calculation, as suggested by grabner et al. (2012), it needs to be emphasized that only about 12% of the trials in the language-switching condition had been reported to be solved through procedural strategies. in addition, lsc were found for artificial problems, which can only be retrieved from memory to the same extent as for multiplication and subtraction. therefore, it is unlikely that procedural strategies alone can account for the overall lsc found in our sample. rather, our findings suggest that translation processes play a major role in lsc (venkatraman et al., 2006). approximately 46% of the trials in the language-switching condition were reported as translation trials. these trials also showed significantly longer response latencies than its counterpart (i.e., no translation). further, the multiple regression analysis including both types of self-reports revealed that only the amount of translation trials is a significant predictor for overall lsc in response latencies. lsc for artificial problems were assumed to be only due to translation processes. however, the analyses of the translation reports revealed that about 50% of the artificial trials in the languageswitching condition were indicated not to include translation processes. there are at least three explanations for this finding. first, when considering that all problems were presented six times in the test session, participants might have had a training in the switching condition during the test session itself. in other words, at some point during the test session (e.g., after solving an arithmetic problem two or three times in the switching condition) participants have acquired the answer to a problem in the previously untrained language and did not require translation any longer. however, in our sample, the use of translation strategies solving items in the untrained language was mixed from the beginning. we analysed the percentage of translation reports across blocks and indeed observed a decrease (block 1+2: 59 %, block 3+4: 53 %, block 5+6: 41 %). still, if the above-mentioned explanation was true, the translation percentage in blocks 1+2 should be much higher. second, it is certainly possible that there are participants who trained problem equations in their native language after the end of a training session or before a new training session the next day, regardless of the fact that the actual training language was english. in these cases, a bond between problem equations and training language would be distorted. this possibility seems likely, since a large number of participants articulated their ambition and gave the impression of being upset about having a low solution rate in the first training sessions. finally, the validity of the translation reports may be limited. in contrast to the problem-solving strategy reports in arithmetic, which are already well-established and validated (campbell & xue, 2011; grabner & de smedt, 2011; lefevre et al., 1996), the present study is the first in which trial-by-trial translation reports were required in the domain of arithmetic. even though translation trials were associated with longer response latencies, it remains elusive to what extent participants accurately reported the occurrence of these processes. in the test session, in which participants had to constantly switch between languages and the three operations, it appears likely that some participants might have had a hard time reliably indicating for each trial what exactly had taken place. in spite of these potential pitfalls, the self-report data provides evidence for translation processes playing a key role in the appearance of lsc. hahn, saalbach, brunner & grabner 17 | f l r in addition to self-reports, we collected eeg data to test the link between lsc and calculation processes at the neurophysiological level. based on the sensitivity of eeg theta activity to arithmetic problem-solving strategies (e.g., grabner & de smedt, 2011, 2012; tschentscher & hauk, 2016), we hypothesized to find lower theta ers in the switching (compared to the no-switching) condition because we assume switching to be accompanied by the stronger application of calculation procedures. in line with this assumption, we observed an effect of language switching consisting of lower theta ers when switching was required. this finding corroborates the results from the strategy reports suggesting that lsc are partly due to additional calculation processes. a closer look at the switching effect sizes, however, revealed that the effects were generally small and only slightly differed between the operations. the largest (but still small) effect of d = 0.18 was observed for subtraction. as expected, the effect size for the artificial numerical facts was practically zero (d = 0.03). since this field of research is still in its infancy, it is currently difficult to derive clear implications for practice. the present study can only reflect the actual classroom situation to a very limited extent, since it is a laboratory study focusing on only a fraction of actual school content (i.e., arithmetic fact knowledge). up to this point, the majority of studies have focused on mathematical knowledge and, likewise within those studies, primarily on simple learning demands (i.e., factual knowledge). to better match real-life learning context, a next step towards procedural and conceptual knowledge would be highly desirable. moreover, there has been insufficient discussion about the extent to which lsc are of temporal persistence. thus, we might ask what happens when the instructional language changes during learning phases. the very few studies that have been conducted to evaluate content knowledge acquisition in clil instruction suggest that clil students perform more poorly (lo & lo, 2014; piesche et al., 2016) or need to spend more time to meet the learning gains of non-clil students (dallinger, jonkmann, hollm & fiege, 2016). in the study of piesche et al. (2016), for instance, it was shown in six-graders that monolingually educated groups outperformed bilingually educated groups regarding learning gains directly after an intervention (i.e. five 90min-lessons on “floating and sinking”) as well as at follow-up six weeks later (small effect-sizes). such evidence raises the question whether basic concepts (e.g. “floating and sinking”) or basic arithmetic shall be learned in the language in which the knowledge will be applied. we might not kill two birds with one stone but even perhaps create little performance gaps we do not see yet, when time efficiency stays the primary concern, with quality of content falling by the wayside. this concern might be especially true considering elementary knowledge which builds the foundation for later study. learning content and foreign language together may put unnecessary load on the working memory (sweller, ayres & kalyuga, 2011). yet, a recent study evaluating a long-term immersion program failed to find lsc when language of instruction and language of testing differed (fleckenstein, gebauer & möller, 2019). however, this may also be the result of selection effects having more intelligent students in immersion classes compared to conventional classes. last, there has not been ample research focusing on individual characteristics. overall, the research in the area of knowledge acquisition and the understanding on how subject matter content and language of acquisition interact remains important and at its beginning. thus, there remains a great need for more experimental and ecologically valid research on the topic. to conclude, in the present study lsc were observed for multiplication, subtraction, and a pure fact learning task using auditory stimuli and an oral response task. by analysing self-reports (i.e., strategy and translation reports), we were able to shed new light on the question of why lsc in arithmetic learning appear. the evidence suggests that translation processes play a key role for lsc in fact knowledge and that this type of knowledge indeed is strongly tied to the language of acquisition. in addition to translation processes, lsc may at least partly be due to the stronger use of calculation procedures, which was observed in both strategy reports and eeg data. thus, self-reports appear to be a promising way to further elucidate the cognitive mechanisms underlying performance decrements in educational settings in which instruction is provided in a different language than the mother tongue. hahn, saalbach, brunner & grabner 18 | f l r keypoints it explores learning and instruction in the context of bilingual learning that is of high current societal relevance it introduces the methodology of different self-reports into research on language-switching costs behavioural and neurophysiological methods are applied to investigate cognitive mechanisms of a well-established but poorly understood phenomenon references baker, c. (2011). foundations of bilingual education and bilingualism. bristol, uk: multilingual matters. biosemi, amsterdam, the netherlands. https://www.biosemi.com/ campbell, j. i., & xue, q. (2001). cognitive arithmetic across cultures. journal of experimental psychology: general, 130(2), 299-315. https://doi.org/10.1037/0096-3445.130.2.299 dallinger, s., jonkmann, k., hollm, j., & fiege, c. (2016). the effect of content and language integrated learning on students' english and history competences–killing two birds with one stone?. learning and instruction, 41, 23-31. https://doi.org/10.1016/j.learninstruc.2015.09.003 dalton-puffer, c. (2007). discourse in content and language integrated learning (clil) classrooms (vol. 20). john benjamins publishing. dehaene, s., molko, n., cohen, l., & wilson, a. j. (2004). arithmetic and the brain. current opinion in neurobiology, 14(2), 218-224. https://doi.org/10.1016/j.conb.2004.03.008 de smedt, b., grabner, r. h., & studer, b. (2009). oscillatory eeg correlates of arithmetic strategy use in addition and subtraction. experimental brain research, 195(4), 635-642. https://doi.org/10.1007/s00221-009-1839-9 eacea, eurydice, & eurostat (2012). key data on teaching languages at school in europe. brussels: eurydice. fleckenstein, j., gebauer, s. k., & möller, j. (2019). promoting mathematics achievement in one-way immersion: performance development over four years of elementary school. contemporary educational psychology, 56, 228-235. https://doi.org/10.1016/j.cedpsych.2019.01.010 french, j. w., ekstrom, r. b., & price, l. a. (1963). manual for kit of reference tests for cognitive factors (revised 1963). educational testing service princeton nj. gentner, d., & goldin-meadow, s. (eds.). (2003). language in mind: advances in the study of language and thought. mit press. https://doi.org/10.7551/mitpress/4117.001.0001 grabner, r. h., & de smedt, b. (2011). neurophysiological evidence for the validity of verbal strategy reports in mental arithmetic. biological psychology, 87(1), 128-136. https://doi.org/10.1016/j.biopsycho.2011.02.019 grabner, r. h., & de smedt, b. (2012). oscillatory eeg correlates of arithmetic strategies: a training study. frontiers in psychology, 3, 428. https://doi.org/10.3389/fpsyg.2012.00428 https://doi.org/10.1037/0096-3445.130.2.299 https://doi.org/10.1016/j.learninstruc.2015.09.003 https://doi.org/10.1016/j.conb.2004.03.008 https://doi.org/10.1007/s00221-009-1839-9 https://doi.org/10.1016/j.cedpsych.2019.01.010 https://doi.org/10.7551/mitpress/4117.001.0001 https://doi.org/10.1016/j.biopsycho.2011.02.019 https://doi.org/10.3389/fpsyg.2012.00428 hahn, saalbach, brunner & grabner 19 | f l r grabner, r. h., saalbach, h., & eckstein, d. (2012). language‐switching costs in bilingual mathematics learning. mind, brain, and education, 6(3), 147-155. https://psycnet.apa.org/doi/10.1111/j.1751-228x.2012.01150.x gramfort, a., luessi, m., larson, e., engemann, d. a., strohmeier, d., brodbeck, c., goj, r., jas, m., brooks, t., parkkonen, l., & hämäläinen, m. (2013). meg and eeg data analysis with mnepython. frontiers in neuroscience, 7, 267. https://doi.org/10.3389/fnins.2013.00267 gratton, g., coles, m. g. h., & donchin, e. (1983). a new method for off-line removal of ocular artifacts. electroencephalography and clinical neurophysiology, 55(4), 468–484. https://doi.org/10.1016/0013-4694(83)90135-9 hahn, c. g., saalbach, h., & grabner, r. h. (2017). language-dependent knowledge acquisition: investigating bilingual arithmetic learning. bilingualism: language and cognition, 22(1), 1-11. http://dx.doi.org/10.1017/s1366728917000530 huhta, a., luoma, s., oscarson, m., sajavaara, k., takala, s., & teasdale, a. (2002). dialang: a diagnostic language assessment system for learners. common european framework of reference for languages: learning, teaching, assessment. case studies, 130-145. imbo, i., & vandierendonck, a. (2007). the development of strategy use in elementary school children: working memory and individual differences. journal of experimental child psychology, 96(4), 284-309. https://psycnet.apa.org/doi/10.1016/j.jecp.2006.09.001 ischebeck, a., zamarian, l., siedentopf, c., koppelstätter, f., benke, t., felber, s., & delazer, m. (2006). how specifically do we learn? imaging the learning of multiplication and subtraction. neuroimage, 30(4), 1365-1375. https://doi.org/10.1016/j.neuroimage.2005.11.016 jäger, a. o., süß, h.-m., & beauducel, a. (1997). berliner intelligenzstruktur-test: bis-test form 4. göttingen: hogrefe. kirk, e.p., and m.h. ashcraft. 2001. telling stories: the perils and promise of using verbal reports to study math strategies. journal of experimental psychology. learning, memory, and cognition 27(1): 157–175. https://psycnet.apa.org/doi/10.1037/0278-7393.27.1.157 lefevre, j. a., sadesky, g. s., & bisanz, j. (1996). selection of procedures in mental addition: reassessing the problem size effect in adults. journal of experimental psychology: learning, memory, and cognition, 22(1), 216. http://dx.doi.org/10.1037/0278-7393.22.1.216 lemaire, p., & reder, l. (1999). what affects strategy selection in arithmetic? the example of parity and five effects on product verification. memory & cognition, 27(2), 364-382. https://doi.org/10.3758/bf03211420 lemhöfer, k., & broersma, m. (2012). introducing lextale: a quick and valid lexical test for advanced learners of english. behavior research methods, 44(2), 325-343. https://psycnet.apa.org/doi/10.3758/s13428-011-0146-0 linguatec (2015). retrieved from http://www.linguatec.de/en/text-to-speech/voice-reader-studio-15/. lo, y.y., & lo, e. s. c. (2014). a meta-analysis of the effectiveness of english-medium education in hong kong. review of educational research, 84(1), 47–73. https://doi.org/10.3102/0034654313499615 malt, b., & wolff, p. (eds.). (2010). words and the mind: how words capture human experience. new york: oxford university press. https://psycnet.apa.org/doi/10.1111/j.1751-228x.2012.01150.x https://doi.org/10.3389/fnins.2013.00267 https://doi.org/10.1016/0013-4694(83)90135-9 http://dx.doi.org/10.1017/s1366728917000530 https://psycnet.apa.org/doi/10.1016/j.jecp.2006.09.001 https://doi.org/10.1016/j.neuroimage.2005.11.016 https://psycnet.apa.org/doi/10.1037/0278-7393.27.1.157 http://dx.doi.org/10.1037/0278-7393.22.1.216 https://doi.org/10.3758/bf03211420 https://psycnet.apa.org/doi/10.3758/s13428-011-0146-0 https://doi.org/10.3102/0034654313499615 hahn, saalbach, brunner & grabner 20 | f l r marian, v., & fausey, c. m. (2006). language‐dependent memory in bilingual learning. applied cognitive psychology, 20(8), 1025-1047. https://doi.org/10.1002/acp.1242 marian, v., & neisser, u. (2000). language-dependent recall of autobiographical memories. journal of experimental psychology: general, 129(3), 361. https://doi.org/10.1037/0096-3445.129.3.361 mochida, k., & harrington, m. (2006). the yes/no test as a measure of receptive vocabulary knowledge. language testing, 2, 73–98. https://doi.org/10.1191/0265532206lt321oa pérez-cañado, m. l. (2012). clil research in europe: past, present, and future. international journal of bilingual education and bilingualism, 15(3), 315-341. https://doi.org/10.1080/13670050.2011.630064 pfurtscheller, g., & lopes da silva, f. h. (1999). event-related eeg/meg synchronization and desynchronization: basic principles. clinical neurophysiology, 110(11), 1842–1857. https://doi.org/10.1016/s1388-2457(99)00141-8 piesche, n., jonkmann, k., fiege, c., & keßler, j. u. (2016). clil for all? a randomised controlled field experiment with sixth-grade students on the effects of content and language integrated science learning. learning and instruction, 44, 108-116. https://doi.org/10.1016/j.learninstruc.2016.04.001 saalbach, h., eckstein, d., andri, n., hobi, r., & grabner, r. h. (2013). when language of instruction and language of application differ: cognitive costs of bilingual mathematics learning. learning and instruction, 26, 36-44. https://doi.org/10.1016/j.learninstruc.2013.01.002 smith-chant, b. l., & lefevre, j. a. (2003). doing as they are told and telling it like it is: self-reports in mental arithmetic. memory & cognition, 31(4), 516-528. https://doi.org/10.3758/bf03196094 spelke, e. s., & tsivkin, s. (2001). language and number: a bilingual training study. cognition, 78(1), 45-88. https://doi.org/10.1016/s0010-0277(00)00108-6 schmitt, d. (2005). test of english for international communication (toeic). esol tests and testing, 100-102. sweller, j., ayres, p., & kalyuga, s. (2011). measuring cognitive load. in cognitive load theory (pp. 71-85). springer, new york, ny. https://doi.org/10.1007/978-1-4419-8126-4_6 syndicate, u.c.l.e. (2001). oxford quick placement test. oxford: oxford university press. tschentscher, n., & hauk, o. (2016). frontal and parietal cortices show different spatiotemporal dynamics across problem-solving stages. journal of cognitive neuroscience, 28(8), 1098-1110. https://doi.org/10.1162/jocn_a_00958 vanbinst, k., ghesquiere, p., & de smedt, b. (2012). numerical magnitude representations and individual differences in children's arithmetic strategy use. mind, brain, and education, 6(3), 129136. https://doi.org/10.1111/j.1751-228x.2012.01148.x venkatraman, v., siong, s. c., chee, m. w., & ansari, d. (2006). effect of language switching on arithmetic: a bilingual fmri study. journal of cognitive neuroscience, 18(1), 64-74. https://doi.org/10.1162/089892906775250067 volmer, e., grabner, r. h., & saalbach, h. (2018). language switching costs in bilingual mathematics learning: transfer effects and individual differences. zeitschrift für erziehungswissenschaft, 21(1), 71-96. https://doi.org/10.1007/s11618-017-0761-0 https://doi.org/10.1002/acp.1242 https://doi.org/10.1037/0096-3445.129.3.361 https://doi.org/10.1191/0265532206lt321oa https://doi.org/10.1080/13670050.2011.630064 https://doi.org/10.1016/s1388-2457(99)00141-8 https://doi.org/10.1016/j.learninstruc.2016.04.001 https://doi.org/10.1016/j.learninstruc.2016.04.001 https://doi.org/10.1016/j.learninstruc.2013.01.002 https://doi.org/10.3758/bf03196094 https://doi.org/10.1016/s0010-0277(00)00108-6 https://doi.org/10.1007/978-1-4419-8126-4_6 https://doi.org/10.1162/jocn_a_00958 https://doi.org/10.1111/j.1751-228x.2012.01148.x https://doi.org/10.1162/089892906775250067 https://doi.org/10.1007/s11618-017-0761-0 hahn, saalbach, brunner & grabner 21 | f l r wolff, d. (2011). der bilinguale sachfachunterricht (clil): was dafür spricht, ihn als innovatives didaktisches konzept zu bezeichnen. in forum sprache, 6, 74-83. microsoft word nyberg_proofs.docx frontline learning research vol. 10 no. 1 (2022) 25 45 issn 2295-3159 corresponding author: kristin nyberg, university of education, institute of psychology, kunzenweg 21, 79117 freiburg, germany, kristin.nyberg@ph-freiburg doi: https://doi.org/10.14786/flr.v10i1.955 self-effective scientific reasoning? differences between elementary and secondary school students kristin nyberg1, susanne koerber1, christopher osterhaus2 1 university of education freiburg, germany 2 university of vechta, germany article received 21 september 2021 / article revised 14 may 2022 / accepted 13 june / available online 24 june abstract although scientific reasoning is not a formal, independent school subject, it is an increasingly important skill, especially for student learning in science, technology, engineering, and mathematics (stem) subjects. to promote scientific reasoning effectively, it is important to know its influencing factors. while cognitive influences have been investigated, affective-motivational factors, particularly self-efficacy, have rarely been considered in studies on scientific reasoning. to examine, for the first time, whether self-efficacy can be measured in a task-specific way and whether self-efficacy correlates with students’ scientific reasoning performance, the study assessed performance in scientific reasoning and self-efficacy (academic and task-specific) in a sample of 140 fourth graders and 148 eighth graders. as expected, higher correlations emerged for task-specific self-efficacy in both grades. a hierarchical cluster analysis showed that the correlational patterns were not the same across grade levels, with differences in self-estimated performance prevailing between the two grade levels: the largest cluster in grade 4 (41%) comprised children who significantly overestimated their performance, whereas the largest cluster in grade 8 (39%) comprised students who gave a realistic estimate of their own performance in scientific reasoning. this cluster was not present in grade 4. additional clusters of students who overestimated or underestimated their performance emerged in both grades. the results support the conclusion that self-efficacy expectations are important to consider when fostering scientific reasoning, and the large number of elementary school students who overestimated their performance suggests that not all students might benefit from interventions targeted at increasing self-efficacy. keywords: scientific reasoning, affective-motivational factor, self-efficacy, elementary school, secondary school nyberg, koerber & osterhaus 26 | f l r 1. introduction fostering science, technology, engineering and mathematics (stem) education in school and integrating technology and engineering into science education is a major challenge in recent years (bybee, 2010). nevertheless, stem education should not be reduced to content knowledge but instead should incorporate the development of scientific attitudes and critical, scientific reasoning or thinking (osborne, 2013). this broader conceptualization of science skills is mirrored in the conceptualization of scientific literacy in the pisa studies (oecd, 2006, 2015), which encompasses, apart from science content knowledge, scientific reasoning and personal affective-motivational factors. scientific reasoning can be viewed as a complex construct which includes several components, such as experimentation skills or understanding the nature of science. previous studies showed evidence for a common conceptual core in scientific reasoning (koerber et al., 2015) and validated group tests exist to reliably measure scientific reasoning in elementary school and above (e.g., koerber et al., 2015; osterhaus et al., 2015). for fostering scientific reasoning skills, it is important to know factors influencing scientific reasoning. most research in the area addresses the impact of general cognitive factors like language, intelligence, problem solving, executive functioning, and specific variables (e.g., advanced theory of mind) on the development of scientific reasoning (koerber et al., 2015; for an overview see zimmerman, 2007). research on the influence of affective-motivational factors in scientific reasoning, has been scarce, despite that affective-motivational constructs like self-beliefs and motivation are essential for academic achievement (cai et al., 2018; pajares, 1996; pajares & valiante, 1997). among affectivemotivational variables, the impact of self-efficacy expectations on academic performance is particularly strong, predicting up to 9% of academic performance in university students (richardson et al., 2012). self-efficacy expectations are classified as competence beliefs based on bandura’s social cognition theory (bandura, 1986; 1997). self-efficacy expectations are understood as the belief in one’s own ability to cope with a future situation or task. according to bandura (1997), the precise measurement of self-efficacy should include items that measure beliefs about expectations and performance as close as possible to the identical future tasks or situations, and the more specifically they are assessed, the more accurate the obtained action predictions will be (bandura, 1997; schwarzer & jerusalem, 2002). students’ expectations of self-efficacy are considered to play an important role in the school context. self-efficacy expectations positively influence academic performance, motivational processes, self-regulation, self-perception, and interest (bandura & schunk, 1981; klassen & usher, 2010; pajares & valiante, 1997; schunk, 1995). students with positive self-efficacy expectations spend more time on challenging tasks and situations. hence, they have the chance to receive more feedback compared to their classmates with lower self-efficacy expectations, who consequently avoid tasks and situations in this domain (britner & pajares, 2001; zeldin & pajares, 2000). when asking students in grades 6, 7 and 8 how skilled they would be at solving a math problem, they mainly use their experience with similar math problems to this point, as a basis for their judgment (usher & pajares, 2009). therefore, it is expected that the relation between performance and self-efficacy may be stronger at higher grade levels based on experience. researchers agree on the positive correlation between self-efficacy expectations and academic success, mostly referring to typical subjects of the curriculum such as mathematics (siefer et al., 2020) or writing (pajares & valiante, 1997). a meta-analysis by multon and brown (1991) with a total sample of roughly 5000 participants reported a correlation of r = .38 between academic performance and selfefficacy expectations. this result is backed up years later by a meta-analysis of honicke and broadbent (2016). they found a moderate positive correlation between the study performance of university students and their self-efficacy of r =.33 across 59 studies. according to bong (2006), the correlation between self-efficacy expectations and academic performance appears to be higher when measuring self-efficacy more specifically, which is consistent with bandura’s (1997) theory that self-efficacy expectations are context specific. in the following study, the assessment of self-efficacy is based on bandura's theory of measuring self-efficacy as specifically as possible, i.e. task-specific. nyberg, koerber & osterhaus 27 | f l r given this context-dependent requirement of a self-efficacy measure, generalizing results to other disciplines seems difficult, which consequently requires investigation in other domains, subjects, and skills. research dealing with self-efficacy and skills that are not included as an own subject in the curriculum but have an important relation to other subjects and skills like scientific reasoning is of great interest. liu et al. (2006), for example, found that science-class self-efficacy expectations of sixth-grade students (e.g., “i am confident i can learn the basic concepts taught in this science class”) correlated significantly with understanding science concepts (r =.28). jansen et al. (2015) investigated performance in scientific literacy and self-efficacy. the sample, which originated from the 2006 pisa survey in germany, included roughly 5000 secondary school students, most of whom were grade 9 at the time of the survey. the self-efficacy scale used in the study included eight items that assessed how confident they thought they were in solving a task in scientific literacy (e.g., predict how changes in an environment will affect the survival of certain species). the items that measured performance included real-world science tasks from the field of science, resembling the conceptualization of scientific literacy of the pisa study (bybee et al., 2010; oecd, 2006). the results of the study suggest that self-efficacy is a significant predictor of scientific literacy (jansen et al., 2015). even though a positive relation between performance and self-efficacy has been found in many domains, it does not necessarily imply that higher self-efficacy is always related to higher performance. students may overor underestimate their performance, which can hinder good performance (bandura, 1986). students who (strongly) overestimate their performance may not invest in learning the specific task or acquiring metacognitive skills (e.g., strategies to achieve goals), because they are confident to master the situation or task without investing more time and effort (hadwin & webster, 2013). identifying these groups of students is particularly important given that not all students might equally benefit from a direct increase in self-efficacy. the present study charters new territory: it investigates whether self-efficacy in scientific reasoning can be measured in a task-specific way. while this approach is similar to what has been shown in a specific area of mathematics (siefer et al., 2020), it is the first time that self-efficacy expectations are assessed in scientific reasoning in two ways of specificity (academic and task-specific). in addition, the study investigates whether there are groups of students with different levels of task-specific selfefficacy and performance in scientific reasoning. a sample of fourth and eighth graders was selected. eight graders were included because they had more opportunities to practice scientific reasoning and gain feedback on it, despite being more limited than in mathematics or sports. the fourth graders were chosen because previous studies, examining self-efficacy expectations had predominantly recruited secondary school and university students. moreover, studies show that perceived academic self-efficacy declines between sixth and eighth grade (harter, 1985). the decline can be described by a tendency of elementary school children to have unrealistically high beliefs about their competence (overestimate their competence) whereas as they get older, their self-judgement increasingly matches external evaluations (nicholls, 1978; stipek & hofman, 1980; pajares & schunk, 2001). this study intended to shed light on the beginnings of this relation. the study focuses on five key questions: (1) is our self-efficacy scale suitable to measure taskspecific self-efficacy in scientific reasoning for grades 4 and 8? (2) what is the relation between students’ performance in scientific reasoning and their self-efficacy expectations, specifically is this relation already observable in elementary school? (3) are the correlations higher when measuring taskspecific self-efficacy instead of academic self-efficacy? (4) and, are there differences in the strength of the relation between both grade levels? (5) finally, do diverse clusters exist which categorize students along the task-specific self-efficacy-performance relation (e.g., overor underestimate and high and low performer)? this last research question pertains to the question whether the relations within the sample are homogeneous or if interindividual differences exist. since on the one hand, both, self-efficacy and scientific reasoning, can be furthered through intervention or training (sodian et al., 2002; margolis & mccabe, 2006), and on the other hand, task-specific self-efficacy measurement provides an accurate insight into self-assessment and related task performance, this allows to specifically focusing on what nyberg, koerber & osterhaus 28 | f l r needs to be fostered: self-efficacy or scientific reasoning. getting better at scientific reasoning skills is especially crucial in relation to the performance in scientific literacy and the important stem subjects. 2. method 2.1 participants the sample consisted of n = 288 students (155 girls, 133 boys), n = 140 fourth graders (m = 9 years, 2 months, sd = 4 months) and n = 148 eighth graders (m = 13 years, 3 months, sd = 6 months). the students went mostly to 14 middle-class schools close to a mid-sized city in southern germany. the eighth graders all attended academic track schools (gymnasium). the data were collected in the first half of the academic year between october 2018 and january 2019. of the 288 children, 68 (23.6%) spoke at least one language other than german at home (the most frequently reported languages were russian, english, and french). student assent and written consent from caretakers were obtained for all participants. institutional review board (irb) approval was not sought for the study because the host institution did not have an established irb. 2.2 materials 2.2.1 scientific reasoning the performance in scientific reasoning was measured with six multiple-select tasks, three tasks testing students’ understanding of the nature of science (nos; cronbach’s a grade 4 = .53, grade 8 = .46) and three tasks testing their metaconceptual understanding of experimentation (unex cronbach’s a grade 4 = .52, grade 8 = .51). since previous studies revealed that the subcomponents of scientific reasoning skills share a common conceptual core among the subcomponents (koerber et al., 2015) and there is a correlation between nos and experimentation skills (osterhaus et al., 2017), the two scientific reasoning scales were collapsed and used as a single performance measure for scientific reasoning in the subsequent analysis (entire scale cronbach’s a grade 4 = .49, grade 8 = .48). further support for the use of the scale is the result of the study by osterhaus et al. (2020), who examined a short scale (spr-i(7)) for validity and reliability and showed similar reliabilities to the six items we used. nos. three nos tasks (a01 scientist, a03 middle ages and a11 mistakes in science) were selected from the science-p reasoning inventory (koerber et al., 2015; osterhaus et al., 2020). these three nos tasks assessed children’s understanding of what scientists do and the understanding of the hypothesis-evidence relation (see appendix for a sample item). for each task, three answer options were presented to the students (on a naïve level = 0 points, an intermediate level = 1 point, a scientifically advanced level = 2 points), and they were asked to agree or disagree with each of the answer options. the lowest level answer selected was taken as the final score on the entire item. thus, the children could obtain a maximum of two points per task and a maximum sum score of six (see osterhaus et al., 2020 for further coding details). unex. the unex tasks were taken from the study of osterhaus et al. (2015). the students were given three imaginary described experiments (trees u1, math textbook u2, classroom u3). the students were asked to decide whether it was a good or bad experiment to test the different assumptions (see appendix for a sample item). children were assigned 2 points, when they selected the correct answer and simultaneously rejected the wrong alternatives. one point was given for selecting the correct answer and the intermediate level but rejecting the naïve level (see coding of the nos tasks). nyberg, koerber & osterhaus 29 | f l r 2.2.2 self-efficacy expectations the self-efficacy expectations were measured on two levels: academic self-efficacy and task-specific self-efficacy expectations. academic self-efficacy. the academic self-efficacy expectations were measured with three items, used from jerusalem and satow (1999). the scale (wirkschul) originally comprises 7 items and was validated on 3000 secondary school students. the entire scale was used in our pilot study. from the subsequent interviews and the reliability tests, two items emerged as having poor reliability. in order to achieve a sufficient reliability in both grade levels for the analysis of the main survey, two additional items had to be excluded. the children were asked to indicate their agreement on a 4-point likert scale, ranging from “strongly disagree” (1) to “strongly agree” (4) with the following 1) “i can solve even the difficult tasks in class if i exert myself.” 2) “even if i am sick for a longer time, i am still able to perform well” 3) “if a teacher doubts my skills, i am sure that i can still perform well”. the scale was applied in german and the description of the items in the text are own translations. cronbach’s a was .61 for grade 4 and for grade 8 =.55. the reported reliabilities of > 5 can be interpreted for scales especially with few items (e.g., nunnally & bernstein, 1978). the corrected item-total correlation (rit) of .56-58 also point to the reliability of the scale. criterion validity is provided by the significant relations to task-specific self-efficacy (grade 8 r= .488 and grade 4 r= .360), school self-concept (grade 8 r= .690 and grade 4 r= .446), and interest (grade 8 r= .229 and grade 4 r= .180). task-specific self-efficacy. in this study a task-specific self-efficacy scale was used. in contrast to mathematics and other familiar academic areas, scientific reasoning tasks are less familiar to especially fourth graders, thus their respective performance might seem less predictable to them, more so since this competence is usually not explicitly taught in school as an own subject. the results of a pilot study with 57 fourth and eighth graders corroborated this impression. based on the participants feedback and in accordance with general guidelines on scale construction (bandura, 2006; moosbrugger & kleava, 2012) we designed a very task specific self-efficacy scale (3 items), in which the respective scientific reasoning task and the self-efficacy scale were presented together. the task-specific self-efficacy were assessed with the following three items together with the scientific reasoning task before the students were asked to solve each of the six scientific reasoning tasks: 1) “i know how to deal with the task” 2) “i am very familiar with such tasks and know how to solve the task” 3) “i would need help to solve the task.”. the likert scale ranged from “strongly disagree” (1) to “strongly agree” (4). 2.2.3 control variables with a proficiency test (elfe 1-6; lenhard & schneider, 2006) text comprehension was measured. to measure nonverbal intelligence a subtest of the cultural fair intelligence test (cft; weiß, 2006) was applied. 2.3 procedure in both age groups, the testing was conducted as a whole-class testing procedure, with each person working individually on their own booklet. before answering each scientific reasoning task, the students were asked to fill in the items of the task-specific self-efficacy. in a first step, they were instructed to look at the scientific reasoning task about 25 seconds but not to try to solve the task. the time of 25 seconds appeared in a pilot study to be the best time to give the students enough time to look over the task but not to try to solve it. after they rated the task-specific self-efficacy item, they solved the scientific reasoning task. to avoid confounding effects of reading ability, the items were presented by a powerpoint presentation and read aloud by an experimenter. the testing took about 60 minutes. nyberg, koerber & osterhaus 30 | f l r 3. results 3.1 core performance and suitability of the task-specific self-efficacy scale 3.1.1 scientific reasoning fourth graders scored an average of 5.23 (sd = 3.06) out of 12 points (43.6%), whereas eighth graders performed significantly better, scoring an average of 7.96 (sd = 2.46) out of 12 points (66.3%), t(277) = -6.03, p < .01. according to cohen (1992), the effect size of r = .60 is strong (see figure 1). 3.1.2 self-efficacy figure 1 shows the mean percent of academic and task-specific self-efficacy expectations transformed to a percentage scale (low to high feelings of self-efficacy). no significant differences in academic self-efficacy were found between grades (t(277) = 1.46, p = .15), between males and females in grade 8 (t(148) = -.294, p = .77) and grade 4 (t(140) = -1.25, p = .13). in task-specific self-efficacy, the eighth graders had significantly higher values than the fourth graders, t(279) = -3.36, p < .01 (see figure 1). the effect size was r = .04, representing a small effect (cohen, 1992). no significant difference was found between males and females in their task-specific self-efficacy in grade 8, t(146) = 1.92, p = .06, and in grade 4, t(139) = -.113, p = .91. figure 1. comparison of mean performance in scientific reasoning, task-specific and academic selfefficacy between grade 4 and 8. **p< .01, sr= scientific reasoning, se= self-efficacy the internal consistency in both grades (cronbach’s a grade 4 = .89, grade 8 = .88) were high as well as the corrected item-total correlations (rit grade 4 = .59-77, rit grade 8= .60-.82) which indicate a good reliability of the scale in both grade levels (table 1; bortz & döring, 2006). nyberg, koerber & osterhaus 31 | f l r table 1 corrected-item-total correlations of the task-specific self-efficacy for grades 4 and 8 rit-i nos unex 1_1 1_2 1_3 2_1 2_2 2_3 3_1 3_2 3_3 4_1 4_2 4_3 5_1 5_2 5_2 6_1 6_2 6_3 g 4 .70 .64 .68 .73 .73 .77 .77 .76 .75 .62 .59 .65 .59 .63 .67 .68 .64 .65 g 8 .62 .60 .61 .75 .81 .78 .68 .73 .76 .74 .76 .77 .65 .60 .82 .79 .80 .75 g=grade, 1= a01 scientist, 2= a03 middle ages, 3= a11 mistakes in science, 4= u1 trees, 5= u2 math textbook, 6= u3 classroom content and criterion validity are ensured by the process of scale construction and correlations in line with theoretical findings (marsh, 2018; schukajlow et al., 2012). significant relations were found with academic self-efficacy (grade 4 r= .360, grade 8 r= .488), academic self-concept (grade 4 r= .395, grade 8 r= .329), and interest (grade 4 r= .542, grade 8 r= .467). furthermore, we tested whether there was measurement invariance for the task-specific selfefficacy scale across grades 4 and 8, using r and the lavaan package (rosseel, 2012). the model shows scalar measurement invariance, following the criterion that maximum change per level in comparative fit index [cfi] and root-mean-square error of approximation [rmsea] should not exceed a certain threshold (cfi ≤ -.010 and rmsea ≤ .015; chen, 2007; for further discussions see cheung & rensvold, 2002). fit indices were as follows: configural (cfi= .997; rmsea= .043), metric (cfi= .994; rmsea= .051) and scalar (cfi= .989; rmsea= .066). all fit indices of the model accuracy remain in a good range (e.g., hu & bentler, 1999), and the maximum-change thresholds for cfi and rmsea are not exceeded by the scalar model. scalar invariance implies that the factor structure, factor loadings, and intercepts remain invariant across the two grade levels. thus, a comparison of items between participants from different groups (here grade 4 and 8) is possible with regard to the latent variable measured (here task-specific self-efficacy). 3.2 relation between scientific reasoning and self-efficacy 3.2.1 academic self-efficacy and scientific reasoning as shown in table 2, academic self-efficacy expectations correlated significantly with scientific reasoning in grade 8 but not in grade 4. no significant differences in the correlation in grade 8 between females and males were found (p = .34). the correlations were controlled for intelligence and reading ability. 3.2.2 task-specific self-efficacy and scientific reasoning performance in scientific reasoning and task-specific self-efficacy correlated significantly in both grades (table 2). again, no significant differences in the correlations between females and males were observed (grade 4, p = .45; grade 8, p =.24). the correlations were controlled for intelligence and reading ability. nyberg, koerber & osterhaus 32 | f l r table 2 correlations between scientific reasoning and academic self-efficacy or task-specific self-efficacy ***p<.001, *p<.05, se = self-efficacy 3.3 cluster analysis the correlation analysis showed a positive relation of task-specific self-efficacy and the performance in scientific reasoning, but the magnitude of the correlation could point to heterogenous patterns within the correlation. a hierarchical cluster analysis was performed to investigate whether there are clusters of students who underestimated or overestimated their performance in relation to their actual performance. therefore, the values of the performance in scientific reasoning and the taskspecific self-efficacy were z-standardized. for the performance, the z-standardized residuals were applied to eliminate the common variance between the performance and task-specific self-efficacy. the squared euclidean distance was taken as an approximate measure. as a clustering algorithm, the ward method was selected. ward's method has been shown to give better results with small clusters (e.g., everitt et al., 2001). this procedure was similar to the approach of previous studies (e.g., hallet et al., 2010; siefer et al., 2020). the cluster analysis was conducted separately for grades 4 and 8 because different profiles in the agreement between scientific reasoning and self-efficacy could be expected in each age group. the analysis revealed a four-cluster solution for the fourth graders and a five-cluster solution for the eighth graders. the number of clusters was determined based on dendrograms and cophenetic correlations. preliminary, on the whole no gender differences were found, except in the very small cluster 1 and 5, in grade 4 and grade 8 respectively. grade 4. cluster 1 (underestimators/good performance) included 18 students (14.5%) who showed the highest performance in scientific reasoning across all clusters. task-specific self-efficacy was also rated high, but it was below their performance. in other words, this group of students slightly underestimated their performance (see figure 2). the cluster contained significant more females than males, ꭕ2(1, n=17) = 3.56, p < .05). cluster 2 (overestimators/poor performance) was the most frequent cluster (n = 51, 41.1%). students performed below average, but despite their weak performance, they showed high task-specific self-efficacy indicating overestimation of their performance (see figure 2). no significant difference between the number of female or male students in this cluster was observed. cluster 3 (strong underestimators/good performance) (n = 24, 19.4%) could best be described as having high-performance and low task-specific self-efficacy. the students seemed to largely underestimate their performance (see figure 2). again, no significant difference was found in the number of female and male students. the last cluster 4 (slight underestimators/poor performance) contained n = 31 (25%) students showing poor performance and low task-specific self-efficacy. the task-specific self-efficacy was even poorer than the performance, which showed that students underestimate their performance (see figure 2). again, there was no significant difference between the numbers of female and male students. academic se task-specific se grade 4 scientific reasoning .076 .194* grade 8 scientific reasoning .253*** .290*** nyberg, koerber & osterhaus 33 | f l r figure 2. showing the resulting clusters with performance in scientific reasoning (z-standardized) and task-specific self-efficacy (z-standardized) in grade 4. cluster 2 is the most frequent cluster. se = self-efficacy a one-way analysis of variance (anova) was performed to statistically confirm the differences between the clusters. in grade 4, cluster assignment had a significant effect on the performance in scientific reasoning, f(4,131) = 90.567, p < .001. the performance of the different clusters differed significantly among students across different clusters, with only the performance in cluster 2 (m = -0.49, sd = 0.56) not being significantly different from the performance in cluster 4 (m = -0.70, sd = 0.50). cluster assignment also significantly affected the level of task-specific self-efficacy (f(4,130) = 91.05, p <.001). all clusters varied significantly in their level of task-specific self-efficacy, except for cluster 1 (m = -0.74 sd = 0.48) and 2 (m = -0.68 sd = 0.67), which were not significantly different. grade 8. the most frequent cluster in grade 8 was cluster 1 (realistic estimators/good performance) with n = 52 (39%) students. the students showed an above-average performance and a high task-specific self-efficacy, indicating that the students’ self-evaluated self-efficacy was close to their actual performance (see figure 3). no significant difference between the number of female or male students in this cluster was observed. cluster 2 (strong overestimators/very poor performance) contained n = 9 (6.7%) students with a poor performance and a high task-specific self-efficacy, indicating a high overestimation of their performance (see figure 3). again, no significant difference was found in the number of female and male students. in cluster 3 (strong underestimators/good performance) n = 21 (15.8%), students had a high performance and a poor task-specific self-efficacy. the students were high performance underestimators (see figure 3). there was no significant difference between the numbers of female and male students in this cluster. students (n = 25, 18.8%) in cluster 4 (overestimators/very poor performance) demonstrated poor performance and average task-specific self-efficacy, indicated below-average overestimation. also, no significant difference between the number of female or male students in this cluster was observed. cluster 5 (underestimators/average performance) with n = 28 (19.5%) included students with an average performance and a poor task-specific self-efficacy. the students underestimated their performance. cluster 5 was the only cluster to have a significant difference between the number of female and male students with significantly more males in the cluster, ꭕ2(1, n = 25) = 3.85. p < .05. nyberg, koerber & osterhaus 34 | f l r anovas were also conducted for grade 8 to test for differences between the clusters. cluster assignment had a significant effect on the performance in scientific reasoning, f(5,140) = 79.90, p < .001. the performance of the clusters differed significantly, except between cluster 2 (m = -1.28, sd = 0.42) and cluster 4 (m = 1.31, sd = 0.65). cluster assignment also significantly affected the level of task-specific self-efficacy, f(5,139) = 82.08, p < .001. all clusters varied significantly in their level of task-specific self-efficacy, except between cluster 3 (m = -0.75 sd = 0.53) and 5 (m = -0.88 sd = 0.40). figure 3. showing the resulting clusters with performance in scientific reasoning (z-standardized) and task-specific self-efficacy (z-standardized) in grade 8. cluster 1 is the most frequent cluster. se = selfefficacy 4. discussion the present study addressed five questions (1) is our self-efficacy scale suitable to measure task-specific self-efficacy in scientific reasoning for grades 4 and 8? (2) what is the relation between students’ performance in scientific reasoning and their self-efficacy expectations, specifically is this relation already observable in elementary school? (3) are the correlations higher when measuring taskspecific self-efficacy instead of academic self-efficacy? (4) and, are there differences in the strength of the relation between both grade levels? (5) finally, do diverse cluster exist which categorize students along the task-specific self-efficacy-performance relation (e.g., overor underestimate and high and low performer)? 4.1 core performance and suitability of the used scales as expected, performance in scientific reasoning was significantly higher in grade 8 than in grade 4. nonetheless, no ceiling effect was observed in the performance of the eighth graders. this finding suggests that performance in scientific reasoning continues to develop through the elementary school years into the secondary school years and may not be completed by the end of the secondary school years, which is in line with previous findings (bullock et al., 2009; bullock & ziegler, 1999; koerber et al., 2015). when studying students’ self-efficacy, it should be emphasized that self-efficacy could be measured task-specifically for scientific reasoning skills– in elementary and secondary school students. and, in contrast to many other studies, the present study investigated two levels of self-efficacy: nyberg, koerber & osterhaus 35 | f l r academic self-efficacy and task-specific self-efficacy. the item characteristics and criterion-related correlations show comparable values for grade 4 and 8 and indicate that the scales can be applied in both grades. the results related to academic self-efficacy should be interpreted carefully, as the scale does not show high internal consistency, nevertheless (significant) correlations were found, which indeed could have been stronger with a scale having higher internal consistency. nevertheless, for a replication of the results, especially in relation to the academic self-efficacy scale, the low internal consistency should be included in further considerations. the item characteristics, correlations of the task-specific self-efficacy scale, the high internal consistency and scalar measurement invariance across both grade levels suggest the reliable and valid use of the scale for grades 4 and 8. the two grades revealed similar academic self-efficacy with no significant difference (grade 8: 64% vs. grade 4: 62%). the task-specific self-efficacy, however, differed significantly between the two grades. eighth graders reported significantly higher task-specific self-efficacy than fourth graders. this finding is consistent with the development of self-efficacy expectations as described by bandura (1997) who identified different sources of self-efficacy. two of the crucial sources for building self-efficacy expectations may be previous mastery experiences and vicarious experiences (usher, 2009). mastery experiences occur when students experience feelings of success on a particular task, which engenders the belief that they can succeed in the task again. vicarious experiences refer to observing someone else perform the task and solve it. students learn from this observation that they may also succeed at the task. young students often have little opportunity to benefit from such experiences in building their selfefficacy beliefs in this specific area because they are not so familiar with scientific reasoning tasks yet. however, eighth graders are more likely to have the opportunity in school to work on science problems in physics or nwt (science and technology) than fourth graders. therefore, it is not surprising that eighth graders rated themselves significantly higher in task-specific self-efficacy. 4.2 relation between scientific reasoning and self-efficacy the correlation between academic self-efficacy and scientific reasoning was found only in grade 8. task-specific self-efficacy and performance in scientific reasoning correlated significantly across grades 4 and 8, with a higher correlation emerging for eighth graders. compared to other studies, we found correlations that tended to be lower. multon and brown (1991) reported a mean correlation of r = .38 in their meta-analysis, and honicke and broadbent (2016) found a similar result of r = .33. however, studies with different samples and heterogenous instruments were included in the studies. in a study by siefer et al. (2020), task-specific self-efficacy correlated with the performance in a specific mathematical area. they reported a correlation of r = .39 between task-specific self-efficacy and the students’ performance on a test of linear functions in grade 8 and 9. these studies also would have hinted towards a higher agreement between performance and self-efficacy if self-efficacy was measured specifically. the lower correlations in our study may have resulted from scientific reasoning not being assigned as an independent subject to the curriculum. receiving feedback and benefiting from mastery experiences in scientific reasoning is more difficult for scientific reasoning tasks (as opposed to mathematics). consequently, students receive less formal feedback on their performance (e.g., in the form of grades). fourth-grade students especially have little opportunity to profit from mastery experiences in scientific reasoning because the opportunities to work on science problems, as required in the here used scientific reasoning tests, are missing. in general, in contrast to self-efficacy measures concerning science content knowledge (e.g., in the 2015 pisa study; schiepe-tiska et al., 2016), we observed no gender differences neither in grade 4 nor in grade 8 in the (scientific reasoning) task-specific self-efficacy and the correlations. this is consistent with previous findings (e.g., koerber et al., 2015) who also reported no gender differences in scientific reasoning. nyberg, koerber & osterhaus 36 | f l r 4.3 differences in the agreement between scientific reasoning and task-specific self-efficacy the positive correlations between performance and task-specific self-efficacy in grades 4 and 8 suggested that higher performance is associated with higher task-specific self-efficacy. however, the rather low correlation coefficients could also indicate interindividual differences between students. for that purpose, a hierarchical cluster analysis was conducted. looking at the found clusters the most noteworthy result was the largest cluster in grade 8, which contained the students who judged their performance realistically and showed a high performance. this fits with the findings of siefer et al. (2020), which also showed that the largest group of the sample (33%) realistically assessed themselves in the domain of linear functions in grades 8 and 9. the cluster with the students, who realistically judge their performance, however, was nonexistent in grade 4. many fourth-grade students were assigned to the cluster of overestimators who performed poorly. this in line with previous research findings that elementary school students are more likely to overestimate themselves, in contrast to the beginning of secondary school, where self-efficacy appears to decline (harter, 1985; zimmerman, 1995). students in grade 4 seemed to mostly overestimate or underestimate their abilities. being able to realistically judge one’s performance requires the skill to realistically judge the required task at an abstract level, which could be difficult for fourth graders. thus, the result for grade 4 students could have reflected in a nonrealistic judgment of their performance. for example, a study by kruger and dunning (1999) suggests the skills needed for a certain task or domain are the same skills needed to evaluate one’s performance. a reasonable assumption is that the fourth graders, who are not as good at scientific reasoning as the eighth graders, may not be as skilled at judging their skills and may consequently overestimate or underestimate their scientific reasoning skills. students become more realistic in their self-efficacy evaluations over time. nevertheless, clusters are more divergent in grade 8 in their levels of performance and self-efficacy. it is well established in literature that both great overestimation and underestimation of one's own abilities can be a barrier to performance (e.g., bandura, 1986; hattie, 2013). if students are overconfident that they can complete the task or situation well, this may lead them to not take proper strategies or other actions to master the task (hadwin & webster, 2013). in contrast, learners with insufficient confidence may be absorbed with cognitive resources that they have to invest more effort in achieving the goal than they should have (hattie, 2013). the generally high number of underestimator distributed across grade 4 and 8 in this study could find themselves in a helpless position, which then results in poorer performance in scientific reasoning, which in turn is critical for performance in scientific literacy and the important stem subjects. this differentiated information is especially relevant for teachers, as it is crucial for them to 1) realize at all, that there might be a substantive heterogeneity in their class with respect to students’ estimation of their own performance. 2) this differentiated information helps teacher providing feedback that results in a realistic confidence with at most minimal overestimation, as this allows for the best student performance (e.g., bandura, 1977) 3) this in turn, might support students in their self-regulation behavior. and finally, 4) this promotional opportunity appears useful in the context that self-efficacy is a measure that can be improved through intervention or training (sodian et al., 2002; margolis & mccabe, 2006). 4.4 limitations and future directions the cross-sectional design used in the present study allows for a thorough investigation of selfefficacy expectations and their relation to scientific reasoning skills. when interpreting the results, the rather low reliability of the academic self-efficacy scale should be taken into account. to understand and explain interindividual differences in the agreement between performance in scientific reasoning and task-specific self-efficacy, variables that provide information about the use of feedback from teachers and parents would be important. furthermore, we cannot conclude whether self-efficacy influences scientific reasoning or, conversely, whether scientific reasoning influences self-efficacy. bandura (1997) argued for a reciprocal relation between skills and self-efficacy. lawson et al. (2007) found support for the hypothesis that reasoning skills are a good predictor of self-efficacy but not the other way around. the fact that results from specific domains cannot simply be transferred across nyberg, koerber & osterhaus 37 | f l r domains was shown in the findings from schöber et al. (2018). for the math domain, the authors found support for a self-enhancement approach (i.e. self-efficacy influences math performance), whereas for writing skills, a skill-development approach was supported (i.e. writing skills influence self-efficacy). only longitudinal studies can reveal the direction and strength of the relations in the context of scientific reasoning skills. the longitudinal design also allows to examine how self-efficacy and scientific reasoning develop over time. the outcomes could provide relevant implications for a potential intervention. identifying where to target an intervention is critical, whether to focus on scientific reasoning skills or self-efficacy expectations and to identify the developmental stage that may be most sensitive to intervention. 4.5 conclusion the present study showed, for the first time, that self-efficacy can be measured in a task-specific manner in scientific reasoning and found a positive correlation between (task-specific) self-efficacy and performance in scientific reasoning, already in students at the end of elementary school. this suggests that a precise and task-specific measurement of self-efficacy can detect effects already in elementary school children. in addition, our results show substantial interindividual differences and differences in age groups in the agreement between self-efficacy and scientific reasoning. this is an important outcome, especially regarding the influence of scientific reasoning skills on science content knowledge and the relevant stem subjects. therefore, this is a possible starting point for promoting this academic discipline. keypoints self-efficacy can be measured in a task-specific manner in scientific reasoningfor both elementary and secondary school students stronger correlation between self-efficacy and scientific reasoning with more precisely measured self-efficacy differences in the agreement between scientific reasoning and task-specific self-efficacy among and across grades references bandura, a. (1986). social foundations of thought and action: a social cognitive theory. engelwood cliffs. bandura, a. (1997). self efficacy: the exercise of control. freeman. bandura, a (2006). guide to the construction of self-efficacy scales. in f. pajares & t. urdan (eds.), self-efficacy beliefs of adolescents (pp. 307-337). information age. bandura, a., & schunk, d. h. (1981). cultivating competence, self-efficacy and intrinsic interest through proximal self-motivation. journal of personality and social psychology, 41, 586–598. https://doi.org/10.1037/0022-3514.41.3.586 bong, m. (2006). asking the right question. how confident are you thatyou could successfully perform these tasks? in f. pajares & t. c. urdan (eds.), self-efficacy beliefs of adolescents (pp.287–305). information age. nyberg, koerber & osterhaus 38 | f l r bortz, j., & döring, n. (2006). forschungsmethoden und evaluation für humanund sozialwissenschaftler [research methods and evaluation for social scientists]. springer. britner, s. l., & pajares, f. (2001). self-efficacy beliefs, motivation, race, and gender in middle school science. journal of women and minorities in science and engineering, 7, 271–285. https://doi.org/10.1615/jwomenminorscieneng.v7.i4.10 bullock, m., sodian, b., & koerber, s. (2009). doing experiments and understanding science: development of scientific reasoning from childhood to adulthood. in w. schneider & m. bullock (eds.). human development from early childhood to early adulthood: findings from a 20 year longitudinal study (pp.173–197). psychology press. bullock, m., & ziegler, a. (1999). scientific reasoning: developmental and individual differences. in f. e. weinert & w. schneider (eds.), individual development from 3 to 12: findings from the munich longitudinal study (pp. 38–54). cambridge university press. bybee, r. w. (2010). advancing stem education: a 2020 vision. technology and engineering teacher, 70, 30–35. cai, d., viljaranta, j., & georgiou, g. k. (2018). direct and indirect effects of self-concept of ability on math skills. learning and individual differences, 61, 1–38. https://doi.org/10.1016/j.lindif.2017.11.009 chen, f.f. (2007). sensitivity of goodness of fit indexes to lack of measurement invariance. structural equation modeling, 14, 464-504. https://dx.doi.org/10.1080/10705510701301834 cheung, g. w., & rensvold, r. b. (2002). evaluating goodness-of-fit indexes for testing measurement invariance. structural equation modeling, 9, 233-255. https://doi.org/10.1207/s15328007sem0902_5 cohen, j. (1992). a power primer. psychological bulletin, 112, 155-159. https://doi.org/10.1037/0033-2909.112.1.155 everitt, b. s., landau, s., & leese, m. (2001). cluster analysis (4th ed.). oxford university press. hadwin, a. f., & webster, e.a. (2013). calibration in goal setting: examining the nature of judgments of confidence. learning and instruction, 24, 37-47. https://doi.org/10.1016/j.learninstruc.2012.10.001 hattie, j. (2013). calibration and confidence. where to next? learning and instruction, 24, 62-66. https://doi.org/10.1016/j.learninstruc.2012.05.009 hallet, d., nunes t., & bryant, p. (2010). individual differences in conceptual and procedural knowledge when learning fractions. journal of educational psychology, 102, 395-406. https://doi.org/10.1037/a0017486 harter, s. (1985). manual for the self-perception profile for children. university of denver honicke, t., & broadbent, j. (2016). the influence of academic self-efficacy on academic performance: a systematic review. educational research review, 17, 63–84. https://doi.org/10.1016/j.edurev.2015.11.002 hu, l., & bentler, p.m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives. structural equation modeling, 6, 1-55. https://doi.org/10.1080/10705519909540118 nyberg, koerber & osterhaus 39 | f l r jansen, m., scherer, r., & schroeders, u. (2015). students’ self-concept and self-efficacy in the sciences: differential relations to antecedents and educational outcomes. contemporary educational psychology, 41, 13–24. https://doi.org/10.1016/j.cedpsych.2014.11.002 jerusalem, m., & satow, l. (1999). schulbezogene selbstwirksamkeitserwartung. in r. schwarzer & m. jerusalem (eds.), skalen zur erfassung von lehrerund schülermerkmalen. freie universität berlin. klassen, r.m., & usher, e. l. (2010). self-efficacy in educational settings: recent research and emerging directions. in t. c. urdan & s. a. karabenick (eds.), the decade ahead: theoretical perspectives on motivation and achievement (pp.1–33). emerald. koerber, s., mayer, d., osterhaus, c., schwippert, k., & sodian, b. (2015). the development of scientific thinking in elementary school: a comprehensive inventory. child development, 86, 327– 336. https://doi.org/10.1111/cdev.12298 kruger, j., & dunning, d. (1999). unskilled and unaware of it: how difficulties in recognizing one’s own incompetence lead to inflated self-assessments. journal of personality and social psychology, 77, 1121–1134. https://doi.org/10.1037/0022-3514.77.6.1121 lawson, a.e., banks, d.l., & lovgin, m. (2007). self-efficacy, reasoning ability, and achievement in college biology. journal of research in science teaching, 44, 706-724. https://doi.org/10.1002/tea.20172 lenhard, w., & schneider, w. (2006). elfe 1-6: ein leseverständnistest für erstbis sechstklässler. hogrefe. liu, m., hsieh, p., cho, y., & schallert, d. l. (2006). middle school students’ self-efficacy, attitudes, and achievement in a computer-enhanced problem-based learning environment. journal of interactive learning research, 17, 225–242. margolis, h., & mccabe, p. p. (2006). improving self-efficacy and motivation: what to do, what to say. intervention in school & clinic, 41, 218–227. https://10.1177/10534512060410040401 marsh, h.w., pekrun, r., parker, p. d., murayama, k., guo, j., dicke, t., & arens, a.k. (2018). the murky distinction between self-concept and self-efficacy: beware of lurking jingle-jangle fallacies. journal of educational psychology, 111, 331-353. https://doi.org/10.1037/edu0000281 moosbrugger, h., & kelava, a. (2012). testtheorie und fragebogenkonstruktion [test theory and questionnaire construction]. springer. multon, k. d., & brown, s. d. (1991). relation of self-efficacy beliefs to academic outcomes: a meta-analytic investigation. journal of counseling psychology, 18, 30–38. nagengast, b., marsh, h. w., scalas l.f., xu, m. k., hau, k. t., & trautwein, u. (2011). who took the “x” out of expectancy-value theory? a psychological mystery, a substantive-methodological synergy, and a cross-national generalization. psychological science, 22, 1058–1066. https://doi.org/10.1177/0956797611415540 nicholls, j. g. (1979). development of perception of own attainment and causal attributions for success and failure in reading. journal of educational psychology, 71, 94–99. https://doi.org/10.1037/00220663.71.1.94 nunnally, j.c. & bernstein i.h. (1978). psychometric theory. new york: mcgraw-hill oecd (2006). assessing scientific, reading and mathematical literacy: a framework for pisa 2006. paris: oecd. nyberg, koerber & osterhaus 40 | f l r oecd (2015). oecd science, technology and industry scoreboard 2015: innovation for growth and society. paris: oecd. osborne, j. (2013). the 21st century challenge for science education: assessing scientific reasoning. thinking skills and creativity, 10, 265–279. https://doi.org/10.1016/j.tsc.2013.07.006 osterhaus, c., koerber, s., & sodian, b. (2020). the science-p reasoning inventory (spr-i): measuring emerging scientific reasoning skills in primary school. international journal of science education, 42, 1087-1107. https://doi.org/10.1080/09500693.2020.1748251 osterhaus, c., koerber, s., & sodian, b. (2017). scientific thinking in elementary school: children’s social cognition and their epistemological understanding promote experimentation skills. developmental psychology, 53, 450-462. https://doi.org/10.1037/dev0000260 osterhaus, c., koerber, s., & sodian, b. (2015). children’s understanding of experimental contrast and experimental control: an inventory for primary school. frontline learning research, 3, 56–94. pajares, f. (1996). self-efficacy beliefs in academic settings. review of educational research, 66, 543–578. https://doi.org/10.3102/00346543066004543 pajares, f., & schunk, d.h. (2001). the development of academic self-efficacy. in a. wigfield & j.s. eccles (eds.), development of achievement motivation (pp. 15-31). academic press. pajares, f., & valiante, g. (1997). influence of self-efficacy on elementary students’ writing. the journal of educational research, 90, 353–360. https://doi.org/10.1080/00220671.1997.10544593 richardson, m., abraham, c., & bond, r. (2012). psychological correlates of university students’ academic performance: a systematic review and meta-analysis. psychological bulletin, 138, 353– 387. https://doi.org/10.1037/a0026838 rosseel, y. (2012). lavaan: an r package for structural equation modeling. journal of statistical software, 48, 1-36. https://doi.org/10.18637/jss.v048.i02 schiepe-tiska, a., simm, i., & schmidtner, s. (2016). motivationale orientierungen, selbstbilder und berufserwartungen in den naturwissenschaften in pisa 2015. in k. reiss, c. sälzer, a. schiepetiska, e. klieme, & o. köller (eds.), pisa 2015. eine studie zwischen kontinuität und innovation (pp. 99–132). waxmann. schöber, c., schütte, k., köller, o., mcelvany, n., & gebauer, m. m. (2018). reciprocal effects between self-efficacy and achievement in mathematics and reading. learning and individual differences, 63, 1–11. https://doi.org/10.1016/j.lindif.2018.01.008 schunk, d. h. (1995). self-efficacy and education and instruction. in j. e. maddux (ed.), self efficacy, adaptation, and adjustment: theory, research, and application (pp. 281–303). plenum press. schukajlow, s., leiss, d., pekrun, r., blum, w., muller, m., & messner, r. (2012). teaching methods for modelling problems and students’ task-specific enjoyment, value, interest and selfefficacy expectations. educational studies in mathematics, 79, 215–237. https://10.1007/s10649011-9341-2 schwarzer, r., & jerusalem, m. (2002). das konzept der selbstwirksamkeit. zeitschrift für pädagogik, 44, 28–53. siefer, k., leuders, t., & obersteiner a. (2020). leistung und selbstwirksamkeitserwartung als kompetenzdimension: eine erfassung individueller ausprägungen im themenbereich linearer funktionen. journal für mathematik-didaktik, 41, 267-299. https://doi: 10.1007/s13138-01900147-x nyberg, koerber & osterhaus 41 | f l r stipek, d.c. & hoffman, j.h. (1980). children’s achievement-related expectancies as a function of academic performance histories and sex. journal of educational psychology, 70, 154-166. sodian, b., thoermer, c., kircher, e., grygier, p., & günther, j. (2002). vermittlung von wissenschaftsverständnis in der grundschule [teaching understanding the nature of science in elementary school]. zeitschrift für pädagogik, 45, 192–206. usher, e. l., & pajares, f. (2009). sources of self-efficacy in mathematics: a validation study. contemporary educational psychology, 34, 89–101. https://doi.org/10.1016/j.cedpsych.2008.09.002 weiß, r.h. (2006). grundintelligenztest skala 2 revision (cft 20-r). hogrefe zeldin, a. l., & pajares, f. (2000). against the odds: self-efficacy beliefs of women in mathematical, scientific, and technological careers. american educational research journal, 37, 215–246. https://doi.org/10.3102%2f00028312037001215 zimmerman, b. (1995). self-efficacy and educational development. in a. bandura (ed.). self-efficacy in changing societies (pp. 202-231). cambridge university press. zimmerman, c. (2007). the development of scientific thinking skills in elementary and middle school. developmental review, 27, 172–223. https://doi.org/10.1016/j.dr2006.12.001 nyberg, koerber & osterhaus 42 | f l r this is not about solving the task, but you are asked to complete the question below in terms of yourself. appendix 1 1. sample items 1.1 task-specific self-efficacy item and nos task (a03 middle ages; osterhaus et al., 2020) strongly disagree disagree agree strongly agree 1. i know how to deal with the task. o o o o 2. i am very familiar with such tasks and know how to solve the task. o o o o 3. i would need help to solve the task. o o o o long ago, in the middle ages, people believed there are witches who could make people sick. a modern-day scientist traveled back to the middle ages with a time machine. scientists in the middle ages thought that witches can make people sick. the modern-day scientist believes that bacteria can make people sick. the modern-day scientist shows the scientist from the middle ages the bacteria under the microscope and explains: “these bacteria are the reason why people get sick!” what will the scientist from the middle ages say to this? nyberg, koerber & osterhaus 43 | f l r now you may solve the task. what will the scientist from the middle ages say to this? he would say this he would not say this 1. “of course, you’re right. bacteria make people sick, not witches.” ¨ ¨ 2. “bacteria could be the witches’ little helpers.” ¨ ¨ 3. “it may be true that there are bacteria here, but witches are still the ones who make people sick.” ¨ ¨ which is the best answer? no. nyberg, koerber & osterhaus 44 | f l r this is not about solving the task, but you are asked to complete the question below in terms of yourself. 1.2 task-specific self-efficacy item and unex task (u1 trees; osterhaus et al., 2015) a scientist travels to a faraway planet, planet ogi. there, he observes that all trees are very small. the scientist develops a tree medicine that is supposed to help the trees grow. he calls this medicine supergrow. the scientist wants to find out whether his supergrow works and whether it really makes the trees grow. therefore, he conducts an experiment. he gives his supergrow to all trees on planet ogi. six months later, the scientist travels back to planet ogi. he observes that all trees are huge now he is convinced: “my supergrow works!” was this a good experiment? strongly disagree disagree agree strongly agree 1. i know how to deal with the task. o o o o 2. i am very familiar with such tasks and know how to solve the task. o o o o 3. i would need help to solve the task. o o o o nyberg, koerber & osterhaus 45 | f l r now you may solve the task. was this a good experiment? ¨ yes ¨ no susan, lisa, and vera wonder whether this was a good experiment. who is right and who is not? is right is not right 1. susan says: „it was not a good experiment because he does not know how big the trees would have grown without his supergrow.” ¨ ¨ 2. lisa says: „it was a good experiment because he found that all trees have grown huge after receiving his supergrow.” ¨ ¨ 3. vera says: „it was a good experiment because you can only see whether things work if you test them.” ¨ ¨ which of the three girls has the best answer? no.______ huber publication frontline learning research vol.13 no. 3 (2025) 53 82 issn 2295-3159 addressing boundary conditions of cognitive and motivational effects of gamified learning stefan huber1, elizabeth cloude2, lukas ober1, moritz edlinger1, antero lindstedt3, kristian kiili3 & manuel ninaus1,4 a university of graz, austria bmichigan state university, usa ctampere university, finland d lead graduate school and research network, university of tübingen, germany> article received 24 december 2025 / article revised 30 may 2025 / accepted 19 august 2025/ available online 12 september 2025 abstract there is a growing interest in developing gamified learning solutions to address educational challenges. however, learning is highly influenced by the conditions in which it takes place (e.g., does gamified learning in a laboratory setting replicate the outcomes of gamified learning online at home?). hence, it is crucial to understand the boundary conditions of different learning contexts to effectively implement gamified interventions that provide optimal learner support. this work contributes to such an understanding by assessing how general contextual aspects of three studies on gamified learning influence cognitive learning and motivational outcomes. therefore, we re-examined the results of two earlier published online studies (study 1: n=285; study 2: n=61) and compared the results to a recently conducted laboratory study (study 3: n=121), all of which employed the same associative learning task. comparing results through a bayesian lens, we find that motivational outcomes induced by gamification differ substantially between contexts. in contrast, cognitive learning outcomes seem comparatively robust across different contextual factors, with some indication of subtle influences in agreement with cognitive learning theories. implications are discussed for future empirical research on learning, highlighting how a better understanding of boundary conditions of gamified learning interventions could open perspectives for context-aware educational interventions. keywords: gamification, game-based learning, context, boundary conditions, cognition, motivation info corresponding author email: stefan.huber@uni-graz.at doi: https://doi.org/10.14786/flr.v13i3.1653 1. introduction research on game-based pedagogies has repeatedly shown that the common factor across various forms of these pedagogies – namely, game elements – can effectively influence learning (barz et al., 2023; sailer & homner, 2020; wouters et al., 2013; zainuddin et al., 2020). however, studies find that the efficacy of game elements can substantially differ, depending on which learning outcome is under investigation. that is, game elements might have positive effects on cognitive learning outcomes, but not on motivational or behavioral outcomes (sailer & homner, 2020). under different conditions, game elements may have positive effects on motivational outcomes yet leave cognitive outcomes unaffected (huber et al., 2023, 2024). furthermore, these effects can depend substantially on the exact contextual features of a learning situation (schlag et al., 2024). in particular, the efficacy of game elements may depend on education level (arztmann et al., 2023; sailer & homner, 2020), gender (conte, 2019; mavridis et al., 2017; l. zhang et al., 2024), subject area (ritzhaupt et al., 2021), research context (huang et al., 2020; sailer & homner, 2020), or sampled populations (ritzhaupt et al., 2021). these findings suggest that the efficacy of game elements on learning depends on specific boundary conditions under which game elements become especially effective for particular learning outcomes. with the term “boundary conditions” we refer to mayer’s (2024) notion that each multimedia learning principle “is subject to boundary conditions including for whom the principle applies, for which kind of lesson the principle applies, and under what circumstances the principle applies” (p. 19). that is, mayer recognizes that the efficacy of each multimedia learning principle depends on several factors: learner characteristics (“for whom the principle applies”), contextual aspects in a narrower sense (“which kind of lesson”, i.e., which subject), and in a broader sense (“under what circumstances”), the environment (e.g., online vs. classroom), or the intention (e.g., related to grading or on a voluntary basis) in which learning occurs. in the present study, we build upon this definition of boundary conditions and the opportunity that we recently used the same associative learning task in several experiments on gamified learning under varying contextual conditions (e.g., conducting the experiments online or in the laboratory). to compare the outcomes of those experiments, we conducted secondary analyses of the findings from two previously published online studies (huber et al., 2023, 2024), in conjunction with the results of a recent laboratory study hence allowing an assessment of how contextual features, not inherent to the learning task per se, shape the efficacy of gamified learning. as both cognitive learning and motivational outcomes were investigated in all three juxtaposed studies, the comparison yields a perspective regarding the boundary conditions of gamified learning on two learning outcomes. the implications of the findings may provide guidance for future research dedicated to further resolve the exact impact of specific contextual features on gamified learning outcomes. gamified learning is “the use of game design elements in non-game contexts” (deterding et al., 2011, p. 9). we focus on this type of game-based pedagogy because game design elements are one essential and common factor among various game-based pedagogies, such as game-based learning, gamification, or playful learning (plass et al., 2020). the focus on the efficacy of game elements on learning outcomes, as in the re-assessed previous work (huber et al., 2023, 2024), may warrant some generalizability across game-based pedagogies. whereas we use the term game-based pedagogy here rather as an umbrella term subsuming game-based, gamified, and playful learning under one term, it generally refers to a teaching approach that integrates games or game design elements into the learning processes and practices to promote motivation, learning, and assessment. as such, the term includes many aspects that go beyond the mere use of game elements like teachers’ roles, integration with the curriculum, educational objectives, or classroom practices (palha & jukić matić, 2025). while not considered in the present work, these aspects might reveal further boundary conditions in practice. the remainder of the introduction section is organized as follows. in section 1.1, we provide a theoretical account for why boundary conditions are important in game-based pedagogies. in section 1.2, we provide an overview of related empirical work, suggesting the importance of considering boundary conditions within the framework of game-based pedagogies. in section 1.3, we summarize the specific aims and research questions of the present study and provide short summaries of two previously published online studies (huber et al., 2023, 2024) as we conducted secondary analyses of the data in the present work. 1.1 theoretical key concepts in psychology, the perspective that a person’s behavior (including learning) results from the interplay between endogenous processes originating from within the person and exogenous characteristics and processes determined by the person’s environment is not new. already in 1936, kurt lewin coined the famous quasi-formula that behavior is a function of a person and their environment (lewin, 1936). that is, according to lewin (1946), human behavior needs to be regarded “as a function of the total situation” (p. 791), in which persons and their contexts constitute a unified whole (plass & kaplan, 2016). as a fundamental key concept, this notion is at the core of many theoretical frameworks in learning research and also applies to game-based pedagogies. although game-based pedagogies are a relatively recent addition to multimedia learning, mayer’s (2019) cognitive model of multimedia learning provides a useful starting point to describe their conceptual foundation. at the core of mayer’s model are three basic principles (mayer, 2011, 2014): (i) the dual channel principle, according to which people have separate channels for processing visual and verbal material; (ii) the limited capacity principle, which states that people can only process a small amount of material in each channel at one time; and (iii) the active processing principle, according to which deep learning occurs when people engage in active cognitive processing during multimedia learning. according to mayer (2019), these core principles need to be extended beyond cognition to describe game-based pedagogies with two further elements: motivation and meta-cognition. a modern framework integrating both of these elements with mayer’s cognitive model is provided by plass and kaplan (2016): the integrated cognitive affective model of learning with multimedia (icalm). a simplified version of the icalm is depicted in fig. 1(a). an essential ingredient of this model is that learners make meaning of multimedia material via a continuous, dynamic interaction between cognition and affect. the latter is rooted in environmental, contextual factors of the situation, and moderated by learner characteristics (e.g., meta-cognitive skills, motivation) and predispositions (e.g., expectations brought into the situation by the learner). we added the dashed arrows in fig. 1(a) to highlight that learning is not regarded as a unidirectional cause-and-effect relation. instead, it is viewed as a closed-loop process in which learners constantly engage in “meaning-making” (schwartz & plass, 2020). thus, in the icalm, the high sensitivity towards specific contextual factors influencing learning is an essential feature. the icalm is built upon the notion that learning occurs in specific contexts, driven by dynamic cognition-emotion interactions. these interactions themselves “emerge and operate in ways that are highly contextualized (and hence sensitive to contextual factors)” (schwartz & plass, 2020, p. 149). learning emerges from the interactions among cognition and affect, further providing a conceptual link to learning engagement and motivation. this is especially important as game-based pedagogies are frequently employed for their capabilities in fostering sustained engagement and motivation in learners (greipl et al., 2020; ryan & rigby, 2020). as outlined by urhahne and wijnia (2023), many, if not all, motivational theories share the same basic motivational model depicted in fig. 1(b), based on the general model of motivation by heckhausen and heckhausen (2018). similar to the icalm (plass & kaplan, 2016), the model emphasizes that motivated behavior emerges from interactions between learners and the situation in which they find themselves. all subsequent concepts to describe motivated behavior [goals, actions, outcomes, consequences in fig. 1(b)] emerge from this interaction and are evaluated in relation to it. motivational theories like expectancy-value theory (eccles & wigfield, 2020), self-efficacy theory (bandura, 1997), or self-determination theory (ryan & deci, 2019) then differ in how they conceptually fill, relate, and weigh the different parts of the motivation-action cycle against each other. however, they share their roots in this interaction between learners and the learning situation. in summary, both cognition and motivation should be sensitive to the exact conditions experienced during learning. as recognized by mayer (2024) in his implicit definition of boundary conditions quoted above, this has important implications for game-based pedagogies, as well as multimedia learning principles in general, which are expected to be sensitive to those exact conditions. this is suggested also by a plethora of empirical research as outlined next. figure 1 outline of typical cognitive and motivational models note. (a) simplified illustration of the integrated cognitive affective model of learning with multimedia (icalm) by plass and kaplan (2016). (b) basic motivational model proposed by urhahne and wijnia (2023). the dashed arrows were added by us in both illustrations, with the rationale for them explained in the text. 1.2 related empirical work while a variety of meta-analyses converge on the conclusion that game-based pedagogies can positively affect both cognitive and motivational learning outcomes, they also typically yield a high heterogeneity of effect sizes. point estimates for cognitive learning outcomes range from medium to large effect sizes, specifically hedge’s g = 0.46-0.85 (alotaibi, 2024; arztmann et al., 2023; bai et al., 2020; hu et al., 2022; huang et al., 2020; sailer & homner, 2020; q. zhang & yu, 2022). substantial heterogeneity between individual studies is frequently noted (e.g., hu et al., 2022, i2 = 86%; huang et al., 2020, i2 = 88%). reported point estimates for motivational outcomes result in an even wider range of small to large effect sizes, specifically hedge’s g = 0.25-0.92 (alotaibi, 2024; arztmann et al., 2023; fadda et al., 2022; hu et al., 2022; li et al., 2024; sailer & homner, 2020; q. zhang & yu, 2022) with, again, substantial heterogeneity between individual studies, i2 = 83% (arztmann et al., 2023; li et al., 2024). effect sizes in the scope of motivation can further vary substantially depending on which aspect or construct of motivation is under investigation. while li et al. (2024) obtained an overall effect size of g = 0.26 for intrinsic motivation, they obtain a large variation in effect sizes for perceptions of autonomy (g = 0.64, 95%-ci [0.14, 1.14]), relatedness (1.78, 95% ci [0.74, 2.81]), and competence (0.28, 95%-ci [0.00, 0.55]). the sensitivity of effect size to the exact subfactor of motivation is supported further by fadda et al. (2022), providing additional evidence that the specific type of motivational outcome moderates the obtained effect size. moderator analyses provide some indications about which underlying factors are driving the heterogeneity in effect sizes for both types of learning outcomes. at large, the suggested factors can be grouped into person factors and contextual factors. among person factors, education level (and hence, indirectly related, age) was repeatedly identified as a factor moderating learning outcomes (arztmann et al., 2023; sailer & homner, 2020). while gender was not a significant moderator in a meta-analysis (arztmann et al., 2023), the authors caution against over-generalization of that result given the limited number of studies in their meta-analysis. in fact, some empirical indications exist suggesting gender differences in game-based learning outcomes (conte, 2019; mavridis et al., 2017; rodrigues et al., 2022; l. zhang et al., 2024). on the other hand, evidence has accumulated regarding the role of contextual factors on learning and its outcomes within the framework of game-based pedagogies. one such contextual factor is that the mentioned meta-analyses represent a mix of investigations into different game-based pedagogies, such as gamification, game-based learning, or serious games. conceptual differences exist between those pedagogies (see, e.g., deterding et al., 2011; plass et al., 2020), potentially underlying some of the obtained empirical differences like gamification affecting extrinsic rather than intrinsic motivation, whereas the opposite holds true for game-based learning (q. zhang & yu, 2022). another important contextual factor is subject area (ritzhaupt et al., 2021). this is also endorsed by research in the strongly related field of learning analytics, and especially in predictive modeling of academic achievement (alyahyan & düştegör, 2020; prenkaj et al., 2021), where it has been found that predictive models produce substantially different results across different course topics (conijn et al., 2017). that is, the transferability of predictive modeling was found to be low, which means that a model, for instance, capable of accurately predicting grades in a mathematics course could perform poorly in a language course and vice versa. further, it was found that generalized models performed worse than course specific models. recent research also showed that different course-specific learning variables (e.g., learning design, student sample composition, or class size) could affect the prediction of outcomes differently (xu et al., 2024). this is in line with the notion that learning outcomes in educational research might depend considerably on the sampled population, which is often comprised of undergraduate university students (ritzhaupt et al., 2021). finally, the research context in which learning is studied has been noted as an influential factor (huang et al., 2020; sailer & homner, 2020). whereas huang et al. (2020) found that learning outcomes reported in dissertations and theses differed significantly from outcomes reported in journal articles and conference proceedings, sailer and homner (2020) noted significant differences in outcomes between quasi-experimental and experimental research designs. sailer and homner (2020) suggested differences in methodological rigor as a possible moderator of effect sizes to explain part of those discrepancies. however, another possible explanation for this finding could be, in part, that meta-analyses typically combine studies that differ widely in many aspects (e.g., population, methodology, etc.; borenstein et al., 2009), which are not yet identified as important determinants of learning outcomes and which can only be addressed by moderator analyses to some extent. to give another, additional perspective, learners’ engagement is considered an essential prerequisite of learning (chi & wylie, 2014; fiedler & beier, 2014; wong & liem, 2022). however, engagement is, like cognition and motivation, a multi-faceted psychological construct influenced strongly by context and timescale (sinatra et al., 2015). that entails that an understanding of engagement depends also on the perspectives taken by the researcher, which can be categorized into person-oriented, context-oriented and person-in-context studies according to the taxonomy proposed by sinatra et al. (2015). however, in comparison to person-oriented and person-in-context oriented research, context-oriented studies seem relatively scarce (booth et al., 2023). moreover, context-oriented studies typically consider the context only at the group level, for instance, investigating if students, on average, learn more efficiently with instructional videos in online or in blended learning courses (seo et al., 2021). in contrast, person-in-context studies usually investigate how changes of particular features within a learning task affect learners, for instance, via emotional or gamified design of certain features (greipl et al., 2021; ninaus et al., 2019). that is, the term context in the person-in-context category of sinatra et al. (2015) is rather narrowly defined. what is mostly lacking, to our best knowledge, are studies investigating how context properties (in the broader sense of context-oriented studies) affect the influence of features used typically in person-in-context studies. that is, e.g., how contextual features affect the effectiveness of gamified task design for different learning outcomes. this is precisely what mayer’s notion of boundary conditions is aiming at and what we are addressing in the present work, as outlined next. 1.3 present work in the present work, we take the opportunity to (re-)assess and compare the results of three studies using the same associative learning task to study the effect of gamified task design in different contexts. that is, all three studies were originally construed as individual, person-in-context studies aiming to investigate the effectiveness of gamified task design on cognitive and motivational learning outcomes. by combining studies across different contexts in the present work, we can shed some light on how cognitive and motivational outcomes are influenced by different contexts during gamified learning. by combining the different results, we end up with a context-oriented analysis to explore possible boundary conditions of gamified learning. by reassessing and comparing the same cognitive learning and motivational outcomes under three different, contextual conditions realized in the three respective studies, we aim to improve our understanding of gamified learning at two different methodological levels: 1. to what extent do cognitive learning and motivational outcomes differ based on the contextual conditions realized in each of the compared studies? 2. are there especially important contextual features that give rise to boundary conditions for the efficacy of gamified learning for cognitive learning and motivational outcomes? since the first two studies in our present work are re-analyses of already-published research (huber et al., 2023, 2024), we provide short summaries of the respective two studies below. study 1. in study 1, huber et al. (2023) explored whether gamifying an associative learning task via a specific set of game elements would affect learners’ engagement more than a non-gamified learning task. the study focused on attrition, a subcomponent of behavioral engagement, in an online learning environment. by comparing the gamified with the non-gamified version of the learning task, game elements were found to reduce dropout rates and enhance engagement. while 1688 persons accessed the landing page of the study, only 685 (about 40%) proceeded to sign the online consent form, and from there, only 312 (about 18%) finished the learning task with significantly more persons (δn = 50) dropping out with the non-gamified version than the gamified version of the task (δn = 23), χ2(1) = 12.28, p < .001. game elements also improved learning efficacy and efficiency indirectly through their ability to increase task attractivity. both the task’s attractivity and the perceived stimulation from the task – used to assess learners’ experience of the learning task – were positively associated with the motivational subcomponents of intrinsic motivation and identified regulation. conversely, they were negatively correlated with the subcomponents of external regulation and amotivation. study 2. in study 2, huber et al. (2024) explored relations between game elements, learners’ experience and motivation, and cognitive learning outcomes. the study aimed to replicate the positive, indirect association between gamified learning and cognitive learning outcomes through learner experience and motivation, as noted in study 1, using a small convenience sample (via a university course) and a slightly more difficult version of the associative learning task. the study was again conducted online and compared two versions of the learning task, a gamified one and a non-gamified one. the results showed no significant differences in cognitive learning outcomes, but medium and large positive effects on motivational outcomes and learner experience, respectively. mediation analysis revealed that while game elements slightly hindered cognitive outcomes directly, this was offset by a stronger positive influence through increased motivation and learner experience. learner experience was further shown to be positively associated with two subcomponents of intrinsic motivation, namely perceived interest and competence. study 3. study 3 was conducted recently and will be described in detail in the present work. 2. material and methods three datasets were collected from studies conducted at two different austrian universities, guided by a value-added research design according to mayer's taxonomy (mayer, 2019, 2020). that is, participants (section 2.1) were randomly assigned to one of two learning task versions in each study within an overall similar experimental procedure (section 2.2). the two different learning task versions consisted of a control version (i.e., non-gamified = without additional game elements) and a gamified version (section 2.3). various learning outcome measures, including cognitive and motivational outcomes, were assessed after the learning task (section 2.4). the results obtained in the two conditions (i.e., the two task versions) were then compared and differences were quantified using statistical analyses (section 2.5). all studies were approved by the respective local universities’ ethics committees. although all studies were similar to each other, they also differed in several respects. this will be clarified in the following sections. table 1 provides an overview of features by which the compared studies differed. except for one task-inherent feature (task difficulty and directly related to it: time-on-task) all features in table 1 refer to the broader study context and are not inherent to the learning task common to all three studies. 2.1 participants table 2 summarizes the demographic information available for the participants of the three studies. regarding demographic information, age, gender, and whether participants were currently enrolled university students, were assessed in all three studies. considering student status, we did not differentiate between graduate and undergraduate students. furthermore, the information was entirely based on self-report, that is, we did not verify with the universities about participants’ student or enrollment status. study 1 . the first of the three compared studies was conducted online with a data acquisition period spanning about 7 weeks from mid-february 2021 to the beginning of april 2021. participants were compensated by being given the option to enter a raffle for five vouchers, each worth 10 eur for an online retailer, which were awarded at the conclusion of data collection (huber et al., 2023). the compensation was very low on purpose to reduce the influence of incentive bias on engagement, especially dropout, which was the research focus of the study. however, some compensation was necessary to improve response rates, which can be especially low in online studies that exceed 10 minutes in duration (sammut et al., 2021). at the same time, the study was broadly advertised using university-wide listservs of the local university to acquire a sufficient number of participants allowing a statistical analysis of attrition over the course of the study. due to this primary aim of the study, many participants dropped out over the course of the learning task (n = 244 overall and significantly more during the non-gamified than the gamified task version; for details, see huber et al., 2023) and hence, never advanced to the questionnaires addressing motivational outcomes. as a result, in this paper we reassess only the data from those n = 285 participants who completed each phase of the study. study 2. the second of the three compared studies was conducted online. in this case, however, participants were acquired during a shorter, about 4-week period in may 2022 within the framework of a university course on empirical research at a different university than in the case of the first study, see table 1. furthermore, participants were compensated by course credit, whereas no course credit could be obtained in the first study. the rationale for changing compensation from study 1 was to increase the chance that participants would stay engaged throughout the study in both gamified and non-gamified conditions, as this was important for an unbiased quantification of the motivational effect of gamification on learning. therefore, the incentive to stay engaged in the learning task was increased by providing course credit for students enrolled in the psychology program, which reduced dropout rates overall (n = 35), and eliminated the difference in dropout rate between gamified and non-gamified conditions (huber et al., 2024). although again university-wide listservs were used for participant recruitment and course credit was provided for participating psychology students, only n = 61 participants completed the study, whose data will be reassessed in the present study. study 3. the third study was conducted on-site, in a laboratory at the same university at which study 2 was conducted. in this case, the data acquisition period lasted about 13 months from the end of april 2023 to beginning of june 2024. the reason for this long data acquisition period was related to the requirement of recruiting a minimum of 120 participants to be able to replicate the motivational effect of gamified learning found in study 2 with sufficient statistical power assuming a medium-to-large effect size. in this study, participants were compensated in the form of course credit for students enrolled in the psychology program in addition to the option to enter a raffle for five vouchers, each worth 50 eur for an online retailer, awarded at the conclusion of data collection. a total of n = 121 participants were recruited, again using university-wide email broadcasts, in addition to word-of-mouth and by advertising the study across social networks. 2.2 procedure in all three studies, participants were administered several questionnaires before and after the learning task to obtain demographic information and motivational outcomes, in addition to other data not discussed in this work. in the present work, we focus only on the cognitive learning and motivational outcomes that were commonly used in all three studies to explore the influence of the different contextual features on the same cognitive and motivational outcomes of gamified learning. these will be described in detail in section 2.4. in the case of the laboratory study (study 3), we further recorded different physiological sensor data (electrodermal activity, heart rate, eye-tracking, facial expression analysis) over the course of the learning task. such data were not recorded in the two online studies (study 1 and study 2). this also means that in the case of study 3, some additional time before and after the learning task was required to attach and detach the respective sensors. table 1 different contextual features realized in the three studies table 2 demographic information on the three study samples note. sd = standard deviation. mad = median absolute deviation adjusted for asymptotically normal consistency. all variables referring to age are given in units of years. 2.3 learning task all studies used an associative learning task in which participants needed to learn associations between symbols and numbers over the course of five levels of the task. at the beginning of the first level of the task, a symbol was presented to a participant in the upper left corner of the computer display (see fig. 2). the participant then had to guess which number, presented on a number line at the bottom of the interface, would be associated with that symbol. the selection was made by moving a slider with the keyboard’s arrow keys to the desired position on the number line and confirmed by pressing the spacebar. after participants made their selection, corrective feedback was provided. that is, the correct number was always indicated for the given symbol with a green vertical bar to allow for learning of the association between symbol and number. this procedure was repeated for all remaining symbols. note that at the first level, participants could only guess which number would correspond to which symbol. that is, the purpose of the first level was entirely to present all the associations once to the participants. in the second level, all symbols would be shown again subsequently. at this point, the participants would typically be able to remember already some of the associations but would still receive corrective feedback for each symbol, allowing them to learn the rest. the overall goal of the learning task was to learn as many associations as possible over the course of five levels. in each level, symbols were presented in random order. selecting a number had to take place within 20 seconds after each symbol was displayed; otherwise, corrective feedback would be provided without having selected a number. study 1 differed from studies 2 and 3 with respect to the total number of symbol-number associations that needed to be learned by the participants. this number was 14 in study 1 (huber et al., 2023), while it was increased to 20 in study 2 (huber et al., 2024) and was kept at 20 also in study 3 (present work). hence, the difficulty of the learning task was also somewhat lower in the first of the three studies (see table 1). the choice to increase the number of symbol-number associations was justified in study 2 (huber et al., 2024) by the aim to increase the discriminatory power of the cognitive outcome measures and mitigate ceiling effects at the fifth level of the task reported in study 1 (huber et al., 2023). however, the increase of presented symbols per level also affects the time participants engage with the task (time-on-task). in table 1, we provide approximate minimum and maximum times participants engage with the task based on the task’s inherent mechanics. figure 2 non-gamified and gamified learning task versions note. screenshots of the learning task in the (a) non-gamified version and (b) in the gamified version when an incorrect answer was provided. corrective feedback was always provided, regardless of correct or incorrect responses. if responses were correct, a green check mark was provided in the non-gamified version of the task whereas the bone count was increased by 1 and the dog wagged its tail and held a bone in its mouth in the gamified version of the task (not shown here). if responses were incorrect, a red x-sign was provided in the non-gamified version of the task (see panel a) whereas the dog cried in the gamified version of the task (see panel b). regarding the design of the learning task, all three studies used two identical design versions, one employing some additional game elements (= gamified version), one representing a control version without those game elements (= non-gamified version). screenshots of the two versions are provided in fig. 2. the difference between gamified and non-gamified task versions was entirely based upon the inclusion of game elements, a core aspect also in game-based learning (plass et al., 2015). the game elements comprised of (i) visual aesthetics (a scenery with trees, sky, mountains, etc.), (ii) a narrative of a dog taking a stroll in the woods looking for bones, and (iii) an incentive system (i.e., the “bone count” displayed in fig. 2(b) in the upper left corner). in the gamified version, the slider’s movement and placement were also accompanied by animations of the dog walking and digging (searching for a bone), respectively. choosing a correct number would result in the dog wagging its tail in the gamified task version, while in the non-gamified task version, only a green check mark would appear. in the gamified task version, choosing an incorrect number would result in the dog shedding tears, while a red x-sign would appear in the non-gamified task version (see also fig. 2). both versions of the learning task were created using the numbertrace game engine [see e.g., koskinen et al. (2023); a short video demonstration may further be viewed at the following url:doi: https://www.youtube.com/watch?v=t7s7xsllrac which was originally developed for fraction instruction using javascript. 2.4 outcome measures regarding cognitive learning outcomes, we considered both learning efficacy and efficiency in all three studies. efficacy refers to how many associations participants had learned at the end of the learning task, that is, in level 5 of the task. because studies differed in the total number of symbol-number associations, see section 2.2, we refer to the proportion of correctly learned associations to compare studies with each other with respect to efficacy. this proportion is given by the number of correctly learned associations in level 5 divided by the total number of symbol-number associations in the respective study (i.e., 14 in study 1, 20 in studies 2 and 3), resulting in a number between 0 and 1 for the learning efficacy of each learner. efficiency refers to how fast participants learned symbol-number associations. originally, in the two online studies (i.e., studies 1 and 2), the authors fitted an exponential learning curve to the number of correct responses per level for each individual participant, and subsequently analyzed the rate constants of these learning curves (huber et al., 2023, 2024). in the present work, we follow a simpler approach by simply adding the numbers of correct responses per level for levels 2 to 5. the rationale behind this approach is that if a participant learns very efficiently, they approach the maximum number of correct responses very fast and the sum of correct responses per level corresponds to a large number. participants learning less efficiently approach their final performance at the fifth level slower and the sum of correct responses corresponds to a lower value. by dividing this sum score again by the maximally possible value (56 for study 1, 80 for studies 2 and 3) we again obtain a number between 0 and 1 for the learning efficiency of each learner. regarding motivational outcomes, studies 1 and 2 used different questionnaires to assess specific facets of motivation. however, both studies complemented these questionnaires by the subscales attractivity and stimulation of the user experience questionnaire (ueq; laugwitz et al., 2008; schrepp et al., 2017). although not targeting motivation at face value, both studies provided evidence that these subscales are associated with the construct motivation. study 1 (huber et al., 2023) showed that attractivity and stimulation were strongly and positively associated with intrinsic motivation (ρpb = 0.69 and ρpb = 0.66, respectively, using the percentage bend correlation described, e.g., by wilcox (2022)). further, identified regulation was moderately and positively associated with attractivity and stimulation (ρpb = 0.45 and ρpb = 0.48, respectively), and somewhat, negatively associated with external regulation (ρpb = -0.25 and ρpb = -0.18, respectively). a strong negative association was found with amotivation (ρpb = -0.49 and ρpb = -0.50, respectively). these facets of motivation were measured via the situational motivation scale by guay et al. (2000) in study 1. study 2 (huber et al., 2024) further revealed that the two subscales of the ueq were positively associated with the subscales interest (ρpb = 0.32 and ρpb = 0.29, respectively) and competence (ρpb = 0.78 and ρpb = 0.82, respectively) of the short scale for intrinsic motivation (wilde et al., 2009). hence, we use these two subscales of the ueq (i.e., attractivity and stimulation) in study 3 to compare the results of studies 1 and 2 regarding motivational outcomes with the results obtained in study 3 and, as such, as a proxy for motivation. both the attractivity and the stimulation subscale of the ueq (laugwitz et al., 2008; schrepp et al., 2017) consist of several items, each presenting two opposing adjectives, forming the endpoints of a 7-point, bipolar rating scale (i.e., from -3 to 3 in steps of 1), on which participants indicate their experience of the task. in particular, attractivity aims to assess the general appeal of a task (or product) by asking, for instance, how “enjoyable” it is perceived (with the poles: “annoying”/“enjoyable”), or “pleasing” (with the poles: “unlikable”/“pleasing”). the subscale uses six items in total. stimulation assesses the capability of a task (or product) to motivate or captivate by asking, for instance, how “interesting” (poles: “not interesting”/“interesting”), or “motivating” (poles: “not motivating”/“motivating”) it is perceived. the subscale uses four items in total. all studies yielded good to excellent reliabilities for the two subscales attractivity and stimulation of the ueq. in particular, we obtained cronbach’s α = 0.92 and α = 0.89, respectively, for study 1, α = 0.93 and α = 0.91, respectively, for study 2, and α = 0.89 and α = 0.82, respectively, for study 3. 2.5 statistical analyses for statistical analyses, we employ bayesian hypothesis testing and estimation as provided by jasp (version 0.19.1.0; jasp team, 2024). in the present work, we apply bayesian inference because it provides some advantages over classical inference (wagenmakers et al., 2018) that are important for the present work. first, bayes factor analysis can quantify the degree of evidence that the data support the null model (with a fixed parameter) compared to alternative models (with a continuous range of possible parameter values that are estimated through conditioning on the data). second, bayesian inference is not biased against the null model. third, it does not depend on sampling plans varying between studies. while our research questions (section 1.3) fall into the area of estimation, the re-assessed previous studies (huber et al., 2023, 2024) have shown that gamified learning can have insignificant effects on cognitive learning outcomes, which would be reflected in high compatibility between a null model (assuming an effect of exactly zero) and the data. this implies that investigating the dependence of respective bayes factors on contextual features across the three studies is informative also for investigating our present research questions. for estimating the cognitive and motivational effects of gamified learning and comparing these effects across the three studies, bayesian estimation provides advantages over classical inference because it can gradually quantify the degree of confidence where effect sizes lie in a specific interval while conditioning on what is factually known (wagenmakers et al., 2018). all respective analyses and data, are available at the open science framework (osf), see https://osf.io/snvw8 and https://osf.io/d97bp , respectively. the specific context of each study was included in each model as a categorical predictor (denoted as “boundary” in files provided at osf) with three categories corresponding to studies 1-3. to assess the effect of contextual factors on cognitive learning and motivational outcomes, we compared bayes factors resulting from bayesian anovas for the following five models for each outcome variable (efficacy, efficiency, attractivity, stimulation): 1) a null model (containing neither task version nor context as predictors); 2) two single predictor models with either task version or context as predictor; and 3) two models including both predictors once without their interaction and once with their interaction included. these models are denoted below simply as “null”, “version”, “context”, “version + context” and “version * context”, respectively. prior odds were assigned uniformly over all five models (i.e., prior odds of 20% for each model). models were compared to the best resulting model in our analyses. all models included age, gender, and student status as random factors to provide some control for their influence because we noted that samples differed across studies (see table 2), and multimedia learning theory suggests that these factors play a role too for cognitive learning and motivational outcomes (mayer, 2024). bayesian anovas were followed up by pairwise comparisons of results obtained with gamified and non-gamified task versions for each of the three studies. for these comparisons, first the assumptions about normality of outcome variables and equality of variances for gamified and non-gamified task versions were checked with shapiro-wilk and brown-forsythe tests, respectively. based on those tests, either jasp’s bayesian implementations of student’s t or mann-whitney tests were used for pairwise comparisons. for each comparison, default cauchy priors centered on zero and with an interquartile range of 1/√2 were used. if not stated explicitly otherwise, credible intervals refer to 95% percentile intervals and are reported directly after parameters in parenthesis. for the pairwise comparisons, null hypotheses always referred to typical nil hypotheses, that is, an exact equality between outcome variables for gamified and non-gamified task versions, whereas alternatives referred to allowing outcome variables to be unequal. 3. results 3.1 effects of context on cognitive learning and motivational outcomes in fig. 3, we depict the means and their 95% credible intervals for (a) efficacy, (b) efficiency, (c) attractivity, and (d) stimulation depending on context and task version. we find that for efficacy and efficiency, neither context nor task version seem to matter much, because all credible intervals yield substantial overlap with each other. the data does not provide any evidence for an interaction between context and task version, either for efficacy or efficiency (see tables 3-4), respectively. in fact, for efficacy (table 3), the best model is the null model that does not include any of the two predictors. except for the single predictor model with task version as a predictor (2.6 times less likely than the null model according to our data), the data provides strong to decisive evidence against all other models. the same is the case for efficiency (table 4), in which case again the single predictor model with task version as predictor and the null model are most likely from all considered models and for the given data, providing strong to decisive evidence against all other models. overall, we conclude that for predicting efficacy and efficiency in the used learning task contextual aspects do not contribute much if anything at all. regarding our first research question (section 1.4), we thus conclude that cognitive learning outcomes resulting from this task do not differ much between the different contexts realized by the three experimental settings. regarding our second research question, the results further suggest that the influence of gamification on cognitive learning outcomes does not depend much on those contextual aspects. figure 3 means and 95% credible intervals for the considered outcome variables note. means and their 95% credible intervals for (a) efficacy, (b) efficiency, (c) attractivity, and (d) stimulation for the three considered studies (from left to right in each panel) and each task version (non-gamified: open circles; gamified: full circles). instead, in the case of motivational outcomes, the depicted means and their 95% credible intervals in figs. 3(c) and (d) indicate that these outcomes depend on contextual aspects. in the case of attractivity (table 5), both two predictor models (i.e., the two models with two predictors, either including or excluding the interaction term), turn out best. in particular, the data allow no clear distinction between both two predictor models, yielding almost identical posterior probabilities. however, against all other models, moderate to strong evidence is provided by our data. in particular, the null model is more than 40 times less likely than both two predictor models. the evidence against the null model is even stronger in the case of stimulation, with the model being more than 100 times less likely than any model, including context as a predictor. overall, this means that for motivational outcomes the data suggest that contextual aspects seem to matter substantially. regarding our first research question (section 1.4), we thus conclude that motivational outcomes resulting from this task differ substantially depending on the different contexts realized by the three experimental settings. regarding our second research question, the relative importance of the interaction between task version and context for predictive quality of the resulting model further suggests that the influence of gamification on motivational outcomes also depends somewhat on those contextual aspects. to resolve this contextual dependence better, pairwise comparisons between gamified and non-gamified task versions for each study context are presented next. table 3 model comparison for efficacy note. all models include age, gender, and student status. models are always compared against the best model (denoted by 0 in the bayes factor) in ascending order. table 4 model comparison for efficiency note. all models include age, gender, and student status. models are always compared against the best model (denoted by 0 in the bayes factor) in ascending order. table 5 model comparison for attractivity note. all models include age, gender, and student status. models are always compared against the best model (denoted by 0 in the bayes factor) in ascending order. table 6 model comparison for stimulation note. all models include age, gender, and student status. models are always compared against the best model (denoted by 0 in the bayes factor) in ascending order. 3.2 effects of context on the influence of gamification on cognitive learning and motivational outcomes study 1. regarding study 1, we depict the prior and posterior distributions for the effects of gamification on all four considered outcome measures in fig. 4. the bayes factor analyses provide evidence for the null hypothesis in the case of efficacy (bf01 = 5.1) and stimulation (bf01 = 5.9), yet for the alternative hypothesis in the case of attractivity (bf10 = 4.1). the median effect size in the latter case is -0.30 (-0.53, -0.08), where the negative sign indicates that attractivity is, on average, higher for the gamified task version than for the non-gamified one. in the case of efficiency, the data are undecisive with respect to both hypotheses (bf10 = 1.4). still, more than 95% of the mass of the posterior distribution is located above zero, indicating that if there is some population effect on efficiency, then efficiency is probably somewhat larger for the non-gamified task version than for the gamified one. the median effect size in that case is 0.24 (0.02, 0.48). figure 4 effects of gamification in study 1 note. prior and posterior distributions for effects of gamification for (a) efficiency, (b) efficacy, (c) attractivity, and (d) stimulation in the first online study. note that a negative value of effect size means that the respective outcome is, on average, larger for the gamified than for the non-gamified task version. study 2. comparison of these results with those obtained for study 2, depicted in fig. 5, suggests that study context might matter for the influence of gamification, especially for motivational outcomes. in particular, we find evidence for the null hypothesis in the cases of both efficacy (bf01 = 3.1) and efficiency (bf01 = 3.5). for attractivity, we obtain strong evidence in favor of the alternative hypothesis (bf10 = 10.4) with a medium median effect size of -0.68 (-1.21, -0.18), indicating the gamified task version to be considerably more appealing to the participants than the non-gamified task version. for stimulation, we also obtain evidence in favor of the alternative hypothesis (bf10 = 8.1), yielding a medium median effect size of -0.66 (-1.18, -0.16), indicating the gamified task version to raise more interest and motivation in the participants than the non-gamified task version. study 3. strong evidence for the existence of boundary conditions for the effects of gamification is found by comparing study 2 with study 3. note that the two studies are closest to each other out of all three possible pairs of studies not only by design, but also by the resulting participant samples, see section 2.1 and table 1. nevertheless, in contrast to study 2, the laboratory data are undecisive regarding both hypotheses for both cognitive learning outcomes, that is, efficacy, bf01 = 1.4 and efficiency, bf01 = 1.1. in addition, the results yield posterior distributions in both cases with almost 90% of their probability mass above zero. that is, if gamification affects efficacy and efficiency at all, then, according to the data, it is more reasonable to assume that they are higher, on average, for the non-gamified task version than for the gamified one. the respective median effect sizes are 0.29 (-0.08, 0.65) and 0.30 (-0.03, 0.66). most importantly and in stark contrast to study 1, study 3 provides evidence for the null hypothesis in the case of both motivational outcomes, attractivity, bf01 = 5.2, and stimulation, bf01 = 5.0. the median effect sizes account for practically negligible to small effects yielding -0.01 (-0.35, 0.33) for attractivity and -0.04 (-0.38, 0.30) for stimulation. figure 5 effects of gamification in study 2 note. prior and posterior distributions for effects of gamification for (a) efficiency, (b) efficacy, (c) attractivity, and (d) stimulation in the second online study. note that a negative value of effect size means that the respective outcome is, on average, larger for the gamified than for the non-gamified task version. figure 6 effects of gamification in study 3 note. prior and posterior distributions for effects of gamification for (a) efficiency, (b) efficacy, (c) attractivity, and (d) stimulation in the laboratory study. note that a negative value of effect size means that the respective outcome is, on average, larger for the gamified than for the non-gamified task version. 4. discussion 4.1 main findings in this paper, we focused specifically on the concepts of context and environment as part of the broader conceptualization of boundary conditions in multimedia learning (mayer, 2024). we aimed to identify the contextual features, in which the three compared studies differed, that give rise to boundary conditions influencing the effect of gamified learning on cognitive learning and motivational outcomes. our findings indicate the existence of such boundary conditions for motivational outcomes, whereas cognitive learning outcomes are influenced by the considered contextual features, if at all, only to a substantially smaller extent. this result is supported by theory as well as meta-analyses and systematic reviews on gamification and game-based learning, consistently showing that employing game elements can have a beneficial effect on both motivational and cognitive outcomes, but revealing also a substantial heterogeneity of effect sizes across studies (arztmann et al., 2023; barz et al., 2023; hu et al., 2022; huang et al., 2020; lampropoulos & kinshuk, 2024; ritzhaupt et al., 2021; ryan & rigby, 2020; sailer & homner, 2020; schlag et al., 2024; wouters et al., 2013). this result is further aligned with plass and kaplan’s (plass & kaplan, 2016) integrated cognitive affective model of learning with multimedia (icalm), which specifically describes that there is a high sensitivity towards contextual factors on cognitive processing and motivation when learning with multimedia. the most obvious contextual difference between the three considered studies is that study 3 was carried out in the laboratory, whereas studies 1 and 2 were carried out online. this difference suggests at least two possible explanations for the substantial motivational difference between studies 1 and 2 and study 3. both explanations are based on differences in expectations and norms between those contexts. first, the motivational difference between the laboratory (study 3) and online studies (studies 1 and 2) might stem from participants coping with cognitive dissonance (harmon-jones & mills, 2019) through reevaluation or self-justification. in the laboratory setting, participants invest considerably more effort compared to an equivalent online setting—they must register, schedule an appointment, travel to the lab, and navigate a social situation with stricter norms and expectations (hoerger, 2010; hoerger & currell, 2012; skitka & sargis, 2006). given this level of effort for completing a relatively unappealing task (i.e., memorizing number-symbol associations), participants may experience cognitive dissonance, making it difficult to justify the experience to themselves. to resolve this discomfort, they might retrospectively reassess the task in the laboratory as less unappealing than it initially seemed. in other words, to avoid cognitive dissonance (likely more prominent with the non-gamified task version) participants may reevaluate the task's attractiveness and their interest in it more favorably after its completion in the laboratory. this is indeed supported by our findings revealing that the non-gamified task is evaluated as more appealing and stimulating in the laboratory than online, whereas the gamified tasks are not evaluated considerably different between those two settings, see fig. 3. however, another explanation considers the possibility that both the non-gamified and gamified task versions are not merely evaluated differently post-hoc but may be genuinely experienced as equally motivating in the laboratory setting. this could be understood as a form of a figure-ground effect. hypothetically, the controlled laboratory conditions (amplified eventually by application of sensors and being monitored), in addition to the absence of other, readily available action options, which might enhance participants' focus on the tasks, possibly reduce the distinction between the gamified and non-gamified versions in the laboratory. in other words, in the laboratory, the non-gamified task version might be experienced as appealing and interesting, given the sociallyand contextually-amplified attention that it receives in that setting, in addition to the simple fact that there is nothing else to do. this is aligned with prior research which has repeatedly shown that participants underestimate the enjoyment of supposedly boring tasks if experienced in the context of a scientific experiment (hatano et al., 2022; kuratomi et al., 2023). in contrast, in the online environment, participants may take part in the study considerably more spontaneously, and more likely than not, also have other action options at their disposal. under these circumstances, the non-gamified task possibly starts out with less appeal to the participants (relative to the laboratory setting) and, in addition to that, is perceived as even less appealing relative to other, alternative action options. given such options, bernecker and ninaus (2021) showed that gamifying a strenuous and tedious working memory task can indeed reduce disengagement and the motivational conflict to do something else. that implies that there might be much more to gain for gamification regarding an influence on learners’ motivation in an online setting than in the laboratory. the present study also tentatively supports that the gamified and non-gamified task versions were truly equivalent regarding their motivational outcomes in the laboratory. to explicate this, the more subtle influence of contextual aspects on cognitive learning outcomes needs to be considered in conjunction with their influence on motivational outcomes. it was argued already in both earlier online studies (huber et al., 2023, 2024), that the effects of game elements on cognitive learning outcomes are aligned with what would be expected from cognitive load theory (sweller, 2011) or cognitive theory of multimedia learning (mayer, 2019, 2024). in general, the additional design elements employed in the gamified task version included more stimuli requiring additional cognitive processing, but non-essential for completion of the task. in addition, the simplicity of the task may not leave much room for generative processing (or germane load). however, even if they represent merely extraneous processing demands, ignoring and discarding the game elements as irrelevant yet puts some demand on information processing in the gamified task version, which is absent in the non-gamified one. in fact, our resulting estimates for effect sizes regarding cognitive learning outcomes show that the difference between gamified and non-gamified task versions is, descriptively, largest in the laboratory study (fig. 6), smallest in the second online study (fig. 5), and somewhat in between in the first online study (fig. 4). furthermore, the sign of the estimates indicates that in the laboratory and the first online study (study 1), the game elements are cognitively demanding in the sense that cognitive outcomes are slightly lower in the gamified than in the non-gamified task version, whereas the opposite is the case for the second online study (study 2). that is, if some cognitive effect of gamification is assumed at all, which is a reasonable assumption based on cognitive learning theories (mayer, 2019, 2024; sweller, 2011), then our analyses suggest that the influence of gamification on cognitive learning outcomes is most detrimental in the laboratory where the two task versions are most similar in motivation, and less detrimental in online settings where some motivational difference exists between the two task versions. however, this pattern of cognitive and motivational differences between task versions and contexts can be entirely explained by the indirect positive effect of game elements on cognitive learning outcomes through enhanced motivation, as previously suggested in both earlier online studies (huber et al., 2023, 2024). as outlined in the icalm framework (plass & kaplan, 2016), game elements can boost motivation, which in turn positively influences cognitive processing, which can explain the subtle differences in cognitive learning outcomes observed across the varying study contexts. in study 3, the motivational difference between gamified and non-gamified task versions was nearly absent. hence, in theory, only the presumed higher cognitive demand due to the additional game elements remains, and in fact, the difference in cognitive learning outcomes turns out largest. the other way around, the motivational difference between gamified and non-gamified task versions was largest in study 2. here, in theory, the positive indirect effect of game elements via motivation is supposedly the largest among the three studies and can compensate for the negative, direct effect due to cognitive demand. in fact, in this case, cognitive learning outcomes are, across all three studies, mostly in favor of a null difference between gamified and non-gamified task versions, descriptively favoring the gamified version over the non-gamified one, see fig. 5. although the most obvious, the difference between laboratory and online settings is not the only feature in which the three studies differ, see table 1. hence, alternative mechanisms could underlie the noted differences in learning outcomes. another difference between the three studies that needs to be addressed is task difficulty. task difficulty, and in direct consequence also the time spent on the learning task, was lower in study 1, requiring participants to learn 14 symbol-number associations, than in studies 2 and 3, both requiring participants to learn 20 symbol-number associations. clearly, task difficulty can affect learning outcomes and a quadratic relationship between difficulty and both cognitive and motivational outcomes can be expected based on conceptual grounds (greipl et al., 2020). according to the limited capacity principle (mayer, 2011, 2014) cognitive resources are limited and can hence be overstressed and lead to cognitive overload. perceiving a task as too hard may then decrease learners’ engagement (baten et al., 2020) and motivation (sailer & sailer, 2021). interestingly, engagement also decreases for tasks perceived as too easy (baten et al., 2020), perhaps related to experiences of boredom and amotivation (engeser & rheinberg, 2008). hence, a part of the motivational difference between study 1 and study 2 might also be due to the difference in task difficulty between those two studies, providing more potential for the effectiveness of game elements in the case of a task that affords more engagement in the first place. some support for such a mechanism is provided by the fact that in both studies 2 and 3 again about a third of all participants (study 2: 34%; study 3: 32%) reached maximum efficacy at the final task level not differing much from study 1 (40%) in that respect. moreover, in all three studies about half of the participants manage to correctly recall all, or all but one, symbol-number associations at the final level (study 1: 54%; study 2: 51%; study 3: 45%). overall, the empirical distribution functions of all three studies are surprisingly similar given the clear difference in difficulty between study 1 and studies 2 and 3 (more details are provided in the respective analysis files provided at osf, see section 2.5). hence, it seems that the increased difficulty has indeed been met with increased engagement on behalf of the participants to some extent. while this may explain some part of the motivational difference between study 1 and 2, it can hardly explain the substantial difference in motivational outcomes between studies 2 and 3 which had the same difficulty level. another difference between all three studies concerns the equivalence of the sampled populations. individual differences have been noted as a potential factor influencing the effectiveness of gamification (ritzhaupt et al., 2021). while accounting for gender, age and student status as covariates may provide some control for interindividual differences, further person characteristics that may make a difference (for instance, prior knowledge, learning strategies, or generally, meta-cognitive abilities, see e.g., azevedo & wiedbusch, 2023) cannot be excluded. in fact, table 1 shows that study 1 differs substantially from studies 2 and 3 regarding the student status of the participants. being a student is not only related to how well one is currently attuned to learning or performing well in cognitive memory tasks but is also closely related to the appeal of the incentives provided for participation (e.g., course credit). however, these incentives also differed somewhat between the studies. incentives are a form of external regulation, which itself is a type of extrinsic motivation, and known to affect the effect of task design elements on learning (wesenberg et al., 2025). hence, incentives potentially influence the efficacy of gamified learning via several pathways. together with additional contextual features like the geographic location and the data acquisition period of the study (table 1), they first affect selection of participants and hence, the composition of the sample. second, they may directly affect the efficacy of gamified learning. while we do not think that the differences in incentives can fully explain the reported differences in learning outcomes between the three studies, more targeted future studies will need to further disentangle the combined effect of the various contextual and individual features on the efficacy of gamified learning, particularly for motivational outcomes. while our present study warrants further investigation, future studies should take special care about varying only one contextual feature at a time. a typical criticism of meta-analyses is that they often compare apples with oranges (i.e., studies that vary considerably in design, methods, contexts, etc.; borenstein et al., 2009). whereas our present study may represent some progress in that regard by comparing apples with pears instead, dedicated research that actually compares apples with apples will be required to sort out those contextual features that are especially relevant for the efficacy of gamified learning from those that do not matter much if at all. however, the importance of identifying such boundary conditions goes beyond the specific case of gamified learning considered in the present work. a rich body of literature exists showing that basic cognitive functions like memory (neath & suprenant, 2003) can be highly sensitive to subtle changes in contextual factors resulting in environmental context-dependent memory effects (smith & vela, 2001). owing to the definition of learning as an “enduring change of the mechanisms of behavior” (domjan, 2010, p. 17, emphasis by us), memory is a crucial factor in learning per se. hence, at a fundamental level, learning is sensitive to the exact conditions under which it occurs, and research on boundary conditions of fundamental learning mechanisms like memory remains of utmost importance to this day (krauspe et al., 2025). however, the findings of our present study suggest that the sensitivity on specific contextual features might be even stronger in the case of affective (including motivational) than cognitive aspects of learning. this implies that future research should delve deeper into investigating the boundary conditions of learning interventions that focus particularly on learners’ motivation. 4.2 limitations the scope of contextual features considered in the present work and the conclusions it allows regarding boundary conditions of gamified learning have several limitations. first, our samples were represented mostly by university students, who likely do not represent the entire population of adult learners. in addition to that, studies differed in the fraction of participants who were not currently enrolled as students at a university. future studies will need to assess the personal background with more scrutiny to resolve the impact of sample composition or more rigorously control the population from which participants are sampled (e.g., only undergraduate university students versus primaryor secondary-school students). second, we solely relied on self-report instruments to measure motivational outcomes. this approach assumes that motivation was always conscious and accessible to the individual, which is sometimes not the case. some studies find that participants report inaccurate data for several reasons, including failure to accurately monitor their internal states (fryer & dinsmore, 2020). in a similar vein, motivation is a multi-dimensional construct, comprising of biological, physiological, social, and cognitive factors that influence behavior (fulmer & frijters, 2009). however, our research only measured a single dimension of motivation, which was itself only indirectly assessed via the two aspects of learner experience, that is, task attractivity and perceived stimulation. while both are at face value obviously associated with self-report measures of intrinsic motivation (wilde et al., 2009), reflected also in previously noted associations with intrinsic and extrinsic motivation scales (huber et al., 2023, 2024), their use as a proxy for intrinsic motivation was solely based on the pragmatic reason of being the only motivation-related measure common to all three compared studies and may not be optimal for providing a complete picture of the potential impact of contextual features on motivational outcomes. future studies should go beyond this limitation by at least complementing those measures with validated instruments designed to measure motivational constructs, resolving, for instance, also aspects of extrinsic motivation (e.g., guay et al., 2000), or subcomponents of motivation like autonomy or competence experiences (e.g., wilde et al., 2009). assessing the task and utility values of the learning task (e.g., gaspard et al., 2020) in its different design versions and contexts could give an additional perspective that provides a link between aspects of motivation and provided participation incentives on the basis of expectancy-value theory (eccles & wigfield, 2020). it can further not be excluded that participants may respond differently to self-report items in an anonymous online setting compared to a laboratory setting (hoerger & currell, 2012), which is on the one hand intricately related to our research questions but, on the other hand, goes beyond the specific scope of gamified learning. lastly, the three studies differed in more contextual features than being conducted in the laboratory or online as discussed in detail in the previous section. as outlined above, future studies will need to further disentangle the combined impact on learning outcomes of the various contextual features realized by the studies compared in the present study. nevertheless, it remains noteworthy what difference the change from an online to a laboratory setting under the most similar study conditions (studies 2 and 3) already made for the considered self-report and performance-based outcomes alone. informed by adjacent research domains (d’mello & booth, 2023), similar or even more substantial differences can be expected when going from the laboratory further into the field. 4.3 implications for building context-aware educational interventions our results emphasize that for an intervention to support learning as intended the context in which learning takes place needs to be explicitly considered. the findings are relevant to the educational community and present implications for designing context-aware educational interventions that benefit both cognitive learning and motivational outcomes with gamified learning tasks (sailer et al., 2024). based on a comparison of three empirical studies on gamified learning under varying contextual conditions, we subsequently illustrate what challenges and difficulties existing boundary conditions can impose on a proper understanding and appropriate interpretation of outcomes under general conditions. our results suggest that in an online setting, gamification may effectively motivate learners to stay engaged with a task. however, in a more controlled environment, perhaps closer to the general outline of a controlled laboratory environment, this influence might cease to have any effect. this suggests that learning interventions utilizing gamification should be context-specific rather than universally applied. we envision that context-aware interventions should be designed to dynamically adapt based on real-time feedback from learners and the learning environment. using data collected in-situ may provide a means to provide interventions that can be adjusted on-the-fly to maintain engagement and optimize cognitive learning and motivational outcomes, ensuring they remain effective across different boundary conditions. furthermore, different learners may respond differently to the same interventions depending on the context, highlighting the need for personalization. context-aware interventions should incorporate personalization strategies, adapting the intervention based on factors such as individual learning preferences, engagement levels, and the specific learning environment (e.g., online vs. laboratory vs. classroom setting; ninaus & sailer, 2022; sailer et al., 2024). to identify patterns that can be harnessed to drive adaptive learning support across conditions, we need to understand which data are informative about which aspect of learning and under which circumstances. one research area fully dedicated to this goal is learning analytics, which seeks to optimize learning and the environments in which it occurs (long & siemens, 2011). in learning analytics, the learning process is envisaged as a dynamic, closed-loop system (clow, 2012; ninaus & sailer, 2022; sailer et al., 2024), similarly to the theoretical considerations outlined in section 1.1. this loop connects learners with behavioral traces (i.e., data) they leave upon interacting with a digital learning system, the patterns that are identified within those traces that influence a learning outcome, and the interventions derived from those patterns to enhance that learning outcome. in the final step of the cycle, the learning environment adapts with respect to identified patterns in a learner’s traces and provides the learner with a new, adapted learning interface, through which the learning process continues in the next iteration of the cycle. a good understanding about what learning processes and outcomes can be expected under what circumstances is of utmost importance for devising such adaptive, closed-loop learning systems. this begins already at the data acquisition stage, because different contexts may dictate how often and what types of data can be collected during learning. for instance, it is well-known that affect expression varies over contexts (barrett et al., 2019; d’mello et al., 2018). if, for example, automated affect detection is then used to probe affective dynamics, for instance, for the detection of signs of confusion or frustration during learning (cloude et al., 2022), it might well turn out that insights from studies conducted in the laboratory are hardly, if at all, transferable into the field (d’mello & booth, 2023). a laboratory setting is not a typical setting for learning. however, it is a typical setting for conducting highly controlled research about learning (booth et al., 2023). our present study also shows that results may considerably differ between laboratory and other (here, online learning) contexts. hence, for the development of tailored educational interventions, it seems essential to know the boundary conditions under which a certain result will be obtained. a result obtained in the laboratory may turn out practically useless in the field and vice versa (d’mello & booth, 2023). as such, the development of context-aware interventions requires the creation of context-specific metrics for evaluating their effectiveness and to gain a more holistic understanding of how learners interact with different learning designs. for example, metrics that work well in online environments (e.g., time-on-task, click-through rates) may not be appropriate for controlled lab settings, where other factors (e.g., effort investment or self-reported engagement) may be more relevant. another area of consideration for future research should be on properly capturing cognitive and affective (also including motivational) fluctuations within and between contexts. while much of the research has focused on studying cognitive and affective processes in situ to understand cognitive changes in single tasks and contexts, future studies should also consider how learner cognition and affect fluctuates across contexts. while notoriously difficult, a promising approach for capturing cognitive and affective fluctuations is through the use of multimodal data collected across different tasks and learning environments (molenaar et al., 2023; ninaus & sailer, 2022; sailer et al., 2024). these multimodal learning analytics (blikstein, 2013) could include eye movements, physiological signals, facial expressions, interactions with the learning system, and repeated self-reported motivation levels. by integrating and triangulating multiple data channels, researchers may gain a deeper understanding of how cognition, affect, and motivation change in real time and how varying contexts affect their outcomes. 6. conclusion in the present work, we set out to investigate how different contextual study features can influence the impact of gamified learning on cognitive learning and motivational outcomes in an associative learning task. we found a substantial influence of the considered contextual features for motivational outcomes. in particular, whereas gamified learning tasks were perceived as substantially more appealing and stimulating in online settings, they exhibited no motivational benefits over non-gamified tasks in a controlled, laboratory setting under otherwise similar conditions. the influence on cognitive learning outcomes appeared relatively more nuanced but aligned well with cognitive and cognitive-affective learning theories, indicating that if cognitive outcomes are affected by gamification, these influences are best illuminated in the laboratory for the given learning task. although only a limited scope of varying contexts, a single learning task, and motivational outcomes solely based on self-reports were considered in this study, we nevertheless think that the present work shows the importance of considering boundary conditions of gamified learning in educational research. keypoints boundary conditions of cognitive and motivational effects of gamified learning are explored. indications for boundary conditions of motivational effects are identified. motivational effects differ substantially between laboratory and online settings. cognitive learning outcomes are comparatively robust across contexts. cognitive effects of gamification can be affected by study context to some extent via motivational effects. acknowledgments this work was supported by the strategic research council (grant no. 358250) within the research council of finland. the authors further acknowledge financial support by the university of graz. references alotaibi, m. s. (2024). game-based learning in early childhood education: a systematic review and meta-analysis. frontiers in psychology, 15, 1307881. https://doi.org/10.3389/fpsyg.2024.1307881 alyahyan, e., & düştegör, d. (2020). predicting academic success in higher education: literature review and best practices. international journal of educational technology in higher education, 17(1), 3. https://doi.org/10.1186/s41239-020-0177-7 arztmann, m., hornstra, l., jeuring, j., & kester, l. (2023). effects of games in stem education: a meta-analysis on the moderating role of student background characteristics. studies in science education, 59(1), 109–145. https://doi.org/10.1080/03057267.2022.2057732 azevedo, r., & wiedbusch, m. (2023). theories of metacognition and pedagogy applied to aied systems. in b. du boulay, a. mitrovic, & k. yacef (eds.), handbook of artificial intelligence in education (pp. 45–67). edward elgar publishing. https://doi.org/10.4337/9781800375413.00013 bai, s., hew, k. f., & huang, b. (2020). does gamification improve student learning outcome? evidence from a meta-analysis and synthesis of qualitative data in educational contexts. educational research review, 30, 100322. https://doi.org/10.1016/j.edurev.2020.100322 bandura, a. (1997). self-efficacy: the exercise of control. w. h. freeman. barrett, l. f., adolphs, r., marsella, s., martinez, a. m., & pollak, s. d. (2019). emotional expressions reconsidered: challenges to inferring emotion from human facial movements. psychological science in the public interest, 20(1), 1–68. https://doi.org/10.1177/1529100619832930 barz, n., benick, m., dörrenbächer-ulrich, l., & perels, f. (2023). the effect of digital game-based learning interventions on cognitive, metacognitive, and affective-motivational learning outcomes in school: a meta-analysis. review of educational research, 003465432311677. https://doi.org/10.3102/00346543231167795 baten, e., vansteenkiste, m., de muynck, g.-j., de poortere, e., & desoete, a. (2020). how can the blow of math difficulty on elementary school children’s motivational, cognitive, and affective experiences be dampened? the critical role of autonomy-supportive instructions. journal of educational psychology, 112(8), 1490–1505. https://doi.org/10.1037/edu0000444 bernecker, k., & ninaus, m. (2021). no pain, no gain? investigating motivational mechanisms of game elements in cognitive tasks. computers in human behavior, 114, 106542. https://doi.org/10.1016/j.chb.2020.106542 blikstein, p. (2013). multimodal learning analytics. proceedings of the third international conference on learning analytics and knowledge, 102–106. https://doi.org/10.1145/2460296.2460316 booth, b. m., bosch, n., & d’mello, s. k. (2023). engagement detection and its applications in learning: a tutorial and selective review. proceedings of the ieee, 111(10), 1398–1422. https://doi.org/10.1109/jproc.2023.3309560 borenstein, m., hedges, l. v., higgins, j. p. t., & rothstein, h. r. (2009). introduction to meta‐analysis (1st ed.). wiley. https://doi.org/10.1002/9780470743386 chi, m. t. h., & wylie, r. (2014). the icap framework: linking cognitive engagement to active learning outcomes. educational psychologist, 49(4), 219–243. https://doi.org/10.1080/00461520.2014.965823 cloude, e. b., dever, d. a., hahs-vaughn, d. l., emerson, a. j., azevedo, r., & lester, j. (2022). affective dynamics and cognition during game-based learning. ieee transactions on affective computing, 13(4), 1705–1717. https://doi.org/10.1109/taffc.2022.3210755 clow, d. (2012). the learning analytics cycle: closing the loop effectively. proceedings of the 2nd international conference on learning analytics and knowledge, 134–138. https://doi.org/10.1145/2330601.2330636 conijn, r., snijders, c., kleingeld, a., & matzat, u. (2017). predicting student performance from lms data: a comparison of 17 blended courses using moodle lms. ieee transactions on learning technologies, 10(1), 17–29. https://doi.org/10.1109/tlt.2016.2616312 conte, p. d. (2019). a gender study on the effects of the “high five game” on the math learning performance of children. international journal of scientific and technology research, 8(12), 2063–2066. deterding, s., dixon, d., khaled, r., & nacke, l. (2011). from game design elements to gamefulness: defining “gamification.” proceedings of the 15th international academic mindtrek conference: envisioning future media environments, 9–15. https://doi.org/10.1145/2181037.2181040 d’mello, s. k., & booth, b. m. (2023). affect detection from wearables in the “real” wild: fact, fantasy, or somewhere in between? ieee intelligent systems, 38(1), 76–84. https://doi.org/10.1109/mis.2022.3221854 d’mello, s. k., kappas, a., & gratch, j. (2018). the affective computing approach to affect measurement. emotion review, 10(2), 174–183. https://doi.org/10.1177/1754073917696583 domjan, m. (2010). the principles of learning and behavior. wadsworth cengage learning. eccles, j. s., & wigfield, a. (2020). from expectancy-value theory to situated expectancy-value theory: a developmental, social cognitive, and sociocultural perspective on motivation. contemporary educational psychology, 61, 101859. https://doi.org/10.1016/j.cedpsych.2020.101859 engeser, s., & rheinberg, f. (2008). flow, performance and moderators of challenge-skill balance. motivation and emotion, 32(3), 158–172. https://doi.org/10.1007/s11031-008-9102-4 fadda, d., pellegrini, m., vivanet, g., & zandonella callegher, c. (2022). effects of digital games on student motivation in mathematics: a meta‐analysis in k‐12. journal of computer assisted learning, 38(1), 304–325. https://doi.org/10.1111/jcal.12618 fiedler, k., & beier, s. (2014). affect and cognitive processes in educational contexts. in r. pekrun & l. linnenbrank-garcia (eds.), international handbook of emotions in education (1st ed.). routledge. fryer, l. k., & dinsmore, d. l. (2020). the promise and pitfalls of self-report. frontline learning research, 8(3), 1–9. https://doi.org/10.14786/flr.v8i3.623 fulmer, s. m., & frijters, j. c. (2009). a review of self-report and alternative approaches in the measurement of student motivation. educational psychology review, 21(3), 219–246. https://doi.org/10.1007/s10648-009-9107-x gaspard, h., jiang, y., piesch, h., nagengast, b., jia, n., lee, j., & bong, m. (2020). assessing students’ values and costs in three countries: gender and age differences within countries and structural differences across countries. learning and individual differences, 79, 101836. https://doi.org/10.1016/j.lindif.2020.101836 greipl, s., klein, e., lindstedt, a., kiili, k., moeller, k., karnath, h.-o., bahnmueller, j., bloechle, j., & ninaus, m. (2021). when the brain comes into play: neurofunctional correlates of emotions and reward in game-based learning. computers in human behavior, 125, 106946. https://doi.org/10.1016/j.chb.2021.106946 greipl, s., moeller, k., & ninaus, m. (2020). potential and limits of game-based learning. international journal of technology enhanced learning, 12(4), 363. https://doi.org/10.1504/ijtel.2020.110047 guay, f., vallerand, r. j., & blanchard, c. (2000). on the assessment of situational intrinsic and extrinsic motivation: the situational motivation scale (sims). motivation and emotion, 24(3), 175–213. https://doi.org/10.1023/a:1005614228250 harmon-jones, e., & mills, j. (2019). an introduction to cognitive dissonance theory and an overview of current perspectives on the theory. in e. harmon-jones (ed.), cognitive dissonance: reexamining a pivotal theory in psychology (2nd ed.). (pp. 3–24). american psychological association. https://doi.org/10.1037/0000135-001 hatano, a., ogulmus, c., shigemasu, h., & murayama, k. (2022). thinking about thinking: people underestimate how enjoyable and engaging just waiting is. journal of experimental psychology: general, 151(12), 3213–3229. https://doi.org/10.1037/xge0001255 heckhausen, j., & heckhausen, h. (2018). motivation and action: introduction and overview. in j. heckhausen & h. heckhausen (eds.), motivation and action (pp. 1–14). springer international publishing. https://doi.org/10.1007/978-3-319-65094-4_1 hoerger, m. (2010). participant dropout as a function of survey length in internet-mediated university studies: implications for study design and voluntary participation in psychological research. cyberpsychology, behavior, and social networking, 13(6), 697–700. https://doi.org/10.1089/cyber.2009.0445 hoerger, m., & currell, c. (2012). ethical issues in internet research. in s. j. knapp, m. c. gottlieb, m. m. handelsman, & l. d. vandecreek (eds.), apa handbook of ethics in psychology, vol 2: practice, teaching, and research (pp. 385–400). american psychological association. hu, y., gallagher, t., wouters, p., van der schaaf, m., & kester, l. (2022). game‐based learning has good chemistry with chemistry education: a three‐level meta‐analysis. journal of research in science teaching, 59(9), 1499–1543. https://doi.org/10.1002/tea.21765 huang, r., ritzhaupt, a. d., sommer, m., zhu, j., stephen, a., valle, n., hampton, j., & li, j. (2020). the impact of gamification in educational settings on student learning outcomes: a meta-analysis. educational technology research and development, 68(4), 1875–1901. https://doi.org/10.1007/s11423-020-09807-z huber, s. e., cortez, r., kiili, k., lindstedt, a., & ninaus, m. (2023). game elements enhance engagement and mitigate attrition in online learning tasks. computers in human behavior, 149, 107948. https://doi.org/10.1016/j.chb.2023.107948 huber, s. e., edlinger, m., lindstedt, a., kiili, k., & ninaus, m. (2024). game elements improve affect and motivation in a learning task. international journal of serious games, 11(4), 103–126. https://doi.org/10.17083/ijsg.v11i4.769 jasp team. (2024). jasp (version 0.19.1.0) [computer software]. koskinen, a., mcmullen, j., ninaus, m., & kiili, k. (2023). does the emotional design of scaffolds enhance learning and motivational outcomes in game‐based learning? journal of computer assisted learning, 39(1), 77–93. https://doi.org/10.1111/jcal.12728 krauspe, j., ebersbach, m., ludwig, a., & scharf, f. (2025). do worked examples boost the spacing effect on lasting learning? learning and instruction, 97, 102103. https://doi.org/10.1016/j.learninstruc.2025.102103 kuratomi, k., johnsen, l., kitagami, s., hatano, a., & murayama, k. (2023). people underestimate their capability to motivate themselves without performance-based extrinsic incentives. motivation and emotion, 47(4), 509–523. https://doi.org/10.1007/s11031-022-09996-5 lampropoulos, g. & kinshuk. (2024). virtual reality and gamification in education: a systematic review. educational technology research and development, 72(3), 1691–1785. https://doi.org/10.1007/s11423-024-10351-3 laugwitz, b., held, t., & schrepp, m. (2008). construction and evaluation of a user experience questionnaire. in a. holzinger (ed.), hci and usability for education and work (vol. 5298, pp. 63–76). springer berlin heidelberg. https://doi.org/10.1007/978-3-540-89350-9_6 lewin, k. (1936). principles of topological psychology. mcgraw-hill. lewin, k. (1946). behavior and development as a function of the total situation. in l. carmichael (ed.), manual of child psychology. (pp. 791–844). john wiley & sons inc. https://doi.org/10.1037/10756-016 li, l., hew, k. f., & du, j. (2024). gamification enhances student intrinsic motivation, perceptions of autonomy and relatedness, but minimal impact on competency: a meta-analysis and systematic review. educational technology research and development. https://doi.org/10.1007/s11423-023-10337-7 long, p. d., & siemens, g. (2011). penetrating the fog: analytics in learning and education. educause review. https://er.educause.edu/articles/2011/9/penetrating-the-fog-analytics-in-learning-and-education mavridis, a., katmada, a., & tsiatsos, t. (2017). impact of online flexible games on students’ attitude towards mathematics. educational technology research and development, 65(6), 1451–1470. https://doi.org/10.1007/s11423-017-9522-5 mayer, r. e. (2011). applying the science of learning. pearson. mayer, r. e. (2014). computer games for learning: an evidence-based approach. mit press. mayer, r. e. (2019). computer games in education. annual review of psychology, 70(1), 531–549. https://doi.org/10.1146/annurev-psych-010418-102744 mayer, r. e. (2020). cognitive foundations of game-based learning. in j. l. plass, r. e. mayer, & b. d. homer (eds.), handbook of game-based learning (pp. 83–110). mit press. mayer, r. e. (2024). the past, present, and future of the cognitive theory of multimedia learning. educational psychology review, 36(1), 8. https://doi.org/10.1007/s10648-023-09842-1 molenaar, i., mooij, s. d., azevedo, r., bannert, m., järvelä, s., & gašević, d. (2023). measuring self-regulated learning and the role of ai: five years of research using multimodal multichannel data. computers in human behavior, 139, 107540. https://doi.org/10.1016/j.chb.2022.107540 neath, i., & suprenant, a. m. (2003). human memory. wadsworth/thomson learning. ninaus, m., greipl, s., kiili, k., lindstedt, a., huber, s., klein, e., karnath, h.-o., & moeller, k. (2019). increased emotional engagement in game-based learning – a machine learning approach on facial emotion detection data. computers & education, 142, 103641. https://doi.org/10.1016/j.compedu.2019.103641 ninaus, m., & sailer, m. (2022). closing the loop – the human role in artificial intelligence for education. frontiers in psychology, 13, 956798. https://doi.org/10.3389/fpsyg.2022.956798 palha, s., & jukić matić, l. (2025). what do teachers anticipate from education in game-based pedagogy? technology, pedagogy and education, 1–14. https://doi.org/10.1080/1475939x.2025.2454453 plass, j. l., homer, b. d., & kinzer, c. k. (2015). foundations of game-based learning. educational psychologist, 50(4), 258–283. https://doi.org/10.1080/00461520.2015.1122533 plass, j. l., homer, b. d., mayer, r. e., & kinzer, c. k. (2020). theoretical foundations of game-based and playful learning. in j. l. plass, r. e. mayer, & b. d. homer (eds.), handbook of game-based learning (pp. 3–24). mit press. plass, j. l., & kaplan, u. (2016). emotional design in digital media for learning. in emotions, technology, design, and learning (pp. 131–161). elsevier. https://doi.org/10.1016/b978-0-12-801856-9.00007-4 prenkaj, b., velardi, p., stilo, g., distante, d., & faralli, s. (2021). a survey of machine learning approaches for student dropout prediction in online courses. acm computing surveys, 53(3), 1–34. https://doi.org/10.1145/3388792 ritzhaupt, a. d., huang, r., sommer, m., zhu, j., stephen, a., valle, n., hampton, j., & li, j. (2021). a meta-analysis on the influence of gamification in formal educational settings on affective and behavioral outcomes. educational technology research and development, 69(5), 2493–2522. https://doi.org/10.1007/s11423-021-10036-1 rodrigues, l., pereira, f., toda, a., palomino, p., oliveira, w., pessoa, m., carvalho, l., oliveira, d., oliveira, e., cristea, a., & isotani, s. (2022). are they learning or playing? moderator conditions of gamification’s success in programming classrooms. acm transactions on computing education, 22(3), 1–27. https://doi.org/10.1145/3485732 ryan, r. m., & deci, e. l. (2019). brick by brick: the origins, development, and future of self-determination theory. in advances in motivation science (vol. 6, pp. 111–156). elsevier. https://doi.org/10.1016/bs.adms.2019.01.001 ryan, r. m., & rigby, c. s. (2020). motivational foundations of game-based learning. in j. l. plass, r. e. mayer, & b. d. homer (eds.), handbook of game-based learning (pp. 153–176). mit press. sailer, m., & homner, l. (2020). the gamification of learning: a meta-analysis. educational psychology review, 32(1), 77–112. https://doi.org/10.1007/s10648-019-09498-w sailer, m., ninaus, m., huber, s. e., bauer, e., & greiff, s. (2024). the end is the beginning is the end: the closed-loop learning analytics framework. computers in human behavior, 158, 108305. https://doi.org/10.1016/j.chb.2024.108305 sailer, m., & sailer, m. (2021). gamification of in‐class activities in flipped classroom lectures. british journal of educational technology, 52(1), 75–90. https://doi.org/10.1111/bjet.12948 sammut, r., griscti, o., & norman, i. j. (2021). strategies to improve response rates to web surveys: a literature review. international journal of nursing studies, 123, 104058. https://doi.org/10.1016/j.ijnurstu.2021.104058 schlag, r., sailer, m., tolks, d., ninaus, m., & sailer, m. (2024). effectiveness of gamification in education. in a. gegenfurtner & i. kollar (eds.), designing effective digital learning environments (pp. 143–159). routledge. schrepp, m., hinderks, a., & thomaschewski, j. (2017). die ux kpi wunsch und wirklichkeit. https://doi.org/10.18420/muc2017-up-0100 schwartz, r. n., & plass, j. l. (2020). types of engagement in learning with games. in j. l. plass, r. e. mayer, & b. d. homer (eds.), handbook of game-based learning (pp. 53–80). mit press. seo, k., dodson, s., harandi, n. m., roberson, n., fels, s., & roll, i. (2021). active learning with online video: the impact of learning context on engagement. computers & education, 165, 104132. https://doi.org/10.1016/j.compedu.2021.104132 sinatra, g. m., heddy, b. c., & lombardi, d. (2015). the challenges of defining and measuring student engagement in science. educational psychologist, 50(1), 1–13. https://doi.org/10.1080/00461520.2014.1002924 skitka, l. j., & sargis, e. g. (2006). the internet as psychological laboratory. annual review of psychology, 57(1), 529–555. https://doi.org/10.1146/annurev.psych.57.102904.190048 smith, s. m., & vela, e. (2001). environmental context-dependent memory: a review and meta-analysis. psychonomic bulletin & review, 8(2), 203–220. https://doi.org/10.3758/bf03196157 sweller, j. (2011). cognitive load theory. in psychology of learning and motivation (vol. 55, pp. 37–76). elsevier. https://doi.org/10.1016/b978-0-12-387691-1.00002-8 urhahne, d., & wijnia, l. (2023). theories of motivation in education: an integrative framework. educational psychology review, 35(2), 45. https://doi.org/10.1007/s10648-023-09767-9 wesenberg, l., jansen, s., krieglstein, f., schneider, s., & rey, g. d. (2025). the influence of seductive details in learning environments with low and high extrinsic motivation. learning and instruction, 96, 102054. https://doi.org/10.1016/j.learninstruc.2024.102054 wilcox, r. r. (2022). introduction to robust estimation and hypothesis testing (5th ed.). academic press. wilde, m., bätz, k., kovaleva, a., & urhahne, d. (2009). überprüfung einer kurzskala intrinsicher motivation (kim). zeitschrift für didaktik der naturwissenschaften, 15, 31–45. wong, z. y., & liem, g. a. d. (2022). student engagement: current state of the construct, conceptual refinement, and future research directions. educational psychology review, 34(1), 107–138. https://doi.org/10.1007/s10648-021-09628-3 wouters, p., van nimwegen, c., van oostendorp, h., & van der spek, e. d. (2013). a meta-analysis of the cognitive and motivational effects of serious games. journal of educational psychology, 105(2), 249–265. https://doi.org/10.1037/a0031311 xu, z., olson, j., pochinki, n., zheng, z., & yu, r. (2024). contexts matter but how? course-level correlates of performance and fairness shift in predictive model transfer. proceedings of the 14th learning analytics and knowledge conference, 713–724. https://doi.org/10.1145/3636555.3636936 zainuddin, z., chu, s. k. w., shujahat, m., & perera, c. j. (2020). the impact of gamification on learning and instruction: a systematic review of empirical evidence. educational research review, 30, 100326. https://doi.org/10.1016/j.edurev.2020.100326 zhang, l., lei, y., pelton, t., pelton, l. f., & shang, j. (2024). an exploration of gendered differences in cognitive, motivational and emotional aspects of game‐based math learning. journal of computer assisted learning, 40(6), 2633–2649. https://doi.org/10.1111/jcal.12956 zhang, q., & yu, z. (2022). meta-analysis on investigating and comparing the effects on learning achievement and motivation for gamification and game-based learning. education research international, 2022, 1–19. https://doi.org/10.1155/2022/1519880 frontline learning research 12 no. 4 (2024) 22 54 issn 2295-3159 corresponding author: susannah c. davis, university of new mexico albuquerque, nm 87131, usa email:, scdavis@unm.edu doi: https://doi.org/10.14786/flr.v12i4.1331 inclusive excellence in practice: integrating equitable, consequential learning and an inclusive climate in higher education classrooms and institutions susannah c. davis1, susan bobbitt nolen2, & milo d. koretsky3 1university of new mexico, usa 2university of washington, usa 3tufts university, usa article received 19 july 2023 / article revised 7 april 2024 / accepted 23 september 2024 / available online 19 november 2024 abstract in education, initiatives aimed at improving diversity, equity, inclusivity, and justice (deij) are often conceptualized and implemented separately from those addressing students’ and faculty’s learning — and the reverse is also true. in this theoretical paper with an empirical illustration, we present a holistic framework based on our experience with a comprehensive change initiative. the i2 framework posits that deij and learning goals need to be addressed simultaneously and at multiple, intersecting organizational levels. through a systems approach, i2 integrates change activity across two dimensions: one representing goals of reform (deij and improved learning) and another representing levels of organizational change (classroom and department/organization). i2 integrates the work of creating equitable, consequential learning opportunities in the classroom and the work of creating an inclusive climate at the departmental/organizational level, emphasizing their inherent relatedness. we provide an empirical example based on design-based implementation research and related mixed methods analyses of a multi-year change project in an engineering department at a large, public university in the united states. the example highlights a need to shift the nature of this work, how we do this work, and the environment and culture within which we do this work at both the classroom level and the department level. the example also illustrates ways that elements of the change initiative intersected with existing institutional practices, leading some innovations to succeed and others to be resisted. the i2 framework provides guidance to practitioners, policymakers, and leaders working towards equitable, consequential learning at the classroom level and an inclusive climate at departmental and institutional levels. keywords: higher education; inclusive excellence; equity; learning; organizational change https://doi.org/10.14786/flr.v12i4.1331 davis, nolen & koretsky 23 | f l r 1. introduction in education, initiatives that address improving diversity, equity, inclusivity, and justice (deij) too often are conceptualized and implemented separately from those addressing improved learning opportunities and experiences (bauman et al., 2005; milem et al., 2005; williams et al., 2005). likewise, most research on improving learning focuses on changing classroom activity, without sufficient consideration of departmental or organizational contexts, policies, and practices, or faculty members’ learning, development, and experiences (maass et al., 2019). in this paper, we argue that fostering deij and improving learning opportunities are mutually constitutive and synergistic and should be addressed using a systemic, multilevel approach that considers classroom, department, and organizational contexts. deij should be central to research and practice on instructional innovation aimed at students’ learning, and students’ and faculty’s learning, development, and experiences should be central to efforts to improve deij in departmental and institutional cultures. in support of this argument, we present the i2 framework, termed i2 because it integrates organizational change efforts within and across two dimensions: one representing goals of reform (deij and improved learning) and another representing levels of organizational change (classroom and department/organization). i2 integrates equitable, consequential learning opportunities in the classroom and an inclusive climate at the departmental/organizational level, emphasizing their inherent relatedness. our systems approach to organizational change accounts for the interconnected social, cultural, and organizational processes that shape reform efforts (engeström, 2001; greeno & engeström, 2014; holland, 2010), and treats the educational institution as a complex system, seeking to develop a strategy compatible with that system (henderson et al., 2011). we draw upon cultural-historical activity theory (chat; engeström, 2001) in considering organizations as activity systems, in which different groups of people, structures, rules, norms, and behaviors serve to maintain or change the status quo. using chat and sociocultural views of learning and identity, we illustrate the i2 framework through an empirical case, describing the design and implementation of our own reform project aimed at improving learning and deij at classroom and departmental levels within a multidisciplinary engineering department in a large, public, research-intensive university in the united states. there is increased global attention to broadening access and supporting student learning and retention in higher education (blackie et al., 2016; direito et al., 2021; mejia & martin, 2023; pineda & mishra, 2023; siri et al., 2022), despite variation in how programs and countries across the world conceptualize and prioritize both equity-centered efforts and pedagogical change. the i2 framework provides a context-sensitive tool for designing, implementing, and evaluating multi-level systems change that addresses both learning and deij goals. in the literature review in the following section, we describe the need for this framework. next, we provide the framework’s theoretical underpinnings and describe the i2 framework itself. finally, we operationalize the framework with an empirical illustration from our own change project. 2. literature review to contextualize the i2 framework, this section summarizes prior research, first on students’ opportunities to learn and meaningfully engage with disciplinary knowledge and practices in classrooms, and then on departmental and campus climate. next, we advocate for an integrated, multilevel model to incorporate learning and deij goals by reviewing how inclusive excellence has been addressed in higher education. there is increasing attention across the globe in diversifying higher education and expanding access and success for students from historically marginalized communities in higher education, though how diversity is conceptualized and supported varies by context (langholz, 2014; mejia & martin, davis, nolen & koretsky 24 | f l r 2023; pineda & mishra, 2023). rationales for increasing diversity include economic advancement, social justice, equity, and internationalization (pineda & mishra, 2022). there are also regional variations in the aspects of diversity (e.g., gender, race/ethnicity, cultural diversity, and inclusion) prioritized (pineda & mishra, 2022). even within europe, there is variation in the ways that ideas about diversity, equity, and inclusion are conceptualized and implemented (direito et al., 2021; pineda & mishra, 2023). in higher education reform in the united states, europe, and beyond, research, organizational policies and practices, and funding opportunities often foreground either deij goals or student learning goals. some efforts to improve access, experiences, and outcomes for minoritized and underserved groups have focused on recruitment, institutional climate, and reducing harm through policy change and faculty development (e.g., bensimon et al., 2019; european commission/ecea/eurydice, 2022; milem et al., 2005; rankin & reason, 2008). efforts to improve learning outcomes for all students have primarily focused on increasing the use of active learning and associated research-based instructional practices (christie & de graff, 2017; lima et al., 2017; lombardi et al., 2021; prince, 2004; raver & maydosz, 2010). these goals—advancing deij in higher education and creating more effective learning opportunities for students—are both necessary to improve the experiences of students and faculty in higher education. moreover, students’ classroom learning and the inclusivity of the departmental and institutional climate in which they learn are connected, but often researched separately or with only vague associations. 2.1 classroom learning much research in education, including in stem, has connected pedagogical practices and instructional innovations with student learning (christersson et al., 2019; national research council, 2012). in general, calls for reform support a shift from teacher-centered to student-centered classroom learning (barr & tagg, 1995), commonly conceptualized in terms of active learning pedagogies (freeman et al., 2014; lima et al., 2017; lombardi et al., 2021). rather than a transmission model of traditional lecture-based instruction, active learning “requires students to do meaningful learning activities and think about what they are doing” (prince, 2004, p. 223). for example, in the flipped classroom approach, students read the textbook or watch lectures outside of class and work together during class, often on the same homework problems as the corresponding traditional lecture course (bishop & verleger, 2013). when equity issues arise in active learning interventions, some researchers have looked to characterize the quality of instruction without considering how instruction positions students in relation to their social and professional identities. for example, theobald et al. (2020) relate the intensity of active learning—the percentage of class time students are engaged in active learning—to the performance of minoritized students, showing high intensity approaches lead to more equitable achievement outcomes. however, this work is framed in terms of the achievement gaps minoritized students need to overcome rather than restructuring classroom activities in ways where these students’ own experiences and perspectives are assets in doing the work. similarly, others frame inequitable participation patterns as an issue of individual students’ confidence or motivation (brown et al., 2015; kurth et al., 2002). instructors may then aim to change individual behavior through psychological or emotional interventions such as values affirmation interventions (hulleman et al., 2017; jordt et al., 2017), while ignoring structural reasons for uneven participation. in contrast, we advocate for a sociocultural framework to create equitable learning opportunities that account for how the activity socially positions the learners. for example, in cases where problems have single correct answers, groups often coalesce around a high-status student whom the others rely on to direct the work (cohen & lotan, 1997; horn, 2012). high status in small group work is often conferred by membership in the dominant social identity group, facility with the language of instruction, quick responses, and social strategies including ignoring or shutting down attempts by lower-status students to participate (horn, 2012; kurth et al., 2002). in contrast, problems structured to allow for davis, nolen & koretsky 25 | f l r multiple approaches or that have more than one acceptable answer can encourage multiple perspectives and more equitable participation (nolen et al., 2024), yet this type of work is uncommon in u.s. university-level stem courses. 2.2 campus and departmental climate oppressive social structures and relations within institutions are produced and reproduced through actions, interactions, practices, and policies at classroom, department, and institutional levels. the ruling relations (smith, 1999), or everyday norms, assumptions, and social interactions that structure higher education have historically been created by and for white men and create and perpetuate systems of oppression and privilege (direito et al., 2021; o’meara et al., 2018; pawley, 2019). these inform classroom, department, and broader campus climates through norms, attitudes, social interactions, and behaviors (o’meara et al., 2018; rankin & reason, 2008). for example, in a study by rankin and reason (2005), students of color experienced racial harassment at higher rates than white students and female students reported higher rates of gender harassment. the negative experiences of students from historically marginalized groups, including harassment, bias, and microaggressions, may be amplified in stem fields (rolin, 2008), which have historically centered the experiences of people who identify as white, male, and heterosexual (direito et al., 2021; pawley, 2019; rincón & georgejackson, 2016; secules, 2019; slaton, 2010). in engineering, researchers have described this phenomenon as the “chilly climate” (hall & sandler, 1982; sandler et al., 1996; walton et al., 2015) or “climate of intimidation” (palmer et al., 2011). perceptions of unwelcoming and inequitable campus climates, feelings of isolation, and negative experiences with bias, harassment, and microaggressions have a detrimental impact on the learning experiences, sense of belonging, disciplinary identification, and persistence of students from historically marginalized communities (beasley & fischer, 2012; chang et al., 2014; hausmann et al., 2007; hurtado et al., 1998; marra et al., 2012; ong et al., 2011, 2018; pawley, 2019; rankin & reason, 2005; seymour & hewitt, 1997; tonso, 2007; yosso et al., 2009). faculty in higher education also experience the effects of non-inclusive campus and department climates and inequitable policies, practices, and norms that perpetuate systems of oppression (garvey & rankin, 2018; harris, 2020; hart, 2016). women and faculty of color remain underrepresented, despite attempts to diversify university faculty (blackburn, 2017; siri et al., 2022; turner et al., 2008; u.s. department of education, 2018). women and faculty of color face inequities in hiring, retention, and promotion (blackburn, 2017; white-lewis, 2020). while employed, these faculty also report more negative experiences, including marginalization, injustice, scholarly isolation, tokenism, social exclusion, lack of belonging, and epistemic exclusion (diggs et al., 2009; minnotte & pederson, 2021; settles et al., 2019, 2022; tippeconnic fox, 2005; turner et al., 2008; zambrana et al., 2017). research examining students’ views of their institutions’ climate (e.g., diversity, equity, discrimination, support, safety) has highlighted important aspects of campus environments that correlate with positive outcomes for diverse students (e.g., museus, 2014; museus et al., 2017). however, this approach lacks consideration of how practices, policies, and experiences at the institutional level interact with those at course and department levels. hurtado and colleagues’ (2012, 2013) model for diverse learning environments more explicitly draws attention to the multiple contexts that influence institutions of higher education and student outcomes, linking student educational outcomes with “the mesolevel dynamics of teaching and learning (inclusive of cocurricular environments) within institutions, and also with these larger macrolevel constraints and processes” (hurtado et al., 2012, p. 49). this work provides an important multicontextual vision of diverse learning environments that highlights “the interaction of systems and reciprocal influences that constrain or lead to an institution’s role in producing social transformation or the reproduction of inequality” (hurtado et al., 2012, p. 103). departmental climates are a “microcosm of the larger institution” (rincon & george-jackson, 2016, p. 743), and unwelcoming campus climates permeate the departments down to classroom interactions (hurtado et al., 2012, 2013). whether it happens within a classroom or departmental davis, nolen & koretsky 26 | f l r context, when relationships and interactions reflect and reproduce the systems of sexism, racism, classism, heterosexism, ableism, and ageism in the larger society and institutions of higher education, students and faculty experience the climate as unwelcoming and their learning, sense of belonging, and persistence suffer. therefore, organizational change efforts with learning and deij goals must consider and address the ways that students and faculty experience campus climate at both classroom and department levels, as well as how to design interventions that attend to the multidimensional nature of climate and its effects on students and faculty. 2.3 inclusive excellence while there is variation across the globe, both deij and student learning receive international attention, though usually through separate efforts (beddoes et al., 2018; blackie et al., 2016; direito et al., 2021; walden et al., 2020). equity and inclusion are core to the european union’s vision for a european education area (european commission/ecea/eurydice, 2022) as the european union works on the “twin challenge of equity and excellence” (european commission, 2024, p. 5). in addition, gender equality is a european research area priority, including consideration of intersections between gender and other aspects of identity, such as ethnicity, disability, and sexual orientation (european commission, 2021; palmén et al., 2020). recently in u.s. higher education there have been attempts to connect deij with student learning, often adopting the language of “inclusive excellence” found in three association for american colleges and universities (aac&u) reports (bauman et al., 2005; milem et al., 2005; williams et al., 2005). the aac&u reports highlight structural barriers to student success, including the preponderance of isolated efforts on campus and the disconnect between deij and educational excellence. as milem and colleagues (2005) state, “education leaders routinely work on diversity initiatives within one committee on campus and work on strengthening the quality of the educational experience within another. this disconnect serves students—and all of education—poorly” (p. vii). calls to foster inclusive excellence have spanned many fields, including nursing education (bleich et al., 2015), kinesiology (mahar et al., 2021), and teacher education (everett & grey, 2016). some have highlighted the need to incorporate ideas about inclusive excellence into faculty development (bryson et al., 2020; forde & carpenter, 2020) and into the classroom (considne et al., 2017; salazar et al., 2010). the aac&u reports simultaneously address inclusion and excellence; however, they lack specific guidance connecting the classroom and department levels, and thus do not adequately address the ways in which students and faculty experience everyday life within their institutions. students and faculty who hold a variety of marginalized and nondominant identities do not artificially separate what happens in the classroom, in the department, and at their institution. as a student walks through a day in their academic life, they may go to a class where they are talked over and their ideas are co-opted by others with more dominant social identities, followed by an advising appointment where they learn that two of the classes they need for their major meet at times when they are scheduled to work at their financially-necessary job, followed by an office hour where the professor remarks on how “well spoken” they are in a way that feels racialized, followed by a class where they are the only student from their ethnic community and no examples or texts connect with their lived experiences. these everyday experiences operate within overarching structural, cultural, disciplinary, and interpersonal power dynamics that affect but transcend relational, classroom, departmental, and institutional levels (collins & bilge, 2020). within a student’s (or faculty member’s) holistic academic experience, not only are their experiences of learning and inclusion interconnected, but these experiences are affected by what is happening at the micro-interaction, classroom, department, and institutional levels. the aac&u inclusive excellence framework focuses primarily on the institutional level and is aimed at high-level academic leaders. of the three initial reports, implications for classroom environments are mentioned only by milem and colleagues (2005), and only briefly. they argue that “active learning pedagogies provide opportunities for students from different backgrounds to engage with each other around the content of courses—a form of interaction that is restricted in a lecture-based davis, nolen & koretsky 27 | f l r environment. these interactions have a direct impact on climate by breaking down stereotypes and facilitating more nuanced out-of-class interactions” (p. 25). the authors also mention the importance of increasing faculty compositional diversity and integrating diverse perspectives in the curriculum. while this report offers some guideposts, it lacks a rich description of how to integrate learning and inclusion at the classroom level, and doesn’t address the contextual connections between the classroom, department, and institution. the need for an integrated multilevel model can be seen in the trajectory of funding priorities. for example, the engaged student learning track of the u.s. national science foundation’s (nsf) improving undergraduate stem education (iuse) initiative (nsf, n.d.a) focuses mainly on the classroom level while another nsf program, advance: organizational change for gender equity in stem academic professions (advance), operates at the institutional level (nsf, n.d.b). recently, opportunities like the nsf’s revolutionizing engineering & computer science departments (red) program (nsf, n.d.c) and the howard hughes medical institute’s inclusive excellence programs (hhmi, n.d.a; n.d.b), have left room for connections between classroom and departmental and institutional cultures and practices. however, few grantees have framed their change projects to address the connections between learning and deij in ways that cross the boundaries between classrooms and department cultures. based on our analysis of their public abstracts, only four of the 24 red projects explicitly address connections both between learning and deij and also between classroom and departmental levels. the majority foreground some but not all aspects of the i2 framework. for example, if deij is considered at all, the underlying logic is frequently that making the curriculum more oriented to current technical challenges will attract diverse students. the focus of these programs is often on learning at the classroom level through curriculum change without addressing learning at the department level (e.g., how faculty will learn to make curricular and pedagogical changes in line with the project’s goals), and concerns about deij are secondary. the development of institutional models of assessments (e.g., climate surveys, focus groups) is an important step in assessing the current climate at particular institutions, establishing which inputs are correlated with desired outputs (i.e., student outcomes), and providing a vision for campus environments that support diverse students to thrive. however, such causal models, which suggest linear cause and effect relations, do not suffice when dealing with the complexity inherent in advancing deij and learning goals (hurtado, 2012, 2013). instead, we take a socioculturally-grounded systems approach to capture the complexities of learning, identity, and the multidimensional, interconnected contexts within which these develop. additionally, survey research leaves open the question of how institutions and programs within them might make changes that advance the positive outcomes associated with more equitable campus climates (hurtado et al., 2008). by taking an activity systems approach and using design-based implementation research (dbir; penuel et al., 2011; sabelli & dede; 2013) to iteratively inform the development, implementation, and evaluation a comprehensive departmental reform using the i2 framework, our research project seeks to contribute to the field’s understanding of the connectedness of deij and learning, as well as the processes, tools, and practices that advance both deij and learning for students and faculty. the i2 framework provides guidance for programs, institutions, and funding agencies that are trying to take a systemic, holistic approach to advancing inclusive excellence in practice while also addressing the interconnections between deij and learning goals. next, we describe the sociocultural approach that forms the spine of the i2 framework. 3. theoretical framework here, we outline the activity systems perspective underlying the i2 framework, taking a sociocultural or situative approach to learning and identity (holland et al., 1998; tonso, 2007; turner davis, nolen & koretsky 28 | f l r & nolen, 2015; wenger, 1998). this approach views learning and identity as socially constructed within and across particular contexts. in an engineering classroom, for example, learning processes (interactions with peers and instructors, problem-solving activities, etc.) and outcomes (getting a “correct” answer, high scores on tests, knowing how to approach a difficult problem) position students in particular ways (e.g., as more or less able, as desirable group members, as hard to work with, as someone with good ideas, etc.). at the same time, those social identities can open or close off opportunities to learn, including regulating access to materials, having one’s ideas taken seriously or ignored, being trusted to figure things out or told to follow directions. much research in higher education considers individuals or groups to be influenced by their contexts (e.g., cabrera et al., 1999; espinosa, 2011; hurtado et al., 1998; museus et al., 2017). changing a department or classroom environment to be more inclusive, for example, would be expected to affect student engagement. in a situative view, individuals and their contexts are not seen as separate entities but as cultural-historical activity systems (engeström, 1987, 2001). individuals are part of their social contexts; learning and identity construction occur within and across those contexts, and the contexts themselves are changed through the activity of the individuals within them through social practice. in chat, individual and group motives and actions mutually influence each other and evolve over time as groups pursue a shared object (engeström, 2001; miettinen, 2005). the object of activity refers to the “ultimate reason” behind collective activity (kaptelinin, 2005, p. 5). individual and collective needs and motives, shaped by the available cultural resources and tools, are negotiated through collective activity (engeström, 2001; miettinen, 2005; stetsenko, 2005). classroom, departmental, and institutional environments are transformed through changes in the practices of the people in them, including changes in policies, social structures, norms, values, and goals. multiple voices and perspectives within the activity system lead to “contradictions” or “historically accumulating structural tensions” (engeström, 2001, p. 137) within the system as it evolves over time (engeström, 1987, 2001). joint work to resolve these contradictions towards shared goals has the potential to lead to organizational learning through “collaborative envisioning and a deliberate collective change effort” (engeström, 2001, p. 137). a situative approach includes consideration of the ongoing cultural histories of people and organizations, including systems of power and oppression (collins & bilge, 2020; slaton, 2010). for example, engineering has a cultural history of participation and status largely limited to white men, leading to structures, practices, and values that reflect those participants (pawley, 2019; riley et al., 2014; secules, 2019; slaton, 2010). as participation broadens, shifts in departmental culture are negotiated. systems of rank and tenure resist change from newcomers, concentrating power in the hands of those (historically, white males) with more seniority. a change in the system, say an institutional policy change to increase diversity through hiring, may only slowly lead to changes in culture. a more diverse group of junior faculty likely brings different histories and values, but must still initially work within the existing structures, producing in ways historically valued by their departments and institutions. bringing those divergent histories to participation in existing practices may lead to changes in those practices and values for both the newer faculty and their departments, but movement toward departmental change is likely to face resistance from those with power gained through established means. newer faculty with less power may be positioned in ways that limit their influence on the department. those pushing for change may be marginalized or even removed. at the same time, newcomers of senior rank (externally hired administrators, for example) may exert more influence because of their status within the larger institutional system. similarly, a situative perspective proposes that changes in the classroom require changes on the part of instructors and students who bring their histories with them. instructors may wish to make their classrooms more just and inclusive, but their instructional practices (e.g., assigning problems with single correct answers or standard approaches), gained through their own participation in traditional classrooms, may work against this goal. if changes are made (e.g., to complex problems encouraging multiple possible approaches), students who have been successful in traditional classrooms may bring inappropriate strategies or re-create status hierarchies that marginalize divergent approaches. davis, nolen & koretsky 29 | f l r finally, classrooms are located within departments, part of the same systems of higher education. departmental norms and policies may support or work against changes to instructional practice. to the extent that innovative instructional practice is more widespread among junior faculty and non-tenure-track instructors, marginalization and conflicting value systems mitigate against investing time and energy in instructional change. structures that open opportunities for faculty collaboration or provide time and other resources for instructional innovation, on the other hand, can support changes to instruction that lead to a more welcoming climate for students as well as faculty. 4. the i2 framework in this section, we describe how the i2 framework integrates equitable, consequential learning in the classroom with the development of an inclusive climate. to do this, we use the situative approach to learning and identity, as well as an understanding of organizations as multi-voiced activity systems described above, to describe our systems-oriented framework for institutions of higher education to meet learning and deij goals through improvement efforts. the i2 framework integrates learning and equity goals across two dimensions: by fostering equitable, consequential learning in the classroom, and by connecting classroom and department/organization levels through the development of an inclusive climate (see figure 1). although this framework is general enough to be used across disciplines, given the context of this study, we draw on work in stem education to illustrate how it works. figure 1: the i2 framework integrates goals of change (deij and learning) and levels of organizational change (department and classroom) 4.1. dimension 1: equitable, consequential work in the classroom equitable, consequential work aims to engage students in more authentic disciplinary practices, and through these practices facilitate the development of students’ disciplinary identification and ability to work within and contribute to equitable and inclusive professional teams (collins et al., 1991; engle, 2012; engle & conant, 2002; hall & jurow, 2015). work is “consequential” when it addresses authentic problems. it is “equitable” when it creates meaningful opportunities for all learners to access and engage in valued practices (hall & jurow, 2015). with susan jurow and colleagues (2016), we argue that consequential work and equitable social practices are mutually constitutive. shifting from instruction as department classroom deij learning equitable student interaction, diverse perspectives equitable & just structures, procedures, faculty interaction, diverse perspectives support for faculty & admin learning & development equitable opportunity to develop professional skills, knowledge davis, nolen & koretsky 30 | f l r usual to equitable, consequential work entails a change in both the nature of the problems and the social practices of students. for example, problem-based learning (pbl) and model eliciting activities (mea) are pedagogies that follow the principles of equitable, consequential work and have been adapted by engineering educators. these pedagogies place students in small teams (e.g., 3-5 students) with realworld client or manufacturing-driven problems. to respond, students need to construct and organize knowledge, consider alternatives, engage in analysis, inquiry, and design, and critique their own reasoning and that of others (barron & darling-hammond, 2008). such pedagogies situate students’ learning activities and sense-making at least partially within the world of engineering practice (e.g., solving a relatively messy, open-ended problem with real-world constraints) rather than solely in the context of disconnected classroom practices (e.g., lectures, exams, solving decontextualized equations), thus promoting conceptual understanding rather than procedural learning (lesh et al., 2000). in addition to supporting students’ disciplinary knowledge and skills, these situated pedagogies resist the ideology of depoliticization, which frames “non-technical” concerns as irrelevant to “real” engineering work, devalues social competencies relative to technical competencies, and frames social structures as just by default (alpay et al., 2008; cech, 2014). learning environments that place students in realistic contexts have increased participation of women and students of color in science and engineering (arastoopour et al., 2014; basu & calabrese barton, 2007). such authentic contexts support professional discussion and reflection on the political, social, and ethical dimensions of science and technology, toward a vision of just engineering practice. it is insufficient, however, to merely place students in groups, even when tasks are presented as realistic engineering work. the nature of the tasks, including the extent to which diverse perspectives are framed as essential to the group’s success and supported through task design, is critical (nolen et al., 2024). equally important, instructional facilitation must align with and support equitable participation. to develop professionally, students need to not only learn about engineering concepts and calculations, but also to participate in engineering practices and to see themselves as belonging to the engineering community (gilbuena et al., 2015; herrenkohl & mertl, 2010; stevens et al., 2008). pedagogies that situate learning in authentic contexts provide opportunities to reason about key subject matter ideas, participate in the discourses of the discipline, and solve authentic problems (windschitl & calabrese barton, 2016). they help students learn knowledge and practices that have meaning and are valued in a professional context, supporting students’ integration into a professional community of practice (lave & wenger, 1991). identification with engineering is associated with improved student learning, academic success, and persistence (rodriguez et al., 2018). windschitl and calabrese barton (2016) argue for a framework that integrates rigor and equity, wherein “rigor is codetermined by standards of performance particular to a task, the quality of support offered by the teacher, and the intellectual activity engaged in by learners... in broad strokes, equity in classroom instruction means providing opportunities for all students to learn challenging ideas, to participate in the characteristic activities of the discipline, and to be valued as important and fully human members of the science [and engineering] learning community” (p. 1101). this integrated conceptualization of rigor and equity in science communities focuses on teaching practices, but not on the broader socio-historical power structures that lead to and perpetuate inequity. in the next section, we describe how the i2 framework connects equitable, consequential learning with the creation of an inclusive culture. 4.2. dimension 2: connecting classroom and department/organization levels through inclusive climate we define inclusive climate as an environment that values and advances justice, equity, diversity, inclusion, student learning and development, and departmental/organizational community. our framework connects the inclusion and learning of students and department members (e.g., faculty davis, nolen & koretsky 31 | f l r and staff) at two levels: the classroom level and department level. at a classroom level, an inclusive climate brings explicit focus to issues of justice, equity, diversity, and inclusion to curriculum and instruction through equitable, consequential learning. for example, an inclusive climate attends to equitable access to learning opportunities that help all students—particularly those traditionally marginalized in engineering and society—to develop a sense of belonging and disciplinary identification. it does so through the development and implementation of complex, open-ended problems, activities, and assignments that benefit from the collaborative attention of multiple individuals with diverse identities and experiences. in addition, an inclusive classroom climate requires equitable and just interpersonal practices, including supportive relationships with faculty and peers, appropriate facilitation of group work, and an absence of (and appropriate acknowledgement and response to, when present) harassment, bias, and microaggressions. an inclusive classroom climate also attends to just and equitable student outcomes through the development and implementation of equitable grading policies that value and promote consequential learning; explicit valuation of inclusive teamwork skills, distributed expertise, and social and political responsibility; and ensuring that all students (again, especially those from marginalized communities) gain the knowledge, skills, and dispositions valued by the discipline and disciplinary community. at a department/organization level, an inclusive climate attends to issues of access, experiences, and outcomes by bringing explicit focus to issues of justice, equity, diversity, and inclusion to departmental policies, practices, and norms. for students, this includes access to and ongoing support for equitable, consequential learning opportunities as well as equitable departmental student outcomes (for example, retention, graduation, and graduate employment). students should leave with the knowledge, skills, and dispositions valued by their discipline or field and that contribute to a just global society. this includes the knowledge and skills to disrupt inequitable and unjust norms and practices within workplaces and the field. achieving these goals requires equitable practices for entry and continuation in the major or program; additional resources and supports for students from marginalized groups; equitable access to co-curricular opportunities, including internships, labs, and research experiences; and equitable access to the classes needed for the program. for faculty, an inclusive climate embraces equitable policies and practices, beginning with more equitable access to graduate school, mentoring opportunities, and faculty hiring processes (liera, 2020; posselt, 2020; white-lewis, 2020). an inclusive climate also facilitates equitable, consequential learning opportunities for faculty, including high quality mentoring and professional development opportunities for women and faculty of color (kachchaf et al., 2015; o’meara et al., 2017; o’meara & terosky, 2010). in an inclusive and just climate, students and faculty experience equitable and just interpersonal interactions and practices with peers, academic leaders, and other stakeholders. in the following sections, we describe a study that provides an empirical example of the i2 framework in action, revealing both the promise and challenge of an approach that integrates learning and deij goals across classroom and department/organizational levels. 5. context and methodology we illustrate the i2 framework using an empirical example from a multi-year organizational and instructional change project in a multidisciplinary engineering department at a large, research-focused public university in the united states. as co-designers and researchers on this project, we began with the foundational framing that this initiative would be a collaborative partnership among faculty, students, change advocates, administrators, staff, and researchers to improve deij and learning. the empirical example highlights a need to shift the nature of the work of students, faculty, administrators, staff, and researchers; how we (the change community) do the work; and the environment and culture within which we do the work at both the classroom level and the departmental/organizational level. we davis, nolen & koretsky 32 | f l r (the authors) emphasize how work at the departmental/organizational level is needed to disrupt ingrained social practice as small groups work collaboratively in the classroom. finally, we discuss implementation challenges and implications for integrating deij and learning opportunities through research and practice. the twin objects of the project were to create: (1) a culture where everyone in the departmental community feels a sense of being valued and belonging, and (2) a learning environment in which students and faculty meaningfully relate learning activities and experiences to each other and to professional practice (koretsky et al., 2018). we (the authors) were part of the change team and related change effort. using a design-based implementation research (dbir) approach (penuel et al., 2011; sabelli & dede; 2013) and our evolving sociocultural theoretical framework, we pursued these objects simultaneously, collecting and analyzing data that was used to inform the ongoing change efforts (davis et al., 2023; koretsky et al., 2018; lutz et al., 2019; michor et al., 2019; nolen et al., 2024). these iterative analyses informed change strategies, which included professional development for faculty, changes to curricula and instructional practices, and changes to the reward structure to support faculty working to advance learning and deij goals. the i2 framework was developed through this iterative, collaborative change process. for purposes of illustration, in this paper we describe the case of one instructional innovation (studio 2.0) that spanned the department’s core classes. the change project built on a previous departmental reform that aimed to promote active learning by shifting core engineering courses to a new studio structure (koretsky, 2015; koretsky et al., 2018). here, large undergraduate lecture courses (100350 students) were complemented by smaller group “studio” sessions with approximately 24 students where graduate student teaching assistants (gtas) facilitated undergraduate students’ collaborative work on engineering problems (mainly in groups of three). in the original innovation, known as studio 1.0, students largely worked on sequestered, conceptually oriented worksheet-type problems typical of many reform stem classrooms (finkelstein & pollock, 2005; koretsky, 2015; koretsky et al., 2018). the studio 2.0 reform aimed to engage students in more equitable, consequential work that would help them develop valued professional knowledge, skills, attitudes, and behaviors. the data that underlie this analysis were collected over five years and include faculty and administrator interviews; a student survey about perceptions of learning; an annual student climate survey; student focus groups; observations at departmental meetings and workshops; classroom observations and video recording; and artifacts from classrooms, the department, and the institution. this was a large project on multiple levels, with iterative data collection and analysis informing the change process throughout the project following our dbir methodology. we draw on several previously published analyses of particular aspects of the study (e.g., davis et al., 2023; koretsky et al., 2018; nolen et al., 2024), supplemented by more recent analyses of additional faculty interviews and observational and documentary evidence to create a more holistic picture of how this framework both derived from and informed the ongoing departmental reform. these analyses illuminated the ways in which creating and sustaining change in deij-centered problem-based learning at the classroom level necessitated the development of new pedagogical practices and supports for these changes at the departmental level. in the following sections, we illustrate the i2 framework using this empirical example, describing how the studio 2.0 reform worked to shift, at both the classroom and department level, (1) the nature of the work, (2) the ways in which faculty engaged in the work, and (3) the environment and culture within which the work was accomplished. davis, nolen & koretsky 33 | f l r 6. the i2 framework in action: an empirical example 6.1. change at the classroom level 6.1.1. shifting the nature of the work at the classroom level, the studio 2.0 project aimed to better engage students in disciplinary practice, changing the nature of students’ work, as well the practices of instructors. the primary aims of the department’s shift to studio 2.0 were to engage students in more authentic engineering problems better suited to small groups and to facilitate students’ development of the practices and identities of social justice-minded engineers able to work within equitable and inclusive teams (koretsky et al., 2018). studio 2.0 problems were changed, distinguishing them from studio 1.0 problems as follows: they were situated in the world of engineering professional practice (i.e., contextualized problems within specific engineering workplaces rather than abstract and decontextualized problems); they were often open-ended, having multiple possible solution paths (rather than instructor-prompted sequential steps to get to “the right answer”), and required more collaboration (e.g., asking students to collaboratively make design recommendations). we encouraged changes in assessment within studios to make concrete a shift in the object of activity. assessment shifted from a focus on individual summative assessment (e.g., a worksheet graded for completion and accuracy) to group formative assessment emphasizing learning, collaboration, and progress on the task (e.g., students’ collaborative engagement in and progress on substantive, realistic problems). this shift in assessment further supported students’ developing disciplinary identities and their understanding of how engineering knowledge and concepts might be usable and used within professional practice. by shifting the object of the activity and nature of related assessments, studio work helped students (and instructors) shift from “school world” to “engineering world” and thus positioned students as developing engineers (koretsky et al., 2018; nolen et al., 2024). through video analysis of small groups, we have identified two distinct forms of engagement. in “school world” mode, learning activity is seen for its transaction value (to “get the points” or “satisfy the instructor.”) in contrast, when engaging in “engineering world” (or, alternately, a “real” world aligned with a different discipline or profession), the object of a team’s activity is to apply disciplinary concepts and practices to create, design, analyze, and optimize processes (or other correspondingly disciplinary-valid ways). the premise here is that the ways of thinking and knowing in engineering world better align with the activity students will undertake in professional practice. by eliciting engineering world engagement, the work that we ask students to do will more likely result in adaptable, flexible, and transferable knowledge and skills. correspondingly, engagement in school world and engineering world depends on different social arrangements. in school world, groups approach the task as if it is best tackled from a single perspective or a “divide and conquer” approach wherein students divide tasks and therefore have little engagement as a group, reinforcing existing social hierarchies and privileging dominant interaction styles. in engineering world, students appear open to alternate perspectives and strategies, are willing to listen to and question each other in pursuit of an initial approach, and are flexible enough to change tactics when necessary. group members need to develop and employ the social skills that foster participation by all, including those with social identities that otherwise can be marginalized in engineering school. engaging students in disciplinary practice in this way has the potential to fundamentally address issues of broad dissatisfaction with schooling and inequitable participation and opportunity to learn. because the wide array of engineering practices offers numerous avenues for legitimate engagement of learners, learning environments that engage students in engineering practice can support access by a more diverse set of learners. through subsequent participation in such activities, learning in engineering and identity development in engineering become linked and inseparable. to become successful engineers, students must learn to engage productively with diverse stakeholders, multiple perspectives, davis, nolen & koretsky 34 | f l r and others with different funds of knowledge (gonzález et al., 2005). coursework, such as studio 2.0 problems, that positions students on inclusive teams as engineers doing consequential work helps students identify equity and social justice as central aspects of engineering itself. this perspective aligns with broader notions of academic rigor discussed above (windschitl & calabrese barton, 2016). this type of inclusive, collaborative work must be supported by instructional practices within an environment of social caring (appleby et al., 2021). thus, changing the nature of the students’ work also requires changes to how we—all stakeholders in higher education systems, and particularly instructors—do the work of education. 6.1.2. shifting how we do the work instructors needed to develop new pedagogical practices, changing the nature of their work to support student learning and new, more equitable forms of engagement (murtonen et al., 2022). in our project, changing the characteristics of classroom activity to be more equitable and consequential required more from instructors than simply adopting the kind of research-based instructional practices (e.g., active learning; use of complex problems) that receive so much attention in higher education (freeman et al., 2014; lombardi et al., 2021), including in the inclusive excellence literature (milem et al., 2005). though adopting research-based instructional practices likely requires professional development (murtonen et al., 2022), creating classroom activities that are more open-ended, contextualized, and situated in “engineering world” (or a different disciplinary “real world”) is a more complex endeavor that involves shifts in an instructor’s knowledge of how people learn and their identity as an instructor. for example, complex, discipline-relevant problems require more collaboration among students with differing lived experiences, cultural backgrounds, epistemologies, knowledge, and skills. when studio 2.0 problems required multiple perspectives, instructors had new responsibilities to disrupt longstanding participation patterns. students have spent years in traditional “school world” activities that reward having the “smartest,” fastest, most high-status students take the lead in group work and discussions. to help students truly benefit from complex problems in the discipline, instructors needed to also help them value multiple perspectives and develop inclusive collaboration skills (cohen & lotan, 1997; horn, 2012; nolen et al., 2024). changes in assessment practice were important in highlighting changes in the value framework for classroom activity. productive, inclusive teamwork involves equitable participation patterns, group-wide engagement, collaborative thinking and co-construction, and the use and production of shared work objects or representations (horn, 2012; windschitl & calabrese barton, 2016; kang et al, 2016). furthermore, it requires that instructors start to value productive friction, where the dilemmas and discrepancies that a team might face can lead to new ideas. such “glorious confusion” is the necessary precursor to deeper learning as well as immersion in engineering (or another disciplinary) world (horn, 2012; michor et al., 2019). just as the work becomes more complex and open-ended for students, so it becomes more complex and open-ended for instructors. these shifts in instructional practices required department-level support, as discussed in section 6.2. 6.1.3. shifting the environment and culture within which we do the work changing both the nature of the work and instructors’ pedagogical approaches also necessitated changes to classroom and department structures, policies, practices, values, and cultures—particularly given that advancing equity and justice was embedded within these goals and practices. at the classroom level, this shift required a rethinking of what counts as work and how to measure progress (for both instructors and students); a shift in values towards equity, justice, and realistic disciplinary practice; an acknowledgement of the assets and resources diverse students bring to the classroom (gonzález et al., 2005); and respect for the whole individual. previous research has found that environments where students feel cared for by peers and instructors support learning and the development of disciplinary practices (appleby et al., 2021). classrooms needed to become spaces where students felt welcome and that they belonged. davis, nolen & koretsky 35 | f l r within our own program, we developed and implemented an annual undergraduate survey to help understand the departmental climate, including attention to classroom learning and learning environments. peer relations and microaggressions predicted students’ identification with engineering. students’ open-ended responses also highlighted the importance of peer relations for their sense of belonging in the classroom and the department (davis et al., 2023). it follows that creating classroom environments where peers interact inclusively and recognize the assets diverse students bring may increase students’ disciplinary identification. climate survey results were reported to department faculty and integrated into professional development opportunities, an example of change at the department level, discussed next. 6.2. change at the department level over the course of the project, our change team initiated a series of department-based supports to facilitate changes at the classroom level in the nature of the work of engineering education, how we do the work, and the environment and culture within which we do the work. while the department and classroom levels are presented separately for clarity, they are mutually supportive and interdependent. 6.2.1. studio 2.0 classroom supports department-level supports were necessary to help faculty successfully create and implement the kinds of complex, contextualized, collaborative problems that shift the nature of the activity and create more equitable participation patterns within classrooms. for example, changing the nature of the work required reorganization of the course and program structure to accommodate studio 2.0 problems and activities. for each studio course, the program added smaller sections of about 24 students that met in between lectures to facilitate small group collaborative activities. to further support this major reorganization of the program curriculum, members of the change team worked with the registrar and the university space committee to secure two dedicated classrooms for studio 2.0. classrooms were adjacent to enable instructors to visit two studios within a single class period. space within classrooms was reconfigured to support small group collaboration (e.g., installing movable tables that could support shared work and be rearranged as needed). graduate teaching assistants (gtas) supported these studio sections. further, the department adopted a new undergraduate learning assistant (la) program modeled after the program from the colorado learning assistant alliance (gray et al., 2016). las provided additional “near peer” support for students’ learning in studios (koretsky et al., 2018). importantly, professional development programs were implemented for both gtas and las where they learned pedagogical principles and social practices to support equitable, consequential learning. 6.2.2. a departmental faculty community of practice the change team established a studio 2.0 community of practice (cop) to support the development of new pedagogical practices, values, and norms that would allow faculty to implement changes to the nature of classroom activity to facilitate equitable, consequential learning. the cop was designed to help departmental faculty develop shared values and goals for the studio model. several studio course instructors met regularly for a term and created a set of instructional design principles for studio 2.0, informed by the learning and deij goals of the overall change project as well as program objectives. participating faculty collectively determined the specific instructional principles that should guide curricular and pedagogical decision making and implementation. these design principles, shown in table 1, were based on three overarching propositions: (1) there are multiple ways to contribute productively to a team, (2) engineering problems have multiple solution paths, and (3) engineers need to make progress despite incomplete knowledge of the problem. davis, nolen & koretsky 36 | f l r table 1 instructional design principles for studio 2.0 identified by instructors during the cop design principle example group-worthy problems as much as possible, make problems challenging enough so that multiple perspectives become valued. include some problems that have multiple solution paths. practice first learn principles by doing cooperative learning facilitate inclusive interactions and ‘situated’ learning looping revisiting concepts within a course and between courses in the program revisit context weave the same context into studios for multiple courses… [to] further develop previously learned knowledge and skills assessment emphasis should be placed on the process of making progress and less emphasis on getting the answer formatting for cognitive load align studio delivery so it is as similar as possible between sections manageable change take baby steps in transitioning from studio 1.0 to 2.0 the cop also created a set of proposed “components of disciplinary knowledge” to inform the design and implementation of program curriculum. the components emphasized the broader set of competencies needed in professional practice (treveylen, 2014); for example, they included open-ended design, computational tools, communication and writing, hands-on experience, and inclusive teamwork. in this way, deij goals were supported by explicitly recognizing the distributed assets and experiences that are needed for engineering work. the department’s curriculum committee then discussed and iterated upon the proposal. finally, faculty voted to adopt the components at a faculty meeting, reifying departmental goals related to learning and deij. together, the instructional design principles and components of disciplinary knowledge created by the cop constituted an agreed-upon framework to guide faculty members as they shifted the nature of their course activities. 6.2.3. professional development opportunities to further support faculty members’ ability to design and implement equitable, consequential studio 2.0 problems, the change team created a week-long studio 2.0 workshop held the summer following the studio 2.0 cop. faculty learned more about the education research that supported the instructional design principles and components of disciplinary knowledge and had opportunities to (re)design a studio activity for a course they taught in line with this framework. participants received feedback during this design process from peers in the workshop, as well as a learning scientist and a faculty member from the university’s education department specializing in collaborative learning. the workshop emphasized the relationship between problem-based learning and deij through attention to inclusive teaming practices, different forms of collaborative engagement, shifting assessment strategies, and valuing students’ diverse funds of knowledge. davis, nolen & koretsky 37 | f l r two institutional resources helped department faculty more fully appreciate the connections between equitable, consequential learning and deij. the university annually offered two 60-hour summer workshops about understanding and applying knowledge of systems of oppression to professional practice: the advance workshop, focusing on equity-minded leadership, and the difference, power, and discrimination (dpd) academy, which focused on developing more inclusive and equitable curricular and instructional practices. seventeen faculty members in the department, including several who participated in the studio 2.0 cop and summer workshop, went through either the advance or dpd workshops, in many cases with financial support from the change project. three departmental faculty members and a change team researcher (the first author) went through the dpd academy together with the goal of learning how to support more inclusive and equitable teamwork in their courses and department. they then created, along with two more faculty members and a postdoctoral researcher, an ongoing inclusive teaming professional learning community (plc). members of the inclusive teaming plc met regularly and designed curricular content, pedagogies, and assessment tools and metrics for inclusive, socially just teaming practices. plc members tried these new tools and approaches in their own classes, evaluated results individually and as a group, and used their experiences to iteratively improve these tools and practices (lutz et al., 2019). 6.2.4. department-level policy changes to implement these changes to the nature of the work and how the work is done, faculty needed recognition for new forms of work, including time and effort on professional development and course development. several strategies were developed to shift the departmental environment and culture to support new forms of activity aimed at equitable, consequential learning. 6.2.4.1. position descriptions one strategy allowed faculty to modify and tailor their position descriptions to align with their professional activity. the goal was to appropriately reward faculty who engaged in transformative work by revising the reward structure of the department. at the university level, annual review of faculty is based on individualized position descriptions (pds). our change team worked with faculty and the department head to allow faculty to customize their pds to better match their current and desired activities, including involvement in department reform activities as well as other deij-focused work. we worked with administrators to emphasize the value of pds as a tool to help interested faculty pursue deij and teaching-related interests and be appropriately rewarded for those contributions to the department and the institution. we imagined that the individualized pds might carry weight during annual reviews and help faculty get appropriate credit for their curricular transformation work, including adapting their courses guided by the studio 2.0 framework. 6.2.4.2. alternating leads model change team members, in collaboration with the curriculum committee, developed a new departmental teaching structure. the alternating leads co-teaching model was designed to support faculty members to develop and improve innovative course activities and delivery. in this model, faculty pairs share a course assignment, with one focused on course delivery and the other on curricular development and integration of key skills and approaches that are interwoven throughout the program curriculum as articulated through the components of disciplinary knowledge (e.g., working in teams, writing, using computational tools). this model had several goals: to institutionalize and support continuous curricular and pedagogical innovation; to give faculty time and credit for ongoing course development; and to have the instructors model to students collaborative and inclusive practices through their work with each other. in practice, alternating leads also provided a new support structure for more inclusive and meaningful interactions among students and between students and instructors. individual pairs had the autonomy to determine more specifically how they would share roles and responsibilities for course development and delivery, but both instructors were expected to attend class regularly. davis, nolen & koretsky 38 | f l r 6.2.4.3. teaching 10s to have more regular, visible opportunities for faculty members to reflect on, discuss, and learn more about effective curricular and pedagogical practices, the department instituted a “teaching 10”— ten minutes of teaching-focused time at the start of each faculty meeting. during each teaching 10, a lead facilitator would focus on a specific topic, discussing education research related to the topic and/or how they approached the topic in their own classes. discussion among the whole faculty followed. sometimes a provocative statement was simply provided (e.g., “most exams students take are too long”) followed by 10 minutes of discussion. 6.2.5. providing multiple entry points for stakeholders shifts were required in both classroom and departmental environments to generate the kinds of systemic changes that would advance both learning and deij. the suite of activities described above worked together to support new and more equitable forms of collaborative engagement for students and instructors. importantly, it addressed a significant challenge of systemic change: engaging a majority of stakeholders, who have different interests, expertise, and perspectives (davis, 2023; kezar, 2018). this suite of change activities provided multiple entry points for faculty, staff, and administrators, and required that both members of our change team and department leadership embrace the essential linkages between deij and learning. while specific activities within the approach might foreground deij or learning, attending them in an integrated way enhanced the potential to positively impact both goals. though there is extra work involved in coordinating multiple efforts, including negotiating among different perspectives, there is also value in creating time and space for people with different perspectives and motives to engage in joint work to determine a shared vision and work towards it (davis, 2023; wenger, 1998). in our project, different people engaged in different change activities (e.g., advance and dpd workshops, the studio 2.0 cop and summer workshop, alternating leads, and the inclusive teaming plc). though some department members participated in several of these activities, most were actively engaged in only one or two. activities were designed to be synergistic, such that participation in any activity helped advance learning and deij goals. the multiple entry points provided by this multilayer framework made room for people with different motivations and interests to learn and engage with change efforts. 6.3. implementation challenges and successes creating systemic and sustainable change in higher education is notoriously difficult, and frequently fails (curry, 1992; kezar, 2018; rowley & sherman, 2001). one challenge of this work in higher education is the need for collaboration among faculty, staff, and administrators to support both individual and organizational learning (allen, 2004; banta, 1996; banta & palomba, 2015; davis, 2023; kezar, 2018; maki, 2010; mentkowski & loacker, 2002). collaboration and distributed engagement are at odds with longstanding traditions of autonomy and academic freedom and run counter to a reward structure that incentivizes individual achievement over collective growth (goodlad et al., 1990; hamilton, 2002; kezar, 2018; olivas, 1993). collective goals (in chat terms, objects) such as creating more equitable, consequential learning experiences, environments, and outcomes within an academic program, require the engagement of many, if not all, stakeholders. for example, if only some faculty, staff, and administrators participate in learning and incorporating more inclusive practices, there may be countervailing policies and practices at work in the department that undermine the reform effort. for example, students may still be subject to culturally offensive comments from nonparticipating faculty or feel frustrated or undermined by advising policies that do not take their lived experiences into account, especially if their advisors have not been engaged in the change effort. the net effect of less-than-full participation may well be that students do not experience a more broadly inclusive climate despite a department’s significant investment and efforts. davis, nolen & koretsky 39 | f l r the higher education environment creates many challenges for creating a truly collaborative endeavor that engages most, if not all, stakeholders and also attends to interactions among multiple aspects of the system, such as instructional practices and interactions, curricular development, assessment practices, norms related to engagement within group work, learning opportunities for faculty, and incentives and rewards for faculty to change their practices (davis, 2023). in chat terms, this kind of collaborative work creates a contradiction between the traditional division of labor in higher education (e.g., one instructor in charge of one class; a committee created to make recommendations for admissions criteria, which an administrator then has decision-making authority to implement or not) and what must be collective efforts reflecting common values in order to reach a shared object. this tension between the division of labor and the object proved salient in our own change project; working collectively to create and advance a shared vision was a challenging task. we provide three illustrative examples next. 6.3.1. alternating leads the alternating leads model encountered fundamental tensions with the existing division of labor, rules, tools, and values that made this innovation significantly harder to garner support for and sustain than other change efforts. the rules, policies, and norms of the institution reflected those in higher education, notably that promotion and tenure guidelines were designed to support a singleinstructor model. the historic practice in this department—following that of the institution—was to assign individual faculty to teach specific courses. each instructor was given significant leeway about how to teach their assigned course and was evaluated based on students’ and peers’ perceptions of their effectiveness as sole instructor. both administrators and students lacked experience evaluating coinstructors, particularly where one instructor served the more visible role, leading class lectures, while the other instructor worked more in the background, developing course content and activities and working with small groups in studio. preexisting norms within the department (and the larger institution within which it operated) valued teaching lecture-style classes over facilitating small group work, putting the instructors working on course redesign at a perceived disadvantage. we found in interviews that both faculty and administrators held the belief that the institution valued studio development and implementation less than lecturing. junior faculty—worried about implications for tenure and promotion—felt pressure to be the “lead” instructor (implementing the lectures). thus, in several classes using the alternating leads model, junior tenure-track faculty led the lecture section and non-tenure-track instructors focused on studio course development and implementation, and the leads did not alternate, reproducing unjust status hierarchies in the institution. from an administrative perspective, the alternating leads model also created perceived challenges with course coverage across the program (another division of labor issue). ultimately, the alternating leads model clashed with too many pre-existing norms, policies, and values at department, college, and university levels, and the administrative support it received was insufficient to allow the model to survive. in our assessment, further changes within the system would have been required for full implementation of the alternating leads model: updated teaching evaluation tools, protocols, and values; updated promotion and tenure guidelines that valued the work of studio development and implementation; and other solutions for course coverage. while there is still some co-teaching happening in the department, the implementation of the alternating leads model fell short of the scale envisioned. 6.3.2. learning assistant program in contrast, the la program has been significantly more successful than alternating leads at garnering the administrative and policy support necessary to reach its stated goals. prior to studio 2.0, the department often used undergraduate students to support instruction, especially for grading. with studio 2.0, undergraduate las shifted from grading support to helping facilitate small group collaborative learning. though the implementation of the la program increased costs for undergraduate davis, nolen & koretsky 40 | f l r instructional support, this program represented a minor shift (from the administrative perspective) and expansion of existing practices rather than a departure from them. instructors, particularly those least comfortable with the new emphasis on small group problem-based learning, appreciated additional facilitation support. administrators found additional financial support when needed to support the program. undergraduate students willingly attended additional training that helped them develop new skills. unlike the alternating leads program, the la program did not conflict with existing norms, policies, and practices, and thus was taken up more readily by faculty and administrators. 6.3.3. timescale and funding tensions another consideration for complex change projects, especially those working to change culture, is the timescale of implementation. the la program could be accomplished within a shorter period of time and within the timescale of the grant because it worked within existing structures and norms at both the departmental and institutional levels. initiatives like alternating leads that challenge existing structures, policies, cultures, and norms require a longer timescale for change to take hold, which also means they are more likely to face the additional challenge of leadership turnover. this does not mean that these more transformative projects should not be undertaken, but rather that change teams should design a change trajectory that accounts for the types of interventions they need to meet their goals and for inevitable pushback against any fundamental challenge to the status quo. thus, interventions that can be accomplished in shorter periods of time and within existing structures can be a starting point that can help gain traction, build a foundation, and provide interim successes on a longer-term trajectory of institutional and culture change (reay et al., 2006; termeer & dewulf, 2019). this conclusion suggests the need for longer-term grant opportunities from funders, given the significant amount of time it takes for multilevel systemic change that involves transformation of existing structures, policies, practices, and norms, and values. short of funders providing longer-term support for cultural and institutional change, change teams will need to plan to secure multiple grants sequentially to provide resources for longer-term change goals. 7. conclusion and implications for nearly 20 years in the united states and beyond, increasing attention has been paid to supporting deij in higher education, as well as improving students’ learning (bauman et al., 2005; blackie et al., 2016; european commission/ecea/eurydice, 2022; langholz, 2014; mejia & martin, 2023; milem et al., 2005; pineda & mishra, 2023; williams et al., 2005). at the same time, scholars have increasingly emphasized the integration of learning and identity (agarwal & sengupta-irving, 2019; rahm & moore, 2016), including in higher education, where students are often learning to become particular kinds of professional people (davis et al., 2023; gilbuena et al, 2015; horn et al., 2008; turner & nolen, 2015). in addition to learning concepts and principles, university students are learning particular social practices as part of their emerging professional identities. understanding the integrated nature of learning and identity provides a way to approach organizational change that can advance both learning and deij goals. the i2 framework presented herein integrates equity and learning goals at classroom and department/institution levels. by integrating efforts to advance learning with efforts to advance deij, i2 provides a model for a systemic, multilevel approach to inclusive excellence through the development of equitable, consequential learning opportunities and an inclusive climate at classroom and departmental levels. we argue that one cannot think about the environment where learning occurs (whether it be student or faculty/administrator learning) as separate from what is being learned and how (e.g., the learning activities and social relationships within which those activities are carried out). davis, nolen & koretsky 41 | f l r i2 prompts us to rethink our conceptions of learning and “student success,” moving beyond conversations about persistence and graduation metrics to attend to the interdependent nature of learning and identity. organizational change guided by i2 facilitates these interwoven learning and identityforming processes through changes to the nature of the work of education, how we accomplish the work, and the environment and culture within which we do the work. i2 accounts for instructors’ coconstitutive learning and identity formation processes as well, and how those can be supported at the classroom and departmental/institutional levels. without integration, activities aimed at deij goals may fall short on learning goals, and vice versa. for example, we can teach faculty about the ways in which systems of oppression operate and how to become more equitable in their interactions, but if the work students do in their classes is oriented towards a single, canonical answer, students will likely arrange themselves in standard, inequitable ways. in our change initiative, instructors who had taken a 60-hour deij training needed further support, through the studio 2.0 summer workshop or the inclusive teaming plc, to understand how designing more complex, open-ended problems for students to work collaboratively could help them pursue deij in practice. when we change the nature of the work so that it is open-ended, complex, and challenging, group members’ diverse competencies and funds of knowledge become resources for one another. thus, opportunities to engage in just social practice in the context of equitable, consequential work position historically marginalized students as engineers, by both peers and faculty, contributing to a sense of belonging and identification with the discipline as well as their learning (adams, 2001; davis et al., 2023; rogelberg & rumery, 1996). such pedagogical shifts towards equitable, consequential learning require attention to departmental (and institutional) policies, practices, and tools. in our own program, this transition required changing the structure of core classes to include smaller-group “studio” sections to support what students learned in lecture each week. within these studios, students engaged in the kind of situated pedagogies described above, facilitated by gtas and undergraduate las overseen by course instructors. changing the structure of classes alone was not enough. faculty, gtas, and las needed support and training to understand and be able to implement situated, cooperative, and inclusive pedagogical practices. other changes to practices, such as increased co-teaching, also supported such pedagogical shifts, but ran into challenges where they conflicted with more ingrained institutional policies, such as instructor evaluation and tenure and promotion policies, that proved harder to change. i2 also addresses the role of intersectional power dynamics (collins & bilge, 2020; svihla et al., 2023). workplaces reflect and perpetuate systems of oppression in the greater society, producing through organizational policies and practices “inequality regimes” that marginalize and devalue certain stakeholders based on their social identities or professional positionality (acker, 2006). organizational change aimed at deij goals therefore requires deliberate attention to the ways in which structural and cultural practices within an organization distribute power and privilege (acker, 2006; armstrong & jovanovic, 2015; svihla et al., 2022). equitable, consequential learning requires changes to occur and persist at multiple levels, from the classroom to the department to the institution and beyond—for example, the larger sociopolitical environment, including national and global political forces (dahlberg et al., 2021; de clercq et al., 2021). for example, teaching workload models or tenure and promotion policies can lead faculty to feel pressure to spend less time on teaching or not try new pedagogical approaches because of competing expectations. these policies will impact faculty’s experiences within their department and their approach to teachingand research-related workplace practices, impacting students’ learning and experiences in the classroom. systems-oriented change in primary and secondary schools and in higher education is difficult work and prone to failure (kezar, 2018; sarason, 1990). those attempting to address both deij and learning goals in an integrated way, and through an approach that addresses both classroom and departmental levels, are likely to face challenges that pull them back towards the status quo. goals, values, policies, and practices originating from a change initiative may meet resistance when they disrupt policies and practices central to an institution’s existing practice, as we saw with the attempted implementation of the alternating leads co-teaching model. furthermore, though we did not focus on davis, nolen & koretsky 42 | f l r the larger sociopolitical environment, forces such as the black lives matter movement and political resistance to critical race theory in education matter in shaping policies and pressures within institutions, departments, and classrooms (dahlberg et al., 2021; de clercq et al., 2021). though the i2 framework was crafted through analysis of a major change effort at only one institution, i2 provides a model for others to think about how to design such change and how it might play out in their unique context. for example, in the review of red abstracts discussed above and in a cross-site study of red programs (davis et al., 2024; svihla et al., 2023), most programs foregrounded learning or deij, rather than incorporating these goals in a meaningful way at both classroom and departmental levels. notably, i2 asks change agents to consider deij and learning as part of the same phenomenon rather than separate, parallel goals. this linkage has implications for change efforts both within and beyond engineering, and also within and beyond higher education. though primary and secondary school systems have unique cultures and challenges (sarason, 1990), schools could also benefit from taking a systems-oriented, multilevel approach that holistically integrates learning and deij goals in classrooms and the larger school community. despite regional variation in conceptualizations and foci of deij and learning efforts, these issues have increasing relevance internationally, suggesting utility of this framework within and beyond the united states (blackie et al., 2016; direito et al., 2021; european commission/ecea/eurydice, 2022; mejia & martin, 2023; pineda & mishra, 2023; walden et al., 2020). i2 prompts users to consider the ways in which the nature of the work, how the work is accomplished, and its environment are all connected and should be approached in an integrated way. this multidimensional framework for equitable, consequential learning guides those trying to change education environments to be more just and equitable, and to concurrently further consequential learning for all, by designing and implementing changes to the policies, practices, and values that guide classroom, department, and institutional levels. keypoints in education, initiatives that address improving diversity, equity, inclusion, and justice (deij) too often are conceptualized and implemented separately from those addressing improved learning. likewise, research on classroom reform is often pursued without consideration of departmental or organizational contexts, policies, and practices, or teachers’/faculty members’ learning, development, and experiences. deij and improving learning opportunities are mutually constitutive and synergistic and should be addressed using a systemic, multilevel approach that considers classroom, department, and organizational contexts. we present the i2 framework, termed i2 for integration within and across two dimensions: one representing goals of reform (deij and improved learning) and another representing levels of organizational change (classroom and department/organization). i2 integrates equitable, consequential learning opportunities in the classroom and an inclusive climate at the departmental/organizational level, emphasizing their inherent relatedness. we illustrate i2 with an empirical example from a systemic change initiative in a multi-program engineering department at a public university in the united states. the example highlights a need to shift the nature of the work of education, how we (the change community) do that work, and the environment and culture within which we do the work at both the classroom level and the department/organization level. davis, nolen & koretsky 43 | f l r our experience with a systemic change project highlighted the difficulty of change initiatives that challenge existing structures, policies, cultures, and norms at departmental and/or organizational levels. this does not mean that these more transformative projects should not be undertaken, but rather that change teams should design a change trajectory that accounts for the types of interventions they need to meet their goals and for inevitable pushback against any fundamental challenge to the status quo. acknowledgments this material is based upon work supported by the national science foundation under grant # 1519467 and 2236163. any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the national science foundation. the authors are grateful for generative conversations with michelle bothwell and acknowledge the influence of her deep commitment to social justice. we appreciate the students, faculty, and staff who contributed to this project in so many ways. references acker, j. (2006). inequality regimes: gender, class, and race in organizations. gender and society, 20(4), 441–464. https://doi.org/10.1177/0891243206289499 adams, s. g. (2001). the effectiveness of the e-team approach to invention and innovation. journal of engineering education, 90(4), 597-600. https://doi.org/10.1002/j.2168-9830.2001.tb00645.x agarwal, p., & sengupta-irving, t. (2019). integrating power to advance the study of connective and productive disciplinary engagement in mathematics and science. cognition and instruction, 37(3), 349–366. https://doi.org/10.1080/07370008.2019.1624544 allen, m. j. (2004). assessing academic programs in higher education. anker publishing. alpay, e., ahearn, a. l., graham, r. h., & bull, a. m. j. (2008). student enthusiasm for engineering: charting changes in student aspirations and motivation. european journal of engineering education, 33(5-6), 573-585. https://doi.org/10.1080/03043790802585454 appleby, l., dini, v., withington, l., lamotte, e., & hammer, d. (2021). disciplinary significance of social caring in postsecondary science, technology, engineering, and mathematics. physical review physics education research, 17(2), 23106. https://doi.org/10.1103/physrevphyseducres.17.023106 arastoopour, g., chesler, n. c., & shaffer, d. w. (2014). epistemic persistence: a simulation-based approach to increasing participation of women in engineering. journal of women and minorities in science and engineering, 20(3), 211–234. https://doi.org/10.1615/jwomenminorscieneng.2014007317 armstrong, m. a., & jovanovic, j. (2015). starting at the crossroads: intersectional approaches to institutionally supporting underrepresented minority women stem faculty. journal of women and minorities in science and engineering, 21(2). https://doi.org/10.1615/jwomenminorscieneng.2015011275 banta, t. w. (1996). assessment in practice: putting principles to work on college campuses. josseybass. https://doi.org/10.1002/j.2168-9830.2001.tb00645.x https://doi.org/10.1103/physrevphyseducres.17.023106 https://doi.org/10.1103/physrevphyseducres.17.023106 https://doi.org/10.1103/physrevphyseducres.17.023106 https://doi.org/10.1615/jwomenminorscieneng.2015011275 https://doi.org/10.1615/jwomenminorscieneng.2015011275 https://doi.org/10.1615/jwomenminorscieneng.2015011275 davis, nolen & koretsky 44 | f l r banta, t. w., & palomba, c. a. (2015). assessment essentials: planning, implementing, and improving assessment in higher education (2nd edition). jossey-bass. barr, r. b., & tagg, j. (1995). from teaching to learning—a new paradigm for undergraduate education. change: the magazine of higher learning, 27(6), 12-26. https://doi.org/10.1080/00091383.1995.10544672 barron, b., & darling-hammond, l. (2008). teaching for meaningful learning. in l. darlinghammond, b. barron, p. d. pearson, a. h. schoenfeld, e. k. stage, t. d. zimmerman, g. n. cervetti, & j. tilson (eds.), powerful learning: what we know about teaching for understanding. jossey-bass. basu, s. j., & calabrese barton, a. (2007). developing a sustained interest in science among urban minority youth. journal of research in science teaching, 44(3), 466–489. https://doi.org/10.1002/tea.20143 bauman, g. l., bustillos, l. t., bensimon, e. m., christopher brown ii, m., & bartee, r. d. (2005). achieving equitable educational outcomes with all students: the institution’s roles and responsibilities. association of american colleges and universities. beasley, m. a., & fischer, m. j. (2012). why they leave: the impact of stereotype threat on the attrition of women and minorities from science, math and engineering majors. social psychology of education, 15(4), 427–448. https://doi.org/10.1007/s11218-012-9185-3 beddoes, k., ihsen, s., vigild, m. e., mitchell, j., panther, g., murphy, m., williams, b., & sanchez ruiz, l. m. (2018). sefi position paper on diversity, equality and inclusiveness in engineering education. https://www.sefi.be/publication/sefi-position-paper-on-diversityequality-and-inclusiveness-in-engineering-education/ bensimon, e. m., dowd, a. c., stanton-salazar, r., & dávila, b. a. (2019). the role of institutional agents in providing institutional support to latinx students in stem. review of higher education, 42(4), 1689–1721. https://doi.org/10.1353/rhe.2019.0080 bishop, j., & verleger, m. a. (2013, june). the flipped classroom: a survey of the research. in american society of engineering education 2013 annual conference & exposition. blackburn, h. (2017). the status of women in stem in higher education: a review of the literature 2007-2017. science and technology libraries, 36, 235-273. https://doi.org/10.1080/0194262x.2017.1371658 blackie, m., le roux, k., & mckenna, s. (2016). possible futures for science and engineering education. higher education, 71(6), 755–766. https://doi.org/10.1007/s10734-015-9962-y bleich, m. r., macwilliams, b. r., & schmidt, b. j. (2015). advancing diversity through inclusive excellence in nursing education. journal of professional nursing, 31(2), 8994. https://doi.org/10.1016/j.profnurs.2014.09.003 brown, p. r., mccord, r. e., matusovich, h. m., & kajfez, r. l. (2015). the use of motivation theory in engineering education research: a systematic review of literature. european journal of engineering education, 40(2), 186-205. https://doi.org/10.1080/03043797.2014.941339 bryson, b. s., masland, l., & colby, s. (2020). strategic faculty development: fostering buy-in for inclusive excellence in teaching. the journal of faculty development, 34(3), 107-116. cabrera, a. f., nora, a., terenzini, p. t., pascarella, e., & hagadorn, l. s. (1999). campus racial climate and the adjustment of students to college: a comparison between white students and african-american students. the journal of higher education, 70(2), 134–160. https://doi.org/10.1080/00221546.1999.11780759 https://doi.org/10.1080/00091383.1995.10544672 https://doi.org/10.1002/tea.20143 https://doi.org/10.1016/j.profnurs.2014.09.003 davis, nolen & koretsky 45 | f l r cech, e. a. (2014). culture of disengagement in engineering education? science, technology, & human values, 39(1), 42–72. https://doi.org/10.1177/0162243913504305 chang, m. j., sharkness, j., hurtado, s., & newman, c. b. (2014). what matters in college for retaining aspiring scientists and engineers from underrepresented racial groups. journal of research in science teaching, 51(5), 555–580. https://doi.org/10.1002/tea.21146 christersson, c., staaf, p., braekhus, s., stjernqvist, r., pusineri, a. g., giovani, c., ... & zhang, t. (2019). promoting active learning in universities. thematic peer group report. european university association. https://eua.eu/downloads/publications/eua%20tpg%20report%205%20promoting%20active%20learning %20in%20universities.pdf christie, m., & de graaff, e. (2017). the philosophical and pedagogical underpinnings of active learning in engineering education. european journal of engineering education, 42(1), 5-16. https://doi.org/10.1080/03043797.2016.1254160 cohen, e. g., & lotan, r. a. (1997). raising expectations for competence: the effectiveness of status interventions. in e. g. cohen & r. a. lotan (eds.), working for equity in heterogeneous classrooms: sociological theory in practice. teachers college press. collins, a., brown, j. s., & holum, a. (1991). cognitive apprenticeship: making thinking visible. american educator, 15(3), 6-11. collins, p. h., & bilge, s. (2020). intersectionality (2nd ed.). polity press. https://www.wiley.com/enus/intersectionality%2c+2nd+edition-p-9781509539673 considine, j. r., mihalick, j. e., mogi‐hein, y. r., penick‐parks, m. w., & van auken, p. m. (2017). how do you achieve inclusive excellence in the classroom? new directions for teaching and learning, 151, 171-187. https://doi.org/10.1002/tl.20255 curry, b. (1992). instituting enduring innovations: achieving continuity of change in higher education. george washington university. dahlberg, g. m., vigmo, s., & surian, a. (2021). widening participation? (re)searching institutional pathways in higher education for migrant students-the cases of sweden and italy. frontline learning research, 9(2), 145–169. https://doi.org/10.14786/flr.v9i2.655 davis, s. c. (2023). engaging faculty in data use for program improvement in teacher education: how leaders bridge individual and collective development. teaching and teacher education, 129, 1–14. https://doi.org/10.1016/j.tate.2023.104147 davis, s. c., kellam, n., sanders, j., & svihla, v. (2024). integrating theories of intersectional power, learning, and change to explore faculty experiences on equity-centered change projects. journal of diversity in higher education. advance online publication. https://doi.org/10.1037/dhe0000601 davis, s. c., nolen, s. b., cheon, n., moise, e., & hamilton, e. w. (2023). engineering climate for marginalized groups: connections to peer relations and engineering identity. journal of engineering education, 112(2), 284–315. https://doi.org/10.1002/jee.20515 de clercq, m., jansen, e., brahm, t., & bosse, e. (2021). from micro to macro: widening the investigation of diversity in the transition to higher education. frontline learning research, 9(2), 1–8. https://doi.org/10.14786/flr.v9i2.783 diggs, g. a., garrison-wade, d. f., estrada, d., & galindo, r. (2009). smiling faces and colored spaces: the experiences of faculty of color pursuing tenure in the academy. the urban review, 41(4), 312–333. https://doi.org/10.1007/s11256-008-0113-y https://doi.org/10.1002/tea.21146 https://doi.org/10.1080/03043797.2016.1254160 https://doi.org/10.1002/tl.20255 https://doi.org/10.14786/flr.v9i2.655 https://doi.org/10.1037/dhe0000601 https://doi.org/10.1007/s11256-008-0113-y davis, nolen & koretsky 46 | f l r direito, i., chance, s., clemmensen, l., craps, s., economides, s. b., isaac, s. r., jolly, a. m., truscott, f. r., & wint, n. (2021). diversity, equity, and inclusion in engineering education: an exploration of european higher education institutions’ strategic frameworks, resources, and initiatives. proceedings of the sefi 49th annual conference: blended learning in engineering education, 189–193. engeström, y. (1987). learning by expanding: an activity-theoretical approach to developmental research. orienta-konsultit oy. engeström, y. (2001). expansive learning at work: toward an activity theoretical reconceptualization. journal of education and work, 14(1), 133–156. https://doi.org/10.1080/13639080123238 engle, r. a. (2012). the productive disciplinary engagement framework: origins, key concepts, and developments. in d. yun dai (ed.), design research on learning and thinking in educational settings (pp. 170-209). routledge. engle, r. a., & conant, f. r. (2002). guiding principles for fostering productive disciplinary engagement: explaining an emergent argument in a community of learners classroom. cognition and instruction, 20(4), 399-483. https://doi.org/10.1207/s1532690xci2004_1 espinosa, l. l. (2011). pipelines and pathways: women of color in undergraduate stem majors and the college experiences that contribute to persistence. harvard educational review, 81(2), 209– 241. https://doi.org/10.17763/haer.81.2.92315ww157656k3u european commission. (2021). european research area policy agenda: overview of actions for the period 2022-2024. https://doi.org/10.2777/52110 european commission, directorate-general for education, youth, sport, and culture. (2024). the twin challenge of equity and excellence in basic skills in the eu – an eu comparative analysis of the pisa 2022 results. publications office of the european union. https://data.europa.eu/doi/10.2766/881521 european commission/ecea/eurydice. (2022). towards equity and inclusion in higher education in europe. publications office of the european union. https://doi.org/10.2797/046055 everett, s., & grey, t. g. (2016). creating inclusive excellence: a model for culturally relevant teacher education. urban education research & policy annuals, 4(2), 72-88. finkelstein, n. d., & pollock, s. j. (2005). replicating and understanding successful innovations: implementing tutorials in introductory physics. physical review special topics-physics education research, 1(1), 010101. https://doi.org/10.1103/physrevstper.1.010101 forde, t., & carpenter, r. (2020). situating inclusive excellence in faculty development programs and practices. journal of faculty development, 34(3). https://link.gale.com/apps/doc/a651906895/aone?u=oregon_oweb&sid=googlescholar&xid =9bcbeec1 freeman, s., eddy, s. l., mcdonough, m., smith, m. k., okoroafor, n., jordt, h., & wenderoth, m. p. (2014). active learning increases student performance in science, engineering, and mathematics. proceedings of the national academy of sciences, 111(23), 8410-8415. https://doi.org/10.1073/pnas.1319030111 garvey, j. c., & rankin, s. s. (2018). the influence of campus climate and urbanization on queerspectrum and trans-spectrum faculty intent to leave. journal of diversity in higher education, 11(1), 67–81. https://doi.org/10.1037/dhe0000035 gilbuena, d. m., sherrett, b. u., gummer, e. s., champagne, a. b., & koretsky, m. d. (2015). feedback on professional skills as enculturation into communities of practice. journal of engineering education, 104(1), 7-34. https://doi.org/10.1002/jee.20061 https://doi.org/10.1080/13639080123238 https://doi.org/10.17763/haer.81.2.92315ww157656k3u https://data.europa.eu/doi/10.2766/881521 https://doi.org/10.2797/046055 https://doi.org/10.1073/pnas.1319030111 https://doi.org/10.1002/jee.20061 davis, nolen & koretsky 47 | f l r gonzález, n., moll, l. c., & amanti, c. (eds.). (2005). funds of knowledge: theorizing practices in households, communities, and classrooms. lawrence earlbaum associates. goodlad, j. i., soder, r., & sirotnik, k. (1990). places where teachers are taught. jossey-bass. gray, k. e., webb, d. c., & otero, v. k. (2016). effects of the learning assistant model on teacher practice. physical review physics education research, 12(2), 020126. https://doi.org/10.1103/physrevphyseducres.12.020126 greeno, j. g., & engeström, y. (2014). learning in activity. in r. k. sawyer (ed.), the cambridge handbook of the learning sciences (2nd edition, pp. 128–148). cambridge university press. https://doi.org/10.1017/cbo9781139519526.009 hall, r., & jurow, a. s. (2015). changing concepts in activity: descriptive and design studies of consequential learning in conceptual practices. educational psychologist, 50(3), 173–189. https://doi.org/10.1080/00461520.2015.1075403 hall, r. m. & sandler, b. r. (1982). the classroom climate: a chilly climate for women? association of american colleges. hamilton, n. w. (2002). academic ethics: problems and materials on professional conduct and shared governance. american council on education/praeger. harris, j. c. (2020). multiracial faculty members’ experiences with teaching, research, and service. journal of diversity in higher education, 13(3), 228–239. https://doi.org/10.1037/dhe0000123 hart, j. (2016). dissecting a gendered organization: implications for career trajectories for mid-career faculty women in stem. journal of higher education, 87(5), 605–634. https://doi.org/10.1353/jhe.2016.0024 hausmann, l. r. m., schofield, j. w., & woods, r. l. (2007). sense of belonging as a predictor of intentions to persist among african american and white first-year college students. research in higher education, 48(7), 803-839. https://doi.org/10.1007/s11162-007-9052-9 henderson, c., beach, a., & finkelstein, n. (2011). facilitating change in undergraduate stem instructional practices: an analytic review of the literature. journal of research in science teaching, 48(8), 952–984. https://doi.org/10.1002/tea.20439 herrenkohl, l., & mertl, v. (2010). how students come to be, know, and do: a case for a broad view of learning. cambridge university press. holen, a., & sortland, b. (2022). the teamwork indicator–a feedback inventory for students in active group learning or team projects. european journal of engineering education, 47(2), 230–244. https://doi.org/10.1080/03043797.2021.1985435 holland, d. (2010). symbolic worlds in time/spaces of practice: identities and transformations. in b. wagoner & b. wagoner (eds.), symbolic transformation: the mind in movement through culture and society (pp. 269-283). routledge/taylor & francis group. holland, d., skinner, d., lachiotte, w., jr., & cain, c. (1998). identity and agency in cultural worlds. harvard educational press. horn, i. s. (2012). strength in numbers. national council of teachers. howard hughes medical institute. (n.d.a). inclusive excellence 1 & 2. https://www.hhmi.org/scienceeducation/programs/inclusive-excellence-1-2 howard hughes medical institute. (n.d.b). inclusive excellence 3 learning community. https://www.hhmi.org/science-education/programs/inclusive-excellence-3-learningcommunity https://www.hhmi.org/science-education/programs/inclusive-excellence-1-2 https://www.hhmi.org/science-education/programs/inclusive-excellence-1-2 https://www.hhmi.org/science-education/programs/inclusive-excellence-3-learning-community https://www.hhmi.org/science-education/programs/inclusive-excellence-3-learning-community davis, nolen & koretsky 48 | f l r hulleman, c. s., kosovich, j. j., barron, k. e., & daniel, d. b. (2017). making connections: replicating and extending the utility value intervention in the classroom. journal of educational psychology, 109(3), 387-404. https://doi.org/10.1037/edu0000146. hurtado, s., alvarez, c. l., guillermo-wann, c., cueller, m., & arellano, l. (2012). a model for diverse learning environments: the scholarship on creating and assessing conditions for student success. in j. c. smart & m. b. paulsen (eds.), higher education: handbook of theory and research (vol. 27, pp. 41-). springer science & business media. https://doi.org/10.1007/97894-007-2950-6 hurtado, s., griffin, k. a., arellano, l., & cuellar, m. (2008). assessing the value of climate assessments: progress and future directions. journal of diversity in higher education, 1(4), 204–221. https://doi.org/10.1037/a0014009 hurtado, s., & guillermo-wann, c. (2013). diverse learning environments: assessing and creating conditions for student success final report to the ford foundation. university of california, los angeles: higher education research institute. hurtado, s., milem, j. f., clayton-pedersen, a. r., & allen, w. r. (1998). enhancing campus climates for racial/ethnic diversity: educational policy and practice. review of higher education, 21(3), 279-302. https://dx.doi.org/10.1353/rhe.1998.0003 hurtado, s. h. & ponjuan, l. (2005). latino educational outcomes and the campus climate. journal of hispanic higher education, 4(3), 235–251. https://doi.org/10.1177/1538192705276548 jordt, h., eddy, s. l., brazil, r., lau, i., mann, c., brownell, s. e., ... & freeman, s. (2017). values affirmation intervention reduces achievement gap between underrepresented minority and white students in introductory biology classes. cbe—life sciences education, 16(3). https://doi.org/10.1187/cbe.16-12-0351 jurow, a. s., teeters, l., shea, m., & van steenis, e. (2016). extending the consequentiality of 'invisible work' in the food justice movement. cognition and instruction, 34(3), 210-221. https://doi.org/10.1080/07370008.2016.1172833 kachchaf, r., hodari, a., ko, l., & ong, m. (2015). career-life balance for women of color: experiences in science and engineering academia. journal of diversity in higher education, 8(3), 175–191. https://doi.org/10.1037/a0039068 kang, h., windschitl, m., stroupe, d., & thompson, j. (2016). designing, launching, and implementing high quality learning opportunities for students that advance scientific thinking. journal of research in science teaching, 53(9), 1316-1340. http://dx.doi.org/10.1002/tea.21329 kaptelinin, v. (2005). the object of activity: making sense of the sense-maker. mind, culture, and activity, 12(1), 4–18. https://doi.org/10.1207/s15327884mca1201 kezar, a. (2018). how colleges change: understanding, leading, and enacting change (second edition). routledge. koretsky, m. d. (2015). program level curriculum reform at scale: using studios to flip the classroom. chemical engineering education, 49(1), 47-57. https://journals.flvc.org/cee/article/view/84260 koretsky, m. d., montfort, d., nolen, s. b., bothwell, m., davis, s. c., & sweeney, j. (2018). towards a stronger covalent bond: pedagogical change for inclusivity and equity. chemical engineering education, 52(2), 117–127. https://journals.flvc.org/cee/article/view/105859 kurth, l. a., anderson, c. w., & palincsar, a. s. (2002). the case of carla: dilemmas of helping all students to understand science. science education, 86(3), 287-313. https://doi.org/10.1002/sce.10009 https://dx.doi.org/10.1353/rhe.1998.0003 https://doi.org/10.1177/1538192705276548 https://doi.org/10.1187/cbe.16-12-0351 https://doi.org/10.1080/07370008.2016.1172833 https://doi.org/10.1037/a0039068 https://doi.org/10.1037/a0039068 http://dx.doi.org/10.1002/tea.21329 https://doi.org/10.1002/sce.10009 davis, nolen & koretsky 49 | f l r langholz, m. (2014). the management of diversity. management revue, 25(3), 207–226. https://doi.org/10.4324/9780367824044-7 lave, j., & wenger, e. (1991). situated learning: legitimate peripheral participation. cambridge university press. lesh, r., hoover, m., hole, b., kelly, a., &post, t. (2000). principles for developing thought revealing activities for students and teachers. in a. kelly & r. lesh (eds.), the handbook of research design in mathematics and science education, (pp. 591–646). lawrence erlbaum associates. liera, r. (2020). moving beyond a culture of niceness in faculty hiring to advance racial equity. american educational research journal, 57(5), 1954–1994. https://doi.org/10.3102/0002831219888624 lima, r. m., andersson, p. h., & saalman, e. (2017). active learning in engineering education: a (re) introduction. european journal of engineering education, 42(1), 1-4. https://doi.org/10.1080/03043797.2016.1254161 lombardi, d., shipley, t. f., bailey, j. m., bretones, p. s., prather, e. e., ballen, c. j., knight, j. k., smith, m. k., stowe, r. l., cooper, m. m., prince, m., atit, k., uttal, d. h., ladue, n. d., mcneal, p. m., ryker, k., st. john, k., van der hoeven kraft, k. j., & docktor, j. l. (2021). the curious construct of active learning. psychological science in the public interest, 22(1), 8– 43. https://doi.org/10.1177/1529100620973974 lutz, b. d., bothwell, m. k., auyeung, n., carlisle, t. k., mallette, n., & davis, s. c. (2019, april). practitioner learning community: design of instructional content, pedagogy, and assessment metrics for inclusive and socially just teaming practices. proceedings of the conference of the collaborative network for engineering and computing diversity, crystal city, va. https://peer.asee.org/31781 maass, k., cobb, p., krainer, k., & potari, d. (2019). different ways to implement innovative teaching approaches at scale. educational studies in mathematics, 102, 303-318. https://doi.org/10.1007/s10649-019-09920-8 mahar, m. t., baweja, h., atencio, m., barkhoff, h., duley, h. y., makuakāne-lundin, g., ... & russell, j. (2021). inclusive excellence in kinesiology units in higher education. kinesiology review, 1, 1-8. https://doi.org/10.1123/kr.2021-0042 maki, p. l. (2010). assessing for learning: building a sustainable commitment across the institution. stylus publishing. marra, r. m., rodgers, k. a., shen, d., & bogue, b. (2012). leaving engineering: a multi-year single institution study. journal of engineering education, 101(1), 6–27. https://doi.org/10.1002/j.2168-9830.2012.tb00039.x mejia, j. a., & martin, j. p. (2023). critical perspectives on diversity, equity, and inclusion research in engineering education. in a. johri (ed.), international handbook of engineering education research (pp. 218–238). taylor & francis. https://doi.org/10.4324/9781003287483-13 mentkowski, m., & loacker, g. (2002). enacting a collaborative scholarship of assessment. in t. w. banta & associates (eds.), building a scholarship of assessment. jossey-bass. michor, e. l., koretsky, m., & nolen, s. b. (2019, june). destigmatizing confusion–a path toward professional practice. 2019 american society of engineering education (asee) annual conference & exposition, tampa, fl, united states. miettinen, r. (2005). object of activity and individual motivation. mind, culture, and activity, 12(1), 17024141–17024142. https://doi.org/10.1207/s15327884mca1201 https://doi.org/10.4324/9780367824044-7 https://doi.org/10.3102/0002831219888624 https://doi.org/10.3102/0002831219888624 https://doi.org/10.3102/0002831219888624 https://doi.org/10.1080/03043797.2016.1254161 https://doi.org/10.1177/1529100620973974 https://peer.asee.org/31781 https://doi.org/10.1123/kr.2021-0042 https://doi.org/10.1002/j.2168-9830.2012.tb00039.x https://doi.org/10.1207/s15327884mca1201 davis, nolen & koretsky 50 | f l r milem, j. f., chang, m. j., & antonio, a. l. (2005). making diversity work on campus: a researchbased perspective. association of american colleges and universities. https://doi.org/10.1016/j.neurenf.2012.04.589 minnotte, k. l., & pedersen, d. e. (2021). turnover intentions in the stem fields: the role of departmental factors. innovative higher education, 46(1), 77–93. https://doi.org/10.1007/s10755-020-09524-8 murtonen, m., anto, e., laakkonen, e., & vilppu, h. (2023). university teachers’ focus on students: examining the relationships between visual attention, conceptions of teaching and pedagogical training. frontline learning research, 10(2), 64–85. https://doi.org/10.14786/flr.v10i2.1031 museus, s. d. (2014). the culturally engaging campus environments (cece) model: a new theory of college success among racially diverse student populations. in m. b. paulsen (ed.), higher education: handbook of theory and research. springer. museus, s. d., yi, v., & saelua, n. (2017). the impact of culturally engaging campus environments on sense of belonging. the review of higher education, 40(2), 187–215. https://dx.doi.org/10.1353/rhe.2017.0001 national research council. (2012). discipline-based education research: understanding and improving learning in undergraduate science and engineering (s.r. singer, n.r. nielsen, and h.a. schweingruber, eds.). national academies press. national science foundation. (n.d.a). improving undergraduate stem education: education and human resources (iuse: ehr). https://beta.nsf.gov/funding/opportunities/improvingundergraduate-stem-education-education-and-human-resources-iuse-ehr national science foundation. (n.d.b). advance: organizational change for gender equity in stem academic professions (advance). https://beta.nsf.gov/funding/opportunities/advanceorganizational-change-gender-equity-stem-academic-professions-advance national science foundation. (n.d.c). iuse/professional formation of engineers: revolutionizing engineering departments (iuse/pfe: red). https://beta.nsf.gov/funding/opportunities/iuseprofessional-formation-engineersrevolutionizing-engineering-departments nolen, s. b., michor, e. l., & koretsky, m. d. (2024). engineers, figuring it out: collaborative learning in cultural worlds. journal of engineering education, 113(1), 164– 194. https://doi.org/10.1002/jee.20576 nora, a., & cabrera, a. f. (1996). the role of perceptions of prejudice and discrimination on the adjustment of minority students to college. journal of higher education, 67, 119148. https://doi.org/10.1080/00221546.1996.11780253 olivas, m. a. (1993). reflections on professorial academic freedom: second thoughts on the third “essential freedom.” stanford law review, 45(6), 1835–1858. https://doi.org/10.2307/1229129 o’meara, k. a., griffin, k. a., nyunt, g., & lounder, a. (2018). disrupting ruling relations: the role of the promise program as a third space. journal of diversity in higher education, 12(3), 205–218. https://doi.org/10.1037/dhe0000095 o’meara, k. a., rivera, m., kuvaeva, a., & corrigan, k. (2017). faculty learning matters: organizational conditions and contexts that shape faculty learning. innovative higher education, 42(4), 355–376. https://doi.org/10.1007/s10755-017-9389-8 o’meara, k., & terosky, a. l. (2010). engendering faculty professional growth. change, 42(6), 44– 51. https://doi.org/10.1080/00091383.2010.523408 https://doi.org/10.1007/s10755-020-09524-8 https://doi.org/10.14786/flr.v10i2.1031 https://dx.doi.org/10.1353/rhe.2017.0001 https://beta.nsf.gov/funding/opportunities/improving-undergraduate-stem-education-education-and-human-resources-iuse-ehr https://beta.nsf.gov/funding/opportunities/improving-undergraduate-stem-education-education-and-human-resources-iuse-ehr https://beta.nsf.gov/funding/opportunities/advance-organizational-change-gender-equity-stem-academic-professions-advance https://beta.nsf.gov/funding/opportunities/advance-organizational-change-gender-equity-stem-academic-professions-advance https://beta.nsf.gov/funding/opportunities/iuseprofessional-formation-engineers-revolutionizing-engineering-departments https://beta.nsf.gov/funding/opportunities/iuseprofessional-formation-engineers-revolutionizing-engineering-departments https://doi.org/10.1002/jee.20576 davis, nolen & koretsky 51 | f l r ong, m., smith, j. m., & ko, l. t. (2018). counterspaces for women of color in stem higher education: marginal and central spaces for persistence and success. journal of research in science teaching, 55(2), 206-245. https://doi.org/10.1002/tea.21417 ong, m., wright, c., espinosa, l., & orfield, g. (2011). inside the double bind: a synthesis of empirical research on undergraduate and graduate women of color in science, technology, engineering, and mathematics. harvard educational review, 81(2), 172-209. https://doi.org/10.17763/haer.81.2.t022245n7x4752v2 palmén, r., arroyo, l., müller, j., reidl, s., caprile, m., & unger, m. (2020). integrating the gender dimension in teaching, research content & knowledge and technology transfer: validating the efforti evaluation framework through three case studies in europe. evaluation and program planning, 79(101751), 1-10. https://doi.org/10.1016/j.evalprogplan.2019.101751 palmer, r. t., maramba, d. c., & dancy, t. e. (2011). a qualitative investigation of factors promoting the retention and persistence of students of color in stem. the journal of negro education, 80(4), 491–504. https://muse.jhu.edu/article/806879 pawley, a. l. (2019). learning from small numbers: studying ruling relations that gender and race the structure of u.s. engineering education. journal of engineering education, 108(1), 13–31. https://doi.org/10.1002/jee.20247 penuel, w. r., fishman, b. j., cheng, b. h., & sabelli, n. (2011). organizing research and development at the intersection of learning, implementation, and design. educational researcher, 40(7), 331337. https://doi.org/10.3102/0013189x11421826 pineda, p., & mishra, s. (2023). the semantics of diversity in higher education: differences between the global north and global south. higher education, 85(4), 865–886. https://doi.org/10.1007/s10734-022-00870-4 posselt, j. r. (2020). equity in science: representation, culture, and the dynamics of change in graduate education. stanford university press. prince, m. (2004). does active learning work? a review of the research. journal of engineering education, 93(3), 223-231. https://doi.org/10.1002/j.2168-9830.2004.tb00809.x rahm, j., & moore, j. c. (2016). a case study of long-term engagement and identity-in-practice: insights into the stem pathways of four underrepresented youths. journal of research in science teaching, 53(5), 768–801. https://doi.org/10.1002/tea.21268 rankin, s. r., & reason, r. d. (2005). differing perceptions: how students of color and white students perceive campus climate for underrepresented groups. journal of college student development, 46(1), 43–61. https://dx.doi.org/10.1353/csd.2005.0008 rankin, s., & reason, r. (2008). transformational tapestry model: a comprehensive approach to transforming campus climate. journal of diversity in higher education, 1(4), 262–274. https://doi.org/10.1037/a0014018 raver, s. a., & maydosz, a. s. (2010). impact of the provision and timing of instructor-provided notes on university students’ learning. active learning in higher education, 11(3), 189-200. https://doi.org/10.1177/1469787410379682 reay, t., golden-biddle, k., & germann, k. (2006). legitimizing a new role: small wins and microprocesses of change. academy of management journal, 49(5), 977-998. https://doi.org/10.5465/amj.2006.22798178 riley, d., slaton, a. e., & pawley, a. l. (2014). social justice and inclusion: women and minorities in engineering. in a. johri & b. m. olds (eds.), cambridge handbook of engineering education https://doi.org/10.1002/tea.21417 https://doi.org/10.17763/haer.81.2.t022245n7x4752v2 https://muse.jhu.edu/article/806879 https://doi.org/10.3102/0013189x11421826 https://doi.org/10.1002/j.2168-9830.2004.tb00809.x https://dx.doi.org/10.1353/csd.2005.0008 https://doi.org/10.1177/1469787410379682 https://doi.org/10.5465/amj.2006.22798178 davis, nolen & koretsky 52 | f l r research (pp. 335–356). new york: cambridge university press. https://doi.org/10.1017/cbo9781139013451.022 rincón, b. e., & george-jackson, c. e. (2016). examining department climate for women in engineering: the role of stem interventions. journal of college student development, 57(6), 742–747. https://doi.org/10.1353/csd.2016.0072 rodriguez, s. l., lu, c., & bartlett, m. (2018). engineering identity development: a review of the higher education literature. international journal of education in mathematics, science and technology, 6(3), 254–265. https://doi.org/10.18404/ijemst.428182 rogelberg, s. g. & rumery, s. m. (1996). team decision quality, time on task, and interpersonal cohesion. small group research, 27(1), 79-90. https://doi.org/10.1177/1046496496271004 rolin, k. (2008). gender and physics: feminist philosophy and science education. science & education, 17, 1111-1125. https://doi.org/10.1007/s11191-006-9065-3 rowley, d. j., & sherman, h. (2001). from strategy to change: implementing the plan in higher education. san francisco: jossey-bass. salazar, m. d. c., norton, a. s., & tuitt, f. a. (2010). weaving promising practices for inclusive excellence into the higher education classroom. to improve the academy, 28(1), 208-226. https://doi.org/10.1002/j.2334-4822.2010.tb00604.x sandler, b., silverberg, l., & hall, r. (1996). the chilly classroom climate: a guide to improve the education of women. national association of women in education. sabelli, n., & dede, c. (2013). empowering design-based implementation research: the need for infrastructure. in b. fishman & w. r. penuel (eds.), design-based implementation research: theories, methods, and exemplars (vol. 112, pp. 464-480). new york: national society for the study of education. sarason, s. (1990). the predictable failure of educational reform: can we change course before it’s too late? jossey-bass. secules, s. (2019). making the familiar strange: an ethnographic scholarship of integration contextualizing engineering educational culture as masculine and competitive. engineering studies, 11(3), 196–216. https://doi.org/10.1 080/19378629.2019.1663200 settles, i. h., buchanan, n. t., & dotson, k. (2019). scrutinized but not recognized: (in)visibility and hypervisibility experiences of faculty of color. journal of vocational behavior, 113, 62–74. https://doi.org/10.1016/j.jvb.2018.06.003 settles, i. h., jones, m. k., buchanan, n. t., & brassel, s. t. (2022). epistemic exclusion of women faculty and faculty of color: understanding scholar(ly) devaluation as a predictor of turnover intentions. journal of higher education, 93(1), 31–55. https://doi.org/10.1080/00221546.2021.1914494 seymour, e., & hewitt, n. m. (1997). talking about leaving: why undergraduates leave the sciences. westview press. siri, a., leone, c., & bencivenga, r. (2022.) equality, diversity, and inclusion strategies adopted in a european university alliance to facilitate the higher education-to-work transition" societies, 12(5), 140. https://doi.org/10.3390/soc12050140 slaton, a. e. (2010). race, rigor and selectivity in u.s. engineering: the history of an occupational color line. harvard university press. smith, d. e. (1999). writing the social: critique, theory, and investigations. university of toronto press. https://doi.org/10.1017/cbo9781139013451.022 https://doi.org/10.1017/cbo9781139013451.022 https://doi.org/10.1017/cbo9781139013451.022 https://doi.org/10.1177/1046496496271004 https://doi.org/10.1002/j.2334-4822.2010.tb00604.x https://doi.org/10.1%2520080/19378629.2019.1663200 https://doi.org/10.1%2520080/19378629.2019.1663200 https://doi.org/10.1016/j.jvb.2018.06.003 https://doi.org/10.1080/00221546.2021.1914494 davis, nolen & koretsky 53 | f l r stetsenko, a. (2005). activity as object-related: resolving the dichotomy of individual and collective planes of activity. mind, culture, and activity, 12(1), 70–88. https://doi.org/10.1207/s15327884mca1201 stevens, r., o’connor, k., garrison, l., jocuns, a., & amos, d. m. (2008). becoming an engineer: toward a three dimensional view of engineering learning. journal of engineering education, 97(3), 355-368. https://doi.org/10.1002/j.2168-9830.2008.tb00984.x svihla, v., davis, s. c., & kellam, n. (2023). the triple change framework: merging theories of intersectional power, learning, and change to enable just, equitable, diverse, and inclusive engineering education. studies in engineering education, 4(2), 38–63. https://doi.org/10.21061/see.87 termeer, c. j. a. m., & dewulf, a. (2019). a small wins framework to overcome the evaluation paradox of governing wicked problems. policy and society, 38(2), 298–314. https://doi.org/10.1080/14494035.2018.1497933 theobald, e. j., hill, m. j., tran, e., agrawal, s., arroyo, e. n., behling, s., ... & freeman, s. (2020). active learning narrows achievement gaps for underrepresented students in undergraduate science, technology, engineering, and math. proceedings of the national academy of sciences, 117(12), 6476-6483. https://doi.org/10.1073/pnas.1916903117 tippeconnic fox, m. j. (2005). voices from within: native american faculty and staff on campus. new directions for student services, 2005(109), 49-59. https://doi.org/10.1002/ss.153 tonso, k. l. (2007). on the outskirts of engineering: learning identity, gender, and power via engineering practice. sense publishers. trevelyan, j. (2014). the making of an expert engineer. crc press. turner, c. s. v., gonzález, j. c., & wood, j. l. (2008). faculty of color in academe: what 20 years of literature tells us. journal of diversity in higher education, 1(3), 139–168. http://dx. doi.org/10.1037/a0012837 turner, j. c., & nolen, s. b. (2015). introduction: the relevance of the situative perspective in educational psychology. educational psychologist, 50(3), 167-172, https://doi.org/10.1080/00461520.2015.1075404 walden, s. e., direito, i., berhan, l., clavero, s., galligan, y., jolly, a.-m., specking, e., & vanasupa, l. (2020). asee & sefi joint statement on diversity, equity, and inclusion: a call and pledge for action. https://www.sefi.be/publication/asee-sefi-joint-statement-on-diversity-equity-andinclusion/ walton, g. m., peach, j. m., logel, c., spencer, s. j., & zanna, m. p. (2015). two brief interventions to mitigate a “chilly climate” transform women’s experience, relationships, and achievement in engineering. journal of educational psychology, 107(2), 468-485. http://dx.doi.org/10.1037/a0037461 wenger, e. (1998). communities of practice: learning, meaning, and identity. cambridge university press. white-lewis, d. k. (2020). the facade of fit in faculty search processes. journal of higher education, 91(6), 833–857. https://doi.org/10.1080/00221546.2020.1775058 williams, d. a., berger, j. b., & mcclendon, s. a. (2005). toward a model of inclusive excellence and change in postsecondary institutions. association of american colleges and universities. https://doi.org/10.1.1.129.2597 https://doi.org/10.1207/s15327884mca1201 https://doi.org/10.1002/j.2168-9830.2008.tb00984.x https://doi.org/10.1073/pnas.1916903117 https://doi.org/10.1002/ss.153 https://doi.org/10.1080/00221546.2020.1775058 davis, nolen & koretsky 54 | f l r windschitl, m., & calabrese barton, a. (2016). rigor and equity by design: seeking a core of practices for the science education community. in d. gitomer & c. bell (eds.), aera handbook of research on teaching (5th ed., pp. 1099-1158). aera press. yosso, t., smith, w., ceja, m., & solórzano, d. (2009). critical race theory, racial microaggressions, and campus racial climate for latina/o undergraduates. harvard educational review, 79(4), 659-691. https://doi.org/10.17763/haer.79.4.m6867014157m707l zambrana, r. e., harvey wingfield, a., lapeyrouse, l. m., dávila, b. a., hoagland, t. l., & valdez, r. b. (2017). blatant, subtle, and insidious: urm faculty perceptions of discriminatory practices in predominantly white institutions. sociological inquiry, 87(2), 207–232. https://doi.org/10.1111/soin.12147 https://doi.org/10.17763/haer.79.4.m6867014157m707l https://doi.org/10.1111/soin.12147 https://doi.org/10.1111/soin.12147 https://doi.org/10.1111/soin.12147 frontline learning research vol. 12 no. 4 (2024) 55 84 issn 2295-3159 corresponding author: mathias mejeh, university of bern, department of research in school and learning, institute of educational science, fabrikstrasse 8, 3012 bern, switzerland email: mathias.mejeh@unibe.ch doi: https://doi.org/10.14786/flr.v12i4.1387 understanding the promotion of self-regulated learning in upper secondary schools: how can teaching quality criteria contribute? mathias mejeh1,3, barbara stampfli2 & tina hascher1 1university of bern, switzerland 2bern university of teacher education, switzerland 3zurich university of teacher education, switzerland article received 15 october 2023 / article revised 5 august 2024 / accepted 8 november 2024 / available online 29 november 2024 abstract self-regulated learning (srl) has gained increasing attention in educational science over the past four decades. especially within the context of the lifelong learning debate, regulatory strategies play a crucial role, as they are essential not only within schools and classrooms but also in lifelong learning contexts. concurrently, the discussion on teaching quality holds an equally central position in educational science. however, these two lines of discourse have, so far, been treated largely independently of each other and lack an alignment. in our exploratory study, we applied the three basic dimensions of teaching quality to assess an srl-promoting learning environment from the perspective of students. we conducted seven focus groups involving a total of n = 49 secondary school students, and the data was analyzed using qualitative content analysis. our analyses demonstrate that these basic dimensions can contribute to analyze the quality of an srl environment. however, further adaptation is required as the three dimensions seem still to be interwoven with a more traditional conceptualization of teaching. keywords: self-regulated learning; teaching quality; focus groups; qualitative content analysis; instructional development mailto:mathias.mejeh@unibe.ch mejeh, stampfli & hascher 56 | f l r 1. introduction school instruction, as a central pillar for the education of all children and adolescents, serves not only as a means for the systematic, intergenerational transfer of knowledge, skills, abilities, norms, and values but also for fostering individual and collective competence and personality development. in this context, both what students learn and how they learn it are equally relevant. the concept of selfregulated learning (srl) is particularly significant, as it empowers children and adolescents to intentionally shape their competence and personality development (oecd, 2018). consequently, exploring strategies to promote srl lies at the heart of both research and educational practice, making it the central focus of this paper. we aim to contribute to a better understanding of the quality of srl in classrooms and schools by analyzing an srl-conducive learning environment through the lens of teaching quality. drawing from a qualitative study involving 49 upper secondary school students, we employ focus groups to investigate which elements of a learning environment tailored to srl significantly contribute to its effectiveness. 2. self-regulated learning: direct promotion and indirect activation in the school classroom srl is a fundamental educational principle for achieving successful student learning (oecd, 2019). the origins of srl research in educational settings are diverse, with roots tracing back to the 1960s and 1970s. srl draws from multiple theoretical frameworks and has been analyzed across various contexts, with key contributions from vygotsky (1962), flavell (1971), and bandura (1986), among others. as panadero (2017) demonstrates in his review article, contemporary srl models are strongly characterized by either a (meta-)cognitive (efklides, 2011; winne & hadwin, 1998), motivational (pintrich, 2004; zimmerman, 2000), or emotional (boekaerts, 2011) perspective. srl can be defined as the behavior exhibited by individuals who aim to enhance their knowledge and skills by actively monitoring and regulating their learning activities (paris & paris, 2001). specifically, srl is understood as a hierarchically organized, temporal, and adaptive process in which learners conduct a task analysis, develop goals, and create plans to solve the task. the achievement of these goals is facilitated using various learning strategies, with motivational and affective factors playing a critical role in initiating and sustaining goal attainment. learners continuously monitor and reflect on their learning processes and goal achievement through metacognitive strategies. the effective use of selfregulatory strategies depends on the specific tasks and the contextual factors of the learning environment (greene et al., 2021). against this backdrop, it becomes clear how important it is to create suitable learning environments where students can actively shape their learning processes. despite the valuable approaches described above, repeated criticism has lined out that models of srl have not dedicated enough attention to the interplay between individual learners and their learning environments (martin, 2007; mccaslin & good, 1996; perry & rahim, 2011). panadero (2017, p. 22) states that “...with the exception of hadwin, järvelä, and miller’s work, not much research has been conducted by the others in exploring how significantly other contexts or the task context affect srl.” moreover, a systematic definition on the quality of srl conducive learning environment remains open. this study aims to make an initial contribution toward addressing this gap. the novelty of our research lies not only in investigating how effectively learners perceive an srl-supportive learning environment but also in systematically identifying and analyzing the quality of such an environment through the lens of teaching quality. to address these questions, it is valuable to explore the distinction between direct strategy instruction and indirect activation of srl (dignath & veenman, 2020). promoting srl involves distinguishing between direct and indirect methods (dignath-van ewijk et al., 2013). direct support involves teachers explicitly teaching self-regulation strategies. consequently, recent studies increasingly emphasize the teacher's role in promoting srl in the classroom (e.g., perry, 2013; karlen et al., 2020; veenman, 2017; vosniadou et al., 2024). for instance, research by perels and colleagues mejeh, stampfli & hascher 57 | f l r (leidinger & perels, 2012; perels et al., 2009; venitz & perels, 2019) shows that teachers can significantly enhance students' srl by using appropriate teaching materials. moreover, directly training teachers to promote students' srl has shown to be effective (e.g., finsterwald et al., 2013; kistner et al., 2010), with teacher beliefs and self-efficacy playing a crucial role (e.g., heirweg et al., 2021; karlen et al., 2020). in this context, the significance of integrating srl promotion into teacher education programs at an early stage becomes particularly clear (kramarski, 2018; kramarski et al., 2013). in this regard, strategy teaching can be further differentiated into explicit and implicit approaches, with direct strategy instruction classified according to its degree of explicitness (dignath & büttner, 2008). brown and colleagues (1981) distinguish three levels of direct strategy instruction: blind, informed, and explicit self-control training. research consistently highlights the effectiveness of direct strategy training (donker et al., 2014). this effectiveness is evident both in the application of cognitive learning strategies (hattie et al., 1996) and in the integration of cognitive and metacognitive strategies (schuster et al., 2023; souvignier & mokhlesgerami, 2006). although numerous studies demonstrate the effectiveness of direct strategy instruction, the role of the learning environment in fostering srl skills is equally significant (karlen & hertel, 2024). a prominent example is the clia model (competence, learning, intervention, assessment), a framework designed to create learning environments that promote sustained knowledge acquisition and the development of students as competent learners and critical thinkers. the clia model emphasizes the promotion of self-regulation as a fundamental component of learning (de corte et al., 2004), with its effectiveness in enhancing various aspects of srl demonstrated across diverse educational contexts. for example, this learning environment has been shown to enhance elementary school students' mathematical problem-solving skills, beliefs, and attitudes (de corte et al., 2004), as well as their collaborative learning processes (de corte, 2012). furthermore, students have been observed to perform better and engage more intensively in metacognitive regulation strategies within srl-conducive learning environments (de corte, 2016; masui & de corte, 2005). a similar proposal was developed by perry (perry, 2013; perry et al., 2018; perry et al., 2020) highlighting various aspects of a classroom environment that emphasize srl. this framework categorizes classroom elements conducive to srl into four main groups: supportive structures for srl, student autonomy and influence, facilitating, guiding, and co-regulating, and functioning as a community. these macro categories are further divided into several micro categories that reflect specific actions teachers take in the classroom. teachers provide srl-supportive structures (1) by offering various activities, routines, and participation structures that support both individual and collaborative learning while ensuring inclusivity for diverse student needs and abilities. these structures include tasks, instructions, familiar routines, as well as visual prompts. for example, task setting involves diverse, advanced learning objectives achieved through real-world, extended, and multifaceted tasks. these tasks foster deep cognitive engagement and metacognition while offering flexibility in learning approaches and representations. in this context, perry et al. (2020) collaborated with teachers to create diverse tasks and evaluation methods that support srl. their findings show that srl-focused instruction promotes deep learning, positive emotions, and enhanced student achievement. student influence and autonomy (2) are fostered by acknowledging learners' perspectives and experiences, providing opportunities to take control of their learning. various forms of self-assessment, along with involvement in decision-making, are crucial to this process. in autonomy-enhancing learning environments, students experience more positive emotions about their learning process. they exhibit increased autonomous motivation (de naeghel et al., 2016), greater engagement, reduced amotivation (cheon & reeve, 2015), and a tendency to seek out more challenging tasks (su & reeve, 2011). support, scaffolding, and co-regulation (3) describe how teachers and peers can act key as key learning supports for students. srl-conducive learning environments feature metacognitive and motivational dialogue, modeling, demonstrations, and differentiated, reciprocal feedback. the importance of the social context has been demonstrated several times, for example, regarding scaffolding through teachers and peers (van leeuwen & janssen, mejeh, stampfli & hascher 58 | f l r 2019; molenaar et al., 2014; salonen et al., 2005) or collaborative learning (hadwin et al., 2018; järvenoja et al., 2018; mccaslin & vriesema, 2018; panadero et al., 2015; vriesema & mccaslin, 2020). creating a community of learners (4) refers to fostering a sense of belonging and group cohesion through participation structures. when a classroom cultivates a positive climate marked by recognition of individuality, mutual support, shared knowledge, and respectful communication, it effectively operates as a community. research indicates that fostering a learning community benefits srl by encouraging students to actively seek help and peer support (perry & drummond, 2002). from the perspective of school and instructional development, a central question is how schools and teachers can foster students' srl. building on previous research into the indirect activation of srl, it is crucial to examine which aspects of instruction students find effective and how they perceive specific design features of the learning environment as conducive to srl. while de corte and perry have contributed to conceptualizing learning environments conducive to srl, a systematic evaluation of their quality remains outstanding (dignath & veenman, 2020; muijs et al., 2014). 3. teaching quality research on teaching's impact on student learning, often termed “teaching quality” or “teaching effectiveness” (seidel & shavelson, 2007), has a longstanding tradition. research on effective teaching encompasses two fundamental objectives. first, it seeks to deconstruct complex teaching processes to identify core instructional characteristics. second, these characteristics are designed to encapsulate the essential elements of teaching that enable students to achieve their learning objectives (praetorius et al., 2020). this approach is further informed by two fundamental dimensions of instructional design: the “deep” structure and the “surface” structure (klieme et al., 2009). surface structures include organizational forms, social formats, and teaching methods (e.g., class groups, project work, group work, etc.). deep structures refer to teaching-learning processes, encompassing learning content examination, teacher-student interaction, and the overall professionalism of teaching and learning (seidel & shavelson, 2007). effective teaching is primarily identified through analyzing deep structures that underlie the different methods and formats teachers use (oser & baeriswyl, 2001). studies have identified criteria for teaching quality, such as a supportive classroom climate, meaningful discourse, effective task engagement scaffolding, explicit strategy instruction, and clear achievement expectations. these criteria were initially outlined by brophy and good (1986) and later expanded upon by brophy (2010), grouping them into primary categories within models such as the quait-model by slavin (1994) and the class-model by pianta and colleagues (2008). given that teaching serves multiple objectives, such as student knowledge acquisition, fostering student interest, and enhancing student social skills (brophy, 1999), a diverse range of models for evaluating teaching quality and effectiveness exists (e.g., ferguson & danielson, 2014; klieme et al., 2009; slavin, 1987, 1994). despite differences in focus, abstraction, and domain specificity, models generally agree on core characteristics of teaching quality, such as time on task or classroom management (praetorius et al., 2018). in the german-speaking academic context, the model of the three basic dimensions of teaching quality, as proposed by klieme and colleagues (2009), has gained prominence. this model consists of three core dimensions: (1) classroom management, (2) cognitive activation, and (3) student support. the model distinguishes between characteristics that enhance motivation and those that improve academic performance. an optimal classroom environment that fosters both motivation and performance seamlessly integrates all three dimensions. classroom management, as defined by praetorius et al. (2018), involves effective handling and prevention of disruptions, efficient use of time, monitoring, teacher presence, clear rules, and established routines. cognitive activation involves engaging students with challenging tasks, connecting to prior knowledge, and encouraging diverse mejeh, stampfli & hascher 59 | f l r problem-solving strategies. it also emphasizes moving away from uni-directional, teacher-focused instruction toward interactive, co-constructive learning, incorporating socratic methods, and fostering students' metacognitive processes. the supportive class climate dimension centers on students' perceptions of competence, autonomy, and social inclusion. extensive empirical research has been conducted on the basic dimensions of teaching quality. research shows that cognitive activation improves academic performance, while student support significantly boosts students' interest in the subject matter (e.g., klieme & rakoczy, 2003). classroom management has the most significant impact on student learning across all subjects (e.g., praetorius et al., 2014). this is likely because effective classroom management maximizes students' learning time (seidel & shavelson, 2007). in summary, teaching quality refers to the set of measurable teaching attributes that directly correlate with students' progress toward achieving educational objectives (klieme, 2019). the three basic dimensions are conceptualized as generic factors of teaching quality that apply across disciplines, making them independent of subject-specific content. it is reasonable to assume that these criteria also provide a framework for analyzing successful srl. therefore, it is crucial to investigate how students perceive the quality of learning environments designed to foster srl and which instructional features they regard as relevant and effective in this context. 4. analyzing a srl learning environment from the perspective of teaching quality research on successful srl has revealed the high importance of both the individual development of the students and the development of the learning environment for successful teaching. moreover, their interaction is of particular interest (dignath-van ewijk et al., 2013). as dignath and veenman (2020, p. 523) pointed out, “…researchers investigating teachers’ srl practice should consider results from generic teaching effectiveness research” and aim to clarify the extent to which specific instructional dimensions contribute to an understanding of effective srl. therefore, we suggest a systematic quality analysis of an srl-enhancing learning environment in using the three basic dimensions of teaching quality common in research on learning and instruction. this approach is innovative, as studies on teaching quality typically do not specifically target a particular teaching concept, such as srl, while research on srl only indirectly addresses the basic dimensions of effective teaching, if at all. by linking these dimensions to srl, there is potential for expanding our understanding of effective instruction. addressing these questions could make a lasting contribution to the development and evaluation of srl-promoting learning environments, while also enhancing overall teaching quality. 5. the present study the importance of srl at the upper secondary school level of education is particularly evident due to its crucial role in the sustained acquisition of knowledge and skills. moreover, it holds significant relevance for postsecondary education (vosniadou, 2020), as students need to be prepared for successful higher education completion, which necessitates the mastery of appropriate selfregulatory strategies (jansen et al., 2019; kitsantas et al., 2008; nandagopal & ericsson, 2012; peverly et al., 2003). teachers lay the foundation for learners to control, shape, and develop their own learning by intentionally designing srl-conducive learning environments (de corte, 2004; perry, 2013). however, despite srl being a central focus of classroom research for the past four decades (panadero, 2017), questions about the specific qualities of srl-conducive learning environments remain open (dignath & veenman, 2020; muijs et al., 2014). thus, there is a clear need for a systematic definition of the quality aspects of an srl-conducive learning environment. this paper aims to address this gap by exploring the following question: mejeh, stampfli & hascher 60 | f l r what aspects of the three basic dimensions do upper secondary school students describe in a learning environment conducive to self-regulated learning? in addressing this broad question, our study is exploratory in nature and aims to identify the potential benefits of analyzing the quality of srl-enhancing learning environments through the lens of the three basic dimensions of teaching quality. thus, we align two important yet distinct discourses in educational research as, to date, there has been little systematic effort to describe the quality characteristics of learning environments that promote srl. 6. methodology and data basis 6.1 study context this study was conducted in switzerland, focusing on students in upper secondary school. in switzerland, upper secondary school typically spans three to five years, catering to students aged 14 to 19, similar to other european upper secondary systems. our research involved examining one class each at the beginning, middle, and end of this educational phase. however, this was not a longitudinal study; instead, we conducted a cross-sectional analysis of different cohorts at various points in their upper secondary education. our partner school, gymnasium hofwil, initiated discussions about fostering srl during an internal teacher workshop involving all staff members. the discussions arose from the school’s dissatisfaction with its current narrow structure and students’ learning outcomes, alongside debates among teaching staff about the need for self-regulation and the prerequisites for successful learning in upper secondary education. the school drew inspiration from a comparable institution that had successfully implemented srl as part of its school development process. the sounding board within the school, comprising teachers, students, and school leaders, recognized the need for a more objective and multi-informant evaluation. after several rounds of internal feedback and evaluation, the school principal contacted one of the researchers to request an external, formative evaluation of the newly implemented teaching structure. this initiated a collaboration between the school and researchers from the university of bern. using the macro and micro categories of srl-conducive learning environments developed by perry and colleagues (perry et al., 2018), the upper secondary school and the researchers collaboratively evaluated and refined the teaching structure in a participatory manner (perry et al., 2020). to accomplish this, we applied three fundamental principles—separation of learning phases and assessment phases, fewer subjects per learning phase, and individual learning time—to perry’s framework (table 1). mejeh, stampfli & hascher 61 | f l r table 1 assignment of macro and micro categories of srl conducive learning environment according to perry et al. (2018) macro categories micro categories instructional setting in our study supportive structures for srl tasks/activities fewer subjects per phase / individual learning time expectations/instructions learning and assessment phases familiar routines and participation structures learning and assessment phases visual prompts fewer subjects per phase / individual learning time student autonomy and influence involvement in decision making/meaningful choices individual learning time control over challenge fewer subjects per phase self-assessment learning and assessment phases facilitating, guiding, and coregulating modeling/demonstrating fewer subjects per phase questioning fewer subjects per phase feedback learning and assessment phases / fewer subjects per phase metacognitive language fewer subjects per phase motivational messages fewer subjects per phase functioning as a community co-constructing knowledge individual learning time positive/non-threatening communication individual learning time supporting/celebrating one another’s learning individual learning time accommodations for individual differences individual learning time mejeh, stampfli & hascher 62 | f l r firstly, the school year was divided into learning and assessment phases, and the framework timetable was transformed into an annual schedule. to establish an srl-supportive learning environment, learning objectives and assignments were presented to students at the beginning of each phase. this ensured that students clearly understood what was expected of them and how to organize and plan their learning. to promote students' autonomy and influence, also a routine was established within the phases. after every five weeks of learning, an assessment phase was held, allowing students to demonstrate what they had learned and engage in self-assessment of their progress. this approach enabled learners to develop consistent learning routines and effective strategies for achieving learning objectives. accordingly, there were no assessments during the learning phases; students focused solely on learning subject content during joint lessons and working independently on assignments during individual learning. in the assessment phase, which lasted one week, summative examinations were conducted to assess the students' level of knowledge and development. secondly, only 3-4 subjects were taught simultaneously in each learning phase (fewer subjects per phase). traditionally, the weekly timetable remains the same every week. this structure was modified so that subjects were scheduled at specific times with greater intensity, ensuring that all subjects (e.g., languages, science, math) were still covered but with more hours per subject during each phase. although this changed the distribution of lessons throughout the school year, the school's educational administration requirements were still met. the adapted learning environment was designed to facilitate deeper immersion in subjects while heavily supporting, guiding, and coregulating in the classroom. with fewer subjects per phase, teachers met students more often, allowing them to use various forms of feedback (formative, summative, prognostic) during the learning phase, which supported the students in their learning in a more targeted way. at the same time, teachers used metacognitive language and supported students by setting learning goals at the beginning of each learning phase and by teaching specific metacognitive strategies. to further promote metacognitive strategies, students were encouraged to maintain learning journals. teachers acted as role models by using strategies such as mind mapping or in-class discussions to address task challenges together, creating srl-supportive structures. thirdly, 30% of instructional time was allocated to individual learning time to establish supportive structures for srl and increase student autonomy and influence (individual learning time). this allowed students to plan and decide independently when, with which tools, and in which social context they would complete their tasks, enhancing their involvement in decision-making. this approach aimed to enhance students' engagement with the subject matter and foster explicit selfregulation through the completion of more complex tasks. students could delve deeper into subjects at their own pace and according to their needs, promoting a classroom community where individual differences were valued. knowledge was co-constructed as teachers acted as learning coaches, guiding students through their learning processes. in case of challenges or questions, students were encouraged to contact the teachers but had to decide for themselves when and in what respect they needed support. the sense of community was further enhanced by offering various types of learning spaces. some spaces were designated for quiet, individual work, while others were intended for discussions and collaborative group work. this arrangement fostered lively discussions, shared learning experiences, and mutual support. students had opportunities to share insights and questions, give positive feedback, and offer mutual explanations. since some students excelled in certain subjects while others excelled in different areas, they could assist each other during individual learning time. 6.2 case selection the dataset comprises a total of seven group discussions. the sample includes 49 students, consisting of 29 females and 20 males, from three different school classes (9th, 11th, and 13th grades) where the srl-supportive learning environment was implemented. accordingly, students aged 14 to 19 were interviewed. to address our research question, we intentionally selected student group mejeh, stampfli & hascher 63 | f l r discussions as our primary method, as learners are central to educational and instructional efforts (scherer et al., 2016; wisniewski et al., 2020). to gain a comprehensive understanding of the learning environment across all grade levels, students with varying levels of experience were purposefully selected according to a predetermined qualitative sampling plan (patton, 2015). this involved examining one class at the beginning, middle, and end of upper secondary school. the participant selection was exhaustive and homogeneous within school classes, encompassing potentially all students within a class. however, between school classes, the selection was heuristic and heterogeneous, intentionally including specific classes in the sample. to facilitate year-specific analyses, group discussions were conducted within individual classes (lamnek, 2005). 6.3 data collection and data analysis the seven group discussions followed a semi-structured guide and were conducted in a conversational setting to replicate an everyday school environment as closely as possible (onwuegbuzie et al., 2009). each discussion involved five to nine participants. before the group discussions, participants were informed about the study and provided their consent for participation and data usage. the discussions were recorded and lasted between 55 and 80 minutes (mean duration = 70 minutes). the audio recordings were transcribed using f4/f5 software, following established guidelines (kuckartz & rädiker, 2019). verbatim transcription was employed, with minor language and punctuation corrections for readability. to maintain student anonymity, each participant was assigned a unique code. these codes prevented the identification of individual participants in the discussions but allowed for individual attribution (e.g., s1). the collected data were analyzed using a structured qualitative content analysis approach with maxqda 20. initially, a category system was deductively developed from existing theory. this category system was then inductively tested and expanded using the collected data (kuckartz & rädiker, 2019). the categories were based on the category system of praetorius et al. (2018) and were slightly modified through translation into german. the category system was organized into three main categories: class management, cognitive activation, and student support, mirroring the three basic dimensions of teaching quality (klieme et al., 2009). table 2 provides an example of the category system, describing one sub-category along with its definition and an anchor example. sense units were defined as coding units, allowing not only single sentences but also entire paragraphs to be assigned to the same code. to ensure inter-subjective comprehensibility, the codings were collectively discussed and compared using an interview as an illustrative example (kuckartz & rädiker, 2019; mayring, 2015). furthermore, interrater reliability was calculated, demonstrating an acceptable level of agreement, with a cohen's kappa value of 0.67 (mcdonald et al., 2019). mejeh, stampfli & hascher 64 | f l r table 2 excerpt from the category system subcategory definition anchor example c la ss ro o m m an ag em en t rule clarity and routines clarity and structure in classroom procedures; rules must be defined and handled equally by all. “i think it's important that the assignments are clear. and that you really know what you have to do and by when you have to do something.” c o g n it iv e a ct iv at io n activation of and link to previous knowledge building on students' existing knowledge, linking of different topics. “it is also very difficult to always find the entry point. because if you repeat something now to get back into it a bit, half of the phase is already over again and then you should actually be much further along and have started a new topic.” s tu d en t s u p p o rt differentiation and adaptive support individualization of assignments, e.g. different levels of difficulty and assistance; more advanced assignments for faster students; support for students who experience difficulties. “in the beginning, there were more teachers, and i had them explain it to me a few times, and then it actually worked. but you do need two or three weeks to find your way in.” 7. results a total of 503 codes were assigned to the three main categories: classroom management (n=160), cognitive activation (n=168), and student support (n=175), which were somewhat evenly distributed among the basic dimensions of teaching quality. the findings are presented and interpreted in relation to the three core characteristics of the studied learning environment: “separation of learning and assessment phase”, “fewer subjects per learning phase”, and “individual learning time”. the categorization and analysis of the interview material were conducted within the context of the three basic dimensions. when categorizing students' statements into these dimensions, the unique aspects of the srl-conducive learning environment were explicitly considered, leading to adjustments in the subcategories. the results are further explored with respect to each class level only when differences between the class levels were identified; otherwise, they represent an overall perspective based on the statements of all learners. for readability, student quotes were edited to remove repetition, pauses, or filler words. mejeh, stampfli & hascher 65 | f l r 7.1 separation of learning and assessment phase table 3 summarizes the results regarding the separation of the learning and assessment phases. the following sections present the results for the three main categories in more detail, clarified through concise interview statements. table 3 summary of results on separation of learning and assessment phase classroom management ➢ lesson organization is vital for students' srl ➢ the structure of the lessons needs to have a certain degree of flexibility ➢ the learning environment encourages students to independently organize their learning, necessitating rules and structures cognitive activation ➢ learning environment activates and stimulates students’ cognitive and metacognitive processes ➢ block schedules support srl processes ➢ separating of school and exam weeks enables better time management and focused exam preparation ➢ learning environment allows for in-depth and focused engagement with learning materials, reduces stress, and improves performance student support ➢ access to teachers is important for addressing questions and clarifications, especially during intensive learning periods ➢ the impact of student support is influenced by the amount of subject matter and the pace of lessons ➢ a balanced workload is necessary, as overwhelming demand or too little material can lead to boredom and reduced motivation ➢ reduced performance stress in this learning environment allows for more focused and in-depth learning, enhancing overall performance in the context of separating the learning and assessment phases, it becomes clear that the organization of lessons plays a crucial role in classroom management across all grades, which in turn supports students' srl. simultaneously, the two primary categories of classroom management and srl demonstrate significant interdependencies in this context. the learning environment encourages students to independently organize their learning processes as much as possible, necessitating appropriate rules and structures. the idea that students have to attend fewer classes is generally viewed positively, but concerns arise regarding the organization of lessons. for example, students' questions about upcoming exams may sometimes go unanswered immediately by their teachers. “i've noticed that, especially during these more intensive learning weeks, questions often arise. in such cases, it's quite practical to be able to approach a teacher for clarification. if you've mejeh, stampfli & hascher 66 | f l r already prepared a summary the week before, that's not a problem. however, during the exam week itself, you have to rely on organizing solutions among yourselves.” (student_13_2_3)1 this demonstrates that students were not only tasked with planning their learning processes over an extended period due to the altered lesson structure, but they also needed to motivate and regulate themselves. challenges arose when students were assigned an excessive number of diverse tasks. in the context of cognitive activation, the srl-conducive learning environment effectively stimulates students' cognitive and metacognitive processes. the separation of school and exam weeks, for instance, is mostly positively evaluated, as it enables students to manage their time more effectively, handle learning materials, and focus on exam preparations and studying. the introduction of block schedules provides structural support for srl processes, allowing more time for reviewing the learning materials before exams. from the students' perspective, the assessment of the impact on student support depended on the amount of subject matter covered. in particular, the pace of lessons played a crucial role: too much demand quickly led to overwhelm, while too little learning material caused boredom and reduced motivation. this situation was further complicated by variations in how teachers conducted their lessons. however, the pressure to perform was significantly alleviated in this teaching environment as students could engage with different learning materials in a more focused and in-depth manner. consequently, this reduction in stress appeared to enhance overall performance across various subjects. “yes, it's just more manageable in a way when you can concentrate on one week, and, of course, it leads to much less stress.” (student_13_2_6) this means that students were primarily empowered to make individual choices, as they could independently organize their learning and working time for the most part. this flexible structure strengthened their srl and allowed them to experience a greater sense of autonomy over extended periods. 7.2 fewer subjects per learning phase table 4 summarizes the results regarding fewer subjects per learning phase. the following sections present the results for the three main categories in more detail, clarified through concise interview statements. 1 students' statements are labeled by grade_interview_number of student. mejeh, stampfli & hascher 67 | f l r table 4 summary of results on fewer subjects per learning phase classroom management ➢ well-organized lessons with minimal interruptions aid srl and concentration ➢ teaching structure should adapt based on subject content, prerequisites, and student interests to enhance motivation cognitive activation ➢ reducing the number of subjects allows for longer durations per subject, promoting deeper learning ➢ the effectiveness of srl depends on the specific subject and students' familiarity with it (e.g., ease in german vs. challenges in mathematics) ➢ in-depth tasks (e.g., reading books) help manage interruptions and promote sustained learning student support ➢ reducing the number of subjects adjusts the lesson pace to support srl ➢ student perspectives on performance pressure varies: some students miss the competitiveness, while others find the reduced subjects alleviate pressure and are beneficial ➢ high-achieving classmates and study groups support srl by aiding comprehension and motivation through interactive learning and collaboration given the significance of reducing the number of subjects per learning phase as a key element in creating a lesson structure conducive to srl, alongside the continued importance of teacher instruction, it can be concluded that well-organized lessons support srl in classroom management. minimizing interruptions during lessons is highly valued – fewer interruptions and subject changes promote greater concentration and offer more opportunities to delve deeper into subject content. teaching longer blocks of the same subject allows students to explore individual topics in greater depth, allowing them to build on existing knowledge. however, the reduction of subjects per learning phase should depend on subject-specific content, recognizing that students have varying prerequisites and interests. a more flexible teaching structure has the potential to enhance students' learning motivation. in terms of cognitive activation, as previously mentioned, students were able to engage with the learning content more intensively due to the reduction in subjects, allowing for a longer learning duration per subject. however, the success of students' srl was largely dependent on the subject matter itself. the effectiveness of srl varied based on both the specific subject and the students' familiarity with it. specifically, students' statements regarding different subjects revealed the following: in german, they found it easier to connect with their prior knowledge or reactivate it. this ease was attributed to german being their native language. however, perceptions of lesson structure varied when it came to mathematics. learners highlighted that the possibilities for srl were influenced by their performance in the subject. “it might vary on an individual basis. for instance, i was never strong in math, so i can't pinpoint whether it was due to the interruptions or simply the subject itself, but, in any case, it made my learning a bit more challenging.” (student_13_2_2) mejeh, stampfli & hascher 68 | f l r in mathematics, establishing connections between the different learning phases posed a significant challenge for students. short repetition periods were evidently insufficient to reactivate previously acquired knowledge. this issue was particularly noticeable in the upper grades for foreign languages, where the condensed nature of lessons seemed to reduce their relevance for students. the learning environment designed to promote srl encouraged more in-depth learning, with less emphasis on memorization, by assigning demanding tasks such as reading and engaging with books. this, in turn, appeared to help learners better manage interruptions better during the individual learning phases. it is evident that, while promoting srl is the intended goal of the developed learning environment, students must already possess the ability to apply self-regulatory strategies to effectively learn in this instructional setting. regarding student support, students noted that the pace of the lessons was adjusted through the reduction of subjects allowing srl to take place. however, an ambivalent picture emerged among the learners regarding the pressure to perform. “in all honesty, i somewhat miss the competitiveness and performance pressure due to having only four subjects.” (student_9_1_8) this quote clearly illustrates that some students associated the reduction in subjects with reduced performance pressure, which lead to a sense of underchallenge. conversely, other students viewed the reduction of subjects as a relief, as it provided a moderately positive level of performance pressure. “i believe it's somewhat related to how each person handles performance pressure. when you receive everything at the beginning of a phase, like two german books, two english books, all the math and biology materials, and the learning objectives, you start thinking, “okay...” — you feel like you should be able to handle at least three-quarters of it. so, it does stress me out a bit at times, but in a positive way” (student_9_1_6) collaboration with high-achieving classmates was perceived as beneficial for srl, as it helped learners grasp connections within specific subjects. studying together had a motivating impact. through interactive discussions and mutual explanations, learning material was more thoroughly internalized and interconnected. 7.3 individual learning time table 5 summarizes the results regarding individual learning time (il). the following sections present the results for the three main categories in more detail, clarified through concise interview statements. mejeh, stampfli & hascher 69 | f l r table 5 summary of results on individual learning time classroom management ➢ clarity of rules and routines is crucial for individual learning time in classroom management, but differing interpretations by teachers can hinder students' srl ➢ teachers should provide clear, well-defined assignments for il to ensure structured learning ➢ allowing students to choose between group and individual work based on their preferences enhances effectiveness of learning ➢ different classes show distinctions in the structure of individual learning time. younger students often find it inflexible and lacking support, while older students value the autonomy it offers cognitive activation ➢ the lesson structure often fails to explore students' thought processes, leading to unclear solution paths in exams ➢ students generally find tasks during individual learning time to be balanced. older students, in particular, are engaged in reflective processes ➢ il fosters co-constructive learning and a strong sense of social connectedness through group work, though distractions can sometimes make it stressful ➢ il supports metacognitive processes, especially when coach roles are clear. effective il requires students to have srl skills student support ➢ students engage in group learning through reciprocal interactions, gradually learning to actively utilize this opportunity. especially older students value collaborative learning and co-regulation ➢ subject-related support is seen as essential for individual learning, with structured interactions with teachers helpful for content-related questions, while process-oriented support from coaches is welcomed for extracurricular issues ➢ unclear assignments and restricted choices can be demotivating for students in the learning environment, emphasizing the need for clear guidelines and teacher support ➢ feedback on individual learning time during lessons varies in effectiveness in the realm of classroom management, the clarity of rules and routines is a pivotal factor for il, during which students are expected to engage in independent learning during class hours. varying interpretations of these rules by different teachers were seen as highly detrimental, as students could only benefit from il sessions for their srl if they were appropriately structured by the teachers. “i would also question how specific these implementation rules are because teachers approach it so differently. some teachers assign new tasks every week, expecting them to be completed by a particular lesson and then discussed in class. on the other hand, there are teachers who provide assignments for the entire phase, requiring students to work on them alongside regular lessons.” (student_11_2_5) mejeh, stampfli & hascher 70 | f l r teachers interpreted the timeframes and content of il differently, prompting students to continually adjust their learning approaches based to each teacher's specifications. through the lens of il, notable distinctions between the three classes became evident. the gradual reduction in the structuring of il was accompanied by a more pronounced organization of learning processes by the students. this instructional approach required clear initial guidance in the form of well-defined assignments. il received mixed feedback from younger students, with the primary criticism revolving around its perceived inflexible structure, which did not align with their preferred learning context. they also expressed concern that, without teacher oversight and monitoring, they felt insufficiently supported, leading to reduced motivation to complete their assignments. “i believe you need a lot of self-discipline to ensure you actually accomplish tasks during il.” (student_9_1_5) conversely, older learners valued the increased autonomy they had during il, which provided them new opportunities to structure their learning process and prepare more flexibly for graduation. however, an ambivalent evaluation by the students remains evident, though the criticism takes on a different tone compared to younger learners. effective time management appears to occur only when upper secondary students are not overly controlled and are given ample freedom to independently select their learning time and environment. this is further emphasized by students' growing recognition of the importance of il over time and their increasing ability to use it more efficiently. “well, i also believe that everyone has different learning paces. some need more time than six lessons per week, while others require less. you can't compel someone to study for six lessons or any specific duration. additionally, when it comes to group assignments, like the geography one you mentioned earlier, students might voluntarily gather at school. it's convenient to meet after school since everyone is already here. i don't think it necessarily requires an official, mandatory lesson to facilitate such discussions. students can organize it themselves. moreover, if you have something to study and need to ask a teacher a question, you might willingly stay at school.” (student_13_1_1) regarding cognitive activation, the analyses indicate a lesson structure that does not fully explore students' thought processes. this is reflected in the learners' perceptions of examinations for example, where they often felt the solution path was unclear, particularly when teachers failed to recognize their struggle to grasp the lesson content. however, overall, the tasks assigned during il were deemed balanced by the students. for simpler tasks, students could use their remaining time to complete additional assignments or seek assistance from classmates, especially in younger grades. the lesson structure encouraged older students to engage in reflective processes about the level of task difficulty, especially in preparation for graduation exams. furthermore, co-constructive learning among upper secondary students was fostered, particularly when il was used for collaborative group work. learning in groups provided students with opportunities to develop srl by learning from one another. “well, not getting help directly, but you observe what other classmates are doing, sit down with them, compare notes, discuss tasks together, and solve them collectively, which in itself is a form of support.” (student_9_2_3) many students, regardless of their grade level, reported a strong sense of social connectedness through group learning. however, challenges arose in specific situations, particularly when distractions diverted their attention from the intended tasks and objectives. in such cases, group work was more likely to be perceived as time-consuming and stressful. il was viewed as supportive of metacognitive processes by the learners, particularly when they engaged reflectively with the learning content or when the roles of the coaches were clarified and deemed helpful. for il to be effective, it was crucial for students to possess a repertoire of self-regulation strategies and have knowledge about srl. mejeh, stampfli & hascher 71 | f l r in terms of student support, students had the opportunity to engage in group learning through reciprocal interactions, gradually learning how to actively utilize this approach. however, when il was perceived as coercive, it occasionally conflicted with other student needs. “in the beginning, it was more common for us to sit together in groups. we had a designated area for that, like an island where we would learn together. however, over time, it evolved to the point where if you wanted to study together, you could, but you could also study effectively on your own. additionally, we were permitted to use other rooms, which i found beneficial if you preferred studying alone.” (student_13_2_5) primarily, students in higher grades mentioned that, with the freedom to work independently in groups, they organized themselves accordingly and had a positive experience with collaborative learning. this led to learning alongside each other and active co-regulation over time. consequently, students had the opportunity to shape their learning processes individually. this suggests that learners were given the option to self-organize for group work when necessary. furthermore, students sought cooperation with classmates, especially when they needed help in a particular subject, highlighting the importance of group work within a subject context. from the students' perspective, structured interactions between teachers and students were particularly helpful when addressing content-related questions. subject-related support was considered more essential for individual learning, whereas process-oriented support (such as extracurricular issues, talent area organization, or school and leisure time) provided by assigned coaches was highly welcomed. “it might be better to have a subject teacher available for each subject once or twice a week, allowing students to work independently, and if they have questions, they can simply ask the teacher in person.” (student_9_2_1) students' dependence on a differentiated lesson structure and adaptive teacher support became evident in instances where unclear assignments resulted in overwhelming demands. similarly, the srl environment had a demotivating impact when students' choices were restricted. this limitation affected both assignment selection and completion, as well as participation in il, especially during marginal or intermediate hours. students expressed mixed views regarding the connection between il and lessons. in some instances, lessons provided opportunities for constructive feedback on tasks students had completed during il. however, in cases where feedback was absent or insufficient, the content of the various teaching settings appeared related in terms of subject matter but lacked alignment in their didactical approaches. “i believe it's important for self-assessment to integrate individual learning time (il) into the lessons to some extent. not necessarily discussing it directly, but laying a foundation in il, for instance, and then revisiting it after some time has passed. this way, you can build upon it in class and have a discussion to ensure that you've learned the material correctly and that your focus during il was on the right topics.” (student_9_2_5) 8. discussion despite srl being a central focus of classroom research for the past four decades (panadero, 2017) and the significant emphasis placed on the learning context, a lack of systematic alignment between research on srl and teaching quality exists, both from theoretical and empirical perspectives (dignath & veenman, 2020; dignath-van ewijk et al., 2013). questions regarding the specific qualities of srl-conducive learning environments remain unanswered (muijs et al., 2014; dignath & veenman, 2020). in the presented study, we investigated a learning environment conducive to srl at the upper secondary level. the objective of this instructional setting was to empower students to engage in more srl by implementing a distinct separation between learning phases and assessment phases, reducing the number of subjects per learning phase, and allowing for individualized learning time (perry et al., 2004, 2018, 2020). our study is unique to date as it represents the first systematic mejeh, stampfli & hascher 72 | f l r attempt to analyze an srl-conducive learning environment through the lens of the three basic dimensions of teaching quality. our analysis revealed that these three dimensions–classroom management, cognitive activation and student support– are also relevant for a learning environment intentionally designed to foster srl, even if they require partial reinterpretation. 8.1 the generic nature of srl-promoting learning environments the results suggest srl-supportive learning environments can largely be assessed based on the three basic dimensions. however, further specification of the generic nature of these dimensions for srl-supportive instruction is necessary. our analyses showed that almost all categories of perry's framework model could be reconstructed, indicating that the three basic dimensions could be effectively mapped within the srl-promoting learning environment of our study. traditionally, the focus of classroom management has been on identifying and reinforcing desirable behavior and preventing undesirable behavior through clear rules and the establishment of routines (hattie, 2009; hochweber et al., 2014; praetorius et al., 2018; seidel & shavelson, 2007). notably, classroom management is primarily embedded within the supportive structures for srl, as familiar routines, participation structures, expectations, instructions, and tasks play a central role in all three elements of the learning environment. although classroom management is considered generic for teaching (charalambous & praetorius, 2020), our study reveals that some generic aspects must be specified in terms of content to meet the requirements of an effective srl-promoting learning environment. for example, the results indicate that the focus in srl-promoting learning environments shifts toward supporting learners in managing their learning processes (de corte et al., 2004; perry, 2013). in this context, the effective handling and prevention of interruptions remain relevant (praetorius et al., 2018), although the definition of a disruption in the classroom needs reinterpretation. activities perceived as disruptive in traditional learning environments (e.g., a spontaneous exchange with a peer) can be seen as productive learning processes in srl learning environments, requiring flexible time management by teachers to allow students to work and reflect independently (dignath & veenman, 2020; dignath-van ewijk et al., 2013). cognitive activation plays an important role in a learning environment conducive to srl and is distributed across all three macro-categories (perry et al., 2018). for students to be cognitively engaged in an srl-promoting environment, it is essential not only to provide challenging tasks (künsting et al., 2016; praetorius et al., 2014) and activate prior knowledge (decristan et al., 2015), but also to create opportunities for in-depth exploration of the subject matter. our results indicate that cognitive activation in srl-promoting learning environments must be expanded to allow students to engage with a subject over an extended period, enabling them to grasp the content more thoroughly (veenman, 2013, 2017). the school’s practice of the separation of learning and assessment phases also proves to be beneficial. it encourages learners to apply relevant srl skills during learning, as highlighted in other studies (bernacki, 2017; järvelä & bannert, 2021; mccardle & hadwin, 2015; moos & azevedo, 2008; winne, 2019). equally important is the support of students in their learning processes as a basic dimension of effective teaching. this includes critical elements such as lesson pacing (baumert et al., 2010; kunter et al., 2005; praetorius et al., 2012), subject content (charalambous & kyriakides, 2017; klette & blikstad-balas, 2018; praetorius et al., 2020), and feedback (lipowsky et al., 2009; praetorius et al., 2014). however, a closer look at our results reveals some differences. for example, students confirmed the specific role of co-regulated learning in srl-promoting learning environments (hadwin et al., 2018; järvenoja et al., 2018; mccaslin & vriesema, 2018; panadero et al., 2015; vriesema & mccaslin, 2020) when they reported that collaboration with high-performing classmates and structured interactions with teachers were crucial for their learning progress. this finding highlights the importance of understanding a class as a community in an srl-promoting learning environment (perry et al., 2020). learners were able to rely on their classmates as a social resource when faced with mejeh, stampfli & hascher 73 | f l r incomprehensible tasks or inadequate feedback, while also practicing help-seeking as an important srl strategy (karabenick & gonida, 2017). 8.2 teacher support and student behavior in srl-promoting learning environments another key finding of our study is the limitation of analyzing an srl-supportive learning environment through the lens of the three basic dimensions, particularly when examining students' specific learning behaviors. these limitations become evident when relationships between srl and the basic dimensions are observed, as this approach identifies critical elements within the learning environment. for example, students often emphasize the organizational aspects of the learning environment over their individual learning experiences. this observation underscores a shortcoming in the three basic dimensions, as it does not explicitly account for motivational and emotional aspects of srl, lacking categories to address these factors. but, even indications of emotional and motivational aspects of students' srl were minimal and rarely emerged during the interviews. this finding aligns with earlier studies, which have consistently noted the subordinate role of motivational and emotional factors in the context of srl (azevedo & feyzi-behnagh, 2011; azevedo et al., 2017; gurtner et al., 2012; peetsma et al., 2017). while students appreciate the open spaces for learning, these open spaces also present substantial challenges. students either lack the appropriate self-regulatory strategies to effectively utilize the provided freedom, or the organization of the learning environment exhibits deficiencies. in the first scenario, it becomes evident that the promotion of srl can only be sustainably effective when direct and indirect forms of support complement each other (dignath & veenman, 2020; schuster et al., 2023; vosniadou et al., 2024). in the second scenario, it underscores the critical role teachers play in instructional design (heirweg et al., 2021; perels et al., 2009; perry, 2013; veenman, 2017). the learning environment is perceived as problematic when it is inflexible and unsuitable for students, leading to high levels of external regulation, which corresponds to the concept of the autonomy antinomy (helsper, 2002): students desire more creative freedom to engage in srl. initially, the simultaneous demand for more teacher support may seem contradictory, but upon closer examination, it highlights the contextand subject-specific nature of srl and its dependence on external regulation by teachers and peers to facilitate learning. our study is pioneering in drawing clear parallels between the discussion on the quality of teaching and the promotion of srl. similar to the debate on the distinction between generic and subject-specific teaching quality (e.g., praetorius et al., 2018; charalambous & praetorius, 2020), research on the promotion of srl also distinguishes between srl as a stable personality (trait) and as context-dependent, task-related behavior (state) of learners (e.g., matthews et al., 2000; winne & perry, 2000). 8.3 distinctions in srl between different grade levels differences between grade levels become apparent, particularly concerning individual learning time. for instance, students in grade 13 demonstrate a greater ability to adapt their srl to the lesson structure over time, resulting in more effective learning. in contrast, 9th-grade students express a desire for greater autonomy in shaping their srl. comparing these findings reveals that the gradual reduction of external regulation by teachers is positively perceived by students, as it allows them to take on greater responsibility for their own learning (corno, 2008; karlen et al., 2023). with an increasing number of choices available to individuals, new srl opportunities emerge, and students assume greater responsibility for their learning. consequently, the learning environment initially causes learners to experience a certain level of overload, but over time, they learn to navigate the open spaces and the altered lesson structure effectively. this development is characterized by a shift toward more effective learning in terms of planning and organization, and these changes are observable within a learning environment explicitly designed to support srl. mejeh, stampfli & hascher 74 | f l r 9. implications & future steps the findings from this study can have significant implications for both theory and practice in the realm of srl and teaching quality. 9.1 theoretical implications our findings underscore the necessity of refining srl theory to integrate the specific nuances of the three basic dimensions in srl-promoting environments. this contributes to the theoretical landscape by expanding existing srl models to include these critical factors, addressing gaps identified by previous research (dignath & veenman, 2020; dignath-van ewijk et al., 2013). our study contributes to the ongoing discourse on teaching quality by highlighting the distinction between generic and srl-specific teaching practices (praetorius et al., 2018). our research underscores that promoting srl requires tailored instructional strategies, thereby advancing the debate on effective teaching. our study also affirms the essential role of co-regulated learning, emphasizing the importance of social interactions and structured peer collaboration for srl (hadwin et al., 2018; järvenoja et al., 2018). this finding enhances srl theory by integrating social dimensions more explicitly, aligning with perry et al.'s (2020) concept of the classroom as a community. additionally, our results call for a theoretical refinement to include motivational and emotional factors when constructing an srl-supportive learning environment (azevedo et al., 2017). this holistic approach is crucial, as it captures the full spectrum of influences on learners' srl. consequently, future research should give more attention to motivational and emotional factors and their role in the design of srl-promoting learning environments. 9.2 practical implications practically, our study highlights the urgent need for enhanced teacher training and professional development to foster srl (kramarski, 2018; kramarski et al., 2013). teachers should be equipped to design and manage srl-promoting environments effectively, balancing autonomy with support, and redefining classroom disruptions to recognize productive learning processes. in this regard, curriculum design should incorporate structures that facilitate distinct learning and assessment phases, promoting deeper engagement with content and the more effective application of srl skills (perry et al., 2018). simultaneously, our study also contributes to a nuanced understanding of the promotion of srl, which has practical consequences. for the sustainable promotion of srl, srl learning environments should include the direct teaching of srl strategies by teachers (karlen et al., 2020; vosniadou et al., 2024), as students can quickly become overwhelmed if they lack srl skills. regarding classroom management, our findings indicate that establishing clear routines and expectations while maintaining flexibility is crucial for creating an srl-supportive learning environment (hochweber et al., 2014; praetorius et al., 2018). moreover, revising assessment methods to align with srl principles is another key implication. our results demonstrate that separating learning and assessment phases combined with continuous, formative feedback are essential steps to support srl development (bernacki, 2017; järvelä & bannert, 2021). finally, acknowledging and addressing the diverse srl needs of students at different grade levels is essential. teachers should progressively reduce external regulation and grant more autonomy as students become increasingly capable of managing their own learning. mejeh, stampfli & hascher 75 | f l r 9.3 limitations in our study, all three basic dimensions were successfully identified, although not all of their subcategories were specifically categorized. however, our study has certain limitations. the first limitation concerns our focus on students as the primary unit of study. while it was a deliberate choice to analyze the students' perspective to address the question at hand (scherer et al., 2016; wisniewski et al., 2020), it is important to acknowledge that our assessment of the didactic setting's design is presented exclusively from this perspective. this highlights a key insight in the context of systematically linking srl and teaching quality: while discussions on teaching quality often originate from the teacher's standpoint, research on srl predominantly adopts the students' viewpoint. a second limitation arises from our somewhat narrow definition of successful teaching in this study. berliner (2005) combines two approaches to “good teaching” (normative standards and professional action) and “effective teaching” (achievement of learning goals in terms of students' skills, abilities, and knowledge) into a comprehensive perspective of “quality teaching”. in this context, we have assumed three basic dimensions of quality teaching, whereas elsewhere, five basic dimensions are defined, with motivation, for instance, explicitly recognized as one of them (wisniewski & zierer, 2020). a more expansive perspective, such as the main-teach model (charalambous & praetorius, 2020), would be valuable for future analyses. 10. conclusion in summary, it is evident from students' responses how crucial it is for teachers to skillfully design and guide their students. in the case under examination, the weaknesses of this learning environment, as perceived by students, became apparent whenever there was a disconnect between its interpretation and implementation by teachers. from this perspective, the question of implementation quality and, consequently, the quality of the school itself becomes relevant when assessing teaching quality (muijs et al., 2014). furthermore, this study emphasizes the importance of considering srl from two interconnected perspectives: how students learn and how teachers teach, including how they learn to teach srl. this dual viewpoint underscores the relationship between teaching practices and the school context, highlighting the need to dismantle existing structures and the significance of an ongoing commitment to school development. a deeper understanding of these dynamics can cultivate environments where both students and teachers continuously grow and improve. keypoints the study aimed to assess the quality of an srl-promoting learning environment based on the three basic dimensions. the quality of srl-enhancing learning environments can be assessed and analyzed through the three basic dimensions, albeit not comprehensively. differences in srl adaptability were noted among grade levels, with older students demonstrating higher adaptability. teachers play a pivotal role in guiding and supporting students' srl. mejeh, stampfli & hascher 76 | f l r acknowledgments we sincerely thank gymnasium hofwil for the productive collaboration. special thanks go to the principal, peter stalder, whose support and commitment have been greatly appreciated. references azevedo, r., & feyzi-behnagh, r. (2011). dysregulated learning with advanced learning technologies. invited papers, 7(2). azevedo, r., taub, m., & mudrick, n. v. (2017). understanding and reasoning about real-time cognitive, affective, and metacognitive processes to foster self-regulation with advanced learning technologies. in d. h. schunk & j. a. greene (eds.), handbook of self-regulation of learning and performance (2nd ed., pp. 254–270). routledge. https://doi.org/10.4324/9781315697048-17 bandura, a. (1986). social foundations of thought and action: a social cognitive theory. prentice-hall. baumert, j., kunter, m., blum, w., brunner, m., voss, t., jordan, a., & tsai, y.-m. (2010). teachers’ mathematical knowledge, cognitive activation in the classroom, and student progress. american educational research journal, 47, 133–180. https://doi.org/10.3102/0002831209345157 berliner, d. c. (2005). the near impossibility of testing for teacher quality. journal of teacher education, 56(3), 205–213. https://doi.org/10.1177/0022487105275904 bernacki, m. l. (2017). examining the cyclical, loosely sequenced, and contingent features of selfregulated learning: trace data and their analysis. in d. h. schunk & j. a. greene (eds.), handbook of self-regulation of learning and performance (pp. 370–387). routledge. https://doi.org/10.4324/9781315697048-24 boekaerts, m. (2011). emotions, emotion regulation, and self-regulation of learning. in b. j. zimmerman & d. h. schunk (eds.), handbook of self-regulation of learning and performance (pp. 408–425). routledge. https://doi.org/10.4324/9780203839010-34 brophy, j. (1999). teaching (educational practices series, vol. 1). international bureau of education. brophy, j. (2010). teacher effects research and teacher quality. the journal of classroom interaction, 45(1), 32–40. brophy, j., & good, t. l. (1986). teacher behavior and student achievement. in m.c. wittrock (ed.), handbook of research on teaching (pp. 328–375). macmillan. brown, a. l., campione, j. c., & day, j. d. (1981). learning to learn: on training students to learn from texts. educational researcher, 10(2), 14–21. https://doi.org/10.3102/0013189x010002014 charalambous, c. y., & kyriakides, e. (2017). working at the nexus of generic and content-specific teaching practices: an exploratory study based on timss secondary analyses. the elementary school journal, 117(3), 423–454. https://doi.org/10.1086/690221 charalambous, c. y., & praetorius, a.-k. (2020). creating a forum for researching teaching and its quality more synergistically. studies in educational evaluation, 67, 100894. https://doi.org/10.1016/j.stueduc.2020.100894 cheon, s. h., & reeve, j. (2015). a classroom-based intervention to help teachers decrease students’ amotivation. contemporary educational psychology, 40, 99–111. https://doi.org/10.1016/j.cedpsych.2014.06.004 https://doi.org/10.4324/9781315697048-17 https://doi.org/10.3102/0002831209345157 https://doi.org/10.1177/0022487105275904 https://doi.org/10.4324/9781315697048-24 https://doi.org/10.4324/9780203839010-34 https://doi.org/10.3102/0013189x010002014 https://doi.org/10.1086/690221 https://doi.org/10.1016/j.stueduc.2020.100894 https://doi.org/10.1016/j.cedpsych.2014.06.004 mejeh, stampfli & hascher 77 | f l r corno, l. (2008). on teaching adaptively. educational psychologist, 43(3), 161–173. https://doi.org/10.1080/00461520802178466 de corte, e. (2012). constructive, self-regulated, situated, and collaborative learning: an approach for the acquisition of adaptive competence. journal of education, 192(2-3), 33-47. https://doi.org/10.1177/0022057412192002-307 de corte, e. (2016). improving higher education students’ learning proficiency by fostering their selfregulation skills. european review, 24(2), 264-276. https://doi.org/ 10.1017/s1062798715000617 de corte, e., verschaffel, l., & masui, c. (2004). the clia-model: a framework for designing powerful learning environments for thinking and problem solving. european journal of psychology of education, 19(4), 365–384. https://doi.org/10.1007/bf03173216 de naeghel, j., van keer, h., vansteenkiste, m., haerens, l., & aelterman, n. (2016). promoting elementary school students’ autonomous reading motivation: effects of a teacher professional development workshop. the journal of educational research, 109(3), 232–252. https://doi.org/10.1080/00220671.2014.942032 decristan, j., klieme, e., kunter, m., hochweber, j., büttner, g., fauth, b., ... & hardy, i. (2015). embedded formative assessment and classroom process quality: how do they interact in promoting science understanding? american educational research journal, 52(6), 11331159. https://doi.org/ 10.3102/0002831215596412 dignath, c., & büttner, g. (2008). components of fostering self-regulated learning among students. a meta-analysis on intervention studies at primary and secondary school level. metacognition and learning, 3(3), 231–264. https://doi.org/10.1007/s11409-008-9029-x dignath, c., & veenman, m. v. (2020). the role of direct strategy instruction and indirect activation of self-regulated learning evidence from classroom observation studies. educational psychology review, 33(2), 489–533. https://doi.org/10.1007/s10648-020-09534-0 dignath-van ewijk, c., dickhäuser, o., & büttner, g. (2013). assessing how teachers enhance selfregulated learning: a multiperspective approach. journal of cognitive education and psychology, 12(3), 338–358. https://doi.org/10.1891/1945-8959.12.3.338 donker, a. s., de boer, h., kostons, d., van ewijk, c. d., & van der werf, m. p. (2014). effectiveness of learning strategy instruction on academic performance: a meta-analysis. educational research review, 11, 1–26. https://doi.org/10.1016/j.edurev.2013.11.002 efklides, a. (2011). interactions of metacognition with motivation and affect in self-regulated learning: the masrl model. educational psychologist, 46(1), 6–25. https://doi.org/10.1080/00461520.2011.538645 ferguson, r. f., & danielson, c. (2014). how framework for teaching and tripod 7cs evidence distinguish key components of effective teaching. in t. j. kane, k. a. kerr, & r. c. pianta (eds.), designing teacher evaluation systems (1st ed., pp. 98–143). wiley. https://doi.org/10.1002/9781119210856.ch4 finsterwald, m., wagner, p., schober, b., lüftenegger, m., & spiel, c. (2013). fostering lifelong learning. evaluation of a teacher education program for professional teachers. teaching and teacher education, 29, 144–155. https://doi.org/10.1016/j.tate.2012.08.009 flavell, j. h. (1971). first discussant's comments: what is memory development the development of? human development, 14(4), 272–278. https://doi.org/10.1080/00461520802178466 https://doi.org/10.1177/0022057412192002-307 https://doi.org/10.1007/bf03173216 https://doi.org/10.1080/00220671.2014.942032 https://doi.org/10.1007/s11409-008-9029-x https://doi.org/10.1007/s10648-020-09534-0 https://doi.org/10.1891/1945-8959.12.3.338 https://doi.org/10.1016/j.edurev.2013.11.002 https://doi.org/10.1080/00461520.2011.538645 https://doi.org/10.1002/9781119210856.ch4 mejeh, stampfli & hascher 78 | f l r greene, j. a., plumley, r. d., urban, c. j., bernacki, m. l., gates, k. m., hogan, k. a., demetriou, c., & panter, a. t. (2021). modeling temporal self-regulatory processing in a higher education biology course. learning and instruction, 72, 101201. https://doi.org/10.1016/j.learninstruc.2019.04.002 gurtner, j.-l., gulfi, a., genoud, p. a., de rocha trindade, b., & schumacher, j. (2012). learning in multiple contexts: are there intra-, crossand transcontextual effects on the learner’s motivation and help seeking? european journal of psychology of education, 27(2), 213–225. https://doi.org/10.1007/s10212-011-0083-4 hadwin, a., järvelä, s., & miller, m. (2018). self-regulation, co-regulation, and shared regulation in collaborative learning environments. in d. h. schunk & j. a. greene (eds.), handbook of self-regulation of learning and performance (2nd ed., pp. 83–106). routledge/taylor & francis group. https://doi.org/10.4324/9781315697048-6 hattie, j. (2009). visible teaching – visible learning: a synthesis of 800+ meta-analyses on achievement. routledge. hattie, j., biggs, j., & purdie, n. (1996). effects of learning skills interventions on student learning: a meta-analysis. review of educational research, 66(2), 99–136. heirweg, s., de smul, m., merchie, e., devos, g., & van keer, h. (2021). do you reap what you sow? the relationship between primary school students’ self-regulated learning and student, teacher, and school determinants. school effectiveness and school improvement, 32(1), 118– 140. https://doi.org/10.1080/09243453.2020.1797829 helsper, w. (2002). lehrerprofessionalität als antinomische handlungsstruktur [teacher professionalism as an antinomian structure of action]. in m. kraul, w. marotzki, & c. schweppe (hrsg.), biographie und profession (s. 64–102). julius klinkhardt. hochweber, j., hosenfeld, i., & klieme, e. (2014). classroom composition, classroom management, and the relationship between student attributes and grades. journal of educational psychology, 106(1), 289–300. https://doi.org/10.1037/a0033829 jansen, r. s., van leeuwen, a., janssen, j., jak, s., & kester, l. (2019). self-regulated learning partially mediates the effect of self-regulated learning interventions on achievement in higher education: a meta-analysis. educational research review, 28, 100292. https://doi.org/10.1016/j.edurev.2019.100292 järvelä, s., & bannert, m. (2021). temporal and adaptive processes of regulated learning—what can multimodaldata tell? learning and instruction, 72, 101268. https://doi.org/10.1016/j.learninstruc.2019.101268. järvenoja, h., järvelä, s., törmänen, t., näykki, p., malmberg, j., kurki, k., mykkänen, a., & isohätälä, j. (2018). capturing motivation and emotion regulation during a learning process. frontline learning research, 6(3), 85–104. https://doi.org/10.14786/flr.v6i3.369310 karabenick, s. a., & gonida, e. n. (2017). academic help seeking as a self-regulated learning strategy: current issues, future directions. in d. h. schunk & j. a. greene (eds.), handbook of self-regulation of learning and performance (pp. 421-433). routledge. https://doi.org/10.4324/9781315697048-27 https://doi.org/10.1016/j.learninstruc.2019.04.002 https://doi.org/10.1007/s10212-011-0083-4 https://doi.org/10.4324/9781315697048-6 https://doi.org/10.1080/09243453.2020.1797829 https://doi.org/10.1037/a0033829 https://doi.org/10.1016/j.edurev.2019.100292 https://doi.org/10.1016/j.learninstruc.2019.101268 https://doi.org/10.14786/flr.v6i3.369310 https://doi.org/10.4324/9781315697048-27 mejeh, stampfli & hascher 79 | f l r karlen, y., & hertel, s. (2024). inspiring self-regulated learning in everyday classrooms: teachers’ professional competences and promotion of self-regulated learning. unterrichtswissenschaft, 52(1), 1–13. https://doi.org/10.1007/s42010-024-00196-3 karlen, y., hertel, s., & hirt, c. n. (2020). teachers’ professional competences in self-regulated learning: an approach to integrate teachers’ competences as self-regulated learners and as agents of self-regulated learning in a holistic manner. frontiers in education, 5, article 159. https://doi.org/10.3389/feduc.2020.00159 karlen, y., hirt, c. n., jud, j., rosenthal, a., & eberli, t. d. (2023). teachers as learners and agents of self-regulated learning: the importance of different teachers competence aspects for promoting metacognition. teaching and teacher education, 125, 104055. https://doi.org/10.1016/j.tate.2023.104055 kistner, s., rakoczy, k., otto, b., dignath-van ewijk, c., büttner, g., & klieme, e. (2010). promotion of self-regulated learning in the classrooms: investigating frequency, quality, and consequences for student performance. metacognition and learning, 5(2), 157–171. https://doi.org/10.1007/s11409-010-9055-3 kitsantas, a., winsler, a., & huie, f. (2008). self-regulation and ability predictors of academic success during college: a predictive validity study. journal of advanced academics, 20(1), 4268. https://doi.org/10.4219/jaa-2008-867 klette, k., & blikstad-balas, m. (2018). observation manuals as lenses to classroom teaching: pitfalls and possibilities. european educational research journal, 17(1), 129–146. https://doi.org/10.1177/1474904117703228 klieme, e. (2019). unterrichtsqualität [teaching quality]. in m. harring, c. rohlfs & m. gläserzikuda (hrsg.), handbuch schulpädagogik (s. 393–408). waxmann. klieme, e., & rakoczy, k. (2003). unterrichtsqualität aus schülerperspektive: kulturspezifische profile, regionale unterschiede und zusammenhänge mit effekten von unterricht [teaching quality from a student perspective: culture-specific profiles, regional differences, and correlations with effects of teaching]. in pisa 2000—ein differenzierter blick auf die länder der bundesrepublik deutschland (s. 333–359). wiesbaden. klieme, e., pauli, c. & reusser, k. (2009). the pythagoras study. investigating effects of teaching and learning in swiss and german mathematics classrooms. in: t. janik (ed.), the power of video studies in investigating teaching and learning in the classroom. (pp. 137-160). waxmann. kramarski, b. (2018). teachers as agents in promoting students' srl and performance: applications for teachers’ dual-role training program. in d. h. schunk & j. a. greene (eds.), handbook of self-regulation of learning and performance 2nd ed. (pp. 223-239). routledge. https://doi.org/10.4324/9781315697048-15 kramarski, b., desoete, a., bannert, m., narciss, s., & perry, n. (2013). new perspectives on integrating self-regulated learning at school. education research international, 2013, 498214. https://doi.org/10.1155/2013/498214 kuckartz, u., & rädiker, s. (2019). analyzing qualitative data with maxqda: text, audio, and video. springer. https://doi.org/10.1007/978-3-030-15671-8 https://doi.org/10.1007/s42010-024-00196-3 https://doi.org/10.3389/feduc.2020.00159 https://doi.org/10.1016/j.tate.2023.104055 https://doi.org/10.1007/s11409-010-9055-3 https://doi.org/10.4219/jaa-2008-867 https://doi.org/10.1177/1474904117703228 https://doi.org/10.4324/9781315697048-15 https://doi.org/10.1155/2013/498214 https://doi.org/10.1007/978-3-030-15671-8 mejeh, stampfli & hascher 80 | f l r künsting, j., neuber, v., & lipowsky, f. (2016). teacher self-efficacy as a long-term predictor of instructional quality in the classroom. european journal psychology of education, 31, 299– 322. https ://doi.org/10.1007/s10212-015-0272-7 kunter, m., klusmann, u., baumert, s., richter, d., voss, t., & hachfeld, a. (2013). professional competence of teachers: effects on instructional quality and student development. journal of educational psychology, 105(3), 805–820. https://doi.org/10.1037/a0032583 lamnek, s. (2005). gruppendiskussion – theorie und praxis [focus group theory and practice]. beltz. leidinger, m., & perels, f. (2012). training self-regulated learning in the classroom: development and evaluation of learning materials to train self-regulated learning during regular mathematics lessons at primary school. education research international, 2012, 1–14. https://doi.org/10.1155/2012/735790 leutwyler, b., & maag merki, k. (2009). school effects on students’ self-regulated learning. a multivariate analysis of the relationship between individual perceptions of school processes and cognitive, metacognitive, and motivational dimensions of self-regulated learning. journal for educational research online, 1(1), 197–223. https://doi.org/10.5167/uzh-29038 lipowsky, f., rakoczy, k., pauli, c., drollinger-vetter, b., klieme, e., & reusser, k. (2009). quality of geometry instruction and its short-term impact on students’ understanding of the pythagorean theorem. learning and instruction, 19, 527–537. https://doi.org/10.1016/j.learninstruc.2008.11.001 martin, j. (2007). the selves of educational psychology: conceptions, contexts, and critical considerations. educational psychologist, 42, 79–89. https://doi.org/10.1080/00461520701263244 masui, c., & de corte, e. (2005). learning to reflect and to attribute constructively as basic components of self-regulated learning. british journal of educational psychology, 75(3), 351– 372. http://dx.doi.org/10.1348/000709905x25030 matthews, g., schwean, v. l., campbell, s. e., saklofske, d. h., & mohamed, a. a. r. (2000). personality, self-regulation, and adaptation: a cognitive-social framework. in m. boekaerts, p. r. pintrich, & m. zeidner (eds.), handbook of self-regulation (pp. 171–207). academic press. https://doi.org/10.1016/b978-012109890-2/50035-4 mayring, p. (2015). qualitative inhaltsanalyse [qualitative content analysis]. beltz. mccardle, l., & hadwin, a. f. (2015). using multiple, contextualized data sources to measure learners’ perceptions of their self-regulated learning. metacognition and learning, 10, 43-75. mccaslin, m. & good, t. l. (1996). the informal curriculum. in d. c. berliner & r. c. calfee (eds.), handbook of educational psychology (pp. 622–670). simon & schuster macmillan. mccaslin, m., & vriesema, c. c. (2018). co-regulation: a model for classroom research in a. vygotskian perspective. in g. a. d. liem & d. m. mcinerey (eds.), big theories revised 2: research on sociocultural influences on motivation and learning (pp. 319–352). information age publishing. mcdonald, n., schoenebeck, s., & forte, a. (2019). reliability and inter-rater reliability in qualitative research: norms and guidelines for cscw and hci practice. proceedings of the acm on human-computer interaction, 3(cscw), 1–23. https://doi.org/10.1145/3359174 michalsky, t. (2021). preservice and inservice teachers’ noticing of explicit instruction for selfregulated learning strategies. frontiers in psychology, 12, 630197. https://doi.org/10.3389/fpsyg.2021.630197 https://doi.org/10.1037/a0032583 https://doi.org/10.1155/2012/735790 https://doi.org/10.5167/uzh-29038 https://doi.org/10.1016/j.learninstruc.2008.11.001 https://doi.org/10.1080/00461520701263244 http://dx.doi.org/10.1348/000709905x25030 https://doi.org/10.1016/b978-012109890-2/50035-4 https://doi.org/10.1145/3359174 https://doi.org/10.3389/fpsyg.2021.630197 mejeh, stampfli & hascher 81 | f l r molenaar, i., sleegers, p., & van boxtel, c. (2014). metacognitive scaffolding during collaborative learning: a promising combination. metacognition & learning, 9, 309–332. https://doi.org/10.1007/s11409-014-9118-y moos, d. c., & azevedo, r. (2008). self-regulated learning with hypermedia: the role of prior domain knowledge. contemporary educational psychology, 33(2), 270–298. https://doi.org/10.1016/j.cedpsych.2007.03.001 muijs, d., kyriakides, l., van der werf, g., creemers, b., timperley, h., & earl, l. (2014). state of the art – teacher effectiveness and professional learning. school effectiveness and school improvement, 25(2), 231–256. https://doi.org/10.1080/09243453.2014.885451 nandagopal, k., & ericsson, k. a. (2012). an expert performance approach to the study of individual differences in self-regulated learning activities in upper-level college students. learning and individual differences, 22(5), 597-609. https://doi.org/10.1016/j.lindif.2011.11.018 oecd, (2018). the future of education and skills: education 2030, the future we want. https://www.oecd.org/education/2030/e2030%20position%20paper%20(05.04.2018). oecd, (2019). oecd skills outlook 2019: thriving in a digital world. https://doi.org/10.1787/df80bc12-en onwuegbuzie, a. j., dickinson, w. b., leech, n. l., & zoran, a. g. (2009). a qualitative framework for collecting and analyzing data in focus group research. international journal of qualitative methods, 8(3), 1–21. https://doi.org/10.1177/160940690900800301 oser, f. k., & baeriswyl, f. j. (2001). choreographies of teaching: bridging instruction to learning. handbook of research on teaching, 4, 1031-1065. panadero, e. (2017). a review of self-regulated learning: six models and four directions for research. frontiers in psychology, 8(422), 1–28. https://doi.org/10.3389/fpsyg.2017.00422 panadero, e., kirschner, p. a., järvelä, s., malmberg, j., & järvenoja, h. (2015). how individual selfregulation affects group regulation and performance: a shared regulation intervention. small group research, 46(4), 431–454. https://doi.org/10.1177/1046496415591219 paris, s. g., & paris, a. h. (2001). classroom applications of research on self-regulated learning. educational psychologist, 36(2), 89–101. https://doi.org/10.1207/s15326985ep3602_4 patton, m. q. (2015). qualitative research & evaluation methods: integrating theory and practice. sage. peetsma, t., van der veen, i., & schuitema, j. (2017). use of time: time perspective intervention of motivation enhancement. palgrave macmillan. https://doi.org/10.1057/978-1-137-60191-9_10 perels, f., dignath, c., & schmitz, b. (2009). is it possible to improve mathematical achievement by means of self-regulation strategies? evaluation of an intervention in regular math classes. european journal of psychology of education, 24(1), 17–31. https://doi.org/10.1007/bf03173472 perry, n. e. & rahim, a. (2011). studying self-regulated learning in classrooms. in zimmerman, b. & schunk, d. (eds.), handbook of self-regulation of learning and performance (pp. 122–136). routledge. perry, n. e. (2013). understanding classroom processes that support children’s self-regulation of learning. in n. e. perry (ed.), self-regulation and dialogue in primary classrooms (pp. 45– 68). the british psychological society. perry, n. e., & drummond, l. (2002). helping young students become self-regulated researchers and writers. the reading teacher, 56(3), 298–310. https://doi.org/10.1007/s11409-014-9118-y https://doi.org/10.1016/j.cedpsych.2007.03.001 https://doi.org/10.1080/09243453.2014.885451 https://doi.org/10.1016/j.lindif.2011.11.018 https://doi.org/10.1787/df80bc12-en https://doi.org/10.1177/160940690900800301 https://doi.org/10.3389/fpsyg.2017.00422 https://doi.org/10.1177/1046496415591219 https://doi.org/10.1207/s15326985ep3602_4 https://doi.org/10.1057/978-1-137-60191-9_10 https://doi.org/10.1007/bf03173472 mejeh, stampfli & hascher 82 | f l r perry, n. e., lisaingo, s., yee, n., parent, n., wan, x., & muis, k. (2020). collaborating with teachers to design and implement assessments for self-regulated learning in the context of authentic classroom writing tasks. assessment in education: principles, policy & practice, 27(4), 416–443. https://doi.org/10.1080/0969594x.2020.1801576 perry, n. e., mazabel, s., dantzer, b., & winne, p. (2018). supporting self-regulation and selfdetermination in the context of music education. in g. a. d. liem & d. m. mcinerney (eds.), big theories revisited 2: a volume of research on sociocultural influences on motivation and learning (pp. 295–318). information age press. perry, n.e., phillips, l., & dowler, j. (2004). examining features of tasks and their potential to promote self-regulated learning. teachers college record: the voice of scholarship in education, 106(9), 1854–1878. https://doi.org/10.1111/j.1467-9620.2004.00408.x peverly, s. t., brobst, k. e., graham, m., & shaw, r. (2003). college adults are not good at selfregulation: a study on the relationship of self-regulation, note taking, and test taking. journal of educational psychology, 95(2), 335–346. https://doi.org/10.1037/0022-0663.95.2.335 pianta, r. c., la paro, k., & hamre, b. k. (2008). classroom assessment scoring system (class). paul h. brookes. pintrich, p. r. (2004). a conceptual framework for assessing motivation and self-regulated learning in college students. educational psychology review, 16(4), 385–407. https://doi.org/10.1007/s10648-004-0006-x praetorius, a. k., herrmann, c., gerlach, e., zülsdorf-kersting, m., heinitz, b., & nehring, a. (2020). unterrichtsqualität in den fachdidaktiken im deutschsprachigen raum–zwischen generik und fachspezifik [teaching quality in subject didactics in german-speaking countries between generics and subject specifics]. unterrichtswissenschaft, 48(3), 409–446. https://doi.org/10.1007/s42010-020-00082-8 praetorius, a. k., klieme, e., herbert, b., & pinger, p. (2018). generic dimensions of teaching quality: the german framework of three basic dimensions. zdm, 50(3), 407–426. https://doi.org/10.1007/s11858-018-0918-4 praetorius, a. k., pauli, c., reusser, k., rakoczy, k., & klieme, e. (2014). one lesson is all you need? stability of instructional quality across lessons. learning and instruction, 31, 1–12. https://doi.org/10.1016/j.learninstruc.2013.12.002 praetorius, a., lenske, g., & helmke, a. (2012). observer ratings of instructional quality: do they fulfill what they promise? learning and instruction, 22, 387–400. https://doi.org/10.1016/j.learninstruc.2012.03.002 praetorius, a.-k., klieme, e., kleickmann, t., brunner, e., lindmeier, a., taut, s. & charalambous, c. y. (2020). towards developing a theory of generic teaching quality: origin, current status, and necessary next steps regarding the three basic dimensions model. zeitschrift für pädagogik, 66, 15–36. https://doi.org/10.25656/01:25861 salonen, p., vauras, m., & efklides, a. (2005). social interaction—what can it tell us about metacognition and coregulation in learning? european psychologist, 10(3), 199–208. https://doi.org/10.1027/1016-9040.10.3.199 scherer, r., nilsen, t., & jansen, m. (2016). evaluating individual students’ perceptions of instructional quality: an investigation of their factor structure, measurement invariance, and relations to educational outcomes. frontiers in psychology, 7. https://doi.org/10.3389/fpsyg.2016.00110 schreier, m. (2017). sampling and generalization. in u. flick (eds.), the sage handbook of qualitative methods for data collection (s. 84–98). sage. https://doi.org/10.1080/0969594x.2020.1801576 https://doi.org/10.1111/j.1467-9620.2004.00408.x https://doi.org/10.1037/0022-0663.95.2.335 https://doi.org/10.1007/s10648-004-0006-x https://doi.org/10.1007/s42010-020-00082-8 https://doi.org/10.1007/s11858-018-0918-4 https://doi.org/10.1016/j.learninstruc.2013.12.002 https://doi.org/10.1016/j.learninstruc.2012.03.002 https://doi.org/10.25656/01:25861 https://doi.org/10.1027/1016-9040.10.3.199 https://doi.org/10.3389/fpsyg.2016.00110 mejeh, stampfli & hascher 83 | f l r schuster, c., stebner, f., geukes, s., jansen, m., leutner, d., & wirth, j. (2023). the effects of direct and indirect training in metacognitive learning strategies on near and far transfer in selfregulated learning. learning and instruction, 83, 101708. https://doi.org/10.1016/j.learninstruc.2022.101708 seidel, t. & shavelson, r. (2007). teaching effectiveness research in the past decade. review of educational research, 77, 454–499. https://doi.org/10.3102/0034654307310317 slavin, r. e. (1987). a theory of school and classroom organization. educational psychologist, 22, 89-108. https://doi.org/10.1207/s15326985ep2202_1 slavin, r. e. (1994). quality, appropriateness, incentive, and time: a model of instructional effectiveness. international journal of educational research, 21(2), 141–157. https://doi.org/10.1016/0883-0355(94)90029-9 souvignier, e., & mokhlesgerami, j. (2006). using self-regulation as a framework for implementing strategy instruction to foster reading comprehension. learning and instruction, 16(1), 57–71. https://doi.org/10.1016/j.learninstruc.2005.12.006 su, y.-l., & reeve, j. (2011). a meta-analysis of the effectiveness of intervention programs designed to support autonomy. educational psychology review, 23(1), 159–188. https://doi.org/10.1007/s10648-010-9142-7 van leeuwen, a., & janssen, j. (2019). a systematic review of teacher guidance during collaborative learning in primary and secondary education. educational research review, 27, 71–89. https://doi.org/10.1016/j.edurev.2019.02.001 veenman, m. v. j. (2013). training metacognitive skills in students with availability and production deficiencies. in h. bembenutty, t. cleary & a. kitsantas (eds.), applications of self-regulated learning across diverse disciplines: a tribute to barry j. zimmerman (pp. 299–324). information age publishing. veenman, m. v. j. (2017). learning to self-monitor and self-regulate. in r. mayer & p. alexander (eds.), handbook of research on learning and instruction (2nd ed., pp. 233–257). routledge. venitz, l., & perels, f. (2019). the promotion of self-regulated learning by kindergarten teachers: differential effects of an indirect intervention. international electronic journal of elementary education, 11(5), 437-448. https://doi.org/10.26822/iejee.2019553340 vosniadou, s. (2020). bridging secondary and higher education. the importance of self-regulated learning. european review, 28(s1), 94–103. https://doi.org/10.1017/s1062798720000939 vosniadou, s., bodner, e., stephenson, h., jeffries, d., lawson, m. j., darmawan, ig. n., kang, s., graham, l., & dignath, c. (2024). the promotion of self-regulated learning in the classroom: a theoretical framework and an observation study. metacognition and learning, 19(1), 381– 419. https://doi.org/10.1007/s11409-024-09374-1 vriesema, c. c., & mccaslin, m. (2020). experience and meaning in small-group contexts. frontline learning research, 8(3), 126–139. https://doi.org/10.14786/flr.v8i3.493 vygotsky, l. s. (1962). thought and language. mit press. winne, p. h. (2019). paradigmatic dimensions of instrumentation and analytic methods in research on self-regulated learning. computers in human behavior, 96, 285-289. https://doi.org/10.1016/j.chb.2019.03.026 winne, p. h., & hadwin, a. f. (1998). studying as self-regulated learning. in d. j. hacker, j. dunlosky, & a. c. graesser (eds.), metacognition in educational theory and practice (pp. 277–304). lawrence erlbaum associates publishers. https://doi.org/10.1016/j.learninstruc.2022.101708 https://doi.org/10.3102/0034654307310317 https://doi.org/10.1207/s15326985ep2202_1 https://doi.org/10.1016/0883-0355(94)90029-9 https://doi.org/10.1016/j.learninstruc.2005.12.006 https://doi.org/10.1007/s10648-010-9142-7 https://doi.org/10.1016/j.edurev.2019.02.001 https://doi.org/10.26822/iejee.2019553340 https://doi.org/10.1017/s1062798720000939 https://doi.org/10.1007/s11409-024-09374-1 https://doi.org/10.14786/flr.v8i3.493 mejeh, stampfli & hascher 84 | f l r winne, p. h., & perry, n. e. (2000). measuring self-regulated learning. in in m. boekaerts, p. r. pintrich, & m. zeidner (eds.), handbook of self-regulation (pp. 531–566). elsevier. https://doi.org/10.1016/b978-012109890-2/50045-7 wisniewski, b., & zierer, k. (2020). entwicklung eines online-fragebogens zur erhebung von unterrichtsqualität durch lernendenfeedback und erste validierungsschritte [development of an online questionnaire to survey teaching quality through learner feedback and initial validation steps.]. psychologie in erziehung und unterricht, 67(2), 138–155. http://dx.doi.org/10.2378/peu2020.art10d wisniewski, b., zierer, k., dresel, m., & daumiller, m. (2020). obtaining secondary students’ perceptions of instructional quality: two-level structure and measurement invariance. learning and instruction, 66, 101303. https://doi.org/10.1016/j.learninstruc.2020.101303 zimmerman, b. j. (2000). attaining self-regulation: a social cognitive perspective. in m. boekaerts, p. r. pintrich, & m. zeidner (eds.), handbook of self-regulation (pp. 13–39). academic press. https://doi.org/10.1016/b978-012109890-2/50045-7 https://doi.org/10.1016/j.learninstruc.2020.101303 microsoft word maatta_publication.docx frontline learning research vol. 9 no. 4 (2021) 92 115 issn 2295-3159 corresponding author: olli määttä, department of education, university of helsinki, finland, email address: olli.maatta@helsinki.fi doi: https://doi.org/10.14786/flr.v9i4.965 students in sight: using mobile eye-tracking to investigate mathematics teachers’ gaze behaviour during task instructiongiving olli maatta1, nora mcintyre2, jussi palomäki1, markku s. hannula1, patrik scheinin1 & petri ihantola1 1department of education, university of helsinki, finland 2 university of southampton, united kingdom article received 16 march 2021 / article revised 5 october / accepted 21 october / available online 1 december abstract mobile eye-tracking research has provided evidence both on teachers' visual attention in relation to their intentions and on teachers’ student-centred gaze patterns. however, the importance of a teacher’s eye-movements when giving instructions is unexplored. in this study we used mobile eye-tracking to investigate six teachers’ gaze patterns when they are giving task instructions for a geometry problem in four different phases of a mathematical problem-solving lesson. we analysed the teachers’ eye-tracking data, their verbal data, and classroom video recordings. our paper brings forth a novel interpretative lens for teacher’s pedagogical intentions communicated by gaze during teacher-led moments such as when introducing new tasks, reorganizing the social structures of students for collaboration, and lesson wrap-ups. a change in the students’ task changes teachers’ gaze patterns, which may indicate a change in teacher’s pedagogical intention. we found that teachers gazed at students throughout the lesson, whereas teachers’ focus was at task-related targets during collaborative instruction-giving more than during the introductory and reflective task instructions. hence, we suggest two previously not detected gaze types: contextualizing gaze for task readiness and collaborative gaze for task focus to contribute to the present discussion on teacher gaze. keywords: mobile eye-tracking; eye-movement; mathematical problem-solving; teacher gaze; intention. maatta et al 93 | f l r 1. introduction every lesson offers a new stage. the teacher’s first steps on the scene set the tone for the rest of the play. the way a teacher behaves is strongly related to the teaching-learning environment, classroom dynamics and learning outcomes (hattie, 2003; van der want et al., 2015; van tartwijk et al., 1998; witt et al., 2004). moreover, a teacher's tacit pedagogical knowledge feeds into her responses to rapidly changing and unanticipated classroom moments. this knowledge forms the essence of the teaching practice and enables the teacher's professional actions in the classroom (toom, 2006). further, the classroom situation provides memory cues for the teacher about similar situations encountered earlier (simon, 1992). thus, teachers’ reactions result from dynamic, ever-changing classroom situations that demand a situation-specific expression of teacher tact, defined also as mindful action, teacherly awareness, and “the practical language of the body it is the language of acting in pedagogical moments” (van manen, 1991). teachers’ tactful behaviors are realized in teacher presence which is achieved via verbal and nonverbal teacher immediacy (mehrabian, 1970). immediacy behaviors as eye contact induce experiences of closeness and warmth in other people. related to teacher presence, jennings and greenberg (2009) refer to the constructs of self-awareness and social awareness forming teacher’s social and emotional competence in their prosocial classroom model. socially and emotionally competent teachers are effective classroom managers (jennings & greenberg, 2009). thus, the multilevel actions of classroom management a teacher engages in to create, support, and facilitate the goals of learning in the classroom (wolff et al., 2017) requires uninterrupted teacher presence. especially eye-tracking research has revealed the importance of teacher presence manifested by teacher gaze: expert teachers’ gaze direction indicates that students are being focused on, and that they are important (mcintyre et al., 2017; mcintyre et al., 2019). visual attention consists of conscious and unconscious gaze behaviour (tatler et al., 2014). human eye-movements can assist in providing information for inferring intentions and next actions (lukander et al., 2017), e.g., the teacher's pedagogical intention in the classroom situation. humans seem to have an innate tendency to learn by understanding where adults look (csibra & gergely, 2009). this has been interpreted as evidence of interaction between a teacher’s visual attention and intentions during student problem-solving (haataja, garcia moreno-esteva, et al., 2019); or how teachers’ gaze patterns differ during communicative (information-giving) and attentional (information-seeking) gaze (mcintyre et al., 2017). however, little evidence exists on teacher gaze behaviour during teacher-led sequences, such as when presenting a new activity, when a change in the social structure of collaborative work is introduced and during lesson wrap-ups. teacher presence is critical to giving task instructions where the teacher at best orients all students around a new learning task and its different phases. the present article reports on analyses of six teachers' eye-tracking data when they give similar task instructions on a geometry problem in a real-world classroom setting. we focused on the effect of the introductory (contextualizing the task and preparing students for it), collaborative (organizing the social structure of the activity) and reflective (looking back at the process and its outcome) instructiongiving on teacher gaze patterns. knowing where to look during different phases of the lesson gives teachers an advantage and yet another tool to promote learning. our work contributes to existing knowledge by addressing how instructions during different types of learning activities affect teacher gaze, and how gaze patterns may reveal teacher’s pedagogical intention. 1.1 teacher presence is important for classroom climate teachers attend to the classroom climate by their gaze behaviour. teachers show their presence by gazing towards students when they listen to them; this increases students’ experiences of closeness with the teacher (mcintyre et al., 2017). also, teachers make conscious and unconscious choices regarding eye-contact based on the prevailing social structures and personal intentions (tatler & land, 2016). maatta et al 94 | f l r previous studies have provided a comprehensive overview of teachers’ nonverbal immediacy behaviours, which abridge the physical and psychological distance between the interactants (mehrabian, 1970). these behaviors include smiling, nods, relaxed body posture, forward leans, movement, gestures, vocal variety, and eye contact in the teaching-learning context (andersen, 1979; mehrabian, 1970; witt et al., 2004). teacher immediacy influences student motivation, which, in turn, results in increased positive affect towards the teacher and the content studied (christophel, 1990; frymier, 1994; mccroskey et al., 1995; richmond, 1990). moreover, in the context of studying mathematics, teachers’ nonverbal immediacy behaviors increase students’ positive affect towards the subject and their perceived learning (mccluskey et al., 2017). thus, in the scope of our study, teacher presence is the realization of teacher immediacy behaviors, such as teacher gaze. related to teacher presence, kounin (1970) introduces concepts as overlapping and withitness to conceptualize teacher’s classroom management behaviours: overlapping meaning teacher’s ability to spread attention over different things in the classroom, and withitness comprising teacher awareness of what goes on in class, as if having eyes in the back of the head. consequently, teachers exert differing classroom management actions and attention when responding to the unanticipated situations happening in the classrooms as opposed to striving to meet the aims of the lesson (wolff et al., 2017). generally, teacher support and warmth in the classroom contributes to healthy learning relationships, and students implicitly adopt similar behaviour among their peers (gest & rodkin, 2011; hendrickx et al., 2016). yet, the slightest difference in students’ or the teacher’s behaviour may disrupt or enhance their attention and responsiveness. the dynamics of the classroom are complex. real-world settings differ by default from static or laboratory circumstances where for example the movement of informants, their position and number can be controlled. in classroom settings, experienced teachers apply their intuitive skills in deciding when to react or not to react in response to classroom events (berliner, 2004). teachers are monitoring their students in a nearly subliminal manner (mcintyre, 2016) by gazing toward students (mcintyre et al., 2017). flexible gaze, moving from emerging classroom problems quicker but visiting them more often typify the gaze behaviour of experienced teachers (cortina et al., 2015; van den bogert et al., 2014). gaze plays a pivotal role in social interaction because gaze signals enable understanding others’ mental states (baron-cohen, 1997; farroni et al., 2002). further, tomasello and carpenter (2007) discuss shared intentionality as a result of joint attention when interactants are experiencing the same thing at the same time and acknowledging that they are doing this. gaze following and joint attention create a shared space of psychological common ground enabling a shared goal and action plans to emerge in for example problem-solving activities (tomasello & carpenter, 2007). thus, we anticipate that the teacher's presence and the pedagogical intentions are manifested in the teacher's gazes at specific task-related targets. 1.2 teacher gaze reveals teacher’s pedagogical intention intention is characterized as a choice with commitment (cohen & levesque, 1990). to identify a teacher's pedagogical intention, we need to examine teacher gaze, as teacher gaze is one important channel for recognizing task-relevant areas in the teaching context. yet, vision literature identifying other teacher intentions than attentional and communicative underlying teacher gaze (mcintyre et al., 2017) is scarce. however, it is evident from practical experience that teachers do more than ask questions and talk when they teach. accumulated classroom experience makes teachers’ visual attention more intentional, and they rely more on top-down mechanisms (tasks influencing the gaze behavior) of perception (haataja, garcia moreno-esteva, et al., 2019). however, teachers are constantly interpreting the dynamic classroom situation and its visual features by focusing their gaze on targets that are important for successful task execution (wolff et al., 2016). thus, teachers’ gaze behavior is guided by the bottom-up (stimulus maatta et al 95 | f l r driven) mechanisms of perception, the top-down mechanisms of visual attention, and the goal-oriented intrinsic intentions (tatler & land, 2016). there are two aspects of the teacher’s pedagogical intention. first, there is a general strategic aspect, when the teacher aspires to meet the learning goals of the lesson. second, there is also a more tactical situated aspect, when the teacher is noticing, interpreting, and reacting to the situational changes in the classroom. van manen (1991) describes this as teachers being constantly active in the immediate interactive processes that aim at maintaining an authentic presence, and a personal relationship with the students. thus, the general and situated aspects of a teacher's pedagogical intention are intertwined in the teaching-learning context. we conceptualize teacher’s pedagogical intention in the mathematical collaborative problemsolving context in accordance with tomasello et al. (2005) where the actor has a goal toward which they choose an intended course of actions. here, we define the teacher's general goal of the lesson as executing the predefined problem-solving task with their students. further, the teacher chooses a set of actions (a plan) towards the goal based on teacher skills and knowledge employing their situationspecific professional vision in interpretation of the current situation and the problem-solving task. first, the intention is to accommodate students around a new learning activity; and second, to make the students focus on the problem-solving task. we do acknowledge the difficulty in fully knowing the intentions of others (johnson et al., 2017). also, as indicated by frey and fisher (2010) much of the instructional moves are part of an internalized decision-making process which teachers are not always able to articulate about. however, we assume that in mathematical problem-solving both the task design and the teacher’s plan of implementing the task accord with the educational goals of the situation. expert gaze has been shown to be task dependent: a meta-analysis has revealed that experts, compared to novices, focus their gaze for shorter periods of time, but more often on areas relevant to the task instead of non-relevant areas (gegenfurtner et al., 2011). teachers' gaze patterns vary depending on what students have previously done, what they are currently doing, and what they are going to do next. thus, we expect to discover previously unseen features of task-dependent teacher gaze. 1.3 teacher’s gaze depends on instructional type in the following, we are exemplifying the effect of the introductory, collaborative, and reflective instruction-giving on teacher gaze patterns. during introductory instruction-giving, teachers contextualize the task and prepare students for it, whereas collaborative instructions aim at organizing the social structure of the activity, and reflective instructions are given to look back at the process and its outcome. human eye-movements are highly task and content specific (rothkopf et al., 2016). teacher gaze is likely to differ depending on the task teachers are instructing students to complete (mcintyre et al., 2017). especially during introductory instruction-giving teacher gaze is crucial in eliciting the feeling of togetherness and placing importance on targets related to learning (mcintyre et al., 2017; wolff et al., 2016). further, shared attention is critical when signalling what is important when executing the task. for shared-attention to emerge, students engaged in problem-solving need to be aware of oneanother and to anticipate when future collective action is likely to happen (shteynberg, 2015). also, teachers deliberately share attention with their students towards critical learning resources of the task, and teacher presence re-engages students to be involved in the shared attention. as shteynberg (2015) concludes, co-attending a task implies that people also devote more cognitive resources to that task. in triadic engagement an individual interacts together with a goal-directed agent toward a shared goal (tomasello et al., 2005). in the present study we are assuming a novel third component in the referential triangle depending on the teacher gaze type. for contextualising gaze aiming at task readiness, we assume the third component alongside with the teacher and student to be conceptual i.e., the shared goal of starting a new learning task. for collaborative gaze that aims at task focus we anticipate the third component maatta et al 96 | f l r to consist either of the solution paper, the laptop, the calculator, or students' personal items not directly related to the task. especially critical to collaborative instruction-giving, teachers are able to read and interpret gaze direction to promote shared attention (shteynberg, 2015; baron-cohen, 1997). moreover, research delineating the relation between the importance of the gaze target and gaze frequency (cortina et al., 2015), student-centeredness and gaze flexibility throughout a lesson (mcintyre, 2016), and the role of teacher’s intentions on gaze behaviour when scaffolding students’ problem-solving (haataja, garcia moreno-esteva et al., 2019) resonate with teachers’ aim of creating connections in the classroom. according to another study (haataja, toivanen et al., 2019) students follow their teacher's gaze and especially the dyadic eye contacts between the teacher and students serve two main aims: first, teachers' eye contact enhances immediacy, creates positive emotions, and prevents task-related frustration. second, teachers' responses to student-initiated eye contacts promote collaboration within the peer group as well as steer students’ attention to task-relevant targets. interestingly, an inclination may be that teachers’ pedagogical intentions of support go beyond the actual problem-solving phase and teachers convey their intentions already when giving the collaborative task instructions. the affective and cognitive scaffolding categories derived by van de pol et al. (2010) align with notions of polya (1948) who instructs teachers to stir up their students’ curiosity and to enhance their desire to solve the problem. haataja, garcia moreno-esteva et al. (2019) conceptualize teacher-student eye contact through scaffolding categories during collaborative problem solving, whereas the present study embraces the role of teacher gaze when giving instructions during lesson phases that both precede and follow the actual problem-solving activity. for students to succeed in the problem-solving task both applying a model for mathematical problem-solving in general (lester, 1989; polya, 1948), and comprehension monitoring in particular (schurter, 2002) are emphasised. in schurter (2002), the last phase of comprehension monitoring consists of teacher prompts such as posing questions of the accuracy of the answer, discussions on whether the answer is reasonable or not, and looking for alternate procedures to find the solution. in our study we expect teachers not only to monitor the comprehension but also to steer student engagement by gazing at problem-solving tools, solution papers, and students' personal belongings during collaborative instruction-giving. 1.4 a mixed analytic approach in examining teacher gaze 1.4.1 durational and proportional analyses our research focus, how teachers’ pedagogical intention during different types of instructiongiving are revealed by gaze, was analysed by aggregated teacher gaze measures comparing mean dwell durations and proportional durations for areas of interest (later aois) representing student presence and student engagement. we expanded the view on communicative and attentional gaze (mcintyre et al., 2017) by addressing additional teacher intentions underlying gaze. we labelled the two teacher gaze types for task readiness as contextualizing gaze and for task focus as collaborative gaze. thus, we expect to capture the role of teacher gaze for task readiness and how task focus is sustained by placing importance in task-relevant areas. 1.4.2 qualitative analysis of scanpaths and teacher verbalizations investigations into teacher intention were continued by qualitative analysis. we wanted to explore how the introductory instruction-giving and teachers’ intentions account for teacher scanpaths in real world dynamic classrooms, for scanpath analysis see e.g. (anderson et al., 2015). in teaching contexts, scanpaths have been analysed for differences between expert and novice teachers (mcintyre & foulsham, 2018). in this study, analysing teacher scanpaths qualitatively offers an insight into teachers’ gaze behaviour critical to classroom management. we wanted to depict in detail how teachers use their gaze for enhancing task readiness when starting a lesson with an introductory task instruction. maatta et al 97 | f l r examining the very beginning of the instruction-giving resulted in qualitative analysis of the first 20 gaze targets for each teacher. scanpaths are visualizations of how the eye physically moves through space (holmqvist & andersson, 2017). the seminal work by noton and stark (1971) presenting “scanpath theory” argues for the top-down character of eye-movements, especially when viewing a previously seen image. consequently, the visual attention of a participant is guided by the stored image and the scanpath used to view it. further, foulsham and underwood (2008) tested this notion concluding that greater similarity can be found in scanpaths of the same participant viewing the same image twice than a different participant viewing the same image. thus, top-down and bottom-up explanations end up with the same prediction for viewing static images (foulsham & underwood, 2008). drawing from research on real world stimuli and tasks holmqvist and andersson (2017) conclude that for example task driven plans are behind our eye-movements. also, scanpath planning and look-ahead fixations precede the actual behaviour related to the task (mennie et al., 2007). moreover, additional research on teacher scanpaths is needed to shed light on expert teacher vision (kaakinen, 2020). verbal data and eye-tracking data can be analysed separately to see if the inferences drawn from both datastreams align (holmqvist & andersson, 2017). in order to make more concrete statements about the cognitive structures underlying teachers’ gaze patterns we linked the analysis of verbal and eye-tracking data (jarodzka et al., 2017). both types of data, eye-tracking data and verbal data consist of datastreams over time, and they are composed of trackable events (holmqvist & andersson, 2017). the best evidence for teachers starting and ending their gazing at the predefined aois of interest is their concurrent verbal data (ericsson & simon, 1980). by analysing teacher verbalizations together with teachers’ scanpaths we aim to recognize two teacher’s pedagogical intentions related to the introduction of the problem-solving task, first assuming teachers to accommodate students around a new learning activity; and second, to make the students to focus on the instructional resources (black/whiteboard) by gazing at them and suggesting shared attention. we transcribed the separate recordings of teachers’ voices by typewriting them. then we transferred the transcribed written text to the elan annotation program (elan, version 5.3, 2018) as a separate tier and aligned it with the gaze coding onsets. 1.5 research questions we use eye-tracking to examine whether teacher given instructions affect how much and how long teachers look at certain critical targets. moreover, we want to delineate the momentary gaze behaviour by analysing teacher scanpaths during the starts of the introductory instruction-giving of all six teachers. we expect teachers' pedagogical intention to change with the students' tasks and this, in turn, changes how teacher presence is expressed as well as how task focus is maintained through gaze. thus, we report on teachers’ gaze behaviour for aggregated variables named student presence and student engagement in research question (rq) 1 and for detailed variables included in student engagement in rq 2. however, in rq 3 we will also report on teachers' gaze behaviour both for instructional resources as black/whiteboard and for targets that guide teachers’ work in the classroom i.e., lesson plans which we call as critical activity targets. specifically, rq 1 asks: how do instructional types predict teacher gaze for (a) student presence and (b) student engagement during instructions for problem-solving tasks? as teachers monitor the classroom by their gaze, they create a sense of togetherness and prepare students for the upcoming learning session. by doing this, teachers ensure that students are both opt to listen to the instructions given by the teacher, and ready for a new activity. we expect teachers to strive for student presence and look more at student faces, bodies, and hands when giving introductory instructions compared with instructions given for collaborative tasks. we anticipate that teachers’ gaze behaviour would place more importance on student engagement during collaborative instruction-giving than during introduction of the task and reflection. we assume that this is done to guide students' attention maatta et al 98 | f l r to critical problem-solving tools and simultaneously make it clear to students to keep away from distractions. additionally, teachers’ gaze behaviour should also reveal teachers’ interest in how students use their learning tools (calculators, rulers, laptops etc.) and do not use their distracting objects (headphones, mobile phones) indicated by longer dwells at these targets during collaborative task instructions compared with introductory work instructions. thus, rq 2 asks: how do instructional types predict teacher gaze for the three gaze targets that student engagement comprises: (a) solution paper, (b) problem-solving tools and (c) distracting objects? we expected that teachers spend longer looking at student solutions when giving instructions for collaborative work compared with introductory tasks. moreover, we anticipate teachers to place importance on the solution paper by their visual attention when instructing on the reflective phase of the lesson compared with collaborative task instructions. we think that teachers do this to promote student reflection on their solutions, to encourage students to look for alternate ways to solve the problem and reinforce correct answers by concentrating on the solution. finally, rq 3 asks: do instructional types predict teacher gaze for (a) instructional resources and (b) critical activity target? we expect teachers to spend longer looking at the black/whiteboard (instructional resources) during introduction and reflection than during collaborative task instructions. also, teachers’ gaze behaviour should reflect their need to reference lesson plans throughout the lesson and thus not yield any differences in the length of dwells toward the teacher desk and teacher notes (critical activity target). 2. method 2.1 participants and procedure the data were obtained from six mathematics teachers (four females, two males), each teacher giving a 45 min lesson with similar content for different students in five lower secondary schools in southern finland. each class was visited twice, and the data are from the second visit. teachers were between 30 and 56 years old, (m=46, sd=10.79) and with 3-31 years of teaching experience (m=13.83, sd=9.50). the students in the six classes were all 9th graders (15-16 yrs). the class size varied between 9 and 19 students (total of 94 students). schools and teachers alike were selected based on their willingness and voluntariness to take part in eye-tracking research. additionally, a written consent was obtained from the participating teachers and students. also, students’ parents were informed. the data collection took place in spring 2017. all participating teachers were given identical task instructions to be delivered along with the problem-solving lesson. the objective of the problem-solving task was to find out the optimal solution to a geometry problem. the research team set up the eye-tracking glasses that had been calibrated at the end of the first lesson and teachers received written task instructions for the lesson (see table 1 for the complete lesson plan). the task instructions referred to in this research are presented in table 2. maatta et al 99 | f l r table 1 the complete lesson plan for mathematical problem-solving phase of the lesson student activity teacher activity 1 preparations teacher informs on the lesson structure, students fetch papers and rulers 2 individual seat-work teacher poses the problem and gives instructions for individual work. 3 pair work teacher gives instructions for pair work. 4 group work teacher gives instructions for group work. 5 individual presentation teacher gives instructions for presenting the solutions individually. 6 whole class discussion teacher leads the discussion maatta et al 100 | f l r table 2 model instructions for mathematical problem-solving task note. the teachers delivered the instructions in their personal style. 2.2 collection and preprocessing of data our target lessons were recorded by three stationary video cameras capturing the action and verbal communication in the classrooms; one camera followed the teacher; the two other cameras were directed towards students wearing the eye trackers. in addition, the teacher's voice was recorded by clipon microphones. further, mobile eye-tracking devices recorded teachers’ eye-movements. the device resembles a protective eyewear equipped with two eye cameras, a scene camera together with electronics connected to a laptop in a backpack letting the teacher move around in the classroom. the eye-tracking devices together with the algorithms and a software for processing the eye-tracking data with an average accuracy of 1.5 degrees of visual angle were designed at the finnish institute of occupational health (lukander et al., 2017). the software records the video frames and produces a video of the scene camera, superimposed with a gaze point. the frame rate of the video camera varies depending on lighting conditions being optimally at 30 fps. in this study we analysed eye-tracking data that was collected when teachers were giving introductory, collaborative, and reflective task instructions for a mathematical problem-solving task. to map the teacher's eye-movement behaviour we defined our main metric as a dwell. researchers agree phase of the lesson type of instruction model instruction 1 introduction you have four cities; they lie on the corners of a square. you have most likely seen how two places are connected by cable or optical fiber. first, work alone and try to find at least three different ways how the cities might be connected. think which of these is the best. 2 collaboration (pair work) join to work pairwise and discuss what is the most effective way of connecting the four cities, so that the least amount of cable is used i.e., so that the total length of the connection is as short as possible 3 collaboration (group work) get into groups of four and continue searching for the best possible solution. 4 reflection now present your best solution on the blackboard, please. maatta et al 101 | f l r that dwells do not exist without fixations, yet literature lacks a unison definition of a dwell. orquin and holmqvist (2018) define dwells as one or more consecutive fixations in an aoi. the dwell is a visit in an aoi, from entry to exit (holmqvist et al., 2011). dwell time in turn is the sum of all fixation durations during a dwell in an aoi. further, as holmqvist et al. (2011) point out, dwells are oftentimes more dispersed than fixations, longer than fixations and can only be calculated if the stimulus has been categorised as aois. as we were interested in specific predefined targets in the classroom upon which teachers were looking, we coded teachers’ gaze tracking data in terms of areas of interest (aois). accordingly, the first author annotated all dwells lasting ≥3 consecutive frames (120 ms) of gaze on an aoi with elan software (elan, version 5.3, 2018) using dwell time as our coding unit. the predefined gaze targets i.e., aois that we were interested in, were derived from the dependent variables in our research questions. when coding, the areas of aois were defined from the superimposed gaze points in the scene camera video. gaps between two dwells on the aois were signalled as a blink or a technical error (for example, eye-tracker losing the pupil of the other eye) by the eye-tracking device and were not included in the coding. informed by the research questions of our study, all dwells were manually labelled indicating the gaze targets: student presence (student face, body, hands, or gestures), student engagement (students’ solution paper, problem-solving tools, and distracting objects as headphones and mobile phones), instructional resources (black/whiteboard and clock), and critical activity target (teacher notes and teacher desk). the detailed coding scheme is provided in table a3 in the appendix. 2.3 measures and analyses our data consisted of 2 108 teacher gaze dwells where the shortest dwell coded was 120 ms and the longest was 16564 ms (mdn = 292 ms, m = 502 ms, sd = 789.00). on average, each teacher's gaze behaviour comprised 351 dwells (sd=168). the total coded dwell duration for all participants was 17 minutes and 32 seconds. errors, blinks, and saccades that were not coded as dwells accounted for 8 minutes and 20 seconds (32%) of the total gaze behaviour, resulting in the overall duration of 25 minutes and 52 seconds of giving instructions during the six lessons. the dwells were further divided between the gaze targets (see table a3 in the appendix), and for each teacher dwell target (e.g., student engagement), we calculated the mean dwell duration, as well as the proportional duration (how long a teacher looks at one area of interest relative to the length of dwells to other possible regions.) observations of not looking at a specific target (i.e., mean dwell duration = 0) were also included when forming the dependent variables. this is done because not looking at all is a signal of the teacher's decision to target the attention toward didactically more crucial targets. omission of those observations that indicate not looking would only be justified in cases when the gaze target is not present in the classroom at the time of recording. thus, the teacher's decision not to look at certain targets is considered as a didactic choice. the durations were non-normally distributed, with skewness of 8.32 (se = 0.053) and kurtosis of 113.15 (se = 0.107). aggregating the gaze dwells into seven dependent variables mitigated the skew but did not remove it, in part due to the number of zero-value observations for mean dwell durations. moreover, there is almost always a positive skew in fixation duration data (holmqvist et al., 2011), and this produces a similar skewness in dwell data. instruction type was used as a categorical independent variable (three levels: introductory, collaborative [collapsed over the phases 2 and 3, see table 2], and reflective). our dependent variables included both the meanand proportional dwell durations of i) student presence and ii) student engagement (rq 1), iii) solution paper, iv) problem-solving tools and v) distracting objects (rq 2), vi) instructional resources, and vii) critical activity target (rq 3). note that the dwell durations assessed in rq 1 are composed of the dwell durations of “student presence” (i.e., student presence = student face + body + hands + gestures) and “student engagement” (i.e., student engagement = solution paper + problem-solving tools + distracting objects). for brevity, all the results on mean dwell durations, and a summary of proportional dwell duration results, are presented in the appendix (figure a2 and table a5), while the main text focuses on proportional dwell durations (the results for both variables were maatta et al 102 | f l r very similar, but proportional dwell durations account for the uneven number of dwells toward each aoi). to evaluate research questions 1-3 (i.e., do instructional types influence teachers’ visual attention?), we used the friedman test, which is a well-known non-parametric test for comparing three or more groups in repeated measurements. in our study, six teachers gave task instructions for introductory, collaborative, and reflective student work, corresponding to three repeated measurements for each teacher. post hoc comparisons were done using the conover-nemenyi method with bonferroniholm corrections for multiple comparisons. human activities (e.g., teachers giving task instructions) draw on visual information acquired by pointing the high-resolution fovea at the gaze targets from which information is required (tatler et al., 2014). during the relocations, the eye stops for fixations approximately three times a second (holmqvist et al., 2011). by investigating teacher scanpaths during the critical lesson phases as when introducing new tasks and reorganizing the social structures for collaboration, we expect to extend the previously discovered features of scanpaths to the crucial teacher-led moments of giving instructions. real-world eye tracking research on eye-movement patterns and differences between teachers’ scanpaths aid understanding of how teachers use gaze to create an optimal learning climate as well as to sustain task focus. thus, to complement the quantitative analyses, we examined the scanpaths of all six individual teachers for their first 20 gaze targets when they were giving the introductory task instructions. these findings are interpreted qualitatively. an overview of analyses is summarized in table 3. table 3 the analyses used in this research research question/aim of the study target of analysis data source coding method analytic method rq1-rq3 proportional dwell durations gaze recordings elan annotation friedman test and post-hoc comparisons rq1-rq3 mean dwell durations gaze recordings elan annotation friedman test and post-hoc comparisons teachers’ pedagogical intention teacher scanpaths classroom video, gaze recordings selection of the first 20 and 5 aois during introductory instructions qualitative descriptions teachers’ pedagogical intention teachers’ verbalizations classroom voice recordings, gaze recordings  transcriptions of teachers’ introductory instructions qualitative descriptions note. results for mean dwell durations are presented in the appendix, figure a2  maatta et al 103 | f l r 3. results 3.1 quantitative findings on proportional dwell durations rq 1: instructional type predicting teacher gaze for (a) student presence and (b) student engagement in problem-solving the friedman test was non-significant for student presence between differences in the "introduction", "collaboration," and "reflection" parts of the teaching event. regarding student engagement the test showed statistically significant differences (χ2(2) = 7, p = .03, w = 0.58) representing a large effect (see figure 1, top row and table a5 in the appendix for details). post hoc tests found significant differences between collaboration and introduction (p = .003) and between collaboration and reflection (p = .024). rq 2: instructional type predicting teacher gaze for (a) solution paper, (b) problem-solving tools and (c) distracting objects in problem-solving the friedman test was found statistically significant between differences in the "introduction", "collaboration," and "reflection" parts of the teaching event for solution paper (χ2(2) = 6.9, p = .03, w = 0.57), problem-solving tools (χ2(2) = 10.2, p = .006, w = 0.85), and distracting objects (χ2(2) = 10, p = .006, w = 0.83), representing large effects. again, the post-hoc tests showed significant differences when comparing the introduction and reflection periods of the teaching event against the collaboration period (see figure 1 and table a5 in the appendix for details). rq 3: instructional types predicting teacher gaze for (a) instructional resources and (b) critical activity target the friedman test indicated statistically significant differences in the "introduction", "collaboration," and "reflection" parts of the teaching event for instructional resources (χ2(2) = 8.43, p = .014, w = 0.70), representing a large effect. moreover, post hoc tests indicated the differences to be significant between collaboration and introduction (p = .004) and between collaboration and reflection (p < .001). there were no statistically significant differences between the "introduction", "collaboration," and "reflection" parts of the teaching event for critical activity target (see figure 1, bottom row and table a5 in the appendix). maatta et al 104 | f l r figure 1. results for research questions (rqs) 1-3 for proportional dwell durations (pdds): how do instruction types predict teacher gaze for student presence and engagement (rq 1, top row); solution paper, problem-solving tools and distracting objects (rq 2, middle row); instructional resources and maatta et al 105 | f l r critical activity target (rq 3, bottom row)? instruction type refers to a teacher’s instruction given for introductory (individual seat-work) collaborative (pair work and group work with three to four students) and reflective (solution presentation) student work. blue dots and lines represent the group median values, with the friedman test p-value, when comparing the three groups, indicated at the top. the faded lines are individual observations from six teachers. * p < .05; ** p < .01; *** p < .001 (post-hoc tests using conover-nemenyi with bonferroni-holm corrections for multiple comparisons). 3.2 qualitative findings next, we are using both qualitative interpretations of scanpaths and descriptions of concurrent verbalizations to shed additional light on what teachers look at when they start a new activity and how teacher speech is linked to their gaze behaviour. 3.2.1 overview of teachers’ scanpaths a general notion of teachers’ gaze patterns when they are introducing the problem-solving task is that teachers mostly look at student targets, mainly at student bodies and faces. thus, when exploring the first 20 areas of interest the most common gaze target was the student, either the student’s face or body except for one teacher who chose to start the lesson by projecting the task instruction on the whiteboard. that decision resulted in most teacher gazes at that specific aoi. an overview of teacher scanpaths and the percentage of dwell counts for the first 20 aois are shown in figure 2. before suggesting more general tendencies we grouped the first 20 aois into four bins. at this stage, we noted that, overall, teachers looked at critical activity targets (teacher notes) first, then at student bodies and faces, followed by looks at the instructional resources (the whiteboard), and the critical activity targets (teacher desk) before they looked at another critical activity target (teacher notes) again. thus, the teachers’ gaze behaviour can be interpreted as flexible, demonstrating a readiness to respond to situational differences arising in the beginning of a lesson. the scanpaths of teachers’ first 20 aois also show that teachers in general demonstrate their pedagogical intention in prioritizing students, embracing them around a new learning activity by maintaining eye contact with students and by placing importance at task-relevant targets. moreover, since gazes targeted both at the student face and student body are the most frequent in the sample for the first 20 aois, there seems to be a face-type and body-type gaze pattern. distinctive for the face-type gaze pattern are consecutive gazes at the student face, whereas characteristic for the body-type gaze is that the gazes targeted at student bodies precede gazes at student faces. one explanation for the face-type gaze is the situational dynamics in the classroom. teachers one and four are both about to start the instructions for the problem-solving task when they track restlessness, movement, and hear something that catches their visual attention. especially, teacher one, who after noticing the gesture of a student calls another student by name and at the same time directs the gaze at the face of the student whose name was called. the gaze sequence then continues with repeated gazes at the student’s face. here, the successive dwells on a student’s face can be more effective to restore the working conditions than the single dwells at individual students. however, teachers one and four are experienced in classroom teaching and they seem to be aware that paying attention to students across the classroom is important. an alternative explanation for the face-type gaze is that teachers build their momentary interaction on the eye contact between them and the student. we found the scanpath of teacher five to reflect a consistency in gazing repeatedly at the student face before moving to task-relevant targets possibly suggesting shared attention. the examples demonstrate teachers’ visual attention as a top-down process where their pedagogical intentions guide the gaze behaviour. maatta et al 106 | f l r figure 2. a. teacher scanpaths of the first 20 aois and b. proportional dwell counts across teachers during introductory task instructions. also, a few (between three to five) successive dwells at the same aoi (student body) exemplify the body-type gaze. there are a few possible explanations accounting for the transitions between the face-type and body-type gaze. overall, the abundance of transitions emerged in gaze sequences when teachers introduced the problem-solving task may be a result of teachers’ adaptation to the complexity of the classroom situation. moreover, classrooms contain by default movement which is likely to bring about teacher’s reflexive eye-movements. teacher four for example uses check-ins for a student, whose hand movement is noticed resulting in a body-face-body gaze chain. teachers with experience tend to pay attention to students’ posture and bodily movements featuring bottom-up processing of gaze behavior (wolff et al., 2016). it may be that teachers track waving or gestures and the first target aois are oftentimes students’ hands followed by identifying the student face. it is characteristic for teachers three and five to visually scan the classroom when starting to give the introductory instruction. teacher three seems to distribute the visual attention across seven different aois when we examine the first 20 gaze targets, whereas teacher five looks at five aois. for teacher three the student targets gazed at represent ten different students, while teacher five only looked at four students. interestingly, this difference cannot be explained by class size as both classes contain 19 students, neither does the duration of the excerpt favour teacher three when explaining the number of students gazed at. still, our preliminary notions of how the gaze patterns of teachers three and five differ, concentrate on the number of gaze targets and how many individual students are looked at. one plausible interpretation is that teacher five is confident with how to cope with the situations arising, without the need to check what is going on all the time. this in turn could interestingly also be a social cue for the students, indicating that all is under control. teachers with less classroom experience knowing neither the routines of school, the subject in depth, nor the students would need to actively use check-ins much more often. and, vice versa, this would again be a signal to the students (together with hesitance, faltering voice etc) of less control. considering the reasons behind the specifics of teachers' visual attention, teacher five acknowledges the power of a direct gaze when conveying the maatta et al 107 | f l r need to listen to the instructions. five consecutive gazes at the student’s face seem effective enough to convert an emerging disruptive incident back to focus on the task. moreover, teacher five uses gaze to gather information on the classroom structure and its dynamics. the very first aoi in the scanpath for teacher five is a novel on a student’s desk. noticing the book at the beginning of the lesson is crucial for a later incident where teacher five spots the book again, tells the student to put the book away and instructs the student to concentrate on the problem-solving task instead. thus, teachers’ scanpaths interpreted in the context of the lesson help us to identify which decisions seem to be important and which are less important, and when they are best acted upon. 3.2.2 concurrent verbalizations during introductory instruction-giving in figure 3 we present the scanpaths of the first five aois complemented with momentary descriptions of teachers’ concurrent verbalizations. the onset for transcribing teacher speech was set when the teachers started to talk or signalled the start of the lesson with gestures, started to gaze at students or by using discourse markers as uh, um, or hmm. figure 3. teacher scanpaths of the five first aois and verbalizations. in addition, the descriptions of the classroom situations for each teacher are provided as they can be seen in the video. t1: teacher stands next to own desk, looks at the lesson plan and counts to three while scanning the classroom. the teacher looks at the face of an inattentive student, steps toward them and says their name out loud. t2: teacher directs their first glance at the lesson notes under the document camera; then gazes at the ready figure on the whiteboard while illustrating cities placed in vertices of a square. maatta et al 108 | f l r t3: teacher begins by standing in the front, gazing at student bodies, faces and desks. they then walk around while hesitating on how to explain the main task. t4: teacher stands in front of students, gazes first at the lesson plan, and then back-and-forth between one inattentive student's body and face while calling them by name. t5: teacher stands in front of students and looks at a student’s desk and the lesson plan followed by consecutive gazes at the face of a disruptive student. t6: teacher stands behind the teacher’s desk and looks at a student’s face. the teacher turns around toward the whiteboard and marks the cities as dots in the vertices of a square. noticing the difficulty to work without glasses, the teacher turns facing the students and draws a square in the air. at the same time, the teacher gazes at the classroom cabinets ending up with looking at the researchers in the back of the classroom. evidence for the contextualizing gaze and its conceptual component (a mutual goal) in the referential triangle in classroom teaching for problem-solving was particularly demonstrated by teachers one and four whose gazes were first directed at the lesson plans but then moved to student bodies and faces suggesting that teachers are actively trying to create student presence and a sense of togetherness. the two excerpts contain a typical situation in the classroom when students are involved in something else, and the teacher decides to address an individual student and demonstrates teacher presence. also, the classroom videos indicate how teachers’ postures, body orientations, and forward leans accompany teachers’ gaze at students’ faces together with verbalizations identifying the gaze target. thus, disruptive behaviour gains teachers’ attention. further, teachers two and six demonstrate gaze behaviour where the intention of the activity is conveyed by gazing at the whiteboard. moreover, the gazing behaviour of teachers two and six highlight that teachers place importance at instructional resources when introducing a new learning activity. teachers three and five react differently to disruptive student behaviour and classroom situations in general when they give the introductory instructions. teacher five tracks the restless student who is turning around in the chair, however, teacher five does not intervene verbally, but by gaze behaviour. teacher experience may explain this notable difference. teacher five is demonstrating a calm, and friendly atmosphere by greeting the students, explaining the task of the day by using taskrelated terminology and providing the students with hints of the lesson structure. on the contrary, teacher three starts by verbalizing the main idea of the problem-solving task using pronouns, walks around in the front of the class and ends up intervening two boys engaged in off-task behaviour. 4. discussion we examined aggregated teacher gaze measures between instructional types analysing proportional and mean dwell durations for gaze targets comprising student presence (face, body, and hands), and student engagement (solution paper, problem-solving tools, and distracting objects). moreover, we used the same measures analysing the detailed components of student engagement, the gaze targets in instructional resources (black/whiteboard, the clock) and the targets that guide a teacher’s work in the classroom which we call as critical activity targets (lesson plans). additionally, to identify teacher gaze patterns for the first 20 aois, we analysed teacher scanpaths and teachers’ concurrent verbalization qualitatively. here, we intend to identify the most frequent gaze target for each teacher and the general tendencies of what the teachers gaze at when they maatta et al 109 | f l r introduce a new task. further, we are interested in how teacher scanpaths and verbal data further illuminate our research foci, teacher gaze for task readiness and task focus. first, the current study adds to the existing process-oriented literature describing teacher behavior by tracking teacher gaze. second, our investigations delve into the role of context in influencing teachers’ gaze behaviour. third, we have analysed the effect of instruction type on teacher gaze for detailed gaze targets suggesting a previously not detected conceptual lens for teacher intentions underlying teacher gaze. in general, our study echoes the notions of previous research about prioritising students (mcintyre et al., 2017; haataja, garcia moreno-esteva, et al., 2019). as concluded by haataja, garcia moreno-esteva et al. (2019) teachers focus more on student faces during affective scaffolding when teacher intentions are characterized by motivating students and reducing their frustration. our results show that teachers gaze toward students regardless of the instruction type. thus, we were able to identify that teachers regard student presence as important during all three types of instruction-giving. building on that, we are suggesting that teachers’ pedagogical intentions in our research are first, to prepare students for a new task and second, to keep them focused and engaged in the task at hand as shown next. our findings also support the interpretation that teachers gaze at targets intended for student engagement (solution paper, problem-solving tools and distracting objects) result in task focus, and that the teachers thus aimed for a state of shared attention. the present study suggests that teachers should place importance at targets that are not linked to the task i.e., the distracting objects on students’ desks during collaborative instruction-givings. as shown by tomasello and carpenter (2007), a psychological common ground of a shared goal and plans were mediated to the students by teacher gaze. shteynberg (2015) also notes that teachers deliberately share attention with their students towards critical learning resources of the task. hence, teacher presence re-engages students to be involved in the shared attention. our study suggests that teachers can monitor the process of shared attention by placing importance on the solution when looking at the solution paper at the same time when instructions are given. indeed, teachers have been reported to focus on students' solution papers during cognitive support when solving mathematical problems (haataja, garcia moreno-esteva et al., 2019). however, teachers need to be well informed of the shared-attention mechanisms as the focus of attention can easily shift from the task to the performance instead and rather impede than facilitate the problem-solving itself. contrary to our hypothesis, we did not notice any statistically significant differences for teacher gaze at targets indicating student presence between the three studied instruction types. previous research on gaze duration has highlighted teachers’ tendency to direct long gazes toward students they are helping (mcintyre et al., 2017). however, brief gazes in turn may either indicate expertise, and skilful teachers are able to allocate their attention between a number of students; or short glances can also reflect uncertainty or hesitation to react to what happens in the classroom (van den bogert et al., 2014). overall, teachers target frequent gazes at the student body in the very beginning of the lesson when a general overview of the present students is achieved. accordingly, teachers looking at student faces preferably engaging in eye-contact indicates teachers’ interest in individual needs of the students, as well as motivating and embracing them in the ongoing learning activity. moreover, there is research evidence to suggest teacher expertise to be linked with more gaze at students’ posture and bodily movements when viewing problematic classroom management events (wolff et al., 2016). also, long stares at restless students in comparison to repeated short glances at individual students and student groups serve an identical purpose noted in the current study: to make students ready for instructions for a new learning assignment. this may be an interpretation of why we did not find any statistically significant differences in the mean dwell durations nor proportions of contextualizing gaze that aims at student presence during introductory, collaborative, or reflective instruction-givings. indeed, teachers’ gaze behaviour can be explained by the innate tendency to rely on eye contact between humans (csibra & gergely, 2009). it must be noted that the seating schemes in real classrooms may account for differences in the suggested face-type vs. body-type gaze. if the teacher sees the faces of the students, and the students are sitting maatta et al 110 | f l r facing the teacher, an eye-contact is more evident than in classrooms where the students are sitting in groups facing their peers. teachers seem to keep students in focus throughout the lesson. teachers’ aim is to monitor the classroom by their gaze, to create a sense of togetherness and to prepare students for the changes in the learning activities. taken together, we incline that the teachers would be featuring a specific contextualizing gaze when introducing the task to ensure that students are listening to the instructions and understanding them. teacher intentions behind contextualizing gaze are the same as for teachers’ general didactic behaviour comprising motivation, drawing attention and raising interest toward new learning activities. overall, teachers' gaze behaviour reveals the essence of their intentions yielding interest in students’ problem-solving tools and results, and we see teachers spending longer looking at students’ solution papers, desks, and calculators when they are giving instructions for collaborative work compared with the two other instruction types. the qualitative analyses show that teachers react by their gaze behaviour to information that has instructional significance. it may be that they are able to predict more accurately the possible future events based on their representations of the classroom. thus, we speculate that teachers are featuring a specific collaborative gaze when instructing on the collaborative tasks to promote task focus indicated by more and longer mean dwell durations at solution papers, problem-solving tools and distracting objects. in turn, the statistically significant differences in mean dwell durations for solution paper between introductory and reflective task instruction is a research artefact; there is nothing to be found in the solution paper yet when introducing the problem-solving task. further, the crucial phase of the problem-solving task when discussing alternate solutions, reinforcing correct answers to the problem and students motivating their answers could have yielded longer gazes at the solution paper during the reflective instruction-giving, but that was not the case. 4.1 limitations there are a number of limitations that need to be addressed for the present study. first, the small sample size (n = 6; 3 repeated measures) contributes to the limited generalizability of our findings. despite the relatively small sample size and its contextuality, we were able to sample an interesting data set with teachers who utilized similar task instructions throughout the studied lessons. the controlled setting (the given frame for the lesson), as well as the natural settings involved (as compared with laboratory settings) add to the ecological validity. hence, actual conclusions can be drawn for educational practice (jarodzka et al., 2017). on the other hand, volunteer teachers may not be wholly representative, and the controlled setting probably also limits the teacher’s authentic actions. second, there is no existing framework to entirely depict a teacher's didactic (gaze) behaviours in real-world teacher-led moments. on one hand the frameworks are created for detailed representation of the problem-solving phase (haataja, garcia moreno-esteva et al., 2019). on the other hand, they are limited to the communicative and attentional gaze (mcintyre et al., 2019). nevertheless, our measures were designed to support the assumption that certain aois are detected and processed after which the interpretation drew from other data sources to triangulate the data collection. analysing teacher scanpaths and verbalizations helped us to delineate teachers’ intentions. third, the coding of gaze data is a lengthy process; only the first author annotated the data. however, the coding procedures were followed closely by the co-authors and miscodings were addressed consistently for all lessons before statistical analyses. further, the detailed coding of dwells representing research specified gaze targets allowed us to aggregate teachers’ gaze behaviour resulting in a model of explanatory power. while the data does not support exact statistical conclusions and such, this type of data, however, does provide valid opportunities for observing the phenomena of task readiness and task focus. maatta et al 111 | f l r fourth, the categorization to introductory, collaborative, and reflective task instructions provides an appropriate basis for aggregation of data for future studies. the categorization was done to give an overview of actual teacher work in the classroom, i.e., introducing a new activity, reorganizing the social structures for collaborative activities, and reflective presentation of the results. however, it should be noted that not merging pair work and group work into a single collaborative instruction-giving would allow further analysis of how shared attention emerges in the different types of collaboration. this would certainly be an interesting future research avenue. we contend our methodological design to identify teacher intentions revealed by gaze patterns worth researching in the future. 4.2 implications teachers may benefit from empirical evidence of how their gaze behaviour can influence their instruction. for example, how an optimal learning climate may be influenced by guiding attention to certain students, and away from others. teachers may also guide the students’ attention during lesson transitions by looking at task related targets. sometimes attention also needs to be directed to distracting behaviour or objects, showing that they have been noticed but are not necessarily commented upon. teachers may also aim at and enhance shared attention by gazing at e.g., student solutions. 4.2.1 contributions to eye-tracking research our findings contribute to eye-tracking research derived from real-world classrooms and the results both underscore teachers’ student-centred attention despite type of instruction and result-centred attention during collaborative instruction enabling shared attention to emerge. consequently, an interesting topic for further work in examining teachers’ gaze sequences during teacher-led moments would be controlling e.g., teacher position, class size, and students’ sitting schemes. 4.2.2 contributions to teachers’ professional development and teacher education new insight into teachers' gaze behaviour is meaningful for both preservice and in-service teacher training purposes. we see an opportunity to apply the results in lesson feedback sessions, part of mentoring processes and when in need of getting an overview of one-way teacher gaze behaviour. however, it is not reasonable to think that merely telling preservice or in-service teachers the fact that gazing itself at for example students would improve their professional performance. rather, the results of our study could be applied in programs such as the one reported by stuermer et al. (2016) where conceptual and practical knowledge are integrated into a rigorous training to describe, explain, and predict classroom situations. moreover, didactizing the expert teacher eye-movement modeling example to be applied in teachers’ professional development is certainly worth investigating (jarodzka et al., 2017). hence, teachers’ awareness of the top-down and bottom-up mechanisms of visual perception is important. moreover, knowledge of the importance of gazing at students in general, and task-relevant targets in particular are critical to the success in teacher-led phases of the lesson. it is also important that teachers should strive to enable a shared attention state. that can be done by fusing the individual students into “we” when giving instructions for collaborative problem solving. students allocate more of their cognitive capacity to those details of their environment that are thought to be co-attended with their peers (shteynberg et al., 2014). in recent research experienced teachers have been found attempting at shared attention when focusing students’ attention on important features that enhance their conceptual understanding of mathematical concepts (pouta et al., 2020). an implication from our study is to continue researching potential gaze sequences or patterns that appear repeatedly in time series analysis mapped into the instruction types throughout our data. therefore, in future studies there is a call for sequence mining of extended gaze data to identify teachers’ attentional patterns and thus validate additional pedagogical intentions underlying gaze. moreover, in the future studies we would like to break down the teachers’ eye contact with students as it was apparent that teachers’ gaze at students' faces did not change in accordance with the maatta et al 112 | f l r types of instruction-giving we were investigating. thus, sequential analysis could be utilized to shed light on student presence which yielded non-significant differences between the types of instructiongiving. we anticipate that eye contact (gaze at student face) should differ regarding teacher gaze sequences during readiness for task and focus on task (introductory vs. collaborative instruction-giving). creating readiness for a task would initiate with eye contact body eye contact gaze chains whereas focus on task would comprise repeated problem-solving tools solution paper gaze cycles. keypoints disentangling the effect of instruction-giving on teacher gaze, we suggest a previously not detected interpretative lens for teacher intentions. teacher presence and the pedagogical intentions are manifested in the teacher's gazes at specific task-related targets. two novel teacher gaze types are suggested: contextualizing gaze for students’ task readiness and collaborative gaze for students’ task focus. the results reveal teachers’ student-centred attention despite instruction type and task-related attention during collaborative instruction-giving. acknowledgments we would like to thank the participant teachers and students as well as dr. enrique garcia moreno-esteva, dr. eeva haataja, dr. anu laine, mrs. jessica salminen-saari, mr. visajaani salonen, and dr. miika toivanen for their contribution in the data collection. references andersen, j. f. (1979). teacher immediacy as a predictor of teaching effectiveness. communication yearbook, 3, 543. https://doi.org/10.1080/23808985.1979.11923782 anderson, n. c., anderson, f., kingstone, a., & bischof, w. f. (2015). a comparison of scanpath comparison methods. behaviour research methods, 47(4), 1377-1392. https://doi.org/10.3758/s13428014-0550-3 baron-cohen, s. (1997). how to build a baby that can read minds: cognitive mechanisms in mindreading. the maladapted mind: classic readings in evolutionary psychopathology, 207-239. berliner, d. c. (2004). describing the behaviour and documenting the accomplishments of expert teachers. bulletin of science, technology & society, 24(3), 200-212. https://doi.org/10.1177/0270467604265535 christophel, d. m. (1990). the relationships among teacher immediacy behaviours, student motivation, and learning. communication education, 39(4), 323-340. https://doi.org/10.1080/03634529009378813 cohen, p. r., & levesque, h. j. (1990). intention is choice with commitment. artificial intelligence, 42(23), 213-261. https://doi.org/10.1016/0004-3702(90)90055-5 cortina, k. s., miller, k. f., mckenzie, r., & epstein, a. (2015). where low and high inference data converge: validation of class assessment of mathematics instruction using mobile eye tracking with expert and novice teachers. international journal of science and mathematics education, 13(2), 389403. https://doi.org/10.1007/s10763-014-9610-5 csibra, g., & gergely, g. (2009). natural pedagogy. trends in cognitive sciences, 13(4), 148-153. https://doi.org/10.1016/j.tics.2009.01.005 maatta et al 113 | f l r elan (version 5.3) [computer software]. (2018). nijmegen: max planck institute for psycholinguistics, the language archive. retrieved from https://archive.mpi.nl/tla/elan ericsson, k. a., & simon, h. a. (1980). verbal reports as data. psychological review, 87(3), 215. https://doi.org/10.1037/0033-295x.87.3.215 farroni, t., csibra, g., simion, f., & johnson, m. h. (2002). eye contact detection in humans from birth. proceedings of the national academy of sciences, 99(14), 96029605. https://doi.org/10.1073/pnas.152159999 foulsham, t., & underwood, g. (2008). what can saliency models predict about eye movements? spatial and sequential aspects of fixations during encoding and recognition. journal of vision, 8(2), 6. https://doi.org/10.1167/8.2.6 frey, n., & fisher, d. (2010). identifying instructional moves during guided learning. the reading teacher, 64(2), 84-95. https://doi.org/10.1598/rt.64.2.1 frymier, a. b. (1994). a model of immediacy in the classroom. communication quarterly, 42(2), 133-144. https://doi.org/10.1080/01463379409369922 gegenfurtner, a., lehtinen, e., & säljö, r. (2011). expertise differences in the comprehension of visualizations: a meta-analysis of eye tracking research in professional domains. educational psychology review, 23(4), 523-552. https://doi.org/10.1007/s10648-011-9174-7 gest, s. d., & rodkin, p. c. (2011). teaching practices and elementary classroom peer ecologies. journal of applied developmental psychology, 32(5), 288-296. https://doi.org/10.1016/j.appdev.2011.02.004 haataja, e., garcia moreno-esteva, e., salonen, v., laine, a., toivanen, m., & hannula, m. s. (2019). teacher's visual attention when scaffolding collaborative mathematical problem solving. teaching and teacher education, 86, 102877. https://doi.org/10.1016/j.tate.2019.102877 haataja, e., toivanen, m., laine, a., & hannula, m. s. (2019). teacher-student eye contact during scaffolding collaborative mathematical problem-solving. lumat: international journal on math, science and technology education, 7(2), 9-26. hattie, j. (2003). teachers make a difference: what is the research evidence? paper presented at the building teacher quality: what does research tell us acer research conference, melbourne, australia. retrieved from http://research.acer.edu.au/research_conference_2003/4/ hendrickx, m. m., mainhard, m. t., boor-klip, h. j., cillessen, a. h., & brekelmans, m. (2016). social dynamics in the classroom: teacher support and conflict and the peer ecology. https://doi.org/10.1016/j.tate.2015.10.004 holmqvist, k., & andersson, r. (2017). eye tracking: a comprehensive guide to methods, paradigms, and measures. lund: eye-tracking research institute. holmqvist, k., nyström, m., andersson, r., dewhurst, r., jarodzka, h., & van de weijer, j. (2011). eye tracking: a comprehensive guide to methods and measures. oup oxford. jarodzka, h., holmqvist, k., & gruber, h. (2017). eye tracking in educational science: theoretical frameworks and research agendas. journal of eye movement research, 10(1):3, 1-18. https://doi.org/10.16910/jemr.10.1.3 jennings, p. a., & greenberg, m. t. (2009). the prosocial classroom: teacher social and emotional competence in relation to student and classroom outcomes. review of educational research, 79(1), 491525. https://doi.org/10.3102/0034654308325693 johnson, h. l., coles, a., & clarke, d. (2017). mathematical tasks and the student: navigating “tensions of intentions” between designers, teachers, and students. zdm, 49(6), 813-822. https://doi.org/10.1007/s11858-017-0894-0 kaakinen, j. k. (2020). what can eye movements tell us about visual perception processes in classroom contexts? commentary on a special issue. educational psychology review, 111. https://doi.org/10.1007/s10648-020-09573-7 kounin, j. s. (1970). discipline and group management in classrooms. new york: holt, rinehart & winston. lester jr, f. k. (1989). the role of metacognition in mathematical problem solving: a study of two grade seven classes. final report. maatta et al 114 | f l r lukander, k., toivanen, m., & puolamäki, k. (2017). inferring intent and action from gaze in naturalistic behavior: a review. international journal of mobile human computer interaction (ijmhci), 9(4), 4157. https://doi.org/10.4018/ijmhci.2017100104 lukander, k., toivanen, m., & puolamäki, k. (2017). probabilistic approach to robust wearable gaze tracking. https://doi.org/10.16910/jemr.10.4.2 mccroskey, j. c., richmond, v. p., sallinen, a., fayer, j. m., & barraclough, r. a. (1995). a cross-cultural and multi-behavioural analysis of the relationship between nonverbal immediacy and teacher evaluation. communication education, 44(4), 281-291. https://doi.org/10.1080/03634529509379019 mccluskey, r., dwyer, j., & sherrod, s. (2017). teacher immediacy and learning mathematics: effects on students with divergent mathematical aptitudes. investigations in mathematics learning, 9(4), 157-170. https://doi.org/10.1080/19477503.2016.1245047 mcintyre, n. a., & foulsham, t. (2018). scanpath analysis of expertise and culture in teacher gaze in realworld classrooms. instructional science, 46(3), 435-455. https://doi.org/10.1007/s11251-017-9445-x mcintyre, n. a., jarodzka, h., & klassen, r. m. (2019). capturing teacher priorities: using real-world eyetracking to investigate expert teacher priorities across two cultures. learning and instruction, 60, 215224. https://doi.org/10.1016/j.learninstruc.2017.12.003 mcintyre, n. a., mainhard, m. t., & klassen, r. m. (2017). are you looking to teach? cultural, temporal and dynamic insights into expert teacher gaze. learning and instruction, 49, 41-53. http://dx.doi.org/10.1016/j.learninstruc.2016.12.005 mcintyre, n. a. (2016). teach at first sight: expert teacher gaze across two cultural settings (doctoral dissertation). university of york, york. mehrabian, a. (1970). a semantic space for nonverbal behaviour. journal of consulting and clinical psychology, 35(2), 248-257. https://doi.org/10.1037/h0030083 mennie, n., hayhoe, m., & sullivan, b. (2007). look-ahead fixations: anticipatory eye movements in natural tasks. experimental brain research, 179(3), 427-442. https://doi.org/10.1007/s00221-006-0804-0 noton, d., & stark, l. (1971). scanpaths in saccadic eye movements while viewing and recognizing patterns. vision research, 11(9), 929-in8. https://doi.org/10.1016/0042-6989(71)90213-6 orquin, j. l., & holmqvist, k. (2018). threats to the validity of eye-movement research in psychology. behaviour research methods, 50(4), 1645-1656. https://doi.org/10.3758/s13428-017-0998-z polya, g. (1948). how to solve it. princeton, nj. princeton university press. pouta, m., lehtinen, e., & palonen, t. (2020). student teachers’ and experienced teachers’ professional vision of students’ understanding of the rational number concept. educational psychology review, 120. https://doi.org/10.1007/s10648-020-09536-y richmond, v. p. (1990). communication in the classroom: power and motivation. communication education, 39(3), 181-195. https://doi.org/10.1080/03634529009378801 rothkopf, c. a., ballard, d. h., & hayhoe, m. m. (2007). task and context determine where you look. journal of vision, 7(14), 16-16. https://doi.org/10.1167/7.14.16 schurter, w. a. (2002). comprehension monitoring: an aid to mathematical problem solving. journal of developmental education, 26(2), 22. shteynberg, g. (2015). shared attention. perspectives on psychological science, 10(5), 579-590. https://doi.org/10.1177/1745691615589104 shteynberg, g., hirsh, j. b., apfelbaum, e. p., larsen, j. t., galinsky, a. d., & roese, n. j. (2014). feeling more together: group attention intensifies emotion. emotion (washington, d.c.), 14(6), 1102-1114. https://doi.org/10.1037/a0037697 simon, h. a. (1992). what is an “explanation” of behaviour? psychological science, 3(3), 150-161. https://doi.org/10.1111/j.1467-9280.1992.tb00017.x stuermer, k., seidel, t., & holzberger, d. (2016). intra-individual differences in developing professional vision: preservice teachers’ changes in the course of an innovative teacher education program. instructional science, 44(3), 293-309. https://doi.org/10.1007/s11251-016-9373-1 tatler, b. w., kirtley, c., macdonald, r. g., mitchell, k. m., & savage, s. w. (2014). the active eye: perspectives on eye movement research. in current trends in eye tracking research (pp. 3-16). springer, cham. https://doi.org/10.1007/978-3-319-02868-2_1 maatta et al 115 | f l r tatler, b. w., & land, m. f. (2016). everyday visual attention. handbook of attention, 391-422. tomasello, m., & carpenter, m. (2007). shared intentionality. developmental science, 10(1), 121-125. https://doi.org/10.1111/j.1467-7687.2007.00573.x tomasello, m., carpenter, m., call, j., behne, t., & moll, h. (2005). understanding and sharing intentions: the origins of cultural cognition. behavioural and brain sciences, 28(5), 675-735. https://doi.org/10.1017/s0140525x05000129 toom, a. (2006). tacit pedagogical knowing at the core of teacher's professionality. van de pol, j., volman, m., & beishuizen, j. (2010). scaffolding in teacher-student interaction: a decade of research. educational psychology review, 22(3), 271-296. https://doi.org/10.1007/s10648-010-9127-6 van den bogert, n., van bruggen, j., kostons, d., & jochems, w. (2014). first steps into understanding teachers' visual perception of classroom events. teaching and teacher education, 37, 208-216. https://doi.org/10.1016/j.tate.2013.09.001 van der want, a. c., den brok, p., beijaard, d., brekelmans, m., claessens, l. c., & pennings, h. j. (2015). teachers' interpersonal role identity. scandinavian journal of educational research, 59(4), 424-442. https://doi.org/10.1080/00313831.2014.904428 van manen, m. (1991). reflectivity and the pedagogical moment: the normativity of pedagogical thinking and acting. journal of curriculum studies, 23(6), 507-536. https://doi.org/10.1080/0022027910230602 van tartwijk, j., brekelmans, m., wubbels, t., fisher, d. l., & fraser, b. j. (1998). students’ perceptions of teacher interpersonal style: the front of the classroom as the teacher’s stage. teaching and teacher education, 14(6), 607-617. https://doi.org/10.1016/s0742-051x(98)00011-0 witt, p. l., wheeless, l. r., & allen, m. (2004). a meta-analytical review of the relationship between teacher immediacy and student learning. communication monographs, 71(2), 184-207. https://doi.org/10.1080/036452042000228054 wolff, c. e., jarodzka, h., van den bogert, n., & boshuizen, h. p. (2016). teacher vision: expert and novice teachers’ perception of problematic classroom management scenes. instructional science, 44(3), 243265. https://doi.org/10.1007/s11251-016-9367-z wolff, c. e., jarodzka, h., & boshuizen, h. p. (2017). see and tell: differences between expert and novice teachers’ interpretations of problematic classroom management events. teaching and teacher education, 66, 295-308. https://doi.org/10.1016/j.tate.2017.04.015 frontline learning research vol. 12 no. 2 (2024) 27 50 issn 2295-3159 corresponding author: siltavuorenpenger 5b, pl9, 00014 helsingin yliopisto iina.hyyppa@helsinki.fi doi: https://doi.org/10.14786/flr.v12i2.1301 fostering students’ systems thinking through futures education iina hyyppä, tapio rasa & antti laherto university of helsinki, finland article received 5th june 2023 / article revised 15 march 2024 / accepted 10 may / available online 14 june abstract in an era of worsening environmental crises, students may not perceive themselves as able to impact and change the inevitably upcoming futures. accordingly, a common goal of educational systems has been to develop students' agency beliefs and sensemaking in a complex world. simultaneously, students are facing unprecedented levels of future anxiety, and educational institutions undervalue the importance of futures thinking. to take on a constructive approach on futures thinking, we examine how students’ systems thinking skills develop during a futures education course in which they write their own visions of a hopeful future. by looking at the thematic spheres of society, nature, and technology, we analyse how students develop systemic understandings of the complex system that is the context of the study: the city of the future. the study examines how students’ written future visions develop throughout the course, and how those changes indicate development in systems thinking. the results show that the futures education course allowed students to improve their understandings of the interconnectedness of the topics they raised, fostering more complete and active understandings of the futures, here shown through the multidimensional development of systems thinking. students developed a deeper understanding of the interrelationships of society, nature, and technology, and advanced understandings of the pathways to change and the actions needed to achieve their futures. keywords: systems thinking; future visions; futures thinking; futures education; secondary education mailto:iina.hyyppa@helsinki.fi https://doi.org/10.14786/flr.v12i2 hyyppä, rasa, & laherto 28 | f l r 1. introduction in a rapidly changing world, thinking about the future is increasingly important. anticipation, forecasting, and scenario building, however, are challenged by uncertainty and complexity: from meteorological to political or social futures, the evolution of complex systems can be predicted only to a limited extent. in an era of environmental crisis, longer future predictions are often problematic to the point of pessimism (united nations, 2023). although such future predictions are highly researched and calculated scenarios, their intentions in evoking awareness and actions can backfire due to overwhelming sensations of helplessness and impossibility (hickman et al., 2021). moving between levels of global challenges and individual actions can be difficult as simple causalities are governed by systemic ones. addressing the need to help students orient towards uncertain and complex futures, the field of futures education aims to bring futures thinking into classrooms today (see e.g. page, 1996). to build these capacities, futures education, as discussed by hicks (2003), entails a conceptual framework of education about the future, focusing on the core skills of visioning, thinking, and systemizing. futures education may e.g. evoke thoughts of hopeful futures and work on understanding the ways in which to achieve them (rasa et al., 2023), and to envision the different complex components that create and impact change (ahvenharju et al., 2018; facer, 2011; hicks, 2003), skills deemed crucial for sustainable development of our future (lotz-sisitka et al., 2015). as eleanor roosevelt iterated in 1978, "the future belongs to those who believe in the beauty of their dreams”, and although seemingly simplistic at first, dreaming of a hopeful future remains a concept very estranged to most students (hickman et al., 2021). educational curricula for secondary education across the globe speak of the future and of teaching students the necessary skills to strive in a future society. in highlighting the importance of future skills, they fail to mention concrete ways of incorporating futures into practice (finnish national agency for education, 2019; secretary-general of the oecd, 2018; see also poli, 2021). history is studied partially as a means of understanding past chains of events and learning from previous successes and mistakes, which brings to question the matter of why futures are not taught as a way of putting historical, scientific, and social knowledge into a context in which it can impact the world of tomorrow. strongly tied in with futures learning, systems thinking involves an ability to conceptualize and frame topics as a part of their larger concepts, whilst understanding both details and wholes (whitcomb et al., 2020). in contemporary competency frameworks (unesco, 2017), both futures thinking and systems thinking are commonly stated as key competences for future citizens and professionals. however, the inherent connections between these competence areas have not yet been sufficiently studied. to evaluate a perspective on futures competencies and their applicability into educational settings, this paper studies a perspective in which students are given the possibility to create their futures. the paper studies the effects of futures education and future visions in developing students’ systems thinking. our focus is not on assessing students’ overall systems thinking competency, but rather on exploring how different aspects of students’ systems thinking are apparent in students’ future visions, and how those aspects develop in a futures education course – i.e. the potential of futures education in developing systems thinking. centred around the theme of “helsinki in 2050”, students vision the future of the capital city of finland. students’ future visions, and the factors that construct their future city, are used to map out how students perceive the future, and how those perceptions can develop through futures education. furthermore, this paper examines the extent to which the implementation of a futures education course impacts students’ abilities to conceptualize the relationships between, and develop a systemic understanding of, societal organizations, natural biospheres, and built environments. 2. theoretical background 2.1. defining futures thinking unlike traditional school subjects, futures are not pre-existent concepts that can be taught through textbooks and ready materials, but rather they act as frameworks to develop thinking and awareness. local, national, and global levels of initiative have been taken in creating multitudes of guidelines hyyppä, rasa, & laherto 29 | f l r with the aims of promoting sustainable futures, for example through increased and improved sustainability and futures education. futures thinking is promoted through key future competencies, evoking mindset and thinking model changes, and creating awareness for and understanding of futures sustainability science and research. (see city of helsinki, 2018; secretary-general of the oecd, 2018; unesco, 2017; united nations, 2023) we presently live in a society created by our past, creating our future. danskernes (1996, as cited by remes & rubin, 1996) defines three ideological orientations that divide society’s thoughts into chronological categories based on the past, present, and future. individuals’ social and ideological surroundings shape their concepts of time, and the values that arise from those understandings of time. the various roles of the future, and how we think about it, have in turn been extensively studied in the field of futures studies. roy amara (1981, as cited by aalto et al., 2022) defines a set of three core requisites to explore the field of futures studies. first, amara (1981) defines states that the future cannot be predicted, but only imagined and speculated. secondly, he asserts that futures are not predefined, and thirdly that societies and individuals can influence the future through their thoughts and actions. amara (1981) emphasizes that the field of futures studies cannot be excluded from the values, principals, and ethics of the researchers in control of the thought processes. likewise, valciukas and bell (2003) state the importance and the development of the field of future studies, yet found that until the 1990’s, little importance was paid to its philosophical framework. valciukas and bell (2003) assert that the future cannot directly be studied, due to its non-existent nature as compared with the present and/or past. they also found that the future exists within our present-day intentions and can only be studied through matters and realities which could impact the future. perhaps due to these abstractions, while education is by nature future-oriented, this relationship is often implicit (poli, 2021). a notable departure from this can be seen in sustainability education: for example, greencomp, the european sustainability competence framework, highlights “envisioning sustainable futures” as a key sustainability competence (bianchi et al., 2022; see also laherto et al., 2023). in greencomp, this “competence area” involves systemic and critical thinking, futures literacy, exploratory thinking, problem framing and political agency. on the other hand, sustainability is a central concern in futures education literature (see e.g. häggström & schmidt, 2021). as rasa (2023, p. 55) argues, “sustainability and futures are, in a sense, two sides of the same coin”. a broad consensus exists that such thinking skills are needed: the education for sustainable development goals (unesco, 2017) report emphasizes the need for systems thinking, anticipation, normative reflection, collaboration, critical thinking, self-awareness, and problem solving. combining the above competencies, unesco (2017) views the thinking skills needed for futures as ones that understand complexity, accept uncertainty, assess consequentiality, question normativity, and work collaboratively with aims of creating better solutions for sustainable education. futures thinking and the general field of futures studies are also strongly correlated with the conceptual framework of future consciousness, coined by sande (1972). the six dimensions of future that are acknowledged and used among members of society, as defined by sande (1972), revolve around time frames, evaluating how far into the future individuals are able and willing to see and plan. they also go on to evaluate the level of optimism and topics of interest. likewise, sande’s (1972) study evaluated a sense of influence, agency, and power in having the ability to change the future, along with an evaluation of expectations in how individuals truly consider the future to look. the last dimension of sande’s (1972) study focuses on values that individuals indicate within their desired futures. revising sande’s (1972) framework, ahvenharju et al. (2018) compile a dimensional framework of five core aspects of futures consciousness and define the components as time perspective, agency beliefs, openness to alternatives, concern for others, and systems perceptions. to narrow the scope of this study, we focus mainly on systems perceptions. related to interactions, decisions and their consequences, and complexity, this fifth dimension of the broader concept of futures consciousness is another conceptual interface between futures, complexity, and agency. this is mirrored in the similar hyyppä, rasa, & laherto 30 | f l r concept of futures literacy, studied as reflexivity of future attitudes, studies, and pathways for action (mangnus et al., 2021). although the explicit focus of this paper lies on systems perceptions in future visions, agency beliefs are considered as an over-arching motivator for the need for futures thinking in education. while agency is here not being considered as an active goal, the concept of agency in terms of ahvenharju et al.’s (2018) definition of agency beliefs importantly motivates future-oriented educational approaches (laherto et al., 2023; rasa et al., 2023). as ahvenharju et al. (2018) define it, agency in futures thinking is crucial for one’s “sense of being able to influence how the future will unfold”. thus, this paper does not explore how participants make concrete efforts to achieve their futures, but rather how students perceive their impact and role on futures and the systems within them. 2.2. developing futures education unesco (2023) states its core mission as “to build peace, eradicate poverty and drive sustainable development”. all centred around future developments, the missions define the core purpose of education as an all-round tool to enable equality and sustainability. how are students to change the future, if schooling practices overemphasize historical and current knowledge and undervalue providing students with the necessary tools to imagine, develop, and create a world and society in which all beings can live sustainably and equitably? futures education provides tools and grounding for promoting students’ thinking skills towards deepened consideration of several possible futures as well as towards the impact of different parties’ agency. although educational curricula set a national guideline for the integration of future related topics into teaching (see e.g. finnish national agency for education, 2019), a multidisciplinary concept such as futures education is only concretized within teaching and learning customs. futures pedagogy can be approached from multiple perspectives by focusing on future visions, dreaming, or practical scientific experimentation and its applicability within future scenarios, to name a few. in evaluating the need for futures-oriented education, fitch and svengalis (1995, as cited by hicks & holden, 1995, p.3) state that: “by adding a future dimension to the learning process, we help to provide direction, purpose, and greater meaning to whatever is being studied. by integrating past, present, and future we act to strengthen a neglected link in the learning process.” futures education has previously also been approached from the perspective of future-scaffolding skills. levrini et al. (2021) study aspects of futures-oriented education centred around organization of present knowledge, imagination of futures, and dynamic and conscious movement within the futures related space and time continuums (tasquier et al., 2019). within the same international research project on teaching and developing students’ futures thinking skills, bol et al. (2023) create guidelines on for futures thinking, involving the constancy of change, reflections on the meanings of time, accepting uncertainty and ambiguity, stimulating long-term and systems thinking and curiosity, promoting imagination and plural, open thinking, as well as consciously understanding the impacts of choices and engaging with the thought-up futures. rasa et al. (2023) go on to explore futures thinking skills in education through the lens of technology and agency, contrasting static and transformational futures and examining students’ ways of complexifying future societal change. in this paper, futures are discussed in plural, following the typical convention in futures studies. the reasoning for this is to emphasise thinking of futures as potential worlds, not a singular, predetermined world to be forecasted; after all, this is the basis of futures literacy and futures consciousness. the singular form is used when speaking of one possible future scenario or vision. hyyppä, rasa, & laherto 31 | f l r 2.3. understanding systems thinking by understanding the interrelatedness of our systems at hand, systems thinking aims to understand full entities without breaking them down into small parts for separate inspection and analysis (gharajedaghi, 2011). systems thinking is to be regarded as a process rather than a result within the continuums of chaos to organization, and of simpleness to complexity. as a core component of ahvenharju et al.’s (2018) futures consciousness framework, system thinking is also depicted as the first central competency required from educational systems and students in order to work towards more sustainable futures. the learning of systems thinking involves the recognition and understanding of intricate relationships, along with the analysis of complexity within systems. likewise, it involves understanding how systems and matters are embedded within themselves and outside factors in different ways. as such, the ability to cope with uncertainty and complexity of the cognitive steps required to develop one’s systems thinking skills are factors apparent in different aspects of education, as implicit learning goals, guidelines, or practices. (unesco, 2017) similarly, shaked and schechter (2019) explore systems thinking in education as the understanding and improvement of complex systems, and their examination as wholes. they highlight the crucial element of understanding interrelatedness of matters as a key focus point of systems thinking. in connecting systems thinking to educational contexts, hofman-bergholm (2018) explores the extent to which the relationships of systems thinking and sustainability education can mutually create synergy in learning, emphasizing the common roles of transdisciplinary, value discussions and action and agency as key skills fundamental for meaningful learning. this allows for recognition of interconnections, identification of feedback, understanding of dynamic behaviour, using conceptual models, testing policies, and acquiring knowledge on root causes, among others, which can all be extracted as skills for broader learning scenarios. 2.4 conceptualizing the city as a complex system thinking about the future naturally involves complexity: firstly, because the world is complex, and secondly because many of current societal challenges relate to complex systems from the climate and ecosystems to cities and urbanisation. conversely, while sustainability transitions involve assessment of decade-scale projections of climate change, city planning involves future-orientedness on various levels. cities require planning and building, anticipation of trends (e.g. immigration, household size) and reacting to newly arising issues in technology, work, transportation and so on (toivonen et al., 2021); thus, complexities naturally arise when discussing images of desirable future cities (see höjer et al., 2011). the teaching module that is the context for the present study tapped into these fruitful connections by bringing students to imagine the future of their city. some educational approaches aiming to build futures literacy in the context of the city have been reported (toivonen et al., 2021), with results supporting the claim that thinking about futures may promote empowerment. the city is a complex system, with a rich body of literature devoted to understanding it (for an introduction, see e.g. moroni & cozzolino, 2019), and additional complexities emerge as societies and cities attempt to undergo sustainability transitions (wiek et al., 2006). the city is, at the minimum, both a place and a locus of numerous (inter)actions (moroni & cozzolino, 2019). analysed further, the systemic nature of the city can be seen as consisting of multiple subsystems. such a division of a system into subsystems is dependent on perspective and context. in this study, the city is seen through students’ eyes (as opposed to e.g. city planning professionals) in the context of imagining sustainable futures. in the context of sustainability, a common and useful (even if problematic) heuristic is to separate human and “natural systems”, or humans and technology (see e.g. ahlqvist & rhisiart, 2015); in this paper, the city is foremostly seen as an entanglement of natural systems, built and technological environments, and humans and social activity (see research methods and processes). similar perspectives of the systemic nature of the city in relation to futures thinking are considered by kivistö (1985). considering futures thinking from an urban developmental perspective, kivistö hyyppä, rasa, & laherto 32 | f l r (1985) names eight core components for analysis of today and tomorrow: society, needs, urban infrastructures, natural resources, engineering and technology, economy, livelihood and services, and other factors. this approach somewhat differs from e.g. seeing the city as like an organism (bettencourt et al., 2007). clearly, systems can be modelled in multiple ways. nevertheless, the city and the body are both pedagogically interesting examples of systems; see (tripto et al., 2017) for a study that, in the context of the body, examines the development of students’ systems thinking, partly mirroring our urban approach. 3. aims of the study this study aims to examine the effectiveness of a future learning course on the development of students’ futures thinking and the systems perception thereof. by examining students’ development of systems thinking during a continuous elaboration of future visions, we evaluate how students’ perceptions of the city as a system manifest complexity in futures thinking. as there exist few studies that analyse students’ futures or systems thinking developing over time, we focus on this specific issue. namely, we describe patterns in students’ thinking as they immerse in more extensive futures thinking, supported by a futures education course. with this somewhat longitudinal approach, the study aims to perceive how constructing future visions can be promoted by providing space and support for that end. the research question of this paper aims to explore an analytical approach on systems thinking and its development: q1: how is students’ systems thinking supported by a course on visioning the city of the future? 4. context of the study the data of this study stems from the future-oriented science education to enhance responsibility and engagement in the society of acceleration and uncertainty (fedora) research project at the university of helsinki (https://fedora-project.eu). the fedora-module, developed in collaboration with helsinki school of natural sciences (an upper secondary school with a science focus), engaged students in creating a sustainable future for the city of helsinki, finland. this experimental science course, titled “my city of the future”, was attended by 11 upper secondary school students aged between 16-18 years. the course consisted of 7 lessons over the course of 2 months. the course began with an introduction into futures thinking: how it often fails to predict the future, yet one can improve and systematize one’s visions, for instance, by distinguishing between thinking about possible, probable, and desirable futures. over the course, students worked on their visions for helsinki in the year 2050, writing evocative, hopeful future descriptions in 4 small groups. similar approaches of gathering students’ images of the future have been explored by angheloiu et al. (2020) and rasa and laherto (2022). the texts were continually challenged by the teachers as well as three invited consulting experts (smart city anthropology, values in futures thinking, energy and sustainability transitions), who posed unscripted questions based on their field of expertise regarding microand macro-level correlations and consequences, challenged specific decisions and actions and proposed alternative ways of thinking about and constructing future cities. over the process, students wrote 4 versions of their future visions. the students also built timelines between today and their vision, mapping central actions to take to reach their desired future, paying special attention to systemic perspectives and the role of technology, science, and built environments in creating sustainability (e.g. energy production) and shaping the city of the future (e.g. new technologies). pedagogical futures education methods, such as visioning and backcasting (see e.g. laherto & rasa, 2022; rasa et al., 2022) were used with aims of promoting future-orientedness, process thinking, and understandings of causalities and (un)certainties. then, the students familiarized themselves with the publicly available “carbon neutral helsinki 2035 action plan [cnh]” (city of helsinki, 2018), guided by a pedagogical workshop on analysing https://fedora-project.eu/ hyyppä, rasa, & laherto 33 | f l r values and assumptions in future scenarios. after this they met with one of cnh’s authors to discuss the rationale for the environmental policies of the city of helsinki. during these activities, students compared their own thinking with official policies and contrasted the actions they wished to see taken with those currently planned or executed. finally, guided by the teachers of the course, the students collected their written visions of helsinki in 2050 into a small pamphlet. the course ended with a discussion panel between the students, the head of the city of helsinki’s climate team, and other students from the school in the audience during which the finalized pamphlet was handed over to the city. while this course was conducted in collaboration between an upper secondary school and the fedora research project, similar futures education courses can be conducted using the above presented methods and approaches in a non-research setting. 5. research methods and processes to approach the research question of this paper, systems thinking was evaluated from a developmental perspective, focusing on how this futures education course influenced the prevalence of systems thinking in students’ written future visions. the dataset examined in this analysis consisted of written future visions from four student groups. four versions of each future vision were used to analyse the revisions occurring from the first version (v1) to the final version (v4), evaluating the content that was added to or developed from v1 in reaction to the activities of the course. the analysed excerpts thus show the alterations and additions from v1 to v4. no group removed content from v1 and no overlapping content apparent in both versions was analysed if it showed no revision between versions. as a first step, all additions to the content of the written future visions were extracted. from these changes, units interpreted as relevant to systems thinking based on the literature above presented were analysed. as an example, consider the following two quotations. the first one, taken from v1 of group 4, envisions the presence of new technology in personal use: v1: “i open my computer and quantum computer, where i develop a graphic operating software.” in v4, the effect of the technologies mentioned in v1 upon personal work life and workloads in general is explored, combining them with the aspects of ease of access and communication through developed technologies. this is perceived as systems thinking as the group further analyses the benefits of the technologies mentioned in v1 and how they impact human life. systems thinking is shown through an increased understanding of the interconnectedness between humans and technology: added to v4: “i did all necessary manual work on the software today, as ai does most of the code. simultaneously i was invited for a virtual meeting … [which] can be accessed through ar or vr glasses.” such revisions between the versions formed the units of analysis for the following qualitative content analysis. the revisions were coded employing the method of inductive thematic analysis (braun & clarke, 2006). initial codes were formulated based on the themes apparent from the future visions, such as digitalization, employment, wellbeing, nature, which arose in the future visions verbatim. the initial codes were then revised, restructured, and formulated into inductively apparent thematic spheres, categorizing each initial code into one or more of the thematic spheres. a review of the inductive themes was followed by defining and naming the three main thematic spheres from which the students approached the city of the future: social sphere; technological sphere; and natural sphere. a comparable exploration of social, natural, and technological themes has been previously stated as a learning goal for the fedora -module, as well as analysed in the context of sustainable development competencies and urban geography (moss et al., 2021; unesco, 2017). as noted earlier (see 2.4), these three spheres are one way to view the city as a system; these spheres were selected for the analysis due to their clear presence in the data. hyyppä, rasa, & laherto 34 | f l r the table below indicates and justifies the thematic spheres formulated through the inductive analysis of the data. they portray the types of excerpts from the visions that went into each of the three thematic spheres to visualize the meanings of each thematic sphere. the justifications provide insights into the analytical processes within this paper. table 1 examples of excerpts and justifications of thematic spheres thematic spheres example of excerpts from the future visions justifications of inclusion social sphere: society and human organizations “algorithms and hundreds of employees spread awareness of municipal matters, and locals act as special experts of their own areas.” “the idea of encouraging people to recycle using external motivators has slowly made its way to finland. a similar concept was used in china in the 2020’s, but instead of rewards, people were forced to recycle under a threat of being fined. finland is trying to turn this idea into a reward.” excerpts include topics of agency and social life, looking into who has a say in causing and making change. matters including the city and its decisions are considered in their humane approach on change through city councils and citizens. social constructs and structures, as well as personal and social agency are included, portraying the benevolence and will for change of the people. preliminary phase codes included i.e. employment, activism, politics, and citizens. technological sphere: technology, science, and built environments “initially technology was believed to fix everything: climate change, environmental crises, political unrest, criminality, marginalization, drug use etc. countless hours and resources were invested in its development, but technology didn’t magically fix everything…” “technology has been deployed for peoples’ benefit and aid, it is used by all, and it is not expensive. there are new ways to manufacture electronic devices with scarce natural resources. ai has been improved and it can be seen in everyday life.” excerpts include topics related to scientific and technological development, such as transportation, battery life and energy production. emissions and sustainable energy production methods, alongside consideration of the city through its infrastructures and non-natural environments, such as roads and power plants, are also explored. ponderings on the ultimate purpose of technological and scientific development are considered. preliminary phase codes included i.e. digitalization, innovation, infrastructure, and artificial intelligence. natural sphere: natural world and biosphere “a nearby apartment building’s dark, wooden walls and extensively covering wall solar panels flicker some hundred meters away. the low sun rays fill the surrounding parks and light up the lower buildings’ vernal green roofs…” “the most extreme nature advocates have of course been against this, but disadvantages have been compensated for by giving space for nature inside the city limits.” excerpts show themes of nature and climate, painting a picture of how the natural world looks. matters consider change within the biosphere, including topics related to wildlife and sustainability, looking at an environmentalist view of the city. the role of true nature in a modern environment where nature is not the only factor contributing to the state and wellbeing of the physical surroundings is also explored. preliminary phase codes included i.e. green sustainability, nature, and wildlife. hyyppä, rasa, & laherto 35 | f l r the excerpts in table 1 are justified as belonging to one thematic sphere, yet some excerpts clearly contain thematical matters from other thematic spheres too. the intention of this paper is to look exactly into those overlapping areas where two or more thematic spheres meet (e.g. by combining social and natural themes). hence, systems thinking is here analysed from a perspective of complexity between the thematic spheres as a means of evaluating systems thinking through the ability to conceptualize topics as broad entities relating to and impacting other matters. as such, systems thinking is perceived through the overlaps of the three spheres (social, technological, and natural) with the main focus on the central overlap of all three thematic spheres. the results are thus explored first through the overlap of all three categories, followed by the socio-technological overlap, then the socio-natural overlap, and lastly the techno-natural overlap. 6. results the students’ future visions indicated deep reevaluation of future-related values, structures, and changes through the revisions between v1 and v4. the future visions included general analysis of multiple different spheres overlapping together, as well as more indepth examination of the dimensions of singular topics and their relations to the other thematic spheres. the revisions made between v1 and v4 can be represented as venn diagrams. note that figures 1-5 are not to scale (nts). the venn diagram in figure 1, representing the city in terms of the three thematic spheres (see 5. research methods and processes), portrays the revisions from v1 to v4 of the groups’ future visions. each number in the figure represents the number of text excerpts from the future visions that discuss topics relating to that thematic sphere (or of the overlap thereof). thus, fig. 1 gives an overview of the general revisions within all groups’ future stories. overall, the revisions of the stories focused largely on a more systematic and full understanding of the different spheres apparent within the students’ future visions. the most frequently discussed thematic sphere, society and human organizations, showed the social nature of the future visions. the societal angle indicated an understanding of the relevance of the human and social components in all dimensions of future developments. in the following, the development of systems thinking in students’ future visions is manifested in increasing overlap between the three main thematic spheres. the results are presented group by group to illustrate the qualitative differences between each group’s revision process. the analysis begins with the full overlap of all three thematic spheres, and then moves on to evaluating the overlaps of two thematic spheres. figure 1. revisions from v1 to v4, all groups (nts) hyyppä, rasa, & laherto 36 | f l r 6.1. group 1 – an environmentalist future (most revisions) group 1’s future vision, titled “green helsinki”, described a sustainable future built around values of environmentalism, social and natural wellbeing, as well as social awareness. the vision explored the human aspects of change and its uncertain nature by considering the agency needed to create change, and the causalities thereof. the revisions of the future vision during the course entailed increased complexity among the overlap of the thematic spheres, as can be seen from the numbers in figure 2. as fig. 2 displays, the most revisions were made in the overlaps of the social sphere, with 23 excerpts. few excerpts were considered only within one singular sphere, showing broad development of systems thinking when comparing v1 to v4. the overall future developments of “green helsinki” remain rather conservative, especially from the technological and scientific perspective, the downsides of which are explored alongside their possibilities. of all groups, this vision most emphasized the need for individual and communal agency, but nonetheless agency is discussed passively, and its drivers remain unexplored. 6.1.1. connections between all 3 themes “green helsinki” shows a compilation of spheres and topics all widely relevant to society and nature, as well as technological and scientific innovation. the progression of systems thinking between v1 and v4 is apparent as seen in the overlaps of the three thematic spheres. whilev1 merely speaks of sustainable energies, v4 explores a multitude of ways of sustainable energy production and considers both their benefits and challenges. v4 analyses the complexity of energy production choices and recognizes the possibility of alternative solutions, seeing energy production as a greater entity and showing possibility of its continued development in relation to its use by society and dependency on renewable natural resources: v4: “the city of helsinki has installed windmills in the windiest places around helsinki. the electricity they produce is used to maintain beneficial services for the city of helsinki and other citizens. … on the other side, i see an apartment building, which as its walls covered by solar panel material. this is still not very common due to its expensive price and need for a communal decision from all the building’s inhabitants. … these solar panels are only being taken into use in the sunniest areas, but i have heard that they should be becoming more common in the near future.” likewise, in v1, no topic is considered from all three spheres. v4 includes reasonings for revisions from v1 and considers the implications of these changes on personal life. for example, v4 attributes decreased environmental footprints to scientific development. the addition indicates developed understanding of the differences and similarities between one’s own needs and those of the community. the same topic is shortly considered from an energy and material efficiency perspective, implicitly showing their need in scientific development. this connection of all three spheres is apparent from: v1: “apartments do not have their own laundry equipment, but apartment complexes have a communal laundry room for everyone’s use.” continued in v4: “the transition to communal laundry rooms was made to save energy and raw materials. people no longer need to buy their own washing machines.” figure 2. revisions from v1 to v4, group 1 (nts) hyyppä, rasa, & laherto 37 | f l r 6.1.2. socio-technological overlap the scientific and technological developments prevalent within the story intertwine with everyday life, serving to ease routine tasks while increasing their sustainability. the relationship between technology and everyday life is emphasized in v4; while v1 merely mentions a technological detail, v4 explicitly names the connection between technology and routine tasks impacting society’s everyday life, as can be seen below: v1: “when i get my laundry in the machine, i check the contents of my refrigerator from my phone…” continued in v4: “linking mobile phones to kitchen appliances has eased everyday life. it is possible to check the fridge or the laundry room’s washing machine without having to go there.” similarly, v4 explores the effects of battery technology development on society’s access to sustainable and functional transportation methods. the following excerpt indicates deeper conceptualization of the impact of scientific development on society. the excerpt discusses the technological sphere from a wide array of perspectives to grasp its fundamental function in being a social necessity: v4: “the battery life of electric cars has been prolonged through developed technology. nowadays, it is possible to drive the same distance with an electric car as with a tankful of gasoline. the price of gasoline has increased drastically, and a normal person can no longer afford to drive a car with a combustion engine…” 6.1.3. socio-natural overlap the interactions of society and human organizations with the environment are addressed through, for example, the future city promoting sustainability through increased recycling. emphasizing the role of governance, v4 imagines society with external motivators inviting each member of society to do their part in improving the environment, examining the role of social constructs and regulations in promoting a cleaner environment: v1: “due to the development of recycling, there is almost no more mixed waste being produced. there are also more recycling points on the streets, so garbage must be carried home. the increased number of recycling points and the instruction of their use has made the environment cleaner as people know how to recycle.” continued in v4: “the idea of encouraging people to recycle using external motivators has slowly made its way … to finland. a similar concept was used in china in the 2020’s, but instead of rewards, people were forced to recycle under a threat of being fined. finland is trying to turn this idea into a reward.” 6.1.4. techno-natural overlap “green helsinki” combines the spheres of built and natural environments by discussing issues such as sustainable energy productions and green infrastructures. mostly, however, the development of the relationship between nature and science can be viewed through the comparison of the group’s vision to the city’s carbon neutral helsinki program (see next excerpt). the students’ story implies that scientific development is necessary for slowing climate change and developing sustainable energies, which is linked with explicit goals on improving nature’s condition. additionally, in speaking of environmentalism and science, the excerpt recognizes the limits of the group’s future vision in not being based on statistics, touching upon a perspective of the non-scientific nature of futures: v4: “we found many similarities in the presumptions and solutions between the cnh goals and our future. both assumed the importance of slowing down climate change and normalizing green energy. … however, … we did not base it on statistics or precise predictions.” hyyppä, rasa, & laherto 38 | f l r 6.2. group 2 – a conservative, simple future (least revisions) group 2’s future vision, “tomorrow in helsinki”, painted a picture of an environmentalist, green future centred around developments and simplifications of everyday life. the future vision maintained a conservative stance on employing largely pre-existent technologies yet normalizing their use. as can be seen from figure 3, the story was revised quite marginally between the two versions. most revisions related to social and/or natural spheres, which both had 5 revised excerpts. the values of the vision revolved around societal wellbeing, with changes and technologies employed to simplify human life. even though the vision discusses the overcoming of the modern-day digital craze, it contradicts itself in continuing to incorporate many new technologies into the future. while some paths of change are explored, the systems perception and agency behind change remains unexplored in this future vision. 6.2.1. overlap of all 3 spheres as with group 1, group 2’s initial story (v1) does not combine all three spheres, which are later developed and visible in v4. the example below involves systems thinking in combining issues of infrastructure, nature and society, and even considers the challenging nature of city development. by recognizing the change needed in habitational structures to accommodate the green future, the story shows a deeper analysis of the change needed to achieve the future set out in v1. v1 generally discussed buildings becoming larger, but the topic below was added newly to v4. the recognition of the imperfect characteristics of change and the debate of their own future vision shows critical thinking towards own future ideologies, and also brings to light the importance of agency in achieving a communal resolution: v4: “even though space constraints have meant that buildings have been built even bigger and taller, housing complexes have become singular massive towers surrounded by large land plots. more space has been given to nature and for people to breathe between the walls. this restructuring, along with the ever-continuing urban migration, has expanded the city limits past their precedent lines. the most extreme climate activists have of course been against this, but the harm has been compensated for by giving more space for nature inside the city.” the relation between the implicit agency of the city and its impact on the environment is one that combines society with its surrounding nature, looking at the relationship from a perspective of built environments. agency is in this vision explored merely from a passive perspective in which the city makes changes, yet the drivers or contributing factors behind change are not explored. the following excerpt recognizes the underlying environmental values within future society, exploring the greater value of nature against societal infrastructures within municipal governance. the excerpt in v4 adds a deepened explanation of the built, physical structures of society’s and nature’s cohabitation: v1: “large municipal plans have paid more attention to green routes, forcing highways and residential areas to make room for forest networks.” continued in v4: “the entire city structure works around this ‘cell network’. the city is divided into [large] residential areas that are surrounded by forest zones of at least a hundred meters.” figure 3. revisions from v1 to v4, group 2 (nts) hyyppä, rasa, & laherto 39 | f l r 6.2.2. socio-technological overlap the technological developments in the story were mostly apparent already from v1, which depicted the use of technology as an aid for everyday life activities and routines. v4 does, however, also portray insights into the development of society’s relationship towards technology. what is considered evident or obvious change in v1 is further explored in v4, and the story concludes that technology can have disadvantages alongside its benefits. here, v4 also considers the value of social equity in allowing all members of society the same access to and understanding of the technologies required to keep up with environmental progress, which can be seen from: v4: “even though everything can be done on the internet nowadays, society has been able to overcome the digital craze of the 2020’s and find a golden middle ground where everyone is given the possibility to be a member of society without requiring digital wizardry. technology is a fabulous helper, but a bad host, and … we also took some time to realize this”. although technology and science are highly prevalent within both versions of the vision, it states that society is distancing itself from digital dependency and working to find a balance between maximal benefit and maximal accessibility and ease of use. technology, science, and built environments are not only presented as solutions to create a more sustainable environment, but also as a means of increasing societal agency by spreading civic awareness. technology is shown to act as a major factor in facilitating communication and spread of awareness, accentuating the need for social equity and accessibility: v4: “algorithms and hundreds of employees spread awareness of municipal matters, and locals act as special experts of their own areas.” 6.2.3. socio-natural overlap society’s ethical behaviour within nature is explored through the thematic overlap of social and natural spheres. the driving forces behind all past change, and that yet to come, is explicitly stated to be personal agency and society’s agency, whilst also the city’s agency remains an implicit, yet important factor of development. in this future vision, personal agency is shown through the attendance of historical preservation events. the excerpt concerning the importance of agency and actions in achieving wanted changes is a key factor showing overlap of social and natural spheres: v4: “we have decided to influence and participate in the king’s road preservation event.” likewise, the social sphere’s overlap with nature is highlighted through the importance of voicing one’s opinion and being able to act for it in matters of environmentalism. the excerpt below portrays the effects of agency and shows the positive outcome of the collaboration of differing perspectives. additionally, the text again voices the complex nature of change through the wide array of alternatives, and of no solution being the sole good one. the general social agency in creating and promoting change is one that is highly prevalent within the vision yet remains rather implicit through the use of the passive tense rather than clearly depicting the group leading the change. in speaking of activism against the new developments regarding habitational structures, v4 explores the complex balance between social and natural spheres of wellbeing: v4: “the most extreme climate activists have of course been against this, but the harm has been compensated for by giving more space for nature inside the city.” 6.2.3. techno-natural overlap technology is widely explored from the perspective of sustainability and environmentalism, which shown through overlap of technological and natural spheres. the relationship between them is readily explored in v1, with the final only making a small addition to the text. the addition can be seen from the underlined words in the excerpt below. the addition shows slight reconsideration of the topic from v1 regarding the extent of the presence of technology within society, yet mainly v1 demonstrates how the students already consider the impact of the technological sphere upon nature before revising their visions: hyyppä, rasa, & laherto 40 | f l r v4: “a nearby apartment building’s dark, wooden walls and extensively covering wall solar panels flicker some hundred meters away. the low sun rays fill the surrounding parks and light up the lower budlings’ vernal green roofs…” 6.3. group 3 – a future of minimalism and nature (substantial revisions) titled “life amidst climate change”, group 3 visioned a green future that promotes the wellbeing of nature and wildlife. the vision revolves around minimalism in individuals’ future needs, promoting greater environmental goals over society’s comfort. the technologies and scientific developments explored are conservative and further exemplify the green developments of the story. as can be seen from figure 4, the three spheres are explored fairly equally. the greatest developments occur in the broadest systems thinking, the overlap of all three spheres, with 7 excerpts. the agency behind change is explored implicitly, yet the vision emphasizes the importance of small changes within each individual’s quotidian life, taking on a first-person narrative. the developments from v1 to v4 largely involve all spheres, with v4 revisiting the topics of the first version and linking them to other spheres and exploring their impacts in relation to each other. 6.3.1. overlap of all 3 spheres the development of systems thinking, and its complexity, is exemplified through the revisitation and connection of topics initially explored in v1. in the following excerpts, v1 recognizes the positive impact of the change in comparison to the situation beforehand, while in v4, the same excerpt receives a more extensive analysis. v4 recognizes the reason behind change, the social aspect of transport, the availability of shareable transport, and their functions. the revisitation of the single sentence shows how the group conceptualized the initial topic of traffic from technological, natural, and social spheres through consideration of causality and functionality: v1: “traffic noise is, however, lower than it has been in the past…” continued in v4: “this is a result of the electrification of traffic, for one. i don’t own a car, but i use electric rental cars that are easily available. this works quite like renting an electric scooter. the app shows you where the cars are…” 6.3.2. socio-technological overlap the relationship between technology, science, built environments, and society is, in this vision, commonly depicted through the development of technological devices to aid society and its members with everyday life tasks and activities. the text itself (see following excerpt) explicitly describes a development that occurred with the function of and relation towards technology, adding an ethical perspective of accessibility and sustainability to technological devices. while v1 discusses health and privacy as sociotechnical points of interest, v4 complexifies the sociotechnical system by including issues of ethical technology development, sustainable resource use, job markets, and even (the uncertain nature of) risk assessment. both excerpts also discuss the safety of technological devices in relation to personal information, and v4 even adds a further challenge linked to society’s attitudes towards technological development, namely that of it stealing society’s jobs. by addressing both flaws and advantages of technology, the text depicts many conceptual levels of technology and its impacts on a wide array of different societal matters: v1: “i can check whether my food needs more minerals or vitamins from my smart watch. i can also see other important health information, and it has a connection to my home appliances. figure 4. revisions from v1 to v4, group 3 (nts) hyyppä, rasa, & laherto 41 | f l r most of my furniture is electric and can be voice-controlled, which i find very useful. at some point this was an information security risk, but it has fortunately been fixed…” continued in v4: “it guides towards a healthy lifestyle without encouraging obsessive use. … technology has been deployed for peoples’ benefit and aid, it is used by all, and it is not expensive. there are new ways to manufacture electronic devices with scarce natural resources. ai has been improved and it can be seen in everyday life. … ai does some of our old work, which was not effective by manual labour. even though it was suspected at first, ai has not taken jobs from anyone, but created more.” 6.3.3. socio-natural overlap the future vision describes the relationship between social and natural spheres through depictions of society’s ethical behaviour within nature. as previous topics within the text, both versions of the vision speak of the relationship between nature and society both explicitly and implicitly, however only v4 analyses the past change that has occurred to arrive at the future that is spoken of. this can be seen as the vision brings to light the humane drivers of change, and how that change must be reflected in society’s behaviour. the excerpt also enforces the idea that the change has been driven forward through quality education and generational progression, while also adding the technological element as a benefit for increased communication and awareness, which lead to increased consciousness of actions. the story portrays additions to the overlap of social and natural spheres in v4 by discussing a topic area unexplored in v1: v4: “i am happy that more people have begun to do their best for the climate in collaboration. this has come as a result of generational change and good education. good global connections through video calls have spread information on different ways to slow climate change…” 6.3.4. techno-natural overlap in speaking of the future in a very modernized yet environmental manner, the vision combines the thematic spheres of technology and nature to create an image of scientific and technological development being leading factors in driving change towards a green, sustainable future. one of the main aspects of technology-related environmentalism is that of green energy production, that is shortly explored in v1 of the text (see below), showing the emphasis on renewable energies and recycling, and their use and impacts of smaller communities as well as on the city as a whole. the excerpt is further developed in v4, portraying development of complexity and systemics in the recognition of different renewable energy sources and the realism of their use. the extent of the reach of sustainable energies is further explored through the analysis on green public transport and its reduction in noise emissions. by bringing into discussion alternative energy sources, such as fusion, and its effects on energy production as well as its challenges in becoming a normalized energy source, the story progresses between the two versions of the stories through recognition of choice consciousness and the multitudes of options available to create the future of their v1: v1: “electricity and heating in housing cooperatives comes from renewable energy sources and recycling is an important part of life in our housing cooperative and in helsinki…” continued in v4: “solar panels provide habitants with daily electricity and have been installed on each house. energy can be stored for cloudy days. traffic has become electric and does not make noise. electricity for traffic comes from small solar and the long-awaited fusion plants. part of the city works on fusion energy, which has decreased electricity prices considerably despite only having one power plant. therefore, it has not yet reached markets properly.” hyyppä, rasa, & laherto 42 | f l r 6.4. group 4 – a future of efficiency (substantial revisions with most new text) with their future vision, “helsinki in 2050”, group 4 explored the most drastic and complex of changes among all groups. the group partially rewrote entire sections of the vision unlike other groups, which mostly made revisions onto the initial text. the story revolved around efficiency from the perspectives of personal ethics, transport and energy production, and social structures. with a strong perspective on the overlap of social and technological matters, with 7 excerpts, the story mainly developed in its perceptions of technology as an all-round tool for social change. the initial story was already largely environmentalist, and hence the development of systems thinking is least explored through the nature category. the described developments are complex to the point of discrepancy and lack of causality within the text, indicating the changes visioned could have been even further developed. the story involves socialist and communist social structures, yet within its social changes, the agency behind them is considered merely implicitly. this contrasts with group 3, for which technological change seemed to be the primary point of entry to an imagined future; for group 4, change was primarily socio-political. 6.4.1. overlap of all 3 spheres it is only in v4 of the text that all three spheres are deeply considered among the same topical excerpts (see example below), showing increased complexity even on a surface level. combining the themes of technology, material efficiency, and community, v4 recognizes the changes needed to achieve the future described, focusing on the societal aspect of true need as compared to use. the technological development is considered alongside its effects on society, limiting the use of devices for necessities, not recreational use. the change is reasoned through the excessive materialism caused by unnecessary use of technology, in turn causing excessive production and emissions. in contradicting the role of technology present in other excerpts of the story, the following excerpt from v4 thus incorporates issues concerning the sustainability of technology from both a societal as well as environmental perspective, which were unexamined in v1: v4: “technology exists, but no longer in recreational use. it is used for communication and in some cases work. … the issue with ai technology is its material aspect, as it consumes natural resources, production is untransparent and unjust, and it must be produced constantly due to its short lifespan. the solution was to make devices tools of community, with smaller batteries, lighter operating software and changeable parts.” 6.4.2. socio-technological overlap “helsinki in 2050” highlights the societal aspect of technological development, shedding light on the need for critical thinking in achieving the full benefit of technology, without excess use, production, or reliance. v4 adds the perspective of the possible downsides of technology, emphasizing the importance of recognizing changes in attitudes and allowing for alternative solutions when a previous plan has proved ineffective: v4: “initially technology was believed to fix everything: climate change, environmental crises, political unrest, criminality, marginalization, drug use etc. countless hours and resources were invested in its development, but technology didn’t magically fix everything…” figure 5. revisions from v1 to v4, group 4 (nts) hyyppä, rasa, & laherto 43 | f l r the critical perspective takes a step back from technology-oriented future narratives regarding society and its needs and wants. v4 proceeds to explore the equality issue of technological development. starting off from a technological perspective, the excerpt below continues to explore the importance of social equity brought on by the communicative aspects of technology and their accessibility. the sense of fairness and community is further explored through the addition of increased collaboration and decreased competition between producers of technological products and services. simultaneously, the excerpt makes a statement on the structures of the economy and its businesses, implicating pure collaboration as a better means of production than competition, later stating the profit remains the same. although the aspect is not fully explored, it aims to consider a further, complex perspective of technology in relation to the people that make and use it, depicting a grand increased in systems thinking related to how technology intertwines with the creating society: v4: “it has been realized that all matters, excluding vital and just life and communication, only spark momentary joy, and have nothing to give. hence all joint ownership production technology firms collaborate, as it improved the work and product. they plan collectively without competition. devices are made for people, not profit. otherwise, devices would become jaded, and the hard work would be lost.” 6.4.3. socio-natural overlap the development of systems thinking within the vision concerning society is not only linked to community and social wellbeing, but also to the central values of the story, environmentalism, and sustainability. the society of the story is shown to take agency in promoting an environmentally friendly lifestyle and environment, creating a sense of social pressure for all citizens to act sustainably. in adding a social taboo on unsustainable transportation methods to v4, the vision depicts the importance of choice consciousness among all of society, emphasizing the importance of community in leading each other towards a common goal: v4: “a social ban has been placed on private car travel … due to resources and the contami nation of cars.” the display of social ethics within nature is further explored through the city’s and society’s agency within the city environment and the greater world, as can be seen through the excerpt discussing immigration and climate change. while stating the failure to achieve sufficient change in time, the excerpt speaks of the failure as a communal doing, touching upon the need for communal actions in creating change. in speaking of the destruction of nature due to society’s previous actions and the inability to undo them, the vision also approaches the importance of social togetherness in times of natural disaster. the developments within the systems thinking related to human ethics and nature consider different perspectives of natural change, bringing to light the destructive realities caused by society’s inability to act in time and enforcing a sense of societal responsibility over natural wellbeing: v4: “helsinki’s population has grown only through immigration, as many areas were destroyed by climate change. we were not able to halt it in time, despite efforts.” 6.4.4. techno-natural overlap systemic development in this vision can also be analysed through the overlap of natural and built environments, exploring how the visions show evidence of interconnectedness between the environment and science. v4 (see excerpt below) discusses the impacts of changes to the built city environments upon the wellbeing of nature. the excerpt explores the reduction of car traffic upon the decreased need for asphalt roads, allowing for the increased prevalence of dirt roads, in turn improving the natural water cycle. in exploring the effects of built environments upon the conditions of the nature, the vision depicts the effects society has, directly or indirectly, upon nature, emphasizing the idea of technology, science, and built environments existing within nature, not above or alongside them: v4: “car travel and its pollution have decreased. therefore, many asphalt roads have been converted to sand, as they no longer have polluting … traffic. fixing asphalt roads is expensive hyyppä, rasa, & laherto 44 | f l r and releases bitumen into the environment, so their upkeep was no longer wise. with sand roads, urban runoffs were deployed as they were only needed to guide waters into the right places on asphalt roads. ridding urban runoffs enabled a natural water cycle in the city, influencing vegetation area wellbeing and decreasing minor flood risks.” 7. discussion in examining the developments of students’ future visions during a futures education course from the perspective of systems thinking, this study took on a constructive approach in viewing futures thinking through hopeful futures. the course focused on evoking hopeful futures thinking as a means of directing the students’ ways of thinking towards the possibilities of solutions to ongoing challenges. while the focus in no means aimed to yield purely positive future images, the notion of hope within the course ties in strongly with the motivation of developing students' agency beliefs to create better futures rather than live in perpetual climate and future anxiety. the revisions of the students’ future visions gave insights into how they were able to revisit topics and reconceptualize them in a more interconnected, complex structure. the development of systems thinking between the two versions of the groups’ future visions encloses multitudes of topic areas and analytical perspectives. through them it can be observed how the interconnectedness of the topics within the stories grows and portrays an increased understanding of the complexity of social, natural, and technological systems and thematic spheres. the development of systems thinking was in this study analysed from the perspective of the city as a whole made up of three thematic spheres: society and human organizations; natural world and biosphere; and technology, science, and built environments. the analysis here portrays the development of the students’ abilities to conceptualize complex matters on both detailed and abstract levels. the results bear resemblance to those of rasa et al. (2023), who analysed complex change in students' views of the future and identified varying extents of systemic thinking. however, their study did not examine the effects of futures learning on these conceptions. thus, our study points towards the potential of futures education in addressing some of the concerns raised by rasa et al (2023). in showing most development in the overlap of all three spheres, students showed development of complete systems thinking through increased understandings and analyses of the relations between human agency as a driver of change, technology, and science as tools to build more sustainable societies, and the constant impact of the biosphere on all actions and decisions. similarly, hofman-bergholm (2018) concludes the improved recognition of interconnections to be a central benefit of the systems thinking approach. the development involves understandings of how technology can act as a tool, yet only when used with precautions and limitations, in reverting to a more natural, green environment. furthermore, the development of systems thinking links to how all change is dependent on social agency, which is shown through notable developments within the thinking models employed by the students during the futures education course. in general, all groups together most emphasized the social sphere, second the technological sphere, and least the natural sphere. the focus on social aspects of the future is one that adds onto the findings of levitas’ (2013) studies. the number of excerpts speak to the development of the students’ future visions throughout the course, and as such show the spheres most central within each groups’ visions. as fig. 1 displays, most revisions were made in the overlapping themes, speaking to the broadened conceptualization skills of the students. with all groups put together, the most frequently apparent overlaps were those of all three spheres, and the socio-technological sphere. as previously indicated, systems thinking is here perceived through the interconnectedness of the three thematic spheres, and as such the full overlap of all three spheres shows significant reconsideration of the topics within students’ future visions and revision of their implications upon other topics. additionally, sustainability and environmentalism were central themes promoted to varying extents in all groups’ future visions. linked in with aspects of social agency, the needs for sustainable technological development and the relationships between natural and built environments, environmental hyyppä, rasa, & laherto 45 | f l r matters permeated all texts. adding on to hofman-bergholm’s (2018) findings on the importance of a systems thinking approach in sustainability education, the results point to the fact that sustainability cannot be excluded from the themes of society and technology. even though the natural sphere is the one with least revisions between v1 and v4 of all groups together (43 excerpts), it is the sphere most tightly interconnected with the other thematic spheres. merely two excerpts were considered from only the natural sphere, whilst all others intertwined natural matters into their social and/or technological contexts. this points to the developed understanding of the role of nature as a complex, intertwined entity impacted by, and impacting, the areas of social and technological development. the revisions showed more focus on the natural sphere, indicating deepened understandings of how natural environments, resources, and biospheres impact the thematic spheres of society and technology. technological themes were present in each group’s vision from the initial version, indicating the current importance of the topic area, as the course did not need to externally raise the importance of the technological sphere. one central finding of this paper is that the revisions showed how students were able to interconnect the technological topics present in v1 to their impacts on society and the natural environments, an aspect equivalent to the findings of rasa and laherto (2022). the intertwined nature of the thematic spheres was only present in the v4s of each groups’ texts, pointing to noteworthy revisions of technological topics during the course. whilst the general line of development of systems thinking was similar between groups, each group portrayed a slightly different emphasis within their stories. group 1 and 4 had most development within the social sphere, group 3 within the technological sphere, and group 2 within both the social and the natural sphere. within the same futures education course, groups were able to elaborate their visions based on their individual topics of interest of concern, which is visible through the differences in emphasis of the thematic spheres. the finding is positive in considering the narrative nature of the results. the aim of the futures education course was in no means to impose a certain vision of the future upon students, but rather to promote futures awareness, critical thinking, and agency for change, and allow students to dream and conceptualize their own futures. by further looking into the thematic spheres and topics discussed within the individual groups’ visions, the study found that even in a small sample size study, student groups wrote visions that differed entirely from others. emphasizing the concept of futures as undefined, pluralistic scenarios, each group wrote a personal future vision whilst attending the same futured education course as the other groups. group 1 showed development of systems thinking by attributing the decreased environmental footprints to scientific development, exploring the relationship between science and nature in creating their tomorrow. group 1 also emphasized the role of governance and civic responsibility, showing a deepened understanding of how ultimately it is society and its individual people who enforce the change needed for a better future. as for group 2, whilst not revising their story as much as other groups, they touched upon a meta-analysis of their own text, looking for faults and imperfections within their future changes. in speaking of climate activism against their future city, group 2 showed deep consideration of the nature of humanity being made up of its differences. furthermore, group 2 visited the topic of artificial intelligence being employed for communication and spreading of awareness, indicating increased complexity between the roles of society and technology. group 3 had a strong stance on the pathways leading to change, showing systemic consideration of what changes society must make in order to improve the state of the environment. whilst promoting topics of social activeness and environmental wellbeing, group 3 approached their story from a very technological perspective, creating a vision of a complex society in which technology has been tied into most aspects of life. contrarily, group 4 approached the changes in their future vision from a socio-political perspective, examining how a change in social and governmental structures creates their tomorrow. as with group 2, group 4 also explores criticism towards their own vision in discussing the downsides of technology and in exploring how they intend to solve those issues. whilst each group visioned an entirely different future city, all groups explored the interconnectedness of their topics, as well as new ones, through the revisions of their texts. it is also noteworthy that at times, students’ efforts in visioning complex entities and understanding their roles in the city came through as discrepancies within the visions. the group that made hyyppä, rasa, & laherto 46 | f l r most revisions to their future, which already from the beginning was more on the radical side of changes, group 4, explicitly stated that their futures revolved around humane and natural values. they stated that the technological craze had been overcome and that the role of technology in people’s lives had substantially decreased. an analysis of their vision, however, showed that technology pertained a fundamental role in their city of the future. technological developments were intricately woven into areas of environment, transportation, and even basic everyday matters such as automatized streetlights. the discrepancy in itself portrays development of systems thinking in that students recognize the multiple areas of impact that technologies have, even if their vision is not able to completely conceptualize the full interconnectedness of all topics within the thematic spheres. rubin (2013) explores similar topics regarding students’ confusion about matters related to the future in a time of constant change. likewise, the result adds to the findings of jacobson (2000) in showing that problem solving of complex systems, such as the city here, is not immaculate among novice-level systems thinkers, such as the participants of this study. as this study was conducted within one futures education course in one local upper secondary school, it must be noted that the results cannot be generalized to all students due to the small sample size of the data. likewise, the future visions themselves are not representative of a larger sample, but the systems thinking apparent in the revisions of the stories can be taken as an indication of the forms of fundamental skills that can be strengthened through futures education. as such, this paper indicates possible results that can stem from employing futures education pedagogies and challenging students’ future visions during a futures education course. further larger scale and broader studies would be needed to create a generalizable, representative dataset. the present findings provide illustration of how futures education may contribute to learning systems thinking. specifically, the results capture some of the dynamics between futures and systems thinking in a suitable context (the city). still, this small-scale study is not generalizable – and furthermore, futures education courses are not set to produce uniform visions of the future. this approach should be expanded to explore how futures education can address systems and their complexities and interrelatedness, alongside a deeper consideration of agency beliefs towards creating more sustainable futures. likewise, this study could be complemented by research on the long-term effects of the futures education course. a discourse analysis on the same data could also provide deeper insights into singular future visions and their depictions of the city of the future. nevertheless, the findings of this study show that students’ understandings of the complex systems and their interrelatedness developed during the course, showing positive predictions for the success of similar futures education courses. it would also be scientifically relevant to study the perceptions of certainty and uncertainty towards the future of the same, or similar texts, to understand how futures education can shape not only students’ ways of thinking systematically and conceptualizing systems, but also analysing how strongly students believe in their desired future. methods of analysing uncertainties of futures, such as those employed by maier et al. (2016), could provide deeper insights into students’ perceptions of plausibility and agency. similar to the examination of uncertainties, this study serves as a robust foundation for studying aspects of sense-making and strange-making of the future. previously explored by bol and de wolf (2023), their approach on this data could yield more results on perspectives on anticipation, empathic thinking, and imagination. 8. conclusions overall, the resulted development of systems thinking in students’ future visions during the futures education course provided new insights into the implications of such futures education courses: the development of systems thinking sheds light onto the progressive benefits of allowing, encouraging, and challenging students to work to build and solve a future in which they would wish to live. in merely a short-term course, students’ understanding of how future perceptions can change from abstract ideologies to concrete plans, as per rajala and cole et al. (2022), deepened. the finding differs from the current realities in schools and curricula (see e.g. finnish national agency for education, 2019), that hyyppä, rasa, & laherto 47 | f l r overemphasize skills and knowledge that students will supposedly need in the future, yet undervalue thinking about and understanding that future that they are heading towards. by visioning a hopeful future and backcasting to concretize the steps between now and an imagined future, this study adds onto that of rasa et al. (rasa et al., 2023) in showing how students were challenged to better conceptualize the need for agency and action in creating change. students showed development of complete systems thinking through increased understandings and analyses of the relations between human agency as a driver of change, technology and science as tools to build more sustainable societies, and the constant impact of the biosphere on all actions and decisions. in understanding the steps needed to achieve a possible future, and in conceiving their interrelatedness among matters of society, nature, and technology, visioning the city of the future was shown to develop both topic-area specific understandings. the development sheds light onto the importance of cross curricular skills and on vast development that can occur even in a short-term course. futures thinking in schools not only allows students to understand the possible futures of the world, but also evaluate their own thinking skills and models. systems thinking and futures education can promote an understanding of an individual’s role within an ever more complex world and foster the perception of their ability to impact the future. by understanding the complex roles of society, nature, science, and technology, students can begin to concretize their hopeful futures into actions in aims of creating a more environmentally and socially sustainable future. keypoints the study takes an approach on analysing education about the future with aims of exploring the development of cross-curricular competencies. the study explores students written visions of a hopeful future from the perspective of the city of the future. through a qualitative thematic analysis, this study finds that futures education expands students’ systems thinking. the development of students’ future visions during a futures education course emerges understandings of complex systems, interrelatedness, and causalities related to futures. the development of systems thinking is shown through the increased interconnectedness of social, natural, and technological thematic spheres that make up the city of the future. students’ systems thinking mostly developed in the complete interconnectedness of all thematic spheres. acknowledgments the authors of this article would like to thank the helsinki school of natural sciences, the teachers, the experts, as well as the students of the course for their participation. we also wish to thank the fedora research project partners, led by the university of bologna, for this research initiative. funding the research project (fedora, https://fedora-project.eu) from which the data of this study stems was funded by the european union’s horizon 2020 research and innovation programme under grant agreement no 872841. ethics approval and consent to participate participation in the course and in this study was voluntary for all students. all participants gave consent for the collection, storage, and use of their data, as well as its anonymized publishing. https://fedora-project.eu/ hyyppä, rasa, & laherto 48 | f l r references aalto, h.-k., heikkilä, k., keski-pukkila, p., mäki, m., & pöllänen, m. (2022). tulevaisuudentutkimus tutuksi perusteita ja menetelmiä. tulevaisuudentutkimuksen verkostoakatemia, tulevaisuuden tutkimuskeskus, turun yliopisto. https://urn.fi/urn:isbn:978-952-249-5631. ahlqvist, t., & rhisiart, m. (2015). emerging pathways for critical futures research: changing contexts and impacts of social theory. futures, 71, 91–104. https://doi.org/https://doi.org/10.1016/j.futures.2015.07.012 ahvenharju, s., minkkinen, m., & lalot, f. (2018). the five dimensions of futures consciousness. futures, 104, 1–13. https://doi.org/10.1016/j.futures.2018.06.010 angheloiu, c., sheldrick, l., & tennant, m. (2020). future tense: exploring dissonance in young people’s images of the future through design futures methods. futures, 117, 102527. https://doi.org/https://doi.org/10.1016/j.futures.2020.102527 bettencourt, l. m., lobo, j., helbing, d., kühnert, c., & west, g. b. (2007). growth, innovation, scaling, and the pace of life in cities. https://doi.org/https://doi.org/10.1073/pnas.0610172104 bol, e., & de wolf, m. (2023). developing futures literacy in the classroom. futures, 146, 103082. https://doi.org/https://doi.org/10.1016/j.futures.2022.103082 bol, e., laherto, a., levrini, o., erduran, s., conti, f., tola, e., troncoso, a., tasquier, g., rasa, t., & pucetaite, r. (2023). future-oriented science education manifesto. https://doi.org/10.5281/zenodo.7519150 braun, v., & clarke, v. (2006). using thematic analysis in psychology. qualitative research in psychology, 3(2), 77–101. https://doi.org/10.1191/1478088706qp063oa city of helsinki. (2018). the carbon-neutral helsinki 2035 action plan. facer, k. (2011). learning futures : education, technology and social change. taylor & francis group. https://doi.org/https://doi.org/10.4324/9780203817308 finnish national agency for education. (2019). lukion opetussuunnitelman perusteet 2019. https://www.oph.fi/sites/default/files/documents/lukion_opetussuunnitelman_perusteet_2019.pdf gharajedaghi, jamshid. (2011). systems thinking: managing chaos and complexity: a platform for designing business architecture (3rd ed.). morgan kaufmann. https://doi.org/https://doi.org/10.1016/c2010-0-66301-2 hickman, c., marks, e., pihkala, p., clayton, s., lewandowski, r. e., mayall, e. e., wray, b., mellor, c., & van susteren, l. (2021). climate anxiety in children and young people and their beliefs about government responses to climate change: a global survey. the lancet planetary health, 5(12), e863–e873. https://doi.org/10.1016/s2542-5196(21)00278-3 hicks, d. (2003). lessons for the future : the missing dimension in education. routledge. https://doi.org/https://doi.org/10.1016/s0016-3287(02)00060-5 hicks, d., & holden, c. (1995). visions of the future : why we need to teach for tomorrow. trentham books. https://doi.org/https://doi.org/10.1080/1350462950010205 hofman-bergholm, m. (2018). could education for sustainable development benefit from a systems thinking approach? systems, 6, 43. https://doi.org/10.3390/systems6040043 höjer, m., gullberg, a., & pettersson, r. (2011). backcasting images of the future city—time and space for sustainable development in stockholm. technological forecasting and social change, 78, 819–834. https://doi.org/10.1016/j.techfore.2011.01.009 jacobson, m. j. (2000). problem solving about complex systems: differences between experts and novices. in b. fishman & s. o’connor-divelbiss (eds.), fourth international conference of the learning sciences (pp. 14–21). kivistö, t. (1985). kaupunkien kehityksestä pitkällä aikavälillä. in p. malaska & m. mannermaa (eds.), tulevaisuuden tutkimus suomessa (pp. 201–221). gaudeamus. hyyppä, rasa, & laherto 49 | f l r laherto, a., & rasa, t. (2022). facilitating transformative science education through futures thinking. on the horizon, 30(2), 96–103. https://doi.org/10.1108/oth-09-2021-0114 laherto, a., rasa, t., miani, l., levrini, o., & erduran, s. (2023). future-oriented science education building sustainability competences: an approach to the european greencomp framework bt. in fazio, x. (ed.). science curriculum for the anthropocene, volume 2: curriculum models for our collective future (pp. 83–105). springer international publishing. https://doi.org/10.1007/978-3-031-37391-6_5 levitas, r. (2013). utopia as method: the imaginary reconstitution of society. palgrave macmillan. https://doi.org/https://doi.org/10.1057/9781137314253 levrini, o., tasquier, g., barelli, e., laherto, a., palmgren, e., branchetti, l., & wilson, c. (2021). recognition and operationalization of future-scaffolding skills: results from an empirical study of a teaching–learning module on climate change and futures thinking. science education, 105(2), 281–308. https://doi.org/10.1002/sce.21612 lotz-sisitka, h., wals, a. e. j., kronlid, d., & mcgarry, d. (2015). transformative, transgressive social learning: rethinking higher education pedagogy in times of systemic global dysfunction. current opinion in environmental sustainability, 16, 73–80. https://doi.org/10.1016/j.cosust.2015.07.018 maier, h. r., guillaume, j. h. a., van delden, h., riddell, g. a., haasnoot, m., & kwakkel, j. h. (2016). an uncertain future, deep uncertainty, scenarios, robustness and adaptation: how do they fit together? environmental modelling & software, 81, 154–164. https://doi.org/https://doi.org/10.1016/j.envsoft.2016.03.014 mangnus, a. c., oomen, j., vervoort, j. m., & hajer, m. a. (2021). futures literacy and the diversity of the future. futures, 132, 102793. https://doi.org/https://doi.org/10.1016/j.futures.2021.102793 moroni, s., & cozzolino, s. (2019). action and the city. emergence, complexity, planning. cities, 90, 42–51. https://doi.org/10.1016/j.cities.2019.01.039 moss, t., voigt, f., & becker, s. (2021). digital urban nature: probing a void in the smart city discourse. city, 25(3–4), 255–276. https://doi.org/10.1080/13604813.2021.1935513 page, j. (1996). education systems as agents of change: an overview of futures education. in new thinking for a new millennium the knowledge base of futures studies (pp. 126–136). routledge. poli, r. (2021). the challenges of futures literacy. futures, 132, 102800. https://doi.org/https://doi.org/10.1016/j.futures.2021.102800 rajala, a., cole, m., & esteban-guitart, m. (2022). utopian methodology: researching educational interventions to promote equity over multiple timescales. journal of the learning sciences, 1–27. https://doi.org/10.1080/10508406.2022.2144736 rasa, t. (2023). futurising science education: technology, agency and scientific literacy. helsinki studies in education, no. 166 [doctoral dissertation, university of helsinki]). https://doi.org/http://dx.doi.org/10.13140/rg.2.2.23545.24164 rasa, t., & laherto, a. (2022). young people’s technological images of the future: implications for science and technology education. european journal of futures research, 10(1), 4. https://doi.org/10.1186/s40309-022-00190-x rasa, t., lavonen, j., & laherto, a. (2023). agency and transformative potential of technology in students’ images of the future: futures thinking as critical scientific literacy. science and education. https://doi.org/10.1007/s11191-023-00432-9 rasa, t., palmgren, e., & laherto, a. (2022). futurising science education: students’ experiences from a course on futures thinking and quantum computing. instructional science, 50(3), 425– 447. https://doi.org/10.1007/s11251-021-09572-3 remes, pirkko., & rubin, anita. (1996). tulevaisuutta etsimässä: tulevaisuusteema kouluopetuksessa. opetushallitus. rubin, a. (2013). hidden, inconsistent, and influential: images of the future in changing times. futures, 45, s38–s44. https://doi.org/https://doi.org/10.1016/j.futures.2012.11.011 hyyppä, rasa, & laherto 50 | f l r sande, ö. (1972). future consciousness. journal of peace research, 9(3), 271–278. https://doi.org/10.1177/002234337200900307 secretary-general of the oecd. (2018). the future of education and skills education 2030 the future we want. https://www.oecd.org/education/2030/e2030%20position%20paper%20(05.04.2018).pdf shaked, h., & schechter, c. (2019). systems thinking for principals of learning-focused schools. journal of school administration research and development, 4(1), 18–23. https://doi.org/http://dx.doi.org/10.32674/jsard.v4i1.1939 tasquier, g., branchetti, l., & levrini, o. (2019). frantic standstill and lack of future: how can science education take care of students’ distopic perceptions of time? in e. mcloughlin, o. e. finlayson, s. erduran, & p. e. childs (eds.), bridging research and practice in science education: selected papers from the esera 2017 conference (pp. 205–224). springer international publishing. https://doi.org/10.1007/978-3-030-17219-0_13 toivonen, s., rashidfarokhi, a., & kyrö, r. (2021). empowering upcoming city developers with futures literacy. futures, 129. https://doi.org/10.1016/j.futures.2021.102734 tripto, j., assaraf, o. b. z., snapir, z., & amit, m. (2017). how is the body’s systemic nature manifested amongst high school biology students? instructional science, 45(1), 73–98. https://doi.org/10.1007/s11251-016-9390-0 unesco. (2017). education for sustainable development goals : learning objectives. https://www.unesco.org/en/articles/education-sustainable-development-goals-learning-objectives united nations. (2023). provisional state of the global climate 2023. https://wmo.int/publicationseries/provisional-state-of-global-climate-2023 valciukas, j., & bell, w. (2003). foundations of futures studies: volume 1 : history, purposes, and knowledge (vol. 1). taylor & francis group. https://doi.org/http://dx.doi.org/10.2307/2655498 whitcomb, c., davidz, h., & groesser, s. (2020). systems thinking. https://doi.org/https://doi.org/10.3390/books978-3-03936-797-9 wiek, a., binder, c., & scholz, r. (2006). functions of scenarios in transition processes. futures, 38, 740–766. https://doi.org/10.1016/j.futures.2005.12.003 microsoft word knoopvancampenkoketal_finalproof.docx frontline learning research vol. 9 no. 4 (2021) 116 140 issn 2295-3159 * shared first authorship** authors contributed equally. corresponding author: carolien knoop-van campen, carolien.knoop-vancampen@ru.nl doi: https://doi.org/10.14786/flr.v9i4.881 how teachers interpret displays of students’ gaze in reading comprehension assignments carolien a. n. knoop-van campen1*, ellen kok2,3*, roos van doornik2**, pam de vries2**, marleen immink2, halszka jarodzka3 & tamara van gog2 1 behavioral science institute, radboud university; netherlands 2 department of education, utrecht university, netherlands 3 department of online learning and instruction, open university of the netherlands article received 3 june 2021 / article revised 16 september / accepted 4 november / available online 7 december abstract reading comprehension is a central skill in secondary education. to be able to provide adaptive instruction, teachers need to be able to accurately estimate students’ reading comprehension. however, they tend to experience difficulties doing so. eye tracking can uncover these reading processes by visualizing what a student looked at, in what order, and for how long, in a gaze display. the question is, however, whether teachers could interpret such displays. we, therefore, examined how teachers interpret gaze displays and perceived their potential use in education to foster tailored support for reading comprehension. sixty teachers in secondary education were presented with three static gaze displays of students performing a reading comprehension task. teachers were asked to report how they interpreted these gaze displays and what they considered to be the promises and pitfalls of gaze displays for education. teachers interpreted in particular reading strategies in the gaze displays quite well, and also interpreted the displays as reflecting other concepts, such as motivation and concentration. results showed that teachers’ interpretations of the gaze displays were generally consistent across teachers and that teachers discriminated well between displays of different strategies. teachers were generally positive about potential applications in educational practice. this study provides first insights into how teachers experience the utility of gaze displays as an innovative tool to support reading instruction, which is timely as rapid technological developments already enable eye tracking through webcams on regular laptops. thus, using gaze displays in an educational setting seems to be an increasingly feasible scenario. keywords: gaze displays; teachers; reading strategies; secondary education; eye tracking knoop-van campen, kok et al 117 | flr 1. introduction reading comprehension is a crucial skill for academic success (murnane et al., 2012). as such, reading comprehension plays an important role in nearly all subjects in secondary education. to optimally support students’ reading comprehension, teachers should ideally provide students with adaptive support that is tailored to their personal needs (van de pol et al., 2019). however, since reading is mainly a covert process, in practice it can be difficult for teachers to get insight into students’ needs and tailor their instructions accordingly. eye tracking might provide an innovative tool to visualize those covert processes, as it allows for constructing gaze displays. gaze displays are visualizations of students’ eye movements during reading. such displays have the potential to provide teachers with insights in the reading strategy used by a student, and as such, in the efficiency of the reading process. while eye tracking is currently not yet widely available to teachers, it can be expected to become a viable option in the near future: for instance, rapid technological developments already enable eye tracking through webcams on regular laptops (madsen et al., 2021). eye-tracking is in fact already being used in the educational practice in a school in germany (böhm, 2021). as such, investigating whether gaze displays could be a useful tool to support reading instruction is timely. whether it can be a useful tool for educational practice, however, depends first of all on how teachers interpret the displays and how they translate the information provided by the gaze display into action (molenaar & knoop-van campen, 2019), and secondly, on what promises and pitfalls of gaze displays teachers perceive. therefore, the present study aims to explore how teachers interpret students’ gaze displays of reading assignments and perceive their usefulness. 1.1 reading comprehension reading comprehension can be defined as “the understanding of the meanings of words as they are used in sentence contexts, comprehension of sentences, and the acquisition of new information from passages of prose” (guthrie & mosenthal, 1987, pp. 291-292). in this process, a mental construction of the meaning of the text is built with new information acquired from the text (guthrie & mosenthal, 1987; pikulski & chard, 2005). efficient reading comprehension combines accuracy and reading time: students quickly and correctly comprehend a text. in building such a mental construction, reading strategies are important (afflerbach et al., 2008). reading strategies can be defined as purposeful, goal-directed actions which readers perform to achieve an established reading objective (afflerbach et al., 2008). the choice for a certain reading strategy depends largely on the reading objective (guthrie & mosenthal, 1987; urquhart & weir, 1998). reading strategies can be divided into two main categories: selective reading strategies and intensive reading strategies (liu, 2010). the selective strategy is used when the reader is looking for certain information (e.g., the answer to a specific question) without having to understand the rest of the text (guthrie & mosenthal, 1987; krishnan, 2011). the parts of the text that do not contain the desired information are skipped (liu, 2010). the intensive reading strategy is often used for making a summary of the text or when remembering a text for learning purposes (katalayi & sivasubramaniam, 2013; krishnan, 2011; liu, 2010). when applying this strategy, the reader reads the entire text as all information is necessary (liu, 2010; urquhart & weir, 1998). effective regulation of reading strategies is an important aspect of academic achievement (andreassen et al., 2017; salmerón et al., 2017). there are various evidence-based best practices for reading comprehension instruction (gambrell et al., 2011). ideally, teachers have a balanced reading comprehension program (duke & pearson, 2009) which includes both explicit strategy instruction (duffy, 2002) as well as support in the form of modelling reading strategies in action (schutz & rainey, 2020) and is tailored to students’ needs. knoop-van campen, kok et al 118 | flr 1.2 visualizing reading comprehension with gaze displays in practice, it can be difficult for teachers to tailor their instruction to the needs of the students. since reading is a covert process, it is difficult for teachers to determine how students read texts and whether they apply reading strategies correctly and efficiently. the automaticity of viewing behavior makes it hard to report (kok et al., 2017; võ et al., 2016), so students’ self-reports are relatively susceptible to either purposefully or accidentally false answers (benfatto et al., 2016; shute & zapatarivera, 2012). teachers can have students read text aloud, but this is a time-consuming intervention and not suitable for classroom lessons. therefore, the information that teachers can draw upon to tailor their instruction is often limited. a way to uncover reading strategies is by capturing readers’ eye movements during reading using eye tracking and then visualize their eye movements in gaze displays (de koning & jarodzka, 2017; knoop-van campen et al., 2021; van gog & jarodzka, 2013). a gaze display of a student’s reading behavior is a condensed visualization of large data sets of his/her eye-movement recordings, which are x& y-coordinates of the focus of the eyes on a computer screen. like learning analytics, they provide condensed summaries of large amounts of information (almosallam & ouertani, 2013). in recent years, gaze displays have found several applications in education and educational research. for instance, visualizations of teachers’ gaze have been used to show learners where the teacher is looking (eye movement modeling examples), and this has been shown to often improve their learning from video examples (chisari et al., 2020; jarodzka et al., 2013; mason et al., 2015; scheiter et al., 2018; van gog et al., 2009). furthermore, displays of students’ own gaze have been used to improve their self-assessment and self-regulated learning (donovan et al., 2008; eder et al., 2020; henneman et al., 2014; kok et al., 2017; kostons et al., 2009). thus far, however, research on the use of gaze displays as a tool for teachers is limited. existing research shows that people can interpret gaze displays of others to a certain extent, but interpretation performance differs between different tasks and stimuli. specifically, studies indicate that people can identify search goals (zelinsky et al., 2013), preferences (foulsham & lock, 2015), and deduce a given task (bahle et al., 2017; van wermeskerken et al., 2018) from displays of someone else’s gaze. emhardt et al. (2020) presented participants with gaze displays of students who answered multiple-choice questions about graphs. they found that participants could infer from those displays which answer a student selected, and if students were confident in their answers. in a study with teachers, špakov et al. (2017) developed several different visualizations of gaze behavior of students learning to read and found that overall, teachers appreciated them, and considered them informative. however, they did not investigate how teachers interpreted those gaze displays, i.e., which information they extracted from the display, whereas teachers’ interpretations of this type of information are critical as they influence how they integrate the information in their professional routines (cf. knoop-van campen & molenaar, 2020). to dive deeper into this question, we approach the interpretation of gaze displays with the learning analytics process model (verbert et al., 2014). this provides a suitable framework for systematically analyzing teachers’ interpretations (verbert et al., 2014). the learning analytics process model distinguishes various stages (van leeuwen et al., 2021; verbert et al., 2014). the awareness stage entails teachers becoming aware of the information. the interpretation stage involves teachers asking themselves questions and reflecting on the information they see, and trying to provide answers to these questions by interpreting and creating new insights into the data. lastly, the enactment stage entails teachers using their interpretations and insights to induce new meaning or even change behavior (verbert et al., 2014). how teachers interpret and use (learning analytics) information is strongly impacted by their experiences and routines (molenaar & knoop-van campen, 2019). gaze displays are currently not used in secondary education, so teachers do not have any experience in the use of gaze displays. their interpretations might be arbitrary and undirected, possibly resulting in varied and contradicting interpretations of gaze displays. these interpretations are important, as they determine the contents of teachers’ actions realizing tailored reading instruction (cf. jivet et al., 2017; molenaar & knoop-van campen, 2019). knoop-van campen, kok et al 119 | flr in sum, gaze displays seem to be a promising tool to provide information on covert processes like reading comprehension. findings on how people in general interpret gaze displays, suggest that teachers might be able to extract information on students’ reading strategies from their gaze displays and compare that to their idea of an effective (desired) strategy. however, making inferences about students’ reading comprehension is arguably a much more complex task than e.g. inferring their preference for an object or multiple-choice answer from a gaze display (in which people tend to look more towards the preferred object or answer option), as it requires interpreting patterns of eye movements. in addition, teachers are known to differ largely in their way they interpret process data (molenaar & knoop-van campen, 2019) which, in the case of reading comprehension tasks, may depend even more based on their own vision of (the importance of) efficient reading and reading strategies. thus, it is unclear which information teachers would be able to extract from gaze displays of students’ reading behavior, and whether this is consistent between teachers. 1.3 the present study the present study investigated how teachers interpret gaze displays and whether and how they would want to use gaze displays in teaching reading comprehension. gaze displays are an innovative tool that could potentially support reading instruction, and the questions addressed in this study are timely as eye-tracking technology is rapidly developing and might become available for use in schools in the near future (as mentioned earlier rudimentary eye tracking is presently already enabled through webcams on regular laptops). moreover, these questions are not only relevant for reading comprehension research and practice, but will also add to eye-tracking research on the use of gaze displays, by investigating whether people can also interpret gaze displays of more complex tasks. specifically, we examined whether teachers could discriminate between gaze displays that show different reading strategies and whether there was consistency among teachers in interpreting the same gaze displays. if teachers discriminate between gaze displays that reflect different reading strategies (i.e., distinguish between displays in their interpretation of them), this reflects that the visual information in the displays is interpreted in terms of (reading) processes that differ between students. if teachers are consistent in their interpretation of the same display (i.e., teachers agree in their interpretation of a display), this indicates that extracted information is interpreted similarly between teachers. finally, we examined the promises and pitfalls teachers reported in working with gaze displays. 2. method 2.1 participants participants were teachers who teach the subject dutch at secondary schools (in the senior general and/or university preparatory tracks) across the netherlands and as such, are knowledgeable ofreading comprehension strategies. all schools in the netherlands in which senior general and/or university preparatory is taught were invited by e-mail to participate. informed consent was provided by 125 participants. only participants who answered all questions for at least one display were included, which totals 60 participants (46 female). the exact number of participants differs between analyses and is reported in the results section. for the full sample, teachers’ mean age was 45.4 (sd = 12.8) and they reported an average of 13.1 (sd = 9.2) years of working experience (range 0-38 years). 2.2 design an explorative approach with open questions was used to elicit thoughts from teachers about gaze displays. in a questionnaire, secondary school teachers were presented with three different gaze knoop-van campen, kok et al 120 | flr displays of students who completed a reading comprehension assignment (see figure 1). teachers typed a response to three open questions per display and two open questions about displays in general. the learning analytics process model (verbert et al., 2014) was used to query teachers’ awareness of the gaze displays (what did teachers observe), their interpretation (how did teachers interpret this information), and their enactment (how did teachers intend to act upon this information). concerning the question of how teachers interpret gaze displays, we coded their answers and examined whether teachers discriminate between gaze displays of different students and whether there is consistency among teachers regarding interpreting the same gaze displays. finally, we examined the promises and pitfalls teachers reported in working with gaze displays by asking them about their views on implementing gaze displays in general in their reading comprehension education. figure 1a. gaze display x. figure 1b. gaze display y. figure 1c. gaze display z. 2.2 instruments 2.2.1 gaze displays three static displays were used (see figure 1). the displays were images of the eye-movement locations during the full trial. the gaze displays were taken from an earlier study (knoop-van campen et al., 2021). in that study, students received reading comprehension assignments which they made voluntary without time limit. students received a reading comprehension task (“which advantages of self-driving trucks are mentioned in the text?”), and subsequently were presented with the text in which they had to find the answer. the text consisted of three paragraphs (three areas of interest: aois) and was preceded by an open-ended question of which the answer is situated in the second paragraph (bottom left). after reading the text, they typed in their answer (without the opportunity to visit the text again). an smi red-500 eye-tracker (smi vision, berlin, germany) captured students’ eye movements during reading at 250 hz. gaze displays were developed in begaze 3.7 (smi vision, berlin, germany) using the scanpath utility. fixations were shown as circles with the size of the circle corresponding to the fixation duration (500 ms = 1 degree). saccades were shown as lines between knoop-van campen, kok et al 121 | flr fixations. fixations and saccades were defined using smi’s standard velocity-based event-detection algorithm with peak velocity threshold of 40 degrees/second and minimum fixation duration of 50 ms. the three gaze displays varied on the type of reading strategy used. the reading strategy (intensive or selective) was quantified with a so-called disparity score, which captures the duration of fixations across different parts of the text (see (knoop-van campen et al., 2021, for more information). this disparity score was calculated based on the standard deviation of the weighted fixation duration times in the three paragraphs. when the standard deviation is low, an equal amount of attention was given to all three paragraphs, while a high standard deviation indicates that the student focused mostly on one part of the text. this disparity score could vary between 0 (attention evenly divided over the three aoi’s: sd of 33%, 33%, and 33% is 0) and 57 (attention 100% focused on one aoi: sd of 100%, 0%, and 0% is 57). based on an experts’ cut-off point and hand-coded validation, disparity scores below 14 were considered intensive reading strategies, and above 14 were considered selective reading strategies. the three selected displays showed a clear intensive strategy (display x: disparity score = 1.42), a clear selective strategy (display y: disparity score = 25.25) and a borderline intensive strategy (display z: disparity score = 11.73). data quality was expressed in the tracking ratio (percentage of time samples where a valid x& y-coordinate was recorded: display x: 97%, display y: 98%, display z: 100%) and deviation scores as retrieved using a four-point validation procedure after nine-point calibration (m = .62, sd = .53. 2.2.2 measures to elicit thoughts from teachers about their awareness, interpretation, and enactment of the gaze displays, three open-ended questions reflected these phases of the learning analytics process model (verbert et al., 2014), see table 1. to optimally grasp teachers’ motives, opinions on and attitudes towards gaze displays, the open-ended questions were phrased according to the interpretive paradigm (tijmstra & boeije, 2011) in such a way that allowed teachers to elaborate on their thoughts (e.g., “please feel free to describe everything that comes to mind.”). table 1 questions for awareness, interpretation, and enactment phase question awareness phase “please take your time to look at the display and explain what you see. what draws your attention? please be as specific as possible.” interpretation phase “you have just described the display, what does this tell you about the reading behavior of the student? please write down everything that you think can be seen in this display.” enactment phase “if this were your student, after viewing and analyzing this display, what would you like to do or do next? consider, for example, adjustments in your education, instruction, or supervision. describe your considerations for whether or not to initiate these actions.” to collect teachers’ perceived promises and pitfalls of gaze displays, also open-ended questions asked whether they thought gaze displays could be useful in their educational practices. first, they were asked “do you feel that gaze displays of your students would help you as a teacher to have a deeper, better, or other insight into the reading behavior of your students? why do you think so?”. then, they were asked “would you use gaze displays of your students (in the future) if this would be easy to realize? knoop-van campen, kok et al 122 | flr for which goal would you use those displays? which concrete actions would you want to realize?”. again, teachers were invited to write down their answers. 2.2.4 procedure the online survey was developed in qualtrics (qualtrics, provo, ut). the online survey mode enabled teachers to complete the survey at any time and place that was suitable for them within three weeks after receiving the invitation. teachers spent on average approximately 23 minutes to complete the survey. in this questionnaire, informed consent was given, following which participants answered demographic questions. next, the different components of gaze displays were shortly explained (i.e., circles show where a person looks, larger fixations represent longer reading times, lines shown the route over the stimulus). then the reading comprehension question and the accompanying text were displayed so teachers could become familiar with the material and knew students’ reading objective. next, teachers were presented with consecutively the three gaze displays with corresponding questions (displays were anonymized, so teachers did not receive any student characteristic). the order in which the three gaze displays were presented, was randomized. last, teachers were asked to comment on the usability of gaze displays in education in general. 2.2.5 data analyses the learning analytics process model (verbert et al., 2014) was used as the basis for coding the written responses of awareness, interpretation, and enactment of gaze displays. additionally, open coding of all answers was conducted, to which (if necessary), theoretically relevant codes were added (such as those for strategies). to enable expressing consistency and discrimination, we created contrasting code scores for awareness and interpretation. for example, a score of 1 was given on the code ‘title’ if the teacher mentioned that the student looked at the title, and a score of -1 was given if the teacher mentioned that the student did not look at the title. if the teacher did not mention the title at all, no score was given for that code. see appendix a for the codebook including coding instructions and examples. coding took place in three rounds and by four coders. first, all four coders each coded the same ten cases and compared and discussed their answers, based on which the codebook was adapted. in round two, each coder coded 25% of the data plus 20% (in total 31 cases) as a second coder. the average krippendorffs alpha1 was 0.56 and the average percentage overlap was 92%. all variables (12 variables in total) with a krippendorffs alpha <.60 or percentage overlap <.80 were discussed. in round 3, these variables were recoded and a new 20% of overlapping cases were coded. one variable with a krippendorffs alpha of 0 was removed because only a few participants mentioned it. the final average krippendorffs alpha was 0.7 and the average percentage overlap was 94.1%2. for analyses, scores of 0 were defined as system-missing values, and subsequently, chi-squared tests were used to quantify teachers’ discrimination of displays. we considered a significant chi-square test as support that teachers discriminate between gaze displays on that code because it means that the pattern of -1 scores and 1 scores differs between gaze displays and that enough teachers mentioned this code to have the statistical power to detect differences. the average consistency within a code was the percentage of teachers who mentioned the most-mentioned score (either -1 or 1), averaged over the three displays. consistency ranged between 50% (equal numbers of answers were scored as -1 or 1) to 100% (all answers were scored the same, as either -1 or 1). the coding process for the perceived promises and pitfalls of gaze displays was similar to that of awareness, interpretation, and enactment. each of the four coders coded 25% of the data and an 1 krippendorffs alpha was calculated using the kalpha macro in spss at the nominal measurement level (hayes & krippendorff, 2007) because it enables the inclusion of four coders. 2 note that krippendorffs alpha is sensitive to low-prevalent scores, especially with binary variables (like our enactment codes). in this study, almost all codes have low-prevalent scores. so, in combination with the high percentage overlap, we consider the reliability acceptable. knoop-van campen, kok et al 123 | flr additional 31 cases as a second coder (63% of data double coded). the average krippendorffs alpha was 0.71 and the percentage overlap was 93%. the three variables with a krippendorffs alpha < .60 or a percentage overlap of < .80 were discussed. two of those were removed because the code was only used four times. the third variable was coded again and all cases were double-coded. however, krippendorffs alpha was negative after recoding, and since this code was also quite uncommon, it was dropped too. the final krippendorffs alpha thus stayed 0.71 with 93% overlap. answers to the two questions were originally coded separately for each question, but most codes were combined since teachers reported pitfalls and promises as a response to both questions. see appendix a for the final codebook. 3. results 3.1 awareness phase 3.1.1 discrimination table 2 provides an overview of the frequency of scores per code per display. table 2 frequency of scores per code about awareness display difference between x y z displays -1 1 -1 1 -1 1 n χ2 p title 0 6 6 0 0 6 18 18.00 < .001 subheading 1 2 2 2 2 6 1 15 2.14 .34 subheading 2 2 2 3 2 6 1 16 1.77 .41 paragraph 1 2 4 1 26 1 18 52 6.32 .04 paragraph 2 0 8 1 29 0 29 67 1.25 .54 paragraph 3 3 5 31 3 23 7 72 11.51 .003 image 10 0 3 0 1 11 25 21.28 < .001 everything 0 26 9 1 2 6 44 31.20 < .001 important words 7 6 2 9 6 9 39 3.23 .20 rereading 4 12 1 9 0 10 36 3.39 .18 total 30 71 59 81 45 98 note. a significant chi-squared test is interpreted as support that teachers discriminated between gaze displays on that code. a total of 55 answers were included for display x, 56 for display y, and 55 for display z. score -1 is used if the teacher mentions that the code is not or rarely looked at, score 1 is used if the teacher mentions that the code is looked at, or often looked at. the main differences between the three displays were in the way the respective students looked at the title, paragraph 3, and the image (see figure 1). these differences were reflected in the remarks of the teachers (see table 2): they discriminated between the three gaze displays for the title, paragraph 3, the image, and whether or not the student looked at everything, but not for subheading 1, subheading 2, paragraph 2, important words, and rereading. surprisingly, teachers also discriminated between the three displays for paragraph 1, whereas the displays showed no obvious difference. this seems to be caused by more teachers mentioning paragraph 1 in displays y and z versus only a few teachers mentioning paragraph 1 in display x. knoop-van campen, kok et al 124 | flr 3.1.2 consistency figure 2 shows the consistency (averaged over three displays) in each of the ‘awareness’ codes. a high consistency indicates that most teachers expressed observing the same viewing behavior. the three most consistent codes were image, paragraph 2, and title. for each of those, almost 100 % of teachers’ remarks were in the same direction. variables in which teachers are least consistent were subheadings 1 and 2, and important words. figure 2. the average consistency within awareness codes. note. average consistency within a code is expressed as the average percentage of teachers who mention the most-mentioned score (either -1 or 1). a higher percentage means more consistency in answers, minimum consistency is 50% (equal number of remarks scores as -1 and as 1). 3.2 interpretation phase 3.2.1 discrimination the main difference between gaze displays was in the reading strategy (see method section). table 3 shows the frequencies with which each of the reading strategies were mentioned for each display. a chi-squared test showed a significant discrimination between displays, χ2(6) = 37.12, p < .001, which is in line with the eye-tracking data: display x shows an intensive strategy, display y shows a selective strategy, and display z showed a borderline intensive strategy. table 3 frequencies for the reading strategies display x y z selective strategy 5 26 19 intensive strategy 11 0 4 unspecified strategy 1 4 0 no strategy 12 4 3 total 29 34 26 knoop-van campen, kok et al 125 | flr apart from those overall reading strategies, teachers made remarks about more specific reading strategies, like the use of headings and signal words as strategies (see table 4). teachers discriminated between displays for the use of signal words but not for the use of headings. next to reading strategies, teachers also wrote down other interpretations of the gaze displays (see table 4). teachers discriminated between displays for confidence and reading time but not for task performance, efficiency, concentration, and motivation. table 4 frequency of scores for each code about interpretation display difference between x y z displays low high low high low high n χ2 p use signal words 7 2 1 8 3 4 25 8.12 .02 use headings 1 2 1 6 2 1 13 2.72 .26 confidence 5 1 1 5 5 0 17 9.70 .008 task performance 0 1 2 3 2 2 10 0.83 .66 reading time* 6 3 1 4 8 0 22 9.09 .01 efficiency 1 0 1 2 2 0 6 3.00 .22 concentration 7 1 6 0 10 1 25 0.76 .68 motivation 1 2 2 0 1 0 6 3.00 .22 total 45 24 45 32 55 12 note. *for reading time, the score low was used if the student was mentioned to be slow, and the score high was used if the student was mentioned to be quick. 3.2.2 consistency figure 3 shows the average consistency within the interpretation codes. consistency was highest for concentration, motivation, and confidence. consistency was lowest (but still relatively high) for task performance, use of headings, and the use of signal words. figure 3. the average consistency within interpretation codes. knoop-van campen, kok et al 126 | flr note. consistency (averaged over the three displays) within a code is expressed as the average percentage of teachers who mention the most-mentioned score (-1 or 1). higher means more consistency in answers. table 3 shows consistency in the remarks regarding reading strategies. display x shows relatively low consistency in codes (41%), in contrast to displays y and z (consistency respectively 76% and 73%), where most teachers remarked that a selective strategy was used. the relative low consistency on display x could be explained by the finding that the score ‘no strategy’ was given to sentences like ‘reads the full text without a clear strategy’ (p017x). in display x, the scores ‘intensive strategy’ and ‘no strategy’ together make up 79% of the remarks. 3.3 enactment phase in the enactment phase, the main contrast was found between teachers who deem no actions necessary and those that would act on the gaze display. additionally, several teachers mentioned that they need more information than just the gaze display. table 5 provides an overview of the scores per display. table 5 frequency of enactment codes x y z no action needed 1 11 3 give compliment 2 1 0 total (no action needed) 3 12 3 instruction-explanation 30 23 30 instruction-modelling 3 0 2 instruction-other 2 3 1 adapt task 16 18 13 other action 1 0 3 total (any action) 40 32 43 need more information from a conversation 5 10 7 need more information on performance 2 1 0 need more information (other) 1 4 3 total (need any additional information) 7 13 8 total main categories 50 57 54 do not know 2 1 1 n 53 55 54 note. four remarks included both that actions would not be needed and that action would be taken (for example: depending on the goal of the learning activity, they would provide instruction or do nothing). those participants were excluded from the table and the chi-squared test to meet the assumption that each participant only contributes to one row. 3.3.1 discrimination a chi-squared test was executed to investigate whether teachers would discriminate between displays regarding either initiating or not initiating action (any action vs. no action needed). the chisquared test showed significant discrimination between the three displays, χ2(2) = 10.610, p = .005. teachers were most likely to act on displays x and z, and less so on display y. knoop-van campen, kok et al 127 | flr 3.3.2 consistency the consistency was lowest for display y, 72% of teachers would act on the displays, whereas for both displays x and z almost all teachers (93%) would act on the displays. 3.4. promises and pitfalls of gaze displays in reading education finally, we asked teachers which general promises and pitfalls they saw in working with gaze displays to support reading instruction. most teachers responded positively (n = 31, 61%) or moderately positive (n = 15, 29%) to the question of whether the gaze displays would help them gather a deeper, better, or other insight into the reading behavior of their students. only five teachers (10%) said displays would not support them. furthermore, most teachers expressed a willingness to use gaze displays if this would be easy to realize. a total of 29 teachers (57%) would use them, 17 (33%) would maybe use them, 3 (6%) would not use them, and 2 teachers (4%) did not answer this question. promises of gaze displays mentioned by teachers were to give them insight into students reading behavior (n = 44, 86%). additionally, part of the teachers said they would use displays to provide students with an insight into their own reading behavior, for example as feedback (n = 18, 35%). next to that, several other applications were mentioned, such as supporting reading motivation and stimulating parent engagement. an important pitfall voiced by teachers was that they lack knowledge about eye tracking and feel that they need more training in gaze display interpretation (n = 11, 22%). teachers also voiced practical concerns such as lack of time and resources (n = 9, 18%). 4. discussion we investigated how teachers interpret and would implement gaze displays in teaching reading comprehension. using the learning analytics process model (verbert et al., 2014) we examined whether teachers discriminated between gaze displays and whether there was consistency among teachers. 4.1 discrimination and consistency in remarks about gaze displays in the awareness phase, the teachers’ pattern of remarks largely mirrors the actual pattern of eye-tracking data. this is reflected in the discrimination between gaze displays that differ and the limited discrimination in codes if gaze displays are minimally different. this suggests that teachers are indeed able to interpret the gaze displays in terms of the coded concepts. inconsistency among teachers seems to stem from expectations of teachers and/or from data-quality issues. for example, whereas most teachers said that paragraph 3 is hardly read, some teachers said that paragraph 3 was read more than expected. e.g., “[the student] looks at section 4 of 5 more intensively, whereas those are not necessary to answer the question” (p004z). thus, whereas remarks coded as awareness might reflect just what can be seen, they are already influenced by the expectations that teachers have about how the student should look at the text. this is in line with the findings of knoop-van campen and molenaar (2020) that teachers’ use of learning analytics information is embedded in their professional routines and expectations. in the interpretation phase, many teachers interpreted gaze displays in terms of reading strategies. they discriminated correctly between displays and they were quite consistent in that. this supports the work of knoop-van campen and colleagues (2021) by showing that teachers can extract reading strategies from gaze displays. more detailed information about reading strategies such as the use of headings and signal words is gathered from the gaze displays, although not very consistently. knoop-van campen, kok et al 128 | flr interestingly, teachers also made inferences about other aspects of students’ performance from the displays. for instance, concentration was quite consistently mentioned, although teachers did not discriminate between displays and generally interpreted the displays as showing low concentration. as we did not have information about students’ concentration, it is not possible to judge the accuracy of these aspects of teachers’ interpretations. it is important to investigate in future research to what extent gaze displays reliably reflect other aspects of study behavior or task performance, such as concentration, confidence, or invested effort and to what extent teachers can reliably pick up on this. on the one hand, it is not unlikely that these aspects of study behaviour can be inferred to a certain extent from specific eye movement patterns (cf. emhardt et al., 2020; rosengrant et al., 2021). on the other hand, we also know from research on teachers’ judgments of students’ performance, that teachers have a strong inclination to use more generic student characteristics in making those judgments (e.g., the motivation or effort they generally display in class), and that they will even ‘fabricate’ student characteristics when they lack information on a student’s identity. that is, when asked to judge students’ text comprehension or math performance based on performance data (i.e., a causal diagram the student filled out about the text), teachers tend to use the information they have about student characteristics, such as motivation, gender, intelligence, or concentration in class, to inform their judgments. however, teachers still do so even when they are presented with anonymized performance data, making statements about students having sloppy habits, having low concentration, being clever or uncertain, even though they did not know who the student was (oudman et al., 2018; van de pol et al., 2021). previous research in the classroom has shown that learning analytics can decrease teachers’ bias towards students (knoop-van campen et al., 2021). if teachers could reliably infer other aspects of study behavior or task performance (such as students concentration or effort on that specific task) from gaze displays, this could support teachers in providing better feedback and guidance on students learning processes. in the enactment phase, most teachers would act on the displayed viewing behavior in some way, although teachers did discriminate between the three displays. consistency was thus relatively high. most actions that teachers mentioned concerned reading strategies. for example, explaining or modelling reading strategies to the student or adapting the task such that other reading strategies would be useful. these results support the notion of xhakaj et al. (2017) that teachers can act on information on a display by altering their instruction, and that this positively affects their teaching (knoop-van campen et al., 2021). apart from that, several teachers mention that the use of gaze displays alone does not provide enough information to act on. whereas teachers showed a certain amount of discrimination between displays and consistency within displays in all three stages, discrimination and consistency are still far from perfect in our sample. however, note that we did not train teachers to interpret gaze displays. most teachers had never seen gaze displays before. this interpretation should thus be considered the lower boundary, and with more information, training, and practice, performance is likely to improve (rienties et al., 2018). 4.2 promises and pitfalls of gaze displays in reading education overall, the teachers were positive about the application of gaze displays in practice. they see many possible applications, but also acknowledge practical problems, even though these might decrease as time passes by and with technological improvements. teachers saw applications of gaze displays both for their monitoring of student’s reading behavior (our intended application), but also considered gaze displays as potential feedback to students, which is an interesting idea that warrants further research (see e.g., henneman et al., 2014, for an example of gaze as feedback in emergency medicine). 4.3 limitations and future research the present study had some limitations and also leads to suggestions for future research. first, interpretation performance is dependent on the display itself (van leeuwen et al., 2019). in this study, we selected displays of data that were not perfect. this is likely to impacte performance, but is also knoop-van campen, kok et al 129 | flr likely to happen in classroom situations, thus providing a realistic reflection of classroom practice. also, the design of the display may have impacted how teachers interpreted it. for example, van wermeskerken and colleagues (2018) found differences between static and dynamic displays regarding participants’ interpretation performance. furthermore, gaze behavior is not always straightforward. for example, displays x and y showed clear examples of the two reading strategies, and indeed, teachers were relatively consistent in describing these strategies. display z showed more of a borderline strategy, which was reflected in a lower consistency in teachers’ remarks. as such, it would be relevant to investigate how characteristics of gaze displays impact teachers’ interpretation performance. secondly, due to the current set-up with open questions, it was not possible to assess the accuracy of most statements that teachers made regarding the eye-tracking data. since the study aimed to unveil teachers’ thoughts about gaze displays (i.e., what information they would extract from a gaze display), we do not have the information from students regarding, for example, their motivation or concentration. in addition, as we asked teachers to elaborate on the presented displays, it was not feasible to use more than a few displays in the present study. future research could seek a more confirmative approach, in which more displays are used that vary systematically over a larger set of measurements (e.g., strategies, reading performance, motivation, confidence) and in which teachers are asked to rate these measures (see for example the work of emhardt et al., 2020). future research could examine whether patterns of viewing behavior teachers distinguish relate to actual student behavior. for instance, teachers interpreted the use of headings and signal words from the gaze displays. think-aloud studies could be used to investigate whether students used those strategies (ericsson & simon, 1993). third, our use of contrasting codes allowed us to make tentative statements about teachers’ judgment accuracy: high consistency combined with high discrimination is necessary if teachers were to validly use gaze displays to act upon. while this approach sensitized us to those contrasts, at the same time it may have diminished our attention to comments that were hard to capture in contrasts. finally, teachers were generally very positive about the gaze displays. however, participants who did not like the displays had probably selected themselves out earlier in the questionnaire. indeed, only 41% of the participants who provided informed consent finished the survey. further research is required to understand the willingness to use gaze displays in a representative sample of the population. 4.4 practical applications this study provides initial steps towards using gaze displays in the classroom to improve students’ reading comprehension, a skill that is central in many school subjects. the findings are promising, showing that even untrained teachers seem able to interpret gaze displays of reading assignments to some extent. training and guidelines could be developed to aid teachers in the interpretation of gaze displays as many teachers mentioned that they would like to know more about how to interpret gaze displays. although future research is needed (e.g., on optimal display design, impact of gaze displays on teachers’ actions in practice, and consecutive on student achievement), our results suggest that in the near future, gaze displays may indeed be used to support teachers in determining which reading strategy a student uses, information that cannot easily or reliably be retrieved another way. although eye tracking technology is at present not yet widely available to schools, this can be expected to change in the near future, as low-cost webcam-based eye tracking solutions are being developed that can be used with regular laptops with camera-function (e.g., madsen et al., 2021) and the quality of such build-in cameras increases swiftly. it is thus certainly plausible that in the near future, students’ eye movements can be tracked during reading tasks (or other learning tasks) and that information on their eye movements can be visualized for teachers and used in educational practices. knoop-van campen, kok et al 130 | flr 4.5 conclusion this study provides first insights into how teachers interpret gaze displays as a form of learning analytics in educational practice. earlier literature showed that gaze displays can be interpreted in terms of the search target, viewing task, preference, or answer to a multiple-choice question (bahle et al., 2017; emhardt et al., 2020; foulsham & lock, 2015; van wermeskerken et al., 2018; zelinsky et al., 2013). our results add to and extend findings from earlier research by showing that gaze displays of reading strategies can be meaningfully interpreted by teachers, which is arguably a substantially more complex task than making inferences about people’s objects or answer preferences from gaze displays. teachers’ remarks about what they saw in the gaze displays (awareness phase) mostly overlapped with the actual pattern of eye-tracking data. we found that teachers interpreted reading strategies in the gaze displays (interpretation phase) quite well, and also interpreted them as reflecting other concepts, such as motivation and concentration. teachers based their initiated actions on the gaze displays (enactment phase). participants were generally positive about using gaze displays in education, both to inform themselves as well as feedback for students, although they mentioned practical issues and a need for more information as areas for improvement. to conclude, gaze displays in reading education seem to be a promising form of learning analytics in education, by providing teachers insight into the reading process of their students, and teachers seem to be willing to incorporate gaze displays in their (future) education. future research should further explore the possibilities of the use of gaze displays in educational practice. keypoints gaze displays are visualizations of viewing behavior, for example showing how students read a text teachers could detect the reading strategy used by the student whose eye movements were displayed quite well teachers inferences about the reading strategies were generally consistent across teachers and discriminated between displays teachers’ remarks about gaze displays provide important information about the feasibility of implementing gaze displays in practice teachers were generally positive about the gaze displays and are willing to use them in their practice as additional information about their students knoop-van campen, kok et al 131 | flr references afflerbach, p., pearson, p. d., & paris, s. g. (2008). clarifying differences between reading skills and reading strategies. the reading teacher, 61(5), 364-373. https://doi.org/https://doi.org/10.1598/rt.61.5.1 almosallam, e. a., & ouertani, h. c. (2013). learning analytics: definitions, applications and related fields. proceedings of the first international conference on advanced data and information engineering (daeng-2013), andreassen, r., jensen, m. s., & bråten, i. (2017). investigating self-regulated study strategies among postsecondary students with and without dyslexia: a diary method study. reading and writing, 30(9), 1891-1916. https://doi.org/10.1007/s11145-017-9758-9 bahle, b., mills, m., & dodd, m. d. (2017). human classifier: observers can deduce task solely from eye movements. attention perception & psychophysics, 79(5), 1415-1425. https://doi.org/10.3758/s13414-017-1324-7 benfatto, m. n., öqvist seimyr, g., ygge, j., pansell, t., rydberg, a., & jacobson, c. (2016). screening for dyslexia using eye tracking during reading. plos one, 11(12), e0165508. https://doi.org/10.1371/journal.pone.0165508 böhm, m. (2021, september 14). lesediagnostik und leseförderung. https://www.lesediagnostik.de/ chisari, l. b., mockevičiūtė, a., ruitenburg, s. k., van vemde, l., kok, e. m., & van gog, t. (2020). effects of prior knowledge and joint attention on learning from eye movement modelling examples. journal of computer assisted learning, 36(4), 569-579. https://doi.org/10.1111/jcal.12428 de koning, b. b., & jarodzka, h. (2017). attention guidance strategies for supporting learning from dynamic visualizations. in r. lowe & r. ploetzner (eds.), learning from dynamic visualization: innovations in research and application (pp. 255-278). springer international publishing. https://doi.org/10.1007/978-3-319-56204-9_11 donovan, t., manning, d. j., & crawford, t. (2008). performance changes in lung nodule detection following perceptual feedback of eye movements. proceedings of spie the international society for optical engineering in san diego, ca, http://dx.doi.org/10.1117/12.768503 duffy, g. g. (2002). the case for direct explanation of strategies. in c. c. block & m. pressley (eds.), comprehension instruction: research-based best practices (pp. 28-41). guilford. duke, n. k., & pearson, p. d. (2009). effective practices for developing reading comprehension. journal of education, 189(1-2), 107-122. https://doi.org/10.1177/0022057409189001-208 eder, t. f., richter, j., scheiter, k., keutel, c., castner, n., kasneci, e., & huettig, f. (2020). how to support dental students in reading radiographs: effects of a gaze‑based compare‑and‑contrast intervention. advances in health sciences education: theory and practice. https://doi.org/10.1007/s10459-020-09975-w emhardt, s. n., van wermeskerken, m., scheiter, k., & van gog, t. (2020). inferring task performance and confidence from displays of eye movements. applied cognitive psychology, 34(6), 1430-1443. https://doi.org/https://doi.org/10.1002/acp.3721 ericsson, k. a., & simon, h. a. (1993). protocol analysis: verbal reports as data. mit press. http://books.google.nl/books?id=z4hqqgaacaaj foulsham, t., & lock, m. (2015). how the eyes tell lies: social gaze during a preference task. cognitive science, 39(7), 1704-1726. https://doi.org/10.1111/cogs.12211 gambrell, l. b., malloy, j. a., & mazzoni, s. a. (2011). evidence-based best practices for comprehensive literacy instruction. in l. m. morrow & l. b. gambrell (eds.), best practices in literacy instruction (fourth edition) (vol. 4, pp. 11-56). the guilford press. guthrie, j. t., & mosenthal, p. (1987). literacy as multidimensional: locating information and reading comprehension. educational psychologist, 22(3-4), 279-297. https://doi.org/10.1080/00461520.1987.9653053 hayes, a. f., & krippendorff, k. (2007). answering the call for a standard reliability measure for coding data. communication methods and measures, 1(1), 77-89. https://doi.org/10.1080/19312450709336664 knoop-van campen, kok et al 132 | flr henneman, e. a., cunningham, h., fisher, d. l., plotkin, k., nathanson, b. h., roche, j. p., marquard, j. l., reilly, c. a., & henneman, p. l. (2014). eye tracking as a debriefing mechanism in the simulated setting improves patient safety practices. dimensions of critical care nursing, 33(3), 129135. https://doi.org/10.1097/dcc.0000000000000041 jarodzka, h., van gog, t., dorr, m., scheiter, k., & gerjets, p. (2013). learning to see: guiding students' attention via a model's eye movements fosters learning. learning and instruction, 25(0), 62-70. https://doi.org/10.1016/j.learninstruc.2012.11.004 jivet, i., scheffel, m., drachsler, h., & specht, m. (2017). awareness is not enough: pitfalls of learning analytics dashboards in the educational practice. european conference on technology enhanced learning, tallinn, estonia. katalayi, g. b., & sivasubramaniam, s. (2013). careful reading versus expeditious reading: investigating the construct validity of a multiple-choice reading test. theory and practice in language studies, 3, 877-884. https://doi.org/10.4304/tpls.3.6.877-884 knoop-van campen, c. a. n., doest., d. ter, verhoeven, l., & segers, e. (2021). the effect of audiosupport on strategy, time, and performance on reading comprehension in secondary school students with dyslexia. annals of dyslexia, 1-20 knoop-van campen, c. a. n. & molenaar, i. (2020). how teachers integrate dashboards into their feedback practices. frontline learning research, 8(4), 37-51. https://doi.org/10.14786/flr.v8i4.641 knoop-van campen, c. a. n., wise, a., & molenaar, i. (2021). the equalizing effect of teacher dashboards on feedback in k-12 classrooms. interactive learning environments, 1-17. https://doi.org/10.1080/10494820.2021.1931346 kok, e. m., aizenman, a. m., võ, m. l.-h., & wolfe, j. m. (2017). even if i showed you where you looked, remembering where you just looked is hard. journal of vision, 17(12), 1-11. https://doi.org/10.1167/17.12.2 kostons, d., van gog, t., & paas, f. (2009). how do i do? investigating effects of expertise and performance-process records on self-assessment. applied cognitive psychology, 23(9), 1256–1265. https://doi.org/10.1002/acp.1528 krishnan, k. s. d. (2011). careful versus expeditious reading: the case of the ielts reading test. academic research international, 1(3), 25. liu, f. (2010). reading abilities and strategies: a short introduction. international education studies, 3(3), 153-157. https://doi.org/10.5539/ies.v3n3p153 madsen, j., julio, s. u., gucik, p. j., steinberg, r., & parra, l. c. (2021). synchronized eye movements predict test scores in online video education. proceedings of the national academy of sciences, 118(5). https://doi.org/10.1073/pnas.2016980118 mason, l., pluchino, p., & tornatora, m. c. (2015). eye-movement modeling of integrative reading of an illustrated text: effects on processing and learning. contemporary educational psychology, 41, 172187. https://doi.org/10.1016/j.cedpsych.2015.01.004 molenaar, i., & knoop-van campen, c. (2019). how teachers make dashboard information actionable. ieee transactions on learning technologies, 12(3), 347-355. https://doi.org/10.1109/tlt.2018.2851585 murnane, r., sawhill, i., & snow, c. (2012). literacy challenges for the twenty-first century: introducing the issue. the future of children, 3-15. https://doi.org/10.1353/foc.2012.0013 oudman, s., van de pol, j., bakker, a., moerbeek, m., & van gog, t. (2018). effects of different cue types on the accuracy of primary school teachers' judgments of students' mathematical understanding. teaching and teacher education, 76, 214-226. https://doi.org/https://doi.org/10.1016/j.tate.2018.02.007 pikulski, j. j., & chard, d. j. (2005). fluency: bridge between decoding and reading comprehension. the reading teacher, 58(6), 510-519. https://doi.org/10.1598/rt.58.6.2 rienties, b., herodotou, c., olney, t., schencks, m., & boroowa, a. (2018). making sense of learning analytics dashboards: a technology acceptance perspective of 95 teachers. international review of research in open and distributed learning, 19(5). https://doi.org/10.19173/irrodl.v19i5.3493 knoop-van campen, kok et al 133 | flr rosengrant, d., hearrington, d., & o’brien, j. (2021). investigating student sustained attention in a guided inquiry lecture course using an eye tracker. educational psychology review, 33(1), 1126. https://doi.org/10.1007/s10648-020-09540-2 salmerón, l., naumann, j., garcía, v., & fajardo, i. (2017). scanning and deep processing of information in hypertext: an eye tracking and cued retrospective think‐aloud study. journal of computer assisted learning, 33(3), 222-233. https://doi.org/10.1111/jcal.12152 scheiter, k., schubert, c., & schüler, a. (2018). self-regulated learning from illustrated text: eye movement modelling to support use and regulation of cognitive processes during learning from multimedia. british journal of educational psychology, 88(1), 80-94. https://doi.org/doi:10.1111/bjep.12175 schutz, k. m., & rainey, e. c. (2020). making sense of modeling in elementary literacy instruction. the reading teacher, 73(4), 443-451. https://doi.org/10.1002/trtr.1863 shute, v. j., & zapata-rivera, d. (2012). adaptive educational systems. in p. j. durlach & a. lesgold (eds.), adaptive technologies for training and education (vol. 7, pp. 1-35). cambridge university press. https://doi.org/10.1017/cbo9781139049580.004 špakov, o., siirtola, h., istance, h., & räihä, k. (2017). visualizing the reading activity of people learning to read. journal of eye movement research, 10(5), 1-12. https://doi.org/10.16910/jemr.10.5.5 tijmstra, j., & boeije, h. (2011). wetenschapsfilosofie in de context van de sociale wetenschappen. boom lemma. urquhart, s., & weir, c. (1998). reading in a second language: process, product and practice.routledge. van de pol, j., de bruin, a. b. h., van loon, m. h., & van gog, t. (2019). students’ and teachers’ monitoring and regulation of students’ text comprehension: effects of comprehension cue availability. contemporary educational psychology, 56, 236-249. https://doi.org/10.1016/j.cedpsych.2019.02.001 van de pol, j., van den boom-muilenburg, s. n., & van gog, t. (2021). exploring the relations between teachers’ cue-utilization, monitoring and regulation of students’ text learning. metacognition and learning, 1-31. van gog, t., & jarodzka, h. (2013). eye tracking as a tool to study and enhance cognitive and metacognitive processes in computer-based learning environments. in r. azevedo & v. aleven (eds.), international handbook of metacognition and learning technologies (pp. 143-156). springer science+ business media. van gog, t., jarodzka, h., scheiter, k., gerjets, p., & paas, f. (2009). attention guidance during example study via the model’s eye movements. computers in human behavior, 25(3), 785-791. https://doi.org/10.1016/j.chb.2009.02.007 van leeuwen, a. van, knoop-van campen, c. a. n., molenaar, i., & rummel, n. (2021). how teacher characteristics relate to how teachers use dashboards: results from two case studies in k-12. journal of learning analytics, 8(2), 6-21. https://doi.org/10.18608/jla.2021.7325 van leeuwen, a., rummel, n., & van gog, t. (2019). what information should cscl teacher dashboards provide to help teachers interpret cscl situations? international journal of computersupported collaborative learning, 14(3), 261-289. https://doi.org/10.1007/s11412-019-09299-x van wermeskerken, m., litchfield, d., & van gog, t. (2018). what am i looking at? interpreting dynamic and static gaze displays. cognitive science, 42(1), 220-252. https://doi.org/10.1111/cogs.12484 verbert, k., govaerts, s., duval, e., santos, j. l., van assche, f., parra, g., & klerkx, j. (2014). learning dashboards: an overview and future research opportunities. personal and ubiquitous computing, 18(6), 1499-1514. https://doi.org/10.1007/s00779-013-0751-2 võ, m. l. h., aizenman, a. m., & wolfe, j. m. (2016). you think you know where you looked? you better look again. journal of experimental psychology: human perception and performance, 42(10), 1477-1481. https://doi.org/10.1037/xhp0000264 knoop-van campen, kok et al 134 | flr xhakaj, f., aleven, v., & mclaren, b. m. (2017). effects of a teacher dashboard for an intelligent tutoring system on teacher knowledge, lesson planning, lessons and student learning. european conference on technology enhanced learning, zelinsky, g. j., peng, y., & samaras, d. (2013). eye can read your mind: decoding gaze fixations to reveal categorical search targets. journal of vision, 13(14), 1-13. https://doi.org/10.1167/13.14.10 knoop-van campen, kok et al 135 | flr appendix a: codebook and detailed overview coding process for awareness, we coded whether teachers mentioned that the student looked at or paid attention to the separate aois (i.e., title, subheadings 1 and 2, section 1, 2, 3, and the image). a score of 1 was given if the teacher mentioned that the student looked (a lot) at this part of the stimulus, and a score of -1 was given if the teacher mentions that the student did not look at this part of the stimulus (also if the teacher mentioned ‘missing’, ‘looking only a little bit at’). in addition to the aois, we also coded looking at ‘everything’, ‘important words/sentences’, ‘rereading’ with scores 1 and -1. if a code was not mentioned or if, in rare cases, both scores could be applied, the score 0 (missing value) was applied. a similar approach was used for interpretation, but here 1 refers to high (e.g., high performance) and -1 refers to low (e.g., low confidence). these scores were used for ‘confidence’, ‘performance’, ‘reading time’, ‘efficiency’, ‘use of signal words’, ‘use of headings’, ‘concentration’, and ‘motivation’. note that for reading time, score 1 refers to quick reading and score -1 refers to slow reading. finally, the type of strategy was scored as intensive, selective, general (any strategy, but unclear which strategy), or no strategy. for enactment, the form of contrast described above was not possible, so for those codes, only 1 (code mentioned) or 0 (code not mentioned) were scored. the codes were ‘no action is needed’, ‘give a compliment’, ‘provide instruction (in the form of an explanation / modelling / another form)’, ‘adapt the assignment’, ‘need more information (from a conversation with student / the performance / other information)’, ‘i do not know how to act on this’, ‘other actions’. to make contrasting codes, we additionally recoded codes into ‘any action’ (any form of enactment mentioned) versus ‘no action’ (teacher writes that no action is needed or only the action ‘give a compliment’). instructions awareness, interpretation, and enactment code the three questions together, i.e., codes for awareness can also be scored if this information is provided in the second question. code what is actually said and avoid making inferences, i.e., do not code ‘everything’ if a participant summed up all different parts, but only if the person actually said that ‘everything’ was read. for awareness and interpretation, two scores are available: -1 and 1. score 0 if this code is not mentioned (missing value). score -1 if the negative form (see definitions) is mentioned and 1 if the positive form (see definitions) is mentioned. if something is mentioned, but it is unclear whether -1 or 1 should be coded, score as 0. for enactment, only 0 (not mentioned) and 1 (mentioned) are possible. if a person refers back to a previous answer, simply code that as everything not mentioned (as it is impossible to know which aspects of the other answer the person refers to). all codes are identical for displays x, y, and z. knoop-van campen, kok et al 136 | flr codebook actual quotes as examples are italicized. translations to dutch can be found in brackets. awareness codes that answer: what do i see? codes under this heading are about the stimulus and gaze behavior without a reference to the cognitive processes of the student. words like “looking at” [kijken naar], “pay attention to” [aandacht hebben voor/letten op] should be coded as awareness because they do not refer to the cognitive conscious processing. 1: score 1 if the teacher mentions looking (a lot) at this part of the display. -1: score -1 if the teacher mentions not looking at this part of the display (also: missing [missen], not looking at [niet kijken naar], looking only a little bit at [weinig kijken naar] 0: score 0 if the part of the display is not mentioned (missing). code explanation example 1 example 1 title (trucks without drivers) looking at the title yes/no student looked at the title looks at/spends (a lot of) attention on the header student ignored/did not look at title does not look at/spend little or no attention on the header subheading 1 looking at the first subheading (note. when "looking at the subheadings" both subheading 1 and subheading 2 are coded). looks at/spends (a lot of) attention on the first subheading does not look at/spend little or no attention on the first subheading subheading 2 looking at the second subheading (note. when "looking at the subheadings" both subheading 1 and subheading 2 are coded). looks at/spends (a lot of) attention on the second subheading does not look at/spend little or no attention on the second subheading first paragraph looking at the first paragraph. explicitly mentioned as such, or if it cannot interpreted otherwise looks at/spends (a lot of) attention on the first paragraph does not look at/spend little or no attention on the first paragraph. skips part/section in this paragraph. second paragraph (contains answer) looking at the second paragraph. explicitly mentioned as such, or if it cannot interpreted otherwise looks at/spends (a lot of) attention on the second paragraph does not look at/spend little or no attention on the second paragraph. skips part/section in this paragraph. third paragraph (and/or ‘final) looking at the third paragraph. explicitly mentioned as such, or if it cannot interpreted otherwise looks at/spends (a lot of) attention on the third/last paragraph. does not look at/spend little or no attention on the third paragraph. skips part/section in this paragraph. image looking at the picture. looks at/spends (a lot of) attention on the picture does not look at/spend little or no attention on the picture. everything evenly read. read the entire text. only code if it is explicitly mentioned that all/not all was read reads everything/ looks at everything reads only part of the text/ does not read everything knoop-van campen, kok et al 137 | flr (not if you can infer it by information about the individual paragraphs). important words/phrases & difficult words when mentioned looking at/ paying attention to important words, important phrases, difficult words or difficult phrases, and keywords or keyphrases looks at/spends (a lot of) attention on the keywords, focus on important words. more circles at keywords. the largest circles are at keywords (driver, highway, safety, stop, bill, self-driving trucks). does not look at/spend little or no attention on keywords, important words. no (extra) circles at keywords. rereading rereading versus no rereading reads back a lot/ rereading little rereading interpretation sensemaking, what are the analytics telling me? interpretation involves the link to the learner's cognitive process, conscious actions, processing, and interpretation. it also includes judgments from the teacher. 1: they refer to high (e.g., high performance) -1: remarks refer to low (e.g., low confidence) 0: the interpretation is not mentioned (missing). code explanation example 1 example 1 confidence remarks on the student's selfconfidence. feels that he has understood the text after paragraph 2. uncertain perhaps. performance comments on student performance/ response/ efficacy. only addresses the answer/reading comprehension outcome. the student got it right / has the right answer. ineffective reading. time remarks on the amount of time the student spent on the text. caution: remarks about (in)efficiency should not be coded here but only be coded under ‘efficiency’. quick reader/ fast slow reader/ slow the student spends a very long time on the text efficiency remarks about the learners’ efficiency (combination of time and performance). only code if the teacher writes: efficiently. very efficient read/ efficient reading behavior not efficient read/ inefficient reading behavior type of strategy here, the type of strategy is coded as intensive or searching. use score 0 if the type of strategy is not mentioned, and see next code if a strategy is mentioned, but it’s not clear which strategy. intensive strategy: everything/intensive reading (do not code here if the behavior is descriptive, then code as awareness). this student reads the entire text accurately. searching/selective strategy: reads only part of the text. relevant synonyms are: searching/ global/ scanning/ reading headings/ focusing/ directed. the student is searching for the answer. knoop-van campen, kok et al 138 | flr general strategy versus no strategy this code can only be used if the type of strategy does not indicate which strategy was used or when no strategy was applied / it does not appear that the students had a deliberate approach. plan, strategy, approach. this score cannot co-occur with a -1 or 1 on type or strategy. “structure/structure element” alone is not sufficient to code as a strategy he reads strategically. the student has read analytically. that the student is not reading in a very structured way / the student needs to read in a more structured way. he just does something, instead of following a reading plan / student does not use a reading strategy / apparently, the student is not searching in a focused way based on a thought process but is searching randomly in the text. use of signal/key / important words remarks on the use of signal words. note that this is about using/ processing/ seeking, and therefore not about seeing/noticing/look at (= awareness). what i see is that for keywords, the student has looked back in the text to see what those keywords refer to. after reading the introduction, this student was not guided by the signal words use of (sub)headings remarks on the use of headings and/or title. note that this is about using/ processing/ seeking, and therefore not about seeing/noticing/look at (= awareness). has (presumably) used the subheads well. doesn't use headings to make sense of the text. concentration remarks on the student's concentration. the opposite that also belongs under this code is 'being distracted'. note: ‘attention wanes' depends on context. that can also belong to awareness (e.g., when teachers mentioned that the last paragraph was not read due to lack of attention). the student is very focused/concentrated. he seems distracted, struggling to stay on task / thoughtless. attention wanes. motivation remarks on being motivated. student is motivated/ interested/ appears interested. the student is not motivated, drops out, has no interest. enactment translate interpretation into pedagogical actions. what would the teacher (want to) do now, after seeing the display? 1: code is mentioned 0: code is not mentioned code explanation no action needed the teacher sees no need to act upon the gaze display. e.g., "he did fine, i don't need to do anything". compliment the teacher indicates that he/she wants to compliment the student/ tell the student that something was done well. e.g., "i want to compliment him for focussing on the second paragraph". knoop-van campen, kok et al 139 | flr instruction: explanation the teacher wants to provide instruction, explanation, clarification, support, guidance to students on addressing (similar) texts. e.g., "i would explain that here he had better looked at the headings first and only then at the text". "i would especially emphasize reading headings". instruction: modelling the teacher wants to model (demonstrate) the text/comparative text to the student. e.g., "i would show how i / to approach a text," or "show how he can scan and then read the important part." instruction: other other instructions/actions toward the learner that are not included in explanation or modeling. adapt assignment remarks that they would modify the assignment. e.g., "give an easier text" / "from now on give the questions afterwards and not before the text". this code can be given when the current assignment should be adjusted according to the teacher, or when a follow-up assignment is given. it also includes giving instruction to collaborate and having the student do something different (like reading out loud) need more information: conversation with student the teacher indicates that he wants more information before he can take an action. the teacher wants to talk to the student first, ask what/why a student did something. e.g., "discuss with the student to find out how he handled this one" or "i would like to know more about what the student thought about it first". e.g., “asking student if ...” need more information: need performance score the teacher indicates that he would like more information before taking any action. the teacher first wants to know if a student gave the right answer, and depending on that information decide to do something. e.g., "what did he score? that's important to know in combination with how he handled it. otherwise i can't give proper feedback". need more information: other the teacher indicates that he would like more information before taking any action. this information does not include talking to the student or checking the answer/score given. e.g., is there a diagnosis of learning/ reading problems? don’t know when a teacher indicates that he does not know what action he would like or could take after seeing the gaze display. e.g., "i have no idea." other actions other actions a teacher would do based on the gaze display that do not fit into the other categories. remarks that are too vague or too short to categorize with certainty under another code. e.g. i could possibly use parent support or a class assistant. promises and pitfalls code explanation scores usefulness answer to the question: “do you feel that gaze display of your students would help you as a teacher to have a deeper, better, or other insight into the reading behavior of your students?” negative: explicit ‘no’ or if only pitfalls are mentioned moderately positive: explicit ‘maybe’ or if both pitfalls and promises are mentioned positive: explicit ‘yes’, or if only promises are mentioned use answer to the question: “would you use gaze displays of your students (in the future) if this would be easy to realize?” negative: explicit ‘no’ or if only pitfalls are mentioned moderately positive: explicit ‘maybe’ or if both pitfalls and promises are mentioned positive: explicit ‘yes’, or if only promises knoop-van campen, kok et al 140 | flr are mentioned promise information for teachers teachers see the promise of gaze displays for informing the teacher (they keep the information and act on it) 1: mentioned 0: not mentioned promise information for student teachers see the promise of gaze displays for informing the student (share the information, e.g., as feedback). 1: mentioned 0: not mentioned pitfall knowledge about eye tracking teachers see their lack of knowledge of eye tracking as a pitfall or would only implement this if they receive more information about gaze displays. 1: mentioned 0: not mentioned pitfall practical problems teachers see practical problems as a pitfall, for example that a lot of time or money needs to be invested. 1: mentioned 0: not mentioned frontline learning research vol. 12 no. 1 (2024) 49 65 issn 2295-3159 corresponding author: oscar van den wijngaard, edlab maastricht university centre for teaching & learning, maastricht university, po box 616, 6200 md maastricht, the netherlands. oscarvandenwijngaard@maastrichtuniversity.nl doi: https://doi.org/10.14786/flr.v12i1.951 new perspectives on civic engagement as an outcome of higher education: an exploratory case study oscar van den wijngaard, simon beausaert, wim gijselaers & mien segers maastricht university, the netherlands article received 15 september 2021 / article revised 14 august 2023 / accepted 21 february 2024/ available online 27 february abstract this study explores the potential of a new perspective on research into the impact of higher education on students' civic engagement. we propose shifting from viewing engagement as the key dependent variable to two 'fundamental constituents': political interest and agency. both constituents have been presented as either static or determined entirely by factors external to education, such as maturation, but also as dynamic and affected by various aspects of the educational experience in higher education. furthermore, as analyses of these effects based on sample means do not account sufficiently for the intersectionality of background variables that define the student experience, we propose that data are explored through cluster analysis. employing this type of analysis, a case study conducted at a small international liberal arts college in the netherlands showed four distinctly different patterns in the development of both constituents of civic engagement. based on further data obtained from the same sample, we offer suggestions for specific foci in further research about the impact of higher education on the development of civic engagement. keywords: civic engagement; higher education; cluster analysis mailto:oscarvandenwijngaard@maastrichtuniversity.nl https://doi.org/10.14786/flr.v12i1.951 van den wijngaard et al. 50 | f l r 1. introduction a recurring theme in the discourse on higher education is how one of its main contributions to society is to instil in students a sense of civic engagement. over the last few decades, educators, (former) administrators, and philosophers have reflected on this specific aspect of higher education’s mission. more often than not, they conclude the university had neglected its role in fostering such engagement, while many also provided broad or specific suggestions as to how the university could resume its responsibility (e.g., aronowitz, 2000; barnett 1990, 2000; bok 2005, 2008; boyer, 1990; boyte, 2008; boyte & finders, 2016; checkoway, 2001; colby et al., 2003; hartley et al., 2010; lewis, 2007; nussbaum, 2010; readings 1997). stating that civic engagement is considered an essential outcome of higher education raises the question of whether it is something that can be learned. understanding the potential role education can play in developing civic engagement implies understanding it as a learning outcome, regardless of its specific conceptualisation. does civic engagement evolve and change, or is it innate, predetermined, and static? furthermore, if it evolves, does it do so in ways similar for everyone (e.g., as a maturation effect) or more varied, in which case changes could be attributed to factors that are not uniform but determined by specific contexts or interventions, including educational ones? if civic engagement cannot be furthered by education, there would be little sense in making it a desired learning outcome, and it would be beside the point to judge education by its (lack of) effectiveness in helping students attain it. with this study, we aim to contribute to the discourse on these issues, which can be summarised as the question about the nature of civic engagement and the role higher education might play in fostering it. we combine exploring the literature on civic engagement as an outcome of higher education and a case study conducted at a liberal arts college in the netherlands to propose two related, new perspectives. first, we suggest that the impact of higher education on civic engagement can be understood better when more attention is given to the development of fundamental constituents of civic engagement. in this study, we present arguments for political interest and agency as such fundamental constituents. secondly, we propose that various intersecting background variables affect this development in diverse ways in individuals. to capture this effect, quantitative research into the impact of education on engagement should use cluster analyses rather than sample means. 2. research review: civic engagement as an outcome of learning 2.1 conceptualisations of civic engagement civic engagement is often understood as an expression of social responsibility and active participation in the context of politics and established democratic structures and initiatives. examples are membership in social organisations or volunteer work (dee, 2004; egerton, 2002; helliwell & putnam, 2007; hillygus, 2005; hylton, 2018; yang & hoskins, 2020; milligan et al., 2004; myers et al., 2019), or more broadly “prosocial and political contributions to community and society” (wray-lake & schubert, 2019, p. 2169). in recent years, the terminology within this discourse has shifted from ‘civic’ or ‘social’ engagement to concepts like global citizenship education or gce (palmer, 2018) and ‘civic agency’ (boyte, 2008; boyte & finders, 2016). in the case of gce, palmer (2018, p. 135) concludes that while there are varying interpretations of the concept, these all have in common that they rely on 'interrelation, inclusivity, curiosity, creativity, and criticality.' boyte’s conceptualisation of ‘civic agency’ takes engagement beyond participation in political structures and communities, as it stresses the role citizens can play in actively shaping those structures and communities (boyte, 2008; boyte & finders, 2016). van den wijngaard et al. 51 | f l r 2.2 civic engagement as an outcome of higher education as the conceptualisation of civic engagement evolved, so did research into civic engagement as an outcome of learning. this section explores a range of studies conducted between 1987 and 2020. socioeconomic background and maturation figure prominently and frequently as explanations of the extent and development of civic engagement, as do years in school and specific courses followed. the various interactions between these potential determinants make research into the impact of education on civic engagement a delicate affair. pascarella and terenzini (2005) refer to the maturation problem when they review the research on attitudes and values as outcomes of education. in that same publication, they stress the importance of distinguishing between effects during college and due to college. they issued the same caveat more than two decades earlier in the 1991 article summarising the educational research observations they made while writing the first edition of their meta-study (pascarella & terenzini, 1991). the case against the effect of education on civic engagement has been put forward in studies that suggest an absence of any notable change. these studies primarily identify background variables as the main determinants of one's level of civic engagement. in what was essentially a comparative crosssectional study, looking at first-year students, juniors, and recent graduates in three different medical schools in michigan (u.s.), maheux and beland (1987) found little effect of education nor maturation on attitudes concerning medical-ethical issues in students attending medical school. they concluded that differences in attitudes were most likely a function of selection. milner et al. (1999) surveyed students in the second, third, and final year of their undergraduate studies, presenting them with ethical issues and asking them to score their attitudes toward these on a 1-10 likert scale. looking at students in business education and comparing these to a control group of non-business students, they found no significant effects on their moral development of education nor maturation, only of starting position in terms of their attitudes towards moral issues (milner et al., 1999). in both studies, no significant change was found in values or attitudes. effects based on social background and selection play an important role in longitudinal and crosssectional studies comparing groups based on attainment, defined as time spent in higher education. social background and (self) selection seem to be related to holland’s theory of congruence, developed further by, e.g., smart, feldman, and ethington (2000), suggesting that students seek out for themselves an environment that they expect to be congruent with their values and interests. a similar explanation comes from egerton (2002), who found that differences in civic engagement, in the form of membership, active or not, of a ‘civic organisation’, exist before entering higher education and change little, with only minor positive effects of higher education within specific groups. a wide array of studies exist on the other side of the argument. these differ in the depth of analysis in establishing what it is about education that makes it an essential factor in developing civic engagement. the broadest observations concern the overall time spent in education without further specification. milligan, moretti, and oreopoulos (2004) compared the impact of educational attainment, in terms of years spent in school, particularly on voting behaviour, between the u.s. and the u.k. for their study, milligan et al. (2004) used data gathered through the national election studies, covering 1948-2000, and the november voting supplement to the current populations surveys from 1978-2000 for the u.s. for the u.k., they combined data from the british election studies for elections between 1964 and 1997 and eurobarometer surveys between 1973 and 1998. they found that educational attainment correlates significantly with voting behaviour in the u.s. but not in the u.k. even without providing more insight into what it is about education that seems to have a positive impact on this specific aspect of civic engagement, the fact that attainment in years did have an effect, at least in the u.s., goes against the idea that the development of a stronger inclination to vote would merely be a matter of maturation. milligan et al. (2004) found the same correlation between educational attainment and the extent to which citizens would be interested in public affairs and politics in both countries. focusing on the u.s. only, dee (2004) looked at the effects of education on voter and volunteer participation and 'civic awareness'. dee drew conclusions similar to those of milligan et al. (2014) based van den wijngaard et al. 52 | f l r on data from the u.s. department of education's longitudinal high school and beyond study, including questions on civic engagement and educational attainment. dee concluded that “[…] educational attainment, both at the post-secondary and secondary levels, has large and independent effects on most measures of civic engagement and attitudes” (dee 2004, p. 1717). helliwell and putnam (2007) defined ‘social capital’ as a combination of the level of trust in fellow citizens and indicators of civic engagement such as membership in clubs and participation in community activities. they found that time spent in higher education has positive relative and absolute effects on social capital (helliwell & putnam, 2007). other studies venture into finding specific aspects of education that may play a role in promoting civic engagement. astin (1997) employed longitudinal data from the cooperative institutional research program (cirp) as input for analysis based on his input-environment-output model. he found most changes in socio-political attitudes attributable to social change, i.e., societal trends and the influence of peers and faculty (astin, 1997). the latter was also found by yang and hoskins (2020), who used data from the citizenship education longitudinal study conducted in the u.k. between 2009 and 2014. they concluded that socialisation through interaction with peers and staff positively affected political participation (such as voting). however, political activism (protest) was not affected significantly, and volunteerism was negatively affected (yang & hoskins, 2020). laird (2005), building on research done by gurin et al. (2002), assessed the impact of diversity and diversity-related programmes on three outcomes related to civic engagement, one of which being social agency, and found significant changes attributable to diversity. the studies by laird (2005) and gurin et al. (2002) were both survey-based and focused on the impact of diversity as an aspect of the opportunity structure provided by the higher education setting. students were asked to indicate which level and type of interactional diversity they had experienced and to respond to items related to scales measuring, e.g., ‘citizenship engagement’, ‘perspective taking’, and ‘active thinking’ (gurin et al., 2002), and ‘social agency’ and ‘critical thinking’ (laird, 2005). hillygus (2005) used longitudinal data from the baccalaureate and beyond longitudinal study, collected during interviews one and four years after graduation, with 9274 students who obtained their bachelor’s degree in the academic year 19921993. hillygus (2005) operationalised civic engagement as voting and other forms of political participation (e.g., attending rallies, volunteering for community action groups, or supporting political campaigns). they looked at the effect of pre-university variables such as parental education, high school test results, and aspects of the curriculum. she found positive significant effects on voting and political participation of curricula that stimulate the development of verbal skills and courses in social sciences (hillygus, 2005). myers et al. (2019) also found a specific impact of intentional, designed aspects of the education experience. they used astin’s input-environment-output model in their analysis of the 20022012 education longitudinal study, following a cohort of u.s. 10th graders starting in 2002. this study focused on the long-term effects of six high impact practices (hip) on civic engagement and suggests a causal relationship between participation in four of them (mentoring, research with a faculty member, internships, and community-based activities) and higher civic engagement in adult life, with a more substantial effect in students entering college with lower civic orientations (myers et al., 2019). at the intersection of background variables and educational experience, we find a series of studies by wray-lake et al. (2014), wray-lake and schubert (2019), and wray-lake, arruda, and schulenberg (2020). wray-lake et al. (2014) used data from the longitudinal study of american youth, which followed adolescents from the age of 13-14 to age 18-19 in school and phone surveyed the same cohort one year after high school and again as adults twenty years after the first measurement. employing a typology of four classes, effectively representing levels of engagement, they found that two-thirds of the subjects remained within their class and that the remaining third displayed upward as well as downward mobility, which "provided only modest support for age-related gains," nor overwhelming evidence for strong homogenous effects of education (wray-lake et al., 2014, p. 95). this general image was nuanced in two later studies. wray-lake and schubert (2019), using the same data as wray-lake et al. (2014), saw gender, race/ethnicity, and parent education as significant determinants of specific levels (or types) of civic engagement in adolescents. as in the 2014 study, there was little mobility in types of engagement, but based on – in this case – the significant role of having 'civic discussions' with van den wijngaard et al. 53 | f l r parents (and, to a lesser extent, friends), they also found that civic engagement may need to be nurtured to maintain a certain level. furthermore, in their 2020 study, based on the u.s. monitoring the future data, wray-lake et al. observed how trajectories in the development of levels and types of civic engagement diverged from general patterns associated with the 'transition to adolescence' due to the influence of race/ethnicity, parent education, and gender, and their intersections (wray-lake et al., 2020). the intersectionality between various background variables and their interaction with environmental context is acknowledged in all of the studies presented above and addressed explicitly by wray-lake et al. (2014), wray-lake and schubert (2019), and wray-lake, arruda, and schulenberg (2020). their studies imply that for a better understanding of the impact of education on civic engagement, we need to look beyond effects at the level of an entire sample and instead zoom in on specific groups within one sample. lerner et al. (2014) used longitudinal data obtained in their 4-h study, which followed young american adolescents from grades 5 through 12 (ages 10-18), to model their development in terms of civic engagement. in this model on 'positive youth development,' political interest and agency are subsumed as two elements under one concept of 'positive and active civic engagement,' for which positive youth development serves as a precondition or a moderator for socioeconomic background variables and attitudes. their model presents a relational developmental systems perspective, which revolves around the dynamic between individual and context (lerner et al., 2014), not unlike the watts and flanagan (2007) model of youth development within opportunity structures. similar to wray-lake et al. (2014), several authors see education as just one context amongst many (family, friends, place) that all affect the development of civic engagement in youth (wray-lake & schubert, 2019; wray-lake et al., 2020; lerner et al., 2014). the intricate interplay between background variables leads lerner et al. (2014) to conclude that a better understanding of the development of civic engagement in youth would benefit from a shift in the research from group to individual and from variable to class. 2.3 fundamental constituents as common characteristics across conceptualisations of civic engagement the conceptualisation of civic engagement (and how it is understood as an outcome of higher education) is multifaceted. it has evolved from forms of participation such as voter behaviour and social and political volunteerism to forms of activism aimed at shaping social structures and communities. one could argue that there is a ‘family resemblance’ between these various conceptualisations, in the sense of wittgenstein’s familienähnlichkeiten (wittgenstein, 1984). while such a resemblance is at once evident and elusive (the notion of ‘games’ is wittgenstein’s well-known example), it would benefit the study and understanding of this family of concepts as an outcome of higher education if specific essential shared characteristics could be identified. we propose that such shared characteristics can be found in what we call ‘fundamental constituents’ for the various interpretations of civic engagement. boyte and finders (2016) discuss ‘civic agency’ as the outcome of an interaction between an intellectual engagement with one's community or society and the desire to take action and belief in one’s efficacy. a similar view can be found in a much earlier study on the civic engagement of youth in general. watts and flanagan (2007) identified two mutually reinforcing factors in civic engagement: on the one hand, an interest in social and political affairs, and the other hand, a sense of agency related specifically to those affairs. in a comparative study of several developmental theories, wilkenfeld et al. (2010) found the same interaction between awareness and understanding and readiness and self-efficacy as two distinct features within the development of civic engagement in adolescents. katzarska-miller and reysen (2013, 2018), in their discussion of a typology for global citizenship based on meta-analyses of their own and others, conclude that “the main content of global citizenship is a concern for the environment, valuing of diversity, empathy for others beyond the local environment, and a sense of responsibility to act” (katzarska-miller & reysen, 2018, p. 3). all these perspectives on engagement share the idea that the presence of an interaction between two key traits is essential for the emergence van den wijngaard et al. 54 | f l r and development of civic engagement of any kind: an interest in, or intellectual engagement with social and political issues and a form of agency that combines a sense of responsibility and self-efficacy. consequently, in this study, rather than using the elusive concept of civic engagement as an outcome of education, we suggest focusing on two fundamental constituents of civic engagement instead: political interest (as a concise expression for having an interest in social and political issues) and agency. given the essential role as fundamental constituents of civic engagement of both political interest and agency, we suggest they are meaningful alternatives to the various indicators often used if one wants to understand the role of higher education in enhancing civic engagement. this role should ultimately be understood as the degree to which political interest and agency are affected by education. the question is whether education plays a role in developing students' civic engagement and whether both constituents of civic engagement, i.e., political interest and agency, are static or dynamic. if the latter, the next question would be whether any changes in these constituents would be uniform or varied. uniform change would suggest that a positive development in political interest and agency would merely be an effect of years of educational attainment, maturation, or both – leading us back to square one. variations in change between groups or individuals suggest a more intricate interplay between variables. in that case, the studies cited above would suggest a role for background variables and specific aspects of the educational context. 3. research aims and questions the following empirical part of this study presents an exploratory case study that focuses on two questions: (1) are political interest and agency static or dynamic, and more importantly, (2) if they evolve, do they do so in similar ways for every student, or are there diverse patterns of development? we hypothesised that we would find variation in how political interest and agency develop over time, indicating that any development in civic engagement, or rather in its constituents, is not simply a function of maturation or having spent a certain amount of time in education. this opens up the possibility that (constituents of) civic engagement are outcomes of higher education. to our knowledge, this is the first study that approaches civic engagement through a quantitative analysis of two variables often identified as essential or fundamental constituents. a better understanding of how these two variables can and do evolve, jointly or separately, in one specific educational environment will help set the agenda for future studies on the impact of aspects of the learning environment and experience on civic engagement across various contexts of higher education. 4. method 4.1 participants students with a fair degree of similarity in selection and self-selection were included in the study, all starting and graduating at a similar age and all participating in the same liberal arts and sciences, open curriculum programme. we collected data from students at a small, three-year liberal arts and open curriculum university college in the netherlands. we used an online survey that was administered twice. the first time was within several weeks after the start of the first year, as an integrated part of a mandatory introductory academic skills course (t1). we tested again two or three months before graduation while the students worked on their final thesis (t2). data were anonymised in both cases. with the first-year students, the survey was administered along with several other surveys on motivation and self-regulated learning. as compensation for participation, students were offered a workshop, also part of the same skills course, in which they could reflect on outcomes based on data at the aggregate level. the survey was not part of a course module in the final year. students who participated received van den wijngaard et al. 55 | f l r a gift voucher as compensation. not being embedded in a course within the curriculum, the final-year survey yielded fewer results than the first-year survey. data collection began in the fall of 2009 and ended in the spring of 2016. 951 t1-measurements (69.8% from 1362 first-year students) and 230 t2measurements (19.6% from 1173 last year students) were obtained. as we were interested in developments over time, we isolated the cases for which we had both a t1 and a t2 measurement. the subset from the data that resulted, our working sample, contained 190 cases. the mean age of participants at t1 was 19.6 years, with a standard deviation of 1.4 years; 38.7% of these students are german, 38.7 are dutch, and 20.9% are from other european countries. 4.2 setting all students study at a small liberal arts college (approximately 650 students in total at the time of data collection) in the netherlands. the college only confers bachelor degrees and offers an 'open curriculum': students are free to design their individual three-year curriculum. the open curriculum assumes that, if adequately supported by engaged and committed teaching and support staff, it enhances student motivation and a sense of responsibility and ownership regarding their studies (teagle, 2006). however, several specific criteria apply about the level and distribution of courses and other educational modules students select, which guarantee that students take a specific amount of advanced level courses and that there is a balance between a broad exploration of several academic disciplines (‘general education’) and a more in-depth focus within a broadly defined field of inquiry. the three broad fields of inquiry are called ‘concentrations’ (humanities, social sciences, and (life) sciences). the curriculum is geared predominantly toward the social sciences and the humanities. within this general framework, students have considerable freedom to define their own individual pathways. even within the 'concentrations,' there is room for various combinations of courses, making the concentration different from a 'major' where there is often less room for individual curriculum design. all courses employ the same pedagogical method of problem-based learning, and seventy per cent of all contact hours are spent in tutorial sessions with no more than twelve students under the guidance of one faculty. motivation and civic engagement on the part of prospective students play a notable role in the recruitment, application, and selection processes. eligible candidates (based on the level of the secondary education diploma, gpa, and motivation letter) participate in an admissions interview that includes questions about civic engagement and involvement. despite the individual freedom within the curricular structure and based on the emphases within the course offering as well as how the college selects and accepts its students, we assume that the majority of students share a distinct degree of interest in and engagement with social issues, while also being more or less the same age when they start their higher education. 4.3 measures we used the improved version of the conditions for civic engagement questionnaire (van den wijngaard et al., 2015), which consists of two scales, one for political interest and one for agency. the original validated questionnaire measured political interest and agency distributed across four scales. for measuring political interest, a distinction was made between two subscales: political interest and social analysis. agency consisted of the subscales valuing applicability and self-efficacy. all scales employed a likert scale ranging from 1 (disagree strongly) to 7 (agree strongly). after carefully examining the questionnaire, we omitted several items referring to past attitudes or behaviour as these would not allow for measuring current attitudes. five items remained in order to measure political interest. consequently, the subscales of political interest and social analysis were grouped into one factor: political interest. furthermore, taking into account previous internal reliability scores for the subscale self-efficacy (α = .57), it was decided to delete the scale. van den wijngaard et al. 56 | f l r the fit of the shortened cseq was tested through confirmatory factor analysis (cfa). the cfa was conducted on the complete sample of 951 respondents, of which the data used for this study are a subset, using ibm/amos 23. the fit of the model was established through several key indicators. the normed chi-square (χ2/df), to minimise the effect of the sample size, should be < 5 (bollen, 1989). hu and bentler (1999) suggest a combinational rule that requires the root means square error approximation (rmsea) to be < 0.06, the standardized root mean squared residual (srmr) < 0.08, and the comparative fit index (cfi) > 0.90, and preferably even > 0.95. following our conceptual framework, we assumed that the two factors would correlate. the initial fit of our model was moderate, as indicated by the fit indices χ2/df=4.87, cfi=0.87, rmsea=0.06, and srmr=0.06. based on the largest standardised residuals, we allowed the error terms of several items within the factors to correlate: in our factor political interest, the item "i usually like to discuss social issues and politics" with "i consider taking one or more courses in political science at the ucm." in the factor agency, the error term of the item “theory is only relevant if it can be applied in a practical way” correlated with “what and how university teaches should reflect the needs of society” and “i expect that the knowledge and skills i acquire in university can be easily applied in society” with “it should be possible to do volunteer or charity work for credit at the college” respectively. these adjustments resulted in values representing a good model fit: χ2/df=3.51, cfi=0.92, rmsea=0.05, and srmr=0.05. cronbach’s alphas were .70 for political interest and .67 for agency. table 1 lists the items of each scale. by measuring both variables twice, we generated four measures in total, two for each of our two variables, political interest (p.i.) and agency (a.g.), at t1 and t2, respectively. table 1 scales, items, and reliability political interest cronbach's α: .70 i usually like to discuss social issues and politics i consider taking one or more courses in political science at ucm i like to analyse the structure and workings of society governments should be followed critically university should help students develop a critical view of society agency cronbach's α: .67 theory is only relevant if it can be applied in a practical way university should actively encourage students to use their knowledge and skills for the benefit of others i believe individuals can make a real difference with regard to social issues social engagement can be taught and should be explicitly addressed in the university curriculum what and how university teaches should reflect the needs of society in fixing social problems and issues, the ‘hands-on’ approach usually works best it should be possible to do volunteer or charity work for credit at the college i expect that the knowledge and skills i acquire in university can be easily applied in society van den wijngaard et al. 57 | f l r 4.4 data analyses we first established the descriptives and pearson correlations for the variables indicated in table 2. then, we performed a paired samples t-test to determine whether our variables are dynamic, the outcomes of which should provide the basis for the analyses that allow us to establish the presence of specific patterns of development between the two variables. in order to test for patterns, we then performed a hierarchical cluster analysis for the change variables for political interest and agency, using the squared euclidian distance as a measure of similarity and applying ward's method in ibm spss version 24 (2016). we determined the number of clusters best fitting the data through hierarchical cluster analysis. based on the outcome of the hierarchical cluster analysis, we then performed a k-means analysis to arrive at an optimal assignment of subjects to clusters. 5. results table 2 presents the descriptives and pearson correlations for both variables, at both moments of measurement. table 2 descriptive statistics and pearson correlations for agency and political interest at t1 and t2 variable descriptives correlations mean sd range min max agency t1 agency t2 political interest t1 agency t1 5.01 .71 4.63 2.13 6.75 agency t2 4.92 .76 3.88 3.13 7.00 .49** political interest t1 5.72 .85 4.20 2.80 7.00 .13 .04 political interest t2 5.74 .95 4.60 2.40 7.00 -.04 .16* .58** *: correlation is significant at the .05 level (2-tailed) **: correlation is significant at the .01 level (2-tailed) a paired samples t-test was performed, which showed that there were no significant changes at the level of the total sample, as shown in table 3. table 3 paired samples test for agency and political interest at t1 and t2 variable mean sd se 95% confidence interval of the difference t df sig. (2tailed) lower upper agency .0928 .7459 .0541 -.01398 .19951 1.714 189 .726 political interest -.0211 .8273 .0600 -.13944 .09734 -.351 189 .088 van den wijngaard et al. 58 | f l r as argued above, more than the absence of significant change in the sample means is needed as evidence for the absence of meaningful patterns at the level of individuals or groups within the sample. this is where we took a decision that we believe represents a new approach in analysing this type of data in the context of research on the impact of education on civic engagement – as we have found no examples of it in the literature. in order to create the possibility of identifying change at the individual level and potentially relating this to other variables, we took two steps: we created change variables and analysed these through cluster analysis. first, we constructed change variables for both our independent variables. variables δpi and δag represent the difference in values t2-t1 for political interest and agency, respectively. their means are the same as those of the paired samples t-test presented in table 3. table 4 statistics for the change variables for agency and political interest variable δag δpi n 190 190 mean -.09 .02 std. error of mean .05 .06 std. deviation .75 .83 variance .56 .68 skewness .11 .11 std. error of skewness .18 .18 kurtosis .06 .69 std. error of kurtosis .35 .35 range 4.31 5.40 minimum -2.13 -2.40 maximum 2.00 3.00 we then established the normalcy of the distribution based on the key statistics for both change variables; the results are shown in table 4. the distribution for δpi is slightly leptokurtic, but otherwise, both variables display a normal distribution within a broad range, suggesting that both independent variables change over time in both positive and negative directions. both variables correlated weakly yet significantly, pearson’s r= .311, p<.001. checking for ceiling effects, we found that this is partially the case for political interest, but not in a strictly linear fashion. as presented in table 5, the most significant increase for both variables occurred where the starting scores were the lowest. however, the highest starting scores did not lead to the highest decrease, and each variable developed differently in the four clusters, so ceiling effects do not explain all variance in change. the distribution and correlation values for the two change variables showed that both evolve in various ways. van den wijngaard et al. 59 | f l r table 5 checking for ceiling effects: comparing t1 and change values agency t1 δag political interest t1 δpi means from lowest to highest value 4.643 .565 5.175 1.090 4.850 .438 5.742 .117 5.181 -.612 5.951 -1.130 5.362 -.608 5.958 -.113 we submitted both variables to a hierarchical cluster analysis using the change variables δpi and δag to determine the presence of significant patterns. based on the agglomeration coefficients shown in table 6, we established the cutoff point at four clusters. table 6 establishing cutoff point for cluster analysis ward's method number of clusters coefficient agglomeration change this step previous step 2 234,518 161,282 73,236 3 161,282 122,270 39,012 4 122,270 90,840 31,430 5 90,840 74,090 16,750 6 74,090 62,251 11,839 7 62,251 55,378 6,873 8 55,378 49,122 6,256 9 49,122 43,845 5,277 10 43,845 38,669 5,176 we then performed a k-means analysis restricted to a single solution of four clusters. analysis of variance (anova) confirmed that the differences between the clusters we obtained were significant for both change variables; for the change variable for political interest: f(3,186) = 179,589, p = .000), for the change variable for agency: f(3,186)= 80,293, p=.000). we thus found that the largest cluster (n=65) consisted of students whose political interest had only a marginally larger value (-.61). the second cluster (n=48) showed a decrease in political interest (-.11) and an increase in agency (.57). the third cluster (n=40) was populated with students whose political interest and agency both increased rather substantially (1.09 and.44 respectively). the fourth and smallest cluster (n=37) was made up of students with a substantially decreased level of political interest (-1.13) and agency (-.61). figure 1 offers a visual representation of the varying degrees of change for both variables for all four clusters. van den wijngaard et al. 60 | f l r 6. discussion and conclusion our analysis of the change variables for political interest and agency showed significant variation in how these two variables developed over time. this confirmed our hypothesis that this would be the case and supported using these fundamental constituents as more specific and operationalisable variables in studies on civic engagement as an outcome of (higher) education. furthermore, the outcome supports our suggestions that cluster analysis rather than analyses of sample means provides helpful insights into the development of these variables over time. with this study, we aimed to assess whether the impact of higher education on civic engagement could be understood better if the focus shifted from studying civic engagement to its constituents. based on a literature review, we proposed two such constituents: political interest and agency. as important as this shift toward different dependent variables is the assessment of the nature of these variables: are they static or dynamic? furthermore, we wanted to know whether, if dynamic, such variables would evolve in a uniform or a diverse fashion, in which case those changes could be attributed to factors that are not uniform (much as maturation) but are determined by specific contexts, including educational ones. from the studies included in this review, egerton (2002) takes the most radical stance, proposing that attitudes towards civic education are primarily developed and defined before entering higher education and that, therefore, these are hardly affected by education at that level. if maturation or other variables that apply to all students in a cohort were the most important factors influencing our dependent variables, as milner et al. (1999) suggest, we would expect to see a uniform development with no significant differences between individuals or groups progressing through time within a cohort. similar outcomes, yet for different reasons, should be expected if constituents of civic engagement would be primarily a function of (self) selection, as maheux and beland (1987) suggest, social background and self-selection, as proposed by smart, feldman, and ethington (2000), in their application of holland’s theory on congruence. even if we assume an effect of education on civic engagement, there may be different rationales for that expectation, yet with similar consequences for how dependent variables are affected. the relationship between education and civic engagement that milligan et al. (2004), dee (2004), and helliwell and putnam (2007) found is mainly a correlation between years in higher education and levels of civic engagement. astin (1997) adds changes in the overall social climate as a factor to years in college, yet in all of these cases, one would expect similar straightforward correlational effects for an entire cohort. figure 1: change in means (y-axis) per cluster (x-axis) for political interest and agency from t1 to t2 van den wijngaard et al. 61 | f l r if other factors, within and outside the educational setting, would affect different individuals or groups differently, we would expect to find two or more diverging patterns within a cohort. focusing on the impact of experiences of interactional diversity, the studies by gurin et al. (2002) and laird (2005) suggest that specific aspects of the educational experience that are not shared by an entire cohort lead to different outcomes regarding levels of civic engagement. we collected data from first-year students and students near graduation at a three-year liberal arts undergraduate programme. this programme has several features that provide the possibility of diverging educational paths due to the openness of the curriculum, in which students design their individual programmes within certain boundaries related to level and focus. at the same time, the way the programme recruits and selects its students strongly emphasises social responsibility and engagement. in terms of the starting position, as milner et al. (1999) discuss in their study on business students, while not necessarily being homogenous, incoming students at this particular college may be expected to have a certain level of civic engagement in common. this relatively high level of ‘civic orientation’ (myers et al., 2019) is reflected by the values of our two variables at t1, the beginning of the first year, for the entire sample: 5.72 (sd 0.85) for political interest and 5.01 (sd 0.71) for agency on a 1-7 likert scale. despite this apparent homogeneity, which seemed to be confirmed by similar changes in both dependent variables to the extent of moderate yet significant correlation, upon cluster analysis, we found four distinct groups, significantly different from each other, with varying combinations of increasing or decreasing values for political interest and agency. the range of change is slightly larger for political interest (between -1.13 and 1.09) than for agency (between -.61 and .57), yet both are substantial given the 1-7 scale. patterns of change are somewhat diverse, as is also suggested by the weak correlation between the two change variables (r= .311, p<.001) between these groups; increasing and decreasing values for both variables combine in different ways. the outcomes of this case study suggest that it adds to our understanding of the impact of higher education on civic engagement if more attention is given to its impact on political interest and agency, which we presented as two important constituents of civic engagement. we found that, by applying cluster analysis rather than analyses based on sample means, these constituents are not only dynamic but, while correlating, evolve separately in diverse ways. this implies that maturation or years in college do not provide an adequate explanation for the change, or at least not to the extent that would render research into other potential determinants irrelevant. likewise, the effect of a similar starting position, at least in this group of students at a college with a reasonably distinct profile emphasising civic engagement, does not seem to be sufficient to explain individual developments in political interest and agency or the dynamic between the two constituents. 7. limitations and further research this study aimed to establish the nature of political interest and agency, seen as dependent variables in models describing the impact of higher education on civic engagement. the sample size and setting scope were limited to one liberal arts college with a specific signature that emphasised civic engagement. while this allowed us to see whether there are distinct variations in the development of both variables even in a potentially homogenous sample, confirmatory studies in other settings could further establish the nature of these variables and, with that, raise the level of relevance of research exploring the reasons for such variation. in the absence of background variables such as socio-economic status, parental education, and ethnicity, we refrained from drawing any conclusions about the potential interaction between the input and the environment from astin’s (1997) input-environment-output model, as myers et al. (2019) did in their study on the effect of high impact practices on civic engagement. in the absence of substantial community or service-oriented activities within or around the curriculum that could serve the ‘third mission of the university’ (compagnucci & spigarelli, 2020), we were not able to make a meaningful appraisal of the role such potential aspects of the environment may have played towards the observed van den wijngaard et al. 62 | f l r outcomes. likewise, an analysis of the role of an intersectional starting position, as explored by wraylake et al. (2014), wray-lake and schubert (2019), and wray-lake et al. (2020), could not be performed. looking at nationality, the distribution of students within each cluster did not deviate significantly or even remarkably from that over the entire sample, with a similar clear and dominant presence of dutch and german students. we found, however, several interesting characteristics of the membership of each of the four clusters regarding gender and curricular reorientation. in a sample of n=190 and a 3.41:1 female/male ratio, the effect of gender is difficult to establish with statistical significance. furthermore, the binary gender categories used in this study do not allow for a sophisticated understanding of the potential impact of gender on the development of the fundamental constituents of civic engagement. acknowledging these limitations, we noticed that the third cluster, which shows an increase in both political interest and agency, had a high female/male ratio of 4.7:1. the fourth cluster, in which both variables decrease, had a ratio of 2.4:1, considerably lower than the sample ratio. a second set of observations concerned the academic focus of the members of each cluster and relates to the positive effect described by hillygus (2005) and hoyte (2018) of social science courses on civic engagement. as explained above, the curriculum at the college under study is open. it does not have defined majors, but students cluster their courses within the ‘concentrations’ social sciences, humanities, and (life) sciences or combine one (and occasionally all) of these ‘concentrations’. as 51.6% of students in this study graduated with a social science curriculum, and a further 30.0% with a focus within the social sciences and humanities, combined with the relatively small sample, it was impossible to draw statistically significant conclusions. interestingly, however, we saw that the clusters with the highest combined increase and the highest combined decrease in values for both variables had the highest percentage of social science graduates (57.5% and 59.5 respectively). comparing these two clusters would suggest that a strong presence of social sciences courses in a curriculum does not guarantee a positive development of political interest and agency. a final observation concerned the difference between clusters regarding students shifting their academic focus between the first and the final semesters. on average, 13.2% of all students in the sample graduated with a different academic focus than they identified at the beginning of their studies. the same caveats apply, but here we found that the cluster of members whose political interest and agency grew significantly also had the largest percentage of switches in curricular focus (17.5%). in comparison, the cluster in which the values of both variables dropped substantially shifted the least (5.4%). the limitations of this study lie in its limited scope, as we used data obtained from students at a single institution, and we were able to perform a longitudinal analysis on only a subset of that data. the fact that, even within a relatively small and homogenous sample, we could establish four classes or clusters based on distinctly different patterns justified one crucial conclusion: political interest and agency, as fundamental constituents of civic engagement, clearly are dynamic and can evolve in similar or opposite directions. while our sample did not allow for specific correlation, let alone causation, between other environmental or background variables and these trends, our tentative findings suggest the possibility of explanations, some of which are, and some of which are not in line with the existing literature. research to provide such explanations would have to focus on the individual, as lerner et al. (2014) suggested. a robust approach to intersectionality advocated by, e.g., wray-lake et al. (2020) would have to do justice to the potential interaction between input and environmental factors, thus extending the notion of intersectionality beyond various personal background variables. as noted needing a stronger quantitative foundation, our three tentative musings suggest an interplay between (at least) gender, academic focus, and stability in academic focus. the latter two aspects could be explored by reviewing the role of disciplines and intended outcomes of programmes and institutions, service and community-oriented components of the learning environment (curricular and extracurricular), as well as the implicit or explicit emphasis on civic engagement in curricula (the visibility of a ‘third mission’), van den wijngaard et al. 63 | f l r in the context of civic engagement. the idea that students tend to enrol in programmes that match their interests and personal values, as developed by smart, feldman, and ethington (2000), building on congruence theories developed by holland, may provide an attractive linking pin between the role of personal attitudes and preferences on the one hand, and of specific educational programmes or interventions on the other. this connection could be given further explanatory power if differences between developments in constituents of civic engagement could be related to differences between academic disciplines in terms of concepts of knowledge and approaches to teaching and learning (e.g., neumann, 2001; smart et al., 2000). in addition to quantitative methods, as explored in this article, further understanding of the interaction between various variables could be obtained through qualitative research by interviewing students individually or in groups. 8. practical implications next to the implications for future research discussed above, on a more practical level, our findings suggest that specific constituents of civic engagement can be seen as outcomes of learning. how these outcomes evolve may depend on many factors within and outside the curriculum. developing or maintaining a sense of civic engagement (or its constituent) is a volatile process that requires careful attention. if the promotion of civic engagement in students is part of the mission of a course (or programme or institution), the identification of such factors and an assessment of their impact will need to be part of the educational design process. the fact that civic engagement as an outcome of learning can be broken down into two specific, distinct fundamental constituents may help in that process. rather than having to choose between precise outcomes, most of which may potentially only manifest themselves after graduation (e.g., voting, party membership, participation in civic or political action), or broad, hard-to-grasp objectives (an attitude of 'civic engagement'), didactics and pedagogy can be applied towards more tangible, measurable outcomes, that may manifest themselves while students are still in college. key points we show that the varying conceptualisations of ‘civic engagement’ share political interest and agency as ‘fundamental constituents’. we propose that using such constituents as dependent variables will allow for a better understanding of the impact of higher education on civic engagement of students. we show that an analysis of profiles of higher education students shows a more complex and nuanced picture of how their civic engagement evolves than when analysing the relation between education and the evolution of civic engagement on a group level. references aronowitz, s. (2000). the knowledge factory. dismantling the corporate university and creating true higher learning. boston: beacon press astin, a.w. (1997). what matters in college? four critical years reviewed. san francisco: jossey-bass barnett, r. (1990). the idea of higher education. bristol, pa: society for research into higher education and open university press barnett, r. (2000). realising the university in an age of supercomplexity. bristol, pa: society for research into higher education and open university press. van den wijngaard et al. 64 | f l r bok, d. (2005). universities in the marketplace. the commercialisation of higher education. princeton, nj: princeton university press bok, d. (2008). our underachieving colleges. a candid look at how much students learn and why they should be learning more. princeton, nj: princeton university press bollen, k. (1989). structural equations with latent variables. new york: john wiley and sons boyer, e l. (1990). scholarship reconsidered: priorities of the professoriate. (the carnegie foundation for the advancement of teaching) san francisco: jossey-bass boyte h.c., (2008). against the current: developing the civic agency of students. change: the magazine of higher learning, 40(3), 8-15. https://doi.org/10.3200/chng.40.3.8-15 boyte, h.c., & finders, m.j. (2016). "a liberation of powers": agency and education for democracy. educational theory, 66(1-2), 127-145. https://doi.org/10.1111/edth.12158 compagnucci, l., & spigarelli, f. (2020). the third mission of the university: a systematic literature review on potentials and constraints. technological forecasting and social change, 161. https://doi.org/10.1016/j.techfore.2020.120284 checkoway, b. (2001). renewing the civic mission of the american research university. the journal of higher education, 72(2), 125-147. https://doi.org/10.1080/00221546.2001.11778875 colby, a., ehrlich, t., beaumont, e., & stephens, j. (2003). educating citizens: preparing america’s undergraduates of lives of moral and civic responsibility. san francisco: jossey-bass dee, t. (2004). are there civic returns to education? journal of public economics, 88(9-10), 1697-1720. https://doi-org.ezproxy.ub.unimaas.nl/10.1016/j.jpubeco.2003.11.002 egerton, m. (2002). higher education and civic engagement. british journal of sociology, 53(4), 603620. https://doi.org/10.1080/0007131022000021506 gurin, p., dey, e.l., hurtado, s., & gurin, g. (2002). diversity and higher education: theory and impact on educational outcomes. harvard educational review, 72(3), 330-366. https://doi.org/10.17763/haer.72.3.01151786u134n051 hartley, m., saltmarsh, j., & clayton, p. (2010). is the civic engagement movement changing higher education? british journal of educational studies, 58(4), 391-406. https://doi.org /10.1080/00071005.2010.527660 helliwell, j. f., & r.d. putnam (2007). education and social capital. eastern economic journal, 33, 119. https://doi.org/10.3386/w7121 hillygus, s. (2005). the missing link: exploring the relationship between higher education and political engagement. political behaviour, 27(1), 25-47. https://doi.org/10.1007/s11109-005-3075-8 hu, l., & bentler, p.m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives. structural equation modeling: a multidisciplinary journal, 6(1), 1-55. http://dx.doi.org/10.1080/10705519909540118 hylton, m.e. (2018). the role of civic literacy and social empathy on rates of civic engagement among university students. journal of higher education outreach and engagement, 22(1), 87-106 katzarska-miller, i., & reysen, s. (2018). inclusive global citizenship education: measuring types of global citizens. journal of global citizenship and equity education, 6(1), 1-23 laird, t.f.n. (2005). college students’ experiences with diversity and their effects on academic selfconfidence, social agency and disposition toward critical thinking. research in higher education, 46(4), 365-387. https://doi.org/10.1007/s11162-005-2966-1 lewis, h. r. (2007). excellence without a soul. does liberal education have a future? new york, ny: public affairs maheux, b., & beland, f. (1987). changes in students’ sociopolitical attitudes during medical school: socialisation or maturation effect? social science and medicine, 24(7), 619-624. https://doi.org/10.1016/0277-9536(87)90067-0 milligan, k., moretti, e., & oreopoulos, p. (2004). does education improve citizenship? evidence from the united states and the united kingdom. journal of public economics, 88(9-10), 1667-1695. https://doi-org.ezproxy.ub.unimaas.nl/10.1016/j.jpubeco.2003.10.005 https://doi.org/10.1080/00221546.2001.11778875 van den wijngaard et al. 65 | f l r milner, d., mahaffey, t., maccaulay, k., & hynes, t. (1999). the effect of business education on the ethics of students: an empirical assessment controlling for maturation. teaching business ethics, 3. 255-267 myers, c.b., myers, s.m., & peters, m. (2019). the longitudinal connections between undergraduate high impact curriculum practices and civic engagement in adulthood. research in higher education, 60, 83-110 neumann, r. (2001). disciplinary differences and university teaching. studies in higher education, 26(2), 135-147. https://doi.org/10.1080/03075070120052071 nussbaum, m. c. (2010). not for profit. why democracy needs the humanities. princeton: princeton university press palmer, n. (2018). emergent constellations: global citizenship education and outrospective fluency. journal of research in international education, 17(2), 134-147. https://doi.org/10.1177/1475240918793963 pascarella, e.t., & terenzini, p.t. (1991). twenty years of research on college students: lessons for future research. research in higher education, 32(1), 83–92 pascarella, e.t., & terenzini, p.t. (2005). how college affects students. volume 2; a third decade of research. san francisco: john wiley & sons. readings, b. (1997). the university in ruins. cambridge ma: harvard university press reysen, s., & katzarska-miller, i. (2013). a model of global citizenship: antecedents and outcomes. international journal of psychology, 48(5), 858-870. https://doi.org/10.1080/00207594.2012.701749 smart, j.c., feldman, k.a., & ethington, c.a. (2000). academic disciplines. holland’s theory and the study of college students and faculty. nashville: vanderbilt university press teagle foundation. (2006). the values of the open curriculum: an alternative tradition in liberal education [white paper]. new york watts, r.j., & flanagan, c. (2007). pushing the envelope on civic youth engagement: a developmental and liberation psychology perspective. journal of community psychology, 35(6), 779-792. https://doi.org/10.1002/jcop.20178 wijngaard, o. van den, beausaert, s., segers, m., & gijselaers, w. (2015). the development and validation of an instrument to measure conditions for social engagement of students in higher education. studies in higher education, 40(4), 704-720. https://doi.org/10.1080/03075079.2013.842214 wilkenfeld, b., lauckhardt, j., & torney-putra, j. (2010). the relation between developmental theory and measures of civic engagement in research on adolescents. in sherrod, l.r., torney-putra, j., & flanagan, c. (eds.) handbook of research on civic engagement in youth (pp. 193-208). hoboken, nj: john wiley & sons. wittgenstein, l. (1984). tractatus logico-philosophicus. tagebücher 1914-1916. philosophische untersuchungen. frankfurt am main: suhrkamp wray-lake, l., arruda, e.h., & schulenberg, j.e. (2020). civic development across the transition to adulthood in a national u.s. sample: variations by race/ethnicity, parent education, and gender. developmental psychology, 56(10), 1948–1967. https://psycnet.apa.org/doi/10.1037/dev0001101 wray-lake, l., rote, w. m., benavides, c. m., & victorino, c. (2014). examining developmental transitions in civic engagement across adolescence: evidence from a national u.s. sample. international journal of developmental science, 8, 95–104. https://doi.org/10.3233/dev14142 wray-lake, l., & shubert, j. (2019). understanding stability and change in civic engagement across adolescence: a typology approach. developmental psychology, 55(10), 2169-2180. http://dx.doi.org/10.1037/dev0000772 yang, j., & hoskins, b. (2020). does university have an effect on young people’s active citizenship in england? higher education, 80, 839–856. https://doi.org/10.1007/s10734-020-00518-1 frontline learning research vol. 12 no. 1 (2024) 34 -48 issn 2295-3159 corresponding author: tine nielsen, ucl university college, niels bohrs allé 1, 5230 odense m, tini@ucl.dk doi: https://doi.org/10.14786/flr.v12i1.1347 student teachers’ opportunities to learn through observation, own practice and feedback on the practice while in field practice placements: a graphical model approach tine nielsen ucl university college, denmark article received 15 august 2023 / article revised 9 january 2024 / accepted 10 january / available online 25 january abstract field practice placement is a crucial part of teacher education, as it affords a real-life context, where teacher and teacher-related skills can be enacted and trained. the present study examined the associations between student teacher opportunities to learn through observation, own practice and the receiving of feedback of said practice, while in field practice placements through a teacher education programme. chain graph models were used to analyse data from 560 danish student teachers who had just completed field practice at one of three levels. results showed that opportunities to learn through observation of fellow students and other teachers was negatively associated with level of field practice, and thus was reported less and less the further along students were in the programme, while opportunities to learn through own practice was positively associated with level of field practice. opportunities to learn through receiving feedback on own practice was associated with level of field practice only via opportunities to learn through own practice. results did not reveal gender or age-wise inequity in the opportunities to learn afforded in the field practice. teacher education programmes could benefit from placing additional focus on opportunities to learn through observation in the later field practice placements. keywords: teacher education; field practice placement; opportunities to learn; chain graph model https://doi.org/10.14786/flr.v12i1.1347 nielsen 35 | f l r 1. introduction teacher education differs across cultures and countries, but a common denominator is that it consists of two parts; the academic (on-campus) part and the non-academic or skills (in schools) part. the manner in which these parts are organized to form a teacher education programme also differs, and the most common form appears to be a subject matter bachelor degree followed by an education degree at the masters level, which includes in-school training (weisdorf, 2020). teacher education has been a subject of study for many decades (menter, 2022). two of the prevailing issues studied are the quality of teacher education, with a more recent strand focusing on the coherence of the academic and nonacademic parts of teacher education (e.g. canrinus et al., 2017; grossman et al., 2008; youngs et al., 2022), and the non-academic parts in themselves (e.g. more, 2003; ulvik et al, 2021). the non-academic parts of teacher education, also denoted clinical experience, pre-service teaching, student teaching placements, practicum, or field practice placement (the latter term will be used throughout this article) is an integral part of teacher education, as it offers unique opportunities to learn the non-academic skills needed to become a teacher. the research on field practice placements in teacher education appears to be focused mainly on the role of universities and schools (for field practice), what is best learned where and from whom, and teacher development and identity (menter, 2022). only more recently, research on opportunities to learn while in field practice has appeared (e.g. cohen & berlin, 2020; nielsen, 2021; youngs et al., 2022). opportunities to learn while in field practice are here used to mean opportunities to learn by engaging in real-life teaching activities in the field practice placements, thus focusing on the training of teacher-skills aspects of the education. cohen and berlin (2020) divide these into opportunities to engage with representations and decompositions of practice and opportunities to approximate or enact teaching practices. in the current study, both aspects are included as distinct opportunities to learn in field practice: representations of practice through opportunities to learn through observing other teachers, enactment of teaching practice though own practice, and decompositions of practice through receiving feedback on own practice (c.f. nielsen, 2021). hammerness et al., (2020) studied opportunities to learn in teacher education through study, practice, and rehearsal of teaching in five countries. however, the context was not field practice placements, but campus coursework. opportunities to learn in field practice placement has however, been studied by e.g. youngs et al. (2022), who had mixed findings, as they found that in mathematics, opportunities to learn about and practise content-specific ambitious instructional practices, during student teaching, were positively associated to their first-year teaching practice through their representation of content and instructional scaffolding. on the other hand, youngs et al. (2022) also found that these opportunities to learn were negatively associated with the students’ ability to create or maintain a productive learning environment. in the context of danish teacher education, it appears that only two single studies have been conducted on opportunities to learn while in field practice. nielsen (2021) used the field practice experience scales (fpe-dk) and found that students who had completed the first two (of three) field practice placements scored significantly and substantially higher on opportunities to learn through observation than did students who had completed the third and thus all field practice placement. with regard to opportunities to learn through own practice, students who had just completed the last two field practice placements scored significantly, but not substantially higher than did students who had completed only the first field practice. lastly, students who had completed the last field practice placement just prior to taking the fpe-dk scored significantly although not substantially higher on opportunities to learn through receiving feedback on their practice than students who completed the first field practice did. nielsen and graf (2021) explored the specific teaching and teaching-related activities in the fpe-dk scales and found differences in which specific opportunities to learn were experienced in the first, second and third field practice placement. nielsen 36 | f l r in denmark, teacher education is an integrated programme focusing on coherence between the teaching subjects, the pedagogical and didactical subjects, and the field practice placements, within a single four-year long teacher education programme (weisdorf, 2020). teacher education in denmark awards a so-called professional bachelor’s degree. in the danish teacher education programme, the ministry of education and research regulates the field practice placement, which amounts to a total of 30 ects (european credit transfer system), which is equivalent to half a year’s worth of study intensity. within a specific teacher education programme, there can be as many as six field practice placements. however, independently of the number of placements, these should demonstrate an education-wise progression corresponding to the nationally defined skills and knowledge objectives within three defined areas of competence for each of three levels of field practice. the competence areas and the included skills and knowledge objectives for each level of field practice are described in nielsen (2021, the s1 file at https://doi.org/10.1371/journal.pone.0258459.s009). students obtain teaching competence in usually three and at least two teaching subjects (ministry of education and research, 2015). one teaching subject has to be danish or mathematics, while the remaining teaching subject(s) can be any subject taught in primary and lower secondary school. in the teacher education programme studied in the current study, field practice is placed within the first, third and fourth years of study and at varying times in the academic year, depending on the time of admission to the teacher education programme (summer or winter). furthermore, at the university college in question, there are two campi each with a summer and a winter intake of students, which follow somewhat different study plans, where the time-wise relationship between the various teaching subjects and the field practice is not the same. in the part of the curriculum for the field practice specific to the university college it is explicitly stated that the students should be provided the opportunity both to observe teachers teaching and to practise teaching themselves in all placements (ucl erhvervsakademi og professionshøjskole, 2022). in addition, the field-practice handbook at this university college states that it is important in relation to the students’ learning processes that they provide each other with feedback and that a three-party (i.e., students, campus teacher and field-practice teacher) supervisory talk is mandatory midway through the placement (larsen, 2021). lastly, there is a contractual agreement with the field-practice schools that they should provide a minimum of one hour’s supervision and feedback each week for the students. recently nielsen (2021) introduced and validated the three field-practice experience scales (fpe-dk). this instrument is the first danish instrument to measure specifically the learning opportunities the student teachers experience through observing, practising and receiving feedback on certain teaching-related activities in field practice placement. thus, the fpe-dk provides the means for investigating students’ experienced opportunities to learn through both observation, own practice and feedback on this practice in a standardized manner in the different field practice placements in the danish teacher education programme. 1.1 the current study the aim of the current study was thus to conduct a first study of the relationships between opportunities to learn through observation, own practice and receiving feedback while in field practice, as measured with the fpe-dk, with the level of field practice students had just completed as well as the interrelationship between the three field experience scales themselves. this will be investigated while taking into account the dependence (or independence) of the three types of opportunities to learn on the type of teacher education programme students were enrolled in, the campus they studied at, as well as their gender and age, and the relationships between these educational and background variables. specifically, it was expected that two of the field practice experience scales would be positively associated with the level of field practice, so that the “rate” of observation and own practice would increase with the level of field practice, as students become more advanced learners and therefore can https://doi.org/10.1371/journal.pone.0258459.s009 nielsen 37 | f l r engage more and more in these processes. no association was expected between level of field practice and the third scale; receiving feedback on own practice, as there is an equal expectation of the amount of feedback provided at each level of field practice and the degree of feedback would then rather be associated directly to the degree of opportunities to learn through own practice. lastly it was expected that the three field practice experience scales would be positively associated with each other. 2. methods 2.1 participants and data collection participants were student teachers (n = 560) who had just completed 6 weeks of field practice placement in a danish public school (primary and lower secondary school) as part of the danish teachertraining programme at one of the danish university colleges. data were collected using an online survey during four weeks immediately after students had completed field practice placements. the majority of the students were enrolled in the regular bachelor of education programme (85.1%) at one of the two campi (72.9% versus 27.1%) of the university college (table 1). the majority of the sample identified as female (70.5%), and the mean age of the sample was 26.8 years. these numbers match the distribution of students admitted to this particular university college. information on the level of the students’ latest field practice (i.e. the one in question), was collected from the study administration at the university college. the distribution of field practice levels was uniform in the study sample, one third at each level (table 1). table 1. characteristics of the study sample (n = 560). frequency (%) campus campus a campus b 408 (72.9) 152 (27.1) ba education programme regular 480 (85.7) other 80 (14.3) latest field practice placement level 1 level 2 level 3 185 (33.0) 188 (33.6) 187 (33.4) gender female male 395 (70.5) 165 (29.5) age groups 23 years and younger 24-26 years 27 years and older 190 (33.9) 198 (35.4) 172 (30.7) mean age (sd), range 26.8 (6.9), 19-65 nielsen 38 | f l r 2.2 instruments the three field practice experience scales each measure student teachers’ opportunities to learn through observation, own practice (i.e. enactment) and feedback on their practice of 12 teacher practices while in field practice placement, as part of their teacher education programme (nielsen, 2021). student teachers report whether or not they have had the opportunity to observe, practice and/or receive feedback on the 12 teacher practices. eleven of the 12 teacher practices originated from the development of ambitious instruction (dai) project (available at www.daiproject.weebly.com). these items were changed by nielsen (2021) to not refer to teaching mathematics, but instead to refer to subject-specific teaching or to have no specific reference, and a new response scale was designed. these changes were made by nielsen (2021) so that the instrument could be used with a student-teacher population with diverse teaching subjects and several of them (c.f. the introduction) – see cohen and berlin (2020) for the original items directed towards mathematics education). the 12th teacher practice “facilitation of a good socio-emotional learning environment” was suggested by a group of norwegian researchers to tap into the more relational side of classroom management (nielsen, 2021). the three resulting 12-item field-practice experience scales were named observed scale, practiced scale and received feedback scale to signal the type of learning opportunities the students experience through the teaching-related activities in field practice placement, while the instrument was named fpe-dk for field practice experience danish language version (nielsen, 2021). the items of the three field practice experience scales are available in both english and danish in nielsen (2021). as the three field practice experience scales have previously been shown to fit the rasch model (nielsen, 2021), the sum score is considered to be a sufficient statistic for the estimated person parameter. in the current study, the choice was thus made to use the sum scores of the three scales in the chain graph model. the sum scores are in reality counting scales where the scores signify the number of opportunities to learn in field practice though observation, through practice and through feedback related to 12 teaching practices, as experienced by the teacher students. the score distributions of the three scales are shown in figure 1. reliabilities reported by nielsen (2021) were: observed scale 0.92, practised scale 0.65, received feedback scale 0.87. in the current study, reliabilities were similar (table 2). table 2. mean (sd) and reliabilities of the three field practice experience scales. scale min max mean sd cronbach’s alpha otl through observation 0 12 7.29 4.15 0.92 otl through own practice 0 12 10.36 1.92 0.69 otl through feedback on practice 0 12 8.49 3.51 0.88 http://www.daiproject.weebly.com/ nielsen 39 | f l r figure 1. distribution of scores on the three field experience scales; opportunities to learn through observation (top), own practice (middle), and receiving feedback on practice (bottom). nielsen 40 | f l r 2.3. statistical methods chain graph models (lauritsen, 1996) consist of nodes representing variables, directed arrows representing causal associations and undirected edges representing non-causal associations (the latter are not present in the more commonly used directed acyclic graphs, dags). chain graph models have a block-recursive structure where arrows are present between blocks, and edges are present within blocks. all paths between any two variables can be determined by the graph structure. thus, it is possible to identify a minimum set of variables to condition on when estimating the direct association between two variables, and thereby simplify the analysis (for further details on analysis by graphical models see lauritzen, 1996, and kreiner et al., 2009). the use of chain graph models allowed assessment not only of the relationships between each of the three field practice experience scales and the level of field practice, and the additional background variables; the type of teacher education programme students were enrolled in, the campus they studied at, and the gender and age of the students. the method at the same time allowed assessment of the associations between the three field practice experience scales themselves. log-linear chain graph models (lauritzen, 1996) were used and not structural equation models, as the latter would presume at least interval level scale and normally distributed data, which was clearly not the case here (figure 1), while the log-linear chain graph models are appropriate for counting and ordinal level scales. figure 2 shows the block-recursive structure underlying the analysis. block a consisted of age and gender time with arrows pointing to the blocks occurring later in time. block b consisted of the campi students were studying at and the type of teacher education programme they were enrolled in, again with arrows pointing to the blocks occurring later in time. block c consisted solely of the level of field practice of the students. finally, block d consisted of the three field practice experience scales. figure 2. the block structure of the chain graph model used in the study. notes. block a: age and gender. block b: campus and type of teacher education programme. block c: level of field practice. block d: the three field practice experience scales. the correlation structure was determined based on statistically significant correlations using partial goodman-kruskal gamma (γ) correlations (davis, 1967; goodman & wallis, 1954; kreiner, 1987). the gamma coefficients are rank correlation coefficients for ordinal categorical data, where a γcoefficient > 0.30 is regarded as a strong association, and a γ-coefficient < 0.10 a weak association. the digram software package was used to defines and test the chain graph model (kreiner, 2003). an automated screening procedure for high-dimensional contingency tables was used first to define a somewhat simpler starting model than the full block-recursive model. this was followed by a stepwise manual model selection strategy aimed at improving the starting model and finally identifying an adequate model for data (kreiner, 1986). decisions about including or eliminating interactions in the manual strategy were based on both the strength of the associations and the strength of the evidence (i.e. p-values). thus, on the one hand, weak associations (i.e. γ < 0.10) were not considered unless the evidence was very strong. on the other hand, strong associations were not considered if the evidence was very weak. the strength of the evidence was evaluated on a continuum distinguishing between weak (p < 0.05), moderate (p < 0.01) and strong (p < 0.001) evidence, as recommended by cox and colleagues (1977). nielsen 41 | f l r having reached a final model, this was confirmed by testing the necessity of including the associations in the final model as well as testing the adequacy of the associations in the final model. subsequently, partial correlations were estimated for all associations in the final chain graph model. the issue of multiple testing was dealt with by controlling the false discovery rate (fdr) using the benjamini-hochberg procedure (benjamini & hochberg, 1995). the problem of estimating the γcoefficients and p-values using asymptotic methods was dealt with by using a monte carlo procedure with 400 samples to obtain exact p-values. 3. results the results in the form of a final chain graph model showing associations and the lack thereof (i.e., conditional dependence/independence) are shown in figure 3 including the partial correlations. no causality other than that imposed by time is implied in the graph (from right to left, cf. the recursive structure in figure 2). the three field practice experience scales are mutually associated, so that there are is a strong positive association (γ = 0.30) between opportunities to learn through observation and through receiving feedback on own practice. also, there is an even stronger positive association between opportunities to learn through own practice and receiving feedback on this practice (γ = 0.67). it should be noted that the latter association is to some degree artificially high, as it is, of course, not possible to score highly on opportunities to learn through feedback on own practice, if you have not also scored highly on opportunities to learning through own practice, but it is possible to score highly on practice, but not so on feedback (cf. the score distribution in figure 1). these findings are not entirely in accordance with the a-priory expectations, as the graphical model revealed that there was no positive association between the degree to which students reported opportunities to learn through observation and through own practice, in fact these were conditionally independent given the remaining associations in the model. with regard to conditional dependence (or independence) of the three field-practice experience scores on the included education, the substantial results were first and foremost in relation to the level of the field practice just completed. opportunities to learn through observation was strongly and negatively associated with level of field practice (γ = -0.37), opportunities to learn through own practice was moderately and positively associated with level of field practice (γ = 0.25), while opportunities to learn through feedback on own practice was not associated with level of field practice. thus, the apriory expectation of positive associations between the level of field practice and opportunities to learn through observation and own practice was only confirmed for own practice, while a negative association was found for observation. in addition, there was a weak negative association between the type of teacher education programme students were enrolled in and opportunities to learn through observation, signifying that students in the regular programme experienced fewer opportunities to learn through observation than did students enrolled in for example, a trainee programme or a programme for professionally trained students. with regard to student characteristics such as gender and age, there were no associations (direct or indirect) between the three field practice experience scores and gender. thus, the degree to which the students experienced having had opportunities to learn through observation, own practice and feedback on own practice, was independent of gender. age on the other hand was indirectly associated with students’ reported opportunities to learn through observation, own practice or feedback on own practice, as all paths from age to the three field practice experience scores passed through either the type of teacher education programme students were enrolled in or the level of the field practice they had just completed. nielsen 42 | f l r figure 3. the final chain graph model notes. all lines represent significant conditional associations between variables. correlations are partial goodman-kruskal gamma-coefficients (γ). the results of testing the necessity of the associations in the final model are included in the appendix (table a1), as are the results of testing the adequacy of the associations in the model (table a2). 4. discussion and implications the study was primarily aimed at investigating the relationships between the three field practice experience scales; opportunities to learn through observation, opportunities to learn through own practice, and opportunities to learn through receiving feedback on practice, and the levels of field practice students had just completed, as well as the interrelationship between the three field experience scales themselves, to confirm or reject the a-priory expectations of these associations (see current study for details). the expected positive association between opportunities to learn through observation and opportunities to learn through own practice on the one side and level of field practice just completed on the other was only partially met. only opportunities to learn through own practice was positively correlated to the level of field practice just completed, while opportunities to learn through observation, on the other hand, was negatively correlated to the level of field practice just completed. that inexperienced students would observe more than more experienced students might seem an obvious finding. however, when considering the fact that student teachers become more advanced thinkers and nielsen 43 | f l r learners through the education, as they learn methods of observation and methods of reflection on these observations as applied to their own pupils, it can also be interpreted as a negative finding. negative, because the students’ advancement as learners are not exploited through increased observation of other teachers, and thus the scaffolding of their teaching self-efficacy through vicarious experience and reinforcement (bandura et al., 1963) is not optimized. in previous research on observation by student teachers, the focus is more often on the benefit for the student who is being observed and receives feedback from the observing student (e.g. baeten & simons, 2014), and not the vicarious benefit of the observing students. the expected lack of a direct association between opportunities to learn through feedback and the level of field practice was confirmed. thus, opportunities to learn through receiving feedback is only associated with the level of field practice indirectly through opportunities to learn through own practice. while the correlation between opportunities to learn through own practice and receiving feedback on this practice is very high, it is also apparent that not all students are provided this opportunity to learn (c.f. figure 1 and 3). this finding is in line with the findings of nielsen and graf (2021), who found that between 9% and 30% of 345 danish student teachers reported not having received feedback on 11 of the 12 teacher and teaching activities they had practised during the field practice placement. thus, it appears that not all student teachers are provided with the classical opportunity of learning and relearning through a feedback-feedforward loop (hermansen, 2003) in relation to their own practice, but rather some are left to construct their own learning based on their practice. baeten and simons (2014) found that student teachers benefitted from being observed when enacting and practicing teaching and subsequently receiving feedback on this practice from fellow student teachers. this is further supported by hill and grossman’s work on learning through teacher observation, which promotes observation and feedback as significant tools in teacher education as well as teachers’ post-degree development (hill & grossman, 2013). all three types of opportunities to learn were conditionally independent of gender as well as conditionally independent of age given the level of field practice. while this cannot be considered evidence of equity, it shows that there is no evidence of gender or age inequity in the opportunities for learning through observation, own practice and receiving feedback on this practice in the field practice placements of the teacher education. thus, the teacher education programme appears on track to comply with the education 2030 framework for action (unesco, 2015) on gender and age equity with regard to the field practice placement parts of the programme. research has previously been done on the subject of equity, when it concerns the academic knowledge parts of teacher education aimed at enabling future teacher to ensure gender equity in their own teaching of pupils (e.g. kollmayer et al., 2020; lucaspalacios et al., 2022). however, no such research has been identified in relation to teacher training programmes ensuring gender equity in the teaching or opportunities to learn of the student teachers themselves. nielsen (2021) called for an extension of the fpe-dk instrument to provide better coverage in regard to a wider set of teacher and teaching-related skills objectives for field practice in the danish teacher education, thus covering more and more diverse opportunities to learn in this context. in the summer of 2023, a reform of the teacher education programme was implemented in denmark and field practice will be expanded. thus, while the 12-item version of the three field practice experience scales provided new insights into the opportunities to learn while in field practice in the current danish teacher education, it is expected that an expanded version of the fpe-dk will provide new insights into the reformed teacher education field practice. in addition, the 12-item version of the fpe-dk can be used for comparative studies of the reformed teacher education field practice with the current one, and it can be used for international comparisons, as it is available in english. 4.1. limitations one limitation of the study consists of the predefinition of the recursive structure of the model. thus, it might be argued that the three variables denoting the level of field practice, the type of teacher nielsen 44 | f l r programme followed and the campus in which students were enrolled, should have been placed at the same recursive level in the model, thus allowing them to have undirected associations. the choice of the current structure was based on the time-wise presence of these variables, as campus and the type of teacher education programme is determined already upon application, and level of field practice follows after that. another circumstance of the study, which could be considered a limitation, is that the teaching subject(s) of the students were not included as a background variable in the model. as the danish student teachers obtain teaching competence in usually three (sometimes only two) subjects out of all the subjects taught in the public schools, and these subjects are followed at different times and in varying sequences in the teacher education programme, it would have required many questions in the survey to construct a variable reflecting this. it was deemed unlikely that such a variable would be useful in the analysis, as it would have too many categories to makes sense. however, in hindsight, information on which of the two “forced choice” subjects that was chosen by each student (i.e. danish language and mathematics) could have been included and may have added to the model. another limitation of the study consists of the low reliability of the own practice scale compared to the other scales. it is not an easily remedied limitation, as it stems from the lower variability in the study sample on this scale compared to the other scales (figure 1). inclusion of student teachers from more university colleges and thus other field practice schools might increase the variability in the scores somewhat. however, as the main activity in field practice placements in the teacher education is to practise teaching and other teaching-related skills, it is likely that the own practice scores will remain right skewed as variability will not increase. it is more likely that the low reliability could be remedied by extending the scales, as already suggested by nielsen (2021). a last limitation consists of the limited scope of the findings, as they currently address the danish teacher education and may extend to countries with a similar structure of teacher education (e.g. sweden and norway). this is a very common limitation in research on teacher education, as most is situated in a single-country context. however, as the instrument is available in english and is easily translated to other languages, future research could have a wider and a cross-cultural scope, by including countries in the model. keypoints opportunities to learn through observation, own practice and feedback on practice, while in teacher education field practice is studied such opportunities to learn have not previously been studied to determine their relationships with the progression of teacher education programmes chain graph models for ordinal data are used to study these relationships in the danish teacher education context opportunities to learn through observation of fellow students and other teachers declined as the programme progressed opportunities to learn through own practice increased as the programme progressed acknowledgments participating students are acknowledged for their willingness to participate. field practice coordinator gitte gorm larsen is acknowledged for her continued inspiration to conduct this work. nielsen 45 | f l r ethics no ethical approval is needed in denmark for research involving only survey data. participating students were informed of their right to withdraw from the study at any time prior to data anonymization, as well as of all other rights and of how their data would be treated in accordance with current european data protection regulations. data availability data is available at zenodo.org, doi: 10.5281/zenodo.8123855 references baeten, m. & simons, m. (2014). student teachers’ team teaching: models, effect, and conditions for implementation. teaching and teacher education, 41, 92-110. https://doi.org/10.1016/j.tate.2014.03.010 bandura, a., ross, d. & ross, s. (1963). vicarious reinforcement and imitative learning. journal of abnormal and social psychology, 67(6), 601–607. https://doi.org/10.1037/h0045550 benjamini, y., & hochberg, y. (1995). controlling the false discovery rate: a practical and powerful approach to multiple testing. journal of the royal statistical society: series b (methodological), 57(1), 289-300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x canrinus, e. t., bergem, o. k., klette, k., & hammerness, k. (2017). coherent teacher education programmes: taking a student perspective. journal of curriculum studies, 49(3), 313-333. https://doi.org/10.1080/00220272.2015.1124145 cohen, j. & berlin, r. (2020). what constitutes an “opportunity to learn” in teacher preparation? journal of teacher education, 71, 434–448. https://doi.org/10.1177/0022487119879893 cox, d. r., spjøtvoll, e., johansen, s., van zwet, w. r., bithell, j. f., barndorff-nielsen, o., & keuls, m. (1977). the role of significance tests [with discussion and reply]. scandinavian journal of statistics, 4(2), 49-70. https://www.jstor.org/stable/4615652 davis, j. a. (1967). a partial coefficient for goodman and kruskal’s gamma. journal of the american statistical association, 62, 189 193. https://doi.org/10.2307/2282922 goodman, l.a. & wallis, w. h. (1954). measures of association for cross classifications. journal of the american statistical association, 49, 732–764. https://doi.org/10.2307/2281536 grossman, p., hammerness, k. m., mcdonald, m., & ronfeldt, m. (2008). constructing coherence: structural predictors of perceptions of coherence in nyc teacher education programs. journal of teacher education, 59(4), 273-287. https://doi.org/10.1177/002248710832212 hammerness, k., klette, k., jenset, i. s., & canrinus, e. t. (2020). opportunities to study, practice, and rehearse teaching in teacher preparation: an international perspective. teachers college record, 122(11), 1–46. https://doi.org/10.1177/016146812012201108 hermansen, m. (2005). re-learning. cbs press, copenhagen. isbn 9788763001687 hill, h. & grossman, p. (2013). learning from teacher observations: challenges and opportunities posed by new teacher evaluation systems. harvard educational review, 83(2), 371-384. https://doi.org/10.17763/haer.83.2.d11511403715u376 kollmayer, m., schultes, m-t., lüftenegger, m., finsterwald, m., spiel, c. & schober, b. (2020). reflect – a teacher training program to promote gender equality in schools. frontiers in education, 5(136). https://doi.org/10.3389/feduc.2020.00136 kreiner, s. (1986). computerized exploratory screening of large dimensional contingency tables. compstat. heidelberg: physica verlag, pp. 43–48. kreiner, s. (1987). analysis of multidimensional contingency tables by exact conditional tests: techniques and strategies. scandinavian journal of statistics, 14, 97–112. https://www.jstor.org/stable/4616054 kreiner, s. (2003). introduction to digram. research report 03/10. copenhagen. http:// publichealth.ku.dk/sections/biostatistics/reports/2003/ https://doi.org/10.1016/j.tate.2014.03.010 https://doi.org/10.1037/h0045550 https://doi.org/10.1111/j.2517-6161.1995.tb02031.x https://doi.org/10.1080/00220272.2015.1124145 https://doi.org/10.1177/0022487119879893 https://www.jstor.org/stable/4615652 https://doi.org/10.2307/2282922 https://doi.org/10.2307/2281536 https://doi.org/10.1177/002248710832212 https://doi.org/10.1177/016146812012201108 https://doi.org/10.17763/haer.83.2.d11511403715u376 https://doi.org/10.3389/feduc.2020.00136 https://www.jstor.org/stable/4616054 nielsen 46 | f l r kreiner, s., petersen, j. h. & siersma, v. (2009). deriving and testing hypotheses in chain graph models. research report 09/09. copenhagen. https://ifsv.sund.ku.dk/biostat/annualreport/images/3/34/researchreport-2009-9.pdf larsen, g. g. (2021). praktikhåndbog læreruddannelsen på fyn: studieår 2021-2022 [handbook of field practice – the teacher education in funen: academic year 2021/2022]. https://www.ucviden.dk/ws/portalfiles/portal/124426918/praktikh_ndbog_l_reruddannelsen_p_fyn_20 21_inkl._bilag.pdf lauritzen, s. l. (1996). graphical models. london: clarendon press. isbn: 9780198522195 lucas-palacios, l., garcía-luque, a. & delgado-algarra, e. j. (2022). gender equity in initial teacher training: descriptive and factorial study of students' conceptions in a spanish educational context. international journal of environmental research in public health, 19(14). https://doi.org/10.3390/ijerph19148369 menter, i. (2022). teacher education research in the twenty-first century. in ian. menter (ed.), the palgrave handbook of teacher education research. palgrave macmillan, cham. pp. 1-29. https://doi.org/10.1007/978-3-030-59533-3_85-1 ministry of education and research (2015). bekendtgørelse om uddannelsen til professionsbachelor som lærer i folkeskolen. bek nr 1068 af 08/09/2015. https://www.retsinformation.dk/eli/lta/2015/1068 more, r. (2003). reexamining the field experiences of preservice teachers. journal of teacher education, 54(1), 31-42. https://doi.org/10.1177/0022487102238 nielsen, t. (2021). psychometric evaluation of the danish language version of the field practice experiences questionnaire for teacher students (fpe-dk) using item analysis according to the rasch model. plos one, 16(10):e0258459. https://doi.org/10.1371/journal.pone.0258459 nielsen, t. & graf, s. t. (2021). muligheder for læring i læreruddannelsespraktikken ved ucl erhvervsakademi og professionshøjskole. ucl erhvervsakademi og professionshøjskole. isbn: 97887-93067-54-7. https://www.ucviden.dk/ws/portalfiles/portal/138101111/muligheder_for_l_ring_i_l_reruddannelsespra ktik_191021_1_.pdf ucl erhvervsakademi og professionshøjskole (2022). studieordning læreruddannelsen. insitutionel del [teacher education program curriculum, university college specific part]. https://esdhweb.ucl.dk/d22-2039047.pdf?_ga=2.71524627.1269598887.16795725851669244073.1617210971 ulvik, m., eide, l., helleve, i., & kvam, e. k. (2021). praksisopplæringens oppfattende og erfarte formål sett fra ulike aktørperspektiv. nordisk tidsskrift for utdanning og praksis, 15(3), 19-37. https://doi.org/10.23865/up.v15.2949 unesco (2015). incheon declaration and framework for action for the implementation of sustainable development goal 4. ensure inclusive and equitable quality education and promote lifelong learning opportunities for all. https://uis.unesco.org/sites/default/files/documents/education-2030-incheonframework-for-action-implementation-of-sdg4-2016-en_2.pdf weisdorf, a. k. (2020). læreruddannelsen i globalt perspektiv—et komparativt studie af læreruddannelsen i danmark, england, finland, holland, new zealand, norge, ontario, singapore, sverige og tyskland [teacher education i a global perspective – a comparative study of teacher education in denmark, the uk, finland, the netherlands, new zealand, norway, ontario, singapore, sweden and germany]. denmark: danske professionshøjskoler. https://xn--danskeprofessionshjskoler-xtc.dk/wpcontent/uploads/2022/01/laereruddannelsen-i-globalt-perspektiv.-et-komparativt-studie.pdf youngs, p., elreda, l. m., anagnostopoulos, d., cohen, j., drake, c. & konstantopoulos, s. (2022). the development of ambitious instruction: how beginning elementary teachers’ preparation experiences are associated with their mathematics and english language arts instructional practices. teaching and teacher education, 110. https://doi.org/10.1016/j.tate.2021.103576 https://ifsv.sund.ku.dk/biostat/annualreport/images/3/34/researchreport-2009-9.pdf https://www.ucviden.dk/ws/portalfiles/portal/124426918/praktikh_ndbog_l_reruddannelsen_p_fyn_2021_inkl._bilag.pdf https://www.ucviden.dk/ws/portalfiles/portal/124426918/praktikh_ndbog_l_reruddannelsen_p_fyn_2021_inkl._bilag.pdf https://doi.org/10.3390/ijerph19148369 https://doi.org/10.1007/978-3-030-59533-3_85-1 https://www.retsinformation.dk/eli/lta/2015/1068 https://doi.org/10.1177/0022487102238 https://doi.org/10.1371/journal.pone.0258459 https://www.ucviden.dk/ws/portalfiles/portal/138101111/muligheder_for_l_ring_i_l_reruddannelsespraktik_191021_1_.pdf https://www.ucviden.dk/ws/portalfiles/portal/138101111/muligheder_for_l_ring_i_l_reruddannelsespraktik_191021_1_.pdf https://esdhweb.ucl.dk/d22-2039047.pdf?_ga=2.71524627.1269598887.1679572585-1669244073.1617210971 https://esdhweb.ucl.dk/d22-2039047.pdf?_ga=2.71524627.1269598887.1679572585-1669244073.1617210971 https://doi.org/10.23865/up.v15.2949 https://uis.unesco.org/sites/default/files/documents/education-2030-incheon-framework-for-action-implementation-of-sdg4-2016-en_2.pdf https://uis.unesco.org/sites/default/files/documents/education-2030-incheon-framework-for-action-implementation-of-sdg4-2016-en_2.pdf https://xn--danskeprofessionshjskoler-xtc.dk/wp-content/uploads/2022/01/laereruddannelsen-i-globalt-perspektiv.-et-komparativt-studie.pdf https://xn--danskeprofessionshjskoler-xtc.dk/wp-content/uploads/2022/01/laereruddannelsen-i-globalt-perspektiv.-et-komparativt-studie.pdf https://doi.org/10.1016/j.tate.2021.103576 nielsen 47 | f l r appendix table a1 testing the necessity of the associations in the final chain graph model. testing 13 separation hypotheses related to existing edges ------------------------------------------------------------------------------------------- p-values p-values (1-sided) 95% confidence hypothesis x² df asymp exact gamma asymp exact interval nsim n ------------------------------------------------------------------------------------------- 1:a&c|b 913.3 712 0.000 0.005 0.25 0.000 0.000 [0.15 0.34] 1000 554 xx ++ 2:a&c|d 637.4 432 0.000 0.000 0.36 0.000 0.000 [0.30 0.43] 1000 560 xx ++ 3:a&d|be 250.9 206 0.018 0.031 -0.34 0.000 0.000 [-0.44 -0.24] 1000 555 x - 4:a&d|ce 365.9 312 0.019 0.033 -0.40 0.000 0.000 [-0.50 -0.29] 1000 558 x - 5:a&e|d 34.0 36 0.565 0.578 -0.14 0.053 0.041 [-0.31 0.03] 1000 560 6:b&c|a 1415.1 789 0.000 0.000 0.67 0.000 0.000 [0.59 0.75] 1000 560 xx ++ 7:b&c|d 1030.2 312 0.000 0.000 0.67 0.000 0.000 [0.62 0.73] 1000 560 xx ++ 8:b&d|a 188.7 158 0.048 0.044 0.38 0.000 0.000 [0.26 0.50] 1000 560 x ++ 9:b&d|c 137.3 130 0.314 0.439 0.12 0.046 0.049 [-0.02 0.27] 1000 560 10:d&f|h 17.6 6 0.007 0.007 0.22 0.003 0.006 [0.06 0.37] 1000 560 xx ++ 11:d&h|f 87.8 8 0.000 0.000 0.39 0.000 0.000 [0.28 0.50] 1000 560 xx ++ 12:e&f 36.6 1 0.000 0.000 0.62 0.000 0.000 [0.37 0.86] 1000 560 xx ++ 13:e&h 12.0 2 0.003 0.003 0.33 0.000 0.000 [0.14 0.52] 1000 560 xx ++ ------------------------------------------------------------------------------------------- benjamini hochberg rejects if p < 0.044 for fdr = 0.05 and p < 0.007 for fdr = 0.01 significance of x² xx : fdr = 0.01 x : fdr = 0.05 gamma ++/-: fdr = 0.01 +/: fdr = 0.05 nielsen 48 | f l r table a2 testing the adequacy of the associations in the final chain graph model. testing 34 separation hypotheses relating to missing edges ------------------------------------------------------------------------------------------- p-values p-values (1-sided) 95% confidence hypothesis x² df asymp exact gamma asymp exact interval nsim n ------------------------------------------------------------------------------------------- 1:a&b|cd 835.4 766 0.041 0.614 0.10 0.049 0.041 [-0.02 0.22] 1000 529 2:a&f|de 73.0 67 0.288 0.287 0.14 0.015 0.016 [0.01 0.27] 1000 560 3:a&g|de 75.4 67 0.224 0.216 0.13 0.022 0.019 [0.00 0.25] 1000 560 4:a&h|de 131.7 134 0.541 0.610 -0.08 0.050 0.043 [-0.18 0.02] 1000 560 5:b&e|ad 109.8 121 0.758 0.810 0.04 0.378 0.524 [-0.19 0.26] 21 517 6:b&e|cd 117.4 97 0.078 0.157 0.14 0.158 0.196 [-0.14 0.43] 102 456 7:b&f|ad 130.3 135 0.599 0.875 0.07 0.208 0.292 [-0.10 0.24] 24 549 8:b&f|cd 88.8 102 0.821 1.000 -0.05 0.328 0.333 [-0.28 0.18] 21 477 9:b&f|de 35.1 40 0.691 0.429 0.01 0.439 0.476 [-0.14 0.16] 21 560 10:b&g|ad 139.3 137 0.430 0.750 0.08 0.167 0.208 [-0.08 0.24] 48 554 11:b&g|cd 124.1 102 0.068 0.177 0.13 0.100 0.145 [-0.07 0.34] 124 491 12:b&g|de 47.0 40 0.206 0.217 0.13 0.026 0.030 [-0.00 0.27] 1000 560 13:b&h|ad 273.4 267 0.381 0.762 0.01 0.433 0.429 [-0.12 0.14] 21 555 14:b&h|cd 259.6 229 0.080 0.381 -0.05 0.260 0.476 [-0.20 0.10] 21 523 15:b&h|de 75.7 80 0.615 0.571 -0.05 0.180 0.286 [-0.16 0.06] 21 560 16:c&d|ab 347.0 302 0.038 0.667 0.06 0.277 0.292 [-0.13 0.25] 24 509 17:c&e|ab 150.1 120 0.033 0.204 -0.15 0.183 0.204 [-0.47 0.17] 54 382 18:c&e|ad 207.9 191 0.190 0.381 -0.03 0.396 0.333 [-0.23 0.18] 21 517 19:c&f|ab 186.0 160 0.078 0.429 0.09 0.237 0.286 [-0.15 0.33] 21 469 20:c&f|ad 223.7 211 0.262 0.691 0.11 0.077 0.079 [-0.04 0.27] 1000 549 21:c&f|de 77.2 61 0.079 0.062 0.09 0.081 0.100 [-0.04 0.23] 1000 560 22:c&g|ab 177.1 162 0.197 0.810 0.07 0.300 0.333 [-0.18 0.32] 21 476 23:c&g|ad 224.9 215 0.307 0.714 0.07 0.175 0.333 [-0.08 0.22] 21 554 24:c&g|de 81.1 61 0.043 0.034 0.14 0.014 0.015 [0.01 0.26] 1000 560 25:c&h|ab 404.7 326 0.002 0.163 -0.04 0.344 0.380 [-0.22 0.15] 92 531 26:c&h|ad 456.4 416 0.084 0.289 0.04 0.245 0.237 [-0.07 0.16] 38 555 27:c&h|de 135.1 122 0.197 0.225 0.00 0.492 0.525 [-0.10 0.10] 40 560 28:d&e|fh 6.6 12 0.881 0.952 0.07 0.289 0.333 [-0.17 0.30] 21 560 29:d&g|fh 16.2 12 0.184 0.190 -0.13 0.062 0.058 [-0.29 0.04] 1000 560 30:e&g|h 4.2 3 0.239 0.333 0.01 0.481 0.571 [-0.25 0.27] 21 560 31:f&g|e 2.8 2 0.252 0.228 -0.19 0.049 0.054 [-0.41 0.03] 1000 560 32:f&g|h 3.2 3 0.368 0.317 -0.12 0.128 0.118 [-0.32 0.09] 246 560 33:f&h|e 5.4 4 0.249 0.242 0.04 0.321 0.333 [-0.13 0.22] 33 560 34:g&h 3.1 2 0.209 0.190 0.11 0.072 0.072 [-0.04 0.25] 1000 560 ------------------------------------------------------------------------------------------- benjamini hochberg rejects if p < 0.001 for fdr = 0.05 and p < 0.000 for fdr = 0.01 significance of x² xx : fdr = 0.01 x : fdr = 0.05 gamma ++/-: fdr = 0.01 +/: fdr = 0.05 frontline learning research vol. 12 no. 1 (2024) 66 123 issn 2295-3159 corresponding authors: ayşenur alp christ, freiestrasse 36, 8032 zürich, switzerland, aysenur.alpchrist@uzh.ch & vanda capon-sieber, freiestrasse 36, 8032 zürich, switzerland, vanda.sieber@ife.uzh.ch doi: https://doi.org/10.14786/flr.v12i1.1349 revisiting the three basic dimensions model: a critical empirical investigation of the indirect effects of student-perceived teaching quality on student outcomes ayşenur alp christ1*, vanda capon-sieber1*, carmen köhler2, eckhard klieme2 & anna-katharina praetorius1 1 university of zurich, switzerland 2 leibniz institute for research and information in education (dipf), germany *the first two authors have a shared contribution in conceptualization, writing, and revising the manuscript and should be considered co-first authors. article received 20 august 2023 / article revised 30 november 2023 / accepted 8 february 2024 / available online 8 march abstract the three basic dimensions model, theorizes three mediators for the effect of teaching quality dimensions on student outcomes. however, the proposed mediating paths and their effects have largely not been empirically tested. this study investigated the mediating role of depthof-processing, time-on-task, and need satisfaction between student-perceived teaching quality and student mathematics achievement and interest, expanding the tbd model to include mediation paths suggested by theories of motivation, cognition, and effort. data from the talis video study for germany, comprising 958 secondary school students in 41 classrooms, were used to run multilevel longitudinal and correlational mediation analyses. the results only found mediation effects at the student level; there were no mediating effects at the classroom level. not all of the hypothesized relationships thought to exist between the mediators and achievement and interest outcomes were confirmed. the conceptual sequence of the variables, the choice of correlational vs. longitudinal evidence, and the level of analysis were all shown to have an impact on the results. the study thus confirms some of the assumptions of the tbd model, identifies new paths between teaching quality and student outcomes, and provides suggestions for how to proceed with further investigation of a model which should be expanded and more thoroughly empirically tested. keywords: teaching quality, learning processes, mediation, interest, achievement, three basic dimensions mailto:aysenur.alpchrist@uzh.ch mailto:vanda.sieber@ife.uzh.ch https://doi.org/10.14786/flr.v12i1.1349 alp christ, capon-sieber et al. 67 | f l r 1. introduction the three basic dimensions (tbd) model of teaching quality is influential and widely used by researchers in the field of teaching quality, particularly in german-speaking countries (klieme et al., 2006, 2009; kunter & trautwein, 2013; praetorius et al., 2018; reusser et al., 2010). this model, developed by klieme et al. (2001, 2009), offers a concise framework for understanding the aspects of teaching quality by categorizing them into three key dimensions: cognitive activation, classroom management, and student support. one of its principal advantages over the many other models and frameworks is that it integrates student learning processes, focusing on the mediating role that they play between teaching quality and student outcomes. for example, it hypothesizes that depth of processing mediates between cognitive activation and student achievement, which suggests that cognitive activation only has a significant impact on learning outcomes when students engage in deep processing (klieme et al., 2006, see figure 1). however, researchers have rarely conducted systematic empirical examinations of the assumed mediators. although the role of mediators has been supported by results from studies focusing on specific paths in the model (for cognitive activation, e.g., hiebert & grouws, 2007; stein & lane, 1996; for classroom management, e.g., hospel & galand, 2016; kunter, et al., 2007; and for student support, e.g., kiemer et al., 2018; mouratidis et al., 2013), overall, empirical research on these mediators remains very limited. the incorporation of mediators between teaching quality and student outcomes in the tbd model was guided by selected theoretical considerations, primarily rooted in self-determination theory (sdt, ryan & deci, 2000) and constructivism (de corte, 2004; pauli & reusser, 2006). this selection of references may miss other valid theoretical perspectives that might explain the mediators and observed student outcomes. moreover, the reasoning behind the choice of theory to explain mediation pathways in the tbd model is not well-articulated in the literature. the lack of clarity creates potential gaps in our understanding of the model and could lead to an incomplete representation of the role of mediators between teaching quality and student outcomes. it is therefore important to consider the possibility of additional theoretically likely relations between the variables in the model. for example, in the context of the original model, depth of processing is influenced by cognitive activation and classroom management. however, by incorporating theoretical insights from other theories, such as the elaboration likelihood model (elm, petty & cacioppo, 1986), we propose that elements of student support, such as activities that accentuate the relevance of tasks, may also contribute to increased depth of processing. this paper aims to address these gaps in the theoretical and empirical foundations of the tbd model. it seeks to comprehensively test the assumptions of the model, including additional possible mediation pathways, to provide a more robust understanding of the relationship between teaching quality and student outcomes. 1.1 the three basic dimensions model of teaching quality the tbd model (figure 1) identifies cognitive activation, classroom management, and student support as the key aspects of teaching quality that affect student outcomes such as achievement and motivation. in particular, cognitive activation and classroom management are assumed to have an effect on student achievement and student support is linked to student motivation. the results of an empirical analysis conducted by klieme et al. (2001) provide support for this idea. emphasizing the role of student understanding, attentiveness, and motivation in the learning process (diederich & tenorth, 1997), the basic dimensions have been theoretically linked to students’ depth of processing, time-on-task, and need satisfaction (i.e., student use of learning opportunities) (klieme et al., 2006, 2009; klieme & rakoczy, 2008; see figure 1). specifically, it has been hypothesized that cognitive activation is linked to depth of processing, that student support has an effect on need satisfaction, and that classroom management is linked to depth of processing, time-on-task, and need satisfaction. for simplicity and parsimony, the original tbd model focused on pathways that included mediators between teaching quality and student outcomes based on the specific theoretical considerations used to formulate the basic dimensions (de alp christ, capon-sieber et al. 68 | f l r corte, 2004; pauli & reusser, 2006; ryan & deci, 2000). to better explain how these dimensions relate to the use of learning opportunities by students, we first describe the basic dimensions in section 1.1.2, then explain the mediators and outcomes in detail. figure 1. the relations between the three basic dimensions and student achievement and motivation according to the tbd model (adapted from klieme et al., 2009). 1.1.2 the three basic dimensions of teaching quality cognitive activation. this dimension is based on different (socio-)constructivist learning theories (aebli, 2011; piaget, 1992; vygotsky, 1978), which emphasize the independent construction of knowledge and interaction with others within the zone of proximal development (zpd, vygotsky, 1978). the current understanding of cognitive activation encompasses multiple facets that aim to stimulate higher-order cognitive processes (lipowsky & hess, 2019; ziegelbauer, 2009). this includes encouraging students to understand learning content by providing challenging tasks and to activate prior knowledge, practicing content-related discourse, and fostering active participation in critical class discussions (förtsch et al., 2018; klieme et al., 2009; lipowsky et al., 2009; lotz, 2016; praetorius et al., 2014, 2018; rakoczy & pauli, 2006). some authors propose that the teaching behaviors should support students’ independent engagement with the learning content (lotz, 2016), aspects of selfregulation and metacognition (praetorius et al., 2018; rieser et al., 2016). classroom management. this dimension encapsulates the “effective strategies for organizing classrooms” proposed by several researchers (doyle, 1986; emmer & stough, 2001; evertson, 1989; kounin, 1970a, 1970b; kunter et al., 2007). the strategies result in increased learning time. this is, among others, the result of the “withitness” of a teacher, which means that a teacher is omnipresent during a lesson and informed about all that is happening in the classroom. with efficient time use, making effective transitions between topics and having clear rules and routines, a teacher can ensure the smooth running of the classroom. successful classroom management also includes early, prompt, intervention to prevent disruptions and discipline problems (kounin, 1970a; kuger, 2016; praetorius et al., 2018). student support. this dimension is based on sdt (ryan & deci, 2000, 2017) and comprises the support of student competence, autonomy, and relatedness (klieme et al., 2009). student support includes giving constructive feedback, addressing student errors and misconceptions in a positive manner, and nurturing an atmosphere of mutual care and respect in the classroom (see fauth et al., 2014, alp christ, capon-sieber et al. 69 | f l r 2019; lipowsky et al., 2009; praetorius et al., 2018). according to sdt, it also involves understanding student needs, helping them when needed, providing them with suitable options and explaining the relevance of the tasks. 1.1.3 mediators between teaching quality and student outcomes depth of processing. based on cognitive constructivist learning theory (de corte, 1995), depth of processing, or high-level thinking, is a student’s reaction to cognitively activating teaching (klieme & rakoczy, 2008). the concept of depth of processing – the level at which a student processes what they are taught – encompasses critical thinking, reasoning, making sense, finding patterns, solving nonroutine problems, as well as some aspects of self-regulation and metacognition (baumert et al., 2010; boston & candela, 2018; klieme et al., 2009; lipowsky et al., 2009; praetorius et al., 2018). mathematics teaching, in particular, should incorporate challenging tasks that are neither too easy nor too hard so that students can develop an in-depth understanding of concepts, not just memorize facts (hiebert & grouws, 2007; silver & stein, 1996; stein et al., 1996; stein & lane, 1996). depth of processing has been empirically linked to student achievement (e.g., chi & wylie, 2014; clifford, 1990; lipowksy et al., 2009) and conceptual development (stein et al., 1996; stein & lane, 1996). in the tbd model depth of processing mediates the relation between cognitive activation and student achievement, and classroom management is assumed to be directly related to depth of processing since a learning environment that helps students to pay attention is seen as an important prerequisite for in-depth engagement with a task (e.g., lipowsky & hess, 2019). time-on-task. time-on-task is the class time during which students are actually engaged in learning activities contributing to learning gains and performance (brophy, 2006; emmer & stough, 2001; finn & zimmer, 2012; fisher et al., 1981; rakoczy, 2006; wang et al., 1993). in the tbd model, time-on-task is a response to classroom management, which in turn is a strong predictor of student learning and achievement (böheim et al., 2020; brophy, 2000; hattie, 2009; klieme et al., 2009; seidel & shavelson, 2007). need satisfaction. research based on sdt resulted in the addition of the satisfaction of the three basic needs for autonomy, competence, and relatedness as a mediator between student support and motivation (klieme & rakoczy, 2003). the need for autonomy is the need to experience personal freedom, volition, and choice (vansteenkiste et al., 2010). the need for competence is the student’s desire for mastery and effectiveness during tasks (ryan & deci, 2002). the need for relatedness refers to the desire for close and warm relationships (baumeister & leary, 1995; deci & ryan, 2002). according to sdt teaching behaviors can influence whether student needs are satisfied (black & deci, 2000). additionally, within the tbd model classroom management is an important prerequisite for the satisfaction of students’ basic needs because, for example, well-organized, undisturbed classrooms may mean students feel more effective when performing tasks (kunter et al., 2007). 1.2 revisiting the tbd model the tbd model assumes that classroom management has an influence on all three mediators (depth of thinking, time-on-task, and need satisfaction), and cognitive activation and student support affect depth of processing and need satisfaction respectively (see figure 1). there is empirical evidence for the role played by single mediators (e.g., hiebert & grouws, 2007; stein & lane, 1996 for cognitive activation; hospel & galand, 2016; kunter, et al., 2007 for classroom management; and kiemer et al., 2018; mouratidis et al., 2013 for student support). however, the current tbd model proposes a complex web of influences and theoretical assumptions which have been added incrementally over time. it is therefore important to periodically review and possibly revise these assumptions and the mediation paths proposed by klieme et al. (2006). the need for a review has been underscored by recent evidence that several of the assumptions are not empirically supported (praetorius et al., 2018). therefore, robust model and theory building warrants a thorough revisit and in-depth investigation of the entire tbd model (praetorius et al., 2020a). alp christ, capon-sieber et al. 70 | f l r section 1.2.1 is a discussion of the possible alternate paths derived from established theories of motivation and cognition, such as expectancy-value theory (evt; wigfield & eccles, 1992), that were not explicitly considered in the formulation of the tbd model but have considerable overlap with its core assumptions. relevant theories were systematically selected, by using the definitions of the dimensions and mediators within the tbd model and conducting a literature search for studies that assessed those constructs, including their subdimensions. klieme et al. (2009) highlighted that while constructivist ideas play a crucial role in understanding teaching quality, they alone cannot fully explain the utilization of learning opportunities and the reasons behind such usage. hence, the integration of motivational and cognitive theories seems essential to comprehensively grasp the learning processes involved. the objective was to improve the theoretical basis of the tbd model and provide a more comprehensive understanding of the underlying processes that affect how teaching quality impacts student outcomes. when the theoretical views and their empirical insights were incorporated into the tbd model, it became evident that additional mediation paths may exist. for example, several theories in the domain of achievement motivation, such as interest theory (it; hidi & renninger, 2006) and the control-value theory of achievement emotions (cvt; pekrun, 2006), suggest that an optimal challenge or even being engaged in a task may affect not only achievement, but also motivational and cognitive processes (see for example vu et al., 2022; wentzel & miele, 2016). we elaborate on these additional assumptions in the following. figure 2. possible relationships between the different parts of the tbd model. 1.2.1 mediating paths for cognitive activation the tbd model assumes a relation between cognitive activation and depth of processing (figure 2, path-a). however, theoretical and empirical evidence suggests that cognitive activation might also affect time-on-task (figure 2, path-b) and need satisfaction (figure 2, path-c). cognitively challenging activities or tasks can direct student attention to particular aspects of content and specify methods by which information is processed and thus influence time-on-task (doyle, 1983). this idea was also explored for mathematics teaching by stein et al. (1996). for example, when a teacher asks questions or presents problems without obvious solutions, students are more likely to pay close attention. these arguments are consistent with influential views on achievement motivation. according to evt student behavior can be seen as a product of the expectancy of success and value of reward (atkinson, 1957; heckhausen, 1991; wigfield & eccles, 1992). the theory of motivational intensity (mit) distinguishes between mere willingness to engage in a task and actual effort (brehm & self, 1989; richter et al., 2016). according to this theory, conditions are identified which determine how much resource is allocated for engaging in a task. moreover, a principle of resource conservation is proposed where it is assumed that even if the willingness to engage in a task is high, only as much effort as needed to succeed in a task will be allocated (brehm & self, 1989). if a task is very easy, effort will be low. when a task is too difficult or when the difficulty exceeds alp christ, capon-sieber et al. 71 | f l r the value of a given reward, a student is likely to disengage from the task, resulting in diminished timeon-task. given the fact that optimal task difficulty (hiebert & grouws, 2007), as well as adaptivity and individualization (helm, 2016; lotz, 2016; rakoczy & pauli, 2006) are often seen as parts of cognitive activation, an effect on time-on-task is also highly probable. theoretical and empirical evidence also suggests that cognitive activation can be related to students’ satisfaction of basic psychological needs (figure 2, path-c). according to evt (wigfield & eccles, 1992) and sdt (ryan & deci, 2017), when teachers give optimally challenging tasks, students’ expectancies for success can be fostered (evt; wigfield & eccles, 2002) and in a similar vein their competence need can be satisfied (sdt; reeve, 2006; 2016). similarly, the basic need for autonomy can be satisfied when teachers present non-routine problems, as it fosters students’ critical thinking and encourages them to solve the tasks using their own methods, which is an important aspect of autonomy in the classroom (sdt; reeve 2009; reeve & jang, 2006). if the students perceive the tasks as valuable and relevant, their basic psychological need for autonomy will be satisfied (sdt; reeve & jang, 2006). empirical studies based on sdt support this link. for example, cognitive activation indirectly affected student interest and self-efficacy through autonomy and competence need satisfaction (schukajlow et al., 2019; schukajlow & krug, 2014). another study argued that a potential underlying mechanism between cognitive activation and student enjoyment in mathematics could be autonomy and competence need satisfaction (lazarides & buchholz, 2019). moreover, cognitively activating behaviors such as aiming to foster independent engagement with the learning content, directly affect autonomy (lotz, 2016). since co-construction of knowledge is an important part of cognitive activation, the experience of relatedness could also be affected (see ryan & powelson, 1991; sun & chen, 2010 for the interplay and similarity of those constructs). 1.2.2 mediating paths for classroom management within the tbd model classroom management is expected to affect all three mediators. classroom management has been shown to affect time-on-task (emmer & stough, 2001; finn & zimmer, 2012; fisher et al., 1981; rakoczy, 2006; wang et al., 1993) (figure 2, path-e), aspects of cognitive engagement (i.e., use of learning and self-regulation strategies) (hospel & galand, 2016) (figure 2, path-d), and students’ need satisfaction (kunter et al., 2007) (figure 2, path-f). thus, classroom management should be relevant for all student learning processes (i.e., depth of processing, time-on-task, need satisfaction) in the classroom. 1.2.3 mediating paths for student support according to sdt, student support has a positive effect on the satisfaction of students’ basic psychological needs (ahn et al., 2021; deci & ryan, 2000; jang et al., 2012; kiemer et al., 2018; mouratidis et al., 2013; zhang et al., 2011) (figure 2, path-i). however, student support is also likely to be related to depth of processing (figure 2, path-g) and time-on-task (figure 2, path-h), which differs from what is postulated in the tbd model. by engaging in supportive teaching behavior, characterized by mutual respect, teachers actively promote a positive learning environment. students are not distracted by a negative teacher-student relationship that could elicit emotions that interfere with attention and self-regulation (blair, 2002; murray & pianta, 2007). a good relationship between teachers and students also allows students to actively participate in their learning environment (hughes et al., 2008; pianta & steinberg, 1992). similarly, by giving constructive feedback, approaching student errors and misconceptions in a positive way, and monitoring student progress, teachers increase active learning time (grabinger & dunlap, 1995; grabinger et al., 1997). cognitive information processing theory (ipt; atkinson & shiffrin, 1968; driscoll, 2005) states that students are attentive when they select and process information that is very important and meaningful for them. according to sdt, one key aspect of student support is making the information relevant and meaningful to the students (see also ahmadi et al., 2023). for example, when teachers engage in autonomy supportive behaviors such as providing rationales for the content and personal relevance, then students are more likely to pay attention during the lesson because the information is useful, meaningful, and important to them (lietaert et al., 2015). alp christ, capon-sieber et al. 72 | f l r a positive climate also allows students to try new and creative solutions without reservations (chan & yuen, 2014), an important aspect of depth of processing. this is because an encouraging, respectful, supportive, and positive learning environment that is open to creativity and improvement, encourages students to seek challenges (turner & meyer, 2004). in addition, according to the elaboration likelihood model (elm, petty & cacioppo, 1986) when teachers highlight the relevance of the tasks to students, the students’ personal involvement increases, which in turn fosters depth of processing (illies & reiter-palmon, 2004; petty et al., 1983; mitchell, 1993). interest theory has been used to describe the relation between personal involvement, depth of processing, and time-on-task (hidi & renninger, 2006; renninger & hidi, 2002). when a student’s attention is triggered by relevant tasks and personal involvement, they will also become interested in content. several studies confirm the link between aspects of student support and aspects of depth of processing such as self-regulation and deep learning strategies (hospel & galand, 2016; rieser et al., 2016; ruiz-alfonso & león, 2019; wang & eccles, 2013), higher analytical problem-solving skills, and student challenge preferences (boggiano et al., 1988; 1993; guay et al., 2008). positive relations have also been identified between student support and time-on-task (chiu, 2004; deci et al., 1994; stallings, 1980). all these studies lend weight to the hypothesis that student support can predict depth of processing and time-on-task. 1.2.4 student use of opportunities and student outcomes because the mediators are interrelated, the relationship between the mediators and the outcomes might also be less discrete than how they are shown in the original model (figure 1); the original model already indicated the relationship between motivational outcomes and achievement (figure 2, path-p). other important theories, such as cvt (pekrun, 2006), also suggest that depth of processing and timeon-task could be related to motivational outcomes (figure 2, paths k and m). for example, students who think critically and solve modelling problems by constructing multiple solutions have a greater interest in the subject (schukajlow & krug, 2014) and higher self-efficacy (schukajlow et al., 2019). interest and self-efficacy have been treated as motivational outcomes in tbd research (figure 2, path-k) (dorfner et al., 2018; fauth et al., 2014, 2019; förtsch et al., 2017; li et al., 2020). time-on-task not only promotes academic achievement (evertson & harris, 1992; good & brophy, 2003), but also appears to be relevant for fostering student motivation (butler & shibaz, 2008; lazarides & buchholz, 2019; rakoczy, 2006) (figure 2, path-m). this relationship is also suggested by other motivation theories such as interest theory and cvt. in these instances, it is hypothesized that being on task or processing information at a deep level creates positive emotions for students (i.e., activity emotions), which in turn fosters their interest and motivation. finally, as proposed in the tbd model, it is hypothesized that need satisfaction affects motivational outcomes which in turn affect achievement (figure 2, path-o-p). studies have shown a link between the satisfaction of a student’s needs and their autonomous motivation (e.g., mouratidis et al., 2015; ryan & deci, 2009), interest (e.g., kunter et al., 2007), and self-efficacy (e.g., sun et al., 2020; zhen et al., 2017) (figure 2, path-o). however, according to sdt, when students’ basic psychological needs are satisfied, they display improved academic performance and achievement (ryan & deci, 2017). theoretical considerations based on sdt, in combination with the studies which found positive relationships between need satisfaction and student achievement (badri et al., 2014; wang et al., 2019; zhou et al., 2021), lead us to hypothesize that need satisfaction is directly positively related not only to motivational outcomes, but also to achievement. depth of processing and time-on-task are likely to be linked to motivational outcomes (figure 2, paths k and m) and need satisfaction can be directly related to achievement (figure 2, path-n). 1.3 study a review of theories in the field of cognitive and motivational psychology and related empirical evidence strongly suggests that there should be more mediation paths than those which have been discussed in the tbd literature to date. our assumptions will be tested by constructing models which alp christ, capon-sieber et al. 73 | f l r consider the assumptions of the original tbd model and additional possible paths. our concrete hypotheses are as follows: h1: the three basic dimensions of teaching quality are all related to the development of student achievement and interest. h2: cognitive activation indirectly predicts the development of student achievement and interest through depth of processing, time-on-task, and need satisfaction. h3: classroom management indirectly predicts the development of student achievement and interest through depth of processing, time-on-task, and need satisfaction. h4: student support indirectly predicts the development of student achievement and interest through depth of processing, time-on-task, and need satisfaction. 2. method this study investigates whether student perceptions of cognitive activation, classroom management, and student support indirectly affect student achievement and interest in mathematics through depth of processing, time-on-task, and need satisfaction. we analyzed data collected in germany as a part of the teaching and learning international survey (talis) video study conducted by the organisation for economic co-operation and development (oecd, 2020). 2.1 participants and procedures the study sample was selected from participants in the talis video study for germany using convenience sampling. the initial sample consisted of 1143 students from 50 classrooms and 39 schools. there are big differences in learning goals, school curricula, student achievement levels, class composition and individual student characteristics between school tracks in germany (hachfeld & lazarides, 2020). as most participating classrooms were from the academic track (“gymnasium”) and we were interested in a homogeneous sample so that the data can be interpreted more unambiguously, we removed participants from all other school forms (e.g., lower-track schools and vocational schools). the final sample consists of 958 students from 41 classrooms and 30 schools (mage = 14.82, sd = 0.62; 50.5 % females; 5.3% did not report their gender). the average number of students per classroom was 23.37 (sd = 4.73, min = 11 max = 31). of 41 classrooms, the majority (35) were 9th grade level and six were 8th grade. most of the students reported that they were born in germany (n = 869), n = 36 students reported that they were born in other countries, and n = 53 did not report their country of birth. the talis video study conformed to ethical standards (oecd, 2020). school principals, teachers, students, and their parents were informed about the purpose of the study. the participants were assured that their participation was anonymous and voluntary and that their information would be secure and confidential. 2.2 instruments and measures the student survey asked about family and peer circumstances and aspects of students’ cognitive, motivational, and emotional learning. it also asked students for their perceptions of teaching quality in the mathematics lessons at the beginning of a specific teaching unit, quadratic equations (mccaffrey et al., 2020; praetorius et al., 2020b). in the talis video study, the constructs measured in the pre-test (t1) are operationalized in terms of mathematics in general, whereas the constructs measured in the post-test (t2) are operationalized only in terms of quadratic equations. to test our hypotheses, variables from the first and second measurement points were used. the unit included between 360 and 1080 minutes of lesson time (m = 797.20, sd = 173.59), spread over a period of 22 to 130 days (m = 58.83, sd = 26.01) (see supplementary material). alp christ, capon-sieber et al. 74 | f l r teaching quality dimensions were measured using student rating, which is considered a valid, reliable, and efficient measure of teaching quality (van der scheer et al., 2019). depth of processing, time-on-task, need satisfaction, and interest were assessed using student self-reports. self-reports are useful for assessing constructs that are not directly observable, such as student use of learning opportunities (appleton et al., 2006; fredricks & mccolskey, 2012). many items in the talis video study questionnaire were based on previous talis and programme for international student assessment (pisa) studies (oecd, 2020; praetorius et al., 2020 b ). the concrete item wordings of the assessed constructs are shown in appendix a. each item was assessed using a four-point likert scale. negative items in the questionnaire were reverse-coded. to account for level-specific reliability (geldhof et al., 2014), we calculated mcdonald’s omega (ω; mcdonald, 1999) for both the within and between levels; these are reported in table 1. we also calculated the descriptive statistics for each item and each subscale (see supplementary material). 2.2.1 independent variables: three dimensions of teaching quality (tbd) in the talis video study for germany, student perceptions of cognitive activation, classroom management, and student support were assessed using items similar to those used in previous talis and pisa studies (oecd, 2020). student-reported cognitive activation was assessed with seven items designed to reveal their perceptions of whether teachers presented tasks and their solutions in a manner that would promote conceptual understanding and content-based discourse (e.g., “our mathematics teacher gives tasks that require us to think critically”). student-reported classroom management was initially assessed using 10 items related to disruptions, transitions, monitoring, and clarity of rules (e.g., “in the lesson, our teacher is clear to us why certain rules are important”). however, we excluded two items from the study because of negative correlations between them and other classroom management items, resulting in eight items for measuring classroom management (see table 1). student-reported student support was assessed using 11 items including three items covering teacher support, four items covering autonomy support, and four items covering competence support (e.g., “our mathematics teacher makes me feel confident in my ability to learn the material”). 2.2.2 mediators: student use of learning opportunities student use of learning opportunities was assessed using three scales (oecd, 2020; vieluf et al., 2020). student self-reported depth of processing was assessed with three items (e.g., “i keep thinking about tasks until i really understand them”). student self-reported time-on-task was assessed using three items (e.g., “i pay attention in mathematics class”). student self-reported need satisfaction was assessed using three items (e.g., “i feel i can decide on things on my own”). in designing our analyses, we had to decide on whether using learning processes at either t1 or t2 as mediators. the primary goal of this study was to examine whether the longitudinal effects of teaching quality on student outcomes were mediated by learning processes. teaching quality was measured focusing on teaching in a class in general (t1) whereas the outcomes were focusing on the unit on quadratic equations (t2). in choosing the appropriate time point for the mediators, we therefore had to decide to either assess learning processes closer to the outcomes or closer to teaching quality. we opted for the latter due to the close interplay between opportunities and use. 2.2.3 outcomes: interest and achievement we chose student individual interest in mathematics classes, which was also used as an outcome in the report of talis video study and by other studies, as a motivational outcome (herbert et al., 2022; zhu & kaiser, 2022). student self-reported interest in mathematics classes was assessed at t1 using three items (e.g., “i often think that what we are talking about in my mathematics class is interesting.”). student self-reported interest in the instructional unit was assessed after the unit, at t2, using three items (e.g., “i was interested in the topic of quadratic equations”). students’ knowledge of mathematics was assessed using 30 multiple-choice items. the pre-test focused on the key prerequisites for the conceptual understanding of quadratic equations. items covered students’ precursors to understanding quadratic equations such as numbers, algebraic expressions, and alp christ, capon-sieber et al. 75 | f l r algebraic equations. the post-test (t2) focused on students’ knowledge of quadratic equations and its applications (mccaffrey et al., 2020). 2.3 data analysis we tested our hypothesized mediation paths with correlative (preliminary) and longitudinal (main) analyses using the lavaan package (v0.6-8; rosseel, 2012) in the r programming software (r development core team, 2020). the r code for all the analyses is included in the supplementary material. to assess the reliability of the aggregated student variables, intraclass correlation coefficients (icc1 and icc2) were computed for all model variables (see table 1). icc1 ranged between 4% and 37%. this range shows the extent to which the individual ratings of the variables are attributable to classroom membership (lebreton & senter, 2008). icc2 is the reliability of the class-average constructs and ranged between .45 and .93. icc2 values between .70 and .85 indicate acceptable levels of reliability (lebreton & senter, 2008; lüdtke et al., 2009). to account for the hierarchical structure of the data, the main analyses were multilevel longitudinal path analyses. due to the complexity of the tbd model, we investigated the mediating effects of three mediators for each dimension of teaching quality in separate models. in keeping with the methodology employed in other empirical studies investigating mediation between teaching quality and student achievement (e.g., león et al., 2017; ruiz alfonso & león, 2017; theis et al., 2020), we used 1-1-1 and 2-2-2 models so that between groups effects and within-group effects were separated (preacher et al., 2011). because the cluster size was too small to apply latent models and due to the complexity of the models, we averaged the items per scale and used the resulting mean scores as manifest variables in our path models. moreover, due to the non-normality of the assessed variables, we used the maximum likelihood with robust standard errors (mlr) estimator (savalei & rosseel, 2021). fifty-one students for whom all values were missing for all the assessed variables were removed from the analyses, leaving a total sample of n = 907. the percentage of missing values for the assessed scales in the total sample ranged from 0.1% to 0.7%. we did not apply a special missing value treatment because of this low percentage (kline, 2011). we used the comparative fit index (cfi), the tucker-lewis index (tli), the standardized root mean square residual (srmr), and the root mean square error of approximation (rmsea) to evaluate model fit. adequate and very good fit is achieved when the cfi and tli are greater than .90 and .95 respectively; rmsea and srmr show adequate fit when they are between .05 and .08, and they show very good fit when they are less than .05 (hu & bentler, 1999). because distributions of indirect effects could be non-normal, the bootstrapping method was used to calculate confidence intervals for the indirect effects (n = 1000 bootstrap samples; preacher & hayes, 2008). the indirect effects are considered statistically significant when the 95% confidence intervals do not include zero (cheung & lau, 2008; mackinnon et al., 2004). 3. results 3.1 descriptive statistics and bivariate correlations descriptive statistics and bivariate pearson’s correlations for all observed variables at the classroom and student level are presented in table 1. positive correlations between the independent variables (the three basic dimensions of teaching quality) and all the mediators (depth of processing, time-on-task, and need satisfaction) were found at both classroom and student levels. however, not all the expected correlations between independent variables and outcomes (interest and achievement), as well as between the mediators and outcomes were found. for example, at the student level, the three basic dimensions of teaching quality were not correlated with achievement at t1. alp christ, capon-sieber et al. 76 | f l r 3.2 preliminary analyses 3.2.1 relationships between teaching quality, learning processes, and student outcomes as a first step, we conducted multilevel path analyses using all the variables that had been assessed at the same point in time (i.e., t1). the results of the three separate direct effect models indicated that, at the student level, all the three basic dimensions of teaching quality were positively related to student interest and only student support was positively related to student achievement. at the classroom level, classroom management and student support were positively related to student interest, and classroom management was positively related to achievement. in a second step, we tested the three mediation models. the correlational mediation analyses indicated that all the three basic dimensions were positively related to all the mediators at both student and classroom levels. furthermore, positive associations were found between all the mediators and outcomes at both levels, except for time-on-task and achievement (for details, see supplementary material). 3.3 main analyses 3.3.1 longitudinal relationships between teaching quality and student outcomes we estimated three multilevel longitudinal path analyses. by conducting three separate direct effect models using longitudinal data, we tested the direct relationship between the three dimensions of teaching quality at t1 and student achievement and interest at t2 while controlling for student achievement and interest, respectively, at t1. the model fit indices are sufficient except for tli, which is slightly lower than the acceptable values for two of the three models (see table 2). the results of the three multilevel path models indicate that the three basic dimensions of teaching quality were not directly associated with mathematics achievement and interest at t2, neither at the classroom nor at the student level, controlling for student achievement and interest at t1 (see figure 3). strong positive relationships between t1 interest and t2 interest and between t1 achievement and t2 achievement were found at both the classroom and student levels. at the student level, the three dimensions of teaching quality at t1 were positively associated with student interest at t1, whereas at the classroom level, only student support at t1 was found to be positively related to student interest at t1. alp christ, capon-sieber et al. 77 | f l r table 1 descriptives, iccs, reliability estimates (ω), and within and between level intercorrelations between the measured variables 1 2 3 4 5 6 7 8 9 10 1. cognitive activation (t1) 1 .15*** .48*** .26*** .17*** .32*** .27*** -.02 .16*** -.01 2. classroom management (t1) .14*** 1 .31*** .07 .18*** .22*** .16*** -.02 .12** .00 3. student support (t1) .49*** .22*** 1 .32*** .26*** .60*** .46*** .08 .28*** .10 4. depth of processing (t1) .24*** .09 .32*** 1 .34*** .45*** .54*** .25*** .36*** .25*** 5. time-on-task (t1) .14*** .21*** .21*** .34*** 1 .26*** .33*** .07 .27*** .06 6. need satisfaction (t1) .33*** .20*** .62*** .47*** .26*** 1 .52*** .25*** .28*** .21*** 7. interest (t1) .24*** .18*** .48*** .54*** .35*** .54*** 1 .22*** .56*** .25*** 8. achievement (t1) -.03 .08 .03 .27*** .10* .22*** .18*** 1 .09 .58*** 9. interest (t2) .15*** .12** .30*** .38*** .28*** .30*** .58*** .06 1 .18*** 10. achievement (t2) -.05 .10* .08 .27*** .08 .20*** .24*** .61*** .21*** 1 meanwithin 2.62 3.01 2.98 2.66 3.10 2.86 2.40 .73 2.17 .50 sdwithin .49 .47 .56 .63 .54 .64 .78 .15 .73 .19 icc1 .10 .37 .26 .04 .04 .11 .13 .20 .12 .16 icc2 .72 .93 .89 .50 .45 .73 .77 .85 .73 .80 ωwithin .65 .60 .86 .67 .73 .63 .85 .81 ωbetween .74 .95 .98 1.00 .95 .90 .98 .99 note. *p < .05. **p < .01. ***p < .001. student-level correlations are displayed below the diagonal and classroom-level correlations are displayed above the diagonal. alp christ, capon-sieber et al. 78 | f l r figure 3. models of the effects of three basic dimensions of teaching quality at t1 on student achievement and interest at t2 controlling for student achievement and interest at t1. note. *p < .05. **p < .01. ***p < .001. alp christ, capon-sieber et al. 79 | f l r table 2 model fit indices model fit indices χ2 p cfi tli rmsea [90%-ci] srmrwithin srmrbetween direct effects models 1. cognitive activation 34.841 < .000 .958 .790 .098 [.067 .132] .037 .073 2. classroom management 21.900 < .000 .977 .885 .075 [.047 .106] .036 .055 3. student support 17.571 = .001 .985 .924 .065 [.037 .096] .030 .024 mediation models 4. cognitive activation 9.294 = .054 .997 .952 .041 [.000 .075] .013 .026 5. classroom management 7.624 = .106 .998 .967 .034 [.000 .068] .012 .022 6. student support 8.557 = .073 .997 .964 .038 [.000 .072] .012 .015 alp christ, capon-sieber et al. 80 | f l r 3.3.2 the mediating role of student use of learning opportunities after investigating the direct effect models, we constructed three separate multilevel longitudinal mediation models for each of the three dimensions of teaching quality: cognitive activation, classroom management, and student support. as with the direct effect models, strong positive relationships between interest at t1 and at t2 and between achievement at t1 and t2 were found at both the classroom and student level. all three mediators depth of processing, time-on-task, and need satisfaction were added to the direct effect models. the multilevel mediation models yielded satisfying fit indices (see table 2). we also conducted bootstrap analyses to calculate the mediating effects of depth of processing, timeon-task, and need satisfaction because the distribution of the mediating effects can be non-normal (preacher & hayes, 2008). the results of the bootstrap analyses are shown in table 3. the three multilevel mediation models generated the following results: cognitive activation: cognitive activation at t1 was positively related to need satisfaction at t1 at both the student and classroom level, and depth of processing at t1 at the student level (see figure 4). mediation analyses revealed that depth of processing at t1 mediated the relation between t1 cognitive activation and t2 achievement (β = .02; b = .01; seb = .00; 95%-ci: [.00, .01]) at the student level. classroom management: classroom management at t1 was positively related to time-on-task at t1 at both the student and the classroom level, and need satisfaction at t1 at the student level (figure 5). mediation analyses also confirmed the mediating role of t1 time-on-task between t1 classroom management and t2 interest (β = .01; b = .02; seb = .01; 95%-ci: [.00, .04]) student support: for student support, a positive relation between t1 student support and t1 need satisfaction was found both at the classroom level and at the student level (see figure 6). relations with t1 depth of processing and t1 time-on-task were also found at the student level. mediation analyses showed that t1 depth of processing mediated the relation between t1 student support and t2 achievement (β = .01; b = .00; seb = .00; 95%-ci: [.00, .01]), whereas t1 time-on-task mediated the relation between t1 student support and t2 interest (β = .01; b = .02; seb = .01; 95%-ci: [.00, .03]) at the student level. it is important to note that some of the standardized regression (beta) coefficients in the models are larger than 1.00. this is primarily due to the multicollinearity and low variance of the variables at the classroom level (i.e., iccs of depth of processing and time-on-task were .04). in summary, in our analyses, none of the mediation assumptions were supported at the classroom level, but some of them were supported at the student level. alp christ, capon-sieber et al. 81 | f l r table 3 results of bootstrap analyses mediation models classroom level student level model 4. cognitive activation à β b seb %95cis β b seb %95cis depth of processing à achievement .00 .00 .02 [-.03, .03] .02 .01 .00 [.00, .01] depth of processing à interest .00 .00 .01 [-.02, .02] .01 .02 .01 [-.00, .04] time-on-task à achievement -.10 -.04 .06 [-.15, .07] -.00 -.00 .00 [-.00, .00] time-on-task à interest .02 .02 .11 [-.19, .23] .01 .01 .01 [.00, .02] need satisfaction à achievement -.02 -.01 .08 [-.17, .16] .01 .00 .00 [.00, .01] need satisfaction à interest -.13 -.19 .33 [-.84, .45] -.01 -.01 .01 [-.03, .01] model 5. classroom management à depth of processing à achievement -.01 -.00 .01 [-.02, .02] .00 .00 .00 [-.00, .00] depth of processing à interest -.00 -.00 .02 [-.03, .03] .00 .00 .01 [-.01, .01] time-on-task à achievement .30 .07 .06 [-.04, .19] -.00 -.00 .00 [-.01, .00] time-on-task à interest .05 .04 .18 [-.32, .39] .01 .02 .01 [.00, .04] need satisfaction à achievement .00 .00 .00 [-.01, .01] .01 .00 .00 [-.00, .01] need satisfaction à interest .01 .01 .04 [-.07, .09] -.01 -.01 .01 [-.03, .01] model 6. student support à depth of processing à achievement -.00 -.00 .02 [-.04, .04] .01 .00 .00 [.00, .01] depth of processing à interest .00 .00 .01 [-.02, .02] .01 .01 .01 [-.00, .03] time-on-task à achievement .01 .00 .04 [-.08, .08] -.01 -.00 .00 [-.01, .00] time-on-task à interest .03 .03 .10 [-.18, .23] .01 .02 .01 [.00, .03] need satisfaction à achievement -1.17 -.28 .20 [-.67, .11] .02 .01 .01 [-.00, .02] need satisfaction à interest -.29 -.23 .68 [-1.56, 1.10] -.02 -.02 .02 [-.07, .02] alp christ, capon-sieber et al. | f l r 82 figure 4. mediation model for cognitive activation with standardized coefficients. note. *p < .05. **p < .01. ***p < .001. alp christ, capon-sieber et al. | f l r 83 figure 5. mediation model for classroom management with standardized coefficients. note. *p < .05. **p < .01. ***p < .001. alp christ, capon-sieber et al. | f l r 84 figure 6. mediation model for student support with standardized coefficients. note. *p < .05. **p < .01. ***p < .001. 4. discussion this study aimed to investigate the mediating role of student learning processes in the relationship between the three basic dimensions of teaching quality and student outcomes in mathematics by focusing on and extending on the hypotheses of the tbd model (klieme et al., 2009). contrary to the premise of the tbd model and our reasoning, the results of our study showed no statistically significant direct longitudinal effects of the teaching quality dimensions on student outcomes at either the classroom or individual level. positive associations were found between teaching quality dimensions and mediators, but the partial mediation models at the classroom level failed to confirm the mediation hypotheses. at the student level, consistent with the tbd model, depth of processing mediated the relationship between cognitive activation and achievement. however, contrary to the predictions of the tbd model and consistent with our new hypotheses, the following relationships were found at the student level: time-on-task mediated the relationship between classroom management and interest, and between student support and interest. depth of processing mediated the relationship between student support and achievement. these results suggest that the tbd model could benefit from an expansion of its hypotheses about mediators using relevant theoretical approaches such as evt (wigfield & eccles, 1992) and elm (petty & cacioppo, 1983). alp christ, capon-sieber et al. | f l r 85 4.1 conceptual expansion of the tbd model the varied results of this study suggest that the relationships between the variables are more complicated than we had initially predicted. in its current form, the tbd model is based on the assumption that there is a clearly defined sequence of teaching quality dimensions over their associated mediators on student outcomes. the simple structure of the model was possibly deliberate, but it has resulted in a model that struggles to reflect the full complexity of teacher-student interactions (vieluf & klieme, 2023). it is therefore important that the model is expanded and refined. based on our results, we believe that subsequent research needs to reassess three key assumptions of the current model: the first fundamental assumption of the tbd model is that the variables are related in a specific, predefined manner, with quality dimensions preceding mediators in the model’s structure. however, within the dimensions and mediators, there is no hierarchical distinction, implying equivalence among the entities within those categories. our study found that not all the hypothesized relationships we thought might exist between the mediators (depth of processing, time-on-task, and need satisfaction) and achievement and interest outcomes were supported. specifically, at the student level, we found that time-on-task predicted student interest and depth of processing predicted student achievement, but none of the other postulated relationships were observed. these findings could suggest that the mediators are not working similarly and in parallel. for instance, according to sdt (ryan & deci, 2017) need satisfaction may be a pre-condition for time-on-task and depth of processing, schlesinger and jentsch (2016) argued that time-on-task might be necessary for deep processing to occur, and according to brown and ryan (2003), conscious attention might be needed to meet psychological needs. we recommend that future research addresses the relationship between the three mediators such as considering the possibility that the mediators are sequential; one acts as a precondition for another. similarly, there might be a hierarchical sequence also for teaching quality dimensions. for example, we were not able to confirm our assumption on the relation between classroom management and depth of processing at both student and class level. one explanation could be that classroom management only indirectly influences depth of processing through cognitive activation (charalambous & praetorius, 2020; klieme et al., 2001). it acts as a pre-condition for the other teaching quality dimensions. this idea is supported by empirical evidence that classroom management predicts cognitive activation at the classroom level (dorfner et al., 2018). classroom management alone may not be enough to promote depth of processing, but it may play an enabling role. researchers should continue to explore the existence of potential mediators between classroom management and depth of processing. the second assumption of the current tbd model that our findings cast doubt on is that only three aspects of learning processes mediate the relationship between the three basic dimensions of teaching quality and student outcomes. being based on the tbd, our paper focused on an analysis of these aspects. however, given the various theoretical approaches explored in this study, other aspects such as emotions related to achievement (cvt; pekrun, 2006) and expectancies and values (evt; wigfield & eccles, 1992) could also be included in the model. the inclusion of these mediators in particular could be productive since studies supporting the relationship between the three basic dimensions of teaching quality and these learning processes already exist (e.g., burić & kim, 2020; lazarides & buchholz, 2019). the third assumption of the tbd model is that some variables are related to others in only one direction. we investigated the relationship between teaching quality at t1 and mediators at t1 and the relationship between mediators at t1 and achievement and interest at t2. we found depth of processing at t1 predicted student interest at t2. but this relationship might be bi-directional over the long term (hidi & renninger, 2006), i.e., when students are interested in mathematics classes, they tend to think more critically and try to solve more challenging problems. likewise, our study showed that interest and achievement at t1 are related to mediators at t1 but the direction of the effects could not be ascertained because they were investigated at the same point in time. this is also true for the relations alp christ, capon-sieber et al. | f l r 86 between teaching quality dimensions and mediators. therefore, future studies should consider using more suitable designs such as cross lagged models and three measurement points to separately investigate the longitudinal mediating effects of each mediator. moreover, the relationships may also be influenced by control variables and moderators, such as student personality traits, adding yet more complexity to any analysis. study designs should test and expand theoretical assumptions, using robust experimental or intervention designs. it is also vital to acknowledge the complexity of an educational reality encompassing countless interactions between teachers and students, not unfairly described as a “hall of mirrors” (berliner, 2002, cronbach, 1975). it is essential to recognize that continued exclusive reliance on quantitative methods such as mediation analysis may not capture the full complexity of the system. qualitative approaches and mixed method studies are needed to further develop the tbd model (vieluf & klieme, 2023). 4.2 correlational vs. longitudinal evidence our study highlighted that choosing whether to use correlational or longitudinal analyses can have a significant impact on the results. the direct effect models using a correlational design resulted in mostly positive direct associations. contrary to our hypotheses, direct effect models with a longitudinal design revealed that at both levels, cognitive activation, classroom management, and student support did not directly predict achievement or interest. in correlational mediation models all paths, with the exception of the relationship between time-on-task and achievement, showed positive associations at both levels. however, in longitudinal mediation models, mediating effects were found only at the student level and only few of them could be identified. correlational research design is frequently used to confirm theoretically predicted relationships between variables in educational research because it is a practical approach. most of the relationships we found using a correlational design were positive but the same was not true when the data were analyzed longitudinally. this discrepancy is important and researchers should investigate differences between correlational and longitudinal data in other settings. correlational results do not establish causality or the direction of effects. to avoid potential misconceptions, researchers should not rely only on correlational designs for research that may have practical implications for teachers. also, the interpretation of correlational findings needs careful framing. for instance, correlational studies should avoid using directional language such as “affect” or “predict” to minimize potential misinterpretations. although correlational studies can be a practical tool in the early stages of a new area of research, helping to identify any relationships, when the research field is saturated with the correlational studies, as it is in teaching quality research, we recommend the use of stronger methods such as longitudinal or experimental designs so that the directionality of effects can be established. although less used, longitudinal designs have the advantage of being able to reveal the direction of effects. they do, however, pose challenges. firstly, using short time intervals between measurement points in longitudinal studies often results in a high stability of the variables over time (begrich et al., 2023). this issue was observed in the analysis of student achievement and interest in our study. the high stability of outcome variables implies that the remaining variables, such as the dimensions of teaching quality at t1 or mediators at t1, only explain a little of the variance (adachi & willoughby, 2015; praetorius et al., 2018; warner et al., 2017). in the future, researchers could mitigate this effect by having longer intervals between measurement points. extended intervals would also enable the monitoring of significant transitions, such as a change of teacher or shifts in classroom dynamics (begrich et al., 2023). another interesting avenue for future research would be to examine whether study outcomes are affected by time between measurements. in our study there was considerable variance in intervals, from 22 to 130 days. it would be interesting to analyze the differences between classes where the interval was larger and those where it was smaller by for example, dividing data at the median time interval, to determine the effect of time intervals on the stability of outcomes. however, due to the limitations imposed by the relatively small size of our study sample and the limited number of classrooms, it was not possible to run such a complex model (hox & mcneish, 2020; maas & hox, alp christ, capon-sieber et al. | f l r 87 2005). secondly, despite providing valuable insights, longitudinal studies only assess specific time points and do not denote any causal link between variables. therefore, experiments or interventions are required to confirm the effects between the variables or determine the absence of effects in certain contexts and settings. for example, teaching quality could be manipulated by training a group of teachers to set optimally challenging tasks that cater to the level of each student. the results from this group could then be compared to a control group providing regular lessons using an experience sampling approach to investigate what effect the treatment had on their learning processes and outcomes (see schukajlow et al., 2023; talić et al., 2022). while correlational studies give an initial indication of the relationships between variables, sometimes, these relationships are not confirmed by a longitudinal study, as is the case here. longitudinal design in tbd research can be improved by having longer intervals between measurement points. however, more rigorous and holistic approaches are necessary in order to be able to show the effect of teaching quality on learning processes and then, in turn, on student outcomes. 4.3 level of analyses we ran the models at both the student and classroom levels and the results differed, depending on the level of analysis. in order to better understand how much individual student perceptions differed from the shared class perception, we separated within group and between group effects (fauth et al., 2014; marsh et al, 2012). this was also helpful for identifying the most suitable constructs for each level. for example, at the classroom level the low icc1 of depth of processing and time-on-task showed that only 4% of the variance in those variables could be attributed to classroom membership. these two variables are also problematic regarding their low icc2. although considering between-level effects for variables with low iccs is possible when intraclass correlations are nonzero and number of individuals per group is high (julian, 2001; lazarides & buchholz, 2019), these effects must be interpreted with caution. given the low iccs, especially for depth of processing and time-on-task, it appears that these constructs might be more idiosyncratic. while students in the same classroom are taught by a single teacher and may share some learning processes related to their common activities, such as solving specific mathematical problems (e.g., hill & rowe, 1996), each student-teacher interaction remains unique. this is because each student has different personality traits, beliefs, values, and a different ability level, prior knowledge, and family background (helmke, 2012; seidel, 2014), all of which will probably influence how they perceive any activity or teaching approach. this observation has important implications. although researchers have been mostly treating teaching quality as a classroom level construct, it might be more important to consider teaching at both levels, paying attention to the individual level effects. studies which mostly assessed teaching quality at classroom level could have missed the effect of individual variables. a study can be interested in relations at the classroom level, the student level, or both (senden et al., 2023; stapleton et al., 2016). while the levels of analysis in a study depend on the data and research questions (marsh et al., 2012), teaching and learning occur at both student and classroom levels and the effects at each level might be different. although studies consider teaching quality most often at the classroom level, it is important that we do not ignore the effect of student-perceived teaching quality at the individual level. researchers should consider refining operationalizations of teaching quality to include aspects such as differentiation and adaptivity (vieluf & klieme, 2023). this would allow teaching quality measures to encompass individual and unique interactions with students, thus enhancing their relevance at the classroom level. moreover, future methodological studies could explore the role of teaching quality and learning processes, particularly by using qualitative interviews, to develop more adequate measures. 4.4 implications for teaching practice the primary focus of our study was to improve the conceptual understanding of teaching and its effects on student outcomes for future research on teaching quality. the findings demonstrate that we alp christ, capon-sieber et al. | f l r 88 are a long way from fully understanding the mediating mechanisms that underlie how teaching affects student outcomes. while these factors make it more difficult to suggest implications for practice than, for example, with an intervention study, we do believe that the results are relevant to teaching practice in two ways. first, that the study found no mediation effects at the classroom level, but several at the student level, suggests that teachers might need shift their focus from the class to the student. this requires a more adaptive approach to teaching, one that is responsive to the evolving dynamics of the class and addresses not just the collective needs of the class but also the unique needs of each student (vieluf & klieme, 2023). while this is not an original recommendation – researchers have been discussing the idea, mostly at a theoretical level, for decades – our study provides supporting empirical evidence. clearly, this is a challenging remit for teachers. support could include developing formative tools to help teachers gather and interpret student perceptions of teaching and use of learning materials, and providing concrete guidance on how to incorporate this information into daily lesson plans (decristan et al., 2015; pinger et al., 2018). second, results suggest that the mechanisms through which teaching shapes learning are far more complex than teaching effectiveness researchers had hitherto hypothesized. not only has our study uncovered a more intricate array of mediation pathways within the original model than previously identified, but it also suggests that an expansion to encompass adjacent theoretical frameworks such as evt (wigfield & eccles, 1992) may reveal yet more mediators. this complexity means that there can be no standard teaching “recipes” that work for all students (see vieluf, 2022). of course teachers, especially trainees, find recipes appealing but these results suggest that teaching is too complex and constrained by context for such prescriptions. to conclude, our study highlights the importance of focusing on the individual student’s use of learning opportunities, resonating with constructivist principles that emphasize the importance of the individual’s construction of knowledge (aebli, 2011; piaget, 1992). this in turn underscores the value of an adaptive and flexible approach to teaching; one that is responsive to the evolving dynamics of the classroom (vieluf & klieme, 2023). 4.5 limitations and future directions first, contrary to our expectations, cognitive activation was not related to time-on-task at either the classroom or student level. looking at the operationalization of the constructs in more detail, it becomes evident that our study operationalized cognitive activation specifically as teachers providing tasks that require critical thinking and presenting problems with no obvious solutions, whereas timeon-task was defined and operationalized as paying attention during the mathematics lesson, but not specifically when solving complex tasks or undertaking critical thinking. when comparing the operationalization of the constructs, it appears that time-on-task was more generally operationalized than cognitive activation. in a similar vein, the variables at t2 referred to specific mathematics lessons on quadratic equations, whereas teaching quality and the mediators referred to mathematics in general. those issues with operationalization could have resulted in greater variability in the way students interpret and respond to the items, which may not have been evident when analyzing the results. some students might usually listen to instructions and pay attention during the lessons but perhaps not be particularly attentive when solving complex problems or vice versa. some students may also be more interested or more successful in some academic domains than in others (jansen et al., 2019). considering all these issues, it would be fruitful to compare the effects by using different operationalizations of the constructs, ideally within one study. second, in the talis video study space restrictions in the questionnaire meant the item for assessing some variables did not permit a detailed investigation of the subdimensions. for example, need satisfaction was assessed by three items, one item for each need: autonomy, competence, and relatedness. research in sdt has begun to focus on the negative impact of need frustration, not just the positive effect of need satisfaction. we suggest that future studies on the tbd model incorporate these theoretical developments by using more comprehensive and well-established questionnaires such as the alp christ, capon-sieber et al. | f l r 89 basic psychological needs satisfaction scales (bpnss; deci & ryan, 2000; gagné, 2003), the balanced measure of psychological needs (bpmn; sheldon & hilpert, 2012), and the basic psychological need satisfaction and frustration scale (bpnsfs; chen et al., 2015, van der kaapdeeder et al., 2020). third, the achievement test used in the talis video study focused on lowto medium-level cognitive demands such as memorization, procedures and simple applications, and did not adequately assess students’ high-level thinking such as using multiple representations and modelling authentic situations. this limitation was due in part to the difficulty of optimizing the achievement test for the diverse curricula in the countries participating in the talis video study (herbert et al., 2022). it is also important to consider that the way the test was administered during the study differed from usual classroom procedures and this may have affected the performance of some students. for example, some students may have been less attentive or more anxious during the test, which could have influenced their performance. it is therefore important to carefully consider the selection and adaptation of measures to accurately capture the constructs of interest in research and to also consider the potential impact of situational factors on performance. fourth, the sample in our study was recruited from secondary schools in germany and is highly selective as teachers decided in whether to participate in the study or not. therefore, the results cannot be generalized to the secondary schools in germany in general, and, even more so, not to other age groups (e.g., primary school or university) or other countries. the results also differed when the effects of the three basic dimensions of teaching quality on student outcomes was analyzed for the different countries which participated in the talis video study (herbert et al., 2022). therefore, cross-cultural studies should investigate mediating effects to strengthen the generalizability of our findings. the fifth limitation relates to measurement perspective, which can have an impact on study results (e.g., zee et al., 2013). this study used student ratings, which are considered valid and are commonly used in the field (appleton et al., 2008; de jong & westerhof, 2001; fredricks, 2022; lüdtke et al., 2009). student ratings of teaching quality have been found in some studies to be a better predictor of student variables than teacher and observer ratings (e.g., kunter & baumert, 2006; styck et al., 2020; wagner et al., 2016). however, observer ratings have the potential to be a more objective measure of teaching quality (clausen, 2002) and it may be that using student data for assessing teaching quality as well as student learning processes introduces a risk of common method bias, particularly in correlational analyses (podsakoff et al., 2003; 2012). although it is difficult to identify this bias empirically, future studies might consider using a marker variable which is theoretically unrelated to the other variables of the study (williams et al., 2010). self-report surveys might also be affected by what is socially desirable, leading to overor under-representation of the participant’s actual behavior (fredricks, 2022). we suggest future studies incorporate multiple perspectives to measure variables and, for simplicity, when comparing mediating effects according to rater perspectives, focus only on any specific paths within the same study (fauth et al., 2020). to summarize, neither the direct nor the indirect effect models in our study provide clear answers about the hypothesized relationships. there could be multiple reasons for these findings. our results reveal the critical importance of certain choices made while designing and analyzing a study. it seems that the conceptual sequence of the variables, the choice of correlational vs. longitudinal evidence, and the level of analysis all have an impact on the results. our finding is in line with recent reviews that have also revealed inconsistent results (see alp christ et al., 2022; praetorius et al., 2018). thus, one important take home message is that current quantitative results on the direct and indirect effects of teaching quality on student outcomes are not easy to interpret. instead of considering the entire chain at once, it might be more productive to focus exclusively on understanding the interplay between teaching quality and student learning processes better for a while (see also hiebert & stigler, 2023). alp christ, capon-sieber et al. | f l r 90 5. conclusions this study is the first to investigate relationships within the entire tbd model using a longitudinal design. it does so by enriching the tbd model with well-established cognitive and motivational theories. the multilevel mediation analyses using both correlational and longitudinal designs revealed varied results which depended on study design and level of analysis and once again highlighted the complexity of the relationships between teaching quality, student learning processes, and student outcomes (see also alp christ et al., 2022). our study contributes to the literature by supporting some of the assumptions of the tbd model and finding new paths between teaching quality and student outcomes. in line with recent appeals in the field (praetorius & charalambous, 2023; vieluf & klieme, 2023), our study advocates for augmenting the current model with supplementary theories pertaining to cognition, motivation, and effort, to advance the field. keypoints the assumptions of the tbd model are revisited and expanded using leading motivational and cognitive theories. first longitudinal investigation of the entire tbd model, integrating new possible mediating paths. multilevel mediation analyses show diverse findings for direct and indirect effects, highlighting model intricacies. conceptual and methodological choices can have a significant influence on the results. supplementary material supplementary material for this article can be found online. funding ayşenur alp christ is funded by the swiss government excellence scholarship (eskas no. 2019.0503). the talis video study germany was supported by the leibniz association. author contributions ayşenur alp christ: conceptualization, literature review, writing – original draft, writing – review & editing, data analysis, visualization, supplementary materials. vanda capon-sieber: conceptualization, literature review, writing – original draft, supervision, writing – review & editing. carmen köhler: supervision of the data analyses, writing – review & editing. eckhard klieme: design of the talis-video study, data collection and curation, writing – review & editing. anna-katharina praetorius: conceptualization, design of the study, data collection and curation, supervision, writing – review & editing. alp christ, capon-sieber et al. | f l r 91 appendix items assessed in the talis video study student questionnaire cognitive activation with discourse (t1) (1= never or almost never , 2 = occasionally, 3 = frequently, 4 = always) our mathematics teacher presents tasks for which there is no obvious solution. our mathematics teacher presents tasks that require us to apply what we have learned to new contexts. our mathematics teacher gives tasks that require us to think critically. our mathematics teacher asks us to decide on our own procedures for solving complex tasks. our mathematics teacher gives us opportunities to explain our ideas. our mathematics teacher encourages us to question and critique arguments made by other students. our mathematics teacher requires us to engage in discussions among ourselves. classroom management (t1) (1 = strongly disagree, 2 = disagree, 3 = agree, 4 = strongly agree) when the lesson begins, our mathematics teacher has to wait quite a long time for us to quieten down. we lose quite a lot of time because of students interrupting the lesson. there is much disruptive noise in this classroom. in our teacher’s class, we are aware of what is allowed and what is not allowed. in our teacher’s class, we know why certain rules are important. our teacher manages to stop disruptions quickly. our teacher reacts to disruptions in such a way that the students stop disturbing learning. in our teacher’s class, transitions from one phase of the lesson to the other (e.g., from discussions to individual work) take a lot of time. student support (t1) (1 = strongly disagree, 2 = disagree, 3 = agree, 4 = strongly agree) our mathematics teacher gives extra help when we need it. our mathematics teacher continues teaching until we understand. our mathematics teacher helps us with our learning. our mathematics teacher makes me feel confident in my ability to do well in the . our mathematics teacher listens to my view on how to do things. i feel that our mathematics teacher understands me. our mathematics teacher makes me feel confident in my ability to learn the material. our mathematics teacher provides me with different alternatives (e.g. learning materials or tasks). our mathematics teacher encourages me to find the best way to proceed by myself. our mathematics teacher lets me work on my own. our mathematics teacher appreciates it when different solutions come up for discussion. depth of processing (t1) (1 = strongly disagree, 2 = disagree, 3 = agree, 4 = strongly agree) i keep thinking about tasks until i really understand them. i think intensively about the mathematical content. i develop my own ideas regarding the topic taught. need satisfaction (t1) (1 = strongly disagree, 2 = disagree, 3 = agree, 4 = strongly agree) i feel that i can decide on things on my own. i feel understood by my mathematics teacher. i feel confident in my ability to learn this material. time-on-task (t1) (1 = strongly disagree, 2 = disagree, 3 = agree, 4 = strongly agree) i pay attention in mathematics class. i listen to the instruction given in class. i let my mind wander during the lessons. alp christ, capon-sieber et al. | f l r 92 interest (t1) (1 = strongly disagree, 2 = disagree, 3 = agree, 4 = strongly agree) i am interested in mathematics. i often think that what we are talking about in my mathematics class is interesting. after mathematics class i am often already curious about the next mathematics class. interest (t2) (1 = strongly disagree, 2 = disagree, 3 = agree, 4 = strongly agree) i was interested in the topic of quadratic equations. i often thought that what we were talking about in my mathematics class during the unit on quadratic equations was interesting. after my mathematics class on the topic of quadratic equations i was often already curious about the next mathematics class. alp christ, capon-sieber et al. | f l r 93 references adachi, p., & willoughby, t. (2015). interpreting effect sizes when controlling for stability effects in longitudinal autoregressive models: implications for psychological science. european journal of developmental psychology, 12, 116–128. https://doi.org/10.1080/17405629.2014.963549 aebli, h. (2011). zwölf grundformen des lehrens: eine allgemeine didaktik aufpsychologischer grundlage. medien und inhalte didaktischer kommunikation, der lernzyklus [twelve basic forms of teaching: a general didactics based on psychology: media and content of didactic communication, the learning cycle] (14. auflage). klett-cotta. ahmadi, a., noetel, m., parker, p., ryan, r.m., ntoumanis, n., reeve, j., beauchamp, m.,dicke, t., yeung, a., ahmadi, m., bartholomew, k., chiu, t.k.f., curran, t., erturan, g., flunger, b., frederick, c., froiland, j. m., gonzález-cutre, d., haerens, l., . . . lonsdale, c. (2023). a classification system for teachers’ motivational behaviors recommended in self-determination theory interventions. journal of educational psychology. 115(8), 1158–1176. https://doi.org/10.1037/edu0000783 ahn, i., ming chiu, m., & patrick, h. (2021). connecting teacher and student motivation: studentperceived teacher need-supportive practices and student need satisfaction. contemporary educational psychology, 64, 101950. https://doi.org/10.1016/j.cedpsych.2021.101950 alp christ, a., capon-sieber, v., grob, u.w., & praetorius, a. k. (2022). learning processes and their mediating role between teaching quality and student achievement: a systematic review. studies in educational evaluation, 75, 101209. https://doi.org/10.1016/j.stueduc.2022.101209 appleton, j.j., christenson, s.l., & furlong, m.j. (2008). student engagement with school: critical conceptual and methodological issues of the construct. psychology in the schools, 45, 369–386. https://doi.org/10.1002/pits.20303 appleton, j. j., christenson, s.l., kim, d., & reschly, a.l. (2006). measuring cognitive and psychological engagement: validation of the student engagement instrument. journal of school psychology, 44(5), 427–445. https://doi.org/10.1016/j.jsp.2006.04.002 atkinson, j. w. (1957). motivational determinants of risk-taking behavior. psychological review, 64, part 1(6), 359–372. https://doi.org/10.1037/h0043445 atkinson, r. c., & shiffrin, r. m. (1968). human memory: a proposed system and its control processes. in k.w. spence & j.t. spence (eds.), the psychology of learning and motivation: advances in research and theory. (vol. 2, pp. 89-195). academic press. badri, r., amani-saribaglou, j., ahrari, g., jahadi, n., & mahmoudi, h. (2014). school culture, basic psychological needs, intrinsic motivation and academic achievement: testing a casual model. mathematics education trends and research, 4, 1–13. https://doi.org/10.5899/2014/metr-00050 baumeister, r. f., & leary, m. r. (1995). the need to belong: desire for interpersonal attachments as a fundamental human motivation. psychological bulletin, 117(3), 497–529. https://doi.org/10.1037/0033-2909.117.3.497 baumert, j., kunter, m., blum, w., brunner, m., voss, t., jordan, a., klusmann, u., krauss, s., neubrand, m., & tsai, y.‑m. (2010). teachers’ mathematical knowledge, cognitive activation in the classroom, and student progress. american educational research journal, 47(1), 133–180. https://doi.org/10.3102/0002831209345157 begrich, l., praetorius, a. k., decristan, j., fauth, b., göllner, r., herrmann, c., kleinknecht, m., taut, s., & kunter, m. (2023). was tun? perspektiven für eine unterrichtsqualitätsforschungder zukunft. [what to do? perspectives for a teaching quality research of the future]. unterrichtswissenschaft, 51, 63–97. https://doi.org/10.1007/s42010-023-00163-4 berliner, d. c. (2002). comment: educational research: the hardest science of all. educational researcher, 31(8), 18–20. https://doi.org/10.3102/0013189x031008018 black, a. e., & deci, e. l. (2000). the effects of instructors' autonomy support and students’ autonomous motivation on learning organic chemistry: a self‐determination theory perspective. science education, 84(6), 740–756. https://doi.org/10.1002/1098237x(200011)84:6<740::aidsce4>3.0.co;2-3 https://doi.org/10.1080/17405629.2014.963549 https://doi.org/10.1037/edu0000783 https://doi.org/10.1016/j.cedpsych.2021.101950 https://doi.org/10.1016/j.stueduc.2022.101209 https://doi.org/10.1002/pits.20303 https://doi.org/10.1016/j.jsp.2006.04.002 https://doi.org/10.1037/h0043445 https://doi.org/10.5899/2014/metr-00050 https://doi.org/10.1037/0033-2909.117.3.497 https://doi.org/10.3102/0002831209345157 https://doi.org/10.1007/s42010-023-00163-4 https://doi.org/10.3102/0013189x031008018 https://doi.org/10.1002/1098237x(200011)84:6%3c740::aid-sce4%3e3.0.co;2-3 https://doi.org/10.1002/1098237x(200011)84:6%3c740::aid-sce4%3e3.0.co;2-3 alp christ, capon-sieber et al. | f l r 94 blair, c. (2002). school readiness: integrating cognition and emotion in a neurobiological conceptualization of children’s functioning at school entry. american psychologist, 57, 111–127. https://doi.org/10.1037//0003-066x.57.2.111 boggiano, a. k., main, d. s., & katz, p. a. (1988). children's preference for challenge: the role of perceived competence and control. journal of personality and social psychology, 54(1), 134– 141. https://psycnet.apa.org/doi/10.1037/0022-3514.54.1.134 boggiano, a. k., flink, c., shields, a., seelbach, a., & barrett, m. (1993). use of techniques promoting students' self-determination: effects on students' analytic problem-solving skills. motivation and emotion, 17(4), 319–336. https://doi.org/10.1007/bf00992323 boston, m. d., & candela, a. g. (2018). the instructional quality assessment as a tool for reflecting on instructional practice. zdm, 50(3), 427–444. https://doi.org/10.1007/s11858-018-0916-6 böheim, r., knogler, m., kosel, c., & seidel, t. (2020). exploring student hand-raising across two school subjects using mixed methods: an investigation of an everyday classroom behavior from a motivational perspective. learning and instruction, 65, 101250. https://doi.org/10.1016/j.learninstruc.2019.101250 brehm, j. w., & self, e. a. (1989). the intensity of motivation. annual review of psychology, 40, 109–131. https://doi.org/10.1146/annurev.ps.40.020189.000545 brophy, j. (2000). teaching. educational practices series: vol. 1. internationalacademy of education. brophy, j. (2006). history of research on classroom management. in c. m. evertson & c. s. weinstein (eds.), handbook of classroom management: research, practice and contemporary issues (pp. 17–43). lawrence erlbaum. brown, k. w., & ryan, r. m. (2003). the benefits of being present: mindfulness and its role in psychological well-being. journal of personality and social psychology, 84(4), 822–848. https://doi.org/10.1037/0022-3514.84.4.822 burić, i., & kim, l. e. (2020). teacher self-efficacy, instructional quality, and student motivational beliefs: an analysis using multilevel structural equation modeling. learning and instruction, 66, 101302. https://doi.org/10.1016/j.learninstruc.2019.101302 butler, r., & shibaz, l. (2008). achievement goals for teaching as predictors of students’ perceptions of instructional practices and students’ help seeking and cheating. learning and instruction, 18(5), 453–467. https://doi.org/10.1016/j.learninstruc.2008.06.004 chan, s., & yuen, m. (2014). personal and environmental factors affecting teachers’ creativityfostering practices in hong kong. thinking skills and creativity, 12, 69-77. https://doi.org/10.1016/j.tsc.2014.02.003 charalambous, c. y., & praetorius, a.-k. (2020). creating a forum for researching teaching and its quality more synergistically. studies in educational evaluation, 67, 100894. https://doi.org/10.1016/j.stueduc.2020.100894 chen, b., vansteenkiste, m., beyers, w., boone, l., deci, e. l., van der kaap-deeder, j., duriez, b., lens, w., matos, l., mouratidis, a., ryan, r. m., sheldon, k. m., soenens, b., van petegem, s., & verstuyf, j. (2015). basic psychological need satisfaction, need frustration, and need strength across four cultures. motivation and emotion, 39(2), 216–236. https://doi.org/10.1007/s11031-014-9450-1 cheung, g. w., & lau, r. s. (2008). testing mediation and suppression effects of latent variables. organizational research methods, 11(2), 296–325. https://doi.org/10.1177/1094428107300343 chi, m. t. h., & wylie, r. (2014). the icap framework: linking cognitive engagement to active learning outcomes. educational psychologist, 49(4), 219– 243.https://doi.org/10.1080/00461520.2014.965823 chiu, m. m. (2004). adapting teacher interventions to student needs during cooperativelearning: how to improve student problem solving and time on-task. american educational research journal, 41(2), 365–399. https://doi.org/10.3102/00028312041002365 clausen, m. (2002). unterrichtsqualität: eine frage der perspektive? [instructional quality: a question of perspectives?]. waxmann. https://doi.org/10.1037//0003-066x.57.2.111 https://psycnet.apa.org/doi/10.1037/0022-3514.54.1.134 https://doi.org/10.1007/bf00992323 https://doi.org/10.1007/s11858-018-0916-6 https://doi.org/10.1016/j.learninstruc.2019.101250 https://doi.org/10.1146/annurev.ps.40.020189.000545 https://doi.org/10.1037/0022-3514.84.4.822 https://doi.org/10.1016/j.learninstruc.2019.101302 https://doi.org/10.1016/j.learninstruc.2008.06.004 https://doi.org/10.1016/j.tsc.2014.02.003 https://doi.org/10.1016/j.stueduc.2020.100894 https://doi.org/10.1007/s11031-014-9450-1 https://doi.org/10.1177/1094428107300343 https://doi.org/10.1080/00461520.2014.965823 https://doi.org/10.3102/00028312041002365 alp christ, capon-sieber et al. | f l r 95 clifford, m. (1990). students need challenge, not easy success: only by teaching students to tolerate failure for sake of true success can educators control the national epidemic of “educational suicide”. educational leadership, 48, 22–26. cronbach, l. j. (1975). beyond the two disciplines of scientific psychology. american psychologist, 30, 116–127. https://doi.org/10.1037/h0076829 de corte, e. (1995). fostering cognitive growth: a perspective from research on mathematics learning and instruction. educational psychologist, 30(1), 37–46. https://doi.org/10.1207/s15326985ep3001_4 de corte, e. (2004). mainstreams and perspectives in research on learning (mathematics) from instruction. applied psychology, 53(2), 279–310. https://doi.org/10.1111/j.14640597.2004.00172.x de jong, r., & westerhof, k. j. (2001). the quality of student ratings of teacher behaviour. learning environments research, 4(1), 51–85. https://doi.org/10.1023/a:1011402608575 deci, e. l., eghrari, h., patrick, b. c., & leone, d. r. (1994). facilitating internalization: the selfdetermination theory perspective. journal of personality, 62(1), 119–142. https://doi.org/10.1111/j.1467-6494.1994.tb00797.x deci, e. l., & ryan, r. m. (2000). the “what” and “why” of goal pursuits: human needs and the self-determination of behavior. psychological inquiry, 11, 227–268. https://doi.org/10.1207/s15327965pli1104_01 deci, e. l., & ryan, r. m. (eds.). (2002). handbook of self-determination research. the university of rochester press. decristan, j., klieme, e., kunter, m., hochweber, j., büttner, g., fauth, b., hondrich, a.l., rieser, s., hertel, s. & hardy, i. (2015). embedded formative assessment and classroom process quality: how do they interact in promoting science understanding? american educational research journal, 52(6), 1133-1159. https://doi.org/10.3102/0002831215596412 diederich, j., & tenorth, h.‑e. (1997). theorie der schule: ein studienbuch zu geschichte, funktionen und gestaltung [theory of school: a study book on history, functions and design]. berlin, germany: cornelsen scriptor. https://doi.org/10.25656/01:11675 dorfner, t., förtsch, c., & neuhaus, b. j. (2018). effects of three basic dimensions of instructional quality on students’ situational interest in sixth-grade biology instruction. learning and instruction, 56, 42–53. https://doi.org/10.1016/j.learninstruc.2018.03.001 doyle, w. (1983). academic work. review of educational research, 53, 159–200. https://doi.org/10.3102/00346543053002159 doyle, w. (1986). classroom organization and management. in m. wittrock (ed.), handbook of research on teaching (pp. 392–431). macmillan. driscoll, m. p. (2005). psychology of learning for instruction. peterson. emmer, e. t., & stough, l. m. (2001). classroom management: a critical part of educational psychology, with implications for teacher education. educational psychologist, 36(2), 103–112. https://doi.org/10.1207/s15326985ep3602_5 evertson, c. m. (1989). improving elementary classroom management: a school-based training program for beginning the year. the journal of educational research, 83(2), 82–90. https://doi.org/10.1080/00220671.1989.10885935 evertson, c. m., & harris, a. m. (1992). what we know about managing classrooms. educational leadership, 49, 74–78. fauth, b., decristan, j., decker, a.‑t., büttner, g., hardy, i., klieme, e., & kunter, m. (2019). the effects of teacher competence on student outcomes in elementary science education: the mediating role of teaching quality. teaching and teacher education, 86, 102882. https://doi.org/10.1016/j.tate.2019.102882 fauth, b., decristan, j., rieser, s., klieme, e., & büttner, g. (2014). student ratings of teaching quality in primary school: dimensions and prediction of student outcomes. learning and instruction, 29, 1–9. https://doi.org/10.1016/j.learninstruc.2013.07.001 https://doi.org/10.1037/h0076829 https://doi.org/10.1207/s15326985ep3001_4 https://doi.org/10.1111/j.1464-0597.2004.00172.x https://doi.org/10.1111/j.1464-0597.2004.00172.x https://doi.org/10.1023/a:1011402608575 https://doi.org/10.1111/j.1467-6494.1994.tb00797.x https://doi.org/10.1207/s15327965pli1104_01 https://doi.org/10.3102/0002831215596412 https://doi.org/10.25656/01:11675 https://doi.org/10.1016/j.learninstruc.2018.03.001 https://doi.org/10.3102/00346543053002159 https://doi.org/10.1207/s15326985ep3602_5 https://doi.org/10.1080/00220671.1989.10885935 https://doi.org/10.1016/j.tate.2019.102882 https://doi.org/10.1016/j.learninstruc.2013.07.001 alp christ, capon-sieber et al. | f l r 96 fauth, b., göllner, r., lenske, l., praetorius, a., & wagner, w. (2020). who sees what? theoretical considerations on the measurement of teaching quality from different perspectives. zeitschrift für pädagogik, 66(1), 138–155. https://doi.org/10.25656/01:25870 finn, j. d., & zimmer, k. s. (eds.). (2012). student engagement: what is it? why does it matter? in s. l. christenson, a. l. reschly, & c. wylie (eds.) handbook of research on student engagement (pp 97–131). springer. https://doi.org/10.1007/978-1-4614-2018-7_5 fisher, c., berliner, d., filby, n., marliave, r., cahen, l., & dishaw, m. (1981). teaching behaviors, academic learning time, and student achievement: an overview. the journal of classroom interaction, 17(1), 2–15. förtsch, c., werner, s., dorfner, t., kotzebue, l. von, & neuhaus, b. j. (2017). effects of cognitive activation in biology lessons on students’ situational interest and achievement. research in science education, 47(3), 559–578. https://doi.org/10.1007/s11165-016-9517-y förtsch, c., werner, s., kotzebue, l. von, & neuhaus, b. j. (2018). effects of high-complexity and high-cognitive-level instructional tasks in biology lessons on students’ factual and conceptual knowledge. research in science & technological education, 36(3), 353–374. https://doi.org/10.1080/ 02635143.2017.1394286 fredricks, j. a., & mccolskey, w. (2012). the measurement of student engagement: a comparative analysis of various methods and student self-report instruments. in s. l. christenson, a. l. reschly, & c. wylie (eds.), handbook of research on student engagement (pp. 763–782). springer us. https://doi.org/10.1007/978-1-4614-2018-7_37 fredricks, j. a. (2022). the measurement of student engagement: methodological advances and comparison of new self-report instruments. in a. reschly & s. christenson (eds.), handbook of research on student engagement (pp. 597–616). springer international publishing. https://doi.org/10.1007/ 978-3-031-07853-8 gagné, m. (2003). the role of autonomy support and autonomy orientation in prosocial behavior engagement. motivation and emotion, 27, 199–223. https://doi.org/10.1023/a:1025007614869 geldhof, g. j., preacher, k. j., & zyphur, m. j. (2014). reliability estimation in a multilevel confirmatory factor analysis framework. psychological methods, 19(1), 72–91. https://doi.org/10.1037/a0032138 good, t. l., & brophy, j. e. (2003). looking in classrooms (9th ed.). pearson education. grabinger, r. s., & dunlap, j. c. (1995). rich environments for active learning: a definition. research in learning technology , 3(2), 5–34. https://doi.org/10.3402/rlt.v3i2.9606 grabinger, s., dunlap, j. c., & duffield, j. a. (1997). rich environments for active learning in action: problem-based learning. research in learning technology, 5(2), 5–17. https://doi.org/10.3402/rlt.v5i2.10558 guay, f., ratelle, c. f., & chanal, j. (2008). optimal learning in optimal contexts: the role of selfdetermination in education. canadian psychology/psychologie canadienne, 49(3), 233–240. https://doi.org/10.1037/a0012758 hachfeld, a., & lazarides, r. (2020). the relation between teacher self-reported individualization and student-perceived teaching quality in linguistically heterogeneous classes: an exploratory study. european journal of psychology of education, 36, 1159–1179. https://doi.org/10.1007/s10212-020-00501-5 hattie, j. (2009). visible learning: a synthesis of over 800 meta-analyses relating to achievement. routledge taylor & francis group. heckhausen, h. (1991). motivation and action (p. k. leppman, trans). springer-verlag. helm, c. (2016). zentrale qualitätsdimensionen von unterricht und ihre effekte auf schüleroutcomes im fach rechnungswesen. [key quality dimensions of teaching and their effects on student outcomes in accounting] zeitschrift für bildungsforschung, 6(2), 101–119. https://doi.org/10.1007/s35834-016-0154-3 helmke, a. (2012). unterrichtsqualität und lehrerprofessionalität: diagnose, evaluation und verbesserung des unterrichts [teaching quality and teacher professionalism: diagnosis, evaluation, and improvement of teaching] (4. überarbeitete aufl.). klett/kallmeyer. https://books.google.ch/books?id=mhdxowaacaaj https://doi.org/10.25656/01:25870 https://doi.org/10.1007/978-1-4614-2018-7_5 https://doi.org/10.1007/s11165-016-9517-y https://doi.org/10.1007/978-1-4614-2018-7_37 https://doi.org/10.1023/a:1025007614869 https://doi.org/10.1037/a0032138 https://doi.org/10.3402/rlt.v3i2.9606 https://doi.org/10.3402/rlt.v5i2.10558 https://doi.org/10.1037/a0012758 https://doi.org/10.1007/s10212-020-00501-5 https://doi.org/10.1007/s35834-016-0154-3 https://books.google.ch/books?id=mhdxowaacaaj alp christ, capon-sieber et al. | f l r 97 herbert, b., fischer, j., & klieme, e. (2022). how valid are student perceptions of teaching quality across education systems? learning and instruction, 82, 101652. https://doi.org/10.1016/j.learninstruc .2022.101652 hidi, s., & renninger, k. a. (2006). the four-phase model of interest development. educational psychologist, 41(2), 111–127. https://doi.org/10.1207/s15326985ep4102_4 hiebert, j., & grouws, d. a. (2007). the effects of classroom mathematics teaching on students’ learning. in f. k. lester (ed.), second handbook of research on mathematics teaching and learning: a project of the national council of teachers of mathematics (pp. 371–404). information age pub. hiebert, j., & stigler, j. w. (2023). creating practical theories of teaching. in a.-k. praetorius & c. y. charalambous (eds.), theorizing teaching: current status and open issues (pp. 23–56). springer. hill, p., & rowe, k. (1996). multilevel modeling in school effectiveness research. school effectiveness and school improvement, 7, 1–34. https://www.tandfonline.com/doi/pdf/10.1080/0924345960070101 hospel, v., & galand, b. (2016). are both classroom autonomy support and structure equally important for students’ engagement? a multilevel analysis. learning and instruction, 41, 1–10. https://doi.org/10.1016/j.learninstruc.2015.09.001 hox, j., & mcneish, d. (2020). small samples in multilevel modeling. in r. van de schoot, & m. miočević (eds.), small sample size solutions: a guide for applied researchers and practitioners (pp. 215–225). routledge. https://doi.org/10.4324/9780429273872-18 . hu, l., & bentler, p. m. (1999). cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives. structural equation modeling: a multidisciplinary journal, 6(1), 1– 55. https://doi.org/10.1080/10705519909540118 hughes, j. n., luo, w., kwok, o. m., & loyd, l. k. (2008). teacher-student support, effortful engagement, and achievement: a 3-year longitudinal study. journal of educational psychology, 100(1), 1–14. https://doi.org/10.1037/0022-0663.100.1.1 illies, j. j., & reiter‐palmon, r. (2004). the effects of type and level of personal involvement on information search and problem solving. journal of applied social psychology, 34(8), 1709– 1729. https://doi.org/10.1111/j.1559-1816.2004.tb02794.x jang, h., kim, e. j., & reeve, j. (2012). longitudinal test of self-determination theory’s motivation mediation model in a naturally occurring classroom context. journal of educational psychology, 104(4), 1175–1188. https://doi.org/10.1037/a0028089 jansen, m., schroeders, u., lüdtke, o., & marsh, h. w. (2019). the dimensional structure of students’ self-concept and interest in science depends on course composition. learning and instruction, 60, 20–28. https://doi.org/10.1016/j.learninstruc.2018.11.001 julian, m. w. (2001). the consequences of ignoring multilevel data structures in nonhierarchical covariance modeling. structural equation modeling: a multidisciplinary journal, 8(3), 325– 352. https://doi.org/10.1207/s15328007sem0803_1 kiemer, k., gröschner, a., kunter, m., & seidel, t. (2018). instructional and motivational classroom discourse and their relationship with teacher autonomy and competence support—findings from teacher professional development. european journal of psychology of education, 33(2), 377– 402. https://doi.org/10.1007/s10212-016-0324-7 klieme, e., lipowsky, f., rakoczy, k., & ratzka, n. (2006). qualitätsdimensionen und wirksamkeit von mathematikunterricht: theoretische grundlagen und ausgewählte ergebnisse des projekts „pythagoras“.[quality dimensions and effectiveness of mathematics education: theoretical foundations and selected results of the “pythagoras” project.] in m. prenzel & l. allolio-näcke (eds.), untersuchungen zur bildungsqualität von schule: abschlussbericht des dfgschwerpunktprogramms (pp. 127–146). waxmann. klieme, e., pauli, c., & reusser, k. (2009). the pythagoras study: investigating effects of teaching and learning in swiss and german mathematics classrooms. in t. janík & t. seidel (eds.), the power of video studies in investigating teaching and learning in the classroom (pp. 137–160). waxmann. https://doi.org/10.1207/s15326985ep4102_4 https://www.tandfonline.com/doi/pdf/10.1080/0924345960070101 https://doi.org/10.1016/j.learninstruc.2015.09.001 https://doi.org/10.4324/9780429273872-18 https://doi.org/10.1080/10705519909540118 https://doi.org/10.1037/0022-0663.100.1.1 https://doi.org/10.1111/j.1559-1816.2004.tb02794.x https://doi.org/10.1037/a0028089 https://doi.org/10.1016/j.learninstruc.2018.11.001 https://doi.org/10.1207/s15328007sem0803_1 https://doi.org/10.1007/s10212-016-0324-7 alp christ, capon-sieber et al. | f l r 98 klieme, e., & rakoczy, k. (2003). unterrichtsqualität aus schülerperspektive: kulturspezifische profile, regionale unterschiede und zusammenhänge mit effekten von unterricht [instructional quality from a student perspective: culture-specific profiles, regional differences, and associations with effects of instruction]. in j. baumert, c. artelt, e. klieme, m. neubrand, m. prenzel, u. schiefele, w. schneider, k.-j. tillmann, & m. weiß (eds.), pisa 2000 — ein differenzierter blick auf die länder der bundesrepublik deutschland (pp. 333–359). leske + budrich; vs verlag für sozialwissenschaften. klieme, e., & rakoczy, k. (2008). empirische unterrichtsforschung und fachdidaktik: outcomeorientierte messung und prozessqualität des unterrichts. [empirical classroom research and subject didactics: outcome-oriented measurement and process quality of teaching. zeitschrift für pädagogik, 54(2), 222–237. https://doi.org/10.25656/01:4348 klieme, e., schümer, g., & knoll, s. (2001). mathematikunterricht in der sekundarstufe i: «aufgabenkultur» und unterrichtsgestaltung [mathematics teaching in lower secondary schools: “task culture” and lesson design]. in e. klieme, & j. baumert (eds.), timss – impulse für schule und unterricht (pp. 43-57). bundesministerium für bildung und forschung. kline, r. b. (2015). principles and practice of structural equation modeling (3rd ed.). guilford press kounin, j. s. (1970a). discipline and group management in classrooms. holt rinehart &winston. kounin, j. s. (1970b). observing and delineating technique of managing behavior in classrooms. journal of research and development in education, 4(1), 62–67. kuger, s. (2016). curriculum and learning time in international school achievement studies. in s. kuger, e. klieme, n. jude, & d. kaplan (eds.), assessing contexts of learning (pp. 395–422). springer international publishing. https://doi.org/10.1007/978-3-319-45357-6_16 kunter, m., & baumert, j. (2006). who is the expert? construct and criteria validity of student and teacher ratings of instruction. learning environment research 9, 231–251. https://doi.org/10.1007/s10984-006-9015-7 kunter, m., baumert, j., & köller, o. (2007). effective classroom management and the development of subject-related interest. learning and instruction, 17(5), 494–509. https://doi.org/10.1016/j.learninstruc.2007.09.002 kunter, m., & trautwein, u. (2013). psychologie des unterrichts [psychology of teaching]. standardwissen lehramt: vol. 3895. ferdinand schöningh. https://doi.org/10.36198/9783838538952 lazarides, r., & buchholz, j. (2019). student-perceived teaching quality: how is it related to different achievement emotions in mathematics classrooms? learning and instruction, 61, 45– 59. https://doi.org/10.1016/j.learninstruc.2019.01.001 lebreton, j. m., & senter, j. l. (2008). answers to 20 questions about interrater reliability and interrater agreement. organizational research methods, 11(4), 815–852. https://doi.org/10.1177/1094428106296642 león, j., medina-garrido, e., & núñez, j. l. (2017). teaching quality in math class: the development of a scale and the analysis of its relationship with engagement and achievement. frontiers in psychology, 8, 895. https://doi.org/10.3389/fpsyg.2017.00895 li, h., liu, j., zhang, d., & liu, h. (2020). examining the relationships between cognitive activation, self-efficacy, socioeconomic status, and achievement in mathematics: a multi-level analysis. the british journal of educational psychology, 91, 101–126. https://doi.org/10.1111/bjep.12351 lietaert, s., roorda, d., laevers, f., verschueren, k., & de fraine, b. (2015). the gender gap in student engagement: the role of teachers’ autonomy support, structure, and involvement. british journal of educational psychology, 85(4), 498–518. https://doi.org/10.1111/bjep.12095 lipowsky, f., & hess, m. (2019). warum es manchmal hilfreich sein kann, das lernen schwerer zu machen – kognitive aktivierung und die kraft des vergleichens [why it can sometimes be helpful to make learning harder – cognitive activation and the power of comparison]. in k. schöppe & f. schulz (hrsg.), kreativität & bildung –nachhaltiges lernen (s. 77–132). kopaed. lipowsky, f., rakoczy, k., pauli, c., drollinger-vetter, b., klieme, e., & reusser, k. (2009). quality of geometry instruction and its short-term impact on students’ understanding of the https://doi.org/10.25656/01:4348 https://doi.org/10.1007/978-3-319-45357-6_16 https://doi.org/10.1007/s10984-006-9015-7 https://doi.org/10.1016/j.learninstruc.2007.09.002 https://doi.org/10.36198/9783838538952 https://doi.org/10.1016/j.learninstruc.2019.01.001 https://doi.org/10.1177/1094428106296642 https://doi.org/10.3389/fpsyg.2017.00895 https://doi.org/10.1111/bjep.12351 https://doi.org/10.1111/bjep.12095 alp christ, capon-sieber et al. | f l r 99 pythagorean theorem. learning and instruction, 19(6), 527–537. https://doi.org/10.1016/j.learninstruc.2008.11.001 lotz, m. (2016). grundlagen des unterrichtens und der unterrichtsforschung [fundamentals of teaching and classroom research]. in m. lotz (ed.), kognitive aktivierung im leseunterricht der grundschule: eine videostudie zur gestaltung und qualität von leseübungen im ersten schuljahr (pp. 7–22). springer fachmedien wiesbaden. https://doi.org/10.1007/978-3-65810436-8_2 lüdtke, o., robitzsch, a., trautwein, u., & kunter, m. (2009). assessing the impact of learning environments: how to use student ratings of classroom or school characteristics in multilevel modelling. contemporary educational psychology, 34(2), 120–131. https://doi.org/10.1016/j.cedpsych.2008.12.001 maas, c. j. m., & hox, j. j. (2005). sufficient sample sizes for multilevel modeling. methodology: european journal of research methods for the behavioral and social sciences, 1(3), 86–92. https://doi.org/10.1027/1614-2241.1.3.86 mackinnon, d. p., lockwood, c. m., & williams, j. (2004). confidence limits for the indirect effect: distribution of the product and resampling methods. multivariate behavioral research, 39(1), 99–128. https://doi.org/10.1207/s15327906mbr3901_4 marsh, h. w., lüdtke, o., nagengast, b., trautwein, u., morin, a. j. s., abduljabbar, a. s., et al. (2012). classroom climate and contextual effects: conceptual and methodological issues in the evaluation of group-level effects. educational psychologist, 47, 106–124. http://dx.doi.org/10.1080/00461520.2012.670488 mccaffrey, d. f., castellano k. e., van essen, t. (2020). student test development. in global teaching insights: technical report. section ii: instrument development. oecd. mcdonald, r. p. (1999). test theory: a unified treatment. lawrence erlbaum. mitchell, m. (1993). situational interest: its multifaceted structure in the secondary school mathematics classroom. journal of educational psychology, 85, 424–436. https://doi.org/10.1037/0022-0663.85.3.424 mouratidis, a., barkoukis, v., & tsorbatzoudis, c. (2015). the relation between balanced need satisfaction and adolescents’ motivation in physical education. european physical education review, 21(4), 421–431. https://doi.org/10.1177/1356336x15577222 mouratidis, a., vansteenkiste, m., michou, a., & lens, w. (2013). perceived structure and achievement goals as predictors of students’ self-regulated learning and affect and the mediating role of competence need satisfaction. learning and individual differences, 23, 179–186. https://doi.org/10.1016/j.lindif.2012.09.001 murray, c., & pianta, r. c. (2007). the importance of teacher-student relationships for adolescents with high incidence disabilities. theory into practice, 46(2), 105–112. https://doi.org/10.1080/00405840701232943 oecd. (2020). global teaching insights: a video study of teaching. oecd. https://doi.org/ 10.1787/20d6f36b-en pauli, c., & reusser, k. (2006). von international vergleichenden video surveys zur videobasierten unterrichtsforschung und -entwicklung. [from international comparative video surveys to videobased teaching research and development] zeitschrift für pädagogik, 52, 774 – 798. https://doi.org/10.25656/01:4488 pekrun, r. (2006). the control-value theory of achievement emotions: assumptions, corollaries, and implications for educational research and practice. educational psychology review, 18, 315– 341. https://doi.org/10.1007/s10648-006-9029-9 petty, r. e., & cacioppo, j. t. (1986): the elaboration likelihood model of persuasion. in l. berkowitz (eds.), 19, advances in experimental social psychology (pp. 123 – 205). new york: academic press. https://doi.org/10.1016/s0065-2601(08)60214-2 petty, r. e., cacioppo, j. t., & schumann, d. (1983). central and peripheral routes to advertising effectiveness: the moderating role of involvement. journal of consumer research, 10(2), 135– 146. https://doi.org/10.1086/208954 piaget, j. (1992). psychologie der intelligenz [psychology of the intelligence]. klett-cotta. https://doi.org/10.1016/j.learninstruc.2008.11.001 https://doi.org/10.1007/978-3-658-10436-8_2 https://doi.org/10.1007/978-3-658-10436-8_2 https://doi.org/10.1016/j.cedpsych.2008.12.001 https://doi.org/10.1027/1614-2241.1.3.86 https://doi.org/10.1207/s15327906mbr3901_4 http://dx.doi.org/10.1080/00461520.2012.670488 https://doi.org/10.1037/0022-0663.85.3.424 https://doi.org/10.1177/1356336x15577222 https://doi.org/10.1016/j.lindif.2012.09.001 https://doi.org/10.1080/00405840701232943 https://doi.org/10.25656/01:4488 https://doi.org/10.1007/s10648-006-9029-9 https://doi.org/10.1016/s0065-2601(08)60214-2 https://doi.org/10.1086/208954 alp christ, capon-sieber et al. | f l r 100 pianta, r. c., & steinberg, m. (1992). teacher–child relationships and the process of adjusting to school. new directions for child and adolescent development, 57, 61–80. https://doi.org/10.1002/cd.23219925706 pinger, p., rakoczy, k., besser, m., & klieme, e. (2018). interplay of formative assessment and instructional quality—interactive effects on students’ mathematics achievement. learning environments research, 21, 61-79. https://doi.org/10.25656/01:17407 podsakoff, p. m., mackenzie, s. b., lee, j.-y., & podsakoff, n. p. (2003). common method biases in behavioral research: a critical review of the literature and recommended remedies. journal of applied psychology, 88(5), 879-903. https://doi.org/10.1037/0021-9010.88.5.879 podsakoff, p. m., mackenzie, s. b., & podsakoff, n. p. (2012). sources of method bias in social science research and recommendations on how to control it. annual review of psychology, 63, 539–569. https://doi.org/10.1146/annurev-psych-120710-100452 praetorius, a.-k., & charalambous, c. y. (2023). where are we on theorizing teaching? a literature overview. in a.-k. praetorius & c. y. charalambous (eds.), theorizing teaching: current status and open issues (pp. 1–22). springer. https://doi.org/10.1007/978-3-031-25613-4_1 praetorius, a.‑k., klieme, e., herbert, b., & pinger, p. (2018). generic dimensions of teaching quality: the german framework of three basic dimensions. zdm, 50(3), 407–426. https://doi.org/10.1007/s11858-018-0918-4 praetorius, a.-k., fischer, j., & klieme, e. (2020b). teacher and student questionnaire development. in global teaching insights technical report. section ii: instrument development. paris: oecd publishing. praetorius, a. k., grünkorn, j., & klieme, e. (2020a). towards developing a theory of generic teaching quality: origin, current status, and necessary next steps regarding the three basic dimensions model. zeitschrift für pädagogik. beiheft, 66(1), 15–36. https://doi.org/10.3262/zpb2001015 praetorius, a.‑k., pauli, c., reusser, k., rakoczy, k., & klieme, e. (2014). one lesson is all you need? stability of instructional quality across lessons. learning and instruction, 31, 2–12. https://doi.org/10.1016/j.learninstruc.2013.12.002 preacher, k. j., & hayes, a. f. (2008). asymptotic and resampling strategies for assessing and comparing indirect effects in multiple mediator models. behavior research methods, 40(3), 879–891. https://doi.org/10.3758/brm.40.3.879 preacher, k. j., zhang, z., & zyphur, m. j. (2011). alternative methods for assessingmediation in multilevel data: the advantages of multilevel sem. structural equation modeling, 18, 161–182. https://doi.org/10.1080/10705511.2011.557329 r development core team. (2020). r: a language and environment for statistical computing [computer software]. http://www.r-project.org/ rakoczy, k. (2006). motivationsunterstützung im mathematikunterricht. zur bedeutung von unterrichtsmerkmalen für die wahrnehmung von schülerinnen und schüler. [motivational support in mathematics education. on the importance of instructional features for the perception of students]. zeitschrift für pädagogik, 52(6), 822–843. https://doi.org/10.25656/01:4490 rakoczy, k., & pauli, c. (2006). hoch inferentes rating. beurteilung der qualitätunterrichtlicher prozesse [high-inference rating. assessment of instructional quality]. in i. hugener, c. pauli, & k. reusser (eds.), video-analysen. dokumentation der erhebungsund auswertungsinstrumente zurschweizerisch-deutschen videostudie "unterrichtsqualität, lernverhaltenund mathematisches verständnis". materialien zur bildungsforschung. gfpf, 15, 206–233. reeve, j. (2006). teachers as facilitators: what autonomy-supportive teachers do and why their students benefit. the elementary school journal, 106(3), 225–236. http://dx.doi.org/10.1086/501484 reeve, j., & jang, h. (2006). what teachers say and do to support students' autonomy during a learning activity. journal of educational psychology, 98(1), 209–218. https://doi.org/10.1037/0022-0663.98.1.209 https://doi.org/10.1002/cd.23219925706 https://doi.org/10.25656/01:17407 https://doi.org/10.1037/0021-9010.88.5.879 https://doi.org/10.1146/annurev-psych-120710-100452 https://doi.org/10.1007/978-3-031-25613-4_1 https://doi.org/10.1007/s11858-018-0918-4 https://doi.org/10.3262/zpb2001015 https://doi.org/10.1016/j.learninstruc.2013.12.002 https://doi.org/10.3758/brm.40.3.879 https://doi.org/10.1080/10705511.2011.557329 http://www.r-project.org/ https://doi.org/10.25656/01:4490 http://dx.doi.org/10.1086/501484 https://doi.org/10.1037/0022-0663.98.1.209 alp christ, capon-sieber et al. | f l r 101 reeve, j. (2009). why teachers adopt a controlling motivating style toward students and how they can become more autonomy supportive. educational psychologist, 44(3), 159–175. https://doi.org/10.1080/00461520903028990 reeve, j. (2016). autonomy-supportive teaching: what it is, how to do it. in w. c. liu, j. c. k. wang, and r. m. ryan (eds.), building autonomous learners: perspectives from research and practice using self-determination theory (pp. 129–152). singapore: springer singapore. renninger, k. a., & hidi, s. (2002). student interest and achievement: developmental issues raised by a case study. in a. wigfield & j. s. eccles (eds.), development of achievement motivation (pp. 173–195). academic press. https://doi.org/10.1016/b978-012750053-9/50009-7 reusser, k., pauli, c., & waldis, m. (eds.). (2010). unterrichtsgestaltung und unterrichtsqualität: ergebnisse einer internationalen und schweizerischen videostudie zum mathematikunterricht. [lesson design and teaching quality: results of an international and swiss video study on mathematics teaching]. waxmann. richter, m., gendolla, g., & wright, r. a. (2016). three decades of research on motivational intensity theory. in advances in motivation science (vol. 3, pp. 149–186). https://doi.org/10.1016/bs.adms.2016.02.001 rieser, s., naumann, a., decristan, j., fauth, b., klieme, e., & büttner, g. (2016). the connection between teaching and learning: linking teaching quality and metacognitive strategy use in primary school. british journal of educational psychology, 86(4), 526–545. https://doi.org/10.1111/bjep.12121 rosseel, y. (2012). lavaan: an r package for structural equation modeling. journal of statistical software, 48 (2), 1–36. https://doi.org/10.18637/jss.v048.i02 ruiz-alfonso, z., & león, j. (2017). passion for math: relationships between teachers’ emphasis on class contents usefulness, motivation and grades. contemporary educational psychology, 51, 284–292. https://doi.org/10.1016/j.cedpsych.2017.08.010 ruiz-alfonso, z., & león, j. (2019). teaching quality: relationships between passion, deep strategy to learn, and epistemic curiosity. school effectiveness and school improvement, 30(2), 212–230. https://doi.org/10.1080/09243453.2018.1562944 ryan, r. m., & deci, e. l. (2000). self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. american psychologist, 55(1), 68–78. https://doi.org/ 10.1037//0003-066x.55.1.68 ryan, r. m., & deci, e. l. (2002). overview of self-determination theory: an organismic dialectical perspective. in e. l. deci & r. m. ryan (eds.), handbook of self-determination research (pp. 3–33). university of rochester press. ryan, r. m., & deci, e. l. (2009). promoting self-determined school engagement: motivation, learning, and well-being. in k. r. wentzel & a. wigfield (eds.), educational psychology handbook series. handbook of motivation at school (pp. 171–196). routledge. ryan, r. m., & deci, e. l. (2017). self-determination theory: basic psychological needs in motivation, development, and wellness. guilford publications. ryan, r. m., & powelson, c. l. (1991). autonomy and relatedness as fundamental to motivation and education. the journal of experimental education, 60(1), 49–66. https://doi.org/ 10.1080/00220973.1991.10806579 savalei, v., & rosseel, y. (2021). computational options for standard errors and test statistics with incomplete normal and nonnormal data in sem. structural equation modeling: a multidisciplinary journal, 29(2), 163–181. https://doi.org/10.31234/osf.io/wmuqj schlesinger, l., & jentsch, a. (2016). theoretical and methodological challenges in measuring instructional quality in mathematics education using classroom observations. zdm, 48, 29–40. https://doi.org/ 10.1007/s11858-016-0765-0 schlesinger, l., jentsch, a., kaiser, g., könig, j., & blömeke, s. (2018). subject-specific characteristics of instructional quality in mathematics education. zdm, 50, 475–490. https://doi.org/10.1007/s11858-018-0917-5 schukajlow, s., achmetli, k., & rakoczy, k. (2019). does constructing multiple solutions for realworld problems affect self-efficacy? educational studies in mathematics, 100(1), 43–60. https://doi.org/10.1080/00461520903028990 https://doi.org/10.1016/b978-012750053-9/50009-7 https://doi.org/10.1016/bs.adms.2016.02.001 https://doi.org/10.1111/bjep.12121 https://doi.org/10.18637/jss.v048.i02 https://doi.org/10.1016/j.cedpsych.2017.08.010 https://doi.org/10.1080/09243453.2018.1562944 https://doi.org/10.31234/osf.io/wmuqj https://doi.org/10.1007/s11858-018-0917-5 alp christ, capon-sieber et al. | f l r 102 https://doi.org/10.1007/s10649-018-9847-y schukajlow, s., & krug, a. (2014). do multiple solutions matter? prompting multiple solutions, interest, competence, and autonomy. journal for research in mathematics education, 45, 497– 533. https://doi.org/10.5951/jresematheduc.45.4.0497 schukajlow, s., rakoczy, k., & pekrun, r. (2023). emotions and motivation in mathematics education: where we are today and where we need to go. zdm–mathematics education, 55, 249–267. https://doi.org/10.1007/s11858-022-01463-2 seidel, t. (2014). angebots-nutzungs-modelle in der unterrichtspsychologie: integration von strukturund prozessparadigma [the opportunity-use model in instructional psychology: integrating structure and process paradigms]. zeitschrift für pädagogik, 60(6), 850–866. https://doi.org/10.25656/01:14686 seidel, t., & shavelson, r. j. (2007). teaching effectiveness research in the past decade: the role of theory and research design in disentangling meta-analysis results. review of educational research, 77(4), 454–499. https://doi.org/10.3102/0034654307310317 senden, b., nilsen, t., & teig, n. (2023). the validity of student ratings of teaching quality: factorial structure, comparability, and the relation to achievement. studies in educational evaluation, 78, 101274. https://doi.org/10.1016/j.stueduc.2023.101274 sheldon, k. m., & hilpert, j. c. (2012). the balanced measure of psychological needs (bmpn) scale: an alternative domain general measure of need satisfaction. motivation and emotion, 36, 439– 451. https://doi.org/10.1007/s11031-012-9279-4 silver, e. a., & stein, m. k. (1996). the quasar project: the "revolution of the possible" in mathematics instructional reform in urban middle schools. urban education, 30(4), 476–521. https://doi.org/10.1177/0042085996030004006 stallings, j. (1980). allocated academic learning time revisited, or beyond time on task. educational researcher, 9(11), 11–16. https://doi.org/10.3102/0013189x009011011 stapleton, l. m., yang, j. s., & hancock, g. r. (2016). construct meaning in multilevel settings. journal of educational and behavioral statistics, 41(5), 481–520. https://doi.org/10.3102/1076998616646200 stein, m. k., grover, b. w., & henningsen, m. (1996). building student capacity for mathematical thinking and reasoning: an analysis of mathematical tasks used in reform classrooms. american educational research journal, 33(2), 455-488. https://doi.org/10.2307/1163292 stein, m. k., & lane, s. (1996). instructional tasks and the development of student capacity to think and reason: an analysis of the relationship between teaching and learning in a reform mathematics project. educational research and evaluation, 2(1), 50–80. https://doi.org/10.1080/ 1380361960020103 styck, k. m., anthony, c. j., sandilos, l. e., & diperna, j. c. (2020). examining rater effects on the classroom assessment scoring system. child development, 92(3), 976–993. https://doi.org/10.1111 /cdev.13460 sun, h., & chen, a. (2010). a pedagogical understanding of the self-determination theory in physical education. quest, 62(4), 364–384. https://doi.org/10.1080/00336297.2010.10483655 sun, y., liu, r.‑d., oei, t.‑p., zhen, r., ding, y., & jiang, r. (2020). perceived parental warmth and adolescents' math engagement in china: the mediating roles of need satisfaction and math selfefficacy. learning and individual differences, 78, 101837. https://doi.org/10.1016/ j.lindif.2020.101837 talić, i., scherer, r., marsh, h. w., greiff, s., möller, j., & niepel, c. (2022). uncovering everyday dynamics in students’ perceptions of instructional quality with sampling. learning and instruction, 81, 101594. https://doi.org/10.1016/j.learninstruc.2022.101594 theis, d., sauerwein, m., & fischer, n. (2020). perceived quality of instruction: the relationship among indicators of students’ basic needs, mastery goals, and academic achievement. british journal of educational psychology, 90, 176–192. https://doi.org/10.1111/bjep.12313 https://doi.org/10.1007/s10649-018-9847-y https://doi.org/10.5951/jresematheduc.45.4.0497 https://doi.org/10.1007/s11858-022-01463-2 https://doi.org/10.25656/01:14686 https://doi.org/10.3102/0034654307310317 https://doi.org/10.1016/j.stueduc.2023.101274 https://doi.org/10.1007/s11031-012-9279-4 https://doi.org/10.1177/0042085996030004006 https://doi.org/10.3102/0013189x009011011 https://doi.org/10.3102/1076998616646200 https://doi.org/10.2307/1163292 https://doi.org/10.1080/00336297.2010.10483655 https://doi.org/10.1016/j.learninstruc.2022.101594 https://doi.org/10.1111/bjep.12313 alp christ, capon-sieber et al. | f l r 103 turner, j. c., & meyer, d. k. (2004). a classroom perspective on the principle of moderate challenge in mathematics. the journal of educational research, 97(6), 311–318. https://psycnet.apa.org/doi/10.3200/joer.97.6.311-318 van der kaap-deeder, j., soenens, b., ryan, r. m., & vansteenkiste, m. (2020). manual of the basic psychological need satisfaction and frustration scale (bpnsfs). ghent university, belgium. van der scheer, e. a., bijlsma, h. j. e., & glas, c. a. w. (2019). validity and reliability ofstudent perceptions of teaching quality in primary education. school effectiveness and school improvement, 30(1), 30–50. https://doi.org/10.1080/09243453.2018.1539015 vansteenkiste, m., niemiec, c. p., & soenens, b. (2010). the development of the five mini-theories of self-determination theory: an historical overview, emerging trends, and future directions. in t. c. urdan & s. a. karabenick (eds.), advances in motivation and achievement: vol. 16. the decade ahead: theoretical perspectives on motivation and achievement (1st ed., pp. 105–165). emerald group publishing limited. https://doi.org/10.1108/s0749-7423(2010)000016a007 vieluf, s., praetorius, a.-k., rakoczy, k., kleinknecht, m., & pietsch, m. (2020). angebotsnutzungs-modelle der wirkweise des unterrichts: ein kritischer vergleich verschiedener modellvarianten [opportunity-use model of effective teaching: a critical comparison of different models]. zeitschrift für pädagogik, 66(1), 63–80. https://doi.org/10.3262/zpb2001063 vieluf, s. (2022). wie, wann und warum nutzen schüler*innen lerngelegenheiten im unterricht? eine übergreifende diskussion der beiträge zum thementeil. [how, when, and why do students use learning opportunities in the classroom? a comprehensive discussion of the contributions to the thematic section]. unterrichtswissenschaft, 50(2), 265-286. https://doi.org/10.1007/s42010022-00144-z vieluf, s., & klieme, e. (2023). teaching effectiveness revisited through the lens of practice theories. in a-k. praetorius & c. y. charalambous (eds.), theorizing teaching: current status and open issues (pp. 57–95). cham: springer international publishing. vu, t., magis-weinberg, l., jansen, b. r., van atteveldt, n., janssen, t. w., lee, n. c., ... & meeter, m. (2022). motivation-achievement cycles in learning: a literature review and research agenda. educational psychology review, 34(1), 39–71. https://doi.org/10.1007/s10648-021-09616-7 vygotsky, l. s. (1978). mind in society: the development of higher psychological processes. harvard university press. wagner, w., göllner, r., werth, s., voss, t., schmitz, b., & trautwein, u. (2016). student and teacher ratings of instructional quality: consistency of ratings over time, agreement, and predictive power. journal of educational psychology, 108(5), 705–721. https://doi.org/10.1037/edu0000075 wang, m. t., & eccles, j. s. (2013). school context, achievement motivation, and academic engagement: a longitudinal study of school engagement using a multidimensional perspective. learning and instruction, 28, 12–23. https://doi.org/10.1016/j.learninstruc.2013.04.002 wang, y., tian, l., & huebner, e. s. (2019). basic psychological needs satisfaction at school, behavioral school engagement, and academic achievement: longitudinal reciprocal relations among elementary school students. contemporary educational psychology, 56, 130–139. https://doi.org/10.1016/j.cedpsych.2019.01.003 wang, m. c., haertel, g. d., & walberg, h. j. (1993). toward a knowledge base for school learning. review of educational research, 63(3), 249–294. https://doi.org/10.3102/00346543063003249 warner, g. j., fay, d., & spörer, n. (2017). relations among personal initiative and the development of reading strategy knowledge and reading comprehension. frontline learning research, 5(2), 1–23. https://doi.org/10.14786/flr.v5i2.272 wentzel, k. r., & miele, d. b. (2016). handbook of motivation at school (2nd ed.). routledge. wigfield, a., & eccles, j. s. (1992). the development of achievement task values: a theoretical analysis. developmental review, 12(3), 265–310. https://doi.org/10.1016/0273-2297(92)90011-p wigfield, a., & eccles, j. (2002). development of achievement motivation. academic press. williams, l. j., hartman, n., & cavazotte, f. (2010). method variance and marker variables: a review and comprehensive cfa marker technique. organizational research methods, 13(3), 477–514. https://psycnet.apa.org/doi/10.3200/joer.97.6.311-318 https://doi.org/10.1080/09243453.2018.1539015 https://doi.org/10.1108/s0749-7423(2010)000016a007 https://doi.org/10.3262/zpb2001063 https://doi.org/10.1007/s42010-022-00144-z https://doi.org/10.1007/s42010-022-00144-z https://doi.org/10.1007/s10648-021-09616-7 https://doi.org/10.1037/edu0000075 https://doi.org/10.1016/j.learninstruc.2013.04.002 https://doi.org/10.1016/j.cedpsych.2019.01.003 https://doi.org/10.3102/00346543063003249 https://doi.org/10.14786/flr.v5i2.272 https://doi.org/10.1016/0273-2297(92)90011-p alp christ, capon-sieber et al. | f l r 104 https://doi.org/10.1177/1094428110366036 zhang, t., solmon, m. a., kosma, m., carson, r. l., & gu, x. (2011). need support, need satisfaction, intrinsic motivation, and physical activity participation among middle school students. journal of teaching in physical education, 30(1), 51–68. https://doi.org/10.1123/jtpe.30.1.51 zee, m., koomen, h. m. y., & van der veen, i. (2013). student-teacher relationship quality and academic adjustment in upper elementary school: the role of student personality. journal of school psychology, 51(4), 517–533. https://doi.org/10.1016/j.jsp.2013.05.003 zhen, r., liu, r.‑d., ding, y., wang, j., liu, y., & le xu (2017). the mediating roles of academic self-efficacy and academic emotions in the relation between basic psychological needs satisfaction and learning engagement among chinese adolescent students. learning and individual differences, 54, 210–216. https://doi.org/10.1016/j.lindif.2017.01.017 zhu, y., & kaiser, g. (2022). impacts of classroom teaching practices on students’ mathematics learning interest, mathematics self-efficacy and mathematics test achievements: a secondary analysis of shanghai data from the international video study global teaching insights. zdm– mathematics education, 54(3), 581–593. https://doi.org/10.1007/s11858-022-01343-9 zhou, j., huebner, e. s., & tian, l. (2021). the reciprocal relations among basic psychological need satisfaction at school, positivity and academic achievement in chinese early adolescents. learning and instruction, 71, 101370. https://doi.org/10.1016/j.learninstruc.2020.101370 ziegelbauer, s. (2009). denkprozesse lernwirksam anregen. sensortechnik im modernen physikunterricht. [stimulating thought processes in an effective way. sensor technology in modern physics lessons.] tectum https://doi.org/10.1177/1094428110366036 https://doi.org/10.1123/jtpe.30.1.51 https://doi.org/10.1016/j.jsp.2013.05.003 https://doi.org/10.1016/j.lindif.2017.01.017 https://doi.org/10.1007/s11858-022-01343-9 https://doi.org/10.1016/j.learninstruc.2020.101370 alp christ, capon-sieber et al. | f l r 105 supplementary material revisiting the three basic dimensions model: a critical empirical investigation of the indirect effects of student-perceived teaching quality on student outcomes 1. scales and items used in the study 1.1. items, reliabilities (ω), and descriptive statistics for the subscales items id-pre ωwithin ωbetween m sd cognitive activation with discourse (t1) sqa_cogdisc .65 .74 2.62 .49 cognitive activation sqa_cogact .52 .85 2.60 .53 our mathematics teacher presents tasks for which there is no obvious solution. sqa18e 2.26 .88 our mathematics teacher presents tasks that require us to apply what we have learned to new contexts. sqa18f 3.09 .72 our mathematics teacher gives tasks that require us to think critically. sqa18g 2.53 .79 our mathematics teacher asks us to decide on our own procedures for solving complex tasks. sqa18h 2.53 .86 discourse sqa_discourse .64 .87 2.65 .71 our mathematics teacher gives us opportunities to explain our ideas. sqa18i 3.13 .86 our mathematics teacher encourages us to question and critique arguments made by other students. sqa18j 2.73 .93 our mathematics teacher requires us to engage in discussions among ourselves. sqa18k 2.08 .96 classroom management (t1) sqa_classman_mon_ removed .60 .95 3.01 .47 disruptions sqa_disruptions .76 .99 2.79 .72 when the lesson begins, our mathematics teacher has to wait quite a long time for us to quieten down. sqa20a_rec 2.83 .79 we lose quite a lot of time because of students interrupting the lesson. sqa20b_rec 2.78 .83 there is much disruptive noise in this classroom. sqa20c_rec 2.77 .85 rule clarity sqa_rulecl 2.92 .72 in our teacher’s class, we are aware of what is allowed and what is not allowed. sqa20d 3.37 .73 in our teacher’s class, we know why certain rules are important sqa20e 3.14 .74 teachers managing disruptions sqa_tmd 3.19 .70 our teacher manages to stop disruptions quickly. sqa20f 3.25 .75 our teacher reacts to disruptions in such a way that the students stop disturbing learning. sqa20g 3.14 .79 transitions in our teacher’s class, transitions from one phase of the lesson to the other (e.g., from discussions to individual work) take a lot of time. sqa20h_rec 2.77 .76 monitoring (removed from the analyses) sqa_monitoring 2.78 .71 alp christ, capon-sieber et al. | f l r 106 our teacher is immediately aware of students doing something else. sqa20i 2.77 .81 our teacher is aware of what is happening in the classroom, even if he or she is busy with an individual student. sqa20j 2.80 .80 student support (t1) sqa_support_all .86 .98 2.98 .56 teacher support for learning sqa_tesup .78 .97 3.03 .71 our mathematics teacher gives extra help when we need it. sqa21a 3.21 .77 our mathematics teacher continues teaching until we understand. sqa21b 2.96 .85 our mathematics teacher helps us with our learning. sqa21c 2.93 .82 competence support sqa_supcom .82 .97 2.92 .72 our mathematics teacher makes me feel confident in my ability to do well in the . sqa21d 2.85 .93 our mathematics teacher listens to my view on how to do things. sqa21e 2.99 .80 i feel that our mathematics teacher understands me. sqa21f 2.93 .91 our mathematics teacher makes me feel confident in my ability to learn the material. sqa21g 2.93 .64 autonomy support sqa_supaut .63 .82 2.98 .55 our mathematics teacher provides me with different alternatives (e.g. learning materials or tasks). sqa21h 2.67 .90 our mathematics teacher encourages me to find the best way to proceed by myself. sqa21i 2.83 .82 our mathematics teacher lets me work on my own. sqa21j 3.32 .67 our mathematics teacher appreciates it when different solutions come up for discussion. sqa21k 3.12 .76 depth of processing (t1) sqa_usecogact .67 1.00 2.66 .63 i keep thinking about tasks until i really understand them. sqa17d 2.96 .81 i think intensively about the mathematical content. sqa17e 2.54 .80 i develop my own ideas regarding the topic taught. sqa17f 2.47 .85 time-on-task (t1) sqa_usetot .73 .95 3.10 .54 i pay attention in mathematics class. sqa17j 3.19 .65 i listen to the instruction given in class. sqa17k 3.33 .60 i let my mind wander during the lessons. sqa17l_rec 2.76 .75 need satisfaction (t1) sqa_useselfdet .63 .90 2.86 .64 i feel i can decide on things on my own. (autonomy) sqa17g 2.50 .88 i feel understood by my mathematics teacher. (relatedness) sqa17h 2.98 .90 i feel confident in my ability to learn this material. (competence) sqa17i 3.11 .74 interest (t1) sqa_interest .85 .98 2.40 .78 i am interested in mathematics. sqa14a 2.73 .90 i often think that what we are talking about in my mathematics class is interesting. sqa14b 2.43 .88 after mathematics class i am often already curious about the next mathematics class. sqa14c 2.05 .86 alp christ, capon-sieber et al. | f l r 107 interest (t2) sqb_pint .81 .99 2.17 .73 i was interested in the topic of quadratic equations. sqb03a 2.49 .88 i often thought that what we were talking about in my mathematics class during the unit on quadratic equations was interesting. sqb03b 2.19 .84 after my mathematics class on the topic of quadratic equations i was often already curious about the next mathematics class. sqb03c 1.83 .80 note. for classroom management, subdimensions are determined according to the gti technical report (praetorius et al., pp. 2020b 1.2. the links to the talis video study https://www.oecd.org/education/school/global-teaching-insights-technical-documents.htm 1.2.1. student questionnaires https://www.oecd.org/education/school/gti-techreport-annexd1.pdf https://www.oecd.org/education/school/gti-techreport-annexd2.pdf 1.2.2. student mathematics tests https://www.oecd.org/education/school/gti-techreport-annexe2.pdf https://www.oecd.org/education/school/gti-techreport-annexe3.pdf 2. r code for the analyses 2.1. r code for the preliminary analysis 2.1.1. descriptive analysis descriptives for the variables mean(mydata$sqa_cogdisc, na.rm=t) sd(mydata$sqa_cogdisc, na.rm=t) mean(mydata$sqa_classman_mon_removed, na.rm=t) sd(mydata$sqa_classman_mon_removed, na.rm=t) mean(mydata$sqa_support_all, na.rm=t) sd(mydata$sqa_support_all, na.rm=t) mean(mydata$sqa_usecogact, na.rm=t) sd(mydata$sqa_usecogact, na.rm=t) mean(mydata$sqa_usetot, na.rm=t) sd(mydata$sqa_usetot, na.rm=t) https://www.oecd.org/education/school/global-teaching-insights-technical-documents.htm https://www.oecd.org/education/school/gti-techreport-annexd1.pdf https://www.oecd.org/education/school/gti-techreport-annexd2.pdf https://www.oecd.org/education/school/gti-techreport-annexe2.pdf https://www.oecd.org/education/school/gti-techreport-annexe3.pdf alp christ, capon-sieber et al. | f l r 108 mean(mydata$sqa_useselfdet, na.rm=t) sd(mydata$sqa_useselfdet, na.rm=t) mean(mydata$sqa_interest, na.rm=t) sd(mydata$sqa_interest, na.rm=t) mean(mydata$sqa_ach, na.rm=t) sd(mydata$sqa_ach, na.rm=t) mean(mydata$sqb_pint, na.rm=t) sd(mydata$sqb_pint, na.rm=t) mean(mydata$sqb_ach, na.rm=t) sd(mydata$sqb_ach, na.rm=t) 2.1.2 bivariate correlations mydata_3